AlexandrusPS: A User-Friendly Pipeline for the Automated Detection of Orthologous Gene Clusters and Subsequent Positive Selection Analysis.


Journal

Genome biology and evolution
ISSN: 1759-6653
Titre abrégé: Genome Biol Evol
Pays: England
ID NLM: 101509707

Informations de publication

Date de publication:
06 10 2023
Historique:
accepted: 06 10 2023
medline: 30 10 2023
pubmed: 13 10 2023
entrez: 13 10 2023
Statut: ppublish

Résumé

The detection of adaptive selection in a system approach considering all protein-coding genes allows for the identification of mechanisms and pathways that enabled adaptation to different environments. Currently, available programs for the estimation of positive selection signals can be divided into two groups. They are either easy to apply but can analyze only one gene family at a time, restricting system analysis; or they can handle larger cohorts of gene families, but require considerable prerequisite data such as orthology associations, codon alignments, phylogenetic trees, and proper configuration files. All these steps require extensive computational expertise, restricting this endeavor to specialists. Here, we introduce AlexandrusPS, a high-throughput pipeline that overcomes technical challenges when conducting transcriptome-wide positive selection analyses on large sets of nucleotide and protein sequences. The pipeline streamlines 1) the execution of an accurate orthology prediction as a precondition for positive selection analysis, 2) preparing and organizing configuration files for CodeML, 3) performing positive selection analysis using CodeML, and 4) generating an output that is easy to interpret, including all maximum likelihood and log-likelihood test results. The only input needed from the user is the CDS and peptide FASTA files of proteins of interest. The pipeline is provided in a Docker image, requiring no program or module installation, enabling the application of the pipeline in any computing environment. AlexandrusPS and its documentation are available via GitHub (https://github.com/alejocn5/AlexandrusPS).

Identifiants

pubmed: 37831426
pii: 7311098
doi: 10.1093/gbe/evad187
pmc: PMC10612477
pii:
doi:

Substances chimiques

Codon 0
Proteins 0

Types de publication

Journal Article Research Support, Non-U.S. Gov't

Langues

eng

Sous-ensembles de citation

IM

Informations de copyright

© The Author(s) 2023. Published by Oxford University Press on behalf of Society for Molecular Biology and Evolution.

Références

Mol Biol Evol. 2012 Oct;29(10):2889-93
pubmed: 22490825
G3 (Bethesda). 2020 Dec 3;10(12):4369-4372
pubmed: 33093185
Mol Biol Evol. 2013 Dec;30(12):2723-4
pubmed: 24105918
Mol Biol Evol. 2013 Jul;30(7):1675-86
pubmed: 23558341
PLoS Genet. 2008 Apr 11;4(4):e1000046
pubmed: 18404212
Mol Biol Evol. 2007 Aug;24(8):1586-91
pubmed: 17483113
Proc Natl Acad Sci U S A. 2009 Sep 8;106(36):E95; author reply E96
pubmed: 19805241
Nucleic Acids Res. 2006 Jul 1;34(Web Server issue):W609-12
pubmed: 16845082
BMC Bioinformatics. 2008 Dec 08;9:524
pubmed: 19061522
Mol Biol Evol. 2007 May;24(5):1219-28
pubmed: 17339634
PLoS One. 2013;8(3):e60440
pubmed: 23555972
Science. 2003 Dec 12;302(5652):1960-3
pubmed: 14671302
BMC Genomics. 2015 Aug 01;16:567
pubmed: 26231214
BMC Genomics. 2019 Dec 12;20(1):977
pubmed: 31842731
BMC Bioinformatics. 2011 Apr 28;12:124
pubmed: 21526987
Nucleic Acids Res. 2007 Jul;35(Web Server issue):W506-11
pubmed: 17586822
Methods Mol Biol. 2014;1079:155-70
pubmed: 24170401
Nucleic Acids Res. 2011 Jul;39(Web Server issue):W470-4
pubmed: 21646336
Mol Biol Evol. 2018 Jul 1;35(7):1668-1677
pubmed: 29659991
PLoS One. 2014 Oct 20;9(10):e96243
pubmed: 25329307
Evol Bioinform Online. 2013 Nov 24;9:487-90
pubmed: 24324324
Nucleic Acids Res. 2011 Jul;39(Web Server issue):W479-85
pubmed: 21531699
BMC Bioinformatics. 2010 May 27;11:284
pubmed: 20507581
Virus Res. 2008 Nov;137(2):253-6
pubmed: 18761043
Mol Biol Evol. 2021 Jan 23;38(2):589-605
pubmed: 32986833
BMC Genomics. 2013 Dec 27;14:924
pubmed: 24373418
BMC Bioinformatics. 2016 Sep 06;17(1):354
pubmed: 27597435
Genes (Basel). 2022 Jun 18;13(6):
pubmed: 35741852
Mol Biol Evol. 2019 Oct 1;36(10):2157-2164
pubmed: 31241141
PLoS One. 2012;7(1):e29903
pubmed: 22253821
PLoS Biol. 2004 Feb;2(2):E29
pubmed: 14966531
Genome Res. 2023 Jan;33(1):112-128
pubmed: 36653121
Gigascience. 2014 Dec 12;3(1):27
pubmed: 25671092
Ecol Evol. 2021 Sep 03;11(19):13029-13035
pubmed: 34646450
Nature. 2013 Oct 10;502(7470):228-31
pubmed: 24005325
Annu Rev Anim Biosci. 2015;3:57-111
pubmed: 25689317
Nature. 2007 Nov 8;450(7167):219-32
pubmed: 17994088

Auteurs

Alejandro Ceron-Noriega (A)

Institute of Molecular Biology (IMB), Quantitative Proteomics, Mainz, Germany.
Institute of Human Genetics, University Medical Center of the Johannes Gutenberg University Mainz, Department of Human Genetics, Mainz, Germany.

Vivien A C Schoonenberg (VAC)

Institute of Molecular Biology (IMB), Quantitative Proteomics, Mainz, Germany.
Present address: Division of Hematology/Oncology, Boston Children's Hospital, Harvard Medical School, Boston, Massachusetts, USA.
Present address: Molecular Pathology Unit and Center for Cancer Research, Massachusetts General Hospital, Department of Pathology, Harvard Medical School, Boston, Massachusetts, USA.

Falk Butter (F)

Institute of Molecular Biology (IMB), Quantitative Proteomics, Mainz, Germany.
Institute of Molecular Virology and Cell Biology, Friedrich-Loeffler-Institute, Greifswald, Germany.

Michal Levin (M)

Institute of Molecular Biology (IMB), Quantitative Proteomics, Mainz, Germany.

Articles similaires

Genome, Chloroplast Phylogeny Genetic Markers Base Composition High-Throughput Nucleotide Sequencing

Selecting optimal software code descriptors-The case of Java.

Yegor Bugayenko, Zamira Kholmatova, Artem Kruglov et al.
1.00
Software Algorithms Programming Languages
Databases, Protein Protein Domains Protein Folding Proteins Deep Learning
Animals Hemiptera Insect Proteins Phylogeny Insecticides

Classifications MeSH