An efficient numerical representation of genome sequence: natural vector with covariance component.
Archaea
Bacteria
Convex hull classification
Fungi
Giant virus
Natural vector with covariance component
Phylogeny
The nearest neighbor classification
Virus
Journal
PeerJ
ISSN: 2167-8359
Titre abrégé: PeerJ
Pays: United States
ID NLM: 101603425
Informations de publication
Date de publication:
2022
2022
Historique:
received:
16
11
2021
accepted:
16
05
2022
entrez:
22
6
2022
pubmed:
23
6
2022
medline:
23
6
2022
Statut:
epublish
Résumé
The characterization and comparison of microbial sequences, including archaea, bacteria, viruses and fungi, are very important to understand their evolutionary origin and the population relationship. Most methods are limited by the sequence length and lack of generality. The purpose of this study is to propose a general characterization method, and to study the classification and phylogeny of the existing datasets. We present a new alignment-free method to represent and compare biological sequences. By adding the covariance between each two nucleotides, the new 18-dimensional natural vector successfully describes 24,250 genomic sequences and 95,542 DNA barcode sequences. The new numerical representation is used to study the classification and phylogenetic relationship of microbial sequences. First, the classification results validate that the six-dimensional covariance vector is necessary to characterize sequences. Then, the 18-dimensional natural vector is further used to conduct the similarity relationship between giant virus and archaea, bacteria, other viruses. The nearest distance calculation results reflect that the giant viruses are closer to bacteria in distribution of four nucleotides. The phylogenetic relationships of the three representative families, Mimiviridae, Pandoraviridae and Marsellieviridae from giant viruses are analyzed. The trees show that ten sequences of Mimiviridae are clustered with Pandoraviridae, and Mimiviridae is closer to the root of the tree than Marsellieviridae. The new developed alignment-free method can be computed very fast, which provides an effective numerical representation for the sequence of microorganisms.
Sections du résumé
Background
The characterization and comparison of microbial sequences, including archaea, bacteria, viruses and fungi, are very important to understand their evolutionary origin and the population relationship. Most methods are limited by the sequence length and lack of generality. The purpose of this study is to propose a general characterization method, and to study the classification and phylogeny of the existing datasets.
Methods
We present a new alignment-free method to represent and compare biological sequences. By adding the covariance between each two nucleotides, the new 18-dimensional natural vector successfully describes 24,250 genomic sequences and 95,542 DNA barcode sequences. The new numerical representation is used to study the classification and phylogenetic relationship of microbial sequences.
Results
First, the classification results validate that the six-dimensional covariance vector is necessary to characterize sequences. Then, the 18-dimensional natural vector is further used to conduct the similarity relationship between giant virus and archaea, bacteria, other viruses. The nearest distance calculation results reflect that the giant viruses are closer to bacteria in distribution of four nucleotides. The phylogenetic relationships of the three representative families, Mimiviridae, Pandoraviridae and Marsellieviridae from giant viruses are analyzed. The trees show that ten sequences of Mimiviridae are clustered with Pandoraviridae, and Mimiviridae is closer to the root of the tree than Marsellieviridae. The new developed alignment-free method can be computed very fast, which provides an effective numerical representation for the sequence of microorganisms.
Identifiants
pubmed: 35729905
doi: 10.7717/peerj.13544
pii: 13544
pmc: PMC9206847
doi:
Substances chimiques
Nucleotides
0
Types de publication
Journal Article
Research Support, Non-U.S. Gov't
Langues
eng
Pagination
e13544Informations de copyright
© 2022 Sun et al.
Déclaration de conflit d'intérêts
The authors declare that they have no competing interests.
Références
BMC Evol Biol. 2018 Dec 27;18(1):200
pubmed: 30587116
Nucleic Acids Res. 1990 Apr 25;18(8):2163-70
pubmed: 2336393
Bioinformatics. 2014 Jul 15;30(14):2000-8
pubmed: 24828656
Front Plant Sci. 2012 Aug 29;3:192
pubmed: 22952468
Brief Bioinform. 2014 May;15(3):369-75
pubmed: 24162172
Science. 2004 Nov 19;306(5700):1344-50
pubmed: 15486256
PeerJ. 2020 Aug 03;8:e9625
pubmed: 32832270
Appl Microbiol Biotechnol. 2010 Jun;87(1):99-108
pubmed: 20405123
PLoS One. 2011 Mar 02;6(3):e17293
pubmed: 21399690
Int J Mol Sci. 2020 May 29;21(11):
pubmed: 32485813
Proc Natl Acad Sci U S A. 2012 Apr 17;109(16):6241-6
pubmed: 22454494
Comput Struct Biotechnol J. 2021 Jul 27;19:4226-4234
pubmed: 34429843
Adv Virus Res. 2013;85:25-56
pubmed: 23439023
Lancet. 1997 Mar 29;349(9056):925-6
pubmed: 9093261
Bioinformatics. 2007 Nov 1;23(21):2947-8
pubmed: 17846036
Brief Bioinform. 2014 May;15(3):376-89
pubmed: 24058049
Viruses. 2019 Apr 30;11(5):
pubmed: 31052218
Virol J. 2009 Oct 27;6:178
pubmed: 19860921
Genomics. 2019 Dec;111(6):1777-1784
pubmed: 30529533
BMC Bioinformatics. 2004 Aug 19;5:113
pubmed: 15318951
Commun Integr Biol. 2012 Jan 1;5(1):102-6
pubmed: 22482024
J Theor Biol. 2014 Oct 21;359:18-28
pubmed: 24911780
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
pubmed: 9254694
Bioinformatics. 2008 Oct 15;24(20):2296-302
pubmed: 18710871
IEEE/ACM Trans Comput Biol Bioinform. 2022 May-Jun;19(3):1782-1793
pubmed: 33237867
J Mol Biol. 1990 Oct 5;215(3):403-10
pubmed: 2231712
Nucleic Acids Res. 2004 Mar 19;32(5):1792-7
pubmed: 15034147
Science. 2013 Jul 19;341(6143):281-6
pubmed: 23869018