An efficient numerical representation of genome sequence: natural vector with covariance component.

Archaea Bacteria Convex hull classification Fungi Giant virus Natural vector with covariance component Phylogeny The nearest neighbor classification Virus

Journal

PeerJ
ISSN: 2167-8359
Titre abrégé: PeerJ
Pays: United States
ID NLM: 101603425

Informations de publication

Date de publication:
2022
Historique:
received: 16 11 2021
accepted: 16 05 2022
entrez: 22 6 2022
pubmed: 23 6 2022
medline: 23 6 2022
Statut: epublish

Résumé

The characterization and comparison of microbial sequences, including archaea, bacteria, viruses and fungi, are very important to understand their evolutionary origin and the population relationship. Most methods are limited by the sequence length and lack of generality. The purpose of this study is to propose a general characterization method, and to study the classification and phylogeny of the existing datasets. We present a new alignment-free method to represent and compare biological sequences. By adding the covariance between each two nucleotides, the new 18-dimensional natural vector successfully describes 24,250 genomic sequences and 95,542 DNA barcode sequences. The new numerical representation is used to study the classification and phylogenetic relationship of microbial sequences. First, the classification results validate that the six-dimensional covariance vector is necessary to characterize sequences. Then, the 18-dimensional natural vector is further used to conduct the similarity relationship between giant virus and archaea, bacteria, other viruses. The nearest distance calculation results reflect that the giant viruses are closer to bacteria in distribution of four nucleotides. The phylogenetic relationships of the three representative families, Mimiviridae, Pandoraviridae and Marsellieviridae from giant viruses are analyzed. The trees show that ten sequences of Mimiviridae are clustered with Pandoraviridae, and Mimiviridae is closer to the root of the tree than Marsellieviridae. The new developed alignment-free method can be computed very fast, which provides an effective numerical representation for the sequence of microorganisms.

Sections du résumé

Background
The characterization and comparison of microbial sequences, including archaea, bacteria, viruses and fungi, are very important to understand their evolutionary origin and the population relationship. Most methods are limited by the sequence length and lack of generality. The purpose of this study is to propose a general characterization method, and to study the classification and phylogeny of the existing datasets.
Methods
We present a new alignment-free method to represent and compare biological sequences. By adding the covariance between each two nucleotides, the new 18-dimensional natural vector successfully describes 24,250 genomic sequences and 95,542 DNA barcode sequences. The new numerical representation is used to study the classification and phylogenetic relationship of microbial sequences.
Results
First, the classification results validate that the six-dimensional covariance vector is necessary to characterize sequences. Then, the 18-dimensional natural vector is further used to conduct the similarity relationship between giant virus and archaea, bacteria, other viruses. The nearest distance calculation results reflect that the giant viruses are closer to bacteria in distribution of four nucleotides. The phylogenetic relationships of the three representative families, Mimiviridae, Pandoraviridae and Marsellieviridae from giant viruses are analyzed. The trees show that ten sequences of Mimiviridae are clustered with Pandoraviridae, and Mimiviridae is closer to the root of the tree than Marsellieviridae. The new developed alignment-free method can be computed very fast, which provides an effective numerical representation for the sequence of microorganisms.

Identifiants

pubmed: 35729905
doi: 10.7717/peerj.13544
pii: 13544
pmc: PMC9206847
doi:

Substances chimiques

Nucleotides 0

Types de publication

Journal Article Research Support, Non-U.S. Gov't

Langues

eng

Pagination

e13544

Informations de copyright

© 2022 Sun et al.

Déclaration de conflit d'intérêts

The authors declare that they have no competing interests.

Références

BMC Evol Biol. 2018 Dec 27;18(1):200
pubmed: 30587116
Nucleic Acids Res. 1990 Apr 25;18(8):2163-70
pubmed: 2336393
Bioinformatics. 2014 Jul 15;30(14):2000-8
pubmed: 24828656
Front Plant Sci. 2012 Aug 29;3:192
pubmed: 22952468
Brief Bioinform. 2014 May;15(3):369-75
pubmed: 24162172
Science. 2004 Nov 19;306(5700):1344-50
pubmed: 15486256
PeerJ. 2020 Aug 03;8:e9625
pubmed: 32832270
Appl Microbiol Biotechnol. 2010 Jun;87(1):99-108
pubmed: 20405123
PLoS One. 2011 Mar 02;6(3):e17293
pubmed: 21399690
Int J Mol Sci. 2020 May 29;21(11):
pubmed: 32485813
Proc Natl Acad Sci U S A. 2012 Apr 17;109(16):6241-6
pubmed: 22454494
Comput Struct Biotechnol J. 2021 Jul 27;19:4226-4234
pubmed: 34429843
Adv Virus Res. 2013;85:25-56
pubmed: 23439023
Lancet. 1997 Mar 29;349(9056):925-6
pubmed: 9093261
Bioinformatics. 2007 Nov 1;23(21):2947-8
pubmed: 17846036
Brief Bioinform. 2014 May;15(3):376-89
pubmed: 24058049
Viruses. 2019 Apr 30;11(5):
pubmed: 31052218
Virol J. 2009 Oct 27;6:178
pubmed: 19860921
Genomics. 2019 Dec;111(6):1777-1784
pubmed: 30529533
BMC Bioinformatics. 2004 Aug 19;5:113
pubmed: 15318951
Commun Integr Biol. 2012 Jan 1;5(1):102-6
pubmed: 22482024
J Theor Biol. 2014 Oct 21;359:18-28
pubmed: 24911780
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
pubmed: 9254694
Bioinformatics. 2008 Oct 15;24(20):2296-302
pubmed: 18710871
IEEE/ACM Trans Comput Biol Bioinform. 2022 May-Jun;19(3):1782-1793
pubmed: 33237867
J Mol Biol. 1990 Oct 5;215(3):403-10
pubmed: 2231712
Nucleic Acids Res. 2004 Mar 19;32(5):1792-7
pubmed: 15034147
Science. 2013 Jul 19;341(6143):281-6
pubmed: 23869018

Auteurs

Nan Sun (N)

Department of Mathematical Sciences, Tsinghua University, Beijing, China.

Xin Zhao (X)

Beijing Electronic Science and Technology Institute, Beijing, China.

Stephen S-T Yau (SS)

Department of Mathematical Sciences, Tsinghua University, Beijing, China.
Yanqi Lake Beijing Institute of Mathematical Sciences and Applications, Beijing, China.

Articles similaires

Genome, Chloroplast Phylogeny Genetic Markers Base Composition High-Throughput Nucleotide Sequencing

[Redispensing of expensive oral anticancer medicines: a practical application].

Lisanne N van Merendonk, Kübra Akgöl, Bastiaan Nuijen
1.00
Humans Antineoplastic Agents Administration, Oral Drug Costs Counterfeit Drugs

Smoking Cessation and Incident Cardiovascular Disease.

Jun Hwan Cho, Seung Yong Shin, Hoseob Kim et al.
1.00
Humans Male Smoking Cessation Cardiovascular Diseases Female
Humans United States Aged Cross-Sectional Studies Medicare Part C

Classifications MeSH