Principal Component Analysis Applications in COVID-19 Genome Sequence Studies.
COVID-19
Genome Sequences
Mutation
Principle Component Analysis
Journal
Cognitive computation
ISSN: 1866-9956
Titre abrégé: Cognit Comput
Pays: United States
ID NLM: 101499358
Informations de publication
Date de publication:
13 Jan 2021
13 Jan 2021
Historique:
received:
07
07
2020
accepted:
06
11
2020
entrez:
18
1
2021
pubmed:
19
1
2021
medline:
19
1
2021
Statut:
aheadofprint
Résumé
RNA genomes from coronavirus have a length as long as 32 kilobases, and the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) that caused the outbreak of coronavirus disease 2019 (COVID-19) pandemic has long sequences which made the analysis difficult. Over 20,000 sequences have been submitted to GISAID, and the number is growing fast each day which increased the difficulties in data analysis; however, genome sequence analysis is critical in understanding the COVID-19 and preventing the spread of the disease. In this study, a principal component analysis (PCA) was applied to the aligned large size genome sequences and the numerical numbers were converted from the letters using a published method designed for protein sequence cluster analysis. The study initialized with a shortlist sequence testing, and the PCA score plot showed high tolerance with low-quality data, and the major virus sequences from humans were separated from the pangolin and bat samples. Our study also successfully built a model for a large number of sequences with more than 20,000 sequences which indicate the potential mutation directions for the COVID-19 which can be served as a pretreatment method for detailed studies such as decision tree-based methods. In summary, our study provided a fast tool to analyze the high-volume genome sequences such as the COVID-19 and successfully applied to more than 20,000 sequences which may provide mutation direction information for COVID-19 studies.
Identifiants
pubmed: 33456620
doi: 10.1007/s12559-020-09790-w
pii: 9790
pmc: PMC7804214
doi:
Types de publication
Journal Article
Langues
eng
Pagination
1-12Informations de copyright
© Springer Science+Business Media, LLC, part of Springer Nature 2021.
Déclaration de conflit d'intérêts
Conflict of InterestThe authors declare that they have no conflict of interest.
Références
Int J Pharm. 2001 Jan 5;212(1):41-53
pubmed: 11165819
Genomics. 2020 Nov;112(6):5204-5213
pubmed: 32966857
Sci Rep. 2019 Dec 17;9(1):19297
pubmed: 31848355
Proc Natl Acad Sci U S A. 2020 Apr 28;117(17):9241-9243
pubmed: 32269081
Euro Surveill. 2017 Mar 30;22(13):
pubmed: 28382917
Mol Biol Evol. 2020 Aug 1;37(8):2155-2172
pubmed: 32359163
Bioinformatics. 2019 Jun 1;35(11):1963-1965
pubmed: 30358807
J Integr Bioinform. 2015 Sep 03;12(1):257
pubmed: 26527191
Nat Struct Biol. 1995 Feb;2(2):171-8
pubmed: 7749921
Environ Pollut. 2009 Aug-Sep;157(8-9):2275-81
pubmed: 19410344
Cell. 2020 Sep 17;182(6):1401-1418.e18
pubmed: 32810439
Glob Chall. 2017 Jan 10;1(1):33-46
pubmed: 31565258
Bioinformatics. 2019 Mar 1;35(5):743-752
pubmed: 30102339
J Struct Funct Genomics. 2014 Mar;15(1):1-11
pubmed: 24496727
Curr Biol. 2020 Apr 20;30(8):1578
pubmed: 32315626
BMC Bioinformatics. 2007 Apr 23;8:135
pubmed: 17451607
Nucleic Acids Res. 2002 Jul 15;30(14):3059-66
pubmed: 12136088
Osong Public Health Res Perspect. 2020 Feb;11(1):3-7
pubmed: 32149036
Int J Antimicrob Agents. 2020 Mar;55(3):105924
pubmed: 32081636
BMC Bioinformatics. 2014 Feb 21;15:51
pubmed: 24555693
J Med Virol. 2020 Jun;92(6):667-674
pubmed: 32167180
J Med Virol. 2020 Sep;92(9):1634-1636
pubmed: 32190908
Chemometr Intell Lab Syst. 2011 Dec 15;109(2):162-170
pubmed: 26246647
Int J Antimicrob Agents. 2020 May;55(5):105951
pubmed: 32234466
IEEE Trans Neural Netw. 1995;6(1):131-43
pubmed: 18263293
Nucleic Acids Res. 2018 Nov 16;46(20):e120
pubmed: 30169659