Centrifuger: lossless compression of microbial genomes for efficient and accurate metagenomic sequence classification.
Compact data structure
FM-index
Metagenomic
r-index
Journal
Genome biology
ISSN: 1474-760X
Titre abrégé: Genome Biol
Pays: England
ID NLM: 100960660
Informations de publication
Date de publication:
25 Apr 2024
25 Apr 2024
Historique:
received:
14
11
2023
accepted:
10
04
2024
medline:
26
4
2024
pubmed:
26
4
2024
entrez:
25
4
2024
Statut:
epublish
Résumé
Centrifuger is an efficient taxonomic classification method that compares sequencing reads against a microbial genome database. In Centrifuger, the Burrows-Wheeler transformed genome sequences are losslessly compressed using a novel scheme called run-block compression. Run-block compression achieves sublinear space complexity and is effective at compressing diverse microbial databases like RefSeq while supporting fast rank queries. Combining this compression method with other strategies for compacting the Ferragina-Manzini (FM) index, Centrifuger reduces the memory footprint by half compared to other FM-index-based approaches. Furthermore, the lossless compression and the unconstrained match length help Centrifuger achieve greater accuracy than competing methods at lower taxonomic levels.
Identifiants
pubmed: 38664753
doi: 10.1186/s13059-024-03244-4
pii: 10.1186/s13059-024-03244-4
doi:
Types de publication
Journal Article
Research Support, Non-U.S. Gov't
Research Support, U.S. Gov't, Non-P.H.S.
Langues
eng
Sous-ensembles de citation
IM
Pagination
106Subventions
Organisme : NIGMS NIH HHS
ID : P20GM130454
Pays : United States
Organisme : NIGMS NIH HHS
ID : R35GM139602
Pays : United States
Organisme : NIGMS NIH HHS
ID : 3P20GM130454-05WS
Pays : United States
Organisme : NHGRI NIH HHS
ID : R01HG011392
Pays : United States
Informations de copyright
© 2024. The Author(s).
Références
Tringe SG, Rubin EM. Metagenomics: DNA sequencing of environmental samples. Nat Rev Genet. 2005;6:805–14.
doi: 10.1038/nrg1709
pubmed: 16304596
Zhang L, Chen F, Zeng Z, Xu M, Sun F, Yang L, et al. Advances in Metagenomics and Its Application in Environmental Microorganisms. Frontiers in Microbiology. 2021;12. Available from: https://www.frontiersin.org/articles/10.3389/fmicb.2021.766364 . Cited 2023 Oct 13
Chiu CY, Miller SA. Clinical metagenomics. Nat Rev Genet. 2019;20:341–55.
doi: 10.1038/s41576-019-0113-7
pubmed: 30918369
pmcid: 6858796
De Filippis F, Paparo L, Nocerino R, Della Gatta G, Carucci L, Russo R, et al. Specific gut microbiome signatures and the associated pro-inflamatory functions are linked to pediatric allergy and acquisition of immune tolerance. Nat Commun. 2021;12:5958.
doi: 10.1038/s41467-021-26266-z
pubmed: 34645820
pmcid: 8514477
Thomas AM, Manghi P, Asnicar F, Pasolli E, Armanini F, Zolfo M, et al. Metagenomic analysis of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation. Nat Med. 2019;25:667–78.
doi: 10.1038/s41591-019-0405-7
pubmed: 30936548
pmcid: 9533319
Knight R, Vrbanac A, Taylor BC, Aksenov A, Callewaert C, Debelius J, et al. Best practices for analysing microbiomes. Nat Rev Microbiol. 2018;16:410–22.
doi: 10.1038/s41579-018-0029-9
pubmed: 29795328
Pruitt KD, Tatusova T, Maglott DR. NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins. Nucleic Acids Res. 2007;35:D61–5.
doi: 10.1093/nar/gkl842
pubmed: 17130148
Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, et al. GenBank. Nucleic Acids Res. 2013;41:D36–42.
doi: 10.1093/nar/gks1195
pubmed: 23193287
Parks DH, Chuvochina M, Rinke C, Mussig AJ, Chaumeil P-A, Hugenholtz P. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res. 2022;50:D785–94.
doi: 10.1093/nar/gkab776
pubmed: 34520557
Wood DE, Lu J, Langmead B. Improved metagenomic analysis with Kraken 2. Genome Biol. 2019;20:257.
doi: 10.1186/s13059-019-1891-0
pubmed: 31779668
pmcid: 6883579
Roberts M, Hayes W, Hunt BR, Mount SM, Yorke JA. Reducing storage requirements for biological sequence comparison. Bioinformatics. 2004;20:3363–9.
doi: 10.1093/bioinformatics/bth408
pubmed: 15256412
Wood DE, Salzberg SL. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome Biol. 2014;15:R46.
doi: 10.1186/gb-2014-15-3-r46
pubmed: 24580807
pmcid: 4053813
Blanco-Míguez A, Beghini F, Cumbo F, McIver LJ, Thompson KN, Zolfo M, et al. Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nat Biotechnol. 2023;41:1633–44.
doi: 10.1038/s41587-023-01688-w
pubmed: 36823356
pmcid: 10635831
Segata N, Waldron L, Ballarini A, Narasimhan V, Jousson O, Huttenhower C. Metagenomic microbial community profiling using unique clade-specific marker genes. Nat Methods. 2012;9:811–4.
doi: 10.1038/nmeth.2066
pubmed: 22688413
pmcid: 3443552
Ounit R, Wanamaker S, Close TJ, Lonardi S. CLARK: fast and accurate classification of metagenomic and genomic sequences using discriminative k-mers. BMC Genomics. 2015;16:236.
doi: 10.1186/s12864-015-1419-2
pubmed: 25879410
pmcid: 4428112
Piro VC, Dadi TH, Seiler E, Reinert K, Renard BY. ganon: precise metagenomics classification against large and up-to-date sets of reference sequences. Bioinformatics. 2020;36:i12–20.
doi: 10.1093/bioinformatics/btaa458
pubmed: 32657362
pmcid: 7355301
Shen W, Xiang H, Huang T, Tang H, Peng M, Cai D, et al. KMCP: accurate metagenomic profiling of both prokaryotic and viral populations by pseudo-mapping. Bioinformatics. 2023;39:btac845.
doi: 10.1093/bioinformatics/btac845
pubmed: 36579886
Kim D, Song L, Breitwieser FP, Salzberg SL. Centrifuge: rapid and sensitive classification of metagenomic sequences. Genome Res. 2016;26:1721–9.
doi: 10.1101/gr.210641.116
pubmed: 27852649
pmcid: 5131823
Burrows M, Wheeler DJ. A block-sorting lossless data compression algorithm. SRS Research Report. 1994;124. Available from: https://cir.nii.ac.jp/crid/1571417124717214720 . Cited 2023 Oct 13.
Ferragina P, Manzini G. Opportunistic data structures with applications. Proceedings 41st Annual Symposium on Foundations of Computer Science. 2000. p. 390–8. Available from: https://ieeexplore.ieee.org/abstract/document/892127 . Cited 2023 Oct 13.
Nasko DJ, Koren S, Phillippy AM, Treangen TJ. RefSeq database growth influences the accuracy of k-mer-based lowest common ancestor species identification. Genome Biol. 2018;19:165.
doi: 10.1186/s13059-018-1554-6
pubmed: 30373669
pmcid: 6206640
Kreft S, Navarro G. On compressing and indexing repetitive sequences. Theoret Comput Sci. 2013;483:115–33.
doi: 10.1016/j.tcs.2012.02.006
Gagie T, Gawrychowski P, Kärkkäinen J, Nekrich Y, Puglisi SJ. A Faster Grammar-Based Self-Index. arXiv; 2012. Available from: http://arxiv.org/abs/1109.3954 . Cited 2023 Oct 13.
Mäkinen V, Navarro G, Sirén J, Välimäki N. Storage and Retrieval of Highly Repetitive Sequence Collections. J Comput Biol. 2010;17:281–308.
doi: 10.1089/cmb.2009.0169
pubmed: 20377446
Nishimoto T, Tabei Y. Optimal-Time Queries on BWT-runs Compressed Indexes. arXiv; 2021. Available from: http://arxiv.org/abs/2006.05104 . Cited 2023 Nov 5.
Gagie T, Navarro G, Prezza N. Optimal-Time Text Indexing in BWT-runs Bounded Space. arXiv; 2017. Available from: http://arxiv.org/abs/1705.10382 . Cited 2023 Sep 20.
Grossi R, Gupta A, Vitter JS. High-order entropy-compressed text indexes. Proceedings of the fourteenth annual ACM-SIAM symposium on Discrete algorithms. USA: Society for Industrial and Applied Mathematics; 2003. p. 841–50.
Prezza N. r-index: the run-length BWT index. 2023. Available from: https://github.com/nicolaprezza/r-index . Cited 2023 Oct 14.
Gog S, Beller T, Moffat A, Petri M. From Theory to Practice: Plug and Play with Succinct Data Structures. arXiv; 2013. Available from: http://arxiv.org/abs/1311.1249 . Cited 2023 Nov 12.
Holtgrewe M. Mason – A Read Simulator for Second Generation Sequencing Data. Technical Report FU Berlin. 2010.Available from: https://publications.imp.fu-berlin.de/962/ . Cited 2023 Oct 6.
Huang W, Li L, Myers JR, Marth GT. ART: a next-generation sequencing read simulator. Bioinformatics. 2012;28:593–4.
doi: 10.1093/bioinformatics/btr708
pubmed: 22199392
Muggli MD, Bowe A, Noyes NR, Morley PS, Belk KE, Raymond R, et al. Succinct colored de Bruijn graphs. Bioinformatics. 2017;33:3181–7.
doi: 10.1093/bioinformatics/btx067
pubmed: 28200001
pmcid: 5872255
Alanko JN, Vuohtoniemi J, Mäklin T, Puglisi SJ. Themisto: a scalable colored k-mer index for sensitive pseudoalignment against hundreds of thousands of bacterial genomes. Bioinformatics. 2023;39:i260–9.
doi: 10.1093/bioinformatics/btad233
pubmed: 37387143
pmcid: 10311346
Alanko JN, Puglisi SJ, Vuohtoniemi J. Succinct k-mer Sets Using Subset Rank Queries on the Spectral Burrows-Wheeler Transform *. bioRxiv; 2022. p. 2022.05.19.492613. Available from: https://www.biorxiv.org/content/10.1101/2022.05.19.492613v2 . Cited 2024 Feb 5.
Meyer F, Fritz A, Deng Z-L, Koslicki D, Lesker TR, Gurevich A, et al. Critical Assessment of Metagenome Interpretation: the second round of challenges. Nat Methods. 2022;19:429–40.
doi: 10.1038/s41592-022-01431-4
pubmed: 35396482
pmcid: 9007738
Dilthey AT, Jain C, Koren S, Phillippy AM. Strain-level metagenomic assignment and compositional estimation for long reads with MetaMaps. Nat Commun. 2019;10:3066.
doi: 10.1038/s41467-019-10934-2
pubmed: 31296857
pmcid: 6624308
Ulrich J-U, Renard BY. Taxor: Fast and space-efficient taxonomic classification of long reads with hierarchical interleaved XOR filters. bioRxiv; 2023. p. 2023.07.20.549822. Available from: https://www.biorxiv.org/10.1101/2023.07.20.549822v1 . Cited 2024 Jan 29.
Ahmed O, Rossi M, Boucher C, Langmead B. Efficient taxa identification using a pangenome index. Genome Res. 2023;33(7):1069–77. https://doi.org/10.1101/gr.277642.123 .
doi: 10.1101/gr.277642.123
pubmed: 37258301
pmcid: 10538492
Gagie T, Kashgouli S, Langmead B. KATKA: A KRAKEN-Like Tool with k Given at Query Time. In: Arroyuelo D, Poblete B, editors. String Processing and Information Retrieval. Cham: Springer International Publishing; 2022. p. 191–7.
doi: 10.1007/978-3-031-20643-6_14
Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7:11257.
doi: 10.1038/ncomms11257
pubmed: 27071849
pmcid: 4833860
Li H. Fast construction of FM-index for long sequence reads. Bioinformatics. 2014;30:3274–5.
doi: 10.1093/bioinformatics/btu541
pubmed: 25107872
pmcid: 4221129
Li H. Exploring single-sample SNP and INDEL calling with whole-genome de novo assembly. Bioinformatics. 2012;28:1838–44.
doi: 10.1093/bioinformatics/bts280
pubmed: 22569178
pmcid: 3389770
Schaeffer L, Pimentel H, Bray N, Melsted P, Pachter L. Pseudoalignment for metagenomic read assignment. Bioinformatics. 2017;33:2082–8.
doi: 10.1093/bioinformatics/btx106
pubmed: 28334086
pmcid: 5870846
Shaw J, Yu YW. Metagenome profiling and containment estimation through abundance-corrected k-mer sketching with sylph. bioRxiv; 2024. p. 2023.11.20.567879. Available from: https://www.biorxiv.org/content/10.1101/2023.11.20.567879v2 . Cited 2024 Jan 28.
Lu J, Breitwieser FP, Thielen P, Salzberg SL. Bracken: estimating species abundance in metagenomics data. PeerJ Comput Sci. 2017;3: e104.
doi: 10.7717/peerj-cs.104
Dempster AP, Laird NM, Rubin DB. Maximum Likelihood from Incomplete Data Via the EM Algorithm. J Roy Stat Soc: Ser B (Methodol). 1977;39:1–22.
doi: 10.1111/j.2517-6161.1977.tb01600.x
Skoufos G, Almodaresi F, Zakeri M, Paulson JN, Patro R, Hatzigeorgiou AG, et al. AGAMEMNON: an Accurate metaGenomics And MEtatranscriptoMics quaNtificatiON analysis suite. Genome Biol. 2022;23:39.
doi: 10.1186/s13059-022-02610-4
pubmed: 35101114
pmcid: 8802518
Liu J, Ma Y, Ren Y, Guo H. Centrifuge+: improving metagenomic analysis upon Centrifuge. bioRxiv; 2023. p. 2023.02.27.530134. Available from: https://www.biorxiv.org/content/10.1101/2023.02.27.530134v1 . Cited 2024 Jan 29.
Morgulis A, Gertz EM, Schäffer AA, Agarwala R. A fast and symmetric DUST implementation to mask low-complexity DNA sequences. J Comput Biol. 2006;13:1028–40.
doi: 10.1089/cmb.2006.13.1028
pubmed: 16796549
Piro VC, Reinert K. ganon2: up-to-date and scalable metagenomics analysis. bioRxiv; 2023. p. 2023.12.07.570547. Available from: https://www.biorxiv.org/content10.1101/2023.12.07.570547v1 . Cited 2024 Jan 29.
Barbay J, Navarro G. On compressing permutations and adaptive sorting. Theoret Comput Sci. 2013;513:109–23.
doi: 10.1016/j.tcs.2013.10.019
Kärkkäinen J. Fast BWT in small space by blockwise suffix sorting. Theoret Comput Sci. 2007;387:249–57.
doi: 10.1016/j.tcs.2007.07.018
Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009;10:R25.
doi: 10.1186/gb-2009-10-3-r25
pubmed: 19261174
pmcid: 2690996
Gog S, Kärkkäinen J, Kempa D, Petri M, Puglisi SJ. Fixed Block Compression Boosting in FM-Indexes: Theory and Practice. Algorithmica. 2019;81:1370–91.
doi: 10.1007/s00453-018-0475-9
Song L, Langmead B. Centrifuger. Github; 2024. https://github.com/mourisl/centrifuger . Accessed 8 Feb 2024.
Song L, Langmead B. Centrifuger v1.0.1. Zenodo; 2024. https://doi.org/10.5281/zenodo.10938378 . Accessed 7 Apr 2024.
Song L, Langmead B. Centrifuger evaluations. Github; 2024. https://github.com/mourisl/centrifuger_evaluations . Accessed 9 Mar 2024.