The GIAB genomic stratifications resource for human reference genomes.


Journal

Nature communications
ISSN: 2041-1723
Titre abrégé: Nat Commun
Pays: England
ID NLM: 101528555

Informations de publication

Date de publication:
19 Oct 2024
Historique:
received: 07 11 2023
accepted: 07 10 2024
medline: 19 10 2024
pubmed: 19 10 2024
entrez: 18 10 2024
Statut: epublish

Résumé

Despite the growing variety of sequencing and variant-calling tools, no workflow performs equally well across the entire human genome. Understanding context-dependent performance is critical for enabling researchers, clinicians, and developers to make informed tradeoffs when selecting sequencing hardware and software. Here we describe a set of "stratifications," which are BED files that define distinct contexts throughout the genome. We define these for GRCh37/38 as well as the new T2T-CHM13 reference, adding many new hard-to-sequence regions which are critical for understanding performance as the field progresses. Specifically, we highlight the increase in hard-to-map and GC-rich stratifications in CHM13 relative to the previous references. We then compare the benchmarking performance with each reference and show the performance penalty brought about by these additional difficult regions in CHM13. Additionally, we demonstrate how the stratifications can track context-specific improvements over different platform iterations, using Oxford Nanopore Technologies as an example. The means to generate these stratifications are available as a snakemake pipeline at https://github.com/usnistgov/giab-stratifications . We anticipate this being useful in enabling precise risk-reward calculations when building sequencing pipelines for any of the commonly-used reference genomes.

Identifiants

pubmed: 39424793
doi: 10.1038/s41467-024-53260-y
pii: 10.1038/s41467-024-53260-y
doi:

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Pagination

9029

Informations de copyright

© 2024. The Author(s).

Références

Olson, N. D. et al. PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult-to-map regions. Cell Genom. 2, 100129 (2022).
doi: 10.1016/j.xgen.2022.100129 pubmed: 35720974 pmcid: 9205427
Krusche, P. et al. Best practices for benchmarking germline small-variant calls in human genomes. Nat. Biotechnol. 37, 555–560 (2019).
doi: 10.1038/s41587-019-0054-x pubmed: 30858580 pmcid: 6699627
Zook, J. M. et al. Integrating human sequence data sets provides a resource of benchmark SNP and indel genotype calls. Nat. Biotechnol. 32, 246–251 (2014).
doi: 10.1038/nbt.2835 pubmed: 24531798
Wagner, J. et al. Benchmarking challenging small variants with linked and long reads. Cell Genom. 2, 100128 (2022).
doi: 10.1016/j.xgen.2022.100128 pubmed: 36452119 pmcid: 9706577
Xiao, C., Zook, J., Trask, S. & Sherry, S. Abstract 5328: GIAB: Genome reference material development resources for clinical sequencing. Cancer Res. 74, 5328–5328 (2014).
doi: 10.1158/1538-7445.AM2014-5328
Wagner, J. et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat. Biotechnol. 40, 672–680 (2022).
doi: 10.1038/s41587-021-01158-1 pubmed: 35132260 pmcid: 9117392
Majidian, S., Agustinho, D. P., Chin, C.-S., Sedlazeck, F. J. & Mahmoud, M. Genomic variant benchmark: if you cannot measure it, you cannot improve it. Genome Biol. 24, 221 (2023).
doi: 10.1186/s13059-023-03061-1 pubmed: 37798733 pmcid: 10552390
Olson, N. D. et al. Variant calling and benchmarking in an era of complete human genome sequences. Nat. Rev. Genet. 24, 464–483 (2023).
doi: 10.1038/s41576-023-00590-0 pubmed: 37059810
English, A. C., Menon, V. K., Gibbs, R. A., Metcalf, G. A. & Sedlazeck, F. J. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol. 23, 271 (2022).
doi: 10.1186/s13059-022-02840-6 pubmed: 36575487 pmcid: 9793516
O’Leary, N. A. et al. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res. 44, D733–D745 (2016).
doi: 10.1093/nar/gkv1189 pubmed: 26553804
Roy, S. et al. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: A joint recommendation of the Association for Molecular Pathology and the College of American Pathologists. J. Mol. Diagn. 20, 4–27 (2018).
doi: 10.1016/j.jmoldx.2017.11.003 pubmed: 29154853
Nurk, S. et al. The complete sequence of a human genome. Science 376, 44–53 (2022).
doi: 10.1126/science.abj6987 pubmed: 35357919 pmcid: 9186530
Rhie, A. et al. The complete sequence of a human Y chromosome. Nature 621, 344–354 (2023).
doi: 10.1038/s41586-023-06457-y pubmed: 37612512 pmcid: 10752217
Antonarakis, S. E. Short arms of human acrocentric chromosomes and the completion of the human genome sequence. Genome Res. 32, 599–607 (2022).
doi: 10.1101/gr.275350.121 pubmed: 35361624 pmcid: 8997349
Foox, J. et al. Performance assessment of DNA sequencing platforms in the ABRF Next-Generation Sequencing Study. Nat. Biotechnol. 39, 1129–1140 (2021).
doi: 10.1038/s41587-021-01049-5 pubmed: 34504351 pmcid: 8985210
Pyke, R. M. et al. Computational KIR copy number discovery reveals interaction between inhibitory receptor burden and survival. Pac. Symp. Biocomput. 24, 148–159 (2019).
pubmed: 30864318 pmcid: 6417817
Aganezov, S. et al. A complete reference genome improves analysis of human genetic variation. Science 376, eabl3533 (2022).
doi: 10.1126/science.abl3533 pubmed: 35357935 pmcid: 9336181
Behera, S. et al. FixItFelix: improving genomic analysis by fixing reference errors. Genome Biol. 24, 31 (2023).
doi: 10.1186/s13059-023-02863-7 pubmed: 36810122 pmcid: 9942314
Cleary, J. G. et al. Comparing variant call files for performance benchmarking of next-generation sequencing variant calling pipelines. bioRxiv 023754. https://doi.org/10.1101/023754 (2015).
Dunn, T. & Narayanasamy, S. vcfdist: accurately benchmarking phased small variant calls in human genomes. Nat. Commun. 14, 8149 (2023).
doi: 10.1038/s41467-023-43876-x pubmed: 38071244 pmcid: 10710436
Zheng, Z. et al. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nat. Comput Sci. 2, 797–803 (2022).
doi: 10.1038/s43588-022-00387-x pubmed: 38177392
English, A. C. et al. Analysis and benchmarking of small and large genomic variants across tandem repeats. Nat. Biotechnol. https://doi.org/10.1038/s41587-024-02225-z (2024).
Jarvis, E. D. et al. Semi-automated assembly of high-quality diploid human reference genomes. Nature 611, 519–531 (2022).
doi: 10.1038/s41586-022-05325-5 pubmed: 36261518 pmcid: 9668749
Smolka, M., Rescheneder, P., Schatz, M. C., von Haeseler, A. & Sedlazeck, F. J. Teaser: Individualized benchmarking and optimization of read mapping results for NGS data. Genome Biol. 16, 235 (2015).
doi: 10.1186/s13059-015-0803-1 pubmed: 26494581 pmcid: 4618857
Chen, N.-C., Solomon, B., Mun, T., Iyer, S. & Langmead, B. Reference flow: reducing reference bias using multiple population genomes. Genome Biol. 22, 8 (2021).
doi: 10.1186/s13059-020-02229-3 pubmed: 33397413 pmcid: 7780692
Vollger, M. R. et al. Segmental duplications and their variation in a complete human genome. Science 376, eabj6965 (2022).
doi: 10.1126/science.abj6965 pubmed: 35357917 pmcid: 8979283
Majidian, S., Kahaei, M. H. & de Ridder, D. Hap10: reconstructing accurate and long polyploid haplotypes using linked reads. BMC Bioinforma. 21, 253 (2020).
doi: 10.1186/s12859-020-03584-5
Chin, C.-S. et al. Phased diploid genome assembly with single-molecule real-time sequencing. Nat. Methods 13, 1050–1054 (2016).
doi: 10.1038/nmeth.4035 pubmed: 27749838 pmcid: 5503144
Sedlazeck, F. J., Lee, H., Darby, C. A. & Schatz, M. C. Piercing the dark matter: bioinformatics of long-range sequencing and mapping. Nat. Rev. Genet. 19, 329–346 (2018).
doi: 10.1038/s41576-018-0003-4 pubmed: 29599501
Dwarshuis, N. et al. StratoMod: Predicting sequencing and variant calling errors with interpretable machine learning. Comm. Bio. 7, 1613 (2024).
Wagner, J. et al. Small variant benchmark from a complete assembly of X and Y chromosomes. Nat. Commun. in press. bioRxiv 2023.10.31.564997. https://doi.org/10.1101/2023.10.31.564997 (2023).
Pedersen, B. S. et al. Effective variant filtering and expected candidate variant yield in studies of rare human disease. NPJ Genom. Med 6, 60 (2021).
doi: 10.1038/s41525-021-00227-3 pubmed: 34267211 pmcid: 8282602
Majidian, S. & Sedlazeck, F. J. PhaseME: Automatic rapid assessment of phasing quality and phasing improvement. Gigascience 9, giaa078 (2020).
Gurevich, A., Saveliev, V., Vyahhi, N. & Tesler, G. QUAST: quality assessment tool for genome assemblies. Bioinformatics 29, 1072–1075 (2013).
doi: 10.1093/bioinformatics/btt086 pubmed: 23422339 pmcid: 3624806
Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol. 21, 245 (2020).
doi: 10.1186/s13059-020-02134-9 pubmed: 32928274 pmcid: 7488777
Benjamini, Y. & Speed, T. P. Summarizing and correcting the GC content bias in high-throughput sequencing. Nucleic Acids Res. 40, e72 (2012).
doi: 10.1093/nar/gks001 pubmed: 22323520 pmcid: 3378858
Cheung, M.-S., Down, T. A., Latorre, I. & Ahringer, J. Systematic bias in high-throughput sequencing data and its correction by BEADS. Nucleic Acids Res. 39, e103 (2011).
doi: 10.1093/nar/gkr425 pubmed: 21646344 pmcid: 3159482
Yip, K. Y., Cheng, C. & Gerstein, M. Machine learning and genome annotation: a match meant to be? Genome Biol. 14, 205 (2013).
doi: 10.1186/gb-2013-14-5-205 pubmed: 23731483 pmcid: 4053789
Fotsing, S. F. et al. The impact of short tandem repeat variation on gene expression. Nat. Genet. 51, 1652–1659 (2019).
doi: 10.1038/s41588-019-0521-9 pubmed: 31676866 pmcid: 6917484
Turner, S. et al. Quality control procedures for genome-wide association studies. Curr. Protoc. Hum. Genet. Chapter 1, Unit1.19 (2011).
Rautiainen, M. et al. Telomere-to-telomere assembly of diploid chromosomes with Verkko. Nat. Biotechnol. 41, 1474–1482 (2023).
doi: 10.1038/s41587-023-01662-6 pubmed: 36797493 pmcid: 10427740
Shumate, A. et al. Assembly and annotation of an Ashkenazi human reference genome. Genome Biol. 21, 129 (2020).
doi: 10.1186/s13059-020-02047-7 pubmed: 32487205 pmcid: 7265644
Derrien, T. et al. Fast computation and applications of genome mappability. PLoS One 7, e30377 (2012).
doi: 10.1371/journal.pone.0030377 pubmed: 22276185 pmcid: 3261895
Li, H. et al. A synthetic-diploid benchmark for accurate variant-calling evaluation. Nat. Methods 15, 595–597 (2018).
doi: 10.1038/s41592-018-0054-7 pubmed: 30013044 pmcid: 6341484
Baid, G. et al. An Extensive sequence dataset of gold-standard samples for benchmarking and development. bioRxiv 2020.12.11.422022. https://doi.org/10.1101/2020.12.11.422022 (2020).

Auteurs

Nathan Dwarshuis (N)

Material Measurement Laboratory, National Institute of Standards and Technology, Gaithersburg, MD., USA.

Divya Kalra (D)

Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, USA.

Jennifer McDaniel (J)

Material Measurement Laboratory, National Institute of Standards and Technology, Gaithersburg, MD., USA.

Philippe Sanio (P)

University of Applied Sciences Upper Austria - FH Hagenberg, Hagenberg im Mühlkreis, Austria.

Pilar Alvarez Jerez (P)

Center for Alzheimer's and Related Dementias (CARD), National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, 20892, USA.
Department of Neurodegenerative Disease, UCL Queen Square Institute of Neurology, University College London, London, UK.

Bharati Jadhav (B)

Department of Genetics and Genomic Sciences and Mindich Child Health and Development Institute, Icahn School of Medicine at Mount, Hess Center for Science and Medicine, New York, NY, USA.

Wenyu Eddy Huang (WE)

Department of Computer Science, College of Engineering, Rice University, Houston, TX, USA.

Rajarshi Mondal (R)

Department of Bioinformatics, Pondicherry University, Pondicherry, India.

Ben Busby (B)

DNA Nexus, Mountain View, CA, USA.

Nathan D Olson (ND)

Material Measurement Laboratory, National Institute of Standards and Technology, Gaithersburg, MD., USA.

Fritz J Sedlazeck (FJ)

Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, USA.
Department of Computer Science, College of Engineering, Rice University, Houston, TX, USA.

Justin Wagner (J)

Material Measurement Laboratory, National Institute of Standards and Technology, Gaithersburg, MD., USA.

Sina Majidian (S)

Department of Computational Biology, University of Lausanne, Lausanne, Switzerland. sina.majidian@unil.ch.
SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland. sina.majidian@unil.ch.

Justin M Zook (JM)

Material Measurement Laboratory, National Institute of Standards and Technology, Gaithersburg, MD., USA. justin.zook@nist.gov.

Articles similaires

Genome, Chloroplast Phylogeny Genetic Markers Base Composition High-Throughput Nucleotide Sequencing

[Redispensing of expensive oral anticancer medicines: a practical application].

Lisanne N van Merendonk, Kübra Akgöl, Bastiaan Nuijen
1.00
Humans Antineoplastic Agents Administration, Oral Drug Costs Counterfeit Drugs

Smoking Cessation and Incident Cardiovascular Disease.

Jun Hwan Cho, Seung Yong Shin, Hoseob Kim et al.
1.00
Humans Male Smoking Cessation Cardiovascular Diseases Female
Humans United States Aged Cross-Sectional Studies Medicare Part C

Classifications MeSH