The GIAB genomic stratifications resource for human reference genomes.
Journal
Nature communications
ISSN: 2041-1723
Titre abrégé: Nat Commun
Pays: England
ID NLM: 101528555
Informations de publication
Date de publication:
19 Oct 2024
19 Oct 2024
Historique:
received:
07
11
2023
accepted:
07
10
2024
medline:
19
10
2024
pubmed:
19
10
2024
entrez:
18
10
2024
Statut:
epublish
Résumé
Despite the growing variety of sequencing and variant-calling tools, no workflow performs equally well across the entire human genome. Understanding context-dependent performance is critical for enabling researchers, clinicians, and developers to make informed tradeoffs when selecting sequencing hardware and software. Here we describe a set of "stratifications," which are BED files that define distinct contexts throughout the genome. We define these for GRCh37/38 as well as the new T2T-CHM13 reference, adding many new hard-to-sequence regions which are critical for understanding performance as the field progresses. Specifically, we highlight the increase in hard-to-map and GC-rich stratifications in CHM13 relative to the previous references. We then compare the benchmarking performance with each reference and show the performance penalty brought about by these additional difficult regions in CHM13. Additionally, we demonstrate how the stratifications can track context-specific improvements over different platform iterations, using Oxford Nanopore Technologies as an example. The means to generate these stratifications are available as a snakemake pipeline at https://github.com/usnistgov/giab-stratifications . We anticipate this being useful in enabling precise risk-reward calculations when building sequencing pipelines for any of the commonly-used reference genomes.
Identifiants
pubmed: 39424793
doi: 10.1038/s41467-024-53260-y
pii: 10.1038/s41467-024-53260-y
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
9029Informations de copyright
© 2024. The Author(s).
Références
Olson, N. D. et al. PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult-to-map regions. Cell Genom. 2, 100129 (2022).
doi: 10.1016/j.xgen.2022.100129
pubmed: 35720974
pmcid: 9205427
Krusche, P. et al. Best practices for benchmarking germline small-variant calls in human genomes. Nat. Biotechnol. 37, 555–560 (2019).
doi: 10.1038/s41587-019-0054-x
pubmed: 30858580
pmcid: 6699627
Zook, J. M. et al. Integrating human sequence data sets provides a resource of benchmark SNP and indel genotype calls. Nat. Biotechnol. 32, 246–251 (2014).
doi: 10.1038/nbt.2835
pubmed: 24531798
Wagner, J. et al. Benchmarking challenging small variants with linked and long reads. Cell Genom. 2, 100128 (2022).
doi: 10.1016/j.xgen.2022.100128
pubmed: 36452119
pmcid: 9706577
Xiao, C., Zook, J., Trask, S. & Sherry, S. Abstract 5328: GIAB: Genome reference material development resources for clinical sequencing. Cancer Res. 74, 5328–5328 (2014).
doi: 10.1158/1538-7445.AM2014-5328
Wagner, J. et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat. Biotechnol. 40, 672–680 (2022).
doi: 10.1038/s41587-021-01158-1
pubmed: 35132260
pmcid: 9117392
Majidian, S., Agustinho, D. P., Chin, C.-S., Sedlazeck, F. J. & Mahmoud, M. Genomic variant benchmark: if you cannot measure it, you cannot improve it. Genome Biol. 24, 221 (2023).
doi: 10.1186/s13059-023-03061-1
pubmed: 37798733
pmcid: 10552390
Olson, N. D. et al. Variant calling and benchmarking in an era of complete human genome sequences. Nat. Rev. Genet. 24, 464–483 (2023).
doi: 10.1038/s41576-023-00590-0
pubmed: 37059810
English, A. C., Menon, V. K., Gibbs, R. A., Metcalf, G. A. & Sedlazeck, F. J. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol. 23, 271 (2022).
doi: 10.1186/s13059-022-02840-6
pubmed: 36575487
pmcid: 9793516
O’Leary, N. A. et al. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res. 44, D733–D745 (2016).
doi: 10.1093/nar/gkv1189
pubmed: 26553804
Roy, S. et al. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: A joint recommendation of the Association for Molecular Pathology and the College of American Pathologists. J. Mol. Diagn. 20, 4–27 (2018).
doi: 10.1016/j.jmoldx.2017.11.003
pubmed: 29154853
Nurk, S. et al. The complete sequence of a human genome. Science 376, 44–53 (2022).
doi: 10.1126/science.abj6987
pubmed: 35357919
pmcid: 9186530
Rhie, A. et al. The complete sequence of a human Y chromosome. Nature 621, 344–354 (2023).
doi: 10.1038/s41586-023-06457-y
pubmed: 37612512
pmcid: 10752217
Antonarakis, S. E. Short arms of human acrocentric chromosomes and the completion of the human genome sequence. Genome Res. 32, 599–607 (2022).
doi: 10.1101/gr.275350.121
pubmed: 35361624
pmcid: 8997349
Foox, J. et al. Performance assessment of DNA sequencing platforms in the ABRF Next-Generation Sequencing Study. Nat. Biotechnol. 39, 1129–1140 (2021).
doi: 10.1038/s41587-021-01049-5
pubmed: 34504351
pmcid: 8985210
Pyke, R. M. et al. Computational KIR copy number discovery reveals interaction between inhibitory receptor burden and survival. Pac. Symp. Biocomput. 24, 148–159 (2019).
pubmed: 30864318
pmcid: 6417817
Aganezov, S. et al. A complete reference genome improves analysis of human genetic variation. Science 376, eabl3533 (2022).
doi: 10.1126/science.abl3533
pubmed: 35357935
pmcid: 9336181
Behera, S. et al. FixItFelix: improving genomic analysis by fixing reference errors. Genome Biol. 24, 31 (2023).
doi: 10.1186/s13059-023-02863-7
pubmed: 36810122
pmcid: 9942314
Cleary, J. G. et al. Comparing variant call files for performance benchmarking of next-generation sequencing variant calling pipelines. bioRxiv 023754. https://doi.org/10.1101/023754 (2015).
Dunn, T. & Narayanasamy, S. vcfdist: accurately benchmarking phased small variant calls in human genomes. Nat. Commun. 14, 8149 (2023).
doi: 10.1038/s41467-023-43876-x
pubmed: 38071244
pmcid: 10710436
Zheng, Z. et al. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nat. Comput Sci. 2, 797–803 (2022).
doi: 10.1038/s43588-022-00387-x
pubmed: 38177392
English, A. C. et al. Analysis and benchmarking of small and large genomic variants across tandem repeats. Nat. Biotechnol. https://doi.org/10.1038/s41587-024-02225-z (2024).
Jarvis, E. D. et al. Semi-automated assembly of high-quality diploid human reference genomes. Nature 611, 519–531 (2022).
doi: 10.1038/s41586-022-05325-5
pubmed: 36261518
pmcid: 9668749
Smolka, M., Rescheneder, P., Schatz, M. C., von Haeseler, A. & Sedlazeck, F. J. Teaser: Individualized benchmarking and optimization of read mapping results for NGS data. Genome Biol. 16, 235 (2015).
doi: 10.1186/s13059-015-0803-1
pubmed: 26494581
pmcid: 4618857
Chen, N.-C., Solomon, B., Mun, T., Iyer, S. & Langmead, B. Reference flow: reducing reference bias using multiple population genomes. Genome Biol. 22, 8 (2021).
doi: 10.1186/s13059-020-02229-3
pubmed: 33397413
pmcid: 7780692
Vollger, M. R. et al. Segmental duplications and their variation in a complete human genome. Science 376, eabj6965 (2022).
doi: 10.1126/science.abj6965
pubmed: 35357917
pmcid: 8979283
Majidian, S., Kahaei, M. H. & de Ridder, D. Hap10: reconstructing accurate and long polyploid haplotypes using linked reads. BMC Bioinforma. 21, 253 (2020).
doi: 10.1186/s12859-020-03584-5
Chin, C.-S. et al. Phased diploid genome assembly with single-molecule real-time sequencing. Nat. Methods 13, 1050–1054 (2016).
doi: 10.1038/nmeth.4035
pubmed: 27749838
pmcid: 5503144
Sedlazeck, F. J., Lee, H., Darby, C. A. & Schatz, M. C. Piercing the dark matter: bioinformatics of long-range sequencing and mapping. Nat. Rev. Genet. 19, 329–346 (2018).
doi: 10.1038/s41576-018-0003-4
pubmed: 29599501
Dwarshuis, N. et al. StratoMod: Predicting sequencing and variant calling errors with interpretable machine learning. Comm. Bio. 7, 1613 (2024).
Wagner, J. et al. Small variant benchmark from a complete assembly of X and Y chromosomes. Nat. Commun. in press. bioRxiv 2023.10.31.564997. https://doi.org/10.1101/2023.10.31.564997 (2023).
Pedersen, B. S. et al. Effective variant filtering and expected candidate variant yield in studies of rare human disease. NPJ Genom. Med 6, 60 (2021).
doi: 10.1038/s41525-021-00227-3
pubmed: 34267211
pmcid: 8282602
Majidian, S. & Sedlazeck, F. J. PhaseME: Automatic rapid assessment of phasing quality and phasing improvement. Gigascience 9, giaa078 (2020).
Gurevich, A., Saveliev, V., Vyahhi, N. & Tesler, G. QUAST: quality assessment tool for genome assemblies. Bioinformatics 29, 1072–1075 (2013).
doi: 10.1093/bioinformatics/btt086
pubmed: 23422339
pmcid: 3624806
Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol. 21, 245 (2020).
doi: 10.1186/s13059-020-02134-9
pubmed: 32928274
pmcid: 7488777
Benjamini, Y. & Speed, T. P. Summarizing and correcting the GC content bias in high-throughput sequencing. Nucleic Acids Res. 40, e72 (2012).
doi: 10.1093/nar/gks001
pubmed: 22323520
pmcid: 3378858
Cheung, M.-S., Down, T. A., Latorre, I. & Ahringer, J. Systematic bias in high-throughput sequencing data and its correction by BEADS. Nucleic Acids Res. 39, e103 (2011).
doi: 10.1093/nar/gkr425
pubmed: 21646344
pmcid: 3159482
Yip, K. Y., Cheng, C. & Gerstein, M. Machine learning and genome annotation: a match meant to be? Genome Biol. 14, 205 (2013).
doi: 10.1186/gb-2013-14-5-205
pubmed: 23731483
pmcid: 4053789
Fotsing, S. F. et al. The impact of short tandem repeat variation on gene expression. Nat. Genet. 51, 1652–1659 (2019).
doi: 10.1038/s41588-019-0521-9
pubmed: 31676866
pmcid: 6917484
Turner, S. et al. Quality control procedures for genome-wide association studies. Curr. Protoc. Hum. Genet. Chapter 1, Unit1.19 (2011).
Rautiainen, M. et al. Telomere-to-telomere assembly of diploid chromosomes with Verkko. Nat. Biotechnol. 41, 1474–1482 (2023).
doi: 10.1038/s41587-023-01662-6
pubmed: 36797493
pmcid: 10427740
Shumate, A. et al. Assembly and annotation of an Ashkenazi human reference genome. Genome Biol. 21, 129 (2020).
doi: 10.1186/s13059-020-02047-7
pubmed: 32487205
pmcid: 7265644
Derrien, T. et al. Fast computation and applications of genome mappability. PLoS One 7, e30377 (2012).
doi: 10.1371/journal.pone.0030377
pubmed: 22276185
pmcid: 3261895
Li, H. et al. A synthetic-diploid benchmark for accurate variant-calling evaluation. Nat. Methods 15, 595–597 (2018).
doi: 10.1038/s41592-018-0054-7
pubmed: 30013044
pmcid: 6341484
Baid, G. et al. An Extensive sequence dataset of gold-standard samples for benchmarking and development. bioRxiv 2020.12.11.422022. https://doi.org/10.1101/2020.12.11.422022 (2020).