Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved variation and methylation.
Journal
Nature methods
ISSN: 1548-7105
Titre abrégé: Nat Methods
Pays: United States
ID NLM: 101215604
Informations de publication
Date de publication:
10 2023
10 2023
Historique:
received:
28
01
2023
accepted:
04
08
2023
medline:
9
10
2023
pubmed:
15
9
2023
entrez:
14
9
2023
Statut:
ppublish
Résumé
Long-read sequencing technologies substantially overcome the limitations of short-reads but have not been considered as a feasible replacement for population-scale projects, being a combination of too expensive, not scalable enough or too error-prone. Here we develop an efficient and scalable wet lab and computational protocol, Napu, for Oxford Nanopore Technologies long-read sequencing that seeks to address those limitations. We applied our protocol to cell lines and brain tissue samples as part of a pilot project for the National Institutes of Health Center for Alzheimer's and Related Dementias. Using a single PromethION flow cell, we can detect single nucleotide polymorphisms with F1-score comparable to Illumina short-read sequencing. Small indel calling remains difficult within homopolymers and tandem repeats, but achieves good concordance to Illumina indel calls elsewhere. Further, we can discover structural variants with F1-score on par with state-of-the-art de novo assembly methods. Our protocol phases small and structural variants at megabase scales and produces highly accurate, haplotype-specific methylation calls.
Identifiants
pubmed: 37710018
doi: 10.1038/s41592-023-01993-x
pii: 10.1038/s41592-023-01993-x
doi:
Types de publication
Journal Article
Research Support, N.I.H., Intramural
Research Support, N.I.H., Extramural
Research Support, Non-U.S. Gov't
Langues
eng
Sous-ensembles de citation
IM
Pagination
1483-1492Subventions
Organisme : NHGRI NIH HHS
ID : U24 HG010262
Pays : United States
Organisme : NHGRI NIH HHS
ID : U24 HG011853
Pays : United States
Organisme : NHGRI NIH HHS
ID : U01 HG010961
Pays : United States
Organisme : NHGRI NIH HHS
ID : R01 HG010485
Pays : United States
Organisme : NIH HHS
ID : OT2 OD026682
Pays : United States
Organisme : NHGRI NIH HHS
ID : T32 HG012344
Pays : United States
Organisme : Intramural NIH HHS
ID : ZIA NS003154
Pays : United States
Organisme : Intramural NIH HHS
ID : ZIA AG000538
Pays : United States
Organisme : NINDS NIH HHS
ID : U24 NS072026
Pays : United States
Organisme : NIA NIH HHS
ID : P30 AG019610
Pays : United States
Organisme : NIA NIH HHS
ID : P30 AG072980
Pays : United States
Organisme : NHLBI NIH HHS
ID : OT3 HL142481
Pays : United States
Organisme : NIH HHS
ID : OT2 OD033761
Pays : United States
Commentaires et corrections
Type : UpdateOf
Informations de copyright
© 2023. The Author(s), under exclusive licence to Springer Nature America, Inc.
Références
DePristo, M. A. et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat. Genet. 43, 491–498 (2011).
pubmed: 21478889
pmcid: 3083463
doi: 10.1038/ng.806
1000 Genomes Project Consortium et al. An integrated map of genetic variation from 1,092 human genomes. Nature 491, 56–65 (2012).
doi: 10.1038/nature11632
100,000 Genomes Project Pilot Investigators et al. 100,000 Genomes pilot on rare-disease diagnosis in health care—preliminary report. N. Engl. J. Med. 385, 1868–1880 (2021).
Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020).
pubmed: 32461654
pmcid: 7334197
doi: 10.1038/s41586-020-2308-7
Huang, K.-L. et al. Pathogenic germline variants in 10,389 adult cancers. Cell 173, 355–370.e14 (2018).
pubmed: 29625052
pmcid: 5949147
doi: 10.1016/j.cell.2018.03.039
ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium. Pan-cancer analysis of whole genomes. Nature 578, 82–93 (2020).
doi: 10.1038/s41586-020-1969-6
Sedlazeck, F. J., Lee, H., Darby, C. A. & Schatz, M. C. Piercing the dark matter: bioinformatics of long-range sequencing and mapping. Nat. Rev. Genet. 19, 329–346 (2018).
pubmed: 29599501
doi: 10.1038/s41576-018-0003-4
Chen, X. et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics 32, 1220–1222 (2016).
pubmed: 26647377
doi: 10.1093/bioinformatics/btv710
Mahmoud, M. et al. Structural variant calling: the long and the short of it. Genome Biol. 20, 246 (2019).
pubmed: 31747936
pmcid: 6868818
doi: 10.1186/s13059-019-1828-7
Zarate, S. et al. Parliament2: accurate structural variant calling at scale. Gigascience 9, giaa145 (2020).
pubmed: 33347570
pmcid: 7751401
doi: 10.1093/gigascience/giaa145
Zook, J. M. et al. A robust benchmark for detection of germline large deletions and insertions. Nat. Biotechnol. 38, 1347–1355 (2020).
pubmed: 32541955
pmcid: 8454654
doi: 10.1038/s41587-020-0538-8
Wagner, J. et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat. Biotechnol. 40, 672–680 (2022).
pubmed: 35132260
pmcid: 9117392
doi: 10.1038/s41587-021-01158-1
Wagner, J. et al. Benchmarking challenging small variants with linked and long reads. Cell Genom. 2, 100128 (2022).
pubmed: 36452119
pmcid: 9706577
doi: 10.1016/j.xgen.2022.100128
Lee, H. & Schatz, M. C. Genomic dark matter: the reliability of short read mapping illustrated by the genome mappability score. Bioinformatics 28, 2097–2105 (2012).
pubmed: 22668792
pmcid: 3413383
doi: 10.1093/bioinformatics/bts330
Martin, M. et al. WhatsHap: fast and accurate read-based phasing. Preprint at bioRxiv https://doi.org/10.1101/085050 (2016).
Loh, P.-R. et al. Reference-based phasing using the Haplotype Reference Consortium panel. Nat. Genet. 48, 1443–1448 (2016).
pubmed: 27694958
pmcid: 5096458
doi: 10.1038/ng.3679
Sedlazeck, F. J. et al. Accurate detection of complex structural variations using single-molecule sequencing. Nat. Methods 15, 461–468 (2018).
pubmed: 29713083
pmcid: 5990442
doi: 10.1038/s41592-018-0001-7
Jiang, T. et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 21, 189 (2020).
pubmed: 32746918
pmcid: 7477834
doi: 10.1186/s13059-020-02107-y
Shafin, K. et al. Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads. Nat. Methods 18, 1322–1332 (2021).
pubmed: 34725481
pmcid: 8571015
doi: 10.1038/s41592-021-01299-w
Lin, J.-H., Chen, L.-C., Yu, S.-C. & Huang, Y.-T. LongPhase: an ultra-fast chromosome-scale phasing algorithm for small and large variants. Bioinformatics 38, 1816–1822 (2022).
pubmed: 35104333
doi: 10.1093/bioinformatics/btac058
Mahmoud, M., Doddapaneni, H., Timp, W. & Sedlazeck, F. J. PRINCESS: comprehensive detection of haplotype resolved SNVs, SVs, and methylation. Genome Biol. 22, 268 (2021).
pubmed: 34521442
pmcid: 8442460
doi: 10.1186/s13059-021-02486-w
Logsdon, G. A., Vollger, M. R. & Eichler, E. E. Long-read human genome sequencing and its applications. Nat. Rev. Genet. 21, 597–614 (2020).
pubmed: 32504078
pmcid: 7877196
doi: 10.1038/s41576-020-0236-x
Rhie, A. et al. Towards complete and error-free genome assemblies of all vertebrate species. Nature 592, 737–746 (2021).
pubmed: 33911273
pmcid: 8081667
doi: 10.1038/s41586-021-03451-0
Nurk, S. et al. The complete sequence of a human genome. Science 376, 44–53 (2022).
pubmed: 35357919
pmcid: 9186530
doi: 10.1126/science.abj6987
Liao, W.-W. et al. A draft human pangenome reference. Nature 617, 312–324 (2023).
pubmed: 37165242
pmcid: 10172123
doi: 10.1038/s41586-023-05896-x
Jarvis, E. D. et al. Automated assembly of high-quality diploid human reference genomes. Nature 611, 519–531 (2022).
pubmed: 36261518
pmcid: 9668749
doi: 10.1038/s41586-022-05325-5
Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive -mer weighting and repeat separation. Genome Res. 27, 722–736 (2017).
pubmed: 28298431
pmcid: 5411767
doi: 10.1101/gr.215087.116
Jain, M. et al. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nat. Biotechnol. 36, 338–345 (2018).
pubmed: 29431738
pmcid: 5889714
doi: 10.1038/nbt.4060
Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P. A. Assembly of long, error-prone reads using repeat graphs. Nat. Biotechnol. 37, 540–546 (2019).
pubmed: 30936562
doi: 10.1038/s41587-019-0072-8
Shafin, K. et al. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes. Nat. Biotechnol. 38, 1044–1053 (2020).
pubmed: 32686750
pmcid: 7483855
doi: 10.1038/s41587-020-0503-6
Rautiainen, M. et al. Verkko: telomere-to-telomere assembly of diploid chromosomes. Nat. Biotechnol. https://doi.org/10.1038/s41587-023-01662-6 (2023).
Billingsley, K. J. et al. Processing human frontal cortex brain tissue for population-scale Oxford Nanopore long-read DNA sequencing SOP v2. protocols.io https://doi.org/10.17504/protocols.io.kxygxzmmov8j/v2 (2022).
Baker, B. et al. Processing human frontal cortex brain tissue for population-scale SQK-LSK114 Oxford Nanopore long-read DNA sequencing SOP v1. protocols.io https://doi.org/10.17504/protocols.io.kxygx3zzog8j/v1 (2022).
Alvarez Jerez, P. et al. Processing frozen cells for population-scale Oxford Nanopore long-read DNA sequencing SOP v1. protocols.io https://doi.org/10.17504/protocols.io.5jyl8pnk7g2w/v1 (2022).
Gibbs, J. R. et al. Abundant quantitative trait loci exist for DNA methylation and gene expression in human brain. PLoS Genet. 6, e1000952 (2010).
pubmed: 20485568
pmcid: 2869317
doi: 10.1371/journal.pgen.1000952
Schatz, M. C. et al. Inverting the model of genomics data sharing with the NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space. Cell Genom. 2, 100085 (2022).
pubmed: 35199087
pmcid: 8863334
doi: 10.1016/j.xgen.2021.100085
Li, H. yak: yet another k-mer analyzer. GitHub https://github.com/lh3/yak (2023).
Smolka, M. et al. Comprehensive structural variant detection: from mosaic to population-level. Preprint at bioRxiv https://doi.org/10.1101/2022.04.04.487055 (2022).
English, A. C., Menon, V. K., Gibbs, R. A., Metcalf, G. A. & Sedlazeck, F. J. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol. 23, 271 (2022).
pubmed: 36575487
pmcid: 9793516
doi: 10.1186/s13059-022-02840-6
Yang, J. & Chaisson, M. J. P. TT-Mars: structural variants assessment based on haplotype-resolved assemblies. Genome Biol. 23, 110 (2022).
pubmed: 35524317
pmcid: 9077962
doi: 10.1186/s13059-022-02666-2
Vollger, M. R. et al. Long-read sequence and assembly of segmental duplications. Nat. Methods 16, 88–94 (2019).
pubmed: 30559433
doi: 10.1038/s41592-018-0236-3
Kirsche, M. et al. Jasmine: population-scale structural variant comparison and analysis. Nat. Methods 20, 408–417 (2023).
pubmed: 36658279
pmcid: 10006329
doi: 10.1038/s41592-022-01753-3
Chowdhury, M., Pedersen, B. S., Sedlazeck, F. J., Quinlan, A. R. & Layer, R. M. Searching thousands of genomes to classify somatic and novel structural variants using STIX. Nat. Methods 19, 445–448 (2022).
pubmed: 35396485
pmcid: 9007735
doi: 10.1038/s41592-022-01423-4
Li, H. et al. A synthetic-diploid benchmark for accurate variant-calling evaluation. Nat. Methods 15, 595–597 (2018).
pubmed: 30013044
pmcid: 6341484
doi: 10.1038/s41592-018-0054-7
Byrska-Bishop, M. et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426–3440.e19 (2022).
pubmed: 36055201
pmcid: 9439720
doi: 10.1016/j.cell.2022.08.004
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018).
pubmed: 29750242
pmcid: 6137996
doi: 10.1093/bioinformatics/bty191
Lin, Y. et al. Assembly of long error-prone reads using de Bruijn graphs. Proc. Natl Acad. Sci. USA 113, E8396–E8405 (2016).
pubmed: 27956617
pmcid: 5206522
doi: 10.1073/pnas.1604560113
Mikheenko, A., Prjibelski, A., Saveliev, V., Antipov, D. & Gurevich, A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics 34, i142–i150 (2018).
pubmed: 29949969
pmcid: 6022658
doi: 10.1093/bioinformatics/bty266
Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–175 (2021).
pubmed: 33526886
pmcid: 7961889
doi: 10.1038/s41592-020-01056-5
Cheng, H. et al. Haplotype-resolved assembly of diploid genomes without parental data. Nat. Biotechnol. 40, 1332–1335 (2022).
pubmed: 35332338
doi: 10.1038/s41587-022-01261-x
Li, H., Feng, X. & Chu, C. The design and construction of reference pangenome graphs with minigraph. Genome Biol. 21, 265 (2020).
pubmed: 33066802
pmcid: 7568353
doi: 10.1186/s13059-020-02168-z
Wick, R. R., Schultz, M. B., Zobel, J. & Holt, K. E. Bandage: interactive visualization of de novo genome assemblies. Bioinformatics 31, 3350–3352 (2015).
pubmed: 26099265
pmcid: 4595904
doi: 10.1093/bioinformatics/btv383
Heller, D. & Vingron, M. SVIM-asm: structural variant detection from haploid and diploid genome assemblies. Bioinformatics 36, 5519–5521 (2020).
pmcid: 8016491
doi: 10.1093/bioinformatics/btaa1034
Razaghi, R. et al. Modbamtools: analysis of single-molecule epigenetic data for long-range profiling, heterogeneity, and clustering. Preprint at bioRxiv https://doi.org/10.1101/2022.07.07.499188 (2022).