Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved variation and methylation.


Journal

Nature methods
ISSN: 1548-7105
Titre abrégé: Nat Methods
Pays: United States
ID NLM: 101215604

Informations de publication

Date de publication:
10 2023
Historique:
received: 28 01 2023
accepted: 04 08 2023
medline: 9 10 2023
pubmed: 15 9 2023
entrez: 14 9 2023
Statut: ppublish

Résumé

Long-read sequencing technologies substantially overcome the limitations of short-reads but have not been considered as a feasible replacement for population-scale projects, being a combination of too expensive, not scalable enough or too error-prone. Here we develop an efficient and scalable wet lab and computational protocol, Napu, for Oxford Nanopore Technologies long-read sequencing that seeks to address those limitations. We applied our protocol to cell lines and brain tissue samples as part of a pilot project for the National Institutes of Health Center for Alzheimer's and Related Dementias. Using a single PromethION flow cell, we can detect single nucleotide polymorphisms with F1-score comparable to Illumina short-read sequencing. Small indel calling remains difficult within homopolymers and tandem repeats, but achieves good concordance to Illumina indel calls elsewhere. Further, we can discover structural variants with F1-score on par with state-of-the-art de novo assembly methods. Our protocol phases small and structural variants at megabase scales and produces highly accurate, haplotype-specific methylation calls.

Identifiants

pubmed: 37710018
doi: 10.1038/s41592-023-01993-x
pii: 10.1038/s41592-023-01993-x
doi:

Types de publication

Journal Article Research Support, N.I.H., Intramural Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Langues

eng

Sous-ensembles de citation

IM

Pagination

1483-1492

Subventions

Organisme : NHGRI NIH HHS
ID : U24 HG010262
Pays : United States
Organisme : NHGRI NIH HHS
ID : U24 HG011853
Pays : United States
Organisme : NHGRI NIH HHS
ID : U01 HG010961
Pays : United States
Organisme : NHGRI NIH HHS
ID : R01 HG010485
Pays : United States
Organisme : NIH HHS
ID : OT2 OD026682
Pays : United States
Organisme : NHGRI NIH HHS
ID : T32 HG012344
Pays : United States
Organisme : Intramural NIH HHS
ID : ZIA NS003154
Pays : United States
Organisme : Intramural NIH HHS
ID : ZIA AG000538
Pays : United States
Organisme : NINDS NIH HHS
ID : U24 NS072026
Pays : United States
Organisme : NIA NIH HHS
ID : P30 AG019610
Pays : United States
Organisme : NIA NIH HHS
ID : P30 AG072980
Pays : United States
Organisme : NHLBI NIH HHS
ID : OT3 HL142481
Pays : United States
Organisme : NIH HHS
ID : OT2 OD033761
Pays : United States

Commentaires et corrections

Type : UpdateOf

Informations de copyright

© 2023. The Author(s), under exclusive licence to Springer Nature America, Inc.

Références

DePristo, M. A. et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat. Genet. 43, 491–498 (2011).
pubmed: 21478889 pmcid: 3083463 doi: 10.1038/ng.806
1000 Genomes Project Consortium et al. An integrated map of genetic variation from 1,092 human genomes. Nature 491, 56–65 (2012).
doi: 10.1038/nature11632
100,000 Genomes Project Pilot Investigators et al. 100,000 Genomes pilot on rare-disease diagnosis in health care—preliminary report. N. Engl. J. Med. 385, 1868–1880 (2021).
Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020).
pubmed: 32461654 pmcid: 7334197 doi: 10.1038/s41586-020-2308-7
Huang, K.-L. et al. Pathogenic germline variants in 10,389 adult cancers. Cell 173, 355–370.e14 (2018).
pubmed: 29625052 pmcid: 5949147 doi: 10.1016/j.cell.2018.03.039
ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium. Pan-cancer analysis of whole genomes. Nature 578, 82–93 (2020).
doi: 10.1038/s41586-020-1969-6
Sedlazeck, F. J., Lee, H., Darby, C. A. & Schatz, M. C. Piercing the dark matter: bioinformatics of long-range sequencing and mapping. Nat. Rev. Genet. 19, 329–346 (2018).
pubmed: 29599501 doi: 10.1038/s41576-018-0003-4
Chen, X. et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics 32, 1220–1222 (2016).
pubmed: 26647377 doi: 10.1093/bioinformatics/btv710
Mahmoud, M. et al. Structural variant calling: the long and the short of it. Genome Biol. 20, 246 (2019).
pubmed: 31747936 pmcid: 6868818 doi: 10.1186/s13059-019-1828-7
Zarate, S. et al. Parliament2: accurate structural variant calling at scale. Gigascience 9, giaa145 (2020).
pubmed: 33347570 pmcid: 7751401 doi: 10.1093/gigascience/giaa145
Zook, J. M. et al. A robust benchmark for detection of germline large deletions and insertions. Nat. Biotechnol. 38, 1347–1355 (2020).
pubmed: 32541955 pmcid: 8454654 doi: 10.1038/s41587-020-0538-8
Wagner, J. et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat. Biotechnol. 40, 672–680 (2022).
pubmed: 35132260 pmcid: 9117392 doi: 10.1038/s41587-021-01158-1
Wagner, J. et al. Benchmarking challenging small variants with linked and long reads. Cell Genom. 2, 100128 (2022).
pubmed: 36452119 pmcid: 9706577 doi: 10.1016/j.xgen.2022.100128
Lee, H. & Schatz, M. C. Genomic dark matter: the reliability of short read mapping illustrated by the genome mappability score. Bioinformatics 28, 2097–2105 (2012).
pubmed: 22668792 pmcid: 3413383 doi: 10.1093/bioinformatics/bts330
Martin, M. et al. WhatsHap: fast and accurate read-based phasing. Preprint at bioRxiv https://doi.org/10.1101/085050 (2016).
Loh, P.-R. et al. Reference-based phasing using the Haplotype Reference Consortium panel. Nat. Genet. 48, 1443–1448 (2016).
pubmed: 27694958 pmcid: 5096458 doi: 10.1038/ng.3679
Sedlazeck, F. J. et al. Accurate detection of complex structural variations using single-molecule sequencing. Nat. Methods 15, 461–468 (2018).
pubmed: 29713083 pmcid: 5990442 doi: 10.1038/s41592-018-0001-7
Jiang, T. et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 21, 189 (2020).
pubmed: 32746918 pmcid: 7477834 doi: 10.1186/s13059-020-02107-y
Shafin, K. et al. Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads. Nat. Methods 18, 1322–1332 (2021).
pubmed: 34725481 pmcid: 8571015 doi: 10.1038/s41592-021-01299-w
Lin, J.-H., Chen, L.-C., Yu, S.-C. & Huang, Y.-T. LongPhase: an ultra-fast chromosome-scale phasing algorithm for small and large variants. Bioinformatics 38, 1816–1822 (2022).
pubmed: 35104333 doi: 10.1093/bioinformatics/btac058
Mahmoud, M., Doddapaneni, H., Timp, W. & Sedlazeck, F. J. PRINCESS: comprehensive detection of haplotype resolved SNVs, SVs, and methylation. Genome Biol. 22, 268 (2021).
pubmed: 34521442 pmcid: 8442460 doi: 10.1186/s13059-021-02486-w
Logsdon, G. A., Vollger, M. R. & Eichler, E. E. Long-read human genome sequencing and its applications. Nat. Rev. Genet. 21, 597–614 (2020).
pubmed: 32504078 pmcid: 7877196 doi: 10.1038/s41576-020-0236-x
Rhie, A. et al. Towards complete and error-free genome assemblies of all vertebrate species. Nature 592, 737–746 (2021).
pubmed: 33911273 pmcid: 8081667 doi: 10.1038/s41586-021-03451-0
Nurk, S. et al. The complete sequence of a human genome. Science 376, 44–53 (2022).
pubmed: 35357919 pmcid: 9186530 doi: 10.1126/science.abj6987
Liao, W.-W. et al. A draft human pangenome reference. Nature 617, 312–324 (2023).
pubmed: 37165242 pmcid: 10172123 doi: 10.1038/s41586-023-05896-x
Jarvis, E. D. et al. Automated assembly of high-quality diploid human reference genomes. Nature 611, 519–531 (2022).
pubmed: 36261518 pmcid: 9668749 doi: 10.1038/s41586-022-05325-5
Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive -mer weighting and repeat separation. Genome Res. 27, 722–736 (2017).
pubmed: 28298431 pmcid: 5411767 doi: 10.1101/gr.215087.116
Jain, M. et al. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nat. Biotechnol. 36, 338–345 (2018).
pubmed: 29431738 pmcid: 5889714 doi: 10.1038/nbt.4060
Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P. A. Assembly of long, error-prone reads using repeat graphs. Nat. Biotechnol. 37, 540–546 (2019).
pubmed: 30936562 doi: 10.1038/s41587-019-0072-8
Shafin, K. et al. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes. Nat. Biotechnol. 38, 1044–1053 (2020).
pubmed: 32686750 pmcid: 7483855 doi: 10.1038/s41587-020-0503-6
Rautiainen, M. et al. Verkko: telomere-to-telomere assembly of diploid chromosomes. Nat. Biotechnol. https://doi.org/10.1038/s41587-023-01662-6 (2023).
Billingsley, K. J. et al. Processing human frontal cortex brain tissue for population-scale Oxford Nanopore long-read DNA sequencing SOP v2. protocols.io https://doi.org/10.17504/protocols.io.kxygxzmmov8j/v2 (2022).
Baker, B. et al. Processing human frontal cortex brain tissue for population-scale SQK-LSK114 Oxford Nanopore long-read DNA sequencing SOP v1. protocols.io https://doi.org/10.17504/protocols.io.kxygx3zzog8j/v1 (2022).
Alvarez Jerez, P. et al. Processing frozen cells for population-scale Oxford Nanopore long-read DNA sequencing SOP v1. protocols.io https://doi.org/10.17504/protocols.io.5jyl8pnk7g2w/v1 (2022).
Gibbs, J. R. et al. Abundant quantitative trait loci exist for DNA methylation and gene expression in human brain. PLoS Genet. 6, e1000952 (2010).
pubmed: 20485568 pmcid: 2869317 doi: 10.1371/journal.pgen.1000952
Schatz, M. C. et al. Inverting the model of genomics data sharing with the NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space. Cell Genom. 2, 100085 (2022).
pubmed: 35199087 pmcid: 8863334 doi: 10.1016/j.xgen.2021.100085
Li, H. yak: yet another k-mer analyzer. GitHub https://github.com/lh3/yak (2023).
Smolka, M. et al. Comprehensive structural variant detection: from mosaic to population-level. Preprint at bioRxiv https://doi.org/10.1101/2022.04.04.487055 (2022).
English, A. C., Menon, V. K., Gibbs, R. A., Metcalf, G. A. & Sedlazeck, F. J. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol. 23, 271 (2022).
pubmed: 36575487 pmcid: 9793516 doi: 10.1186/s13059-022-02840-6
Yang, J. & Chaisson, M. J. P. TT-Mars: structural variants assessment based on haplotype-resolved assemblies. Genome Biol. 23, 110 (2022).
pubmed: 35524317 pmcid: 9077962 doi: 10.1186/s13059-022-02666-2
Vollger, M. R. et al. Long-read sequence and assembly of segmental duplications. Nat. Methods 16, 88–94 (2019).
pubmed: 30559433 doi: 10.1038/s41592-018-0236-3
Kirsche, M. et al. Jasmine: population-scale structural variant comparison and analysis. Nat. Methods 20, 408–417 (2023).
pubmed: 36658279 pmcid: 10006329 doi: 10.1038/s41592-022-01753-3
Chowdhury, M., Pedersen, B. S., Sedlazeck, F. J., Quinlan, A. R. & Layer, R. M. Searching thousands of genomes to classify somatic and novel structural variants using STIX. Nat. Methods 19, 445–448 (2022).
pubmed: 35396485 pmcid: 9007735 doi: 10.1038/s41592-022-01423-4
Li, H. et al. A synthetic-diploid benchmark for accurate variant-calling evaluation. Nat. Methods 15, 595–597 (2018).
pubmed: 30013044 pmcid: 6341484 doi: 10.1038/s41592-018-0054-7
Byrska-Bishop, M. et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426–3440.e19 (2022).
pubmed: 36055201 pmcid: 9439720 doi: 10.1016/j.cell.2022.08.004
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018).
pubmed: 29750242 pmcid: 6137996 doi: 10.1093/bioinformatics/bty191
Lin, Y. et al. Assembly of long error-prone reads using de Bruijn graphs. Proc. Natl Acad. Sci. USA 113, E8396–E8405 (2016).
pubmed: 27956617 pmcid: 5206522 doi: 10.1073/pnas.1604560113
Mikheenko, A., Prjibelski, A., Saveliev, V., Antipov, D. & Gurevich, A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics 34, i142–i150 (2018).
pubmed: 29949969 pmcid: 6022658 doi: 10.1093/bioinformatics/bty266
Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–175 (2021).
pubmed: 33526886 pmcid: 7961889 doi: 10.1038/s41592-020-01056-5
Cheng, H. et al. Haplotype-resolved assembly of diploid genomes without parental data. Nat. Biotechnol. 40, 1332–1335 (2022).
pubmed: 35332338 doi: 10.1038/s41587-022-01261-x
Li, H., Feng, X. & Chu, C. The design and construction of reference pangenome graphs with minigraph. Genome Biol. 21, 265 (2020).
pubmed: 33066802 pmcid: 7568353 doi: 10.1186/s13059-020-02168-z
Wick, R. R., Schultz, M. B., Zobel, J. & Holt, K. E. Bandage: interactive visualization of de novo genome assemblies. Bioinformatics 31, 3350–3352 (2015).
pubmed: 26099265 pmcid: 4595904 doi: 10.1093/bioinformatics/btv383
Heller, D. & Vingron, M. SVIM-asm: structural variant detection from haploid and diploid genome assemblies. Bioinformatics 36, 5519–5521 (2020).
pmcid: 8016491 doi: 10.1093/bioinformatics/btaa1034
Razaghi, R. et al. Modbamtools: analysis of single-molecule epigenetic data for long-range profiling, heterogeneity, and clustering. Preprint at bioRxiv https://doi.org/10.1101/2022.07.07.499188 (2022).

Auteurs

Mikhail Kolmogorov (M)

Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, MD, USA. mikhail.kolmogorov@nih.gov.

Kimberley J Billingsley (KJ)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA. kimberley.billingsley@nih.gov.
Laboratory of Neurogenetics, National Institute on Aging, National Institutes of Health, Bethesda, MD, USA. kimberley.billingsley@nih.gov.

Mira Mastoras (M)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Melissa Meredith (M)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Jean Monlong (J)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Ryan Lorig-Roach (R)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Mobin Asri (M)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Pilar Alvarez Jerez (P)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.

Laksh Malik (L)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.

Ramita Dewan (R)

Laboratory of Neurogenetics, National Institute on Aging, National Institutes of Health, Bethesda, MD, USA.

Xylena Reed (X)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.

Rylee M Genner (RM)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.

Kensuke Daida (K)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.
Laboratory of Neurogenetics, National Institute on Aging, National Institutes of Health, Bethesda, MD, USA.

Sairam Behera (S)

Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, USA.

Kishwar Shafin (K)

Google LLC, Mountain View, CA, USA.

Trevor Pesout (T)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Jeshuwin Prabakaran (J)

Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, MD, USA.
Department of Biological Sciences, University of Maryland Baltimore County, Baltimore, MD, USA.

Paolo Carnevali (P)

Chan Zuckerberg Initiative, Redwood City, CA, USA.

Jianzhi Yang (J)

Department of Quantitative and Computational Biology, University of Southern California, Los Angeles, CA, USA.

Arang Rhie (A)

Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, National Institutes of Health, Bethesda, MD, USA.

Sonja W Scholz (SW)

Neurodegenerative Diseases Research Unit, National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.
Department of Neurology, Johns Hopkins University Medical Center, Baltimore, MD, USA.

Bryan J Traynor (BJ)

Laboratory of Neurogenetics, National Institute on Aging, National Institutes of Health, Bethesda, MD, USA.
Department of Neurology, Johns Hopkins University Medical Center, Baltimore, MD, USA.

Karen H Miga (KH)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA.

Miten Jain (M)

Department of Bioengineering, Northeastern University, Boston, MA, USA.
Department of Physics, Northeastern University, Boston, MA, USA.

Winston Timp (W)

Department of Biomedical Engineering, Johns Hopkins University, Baltimore, MD, USA.

Adam M Phillippy (AM)

Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, National Institutes of Health, Bethesda, MD, USA.

Mark Chaisson (M)

Department of Quantitative and Computational Biology, University of Southern California, Los Angeles, CA, USA.

Fritz J Sedlazeck (FJ)

Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, USA.
Department of Computer Science, Rice University, Houston, TX, USA.

Cornelis Blauwendraat (C)

Center for Alzheimer's and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA. cornelis.blauwendraat@nih.gov.
Laboratory of Neurogenetics, National Institute on Aging, National Institutes of Health, Bethesda, MD, USA. cornelis.blauwendraat@nih.gov.

Benedict Paten (B)

UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA. bpaten@ucsc.edu.

Articles similaires

Genome, Chloroplast Phylogeny Genetic Markers Base Composition High-Throughput Nucleotide Sequencing

[Redispensing of expensive oral anticancer medicines: a practical application].

Lisanne N van Merendonk, Kübra Akgöl, Bastiaan Nuijen
1.00
Humans Antineoplastic Agents Administration, Oral Drug Costs Counterfeit Drugs

Smoking Cessation and Incident Cardiovascular Disease.

Jun Hwan Cho, Seung Yong Shin, Hoseob Kim et al.
1.00
Humans Male Smoking Cessation Cardiovascular Diseases Female
Humans United States Aged Cross-Sectional Studies Medicare Part C

Classifications MeSH