A cost-effective sequencing method for genetic studies combining high-depth whole exome and low-depth whole genome.


Journal

NPJ genomic medicine
ISSN: 2056-7944
Titre abrégé: NPJ Genom Med
Pays: England
ID NLM: 101685193

Informations de publication

Date de publication:
07 Feb 2024
Historique:
received: 12 05 2023
accepted: 07 12 2023
medline: 8 2 2024
pubmed: 8 2 2024
entrez: 7 2 2024
Statut: epublish

Résumé

Whole genome sequencing (WGS) at high-depth (30X) allows the accurate discovery of variants in the coding and non-coding DNA regions and helps elucidate the genetic underpinnings of human health and diseases. Yet, due to the prohibitive cost of high-depth WGS, most large-scale genetic association studies use genotyping arrays or high-depth whole exome sequencing (WES). Here we propose a cost-effective method which we call "Whole Exome Genome Sequencing" (WEGS), that combines low-depth WGS and high-depth WES with up to 8 samples pooled and sequenced simultaneously (multiplexed). We experimentally assess the performance of WEGS with four different depth of coverage and sample multiplexing configurations. We show that the optimal WEGS configurations are 1.7-2.0 times cheaper than standard WES (no-plexing), 1.8-2.1 times cheaper than high-depth WGS, reach similar recall and precision rates in detecting coding variants as WES, and capture more population-specific variants in the rest of the genome that are difficult to recover when using genotype imputation methods. We apply WEGS to 862 patients with peripheral artery disease and show that it directly assesses more known disease-associated variants than a typical genotyping array and thousands of non-imputable variants per disease-associated locus.

Identifiants

pubmed: 38326393
doi: 10.1038/s41525-024-00390-3
pii: 10.1038/s41525-024-00390-3
doi:

Types de publication

Journal Article

Langues

eng

Pagination

8

Informations de copyright

© 2024. The Author(s).

Références

Halldorsson, B. V. et al. The sequences of 150,119 genomes in the UK Biobank. Nature 607, 732–740 (2022).
doi: 10.1038/s41586-022-04965-x pubmed: 35859178 pmcid: 9329122
Taliun, D. et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590, 290–299 (2021).
doi: 10.1038/s41586-021-03205-y pubmed: 33568819 pmcid: 7875770
Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020).
doi: 10.1038/s41586-020-2308-7 pubmed: 32461654 pmcid: 7334197
Backman, J. D. et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature 599, 628–634 (2021).
doi: 10.1038/s41586-021-04103-z pubmed: 34662886 pmcid: 8596853
McCarthy, S. et al. A reference panel of 64,976 haplotypes for genotype imputation. Nat. Genet. 48, 1279–1283 (2016).
doi: 10.1038/ng.3643 pubmed: 27548312 pmcid: 5388176
1000 Genomes Project Consortium. et al. A global reference for human genetic variation. Nature 526, 68–74 (2015).
doi: 10.1038/nature15393
Quick, C. et al. Sequencing and imputation in GWAS: Cost‐effective strategies to increase power and genomic coverage across diverse populations. Genet. Epidemiol. 44, 537–549 Preprint at https://doi.org/10.1002/gepi.22326 (2020).
doi: 10.1002/gepi.22326 pubmed: 32519380 pmcid: 7449570
Mitt, M. et al. Improved imputation accuracy of rare and low-frequency variants using population-specific high-coverage WGS-based imputation reference panel. Eur. J. Hum. Genet. 25, 869–876 (2017).
doi: 10.1038/ejhg.2017.51 pubmed: 28401899 pmcid: 5520064
Kurki, M. I. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature 613, 508–518 (2023).
doi: 10.1038/s41586-022-05473-8 pubmed: 36653562 pmcid: 9849126
Pistis, G. et al. Rare variant genotype imputation with thousands of study-specific whole-genome sequences: implications for cost-effective study designs. Eur. J. Hum. Genet. 23, 975–983 (2015).
doi: 10.1038/ejhg.2014.216 pubmed: 25293720
Martin, A. R. et al. Low-coverage sequencing cost-effectively detects known and novel variation in underrepresented populations. Am. J. Hum. Genet. 108, 656–668 (2021).
doi: 10.1016/j.ajhg.2021.03.012 pubmed: 33770507 pmcid: 8059370
Homburger, J. R. et al. Low coverage whole genome sequencing enables accurate assessment of common variants and calculation of genome-wide polygenic scores. Genome Med. 11 (2019).
Li, J. H., Mazur, C. A., Berisa, T. & Pickrell, J. K. Low-pass sequencing increases the power of GWAS and decreases measurement error of polygenic risk scores compared to genotyping arrays. Genome Res. 31 (2021).
Gilly, A. et al. Very low-depth whole-genome sequencing in complex trait association studies. Bioinformatics 35 (2019).
Darst, B. F. et al. Combined Effect of a Polygenic Risk Score and Rare Genetic Variants on Prostate Cancer Risk. Eur. Urol. 80 (2021).
Lali, R. et al. Calibrated rare variant genetic risk scores for complex disease prediction using large exome sequence repositories. Nat. Commun. 12 (2021).
Dornbos, P. et al. A combined polygenic score of 21,293 rare and 22 common variants improves diabetes diagnosis based on hemoglobin A1C levels. Nat. Genet. 54 (2022).
Sollis, E. et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 51, D977–D985 (2023).
doi: 10.1093/nar/gkac1010 pubmed: 36350656
Wong, K. H., Jin, Y. & Moqtaderi, Z. Multiplex Illumina Sequencing Using DNA Barcoding. Curr. Protoc. Mol. Biol. Chapter 7 Unit 7.11, (2013).
Vodák, D. et al. Sample-Index Misassignment Impacts Tumour Exome Sequencing. Sci. Rep. 8, 5307 (2018).
doi: 10.1038/s41598-018-23563-4 pubmed: 29593270 pmcid: 5871786
Marx, V. How to deduplicate PCR. Nat. Methods 14, 473–476 (2017).
doi: 10.1038/nmeth.4268 pubmed: 28448070
Kivioja, T. et al. Counting absolute numbers of molecules using unique molecular identifiers. Nat. Methods 9, 72–74 (2011).
doi: 10.1038/nmeth.1778 pubmed: 22101854
Tsagiopoulou, M. et al. UMIc: A Preprocessing Method for UMI Deduplication and Reads Correction. Front. Genet. 12, 660366 (2021).
doi: 10.3389/fgene.2021.660366 pubmed: 34122513 pmcid: 8193862
Chou, W.-C. et al. A combined reference panel from the 1000 Genomes and UK10K projects improved rare variant imputation in European and Chinese samples. Sci. Rep. 6, 39313 (2016).
doi: 10.1038/srep39313 pubmed: 28004816 pmcid: 5177868
Roshyara, N. R. & Scholz, M. Impact of genetic similarity on imputation accuracy. BMC Genet. 16, 90 (2015).
doi: 10.1186/s12863-015-0248-2 pubmed: 26193934 pmcid: 4509609
Rubinacci, S., Ribeiro, D. M., Hofmeister, R. J. & Delaneau, O. Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat. Genet. 53, 120–126 (2021).
doi: 10.1038/s41588-020-00756-0 pubmed: 33414550
Das, S. et al. Next-generation genotype imputation service and methods. Nat. Genet. 48, 1284–1287 (2016).
doi: 10.1038/ng.3656 pubmed: 27571263 pmcid: 5157836
Klarin, D. et al. Genome-wide association study of peripheral artery disease in the Million Veteran Program. Nat. Med. 25, 1274–1279 (2019).
doi: 10.1038/s41591-019-0492-5 pubmed: 31285632 pmcid: 6768096
DePristo, M. A. et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat. Genet. 43, 491–498 (2011).
doi: 10.1038/ng.806 pubmed: 21478889 pmcid: 3083463
Trost, B. et al. Impact of DNA source on genetic variant detection from human whole-genome sequencing data. J. Med. Genet 56, 809–817 (2019).
doi: 10.1136/jmedgenet-2019-106281 pubmed: 31515274
Van der Auwera, G. A. et al. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline. Curr. Protoc. Bioinforma. 43, 11.10.1–11.10.33 (2013).
De Summa, S. et al. GATK hard filtering: tunable parameters to improve variant calling for next generation sequencing targeted gene panel data. BMC Bioinformatics 18 (2017).
Koboldt, D. C. Best practices for variant calling in clinical sequencing. Genome Med. 12 (2020).
Zheng, J. et al. A comprehensive assessment of Next‐Generation Sequencing variants validation using a secondary technology. Mol. Genet. Genomic Med. 7 (2019).
Marchini, J. & Howie, B. Genotype imputation for genome-wide association studies. Nat. Rev. Genet. 11, 499–511 (2010).
doi: 10.1038/nrg2796 pubmed: 20517342
Sun, Q. et al. MagicalRsq: Machine-learning-based genotype imputation quality calibration. Am. J. Hum. Genet. 109, 1986–1997 (2022).
doi: 10.1016/j.ajhg.2022.09.009 pubmed: 36198314 pmcid: 9674945
Ball, M. P. et al. A public resource facilitating clinical use of genomes. PNAS 109, 11920–11927 (2012).
doi: 10.1073/pnas.1201904109 pubmed: 22797899 pmcid: 3409785
Zook, J. M. et al. Extensive sequencing of seven human genomes to characterize benchmark reference materials. Sci. Data 3, 1–26 (2016).
doi: 10.1038/sdata.2016.25
Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. https://doi.org/10.48550/arXiv.1303.3997 (2013).
Jiang, H., Lei, R., Ding, S.-W. & Zhu, S. Skewer: a fast and accurate adapter trimmer for next-generation sequencing paired-end reads. BMC Bioinforma. 15, 182 (2014).
doi: 10.1186/1471-2105-15-182
Olson, N. D. et al. PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult-to-map regions. Cell Genom 2 (2022).
Wagner, J. et al. Benchmarking challenging small variants with linked and long reads. Cell Genom. 2, 100128 (2022).
doi: 10.1016/j.xgen.2022.100128 pubmed: 36452119 pmcid: 9706577
Zook, J. M. et al. An open resource for accurately benchmarking small variant and reference calls. Nat. Biotechnol. 37, 561–566 (2019).
doi: 10.1038/s41587-019-0074-6 pubmed: 30936564 pmcid: 6500473
Zook, J. M. et al. A robust benchmark for detection of germline large deletions and insertions. Nat. Biotechnol. 38, 1347–1355 (2020).
doi: 10.1038/s41587-020-0538-8 pubmed: 32541955 pmcid: 8454654
Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079 (2009).
doi: 10.1093/bioinformatics/btp352 pubmed: 19505943 pmcid: 2723002
Virtanen, P. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods 17, 261–272 (2020).
doi: 10.1038/s41592-019-0686-2 pubmed: 32015543 pmcid: 7056644
Van der Auwera, G. A. & O’Connor, B. D. Genomics in the Cloud: Using Docker, GATK, and WDL in Terra. (O’Reilly Media, 2020).
Cavalli-Sforza, L. L. The Human Genome Diversity Project: past, present and future. Nat. Rev. Genet. 6, 333–340 (2005).
doi: 10.1038/nrg1579 pubmed: 15803201
Delaneau, O., Zagury, J.-F., Robinson, M. R., Marchini, J. L. & Dermitzakis, E. T. Accurate, scalable and integrative haplotype estimation. Nat. Commun. 10, 1–10 (2019).
doi: 10.1038/s41467-019-13225-y
Wang, C., Zhan, X., Liang, L., Abecasis, G. R. & Lin, X. Improved ancestry estimation for both genotyping and sequencing data using projection Procrustes analysis and genotype imputation. Am. J. Hum. Genet. 96, 926–937 (2015).
doi: 10.1016/j.ajhg.2015.04.018 pubmed: 26027497 pmcid: 4457959
Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. arXiv [cs.LG] (2012).

Auteurs

Claude Bhérer (C)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada.

Robert Eveleigh (R)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canadian Centre for Computational Genomics, McGill University, Montréal, Québec, Canada.

Katerina Trajanoska (K)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada.

Janick St-Cyr (J)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.

Antoine Paccard (A)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.

Praveen Nadukkalam Ravindran (P)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada.

Elizabeth Caron (E)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.

Nimara Bader Asbah (N)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.

Peyton McClelland (P)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada.

Clare Wei (C)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada.

Iris Baumgartner (I)

Division of Angiology, Swiss Cardiovascular Center, Inselspital, Bern University Hospital, University of Bern, Bern, Switzerland.
Department for BioMedical Research (DBMR), University of Bern, Bern, Switzerland.

Marc Schindewolf (M)

Division of Angiology, Swiss Cardiovascular Center, Inselspital, Bern University Hospital, University of Bern, Bern, Switzerland.
Department for BioMedical Research (DBMR), University of Bern, Bern, Switzerland.

Yvonne Döring (Y)

Division of Angiology, Swiss Cardiovascular Center, Inselspital, Bern University Hospital, University of Bern, Bern, Switzerland.
Department for BioMedical Research (DBMR), University of Bern, Bern, Switzerland.
Institute for Cardiovascular Prevention (IPEK), Ludwig-Maximilians University Munich, Pettenkoferstr 9, 80336, Munich, Germany.

Danielle Perley (D)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canadian Centre for Computational Genomics, McGill University, Montréal, Québec, Canada.

François Lefebvre (F)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canadian Centre for Computational Genomics, McGill University, Montréal, Québec, Canada.

Pierre Lepage (P)

Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.

Mathieu Bourgey (M)

Mila - Quebec AI Institute, Montréal, Québec, Canada.

Guillaume Bourque (G)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canadian Centre for Computational Genomics, McGill University, Montréal, Québec, Canada.

Jiannis Ragoussis (J)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.

Vincent Mooser (V)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada.

Daniel Taliun (D)

Department of Human Genetics, Faculty of Medicine and Health Sciences, McGill University, Montréal, Québec, Canada. daniel.taliun@mcgill.ca.
Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec, Canada. daniel.taliun@mcgill.ca.
Canada Excellence Research Chair in Genomic Medicine, McGill University, Montréal, Québec, Canada. daniel.taliun@mcgill.ca.

Classifications MeSH