Rare coding variant analysis for human diseases across biobanks and ancestries.
Journal
Nature genetics
ISSN: 1546-1718
Titre abrégé: Nat Genet
Pays: United States
ID NLM: 9216904
Informations de publication
Date de publication:
29 Aug 2024
29 Aug 2024
Historique:
received:
26
02
2023
accepted:
01
08
2024
medline:
31
8
2024
pubmed:
31
8
2024
entrez:
29
8
2024
Statut:
aheadofprint
Résumé
Large-scale sequencing has enabled unparalleled opportunities to investigate the role of rare coding variation in human phenotypic variability. Here, we present a pan-ancestry analysis of sequencing data from three large biobanks, including the All of Us research program. Using mixed-effects models, we performed gene-based rare variant testing for 601 diseases across 748,879 individuals, including 155,236 with ancestry dissimilar to European. We identified 363 significant associations, which highlighted core genes for the human disease phenome and identified potential novel associations, including UBR3 for cardiometabolic disease and YLPM1 for psychiatric disease. Pan-ancestry burden testing represented an inclusive and useful approach for discovery in diverse datasets, although we also highlight the importance of ancestry-specific sensitivity analyses in this setting. Finally, we found that effect sizes for rare protein-disrupting variants were concordant between samples similar to European ancestry and other genetic ancestries (β
Identifiants
pubmed: 39210047
doi: 10.1038/s41588-024-01894-5
pii: 10.1038/s41588-024-01894-5
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Subventions
Organisme : Foundation for the National Institutes of Health (Foundation for the National Institutes of Health, Inc.)
ID : 1RO1HL092577
Organisme : Foundation for the National Institutes of Health (Foundation for the National Institutes of Health, Inc.)
ID : 1R01HL157635
Organisme : Foundation for the National Institutes of Health (Foundation for the National Institutes of Health, Inc.)
ID : K08HL159346
Organisme : Foundation for the National Institutes of Health (Foundation for the National Institutes of Health, Inc.)
ID : 1K08HL153937
Organisme : American Heart Association (American Heart Association, Inc.)
ID : 961045
Organisme : American Heart Association (American Heart Association, Inc.)
ID : 18SFRN34110082
Organisme : American Heart Association (American Heart Association, Inc.)
ID : 18SFRN34110082
Organisme : American Heart Association (American Heart Association, Inc.)
ID : 862032
Organisme : Hartstichting (Dutch Heart Foundation)
ID : 03-007-2022-0035
Informations de copyright
© 2024. The Author(s), under exclusive licence to Springer Nature America, Inc.
Références
Backman, J. D. et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature 599, 628–634 (2021).
pubmed: 34662886
pmcid: 8596853
doi: 10.1038/s41586-021-04103-z
Taliun, D. et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590, 290–299 (2021).
pubmed: 33568819
pmcid: 7875770
doi: 10.1038/s41586-021-03205-y
Wang, Q. et al. Rare variant contribution to human disease in 281,104 UK Biobank exomes. Nature 597, 527–532 (2021).
pubmed: 34375979
pmcid: 8458098
doi: 10.1038/s41586-021-03855-y
Karczewski, K. J. et al. Systematic single-variant and gene-based association testing of thousands of phenotypes in 394,841 UK Biobank exomes. Cell Genomics 2, 100168 (2022).
pubmed: 36778668
pmcid: 9903662
doi: 10.1016/j.xgen.2022.100168
Jurgens, S. J. et al. Analysis of rare genetic variation underlying cardiometabolic diseases and traits among 200,000 individuals in the UK Biobank. Nat. Genet. 54, 240–250 (2022).
pubmed: 35177841
pmcid: 8930703
doi: 10.1038/s41588-021-01011-w
Gudbjartsson, D. F. et al. Large-scale whole-genome sequencing of the Icelandic population. Nat. Genet. 47, 435–444 (2015).
pubmed: 25807286
doi: 10.1038/ng.3247
Sun, B. B. et al. Genetic associations of protein-coding variants in human disease. Nature 603, 95–102 (2022).
pubmed: 35197637
pmcid: 8891017
doi: 10.1038/s41586-022-04394-w
Heyne, H. O. et al. Mono- and biallelic variant effects on disease at biobank scale. Nature 613, 519–525 (2023).
pubmed: 36653560
pmcid: 9849130
doi: 10.1038/s41586-022-05420-7
Need, A. C. & Goldstein, D. B. Next generation disparities in human genomics: concerns and remedies. Trends Genet. 25, 489–494 (2009).
pubmed: 19836853
doi: 10.1016/j.tig.2009.09.012
Popejoy, A. B. & Fullerton, S. M. Genomics is failing on diversity. Nature 538, 161–164 (2016).
pubmed: 27734877
pmcid: 5089703
doi: 10.1038/538161a
Hindorff, L. A. et al. Prioritizing diversity in human genomics research. Nat. Rev. Genet. 19, 175–185 (2018).
pubmed: 29151588
doi: 10.1038/nrg.2017.89
Ramirez, H. A. et al. The All of Us Research Program: data quality, utility, and diversity. Patterns 3, 100570 (2022).
pubmed: 36033590
pmcid: 9403360
doi: 10.1016/j.patter.2022.100570
Gurdasani, D. et al. The African Genome Variation Project shapes medical genetics in Africa. Nature 517, 327–332 (2015).
pubmed: 25470054
doi: 10.1038/nature13997
Gaziano, J. M. et al. Million Veteran Program: a mega-biobank to study genetic influences on health and disease. J. Clin. Epidemiol. 70, 214–223 (2016).
pubmed: 26441289
doi: 10.1016/j.jclinepi.2015.09.016
All of Us Research Program Genomics Investigators. Genomic data in the All of Us research program. Nature 627, 340–346 (2024).
doi: 10.1038/s41586-023-06957-x
Koyama, S. et al. Decoding genetics, ancestry, and geospatial context for precision health. Preprint at medRxiv https://doi.org/10.1101/2023.10.24.23297096 (2023).
Denny, J. C. et al. The ‘All of Us’ research program. N. Engl. J. Med. 381, 668–676 (2019).
pubmed: 31412182
doi: 10.1056/NEJMsr1809937
Ding, Y. et al. Polygenic scoring accuracy varies across the genetic ancestry continuum. Nature 618, 774–781 (2023).
pubmed: 37198491
pmcid: 10284707
doi: 10.1038/s41586-023-06079-4
Auton, A. et al. A global reference for human genetic variation. Nature 526, 68–74 (2015).
pubmed: 26432245
doi: 10.1038/nature15393
Janssen, F., Bardoutsos, A. & Vidra, N. Obesity prevalence in the long-term future in 18 European countries and in the USA. Obes. Facts 13, 514–527 (2020).
pubmed: 33075798
pmcid: 7670332
doi: 10.1159/000511023
Marshall, A. et al. Comparison of hypertension healthcare outcomes among older people in the USA and England. J. Epidemiol. Community Health 70, 264–270 (2016).
pubmed: 26598759
doi: 10.1136/jech-2014-205336
Joffres, M. et al. Hypertension prevalence, awareness, treatment and control in national surveys from England, the USA and Canada, and correlation with stroke and ischaemic heart disease mortality: a cross-sectional study. BMJ Open 3, e003423 (2013).
pubmed: 23996822
pmcid: 3758966
doi: 10.1136/bmjopen-2013-003423
Matyori, A., Brown, C. P., Ali, A. & Sherbeny, F. Statins utilization trends and expenditures in the U.S. before and after the implementation of the 2013 ACC/AHA guidelines. Saudi Pharm. J. 31, 795–800 (2023).
pubmed: 37228328
pmcid: 10203693
doi: 10.1016/j.jsps.2023.04.002
Gao, Y., Shah, L. M., Ding, J. & Martin, S. S. US trends in cholesterol screening, lipid levels, and lipid-lowering medication use in US adults, 1999 to 2018. J. Am. Heart Assoc. 12, e028205 (2023).
pubmed: 36625302
pmcid: 9973640
doi: 10.1161/JAHA.122.028205
Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020).
pubmed: 32461654
pmcid: 7334197
doi: 10.1038/s41586-020-2308-7
Lek, M. et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature 536, 285–291 (2016).
pubmed: 27535533
pmcid: 5018207
doi: 10.1038/nature19057
Kasar, S. et al. Whole-genome sequencing reveals activation-induced cytidine deaminase signatures during indolent chronic lymphocytic leukaemia evolution. Nat. Commun. 6, 8866 (2015).
pubmed: 26638776
doi: 10.1038/ncomms9866
Jurgens, S. J. et al. Adjusting for common variant polygenic scores improves yield in rare variant association analyses. Nat. Genet. 55, 544–548 (2023).
pubmed: 36959364
pmcid: 11078202
doi: 10.1038/s41588-023-01342-w
Jaiswal, S. Clonal hematopoiesis and nonhematologic disorders. Blood 136, 1606–1614 (2020).
pubmed: 32736379
pmcid: 8209629
Asada, S. & Kitamura, T. Clonal hematopoiesis and associated diseases: a review of recent findings. Cancer Sci. 112, 3962–3971 (2021).
pubmed: 34328684
pmcid: 8486184
doi: 10.1111/cas.15094
Mitchell, E. et al. Clonal dynamics of haematopoiesis across the human lifespan. Nature 606, 343–350 (2022).
pubmed: 35650442
pmcid: 9177428
doi: 10.1038/s41586-022-04786-y
Ingles, J. et al. Evaluating the clinical validity of hypertrophic cardiomyopathy genes. Circ. Genom. Precis Med 12, e002460 (2019).
pubmed: 30681346
pmcid: 6410971
doi: 10.1161/CIRCGEN.119.002460
Walsh, R. et al. Reassessment of Mendelian gene pathogenicity using 7,855 cardiomyopathy cases and 60,706 reference samples. Genet. Med. 19, 192–203 (2017).
pubmed: 27532257
doi: 10.1038/gim.2016.90
National Academies of Sciences, Engineering, and Medicine; Policy and Global Affairs; Committee on Women in Science, Engineering, and Medicine; Committee on Improving the Representation of Women and Underrepresented Minorities in Clinical Trials Research. Improving Representation in Clinical Trials and Research: Building Research Equity for Women and Underrepresented Groups (National Academies Press, 2022).
Ward, E. et al. Cancer disparities by race/ethnicity and socioeconomic status. CA Cancer J. Clin. 54, 78–93 (2004).
pubmed: 15061598
doi: 10.3322/canjclin.54.2.78
Suther, S. & Kiros, G. E. Barriers to the use of genetic testing: a study of racial and ethnic disparities. Genet. Med. 11, 655–662 (2009).
pubmed: 19752639
doi: 10.1097/GIM.0b013e3181ab22aa
Wojcik, G. L. et al. Genetic analyses of diverse populations improves discovery for complex traits. Nature 570, 514–518 (2019).
pubmed: 31217584
pmcid: 6785182
doi: 10.1038/s41586-019-1310-4
Vujkovic, M. et al. Discovery of 318 new risk loci for type 2 diabetes and related vascular outcomes among 1.4 million participants in a multi-ancestry meta-analysis. Nat. Genet. 52, 680–691 (2020).
pubmed: 32541925
pmcid: 7343592
doi: 10.1038/s41588-020-0637-y
Graham, S. E. et al. The power of genetic diversity in genome-wide association studies of lipids. Nature 600, 675–679 (2021).
pubmed: 34887591
pmcid: 8730582
doi: 10.1038/s41586-021-04064-3
Wall, J. D. et al. South Asian medical cohorts reveal strong founder effects and high rates of homozygosity. Nat. Commun. 14, 3377 (2023).
pubmed: 37291107
pmcid: 10250394
doi: 10.1038/s41467-023-38766-1
Van Hout, C. V. et al. Exome sequencing and characterization of 49,960 individuals in the UK Biobank. Nature 586, 749–756 (2020).
pubmed: 33087929
pmcid: 7759458
doi: 10.1038/s41586-020-2853-0
Deaton, A. M. et al. Gene-level analysis of rare variants in 379,066 whole exome sequences identifies an association of GIGYF1 loss of function with type 2 diabetes. Sci. Rep. 11, 21565 (2021).
pubmed: 34732801
pmcid: 8566487
doi: 10.1038/s41598-021-99091-5
Välimäki, N. et al. Inherited mutations affecting the SRCAP complex are central in moderate-penetrance predisposition to uterine leiomyomas. Am. J. Hum. Genet. 110, 460–474 (2023).
pubmed: 36773604
pmcid: 10027472
doi: 10.1016/j.ajhg.2023.01.009
Haas, M. E. et al. Machine learning enables new insights into genetic contributions to liver fat accumulation. Cell Genom. 1, 100066 (2021).
pubmed: 34957434
pmcid: 8699145
doi: 10.1016/j.xgen.2021.100066
Khera, A. V. et al. Gene sequencing identifies perturbation in nitric oxide signaling as a nonlipid molecular subtype of coronary artery disease. Circ. Genom. Precis. Med. 15, e003598 (2022).
pubmed: 36215124
pmcid: 9771961
doi: 10.1161/CIRCGEN.121.003598
Ward, J. et al. Genome-wide analysis in UK Biobank identifies four loci associated with mood instability and genetic correlation with major depressive disorder, anxiety disorder and schizophrenia. Transl. Psychiatry 7, 1264 (2017).
pubmed: 29187730
pmcid: 5802589
doi: 10.1038/s41398-017-0012-7
Luciano, M. et al. Association analysis in over 329,000 individuals identifies 116 independent variants influencing neuroticism. Nat. Genet. 50, 6–11 (2018).
pubmed: 29255261
doi: 10.1038/s41588-017-0013-8
Nagel, M. et al. Meta-analysis of genome-wide association studies for neuroticism in 449,484 individuals identifies novel genetic loci and pathways. Nat. Genet. 50, 920–927 (2018).
pubmed: 29942085
doi: 10.1038/s41588-018-0151-7
Mountjoy, E. et al. An open approach to systematically prioritize causal variants and genes at all published human GWAS trait-associated loci. Nat. Genet. 53, 1527–1533 (2021).
pubmed: 34711957
pmcid: 7611956
doi: 10.1038/s41588-021-00945-5
Liu, F. R. et al. Pedigree-based study to identify GOLGB1 as a risk gene for bipolar disorder. Transl. Psychiatry 12, 390 (2022).
pubmed: 36115840
pmcid: 9482626
doi: 10.1038/s41398-022-02163-x
Palmer, D. S. et al. Exome sequencing in bipolar disorder identifies AKAP11 as a risk gene shared with schizophrenia. Nat. Genet. 54, 541–547 (2022).
pubmed: 35410376
pmcid: 9117467
doi: 10.1038/s41588-022-01034-x
Cui, J. et al. Disruption of Gpr45 causes reduced hypothalamic POMC expression and obesity. J. Clin. Invest. 126, 3192–3206 (2016).
pubmed: 27500489
pmcid: 5004970
doi: 10.1172/JCI85676
Akbari, P. et al. Sequencing of 640,000 exomes identifies GPR75 variants associated with protection from obesity. Science 373, eabf8683 (2021).
pubmed: 34210852
pmcid: 10275396
doi: 10.1126/science.abf8683
Yamazaki, O., Hirohama, D., Ishizawa, K. & Shibata, S. Role of the ubiquitin proteasome system in the regulation of blood pressure: a review. Int. J. Mol. Sci. 21, 5358 (2020).
pubmed: 32731518
pmcid: 7432568
doi: 10.3390/ijms21155358
Li, X. Y., Zhai, W. J. & Teng, C. B. Notch signaling in pancreatic development. Int. J. Mol. Sci. 17, 48 (2015).
pubmed: 26729103
pmcid: 4730293
doi: 10.3390/ijms17010048
Horn, S. et al. Mind bomb 1 is required for pancreatic β-cell formation. Proc. Natl Acad. Sci. USA 109, 7356–7361 (2012).
pubmed: 22529374
pmcid: 3358874
doi: 10.1073/pnas.1203605109
Potter, G. B., Facchinetti, F., Beaudoin, G. M. & Thompson, C. C. Neuronal expression of synaptotagmin-related gene 1 is regulated by thyroid hormone during cerebellar development. J. Neurosci. 21, 4373–4380 (2001).
pubmed: 11404423
pmcid: 6762732
doi: 10.1523/JNEUROSCI.21-12-04373.2001
Moghadam, P. K. & Jackson, M. B. The functional significance of synaptotagmin diversity in neuroendocrine secretion. Front Endocrinol. (Lausanne) 4, 124 (2013).
pubmed: 24065953
doi: 10.3389/fendo.2013.00124
Brown, B. C., Asian Genetic Epidemiology Network Type 2 Diabetes Consortium, Ye, C. J., Price, A. L. & Zaitlen, N. Transethnic genetic-correlation estimates from summary statistics. Am. J. Hum. Genet 99, 76–88 (2016).
pubmed: 27321947
pmcid: 5005434
doi: 10.1016/j.ajhg.2016.05.001
Galinsky, K. J. et al. Estimating cross-population genetic correlations of causal effect sizes. Genet. Epidemiol. 43, 180–188 (2019).
pubmed: 30474154
doi: 10.1002/gepi.22173
Yengo, L. et al. A saturated map of common genetic variants associated with human height. Nature 610, 704–712 (2022).
pubmed: 36224396
pmcid: 9605867
doi: 10.1038/s41586-022-05275-y
Hou, K. et al. Causal effects on complex traits are similar for common variants across segments of different continental ancestries within admixed individuals. Nat. Genet. 55, 549–558 (2023).
pubmed: 36941441
pmcid: 11120833
doi: 10.1038/s41588-023-01338-6
Ziyatdinov, A. et al. Genotyping, sequencing and analysis of 140,000 adults from the Mexico City Prospective Study. Nature 622, 784–793 (2023).
pubmed: 37821707
pmcid: 10600010
doi: 10.1038/s41586-023-06595-3
Fatumo, S. & Inouye, M. African genomes hold the key to accurate genetic risk prediction. Nat. Hum. Behav. 7, 295–296 (2023).
pubmed: 36823371
doi: 10.1038/s41562-023-01549-1
Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209 (2018).
pubmed: 30305743
pmcid: 6786975
doi: 10.1038/s41586-018-0579-z
Sudlow, C. et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 12, e1001779 (2015).
pubmed: 25826379
pmcid: 4380465
doi: 10.1371/journal.pmed.1001779
Szustakowski, J. D. et al. Advancing human genetics research and drug discovery through exome sequencing of the UK Biobank. Nat. Genet. 53, 942–948 (2021).
pubmed: 34183854
doi: 10.1038/s41588-021-00885-0
Cronin, R. M. et al. Development of the initial surveys for the All of Us Research Program. Epidemiology 30, 597–608 (2019).
pubmed: 31045611
pmcid: 6548672
doi: 10.1097/EDE.0000000000001028
Karlson, E. W., Boutin, N. T., Hoffnagle, A. G. & Allen, N. L. Building the Partners HealthCare Biobank at Partners Personalized Medicine: informed consent, return of research results, recruitment lessons and operational considerations. J. Pers. Med. 6, 2 (2016).
pubmed: 26784234
pmcid: 4810381
doi: 10.3390/jpm6010002
Boutin, N. T. et al. Implementation of electronic consent at a biobank: an opportunity for precision medicine research. J. Pers. Med. 6, 17 (2016).
pubmed: 27294961
pmcid: 4932464
doi: 10.3390/jpm6020017
Wu, P. et al. Mapping ICD-10 and ICD-10-CM Codes to Phecodes: Workflow Development and Initial Evaluation. JMIR Med. Inf. 7, e14325 (2019).
doi: 10.2196/14325
Liu, X., Wu, C., Li, C. & Boerwinkle, E. dbNSFP v3.0: a one-stop database of functional predictions and annotations for human nonsynonymous and splice-site SNVs. Hum. Mutat. 37, 235–241 (2016).
pubmed: 26555599
pmcid: 4752381
doi: 10.1002/humu.22932
McLaren, W. et al. The ensembl variant effect predictor. Genome Biol. 17, 122 (2016).
pubmed: 27268795
pmcid: 4893825
doi: 10.1186/s13059-016-0974-4
Gogarten, S. M. et al. Genetic association testing using the GENESIS R/Bioconductor package. Bioinformatics 35, 5346–5348 (2019).
pubmed: 31329242
pmcid: 7904076
doi: 10.1093/bioinformatics/btz567
Zhou, W. et al. Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies. Nat. Genet. 50, 1335–1341 (2018).
pubmed: 30104761
pmcid: 6119127
doi: 10.1038/s41588-018-0184-y
Heinze, G. A comparative investigation of methods for logistic regression with separated or nearly separated data. Stat. Med. 25, 4216–4226 (2006).
pubmed: 16955543
doi: 10.1002/sim.2687
Mbatchou, J. et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet. 53, 1097–1103 (2021).
pubmed: 34017140
doi: 10.1038/s41588-021-00870-7
Tang, Z. Z. & Lin, D. Y. MASS: meta-analysis of score statistics for sequencing studies. Bioinformatics 29, 1803–1805 (2013).
pubmed: 23698861
pmcid: 3702254
doi: 10.1093/bioinformatics/btt280
Zhao, Z. et al. UK Biobank whole-exome sequence binary phenome analysis with robust region-based rare-variant test. Am. J. Hum. Genet. 106, 3–12 (2020).
pubmed: 31866045
doi: 10.1016/j.ajhg.2019.11.012
Liu, Y. et al. ACAT: a fast and powerful p value combination method for rare-variant analysis in sequencing studies. Am. J. Hum. Genet. 104, 410–421 (2019).
pubmed: 30849328
pmcid: 6407498
doi: 10.1016/j.ajhg.2019.01.002
Muchinsky, P. M. The correction for attenuation.Educ. Psychol. Meas. 56, 63–75 (1996).
doi: 10.1177/0013164496056001004
Deming, W. E. Statistical Adjustment of Data (Wiley, 1943).