Multiple imputation for systematically missing confounders within a distributed data drug safety network: A simulation study and real-world example.
Cohort Studies
Computer Simulation
Confounding Factors, Epidemiologic
Data Interpretation, Statistical
Databases, Factual
Diabetes Mellitus, Type 2
/ drug therapy
Female
Humans
Hypoglycemic Agents
/ adverse effects
Male
Middle Aged
Myocardial Infarction
/ epidemiology
Pharmacoepidemiology
Retrospective Studies
United Kingdom
/ epidemiology
bias
cohort study
confounding
distributed data network
missing data
multiple imputation
pharmacoepidemiology
simulation study
Journal
Pharmacoepidemiology and drug safety
ISSN: 1099-1557
Titre abrégé: Pharmacoepidemiol Drug Saf
Pays: England
ID NLM: 9208369
Informations de publication
Date de publication:
01 2020
01 2020
Historique:
received:
03
08
2018
revised:
22
03
2019
accepted:
09
07
2019
pubmed:
6
9
2019
medline:
17
12
2020
entrez:
6
9
2019
Statut:
ppublish
Résumé
In distributed data networks, some data sites may be systematically missing important confounders that are captured by other sites in the network (eg, body mass index [BMI]). Multiple imputation may help repair bias in these scenarios. However, multiple imputation has not been described for distributed data networks where data access restrictions prevent centralized analysis. We conducted a simulation study and a real-world analysis using the UK's Clinical Practice Research Datalink to evaluate multiple imputation for confounders that are systematically missing from a subset of data sites in mock distributed data networks. The simulation study addressed univariate missing data, while the real-world analysis addressed multivariate missing data. Both studies were designed as retrospective cohort studies of the effect of current statin use on the risk of myocardial infarction among patients with newly treated type 2 diabetes. In our simulation study, multiple imputation repaired bias from missing BMI in all scenarios, with a median bias reduction of 118% in the default scenario. In our real-world study, the multiply imputed analysis (hazard ratio [HR]: 0.86; 95% confidence interval [CI], 0.69-1.08) was closer to the analysis that considered the true confounder values (HR: 0.85; 95% CI, 0.66-1.10) than the analysis that ignored them (HR: 0.93; 95% CI, 0.73-1.20). Multiple imputation adapted to distributed data settings is a feasible method to reduce bias from unmeasured but measurable confounders when at least one database contains the variables of interest. Further research is needed to evaluate its validity in real distributed data networks.
Substances chimiques
Hypoglycemic Agents
0
Types de publication
Journal Article
Research Support, Non-U.S. Gov't
Langues
eng
Sous-ensembles de citation
IM
Pagination
35-44Subventions
Organisme : CIHR
ID : DSE-146021
Pays : Canada
Informations de copyright
© 2019 John Wiley & Sons, Ltd.
Références
Suissa S, Henry D, Caetano P, et al. CNODES: the Canadian Network for Observational Drug Effect Studies. Open Med. 2012;6:134-140.
Riley RD, Lambert PC, Abo-Zaid G. Meta-analysis of individual participant data: rationale, conduct, and reporting. Bmj. 2010;340(feb05 1):c221.
Burgess S, White IR, Resche-Rigon M, Wood AM. Combining multiple imputation and meta-analysis with individual participant data. Statistics in medicine. 2013;32(26):4499-4514.
Jolani S, Debray TP, Koffijberg H, van Buuren S, Moons KG. Imputation of systematically missing predictors in an individual participant data meta-analysis: a generalized approach using MICE. Statistics in medicine. 2015;34(11):1841-1863.
Koopman L, van der Heijden GJ, Grobbee DE, Rovers MM. Comparison of methods of handling missing data in individual patient data meta-analyses: an empirical example on antibiotics in children with acute otitis media. American journal of epidemiology. 2008;167(5):540-545.
Quartagno M, Carpenter J. Multiple imputation for IPD meta-analysis: allowing for heterogeneity and studies with missing covariates. Statistics in medicine. 2016;35(17):2938-2954.
Resche-Rigon M, White IR. Multiple imputation by chained equations for systematically and sporadically missing multilevel data. Statistical methods in medical research. 2018;27(6):1634-1649.
Resche-Rigon M, White IR, Bartlett JW, Peters SAE, Thompson SG, on behalf of the PROG-IMT Study Group. Multiple imputation for handling systematically missing confounders in meta-analysis of individual participant data. Statistics in Medicine. 2013;32(28):4890-4905.
Van Buuren S, Brand JP, Groothuis-Oudshoorn CG, et al. Fully conditional specification in multivariate imputation. Journal of statistical computation and simulation. 2006;76(12):1049-1064.
White IR, Royston P, Wood AM. Multiple imputation using chained equations: issues and guidance for practice. Stat Med. 2011;30(4):377-399.
Taylor F, Huffman MD, Macedo AF, et al. Statins for the primary prevention of cardiovascular disease. The Cochrane Library. 2013;(1): CD004816. https://doi.org/10.1002/14651858.CD004816.pub5.
National Institute for Health and Care Excellence (UK). Cardiovascular disease: risk assessment and reduction, including lipid modification 2014.
Stone NJ, Robinson JG, Lichtenstein AH, et al. 2013 ACC/AHA guideline on the treatment of blood cholesterol to reduce atherosclerotic cardiovascular risk in adults. Circulation. 2014;129(25 suppl 2):S1-S45.
Yusuf S, Hawken S, Ôunpuu S, et al. Effect of potentially modifiable risk factors associated with myocardial infarction in 52 countries (the INTERHEART study): case-control study. Lancet. 2004;364(9438):937-952.
Stratton IM, Adler AI, Neil HAW, et al. Association of glycaemia with macrovascular and microvascular complications of type 2 diabetes (UKPDS 35): prospective observational study. BMJ. 2000;321(7258):405-412.
UK Prospective Diabetes Study Group. Tight blood pressure control and risk of macrovascular and microvascular complications in type 2 diabetes: UKPDS 38. BMJ. 1998;317(7160):703-713.
Little RJ, Rubin DB. Statistical analysis with missing data. John Wiley & Sons; 2014.
MacCallum RC, Zhang S, Preacher KJ, Rucker DD. On the practice of dichotomization of quantitative variables. Psychological methods. 2002;7(1):19-40.
Austin PC. Generating survival times to simulate Cox proportional hazards models with time-varying covariates. Statistics in medicine. 2012;31(29):3946-3958.
Herrett E, Gallagher AM, Bhaskaran K, et al. Data resource profile: Clinical Practice Research Datalink (CPRD). Int J Epidemiol. 2015;44(3):827-836.
Hughes RA, White IR, Seaman SR, Carpenter JR, Tilling K, Sterne JAC. Joint modelling rationale for chained equations. BMC medical research methodology. 2014;14(1):28.
Bhatnagar S, Turgeon M, Saarela O, et al. casebase: fitting flexible smooth-in-time hazards and risk functions via logistic and multinomial regression. R package version 0.1.0 ed: CRAN; 2017.