Machine Learning Did Not Outperform Conventional Competing Risk Modeling to Predict Revision Arthroplasty.
Journal
Clinical orthopaedics and related research
ISSN: 1528-1132
Titre abrégé: Clin Orthop Relat Res
Pays: United States
ID NLM: 0075674
Informations de publication
Date de publication:
12 Mar 2024
12 Mar 2024
Historique:
received:
11
07
2023
accepted:
01
02
2024
medline:
12
3
2024
pubmed:
12
3
2024
entrez:
12
3
2024
Statut:
aheadofprint
Résumé
Estimating the risk of revision after arthroplasty could inform patient and surgeon decision-making. However, there is a lack of well-performing prediction models assisting in this task, which may be due to current conventional modeling approaches such as traditional survivorship estimators (such as Kaplan-Meier) or competing risk estimators. Recent advances in machine learning survival analysis might improve decision support tools in this setting. Therefore, this study aimed to assess the performance of machine learning compared with that of conventional modeling to predict revision after arthroplasty. Does machine learning perform better than traditional regression models for estimating the risk of revision for patients undergoing hip or knee arthroplasty? Eleven datasets from published studies from the Dutch Arthroplasty Register reporting on factors associated with revision or survival after partial or total knee and hip arthroplasty between 2018 and 2022 were included in our study. The 11 datasets were observational registry studies, with a sample size ranging from 3038 to 218,214 procedures. We developed a set of time-to-event models for each dataset, leading to 11 comparisons. A set of predictors (factors associated with revision surgery) was identified based on the variables that were selected in the included studies. We assessed the predictive performance of two state-of-the-art statistical time-to-event models for 1-, 2-, and 3-year follow-up: a Fine and Gray model (which models the cumulative incidence of revision) and a cause-specific Cox model (which models the hazard of revision). These were compared with a machine-learning approach (a random survival forest model, which is a decision tree-based machine-learning algorithm for time-to-event analysis). Performance was assessed according to discriminative ability (time-dependent area under the receiver operating curve), calibration (slope and intercept), and overall prediction error (scaled Brier score). Discrimination, known as the area under the receiver operating characteristic curve, measures the model's ability to distinguish patients who achieved the outcomes from those who did not and ranges from 0.5 to 1.0, with 1.0 indicating the highest discrimination score and 0.50 the lowest. Calibration plots the predicted versus the observed probabilities; a perfect plot has an intercept of 0 and a slope of 1. The Brier score calculates a composite of discrimination and calibration, with 0 indicating perfect prediction and 1 the poorest. A scaled version of the Brier score, 1 - (model Brier score/null model Brier score), can be interpreted as the amount of overall prediction error. Using machine learning survivorship analysis, we found no differences between the competing risks estimator and traditional regression models for patients undergoing arthroplasty in terms of discriminative ability (patients who received a revision compared with those who did not). We found no consistent differences between the validated performance (time-dependent area under the receiver operating characteristic curve) of different modeling approaches because these values ranged between -0.04 and 0.03 across the 11 datasets (the time-dependent area under the receiver operating characteristic curve of the models across 11 datasets ranged between 0.52 to 0.68). In addition, the calibration metrics and scaled Brier scores produced comparable estimates, showing no advantage of machine learning over traditional regression models. Machine learning did not outperform traditional regression models. Neither machine learning modeling nor traditional regression methods were sufficiently accurate in order to offer prognostic information when predicting revision arthroplasty. The benefit of these modeling approaches may be limited in this context.
Sections du résumé
BACKGROUND
BACKGROUND
Estimating the risk of revision after arthroplasty could inform patient and surgeon decision-making. However, there is a lack of well-performing prediction models assisting in this task, which may be due to current conventional modeling approaches such as traditional survivorship estimators (such as Kaplan-Meier) or competing risk estimators. Recent advances in machine learning survival analysis might improve decision support tools in this setting. Therefore, this study aimed to assess the performance of machine learning compared with that of conventional modeling to predict revision after arthroplasty.
QUESTION/PURPOSE
OBJECTIVE
Does machine learning perform better than traditional regression models for estimating the risk of revision for patients undergoing hip or knee arthroplasty?
METHODS
METHODS
Eleven datasets from published studies from the Dutch Arthroplasty Register reporting on factors associated with revision or survival after partial or total knee and hip arthroplasty between 2018 and 2022 were included in our study. The 11 datasets were observational registry studies, with a sample size ranging from 3038 to 218,214 procedures. We developed a set of time-to-event models for each dataset, leading to 11 comparisons. A set of predictors (factors associated with revision surgery) was identified based on the variables that were selected in the included studies. We assessed the predictive performance of two state-of-the-art statistical time-to-event models for 1-, 2-, and 3-year follow-up: a Fine and Gray model (which models the cumulative incidence of revision) and a cause-specific Cox model (which models the hazard of revision). These were compared with a machine-learning approach (a random survival forest model, which is a decision tree-based machine-learning algorithm for time-to-event analysis). Performance was assessed according to discriminative ability (time-dependent area under the receiver operating curve), calibration (slope and intercept), and overall prediction error (scaled Brier score). Discrimination, known as the area under the receiver operating characteristic curve, measures the model's ability to distinguish patients who achieved the outcomes from those who did not and ranges from 0.5 to 1.0, with 1.0 indicating the highest discrimination score and 0.50 the lowest. Calibration plots the predicted versus the observed probabilities; a perfect plot has an intercept of 0 and a slope of 1. The Brier score calculates a composite of discrimination and calibration, with 0 indicating perfect prediction and 1 the poorest. A scaled version of the Brier score, 1 - (model Brier score/null model Brier score), can be interpreted as the amount of overall prediction error.
RESULTS
RESULTS
Using machine learning survivorship analysis, we found no differences between the competing risks estimator and traditional regression models for patients undergoing arthroplasty in terms of discriminative ability (patients who received a revision compared with those who did not). We found no consistent differences between the validated performance (time-dependent area under the receiver operating characteristic curve) of different modeling approaches because these values ranged between -0.04 and 0.03 across the 11 datasets (the time-dependent area under the receiver operating characteristic curve of the models across 11 datasets ranged between 0.52 to 0.68). In addition, the calibration metrics and scaled Brier scores produced comparable estimates, showing no advantage of machine learning over traditional regression models.
CONCLUSION
CONCLUSIONS
Machine learning did not outperform traditional regression models.
CLINICAL RELEVANCE
CONCLUSIONS
Neither machine learning modeling nor traditional regression methods were sufficiently accurate in order to offer prognostic information when predicting revision arthroplasty. The benefit of these modeling approaches may be limited in this context.
Identifiants
pubmed: 38470976
doi: 10.1097/CORR.0000000000003018
pii: 00003086-990000000-01528
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Informations de copyright
Copyright © 2024 by the Association of Bone and Joint Surgeons.
Déclaration de conflit d'intérêts
All ICMJE Conflict of Interest Forms for authors and Clinical Orthopaedics and Related Research® editors and board members are on file with the publication and can be viewed on request.
Références
Aalen OO, Johansen S. An empirical transition matrix for non-homogeneous Markov chains based on censored observations. Scand J Stat. 1978;5:141-150.
Aram P, Trela-Larsen L, Sayers A, et al. Estimating an individual’s probability of revision surgery after knee replacement: a comparison of modeling approaches using a national data set. Am J Epidemiol. 2018;187:2252-2262.
Austin PC, Steyerberg EW, Putter H. Fine-Gray subdistribution hazard models to simultaneously estimate the absolute risk of different event types: cumulative total failure probability may exceed 1. Stat Med. 2021;40:4200-4212.
Bloemheuvel EM, van Steenbergen LN, Swierstra BA. Dual mobility cups in primary total hip arthroplasties: trend over time in use, patient characteristics, and mid-term revision in 3,038 cases in the Dutch Arthroplasty Register (2007-2016). Acta Orthop. 2019;90:11-14.
Bloemheuvel EM, van Steenbergen LN, Swierstra BA. Lower 5-year cup re-revision rate for dual mobility cups compared with unipolar cups: report of 15,922 cup revision cases in the Dutch Arthroplasty Register (2007-2016). Acta Orthop. 2019;90:338-341.
Burger JA, Kleeblad LJ, Sierevelt IN, et al. A comprehensive evaluation of lateral unicompartmental knee arthroplasty short to mid-term survivorship, and the effect of patient and implant characteristics: an analysis of data from the Dutch Arthroplasty Register. J Arthroplasty. 2020;35:1813-1818.
Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD Statement. BMC Med. 2015;13:1-10.
Cox DR. Two further applications of a model for binary regression. Biometrika. 1958;45:562-565
Dutch Arthroplasty Register. LROI Report 2020. Available at: https://www.lroi-report.nl/app/uploads/2022/04/PDF-LROI-annual-report-2021.pdf. Accessed July 4, 2022.
Fine JP, Gray RJ. A proportional hazards model for the subdistribution of a competing risk. J Am Stat Assoc. 1999;94:496-509.
Ishwaran H, Gerds TA, Kogalur UB, Moore RD, Gange SJ, Lau BM. Random survival forests for competing risks. Biostatistics. 2014;15:757-773.
Janssen L, Wijnands KAP, Janssen D, Janssen MWHE, Morrenhof JW. Do stem design and surgical approach influence early aseptic loosening in cementless THA? Clin Orthop Relat Res. 2018;476:1212-1220.
Keurentjes JC, Fiocco M, Schreurs BW, Pijls BG, Nouta KA, Nelissen RGHH. Revision surgery is overestimated in hip replacement. Bone Joint Res. 2012;1:258-262.
Kuijpers MFL, Hannink G, van Steenbergen LN, Schreurs BW. Outcome of revision hip arthroplasty in patients younger than 55 years: an analysis of 1,037 revisions in the Dutch Arthroplasty Register. Acta Orthop. 2020;91:165-170.
Kuijpers MFL, Hannink G, Vehmeijer SBW, van Steenbergen LN, Schreurs BW. The risk of revision after total hip arthroplasty in young patients depends on surgical approach, femoral head size and bearing type; an analysis of 19,682 operations in the Dutch arthroplasty register. BMC Musculoskelet Disord. 2019;20:385.
Labek G, Thaler M, Janda W, Agreiter M, Stöckl B. Revision rates after total joint replacement: cumulative results from worldwide joint register datasets. J Bone Joint Surg Br. 2011;93:293-297.
Luo W, Phung D, Tran T, et al. Guidelines for developing and reporting machine learning predictive models in biomedical research: a multidisciplinary view. J Med Internet Res. 2016;18:e323.
Martin RK, Wastvedt S, Lange J, Pareek A, Wolfson J, Lund B. Limited clinical utility of a machine learning revision prediction model based on a national hip arthroscopy registry. Knee Surg Sports Traumatol Arthrosc. 2023;31:2079-2089
Moerman S, Mathijssen NMC, Tuinebreijer WE, Vochteloo AJH, Nelissen RGHH. Hemiarthroplasty and total hip arthroplasty in 30,830 patients with hip fractures: data from the Dutch Arthroplasty Register on revision and risk factors for revision. Acta Orthop. 2018;89:509-514.
Moncada-Torres A, van Maaren MC, Hendriks MP, Siesling S, Geleijnse G. Explainable machine learning can outperform Cox regression predictions and provide insights in breast cancer survival. Sci Rep. 2021;11:6968.
Oosterhoff JHF, Gravesteijn BY, Karhade AV, et al. Feasibility of machine learning and logistic regression algorithms to predict outcome in orthopaedic trauma surgery. J Bone Joint Surg Am. 2022;104:544-551.
Peters RM, van Steenbergen LN, Bulstra SK, et al. Nationwide review of mixed and non-mixed components from different manufacturers in total hip arthroplasty. Acta Orthop. 2016;87:356-362.
Peters RM, van Steenbergen LN, Stevens M, Rijk PC, Bulstra SK, Zijlstra WP. The effect of bearing type on the outcome of total hip arthroplasty. Acta Orthop. 2018;89:163-169.
Peters RM, van Steenbergen LN, Stewart RE, et al. Patient characteristics influence revision rate of total hip arthroplasty: American Society of Anesthesiologists score and body mass index were the strongest predictors for short-term revision after primary total hip arthroplasty. J Arthroplasty. 2020;35:188-192.e2.
Pickett KL, Suresh K, Campbell KR, Davis S, Juarez-Colunga E. Random survival forests for dynamic predictions of a time-to-event outcome using a longitudinal biomarker. BMC Med. Res Methodol. 2021;21:216.
Putter H, Fiocco M, Geskus RB. Tutorial in biostatistics: competing risks and multi-state models. Stat Med. 2007;26:2389–2430.
Rubin D. Multiple Imputation for nonresponse in surveys. John Wiley & Sons Inc; 1987.
Sorel JC, Veltman ES, Honig A, Poolman RW. The influence of preoperative psychological distress on pain and function after total knee arthroplasty: a systematic review and meta-analysis. Bone Joint J. 2019;101:7-14.
Spekenbrink-Spooren A, Van Steenbergen LN, Denissen GAW, Swierstra BA, Poolman RW, Nelissen RGHH. Higher mid-term revision rates of posterior stabilized compared with cruciate retaining total knee arthroplasties: 133,841 cemented arthroplasties for osteoarthritis in the Netherlands in 2007-2016. Acta Orthop. 2018;89:640-645.
Steyerberg EW, Vergouwe Y. Towards better clinical prediction models: seven steps for development and an ABCD for validation. Eur. Heart J. 2014;35:1925-1931.
Steyerberg EW, Vickers AJ, Cook NR, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21:128-138.
van Buuren S, Groothuis-Oudshoorn K. Mice: multivariate imputation by chained equations in R. J Stat Software Artic. 2011;45:1-67.
van Calster B, Vickers AJ. Calibration of risk prediction models: impact on decision-analytic performance. Med Decis Making. 2015;35:162-169.
van den Goorbergh R, van Smeden M, Timmerman D, Van Calster B. The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression. J Am Med Informatics Assoc. 2022;29:1525-1534.
van der Pas S, Nelissen R, Fiocco M. Different competing risks models for different questions may give similar results in arthroplasty registers in the presence of few events. Acta Orthop. 2018;89:145-151.
van Geloven N, Giardiello D, Bonneville EF, et al. Validation of prediction models in the presence of competing risks: a guide through modern methods. BMJ. 2022;377:e069249.
van Oost I, Koenraadt KLM, van Steenbergen LN, Bolder SBT, van Geenen RCI. Higher risk of revision for partial knee replacements in low absolute volume hospitals: data from 18,134 partial knee replacements in the Dutch Arthroplasty Register. Acta Orthop. 2020;91:426-432.
van Steenbergen L, Denissen G, Schreurs B, Zijlstra W, Koot H, Nelissen R. Dutch advice not to use large head metal-on-metal hip arthroplasties justifiable – results from the Dutch Arthroplasty Register. Ned Tijdschr voor Orthop. 2020;27:4-11.
Zijlstra WP, De Hartog B, Van Steenbergen LN, Scheurs BW, Nelissen RGHH. Effect of femoral head size and surgical approach on risk of revision for dislocation after total hip arthroplasty. Acta Orthop. 2017;88:395-401.