Machine Learning Electronic Health Record Identification of Patients with Rheumatoid Arthritis: Algorithm Pipeline Development and Validation Study.

Electronic Health Records Gradient Boosting Natural Language Processing Rheumatoid Arthritis Supervised machine learning Support Vector Machine

Journal

JMIR medical informatics
ISSN: 2291-9694
Titre abrégé: JMIR Med Inform
Pays: Canada
ID NLM: 101645109

Informations de publication

Date de publication:
30 Nov 2020
Historique:
received: 28 08 2020
accepted: 24 10 2020
revised: 18 10 2020
entrez: 30 11 2020
pubmed: 1 12 2020
medline: 1 12 2020
Statut: epublish

Résumé

Financial codes are often used to extract diagnoses from electronic health records. This approach is prone to false positives. Alternatively, queries are constructed, but these are highly center and language specific. A tantalizing alternative is the automatic identification of patients by employing machine learning on format-free text entries. The aim of this study was to develop an easily implementable workflow that builds a machine learning algorithm capable of accurately identifying patients with rheumatoid arthritis from format-free text fields in electronic health records. Two electronic health record data sets were employed: Leiden (n=3000) and Erlangen (n=4771). Using a portion of the Leiden data (n=2000), we compared 6 different machine learning methods and a naïve word-matching algorithm using 10-fold cross-validation. Performances were compared using the area under the receiver operating characteristic curve (AUROC) and the area under the precision recall curve (AUPRC), and F1 score was used as the primary criterion for selecting the best method to build a classifying algorithm. We selected the optimal threshold of positive predictive value for case identification based on the output of the best method in the training data. This validation workflow was subsequently applied to a portion of the Erlangen data (n=4293). For testing, the best performing methods were applied to remaining data (Leiden n=1000; Erlangen n=478) for an unbiased evaluation. For the Leiden data set, the word-matching algorithm demonstrated mixed performance (AUROC 0.90; AUPRC 0.33; F1 score 0.55), and 4 methods significantly outperformed word-matching, with support vector machines performing best (AUROC 0.98; AUPRC 0.88; F1 score 0.83). Applying this support vector machine classifier to the test data resulted in a similarly high performance (F1 score 0.81; positive predictive value [PPV] 0.94), and with this method, we could identify 2873 patients with rheumatoid arthritis in less than 7 seconds out of the complete collection of 23,300 patients in the Leiden electronic health record system. For the Erlangen data set, gradient boosting performed best (AUROC 0.94; AUPRC 0.85; F1 score 0.82) in the training set, and applied to the test data, resulted once again in good results (F1 score 0.67; PPV 0.97). We demonstrate that machine learning methods can extract the records of patients with rheumatoid arthritis from electronic health record data with high precision, allowing research on very large populations for limited costs. Our approach is language and center independent and could be applied to any type of diagnosis. We have developed our pipeline into a universally applicable and easy-to-implement workflow to equip centers with their own high-performing algorithm. This allows the creation of observational studies of unprecedented size covering different countries for low cost from already available data in electronic health record systems.

Sections du résumé

BACKGROUND BACKGROUND
Financial codes are often used to extract diagnoses from electronic health records. This approach is prone to false positives. Alternatively, queries are constructed, but these are highly center and language specific. A tantalizing alternative is the automatic identification of patients by employing machine learning on format-free text entries.
OBJECTIVE OBJECTIVE
The aim of this study was to develop an easily implementable workflow that builds a machine learning algorithm capable of accurately identifying patients with rheumatoid arthritis from format-free text fields in electronic health records.
METHODS METHODS
Two electronic health record data sets were employed: Leiden (n=3000) and Erlangen (n=4771). Using a portion of the Leiden data (n=2000), we compared 6 different machine learning methods and a naïve word-matching algorithm using 10-fold cross-validation. Performances were compared using the area under the receiver operating characteristic curve (AUROC) and the area under the precision recall curve (AUPRC), and F1 score was used as the primary criterion for selecting the best method to build a classifying algorithm. We selected the optimal threshold of positive predictive value for case identification based on the output of the best method in the training data. This validation workflow was subsequently applied to a portion of the Erlangen data (n=4293). For testing, the best performing methods were applied to remaining data (Leiden n=1000; Erlangen n=478) for an unbiased evaluation.
RESULTS RESULTS
For the Leiden data set, the word-matching algorithm demonstrated mixed performance (AUROC 0.90; AUPRC 0.33; F1 score 0.55), and 4 methods significantly outperformed word-matching, with support vector machines performing best (AUROC 0.98; AUPRC 0.88; F1 score 0.83). Applying this support vector machine classifier to the test data resulted in a similarly high performance (F1 score 0.81; positive predictive value [PPV] 0.94), and with this method, we could identify 2873 patients with rheumatoid arthritis in less than 7 seconds out of the complete collection of 23,300 patients in the Leiden electronic health record system. For the Erlangen data set, gradient boosting performed best (AUROC 0.94; AUPRC 0.85; F1 score 0.82) in the training set, and applied to the test data, resulted once again in good results (F1 score 0.67; PPV 0.97).
CONCLUSIONS CONCLUSIONS
We demonstrate that machine learning methods can extract the records of patients with rheumatoid arthritis from electronic health record data with high precision, allowing research on very large populations for limited costs. Our approach is language and center independent and could be applied to any type of diagnosis. We have developed our pipeline into a universally applicable and easy-to-implement workflow to equip centers with their own high-performing algorithm. This allows the creation of observational studies of unprecedented size covering different countries for low cost from already available data in electronic health record systems.

Identifiants

pubmed: 33252349
pii: v8i11e23930
doi: 10.2196/23930
pmc: PMC7735897
doi:

Types de publication

Journal Article

Langues

eng

Pagination

e23930

Informations de copyright

©Tjardo D Maarseveen, Timo Meinderink, Marcel J T Reinders, Johannes Knitza, Tom W J Huizinga, Arnd Kleyer, David Simon, Erik B van den Akker, Rachel Knevel. Originally published in JMIR Medical Informatics (http://medinform.jmir.org), 30.11.2020.

Références

Arthritis Care Res (Hoboken). 2010 Aug;62(8):1120-7
pubmed: 20235204
Sci Data. 2016 Mar 15;3:160018
pubmed: 26978244
Arthritis Rheum. 1988 Mar;31(3):315-24
pubmed: 3358796
J Am Med Inform Assoc. 2016 Nov;23(6):1046-1052
pubmed: 27026615
Clin Exp Rheumatol. 2014 Sep-Oct;32(5 Suppl 85):S-135-40
pubmed: 25365103
PLoS One. 2015 Mar 04;10(3):e0118432
pubmed: 25738806
IEEE Trans Pattern Anal Mach Intell. 2010 Mar;32(3):569-75
pubmed: 20075479
Neural Comput. 1998 Sep 15;10(7):1895-1923
pubmed: 9744903
Arthritis Res Ther. 2019 Dec 30;21(1):305
pubmed: 31888720
Arthritis Rheum. 2010 Sep;62(9):2569-81
pubmed: 20872595

Auteurs

Tjardo D Maarseveen (TD)

Department of Rheumatology, Leiden University Medical Center, Leiden, Netherlands.

Timo Meinderink (T)

Department of Internal Medicine 3, Friedrich-Alexander University Erlangen-Nuremberg, Erlangen, Germany.
Deutsches Zentrum für Immuntherapie, Erlangen-Nuremberg and Universitätsklinikum, Erlangen, Germany.

Marcel J T Reinders (MJT)

Leiden Computational Biology Centre, Leiden University Medical Center, Leiden, Netherlands.
Molecular Epidemiology, Leiden University Medical Center, Leiden, Netherlands.

Johannes Knitza (J)

Department of Internal Medicine 3, Friedrich-Alexander University Erlangen-Nuremberg, Erlangen, Germany.
Deutsches Zentrum für Immuntherapie, Erlangen-Nuremberg and Universitätsklinikum, Erlangen, Germany.

Tom W J Huizinga (TWJ)

Department of Rheumatology, Leiden University Medical Center, Leiden, Netherlands.

Arnd Kleyer (A)

Department of Internal Medicine 3, Friedrich-Alexander University Erlangen-Nuremberg, Erlangen, Germany.
Deutsches Zentrum für Immuntherapie, Erlangen-Nuremberg and Universitätsklinikum, Erlangen, Germany.

David Simon (D)

Department of Internal Medicine 3, Friedrich-Alexander University Erlangen-Nuremberg, Erlangen, Germany.
Deutsches Zentrum für Immuntherapie, Erlangen-Nuremberg and Universitätsklinikum, Erlangen, Germany.

Erik B van den Akker (EB)

Leiden Computational Biology Centre, Leiden University Medical Center, Leiden, Netherlands.
Molecular Epidemiology, Leiden University Medical Center, Leiden, Netherlands.

Rachel Knevel (R)

Department of Rheumatology, Leiden University Medical Center, Leiden, Netherlands.
Division of Rheumatology, Inflammation and Immunity, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, United States.

Classifications MeSH