Information Extraction from Echocardiography Reports for a Clinical Follow-up Study-Comparison of Extracted Variables Intended for General Use in a Data Warehouse with Those Intended Specifically for the Study.

Data Warehousing Echocardiography Follow-Up Studies Hospital Information Systems Humans Information Storage and Retrieval

Journal

Methods of information in medicine

ISSN: 2511-705X

Titre abrégé: Methods Inf Med

Pays: Germany

ID NLM: 0210453

Informations de publication

Date de publication:
Nov 2019

Historique:

pubmed: 31 1 2020

medline: 8 8 2020

entrez: 31 1 2020

Statut: ppublish

Résumé

The interest in information extraction from clinical reports for secondary data use is increasing. But experience with the productive use of information extraction processes over time is scarce. A clinical data warehouse has been in use at our university hospital for several years, which also provides an information extraction of echocardiography reports developed for general use. This study aims to illustrate the difficulties encountered, while using data from a preexisting information extraction process for a large clinical study. To compare the data from the preexisting process with the data obtained from a specially developed process designed to improve the quality and completeness of the study data. We extracted the echocardiography variables for 440 patients from the general-use information extraction of the data warehouse (678 reports). Then we developed an information extraction process for the same variables but specifically for this study, with the aim to extract as much information as possible from the text. The extracted data of both processes were compared with a newly created gold standard defined by a cardiologist with long-standing experience in heart failure. Among 57 echocardiography variables considered relevant for the study, 50 were documented in the routine text reports and could be extracted. Twenty of the required variables were not provided by the general-use extraction process, some others were not provided correctly. The median macro F1-score (precision, recall) across the 30 variables for which values were extracted was 0.81 (0.94, 0.77). Across all 50 variables, as relevant for the study, median macro F1-score was only 0.49 (0.56, 0.46). Employing the study-specific approach considerably improved the quality and completeness of the variables, resulting in F1-scores of 0.97 (0.98, 0.96) across all variables. Data from information extractions can be used for large clinical studies. However, preexisting information extraction processes should be treated with caution, as the time and effort spent defining each variable in the information extraction process may not be clear.

Sections du résumé

BACKGROUND BACKGROUND

OBJECTIVES OBJECTIVE

This study aims to illustrate the difficulties encountered, while using data from a preexisting information extraction process for a large clinical study. To compare the data from the preexisting process with the data obtained from a specially developed process designed to improve the quality and completeness of the study data.

METHODS METHODS

We extracted the echocardiography variables for 440 patients from the general-use information extraction of the data warehouse (678 reports). Then we developed an information extraction process for the same variables but specifically for this study, with the aim to extract as much information as possible from the text. The extracted data of both processes were compared with a newly created gold standard defined by a cardiologist with long-standing experience in heart failure.

RESULTS RESULTS

Among 57 echocardiography variables considered relevant for the study, 50 were documented in the routine text reports and could be extracted. Twenty of the required variables were not provided by the general-use extraction process, some others were not provided correctly. The median macro F1-score (precision, recall) across the 30 variables for which values were extracted was 0.81 (0.94, 0.77). Across all 50 variables, as relevant for the study, median macro F1-score was only 0.49 (0.56, 0.46). Employing the study-specific approach considerably improved the quality and completeness of the variables, resulting in F1-scores of 0.97 (0.98, 0.96) across all variables.

CONCLUSION CONCLUSIONS

Data from information extractions can be used for large clinical studies. However, preexisting information extraction processes should be treated with caution, as the time and effort spent defining each variable in the information extraction process may not be clear.

Identifiants

DOI: 10.1055/s-0039-3402069 PMID: 32000268

pubmed: 32000268

doi: 10.1055/s-0039-3402069

doi:

Types de publication

Comparative Study Journal Article

Langues

eng

Sous-ensembles de citation

Pagination

140-150

Subventions

Organisme : #01EO1004

ID : German Ministry of Education and Research (BMBF), Berlin

Organisme : #01EO1504

ID : German Ministry of Education and Research (BMBF), Berlin

Informations de copyright

Georg Thieme Verlag KG Stuttgart · New York.

Déclaration de conflit d'intérêts

None declared.

Information Extraction from Echocardiography Reports for a Clinical Follow-up Study-Comparison of Extracted Variables Intended for General Use in a Data Warehouse with Those Intended Specifically for the Study.

Journal

Informations de publication

Résumé

Sections du résumé

Identifiants

Types de publication

Langues

Sous-ensembles de citation

Pagination

Subventions

Informations de copyright

Déclaration de conflit d'intérêts

Auteurs

Mathias Kaspar (M)

Caroline Morbach (C)

Georg Fette (G)

Maximilian Ertl (M)

Lea K Seidlmayer (LK)

Jonathan Krebs (J)

Georg Dietrich (G)

Leon Liman (L)

Frank Puppe (F)

Stefan Störk (S)

Articles similaires

[Redispensing of expensive oral anticancer medicines: a practical application].

Smoking Cessation and Incident Cardiovascular Disease.

Evaluation of Low-Value Services Across Major Medicare Advantage Insurers and Traditional Medicare.

Effectiveness of Virtual Yoga for Chronic Low Back Pain: A Randomized Clinical Trial.

Classifications MeSH