The virtual reference radiologist: comprehensive AI assistance for clinical image reading and interpretation.

Artificial intelligence Diagnostic errors Diagnostic imaging Radiology

Journal

European radiology
ISSN: 1432-1084
Titre abrégé: Eur Radiol
Pays: Germany
ID NLM: 9114774

Informations de publication

Date de publication:
16 Apr 2024
Historique:
received: 12 10 2023
accepted: 08 03 2024
revised: 27 02 2024
medline: 17 4 2024
pubmed: 17 4 2024
entrez: 16 4 2024
Statut: aheadofprint

Résumé

Large language models (LLMs) have shown potential in radiology, but their ability to aid radiologists in interpreting imaging studies remains unexplored. We investigated the effects of a state-of-the-art LLM (GPT-4) on the radiologists' diagnostic workflow. In this retrospective study, six radiologists of different experience levels read 40 selected radiographic [n = 10], CT [n = 10], MRI [n = 10], and angiographic [n = 10] studies unassisted (session one) and assisted by GPT-4 (session two). Each imaging study was presented with demographic data, the chief complaint, and associated symptoms, and diagnoses were registered using an online survey tool. The impact of Artificial Intelligence (AI) on diagnostic accuracy, confidence, user experience, input prompts, and generated responses was assessed. False information was registered. Linear mixed-effect models were used to quantify the factors (fixed: experience, modality, AI assistance; random: radiologist) influencing diagnostic accuracy and confidence. When assessing if the correct diagnosis was among the top-3 differential diagnoses, diagnostic accuracy improved slightly from 181/240 (75.4%, unassisted) to 188/240 (78.3%, AI-assisted). Similar improvements were found when only the top differential diagnosis was considered. AI assistance was used in 77.5% of the readings. Three hundred nine prompts were generated, primarily involving differential diagnoses (59.1%) and imaging features of specific conditions (27.5%). Diagnostic confidence was significantly higher when readings were AI-assisted (p > 0.001). Twenty-three responses (7.4%) were classified as hallucinations, while two (0.6%) were misinterpretations. Integrating GPT-4 in the diagnostic process improved diagnostic accuracy slightly and diagnostic confidence significantly. Potentially harmful hallucinations and misinterpretations call for caution and highlight the need for further safeguarding measures. Using GPT-4 as a virtual assistant when reading images made six radiologists of different experience levels feel more confident and provide more accurate diagnoses; yet, GPT-4 gave factually incorrect and potentially harmful information in 7.4% of its responses.

Identifiants

pubmed: 38627289
doi: 10.1007/s00330-024-10727-2
pii: 10.1007/s00330-024-10727-2
doi:

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Subventions

Organisme : Horizon 2020 Framework Programme
ID : 101057091
Organisme : Horizon 2020 Framework Programme
ID : 101057091
Organisme : Deutsche Forschungsgemeinschaft
ID : TR 1700/7-1
Organisme : BMBF
ID : 01KD2215A, 031L0312A

Informations de copyright

© 2024. The Author(s).

Références

Nav N (2023) 97+ ChatGPT Statistics & User Numbers in May 2023 (New Data). Available via https://nerdynav.com/chatgpt-statistics/ . Accessed 25 May 2023
De Angelis L, Baglivo F, Arzilli G et al (2023) ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health. Front Public Health 11:1567
doi: 10.3389/fpubh.2023.1166120
Elkassem AA, Smith AD (2023) Potential use cases for ChatGPT in radiology reporting. AJR Am J Roentgenol 221:373–376
Adams LC, Truhn D, Busch F et al (2023) Leveraging GPT-4 for post hoc transformation of free-text radiology reports into structured reporting: a multilingual feasibility study. Radiology 307:e230725
Rao A, Kim J, Kamineni M, Pang M, Lie W, Succi MD (2023) Evaluating ChatGPT as an adjunct for radiologic decision-making. medRxiv:2023.2002. 2002.23285399
Nori H, King N, McKinney SM, Carignan D, Horvitz E (2023) Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:230313375
Bhayana R, Krishna S, Bleakney RR (2023) Performance of ChatGPT on a radiology board-style examination: Insights into current strengths and limitations. Radiology 307:e230582
Bajaj S, Gandhi D, Nayar D (2023) Potential applications and impact of ChatGPT in radiology. Acad Radiol S1076-6332(23)00460-9. https://doi.org/10.1016/j.acra.2023.08.039
Akinci D’Antonoli T, Stanzione A, Bluethgen C et al (2023) Large language models in radiology: fundamentals, applications, ethical considerations, risks, and future directions. Diagn Interv Radiol 30:80–90
Bera K, O’Connor G, Jiang S, Tirumani SH, Ramaiya N (2023) Analysis of ChatGPT publications in radiology: literature so far. Curr Probl Diagn Radiol 53:215–225
Suthar PP, Kounsal A, Chhetri L, Saini D, Dua SG (2023) Artificial Intelligence (AI) in Radiology: A Deep Dive Into ChatGPT 4.0’s Accuracy with the American Journal of Neuroradiology’s (AJNR) “Case of the Month”. Cureus 15(8):e43958
Ueda D, Mitsuyama Y, Takita H et al (2023) Diagnostic Performance of ChatGPT from Patient History and Imaging Findings on the Diagnosis Please Quizzes. Radiology 308:e231040
doi: 10.1148/radiol.231040 pubmed: 37462501
Rau A, Rau S, Zoeller D et al (2023) A Context-based Chatbot Surpasses Trained Radiologists and Generic ChatGPT in Following the ACR Appropriateness Guidelines. Radiology 308:e230970
doi: 10.1148/radiol.230970 pubmed: 37489981
Finck T, Moosbauer J, Probst M et al (2022) Faster and Better: How Anomaly Detection Can Accelerate and Improve Reporting of Head Computed Tomography. Diagnostics 12:452
doi: 10.3390/diagnostics12020452 pubmed: 35204543 pmcid: 8871235
Faul F, Erdfelder E, Lang AG, Buchner A (2007) G*Power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav Res Methods 39:175–191
doi: 10.3758/BF03193146 pubmed: 17695343
Kung TH, Cheatham M, Medenilla A et al (2023) Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLoS Digital Health 2:e0000198
doi: 10.1371/journal.pdig.0000198 pubmed: 36812645 pmcid: 9931230
Kanjee Z, Crowe B, Rodman A (2023) Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA 330:78–80
Dratsch T, Chen X, Rezazade Mehrizi M et al (2023) Automation bias in mammography: The impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology 307:e222176
doi: 10.1148/radiol.222176 pubmed: 37129490
Lee P, Bubeck S, Petro J (2023) Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N Engl J Med 388:1233–1239
doi: 10.1056/NEJMsr2214184 pubmed: 36988602
Lightman H, Kosaraju V, Burda Y et al (2023) Let’s Verify Step by Step. arXiv:230520050 https://doi.org/10.48550/arXiv.2305.20050
Chen L, Zaharia M, Zou J (2023) How is ChatGPT’s behavior changing over time? arXiv:230709009 https://doi.org/10.48550/arXiv.2307.09009
White J, Fu Q, Hays S et al (2023) A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv:230211382 https://doi.org/10.48550/arXiv.2302.11382

Auteurs

Robert Siepmann (R)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Marc Huppertz (M)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Annika Rastkhiz (A)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Matthias Reen (M)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Eric Corban (E)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Christian Schmidt (C)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Stephan Wilke (S)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Philipp Schad (P)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Can Yüksel (C)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Christiane Kuhl (C)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Daniel Truhn (D)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.

Sven Nebelung (S)

Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. snebelung@ukaachen.de.

Classifications MeSH