Are large language models valid tools for patient information on lumbar disc herniation? The spine surgeons' perspective.

AI evaluation ChatGPT Google bard Large language models Lumbar disc herniation Patient education

Journal

Brain & spine
ISSN: 2772-5294
Titre abrégé: Brain Spine
Pays: Netherlands
ID NLM: 9918470888906676

Informations de publication

Date de publication:
2024
Historique:
received: 27 12 2023
revised: 19 02 2024
accepted: 04 04 2024
medline: 6 5 2024
pubmed: 6 5 2024
entrez: 6 5 2024
Statut: epublish

Résumé

Generative AI is revolutionizing patient education in healthcare, particularly through chatbots that offer personalized, clear medical information. Reliability and accuracy are vital in AI-driven patient education. How effective are Large Language Models (LLM), such as ChatGPT and Google Bard, in delivering accurate and understandable patient education on lumbar disc herniation? Ten Frequently Asked Questions about lumbar disc herniation were selected from 133 questions and were submitted to three LLMs. Six experienced spine surgeons rated the responses on a scale from "excellent" to "unsatisfactory," and evaluated the answers for exhaustiveness, clarity, empathy, and length. Statistical analysis involved Fleiss Kappa, Chi-square, and Friedman tests. Out of the responses, 27.2% were excellent, 43.9% satisfactory with minimal clarification, 18.3% satisfactory with moderate clarification, and 10.6% unsatisfactory. There were no significant differences in overall ratings among the LLMs (p = 0.90); however, inter-rater reliability was not achieved, and large differences among raters were detected in the distribution of answer frequencies. Overall, ratings varied among the 10 answers (p = 0.043). The average ratings for exhaustiveness, clarity, empathy, and length were above 3.5/5. LLMs show potential in patient education for lumbar spine surgery, with generally positive feedback from evaluators. The new EU AI Act, enforcing strict regulation on AI systems, highlights the need for rigorous oversight in medical contexts. In the current study, the variability in evaluations and occasional inaccuracies underline the need for continuous improvement. Future research should involve more advanced models to enhance patient-physician communication.

Identifiants

pubmed: 38706800
doi: 10.1016/j.bas.2024.102804
pii: S2772-5294(24)00060-2
pmc: PMC11067000
doi:

Types de publication

Journal Article

Langues

eng

Pagination

102804

Informations de copyright

© 2024 The Authors.

Déclaration de conflit d'intérêts

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Auteurs

Siegmund Lang (S)

Department of Trauma Surgery, University Hospital Regensburg, Regensburg, Germany.

Jacopo Vitale (J)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Tamás F Fekete (TF)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Daniel Haschtmann (D)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Raluca Reitmeir (R)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Mario Ropelato (M)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Jani Puhakka (J)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Fabio Galbusera (F)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Markus Loibl (M)

Spine Center, Schulthess Klinik, Zurich, Switzerland.

Classifications MeSH