Usefulness and Accuracy of Artificial Intelligence Chatbot Responses to Patient Questions for Neurosurgical Procedures.
Journal
Neurosurgery
ISSN: 1524-4040
Titre abrégé: Neurosurgery
Pays: United States
ID NLM: 7802914
Informations de publication
Date de publication:
14 Feb 2024
14 Feb 2024
Historique:
received:
30
08
2023
accepted:
17
12
2023
medline:
14
2
2024
pubmed:
14
2
2024
entrez:
14
2
2024
Statut:
aheadofprint
Résumé
The Internet has become a primary source of health information, leading patients to seek answers online before consulting health care providers. This study aims to evaluate the implementation of Chat Generative Pre-Trained Transformer (ChatGPT) in neurosurgery by assessing the accuracy and helpfulness of artificial intelligence (AI)-generated responses to common postsurgical questions. A list of 60 commonly asked questions regarding neurosurgical procedures was developed. ChatGPT-3.0, ChatGPT-3.5, and ChatGPT-4.0 responses to these questions were recorded and graded by numerous practitioners for accuracy and helpfulness. The understandability and actionability of the answers were assessed using the Patient Education Materials Assessment Tool. Readability analysis was conducted using established scales. A total of 1080 responses were evaluated, equally divided among ChatGPT-3.0, 3.5, and 4.0, each contributing 360 responses. The mean helpfulness score across the 3 subsections was 3.511 ± 0.647 while the accuracy score was 4.165 ± 0.567. The Patient Education Materials Assessment Tool analysis revealed that the AI-generated responses had higher actionability scores than understandability. This indicates that the answers provided practical guidance and recommendations that patients could apply effectively. On the other hand, the mean Flesch Reading Ease score was 33.5, suggesting that the readability level of the responses was relatively complex. The Raygor Readability Estimate scores ranged within the graduate level, with an average score of the 15th grade. The artificial intelligence chatbot's responses, although factually accurate, were not rated highly beneficial, with only marginal differences in perceived helpfulness and accuracy between ChatGPT-3.0 and ChatGPT-3.5 versions. Despite this, the responses from ChatGPT-4.0 showed a notable improvement in understandability, indicating enhanced readability over earlier versions.
Sections du résumé
BACKGROUND AND OBJECTIVES
OBJECTIVE
The Internet has become a primary source of health information, leading patients to seek answers online before consulting health care providers. This study aims to evaluate the implementation of Chat Generative Pre-Trained Transformer (ChatGPT) in neurosurgery by assessing the accuracy and helpfulness of artificial intelligence (AI)-generated responses to common postsurgical questions.
METHODS
METHODS
A list of 60 commonly asked questions regarding neurosurgical procedures was developed. ChatGPT-3.0, ChatGPT-3.5, and ChatGPT-4.0 responses to these questions were recorded and graded by numerous practitioners for accuracy and helpfulness. The understandability and actionability of the answers were assessed using the Patient Education Materials Assessment Tool. Readability analysis was conducted using established scales.
RESULTS
RESULTS
A total of 1080 responses were evaluated, equally divided among ChatGPT-3.0, 3.5, and 4.0, each contributing 360 responses. The mean helpfulness score across the 3 subsections was 3.511 ± 0.647 while the accuracy score was 4.165 ± 0.567. The Patient Education Materials Assessment Tool analysis revealed that the AI-generated responses had higher actionability scores than understandability. This indicates that the answers provided practical guidance and recommendations that patients could apply effectively. On the other hand, the mean Flesch Reading Ease score was 33.5, suggesting that the readability level of the responses was relatively complex. The Raygor Readability Estimate scores ranged within the graduate level, with an average score of the 15th grade.
CONCLUSION
CONCLUSIONS
The artificial intelligence chatbot's responses, although factually accurate, were not rated highly beneficial, with only marginal differences in perceived helpfulness and accuracy between ChatGPT-3.0 and ChatGPT-3.5 versions. Despite this, the responses from ChatGPT-4.0 showed a notable improvement in understandability, indicating enhanced readability over earlier versions.
Identifiants
pubmed: 38353558
doi: 10.1227/neu.0000000000002856
pii: 00006123-990000000-01053
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Informations de copyright
Copyright © Congress of Neurological Surgeons 2024. All rights reserved.
Références
Zhou S, Zhou F, Sun Y, et al. The application of artificial intelligence in spine surgery. Front Surg. 2022;9:885599.
Mofatteh M. Neurosurgery and artificial intelligence. AIMS Neurosci. 2021;8(4):477-495.
Emblem KE, Nedregaard B, Hald JK, Nome T, Due-Tonnessen P, Bjornerud A. Automatic glioma characterization from dynamic susceptibility contrast imaging: brain tumor segmentation using knowledge-based fuzzy clustering. J Magn Reson Imaging. 2009;30(1):1-10.
Shi Z, Hu B, Schoepf UJ, et al. Artificial intelligence in the management of intracranial aneurysms: current status and future perspectives. AJNR Am J Neuroradiol. 2020;41(3):373-379.
Doerr SA, Weber-Levine C, Hersh AM, et al. Automated prediction of the thoracolumbar injury classification and severity score from CT using a novel deep learning algorithm. Neurosurg Focus. 2022;52(4):E5.
Yanni DS, Ozgur BM, Louis RG, et al. Real-time navigation guidance with intraoperative CT imaging for pedicle screw placement using an augmented reality head-mounted display: a proof-of-concept study. Neurosurg Focus. 2021;51(2):e11.
Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2):e0000198.
Hopkins BS, Nguyen VN, Dallas J, et al. ChatGPT versus the neurosurgical written boards: a comparative analysis of artificial intelligence/machine learning performance on neurosurgical board-style questions. J Neurosurg. 2023;139(3):904-911.
Stokel-Walker C. ChatGPT listed as author on research papers: many scientists disapprove. Nature. 2023;613(7945):620-621.
Dubin JA, Bains SS, Chen Z, et al. Using a Google Web search analysis to assess the utility of ChatGPT in total joint arthroplasty. J Arthroplasty. 2023;38(7):1195-1202.
Sun Y, Zhang Y, Gwizdka J, Trace CB. Consumer evaluation of the quality of online health information: systematic literature review of relevant criteria and indicators. J Med Internet Res. 2019;21(5):e12522.
Wong K, Keefe KR, Gilad A, Chong-Sun Li CJ, Levi JR. Parental actionability of educational materials regarding laryngotracheal reconstruction. JAMA Otolaryngol Head Neck Surg. 2017;143(9):953-954.
Agarwal N, Funahashi R, Taylor T, et al. Patient education and engagement through multimedia: a prospective pilot study on health literacy in patients with cerebral aneurysms. World Neurosurg. 2020;138:e819-e826.
Hansberry DR, Agarwal N, John ES, et al. Evaluation of internet-based patient education materials from internal medicine subspecialty organizations: will patients understand them? Intern Emerg Med. 2017;12(4):535-543.
Hansberry DR, D'Angelo M, White MD, et al. Quantitative analysis of the level of readability of online emergency radiology-based patient education resources. Emerg Radiol. 2018;25(2):147-152.
Kamath P, Zheng R, Narasimman M, et al. Evaluation of online patient education materials concerning skin cancers. J Am Acad Dermatol. 2021;84(1):190-191.
Kim C, Prabhu AV, Hansberry DR, Agarwal N, Heron DE, Beriwal S. Digital era of mobile communications and smartphones: a novel analysis of patient comprehension of cancer-related information available through mobile applications. Cancer Invest. 2019;37(3):127-133.
Para A, Thelmo F, Rynecki ND, et al. Evaluating the readability of online patient education materials related to orthopedic oncology. Orthopedics. 2021;44(1):38-42.
Prabhu AV, Donovan AL, Crihalmeanu T, et al. Radiology online patient education materials provided by major university hospitals: do they conform to NIH and AMA guidelines? Curr Probl Diagn Radiol. 2018;47(2):75-79.
Rooney MK, Santiago G, Perni S, et al. Readability of patient education materials from high-impact medical journals: a 20-year analysis. J Patient Exp. 2021;8:2374373521998847.
Oteri V, Martinelli A, Crivellaro E, Gigli F. The impact of preoperative anxiety on patients undergoing brain surgery: a systematic review. Neurosurg Rev. 2021;44(6):3047-3057.
Quivr—Get a Second Brain with Generative AI. Accessed December 3, 2023. https://www.quivr.app/