Performance of ChatGPT on Nephrology Test Questions.
Journal
Clinical journal of the American Society of Nephrology : CJASN
ISSN: 1555-905X
Titre abrégé: Clin J Am Soc Nephrol
Pays: United States
ID NLM: 101271570
Informations de publication
Date de publication:
18 Oct 2023
18 Oct 2023
Historique:
received:
01
06
2023
accepted:
12
10
2023
pubmed:
18
10
2023
medline:
18
10
2023
entrez:
18
10
2023
Statut:
aheadofprint
Résumé
ChatGPT is a novel tool that allows people to engage in conversations with an advanced machine learning model. ChatGPT's performance in the US Medical Licensing Examination is comparable with a successful candidate's performance. However, its performance in the nephrology field remains undetermined. This study assessed ChatGPT's capabilities in answering nephrology test questions. Questions sourced from Nephrology Self-Assessment Program and Kidney Self-Assessment Program were used, each with multiple-choice single-answer questions. Questions containing visual elements were excluded. Each question bank was run twice using GPT-3.5 and GPT-4. Total accuracy rate, defined as the percentage of correct answers obtained by ChatGPT in either the first or second run, and the total concordance, defined as the percentage of identical answers provided by ChatGPT during both runs, regardless of their correctness, were used to assess its performance. A comprehensive assessment was conducted on a set of 975 questions, comprising 508 questions from Nephrology Self-Assessment Program and 467 from Kidney Self-Assessment Program. GPT-3.5 resulted in a total accuracy rate of 51%. Notably, the employment of Nephrology Self-Assessment Program yielded a higher accuracy rate compared with Kidney Self-Assessment Program (58% versus 44%; P < 0.001). The total concordance rate across all questions was 78%, with correct answers exhibiting a higher concordance rate (84%) compared with incorrect answers (73%) ( P < 0.001). When examining various nephrology subfields, the total accuracy rates were relatively lower in electrolyte and acid-base disorder, glomerular disease, and kidney-related bone and stone disorders. The total accuracy rate of GPT-4's response was 74%, higher than GPT-3.5 ( P < 0.001) but remained below the passing threshold and average scores of nephrology examinees (77%). ChatGPT exhibited limitations regarding accuracy and repeatability when addressing nephrology-related questions. Variations in performance were evident across various subfields.
Sections du résumé
BACKGROUND
BACKGROUND
ChatGPT is a novel tool that allows people to engage in conversations with an advanced machine learning model. ChatGPT's performance in the US Medical Licensing Examination is comparable with a successful candidate's performance. However, its performance in the nephrology field remains undetermined. This study assessed ChatGPT's capabilities in answering nephrology test questions.
METHODS
METHODS
Questions sourced from Nephrology Self-Assessment Program and Kidney Self-Assessment Program were used, each with multiple-choice single-answer questions. Questions containing visual elements were excluded. Each question bank was run twice using GPT-3.5 and GPT-4. Total accuracy rate, defined as the percentage of correct answers obtained by ChatGPT in either the first or second run, and the total concordance, defined as the percentage of identical answers provided by ChatGPT during both runs, regardless of their correctness, were used to assess its performance.
RESULTS
RESULTS
A comprehensive assessment was conducted on a set of 975 questions, comprising 508 questions from Nephrology Self-Assessment Program and 467 from Kidney Self-Assessment Program. GPT-3.5 resulted in a total accuracy rate of 51%. Notably, the employment of Nephrology Self-Assessment Program yielded a higher accuracy rate compared with Kidney Self-Assessment Program (58% versus 44%; P < 0.001). The total concordance rate across all questions was 78%, with correct answers exhibiting a higher concordance rate (84%) compared with incorrect answers (73%) ( P < 0.001). When examining various nephrology subfields, the total accuracy rates were relatively lower in electrolyte and acid-base disorder, glomerular disease, and kidney-related bone and stone disorders. The total accuracy rate of GPT-4's response was 74%, higher than GPT-3.5 ( P < 0.001) but remained below the passing threshold and average scores of nephrology examinees (77%).
CONCLUSIONS
CONCLUSIONS
ChatGPT exhibited limitations regarding accuracy and repeatability when addressing nephrology-related questions. Variations in performance were evident across various subfields.
Identifiants
pubmed: 37851468
doi: 10.2215/CJN.0000000000000330
pii: 01277230-990000000-00266
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Informations de copyright
Copyright © 2023 by the American Society of Nephrology.
Références
Krisanapan P, Tangpanithandee S, Thongprayoon C, Pattharanitima P, Cheungpasitporn W. Revolutionizing chronic kidney disease management with machine learning and artificial intelligence. J Clin Med. 2023;12(8):3018. doi: 10.3390/jcm12083018
doi: 10.3390/jcm12083018
Thongprayoon C, Kaewput W, Kovvuru K, et al. Promises of Big data and artificial intelligence in nephrology and transplantation. J Clin Med. 2020;9(4):1107. doi: 10.3390/jcm9041107
doi: 10.3390/jcm9041107
Thongprayoon C, Vaitla P, Jadlowiec CC, et al. Use of machine learning consensus clustering to identify distinct subtypes of Black kidney transplant recipients and associated outcomes. JAMA Surg. 2022;157(7):e221286. doi: 10.1001/jamasurg.2022.1286
doi: 10.1001/jamasurg.2022.1286
OpenAI. Introducing ChatGPT. Accessed November 30, 2023. https://openai.com/blog/chatgpt
Eysenbach G. The role of ChatGPT, generative language models, and artificial intelligence in medical education: a conversation with ChatGPT and a call for papers. JMIR Med Educ. 2023;9:e46885. doi: 10.2196/46885
doi: 10.2196/46885
Sallam M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare (Basel). 2023;11(6):887. doi: 10.3390/healthcare11060887
doi: 10.3390/healthcare11060887
Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ. 2023;9:e45312. doi: 10.2196/45312
doi: 10.2196/45312
Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2:e0000198. doi: 10.1371/journal.pdig.0000198
doi: 10.1371/journal.pdig.0000198
Munoz-Zuluaga C, Zhao Z, Wang F, Greenblatt MB, Yang HS. Assessing the accuracy and clinical utility of ChatGPT in laboratory medicine. Clin Chem. 2023;69(8):939–940. doi: 10.1093/clinchem/hvad058
doi: 10.1093/clinchem/hvad058
Hoffer EK. ChatGPT provides references that are real, inappropriate, or (most often) fake. J Vasc Interv Radiol. 2023. doi: 10.1016/j.jvir.2023.07.001
doi: 10.1016/j.jvir.2023.07.001
Thirunavukarasu AJ, Hassan R, Mahmood S, et al. Trialling a large language model (ChatGPT) in general practice with the applied knowledge test: observational study demonstrating opportunities and limitations in primary care. JMIR Med Educ. 2023;9:e46599. doi: 10.2196/46599
doi: 10.2196/46599
Miao J, Thongprayoon C, Cheungpasitporn W. Assessing the accuracy of ChatGPT on core questions in glomerular disease. Kidney Int Rep. 2023;8(8):1657–1659. doi: 10.1016/j.ekir.2023.05.014
doi: 10.1016/j.ekir.2023.05.014
Oh N, Choi GS, Lee WY. ChatGPT goes to the operating room: evaluating GPT-4 performance and its potential in surgical education and training in the era of large language models. Ann Surg Treat Res. 2023;104(5):269–273. doi: 10.4174/astr.2023.104.5.269
doi: 10.4174/astr.2023.104.5.269
Rosol M, Gasior JS, Łaba J, Korzeniewski K, Młyńczak M. Evaluation of the performance of GPT-3.5 and GPT-4 on the medical final examination. medRxiv. 2023. doi: 10.1101/2023.06.04.23290939
doi: 10.1101/2023.06.04.23290939
Deebel NA, Terlecki R. ChatGPT performance on the American Urological Association self-assessment study program and the potential influence of artificial intelligence in urologic training. Urology. 2023;177:29–33. doi: 10.1016/j.urology.2023.05.010
doi: 10.1016/j.urology.2023.05.010
Mihalache A, Popovic MM, Muni RH. Performance of an artificial intelligence chatbot in ophthalmic knowledge assessment. JAMA Ophthalmol. 2023;141(6):589–597. doi: 10.1001/jamaophthalmol.2023.1144
doi: 10.1001/jamaophthalmol.2023.1144
Suchman K, Garg S, Trindade AJ. Chat generative pretrained transformer fails the multiple-choice American College of Gastroenterology self-assessment test. Am J Gastroenterol. 2023. doi: 10.14309/ajg.0000000000002320
doi: 10.14309/ajg.0000000000002320
Huh S. Are ChatGPT's knowledge and interpretation ability comparable to those of medical students in Korea for taking a parasitology examination?: a descriptive study. J Educ Eval Health Prof. 2023;20(1):1. doi: 10.3352/jeehp.2023.20.1
doi: 10.3352/jeehp.2023.20.1
Wang YM, Shen HW, Chen TJ. Performance of ChatGPT on the pharmacist licensing examination in Taiwan. J Chin Med Assoc. 2023;86(7):653–658. doi: 10.1097/JCMA.0000000000000942
doi: 10.1097/JCMA.0000000000000942
Giannos P, Delardas O. Performance of ChatGPT on UK standardized admission tests: insights from the BMAT, TMUA, LNAT, and TSA examinations. JMIR Med Educ. 2023;9:e47737. doi: 10.2196/47737
doi: 10.2196/47737
Takagi S, Watari T, Erabi A, Sakaguchi K. Performance of GPT-3.5 and GPT-4 on the Japanese medical licensing examination: comparison study. JMIR Med Educ. 2023;9:e48002. doi: 10.2196/48002
doi: 10.2196/48002
Bhayana R, Krishna S, Bleakney RR. Performance of ChatGPT on a radiology board-style examination: insights into current strengths and limitations. Radiology. 2023;307(5):e230582. doi: 10.1148/radiol.230582
doi: 10.1148/radiol.230582
Nori H, King N, Mckinney SM, Carignan D, Horvitz E. Capabilities of GPT-4 on medical challenge problems. arXiv. 2023. doi: 10.48550/arXiv.2303.13375
doi: 10.48550/arXiv.2303.13375
Singhal K, Tu T, Gottweis J, et al. Towards expert-level medical question answering with large language models. arXiv. 2023. doi: 10.48550/arXiv.2305.09617
doi: 10.48550/arXiv.2305.09617
Anil R, Dai AM, Firat O, et al. PaLM 2 technical report. arXiv. 2023. doi: 10.48550/arXiv.2305.10403
doi: 10.48550/arXiv.2305.10403
Suppadungsuk S, Thongprayoon C, Krisanapan P, et al. Examining the validity of ChatGPT in identifying relevant nephrology literature: findings and implications. J Clin Med. 2023;12(17):5550. doi: 10.3390/jcm12175550
doi: 10.3390/jcm12175550
Garcia Valencia OA, Suppadungsuk S, Thongprayoon C, et al. Ethical implications of chatbot utilization in nephrology. J Personalized Med. 2023;13(9):1363. doi: 10.3390/jpm13091363
doi: 10.3390/jpm13091363