A Language Model-Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study.

Humans Prospective Studies Medical History Taking / methods Students, Medical / psychology Patient Simulation Female Male Clinical Competence / standards Artificial Intelligence Feedback Reproducibility of Results Education, Medical, Undergraduate / methods

ChatGPT GPT: LLM LLMs NLP TEL artificial intelligence chatbot chatbots communication communication skills conversational agent conversational agents histories history interaction interactions language model language models machine learning medical education natural language processing relationship relationships simulated student students technology enhanced education virtual patients communication

Journal

JMIR medical education

ISSN: 2369-3762

Titre abrégé: JMIR Med Educ

Pays: Canada

ID NLM: 101684518

Informations de publication

Date de publication:
16 Aug 2024

Historique:

received: 05 04 2024

accepted: 27 06 2024

revised: 21 05 2024

medline: 16 8 2024

pubmed: 16 8 2024

entrez: 16 8 2024

Statut: epublish

Résumé

Although history taking is fundamental for diagnosing medical conditions, teaching and providing feedback on the skill can be challenging due to resource constraints. Virtual simulated patients and web-based chatbots have thus emerged as educational tools, with recent advancements in artificial intelligence (AI) such as large language models (LLMs) enhancing their realism and potential to provide feedback. In our study, we aimed to evaluate the effectiveness of a Generative Pretrained Transformer (GPT) 4 model to provide structured feedback on medical students' performance in history taking with a simulated patient. We conducted a prospective study involving medical students performing history taking with a GPT-powered chatbot. To that end, we designed a chatbot to simulate patients' responses and provide immediate feedback on the comprehensiveness of the students' history taking. Students' interactions with the chatbot were analyzed, and feedback from the chatbot was compared with feedback from a human rater. We measured interrater reliability and performed a descriptive analysis to assess the quality of feedback. Most of the study's participants were in their third year of medical school. A total of 1894 question-answer pairs from 106 conversations were included in our analysis. GPT-4's role-play and responses were medically plausible in more than 99% of cases. Interrater reliability between GPT-4 and the human rater showed "almost perfect" agreement (Cohen κ=0.832). Less agreement (κ<0.6) detected for 8 out of 45 feedback categories highlighted topics about which the model's assessments were overly specific or diverged from human judgement. The GPT model was effective in providing structured feedback on history-taking dialogs provided by medical students. Although we unraveled some limitations regarding the specificity of feedback for certain feedback categories, the overall high agreement with human raters suggests that LLMs can be a valuable tool for medical education. Our findings, thus, advocate the careful integration of AI-driven feedback mechanisms in medical training and highlight important aspects when LLMs are used in that context.

Sections du résumé

BACKGROUND BACKGROUND

OBJECTIVE OBJECTIVE

In our study, we aimed to evaluate the effectiveness of a Generative Pretrained Transformer (GPT) 4 model to provide structured feedback on medical students' performance in history taking with a simulated patient.

METHODS METHODS

We conducted a prospective study involving medical students performing history taking with a GPT-powered chatbot. To that end, we designed a chatbot to simulate patients' responses and provide immediate feedback on the comprehensiveness of the students' history taking. Students' interactions with the chatbot were analyzed, and feedback from the chatbot was compared with feedback from a human rater. We measured interrater reliability and performed a descriptive analysis to assess the quality of feedback.

RESULTS RESULTS

Most of the study's participants were in their third year of medical school. A total of 1894 question-answer pairs from 106 conversations were included in our analysis. GPT-4's role-play and responses were medically plausible in more than 99% of cases. Interrater reliability between GPT-4 and the human rater showed "almost perfect" agreement (Cohen κ=0.832). Less agreement (κ<0.6) detected for 8 out of 45 feedback categories highlighted topics about which the model's assessments were overly specific or diverged from human judgement.

CONCLUSIONS CONCLUSIONS

The GPT model was effective in providing structured feedback on history-taking dialogs provided by medical students. Although we unraveled some limitations regarding the specificity of feedback for certain feedback categories, the overall high agreement with human raters suggests that LLMs can be a valuable tool for medical education. Our findings, thus, advocate the careful integration of AI-driven feedback mechanisms in medical training and highlight important aspects when LLMs are used in that context.

Identifiants

DOI: 10.2196/59213 PMID: 39150749

pubmed: 39150749

pii: v10i1e59213

doi: 10.2196/59213

doi:

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

Pagination

e59213

Informations de copyright

©Friederike Holderried, Christian Stegemann-Philipps, Anne Herrmann-Werner, Teresa Festl-Wietek, Martin Holderried, Carsten Eickhoff, Moritz Mahling. Originally published in JMIR Medical Education (https://mededu.jmir.org), 16.08.2024.

A Language Model-Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study.

Journal

Informations de publication

Résumé

Sections du résumé

Identifiants

Types de publication

Langues

Sous-ensembles de citation

Pagination

Informations de copyright

Auteurs

Friederike Holderried (F)

Christian Stegemann-Philipps (C)

Anne Herrmann-Werner (A)

Teresa Festl-Wietek (T)

Martin Holderried (M)

Carsten Eickhoff (C)

Moritz Mahling (M)

Articles similaires

[Redispensing of expensive oral anticancer medicines: a practical application].

Smoking Cessation and Incident Cardiovascular Disease.

Evaluation of Low-Value Services Across Major Medicare Advantage Insurers and Traditional Medicare.

Effectiveness of Virtual Yoga for Chronic Low Back Pain: A Randomized Clinical Trial.

Classifications MeSH