An overview of diagnostics and therapeutics using large language models.
Journal
Journal of traumatic stress
ISSN: 1573-6598
Titre abrégé: J Trauma Stress
Pays: United States
ID NLM: 8809259
Informations de publication
Date de publication:
18 Jul 2024
18 Jul 2024
Historique:
revised:
04
06
2024
received:
16
04
2024
accepted:
07
06
2024
medline:
18
7
2024
pubmed:
18
7
2024
entrez:
18
7
2024
Statut:
aheadofprint
Résumé
There is an acute need for solutions to treat stress and trauma-related sequelae, and there are well-documented shortages of qualified human professionals. Artificial intelligence (AI) presents an opportunity to create advanced screening, diagnosis, and treatment solutions that relieve the burden on people and can provide just-in-time interventions. Large language models (LLMs), in particular, are promising given the role language plays in understanding and treating traumatic stress and other mental health conditions. In this article, we provide an overview of the state-of-the-art LLMs applications in diagnostic assessments, clinical note generation, and therapeutic support. We discuss the open research direction and challenges that need to be overcome to realize the full potential of deploying language models for use in clinical contexts. We highlight the need for increased representation in AI systems to ensure there are no disparities in access. Public datasets and models will help lead progress toward better models; however, privacy-preserving model training will be necessary for protecting patient data.
Types de publication
Journal Article
Review
Langues
eng
Sous-ensembles de citation
IM
Subventions
Organisme : NIMH NIH HHS
ID : K23MH134068
Pays : United States
Informations de copyright
© 2024 International Society for Traumatic Stress Studies.
Références
Acerbi, A., & Stubbersfield, J. M. (2023). Large language models show human‐like content biases in transmission chain experiments. Proceedings of the National Academy of Sciences, 120(44), Article e2313790120. https://doi.org/10.1073/pnas.2313790120
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Jauregui Ruano, R., Jeffrey, K., … Zeng, A. (2022). Do as I can, not as I say: Grounding language in robotic affordances. arXiv. https://doi.org/10.48550/arXiv.2204.01691
American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). https://doi.org/10.1176/appi.books.9780890425596
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv. https://doi.org/10.48550/arXiv.1409.0473
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, A., Herbert‐Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few‐shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds). Advances in Neural Information Processing Systems 33 (pp. 1877–1901). NeurIPS.
Cambon, A., Hecht, B., Edelman, B., Ngwe, D., Jaffe, S., Heger, A., Vorvoreanu, M., Peng, S., Hofman, J., Farach, A., Bermejo‐Cano, M., Knudsen, E., Bono, J., Sanghavi, H., Spatharioti, S., Rothschild, D., Goldstein, D. G., Kalliamvakou, E., Cichon, P., … Teevan, J. (2023). Early llm‐based tools for enterprise information workers likely provide meaningful boosts to productivity [Technical report]. https://www.microsoft.com/en‐us/research/publication/early‐llm‐based‐tools‐for‐enterprise‐information‐workers‐likely‐provide‐meaningful‐boosts‐to‐productivity
Chiu, Y. Y., Sharma, A., Lin, I. W., & Althoff, T. (2024). A computational framework for behavioral assessment of llm therapists. arXiv. https://doi.org/10.48550/arXiv.2401.00820
Gabriel, S., Puri, I., Xu, X., Malgaroli, M., & Ghassemi, M. (2024). Can AI relate: Testing large language model response for mental health support. arXiv. https://doi.org/10.48550/arXiv.2405.12021
Galatzer‐Levy, I. R., McDuff, D., Natarajan, V., Karthikesalingam, A., & Malgaroli, M. (2023). The capability of large language models to measure psychiatric functioning. arXiv. https://doi.org/10.48550/arXiv.2308.01834
Garcia, P., Ma, S. P., Shah, S., Smith, M., Jeong, Y., Devon‐Sand, A., Tai‐Seale, M., Takazawa, K., Clutter, D., Vogt, K., Lugtu, C., Rojo, M., Lin, S., Shanafelt, T., Pfeffer, M. A., & Sharp, C. (2024). Artificial intelligence–generated draft replies to patient inbox messages. JAMA Network Open, 7(3), Article e243201. https://doi.org/10.1001/jamanetworkopen.2024.3201
Glenberg, A. M., Havas, D., Becker, R., & Rinck, M. (2005). Grounding language in bodily states. In D. Pecher & R. A. Zwaan (Eds). Grounding cognition: The role of perception and action in memory, language, and thinking (pp. 115–128). Cambridge University Press.
Huber, B., McDuff, D., Brockett, C., Galley, M., & Dolan, B. (2018). Emotional dialogue generation using imagegrounded language models. In CHI ’18: Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (pp. 1–12). https://doi.org/10.1145/3173574.3173851
Jin, Y., Chandra, M., Verma, G., Hu, Y., Choudhury, M. D., & Kumar, S. (2023). Better to ask in English: Cross‐lingual evaluation of large language models for healthcare queries. arXiv. https://doi.org/10.48550/arXiv.2310.13132
Kim, Y., Xu, X., McDuff, D., Breazeal, C., & Park, H. W. (2024). Health‐llm: Large language models for health prediction via wearable sensor data. arXiv. https://doi.org/10.48550/arXiv.2401.06866
Le Scao, T., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., Tow, J., Rush, A. M., Biderman, S., Webson, A., Ammanamanchi, P. S., Wang, T., Sagot, B., Muennighoff, N., Villanova del Moral, A., … Wolf, T. (2022). BLOOM: A 176b‐parameter open‐access multilingual language model. arXiv. https://doi.org/10.48550/arXiv.2211.05100
Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J. ‐N., Chang, K. ‐W., & Gao, J. (2022). Grounded language‐image pre‐training. arXiv. https://doi.org/10.48550/arXiv.2112.03857
Liu, X., McDuff, D., Kovacs, G., Galatzer‐Levy, I., Sunshine, J., Zhan, J., Poh, M. ‐Z., Liao, S., Di Achille, P., & Patel, S. (2023). Large language models are few‐shot health learners. arXiv. https://doi.org/10.48550/arXiv.2305.15525
Liu, Y., Jain, A., Eng, C., Way, D. H., Lee, K., Bui, P., Kanada, K., de Oliveira Marinho, G., Gallegos, J., Gabriele, S., Gupta, V., Singh, N., Natarajan, V., Hofmann‐Wellenhof, R., Corrado, G. S., Peng, L. H., Webster, D. R., Ai, D., Huang, S. J., & Coz, D. (2020). A deep learning system for differential diagnosis of skin diseases. Nature Medicine, 26(6), 900–908. https://doi.org/10.1038/s41591‐020‐0842‐3
Malgaroli, M., Hull, T. D., Zech, J. M., & Althoff, T. (2023). Natural language processing for mental health interventions: A systematic review and research framework. Translational Psychiatry, 13(1), Article 309. https://doi.org/10.1038/s41398‐023‐02592‐2
Malgaroli, M., & Schultebraucks, K. (2020). Artificial intelligence and posttraumatic stress disorder (PTSD). European Psychologist, 25(4), 272–282. https://doi.org/10.1027/1016‐9040/a000423
Malgaroli, M., Tseng, E., Hull, T. D., Jennings, E., Choudhury, T. K., & Simon, N. M. (2023). Association of health care work with anxiety and depression during the COVID‐19 pandemic: Structural topic modeling study. JMIR AI, 2(1), Article e47223. https://doi.org/10.2196/47223
McDuff, D., Schaekermann, M., Tu, T., Palepu, A., Wang, A., Garrison, J., Singhal, K., Sharma, Y., Azizi, S., Kulkarni, K., Hou, L., Cheng, Y., Liu, Y., Mahdavi, S. S., Prakash, S., Pathak, A., Semturs, C., Patel, S., Webster, D. R., & …Natarajan, V. (2023). Towards accurate differential diagnosis with large language models. arXiv. https://doi.org/10.48550/arXiv.2312.00164
Mohr, D. C., Zhang, M., & Schueller, S. M. (2017). Personal sensing: Understanding mental health using ubiquitous sensors and machine learning. Annual Review Of Clinical Psychology, 13, 23–47. https://doi.org/10.1146/annurev‐clinpsy‐032816‐044949
Mostafazadeh, N., Brockett, C., Dolan, B., Galley, M., Gao, J., Spithourakis, G. P., & Vanderwende, L. (2017). Image‐grounded conversations: Multimodal context for natural question and response generation. arXiv. https://doi.org/10.48550/arXiv.1701.08251
Norman, K. P., Govindjee, A., Norman, S. R., Godoy, M., Cerrone, K. L., Kieschnick, D. W., & Kassler, W. (2020). Natural language processing tools for assessing progress and outcome of two veteran populations: Cohort study from a novel online intervention for posttraumatic growth. JMIR Formative Research, 4(9), Article e17424. https://doi.org/10.2196/17424
Petrov, A., La Malfa, E., Torr, P., & Bibi, A. (2024). Language model tokenizers introduce unfairness between languages. In Advances in neural information processing systems, 36, (pp. 1–28). NeurIPS.
Rauschecker, A. M., Rudie, J. D., Xie, L., Wang, J., Duong, M. T., Botzolakis, E. J., Kovalovich, A. M., Egan, J., Cook, T. C., Bryan, R. N., Nasrallah, I. M., Mohan, S., & Gee, J. C. (2020). Artificial intelligence system approaching neuroradiologist‐level differential diagnosis accuracy at brain MRI. Radiology, 295(3), 626–637. https://doi.org/10.1148/radiol.2020190283
Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage? arXiv. https://doi.org/10.48550/arXiv.2304.15004
Schueller, S. M., & Morris, R. R. (2023). Clinical science and practice in the age of large language models and generative artificial intelligence. Journal of Consulting and Clinical Psychology, 91(10), 559–561. https://doi.org/10.1037/ccp0000848
Schultebraucks, K., Yadav, V., Shalev, A. Y., Bonanno, G. A., & Galatzer‐Levy, I. R. (2022). Deep learning‐based classification of posttraumatic stress disorder and depression following trauma utilizing visual and auditory markers of arousal and mood. Psychological Medicine, 52(5), 957–967. https://doi.org/10.1017/S0033291720002718
Shah, R. S., Holt, F., Hayati, S. A., Agarwal, A., Wang, Y. ‐C., Kraut, R. E., & Yang, D. (2022). Modeling motivational interviewing strategies on an online peer‐to‐peer counseling platform. Proceedings of the ACM on HumanComputer Interaction, 6, Article 527.
Sharma, A., Lin, I. W., Miner, A. S., Atkins, D. C., & Althoff, T. (2023). Human–AI collaboration enables more empathic conversations in text‐based peer‐to‐peer mental health support. Nature Machine Intelligence, 5(1), 46–57. https://doi.org/10.1038/s42256‐022‐00593‐2
Sharma, A., Rushton, K., Lin, I. W., Nguyen, T., & Althoff, T. (2023). Facilitating self‐guided mental health interventions through human‐language model interaction: A Case study of cognitive restructuring. arXiv. https://doi.org/10.48550/arXiv.2310.15461
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Hou, L., Clark, K., Pfohl, S., Cole‐Lewis, H., Neal, D., Schaekermann, M., Wang, A., Amin, M., Lachgar, S., Mansfield, P., Prakash, S., Green, B., Dominowska, E., y Arcas, B. A., … Natarajan, V. (2023). Towards expert‐level medical question answering with large language models. arXiv. https://doi.org/10.48550/arXiv.2305.09617
So, J. ‐ H., Chang, J., Kim, E., Na, J., Choi, J., Sohn, J. ‐Y., Kim, B.‐H., & Chu, S. H. (2024). Aligning large language models for enhancing psychiatric interviews through symptom delineation and summarization. arXiv. https://doi.org/10.48550/arXiv.2403.17428
Son, Y., Clouston, S. A., Kotov, R., Eichstaedt, J. C., Bromet, E. J., Luft, B. J., & Schwartz, H. A. (2023). World Trade Center responders in their own words: Predicting PTSD symptom trajectories with AI‐based language analyses of interviews. Psychological Medicine, 53(3), 918–926. https://doi.org/10.1017/S0033291721002294
Szolovits, P., & Pauker, S. G. (1978). Categorical and probabilistic reasoning in medical diagnosis. Artificial Intelligence, 11(1‐2), 115–144.
United Nations High Commissioner for Refugees. (2023). Mid‐year trends 2023. https://www.unhcr.org/mid‐year‐trends‐report‐2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30. NeurIPS.
Webster, P. (2023). Six ways large language models are changing healthcare. Nature Medicine, 29(12), 2969–2971. https://doi.org/10.1038/s41591‐023‐02700‐1
World Health Organization. (2022). World mental health report: Transforming mental health for all. https://www.who.int/publications/i/item/9789240049338
Xu, X., Dinkel, H., Wu, M., & Yu, K. (2021). Text‐to‐audio grounding: Building correspondence between captions and sound events. In ICASSP 2021‐2021 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 606–610). ICASSP.
Ye, R., Wang, W., Chai, J., Li, D., Li, Z., Xu, Y., Du, Y., Wang, Y., & Chen, S. (2024). Openfedllm: Training large language models on decentralized private data via federated learning. arXiv. https://doi.org/10.48550/arXiv.2402.06954
Zhou, K., Ethayarajh, K., & Jurafsky, D. (2022). Richer countries and richer representations. arXiv. https://doi.org/10.48550/arXiv.2205.05093
Zhou, L., Kalantidis, Y., Chen, X., Corso, J. J., & Rohrbach, M. (2019). Grounded video description. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 6578–6587). IEEE/CVF.