Brazilian Journal of Anesthesiology
https://app.periodikos.com.br/journal/rba/article/doi/10.1016/j.bjane.2026.844768
Brazilian Journal of Anesthesiology
Original Investigation

Evaluation of two large language models for intensive care unit discharge decisions: a prospective observational cohort study

Avaliação de dois grandes modelos de linguagem para decisões de alta da unidade de terapia intensiva: um estudo de coorte observacional prospectivo

Engin İhsan Turan, Abdurrahman Engin Baydemir, Ebru Kaya, Zehra Polat Turan, Ayça Sultan Şahin

Downloads: 0
Views: 8

Abstract

Background

The aim of this study was to evaluate the effectiveness of two general-purpose Large Language Models (LLMs), ChatGPT and Gemini, in predicting Intensive Care Unit (ICU) discharge decisions (discharge vs. non-discharge). By comparing their outputs with decisions made by ICU physicians, we sought to determine the alignment of AI-generated recommendations with expert clinical judgment and assess their potential as decision-support tools in critical care.

Methods

This prospective observational cohort study was conducted in a tertiary ICU between September 2024 and May 2025. Adult patients (≥ 18 years) requiring ICU discharge decisions were included. Standardized clinical prompts were generated from electronic health records and input into ChatGPT and Gemini. The models’ binary discharge decisions were compared to those of ICU physicians. Model performance was assessed using accuracy, sensitivity, specificity, F1 score, Cohen’s kappa, and McNemar’s test. Discharge was defined as the positive class for all diagnostic performance analyses.

Results

A total of 398 patients were analyzed. ChatGPT demonstrated higher accuracy than Gemini (87.2% vs. 66.3%), with higher sensitivity (85.9% vs. 46.9%) and F1 score (0.890 vs. 0.628), whereas Gemini showed higher specificity (96.2% vs. 89.2%). Agreement with clinician decisions was substantial for ChatGPT (κ = 0.737, p = 0.024) and fair for Gemini (κ = 0.379, p < 0.001). Laboratory markers such as lactate, hemoglobin, and procalcitonin significantly differed between discharged and non-discharged patients.

Conclusion

Large language models may support ICU discharge decisions when guided by structured, guideline-informed prompting. ChatGPT achieved higher overall accuracy, sensitivity, and F1 score, whereas Gemini demonstrated higher specificity.

Trial registration

Externation (Discharge) of ICU, NCT06584890, registered 03 September 2024, prospectively registered, https://register.clinicaltrials.gov/prs/beta/studies/S000EVXZ00000029/recordSummary.

Keywords

Artificial intelligence; Decision making; Decision support systems, Clinical; Intensive care units; Natural language processing; Patient discharge 

Resumo

Introdução

O objetivo deste estudo foi avaliar a eficácia de dois grandes modelos de linguagem (LLMs) de uso geral, ChatGPT e Gemini, na predição de decisões de alta da Unidade de Terapia Intensiva (UTI) (alta vs. não alta). Ao comparar seus resultados com as decisões tomadas por médicos intensivistas, buscou-se determinar o alinhamento das recomendações geradas por IA com o julgamento clínico especializado e avaliar seu potencial como ferramentas de apoio à decisão em cuidados intensivos.

Métodos

Este estudo de coorte observacional prospectivo foi realizado em uma UTI terciária entre setembro de 2024 e maio de 2025. Foram incluídos pacientes adultos (≥ 18 anos) que necessitavam de decisões de alta da UTI. Prompts clínicos padronizados foram gerados a partir de prontuários eletrônicos de saúde e inseridos no ChatGPT e no Gemini. As decisões binárias de alta dos modelos foram comparadas com as dos médicos intensivistas. O desempenho do modelo foi avaliado por acurácia, sensibilidade, especificidade, escore F1, kappa de Cohen e teste de McNemar. A alta foi definida como a classe positiva para todas as análises de desempenho diagnóstico.

Resultados

Foram analisados 398 pacientes no total. O ChatGPT demonstrou maior acurácia que o Gemini (87,2% vs. 66,3%), com maior sensibilidade (85,9% vs. 46,9%) e escore F1 (0,890 vs. 0,628), ao passo que o Gemini apresentou maior especificidade (96,2% vs. 89,2%). A concordância com as decisões dos clínicos foi substancial para o ChatGPT (κ = 0,737, p = 0,024) e razoável para o Gemini (κ = 0,379, p < 0,001). Marcadores laboratoriais como lactato, hemoglobina e procalcitonina diferiram significativamente entre pacientes que receberam alta e os que não receberam.

Conclusão

Grandes modelos de linguagem podem apoiar as decisões de alta da UTI quando orientados por prompts estruturados e baseados em diretrizes. O ChatGPT alcançou maior acurácia geral, sensibilidade e escore F1, enquanto o Gemini demonstrou maior especificidade.

Registro de Ensaio Clínico

Externation (Discharge) of ICU, NCT06584890, registrado em 3 de setembro de 2024, registrado prospectivamente, https://register.clinicaltrials.gov/prs/beta/studies/S000EVXZ00000029/recordSummary.

Palavras-chave

Inteligência artificial; Tomada de decisão; Sistemas de apoio à decisão clínica; Unidades de terapia intensiva; Processamento de linguagem natural; Alta do paciente

Referencias

1. Forster GM, Bihari S, Tiruvoipati R, Bailey M, Pilcher D. The Association between Discharge Delay from Intensive Care and Patient Outcomes. Am J Respir Crit Care Med. 2020;202: 1399−406.

2. Vollam S, Gustafson O, Morgan L, Pattison N, Thomas H, Watkinson P. Patient Harm and Institutional Avoidability of Out-ofHours Discharge From Intensive Care: An Analysis Using Mixed Methods*. Crit Care Med. 2022;50:1083−92.

3. Nates JL, Nunnally M, Kleinpell R, et al. ICU Admission, Discharge, and Triage Guidelines: A Framework to Enhance Clinical Operations, Development of Institutional Policies, and Further Research. Crit Care Med. 2016;44:1553−602.

4. Ruppert MM, Loftus TJ, Small C, et al. Predictive Modeling for Readmission to Intensive Care: A Systematic Review. Crit Care Explor. 2023;5:e0848.

5. Turan E, Baydemir AE, Ozcan FG, Şahin AS. Evaluating the accuracy of ChatGPT-4 in predicting ASA scores: A prospective multicentric study ChatGPT-4 in ASA score prediction. J Clin Anesth. 2024;96:111475.

6. Turan E, Baydemir AE, Bal{tatl{ AB, Sahin AS. Assessing the accuracy of ChatGPT in interpreting blood gas analysis results ChatGPT-4 in blood gas analysis. J Clin Anesth. 2025;102:111787.

7. Turan EI, Baydemir AE, Sahin AS, Ozcan FG. Effectiveness of ChatGPT-4 in predicting the human decision to send patients to the postoperative intensive care unit: a prospective multicentric study. Minerva Anestesiol. 2025;91:259−67.

8. Dost B, Turan E, Ayd{n ME, et al. Artificial Intelligence in Anaesthesiology: Current Applications, Challenges, and Future Directions. Turk J Anaesthesiol Reanim. 2025;53:282−92.

9. Plotnikoff KM, Krewulak KD, Hernandez L, et al. Patient dis- charge from intensive care: an updated scoping review to identify tools and practices to inform high-quality care. Crit Care. 2021;25:438.

10. Hiller M, Burisch C, Wittmann M, Bracht H, Kaltwasser A, Bakker J. The current state of intensive care unit discharge practices - Results of an international survey study. Front Med (Lausanne). 2024;11:1377902.

11. Hiller M, Wittmann M, Bracht H, Bakker J. Delphi study to derive expert consensus on a set of criteria to evaluate discharge readiness for adult ICU patients to be discharged to a general ward ‒ European perspective. BMC Health Serv Res. 2022;22:773.

12. You SB, Ulrich CM. Ethical considerations in evaluating discharge readiness from the intensive care unit. Nurs Ethics. 2024;31:896−906.

13. van Sluisveld N, Oerlemans A, Westert G, van der Hoeven JG, Wollersheim H, Zegers M. Barriers and facilitators to improve safety and efficiency of the ICU discharge process: a mixed methods study. BMC Health Serv Res. 2017;17:251.

14. Wu CP, Shirley RB, Milinovich A, et al. Exploring timely and safe discharge from ICU: a comparative study of machine learning predictions and clinical practices. Intensive Care Med Exp. 2025;13:10.

15. Thoral PJ, Fornasa M, de Bruin DP, et al. Explainable Machine Learning on AmsterdamUMCdb for ICU Discharge Decision Support: Uniting Intensivists and Data Scientists. Critical Care Explorations. 2021;3:e0529.

16. Temple MW, Lehmann CU, Fabbri D. Natural Language Processing for Cohort Discovery in a Discharge Prediction Model for the Neonatal ICU. Appl Clin Inform. 2016;7:101−15.

17. Loreto M, Lisboa T, Moreira VP. Early prediction of ICU readmissions using classification algorithms. Comput Biol Med. 2020;118:103636.


Submitted date:
01/08/2025

Accepted date:
16/05/2026

6a625f42a95395638577b134 rba Articles
Links & Downloads

Braz J Anesthesiol

Share this page
Page Sections