Scientific article
Review
OA Policy
English

Advancing Knowledge in Evaluating the Clinical Impact of Large Language Models for Clinical Text Summarization : A Narrative Review

Published inStudies in health technology and informatics, vol. 336, p. 610-614
Publication date2026-05-21
Abstract

Large language models (LLMs) are increasingly explored for clinical text summarization from electronic health records (EHRs), where they could reduce documentation burden. While promising, errors in generated summaries may propagate misinformation, misrepresent clinical facts, and ultimately pose risks to patient safety. This review examined how the clinical impact of LLM-based summarization has been evaluated across four dimensions: utility, failure, patient safety risks, and bias. Literature was retrieved from PubMed between June 2024 and September 2025. Of the 144 retrieved studies, 32 studies were included in the analysis. Failure analysis was the most frequently reported (n=16, 50%), focusing on inaccuracies, omissions, and hallucinations, though definitions and methods varied widely. Utility analysis (n=8, 25%) examined workflow efficiency, readability, and understandability. Bias analysis (n=7, 22%) considered gender, race, and stigmatizing language. Only one study explicitly conducted patient safety risk analysis (n=1, 3%), using risk matrices. Findings indicate that evaluation frameworks are gradually moving beyond technical performance metrics; however, methodologies remain heterogeneous and lack standardization. Future research should establish consistent error taxonomies, structured risk assessment frameworks, and systematic bias detection methods to enable the safe clinical deployment of LLM-based summarization systems.

Keywords
  • Electronic Health Record (EHR)
  • Generative Artificial Intelligence
  • Healthcare
  • Large Language Models (LLMs)
  • Electronic Health Records / organization & administration
  • Humans
  • Natural Language Processing
  • Patient Safety
  • Large Language Models
Citation (ISO format)
BEDNARCZYK, Lydie et al. Advancing Knowledge in Evaluating the Clinical Impact of Large Language Models for Clinical Text Summarization : A Narrative Review. In: Studies in health technology and informatics, 2026, vol. 336, p. 610–614. doi: 10.3233/SHTI260243
Main files (1)
Article (Published version)
Identifiers
Additional URL for this publicationhttps://ebooks.iospress.nl/doi/10.3233/SHTI260243
Journal ISSN0926-9630
6views
12downloads

Technical informations

Creation01/06/2026 08:52:36
First validation01/07/2026 09:29:43
Update01/07/2026 09:29:43
Status update01/07/2026 09:29:43
Last indexation01/07/2026 09:29:44
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack