Scientific article
OA Policy
English

Reliable Enough ? Benchmarking LLMs for Clinical Concept Extraction

Published inStudies in health technology and informatics, vol. 336, p. 786-787
Publication date2026-05-21
Abstract

Language models (LMs) offer promise for automating concept extraction, a task traditionally reliant on manual curation and earlier natural language processing (NLP) methods. We present a benchmarking approach to systematically evaluate on-premises LM performance. Using identification of first-line pharmacological treatments in melanoma patients as a test case, we demonstrate how this method supports structured comparison and error analysis for local models. Results indicate the importance of prompt design, and that small models struggle with layered reasoning tasks, such as sequencing interventions. These findings suggest that LMs are best deployed as supportive tools requiring careful evaluation.

Keywords
  • Benchmarking
  • Computational Linguistics
  • Medical Informatics
  • Responsible AI
  • Natural Language Processing
  • Benchmarking / methods
  • Humans
  • Melanoma / drug therapy
  • Data Mining / methods
  • Data Mining / standards
  • Reproducibility of Results
  • Electronic Health Records
Citation (ISO format)
PIGNAT, Johann et al. Reliable Enough ? Benchmarking LLMs for Clinical Concept Extraction. In: Studies in health technology and informatics, 2026, vol. 336, p. 786–787. doi: 10.3233/SHTI260285
Main files (1)
Article (Published version)
Identifiers
Journal ISSN0926-9630
7views
24downloads

Technical informations

Creation02/06/2026 15:35:19
First validation06/07/2026 07:52:04
Update06/07/2026 07:52:04
Status update06/07/2026 07:52:04
Last indexation06/07/2026 07:52:05
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack