Scientific article
OA Policy
English

HealthContradict : Evaluating biomedical knowledge conflicts in language models

Published innpj digital medicine, vol. 9, no. 1, 152
Publication date2026-01-21
Abstract

How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to reason over long, conflicting biomedical contexts using HealthContradict, an expert-verified dataset comprising 920 unique instances, each consisting of a health-related question, a factual answer supported by scientific evidence, and two documents presenting contradictory stances. We consider several prompt settings, including correct, incorrect or contradictory context, and measure their impact on model outputs. Compared to existing medical question-answering evaluation benchmarks, HealthContradict provides greater distinctions of language models' contextual reasoning capabilities. Our experiments show that the strength of fine-tuned biomedical language models lies not only in their parametric knowledge from pretraining, but also in their ability to exploit correct context while resisting incorrect context.

Citation (ISO format)
ZHANG, Boya et al. HealthContradict : Evaluating biomedical knowledge conflicts in language models. In: npj digital medicine, 2026, vol. 9, n° 1, p. 152. doi: 10.1038/s41746-025-02336-0
Main files (1)
Article (Published version)
Secondary files (1)
Supplemental data
accessLevelPublic
Identifiers
Additional URL for this publicationhttps://www.nature.com/articles/s41746-025-02336-0
Journal ISSN2398-6352
30views
119downloads

Technical informations

Creation28/02/2026 09:17:17
First validation03/03/2026 09:13:24
Update05/03/2026 13:15:07
Status update05/03/2026 13:15:07
Last indexation05/03/2026 13:15:14
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack