Scientific article
English

Item-Level Evaluation of Multimodal Large Language Models in Neuroradiology : Generational Performance and Execution Variability

Published inAmerican journal of neuroradiology, ajnr.A9477
First online date2026-06-13
Abstract

Background: Multimodal large language models have demonstrated consistent generational improvements on medical benchmark tasks, including radiology applications. However, whether these gains represent meaningful convergence toward the expert reference performance in subspecialty imaging domains such as neuroradiology remains uncertain, particularly when evaluated using item-level, human-referenced designs.

Methods: We compared expert neuroradiologists, radiology residents, and four vision-capable large language models (GPT-4, GPT-5, Gemini 1.5, Gemini 2.5) using 106 image-based neuroradiology multiple-choice questions derived from Radiopaedia. Analyses were conducted at the question (item) level, preserving within-question pairing across groups. Mean accuracy differences were estimated using non-parametric bootstrap 95% confidence intervals, and statistical inference was performed using paired permutation tests restricted to pre-specified contrasts with false discovery rate correction for primary comparisons. Repeated model executions were analyzed to characterize execution-level variability. Performance was further contextualized against Radiopaedia community accuracy as a community-level reference.

Results: Expert neuroradiologists achieved the highest mean item-level accuracy (0.915; 95% confidence interval, 0.877-0.953). Second-generation models demonstrated improved mean accuracy relative to earlier versions and approximated or exceeded resident-level performance in selected comparisons. However, a substantial gap relative to the expert reference persisted. GPT-5 and Gemini 2.5 underperformed the expert reference by mean per-item differences of -0.236 and -0.217, respectively. When contextualized against the community-level reference, advanced models aligned more closely with aggregate learner performance than with the expert reference accuracy. Improvements in mean accuracy were not uniformly accompanied by improved execution consistency.

Conclusions: Although multimodal large language models show meaningful generational gains on neuroradiology tasks, these improvements do not constitute convergence toward the expert reference performance. Item-level, paired, human-referenced evaluation provides critical context for interpreting benchmark performance and helps distinguish apparent performance gains from true alignment with the expert reference. Importantly, relative stability in aggregate accuracy does not necessarily imply reliability at the level of individual decisions, underscoring the need to assess both performance magnitude and execution-level consistency.

Keywords
  • Multimodal large language models
  • Neuroradiology
  • Diagnostic accuracy
  • Inter-run variability
Citation (ISO format)
OJEDA ESPARZA, Jose Federico et al. Item-Level Evaluation of Multimodal Large Language Models in Neuroradiology : Generational Performance and Execution Variability. In: American journal of neuroradiology, 2026, p. ajnr.A9477. doi: 10.3174/ajnr.A9477
Main files (1)
Article (Submitted version)
accessLevelRestricted
Identifiers
Additional URL for this publicationhttp://www.ajnr.org/lookup/doi/10.3174/ajnr.A9477
Journal ISSN0195-6108
1views
0downloads

Technical informations

Creation19/08/2026 10:15:44
First validation28/08/2026 07:31:12
Update28/08/2026 07:31:12
Status update28/08/2026 07:31:12
Last indexation28/08/2026 07:31:13
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack