Scientific article
OA Policy
English

The detectability paradox: bilingual medical report generation with open-weight models and the limits of human oversight

Published inJournal of the American Medical Informatics Association, vol. 33, no. 7, p. 1303-1313
Publication date2026-07-01
First online date2026-05-08
Abstract

Objectives: The automation of medical report generation using large language models (LLMs) could significantly reduce physicians' documentation burden while enhancing healthcare efficiency. However, the misuse of generative artificial intelligence in medical reporting can lead to important safety risks for patients. We addressed 2 questions: (1) What is the quality of medical reports generated by LLMs in English and French? and (2) Can we distinguish between human-written and LLM-generated medical reports?

Materials and methods: We evaluated the quality of reports generated by several multilingual, open-weight LLMs using text similarity metrics on 4212 medical reports in English and French across multiple specialties. A bilingual expert panel of certified physicians (n = 4) and medical residents (n = 5) scored accuracy, fluency, and completeness of generated reports using a 1-5 Likert scale. Experts also completed a Turing-like test, blindly identifying reports as human or machine-generated.

Results: Phi-4 achieved the best overall performance (ROUGE-1: 0.70, BERTScore: 0.83). Expert evaluation confirmed high-quality reports in both languages (overall 4.6/5.0). Medical experts performed better than chance but struggled to differentiate human versus machine reports (accuracy: 0.60). Automatic classifiers showed strong performance (accuracy: 0.98).

Discussion: The high quality of LLM-generated reports supports their potential to enhance healthcare efficiency in multilingual settings. However, the discrepancy between human detection difficulty and automated detection success reveals inherent limitations in relying solely on human oversight for quality assurance and misuse prevention.

Conclusions: Deployment of LLMs for medical reporting requires combining automated detection tools with human expertise to ensure patient safety. Dataset and code: https://github.com/ds4dh/medical_report_generation.

Keywords
  • Automated report generation
  • Electronic health records (EHRs)
  • Large language models (LLMs)
  • Medical documentation
  • Multilingual medical reports
Citation (ISO format)
ROUHIZADEH, Hossein et al. The detectability paradox: bilingual medical report generation with open-weight models and the limits of human oversight. In: Journal of the American Medical Informatics Association, 2026, vol. 33, n° 7, p. 1303–1313. doi: 10.1093/jamia/ocag070
Main files (1)
Article (Published version)
Secondary files (1)
Supplemental data
accessLevelPublic
Identifiers
Journal ISSN1067-5027
3views
0downloads

Technical informations

Creation08/07/2026 15:38:15
First validation29/07/2026 08:18:53
Update29/07/2026 08:18:53
Status update29/07/2026 08:18:53
Last indexation29/07/2026 08:18:54
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack