Proceedings chapter
OA Policy
English

Medical Dialogue Audio Transcription : Dataset and Benchmarking of ASR Models

Presented atFortaleza, CE (Brazil), September 29th - October 2nd, 2025
PublisherPorto Alegre, RS : SBC
First online date2025-09-29
Abstract

The development of Automatic Speech Recognition (ASR) technologies for healthcare applications is hindered by the limited availability of publicly accessible speech corpora that reflect both natural medical dialogues and the acoustic conditions typically found in clinical environments. In this study, we present the creation and characterization of MedDialogue-Audio, a new synthetic English-language corpus designed to address this gap. The dataset was derived from the MedDialog-EN transcription set and enriched through a multi-stage processing pipeline that involved text normalization with a large language model, speech synthesis, and the controlled addition of both white noise and hospital ambient sounds. We provide descriptive statistics for the corpus, which comprises more than 10,000 dialogues, as well as benchmarking results from leading ASR models. The experiments assess transcription performance across varying signal-to-noise ratios and establish baseline metrics to support future research in this field.

Keywords
  • Audio
  • Automatic Speech Recognition
  • Medical Dataset
  • Text-to-Speech
Citation (ISO format)
GASSENN, Aline E. et al. Medical Dialogue Audio Transcription : Dataset and Benchmarking of ASR Models. In: Proceedings of the 7th Dataset Showcase Workshop (DSW 2025). Fortaleza, CE (Brazil). Porto Alegre, RS : SBC, 2025. p. 71–82. doi: 10.5753/dsw.2025.248010
Main files (1)
Proceedings chapter (Published version)
accessLevelPublic
Identifiers
11views
333downloads

Technical informations

Creation12/03/2026 08:53:02
First validation23/03/2026 08:08:04
Update23/03/2026 08:08:04
Status update23/03/2026 08:08:04
Last indexation23/03/2026 08:08:05
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack