Proceedings chapter
OA Policy
English

Using Source-Language Transformations to Address Register Mismatches in SMT

Presented atSan Diego (California, USA), Oct 28-Nov 1
Publication date2012
Abstract

Mismatches between training and test data are a ubiquitous problem for real SMT applica- tions. In this paper, we examine a type of mismatch that commonly arises when translat- ing from French and similar languages: avail- able training data is mostly formal register, but test data may well be informal register. We consider methods for defining surface trans- formations that map common informal lan- guage constructions into their formal language counterparts, or vice versa; we then describe two ways to use these mappings, either to cre- ate artificial training data or to pre-process source text at run-time. An initial evalua- tion performed using crowd-sourced compar- isons of alternate translations produced by a French-to-English SMT system suggests that both methods can improve performance, with run-time pre-processing being the more effec- tive of the two.

Research groups
Citation (ISO format)
RAYNER, Emmanuel, BOUILLON, Pierrette, HADDOW, Barry. Using Source-Language Transformations to Address Register Mismatches in SMT. In: Proceedings of the tenth biennial conference of the Association for Machine Translation in the Americas (AMTA-2012). San Diego (California, USA). [s.l.] : [s.n.], 2012.
Main files (1)
Proceedings chapter (Accepted version)
accessLevelPublic
Identifiers
  • PID : unige:30919
623views
343downloads

Technical informations

Creation04/11/2013 18:56:00
First validation04/11/2013 18:56:00
Update14/03/2023 20:35:52
Status update14/03/2023 20:35:52
Last indexation17/12/2024 15:35:26
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack