The Role of human reference translation in machine translation evaluation
Both manual and automatic methods for Machine Translation (MT) evaluation heavily rely on professional human translation. In manual evaluation, human translation is often used instead of the source text in order to avoid the need for bilingual speakers, whereas the majority of automatic evaluation t...
| Autor: | |
|---|---|
| Tipo de recurso: | tesis doctoral |
| Estado: | Versión publicada |
| Fecha de publicación: | 2017 |
| País: | España |
| Institución: | CBUC, CESCA |
| Repositorio: | TDR. Tesis Doctorales en Red |
| OAI Identifier: | oai:www.tdx.cat:10803/404987 |
| Acceso en línea: | http://hdl.handle.net/10803/404987 |
| Access Level: | acceso abierto |
| Palabra clave: | Machine translation Statistical machine translation Machine translation evaluation Machine translation quality Quality estimation Automatic evaluation Monolingual alignment Translation errors Distributional similarity Translation studies Translation shifts Translationese Translation equivalence Traducción automática Traducción automática estadística Evaluación de la traducción automática Calidad de la traducción automática Estimación de calidad Alineamiento monolingüe Errores de traducción Similitud distribucional Estudios de traducción Evaluación automática Equivalencia traductora 81 |
| Sumario: | Both manual and automatic methods for Machine Translation (MT) evaluation heavily rely on professional human translation. In manual evaluation, human translation is often used instead of the source text in order to avoid the need for bilingual speakers, whereas the majority of automatic evaluation techniques measure string similarity between MT output and a human translation (commonly referred to as candidate and reference translations), assuming that the closer they are, the higher the MT quality. In spite of the crucial role of human reference translation in the assessment of MT quality, its fundamental characteristics have been largely disregarded. An inherent property of professional translation is the adaptation of the original text to the expectations of the target audience. As a consequence, human translation can be rather different from the original text, which, as will be shown throughout this work, has a strong impact on the results of MT evaluation. The first goal of our research was to assess the effects of using human translation as a benchmark for MT evaluation. To achieve this goal, we started with a theoretical discussion of the relation between original and translated texts. We identified the presence of optional translation shifts as one of the fundamental characteristics of human translation. We analyzed the impact of translation shifts on automatic and manual MT evaluation showing that in both cases quality assessment is strongly biased by the reference provided. The second goal of our work was to improve the accuracy of automatic evaluation in terms of the correlation with human judgments. Given the limitations of reference-based evaluation discussed in the first part of the work, instead of considering different aspects of similarity we focused on the differences between MT output and reference translation searching for criteria that would allow distinguishing between acceptable linguistic variation and deviations induced by MT errors. In the first place, we explored the use of local syntactic context for validating the matches between candidate and reference words. In the second place, to compensate for the lack of information regarding the MT segments for which no counterpart in the reference translation was found, we enhanced reference-based evaluation with fluency-oriented features. We implemented our approach as a family of automatic evaluation metrics that showed highly competitive performance in a series of well-known MT evaluation campaigns. |
|---|