The impact of tokenization on gender bias in Machine Translation
Treball de fi de màster en Lingüística Teòrica i Aplicada. Directora: Dra. Maite Melero Nogues
| Autor: | |
|---|---|
| Tipo de recurso: | tesis de maestría |
| Fecha de publicación: | 2023 |
| País: | España |
| Institución: | Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| Repositorio: | Recercat. Dipósit de la Recerca de Catalunya |
| OAI Identifier: | oai:recercat.cat:10230/57998 |
| Acceso en línea: | http://hdl.handle.net/10230/57998 |
| Access Level: | acceso abierto |
| Palabra clave: | Machine Translation Neural Machine Translation Sub-word tokenization Gender bias Unigram BPE (Byte Pair Encoding) Character-based tokenization Morfessor |
| id |
ES_bda5f672b155de20ecd569e449fe246a |
|---|---|
| oai_identifier_str |
oai:recercat.cat:10230/57998 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
The impact of tokenization on gender bias in Machine TranslationMash, AudreyMachine TranslationNeural Machine TranslationSub-word tokenizationGender biasUnigramBPE (Byte Pair Encoding)Character-based tokenizationMorfessorTreball de fi de màster en Lingüística Teòrica i Aplicada. Directora: Dra. Maite Melero NoguesThis study examines the impact of tokenization methods on gender bias in Neural Machine Translation (NMT). Unigram, BPE, Character, and Morfessor tokenization approaches are compared in terms of translation quality measured by BLEU scores and gender accuracy. Results show that Unigram achieves the highest BLEU scores, closely followed by BPE and Morfessor, while Character performs lower. However, all models display a bias towards generating masculine forms more frequently than feminine forms in gender accuracy analysis. They also overwhelming generate masculine forms when no context is provided. The Unigram method exhibits the highest accuracy for both feminine and masculine forms, surpassing BPE and Morfessor. These findings emphasize the need to address gender bias in MT systems and the complex relationship between tokenization methods, translation quality, and gender accuracy. Further research is warranted to explore additional factors influencing gender bias. This study contributes to the development of inclusive and unbiased translation technologies.202320232023info:eu-repo/semantics/masterThesisapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/57998reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésLlicència CC Reconeixement-NoComercial-SenseObraDerivada 4.0 Internacional (CC BY-NC-ND 4.0)https://creativecommons.org/licenses/by-nc-nd/4.0/deed.cainfo:eu-repo/semantics/openAccessoai:recercat.cat:10230/579982026-05-29T05:05:01Z |
| dc.title.none.fl_str_mv |
The impact of tokenization on gender bias in Machine Translation |
| title |
The impact of tokenization on gender bias in Machine Translation |
| spellingShingle |
The impact of tokenization on gender bias in Machine Translation Mash, Audrey Machine Translation Neural Machine Translation Sub-word tokenization Gender bias Unigram BPE (Byte Pair Encoding) Character-based tokenization Morfessor |
| title_short |
The impact of tokenization on gender bias in Machine Translation |
| title_full |
The impact of tokenization on gender bias in Machine Translation |
| title_fullStr |
The impact of tokenization on gender bias in Machine Translation |
| title_full_unstemmed |
The impact of tokenization on gender bias in Machine Translation |
| title_sort |
The impact of tokenization on gender bias in Machine Translation |
| dc.creator.none.fl_str_mv |
Mash, Audrey |
| author |
Mash, Audrey |
| author_facet |
Mash, Audrey |
| author_role |
author |
| dc.subject.none.fl_str_mv |
Machine Translation Neural Machine Translation Sub-word tokenization Gender bias Unigram BPE (Byte Pair Encoding) Character-based tokenization Morfessor |
| topic |
Machine Translation Neural Machine Translation Sub-word tokenization Gender bias Unigram BPE (Byte Pair Encoding) Character-based tokenization Morfessor |
| description |
Treball de fi de màster en Lingüística Teòrica i Aplicada. Directora: Dra. Maite Melero Nogues |
| publishDate |
2023 |
| dc.date.none.fl_str_mv |
2023 2023 2023 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/masterThesis |
| format |
masterThesis |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10230/57998 |
| url |
http://hdl.handle.net/10230/57998 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.rights.none.fl_str_mv |
Llicència CC Reconeixement-NoComercial-SenseObraDerivada 4.0 Internacional (CC BY-NC-ND 4.0) https://creativecommons.org/licenses/by-nc-nd/4.0/deed.ca info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
Llicència CC Reconeixement-NoComercial-SenseObraDerivada 4.0 Internacional (CC BY-NC-ND 4.0) https://creativecommons.org/licenses/by-nc-nd/4.0/deed.ca |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf application/pdf |
| dc.source.none.fl_str_mv |
reponame:Recercat. Dipósit de la Recerca de Catalunya instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| instname_str |
Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| reponame_str |
Recercat. Dipósit de la Recerca de Catalunya |
| collection |
Recercat. Dipósit de la Recerca de Catalunya |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869418218192371712 |
| score |
15,198674 |