Processing a learner corpus to identify differences: the influence of task, genre and student background

This master thesis deals with the technical and methodological aspects in creating, cleaning and processing a Brazilian university level learner corpus, the Corpus do Inglês sem Fronteiras (CorIsF) v 1.0. The two main goals of this study consist of making the processing of CorIsF replicable and in i...

Descripción completa

Detalles Bibliográficos
Autor: Andressa Rodrigues Gomide
Tipo de recurso: tesis de maestría
Estado:Versión publicada
Fecha de publicación:2016
País:Brasil
Institución:Universidade Federal de Minas Gerais (UFMG)
Repositorio:Repositório Institucional da UFMG
Idioma:portugués
OAI Identifier:oai:repositorio.ufmg.br:1843/MGSS-A9KGY5
Acceso en línea:http://hdl.handle.net/1843/MGSS-A9KGY5
Access Level:acceso abierto
Palabra clave:Inglês para fins acadêmicos
Corpus de aprendiz
Desenho de corpus
Língua inglesa Estudo e ensino Falantes de português Brasil
Língua inglesa Estudo e ensino Falantes estrangeiros
Lingüística textual
Aquisição da segunda linguagem
Lingua inglesa Gramatica
Linguística de corpus
Descripción
Sumario:This master thesis deals with the technical and methodological aspects in creating, cleaning and processing a Brazilian university level learner corpus, the Corpus do Inglês sem Fronteiras (CorIsF) v 1.0. The two main goals of this study consist of making the processing of CorIsF replicable and in investigating and describing the variation of some linguistic characteristics across different learner groups, tasks andgenres. The procedure was carried in R, a free software environment for statistical computing and graphics, and was divided in four parts: dataset compilation and preprocessing; dataset processing; extraction of the key features; and data visualization. The first step deals with the method used to collect the data and to do the first cleaning process, such as eliminating unwanted data and keeping the relevant ones. In the following step, CorIsF was subset in five small corpora covering different learner profiles, two different tasks, and on genre, and annotated with a part-ofspeech (POS) tagger. In the third step the variability of POS within subcorpora, the frequency of types and tokens, and the usage of n-grams were investigated. In the final step some exploratory data visualization were performed with the creation and analysis of plots and wordclouds. After the preparation of the data, the language used in each subcorpora was contrasted and analysed, suggesting that task, genre and student background are likely to influence learners written production.