Algoritmo de segmentación de habla independiente de texto en uno y dos niveles

Success in the performance of automatic speech recognition depends, among other issues, from an accurate segmentation of the input signal. Such signal may be divided by words, vowels or phonemes, the last being the most popular. Segmentation may be achieved using different techniques, some restricte...

Descripción completa

Detalles Bibliográficos
Autor: RICARDO SANCHEZ JURADO
Tipo de recurso: tesis de maestría
Estado:Versión aceptada para publicación
Fecha de publicación:2008
País:México
Institución:Instituto Nacional de Astrofísica, Óptica y Electrónica
Repositorio:Repositorio Institucional del INAOE
Idioma:español
OAI Identifier:oai:inaoe.repositorioinstitucional.mx:1009/563
Acceso en línea:http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/563
Access Level:acceso abierto
Palabra clave:info:eu-repo/classification/Procesamiento de voz/Speech processing
info:eu-repo/classification/Reconocimiento de voz/Speech recognition
info:eu-repo/classification/Extracción de características/Feature extraction
info:eu-repo/classification/cti/7
info:eu-repo/classification/cti/33
info:eu-repo/classification/cti/3304
info:eu-repo/classification/cti/120308
Descripción
Sumario:Success in the performance of automatic speech recognition depends, among other issues, from an accurate segmentation of the input signal. Such signal may be divided by words, vowels or phonemes, the last being the most popular. Segmentation may be achieved using different techniques, some restricted by text or speaker and others free of restrictions. In this research we present a text and speaker-independent algorithm to obtain phonetic boundaries of a speech signal, using only acoustic features. The signal is divided into segments, called frames, small enough to be handled by coding algorithms as Mel Filter Banks or stationary wavelet transforms. Each feature is converted to a fuzzy representation in order to detect transitions among phonemes that, in other way, could not be clearly identified. In addition, we propose a modification in Euclidian and Chebishev distances to calculate feature distances using four adjacent frames. New strategies to select candidates for boundaries in one and two levels are also presented and analyzed. Genetic algorithms are used to optimize some parameters in the proposed algorithm. The algorithm was tested using two different corpuses, one in English and one in Spanish language. A correct segmentation of 80.28% was obtained for English and 82.58% for Spanish. This performance is similar to results obtained by other research works using English language.