Algoritmo de segmentación de habla independiente de texto en uno y dos niveles
Success in the performance of automatic speech recognition depends, among other issues, from an accurate segmentation of the input signal. Such signal may be divided by words, vowels or phonemes, the last being the most popular. Segmentation may be achieved using different techniques, some restricte...
| Autor: | |
|---|---|
| Tipo de recurso: | tesis de maestría |
| Estado: | Versión aceptada para publicación |
| Fecha de publicación: | 2008 |
| País: | México |
| Institución: | Instituto Nacional de Astrofísica, Óptica y Electrónica |
| Repositorio: | Repositorio Institucional del INAOE |
| Idioma: | español |
| OAI Identifier: | oai:inaoe.repositorioinstitucional.mx:1009/563 |
| Acceso en línea: | http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/563 |
| Access Level: | acceso abierto |
| Palabra clave: | info:eu-repo/classification/Procesamiento de voz/Speech processing info:eu-repo/classification/Reconocimiento de voz/Speech recognition info:eu-repo/classification/Extracción de características/Feature extraction info:eu-repo/classification/cti/7 info:eu-repo/classification/cti/33 info:eu-repo/classification/cti/3304 info:eu-repo/classification/cti/120308 |
| Sumario: | Success in the performance of automatic speech recognition depends, among other issues, from an accurate segmentation of the input signal. Such signal may be divided by words, vowels or phonemes, the last being the most popular. Segmentation may be achieved using different techniques, some restricted by text or speaker and others free of restrictions. In this research we present a text and speaker-independent algorithm to obtain phonetic boundaries of a speech signal, using only acoustic features. The signal is divided into segments, called frames, small enough to be handled by coding algorithms as Mel Filter Banks or stationary wavelet transforms. Each feature is converted to a fuzzy representation in order to detect transitions among phonemes that, in other way, could not be clearly identified. In addition, we propose a modification in Euclidian and Chebishev distances to calculate feature distances using four adjacent frames. New strategies to select candidates for boundaries in one and two levels are also presented and analyzed. Genetic algorithms are used to optimize some parameters in the proposed algorithm. The algorithm was tested using two different corpuses, one in English and one in Spanish language. A correct segmentation of 80.28% was obtained for English and 82.58% for Spanish. This performance is similar to results obtained by other research works using English language. |
|---|