Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis
Speaker diarization is traditionally defined as the problem of determining “who speaks when” given an audio or video stream. This is an important task in many applications for meeting rooms, including automatic transcription of conversations, camera steering or content summarization. When the room i...
| Autores: | , , , , |
|---|---|
| Tipo de documento: | artigo |
| Estado: | Versão publicada |
| Data de publicação: | 2018 |
| País: | España |
| Recursos: | Universidad de Jaén |
| Repositório: | RUJA. Repositorio Institucional de la Producción Científica de la Universidad de Jaén |
| OAI Identifier: | oai:ruja.ujaen.es:10953/2188 |
| Acesso em linha: | https://hdl.handle.net/10953/2188 |
| Access Level: | Acceso aberto |
| Palavra-chave: | Speaker diarization Meeting rooms SRP-PHAT Multimodal processing 621.39 |
| id |
ES_bd17ce5fd86df71a72df26dede21e659 |
|---|---|
| oai_identifier_str |
oai:ruja.ujaen.es:10953/2188 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysisCabañas-Molero, Pablo AntonioLucena, ManuelFuertes, José ManuelVera-Candeas, PedroRuiz-Reyes, NicolásSpeaker diarizationMeeting roomsSRP-PHATMultimodal processing621.39Speaker diarization is traditionally defined as the problem of determining “who speaks when” given an audio or video stream. This is an important task in many applications for meeting rooms, including automatic transcription of conversations, camera steering or content summarization. When the room is equipped with microphone arrays and cameras, speakers can be distinguished according to their location and the problem can be addressed through localization techniques. This article proposes a multimodal speaker diarization system for meeting environments based on a modified SRP-PHAT function evaluated on space volumes rather than discrete points. In our system, this function is used in combination with a circular array, enabling audio-based localization based on the selection of local maxima. Voicing detection is used to detect speech frames, whereas video analysis is introduced to aid in the decision when users move or simultaneously speak. The approach is evaluated on the well-known AMI dataset with approximately 100 hours of realistic meeting recordings and shows an average diarization error rate of 21% – 25%.This work was supported by the Andalusian Economy and Knowledge Council under project 2010-TIC6762, and the Spanish Ministry of Economy and Competitiveness under project TEC2015-67387-C4-2-R.Springer202420242018info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfhttps://hdl.handle.net/10953/2188reponame:RUJA. Repositorio Institucional de la Producción Científica de la Universidad de Jaéninstname:Universidad de JaénInglésMultimedia Tools and Applications 2018; 77, 27685–27707CC0 1.0 Universalhttp://creativecommons.org/publicdomain/zero/1.0/info:eu-repo/semantics/openAccessoai:ruja.ujaen.es:10953/21882026-06-24T12:41:07Z |
| dc.title.none.fl_str_mv |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| title |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| spellingShingle |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis Cabañas-Molero, Pablo Antonio Speaker diarization Meeting rooms SRP-PHAT Multimodal processing 621.39 |
| title_short |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| title_full |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| title_fullStr |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| title_full_unstemmed |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| title_sort |
Multimodal speaker diarization for meetings using volume-evaluated SRP-PHAT and video analysis |
| dc.creator.none.fl_str_mv |
Cabañas-Molero, Pablo Antonio Lucena, Manuel Fuertes, José Manuel Vera-Candeas, Pedro Ruiz-Reyes, Nicolás |
| author |
Cabañas-Molero, Pablo Antonio |
| author_facet |
Cabañas-Molero, Pablo Antonio Lucena, Manuel Fuertes, José Manuel Vera-Candeas, Pedro Ruiz-Reyes, Nicolás |
| author_role |
author |
| author2 |
Lucena, Manuel Fuertes, José Manuel Vera-Candeas, Pedro Ruiz-Reyes, Nicolás |
| author2_role |
author author author author |
| dc.subject.none.fl_str_mv |
Speaker diarization Meeting rooms SRP-PHAT Multimodal processing 621.39 |
| topic |
Speaker diarization Meeting rooms SRP-PHAT Multimodal processing 621.39 |
| description |
Speaker diarization is traditionally defined as the problem of determining “who speaks when” given an audio or video stream. This is an important task in many applications for meeting rooms, including automatic transcription of conversations, camera steering or content summarization. When the room is equipped with microphone arrays and cameras, speakers can be distinguished according to their location and the problem can be addressed through localization techniques. This article proposes a multimodal speaker diarization system for meeting environments based on a modified SRP-PHAT function evaluated on space volumes rather than discrete points. In our system, this function is used in combination with a circular array, enabling audio-based localization based on the selection of local maxima. Voicing detection is used to detect speech frames, whereas video analysis is introduced to aid in the decision when users move or simultaneously speak. The approach is evaluated on the well-known AMI dataset with approximately 100 hours of realistic meeting recordings and shows an average diarization error rate of 21% – 25%. |
| publishDate |
2018 |
| dc.date.none.fl_str_mv |
2018 2024 2024 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/publishedVersion |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.none.fl_str_mv |
https://hdl.handle.net/10953/2188 |
| url |
https://hdl.handle.net/10953/2188 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
Multimedia Tools and Applications 2018; 77, 27685–27707 |
| dc.rights.none.fl_str_mv |
CC0 1.0 Universal http://creativecommons.org/publicdomain/zero/1.0/ info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
CC0 1.0 Universal http://creativecommons.org/publicdomain/zero/1.0/ |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf |
| dc.publisher.none.fl_str_mv |
Springer |
| publisher.none.fl_str_mv |
Springer |
| dc.source.none.fl_str_mv |
reponame:RUJA. Repositorio Institucional de la Producción Científica de la Universidad de Jaén instname:Universidad de Jaén |
| instname_str |
Universidad de Jaén |
| reponame_str |
RUJA. Repositorio Institucional de la Producción Científica de la Universidad de Jaén |
| collection |
RUJA. Repositorio Institucional de la Producción Científica de la Universidad de Jaén |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869418171121795072 |
| score |
15,812429 |