Egocentric video description based on temporally-linked sequences

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the user. A natural topic that arises in egocentric vision is story...

Descripción completa

Detalles Bibliográficos
Autores: Bolaños Solà, Marc, Peris, Álvaro, Casacuberta, Francisco, Soler, Sergi, Radeva, Petia
Tipo de recurso: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2018
País:España
Institución:Universidad de Barcelona
Repositorio:Dipòsit Digital de la UB
OAI Identifier:oai:diposit.ub.edu:2445/143165
Acceso en línea:https://hdl.handle.net/2445/143165
Access Level:acceso abierto
Palabra clave:Aprenentatge visual
Vídeo en l'ensenyament
Visual learning
Video tapes in education
id ES_9fa23d7940c0bf10d6f188fee6db69ae
oai_identifier_str oai:diposit.ub.edu:2445/143165
network_acronym_str ES
network_name_str España
repository_id_str
spelling Egocentric video description based on temporally-linked sequencesBolaños Solà, MarcPeris, ÁlvaroCasacuberta, FranciscoSoler, SergiRadeva, PetiaAprenentatge visualVídeo en l'ensenyamentVisual learningVideo tapes in educationEgocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the user. A natural topic that arises in egocentric vision is storytelling, that is, how to understand and tell the story relying behind the pictures. In this paper, we tackle storytelling as an egocentric sequences description problem. We propose a novel methodology that exploits information from temporally neighboring events, matching precisely the nature of egocentric sequences. Furthermore, we present a new method for multimodal data fusion consisting on a multi-input attention recurrent network. We also release the EDUB-SegDesc dataset. This is the first dataset for egocentric image sequences description, consisting of 1339 events with 3991 descriptions, from 55 days acquired by 11 people. Finally, we prove that our proposal outperforms classical attentional encoder-decoder methods for video description.Elsevier2018info:eu-repo/semantics/articleinfo:eu-repo/semantics/acceptedVersionapplication/pdfhttps://hdl.handle.net/2445/143165Articles publicats en revistes (Matemàtiques i Informàtica)reponame:Dipòsit Digital de la UBinstname:Universidad de BarcelonaInglésVersió postprint del document publicat a:Journal of Visual Communication and Image Representation, 2018, vol. 50, p. 205-216cc-by-nc-nd (c) Academic Press , 2018http://creativecommons.org/licenses/by-nc-nd/3.0/esinfo:eu-repo/semantics/openAccessoai:diposit.ub.edu:2445/1431652026-05-27T06:46:51Z
dc.title.none.fl_str_mv Egocentric video description based on temporally-linked sequences
title Egocentric video description based on temporally-linked sequences
spellingShingle Egocentric video description based on temporally-linked sequences
Bolaños Solà, Marc
Aprenentatge visual
Vídeo en l'ensenyament
Visual learning
Video tapes in education
title_short Egocentric video description based on temporally-linked sequences
title_full Egocentric video description based on temporally-linked sequences
title_fullStr Egocentric video description based on temporally-linked sequences
title_full_unstemmed Egocentric video description based on temporally-linked sequences
title_sort Egocentric video description based on temporally-linked sequences
dc.creator.none.fl_str_mv Bolaños Solà, Marc
Peris, Álvaro
Casacuberta, Francisco
Soler, Sergi
Radeva, Petia
author Bolaños Solà, Marc
author_facet Bolaños Solà, Marc
Peris, Álvaro
Casacuberta, Francisco
Soler, Sergi
Radeva, Petia
author_role author
author2 Peris, Álvaro
Casacuberta, Francisco
Soler, Sergi
Radeva, Petia
author2_role author
author
author
author
dc.subject.none.fl_str_mv Aprenentatge visual
Vídeo en l'ensenyament
Visual learning
Video tapes in education
topic Aprenentatge visual
Vídeo en l'ensenyament
Visual learning
Video tapes in education
description Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the user. A natural topic that arises in egocentric vision is storytelling, that is, how to understand and tell the story relying behind the pictures. In this paper, we tackle storytelling as an egocentric sequences description problem. We propose a novel methodology that exploits information from temporally neighboring events, matching precisely the nature of egocentric sequences. Furthermore, we present a new method for multimodal data fusion consisting on a multi-input attention recurrent network. We also release the EDUB-SegDesc dataset. This is the first dataset for egocentric image sequences description, consisting of 1339 events with 3991 descriptions, from 55 days acquired by 11 people. Finally, we prove that our proposal outperforms classical attentional encoder-decoder methods for video description.
publishDate 2018
dc.date.none.fl_str_mv 2018
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/acceptedVersion
format article
status_str acceptedVersion
dc.identifier.none.fl_str_mv https://hdl.handle.net/2445/143165
url https://hdl.handle.net/2445/143165
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Versió postprint del document publicat a:
Journal of Visual Communication and Image Representation, 2018, vol. 50, p. 205-216
dc.rights.none.fl_str_mv cc-by-nc-nd (c) Academic Press , 2018
http://creativecommons.org/licenses/by-nc-nd/3.0/es
info:eu-repo/semantics/openAccess
rights_invalid_str_mv cc-by-nc-nd (c) Academic Press , 2018
http://creativecommons.org/licenses/by-nc-nd/3.0/es
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv Articles publicats en revistes (Matemàtiques i Informàtica)
reponame:Dipòsit Digital de la UB
instname:Universidad de Barcelona
instname_str Universidad de Barcelona
reponame_str Dipòsit Digital de la UB
collection Dipòsit Digital de la UB
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869414944684900352
score 15,198674