Unsupervised learning from videos using temporal coherency deep networks
In this work we address the challenging problem of unsupervised learning from videos. Existing methods utilize the spatio-temporal continuity in contiguous video frames as regularization for the learning process. Typically, this temporal coherence of close frames is used as a free form of annotation...
| Autores: | , |
|---|---|
| Tipo de recurso: | artículo |
| Fecha de publicación: | 2019 |
| País: | España |
| Institución: | Universidad de Alcalá (UAH) |
| Repositorio: | e_Buah Biblioteca Digital Universidad de Alcalá |
| Idioma: | inglés |
| OAI Identifier: | oai:ebuah.uah.es:10017/63289 |
| Acceso en línea: | http://hdl.handle.net/10017/63289 https://dx.doi.org/10.1016/j.cviu.2018.08.003 |
| Access Level: | acceso abierto |
| Palabra clave: | Unsupervised learning Action discovery Action recognition Object recognition Deep learning Robótica e Informática Industrial Robotics |
| id |
ES_fb53abb666c7a8a6269fc11b7cb20ffb |
|---|---|
| oai_identifier_str |
oai:ebuah.uah.es:10017/63289 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Unsupervised learning from videos using temporal coherency deep networksRedondo Cabrera, CarolinaLópez Sastre, Roberto Javier|||0000-0002-2477-0152Unsupervised learningAction discoveryAction recognitionObject recognitionDeep learningRobótica e Informática IndustrialRoboticsIn this work we address the challenging problem of unsupervised learning from videos. Existing methods utilize the spatio-temporal continuity in contiguous video frames as regularization for the learning process. Typically, this temporal coherence of close frames is used as a free form of annotation, encouraging the learned representations to exhibit small differences between these frames. But this type of approach fails to capture the dissimilarity between videos with different content, hence learning less discriminative features. We here propose two Siamese architectures for Convolutional Neural Networks, and their corresponding novel loss functions, to learn from unlabeled videos, which jointly exploit the local temporal coherence between contiguous frames, and a global discriminative margin used to separate representations of different videos. An extensive experimental evaluation is presented, where we validate the proposed models on various tasks. First, we show how the learned features can be used to discover actions and scenes in video collections. Second, we show the benefits of such an unsupervised learning from just unlabeled videos, which can be directly used as a prior for the supervised recognition tasks of actions and objects in images, where our results further show that our features can even surpass a traditional and heavily supervised pre-training plus fine-tuning strategy.Agencia Estatal de InvestigaciónElsevier20192019-02-01journal articlehttp://purl.org/coar/resource_type/c_6501NAhttp://purl.org/coar/version/c_be7fb7dd8ff6fe43info:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10017/63289https://dx.doi.org/10.1016/j.cviu.2018.08.003reponame:e_Buah Biblioteca Digital Universidad de Alcaláinstname:Universidad de Alcalá (UAH)InglésengAgencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2013-2016 TEC2016-80326-R PROCESADO DIGITAL Y RECONOCIMIENTO DE PATRONES PARA AYUDAS TECNICAS A LA DIVERSIDAD FUNCIONALopen accesshttp://purl.org/coar/access_right/c_abf2Attribution-NonCommercial-NoDerivatives 4.0 Internationalhttp://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccessoai:ebuah.uah.es:10017/632892026-06-18T11:13:07Z |
| dc.title.none.fl_str_mv |
Unsupervised learning from videos using temporal coherency deep networks |
| title |
Unsupervised learning from videos using temporal coherency deep networks |
| spellingShingle |
Unsupervised learning from videos using temporal coherency deep networks Redondo Cabrera, Carolina Unsupervised learning Action discovery Action recognition Object recognition Deep learning Robótica e Informática Industrial Robotics |
| title_short |
Unsupervised learning from videos using temporal coherency deep networks |
| title_full |
Unsupervised learning from videos using temporal coherency deep networks |
| title_fullStr |
Unsupervised learning from videos using temporal coherency deep networks |
| title_full_unstemmed |
Unsupervised learning from videos using temporal coherency deep networks |
| title_sort |
Unsupervised learning from videos using temporal coherency deep networks |
| dc.creator.none.fl_str_mv |
Redondo Cabrera, Carolina López Sastre, Roberto Javier|||0000-0002-2477-0152 |
| author |
Redondo Cabrera, Carolina |
| author_facet |
Redondo Cabrera, Carolina López Sastre, Roberto Javier|||0000-0002-2477-0152 |
| author_role |
author |
| author2 |
López Sastre, Roberto Javier|||0000-0002-2477-0152 |
| author2_role |
author |
| dc.subject.none.fl_str_mv |
Unsupervised learning Action discovery Action recognition Object recognition Deep learning Robótica e Informática Industrial Robotics |
| topic |
Unsupervised learning Action discovery Action recognition Object recognition Deep learning Robótica e Informática Industrial Robotics |
| description |
In this work we address the challenging problem of unsupervised learning from videos. Existing methods utilize the spatio-temporal continuity in contiguous video frames as regularization for the learning process. Typically, this temporal coherence of close frames is used as a free form of annotation, encouraging the learned representations to exhibit small differences between these frames. But this type of approach fails to capture the dissimilarity between videos with different content, hence learning less discriminative features. We here propose two Siamese architectures for Convolutional Neural Networks, and their corresponding novel loss functions, to learn from unlabeled videos, which jointly exploit the local temporal coherence between contiguous frames, and a global discriminative margin used to separate representations of different videos. An extensive experimental evaluation is presented, where we validate the proposed models on various tasks. First, we show how the learned features can be used to discover actions and scenes in video collections. Second, we show the benefits of such an unsupervised learning from just unlabeled videos, which can be directly used as a prior for the supervised recognition tasks of actions and objects in images, where our results further show that our features can even surpass a traditional and heavily supervised pre-training plus fine-tuning strategy. |
| publishDate |
2019 |
| dc.date.none.fl_str_mv |
2019 2019-02-01 |
| dc.type.none.fl_str_mv |
journal article http://purl.org/coar/resource_type/c_6501 NA http://purl.org/coar/version/c_be7fb7dd8ff6fe43 |
| dc.type.openaire.fl_str_mv |
info:eu-repo/semantics/article |
| format |
article |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10017/63289 https://dx.doi.org/10.1016/j.cviu.2018.08.003 |
| url |
http://hdl.handle.net/10017/63289 https://dx.doi.org/10.1016/j.cviu.2018.08.003 |
| dc.language.none.fl_str_mv |
Inglés eng |
| language_invalid_str_mv |
Inglés |
| language |
eng |
| dc.relation.none.fl_str_mv |
Agencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2013-2016 TEC2016-80326-R PROCESADO DIGITAL Y RECONOCIMIENTO DE PATRONES PARA AYUDAS TECNICAS A LA DIVERSIDAD FUNCIONAL |
| dc.rights.none.fl_str_mv |
open access http://purl.org/coar/access_right/c_abf2 Attribution-NonCommercial-NoDerivatives 4.0 International http://creativecommons.org/licenses/by-nc-nd/4.0/ |
| dc.rights.openaire.fl_str_mv |
info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
open access http://purl.org/coar/access_right/c_abf2 Attribution-NonCommercial-NoDerivatives 4.0 International http://creativecommons.org/licenses/by-nc-nd/4.0/ |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf |
| dc.publisher.none.fl_str_mv |
Elsevier |
| publisher.none.fl_str_mv |
Elsevier |
| dc.source.none.fl_str_mv |
reponame:e_Buah Biblioteca Digital Universidad de Alcalá instname:Universidad de Alcalá (UAH) |
| instname_str |
Universidad de Alcalá (UAH) |
| reponame_str |
e_Buah Biblioteca Digital Universidad de Alcalá |
| collection |
e_Buah Biblioteca Digital Universidad de Alcalá |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869425318692913152 |
| score |
15,812455 |