Scene Classification Using a Hybrid Generative/Discriminative Approach

We investigate whether dimensionality reduction using a latent generative model is beneficial for the task of weakly supervised scene classification. In detail, we are given a set of labeled images of scenes (for example, coast, forest, city, river, etc.), and our objective is to classify a new imag...

Descripción completa

Detalles Bibliográficos
Autores: Bosch Rué, Anna, Zisserman, Andrew, Muñoz Pujol, Xavier
Tipo de recurso: artículo
Fecha de publicación:2008
País:España
Institución:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
Repositorio:Recercat. Dipósit de la Recerca de Catalunya
OAI Identifier:oai:recercat.cat:10256/2317
Acceso en línea:http://hdl.handle.net/10256/2317
Access Level:acceso abierto
Palabra clave:Algorismes computacionals
Discriminació visual
Imatges -- Processament
Reconeixement de formes (Informàtica)
Vídeo digital
Computer algorithms
Digital video
Image processing
Pattern recognition systems
Visual discrimination
id ES_4f79e07f8006e35ecfb5b23b6c8204ca
oai_identifier_str oai:recercat.cat:10256/2317
network_acronym_str ES
network_name_str España
repository_id_str
spelling Scene Classification Using a Hybrid Generative/Discriminative ApproachBosch Rué, AnnaZisserman, AndrewMuñoz Pujol, XavierAlgorismes computacionalsDiscriminació visualImatges -- ProcessamentReconeixement de formes (Informàtica)Vídeo digitalComputer algorithmsDigital videoImage processingPattern recognition systemsVisual discriminationWe investigate whether dimensionality reduction using a latent generative model is beneficial for the task of weakly supervised scene classification. In detail, we are given a set of labeled images of scenes (for example, coast, forest, city, river, etc.), and our objective is to classify a new image into one of these categories. Our approach consists of first discovering latent ";topics"; using probabilistic Latent Semantic Analysis (pLSA), a generative model from the statistical text literature here applied to a bag of visual words representation for each image, and subsequently, training a multiway classifier on the topic distribution vector for each image. We compare this approach to that of representing each image by a bag of visual words vector directly and training a multiway classifier on these vectors. To this end, we introduce a novel vocabulary using dense color SIFT descriptors and then investigate the classification performance under changes in the size of the visual vocabulary, the number of latent topics learned, and the type of discriminative classifier used (k-nearest neighbor or SVM). We achieve superior classification performance to recent publications that have used a bag of visual word representation, in all cases, using the authors' own data sets and testing protocols. We also investigate the gain in adding spatial information. We show applications to image retrieval with relevance feedback and to scene classification in videosIEEE2008info:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10256/2317http://hdl.handle.net/10256/2317© IEEE Transactions on Pattern Analysis and Machine Intelligence, 2008, vol. 30, p. 712-727Articles publicats (D-ATC)reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)Inglésinfo:eu-repo/semantics/altIdentifier/doi/10.1109/TPAMI.2007.70716info:eu-repo/semantics/altIdentifier/issn/0162-8828Tots els drets reservatsinfo:eu-repo/semantics/openAccessoai:recercat.cat:10256/23172026-05-29T05:05:01Z
dc.title.none.fl_str_mv Scene Classification Using a Hybrid Generative/Discriminative Approach
title Scene Classification Using a Hybrid Generative/Discriminative Approach
spellingShingle Scene Classification Using a Hybrid Generative/Discriminative Approach
Bosch Rué, Anna
Algorismes computacionals
Discriminació visual
Imatges -- Processament
Reconeixement de formes (Informàtica)
Vídeo digital
Computer algorithms
Digital video
Image processing
Pattern recognition systems
Visual discrimination
title_short Scene Classification Using a Hybrid Generative/Discriminative Approach
title_full Scene Classification Using a Hybrid Generative/Discriminative Approach
title_fullStr Scene Classification Using a Hybrid Generative/Discriminative Approach
title_full_unstemmed Scene Classification Using a Hybrid Generative/Discriminative Approach
title_sort Scene Classification Using a Hybrid Generative/Discriminative Approach
dc.creator.none.fl_str_mv Bosch Rué, Anna
Zisserman, Andrew
Muñoz Pujol, Xavier
author Bosch Rué, Anna
author_facet Bosch Rué, Anna
Zisserman, Andrew
Muñoz Pujol, Xavier
author_role author
author2 Zisserman, Andrew
Muñoz Pujol, Xavier
author2_role author
author
dc.subject.none.fl_str_mv Algorismes computacionals
Discriminació visual
Imatges -- Processament
Reconeixement de formes (Informàtica)
Vídeo digital
Computer algorithms
Digital video
Image processing
Pattern recognition systems
Visual discrimination
topic Algorismes computacionals
Discriminació visual
Imatges -- Processament
Reconeixement de formes (Informàtica)
Vídeo digital
Computer algorithms
Digital video
Image processing
Pattern recognition systems
Visual discrimination
description We investigate whether dimensionality reduction using a latent generative model is beneficial for the task of weakly supervised scene classification. In detail, we are given a set of labeled images of scenes (for example, coast, forest, city, river, etc.), and our objective is to classify a new image into one of these categories. Our approach consists of first discovering latent ";topics"; using probabilistic Latent Semantic Analysis (pLSA), a generative model from the statistical text literature here applied to a bag of visual words representation for each image, and subsequently, training a multiway classifier on the topic distribution vector for each image. We compare this approach to that of representing each image by a bag of visual words vector directly and training a multiway classifier on these vectors. To this end, we introduce a novel vocabulary using dense color SIFT descriptors and then investigate the classification performance under changes in the size of the visual vocabulary, the number of latent topics learned, and the type of discriminative classifier used (k-nearest neighbor or SVM). We achieve superior classification performance to recent publications that have used a bag of visual word representation, in all cases, using the authors' own data sets and testing protocols. We also investigate the gain in adding spatial information. We show applications to image retrieval with relevance feedback and to scene classification in videos
publishDate 2008
dc.date.none.fl_str_mv 2008
dc.type.none.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv http://hdl.handle.net/10256/2317
http://hdl.handle.net/10256/2317
url http://hdl.handle.net/10256/2317
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv info:eu-repo/semantics/altIdentifier/doi/10.1109/TPAMI.2007.70716
info:eu-repo/semantics/altIdentifier/issn/0162-8828
dc.rights.none.fl_str_mv Tots els drets reservats
info:eu-repo/semantics/openAccess
rights_invalid_str_mv Tots els drets reservats
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv IEEE
publisher.none.fl_str_mv IEEE
dc.source.none.fl_str_mv © IEEE Transactions on Pattern Analysis and Machine Intelligence, 2008, vol. 30, p. 712-727
Articles publicats (D-ATC)
reponame:Recercat. Dipósit de la Recerca de Catalunya
instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
instname_str Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
reponame_str Recercat. Dipósit de la Recerca de Catalunya
collection Recercat. Dipósit de la Recerca de Catalunya
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869407822480932864
score 15,812429