Logo Detection with No Priors

In recent years, top referred methods on object detection like R-CNN have implemented this task as a combination of proposal region generation and supervised classification on the proposed bounding boxes. Although this pipeline has achieved state-of-the-art results in multiple datasets, it has inher...

ver descrição completa

Detalhes bibliográficos
Autores: Velazquez Dorta, Diego Alejandro|||0000-0001-6782-4336, Gonfaus, Josep M.|||0000-0003-2079-4103, Rodríguez López, Pau|||0000-0002-1689-8084, Roca, F. Xavier|||0000-0002-7043-7334, Ozawa, Seiichi|||0000-0002-0965-0064, Gonzàlez, Jordi|||0000-0001-8033-0306
Formato: artículo
Fecha de publicación:2021
País:España
Recursos:Universitat Autònoma de Barcelona
Repositorio:Dipòsit Digital de Documents de la UAB
Idioma:inglés
OAI Identifier:oai:ddd.uab.cat:326624
Acesso em linha:https://ddd.uab.cat/record/326624
https://dx.doi.org/urn:doi:10.1109/ACCESS.2021.3101297
Access Level:acceso abierto
Palavra-chave:Attention
Deep learning
Logo detection
Object detection
Transformers
Descrição
Resumo:In recent years, top referred methods on object detection like R-CNN have implemented this task as a combination of proposal region generation and supervised classification on the proposed bounding boxes. Although this pipeline has achieved state-of-the-art results in multiple datasets, it has inherent limitations that make object detection a very complex and inefficient task in computational terms. Instead of considering this standard strategy, in this paper we enhance Detection Transformers (DETR) which tackles object detection as a set-prediction problem directly in an end-to-end fully differentiable pipeline without requiring priors. In particular, we incorporate Feature Pyramids (FP) to the DETR architecture and demonstrate the effectiveness of the resulting DETR-FP approach on improving logo detection results thanks to the improved detection of small logos. So, without requiring any domain specific prior to be fed to the model, DETR-FP obtains competitive results on the OpenLogo and MS-COCO datasets offering a relative improvement of up to 30%, when compared to a Faster R-CNN baseline which strongly depends on hand-designed priors.