Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications

Augmented reality applications involving human interaction with virtual objects often rely on segmentation-based hand detection techniques. Semantic segmentation can then be enhanced with instance-specific information to model complex interactions between objects, but extracting such information typ...

Descripción completa

Detalles Bibliográficos
Autores: Veganzones, Miguel, Cisnal De La Rica, Ana, Fuente López, Eusebio de la, Fraile Marinero, Juan Carlos
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2024
País:España
Institución:Universidad de Valladolid
Repositorio:UVaDOC. Repositorio Documental de la Universidad de Valladolid
OAI Identifier:oai:uvadoc.uva.es:10324/73711
Acceso en línea:https://doi.org/10.3390/app142311357
https://uvadoc.uva.es/handle/10324/73711
Access Level:acceso abierto
Palabra clave:computer vision
convolutional neural networks
deep learning
hand segmentation
semantic segmentation
id ES_c001e781e2f18ffbdb166ab2f766022f
oai_identifier_str oai:uvadoc.uva.es:10324/73711
network_acronym_str ES
network_name_str España
repository_id_str
spelling Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation ApplicationsVeganzones, MiguelCisnal De La Rica, AnaFuente López, Eusebio de laFraile Marinero, Juan Carloscomputer visionconvolutional neural networksdeep learninghand segmentationsemantic segmentationAugmented reality applications involving human interaction with virtual objects often rely on segmentation-based hand detection techniques. Semantic segmentation can then be enhanced with instance-specific information to model complex interactions between objects, but extracting such information typically increases the computational load significantly. This study proposes a training strategy that enables conventional semantic segmentation networks to preserve some instance information during inference. This is accomplished by introducing pixel weight maps into the loss calculation, increasing the importance of boundary pixels between instances. We compare two common fully convolutional network (FCN) architectures, U-Net and ResNet, and fine-tune the fittest to improve segmentation results. Although the resulting model does not reach state-of-the-art segmentation performance on the EgoHands dataset, it preserves some instance information with no computational overhead. As expected, degraded segmentations are a necessary trade-off to preserve boundaries when instances are close together. This strategy allows approximating instance segmentation in real-time using non-specialized hardware, obtaining a unique blob for an instance with an intersection over union greater than 50% in 79% of the instances in our test set. A simple FCN, typically used for semantic segmentation, has shown promising instance segmentation results by introducing per-pixel weight maps during training for light-weight applications.MDPI2024info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfhttps://doi.org/10.3390/app142311357https://uvadoc.uva.es/handle/10324/73711reponame:UVaDOC. Repositorio Documental de la Universidad de Valladolidinstname:Universidad de ValladolidEspañolhttps://www.mdpi.com/2076-3417/14/23/11357info:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by-nc-nd/4.0/oai:uvadoc.uva.es:10324/737112026-06-13T12:44:47Z
dc.title.none.fl_str_mv Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
title Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
spellingShingle Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
Veganzones, Miguel
computer vision
convolutional neural networks
deep learning
hand segmentation
semantic segmentation
title_short Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
title_full Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
title_fullStr Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
title_full_unstemmed Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
title_sort Training Fully Convolutional Neural Networks for Lightweight, Non-Critical Instance Segmentation Applications
dc.creator.none.fl_str_mv Veganzones, Miguel
Cisnal De La Rica, Ana
Fuente López, Eusebio de la
Fraile Marinero, Juan Carlos
author Veganzones, Miguel
author_facet Veganzones, Miguel
Cisnal De La Rica, Ana
Fuente López, Eusebio de la
Fraile Marinero, Juan Carlos
author_role author
author2 Cisnal De La Rica, Ana
Fuente López, Eusebio de la
Fraile Marinero, Juan Carlos
author2_role author
author
author
dc.subject.none.fl_str_mv computer vision
convolutional neural networks
deep learning
hand segmentation
semantic segmentation
topic computer vision
convolutional neural networks
deep learning
hand segmentation
semantic segmentation
description Augmented reality applications involving human interaction with virtual objects often rely on segmentation-based hand detection techniques. Semantic segmentation can then be enhanced with instance-specific information to model complex interactions between objects, but extracting such information typically increases the computational load significantly. This study proposes a training strategy that enables conventional semantic segmentation networks to preserve some instance information during inference. This is accomplished by introducing pixel weight maps into the loss calculation, increasing the importance of boundary pixels between instances. We compare two common fully convolutional network (FCN) architectures, U-Net and ResNet, and fine-tune the fittest to improve segmentation results. Although the resulting model does not reach state-of-the-art segmentation performance on the EgoHands dataset, it preserves some instance information with no computational overhead. As expected, degraded segmentations are a necessary trade-off to preserve boundaries when instances are close together. This strategy allows approximating instance segmentation in real-time using non-specialized hardware, obtaining a unique blob for an instance with an intersection over union greater than 50% in 79% of the instances in our test set. A simple FCN, typically used for semantic segmentation, has shown promising instance segmentation results by introducing per-pixel weight maps during training for light-weight applications.
publishDate 2024
dc.date.none.fl_str_mv 2024
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv https://doi.org/10.3390/app142311357
https://uvadoc.uva.es/handle/10324/73711
url https://doi.org/10.3390/app142311357
https://uvadoc.uva.es/handle/10324/73711
dc.language.none.fl_str_mv Español
language_invalid_str_mv Español
dc.relation.none.fl_str_mv https://www.mdpi.com/2076-3417/14/23/11357
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
http://creativecommons.org/licenses/by-nc-nd/4.0/
eu_rights_str_mv openAccess
rights_invalid_str_mv http://creativecommons.org/licenses/by-nc-nd/4.0/
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv MDPI
publisher.none.fl_str_mv MDPI
dc.source.none.fl_str_mv reponame:UVaDOC. Repositorio Documental de la Universidad de Valladolid
instname:Universidad de Valladolid
instname_str Universidad de Valladolid
reponame_str UVaDOC. Repositorio Documental de la Universidad de Valladolid
collection UVaDOC. Repositorio Documental de la Universidad de Valladolid
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869418440197931008
score 15,812429