A graph-based cache for large-scale similarity search engines

Large-scale similarity search engines are complex systems devised to process unstructured data like images and videos. These systems are deployed on clusters of distributed processors communicated through high-speed networks. To process a new query, a distance function is evaluated between the query...

Descripción completa

Detalles Bibliográficos
Autores: Gil Costa, Graciela Verónica, Marin, Mauricio, Bonacic, Carolina, Solar, Roberto
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2018
País:Argentina
Institución:Consejo Nacional de Investigaciones Científicas y Técnicas
Repositorio:CONICET Digital (CONICET)
Idioma:inglés
OAI Identifier:oai:ri.conicet.gov.ar:11336/93223
Acceso en línea:http://hdl.handle.net/11336/93223
Access Level:acceso abierto
Palabra clave:APPROXIMATE SIMILARITY SEARCH
DISTRIBUTED LARGE-SCALE SEARCH ENGINES
METRIC SPACE CACHE
https://purl.org/becyt/ford/1.2
https://purl.org/becyt/ford/1
id AR_edc01d41ed8906cbb475c0fc0d6b006d
oai_identifier_str oai:ri.conicet.gov.ar:11336/93223
network_acronym_str AR
network_name_str Argentina
repository_id_str
spelling A graph-based cache for large-scale similarity search enginesGil Costa, Graciela VerónicaMarin, MauricioBonacic, CarolinaSolar, RobertoAPPROXIMATE SIMILARITY SEARCHDISTRIBUTED LARGE-SCALE SEARCH ENGINESMETRIC SPACE CACHEhttps://purl.org/becyt/ford/1.2https://purl.org/becyt/ford/1Large-scale similarity search engines are complex systems devised to process unstructured data like images and videos. These systems are deployed on clusters of distributed processors communicated through high-speed networks. To process a new query, a distance function is evaluated between the query and the objects stored in the database. This process relays on a metric space index distributed among the processors. In this paper, we propose a cache-based strategy devised to reduce the number of computations required to retrieve the top-k object results for user queries by using pre-computed information. Our proposal executes an approximate similarity search algorithm, which takes advantage of the links between objects stored in the cache memory. Those links form a graph of similarity among pre-computed queries. Compared to the previous methods in the literature, the proposed approach reduces the number of distance evaluations up to 60%.Fil: Gil Costa, Graciela Verónica. Universidad Nacional de San Luis; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas; ArgentinaFil: Marin, Mauricio. Universidad de Santiago de Chile; ChileFil: Bonacic, Carolina. Universidad de Santiago de Chile; ChileFil: Solar, Roberto. Universidad de Santiago de Chile; ChileSpringer2018-05info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttp://purl.org/coar/resource_type/c_6501info:ar-repo/semantics/articuloapplication/pdfapplication/pdfhttp://hdl.handle.net/11336/93223Gil Costa, Graciela Verónica; Marin, Mauricio; Bonacic, Carolina; Solar, Roberto; A graph-based cache for large-scale similarity search engines; Springer; Journal of Supercomputing; 74; 5; 5-2018; 2006-20340920-85421573-0484CONICET DigitalCONICETenginfo:eu-repo/semantics/altIdentifier/url/https://link.springer.com/article/10.1007/s11227-017-2207-3info:eu-repo/semantics/altIdentifier/doi/10.1007/s11227-017-2207-3info:eu-repo/semantics/openAccesshttps://creativecommons.org/licenses/by-nc-sa/2.5/ar/reponame:CONICET Digital (CONICET)instname:Consejo Nacional de Investigaciones Científicas y Técnicas2024-05-08T13:45:42Zoai:ri.conicet.gov.ar:11336/93223instacron:CONICETInstitucionalhttp://ri.conicet.gov.ar/Organismo científico-tecnológicoNo correspondehttp://ri.conicet.gov.ar/oai/requestdasensio@conicet.gov.ar; lcarlino@conicet.gov.arArgentinaNo correspondeNo correspondeNo correspondeopendoar:34982024-05-08 13:45:42.782CONICET Digital (CONICET) - Consejo Nacional de Investigaciones Científicas y Técnicasfalse
dc.title.none.fl_str_mv A graph-based cache for large-scale similarity search engines
title A graph-based cache for large-scale similarity search engines
spellingShingle A graph-based cache for large-scale similarity search engines
Gil Costa, Graciela Verónica
APPROXIMATE SIMILARITY SEARCH
DISTRIBUTED LARGE-SCALE SEARCH ENGINES
METRIC SPACE CACHE
https://purl.org/becyt/ford/1.2
https://purl.org/becyt/ford/1
title_short A graph-based cache for large-scale similarity search engines
title_full A graph-based cache for large-scale similarity search engines
title_fullStr A graph-based cache for large-scale similarity search engines
title_full_unstemmed A graph-based cache for large-scale similarity search engines
title_sort A graph-based cache for large-scale similarity search engines
dc.creator.none.fl_str_mv Gil Costa, Graciela Verónica
Marin, Mauricio
Bonacic, Carolina
Solar, Roberto
author Gil Costa, Graciela Verónica
author_facet Gil Costa, Graciela Verónica
Marin, Mauricio
Bonacic, Carolina
Solar, Roberto
author_role author
author2 Marin, Mauricio
Bonacic, Carolina
Solar, Roberto
author2_role author
author
author
dc.subject.none.fl_str_mv APPROXIMATE SIMILARITY SEARCH
DISTRIBUTED LARGE-SCALE SEARCH ENGINES
METRIC SPACE CACHE
https://purl.org/becyt/ford/1.2
https://purl.org/becyt/ford/1
topic APPROXIMATE SIMILARITY SEARCH
DISTRIBUTED LARGE-SCALE SEARCH ENGINES
METRIC SPACE CACHE
https://purl.org/becyt/ford/1.2
https://purl.org/becyt/ford/1
description Large-scale similarity search engines are complex systems devised to process unstructured data like images and videos. These systems are deployed on clusters of distributed processors communicated through high-speed networks. To process a new query, a distance function is evaluated between the query and the objects stored in the database. This process relays on a metric space index distributed among the processors. In this paper, we propose a cache-based strategy devised to reduce the number of computations required to retrieve the top-k object results for user queries by using pre-computed information. Our proposal executes an approximate similarity search algorithm, which takes advantage of the links between objects stored in the cache memory. Those links form a graph of similarity among pre-computed queries. Compared to the previous methods in the literature, the proposed approach reduces the number of distance evaluations up to 60%.
publishDate 2018
dc.date.none.fl_str_mv 2018-05
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
http://purl.org/coar/resource_type/c_6501
info:ar-repo/semantics/articulo
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/11336/93223
Gil Costa, Graciela Verónica; Marin, Mauricio; Bonacic, Carolina; Solar, Roberto; A graph-based cache for large-scale similarity search engines; Springer; Journal of Supercomputing; 74; 5; 5-2018; 2006-2034
0920-8542
1573-0484
CONICET Digital
CONICET
url http://hdl.handle.net/11336/93223
identifier_str_mv Gil Costa, Graciela Verónica; Marin, Mauricio; Bonacic, Carolina; Solar, Roberto; A graph-based cache for large-scale similarity search engines; Springer; Journal of Supercomputing; 74; 5; 5-2018; 2006-2034
0920-8542
1573-0484
CONICET Digital
CONICET
dc.language.none.fl_str_mv eng
language eng
dc.relation.none.fl_str_mv info:eu-repo/semantics/altIdentifier/url/https://link.springer.com/article/10.1007/s11227-017-2207-3
info:eu-repo/semantics/altIdentifier/doi/10.1007/s11227-017-2207-3
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
https://creativecommons.org/licenses/by-nc-sa/2.5/ar/
eu_rights_str_mv openAccess
rights_invalid_str_mv https://creativecommons.org/licenses/by-nc-sa/2.5/ar/
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Springer
publisher.none.fl_str_mv Springer
dc.source.none.fl_str_mv reponame:CONICET Digital (CONICET)
instname:Consejo Nacional de Investigaciones Científicas y Técnicas
instname_str Consejo Nacional de Investigaciones Científicas y Técnicas
reponame_str CONICET Digital (CONICET)
collection CONICET Digital (CONICET)
repository.name.fl_str_mv CONICET Digital (CONICET) - Consejo Nacional de Investigaciones Científicas y Técnicas
repository.mail.fl_str_mv dasensio@conicet.gov.ar; lcarlino@conicet.gov.ar
_version_ 1799195188151713792
score 15,812429