A graph-based cache for large-scale similarity search engines
Large-scale similarity search engines are complex systems devised to process unstructured data like images and videos. These systems are deployed on clusters of distributed processors communicated through high-speed networks. To process a new query, a distance function is evaluated between the query...
| Autores: | , , , |
|---|---|
| Tipo de recurso: | artículo |
| Estado: | Versión publicada |
| Fecha de publicación: | 2018 |
| País: | Argentina |
| Institución: | Consejo Nacional de Investigaciones Científicas y Técnicas |
| Repositorio: | CONICET Digital (CONICET) |
| Idioma: | inglés |
| OAI Identifier: | oai:ri.conicet.gov.ar:11336/93223 |
| Acceso en línea: | http://hdl.handle.net/11336/93223 |
| Access Level: | acceso abierto |
| Palabra clave: | APPROXIMATE SIMILARITY SEARCH DISTRIBUTED LARGE-SCALE SEARCH ENGINES METRIC SPACE CACHE https://purl.org/becyt/ford/1.2 https://purl.org/becyt/ford/1 |
| id |
AR_edc01d41ed8906cbb475c0fc0d6b006d |
|---|---|
| oai_identifier_str |
oai:ri.conicet.gov.ar:11336/93223 |
| network_acronym_str |
AR |
| network_name_str |
Argentina |
| repository_id_str |
|
| spelling |
A graph-based cache for large-scale similarity search enginesGil Costa, Graciela VerónicaMarin, MauricioBonacic, CarolinaSolar, RobertoAPPROXIMATE SIMILARITY SEARCHDISTRIBUTED LARGE-SCALE SEARCH ENGINESMETRIC SPACE CACHEhttps://purl.org/becyt/ford/1.2https://purl.org/becyt/ford/1Large-scale similarity search engines are complex systems devised to process unstructured data like images and videos. These systems are deployed on clusters of distributed processors communicated through high-speed networks. To process a new query, a distance function is evaluated between the query and the objects stored in the database. This process relays on a metric space index distributed among the processors. In this paper, we propose a cache-based strategy devised to reduce the number of computations required to retrieve the top-k object results for user queries by using pre-computed information. Our proposal executes an approximate similarity search algorithm, which takes advantage of the links between objects stored in the cache memory. Those links form a graph of similarity among pre-computed queries. Compared to the previous methods in the literature, the proposed approach reduces the number of distance evaluations up to 60%.Fil: Gil Costa, Graciela Verónica. Universidad Nacional de San Luis; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas; ArgentinaFil: Marin, Mauricio. Universidad de Santiago de Chile; ChileFil: Bonacic, Carolina. Universidad de Santiago de Chile; ChileFil: Solar, Roberto. Universidad de Santiago de Chile; ChileSpringer2018-05info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttp://purl.org/coar/resource_type/c_6501info:ar-repo/semantics/articuloapplication/pdfapplication/pdfhttp://hdl.handle.net/11336/93223Gil Costa, Graciela Verónica; Marin, Mauricio; Bonacic, Carolina; Solar, Roberto; A graph-based cache for large-scale similarity search engines; Springer; Journal of Supercomputing; 74; 5; 5-2018; 2006-20340920-85421573-0484CONICET DigitalCONICETenginfo:eu-repo/semantics/altIdentifier/url/https://link.springer.com/article/10.1007/s11227-017-2207-3info:eu-repo/semantics/altIdentifier/doi/10.1007/s11227-017-2207-3info:eu-repo/semantics/openAccesshttps://creativecommons.org/licenses/by-nc-sa/2.5/ar/reponame:CONICET Digital (CONICET)instname:Consejo Nacional de Investigaciones Científicas y Técnicas2024-05-08T13:45:42Zoai:ri.conicet.gov.ar:11336/93223instacron:CONICETInstitucionalhttp://ri.conicet.gov.ar/Organismo científico-tecnológicoNo correspondehttp://ri.conicet.gov.ar/oai/requestdasensio@conicet.gov.ar; lcarlino@conicet.gov.arArgentinaNo correspondeNo correspondeNo correspondeopendoar:34982024-05-08 13:45:42.782CONICET Digital (CONICET) - Consejo Nacional de Investigaciones Científicas y Técnicasfalse |
| dc.title.none.fl_str_mv |
A graph-based cache for large-scale similarity search engines |
| title |
A graph-based cache for large-scale similarity search engines |
| spellingShingle |
A graph-based cache for large-scale similarity search engines Gil Costa, Graciela Verónica APPROXIMATE SIMILARITY SEARCH DISTRIBUTED LARGE-SCALE SEARCH ENGINES METRIC SPACE CACHE https://purl.org/becyt/ford/1.2 https://purl.org/becyt/ford/1 |
| title_short |
A graph-based cache for large-scale similarity search engines |
| title_full |
A graph-based cache for large-scale similarity search engines |
| title_fullStr |
A graph-based cache for large-scale similarity search engines |
| title_full_unstemmed |
A graph-based cache for large-scale similarity search engines |
| title_sort |
A graph-based cache for large-scale similarity search engines |
| dc.creator.none.fl_str_mv |
Gil Costa, Graciela Verónica Marin, Mauricio Bonacic, Carolina Solar, Roberto |
| author |
Gil Costa, Graciela Verónica |
| author_facet |
Gil Costa, Graciela Verónica Marin, Mauricio Bonacic, Carolina Solar, Roberto |
| author_role |
author |
| author2 |
Marin, Mauricio Bonacic, Carolina Solar, Roberto |
| author2_role |
author author author |
| dc.subject.none.fl_str_mv |
APPROXIMATE SIMILARITY SEARCH DISTRIBUTED LARGE-SCALE SEARCH ENGINES METRIC SPACE CACHE https://purl.org/becyt/ford/1.2 https://purl.org/becyt/ford/1 |
| topic |
APPROXIMATE SIMILARITY SEARCH DISTRIBUTED LARGE-SCALE SEARCH ENGINES METRIC SPACE CACHE https://purl.org/becyt/ford/1.2 https://purl.org/becyt/ford/1 |
| description |
Large-scale similarity search engines are complex systems devised to process unstructured data like images and videos. These systems are deployed on clusters of distributed processors communicated through high-speed networks. To process a new query, a distance function is evaluated between the query and the objects stored in the database. This process relays on a metric space index distributed among the processors. In this paper, we propose a cache-based strategy devised to reduce the number of computations required to retrieve the top-k object results for user queries by using pre-computed information. Our proposal executes an approximate similarity search algorithm, which takes advantage of the links between objects stored in the cache memory. Those links form a graph of similarity among pre-computed queries. Compared to the previous methods in the literature, the proposed approach reduces the number of distance evaluations up to 60%. |
| publishDate |
2018 |
| dc.date.none.fl_str_mv |
2018-05 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/publishedVersion http://purl.org/coar/resource_type/c_6501 info:ar-repo/semantics/articulo |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/11336/93223 Gil Costa, Graciela Verónica; Marin, Mauricio; Bonacic, Carolina; Solar, Roberto; A graph-based cache for large-scale similarity search engines; Springer; Journal of Supercomputing; 74; 5; 5-2018; 2006-2034 0920-8542 1573-0484 CONICET Digital CONICET |
| url |
http://hdl.handle.net/11336/93223 |
| identifier_str_mv |
Gil Costa, Graciela Verónica; Marin, Mauricio; Bonacic, Carolina; Solar, Roberto; A graph-based cache for large-scale similarity search engines; Springer; Journal of Supercomputing; 74; 5; 5-2018; 2006-2034 0920-8542 1573-0484 CONICET Digital CONICET |
| dc.language.none.fl_str_mv |
eng |
| language |
eng |
| dc.relation.none.fl_str_mv |
info:eu-repo/semantics/altIdentifier/url/https://link.springer.com/article/10.1007/s11227-017-2207-3 info:eu-repo/semantics/altIdentifier/doi/10.1007/s11227-017-2207-3 |
| dc.rights.none.fl_str_mv |
info:eu-repo/semantics/openAccess https://creativecommons.org/licenses/by-nc-sa/2.5/ar/ |
| eu_rights_str_mv |
openAccess |
| rights_invalid_str_mv |
https://creativecommons.org/licenses/by-nc-sa/2.5/ar/ |
| dc.format.none.fl_str_mv |
application/pdf application/pdf |
| dc.publisher.none.fl_str_mv |
Springer |
| publisher.none.fl_str_mv |
Springer |
| dc.source.none.fl_str_mv |
reponame:CONICET Digital (CONICET) instname:Consejo Nacional de Investigaciones Científicas y Técnicas |
| instname_str |
Consejo Nacional de Investigaciones Científicas y Técnicas |
| reponame_str |
CONICET Digital (CONICET) |
| collection |
CONICET Digital (CONICET) |
| repository.name.fl_str_mv |
CONICET Digital (CONICET) - Consejo Nacional de Investigaciones Científicas y Técnicas |
| repository.mail.fl_str_mv |
dasensio@conicet.gov.ar; lcarlino@conicet.gov.ar |
| _version_ |
1799195188151713792 |
| score |
15,812429 |