Large expert-curated database for benchmarking document similarity detection in biomedical literature search

P.B. participated in the design, carried out the study, implemented the websites and search systems and wrote the manuscript. RELISH consortium annotated the articles. Y.Z. conceived the study, participated in the initial design, assisted in analyzing data and wrote the manuscript. All authors read,...

Descripción completa

Detalles Bibliográficos
Autores: Brown, Peter, RELISH Consortium, Zhou, Yaoqi, González-Prendes, Rayner
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2019
País:España
Institución:Universitat de Lleida (UdL)
Repositorio:Repositori Obert UdL
OAI Identifier:oai:repositori.udl.cat:10459.1/68347
Acceso en línea:https://doi.org/10.1093/database/baz085
http://hdl.handle.net/10459.1/68347
Access Level:acceso abierto
id ES_053500a47eeb84ada76d8115da459b37
oai_identifier_str oai:repositori.udl.cat:10459.1/68347
network_acronym_str ES
network_name_str España
repository_id_str
spelling Large expert-curated database for benchmarking document similarity detection in biomedical literature searchBrown, PeterRELISH ConsortiumZhou, YaoqiGonzález-Prendes, RaynerP.B. participated in the design, carried out the study, implemented the websites and search systems and wrote the manuscript. RELISH consortium annotated the articles. Y.Z. conceived the study, participated in the initial design, assisted in analyzing data and wrote the manuscript. All authors read, contributed to the discussion and approved the manuscript. Rayner Gonzálezlez-Prendes is member of the RELISH consortium.Document recommendation systems for locating relevant literature have mostly relied on methods developed a decade ago. This is largely due to the lack of a large offline gold-standard benchmark of relevant documents that cover a variety of research fields such that newly developed literature search techniques can be compared, improved and translated into practice. To overcome this bottleneck, we have established the RElevant LIterature SearcH consortium consisting of more than 1500 scientists from 84 countries, who have collectively annotated the relevance of over 180 000 PubMed-listed articles with regard to their respective seed (input) article/s. The majority of annotations were contributed by highly experienced, original authors of the seed articles. The collected data cover 76% of all unique PubMed Medical Subject Headings descriptors. No systematic biases were observed across different experience levels, research fields or time spent on annotations. More importantly, annotations of the same document pairs contributed by different scientists were highly concordant. We further show that the three representative baseline methods used to generate recommended articles for evaluation (Okapi Best Matching 25, Term Frequency–Inverse Document Frequency and PubMed Related Articles) had similar overall performances. Additionally, we found that these methods each tend to produce distinct collections of recommended articles, suggesting that a hybrid method may be required to completely capture all relevant articles. The established database server located at https://relishdb.ict.griffith.edu.au is freely available for the downloading of annotation data and the blind testing of new methods. We expect that this benchmark will be useful for stimulating the development of new powerful techniques for title and title/abstract-based search engines for relevant articles in biomedical research.Griffith University Gowonda HPC Cluster; Queensland Cyber Infrastructure Foundation.Oxford University Press2019info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttps://doi.org/10.1093/database/baz085http://hdl.handle.net/10459.1/68347reponame:Repositori Obert UdL instname:Universitat de Lleida (UdL)InglésReproducció del document publicat a: https://doi.org/10.1093/database/baz085Database: the journal of biological databases and curation, 2019, vol. 2019, baz085, p. 1-66cc-by (c) Brown, Peter et al., 2019info:eu-repo/semantics/openAccesshttps://creativecommons.org/licenses/by/4.0/oai:repositori.udl.cat:10459.1/683472026-06-24T12:42:17Z
dc.title.none.fl_str_mv Large expert-curated database for benchmarking document similarity detection in biomedical literature search
title Large expert-curated database for benchmarking document similarity detection in biomedical literature search
spellingShingle Large expert-curated database for benchmarking document similarity detection in biomedical literature search
Brown, Peter
title_short Large expert-curated database for benchmarking document similarity detection in biomedical literature search
title_full Large expert-curated database for benchmarking document similarity detection in biomedical literature search
title_fullStr Large expert-curated database for benchmarking document similarity detection in biomedical literature search
title_full_unstemmed Large expert-curated database for benchmarking document similarity detection in biomedical literature search
title_sort Large expert-curated database for benchmarking document similarity detection in biomedical literature search
dc.creator.none.fl_str_mv Brown, Peter
RELISH Consortium
Zhou, Yaoqi
González-Prendes, Rayner
author Brown, Peter
author_facet Brown, Peter
RELISH Consortium
Zhou, Yaoqi
González-Prendes, Rayner
author_role author
author2 RELISH Consortium
Zhou, Yaoqi
González-Prendes, Rayner
author2_role author
author
author
description P.B. participated in the design, carried out the study, implemented the websites and search systems and wrote the manuscript. RELISH consortium annotated the articles. Y.Z. conceived the study, participated in the initial design, assisted in analyzing data and wrote the manuscript. All authors read, contributed to the discussion and approved the manuscript. Rayner Gonzálezlez-Prendes is member of the RELISH consortium.
publishDate 2019
dc.date.none.fl_str_mv 2019
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv https://doi.org/10.1093/database/baz085
http://hdl.handle.net/10459.1/68347
url https://doi.org/10.1093/database/baz085
http://hdl.handle.net/10459.1/68347
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Reproducció del document publicat a: https://doi.org/10.1093/database/baz085
Database: the journal of biological databases and curation, 2019, vol. 2019, baz085, p. 1-66
dc.rights.none.fl_str_mv cc-by (c) Brown, Peter et al., 2019
info:eu-repo/semantics/openAccess
https://creativecommons.org/licenses/by/4.0/
rights_invalid_str_mv cc-by (c) Brown, Peter et al., 2019
https://creativecommons.org/licenses/by/4.0/
eu_rights_str_mv openAccess
dc.publisher.none.fl_str_mv Oxford University Press
publisher.none.fl_str_mv Oxford University Press
dc.source.none.fl_str_mv reponame:Repositori Obert UdL
instname:Universitat de Lleida (UdL)
instname_str Universitat de Lleida (UdL)
reponame_str Repositori Obert UdL
collection Repositori Obert UdL
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869402840367104000
score 15,812429