Semi‑supervised incremental learning with few examples for discovering medical association rules

Background: Association Rules are one of the main ways to represent structural patterns underlying raw data. They represent dependencies between sets of observations contained in the data. The associations established by these rules are very useful in the medical domain, for example in the predictiv...

Descripción completa

Detalles Bibliográficos
Autores: Sánchez‑de‑Madariaga, Ricardo, Cantero Escribano, José Miguel, Martínez Romo, Juan, Araujo Serna, M. Lourdes
Tipo de recurso: artículo
Fecha de publicación:2022
País:España
Institución:Universidad Nacional de Educación a Distancia
Repositorio:e-spacio. Repositorio Institucional de la UNED
Idioma:inglés
OAI Identifier:oai:e-spacio.uned.es:20.500.14468/23301
Acceso en línea:https://hdl.handle.net/20.500.14468/23301
Access Level:acceso abierto
Palabra clave:Medical records
Association rules discovery
Machine learning
Semi‑supervised approach
id ES_bc4f86705e215313458b0d802b435f4f
oai_identifier_str oai:e-spacio.uned.es:20.500.14468/23301
network_acronym_str ES
network_name_str España
repository_id_str
spelling Semi‑supervised incremental learning with few examples for discovering medical association rulesSánchez‑de‑Madariaga, RicardoCantero Escribano, José MiguelMartínez Romo, JuanAraujo Serna, M. LourdesMedical recordsAssociation rules discoveryMachine learningSemi‑supervised approachBackground: Association Rules are one of the main ways to represent structural patterns underlying raw data. They represent dependencies between sets of observations contained in the data. The associations established by these rules are very useful in the medical domain, for example in the predictive health field. Classic algorithms for association rule mining give rise to huge amounts of possible rules that should be filtered in order to select those most likely to be true. Most of the proposed techniques for these tasks are unsupervised. However, the accuracy provided by unsupervised systems is limited. Conversely, resorting to annotated data for training supervised systems is expensive and time‑consuming. The purpose of this research is to design a new semi‑supervised algorithm that performs like supervised algorithms but uses an affordable amount of training data. Methods: In this work we propose a new semi‑supervised data mining model that combines unsupervised techniques (Fisher’s exact test) with limited supervision. Starting with a small seed of annotated data, the model improves results (F‑measure) obtained, using a fully supervised system (standard supervised ML algorithms). The idea is based on utilising the agreement between the predictions of the supervised system and those of the unsupervised techniques in a series of iterative steps. Results: The new semi‑supervised ML algorithm improves the results of supervised algorithms computed using the F‑measure in the task of mining medical association rules, but training with an affordable amount of manually annotated data. Conclusions: Using a small amount of annotated data (which is easily achievable) leads to results similar to those of a supervised system. The proposal may be an important step for the practical development of techniques for mining association rules and generating new valuable scientific medical knowledge.BioMed Centrale-Spacio UNED20242024-08-2120222022-01-0120222022-01-01journal articlehttp://purl.org/coar/resource_type/c_6501info:eu-repo/semantics/articleapplication/pdfhttps://hdl.handle.net/20.500.14468/23301reponame:e-spacio. Repositorio Institucional de la UNEDinstname:Universidad Nacional de Educación a DistanciaInglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by-nc-nd/4.0oai:e-spacio.uned.es:20.500.14468/233012026-06-06T12:38:31Z
dc.title.none.fl_str_mv Semi‑supervised incremental learning with few examples for discovering medical association rules
title Semi‑supervised incremental learning with few examples for discovering medical association rules
spellingShingle Semi‑supervised incremental learning with few examples for discovering medical association rules
Sánchez‑de‑Madariaga, Ricardo
Medical records
Association rules discovery
Machine learning
Semi‑supervised approach
title_short Semi‑supervised incremental learning with few examples for discovering medical association rules
title_full Semi‑supervised incremental learning with few examples for discovering medical association rules
title_fullStr Semi‑supervised incremental learning with few examples for discovering medical association rules
title_full_unstemmed Semi‑supervised incremental learning with few examples for discovering medical association rules
title_sort Semi‑supervised incremental learning with few examples for discovering medical association rules
dc.creator.none.fl_str_mv Sánchez‑de‑Madariaga, Ricardo
Cantero Escribano, José Miguel
Martínez Romo, Juan
Araujo Serna, M. Lourdes
author Sánchez‑de‑Madariaga, Ricardo
author_facet Sánchez‑de‑Madariaga, Ricardo
Cantero Escribano, José Miguel
Martínez Romo, Juan
Araujo Serna, M. Lourdes
author_role author
author2 Cantero Escribano, José Miguel
Martínez Romo, Juan
Araujo Serna, M. Lourdes
author2_role author
author
author
dc.contributor.none.fl_str_mv e-Spacio UNED
dc.subject.none.fl_str_mv Medical records
Association rules discovery
Machine learning
Semi‑supervised approach
topic Medical records
Association rules discovery
Machine learning
Semi‑supervised approach
description Background: Association Rules are one of the main ways to represent structural patterns underlying raw data. They represent dependencies between sets of observations contained in the data. The associations established by these rules are very useful in the medical domain, for example in the predictive health field. Classic algorithms for association rule mining give rise to huge amounts of possible rules that should be filtered in order to select those most likely to be true. Most of the proposed techniques for these tasks are unsupervised. However, the accuracy provided by unsupervised systems is limited. Conversely, resorting to annotated data for training supervised systems is expensive and time‑consuming. The purpose of this research is to design a new semi‑supervised algorithm that performs like supervised algorithms but uses an affordable amount of training data. Methods: In this work we propose a new semi‑supervised data mining model that combines unsupervised techniques (Fisher’s exact test) with limited supervision. Starting with a small seed of annotated data, the model improves results (F‑measure) obtained, using a fully supervised system (standard supervised ML algorithms). The idea is based on utilising the agreement between the predictions of the supervised system and those of the unsupervised techniques in a series of iterative steps. Results: The new semi‑supervised ML algorithm improves the results of supervised algorithms computed using the F‑measure in the task of mining medical association rules, but training with an affordable amount of manually annotated data. Conclusions: Using a small amount of annotated data (which is easily achievable) leads to results similar to those of a supervised system. The proposal may be an important step for the practical development of techniques for mining association rules and generating new valuable scientific medical knowledge.
publishDate 2022
dc.date.none.fl_str_mv 2022
2022-01-01
2022
2022-01-01
2024
2024-08-21
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://hdl.handle.net/20.500.14468/23301
url https://hdl.handle.net/20.500.14468/23301
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
info:eu-repo/semantics/openAccess
http://creativecommons.org/licenses/by-nc-nd/4.0
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
http://creativecommons.org/licenses/by-nc-nd/4.0
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv BioMed Central
publisher.none.fl_str_mv BioMed Central
dc.source.none.fl_str_mv reponame:e-spacio. Repositorio Institucional de la UNED
instname:Universidad Nacional de Educación a Distancia
instname_str Universidad Nacional de Educación a Distancia
reponame_str e-spacio. Repositorio Institucional de la UNED
collection e-spacio. Repositorio Institucional de la UNED
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869418102205186048
score 15,812455