Feature Selection for Microarray Gene Expression Data Using Simulated Annealing Guided by the Multivariate Joint Entropy

Microarray classification poses many challen- ges for data analysis, given that a gene expression data set may consist of dozens of observations with thousands or even tens of thousands of genes. In this context, feature subset selection techniques can be very useful to reduce the representation spa...

Descripción completa

Detalles Bibliográficos
Autores: Félix Fernando González-Navarro, Lluís A. Belanche-Muñoz
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2014
País:México
Institución:Universidad Autónoma de Baja California
Repositorio:Redalyc-UABC
OAI Identifier:oai:redalyc.org:61531305005
Acceso en línea:https://www.redalyc.org/articulo.oa?id=61531305005
Access Level:acceso abierto
Palabra clave:Computación
sion data
Feature selection
simulated annealing
microarray gene expres
multivariate joint entropy
Descripción
Sumario:Microarray classification poses many challen- ges for data analysis, given that a gene expression data set may consist of dozens of observations with thousands or even tens of thousands of genes. In this context, feature subset selection techniques can be very useful to reduce the representation space to one that is manageable by classification techniques. In this work we use the discretized multivariate joint entropy as the basis for a fast evaluation of gene relevance in a Microarray Gene Expression context. The proposed algorithm com- bines a simulated annealing schedule specially designed for feature subset selection with the incrementally com- puted joint entropy, reusing previous values to compute current feature subset relevance. This combination turns out to be a powerful tool when applied to the maximiza- tion of gene subset relevance. Our method delivers highly interpretable solutions that are more accurate than competing methods. The algorithm is fast, effective and has no critical parameters. The experimental re- sults in several public-domain microarray data sets show a notoriously high classification performance and low size subsets, formed mostly by biologically meaningful genes. The technique is general and could be used in other similar scenarios.