Testing for the existence of clusters

Detecting and determining clusters present in a certain sample has been an important concern, among researchers from different fields, for a long time. In particular, assessing whether the clusters are statistically significant, is a question that has been asked by a number of experimenters. Recentl...

ver descrição completa

Detalhes bibliográficos
Autores: Fuentes, Claudio, Casella, George
Formato: artículo
Fecha de publicación:2009
País:España
Recursos:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2099/8947
Acesso em linha:https://hdl.handle.net/2099/8947
Access Level:acceso abierto
Palavra-chave:Mathematical statistics
Hierarchical models
Bayesian inference
Frequentist calibration
Monte Carlo methods
P-values.
Estadística matemàtica
Classificació AMS::62 Statistics::62F Parametric inference
Àrees temàtiques de la UPC::Matemàtiques i estadística::Estadística matemàtica
id ES_2f674d5e8b1bcd4c58e95a7780fe18ed
oai_identifier_str oai:upcommons.upc.edu:2099/8947
network_acronym_str ES
network_name_str España
repository_id_str
spelling Testing for the existence of clustersFuentes, ClaudioCasella, GeorgeMathematical statisticsHierarchical modelsBayesian inferenceFrequentist calibrationMonte Carlo methodsP-values.Estadística matemàticaClassificació AMS::62 Statistics::62F Parametric inferenceÀrees temàtiques de la UPC::Matemàtiques i estadística::Estadística matemàticaDetecting and determining clusters present in a certain sample has been an important concern, among researchers from different fields, for a long time. In particular, assessing whether the clusters are statistically significant, is a question that has been asked by a number of experimenters. Recently, this question arose again in a study in maize genetics, where determining the significance of clusters is crucial as a primary step in the identification of a genome-wide collection of mutants that may affect the kernel composition. Although several efforts have been made in this direction, not much has been done with the aim of developing an actual hypothesis test in order to assess the significance of clusters. In this paper, we propose a new methodology that allows the examination of the hypothesis test H0 : =1 vs. H1 : =k, where denotes the number of clusters present in a certain population. Our procedure, based on Bayesian tools, permits us to obtain closed form expressions for the posterior probabilities corresponding to the null hypothesis. From here, we calibrate our results by estimating the frequentist null distribution of the posterior probabilities in order to obtain the p-values associated with the observed posterior probabilities. In most cases, actual evaluation of the posterior probabilities is computationally intensive and several algorithms have been discussed in the literature. Here, we propose a simple estimation procedure, based on MCMC techniques, that permits an efficient and easily implementable evaluation of the test. Finally, we present simulation studies that support our conclusions, and we apply our method to the analysis of NIR spectroscopy data coming from the genetic study that motivated this work.Peer ReviewedInstitut d'Estadística de Catalunya20092009-01-0120102010-04-27journal articlehttp://purl.org/coar/resource_type/c_6501NAhttp://purl.org/coar/version/c_be7fb7dd8ff6fe43info:eu-repo/semantics/articleapplication/pdfhttps://hdl.handle.net/2099/8947reponame:UPCommons. Portal del coneixement obert de la UPCinstname:Universitat Politècnica de Catalunya (UPC)Inglésengopen accesshttp://purl.org/coar/access_right/c_abf2Attribution-NonCommercial-NoDerivs 3.0 Spainhttp://creativecommons.org/licenses/by-nc-nd/3.0/es/info:eu-repo/semantics/openAccessoai:upcommons.upc.edu:2099/89472026-05-27T15:37:01Z
dc.title.none.fl_str_mv Testing for the existence of clusters
title Testing for the existence of clusters
spellingShingle Testing for the existence of clusters
Fuentes, Claudio
Mathematical statistics
Hierarchical models
Bayesian inference
Frequentist calibration
Monte Carlo methods
P-values.
Estadística matemàtica
Classificació AMS::62 Statistics::62F Parametric inference
Àrees temàtiques de la UPC::Matemàtiques i estadística::Estadística matemàtica
title_short Testing for the existence of clusters
title_full Testing for the existence of clusters
title_fullStr Testing for the existence of clusters
title_full_unstemmed Testing for the existence of clusters
title_sort Testing for the existence of clusters
dc.creator.none.fl_str_mv Fuentes, Claudio
Casella, George
author Fuentes, Claudio
author_facet Fuentes, Claudio
Casella, George
author_role author
author2 Casella, George
author2_role author
dc.subject.none.fl_str_mv Mathematical statistics
Hierarchical models
Bayesian inference
Frequentist calibration
Monte Carlo methods
P-values.
Estadística matemàtica
Classificació AMS::62 Statistics::62F Parametric inference
Àrees temàtiques de la UPC::Matemàtiques i estadística::Estadística matemàtica
topic Mathematical statistics
Hierarchical models
Bayesian inference
Frequentist calibration
Monte Carlo methods
P-values.
Estadística matemàtica
Classificació AMS::62 Statistics::62F Parametric inference
Àrees temàtiques de la UPC::Matemàtiques i estadística::Estadística matemàtica
description Detecting and determining clusters present in a certain sample has been an important concern, among researchers from different fields, for a long time. In particular, assessing whether the clusters are statistically significant, is a question that has been asked by a number of experimenters. Recently, this question arose again in a study in maize genetics, where determining the significance of clusters is crucial as a primary step in the identification of a genome-wide collection of mutants that may affect the kernel composition. Although several efforts have been made in this direction, not much has been done with the aim of developing an actual hypothesis test in order to assess the significance of clusters. In this paper, we propose a new methodology that allows the examination of the hypothesis test H0 : =1 vs. H1 : =k, where denotes the number of clusters present in a certain population. Our procedure, based on Bayesian tools, permits us to obtain closed form expressions for the posterior probabilities corresponding to the null hypothesis. From here, we calibrate our results by estimating the frequentist null distribution of the posterior probabilities in order to obtain the p-values associated with the observed posterior probabilities. In most cases, actual evaluation of the posterior probabilities is computationally intensive and several algorithms have been discussed in the literature. Here, we propose a simple estimation procedure, based on MCMC techniques, that permits an efficient and easily implementable evaluation of the test. Finally, we present simulation studies that support our conclusions, and we apply our method to the analysis of NIR spectroscopy data coming from the genetic study that motivated this work.
publishDate 2009
dc.date.none.fl_str_mv 2009
2009-01-01
2010
2010-04-27
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
NA
http://purl.org/coar/version/c_be7fb7dd8ff6fe43
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://hdl.handle.net/2099/8947
url https://hdl.handle.net/2099/8947
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
Attribution-NonCommercial-NoDerivs 3.0 Spain
http://creativecommons.org/licenses/by-nc-nd/3.0/es/
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
Attribution-NonCommercial-NoDerivs 3.0 Spain
http://creativecommons.org/licenses/by-nc-nd/3.0/es/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Institut d'Estadística de Catalunya
publisher.none.fl_str_mv Institut d'Estadística de Catalunya
dc.source.none.fl_str_mv reponame:UPCommons. Portal del coneixement obert de la UPC
instname:Universitat Politècnica de Catalunya (UPC)
instname_str Universitat Politècnica de Catalunya (UPC)
reponame_str UPCommons. Portal del coneixement obert de la UPC
collection UPCommons. Portal del coneixement obert de la UPC
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869405475812933632
score 15,301603