Machine learning: how much does it tell about protein folding rates?

The prediction of protein folding rates is a necessary step towards understanding the principles of protein folding. Due to the increasing amount of experimental data, numerous protein folding models and predictors of protein folding rates have been developed in the last decade. The problem has also...

ver descrição completa

Detalhes bibliográficos
Autores: Corrales, Marc, Cuscó, Pol, Usmanova, Dinara R., Chen, Heng-Chang, Bogatyreva, Natalya S., Filion, Guillaume, Ivankov, Dmitry N.
Formato: artículo
Estado:Versión publicada
Fecha de publicación:2015
País:España
Recursos:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
Repositorio:Recercat. Dipósit de la Recerca de Catalunya
OAI Identifier:oai:recercat.cat:10230/58480
Acesso em linha:http://hdl.handle.net/10230/58480
http://dx.doi.org/10.1371/journal.pone.0143166
Access Level:acceso abierto
Palavra-chave:Aprenentatge automàtic
Logaritmes
Proteïnes
Estadística
id ES_728e1fad4cefc545dcecef9ecf015736
oai_identifier_str oai:recercat.cat:10230/58480
network_acronym_str ES
network_name_str España
repository_id_str
spelling Machine learning: how much does it tell about protein folding rates?Corrales, MarcCuscó, PolUsmanova, Dinara R.Chen, Heng-ChangBogatyreva, Natalya S.Filion, GuillaumeIvankov, Dmitry N.Aprenentatge automàticLogaritmesProteïnesEstadísticaThe prediction of protein folding rates is a necessary step towards understanding the principles of protein folding. Due to the increasing amount of experimental data, numerous protein folding models and predictors of protein folding rates have been developed in the last decade. The problem has also attracted the attention of scientists from computational fields, which led to the publication of several machine learning-based models to predict the rate of protein folding. Some of them claim to predict the logarithm of protein folding rate with an accuracy greater than 90%. However, there are reasons to believe that such claims are exaggerated due to large fluctuations and overfitting of the estimates. When we confronted three selected published models with new data, we found a much lower predictive power than reported in the original publications. Overly optimistic predictive powers appear from violations of the basic principles of machine-learning. We highlight common misconceptions in the studies claiming excessive predictive power and propose to use learning curves as a safeguard against those mistakes. As an example, we show that the current amount of experimental data is insufficient to build a linear predictor of logarithms of folding rates based on protein amino acid composition.NSB was supported by the Russian Science Foundation Grant 14-24-00157. DRU and DNI were supported by ERC grant 335980_EinME. The authors acknowledge support of the Spanish Ministry of Economy and Competitiveness, ‘Centro de Excelencia Severo Ochoa 2013-2017’, SEV-2012-0208. HCC and PC were supported by the Spanish Ministry of Economy and Competitiveness (including State Training Subprogram: predoctoral fellowships for the training of PhD students (FPI) 2013). MC and GF were supported by the CRG. The publication cost was covered by CRG.Public Library of Science (PLoS)202320232015info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/58480http://dx.doi.org/10.1371/journal.pone.0143166reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésPLoS ONE. 2015 Nov 25;10(11):e0143166info:eu-repo/grantAgreement/EC/FP7/335980© 2015 Corrales et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.http://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/openAccessoai:recercat.cat:10230/584802026-05-29T05:05:01Z
dc.title.none.fl_str_mv Machine learning: how much does it tell about protein folding rates?
title Machine learning: how much does it tell about protein folding rates?
spellingShingle Machine learning: how much does it tell about protein folding rates?
Corrales, Marc
Aprenentatge automàtic
Logaritmes
Proteïnes
Estadística
title_short Machine learning: how much does it tell about protein folding rates?
title_full Machine learning: how much does it tell about protein folding rates?
title_fullStr Machine learning: how much does it tell about protein folding rates?
title_full_unstemmed Machine learning: how much does it tell about protein folding rates?
title_sort Machine learning: how much does it tell about protein folding rates?
dc.creator.none.fl_str_mv Corrales, Marc
Cuscó, Pol
Usmanova, Dinara R.
Chen, Heng-Chang
Bogatyreva, Natalya S.
Filion, Guillaume
Ivankov, Dmitry N.
author Corrales, Marc
author_facet Corrales, Marc
Cuscó, Pol
Usmanova, Dinara R.
Chen, Heng-Chang
Bogatyreva, Natalya S.
Filion, Guillaume
Ivankov, Dmitry N.
author_role author
author2 Cuscó, Pol
Usmanova, Dinara R.
Chen, Heng-Chang
Bogatyreva, Natalya S.
Filion, Guillaume
Ivankov, Dmitry N.
author2_role author
author
author
author
author
author
dc.subject.none.fl_str_mv Aprenentatge automàtic
Logaritmes
Proteïnes
Estadística
topic Aprenentatge automàtic
Logaritmes
Proteïnes
Estadística
description The prediction of protein folding rates is a necessary step towards understanding the principles of protein folding. Due to the increasing amount of experimental data, numerous protein folding models and predictors of protein folding rates have been developed in the last decade. The problem has also attracted the attention of scientists from computational fields, which led to the publication of several machine learning-based models to predict the rate of protein folding. Some of them claim to predict the logarithm of protein folding rate with an accuracy greater than 90%. However, there are reasons to believe that such claims are exaggerated due to large fluctuations and overfitting of the estimates. When we confronted three selected published models with new data, we found a much lower predictive power than reported in the original publications. Overly optimistic predictive powers appear from violations of the basic principles of machine-learning. We highlight common misconceptions in the studies claiming excessive predictive power and propose to use learning curves as a safeguard against those mistakes. As an example, we show that the current amount of experimental data is insufficient to build a linear predictor of logarithms of folding rates based on protein amino acid composition.
publishDate 2015
dc.date.none.fl_str_mv 2015
2023
2023
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10230/58480
http://dx.doi.org/10.1371/journal.pone.0143166
url http://hdl.handle.net/10230/58480
http://dx.doi.org/10.1371/journal.pone.0143166
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv PLoS ONE. 2015 Nov 25;10(11):e0143166
info:eu-repo/grantAgreement/EC/FP7/335980
dc.rights.none.fl_str_mv http://creativecommons.org/licenses/by/4.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv http://creativecommons.org/licenses/by/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Public Library of Science (PLoS)
publisher.none.fl_str_mv Public Library of Science (PLoS)
dc.source.none.fl_str_mv reponame:Recercat. Dipósit de la Recerca de Catalunya
instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
instname_str Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
reponame_str Recercat. Dipósit de la Recerca de Catalunya
collection Recercat. Dipósit de la Recerca de Catalunya
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869410744290770944
score 15.812429