EP-Pred: A machine learning tool for bioprospecting promiscuous ester hydrolases

When bioprospecting for novel industrial enzymes, substrate promiscuity is a desirable property that increases the reusability of the enzyme. Among industrial enzymes, ester hydrolases have great relevance for which the demand has not ceased to increase. However, the search for new substrate promisc...

Descripción completa

Detalles Bibliográficos
Autores: Xiang, Ruite, Fernandez Lopez, Laura, Robles Martín, Ana, Ferrer, Manuel, Guallar, Víctor|||0000-0002-4580-1114
Tipo de recurso: artículo
Fecha de publicación:2022
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/374835
Acceso en línea:https://hdl.handle.net/2117/374835
https://dx.doi.org/10.3390/biom12101529
Access Level:acceso abierto
Palabra clave:Machine learning
Biocatalysts
Bioprospecting
Esterases/lipases
Hydrolases
Supervised learning
Enzims
Simulació per ordinador
Àrees temàtiques de la UPC::Informàtica::Aplicacions de la informàtica::Bioinformàtica
Descripción
Sumario:When bioprospecting for novel industrial enzymes, substrate promiscuity is a desirable property that increases the reusability of the enzyme. Among industrial enzymes, ester hydrolases have great relevance for which the demand has not ceased to increase. However, the search for new substrate promiscuous ester hydrolases is not trivial since the mechanism behind this property is greatly influenced by the active site’s structural and physicochemical characteristics. These characteristics must be computed from the 3D structure, which is rarely available and expensive to measure, hence the need for a method that can predict promiscuity from sequence alone. Here we report such a method called EP-pred, an ensemble binary classifier, that combines three machine learning algorithms: SVM, KNN, and a Linear model. EP-pred has been evaluated against the Lipase Engineering Database together with a hidden Markov approach leading to a final set of ten sequences predicted to encode promiscuous esterases. Experimental results confirmed the validity of our method since all ten proteins were found to exhibit a broad substrate ambiguity.