Machine Learning Approaches in Bioinformatics: Advances in Transcription and Protein Fitness Prediction

As we move deeper into the information age, bioinformatics has become increasingly important in modern biology, largely due to its critical role in processing and analyzing the vast amounts of complex data generated in the field. Traditional methods are often overwhelmed by the large volume and comp...

Descripción completa

Detalles Bibliográficos
Autor: Barbero Aparicio, José Antonio
Tipo de recurso: tesis doctoral
Estado:Versión aceptada para publicación
Fecha de publicación:2023
País:España
Institución:Universidad de Burgos (UBU)
Repositorio:Repositorio Institucional de la Universidad de Burgos (RIUBU)
OAI Identifier:oai:riubu.ubu.es:10259/9060
Acceso en línea:http://hdl.handle.net/10259/9060
Access Level:acceso embargado
Palabra clave:Bioinformatics
Machine learning
Transcription start site
Deep learning
Protein fitness
Bioinformática
Aprendizaje automático
Sitio de inicio de la transcripción
Aprendizaje profundo
Aptitud de la proteína
Informática
Computer science
1203.04 Inteligencia Artificial
Descripción
Sumario:As we move deeper into the information age, bioinformatics has become increasingly important in modern biology, largely due to its critical role in processing and analyzing the vast amounts of complex data generated in the field. Traditional methods are often overwhelmed by the large volume and complexity of this data, positioning machine learning techniques as an optimal solution. Exploring the intersection of machine learning and bioinformatics offers numerous opportunities to develop and improve computational tools designed to handle and gain insights from these vast datasets. The main objective of this thesis is to develop a thorough exploration of the possibilities of machine learning in the field of bioinformatics, with a particular focus on specific problems such as transcription start and protein fitness prediction. Furthermore, given the similarities between bioinformatics sequence data and the natural language processing domain, the research emphasizes the use of sequence-based methods. Our research has resulted in several contributions to the field in the form of three scientific papers. The first two focus on transcription start prediction. In the first, we discovered that the integration of biophysical simulations in conjunction with the DNA sequence can improve the results of machine learning methods. Additionally, in our second paper we concluded that, while support vector machines have been a traditional choice for transcription start prediction, our research suggests that deep learning methods outperform them, marking a paradigm shift in the field. In addition, we presented custom-built datasets using Ensembl data, providing a valuable resource for future studies. The third paper addresses the issue of protein fitness prediction specifically in scarce dataset scenarios and concludes that deep transfer learning methods get established as the best alternative when compared with other strategies well suited for such situations, such as semi-supervised learning.