Machine Learning Approaches in Bioinformatics: Advances in Transcription and Protein Fitness Prediction
As we move deeper into the information age, bioinformatics has become increasingly important in modern biology, largely due to its critical role in processing and analyzing the vast amounts of complex data generated in the field. Traditional methods are often overwhelmed by the large volume and comp...
| Autor: | |
|---|---|
| Tipo de recurso: | tesis doctoral |
| Estado: | Versión aceptada para publicación |
| Fecha de publicación: | 2023 |
| País: | España |
| Institución: | Universidad de Burgos (UBU) |
| Repositorio: | Repositorio Institucional de la Universidad de Burgos (RIUBU) |
| OAI Identifier: | oai:riubu.ubu.es:10259/9060 |
| Acceso en línea: | http://hdl.handle.net/10259/9060 |
| Access Level: | acceso embargado |
| Palabra clave: | Bioinformatics Machine learning Transcription start site Deep learning Protein fitness Bioinformática Aprendizaje automático Sitio de inicio de la transcripción Aprendizaje profundo Aptitud de la proteína Informática Computer science 1203.04 Inteligencia Artificial |
| Sumario: | As we move deeper into the information age, bioinformatics has become increasingly important in modern biology, largely due to its critical role in processing and analyzing the vast amounts of complex data generated in the field. Traditional methods are often overwhelmed by the large volume and complexity of this data, positioning machine learning techniques as an optimal solution. Exploring the intersection of machine learning and bioinformatics offers numerous opportunities to develop and improve computational tools designed to handle and gain insights from these vast datasets. The main objective of this thesis is to develop a thorough exploration of the possibilities of machine learning in the field of bioinformatics, with a particular focus on specific problems such as transcription start and protein fitness prediction. Furthermore, given the similarities between bioinformatics sequence data and the natural language processing domain, the research emphasizes the use of sequence-based methods. Our research has resulted in several contributions to the field in the form of three scientific papers. The first two focus on transcription start prediction. In the first, we discovered that the integration of biophysical simulations in conjunction with the DNA sequence can improve the results of machine learning methods. Additionally, in our second paper we concluded that, while support vector machines have been a traditional choice for transcription start prediction, our research suggests that deep learning methods outperform them, marking a paradigm shift in the field. In addition, we presented custom-built datasets using Ensembl data, providing a valuable resource for future studies. The third paper addresses the issue of protein fitness prediction specifically in scarce dataset scenarios and concludes that deep transfer learning methods get established as the best alternative when compared with other strategies well suited for such situations, such as semi-supervised learning. |
|---|