Machine learning methods of property prediction in drug discovery
Molecular properties are highly important in drug discovery with binding affinity being crucial for selecting hit and lead candidates. Existing machine learning models often lack speed for large-scale screening or lack generalizability. This thesis develops faster models for binding affinity predict...
| Autor: | |
|---|---|
| Tipo de documento: | tese |
| Estado: | Versão publicada |
| Data de publicação: | 2024 |
| País: | España |
| Recursos: | CBUC, CESCA |
| Repositório: | TDR. Tesis Doctorales en Red |
| OAI Identifier: | oai:www.tdx.cat:10803/692716 |
| Acesso em linha: | http://hdl.handle.net/10803/692716 |
| Access Level: | Acceso aberto |
| Palavra-chave: | Machine learning Active learning Binding affinity Acid-base dissociation constant Interpretability Aprendizaje automático Aprendizaje activo Afinidad de unión Interpretabilidad Constante de disociación ácido-base 577 |
| Resumo: | Molecular properties are highly important in drug discovery with binding affinity being crucial for selecting hit and lead candidates. Existing machine learning models often lack speed for large-scale screening or lack generalizability. This thesis develops faster models for binding affinity prediction, surpassing previous voxel-based models. It compares 2D and 3D models through extensive benchmarks, highlighting scenarios where simpler 2D models outperform complex 3D models and showing how neural networks improve with supervised and unsupervised pretraining. The work also shows how combining 2D and 3D models yields state-of-the-art results in active learning. To address the difficult interpretability of neural network predictions, a new application is introduced to identify key contributing areas of the input space. Additionally, correct protonation states are essential for preparing molecular libraries. The thesis extends 3D graph-based models for micro-pKa estimation, presenting a new model and application that is more robust to diverse chemical compounds than existing methods. |
|---|