Engineering data-sharing practices for a fair and trustworthy AI

Machine learning (ML) technology may discriminate toward specific social groups. For example, recent research have revealed that ML applications are more likely to fail in identifying women than males in hospitals. Recent research has identified the data used to train these models as one of the caus...

Descripción completa

Detalles Bibliográficos
Autor: Giner Miguelez, Joan
Tipo de recurso: tesis doctoral
Estado:Versión publicada
Fecha de publicación:2024
País:España
Institución:CBUC, CESCA
Repositorio:TDR. Tesis Doctorales en Red
OAI Identifier:oai:www.tdx.cat:10803/692115
Acceso en línea:http://hdl.handle.net/10803/692115
Access Level:acceso abierto
Palabra clave:compartició de dades
compartición de datos
data-sharing practices
aprenentatge automàtic
aprendizaje automático
machine learning
IA confiable
trustworthy AI
equitat a la IA
equidad en la IA
fairness
documentació de dades
documentación de datos
data documentation
Ciencies de la computació
004
Descripción
Sumario:Machine learning (ML) technology may discriminate toward specific social groups. For example, recent research have revealed that ML applications are more likely to fail in identifying women than males in hospitals. Recent research has identified the data used to train these models as one of the causes of these issues. The research community has proposed guidelines to detect the dimensions that can generate these discriminatory behaviors. However, these proposals lack a set structure, restricting their computation and the creation of engineering approaches built upon them. This thesis presents a domain-specific language to document data for ML. This language has served as a basis for creating the responsible AI extension of \emph{Croissant}, a standard adopted by major search engines, such as \emph{Google Dataset Search}. Moreover, this thesis studies the use of large language models (LLM) to automatically create data documentation and the readiness of scientific data for its use in ML.