Diffusion models: task and domain adaptation for training efficiency

Diffusion Probabilistic Models represent a novel deep learning architecture that surpasses existing state-of-the-art methodologies across various fields and applications. The current work will exhaustively study and describe how they work, reviewing their limitations and comparing them with other st...

ver descrição completa

Detalhes bibliográficos
Autor: Antona I Pizà, Héctor
Tipo de documento: dissertação
Data de publicação:2023
País:España
Recursos:Universitat Politècnica de Catalunya (UPC)
Repositório:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglês
OAI Identifier:oai:upcommons.upc.edu:2117/401976
Acesso em linha:https://hdl.handle.net/2117/401976
Access Level:Acceso aberto
Palavra-chave:Machine learning
Diffusion Probabilistic Models
Deep Learning
Adaptive learning
Image inpainting
Image colorization
Image-to-image translation
Training efficiency
Aprenentatge automàtic
Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic
Descrição
Resumo:Diffusion Probabilistic Models represent a novel deep learning architecture that surpasses existing state-of-the-art methodologies across various fields and applications. The current work will exhaustively study and describe how they work, reviewing their limitations and comparing them with other state-of-the-art models they compete with. One of the prominent challenges associated with Diffusion Models lies in their computational requirements, demanding substantial computational resources and extended training time. The study acknowledges and focuses on this inconvenience. This study proposes an approach that takes advantage of Diffusion Model's versatility in both task-focused and domain-focused adaptive learning to optimize and accelerate its training. Departing from a base model focused on a human faces inpainting task, it will be adapted to fulfill a human faces colorization task, which is accomplished with significantly greater efficiency compared to training from scratch. Furthermore, the model will be further adapted to widen its scope and be able to colorize images from domains different from human faces, concretely landscapes images. This will once again be accomplished with notable swiftness compared to conventional approaches.