Low-cost training of image-to-image diffusion models with incremental learning and task/domain adaptation

Diffusion models specialized in image-to-image translation tasks, like inpainting and colorization, have outperformed the state of the art, yet their computational requirements are exceptionally demanding. This study analyzes different strategies to train image-to-image diffusion models in a low-res...

ver descrição completa

Detalhes bibliográficos
Autores: Antona Pizà, Héctor, Otero Calviño, Beatriz|||0000-0002-9194-559X, Tous Liesa, Rubén|||0000-0002-1409-5843
Tipo de documento: artigo
Data de publicação:2024
País:España
Recursos:Universitat Politècnica de Catalunya (UPC)
Repositório:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglês
OAI Identifier:oai:upcommons.upc.edu:2117/407762
Acesso em linha:https://hdl.handle.net/2117/407762
https://dx.doi.org/10.3390/electronics13040722
Access Level:Acceso aberto
Palavra-chave:Color computer graphics
Deep learning (Machine learning)
Diffusion probabilistic models
Adaptive learning
Transfer learning
Image inpainting
Image colorization
Image-to-image translation
Training efficiency
Infografia en color
Aprenentatge profund
Àrees temàtiques de la UPC::Informàtica::Infografia
Descrição
Resumo:Diffusion models specialized in image-to-image translation tasks, like inpainting and colorization, have outperformed the state of the art, yet their computational requirements are exceptionally demanding. This study analyzes different strategies to train image-to-image diffusion models in a low-resource setting. The studied strategies include incremental learning and task/domain transfer learning. First, a base model for human face inpainting is trained from scratch with an incremental learning strategy. The resulting model achieves an FID score almost equivalent to that of its batch learning equivalent while significantly reducing the training time. Second, the base model is fine-tuned to perform a different task, image colorization, and, in a different domain, landscape images. The resulting colorization models showcase exceptional performances with a minimal number of training epochs. We examine the impact of different configurations and provide insights into the ability of image-to-image diffusion models for transfer learning across tasks and domains.