A comparative analysis of encoder only and decoder only models in intent classification and sentiment analysis: navigating the trade-offs in model size and performance

Intent classification and sentiment analysis stand as pivotal tasks in natural language understanding (NLU), with applications ranging from virtual assistants to customer service. The advent of transformer-based models has significantly enhanced the performance of various NLP tasks, with encoder-onl...

ver descrição completa

Detalhes bibliográficos
Autores: Benayas Alamos, Alberto José, Sicilia Urbán, Miguel Ángel|||0000-0003-3067-4180, Mora Cantallops, Marçal|||0000-0002-2480-1078
Tipo de documento: artigo
Data de publicação:2024
País:España
Recursos:Universidad de Alcalá (UAH)
Repositório:e_Buah Biblioteca Digital Universidad de Alcalá
Idioma:inglês
OAI Identifier:oai:ebuah.uah.es:10017/68273
Acesso em linha:http://hdl.handle.net/10017/68273
https://dx.doi.org/10.1007/s10579-024-09796-y
Access Level:Acceso aberto
Palavra-chave:Intent classification
Sentiment analysis
Large language models
Conversational AI
Informática
Computer science
Descrição
Resumo:Intent classification and sentiment analysis stand as pivotal tasks in natural language understanding (NLU), with applications ranging from virtual assistants to customer service. The advent of transformer-based models has significantly enhanced the performance of various NLP tasks, with encoder-only architectures gaining prominence for their effectiveness. More recently, there has been a surge in the development of larger and more powerful decoder-only models, traditionally employed for text generation tasks. This paper aims to answer the question of whether the colossal scale of newer decoder-only language models is essential for real-world applications. The investigation involves a performance comparison between these decoder-only models and the well-established encoder-only models specifically in the domains of intent classification and sentiment analysis. The results of our study indicate that, for tasks involving natural language understanding, encoder-only models generally outperform decoder-only models, all while demanding a fraction of the computational resources. This sheds light on the practicality and efficiency of encoder-only architectures in comparison to their decoder-only counterparts in real-world applications, providing valuable insights for the advancement of natural language processing technologies.