SummCoder: An unsupervised framework for extractive text summarization based on deep auto-encoders

[EN] In this paper, we propose SummCoder, a novel methodology for generic extractive text summarization of single documents. The approach generates a summary according to three sentence selection metrics formulated by us: sentence content relevance, sentence novelty, and sentence position relevance....

Descripción completa

Detalles Bibliográficos
Autores: Joshi, Akanksha, Fidalgo Fernández, Eduardo, Alegre Gutiérrez, Enrique, Fernández Robles, Laura
Tipo de recurso: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2019
País:España
Institución:Universidad de León
Repositorio:BULERIA. Repositorio Institucional de la Universidad de León
OAI Identifier:oai:buleria.unileon.es:10612/23128
Acceso en línea:https://hdl.handle.net/10612/23128
Access Level:acceso abierto
Palabra clave:Informática
Ingeniería de sistemas
Extractive text summarization
Auto-encoder
Deep learning
Sentence embedding
TOR darknet
Extractive summarization
3304.05 Sistemas de Reconocimiento de Caracteres
1203.04 Inteligencia Artificial
Descripción
Sumario:[EN] In this paper, we propose SummCoder, a novel methodology for generic extractive text summarization of single documents. The approach generates a summary according to three sentence selection metrics formulated by us: sentence content relevance, sentence novelty, and sentence position relevance. The sentence content relevance is measured using a deep auto-encoder network, and the novelty metric is derived by exploiting the similarity among sentences represented as embeddings in a distributed semantic space. The sentence position relevance metric is a hand-designed feature, which assigns more weight to the first few sentences through a dynamic weight calculation function regulated by the document length. Furthermore, a sentence ranking and a selection technique are developed to generate the document summary by ranking the sentences according to the final score obtained through the fusion of the three sentences selection metrics. We also introduce a new summarization benchmark, Tor Illegal Documents Summarization (TIDSumm) dataset, especially to assist Law Enforcement Agencies (LEAs), that contains two sets of ground truth summaries, manually created, for 100 web documents extracted from onion websites in Tor (The Onion Router) network. Empirical results show that, on DUC 2002, on Blog Summarization, and on TIDSumm datasets, our text summarization approach obtains comparable or better performance than the state-of-the-art methods for different ROUGE metrics.