Finding maximal sequential patterns in text document collections and single documents

In this paper, two algorithms for discovering all the Maximal Sequential Patterns (MSP) in a document collection and in a single document are presented. The proposed algorithms follow the “pattern-growth strategy” where small frequent sequences are found first with the goal of growing them to obtain...

ver descrição completa

Detalhes bibliográficos
Autores: RENE ARNULFO GARCIA HERNANDEZ, JOSE FRANCISCO MARTINEZ TRINIDAD, JESUS ARIEL CARRASCO OCHOA
Tipo de documento: artigo
Estado:Versión aceptada para publicación
Data de publicação:2010
País:México
Recursos:Instituto Nacional de Astrofísica, Óptica y Electrónica
Repositório:Repositorio Institucional del INAOE
Idioma:inglês
OAI Identifier:oai:inaoe.repositorioinstitucional.mx:1009/1403
Acesso em linha:http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/1403
Access Level:Acceso aberto
Palavra-chave:info:eu-repo/classification/Text mining/Text mining
info:eu-repo/classification/Maximal sequential patterns/Maximal sequential patterns
info:eu-repo/classification/cti/1
info:eu-repo/classification/cti/12
info:eu-repo/classification/cti/1203
Descrição
Resumo:In this paper, two algorithms for discovering all the Maximal Sequential Patterns (MSP) in a document collection and in a single document are presented. The proposed algorithms follow the “pattern-growth strategy” where small frequent sequences are found first with the goal of growing them to obtain MSP. Our algorithms process the documents in an incremental way avoiding re-computing all the MSP when new documents are added. Experiments showing the performance of our algorithms and comparing against GSP, DELISP, GenPrefixSpan and cSPADE algorithms over public standard databases are also presented.