USING FREQUENT PATTERN MINING ALGORITHMS IN TEXT ANALYSIS
Author(s) -
Piotr Ożdżyński
Publication year - 2017
Publication title -
information system in management
Language(s) - English
Resource type - Journals
eISSN - 2544-1728
pISSN - 2084-5537
DOI - 10.22630/isim.2017.6.3.19
Subject(s) - computer science , field (mathematics) , data science , process (computing) , software , data mining , tracking (education) , software engineering , operating system , programming language , mathematics , psychology , pedagogy , pure mathematics
In text mining, effectiveness of methods depends on d cument representations. The ones based on frequent word sequences are used in uch tasks as categorization, clustering and topic modelling. In the paper a comp arison of different algorithms for finding frequent word sequences is presented. There are considered techniques dedicated for market basket analysis such as GSP an d PrefixSpan as well as a method based on a suffix array. The investigated te chniques are compared with the new approach of searching maximum frequent word seq uences in document sets. Performance of the algorithms is examined taking in to account execution times for the considered test collections.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom