z-logo
open-access-imgOpen Access
Parallel noise eliminate: A parallel noise elimination algorithm for massive text categorization
Author(s) -
Xiaojuan Hu,
Lei Liu,
Ningjia Qiu,
Meng Li
Publication year - 2018
Publication title -
journal of algorithms and computational technology
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.234
H-Index - 13
eISSN - 1748-3026
pISSN - 1748-3018
DOI - 10.1177/1748301818779047
Subject(s) - categorization , noise (video) , computer science , pattern recognition (psychology) , key (lock) , process (computing) , artificial intelligence , algorithm , noise measurement , speech recognition , data mining , noise reduction , computer security , operating system , image (mathematics)
Noise data in text are one of the main factors affecting the quality of text categorization. A parallel noise data elimination algorithm based on principal component analysis method and term frequency-inverse document frequency method for the noise data issue of massive text categorization is proposed. Five types of noise data which may occur during text categorization process are analyzed and summarized in this paper. Before text categorization, a redundant noise elimination algorithm based on key feature selection is presented for redundant noise features. During the process of text categorization, the error noise detection algorithm is given for inaccurate noise features. The proposed method is compared with other four typical noise processing methods in different noise ratios on two common corpora. The results show that the proposed method is feasible and can maintain more stable and excellent classification performance and lower running time.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom