z-logo
open-access-imgOpen Access
Efficient Pairwise Document Similarity Computation in Big Datasets
Author(s) -
Papias Niyigena,
Zuping Zhang,
Weiqi Li,
Jun Long
Publication year - 2015
Publication title -
international journal of database theory and application
Language(s) - English
Resource type - Journals
eISSN - 2207-9688
pISSN - 2005-4270
DOI - 10.14257/ijdta.2015.8.4.07
Subject(s) - computer science , pairwise comparison , similarity (geometry) , computation , information retrieval , data mining , artificial intelligence , theoretical computer science , algorithm , image (mathematics)
Support Vector Machine (SVM) is extremely powerful and widely accepted classifier in the field of machine learning due to its better generalization capability. However, SVM is not suiTable for large scale dataset due to its high computational complexity. The computation and storage requirement increases tremendously for large dataset. In this paper, we have proposed a MapReduce based SVM for large scale data. MapReduce is a distributed programming model which works on large scale dataset by dividing the huge datasets in smaller chunks. MapReduce distribution model works on several frame works like Hadoop Twister and so on. In this paper, we have analyzed the impact of penalty and kernel parameters on the performance of parallel SVM. The experimental result shows that the number of support vectors and predictive accuracy of SVM is affected by the choice of these parameters. From experimental results, it is also analyzed that the computation time taken by the SVM with multi-node cluster is less as compared to the single node cluster for large dataset.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom