Rough Sets
Author(s) -
Hung Son Nguyen,
Quang-Thuy Ha,
Tianrui Li,
Małgorzata Przybyła-Kasperek
Publication year - 2018
Publication title -
lecture notes in computer science
Language(s) - English
Resource type - Book series
SCImago Journal Rank - 0.249
H-Index - 400
eISSN - 1611-3349
pISSN - 0302-9743
DOI - 10.1007/978-3-319-99368-3
Subject(s) - rough set , computer science , emphasis (telecommunications) , set (abstract data type) , artificial intelligence , operations research , telecommunications , mathematics , programming language
We discuss an approximate database engine that we started designing at Infobright, and now we continue its development for Security On-Demand (SOD). At SOD, it is used in everyday data analytics, allowing for fast approximate execution of ad-hoc queries over tens of billions of data rows [1]. In our engine, queries are run against collections of histograms that represent domains of single columns over groupings of consecutively loaded data rows (so-called packrows). Query execution process corresponds to transformation of such granulated summaries of the input data into summaries reflecting query results [2]. We compare our algorithms that generate histogram descriptions of the original data with data quantization methods that are widely used in data mining. We also introduce a new idea of extending SQL with function hist(a) that produces quantized representation of column a by means of merging a’s histograms corresponding to particular packrows into a unified a’s histogram over the whole data. We refer to our recent works on summary-based data visualization [3] and machine learning [4] in order to illustrate several scenarios of utilizing hist in practice.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom