CASOS: a subspace method for anomaly detection in high dimensional astronomical databases | Zendy

Henrion Marc | Zendy; Hand David J. | Zendy; Gandy Axel | Zendy; Mortlock Daniel J. | Zendy

AI Assistant Blog Pricing

Home ZAIA Blog

Premium

CASOS: a subspace method for anomaly detection in high dimensional astronomical databases

Author(s) -

Henrion Marc,

Hand David J.,

Gandy Axel,

Mortlock Daniel J.

Publication year - 2013

Publication title -

statistical analysis and data mining: the asa data science journal

Language(s) - English

Resource type - Journals

SCImago Journal Rank - 0.381

H-Index - 33

eISSN - 1932-1872

pISSN - 1932-1864

DOI - 10.1002/sam.11167

Subject(s) - anomaly detection , linear subspace , subspace topology , computer science , outlier , anomaly (physics) , masking (illustration) , data mining , object (grammar) , missing data , clustering high dimensional data , pattern recognition (psychology) , algorithm , artificial intelligence , cluster analysis , mathematics , machine learning , art , physics , geometry , visual arts , condensed matter physics

We develop a novel algorithm for detecting anomalies. Our method has been developed to suit the challenging task of detecting anomalous sources in cross‐matched astronomical survey data. Our algorithm computes anomaly scores in lower‐dimensional subspaces of the data. By subspaces we mean, in this work, subsets of the original data variables. Our technique presents several advantages over existing methods: it can work directly on data with missing values; it addresses some of the problems posed by high‐dimensional data spaces; it is less susceptible to a masking effect from irrelevant features; it can be easily adapted to suit specific needs and it allows an easier interpretation of why a given object has a high combined anomaly score. One drawback of our method is that it cannot detect outliers that are only apparent in high‐dimensional spaces. Anomaly scores are computed using a nearest neighbor (NN) technique, but the algorithm works with any other method computing numerical anomaly scores. We demonstrate the properties of our algorithm and evaluate its performance on both simulated and real datasets. We show that it is capable of outperforming state‐of‐the‐art, full‐dimensional approaches in some situations. © 2013 Wiley Periodicals, Inc. Statistical Analysis and Data Mining 6: 53–72, 2013

This content is not available in your region!

Continue researching here.

Having issues? You can contact us here

Accelerating Research