preprint Open access

Data Discovery and Anomaly Detection Using Atypicality: Theory

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
2
References
39
Comments
0
Paper overview

Abstract

A central question in the era of 'big data' is what to do with the enormous amount of information. One possibility is to characterize it through statistics, e.g., averages, or classify it using machine learning, in order to understand the general structure of the overall data. The perspective in this paper is the opposite, namely that most of the value in the information in some applications is in the parts that deviate from the average, that are unusual, atypical. We define what we mean by 'atypical' in an axiomatic way as data that can be encoded with fewer bits in itself rather than using the code for the typical data. We show that this definition has good theoretical properties. We then develop an implementation based on universal source coding, and apply this to a number of real world data sets.

Record transparency

Publication details

DOI
10.48550/arxiv.1709.03189
OpenAlex
W2754781333
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.