preprint Open access

Probabilistic Modeling for Novelty Detection with Applications to Fraud\n Identification

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
3
References
95
Comments
0
Paper overview

Abstract

Novelty detection is the unsupervised problem of identifying anomalies in\ntest data which significantly differ from the training set. Novelty detection\nis one of the classic challenges in Machine Learning and a core component of\nseveral research areas such as fraud detection, intrusion detection, medical\ndiagnosis, data cleaning, and fault prevention. While numerous algorithms were\ndesigned to address this problem, most methods are only suitable to model\ncontinuous numerical data. Tackling datasets composed of mixed-type features,\nsuch as numerical and categorical data, or temporal datasets describing\ndiscrete event sequences is a challenging task. In addition to the supported\ndata types, the key criteria for efficient novelty detection methods are the\nability to accurately dissociate novelties from nominal samples, the\ninterpretability, the scalability and the robustness to anomalies located in\nthe training data.\n In this thesis, we investigate novel ways to tackle these issues. In\nparticular, we propose (i) an experimental comparison of novelty detection\nmethods for mixed-type data (ii) an experimental comparison of novelty\ndetection methods for sequence data, (iii) a probabilistic nonparametric\nnovelty detection method for mixed-type data based on Dirichlet process\nmixtures and exponential-family distributions and (iv) an autoencoder-based\nnovelty detection model with encoder/decoder modelled as deep Gaussian\nprocesses.\n

Record transparency

Publication details

DOI
10.48550/arxiv.1903.01730
OpenAlex
W4288420159
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.