preprint Open access

PULasso: High-dimensional variable selection with presence-only data

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
2
References
19
Comments
0
Paper overview

Abstract

In various real-world problems, we are presented with classification problems with <i>positive and unlabeled data</i>, referred to as presence-only responses. In this article we study variable selection in the context of presence only responses where the number of features or covariates <i>p</i> is large. The combination of <i>presence-only responses</i> and <i>high dimensionality</i> presents both statistical and computational challenges. In this article, we develop the <i>PUlasso</i> algorithm for variable selection and classification with positive and unlabeled responses. Our algorithm involves using the majorization-minimization framework which is a generalization of the well-known expectation-maximization (EM) algorithm. In particular to make our algorithm scalable, we provide two computational speed-ups to the standard EM algorithm. We provide a theoretical guarantee where we first show that our algorithm converges to a stationary point, and then prove that any stationary point within a local neighborhood of the true parameter achieves the minimax optimal mean-squared error under both strict sparsity and group sparsity assumptions. We also demonstrate through simulations that our algorithm outperforms state-of-the-art algorithms in the moderate <i>p</i> settings in terms of classification performance. Finally, we demonstrate that our PUlasso algorithm performs well on a biochemistry example. Supplementary materials for this article are available online.

Record transparency

Publication details

DOI
10.48550/arxiv.1711.08129
OpenAlex
W2770847996
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.