PULasso: High-dimensional variable selection with presence-only data
At a glance
- Citations
- 2
- References
- 19
- Comments
- 0
Abstract
In various real-world problems, we are presented with classification problems with <i>positive and unlabeled data</i>, referred to as presence-only responses. In this article we study variable selection in the context of presence only responses where the number of features or covariates <i>p</i> is large. The combination of <i>presence-only responses</i> and <i>high dimensionality</i> presents both statistical and computational challenges. In this article, we develop the <i>PUlasso</i> algorithm for variable selection and classification with positive and unlabeled responses. Our algorithm involves using the majorization-minimization framework which is a generalization of the well-known expectation-maximization (EM) algorithm. In particular to make our algorithm scalable, we provide two computational speed-ups to the standard EM algorithm. We provide a theoretical guarantee where we first show that our algorithm converges to a stationary point, and then prove that any stationary point within a local neighborhood of the true parameter achieves the minimax optimal mean-squared error under both strict sparsity and group sparsity assumptions. We also demonstrate through simulations that our algorithm outperforms state-of-the-art algorithms in the moderate <i>p</i> settings in terms of classification performance. Finally, we demonstrate that our PUlasso algorithm performs well on a biochemistry example. Supplementary materials for this article are available online.
Publication details
- DOI
- 10.48550/arxiv.1711.08129
- OpenAlex
- W2770847996
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.