KI-Mix: Enhancing Cyber Threat Detection in Incomplete Supervision Setting Through Knowledge-informed Pseudo-anomaly Generation
At a glance
- Citations
- 0
- References
- 23
- Comments
- 0
Abstract
Data-driven methodologies have exhibited remarkable performance in identifying various cyber threats. However, obtaining well-labeled training samples is enormously expensive and often challenging when tackling practical cyber-security problems, due to the cost and difficulties in data annotation. To address this issue, we propose KI-Mix, a novel pseudo-anomaly generation algorithm for cyber threat detection on the basis of the limited labeled anomalies and a large volume of unlabeled data. In a nutshell, KI-Mix incorporates security domain knowledge into data interpolation to capture more labeled data to facilitate semi-supervised detection on cyber anomalies. We compare the performance of KI-Mix with several commonly applied augmentation techniques, such as Mixup and CutMix to evaluate its effectiveness in limited annotated data settings. Through extensive experiments on five security datasets covering various aspects of network threats, we demonstrate that KI-Mix outperforms other methods with equivalent baseline models. Notably, KI-Mix is model-agnostic to enable any data-driven threats detection models to handle incomplete supervision problems in real-world cyber threat detection.
Publication details
- DOI
- 10.1109/smc54092.2024.10830942
- OpenAlex
- W4406611900
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.