conference-paper

Efficient Privacy-Preserving Data Annotation via Active PrivBayes Synthetic Data Generation

Research footprint

At a glance

Citations
0
References
29
Comments
0
Paper overview

Öz

Human involvement is essential for building training datasets and AI models, especially in pervasive computing where sensitive real-world data are used for novel applications. Data annotation is typically performed by humans to prepare data for AI applications. While privacy-preserving techniques such as federated learning and secure computation have been widely studied for AI training and inference, they do not cover the entire AI lifecycle or address human-data interaction. This study pro-poses and demonstrates a method for efficient privacy-preserving human annotation using synthetic data generation integrated with both active learning and differential privacy. Specifically, the method iteratively generates synthetic data containing only explanatory variables from real-world data with differential privacy guarantees, incorporating the acquisition function from active learning into the generation process. Experimental results demonstrate that the proposed method improves the efficiency of human annotation work compared to a simple combination of existing methods. Moreover, it overcomes the trade-off between privacy preservation and AI model performance, achieving both stricter privacy guarantees and higher model accuracy.

Record transparency

Publication details

DOI
10.1109/percomworkshops65533.2025.00159
OpenAlex
W4411447255
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.