conference-paper Open access

An Effective Text Data Augmentation Method for Filtering Inappropriate Webpages

Research footprint

At a glance

Citations
0
References
18
Comments
0
Paper overview

Öz

With the advance of the internet, large-scale inappropriate content can be seen everywhere.It is not easy to use a fixed and static domain name list to filter inappropriate websites because domain names can easily be changed to avoid abnormal detection.The number of normal websites is typically significantly more than that of abnormal websites; therefore, the website data for these two types might be imbalanced.This paper presents a text data augmentation method to augment a blacklist for improving the accuracy of website classification (detection) tasks by synthesizing data for the minority categories using a language model.It also adopts a selection strategy to filter out the toxic output to get suitable synthetic data and then to distill knowledge from a language model into a synthetic dataset.To evaluate the performance of the proposed method, we compared it with easy data augmentation (EDA) and TextSmoothing for a blacklist to filter inappropriate webpages.Experimental results show the proposed data augmentation mechanism can enhance the performance of a blacklist.Also, the proposed method can outperform the other data augmentation methods for filtering inappropriate webpages.

Record transparency

Publication details

DOI
10.1145/3732437.3732756
OpenAlex
W4415167982
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.