conference-paper وصول مفتوح

An Effective Text Data Augmentation Method for Filtering Inappropriate Webpages

Research footprint

At a glance

الاستشهادات
0
المراجع
18
Comments
0
Paper overview

Abstract

With the advance of the internet, large-scale inappropriate content can be seen everywhere.It is not easy to use a fixed and static domain name list to filter inappropriate websites because domain names can easily be changed to avoid abnormal detection.The number of normal websites is typically significantly more than that of abnormal websites; therefore, the website data for these two types might be imbalanced.This paper presents a text data augmentation method to augment a blacklist for improving the accuracy of website classification (detection) tasks by synthesizing data for the minority categories using a language model.It also adopts a selection strategy to filter out the toxic output to get suitable synthetic data and then to distill knowledge from a language model into a synthetic dataset.To evaluate the performance of the proposed method, we compared it with easy data augmentation (EDA) and TextSmoothing for a blacklist to filter inappropriate webpages.Experimental results show the proposed data augmentation mechanism can enhance the performance of a blacklist.Also, the proposed method can outperform the other data augmentation methods for filtering inappropriate webpages.

Record transparency

Publication details

DOI
10.1145/3732437.3732756
OpenAlex
W4415167982
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.