DistilBERT News Article Classification: Overcoming Data Challenges and Improving Deep Learning
At a glance
- Citations
- 1
- References
- 0
- Comments
- 0
Abstract
Deep learning has revolutionized natural language processing by enabling complex models to classify text. This paper investigates using DistilBERT, a pre-trained transformer model, to classify news stories into World, Sport, Business, and Sci/Tech. To ensure model robustness, class imbalance, missing values, and duplicate entries in the dataset were carefully rectified. Five-crop augmentation and random flips were utilized to boost data variability and generalization. Model performance and embeddings were also visualized using t-SNE embeddings. Learning rate modifications and batch size trials optimized the model and balanced computing cost and performance. The project examined manual linear layer configuration to demonstrate mathematical transformations and deep learning architecture adaptability. DistilBERT was developed to extract important textual elements. After processing, these traits were categorized. Model metrics such as accuracy, precision, recall, and the F1 score demonstrated performance improvements. Deep learning pipelines can overcome data challenges and produce cutting-edge NLP results. This study explains gradient behaviour and pre-trained weight sensitivity. This study also reveals the need for hyperparameter tuning to optimize pre-trained models across domains. The case study advances transfer learning research and advises NLP practitioners. By increasing the learning rate, researchers and developers can improve pre-trained language models like DistilBERT.
Publication details
- DOI
- 10.1109/icaeca63854.2025.11012287
- OpenAlex
- W4411206276
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.