conference-paper

Pilot Testing Transformers-based Models for Multi-class Classification of Greek Text

Research footprint

At a glance

Citations
0
References
21
Comments
0
Paper overview

Öz

Text data are generated on a daily basis through various social media platforms. From a computational perspective, these text data can be analyzed using machine learning algorithms. The Bidirectional Encoder Representations from Transformers (BERT) is a model specifically tailored for natural language processing tasks, such as text classification, which requires understanding of word context. GreekBERT is a monolingual, pre-trained transformers-based language model developed for the analysis of Greek text. This paper investigates the fine-tuning of GreekBERT and of other transformers-based models, including few-shot and zero-shot learning techniques, for the multi-class classification of Greek tweets, with the objective of facilitating further informatics research. To fine-tune the transformers models, we perform experiments utilizing a training dataset of 2,000 tweets based on eleven categories, while evaluating on a distinct set of 120 tweets. The experimental results indicate issues such as mislabeled tweets, suboptimal evaluation metrics, with accuracy levels at or below 35%, and a significant class imbalance. This experimentation underscores the challenges associated with adapting pre-trained language models for specific downstream tasks and highlights the need for ongoing experimentation and optimization of the models.

Record transparency

Publication details

DOI
10.1109/dessert65323.2024.11122141
OpenAlex
W4413392912
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.