conference-paper

Sentence Classification with Imbalanced Data for Health Applications

  • UEA Digital Repository (University of East Anglia)
  • University of East Anglia
Research footprint

At a glance

Citations
1
References
40
Comments
0
Paper overview

Abstract

Identifying and extracting reports of medications, their abuse or adverse effects from social media is a challenging task. In social media, relevant reports are very infrequent, causes imbalanced class distribution for machine learning algorithms. Learning algorithms typically designed to optimize the overall accuracy without considering the relative distribution of each class. Thus, imbalanced class distribution is problematic as learning algorithms have low predictive accuracy for the infrequent class. Moreover, social media represents natural linguistic variation in creative language expressions. In this paper, we have used a combination of data balancing and neural language representation techniques to address the challenges. Specifically, we participated the shared tasks 1, 2 (all languages), 4, and 3 (only the span detection, no normalization was attempted) in Social Media Mining for Health applications (SMM4H) 2020 (Klein et al., 2020). The results show that with the proposed methodology recall scores are better than the precision scores for the shared tasks. The recall score is also better compared to the mean score of the total submissions. However, the F1-score is worse than the mean score except for task 2 (French).

Record transparency

Publication details

OpenAlex
W3119894181
Document type
conference-paper
Language
EN
Source
UEA Digital Repository (University of East Anglia)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.