article Open access

Tagging a Norwegian Dialect Corpus

  • DSpace repository (University of Tartu)
  • University of Tartu
Research footprint

At a glance

Citations
0
References
18
Comments
0
Paper overview

Abstract

This paper describes an evaluation of five data-driven Part-of-Speech (PoS) taggers for spoken Norwegian.The taggers all rely on different machine learning mechanisms: decision trees, hidden Markov models (HMMs), conditional random fields (CRFs), long-short term memory networks (LSTMs), and convolutional neural networks (CNNs).We go into some of the challenges posed by the task of tagging spoken, as opposed to written, language, and in particular a wide range of dialects as is found in the recordings of the LIA (Language Infrastructure made Accessible) project.The results show that the taggers based on either conditional random fields or neural networks perform much better than the rest, with the LSTM tagger getting the highest score.

Record transparency

Publication details

OpenAlex
W2982668319
Document type
article
Language
EN
Source
DSpace repository (University of Tartu)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.