Jörg Tiedemann
8 papers in the PaperMetrix corpus
Papers by this author
-
Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign
2018 · Työväentutkimus Vuosikirja
We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …
-
An Analysis of Encoder Representations in Transformer-Based Machine Translation
2018
The attention mechanism is a successful technique in modern NLP, especially in tasks like machine translation. The recently proposed network architecture of the Transformer is based entirely on attention mechanisms and achieves new state of …
-
Analysing concatenation approaches to document-level NMT in two different domains
2019
In this paper, we investigate how different aspects of discourse context affect the performance of recent neural MT systems. We describe two popular datasets covering news and movie subtitles and we provide a thorough analysis …
-
OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles
2016
We present a new major release of the OpenSubtitles collection of parallel corpora.The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion …
-
Efficient Word Alignment with Markov Chain Monte Carlo
2016 · The Prague Bulletin of Mathematical Linguistics
Abstract We present EFMARAL, a new system for efficient and accurate word alignment using a Bayesian model with Markov Chain Monte Carlo (MCMC) inference. Through careful selection of data structures and model architecture we are …
-
Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task
2016 · International Conference on Computational Linguistics
We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …
-
Continuous multilinguality with language vectors
2017
Most existing models for multilingual natural language processing (NLP) treat language as a discrete category, and make predictions for either one language or the other. In contrast, we propose using continuous vector representations of language. …
-
OPUS-MT – Building open translation services for the World
2020 · Työväentutkimus Vuosikirja
This paper presents OPUS-MT a project that focuses on the development of free resources and tools for machine translation. The current status is a repository of over 1,000 pre-trained neural machine translation models that are …