Researcher profile

Jörg Tiedemann

8 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign

    2018 · Työväentutkimus Vuosikirja

    We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …

  2. An Analysis of Encoder Representations in Transformer-Based Machine Translation

    2018

    The attention mechanism is a successful technique in modern NLP, especially in tasks like machine translation. The recently proposed network architecture of the Transformer is based entirely on attention mechanisms and achieves new state of …

  3. Analysing concatenation approaches to document-level NMT in two different domains

    2019

    In this paper, we investigate how different aspects of discourse context affect the performance of recent neural MT systems. We describe two popular datasets covering news and movie subtitles and we provide a thorough analysis …

  4. OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

    2016

    We present a new major release of the OpenSubtitles collection of parallel corpora.The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion …

  5. Efficient Word Alignment with Markov Chain Monte Carlo

    2016 · ˜The œPrague Bulletin of Mathematical Linguistics

    Abstract We present EFMARAL, a new system for efficient and accurate word alignment using a Bayesian model with Markov Chain Monte Carlo (MCMC) inference. Through careful selection of data structures and model architecture we are …

  6. Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task

    2016 · International Conference on Computational Linguistics

    We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …

  7. Continuous multilinguality with language vectors

    2017

    Most existing models for multilingual natural language processing (NLP) treat language as a discrete category, and make predictions for either one language or the other. In contrast, we propose using continuous vector representations of language. …

  8. OPUS-MT – Building open translation services for the World

    2020 · Työväentutkimus Vuosikirja

    This paper presents OPUS-MT a project that focuses on the development of free resources and tools for machine translation. The current status is a repository of over 1,000 pre-trained neural machine translation models that are …