Eiichiro Sumita
14 papers in the PaperMetrix corpus
Papers by this author
-
Converting Continuous-Space Language Models into <i>N</i> -gram Language Models with Efficient Bilingual Pruning for Statistical Machine Translation
2016 · ACM Transactions on Asian and Low-Resource Language Information Processing
The Language Model (LM) is an essential component of Statistical Machine Translation (SMT). In this article, we focus on developing efficient methods for LM construction. Our main contribution is that we propose a Natural N …
-
Hierarchical Phrase-based Stream Decoding
2015
This paper proposes a method for hierarchical phrase-based stream decoding. A stream decoder is able to take a continuous stream of tokens as input, and segments this stream into word sequences that are translated and …
-
A Prototype Automatic Simultaneous Interpretation System
2016 · International Conference on Computational Linguistics
Simultaneous interpretation allows people to communicate spontaneously across language boundaries, but such services are prohibitively expensive for the general public. This paper presents a fully automatic simultaneous interpretation system to address this problem. Though the …
-
MuTUAL: A Controlled Authoring Support System Enabling Contextual Machine Translation.
2016 · International Conference on Computational Linguistics
The paper introduces a web-based authoring support system, MuTUAL, which aims to help writers create multilingual texts. The highlighted feature of the system is that it enables machine translation (MT) to generate outputs appropriate to …
-
Assessing translation ability through vocabulary ability assessment
2016 · International Joint Conference on Artificial Intelligence
Translation ability is known as one of the most difficult language abilities to measure. A typical method of measuring translation ability involves asking translators to translate sentences and to request professional evaluators to grade the …
-
Instance Weighting for Neural Machine Translation Domain Adaptation
2017
Instance weighting has been widely applied to phrase-based machine translation domain adaptation. However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.
-
NOVA
2018 · ACM Transactions on Asian and Low-Resource Language Information Processing
A feasible and flexible annotation system is designed for joint tokenization and part-of-speech (POS) tagging to annotate those languages without natural definitions of words . This design was motivated by the fact that word separators …
-
Improving Neural Machine Translation through Phrase-based Forced Decoding
2017 · International Joint Conference on Natural Language Processing
Compared to traditional statistical machine translation (SMT), neural machine translation (NMT) often sacrifices adequacy for the sake of fluency. We propose a method to combine the advantages of traditional SMT and NMT by exploiting an …
-
A System for Worldwide COVID-19 Information Aggregation
2020 · arXiv (Cornell University)
The global pandemic of COVID-19 has made the public pay close attention to related news, covering various domains, such as sanitation, treatment, and effects on education. Meanwhile, the COVID-19 condition is very different among the …
-
User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization
2021 · arXiv (Cornell University)
Morphological analysis (MA) and lexical normalization (LN) are both important tasks for Japanese user-generated text (UGT). To evaluate and compare different MA/LN systems, we have constructed a publicly available Japanese UGT corpus. Our corpus comprises …
-
YANMTT: Yet Another Neural Machine Translation Toolkit
2021 · arXiv (Cornell University)
In this paper we present our open-source neural machine translation (NMT) toolkit called "Yet Another Neural Machine Translation Toolkit" abbreviated as YANMTT which is built on top of the Transformers library. Despite the growing importance …
-
ASPEC: Asian Scientific Paper Excerpt Corpus
2016
In this paper, we describe the details of the ASPEC (Asian Scientific Paper Excerpt Corpus), which is the first large-size parallel corpus of scientific paper domain.ASPEC was constructed in the Japanese-Chinese machine translation project conducted …
-
Sentence Embedding for Neural Machine Translation Domain Adaptation
2017
Although new corpora are becoming increasingly available for machine translation, only those that belong to the same or similar domains are typically able to improve translation performance. Recently Neural Machine Translation (NMT) has become prominent …
-
Syntax-Directed Attention for Neural Machine Translation
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Attention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window …