Antonios Anastasopoulos
16 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Leveraging Translations for Speech Transcription in Low-resource Settings
2018
Recently proposed data collection frameworks for endangered language documentation aim not only to collect speech in the language of interest, but also to collect translations into a high-resource language that will render the collected resource …
-
Choosing Transfer Languages for Cross-Lingual Learning
2019 · arXiv (Cornell University)
Cross-lingual transfer, where a high-resource transfer language is used to improve the accuracy of a low-resource task language, is now an invaluable tool for improving performance of natural language processing (NLP) on low-resource languages. However, …
-
A small Griko-Italian speech translation corpus
2018 · White Rose Research Online (University of Leeds, The University of Sheffield, University of York)
This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which …
-
An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages
2016
For many low-resource languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Translated speech data is potentially valuable for documenting endangered languages or for training speech translation systems. A …
-
Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks
2019 · arXiv (Cornell University)
Learning general representations of text is a fundamental problem for many natural language understanding (NLU) tasks. Previously, researchers have proposed to use language model pre-training and multi-task learning to learn robust representations. However, these methods …
-
TICO-19: the Translation Initiative for COvid-19
2020
Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, Sylwia …
-
OCR Post Correction for Endangered Language Texts
2020
There is little to no data available to build natural language processing models for most endangered languages. However, textual data in these languages often exists in formats that are not machine-readable, such as paper books …
-
Automatic Interlinear Glossing for Under-Resourced Languages Leveraging Translations
2020
Interlinear Glossed Text (IGT) is a widely used format for encoding linguistic information in language documentation projects and scholarly papers. Manual production of IGT takes time and requires linguistic expertise. We attempt to address this …
-
When is Wall a Pared and when a Muro? -- Extracting Rules Governing Lexical Selection
2021 · arXiv (Cornell University)
Learning fine-grained distinctions between vocabulary items is a key challenge in learning a new language. For example, the noun "wall" has different lexical manifestations in Spanish -- "pared" refers to an indoor wall while "muro" …
-
AUTOLEX: An Automatic Framework for Linguistic Exploration
2022 · arXiv (Cornell University)
Each language has its own complex systems of word, phrase, and sentence construction, the guiding principles of which are often summarized in grammar descriptions for the consumption of linguists or language learners. However, manual creation …
-
Systematic Inequalities in Language Technology Performance across the World's Languages
2021 · arXiv (Cornell University)
Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown enormously over the …
-
A case study on using speech-to-translation alignments for language\n documentation
2017 · arXiv (Cornell University)
For many low-resource or endangered languages, spoken language resources are\nmore likely to be annotated with translations than with transcriptions. Recent\nwork exploits such annotations to produce speech-to-translation alignments,\nwithout access to any text transcriptions. We investigate whether …
-
Noisy Parallel Data Alignment
2023 · arXiv (Cornell University)
An ongoing challenge in current natural language processing is how its major advancements tend to disproportionately favor resource-rich languages, leaving a significant number of under-resourced languages behind. Due to the lack of resources required to …
-
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers
2024 · arXiv (Cornell University)
Identifying linguistic differences between dialects of a language often requires expert knowledge and meticulous human analysis. This is largely due to the complexity and nuance involved in studying various dialects. We present a novel approach …
-
An Attentional Model for Speech Translation Without Transcription
2016
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, Trevor Cohn. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
-
Tied Multitask Learning for Neural Speech Translation
2018
Antonios Anastasopoulos, David Chiang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.