Xiaoyu Shen
6 papers in the PaperMetrix corpus
Papers by this author
-
Data Augmentation for Multiclass Utterance Classification – A Systematic Study
2020
Utterance classification is a key component in many conversational systems. However, classifying real-world user utterances is challenging, as people may express their ideas and thoughts in manifold ways, and the amount of training data for …
-
Neural Data-to-Text Generation with LM-based Text Augmentation
2021 · arXiv (Cornell University)
For many new application domains for data-to-text generation, the main obstacle in training neural models consists of a lack of training data. While usually large numbers of instances are available on the data side, often …
-
A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation
2022 · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen …
-
Meta Self-Refinement for Robust Learning with Weak Supervision
2022 · arXiv (Cornell University)
Training deep neural networks (DNNs) under weak supervision has attracted increasing research attention as it can significantly reduce the annotation cost. However, labels from weak supervision can be noisy, and the high capacity of DNNs …
-
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
2025 · arXiv (Cornell University)
The alt-text generation task produces concise, context-relevant descriptions of images, enabling blind and low-vision users to access online images. Despite the capabilities of large vision-language models, alt-text generation performance remains limited due to noisy user …
-
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
2017 · arXiv (Cornell University)
We develop a high-quality multi-turn dialog dataset, DailyDialog, which is intriguing in several aspects. The language is human-written and less noisy. The dialogues in the dataset reflect our daily communication way and cover various topics …