Yu Cheng
12 papers in the PaperMetrix corpus
Papers by this author
-
Diverse Few-Shot Text Classification with Multiple Metrics
2018
Mo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, Bowen Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human …
-
Contextual Text Style Transfer
2020 · arXiv (Cornell University)
We introduce a new task, Contextual Text Style Transfer - translating a sentence into a desired style with its surrounding context taken into account. This brings two key challenges to existing style transfer approaches: ($i$) …
-
InfoBERT: Improving Robustness of Language Models from An Information\n Theoretic Perspective
2020 · arXiv (Cornell University)
Large-scale language models such as BERT have achieved state-of-the-art\nperformance across a wide range of NLP tasks. Recent studies, however, show\nthat such BERT-based models are vulnerable facing the threats of textual\nadversarial attacks. We aim to address …
-
Dynamic impairment classification through arrayed comparisons
2022 · Statistics in Medicine
The multivariate normative comparison (MNC) method has been used for identifying cognitive impairment. When participants' cognitive brain domains are evaluated regularly, the longitudinal MNC (LMNC) has been introduced to correct for the intercorrelation among repeated …
-
Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers
2023
Previous works reveal that similar to CNNs, vision transformers (ViT) are also vulnerable to universal adversarial patch attacks. In this paper, we empirically reveal and mathematically explain that the shallow tokens in the transformer and …
-
HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?
2025 · arXiv (Cornell University)
Recently, the physical capabilities of (M)LLMs have garnered increasing attention. However, existing benchmarks for physics suffer from two major gaps: they neither provide systematic and up-to-date coverage of real-world physics competitions such as physics Olympiads, …
-
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
2025 · arXiv (Cornell University)
Continual fine-tuning of large language models (LLMs) suffers from catastrophic forgetting. Rehearsal-based methods mitigate this problem by retaining a small set of old data. Nevertheless, they still suffer inevitable performance loss. Although training separate experts …
-
Adversarial Category Alignment Network for Cross-domain Sentiment Classification
2019
Xiaoye Qu, Zhikang Zou, Yu Cheng, Yang Yang, Pan Zhou. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). …
-
Patient Knowledge Distillation for BERT Model Compression
2019
Siqi Sun, Yu Cheng, Zhe Gan, Jingjing Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Domain Adaptive Text Style Transfer
2019
Dianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …
-
FreeLB: Enhanced Adversarial Training for Natural Language Understanding
2019 · arXiv (Cornell University)
Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that …
-
Discourse-Aware Neural Extractive Text Summarization
2020
Recently BERT has been adopted for document encoding in state-of-the-art text summarization models. However, sentence-based extractive models often result in redundant or uninformative phrases in the extracted summaries. Also, long-range dependencies throughout a document are …