Researcher profile

Yu Cheng

12 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Diverse Few-Shot Text Classification with Multiple Metrics

    2018

    Mo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, Bowen Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human …

  2. Contextual Text Style Transfer

    2020 · arXiv (Cornell University)

    We introduce a new task, Contextual Text Style Transfer - translating a sentence into a desired style with its surrounding context taken into account. This brings two key challenges to existing style transfer approaches: ($i$) …

  3. InfoBERT: Improving Robustness of Language Models from An Information\n Theoretic Perspective

    2020 · arXiv (Cornell University)

    Large-scale language models such as BERT have achieved state-of-the-art\nperformance across a wide range of NLP tasks. Recent studies, however, show\nthat such BERT-based models are vulnerable facing the threats of textual\nadversarial attacks. We aim to address …

  4. Dynamic impairment classification through arrayed comparisons

    2022 · Statistics in Medicine

    The multivariate normative comparison (MNC) method has been used for identifying cognitive impairment. When participants' cognitive brain domains are evaluated regularly, the longitudinal MNC (LMNC) has been introduced to correct for the intercorrelation among repeated …

  5. Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers

    2023

    Previous works reveal that similar to CNNs, vision transformers (ViT) are also vulnerable to universal adversarial patch attacks. In this paper, we empirically reveal and mathematically explain that the shallow tokens in the transformer and …

  6. HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?

    2025 · arXiv (Cornell University)

    Recently, the physical capabilities of (M)LLMs have garnered increasing attention. However, existing benchmarks for physics suffer from two major gaps: they neither provide systematic and up-to-date coverage of real-world physics competitions such as physics Olympiads, …

  7. SEE: Continual Fine-tuning with Sequential Ensemble of Experts

    2025 · arXiv (Cornell University)

    Continual fine-tuning of large language models (LLMs) suffers from catastrophic forgetting. Rehearsal-based methods mitigate this problem by retaining a small set of old data. Nevertheless, they still suffer inevitable performance loss. Although training separate experts …

  8. Adversarial Category Alignment Network for Cross-domain Sentiment Classification

    2019

    Xiaoye Qu, Zhikang Zou, Yu Cheng, Yang Yang, Pan Zhou. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). …

  9. Patient Knowledge Distillation for BERT Model Compression

    2019

    Siqi Sun, Yu Cheng, Zhe Gan, Jingjing Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  10. Domain Adaptive Text Style Transfer

    2019

    Dianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …

  11. FreeLB: Enhanced Adversarial Training for Natural Language Understanding

    2019 · arXiv (Cornell University)

    Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that …

  12. Discourse-Aware Neural Extractive Text Summarization

    2020

    Recently BERT has been adopted for document encoding in state-of-the-art text summarization models. However, sentence-based extractive models often result in redundant or uninformative phrases in the extracted summaries. Also, long-range dependencies throughout a document are …