Siqi Sun
6 papers in the PaperMetrix corpus
Papers by this author
-
Graphical Model Sketch
2016 · arXiv (Cornell University)
Structured high-cardinality data arises in many domains, and poses a major challenge for both modeling and inference. Graphical models are a popular approach to modeling structured data but they are unsuitable for high-cardinality variables. The …
-
Identifying nonlinear relations among random variables: A network analytic approach
2024 · arXiv (Cornell University)
Nonlinear relations, such as the curvilinear relationship between childhood trauma and resilience in patients with schizophrenia and the moderation relationship between mentalizing, and internalizing and externalizing symptoms and quality of life in youths, are more …
-
Patient Knowledge Distillation for BERT Model Compression
2019
Siqi Sun, Yu Cheng, Zhe Gan, Jingjing Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation
2020
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, Bill Dolan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. 2020.
-
FreeLB: Enhanced Adversarial Training for Natural Language Understanding
2019 · arXiv (Cornell University)
Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that …
-
DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation
2019 · arXiv (Cornell University)
We present a large, tunable neural conversational response generation model, DialoGPT (dialogue generative pre-trained transformer). Trained on 147M conversation-like exchanges extracted from Reddit comment chains over a period spanning from 2005 through 2017, DialoGPT extends …