Fei Mi
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
COLD: A Benchmark for Chinese Offensive Language Detection
2022 · arXiv (Cornell University)
Offensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models. However, this task in Chinese is still under exploration due to the scarcity of reliable datasets. To …
-
PanGu-Bot: Efficient Generative Dialogue Pre-training from Pre-trained Language Model
2022 · arXiv (Cornell University)
In this paper, we introduce PanGu-Bot, a Chinese pre-trained open-domain dialogue generation model based on a large pre-trained language model (PLM) PANGU-alpha (Zeng et al.,2021). Different from other pre-trained dialogue models trained over a massive …
-
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
2024 · arXiv (Cornell University)
Large Language Models (LLMs) has shown exceptional capabilities in many natual language understanding and generation tasks. However, the personalization issue still remains a much-coveted property, especially when it comes to the multiple sources involved in …
-
CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
2024 · arXiv (Cornell University)
As large language models (LLMs) constantly evolve, ensuring their safety remains a critical research problem. Previous red-teaming approaches for LLM safety have primarily focused on single prompt attacks or goal hijacking. To the best of …
-
SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation
2021 · arXiv (Cornell University)
Code representation learning, which aims to encode the semantics of source code into distributed vectors, plays an important role in recent deep-learning-based models for code intelligence. Recently, many pre-trained language models for source code (e.g., …