Jie Fu
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives
2019 · arXiv (Cornell University)
This paper tackles the problem of reading comprehension over long narratives where documents easily span over thousands of tokens. We propose a curriculum learning (CL) based Pointer-Generator framework for reading/sampling over large documents, enabling diverse …
-
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
2023 · arXiv (Cornell University)
The prompt-based learning paradigm, which bridges the gap between pre-training and fine-tuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings. Despite being widely applied, prompt-based learning is vulnerable to backdoor attacks. Textual …
-
Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility
2023 · arXiv (Cornell University)
The recent popularity of large language models (LLMs) has brought a significant impact to boundless fields, particularly through their open-ended ecosystem such as the APIs, open-sourced models, and plugins. However, with their widespread deployment, there …
-
Information-Theoretic Opacity-Enforcement in Markov Decision Processes
2024 · arXiv (Cornell University)
The paper studies information-theoretic opacity, an information-flow privacy property, in a setting involving two agents: A planning agent who controls a stochastic system and an observer who partially observes the system states. The goal of …
-
Differentially Private Federated Learning: A Systematic Review
2024 · arXiv (Cornell University)
In recent years, privacy and security concerns in machine learning have promoted trusted federated learning to the forefront of research. Differential privacy has emerged as the de facto standard for privacy protection in federated learning …
-
MIO: A Foundation Model on Multimodal Tokens
2024 · arXiv (Cornell University)
In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language …
-
Thinker: Learning to Think Fast and Slow
2025
Recent studies show that the reasoning capabilities of Large Language Models (LLMs) can be improved by applying Reinforcement Learning (RL) to question-answering (QA) tasks in areas such as math and coding. With a long context …
-
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
2025 · arXiv (Cornell University)
Grokking is proposed and widely studied as an intricate phenomenon in which generalization is achieved after a long-lasting period of overfitting. In this work, we propose NeuralGrok, a novel gradient-based approach that learns an optimal …