ملف الباحث
Matthieu Zimmer
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning
2023 · arXiv (Cornell University)
A key method for creating Artificial Intelligence (AI) agents is Reinforcement Learning (RL). However, constructing a standalone RL policy that maps perception to action directly encounters severe problems, chief among them being its lack of …
-
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
2025 · arXiv (Cornell University)
We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one. Our approach models the generation of safe responses as a constrained Markov Decision …