Researcher profile
Huaxiu Yao
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
2024 · arXiv (Cornell University)
Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In reinforcement learning from human feedback …
-
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
2025 · arXiv (Cornell University)
Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across diverse medical specialties, limiting their performance. Recent efforts introduce multi-agent collaboration frameworks inspired by …