ملف الباحث

Huaxiu Yao

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling

    2024 · arXiv (Cornell University)

    Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In reinforcement learning from human feedback …

  2. MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

    2025 · arXiv (Cornell University)

    Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across diverse medical specialties, limiting their performance. Recent efforts introduce multi-agent collaboration frameworks inspired by …