ملف الباحث
Adam Khoja
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
2025 · arXiv (Cornell University)
As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced reasoning and knowledge …
-
Multi-Agent Inverse Q-Learning from Demonstrations
2025
When reward functions are hand-designed, deep reinforcement learning algorithms often suffer from reward misspecification, causing them to learn suboptimal policies in terms of the intended task objectives. In the single-agent case, inverse reinforcement learning (IRL) …