Researcher profile

Qingpeng Cai

2 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Two-Stage Constrained Actor-Critic for Short Video Recommendation

    2023

    The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users sequentially interact with the system and provide complex and multi-faceted responses, including …

  2. Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration

    2025 · arXiv (Cornell University)

    Reinforcement Learning (RL) has become a key approach for enhancing the reasoning capabilities of large language models. However, prevalent RL approaches like proximal policy optimization and group relative policy optimization suffer from sparse, outcome-based rewards …