conference-paper
وصول مفتوح
SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets
Research footprint
At a glance
- الاستشهادات
- 116
- المراجع
- 37
- Comments
- 0
Paper overview
Abstract
Reinforcement learning methods for recommender systems optimize recommendations for long-term user engagement. However, since users are often presented with slates of multiple items---which may have interacting effects on user choice---methods are required to deal with the combinatorics of the RL action space. We develop SlateQ, a decomposition of value-based temporal-difference and Q-learning that renders RL tractable with slates. Under mild assumptions on user choice behavior, we show that the long-term value (LTV) of a slate can be decomposed into a tractable function of its component item-wise LTVs. We demonstrate our methods in simulation, and validate the scalability and effectiveness of decomposed TD-learning on YouTube.
Record transparency
Publication details
- DOI
- 10.24963/ijcai.2019/360
- OpenAlex
- W2948345531
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.