Shuai Ma
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Transition-based versus State-based Reward Functions for MDPs with Value-at-Risk
2016 · arXiv (Cornell University)
In reinforcement learning, the reward function on current state and action is widely used. When the objective is about the expectation of the (discounted) total reward only, it works perfectly. However, if the objective involves …
-
Dual-View Variational Autoencoders for Semi-Supervised Text Matching
2019
Semantically matching two text sequences (usually two sentences) is a fundamental problem in NLP. Most previous methods either encode each of the two sentences into a vector representation (sentence-level embedding) or leverage word-level interaction features …
-
Modeling Adaptive Expression of Robot Learning Engagement and Exploring Its Effects on Human Teachers
2022 · ACM Transactions on Computer-Human Interaction
Robot Learning from Demonstration (RLfD) allows non-expert users to teach a robot new skills or tasks directly through demonstrations. Although modeled after human–human learning and teaching, existing RLfD methods make robots act as passive observers …