Researcher profile
Ethan Mendes
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Language Models can Self-Improve at State-Value Estimation for Better Search
2025
Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, particularly in interactive domains such as web tasks. We introduce Self-Taught Lookahead (STL), a reward-free framework that improves language model-based value …