Shenao Zhang
Shenao Zhang on Hugging Face Daily Papers: 4 papers, 1 in the top 3 of their day, 73 upvotes.
- Learning to Reason as Action Abstractions with Scalable Mid-Training RL 5 upvotes, #35 of 2025-10-01
- Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning 6 upvotes, #40 of 2025-05-28
- Offline Reinforcement Learning for LLM Multi-Step Reasoning 33 upvotes, #2 of 2024-12-23
- Self-Exploring Language Models: Active Preference Elicitation for Online Alignment 12 upvotes, #4 of 2024-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.