haoyu wang
haoyu wang on Hugging Face Daily Papers: 3 papers, 0 in the top 3 of their day, 38 upvotes.
- RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains 13 upvotes, #29 of 2026-05-29
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training 14 upvotes, #31 of 2026-02-03
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment 7 upvotes, #33 of 2025-10-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.