ChengQ
ChengQ on Hugging Face Daily Papers: 13 papers, 3 in the top 3 of their day, 736 upvotes.
- Code as Agent Harness 204 upvotes, #1 of 2026-05-19
- CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing 21 upvotes, #10 of 2026-05-07
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe 85 upvotes, #2 of 2026-04-15
- How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
- Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data 10 upvotes, #17 of 2026-03-03
- Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs 10 upvotes, #27 of 2026-02-11
- Agentic Reasoning for Large Language Models 182 upvotes, #1 of 2026-01-22
- From Word to World: Can Large Language Models be Implicit Text-based World Models? 11 upvotes, #12 of 2025-12-25
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
- LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering 2 upvotes, #24 of 2025-11-18
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents 20 upvotes, #4 of 2025-11-06
- Self-Improving LLM Agents at Test-Time 9 upvotes, #30 of 2025-10-14
- Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts 4 upvotes, #16 of 2025-09-15
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.