Cheng Qian
Cheng Qian on Hugging Face Daily Papers: 14 papers, 5 in the top 3 of their day, 665 upvotes.
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses 111 upvotes, #3 of 2026-08-13
- Trimming the Long-Tail of Visual World Modeling Evaluation 42 upvotes, #7 of 2026-06-30
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems 95 upvotes, #1 of 2026-06-23
- AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints 40 upvotes, #4 of 2026-06-05
- Advancing Creative Physical Intelligence in Large Multimodal Models 19 upvotes, #21 of 2026-05-28
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning 10 upvotes, #16 of 2025-09-26
- LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering 6 upvotes, #16 of 2025-09-12
- UserBench: An Interactive Gym Environment for User-Centric Agents 29 upvotes, #10 of 2025-08-12
- A Survey of Self-Evolving Agents: On Path to Artificial Super Intelligence 72 upvotes, #2 of 2025-07-29
- RM-R1: Reward Modeling as Reasoning 66 upvotes, #3 of 2025-05-06
- OTC: Optimal Tool Calls via Reinforcement Learning 33 upvotes, #5 of 2025-04-22
- ToolRL: Reward is All Tool Learning Needs 41 upvotes, #4 of 2025-04-22
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents 23 upvotes, #2 of 2025-03-05
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents 32 upvotes, #4 of 2025-02-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.