Chengsong Huang
Chengsong Huang on Hugging Face Daily Papers: 23 papers, 7 in the top 3 of their day, 1,378 upvotes.
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses 216 upvotes, #2 of 2026-09-22
- EnvHarness: Awakening Static Worlds for Agent Learning 265 upvotes, #1 of 2026-08-21
- Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling 14 upvotes, #19 of 2026-06-03
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories 49 upvotes, #7 of 2026-05-21
- Process Rewards with Learned Reliability 52 upvotes, #7 of 2026-05-20
- G-Zero: Self-Play for Open-Ended Generation from Zero Data 16 upvotes, #16 of 2026-05-12
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling 64 upvotes, #5 of 2026-05-11
- Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration 36 upvotes, #9 of 2026-05-08
- MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data 51 upvotes, #3 of 2026-03-11
- Training Data Efficiency in Multimodal Process Reward Models 75 upvotes, #4 of 2026-02-05
- Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing 24 upvotes, #13 of 2026-02-04
- TTCS: Test-Time Curriculum Synthesis for Self-Evolving 33 upvotes, #7 of 2026-02-02
- RelayLLM: Efficient Reasoning via Collaborative Decoding 27 upvotes, #6 of 2026-01-09
- Benchmark^2: Systematic Evaluation of LLM Benchmarks 33 upvotes, #4 of 2026-01-08
- Guided Self-Evolving LLMs with Minimal Human Supervision 48 upvotes, #5 of 2025-12-03
- VisPlay: Self-Evolving Vision-Language Models from Images 41 upvotes, #4 of 2025-11-20
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning 95 upvotes, #2 of 2025-09-10
- Self-Rewarding Vision-Language Model via Reasoning Decomposition 77 upvotes, #2 of 2025-08-28
- R-Zero: Self-Evolving Reasoning LLM from Zero Data 107 upvotes, #2 of 2025-08-08
- POSS: Position Specialist Generates Better Draft for Speculative Decoding 6 upvotes, #25 of 2025-06-05
- CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation 8 upvotes, #15 of 2025-04-09
- Efficient Test-Time Scaling via Self-Calibration 13 upvotes, #10 of 2025-03-04
- LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition 34 upvotes, #1 of 2023-07-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.