Chengsong Huang

Chengsong Huang on Hugging Face Daily Papers: 23 papers, 7 in the top 3 of their day, 1,378 upvotes.

  1. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses 216 upvotes, #2 of 2026-09-22
  2. EnvHarness: Awakening Static Worlds for Agent Learning 265 upvotes, #1 of 2026-08-21
  3. Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling 14 upvotes, #19 of 2026-06-03
  4. You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories 49 upvotes, #7 of 2026-05-21
  5. Process Rewards with Learned Reliability 52 upvotes, #7 of 2026-05-20
  6. G-Zero: Self-Play for Open-Ended Generation from Zero Data 16 upvotes, #16 of 2026-05-12
  7. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling 64 upvotes, #5 of 2026-05-11
  8. Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration 36 upvotes, #9 of 2026-05-08
  9. MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data 51 upvotes, #3 of 2026-03-11
  10. Training Data Efficiency in Multimodal Process Reward Models 75 upvotes, #4 of 2026-02-05
  11. Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing 24 upvotes, #13 of 2026-02-04
  12. TTCS: Test-Time Curriculum Synthesis for Self-Evolving 33 upvotes, #7 of 2026-02-02
  13. RelayLLM: Efficient Reasoning via Collaborative Decoding 27 upvotes, #6 of 2026-01-09
  14. Benchmark^2: Systematic Evaluation of LLM Benchmarks 33 upvotes, #4 of 2026-01-08
  15. Guided Self-Evolving LLMs with Minimal Human Supervision 48 upvotes, #5 of 2025-12-03
  16. VisPlay: Self-Evolving Vision-Language Models from Images 41 upvotes, #4 of 2025-11-20
  17. Parallel-R1: Towards Parallel Thinking via Reinforcement Learning 95 upvotes, #2 of 2025-09-10
  18. Self-Rewarding Vision-Language Model via Reasoning Decomposition 77 upvotes, #2 of 2025-08-28
  19. R-Zero: Self-Evolving Reasoning LLM from Zero Data 107 upvotes, #2 of 2025-08-08
  20. POSS: Position Specialist Generates Better Draft for Speculative Decoding 6 upvotes, #25 of 2025-06-05
  21. CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation 8 upvotes, #15 of 2025-04-09
  22. Efficient Test-Time Scaling via Self-Calibration 13 upvotes, #10 of 2025-03-04
  23. LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition 34 upvotes, #1 of 2023-07-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.