Chinese Academic of Science Institute of Automation

Chinese Academic of Science Institute of Automation on Hugging Face Daily Papers: 20 papers, 2 in the top 3 of their day, 0 paper of the day.

  1. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? 20 upvotes, #11 of 2026-09-10
  2. SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs 53 upvotes, #2 of 2026-08-10
  3. Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning 42 upvotes, #4 of 2026-08-10
  4. Continual Learning in Transition 23 upvotes, #19 of 2026-08-07
  5. SWE-Touch: Benchmarking Coding Agents When Users Touch the Code 24 upvotes, #14 of 2026-08-04
  6. PhiZero: A World Model Built Around Physical Language 167 upvotes, #5 of 2026-07-31
  7. The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation 26 upvotes, #10 of 2026-07-28
  8. Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It 18 upvotes, #10 of 2026-06-26
  9. Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do 9 upvotes, #15 of 2026-06-25
  10. Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution 7 upvotes, #36 of 2026-05-26
  11. Uncovering Entity Identity Confusion in Multimodal Knowledge Editing 1 upvotes, #51 of 2026-05-12
  12. ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning 7 upvotes, #14 of 2026-05-07
  13. From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space 28 upvotes, #6 of 2026-04-16
  14. PLUME: Latent Reasoning Based Universal Multimodal Embedding 15 upvotes, #19 of 2026-04-07
  15. World Models for Policy Refinement in StarCraft II 6 upvotes, #14 of 2026-02-20
  16. Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies 60 upvotes, #3 of 2025-12-24
  17. VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression? 6 upvotes, #22 of 2025-12-18
  18. IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting 3 upvotes, #16 of 2025-12-11
  19. Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences 26 upvotes, #12 of 2025-10-28
  20. Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective 3 upvotes, #41 of 2025-10-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.