Daily Papers of 2025-09-26

  1. VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models 115 upvotes, #1 of 2025-09-26
  2. MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources 98 upvotes, #2 of 2025-09-26
  3. SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines 93 upvotes, #3 of 2025-09-26
  4. Tree Search for LLM Agent Reinforcement Learning 84 upvotes, #4 of 2025-09-26
  5. Seedream 4.0: Toward Next-generation Multimodal Image Generation 68 upvotes, #5 of 2025-09-26
  6. Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets 36 upvotes, #6 of 2025-09-26
  7. AutoIntent: AutoML for Text Classification 28 upvotes, #7 of 2025-09-26
  8. TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them 26 upvotes, #8 of 2025-09-26
  9. Thinking Augmented Pre-training 22 upvotes, #9 of 2025-09-26
  10. Residual Off-Policy RL for Finetuning Behavior Cloning Policies 17 upvotes, #10 of 2025-09-26
  11. CE-GPPO: Controlling Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning 16 upvotes, #11 of 2025-09-26
  12. Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution 15 upvotes, #12 of 2025-09-26
  13. CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive Modeling 15 upvotes, #12 of 2025-09-26
  14. Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory 13 upvotes, #14 of 2025-09-26
  15. Does FLUX Already Know How to Perform Physically Plausible Image Composition? 13 upvotes, #14 of 2025-09-26
  16. UserRL: Training Interactive User-Centric Agent via Reinforcement Learning 10 upvotes, #16 of 2025-09-26
  17. V-GameGym: Visual Game Generation for Code Large Language Models 9 upvotes, #17 of 2025-09-26
  18. ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning 9 upvotes, #17 of 2025-09-26
  19. SD3.5-Flash: Distribution-Guided Distillation of Generative Flows 9 upvotes, #17 of 2025-09-26
  20. SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent 8 upvotes, #20 of 2025-09-26
  21. Quantized Visual Geometry Grounded Transformer 8 upvotes, #20 of 2025-09-26
  22. OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps 6 upvotes, #22 of 2025-09-26
  23. When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity 6 upvotes, #22 of 2025-09-26
  24. Behind RoPE: How Does Causal Mask Encode Positional Information? 6 upvotes, #22 of 2025-09-26
  25. BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback 6 upvotes, #22 of 2025-09-26
  26. Interactive Recommendation Agent with Active User Commands 5 upvotes, #26 of 2025-09-26
  27. CompLLM: Compression for Long Context Q&A 4 upvotes, #27 of 2025-09-26
  28. MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning 4 upvotes, #27 of 2025-09-26
  29. Thinking While Listening: Simple Test Time Scaling For Audio Classification 3 upvotes, #29 of 2025-09-26
  30. Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving 3 upvotes, #29 of 2025-09-26
  31. StyleBench: Evaluating thinking styles in Large Language Models 3 upvotes, #29 of 2025-09-26
  32. Blueprints of Trust: AI System Cards for End to End Transparency and Governance 2 upvotes, #32 of 2025-09-26
  33. MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model 2 upvotes, #32 of 2025-09-26
  34. The Unanticipated Asymmetry Between Perceptual Optimization and Assessment 2 upvotes, #32 of 2025-09-26
  35. Evaluating Large Language Models for Detecting Antisemitism 1 upvotes, #35 of 2025-09-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.