Daily Papers of 2025-10-10

  1. Agent Learning via Early Experience 223 upvotes, #1 of 2025-10-10
  2. MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization 103 upvotes, #2 of 2025-10-10
  3. DreamOmni2: Multimodal Instruction-based Editing and Generation 72 upvotes, #3 of 2025-10-10
  4. MemMamba: Rethinking Memory Patterns in State Space Model 67 upvotes, #4 of 2025-10-10
  5. UniVideo: Unified Understanding, Generation, and Editing for Videos 64 upvotes, #5 of 2025-10-10
  6. VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning 60 upvotes, #6 of 2025-10-10
  7. Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning 54 upvotes, #7 of 2025-10-10
  8. From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning 48 upvotes, #8 of 2025-10-10
  9. When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs 44 upvotes, #9 of 2025-10-10
  10. Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward 43 upvotes, #10 of 2025-10-10
  11. Training-Free Group Relative Policy Optimization 40 upvotes, #11 of 2025-10-10
  12. The Alignment Waltz: Jointly Training Agents to Collaborate for Safety 39 upvotes, #12 of 2025-10-10
  13. ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation 31 upvotes, #13 of 2025-10-10
  14. Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense 30 upvotes, #14 of 2025-10-10
  15. NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents 27 upvotes, #15 of 2025-10-10
  16. First Try Matters: Revisiting the Role of Reflection in Reasoning Models 24 upvotes, #16 of 2025-10-10
  17. DeepPrune: Parallel Scaling without Inter-trace Redundancy 23 upvotes, #17 of 2025-10-10
  18. Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks 22 upvotes, #18 of 2025-10-10
  19. LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions 22 upvotes, #18 of 2025-10-10
  20. PickStyle: Video-to-Video Style Transfer with Context-Style Adapters 20 upvotes, #20 of 2025-10-10
  21. UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution 20 upvotes, #20 of 2025-10-10
  22. NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints 19 upvotes, #22 of 2025-10-10
  23. CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards 18 upvotes, #23 of 2025-10-10
  24. InstructX: Towards Unified Visual Editing with MLLM Guidance 16 upvotes, #24 of 2025-10-10
  25. UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG 15 upvotes, #25 of 2025-10-10
  26. LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling 14 upvotes, #26 of 2025-10-10
  27. Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction 10 upvotes, #27 of 2025-10-10
  28. Reinforcing Diffusion Models by Direct Group Preference Optimization 10 upvotes, #27 of 2025-10-10
  29. Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window 9 upvotes, #29 of 2025-10-10
  30. UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections 8 upvotes, #30 of 2025-10-10
  31. Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency 8 upvotes, #30 of 2025-10-10
  32. SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models 8 upvotes, #30 of 2025-10-10
  33. OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment 7 upvotes, #33 of 2025-10-10
  34. Memory Retrieval and Consolidation in Large Language Models through Function Tokens 7 upvotes, #33 of 2025-10-10
  35. Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints 6 upvotes, #35 of 2025-10-10
  36. OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction 5 upvotes, #36 of 2025-10-10
  37. GCPO: When Contrast Fails, Go Gold 5 upvotes, #36 of 2025-10-10
  38. Recycling Pretrained Checkpoints: Orthogonal Growth of Mixture-of-Experts for Efficient Large Language Model Pre-Training 5 upvotes, #36 of 2025-10-10
  39. DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model 5 upvotes, #36 of 2025-10-10
  40. SViM3D: Stable Video Material Diffusion for Single Image 3D Generation 4 upvotes, #40 of 2025-10-10
  41. R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation 4 upvotes, #40 of 2025-10-10
  42. Towards Scalable and Consistent 3D Editing 3 upvotes, #42 of 2025-10-10
  43. Search-R3: Unifying Reasoning and Embedding Generation in Large Language Models 3 upvotes, #42 of 2025-10-10
  44. Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs 3 upvotes, #42 of 2025-10-10
  45. A^2Search: Ambiguity-Aware Question Answering with Reinforcement Learning 3 upvotes, #42 of 2025-10-10
  46. Beyond Outliers: A Study of Optimizers Under Quantization 2 upvotes, #46 of 2025-10-10
  47. Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models 2 upvotes, #46 of 2025-10-10
  48. GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations 2 upvotes, #46 of 2025-10-10
  49. Fidelity-Aware Data Composition for Robust Robot Generalization 1 upvotes, #49 of 2025-10-10
  50. Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning 1 upvotes, #49 of 2025-10-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.