Daily Papers of 2025-10-10
- Agent Learning via Early Experience 223 upvotes, #1 of 2025-10-10
- MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization 103 upvotes, #2 of 2025-10-10
- DreamOmni2: Multimodal Instruction-based Editing and Generation 72 upvotes, #3 of 2025-10-10
- MemMamba: Rethinking Memory Patterns in State Space Model 67 upvotes, #4 of 2025-10-10
- UniVideo: Unified Understanding, Generation, and Editing for Videos 64 upvotes, #5 of 2025-10-10
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning 60 upvotes, #6 of 2025-10-10
- Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning 54 upvotes, #7 of 2025-10-10
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning 48 upvotes, #8 of 2025-10-10
- When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs 44 upvotes, #9 of 2025-10-10
- Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward 43 upvotes, #10 of 2025-10-10
- Training-Free Group Relative Policy Optimization 40 upvotes, #11 of 2025-10-10
- The Alignment Waltz: Jointly Training Agents to Collaborate for Safety 39 upvotes, #12 of 2025-10-10
- ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation 31 upvotes, #13 of 2025-10-10
- Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense 30 upvotes, #14 of 2025-10-10
- NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents 27 upvotes, #15 of 2025-10-10
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models 24 upvotes, #16 of 2025-10-10
- DeepPrune: Parallel Scaling without Inter-trace Redundancy 23 upvotes, #17 of 2025-10-10
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks 22 upvotes, #18 of 2025-10-10
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions 22 upvotes, #18 of 2025-10-10
- PickStyle: Video-to-Video Style Transfer with Context-Style Adapters 20 upvotes, #20 of 2025-10-10
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution 20 upvotes, #20 of 2025-10-10
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints 19 upvotes, #22 of 2025-10-10
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards 18 upvotes, #23 of 2025-10-10
- InstructX: Towards Unified Visual Editing with MLLM Guidance 16 upvotes, #24 of 2025-10-10
- UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG 15 upvotes, #25 of 2025-10-10
- LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling 14 upvotes, #26 of 2025-10-10
- Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction 10 upvotes, #27 of 2025-10-10
- Reinforcing Diffusion Models by Direct Group Preference Optimization 10 upvotes, #27 of 2025-10-10
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window 9 upvotes, #29 of 2025-10-10
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections 8 upvotes, #30 of 2025-10-10
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency 8 upvotes, #30 of 2025-10-10
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models 8 upvotes, #30 of 2025-10-10
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment 7 upvotes, #33 of 2025-10-10
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens 7 upvotes, #33 of 2025-10-10
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints 6 upvotes, #35 of 2025-10-10
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction 5 upvotes, #36 of 2025-10-10
- GCPO: When Contrast Fails, Go Gold 5 upvotes, #36 of 2025-10-10
- Recycling Pretrained Checkpoints: Orthogonal Growth of Mixture-of-Experts for Efficient Large Language Model Pre-Training 5 upvotes, #36 of 2025-10-10
- DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model 5 upvotes, #36 of 2025-10-10
- SViM3D: Stable Video Material Diffusion for Single Image 3D Generation 4 upvotes, #40 of 2025-10-10
- R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation 4 upvotes, #40 of 2025-10-10
- Towards Scalable and Consistent 3D Editing 3 upvotes, #42 of 2025-10-10
- Search-R3: Unifying Reasoning and Embedding Generation in Large Language Models 3 upvotes, #42 of 2025-10-10
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs 3 upvotes, #42 of 2025-10-10
- A^2Search: Ambiguity-Aware Question Answering with Reinforcement Learning 3 upvotes, #42 of 2025-10-10
- Beyond Outliers: A Study of Optimizers Under Quantization 2 upvotes, #46 of 2025-10-10
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models 2 upvotes, #46 of 2025-10-10
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations 2 upvotes, #46 of 2025-10-10
- Fidelity-Aware Data Composition for Robust Robot Generalization 1 upvotes, #49 of 2025-10-10
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning 1 upvotes, #49 of 2025-10-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.