Daily Papers of 2025-10-13

  1. D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI 128 upvotes, #1 of 2025-10-13
  2. Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation 114 upvotes, #2 of 2025-10-13
  3. KORMo: Korean Open Reasoning Model for Everyone 68 upvotes, #3 of 2025-10-13
  4. AutoPR: Let's Automate Your Academic Promotion! 48 upvotes, #4 of 2025-10-13
  5. TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffusion Sampling 46 upvotes, #5 of 2025-10-13
  6. Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs 46 upvotes, #5 of 2025-10-13
  7. StreamingVLM: Real-Time Understanding for Infinite Video Streams 45 upvotes, #7 of 2025-10-13
  8. BEAR: Benchmarking and Enhancing Multimodal Language Models for Atomic Embodied Capabilities 44 upvotes, #8 of 2025-10-13
  9. Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels 31 upvotes, #9 of 2025-10-13
  10. BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution 28 upvotes, #10 of 2025-10-13
  11. R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 25 upvotes, #11 of 2025-10-13
  12. Which Heads Matter for Reasoning? RL-Guided KV Cache Compression 21 upvotes, #12 of 2025-10-13
  13. SpaceVista: All-Scale Visual Spatial Reasoning from mm to km 17 upvotes, #13 of 2025-10-13
  14. DISCO: Diversifying Sample Condensation for Efficient Model Evaluation 14 upvotes, #14 of 2025-10-13
  15. Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting 13 upvotes, #15 of 2025-10-13
  16. ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping 12 upvotes, #16 of 2025-10-13
  17. Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy Prediction 9 upvotes, #17 of 2025-10-13
  18. Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization 9 upvotes, #17 of 2025-10-13
  19. PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs 9 upvotes, #17 of 2025-10-13
  20. LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? 7 upvotes, #20 of 2025-10-13
  21. MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval 7 upvotes, #20 of 2025-10-13
  22. TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control 7 upvotes, #20 of 2025-10-13
  23. Instant4D: 4D Gaussian Splatting in Minutes 6 upvotes, #23 of 2025-10-13
  24. Understanding DeepResearch via Reports 6 upvotes, #23 of 2025-10-13
  25. Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition 6 upvotes, #23 of 2025-10-13
  26. Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models 6 upvotes, #23 of 2025-10-13
  27. StatEval: A Comprehensive Benchmark for Large Language Models in Statistics 6 upvotes, #23 of 2025-10-13
  28. Dyna-Mind: Learning to Simulate from Experience for Better AI Agents 6 upvotes, #23 of 2025-10-13
  29. Parallel Test-Time Scaling for Latent Reasoning Models 5 upvotes, #29 of 2025-10-13
  30. Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols 5 upvotes, #29 of 2025-10-13
  31. One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework 4 upvotes, #31 of 2025-10-13
  32. ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review 4 upvotes, #31 of 2025-10-13
  33. Mitigating Overthinking through Reasoning Shaping 4 upvotes, #31 of 2025-10-13
  34. Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models 4 upvotes, #31 of 2025-10-13
  35. A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks 3 upvotes, #35 of 2025-10-13
  36. Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation 3 upvotes, #35 of 2025-10-13
  37. ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization 2 upvotes, #37 of 2025-10-13
  38. ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL 2 upvotes, #37 of 2025-10-13
  39. Temporal Prompting Matters: Rethinking Referring Video Object Segmentation 2 upvotes, #37 of 2025-10-13
  40. LLM4Cell: A Survey of Large Language and Agentic Models for Single-Cell Biology 2 upvotes, #37 of 2025-10-13
  41. How to Teach Large Multimodal Models New Skills 2 upvotes, #37 of 2025-10-13
  42. GTAlign: Game-Theoretic Alignment of LLM Assistants for Mutual Welfare 2 upvotes, #37 of 2025-10-13
  43. MONKEY: Masking ON KEY-Value Activation Adapter for Personalization 1 upvotes, #43 of 2025-10-13
  44. ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall 1 upvotes, #43 of 2025-10-13
  45. Formalizing Style in Personal Narratives 1 upvotes, #43 of 2025-10-13
  46. Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation 1 upvotes, #43 of 2025-10-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.