Daily Papers of 2025-10-29

  1. InteractComp: Evaluating Search Agents With Ambiguous Queries 96 upvotes, #1 of 2025-10-29
  2. Tongyi DeepResearch Technical Report 89 upvotes, #2 of 2025-10-29
  3. AgentFold: Long-Horizon Web Agents with Proactive Context Management 65 upvotes, #3 of 2025-10-29
  4. RoboOmni: Proactive Robot Manipulation in Omni-modal Context 52 upvotes, #4 of 2025-10-29
  5. Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents 50 upvotes, #5 of 2025-10-29
  6. Uniform Discrete Diffusion with Metric Path for Video Generation 39 upvotes, #6 of 2025-10-29
  7. From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors 25 upvotes, #7 of 2025-10-29
  8. Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents 24 upvotes, #8 of 2025-10-29
  9. Batch Speculative Decoding Done Right 22 upvotes, #9 of 2025-10-29
  10. OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents 22 upvotes, #9 of 2025-10-29
  11. Group Relative Attention Guidance for Image Editing 22 upvotes, #9 of 2025-10-29
  12. Repurposing Synthetic Data for Fine-grained Search Agent Supervision 22 upvotes, #9 of 2025-10-29
  13. AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis 22 upvotes, #9 of 2025-10-29
  14. VisCoder2: Building Multi-Language Visualization Coding Agents 20 upvotes, #14 of 2025-10-29
  15. Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs 20 upvotes, #14 of 2025-10-29
  16. WebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking 20 upvotes, #14 of 2025-10-29
  17. ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking 20 upvotes, #14 of 2025-10-29
  18. ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality 18 upvotes, #18 of 2025-10-29
  19. Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning 18 upvotes, #18 of 2025-10-29
  20. STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence 18 upvotes, #18 of 2025-10-29
  21. Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance 18 upvotes, #18 of 2025-10-29
  22. Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models 14 upvotes, #22 of 2025-10-29
  23. VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations 14 upvotes, #22 of 2025-10-29
  24. UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset 13 upvotes, #24 of 2025-10-29
  25. SPICE: Self-Play In Corpus Environments Improves Reasoning 12 upvotes, #25 of 2025-10-29
  26. Latent Chain-of-Thought for Visual Reasoning 9 upvotes, #26 of 2025-10-29
  27. Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures 9 upvotes, #26 of 2025-10-29
  28. MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion 6 upvotes, #28 of 2025-10-29
  29. Rethinking Visual Intelligence: Insights from Video Pretraining 5 upvotes, #29 of 2025-10-29
  30. FunReason-MT Technical Report: Overcoming the Complexity Barrier in Multi-Turn Function Calling 5 upvotes, #29 of 2025-10-29
  31. PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding 4 upvotes, #31 of 2025-10-29
  32. ATOM: AdapTive and OptiMized dynamic temporal knowledge graph construction using LLMs 4 upvotes, #31 of 2025-10-29
  33. SAO-Instruct: Free-form Audio Editing using Natural Language Instructions 4 upvotes, #31 of 2025-10-29
  34. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? 4 upvotes, #31 of 2025-10-29
  35. Generalization or Memorization: Dynamic Decoding for Mode Steering 3 upvotes, #35 of 2025-10-29
  36. VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set 2 upvotes, #36 of 2025-10-29
  37. GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping 2 upvotes, #36 of 2025-10-29
  38. S-Chain: Structured Visual Chain-of-Thought For Medicine 2 upvotes, #36 of 2025-10-29
  39. Optimize Any Topology: A Foundation Model for Shape- and Resolution-Free Structural Topology Optimization 2 upvotes, #36 of 2025-10-29
  40. PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding 1 upvotes, #40 of 2025-10-29
  41. Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language 2 upvotes, #41 of 2025-10-29

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.