Daily Papers of 2025-10-29
- InteractComp: Evaluating Search Agents With Ambiguous Queries 96 upvotes, #1 of 2025-10-29
- Tongyi DeepResearch Technical Report 89 upvotes, #2 of 2025-10-29
- AgentFold: Long-Horizon Web Agents with Proactive Context Management 65 upvotes, #3 of 2025-10-29
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context 52 upvotes, #4 of 2025-10-29
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents 50 upvotes, #5 of 2025-10-29
- Uniform Discrete Diffusion with Metric Path for Video Generation 39 upvotes, #6 of 2025-10-29
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors 25 upvotes, #7 of 2025-10-29
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents 24 upvotes, #8 of 2025-10-29
- Batch Speculative Decoding Done Right 22 upvotes, #9 of 2025-10-29
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents 22 upvotes, #9 of 2025-10-29
- Group Relative Attention Guidance for Image Editing 22 upvotes, #9 of 2025-10-29
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision 22 upvotes, #9 of 2025-10-29
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis 22 upvotes, #9 of 2025-10-29
- VisCoder2: Building Multi-Language Visualization Coding Agents 20 upvotes, #14 of 2025-10-29
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs 20 upvotes, #14 of 2025-10-29
- WebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking 20 upvotes, #14 of 2025-10-29
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking 20 upvotes, #14 of 2025-10-29
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality 18 upvotes, #18 of 2025-10-29
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning 18 upvotes, #18 of 2025-10-29
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence 18 upvotes, #18 of 2025-10-29
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance 18 upvotes, #18 of 2025-10-29
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models 14 upvotes, #22 of 2025-10-29
- VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations 14 upvotes, #22 of 2025-10-29
- UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset 13 upvotes, #24 of 2025-10-29
- SPICE: Self-Play In Corpus Environments Improves Reasoning 12 upvotes, #25 of 2025-10-29
- Latent Chain-of-Thought for Visual Reasoning 9 upvotes, #26 of 2025-10-29
- Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures 9 upvotes, #26 of 2025-10-29
- MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion 6 upvotes, #28 of 2025-10-29
- Rethinking Visual Intelligence: Insights from Video Pretraining 5 upvotes, #29 of 2025-10-29
- FunReason-MT Technical Report: Overcoming the Complexity Barrier in Multi-Turn Function Calling 5 upvotes, #29 of 2025-10-29
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding 4 upvotes, #31 of 2025-10-29
- ATOM: AdapTive and OptiMized dynamic temporal knowledge graph construction using LLMs 4 upvotes, #31 of 2025-10-29
- SAO-Instruct: Free-form Audio Editing using Natural Language Instructions 4 upvotes, #31 of 2025-10-29
- ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? 4 upvotes, #31 of 2025-10-29
- Generalization or Memorization: Dynamic Decoding for Mode Steering 3 upvotes, #35 of 2025-10-29
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set 2 upvotes, #36 of 2025-10-29
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping 2 upvotes, #36 of 2025-10-29
- S-Chain: Structured Visual Chain-of-Thought For Medicine 2 upvotes, #36 of 2025-10-29
- Optimize Any Topology: A Foundation Model for Shape- and Resolution-Free Structural Topology Optimization 2 upvotes, #36 of 2025-10-29
- PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding 1 upvotes, #40 of 2025-10-29
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language 2 upvotes, #41 of 2025-10-29
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.