Daily Papers of 2025-10-21

  1. DeepAnalyze: Agentic Large Language Models for Autonomous Data Science 90 upvotes, #1 of 2025-10-21
  2. Glyph: Scaling Context Windows via Visual-Text Compression 61 upvotes, #2 of 2025-10-21
  3. PICABench: How Far Are We from Physically Realistic Image Editing? 60 upvotes, #3 of 2025-10-21
  4. FineVision: Open Data Is All You Need 59 upvotes, #4 of 2025-10-21
  5. RL makes MLLMs see better than SFT 46 upvotes, #5 of 2025-10-21
  6. TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model 34 upvotes, #6 of 2025-10-21
  7. When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling 32 upvotes, #7 of 2025-10-21
  8. Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation 32 upvotes, #7 of 2025-10-21
  9. QueST: Incentivizing LLMs to Generate Difficult Problems 31 upvotes, #9 of 2025-10-21
  10. AION-1: Omnimodal Foundation Model for Astronomical Sciences 27 upvotes, #10 of 2025-10-21
  11. Annotation-Efficient Universal Honesty Alignment 20 upvotes, #11 of 2025-10-21
  12. Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling 19 upvotes, #12 of 2025-10-21
  13. Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback 18 upvotes, #13 of 2025-10-21
  14. Chronos-2: From Univariate to Universal Forecasting 14 upvotes, #14 of 2025-10-21
  15. Executable Knowledge Graphs for Replicating AI Research 12 upvotes, #15 of 2025-10-21
  16. ConsistEdit: Highly Consistent and Precise Training-free Visual Editing 12 upvotes, #15 of 2025-10-21
  17. Deep Self-Evolving Reasoning 10 upvotes, #17 of 2025-10-21
  18. Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics 7 upvotes, #18 of 2025-10-21
  19. Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset 6 upvotes, #19 of 2025-10-21
  20. Beyond Pipelines: A Survey of the Paradigm Shift toward Model-Native Agentic AI 6 upvotes, #19 of 2025-10-21
  21. Constantly Improving Image Models Need Constantly Improving Benchmarks 5 upvotes, #21 of 2025-10-21
  22. MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models 4 upvotes, #22 of 2025-10-21
  23. Agentic Reinforcement Learning for Search is Unsafe 4 upvotes, #22 of 2025-10-21
  24. UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action 4 upvotes, #22 of 2025-10-21
  25. Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering 3 upvotes, #25 of 2025-10-21
  26. Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense 3 upvotes, #25 of 2025-10-21
  27. What Limits Agentic Systems Efficiency? 3 upvotes, #25 of 2025-10-21
  28. Balanced Multi-Task Attention for Satellite Image Classification: A Systematic Approach to Achieving 97.23% Accuracy on EuroSAT Without Pre-Training 2 upvotes, #28 of 2025-10-21
  29. On Non-interactive Evaluation of Animal Communication Translators 2 upvotes, #28 of 2025-10-21
  30. GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer 2 upvotes, #28 of 2025-10-21
  31. Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection 2 upvotes, #28 of 2025-10-21
  32. Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains 2 upvotes, #28 of 2025-10-21
  33. AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning 1 upvotes, #33 of 2025-10-21
  34. Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models 1 upvotes, #33 of 2025-10-21
  35. Test-Time Scaling of Reasoning Models for Machine Translation 2 upvotes, #35 of 2025-10-21
  36. MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes 2 upvotes, #35 of 2025-10-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.