Daily Papers of 2026-01-13

  1. Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning 202 upvotes, #1 of 2026-01-13
  2. BabyVision: Visual Reasoning Beyond Language 184 upvotes, #2 of 2026-01-13
  3. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning 77 upvotes, #3 of 2026-01-13
  4. MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head 46 upvotes, #4 of 2026-01-13
  5. X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests 40 upvotes, #5 of 2026-01-13
  6. Lost in the Noise: How Reasoning Models Fail with Contextual Distractors 29 upvotes, #6 of 2026-01-13
  7. GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts 27 upvotes, #7 of 2026-01-13
  8. OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent 25 upvotes, #8 of 2026-01-13
  9. Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models 24 upvotes, #9 of 2026-01-13
  10. Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction 21 upvotes, #10 of 2026-01-13
  11. MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era 19 upvotes, #11 of 2026-01-13
  12. DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving 18 upvotes, #12 of 2026-01-13
  13. Dr. Zero: Self-Evolving Search Agents without Training Data 17 upvotes, #13 of 2026-01-13
  14. Boosting Latent Diffusion Models via Disentangled Representation Alignment 16 upvotes, #14 of 2026-01-13
  15. What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models 15 upvotes, #15 of 2026-01-13
  16. ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration 15 upvotes, #15 of 2026-01-13
  17. TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning 9 upvotes, #17 of 2026-01-13
  18. Forest Before Trees: Latent Superposition for Efficient Visual Reasoning 9 upvotes, #17 of 2026-01-13
  19. RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction 7 upvotes, #19 of 2026-01-13
  20. OpenTinker: Separating Concerns in Agentic Reinforcement Learning 5 upvotes, #20 of 2026-01-13
  21. How Do Large Language Models Learn Concepts During Continual Pre-Training? 3 upvotes, #21 of 2026-01-13
  22. e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings 3 upvotes, #21 of 2026-01-13
  23. Sci-Reasoning: A Dataset Decoding AI Innovation Patterns 3 upvotes, #21 of 2026-01-13
  24. Structured Episodic Event Memory 3 upvotes, #21 of 2026-01-13
  25. Are LLM Decisions Faithful to Verbal Confidence? 3 upvotes, #21 of 2026-01-13
  26. Artificial Entanglement in the Fine-Tuning of Large Language Models 2 upvotes, #26 of 2026-01-13
  27. Codified Foreshadowing-Payoff Text Generation 2 upvotes, #26 of 2026-01-13
  28. ShowUI-Aloha: Human-Taught GUI Agent 2 upvotes, #26 of 2026-01-13
  29. "TODO: Fix the Mess Gemini Created": Towards Understanding GenAI-Induced Self-Admitted Technical Debt 2 upvotes, #26 of 2026-01-13
  30. FlyPose: Towards Robust Human Pose Estimation From Aerial Views 1 upvotes, #30 of 2026-01-13
  31. On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation 1 upvotes, #30 of 2026-01-13
  32. Does Inference Scaling Improve Reasoning Faithfulness? A Multi-Model Analysis of Self-Consistency Tradeoffs 1 upvotes, #30 of 2026-01-13
  33. Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths 1 upvotes, #30 of 2026-01-13
  34. SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models 1 upvotes, #30 of 2026-01-13
  35. Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? 1 upvotes, #30 of 2026-01-13
  36. Stochastic CHAOS: Why Deterministic Inference Kills, and Distributional Variability Is the Heartbeat of Artifical Cognition 1 upvotes, #30 of 2026-01-13
  37. On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training 1 upvotes, #30 of 2026-01-13
  38. Benchmarking Small Language Models and Small Reasoning Language Models on System Log Severity Classification 1 upvotes, #30 of 2026-01-13
  39. SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers 1 upvotes, #39 of 2026-01-13
  40. A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality 1 upvotes, #39 of 2026-01-13
  41. 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence 1 upvotes, #39 of 2026-01-13
  42. FinForge: Semi-Synthetic Financial Benchmark Generation 2 upvotes, #39 of 2026-01-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.