Daily Papers of 2026-01-13
- Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning 202 upvotes, #1 of 2026-01-13
- BabyVision: Visual Reasoning Beyond Language 184 upvotes, #2 of 2026-01-13
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning 77 upvotes, #3 of 2026-01-13
- MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head 46 upvotes, #4 of 2026-01-13
- X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests 40 upvotes, #5 of 2026-01-13
- Lost in the Noise: How Reasoning Models Fail with Contextual Distractors 29 upvotes, #6 of 2026-01-13
- GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts 27 upvotes, #7 of 2026-01-13
- OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent 25 upvotes, #8 of 2026-01-13
- Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models 24 upvotes, #9 of 2026-01-13
- Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction 21 upvotes, #10 of 2026-01-13
- MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era 19 upvotes, #11 of 2026-01-13
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving 18 upvotes, #12 of 2026-01-13
- Dr. Zero: Self-Evolving Search Agents without Training Data 17 upvotes, #13 of 2026-01-13
- Boosting Latent Diffusion Models via Disentangled Representation Alignment 16 upvotes, #14 of 2026-01-13
- What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models 15 upvotes, #15 of 2026-01-13
- ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration 15 upvotes, #15 of 2026-01-13
- TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning 9 upvotes, #17 of 2026-01-13
- Forest Before Trees: Latent Superposition for Efficient Visual Reasoning 9 upvotes, #17 of 2026-01-13
- RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction 7 upvotes, #19 of 2026-01-13
- OpenTinker: Separating Concerns in Agentic Reinforcement Learning 5 upvotes, #20 of 2026-01-13
- How Do Large Language Models Learn Concepts During Continual Pre-Training? 3 upvotes, #21 of 2026-01-13
- e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings 3 upvotes, #21 of 2026-01-13
- Sci-Reasoning: A Dataset Decoding AI Innovation Patterns 3 upvotes, #21 of 2026-01-13
- Structured Episodic Event Memory 3 upvotes, #21 of 2026-01-13
- Are LLM Decisions Faithful to Verbal Confidence? 3 upvotes, #21 of 2026-01-13
- Artificial Entanglement in the Fine-Tuning of Large Language Models 2 upvotes, #26 of 2026-01-13
- Codified Foreshadowing-Payoff Text Generation 2 upvotes, #26 of 2026-01-13
- ShowUI-Aloha: Human-Taught GUI Agent 2 upvotes, #26 of 2026-01-13
- "TODO: Fix the Mess Gemini Created": Towards Understanding GenAI-Induced Self-Admitted Technical Debt 2 upvotes, #26 of 2026-01-13
- FlyPose: Towards Robust Human Pose Estimation From Aerial Views 1 upvotes, #30 of 2026-01-13
- On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation 1 upvotes, #30 of 2026-01-13
- Does Inference Scaling Improve Reasoning Faithfulness? A Multi-Model Analysis of Self-Consistency Tradeoffs 1 upvotes, #30 of 2026-01-13
- Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths 1 upvotes, #30 of 2026-01-13
- SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models 1 upvotes, #30 of 2026-01-13
- Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? 1 upvotes, #30 of 2026-01-13
- Stochastic CHAOS: Why Deterministic Inference Kills, and Distributional Variability Is the Heartbeat of Artifical Cognition 1 upvotes, #30 of 2026-01-13
- On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training 1 upvotes, #30 of 2026-01-13
- Benchmarking Small Language Models and Small Reasoning Language Models on System Log Severity Classification 1 upvotes, #30 of 2026-01-13
- SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers 1 upvotes, #39 of 2026-01-13
- A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality 1 upvotes, #39 of 2026-01-13
- 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence 1 upvotes, #39 of 2026-01-13
- FinForge: Semi-Synthetic Financial Benchmark Generation 2 upvotes, #39 of 2026-01-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.