Daily Papers of 2026-01-23

  1. EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience 89 upvotes, #1 of 2026-01-23
  2. LLM-in-Sandbox Elicits General Agentic Intelligence 82 upvotes, #2 of 2026-01-23
  3. HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding 73 upvotes, #3 of 2026-01-23
  4. The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models 68 upvotes, #4 of 2026-01-23
  5. BayesianVLA: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries 54 upvotes, #5 of 2026-01-23
  6. Qwen3-TTS Technical Report 54 upvotes, #5 of 2026-01-23
  7. Stable-DiffCoder: Pushing the Frontier of Code Diffusion Large Language Model 53 upvotes, #7 of 2026-01-23
  8. Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders 51 upvotes, #8 of 2026-01-23
  9. SAMTok: Representing Any Mask with Two Words 41 upvotes, #9 of 2026-01-23
  10. Learning to Discover at Test Time 40 upvotes, #10 of 2026-01-23
  11. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces 32 upvotes, #11 of 2026-01-23
  12. OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation 18 upvotes, #12 of 2026-01-23
  13. Towards Automated Kernel Generation in the Era of LLMs 16 upvotes, #13 of 2026-01-23
  14. VideoMaMa: Mask-Guided Video Matting via Generative Prior 13 upvotes, #14 of 2026-01-23
  15. Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing 13 upvotes, #14 of 2026-01-23
  16. Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning 13 upvotes, #14 of 2026-01-23
  17. PROGRESSLM: Towards Progress Reasoning in Vision-Language Models 12 upvotes, #17 of 2026-01-23
  18. ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion 12 upvotes, #17 of 2026-01-23
  19. Agentic Uncertainty Quantification 8 upvotes, #19 of 2026-01-23
  20. 360Anything: Geometry-Free Lifting of Images and Videos to 360° 8 upvotes, #19 of 2026-01-23
  21. Agentic Confidence Calibration 5 upvotes, #21 of 2026-01-23
  22. VIOLA: Towards Video In-Context Learning with Minimal Annotations 4 upvotes, #22 of 2026-01-23
  23. From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models 4 upvotes, #22 of 2026-01-23
  24. MirrorBench: An Extensible Framework to Evaluate User-Proxy Agents for Human-Likeness 2 upvotes, #24 of 2026-01-23
  25. Numba-Accelerated 2D Diffusion-Limited Aggregation: Implementation and Fractal Characterization 1 upvotes, #25 of 2026-01-23
  26. Wigner's Friend as a Circuit: Inter-Branch Communication Witness Benchmarks on Superconducting Quantum Hardware 1 upvotes, #25 of 2026-01-23
  27. LLM Prompt Evaluation for Educational Applications 1 upvotes, #25 of 2026-01-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.