Daily Papers of 2026-01-16

  1. STEP3-VL-10B Technical Report 183 upvotes, #1 of 2026-01-16
  2. Urban Socio-Semantic Segmentation with Vision-Language Reasoning 154 upvotes, #2 of 2026-01-16
  3. Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs 140 upvotes, #3 of 2026-01-16
  4. Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning 82 upvotes, #4 of 2026-01-16
  5. VIBE: Visual Instruction Based Editor 61 upvotes, #5 of 2026-01-16
  6. Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning 44 upvotes, #6 of 2026-01-16
  7. DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset 36 upvotes, #7 of 2026-01-16
  8. Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering 36 upvotes, #7 of 2026-01-16
  9. HeartMuLa: A Family of Open Sourced Music Foundation Models 35 upvotes, #9 of 2026-01-16
  10. FlowAct-R1: Towards Interactive Humanoid Video Generation 33 upvotes, #10 of 2026-01-16
  11. Transition Matching Distillation for Fast Video Generation 31 upvotes, #11 of 2026-01-16
  12. CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation 30 upvotes, #12 of 2026-01-16
  13. Alterbute: Editing Intrinsic Attributes of Objects in Images 29 upvotes, #13 of 2026-01-16
  14. Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders 28 upvotes, #14 of 2026-01-16
  15. Action100M: A Large-scale Video Action Dataset 26 upvotes, #15 of 2026-01-16
  16. Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding 26 upvotes, #15 of 2026-01-16
  17. ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback 24 upvotes, #17 of 2026-01-16
  18. MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching 24 upvotes, #17 of 2026-01-16
  19. A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Doubao 1.8, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5 23 upvotes, #19 of 2026-01-16
  20. PACEvolve: Enabling Long-Horizon Progress-Aware Consistent Evolution 19 upvotes, #20 of 2026-01-16
  21. M^4olGen: Multi-Agent, Multi-Stage Molecular Generation under Precise Multi-Property Constraints 17 upvotes, #21 of 2026-01-16
  22. TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts 13 upvotes, #22 of 2026-01-16
  23. Inference-time Physics Alignment of Video Generative Models with Latent World Models 12 upvotes, #23 of 2026-01-16
  24. LSRIF: Logic-Structured Reinforcement Learning for Instruction Following 11 upvotes, #24 of 2026-01-16
  25. LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning 11 upvotes, #24 of 2026-01-16
  26. RigMo: Unifying Rig and Motion Learning for Generative Animation 9 upvotes, #26 of 2026-01-16
  27. EvasionBench: Detecting Evasive Answers in Financial Q&A via Multi-Model Consensus and LLM-as-Judge 9 upvotes, #26 of 2026-01-16
  28. V-DPM: 4D Video Reconstruction with Dynamic Point Maps 9 upvotes, #26 of 2026-01-16
  29. PRL: Process Reward Learning Improves LLMs' Reasoning Ability and Broadens the Reasoning Boundary 8 upvotes, #29 of 2026-01-16
  30. Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL 6 upvotes, #30 of 2026-01-16
  31. Deriving Character Logic from Storyline as Codified Decision Trees 6 upvotes, #30 of 2026-01-16
  32. Enhancing Sentiment Classification and Irony Detection in Large Language Models through Advanced Prompt Engineering Techniques 5 upvotes, #32 of 2026-01-16
  33. Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale 5 upvotes, #32 of 2026-01-16
  34. CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents 4 upvotes, #34 of 2026-01-16
  35. VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation 4 upvotes, #34 of 2026-01-16
  36. Demystifying the Slash Pattern in Attention: The Role of RoPE 3 upvotes, #36 of 2026-01-16
  37. WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments 3 upvotes, #36 of 2026-01-16
  38. Memory Bank Compression for Continual Adaptation of Large Language Models 2 upvotes, #38 of 2026-01-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.