Daily Papers of 2026-02-03

  1. Green-VLA: Staged Vision-Language-Action Model for Generalist Robots 266 upvotes, #1 of 2026-02-03
  2. Kimi K2.5: Visual Agentic Intelligence 219 upvotes, #2 of 2026-02-03
  3. Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models 149 upvotes, #3 of 2026-02-03
  4. Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models 124 upvotes, #4 of 2026-02-03
  5. Closing the Loop: Universal Repository Representation with RPG-Encoder 82 upvotes, #5 of 2026-02-03
  6. UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing 75 upvotes, #6 of 2026-02-03
  7. SWE-Universe: Scale Real-World Verifiable Environments to Millions 59 upvotes, #7 of 2026-02-03
  8. FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents 45 upvotes, #8 of 2026-02-03
  9. SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning 44 upvotes, #9 of 2026-02-03
  10. PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss 41 upvotes, #10 of 2026-02-03
  11. WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora 40 upvotes, #11 of 2026-02-03
  12. Rethinking Generative Recommender Tokenizer: Recsys-Native Encoding and Semantic Quantization Beyond LLMs 40 upvotes, #11 of 2026-02-03
  13. Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning 39 upvotes, #13 of 2026-02-03
  14. Generative Visual Code Mobile World Models 39 upvotes, #13 of 2026-02-03
  15. Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling 35 upvotes, #15 of 2026-02-03
  16. Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles 32 upvotes, #16 of 2026-02-03
  17. RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System 31 upvotes, #17 of 2026-02-03
  18. SLIME: Stabilized Likelihood Implicit Margin Enforcement for Preference Optimization 29 upvotes, #18 of 2026-02-03
  19. Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention 27 upvotes, #19 of 2026-02-03
  20. Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation 24 upvotes, #20 of 2026-02-03
  21. PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards 23 upvotes, #21 of 2026-02-03
  22. Rethinking Selective Knowledge Distillation 22 upvotes, #22 of 2026-02-03
  23. Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation 22 upvotes, #22 of 2026-02-03
  24. FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space 18 upvotes, #24 of 2026-02-03
  25. Ebisu: Benchmarking Large Language Models in Japanese Finance 17 upvotes, #25 of 2026-02-03
  26. RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 17 upvotes, #25 of 2026-02-03
  27. Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning 16 upvotes, #27 of 2026-02-03
  28. Toward Cognitive Supersensing in Multimodal Large Language Model 16 upvotes, #27 of 2026-02-03
  29. How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing 16 upvotes, #27 of 2026-02-03
  30. Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars 15 upvotes, #30 of 2026-02-03
  31. Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training 14 upvotes, #31 of 2026-02-03
  32. CUA-Skill: Develop Skills for Computer Using Agent 13 upvotes, #32 of 2026-02-03
  33. Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics 13 upvotes, #32 of 2026-02-03
  34. LoopViT: Scaling Visual ARC with Looped Transformers 11 upvotes, #34 of 2026-02-03
  35. AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios 10 upvotes, #35 of 2026-02-03
  36. Show, Don't Tell: Morphing Latent Reasoning into Image Generation 10 upvotes, #35 of 2026-02-03
  37. Sparse Reward Subsystem in Large Language Models 9 upvotes, #37 of 2026-02-03
  38. TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios 9 upvotes, #37 of 2026-02-03
  39. Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 9 upvotes, #37 of 2026-02-03
  40. PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding 8 upvotes, #40 of 2026-02-03
  41. PromptRL: Prompt Matters in RL for Flow-Based Image Generation 8 upvotes, #40 of 2026-02-03
  42. CoDiQ: Test-Time Scaling for Controllable Difficult Question Generation 7 upvotes, #42 of 2026-02-03
  43. VoxServe: Streaming-Centric Serving System for Speech Language Models 6 upvotes, #43 of 2026-02-03
  44. Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry 5 upvotes, #44 of 2026-02-03
  45. A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation 5 upvotes, #44 of 2026-02-03
  46. VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration 5 upvotes, #44 of 2026-02-03
  47. Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning 5 upvotes, #44 of 2026-02-03
  48. Enhancing Multi-Image Understanding through Delimiter Token Scaling 5 upvotes, #44 of 2026-02-03
  49. Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models 5 upvotes, #44 of 2026-02-03
  50. An Empirical Study of World Model Quantization 5 upvotes, #44 of 2026-02-03
  51. Interacted Planes Reveal 3D Line Mapping 4 upvotes, #51 of 2026-02-03
  52. On the Limits of Layer Pruning for Generative Reasoning in LLMs 4 upvotes, #51 of 2026-02-03
  53. Clipping-Free Policy Optimization for Large Language Models 3 upvotes, #53 of 2026-02-03
  54. On the Relationship Between Representation Geometry and Generalization in Deep Neural Networks 3 upvotes, #53 of 2026-02-03
  55. PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers 3 upvotes, #53 of 2026-02-03
  56. Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language Models 3 upvotes, #53 of 2026-02-03
  57. OVD: On-policy Verbal Distillation 2 upvotes, #57 of 2026-02-03
  58. Mano: Restriking Manifold Optimization for LLM Training 2 upvotes, #57 of 2026-02-03
  59. YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation 2 upvotes, #57 of 2026-02-03
  60. SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia 2 upvotes, #57 of 2026-02-03
  61. Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models 2 upvotes, #57 of 2026-02-03
  62. Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning 2 upvotes, #57 of 2026-02-03
  63. Competing Visions of Ethical AI: A Case Study of OpenAI 1 upvotes, #63 of 2026-02-03
  64. Influence Guided Sampling for Domain Adaptation of Text Retrievers 1 upvotes, #63 of 2026-02-03
  65. ParalESN: Enabling parallel information processing in Reservoir Computing 1 upvotes, #63 of 2026-02-03
  66. AI-Generated Image Detectors Overrely on Global Artifacts: Evidence from Inpainting Exchange 1 upvotes, #63 of 2026-02-03
  67. INDIBATOR: Diverse and Fact-Grounded Individuality for Multi-Agent Debate in Molecular Discovery 1 upvotes, #63 of 2026-02-03
  68. Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages 1 upvotes, #63 of 2026-02-03
  69. Implicit neural representation of textures 1 upvotes, #63 of 2026-02-03
  70. Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation 2 upvotes, #70 of 2026-02-03
  71. Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory 0 upvotes, #70 of 2026-02-03
  72. Where to Attend: A Principled Vision-Centric Position Encoding with Parabolas 0 upvotes, #70 of 2026-02-03
  73. Internal Flow Signatures for Self-Checking and Refinement in LLMs 0 upvotes, #70 of 2026-02-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.