Daily Papers of 2026-02-03
- Green-VLA: Staged Vision-Language-Action Model for Generalist Robots 266 upvotes, #1 of 2026-02-03
- Kimi K2.5: Visual Agentic Intelligence 219 upvotes, #2 of 2026-02-03
- Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models 149 upvotes, #3 of 2026-02-03
- Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models 124 upvotes, #4 of 2026-02-03
- Closing the Loop: Universal Repository Representation with RPG-Encoder 82 upvotes, #5 of 2026-02-03
- UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing 75 upvotes, #6 of 2026-02-03
- SWE-Universe: Scale Real-World Verifiable Environments to Millions 59 upvotes, #7 of 2026-02-03
- FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents 45 upvotes, #8 of 2026-02-03
- SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning 44 upvotes, #9 of 2026-02-03
- PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss 41 upvotes, #10 of 2026-02-03
- WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora 40 upvotes, #11 of 2026-02-03
- Rethinking Generative Recommender Tokenizer: Recsys-Native Encoding and Semantic Quantization Beyond LLMs 40 upvotes, #11 of 2026-02-03
- Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning 39 upvotes, #13 of 2026-02-03
- Generative Visual Code Mobile World Models 39 upvotes, #13 of 2026-02-03
- Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling 35 upvotes, #15 of 2026-02-03
- Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles 32 upvotes, #16 of 2026-02-03
- RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System 31 upvotes, #17 of 2026-02-03
- SLIME: Stabilized Likelihood Implicit Margin Enforcement for Preference Optimization 29 upvotes, #18 of 2026-02-03
- Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention 27 upvotes, #19 of 2026-02-03
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation 24 upvotes, #20 of 2026-02-03
- PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards 23 upvotes, #21 of 2026-02-03
- Rethinking Selective Knowledge Distillation 22 upvotes, #22 of 2026-02-03
- Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation 22 upvotes, #22 of 2026-02-03
- FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space 18 upvotes, #24 of 2026-02-03
- Ebisu: Benchmarking Large Language Models in Japanese Finance 17 upvotes, #25 of 2026-02-03
- RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 17 upvotes, #25 of 2026-02-03
- Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning 16 upvotes, #27 of 2026-02-03
- Toward Cognitive Supersensing in Multimodal Large Language Model 16 upvotes, #27 of 2026-02-03
- How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing 16 upvotes, #27 of 2026-02-03
- Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars 15 upvotes, #30 of 2026-02-03
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training 14 upvotes, #31 of 2026-02-03
- CUA-Skill: Develop Skills for Computer Using Agent 13 upvotes, #32 of 2026-02-03
- Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics 13 upvotes, #32 of 2026-02-03
- LoopViT: Scaling Visual ARC with Looped Transformers 11 upvotes, #34 of 2026-02-03
- AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios 10 upvotes, #35 of 2026-02-03
- Show, Don't Tell: Morphing Latent Reasoning into Image Generation 10 upvotes, #35 of 2026-02-03
- Sparse Reward Subsystem in Large Language Models 9 upvotes, #37 of 2026-02-03
- TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios 9 upvotes, #37 of 2026-02-03
- Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 9 upvotes, #37 of 2026-02-03
- PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding 8 upvotes, #40 of 2026-02-03
- PromptRL: Prompt Matters in RL for Flow-Based Image Generation 8 upvotes, #40 of 2026-02-03
- CoDiQ: Test-Time Scaling for Controllable Difficult Question Generation 7 upvotes, #42 of 2026-02-03
- VoxServe: Streaming-Centric Serving System for Speech Language Models 6 upvotes, #43 of 2026-02-03
- Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry 5 upvotes, #44 of 2026-02-03
- A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation 5 upvotes, #44 of 2026-02-03
- VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration 5 upvotes, #44 of 2026-02-03
- Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning 5 upvotes, #44 of 2026-02-03
- Enhancing Multi-Image Understanding through Delimiter Token Scaling 5 upvotes, #44 of 2026-02-03
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models 5 upvotes, #44 of 2026-02-03
- An Empirical Study of World Model Quantization 5 upvotes, #44 of 2026-02-03
- Interacted Planes Reveal 3D Line Mapping 4 upvotes, #51 of 2026-02-03
- On the Limits of Layer Pruning for Generative Reasoning in LLMs 4 upvotes, #51 of 2026-02-03
- Clipping-Free Policy Optimization for Large Language Models 3 upvotes, #53 of 2026-02-03
- On the Relationship Between Representation Geometry and Generalization in Deep Neural Networks 3 upvotes, #53 of 2026-02-03
- PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers 3 upvotes, #53 of 2026-02-03
- Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language Models 3 upvotes, #53 of 2026-02-03
- OVD: On-policy Verbal Distillation 2 upvotes, #57 of 2026-02-03
- Mano: Restriking Manifold Optimization for LLM Training 2 upvotes, #57 of 2026-02-03
- YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation 2 upvotes, #57 of 2026-02-03
- SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia 2 upvotes, #57 of 2026-02-03
- Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models 2 upvotes, #57 of 2026-02-03
- Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning 2 upvotes, #57 of 2026-02-03
- Competing Visions of Ethical AI: A Case Study of OpenAI 1 upvotes, #63 of 2026-02-03
- Influence Guided Sampling for Domain Adaptation of Text Retrievers 1 upvotes, #63 of 2026-02-03
- ParalESN: Enabling parallel information processing in Reservoir Computing 1 upvotes, #63 of 2026-02-03
- AI-Generated Image Detectors Overrely on Global Artifacts: Evidence from Inpainting Exchange 1 upvotes, #63 of 2026-02-03
- INDIBATOR: Diverse and Fact-Grounded Individuality for Multi-Agent Debate in Molecular Discovery 1 upvotes, #63 of 2026-02-03
- Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages 1 upvotes, #63 of 2026-02-03
- Implicit neural representation of textures 1 upvotes, #63 of 2026-02-03
- Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation 2 upvotes, #70 of 2026-02-03
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory 0 upvotes, #70 of 2026-02-03
- Where to Attend: A Principled Vision-Centric Position Encoding with Parabolas 0 upvotes, #70 of 2026-02-03
- Internal Flow Signatures for Self-Checking and Refinement in LLMs 0 upvotes, #70 of 2026-02-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.