Daily Papers of 2026-04-10
- Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability 314 upvotes, #1 of 2026-04-10
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver 276 upvotes, #2 of 2026-04-10
- ClawBench: Can AI Agents Complete Everyday Online Tasks? 255 upvotes, #3 of 2026-04-10
- HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents 182 upvotes, #4 of 2026-04-10
- When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models 114 upvotes, #5 of 2026-04-10
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents 108 upvotes, #6 of 2026-04-10
- MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping 96 upvotes, #7 of 2026-04-10
- LPM 1.0: Video-based Character Performance Model 71 upvotes, #8 of 2026-04-10
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering 50 upvotes, #9 of 2026-04-10
- DMax: Aggressive Parallel Decoding for dLLMs 50 upvotes, #9 of 2026-04-10
- OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks 48 upvotes, #11 of 2026-04-10
- KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation 47 upvotes, #12 of 2026-04-10
- MolmoWeb: Open Visual Web Agent and Open Data for the Open Web 41 upvotes, #13 of 2026-04-10
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models 41 upvotes, #13 of 2026-04-10
- OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence 39 upvotes, #15 of 2026-04-10
- OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering 25 upvotes, #16 of 2026-04-10
- Graph of Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills 21 upvotes, #17 of 2026-04-10
- Structured Distillation of Web Agent Capabilities Enables Generalization 20 upvotes, #18 of 2026-04-10
- Small Vision-Language Models are Smart Compressors for Long Video Understanding 20 upvotes, #18 of 2026-04-10
- FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On 20 upvotes, #18 of 2026-04-10
- Automating Database-Native Function Code Synthesis with LLMs 17 upvotes, #21 of 2026-04-10
- ViVa: A Video-Generative Value Model for Robot Reinforcement Learning 17 upvotes, #21 of 2026-04-10
- Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference 16 upvotes, #23 of 2026-04-10
- SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds 16 upvotes, #23 of 2026-04-10
- Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces 15 upvotes, #25 of 2026-04-10
- Training a Student Expert via Semi-Supervised Foundation Model Distillation 10 upvotes, #26 of 2026-04-10
- Lighting-grounded Video Generation with Renderer-based Agent Reasoning 10 upvotes, #26 of 2026-04-10
- Personalizing Text-to-Image Generation to Individual Taste 8 upvotes, #28 of 2026-04-10
- ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models 8 upvotes, #28 of 2026-04-10
- PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models 8 upvotes, #28 of 2026-04-10
- Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization 8 upvotes, #28 of 2026-04-10
- The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment 7 upvotes, #32 of 2026-04-10
- Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics 7 upvotes, #32 of 2026-04-10
- AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors 6 upvotes, #34 of 2026-04-10
- POS-ISP: Pipeline Optimization at the Sequence Level for Task-aware ISP 6 upvotes, #34 of 2026-04-10
- QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration 5 upvotes, #36 of 2026-04-10
- Structural Graph Probing of Vision-Language Models 5 upvotes, #36 of 2026-04-10
- Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images 5 upvotes, #36 of 2026-04-10
- Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search 5 upvotes, #36 of 2026-04-10
- On the Global Photometric Alignment for Low-Level Vision 5 upvotes, #36 of 2026-04-10
- RewardFlow: Generate Images by Optimizing What You Reward 5 upvotes, #36 of 2026-04-10
- CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation 2 upvotes, #42 of 2026-04-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.