Daily Papers of 2026-04-10

  1. Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability 314 upvotes, #1 of 2026-04-10
  2. SkillClaw: Let Skills Evolve Collectively with Agentic Evolver 276 upvotes, #2 of 2026-04-10
  3. ClawBench: Can AI Agents Complete Everyday Online Tasks? 255 upvotes, #3 of 2026-04-10
  4. HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents 182 upvotes, #4 of 2026-04-10
  5. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models 114 upvotes, #5 of 2026-04-10
  6. GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents 108 upvotes, #6 of 2026-04-10
  7. MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping 96 upvotes, #7 of 2026-04-10
  8. LPM 1.0: Video-based Character Performance Model 71 upvotes, #8 of 2026-04-10
  9. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering 50 upvotes, #9 of 2026-04-10
  10. DMax: Aggressive Parallel Decoding for dLLMs 50 upvotes, #9 of 2026-04-10
  11. OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks 48 upvotes, #11 of 2026-04-10
  12. KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation 47 upvotes, #12 of 2026-04-10
  13. MolmoWeb: Open Visual Web Agent and Open Data for the Open Web 41 upvotes, #13 of 2026-04-10
  14. Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models 41 upvotes, #13 of 2026-04-10
  15. OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence 39 upvotes, #15 of 2026-04-10
  16. OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering 25 upvotes, #16 of 2026-04-10
  17. Graph of Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills 21 upvotes, #17 of 2026-04-10
  18. Structured Distillation of Web Agent Capabilities Enables Generalization 20 upvotes, #18 of 2026-04-10
  19. Small Vision-Language Models are Smart Compressors for Long Video Understanding 20 upvotes, #18 of 2026-04-10
  20. FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On 20 upvotes, #18 of 2026-04-10
  21. Automating Database-Native Function Code Synthesis with LLMs 17 upvotes, #21 of 2026-04-10
  22. ViVa: A Video-Generative Value Model for Robot Reinforcement Learning 17 upvotes, #21 of 2026-04-10
  23. Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference 16 upvotes, #23 of 2026-04-10
  24. SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds 16 upvotes, #23 of 2026-04-10
  25. Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces 15 upvotes, #25 of 2026-04-10
  26. Training a Student Expert via Semi-Supervised Foundation Model Distillation 10 upvotes, #26 of 2026-04-10
  27. Lighting-grounded Video Generation with Renderer-based Agent Reasoning 10 upvotes, #26 of 2026-04-10
  28. Personalizing Text-to-Image Generation to Individual Taste 8 upvotes, #28 of 2026-04-10
  29. ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models 8 upvotes, #28 of 2026-04-10
  30. PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models 8 upvotes, #28 of 2026-04-10
  31. Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization 8 upvotes, #28 of 2026-04-10
  32. The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment 7 upvotes, #32 of 2026-04-10
  33. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics 7 upvotes, #32 of 2026-04-10
  34. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors 6 upvotes, #34 of 2026-04-10
  35. POS-ISP: Pipeline Optimization at the Sequence Level for Task-aware ISP 6 upvotes, #34 of 2026-04-10
  36. QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration 5 upvotes, #36 of 2026-04-10
  37. Structural Graph Probing of Vision-Language Models 5 upvotes, #36 of 2026-04-10
  38. Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images 5 upvotes, #36 of 2026-04-10
  39. Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search 5 upvotes, #36 of 2026-04-10
  40. On the Global Photometric Alignment for Low-Level Vision 5 upvotes, #36 of 2026-04-10
  41. RewardFlow: Generate Images by Optimizing What You Reward 5 upvotes, #36 of 2026-04-10
  42. CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation 2 upvotes, #42 of 2026-04-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.