Daily Papers of 2025-12-18

  1. Step-GUI Technical Report 121 upvotes, #1 of 2025-12-18
  2. Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition 55 upvotes, #2 of 2025-12-18
  3. DEER: Draft with Diffusion, Verify with Autoregressive Models 41 upvotes, #3 of 2025-12-18
  4. HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices 39 upvotes, #4 of 2025-12-18
  5. Fast and Accurate Causal Parallel Decoding using Jacobi Forcing 39 upvotes, #4 of 2025-12-18
  6. Universal Reasoning Model 37 upvotes, #6 of 2025-12-18
  7. Puzzle Curriculum GRPO for Vision-Centric Reasoning 33 upvotes, #7 of 2025-12-18
  8. MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence 21 upvotes, #8 of 2025-12-18
  9. IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning 19 upvotes, #9 of 2025-12-18
  10. VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs 17 upvotes, #10 of 2025-12-18
  11. Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning 17 upvotes, #10 of 2025-12-18
  12. SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning 16 upvotes, #12 of 2025-12-18
  13. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning 16 upvotes, #12 of 2025-12-18
  14. Robust and Calibrated Detection of Authentic Multimedia Content 15 upvotes, #14 of 2025-12-18
  15. DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models 15 upvotes, #14 of 2025-12-18
  16. FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition 14 upvotes, #16 of 2025-12-18
  17. End-to-End Training for Autoregressive Video Diffusion via Self-Resampling 14 upvotes, #16 of 2025-12-18
  18. Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future 12 upvotes, #18 of 2025-12-18
  19. VABench: A Comprehensive Benchmark for Audio-Video Generation 7 upvotes, #19 of 2025-12-18
  20. Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets 7 upvotes, #19 of 2025-12-18
  21. In Pursuit of Pixel Supervision for Visual Pre-training 7 upvotes, #19 of 2025-12-18
  22. VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression? 6 upvotes, #22 of 2025-12-18
  23. WAY: Estimation of Vessel Destination in Worldwide AIS Trajectory 5 upvotes, #23 of 2025-12-18
  24. SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations 5 upvotes, #23 of 2025-12-18
  25. Understanding and Improving Hyperbolic Deep Reinforcement Learning 5 upvotes, #23 of 2025-12-18
  26. SCOPE: Prompt Evolution for Enhancing Agent Effectiveness 5 upvotes, #23 of 2025-12-18
  27. FrontierCS: Evolving Challenges for Evolving Intelligence 5 upvotes, #23 of 2025-12-18
  28. LikeBench: Evaluating Subjective Likability in LLMs for Personalization 2 upvotes, #28 of 2025-12-18
  29. Hybrid Attribution Priors for Explainable and Robust Model Training 2 upvotes, #28 of 2025-12-18
  30. Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation 1 upvotes, #30 of 2025-12-18
  31. Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics 3 upvotes, #31 of 2025-12-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.