Daily Papers of 2025-12-18
- Step-GUI Technical Report 121 upvotes, #1 of 2025-12-18
- Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition 55 upvotes, #2 of 2025-12-18
- DEER: Draft with Diffusion, Verify with Autoregressive Models 41 upvotes, #3 of 2025-12-18
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices 39 upvotes, #4 of 2025-12-18
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing 39 upvotes, #4 of 2025-12-18
- Universal Reasoning Model 37 upvotes, #6 of 2025-12-18
- Puzzle Curriculum GRPO for Vision-Centric Reasoning 33 upvotes, #7 of 2025-12-18
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence 21 upvotes, #8 of 2025-12-18
- IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning 19 upvotes, #9 of 2025-12-18
- VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs 17 upvotes, #10 of 2025-12-18
- Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning 17 upvotes, #10 of 2025-12-18
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning 16 upvotes, #12 of 2025-12-18
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning 16 upvotes, #12 of 2025-12-18
- Robust and Calibrated Detection of Authentic Multimedia Content 15 upvotes, #14 of 2025-12-18
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models 15 upvotes, #14 of 2025-12-18
- FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition 14 upvotes, #16 of 2025-12-18
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling 14 upvotes, #16 of 2025-12-18
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future 12 upvotes, #18 of 2025-12-18
- VABench: A Comprehensive Benchmark for Audio-Video Generation 7 upvotes, #19 of 2025-12-18
- Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets 7 upvotes, #19 of 2025-12-18
- In Pursuit of Pixel Supervision for Visual Pre-training 7 upvotes, #19 of 2025-12-18
- VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression? 6 upvotes, #22 of 2025-12-18
- WAY: Estimation of Vessel Destination in Worldwide AIS Trajectory 5 upvotes, #23 of 2025-12-18
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations 5 upvotes, #23 of 2025-12-18
- Understanding and Improving Hyperbolic Deep Reinforcement Learning 5 upvotes, #23 of 2025-12-18
- SCOPE: Prompt Evolution for Enhancing Agent Effectiveness 5 upvotes, #23 of 2025-12-18
- FrontierCS: Evolving Challenges for Evolving Intelligence 5 upvotes, #23 of 2025-12-18
- LikeBench: Evaluating Subjective Likability in LLMs for Personalization 2 upvotes, #28 of 2025-12-18
- Hybrid Attribution Priors for Explainable and Robust Model Training 2 upvotes, #28 of 2025-12-18
- Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation 1 upvotes, #30 of 2025-12-18
- Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics 3 upvotes, #31 of 2025-12-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.