Daily Papers of 2025-12-09
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning 72 upvotes, #1 of 2025-12-09
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs 55 upvotes, #2 of 2025-12-09
- Unified Video Editing with Temporal Reasoner 43 upvotes, #3 of 2025-12-09
- Voxify3D: Pixel Art Meets Volumetric Rendering 39 upvotes, #4 of 2025-12-09
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models 31 upvotes, #5 of 2025-12-09
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing 28 upvotes, #6 of 2025-12-09
- Scaling Zero-Shot Reference-to-Video Generation 28 upvotes, #6 of 2025-12-09
- Distribution Matching Variational AutoEncoder 27 upvotes, #8 of 2025-12-09
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems 25 upvotes, #9 of 2025-12-09
- Relational Visual Similarity 23 upvotes, #10 of 2025-12-09
- Multi-view Pyramid Transformer: Look Coarser to See Broader 20 upvotes, #11 of 2025-12-09
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation 19 upvotes, #12 of 2025-12-09
- OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation 17 upvotes, #13 of 2025-12-09
- LongCat-Image Technical Report 17 upvotes, #13 of 2025-12-09
- OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation 17 upvotes, #13 of 2025-12-09
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation 16 upvotes, #16 of 2025-12-09
- SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning 15 upvotes, #17 of 2025-12-09
- Visual Generation Tuning 13 upvotes, #18 of 2025-12-09
- ReCamDriving: LiDAR-Free Camera-Controlled Novel Trajectory Video Generation 8 upvotes, #19 of 2025-12-09
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning 8 upvotes, #19 of 2025-12-09
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning 6 upvotes, #21 of 2025-12-09
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation 5 upvotes, #22 of 2025-12-09
- Embodied Referring Expression Comprehension in Human-Robot Interaction 3 upvotes, #23 of 2025-12-09
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning 3 upvotes, #23 of 2025-12-09
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators 3 upvotes, #23 of 2025-12-09
- Group Representational Position Encoding 3 upvotes, #23 of 2025-12-09
- Small-Gain Nash: Certified Contraction to Nash Equilibria in Differentiable Games 2 upvotes, #27 of 2025-12-09
- DZ-TDPO: Non-Destructive Temporal Alignment for Mutable State Tracking in Long-Context Dialogue 1 upvotes, #28 of 2025-12-09
- Structured Document Translation via Format Reinforcement Learning 1 upvotes, #28 of 2025-12-09
- Vector Quantization using Gaussian Variational Autoencoder 1 upvotes, #28 of 2025-12-09
- JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention 1 upvotes, #28 of 2025-12-09
- The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation 1 upvotes, #32 of 2025-12-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.