Daily Papers of 2025-12-04
- Qwen3-VL Technical Report 120 upvotes, #1 of 2025-12-04
- PretrainZero: Reinforcement Active Pretraining 44 upvotes, #2 of 2025-12-04
- Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach 39 upvotes, #3 of 2025-12-04
- OneThinker: All-in-one Reasoning Model for Image and Video 30 upvotes, #4 of 2025-12-04
- ViDiC: Video Difference Captioning 26 upvotes, #5 of 2025-12-04
- RELIC: Interactive Video World Model with Long-Horizon Memory 23 upvotes, #6 of 2025-12-04
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL 21 upvotes, #7 of 2025-12-04
- Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation 18 upvotes, #8 of 2025-12-04
- Thinking with Programming Vision: Towards a Unified View for Thinking with Images 15 upvotes, #9 of 2025-12-04
- PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design 13 upvotes, #10 of 2025-12-04
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment 12 upvotes, #11 of 2025-12-04
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation 12 upvotes, #11 of 2025-12-04
- Jina-VLM: Small Multilingual Vision Language Model 12 upvotes, #11 of 2025-12-04
- Light-X: Generative 4D Video Rendering with Camera and Illumination Control 10 upvotes, #14 of 2025-12-04
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment 7 upvotes, #15 of 2025-12-04
- AutoNeural: Co-Designing Vision-Language Models for NPU Inference 5 upvotes, #16 of 2025-12-04
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem 4 upvotes, #17 of 2025-12-04
- AlignBench: Benchmarking Fine-Grained Image-Text Alignment with Synthetic Image-Caption Pairs 3 upvotes, #18 of 2025-12-04
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs 3 upvotes, #18 of 2025-12-04
- In-Context Representation Hijacking 3 upvotes, #18 of 2025-12-04
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition 3 upvotes, #18 of 2025-12-04
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding 3 upvotes, #18 of 2025-12-04
- SkillFactory: Self-Distillation For Learning Cognitive Behaviors 3 upvotes, #18 of 2025-12-04
- BlurDM: A Blur Diffusion Model for Image Deblurring 2 upvotes, #24 of 2025-12-04
- PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation 2 upvotes, #24 of 2025-12-04
- Adversarial Confusion Attack: Disrupting Multimodal Large Language Models 1 upvotes, #26 of 2025-12-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.