Daily Papers of 2025-06-30
- BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing 60 upvotes, #1 of 2025-06-30
- LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs 34 upvotes, #2 of 2025-06-30
- XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation 28 upvotes, #3 of 2025-06-30
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation 27 upvotes, #4 of 2025-06-30
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models 22 upvotes, #5 of 2025-06-30
- From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios 18 upvotes, #6 of 2025-06-30
- Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity 16 upvotes, #7 of 2025-06-30
- Ark: An Open-source Python-based Framework for Robot Learning 14 upvotes, #8 of 2025-06-30
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs 13 upvotes, #9 of 2025-06-30
- Spatial Mental Modeling from Limited Views 12 upvotes, #10 of 2025-06-30
- Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy 12 upvotes, #10 of 2025-06-30
- The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements 11 upvotes, #12 of 2025-06-30
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning 10 upvotes, #13 of 2025-06-30
- In-Context Learning Strategies Emerge Rationally 9 upvotes, #14 of 2025-06-30
- SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning 8 upvotes, #15 of 2025-06-30
- Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training 6 upvotes, #16 of 2025-06-30
- Jan-nano Technical Report 6 upvotes, #16 of 2025-06-30
- Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning 4 upvotes, #18 of 2025-06-30
- Noise Consistency Training: A Native Approach for One-Step Generator in Learning Additional Controls 4 upvotes, #18 of 2025-06-30
- Performance Prediction for Large Systems via Text-to-Text Regression 4 upvotes, #18 of 2025-06-30
- RetFiner: A Vision-Language Refinement Scheme for Retinal Foundation Models 3 upvotes, #21 of 2025-06-30
- Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute 2 upvotes, #22 of 2025-06-30
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling 2 upvotes, #22 of 2025-06-30
- Adaptive Domain Modeling with Language Models: A Multi-Agent Approach to Task Planning 1 upvotes, #24 of 2025-06-30
- Global and Local Entailment Learning for Natural World Imagery 1 upvotes, #24 of 2025-06-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.