Daily Papers of 2025-06-30

  1. BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing 60 upvotes, #1 of 2025-06-30
  2. LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs 34 upvotes, #2 of 2025-06-30
  3. XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation 28 upvotes, #3 of 2025-06-30
  4. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation 27 upvotes, #4 of 2025-06-30
  5. ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models 22 upvotes, #5 of 2025-06-30
  6. From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios 18 upvotes, #6 of 2025-06-30
  7. Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity 16 upvotes, #7 of 2025-06-30
  8. Ark: An Open-source Python-based Framework for Robot Learning 14 upvotes, #8 of 2025-06-30
  9. Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs 13 upvotes, #9 of 2025-06-30
  10. Spatial Mental Modeling from Limited Views 12 upvotes, #10 of 2025-06-30
  11. Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy 12 upvotes, #10 of 2025-06-30
  12. The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements 11 upvotes, #12 of 2025-06-30
  13. MiCo: Multi-image Contrast for Reinforcement Visual Reasoning 10 upvotes, #13 of 2025-06-30
  14. In-Context Learning Strategies Emerge Rationally 9 upvotes, #14 of 2025-06-30
  15. SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning 8 upvotes, #15 of 2025-06-30
  16. Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training 6 upvotes, #16 of 2025-06-30
  17. Jan-nano Technical Report 6 upvotes, #16 of 2025-06-30
  18. Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning 4 upvotes, #18 of 2025-06-30
  19. Noise Consistency Training: A Native Approach for One-Step Generator in Learning Additional Controls 4 upvotes, #18 of 2025-06-30
  20. Performance Prediction for Large Systems via Text-to-Text Regression 4 upvotes, #18 of 2025-06-30
  21. RetFiner: A Vision-Language Refinement Scheme for Retinal Foundation Models 3 upvotes, #21 of 2025-06-30
  22. Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute 2 upvotes, #22 of 2025-06-30
  23. GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling 2 upvotes, #22 of 2025-06-30
  24. Adaptive Domain Modeling with Language Models: A Multi-Agent Approach to Task Planning 1 upvotes, #24 of 2025-06-30
  25. Global and Local Entailment Learning for Natural World Imagery 1 upvotes, #24 of 2025-06-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.