Daily Papers of 2025-11-07

  1. Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 187 upvotes, #1 of 2025-11-07
  2. V-Thinker: Interactive Thinking with Images 93 upvotes, #2 of 2025-11-07
  3. Scaling Agent Learning via Experience Synthesis 72 upvotes, #3 of 2025-11-07
  4. Cambrian-S: Towards Spatial Supersensing in Video 34 upvotes, #4 of 2025-11-07
  5. NVIDIA Nemotron Nano V2 VL 25 upvotes, #5 of 2025-11-07
  6. The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms 15 upvotes, #6 of 2025-11-07
  7. GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents 14 upvotes, #7 of 2025-11-07
  8. Contamination Detection for VLMs using Multi-Modal Semantic Perturbation 12 upvotes, #8 of 2025-11-07
  9. Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts 7 upvotes, #9 of 2025-11-07
  10. RDMA Point-to-Point Communication for LLM Systems 5 upvotes, #10 of 2025-11-07
  11. EVTAR: End-to-End Try on with Additional Unpaired Visual Reference 4 upvotes, #11 of 2025-11-07
  12. SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding 4 upvotes, #11 of 2025-11-07
  13. SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning 3 upvotes, #13 of 2025-11-07
  14. How to Evaluate Speech Translation with Source-Aware Neural MT Metrics 3 upvotes, #13 of 2025-11-07
  15. Learning Vision-Driven Reactive Soccer Skills for Humanoid Robots 3 upvotes, #13 of 2025-11-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.