Daily Papers of 2025-11-04

  1. Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation 81 upvotes, #1 of 2025-11-04
  2. EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities 47 upvotes, #2 of 2025-11-04
  3. Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph 40 upvotes, #3 of 2025-11-04
  4. World Simulation with Video Foundation Models for Physical AI 39 upvotes, #4 of 2025-11-04
  5. UniREditBench: A Unified Reasoning-based Image Editing Benchmark 36 upvotes, #5 of 2025-11-04
  6. The Underappreciated Power of Vision Models for Graph Structural Understanding 34 upvotes, #6 of 2025-11-04
  7. UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback 33 upvotes, #7 of 2025-11-04
  8. MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models 31 upvotes, #8 of 2025-11-04
  9. ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation 31 upvotes, #8 of 2025-11-04
  10. PHUMA: Physically-Grounded Humanoid Locomotion Dataset 27 upvotes, #10 of 2025-11-04
  11. MotionStream: Real-Time Video Generation with Interactive Motion Controls 25 upvotes, #11 of 2025-11-04
  12. ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use 22 upvotes, #12 of 2025-11-04
  13. LongCat-Flash-Omni Technical Report 21 upvotes, #13 of 2025-11-04
  14. OpenSIR: Open-Ended Self-Improving Reasoner 20 upvotes, #14 of 2025-11-04
  15. Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum 16 upvotes, #15 of 2025-11-04
  16. TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning 15 upvotes, #16 of 2025-11-04
  17. NaviTrace: Evaluating Embodied Navigation of Vision-Language Models 13 upvotes, #17 of 2025-11-04
  18. left|,circlearrowright,text{BUS},right|: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles 12 upvotes, #18 of 2025-11-04
  19. Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench 11 upvotes, #19 of 2025-11-04
  20. Trove: A Flexible Toolkit for Dense Retrieval 10 upvotes, #20 of 2025-11-04
  21. Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models 9 upvotes, #21 of 2025-11-04
  22. Data-Efficient RLVR via Off-Policy Influence Guidance 8 upvotes, #22 of 2025-11-04
  23. Towards Robust Mathematical Reasoning 7 upvotes, #23 of 2025-11-04
  24. Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process 6 upvotes, #24 of 2025-11-04
  25. How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment 6 upvotes, #24 of 2025-11-04
  26. UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings 5 upvotes, #26 of 2025-11-04
  27. GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding 3 upvotes, #27 of 2025-11-04
  28. AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence 3 upvotes, #27 of 2025-11-04
  29. Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers 2 upvotes, #29 of 2025-11-04
  30. Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement 2 upvotes, #29 of 2025-11-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.