Daily Papers of 2025-11-04
- Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation 81 upvotes, #1 of 2025-11-04
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities 47 upvotes, #2 of 2025-11-04
- Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph 40 upvotes, #3 of 2025-11-04
- World Simulation with Video Foundation Models for Physical AI 39 upvotes, #4 of 2025-11-04
- UniREditBench: A Unified Reasoning-based Image Editing Benchmark 36 upvotes, #5 of 2025-11-04
- The Underappreciated Power of Vision Models for Graph Structural Understanding 34 upvotes, #6 of 2025-11-04
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback 33 upvotes, #7 of 2025-11-04
- MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models 31 upvotes, #8 of 2025-11-04
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation 31 upvotes, #8 of 2025-11-04
- PHUMA: Physically-Grounded Humanoid Locomotion Dataset 27 upvotes, #10 of 2025-11-04
- MotionStream: Real-Time Video Generation with Interactive Motion Controls 25 upvotes, #11 of 2025-11-04
- ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use 22 upvotes, #12 of 2025-11-04
- LongCat-Flash-Omni Technical Report 21 upvotes, #13 of 2025-11-04
- OpenSIR: Open-Ended Self-Improving Reasoner 20 upvotes, #14 of 2025-11-04
- Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum 16 upvotes, #15 of 2025-11-04
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning 15 upvotes, #16 of 2025-11-04
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models 13 upvotes, #17 of 2025-11-04
- left|,circlearrowright,text{BUS},right|: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles 12 upvotes, #18 of 2025-11-04
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench 11 upvotes, #19 of 2025-11-04
- Trove: A Flexible Toolkit for Dense Retrieval 10 upvotes, #20 of 2025-11-04
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models 9 upvotes, #21 of 2025-11-04
- Data-Efficient RLVR via Off-Policy Influence Guidance 8 upvotes, #22 of 2025-11-04
- Towards Robust Mathematical Reasoning 7 upvotes, #23 of 2025-11-04
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process 6 upvotes, #24 of 2025-11-04
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment 6 upvotes, #24 of 2025-11-04
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings 5 upvotes, #26 of 2025-11-04
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding 3 upvotes, #27 of 2025-11-04
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence 3 upvotes, #27 of 2025-11-04
- Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers 2 upvotes, #29 of 2025-11-04
- Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement 2 upvotes, #29 of 2025-11-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.