Daily Papers of 2025-10-16
- FlashWorld: High-quality 3D Scene Generation within Seconds 67 upvotes, #1 of 2025-10-16
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE 60 upvotes, #2 of 2025-10-16
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization 54 upvotes, #3 of 2025-10-16
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs 49 upvotes, #4 of 2025-10-16
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models 42 upvotes, #5 of 2025-10-16
- PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning 36 upvotes, #6 of 2025-10-16
- Trace Anything: Representing Any Video in 4D via Trajectory Fields 30 upvotes, #7 of 2025-10-16
- The Art of Scaling Reinforcement Learning Compute for LLMs 29 upvotes, #8 of 2025-10-16
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue 28 upvotes, #9 of 2025-10-16
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs 26 upvotes, #10 of 2025-10-16
- Stronger Together: On-Policy Reinforcement Learning for Collaborative LLMs 25 upvotes, #11 of 2025-10-16
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving 24 upvotes, #12 of 2025-10-16
- Generative Universal Verifier as Multimodal Meta-Reasoner 24 upvotes, #12 of 2025-10-16
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy 16 upvotes, #14 of 2025-10-16
- The Role of Computing Resources in Publishing Foundation Model Research 14 upvotes, #15 of 2025-10-16
- Reasoning in Space via Grounding in the World 14 upvotes, #15 of 2025-10-16
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model 13 upvotes, #17 of 2025-10-16
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning 11 upvotes, #18 of 2025-10-16
- What Generative Search Engines Like and How to Optimize Web Content Cooperatively 10 upvotes, #19 of 2025-10-16
- Universal Image Restoration Pre-training via Masked Degradation Classification 10 upvotes, #19 of 2025-10-16
- Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark 9 upvotes, #21 of 2025-10-16
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model 8 upvotes, #22 of 2025-10-16
- Revisiting Model Interpolation for Efficient Reasoning 8 upvotes, #22 of 2025-10-16
- Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention 6 upvotes, #24 of 2025-10-16
- Direct Multi-Token Decoding 5 upvotes, #25 of 2025-10-16
- HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication 4 upvotes, #26 of 2025-10-16
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving 4 upvotes, #26 of 2025-10-16
- Learning to Grasp Anything by Playing with Random Toys 4 upvotes, #26 of 2025-10-16
- NOSA: Native and Offloadable Sparse Attention 4 upvotes, #26 of 2025-10-16
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math 4 upvotes, #26 of 2025-10-16
- Don't Throw Away Your Pretrained Model 2 upvotes, #31 of 2025-10-16
- GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search 2 upvotes, #31 of 2025-10-16
- Point Prompting: Counterfactual Tracking with Video Diffusion Models 2 upvotes, #31 of 2025-10-16
- MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training 2 upvotes, #31 of 2025-10-16
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems 2 upvotes, #31 of 2025-10-16
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation 1 upvotes, #36 of 2025-10-16
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning 1 upvotes, #36 of 2025-10-16
- EAGER: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling 1 upvotes, #36 of 2025-10-16
- MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model 1 upvotes, #36 of 2025-10-16
- Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain 1 upvotes, #36 of 2025-10-16
- Dedelayed: Deleting remote inference delay via on-device correction 1 upvotes, #36 of 2025-10-16
- Evaluating Language Models' Evaluations of Games 1 upvotes, #42 of 2025-10-16
- Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs 1 upvotes, #42 of 2025-10-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.