Daily Papers of 2026-03-24
- Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models 123 upvotes, #1 of 2026-03-24
- Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model 119 upvotes, #2 of 2026-03-24
- OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis 91 upvotes, #3 of 2026-03-24
- Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs 85 upvotes, #4 of 2026-03-24
- LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning 75 upvotes, #5 of 2026-03-24
- VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding 50 upvotes, #6 of 2026-03-24
- SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning 45 upvotes, #7 of 2026-03-24
- Repurposing Geometric Foundation Models for Multi-view Diffusion 45 upvotes, #7 of 2026-03-24
- mSFT: Addressing Dataset Mixtures Overfiting Heterogeneously in Multi-task SFT 37 upvotes, #9 of 2026-03-24
- Manifold-Aware Exploration for Reinforcement Learning in Video Generation 33 upvotes, #10 of 2026-03-24
- F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting 32 upvotes, #11 of 2026-03-24
- On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation 27 upvotes, #12 of 2026-03-24
- Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection 26 upvotes, #13 of 2026-03-24
- RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models 23 upvotes, #14 of 2026-03-24
- BubbleRAG: Evidence-Driven Retrieval-Augmented Generation for Black-Box Knowledge Graphs 21 upvotes, #15 of 2026-03-24
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost 18 upvotes, #16 of 2026-03-24
- SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models 17 upvotes, #17 of 2026-03-24
- REVERE: Reflective Evolving Research Engineer for Scientific Workflows 17 upvotes, #17 of 2026-03-24
- Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation 16 upvotes, #19 of 2026-03-24
- The Universal Normal Embedding 15 upvotes, #20 of 2026-03-24
- Scaling DoRA: High-Rank Adaptation via Factored Norms and Fused Kernels 13 upvotes, #21 of 2026-03-24
- Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models 12 upvotes, #22 of 2026-03-24
- Generalized Discrete Diffusion from Snapshots 11 upvotes, #23 of 2026-03-24
- Agentic AI and the next intelligence explosion 8 upvotes, #24 of 2026-03-24
- ToolRosetta: Bridging Open-Source Repositories and Large Language Model Agents through Automated Tool Standardization 6 upvotes, #25 of 2026-03-24
- Scalable Prompt Routing via Fine-Grained Latent Task Discovery 6 upvotes, #25 of 2026-03-24
- Effective Strategies for Asynchronous Software Engineering Agents 6 upvotes, #25 of 2026-03-24
- Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation 5 upvotes, #28 of 2026-03-24
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe 5 upvotes, #28 of 2026-03-24
- WorldCache: Content-Aware Caching for Accelerated Video World Models 4 upvotes, #30 of 2026-03-24
- Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies 3 upvotes, #31 of 2026-03-24
- SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection 3 upvotes, #31 of 2026-03-24
- FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models 3 upvotes, #31 of 2026-03-24
- AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference 3 upvotes, #31 of 2026-03-24
- MemDLM: Memory-Enhanced DLM Training 3 upvotes, #31 of 2026-03-24
- Aperiodic Structures Never Collapse: Fibonacci Hierarchies for Lossless Compression 2 upvotes, #36 of 2026-03-24
- In-the-Wild Camouflage Attack on Vehicle Detectors through Controllable Image Editing 2 upvotes, #36 of 2026-03-24
- Semantic Audio-Visual Navigation in Continuous Environments 2 upvotes, #36 of 2026-03-24
- Understanding Behavior Cloning with Action Quantization 2 upvotes, #36 of 2026-03-24
- AdditiveLLM2: A Multi-modal Large Language Model for Additive Manufacturing 2 upvotes, #36 of 2026-03-24
- Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs 1 upvotes, #41 of 2026-03-24
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.