Daily Papers of 2026-05-12
- Qwen-Image-2.0 Technical Report 106 upvotes, #1 of 2026-05-12
- Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs 77 upvotes, #2 of 2026-05-12
- CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models 68 upvotes, #3 of 2026-05-12
- Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training 50 upvotes, #4 of 2026-05-12
- TMAS: Scaling Test-Time Compute via Multi-Agent Synergy 49 upvotes, #5 of 2026-05-12
- Model Merging Scaling Laws in Large Language Models 44 upvotes, #6 of 2026-05-12
- PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents 32 upvotes, #7 of 2026-05-12
- WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors 30 upvotes, #8 of 2026-05-12
- Pixal3D: Pixel-Aligned 3D Generation from Images 30 upvotes, #8 of 2026-05-12
- SEIF: Self-Evolving Reinforcement Learning for Instruction Following 29 upvotes, #10 of 2026-05-12
- Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models 29 upvotes, #10 of 2026-05-12
- Key-Value Means 24 upvotes, #12 of 2026-05-12
- Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria 23 upvotes, #13 of 2026-05-12
- X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction 22 upvotes, #14 of 2026-05-12
- LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? 21 upvotes, #15 of 2026-05-12
- G-Zero: Self-Play for Open-Ended Generation from Zero Data 16 upvotes, #16 of 2026-05-12
- Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR 16 upvotes, #16 of 2026-05-12
- RigidFormer: Learning Rigid Dynamics using Transformers 14 upvotes, #18 of 2026-05-12
- ELF: Embedded Language Flows 14 upvotes, #18 of 2026-05-12
- Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control 13 upvotes, #20 of 2026-05-12
- SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training 13 upvotes, #20 of 2026-05-12
- NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation 13 upvotes, #20 of 2026-05-12
- Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning 13 upvotes, #20 of 2026-05-12
- A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models 11 upvotes, #24 of 2026-05-12
- Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction 11 upvotes, #24 of 2026-05-12
- jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition 10 upvotes, #26 of 2026-05-12
- SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding 9 upvotes, #27 of 2026-05-12
- Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions 9 upvotes, #27 of 2026-05-12
- AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems 8 upvotes, #29 of 2026-05-12
- Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization 8 upvotes, #29 of 2026-05-12
- Reinforcing Multimodal Reasoning Against Visual Degradation 7 upvotes, #31 of 2026-05-12
- Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding 7 upvotes, #31 of 2026-05-12
- Mela: Test-Time Memory Consolidation based on Transformation Hypothesis 7 upvotes, #31 of 2026-05-12
- Conformal Agent Error Attribution 6 upvotes, #34 of 2026-05-12
- FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration 6 upvotes, #34 of 2026-05-12
- DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification 6 upvotes, #34 of 2026-05-12
- Can Muon Fine-tune Adam-Pretrained Models? 6 upvotes, #34 of 2026-05-12
- MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation 5 upvotes, #38 of 2026-05-12
- Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon 5 upvotes, #38 of 2026-05-12
- Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient? 5 upvotes, #38 of 2026-05-12
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why 4 upvotes, #41 of 2026-05-12
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark 4 upvotes, #41 of 2026-05-12
- Training-Free Dense Hand Contact Estimation with Multi-Modal Large Language Models 3 upvotes, #43 of 2026-05-12
- Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms 3 upvotes, #43 of 2026-05-12
- FORTIS: Benchmarking Over-Privilege in Agent Skills 3 upvotes, #43 of 2026-05-12
- Crosslingual On-Policy Self-Distillation for Multilingual Reasoning 3 upvotes, #43 of 2026-05-12
- DeepRefine: Agent-Compiled Knowledge Refinement via Reinforcement Learning 3 upvotes, #43 of 2026-05-12
- GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs 3 upvotes, #43 of 2026-05-12
- DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices 3 upvotes, #43 of 2026-05-12
- PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning 2 upvotes, #50 of 2026-05-12
- Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression 1 upvotes, #51 of 2026-05-12
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition 1 upvotes, #51 of 2026-05-12
- Uncovering Entity Identity Confusion in Multimodal Knowledge Editing 1 upvotes, #51 of 2026-05-12
- SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis 1 upvotes, #51 of 2026-05-12
- 100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts 1 upvotes, #51 of 2026-05-12
- Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimization 1 upvotes, #51 of 2026-05-12
- LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language 1 upvotes, #51 of 2026-05-12
- Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models 1 upvotes, #51 of 2026-05-12
- SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning 1 upvotes, #51 of 2026-05-12
- Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models 1 upvotes, #51 of 2026-05-12
- Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inference 1 upvotes, #51 of 2026-05-12
- Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace 1 upvotes, #51 of 2026-05-12
- A Closed-Form Upper Bound for Admissible Learning-Rate Steps in Belief-Space Dynamics 1 upvotes, #63 of 2026-05-12
- Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents 2 upvotes, #63 of 2026-05-12
- Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do) 0 upvotes, #63 of 2026-05-12
- TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation 1 upvotes, #63 of 2026-05-12
- The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection 0 upvotes, #63 of 2026-05-12
- CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models 0 upvotes, #63 of 2026-05-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.