Daily Papers of 2026-05-12

  1. Qwen-Image-2.0 Technical Report 106 upvotes, #1 of 2026-05-12
  2. Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs 77 upvotes, #2 of 2026-05-12
  3. CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models 68 upvotes, #3 of 2026-05-12
  4. Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training 50 upvotes, #4 of 2026-05-12
  5. TMAS: Scaling Test-Time Compute via Multi-Agent Synergy 49 upvotes, #5 of 2026-05-12
  6. Model Merging Scaling Laws in Large Language Models 44 upvotes, #6 of 2026-05-12
  7. PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents 32 upvotes, #7 of 2026-05-12
  8. WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors 30 upvotes, #8 of 2026-05-12
  9. Pixal3D: Pixel-Aligned 3D Generation from Images 30 upvotes, #8 of 2026-05-12
  10. SEIF: Self-Evolving Reinforcement Learning for Instruction Following 29 upvotes, #10 of 2026-05-12
  11. Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models 29 upvotes, #10 of 2026-05-12
  12. Key-Value Means 24 upvotes, #12 of 2026-05-12
  13. Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria 23 upvotes, #13 of 2026-05-12
  14. X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction 22 upvotes, #14 of 2026-05-12
  15. LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? 21 upvotes, #15 of 2026-05-12
  16. G-Zero: Self-Play for Open-Ended Generation from Zero Data 16 upvotes, #16 of 2026-05-12
  17. Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR 16 upvotes, #16 of 2026-05-12
  18. RigidFormer: Learning Rigid Dynamics using Transformers 14 upvotes, #18 of 2026-05-12
  19. ELF: Embedded Language Flows 14 upvotes, #18 of 2026-05-12
  20. Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control 13 upvotes, #20 of 2026-05-12
  21. SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training 13 upvotes, #20 of 2026-05-12
  22. NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation 13 upvotes, #20 of 2026-05-12
  23. Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning 13 upvotes, #20 of 2026-05-12
  24. A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models 11 upvotes, #24 of 2026-05-12
  25. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction 11 upvotes, #24 of 2026-05-12
  26. jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition 10 upvotes, #26 of 2026-05-12
  27. SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding 9 upvotes, #27 of 2026-05-12
  28. Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions 9 upvotes, #27 of 2026-05-12
  29. AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems 8 upvotes, #29 of 2026-05-12
  30. Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization 8 upvotes, #29 of 2026-05-12
  31. Reinforcing Multimodal Reasoning Against Visual Degradation 7 upvotes, #31 of 2026-05-12
  32. Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding 7 upvotes, #31 of 2026-05-12
  33. Mela: Test-Time Memory Consolidation based on Transformation Hypothesis 7 upvotes, #31 of 2026-05-12
  34. Conformal Agent Error Attribution 6 upvotes, #34 of 2026-05-12
  35. FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration 6 upvotes, #34 of 2026-05-12
  36. DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification 6 upvotes, #34 of 2026-05-12
  37. Can Muon Fine-tune Adam-Pretrained Models? 6 upvotes, #34 of 2026-05-12
  38. MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation 5 upvotes, #38 of 2026-05-12
  39. Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon 5 upvotes, #38 of 2026-05-12
  40. Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient? 5 upvotes, #38 of 2026-05-12
  41. Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why 4 upvotes, #41 of 2026-05-12
  42. RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark 4 upvotes, #41 of 2026-05-12
  43. Training-Free Dense Hand Contact Estimation with Multi-Modal Large Language Models 3 upvotes, #43 of 2026-05-12
  44. Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms 3 upvotes, #43 of 2026-05-12
  45. FORTIS: Benchmarking Over-Privilege in Agent Skills 3 upvotes, #43 of 2026-05-12
  46. Crosslingual On-Policy Self-Distillation for Multilingual Reasoning 3 upvotes, #43 of 2026-05-12
  47. DeepRefine: Agent-Compiled Knowledge Refinement via Reinforcement Learning 3 upvotes, #43 of 2026-05-12
  48. GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs 3 upvotes, #43 of 2026-05-12
  49. DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices 3 upvotes, #43 of 2026-05-12
  50. PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning 2 upvotes, #50 of 2026-05-12
  51. Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression 1 upvotes, #51 of 2026-05-12
  52. InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition 1 upvotes, #51 of 2026-05-12
  53. Uncovering Entity Identity Confusion in Multimodal Knowledge Editing 1 upvotes, #51 of 2026-05-12
  54. SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis 1 upvotes, #51 of 2026-05-12
  55. 100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts 1 upvotes, #51 of 2026-05-12
  56. Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimization 1 upvotes, #51 of 2026-05-12
  57. LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language 1 upvotes, #51 of 2026-05-12
  58. Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models 1 upvotes, #51 of 2026-05-12
  59. SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning 1 upvotes, #51 of 2026-05-12
  60. Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models 1 upvotes, #51 of 2026-05-12
  61. Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inference 1 upvotes, #51 of 2026-05-12
  62. Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace 1 upvotes, #51 of 2026-05-12
  63. A Closed-Form Upper Bound for Admissible Learning-Rate Steps in Belief-Space Dynamics 1 upvotes, #63 of 2026-05-12
  64. Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents 2 upvotes, #63 of 2026-05-12
  65. Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do) 0 upvotes, #63 of 2026-05-12
  66. TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation 1 upvotes, #63 of 2026-05-12
  67. The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection 0 upvotes, #63 of 2026-05-12
  68. CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models 0 upvotes, #63 of 2026-05-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.