Daily Papers of 2026-07-02

  1. PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception 41 upvotes, #1 of 2026-07-02
  2. TurboServe: Serving Streaming Video Generation Efficiently and Economically 34 upvotes, #2 of 2026-07-02
  3. ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving 31 upvotes, #3 of 2026-07-02
  4. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory 29 upvotes, #4 of 2026-07-02
  5. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity 28 upvotes, #5 of 2026-07-02
  6. Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning 27 upvotes, #6 of 2026-07-02
  7. AutoTrainess: Teaching Language Models to Improve Language Models Autonomously 22 upvotes, #7 of 2026-07-02
  8. ASPIRE: Agentic /Skills Discovery for Robotics 22 upvotes, #7 of 2026-07-02
  9. Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts 22 upvotes, #7 of 2026-07-02
  10. ABot-M0.5: Unified Mobility-and-Manipulation World Action Model 19 upvotes, #10 of 2026-07-02
  11. CausalMix: Data Mixture as Causal Inference for Language Model Training 19 upvotes, #10 of 2026-07-02
  12. Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning 18 upvotes, #12 of 2026-07-02
  13. Valdi: Value Diffusion World Models 15 upvotes, #13 of 2026-07-02
  14. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors 14 upvotes, #14 of 2026-07-02
  15. BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery 13 upvotes, #15 of 2026-07-02
  16. PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking 12 upvotes, #16 of 2026-07-02
  17. Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks 11 upvotes, #17 of 2026-07-02
  18. Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination 11 upvotes, #17 of 2026-07-02
  19. The State-Prediction Separation Hypothesis 11 upvotes, #17 of 2026-07-02
  20. Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising 10 upvotes, #20 of 2026-07-02
  21. NoPA: Non-Parametric Online 3D Scene Graph Generation 9 upvotes, #21 of 2026-07-02
  22. AI translation of literary texts is "fine", but readers still prefer human translations 8 upvotes, #22 of 2026-07-02
  23. Building to the Test: Coding Agents Deliver What You Check, Not What You Requested 8 upvotes, #22 of 2026-07-02
  24. When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling 8 upvotes, #22 of 2026-07-02
  25. AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation 8 upvotes, #22 of 2026-07-02
  26. GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity 8 upvotes, #22 of 2026-07-02
  27. Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? 8 upvotes, #22 of 2026-07-02
  28. Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation 7 upvotes, #28 of 2026-07-02
  29. SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation 7 upvotes, #28 of 2026-07-02
  30. Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue 7 upvotes, #28 of 2026-07-02
  31. Autonomous Scientific Discovery via Iterative Meta-Reflection 7 upvotes, #28 of 2026-07-02
  32. CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion 6 upvotes, #32 of 2026-07-02
  33. HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents 6 upvotes, #32 of 2026-07-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.