Daily Papers of 2026-08-05

  1. MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations 96 upvotes, #1 of 2026-08-05
  2. JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion 91 upvotes, #2 of 2026-08-05
  3. Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing 88 upvotes, #3 of 2026-08-05
  4. AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling 79 upvotes, #4 of 2026-08-05
  5. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 50 upvotes, #5 of 2026-08-05
  6. Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation 47 upvotes, #6 of 2026-08-05
  7. PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning 38 upvotes, #7 of 2026-08-05
  8. Quo Vadis, World Modeling? 37 upvotes, #8 of 2026-08-05
  9. PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents 33 upvotes, #9 of 2026-08-05
  10. LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models 32 upvotes, #10 of 2026-08-05
  11. Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging 26 upvotes, #11 of 2026-08-05
  12. OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models 26 upvotes, #11 of 2026-08-05
  13. CAPEval: A Decoupled Caption Evaluation across Understanding and Generation 25 upvotes, #13 of 2026-08-05
  14. SkillJack: Persistent Skill Backdoors in Self-Evolving Agents 23 upvotes, #14 of 2026-08-05
  15. UniWorld-Design: From Pixel Generation to Layer-Native Design 21 upvotes, #15 of 2026-08-05
  16. MiniWorld: Democratizing the Training of Video World Models from Scratch 19 upvotes, #16 of 2026-08-05
  17. TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning 19 upvotes, #16 of 2026-08-05
  18. Decoding Children's Gait Behavior 15 upvotes, #18 of 2026-08-05
  19. GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience 15 upvotes, #18 of 2026-08-05
  20. Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements 14 upvotes, #20 of 2026-08-05
  21. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? 14 upvotes, #20 of 2026-08-05
  22. ExplainBench: Evaluating Code Explanations from Agents 13 upvotes, #22 of 2026-08-05
  23. RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction 11 upvotes, #23 of 2026-08-05
  24. When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills 9 upvotes, #24 of 2026-08-05
  25. Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories 8 upvotes, #25 of 2026-08-05
  26. When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings 8 upvotes, #25 of 2026-08-05
  27. Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking 7 upvotes, #27 of 2026-08-05
  28. ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts 7 upvotes, #27 of 2026-08-05
  29. PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs 7 upvotes, #27 of 2026-08-05
  30. ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads 6 upvotes, #30 of 2026-08-05
  31. Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation 5 upvotes, #31 of 2026-08-05
  32. CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning 4 upvotes, #32 of 2026-08-05
  33. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels 4 upvotes, #32 of 2026-08-05
  34. LegalPincite: Multi-level Legal Information Retrieval Dataset 4 upvotes, #32 of 2026-08-05
  35. ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning 4 upvotes, #32 of 2026-08-05
  36. Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents 3 upvotes, #36 of 2026-08-05
  37. Multi-Task Multi-Frame Visual Piano Transcription 3 upvotes, #36 of 2026-08-05
  38. When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs 3 upvotes, #36 of 2026-08-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.