Daily Papers of 2026-08-05
- MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations 96 upvotes, #1 of 2026-08-05
- JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion 91 upvotes, #2 of 2026-08-05
- Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing 88 upvotes, #3 of 2026-08-05
- AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling 79 upvotes, #4 of 2026-08-05
- Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 50 upvotes, #5 of 2026-08-05
- Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation 47 upvotes, #6 of 2026-08-05
- PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning 38 upvotes, #7 of 2026-08-05
- Quo Vadis, World Modeling? 37 upvotes, #8 of 2026-08-05
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents 33 upvotes, #9 of 2026-08-05
- LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models 32 upvotes, #10 of 2026-08-05
- Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging 26 upvotes, #11 of 2026-08-05
- OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models 26 upvotes, #11 of 2026-08-05
- CAPEval: A Decoupled Caption Evaluation across Understanding and Generation 25 upvotes, #13 of 2026-08-05
- SkillJack: Persistent Skill Backdoors in Self-Evolving Agents 23 upvotes, #14 of 2026-08-05
- UniWorld-Design: From Pixel Generation to Layer-Native Design 21 upvotes, #15 of 2026-08-05
- MiniWorld: Democratizing the Training of Video World Models from Scratch 19 upvotes, #16 of 2026-08-05
- TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning 19 upvotes, #16 of 2026-08-05
- Decoding Children's Gait Behavior 15 upvotes, #18 of 2026-08-05
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience 15 upvotes, #18 of 2026-08-05
- Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements 14 upvotes, #20 of 2026-08-05
- ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? 14 upvotes, #20 of 2026-08-05
- ExplainBench: Evaluating Code Explanations from Agents 13 upvotes, #22 of 2026-08-05
- RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction 11 upvotes, #23 of 2026-08-05
- When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills 9 upvotes, #24 of 2026-08-05
- Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories 8 upvotes, #25 of 2026-08-05
- When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings 8 upvotes, #25 of 2026-08-05
- Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking 7 upvotes, #27 of 2026-08-05
- ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts 7 upvotes, #27 of 2026-08-05
- PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs 7 upvotes, #27 of 2026-08-05
- ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads 6 upvotes, #30 of 2026-08-05
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation 5 upvotes, #31 of 2026-08-05
- CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning 4 upvotes, #32 of 2026-08-05
- ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels 4 upvotes, #32 of 2026-08-05
- LegalPincite: Multi-level Legal Information Retrieval Dataset 4 upvotes, #32 of 2026-08-05
- ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning 4 upvotes, #32 of 2026-08-05
- Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents 3 upvotes, #36 of 2026-08-05
- Multi-Task Multi-Frame Visual Piano Transcription 3 upvotes, #36 of 2026-08-05
- When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs 3 upvotes, #36 of 2026-08-05
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.