Daily Papers of 2026-08-04

  1. LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks 166 upvotes, #1 of 2026-08-04
  2. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks 155 upvotes, #2 of 2026-08-04
  3. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs 141 upvotes, #3 of 2026-08-04
  4. DAPD: Dual-Anchored Policy Distillation 108 upvotes, #4 of 2026-08-04
  5. InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis 66 upvotes, #5 of 2026-08-04
  6. Progressive Agent Skill Generation via Reinforcement Learning 58 upvotes, #6 of 2026-08-04
  7. UEmbed: Unified Sparse and Dense Multimodal Embeddings 50 upvotes, #7 of 2026-08-04
  8. VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation 46 upvotes, #8 of 2026-08-04
  9. CADENA: Stepwise CAD Reverse Engineering 38 upvotes, #9 of 2026-08-04
  10. DiffusionGemma Technical Report 36 upvotes, #10 of 2026-08-04
  11. WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity 34 upvotes, #11 of 2026-08-04
  12. SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation 31 upvotes, #12 of 2026-08-04
  13. WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning 27 upvotes, #13 of 2026-08-04
  14. SWE-Touch: Benchmarking Coding Agents When Users Touch the Code 24 upvotes, #14 of 2026-08-04
  15. GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning 24 upvotes, #14 of 2026-08-04
  16. MemSFT: Mitigating Alignment Tax with an External Parametric Memory 23 upvotes, #16 of 2026-08-04
  17. Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations 22 upvotes, #17 of 2026-08-04
  18. GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding 21 upvotes, #18 of 2026-08-04
  19. To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing 20 upvotes, #19 of 2026-08-04
  20. AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? 20 upvotes, #19 of 2026-08-04
  21. LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation 18 upvotes, #21 of 2026-08-04
  22. 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering 18 upvotes, #21 of 2026-08-04
  23. DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents 17 upvotes, #23 of 2026-08-04
  24. Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis 16 upvotes, #24 of 2026-08-04
  25. DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents 15 upvotes, #25 of 2026-08-04
  26. StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field 13 upvotes, #26 of 2026-08-04
  27. RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems 11 upvotes, #27 of 2026-08-04
  28. Zero-Mem: Zero-Token Memory Operations for LLM Agents 11 upvotes, #27 of 2026-08-04
  29. ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step 11 upvotes, #27 of 2026-08-04
  30. Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures 10 upvotes, #30 of 2026-08-04
  31. Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis 10 upvotes, #30 of 2026-08-04
  32. SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space 10 upvotes, #30 of 2026-08-04
  33. Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts 8 upvotes, #33 of 2026-08-04
  34. ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures 6 upvotes, #34 of 2026-08-04
  35. Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge 6 upvotes, #34 of 2026-08-04
  36. Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV 4 upvotes, #36 of 2026-08-04
  37. A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples 4 upvotes, #36 of 2026-08-04
  38. Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI 4 upvotes, #36 of 2026-08-04
  39. Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 3 upvotes, #39 of 2026-08-04
  40. GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation 2 upvotes, #40 of 2026-08-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.