Daily Papers of 2026-08-04
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks 166 upvotes, #1 of 2026-08-04
- SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks 155 upvotes, #2 of 2026-08-04
- Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs 141 upvotes, #3 of 2026-08-04
- DAPD: Dual-Anchored Policy Distillation 108 upvotes, #4 of 2026-08-04
- InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis 66 upvotes, #5 of 2026-08-04
- Progressive Agent Skill Generation via Reinforcement Learning 58 upvotes, #6 of 2026-08-04
- UEmbed: Unified Sparse and Dense Multimodal Embeddings 50 upvotes, #7 of 2026-08-04
- VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation 46 upvotes, #8 of 2026-08-04
- CADENA: Stepwise CAD Reverse Engineering 38 upvotes, #9 of 2026-08-04
- DiffusionGemma Technical Report 36 upvotes, #10 of 2026-08-04
- WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity 34 upvotes, #11 of 2026-08-04
- SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation 31 upvotes, #12 of 2026-08-04
- WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning 27 upvotes, #13 of 2026-08-04
- SWE-Touch: Benchmarking Coding Agents When Users Touch the Code 24 upvotes, #14 of 2026-08-04
- GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning 24 upvotes, #14 of 2026-08-04
- MemSFT: Mitigating Alignment Tax with an External Parametric Memory 23 upvotes, #16 of 2026-08-04
- Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations 22 upvotes, #17 of 2026-08-04
- GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding 21 upvotes, #18 of 2026-08-04
- To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing 20 upvotes, #19 of 2026-08-04
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? 20 upvotes, #19 of 2026-08-04
- LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation 18 upvotes, #21 of 2026-08-04
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering 18 upvotes, #21 of 2026-08-04
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents 17 upvotes, #23 of 2026-08-04
- Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis 16 upvotes, #24 of 2026-08-04
- DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents 15 upvotes, #25 of 2026-08-04
- StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field 13 upvotes, #26 of 2026-08-04
- RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems 11 upvotes, #27 of 2026-08-04
- Zero-Mem: Zero-Token Memory Operations for LLM Agents 11 upvotes, #27 of 2026-08-04
- ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step 11 upvotes, #27 of 2026-08-04
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures 10 upvotes, #30 of 2026-08-04
- Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis 10 upvotes, #30 of 2026-08-04
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space 10 upvotes, #30 of 2026-08-04
- Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts 8 upvotes, #33 of 2026-08-04
- ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures 6 upvotes, #34 of 2026-08-04
- Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge 6 upvotes, #34 of 2026-08-04
- Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV 4 upvotes, #36 of 2026-08-04
- A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples 4 upvotes, #36 of 2026-08-04
- Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI 4 upvotes, #36 of 2026-08-04
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 3 upvotes, #39 of 2026-08-04
- GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation 2 upvotes, #40 of 2026-08-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.