Daily Papers of 2026-10-07

  1. DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation 21 upvotes, #1 of 2026-10-07
  2. AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents 15 upvotes, #2 of 2026-10-07
  3. HuatuoGPT-3: RL-Only Domain Adaptation from Base Models 14 upvotes, #3 of 2026-10-07
  4. Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents 10 upvotes, #4 of 2026-10-07
  5. World Action Learning via Interaction-Centric Spectral Latent Guidance 10 upvotes, #4 of 2026-10-07
  6. EVISKILL: Grounding Skill Evolution in Replayable Evidence 10 upvotes, #4 of 2026-10-07
  7. DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling 7 upvotes, #7 of 2026-10-07
  8. Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment 6 upvotes, #8 of 2026-10-07
  9. GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution 5 upvotes, #9 of 2026-10-07
  10. TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models 4 upvotes, #10 of 2026-10-07
  11. ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing 3 upvotes, #11 of 2026-10-07
  12. Personal-Agent Mediated Recommendation with Cross-Platform User History 2 upvotes, #12 of 2026-10-07
  13. Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution 2 upvotes, #12 of 2026-10-07
  14. EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation 2 upvotes, #12 of 2026-10-07
  15. Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions 1 upvotes, #15 of 2026-10-07
  16. DistScene: Object-to-Scene Distillation for 3D Scene Generation 1 upvotes, #15 of 2026-10-07
  17. HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models 1 upvotes, #15 of 2026-10-07
  18. Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents 1 upvotes, #15 of 2026-10-07
  19. Harness Engineering for Software Engineering via Modular Executable Dev-Primitives 1 upvotes, #15 of 2026-10-07
  20. VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning 1 upvotes, #15 of 2026-10-07
  21. Building Rome from a Single Image 1 upvotes, #15 of 2026-10-07
  22. World Models' Last Exam in Physics 1 upvotes, #15 of 2026-10-07
  23. Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training 0 upvotes, #23 of 2026-10-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.