Daily Papers of 2026-09-09

  1. NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 419 upvotes, #1 of 2026-09-09
  2. AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing 218 upvotes, #2 of 2026-09-09
  3. Omni Interaction Agent Technical Report 134 upvotes, #3 of 2026-09-09
  4. Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation 94 upvotes, #4 of 2026-09-09
  5. OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining 77 upvotes, #5 of 2026-09-09
  6. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation 70 upvotes, #6 of 2026-09-09
  7. DriveZero: End-to-End Driving Beyond Human Demonstrations 63 upvotes, #7 of 2026-09-09
  8. GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation 59 upvotes, #8 of 2026-09-09
  9. Miles v0.1: Production-Level Post-Training 57 upvotes, #9 of 2026-09-09
  10. Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout 56 upvotes, #10 of 2026-09-09
  11. SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution 47 upvotes, #11 of 2026-09-09
  12. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents 41 upvotes, #12 of 2026-09-09
  13. BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference 40 upvotes, #13 of 2026-09-09
  14. Reason Through the Latent! Making Latent Visual Reasoning Necessary 36 upvotes, #14 of 2026-09-09
  15. Agentic Visual Generation: From Generative Models to Agentic Control 36 upvotes, #14 of 2026-09-09
  16. Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training 35 upvotes, #16 of 2026-09-09
  17. CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements 35 upvotes, #16 of 2026-09-09
  18. VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification 34 upvotes, #18 of 2026-09-09
  19. Steering Geometry: Validating Human Value Geometry in LLM Steering Space 30 upvotes, #19 of 2026-09-09
  20. EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? 29 upvotes, #20 of 2026-09-09
  21. Kalman Delta Networks: Uncertainty-aware Associative Memory 29 upvotes, #20 of 2026-09-09
  22. RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? 28 upvotes, #22 of 2026-09-09
  23. Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy 25 upvotes, #23 of 2026-09-09
  24. CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs 25 upvotes, #23 of 2026-09-09
  25. Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks 25 upvotes, #23 of 2026-09-09
  26. What Did I Just Say? Self-Listening for Full-Duplex Speech Models 24 upvotes, #26 of 2026-09-09
  27. TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model 23 upvotes, #27 of 2026-09-09
  28. What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets 22 upvotes, #28 of 2026-09-09
  29. Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation 22 upvotes, #28 of 2026-09-09
  30. VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes 22 upvotes, #28 of 2026-09-09
  31. SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions 21 upvotes, #31 of 2026-09-09
  32. Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model 21 upvotes, #31 of 2026-09-09
  33. SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation 20 upvotes, #33 of 2026-09-09
  34. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? 19 upvotes, #34 of 2026-09-09
  35. TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation 19 upvotes, #34 of 2026-09-09
  36. MOLE: Detecting Insider Threats in AI Agents 19 upvotes, #34 of 2026-09-09
  37. ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding 19 upvotes, #34 of 2026-09-09
  38. Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection 18 upvotes, #38 of 2026-09-09
  39. A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM 18 upvotes, #38 of 2026-09-09
  40. Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions 17 upvotes, #40 of 2026-09-09
  41. Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise 16 upvotes, #41 of 2026-09-09
  42. RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting 14 upvotes, #42 of 2026-09-09
  43. MasterControl Seventeen Every Time 8 upvotes, #43 of 2026-09-09
  44. Graph Machine: Towards Better Pretraining via Edges 7 upvotes, #44 of 2026-09-09
  45. Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions 7 upvotes, #44 of 2026-09-09
  46. RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives 6 upvotes, #46 of 2026-09-09
  47. Learning 3D Editing without Paired Supervision via Generative Prior Distillation 4 upvotes, #47 of 2026-09-09
  48. NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting 3 upvotes, #48 of 2026-09-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.