Daily Papers of 2026-09-03

  1. Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills 540 upvotes, #1 of 2026-09-03
  2. HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? 264 upvotes, #2 of 2026-09-03
  3. Aspire: Can Models Self-Evolve from Vague Goals? 228 upvotes, #3 of 2026-09-03
  4. SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models 147 upvotes, #4 of 2026-09-03
  5. EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction 118 upvotes, #5 of 2026-09-03
  6. It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning 72 upvotes, #6 of 2026-09-03
  7. Language Models Can Control Their Own Attention 68 upvotes, #7 of 2026-09-03
  8. S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? 39 upvotes, #8 of 2026-09-03
  9. WHALE: A Simple Recipe for Joint Harness-Weight Optimization 33 upvotes, #9 of 2026-09-03
  10. On the Design Fundamentals of Pixel Text Representation Learning 32 upvotes, #10 of 2026-09-03
  11. Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering 29 upvotes, #11 of 2026-09-03
  12. NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference 27 upvotes, #12 of 2026-09-03
  13. ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes 25 upvotes, #13 of 2026-09-03
  14. Cliff: Learning Process Rewards from the First Mistake 20 upvotes, #14 of 2026-09-03
  15. A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss 18 upvotes, #15 of 2026-09-03
  16. VibeVoice-ASR-Streaming Technical Report 17 upvotes, #16 of 2026-09-03
  17. Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation 16 upvotes, #17 of 2026-09-03
  18. Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers 12 upvotes, #18 of 2026-09-03
  19. Post-Training Language Models for Gold-Medal Performance in Coding Competitions 11 upvotes, #19 of 2026-09-03
  20. MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval 9 upvotes, #20 of 2026-09-03
  21. ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval 8 upvotes, #21 of 2026-09-03
  22. CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing 8 upvotes, #21 of 2026-09-03
  23. PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation 7 upvotes, #23 of 2026-09-03
  24. SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions 5 upvotes, #24 of 2026-09-03
  25. Exploring Collaboration between a language and a non-language agent 5 upvotes, #24 of 2026-09-03
  26. Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents 4 upvotes, #26 of 2026-09-03
  27. Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models 4 upvotes, #26 of 2026-09-03
  28. FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos 4 upvotes, #26 of 2026-09-03
  29. Kirin: Animal Motion Generation from In-the-Wild Video 4 upvotes, #26 of 2026-09-03
  30. Small Language Models as Judges for Rubric-Based Reinforcement Learning 3 upvotes, #30 of 2026-09-03
  31. Replacing Training with Memory: Listwise Selection for Text-to-SQL 3 upvotes, #30 of 2026-09-03
  32. Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens 3 upvotes, #30 of 2026-09-03
  33. Debias-SparseGPT: Bias-Aware Pruning for Large Language Models 3 upvotes, #30 of 2026-09-03
  34. Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations 2 upvotes, #34 of 2026-09-03
  35. Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations 2 upvotes, #34 of 2026-09-03
  36. An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems 1 upvotes, #36 of 2026-09-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.