Daily Papers of 2026-09-03
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills 540 upvotes, #1 of 2026-09-03
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? 264 upvotes, #2 of 2026-09-03
- Aspire: Can Models Self-Evolve from Vague Goals? 228 upvotes, #3 of 2026-09-03
- SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models 147 upvotes, #4 of 2026-09-03
- EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction 118 upvotes, #5 of 2026-09-03
- It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning 72 upvotes, #6 of 2026-09-03
- Language Models Can Control Their Own Attention 68 upvotes, #7 of 2026-09-03
- S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? 39 upvotes, #8 of 2026-09-03
- WHALE: A Simple Recipe for Joint Harness-Weight Optimization 33 upvotes, #9 of 2026-09-03
- On the Design Fundamentals of Pixel Text Representation Learning 32 upvotes, #10 of 2026-09-03
- Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering 29 upvotes, #11 of 2026-09-03
- NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference 27 upvotes, #12 of 2026-09-03
- ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes 25 upvotes, #13 of 2026-09-03
- Cliff: Learning Process Rewards from the First Mistake 20 upvotes, #14 of 2026-09-03
- A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss 18 upvotes, #15 of 2026-09-03
- VibeVoice-ASR-Streaming Technical Report 17 upvotes, #16 of 2026-09-03
- Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation 16 upvotes, #17 of 2026-09-03
- Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers 12 upvotes, #18 of 2026-09-03
- Post-Training Language Models for Gold-Medal Performance in Coding Competitions 11 upvotes, #19 of 2026-09-03
- MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval 9 upvotes, #20 of 2026-09-03
- ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval 8 upvotes, #21 of 2026-09-03
- CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing 8 upvotes, #21 of 2026-09-03
- PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation 7 upvotes, #23 of 2026-09-03
- SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions 5 upvotes, #24 of 2026-09-03
- Exploring Collaboration between a language and a non-language agent 5 upvotes, #24 of 2026-09-03
- Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents 4 upvotes, #26 of 2026-09-03
- Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models 4 upvotes, #26 of 2026-09-03
- FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos 4 upvotes, #26 of 2026-09-03
- Kirin: Animal Motion Generation from In-the-Wild Video 4 upvotes, #26 of 2026-09-03
- Small Language Models as Judges for Rubric-Based Reinforcement Learning 3 upvotes, #30 of 2026-09-03
- Replacing Training with Memory: Listwise Selection for Text-to-SQL 3 upvotes, #30 of 2026-09-03
- Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens 3 upvotes, #30 of 2026-09-03
- Debias-SparseGPT: Bias-Aware Pruning for Large Language Models 3 upvotes, #30 of 2026-09-03
- Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations 2 upvotes, #34 of 2026-09-03
- Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations 2 upvotes, #34 of 2026-09-03
- An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems 1 upvotes, #36 of 2026-09-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.