Daily Papers of 2026-08-25
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work 204 upvotes, #1 of 2026-08-25
- EchoWM: Open and Enterable Omnimodal World Models 79 upvotes, #2 of 2026-08-25
- TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming 58 upvotes, #3 of 2026-08-25
- Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision 51 upvotes, #4 of 2026-08-25
- Prime Agent: A Self-Improving RLM Harness 47 upvotes, #5 of 2026-08-25
- MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks 41 upvotes, #6 of 2026-08-25
- Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion 33 upvotes, #7 of 2026-08-25
- The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models 31 upvotes, #8 of 2026-08-25
- Towards a Densing Law for User Representation Learning at Billion-Scale Capacity 28 upvotes, #9 of 2026-08-25
- RISE: Adaptive Imagination for World Action Models 26 upvotes, #10 of 2026-08-25
- ReWorld: An Interactive World Model with Long-Horizon Memory 24 upvotes, #11 of 2026-08-25
- GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? 17 upvotes, #12 of 2026-08-25
- ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction 16 upvotes, #13 of 2026-08-25
- Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization 16 upvotes, #13 of 2026-08-25
- LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures 15 upvotes, #15 of 2026-08-25
- One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows 12 upvotes, #16 of 2026-08-25
- Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection 11 upvotes, #17 of 2026-08-25
- Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs 11 upvotes, #17 of 2026-08-25
- AutoResearch: Insight In, Hallucination Out 10 upvotes, #19 of 2026-08-25
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress 8 upvotes, #20 of 2026-08-25
- Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors 7 upvotes, #21 of 2026-08-25
- LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks 7 upvotes, #21 of 2026-08-25
- One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders 5 upvotes, #23 of 2026-08-25
- TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration 5 upvotes, #23 of 2026-08-25
- ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts 5 upvotes, #23 of 2026-08-25
- What AstroPT knows about galaxies, and what that can teach us about LLMs 5 upvotes, #23 of 2026-08-25
- Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports 5 upvotes, #23 of 2026-08-25
- RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling 4 upvotes, #28 of 2026-08-25
- Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA 4 upvotes, #28 of 2026-08-25
- From Generation to Simulation: How Far Are World Models from Being True Simulators? 4 upvotes, #28 of 2026-08-25
- The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search 4 upvotes, #28 of 2026-08-25
- Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning 3 upvotes, #32 of 2026-08-25
- Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection 3 upvotes, #32 of 2026-08-25
- EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment 1 upvotes, #34 of 2026-08-25
- WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning 1 upvotes, #34 of 2026-08-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.