Daily Papers of 2026-06-24

  1. Qwen-AgentWorld: Language World Models for General Agents 144 upvotes, #1 of 2026-06-24
  2. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? 62 upvotes, #2 of 2026-06-24
  3. Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning 52 upvotes, #3 of 2026-06-24
  4. OpenThoughts-Agent: Data Recipes for Agentic Models 46 upvotes, #4 of 2026-06-24
  5. MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization 43 upvotes, #5 of 2026-06-24
  6. MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management 42 upvotes, #6 of 2026-06-24
  7. AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction 33 upvotes, #7 of 2026-06-24
  8. LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis 23 upvotes, #8 of 2026-06-24
  9. FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation 22 upvotes, #9 of 2026-06-24
  10. Critique of Agent Model 21 upvotes, #10 of 2026-06-24
  11. Semantic Browsing: Controllable Diversity for Image Generation 20 upvotes, #11 of 2026-06-24
  12. FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs 11 upvotes, #12 of 2026-06-24
  13. Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning 11 upvotes, #12 of 2026-06-24
  14. Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning 11 upvotes, #12 of 2026-06-24
  15. DiffusionBench: On Holistic Evaluation of Diffusion Transformers 11 upvotes, #12 of 2026-06-24
  16. World Value Models for Robotic Manipulation 7 upvotes, #16 of 2026-06-24
  17. EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies 6 upvotes, #17 of 2026-06-24
  18. VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct 6 upvotes, #17 of 2026-06-24
  19. DREAM: Dense Retrieval Embeddings via Autoregressive Modeling 6 upvotes, #17 of 2026-06-24
  20. AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning 4 upvotes, #20 of 2026-06-24
  21. QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging 3 upvotes, #21 of 2026-06-24
  22. ChartWalker: Benchmarking the Cross-Chart RAG Task 3 upvotes, #21 of 2026-06-24
  23. ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection 3 upvotes, #21 of 2026-06-24
  24. FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning 1 upvotes, #24 of 2026-06-24
  25. MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery 1 upvotes, #24 of 2026-06-24
  26. FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation 1 upvotes, #24 of 2026-06-24
  27. An Efficient Method for the Optimal Control of Microgrids Under Uncertainties using Local Reduction 0 upvotes, #27 of 2026-06-24
  28. Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation 2 upvotes, #27 of 2026-06-24
  29. InSight: Self-Guided Skill Acquisition via Steerable VLAs 1 upvotes, #27 of 2026-06-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.