Daily Papers of 2025-11-11

  1. Grounding Computer Use Agents on Human Demonstrations 98 upvotes, #1 of 2025-11-11
  2. HaluMem: Evaluating Hallucinations in Memory Systems of Agents 88 upvotes, #2 of 2025-11-11
  3. IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction 69 upvotes, #3 of 2025-11-11
  4. DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation 49 upvotes, #4 of 2025-11-11
  5. The Station: An Open-World Environment for AI-Driven Discovery 34 upvotes, #5 of 2025-11-11
  6. Robot Learning from a Physical World Model 26 upvotes, #6 of 2025-11-11
  7. Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs 24 upvotes, #7 of 2025-11-11
  8. Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions 22 upvotes, #8 of 2025-11-11
  9. RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services 18 upvotes, #9 of 2025-11-11
  10. MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs 17 upvotes, #10 of 2025-11-11
  11. Reasoning with Confidence: Efficient Verification of LLM Reasoning Steps via Uncertainty Heads 16 upvotes, #11 of 2025-11-11
  12. SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization 15 upvotes, #12 of 2025-11-11
  13. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence 15 upvotes, #12 of 2025-11-11
  14. RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments 12 upvotes, #14 of 2025-11-11
  15. NURBGen: High-Fidelity Text-to-CAD Generation through LLM-Driven NURBS Modeling 11 upvotes, #15 of 2025-11-11
  16. Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks 10 upvotes, #16 of 2025-11-11
  17. FLEX: Continuous Agent Evolution via Forward Learning from Experience 9 upvotes, #17 of 2025-11-11
  18. RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization 7 upvotes, #18 of 2025-11-11
  19. Reinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMs 7 upvotes, #18 of 2025-11-11
  20. Long Grounded Thoughts: Distilling Compositional Visual Reasoning Chains at Scale 6 upvotes, #20 of 2025-11-11
  21. Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models 5 upvotes, #21 of 2025-11-11
  22. 10 Open Challenges Steering the Future of Vision-Language-Action Models 5 upvotes, #21 of 2025-11-11
  23. LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs 5 upvotes, #21 of 2025-11-11
  24. MPJudge: Towards Perceptual Assessment of Music-Induced Paintings 5 upvotes, #21 of 2025-11-11
  25. DigiData: Training and Evaluating General-Purpose Mobile Control Agents 5 upvotes, #21 of 2025-11-11
  26. Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries 4 upvotes, #26 of 2025-11-11
  27. SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads? 4 upvotes, #26 of 2025-11-11
  28. Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning 4 upvotes, #26 of 2025-11-11
  29. VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models 4 upvotes, #26 of 2025-11-11
  30. DIMO: Diverse 3D Motion Generation for Arbitrary Objects 4 upvotes, #26 of 2025-11-11
  31. Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models 2 upvotes, #31 of 2025-11-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.