Daily Papers of 2025-11-11
- Grounding Computer Use Agents on Human Demonstrations 98 upvotes, #1 of 2025-11-11
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents 88 upvotes, #2 of 2025-11-11
- IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction 69 upvotes, #3 of 2025-11-11
- DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation 49 upvotes, #4 of 2025-11-11
- The Station: An Open-World Environment for AI-Driven Discovery 34 upvotes, #5 of 2025-11-11
- Robot Learning from a Physical World Model 26 upvotes, #6 of 2025-11-11
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs 24 upvotes, #7 of 2025-11-11
- Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions 22 upvotes, #8 of 2025-11-11
- RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services 18 upvotes, #9 of 2025-11-11
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs 17 upvotes, #10 of 2025-11-11
- Reasoning with Confidence: Efficient Verification of LLM Reasoning Steps via Uncertainty Heads 16 upvotes, #11 of 2025-11-11
- SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization 15 upvotes, #12 of 2025-11-11
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence 15 upvotes, #12 of 2025-11-11
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments 12 upvotes, #14 of 2025-11-11
- NURBGen: High-Fidelity Text-to-CAD Generation through LLM-Driven NURBS Modeling 11 upvotes, #15 of 2025-11-11
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks 10 upvotes, #16 of 2025-11-11
- FLEX: Continuous Agent Evolution via Forward Learning from Experience 9 upvotes, #17 of 2025-11-11
- RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization 7 upvotes, #18 of 2025-11-11
- Reinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMs 7 upvotes, #18 of 2025-11-11
- Long Grounded Thoughts: Distilling Compositional Visual Reasoning Chains at Scale 6 upvotes, #20 of 2025-11-11
- Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models 5 upvotes, #21 of 2025-11-11
- 10 Open Challenges Steering the Future of Vision-Language-Action Models 5 upvotes, #21 of 2025-11-11
- LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs 5 upvotes, #21 of 2025-11-11
- MPJudge: Towards Perceptual Assessment of Music-Induced Paintings 5 upvotes, #21 of 2025-11-11
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents 5 upvotes, #21 of 2025-11-11
- Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries 4 upvotes, #26 of 2025-11-11
- SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads? 4 upvotes, #26 of 2025-11-11
- Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning 4 upvotes, #26 of 2025-11-11
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models 4 upvotes, #26 of 2025-11-11
- DIMO: Diverse 3D Motion Generation for Arbitrary Objects 4 upvotes, #26 of 2025-11-11
- Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models 2 upvotes, #31 of 2025-11-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.