alibaba
alibaba on Hugging Face Daily Papers: 15 papers, 3 in the top 3 of their day, 2 paper of the day.
- Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States 83 upvotes, #6 of 2026-10-02
- HappyWorld-Bench 56 upvotes, #4 of 2026-09-24
- One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents 50 upvotes, #10 of 2026-09-22
- DREAM Technical Report 5 upvotes, #19 of 2026-08-26
- Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization 16 upvotes, #13 of 2026-08-25
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation 12 upvotes, #16 of 2026-08-19
- CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing 16 upvotes, #12 of 2026-08-17
- DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation 98 upvotes, #4 of 2026-08-14
- Business Arena: Benchmarking LLM Agents in a Realistic Marketplace 21 upvotes, #14 of 2026-08-11
- MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations 96 upvotes, #1 of 2026-08-05
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks 166 upvotes, #1 of 2026-08-04
- How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs 6 upvotes, #25 of 2026-06-10
- Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics 13 upvotes, #32 of 2026-02-03
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization 54 upvotes, #3 of 2025-10-16
- Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning 65 upvotes, #4 of 2025-10-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.