Xing Han Lù

Xing Han Lù on Hugging Face Daily Papers: 12 papers, 3 in the top 3 of their day, 422 upvotes.

  1. OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks 21 upvotes, #14 of 2026-06-30
  2. Structured Distillation of Web Agent Capabilities Enables Generalization 20 upvotes, #18 of 2026-04-10
  3. Grounding Computer Use Agents on Human Demonstrations 98 upvotes, #1 of 2025-11-11
  4. FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents 3 upvotes, #22 of 2025-10-06
  5. Build the web for agents, not agents for the web 19 upvotes, #15 of 2025-06-13
  6. AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories 27 upvotes, #8 of 2025-04-15
  7. DeepSeek-R1 Thoughtology: Let's <think> about LLM Reasoning 79 upvotes, #2 of 2025-04-11
  8. SafeArena: Evaluating the Safety of Autonomous Web Agents 18 upvotes, #11 of 2025-03-10
  9. MMTEB: Massive Multilingual Text Embedding Benchmark 31 upvotes, #5 of 2025-02-20
  10. The BrowserGym Ecosystem for Web Agent Research 18 upvotes, #6 of 2024-12-12
  11. BM25S: Orders of magnitude faster lexical search via eager sparse scoring 9 upvotes, #9 of 2024-07-10
  12. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue 39 upvotes, #2 of 2024-02-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.