Xiangru Tang

Xiangru Tang on Hugging Face Daily Papers: 15 papers, 8 in the top 3 of their day, 654 upvotes.

  1. EvoClaw: Evaluating AI Agents on Continuous Software Evolution 20 upvotes, #13 of 2026-03-17
  2. Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 66 upvotes, #3 of 2025-07-08
  3. LocAgent: Graph-Guided LLM Agents for Code Localization 6 upvotes, #24 of 2025-03-12
  4. MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning 14 upvotes, #17 of 2025-03-11
  5. MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents 23 upvotes, #2 of 2025-03-05
  6. MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 79 upvotes, #2 of 2025-01-22
  7. ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning 8 upvotes, #11 of 2025-01-14
  8. OpenDevin: An Open Platform for AI Software Developers as Generalist Agents 62 upvotes, #1 of 2024-07-25
  9. StarCoder 2 and The Stack v2: The Next Generation 160 upvotes, #1 of 2024-03-01
  10. ChatCell: Facilitating Single-Cell Analysis with Natural Language 13 upvotes, #7 of 2024-02-14
  11. Weaver: Foundation Models for Creative Writing 46 upvotes, #1 of 2024-01-31
  12. ML-Bench: Large Language Models Leverage Open-source Libraries for Machine Learning Tasks 10 upvotes, #6 of 2023-11-17
  13. Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? 11 upvotes, #10 of 2023-09-19
  14. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs 102 upvotes, #1 of 2023-08-01
  15. RWKV: Reinventing RNNs for the Transformer Era 21 upvotes, #1 of 2023-05-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.