Xiangru Tang
Xiangru Tang on Hugging Face Daily Papers: 15 papers, 8 in the top 3 of their day, 654 upvotes.
- EvoClaw: Evaluating AI Agents on Continuous Software Evolution 20 upvotes, #13 of 2026-03-17
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 66 upvotes, #3 of 2025-07-08
- LocAgent: Graph-Guided LLM Agents for Code Localization 6 upvotes, #24 of 2025-03-12
- MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning 14 upvotes, #17 of 2025-03-11
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents 23 upvotes, #2 of 2025-03-05
- MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 79 upvotes, #2 of 2025-01-22
- ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning 8 upvotes, #11 of 2025-01-14
- OpenDevin: An Open Platform for AI Software Developers as Generalist Agents 62 upvotes, #1 of 2024-07-25
- StarCoder 2 and The Stack v2: The Next Generation 160 upvotes, #1 of 2024-03-01
- ChatCell: Facilitating Single-Cell Analysis with Natural Language 13 upvotes, #7 of 2024-02-14
- Weaver: Foundation Models for Creative Writing 46 upvotes, #1 of 2024-01-31
- ML-Bench: Large Language Models Leverage Open-source Libraries for Machine Learning Tasks 10 upvotes, #6 of 2023-11-17
- Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? 11 upvotes, #10 of 2023-09-19
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs 102 upvotes, #1 of 2023-08-01
- RWKV: Reinventing RNNs for the Transformer Era 21 upvotes, #1 of 2023-05-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.