Yu Su
Yu Su on Hugging Face Daily Papers: 24 papers, 9 in the top 3 of their day, 1,089 upvotes.
- Automatic Image-Level Morphological Trait Annotation for Organismal Images 5 upvotes, #41 of 2026-04-03
- Agent Learning via Early Experience 223 upvotes, #1 of 2025-10-10
- Watch and Learn: Learning to Use Computers from Online Videos 9 upvotes, #18 of 2025-10-07
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge 45 upvotes, #1 of 2025-06-27
- Is Extending Modality The Right Path Towards Omni-Modality? 21 upvotes, #8 of 2025-06-09
- ARM: Adaptive Reasoning Model 43 upvotes, #7 of 2025-05-27
- SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills 11 upvotes, #10 of 2025-04-09
- Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems 230 upvotes, #1 of 2025-04-04
- From RAG to Memory: Non-Parametric Continual Learning for Large Language Models 11 upvotes, #16 of 2025-02-21
- On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective 44 upvotes, #2 of 2025-02-20
- Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents 9 upvotes, #18 of 2025-02-18
- Sparse Autoencoders for Scientifically Rigorous Interpretation of Vision Models 6 upvotes, #22 of 2025-02-12
- Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents 11 upvotes, #6 of 2024-11-21
- Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents 16 upvotes, #7 of 2024-10-08
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery 18 upvotes, #5 of 2024-10-08
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
- VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images 7 upvotes, #8 of 2024-09-02
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents 12 upvotes, #7 of 2024-08-13
- Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization 30 upvotes, #3 of 2024-05-27
- LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error 19 upvotes, #5 of 2024-03-08
- TravelPlanner: A Benchmark for Real-World Planning with Language Agents 38 upvotes, #3 of 2024-02-05
- GPT-4V(ision) is a Generalist Web Agent, if Grounded 23 upvotes, #3 of 2024-01-04
- AgentBench: Evaluating LLMs as Agents 26 upvotes, #2 of 2023-08-08
- MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing 37 upvotes, #1 of 2023-06-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.