Shilong Liu

Shilong Liu on Hugging Face Daily Papers: 17 papers, 5 in the top 3 of their day, 572 upvotes.

  1. EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents 18 upvotes, #16 of 2026-06-10
  2. LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding 135 upvotes, #1 of 2026-05-27
  3. MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory 60 upvotes, #6 of 2026-05-15
  4. Web World Models 22 upvotes, #10 of 2025-12-30
  5. UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers 15 upvotes, #13 of 2025-12-05
  6. A Survey of Self-Evolving Agents: On Path to Artificial Super Intelligence 72 upvotes, #2 of 2025-07-29
  7. Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution 6 upvotes, #40 of 2025-05-28
  8. TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video 19 upvotes, #8 of 2024-12-03
  9. DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding 10 upvotes, #11 of 2024-11-22
  10. TAPTRv2: Attention-based Position Update Improves Tracking Any Point 10 upvotes, #20 of 2024-07-30
  11. Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection 21 upvotes, #5 of 2024-05-17
  12. Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
  13. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
  14. Visual In-Context Prompting 18 upvotes, #7 of 2023-11-23
  15. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 50 upvotes, #2 of 2023-11-10
  16. Semantic-SAM: Segment and Recognize Anything at Any Granularity 23 upvotes, #3 of 2023-07-11
  17. Recognize Anything: A Strong Image Tagging Model 12 upvotes, #1 of 2023-06-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.