Hanoona Rasheed
Hanoona Rasheed on Hugging Face Daily Papers: 9 papers, 2 in the top 3 of their day, 191 upvotes.
- Training-Free Speech-Centric Omni Understanding with Frozen VLMs 7 upvotes, #23 of 2026-09-07
- Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding 19 upvotes, #14 of 2026-08-31
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos 24 upvotes, #11 of 2025-06-06
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding 17 upvotes, #12 of 2025-04-18
- Perception Encoder: The best visual embeddings are not at the output of the network 31 upvotes, #5 of 2025-04-18
- PALO: A Polyglot Large Multimodal Model for 5B People 24 upvotes, #3 of 2024-02-23
- PG-Video-LLaVA: Pixel Grounding Large Video-Language Models 18 upvotes, #7 of 2023-11-23
- GLaMM: Pixel Grounding Large Multimodal Model 36 upvotes, #1 of 2023-11-07
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models 7 upvotes, #4 of 2023-06-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.