liuzuyan
liuzuyan on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 451 upvotes.
- ViQ: Text-Aligned Visual Quantized Representations at Any Resolution 38 upvotes, #6 of 2026-06-26
- GEM: Generative Supervision Helps Embodied Intelligence 41 upvotes, #8 of 2026-05-28
- HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents 182 upvotes, #4 of 2026-04-10
- PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning 18 upvotes, #12 of 2026-04-02
- Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models 12 upvotes, #22 of 2026-03-24
- GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization 89 upvotes, #2 of 2025-11-24
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs 16 upvotes, #17 of 2025-06-06
- Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment 21 upvotes, #7 of 2025-02-07
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 19 upvotes, #7 of 2024-11-22
- Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 22 upvotes, #5 of 2024-09-20
- Efficient Inference of Vision Instruction-Following Models with Elastic Cache 15 upvotes, #8 of 2024-07-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.