Zhuoran Zhang
Zhuoran Zhang on Hugging Face Daily Papers: 10 papers, 0 in the top 3 of their day, 336 upvotes.
- FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory 14 upvotes, #18 of 2026-08-06
- OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models 26 upvotes, #11 of 2026-08-05
- Beacon: Knowing When and How to Perform Agentic Visual Reasoning 54 upvotes, #8 of 2026-07-31
- RefCaptioner: Multi-Reference Image-Grounded Video Captioning 30 upvotes, #14 of 2026-07-31
- KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
- LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
- Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos 22 upvotes, #12 of 2026-05-20
- When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs 24 upvotes, #5 of 2025-11-05
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.