Ran Xu
Ran Xu on Hugging Face Daily Papers: 12 papers, 4 in the top 3 of their day, 427 upvotes.
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 80 upvotes, #1 of 2025-05-15
- DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs 13 upvotes, #11 of 2025-04-25
- BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions 21 upvotes, #4 of 2024-11-13
- xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs 14 upvotes, #4 of 2024-10-23
- xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations 32 upvotes, #5 of 2024-08-23
- xGen-MM (BLIP-3): A Family of Open Large Multimodal Models 91 upvotes, #1 of 2024-08-19
- MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 10 upvotes, #18 of 2024-06-18
- TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
- BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents 20 upvotes, #4 of 2023-08-14
- Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization 21 upvotes, #1 of 2023-08-07
- UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild 4 upvotes, #6 of 2023-05-19
- ULIP-2: Towards Scalable Multimodal Pre-training For 3D Understanding 2 upvotes, #8 of 2023-05-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.