Saining Xie
Saining Xie on Hugging Face Daily Papers: 27 papers, 7 in the top 3 of their day, 984 upvotes.
- PaintBench: Deterministic Evaluation of Precise Visual Editing 3 upvotes, #34 of 2026-06-04
- Self-Refining Video Sampling 24 upvotes, #8 of 2026-01-27
- Transition Matching Distillation for Fast Video Generation 31 upvotes, #11 of 2026-01-16
- Cambrian-S: Towards Spatial Supersensing in Video 34 upvotes, #4 of 2025-11-07
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts 7 upvotes, #9 of 2025-11-07
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding 4 upvotes, #11 of 2025-11-07
- Diffusion Transformers with Representation Autoencoders 155 upvotes, #2 of 2025-10-14
- MetaCLIP 2: A Worldwide Scaling Recipe 22 upvotes, #5 of 2025-07-31
- BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing 60 upvotes, #1 of 2025-06-30
- Spatial Mental Modeling from Limited Views 12 upvotes, #10 of 2025-06-30
- LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? 20 upvotes, #5 of 2025-06-16
- Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis 9 upvotes, #16 of 2025-05-16
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 80 upvotes, #1 of 2025-05-15
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers 20 upvotes, #7 of 2025-04-17
- Scaling Language-Free Visual Representation Learning 24 upvotes, #10 of 2025-04-02
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training 100 upvotes, #1 of 2025-01-29
- Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps 65 upvotes, #1 of 2025-01-17
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces 22 upvotes, #5 of 2024-12-19
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark 3 upvotes, #14 of 2024-10-07
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs 48 upvotes, #2 of 2024-06-25
- MoDE: CLIP Data Experts via Clustering 11 upvotes, #3 of 2024-04-25
- V-IRL: Grounding Virtual Intelligence in Real Life 16 upvotes, #11 of 2024-02-06
- Deconstructing Denoising Diffusion Models for Self-Supervised Learning 18 upvotes, #7 of 2024-01-26
- SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers 13 upvotes, #8 of 2024-01-18
- Image Sculpting: Precise Object Editing with 3D Geometry Control 19 upvotes, #4 of 2024-01-04
- Demystifying CLIP Data 20 upvotes, #8 of 2023-09-29
- Going Denser with Open-Vocabulary Part Segmentation 2 upvotes, #12 of 2023-05-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.