Ziwei Liu
Ziwei Liu on Hugging Face Daily Papers: 65 papers, 19 in the top 3 of their day, 1,817 upvotes.
- EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos 57 upvotes, #18 of 2026-10-02
- Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning 49 upvotes, #13 of 2026-09-29
- Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation 68 upvotes, #7 of 2026-03-18
- HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions 149 upvotes, #3 of 2026-03-17
- ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors 21 upvotes, #6 of 2026-03-05
- EgoTwin: Dreaming Body and View in First Person 18 upvotes, #4 of 2025-08-25
- PhysX: Physical-Grounded 3D Asset Generation 37 upvotes, #2 of 2025-07-17
- 3D Scene Generation: A Survey 16 upvotes, #6 of 2025-05-09
- GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography 21 upvotes, #6 of 2025-04-10
- Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency 20 upvotes, #8 of 2025-03-31
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness 30 upvotes, #5 of 2025-03-28
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey 28 upvotes, #6 of 2025-03-18
- EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
- WHAC: World-grounded Humans and Cameras 2 upvotes, #26 of 2025-02-24
- Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment 21 upvotes, #7 of 2025-02-07
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 22 upvotes, #5 of 2025-01-24
- CityDreamer4D: Compositional Generative Model of Unbounded 4D Cities 19 upvotes, #3 of 2025-01-16
- RepVideo: Rethinking Cross-Layer Representation for Video Generation 15 upvotes, #4 of 2025-01-16
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives 23 upvotes, #4 of 2025-01-10
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control 22 upvotes, #6 of 2025-01-08
- Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models 35 upvotes, #2 of 2024-12-17
- FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion 19 upvotes, #6 of 2024-12-16
- FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 20 upvotes, #7 of 2024-12-11
- Imagine360: Immersive 360 Video Generation from Perspective Anchor 26 upvotes, #4 of 2024-12-05
- SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters 21 upvotes, #6 of 2024-12-03
- Material Anything: Generating Materials for Any 3D Object via Diffusion 40 upvotes, #1 of 2024-11-26
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 19 upvotes, #7 of 2024-11-22
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models 28 upvotes, #2 of 2024-11-21
- FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality 20 upvotes, #4 of 2024-10-28
- DynamicCity: Large-Scale LiDAR Generation from Dynamic Scenes 12 upvotes, #6 of 2024-10-24
- Video Instruction Tuning With Synthetic Data 33 upvotes, #5 of 2024-10-04
- Disco4D: Disentangled 4D Human Generation and Animation from a Single Image 8 upvotes, #9 of 2024-09-27
- 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion 17 upvotes, #8 of 2024-09-20
- Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 22 upvotes, #5 of 2024-09-20
- Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion 22 upvotes, #4 of 2024-09-18
- ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer 8 upvotes, #9 of 2024-08-07
- LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
- VEnhancer: Generative Space-Time Enhancement for Video Generation 8 upvotes, #7 of 2024-07-11
- FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models 10 upvotes, #13 of 2024-06-26
- Long Context Transfer from Language to Vision 30 upvotes, #4 of 2024-06-25
- TC4D: Trajectory-Conditioned Text-to-4D Generation 13 upvotes, #4 of 2024-03-27
- ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars 11 upvotes, #5 of 2024-03-25
- FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation 6 upvotes, #10 of 2024-03-20
- ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance 8 upvotes, #9 of 2024-03-20
- LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation 29 upvotes, #5 of 2024-02-08
- URHand: Universal Relightable Hands 25 upvotes, #3 of 2024-01-11
- GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation 21 upvotes, #5 of 2024-01-09
- DreamGaussian4D: Generative 4D Gaussian Splatting 19 upvotes, #6 of 2023-12-29
- InstructVideo: Instructing Video Diffusion Models with Human Feedback 18 upvotes, #5 of 2023-12-21
- FreeInit: Bridging Initialization Gap in Video Diffusion Models 26 upvotes, #1 of 2023-12-13
- HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image 22 upvotes, #3 of 2023-12-08
- OtterHD: A High-Resolution Multi-modality Model 34 upvotes, #1 of 2023-11-08
- FreeNoise: Tuning-Free Longer Video Diffusion Via Noise Rescheduling 10 upvotes, #4 of 2023-10-24
- Octopus: Embodied Vision-Language Programmer from Environmental Feedback 37 upvotes, #2 of 2023-10-13
- HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion 16 upvotes, #6 of 2023-10-13
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation 48 upvotes, #3 of 2023-09-29
- LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models 43 upvotes, #2 of 2023-09-27
- MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation 9 upvotes, #3 of 2023-09-25
- FreeU: Free Lunch in Diffusion U-Net 66 upvotes, #2 of 2023-09-21
- CityDreamer: Compositional Generative Model of Unbounded 3D Cities 21 upvotes, #5 of 2023-09-04
- Link-Context Learning for Multimodal LLMs 17 upvotes, #4 of 2023-08-16
- DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering 6 upvotes, #9 of 2023-07-20
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation 26 upvotes, #3 of 2023-07-14
- Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation 113 upvotes, #1 of 2023-06-14
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning 12 upvotes, #2 of 2023-06-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.