Ziwei Liu

Ziwei Liu on Hugging Face Daily Papers: 65 papers, 19 in the top 3 of their day, 1,817 upvotes.

  1. EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos 57 upvotes, #18 of 2026-10-02
  2. Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning 49 upvotes, #13 of 2026-09-29
  3. Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation 68 upvotes, #7 of 2026-03-18
  4. HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions 149 upvotes, #3 of 2026-03-17
  5. ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors 21 upvotes, #6 of 2026-03-05
  6. EgoTwin: Dreaming Body and View in First Person 18 upvotes, #4 of 2025-08-25
  7. PhysX: Physical-Grounded 3D Asset Generation 37 upvotes, #2 of 2025-07-17
  8. 3D Scene Generation: A Survey 16 upvotes, #6 of 2025-05-09
  9. GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography 21 upvotes, #6 of 2025-04-10
  10. Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency 20 upvotes, #8 of 2025-03-31
  11. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness 30 upvotes, #5 of 2025-03-28
  12. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey 28 upvotes, #6 of 2025-03-18
  13. EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
  14. WHAC: World-grounded Humans and Cameras 2 upvotes, #26 of 2025-02-24
  15. Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment 21 upvotes, #7 of 2025-02-07
  16. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 22 upvotes, #5 of 2025-01-24
  17. CityDreamer4D: Compositional Generative Model of Unbounded 4D Cities 19 upvotes, #3 of 2025-01-16
  18. RepVideo: Rethinking Cross-Layer Representation for Video Generation 15 upvotes, #4 of 2025-01-16
  19. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives 23 upvotes, #4 of 2025-01-10
  20. Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control 22 upvotes, #6 of 2025-01-08
  21. Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models 35 upvotes, #2 of 2024-12-17
  22. FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion 19 upvotes, #6 of 2024-12-16
  23. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 20 upvotes, #7 of 2024-12-11
  24. Imagine360: Immersive 360 Video Generation from Perspective Anchor 26 upvotes, #4 of 2024-12-05
  25. SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters 21 upvotes, #6 of 2024-12-03
  26. Material Anything: Generating Materials for Any 3D Object via Diffusion 40 upvotes, #1 of 2024-11-26
  27. Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 19 upvotes, #7 of 2024-11-22
  28. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models 28 upvotes, #2 of 2024-11-21
  29. FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality 20 upvotes, #4 of 2024-10-28
  30. DynamicCity: Large-Scale LiDAR Generation from Dynamic Scenes 12 upvotes, #6 of 2024-10-24
  31. Video Instruction Tuning With Synthetic Data 33 upvotes, #5 of 2024-10-04
  32. Disco4D: Disentangled 4D Human Generation and Animation from a Single Image 8 upvotes, #9 of 2024-09-27
  33. 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion 17 upvotes, #8 of 2024-09-20
  34. Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 22 upvotes, #5 of 2024-09-20
  35. Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion 22 upvotes, #4 of 2024-09-18
  36. ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer 8 upvotes, #9 of 2024-08-07
  37. LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
  38. VEnhancer: Generative Space-Time Enhancement for Video Generation 8 upvotes, #7 of 2024-07-11
  39. FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models 10 upvotes, #13 of 2024-06-26
  40. Long Context Transfer from Language to Vision 30 upvotes, #4 of 2024-06-25
  41. TC4D: Trajectory-Conditioned Text-to-4D Generation 13 upvotes, #4 of 2024-03-27
  42. ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars 11 upvotes, #5 of 2024-03-25
  43. FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation 6 upvotes, #10 of 2024-03-20
  44. ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance 8 upvotes, #9 of 2024-03-20
  45. LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation 29 upvotes, #5 of 2024-02-08
  46. URHand: Universal Relightable Hands 25 upvotes, #3 of 2024-01-11
  47. GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation 21 upvotes, #5 of 2024-01-09
  48. DreamGaussian4D: Generative 4D Gaussian Splatting 19 upvotes, #6 of 2023-12-29
  49. InstructVideo: Instructing Video Diffusion Models with Human Feedback 18 upvotes, #5 of 2023-12-21
  50. FreeInit: Bridging Initialization Gap in Video Diffusion Models 26 upvotes, #1 of 2023-12-13
  51. HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image 22 upvotes, #3 of 2023-12-08
  52. OtterHD: A High-Resolution Multi-modality Model 34 upvotes, #1 of 2023-11-08
  53. FreeNoise: Tuning-Free Longer Video Diffusion Via Noise Rescheduling 10 upvotes, #4 of 2023-10-24
  54. Octopus: Embodied Vision-Language Programmer from Environmental Feedback 37 upvotes, #2 of 2023-10-13
  55. HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion 16 upvotes, #6 of 2023-10-13
  56. DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation 48 upvotes, #3 of 2023-09-29
  57. LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models 43 upvotes, #2 of 2023-09-27
  58. MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation 9 upvotes, #3 of 2023-09-25
  59. FreeU: Free Lunch in Diffusion U-Net 66 upvotes, #2 of 2023-09-21
  60. CityDreamer: Compositional Generative Model of Unbounded 3D Cities 21 upvotes, #5 of 2023-09-04
  61. Link-Context Learning for Multimodal LLMs 17 upvotes, #4 of 2023-08-16
  62. DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering 6 upvotes, #9 of 2023-07-20
  63. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation 26 upvotes, #3 of 2023-07-14
  64. Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation 113 upvotes, #1 of 2023-06-14
  65. MIMIC-IT: Multi-Modal In-Context Instruction Tuning 12 upvotes, #2 of 2023-06-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.