Zeyu Zhang

Zeyu Zhang on Hugging Face Daily Papers: 66 papers, 3 in the top 3 of their day, 847 upvotes.

  1. DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation 6 upvotes, #75 of 2026-10-02
  2. WorldAttention: An Efficient Attention Architecture for Interactive Video World Models 49 upvotes, #21 of 2026-09-30
  3. DeltaWAM: Delta World Action Models for Bimanual Manipulation 19 upvotes, #13 of 2026-09-25
  4. ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation 1 upvotes, #24 of 2026-09-14
  5. MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control 8 upvotes, #13 of 2026-09-14
  6. ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes 25 upvotes, #13 of 2026-09-03
  7. Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion 33 upvotes, #7 of 2026-08-25
  8. VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery 10 upvotes, #12 of 2026-07-13
  9. SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction 3 upvotes, #10 of 2026-06-22
  10. GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning 3 upvotes, #10 of 2026-06-22
  11. DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects 72 upvotes, #2 of 2026-06-19
  12. MotionVLA: Vision-Language-Action Model for Humanoid Motion 4 upvotes, #25 of 2026-06-17
  13. WorldOlympiad: Can Your World Model Survive a Triathlon? 31 upvotes, #13 of 2026-06-10
  14. Latent Spatial Memory for Video World Models 66 upvotes, #4 of 2026-06-09
  15. PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps 9 upvotes, #24 of 2026-06-03
  16. TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction 51 upvotes, #5 of 2026-05-26
  17. PresentAgent-2: Towards Generalist Multimodal Presentation Agents 8 upvotes, #22 of 2026-05-14
  18. EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory 1 upvotes, #69 of 2026-05-13
  19. Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction 1 upvotes, #56 of 2026-05-13
  20. World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 34 upvotes, #2 of 2026-04-28
  21. UniMesh: Unifying 3D Mesh Understanding and Generation 10 upvotes, #16 of 2026-04-22
  22. HSG: Hyperbolic Scene Graph 1 upvotes, #49 of 2026-04-21
  23. Less Detail, Better Answers: Degradation-Driven Prompting for VQA 13 upvotes, #23 of 2026-04-07
  24. LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models 2 upvotes, #35 of 2026-03-10
  25. MWM: Mobile World Models for Action-Conditioned Consistent Prediction 0 upvotes, #48 of 2026-03-10
  26. GeoWorld: Geometric World Models 7 upvotes, #15 of 2026-02-27
  27. OmniOCR: Generalist OCR for Ethnic Minority Languages 2 upvotes, #25 of 2026-02-25
  28. OCR-Agent: Agentic OCR with Capability and Memory Reflection 2 upvotes, #25 of 2026-02-25
  29. StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation 0 upvotes, #23 of 2026-02-20
  30. MMA: Multimodal Memory Agent 8 upvotes, #14 of 2026-02-19
  31. MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation 3 upvotes, #21 of 2026-02-17
  32. Light4D: Training-Free Extreme Viewpoint 4D Video Relighting 2 upvotes, #23 of 2026-02-16
  33. Code2Worlds: Empowering Coding LLMs for 4D World Generation 4 upvotes, #18 of 2026-02-16
  34. GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning 1 upvotes, #27 of 2026-02-16
  35. V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval 8 upvotes, #29 of 2026-02-06
  36. 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence 1 upvotes, #39 of 2026-01-13
  37. AnyDepth: Depth Estimation Made Easy 9 upvotes, #14 of 2026-01-12
  38. CoV: Chain-of-View Prompting for Spatial Reasoning 8 upvotes, #15 of 2026-01-09
  39. DragMesh: Interactive 3D Generation Made Easy 1 upvotes, #25 of 2025-12-12
  40. EgoLCD: Egocentric Video Generation with Long Context Diffusion 5 upvotes, #29 of 2025-12-05
  41. BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation 3 upvotes, #32 of 2025-12-03
  42. Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation 45 upvotes, #3 of 2025-11-27
  43. MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots 5 upvotes, #12 of 2025-11-27
  44. EvoVLA: Self-Evolving Vision-Language-Action Model 4 upvotes, #24 of 2025-11-25
  45. DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion 1 upvotes, #28 of 2025-10-20
  46. VLA-R1: Enhancing Reasoning in Vision-Language-Action Models 7 upvotes, #28 of 2025-10-03
  47. UniVid: The Open-Source Unified Video Model 3 upvotes, #64 of 2025-09-30
  48. VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction 23 upvotes, #8 of 2025-09-24
  49. VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery 1 upvotes, #33 of 2025-09-23
  50. StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes 1 upvotes, #33 of 2025-09-23
  51. Nav-R1: Reasoning and Navigation in Embodied Scenes 6 upvotes, #10 of 2025-09-16
  52. ReMoMask: Retrieval-Augmented Masked Motion Generation 4 upvotes, #20 of 2025-08-05
  53. 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding 15 upvotes, #4 of 2025-08-04
  54. PresentAgent: Multimodal Agent for Presentation Video Generation 8 upvotes, #18 of 2025-07-08
  55. Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting 11 upvotes, #22 of 2025-06-06
  56. ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS 4 upvotes, #49 of 2025-05-30
  57. Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models 5 upvotes, #29 of 2025-05-22
  58. MediAug: Exploring Visual Augmentation in Medical Imaging 6 upvotes, #10 of 2025-05-02
  59. DiffuMural: Restoring Dunhuang Murals with Multi-scale Diffusion 1 upvotes, #28 of 2025-04-15
  60. 3D CoCa: Contrastive Learners are 3D Captioners 5 upvotes, #23 of 2025-04-15
  61. PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 1 upvotes, #26 of 2025-03-27
  62. Motion Anything: Any to Motion Generation 22 upvotes, #5 of 2025-03-13
  63. DOEI: Dual Optimization of Embedding Information for Attention-Enhanced Class Activation Maps 2 upvotes, #23 of 2025-02-27
  64. Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework 3 upvotes, #26 of 2025-02-21
  65. KMM: Key Frame Mask Mamba for Extended Motion Generation 3 upvotes, #15 of 2024-11-12
  66. Motion Mamba: Efficient and Long Sequence Motion Generation with Hierarchical and Bidirectional Selective SSM 11 upvotes, #5 of 2024-03-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.