Zeyu Zhang
Zeyu Zhang on Hugging Face Daily Papers: 66 papers, 3 in the top 3 of their day, 847 upvotes.
- DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation 6 upvotes, #75 of 2026-10-02
- WorldAttention: An Efficient Attention Architecture for Interactive Video World Models 49 upvotes, #21 of 2026-09-30
- DeltaWAM: Delta World Action Models for Bimanual Manipulation 19 upvotes, #13 of 2026-09-25
- ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation 1 upvotes, #24 of 2026-09-14
- MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control 8 upvotes, #13 of 2026-09-14
- ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes 25 upvotes, #13 of 2026-09-03
- Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion 33 upvotes, #7 of 2026-08-25
- VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery 10 upvotes, #12 of 2026-07-13
- SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction 3 upvotes, #10 of 2026-06-22
- GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning 3 upvotes, #10 of 2026-06-22
- DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects 72 upvotes, #2 of 2026-06-19
- MotionVLA: Vision-Language-Action Model for Humanoid Motion 4 upvotes, #25 of 2026-06-17
- WorldOlympiad: Can Your World Model Survive a Triathlon? 31 upvotes, #13 of 2026-06-10
- Latent Spatial Memory for Video World Models 66 upvotes, #4 of 2026-06-09
- PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps 9 upvotes, #24 of 2026-06-03
- TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction 51 upvotes, #5 of 2026-05-26
- PresentAgent-2: Towards Generalist Multimodal Presentation Agents 8 upvotes, #22 of 2026-05-14
- EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory 1 upvotes, #69 of 2026-05-13
- Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction 1 upvotes, #56 of 2026-05-13
- World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 34 upvotes, #2 of 2026-04-28
- UniMesh: Unifying 3D Mesh Understanding and Generation 10 upvotes, #16 of 2026-04-22
- HSG: Hyperbolic Scene Graph 1 upvotes, #49 of 2026-04-21
- Less Detail, Better Answers: Degradation-Driven Prompting for VQA 13 upvotes, #23 of 2026-04-07
- LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models 2 upvotes, #35 of 2026-03-10
- MWM: Mobile World Models for Action-Conditioned Consistent Prediction 0 upvotes, #48 of 2026-03-10
- GeoWorld: Geometric World Models 7 upvotes, #15 of 2026-02-27
- OmniOCR: Generalist OCR for Ethnic Minority Languages 2 upvotes, #25 of 2026-02-25
- OCR-Agent: Agentic OCR with Capability and Memory Reflection 2 upvotes, #25 of 2026-02-25
- StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation 0 upvotes, #23 of 2026-02-20
- MMA: Multimodal Memory Agent 8 upvotes, #14 of 2026-02-19
- MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation 3 upvotes, #21 of 2026-02-17
- Light4D: Training-Free Extreme Viewpoint 4D Video Relighting 2 upvotes, #23 of 2026-02-16
- Code2Worlds: Empowering Coding LLMs for 4D World Generation 4 upvotes, #18 of 2026-02-16
- GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning 1 upvotes, #27 of 2026-02-16
- V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval 8 upvotes, #29 of 2026-02-06
- 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence 1 upvotes, #39 of 2026-01-13
- AnyDepth: Depth Estimation Made Easy 9 upvotes, #14 of 2026-01-12
- CoV: Chain-of-View Prompting for Spatial Reasoning 8 upvotes, #15 of 2026-01-09
- DragMesh: Interactive 3D Generation Made Easy 1 upvotes, #25 of 2025-12-12
- EgoLCD: Egocentric Video Generation with Long Context Diffusion 5 upvotes, #29 of 2025-12-05
- BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation 3 upvotes, #32 of 2025-12-03
- Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation 45 upvotes, #3 of 2025-11-27
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots 5 upvotes, #12 of 2025-11-27
- EvoVLA: Self-Evolving Vision-Language-Action Model 4 upvotes, #24 of 2025-11-25
- DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion 1 upvotes, #28 of 2025-10-20
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models 7 upvotes, #28 of 2025-10-03
- UniVid: The Open-Source Unified Video Model 3 upvotes, #64 of 2025-09-30
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction 23 upvotes, #8 of 2025-09-24
- VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery 1 upvotes, #33 of 2025-09-23
- StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes 1 upvotes, #33 of 2025-09-23
- Nav-R1: Reasoning and Navigation in Embodied Scenes 6 upvotes, #10 of 2025-09-16
- ReMoMask: Retrieval-Augmented Masked Motion Generation 4 upvotes, #20 of 2025-08-05
- 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding 15 upvotes, #4 of 2025-08-04
- PresentAgent: Multimodal Agent for Presentation Video Generation 8 upvotes, #18 of 2025-07-08
- Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting 11 upvotes, #22 of 2025-06-06
- ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS 4 upvotes, #49 of 2025-05-30
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models 5 upvotes, #29 of 2025-05-22
- MediAug: Exploring Visual Augmentation in Medical Imaging 6 upvotes, #10 of 2025-05-02
- DiffuMural: Restoring Dunhuang Murals with Multi-scale Diffusion 1 upvotes, #28 of 2025-04-15
- 3D CoCa: Contrastive Learners are 3D Captioners 5 upvotes, #23 of 2025-04-15
- PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 1 upvotes, #26 of 2025-03-27
- Motion Anything: Any to Motion Generation 22 upvotes, #5 of 2025-03-13
- DOEI: Dual Optimization of Embedding Information for Attention-Enhanced Class Activation Maps 2 upvotes, #23 of 2025-02-27
- Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework 3 upvotes, #26 of 2025-02-21
- KMM: Key Frame Mask Mamba for Extended Motion Generation 3 upvotes, #15 of 2024-11-12
- Motion Mamba: Efficient and Long Sequence Motion Generation with Hierarchical and Bidirectional Selective SSM 11 upvotes, #5 of 2024-03-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.