Siteng Huang
Siteng Huang on Hugging Face Daily Papers: 22 papers, 7 in the top 3 of their day, 742 upvotes.
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model 197 upvotes, #1 of 2026-07-21
- RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation 76 upvotes, #3 of 2026-07-08
- RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation 92 upvotes, #1 of 2026-07-08
- VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon 30 upvotes, #5 of 2026-07-06
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation 5 upvotes, #24 of 2026-04-02
- RynnBrain: Open Embodied Foundation Models 42 upvotes, #3 of 2026-02-19
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models 11 upvotes, #9 of 2025-12-11
- RynnVLA-002: A Unified Vision-Language-Action and World Model 24 upvotes, #5 of 2025-11-24
- High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting 11 upvotes, #26 of 2025-10-14
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation 20 upvotes, #9 of 2025-09-19
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model 189 upvotes, #1 of 2025-09-12
- Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors 10 upvotes, #17 of 2025-08-13
- WorldVLA: Towards Autoregressive Action World Model 35 upvotes, #3 of 2025-06-27
- VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL 5 upvotes, #29 of 2025-05-22
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning 10 upvotes, #20 of 2025-05-21
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation 8 upvotes, #12 of 2025-05-08
- Unicorn: Text-Only Data Synthesis for Vision Language Model Training 37 upvotes, #6 of 2025-04-01
- Exploring the Evolution of Physics Cognition in Video Generation: A Survey 11 upvotes, #15 of 2025-03-28
- CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction 6 upvotes, #12 of 2024-12-10
- Rethinking Token Reduction in MLLMs: Towards a Unified Paradigm for Training-Free Acceleration 18 upvotes, #4 of 2024-11-27
- PiTe: Pixel-Temporal Alignment for Large Video-Language Model 11 upvotes, #7 of 2024-09-13
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference 30 upvotes, #3 of 2024-03-22
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.