Siteng Huang

Siteng Huang on Hugging Face Daily Papers: 22 papers, 7 in the top 3 of their day, 742 upvotes.

  1. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model 197 upvotes, #1 of 2026-07-21
  2. RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation 76 upvotes, #3 of 2026-07-08
  3. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation 92 upvotes, #1 of 2026-07-08
  4. VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon 30 upvotes, #5 of 2026-07-06
  5. MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation 5 upvotes, #24 of 2026-04-02
  6. RynnBrain: Open Embodied Foundation Models 42 upvotes, #3 of 2026-02-19
  7. HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models 11 upvotes, #9 of 2025-12-11
  8. RynnVLA-002: A Unified Vision-Language-Action and World Model 24 upvotes, #5 of 2025-11-24
  9. High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting 11 upvotes, #26 of 2025-10-14
  10. RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation 20 upvotes, #9 of 2025-09-19
  11. VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model 189 upvotes, #1 of 2025-09-12
  12. Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors 10 upvotes, #17 of 2025-08-13
  13. WorldVLA: Towards Autoregressive Action World Model 35 upvotes, #3 of 2025-06-27
  14. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL 5 upvotes, #29 of 2025-05-22
  15. SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning 10 upvotes, #20 of 2025-05-21
  16. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation 8 upvotes, #12 of 2025-05-08
  17. Unicorn: Text-Only Data Synthesis for Vision Language Model Training 37 upvotes, #6 of 2025-04-01
  18. Exploring the Evolution of Physics Cognition in Video Generation: A Survey 11 upvotes, #15 of 2025-03-28
  19. CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction 6 upvotes, #12 of 2024-12-10
  20. Rethinking Token Reduction in MLLMs: Towards a Unified Paradigm for Training-Free Acceleration 18 upvotes, #4 of 2024-11-27
  21. PiTe: Pixel-Temporal Alignment for Large Video-Language Model 11 upvotes, #7 of 2024-09-13
  22. Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference 30 upvotes, #3 of 2024-03-22

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.