yl-1993

yl-1993 on Hugging Face Daily Papers: 26 papers, 6 in the top 3 of their day, 1,512 upvotes.

  1. SenseNova-U1.5: Towards Native Unified Visual Intelligence 262 upvotes, #2 of 2026-09-11
  2. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 266 upvotes, #1 of 2026-08-27
  3. Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence 43 upvotes, #9 of 2026-07-21
  4. Vision as Unified Multimodal Generation 46 upvotes, #6 of 2026-07-08
  5. From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
  6. SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
  7. Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer 41 upvotes, #7 of 2026-03-20
  8. Demystifing Video Reasoning 356 upvotes, #1 of 2026-03-18
  9. EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents 11 upvotes, #12 of 2026-02-27
  10. A Very Big Video Reasoning Suite 503 upvotes, #1 of 2026-02-24
  11. Scaling Spatial Intelligence with Multimodal Foundation Models 41 upvotes, #6 of 2025-11-21
  12. Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals 21 upvotes, #8 of 2025-11-03
  13. The Quest for Generalizable Motion Generation: Data, Model, and Evaluation 27 upvotes, #10 of 2025-10-31
  14. Has GPT-5 Achieved Spatial Intelligence? An Empirical Study 31 upvotes, #8 of 2025-08-19
  15. DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior 6 upvotes, #23 of 2025-08-07
  16. TokensGen: Harnessing Condensed Tokens for Long Video Generation 6 upvotes, #21 of 2025-07-22
  17. EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
  18. WHAC: World-grounded Humans and Cameras 2 upvotes, #26 of 2025-02-24
  19. SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters 21 upvotes, #6 of 2024-12-03
  20. Trajectory Attention for Fine-grained Video Motion Control 12 upvotes, #11 of 2024-12-02
  21. Disco4D: Disentangled 4D Human Generation and Animation from a Single Image 8 upvotes, #9 of 2024-09-27
  22. UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model 8 upvotes, #10 of 2024-08-02
  23. MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers 21 upvotes, #6 of 2024-06-18
  24. I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models 15 upvotes, #5 of 2024-05-28
  25. Story-to-Motion: Synthesizing Infinite and Controllable Character Animation from Long Text 29 upvotes, #3 of 2023-11-14
  26. DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering 6 upvotes, #9 of 2023-07-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.