yl-1993
yl-1993 on Hugging Face Daily Papers: 26 papers, 6 in the top 3 of their day, 1,512 upvotes.
- SenseNova-U1.5: Towards Native Unified Visual Intelligence 262 upvotes, #2 of 2026-09-11
- VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 266 upvotes, #1 of 2026-08-27
- Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence 43 upvotes, #9 of 2026-07-21
- Vision as Unified Multimodal Generation 46 upvotes, #6 of 2026-07-08
- From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
- Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer 41 upvotes, #7 of 2026-03-20
- Demystifing Video Reasoning 356 upvotes, #1 of 2026-03-18
- EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents 11 upvotes, #12 of 2026-02-27
- A Very Big Video Reasoning Suite 503 upvotes, #1 of 2026-02-24
- Scaling Spatial Intelligence with Multimodal Foundation Models 41 upvotes, #6 of 2025-11-21
- Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals 21 upvotes, #8 of 2025-11-03
- The Quest for Generalizable Motion Generation: Data, Model, and Evaluation 27 upvotes, #10 of 2025-10-31
- Has GPT-5 Achieved Spatial Intelligence? An Empirical Study 31 upvotes, #8 of 2025-08-19
- DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior 6 upvotes, #23 of 2025-08-07
- TokensGen: Harnessing Condensed Tokens for Long Video Generation 6 upvotes, #21 of 2025-07-22
- EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
- WHAC: World-grounded Humans and Cameras 2 upvotes, #26 of 2025-02-24
- SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters 21 upvotes, #6 of 2024-12-03
- Trajectory Attention for Fine-grained Video Motion Control 12 upvotes, #11 of 2024-12-02
- Disco4D: Disentangled 4D Human Generation and Animation from a Single Image 8 upvotes, #9 of 2024-09-27
- UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model 8 upvotes, #10 of 2024-08-02
- MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers 21 upvotes, #6 of 2024-06-18
- I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models 15 upvotes, #5 of 2024-05-28
- Story-to-Motion: Synthesizing Infinite and Controllable Character Animation from Long Text 29 upvotes, #3 of 2023-11-14
- DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering 6 upvotes, #9 of 2023-07-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.