Kling Team
Kling Team on Hugging Face Daily Papers: 33 papers, 5 in the top 3 of their day, 3 paper of the day.
- ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing 16 upvotes, #25 of 2026-08-07
- Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations 22 upvotes, #17 of 2026-08-04
- Beacon: Knowing When and How to Perform Agentic Visual Reasoning 54 upvotes, #8 of 2026-07-31
- RefCaptioner: Multi-Reference Image-Grounded Video Captioning 30 upvotes, #14 of 2026-07-31
- KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation 34 upvotes, #8 of 2026-07-17
- UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating 26 upvotes, #9 of 2026-06-25
- OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data 106 upvotes, #1 of 2026-06-15
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization 29 upvotes, #7 of 2026-06-08
- VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization 29 upvotes, #13 of 2026-06-02
- DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory 12 upvotes, #27 of 2026-06-01
- LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
- LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning 46 upvotes, #8 of 2026-05-22
- PureCC: Pure Learning for Text-to-Image Concept Customization 9 upvotes, #21 of 2026-03-10
- Kling-MotionControl Technical Report 26 upvotes, #7 of 2026-03-04
- Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers 27 upvotes, #13 of 2026-02-05
- Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks 46 upvotes, #7 of 2026-02-04
- 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation 55 upvotes, #5 of 2026-02-04
- CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation 30 upvotes, #12 of 2026-01-16
- Klear: Unified Multi-Task Audio-Video Joint Generation 13 upvotes, #6 of 2026-01-08
- GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models 24 upvotes, #9 of 2025-12-30
- SemanticGen: Video Generation in Semantic Space 88 upvotes, #1 of 2025-12-24
- Kling-Omni Technical Report 155 upvotes, #1 of 2025-12-19
- KlingAvatar 2.0 Technical Report 40 upvotes, #8 of 2025-12-16
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder 35 upvotes, #3 of 2025-12-15
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation 16 upvotes, #16 of 2025-12-09
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework 62 upvotes, #3 of 2025-12-03
- Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO 31 upvotes, #7 of 2025-11-21
- Latent Diffusion Model without Variational Autoencoder 47 upvotes, #5 of 2025-10-20
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes 16 upvotes, #21 of 2025-10-14
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution 20 upvotes, #20 of 2025-10-10
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning 60 upvotes, #6 of 2025-10-10
- UniVideo: Unified Understanding, Generation, and Editing for Videos 64 upvotes, #5 of 2025-10-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.