Kling Team

Kling Team on Hugging Face Daily Papers: 33 papers, 5 in the top 3 of their day, 3 paper of the day.

  1. ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing 16 upvotes, #25 of 2026-08-07
  2. Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations 22 upvotes, #17 of 2026-08-04
  3. Beacon: Knowing When and How to Perform Agentic Visual Reasoning 54 upvotes, #8 of 2026-07-31
  4. RefCaptioner: Multi-Reference Image-Grounded Video Captioning 30 upvotes, #14 of 2026-07-31
  5. KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
  6. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation 34 upvotes, #8 of 2026-07-17
  7. UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating 26 upvotes, #9 of 2026-06-25
  8. OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data 106 upvotes, #1 of 2026-06-15
  9. AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization 29 upvotes, #7 of 2026-06-08
  10. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization 29 upvotes, #13 of 2026-06-02
  11. DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory 12 upvotes, #27 of 2026-06-01
  12. LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
  13. LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning 46 upvotes, #8 of 2026-05-22
  14. PureCC: Pure Learning for Text-to-Image Concept Customization 9 upvotes, #21 of 2026-03-10
  15. Kling-MotionControl Technical Report 26 upvotes, #7 of 2026-03-04
  16. Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers 27 upvotes, #13 of 2026-02-05
  17. Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks 46 upvotes, #7 of 2026-02-04
  18. 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation 55 upvotes, #5 of 2026-02-04
  19. CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation 30 upvotes, #12 of 2026-01-16
  20. Klear: Unified Multi-Task Audio-Video Joint Generation 13 upvotes, #6 of 2026-01-08
  21. GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models 24 upvotes, #9 of 2025-12-30
  22. SemanticGen: Video Generation in Semantic Space 88 upvotes, #1 of 2025-12-24
  23. Kling-Omni Technical Report 155 upvotes, #1 of 2025-12-19
  24. KlingAvatar 2.0 Technical Report 40 upvotes, #8 of 2025-12-16
  25. SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder 35 upvotes, #3 of 2025-12-15
  26. UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation 16 upvotes, #16 of 2025-12-09
  27. MultiShotMaster: A Controllable Multi-Shot Video Generation Framework 62 upvotes, #3 of 2025-12-03
  28. Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO 31 upvotes, #7 of 2025-11-21
  29. Latent Diffusion Model without Variational Autoencoder 47 upvotes, #5 of 2025-10-20
  30. AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes 16 upvotes, #21 of 2025-10-14
  31. UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution 20 upvotes, #20 of 2025-10-10
  32. VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning 60 upvotes, #6 of 2025-10-10
  33. UniVideo: Unified Understanding, Generation, and Editing for Videos 64 upvotes, #5 of 2025-10-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.