Jaemin Cho

Jaemin Cho on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 174 upvotes.

  1. Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents 6 upvotes, #23 of 2025-08-12
  2. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality 22 upvotes, #10 of 2025-07-11
  3. Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning 6 upvotes, #25 of 2025-06-05
  4. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance 9 upvotes, #30 of 2025-05-29
  5. CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting 5 upvotes, #19 of 2025-04-23
  6. Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems 13 upvotes, #15 of 2025-04-15
  7. Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization 6 upvotes, #15 of 2025-04-14
  8. VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement 8 upvotes, #11 of 2024-11-25
  9. M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding 25 upvotes, #6 of 2024-11-08
  10. DOCCI: Descriptions of Connected and Contrasting Images 5 upvotes, #13 of 2024-05-01
  11. Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model 16 upvotes, #6 of 2024-04-16
  12. VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning 35 upvotes, #3 of 2023-09-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.