Jaemin Cho
Jaemin Cho on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 174 upvotes.
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents 6 upvotes, #23 of 2025-08-12
- A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality 22 upvotes, #10 of 2025-07-11
- Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning 6 upvotes, #25 of 2025-06-05
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance 9 upvotes, #30 of 2025-05-29
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting 5 upvotes, #19 of 2025-04-23
- Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems 13 upvotes, #15 of 2025-04-15
- Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization 6 upvotes, #15 of 2025-04-14
- VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement 8 upvotes, #11 of 2024-11-25
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding 25 upvotes, #6 of 2024-11-08
- DOCCI: Descriptions of Connected and Contrasting Images 5 upvotes, #13 of 2024-05-01
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model 16 upvotes, #6 of 2024-04-16
- VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning 35 upvotes, #3 of 2023-09-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.