Daily Papers of 2024-11-28
- ROICtrl: Boosting Instance Control for Visual Generation 77 upvotes, #1 of 2024-11-28
- CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models 46 upvotes, #2 of 2024-11-28
- Identity-Preserving Text-to-Video Generation by Frequency Decomposition 30 upvotes, #3 of 2024-11-28
- Large Language Model-Brained GUI Agents: A Survey 23 upvotes, #4 of 2024-11-28
- MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation 22 upvotes, #5 of 2024-11-28
- Interleaved Scene Graph for Interleaved Text-and-Image Generation Assessment 20 upvotes, #6 of 2024-11-28
- 3D Convex Splatting: Radiance Field Rendering with 3D Smooth Convexes 16 upvotes, #7 of 2024-11-28
- Diffusion Self-Distillation for Zero-Shot Customized Image Generation 15 upvotes, #8 of 2024-11-28
- DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving 14 upvotes, #9 of 2024-11-28
- Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters 12 upvotes, #10 of 2024-11-28
- Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient 11 upvotes, #11 of 2024-11-28
- UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing 10 upvotes, #12 of 2024-11-28
- DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching 10 upvotes, #12 of 2024-11-28
- ChatRex: Taming Multimodal LLM for Joint Perception and Understanding 9 upvotes, #14 of 2024-11-28
- Video-Guided Foley Sound Generation with Multimodal Controls 7 upvotes, #15 of 2024-11-28
- Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis 7 upvotes, #15 of 2024-11-28
- Draft Model Knows When to Stop: A Self-Verification Length Policy for Speculative Decoding 6 upvotes, #17 of 2024-11-28
- Optimizing Brain Tumor Segmentation with MedNeXt: BraTS 2024 SSA and Pediatrics 5 upvotes, #18 of 2024-11-28
- VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format 5 upvotes, #18 of 2024-11-28
- Adaptive Blind All-in-One Image Restoration 4 upvotes, #20 of 2024-11-28
- Training and Evaluating Language Models with Template-based Data Generation 3 upvotes, #21 of 2024-11-28
- Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing 2 upvotes, #22 of 2024-11-28
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.