Daily Papers of 2024-11-28

  1. ROICtrl: Boosting Instance Control for Visual Generation 77 upvotes, #1 of 2024-11-28
  2. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models 46 upvotes, #2 of 2024-11-28
  3. Identity-Preserving Text-to-Video Generation by Frequency Decomposition 30 upvotes, #3 of 2024-11-28
  4. Large Language Model-Brained GUI Agents: A Survey 23 upvotes, #4 of 2024-11-28
  5. MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation 22 upvotes, #5 of 2024-11-28
  6. Interleaved Scene Graph for Interleaved Text-and-Image Generation Assessment 20 upvotes, #6 of 2024-11-28
  7. 3D Convex Splatting: Radiance Field Rendering with 3D Smooth Convexes 16 upvotes, #7 of 2024-11-28
  8. Diffusion Self-Distillation for Zero-Shot Customized Image Generation 15 upvotes, #8 of 2024-11-28
  9. DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving 14 upvotes, #9 of 2024-11-28
  10. Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters 12 upvotes, #10 of 2024-11-28
  11. Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient 11 upvotes, #11 of 2024-11-28
  12. UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing 10 upvotes, #12 of 2024-11-28
  13. DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching 10 upvotes, #12 of 2024-11-28
  14. ChatRex: Taming Multimodal LLM for Joint Perception and Understanding 9 upvotes, #14 of 2024-11-28
  15. Video-Guided Foley Sound Generation with Multimodal Controls 7 upvotes, #15 of 2024-11-28
  16. Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis 7 upvotes, #15 of 2024-11-28
  17. Draft Model Knows When to Stop: A Self-Verification Length Policy for Speculative Decoding 6 upvotes, #17 of 2024-11-28
  18. Optimizing Brain Tumor Segmentation with MedNeXt: BraTS 2024 SSA and Pediatrics 5 upvotes, #18 of 2024-11-28
  19. VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format 5 upvotes, #18 of 2024-11-28
  20. Adaptive Blind All-in-One Image Restoration 4 upvotes, #20 of 2024-11-28
  21. Training and Evaluating Language Models with Template-based Data Generation 3 upvotes, #21 of 2024-11-28
  22. Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing 2 upvotes, #22 of 2024-11-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.