Daily Papers of 2023-12-06

  1. ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation 33 upvotes, #1 of 2023-12-06
  2. FaceStudio: Put Your Face Everywhere in Seconds 32 upvotes, #2 of 2023-12-06
  3. Analyzing and Improving the Training Dynamics of Diffusion Models 32 upvotes, #2 of 2023-12-06
  4. Relightable Gaussian Codec Avatars 31 upvotes, #4 of 2023-12-06
  5. X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model 26 upvotes, #5 of 2023-12-06
  6. OneLLM: One Framework to Align All Modalities with Language 23 upvotes, #6 of 2023-12-06
  7. Cache Me if You Can: Accelerating Diffusion Models through Block Caching 20 upvotes, #7 of 2023-12-06
  8. LivePhoto: Real Image Animation with Text-guided Motion Control 18 upvotes, #8 of 2023-12-06
  9. Describing Differences in Image Sets with Natural Language 15 upvotes, #9 of 2023-12-06
  10. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
  11. Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models 14 upvotes, #10 of 2023-12-06
  12. LooseControl: Lifting ControlNet for Generalized Depth Conditioning 14 upvotes, #10 of 2023-12-06
  13. Orthogonal Adaptation for Modular Customization of Diffusion Models 13 upvotes, #13 of 2023-12-06
  14. Fine-grained Controllable Video Generation via Object Appearance and Context 13 upvotes, #13 of 2023-12-06
  15. DragVideo: Interactive Drag-style Video Editing 11 upvotes, #15 of 2023-12-06
  16. StableDreamer: Taming Noisy Score Distillation Sampling for Text-to-3D 10 upvotes, #16 of 2023-12-06
  17. MVHumanNet: A Large-scale Dataset of Multi-view Daily Dressing Human Captures 10 upvotes, #16 of 2023-12-06
  18. Training Chain-of-Thought via Latent-Variable Inference 9 upvotes, #18 of 2023-12-06
  19. Axiomatic Preference Modeling for Longform Question Answering 9 upvotes, #18 of 2023-12-06
  20. GPT4Point: A Unified Framework for Point-Language Understanding and Generation 9 upvotes, #18 of 2023-12-06
  21. ReconFusion: 3D Reconstruction with Diffusion Priors 9 upvotes, #18 of 2023-12-06
  22. Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia 9 upvotes, #18 of 2023-12-06
  23. WhisBERT: Multimodal Text-Audio Language Modeling on 100M Words 8 upvotes, #23 of 2023-12-06
  24. Alchemist: Parametric Control of Material Properties with Diffusion Models 8 upvotes, #23 of 2023-12-06
  25. Generating Fine-Grained Human Motions Using ChatGPT-Refined Descriptions 7 upvotes, #25 of 2023-12-06
  26. Language-Informed Visual Concept Learning 6 upvotes, #26 of 2023-12-06
  27. Multimodal Data and Resource Efficient Device-Directed Speech Detection with Large Foundation Models 5 upvotes, #27 of 2023-12-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.