Daily Papers of 2023-12-06
- ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation 33 upvotes, #1 of 2023-12-06
- FaceStudio: Put Your Face Everywhere in Seconds 32 upvotes, #2 of 2023-12-06
- Analyzing and Improving the Training Dynamics of Diffusion Models 32 upvotes, #2 of 2023-12-06
- Relightable Gaussian Codec Avatars 31 upvotes, #4 of 2023-12-06
- X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model 26 upvotes, #5 of 2023-12-06
- OneLLM: One Framework to Align All Modalities with Language 23 upvotes, #6 of 2023-12-06
- Cache Me if You Can: Accelerating Diffusion Models through Block Caching 20 upvotes, #7 of 2023-12-06
- LivePhoto: Real Image Animation with Text-guided Motion Control 18 upvotes, #8 of 2023-12-06
- Describing Differences in Image Sets with Natural Language 15 upvotes, #9 of 2023-12-06
- LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
- Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models 14 upvotes, #10 of 2023-12-06
- LooseControl: Lifting ControlNet for Generalized Depth Conditioning 14 upvotes, #10 of 2023-12-06
- Orthogonal Adaptation for Modular Customization of Diffusion Models 13 upvotes, #13 of 2023-12-06
- Fine-grained Controllable Video Generation via Object Appearance and Context 13 upvotes, #13 of 2023-12-06
- DragVideo: Interactive Drag-style Video Editing 11 upvotes, #15 of 2023-12-06
- StableDreamer: Taming Noisy Score Distillation Sampling for Text-to-3D 10 upvotes, #16 of 2023-12-06
- MVHumanNet: A Large-scale Dataset of Multi-view Daily Dressing Human Captures 10 upvotes, #16 of 2023-12-06
- Training Chain-of-Thought via Latent-Variable Inference 9 upvotes, #18 of 2023-12-06
- Axiomatic Preference Modeling for Longform Question Answering 9 upvotes, #18 of 2023-12-06
- GPT4Point: A Unified Framework for Point-Language Understanding and Generation 9 upvotes, #18 of 2023-12-06
- ReconFusion: 3D Reconstruction with Diffusion Priors 9 upvotes, #18 of 2023-12-06
- Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia 9 upvotes, #18 of 2023-12-06
- WhisBERT: Multimodal Text-Audio Language Modeling on 100M Words 8 upvotes, #23 of 2023-12-06
- Alchemist: Parametric Control of Material Properties with Diffusion Models 8 upvotes, #23 of 2023-12-06
- Generating Fine-Grained Human Motions Using ChatGPT-Refined Descriptions 7 upvotes, #25 of 2023-12-06
- Language-Informed Visual Concept Learning 6 upvotes, #26 of 2023-12-06
- Multimodal Data and Resource Efficient Device-Directed Speech Detection with Large Foundation Models 5 upvotes, #27 of 2023-12-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.