Daily Papers of 2024-01-30
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models 54 upvotes, #2 of 2024-01-30
- Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling 53 upvotes, #3 of 2024-01-30
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling 39 upvotes, #4 of 2024-01-30
- SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning 26 upvotes, #5 of 2024-01-30
- Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance 24 upvotes, #6 of 2024-01-30
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception 20 upvotes, #7 of 2024-01-30
- StableIdentity: Inserting Anybody into Anywhere at First Sight 18 upvotes, #8 of 2024-01-30
- Object-Driven One-Shot Fine-tuning of Text-to-Image Diffusion with Prototypical Embedding 12 upvotes, #9 of 2024-01-30
- Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation 11 upvotes, #10 of 2024-01-30
- Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization 7 upvotes, #11 of 2024-01-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.