Daily Papers of 2024-01-30

  1. InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
  2. MoE-LLaVA: Mixture of Experts for Large Vision-Language Models 54 upvotes, #2 of 2024-01-30
  3. Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling 53 upvotes, #3 of 2024-01-30
  4. Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling 39 upvotes, #4 of 2024-01-30
  5. SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning 26 upvotes, #5 of 2024-01-30
  6. Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance 24 upvotes, #6 of 2024-01-30
  7. Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception 20 upvotes, #7 of 2024-01-30
  8. StableIdentity: Inserting Anybody into Anywhere at First Sight 18 upvotes, #8 of 2024-01-30
  9. Object-Driven One-Shot Fine-tuning of Text-to-Image Diffusion with Prototypical Embedding 12 upvotes, #9 of 2024-01-30
  10. Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation 11 upvotes, #10 of 2024-01-30
  11. Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization 7 upvotes, #11 of 2024-01-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.