Daily Papers of 2024-03-11

  1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 46 upvotes, #1 of 2024-03-11
  2. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment 37 upvotes, #2 of 2024-03-11
  3. DeepSeek-VL: Towards Real-World Vision-Language Understanding 33 upvotes, #3 of 2024-03-11
  4. Personalized Audiobook Recommendations at Spotify Through Graph Neural Networks 18 upvotes, #4 of 2024-03-11
  5. CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion 15 upvotes, #5 of 2024-03-11
  6. VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models 14 upvotes, #6 of 2024-03-11
  7. CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model 12 upvotes, #7 of 2024-03-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.