Daily Papers of 2024-12-09

  1. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
  2. EXAONE 3.5: Series of Large Language Models for Real-world Use Cases 45 upvotes, #2 of 2024-12-09
  3. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
  4. LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment 39 upvotes, #4 of 2024-12-09
  5. APOLLO: SGD-like Memory, AdamW-level Performance 36 upvotes, #5 of 2024-12-09
  6. SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion 30 upvotes, #6 of 2024-12-09
  7. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 20 upvotes, #7 of 2024-12-09
  8. GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration 17 upvotes, #8 of 2024-12-09
  9. CompCap: Improving Multimodal Large Language Models with Composite Captions 17 upvotes, #8 of 2024-12-09
  10. Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction 14 upvotes, #10 of 2024-12-09
  11. PanoDreamer: 3D Panorama Synthesis from a Single Image 9 upvotes, #11 of 2024-12-09
  12. Mind the Time: Temporally-Controlled Multi-Event Video Generation 9 upvotes, #11 of 2024-12-09
  13. 2DGS-Room: Seed-Guided 2D Gaussian Splatting with Geometric Constrains for High-Fidelity Indoor Scene Reconstruction 8 upvotes, #13 of 2024-12-09
  14. BigDocs: An Open and Permissively-Licensed Dataset for Training Multimodal Models on Document and Code Tasks 8 upvotes, #13 of 2024-12-09
  15. DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling 7 upvotes, #15 of 2024-12-09
  16. RL Zero: Zero-Shot Language to Behaviors without any Supervision 4 upvotes, #16 of 2024-12-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.