Daily Papers of 2024-11-18

  1. LLaVA-o1: Let Vision Language Models Reason Step-by-Step 99 upvotes, #1 of 2024-11-18
  2. Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement 32 upvotes, #2 of 2024-11-18
  3. The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use 27 upvotes, #3 of 2024-11-18
  4. GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation 21 upvotes, #4 of 2024-11-18
  5. Xmodel-1.5: An 1B-scale Multilingual LLM 14 upvotes, #5 of 2024-11-18
  6. Number it: Temporal Grounding Videos like Flipping Manga 12 upvotes, #6 of 2024-11-18
  7. MARS: Unleashing the Power of Variance Reduction for Training Large Models 11 upvotes, #7 of 2024-11-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.