Daily Papers of 2024-03-15

  1. MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training 105 upvotes, #1 of 2024-03-15
  2. Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking 47 upvotes, #2 of 2024-03-15
  3. Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset 46 upvotes, #3 of 2024-03-15
  4. GiT: Towards Generalist Vision Transformer through Universal Language Interface 23 upvotes, #4 of 2024-03-15
  5. StreamMultiDiffusion: Real-Time Interactive Generation with Region-Based Semantic Control 20 upvotes, #5 of 2024-03-15
  6. Video Editing via Factorized Diffusion Distillation 20 upvotes, #5 of 2024-03-15
  7. BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences 19 upvotes, #7 of 2024-03-15
  8. Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring 12 upvotes, #8 of 2024-03-15
  9. Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding 10 upvotes, #9 of 2024-03-15
  10. VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding 8 upvotes, #10 of 2024-03-15
  11. Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering 8 upvotes, #10 of 2024-03-15
  12. Veagle: Advancements in Multimodal Representation Learning 7 upvotes, #12 of 2024-03-15
  13. LocalMamba: Visual State Space Model with Windowed Selective Scan 7 upvotes, #12 of 2024-03-15
  14. 3D-VLA: A 3D Vision-Language-Action Generative World Model 6 upvotes, #14 of 2024-03-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.