Daily Papers of 2023-12-15

  1. StemGen: A music generation model that listens 48 upvotes, #1 of 2023-12-15
  2. TinyGSM: achieving >80% on GSM8k with small language models 40 upvotes, #2 of 2023-12-15
  3. CogAgent: A Visual Language Model for GUI Agents 32 upvotes, #3 of 2023-12-15
  4. VideoLCM: Video Latent Consistency Model 23 upvotes, #4 of 2023-12-15
  5. A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions 18 upvotes, #5 of 2023-12-15
  6. Mosaic-SDF for 3D Generative Models 16 upvotes, #6 of 2023-12-15
  7. Pixel Aligned Language Models 15 upvotes, #7 of 2023-12-15
  8. Holodeck: Language Guided Generation of 3D Embodied AI Environments 14 upvotes, #8 of 2023-12-15
  9. Zebra: Extending Context Window with Layerwise Grouped Local-Global Attention 12 upvotes, #9 of 2023-12-15
  10. SEEAvatar: Photorealistic Text-to-3D Avatar Generation with Constrained Geometry and Appearance 12 upvotes, #9 of 2023-12-15
  11. Vision-Language Models as a Source of Rewards 12 upvotes, #9 of 2023-12-15
  12. FineControlNet: Fine-level Text Control for Image Generation with Spatially Aligned Text Control Injection 10 upvotes, #12 of 2023-12-15
  13. ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks 9 upvotes, #13 of 2023-12-15
  14. General Object Foundation Model for Images and Videos at Scale 9 upvotes, #13 of 2023-12-15
  15. LIME: Localized Image Editing via Attention Regularization in Diffusion Models 9 upvotes, #13 of 2023-12-15
  16. UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation 8 upvotes, #16 of 2023-12-15
  17. Modeling Complex Mathematical Reasoning via Large Language Model based MathAgent 8 upvotes, #16 of 2023-12-15
  18. Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking 8 upvotes, #16 of 2023-12-15
  19. VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation 7 upvotes, #19 of 2023-12-15
  20. SHAP-EDITOR: Instruction-guided Latent 3D Editing in Seconds 6 upvotes, #20 of 2023-12-15
  21. TigerBot: An Open Multilingual Multitask LLM 4 upvotes, #21 of 2023-12-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.