Daily Papers of 2024-07-11

  1. PaliGemma: A versatile 3B VLM for transfer 58 upvotes, #1 of 2024-07-11
  2. Inference Performance Optimization for Large Language Models on CPUs 47 upvotes, #2 of 2024-07-11
  3. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 34 upvotes, #3 of 2024-07-11
  4. Controlling Space and Time with Diffusion Models 15 upvotes, #4 of 2024-07-11
  5. Video-to-Audio Generation with Hidden Alignment 11 upvotes, #5 of 2024-07-11
  6. Still-Moving: Customized Video Generation without Customized Video Data 9 upvotes, #6 of 2024-07-11
  7. VEnhancer: Generative Space-Time Enhancement for Video Generation 8 upvotes, #7 of 2024-07-11
  8. Do Vision and Language Models Share Concepts? A Vector Space Alignment Study 7 upvotes, #8 of 2024-07-11
  9. CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging 4 upvotes, #9 of 2024-07-11
  10. On Leakage of Code Generation Evaluation Datasets 4 upvotes, #9 of 2024-07-11
  11. This&That: Language-Gesture Controlled Video Generation for Robot Planning 3 upvotes, #11 of 2024-07-11
  12. An accurate detection is not all you need to combat label noise in web-noisy datasets 2 upvotes, #12 of 2024-07-11
  13. CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation 1 upvotes, #13 of 2024-07-11
  14. BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark 1 upvotes, #13 of 2024-07-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.