Daily Papers of 2024-07-11
- PaliGemma: A versatile 3B VLM for transfer 58 upvotes, #1 of 2024-07-11
- Inference Performance Optimization for Large Language Models on CPUs 47 upvotes, #2 of 2024-07-11
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 34 upvotes, #3 of 2024-07-11
- Controlling Space and Time with Diffusion Models 15 upvotes, #4 of 2024-07-11
- Video-to-Audio Generation with Hidden Alignment 11 upvotes, #5 of 2024-07-11
- Still-Moving: Customized Video Generation without Customized Video Data 9 upvotes, #6 of 2024-07-11
- VEnhancer: Generative Space-Time Enhancement for Video Generation 8 upvotes, #7 of 2024-07-11
- Do Vision and Language Models Share Concepts? A Vector Space Alignment Study 7 upvotes, #8 of 2024-07-11
- CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging 4 upvotes, #9 of 2024-07-11
- On Leakage of Code Generation Evaluation Datasets 4 upvotes, #9 of 2024-07-11
- This&That: Language-Gesture Controlled Video Generation for Robot Planning 3 upvotes, #11 of 2024-07-11
- An accurate detection is not all you need to combat label noise in web-noisy datasets 2 upvotes, #12 of 2024-07-11
- CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation 1 upvotes, #13 of 2024-07-11
- BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark 1 upvotes, #13 of 2024-07-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.