Daily Papers of 2024-08-30
- Law of Vision Representation in MLLMs 87 upvotes, #1 of 2024-08-30
- CogVLM2: Visual Language Models for Image and Video Understanding 55 upvotes, #2 of 2024-08-30
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling 42 upvotes, #3 of 2024-08-30
- ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model 28 upvotes, #4 of 2024-08-30
- SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners 25 upvotes, #5 of 2024-08-30
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems 20 upvotes, #6 of 2024-08-30
- CSGO: Content-Style Composition in Text-to-Image Generation 16 upvotes, #7 of 2024-08-30
- 3D Reconstruction with Spatial Memory 11 upvotes, #8 of 2024-08-30
- StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements 9 upvotes, #9 of 2024-08-30
- Scaling Up Diffusion and Flow-based XGBoost Models 8 upvotes, #10 of 2024-08-30
- Meta Flow Matching: Integrating Vector Fields on the Wasserstein Manifold 7 upvotes, #11 of 2024-08-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.