Daily Papers of 2024-05-28
- An Introduction to Vision-Language Modeling 68 upvotes, #1 of 2024-05-28
- Transformers Can Do Arithmetic with the Right Embeddings 49 upvotes, #2 of 2024-05-28
- Matryoshka Multimodal Models 29 upvotes, #3 of 2024-05-28
- Zamba: A Compact 7B SSM Hybrid Model 17 upvotes, #4 of 2024-05-28
- I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models 15 upvotes, #5 of 2024-05-28
- Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer 13 upvotes, #6 of 2024-05-28
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models 13 upvotes, #6 of 2024-05-28
- Looking Backward: Streaming Video-to-Video Translation with Feature Banks 12 upvotes, #8 of 2024-05-28
- Trans-LoRA: towards data-free Transferable Parameter Efficient Finetuning 12 upvotes, #8 of 2024-05-28
- Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels 11 upvotes, #10 of 2024-05-28
- LoGAH: Predicting 774-Million-Parameter Transformers using Graph HyperNetworks with 1/100 Parameters 10 upvotes, #11 of 2024-05-28
- EM Distillation for One-step Diffusion Models 10 upvotes, #11 of 2024-05-28
- Part123: Part-aware 3D Reconstruction from a Single-view Image 10 upvotes, #11 of 2024-05-28
- Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control 9 upvotes, #14 of 2024-05-28
- Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models 7 upvotes, #15 of 2024-05-28
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.