Daily Papers of 2024-05-28

  1. An Introduction to Vision-Language Modeling 68 upvotes, #1 of 2024-05-28
  2. Transformers Can Do Arithmetic with the Right Embeddings 49 upvotes, #2 of 2024-05-28
  3. Matryoshka Multimodal Models 29 upvotes, #3 of 2024-05-28
  4. Zamba: A Compact 7B SSM Hybrid Model 17 upvotes, #4 of 2024-05-28
  5. I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models 15 upvotes, #5 of 2024-05-28
  6. Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer 13 upvotes, #6 of 2024-05-28
  7. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models 13 upvotes, #6 of 2024-05-28
  8. Looking Backward: Streaming Video-to-Video Translation with Feature Banks 12 upvotes, #8 of 2024-05-28
  9. Trans-LoRA: towards data-free Transferable Parameter Efficient Finetuning 12 upvotes, #8 of 2024-05-28
  10. Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels 11 upvotes, #10 of 2024-05-28
  11. LoGAH: Predicting 774-Million-Parameter Transformers using Graph HyperNetworks with 1/100 Parameters 10 upvotes, #11 of 2024-05-28
  12. EM Distillation for One-step Diffusion Models 10 upvotes, #11 of 2024-05-28
  13. Part123: Part-aware 3D Reconstruction from a Single-view Image 10 upvotes, #11 of 2024-05-28
  14. Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control 9 upvotes, #14 of 2024-05-28
  15. Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models 7 upvotes, #15 of 2024-05-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.