Daily Papers of 2024-06-13

  1. NaRCan: Natural Refined Canonical Image with Integration of Diffusion Prior for Video Editing 43 upvotes, #1 of 2024-06-13
  2. Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing 42 upvotes, #2 of 2024-06-13
  3. MotionClone: Training-Free Motion Cloning for Controllable Video Generation 38 upvotes, #3 of 2024-06-13
  4. Are We Done with MMLU? 35 upvotes, #4 of 2024-06-13
  5. What If We Recaption Billions of Web Images with LLaMA-3? 35 upvotes, #4 of 2024-06-13
  6. PowerInfer-2: Fast Large Language Model Inference on a Smartphone 33 upvotes, #6 of 2024-06-13
  7. Physics3D: Learning Physical Properties of 3D Gaussians via Video Diffusion 29 upvotes, #7 of 2024-06-13
  8. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 28 upvotes, #8 of 2024-06-13
  9. 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination 27 upvotes, #9 of 2024-06-13
  10. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
  11. Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters 21 upvotes, #11 of 2024-06-13
  12. FontStudio: Shape-Adaptive Diffusion Model for Coherent and Consistent Font Effect Generation 17 upvotes, #12 of 2024-06-13
  13. AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation 13 upvotes, #13 of 2024-06-13
  14. Hierarchical Patch Diffusion Models for High-Resolution Video Generation 13 upvotes, #13 of 2024-06-13
  15. Discovering Preference Optimization Algorithms with and for Large Language Models 12 upvotes, #15 of 2024-06-13
  16. VCR: Visual Caption Restoration 10 upvotes, #16 of 2024-06-13
  17. Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models 10 upvotes, #16 of 2024-06-13
  18. Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models 7 upvotes, #18 of 2024-06-13
  19. Hibou: A Family of Foundational Vision Transformers for Pathology 6 upvotes, #19 of 2024-06-13
  20. Large Language Model Unlearning via Embedding-Corrupted Prompts 6 upvotes, #19 of 2024-06-13
  21. Simplified and Generalized Masked Diffusion for Discrete Data 4 upvotes, #21 of 2024-06-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.