Daily Papers of 2024-12-11

  1. STIV: Scalable Text and Image Conditioned Video Generation 67 upvotes, #1 of 2024-12-11
  2. Evaluating and Aligning CodeLLMs on Human Preference 47 upvotes, #2 of 2024-12-11
  3. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation 43 upvotes, #3 of 2024-12-11
  4. ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer 30 upvotes, #4 of 2024-12-11
  5. Hidden in the Noise: Two-Stage Robust Watermarking for Images 28 upvotes, #5 of 2024-12-11
  6. UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics 23 upvotes, #6 of 2024-12-11
  7. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 20 upvotes, #7 of 2024-12-11
  8. Mobile Video Diffusion 19 upvotes, #8 of 2024-12-11
  9. MoViE: Mobile Diffusion for Video Editing 18 upvotes, #9 of 2024-12-11
  10. Granite Guardian 18 upvotes, #9 of 2024-12-11
  11. 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation 18 upvotes, #9 of 2024-12-11
  12. Perception Tokens Enhance Visual Reasoning in Multimodal Language Models 16 upvotes, #12 of 2024-12-11
  13. Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation 16 upvotes, #12 of 2024-12-11
  14. Video Motion Transfer with Diffusion Transformers 16 upvotes, #12 of 2024-12-11
  15. EMOv2: Pushing 5M Vision Model Frontier 13 upvotes, #15 of 2024-12-11
  16. LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation 11 upvotes, #16 of 2024-12-11
  17. ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance 11 upvotes, #16 of 2024-12-11
  18. Fully Open Source Moxin-7B Technical Report 10 upvotes, #18 of 2024-12-11
  19. Chimera: Improving Generalist Model with Domain-Specific Experts 9 upvotes, #19 of 2024-12-11
  20. ObjCtrl-2.5D: Training-free Object Control with Camera Poses 8 upvotes, #20 of 2024-12-11
  21. GraPE: A Generate-Plan-Edit Framework for Compositional T2I Synthesis 4 upvotes, #21 of 2024-12-11
  22. HARP: Hesitation-Aware Reframing in Transformer Inference Pass 4 upvotes, #21 of 2024-12-11
  23. Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment 2 upvotes, #23 of 2024-12-11
  24. A New Federated Learning Framework Against Gradient Inversion Attacks 2 upvotes, #23 of 2024-12-11
  25. Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation 2 upvotes, #23 of 2024-12-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.