Daily Papers of 2024-12-17

  1. Byte Latent Transformer: Patches Scale Better Than Tokens 74 upvotes, #1 of 2024-12-17
  2. Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models 35 upvotes, #2 of 2024-12-17
  3. BrushEdit: All-In-One Image Inpainting and Editing 33 upvotes, #3 of 2024-12-17
  4. RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation 33 upvotes, #3 of 2024-12-17
  5. ColorFlow: Retrieval-Augmented Image Sequence Colorization 26 upvotes, #5 of 2024-12-17
  6. Smaller Language Models Are Better Instruction Evolvers 24 upvotes, #6 of 2024-12-17
  7. Causal Diffusion Transformers for Generative Modeling 23 upvotes, #7 of 2024-12-17
  8. SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models 15 upvotes, #8 of 2024-12-17
  9. Wonderland: Navigating 3D Scenes from a Single Image 14 upvotes, #9 of 2024-12-17
  10. GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs 13 upvotes, #10 of 2024-12-17
  11. VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping 12 upvotes, #11 of 2024-12-17
  12. IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations 12 upvotes, #11 of 2024-12-17
  13. StrandHead: Text to Strand-Disentangled 3D Head Avatars Using Hair Geometric Priors 11 upvotes, #13 of 2024-12-17
  14. SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator 10 upvotes, #14 of 2024-12-17
  15. The Open Source Advantage in Large Language Models (LLMs) 9 upvotes, #15 of 2024-12-17
  16. Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning 8 upvotes, #16 of 2024-12-17
  17. SplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular Video 7 upvotes, #17 of 2024-12-17
  18. Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture 6 upvotes, #18 of 2024-12-17
  19. TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning 5 upvotes, #19 of 2024-12-17
  20. DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes 5 upvotes, #19 of 2024-12-17
  21. MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes 5 upvotes, #19 of 2024-12-17
  22. Whisper-GPT: A Hybrid Representation Audio Large Language Model 4 upvotes, #22 of 2024-12-17
  23. MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization 3 upvotes, #23 of 2024-12-17
  24. Reliable, Reproducible, and Really Fast Leaderboards with Evalica 2 upvotes, #24 of 2024-12-17
  25. Just a Simple Transformation is Enough for Data Protection in Vertical Federated Learning 2 upvotes, #24 of 2024-12-17
  26. GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training 2 upvotes, #24 of 2024-12-17
  27. RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning 1 upvotes, #27 of 2024-12-17
  28. Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion Models 1 upvotes, #27 of 2024-12-17

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.