Daily Papers of 2024-12-17
- Byte Latent Transformer: Patches Scale Better Than Tokens 74 upvotes, #1 of 2024-12-17
- Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models 35 upvotes, #2 of 2024-12-17
- BrushEdit: All-In-One Image Inpainting and Editing 33 upvotes, #3 of 2024-12-17
- RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation 33 upvotes, #3 of 2024-12-17
- ColorFlow: Retrieval-Augmented Image Sequence Colorization 26 upvotes, #5 of 2024-12-17
- Smaller Language Models Are Better Instruction Evolvers 24 upvotes, #6 of 2024-12-17
- Causal Diffusion Transformers for Generative Modeling 23 upvotes, #7 of 2024-12-17
- SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models 15 upvotes, #8 of 2024-12-17
- Wonderland: Navigating 3D Scenes from a Single Image 14 upvotes, #9 of 2024-12-17
- GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs 13 upvotes, #10 of 2024-12-17
- VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping 12 upvotes, #11 of 2024-12-17
- IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations 12 upvotes, #11 of 2024-12-17
- StrandHead: Text to Strand-Disentangled 3D Head Avatars Using Hair Geometric Priors 11 upvotes, #13 of 2024-12-17
- SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator 10 upvotes, #14 of 2024-12-17
- The Open Source Advantage in Large Language Models (LLMs) 9 upvotes, #15 of 2024-12-17
- Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning 8 upvotes, #16 of 2024-12-17
- SplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular Video 7 upvotes, #17 of 2024-12-17
- Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture 6 upvotes, #18 of 2024-12-17
- TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning 5 upvotes, #19 of 2024-12-17
- DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes 5 upvotes, #19 of 2024-12-17
- MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes 5 upvotes, #19 of 2024-12-17
- Whisper-GPT: A Hybrid Representation Audio Large Language Model 4 upvotes, #22 of 2024-12-17
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization 3 upvotes, #23 of 2024-12-17
- Reliable, Reproducible, and Really Fast Leaderboards with Evalica 2 upvotes, #24 of 2024-12-17
- Just a Simple Transformation is Enough for Data Protection in Vertical Federated Learning 2 upvotes, #24 of 2024-12-17
- GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training 2 upvotes, #24 of 2024-12-17
- RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning 1 upvotes, #27 of 2024-12-17
- Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion Models 1 upvotes, #27 of 2024-12-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.