Daily Papers of 2025-01-16

  1. Towards Best Practices for Open Datasets for LLM Training 47 upvotes, #1 of 2025-01-16
  2. MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents 28 upvotes, #2 of 2025-01-16
  3. CityDreamer4D: Compositional Generative Model of Unbounded 4D Cities 19 upvotes, #3 of 2025-01-16
  4. RepVideo: Rethinking Cross-Layer Representation for Video Generation 15 upvotes, #4 of 2025-01-16
  5. Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion 12 upvotes, #5 of 2025-01-16
  6. XMusic: Towards a Generalized and Controllable Symbolic Music Generation Framework 10 upvotes, #6 of 2025-01-16
  7. Multimodal LLMs Can Reason about Aesthetics in Zero-Shot 10 upvotes, #6 of 2025-01-16
  8. Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 7 upvotes, #8 of 2025-01-16
  9. Trusted Machine Learning Models Unlock Private Inference for Problems Currently Infeasible with Cryptography 6 upvotes, #9 of 2025-01-16
  10. MINIMA: Modality Invariant Image Matching 3 upvotes, #10 of 2025-01-16
  11. Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding 2 upvotes, #11 of 2025-01-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.