Daily Papers of 2025-01-03

  1. 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining 91 upvotes, #1 of 2025-01-03
  2. VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control 46 upvotes, #2 of 2025-01-03
  3. CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings 44 upvotes, #3 of 2025-01-03
  4. LTX-Video: Realtime Video Latent Diffusion 40 upvotes, #4 of 2025-01-03
  5. VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM 40 upvotes, #4 of 2025-01-03
  6. Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models 34 upvotes, #6 of 2025-01-03
  7. ProgCo: Program Helps Self-Correction of Large Language Models 24 upvotes, #7 of 2025-01-03
  8. MLLM-as-a-Judge for Image Safety without Human Labeling 23 upvotes, #8 of 2025-01-03
  9. MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models 22 upvotes, #9 of 2025-01-03
  10. A3: Android Agent Arena for Mobile GUI Agents 22 upvotes, #9 of 2025-01-03
  11. Unifying Specialized Visual Encoders for Video Language Models 20 upvotes, #11 of 2025-01-03
  12. Dynamic Scaling of Unit Tests for Code Reward Modeling 16 upvotes, #12 of 2025-01-03
  13. SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration 10 upvotes, #13 of 2025-01-03
  14. Nested Attention: Semantic-aware Attention Values for Concept Personalization 10 upvotes, #13 of 2025-01-03
  15. MapQaTor: A System for Efficient Annotation of Map Query Datasets 8 upvotes, #15 of 2025-01-03
  16. Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing 7 upvotes, #16 of 2025-01-03
  17. Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding 5 upvotes, #17 of 2025-01-03
  18. SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization 5 upvotes, #17 of 2025-01-03
  19. Population Aware Diffusion for Time Series Generation 4 upvotes, #19 of 2025-01-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.