Daily Papers of 2025-04-28

  1. Towards Understanding Camera Motions in Any Video 154 upvotes, #1 of 2025-04-28
  2. Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning 53 upvotes, #2 of 2025-04-28
  3. BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs 41 upvotes, #3 of 2025-04-28
  4. VideoVista-CulturalLingo: 360^circ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension 21 upvotes, #4 of 2025-04-28
  5. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark 17 upvotes, #5 of 2025-04-28
  6. Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation 16 upvotes, #6 of 2025-04-28
  7. Kimi-Audio Technical Report 14 upvotes, #7 of 2025-04-28
  8. The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs 12 upvotes, #8 of 2025-04-28
  9. Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family 12 upvotes, #8 of 2025-04-28
  10. Subject-driven Video Generation via Disentangled Identity and Motion 11 upvotes, #10 of 2025-04-28
  11. DianJin-R1: Evaluating and Enhancing Financial Reasoning in Large Language Models 9 upvotes, #11 of 2025-04-28
  12. DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency 7 upvotes, #12 of 2025-04-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.