Daily Papers of 2025-01-28

  1. Qwen2.5-1M Technical Report 51 upvotes, #1 of 2025-01-28
  2. Baichuan-Omni-1.5 Technical Report 48 upvotes, #2 of 2025-01-28
  3. ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer 23 upvotes, #3 of 2025-01-28
  4. Towards General-Purpose Model-Free Reinforcement Learning 23 upvotes, #3 of 2025-01-28
  5. Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation 15 upvotes, #5 of 2025-01-28
  6. iFormer: Integrating ConvNet and Transformer for Mobile Application 10 upvotes, #6 of 2025-01-28
  7. Are Vision Language Models Texture or Shape Biased and Can We Steer Them? 9 upvotes, #7 of 2025-01-28
  8. Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models 9 upvotes, #7 of 2025-01-28
  9. Visual Generation Without Guidance 8 upvotes, #9 of 2025-01-28
  10. CodeMonkeys: Scaling Test-Time Compute for Software Engineering 7 upvotes, #10 of 2025-01-28
  11. Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity 7 upvotes, #10 of 2025-01-28
  12. OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas 6 upvotes, #12 of 2025-01-28
  13. Feasible Learning 5 upvotes, #13 of 2025-01-28
  14. Return of the Encoder: Maximizing Parameter Efficiency for SLMs 5 upvotes, #13 of 2025-01-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.