Daily Papers of 2025-03-04

  1. Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs 70 upvotes, #1 of 2025-03-04
  2. Visual-RFT: Visual Reinforcement Fine-Tuning 62 upvotes, #2 of 2025-03-04
  3. Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models 39 upvotes, #3 of 2025-03-04
  4. Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs 31 upvotes, #4 of 2025-03-04
  5. DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion 26 upvotes, #5 of 2025-03-04
  6. From Hours to Minutes: Lossless Acceleration of Ultra Long Sequence Generation up to 100K Tokens 24 upvotes, #6 of 2025-03-04
  7. OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment 23 upvotes, #7 of 2025-03-04
  8. When an LLM is apprehensive about its answers -- and when its uncertainty is justified 19 upvotes, #8 of 2025-03-04
  9. Liger: Linearizing Large Language Models to Gated Recurrent Structures 15 upvotes, #9 of 2025-03-04
  10. Efficient Test-Time Scaling via Self-Calibration 13 upvotes, #10 of 2025-03-04
  11. Speculative Ad-hoc Querying 12 upvotes, #11 of 2025-03-04
  12. Qilin: A Multimodal Information Retrieval Dataset with APP-level User Sessions 11 upvotes, #12 of 2025-03-04
  13. DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting 10 upvotes, #13 of 2025-03-04
  14. Large-Scale Data Selection for Instruction Tuning 10 upvotes, #13 of 2025-03-04
  15. Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation 9 upvotes, #15 of 2025-03-04
  16. SampleMix: A Sample-wise Pre-training Data Mixing Strategey by Coordinating Data Quality and Diversity 8 upvotes, #16 of 2025-03-04
  17. CodeArena: A Collective Evaluation Platform for LLM Code Generation 7 upvotes, #17 of 2025-03-04
  18. VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation 7 upvotes, #17 of 2025-03-04
  19. PodAgent: A Comprehensive Framework for Podcast Generation 6 upvotes, #19 of 2025-03-04
  20. AI-Invented Tonal Languages: Preventing a Machine Lingua Franca Beyond Human Understanding 5 upvotes, #20 of 2025-03-04
  21. Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia 5 upvotes, #20 of 2025-03-04
  22. General Reasoning Requires Learning to Reason from the Get-go 4 upvotes, #22 of 2025-03-04
  23. CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic Environments 3 upvotes, #23 of 2025-03-04
  24. Direct Discriminative Optimization: Your Likelihood-Based Visual Generative Model is Secretly a GAN Discriminator 3 upvotes, #23 of 2025-03-04
  25. Teaching Metric Distance to Autoregressive Multimodal Foundational Models 3 upvotes, #23 of 2025-03-04
  26. Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model 2 upvotes, #26 of 2025-03-04
  27. Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis 2 upvotes, #26 of 2025-03-04
  28. RSQ: Learning from Important Tokens Leads to Better Quantized LLMs 2 upvotes, #26 of 2025-03-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.