Daily Papers of 2025-02-19

  1. Soundwave: Less is More for Speech-Text Alignment in LLMs 76 upvotes, #1 of 2025-02-19
  2. Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity 62 upvotes, #2 of 2025-02-19
  3. Phantom: Subject-consistent video generation via cross-modal alignment 50 upvotes, #3 of 2025-02-19
  4. Continuous Diffusion Model for Language Modeling 49 upvotes, #4 of 2025-02-19
  5. Magma: A Foundation Model for Multimodal AI Agents 46 upvotes, #5 of 2025-02-19
  6. Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation 35 upvotes, #6 of 2025-02-19
  7. Rethinking Diverse Human Preference Learning through Principal Component Analysis 34 upvotes, #7 of 2025-02-19
  8. You Do Not Fully Utilize Transformer's Representation Capacity 33 upvotes, #8 of 2025-02-19
  9. FLAG-Trader: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading 31 upvotes, #9 of 2025-02-19
  10. SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation 29 upvotes, #10 of 2025-02-19
  11. SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models 27 upvotes, #11 of 2025-02-19
  12. Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering 27 upvotes, #11 of 2025-02-19
  13. Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities? 16 upvotes, #13 of 2025-02-19
  14. OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning 15 upvotes, #14 of 2025-02-19
  15. RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm 15 upvotes, #14 of 2025-02-19
  16. PAFT: Prompt-Agnostic Fine-Tuning 15 upvotes, #14 of 2025-02-19
  17. Atom of Thoughts for Markov LLM Test-Time Scaling 12 upvotes, #17 of 2025-02-19
  18. Text2World: Benchmarking Large Language Models for Symbolic World Model Generation 12 upvotes, #17 of 2025-02-19
  19. MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections 11 upvotes, #19 of 2025-02-19
  20. YOLOv12: Attention-Centric Real-Time Object Detectors 10 upvotes, #20 of 2025-02-19
  21. HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading 10 upvotes, #20 of 2025-02-19
  22. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation 9 upvotes, #22 of 2025-02-19
  23. Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options 7 upvotes, #23 of 2025-02-19
  24. Eager Updates For Overlapped Communication and Computation in DiLoCo 7 upvotes, #23 of 2025-02-19
  25. Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge 6 upvotes, #25 of 2025-02-19
  26. The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1 5 upvotes, #26 of 2025-02-19
  27. Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey 4 upvotes, #27 of 2025-02-19
  28. Pre-training Auto-regressive Robotic Models with 4D Representations 4 upvotes, #27 of 2025-02-19
  29. FinMTEB: Finance Massive Text Embedding Benchmark 3 upvotes, #29 of 2025-02-19
  30. Harnessing Vision Models for Time Series Analysis: A Survey 2 upvotes, #30 of 2025-02-19
  31. Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages 2 upvotes, #30 of 2025-02-19
  32. Scaling Autonomous Agents via Automatic Reward Modeling And Planning 2 upvotes, #30 of 2025-02-19
  33. Perovskite-LLM: Knowledge-Enhanced Large Language Models for Perovskite Solar Cell Research 2 upvotes, #30 of 2025-02-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.