Daily Papers of 2025-02-21

  1. MLGym: A New Framework and Benchmark for Advancing AI Research Agents 167 upvotes, #1 of 2025-02-21
  2. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features 118 upvotes, #2 of 2025-02-21
  3. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
  4. How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM? 80 upvotes, #4 of 2025-02-21
  5. S*: Test Time Scaling for Code Generation 56 upvotes, #5 of 2025-02-21
  6. Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning 42 upvotes, #6 of 2025-02-21
  7. Discovering highly efficient low-weight quantum error-correcting codes with reinforcement learning 35 upvotes, #7 of 2025-02-21
  8. S^2R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning 27 upvotes, #8 of 2025-02-21
  9. Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information 23 upvotes, #9 of 2025-02-21
  10. LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models 23 upvotes, #9 of 2025-02-21
  11. PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC 17 upvotes, #11 of 2025-02-21
  12. How to Get Your LLM to Generate Challenging Problems for Evaluation 16 upvotes, #12 of 2025-02-21
  13. Dynamic Concepts Personalization from Single Videos 14 upvotes, #13 of 2025-02-21
  14. Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation 13 upvotes, #14 of 2025-02-21
  15. LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention 12 upvotes, #15 of 2025-02-21
  16. RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers 11 upvotes, #16 of 2025-02-21
  17. NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization 11 upvotes, #16 of 2025-02-21
  18. AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO 11 upvotes, #16 of 2025-02-21
  19. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models 11 upvotes, #16 of 2025-02-21
  20. Generating Skyline Datasets for Data Science Models 7 upvotes, #20 of 2025-02-21
  21. Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data 7 upvotes, #20 of 2025-02-21
  22. Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models 7 upvotes, #20 of 2025-02-21
  23. CLIPPER: Compression enables long-context synthetic data generation 7 upvotes, #20 of 2025-02-21
  24. LLM-based User Profile Management for Recommender System 5 upvotes, #24 of 2025-02-21
  25. Generating π-Functional Molecules Using STGG+ with Active Learning 4 upvotes, #25 of 2025-02-21
  26. How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild 3 upvotes, #26 of 2025-02-21
  27. Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework 3 upvotes, #26 of 2025-02-21
  28. Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images 3 upvotes, #26 of 2025-02-21
  29. Unstructured Evidence Attribution for Long Context Query Focused Summarization 3 upvotes, #26 of 2025-02-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.