Daily Papers of 2025-02-27

  1. GHOST 2.0: generative high-fidelity one shot transfer of heads 62 upvotes, #1 of 2025-02-27
  2. Kanana: Compute-efficient Bilingual Language Models 59 upvotes, #2 of 2025-02-27
  3. TheoremExplainAgent: Towards Multimodal Explanations for LLM Theorem Understanding 42 upvotes, #3 of 2025-02-27
  4. Towards an AI co-scientist 39 upvotes, #4 of 2025-02-27
  5. Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance 30 upvotes, #5 of 2025-02-27
  6. Language Models' Factuality Depends on the Language of Inquiry 29 upvotes, #6 of 2025-02-27
  7. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? 26 upvotes, #7 of 2025-02-27
  8. Rank1: Test-Time Compute for Reranking in Information Retrieval 25 upvotes, #8 of 2025-02-27
  9. Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems 21 upvotes, #9 of 2025-02-27
  10. Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs 19 upvotes, #10 of 2025-02-27
  11. Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation 18 upvotes, #11 of 2025-02-27
  12. VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model 11 upvotes, #12 of 2025-02-27
  13. Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator 11 upvotes, #12 of 2025-02-27
  14. CritiQ: Mining Data Quality Criteria from Human Preferences 9 upvotes, #14 of 2025-02-27
  15. BIG-Bench Extra Hard 6 upvotes, #15 of 2025-02-27
  16. Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization 6 upvotes, #15 of 2025-02-27
  17. MolSpectra: Pre-training 3D Molecular Representation with Multi-modal Energy Spectra 5 upvotes, #17 of 2025-02-27
  18. AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement 5 upvotes, #17 of 2025-02-27
  19. FSPO: Few-Shot Preference Optimization of Synthetic Preference Data in LLMs Elicits Effective Personalization to Real Users 5 upvotes, #17 of 2025-02-27
  20. Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications 5 upvotes, #17 of 2025-02-27
  21. Towards Optimal Multi-draft Speculative Decoding 4 upvotes, #21 of 2025-02-27
  22. MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge 3 upvotes, #22 of 2025-02-27
  23. DOEI: Dual Optimization of Embedding Information for Attention-Enhanced Class Activation Maps 2 upvotes, #23 of 2025-02-27
  24. PosterSum: A Multimodal Benchmark for Scientific Poster Summarization 2 upvotes, #23 of 2025-02-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.