Daily Papers of 2025-05-27

  1. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model 212 upvotes, #1 of 2025-05-27
  2. Shifting AI Efficiency From Model-Centric to Data-Centric Compression 141 upvotes, #2 of 2025-05-27
  3. Alchemist: Turning Public Text-to-Image Data into Generative Gold 73 upvotes, #3 of 2025-05-27
  4. BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs 61 upvotes, #4 of 2025-05-27
  5. Embodied Agents Meet Personalization: Exploring Memory Utilization for Personalized Assistance 46 upvotes, #5 of 2025-05-27
  6. PATS: Process-Level Adaptive Thinking Mode Switching 46 upvotes, #5 of 2025-05-27
  7. ARM: Adaptive Reasoning Model 43 upvotes, #7 of 2025-05-27
  8. Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles 40 upvotes, #8 of 2025-05-27
  9. Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective 36 upvotes, #9 of 2025-05-27
  10. B-score: Detecting biases in large language models using response history 30 upvotes, #10 of 2025-05-27
  11. Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers 30 upvotes, #10 of 2025-05-27
  12. Flex-Judge: Think Once, Judge Anywhere 27 upvotes, #12 of 2025-05-27
  13. Learning to Reason without External Rewards 25 upvotes, #13 of 2025-05-27
  14. MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search 24 upvotes, #14 of 2025-05-27
  15. Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps 23 upvotes, #15 of 2025-05-27
  16. Lifelong Safety Alignment for Language Models 23 upvotes, #15 of 2025-05-27
  17. ModernGBERT: German-only 1B Encoder Model Trained from Scratch 20 upvotes, #17 of 2025-05-27
  18. Jodi: Unification of Visual Generation and Understanding via Joint Modeling 20 upvotes, #17 of 2025-05-27
  19. Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 18 upvotes, #19 of 2025-05-27
  20. StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
  21. Hybrid Neural-MPM for Interactive Fluid Simulations in Real-Time 17 upvotes, #21 of 2025-05-27
  22. Discrete Markov Bridge 17 upvotes, #21 of 2025-05-27
  23. REARANK: Reasoning Re-ranking Agent via Reinforcement Learning 17 upvotes, #21 of 2025-05-27
  24. Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration 17 upvotes, #21 of 2025-05-27
  25. Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions 15 upvotes, #25 of 2025-05-27
  26. AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting 14 upvotes, #26 of 2025-05-27
  27. Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI 14 upvotes, #26 of 2025-05-27
  28. WHISTRESS: Enriching Transcriptions with Sentence Stress Detection 13 upvotes, #28 of 2025-05-27
  29. Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression 13 upvotes, #28 of 2025-05-27
  30. Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition 13 upvotes, #28 of 2025-05-27
  31. G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning 12 upvotes, #31 of 2025-05-27
  32. The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation 12 upvotes, #31 of 2025-05-27
  33. Interleaved Reasoning for Large Language Models via Reinforcement Learning 12 upvotes, #31 of 2025-05-27
  34. Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals 11 upvotes, #34 of 2025-05-27
  35. Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models 11 upvotes, #34 of 2025-05-27
  36. InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction 10 upvotes, #36 of 2025-05-27
  37. MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research 10 upvotes, #36 of 2025-05-27
  38. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition 9 upvotes, #38 of 2025-05-27
  39. WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference 9 upvotes, #38 of 2025-05-27
  40. STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs 8 upvotes, #40 of 2025-05-27
  41. LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models 8 upvotes, #40 of 2025-05-27
  42. Dynamic Risk Assessments for Offensive Cybersecurity Agents 7 upvotes, #42 of 2025-05-27
  43. Strong Membership Inference Attacks on Massive Datasets and (Moderately) Large Language Models 7 upvotes, #42 of 2025-05-27
  44. The Coverage Principle: A Framework for Understanding Compositional Generalization 7 upvotes, #42 of 2025-05-27
  45. Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary? 6 upvotes, #45 of 2025-05-27
  46. Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective 6 upvotes, #45 of 2025-05-27
  47. Accelerating Nash Learning from Human Feedback via Mirror Prox 6 upvotes, #45 of 2025-05-27
  48. Hybrid Latent Reasoning via Reinforcement Learning 5 upvotes, #48 of 2025-05-27
  49. DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical Dialogue 5 upvotes, #48 of 2025-05-27
  50. Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs 5 upvotes, #48 of 2025-05-27
  51. Bridging Supervised Learning and Reinforcement Learning in Math Reasoning 4 upvotes, #51 of 2025-05-27
  52. An Embarrassingly Simple Defense Against LLM Abliteration Attacks 4 upvotes, #51 of 2025-05-27
  53. GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes 4 upvotes, #51 of 2025-05-27
  54. EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning 3 upvotes, #54 of 2025-05-27
  55. UFT: Unifying Supervised and Reinforcement Fine-Tuning 3 upvotes, #54 of 2025-05-27
  56. Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation 3 upvotes, #54 of 2025-05-27
  57. Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision 3 upvotes, #54 of 2025-05-27
  58. Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey 2 upvotes, #58 of 2025-05-27
  59. TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification 2 upvotes, #58 of 2025-05-27
  60. InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning 2 upvotes, #58 of 2025-05-27
  61. The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models 2 upvotes, #58 of 2025-05-27
  62. MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models 2 upvotes, #58 of 2025-05-27
  63. FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models 2 upvotes, #58 of 2025-05-27
  64. Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models 2 upvotes, #58 of 2025-05-27
  65. DiSA: Diffusion Step Annealing in Autoregressive Image Generation 2 upvotes, #58 of 2025-05-27
  66. Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning 1 upvotes, #66 of 2025-05-27
  67. Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models 1 upvotes, #66 of 2025-05-27
  68. CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark 1 upvotes, #66 of 2025-05-27
  69. The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models 1 upvotes, #66 of 2025-05-27
  70. MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs 1 upvotes, #66 of 2025-05-27
  71. EgoZero: Robot Learning from Smart Glasses 1 upvotes, #66 of 2025-05-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.