Daily Papers of 2025-05-27
- Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model 212 upvotes, #1 of 2025-05-27
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression 141 upvotes, #2 of 2025-05-27
- Alchemist: Turning Public Text-to-Image Data into Generative Gold 73 upvotes, #3 of 2025-05-27
- BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs 61 upvotes, #4 of 2025-05-27
- Embodied Agents Meet Personalization: Exploring Memory Utilization for Personalized Assistance 46 upvotes, #5 of 2025-05-27
- PATS: Process-Level Adaptive Thinking Mode Switching 46 upvotes, #5 of 2025-05-27
- ARM: Adaptive Reasoning Model 43 upvotes, #7 of 2025-05-27
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles 40 upvotes, #8 of 2025-05-27
- Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective 36 upvotes, #9 of 2025-05-27
- B-score: Detecting biases in large language models using response history 30 upvotes, #10 of 2025-05-27
- Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers 30 upvotes, #10 of 2025-05-27
- Flex-Judge: Think Once, Judge Anywhere 27 upvotes, #12 of 2025-05-27
- Learning to Reason without External Rewards 25 upvotes, #13 of 2025-05-27
- MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search 24 upvotes, #14 of 2025-05-27
- Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps 23 upvotes, #15 of 2025-05-27
- Lifelong Safety Alignment for Language Models 23 upvotes, #15 of 2025-05-27
- ModernGBERT: German-only 1B Encoder Model Trained from Scratch 20 upvotes, #17 of 2025-05-27
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling 20 upvotes, #17 of 2025-05-27
- Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 18 upvotes, #19 of 2025-05-27
- StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
- Hybrid Neural-MPM for Interactive Fluid Simulations in Real-Time 17 upvotes, #21 of 2025-05-27
- Discrete Markov Bridge 17 upvotes, #21 of 2025-05-27
- REARANK: Reasoning Re-ranking Agent via Reinforcement Learning 17 upvotes, #21 of 2025-05-27
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration 17 upvotes, #21 of 2025-05-27
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions 15 upvotes, #25 of 2025-05-27
- AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting 14 upvotes, #26 of 2025-05-27
- Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI 14 upvotes, #26 of 2025-05-27
- WHISTRESS: Enriching Transcriptions with Sentence Stress Detection 13 upvotes, #28 of 2025-05-27
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression 13 upvotes, #28 of 2025-05-27
- Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition 13 upvotes, #28 of 2025-05-27
- G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning 12 upvotes, #31 of 2025-05-27
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation 12 upvotes, #31 of 2025-05-27
- Interleaved Reasoning for Large Language Models via Reinforcement Learning 12 upvotes, #31 of 2025-05-27
- Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals 11 upvotes, #34 of 2025-05-27
- Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models 11 upvotes, #34 of 2025-05-27
- InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction 10 upvotes, #36 of 2025-05-27
- MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research 10 upvotes, #36 of 2025-05-27
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition 9 upvotes, #38 of 2025-05-27
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference 9 upvotes, #38 of 2025-05-27
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs 8 upvotes, #40 of 2025-05-27
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models 8 upvotes, #40 of 2025-05-27
- Dynamic Risk Assessments for Offensive Cybersecurity Agents 7 upvotes, #42 of 2025-05-27
- Strong Membership Inference Attacks on Massive Datasets and (Moderately) Large Language Models 7 upvotes, #42 of 2025-05-27
- The Coverage Principle: A Framework for Understanding Compositional Generalization 7 upvotes, #42 of 2025-05-27
- Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary? 6 upvotes, #45 of 2025-05-27
- Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective 6 upvotes, #45 of 2025-05-27
- Accelerating Nash Learning from Human Feedback via Mirror Prox 6 upvotes, #45 of 2025-05-27
- Hybrid Latent Reasoning via Reinforcement Learning 5 upvotes, #48 of 2025-05-27
- DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical Dialogue 5 upvotes, #48 of 2025-05-27
- Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs 5 upvotes, #48 of 2025-05-27
- Bridging Supervised Learning and Reinforcement Learning in Math Reasoning 4 upvotes, #51 of 2025-05-27
- An Embarrassingly Simple Defense Against LLM Abliteration Attacks 4 upvotes, #51 of 2025-05-27
- GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes 4 upvotes, #51 of 2025-05-27
- EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning 3 upvotes, #54 of 2025-05-27
- UFT: Unifying Supervised and Reinforcement Fine-Tuning 3 upvotes, #54 of 2025-05-27
- Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation 3 upvotes, #54 of 2025-05-27
- Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision 3 upvotes, #54 of 2025-05-27
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey 2 upvotes, #58 of 2025-05-27
- TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification 2 upvotes, #58 of 2025-05-27
- InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning 2 upvotes, #58 of 2025-05-27
- The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models 2 upvotes, #58 of 2025-05-27
- MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models 2 upvotes, #58 of 2025-05-27
- FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models 2 upvotes, #58 of 2025-05-27
- Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models 2 upvotes, #58 of 2025-05-27
- DiSA: Diffusion Step Annealing in Autoregressive Image Generation 2 upvotes, #58 of 2025-05-27
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning 1 upvotes, #66 of 2025-05-27
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models 1 upvotes, #66 of 2025-05-27
- CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark 1 upvotes, #66 of 2025-05-27
- The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models 1 upvotes, #66 of 2025-05-27
- MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs 1 upvotes, #66 of 2025-05-27
- EgoZero: Robot Learning from Smart Glasses 1 upvotes, #66 of 2025-05-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.