Daily Papers of 2026-02-09
- F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare 70 upvotes, #1 of 2026-02-09
- AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders 59 upvotes, #2 of 2026-02-09
- Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making 59 upvotes, #2 of 2026-02-09
- OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions 57 upvotes, #4 of 2026-02-09
- On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models 52 upvotes, #5 of 2026-02-09
- Pisets: A Robust Speech Recognition System for Lectures and Interviews 33 upvotes, #6 of 2026-02-09
- MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration 32 upvotes, #7 of 2026-02-09
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos 30 upvotes, #8 of 2026-02-09
- Self-Improving World Modelling with Latent Actions 29 upvotes, #9 of 2026-02-09
- Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math 22 upvotes, #10 of 2026-02-09
- Self-Improving Multilingual Long Reasoning via Translation-Reasoning Integrated Training 18 upvotes, #11 of 2026-02-09
- Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers 18 upvotes, #11 of 2026-02-09
- POINTS-GUI-G: GUI-Grounding Journey 16 upvotes, #13 of 2026-02-09
- Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities 15 upvotes, #14 of 2026-02-09
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments 13 upvotes, #15 of 2026-02-09
- OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention 12 upvotes, #16 of 2026-02-09
- EgoAVU: Egocentric Audio-Visual Understanding 12 upvotes, #16 of 2026-02-09
- InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning 12 upvotes, #16 of 2026-02-09
- Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models 11 upvotes, #19 of 2026-02-09
- Large Language Model Reasoning Failures 11 upvotes, #19 of 2026-02-09
- QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals 9 upvotes, #21 of 2026-02-09
- OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale 9 upvotes, #21 of 2026-02-09
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing 8 upvotes, #23 of 2026-02-09
- RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs 7 upvotes, #24 of 2026-02-09
- compar:IA: The French Government's LLM arena to collect French-language human prompts and preference data 7 upvotes, #24 of 2026-02-09
- ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution 6 upvotes, #26 of 2026-02-09
- SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks 6 upvotes, #26 of 2026-02-09
- Uncovering Cross-Objective Interference in Multi-Objective Alignment 6 upvotes, #26 of 2026-02-09
- PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks 5 upvotes, #29 of 2026-02-09
- SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue 4 upvotes, #30 of 2026-02-09
- Revisiting the Shape Convention of Transformer Language Models 4 upvotes, #30 of 2026-02-09
- SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees 4 upvotes, #30 of 2026-02-09
- Vision Transformer Finetuning Benefits from Non-Smooth Components 4 upvotes, #30 of 2026-02-09
- SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs 3 upvotes, #34 of 2026-02-09
- Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs 2 upvotes, #35 of 2026-02-09
- SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization 2 upvotes, #35 of 2026-02-09
- Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion 2 upvotes, #35 of 2026-02-09
- Learning a Generative Meta-Model of LLM Activations 2 upvotes, #35 of 2026-02-09
- Avoiding Premature Collapse: Adaptive Annealing for Entropy-Regularized Structural Inference 1 upvotes, #39 of 2026-02-09
- AtlasPatch: An Efficient and Scalable Tool for Whole Slide Image Preprocessing in Computational Pathology 1 upvotes, #39 of 2026-02-09
- Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search 1 upvotes, #39 of 2026-02-09
- Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction 1 upvotes, #39 of 2026-02-09
- Uncertainty Drives Social Bias Changes in Quantized Large Language Models 1 upvotes, #39 of 2026-02-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.