Daily Papers of 2026-02-09

  1. F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare 70 upvotes, #1 of 2026-02-09
  2. AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders 59 upvotes, #2 of 2026-02-09
  3. Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making 59 upvotes, #2 of 2026-02-09
  4. OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions 57 upvotes, #4 of 2026-02-09
  5. On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models 52 upvotes, #5 of 2026-02-09
  6. Pisets: A Robust Speech Recognition System for Lectures and Interviews 33 upvotes, #6 of 2026-02-09
  7. MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration 32 upvotes, #7 of 2026-02-09
  8. DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos 30 upvotes, #8 of 2026-02-09
  9. Self-Improving World Modelling with Latent Actions 29 upvotes, #9 of 2026-02-09
  10. Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math 22 upvotes, #10 of 2026-02-09
  11. Self-Improving Multilingual Long Reasoning via Translation-Reasoning Integrated Training 18 upvotes, #11 of 2026-02-09
  12. Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers 18 upvotes, #11 of 2026-02-09
  13. POINTS-GUI-G: GUI-Grounding Journey 16 upvotes, #13 of 2026-02-09
  14. Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities 15 upvotes, #14 of 2026-02-09
  15. MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments 13 upvotes, #15 of 2026-02-09
  16. OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention 12 upvotes, #16 of 2026-02-09
  17. EgoAVU: Egocentric Audio-Visual Understanding 12 upvotes, #16 of 2026-02-09
  18. InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning 12 upvotes, #16 of 2026-02-09
  19. Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models 11 upvotes, #19 of 2026-02-09
  20. Large Language Model Reasoning Failures 11 upvotes, #19 of 2026-02-09
  21. QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals 9 upvotes, #21 of 2026-02-09
  22. OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale 9 upvotes, #21 of 2026-02-09
  23. Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing 8 upvotes, #23 of 2026-02-09
  24. RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs 7 upvotes, #24 of 2026-02-09
  25. compar:IA: The French Government's LLM arena to collect French-language human prompts and preference data 7 upvotes, #24 of 2026-02-09
  26. ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution 6 upvotes, #26 of 2026-02-09
  27. SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks 6 upvotes, #26 of 2026-02-09
  28. Uncovering Cross-Objective Interference in Multi-Objective Alignment 6 upvotes, #26 of 2026-02-09
  29. PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks 5 upvotes, #29 of 2026-02-09
  30. SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue 4 upvotes, #30 of 2026-02-09
  31. Revisiting the Shape Convention of Transformer Language Models 4 upvotes, #30 of 2026-02-09
  32. SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees 4 upvotes, #30 of 2026-02-09
  33. Vision Transformer Finetuning Benefits from Non-Smooth Components 4 upvotes, #30 of 2026-02-09
  34. SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs 3 upvotes, #34 of 2026-02-09
  35. Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs 2 upvotes, #35 of 2026-02-09
  36. SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization 2 upvotes, #35 of 2026-02-09
  37. Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion 2 upvotes, #35 of 2026-02-09
  38. Learning a Generative Meta-Model of LLM Activations 2 upvotes, #35 of 2026-02-09
  39. Avoiding Premature Collapse: Adaptive Annealing for Entropy-Regularized Structural Inference 1 upvotes, #39 of 2026-02-09
  40. AtlasPatch: An Efficient and Scalable Tool for Whole Slide Image Preprocessing in Computational Pathology 1 upvotes, #39 of 2026-02-09
  41. Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search 1 upvotes, #39 of 2026-02-09
  42. Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction 1 upvotes, #39 of 2026-02-09
  43. Uncertainty Drives Social Bias Changes in Quantized Large Language Models 1 upvotes, #39 of 2026-02-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.