Daily Papers of 2026-02-12

  1. Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters 177 upvotes, #1 of 2026-02-12
  2. VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval 120 upvotes, #2 of 2026-02-12
  3. GENIUS: Generative Fluid Intelligence Evaluation Suite 53 upvotes, #3 of 2026-02-12
  4. PhyCritic: Multimodal Critic Models for Physical AI 51 upvotes, #4 of 2026-02-12
  5. ASA: Training-Free Representation Engineering for Tool-Calling Agents 39 upvotes, #5 of 2026-02-12
  6. Towards Autonomous Mathematics Research 35 upvotes, #6 of 2026-02-12
  7. When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning 28 upvotes, #7 of 2026-02-12
  8. TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions 27 upvotes, #8 of 2026-02-12
  9. How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning 27 upvotes, #8 of 2026-02-12
  10. G-LNS: Generative Large Neighborhood Search for LLM-Based Automatic Heuristic Design 25 upvotes, #10 of 2026-02-12
  11. Internalizing Meta-Experience into Memory for Guided Reinforcement Learning in Large Language Models 19 upvotes, #11 of 2026-02-12
  12. FeatureBench: Benchmarking Agentic Coding for Complex Feature Development 18 upvotes, #12 of 2026-02-12
  13. DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning 18 upvotes, #12 of 2026-02-12
  14. ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression 17 upvotes, #14 of 2026-02-12
  15. Online Causal Kalman Filtering for Stable and Effective Policy Optimization 16 upvotes, #15 of 2026-02-12
  16. GameDevBench: Evaluating Agentic Capabilities Through Game Development 15 upvotes, #16 of 2026-02-12
  17. LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation 15 upvotes, #16 of 2026-02-12
  18. LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation 13 upvotes, #18 of 2026-02-12
  19. The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context 13 upvotes, #18 of 2026-02-12
  20. ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning 12 upvotes, #20 of 2026-02-12
  21. Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards 12 upvotes, #20 of 2026-02-12
  22. Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning 12 upvotes, #20 of 2026-02-12
  23. Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models 11 upvotes, #23 of 2026-02-12
  24. CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion 10 upvotes, #24 of 2026-02-12
  25. Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL 9 upvotes, #25 of 2026-02-12
  26. EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies 9 upvotes, #25 of 2026-02-12
  27. Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models 8 upvotes, #27 of 2026-02-12
  28. QP-OneModel: A Unified Generative LLM for Multi-Task Query Understanding in Xiaohongshu Search 6 upvotes, #28 of 2026-02-12
  29. When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models 6 upvotes, #28 of 2026-02-12
  30. Benchmarking Large Language Models for Knowledge Graph Validation 6 upvotes, #28 of 2026-02-12
  31. Free(): Learning to Forget in Malloc-Only Reasoning Models 5 upvotes, #31 of 2026-02-12
  32. Beyond Correctness: Learning Robust Reasoning via Transfer 5 upvotes, #31 of 2026-02-12
  33. Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent Tokens 5 upvotes, #31 of 2026-02-12
  34. AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions 4 upvotes, #34 of 2026-02-12
  35. Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents 4 upvotes, #34 of 2026-02-12
  36. Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation 4 upvotes, #34 of 2026-02-12
  37. TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments 3 upvotes, #37 of 2026-02-12
  38. ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation 3 upvotes, #37 of 2026-02-12
  39. Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing 2 upvotes, #39 of 2026-02-12
  40. When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents 2 upvotes, #39 of 2026-02-12
  41. UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory 2 upvotes, #39 of 2026-02-12
  42. GoodVibe: Security-by-Vibe for LLM-Based Code Generation 2 upvotes, #39 of 2026-02-12
  43. Weight Decay Improves Language Model Plasticity 2 upvotes, #39 of 2026-02-12
  44. Graph-Enhanced Deep Reinforcement Learning for Multi-Objective Unrelated Parallel Machine Scheduling 1 upvotes, #44 of 2026-02-12
  45. Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative Recommendation 1 upvotes, #44 of 2026-02-12
  46. FedPS: Federated data Preprocessing via aggregated Statistics 1 upvotes, #44 of 2026-02-12
  47. From Features to Actions: Explainability in Traditional and Agentic AI Systems 1 upvotes, #47 of 2026-02-12
  48. StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors 0 upvotes, #47 of 2026-02-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.