Daily Papers of 2026-02-06

  1. CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty 79 upvotes, #1 of 2026-02-06
  2. Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening 69 upvotes, #2 of 2026-02-06
  3. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents 53 upvotes, #3 of 2026-02-06
  4. Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR 48 upvotes, #4 of 2026-02-06
  5. DFlash: Block Diffusion for Flash Speculative Decoding 41 upvotes, #5 of 2026-02-06
  6. Context Forcing: Consistent Autoregressive Video Generation with Long Context 35 upvotes, #6 of 2026-02-06
  7. Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations 28 upvotes, #7 of 2026-02-06
  8. Reinforced Attention Learning 27 upvotes, #8 of 2026-02-06
  9. Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention 26 upvotes, #9 of 2026-02-06
  10. RISE-Video: Can Video Generators Decode Implicit World Rules? 26 upvotes, #9 of 2026-02-06
  11. Privileged Information Distillation for Language Models 25 upvotes, #11 of 2026-02-06
  12. ProAct: Agentic Lookahead in Interactive Environments 25 upvotes, #11 of 2026-02-06
  13. Reinforcement World Model Learning for LLM-based Agents 25 upvotes, #11 of 2026-02-06
  14. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions 22 upvotes, #14 of 2026-02-06
  15. Grounding and Enhancing Informativeness and Utility in Dataset Distillation 19 upvotes, #15 of 2026-02-06
  16. Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities 19 upvotes, #15 of 2026-02-06
  17. Semantic Search over 9 Million Mathematical Theorems 19 upvotes, #15 of 2026-02-06
  18. Steering LLMs via Scalable Interactive Oversight 18 upvotes, #18 of 2026-02-06
  19. SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers 18 upvotes, #18 of 2026-02-06
  20. Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory 16 upvotes, #20 of 2026-02-06
  21. Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning 15 upvotes, #21 of 2026-02-06
  22. LatentMem: Customizing Latent Memory for Multi-Agent Systems 14 upvotes, #22 of 2026-02-06
  23. Multi-Task GRPO: Reliable LLM Reasoning Across Tasks 12 upvotes, #23 of 2026-02-06
  24. SAGE: Benchmarking and Improving Retrieval for Deep Research Agents 12 upvotes, #23 of 2026-02-06
  25. DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers 11 upvotes, #25 of 2026-02-06
  26. Towards Reducible Uncertainty Modeling for Reliable Large Language Model Agents 11 upvotes, #25 of 2026-02-06
  27. BABE: Biology Arena BEnchmark 10 upvotes, #27 of 2026-02-06
  28. SwimBird: Eliciting Switchable Reasoning Mode in Hybrid Autoregressive MLLMs 10 upvotes, #27 of 2026-02-06
  29. V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval 8 upvotes, #29 of 2026-02-06
  30. CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs 7 upvotes, #30 of 2026-02-06
  31. Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better 7 upvotes, #30 of 2026-02-06
  32. Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning 6 upvotes, #32 of 2026-02-06
  33. Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training 5 upvotes, #33 of 2026-02-06
  34. Breaking the Static Graph: Context-Aware Traversal for Robust Retrieval-Augmented Generation 4 upvotes, #34 of 2026-02-06
  35. Adaptive 1D Video Diffusion Autoencoder 4 upvotes, #34 of 2026-02-06
  36. Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization 3 upvotes, #36 of 2026-02-06
  37. Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention 3 upvotes, #36 of 2026-02-06
  38. FastVMT: Eliminating Redundancy in Video Motion Transfer 3 upvotes, #36 of 2026-02-06
  39. Pathwise Test-Time Correction for Autoregressive Long Video Generation 3 upvotes, #36 of 2026-02-06
  40. Failing to Explore: Language Models on Interactive Tasks 2 upvotes, #40 of 2026-02-06
  41. UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization 2 upvotes, #40 of 2026-02-06
  42. Do Vision-Language Models Respect Contextual Integrity in Location Disclosure? 2 upvotes, #40 of 2026-02-06
  43. Fast-SAM3D: 3Dfy Anything in Images but Faster 2 upvotes, #40 of 2026-02-06
  44. A Unified Framework for Rethinking Policy Divergence Measures in GRPO 2 upvotes, #40 of 2026-02-06
  45. Assessing Domain-Level Susceptibility to Emergent Misalignment from Narrow Finetuning 1 upvotes, #45 of 2026-02-06
  46. Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing 1 upvotes, #45 of 2026-02-06
  47. PhysicsAgentABM: Physics-Guided Generative Agent-Based Modeling 0 upvotes, #47 of 2026-02-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.