Daily Papers of 2026-02-18

  1. GLM-5: from Vibe Coding to Agentic Engineering 94 upvotes, #1 of 2026-02-18
  2. Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines? 55 upvotes, #2 of 2026-02-18
  3. SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks 51 upvotes, #3 of 2026-02-18
  4. Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook 26 upvotes, #4 of 2026-02-18
  5. A Trajectory-Based Safety Audit of Clawdbot (OpenClaw) 22 upvotes, #5 of 2026-02-18
  6. jina-embeddings-v5-text: Task-Targeted Embedding Distillation 21 upvotes, #6 of 2026-02-18
  7. ResearchGym: Evaluating Language Model Agents on Real-World AI Research 20 upvotes, #7 of 2026-02-18
  8. UniT: Unified Multimodal Chain-of-Thought Test-time Scaling 19 upvotes, #8 of 2026-02-18
  9. Revisiting the Platonic Representation Hypothesis: An Aristotelian View 11 upvotes, #9 of 2026-02-18
  10. Geometry-Aware Rotary Position Embedding for Consistent Video World Model 9 upvotes, #10 of 2026-02-18
  11. On Surprising Effectiveness of Masking Updates in Adaptive Optimizers 9 upvotes, #10 of 2026-02-18
  12. COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression 7 upvotes, #12 of 2026-02-18
  13. Panini: Continual Learning in Token Space via Structured Memory 6 upvotes, #13 of 2026-02-18
  14. TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models 6 upvotes, #13 of 2026-02-18
  15. Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal Models 6 upvotes, #13 of 2026-02-18
  16. Causal-JEPA: Learning World Models through Object-Level Latent Interventions 5 upvotes, #16 of 2026-02-18
  17. Learning Native Continuation for Action Chunking Flow Policies 3 upvotes, #17 of 2026-02-18
  18. Visual Persuasion: What Influences Decisions of Vision-Language Models? 3 upvotes, #17 of 2026-02-18
  19. STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens 3 upvotes, #17 of 2026-02-18
  20. ClinAlign: Scaling Healthcare Alignment from Clinician Preference 2 upvotes, #20 of 2026-02-18
  21. HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam 2 upvotes, #20 of 2026-02-18
  22. Prescriptive Scaling Reveals the Evolution of Language Model Capabilities 2 upvotes, #20 of 2026-02-18
  23. The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems 2 upvotes, #20 of 2026-02-18
  24. How Much Reasoning Do Retrieval-Augmented Models Add beyond LLMs? A Benchmarking Framework for Multi-Hop Inference over Hybrid Knowledge 1 upvotes, #24 of 2026-02-18
  25. Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation 1 upvotes, #24 of 2026-02-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.