Daily Papers of 2026-02-18
- GLM-5: from Vibe Coding to Agentic Engineering 94 upvotes, #1 of 2026-02-18
- Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines? 55 upvotes, #2 of 2026-02-18
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks 51 upvotes, #3 of 2026-02-18
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook 26 upvotes, #4 of 2026-02-18
- A Trajectory-Based Safety Audit of Clawdbot (OpenClaw) 22 upvotes, #5 of 2026-02-18
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation 21 upvotes, #6 of 2026-02-18
- ResearchGym: Evaluating Language Model Agents on Real-World AI Research 20 upvotes, #7 of 2026-02-18
- UniT: Unified Multimodal Chain-of-Thought Test-time Scaling 19 upvotes, #8 of 2026-02-18
- Revisiting the Platonic Representation Hypothesis: An Aristotelian View 11 upvotes, #9 of 2026-02-18
- Geometry-Aware Rotary Position Embedding for Consistent Video World Model 9 upvotes, #10 of 2026-02-18
- On Surprising Effectiveness of Masking Updates in Adaptive Optimizers 9 upvotes, #10 of 2026-02-18
- COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression 7 upvotes, #12 of 2026-02-18
- Panini: Continual Learning in Token Space via Structured Memory 6 upvotes, #13 of 2026-02-18
- TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models 6 upvotes, #13 of 2026-02-18
- Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal Models 6 upvotes, #13 of 2026-02-18
- Causal-JEPA: Learning World Models through Object-Level Latent Interventions 5 upvotes, #16 of 2026-02-18
- Learning Native Continuation for Action Chunking Flow Policies 3 upvotes, #17 of 2026-02-18
- Visual Persuasion: What Influences Decisions of Vision-Language Models? 3 upvotes, #17 of 2026-02-18
- STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens 3 upvotes, #17 of 2026-02-18
- ClinAlign: Scaling Healthcare Alignment from Clinician Preference 2 upvotes, #20 of 2026-02-18
- HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam 2 upvotes, #20 of 2026-02-18
- Prescriptive Scaling Reveals the Evolution of Language Model Capabilities 2 upvotes, #20 of 2026-02-18
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems 2 upvotes, #20 of 2026-02-18
- How Much Reasoning Do Retrieval-Augmented Models Add beyond LLMs? A Benchmarking Framework for Multi-Hop Inference over Hybrid Knowledge 1 upvotes, #24 of 2026-02-18
- Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation 1 upvotes, #24 of 2026-02-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.