Daily Papers of 2026-04-02
- ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers 181 upvotes, #1 of 2026-04-02
- Terminal Agents Suffice for Enterprise Automation 92 upvotes, #2 of 2026-04-02
- MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome 68 upvotes, #3 of 2026-04-02
- ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners? 43 upvotes, #4 of 2026-04-02
- Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification 42 upvotes, #5 of 2026-04-02
- Embarrassingly Simple Self-Distillation Improves Code Generation 33 upvotes, #6 of 2026-04-02
- QuitoBench: A High-Quality Open Time Series Forecasting Benchmark 31 upvotes, #7 of 2026-04-02
- Reasoning Shift: How Context Silently Shortens LLM Reasoning 29 upvotes, #8 of 2026-04-02
- HippoCamp: Benchmarking Contextual Agents on Personal Computers 27 upvotes, #9 of 2026-04-02
- GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation 25 upvotes, #10 of 2026-04-02
- Brevity Constraints Reverse Performance Hierarchies in Language Models 21 upvotes, #11 of 2026-04-02
- PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning 18 upvotes, #12 of 2026-04-02
- Universal YOCO for Efficient Depth Scaling 17 upvotes, #13 of 2026-04-02
- Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers 15 upvotes, #14 of 2026-04-02
- Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants 13 upvotes, #15 of 2026-04-02
- Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding 12 upvotes, #16 of 2026-04-02
- A Survey of On-Policy Distillation for Large Language Models 9 upvotes, #17 of 2026-04-02
- Do Phone-Use Agents Respect Your Privacy? 9 upvotes, #17 of 2026-04-02
- UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems 8 upvotes, #19 of 2026-04-02
- Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines 7 upvotes, #20 of 2026-04-02
- Understand and Accelerate Memory Processing Pipeline for Disaggregated LLM Inference 6 upvotes, #21 of 2026-04-02
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding 6 upvotes, #21 of 2026-04-02
- S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models 6 upvotes, #21 of 2026-04-02
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation 5 upvotes, #24 of 2026-04-02
- Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy 5 upvotes, #24 of 2026-04-02
- MemRerank: Preference Memory for Personalized Product Reranking 5 upvotes, #24 of 2026-04-02
- When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation 5 upvotes, #24 of 2026-04-02
- Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment 4 upvotes, #28 of 2026-04-02
- AI Generalisation Gap In Comorbid Sleep Disorder Staging 3 upvotes, #29 of 2026-04-02
- AgentWatcher: A Rule-based Prompt Injection Monitor 3 upvotes, #29 of 2026-04-02
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.