Daily Papers of 2026-04-08

  1. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
  2. Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents 114 upvotes, #2 of 2026-04-08
  3. Learning to Retrieve from Agent Trajectories 69 upvotes, #3 of 2026-04-08
  4. ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation 53 upvotes, #4 of 2026-04-08
  5. GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers 45 upvotes, #5 of 2026-04-08
  6. MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU 45 upvotes, #5 of 2026-04-08
  7. Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision 42 upvotes, #7 of 2026-04-08
  8. Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning 41 upvotes, #8 of 2026-04-08
  9. ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement 40 upvotes, #9 of 2026-04-08
  10. How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings 40 upvotes, #9 of 2026-04-08
  11. Watch Before You Answer: Learning from Visually Grounded Post-Training 35 upvotes, #11 of 2026-04-08
  12. General Multimodal Protein Design Enables DNA-Encoding of Chemistry 31 upvotes, #12 of 2026-04-08
  13. Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework 30 upvotes, #13 of 2026-04-08
  14. In-Place Test-Time Training 28 upvotes, #14 of 2026-04-08
  15. ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces 24 upvotes, #15 of 2026-04-08
  16. DARE: Diffusion Large Language Models Alignment and Reinforcement Executor 21 upvotes, #16 of 2026-04-08
  17. Demystifying When Pruning Works via Representation Hierarchies 20 upvotes, #17 of 2026-04-08
  18. Experience Transfer for Multimodal LLM Agents in Minecraft Game 15 upvotes, #18 of 2026-04-08
  19. MedGemma 1.5 Technical Report 14 upvotes, #19 of 2026-04-08
  20. Action Images: End-to-End Policy Learning via Multiview Video Generation 14 upvotes, #19 of 2026-04-08
  21. Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents 10 upvotes, #21 of 2026-04-08
  22. MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control 10 upvotes, #21 of 2026-04-08
  23. Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling? 8 upvotes, #23 of 2026-04-08
  24. REAM: Merging Improves Pruning of Experts in LLMs 8 upvotes, #23 of 2026-04-08
  25. Context-Value-Action Architecture for Value-Driven Large Language Model Agents 8 upvotes, #23 of 2026-04-08
  26. QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization 8 upvotes, #23 of 2026-04-08
  27. Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models 7 upvotes, #27 of 2026-04-08
  28. FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification 7 upvotes, #27 of 2026-04-08
  29. CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation 7 upvotes, #27 of 2026-04-08
  30. Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning 7 upvotes, #27 of 2026-04-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.