Daily Papers of 2026-02-25

  1. On Data Engineering for Scaling LLM Terminal Capabilities 90 upvotes, #1 of 2026-02-25
  2. Query-focused and Memory-aware Reranker for Long Context Processing 55 upvotes, #2 of 2026-02-25
  3. PyVision-RL: Forging Open Agentic Vision Models via RL 29 upvotes, #3 of 2026-02-25
  4. Test-Time Training with KV Binding Is Secretly Linear Attention 29 upvotes, #3 of 2026-02-25
  5. From Perception to Action: An Interactive Benchmark for Vision Reasoning 22 upvotes, #5 of 2026-02-25
  6. Multi-Vector Index Compression in Any Modality 22 upvotes, #5 of 2026-02-25
  7. QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models 16 upvotes, #7 of 2026-02-25
  8. DREAM: Deep Research Evaluation with Agentic Metrics 14 upvotes, #8 of 2026-02-25
  9. See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis 13 upvotes, #9 of 2026-02-25
  10. LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces 12 upvotes, #10 of 2026-02-25
  11. Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation 11 upvotes, #11 of 2026-02-25
  12. PETS: A Principled Framework Towards Optimal Trajectory Allocation for Efficient Test-Time Self-Consistency 8 upvotes, #12 of 2026-02-25
  13. Benchmark Test-Time Scaling of General LLM Agents 8 upvotes, #12 of 2026-02-25
  14. TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model Agents 7 upvotes, #14 of 2026-02-25
  15. RankEvolve: Automating the Discovery of Retrieval Algorithms via LLM-Driven Evolution 6 upvotes, #15 of 2026-02-25
  16. The Art of Efficient Reasoning: Data, Reward, and Optimization 6 upvotes, #15 of 2026-02-25
  17. Aletheia tackles FirstProof autonomously 5 upvotes, #17 of 2026-02-25
  18. One-step Language Modeling via Continuous Denoising 4 upvotes, #18 of 2026-02-25
  19. Implicit Intelligence -- Evaluating Agents on What Users Don't Say 4 upvotes, #18 of 2026-02-25
  20. Generative AI and Machine Learning Collaboration for Container Dwell Time Prediction via Data Standardization 4 upvotes, #18 of 2026-02-25
  21. Communication-Inspired Tokenization for Structured Image Representations 4 upvotes, #18 of 2026-02-25
  22. Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs 4 upvotes, #18 of 2026-02-25
  23. The Diffusion Duality, Chapter II: Ψ-Samplers and Efficient Curriculum 3 upvotes, #23 of 2026-02-25
  24. Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking 3 upvotes, #23 of 2026-02-25
  25. Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization 2 upvotes, #25 of 2026-02-25
  26. SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking 2 upvotes, #25 of 2026-02-25
  27. OmniOCR: Generalist OCR for Ethnic Minority Languages 2 upvotes, #25 of 2026-02-25
  28. OCR-Agent: Agentic OCR with Capability and Memory Reflection 2 upvotes, #25 of 2026-02-25
  29. FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving 1 upvotes, #29 of 2026-02-25
  30. LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency 1 upvotes, #29 of 2026-02-25
  31. Learning to Detect Language Model Training Data via Active Reconstruction 1 upvotes, #29 of 2026-02-25
  32. TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering 1 upvotes, #32 of 2026-02-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.