Daily Papers of 2026-02-25
- On Data Engineering for Scaling LLM Terminal Capabilities 90 upvotes, #1 of 2026-02-25
- Query-focused and Memory-aware Reranker for Long Context Processing 55 upvotes, #2 of 2026-02-25
- PyVision-RL: Forging Open Agentic Vision Models via RL 29 upvotes, #3 of 2026-02-25
- Test-Time Training with KV Binding Is Secretly Linear Attention 29 upvotes, #3 of 2026-02-25
- From Perception to Action: An Interactive Benchmark for Vision Reasoning 22 upvotes, #5 of 2026-02-25
- Multi-Vector Index Compression in Any Modality 22 upvotes, #5 of 2026-02-25
- QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models 16 upvotes, #7 of 2026-02-25
- DREAM: Deep Research Evaluation with Agentic Metrics 14 upvotes, #8 of 2026-02-25
- See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis 13 upvotes, #9 of 2026-02-25
- LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces 12 upvotes, #10 of 2026-02-25
- Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation 11 upvotes, #11 of 2026-02-25
- PETS: A Principled Framework Towards Optimal Trajectory Allocation for Efficient Test-Time Self-Consistency 8 upvotes, #12 of 2026-02-25
- Benchmark Test-Time Scaling of General LLM Agents 8 upvotes, #12 of 2026-02-25
- TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model Agents 7 upvotes, #14 of 2026-02-25
- RankEvolve: Automating the Discovery of Retrieval Algorithms via LLM-Driven Evolution 6 upvotes, #15 of 2026-02-25
- The Art of Efficient Reasoning: Data, Reward, and Optimization 6 upvotes, #15 of 2026-02-25
- Aletheia tackles FirstProof autonomously 5 upvotes, #17 of 2026-02-25
- One-step Language Modeling via Continuous Denoising 4 upvotes, #18 of 2026-02-25
- Implicit Intelligence -- Evaluating Agents on What Users Don't Say 4 upvotes, #18 of 2026-02-25
- Generative AI and Machine Learning Collaboration for Container Dwell Time Prediction via Data Standardization 4 upvotes, #18 of 2026-02-25
- Communication-Inspired Tokenization for Structured Image Representations 4 upvotes, #18 of 2026-02-25
- Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs 4 upvotes, #18 of 2026-02-25
- The Diffusion Duality, Chapter II: Ψ-Samplers and Efficient Curriculum 3 upvotes, #23 of 2026-02-25
- Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking 3 upvotes, #23 of 2026-02-25
- Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization 2 upvotes, #25 of 2026-02-25
- SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking 2 upvotes, #25 of 2026-02-25
- OmniOCR: Generalist OCR for Ethnic Minority Languages 2 upvotes, #25 of 2026-02-25
- OCR-Agent: Agentic OCR with Capability and Memory Reflection 2 upvotes, #25 of 2026-02-25
- FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving 1 upvotes, #29 of 2026-02-25
- LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency 1 upvotes, #29 of 2026-02-25
- Learning to Detect Language Model Training Data via Active Reconstruction 1 upvotes, #29 of 2026-02-25
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering 1 upvotes, #32 of 2026-02-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.