Daily Papers of 2026-05-14
- MinT: Managed Infrastructure for Training and Serving Millions of LLMs 216 upvotes, #1 of 2026-05-14
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image 138 upvotes, #2 of 2026-05-14
- AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation 96 upvotes, #3 of 2026-05-14
- Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context 85 upvotes, #4 of 2026-05-14
- EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents 64 upvotes, #5 of 2026-05-14
- Qwen-Image-VAE-2.0 Technical Report 58 upvotes, #6 of 2026-05-14
- Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling 48 upvotes, #7 of 2026-05-14
- TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking 37 upvotes, #8 of 2026-05-14
- Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling 32 upvotes, #9 of 2026-05-14
- Many-Shot CoT-ICL: Making In-Context Learning Truly Learn 32 upvotes, #9 of 2026-05-14
- The DAWN of World-Action Interactive Models 22 upvotes, #11 of 2026-05-14
- Asymmetric Flow Models 21 upvotes, #12 of 2026-05-14
- FrameSkip: Learning from Fewer but More Informative Frames in VLA Training 21 upvotes, #12 of 2026-05-14
- KL for a KL: On-Policy Distillation with Control Variate Baseline 19 upvotes, #14 of 2026-05-14
- HAGE: Harnessing Agentic Memory via RL-Driven Weighted Graph Evolution 15 upvotes, #15 of 2026-05-14
- Learning Agentic Policy from Action Guidance 12 upvotes, #16 of 2026-05-14
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion 12 upvotes, #16 of 2026-05-14
- Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty? 11 upvotes, #18 of 2026-05-14
- Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge 10 upvotes, #19 of 2026-05-14
- RewardHarness: Self-Evolving Agentic Post-Training 9 upvotes, #20 of 2026-05-14
- Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation 9 upvotes, #20 of 2026-05-14
- PresentAgent-2: Towards Generalist Multimodal Presentation Agents 8 upvotes, #22 of 2026-05-14
- MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning 8 upvotes, #22 of 2026-05-14
- RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation 8 upvotes, #22 of 2026-05-14
- FeatCal: Feature Calibration for Post-Merging Models 7 upvotes, #25 of 2026-05-14
- Context Training with Active Information Seeking 7 upvotes, #25 of 2026-05-14
- RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data 7 upvotes, #25 of 2026-05-14
- Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs 6 upvotes, #28 of 2026-05-14
- Revisiting DAgger in the Era of LLM-Agents 6 upvotes, #28 of 2026-05-14
- Retrieval from Within: An Intrinsic Capability of Attention-Based Models 5 upvotes, #30 of 2026-05-14
- LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models 5 upvotes, #30 of 2026-05-14
- BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Data 5 upvotes, #30 of 2026-05-14
- MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading 4 upvotes, #33 of 2026-05-14
- FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation 3 upvotes, #34 of 2026-05-14
- MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching 3 upvotes, #34 of 2026-05-14
- The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs 3 upvotes, #34 of 2026-05-14
- An Empirical Study of Automating Agent Evaluation 3 upvotes, #34 of 2026-05-14
- Position: LLM Inference Should Be Evaluated as Energy-to-Token Production 3 upvotes, #34 of 2026-05-14
- AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation 3 upvotes, #34 of 2026-05-14
- Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition 3 upvotes, #34 of 2026-05-14
- Towards Self-Evolving Agentic Literature Retrieval 3 upvotes, #34 of 2026-05-14
- SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety 2 upvotes, #42 of 2026-05-14
- AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents 2 upvotes, #42 of 2026-05-14
- Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection 2 upvotes, #42 of 2026-05-14
- Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization 2 upvotes, #42 of 2026-05-14
- From Pixels to Concepts: Do Segmentation Models Understand What They Segment? 2 upvotes, #42 of 2026-05-14
- From Generalist to Specialist Representation 2 upvotes, #42 of 2026-05-14
- F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking 2 upvotes, #42 of 2026-05-14
- IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages 2 upvotes, #42 of 2026-05-14
- PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents 2 upvotes, #42 of 2026-05-14
- Federation of Experts: Communication Efficient Distributed Inference for Large Language Models 1 upvotes, #51 of 2026-05-14
- ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes 1 upvotes, #51 of 2026-05-14
- M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement 1 upvotes, #51 of 2026-05-14
- Frequency Bias and OOD Generalization in Neural Operators under a Variable-Coefficient Wave Equation 1 upvotes, #51 of 2026-05-14
- Active Tabular Augmentation via Policy-Guided Diffusion Inpainting 0 upvotes, #55 of 2026-05-14
- WriteSAE: Sparse Autoencoders for Recurrent State 0 upvotes, #55 of 2026-05-14
- FlowCompile: An Optimizing Compiler for Structured LLM Workflows 1 upvotes, #55 of 2026-05-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.