Daily Papers of 2026-05-14

  1. MinT: Managed Infrastructure for Training and Serving Millions of LLMs 216 upvotes, #1 of 2026-05-14
  2. MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image 138 upvotes, #2 of 2026-05-14
  3. AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation 96 upvotes, #3 of 2026-05-14
  4. Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context 85 upvotes, #4 of 2026-05-14
  5. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents 64 upvotes, #5 of 2026-05-14
  6. Qwen-Image-VAE-2.0 Technical Report 58 upvotes, #6 of 2026-05-14
  7. Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling 48 upvotes, #7 of 2026-05-14
  8. TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking 37 upvotes, #8 of 2026-05-14
  9. Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling 32 upvotes, #9 of 2026-05-14
  10. Many-Shot CoT-ICL: Making In-Context Learning Truly Learn 32 upvotes, #9 of 2026-05-14
  11. The DAWN of World-Action Interactive Models 22 upvotes, #11 of 2026-05-14
  12. Asymmetric Flow Models 21 upvotes, #12 of 2026-05-14
  13. FrameSkip: Learning from Fewer but More Informative Frames in VLA Training 21 upvotes, #12 of 2026-05-14
  14. KL for a KL: On-Policy Distillation with Control Variate Baseline 19 upvotes, #14 of 2026-05-14
  15. HAGE: Harnessing Agentic Memory via RL-Driven Weighted Graph Evolution 15 upvotes, #15 of 2026-05-14
  16. Learning Agentic Policy from Action Guidance 12 upvotes, #16 of 2026-05-14
  17. Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion 12 upvotes, #16 of 2026-05-14
  18. Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty? 11 upvotes, #18 of 2026-05-14
  19. Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge 10 upvotes, #19 of 2026-05-14
  20. RewardHarness: Self-Evolving Agentic Post-Training 9 upvotes, #20 of 2026-05-14
  21. Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation 9 upvotes, #20 of 2026-05-14
  22. PresentAgent-2: Towards Generalist Multimodal Presentation Agents 8 upvotes, #22 of 2026-05-14
  23. MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning 8 upvotes, #22 of 2026-05-14
  24. RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation 8 upvotes, #22 of 2026-05-14
  25. FeatCal: Feature Calibration for Post-Merging Models 7 upvotes, #25 of 2026-05-14
  26. Context Training with Active Information Seeking 7 upvotes, #25 of 2026-05-14
  27. RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data 7 upvotes, #25 of 2026-05-14
  28. Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs 6 upvotes, #28 of 2026-05-14
  29. Revisiting DAgger in the Era of LLM-Agents 6 upvotes, #28 of 2026-05-14
  30. Retrieval from Within: An Intrinsic Capability of Attention-Based Models 5 upvotes, #30 of 2026-05-14
  31. LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models 5 upvotes, #30 of 2026-05-14
  32. BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Data 5 upvotes, #30 of 2026-05-14
  33. MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading 4 upvotes, #33 of 2026-05-14
  34. FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation 3 upvotes, #34 of 2026-05-14
  35. MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching 3 upvotes, #34 of 2026-05-14
  36. The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs 3 upvotes, #34 of 2026-05-14
  37. An Empirical Study of Automating Agent Evaluation 3 upvotes, #34 of 2026-05-14
  38. Position: LLM Inference Should Be Evaluated as Energy-to-Token Production 3 upvotes, #34 of 2026-05-14
  39. AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation 3 upvotes, #34 of 2026-05-14
  40. Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition 3 upvotes, #34 of 2026-05-14
  41. Towards Self-Evolving Agentic Literature Retrieval 3 upvotes, #34 of 2026-05-14
  42. SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety 2 upvotes, #42 of 2026-05-14
  43. AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents 2 upvotes, #42 of 2026-05-14
  44. Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection 2 upvotes, #42 of 2026-05-14
  45. Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization 2 upvotes, #42 of 2026-05-14
  46. From Pixels to Concepts: Do Segmentation Models Understand What They Segment? 2 upvotes, #42 of 2026-05-14
  47. From Generalist to Specialist Representation 2 upvotes, #42 of 2026-05-14
  48. F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking 2 upvotes, #42 of 2026-05-14
  49. IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages 2 upvotes, #42 of 2026-05-14
  50. PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents 2 upvotes, #42 of 2026-05-14
  51. Federation of Experts: Communication Efficient Distributed Inference for Large Language Models 1 upvotes, #51 of 2026-05-14
  52. ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes 1 upvotes, #51 of 2026-05-14
  53. M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement 1 upvotes, #51 of 2026-05-14
  54. Frequency Bias and OOD Generalization in Neural Operators under a Variable-Coefficient Wave Equation 1 upvotes, #51 of 2026-05-14
  55. Active Tabular Augmentation via Policy-Guided Diffusion Inpainting 0 upvotes, #55 of 2026-05-14
  56. WriteSAE: Sparse Autoencoders for Recurrent State 0 upvotes, #55 of 2026-05-14
  57. FlowCompile: An Optimizing Compiler for Structured LLM Workflows 1 upvotes, #55 of 2026-05-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.