Daily Papers of 2026-06-18
- MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction 50 upvotes, #1 of 2026-06-18
- Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games 46 upvotes, #2 of 2026-06-18
- Kairos: A Native World Model Stack for Physical AI 37 upvotes, #3 of 2026-06-18
- Guava: An Effective and Universal Harness for Embodied Manipulation 28 upvotes, #4 of 2026-06-18
- From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning 26 upvotes, #5 of 2026-06-18
- EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts 24 upvotes, #6 of 2026-06-18
- The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL 20 upvotes, #7 of 2026-06-18
- SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior 17 upvotes, #8 of 2026-06-18
- Native Active Perception as Reasoning for Omni-Modal Understanding 17 upvotes, #8 of 2026-06-18
- Reinforcing Dual-Path Reasoning in Spatial Vision Language Models 15 upvotes, #10 of 2026-06-18
- Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding 15 upvotes, #10 of 2026-06-18
- MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model 13 upvotes, #12 of 2026-06-18
- STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability 12 upvotes, #13 of 2026-06-18
- PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation 11 upvotes, #14 of 2026-06-18
- Sumi: Open Uniform Diffusion Language Model from Scratch 11 upvotes, #14 of 2026-06-18
- Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems 10 upvotes, #16 of 2026-06-18
- ViT-Up: Faithful Feature Upsampling for Vision Transformers 9 upvotes, #17 of 2026-06-18
- SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks 8 upvotes, #18 of 2026-06-18
- CEO-Bench: Can Agents Play the Long Game? 8 upvotes, #18 of 2026-06-18
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents 6 upvotes, #20 of 2026-06-18
- Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish 6 upvotes, #20 of 2026-06-18
- Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness 6 upvotes, #20 of 2026-06-18
- Physics-IQ Verified 5 upvotes, #23 of 2026-06-18
- HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing 4 upvotes, #24 of 2026-06-18
- IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products 4 upvotes, #24 of 2026-06-18
- When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? 4 upvotes, #24 of 2026-06-18
- RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents 4 upvotes, #24 of 2026-06-18
- Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning 4 upvotes, #24 of 2026-06-18
- iOSWorld: A Benchmark for Personally Intelligent Phone Agents 3 upvotes, #29 of 2026-06-18
- LLM-Enabled NWDAF: A Step Toward AI-Native 6G Network Intelligence 3 upvotes, #29 of 2026-06-18
- Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities 3 upvotes, #29 of 2026-06-18
- REVES: REvision and VErification--Augmented Training for Test-Time Scaling 3 upvotes, #29 of 2026-06-18
- Learning User Simulators with Turing Rewards 3 upvotes, #29 of 2026-06-18
- Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns 2 upvotes, #34 of 2026-06-18
- Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation 2 upvotes, #34 of 2026-06-18
- Re-Centering Humans in LLM Personalization 1 upvotes, #36 of 2026-06-18
- A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets 1 upvotes, #36 of 2026-06-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.