Daily Papers of 2026-06-16
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence 200 upvotes, #1 of 2026-06-16
- Geometric Action Model for Robot Policy Learning 112 upvotes, #2 of 2026-06-16
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models 110 upvotes, #3 of 2026-06-16
- DreamX-World 1.0: A General-Purpose Interactive World Model 109 upvotes, #4 of 2026-06-16
- FastContext: Training Efficient Repository Explorer for Coding Agents 91 upvotes, #5 of 2026-06-16
- Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale 80 upvotes, #6 of 2026-06-16
- Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models 33 upvotes, #7 of 2026-06-16
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation 29 upvotes, #8 of 2026-06-16
- VisualClaw: A Real-Time, Personalized Agent for the Physical World 28 upvotes, #9 of 2026-06-16
- BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering 27 upvotes, #10 of 2026-06-16
- BadWorld: Adversarial Attacks on World Models 18 upvotes, #11 of 2026-06-16
- OneRank: Unified Transformer-Native Ranking Architecture for Multi-Task Recommendation 18 upvotes, #11 of 2026-06-16
- Memento: Reconstruct to Remember for Consistent Long Video Generation 16 upvotes, #13 of 2026-06-16
- Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time 16 upvotes, #13 of 2026-06-16
- TokenPilot: Cache-Efficient Context Management for LLM Agents 16 upvotes, #13 of 2026-06-16
- MVEB: Massive Video Embedding Benchmark 15 upvotes, #16 of 2026-06-16
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 15 upvotes, #16 of 2026-06-16
- SP^3: Spherical Priors for Plug-and-Play Restoration 15 upvotes, #16 of 2026-06-16
- UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer 14 upvotes, #19 of 2026-06-16
- CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks? 13 upvotes, #20 of 2026-06-16
- Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking 13 upvotes, #20 of 2026-06-16
- GD^2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization 13 upvotes, #20 of 2026-06-16
- PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions 12 upvotes, #23 of 2026-06-16
- Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks 12 upvotes, #23 of 2026-06-16
- You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences 12 upvotes, #23 of 2026-06-16
- Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving 10 upvotes, #26 of 2026-06-16
- Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes 9 upvotes, #27 of 2026-06-16
- Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders 8 upvotes, #28 of 2026-06-16
- Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning 7 upvotes, #29 of 2026-06-16
- Artificial Intelligence Index Report 2026 5 upvotes, #30 of 2026-06-16
- PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory 5 upvotes, #30 of 2026-06-16
- MMDiff: Extending Diffusion Transformers for Multi-Modal Generation 5 upvotes, #30 of 2026-06-16
- ExpRL: Exploratory RL for LLM Mid-Training 5 upvotes, #30 of 2026-06-16
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies 4 upvotes, #34 of 2026-06-16
- Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs 4 upvotes, #34 of 2026-06-16
- Implicit Reasoning for Large Language Model-based Generative Recommendation 3 upvotes, #36 of 2026-06-16
- Attacks on Machine-Text Detectors Retain Stylistic Fingerprints 2 upvotes, #37 of 2026-06-16
- Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks 2 upvotes, #37 of 2026-06-16
- EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video 2 upvotes, #37 of 2026-06-16
- The Ghosts of Polymarket: When Off-Chain Matches Meet On-Chain Reverts 1 upvotes, #40 of 2026-06-16
- TuneJury: An Open Metric for Improving Music Generation Preference Alignment 1 upvotes, #40 of 2026-06-16
- Human Universal Grasping 1 upvotes, #40 of 2026-06-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.