Daily Papers of 2026-06-16

  1. JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence 200 upvotes, #1 of 2026-06-16
  2. Geometric Action Model for Robot Policy Learning 112 upvotes, #2 of 2026-06-16
  3. VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models 110 upvotes, #3 of 2026-06-16
  4. DreamX-World 1.0: A General-Purpose Interactive World Model 109 upvotes, #4 of 2026-06-16
  5. FastContext: Training Efficient Repository Explorer for Coding Agents 91 upvotes, #5 of 2026-06-16
  6. Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale 80 upvotes, #6 of 2026-06-16
  7. Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models 33 upvotes, #7 of 2026-06-16
  8. Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation 29 upvotes, #8 of 2026-06-16
  9. VisualClaw: A Real-Time, Personalized Agent for the Physical World 28 upvotes, #9 of 2026-06-16
  10. BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering 27 upvotes, #10 of 2026-06-16
  11. BadWorld: Adversarial Attacks on World Models 18 upvotes, #11 of 2026-06-16
  12. OneRank: Unified Transformer-Native Ranking Architecture for Multi-Task Recommendation 18 upvotes, #11 of 2026-06-16
  13. Memento: Reconstruct to Remember for Consistent Long Video Generation 16 upvotes, #13 of 2026-06-16
  14. Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time 16 upvotes, #13 of 2026-06-16
  15. TokenPilot: Cache-Efficient Context Management for LLM Agents 16 upvotes, #13 of 2026-06-16
  16. MVEB: Massive Video Embedding Benchmark 15 upvotes, #16 of 2026-06-16
  17. Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 15 upvotes, #16 of 2026-06-16
  18. SP^3: Spherical Priors for Plug-and-Play Restoration 15 upvotes, #16 of 2026-06-16
  19. UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer 14 upvotes, #19 of 2026-06-16
  20. CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks? 13 upvotes, #20 of 2026-06-16
  21. Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking 13 upvotes, #20 of 2026-06-16
  22. GD^2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization 13 upvotes, #20 of 2026-06-16
  23. PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions 12 upvotes, #23 of 2026-06-16
  24. Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks 12 upvotes, #23 of 2026-06-16
  25. You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences 12 upvotes, #23 of 2026-06-16
  26. Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving 10 upvotes, #26 of 2026-06-16
  27. Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes 9 upvotes, #27 of 2026-06-16
  28. Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders 8 upvotes, #28 of 2026-06-16
  29. Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning 7 upvotes, #29 of 2026-06-16
  30. Artificial Intelligence Index Report 2026 5 upvotes, #30 of 2026-06-16
  31. PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory 5 upvotes, #30 of 2026-06-16
  32. MMDiff: Extending Diffusion Transformers for Multi-Modal Generation 5 upvotes, #30 of 2026-06-16
  33. ExpRL: Exploratory RL for LLM Mid-Training 5 upvotes, #30 of 2026-06-16
  34. LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies 4 upvotes, #34 of 2026-06-16
  35. Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs 4 upvotes, #34 of 2026-06-16
  36. Implicit Reasoning for Large Language Model-based Generative Recommendation 3 upvotes, #36 of 2026-06-16
  37. Attacks on Machine-Text Detectors Retain Stylistic Fingerprints 2 upvotes, #37 of 2026-06-16
  38. Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks 2 upvotes, #37 of 2026-06-16
  39. EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video 2 upvotes, #37 of 2026-06-16
  40. The Ghosts of Polymarket: When Off-Chain Matches Meet On-Chain Reverts 1 upvotes, #40 of 2026-06-16
  41. TuneJury: An Open Metric for Improving Music Generation Preference Alignment 1 upvotes, #40 of 2026-06-16
  42. Human Universal Grasping 1 upvotes, #40 of 2026-06-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.