Daily Papers of 2025-10-15

  1. Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model 139 upvotes, #1 of 2025-10-15
  2. Advancing End-to-End Pixel Space Generative Modeling via Self-supervised Pre-training 104 upvotes, #2 of 2025-10-15
  3. DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation 94 upvotes, #3 of 2025-10-15
  4. Scaling Language-Centric Omnimodal Representation Learning 94 upvotes, #3 of 2025-10-15
  5. Robot Learning: A Tutorial 81 upvotes, #5 of 2025-10-15
  6. A Survey of Vibe Coding with Large Language Models 45 upvotes, #6 of 2025-10-15
  7. Detect Anything via Next Point Prediction 42 upvotes, #7 of 2025-10-15
  8. RAG-Anything: All-in-One RAG Framework 35 upvotes, #8 of 2025-10-15
  9. FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution 33 upvotes, #9 of 2025-10-15
  10. Dr.LLM: Dynamic Layer Routing in LLMs 30 upvotes, #10 of 2025-10-15
  11. Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models 29 upvotes, #11 of 2025-10-15
  12. ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning 25 upvotes, #12 of 2025-10-15
  13. R-WoM: Retrieval-augmented World Model For Computer-use Agents 21 upvotes, #13 of 2025-10-15
  14. SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models 19 upvotes, #14 of 2025-10-15
  15. Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity 15 upvotes, #15 of 2025-10-15
  16. UniFusion: Vision-Language Model as Unified Encoder in Image Generation 15 upvotes, #15 of 2025-10-15
  17. Deconstructing Attention: Investigating Design Principles for Effective Language Modeling 14 upvotes, #17 of 2025-10-15
  18. Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks 14 upvotes, #17 of 2025-10-15
  19. Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models 12 upvotes, #19 of 2025-10-15
  20. DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search 12 upvotes, #19 of 2025-10-15
  21. SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model 10 upvotes, #21 of 2025-10-15
  22. HoneyBee: Data Recipes for Vision-Language Reasoners 9 upvotes, #22 of 2025-10-15
  23. ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation 8 upvotes, #23 of 2025-10-15
  24. Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing 5 upvotes, #24 of 2025-10-15
  25. The Geometry of Reasoning: Flowing Logics in Representation Space 5 upvotes, #24 of 2025-10-15
  26. Tensor Logic: The Language of AI 5 upvotes, #24 of 2025-10-15
  27. What If : Understanding Motion Through Sparse Interactions 5 upvotes, #24 of 2025-10-15
  28. MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces 4 upvotes, #28 of 2025-10-15
  29. LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens 4 upvotes, #28 of 2025-10-15
  30. One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration 4 upvotes, #28 of 2025-10-15
  31. Cautious Weight Decay 4 upvotes, #28 of 2025-10-15
  32. Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management 3 upvotes, #32 of 2025-10-15
  33. ExpVid: A Benchmark for Experiment Video Understanding & Reasoning 3 upvotes, #32 of 2025-10-15
  34. SR-Scientist: Scientific Equation Discovery With Agentic AI 3 upvotes, #32 of 2025-10-15
  35. Scaling Long-Horizon LLM Agent via Context-Folding 3 upvotes, #32 of 2025-10-15
  36. Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models 2 upvotes, #36 of 2025-10-15
  37. ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution 2 upvotes, #36 of 2025-10-15
  38. ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability 1 upvotes, #38 of 2025-10-15
  39. SynthID-Image: Image watermarking at internet scale 1 upvotes, #38 of 2025-10-15
  40. Why Do Transformers Fail to Forecast Time Series In-Context? 1 upvotes, #38 of 2025-10-15
  41. Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap 1 upvotes, #38 of 2025-10-15
  42. Information-Preserving Reformulation of Reasoning Traces for Antidistillation 1 upvotes, #38 of 2025-10-15
  43. Bag of Tricks for Subverting Reasoning-based Safety Guardrails 1 upvotes, #38 of 2025-10-15
  44. Deep Research Brings Deeper Harm 1 upvotes, #38 of 2025-10-15
  45. Locket: Robust Feature-Locking Technique for Language Models 1 upvotes, #38 of 2025-10-15
  46. Mitigating the Noise Shift for Denoising Generative Models via Noise Awareness Guidance 1 upvotes, #38 of 2025-10-15
  47. dInfer: An Efficient Inference Framework for Diffusion Language Models 4 upvotes, #47 of 2025-10-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.