Daily Papers of 2025-10-15
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model 139 upvotes, #1 of 2025-10-15
- Advancing End-to-End Pixel Space Generative Modeling via Self-supervised Pre-training 104 upvotes, #2 of 2025-10-15
- DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation 94 upvotes, #3 of 2025-10-15
- Scaling Language-Centric Omnimodal Representation Learning 94 upvotes, #3 of 2025-10-15
- Robot Learning: A Tutorial 81 upvotes, #5 of 2025-10-15
- A Survey of Vibe Coding with Large Language Models 45 upvotes, #6 of 2025-10-15
- Detect Anything via Next Point Prediction 42 upvotes, #7 of 2025-10-15
- RAG-Anything: All-in-One RAG Framework 35 upvotes, #8 of 2025-10-15
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution 33 upvotes, #9 of 2025-10-15
- Dr.LLM: Dynamic Layer Routing in LLMs 30 upvotes, #10 of 2025-10-15
- Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models 29 upvotes, #11 of 2025-10-15
- ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning 25 upvotes, #12 of 2025-10-15
- R-WoM: Retrieval-augmented World Model For Computer-use Agents 21 upvotes, #13 of 2025-10-15
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models 19 upvotes, #14 of 2025-10-15
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity 15 upvotes, #15 of 2025-10-15
- UniFusion: Vision-Language Model as Unified Encoder in Image Generation 15 upvotes, #15 of 2025-10-15
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling 14 upvotes, #17 of 2025-10-15
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks 14 upvotes, #17 of 2025-10-15
- Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models 12 upvotes, #19 of 2025-10-15
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search 12 upvotes, #19 of 2025-10-15
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model 10 upvotes, #21 of 2025-10-15
- HoneyBee: Data Recipes for Vision-Language Reasoners 9 upvotes, #22 of 2025-10-15
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation 8 upvotes, #23 of 2025-10-15
- Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing 5 upvotes, #24 of 2025-10-15
- The Geometry of Reasoning: Flowing Logics in Representation Space 5 upvotes, #24 of 2025-10-15
- Tensor Logic: The Language of AI 5 upvotes, #24 of 2025-10-15
- What If : Understanding Motion Through Sparse Interactions 5 upvotes, #24 of 2025-10-15
- MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces 4 upvotes, #28 of 2025-10-15
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens 4 upvotes, #28 of 2025-10-15
- One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration 4 upvotes, #28 of 2025-10-15
- Cautious Weight Decay 4 upvotes, #28 of 2025-10-15
- Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management 3 upvotes, #32 of 2025-10-15
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning 3 upvotes, #32 of 2025-10-15
- SR-Scientist: Scientific Equation Discovery With Agentic AI 3 upvotes, #32 of 2025-10-15
- Scaling Long-Horizon LLM Agent via Context-Folding 3 upvotes, #32 of 2025-10-15
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models 2 upvotes, #36 of 2025-10-15
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution 2 upvotes, #36 of 2025-10-15
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability 1 upvotes, #38 of 2025-10-15
- SynthID-Image: Image watermarking at internet scale 1 upvotes, #38 of 2025-10-15
- Why Do Transformers Fail to Forecast Time Series In-Context? 1 upvotes, #38 of 2025-10-15
- Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap 1 upvotes, #38 of 2025-10-15
- Information-Preserving Reformulation of Reasoning Traces for Antidistillation 1 upvotes, #38 of 2025-10-15
- Bag of Tricks for Subverting Reasoning-based Safety Guardrails 1 upvotes, #38 of 2025-10-15
- Deep Research Brings Deeper Harm 1 upvotes, #38 of 2025-10-15
- Locket: Robust Feature-Locking Technique for Language Models 1 upvotes, #38 of 2025-10-15
- Mitigating the Noise Shift for Denoising Generative Models via Noise Awareness Guidance 1 upvotes, #38 of 2025-10-15
- dInfer: An Efficient Inference Framework for Diffusion Language Models 4 upvotes, #47 of 2025-10-15
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.