Daily Papers of 2025-10-17
- When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA 107 upvotes, #1 of 2025-10-17
- Agentic Entropy-Balanced Policy Optimization 95 upvotes, #2 of 2025-10-17
- WithAnyone: Towards Controllable and ID Consistent Image Generation 76 upvotes, #3 of 2025-10-17
- AI for Service: Proactive Assistance with AI Glasses 71 upvotes, #4 of 2025-10-17
- PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model 64 upvotes, #5 of 2025-10-17
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale 64 upvotes, #5 of 2025-10-17
- ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints 53 upvotes, #7 of 2025-10-17
- BitNet Distillation 47 upvotes, #8 of 2025-10-17
- LaSeR: Reinforcement Learning with Last-Token Self-Rewarding 37 upvotes, #9 of 2025-10-17
- Attention Is All You Need for KV Cache in Diffusion LLMs 35 upvotes, #10 of 2025-10-17
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents 32 upvotes, #11 of 2025-10-17
- TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar 29 upvotes, #12 of 2025-10-17
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning 22 upvotes, #13 of 2025-10-17
- LLMs Can Get "Brain Rot"! 20 upvotes, #14 of 2025-10-17
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning 17 upvotes, #15 of 2025-10-17
- Large Language Models Do NOT Really Know What They Don't Know 16 upvotes, #16 of 2025-10-17
- LLM-guided Hierarchical Retrieval 15 upvotes, #17 of 2025-10-17
- COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes 13 upvotes, #18 of 2025-10-17
- VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation 13 upvotes, #18 of 2025-10-17
- Qwen3Guard Technical Report 12 upvotes, #20 of 2025-10-17
- Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0 Tech Report 12 upvotes, #20 of 2025-10-17
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild 11 upvotes, #22 of 2025-10-17
- AnyUp: Universal Feature Upsampling 10 upvotes, #23 of 2025-10-17
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures 10 upvotes, #23 of 2025-10-17
- Agentic Design of Compositional Machines 10 upvotes, #23 of 2025-10-17
- SimKO: Simple Pass@K Policy Optimization 9 upvotes, #26 of 2025-10-17
- VLA-0: Building State-of-the-Art VLAs with Zero Modification 8 upvotes, #27 of 2025-10-17
- Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning 8 upvotes, #27 of 2025-10-17
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth 7 upvotes, #29 of 2025-10-17
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation 7 upvotes, #29 of 2025-10-17
- VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator 6 upvotes, #31 of 2025-10-17
- The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models 6 upvotes, #31 of 2025-10-17
- LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning 6 upvotes, #31 of 2025-10-17
- RealDPO: Real or Not Real, that is the Preference 6 upvotes, #31 of 2025-10-17
- Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models 6 upvotes, #31 of 2025-10-17
- Learning an Image Editing Model without Image Editing Pairs 6 upvotes, #31 of 2025-10-17
- On Pretraining for Project-Level Code Completion 5 upvotes, #37 of 2025-10-17
- Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space Reasoning 5 upvotes, #37 of 2025-10-17
- DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation 5 upvotes, #37 of 2025-10-17
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training 5 upvotes, #37 of 2025-10-17
- Budget-aware Test-time Scaling via Discriminative Verification 4 upvotes, #41 of 2025-10-17
- Predicting Task Performance with Context-aware Scaling Laws 3 upvotes, #42 of 2025-10-17
- Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction Animation 3 upvotes, #42 of 2025-10-17
- SCas4D: Structural Cascaded Optimization for Boosting Persistent 4D Novel View Synthesis 2 upvotes, #44 of 2025-10-17
- Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms 2 upvotes, #44 of 2025-10-17
- MoM: Mixtures of Scenario-Aware Document Memories for Retrieval-Augmented Generation Systems 2 upvotes, #44 of 2025-10-17
- RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models 1 upvotes, #47 of 2025-10-17
- RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems 1 upvotes, #47 of 2025-10-17
- Beyond One World: Benchmarking Super Heros in Role-Playing Across Multiversal Contexts 1 upvotes, #47 of 2025-10-17
- GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning 1 upvotes, #47 of 2025-10-17
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference 1 upvotes, #51 of 2025-10-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.