Daily Papers of 2025-10-06
- Apriel-1.5-15b-Thinker 106 upvotes, #1 of 2025-10-06
- Large Reasoning Models Learn Better Alignment from Flawed Thinking 53 upvotes, #2 of 2025-10-06
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation 38 upvotes, #3 of 2025-10-06
- CoDA: Agentic Systems for Collaborative Data Visualization 28 upvotes, #4 of 2025-10-06
- Game-Time: Evaluating Temporal Dynamics in Spoken Language Models 26 upvotes, #5 of 2025-10-06
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization 25 upvotes, #6 of 2025-10-06
- Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition 19 upvotes, #7 of 2025-10-06
- Self-Improvement in Multimodal Large Language Models: A Survey 18 upvotes, #8 of 2025-10-06
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents 16 upvotes, #9 of 2025-10-06
- OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data 14 upvotes, #10 of 2025-10-06
- OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features 12 upvotes, #11 of 2025-10-06
- Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs 10 upvotes, #12 of 2025-10-06
- Efficient Test-Time Scaling for Small Vision-Language Models 9 upvotes, #13 of 2025-10-06
- Triangle Splatting+: Differentiable Rendering with Opaque Triangles 7 upvotes, #14 of 2025-10-06
- REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration 7 upvotes, #14 of 2025-10-06
- SurveyBench: How Well Can LLM(-Agents) Write Academic Surveys? 6 upvotes, #16 of 2025-10-06
- A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning 5 upvotes, #17 of 2025-10-06
- Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling 5 upvotes, #17 of 2025-10-06
- TalkPlay-Tools: Conversational Music Recommendation with LLM Tool Calling 4 upvotes, #19 of 2025-10-06
- Pretraining with hierarchical memories: separating long-tail and common knowledge 4 upvotes, #19 of 2025-10-06
- SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus 4 upvotes, #19 of 2025-10-06
- Personalized Reasoning: Just-In-Time Personalization and Why LLMs Fail At It 3 upvotes, #22 of 2025-10-06
- Align Your Tangent: Training Better Consistency Models via Manifold-Aligned Tangents 3 upvotes, #22 of 2025-10-06
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents 3 upvotes, #22 of 2025-10-06
- LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning 3 upvotes, #22 of 2025-10-06
- FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents 3 upvotes, #22 of 2025-10-06
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping 3 upvotes, #22 of 2025-10-06
- SoundReactor: Frame-level Online Video-to-Audio Generation 2 upvotes, #28 of 2025-10-06
- How Confident are Video Models? Empowering Video Models to Express their Uncertainty 2 upvotes, #28 of 2025-10-06
- Less LLM, More Documents: Searching for Improved RAG 2 upvotes, #28 of 2025-10-06
- Dale meets Langevin: A Multiplicative Denoising Diffusion Model 2 upvotes, #28 of 2025-10-06
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models 2 upvotes, #28 of 2025-10-06
- DiffTester: Accelerating Unit Test Generation for Diffusion LLMs via Repetitive Pattern 1 upvotes, #33 of 2025-10-06
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models 1 upvotes, #33 of 2025-10-06
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces 1 upvotes, #35 of 2025-10-06
- NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving 1 upvotes, #35 of 2025-10-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.