Daily Papers of 2025-06-05
- MiMo-VL Technical Report 70 upvotes, #1 of 2025-06-05
- AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment 45 upvotes, #2 of 2025-06-05
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning 45 upvotes, #2 of 2025-06-05
- OpenThoughts: Data Recipes for Reasoning Models 39 upvotes, #4 of 2025-06-05
- A Controllable Examination for Long-Context Language Models 32 upvotes, #5 of 2025-06-05
- SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models 32 upvotes, #5 of 2025-06-05
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos 29 upvotes, #7 of 2025-06-05
- Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis 26 upvotes, #8 of 2025-06-05
- Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation 24 upvotes, #9 of 2025-06-05
- VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
- Image Editing As Programs with Diffusion Models 22 upvotes, #10 of 2025-06-05
- IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation 21 upvotes, #12 of 2025-06-05
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
- Ψ-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models 16 upvotes, #14 of 2025-06-05
- SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation 14 upvotes, #15 of 2025-06-05
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models 13 upvotes, #16 of 2025-06-05
- LayerFlow: A Unified Model for Layer-aware Video Generation 13 upvotes, #16 of 2025-06-05
- TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence 11 upvotes, #18 of 2025-06-05
- Rectified Sparse Attention 10 upvotes, #19 of 2025-06-05
- TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models 9 upvotes, #20 of 2025-06-05
- Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games 9 upvotes, #20 of 2025-06-05
- BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation 8 upvotes, #22 of 2025-06-05
- Beyond the Surface: Measuring Self-Preference in LLM Judgments 8 upvotes, #22 of 2025-06-05
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers 7 upvotes, #24 of 2025-06-05
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech 6 upvotes, #25 of 2025-06-05
- Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback 6 upvotes, #25 of 2025-06-05
- Robustness in Both Domains: CLIP Needs a Robust Text Encoder 6 upvotes, #25 of 2025-06-05
- Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning 6 upvotes, #25 of 2025-06-05
- POSS: Position Specialist Generates Better Draft for Speculative Decoding 6 upvotes, #25 of 2025-06-05
- Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation 5 upvotes, #30 of 2025-06-05
- Quantitative LLM Judges 5 upvotes, #30 of 2025-06-05
- Adapt before Continual Learning 5 upvotes, #30 of 2025-06-05
- DLP: Dynamic Layerwise Pruning in Large Language Models 4 upvotes, #33 of 2025-06-05
- Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents 4 upvotes, #33 of 2025-06-05
- RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions 4 upvotes, #33 of 2025-06-05
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 3 upvotes, #36 of 2025-06-05
- Small Language Models are the Future of Agentic AI 3 upvotes, #36 of 2025-06-05
- HTSC-2025: A Benchmark Dataset of Ambient-Pressure High-Temperature Superconductors for AI-Driven Critical Temperature Prediction 3 upvotes, #36 of 2025-06-05
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems 3 upvotes, #36 of 2025-06-05
- Unleashing Hour-Scale Video Training for Long Video-Language Understanding 3 upvotes, #36 of 2025-06-05
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning 2 upvotes, #41 of 2025-06-05
- Solving Inverse Problems with FLAIR 2 upvotes, #41 of 2025-06-05
- Robust Neural Rendering in the Wild with Asymmetric Dual 3D Gaussian Splatting 2 upvotes, #41 of 2025-06-05
- VLMs Can Aggregate Scattered Training Patches 2 upvotes, #41 of 2025-06-05
- CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents 2 upvotes, #41 of 2025-06-05
- Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective 2 upvotes, #41 of 2025-06-05
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning 2 upvotes, #41 of 2025-06-05
- RiOSWorld: Benchmarking the Risk of Multimodal Compter-Use Agents 1 upvotes, #48 of 2025-06-05
- Survey of Active Learning Hyperparameters: Insights from a Large-Scale Experimental Grid 1 upvotes, #48 of 2025-06-05
- Sounding that Object: Interactive Object-Aware Image to Audio Generation 1 upvotes, #48 of 2025-06-05
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.