Daily Papers of 2025-08-14
- Story2Board: A Training-Free Approach for Expressive Storyboard Generation 62 upvotes, #1 of 2025-08-14
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory 49 upvotes, #2 of 2025-08-14
- Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation 38 upvotes, #3 of 2025-08-14
- Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery 37 upvotes, #4 of 2025-08-14
- AWorld: Dynamic Multi-Agent System with Stable Maneuvering for Robust GAIA Problem Solving 30 upvotes, #5 of 2025-08-14
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing 29 upvotes, #6 of 2025-08-14
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation 24 upvotes, #7 of 2025-08-14
- Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment 19 upvotes, #8 of 2025-08-14
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models 16 upvotes, #9 of 2025-08-14
- MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models 15 upvotes, #10 of 2025-08-14
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models 14 upvotes, #11 of 2025-08-14
- Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning 11 upvotes, #12 of 2025-08-14
- μ-Parametrization for Mixture of Experts 8 upvotes, #13 of 2025-08-14
- IAG: Input-aware Backdoor Attack on VLMs for Visual Grounding 7 upvotes, #14 of 2025-08-14
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing 6 upvotes, #15 of 2025-08-14
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models 6 upvotes, #15 of 2025-08-14
- Decentralized Aerial Manipulation of a Cable-Suspended Load using Multi-Agent Reinforcement Learning 5 upvotes, #17 of 2025-08-14
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors 5 upvotes, #17 of 2025-08-14
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study 3 upvotes, #19 of 2025-08-14
- ASM-UNet: Adaptive Scan Mamba Integrating Group Commonalities and Individual Variations for Fine-Grained Segmentation 2 upvotes, #20 of 2025-08-14
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage 2 upvotes, #20 of 2025-08-14
- AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance 1 upvotes, #22 of 2025-08-14
- ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering 0 upvotes, #23 of 2025-08-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.