Daily Papers of 2025-08-14

  1. Story2Board: A Training-Free Approach for Expressive Storyboard Generation 62 upvotes, #1 of 2025-08-14
  2. Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory 49 upvotes, #2 of 2025-08-14
  3. Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation 38 upvotes, #3 of 2025-08-14
  4. Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery 37 upvotes, #4 of 2025-08-14
  5. AWorld: Dynamic Multi-Agent System with Stable Maneuvering for Robust GAIA Problem Solving 30 upvotes, #5 of 2025-08-14
  6. Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing 29 upvotes, #6 of 2025-08-14
  7. Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation 24 upvotes, #7 of 2025-08-14
  8. Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment 19 upvotes, #8 of 2025-08-14
  9. Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models 16 upvotes, #9 of 2025-08-14
  10. MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models 15 upvotes, #10 of 2025-08-14
  11. Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models 14 upvotes, #11 of 2025-08-14
  12. Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning 11 upvotes, #12 of 2025-08-14
  13. μ-Parametrization for Mixture of Experts 8 upvotes, #13 of 2025-08-14
  14. IAG: Input-aware Backdoor Attack on VLMs for Visual Grounding 7 upvotes, #14 of 2025-08-14
  15. CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing 6 upvotes, #15 of 2025-08-14
  16. VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models 6 upvotes, #15 of 2025-08-14
  17. Decentralized Aerial Manipulation of a Cable-Suspended Load using Multi-Agent Reinforcement Learning 5 upvotes, #17 of 2025-08-14
  18. GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors 5 upvotes, #17 of 2025-08-14
  19. Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study 3 upvotes, #19 of 2025-08-14
  20. ASM-UNet: Adaptive Scan Mamba Integrating Group Commonalities and Individual Variations for Fine-Grained Segmentation 2 upvotes, #20 of 2025-08-14
  21. The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage 2 upvotes, #20 of 2025-08-14
  22. AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance 1 upvotes, #22 of 2025-08-14
  23. ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering 0 upvotes, #23 of 2025-08-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.