Daily Papers of 2025-09-22
- RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation 117 upvotes, #1 of 2025-09-22
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer 49 upvotes, #2 of 2025-09-22
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification 45 upvotes, #3 of 2025-09-22
- SPATIALGEN: Layout-guided 3D Indoor Scene Generation 24 upvotes, #4 of 2025-09-22
- BaseReward: A Strong Baseline for Multimodal Reward Model 21 upvotes, #5 of 2025-09-22
- A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning 18 upvotes, #6 of 2025-09-22
- Lynx: Towards High-Fidelity Personalized Video Generation 12 upvotes, #7 of 2025-09-22
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent 10 upvotes, #8 of 2025-09-22
- RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes 5 upvotes, #9 of 2025-09-22
- Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems 2 upvotes, #10 of 2025-09-22
- Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents 2 upvotes, #10 of 2025-09-22
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers 1 upvotes, #12 of 2025-09-22
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech 1 upvotes, #12 of 2025-09-22
- Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing 1 upvotes, #12 of 2025-09-22
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue 2 upvotes, #15 of 2025-09-22
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.