Daily Papers of 2025-03-17
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video 117 upvotes, #1 of 2025-03-17
- PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity 80 upvotes, #2 of 2025-03-17
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion 72 upvotes, #3 of 2025-03-17
- Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning 33 upvotes, #4 of 2025-03-17
- API Agents vs. GUI Agents: Divergence and Convergence 31 upvotes, #5 of 2025-03-17
- Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models 26 upvotes, #6 of 2025-03-17
- VGGT: Visual Geometry Grounded Transformer 20 upvotes, #7 of 2025-03-17
- Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers 17 upvotes, #8 of 2025-03-17
- FlowTok: Flowing Seamlessly Across Text and Image Tokens 16 upvotes, #9 of 2025-03-17
- Large-scale Pre-training for Grounded Video Caption Generation 16 upvotes, #9 of 2025-03-17
- TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools 16 upvotes, #9 of 2025-03-17
- Exploring the Vulnerabilities of Federated Learning: A Deep Dive into Gradient Inversion Attacks 15 upvotes, #12 of 2025-03-17
- Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers? 12 upvotes, #13 of 2025-03-17
- ARMOR v0.1: Empowering Autoregressive Multimodal Understanding Model with Interleaved Multimodal Generation via Asymmetric Synergy 8 upvotes, #14 of 2025-03-17
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges 8 upvotes, #14 of 2025-03-17
- Neighboring Autoregressive Modeling for Efficient Visual Generation 8 upvotes, #14 of 2025-03-17
- ETCH: Generalizing Body Fitting to Clothed Humans via Equivariant Tightness 7 upvotes, #17 of 2025-03-17
- Learning Few-Step Diffusion Models by Trajectory Distribution Matching 6 upvotes, #18 of 2025-03-17
- From TOWER to SPIRE: Adding the Speech Modality to a Text-Only LLM 6 upvotes, #18 of 2025-03-17
- Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption 5 upvotes, #20 of 2025-03-17
- Open-World Skill Discovery from Unsegmented Demonstrations 5 upvotes, #20 of 2025-03-17
- Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty? 5 upvotes, #20 of 2025-03-17
- TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing 5 upvotes, #20 of 2025-03-17
- MaRI: Material Retrieval Integration across Domains 4 upvotes, #24 of 2025-03-17
- GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving 3 upvotes, #25 of 2025-03-17
- CHOrD: Generation of Collision-Free, House-Scale, and Organized Digital Twins for 3D Indoor Scenes with Controllable Floor Plans and Optimal Layouts 3 upvotes, #25 of 2025-03-17
- Group-robust Machine Unlearning 1 upvotes, #27 of 2025-03-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.