Daily Papers of 2025-03-17

  1. ReCamMaster: Camera-Controlled Generative Rendering from A Single Video 117 upvotes, #1 of 2025-03-17
  2. PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity 80 upvotes, #2 of 2025-03-17
  3. SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion 72 upvotes, #3 of 2025-03-17
  4. Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning 33 upvotes, #4 of 2025-03-17
  5. API Agents vs. GUI Agents: Divergence and Convergence 31 upvotes, #5 of 2025-03-17
  6. Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models 26 upvotes, #6 of 2025-03-17
  7. VGGT: Visual Geometry Grounded Transformer 20 upvotes, #7 of 2025-03-17
  8. Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers 17 upvotes, #8 of 2025-03-17
  9. FlowTok: Flowing Seamlessly Across Text and Image Tokens 16 upvotes, #9 of 2025-03-17
  10. Large-scale Pre-training for Grounded Video Caption Generation 16 upvotes, #9 of 2025-03-17
  11. TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools 16 upvotes, #9 of 2025-03-17
  12. Exploring the Vulnerabilities of Federated Learning: A Deep Dive into Gradient Inversion Attacks 15 upvotes, #12 of 2025-03-17
  13. Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers? 12 upvotes, #13 of 2025-03-17
  14. ARMOR v0.1: Empowering Autoregressive Multimodal Understanding Model with Interleaved Multimodal Generation via Asymmetric Synergy 8 upvotes, #14 of 2025-03-17
  15. ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges 8 upvotes, #14 of 2025-03-17
  16. Neighboring Autoregressive Modeling for Efficient Visual Generation 8 upvotes, #14 of 2025-03-17
  17. ETCH: Generalizing Body Fitting to Clothed Humans via Equivariant Tightness 7 upvotes, #17 of 2025-03-17
  18. Learning Few-Step Diffusion Models by Trajectory Distribution Matching 6 upvotes, #18 of 2025-03-17
  19. From TOWER to SPIRE: Adding the Speech Modality to a Text-Only LLM 6 upvotes, #18 of 2025-03-17
  20. Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption 5 upvotes, #20 of 2025-03-17
  21. Open-World Skill Discovery from Unsegmented Demonstrations 5 upvotes, #20 of 2025-03-17
  22. Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty? 5 upvotes, #20 of 2025-03-17
  23. TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing 5 upvotes, #20 of 2025-03-17
  24. MaRI: Material Retrieval Integration across Domains 4 upvotes, #24 of 2025-03-17
  25. GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving 3 upvotes, #25 of 2025-03-17
  26. CHOrD: Generation of Collision-Free, House-Scale, and Organized Digital Twins for 3D Indoor Scenes with Controllable Floor Plans and Optimal Layouts 3 upvotes, #25 of 2025-03-17
  27. Group-robust Machine Unlearning 1 upvotes, #27 of 2025-03-17

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.