Daily Papers of 2025-11-24

  1. SAM 3: Segment Anything with Concepts 96 upvotes, #1 of 2025-11-24
  2. GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization 89 upvotes, #2 of 2025-11-24
  3. OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe 88 upvotes, #3 of 2025-11-24
  4. Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story 85 upvotes, #4 of 2025-11-24
  5. RynnVLA-002: A Unified Vision-Language-Action and World Model 24 upvotes, #5 of 2025-11-24
  6. O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents 23 upvotes, #6 of 2025-11-24
  7. Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination 20 upvotes, #7 of 2025-11-24
  8. WorldGen: From Text to Traversable and Interactive 3D Worlds 18 upvotes, #8 of 2025-11-24
  9. VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models 15 upvotes, #9 of 2025-11-24
  10. Parrot: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs 15 upvotes, #9 of 2025-11-24
  11. Loomis Painter: Reconstructing the Painting Process 15 upvotes, #9 of 2025-11-24
  12. Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight 12 upvotes, #12 of 2025-11-24
  13. InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization 11 upvotes, #13 of 2025-11-24
  14. Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models 9 upvotes, #14 of 2025-11-24
  15. MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging 8 upvotes, #15 of 2025-11-24
  16. VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation 7 upvotes, #16 of 2025-11-24
  17. Insights from the ICLR Peer Review and Rebuttal Process 6 upvotes, #17 of 2025-11-24
  18. OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists 6 upvotes, #17 of 2025-11-24
  19. Diversity Has Always Been There in Your Visual Autoregressive Models 6 upvotes, #17 of 2025-11-24
  20. Taming Generative Synthetic Data for X-ray Prohibited Item Detection 2 upvotes, #20 of 2025-11-24
  21. Planning with Sketch-Guided Verification for Physics-Aware Video Generation 2 upvotes, #20 of 2025-11-24
  22. Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations 1 upvotes, #22 of 2025-11-24
  23. Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models 1 upvotes, #22 of 2025-11-24
  24. Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design 1 upvotes, #22 of 2025-11-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.