Daily Papers of 2025-11-24
- SAM 3: Segment Anything with Concepts 96 upvotes, #1 of 2025-11-24
- GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization 89 upvotes, #2 of 2025-11-24
- OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe 88 upvotes, #3 of 2025-11-24
- Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story 85 upvotes, #4 of 2025-11-24
- RynnVLA-002: A Unified Vision-Language-Action and World Model 24 upvotes, #5 of 2025-11-24
- O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents 23 upvotes, #6 of 2025-11-24
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination 20 upvotes, #7 of 2025-11-24
- WorldGen: From Text to Traversable and Interactive 3D Worlds 18 upvotes, #8 of 2025-11-24
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models 15 upvotes, #9 of 2025-11-24
- Parrot: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs 15 upvotes, #9 of 2025-11-24
- Loomis Painter: Reconstructing the Painting Process 15 upvotes, #9 of 2025-11-24
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight 12 upvotes, #12 of 2025-11-24
- InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization 11 upvotes, #13 of 2025-11-24
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models 9 upvotes, #14 of 2025-11-24
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging 8 upvotes, #15 of 2025-11-24
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation 7 upvotes, #16 of 2025-11-24
- Insights from the ICLR Peer Review and Rebuttal Process 6 upvotes, #17 of 2025-11-24
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists 6 upvotes, #17 of 2025-11-24
- Diversity Has Always Been There in Your Visual Autoregressive Models 6 upvotes, #17 of 2025-11-24
- Taming Generative Synthetic Data for X-ray Prohibited Item Detection 2 upvotes, #20 of 2025-11-24
- Planning with Sketch-Guided Verification for Physics-Aware Video Generation 2 upvotes, #20 of 2025-11-24
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations 1 upvotes, #22 of 2025-11-24
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models 1 upvotes, #22 of 2025-11-24
- Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design 1 upvotes, #22 of 2025-11-24
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.