Daily Papers of 2025-12-17
- MMGR: Multi-Modal Generative Reasoning 114 upvotes, #1 of 2025-12-17
- Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans? 63 upvotes, #2 of 2025-12-17
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling 61 upvotes, #3 of 2025-12-17
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling 40 upvotes, #4 of 2025-12-17
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value 38 upvotes, #5 of 2025-12-17
- RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics 36 upvotes, #6 of 2025-12-17
- Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure 28 upvotes, #7 of 2025-12-17
- Reveal Hidden Pitfalls and Navigate Next Generation of Vector Similarity Search from Task-Centric Views 27 upvotes, #8 of 2025-12-17
- MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives 26 upvotes, #9 of 2025-12-17
- Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models 25 upvotes, #10 of 2025-12-17
- Olmo 3 22 upvotes, #11 of 2025-12-17
- Differentiable Evolutionary Reinforcement Learning 18 upvotes, #12 of 2025-12-17
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs 18 upvotes, #12 of 2025-12-17
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement 16 upvotes, #14 of 2025-12-17
- RecGPT-V2 Technical Report 16 upvotes, #14 of 2025-12-17
- Feedforward 3D Editing via Text-Steerable Image-to-3D 13 upvotes, #16 of 2025-12-17
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed 12 upvotes, #17 of 2025-12-17
- SS4D: Native 4D Generative Model via Structured Spacetime Latents 12 upvotes, #17 of 2025-12-17
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse 11 upvotes, #19 of 2025-12-17
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning 10 upvotes, #20 of 2025-12-17
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models 9 upvotes, #21 of 2025-12-17
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models 8 upvotes, #22 of 2025-12-17
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in 7 upvotes, #23 of 2025-12-17
- RePo: Language Models with Context Re-Positioning 7 upvotes, #23 of 2025-12-17
- Spherical Leech Quantization for Visual Tokenization and Generation 7 upvotes, #23 of 2025-12-17
- CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives 6 upvotes, #26 of 2025-12-17
- Janus: Disaggregating Attention and Experts for Scalable MoE Inference 5 upvotes, #27 of 2025-12-17
- TAT: Task-Adaptive Transformer for All-in-One Medical Image Restoration 5 upvotes, #27 of 2025-12-17
- TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning 3 upvotes, #29 of 2025-12-17
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation 2 upvotes, #30 of 2025-12-17
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents 2 upvotes, #30 of 2025-12-17
- Hierarchical Dataset Selection for High-Quality Data Sharing 1 upvotes, #32 of 2025-12-17
- UAGLNet: Uncertainty-Aggregated Global-Local Fusion Network with Cooperative CNN-Transformer for Building Extraction 1 upvotes, #32 of 2025-12-17
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction 1 upvotes, #32 of 2025-12-17
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation 1 upvotes, #35 of 2025-12-17
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates 1 upvotes, #35 of 2025-12-17
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation 1 upvotes, #35 of 2025-12-17
- Unveiling User Perceptions in the Generative AI Era: A Sentiment-Driven Evaluation of AI Educational Apps' Role in Digital Transformation of e-Teaching 1 upvotes, #35 of 2025-12-17
- S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation 1 upvotes, #35 of 2025-12-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.