Daily Papers of 2025-10-28
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations 172 upvotes, #1 of 2025-10-28
- ReCode: Unify Plan and Action for Universal Granularity Control 117 upvotes, #2 of 2025-10-28
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype? 64 upvotes, #3 of 2025-10-28
- FARMER: Flow AutoRegressive Transformer over Pixels 56 upvotes, #4 of 2025-10-28
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
- Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation 41 upvotes, #5 of 2025-10-28
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction 39 upvotes, #7 of 2025-10-28
- ACG: Action Coherence Guidance for Flow-based VLA models 36 upvotes, #8 of 2025-10-28
- E^2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker 31 upvotes, #9 of 2025-10-28
- Open Multimodal Retrieval-Augmented Factual Image Generation 30 upvotes, #10 of 2025-10-28
- Knocking-Heads Attention 28 upvotes, #11 of 2025-10-28
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences 26 upvotes, #12 of 2025-10-28
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity 21 upvotes, #13 of 2025-10-28
- LongCat-Video Technical Report 20 upvotes, #14 of 2025-10-28
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation 20 upvotes, #14 of 2025-10-28
- LightBagel: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation 16 upvotes, #16 of 2025-10-28
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding 14 upvotes, #17 of 2025-10-28
- LimRank: Less is More for Reasoning-Intensive Information Reranking 8 upvotes, #18 of 2025-10-28
- RobotArena infty: Scalable Robot Benchmarking via Real-to-Sim Translation 8 upvotes, #18 of 2025-10-28
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution 8 upvotes, #18 of 2025-10-28
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation 7 upvotes, #21 of 2025-10-28
- Code Aesthetics with Agentic Reward Feedback 7 upvotes, #21 of 2025-10-28
- VoMP: Predicting Volumetric Mechanical Property Fields 6 upvotes, #23 of 2025-10-28
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection 5 upvotes, #24 of 2025-10-28
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling 5 upvotes, #24 of 2025-10-28
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction 4 upvotes, #26 of 2025-10-28
- Language Server CLI Empowers Language Agents with Process Rewards 4 upvotes, #26 of 2025-10-28
- Scaling Laws for Deepfake Detection 3 upvotes, #28 of 2025-10-28
- EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization 3 upvotes, #28 of 2025-10-28
- DiffusionLane: Diffusion Model for Lane Detection 3 upvotes, #28 of 2025-10-28
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis 3 upvotes, #28 of 2025-10-28
- MARS-M: When Variance Reduction Meets Matrices 2 upvotes, #32 of 2025-10-28
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers 2 upvotes, #32 of 2025-10-28
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing 2 upvotes, #32 of 2025-10-28
- Memory-based Language Models: An Efficient, Explainable, and Eco-friendly Approach to Large Language Modeling 2 upvotes, #32 of 2025-10-28
- Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMS 2 upvotes, #32 of 2025-10-28
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.