Daily Papers of 2025-10-28

  1. Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations 172 upvotes, #1 of 2025-10-28
  2. ReCode: Unify Plan and Action for Universal Granularity Control 117 upvotes, #2 of 2025-10-28
  3. A Survey of Data Agents: Emerging Paradigm or Overstated Hype? 64 upvotes, #3 of 2025-10-28
  4. FARMER: Flow AutoRegressive Transformer over Pixels 56 upvotes, #4 of 2025-10-28
  5. VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
  6. Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation 41 upvotes, #5 of 2025-10-28
  7. IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction 39 upvotes, #7 of 2025-10-28
  8. ACG: Action Coherence Guidance for Flow-based VLA models 36 upvotes, #8 of 2025-10-28
  9. E^2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker 31 upvotes, #9 of 2025-10-28
  10. Open Multimodal Retrieval-Augmented Factual Image Generation 30 upvotes, #10 of 2025-10-28
  11. Knocking-Heads Attention 28 upvotes, #11 of 2025-10-28
  12. Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences 26 upvotes, #12 of 2025-10-28
  13. PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity 21 upvotes, #13 of 2025-10-28
  14. LongCat-Video Technical Report 20 upvotes, #14 of 2025-10-28
  15. The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation 20 upvotes, #14 of 2025-10-28
  16. LightBagel: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation 16 upvotes, #16 of 2025-10-28
  17. MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding 14 upvotes, #17 of 2025-10-28
  18. LimRank: Less is More for Reasoning-Intensive Information Reranking 8 upvotes, #18 of 2025-10-28
  19. RobotArena infty: Scalable Robot Benchmarking via Real-to-Sim Translation 8 upvotes, #18 of 2025-10-28
  20. Multi-Agent Evolve: LLM Self-Improve through Co-evolution 8 upvotes, #18 of 2025-10-28
  21. Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation 7 upvotes, #21 of 2025-10-28
  22. Code Aesthetics with Agentic Reward Feedback 7 upvotes, #21 of 2025-10-28
  23. VoMP: Predicting Volumetric Mechanical Property Fields 6 upvotes, #23 of 2025-10-28
  24. PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection 5 upvotes, #24 of 2025-10-28
  25. Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling 5 upvotes, #24 of 2025-10-28
  26. SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction 4 upvotes, #26 of 2025-10-28
  27. Language Server CLI Empowers Language Agents with Process Rewards 4 upvotes, #26 of 2025-10-28
  28. Scaling Laws for Deepfake Detection 3 upvotes, #28 of 2025-10-28
  29. EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization 3 upvotes, #28 of 2025-10-28
  30. DiffusionLane: Diffusion Model for Lane Detection 3 upvotes, #28 of 2025-10-28
  31. Once Upon an Input: Reasoning via Per-Instance Program Synthesis 3 upvotes, #28 of 2025-10-28
  32. MARS-M: When Variance Reduction Meets Matrices 2 upvotes, #32 of 2025-10-28
  33. Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers 2 upvotes, #32 of 2025-10-28
  34. FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing 2 upvotes, #32 of 2025-10-28
  35. Memory-based Language Models: An Efficient, Explainable, and Eco-friendly Approach to Large Language Modeling 2 upvotes, #32 of 2025-10-28
  36. Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMS 2 upvotes, #32 of 2025-10-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.