Daily Papers of 2025-11-26

  1. ROOT: Robust Orthogonalized Optimizer for Neural Network Training 166 upvotes, #1 of 2025-11-26
  2. GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms 117 upvotes, #2 of 2025-11-26
  3. MedSAM3: Delving into Segment Anything with Medical Concepts 48 upvotes, #3 of 2025-11-26
  4. Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning 46 upvotes, #4 of 2025-11-26
  5. SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation 39 upvotes, #5 of 2025-11-26
  6. Soft Adaptive Policy Optimization 33 upvotes, #6 of 2025-11-26
  7. Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward 31 upvotes, #7 of 2025-11-26
  8. iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation 31 upvotes, #7 of 2025-11-26
  9. GigaWorld-0: World Models as Data Engine to Empower Embodied AI 30 upvotes, #9 of 2025-11-26
  10. STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flow 29 upvotes, #10 of 2025-11-26
  11. SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 26 upvotes, #11 of 2025-11-26
  12. HunyuanOCR Technical Report 19 upvotes, #12 of 2025-11-26
  13. MagicWorld: Interactive Geometry-driven Video World Exploration 17 upvotes, #13 of 2025-11-26
  14. UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers 16 upvotes, #14 of 2025-11-26
  15. OmniAlpha: A Sequence-to-Sequence Framework for Unified Multi-Task RGBA Generation 12 upvotes, #15 of 2025-11-26
  16. CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning 11 upvotes, #16 of 2025-11-26
  17. ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding 11 upvotes, #16 of 2025-11-26
  18. Fara-7B: An Efficient Agentic Model for Computer Use 9 upvotes, #18 of 2025-11-26
  19. Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs 9 upvotes, #18 of 2025-11-26
  20. Think Visually, Reason Textually: Vision-Language Synergy in ARC 8 upvotes, #20 of 2025-11-26
  21. Cognitive Foundations for Reasoning and Their Manifestation in LLMs 8 upvotes, #20 of 2025-11-26
  22. MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts 8 upvotes, #20 of 2025-11-26
  23. Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution 7 upvotes, #23 of 2025-11-26
  24. VQ-VA World: Towards High-Quality Visual Question-Visual Answering 7 upvotes, #23 of 2025-11-26
  25. Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion 6 upvotes, #25 of 2025-11-26
  26. Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation 4 upvotes, #26 of 2025-11-26
  27. PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding 4 upvotes, #26 of 2025-11-26
  28. DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection 3 upvotes, #28 of 2025-11-26
  29. Unified all-atom molecule generation with neural fields 2 upvotes, #29 of 2025-11-26
  30. SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System 2 upvotes, #29 of 2025-11-26
  31. Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization 2 upvotes, #29 of 2025-11-26
  32. Future Is Unevenly Distributed: Forecasting Ability of LLMs Depends on What We're Asking 1 upvotes, #32 of 2025-11-26
  33. Concept-Aware Batch Sampling Improves Language-Image Pretraining 1 upvotes, #32 of 2025-11-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.