Daily Papers of 2025-11-26
- ROOT: Robust Orthogonalized Optimizer for Neural Network Training 166 upvotes, #1 of 2025-11-26
- GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms 117 upvotes, #2 of 2025-11-26
- MedSAM3: Delving into Segment Anything with Medical Concepts 48 upvotes, #3 of 2025-11-26
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning 46 upvotes, #4 of 2025-11-26
- SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation 39 upvotes, #5 of 2025-11-26
- Soft Adaptive Policy Optimization 33 upvotes, #6 of 2025-11-26
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward 31 upvotes, #7 of 2025-11-26
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation 31 upvotes, #7 of 2025-11-26
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI 30 upvotes, #9 of 2025-11-26
- STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flow 29 upvotes, #10 of 2025-11-26
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 26 upvotes, #11 of 2025-11-26
- HunyuanOCR Technical Report 19 upvotes, #12 of 2025-11-26
- MagicWorld: Interactive Geometry-driven Video World Exploration 17 upvotes, #13 of 2025-11-26
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers 16 upvotes, #14 of 2025-11-26
- OmniAlpha: A Sequence-to-Sequence Framework for Unified Multi-Task RGBA Generation 12 upvotes, #15 of 2025-11-26
- CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning 11 upvotes, #16 of 2025-11-26
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding 11 upvotes, #16 of 2025-11-26
- Fara-7B: An Efficient Agentic Model for Computer Use 9 upvotes, #18 of 2025-11-26
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs 9 upvotes, #18 of 2025-11-26
- Think Visually, Reason Textually: Vision-Language Synergy in ARC 8 upvotes, #20 of 2025-11-26
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs 8 upvotes, #20 of 2025-11-26
- MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts 8 upvotes, #20 of 2025-11-26
- Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution 7 upvotes, #23 of 2025-11-26
- VQ-VA World: Towards High-Quality Visual Question-Visual Answering 7 upvotes, #23 of 2025-11-26
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion 6 upvotes, #25 of 2025-11-26
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation 4 upvotes, #26 of 2025-11-26
- PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding 4 upvotes, #26 of 2025-11-26
- DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection 3 upvotes, #28 of 2025-11-26
- Unified all-atom molecule generation with neural fields 2 upvotes, #29 of 2025-11-26
- SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System 2 upvotes, #29 of 2025-11-26
- Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization 2 upvotes, #29 of 2025-11-26
- Future Is Unevenly Distributed: Forecasting Ability of LLMs Depends on What We're Asking 1 upvotes, #32 of 2025-11-26
- Concept-Aware Batch Sampling Improves Language-Image Pretraining 1 upvotes, #32 of 2025-11-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.