Daily Papers of 2025-12-19

  1. Kling-Omni Technical Report 155 upvotes, #1 of 2025-12-19
  2. Adaptation of Agentic AI 92 upvotes, #2 of 2025-12-19
  3. Next-Embedding Prediction Makes Strong Vision Learners 78 upvotes, #3 of 2025-12-19
  4. LLaDA2.0: Scaling Up Diffusion Language Models to 100B 77 upvotes, #4 of 2025-12-19
  5. Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model 38 upvotes, #5 of 2025-12-19
  6. StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors 37 upvotes, #6 of 2025-12-19
  7. Generative Refocusing: Flexible Defocus Control from a Single Image 36 upvotes, #7 of 2025-12-19
  8. Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation 33 upvotes, #8 of 2025-12-19
  9. Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection 30 upvotes, #9 of 2025-12-19
  10. REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion 25 upvotes, #10 of 2025-12-19
  11. DeContext as Defense: Safe Image Editing in Diffusion Transformers 24 upvotes, #11 of 2025-12-19
  12. The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text 24 upvotes, #11 of 2025-12-19
  13. JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
  14. N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models 19 upvotes, #14 of 2025-12-19
  15. EasyV2V: A High-quality Instruction-based Video Editing Framework 17 upvotes, #15 of 2025-12-19
  16. Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image 12 upvotes, #16 of 2025-12-19
  17. AdaTooler-V: Adaptive Tool-Use for Images and Videos 11 upvotes, #17 of 2025-12-19
  18. RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing 10 upvotes, #18 of 2025-12-19
  19. FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction 10 upvotes, #18 of 2025-12-19
  20. Exploration v.s. Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward 10 upvotes, #18 of 2025-12-19
  21. ModelTables: A Corpus of Tables about Models 8 upvotes, #21 of 2025-12-19
  22. VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks 8 upvotes, #21 of 2025-12-19
  23. Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 7 upvotes, #23 of 2025-12-19
  24. Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification 7 upvotes, #23 of 2025-12-19
  25. Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language 6 upvotes, #25 of 2025-12-19
  26. Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision 6 upvotes, #25 of 2025-12-19
  27. Bidirectional Normalizing Flow: From Data to Noise and Back 5 upvotes, #27 of 2025-12-19
  28. Improving Recursive Transformers with Mixture of LoRAs 4 upvotes, #28 of 2025-12-19
  29. Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers 4 upvotes, #28 of 2025-12-19
  30. Make-It-Poseable: Feed-forward Latent Posing Model for 3D Humanoid Character Animation 4 upvotes, #28 of 2025-12-19
  31. FrameDiffuser: G-Buffer-Conditioned Diffusion for Neural Forward Frame Rendering 3 upvotes, #31 of 2025-12-19
  32. Coupled Variational Reinforcement Learning for Language Model General Reasoning 2 upvotes, #32 of 2025-12-19
  33. Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space 2 upvotes, #32 of 2025-12-19
  34. Vibe Spaces for Creatively Connecting and Expressing Visual Concepts 1 upvotes, #34 of 2025-12-19
  35. TabReX : Tabular Referenceless eXplainable Evaluation 1 upvotes, #34 of 2025-12-19
  36. MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning 1 upvotes, #34 of 2025-12-19
  37. Sharing State Between Prompts and Programs 1 upvotes, #37 of 2025-12-19
  38. EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration 1 upvotes, #37 of 2025-12-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.