Daily Papers of 2025-12-19
- Kling-Omni Technical Report 155 upvotes, #1 of 2025-12-19
- Adaptation of Agentic AI 92 upvotes, #2 of 2025-12-19
- Next-Embedding Prediction Makes Strong Vision Learners 78 upvotes, #3 of 2025-12-19
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B 77 upvotes, #4 of 2025-12-19
- Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model 38 upvotes, #5 of 2025-12-19
- StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors 37 upvotes, #6 of 2025-12-19
- Generative Refocusing: Flexible Defocus Control from a Single Image 36 upvotes, #7 of 2025-12-19
- Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation 33 upvotes, #8 of 2025-12-19
- Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection 30 upvotes, #9 of 2025-12-19
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion 25 upvotes, #10 of 2025-12-19
- DeContext as Defense: Safe Image Editing in Diffusion Transformers 24 upvotes, #11 of 2025-12-19
- The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text 24 upvotes, #11 of 2025-12-19
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models 19 upvotes, #14 of 2025-12-19
- EasyV2V: A High-quality Instruction-based Video Editing Framework 17 upvotes, #15 of 2025-12-19
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image 12 upvotes, #16 of 2025-12-19
- AdaTooler-V: Adaptive Tool-Use for Images and Videos 11 upvotes, #17 of 2025-12-19
- RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing 10 upvotes, #18 of 2025-12-19
- FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction 10 upvotes, #18 of 2025-12-19
- Exploration v.s. Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward 10 upvotes, #18 of 2025-12-19
- ModelTables: A Corpus of Tables about Models 8 upvotes, #21 of 2025-12-19
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks 8 upvotes, #21 of 2025-12-19
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 7 upvotes, #23 of 2025-12-19
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification 7 upvotes, #23 of 2025-12-19
- Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language 6 upvotes, #25 of 2025-12-19
- Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision 6 upvotes, #25 of 2025-12-19
- Bidirectional Normalizing Flow: From Data to Noise and Back 5 upvotes, #27 of 2025-12-19
- Improving Recursive Transformers with Mixture of LoRAs 4 upvotes, #28 of 2025-12-19
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers 4 upvotes, #28 of 2025-12-19
- Make-It-Poseable: Feed-forward Latent Posing Model for 3D Humanoid Character Animation 4 upvotes, #28 of 2025-12-19
- FrameDiffuser: G-Buffer-Conditioned Diffusion for Neural Forward Frame Rendering 3 upvotes, #31 of 2025-12-19
- Coupled Variational Reinforcement Learning for Language Model General Reasoning 2 upvotes, #32 of 2025-12-19
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space 2 upvotes, #32 of 2025-12-19
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts 1 upvotes, #34 of 2025-12-19
- TabReX : Tabular Referenceless eXplainable Evaluation 1 upvotes, #34 of 2025-12-19
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning 1 upvotes, #34 of 2025-12-19
- Sharing State Between Prompts and Programs 1 upvotes, #37 of 2025-12-19
- EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration 1 upvotes, #37 of 2025-12-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.