Daily Papers of 2025-12-02
- From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence 248 upvotes, #1 of 2025-12-02
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling 150 upvotes, #2 of 2025-12-02
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights 88 upvotes, #3 of 2025-12-02
- Stabilizing Reinforcement Learning with LLMs: Formulation and Practices 83 upvotes, #4 of 2025-12-02
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models 60 upvotes, #5 of 2025-12-02
- How Far Are We from Genuinely Useful Deep Research Agents? 51 upvotes, #6 of 2025-12-02
- What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards 47 upvotes, #7 of 2025-12-02
- Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout 44 upvotes, #8 of 2025-12-02
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment 38 upvotes, #9 of 2025-12-02
- Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models 38 upvotes, #9 of 2025-12-02
- LFM2 Technical Report 34 upvotes, #11 of 2025-12-02
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning 24 upvotes, #12 of 2025-12-02
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference 23 upvotes, #13 of 2025-12-02
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation 23 upvotes, #13 of 2025-12-02
- Rectifying LLM Thought from Lens of Optimization 23 upvotes, #13 of 2025-12-02
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model 17 upvotes, #16 of 2025-12-02
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 15 upvotes, #17 of 2025-12-02
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation 14 upvotes, #18 of 2025-12-02
- Flow Straighter and Faster: Efficient One-Step Generative Modeling via MeanFlow on Rectified Trajectories 14 upvotes, #18 of 2025-12-02
- SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs 14 upvotes, #18 of 2025-12-02
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression 14 upvotes, #18 of 2025-12-02
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision 14 upvotes, #18 of 2025-12-02
- SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling 8 upvotes, #23 of 2025-12-02
- PromptBridge: Cross-Model Prompt Transfer for Large Language Models 8 upvotes, #23 of 2025-12-02
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models 8 upvotes, #23 of 2025-12-02
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought 7 upvotes, #26 of 2025-12-02
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos 7 upvotes, #26 of 2025-12-02
- HiconAgent: History Context-aware Policy Optimization for GUI Agents 5 upvotes, #28 of 2025-12-02
- CauSight: Learning to Supersense for Visual Causal Discovery 5 upvotes, #28 of 2025-12-02
- Asking like Socrates: Socrates helps VLMs understand remote sensing images 4 upvotes, #30 of 2025-12-02
- POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models 4 upvotes, #30 of 2025-12-02
- Seeing the Wind from a Falling Leaf 4 upvotes, #30 of 2025-12-02
- Learning Eigenstructures of Unstructured Data Manifolds 4 upvotes, #30 of 2025-12-02
- OpenREAD: Reinforced Open-Ended Reasoing for End-to-End Autonomous Driving with LLM-as-Critic 3 upvotes, #34 of 2025-12-02
- Agentic Policy Optimization via Instruction-Policy Co-Evolution 3 upvotes, #34 of 2025-12-02
- IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages 2 upvotes, #36 of 2025-12-02
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing 2 upvotes, #36 of 2025-12-02
- Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks 2 upvotes, #36 of 2025-12-02
- ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling 2 upvotes, #36 of 2025-12-02
- The Art of Scaling Test-Time Compute for Large Language Models 2 upvotes, #36 of 2025-12-02
- Structured Extraction from Business Process Diagrams Using Vision-Language Models 1 upvotes, #41 of 2025-12-02
- A Hierarchical Framework for Humanoid Locomotion with Supernumerary Limbs 1 upvotes, #41 of 2025-12-02
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models 1 upvotes, #41 of 2025-12-02
- OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion 1 upvotes, #44 of 2025-12-02
- Doppler-Enhanced Deep Learning: Improving Thyroid Nodule Segmentation with YOLOv5 Instance Segmentation 1 upvotes, #44 of 2025-12-02
- MEGConformer: Conformer-Based MEG Decoder for Robust Speech and Phoneme Classification 1 upvotes, #44 of 2025-12-02
- Generative Video Motion Editing with 3D Point Tracks 4 upvotes, #44 of 2025-12-02
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.