Daily Papers of 2025-06-06
- ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development 65 upvotes, #1 of 2025-06-06
- SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training 56 upvotes, #2 of 2025-06-06
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models 55 upvotes, #3 of 2025-06-06
- Video World Models with Long-term Spatial Memory 50 upvotes, #4 of 2025-06-06
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics 39 upvotes, #5 of 2025-06-06
- Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts 37 upvotes, #6 of 2025-06-06
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text 36 upvotes, #7 of 2025-06-06
- Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights 29 upvotes, #8 of 2025-06-06
- Aligning Latent Spaces with Flow Priors 25 upvotes, #9 of 2025-06-06
- Inference-Time Hyper-Scaling with KV Cache Compression 25 upvotes, #9 of 2025-06-06
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models 24 upvotes, #11 of 2025-06-06
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos 24 upvotes, #11 of 2025-06-06
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs 20 upvotes, #13 of 2025-06-06
- Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations 18 upvotes, #14 of 2025-06-06
- Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design 18 upvotes, #14 of 2025-06-06
- Search Arena: Analyzing Search-Augmented LLMs 17 upvotes, #16 of 2025-06-06
- StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs 16 upvotes, #17 of 2025-06-06
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs 16 upvotes, #17 of 2025-06-06
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? 15 upvotes, #19 of 2025-06-06
- FlexPainter: Flexible and Multi-View Consistent Texture Generation 14 upvotes, #20 of 2025-06-06
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning 12 upvotes, #21 of 2025-06-06
- Language-Image Alignment with Fixed Text Encoders 11 upvotes, #22 of 2025-06-06
- Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting 11 upvotes, #22 of 2025-06-06
- Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack 9 upvotes, #24 of 2025-06-06
- SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers 7 upvotes, #25 of 2025-06-06
- Geometry-Editable and Appearance-Preserving Object Compositon 6 upvotes, #26 of 2025-06-06
- Kinetics: Rethinking Test-Time Scaling Laws 6 upvotes, #26 of 2025-06-06
- FreeTimeGS: Free Gaussians at Anytime and Anywhere for Dynamic Scene Reconstruction 6 upvotes, #26 of 2025-06-06
- Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets 5 upvotes, #29 of 2025-06-06
- RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGS 4 upvotes, #30 of 2025-06-06
- Images are Worth Variable Length of Representations 4 upvotes, #30 of 2025-06-06
- Contextual Integrity in LLMs via Reasoning and Reinforcement Learning 4 upvotes, #30 of 2025-06-06
- MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale 4 upvotes, #30 of 2025-06-06
- FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation 3 upvotes, #34 of 2025-06-06
- Micro-Act: Mitigate Knowledge Conflict in Question Answering via Actionable Self-Reasoning 3 upvotes, #34 of 2025-06-06
- Rectified Point Flow: Generic Point Cloud Pose Estimation 3 upvotes, #34 of 2025-06-06
- Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving 2 upvotes, #37 of 2025-06-06
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios 2 upvotes, #37 of 2025-06-06
- BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations 2 upvotes, #37 of 2025-06-06
- Watermarking Degrades Alignment in Language Models: Analysis and Mitigation 2 upvotes, #37 of 2025-06-06
- Perceptual Decoupling for Scalable Multi-modal Reasoning via Reward-Optimized Captioning 2 upvotes, #37 of 2025-06-06
- FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing 2 upvotes, #37 of 2025-06-06
- MARBLE: Material Recomposition and Blending in CLIP-Space 2 upvotes, #37 of 2025-06-06
- What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training 1 upvotes, #44 of 2025-06-06
- Rethinking Whole-Body CT Image Interpretation: An Abnormality-Centric Approach 1 upvotes, #44 of 2025-06-06
- PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment 1 upvotes, #44 of 2025-06-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.