Daily Papers of 2025-06-06

  1. ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development 65 upvotes, #1 of 2025-06-06
  2. SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training 56 upvotes, #2 of 2025-06-06
  3. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models 55 upvotes, #3 of 2025-06-06
  4. Video World Models with Long-term Spatial Memory 50 upvotes, #4 of 2025-06-06
  5. RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics 39 upvotes, #5 of 2025-06-06
  6. Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts 37 upvotes, #6 of 2025-06-06
  7. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text 36 upvotes, #7 of 2025-06-06
  8. Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights 29 upvotes, #8 of 2025-06-06
  9. Aligning Latent Spaces with Flow Priors 25 upvotes, #9 of 2025-06-06
  10. Inference-Time Hyper-Scaling with KV Cache Compression 25 upvotes, #9 of 2025-06-06
  11. VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models 24 upvotes, #11 of 2025-06-06
  12. VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos 24 upvotes, #11 of 2025-06-06
  13. AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs 20 upvotes, #13 of 2025-06-06
  14. Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations 18 upvotes, #14 of 2025-06-06
  15. Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design 18 upvotes, #14 of 2025-06-06
  16. Search Arena: Analyzing Search-Augmented LLMs 17 upvotes, #16 of 2025-06-06
  17. StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs 16 upvotes, #17 of 2025-06-06
  18. SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs 16 upvotes, #17 of 2025-06-06
  19. EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? 15 upvotes, #19 of 2025-06-06
  20. FlexPainter: Flexible and Multi-View Consistent Texture Generation 14 upvotes, #20 of 2025-06-06
  21. MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning 12 upvotes, #21 of 2025-06-06
  22. Language-Image Alignment with Fixed Text Encoders 11 upvotes, #22 of 2025-06-06
  23. Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting 11 upvotes, #22 of 2025-06-06
  24. Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack 9 upvotes, #24 of 2025-06-06
  25. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers 7 upvotes, #25 of 2025-06-06
  26. Geometry-Editable and Appearance-Preserving Object Compositon 6 upvotes, #26 of 2025-06-06
  27. Kinetics: Rethinking Test-Time Scaling Laws 6 upvotes, #26 of 2025-06-06
  28. FreeTimeGS: Free Gaussians at Anytime and Anywhere for Dynamic Scene Reconstruction 6 upvotes, #26 of 2025-06-06
  29. Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets 5 upvotes, #29 of 2025-06-06
  30. RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGS 4 upvotes, #30 of 2025-06-06
  31. Images are Worth Variable Length of Representations 4 upvotes, #30 of 2025-06-06
  32. Contextual Integrity in LLMs via Reasoning and Reinforcement Learning 4 upvotes, #30 of 2025-06-06
  33. MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale 4 upvotes, #30 of 2025-06-06
  34. FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation 3 upvotes, #34 of 2025-06-06
  35. Micro-Act: Mitigate Knowledge Conflict in Question Answering via Actionable Self-Reasoning 3 upvotes, #34 of 2025-06-06
  36. Rectified Point Flow: Generic Point Cloud Pose Estimation 3 upvotes, #34 of 2025-06-06
  37. Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving 2 upvotes, #37 of 2025-06-06
  38. SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios 2 upvotes, #37 of 2025-06-06
  39. BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations 2 upvotes, #37 of 2025-06-06
  40. Watermarking Degrades Alignment in Language Models: Analysis and Mitigation 2 upvotes, #37 of 2025-06-06
  41. Perceptual Decoupling for Scalable Multi-modal Reasoning via Reward-Optimized Captioning 2 upvotes, #37 of 2025-06-06
  42. FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing 2 upvotes, #37 of 2025-06-06
  43. MARBLE: Material Recomposition and Blending in CLIP-Space 2 upvotes, #37 of 2025-06-06
  44. What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training 1 upvotes, #44 of 2025-06-06
  45. Rethinking Whole-Body CT Image Interpretation: An Abnormality-Centric Approach 1 upvotes, #44 of 2025-06-06
  46. PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment 1 upvotes, #44 of 2025-06-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.