Daily Papers of 2025-09-23

  1. Qwen3-Omni Technical Report 121 upvotes, #1 of 2025-09-23
  2. LIMI: Less is More for Agency 94 upvotes, #2 of 2025-09-23
  3. OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models 63 upvotes, #3 of 2025-09-23
  4. ARE: Scaling Up Agent Environments and Evaluations 33 upvotes, #4 of 2025-09-23
  5. OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System 32 upvotes, #5 of 2025-09-23
  6. TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs 27 upvotes, #6 of 2025-09-23
  7. VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models 25 upvotes, #7 of 2025-09-23
  8. DiffusionNFT: Online Diffusion Reinforcement with Forward Process 20 upvotes, #8 of 2025-09-23
  9. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? 19 upvotes, #9 of 2025-09-23
  10. EpiCache: Episodic KV Cache Management for Long Conversational Question Answering 18 upvotes, #10 of 2025-09-23
  11. GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning 17 upvotes, #11 of 2025-09-23
  12. FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions 13 upvotes, #12 of 2025-09-23
  13. ByteWrist: A Parallel Robotic Wrist Enabling Flexible and Anthropomorphic Motion for Confined Spaces 13 upvotes, #12 of 2025-09-23
  14. Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels 12 upvotes, #14 of 2025-09-23
  15. Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLM 10 upvotes, #15 of 2025-09-23
  16. QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models 9 upvotes, #16 of 2025-09-23
  17. Synthetic bootstrapped pretraining 8 upvotes, #17 of 2025-09-23
  18. Mano Report 8 upvotes, #17 of 2025-09-23
  19. ContextFlow: Training-Free Video Object Editing via Adaptive Context Enrichment 7 upvotes, #19 of 2025-09-23
  20. MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction 7 upvotes, #19 of 2025-09-23
  21. Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications 6 upvotes, #21 of 2025-09-23
  22. Understanding Embedding Scaling in Collaborative Filtering 5 upvotes, #22 of 2025-09-23
  23. Cross-Attention is Half Explanation in Speech-to-Text Models 5 upvotes, #22 of 2025-09-23
  24. Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning 4 upvotes, #24 of 2025-09-23
  25. UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning 4 upvotes, #24 of 2025-09-23
  26. AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing? 3 upvotes, #26 of 2025-09-23
  27. D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models 3 upvotes, #26 of 2025-09-23
  28. V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts 3 upvotes, #26 of 2025-09-23
  29. DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation 2 upvotes, #29 of 2025-09-23
  30. From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem 2 upvotes, #29 of 2025-09-23
  31. From Uniform to Heterogeneous: Tailoring Policy Optimization to Every Token's Nature 2 upvotes, #29 of 2025-09-23
  32. Accurate and Efficient Low-Rank Model Merging in Core Space 2 upvotes, #29 of 2025-09-23
  33. CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects 1 upvotes, #33 of 2025-09-23
  34. FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation 1 upvotes, #33 of 2025-09-23
  35. StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes 1 upvotes, #33 of 2025-09-23
  36. When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs 1 upvotes, #33 of 2025-09-23
  37. VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery 1 upvotes, #33 of 2025-09-23
  38. BeepBank-500: A Synthetic Earcon Mini-Corpus for UI Sound Research and Psychoacoustics Research 1 upvotes, #33 of 2025-09-23
  39. DIWALI - Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context 1 upvotes, #33 of 2025-09-23
  40. Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMs 1 upvotes, #33 of 2025-09-23
  41. SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning 0 upvotes, #41 of 2025-09-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.