Daily Papers of 2025-10-09

  1. Cache-to-Cache: Direct Semantic Communication Between Large Language Models 89 upvotes, #1 of 2025-10-09
  2. Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer 69 upvotes, #2 of 2025-10-09
  3. Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding 49 upvotes, #3 of 2025-10-09
  4. RLinf-VLA: A Unified and Efficient Framework for VLA+RL Training 35 upvotes, #4 of 2025-10-09
  5. MATRIX: Mask Track Alignment for Interaction-aware Video Generation 35 upvotes, #4 of 2025-10-09
  6. SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models 34 upvotes, #6 of 2025-10-09
  7. Vibe Checker: Aligning Code Evaluation with Human Preference 30 upvotes, #7 of 2025-10-09
  8. Multi-Agent Tool-Integrated Policy Optimization 29 upvotes, #8 of 2025-10-09
  9. The Markovian Thinker 27 upvotes, #9 of 2025-10-09
  10. Artificial Hippocampus Networks for Efficient Long-Context Modeling 26 upvotes, #10 of 2025-10-09
  11. Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought 24 upvotes, #11 of 2025-10-09
  12. Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention 22 upvotes, #12 of 2025-10-09
  13. The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP 22 upvotes, #12 of 2025-10-09
  14. OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot 21 upvotes, #14 of 2025-10-09
  15. Revisiting Long-context Modeling from Context Denoising Perspective 20 upvotes, #15 of 2025-10-09
  16. CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling 19 upvotes, #16 of 2025-10-09
  17. Native Hybrid Attention for Efficient Sequence Modeling 16 upvotes, #17 of 2025-10-09
  18. When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation 14 upvotes, #18 of 2025-10-09
  19. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation 12 upvotes, #19 of 2025-10-09
  20. Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs 11 upvotes, #20 of 2025-10-09
  21. TTRV: Test-Time Reinforcement Learning for Vision Language Models 11 upvotes, #20 of 2025-10-09
  22. Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods 11 upvotes, #20 of 2025-10-09
  23. Reinforcement Mid-Training 8 upvotes, #23 of 2025-10-09
  24. NorMuon: Making Muon more efficient and scalable 6 upvotes, #24 of 2025-10-09
  25. Revisiting the Uniform Information Density Hypothesis in LLM Reasoning Traces 6 upvotes, #24 of 2025-10-09
  26. G^2RPO: Granular GRPO for Precise Reward in Flow Models 5 upvotes, #26 of 2025-10-09
  27. MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline 5 upvotes, #26 of 2025-10-09
  28. WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation 5 upvotes, #26 of 2025-10-09
  29. AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning 4 upvotes, #29 of 2025-10-09
  30. Bridging Text and Video Generation: A Survey 3 upvotes, #30 of 2025-10-09
  31. A Single Character can Make or Break Your LLM Evals 3 upvotes, #30 of 2025-10-09
  32. Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent 3 upvotes, #30 of 2025-10-09
  33. Heptapod: Language Modeling on Visual Signals 3 upvotes, #30 of 2025-10-09
  34. Online Generic Event Boundary Detection 3 upvotes, #30 of 2025-10-09
  35. Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models 3 upvotes, #30 of 2025-10-09
  36. U-Bench: A Comprehensive Understanding of U-Net through 100-Variant Benchmarking 3 upvotes, #30 of 2025-10-09
  37. DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents 2 upvotes, #37 of 2025-10-09
  38. FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering 2 upvotes, #37 of 2025-10-09
  39. M3Retrieve: Benchmarking Multimodal Retrieval for Medicine 2 upvotes, #37 of 2025-10-09
  40. TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility 2 upvotes, #37 of 2025-10-09
  41. D^3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection 1 upvotes, #41 of 2025-10-09
  42. PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles 1 upvotes, #41 of 2025-10-09
  43. Glocal Information Bottleneck for Time Series Imputation 1 upvotes, #43 of 2025-10-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.