Daily Papers of 2025-08-26

  1. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
  2. Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation 40 upvotes, #2 of 2025-08-26
  3. MV-RAG: Retrieval Augmented Multiview Diffusion 36 upvotes, #3 of 2025-08-26
  4. Hermes 4 Technical Report 30 upvotes, #4 of 2025-08-26
  5. Understanding Tool-Integrated Reasoning 28 upvotes, #5 of 2025-08-26
  6. T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation 26 upvotes, #6 of 2025-08-26
  7. MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs 26 upvotes, #6 of 2025-08-26
  8. Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling 23 upvotes, #8 of 2025-08-26
  9. Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning 21 upvotes, #9 of 2025-08-26
  10. PosterGen: Aesthetic-Aware Paper-to-Poster Generation via Multi-Agent LLMs 14 upvotes, #10 of 2025-08-26
  11. UQ: Assessing Language Models on Unsolved Questions 13 upvotes, #11 of 2025-08-26
  12. MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment 8 upvotes, #12 of 2025-08-26
  13. TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling 7 upvotes, #13 of 2025-08-26
  14. ST-Raptor: LLM-Powered Semi-Structured Table Question Answering 6 upvotes, #14 of 2025-08-26
  15. Limitations of Normalization in Attention Mechanism 5 upvotes, #15 of 2025-08-26
  16. MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting 4 upvotes, #16 of 2025-08-26
  17. Neither Valid nor Reliable? Investigating the Use of LLMs as Judges 4 upvotes, #16 of 2025-08-26
  18. Explain Before You Answer: A Survey on Compositional Visual Reasoning 3 upvotes, #18 of 2025-08-26
  19. SpotEdit: Evaluating Visually-Guided Image Editing Methods 2 upvotes, #19 of 2025-08-26
  20. Semantic Diffusion Posterior Sampling for Cardiac Ultrasound Dehazing 1 upvotes, #20 of 2025-08-26
  21. German4All - A Dataset and Model for Readability-Controlled Paraphrasing in German 1 upvotes, #20 of 2025-08-26
  22. If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition 1 upvotes, #22 of 2025-08-26
  23. REGEN: Real-Time Photorealism Enhancement in Games via a Dual-Stage Generative Network Framework 1 upvotes, #22 of 2025-08-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.