Daily Papers of 2025-05-29

  1. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models 114 upvotes, #1 of 2025-05-29
  2. SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents 84 upvotes, #2 of 2025-05-29
  3. R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing 68 upvotes, #3 of 2025-05-29
  4. Skywork Open Reasoner 1 Technical Report 52 upvotes, #4 of 2025-05-29
  5. Sherlock: Self-Correcting Reasoning in Vision-Language Models 50 upvotes, #5 of 2025-05-29
  6. Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO 45 upvotes, #6 of 2025-05-29
  7. Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment 44 upvotes, #7 of 2025-05-29
  8. SageAttention2++: A More Efficient Implementation of SageAttention2 41 upvotes, #8 of 2025-05-29
  9. Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start 36 upvotes, #9 of 2025-05-29
  10. RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination 33 upvotes, #10 of 2025-05-29
  11. Fostering Video Reasoning via Next-Event Prediction 27 upvotes, #11 of 2025-05-29
  12. Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems 25 upvotes, #12 of 2025-05-29
  13. DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research 25 upvotes, #12 of 2025-05-29
  14. FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding 22 upvotes, #14 of 2025-05-29
  15. Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs 21 upvotes, #15 of 2025-05-29
  16. Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models 18 upvotes, #16 of 2025-05-29
  17. WebDancer: Towards Autonomous Information Seeking Agency 18 upvotes, #16 of 2025-05-29
  18. Let's Predict Sentence by Sentence 17 upvotes, #18 of 2025-05-29
  19. What Makes for Text to 360-degree Panorama Generation with Stable Diffusion? 15 upvotes, #19 of 2025-05-29
  20. Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness 15 upvotes, #19 of 2025-05-29
  21. Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States 14 upvotes, #21 of 2025-05-29
  22. Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality 14 upvotes, #21 of 2025-05-29
  23. Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach 14 upvotes, #21 of 2025-05-29
  24. CHIMERA: A Knowledge Base of Idea Recombination in Scientific Literature 14 upvotes, #21 of 2025-05-29
  25. SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem 14 upvotes, #21 of 2025-05-29
  26. Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Credit Assignment 13 upvotes, #26 of 2025-05-29
  27. Thinking with Generated Images 13 upvotes, #26 of 2025-05-29
  28. LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling 12 upvotes, #28 of 2025-05-29
  29. VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning 10 upvotes, #29 of 2025-05-29
  30. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance 9 upvotes, #30 of 2025-05-29
  31. RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction 7 upvotes, #31 of 2025-05-29
  32. MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding 6 upvotes, #32 of 2025-05-29
  33. Prot2Token: A Unified Framework for Protein Modeling via Next-Token Prediction 6 upvotes, #32 of 2025-05-29
  34. Pitfalls of Rule- and Model-based Verifiers -- A Case Study on Mathematical Reasoning 6 upvotes, #32 of 2025-05-29
  35. Text2Grad: Reinforcement Learning from Natural Language Feedback 6 upvotes, #32 of 2025-05-29
  36. PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models 6 upvotes, #32 of 2025-05-29
  37. Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods 5 upvotes, #37 of 2025-05-29
  38. One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models 5 upvotes, #37 of 2025-05-29
  39. Zero-Shot Vision Encoder Grafting via LLM Surrogates 5 upvotes, #37 of 2025-05-29
  40. Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking 4 upvotes, #40 of 2025-05-29
  41. GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains 4 upvotes, #40 of 2025-05-29
  42. Efficient Data Selection at Scale via Influence Distillation 4 upvotes, #40 of 2025-05-29
  43. Styl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and Styles 4 upvotes, #40 of 2025-05-29
  44. Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex 3 upvotes, #44 of 2025-05-29
  45. Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph 3 upvotes, #44 of 2025-05-29
  46. HoPE: Hybrid of Position Embedding for Length Generalization in Vision-Language Models 3 upvotes, #44 of 2025-05-29
  47. AITEE -- Agentic Tutor for Electrical Engineering 3 upvotes, #44 of 2025-05-29
  48. FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control 3 upvotes, #44 of 2025-05-29
  49. PixelThink: Towards Efficient Chain-of-Pixel Reasoning 3 upvotes, #44 of 2025-05-29
  50. MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding 2 upvotes, #50 of 2025-05-29
  51. Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities 2 upvotes, #50 of 2025-05-29
  52. Right Side Up? Disentangling Orientation Understanding in MLLMs with Fine-grained Multi-axis Perception Tasks 2 upvotes, #50 of 2025-05-29
  53. Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese 2 upvotes, #50 of 2025-05-29
  54. First Finish Search: Efficient Test-Time Scaling in Large Language Models 1 upvotes, #54 of 2025-05-29
  55. Can Large Language Models Infer Causal Relationships from Real-World Text? 1 upvotes, #54 of 2025-05-29
  56. Towards Scalable Language-Image Pre-training for 3D Medical Imaging 1 upvotes, #54 of 2025-05-29
  57. Precise In-Parameter Concept Erasure in Large Language Models 1 upvotes, #54 of 2025-05-29
  58. IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests 0 upvotes, #58 of 2025-05-29

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.