Daily Papers of 2025-05-30

  1. Table-R1: Inference-Time Scaling for Table Reasoning 88 upvotes, #1 of 2025-05-30
  2. Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence 66 upvotes, #2 of 2025-05-30
  3. The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason 64 upvotes, #3 of 2025-05-30
  4. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 56 upvotes, #4 of 2025-05-30
  5. ZeroGUI: Automating Online GUI Learning at Zero Human Cost 45 upvotes, #5 of 2025-05-30
  6. Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding 39 upvotes, #6 of 2025-05-30
  7. VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? 39 upvotes, #6 of 2025-05-30
  8. D-AR: Diffusion via Autoregressive Models 34 upvotes, #8 of 2025-05-30
  9. AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views 31 upvotes, #9 of 2025-05-30
  10. cadrille: Multi-modal CAD Reconstruction with Online Reinforcement Learning 28 upvotes, #10 of 2025-05-30
  11. Are Reasoning Models More Prone to Hallucination? 24 upvotes, #11 of 2025-05-30
  12. UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning 23 upvotes, #12 of 2025-05-30
  13. Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering 23 upvotes, #12 of 2025-05-30
  14. LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers 23 upvotes, #12 of 2025-05-30
  15. ATLAS: Learning to Optimally Memorize the Context at Test Time 22 upvotes, #15 of 2025-05-30
  16. Multi-Domain Explainability of Preferences 21 upvotes, #16 of 2025-05-30
  17. Train Sparse Autoencoders Efficiently by Utilizing Features Correlation 21 upvotes, #16 of 2025-05-30
  18. FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian 20 upvotes, #18 of 2025-05-30
  19. VidText: Towards Comprehensive Evaluation for Video Text Understanding 20 upvotes, #18 of 2025-05-30
  20. SWE-bench Goes Live! 20 upvotes, #18 of 2025-05-30
  21. Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation 17 upvotes, #21 of 2025-05-30
  22. StressTest: Can YOUR Speech LM Handle the Stress? 17 upvotes, #21 of 2025-05-30
  23. REOrdering Patches Improves Vision Models 16 upvotes, #23 of 2025-05-30
  24. DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning 15 upvotes, #24 of 2025-05-30
  25. On-Policy RL with Optimal Reward Baseline 14 upvotes, #25 of 2025-05-30
  26. Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model 14 upvotes, #25 of 2025-05-30
  27. System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts 12 upvotes, #27 of 2025-05-30
  28. SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents 12 upvotes, #27 of 2025-05-30
  29. PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions 11 upvotes, #29 of 2025-05-30
  30. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control 11 upvotes, #29 of 2025-05-30
  31. Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? 10 upvotes, #31 of 2025-05-30
  32. Differentiable Solver Search for Fast Diffusion Sampling 10 upvotes, #31 of 2025-05-30
  33. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction 9 upvotes, #33 of 2025-05-30
  34. MAGREF: Masked Guidance for Any-Reference Video Generation 9 upvotes, #33 of 2025-05-30
  35. Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction 8 upvotes, #35 of 2025-05-30
  36. ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind 8 upvotes, #35 of 2025-05-30
  37. One-shot Entropy Minimization 7 upvotes, #37 of 2025-05-30
  38. Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape 7 upvotes, #37 of 2025-05-30
  39. ATI: Any Trajectory Instruction for Controllable Video Generation 7 upvotes, #37 of 2025-05-30
  40. Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization 7 upvotes, #37 of 2025-05-30
  41. ZeroSep: Separate Anything in Audio with Zero Training 7 upvotes, #37 of 2025-05-30
  42. CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays 6 upvotes, #42 of 2025-05-30
  43. When Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy 6 upvotes, #42 of 2025-05-30
  44. Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting 5 upvotes, #44 of 2025-05-30
  45. CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting 5 upvotes, #44 of 2025-05-30
  46. UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes 5 upvotes, #44 of 2025-05-30
  47. To Trust Or Not To Trust Your Vision-Language Model's Prediction 5 upvotes, #44 of 2025-05-30
  48. Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint 5 upvotes, #44 of 2025-05-30
  49. A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models 4 upvotes, #49 of 2025-05-30
  50. ChartLens: Fine-grained Visual Attribution in Charts 4 upvotes, #49 of 2025-05-30
  51. Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation 4 upvotes, #49 of 2025-05-30
  52. SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model 4 upvotes, #49 of 2025-05-30
  53. Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates 4 upvotes, #49 of 2025-05-30
  54. ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS 4 upvotes, #49 of 2025-05-30
  55. How Animals Dance (When You're Not Looking) 4 upvotes, #49 of 2025-05-30
  56. TokBench: Evaluating Your Visual Tokenizer before Visual Generation 3 upvotes, #56 of 2025-05-30
  57. Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator 3 upvotes, #56 of 2025-05-30
  58. GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents 3 upvotes, #56 of 2025-05-30
  59. Grounded Reinforcement Learning for Visual Reasoning 3 upvotes, #56 of 2025-05-30
  60. Differential Information: An Information-Theoretic Perspective on Preference Optimization 3 upvotes, #56 of 2025-05-30
  61. MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence 3 upvotes, #56 of 2025-05-30
  62. Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities 2 upvotes, #62 of 2025-05-30
  63. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking 2 upvotes, #62 of 2025-05-30
  64. Model-Preserving Adaptive Rounding 2 upvotes, #62 of 2025-05-30
  65. Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement 2 upvotes, #62 of 2025-05-30
  66. Toward Reliable Biomedical Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models 1 upvotes, #66 of 2025-05-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.