Daily Papers of 2025-05-30
- Table-R1: Inference-Time Scaling for Table Reasoning 88 upvotes, #1 of 2025-05-30
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence 66 upvotes, #2 of 2025-05-30
- The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason 64 upvotes, #3 of 2025-05-30
- VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 56 upvotes, #4 of 2025-05-30
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost 45 upvotes, #5 of 2025-05-30
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding 39 upvotes, #6 of 2025-05-30
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? 39 upvotes, #6 of 2025-05-30
- D-AR: Diffusion via Autoregressive Models 34 upvotes, #8 of 2025-05-30
- AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views 31 upvotes, #9 of 2025-05-30
- cadrille: Multi-modal CAD Reconstruction with Online Reinforcement Learning 28 upvotes, #10 of 2025-05-30
- Are Reasoning Models More Prone to Hallucination? 24 upvotes, #11 of 2025-05-30
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning 23 upvotes, #12 of 2025-05-30
- Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering 23 upvotes, #12 of 2025-05-30
- LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers 23 upvotes, #12 of 2025-05-30
- ATLAS: Learning to Optimally Memorize the Context at Test Time 22 upvotes, #15 of 2025-05-30
- Multi-Domain Explainability of Preferences 21 upvotes, #16 of 2025-05-30
- Train Sparse Autoencoders Efficiently by Utilizing Features Correlation 21 upvotes, #16 of 2025-05-30
- FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian 20 upvotes, #18 of 2025-05-30
- VidText: Towards Comprehensive Evaluation for Video Text Understanding 20 upvotes, #18 of 2025-05-30
- SWE-bench Goes Live! 20 upvotes, #18 of 2025-05-30
- Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation 17 upvotes, #21 of 2025-05-30
- StressTest: Can YOUR Speech LM Handle the Stress? 17 upvotes, #21 of 2025-05-30
- REOrdering Patches Improves Vision Models 16 upvotes, #23 of 2025-05-30
- DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning 15 upvotes, #24 of 2025-05-30
- On-Policy RL with Optimal Reward Baseline 14 upvotes, #25 of 2025-05-30
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model 14 upvotes, #25 of 2025-05-30
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts 12 upvotes, #27 of 2025-05-30
- SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents 12 upvotes, #27 of 2025-05-30
- PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions 11 upvotes, #29 of 2025-05-30
- GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control 11 upvotes, #29 of 2025-05-30
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? 10 upvotes, #31 of 2025-05-30
- Differentiable Solver Search for Fast Diffusion Sampling 10 upvotes, #31 of 2025-05-30
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction 9 upvotes, #33 of 2025-05-30
- MAGREF: Masked Guidance for Any-Reference Video Generation 9 upvotes, #33 of 2025-05-30
- Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction 8 upvotes, #35 of 2025-05-30
- ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind 8 upvotes, #35 of 2025-05-30
- One-shot Entropy Minimization 7 upvotes, #37 of 2025-05-30
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape 7 upvotes, #37 of 2025-05-30
- ATI: Any Trajectory Instruction for Controllable Video Generation 7 upvotes, #37 of 2025-05-30
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization 7 upvotes, #37 of 2025-05-30
- ZeroSep: Separate Anything in Audio with Zero Training 7 upvotes, #37 of 2025-05-30
- CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays 6 upvotes, #42 of 2025-05-30
- When Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy 6 upvotes, #42 of 2025-05-30
- Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting 5 upvotes, #44 of 2025-05-30
- CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting 5 upvotes, #44 of 2025-05-30
- UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes 5 upvotes, #44 of 2025-05-30
- To Trust Or Not To Trust Your Vision-Language Model's Prediction 5 upvotes, #44 of 2025-05-30
- Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint 5 upvotes, #44 of 2025-05-30
- A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models 4 upvotes, #49 of 2025-05-30
- ChartLens: Fine-grained Visual Attribution in Charts 4 upvotes, #49 of 2025-05-30
- Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation 4 upvotes, #49 of 2025-05-30
- SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model 4 upvotes, #49 of 2025-05-30
- Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates 4 upvotes, #49 of 2025-05-30
- ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS 4 upvotes, #49 of 2025-05-30
- How Animals Dance (When You're Not Looking) 4 upvotes, #49 of 2025-05-30
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation 3 upvotes, #56 of 2025-05-30
- Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator 3 upvotes, #56 of 2025-05-30
- GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents 3 upvotes, #56 of 2025-05-30
- Grounded Reinforcement Learning for Visual Reasoning 3 upvotes, #56 of 2025-05-30
- Differential Information: An Information-Theoretic Perspective on Preference Optimization 3 upvotes, #56 of 2025-05-30
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence 3 upvotes, #56 of 2025-05-30
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities 2 upvotes, #62 of 2025-05-30
- Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking 2 upvotes, #62 of 2025-05-30
- Model-Preserving Adaptive Rounding 2 upvotes, #62 of 2025-05-30
- Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement 2 upvotes, #62 of 2025-05-30
- Toward Reliable Biomedical Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models 1 upvotes, #66 of 2025-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.