Daily Papers of 2025-05-28
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows 101 upvotes, #1 of 2025-05-28
- Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers 91 upvotes, #2 of 2025-05-28
- MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs 81 upvotes, #3 of 2025-05-28
- OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data 63 upvotes, #4 of 2025-05-28
- SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond 62 upvotes, #5 of 2025-05-28
- Exploring the Latent Capacity of LLMs for One-Step Text Generation 59 upvotes, #6 of 2025-05-28
- Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning 54 upvotes, #7 of 2025-05-28
- OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation 52 upvotes, #8 of 2025-05-28
- MMMR: Benchmarking Massive Multi-Modal Reasoning Tasks 45 upvotes, #9 of 2025-05-28
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence 44 upvotes, #10 of 2025-05-28
- VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization 41 upvotes, #11 of 2025-05-28
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation 38 upvotes, #12 of 2025-05-28
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents 38 upvotes, #12 of 2025-05-28
- GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning 36 upvotes, #15 of 2025-05-28
- SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use 31 upvotes, #16 of 2025-05-28
- rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset 26 upvotes, #17 of 2025-05-28
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? 26 upvotes, #17 of 2025-05-28
- Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems 24 upvotes, #20 of 2025-05-28
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks 20 upvotes, #21 of 2025-05-28
- Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL 19 upvotes, #22 of 2025-05-28
- MotionPro: A Precise Motion Controller for Image-to-Video Generation 19 upvotes, #22 of 2025-05-28
- HoliTom: Holistic Token Merging for Fast Video Large Language Models 18 upvotes, #24 of 2025-05-28
- NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI 17 upvotes, #25 of 2025-05-28
- ImgEdit: A Unified Image Editing Dataset and Benchmark 17 upvotes, #25 of 2025-05-28
- Frame In-N-Out: Unbounded Controllable Image-to-Video Generation 17 upvotes, #25 of 2025-05-28
- How does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective 17 upvotes, #25 of 2025-05-28
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms 14 upvotes, #29 of 2025-05-28
- Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO 14 upvotes, #29 of 2025-05-28
- FinTagging: An LLM-ready Benchmark for Extracting and Structuring Financial Information 13 upvotes, #31 of 2025-05-28
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction 13 upvotes, #31 of 2025-05-28
- Rendering-Aware Reinforcement Learning for Vector Graphics Generation 11 upvotes, #33 of 2025-05-28
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models 11 upvotes, #33 of 2025-05-28
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection 10 upvotes, #35 of 2025-05-28
- Thinker: Learning to Think Fast and Slow 10 upvotes, #35 of 2025-05-28
- MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation 8 upvotes, #37 of 2025-05-28
- SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning 8 upvotes, #37 of 2025-05-28
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment 8 upvotes, #37 of 2025-05-28
- VideoGameBench: Can Vision-Language Models complete popular video games? 6 upvotes, #40 of 2025-05-28
- Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution 6 upvotes, #40 of 2025-05-28
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness 6 upvotes, #40 of 2025-05-28
- Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning 6 upvotes, #40 of 2025-05-28
- Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning 6 upvotes, #40 of 2025-05-28
- Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs 5 upvotes, #45 of 2025-05-28
- R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 5 upvotes, #45 of 2025-05-28
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression 5 upvotes, #45 of 2025-05-28
- BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases 5 upvotes, #45 of 2025-05-28
- Minute-Long Videos with Dual Parallelisms 5 upvotes, #45 of 2025-05-28
- Sci-Fi: Symmetric Constraint for Frame Inbetweening 5 upvotes, #45 of 2025-05-28
- Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration 5 upvotes, #45 of 2025-05-28
- Reverse Preference Optimization for Complex Instruction Following 5 upvotes, #45 of 2025-05-28
- MLLMs are Deeply Affected by Modality Bias 4 upvotes, #53 of 2025-05-28
- SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline 4 upvotes, #53 of 2025-05-28
- Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval 4 upvotes, #53 of 2025-05-28
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 4 upvotes, #53 of 2025-05-28
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback 3 upvotes, #57 of 2025-05-28
- DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response 3 upvotes, #57 of 2025-05-28
- Capability-Based Scaling Laws for LLM Red-Teaming 3 upvotes, #57 of 2025-05-28
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals 3 upvotes, #57 of 2025-05-28
- Spatial Knowledge Graph-Guided Multimodal Synthesis 3 upvotes, #57 of 2025-05-28
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO 2 upvotes, #62 of 2025-05-28
- PreMoe: Lightening MoEs on Constrained Memory by Expert Pruning and Retrieval 2 upvotes, #62 of 2025-05-28
- SATORI-R1: Incentivizing Multimodal Reasoning with Spatial Grounding and Verifiable Rewards 2 upvotes, #62 of 2025-05-28
- AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery 2 upvotes, #62 of 2025-05-28
- Do RAG Systems Suffer From Positional Bias? 1 upvotes, #66 of 2025-05-28
- Improving Chemical Understanding of LLMs via SMILES Parsing 1 upvotes, #66 of 2025-05-28
- Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms 1 upvotes, #66 of 2025-05-28
- Explaining Sources of Uncertainty in Automated Fact-Checking 1 upvotes, #66 of 2025-05-28
- CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models 1 upvotes, #66 of 2025-05-28
- Absolute Coordinates Make Motion Generation Easy 1 upvotes, #66 of 2025-05-28
- Knowledge Base Construction for Knowledge-Augmented Text-to-SQL 1 upvotes, #66 of 2025-05-28
- An Explainable Diagnostic Framework for Neurodegenerative Dementias via Reinforcement-Optimized LLM Reasoning 0 upvotes, #73 of 2025-05-28
- Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction 0 upvotes, #73 of 2025-05-28
- Ankh3: Multi-Task Pretraining with Sequence Denoising and Completion Enhances Protein Representations 0 upvotes, #73 of 2025-05-28
- Vision Transformers with Self-Distilled Registers 0 upvotes, #73 of 2025-05-28
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.