Daily Papers of 2025-05-28

  1. ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows 101 upvotes, #1 of 2025-05-28
  2. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers 91 upvotes, #2 of 2025-05-28
  3. MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs 81 upvotes, #3 of 2025-05-28
  4. OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data 63 upvotes, #4 of 2025-05-28
  5. SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond 62 upvotes, #5 of 2025-05-28
  6. Exploring the Latent Capacity of LLMs for One-Step Text Generation 59 upvotes, #6 of 2025-05-28
  7. Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning 54 upvotes, #7 of 2025-05-28
  8. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation 52 upvotes, #8 of 2025-05-28
  9. MMMR: Benchmarking Massive Multi-Modal Reasoning Tasks 45 upvotes, #9 of 2025-05-28
  10. Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence 44 upvotes, #10 of 2025-05-28
  11. VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization 41 upvotes, #11 of 2025-05-28
  12. Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation 38 upvotes, #12 of 2025-05-28
  13. MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
  14. UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents 38 upvotes, #12 of 2025-05-28
  15. GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning 36 upvotes, #15 of 2025-05-28
  16. SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use 31 upvotes, #16 of 2025-05-28
  17. rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset 26 upvotes, #17 of 2025-05-28
  18. Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? 26 upvotes, #17 of 2025-05-28
  19. Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
  20. MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems 24 upvotes, #20 of 2025-05-28
  21. Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks 20 upvotes, #21 of 2025-05-28
  22. Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL 19 upvotes, #22 of 2025-05-28
  23. MotionPro: A Precise Motion Controller for Image-to-Video Generation 19 upvotes, #22 of 2025-05-28
  24. HoliTom: Holistic Token Merging for Fast Video Large Language Models 18 upvotes, #24 of 2025-05-28
  25. NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI 17 upvotes, #25 of 2025-05-28
  26. ImgEdit: A Unified Image Editing Dataset and Benchmark 17 upvotes, #25 of 2025-05-28
  27. Frame In-N-Out: Unbounded Controllable Image-to-Video Generation 17 upvotes, #25 of 2025-05-28
  28. How does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective 17 upvotes, #25 of 2025-05-28
  29. Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms 14 upvotes, #29 of 2025-05-28
  30. Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO 14 upvotes, #29 of 2025-05-28
  31. FinTagging: An LLM-ready Benchmark for Extracting and Structuring Financial Information 13 upvotes, #31 of 2025-05-28
  32. DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction 13 upvotes, #31 of 2025-05-28
  33. Rendering-Aware Reinforcement Learning for Vector Graphics Generation 11 upvotes, #33 of 2025-05-28
  34. ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models 11 upvotes, #33 of 2025-05-28
  35. VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection 10 upvotes, #35 of 2025-05-28
  36. Thinker: Learning to Think Fast and Slow 10 upvotes, #35 of 2025-05-28
  37. MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation 8 upvotes, #37 of 2025-05-28
  38. SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning 8 upvotes, #37 of 2025-05-28
  39. Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment 8 upvotes, #37 of 2025-05-28
  40. VideoGameBench: Can Vision-Language Models complete popular video games? 6 upvotes, #40 of 2025-05-28
  41. Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution 6 upvotes, #40 of 2025-05-28
  42. MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness 6 upvotes, #40 of 2025-05-28
  43. Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning 6 upvotes, #40 of 2025-05-28
  44. Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning 6 upvotes, #40 of 2025-05-28
  45. Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs 5 upvotes, #45 of 2025-05-28
  46. R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 5 upvotes, #45 of 2025-05-28
  47. Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression 5 upvotes, #45 of 2025-05-28
  48. BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases 5 upvotes, #45 of 2025-05-28
  49. Minute-Long Videos with Dual Parallelisms 5 upvotes, #45 of 2025-05-28
  50. Sci-Fi: Symmetric Constraint for Frame Inbetweening 5 upvotes, #45 of 2025-05-28
  51. Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration 5 upvotes, #45 of 2025-05-28
  52. Reverse Preference Optimization for Complex Instruction Following 5 upvotes, #45 of 2025-05-28
  53. MLLMs are Deeply Affected by Modality Bias 4 upvotes, #53 of 2025-05-28
  54. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline 4 upvotes, #53 of 2025-05-28
  55. Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval 4 upvotes, #53 of 2025-05-28
  56. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 4 upvotes, #53 of 2025-05-28
  57. ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback 3 upvotes, #57 of 2025-05-28
  58. DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response 3 upvotes, #57 of 2025-05-28
  59. Capability-Based Scaling Laws for LLM Red-Teaming 3 upvotes, #57 of 2025-05-28
  60. Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals 3 upvotes, #57 of 2025-05-28
  61. Spatial Knowledge Graph-Guided Multimodal Synthesis 3 upvotes, #57 of 2025-05-28
  62. R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO 2 upvotes, #62 of 2025-05-28
  63. PreMoe: Lightening MoEs on Constrained Memory by Expert Pruning and Retrieval 2 upvotes, #62 of 2025-05-28
  64. SATORI-R1: Incentivizing Multimodal Reasoning with Spatial Grounding and Verifiable Rewards 2 upvotes, #62 of 2025-05-28
  65. AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery 2 upvotes, #62 of 2025-05-28
  66. Do RAG Systems Suffer From Positional Bias? 1 upvotes, #66 of 2025-05-28
  67. Improving Chemical Understanding of LLMs via SMILES Parsing 1 upvotes, #66 of 2025-05-28
  68. Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms 1 upvotes, #66 of 2025-05-28
  69. Explaining Sources of Uncertainty in Automated Fact-Checking 1 upvotes, #66 of 2025-05-28
  70. CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models 1 upvotes, #66 of 2025-05-28
  71. Absolute Coordinates Make Motion Generation Easy 1 upvotes, #66 of 2025-05-28
  72. Knowledge Base Construction for Knowledge-Augmented Text-to-SQL 1 upvotes, #66 of 2025-05-28
  73. An Explainable Diagnostic Framework for Neurodegenerative Dementias via Reinforcement-Optimized LLM Reasoning 0 upvotes, #73 of 2025-05-28
  74. Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction 0 upvotes, #73 of 2025-05-28
  75. Ankh3: Multi-Task Pretraining with Sequence Denoising and Completion Enhances Protein Representations 0 upvotes, #73 of 2025-05-28
  76. Vision Transformers with Self-Distilled Registers 0 upvotes, #73 of 2025-05-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.