Daily Papers of 2025-04-24
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
- Tina: Tiny Reasoning Models via LoRA 50 upvotes, #2 of 2025-04-24
- DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning 48 upvotes, #3 of 2025-04-24
- Trillion 7B Technical Report 34 upvotes, #4 of 2025-04-24
- PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models 34 upvotes, #4 of 2025-04-24
- I-Con: A Unifying Framework for Representation Learning 29 upvotes, #6 of 2025-04-24
- DreamO: A Unified Framework for Image Customization 19 upvotes, #7 of 2025-04-24
- Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model 18 upvotes, #8 of 2025-04-24
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset 18 upvotes, #8 of 2025-04-24
- Decoupled Global-Local Alignment for Improving Compositional Understanding 15 upvotes, #10 of 2025-04-24
- A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment 13 upvotes, #11 of 2025-04-24
- Rethinking the Generation of High-Quality CoT Data from the Perspective of LLM-Adaptive Question Difficulty Grading 12 upvotes, #12 of 2025-04-24
- RePOPE: Impact of Annotation Errors on the POPE Benchmark 8 upvotes, #13 of 2025-04-24
- CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation 6 upvotes, #14 of 2025-04-24
- Causal-Copilot: An Autonomous Causal Analysis Agent 5 upvotes, #15 of 2025-04-24
- Unchecked and Overlooked: Addressing the Checkbox Blind Spot in Large Language Models with CheckboxQA 4 upvotes, #16 of 2025-04-24
- Progressive Language-guided Visual Learning for Multi-Task Visual Grounding 2 upvotes, #17 of 2025-04-24
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.