Daily Papers of 2025-06-18
- MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation 88 upvotes, #1 of 2025-06-18
- Scaling Test-time Compute for LLM Agents 57 upvotes, #2 of 2025-06-18
- CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following 53 upvotes, #3 of 2025-06-18
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs 43 upvotes, #4 of 2025-06-18
- Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team 38 upvotes, #5 of 2025-06-18
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs 35 upvotes, #6 of 2025-06-18
- Efficient Medical VIE via Reinforcement Learning 27 upvotes, #7 of 2025-06-18
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning 26 upvotes, #8 of 2025-06-18
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model 26 upvotes, #8 of 2025-06-18
- Reasoning with Exploration: An Entropy Perspective 26 upvotes, #8 of 2025-06-18
- Align Your Flow: Scaling Continuous-Time Flow Map Distillation 17 upvotes, #11 of 2025-06-18
- Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure 16 upvotes, #12 of 2025-06-18
- QFFT, Question-Free Fine-Tuning for Adaptive Reasoning 16 upvotes, #12 of 2025-06-18
- From Bytes to Ideas: Language Modeling with Autoregressive U-Nets 13 upvotes, #14 of 2025-06-18
- Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees 11 upvotes, #15 of 2025-06-18
- VideoMolmo: Spatio-Temporal Grounding Meets Pointing 10 upvotes, #16 of 2025-06-18
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models 10 upvotes, #16 of 2025-06-18
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios 10 upvotes, #16 of 2025-06-18
- Optimizing Length Compression in Large Reasoning Models 10 upvotes, #16 of 2025-06-18
- Ambient Diffusion Omni: Training Good Models with Bad Data 9 upvotes, #20 of 2025-06-18
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations 9 upvotes, #20 of 2025-06-18
- Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs 9 upvotes, #20 of 2025-06-18
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning 6 upvotes, #23 of 2025-06-18
- Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders 6 upvotes, #23 of 2025-06-18
- AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents 6 upvotes, #23 of 2025-06-18
- Mixture-of-Experts Meets In-Context Reinforcement Learning 5 upvotes, #26 of 2025-06-18
- Universal Jailbreak Suffixes Are Strong Attention Hijackers 5 upvotes, #26 of 2025-06-18
- CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility Simulation 4 upvotes, #28 of 2025-06-18
- Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning 3 upvotes, #29 of 2025-06-18
- EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction 3 upvotes, #29 of 2025-06-18
- TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Scale-Oriented Contrast 3 upvotes, #29 of 2025-06-18
- Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations 3 upvotes, #29 of 2025-06-18
- Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers 3 upvotes, #29 of 2025-06-18
- VisText-Mosquito: A Multimodal Dataset and Benchmark for AI-Based Mosquito Breeding Site Detection and Reasoning 2 upvotes, #34 of 2025-06-18
- DynaGuide: Steering Diffusion Polices with Active Dynamic Guidance 1 upvotes, #35 of 2025-06-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.