Daily Papers of 2025-06-18

  1. MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation 88 upvotes, #1 of 2025-06-18
  2. Scaling Test-time Compute for LLM Agents 57 upvotes, #2 of 2025-06-18
  3. CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following 53 upvotes, #3 of 2025-06-18
  4. LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs 43 upvotes, #4 of 2025-06-18
  5. Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team 38 upvotes, #5 of 2025-06-18
  6. Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs 35 upvotes, #6 of 2025-06-18
  7. Efficient Medical VIE via Reinforcement Learning 27 upvotes, #7 of 2025-06-18
  8. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning 26 upvotes, #8 of 2025-06-18
  9. Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model 26 upvotes, #8 of 2025-06-18
  10. Reasoning with Exploration: An Entropy Perspective 26 upvotes, #8 of 2025-06-18
  11. Align Your Flow: Scaling Continuous-Time Flow Map Distillation 17 upvotes, #11 of 2025-06-18
  12. Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure 16 upvotes, #12 of 2025-06-18
  13. QFFT, Question-Free Fine-Tuning for Adaptive Reasoning 16 upvotes, #12 of 2025-06-18
  14. From Bytes to Ideas: Language Modeling with Autoregressive U-Nets 13 upvotes, #14 of 2025-06-18
  15. Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees 11 upvotes, #15 of 2025-06-18
  16. VideoMolmo: Spatio-Temporal Grounding Meets Pointing 10 upvotes, #16 of 2025-06-18
  17. EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models 10 upvotes, #16 of 2025-06-18
  18. CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios 10 upvotes, #16 of 2025-06-18
  19. Optimizing Length Compression in Large Reasoning Models 10 upvotes, #16 of 2025-06-18
  20. Ambient Diffusion Omni: Training Good Models with Bad Data 9 upvotes, #20 of 2025-06-18
  21. xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations 9 upvotes, #20 of 2025-06-18
  22. Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs 9 upvotes, #20 of 2025-06-18
  23. Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning 6 upvotes, #23 of 2025-06-18
  24. Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders 6 upvotes, #23 of 2025-06-18
  25. AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents 6 upvotes, #23 of 2025-06-18
  26. Mixture-of-Experts Meets In-Context Reinforcement Learning 5 upvotes, #26 of 2025-06-18
  27. Universal Jailbreak Suffixes Are Strong Attention Hijackers 5 upvotes, #26 of 2025-06-18
  28. CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility Simulation 4 upvotes, #28 of 2025-06-18
  29. Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning 3 upvotes, #29 of 2025-06-18
  30. EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction 3 upvotes, #29 of 2025-06-18
  31. TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Scale-Oriented Contrast 3 upvotes, #29 of 2025-06-18
  32. Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations 3 upvotes, #29 of 2025-06-18
  33. Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers 3 upvotes, #29 of 2025-06-18
  34. VisText-Mosquito: A Multimodal Dataset and Benchmark for AI-Based Mosquito Breeding Site Detection and Reasoning 2 upvotes, #34 of 2025-06-18
  35. DynaGuide: Steering Diffusion Polices with Active Dynamic Guidance 1 upvotes, #35 of 2025-06-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.