Daily Papers of 2025-08-21

  1. DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization 78 upvotes, #1 of 2025-08-21
  2. MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds 62 upvotes, #2 of 2025-08-21
  3. FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction 61 upvotes, #3 of 2025-08-21
  4. From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models 58 upvotes, #4 of 2025-08-21
  5. MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers 39 upvotes, #5 of 2025-08-21
  6. Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization 39 upvotes, #5 of 2025-08-21
  7. From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery 31 upvotes, #7 of 2025-08-21
  8. NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model 31 upvotes, #7 of 2025-08-21
  9. Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs 19 upvotes, #9 of 2025-08-21
  10. RynnEC: Bringing MLLMs into Embodied World 18 upvotes, #10 of 2025-08-21
  11. Virtuous Machines: Towards Artificial General Science 9 upvotes, #11 of 2025-08-21
  12. On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting 6 upvotes, #12 of 2025-08-21
  13. FLARE: Fast Low-rank Attention Routing Engine 6 upvotes, #12 of 2025-08-21
  14. ViExam: Are Vision Language Models Better than Humans on Vietnamese Multimodal Exam Questions? 5 upvotes, #14 of 2025-08-21
  15. Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer 3 upvotes, #15 of 2025-08-21
  16. Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis 3 upvotes, #15 of 2025-08-21
  17. mSCoRe: a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning 1 upvotes, #17 of 2025-08-21
  18. Leuvenshtein: Efficient FHE-based Edit Distance Computation with Single Bootstrap per Cell 1 upvotes, #17 of 2025-08-21
  19. Refining Contrastive Learning and Homography Relations for Multi-Modal Recommendation 1 upvotes, #19 of 2025-08-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.