Yilun Zhao

Yilun Zhao on Hugging Face Daily Papers: 43 papers, 6 in the top 3 of their day, 1,219 upvotes.

  1. GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents 28 upvotes, #8 of 2026-06-26
  2. VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding 35 upvotes, #5 of 2026-06-05
  3. OpenComputer: Verifiable Software Worlds for Computer-Use Agents 57 upvotes, #5 of 2026-05-20
  4. Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems 38 upvotes, #7 of 2026-05-07
  5. Step-level Optimization for Efficient Computer-use Agents 18 upvotes, #14 of 2026-05-01
  6. TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction 6 upvotes, #17 of 2026-04-28
  7. RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation 6 upvotes, #17 of 2026-03-12
  8. ANCHOR: Branch-Point Data Generation for GUI Agents 5 upvotes, #35 of 2026-02-11
  9. SAGE: Benchmarking and Improving Retrieval for Deep Research Agents 12 upvotes, #23 of 2026-02-06
  10. Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing 13 upvotes, #14 of 2026-01-23
  11. Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs 140 upvotes, #3 of 2026-01-16
  12. Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL 6 upvotes, #30 of 2026-01-16
  13. AlphaResearch: Accelerating New Algorithm Discovery with Language Models 13 upvotes, #8 of 2025-11-14
  14. LimRank: Less is More for Reasoning-Intensive Information Reranking 8 upvotes, #18 of 2025-10-28
  15. FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain 5 upvotes, #21 of 2025-10-20
  16. MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval 7 upvotes, #20 of 2025-10-13
  17. PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles 1 upvotes, #41 of 2025-10-09
  18. FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering 2 upvotes, #37 of 2025-10-09
  19. AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research 16 upvotes, #11 of 2025-07-18
  20. Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers 11 upvotes, #6 of 2025-07-16
  21. Efficiency-Effectiveness Reranking FLOPs for LLM-based Rerankers 13 upvotes, #16 of 2025-07-09
  22. Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers 18 upvotes, #10 of 2025-07-04
  23. SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks 42 upvotes, #3 of 2025-07-02
  24. SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification 11 upvotes, #8 of 2025-06-19
  25. Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure 16 upvotes, #12 of 2025-06-18
  26. MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation 88 upvotes, #1 of 2025-06-18
  27. Table-R1: Inference-Time Scaling for Table Reasoning 88 upvotes, #1 of 2025-05-30
  28. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 56 upvotes, #4 of 2025-05-30
  29. Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 52 upvotes, #4 of 2025-05-22
  30. Z1: Efficient Test-time Scaling with Code 25 upvotes, #9 of 2025-04-02
  31. PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving 16 upvotes, #12 of 2025-03-31
  32. MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search 9 upvotes, #12 of 2025-03-27
  33. Survey on Evaluation of LLM-based Agents 78 upvotes, #2 of 2025-03-21
  34. MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning 14 upvotes, #17 of 2025-03-11
  35. IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval 20 upvotes, #8 of 2025-03-07
  36. MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 79 upvotes, #2 of 2025-01-22
  37. ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning 8 upvotes, #11 of 2025-01-14
  38. HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation 9 upvotes, #13 of 2024-12-31
  39. M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models 14 upvotes, #15 of 2024-11-08
  40. TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models 19 upvotes, #4 of 2024-11-04
  41. Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications 42 upvotes, #4 of 2024-08-23
  42. ML-Bench: Large Language Models Leverage Open-source Libraries for Machine Learning Tasks 10 upvotes, #6 of 2023-11-17
  43. Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? 11 upvotes, #10 of 2023-09-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.