Yilun Zhao
Yilun Zhao on Hugging Face Daily Papers: 43 papers, 6 in the top 3 of their day, 1,219 upvotes.
- GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents 28 upvotes, #8 of 2026-06-26
- VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding 35 upvotes, #5 of 2026-06-05
- OpenComputer: Verifiable Software Worlds for Computer-Use Agents 57 upvotes, #5 of 2026-05-20
- Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems 38 upvotes, #7 of 2026-05-07
- Step-level Optimization for Efficient Computer-use Agents 18 upvotes, #14 of 2026-05-01
- TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction 6 upvotes, #17 of 2026-04-28
- RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation 6 upvotes, #17 of 2026-03-12
- ANCHOR: Branch-Point Data Generation for GUI Agents 5 upvotes, #35 of 2026-02-11
- SAGE: Benchmarking and Improving Retrieval for Deep Research Agents 12 upvotes, #23 of 2026-02-06
- Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing 13 upvotes, #14 of 2026-01-23
- Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs 140 upvotes, #3 of 2026-01-16
- Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL 6 upvotes, #30 of 2026-01-16
- AlphaResearch: Accelerating New Algorithm Discovery with Language Models 13 upvotes, #8 of 2025-11-14
- LimRank: Less is More for Reasoning-Intensive Information Reranking 8 upvotes, #18 of 2025-10-28
- FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain 5 upvotes, #21 of 2025-10-20
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval 7 upvotes, #20 of 2025-10-13
- PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles 1 upvotes, #41 of 2025-10-09
- FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering 2 upvotes, #37 of 2025-10-09
- AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research 16 upvotes, #11 of 2025-07-18
- Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers 11 upvotes, #6 of 2025-07-16
- Efficiency-Effectiveness Reranking FLOPs for LLM-based Rerankers 13 upvotes, #16 of 2025-07-09
- Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers 18 upvotes, #10 of 2025-07-04
- SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks 42 upvotes, #3 of 2025-07-02
- SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification 11 upvotes, #8 of 2025-06-19
- Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure 16 upvotes, #12 of 2025-06-18
- MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation 88 upvotes, #1 of 2025-06-18
- Table-R1: Inference-Time Scaling for Table Reasoning 88 upvotes, #1 of 2025-05-30
- VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 56 upvotes, #4 of 2025-05-30
- Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 52 upvotes, #4 of 2025-05-22
- Z1: Efficient Test-time Scaling with Code 25 upvotes, #9 of 2025-04-02
- PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving 16 upvotes, #12 of 2025-03-31
- MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search 9 upvotes, #12 of 2025-03-27
- Survey on Evaluation of LLM-based Agents 78 upvotes, #2 of 2025-03-21
- MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning 14 upvotes, #17 of 2025-03-11
- IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval 20 upvotes, #8 of 2025-03-07
- MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 79 upvotes, #2 of 2025-01-22
- ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning 8 upvotes, #11 of 2025-01-14
- HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation 9 upvotes, #13 of 2024-12-31
- M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models 14 upvotes, #15 of 2024-11-08
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models 19 upvotes, #4 of 2024-11-04
- Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications 42 upvotes, #4 of 2024-08-23
- ML-Bench: Large Language Models Leverage Open-source Libraries for Machine Learning Tasks 10 upvotes, #6 of 2023-11-17
- Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? 11 upvotes, #10 of 2023-09-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.