Seungone Kim
Seungone Kim on Hugging Face Daily Papers: 21 papers, 5 in the top 3 of their day, 765 upvotes.
- K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts 55 upvotes, #6 of 2026-06-02
- Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization 6 upvotes, #43 of 2026-05-28
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists 11 upvotes, #20 of 2026-05-21
- Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs 77 upvotes, #2 of 2026-05-12
- Reasoning over mathematical objects: on-policy reward modeling and test time aggregation 6 upvotes, #25 of 2026-03-20
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists 12 upvotes, #15 of 2025-12-01
- SPICE: Self-Play In Corpus Environments Improves Reasoning 12 upvotes, #25 of 2025-10-29
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning 62 upvotes, #2 of 2025-07-02
- Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability 13 upvotes, #23 of 2025-06-04
- Let's Predict Sentence by Sentence 17 upvotes, #18 of 2025-05-29
- FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS 2 upvotes, #41 of 2025-05-26
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents 96 upvotes, #1 of 2025-05-22
- Reasoning Models Better Express Their Confidence 18 upvotes, #11 of 2025-05-21
- The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think 24 upvotes, #6 of 2025-05-16
- Bridging the Data Provenance Gap Across Text, Speech and Video 6 upvotes, #10 of 2024-12-25
- Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
- Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages 41 upvotes, #8 of 2024-10-22
- Consent in Crisis: The Rapid Decline of the AI Data Commons 9 upvotes, #11 of 2024-07-23
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models 79 upvotes, #2 of 2024-05-03
- Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models 44 upvotes, #3 of 2024-04-04
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets 14 upvotes, #4 of 2023-07-21
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.