Seungone Kim

Seungone Kim on Hugging Face Daily Papers: 21 papers, 5 in the top 3 of their day, 765 upvotes.

  1. K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts 55 upvotes, #6 of 2026-06-02
  2. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization 6 upvotes, #43 of 2026-05-28
  3. On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists 11 upvotes, #20 of 2026-05-21
  4. Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs 77 upvotes, #2 of 2026-05-12
  5. Reasoning over mathematical objects: on-policy reward modeling and test time aggregation 6 upvotes, #25 of 2026-03-20
  6. RefineBench: Evaluating Refinement Capability of Language Models via Checklists 12 upvotes, #15 of 2025-12-01
  7. SPICE: Self-Play In Corpus Environments Improves Reasoning 12 upvotes, #25 of 2025-10-29
  8. Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning 62 upvotes, #2 of 2025-07-02
  9. Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability 13 upvotes, #23 of 2025-06-04
  10. Let's Predict Sentence by Sentence 17 upvotes, #18 of 2025-05-29
  11. FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS 2 upvotes, #41 of 2025-05-26
  12. Web-Shepherd: Advancing PRMs for Reinforcing Web Agents 96 upvotes, #1 of 2025-05-22
  13. Reasoning Models Better Express Their Confidence 18 upvotes, #11 of 2025-05-21
  14. The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think 24 upvotes, #6 of 2025-05-16
  15. Bridging the Data Provenance Gap Across Text, Speech and Video 6 upvotes, #10 of 2024-12-25
  16. Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
  17. Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages 41 upvotes, #8 of 2024-10-22
  18. Consent in Crisis: The Rapid Decline of the AI Data Commons 9 upvotes, #11 of 2024-07-23
  19. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models 79 upvotes, #2 of 2024-05-03
  20. Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models 44 upvotes, #3 of 2024-04-04
  21. FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets 14 upvotes, #4 of 2023-07-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.