GUIJIN SON
GUIJIN SON on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 309 upvotes.
- ResearchMath-14K: Scaling Research-Level Mathematics via Agents 49 upvotes, #6 of 2026-05-28
- Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback 19 upvotes, #12 of 2026-05-25
- Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs 77 upvotes, #2 of 2026-05-12
- Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math 22 upvotes, #10 of 2026-02-09
- What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models 15 upvotes, #15 of 2026-01-13
- Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures 9 upvotes, #26 of 2025-10-29
- Revisiting the Uniform Information Density Hypothesis in LLM Reasoning Traces 6 upvotes, #24 of 2025-10-09
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought 24 upvotes, #11 of 2025-10-09
- From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation 16 upvotes, #8 of 2025-07-15
- BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation 8 upvotes, #22 of 2025-06-05
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research 9 upvotes, #23 of 2025-05-20
- Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning 24 upvotes, #7 of 2025-02-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.