Junjie Ye
Junjie Ye on Hugging Face Daily Papers: 13 papers, 2 in the top 3 of their day, 298 upvotes.
- SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 12 upvotes, #13 of 2026-08-12
- LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening 20 upvotes, #16 of 2026-05-21
- CCTU: A Benchmark for Tool Use under Complex Constraints 2 upvotes, #33 of 2026-03-18
- CL-bench: A Benchmark for Context Learning 22 upvotes, #16 of 2026-02-05
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning 18 upvotes, #18 of 2025-10-29
- Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels 12 upvotes, #14 of 2025-09-23
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning 55 upvotes, #3 of 2025-09-11
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments 16 upvotes, #13 of 2025-08-13
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios 10 upvotes, #16 of 2025-06-18
- A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models 9 upvotes, #7 of 2025-05-14
- Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training 84 upvotes, #1 of 2025-01-22
- ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use 9 upvotes, #14 of 2025-01-07
- MouSi: Poly-Visual-Expert Vision-Language Models 9 upvotes, #13 of 2024-01-31
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.