Junjie Ye

Junjie Ye on Hugging Face Daily Papers: 13 papers, 2 in the top 3 of their day, 298 upvotes.

  1. SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 12 upvotes, #13 of 2026-08-12
  2. LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening 20 upvotes, #16 of 2026-05-21
  3. CCTU: A Benchmark for Tool Use under Complex Constraints 2 upvotes, #33 of 2026-03-18
  4. CL-bench: A Benchmark for Context Learning 22 upvotes, #16 of 2026-02-05
  5. Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning 18 upvotes, #18 of 2025-10-29
  6. Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels 12 upvotes, #14 of 2025-09-23
  7. AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning 55 upvotes, #3 of 2025-09-11
  8. Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments 16 upvotes, #13 of 2025-08-13
  9. CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios 10 upvotes, #16 of 2025-06-18
  10. A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models 9 upvotes, #7 of 2025-05-14
  11. Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training 84 upvotes, #1 of 2025-01-22
  12. ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use 9 upvotes, #14 of 2025-01-07
  13. MouSi: Poly-Visual-Expert Vision-Language Models 9 upvotes, #13 of 2024-01-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.