Jindong Wang
Jindong Wang on Hugging Face Daily Papers: 11 papers, 2 in the top 3 of their day, 215 upvotes.
- Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach 14 upvotes, #21 of 2025-05-29
- Outcome-Refining Process Supervision for Code Generation 16 upvotes, #11 of 2024-12-24
- MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders 5 upvotes, #40 of 2024-10-10
- Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist 18 upvotes, #7 of 2024-07-12
- TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
- Supervised Knowledge Makes Large Language Models Better In-context Learners 9 upvotes, #11 of 2023-12-27
- PromptBench: A Unified Library for Evaluation of Large Language Models 16 upvotes, #4 of 2023-12-14
- How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation 8 upvotes, #13 of 2023-12-13
- A Survey on Evaluation of Large Language Models 43 upvotes, #2 of 2023-07-07
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization 7 upvotes, #4 of 2023-06-09
- PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts 3 upvotes, #5 of 2023-06-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.