Bill Yuchen Lin

Bill Yuchen Lin on Hugging Face Daily Papers: 21 papers, 9 in the top 3 of their day, 656 upvotes.

  1. CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation 8 upvotes, #15 of 2025-04-09
  2. Small Models Struggle to Learn from Strong Reasoners 27 upvotes, #7 of 2025-02-20
  3. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning 14 upvotes, #13 of 2025-02-04
  4. Stronger Models are NOT Stronger Teachers for Instruction Tuning 27 upvotes, #1 of 2024-11-13
  5. On Memorization of Large Language Models in Logical Reasoning 14 upvotes, #6 of 2024-10-31
  6. The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism 19 upvotes, #4 of 2024-07-16
  7. WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs 11 upvotes, #6 of 2024-06-27
  8. MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24
  9. WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 11 upvotes, #16 of 2024-06-18
  10. Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing 42 upvotes, #2 of 2024-06-13
  11. WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild 23 upvotes, #3 of 2024-06-10
  12. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models 79 upvotes, #2 of 2024-05-03
  13. RewardBench: Evaluating Reward Models for Language Modeling 14 upvotes, #9 of 2024-03-21
  14. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 84 upvotes, #1 of 2024-02-23
  15. L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects 18 upvotes, #3 of 2024-02-15
  16. The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning 31 upvotes, #3 of 2023-12-05
  17. Lumos: Learning Agents with Unified Data, Modular Design, and Open-Source LLMs 28 upvotes, #4 of 2023-11-13
  18. LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition 34 upvotes, #1 of 2023-07-26
  19. LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion 6 upvotes, #4 of 2023-06-06
  20. Faith and Fate: Limits of Transformers on Compositionality 9 upvotes, #1 of 2023-05-31
  21. SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks 4 upvotes, #8 of 2023-05-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.