Graham Neubig

Graham Neubig on Hugging Face Daily Papers: 21 papers, 8 in the top 3 of their day, 780 upvotes.

  1. The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think 24 upvotes, #6 of 2025-05-16
  2. Demystifying Long Chain-of-Thought Reasoning in LLMs 49 upvotes, #2 of 2025-02-06
  3. TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks 43 upvotes, #3 of 2024-12-19
  4. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
  5. Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
  6. OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs 25 upvotes, #5 of 2024-11-22
  7. JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation 12 upvotes, #7 of 2024-10-23
  8. Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages 41 upvotes, #8 of 2024-10-22
  9. NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples 35 upvotes, #3 of 2024-10-21
  10. Harnessing Webpage UIs for Text-Rich Visual Understanding 28 upvotes, #7 of 2024-10-18
  11. Agent Workflow Memory 25 upvotes, #3 of 2024-09-12
  12. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
  13. OpenDevin: An Open Platform for AI Software Developers as Generalist Agents 62 upvotes, #1 of 2024-07-25
  14. VIMI: Grounding Video Generation through Multi-modal Instruction 8 upvotes, #13 of 2024-07-10
  15. Training Task Experts through Retrieval Based Distillation 6 upvotes, #13 of 2024-07-09
  16. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models 79 upvotes, #2 of 2024-05-03
  17. SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents 18 upvotes, #4 of 2024-03-14
  18. Instruction-tuned Language Models are Better Knowledge Learners 26 upvotes, #5 of 2024-02-21
  19. Alignment for Honesty 13 upvotes, #5 of 2023-12-13
  20. The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation 7 upvotes, #8 of 2023-08-15
  21. WebArena: A Realistic Web Environment for Building Autonomous Agents 27 upvotes, #3 of 2023-07-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.