Jiaheng Liu

Jiaheng Liu on Hugging Face Daily Papers: 32 papers, 11 in the top 3 of their day, 1,082 upvotes.

  1. Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? 54 upvotes, #4 of 2025-09-05
  2. Flow-GRPO: Training Flow Matching Models via Online RL 68 upvotes, #2 of 2025-05-09
  3. IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs 22 upvotes, #8 of 2025-04-23
  4. Mavors: Multi-granularity Video Representation for Multimodal Large Language Model 30 upvotes, #7 of 2025-04-15
  5. COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values 41 upvotes, #5 of 2025-04-09
  6. YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
  7. CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models 23 upvotes, #8 of 2025-02-25
  8. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
  9. Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
  10. ProgCo: Program Helps Self-Correction of Large Language Models 24 upvotes, #7 of 2025-01-03
  11. PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos 6 upvotes, #21 of 2024-12-03
  12. Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models 33 upvotes, #4 of 2024-11-12
  13. M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation 6 upvotes, #16 of 2024-11-04
  14. AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions 36 upvotes, #2 of 2024-10-30
  15. Can MLLMs Understand the Deep Implication Behind Chinese Images? 7 upvotes, #23 of 2024-10-18
  16. PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment 18 upvotes, #14 of 2024-10-18
  17. MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models 17 upvotes, #7 of 2024-10-16
  18. ING-VP: MLLMs cannot Play Easy Vision-based Games Yet 8 upvotes, #26 of 2024-10-10
  19. MIO: A Foundation Model on Multimodal Tokens 46 upvotes, #2 of 2024-09-30
  20. HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models 39 upvotes, #1 of 2024-09-25
  21. OmniBench: Towards The Future of Universal Omni-Language Models 24 upvotes, #4 of 2024-09-25
  22. FuzzCoder: Byte-level Fuzzing Test via Large Language Model 44 upvotes, #3 of 2024-09-06
  23. TableBench: A Comprehensive and Complex Benchmark for Table Question Answering 48 upvotes, #1 of 2024-08-21
  24. I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm 30 upvotes, #2 of 2024-08-16
  25. DDK: Distilling Domain Knowledge for Efficient Large Language Models 18 upvotes, #4 of 2024-07-25
  26. LongIns: A Challenging Long-context Instruction-based Exam for LLMs 18 upvotes, #6 of 2024-06-26
  27. Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level 13 upvotes, #14 of 2024-06-21
  28. McEval: Massively Multilingual Code Evaluation 38 upvotes, #3 of 2024-06-12
  29. MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series 41 upvotes, #1 of 2024-05-30
  30. MuPT: A Generative Symbolic Music Pretrained Transformer 14 upvotes, #5 of 2024-04-10
  31. Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model 8 upvotes, #7 of 2024-04-08
  32. E^2-LLM: Efficient and Extreme Length Extension of Large Language Models 26 upvotes, #4 of 2024-01-17

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.