Jiaheng Liu
Jiaheng Liu on Hugging Face Daily Papers: 32 papers, 11 in the top 3 of their day, 1,082 upvotes.
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? 54 upvotes, #4 of 2025-09-05
- Flow-GRPO: Training Flow Matching Models via Online RL 68 upvotes, #2 of 2025-05-09
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs 22 upvotes, #8 of 2025-04-23
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model 30 upvotes, #7 of 2025-04-15
- COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values 41 upvotes, #5 of 2025-04-09
- YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
- CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models 23 upvotes, #8 of 2025-02-25
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
- Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
- ProgCo: Program Helps Self-Correction of Large Language Models 24 upvotes, #7 of 2025-01-03
- PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos 6 upvotes, #21 of 2024-12-03
- Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models 33 upvotes, #4 of 2024-11-12
- M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation 6 upvotes, #16 of 2024-11-04
- AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions 36 upvotes, #2 of 2024-10-30
- Can MLLMs Understand the Deep Implication Behind Chinese Images? 7 upvotes, #23 of 2024-10-18
- PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment 18 upvotes, #14 of 2024-10-18
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models 17 upvotes, #7 of 2024-10-16
- ING-VP: MLLMs cannot Play Easy Vision-based Games Yet 8 upvotes, #26 of 2024-10-10
- MIO: A Foundation Model on Multimodal Tokens 46 upvotes, #2 of 2024-09-30
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models 39 upvotes, #1 of 2024-09-25
- OmniBench: Towards The Future of Universal Omni-Language Models 24 upvotes, #4 of 2024-09-25
- FuzzCoder: Byte-level Fuzzing Test via Large Language Model 44 upvotes, #3 of 2024-09-06
- TableBench: A Comprehensive and Complex Benchmark for Table Question Answering 48 upvotes, #1 of 2024-08-21
- I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm 30 upvotes, #2 of 2024-08-16
- DDK: Distilling Domain Knowledge for Efficient Large Language Models 18 upvotes, #4 of 2024-07-25
- LongIns: A Challenging Long-context Instruction-based Exam for LLMs 18 upvotes, #6 of 2024-06-26
- Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level 13 upvotes, #14 of 2024-06-21
- McEval: Massively Multilingual Code Evaluation 38 upvotes, #3 of 2024-06-12
- MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series 41 upvotes, #1 of 2024-05-30
- MuPT: A Generative Symbolic Music Pretrained Transformer 14 upvotes, #5 of 2024-04-10
- Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model 8 upvotes, #7 of 2024-04-08
- E^2-LLM: Efficient and Extreme Length Extension of Large Language Models 26 upvotes, #4 of 2024-01-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.