Wenhu Chen
Wenhu Chen on Hugging Face Daily Papers: 39 papers, 11 in the top 3 of their day, 1,453 upvotes.
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
- VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
- StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning 48 upvotes, #4 of 2025-05-23
- General-Reasoner: Advancing LLM Reasoning Across All Domains 20 upvotes, #10 of 2025-05-21
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning 42 upvotes, #4 of 2025-04-15
- ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
- Towards Trustworthy GUI Agents: A Survey 20 upvotes, #13 of 2025-04-02
- MoCha: Towards Movie-Grade Talking Character Synthesis 103 upvotes, #1 of 2025-04-01
- Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers 17 upvotes, #8 of 2025-03-17
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
- YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
- ABC: Achieving Better Control of Multimodal Embeddings using VLMs 18 upvotes, #6 of 2025-03-06
- PixelWorld: Towards Perceiving Everything as Pixels 14 upvotes, #6 of 2025-02-03
- Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 50 upvotes, #1 of 2025-01-30
- MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
- VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation 24 upvotes, #5 of 2024-12-03
- OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision 42 upvotes, #2 of 2024-11-12
- Harnessing Webpage UIs for Text-Rich Visual Understanding 28 upvotes, #7 of 2024-10-18
- MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks 4 upvotes, #43 of 2024-10-10
- T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design 14 upvotes, #15 of 2024-10-10
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
- LongIns: A Challenging Long-context Instruction-based Exam for LLMs 18 upvotes, #6 of 2024-06-26
- MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24
- LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs 54 upvotes, #1 of 2024-06-24
- WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 11 upvotes, #16 of 2024-06-18
- GenAI Arena: An Open Evaluation Platform for Generative Models 18 upvotes, #4 of 2024-06-10
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark 35 upvotes, #1 of 2024-06-04
- MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series 41 upvotes, #1 of 2024-05-30
- T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback 18 upvotes, #2 of 2024-05-30
- AnyV2V: A Plug-and-Play Framework For Any Video-to-Video Editing Tasks 17 upvotes, #4 of 2024-03-22
- ChatMusician: Understanding and Generating Music Intrinsically with LLM 57 upvotes, #1 of 2024-02-27
- StructLM: Towards Building Generalist Models for Structured Knowledge Grounding 27 upvotes, #7 of 2024-02-27
- OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 84 upvotes, #1 of 2024-02-23
- ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation 27 upvotes, #6 of 2024-02-08
- CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-01-23
- E^2-LLM: Efficient and Extreme Length Extension of Large Language Models 26 upvotes, #4 of 2024-01-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.