Wenhu Chen

Wenhu Chen on Hugging Face Daily Papers: 39 papers, 11 in the top 3 of their day, 1,453 upvotes.

  1. Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
  2. Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
  3. VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
  4. StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
  5. Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning 48 upvotes, #4 of 2025-05-23
  6. General-Reasoner: Advancing LLM Reasoning Across All Domains 20 upvotes, #10 of 2025-05-21
  7. VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning 42 upvotes, #4 of 2025-04-15
  8. ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
  9. Towards Trustworthy GUI Agents: A Survey 20 upvotes, #13 of 2025-04-02
  10. MoCha: Towards Movie-Grade Talking Character Synthesis 103 upvotes, #1 of 2025-04-01
  11. Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers 17 upvotes, #8 of 2025-03-17
  12. VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
  13. YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
  14. ABC: Achieving Better Control of Multimodal Embeddings using VLMs 18 upvotes, #6 of 2025-03-06
  15. PixelWorld: Towards Perceiving Everything as Pixels 14 upvotes, #6 of 2025-02-03
  16. Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 50 upvotes, #1 of 2025-01-30
  17. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
  18. VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation 24 upvotes, #5 of 2024-12-03
  19. OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision 42 upvotes, #2 of 2024-11-12
  20. Harnessing Webpage UIs for Text-Rich Visual Understanding 28 upvotes, #7 of 2024-10-18
  21. MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
  22. VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks 4 upvotes, #43 of 2024-10-10
  23. T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design 14 upvotes, #15 of 2024-10-10
  24. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
  25. LongIns: A Challenging Long-context Instruction-based Exam for LLMs 18 upvotes, #6 of 2024-06-26
  26. MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24
  27. LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs 54 upvotes, #1 of 2024-06-24
  28. WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 11 upvotes, #16 of 2024-06-18
  29. GenAI Arena: An Open Evaluation Platform for Generative Models 18 upvotes, #4 of 2024-06-10
  30. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark 35 upvotes, #1 of 2024-06-04
  31. MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series 41 upvotes, #1 of 2024-05-30
  32. T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback 18 upvotes, #2 of 2024-05-30
  33. AnyV2V: A Plug-and-Play Framework For Any Video-to-Video Editing Tasks 17 upvotes, #4 of 2024-03-22
  34. ChatMusician: Understanding and Generating Music Intrinsically with LLM 57 upvotes, #1 of 2024-02-27
  35. StructLM: Towards Building Generalist Models for Structured Knowledge Grounding 27 upvotes, #7 of 2024-02-27
  36. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 84 upvotes, #1 of 2024-02-23
  37. ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation 27 upvotes, #6 of 2024-02-08
  38. CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-01-23
  39. E^2-LLM: Efficient and Extreme Length Extension of Large Language Models 26 upvotes, #4 of 2024-01-17

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.