HUANG SHAOHAN

HUANG SHAOHAN on Hugging Face Daily Papers: 28 papers, 12 in the top 3 of their day, 2,009 upvotes.

  1. Thinking Augmented Pre-training 22 upvotes, #9 of 2025-09-26
  2. VibeVoice Technical Report 118 upvotes, #1 of 2025-08-27
  3. VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models 6 upvotes, #15 of 2025-08-14
  4. Geometric-Mean Policy Optimization 30 upvotes, #7 of 2025-07-29
  5. Reasoning with Exploration: An Entropy Perspective 26 upvotes, #8 of 2025-06-18
  6. On-Policy RL with Optimal Reward Baseline 14 upvotes, #25 of 2025-05-30
  7. Reward Reasoning Model 32 upvotes, #5 of 2025-05-21
  8. Think Only When You Need with Large Hybrid-Reasoning Models 18 upvotes, #11 of 2025-05-21
  9. BitNet b1.58 2B4T Technical Report 66 upvotes, #1 of 2025-04-17
  10. GeAR: Generation Augmented Retrieval 21 upvotes, #8 of 2025-01-09
  11. Multimodal Latent Language Modeling with Next-Token Diffusion 38 upvotes, #5 of 2024-12-13
  12. On Domain-Specific Post-Training for Multimodal Large Language Models 24 upvotes, #4 of 2024-12-02
  13. MH-MoE:Multi-Head Mixture-of-Experts 22 upvotes, #7 of 2024-11-26
  14. Instruction Pre-Training: Language Models are Supervised Multitask Learners 74 upvotes, #2 of 2024-06-21
  15. MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning 39 upvotes, #2 of 2024-05-21
  16. Multi-Head Mixture-of-Experts 45 upvotes, #2 of 2024-04-24
  17. The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits 630 upvotes, #1 of 2024-02-28
  18. Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models 52 upvotes, #2 of 2024-02-21
  19. Democratizing Reasoning Ability: Tailored Learning from Large Language Model 16 upvotes, #5 of 2023-10-23
  20. BitNet: Scaling 1-bit Transformers for Large Language Models 108 upvotes, #1 of 2023-10-18
  21. Calibrating LLM-Based Evaluator 12 upvotes, #5 of 2023-09-26
  22. Kosmos-2.5: A Multimodal Literate Model 56 upvotes, #4 of 2023-09-21
  23. Adapting Large Language Models via Reading Comprehension 82 upvotes, #2 of 2023-09-19
  24. Retentive Network: A Successor to Transformer for Large Language Models 173 upvotes, #1 of 2023-07-18
  25. LongNet: Scaling Transformers to 1,000,000,000 Tokens 82 upvotes, #2 of 2023-07-06
  26. Kosmos-2: Grounding Multimodal Large Language Models to the World 36 upvotes, #1 of 2023-06-27
  27. Dual-Alignment Pre-training for Cross-lingual Sentence Embedding 1 upvotes, #10 of 2023-05-17
  28. Pre-training Language Model as a Multi-perspective Course Learner 1 upvotes, #6 of 2023-05-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.