HUANG SHAOHAN
HUANG SHAOHAN on Hugging Face Daily Papers: 28 papers, 12 in the top 3 of their day, 2,009 upvotes.
- Thinking Augmented Pre-training 22 upvotes, #9 of 2025-09-26
- VibeVoice Technical Report 118 upvotes, #1 of 2025-08-27
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models 6 upvotes, #15 of 2025-08-14
- Geometric-Mean Policy Optimization 30 upvotes, #7 of 2025-07-29
- Reasoning with Exploration: An Entropy Perspective 26 upvotes, #8 of 2025-06-18
- On-Policy RL with Optimal Reward Baseline 14 upvotes, #25 of 2025-05-30
- Reward Reasoning Model 32 upvotes, #5 of 2025-05-21
- Think Only When You Need with Large Hybrid-Reasoning Models 18 upvotes, #11 of 2025-05-21
- BitNet b1.58 2B4T Technical Report 66 upvotes, #1 of 2025-04-17
- GeAR: Generation Augmented Retrieval 21 upvotes, #8 of 2025-01-09
- Multimodal Latent Language Modeling with Next-Token Diffusion 38 upvotes, #5 of 2024-12-13
- On Domain-Specific Post-Training for Multimodal Large Language Models 24 upvotes, #4 of 2024-12-02
- MH-MoE:Multi-Head Mixture-of-Experts 22 upvotes, #7 of 2024-11-26
- Instruction Pre-Training: Language Models are Supervised Multitask Learners 74 upvotes, #2 of 2024-06-21
- MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning 39 upvotes, #2 of 2024-05-21
- Multi-Head Mixture-of-Experts 45 upvotes, #2 of 2024-04-24
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits 630 upvotes, #1 of 2024-02-28
- Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models 52 upvotes, #2 of 2024-02-21
- Democratizing Reasoning Ability: Tailored Learning from Large Language Model 16 upvotes, #5 of 2023-10-23
- BitNet: Scaling 1-bit Transformers for Large Language Models 108 upvotes, #1 of 2023-10-18
- Calibrating LLM-Based Evaluator 12 upvotes, #5 of 2023-09-26
- Kosmos-2.5: A Multimodal Literate Model 56 upvotes, #4 of 2023-09-21
- Adapting Large Language Models via Reading Comprehension 82 upvotes, #2 of 2023-09-19
- Retentive Network: A Successor to Transformer for Large Language Models 173 upvotes, #1 of 2023-07-18
- LongNet: Scaling Transformers to 1,000,000,000 Tokens 82 upvotes, #2 of 2023-07-06
- Kosmos-2: Grounding Multimodal Large Language Models to the World 36 upvotes, #1 of 2023-06-27
- Dual-Alignment Pre-training for Cross-lingual Sentence Embedding 1 upvotes, #10 of 2023-05-17
- Pre-training Language Model as a Multi-perspective Course Learner 1 upvotes, #6 of 2023-05-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.