Shiwei Liu
Shiwei Liu on Hugging Face Daily Papers: 9 papers, 0 in the top 3 of their day, 196 upvotes.
- SoS1: O1 and R1-Like Reasoning LLMs are Sum-of-Square Solvers 20 upvotes, #6 of 2025-03-03
- Stable-SPAM: How to Train in 4-Bit More Stably than 16-Bit Adam 16 upvotes, #11 of 2025-02-25
- Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More 9 upvotes, #16 of 2025-02-12
- The Curse of Depth in Large Language Models 27 upvotes, #5 of 2025-02-11
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning 20 upvotes, #7 of 2025-01-23
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training 15 upvotes, #9 of 2025-01-14
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN 18 upvotes, #7 of 2024-12-19
- From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank Gradients 6 upvotes, #13 of 2024-07-17
- Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients 28 upvotes, #4 of 2024-07-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.