Colin Raffel
Colin Raffel on Hugging Face Daily Papers: 9 papers, 4 in the top 3 of their day, 606 upvotes.
- Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models 9 upvotes, #22 of 2026-06-12
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior 16 upvotes, #9 of 2025-12-25
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 55 upvotes, #2 of 2025-06-26
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text 36 upvotes, #7 of 2025-06-06
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model 158 upvotes, #1 of 2025-02-06
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale 72 upvotes, #1 of 2024-06-26
- DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows 31 upvotes, #4 of 2024-02-19
- Distributed Inference and Fine-tuning of Large Language Models Over The Internet 27 upvotes, #2 of 2023-12-14
- Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model 11 upvotes, #6 of 2023-10-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.