Guilherme Penedo
Guilherme Penedo on Hugging Face Daily Papers: 6 papers, 5 in the top 3 of their day, 615 upvotes.
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 55 upvotes, #2 of 2025-06-26
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text 36 upvotes, #7 of 2025-06-06
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model 158 upvotes, #1 of 2025-02-06
- Towards Best Practices for Open Datasets for LLM Training 47 upvotes, #1 of 2025-01-16
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale 72 upvotes, #1 of 2024-06-26
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only 45 upvotes, #1 of 2023-06-05
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.