Guilherme Penedo

Guilherme Penedo on Hugging Face Daily Papers: 6 papers, 5 in the top 3 of their day, 615 upvotes.

  1. FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 55 upvotes, #2 of 2025-06-26
  2. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text 36 upvotes, #7 of 2025-06-06
  3. SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model 158 upvotes, #1 of 2025-02-06
  4. Towards Best Practices for Open Datasets for LLM Training 47 upvotes, #1 of 2025-01-16
  5. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale 72 upvotes, #1 of 2024-06-26
  6. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only 45 upvotes, #1 of 2023-06-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.