Luca Soldaini
Luca Soldaini on Hugging Face Daily Papers: 19 papers, 6 in the top 3 of their day, 809 upvotes.
- Bolmo: Byteifying the Next Generation of Language Models 11 upvotes, #14 of 2025-12-22
- Olmo 3 22 upvotes, #11 of 2025-12-17
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research 53 upvotes, #4 of 2025-11-25
- olmOCR 2: Unit Test Rewards for Document OCR 10 upvotes, #13 of 2025-10-23
- FlexOlmo: Open Language Models for Flexible Data Use 5 upvotes, #15 of 2025-07-10
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text 36 upvotes, #7 of 2025-06-06
- Teaching Models to Understand (but not Generate) High-risk Data 4 upvotes, #16 of 2025-05-07
- DataDecide: How to Predict Best Pretraining Data with Small Experiments 15 upvotes, #10 of 2025-04-16
- OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens 69 upvotes, #1 of 2025-04-10
- Establishing Task Scaling Laws via Compute-Efficient Model Ladders 2 upvotes, #30 of 2024-12-06
- TÜLU 3: Pushing Frontiers in Open Language Model Post-Training 55 upvotes, #1 of 2024-11-25
- OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs 25 upvotes, #5 of 2024-11-22
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models 89 upvotes, #1 of 2024-09-26
- OLMoE: Open Mixture-of-Experts Language Models 67 upvotes, #2 of 2024-09-04
- FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions 8 upvotes, #9 of 2024-03-25
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research 66 upvotes, #2 of 2024-02-02
- OLMo: Accelerating the Science of Language Models 86 upvotes, #1 of 2024-02-02
- Paloma: A Benchmark for Evaluating Language Model Fit 12 upvotes, #8 of 2023-12-19
- What's In My Big Data? 11 upvotes, #8 of 2023-11-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.