Abstract
The vast diversity of styles, domains, and quality levels present in language model pre-training corpora is essential in developing general model capabilities, but efficiently learning and deploying the correct behaviors exemplified in each of these heterogeneous data sources is challenging. To address this, we propose a new method, termed Metadata Conditioning then Cooldown (MeCo), to incorporate additional learning cues during pre-training. MeCo first provides metadata (e.g., URLs like en.wikipedia.org) alongside the text during training, then transitions to a cooldown phase using only standard text—enabling the model to perform well even without metadata. MeCo significantly accelerates pre-training across different model scales (600M to 8B parameters) and training corpora (C4, RefinedWeb, and DCLM). Notably, a 1.6B language model trained with MeCo matches the downstream task performance of standard pretraining while using 33% less data. Additionally, MeCo allows us to steer language models by conditioning the inference prompt on either real or fabricated metadata that encodes the desired output properties—for example, prepending wikipedia.org to reduce harmful generations or factquizmaster.com (fabricated) to improve performance on common knowledge tasks. We further demonstrate that MeCo is compatible with various types of metadata, such as model-generated topics. MeCo is remarkably simple, adds no computational overhead, and shows promise for producing more capable and steerable language models. Our models, data, and code are available at https://github.com/princeton-pli/MeCo.
| Original language | English (US) |
|---|---|
| Pages (from-to) | 18612-18629 |
| Number of pages | 18 |
| Journal | Proceedings of Machine Learning Research |
| Volume | 267 |
| State | Published - 2025 |
| Event | 42nd International Conference on Machine Learning, ICML 2025 - Vancouver, Canada Duration: Jul 13 2025 → Jul 19 2025 |
All Science Journal Classification (ASJC) codes
- Software
- Control and Systems Engineering
- Statistics and Probability
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Metadata Conditioning Accelerates Language Model Pre-training'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver