AI glossary
Pre-training
Pre-training is the initial stage of training a model, usually on a large, general, often unlabeled dataset with a self-supervised objective, before the model is further trained (fine-tuned) on a smaller, more specific dataset for a particular downstream task. This process establishes the foundational knowledge base that allows subsequent adaptation to be efficient and accurate.
The pretrain-then-fine-tune pattern
The core idea behind this pattern is to let a model first learn broad, general patterns in language (or another data type) at a massive scale. Once these general patterns are established, fine-tuning on a much smaller task-specific dataset afterward is enough to adapt that general knowledge to a specific job.
This approach contrasts sharply with learning everything from scratch on limited task-specific data. When you start from scratch, the model must simultaneously learn basic grammar, syntax, and world knowledge while trying to solve your specific problem. With pre-training, the model already understands how language works. It only needs to learn how to apply that understanding to your narrow domain. This is why the fine-tuning process is significantly faster and requires far less labeled data than building a model for a single task from the ground up.
Self-supervised pre-training objectives
Self-supervised pre-training objectives allow a model to learn from raw, unlabeled text by generating its own training signal. This eliminates the need for humans to label each example manually, which is often the biggest bottleneck in traditional machine learning.
For instance, the model might be tasked with predicting the next word in a sequence. Alternatively, it might need to predict a word that has been deliberately hidden (“masked”) from elsewhere in a sentence. In both cases, the “label” is derived directly from the input text itself. This falls under the umbrella of self-supervised learning, where the data structure provides the supervision rather than human annotators.
BERT and GPT as examples
Two prominent examples illustrate how different pre-training objectives shape model behavior.
Google researchers introduced BERT (“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Devlin et al., 2018) as a model pre-trained on masked-language-modeling and next-sentence-prediction objectives over large text corpora. It was then fine-tuned for specific tasks such as question answering. BERT’s bidirectional nature allows it to understand context from both directions in a sentence, making it particularly strong for tasks requiring deep semantic understanding.
In contrast, OpenAI’s GPT model family is pre-trained on the objective of predicting the next token in a sequence of text. This same next-token-prediction pre-training objective is what most modern large language model (LLM) architectures are built on. GPT models process text in a single direction, which makes them highly effective for generative tasks where predicting the subsequent word is the primary goal. Both approaches rely on the Transformer architecture to handle long-range dependencies in text efficiently.
Why pre-training works
Pre-training works because the objective (like next-token prediction) can be generated automatically from raw text without manual labeling. This allows the process to be run at a scale of data and compute far larger than would be feasible for a manually labeled dataset.
Because the model can ingest billions of words from diverse sources, it picks up broad language patterns, factual associations, and structural nuances before it ever sees a task-specific example. This broad foundation means that when you later introduce specific data, the model doesn’t need to learn basic concepts. It only needs to adjust its existing weights to prioritize the information relevant to your specific use case. This is the fundamental advantage of the BERT and GPT paradigms: scale in pre-training translates to capability in downstream tasks.
Cost and scale
Pre-training a large language model from scratch requires substantially more data and compute than fine-tuning an already pre-trained model. The computational cost includes not just the training runs, but also the infrastructure to store and process petabytes of text data.
This is why most teams building on LLMs fine-tune or otherwise adapt an existing pre-trained model rather than pre-training their own from the ground up. Unless you have access to massive computational resources and unique, high-quality data that isn’t already well-represented in public corpora, starting from a pre-trained foundation is the more efficient path. Fine-tuning builds on the general knowledge already embedded in the model, requiring only a fraction of the compute power needed for initial pre-training.
FAQ
What is the difference between pre-training and fine-tuning?
Pre-training teaches the model general patterns using vast amounts of unlabeled data. Fine-tuning adapts that pre-trained model to a specific task using a smaller, labeled dataset.
Why do we need self-supervised learning for pre-training?
Self-supervised learning allows models to learn from raw, unlabeled text by creating their own labels (like predicting masked words). This makes it possible to train on massive datasets that would be too expensive or time-consuming to label manually.
Can I pre-train a model myself?
Yes, but it requires significant compute resources and large datasets. Most developers choose to fine-tune an existing pre-trained model instead.