Pre-training is the expensive part of building an LLM. The model processes trillions of tokens from books, websites, code repositories, and other text sources, learning to predict the next token. This phase requires thousands of GPUs and costs millions of dollars.
Pre-training produces a 'foundation model' that understands language broadly but isn't specialized for any task. Fine-tuning then adapts it for specific uses.