Architecture

Transformer

A neural network architecture based on self-attention, introduced in the 2017 paper 'Attention Is All You Need.'

The transformer replaced recurrent neural networks (RNNs) as the dominant architecture for language modeling. Its key innovation is the self-attention mechanism, which allows the model to weigh the relevance of every token in the input when producing each output token.

Transformers process all tokens in parallel (unlike RNNs which process sequentially), making them much faster to train on modern GPUs. Every major LLM — GPT, Claude, Gemini, LLaMA — is built on transformer architecture.

← All terms