Attention lets models dynamically focus on the most relevant context. In self-attention, each token computes a weighted sum of all other tokens, where the weights represent relevance.
Multi-head attention runs multiple attention computations in parallel, allowing the model to attend to different types of relationships simultaneously — syntactic structure in one head, semantic meaning in another. This is the core building block of transformers.