LoRA dramatically reduces the cost of fine-tuning by only training low-rank matrices that are added to existing model weights. Instead of updating billions of parameters, LoRA trains millions — a 1000x reduction.
This makes fine-tuning accessible on consumer hardware. LoRA adapters are small (typically 1-100 MB vs. the full model's 10-100 GB), can be swapped at inference time, and can be combined (merged).