Tokenization is the first step in any LLM pipeline. Common algorithms include Byte Pair Encoding (BPE) and SentencePiece, which learn to split text into subword units based on frequency in the training data.
Good tokenization balances vocabulary size (fewer tokens = faster processing) with representation quality (each token should carry meaningful information). Different models use different tokenizers, which is why token counts vary between providers.