Convert text and other inputs into model-ready units using character, word, subword, byte, and multimodal tokenization. Inspect vocabulary tradeoffs, sequence length, multilingual behavior, and tokenization failures.