Implement self-attention, positional representations, multi-head attention, and encoder or decoder stacks, then train a small transformer.