Implement attention, positional information, encoder and decoder architectures, pretraining objectives, and efficient transformer inference.