Apply patch embeddings, class tokens, hierarchical attention, and image-specific augmentation to transformer-based vision models. Compare vision transformers with convolutional networks under different data and compute limits.