Train representations from unlabeled data using contrastive learning, masked prediction, and pretext tasks, then transfer them to downstream problems.