Use GPUs and other accelerators efficiently through batching, mixed precision, memory management, checkpointing, and optimized kernels. Scale training with data, tensor, pipeline, and expert parallelism while diagnosing communication bottlenecks.