Reduce memory, latency, and energy use with quantization, pruning, distillation, compilation, caching, and efficient model architectures.