Model compression and efficient inference: Deep Learning course | Zoonk
47. Model compression and efficient inference
Reduce model size and latency with pruning, quantization, distillation, low-rank methods, efficient attention, and architecture changes. Benchmark accuracy, memory, throughput, energy use, and hardware compatibility.