Serve models through APIs, batch jobs, streaming responses, and real-time endpoints. Apply quantization, caching, continuous batching, speculative decoding, autoscaling, routing, and capacity planning.