Serve models through batch jobs, streaming systems, and online APIs while managing latency, throughput, accelerators, caching, and cost.