Deploy

This chapter covers creating and scaling inference services, model management and storage, Model as a Service (MaaS), quota and metering at the inference gateway, and model compression.

The components this chapter relies on are listed under Components below.

Model Management

Inference Service

Inference Gateway

LLM Compressor

Model as a Service

Components