Skip to content
Back to student guides
KServeMLOpsModel serving3 levels98 sectionsCovers KServe 0.21

The Complete KServe Guide

Serve models on Kubernetes with autoscaling and canaries using KServe. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
16sections
41examples

This is part one of three. It covers everything a newcomer needs to serve a machine learning model on Kubernetes with KServe, not a teaser. By the end you can install KServe on a local cluster, deploy a trained model as a web endpoint with one YAML file, send it predictions, read its status, find out why it failed when it fails, and tidy up after yourself. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and these ideas only stick once you have watched your own model go from "not ready" to "ready" and answered a real request.

One honest warning before we start. KServe is a young project that has not reached version 1.0, and its minor releases can change things. The newest release when this guide was checked was v0.21.0 (published on 25 September 2026), while the documentation site was still showing 0.20 install commands for a few days afterwards. Wherever a command contains a version number, this guide uses a shell variable so you can put in the version you are installing. When the docs and this guide disagree, the docs win, and the Sources section at the end tells you where to look.

What KServe is, and the problem it solves

KServe is an open-source tool that turns a trained model into a running, network-reachable service on a Kubernetes cluster, and keeps that service healthy, scaled and updatable. You describe the service in a short YAML file, apply it with kubectl, and KServe creates everything underneath: the containers, the networking, the autoscaling rules and the plumbing that downloads your model file.

To see why that is useful, start with what happens without it. You have trained a scikit-learn model and saved it as a file. Someone now needs to call it from a web application. The simplest approach is to write a small web server (Flask or FastAPI), load the file at start-up, expose a /predict route, wrap it in a Docker image, and write a Kubernetes Deployment and Service to run it. That works, once. Then the second model arrives, built with XGBoost, and you write another server. The third is a PyTorch model that needs a GPU. Each server has slightly different request formats, different health checks, different logging, different ways of loading its file, and a different person who understands it. Rolling out a new model version safely (send ten percent of traffic to it, watch, then promote) means more custom work. Scaling down to nothing overnight to save money means even more.

Teams that serve many models end up writing the same glue again and again. KServe is that glue, written once, shared by many companies, and governed as a project of the Cloud Native Computing Foundation. It offers four things that matter to a beginner:

📦

One resource per model

You write a single InferenceService YAML. KServe turns it into pods, services and routes, so you stop hand-writing Deployments.

🔌

Ready-made model servers

Built-in runtimes know how to serve scikit-learn, XGBoost, Hugging Face, TensorFlow and other model formats, so you often write no serving code at all.

📈

Scaling and rollouts

Autoscaling, and in some modes scale-to-zero and canary traffic splitting, come from configuration, not from code you maintain.

🧭

A standard API

Every model gets the same style of endpoint and the same status fields, whichever framework built it.

A short history helps you read older tutorials. The project began inside the Kubeflow community under the name KFServing. It was renamed KServe around version 0.7, and its API group became serving.kserve.io. If you find a blog post that uses kfserving.io or an apiVersion you do not recognise, it is probably years old. The stable API for the main resource today is serving.kserve.io/v1beta1. KServe also grew a newer, separate resource for large language models, LLMInferenceService, which is marked alpha; this beginner guide focuses on the stable InferenceService and mentions the newer one only so the name is not a surprise.

KServe is not a training tool, a model registry or a notebook. It starts after your model exists as a file somewhere (a bucket, a persistent volume, a Hugging Face repository) and ends at an HTTP or gRPC endpoint. If you want to package a model into your own image and deploy it by other means, look at BentoML. If you need the cluster fundamentals first, the Kubernetes guide in this series covers them, and you will find it much easier to follow this guide with that background. If you want to track the experiments that produced the model, see MLflow.

What you need to follow along: a laptop that can run containers, a terminal, around 8 GB of free memory, and patience for a few downloads. You do not need a GPU, and you do not need a cloud account.

Try it
  1. Think of a model you or your team has built. Write down where its file lives and who calls it.
  2. List what would be needed to make it answer HTTP requests reliably: a server, a container, a Deployment, a Service, a way to scale, a way to update it.
  3. Count how many of those items are about your model's logic and how many are plumbing.
most of the list is plumbing. KServe's job is to own the plumbing so you only supply the model file.

The mental model: five nouns

KServe has a lot of vocabulary, but a beginner needs only five words. Learn these and the rest of the documentation becomes readable.

MODEL FILEin a bucket or volume
→
INFERENCESERVICEyour YAML
→
SERVING RUNTIMEthe model server
→
PODrunning endpoint

InferenceService is the main resource and the one you will write. It is a Kubernetes custom resource, which means KServe taught the cluster a new kind of object. Just as Kubernetes understands Deployment and Service out of the box, after installing KServe it also understands InferenceService. One InferenceService describes one model endpoint: which model format it is, where the file lives, and how much CPU and memory it may use. Its short name is isvc, so kubectl get isvc works.

Predictor is the part of the InferenceService that runs the model. Every InferenceService has one. In your YAML it appears as spec.predictor. The same resource can optionally describe a transformer, which cleans or reshapes inputs before they reach the model and outputs after it, and an explainer, which produces explanations of predictions. As a beginner you will use a predictor alone, and that is a perfectly complete setup.

ServingRuntime (and its cluster-wide sibling ClusterServingRuntime) is a template that describes a model server: which container image to run, and which model formats it understands. When you say "this is a scikit-learn model", KServe looks through the available runtimes for one that supports sklearn, and builds the pod from that runtime's template. This is the reason you rarely write a Dockerfile: a ClusterServingRuntime named for scikit-learn already exists after installation. A namespaced ServingRuntime exists only inside one namespace; a ClusterServingRuntime is visible everywhere.

Storage initializer is a small helper that KServe adds to your pod as an init container, which is a container that runs to completion before the main one starts. Its single job is to download your model from the storageUri you gave and place it in a directory called /mnt/models, where the model server expects to find it. When a deployment fails because the model could not be fetched, this is the container whose logs you read.

Deployment mode is how KServe runs your model underneath. The current docs describe three. Standard mode (called RawDeployment in older documentation) uses plain Kubernetes pieces: a Deployment, a Service and a HorizontalPodAutoscaler. Knative (serverless) mode builds on the Knative project and gives you scale-to-zero and revision-based canary rollouts, at the cost of installing and operating Knative as well. ModelMesh is a separate option for packing very many small models onto shared pods. For a first project on a laptop, Standard is the simplest because it needs the fewest extra components, and it is the mode this guide installs.

Put those together and the flow in the diagram reads naturally. You write an InferenceService. The KServe controller notices it, finds a matching ServingRuntime, and creates a pod. The pod's init container (the storage initializer) downloads the model file. The main container (the model server, named kserve-container) loads it and starts listening. Kubernetes networking, and optionally a gateway, expose the endpoint.

Two more words will appear in error messages and docs. Protocol means the shape of the requests the endpoint accepts. KServe supports the v1 REST protocol (the URL looks like /v1/models/<name>:predict) and the v2 protocol, also called the Open Inference Protocol, which is shared with other serving tools and also supports gRPC (a binary request format, faster than JSON). Your first example uses v1. Gateway means the entry point that receives outside traffic and forwards it to the right model. KServe works with the Kubernetes Gateway API, Ingress controllers and Istio; the current docs recommend Gateway API.

Try it
  1. Without looking back, draw the chain from a model file to a running endpoint using the five nouns.
  2. Mark which of the five you write yourself (one) and which KServe supplies or creates (the rest).
you write only the InferenceService. The runtime ships with KServe, the storage initializer is injected, and the mode is a setting.

Before you install: what must already exist

KServe is not a standalone program you download and run. It is an add-on to Kubernetes, so the first thing to understand is what has to be in place underneath it. Skipping this section is the most common reason first installs fail.

According to the official Kubernetes deployment guide, the requirements are:

  • A Kubernetes cluster, version 1.32 or newer. KServe's controller runs inside it, and so do your models.
  • cert-manager, version 1.15.0 or newer. Kubernetes lets the API server call out to extra programs (called webhooks) when objects are created, and those calls must be encrypted. KServe's controller has such a webhook, which checks and fills in defaults for your InferenceService. cert-manager is the tool that creates and renews the certificates for it. If cert-manager is missing, KServe's webhook has no certificate and your first InferenceService is rejected.
  • A network layer. Either the Kubernetes Gateway API (recommended by the current docs), or an Ingress controller. Istio, if you use it, must be 1.22 or newer.
  • kubectl with cluster-admin permissions, because installing KServe creates cluster-wide objects.
  • Helm 3 if you choose the Helm install route. Helm is a package manager for Kubernetes; a chart is a package.
  • Knative Serving, only if you want serverless mode. You do not need it for this guide.

There is no native KServe server for Windows or macOS. KServe runs on Kubernetes, and Kubernetes nodes run Linux. The operating system of your laptop affects only the client tools you install, and how you provide a local cluster:

  • Linux: install kubectl, helm and a local-cluster tool such as kind, minikube or k3d. Docker or Podman is needed for kind.
  • macOS: brew install kubectl helm kind installs the tools. Docker Desktop, Colima or OrbStack provides the container runtime. If you have an Apple Silicon Mac, be aware that some model-server images may be built for Intel (amd64) only; if a pod fails to start with an architecture message, that is the reason.
  • Windows: use WSL2 with an Ubuntu distribution, turn on Docker Desktop's WSL integration, and then follow the Linux steps inside Ubuntu. Native PowerShell works for kubectl and helm against a remote cluster, but the shell snippets in this guide use Bash syntax ($(...), heredocs) and need translating.

The recommended path for a student is a local cluster made by kind (Kubernetes in Docker). It creates a one-node cluster inside a container in about a minute and can be thrown away with one command.

BASH
kind create cluster --name kserve-demo
kubectl cluster-info --context kind-kserve-demo
kubectl get nodes

The first command builds the cluster, and the second confirms that kubectl is pointing at it. The third should print one node with STATUS of Ready. If it prints NotReady, wait a minute and run it again.

Check the version of Kubernetes that kind installed, because KServe asks for 1.32 or newer.

BASH
kubectl version

Look at the Server Version line. If it is older than 1.32, upgrade kind (brew upgrade kind on macOS) and recreate the cluster; do not try to continue.

A laptop cluster is for learning, not for serving real users A kind cluster lives and dies with your machine, has one node, and is not reachable from outside. It teaches you every concept in this guide, but production serving belongs on a managed or properly operated cluster. The Mid-level and Senior parts of this series cover that.
Try it
  1. Install kubectl, helm and kind for your operating system.
  2. Create the kserve-demo cluster and run kubectl get nodes.
  3. Run kubectl get pods -A and read the names of the system pods.
one node with STATUS Ready, and a handful of pods in kube-system in the Running state. That is your empty stage.

Installing KServe

Installation happens in three stages: certificates, networking, then KServe itself. Do them in this order, because each depends on the one before.

Stage 1: cert-manager. The cert-manager project publishes a single manifest per release. The official KServe docs require at least 1.15.0, so install that or a newer release from the cert-manager releases page.

BASH
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.15.0/cert-manager.yaml
kubectl wait --for=condition=Available deployment --all -n cert-manager --timeout=180s

The second command waits until all three cert-manager deployments report Available. Without that wait, the next stage can race ahead of the certificates it needs and give confusing webhook errors.

Stage 2: Gateway API definitions. The Gateway API is a set of Kubernetes object types (Gateway, HTTPRoute and friends) for describing how outside traffic reaches services. The KServe docs use release 1.2.1 of the Gateway API definitions.

BASH
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.2.1/standard-install.yaml

This installs only the definitions of the objects, not a program that acts on them. Actual traffic handling needs a Gateway implementation (a controller such as Envoy Gateway or Istio). For this guide you will test through kubectl port-forward, which does not need one, so you can stop here. In a real cluster, choosing and installing a gateway implementation is a design decision covered in the later parts of the series.

Stage 3: KServe. The official guide offers a Helm route and a plain YAML route. We will show Helm, because it makes upgrades and removal simpler. Helm installs in two steps: first the custom resource definitions, which are the new object types such as InferenceService, then the controller and default runtimes.

Set the version once so it is easy to change. The documentation was pinning 0.20.0 when it was checked, and the latest release was 0.21.0. Use whichever you have confirmed exists on the releases page.

BASH
export KSERVE_VERSION=v0.21.0

helm install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd --version ${KSERVE_VERSION}

helm install kserve oci://ghcr.io/kserve/charts/kserve-resources --version ${KSERVE_VERSION} \
  --set kserve.controller.deploymentMode=Standard \
  --set kserve.controller.gateway.ingressGateway.enableGatewayApi=true \
  --set kserve.controller.gateway.ingressGateway.kserveGateway=kserve/kserve-ingress-gateway

Read the command in pieces. oci://ghcr.io/kserve/charts/... means the charts are stored in a container registry rather than on a website, and Helm 3 can install straight from it. The first chart (kserve-crd) must go first, because the second one contains objects that use the new types. deploymentMode=Standard chooses plain Kubernetes Deployments. The two gateway flags tell the controller to use the Gateway API and name the gateway it should attach routes to.

If helm install kserve-crd complains that the version cannot be found, the release you chose may not have published its charts yet. Fall back to the version shown in the current docs.

There is also a plain YAML route, which applies two manifests from the GitHub release. Notice the --server-side flag.

BASH
kubectl apply --server-side -f https://github.com/kserve/kserve/releases/download/${KSERVE_VERSION}/kserve.yaml
kubectl apply --server-side -f https://github.com/kserve/kserve/releases/download/${KSERVE_VERSION}/kserve-cluster-resources.yaml

The custom resource definitions are very large. A normal kubectl apply stores a copy of each object in an annotation, and that annotation can exceed Kubernetes' size limit. Server-side apply moves the bookkeeping to the API server and avoids the problem. If you ever see an error about annotations being too long, that is the cue to add --server-side.

Use one install route or the other, never both. Mixing them leaves objects owned by two installers and makes later upgrades painful.

Try it
  1. Install cert-manager and wait for its deployments to become available.
  2. Apply the Gateway API definitions.
  3. Install the two KServe Helm charts with your chosen version.
Helm prints STATUS: deployed for both releases. If it does not, read the error and the callout above before retrying.

Checking that the install worked

Never deploy a model onto an install you have not checked. Two commands tell you almost everything.

BASH
kubectl get pods -n kserve

You should see a pod whose name starts with kserve-controller-manager, and its STATUS should be Running with all its containers ready (for example 2/2). If it shows ContainerCreating, wait. If it shows CrashLoopBackOff or stays Pending, run kubectl describe pod on it and read the Events at the bottom; they usually name the cause, such as a missing certificate or too little memory.

BASH
kubectl get clusterservingruntimes

This lists the built-in model servers that came with the install, one row per model format family. You should see runtimes for common formats such as scikit-learn, XGBoost and Hugging Face. The exact list depends on the version, so read it rather than trusting memory. This command is how you will later answer the question "does KServe know how to serve my model format?".

Next, confirm that the new resource type exists, which proves the custom resource definitions are installed.

BASH
kubectl get crd | grep serving.kserve.io
kubectl api-resources --api-group=serving.kserve.io

The first command filters the cluster's list of custom resource definitions for the KServe group. The second shows each resource's short name and whether it is namespaced; you should see inferenceservices with the short name isvc.

If the controller pod is healthy but a later deployment fails with a message about a webhook, the usual suspect is cert-manager. Check it directly.

BASH
kubectl get pods -n cert-manager
kubectl get certificate -n kserve

The first command should show three pods running. The second lists the certificates cert-manager issued for the KServe webhook, and each should show READY as True.

Try it
  1. Run all four checks above and write down what each one proves.
  2. Run kubectl describe pod on the controller pod and find the Containers section. List the container names.
a Running controller, a list of cluster serving runtimes, and the serving.kserve.io resources. Each check isolates a different layer: the controller, the runtimes, the API types, and the certificates.

Your first InferenceService

Now the payoff. You will deploy the example model from the official getting-started guide: a scikit-learn classifier trained on the famous iris flower data set. Given four measurements of a flower (sepal length, sepal width, petal length, petal width, all in centimetres), it predicts one of three species, numbered 0, 1 and 2. The model file is already published in a public Google Cloud Storage bucket, so you do not need to train or upload anything.

First create a namespace. A namespace is a named folder inside the cluster that groups related objects.

BASH
kubectl create namespace kserve-test
Never deploy models into KServe's own namespaces The official docs warn against creating InferenceServices in control-plane namespaces such as kserve. Those namespaces are excluded from the injection that adds the storage initializer, so the model is never downloaded and the pod fails with No such file or directory: '/mnt/models'. Always use a namespace of your own, like kserve-test.

Now write the InferenceService. Save it as a file so you can edit and re-apply it, which is better practice than pasting a heredoc into the terminal.

sklearn-iris.yaml
apiVersion: "serving.kserve.io/v1beta1"
kind: "InferenceService"
metadata:
  name: "sklearn-iris"
spec:
  predictor:
    model:
      modelFormat:
        name: sklearn
      storageUri: "gs://kfserving-examples/models/sklearn/1.0/model"
      resources:
        requests:
          cpu: "100m"
          memory: "512Mi"
        limits:
          cpu: "1"
          memory: "1Gi"

Read it line by line, because every beginner file you write will look like this one.

apiVersion and kind tell Kubernetes which type of object this is. serving.kserve.io/v1beta1 is the stable API group and version; InferenceService is the type. metadata.name is the name you will use in every command, and it also becomes part of the endpoint's address.

Under spec, the predictor is the component that runs the model, and model is the form of predictor that says "use a built-in runtime for this model format". The modelFormat.name of sklearn is how KServe picks a runtime: it searches the available ServingRuntimes for one that lists sklearn as a supported format. The storageUri says where the model lives; the gs:// prefix means Google Cloud Storage, and the storage initializer knows how to read it. Finally resources states how much CPU and memory the container asks for (requests, used by the scheduler to decide where to place it) and how much it may use at most (limits). 100m means one tenth of a CPU core, and 512Mi means 512 mebibytes of memory.

Apply the file and watch it come alive.

BASH
kubectl apply -n kserve-test -f sklearn-iris.yaml
kubectl get inferenceservices sklearn-iris -n kserve-test

The first output usually shows READY as False or empty, because the pod has not started yet. Watch it change with the -w flag, then press Ctrl+C when it turns True.

BASH
kubectl get isvc sklearn-iris -n kserve-test -w

While you wait, look at what KServe created for you. This is the plumbing you did not have to write.

BASH
kubectl get pods -n kserve-test
kubectl get deployments,services -n kserve-test

You will see a pod whose name starts with sklearn-iris-predictor. In Standard mode there is also a Deployment and a Service for it. The pod goes through states you can learn to read: Init:0/1 means the storage initializer is downloading the model, PodInitializing means it finished, and Running with 1/1 ready means the model server loaded the model and passed its readiness check. The first time can take a few minutes because the model-server image must be pulled onto the node.

When READY is True, your model is live. The URL column shows the address KServe assigned. On a laptop cluster without a gateway implementation, that address is not reachable from your machine yet, which is why the next section uses port-forwarding.

The YAML is a desired state, not a command Applying the file does not run a program once. It records what you want, and the KServe controller keeps working to make reality match. If a pod is deleted, the controller recreates it. If you edit the file and apply again, KServe changes the running model to match. This "declare what you want and let the controller reconcile" idea is the heart of Kubernetes, and every KServe feature builds on it.
Try it
  1. Create the namespace and apply sklearn-iris.yaml.
  2. Run kubectl get isvc -n kserve-test -w and note how long it takes to reach READY True.
  3. In another terminal, run kubectl get pods -n kserve-test -w and watch the init stage.
the pod moves from Init to Running, and the InferenceService turns Ready. You have just deployed a model without writing a server.

Sending your first prediction

A model that is ready is worth nothing until you have called it. This section sends a real request and teaches you to read the answer.

The request body for the v1 protocol is a JSON object with an instances list. Each instance is one row of input, here one flower. Save two flowers in a file called iris-input.json.

iris-input.json
{"instances": [[6.8, 2.8, 4.8, 1.4], [6.0, 3.4, 4.5, 1.6]]}

The simplest way to reach the model from a laptop is to open a tunnel from a local port to the service inside the cluster. First find the service name.

BASH
kubectl get svc -n kserve-test

In Standard mode the model's service is normally named after the InferenceService with a -predictor suffix, for example sklearn-iris-predictor. Read your own output to confirm the name. Then open the tunnel, in a terminal you leave running.

BASH
kubectl port-forward -n kserve-test svc/sklearn-iris-predictor 8080:80

8080:80 means "my local port 8080 forwards to port 80 of the service". If your service lists a different port, use that one on the right-hand side. While the tunnel runs, call the model from a second terminal. The URL has a fixed shape: /v1/models/<InferenceService name>:predict.

BASH
curl -s -H "Content-Type: application/json" \
  http://localhost:8080/v1/models/sklearn-iris:predict \
  -d @./iris-input.json

You should get this back.

JSON
{"predictions": [1, 1]}

Read it. The list has one prediction per input row, in the same order. Both flowers were classified as class 1, which in the iris data set is the species Iris versicolor. The model returns only the class number; turning numbers into names is your application's job.

You can also ask the server which models it hosts and whether it is healthy. Those are cheap checks worth knowing.

BASH
curl -s http://localhost:8080/v1/models/sklearn-iris
curl -s http://localhost:8080/v1/models

A healthy model replies with a small JSON document stating that it is ready.

Calling it through a gateway. On a real cluster with a gateway, you do not port-forward to the pod. You send requests to the gateway's address and tell it which model you want through the HTTP Host header. KServe gives each InferenceService a hostname, stored in its status, and the gateway routes by that hostname. The official example, which uses an Istio ingress gateway, looks like this.

BASH
export INGRESS_HOST=$(kubectl -n istio-system get service istio-ingressgateway -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
export INGRESS_PORT=$(kubectl -n istio-system get service istio-ingressgateway -o jsonpath='{.spec.ports[?(@.name=="http2")].port}')
SERVICE_HOSTNAME=$(kubectl get inferenceservice sklearn-iris -n kserve-test -o jsonpath='{.status.url}' | cut -d "/" -f 3)

curl -v -H "Host: ${SERVICE_HOSTNAME}" -H "Content-Type: application/json" \
  "http://${INGRESS_HOST}:${INGRESS_PORT}/v1/models/sklearn-iris:predict" -d @./iris-input.json

The jsonpath expressions pull single values out of Kubernetes objects: the gateway's external IP address, its port, and the hostname from the InferenceService's .status.url. The cut command keeps only the hostname part of the URL. The lesson that matters beyond Istio: the Host header must be the hostname from .status.url, or the gateway has no idea which model you mean and answers with a 404. If you use the Gateway API instead of Istio, the address you call is that gateway's address, but the Host-header rule is the same.

Python users can skip curl. The requests library does the same job.

predict.py
import requests

url = "http://localhost:8080/v1/models/sklearn-iris:predict"
payload = {"instances": [[6.8, 2.8, 4.8, 1.4], [6.0, 3.4, 4.5, 1.6]]}

response = requests.post(url, json=payload, timeout=10)
response.raise_for_status()
print(response.json())
Try it
  1. Create iris-input.json, open the port-forward and run the curl command.
  2. Change one number in the input and send it again. Does the predicted class change?
  3. Send a row with only three numbers and read the error message you get back.
a predictions list for good input, and a readable error for bad input. The error comes from the model server itself, which tells you the request reached the model even though the input was wrong.

Everyday commands, grouped by what you are trying to do

KServe adds very few commands of its own. You work with kubectl, treating the InferenceService as one more Kubernetes object. This section groups the commands by intent, so you can find the one you need at the moment you need it.

Seeing what exists. List the models in one namespace, or in all namespaces. Use the isvc short name to save typing.

BASH
kubectl get isvc -n kserve-test
kubectl get isvc -A
kubectl get isvc sklearn-iris -n kserve-test -o yaml

The columns of the plain listing are the name, the URL, whether it is ready, and (depending on the mode) fields about revisions and traffic. The -o yaml form prints the whole object, including a status section that KServe fills in. When something is wrong, the status section is where the reason hides.

Finding out why. describe prints the object together with its recent events and conditions in human-friendly form.

BASH
kubectl describe isvc sklearn-iris -n kserve-test
kubectl get events -n kserve-test --sort-by=.lastTimestamp

Look in the Conditions part for lines whose status is False; each has a Reason and Message. The events listing is sorted by time, so the newest message, often the most relevant, is at the bottom.

Reading logs. A model pod has more than one container, so you must say which one you mean with -c. The model server container is called kserve-container, and the download helper is called storage-initializer.

BASH
kubectl get pods -n kserve-test
kubectl logs -n kserve-test <pod-name> -c kserve-container
kubectl logs -n kserve-test <pod-name> -c storage-initializer

Replace <pod-name> with the pod's real name from the first command. Use the storage-initializer logs when a pod is stuck at Init, and the kserve-container logs when the pod runs but misbehaves. Add -f to follow logs live.

Changing a model. Edit the YAML and apply it again. Kubernetes compares it with what is running and changes only what differs. For a one-off change, kubectl edit opens the live object in your editor, but editing the file is better because your file stays the source of truth.

BASH
kubectl apply -n kserve-test -f sklearn-iris.yaml
kubectl edit isvc sklearn-iris -n kserve-test

Removing a model. Delete the InferenceService, and KServe removes the pods, services and routes it created for it. You do not delete those by hand.

BASH
kubectl delete isvc sklearn-iris -n kserve-test
kubectl delete namespace kserve-test

Asking what a field means. kubectl explain documents any field of any resource directly from the cluster, so it always matches the installed version.

BASH
kubectl explain inferenceservice.spec.predictor
kubectl explain inferenceservice.spec.predictor.model

This habit is worth more than any cheat sheet, because KServe fields change between releases and the explain output is always the truth for your cluster.

Try it
  1. Run kubectl describe isvc on your model and find the Conditions section.
  2. Read the logs of both containers in the pod.
  3. Run kubectl explain inferenceservice.spec.predictor.model and list three fields you have not used yet.
healthy conditions, a short storage-initializer log that ends with the model being downloaded, and a list of optional fields. You now know where to look first when something breaks.

Model formats and serving runtimes

When you wrote modelFormat: name: sklearn, you relied on KServe to find a matching model server. This section explains how that matching works, because it explains most beginner surprises about "my model format is not supported".

A ServingRuntime is a Kubernetes object that describes one model server: the container image, its command-line arguments, its default resources, and the list of model formats it supports. After installation, a set of ClusterServingRuntimes are available in every namespace. Look at them.

BASH
kubectl get clusterservingruntimes
kubectl get clusterservingruntime -o wide

When you create an InferenceService with a modelFormat, the controller looks for a runtime that supports that format, and uses the first suitable one. The docs list formats such as scikit-learn, XGBoost, LightGBM, TensorFlow, PyTorch (through TorchServe), ONNX and others (served by runtimes like Triton), PMML, Paddle, and Hugging Face for transformer models. Which of these are present in your install depends on the version, so trust your own get clusterservingruntimes output over any list in a guide.

Why does the model format matter at all? Because a saved model file is meaningless without the code that reads it. A scikit-learn model is usually a pickled Python object; an XGBoost model is a booster file; a TensorFlow model is a SavedModel directory. Each needs a server that knows how to load it. The runtime is that server, and the modelFormat is your declaration of what the files are.

If you need a specific runtime rather than the automatic choice, name it.

YAML
spec:
  predictor:
    model:
      modelFormat:
        name: sklearn
      runtime: kserve-sklearnserver
      storageUri: "gs://kfserving-examples/models/sklearn/1.0/model"

The runtime field is the name of a ServingRuntime or ClusterServingRuntime. The name kserve-sklearnserver is an example of the naming style; copy the exact name from your own kubectl get clusterservingruntimes output, because names can differ between versions. Naming the runtime explicitly is also the fix when the controller complains that it cannot find a runtime for your model type.

There is one more decision hiding in the runtime: the protocol. Many runtimes can speak both v1 and v2. The default for the example above is v1, with URLs like /v1/models/sklearn-iris:predict. To use the v2 Open Inference Protocol, which has a uniform request layout across frameworks and also supports gRPC, you add a protocolVersion to the model.

YAML
spec:
  predictor:
    model:
      modelFormat:
        name: sklearn
      protocolVersion: v2
      storageUri: "gs://kfserving-examples/models/sklearn/1.0/model"

With v2 the request URL becomes /v2/models/<name>/infer, and the JSON body is organised as a list of named input tensors (arrays of numbers) with a shape and a data type, instead of the plain instances list. For a first project, stay with v1. Use v2 when a tool you integrate with expects the Open Inference Protocol, and read the runtime's page in the docs to confirm that it supports it.

Try it
  1. List the cluster serving runtimes and write down which model formats your install supports.
  2. Pick one runtime and run kubectl get clusterservingruntime NAME -o yaml. Find the supportedModelFormats list and the container image.
a YAML document that names an image and a list of supported formats. A runtime is just a template; KServe fills in your model's address and resource settings.

Where models come from: the storage URI

The storageUri is the single most important field after modelFormat, and the source of many errors. It tells the storage initializer where to download the model from. The initializer copies the files to /mnt/models inside the pod, so the model server always finds them in the same place regardless of the source.

Supported address schemes include:

Scheme Source Example
gs:// Google Cloud Storage gs://my-bucket/models/churn/1
s3:// Amazon S3 and compatible stores s3://my-bucket/models/churn/1
pvc:// A Kubernetes PersistentVolumeClaim pvc://model-store/churn/1
hf:// Hugging Face Hub hf://org/model-name
https:// A web address https://example.com/model.joblib
oci:// A container image used as a model package see the docs

The docs also list Azure Blob Storage addresses, and v0.21 adds a ms:// scheme for ModelScope. The full list changes between releases, so check the current storage page before relying on a scheme that is not in the table.

What the storage location must contain. The folder you point to must hold the files the runtime expects. For scikit-learn the example folder holds a file called model.joblib; the runtime looks for it by that name. If your folder holds my_model.pkl, the server cannot find it. The runtime's page in the docs states the file names it looks for, so check it before you upload. Many "model fails to load" problems are simply the wrong file name or an extra nested folder.

Public versus private storage. The example bucket is public, so nobody needs a password. Your own buckets are almost always private. The download then needs credentials, and KServe expects them in a Kubernetes Secret attached to a ServiceAccount, and the InferenceService must run as that service account. The exact keys and annotations differ per cloud, and getting them right is the biggest beginner stumbling block with private storage. The Mid-level part walks through it. As a beginner, if you want to practise without credentials, use public examples, or use a pvc:// address with a volume you created yourself.

A PersistentVolumeClaim (PVC) is Kubernetes' way of asking for disk space that survives pod restarts. Pointing storageUri at a PVC is handy in a lab: you copy your model onto the volume once, and every pod reads from it.

Try it
  1. Change the storageUri in a copy of your YAML to a path that does not exist, such as gs://kfserving-examples/models/sklearn/1.0/nothing, and deploy it under a new name.
  2. Read the storage-initializer logs of the new pod.
  3. Delete the broken InferenceService.
the pod stays in an init or error state and the initializer log explains it could not fetch the files. Learning to recognise this failure on purpose makes it far less scary when it happens by accident.

Resources, scaling and the three modes in plain language

Once a model works, three practical questions follow: how much machine does it need, how many copies should run, and how does it handle quiet periods. This section gives the beginner-level answers and flags what is left for later parts.

Requests and limits. You already wrote them in the YAML. A request is a reservation: the Kubernetes scheduler only places the pod on a node with that much spare capacity. A limit is a ceiling: a container that exceeds its memory limit is killed (OOMKilled), and one that exceeds its CPU limit is slowed down. Beginners often set limits too low for the model and then see restarts. A rule of thumb for small models is to measure memory use after the model is loaded and set the limit comfortably above it.

BASH
kubectl top pod -n kserve-test

This command needs the Kubernetes metrics server, which kind does not include by default. If you see an error saying the metrics API is not available, that is why; the command is not essential for this guide.

Replica counts and autoscaling. The predictor accepts minReplicas and maxReplicas.

YAML
spec:
  predictor:
    minReplicas: 1
    maxReplicas: 3
    model:
      modelFormat:
        name: sklearn
      storageUri: "gs://kfserving-examples/models/sklearn/1.0/model"

With these set, KServe keeps at least one copy running and may add up to three when load rises. What "load" means depends on the deployment mode. In Standard mode the scaling is done by a Kubernetes HorizontalPodAutoscaler, which by default reacts to CPU or memory use. In Knative mode, the Knative autoscaler reacts to the number of requests in flight. KServe can also integrate KEDA for scaling on custom metrics; that is an advanced topic.

Scale to zero. Setting minReplicas: 0 lets a model scale down to nothing when no requests arrive, which saves resources. This works only in Knative (serverless) mode. In Standard mode a model always has at least one pod. The cost of scale-to-zero is a cold start: the first request after idle time waits while a pod starts, an image may be pulled, and the model downloads and loads. For a small scikit-learn model that wait is seconds; for a large neural network it can be minutes.

Which mode should a beginner pick? The official docs and the project's release notes describe the trade-off this way. Standard mode has the fewest moving parts, suits steady traffic and GPU-heavy models, and is the simplest to learn. Knative mode adds scale-to-zero, revisions and percentage-based canary traffic, but you must install and operate Knative as well. ModelMesh is for thousands of small models. This guide uses Standard. If you meet a tutorial that uses canaryTrafficPercent or scale-to-zero, it is written for Knative mode.

Canary rollouts are worth one sentence now, because you will hear the word. A canary rollout sends a small share of traffic (say ten percent) to a new model version while the old one handles the rest, so you can watch for problems before switching everyone. KServe supports this by setting canaryTrafficPercent in Knative mode, and the v0.21 release notes describe canary traffic splitting for Standard mode with Gateway API routes too. The details belong to the Mid-level part.

Try it
  1. Add minReplicas: 2 to your predictor, apply it, and run kubectl get pods -n kserve-test.
  2. Delete one pod with kubectl delete pod and watch KServe's underlying Deployment bring it back.
  3. Set minReplicas back to 1 and apply.
two pods appear, a deleted pod is replaced within seconds, and the count settles back to one. That is the reconcile loop doing its job.

Configuration you will meet

Most of what you configure lives in the InferenceService YAML. A few things live elsewhere, and you should know where, even if you never change them in this guide.

The inferenceservice-config ConfigMap. A ConfigMap is a Kubernetes object that holds settings as text. KServe's global settings live in one called inferenceservice-config, in the kserve namespace. It has sections (keys) such as deploy (the default deployment mode), ingress (gateway settings), storageInitializer (the downloader's image and resource limits), logger, batcher and explainers. You read it like this.

BASH
kubectl get configmap inferenceservice-config -n kserve -o yaml

Reading it is safe and educational. Editing it affects every model on the cluster, and the controller may need to restart to pick up the change, so treat that as an administrator's task and prefer Helm values over hand edits.

Annotations. Annotations are key-value labels in metadata.annotations. KServe reads some of them to change behaviour per model. The docs show the deployment mode being selected per model with serving.kserve.io/deploymentMode, for example. Annotation keys are exact strings, and a typo is silently ignored, so copy them from the documentation for your version.

Try it
  1. Print the inferenceservice-config ConfigMap and find the deploy key.
  2. Find the storageInitializer key and note the memory limit.
JSON text inside each key. Seeing where the defaults live demystifies behaviour you would otherwise think is magic.

Reading the common errors

Things will fail, and that is normal. The skill that separates a beginner from a competent operator is not avoiding errors but reading them in the right order. Always go from the outside in: the InferenceService status first, then the pods, then the container logs.

BASH
kubectl describe isvc NAME -n NAMESPACE
kubectl get pods -n NAMESPACE
kubectl describe pod POD -n NAMESPACE
kubectl logs POD -n NAMESPACE -c storage-initializer
kubectl logs POD -n NAMESPACE -c kserve-container

Here are the failures beginners meet most, with the cause and fix of each.

No such file or directory: '/mnt/models'. The model server started but no model was downloaded. The official docs give one documented cause: the InferenceService was created in a control-plane namespace, where the storage initializer is not injected. Fix: delete it and recreate it in an ordinary namespace such as kserve-test.

failed calling webhook ... when you apply. The API server could not reach KServe's webhook, or the webhook's certificate is invalid. Causes: the KServe controller is not running yet, or cert-manager is missing or not issuing certificates. Fix: run kubectl get pods -n kserve, kubectl get pods -n cert-manager and kubectl get certificate -n kserve, wait until all are ready, and apply again. The exact wording of the message differs between versions, so search for the words "webhook" and "kserve" in the output.

A very long annotation error during installation. If you installed with plain kubectl apply and an error says an annotation is too long, the large custom resource definition was applied client-side. Use kubectl apply --server-side.

READY stays False. Run kubectl describe isvc and read the Conditions. Each failing condition has a reason. Typical reasons point at the predictor pod not being ready, or at the networking layer not being configured. In Knative mode, the reasons often refer to revisions and routes, which means Knative or the network layer is not installed correctly. In Standard mode, look at the pod.

ImagePullBackOff or ErrImagePull. The cluster cannot download the container image. Causes: a typo in a custom image name, a private registry without credentials, no internet access from the cluster, or an architecture mismatch (an Intel-only image on an Apple Silicon machine). Fix: read kubectl describe pod events, which print the exact image and the pull error.

Pending pod with Insufficient cpu, Insufficient memory or Insufficient nvidia.com/gpu. The scheduler cannot find a node with room. On a small kind cluster this usually means your requests are too large, or other pods are using the capacity. Lower the requests, delete unused models, or add nodes. For GPUs, the cause may be no GPU capacity, a missing device plugin, or missing tolerations.

Init container OOMKilled. The storage initializer ran out of memory downloading a large model. Raise its memory in the storageInitializer section of the configuration.

Access denied (403) in the initializer log. The bucket is private and no credentials were attached through a service account secret. Fix: configure the credentials as described in the Mid-level part, or switch to a public example while you practise.

A message that no runtime supports the model type. The modelFormat name matches no ServingRuntime. Check spelling against kubectl get clusterservingruntimes, or name a runtime explicitly. The exact wording differs between versions.

curl returns 404 or no healthy upstream through a gateway. The Host header is wrong, the route is not accepted, or the gateway is not ready. Fix: use the hostname from .status.url as the Host header, then check kubectl get httproute -n NAMESPACE and the gateway's status.

Try it
  1. Deploy a copy of the example with a deliberately wrong modelFormat, such as sklearnn.
  2. Read kubectl describe isvc and the controller logs: kubectl logs -n kserve deploy/kserve-controller-manager -c manager.
  3. Fix the typo, apply again and confirm it turns Ready.
an error that names the missing runtime or format, and a recovery by editing the file. Breaking things deliberately in a lab is the fastest way to learn their messages.

Putting it all together

Here is a small end-to-end project that uses everything above in one sitting. It takes about thirty minutes the first time. You will create a clean cluster, install KServe, deploy the iris model with sensible resources, call it from a Python script, scale it, break it on purpose and tear everything down.

Step 1: a fresh cluster and a script for the install. Putting the install in a script means you can recreate the lab whenever you like. Note the versions at the top.

install-kserve.sh
#!/usr/bin/env bash
set -euo pipefail

KSERVE_VERSION=v0.21.0
CERT_MANAGER_VERSION=v1.15.0
GATEWAY_API_VERSION=v1.2.1

kind create cluster --name kserve-demo

kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/${CERT_MANAGER_VERSION}/cert-manager.yaml
kubectl wait --for=condition=Available deployment --all -n cert-manager --timeout=180s

kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/${GATEWAY_API_VERSION}/standard-install.yaml

helm install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd --version ${KSERVE_VERSION}
helm install kserve oci://ghcr.io/kserve/charts/kserve-resources --version ${KSERVE_VERSION} \
  --set kserve.controller.deploymentMode=Standard \
  --set kserve.controller.gateway.ingressGateway.enableGatewayApi=true \
  --set kserve.controller.gateway.ingressGateway.kserveGateway=kserve/kserve-ingress-gateway

kubectl wait --for=condition=Available deployment --all -n kserve --timeout=300s
kubectl get clusterservingruntimes

set -euo pipefail makes the script stop at the first failure instead of ploughing on. The wait commands replace guessing with an explicit pause until things are ready.

Step 2: the model. Use the YAML from earlier, with a minReplicas so you can see scaling.

sklearn-iris.yaml
apiVersion: "serving.kserve.io/v1beta1"
kind: "InferenceService"
metadata:
  name: "sklearn-iris"
spec:
  predictor:
    minReplicas: 1
    maxReplicas: 2
    model:
      modelFormat:
        name: sklearn
      storageUri: "gs://kfserving-examples/models/sklearn/1.0/model"
      resources:
        requests:
          cpu: "100m"
          memory: "512Mi"
        limits:
          cpu: "1"
          memory: "1Gi"
BASH
kubectl create namespace kserve-test
kubectl apply -n kserve-test -f sklearn-iris.yaml
kubectl wait --for=condition=Ready isvc/sklearn-iris -n kserve-test --timeout=300s

The wait command blocks until the InferenceService reports the Ready condition, then returns, which makes the project scriptable.

Step 3: a client script. This version sends several flowers and prints readable names.

classify.py
import requests

URL = "http://localhost:8080/v1/models/sklearn-iris:predict"
SPECIES = {0: "setosa", 1: "versicolor", 2: "virginica"}

flowers = [
    [5.1, 3.5, 1.4, 0.2],
    [6.8, 2.8, 4.8, 1.4],
    [6.3, 3.3, 6.0, 2.5],
]

response = requests.post(URL, json={"instances": flowers}, timeout=10)
response.raise_for_status()

for flower, label in zip(flowers, response.json()["predictions"]):
    print(flower, "->", SPECIES[label])

Open the tunnel in one terminal (kubectl port-forward -n kserve-test svc/sklearn-iris-predictor 8080:80, using the service name you confirmed with kubectl get svc) and run python classify.py in another. Each flower gets one species name, and the first row, with its small petals, should come out as setosa.

Step 4: observe and break. Run kubectl get pods -n kserve-test, read the logs of kserve-container, then change the storageUri to a wrong path and apply. Watch the new pod fail, read the storage-initializer log, then restore the correct path and apply again. Because Kubernetes performs a rolling update, the old pod keeps serving until a new one is ready, so a bad update does not take your endpoint down in Standard mode. Confirm that for yourself by running the classify script while the broken update is stuck.

Step 5: clean up. Delete things in the reverse order of creation.

BASH
kubectl delete isvc sklearn-iris -n kserve-test
kubectl delete namespace kserve-test
helm uninstall kserve
helm uninstall kserve-crd
kind delete cluster --name kserve-demo

Uninstall kserve before kserve-crd, because removing the definitions first would orphan objects that depend on them. Deleting the kind cluster at the end removes everything, including cert-manager, in one step, which is why it is the best lab setup.

Try it
  1. Run the whole project from install-kserve.sh to the cleanup, without copy-pasting from this page.
  2. Add a fourth flower of your own to classify.py.
  3. Time how long each stage takes and note which is slowest.
a working endpoint, species names printed in your terminal, and a clean machine afterwards. The slowest stage is usually the first image pull; that is the cold start you read about.

What you can now do, and what comes next

You started this guide with a model file and ended with a web service you installed, deployed, called, scaled, broke and removed. In concrete terms, you can now:

  • explain what KServe does and why a team prefers it to hand-written servers,
  • name the five core nouns (InferenceService, predictor, ServingRuntime, storage initializer, deployment mode) and say how they connect,
  • install KServe with its prerequisites on a local cluster, and verify the install in four checks,
  • write an InferenceService with a model format, a storage URI and resource settings,
  • call the model with curl and Python, using both a port-forward and a gateway Host header,
  • find the cause of a failing model by working from the InferenceService status to pods to container logs,
  • recognise the most common errors and apply the right fix.

Good neighbours to read alongside this series: Kubernetes for the platform underneath everything here, Docker for images, Helm for the chart-based install, MLflow for the models you will serve, BentoML for an alternative way to package and serve them, and Evidently for watching a deployed model's data and quality over time.

Sources