This is part one of three. It covers everything you need to put a real model behind a real HTTP endpoint with Seldon Core 2, on a cluster you control. By the end you can install Seldon Core 2 on a Kubernetes cluster, deploy a trained model as a Model resource, call it over the Open Inference Protocol, understand why the scheduler sometimes refuses to place your model, run your own pinned inference server, and chain two models into a Pipeline. Mid-level and Senior take the same material further into scaling, Kafka tuning, security and platform ownership; nothing here is throwaway.
Each section ends with a Try it task. Do them in order on a throwaway cluster. Seldon Core 2 is a distributed system, and the only way the vocabulary sticks is to watch your own model go Ready, then break it on purpose and read the condition it reports.
Before anything else, one warning that will save you hours, because it is the single most common way beginners lose an afternoon.
Model, Server, Pipeline and Experiment under the API group mlops.seldon.io/v1alpha1. Seldon Core 1 is an older, different architecture built around a single SeldonDeployment resource and a Python wrapper library. The v1 line is still released — v1.19.0 landed on 23 January 2026 as a maintenance and security release — but it is a separate product, not a step towards v2. Nearly every Stack Overflow answer, conference talk and blog post you find will be about v1. If a tutorial tells you to write kind: SeldonDeployment, or to pip install seldon-core and subclass a Python class, you are reading about v1 and none of it applies here.
What Seldon Core is, and the problem it solves
Training a model is a self-contained piece of work. You have data, you have a notebook, and at the end you have an artifact on disk: a pickled scikit-learn pipeline, an XGBoost booster, a saved TensorFlow graph. That artifact is worth nothing to the business until something can be asked, thousands of times a second, "here is a row of features, what do you predict?" and get an answer back in milliseconds, reliably, with a record of what was asked and what was answered.
Serving is where most machine-learning projects actually stall, and the reasons are depressingly repetitive. Somebody writes a small Flask application that loads the pickle at startup and exposes /predict. It works. Then there are nine models, so there are nine Flask applications, each with a slightly different request format, each pinned to a different scikit-learn version, each running one replica on a node with eight idle gigabytes of RAM. Then somebody asks what the model predicted for customer 48213 last Tuesday, and nobody knows, because the only record is a log line that was rotated away. Then a data scientist wants to A/B test a new version against the old one, and the only mechanism available is an if statement in the application code.
Seldon Core 2 is a Kubernetes-native framework that replaces that pile of bespoke services with a small set of declarative resources. The official documentation describes it as a framework for deploying and managing machine-learning and LLM systems at scale, and organises its own value around four claims worth unpacking in plain terms:
Flexibility
The same resources run on-premises, in a cloud, or in a hybrid setup. It is Kubernetes underneath, so wherever Kubernetes runs, this runs.
Standardisation
One request format, the Open Inference Protocol, for every framework. Learn to call a scikit-learn model and you already know how to call a PyTorch one.
Observability
Prometheus metrics for operations, and — for pipelines — every input and output retained on a Kafka topic, so predictions are auditable and replayable.
Optimisation
Multi-model serving packs many models onto shared server replicas instead of giving each one its own idle pod.
The second and fourth of those are the ones that change your daily life. Standardisation means your client code does not care which framework trained the model. Multi-model serving means a server replica is a host for models rather than a wrapper around one model: a single MLServer pod can hold a dozen small models, loading and unloading them on demand. If you have forty small models, that is the difference between forty pods and three.
The third claim deserves its own sentence, because it is the philosophical centre of the project. Seldon calls its approach data-centric MLOps: the data flowing through your models is treated as a first-class concern rather than an afterthought. That is why Apache Kafka sits in the middle of the architecture rather than at the edge of it. When you build a pipeline, the inputs and outputs of every step are written to Kafka topics, which is what makes an audit trail and a replay possible at all.
- List the models your team or your coursework has trained in the last year.
- For each, write down how it is served today: a notebook, a Flask app, a batch job, or nowhere.
- Count how many distinct request formats a client would have to learn to call all of them.
The version landscape, and what to install
Pinning down which version you are on is not bureaucracy here; it decides which features exist and which bugs you will hit.
The current v2 release line is 2.10.x, with v2.10.2 published on 19 December 2025. The 2.10 series is where a number of things changed in ways a beginner should know about up front, because they contradict older documentation.
Three practical consequences follow.
Install 2.10.2, not an earlier 2.10. The two patch releases after 2.10.0 fix problems you are likely to meet: pipelines that will not create or delete, and models that quietly come back on fewer replicas than you asked for after a node reboot. Starting on the patched version is free; debugging a known bug is not.
Model-native autoscaling was deliberately disabled in 2.10.0 pending wider fixes. If you read an older page describing automatic replica scaling driven by the model itself, that path is not available on a current install. The official advice is to scale inference servers with the Kubernetes Horizontal Pod Autoscaler or with KEDA instead. The Mid-level guide covers how; at beginner level, set replica counts by hand and know that the automatic route is off.
The documentation moved. Older URLs under /user-guide/kubernetes-resources/... now return 404. The current pages are flat — /seldon-core-2/user-guide/models, /servers, /pipelines, /quickstart — and historical versions live under versioned paths such as /seldon-core-2/v2.10/.... If a search result 404s, strip it back to the docs home and navigate down.
LICENSE file in the SeldonIO/seldon-core repository for the authoritative terms before you build a product on it. Nothing in this guide is legal advice, and learning on a throwaway cluster is not the case you need to worry about.
- Open the releases page of the
SeldonIO/seldon-corerepository. - Find the newest tag beginning
v2.and note it down; that is the version this guide's commands target as 2.10.2. - Read its release notes top to bottom, even the parts you do not yet understand.
The four nouns
Almost everything in Seldon Core 2 is one of four Kubernetes custom resources. Learn these four words properly and the rest of the system becomes readable.
| Server | Model | Pipeline | Experiment | |
|---|---|---|---|---|
| Is | A deployment of an inference server | One trained artifact, loaded and addressable | A graph of models wired together | A traffic split between models or pipelines |
| Kubernetes shape | A StatefulSet of replicas | Not a pod of its own; it lives inside a Server | Kafka topics plus a dataflow graph | Routing rules in the data plane |
| You create | Rarely — two come with the install | Constantly | When one call needs several models | For A/B tests and shadow traffic |
| Analogy | An application server | A deployed application | A chain of services | A load balancer with weights |
A Server is a running inference server: a StatefulSet of replicas able to hold models. Two server "farms" ship with the default installation, each with one replica: MLServer, Seldon's own Python-based server, and NVIDIA Triton. You do not usually create a Server to deploy your first model; you use one that is already there.
A Model is the resource you will write most often. It is a Kubernetes object that points at a trained artifact and says what the artifact needs in order to run. Critically, a Model is not a pod. It is a request to the scheduler to load an artifact onto some Server replica that can host it. That indirection is multi-model serving, and it is the thing newcomers most often get wrong when they go looking for "my model's pod" and cannot find one.
A Model is also more general than the word suggests. The documentation is explicit that the same resource represents a predictive model, a drift detector, an outlier detector, an explainer, a feature transformation or a routing model. Anything that takes tensors in and gives tensors out is a Model.
A Pipeline is a directed graph of Models, wired together through Kafka topics. You use one when a single prediction needs more than one model: a feature transformer followed by a classifier, or a classifier with a drift detector watching the same input.
An Experiment splits traffic between candidate models or pipelines by weight, which is how you run an A/B test without putting an if statement in your client. It is listed here for completeness; it belongs to the Mid-level guide.
Four supporting components complete the picture, and their names appear in logs often enough to be worth memorising. The operator, also called the controller manager, watches your custom resources and reconciles them. The scheduler is the brain of the control plane: it decides which Server replica each Model should live on. The agent is a sidecar container on every server pod that downloads artifacts and asks the inference server to load and unload them. And Envoy, exposed through a Kubernetes Service called seldon-mesh, is the single front door for the data plane: every inference request enters there and is routed to the right model.
- Without looking back, write one sentence each for Server, Model, Pipeline and Experiment.
- Then answer: if a Model is not a pod, where does it actually run?
- Check your answers against the table above.
What you need before you install
Seldon Core 2 is a Kubernetes system and only a Kubernetes system. There is no native macOS or Windows installer, and no single binary that serves a model on your laptop. What you need is a cluster.
The production installation page states the requirements plainly: Kubernetes 1.27 or later, plus kubectl and helm. For learning, a local cluster is fine, and kind — Kubernetes in Docker — is the least troublesome option.
# Linux or macOS, with Docker already running
kind create cluster --name seldon
kubectl cluster-info --context kind-seldon
kubectl version
The last command matters: confirm the server version is 1.27 or newer before you go further. Installing onto an older cluster produces Helm and CRD errors that look mysterious until you remember to check.
On macOS, Homebrew gets you the tools: brew install kubectl helm kind. You need a container runtime for kind to use, which means Docker Desktop, Colima or OrbStack. On Apple Silicon, check that the inference server images you intend to use publish an arm64 manifest before you rely on them. MLServer and Triton are large images with native dependencies and architecture support is not something to assume; if a pod sits in a crash loop with an exec format error, that is what happened.
On Windows, use WSL2 with an Ubuntu distribution plus Docker Desktop's WSL integration, and run every command from the WSL shell. kubectl and helm work natively in PowerShell, but the sample scripts and the curl invocations in the documentation are written for bash, so a Linux shell saves you constant translation.
ImagePullBackOff accompanied by x509: certificate signed by unknown authority almost always means something is inspecting TLS between your machine and the registry — a proxy, a VPN client, or a zero-trust gateway. The fix is to export that interceptor's root certificate authority and add it to the trust store your container runtime uses. Do not disable certificate verification to get past it; you will leave that setting behind and it will be the next person's security incident.
Two optional pieces are worth knowing about now. Kafka is required for pipelines and not required for calling models directly, so you can do the first half of this guide without it. A Prometheus stack in a seldon-monitoring namespace gives you the metrics the Mid-level guide builds on. The documentation also describes a learning environment with a self-hosted Kafka setup, which is the shortest path to a laptop cluster that can run pipelines.
- Create a
kindcluster namedseldon. - Run
kubectl versionand confirm the server version is 1.27 or newer. - Run
helm versionandkubectl get nodes.
Ready state and two working CLIs. That is the whole prerequisite list, and everything from here on is Helm and YAML.
Installing Seldon Core 2 with Helm
The installation is four Helm charts applied in order. The order is not cosmetic, and understanding why teaches you something about how Kubernetes extensions work.
kubectl create ns seldon-mesh || echo "Namespace seldon-mesh already exists"
kubectl create ns seldon-monitoring || echo "Namespace seldon-monitoring already exists"
helm repo add seldon-charts https://seldonio.github.io/helm-charts/
helm repo update seldon-charts
Two namespaces: seldon-mesh for the runtime and your models, seldon-monitoring for the metrics stack. Now the charts.
helm upgrade seldon-core-v2-crds seldon-charts/seldon-core-v2-crds \
--namespace default --install
helm upgrade seldon-core-v2-setup seldon-charts/seldon-core-v2-setup \
--namespace seldon-mesh --set controller.clusterwide=true --install
helm upgrade seldon-core-v2-runtime seldon-charts/seldon-core-v2-runtime \
--namespace seldon-mesh --install
helm upgrade seldon-core-v2-servers seldon-charts/seldon-core-v2-servers \
--namespace seldon-mesh --install
kubectl get pods -n seldon-mesh
What each chart does, in the order they must be applied:
seldon-core-v2-crdsinstalls the custom resource definitions — the API extensions that teach your cluster the wordsModel,Server,PipelineandExperiment. CRDs are cluster-scoped, which is why this one goes intodefaultrather thanseldon-mesh. Nothing else can work until the cluster knows these kinds exist.seldon-core-v2-setupinstalls the operator, the controller that watches those resources and acts on them.controller.clusterwide=truetells it to watch every namespace. The alternative is a namespaced operator:--set "controller.watchNamespaces={ns1,ns2}", which is how multi-tenant installations isolate teams.seldon-core-v2-runtimeinstalls the per-namespace runtime: the scheduler, Envoy, the model and pipeline gateways, and the dataflow engine. This is the machinery that actually places and serves models.seldon-core-v2-serversinstalls the default Servers — one MLServer and one Triton, one replica each.
The helm upgrade ... --install form is used throughout rather than helm install, because it is idempotent: the same command creates the release the first time and updates it afterwards, which makes it safe to put in a script or re-run after a typo.
Give the pods a minute, then look:
kubectl get pods -n seldon-mesh
kubectl get servers -n seldon-mesh
You should see the controller manager, the scheduler, Envoy, the dataflow engine, the model gateway, the pipeline gateway, and StatefulSets for mlserver and triton. Exact pod names vary with the version and your values, so read what your cluster actually prints rather than matching a list from a blog post. kubectl get servers is the command to remember: it is how you find out what your models are allowed to be scheduled onto.
--version to each helm upgrade so that re-running your install script months later produces the same cluster. Without it, Helm takes whatever is newest in the repository, and an unplanned minor upgrade is a bad way to start a debugging session.
Uninstalling runs in reverse: servers, runtime, setup, then CRDs. Leave the CRD chart for last and treat it with respect — deleting the CRDs deletes every Model, Server and Pipeline in the cluster, because Kubernetes garbage-collects the custom objects when their definition disappears. More than one person has removed the "harmless CRDs chart" first and taken production with it.
- Run all four
helm upgradecommands against yourkindcluster. - Watch the pods come up with
kubectl get pods -n seldon-mesh -w. - Run
kubectl get servers -n seldon-meshand note the names and replica counts.
kubectl describe pod will usually say your laptop is out of memory.
Your first model
The canonical first model is an iris classifier: a scikit-learn model that takes four flower measurements and predicts a species. Seldon publishes the trained artifact in a public bucket, so you do not need to train anything.
apiVersion: mlops.seldon.io/v1alpha1
kind: Model
metadata:
name: iris
spec:
storageUri: "gs://seldon-models/scv2/samples/mlserver_1.5.0/iris-sklearn"
requirements:
- sklearn
memory: 100Ki
Four lines of spec carry the whole idea, so read each one.
apiVersion: mlops.seldon.io/v1alpha1 is the API group and version. Note the v1alpha1: there is no v1 API for Core 2 yet. Every Core 2 resource you write uses this exact string, and getting it wrong produces no matches for kind "Model", which is the same error you get when the CRDs are not installed.
storageUri points at the artifact. The value is an rclone-style URI, which means Seldon supports the storage backends rclone supports: gs:// for Google Cloud Storage, s3:// for S3 and S3-compatible stores, local paths on a persistent volume, and more. Notice the mlserver_1.5.0 segment in the sample path — Seldon publishes the same sample artifacts rebuilt for different MLServer versions, because a pickle saved by one scikit-learn version may not load in another. That detail is a preview of a real production concern.
requirements is a list of capability tags. This is the most important field to understand, and the next section is devoted to it. For now: sklearn means "this artifact needs a server that can run scikit-learn models".
memory is a declared budget the scheduler uses when deciding how many models fit on a replica. The 100Ki in the sample is a toy value chosen because the iris model genuinely is tiny. For anything real, measure and declare honestly; the consequences of lying are covered below.
Apply it and wait:
kubectl apply -f iris.yaml -n seldon-mesh
kubectl get model iris -n seldon-mesh
kubectl wait --for condition=ready --timeout=300s model iris -n seldon-mesh
kubectl wait is worth adopting as a habit rather than refreshing get in a loop. It blocks until the resource reports itself ready, or fails after the timeout — which makes it usable in scripts and in CI, and makes "it never became ready" an explicit outcome rather than something you give up waiting for.
While you are waiting, look for your model's pod. There isn't one. kubectl get pods -n seldon-mesh shows the same pods as before. The iris model was loaded into the existing MLServer replica by the agent sidecar, which fetched the artifact from the bucket and asked MLServer to load it. That is multi-model serving in action, and it is why the second, third and tenth model you deploy cost you almost nothing.
- Apply
iris.yamland wait for it to become ready. - Run
kubectl get pods -n seldon-meshand confirm no new pod appeared. - Run
kubectl describe model iris -n seldon-meshand read the conditions block.
Calling the model: the Open Inference Protocol
Seldon Core 2 speaks the Open Inference Protocol, or OIP — the standard formerly known as the V2 inference protocol and still widely called KServe V2. It is a shared contract, which is the practical meaning of Seldon's "learn once, deploy anywhere" claim: KServe speaks it too, so client code written against one works against the other.
The protocol is available over HTTP/REST and over gRPC. The inference path has a fixed shape:
POST /v2/models/<model-name>/infer
First you need an address. On a cloud cluster with a load balancer, the mesh has an external IP:
kubectl get svc seldon-mesh -n seldon-mesh \
-o jsonpath='{.status.loadBalancer.ingress[0].ip}'
On kind, minikube or any laptop cluster there is no load-balancer controller, so that command prints nothing. Forward the service port instead:
kubectl get svc seldon-mesh -n seldon-mesh
kubectl port-forward svc/seldon-mesh -n seldon-mesh 8080:80
Run the get svc first and read the port column; map your local 8080 to whatever port the service actually exposes rather than trusting a number from a tutorial. With the forward running in one terminal, call the model from another:
curl -X POST http://localhost:8080/v2/models/iris/infer \
-H "Content-Type: application/json" \
-H "Seldon-Model: iris" \
-d '{"inputs":[{"name":"predict","shape":[1,4],"datatype":"FP32","data":[[1,2,3,4]]}]}'
Two things in that request decide whether it works.
The Seldon-Model header is how Envoy routes. There is one seldon-mesh Service for every model in the namespace, and the header is what tells the mesh which model you mean. Omit it and the request does not reach your model, no matter how correct the URL path is. This trips up almost everyone exactly once, because the path already contains the model name and it feels redundant. It is not: the path is part of the OIP contract that the inference server sees, and the header is the mesh's routing key.
The body is a tensor description, not a row of JSON. Each entry in inputs has a name the model expects, a shape, a datatype from the OIP type list, and data. The datatype must match what the model was trained on: FP32 for the iris floats, INT64 for integer features. Sending FP32 where INT64 is expected gets you an error from the inference server rather than a wrong prediction, which is the right behaviour but still a surprise the first time.
Here is the shape for a different sample, the income classifier, which takes twelve integer features:
{"inputs":[{"name":"income","datatype":"INT64","shape":[1,12],"data":[53,4,0,2,8,4,2,0,0,0,60,9]}]}
Note that the tensor name changed with the model. Tensor names are part of the model's own contract, not Seldon's, and reading them off the model's metadata or its training code is part of deploying it.
When you are done, remove the model the same way you created it:
kubectl delete -f iris.yaml -n seldon-mesh
Declarative all the way through: the file is the source of truth, and apply and delete are the only two verbs you need.
- Port-forward the mesh and call iris successfully.
- Now send the exact same request with the
Seldon-Modelheader removed, and read what comes back. - Send it again with the header present but
"datatype":"INT64", and read that error too.
Capabilities and scheduling
This section explains the error you are most likely to hit, so it is worth reading even if everything is working.
Every Server advertises a set of capabilities: tags describing what it can run. Every Model declares requirements: tags describing what it needs. The scheduler places a Model on a Server only when the Model's requirements are a subset of the Server's capabilities. That is the whole matching rule, and it is string comparison — no clever inference, no fuzzy matching.
The default MLServer advertises these capabilities:
mlserver, alibi-detect, alibi-explain, huggingface, lightgbm,
mlflow, python, sklearn, spark-mlib, xgboost
The default Triton advertises these:
triton, dali, fil, onnx, openvino, python, pytorch,
tensorflow, tensorrt
Read those two lists together and the division of labour is obvious. Classical machine learning — scikit-learn, XGBoost, LightGBM, MLflow-packaged models, Hugging Face transformers — goes to MLServer. Deep-learning graphs and optimised runtimes — ONNX, PyTorch, TensorFlow, TensorRT, OpenVINO — go to Triton. Note spark-mlib, spelled exactly that way in the capability list, which is a reminder that these are opaque strings rather than names with meaning: you match the string that is there, including its typo.
Capability matching is only the first of three filters. The scheduler also checks that enough Server replicas exist for the Model's requested replica count, and that each candidate replica has enough free memory for the Model's declared memory. Fail any of the three and the Model reports a condition ModelReady of False with the reason ScheduleFailed.
Diagnosing ScheduleFailed well
kubectl describe model <name>and read the conditionskubectl get servers -n seldon-meshto see what exists- Compare your
requirementsstrings letter by letter against the server's capabilities - Check whether you asked for more replicas than the Server has
- Check whether your
memoryfigure can fit
Diagnosing it badly
- Looking for the model's pod, finding none, and concluding the install is broken
- Deleting and re-applying the same YAML repeatedly
- Restarting the scheduler
- Assuming the artifact URI is wrong when the condition says scheduling
- Raising
memoryuntil it is placed, without measuring
On memory, one point matters enough to repeat: it is a budget the scheduler reasons with, not a limit the kernel enforces. Declaring 100Ki for a model that really needs two gigabytes does not cap it at 100Ki. It tells the scheduler the model is tiny, so the scheduler happily packs many more models onto that replica than will fit, and then the pod is killed for exceeding its container memory limit. The symptom appears far from the cause: a server replica restarting repeatedly, taking unrelated models down with it. Measure what your model's resident size actually is and declare that.
Since 2.9, the scheduler also supports partial scheduling: when capacity is short, a model may run on fewer replicas than requested rather than failing outright. This is usually what you want — degraded service beats no service — but it means "ready" does not always mean "at full requested capacity", so check the replica count rather than only the condition. The 2.10.1 release fixed a regression in exactly this area where models came back on too few replicas after pod restarts, which is another reason to be on a current patch version.
- Copy
iris.yamltobroken.yaml, rename the model, and changerequirementstoscikit-learn. - Apply it and run
kubectl describe modelon it. - Find
ScheduleFailedin the output, then fix the tag back tosklearnand re-apply.
scikit-learn is not sklearn" is a lesson best learned on purpose.
Running your own Server
The default Servers get you started, but eventually you need control over the image: a specific MLServer version, extra Python dependencies your model imports, or a GPU node. That means writing a Server resource.
apiVersion: mlops.seldon.io/v1alpha1
kind: Server
metadata:
name: mlserver-custom
spec:
serverConfig: mlserver
capabilities:
- income-classifier-deps
podSpec:
containers:
- image: seldonio/mlserver:1.6.0
name: mlserver
Three fields do the work. serverConfig: mlserver names a ServerConfig, a template that defines the pod shape, the agent sidecar and the rclone sidecar. You are not writing an inference server from scratch; you are starting from Seldon's MLServer template and overriding part of it. podSpec is that override, and because it is an ordinary Kubernetes pod spec you can set resources, node selectors, tolerations and environment variables there — which is how GPU scheduling is done, by requesting nvidia.com/gpu and targeting GPU nodes.
capabilities is the field to be careful with. Setting capabilities replaces the defaults entirely. The server above advertises exactly one tag, income-classifier-deps, and nothing else — so no model requiring sklearn can land on it. That is sometimes exactly what you want: a dedicated server for one model with one specific dependency set, with a tag that no other model will accidentally match. When instead you want to keep the defaults and add to them, use extraCapabilities. If both are set, capabilities wins.
A model targeting that server names the matching tag:
apiVersion: mlops.seldon.io/v1alpha1
kind: Model
metadata:
name: income-classifier
spec:
storageUri: "gs://seldon-models/scv2/samples/mlserver_1.4.0/income-sklearn/classifier"
requirements:
- income-classifier-deps
memory: 100Ki
This gives you a clean pattern for version pinning. The Servers documentation shows the same idea with a version tag: a server running image seldonio/mlserver:1.3.4 advertising the capability mlserver-1.3.4, and models that need that exact runtime requiring that exact tag. The capability string becomes a contract between an artifact and the runtime that can load it — which is much better than discovering the mismatch as an unpickling error in a container log.
Private artifact stores need credentials, and credentials belong in Kubernetes Secrets in rclone configuration format, referenced from the Model or configured for the agent — never inline in YAML you commit. Getting the exact field names right is worth doing against the current CRD reference rather than from memory, and the Mid-level guide covers the production pattern alongside external secret managers.
- Apply
mlserver-custom.yamland watch a new StatefulSet appear — this time there is a pod. - Apply
income-classifier.yaml, wait for ready, and call it with the twelve-integer payload above. - Now apply a copy of
iris.yamlwithserver: mlserver-custompinned, and watch it fail to schedule.
capabilities really does remove sklearn. That third step is the lesson.
Your first Pipeline
A Pipeline wires Models into a graph. The classic shape is a preprocessor feeding a classifier, so that the client sends raw input and the transformation happens server-side where it is versioned alongside the model.
apiVersion: mlops.seldon.io/v1alpha1
kind: Pipeline
metadata:
name: income-classifier-app
spec:
steps:
- name: preprocessor
- name: income-classifier
inputs:
- preprocessor
output:
steps:
- income-classifier
Each entry in steps names a Model that is already deployed. A Pipeline does not deploy models; it connects them. If a named step does not exist as a ready Model, the Pipeline will not become ready either, and the fix is upstream.
inputs is where the wiring lives: listing preprocessor under income-classifier means the classifier consumes the preprocessor's output. The general form uses dot notation — <stepName|pipelineName>.<inputs|outputs>.<tensorName> — which lets a step consume one specific named tensor rather than everything upstream produced.
output.steps declares which step's output is the pipeline's answer. It is required for synchronous REST and gRPC calls: without it, there is nothing to return to the caller.
Calling a pipeline looks almost identical to calling a model, with one difference that is easy to miss:
curl -X POST http://localhost:8080/v2/models/income-classifier-app/infer \
-H "Content-Type: application/json" \
-H "Seldon-Model: income-classifier-app.pipeline" \
-d '{"inputs":[{"name":"income","datatype":"INT64","shape":[1,12],"data":[53,4,0,2,8,4,2,0,0,0,60,9]}]}'
The .pipeline suffix in the Seldon-Model header is what tells the mesh you mean the pipeline rather than a model of the same name. The URL path does not carry it. Forget the suffix and you get a routing failure that reads as "model not found", which sends people hunting for a deployment problem that does not exist.
Underneath, each step's input and output is a Kafka topic. That is why Kafka is a hard dependency for pipelines and not for direct model calls, and it is also where the audit story comes from: the data that flowed through every step is sitting on a topic, so it can be inspected, replayed, or fed into Kafka Connect or ksqlDB. Since 2.10 a Kafka Schema Registry integration makes the schema contracts of those topics visible too.
Three more pipeline features are worth knowing by name so you recognise them when you need them. Joins control what happens when a step has several inputs: inner waits for all of them, outer waits for a window set by joinWindowMs and then proceeds with what arrived, and any proceeds on the first arrival. tensorMap renames tensors between steps when the producer's output name does not match the consumer's expected input name — the alternative being to retrain or rewrap a model over a naming mismatch. And triggers start a step when another step completes without passing that step's data along, which is how you express "run this only after that finished".
- Deploy the preprocessor and classifier models from the quickstart, then apply the pipeline.
- Call it with the
.pipelinesuffix in the header, then without it, and compare. - Delete the preprocessor Model and watch what the Pipeline's status does.
Reading errors
Seldon Core 2 reports problems in three places, and knowing which to read saves most of the time people waste on it. Kubernetes conditions tell you about scheduling. The agent container's logs tell you about fetching and loading artifacts. The inference server's logs tell you about running the model.
Here is the routine, in order:
kubectl get model,server,pipeline -n seldon-mesh
kubectl describe model iris -n seldon-mesh
kubectl logs -n seldon-mesh <server-pod> -c agent
kubectl logs -n seldon-mesh <server-pod> -c mlserver
The first command is the overview: it shows every resource and its readiness in one screen, which immediately tells you whether the problem is one model or the whole runtime. The second reads conditions. The third and fourth are the two containers that matter on a server pod — the agent, which fetched your artifact, and the inference server, which tried to load and run it.
The failures you are most likely to meet:
| Symptom | What you see | Cause and fix |
|---|---|---|
| Model never ready | ModelReady: False, reason ScheduleFailed |
No Server advertises all your requirements, or not enough replicas, or not enough free memory. Compare strings; check kubectl get servers. |
| Model not found when calling | A 404 or a not-found error | Missing or wrong Seldon-Model header, a typo in the name, a pipeline called without .pipeline, or a model still loading. |
Server pod in CrashLoopBackOff |
Restarts on mlserver or agent | Often rclone cannot fetch storageUri — wrong bucket path or missing credentials — or the artifact needs a different inference-server version. Read the agent log first, then the server log. |
ImagePullBackOff |
x509: certificate signed by unknown authority or a registry timeout |
TLS inspection or a proxy between the node and the registry, or no route to the registry at all. |
no matches for kind "Model" |
Immediate kubectl apply failure |
The CRDs chart is not installed, or apiVersion is wrong. |
no matches for kind "SeldonDeployment" |
Immediate failure | You are following a Seldon Core 1 tutorial on a Core 2 cluster. |
| Helm or CRD errors during install | Validation failures | Kubernetes older than 1.27. |
| Pipeline will not create or delete | Stuck not-ready pipeline | A known issue fixed in 2.10.2 — check your version first. Otherwise Kafka connectivity; read the dataflow-engine logs. |
| Hanging gRPC streaming calls | The call never returns | Fixed in 2.10.2. Also check your MLServer is 1.6.0 or newer, which streaming requires. |
| Fewer replicas than requested after a restart | Model ready but under-replicated | The partial-scheduling regression fixed in 2.10.1. |
Two of those rows deserve emphasis because they are version answers rather than configuration answers. "Pipelines fail to create or delete" and "gRPC streams block" were both long-standing bugs fixed in 2.10.2. If you are debugging either on an earlier version, you are debugging someone else's already-fixed bug. Check your version before you check your YAML is a good general habit with this project, given how much the 2.10 patch line changed.
no matches for kind — says nothing about versions. Before trusting any Seldon page, look for mlops.seldon.io (v2) or SeldonDeployment (v1) and decide whether it is talking about your system at all.
- Apply a Model whose
storageUripoints at a bucket path that does not exist. - Find the agent container's log and read the fetch failure.
- Note how the Kubernetes condition differs from the one you saw for
ScheduleFailed.
Configuration you will actually touch
Most of Seldon Core 2's configuration lives in Helm values and in two custom resources, and at beginner level you need to recognise them rather than master them.
SeldonRuntime is the resource that installs the per-namespace runtime — scheduler, Envoy, gateways, dataflow engine. The seldon-core-v2-runtime chart creates one for you. If you want Seldon serving models in a second namespace, a second SeldonRuntime is the mechanism.
SeldonConfig holds the runtime's configuration: Kafka connection details, tracing, and since 2.10 a scaling configuration block. The point of centralising it in a resource is that configuration can be changed after the cluster is deployed rather than only at install time. Two settings from 2.10 are worth knowing exist: maxShardCountMultiplier, which lets pipeline components scale horizontally beyond the limit that Kafka partition counts used to impose, and the new scaling config block under config.ScalingConfig. The upgrade notes say to set the new scaling configuration before upgrading to 2.11 — and that with Helm the 2.9-to-2.10 upgrade fills the new options from your existing values automatically, while a non-Helm install must set maxShardCountMultiplier by hand.
Operator scope is the decision you make at install time. controller.clusterwide=true watches everything; controller.watchNamespaces={ns1,ns2} restricts it. The namespaced form is the foundation of multi-tenancy, and it pairs with per-namespace SeldonRuntimes and separate Kafka credentials per team.
Kafka security is the configuration most likely to block your first real pipeline. TLS is configured by referencing a Secret under security.kafka.ssl.client.secret in the Helm values, and SASL_SSL — which is what managed Kafka services such as Confluent Cloud use — needs both the SASL mechanism and the client credentials set. The documentation has worked examples for Confluent SASL and OAuth 2.0. The failure mode is recognisable: broker timeouts and connection errors in the dataflow-engine and model-gateway logs, almost always because a Secret name is wrong or the Secret is in the wrong namespace.
Monitoring comes from Prometheus. The install flow creates a seldon-monitoring namespace for a metrics stack, and Seldon exposes metrics including request rates such as infer_rps, which is also the metric the documented HPA pattern scales on. You do not need this for your first model, but install it before you need to answer "is it slow?" rather than after.
Finally, there is a seldon command-line tool, distributed as a release binary, which talks to the scheduler directly to load, inspect and call models and pipelines. It is genuinely useful — particularly pipeline inspect, which reads the Kafka topics behind a pipeline so you can see what actually flowed through each step. On Kubernetes, though, prefer kubectl apply and let the operator reconcile: that keeps your cluster's state matching files in git, which is the entire point of a declarative system. Check the current CLI documentation for exact syntax and connection flags before scripting against it.
- Run
kubectl get seldonruntime,seldonconfig -n seldon-mesh. - Run
kubectl get seldonconfig -n seldon-mesh -o yamland skim it. - Find the Kafka section and note whether your install points at a self-hosted broker.
Putting it all together
Here is the whole beginner path as one exercise. Start from an empty cluster and end with a two-step pipeline you can call. Do it without scrolling up if you can.
# 1. Cluster
kind create cluster --name seldon
kubectl version
# 2. Namespaces and chart repository
kubectl create ns seldon-mesh || echo "Namespace seldon-mesh already exists"
kubectl create ns seldon-monitoring || echo "Namespace seldon-monitoring already exists"
helm repo add seldon-charts https://seldonio.github.io/helm-charts/
helm repo update seldon-charts
# 3. Four charts, in order
helm upgrade seldon-core-v2-crds seldon-charts/seldon-core-v2-crds \
--namespace default --install
helm upgrade seldon-core-v2-setup seldon-charts/seldon-core-v2-setup \
--namespace seldon-mesh --set controller.clusterwide=true --install
helm upgrade seldon-core-v2-runtime seldon-charts/seldon-core-v2-runtime \
--namespace seldon-mesh --install
helm upgrade seldon-core-v2-servers seldon-charts/seldon-core-v2-servers \
--namespace seldon-mesh --install
# 4. Confirm
kubectl get pods -n seldon-mesh
kubectl get servers -n seldon-mesh
Then the model, the wait, and the call:
kubectl apply -f iris.yaml -n seldon-mesh
kubectl wait --for condition=ready --timeout=300s model iris -n seldon-mesh
kubectl port-forward svc/seldon-mesh -n seldon-mesh 8080:80 &
curl -X POST http://localhost:8080/v2/models/iris/infer \
-H "Content-Type: application/json" \
-H "Seldon-Model: iris" \
-d '{"inputs":[{"name":"predict","shape":[1,4],"datatype":"FP32","data":[[1,2,3,4]]}]}'
Four checkpoints, each with a specific thing to verify. After the charts, kubectl get servers lists two Servers. After the model, kubectl get pods shows no new pod — the model is inside an existing replica. After the call, you get a tensor back. And after you deliberately drop the Seldon-Model header, you get a routing failure that you now recognise on sight.
For the pipeline half, deploy the two quickstart models, apply the Pipeline, and call it with Seldon-Model: income-classifier-app.pipeline. You will need Kafka for this, which is what the learning-environment documentation sets up.
Clean up in reverse order, and remember what the last step does:
kubectl delete -f iris.yaml -n seldon-mesh
helm uninstall seldon-core-v2-servers -n seldon-mesh
helm uninstall seldon-core-v2-runtime -n seldon-mesh
helm uninstall seldon-core-v2-setup -n seldon-mesh
helm uninstall seldon-core-v2-crds -n default
kind delete cluster --name seldon
That CRD uninstall is the one to be careful with on any cluster you did not create for practice.
- Tear your cluster down completely and rebuild it from the block above, from memory where you can.
- Time it. Note which step you had to look up.
- Write the whole thing into a shell script you keep.
What you can now do, and what comes next
You can install Seldon Core 2 on a Kubernetes cluster from four Helm charts in the right order, and explain what each chart contributes. You can deploy a trained artifact as a Model, understand that it lives inside a shared Server replica rather than a pod of its own, and know what each field of its spec means. You can call it over the Open Inference Protocol and you know why the Seldon-Model header exists. You can read a ScheduleFailed condition and name the three filters the scheduler applies. You can write your own Server with a pinned image, and you know the difference between capabilities and extraCapabilities. You can chain models into a Pipeline, and you know why calling one needs the .pipeline suffix. And you can tell Seldon Core 2 apart from Seldon Core 1 on sight, which is more than most search results manage.
What the Mid-level guide takes up next, in the order it matters:
Scaling properly. Native model autoscaling is off in 2.10, so scaling means HPA or KEDA on the Servers. The documented pattern has real constraints worth understanding before you rely on it: only custom Prometheus metrics, matched HPAs on both the Model and the Server applied together, and a one-to-one mapping between Models and Servers that gives up the multi-model-serving benefit. Also: the scheduler does not create Server replicas, so a Model that scales ahead of its Server reports ScheduleFailed while the existing replicas keep serving.
Pipelines in anger. Join semantics, tensorMap, triggers, conditional routing, cross-pipeline externalInputs and externalTriggers, and cyclic pipelines with bounded iterations. Plus the Kafka side: partitions, retention and the maxShardCountMultiplier that lets pipeline components scale past partition limits.
Experiments. Weighted traffic splits for A/B tests, and shadow traffic for validating a candidate against production requests without serving its answers.
Monitoring both kinds. Operational monitoring with Prometheus and Grafana, and data-science monitoring by deploying drift and outlier detectors — which are Models with the alibi-detect capability — inside your pipelines.
Where to go in this catalogue. KServe is the closest comparison and also speaks the Open Inference Protocol; the architectures differ sharply, with Seldon's Kafka pipelines and multi-model servers against KServe's one-InferenceService-per-model model. Triton is the other default server here and is worth knowing directly. MLflow is where the artifacts you deploy usually come from, and MLServer can load MLflow-packaged models through the mlflow capability. BentoML is an alternative packaging-and-serving route worth comparing. Kubernetes and Helm are the ground this all stands on, and everything in this guide is easier if those are solid. Prometheus and Grafana are the monitoring stack, Alibi Detect is the drift and outlier library behind the data-science monitoring story, and Evidently covers the same ground from a different angle.
One closing note on region. If you are deploying for a Gulf or Egyptian employer with data-residency obligations, Seldon Core 2's being plain Kubernetes is the point: it runs wherever your cluster runs, on-premises or in a UAE or Saudi cloud region, with no dependency on a vendor-hosted control plane. But do check your Kafka: a managed Kafka in a distant region moves every pipeline payload across a border, and with pipelines that payload is also retained on topics there. Keep the cluster and its Kafka in the same region, and decide topic retention with your compliance obligations in mind rather than as a performance setting.
Sources
- Seldon Core 2 documentation home
- Concepts
- Core features
- Architecture
- Quickstart
- Production environment installation
- Learning environment installation
- Upgrading
- Models
- Model scheduling
- Servers
- Pipelines
- Custom servers example
- Single-model-serving HPA
- Operational monitoring
- Pipeline advanced configuration
- Seldon Core releases
- Seldon Helm charts repository