This is part one of three. It assumes you have never touched Kubeflow, and it is honest about one thing up front: Kubeflow is not a single program you install and open. It is a family of projects that share a home on Kubernetes, and half of being productive with it is knowing which member of the family you actually need. By the end of this guide you will know what each part is for, you will have written and run a real machine learning pipeline, you will know how to bring up the platform on a laptop or a cluster, and you will be able to read the errors that greet almost every newcomer. Mid-level and Senior take the same ground further; nothing here is thrown away.
Each section ends with a Try it task where it helps. Do them as you go. Kubeflow concepts only stick once you have watched your own pipeline run, fail, and run again.
What Kubeflow is, and the problem it solves
Training a model in a notebook is the easy part of machine learning. The hard part starts the day someone asks you to do it again: with new data, on a schedule, on a bigger machine, with five different learning rates, and with a record of exactly which code and data produced the model now running in production. A notebook cannot answer those questions. A folder of scripts held together by memory and shell history cannot either.
Kubeflow describes itself as the cloud native AI platform: a set of open-source projects that together make the machine learning lifecycle run on Kubernetes, the system most organisations already use to run containers across a cluster of machines. If Kubernetes is new to you, read the Kubernetes guide alongside this one, because almost everything Kubeflow does is expressed as Kubernetes objects. A pipeline step becomes a pod. A training job becomes a set of pods. A notebook server becomes a pod with a disk attached.
The design idea is that each stage of the machine learning lifecycle gets its own project, and each project can be used alone:
Before platforms like this, teams glued the stages together themselves. One engineer wrote a cron job that called a training script. Another kept a spreadsheet of hyperparameter runs. A third copied model files to a server by hand. Each of these worked until the person who built it left. Kubeflow's pitch is that these chores are the same in every company, so they should be shared, open-source infrastructure rather than private folklore.
There is a cost, and you should hear it now. Kubeflow runs on Kubernetes, so you inherit Kubernetes. The full platform is heavy: the reference bundle needs roughly 4.4 CPU cores and 12 GiB of memory just to sit idle, before you run anything. If you only want to track experiments from a laptop, a lighter tool such as MLflow is the better start. If you only want a schedule of data jobs, Airflow may fit better. Kubeflow earns its weight when you already have a cluster, several people sharing it, and a need for reproducible pipelines, distributed training and isolation between teams.
What people use it for:
Repeatable pipelines
Preparing data, training, evaluating and registering a model as a graph of steps you can rerun with one click or one command.
Shared notebooks
Every data scientist gets a notebook server on the cluster, with cluster GPUs and a private disk, instead of a laptop.
Distributed training
Launching a PyTorch job across several machines without hand-writing the pod wiring.
Hyperparameter tuning
Letting an algorithm try many settings in parallel and report the best one.
- Think of the last model you trained. Write down every step from raw data to the moment someone could call it.
- Mark each step that you did by hand and could not rerun tomorrow from a single command.
- Count them.
How Kubeflow is versioned and packaged
Newcomers trip on this before they run a single command, so it comes early. Searching for "Kubeflow 1.12" leads nowhere, and old tutorials quote versions that no longer exist.
Kubeflow is a set of independent subprojects, each with its own repository, release schedule and version number. The ones you will meet in this guide are:
- Kubeflow Pipelines (KFP), for defining and running workflows.
- Kubeflow Trainer, for running training jobs, including distributed ones.
- Katib, for hyperparameter tuning.
- Kubeflow Notebooks, for notebook servers on the cluster.
- Kubeflow Dashboard, the web front door and the place where users and namespaces are managed.
- Kubeflow Hub, formerly called Model Registry, a catalogue of your models and their versions.
- Spark Operator, for running Apache Spark jobs on Kubernetes.
Serving models is handled by KServe. It is now classed as an ecosystem project rather than a Kubeflow subproject, but the standard bundle still ships it. See the KServe guide once you have a model to deploy.
On top of the subprojects sits a curated bundle called the Kubeflow Community Distribution, which lives in the repository kubeflow/community-distribution. You may see it called "the manifests", because it is a collection of Kustomize manifests (see the Kustomize guide) that install compatible versions of everything plus the supporting cast: Istio for networking, cert-manager for certificates, Dex and oauth2-proxy for login, and Knative for serverless serving. The repository used to be named kubeflow/manifests; GitHub redirects the old address.
The bundle moved to calendar versioning in 2026. A version now reads YY.MM, with patches as YY.MM.PATCH. The release this guide is written against is 26.03.1, published in June 2026, following the base 26.03 release in March. The last release under the old scheme was 1.11.0, in December 2025, and there is no 1.12. The project aims for two base releases a year, timed before the two KubeCon conferences, and the community supports each one on a best-effort basis for about six months.
Inside 26.03.1, the components have their own numbers: Pipelines 2.16.1, Trainer 2.2.0, Katib 0.19.0, Notebooks 1.11.0, Dashboard 2.0.0 and KServe 0.18.0. Newer standalone releases of several of these already exist. That is normal, and it is why you always check which version a tutorial used before copying its commands.
- Open the
kubeflow/community-distributionrepository on GitHub and find the list of tags. - Find the tag
26.03.1and read the component version table in its README.
The mental model: the parts you will meet
Kubeflow has a lot of nouns. Learn them in groups and the platform stops feeling like a wall of logos.
The platform layer. The Central Dashboard is the web page you open in a browser. It is a shell that embeds the interface of each component, so you click "Notebooks" or "Pipelines" in a side menu and the page for that component appears inside the dashboard. A Profile is Kubeflow's unit of multi-user isolation. Each person or team gets a profile, and a profile creates a Kubernetes namespace of the same name, plus the permissions that make its owner an administrator of that namespace. Your work lives in your profile's namespace, and other users cannot see it unless you add them as contributors. Authentication happens before you reach any of this: a login proxy checks who you are and Istio attaches your email address to each request in a header called kubeflow-userid, which the web applications trust.
The pipeline layer. A component is one step, packaged as a container. A pipeline is a graph of components, where the output of one step feeds the input of the next. You write both in Python using the kfp library, then compile the pipeline into a YAML file, a portable description that the backend can execute. One execution of a pipeline is a run. Runs are grouped into experiments, which is just a folder for organisation. A run that repeats on a schedule is a recurring run.
The data passed between steps. Small values such as a number or a string are parameters. Files such as a dataset or a trained model are artifacts. Artifacts are stored in an object store (an S3-style bucket) and steps receive them as files. This distinction matters a great deal, and we will return to it in the first pipeline.
The training layer. A TrainJob is a request to train a model, possibly across several machines. It refers to a runtime, which is a blueprint written by an administrator that says how to launch, for example, PyTorch distributed training. You pick the runtime and supply your code. A Katib Experiment is a tuning job: a goal, a search space of settings, an algorithm to explore it, and a budget of trials. Each trial is one try.
The notebook layer. A Notebook in Kubeflow terms is a Jupyter or similar server running as a pod, with a persistent disk so your files survive restarts.
Hold on to this stack. When something breaks, the first question is always which layer: did the request get in (front door), did the component accept the work (components), or did Kubernetes fail to start a pod (bottom)? Most beginners look at the wrong layer for hours.
One more term you will see in official pages: KFP v2. Kubeflow Pipelines had an older SDK generation, v1, and it has been removed. Everything here uses v2, and if a tutorial mentions ContainerOp or "v2 compatible mode", it is out of date.
What you need before you install
Installing the whole platform is a cluster project, so decide what you want before you start. There are three realistic starting points, ordered from lightest to heaviest.
No cluster at all. The Pipelines SDK can run pipelines locally, on your own computer, using a local runner. You need Python and, optionally, Docker. This is where you should begin, and this guide's first project uses it. You learn the pipeline language without touching Kubernetes.
One component on a small cluster. You can install Kubeflow Pipelines on its own, without the dashboard, profiles or login. It needs a Kubernetes cluster and kubectl. A local cluster made with kind (Kubernetes in Docker) is enough. This is the right step after local runs, because you see real pods without the multi-user machinery.
The full distribution. The community distribution recommends 16 GB of memory and 8 CPU cores. You can trim the install file to fit into roughly 4 to 8 GB and 2 to 4 cores by removing components you do not need, but expect a slow, tight machine. The platform expects Kubernetes 1.34 or newer, kubectl, and Kustomize 5.8 or close to it. Kind needs Docker or Podman. You also need a cluster storage class that can create disks on demand, because Pipelines, Katib and Hub each request persistent volumes.
A few notes by operating system. The official documentation describes one path, kind on Linux, and does not have separate macOS or Windows pages. On a Mac, the practical route is to run Docker Desktop, Colima or Podman Machine, give that virtual machine at least 16 GB of memory and 8 CPUs, and run the same scripts. This is inferred from how kind works, not spelled out in the docs, so treat it as a best effort. On Apple Silicon, expect some images to fail with "no matching manifest for linux/arm64", because not every image is built for ARM. On Windows, use WSL2 with an Ubuntu shell and run everything inside it; raise the memory limit for WSL in your .wslconfig file. There is no native PowerShell install path.
On Linux, kind clusters that run many pods hit a low file-watcher limit. Raise it before installing:
sudo sysctl fs.inotify.max_user_instances=2280
sudo sysctl fs.inotify.max_user_watches=1255360
If you skip this, pods crash at random with errors mentioning "too many open files". It is one of the most common first-install failures, and the fix takes ten seconds.
If you work in the Gulf or Egypt, you may also need to think about where the cluster lives. A managed Kubernetes service in a Middle East cloud region keeps training data in-country, which some employers and regulators require. Kubeflow itself has no opinion about the region; it runs wherever the cluster runs. Only your object store and your data sources carry the residency question.
- Check your tools:
python3 --version,docker --version,kubectl version --client. - Decide which of the three starting points your machine supports. Be honest about the memory.
Installing the pieces
The SDK: start here
Install the Pipelines SDK into a virtual environment. Pin the version, because the SDK and the backend must agree:
python3 -m venv .venv
source .venv/bin/activate
pip install kfp==2.17.0
python -c "import kfp; print(kfp.__version__)"
kfp --version
The second-to-last command should print 2.17.0. The kfp package also installs a command-line tool, also named kfp. If you want to build your own component images later with kfp component build, install kfp[all], which adds the Docker client library. You can work through the local-run section below with just this.
Pipelines on its own, on a kind cluster
Create a cluster, then install the standalone Pipelines release. The official steps use a version variable and two applies, with a wait between them so the custom resource definitions are registered before anything uses them:
kind create cluster
export PIPELINE_VERSION=2.17.0
kubectl apply -k "github.com/kubeflow/pipelines/manifests/kustomize/cluster-scoped-resources?ref=$PIPELINE_VERSION"
kubectl wait --for condition=established --timeout=60s crd/applications.app.k8s.io
kubectl apply -k "github.com/kubeflow/pipelines/manifests/kustomize/env/dev?ref=$PIPELINE_VERSION"
kubectl get pods -n kubeflow
kubectl port-forward -n kubeflow svc/ml-pipeline-ui 8080:80
Give it a few minutes; the pods in the kubeflow namespace move from ContainerCreating to Running. When the port-forward is running, open http://localhost:8080 and you see the Pipelines interface. The dev overlay uses a PostgreSQL database that is meant for trying things out. A production install uses the platform-agnostic overlay with MySQL, which is a topic for the Mid-level guide.
The full distribution
The complete platform installs from the community distribution repository. Check out the tag, create a kind cluster with the helper script, and apply everything in a retry loop:
git clone https://github.com/kubeflow/community-distribution.git
cd community-distribution
git checkout 26.03.1
./tests/install_KinD_create_KinD_cluster_install_kustomize.sh
kind get kubeconfig --name kubeflow > /tmp/kubeflow-config
export KUBECONFIG=/tmp/kubeflow-config
while ! kustomize build example | kubectl apply --server-side --force-conflicts -f -; do echo "Retrying to apply resources"; sleep 20; done
The loop is not a hack; it is the documented approach, and it is worth understanding. Kubernetes has two kinds of things in the install: definitions of new object types (CRDs), and objects of those new types. If an object arrives before its type is registered, kubectl fails with no matches for kind ... in version ... or ensure CRDs are installed first. The retry loop simply tries again after 20 seconds, by which time the definitions exist. Seeing those errors in the first few rounds is expected.
If Docker Hub rate limits are a concern, log in and create a pull secret named regcred as shown in the repository's README before you start the loop. The community distribution also documents installing component by component, which is useful when a single piece fails and you want to isolate it.
Checking that it works
Wait for the namespaces to fill with running pods, then look at each in turn:
kubectl get pods -n cert-manager
kubectl get pods -n istio-system
kubectl get pods -n auth
kubectl get pods -n oauth2-proxy
kubectl get pods -n knative-serving
kubectl get pods -n kubeflow
kubectl get pods -n kubeflow-system
kubectl get pods -n kubeflow-user-example-com
Every pod should show Running or Completed. Then open the front door with a port-forward:
kubectl port-forward svc/istio-ingressgateway -n istio-system 8080:80
Browse to http://localhost:8080. You land on the Dex login page. The development account is user@example.com with the password 12341234, and its profile namespace is kubeflow-user-example-com.
Also note that plain HTTP works on localhost only. On any other address, the web applications set secure cookies, so you need HTTPS in front of the gateway. Otherwise logins loop or fail with cookie errors.
- Install the SDK and run
kfp --version. - If your machine has the memory, create a kind cluster and install standalone Pipelines. Open the interface through the port-forward.
Your first pipeline, built step by step
We will build a tiny pipeline that mimics a real one: make some data, train a trivial "model", and report a score. The content is deliberately silly because the point is the shape. Start on your laptop with no cluster.
Step one: a component
A component starts as an ordinary Python function with a decorator. The type hints are required, because Kubeflow uses them to work out what goes in and out:
from kfp import dsl, compiler
@dsl.component(base_image="python:3.11")
def make_numbers(count: int) -> list:
return [float(i) for i in range(count)]
@dsl.component(base_image="python:3.11")
def average(numbers: list) -> float:
return sum(numbers) / len(numbers)
Two things deserve a pause. First, @dsl.component turns the function into something that will run in its own container, using the base image you name. The function body is copied into that container and executed there, so the function must be self-contained: any import it needs goes inside the body, and it cannot see variables from the rest of your file. Second, the return type is the contract. int, float, str, bool, list and dict are parameters that pass by value between steps.
If a step needs a library that is not in the base image, list it, and Kubeflow installs it when the step starts:
@dsl.component(base_image="python:3.11", packages_to_install=["pandas"])
def summarise(path: str) -> str:
import pandas as pd
return str(pd.read_csv(path).describe())
Notice the import inside the function. That is not style; it is required.
Step two: a pipeline
A pipeline is a function too, decorated with @dsl.pipeline. Its body wires components together. This is where beginners are caught out, so read carefully: the body of a pipeline function does not run as ordinary Python. It runs once, at compile time, only to record which steps exist and which outputs feed which inputs. Do not put if statements on values, loops over data, or print calls there and expect them to run at training time. Everything real happens inside components.
@dsl.pipeline(name="average-pipeline", description="Make numbers and average them")
def average_pipeline(count: int = 10) -> float:
numbers = make_numbers(count=count)
result = average(numbers=numbers.output)
return result.output
Calling a component inside the pipeline returns a task. The task's .output is a placeholder that means "whatever this step returns, once it has run". Passing numbers.output into average creates the dependency: average will not start until make_numbers has finished. Steps that do not depend on each other can run at the same time.
Step three: compile
Compiling turns the Python into a YAML file:
compiler.Compiler().compile(average_pipeline, package_path="average_pipeline.yaml")
Open the YAML. You will see the pipeline's parameters, each component's container image and command, and the graph of dependencies. This file is the real product. It is portable, it can be uploaded to any Kubeflow Pipelines backend, and you can commit it to git and diff it in review. The Python is only a convenient way to write it.
Step four: run it locally
Local execution needs no cluster. You initialise a local runner, then call the pipeline like a function:
from kfp import local
local.init(runner=local.SubprocessRunner())
task = average_pipeline(count=5)
print(task.output)
SubprocessRunner runs each step as a subprocess in a virtual environment on your machine, which is quick and needs no Docker. If you want each step in a container just as on a cluster, use local.DockerRunner() instead; it is slower but closer to the real thing. Expect the first run to pause while the runner creates environments. The printed output is 2.0, the average of 0 to 4.
Local mode is the fastest feedback loop you will have. Use it to debug logic before you pay for cluster time.
Step five: run it on a cluster
With standalone Pipelines port-forwarded to port 8080, submit the same pipeline to the cluster:
from kfp.client import Client
from first_pipeline import average_pipeline
client = Client(host="http://localhost:8080")
run = client.create_run_from_pipeline_func(
average_pipeline,
arguments={"count": 20},
experiment_name="first-experiments",
)
client.wait_for_run_completion(run.run_id, timeout=3600)
print("done", run.run_id)
This compiles the pipeline, creates the experiment if needed, starts a run, and waits. In the browser, open the run and you see the graph with each step coloured by state. Click a step to see its logs, its inputs and outputs, and the pod that ran it.
On the full multi-user distribution, the client needs to prove who you are. The simplest place to do this is from a notebook inside the cluster, which is covered in the notebooks section.
kfp dsl compile --py first_pipeline.py --output average_pipeline.yaml, and add --function average_pipeline if the file contains more than one pipeline.- Type in the two components and the pipeline, then compile to YAML and open the file.
- Run it locally with
SubprocessRunnerand confirm it prints2.0for a count of 5. - Add a third component that returns the largest number and wire it in.
Passing files between steps: artifacts
Parameters are fine for a number. A trained model or a table of a million rows is not something you pass by value. For files, Kubeflow uses artifacts, which are stored in the object store and handed to each step as a path.
The type hints change. A step that writes a file declares an Output, and a step that reads it declares an Input:
from kfp import dsl
from kfp.dsl import Dataset, Model, Input, Output, Metrics
@dsl.component(base_image="python:3.11")
def prepare(out_data: Output[Dataset]):
with open(out_data.path, "w") as f:
f.write("x,y\n1,2\n2,4\n3,6\n")
@dsl.component(base_image="python:3.11")
def train(data: Input[Dataset], model: Output[Model], metrics: Output[Metrics]):
rows = open(data.path).read().strip().split("\n")[1:]
slope = sum(float(r.split(",")[1]) / float(r.split(",")[0]) for r in rows) / len(rows)
with open(model.path, "w") as f:
f.write(str(slope))
metrics.log_metric("slope", slope)
@dsl.pipeline(name="artifact-pipeline")
def artifact_pipeline():
d = prepare()
train(data=d.outputs["out_data"])
Read the function signatures closely. out_data: Output[Dataset] is not a value you return; it is a handle. You write your file to out_data.path, and Kubeflow uploads it to the object store afterwards. In the next step, data: Input[Dataset] gives you a local path to the downloaded copy. You never write an S3 client yourself.
The artifact types carry meaning: Dataset, Model and Metrics are the common ones, and there are others for classification metrics, HTML and Markdown. The interface uses the type to show them properly; Metrics values appear in the run's metrics view, and you can compare them across runs.
Notice that when a component has an Output artifact, the pipeline refers to it by name through d.outputs["out_data"]. When a component has a single plain return value, you use .output. That split causes many beginner errors: if Kubeflow complains that a task has more than one output, use the named form.
Where do the files go? The object store named in the pipeline root. On the standard install this is a store reached through a service still called minio-service, although the default store is now SeaweedFS, a replacement for MinIO that speaks the same S3 protocol. You will see minio://mlpipeline/v2/artifacts as the default root. On a cloud cluster you point it at an S3 or GCS bucket instead, which is a configuration topic for later.
list parameter ends up inside Kubernetes objects, which have size limits, and runs fail in strange ways. Anything large belongs in an artifact.- Run the artifact pipeline locally with
SubprocessRunner. - Change the data in
prepareand confirm the logged slope changes.
Runs, caching and reading the interface
Once a pipeline is on a cluster, you spend most of your time in a few screens. Learn what each is telling you.
The pipelines list holds uploaded pipeline definitions. A pipeline can have several versions, each an uploaded YAML. The experiments list groups runs. The runs list shows every execution with its status, duration and the experiment it belongs to. Open a run and the graph view draws your steps as boxes, coloured for pending, running, succeeded, failed or cached. Clicking a box opens a side panel with tabs for input and output, visualisations, details, logs and the events from Kubernetes.
When a step fails, go to the logs tab first. If there are no logs, the step never started, so switch to the events view. A failure before logs exist is almost always a Kubernetes problem rather than a Python problem: the image could not be pulled, the pod could not be scheduled for lack of memory, or a volume could not mount. A failure with logs is usually your own code, and the traceback is right there.
Caching is on by default and is the feature that surprises people most. If a step's definition and inputs are identical to a previous successful run, Kubeflow skips it and reuses the earlier outputs, marking the step as cached. For expensive steps this saves hours. For a step that reads the current time or pulls fresh data from outside, it silently gives you stale answers. You have two controls. Per run, pass enable_caching=False to create_run_from_pipeline_func. Per step, set task.set_caching_options(False) inside the pipeline:
@dsl.pipeline(name="no-cache-demo")
def no_cache_demo():
t = make_numbers(count=5)
t.set_caching_options(False)
You can also set a cache-off default for a whole compile by passing --disable-execution-caching-by-default to kfp dsl compile.
A recurring run repeats a pipeline on a schedule. It is created from the interface, from client.create_recurring_run, or from kfp recurring-run create. Older tutorials call these "jobs"; the SDK methods with "job" in their names are deprecated in favour of the recurring-run names and still print a deprecation warning.
Under the hood, the backend converts your compiled pipeline into an Argo Workflow and runs it. You do not need Argo knowledge to be productive, but it explains two things you will see in pod lists: each step has a pod, and the pod names look like generated workflow names. When you search with kubectl get pods -n <your namespace> during a run, those are your steps.
- Run the same pipeline twice with identical arguments on a cluster.
- Compare the second run's graph with the first.
- Run it again with
enable_caching=False.
Profiles, the dashboard and notebooks
If you installed the full distribution, you work through the Central Dashboard, and the ideas here apply.
After logging in, the top bar shows a namespace selector. That is your profile. Everything you create, such as notebooks, runs and experiments, belongs to the namespace selected. Switching it changes which resources you see. A frequent beginner complaint is "my run disappeared", which almost always means the selector is on a different namespace.
To create a profile for a new user, an administrator applies a small YAML to the cluster:
apiVersion: kubeflow.org/v1
kind: Profile
metadata:
name: aya-team
spec:
owner:
kind: User
name: aya@example.com
kubectl apply -f profile.yaml
kubectl get profiles
The Profile controller reacts by creating the namespace aya-team, a role binding that makes the owner an administrator there, and two service accounts named default-editor and default-viewer. Self-service profile creation through the dashboard is off by default and turned on with the setting CD_REGISTRATION_FLOW=true on the dashboard deployment.
kubectl delete profile aya-team removes the namespace with all notebooks, volumes and runs. Never delete the Profiles CRD itself: that removes every profile and every user namespace on the cluster.Also know the limits of this isolation. The documentation states plainly that profiles give no stronger separation than Kubernetes namespaces do. They keep honest colleagues out of each other's way. They are not a hard security boundary between mutually distrustful tenants.
Notebooks
Open Notebooks in the dashboard and choose New Notebook. You pick a name, an image (Jupyter with various libraries is offered, including GPU builds), CPU and memory, and a workspace volume. The workspace volume is a persistent disk mounted into the pod, and it is the only place files survive a restart. Anything you save elsewhere vanishes when the notebook stops.
Behind the page, a Notebook object is created, and the notebook controller turns it into a pod named <name>-0. If a notebook never becomes ready, inspect it the Kubernetes way:
kubectl get notebooks -n aya-team
kubectl describe notebook my-notebook -n aya-team
kubectl logs my-notebook-0 -n aya-team
The usual reasons are an image that cannot be pulled, a disk that cannot be created, or a namespace that has hit its resource quota. Culling, meaning automatically stopping idle notebooks, is off by default. An administrator enables it with ENABLE_CULLING on the notebook controller, and the idle time defaults to 1440 minutes, one day.
Talking to Pipelines from a notebook
The easiest authenticated client is one created inside a notebook. The distribution defines a PodDefault named access-ml-pipeline: a rule that, for notebooks carrying a matching label, mounts a short-lived token for the Pipelines API. When you create the notebook, tick the configuration called "Allow access to Kubeflow Pipelines" in the form. Then from that notebook:
import kfp
client = kfp.Client()
print(client.list_experiments())
Without the PodDefault you get Unauthenticated: Request header error: there is no user identity header, which means the request reached the API without a login or a token.
- Log in to the dashboard, open Notebooks and create a small notebook with a 5 GB workspace.
- In a terminal inside it, create a file under your home folder, stop the notebook and start it again.
Training jobs and tuning
Not every job is a pipeline step. When you need to train across several machines, or try a hundred learning rates, Kubeflow has purpose-built tools.
Kubeflow Trainer
Trainer is the current way to run training jobs on Kubernetes. It replaces an older project called the Training Operator, which defined objects like PyTorchJob and TFJob. Those old objects still appear in tutorials, and the old operator has been removed from the development branch of the distribution. Write new work for Trainer.
A TrainJob is the object you create. It names a runtime, which an administrator has already installed. A runtime is a template for a framework, such as torch-distributed, so you do not have to know how to wire processes together. The shipped runtimes cover PyTorch, DeepSpeed, JAX, MLX and XGBoost, plus ready-made fine-tuning recipes for a few small language models.
Install the Trainer and its runtimes on a cluster with Kubernetes 1.31 or newer using Helm (see the Helm guide):
helm install kubeflow-trainer oci://ghcr.io/kubeflow/charts/kubeflow-trainer \
--namespace kubeflow-system --create-namespace \
--version 2.3.0 --set runtimes.defaultEnabled=true
The full distribution already includes Trainer, in the kubeflow-system namespace. The easiest way to use it is the Kubeflow SDK, a separate Python package that wraps several components behind one client:
pip install -U kubeflow
from kubeflow.trainer import TrainerClient, CustomTrainer
def train_fn():
import torch
print("hello from", torch.__version__)
client = TrainerClient()
for runtime in client.list_runtimes():
print(runtime.name)
job = client.train(
trainer=CustomTrainer(func=train_fn, num_nodes=2,
resources_per_node={"cpu": 3, "memory": "16Gi"}),
runtime=client.get_runtime("torch-distributed"),
)
client.wait_for_job_status(job)
for line in client.get_job_logs(job, follow=True):
print(line)
You hand Trainer a plain Python function, the number of machines, and the resources each one gets. Trainer packages the function, starts the pods and sets up the environment variables PyTorch needs to find its peers. client.list_runtimes() is how you discover what your administrator installed. If it returns nothing, the runtimes are missing, not your code.
Under the hood, the Kubernetes objects are TrainJob and a JobSet, and you can inspect them:
kubectl get trainjobs -n aya-team
kubectl describe trainjob <name> -n aya-team
kubectl get jobsets -n aya-team
The same job can be written as YAML and applied with kubectl, which suits GitOps workflows. Its API group is trainer.kubeflow.org/v1alpha1. The v1alpha1 label tells you the API is still evolving, and Trainer 2.2 already contained breaking changes, so pin versions and read release notes.
Katib for hyperparameter tuning
A hyperparameter is a setting you choose before training, such as a learning rate, that the training does not learn by itself. Finding a good one by hand is slow. Katib automates it: you define an Experiment with an objective (maximise accuracy, say), a search space, an algorithm such as random search, grid search, Bayesian optimisation or Hyperband, and a limit on trials. Katib's suggestion service proposes settings, launches a trial for each, collects the result and feeds it back.
The key detail for newcomers is how Katib learns the result. By default it reads your program's standard output and looks for lines of the form name=value. If the objective metric is called accuracy, your code must print exactly accuracy=0.93. A trial that never prints in that format ends as MetricsUnavailable.
import kubeflow.katib as katib
def objective(parameters):
result = 4 * int(parameters["a"]) - float(parameters["b"]) ** 2
print(f"result={result}")
client = katib.KatibClient(namespace="aya-team")
client.tune(
name="tune-experiment",
objective=objective,
parameters={
"a": katib.search.int(min=10, max=20),
"b": katib.search.double(min=0.1, max=0.2),
},
objective_metric_name="result",
max_trial_count=12,
resources_per_trial={"cpu": "2"},
)
client.wait_for_experiment_condition(name="tune-experiment")
print(client.get_optimal_hyperparameters("tune-experiment"))
The kubeflow-katib package provides KatibClient and installs with pip install -U kubeflow-katib. Katib's API group is kubeflow.org/v1beta1. Every trial runs as pods in your namespace, so your resource quota matters; twelve trials at two CPUs each can starve a small cluster if they run in parallel.
- With Trainer installed, run
client.list_runtimes()and note the names. - Run the Katib example with
max_trial_count=6. Open the Katib page in the dashboard and watch the trials appear.
Everyday commands, grouped by what you want to do
Keep this section open in a tab. Every command below is covered by the official CLI or client.
Write and compile.
kfp dsl compile --py pipeline.py --output pipeline.yaml
kfp dsl compile --py pipeline.py --output pipeline.yaml --pipeline-parameters '{"count": 10}'
Upload and manage pipelines. These need --endpoint pointing at your API, or the environment set up for it:
kfp --endpoint http://localhost:8080 pipeline create -p my-pipeline pipeline.yaml
kfp --endpoint http://localhost:8080 pipeline list
kfp --endpoint http://localhost:8080 pipeline create-version -p PIPELINE_ID -v v2 pipeline.yaml
kfp --endpoint http://localhost:8080 pipeline delete PIPELINE_ID
Run and watch.
kfp --endpoint http://localhost:8080 experiment create first-experiments
kfp --endpoint http://localhost:8080 run create -e first-experiments -r run-one -f pipeline.yaml -w count=10
kfp --endpoint http://localhost:8080 run list -e EXPERIMENT_ID
kfp --endpoint http://localhost:8080 run get RUN_ID -w
The -w flag makes the command wait until the run finishes; key=value pairs at the end set the pipeline's parameters.
Schedule. kfp recurring-run create, list, get, enable, disable and delete manage scheduled runs.
Inspect Kubernetes objects.
kubectl get pods -n aya-team
kubectl describe pod POD_NAME -n aya-team
kubectl logs POD_NAME -n aya-team
kubectl get profiles
kubectl get notebooks -n aya-team
kubectl get trainjobs -n aya-team
Python client. The same operations exist on kfp.Client: create_run_from_pipeline_func, create_run_from_pipeline_package, list_runs, get_run, wait_for_run_completion, terminate_run, upload_pipeline, list_pipelines, create_experiment and list_experiments. Two renamings to remember: create_recurring_run replaces the old "job" functions, and GPU requests use set_accelerator_type and set_accelerator_limit rather than the deprecated set_gpu_limit.
Ask for resources per step. This is how a step gets memory and a GPU:
train_task = train(data=d.outputs["out_data"])
train_task.set_cpu_limit("4")
train_task.set_memory_limit("8G")
train_task.set_accelerator_type("nvidia.com/gpu")
train_task.set_accelerator_limit(1)
A step that asks for more than any node can give stays Pending forever, with no error in the logs; the Kubernetes events say Insufficient cpu or Insufficient memory.
- Compile your first pipeline with the CLI and upload it with
kfp pipeline create. - Start a run from the CLI with
-wand a parameter override.
Configuration and the most common errors
The few settings worth knowing now
You will not configure much at this level, but four settings explain a lot of behaviour.
The pipeline root says where artifacts are stored. Each profile namespace has a ConfigMap named kfp-launcher with a key defaultPipelineRoot. On the default install it points at the bundled object store. To use a cloud bucket, an administrator changes this value and supplies credentials. You can also override it for a single pipeline with pipeline_root when submitting.
The default base image for components is a recent Python 3.11 image. Naming base_image explicitly, as we did, makes your pipeline independent of that default.
Namespace for the client. On a multi-user install, runs belong to a namespace. Pass namespace="aya-team" when you create runs, or call client.set_user_namespace("aya-team") once; it is remembered in $HOME/.config/kfp/context.json.
Culling and registration for notebooks and profiles, as described above, are switched with environment variables on the respective deployments.
Reading the errors
Kubeflow errors usually come from one of the four layers in the stack diagram. Match the message to the layer and the fix becomes obvious.
| What you see | What it means | What to do |
|---|---|---|
no matches for kind ... ensure CRDs are installed first | Objects applied before their type definitions were registered | Expected on first install. Re-run the apply, or use the retry loop. |
| Pods crash, "too many open files" | Low file-watch limits on the host running kind | Raise the two inotify sysctl values shown earlier. |
toomanyrequests: You have reached your pull rate limit | Docker Hub's anonymous limit | Log in and create the regcred pull secret, or mirror images. |
no matching manifest for linux/arm64 | An image is not built for ARM | Use amd64 nodes or emulation. |
RBAC: access denied in the dashboard | An Istio authorisation policy rejected you | Check you are using the right profile and that the user header is reaching the service. |
Unauthenticated: Request header error: there is no user identity header | API called without a login or token | Go through the gateway with a cookie or token, or use the in-notebook client. |
PermissionDenied: User ... is not authorized | No role in the target namespace | Add the user as a contributor or pass the correct namespace. |
Step stays Pending | Kubernetes cannot place the pod | Read events with kubectl describe pod; lower the request or add capacity. |
Katib trial MetricsUnavailable | Metric not printed as name=value | Print exactly the objective name, an equals sign and a number. |
DeprecationWarning: 'set_gpu_limit' is deprecated | Old SDK call | Rename to set_accelerator_limit. |
Another message deserves its own paragraph: "output_metadata.json": proto: ... unknown field "custom_path". It appears when a pipeline compiled with SDK 2.15.0 or 2.15.1 runs against an older backend. The fix is to upgrade the SDK to 2.15.2 or later, or to upgrade the backend, which is why the SDK and backend versions should stay matched.
When you are lost, use a fixed order. First, is the step in the run graph failing with logs? Read the traceback. Without logs, use kubectl describe pod on the step's pod and read the events at the bottom. If the page itself will not load, check the ingress and login pods in istio-system, oauth2-proxy and auth. That three-step habit solves most problems before any search engine is needed.
- Ask a step for 1000 CPUs with
set_cpu_limit("1000")and run it. - Find the pod with
kubectl get podsand runkubectl describe podon it.
Putting it all together
Here is one small, complete project that uses what you have learned. It is a pipeline that prepares data, trains a toy model with a configurable setting, evaluates it, and only reports success if the score clears a threshold. The project has three files.
from kfp import dsl, compiler
from kfp.dsl import Dataset, Model, Input, Output, Metrics
@dsl.component(base_image="python:3.11")
def prepare(rows: int, out_data: Output[Dataset]):
with open(out_data.path, "w") as f:
f.write("x,y\n")
for i in range(1, rows + 1):
f.write(f"{i},{2 * i}\n")
@dsl.component(base_image="python:3.11")
def train(data: Input[Dataset], scale: float, model: Output[Model]):
lines = open(data.path).read().strip().split("\n")[1:]
ratios = [float(l.split(",")[1]) / float(l.split(",")[0]) for l in lines]
with open(model.path, "w") as f:
f.write(str(scale * sum(ratios) / len(ratios)))
@dsl.component(base_image="python:3.11")
def evaluate(model: Input[Model], metrics: Output[Metrics]) -> float:
slope = float(open(model.path).read())
error = abs(slope - 2.0)
metrics.log_metric("error", error)
return error
@dsl.pipeline(name="toy-training", description="prepare, train, evaluate")
def toy_training(rows: int = 50, scale: float = 1.0):
d = prepare(rows=rows)
t = train(data=d.outputs["out_data"], scale=scale)
evaluate(model=t.outputs["model"])
if __name__ == "__main__":
compiler.Compiler().compile(toy_training, package_path="toy_training.yaml")
from kfp import local
from pipeline import toy_training
local.init(runner=local.SubprocessRunner())
toy_training(rows=20, scale=1.0)
from kfp.client import Client
from pipeline import toy_training
client = Client(host="http://localhost:8080")
for scale in [0.5, 1.0, 1.5]:
client.create_run_from_pipeline_func(
toy_training,
arguments={"rows": 100, "scale": scale},
experiment_name="toy-scale-sweep",
)
The workflow is the one you would use on real work. You iterate with run_local.py until the logic is right, which costs nothing. You then compile and commit toy_training.yaml so the exact definition is reviewed in git. You submit three runs on the cluster, one per scale, all inside one experiment. Open that experiment in the interface, select the runs and compare the error metric: the run with scale=1.0 has the lowest error, which you can see at a glance because the metrics are logged against each run.
Two habits make this scale up. Every step names its base image, so it runs the same tomorrow. Every large output is an artifact, so nothing depends on memory between steps. Replace the toy steps with your real data loading and training code, keep the same skeleton, and you have a real pipeline.
When you want to tune scale automatically instead of by hand, that is Katib's job. When you want to run the training step across several GPUs, that is Trainer's. When the model is good enough to serve, continue to the KServe guide. If you want to log parameters and metrics across many experiments with a lighter tool alongside this, see the MLflow guide.
- Run
run_local.pyand read the logged error. - Compile the pipeline and submit the three-scale sweep on your cluster.
- Compare the three runs in the interface.
What you can now do, and what comes next
You can explain what Kubeflow is and why it is a family of projects rather than one program. You know that version numbers now follow the calendar, that the current bundle is 26.03.1, and that pipelines, Trainer, Katib, Notebooks, Hub and KServe each have a separate job. You can name the layers a request passes through and use that to locate a failure. You have installed the SDK, written components and a pipeline, compiled them to YAML, run them locally and on a cluster, passed files between steps as artifacts, used caching on purpose, and read a run in the interface. You can create a profile, start a notebook that survives restarts, launch a training job with the Kubeflow SDK and tune a function with Katib. And you can read the ten most common error messages.
What to practise next: write one pipeline for your own project rather than a toy, even if it has only three steps. Run it twice and watch caching. Break it deliberately and read the failure. Those three acts teach more than another hour of reading.
The Mid-level guide picks up with control flow (conditionals, loops, parallel fan-out and exit handlers), resource and secret management with kfp-kubernetes, how the backend works internally, production object stores and databases, and connecting your cluster to CI and GitOps. The Senior guide covers running Kubeflow as a platform for other teams: multi-tenancy limits, upgrades, security and capacity. Neighbouring guides in this catalogue worth reading alongside: Kubernetes for the substrate, Docker for the component images, KServe for serving, and MLflow for experiment tracking.
One last piece of advice for learning on your own. Keep a notebook of the failures you meet, with the message, the layer it came from and the fix. Kubeflow is a platform where the same dozen problems recur in different disguises: a missing permission, a pod that cannot start, a stale cache, a mismatched version. After a month, your own list will be a better reference than any guide, including this one, and it is the thing interviewers respond to most when you describe real debugging stories from your projects.
Sources
- Kubeflow documentation: https://www.kubeflow.org/docs/
- Kubeflow Community Distribution (manifests and install): https://github.com/kubeflow/community-distribution
- Kubeflow Pipelines documentation: https://www.kubeflow.org/docs/components/pipelines/
- Kubeflow Pipelines local execution: https://www.kubeflow.org/docs/components/pipelines/user-guides/core-functions/execute-kfp-pipelines-locally/
- Kubeflow Trainer documentation: https://trainer.kubeflow.org/
- Kubeflow SDK documentation: https://sdk.kubeflow.org/
- Katib documentation: https://www.kubeflow.org/docs/components/katib/
- Kubeflow Notebooks documentation: https://www.kubeflow.org/docs/components/notebooks/
- Kubeflow Spark Operator: https://spark.kubeflow.org/