تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
KubeflowMLOpsPipelines & orchestration3 مستويات106 قسمًايغطّي Kubeflow 26.03دليل بالإنجليزية

The Complete Kubeflow Guide

Run ML pipelines, notebooks and training jobs on Kubernetes with Kubeflow. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
15sections
31examples

This is part one of three. It assumes you have never touched Kubeflow, and it is honest about one thing up front: Kubeflow is not a single program you install and open. It is a family of projects that share a home on Kubernetes, and half of being productive with it is knowing which member of the family you actually need. By the end of this guide you will know what each part is for, you will have written and run a real machine learning pipeline, you will know how to bring up the platform on a laptop or a cluster, and you will be able to read the errors that greet almost every newcomer. Mid-level and Senior take the same ground further; nothing here is thrown away.

Each section ends with a Try it task where it helps. Do them as you go. Kubeflow concepts only stick once you have watched your own pipeline run, fail, and run again.

What Kubeflow is, and the problem it solves

Training a model in a notebook is the easy part of machine learning. The hard part starts the day someone asks you to do it again: with new data, on a schedule, on a bigger machine, with five different learning rates, and with a record of exactly which code and data produced the model now running in production. A notebook cannot answer those questions. A folder of scripts held together by memory and shell history cannot either.

Kubeflow describes itself as the cloud native AI platform: a set of open-source projects that together make the machine learning lifecycle run on Kubernetes, the system most organisations already use to run containers across a cluster of machines. If Kubernetes is new to you, read the Kubernetes guide alongside this one, because almost everything Kubeflow does is expressed as Kubernetes objects. A pipeline step becomes a pod. A training job becomes a set of pods. A notebook server becomes a pod with a disk attached.

The design idea is that each stage of the machine learning lifecycle gets its own project, and each project can be used alone:

NOTEBOOKSexplore
→
PIPELINESautomate
→
TRAINER + KATIBtrain and tune
→
HUBregister
→
KSERVEserve

Before platforms like this, teams glued the stages together themselves. One engineer wrote a cron job that called a training script. Another kept a spreadsheet of hyperparameter runs. A third copied model files to a server by hand. Each of these worked until the person who built it left. Kubeflow's pitch is that these chores are the same in every company, so they should be shared, open-source infrastructure rather than private folklore.

There is a cost, and you should hear it now. Kubeflow runs on Kubernetes, so you inherit Kubernetes. The full platform is heavy: the reference bundle needs roughly 4.4 CPU cores and 12 GiB of memory just to sit idle, before you run anything. If you only want to track experiments from a laptop, a lighter tool such as MLflow is the better start. If you only want a schedule of data jobs, Airflow may fit better. Kubeflow earns its weight when you already have a cluster, several people sharing it, and a need for reproducible pipelines, distributed training and isolation between teams.

What people use it for:

🔁

Repeatable pipelines

Preparing data, training, evaluating and registering a model as a graph of steps you can rerun with one click or one command.

🖥️

Shared notebooks

Every data scientist gets a notebook server on the cluster, with cluster GPUs and a private disk, instead of a laptop.

⚡

Distributed training

Launching a PyTorch job across several machines without hand-writing the pod wiring.

🎯

Hyperparameter tuning

Letting an algorithm try many settings in parallel and report the best one.

Try it
  1. Think of the last model you trained. Write down every step from raw data to the moment someone could call it.
  2. Mark each step that you did by hand and could not rerun tomorrow from a single command.
  3. Count them.
a list where most steps are manual. Each manual step is a place where Kubeflow pipelines can replace memory with a file you commit.

How Kubeflow is versioned and packaged

Newcomers trip on this before they run a single command, so it comes early. Searching for "Kubeflow 1.12" leads nowhere, and old tutorials quote versions that no longer exist.

Kubeflow is a set of independent subprojects, each with its own repository, release schedule and version number. The ones you will meet in this guide are:

  • Kubeflow Pipelines (KFP), for defining and running workflows.
  • Kubeflow Trainer, for running training jobs, including distributed ones.
  • Katib, for hyperparameter tuning.
  • Kubeflow Notebooks, for notebook servers on the cluster.
  • Kubeflow Dashboard, the web front door and the place where users and namespaces are managed.
  • Kubeflow Hub, formerly called Model Registry, a catalogue of your models and their versions.
  • Spark Operator, for running Apache Spark jobs on Kubernetes.

Serving models is handled by KServe. It is now classed as an ecosystem project rather than a Kubeflow subproject, but the standard bundle still ships it. See the KServe guide once you have a model to deploy.

On top of the subprojects sits a curated bundle called the Kubeflow Community Distribution, which lives in the repository kubeflow/community-distribution. You may see it called "the manifests", because it is a collection of Kustomize manifests (see the Kustomize guide) that install compatible versions of everything plus the supporting cast: Istio for networking, cert-manager for certificates, Dex and oauth2-proxy for login, and Knative for serverless serving. The repository used to be named kubeflow/manifests; GitHub redirects the old address.

The bundle moved to calendar versioning in 2026. A version now reads YY.MM, with patches as YY.MM.PATCH. The release this guide is written against is 26.03.1, published in June 2026, following the base 26.03 release in March. The last release under the old scheme was 1.11.0, in December 2025, and there is no 1.12. The project aims for two base releases a year, timed before the two KubeCon conferences, and the community supports each one on a best-effort basis for about six months.

Inside 26.03.1, the components have their own numbers: Pipelines 2.16.1, Trainer 2.2.0, Katib 0.19.0, Notebooks 1.11.0, Dashboard 2.0.0 and KServe 0.18.0. Newer standalone releases of several of these already exist. That is normal, and it is why you always check which version a tutorial used before copying its commands.

Do not confuse the bundle with a vendor product. Companies sell packaged distributions built on Kubeflow. The community distribution does not endorse or certify any of them. If your employer uses a vendor's build, its install steps and supported versions come from that vendor, not from this guide.
Try it
  1. Open the kubeflow/community-distribution repository on GitHub and find the list of tags.
  2. Find the tag 26.03.1 and read the component version table in its README.
a table listing Pipelines, Trainer, Katib and the rest, each with its own version number. That table is the truth for that release.

The mental model: the parts you will meet

Kubeflow has a lot of nouns. Learn them in groups and the platform stops feeling like a wall of logos.

The platform layer. The Central Dashboard is the web page you open in a browser. It is a shell that embeds the interface of each component, so you click "Notebooks" or "Pipelines" in a side menu and the page for that component appears inside the dashboard. A Profile is Kubeflow's unit of multi-user isolation. Each person or team gets a profile, and a profile creates a Kubernetes namespace of the same name, plus the permissions that make its owner an administrator of that namespace. Your work lives in your profile's namespace, and other users cannot see it unless you add them as contributors. Authentication happens before you reach any of this: a login proxy checks who you are and Istio attaches your email address to each request in a header called kubeflow-userid, which the web applications trust.

The pipeline layer. A component is one step, packaged as a container. A pipeline is a graph of components, where the output of one step feeds the input of the next. You write both in Python using the kfp library, then compile the pipeline into a YAML file, a portable description that the backend can execute. One execution of a pipeline is a run. Runs are grouped into experiments, which is just a folder for organisation. A run that repeats on a schedule is a recurring run.

The data passed between steps. Small values such as a number or a string are parameters. Files such as a dataset or a trained model are artifacts. Artifacts are stored in an object store (an S3-style bucket) and steps receive them as files. This distinction matters a great deal, and we will return to it in the first pipeline.

The training layer. A TrainJob is a request to train a model, possibly across several machines. It refers to a runtime, which is a blueprint written by an administrator that says how to launch, for example, PyTorch distributed training. You pick the runtime and supply your code. A Katib Experiment is a tuning job: a goal, a search space of settings, an algorithm to explore it, and a budget of trials. Each trial is one try.

The notebook layer. A Notebook in Kubeflow terms is a Jupyter or similar server running as a pod, with a persistent disk so your files survive restarts.

YOUbrowser, SDK, kubectl
→
FRONT DOORIstio, oauth2-proxy, Dex, Dashboard
→
COMPONENTSPipelines, Trainer, Katib, Notebooks
→
KUBERNETESpods, volumes, namespaces

Hold on to this stack. When something breaks, the first question is always which layer: did the request get in (front door), did the component accept the work (components), or did Kubernetes fail to start a pod (bottom)? Most beginners look at the wrong layer for hours.

One more term you will see in official pages: KFP v2. Kubeflow Pipelines had an older SDK generation, v1, and it has been removed. Everything here uses v2, and if a tutorial mentions ContainerOp or "v2 compatible mode", it is out of date.

What you need before you install

Installing the whole platform is a cluster project, so decide what you want before you start. There are three realistic starting points, ordered from lightest to heaviest.

No cluster at all. The Pipelines SDK can run pipelines locally, on your own computer, using a local runner. You need Python and, optionally, Docker. This is where you should begin, and this guide's first project uses it. You learn the pipeline language without touching Kubernetes.

One component on a small cluster. You can install Kubeflow Pipelines on its own, without the dashboard, profiles or login. It needs a Kubernetes cluster and kubectl. A local cluster made with kind (Kubernetes in Docker) is enough. This is the right step after local runs, because you see real pods without the multi-user machinery.

The full distribution. The community distribution recommends 16 GB of memory and 8 CPU cores. You can trim the install file to fit into roughly 4 to 8 GB and 2 to 4 cores by removing components you do not need, but expect a slow, tight machine. The platform expects Kubernetes 1.34 or newer, kubectl, and Kustomize 5.8 or close to it. Kind needs Docker or Podman. You also need a cluster storage class that can create disks on demand, because Pipelines, Katib and Hub each request persistent volumes.

A few notes by operating system. The official documentation describes one path, kind on Linux, and does not have separate macOS or Windows pages. On a Mac, the practical route is to run Docker Desktop, Colima or Podman Machine, give that virtual machine at least 16 GB of memory and 8 CPUs, and run the same scripts. This is inferred from how kind works, not spelled out in the docs, so treat it as a best effort. On Apple Silicon, expect some images to fail with "no matching manifest for linux/arm64", because not every image is built for ARM. On Windows, use WSL2 with an Ubuntu shell and run everything inside it; raise the memory limit for WSL in your .wslconfig file. There is no native PowerShell install path.

On Linux, kind clusters that run many pods hit a low file-watcher limit. Raise it before installing:

BASH
sudo sysctl fs.inotify.max_user_instances=2280
sudo sysctl fs.inotify.max_user_watches=1255360

If you skip this, pods crash at random with errors mentioning "too many open files". It is one of the most common first-install failures, and the fix takes ten seconds.

If you work in the Gulf or Egypt, you may also need to think about where the cluster lives. A managed Kubernetes service in a Middle East cloud region keeps training data in-country, which some employers and regulators require. Kubeflow itself has no opinion about the region; it runs wherever the cluster runs. Only your object store and your data sources carry the residency question.

Try it
  1. Check your tools: python3 --version, docker --version, kubectl version --client.
  2. Decide which of the three starting points your machine supports. Be honest about the memory.
a clear answer. If you have under 16 GB free, choose the local runner first and standalone Pipelines second.

Installing the pieces

The SDK: start here

Install the Pipelines SDK into a virtual environment. Pin the version, because the SDK and the backend must agree:

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install kfp==2.17.0
python -c "import kfp; print(kfp.__version__)"
kfp --version

The second-to-last command should print 2.17.0. The kfp package also installs a command-line tool, also named kfp. If you want to build your own component images later with kfp component build, install kfp[all], which adds the Docker client library. You can work through the local-run section below with just this.

Pipelines on its own, on a kind cluster

Create a cluster, then install the standalone Pipelines release. The official steps use a version variable and two applies, with a wait between them so the custom resource definitions are registered before anything uses them:

BASH
kind create cluster
export PIPELINE_VERSION=2.17.0
kubectl apply -k "github.com/kubeflow/pipelines/manifests/kustomize/cluster-scoped-resources?ref=$PIPELINE_VERSION"
kubectl wait --for condition=established --timeout=60s crd/applications.app.k8s.io
kubectl apply -k "github.com/kubeflow/pipelines/manifests/kustomize/env/dev?ref=$PIPELINE_VERSION"
kubectl get pods -n kubeflow
kubectl port-forward -n kubeflow svc/ml-pipeline-ui 8080:80

Give it a few minutes; the pods in the kubeflow namespace move from ContainerCreating to Running. When the port-forward is running, open http://localhost:8080 and you see the Pipelines interface. The dev overlay uses a PostgreSQL database that is meant for trying things out. A production install uses the platform-agnostic overlay with MySQL, which is a topic for the Mid-level guide.

The full distribution

The complete platform installs from the community distribution repository. Check out the tag, create a kind cluster with the helper script, and apply everything in a retry loop:

BASH
git clone https://github.com/kubeflow/community-distribution.git
cd community-distribution
git checkout 26.03.1
./tests/install_KinD_create_KinD_cluster_install_kustomize.sh
kind get kubeconfig --name kubeflow > /tmp/kubeflow-config
export KUBECONFIG=/tmp/kubeflow-config
while ! kustomize build example | kubectl apply --server-side --force-conflicts -f -; do echo "Retrying to apply resources"; sleep 20; done

The loop is not a hack; it is the documented approach, and it is worth understanding. Kubernetes has two kinds of things in the install: definitions of new object types (CRDs), and objects of those new types. If an object arrives before its type is registered, kubectl fails with no matches for kind ... in version ... or ensure CRDs are installed first. The retry loop simply tries again after 20 seconds, by which time the definitions exist. Seeing those errors in the first few rounds is expected.

If Docker Hub rate limits are a concern, log in and create a pull secret named regcred as shown in the repository's README before you start the loop. The community distribution also documents installing component by component, which is useful when a single piece fails and you want to isolate it.

Checking that it works

Wait for the namespaces to fill with running pods, then look at each in turn:

BASH
kubectl get pods -n cert-manager
kubectl get pods -n istio-system
kubectl get pods -n auth
kubectl get pods -n oauth2-proxy
kubectl get pods -n knative-serving
kubectl get pods -n kubeflow
kubectl get pods -n kubeflow-system
kubectl get pods -n kubeflow-user-example-com

Every pod should show Running or Completed. Then open the front door with a port-forward:

BASH
kubectl port-forward svc/istio-ingressgateway -n istio-system 8080:80

Browse to http://localhost:8080. You land on the Dex login page. The development account is user@example.com with the password 12341234, and its profile namespace is kubeflow-user-example-com.

That default password is public. It is in the documentation. Never expose this install to anyone you do not trust without changing it. Change it before the cluster is reachable from outside your laptop, and connect a real identity provider for a shared cluster.

Also note that plain HTTP works on localhost only. On any other address, the web applications set secure cookies, so you need HTTPS in front of the gateway. Otherwise logins loop or fail with cookie errors.

Try it
  1. Install the SDK and run kfp --version.
  2. If your machine has the memory, create a kind cluster and install standalone Pipelines. Open the interface through the port-forward.
the Pipelines page, with a list of pipelines that is empty or contains a few tutorial samples.

Your first pipeline, built step by step

We will build a tiny pipeline that mimics a real one: make some data, train a trivial "model", and report a score. The content is deliberately silly because the point is the shape. Start on your laptop with no cluster.

Step one: a component

A component starts as an ordinary Python function with a decorator. The type hints are required, because Kubeflow uses them to work out what goes in and out:

first_pipeline.py
from kfp import dsl, compiler

@dsl.component(base_image="python:3.11")
def make_numbers(count: int) -> list:
    return [float(i) for i in range(count)]

@dsl.component(base_image="python:3.11")
def average(numbers: list) -> float:
    return sum(numbers) / len(numbers)

Two things deserve a pause. First, @dsl.component turns the function into something that will run in its own container, using the base image you name. The function body is copied into that container and executed there, so the function must be self-contained: any import it needs goes inside the body, and it cannot see variables from the rest of your file. Second, the return type is the contract. int, float, str, bool, list and dict are parameters that pass by value between steps.

If a step needs a library that is not in the base image, list it, and Kubeflow installs it when the step starts:

PYTHON
@dsl.component(base_image="python:3.11", packages_to_install=["pandas"])
def summarise(path: str) -> str:
    import pandas as pd
    return str(pd.read_csv(path).describe())

Notice the import inside the function. That is not style; it is required.

Step two: a pipeline

A pipeline is a function too, decorated with @dsl.pipeline. Its body wires components together. This is where beginners are caught out, so read carefully: the body of a pipeline function does not run as ordinary Python. It runs once, at compile time, only to record which steps exist and which outputs feed which inputs. Do not put if statements on values, loops over data, or print calls there and expect them to run at training time. Everything real happens inside components.

first_pipeline.py
@dsl.pipeline(name="average-pipeline", description="Make numbers and average them")
def average_pipeline(count: int = 10) -> float:
    numbers = make_numbers(count=count)
    result = average(numbers=numbers.output)
    return result.output

Calling a component inside the pipeline returns a task. The task's .output is a placeholder that means "whatever this step returns, once it has run". Passing numbers.output into average creates the dependency: average will not start until make_numbers has finished. Steps that do not depend on each other can run at the same time.

Step three: compile

Compiling turns the Python into a YAML file:

first_pipeline.py
compiler.Compiler().compile(average_pipeline, package_path="average_pipeline.yaml")

Open the YAML. You will see the pipeline's parameters, each component's container image and command, and the graph of dependencies. This file is the real product. It is portable, it can be uploaded to any Kubeflow Pipelines backend, and you can commit it to git and diff it in review. The Python is only a convenient way to write it.

Step four: run it locally

Local execution needs no cluster. You initialise a local runner, then call the pipeline like a function:

first_pipeline.py
from kfp import local

local.init(runner=local.SubprocessRunner())

task = average_pipeline(count=5)
print(task.output)

SubprocessRunner runs each step as a subprocess in a virtual environment on your machine, which is quick and needs no Docker. If you want each step in a container just as on a cluster, use local.DockerRunner() instead; it is slower but closer to the real thing. Expect the first run to pause while the runner creates environments. The printed output is 2.0, the average of 0 to 4.

Local mode is the fastest feedback loop you will have. Use it to debug logic before you pay for cluster time.

Step five: run it on a cluster

With standalone Pipelines port-forwarded to port 8080, submit the same pipeline to the cluster:

submit.py
from kfp.client import Client
from first_pipeline import average_pipeline

client = Client(host="http://localhost:8080")
run = client.create_run_from_pipeline_func(
    average_pipeline,
    arguments={"count": 20},
    experiment_name="first-experiments",
)
client.wait_for_run_completion(run.run_id, timeout=3600)
print("done", run.run_id)

This compiles the pipeline, creates the experiment if needed, starts a run, and waits. In the browser, open the run and you see the graph with each step coloured by state. Click a step to see its logs, its inputs and outputs, and the pod that ran it.

On the full multi-user distribution, the client needs to prove who you are. The simplest place to do this is from a notebook inside the cluster, which is covered in the notebooks section.

Shortcut. To compile from the command line instead of Python, use kfp dsl compile --py first_pipeline.py --output average_pipeline.yaml, and add --function average_pipeline if the file contains more than one pipeline.
Try it
  1. Type in the two components and the pipeline, then compile to YAML and open the file.
  2. Run it locally with SubprocessRunner and confirm it prints 2.0 for a count of 5.
  3. Add a third component that returns the largest number and wire it in.
a YAML file you can read, a correct local answer, and a pipeline where two steps read from the same upstream step.

Passing files between steps: artifacts

Parameters are fine for a number. A trained model or a table of a million rows is not something you pass by value. For files, Kubeflow uses artifacts, which are stored in the object store and handed to each step as a path.

The type hints change. A step that writes a file declares an Output, and a step that reads it declares an Input:

artifacts.py
from kfp import dsl
from kfp.dsl import Dataset, Model, Input, Output, Metrics

@dsl.component(base_image="python:3.11")
def prepare(out_data: Output[Dataset]):
    with open(out_data.path, "w") as f:
        f.write("x,y\n1,2\n2,4\n3,6\n")

@dsl.component(base_image="python:3.11")
def train(data: Input[Dataset], model: Output[Model], metrics: Output[Metrics]):
    rows = open(data.path).read().strip().split("\n")[1:]
    slope = sum(float(r.split(",")[1]) / float(r.split(",")[0]) for r in rows) / len(rows)
    with open(model.path, "w") as f:
        f.write(str(slope))
    metrics.log_metric("slope", slope)

@dsl.pipeline(name="artifact-pipeline")
def artifact_pipeline():
    d = prepare()
    train(data=d.outputs["out_data"])

Read the function signatures closely. out_data: Output[Dataset] is not a value you return; it is a handle. You write your file to out_data.path, and Kubeflow uploads it to the object store afterwards. In the next step, data: Input[Dataset] gives you a local path to the downloaded copy. You never write an S3 client yourself.

The artifact types carry meaning: Dataset, Model and Metrics are the common ones, and there are others for classification metrics, HTML and Markdown. The interface uses the type to show them properly; Metrics values appear in the run's metrics view, and you can compare them across runs.

Notice that when a component has an Output artifact, the pipeline refers to it by name through d.outputs["out_data"]. When a component has a single plain return value, you use .output. That split causes many beginner errors: if Kubeflow complains that a task has more than one output, use the named form.

Where do the files go? The object store named in the pipeline root. On the standard install this is a store reached through a service still called minio-service, although the default store is now SeaweedFS, a replacement for MinIO that speaks the same S3 protocol. You will see minio://mlpipeline/v2/artifacts as the default root. On a cloud cluster you point it at an S3 or GCS bucket instead, which is a configuration topic for later.

Do not pass big data as a parameter. A list of a million numbers as a list parameter ends up inside Kubernetes objects, which have size limits, and runs fail in strange ways. Anything large belongs in an artifact.
Try it
  1. Run the artifact pipeline locally with SubprocessRunner.
  2. Change the data in prepare and confirm the logged slope changes.
the slope printed or logged as 2.0 for the sample data, changing when you change the rows.

Runs, caching and reading the interface

Once a pipeline is on a cluster, you spend most of your time in a few screens. Learn what each is telling you.

The pipelines list holds uploaded pipeline definitions. A pipeline can have several versions, each an uploaded YAML. The experiments list groups runs. The runs list shows every execution with its status, duration and the experiment it belongs to. Open a run and the graph view draws your steps as boxes, coloured for pending, running, succeeded, failed or cached. Clicking a box opens a side panel with tabs for input and output, visualisations, details, logs and the events from Kubernetes.

When a step fails, go to the logs tab first. If there are no logs, the step never started, so switch to the events view. A failure before logs exist is almost always a Kubernetes problem rather than a Python problem: the image could not be pulled, the pod could not be scheduled for lack of memory, or a volume could not mount. A failure with logs is usually your own code, and the traceback is right there.

Caching is on by default and is the feature that surprises people most. If a step's definition and inputs are identical to a previous successful run, Kubeflow skips it and reuses the earlier outputs, marking the step as cached. For expensive steps this saves hours. For a step that reads the current time or pulls fresh data from outside, it silently gives you stale answers. You have two controls. Per run, pass enable_caching=False to create_run_from_pipeline_func. Per step, set task.set_caching_options(False) inside the pipeline:

PYTHON
@dsl.pipeline(name="no-cache-demo")
def no_cache_demo():
    t = make_numbers(count=5)
    t.set_caching_options(False)

You can also set a cache-off default for a whole compile by passing --disable-execution-caching-by-default to kfp dsl compile.

A recurring run repeats a pipeline on a schedule. It is created from the interface, from client.create_recurring_run, or from kfp recurring-run create. Older tutorials call these "jobs"; the SDK methods with "job" in their names are deprecated in favour of the recurring-run names and still print a deprecation warning.

Under the hood, the backend converts your compiled pipeline into an Argo Workflow and runs it. You do not need Argo knowledge to be productive, but it explains two things you will see in pod lists: each step has a pod, and the pod names look like generated workflow names. When you search with kubectl get pods -n <your namespace> during a run, those are your steps.

Try it
  1. Run the same pipeline twice with identical arguments on a cluster.
  2. Compare the second run's graph with the first.
  3. Run it again with enable_caching=False.
steps marked cached on the second run, finishing in seconds, and ordinary executions again on the third.

Profiles, the dashboard and notebooks

If you installed the full distribution, you work through the Central Dashboard, and the ideas here apply.

After logging in, the top bar shows a namespace selector. That is your profile. Everything you create, such as notebooks, runs and experiments, belongs to the namespace selected. Switching it changes which resources you see. A frequent beginner complaint is "my run disappeared", which almost always means the selector is on a different namespace.

To create a profile for a new user, an administrator applies a small YAML to the cluster:

profile.yaml
apiVersion: kubeflow.org/v1
kind: Profile
metadata:
  name: aya-team
spec:
  owner:
    kind: User
    name: aya@example.com
BASH
kubectl apply -f profile.yaml
kubectl get profiles

The Profile controller reacts by creating the namespace aya-team, a role binding that makes the owner an administrator there, and two service accounts named default-editor and default-viewer. Self-service profile creation through the dashboard is off by default and turned on with the setting CD_REGISTRATION_FLOW=true on the dashboard deployment.

Deleting a profile deletes everything in it. kubectl delete profile aya-team removes the namespace with all notebooks, volumes and runs. Never delete the Profiles CRD itself: that removes every profile and every user namespace on the cluster.

Also know the limits of this isolation. The documentation states plainly that profiles give no stronger separation than Kubernetes namespaces do. They keep honest colleagues out of each other's way. They are not a hard security boundary between mutually distrustful tenants.

Notebooks

Open Notebooks in the dashboard and choose New Notebook. You pick a name, an image (Jupyter with various libraries is offered, including GPU builds), CPU and memory, and a workspace volume. The workspace volume is a persistent disk mounted into the pod, and it is the only place files survive a restart. Anything you save elsewhere vanishes when the notebook stops.

Behind the page, a Notebook object is created, and the notebook controller turns it into a pod named <name>-0. If a notebook never becomes ready, inspect it the Kubernetes way:

BASH
kubectl get notebooks -n aya-team
kubectl describe notebook my-notebook -n aya-team
kubectl logs my-notebook-0 -n aya-team

The usual reasons are an image that cannot be pulled, a disk that cannot be created, or a namespace that has hit its resource quota. Culling, meaning automatically stopping idle notebooks, is off by default. An administrator enables it with ENABLE_CULLING on the notebook controller, and the idle time defaults to 1440 minutes, one day.

Talking to Pipelines from a notebook

The easiest authenticated client is one created inside a notebook. The distribution defines a PodDefault named access-ml-pipeline: a rule that, for notebooks carrying a matching label, mounts a short-lived token for the Pipelines API. When you create the notebook, tick the configuration called "Allow access to Kubeflow Pipelines" in the form. Then from that notebook:

PYTHON
import kfp
client = kfp.Client()
print(client.list_experiments())

Without the PodDefault you get Unauthenticated: Request header error: there is no user identity header, which means the request reached the API without a login or a token.

Try it
  1. Log in to the dashboard, open Notebooks and create a small notebook with a 5 GB workspace.
  2. In a terminal inside it, create a file under your home folder, stop the notebook and start it again.
your file still present, because it lived on the workspace volume.

Training jobs and tuning

Not every job is a pipeline step. When you need to train across several machines, or try a hundred learning rates, Kubeflow has purpose-built tools.

Kubeflow Trainer

Trainer is the current way to run training jobs on Kubernetes. It replaces an older project called the Training Operator, which defined objects like PyTorchJob and TFJob. Those old objects still appear in tutorials, and the old operator has been removed from the development branch of the distribution. Write new work for Trainer.

A TrainJob is the object you create. It names a runtime, which an administrator has already installed. A runtime is a template for a framework, such as torch-distributed, so you do not have to know how to wire processes together. The shipped runtimes cover PyTorch, DeepSpeed, JAX, MLX and XGBoost, plus ready-made fine-tuning recipes for a few small language models.

Install the Trainer and its runtimes on a cluster with Kubernetes 1.31 or newer using Helm (see the Helm guide):

BASH
helm install kubeflow-trainer oci://ghcr.io/kubeflow/charts/kubeflow-trainer \
  --namespace kubeflow-system --create-namespace \
  --version 2.3.0 --set runtimes.defaultEnabled=true

The full distribution already includes Trainer, in the kubeflow-system namespace. The easiest way to use it is the Kubeflow SDK, a separate Python package that wraps several components behind one client:

BASH
pip install -U kubeflow
train_job.py
from kubeflow.trainer import TrainerClient, CustomTrainer

def train_fn():
    import torch
    print("hello from", torch.__version__)

client = TrainerClient()
for runtime in client.list_runtimes():
    print(runtime.name)

job = client.train(
    trainer=CustomTrainer(func=train_fn, num_nodes=2,
                          resources_per_node={"cpu": 3, "memory": "16Gi"}),
    runtime=client.get_runtime("torch-distributed"),
)
client.wait_for_job_status(job)
for line in client.get_job_logs(job, follow=True):
    print(line)

You hand Trainer a plain Python function, the number of machines, and the resources each one gets. Trainer packages the function, starts the pods and sets up the environment variables PyTorch needs to find its peers. client.list_runtimes() is how you discover what your administrator installed. If it returns nothing, the runtimes are missing, not your code.

Under the hood, the Kubernetes objects are TrainJob and a JobSet, and you can inspect them:

BASH
kubectl get trainjobs -n aya-team
kubectl describe trainjob <name> -n aya-team
kubectl get jobsets -n aya-team

The same job can be written as YAML and applied with kubectl, which suits GitOps workflows. Its API group is trainer.kubeflow.org/v1alpha1. The v1alpha1 label tells you the API is still evolving, and Trainer 2.2 already contained breaking changes, so pin versions and read release notes.

Katib for hyperparameter tuning

A hyperparameter is a setting you choose before training, such as a learning rate, that the training does not learn by itself. Finding a good one by hand is slow. Katib automates it: you define an Experiment with an objective (maximise accuracy, say), a search space, an algorithm such as random search, grid search, Bayesian optimisation or Hyperband, and a limit on trials. Katib's suggestion service proposes settings, launches a trial for each, collects the result and feeds it back.

The key detail for newcomers is how Katib learns the result. By default it reads your program's standard output and looks for lines of the form name=value. If the objective metric is called accuracy, your code must print exactly accuracy=0.93. A trial that never prints in that format ends as MetricsUnavailable.

tune.py
import kubeflow.katib as katib

def objective(parameters):
    result = 4 * int(parameters["a"]) - float(parameters["b"]) ** 2
    print(f"result={result}")

client = katib.KatibClient(namespace="aya-team")
client.tune(
    name="tune-experiment",
    objective=objective,
    parameters={
        "a": katib.search.int(min=10, max=20),
        "b": katib.search.double(min=0.1, max=0.2),
    },
    objective_metric_name="result",
    max_trial_count=12,
    resources_per_trial={"cpu": "2"},
)
client.wait_for_experiment_condition(name="tune-experiment")
print(client.get_optimal_hyperparameters("tune-experiment"))

The kubeflow-katib package provides KatibClient and installs with pip install -U kubeflow-katib. Katib's API group is kubeflow.org/v1beta1. Every trial runs as pods in your namespace, so your resource quota matters; twelve trials at two CPUs each can starve a small cluster if they run in parallel.

Try it
  1. With Trainer installed, run client.list_runtimes() and note the names.
  2. Run the Katib example with max_trial_count=6. Open the Katib page in the dashboard and watch the trials appear.
a list of runtimes such as torch-distributed, and six trials each reporting a result, with one marked optimal.

Everyday commands, grouped by what you want to do

Keep this section open in a tab. Every command below is covered by the official CLI or client.

Write and compile.

BASH
kfp dsl compile --py pipeline.py --output pipeline.yaml
kfp dsl compile --py pipeline.py --output pipeline.yaml --pipeline-parameters '{"count": 10}'

Upload and manage pipelines. These need --endpoint pointing at your API, or the environment set up for it:

BASH
kfp --endpoint http://localhost:8080 pipeline create -p my-pipeline pipeline.yaml
kfp --endpoint http://localhost:8080 pipeline list
kfp --endpoint http://localhost:8080 pipeline create-version -p PIPELINE_ID -v v2 pipeline.yaml
kfp --endpoint http://localhost:8080 pipeline delete PIPELINE_ID

Run and watch.

BASH
kfp --endpoint http://localhost:8080 experiment create first-experiments
kfp --endpoint http://localhost:8080 run create -e first-experiments -r run-one -f pipeline.yaml -w count=10
kfp --endpoint http://localhost:8080 run list -e EXPERIMENT_ID
kfp --endpoint http://localhost:8080 run get RUN_ID -w

The -w flag makes the command wait until the run finishes; key=value pairs at the end set the pipeline's parameters.

Schedule. kfp recurring-run create, list, get, enable, disable and delete manage scheduled runs.

Inspect Kubernetes objects.

BASH
kubectl get pods -n aya-team
kubectl describe pod POD_NAME -n aya-team
kubectl logs POD_NAME -n aya-team
kubectl get profiles
kubectl get notebooks -n aya-team
kubectl get trainjobs -n aya-team

Python client. The same operations exist on kfp.Client: create_run_from_pipeline_func, create_run_from_pipeline_package, list_runs, get_run, wait_for_run_completion, terminate_run, upload_pipeline, list_pipelines, create_experiment and list_experiments. Two renamings to remember: create_recurring_run replaces the old "job" functions, and GPU requests use set_accelerator_type and set_accelerator_limit rather than the deprecated set_gpu_limit.

Ask for resources per step. This is how a step gets memory and a GPU:

PYTHON
train_task = train(data=d.outputs["out_data"])
train_task.set_cpu_limit("4")
train_task.set_memory_limit("8G")
train_task.set_accelerator_type("nvidia.com/gpu")
train_task.set_accelerator_limit(1)

A step that asks for more than any node can give stays Pending forever, with no error in the logs; the Kubernetes events say Insufficient cpu or Insufficient memory.

Try it
  1. Compile your first pipeline with the CLI and upload it with kfp pipeline create.
  2. Start a run from the CLI with -w and a parameter override.
the command waiting, then reporting success, with the run visible in the browser.

Configuration and the most common errors

The few settings worth knowing now

You will not configure much at this level, but four settings explain a lot of behaviour.

The pipeline root says where artifacts are stored. Each profile namespace has a ConfigMap named kfp-launcher with a key defaultPipelineRoot. On the default install it points at the bundled object store. To use a cloud bucket, an administrator changes this value and supplies credentials. You can also override it for a single pipeline with pipeline_root when submitting.

The default base image for components is a recent Python 3.11 image. Naming base_image explicitly, as we did, makes your pipeline independent of that default.

Namespace for the client. On a multi-user install, runs belong to a namespace. Pass namespace="aya-team" when you create runs, or call client.set_user_namespace("aya-team") once; it is remembered in $HOME/.config/kfp/context.json.

Culling and registration for notebooks and profiles, as described above, are switched with environment variables on the respective deployments.

Reading the errors

Kubeflow errors usually come from one of the four layers in the stack diagram. Match the message to the layer and the fix becomes obvious.

What you seeWhat it meansWhat to do
no matches for kind ... ensure CRDs are installed firstObjects applied before their type definitions were registeredExpected on first install. Re-run the apply, or use the retry loop.
Pods crash, "too many open files"Low file-watch limits on the host running kindRaise the two inotify sysctl values shown earlier.
toomanyrequests: You have reached your pull rate limitDocker Hub's anonymous limitLog in and create the regcred pull secret, or mirror images.
no matching manifest for linux/arm64An image is not built for ARMUse amd64 nodes or emulation.
RBAC: access denied in the dashboardAn Istio authorisation policy rejected youCheck you are using the right profile and that the user header is reaching the service.
Unauthenticated: Request header error: there is no user identity headerAPI called without a login or tokenGo through the gateway with a cookie or token, or use the in-notebook client.
PermissionDenied: User ... is not authorizedNo role in the target namespaceAdd the user as a contributor or pass the correct namespace.
Step stays PendingKubernetes cannot place the podRead events with kubectl describe pod; lower the request or add capacity.
Katib trial MetricsUnavailableMetric not printed as name=valuePrint exactly the objective name, an equals sign and a number.
DeprecationWarning: 'set_gpu_limit' is deprecatedOld SDK callRename to set_accelerator_limit.

Another message deserves its own paragraph: "output_metadata.json": proto: ... unknown field "custom_path". It appears when a pipeline compiled with SDK 2.15.0 or 2.15.1 runs against an older backend. The fix is to upgrade the SDK to 2.15.2 or later, or to upgrade the backend, which is why the SDK and backend versions should stay matched.

When you are lost, use a fixed order. First, is the step in the run graph failing with logs? Read the traceback. Without logs, use kubectl describe pod on the step's pod and read the events at the bottom. If the page itself will not load, check the ingress and login pods in istio-system, oauth2-proxy and auth. That three-step habit solves most problems before any search engine is needed.

Try it
  1. Ask a step for 1000 CPUs with set_cpu_limit("1000") and run it.
  2. Find the pod with kubectl get pods and run kubectl describe pod on it.
a pod stuck in Pending and an event at the bottom saying Insufficient cpu. Then fix the request and watch it run.

Putting it all together

Here is one small, complete project that uses what you have learned. It is a pipeline that prepares data, trains a toy model with a configurable setting, evaluates it, and only reports success if the score clears a threshold. The project has three files.

pipeline.py
from kfp import dsl, compiler
from kfp.dsl import Dataset, Model, Input, Output, Metrics

@dsl.component(base_image="python:3.11")
def prepare(rows: int, out_data: Output[Dataset]):
    with open(out_data.path, "w") as f:
        f.write("x,y\n")
        for i in range(1, rows + 1):
            f.write(f"{i},{2 * i}\n")

@dsl.component(base_image="python:3.11")
def train(data: Input[Dataset], scale: float, model: Output[Model]):
    lines = open(data.path).read().strip().split("\n")[1:]
    ratios = [float(l.split(",")[1]) / float(l.split(",")[0]) for l in lines]
    with open(model.path, "w") as f:
        f.write(str(scale * sum(ratios) / len(ratios)))

@dsl.component(base_image="python:3.11")
def evaluate(model: Input[Model], metrics: Output[Metrics]) -> float:
    slope = float(open(model.path).read())
    error = abs(slope - 2.0)
    metrics.log_metric("error", error)
    return error

@dsl.pipeline(name="toy-training", description="prepare, train, evaluate")
def toy_training(rows: int = 50, scale: float = 1.0):
    d = prepare(rows=rows)
    t = train(data=d.outputs["out_data"], scale=scale)
    evaluate(model=t.outputs["model"])

if __name__ == "__main__":
    compiler.Compiler().compile(toy_training, package_path="toy_training.yaml")
run_local.py
from kfp import local
from pipeline import toy_training

local.init(runner=local.SubprocessRunner())
toy_training(rows=20, scale=1.0)
run_cluster.py
from kfp.client import Client
from pipeline import toy_training

client = Client(host="http://localhost:8080")
for scale in [0.5, 1.0, 1.5]:
    client.create_run_from_pipeline_func(
        toy_training,
        arguments={"rows": 100, "scale": scale},
        experiment_name="toy-scale-sweep",
    )

The workflow is the one you would use on real work. You iterate with run_local.py until the logic is right, which costs nothing. You then compile and commit toy_training.yaml so the exact definition is reviewed in git. You submit three runs on the cluster, one per scale, all inside one experiment. Open that experiment in the interface, select the runs and compare the error metric: the run with scale=1.0 has the lowest error, which you can see at a glance because the metrics are logged against each run.

Two habits make this scale up. Every step names its base image, so it runs the same tomorrow. Every large output is an artifact, so nothing depends on memory between steps. Replace the toy steps with your real data loading and training code, keep the same skeleton, and you have a real pipeline.

When you want to tune scale automatically instead of by hand, that is Katib's job. When you want to run the training step across several GPUs, that is Trainer's. When the model is good enough to serve, continue to the KServe guide. If you want to log parameters and metrics across many experiments with a lighter tool alongside this, see the MLflow guide.

Try it
  1. Run run_local.py and read the logged error.
  2. Compile the pipeline and submit the three-scale sweep on your cluster.
  3. Compare the three runs in the interface.
three runs in one experiment, with the metric lowest at a scale of 1.0.

What you can now do, and what comes next

You can explain what Kubeflow is and why it is a family of projects rather than one program. You know that version numbers now follow the calendar, that the current bundle is 26.03.1, and that pipelines, Trainer, Katib, Notebooks, Hub and KServe each have a separate job. You can name the layers a request passes through and use that to locate a failure. You have installed the SDK, written components and a pipeline, compiled them to YAML, run them locally and on a cluster, passed files between steps as artifacts, used caching on purpose, and read a run in the interface. You can create a profile, start a notebook that survives restarts, launch a training job with the Kubeflow SDK and tune a function with Katib. And you can read the ten most common error messages.

What to practise next: write one pipeline for your own project rather than a toy, even if it has only three steps. Run it twice and watch caching. Break it deliberately and read the failure. Those three acts teach more than another hour of reading.

The Mid-level guide picks up with control flow (conditionals, loops, parallel fan-out and exit handlers), resource and secret management with kfp-kubernetes, how the backend works internally, production object stores and databases, and connecting your cluster to CI and GitOps. The Senior guide covers running Kubeflow as a platform for other teams: multi-tenancy limits, upgrades, security and capacity. Neighbouring guides in this catalogue worth reading alongside: Kubernetes for the substrate, Docker for the component images, KServe for serving, and MLflow for experiment tracking.

One last piece of advice for learning on your own. Keep a notebook of the failures you meet, with the message, the layer it came from and the fix. Kubeflow is a platform where the same dozen problems recur in different disguises: a missing permission, a pod that cannot start, a stale cache, a mismatched version. After a month, your own list will be a better reference than any guide, including this one, and it is the thing interviewers respond to most when you describe real debugging stories from your projects.

Sources