تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
ZenMLMLOpsPipelines & orchestration3 مستويات106 قسمًايغطّي ZenML 0.97دليل بالإنجليزية

The Complete ZenML Guide

Build portable ML pipelines that run on any stack with ZenML. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
16sections
44examples

This is part one of three. It covers everything you need to do real work with ZenML 0.97, not a teaser. By the end you can install it, write a pipeline out of steps, run it, see every intermediate result stored and versioned, control caching, read the dashboard, swap the infrastructure underneath your code without rewriting it, and diagnose the errors beginners meet first. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas only stick once you have watched your own pipeline run, skip a step because of the cache, and fail in a way you can read.

What ZenML is, and the problem it solves

ZenML is an open-source Python framework for building machine learning workflows that run the same way on your laptop and on production infrastructure. You write ordinary Python functions, mark them with a decorator, and ZenML records what went in, what came out, which code produced it, and where everything is stored. Then you can point the very same code at Kubernetes, SageMaker, Vertex AI or another backend by changing a configuration object instead of the code.

STEPSPython functions
→
PIPELINEsteps wired together
→
STACKwhere it runs
→
TRACKED RUNartifacts, logs, lineage

To see why this matters, think about how a machine learning project usually starts. Someone writes a notebook. It loads a CSV, cleans it, trains a model and prints an accuracy. It works, so the team wants to run it every night on fresh data, on a bigger machine, with a GPU, and to know next month exactly which data and code produced the model that is now in production.

The notebook cannot do any of that. Cells depend on hidden state, so running them in a different order gives a different answer. Intermediate results live in memory and vanish when the kernel dies. Moving to a bigger machine means rewriting the code around a scheduler, a container and a storage bucket. Six months later nobody can say which version of the data trained the model, because nothing recorded it.

The usual fix is to glue together a workflow scheduler such as Airflow, an experiment tracker such as MLflow, a container build, and a pile of scripts that move files around. That works, but the glue is the hard part, and it ties the code to one scheduler. ZenML sits above those tools rather than replacing them. It gives you one small Python interface for describing the work, and it talks to schedulers, container registries and trackers on your behalf.

Three consequences of that design shape everything that follows.

Your code describes the work, not the infrastructure. A step is a plain Python function. It does not import a Kubernetes client or a cloud SDK. Where it runs is decided by something separate, the stack, which you can change without touching the function.

Every output is stored, not passed around in memory. When one step returns a dataframe and the next step receives it, ZenML has written the dataframe to storage in between and read it back. That sounds wasteful until you notice what it buys you: the result survives a crash, can be inspected later, can be reused by a different run, and carries a version and a lineage.

ZenML is a metadata layer, not a data warehouse. A small server keeps records about runs. Your actual data lives in storage you own: a local folder on day one, an S3 bucket or Google Cloud Storage bucket later. That matters for data residency. If your employer in the Gulf or in Egypt needs the data to stay in a particular region, the artifacts stay in the bucket you chose in the region you chose, and only metadata goes to the server.

ZenML is licensed under Apache-2.0. There is a commercial product, ZenML Pro, which adds team features such as role-based access control, but everything in this guide works with the free open-source version.

Try it

Think of one notebook you have written. On paper, list its cells, and for each one write what it needs as input and what it produces. Each cell with a clear input and output is a candidate step, and you have just sketched your first pipeline.

The core mental model: steps, pipelines, artifacts and stacks

ZenML has a small vocabulary. Learn these nouns well and the documentation becomes easy to read.

A step is a Python function decorated with @step. Its inputs and outputs have type annotations, and ZenML uses those annotations to decide how to store each output. One step should do one job: load data, clean it, train, evaluate.

A pipeline is a function decorated with @pipeline that calls steps and passes the output of one into the next. The pipeline body does not do the work itself. It describes the order and the data flow. When you call a pipeline, ZenML first works out the graph of steps, a directed acyclic graph or DAG, and then executes it. This is the default, called a static pipeline. There is also a newer dynamic pipeline, written @pipeline(dynamic=True), whose body runs at runtime so you can use ordinary loops and conditions. It is an advanced tool, so this guide stays with static pipelines.

An artifact is anything a step returns. Each artifact is written to the artifact store, versioned, and linked to the run and step that created it. Downstream steps read it from the store. By default an output is called output, or output_0, output_1 and so on when a step returns several. You can give it a proper name with Annotated[type, "name"], which you will do in every real project.

A materializer is the small piece of code that knows how to write one kind of Python object to storage and read it back. ZenML ships materializers for integers, strings, lists, dictionaries, NumPy arrays and pandas dataframes, and integrations add more, for example scikit-learn models. When nothing matches, ZenML falls back to cloudpickle. That fallback works but is not safe for production: pickles can break across Python versions and can run arbitrary code when loaded, so treat the warning that mentions it as a prompt to fix something.

A stack is a named set of infrastructure pieces, called stack components, that a pipeline runs on. Every stack needs at least two components. The orchestrator decides what runs when and where. The artifact store is where artifacts are written. Others, such as a container registry or an experiment tracker, are optional. The stack you get on day one is called default. It pairs a local orchestrator, which runs steps on your own machine, with a local artifact store, which is a folder on disk.

A flavor is a concrete implementation of a component type. The local orchestrator and the kubernetes orchestrator are two flavors of the orchestrator type. The local artifact store and the s3 artifact store are two flavors of the artifact store type.

A pipeline run is one execution of a pipeline. Each run records its status, its parameters, the artifacts it produced and its logs. Finally, the ZenML server is a small web application with a database and a dashboard. It stores the records of runs, artifacts and stacks, and your client talks to it over HTTP.

How the pieces fit together
Your code — steps and a pipeline, plain Python, no infrastructure in it
ZenML client — the zenml package and CLI you installed
ZenML server — metadata only: runs, artifact records, stacks, secrets
Stack — orchestrator + artifact store (+ optional components), where the work and the data actually live

Keep one sentence in mind: code says what to do, the stack says where, and the server remembers what happened. When something confuses you later, ask which of the three it belongs to.

Try it

Without looking back, write the four nouns in order from smallest to largest: step, artifact, pipeline, stack. Then state in one sentence what the orchestrator and the artifact store each do. Check yourself against the section above.

Installing ZenML and checking the setup

ZenML 0.97, the version this guide targets, supports Python 3.10 through 3.14. Use a virtual environment so that ZenML's dependencies stay separate from everything else on your machine. On Linux or macOS:

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install 'zenml[server]'

The square-bracket part is an extra, an optional bundle of dependencies. It matters more than it looks. Since version 0.90 the bare pip install zenml installs only the client, which can talk to a deployed server but cannot run a database of its own. You choose an extra depending on what you want:

You want Install
Pure local use with a SQLite database and no server process pip install 'zenml[local]'
A local server and dashboard on your laptop pip install 'zenml[server]'
Nicer output inside Jupyter notebooks pip install 'zenml[jupyter]'
Only a client that connects to a team server pip install zenml

If you install the bare package and then try to use it without a server, you will see an error that begins like this:

TEXT
ImportError: It seems like you've installed the `zenml` package without the `local` extra, but are trying to use ZenML with a local database.

The message itself tells you both fixes: install zenml[local], or run zenml login to connect to a server. This is the single most common first error, because older tutorials say pip install zenml and nothing more.

If you prefer the uv tool, uv venv and uv pip install 'zenml[server]' work the same way.

macOS on Apple Silicon. Before you start a local server, set this variable once per shell:

BASH
export OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES

Without it the local server process can crash on the newer macOS fork checks. You do not need it if you only connect to a remote server.

Windows. The ZenML FAQ says Windows is officially supported only through WSL, the Windows Subsystem for Linux. Some commands work natively, but anything that starts a server process does not. Install WSL, open an Ubuntu shell, and follow the Linux steps. This is worth doing even if it feels like a detour, because almost all ML infrastructure targets Linux.

Now verify the installation. Each of these commands answers a different question:

BASH
zenml version
python -c "import zenml; print(zenml.__version__)"
zenml status
zenml stack describe

zenml version prints the installed version, for example 0.97.0. The Python one-liner confirms that the package imported by your interpreter is the same one the command line uses, which catches the classic mistake of installing into one environment and running from another. zenml status reports whether you are connected to a server, which stack is active and which project you are in. zenml stack describe prints the components of the active stack. On a fresh install you should see the default stack with a local orchestrator and a local artifact store.

When you need to ask someone for help, run zenml info -a -s and paste the output. It gathers the versions and settings that maintainers ask for first.

Try it

Create a fresh virtual environment, install zenml[server], and run the four verification commands. Write down the ZenML version and the name of the active stack. If zenml status says you are not connected to anything yet, that is expected. The next section fixes it.

Your first server, and a project folder

A ZenML client keeps its records in a server. For learning, the server runs on your own laptop. Two commands set everything up. Inside a new project folder, run:

BASH
mkdir zenml-first && cd zenml-first
zenml init
zenml login --local

zenml init marks the current folder as the source root of your project. It creates a small hidden .zen directory. ZenML uses the source root to work out how to import your code, and later, to decide which files to package when a pipeline runs in a container. Run it once at the top of each project.

zenml login --local starts a local server as a background process and prints the address of its dashboard, by default on port 8237. The first time you open the dashboard in a browser it asks you to set up an administrator account. Afterwards your client is connected to that server automatically.

Two facts about the local server save an hour of confusion later. First, it does not survive a reboot or a shutdown. After you restart your machine, run zenml login --local again. If you forget, the next pipeline run fails with an error that ends in Connection refused and mentions 127.0.0.1:8237. Second, older tutorials tell you to run zenml up, zenml down and zenml connect. Those commands still exist in 0.97 but are deprecated and print warnings. The current forms are zenml login --local, zenml logout --local and, for a remote server, zenml login https://your-server.

If you have Docker installed and prefer to keep the server in a container, zenml login --local --docker does the same thing inside one. Learn Docker first if the word container is new, using the Docker guide.

Run zenml status again. It should now show that you are connected to a local server. Open the dashboard and look at the empty Pipelines page. By the end of this guide it will be full.

Tip. If a command complains about a missing server, run zenml status first. It tells you in one line whether you are connected, so you can tell a forgotten zenml login --local from a real bug.

Try it

Run zenml init and zenml login --local in a new folder. Open the dashboard address it prints, create the administrator account, and find the Stacks page. Confirm that a stack named default exists. Then close the terminal, open a new one and run zenml status to see that the connection is remembered.

Your first pipeline, step by step

We will build the smallest useful pipeline: it loads a dataset, trains a classifier, and reports its accuracy. First install the scikit-learn integration, which also adds the materializer that knows how to store scikit-learn models:

BASH
zenml integration install sklearn -y --uv

An integration is a bundle of extra packages and ZenML components for one external tool. The -y flag skips the confirmation prompt, and --uv installs with the faster uv installer. If you do not have uv, leave that flag off.

Now create a file named run.py. We build it in three pieces so each idea is visible. First, the imports and the data step:

run.py
from typing import Annotated, Tuple

import pandas as pd
from sklearn.base import ClassifierMixin
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split

from zenml import pipeline, step


@step
def load_data() -> Tuple[
    Annotated[pd.DataFrame, "X_train"],
    Annotated[pd.DataFrame, "X_test"],
    Annotated[pd.Series, "y_train"],
    Annotated[pd.Series, "y_test"],
]:
    """Load the iris dataset and split it into train and test parts."""
    iris = load_iris(as_frame=True)
    X_train, X_test, y_train, y_test = train_test_split(
        iris.data, iris.target, test_size=0.2, random_state=42
    )
    return X_train, X_test, y_train, y_test

Read the decorator and the return annotation closely, because this is where ZenML does its work. @step tells ZenML that load_data is a step. The return type is a Tuple of four Annotated values. Each Annotated[type, "name"] says: this output has this Python type and this artifact name. ZenML uses the type to choose a materializer, for example the pandas one for a dataframe, and the name to label the artifact in the dashboard.

A rule that trips people up: a step has several outputs only when its return statement is a tuple, as it is here. If you annotate a Tuple but return a single object, ZenML treats the whole thing as one output.

The second piece is the training step:

run.py
@step
def train_model(
    X_train: pd.DataFrame, y_train: pd.Series, n_estimators: int = 100
) -> Annotated[ClassifierMixin, "model"]:
    """Train a random forest and return it."""
    model = RandomForestClassifier(n_estimators=n_estimators, random_state=42)
    model.fit(X_train, y_train)
    return model

The inputs X_train and y_train come from artifacts produced by another step. n_estimators is a parameter: a simple JSON-friendly value such as a number or string that ZenML records with the run. The return annotation names the trained model model. The type ClassifierMixin is the scikit-learn base class for classifiers, and the scikit-learn integration's materializer handles it.

The third piece is an evaluation step and the pipeline that connects everything:

run.py
@step
def evaluate(
    model: ClassifierMixin, X_test: pd.DataFrame, y_test: pd.Series
) -> Annotated[float, "accuracy"]:
    """Score the model on held-out data."""
    accuracy = accuracy_score(y_test, model.predict(X_test))
    print(f"Test accuracy: {accuracy:.3f}")
    return float(accuracy)


@pipeline
def training_pipeline(n_estimators: int = 100):
    X_train, X_test, y_train, y_test = load_data()
    model = train_model(X_train, y_train, n_estimators=n_estimators)
    evaluate(model, X_test, y_test)


if __name__ == "__main__":
    training_pipeline()

Inside training_pipeline you see ordinary-looking function calls, but they do not execute the steps immediately. When ZenML compiles the pipeline, each call returns a placeholder that stands for the future output of that step. Passing it to the next step is how you draw an edge in the graph. That is why you cannot do arbitrary Python on those placeholders inside the pipeline body, such as printing the dataframe or calling len() on it. The real values exist only when the steps run.

Run it:

BASH
python run.py

ZenML prints the run name, the step each stage is executing, and a link to the dashboard. You will see lines showing that load_data, train_model and evaluate each ran, your printed accuracy, and the message that the pipeline run finished. The accuracy for this dataset is typically high, near 1.0, because iris is an easy problem.

Notice what you did not write: no file paths, no serialization code, no database calls. Yet every output is now stored and recorded. Open the dashboard, click Pipelines, click training_pipeline, then the run. You will see the graph of three steps with the artifacts between them. Click the model artifact to see its version and where it is stored.

Try it

Type in the three pieces, run python run.py, and open the run in the dashboard. Click each step and find its logs, its output artifacts and the accuracy printed by evaluate. Then change test_size to 0.3 and run again. Compare the two runs in the list.

What just happened: runs, artifacts and the local store

It helps to know exactly where everything went, so that ZenML never feels like magic. Run these commands:

BASH
zenml pipeline list
zenml pipeline runs list
zenml artifact list

Since version 0.95, list commands show the newest items first. That is a change from earlier releases, and old blog posts may describe the opposite order. zenml pipeline list shows training_pipeline. zenml pipeline runs list shows each run with a name and a status. You can filter to one pipeline with --pipeline=training_pipeline, and you can change the output format with --output, which accepts table, json, yaml, csv and tsv. The JSON form is handy for scripts.

zenml artifact list shows the named artifacts you created: X_train, X_test, y_train, y_test, model and accuracy. If you had not used Annotated, each would appear with an unhelpful default name such as training_pipeline::train_model::output. Naming them costs one line and saves real confusion later.

The data itself lives in the local artifact store, a folder inside ZenML's configuration directory on your machine. You do not need to go there, and you should not edit it by hand, but knowing that a real file exists for every artifact demystifies the dashboard. When you later switch to an S3 bucket, the same artifacts simply land in the bucket.

Each run is given a name automatically, made from the pipeline name and a timestamp. Run names must be unique within a project. If you pass your own run_name and reuse it, you get Pipeline run name '...' already exists in this project. The fix is to include the placeholders {date} and {time} in the name, or to omit the name and let ZenML generate one.

You can also fetch results from Python instead of the dashboard, which is how you use a trained model later in a notebook or a service:

PYTHON
from zenml.client import Client

client = Client()
run = client.get_pipeline("training_pipeline").last_run
model = run.steps["train_model"].outputs["model"][0].load()
print(type(model))

There are two details to notice. outputs["model"] returns a list of artifact versions, which is why the [0] is there. This has been the behaviour since 0.92, and older examples that index without the list no longer work. And .load() reads the artifact from the store and gives you back the real Python object, here a fitted scikit-learn model.

An even shorter route, when you just want the latest version of a named artifact, is Client().get_artifact_version("model").load().

Watch out. Every run writes new artifact versions. A project that runs a pipeline hundreds of times with large dataframes fills the store. Learn where your artifact store lives now, and delete old experiments deliberately rather than waiting for a full disk.

Try it

Run zenml pipeline runs list --output json and read the JSON. Then use the Python snippet above to load the model from your last run and call model.predict() on a few rows of the iris data. You have just used a pipeline result outside the pipeline.

Caching, parameters and settings

If you run python run.py twice in a row, the second run is noticeably faster, and the output shows lines saying that steps were cached. ZenML caches every step by default. Before running a step it computes a cache key from the step's code, its parameters and its input artifacts. If an earlier successful run used an identical key, ZenML reuses the stored outputs instead of executing the step again. This is the feature that makes iterating on a long pipeline bearable: change only the last step and only the last step reruns.

Caching helps most when steps are deterministic. It is wrong for steps whose answer depends on something outside the key, such as the current time, a live database, or a random seed you did not pass in. For those, turn caching off for that step:

PYTHON
@step(enable_cache=False)
def fetch_latest_prices() -> Annotated[pd.DataFrame, "prices"]:
    ...

You can also set it for a whole pipeline with @pipeline(enable_cache=False), or for a single run with training_pipeline.with_options(enable_cache=False)(). The more specific setting wins, so a cached pipeline can still contain one uncached step.

Try the experiment: run the pipeline, then change n_estimators from 100 to 200 and run again. load_data is cached, because nothing it depends on changed, while train_model and evaluate rerun. To change a parameter without editing the function default, pass it when you call the pipeline:

PYTHON
training_pipeline(n_estimators=300)

Pipeline parameters must be JSON-serializable: numbers, strings, booleans, lists and dictionaries of these. You cannot pass a dataframe as a pipeline parameter. Data flows between steps as artifacts, and small settings flow in as parameters.

For anything beyond a couple of parameters, put them in a YAML configuration file instead of in Python. Create config.yaml:

config.yaml
enable_cache: true
run_name: "iris_{date}_{time}"
parameters:
  n_estimators: 200
steps:
  train_model:
    parameters:
      n_estimators: 50

and run the pipeline with it:

PYTHON
training_pipeline.with_options(config_path="config.yaml")()

The precedence is worth memorizing, from strongest to weakest: values you set in Python at call time, then step-level YAML, then pipeline-level YAML, then the defaults in the code. In this file, train_model would use 50 even though the pipeline-level value says 200, because the step-level entry is more specific. Files like this let you keep experiment settings in version control next to the code, so a reviewer can see exactly what changed between two runs.

Two other switches come up quickly. enable_step_logs=False stops ZenML from storing a step's printed output, useful when a step prints something huge or sensitive. And retry=StepRetryConfig(max_retries=3, delay=10, backoff=2), imported from zenml.config.retry_config, makes a step try again after a failure, for flaky work such as a network download. Retries are per step, not per pipeline.

Try it

Run the pipeline twice and note which steps are cached the second time. Then change n_estimators and see which steps rerun. Finally add enable_cache=False to load_data and observe that it now runs every time. Check whether the downstream steps rerun too, and explain why, given that the cache key includes the identity of the input artifacts.

Stacks: changing where the same code runs

So far everything ran on the default stack. The promise at the start of this guide was that the same code can run elsewhere. The stack is how. Inspect what you have:

BASH
zenml stack list
zenml stack describe

zenml stack list shows every stack and marks the active one. zenml stack describe lists its components and their flavors. Each component type has its own command group, named after the type, with the same verbs everywhere. You can list the flavors a component type supports:

BASH
zenml orchestrator flavor list
zenml artifact-store flavor list

Registering a new component and a new stack looks like this. The example adds an S3 artifact store, which is where real teams put their artifacts. You would need the S3 integration installed and valid AWS credentials, so read it now and try it later:

BASH
zenml artifact-store register my_s3_store --flavor=s3 --path=s3://my-bucket
zenml stack register cloud_stack -o default -a my_s3_store
zenml stack set cloud_stack

The short flags on stack register name the component slot: -o is the orchestrator and -a the artifact store. Others are -c for a container registry, -i for an image builder, -s for a step operator and -e for an experiment tracker. Here -o default reuses the existing local orchestrator, so the stack keeps running steps on your machine but stores artifacts in S3. zenml stack set makes it the active stack, and the next python run.py uses it. You did not touch run.py.

That is the mental model to hold onto. The pipeline file is the stable part; the stack is the knob. A team commonly keeps a default stack for laptops, a staging stack with a Kubernetes orchestrator and a bucket in one cloud region, and a production stack with its own bucket and registry. Because stacks are named and stored on the server, everyone on the team sees the same definitions.

A few limits matter right now. A stack that contains remote components, such as a Kubernetes orchestrator or an S3 store, needs a remote ZenML server so that the machines doing the work can reach the metadata. If you try it against the local server on your laptop, you get RuntimeError: Stacks with remote components such as remote orchestrators and step operators require a remote ZenML server. Also, a remote orchestrator cannot use a local artifact store or a local container registry, because a pod on another machine cannot see a folder on your disk. The component validators say so plainly, for example that a component "is a local stack component and will not be available in the Kubernetes pipeline step". Deploying a server and a cloud stack is the natural next step after this guide. The command zenml stack deploy -p aws can provision one for you, and Kubernetes is the most common place a remote orchestrator runs.

Two other component types are worth a mention because beginners meet them first. An experiment tracker, for example MLflow, records metrics and parameters in a tracking server, and ZenML can attach one to a stack so steps log to it. See the MLflow guide. A step operator runs a single heavy step, such as training, on specialised hardware while the rest of the pipeline stays on the orchestrator.

Watch out. Stacks can reference components that point to local paths. Before you share a stack with a teammate or run it on a cluster, run zenml stack describe and check that no artifact store or container registry is a local one.

Try it

Run zenml stack list and zenml stack describe. Then register a second local stack with a new name using the existing default components, set it active, run your pipeline, and set default active again. Confirm in the dashboard that both runs show which stack they used.

Running in a container: DockerSettings

When a pipeline runs on your laptop's local orchestrator, steps run straight in your Python environment. The moment you use a remote orchestrator, each step runs in a container image that ZenML builds for you. That image must contain your code and the packages your steps import. You tell ZenML what to install with DockerSettings.

PYTHON
from zenml import pipeline
from zenml.config import DockerSettings

docker_settings = DockerSettings(
    requirements=["pandas", "scikit-learn"],
    apt_packages=["git"],
)

@pipeline(settings={"docker": docker_settings})
def training_pipeline(n_estimators: int = 100):
    ...

The keys you will use first are requirements, a list of pip packages or the path to a requirements file, and apt_packages, for system libraries. You can also set parent_image, the base image to start from, for example python:3.11-slim, and environment, a dictionary of environment variables to set inside the container. Since version 0.85 ZenML installs Python packages inside the image with uv by default, which is faster than pip. The python_package_installer setting lets you choose pip if you need to.

The same settings can live in your YAML file:

config.yaml
settings:
  docker:
    parent_image: python:3.11-slim
    requirements:
      - pandas
      - scikit-learn

You do not need any of this while you use the local orchestrator. Knowing it exists prevents the most common remote failure: a step that works on your laptop and then dies in the container with ModuleNotFoundError, because the package was installed on your machine but never listed in the image. The fix is always the same, add it to requirements.

A related everyday setting is resources. ResourceSettings asks for CPU, memory or a GPU for a step, which only takes effect on orchestrators that can honour it:

PYTHON
from zenml import step

@step(settings={"resources": {"cpu_count": 2, "memory": "4Gb"}})
def train_big_model() -> None:
    ...

If you are new to containers, read the Docker guide before you go remote, because nearly every confusing remote error is really a container question.

Try it

Add a DockerSettings object with a requirements list to your pipeline and run it locally. Nothing changes on the local orchestrator, and that is the lesson: the settings describe the container that a remote orchestrator will build, and you can write them before you ever leave your laptop.

The everyday commands, grouped by what you are trying to do

Once the concepts are in place, daily use is a short list of commands. Grouped by intent:

Connect and check.

BASH
zenml login --local          # start and connect to a local server
zenml login https://my-zenml-server.example.com   # connect to a team server
zenml logout                 # disconnect from the current server
zenml status                 # where am I connected, which stack, which project
zenml version

For a remote server, zenml login <URL> opens a browser window for you to sign in, a flow called a device flow. The resulting token is stored on your machine.

Look at what exists.

BASH
zenml stack list
zenml stack describe [NAME]
zenml pipeline list
zenml pipeline runs list --pipeline=training_pipeline
zenml artifact list
zenml model list
zenml integration list

zenml model list shows entries in the Model Control Plane, a view that groups the artifacts, runs and metadata that belong to one model. You create one by giving a pipeline or a step a model setting. It is optional for beginners, and a useful idea to meet early: the model as the thing you track over time, rather than a pile of files.

Change the active stack.

BASH
zenml stack set <NAME>

Install tool support.

BASH
zenml integration install sklearn -y --uv
zenml integration list

Reset. The following command is destructive. It wipes the local ZenML metadata and stores, so you can start over:

BASH
zenml clean

Use it only on a throwaway local setup. After a clean, run zenml init and zenml login --local again.

Change the output format of any list. Add --output json, yaml, csv or tsv, or set the default once with the environment variable ZENML_DEFAULT_OUTPUT. You can also choose columns, for example zenml pipeline runs list --columns id,index,name --output json.

Get help. Every command group has --help: zenml stack --help, zenml artifact-store register --help. The help text is generated from the code, so it is always the truth for the version you installed. When a tutorial's flag does not exist, trust --help.

From Python, the Client gives you the same information programmatically, which is how you build tools on top of ZenML. You already used it to fetch a run and load an artifact from it.

The dashboard is the best place to browse. On a pipeline run page you can see the graph, click a step to read its logs, open an artifact to see its preview or visualization, and use the timeline to find the slow step.

Try it

Without the dashboard, use only the command line to answer three questions about your project: how many runs does training_pipeline have, what are the names of your artifacts, and which stack is active? Write the exact commands you used.

Secrets and credentials: the beginner's safe habits

Sooner or later a step needs a password or API key, for a database, a cloud service or an experiment tracker. The beginner mistake is to paste it into the code, where it lands in version control and in the recorded run. ZenML has a secret store on the server for this, and a simple habit avoids almost every leak.

Create a secret from the command line:

BASH
zenml secret create my_database --username=admin --password=change-me-now

This stores a group of key-value pairs under one name. In interactive mode, zenml secret create my_database -i, ZenML prompts for the values so they do not appear in your shell history. By default a secret is public to everyone allowed to use the server. Add --private to make it visible only to you.

Read it inside a step, at run time, not at import time:

PYTHON
from zenml import step
from zenml.client import Client

@step
def query_database() -> None:
    secret = Client().get_secret("my_database")
    password = secret.secret_values["password"]
    ...

You can also reference a secret in a component's configuration with the syntax {{my_database.password}}, which lets a stack component such as an experiment tracker use the secret without the value ever appearing in your code. List secrets with zenml secret list.

Three habits protect you. Never commit a .env file or a key to Git. Never print a secret, because step logs are stored. And remember that the local server stores secrets in its own database; a shared team server should have an encryption key configured by whoever runs it, which is a senior-level concern covered in the later parts of this series.

For cloud access, ZenML goes further with service connectors, objects that hold cloud credentials on the server and hand your components short-lived, limited tokens. You do not need them to learn the basics, and they become important the day you connect an S3 bucket from a team server.

Tip. If you ever paste a real credential into a file by accident, rotate it immediately instead of deleting the line. Git history and stored step logs keep copies.

Try it

Create a secret with a fake password, then write a step that reads it with Client().get_secret and prints only the secret's key names, never the values. Run the pipeline and confirm from the step's logs that no value leaked.

Common errors and how to read them

ZenML errors tend to fall into a handful of families. Reading the first and last lines of the message, rather than the middle of a long traceback, usually tells you which family you are in.

The local extra is missing.

TEXT
ImportError: It seems like you've installed the `zenml` package without the `local` extra, but are trying to use ZenML with a local database.

The bare package cannot hold a local database. Run pip install 'zenml[local]', or connect to a server with zenml login.

The local server is gone.

TEXT
RuntimeError: Error initializing rest store with URL 'http://127.0.0.1:8237': HTTPConnectionPool(host='127.0.0.1', port=8237): Max retries exceeded with url: /api/v1/login ... [Errno 61] Connection refused

The local server process ended, usually because the machine restarted. Run zenml login --local again. The number in the last part of the message is the operating system's refusal code, and it simply means nothing is listening on that port.

A remote stack with a local server.

TEXT
RuntimeError: Stacks with remote components such as remote orchestrators and step operators require a remote ZenML server.

The stack you activated needs machines elsewhere to reach the server. Switch back with zenml stack set default, or deploy a real server and run zenml login <URL>.

Client and server versions differ.

TEXT
Your ZenML client version (X) does not match the server version (Y). This version mismatch might lead to errors or unexpected behavior.

Treat this as a real instruction, not noise. Several ZenML releases changed the API in ways that break older clients. Install the client version that matches the server, or upgrade the server. When you use a Kubernetes orchestrator, the version in the container must match exactly.

No materializer for your type.

TEXT
No materializer is registered for type `<T>`, so the default Pickle materializer was used. Pickle is not production ready...

This is a warning, and the run continues. It means ZenML did not know how to store your object and used cloudpickle. Return a supported type, install the integration that adds the right materializer, or write a custom one. A related message says the artifact was materialized under a different Python version, which is what pickle fragility looks like in practice.

A duplicate run name.

TEXT
Pipeline run name '<name>' already exists in this project. Each pipeline run must have a unique name.

You fixed the run_name in code or YAML. Add {date} and {time}, or remove it.

A stack component is missing.

TEXT
AttributeError: 'NoneType' object has no attribute 'name'

This appears when a step asks for, say, the experiment tracker through experiment_tracker.name and the active stack has none. Register a tracker and add it with zenml stack update -e <name>.

A missing secret. A StackValidationError before a run starts usually means a component refers to a secret or a key inside it that does not exist. Create it, or run zenml stack register-secrets, which prompts you for missing ones.

The local database is locked.

TEXT
sqlite3.OperationalError: database is locked

Two processes wrote to the SQLite database at once. On a laptop, wait and retry. If you truly need parallel writers, use a server backed by MySQL instead of the local SQLite file.

ModuleNotFoundError inside a container. The package was not listed in DockerSettings(requirements=...). Add it.

For everything else, turn up the logging and read what ZenML says it is doing:

BASH
export ZENML_LOGGING_VERBOSITY=DEBUG
python run.py

And when you ask for help, include the output of zenml info -a -s, zenml status and zenml stack describe. People can answer in one round instead of four.

Try it

Cause two of these errors on purpose and read them. Stop your local server by running zenml logout --local, then run the pipeline and find the connection error. Then pass the same run_name in a config file twice. For each, write down the family of the error, the cause and the one-line fix.

Putting it all together: a small end-to-end project

Now combine everything into one project you could show in an interview. The goal is a pipeline with a clear structure, a config file for experiments, cached steps where sensible, named artifacts, and a short script that uses the result. Lay out the folder:

TEXT
zenml-first/
  .zen/                 created by zenml init
  steps/
    __init__.py
    data.py
    model.py
  pipelines/
    __init__.py
    training.py
  configs/
    experiment.yaml
  run.py
  requirements.txt

Splitting steps by concern keeps each file short. Put load_data in steps/data.py, and train_model and evaluate in steps/model.py. Each file carries the same decorators and annotations you wrote earlier. The pipeline lives in pipelines/training.py:

pipelines/training.py
from zenml import pipeline

from steps.data import load_data
from steps.model import evaluate, train_model


@pipeline
def training_pipeline(n_estimators: int = 100):
    X_train, X_test, y_train, y_test = load_data()
    model = train_model(X_train, y_train, n_estimators=n_estimators)
    evaluate(model, X_test, y_test)

The entry point, run.py, picks up the configuration from the command line so that a colleague can run an experiment without editing Python:

run.py
import argparse

from pipelines.training import training_pipeline

if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--config", default="configs/experiment.yaml")
    args = parser.parse_args()

    training_pipeline.with_options(config_path=args.config)()

And the configuration file records the experiment:

configs/experiment.yaml
enable_cache: true
run_name: "iris_{date}_{time}"
parameters:
  n_estimators: 200
settings:
  docker:
    requirements:
      - pandas
      - scikit-learn

Run the project from its top folder so imports resolve from the source root you created with zenml init:

BASH
python run.py
python run.py --config configs/experiment.yaml

Then use the result outside the pipeline. This small script loads the most recent accuracy and model, which is what a downstream service or a notebook would do:

load_result.py
from zenml.client import Client

client = Client()
run = client.get_pipeline("training_pipeline").last_run

accuracy = run.steps["evaluate"].outputs["accuracy"][0].load()
model = run.steps["train_model"].outputs["model"][0].load()

print(f"Last accuracy: {accuracy:.3f}")
print(model.predict([[5.1, 3.5, 1.4, 0.2]]))

Finally, review the run in the dashboard and answer four questions, which are exactly the ones ZenML exists to make easy. Which code ran? Which parameters did it use? Where is the model stored? Which earlier run can I compare it to? If you can answer all four from the dashboard in under a minute, your pipeline is doing its job.

To extend it for practice, try these in order: add a step that writes a confusion matrix as a returned value, change the dataset size and watch which steps rerun, register a second stack and run against it, and add a step that raises an error to see how a failed run is displayed. Each one teaches a different corner of the tool.

Try it

Build the project exactly as laid out, run it twice, and make a third run with a different n_estimators in the YAML file. Use load_result.py to print the accuracy of the last run, and then find the same number in the dashboard.

What you can now do, and what comes next

You can now install ZenML and know why the extras exist. You can start a local server, explain what zenml init does, and recover when the server disappears after a reboot. You can write steps with typed, named outputs, compose them into a pipeline, run it, and find every artifact in the dashboard or through the Python client. You understand that artifacts are stored rather than passed in memory, that caching is on by default and keyed on code, parameters and inputs, and that a stack separates your code from your infrastructure. You can read the most common error messages and pick the right fix.

What this guide deliberately left for the next levels:

  • Real infrastructure. Deploying a team server, registering a Kubernetes or cloud orchestrator, connecting cloud storage through service connectors, and running steps on GPUs.
  • Production habits. Schedules, pipeline snapshots and deployments, hooks that send alerts, materializers for your own types, and testing pipelines.
  • Platform ownership. Upgrading servers safely, securing them, and deciding when ZenML is the right tool and when a simpler orchestrator is enough.

The mid-level guide starts from exactly where you are. For neighbouring tools, MLflow pairs naturally with ZenML's experiment tracker component, Kubernetes is where most remote stacks end up, and Docker explains the images ZenML builds for you.

Reading the dashboard like a debugger

The dashboard is more than a list of runs. Learning to read it quickly is what turns ZenML from a recorder into a debugging tool, so it is worth a section of its own.

Start on the Pipelines page, which lists each pipeline by name. Opening one shows its runs with a status: running, completed or failed. A completed run may contain cached steps, and the graph marks those steps so you can tell what actually executed. When a run fails, the failing step is highlighted in red. Click it and read its logs, the printed output and the traceback, before you look anywhere else. Nine times out of ten the error is right there, and it is an ordinary Python error from your own function.

The run graph shows steps as boxes and artifacts as the nodes between them. Click an artifact and the dashboard shows its name, its version, its type, the materializer that stored it and where it lives in the artifact store. For dataframes and similar types ZenML can show a preview or a visualization. This is where the naming habit pays off: a graph full of X_train, model and accuracy reads like a sentence, while a graph full of output does not.

Use the lineage view to answer the question that started this guide, which model came from which data. Follow the arrows backwards from the model artifact to the step that created it, then to the artifacts that step consumed, and so on to the original data load. Each artifact version is immutable, so the answer stays true even after you run the pipeline a hundred more times.

The Stacks page shows which components each stack contains, which is the quickest way to spot a local component hiding inside a stack you meant to share. The Settings area holds users, service accounts and, on the server, secrets. You will not need most of it on day one.

When two runs give different accuracy, compare their parameters side by side, then their step code versions, then their input artifacts. Differences almost always appear in one of those three, and that is exactly why ZenML records them. A useful habit is to put a short, meaningful run_name pattern in your YAML file, such as iris_{date}_{time}, so that the list is readable without opening each run.

Try it

Add a line to evaluate that raises an error, for example dividing by zero, and run the pipeline. In the dashboard, find the failed step, read its log, and identify the exact line. Remove the error and confirm that the next run succeeds and that the earlier steps are cached.

Sources

A dynamic pipeline, declared with @pipeline(dynamic=True), differs at the first stage. It is not compiled ahead of time. Its body runs at runtime in the orchestration environment, so ordinary Python if and for statements can decide what runs next, and .map() can fan out over however many items an earlier step produced. You pay for that flexibility in orchestrator support, which a later section covers. Reach for it when the shape of the work genuinely depends on data, not because it feels more Pythonic.

Try it Run a pipeline on the default stack with ZENML_LOGGING_VERBOSITY=DEBUG set. In the output, find the point where compilation ends and execution begins. Then run zenml pipeline runs list --columns id,index,name --output json and confirm the run you just made is at the top, since lists have been newest first since 0.95.

Artifacts, materializers and the pickle trap

An artifact is a value written to the artifact store by a materializer, a subclass of BaseMaterializer that knows how to serialize and deserialize one Python type. ZenML ships materializers for common types: built-ins, NumPy arrays, pandas dataframes, and the types of many integrations. It picks one by the declared type of the step output. That is why type annotations on steps are not decoration at this level. They choose the serialization format.

When no materializer matches, ZenML falls back to CloudpickleMaterializer and logs a warning: No materializer is registered for type ..., so the default Pickle materializer was used. Pickle is not production ready. Take that warning seriously for three reasons. Pickles are tied to the Python version that wrote them, so an artifact written on 3.11 may not load on 3.13, and ZenML warns you on load when it notices. They are opaque, so you cannot inspect the data without loading it. And unpickling can execute arbitrary code, which matters if anyone other than you can write to the artifact store. Since 0.96.3 ZenML stores a SHA-256 hash with the pickle and validates it on load, which catches tampering after the fact but does not make pickle a format you would choose.

Set ZENML_ENFORCE_TYPE_ANNOTATIONS=True in CI and unannotated steps fail instead of quietly taking the fallback path.

Writing a materializer is short. You declare which types it handles and which artifact type it produces, then implement load and save:

materializers.py
import json
import os
from typing import Any, ClassVar, Tuple, Type

from zenml.enums import ArtifactType
from zenml.io import fileio
from zenml.materializers.base_materializer import BaseMaterializer


class Thresholds:
    def __init__(self, values: dict[str, float]):
        self.values = values


class ThresholdsMaterializer(BaseMaterializer):
    ASSOCIATED_TYPES: ClassVar[Tuple[Type[Any], ...]] = (Thresholds,)
    ASSOCIATED_ARTIFACT_TYPE: ClassVar[ArtifactType] = ArtifactType.DATA

    def load(self, data_type: Type[Thresholds]) -> Thresholds:
        path = os.path.join(self.uri, "thresholds.json")
        with fileio.open(path, "r") as f:
            return Thresholds(json.load(f))

    def save(self, data: Thresholds) -> None:
        path = os.path.join(self.uri, "thresholds.json")
        with fileio.open(path, "w") as f:
            json.dump(data.values, f)

Use fileio.open rather than Python's open, because self.uri may be an s3:// or gs:// path and fileio is the layer that speaks to every artifact store. Attach the materializer either by defining it with the class being handled, which registers it automatically when the module is imported, or explicitly on a step with @step(output_materializers=ThresholdsMaterializer).

Two naming details cause confusion. An unnamed output is called output, or output_0, output_1 for several, and the dashboard shows it as {pipeline_name}::{step_name}::output. Give outputs real names with Annotated[T, "name"], because that name is how you fetch them later. Also, a step returns multiple outputs only when its return statement is a tuple literal. A function annotated -> Tuple[int, int] that returns a variable holding a tuple produces a single artifact. Finally, since 0.97.0, a dict with non-string keys no longer goes through the JSON path, where keys used to be turned into strings silently. If another system reads your dict artifacts as JSON, use string keys.

When you fetch artifacts back through the client, remember that outputs["name"] has been a list of artifact versions since 0.92. That is why the idiom is .outputs["model"][0].load().

Pickle works on your laptop and breaks in the pipeline A step that returns a custom object passes every local test, then fails weeks later when a different Python version in the orchestrator image cannot load it. If the warning about the Pickle materializer appears in a log, treat it as a failing check, not a notice.
Try it Write a step that returns an instance of a class you defined. Run it, read the warning, then add a materializer like the one above and confirm the warning is gone and the artifact version in the dashboard shows a readable file under its URI.

Caching, precisely

Caching is on by default and is the most cost-effective feature in the tool, so it is worth understanding exactly. The problem it solves is re-running an unchanged step, and the danger is that it can hide a change you did make.

A step is skipped, and its previous outputs reused, when its cache key matches a previous successful run. By default the key includes the step's source code, its parameters, the IDs of its input artifacts, and the values of those artifacts. You control the parts through CachePolicy, imported from zenml.config. Its include_step_code, include_step_parameters, include_artifact_values and include_artifact_ids fields all default to True. You can add the source of code the step depends on with source_dependencies=[...], set a custom key function with cache_func (a zero-argument function returning a string), or make cached results expire with expires_after, given in seconds.

The most useful of these for real work is cache_func. A step that reads a table from a warehouse has identical code, parameters and inputs on Monday and on Tuesday, so the default key matches and it reuses Monday's data. The fix is to put something that changes with the data into the key:

steps/load.py
from datetime import date

import pandas as pd
from zenml import step
from zenml.config import CachePolicy


def partition_key() -> str:
    return date.today().isoformat()


@step(cache_policy=CachePolicy(cache_func=partition_key))
def load_daily() -> pd.DataFrame:
    return pd.read_parquet("s3://my-bucket/events/latest.parquet")

The alternative is enable_cache=False, which is correct for steps that genuinely must always run, such as a step that posts to an external system. Using it everywhere throws away the feature, and using it nowhere causes stale data. Decide per step.

There is one subtle interaction with code. The key includes the step's source, so editing a helper function that the step calls does not invalidate the cache unless you list it in source_dependencies. If a fix in a shared module seems not to take effect, that is the first thing to check. Caching also interacts with downstream steps: because the key includes input artifact IDs, a step that re-runs and produces a new artifact version invalidates every consumer, while a cache hit reuses the old artifact version and lets the consumers hit too. The effect is that cache hits propagate down the graph, and so do misses.

See what was cached In the dashboard a cached step is marked as such, and in the run's step list its status reads as cached. After a code change, run twice. If the second run is not entirely cached, something is putting non-determinism into a key: a timestamp in a parameter, or an artifact that changes each run.
Try it Add a random number as a step parameter default and watch the step never hit the cache. Then replace it with a fixed value and confirm the second run is fully cached. Finally, give one step a CachePolicy with expires_after=60, wait a minute, and confirm it re-runs.

Stacks, components and flavors in practice

A stack is a named set of components, and every stack needs at least an orchestrator and an artifact store. The value is that switching where code runs is a configuration change and not a code change. The discipline is deciding what belongs in a stack and what belongs in code.

The component types in 0.97.0 are orchestrator, artifact_store, container_registry, image_builder, step_operator, experiment_tracker, model_registry, model_deployer (deprecated), deployer, alerter, annotator, data_validator, feature_store, log_store and sandbox. Each type has flavors, the concrete implementations. Ask the CLI what is installed:

BASH
zenml orchestrator flavor list
zenml artifact-store flavor list
zenml integration list

Many flavors live in integrations that need extra packages. zenml integration install <name> -y --uv installs them, and zenml stack export-requirements <stack> prints what a given stack needs, which is the right input for a CI image.

A realistic remote stack on AWS looks like this:

BASH
zenml artifact-store register s3_store --flavor=s3 --path=s3://my-ml-bucket
zenml container-registry register ecr --flavor=aws \
  --uri=123456789012.dkr.ecr.me-south-1.amazonaws.com
zenml orchestrator register k8s --flavor=kubernetes
zenml stack register prod -o k8s -a s3_store -c ecr --set

The short flags on zenml stack register are -o for the orchestrator, -a for the artifact store, -c for the container registry, -i for the image builder, -s for a step operator, -e for an experiment tracker, -r for a model registry, -D for a deployer and -l for a log store. Change a stack later with zenml stack update prod -e mlflow_tracker, and take a component out with zenml stack remove-component prod -e.

Two rules keep remote stacks from failing confusingly. First, a stack with remote components needs a remote server: if you point a laptop using only the local SQLite store at it, you get Stacks with remote components such as remote orchestrators and step operators require a remote ZenML server. Second, every component a remote orchestrator uses must itself be reachable from the cluster. The Kubernetes orchestrator's validation says so directly, with messages such as is a local stack component and will not be available in the Kubernetes pipeline step and points to a local container registry. A local artifact store or a localhost registry can never work, because the pod cannot see your disk.

Stacks are data, so treat them like code. zenml stack export prod -f prod.yaml and zenml stack import let you keep them in version control and rebuild them elsewhere, zenml stack copy clones one, and zenml stack register-secrets prompts for any secret a stack references but the server lacks. For a first cloud stack, zenml stack deploy -p aws|gcp|azure provisions infrastructure and registers the stack in one step, which is a good way to see a correct layout before writing your own Terraform. The Terraform guide covers managing that infrastructure properly.

Try it Export your current stack with zenml stack export, open the YAML, and find each component's flavor and configuration. Then run zenml stack describe and compare. If you keep stacks for several environments, commit those files.

Service connectors and the secrets store

Every remote component needs credentials, and the question is where they live. The naive answer is environment variables on each developer's machine and each pod, which leaks, rots and cannot be revoked in one place. ZenML offers two mechanisms that fix this, and they do different jobs.

A service connector is an object stored on the server that holds a credential for AWS, GCP, Azure, Kubernetes, Docker, HyperAI or OAuth2 (added in 0.96). Its point is that clients and workloads never receive the long-lived credential. The server uses it to issue short-lived, down-scoped credentials on demand, such as an STS token limited to one bucket. A component is linked to a connector with --connector.

BASH
zenml service-connector register aws-conn --type aws --auto-configure
zenml service-connector verify aws-conn
zenml service-connector list-resources --resource-type s3-bucket

zenml artifact-store connect s3_store --connector aws-conn
zenml orchestrator register k8s --flavor=kubernetes \
  --connector aws-conn --resource-id my-eks-cluster

The authentication method matters. AWS supports implicit, secret-key, sts-token, iam-role, session-token and federation-token, and zenml service-connector describe-type aws lists them with their requirements. Prefer iam-role, which gives each client a temporary session from a role the long-lived key can assume, over handing out a raw secret key. Implicit authentication, where the connector picks up whatever credentials the server process happens to have, is disabled by default on the server. Enable it only deliberately, with ZENML_ENABLE_IMPLICIT_AUTH_METHODS=true or the Helm value enableImplicitAuthMethods, because it lets any user of the server act with the server's own identity.

A secret is different: a named group of key-value pairs held in the server's secrets store, for values that are not cloud credentials such as an MLflow password or an API key. You create one with zenml secret create mlflow_secret --username=admin --password=abc123, and you reference it from component configuration as {{mlflow_secret.username}}, which the server substitutes at use time, so the value never sits in a stack export. Secrets are public by default, meaning governed by RBAC, or private with --private, visible only to their creator.

Inside a step you can read one with Client().get_secret("name").secret_values["key"], or declare secrets=["name"] on the step or pipeline so the values arrive as environment variables. Before each run ZenML validates that the referenced secrets exist. The level is set by ZENML_SECRET_VALIDATION_LEVEL, one of NONE, SECRET_EXISTS or SECRET_AND_KEY_EXISTS, which is the default.

The default backend stores secrets in the server's SQL database, and one setting deserves an alarm bell. Set ZENML_SECRETS_STORE_ENCRYPTION_KEY, for example from openssl rand -hex 32, or the secrets are stored unencrypted. If you lose the key they cannot be recovered, so keep it somewhere you back up. For regulated workloads you can point the store at AWS Secrets Manager, GCP Secret Manager, Azure Key Vault or HashiCorp Vault with ZENML_SECRETS_STORE_TYPE, which also keeps secret material out of the ZenML database entirely. The Vault guide explains the Vault side.

A stale Docker login can override the connector In 0.96.0 an invalid entry in your local Docker credential helper took precedence over connector credentials, producing denied: Your authorization token has expired. It was fixed in 0.96.2. On any version, docker logout <REGISTRY> is the quick check when a push fails although the connector verifies fine.
Try it Register a connector, run zenml service-connector verify on it, and then list its resources. Link your artifact store to it and run a pipeline with no cloud environment variables set in your shell. If it works, the connector is doing its job.

Configuring a run: YAML and precedence

Configuration scattered across decorators, with_options calls and environment variables becomes impossible to reason about, so ZenML lets you move run configuration into a YAML file and defines a strict order of precedence for the layers. From highest to lowest: runtime Python code, then step-level YAML, then pipeline-level YAML, then the defaults in your code. When a setting does not take effect, the answer is nearly always that something higher in this list overrides it.

configs/prod.yaml
enable_cache: true
run_name: "train_{date}_{time}"
parameters:
  lr: 0.01
settings:
  docker:
    parent_image: python:3.11-slim
    requirements: requirements.txt
    apt_packages: ["git"]
  resources:
    cpu_count: 2
    memory: "4Gb"
steps:
  train:
    parameters:
      lr: 0.001
    settings:
      resources:
        gpu_count: 1
        memory: "16Gb"
model:
  name: churn_classifier
  tags: [prod]

You apply it with training.with_options(config_path="configs/prod.yaml")(), or from the command line with zenml pipeline run run.training --config configs/prod.yaml. If you want a template to start from, zenml pipeline build-configuration my_pipeline > config.yaml generates one.

Keep one file per environment and let the code stay environment-agnostic. The run_name placeholders {date} and {time} are not cosmetic: run names must be unique within a project, and a fixed name fails the second run with Pipeline run name '...' already exists in this project. Environment variables can be interpolated in the Docker environment block as ${MY_API_KEY}, which keeps values out of the file, although a secret reference is the better route for anything sensitive.

The same layering applies to step settings. Settings have keys, "docker", "resources", "orchestrator", and flavor-specific ones such as "experiment_tracker.mlflow". A step-level entry overrides the pipeline-level entry for that step only. That is how you give one training step a GPU and keep the other steps on small CPU pods. The same mechanism picks a step operator for one step with step_operator: vertex_gpu, which sends that step to Vertex while the rest stay on the cluster. See the Vertex AI guide and the SageMaker guide for what those backends then provide.

Try it Put a parameter in three places: the code default, the pipeline YAML, and a step-level YAML block. Run and read which value the step received, then pass it at runtime too. Write down the order you observed and check it against the list above.

Docker settings and image builds

On a remote orchestrator every step runs in a container, and ZenML builds that image for you from DockerSettings. The problem it solves is reproducibility: the pod must have your code and exactly your dependencies. The problem it creates is build time, which is often the slowest part of a pipeline run.

run.py
from zenml import pipeline
from zenml.config import DockerSettings

docker = DockerSettings(
    parent_image="python:3.11-slim",
    requirements="requirements.txt",
    apt_packages=["git"],
    python_package_installer="uv",
    python_package_installer_cache_mount=True,
)


@pipeline(settings={"docker": docker})
def training():
    ...

You can give a list or a file path for requirements, point at pyproject_path, ask for required_integrations, or set replicate_local_python_environment to freeze your current environment, which is convenient and usually a mistake, because it copies in everything on your laptop. The install order is: local environment, then stack requirements, then required_integrations, then pyproject, then requirements. A later source can pin a version an earlier one installed, so if a version surprises you, read the list as an order of events. The default installer has been uv since 0.85, and python_package_installer_cache_mount adds a BuildKit cache mount for it, which makes a rebuild after a one-line dependency change fast. If you use uv with a parent image that has no active virtual environment, pass python_package_installer_args={"system": None}.

The fastest build is no build. With skip_build=True ZenML uses your parent_image directly, and it then requires that image to already contain Python, pip, zenml, and your code in /app. That suits a mature project where CI builds the image once, tests it, and the pipeline just runs it. The alternative is code delivery without a rebuild: allow_download_from_code_repository lets the pod fetch code from a registered GitHub or GitLab repository, and allow_download_from_artifact_store uploads it to the artifact store instead. Both leave dependencies in the image and move only your source at runtime, so a code-only change costs nothing to build. If you do neither, ZenML can include files in the image itself (allow_including_files_in_images, on by default).

On Apple Silicon a locally built image is arm64 and will not run on an amd64 cluster, and layer caching will not work unless you set the platform explicitly: DockerSettings(build_config={"build_options": {"platform": "linux/amd64"}}). The general mechanics of layers, cache keys and multi-stage builds are in the Docker guide, and they apply unchanged, since ZenML's builder produces ordinary Docker images.

Inspect before you pay zenml pipeline build builds the images for a pipeline without running it, and zenml pipeline builds list shows what exists. Run the build in CI first and failures in requirements surface before any cluster time is spent.
Try it Time a cold build and a rebuild after changing one line of code. Then turn on python_package_installer_cache_mount and change one line of requirements.txt. The difference between the second and third timings is what caching buys you.

Running on Kubernetes

The Kubernetes orchestrator is the most common remote target, because a team that already has a cluster does not need a second system. It is also where most of the subtle failures live, so it is worth knowing how the pieces fit.

When you run a pipeline the client builds images, then creates an orchestrator pod in the cluster. That pod walks the DAG and starts each step as a Kubernetes Job, which has been the behaviour since 0.84. Each step pod pulls its own image, loads its inputs from the artifact store and writes its outputs back. Nothing has to be shared between pods except the artifact store and the server, which is why both must be reachable from inside the cluster.

BASH
zenml service-connector register eks-conn --type aws --auto-configure
zenml orchestrator register k8s --flavor=kubernetes \
  --connector eks-conn --resource-id my-eks-cluster
zenml stack update prod -o k8s

The rule that catches people is version matching. With this orchestrator, the ZenML version in your client and the version in the orchestrator pod must be exactly the same, because the pod runs the same entrypoint code your client compiled against. A cached image built last month with an older zenml therefore produces version-mismatch failures. Pin the version in your requirements and rebuild when you upgrade. A more general warning, Your ZenML client version (X) does not match the server version (Y), has the same root cause, since 0.83, 0.90 and 0.94 each broke compatibility between versions.

Resource requests come from ResourceSettings on the step, cpu_count, gpu_count, memory, and node placement and pod-level settings such as tolerations and service accounts go into the orchestrator's settings, which you can set per step in the YAML. A step that needs more than one pod can use the multi-pod Kubernetes step operator, which has had a pod_count setting since 0.96.3. Since 0.96.3, pods for a run that has already ended are no longer retried, so the old Run is already finished. noise is gone.

For failures, distinguish two layers. A step that raises an exception is a ZenML-level failure, visible in the run's step logs. A pod that never starts, because of ImagePullBackOff, an unschedulable resource request or a quota, is a Kubernetes-level failure and will not appear in ZenML's logs at all. Read it with kubectl describe pod and kubectl get events, and compare your requests with what the node pool can provide. The Kubernetes guide covers that diagnosis in depth.

Try it Run a two-step pipeline on a dev cluster and watch it with kubectl get pods -w. You should see the orchestrator pod first, then one step pod per step. Then deliberately set a memory request larger than any node and observe the pod stay Pending while ZenML shows the run as still running.

Scheduling and running from CI

A pipeline that only runs when someone types a command is a script, and the next step is having something else run it. There are two routes at this level: schedules the orchestrator owns, and runs that your CI system starts.

A schedule is attached with with_options:

run.py
from zenml.config.schedule import Schedule

training.with_options(schedule=Schedule(cron_expression="0 3 * * *"))()

Whether a schedule is actually honoured depends on the orchestrator. Local, local_docker, SkyPilot and Tekton do not support scheduling. Among the others, the important distinction is that native schedule management, where updating or deleting the ZenML schedule also updates the backend, exists only for Kubernetes, which uses CronJobs. On every other orchestrator, deleting a schedule in ZenML leaves the job running in the backend until you remove it there, and a schedule you forgot about keeps spending money. The Kubernetes settings concurrency_policy (Allow, Forbid or Replace) and starting_deadline_seconds decide what happens when a run overlaps the next tick, and Forbid is the safe default for a training job that writes to shared state.

BASH
zenml pipeline schedule list
zenml pipeline schedule update <id> --cron-expression='0 4 * * *'
zenml pipeline schedule deactivate <id>
zenml pipeline schedule delete <id> --hard

Plain delete archives a schedule and --hard removes it. After a server upgrade, delete and recreate schedules, since they are tied to the server version like snapshots.

For CI, the question is identity. A pipeline run needs to authenticate to the server, and the wrong answer is a developer's personal login. Create a service account, which prints an API key exactly once, and give that key to the CI system as a secret:

BASH
zenml service-account create ci-bot
zenml service-account api-key ci-bot rotate <key> --retain 60

The CI job then sets ZENML_STORE_URL, ZENML_STORE_API_KEY and ZENML_ACTIVE_PROJECT_ID, or calls zenml login https://<server> --api-key. Rotation with --retain 60 keeps the old key valid for sixty minutes so you can roll it across jobs. On an OSS server without RBAC, only admins can manage service accounts and API keys, a change from 0.96.0. A workflow in GitHub Actions that installs the pinned ZenML version, sets those variables and runs python run.py --config configs/prod.yaml is enough for most teams, and Argo CD or another GitOps tool can own the stack and cluster definitions around it.

Pipeline snapshots are the other route. A snapshot is an immutable capture of the DAG, code, configuration and images, created with zenml pipeline snapshot create run.my_pipeline --name v1 --stack prod. You can deploy one as a service in the open-source edition, but running a snapshot from the dashboard, CLI or API is a ZenML Pro feature, and so are triggers. Snapshots replaced run templates in 0.90, and zenml pipeline create-run-template still prints a deprecation warning. Because a snapshot is tied to a server version, rebuild them as part of every server upgrade.

Try it Create a service account on a test server, export its key into a shell with only the three ZENML_STORE_* variables, and run a pipeline from a clean virtual environment. That is exactly what your CI job does.

Deploying and operating the server

The local daemon from Beginner is for one person. A team needs a shared server, and operating it is mostly about the database and about not breaking the version contract.

The server is a FastAPI app with the dashboard, backed by a SQL database. SQLite is for development only. For anything shared use MySQL 8.0 or later. Remember that the server holds metadata and nothing else, so it is small and stateless apart from the database. The two supported routes are Docker for a single host and the Helm chart for Kubernetes:

BASH
docker run -it -d -p 8080:8080 --name zenml zenmldocker/zenml-server:0.97.0

Always pin the image tag to your client version, and note that without any ZENML_STORE_* variable the database is an ephemeral SQLite file inside the container, so data disappears with it.

BASH
helm -n zenml install zenml-server oci://public.ecr.aws/zenml/zenml \
  --version 0.97.0 --values custom-values.yaml
custom-values.yaml
zenml:
  database:
    url: "mysql://zenml:CHANGEME@mysql.internal:3306/zenml"
  ingress:
    enabled: true

Do not guess the exact key layout for ingress and secrets: take the chart's values.yaml for the version you install, and keep your overrides in one file in Git. The chart also exposes autoscaling values (autoscaling.enabled, minReplicas, maxReplicas), and a recent addition lets you render an HTTPRoute from server.gateway.enabled.

Once you run more than one replica, three settings matter. Set a shared ZENML_SERVER_JWT_SECRET_KEY, because otherwise each replica generates its own random key and a token issued by one is rejected by another. Use the same secrets-store encryption key everywhere. And size the database connection pool against the thread pool: the server thread pool defaults to 40 (ZENML_SERVER_THREAD_POOL_SIZE), and database.poolSize plus database.maxOverflow should be at least that.

Upgrades are where operators get hurt. The rules are firm. Downgrades are not supported. If an upgrade goes wrong, the error Revision not found in the database init logs means a newer schema was applied and an older server was started, and the only way back is a restored backup. Keep the client and server on the same version. Recreate snapshots and schedules after each upgrade. The server takes its own backup before each schema migration, controlled by ZENML_STORE_BACKUP_STRATEGY, but those backups are a safety net for the migration, not a retention policy. Run your own scheduled database backups. A sound pattern is a staging server that mirrors production: upgrade it first, run a smoke-test pipeline, then promote. The Helm guide covers the mechanics of --reuse-values and value files.

Behind a load balancer, remember that since 0.95 login rate limiting no longer trusts a raw X-Forwarded-For header. If you switch rate limiting on with ZENML_SERVER_RATE_LIMIT_ENABLED=1, configure the proxy headers, or every user appears to come from the ingress.

Losing the encryption key loses the secrets Back up ZENML_SECRETS_STORE_ENCRYPTION_KEY separately from the database. A restored database without its key is a table of unreadable blobs.
Try it Start the Docker server with a bind-mounted data directory and a generated encryption key, register a secret, stop and recreate the container, and check that the secret is still readable. Then repeat without the key to see what you would have lost.

Hooks, retries and failure handling

Real pipelines fail for boring reasons: a flaky API, a spot instance reclaimed, a network blip to the artifact store. The aim is for the boring failures to heal themselves and the real ones to tell you loudly.

Retries are per step, configured with StepRetryConfig from zenml.config.retry_config:

steps/fetch.py
from zenml import step
from zenml.config.retry_config import StepRetryConfig


@step(retry=StepRetryConfig(max_retries=3, delay=10, backoff=2))
def fetch_features() -> dict[str, float]:
    ...

The delay is in seconds and backoff multiplies it on each attempt. There is no pipeline-level retry, which is deliberate: a step is the unit you can safely repeat. Retry only steps whose repeat is harmless. A step that appends rows to a table and then fails on the network will append twice, so make such steps idempotent before adding retries.

Hooks run code at points in a step's or pipeline's life. The lifecycle since 0.95 is on_start, on_end, on_success and on_failure, with on_pause and on_resume for dynamic runs and on_init and on_cleanup as setup and teardown. Their scope differs between pipeline types. On a static pipeline, a pipeline-level hook is a default inherited by every step. On a dynamic pipeline, it fires once per run. Decide which you mean, because a failure alert attached at pipeline level on a static pipeline will send one message per failing step.

hooks.py
from zenml import pipeline, step
from zenml.hooks import alerter_failure_hook, alerter_success_hook


@step
def train() -> float:
    ...


@pipeline(on_failure=alerter_failure_hook, on_success=alerter_success_hook)
def training():
    train()

This requires an alerter in your stack, Slack or Discord, configured with a secret for its token. Two behaviours are worth knowing. A hook that raises is swallowed: the exception never fails the run, and the failure is stored as a HookInvocation with status FAILED, so a broken alert hook fails silently unless you look. You can query those with Client().list_hook_invocations(...). Hooks for on_failure and on_end can also take an optional BaseException argument if you want the error in the message.

A related rule came in 0.96.2. start_after became a reserved step keyword, so a step parameter with that name must be renamed before you upgrade. A scan for reserved names before an upgrade is cheaper than discovering it in production.

For controlling the blast radius of a failure, choose the execution mode deliberately. CONTINUE_ON_FAILURE lets independent branches finish, which is what you want when one of ten parallel evaluations fails and the other nine are worth keeping. STOP_ON_FAILURE and FAIL_FAST stop sooner. Remember the earlier warning that only some orchestrators honour them.

Try it Write a step that fails on its first two attempts using a counter file, give it StepRetryConfig(max_retries=3, delay=1), and confirm the run succeeds. Then add an on_failure hook that itself raises, and find the failed HookInvocation afterwards.

Testing pipelines without a stack

Pipelines are hard to test because a run needs a server, a stack and often a cluster. Most of what you want to test is not ZenML at all, it is your function, and ZenML makes that easy to reach.

A step decorated with @step can still be called through its original function with my_step.entrypoint(...), which bypasses ZenML completely. This makes the unit test for a step an ordinary test:

tests/test_steps.py
import pandas as pd

from steps.features import build_features


def test_build_features_drops_nulls():
    raw = pd.DataFrame({"age": [31, None, 45], "spend": [10.0, 5.0, 7.5]})
    out = build_features.entrypoint(raw)
    assert out["age"].notna().all()
    assert len(out) == 2

If you want the real step object executed without a stack, set ZENML_RUN_SINGLE_STEPS_WITHOUT_STACK=true. That is mostly for exploratory work. Keep tests at three levels and be honest about what each proves. Unit tests on .entrypoint prove your logic. A pipeline test on the default local stack, with a tiny dataset, proves the graph is wired correctly, that types line up, and that materializers handle what you return. A scheduled or release-time run against a staging stack proves the infrastructure. The second level is the one teams skip, and it catches the most embarrassing failures: a misnamed output, a step that returns a tuple variable instead of a tuple literal and so yields one artifact, a parameter that is not JSON-serializable.

Give the level-two test its own isolated environment. A local SQLite store is fine for a single process, but mapped dynamic steps and parallel test runners writing concurrently can hit sqlite3.OperationalError: database is locked, for which 0.96.3 raised the lock wait to 60 seconds. If your CI runs tests in parallel, serialize the ZenML ones or point them at a server backed by MySQL.

BASH
export ZENML_ENFORCE_TYPE_ANNOTATIONS=True
export ZENML_LOGGING_VERBOSITY=INFO
pytest tests/ -x

Also test the things that only fail when you deploy. Validate your YAML by loading it in CI with a dry build, using zenml pipeline build, which catches requirements conflicts and bad settings keys. Pin the ZenML version in the test environment to the one in production. A check that compares zenml.__version__ to the server version fails the build early, which is far cheaper than the warning that appears after you have already deployed. Data tests belong in the same suite. If your pipeline uses Great Expectations or Evidently through the data validator component, run those checks on a fixed sample as part of the pipeline test, so a schema change fails a build, not a model.

Try it Pick your most complex step, move its body into a plain function, call that function from the step, and write a test for it. Then add one pipeline-level test that runs the whole thing on a five-row dataset.

Logs, observability and debugging

When something fails you want to know which layer failed before you read any logs. The layers are the client, the server, the orchestrator, and your step code, and each has its own place to look.

For your own code, step logs are captured and shown in the dashboard per step. Where they are stored is decided by the log store component, added in 0.93, with artifact store (the default), OpenTelemetry and Datadog flavors. The default keeps them as files next to your artifacts, which is fine until you want to search across runs, at which point an OTel or Datadog log store puts them in the same place as the rest of your telemetry. Logging can be turned off per step with enable_step_logs=False, or globally with ZENML_DISABLE_STEP_LOGS_STORAGE=true, useful when a step prints sensitive data. Since 0.97.0 the REST API returns logs through a paginated, filterable entries endpoint instead of one large blob, so any script of yours that read the old response has to change.

For the server, zenml logs -f follows a local daemon, docker logs zenml -f a container, and on Kubernetes:

BASH
kubectl -n zenml logs -l app.kubernetes.io/name=zenml
kubectl -n zenml logs -l app.kubernetes.io/name=zenml -c zenml-db-init

The second form reads the database initialization container, which is where database connection problems are visible. Access denied for user there means wrong credentials, Can't connect to MySQL server means a wrong host or network path, and a direct client connection must use 127.0.0.1, not localhost. For the client, export ZENML_LOGGING_VERBOSITY=DEBUG is the first move, and ZENML_CONSOLE_LOGGING_FORMAT=json makes console output parseable in CI. When you need to ask for help, or file an issue, gather zenml info -a -s, zenml status and zenml stack describe.

Try it Break a pipeline three ways on purpose: a bad secret reference, a step that raises, and an unschedulable pod. For each, write down which layer's logs showed the failure. Expect the third to be invisible in ZenML and visible in kubectl describe.

Dynamic pipelines and step operators

Two features widen what a pipeline can express, and both have sharp edges that the marketing does not mention.

Dynamic pipelines let the body run at runtime. They support fan-out with .map() (and .product() for a cartesian product), broadcasting with unmapped(), concurrent execution with .submit(), and child pipelines. The distinction to hold onto is between .load() and .chunk(). .load() brings the actual data into the orchestrating process so your Python can decide something, while .chunk() passes a reference as a DAG edge without loading it:

dynamic.py
from zenml import pipeline, step


@step
def list_shards() -> list[str]:
    return ["a", "b", "c"]


@step
def process(shard: str) -> int:
    return len(shard)


@pipeline(dynamic=True)
def fan_out():
    shards = list_shards()
    process.map(shard=shards)

Mapping works only over artifacts produced in the same run, and the chunk size for mapping is fixed at one. An isolated step, @step(runtime="isolated"), runs in its own container, and works only on Kubernetes, Vertex, SageMaker or AzureML. runtime="inline" runs in the orchestrator environment. On the local and local_docker orchestrators only inline steps are supported, and choosing isolated there fails. Each isolated step is a fresh container, so a fan-out of a thousand is a thousand containers, which is a cost line as well as a feature.

Step operators solve a different problem: most of a pipeline is cheap, and one step needs a GPU. A step operator sends a single step to specialised compute such as SageMaker, Vertex, AzureML, Modal or a multi-pod Kubernetes job, while the rest stay on your orchestrator. You attach one per step with @step(step_operator="vertex_gpu") or in YAML. Custom operators that must work in dynamic pipelines have to implement submit_step and get_step_status, and the Spark operator is not dynamic-compatible.

Since 0.95, CommandStep runs a shell command or self-contained Python as a step, which is useful to wrap an existing training script without rewriting it. It has no inputs, outputs, step context or hooks, and ZenML does not track its logs. In a static pipeline it requires a step operator.

Finally, long-running or interactive flows can use zenml.wait(), which pauses the run and frees compute. On OSS a paused run must be resumed manually, with zenml pipeline runs resume <id>, while Pro resumes automatically.

Try it Write the fan-out above, run it locally, then change process to runtime="isolated" and read the error on the local orchestrator. It is the quickest way to remember the rule.

Working with the neighbouring tools

ZenML is glue, and its value shows in how it sits next to the tools you already run. The integrations are stack components, so adding one means registering a component and updating the stack, not rewriting pipelines.

Experiment tracking. The experiment tracker component supports MLflow, Weights and Biases, Neptune, Comet and Trackio. The pattern with MLflow is to register the tracker with its credentials as secret references and then enable it on a step:

BASH
zenml integration install mlflow -y --uv
zenml secret create mlflow_secret --username=admin --password=abc123
zenml experiment-tracker register mlflow_tracker --flavor=mlflow \
  --tracking_uri=https://mlflow.internal \
  --tracking_username={{mlflow_secret.username}} \
  --tracking_password={{mlflow_secret.password}}
zenml stack update prod -e mlflow_tracker

Inside a step you then write @step(experiment_tracker="mlflow_tracker") and call MLflow's own logging API. ZenML does not replace the tracker, it links each MLflow run to the ZenML run so you can move between the two views. Note that a local MLflow with no tracking_uri set defaults to a SQLite backend since 0.95. See also the guides for Weights and Biases and Comet.

Orchestration neighbours. If you already run Airflow or Kubeflow, ZenML can target them as orchestrators, so pipelines keep their ZenML lineage while the existing scheduler runs them. Airflow 3.0 has been supported since 0.85. Whether that is wise is a question of who owns the scheduler: use it when the platform team mandates one scheduler.

Data quality and features. The data validator component supports Evidently, Great Expectations and Deepchecks, and the feature store component supports Feast. Data versioning tools such as DVC and lakeFS overlap with ZenML artifact versioning, so decide which owns the dataset of record. A common split is to let the data tool own raw data and have ZenML store derived artifacts.

Serving. Model deployers are deprecated since 0.91. The current route is the deployer component, which runs a whole pipeline as an HTTP service in online mode, with local, docker, kubernetes, gcp (Cloud Run), aws (App Runner) and huggingface flavors. zenml pipeline deploy weather_agent.weather_pipeline --name my_dep provisions one, and zenml deployment invoke calls it. For dedicated model serving with autoscaling you may still prefer KServe or BentoML, with ZenML producing the model artifact and a registry entry.

Try it Add one integration to your stack, an alerter or an experiment tracker, and use it from one step. Count the lines of pipeline code you changed. It should be close to zero besides the decorator argument.

Putting it all together

Here is one small project that uses most of what this level covered, in the order you would set it up.

Start with the repository. It holds steps/, pipelines/, configs/dev.yaml and configs/prod.yaml, a requirements.txt with ZenML pinned, and tests/. Steps are thin adapters over plain functions. Every step output has a name and a type that has a materializer, and ZENML_ENFORCE_TYPE_ANNOTATIONS=True runs in CI, so pickle cannot slip in.

Next the platform. A MySQL-backed ZenML server runs from the Helm chart, pinned to the same version as the client, with a generated secrets-store encryption key kept in your secrets manager and a shared JWT key for the replicas. The production stack registers an S3 artifact store, an ECR registry and a Kubernetes orchestrator, all reached through an AWS service connector using the iam-role method. The stack definition is exported to YAML and committed. The MLflow credentials live in a ZenML secret, referenced from the tracker configuration.

Then the run path. A CI workflow installs the pinned client, authenticates with a service account's API key, runs the unit tests with .entrypoint, runs the local-stack pipeline test on a tiny sample, builds the image once with zenml pipeline build, and finally runs python run.py --config configs/prod.yaml against the production stack. Caching is on. A warehouse-reading step has a cache_func keyed on the partition date. The training step has a retry policy, and on_failure alerts Slack through the alerter. The schedule runs on Kubernetes CronJobs with concurrency_policy: Forbid.

Then the operating loop. Logs flow to an OTel log store and the server exports traces to the platform collector, where alerts watch database CPU. Upgrades go to a staging server first, followed by a smoke-test pipeline, then production, after which snapshots and schedules are recreated. Database backups run on their own schedule, separate from the automatic pre-migration backup.

Try it Write your own list of the ten decisions above for a project you own, and mark which are already true. The unmarked ones are your backlog, and the first two or three (pinned versions, a named-output convention, an encryption key) are usually the cheapest.

Where you are now

You can now predict what a pipeline call will do: compile, build, orchestrate, materialize and report. You know why a step re-ran or did not, and you can shape the cache key. You can assemble a remote stack from components, give it credentials through connectors and secrets without leaking them, and keep the stack definition in Git. You can build images quickly, run on Kubernetes without version skew, schedule and trigger runs from CI with a service account, and operate a server that you can back up and upgrade without being afraid of it.

The next level, Senior, treats ZenML as a platform for a team: dynamic pipelines at scale, snapshots and deployments as a delivery mechanism, multi-tenancy, the trust model, and when ZenML is the wrong tool. Before you move on, make sure you have run a pipeline on a remote stack and broken it on purpose in each of the layers above, because that is what makes the next level's failure modes familiar.

Sources