This is part one of three. It covers everything you need to do real work with ZenML 0.97, not a teaser. By the end you can install it, write a pipeline out of steps, run it, see every intermediate result stored and versioned, control caching, read the dashboard, swap the infrastructure underneath your code without rewriting it, and diagnose the errors beginners meet first. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas only stick once you have watched your own pipeline run, skip a step because of the cache, and fail in a way you can read.
What ZenML is, and the problem it solves
ZenML is an open-source Python framework for building machine learning workflows that run the same way on your laptop and on production infrastructure. You write ordinary Python functions, mark them with a decorator, and ZenML records what went in, what came out, which code produced it, and where everything is stored. Then you can point the very same code at Kubernetes, SageMaker, Vertex AI or another backend by changing a configuration object instead of the code.
To see why this matters, think about how a machine learning project usually starts. Someone writes a notebook. It loads a CSV, cleans it, trains a model and prints an accuracy. It works, so the team wants to run it every night on fresh data, on a bigger machine, with a GPU, and to know next month exactly which data and code produced the model that is now in production.
The notebook cannot do any of that. Cells depend on hidden state, so running them in a different order gives a different answer. Intermediate results live in memory and vanish when the kernel dies. Moving to a bigger machine means rewriting the code around a scheduler, a container and a storage bucket. Six months later nobody can say which version of the data trained the model, because nothing recorded it.
The usual fix is to glue together a workflow scheduler such as Airflow, an experiment tracker such as MLflow, a container build, and a pile of scripts that move files around. That works, but the glue is the hard part, and it ties the code to one scheduler. ZenML sits above those tools rather than replacing them. It gives you one small Python interface for describing the work, and it talks to schedulers, container registries and trackers on your behalf.
Three consequences of that design shape everything that follows.
Your code describes the work, not the infrastructure. A step is a plain Python function. It does not import a Kubernetes client or a cloud SDK. Where it runs is decided by something separate, the stack, which you can change without touching the function.
Every output is stored, not passed around in memory. When one step returns a dataframe and the next step receives it, ZenML has written the dataframe to storage in between and read it back. That sounds wasteful until you notice what it buys you: the result survives a crash, can be inspected later, can be reused by a different run, and carries a version and a lineage.
ZenML is a metadata layer, not a data warehouse. A small server keeps records about runs. Your actual data lives in storage you own: a local folder on day one, an S3 bucket or Google Cloud Storage bucket later. That matters for data residency. If your employer in the Gulf or in Egypt needs the data to stay in a particular region, the artifacts stay in the bucket you chose in the region you chose, and only metadata goes to the server.
ZenML is licensed under Apache-2.0. There is a commercial product, ZenML Pro, which adds team features such as role-based access control, but everything in this guide works with the free open-source version.
Try it
Think of one notebook you have written. On paper, list its cells, and for each one write what it needs as input and what it produces. Each cell with a clear input and output is a candidate step, and you have just sketched your first pipeline.
The core mental model: steps, pipelines, artifacts and stacks
ZenML has a small vocabulary. Learn these nouns well and the documentation becomes easy to read.
A step is a Python function decorated with @step. Its inputs and outputs have type annotations, and ZenML uses those annotations to decide how to store each output. One step should do one job: load data, clean it, train, evaluate.
A pipeline is a function decorated with @pipeline that calls steps and passes the output of one into the next. The pipeline body does not do the work itself. It describes the order and the data flow. When you call a pipeline, ZenML first works out the graph of steps, a directed acyclic graph or DAG, and then executes it. This is the default, called a static pipeline. There is also a newer dynamic pipeline, written @pipeline(dynamic=True), whose body runs at runtime so you can use ordinary loops and conditions. It is an advanced tool, so this guide stays with static pipelines.
An artifact is anything a step returns. Each artifact is written to the artifact store, versioned, and linked to the run and step that created it. Downstream steps read it from the store. By default an output is called output, or output_0, output_1 and so on when a step returns several. You can give it a proper name with Annotated[type, "name"], which you will do in every real project.
A materializer is the small piece of code that knows how to write one kind of Python object to storage and read it back. ZenML ships materializers for integers, strings, lists, dictionaries, NumPy arrays and pandas dataframes, and integrations add more, for example scikit-learn models. When nothing matches, ZenML falls back to cloudpickle. That fallback works but is not safe for production: pickles can break across Python versions and can run arbitrary code when loaded, so treat the warning that mentions it as a prompt to fix something.
A stack is a named set of infrastructure pieces, called stack components, that a pipeline runs on. Every stack needs at least two components. The orchestrator decides what runs when and where. The artifact store is where artifacts are written. Others, such as a container registry or an experiment tracker, are optional. The stack you get on day one is called default. It pairs a local orchestrator, which runs steps on your own machine, with a local artifact store, which is a folder on disk.
A flavor is a concrete implementation of a component type. The local orchestrator and the kubernetes orchestrator are two flavors of the orchestrator type. The local artifact store and the s3 artifact store are two flavors of the artifact store type.
A pipeline run is one execution of a pipeline. Each run records its status, its parameters, the artifacts it produced and its logs. Finally, the ZenML server is a small web application with a database and a dashboard. It stores the records of runs, artifacts and stacks, and your client talks to it over HTTP.
zenml package and CLI you installedKeep one sentence in mind: code says what to do, the stack says where, and the server remembers what happened. When something confuses you later, ask which of the three it belongs to.
Try it
Without looking back, write the four nouns in order from smallest to largest: step, artifact, pipeline, stack. Then state in one sentence what the orchestrator and the artifact store each do. Check yourself against the section above.
Installing ZenML and checking the setup
ZenML 0.97, the version this guide targets, supports Python 3.10 through 3.14. Use a virtual environment so that ZenML's dependencies stay separate from everything else on your machine. On Linux or macOS:
python3 -m venv .venv
source .venv/bin/activate
pip install 'zenml[server]'
The square-bracket part is an extra, an optional bundle of dependencies. It matters more than it looks. Since version 0.90 the bare pip install zenml installs only the client, which can talk to a deployed server but cannot run a database of its own. You choose an extra depending on what you want:
| You want | Install |
|---|---|
| Pure local use with a SQLite database and no server process | pip install 'zenml[local]' |
| A local server and dashboard on your laptop | pip install 'zenml[server]' |
| Nicer output inside Jupyter notebooks | pip install 'zenml[jupyter]' |
| Only a client that connects to a team server | pip install zenml |
If you install the bare package and then try to use it without a server, you will see an error that begins like this:
ImportError: It seems like you've installed the `zenml` package without the `local` extra, but are trying to use ZenML with a local database.
The message itself tells you both fixes: install zenml[local], or run zenml login to connect to a server. This is the single most common first error, because older tutorials say pip install zenml and nothing more.
If you prefer the uv tool, uv venv and uv pip install 'zenml[server]' work the same way.
macOS on Apple Silicon. Before you start a local server, set this variable once per shell:
export OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES
Without it the local server process can crash on the newer macOS fork checks. You do not need it if you only connect to a remote server.
Windows. The ZenML FAQ says Windows is officially supported only through WSL, the Windows Subsystem for Linux. Some commands work natively, but anything that starts a server process does not. Install WSL, open an Ubuntu shell, and follow the Linux steps. This is worth doing even if it feels like a detour, because almost all ML infrastructure targets Linux.
Now verify the installation. Each of these commands answers a different question:
zenml version
python -c "import zenml; print(zenml.__version__)"
zenml status
zenml stack describe
zenml version prints the installed version, for example 0.97.0. The Python one-liner confirms that the package imported by your interpreter is the same one the command line uses, which catches the classic mistake of installing into one environment and running from another. zenml status reports whether you are connected to a server, which stack is active and which project you are in. zenml stack describe prints the components of the active stack. On a fresh install you should see the default stack with a local orchestrator and a local artifact store.
When you need to ask someone for help, run zenml info -a -s and paste the output. It gathers the versions and settings that maintainers ask for first.
Try it
Create a fresh virtual environment, install zenml[server], and run the four verification commands. Write down the ZenML version and the name of the active stack. If zenml status says you are not connected to anything yet, that is expected. The next section fixes it.
Your first server, and a project folder
A ZenML client keeps its records in a server. For learning, the server runs on your own laptop. Two commands set everything up. Inside a new project folder, run:
mkdir zenml-first && cd zenml-first
zenml init
zenml login --local
zenml init marks the current folder as the source root of your project. It creates a small hidden .zen directory. ZenML uses the source root to work out how to import your code, and later, to decide which files to package when a pipeline runs in a container. Run it once at the top of each project.
zenml login --local starts a local server as a background process and prints the address of its dashboard, by default on port 8237. The first time you open the dashboard in a browser it asks you to set up an administrator account. Afterwards your client is connected to that server automatically.
Two facts about the local server save an hour of confusion later. First, it does not survive a reboot or a shutdown. After you restart your machine, run zenml login --local again. If you forget, the next pipeline run fails with an error that ends in Connection refused and mentions 127.0.0.1:8237. Second, older tutorials tell you to run zenml up, zenml down and zenml connect. Those commands still exist in 0.97 but are deprecated and print warnings. The current forms are zenml login --local, zenml logout --local and, for a remote server, zenml login https://your-server.
If you have Docker installed and prefer to keep the server in a container, zenml login --local --docker does the same thing inside one. Learn Docker first if the word container is new, using the Docker guide.
Run zenml status again. It should now show that you are connected to a local server. Open the dashboard and look at the empty Pipelines page. By the end of this guide it will be full.
zenml status first. It tells you in one line whether you are connected, so you can tell a forgotten zenml login --local from a real bug.
Try it
Run zenml init and zenml login --local in a new folder. Open the dashboard address it prints, create the administrator account, and find the Stacks page. Confirm that a stack named default exists. Then close the terminal, open a new one and run zenml status to see that the connection is remembered.
Your first pipeline, step by step
We will build the smallest useful pipeline: it loads a dataset, trains a classifier, and reports its accuracy. First install the scikit-learn integration, which also adds the materializer that knows how to store scikit-learn models:
zenml integration install sklearn -y --uv
An integration is a bundle of extra packages and ZenML components for one external tool. The -y flag skips the confirmation prompt, and --uv installs with the faster uv installer. If you do not have uv, leave that flag off.
Now create a file named run.py. We build it in three pieces so each idea is visible. First, the imports and the data step:
from typing import Annotated, Tuple
import pandas as pd
from sklearn.base import ClassifierMixin
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from zenml import pipeline, step
@step
def load_data() -> Tuple[
Annotated[pd.DataFrame, "X_train"],
Annotated[pd.DataFrame, "X_test"],
Annotated[pd.Series, "y_train"],
Annotated[pd.Series, "y_test"],
]:
"""Load the iris dataset and split it into train and test parts."""
iris = load_iris(as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(
iris.data, iris.target, test_size=0.2, random_state=42
)
return X_train, X_test, y_train, y_test
Read the decorator and the return annotation closely, because this is where ZenML does its work. @step tells ZenML that load_data is a step. The return type is a Tuple of four Annotated values. Each Annotated[type, "name"] says: this output has this Python type and this artifact name. ZenML uses the type to choose a materializer, for example the pandas one for a dataframe, and the name to label the artifact in the dashboard.
A rule that trips people up: a step has several outputs only when its return statement is a tuple, as it is here. If you annotate a Tuple but return a single object, ZenML treats the whole thing as one output.
The second piece is the training step:
@step
def train_model(
X_train: pd.DataFrame, y_train: pd.Series, n_estimators: int = 100
) -> Annotated[ClassifierMixin, "model"]:
"""Train a random forest and return it."""
model = RandomForestClassifier(n_estimators=n_estimators, random_state=42)
model.fit(X_train, y_train)
return model
The inputs X_train and y_train come from artifacts produced by another step. n_estimators is a parameter: a simple JSON-friendly value such as a number or string that ZenML records with the run. The return annotation names the trained model model. The type ClassifierMixin is the scikit-learn base class for classifiers, and the scikit-learn integration's materializer handles it.
The third piece is an evaluation step and the pipeline that connects everything:
@step
def evaluate(
model: ClassifierMixin, X_test: pd.DataFrame, y_test: pd.Series
) -> Annotated[float, "accuracy"]:
"""Score the model on held-out data."""
accuracy = accuracy_score(y_test, model.predict(X_test))
print(f"Test accuracy: {accuracy:.3f}")
return float(accuracy)
@pipeline
def training_pipeline(n_estimators: int = 100):
X_train, X_test, y_train, y_test = load_data()
model = train_model(X_train, y_train, n_estimators=n_estimators)
evaluate(model, X_test, y_test)
if __name__ == "__main__":
training_pipeline()
Inside training_pipeline you see ordinary-looking function calls, but they do not execute the steps immediately. When ZenML compiles the pipeline, each call returns a placeholder that stands for the future output of that step. Passing it to the next step is how you draw an edge in the graph. That is why you cannot do arbitrary Python on those placeholders inside the pipeline body, such as printing the dataframe or calling len() on it. The real values exist only when the steps run.
Run it:
python run.py
ZenML prints the run name, the step each stage is executing, and a link to the dashboard. You will see lines showing that load_data, train_model and evaluate each ran, your printed accuracy, and the message that the pipeline run finished. The accuracy for this dataset is typically high, near 1.0, because iris is an easy problem.
Notice what you did not write: no file paths, no serialization code, no database calls. Yet every output is now stored and recorded. Open the dashboard, click Pipelines, click training_pipeline, then the run. You will see the graph of three steps with the artifacts between them. Click the model artifact to see its version and where it is stored.
Try it
Type in the three pieces, run python run.py, and open the run in the dashboard. Click each step and find its logs, its output artifacts and the accuracy printed by evaluate. Then change test_size to 0.3 and run again. Compare the two runs in the list.
What just happened: runs, artifacts and the local store
It helps to know exactly where everything went, so that ZenML never feels like magic. Run these commands:
zenml pipeline list
zenml pipeline runs list
zenml artifact list
Since version 0.95, list commands show the newest items first. That is a change from earlier releases, and old blog posts may describe the opposite order. zenml pipeline list shows training_pipeline. zenml pipeline runs list shows each run with a name and a status. You can filter to one pipeline with --pipeline=training_pipeline, and you can change the output format with --output, which accepts table, json, yaml, csv and tsv. The JSON form is handy for scripts.
zenml artifact list shows the named artifacts you created: X_train, X_test, y_train, y_test, model and accuracy. If you had not used Annotated, each would appear with an unhelpful default name such as training_pipeline::train_model::output. Naming them costs one line and saves real confusion later.
The data itself lives in the local artifact store, a folder inside ZenML's configuration directory on your machine. You do not need to go there, and you should not edit it by hand, but knowing that a real file exists for every artifact demystifies the dashboard. When you later switch to an S3 bucket, the same artifacts simply land in the bucket.
Each run is given a name automatically, made from the pipeline name and a timestamp. Run names must be unique within a project. If you pass your own run_name and reuse it, you get Pipeline run name '...' already exists in this project. The fix is to include the placeholders {date} and {time} in the name, or to omit the name and let ZenML generate one.
You can also fetch results from Python instead of the dashboard, which is how you use a trained model later in a notebook or a service:
from zenml.client import Client
client = Client()
run = client.get_pipeline("training_pipeline").last_run
model = run.steps["train_model"].outputs["model"][0].load()
print(type(model))
There are two details to notice. outputs["model"] returns a list of artifact versions, which is why the [0] is there. This has been the behaviour since 0.92, and older examples that index without the list no longer work. And .load() reads the artifact from the store and gives you back the real Python object, here a fitted scikit-learn model.
An even shorter route, when you just want the latest version of a named artifact, is Client().get_artifact_version("model").load().
Try it
Run zenml pipeline runs list --output json and read the JSON. Then use the Python snippet above to load the model from your last run and call model.predict() on a few rows of the iris data. You have just used a pipeline result outside the pipeline.
Caching, parameters and settings
If you run python run.py twice in a row, the second run is noticeably faster, and the output shows lines saying that steps were cached. ZenML caches every step by default. Before running a step it computes a cache key from the step's code, its parameters and its input artifacts. If an earlier successful run used an identical key, ZenML reuses the stored outputs instead of executing the step again. This is the feature that makes iterating on a long pipeline bearable: change only the last step and only the last step reruns.
Caching helps most when steps are deterministic. It is wrong for steps whose answer depends on something outside the key, such as the current time, a live database, or a random seed you did not pass in. For those, turn caching off for that step:
@step(enable_cache=False)
def fetch_latest_prices() -> Annotated[pd.DataFrame, "prices"]:
...
You can also set it for a whole pipeline with @pipeline(enable_cache=False), or for a single run with training_pipeline.with_options(enable_cache=False)(). The more specific setting wins, so a cached pipeline can still contain one uncached step.
Try the experiment: run the pipeline, then change n_estimators from 100 to 200 and run again. load_data is cached, because nothing it depends on changed, while train_model and evaluate rerun. To change a parameter without editing the function default, pass it when you call the pipeline:
training_pipeline(n_estimators=300)
Pipeline parameters must be JSON-serializable: numbers, strings, booleans, lists and dictionaries of these. You cannot pass a dataframe as a pipeline parameter. Data flows between steps as artifacts, and small settings flow in as parameters.
For anything beyond a couple of parameters, put them in a YAML configuration file instead of in Python. Create config.yaml:
enable_cache: true
run_name: "iris_{date}_{time}"
parameters:
n_estimators: 200
steps:
train_model:
parameters:
n_estimators: 50
and run the pipeline with it:
training_pipeline.with_options(config_path="config.yaml")()
The precedence is worth memorizing, from strongest to weakest: values you set in Python at call time, then step-level YAML, then pipeline-level YAML, then the defaults in the code. In this file, train_model would use 50 even though the pipeline-level value says 200, because the step-level entry is more specific. Files like this let you keep experiment settings in version control next to the code, so a reviewer can see exactly what changed between two runs.
Two other switches come up quickly. enable_step_logs=False stops ZenML from storing a step's printed output, useful when a step prints something huge or sensitive. And retry=StepRetryConfig(max_retries=3, delay=10, backoff=2), imported from zenml.config.retry_config, makes a step try again after a failure, for flaky work such as a network download. Retries are per step, not per pipeline.
Try it
Run the pipeline twice and note which steps are cached the second time. Then change n_estimators and see which steps rerun. Finally add enable_cache=False to load_data and observe that it now runs every time. Check whether the downstream steps rerun too, and explain why, given that the cache key includes the identity of the input artifacts.
Stacks: changing where the same code runs
So far everything ran on the default stack. The promise at the start of this guide was that the same code can run elsewhere. The stack is how. Inspect what you have:
zenml stack list
zenml stack describe
zenml stack list shows every stack and marks the active one. zenml stack describe lists its components and their flavors. Each component type has its own command group, named after the type, with the same verbs everywhere. You can list the flavors a component type supports:
zenml orchestrator flavor list
zenml artifact-store flavor list
Registering a new component and a new stack looks like this. The example adds an S3 artifact store, which is where real teams put their artifacts. You would need the S3 integration installed and valid AWS credentials, so read it now and try it later:
zenml artifact-store register my_s3_store --flavor=s3 --path=s3://my-bucket
zenml stack register cloud_stack -o default -a my_s3_store
zenml stack set cloud_stack
The short flags on stack register name the component slot: -o is the orchestrator and -a the artifact store. Others are -c for a container registry, -i for an image builder, -s for a step operator and -e for an experiment tracker. Here -o default reuses the existing local orchestrator, so the stack keeps running steps on your machine but stores artifacts in S3. zenml stack set makes it the active stack, and the next python run.py uses it. You did not touch run.py.
That is the mental model to hold onto. The pipeline file is the stable part; the stack is the knob. A team commonly keeps a default stack for laptops, a staging stack with a Kubernetes orchestrator and a bucket in one cloud region, and a production stack with its own bucket and registry. Because stacks are named and stored on the server, everyone on the team sees the same definitions.
A few limits matter right now. A stack that contains remote components, such as a Kubernetes orchestrator or an S3 store, needs a remote ZenML server so that the machines doing the work can reach the metadata. If you try it against the local server on your laptop, you get RuntimeError: Stacks with remote components such as remote orchestrators and step operators require a remote ZenML server. Also, a remote orchestrator cannot use a local artifact store or a local container registry, because a pod on another machine cannot see a folder on your disk. The component validators say so plainly, for example that a component "is a local stack component and will not be available in the Kubernetes pipeline step". Deploying a server and a cloud stack is the natural next step after this guide. The command zenml stack deploy -p aws can provision one for you, and Kubernetes is the most common place a remote orchestrator runs.
Two other component types are worth a mention because beginners meet them first. An experiment tracker, for example MLflow, records metrics and parameters in a tracking server, and ZenML can attach one to a stack so steps log to it. See the MLflow guide. A step operator runs a single heavy step, such as training, on specialised hardware while the rest of the pipeline stays on the orchestrator.
zenml stack describe and check that no artifact store or container registry is a local one.
Try it
Run zenml stack list and zenml stack describe. Then register a second local stack with a new name using the existing default components, set it active, run your pipeline, and set default active again. Confirm in the dashboard that both runs show which stack they used.
Running in a container: DockerSettings
When a pipeline runs on your laptop's local orchestrator, steps run straight in your Python environment. The moment you use a remote orchestrator, each step runs in a container image that ZenML builds for you. That image must contain your code and the packages your steps import. You tell ZenML what to install with DockerSettings.
from zenml import pipeline
from zenml.config import DockerSettings
docker_settings = DockerSettings(
requirements=["pandas", "scikit-learn"],
apt_packages=["git"],
)
@pipeline(settings={"docker": docker_settings})
def training_pipeline(n_estimators: int = 100):
...
The keys you will use first are requirements, a list of pip packages or the path to a requirements file, and apt_packages, for system libraries. You can also set parent_image, the base image to start from, for example python:3.11-slim, and environment, a dictionary of environment variables to set inside the container. Since version 0.85 ZenML installs Python packages inside the image with uv by default, which is faster than pip. The python_package_installer setting lets you choose pip if you need to.
The same settings can live in your YAML file:
settings:
docker:
parent_image: python:3.11-slim
requirements:
- pandas
- scikit-learn
You do not need any of this while you use the local orchestrator. Knowing it exists prevents the most common remote failure: a step that works on your laptop and then dies in the container with ModuleNotFoundError, because the package was installed on your machine but never listed in the image. The fix is always the same, add it to requirements.
A related everyday setting is resources. ResourceSettings asks for CPU, memory or a GPU for a step, which only takes effect on orchestrators that can honour it:
from zenml import step
@step(settings={"resources": {"cpu_count": 2, "memory": "4Gb"}})
def train_big_model() -> None:
...
If you are new to containers, read the Docker guide before you go remote, because nearly every confusing remote error is really a container question.
Try it
Add a DockerSettings object with a requirements list to your pipeline and run it locally. Nothing changes on the local orchestrator, and that is the lesson: the settings describe the container that a remote orchestrator will build, and you can write them before you ever leave your laptop.
The everyday commands, grouped by what you are trying to do
Once the concepts are in place, daily use is a short list of commands. Grouped by intent:
Connect and check.
zenml login --local # start and connect to a local server
zenml login https://my-zenml-server.example.com # connect to a team server
zenml logout # disconnect from the current server
zenml status # where am I connected, which stack, which project
zenml version
For a remote server, zenml login <URL> opens a browser window for you to sign in, a flow called a device flow. The resulting token is stored on your machine.
Look at what exists.
zenml stack list
zenml stack describe [NAME]
zenml pipeline list
zenml pipeline runs list --pipeline=training_pipeline
zenml artifact list
zenml model list
zenml integration list
zenml model list shows entries in the Model Control Plane, a view that groups the artifacts, runs and metadata that belong to one model. You create one by giving a pipeline or a step a model setting. It is optional for beginners, and a useful idea to meet early: the model as the thing you track over time, rather than a pile of files.
Change the active stack.
zenml stack set <NAME>
Install tool support.
zenml integration install sklearn -y --uv
zenml integration list
Reset. The following command is destructive. It wipes the local ZenML metadata and stores, so you can start over:
zenml clean
Use it only on a throwaway local setup. After a clean, run zenml init and zenml login --local again.
Change the output format of any list. Add --output json, yaml, csv or tsv, or set the default once with the environment variable ZENML_DEFAULT_OUTPUT. You can also choose columns, for example zenml pipeline runs list --columns id,index,name --output json.
Get help. Every command group has --help: zenml stack --help, zenml artifact-store register --help. The help text is generated from the code, so it is always the truth for the version you installed. When a tutorial's flag does not exist, trust --help.
From Python, the Client gives you the same information programmatically, which is how you build tools on top of ZenML. You already used it to fetch a run and load an artifact from it.
The dashboard is the best place to browse. On a pipeline run page you can see the graph, click a step to read its logs, open an artifact to see its preview or visualization, and use the timeline to find the slow step.
Try it
Without the dashboard, use only the command line to answer three questions about your project: how many runs does training_pipeline have, what are the names of your artifacts, and which stack is active? Write the exact commands you used.
Secrets and credentials: the beginner's safe habits
Sooner or later a step needs a password or API key, for a database, a cloud service or an experiment tracker. The beginner mistake is to paste it into the code, where it lands in version control and in the recorded run. ZenML has a secret store on the server for this, and a simple habit avoids almost every leak.
Create a secret from the command line:
zenml secret create my_database --username=admin --password=change-me-now
This stores a group of key-value pairs under one name. In interactive mode, zenml secret create my_database -i, ZenML prompts for the values so they do not appear in your shell history. By default a secret is public to everyone allowed to use the server. Add --private to make it visible only to you.
Read it inside a step, at run time, not at import time:
from zenml import step
from zenml.client import Client
@step
def query_database() -> None:
secret = Client().get_secret("my_database")
password = secret.secret_values["password"]
...
You can also reference a secret in a component's configuration with the syntax {{my_database.password}}, which lets a stack component such as an experiment tracker use the secret without the value ever appearing in your code. List secrets with zenml secret list.
Three habits protect you. Never commit a .env file or a key to Git. Never print a secret, because step logs are stored. And remember that the local server stores secrets in its own database; a shared team server should have an encryption key configured by whoever runs it, which is a senior-level concern covered in the later parts of this series.
For cloud access, ZenML goes further with service connectors, objects that hold cloud credentials on the server and hand your components short-lived, limited tokens. You do not need them to learn the basics, and they become important the day you connect an S3 bucket from a team server.
Try it
Create a secret with a fake password, then write a step that reads it with Client().get_secret and prints only the secret's key names, never the values. Run the pipeline and confirm from the step's logs that no value leaked.
Common errors and how to read them
ZenML errors tend to fall into a handful of families. Reading the first and last lines of the message, rather than the middle of a long traceback, usually tells you which family you are in.
The local extra is missing.
ImportError: It seems like you've installed the `zenml` package without the `local` extra, but are trying to use ZenML with a local database.
The bare package cannot hold a local database. Run pip install 'zenml[local]', or connect to a server with zenml login.
The local server is gone.
RuntimeError: Error initializing rest store with URL 'http://127.0.0.1:8237': HTTPConnectionPool(host='127.0.0.1', port=8237): Max retries exceeded with url: /api/v1/login ... [Errno 61] Connection refused
The local server process ended, usually because the machine restarted. Run zenml login --local again. The number in the last part of the message is the operating system's refusal code, and it simply means nothing is listening on that port.
A remote stack with a local server.
RuntimeError: Stacks with remote components such as remote orchestrators and step operators require a remote ZenML server.
The stack you activated needs machines elsewhere to reach the server. Switch back with zenml stack set default, or deploy a real server and run zenml login <URL>.
Client and server versions differ.
Your ZenML client version (X) does not match the server version (Y). This version mismatch might lead to errors or unexpected behavior.
Treat this as a real instruction, not noise. Several ZenML releases changed the API in ways that break older clients. Install the client version that matches the server, or upgrade the server. When you use a Kubernetes orchestrator, the version in the container must match exactly.
No materializer for your type.
No materializer is registered for type `<T>`, so the default Pickle materializer was used. Pickle is not production ready...
This is a warning, and the run continues. It means ZenML did not know how to store your object and used cloudpickle. Return a supported type, install the integration that adds the right materializer, or write a custom one. A related message says the artifact was materialized under a different Python version, which is what pickle fragility looks like in practice.
A duplicate run name.
Pipeline run name '<name>' already exists in this project. Each pipeline run must have a unique name.
You fixed the run_name in code or YAML. Add {date} and {time}, or remove it.
A stack component is missing.
AttributeError: 'NoneType' object has no attribute 'name'
This appears when a step asks for, say, the experiment tracker through experiment_tracker.name and the active stack has none. Register a tracker and add it with zenml stack update -e <name>.
A missing secret. A StackValidationError before a run starts usually means a component refers to a secret or a key inside it that does not exist. Create it, or run zenml stack register-secrets, which prompts you for missing ones.
The local database is locked.
sqlite3.OperationalError: database is locked
Two processes wrote to the SQLite database at once. On a laptop, wait and retry. If you truly need parallel writers, use a server backed by MySQL instead of the local SQLite file.
ModuleNotFoundError inside a container. The package was not listed in DockerSettings(requirements=...). Add it.
For everything else, turn up the logging and read what ZenML says it is doing:
export ZENML_LOGGING_VERBOSITY=DEBUG
python run.py
And when you ask for help, include the output of zenml info -a -s, zenml status and zenml stack describe. People can answer in one round instead of four.
Try it
Cause two of these errors on purpose and read them. Stop your local server by running zenml logout --local, then run the pipeline and find the connection error. Then pass the same run_name in a config file twice. For each, write down the family of the error, the cause and the one-line fix.
Putting it all together: a small end-to-end project
Now combine everything into one project you could show in an interview. The goal is a pipeline with a clear structure, a config file for experiments, cached steps where sensible, named artifacts, and a short script that uses the result. Lay out the folder:
zenml-first/
.zen/ created by zenml init
steps/
__init__.py
data.py
model.py
pipelines/
__init__.py
training.py
configs/
experiment.yaml
run.py
requirements.txt
Splitting steps by concern keeps each file short. Put load_data in steps/data.py, and train_model and evaluate in steps/model.py. Each file carries the same decorators and annotations you wrote earlier. The pipeline lives in pipelines/training.py:
from zenml import pipeline
from steps.data import load_data
from steps.model import evaluate, train_model
@pipeline
def training_pipeline(n_estimators: int = 100):
X_train, X_test, y_train, y_test = load_data()
model = train_model(X_train, y_train, n_estimators=n_estimators)
evaluate(model, X_test, y_test)
The entry point, run.py, picks up the configuration from the command line so that a colleague can run an experiment without editing Python:
import argparse
from pipelines.training import training_pipeline
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("--config", default="configs/experiment.yaml")
args = parser.parse_args()
training_pipeline.with_options(config_path=args.config)()
And the configuration file records the experiment:
enable_cache: true
run_name: "iris_{date}_{time}"
parameters:
n_estimators: 200
settings:
docker:
requirements:
- pandas
- scikit-learn
Run the project from its top folder so imports resolve from the source root you created with zenml init:
python run.py
python run.py --config configs/experiment.yaml
Then use the result outside the pipeline. This small script loads the most recent accuracy and model, which is what a downstream service or a notebook would do:
from zenml.client import Client
client = Client()
run = client.get_pipeline("training_pipeline").last_run
accuracy = run.steps["evaluate"].outputs["accuracy"][0].load()
model = run.steps["train_model"].outputs["model"][0].load()
print(f"Last accuracy: {accuracy:.3f}")
print(model.predict([[5.1, 3.5, 1.4, 0.2]]))
Finally, review the run in the dashboard and answer four questions, which are exactly the ones ZenML exists to make easy. Which code ran? Which parameters did it use? Where is the model stored? Which earlier run can I compare it to? If you can answer all four from the dashboard in under a minute, your pipeline is doing its job.
To extend it for practice, try these in order: add a step that writes a confusion matrix as a returned value, change the dataset size and watch which steps rerun, register a second stack and run against it, and add a step that raises an error to see how a failed run is displayed. Each one teaches a different corner of the tool.
Try it
Build the project exactly as laid out, run it twice, and make a third run with a different n_estimators in the YAML file. Use load_result.py to print the accuracy of the last run, and then find the same number in the dashboard.
What you can now do, and what comes next
You can now install ZenML and know why the extras exist. You can start a local server, explain what zenml init does, and recover when the server disappears after a reboot. You can write steps with typed, named outputs, compose them into a pipeline, run it, and find every artifact in the dashboard or through the Python client. You understand that artifacts are stored rather than passed in memory, that caching is on by default and keyed on code, parameters and inputs, and that a stack separates your code from your infrastructure. You can read the most common error messages and pick the right fix.
What this guide deliberately left for the next levels:
- Real infrastructure. Deploying a team server, registering a Kubernetes or cloud orchestrator, connecting cloud storage through service connectors, and running steps on GPUs.
- Production habits. Schedules, pipeline snapshots and deployments, hooks that send alerts, materializers for your own types, and testing pipelines.
- Platform ownership. Upgrading servers safely, securing them, and deciding when ZenML is the right tool and when a simpler orchestrator is enough.
The mid-level guide starts from exactly where you are. For neighbouring tools, MLflow pairs naturally with ZenML's experiment tracker component, Kubernetes is where most remote stacks end up, and Docker explains the images ZenML builds for you.
Reading the dashboard like a debugger
The dashboard is more than a list of runs. Learning to read it quickly is what turns ZenML from a recorder into a debugging tool, so it is worth a section of its own.
Start on the Pipelines page, which lists each pipeline by name. Opening one shows its runs with a status: running, completed or failed. A completed run may contain cached steps, and the graph marks those steps so you can tell what actually executed. When a run fails, the failing step is highlighted in red. Click it and read its logs, the printed output and the traceback, before you look anywhere else. Nine times out of ten the error is right there, and it is an ordinary Python error from your own function.
The run graph shows steps as boxes and artifacts as the nodes between them. Click an artifact and the dashboard shows its name, its version, its type, the materializer that stored it and where it lives in the artifact store. For dataframes and similar types ZenML can show a preview or a visualization. This is where the naming habit pays off: a graph full of X_train, model and accuracy reads like a sentence, while a graph full of output does not.
Use the lineage view to answer the question that started this guide, which model came from which data. Follow the arrows backwards from the model artifact to the step that created it, then to the artifacts that step consumed, and so on to the original data load. Each artifact version is immutable, so the answer stays true even after you run the pipeline a hundred more times.
The Stacks page shows which components each stack contains, which is the quickest way to spot a local component hiding inside a stack you meant to share. The Settings area holds users, service accounts and, on the server, secrets. You will not need most of it on day one.
When two runs give different accuracy, compare their parameters side by side, then their step code versions, then their input artifacts. Differences almost always appear in one of those three, and that is exactly why ZenML records them. A useful habit is to put a short, meaningful run_name pattern in your YAML file, such as iris_{date}_{time}, so that the list is readable without opening each run.
Try it
Add a line to evaluate that raises an error, for example dividing by zero, and run the pipeline. In the dashboard, find the failed step, read its log, and identify the exact line. Remove the error and confirm that the next run succeeds and that the earlier steps are cached.