Skip to content
Back to student guides
KedroMLOpsPipelines & orchestration3 levels104 sectionsCovers Kedro 1.7

The Complete Kedro Guide

Structure maintainable, production-ready data-science code with Kedro. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
18sections
35examples

This is part one of three. It covers everything you need to build and run a real Kedro project, not a teaser. By the end you can create a project from scratch, write small Python functions and wire them into a pipeline, describe your data in a YAML catalog instead of hard-coding file paths, pass parameters without editing code, run the whole pipeline or just a slice of it, look at it as a diagram, and read the error messages you will certainly meet along the way. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. Kedro's ideas look obvious on the page and only become yours after you have watched your own pipeline fail because of a typo in a dataset name and then fixed it.

This guide is written against Kedro 1.7, released at the end of September 2026. Kedro changed a lot between 0.19 and 1.0, and many tutorials on the internet still show the old commands. Where that matters, the text says so and shows the current form.

What Kedro is, and the problem it solves

Kedro is an open-source Python framework for writing data and machine-learning code so that it stays reproducible, modular and maintainable. It is an LF AI & Data project under the Apache 2.0 licence. The key word is framework: Kedro gives you a project layout, a way to describe steps and their dependencies, and a way to describe where data lives. You still write the actual logic (cleaning a table, training a model) in ordinary Python with pandas, scikit-learn, or whatever you like.

What Kedro is not is just as important, because beginners often expect the wrong thing. It is not a scheduler that runs your code every night. It is not a server, a dashboard, or a cloud service. It does not track experiments on its own. It is a way of authoring a pipeline well. When you later need scheduling, retries and monitoring, you hand the Kedro project to an orchestrator such as Airflow, and the project runs there unchanged. That separation is the design: write the logic once, cleanly, then choose where it runs.

FUNCTIONSplain Python
→
PIPELINEa graph of nodes
→
CATALOGwhere data lives
→
RUNlaptop or orchestrator

The diagram is the whole idea. You write functions, you describe how they connect, you describe where the data comes from and goes to, and then Kedro runs them in the right order.

The problem: code that works once

Almost everyone starts a data project in a notebook. You load a CSV, clean it in one cell, engineer a feature in another, train a model in a third, and it works. The trouble starts a month later, or the moment a colleague tries to run it.

The notebook depends on the order you happened to run the cells in. A path like /Users/aya/Downloads/final_v3.csv is baked into cell four. A magic number such as 0.2 for the test split sits in the middle of a function, and nobody remembers why it is not 0.25. Your colleague runs the cells in a different order, gets a different answer, and neither of you can say which is right. Nothing is tested, because nothing is a function you can call on its own. When the project needs to run every night on a server, somebody must rewrite it, and the rewrite introduces new bugs.

These are not signs of a careless person. They are what you get when the structure of the project is an accident. Kedro's answer is to make the structure deliberate, and to make it the same in every project, so that a new team member opens any Kedro project and already knows where the data descriptions, the parameters, the functions and the wiring live.

What problems each part solves

Kedro addresses four separate pains, and it helps to keep them apart.

  • Hard-coded paths and I/O inside functions. Kedro moves all reading and writing out of your functions and into a YAML file called the Data Catalog. Your function receives a DataFrame and returns a DataFrame. Whether that frame came from a local CSV, an S3 bucket or a database is a configuration detail.
  • Hidden order of execution. Instead of remembering which cell to run first, you declare each step's inputs and outputs by name. Kedro works out the order from those names.
  • Magic numbers and secrets in code. Parameters live in parameters.yml, credentials live in a git-ignored file, and code only refers to them by name.
  • Every project laid out differently. kedro new creates the same folder structure each time, so habits transfer between projects and between teams.

There is a fifth, quieter benefit. Because each step is a small function with named inputs and outputs, you can test a step by calling it with a tiny DataFrame, without running anything else. That is the difference between code you can trust and code you can only hope about.

Try it
  1. Think of the last analysis or model you built. Write down every file path, every number such as a threshold or split ratio, and every password or key it contained.
  2. For each one, decide: is this data location, a parameter, or a credential? Kedro has a separate home for each of the three, and you will meet all of them in this guide.
  3. Count how many of them were written directly inside a function body. Those are the ones Kedro will move out.

What came before: scripts, notebooks, and Makefiles

It is worth knowing what people did before, because Kedro only makes sense as a reaction to it.

The first approach is the notebook that grew. It is great for exploration and poor for anything that must run twice with the same result. The second is the folder of scripts numbered 01_clean.py, 02_features.py, 03_train.py, each reading a file the previous one wrote. This is better, because each script can be run alone, but the wiring is in your head. Nothing says that 03_train.py needs the output of 02_features.py, and the paths are repeated as strings in both files. Rename one and the other breaks silently.

The third approach is a build tool, such as a Makefile, that declares which file depends on which. It solves the ordering problem but knows nothing about Python functions, parameters or data formats, so you write a lot of glue by hand.

Kedro keeps the good part of each. From notebooks it keeps the ability to explore (it can open a notebook with your project's catalog already loaded, which we will do later). From scripts it keeps one-step-per-file modularity. From build tools it keeps the dependency graph, but the graph is drawn from the names of the datasets your functions consume and produce, so there is no separate file to keep in sync.

There are also neighbouring tools in this catalogue that people confuse with Kedro. Airflow and Prefect are orchestrators: they schedule and monitor tasks. DVC versions data and pipelines by tracking files. MLflow records experiments and models. Kedro sits at a different level, as the place where you author the code itself, and it cooperates with all of these. If you want to see how Kedro code is scheduled, the Airflow guide is the natural next read, and MLflow is where experiment tracking lives.

The mental model: four nouns

Kedro has a large surface area, but the core is four ideas. Learn these and everything else is a variation.

Node. A node wraps one Python function and gives names to its inputs and outputs. The function should be pure: given the same inputs, it returns the same outputs and does nothing else (no reading files, no printing results to disk). A node says: "call clean_orders with the dataset named orders, and call whatever it returns clean_orders."

Pipeline. A pipeline is a collection of nodes. The crucial fact is that the order of execution is worked out from the data dependencies, not from the order of the list. If node B consumes a dataset that node A produces, Kedro runs A first. You can list them in any order and get the same result. Pipelines can also be added together with +.

Data Catalog. The catalog is the registry of every dataset in the project, built from YAML files named catalog*.yml. A dataset entry says what kind of thing it is (a CSV, a Parquet file, a pickled model), where it lives and how to read and write it. Nodes never do their own I/O; Kedro loads inputs from the catalog before calling your function and saves the outputs after. If a name is used in the pipeline but is not in the catalog, Kedro quietly keeps that dataset in memory for the duration of the run.

Configuration. Everything that is not code: the catalog, the parameters, the credentials and the logging setup. It lives in a folder called conf/, split into environments such as base (shared, committed) and local (yours, git-ignored).

Two smaller nouns complete the picture. A runner decides how the nodes are executed, one after another by default. A session manages one run. You will rarely touch either as a beginner, but the words appear in error messages.

A dataset is a name, not a file In Kedro a dataset is an entry in the catalog that knows how to load and save one thing. The same name appears in the pipeline (as a string) and in catalog.yml (as a key). That shared name is the entire link between your code and your data. A typo in either place is the most common beginner error, and the later sections show exactly how it looks.

Installing Kedro and checking the setup

You need Python 3.10 or later and git. Git matters because kedro new downloads starter templates with it. Kedro 1.x no longer supports Python 3.9, so if python --version prints 3.9 or earlier, install a newer Python first. Kedro works on macOS, Linux and Windows, in cmd.exe and PowerShell.

Use one virtual environment per project. This is not a Kedro rule; it is how Python work stays sane. Two projects that need different versions of pandas cannot share one environment.

The official documentation recommends the tool uv. With it, you can try Kedro without installing anything permanently:

BASH
uvx kedro new --starter spaceflights-pandas --name spaceflights
cd spaceflights
uv run kedro run --pipelines __default__

uvx is an alias for uv tool run: it runs Kedro in a throwaway environment, just long enough to create the project. After that, uv sync and uv run use the project's own .venv.

If you prefer plain Python, create and activate an environment, then install:

BASH
python -m venv .venv
source .venv/bin/activate        # macOS and Linux
# .\.venv\Scripts\activate       # Windows cmd or PowerShell
pip install kedro

Other routes work too: uv pip install kedro, poetry add kedro, or conda install -c conda-forge kedro. On zsh, quote any package with brackets, for example pip install "kedro[server]", or the shell will try to expand the brackets.

Verify the install with:

BASH
kedro info

It prints an ASCII-art banner, the installed Kedro version, and a list of installed plugins, or the words "No plugins installed". You can also run kedro --version (or kedro -V), which prints only the version. python -m kedro works as well. If the command is not found, your environment is not activated; that is nearly always the cause.

Telemetry is on by default Kedro installs a small package called kedro-telemetry, and usage statistics are collected unless you opt out. It records a random ID, the command you ran with arguments masked, version numbers, your operating system, and counts of datasets, nodes and pipelines. It never collects your data. To opt out, set the environment variable KEDRO_DISABLE_TELEMETRY (or DO_NOT_TRACK) to any value, or create a project with kedro new --telemetry=no. Some employers require this, so it is worth knowing on day one.
Try it
  1. Create a fresh virtual environment, activate it, and run pip install kedro.
  2. Run kedro info and find the version line. It should start with 1.
  3. Run kedro starter list and read the aliases it prints. You will use one of them later.
  4. Run deactivate, then kedro info again, and read the error. Now you know what "environment not active" looks like.

Your first project: what kedro new creates

Everything in Kedro starts from kedro new. It asks a few questions and builds a project skeleton. You can answer the questions interactively, or pass them as flags so the command runs without asking. For this guide we will build a small project from an empty skeleton so that every file in it is one you wrote and understand:

BASH
kedro new --name orders_demo --tools=data --example=n

The flags mean: --name is the project name; --tools=data adds the numbered data/ folders (other tools are lint, test, log, docs and pyspark, and you can pass all or none); --example=n says do not include the example pipeline. Kedro refuses project names whose Python package name would shadow something built into Python, such as email, json or import, because that would break imports in confusing ways. Change into the new folder (run ls to see its name), create and activate a virtual environment there, and install the dependencies:

BASH
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install "kedro-datasets[pandas-csvdataset]"

The second install adds the dataset classes we need. Kedro's core ships only a few in-memory datasets; everything that touches a real file format lives in a separate package called kedro-datasets, installed by groups. If you forget it, the first run fails with a message saying the class was not found, and that message is covered in the errors section.

The layout you now have looks like this:

TEXT
orders-demo/
├── conf/
│   ├── base/          # catalog.yml, parameters.yml - committed to git
│   └── local/         # credentials.yml - git-ignored
├── data/              # 01_raw ... 08_reporting
├── notebooks/
├── src/orders_demo/
│   ├── __init__.py
│   ├── __main__.py
│   ├── pipeline_registry.py
│   ├── settings.py
│   └── pipelines/
├── tests/
├── pyproject.toml
├── requirements.txt
└── README.md

Read it as four zones.

conf/ is configuration. base holds settings shared by everyone and committed to git. local holds settings that are yours alone, above all credentials, and is ignored by git so passwords never end up in the repository.

data/ is where datasets live on your machine. The numbered folders are a convention, not a requirement: 01_raw for data exactly as received, 02_intermediate for lightly cleaned data, 03_primary for the tables you model from, 04_feature for features, 05_model_input for what goes into training, 06_models for trained models, 07_model_output for predictions, and 08_reporting for final tables and charts. The numbering tells a reader how far from the source a file is. Raw data is never edited in place.

src/orders_demo/ is your code. pipeline_registry.py lists the pipelines the project knows about. settings.py holds project-wide settings. Each pipeline gets its own folder under pipelines/, containing a nodes.py for the functions and a pipeline.py for the wiring.

pyproject.toml marks the project root and records its identity in a [tool.kedro] section: package_name, project_name and kedro_init_version. Kedro looks for this file to know it is inside a project, which is why you can run kedro run from any subfolder, and why running it from outside a project gives "Could not find the project configuration file 'pyproject.toml'".

Commit before you do anything else Run git init and make a first commit straight after kedro new. The generated .gitignore already excludes conf/local/, so credentials stay out of the commit, and you will have a clean baseline to compare against when something breaks.

Starters, if you want an example

If you would rather begin from a working example, kedro new --starter <alias> copies a finished project. The official aliases in 1.7 are astro-airflow-iris, spaceflights-pandas, spaceflights-pyspark, databricks-iris and support-agent-langgraph. The most widely used is spaceflights-pandas, a three-pipeline project that predicts shuttle prices, and it is the project the official tutorial walks through. Older articles mention a pandas-iris starter; it has not existed since version 0.19.

Try it
  1. Run the kedro new command above, then git init and commit.
  2. Open pyproject.toml and find the [tool.kedro] section. Note the three keys.
  3. Open src/orders_demo/pipeline_registry.py and read it. You do not need to change it yet; just notice that it discovers pipelines automatically.
  4. Run kedro run. It should finish almost at once, because there is nothing to run. Kedro is telling you the skeleton is valid.

Your first pipeline, step by step

We will build a tiny but realistic pipeline. A CSV of shop orders comes in; the pipeline cleans it, then summarises revenue per city. Three datasets and two nodes, which is enough to show every moving part.

Step 1: put in some raw data

Create data/01_raw/orders.csv with this content. It has a few deliberate flaws: a missing city, a missing amount, and inconsistent capitalisation.

data/01_raw/orders.csv
order_id,city,amount,created
1,cairo,120.5,2026-01-03
2,Riyadh,80,2026-01-04
3,  dubai ,200,2026-01-04
4,Cairo,,2026-01-05
5,,45,2026-01-06
6,riyadh,15.25,2026-01-07
7,Cairo,9,2026-01-08

Step 2: scaffold a pipeline

Kedro can create the folder structure for you:

BASH
kedro pipeline create orders

This makes src/orders_demo/pipelines/orders/ with nodes.py and pipeline.py, a conf/base/parameters_orders.yml for this pipeline's parameters, and a matching folder under tests/. The habit this builds is worth keeping: one pipeline, one folder, and its parameters and tests next to it.

Step 3: write the functions

Open src/orders_demo/pipelines/orders/nodes.py and replace its contents with two plain functions:

src/orders_demo/pipelines/orders/nodes.py
import pandas as pd


def clean_orders(orders: pd.DataFrame) -> pd.DataFrame:
    """Drop incomplete rows and tidy the city names."""
    cleaned = orders.dropna(subset=["city", "amount"]).copy()
    cleaned["city"] = cleaned["city"].str.strip().str.title()
    cleaned["amount"] = cleaned["amount"].astype(float)
    return cleaned


def summarise_by_city(orders: pd.DataFrame, min_amount: float) -> pd.DataFrame:
    """Count orders and total revenue per city, ignoring tiny orders."""
    large = orders[orders["amount"] >= min_amount]
    return large.groupby("city", as_index=False).agg(
        orders=("order_id", "count"),
        revenue=("amount", "sum"),
    )

Look at what is missing. There is no pd.read_csv, no file path, no to_csv, and the threshold is an argument rather than a constant. These functions know nothing about Kedro. You could import them into a test and call them with a three-row DataFrame. That is exactly the point: the logic is separate from the plumbing.

Step 4: wire them into a pipeline

Open pipeline.py in the same folder:

src/orders_demo/pipelines/orders/pipeline.py
from kedro.pipeline import Node, Pipeline

from .nodes import clean_orders, summarise_by_city


def create_pipeline(**kwargs) -> Pipeline:
    return Pipeline(
        [
            Node(
                func=clean_orders,
                inputs="orders",
                outputs="clean_orders",
                name="clean_orders_node",
            ),
            Node(
                func=summarise_by_city,
                inputs=["clean_orders", "params:min_amount"],
                outputs="city_summary",
                name="summarise_by_city_node",
            ),
        ]
    )

Read this slowly, because it is the heart of Kedro. Each Node says: call this function, feed it these named inputs, and call its result by this name. The first node takes the dataset called orders and produces clean_orders. The second takes clean_orders (the very name the first produced, and this is the link that makes the order) plus a parameter called min_amount, and produces city_summary. The string "params:min_amount" is Kedro's syntax for "the parameter min_amount from the parameters file".

If a function takes several arguments, inputs is a list, matched to the function's arguments by position. If it returns several values, outputs is a list too. Everything after outputs in Node(...) is keyword-only, and name is optional but strongly recommended, because names appear in logs, in error messages and in --nodes on the command line.

Node or node? Older tutorials write from kedro.pipeline import node, pipeline (lowercase functions). Those wrappers still work in Kedro 1.x, but the capitalised Node and Pipeline classes are the preferred form, and this guide uses them throughout. You will see both in real projects.

Step 5: describe the data

Open conf/base/catalog.yml and add three entries:

conf/base/catalog.yml
orders:
  type: pandas.CSVDataset
  filepath: data/01_raw/orders.csv

clean_orders:
  type: pandas.CSVDataset
  filepath: data/02_intermediate/clean_orders.csv

city_summary:
  type: pandas.CSVDataset
  filepath: data/08_reporting/city_summary.csv

Each key is a dataset name, identical to the strings in pipeline.py. The type is the dataset class; pandas.CSVDataset is shorthand for kedro_datasets.pandas.CSVDataset. Note the spelling: Dataset with a lowercase s. Before kedro-datasets 2.0 it was DataSet, and old tutorials still show that. The old spelling produces a warning and then a "class not found" error.

Step 6: set the parameter

Open conf/base/parameters_orders.yml and set:

conf/base/parameters_orders.yml
min_amount: 10

The pipeline's params:min_amount input now resolves to 10.

Try it
  1. Type in all six steps. Resist copy-pasting for the two code files; typing slows you down enough to notice each name.
  2. Before running anything, say aloud which node will run first and how you know. Then swap the order of the two nodes in the list and ask yourself whether the answer changes. It should not.
  3. Change min_amount to 100 in the YAML file only. Notice that no Python file needs to change.

Running the pipeline and reading the output

Now run it:

BASH
kedro run

With no options, Kedro runs the pipeline named __default__. Where does that come from? Open pipeline_registry.py. The template calls find_pipelines(), which discovers every folder under pipelines/ that exposes a create_pipeline function, registers each under its folder name, and sets __default__ to the sum of all of them. So kedro run runs everything, and kedro run --pipelines orders runs just the orders pipeline.

The output is a stream of log lines. You will see Kedro loading the project configuration, then for each node a line announcing it is running, lines saying it is loading a dataset and saving a dataset, a line saying the node completed, and a final "Pipeline execution completed successfully." The exact wording and formatting differ slightly between versions and with Rich logging, so learn the shape rather than the text: load inputs, run function, save outputs, next node.

When it finishes, look at data/02_intermediate/clean_orders.csv and data/08_reporting/city_summary.csv. The cleaned file has five rows (the two incomplete ones were dropped), the cities are tidy ("Cairo", "Riyadh", "Dubai"), and the summary has one row per city. With min_amount: 10 the order from Cairo worth 9 is excluded, so Cairo shows one counted order rather than two. Change the parameter and run again to see the numbers move.

Notice what you did not have to do: write any code to load the CSV, or tell Kedro which node goes first.

Seeing the plan without running

Two commands show what Kedro intends to do, which is useful when a pipeline grows:

BASH
kedro registry list
kedro registry describe orders

The first lists the pipeline names. The second lists the nodes of one pipeline (or __default__ if you give no name).

Try it
  1. Run kedro run and open both output files. Check the numbers against the raw CSV by hand.
  2. Delete data/02_intermediate/clean_orders.csv and run again. It is recreated.
  3. Run kedro registry list and kedro registry describe orders and read the output.
  4. Introduce a typo in one dataset name in pipeline.py and run. Read the error top to bottom, then fix it.

The Data Catalog in depth

You have used three catalog entries. The catalog can do much more, and most of a beginner's day-to-day Kedro work is editing it.

Anatomy of an entry

Here is one entry with most of the options a beginner will meet:

conf/base/catalog.yml
companies:
  type: pandas.CSVDataset
  filepath: data/01_raw/companies.csv
  load_args:
    sep: ","
  save_args:
    index: false
  metadata:
    kedro-viz:
      layer: raw
  • type names the dataset class. Kedro resolves a short name like pandas.CSVDataset first against kedro.io, then against kedro_datasets, and finally treats it as a full dotted path to a class in your own code.
  • filepath is a path or URL. Local paths are relative to the project root. Because Kedro uses the fsspec library underneath, the same key accepts s3://, gcs://, abfs://, hdfs:// and http(s):// URLs. Moving a dataset from your laptop to cloud storage is therefore a one-line change, with no code change in the nodes.
  • load_args are forwarded to the underlying reader (for the CSV dataset, pandas.read_csv), and save_args to the writer (DataFrame.to_csv). Here index: false stops pandas writing the row numbers as an extra column, a classic annoyance.
  • credentials (not shown) names a block in a credentials file, covered below.
  • metadata is free-form information that Kedro ignores but tools can use. The kedro-viz: layer: entry tells Kedro-Viz which layer to draw the dataset in. In old tutorials the layer: key sat at the top level of the entry; it has since moved under metadata.

Choosing a dataset type

The type depends on what you are storing. A few you will meet early:

You have Use
A CSV file pandas.CSVDataset
A Parquet file (compact, fast, typed) pandas.ParquetDataset
An Excel file pandas.ExcelDataset
A trained scikit-learn model pickle.PickleDataset
A chart made with matplotlib matplotlib.MatplotlibDataset

Each group needs its own dependencies. Install them by naming the group in brackets, for example pip install "kedro-datasets[pandas-csvdataset,pandas-parquetdataset]". If you skip this you will see either No module named '...'. Please install the missing dependencies for kedro_datasets... or Class '...' not found, and both mean the same thing: install the right group.

Pickle runs code when it loads Pickle is convenient for saving a Python object exactly as it is, and it is how many people store models. But loading a pickle file can execute arbitrary code, so never load a pickle that came from someone you do not trust. For models you share, prefer a format designed for data.

Versioning a dataset

Add versioned: true to an entry and Kedro stops overwriting the file. Each save goes into a new timestamped folder, <filepath>/<timestamp>/<filename>, and loads read the latest one by default. This keeps a history of a model or a report without any code. You can load an older one with kedro run --load-versions=city_summary:2026-01-01T00.00.00.000Z. Versioning is not available for http(s):// paths, and asking for it there raises "Versioning is not supported for HTTP protocols."

What happens to datasets that are not in the catalog

A dataset that appears in the pipeline but has no catalog entry is held in memory while the run lasts and discarded afterwards. That is useful for throwaway intermediates and dangerous for anything you want to keep or inspect. The rule of thumb for beginners: if you might want to look at it afterwards, give it a catalog entry. Also remember that memory-only datasets cannot be passed between separate tasks when a pipeline is later split across an orchestrator, so in production every intermediate gets an entry.

Many similar datasets: factories

If you have twenty CSV files that follow a naming pattern, writing twenty entries is tedious. A dataset factory matches names with a pattern, like a reversed f-string:

conf/base/catalog.yml
"{name}_data":
  type: pandas.CSVDataset
  filepath: data/01_raw/{name}_data.csv

Now any dataset whose name ends in _data, such as customers_data or shipments_data, is served by this one entry, with {name} filled in. The quotes are mandatory; without them YAML reads the braces as its own syntax and you get a parser error. Explicit entries always beat patterns, and there may be at most one catch-all pattern of your own. You do not need factories on day one, but you should recognise them when a colleague's catalog seems to be missing entries that the pipeline uses.

To see how Kedro resolves names, use:

BASH
kedro catalog describe-datasets
kedro catalog list-patterns
kedro catalog resolve-patterns

describe-datasets groups the datasets into those defined explicitly, those produced by a factory pattern and those that will default to memory. list-patterns shows the patterns in priority order, and resolve-patterns prints the fully resolved configuration. These replaced the older kedro catalog list and kedro catalog rank in version 1.0.

Try it
  1. Change the city_summary entry to pandas.ParquetDataset with a .parquet filepath, install kedro-datasets[pandas-parquetdataset], and rerun. You changed the file format without touching a line of Python.
  2. Add versioned: true to clean_orders and run twice. Look inside data/02_intermediate/clean_orders.csv and find the timestamped folders.
  3. Run kedro catalog describe-datasets. Which datasets are explicit and which are defaults?

Parameters and credentials

Parameters: numbers that are not code

A parameter is a value that changes the behaviour of a pipeline and that you may want to adjust without editing functions: a threshold, a split ratio, a random seed, a list of columns. Parameters live in YAML files whose names start with parameters, inside conf/base/ (or another environment). You already used one.

Nested values work too:

conf/base/parameters_orders.yml
min_amount: 10
model_options:
  test_size: 0.2
  random_state: 3

In a node, reference a single value as "params:min_amount", a nested one with dots as "params:model_options.test_size", or the whole dictionary as "params:model_options". The special name "parameters" passes every parameter at once. The function receives ordinary Python values: a float, a dict.

You can override a parameter for one run, without editing any file, from the command line:

BASH
kedro run --params="min_amount=50"

The syntax is key=value, separated by commas for several (--params="min_amount=50,model_options.test_size=0.3"), with dots for nested keys. Kedro turns values into integers or floats where it can and leaves the rest as strings. A colon (key:value) was the syntax before 0.19 and some documentation examples still show it; use the equals sign. If a value contains spaces, quote the whole argument.

Credentials: secrets that stay out of git

Some datasets need a password, key or token, for example an S3 bucket or a database. Never put these in catalog.yml, because that file is committed. Put them in conf/local/credentials.yml, which git ignores:

conf/local/credentials.yml
dev_s3:
  client_kwargs:
    aws_access_key_id: your-key-id
    aws_secret_access_key: your-secret

and reference the block by name from the catalog:

conf/base/catalog.yml
orders:
  type: pandas.CSVDataset
  filepath: s3://my-bucket/orders.csv
  credentials: dev_s3

Anything in conf/local/ is ignored by git, and kedro package leaves those files out of the packaged configuration. On a server or in a CI job there is no conf/local/credentials.yml, so the usual approach is to read secrets from environment variables. Inside a credentials file only, you can write ${oc.env:AWS_ACCESS_KEY_ID} and Kedro substitutes the variable's value. For AWS specifically, the standard variables AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY and AWS_SESSION_TOKEN are also picked up directly.

A committed secret is a leaked secret If you ever commit a real key to git, deleting it in the next commit does not help, because the history keeps it. Rotate the key at the provider and treat it as public. Check git status before every commit in a project that has a credentials file.
Try it
  1. Add the nested model_options block to your parameters file. You do not have to use it yet.
  2. Run kedro run --params="min_amount=100" and confirm the summary shrinks. Then run plain kedro run and confirm it returns to the YAML value.
  3. Create conf/local/credentials.yml with a fake block, run git status, and confirm the file does not appear.

Configuration environments

The conf/ folder holds more than one environment, and the idea is simple but easy to misread. An environment is a subfolder. base is the foundation and is always loaded. On top of it Kedro loads a second environment, local by default, and where the same top-level key appears in both, the second one wins. By default this replacement is whole-key: if local defines model_options, it replaces base's entire model_options, not only the keys it mentions.

The effect is that one project can behave differently in different places with no code change. base holds the settings that are right for everyone. Your local folder holds your own tweaks, such as a smaller sample file to keep development fast. You can define further environments (dev, staging, prod) as folders and choose one at run time:

BASH
kedro run --env=staging

Or set the environment variable KEDRO_ENV=staging; if both are present the command-line flag wins.

Two warnings follow from how this works. First, within one environment, the same top-level key must not appear in two files, or Kedro raises ValueError: Duplicate keys found in ... naming both files. That protects you from silently losing a setting. Second, if you name an environment that does not exist, you get a MissingConfigException about a configuration path that does not exist or is not a directory. Create conf/<env>/ first.

Kedro also has a templating feature built on a library called OmegaConf, so a value in YAML can refer to another with ${...}. You do not need it yet; mid-level covers it.

Try it
  1. Create conf/local/parameters_orders.yml containing min_amount: 50 and run kedro run. Which value is used, and why?
  2. Create conf/staging/ with a different catalog path for city_summary and run kedro run --env=staging.
  3. Delete the local file again so your later results match this guide.

The everyday commands, grouped by what you want to do

You do not need to memorise the full command list. Group it by intent.

I want to run things

BASH
kedro run                                  # the default pipeline
kedro run --pipelines orders               # one pipeline
kedro run --pipelines orders,training      # several at once
kedro run --env=staging                    # another configuration environment
kedro run --params="min_amount=50"         # override parameters

Use --pipelines (plural). Kedro 1.2 introduced it so several pipelines can run together, and it deprecated the singular --pipeline flag: the old flag still works but prints a deprecation warning, and using both at once is an error. Most online tutorials still show the old flag.

I want to run only part of a pipeline

While developing you rarely want the whole thing. Kedro offers several ways to pick a slice:

BASH
kedro run --nodes=clean_orders_node        # specific nodes, by name
kedro run --tags=cleaning                  # nodes that carry a tag
kedro run --from-nodes=summarise_by_city_node
kedro run --to-nodes=clean_orders_node
kedro run --from-inputs=clean_orders
kedro run --to-outputs=city_summary

--from-nodes runs that node and everything after it; --to-nodes runs everything up to and including it. --from-inputs and --to-outputs do the same by dataset name. Tags are labels you add to a node in code, Node(..., tags=["cleaning"]), and --tags selects nodes carrying any of the tags you list. Names and tags may contain only letters, digits, hyphens, underscores and full stops.

A very useful flag is --only-missing-outputs, which skips any node whose saved outputs already all exist. After a failure halfway through a long run, it lets you resume without recomputing the early steps.

If your filters match nothing, Kedro says "Pipeline contains no nodes after applying all provided filters." That message means a typo or a wrong name, not a broken pipeline.

I want to look around

BASH
kedro info                                 # version and plugins
kedro registry list                        # pipeline names
kedro registry describe orders             # nodes in a pipeline
kedro catalog describe-datasets            # how each dataset is defined
kedro starter list                         # available starters

I want to create or change things

BASH
kedro new                                  # new project
kedro pipeline create <name>               # new pipeline folder
kedro pipeline delete <name>               # remove one

I want to work interactively

BASH
kedro ipython                              # IPython with the project loaded
kedro jupyter lab                          # JupyterLab with a Kedro kernel
kedro jupyter notebook                     # classic notebook

The next sections cover the interactive and visual tools.

Run from any folder inside the project Kedro finds the project root by looking for pyproject.toml, so you can run project commands from any subdirectory. The only commands that work outside a project are the global ones: kedro new, kedro starter list and kedro info.
Try it
  1. Add tags=["cleaning"] to the first node. Run kedro run --tags=cleaning and confirm only one node runs.
  2. Run kedro run --from-nodes=summarise_by_city_node. It works because clean_orders was saved to disk by an earlier run. Now delete that file and run it again, and read the error.
  3. Try --only-missing-outputs twice in a row and watch what is skipped.

Exploring with notebooks and Kedro-Viz

Notebooks, the Kedro way

Notebooks are the best tool for looking at data, and Kedro does not ask you to give them up. It asks you to keep them as a window onto the project and not as the place where the logic lives. Start a notebook with your project already loaded:

BASH
kedro jupyter lab

This creates a kernel named after your package, shown as Kedro (orders_demo), and opens JupyterLab. In the kernel, these variables already exist: catalog, context, pipelines and session. Try:

PYTHON
catalog.keys()                      # all dataset names
df = catalog.load("clean_orders")   # read a dataset the way a node would
df.head()
catalog.load("parameters")          # every parameter as a dictionary
pipelines["__default__"]            # the pipeline object

catalog.load("name") is the important one. It returns a dataset exactly as your nodes would see it, so you can experiment on real data and, once a snippet works, paste it into nodes.py as a function. That is the healthy loop: explore in the notebook, then promote working code into a node.

If you launch Jupyter some other way, load the project with two magic commands: %load_ext kedro.ipython, then %reload_kedro <path to project root>. If you see "Line magic function '%reload_kedro' not found", the extension was not loaded. kedro ipython does the same for a terminal IPython session. If it complains that the module IPython is not found, install the project requirements first.

A warning for those who read older material: the catalog method catalog.list() no longer exists. Since 1.0 use catalog.keys() or catalog.filter("pattern").

Kedro-Viz: see the pipeline as a diagram

Kedro-Viz is a separate package that draws your pipeline as an interactive graph, with datasets and nodes as shapes and the dependencies as arrows. It is the best way to explain a pipeline to someone else, and also the quickest way to spot a mistake, such as a node with no inputs.

BASH
pip install kedro-viz
kedro viz run

It opens a browser tab on 127.0.0.1:4141. (The shorter kedro viz also works.) Click a node to see its code and parameters, click a dataset to see its type and file path, and use the filters on the side to show only some pipelines, tags or layers. The layers you set under metadata: kedro-viz: layer: in the catalog are drawn as horizontal bands, so a pipeline reads from raw data at one edge to reporting at the other. Use --no-browser if you are on a remote machine and want to open the address yourself.

Viz is a window, not a runner Kedro-Viz reads your project's structure; it does not run the pipeline. If the diagram looks wrong, the cause is in the pipeline definition, not in Viz.
Try it
  1. Run kedro jupyter lab. In a new notebook under the Kedro kernel, run catalog.keys() and load orders. Compare it to the cleaned version.
  2. Install Kedro-Viz and run kedro viz run. Find your two nodes and three datasets.
  3. Add metadata: kedro-viz: layer: raw to orders in the catalog (nested as in the earlier example) and refresh the diagram.

Growing the project: several pipelines, namespaces and reuse

A project of one pipeline is a demo. Real projects have several: data preparation, training, evaluation. Kedro's answer is the modular pipeline: each one lives in its own folder with its own nodes.py and pipeline.py, and exposes a create_pipeline function. You already made one with kedro pipeline create orders. Making a second is the same command:

BASH
kedro pipeline create reporting

Because pipeline_registry.py uses find_pipelines(), the new pipeline is registered the moment its folder has a create_pipeline. Then kedro run runs both (the __default__ pipeline is their sum) and kedro run --pipelines reporting runs only the new one. Datasets are shared between pipelines by name: if reporting consumes city_summary, it finds it in the same catalog, produced by orders. Keeping one pipeline per folder means each can be understood, tested and run on its own.

Namespaces: using a pipeline twice

Sometimes the same pipeline should run twice on different inputs, for example one model trained on two regions. Wrapping it in a namespace prefixes its node names, dataset names and parameters with a label, so two copies do not collide:

PYTHON
from kedro.pipeline import Pipeline

base = create_pipeline()
riyadh = Pipeline(base, namespace="riyadh")
cairo = Pipeline(base, namespace="cairo")

Now the dataset clean_orders becomes riyadh.clean_orders in the one copy and cairo.clean_orders in the other. Namespaces are also how Kedro-Viz collapses groups of nodes into one expandable box. Run one with kedro run --namespaces=riyadh; the flag accepts several, comma-separated. The full dotted-name rules are covered at mid-level; for now remember that the dot character in a dataset name is reserved for namespaces, so avoid dots in your own dataset names.

If you reuse a pipeline without a namespace you will see "Pipeline nodes must have unique names". That message means the same node name appears twice, and a namespace is the fix.

Try it
  1. Create a second pipeline called reporting with one node that reads city_summary and returns the top city. Give its output a catalog entry.
  2. Run kedro registry list and confirm both pipelines appear, then kedro run --pipelines reporting.
  3. Open Kedro-Viz and confirm the two pipelines connect through city_summary.

Testing a node

The payoff of pure functions is that testing is ordinary Python. No Kedro needed:

tests/pipelines/orders/test_nodes.py
import pandas as pd

from orders_demo.pipelines.orders.nodes import clean_orders, summarise_by_city


def test_clean_orders_drops_incomplete_rows():
    raw = pd.DataFrame(
        {"order_id": [1, 2], "city": ["cairo", None], "amount": [10.0, 5.0]}
    )
    result = clean_orders(raw)
    assert len(result) == 1
    assert result.iloc[0]["city"] == "Cairo"


def test_summary_ignores_small_orders():
    orders = pd.DataFrame(
        {"order_id": [1, 2], "city": ["Cairo", "Cairo"], "amount": [100.0, 1.0]}
    )
    summary = summarise_by_city(orders, min_amount=10)
    assert summary.iloc[0]["orders"] == 1

Run it with pytest from the project root. If you created the project with the test tool, pytest is already in the requirements; otherwise pip install pytest. The test builds a three-row DataFrame, calls the function and checks the answer. It does not read a file, does not need a catalog and runs in milliseconds. That is the benefit of keeping I/O out of functions.

Common errors and how to read them

Kedro errors are long because they include the chain of causes. The rule is: read the last line first, then the one above it. The final exception says what went wrong; the chained ones say where. Here are the ones you will meet in your first week.

What you see What it means What to do
Could not find the project configuration file 'pyproject.toml' You are outside a Kedro project cd into the project
DatasetError: Class 'pandas.CSVDataSet' not found, is this a typo? Old capital-S spelling, or kedro-datasets not installed Use CSVDataset and install kedro-datasets[pandas-csvdataset]
No module named '...'. Please install the missing dependencies for kedro_datasets... A dataset group's dependencies are missing pip install "kedro-datasets[<group>]"
DatasetNotFoundError: Dataset 'x' not found in the catalog A name is used but defined nowhere Compare the name in code and in catalog.yml
Pipeline input(s) {'x'} not found in the DataCatalog A pipeline input has no data behind it Add a catalog entry, or fix the name
DatasetError: Failed while loading data from dataset ... The file, path or library is wrong Read the chained exception underneath
DatasetError: Saving 'None' to a 'Dataset' is not allowed A node forgot to return Return the value
Duplicate keys found in <f1> and <f2> One key in two files of the same environment Rename it or move it
ParserError: Invalid YAML or JSON file ... A YAML syntax mistake, often an unquoted {name} pattern Quote it and fix indentation
OutputNotUniqueError Two nodes write the same dataset Rename one output
Failed to save outputs of node ... returned N output(s), whereas the node definition contains M Return values do not match outputs Align them
Pipeline contains no nodes after applying all provided filters. A --tags or --nodes filter matched nothing Fix the filter
Your Kedro project version X does not match Kedro package version Y A 0.19 project opened with Kedro 1.x Migrate the project, then update kedro_init_version
A node must return what it promises If outputs=["a", "b"] the function must return two values, such as return a, b. If it returns one, Kedro stops with the "returned N output(s)" error. The reverse holds too: a function that returns a value while the node was declared with outputs=None produces a warning that the value is being thrown away.

Two habits make errors much cheaper. Always give nodes a name, because then the log line says which step failed in words you chose. And when a dataset fails to load, open a notebook and call catalog.load("that_name") directly. The error is then isolated from the pipeline, and you can try the fix interactively.

Finally, watch for two warnings. KedroDeprecationWarning means you used a flag or function scheduled for removal (the old --pipeline flag is the common one); it still works today, but fix it now. A KedroExperimentalWarning means you touched a feature whose design may still change.

Try it
  1. Break things on purpose, one at a time: misspell a dataset in the catalog, misspell it in the pipeline, return only one value from a function declared with two outputs, and write CSVDataSet with a capital S. For each, find the line of the error that tells you the cause.
  2. Fix each one before moving on, and run kedro run to confirm.

Putting it all together

Here is the whole project in one pass, the way you would set it up for a colleague. Every command is one you have already used.

BASH
kedro new --name orders_demo --tools=data --example=n
cd orders-demo
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install "kedro-datasets[pandas-csvdataset]" pytest
git init && git add . && git commit -m "Kedro skeleton"
kedro pipeline create orders

Then edit five files:

  1. data/01_raw/orders.csv, the raw data.
  2. src/orders_demo/pipelines/orders/nodes.py, the two pure functions.
  3. src/orders_demo/pipelines/orders/pipeline.py, the two nodes.
  4. conf/base/catalog.yml, the three datasets.
  5. conf/base/parameters_orders.yml, the min_amount parameter.

Run it, inspect it and test it:

BASH
kedro run
kedro registry describe orders
kedro run --params="min_amount=50"
pytest
kedro viz run

Step back and notice the properties of what you built. The functions contain no paths and no magic numbers, so they are reusable. The wiring is explicit and named, so a stranger can read it. The data locations are in YAML, so moving from a laptop to cloud storage means editing one file. The parameter can be changed from the command line, so experiments need no code edits. The code is tested without running the pipeline. And the project has the same shape as every other Kedro project in the world, so any Kedro user can find their way around it.

A final check for people working at companies in the Gulf or Egypt: where your data sits matters for regulation. Because locations are just filepath values, you can point a dataset at a storage bucket in a Middle East cloud region and keep the code identical, which is exactly the kind of change this design makes easy. Confirm the residency rules that apply to your employer, but know that Kedro will not stand in the way.

Try it
  1. Rebuild the whole project from an empty folder without looking at this guide. Use it only when stuck, and note each place you got stuck.
  2. Extend it with a third node that counts orders per day from the created column, saved as a new dataset with a catalog entry and a test.
  3. Push it to a git repository and ask a friend to clone it, install the requirements and run kedro run. Fix whatever breaks. That is the real test of reproducibility.

What you can now do, and what comes next

You can now create a Kedro project, explain its four nouns (node, pipeline, catalog, configuration), write pure functions and wire them into a pipeline, describe data in YAML, pass parameters and keep credentials out of git, choose a configuration environment, run all or part of a pipeline, explore it in a notebook, draw it with Kedro-Viz, split a project into modular pipelines, test a node, and read Kedro's common errors.

Here is what to check you can do without looking anything up:

Can you... The answer
Say what decides node order? Data dependencies between names, not list order
Move a file location without touching Python? Edit filepath in catalog.yml
Change a threshold for one run? kedro run --params="key=value"
Keep a secret out of git? conf/local/credentials.yml, referenced by name
Run only the second half of a pipeline? --from-nodes or --from-inputs
Find why a name is not found? Compare the string in code and in the catalog
Check the version and plugins? kedro info

Mid-level takes every one of those topics further: how the catalog, config loader and runners work internally, dataset factories and templating with OmegaConf, parameter and dataset validation, hooks, the choice of runner, packaging a project as a wheel, and deploying it with Docker and orchestrators.

Senior then covers what you own when Kedro is your team's platform: the security model, scaling limits, upgrade and migration planning, custom datasets and runners, and where Kedro stops and another tool should take over.

Natural next reads in this catalogue are Airflow for scheduling a Kedro project, MLflow for experiment tracking, DVC for versioning data, and Docker for packaging it all so it runs anywhere.

Sources