This is part one of three. It covers everything you need to build and run a real Kedro project, not a teaser. By the end you can create a project from scratch, write small Python functions and wire them into a pipeline, describe your data in a YAML catalog instead of hard-coding file paths, pass parameters without editing code, run the whole pipeline or just a slice of it, look at it as a diagram, and read the error messages you will certainly meet along the way. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. Kedro's ideas look obvious on the page and only become yours after you have watched your own pipeline fail because of a typo in a dataset name and then fixed it.
This guide is written against Kedro 1.7, released at the end of September 2026. Kedro changed a lot between 0.19 and 1.0, and many tutorials on the internet still show the old commands. Where that matters, the text says so and shows the current form.
What Kedro is, and the problem it solves
Kedro is an open-source Python framework for writing data and machine-learning code so that it stays reproducible, modular and maintainable. It is an LF AI & Data project under the Apache 2.0 licence. The key word is framework: Kedro gives you a project layout, a way to describe steps and their dependencies, and a way to describe where data lives. You still write the actual logic (cleaning a table, training a model) in ordinary Python with pandas, scikit-learn, or whatever you like.
What Kedro is not is just as important, because beginners often expect the wrong thing. It is not a scheduler that runs your code every night. It is not a server, a dashboard, or a cloud service. It does not track experiments on its own. It is a way of authoring a pipeline well. When you later need scheduling, retries and monitoring, you hand the Kedro project to an orchestrator such as Airflow, and the project runs there unchanged. That separation is the design: write the logic once, cleanly, then choose where it runs.
The diagram is the whole idea. You write functions, you describe how they connect, you describe where the data comes from and goes to, and then Kedro runs them in the right order.
The problem: code that works once
Almost everyone starts a data project in a notebook. You load a CSV, clean it in one cell, engineer a feature in another, train a model in a third, and it works. The trouble starts a month later, or the moment a colleague tries to run it.
The notebook depends on the order you happened to run the cells in. A path like /Users/aya/Downloads/final_v3.csv is baked into cell four. A magic number such as 0.2 for the test split sits in the middle of a function, and nobody remembers why it is not 0.25. Your colleague runs the cells in a different order, gets a different answer, and neither of you can say which is right. Nothing is tested, because nothing is a function you can call on its own. When the project needs to run every night on a server, somebody must rewrite it, and the rewrite introduces new bugs.
These are not signs of a careless person. They are what you get when the structure of the project is an accident. Kedro's answer is to make the structure deliberate, and to make it the same in every project, so that a new team member opens any Kedro project and already knows where the data descriptions, the parameters, the functions and the wiring live.
What problems each part solves
Kedro addresses four separate pains, and it helps to keep them apart.
- Hard-coded paths and I/O inside functions. Kedro moves all reading and writing out of your functions and into a YAML file called the Data Catalog. Your function receives a DataFrame and returns a DataFrame. Whether that frame came from a local CSV, an S3 bucket or a database is a configuration detail.
- Hidden order of execution. Instead of remembering which cell to run first, you declare each step's inputs and outputs by name. Kedro works out the order from those names.
- Magic numbers and secrets in code. Parameters live in
parameters.yml, credentials live in a git-ignored file, and code only refers to them by name. - Every project laid out differently.
kedro newcreates the same folder structure each time, so habits transfer between projects and between teams.
There is a fifth, quieter benefit. Because each step is a small function with named inputs and outputs, you can test a step by calling it with a tiny DataFrame, without running anything else. That is the difference between code you can trust and code you can only hope about.
- Think of the last analysis or model you built. Write down every file path, every number such as a threshold or split ratio, and every password or key it contained.
- For each one, decide: is this data location, a parameter, or a credential? Kedro has a separate home for each of the three, and you will meet all of them in this guide.
- Count how many of them were written directly inside a function body. Those are the ones Kedro will move out.
What came before: scripts, notebooks, and Makefiles
It is worth knowing what people did before, because Kedro only makes sense as a reaction to it.
The first approach is the notebook that grew. It is great for exploration and poor for anything that must run twice with the same result. The second is the folder of scripts numbered 01_clean.py, 02_features.py, 03_train.py, each reading a file the previous one wrote. This is better, because each script can be run alone, but the wiring is in your head. Nothing says that 03_train.py needs the output of 02_features.py, and the paths are repeated as strings in both files. Rename one and the other breaks silently.
The third approach is a build tool, such as a Makefile, that declares which file depends on which. It solves the ordering problem but knows nothing about Python functions, parameters or data formats, so you write a lot of glue by hand.
Kedro keeps the good part of each. From notebooks it keeps the ability to explore (it can open a notebook with your project's catalog already loaded, which we will do later). From scripts it keeps one-step-per-file modularity. From build tools it keeps the dependency graph, but the graph is drawn from the names of the datasets your functions consume and produce, so there is no separate file to keep in sync.
There are also neighbouring tools in this catalogue that people confuse with Kedro. Airflow and Prefect are orchestrators: they schedule and monitor tasks. DVC versions data and pipelines by tracking files. MLflow records experiments and models. Kedro sits at a different level, as the place where you author the code itself, and it cooperates with all of these. If you want to see how Kedro code is scheduled, the Airflow guide is the natural next read, and MLflow is where experiment tracking lives.
The mental model: four nouns
Kedro has a large surface area, but the core is four ideas. Learn these and everything else is a variation.
Node. A node wraps one Python function and gives names to its inputs and outputs. The function should be pure: given the same inputs, it returns the same outputs and does nothing else (no reading files, no printing results to disk). A node says: "call clean_orders with the dataset named orders, and call whatever it returns clean_orders."
Pipeline. A pipeline is a collection of nodes. The crucial fact is that the order of execution is worked out from the data dependencies, not from the order of the list. If node B consumes a dataset that node A produces, Kedro runs A first. You can list them in any order and get the same result. Pipelines can also be added together with +.
Data Catalog. The catalog is the registry of every dataset in the project, built from YAML files named catalog*.yml. A dataset entry says what kind of thing it is (a CSV, a Parquet file, a pickled model), where it lives and how to read and write it. Nodes never do their own I/O; Kedro loads inputs from the catalog before calling your function and saves the outputs after. If a name is used in the pipeline but is not in the catalog, Kedro quietly keeps that dataset in memory for the duration of the run.
Configuration. Everything that is not code: the catalog, the parameters, the credentials and the logging setup. It lives in a folder called conf/, split into environments such as base (shared, committed) and local (yours, git-ignored).
Two smaller nouns complete the picture. A runner decides how the nodes are executed, one after another by default. A session manages one run. You will rarely touch either as a beginner, but the words appear in error messages.
catalog.yml (as a key). That shared name is the entire link between your code and your data. A typo in either place is the most common beginner error, and the later sections show exactly how it looks.
Installing Kedro and checking the setup
You need Python 3.10 or later and git. Git matters because kedro new downloads starter templates with it. Kedro 1.x no longer supports Python 3.9, so if python --version prints 3.9 or earlier, install a newer Python first. Kedro works on macOS, Linux and Windows, in cmd.exe and PowerShell.
Use one virtual environment per project. This is not a Kedro rule; it is how Python work stays sane. Two projects that need different versions of pandas cannot share one environment.
The official documentation recommends the tool uv. With it, you can try Kedro without installing anything permanently:
uvx kedro new --starter spaceflights-pandas --name spaceflights
cd spaceflights
uv run kedro run --pipelines __default__
uvx is an alias for uv tool run: it runs Kedro in a throwaway environment, just long enough to create the project. After that, uv sync and uv run use the project's own .venv.
If you prefer plain Python, create and activate an environment, then install:
python -m venv .venv
source .venv/bin/activate # macOS and Linux
# .\.venv\Scripts\activate # Windows cmd or PowerShell
pip install kedro
Other routes work too: uv pip install kedro, poetry add kedro, or conda install -c conda-forge kedro. On zsh, quote any package with brackets, for example pip install "kedro[server]", or the shell will try to expand the brackets.
Verify the install with:
kedro info
It prints an ASCII-art banner, the installed Kedro version, and a list of installed plugins, or the words "No plugins installed". You can also run kedro --version (or kedro -V), which prints only the version. python -m kedro works as well. If the command is not found, your environment is not activated; that is nearly always the cause.
kedro-telemetry, and usage statistics are collected unless you opt out. It records a random ID, the command you ran with arguments masked, version numbers, your operating system, and counts of datasets, nodes and pipelines. It never collects your data. To opt out, set the environment variable KEDRO_DISABLE_TELEMETRY (or DO_NOT_TRACK) to any value, or create a project with kedro new --telemetry=no. Some employers require this, so it is worth knowing on day one.
- Create a fresh virtual environment, activate it, and run
pip install kedro. - Run
kedro infoand find the version line. It should start with 1. - Run
kedro starter listand read the aliases it prints. You will use one of them later. - Run
deactivate, thenkedro infoagain, and read the error. Now you know what "environment not active" looks like.
Your first project: what kedro new creates
Everything in Kedro starts from kedro new. It asks a few questions and builds a project skeleton. You can answer the questions interactively, or pass them as flags so the command runs without asking. For this guide we will build a small project from an empty skeleton so that every file in it is one you wrote and understand:
kedro new --name orders_demo --tools=data --example=n
The flags mean: --name is the project name; --tools=data adds the numbered data/ folders (other tools are lint, test, log, docs and pyspark, and you can pass all or none); --example=n says do not include the example pipeline. Kedro refuses project names whose Python package name would shadow something built into Python, such as email, json or import, because that would break imports in confusing ways. Change into the new folder (run ls to see its name), create and activate a virtual environment there, and install the dependencies:
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install "kedro-datasets[pandas-csvdataset]"
The second install adds the dataset classes we need. Kedro's core ships only a few in-memory datasets; everything that touches a real file format lives in a separate package called kedro-datasets, installed by groups. If you forget it, the first run fails with a message saying the class was not found, and that message is covered in the errors section.
The layout you now have looks like this:
orders-demo/
├── conf/
│ ├── base/ # catalog.yml, parameters.yml - committed to git
│ └── local/ # credentials.yml - git-ignored
├── data/ # 01_raw ... 08_reporting
├── notebooks/
├── src/orders_demo/
│ ├── __init__.py
│ ├── __main__.py
│ ├── pipeline_registry.py
│ ├── settings.py
│ └── pipelines/
├── tests/
├── pyproject.toml
├── requirements.txt
└── README.md
Read it as four zones.
conf/ is configuration. base holds settings shared by everyone and committed to git. local holds settings that are yours alone, above all credentials, and is ignored by git so passwords never end up in the repository.
data/ is where datasets live on your machine. The numbered folders are a convention, not a requirement: 01_raw for data exactly as received, 02_intermediate for lightly cleaned data, 03_primary for the tables you model from, 04_feature for features, 05_model_input for what goes into training, 06_models for trained models, 07_model_output for predictions, and 08_reporting for final tables and charts. The numbering tells a reader how far from the source a file is. Raw data is never edited in place.
src/orders_demo/ is your code. pipeline_registry.py lists the pipelines the project knows about. settings.py holds project-wide settings. Each pipeline gets its own folder under pipelines/, containing a nodes.py for the functions and a pipeline.py for the wiring.
pyproject.toml marks the project root and records its identity in a [tool.kedro] section: package_name, project_name and kedro_init_version. Kedro looks for this file to know it is inside a project, which is why you can run kedro run from any subfolder, and why running it from outside a project gives "Could not find the project configuration file 'pyproject.toml'".
git init and make a first commit straight after kedro new. The generated .gitignore already excludes conf/local/, so credentials stay out of the commit, and you will have a clean baseline to compare against when something breaks.
Starters, if you want an example
If you would rather begin from a working example, kedro new --starter <alias> copies a finished project. The official aliases in 1.7 are astro-airflow-iris, spaceflights-pandas, spaceflights-pyspark, databricks-iris and support-agent-langgraph. The most widely used is spaceflights-pandas, a three-pipeline project that predicts shuttle prices, and it is the project the official tutorial walks through. Older articles mention a pandas-iris starter; it has not existed since version 0.19.
- Run the
kedro newcommand above, thengit initand commit. - Open
pyproject.tomland find the[tool.kedro]section. Note the three keys. - Open
src/orders_demo/pipeline_registry.pyand read it. You do not need to change it yet; just notice that it discovers pipelines automatically. - Run
kedro run. It should finish almost at once, because there is nothing to run. Kedro is telling you the skeleton is valid.
Your first pipeline, step by step
We will build a tiny but realistic pipeline. A CSV of shop orders comes in; the pipeline cleans it, then summarises revenue per city. Three datasets and two nodes, which is enough to show every moving part.
Step 1: put in some raw data
Create data/01_raw/orders.csv with this content. It has a few deliberate flaws: a missing city, a missing amount, and inconsistent capitalisation.
order_id,city,amount,created
1,cairo,120.5,2026-01-03
2,Riyadh,80,2026-01-04
3, dubai ,200,2026-01-04
4,Cairo,,2026-01-05
5,,45,2026-01-06
6,riyadh,15.25,2026-01-07
7,Cairo,9,2026-01-08
Step 2: scaffold a pipeline
Kedro can create the folder structure for you:
kedro pipeline create orders
This makes src/orders_demo/pipelines/orders/ with nodes.py and pipeline.py, a conf/base/parameters_orders.yml for this pipeline's parameters, and a matching folder under tests/. The habit this builds is worth keeping: one pipeline, one folder, and its parameters and tests next to it.
Step 3: write the functions
Open src/orders_demo/pipelines/orders/nodes.py and replace its contents with two plain functions:
import pandas as pd
def clean_orders(orders: pd.DataFrame) -> pd.DataFrame:
"""Drop incomplete rows and tidy the city names."""
cleaned = orders.dropna(subset=["city", "amount"]).copy()
cleaned["city"] = cleaned["city"].str.strip().str.title()
cleaned["amount"] = cleaned["amount"].astype(float)
return cleaned
def summarise_by_city(orders: pd.DataFrame, min_amount: float) -> pd.DataFrame:
"""Count orders and total revenue per city, ignoring tiny orders."""
large = orders[orders["amount"] >= min_amount]
return large.groupby("city", as_index=False).agg(
orders=("order_id", "count"),
revenue=("amount", "sum"),
)
Look at what is missing. There is no pd.read_csv, no file path, no to_csv, and the threshold is an argument rather than a constant. These functions know nothing about Kedro. You could import them into a test and call them with a three-row DataFrame. That is exactly the point: the logic is separate from the plumbing.
Step 4: wire them into a pipeline
Open pipeline.py in the same folder:
from kedro.pipeline import Node, Pipeline
from .nodes import clean_orders, summarise_by_city
def create_pipeline(**kwargs) -> Pipeline:
return Pipeline(
[
Node(
func=clean_orders,
inputs="orders",
outputs="clean_orders",
name="clean_orders_node",
),
Node(
func=summarise_by_city,
inputs=["clean_orders", "params:min_amount"],
outputs="city_summary",
name="summarise_by_city_node",
),
]
)
Read this slowly, because it is the heart of Kedro. Each Node says: call this function, feed it these named inputs, and call its result by this name. The first node takes the dataset called orders and produces clean_orders. The second takes clean_orders (the very name the first produced, and this is the link that makes the order) plus a parameter called min_amount, and produces city_summary. The string "params:min_amount" is Kedro's syntax for "the parameter min_amount from the parameters file".
If a function takes several arguments, inputs is a list, matched to the function's arguments by position. If it returns several values, outputs is a list too. Everything after outputs in Node(...) is keyword-only, and name is optional but strongly recommended, because names appear in logs, in error messages and in --nodes on the command line.
from kedro.pipeline import node, pipeline (lowercase functions). Those wrappers still work in Kedro 1.x, but the capitalised Node and Pipeline classes are the preferred form, and this guide uses them throughout. You will see both in real projects.
Step 5: describe the data
Open conf/base/catalog.yml and add three entries:
orders:
type: pandas.CSVDataset
filepath: data/01_raw/orders.csv
clean_orders:
type: pandas.CSVDataset
filepath: data/02_intermediate/clean_orders.csv
city_summary:
type: pandas.CSVDataset
filepath: data/08_reporting/city_summary.csv
Each key is a dataset name, identical to the strings in pipeline.py. The type is the dataset class; pandas.CSVDataset is shorthand for kedro_datasets.pandas.CSVDataset. Note the spelling: Dataset with a lowercase s. Before kedro-datasets 2.0 it was DataSet, and old tutorials still show that. The old spelling produces a warning and then a "class not found" error.
Step 6: set the parameter
Open conf/base/parameters_orders.yml and set:
min_amount: 10
The pipeline's params:min_amount input now resolves to 10.
- Type in all six steps. Resist copy-pasting for the two code files; typing slows you down enough to notice each name.
- Before running anything, say aloud which node will run first and how you know. Then swap the order of the two nodes in the list and ask yourself whether the answer changes. It should not.
- Change
min_amountto 100 in the YAML file only. Notice that no Python file needs to change.
Running the pipeline and reading the output
Now run it:
kedro run
With no options, Kedro runs the pipeline named __default__. Where does that come from? Open pipeline_registry.py. The template calls find_pipelines(), which discovers every folder under pipelines/ that exposes a create_pipeline function, registers each under its folder name, and sets __default__ to the sum of all of them. So kedro run runs everything, and kedro run --pipelines orders runs just the orders pipeline.
The output is a stream of log lines. You will see Kedro loading the project configuration, then for each node a line announcing it is running, lines saying it is loading a dataset and saving a dataset, a line saying the node completed, and a final "Pipeline execution completed successfully." The exact wording and formatting differ slightly between versions and with Rich logging, so learn the shape rather than the text: load inputs, run function, save outputs, next node.
When it finishes, look at data/02_intermediate/clean_orders.csv and data/08_reporting/city_summary.csv. The cleaned file has five rows (the two incomplete ones were dropped), the cities are tidy ("Cairo", "Riyadh", "Dubai"), and the summary has one row per city. With min_amount: 10 the order from Cairo worth 9 is excluded, so Cairo shows one counted order rather than two. Change the parameter and run again to see the numbers move.
Notice what you did not have to do: write any code to load the CSV, or tell Kedro which node goes first.
Seeing the plan without running
Two commands show what Kedro intends to do, which is useful when a pipeline grows:
kedro registry list
kedro registry describe orders
The first lists the pipeline names. The second lists the nodes of one pipeline (or __default__ if you give no name).
- Run
kedro runand open both output files. Check the numbers against the raw CSV by hand. - Delete
data/02_intermediate/clean_orders.csvand run again. It is recreated. - Run
kedro registry listandkedro registry describe ordersand read the output. - Introduce a typo in one dataset name in
pipeline.pyand run. Read the error top to bottom, then fix it.
The Data Catalog in depth
You have used three catalog entries. The catalog can do much more, and most of a beginner's day-to-day Kedro work is editing it.
Anatomy of an entry
Here is one entry with most of the options a beginner will meet:
companies:
type: pandas.CSVDataset
filepath: data/01_raw/companies.csv
load_args:
sep: ","
save_args:
index: false
metadata:
kedro-viz:
layer: raw
typenames the dataset class. Kedro resolves a short name likepandas.CSVDatasetfirst againstkedro.io, then againstkedro_datasets, and finally treats it as a full dotted path to a class in your own code.filepathis a path or URL. Local paths are relative to the project root. Because Kedro uses thefsspeclibrary underneath, the same key acceptss3://,gcs://,abfs://,hdfs://andhttp(s)://URLs. Moving a dataset from your laptop to cloud storage is therefore a one-line change, with no code change in the nodes.load_argsare forwarded to the underlying reader (for the CSV dataset,pandas.read_csv), andsave_argsto the writer (DataFrame.to_csv). Hereindex: falsestops pandas writing the row numbers as an extra column, a classic annoyance.credentials(not shown) names a block in a credentials file, covered below.metadatais free-form information that Kedro ignores but tools can use. Thekedro-viz: layer:entry tells Kedro-Viz which layer to draw the dataset in. In old tutorials thelayer:key sat at the top level of the entry; it has since moved undermetadata.
Choosing a dataset type
The type depends on what you are storing. A few you will meet early:
| You have | Use |
|---|---|
| A CSV file | pandas.CSVDataset |
| A Parquet file (compact, fast, typed) | pandas.ParquetDataset |
| An Excel file | pandas.ExcelDataset |
| A trained scikit-learn model | pickle.PickleDataset |
| A chart made with matplotlib | matplotlib.MatplotlibDataset |
Each group needs its own dependencies. Install them by naming the group in brackets, for example pip install "kedro-datasets[pandas-csvdataset,pandas-parquetdataset]". If you skip this you will see either No module named '...'. Please install the missing dependencies for kedro_datasets... or Class '...' not found, and both mean the same thing: install the right group.
Versioning a dataset
Add versioned: true to an entry and Kedro stops overwriting the file. Each save goes into a new timestamped folder, <filepath>/<timestamp>/<filename>, and loads read the latest one by default. This keeps a history of a model or a report without any code. You can load an older one with kedro run --load-versions=city_summary:2026-01-01T00.00.00.000Z. Versioning is not available for http(s):// paths, and asking for it there raises "Versioning is not supported for HTTP protocols."
What happens to datasets that are not in the catalog
A dataset that appears in the pipeline but has no catalog entry is held in memory while the run lasts and discarded afterwards. That is useful for throwaway intermediates and dangerous for anything you want to keep or inspect. The rule of thumb for beginners: if you might want to look at it afterwards, give it a catalog entry. Also remember that memory-only datasets cannot be passed between separate tasks when a pipeline is later split across an orchestrator, so in production every intermediate gets an entry.
Many similar datasets: factories
If you have twenty CSV files that follow a naming pattern, writing twenty entries is tedious. A dataset factory matches names with a pattern, like a reversed f-string:
"{name}_data":
type: pandas.CSVDataset
filepath: data/01_raw/{name}_data.csv
Now any dataset whose name ends in _data, such as customers_data or shipments_data, is served by this one entry, with {name} filled in. The quotes are mandatory; without them YAML reads the braces as its own syntax and you get a parser error. Explicit entries always beat patterns, and there may be at most one catch-all pattern of your own. You do not need factories on day one, but you should recognise them when a colleague's catalog seems to be missing entries that the pipeline uses.
To see how Kedro resolves names, use:
kedro catalog describe-datasets
kedro catalog list-patterns
kedro catalog resolve-patterns
describe-datasets groups the datasets into those defined explicitly, those produced by a factory pattern and those that will default to memory. list-patterns shows the patterns in priority order, and resolve-patterns prints the fully resolved configuration. These replaced the older kedro catalog list and kedro catalog rank in version 1.0.
- Change the
city_summaryentry topandas.ParquetDatasetwith a.parquetfilepath, installkedro-datasets[pandas-parquetdataset], and rerun. You changed the file format without touching a line of Python. - Add
versioned: truetoclean_ordersand run twice. Look insidedata/02_intermediate/clean_orders.csvand find the timestamped folders. - Run
kedro catalog describe-datasets. Which datasets are explicit and which are defaults?
Parameters and credentials
Parameters: numbers that are not code
A parameter is a value that changes the behaviour of a pipeline and that you may want to adjust without editing functions: a threshold, a split ratio, a random seed, a list of columns. Parameters live in YAML files whose names start with parameters, inside conf/base/ (or another environment). You already used one.
Nested values work too:
min_amount: 10
model_options:
test_size: 0.2
random_state: 3
In a node, reference a single value as "params:min_amount", a nested one with dots as "params:model_options.test_size", or the whole dictionary as "params:model_options". The special name "parameters" passes every parameter at once. The function receives ordinary Python values: a float, a dict.
You can override a parameter for one run, without editing any file, from the command line:
kedro run --params="min_amount=50"
The syntax is key=value, separated by commas for several (--params="min_amount=50,model_options.test_size=0.3"), with dots for nested keys. Kedro turns values into integers or floats where it can and leaves the rest as strings. A colon (key:value) was the syntax before 0.19 and some documentation examples still show it; use the equals sign. If a value contains spaces, quote the whole argument.
Credentials: secrets that stay out of git
Some datasets need a password, key or token, for example an S3 bucket or a database. Never put these in catalog.yml, because that file is committed. Put them in conf/local/credentials.yml, which git ignores:
dev_s3:
client_kwargs:
aws_access_key_id: your-key-id
aws_secret_access_key: your-secret
and reference the block by name from the catalog:
orders:
type: pandas.CSVDataset
filepath: s3://my-bucket/orders.csv
credentials: dev_s3
Anything in conf/local/ is ignored by git, and kedro package leaves those files out of the packaged configuration. On a server or in a CI job there is no conf/local/credentials.yml, so the usual approach is to read secrets from environment variables. Inside a credentials file only, you can write ${oc.env:AWS_ACCESS_KEY_ID} and Kedro substitutes the variable's value. For AWS specifically, the standard variables AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY and AWS_SESSION_TOKEN are also picked up directly.
git status before every commit in a project that has a credentials file.
- Add the nested
model_optionsblock to your parameters file. You do not have to use it yet. - Run
kedro run --params="min_amount=100"and confirm the summary shrinks. Then run plainkedro runand confirm it returns to the YAML value. - Create
conf/local/credentials.ymlwith a fake block, rungit status, and confirm the file does not appear.
Configuration environments
The conf/ folder holds more than one environment, and the idea is simple but easy to misread. An environment is a subfolder. base is the foundation and is always loaded. On top of it Kedro loads a second environment, local by default, and where the same top-level key appears in both, the second one wins. By default this replacement is whole-key: if local defines model_options, it replaces base's entire model_options, not only the keys it mentions.
The effect is that one project can behave differently in different places with no code change. base holds the settings that are right for everyone. Your local folder holds your own tweaks, such as a smaller sample file to keep development fast. You can define further environments (dev, staging, prod) as folders and choose one at run time:
kedro run --env=staging
Or set the environment variable KEDRO_ENV=staging; if both are present the command-line flag wins.
Two warnings follow from how this works. First, within one environment, the same top-level key must not appear in two files, or Kedro raises ValueError: Duplicate keys found in ... naming both files. That protects you from silently losing a setting. Second, if you name an environment that does not exist, you get a MissingConfigException about a configuration path that does not exist or is not a directory. Create conf/<env>/ first.
Kedro also has a templating feature built on a library called OmegaConf, so a value in YAML can refer to another with ${...}. You do not need it yet; mid-level covers it.
- Create
conf/local/parameters_orders.ymlcontainingmin_amount: 50and runkedro run. Which value is used, and why? - Create
conf/staging/with a different catalog path forcity_summaryand runkedro run --env=staging. - Delete the local file again so your later results match this guide.
The everyday commands, grouped by what you want to do
You do not need to memorise the full command list. Group it by intent.
I want to run things
kedro run # the default pipeline
kedro run --pipelines orders # one pipeline
kedro run --pipelines orders,training # several at once
kedro run --env=staging # another configuration environment
kedro run --params="min_amount=50" # override parameters
Use --pipelines (plural). Kedro 1.2 introduced it so several pipelines can run together, and it deprecated the singular --pipeline flag: the old flag still works but prints a deprecation warning, and using both at once is an error. Most online tutorials still show the old flag.
I want to run only part of a pipeline
While developing you rarely want the whole thing. Kedro offers several ways to pick a slice:
kedro run --nodes=clean_orders_node # specific nodes, by name
kedro run --tags=cleaning # nodes that carry a tag
kedro run --from-nodes=summarise_by_city_node
kedro run --to-nodes=clean_orders_node
kedro run --from-inputs=clean_orders
kedro run --to-outputs=city_summary
--from-nodes runs that node and everything after it; --to-nodes runs everything up to and including it. --from-inputs and --to-outputs do the same by dataset name. Tags are labels you add to a node in code, Node(..., tags=["cleaning"]), and --tags selects nodes carrying any of the tags you list. Names and tags may contain only letters, digits, hyphens, underscores and full stops.
A very useful flag is --only-missing-outputs, which skips any node whose saved outputs already all exist. After a failure halfway through a long run, it lets you resume without recomputing the early steps.
If your filters match nothing, Kedro says "Pipeline contains no nodes after applying all provided filters." That message means a typo or a wrong name, not a broken pipeline.
I want to look around
kedro info # version and plugins
kedro registry list # pipeline names
kedro registry describe orders # nodes in a pipeline
kedro catalog describe-datasets # how each dataset is defined
kedro starter list # available starters
I want to create or change things
kedro new # new project
kedro pipeline create <name> # new pipeline folder
kedro pipeline delete <name> # remove one
I want to work interactively
kedro ipython # IPython with the project loaded
kedro jupyter lab # JupyterLab with a Kedro kernel
kedro jupyter notebook # classic notebook
The next sections cover the interactive and visual tools.
pyproject.toml, so you can run project commands from any subdirectory. The only commands that work outside a project are the global ones: kedro new, kedro starter list and kedro info.
- Add
tags=["cleaning"]to the first node. Runkedro run --tags=cleaningand confirm only one node runs. - Run
kedro run --from-nodes=summarise_by_city_node. It works becauseclean_orderswas saved to disk by an earlier run. Now delete that file and run it again, and read the error. - Try
--only-missing-outputstwice in a row and watch what is skipped.
Exploring with notebooks and Kedro-Viz
Notebooks, the Kedro way
Notebooks are the best tool for looking at data, and Kedro does not ask you to give them up. It asks you to keep them as a window onto the project and not as the place where the logic lives. Start a notebook with your project already loaded:
kedro jupyter lab
This creates a kernel named after your package, shown as Kedro (orders_demo), and opens JupyterLab. In the kernel, these variables already exist: catalog, context, pipelines and session. Try:
catalog.keys() # all dataset names
df = catalog.load("clean_orders") # read a dataset the way a node would
df.head()
catalog.load("parameters") # every parameter as a dictionary
pipelines["__default__"] # the pipeline object
catalog.load("name") is the important one. It returns a dataset exactly as your nodes would see it, so you can experiment on real data and, once a snippet works, paste it into nodes.py as a function. That is the healthy loop: explore in the notebook, then promote working code into a node.
If you launch Jupyter some other way, load the project with two magic commands: %load_ext kedro.ipython, then %reload_kedro <path to project root>. If you see "Line magic function '%reload_kedro' not found", the extension was not loaded. kedro ipython does the same for a terminal IPython session. If it complains that the module IPython is not found, install the project requirements first.
A warning for those who read older material: the catalog method catalog.list() no longer exists. Since 1.0 use catalog.keys() or catalog.filter("pattern").
Kedro-Viz: see the pipeline as a diagram
Kedro-Viz is a separate package that draws your pipeline as an interactive graph, with datasets and nodes as shapes and the dependencies as arrows. It is the best way to explain a pipeline to someone else, and also the quickest way to spot a mistake, such as a node with no inputs.
pip install kedro-viz
kedro viz run
It opens a browser tab on 127.0.0.1:4141. (The shorter kedro viz also works.) Click a node to see its code and parameters, click a dataset to see its type and file path, and use the filters on the side to show only some pipelines, tags or layers. The layers you set under metadata: kedro-viz: layer: in the catalog are drawn as horizontal bands, so a pipeline reads from raw data at one edge to reporting at the other. Use --no-browser if you are on a remote machine and want to open the address yourself.
- Run
kedro jupyter lab. In a new notebook under the Kedro kernel, runcatalog.keys()and loadorders. Compare it to the cleaned version. - Install Kedro-Viz and run
kedro viz run. Find your two nodes and three datasets. - Add
metadata: kedro-viz: layer: rawtoordersin the catalog (nested as in the earlier example) and refresh the diagram.
Growing the project: several pipelines, namespaces and reuse
A project of one pipeline is a demo. Real projects have several: data preparation, training, evaluation. Kedro's answer is the modular pipeline: each one lives in its own folder with its own nodes.py and pipeline.py, and exposes a create_pipeline function. You already made one with kedro pipeline create orders. Making a second is the same command:
kedro pipeline create reporting
Because pipeline_registry.py uses find_pipelines(), the new pipeline is registered the moment its folder has a create_pipeline. Then kedro run runs both (the __default__ pipeline is their sum) and kedro run --pipelines reporting runs only the new one. Datasets are shared between pipelines by name: if reporting consumes city_summary, it finds it in the same catalog, produced by orders. Keeping one pipeline per folder means each can be understood, tested and run on its own.
Namespaces: using a pipeline twice
Sometimes the same pipeline should run twice on different inputs, for example one model trained on two regions. Wrapping it in a namespace prefixes its node names, dataset names and parameters with a label, so two copies do not collide:
from kedro.pipeline import Pipeline
base = create_pipeline()
riyadh = Pipeline(base, namespace="riyadh")
cairo = Pipeline(base, namespace="cairo")
Now the dataset clean_orders becomes riyadh.clean_orders in the one copy and cairo.clean_orders in the other. Namespaces are also how Kedro-Viz collapses groups of nodes into one expandable box. Run one with kedro run --namespaces=riyadh; the flag accepts several, comma-separated. The full dotted-name rules are covered at mid-level; for now remember that the dot character in a dataset name is reserved for namespaces, so avoid dots in your own dataset names.
If you reuse a pipeline without a namespace you will see "Pipeline nodes must have unique names". That message means the same node name appears twice, and a namespace is the fix.
- Create a second pipeline called
reportingwith one node that readscity_summaryand returns the top city. Give its output a catalog entry. - Run
kedro registry listand confirm both pipelines appear, thenkedro run --pipelines reporting. - Open Kedro-Viz and confirm the two pipelines connect through
city_summary.
Testing a node
The payoff of pure functions is that testing is ordinary Python. No Kedro needed:
import pandas as pd
from orders_demo.pipelines.orders.nodes import clean_orders, summarise_by_city
def test_clean_orders_drops_incomplete_rows():
raw = pd.DataFrame(
{"order_id": [1, 2], "city": ["cairo", None], "amount": [10.0, 5.0]}
)
result = clean_orders(raw)
assert len(result) == 1
assert result.iloc[0]["city"] == "Cairo"
def test_summary_ignores_small_orders():
orders = pd.DataFrame(
{"order_id": [1, 2], "city": ["Cairo", "Cairo"], "amount": [100.0, 1.0]}
)
summary = summarise_by_city(orders, min_amount=10)
assert summary.iloc[0]["orders"] == 1
Run it with pytest from the project root. If you created the project with the test tool, pytest is already in the requirements; otherwise pip install pytest. The test builds a three-row DataFrame, calls the function and checks the answer. It does not read a file, does not need a catalog and runs in milliseconds. That is the benefit of keeping I/O out of functions.
Common errors and how to read them
Kedro errors are long because they include the chain of causes. The rule is: read the last line first, then the one above it. The final exception says what went wrong; the chained ones say where. Here are the ones you will meet in your first week.
| What you see | What it means | What to do |
|---|---|---|
Could not find the project configuration file 'pyproject.toml' |
You are outside a Kedro project | cd into the project |
DatasetError: Class 'pandas.CSVDataSet' not found, is this a typo? |
Old capital-S spelling, or kedro-datasets not installed |
Use CSVDataset and install kedro-datasets[pandas-csvdataset] |
No module named '...'. Please install the missing dependencies for kedro_datasets... |
A dataset group's dependencies are missing | pip install "kedro-datasets[<group>]" |
DatasetNotFoundError: Dataset 'x' not found in the catalog |
A name is used but defined nowhere | Compare the name in code and in catalog.yml |
Pipeline input(s) {'x'} not found in the DataCatalog |
A pipeline input has no data behind it | Add a catalog entry, or fix the name |
DatasetError: Failed while loading data from dataset ... |
The file, path or library is wrong | Read the chained exception underneath |
DatasetError: Saving 'None' to a 'Dataset' is not allowed |
A node forgot to return |
Return the value |
Duplicate keys found in <f1> and <f2> |
One key in two files of the same environment | Rename it or move it |
ParserError: Invalid YAML or JSON file ... |
A YAML syntax mistake, often an unquoted {name} pattern |
Quote it and fix indentation |
OutputNotUniqueError |
Two nodes write the same dataset | Rename one output |
Failed to save outputs of node ... returned N output(s), whereas the node definition contains M |
Return values do not match outputs |
Align them |
Pipeline contains no nodes after applying all provided filters. |
A --tags or --nodes filter matched nothing |
Fix the filter |
Your Kedro project version X does not match Kedro package version Y |
A 0.19 project opened with Kedro 1.x | Migrate the project, then update kedro_init_version |
outputs=["a", "b"] the function must return two values, such as return a, b. If it returns one, Kedro stops with the "returned N output(s)" error. The reverse holds too: a function that returns a value while the node was declared with outputs=None produces a warning that the value is being thrown away.
Two habits make errors much cheaper. Always give nodes a name, because then the log line says which step failed in words you chose. And when a dataset fails to load, open a notebook and call catalog.load("that_name") directly. The error is then isolated from the pipeline, and you can try the fix interactively.
Finally, watch for two warnings. KedroDeprecationWarning means you used a flag or function scheduled for removal (the old --pipeline flag is the common one); it still works today, but fix it now. A KedroExperimentalWarning means you touched a feature whose design may still change.
- Break things on purpose, one at a time: misspell a dataset in the catalog, misspell it in the pipeline, return only one value from a function declared with two outputs, and write
CSVDataSetwith a capital S. For each, find the line of the error that tells you the cause. - Fix each one before moving on, and run
kedro runto confirm.
Putting it all together
Here is the whole project in one pass, the way you would set it up for a colleague. Every command is one you have already used.
kedro new --name orders_demo --tools=data --example=n
cd orders-demo
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install "kedro-datasets[pandas-csvdataset]" pytest
git init && git add . && git commit -m "Kedro skeleton"
kedro pipeline create orders
Then edit five files:
data/01_raw/orders.csv, the raw data.src/orders_demo/pipelines/orders/nodes.py, the two pure functions.src/orders_demo/pipelines/orders/pipeline.py, the two nodes.conf/base/catalog.yml, the three datasets.conf/base/parameters_orders.yml, themin_amountparameter.
Run it, inspect it and test it:
kedro run
kedro registry describe orders
kedro run --params="min_amount=50"
pytest
kedro viz run
Step back and notice the properties of what you built. The functions contain no paths and no magic numbers, so they are reusable. The wiring is explicit and named, so a stranger can read it. The data locations are in YAML, so moving from a laptop to cloud storage means editing one file. The parameter can be changed from the command line, so experiments need no code edits. The code is tested without running the pipeline. And the project has the same shape as every other Kedro project in the world, so any Kedro user can find their way around it.
A final check for people working at companies in the Gulf or Egypt: where your data sits matters for regulation. Because locations are just filepath values, you can point a dataset at a storage bucket in a Middle East cloud region and keep the code identical, which is exactly the kind of change this design makes easy. Confirm the residency rules that apply to your employer, but know that Kedro will not stand in the way.
- Rebuild the whole project from an empty folder without looking at this guide. Use it only when stuck, and note each place you got stuck.
- Extend it with a third node that counts orders per day from the
createdcolumn, saved as a new dataset with a catalog entry and a test. - Push it to a git repository and ask a friend to clone it, install the requirements and run
kedro run. Fix whatever breaks. That is the real test of reproducibility.
What you can now do, and what comes next
You can now create a Kedro project, explain its four nouns (node, pipeline, catalog, configuration), write pure functions and wire them into a pipeline, describe data in YAML, pass parameters and keep credentials out of git, choose a configuration environment, run all or part of a pipeline, explore it in a notebook, draw it with Kedro-Viz, split a project into modular pipelines, test a node, and read Kedro's common errors.
Here is what to check you can do without looking anything up:
| Can you... | The answer |
|---|---|
| Say what decides node order? | Data dependencies between names, not list order |
| Move a file location without touching Python? | Edit filepath in catalog.yml |
| Change a threshold for one run? | kedro run --params="key=value" |
| Keep a secret out of git? | conf/local/credentials.yml, referenced by name |
| Run only the second half of a pipeline? | --from-nodes or --from-inputs |
| Find why a name is not found? | Compare the string in code and in the catalog |
| Check the version and plugins? | kedro info |
Mid-level takes every one of those topics further: how the catalog, config loader and runners work internally, dataset factories and templating with OmegaConf, parameter and dataset validation, hooks, the choice of runner, packaging a project as a wheel, and deploying it with Docker and orchestrators.
Senior then covers what you own when Kedro is your team's platform: the security model, scaling limits, upgrade and migration planning, custom datasets and runners, and where Kedro stops and another tool should take over.
Natural next reads in this catalogue are Airflow for scheduling a Kedro project, MLflow for experiment tracking, DVC for versioning data, and Docker for packaging it all so it runs anywhere.
Sources
- Kedro documentation home
- Install Kedro
- Kedro concepts
- Commands reference
- Create a new project
- Project tools
- Starters
- Configuration basics
- Parameters and credentials
- The Data Catalog
- Dataset factories
- Nodes
- Modular pipelines
- Namespaces
- Run a pipeline
- Kedro and notebooks
- Telemetry
- Migration guide
- Kedro release notes
- Kedro-Viz documentation
- kedro-datasets documentation