This is part one of three. It covers everything you need to write, run, debug and inspect real Metaflow projects on your own laptop, written for someone who has never used the tool. By the end you can turn a messy notebook into a structured pipeline, run steps in parallel, pass parameters, survive failures without losing work, look back at any past run from Python, and build a small report that opens in your browser. Mid-level and Senior take the same topics further, into remote compute, scheduling and running Metaflow as a platform. Nothing here is thrown away.
This guide is written against Metaflow 2.19, the current release line. Metaflow keeps its promises about backwards compatibility, so nearly everything you learn will still work next year, but where a feature is recent the text says which version introduced it.
Each section ends with a Try it task. Do them as you go. Pipelines are something you understand by watching your own one break and recover, not by reading about it.
What Metaflow is, and the problem it solves
Machine learning work usually starts in a notebook. You load data, clean it, train a model, plot a score. It works, and then someone asks for it to run every night, or to train on a bigger machine, or to be reproduced by a colleague who is not you. At that point the notebook becomes the problem: its cells were run in an order nobody wrote down, its state lives in a kernel that dies when you close the laptop, and "the model I showed on Tuesday" is not stored anywhere.
Metaflow is a Python library that turns that work into a flow: a list of named steps that run in a defined order, with everything they produce saved automatically. It was created at Netflix to let data scientists build and deploy their own pipelines without becoming infrastructure engineers, and it is open source. You write ordinary Python. You do not write YAML, you do not learn a new language, and you do not need a server to start.
Three promises explain why people adopt it.
Your work is structured. A flow is a class, and each stage is a method. The order is explicit, so a colleague can read the file and understand the pipeline in a minute.
Your work is saved. Every value you assign to self inside a step is stored automatically and versioned per run. You never write code to save intermediate results, and you can open any old result later.
Your work can grow. The same file that runs on your laptop can later run its steps on a cluster, be scheduled by a production orchestrator, and react to events, without rewriting it. This guide only covers the laptop part, deliberately: the laptop is where you learn the model, and everything after that is the same model with more machinery behind it.
What came before? Teams either wrote one long script (fast to start, impossible to resume), or adopted a general orchestrator such as Airflow and wrote tasks as operators in that system's own vocabulary. Metaflow sits in between: it is designed around the data scientist's workflow (write code, run it, look at the results, iterate) and treats orchestration as something you switch on afterwards. If you have read the guides on MLflow or Prefect, it helps to know the division of labour. MLflow records experiments and models, Prefect and Airflow schedule work, and Metaflow is the authoring and tracking layer in which a pipeline is written, with its own record of every run.
- Take one notebook or script you have written. List its stages in order (load, clean, train, evaluate), one line each.
- For each stage, write what it needs from the previous one and what it produces. Those lists are the steps and artifacts of your first flow.
The mental model: flows, steps, artifacts, runs
Metaflow has a small vocabulary. Learn these nouns well and every later topic is a variation.
A flow is a class that inherits from FlowSpec. It is the smallest thing you can run or schedule. You run it by executing the file as a script: python myflow.py run.
A step is a method marked with the @step decorator. Steps are the nodes of a graph, and the documentation calls a step the smallest resumable unit of computation. That phrase matters: if a step fails, you can fix the code and continue from that step without redoing the ones before it. A sensible step does one meaningful thing and, as a rule of thumb from the documentation, should take well under an hour.
A transition connects steps. The last line of every step is a call to self.next(...), naming the step that comes next. A flow must begin with a step named start and finish with a step named end. (Since version 2.19.23 you can instead mark the entry and exit steps with @step(start=True) and @step(end=True), but the plain names are what you will meet in almost every example, so this guide uses them.)
A run is one execution of a flow from start to end. Each run gets an ID, which for local runs is a simple counting number: 1, 2, 3. A task is one step executing inside a run. Usually a step produces exactly one task, but a step that fans out over a list produces one task per item, which is how Metaflow does parallel work. A task is addressed by a pathspec of the form FlowName/run_id/step_name/task_id, for example HelloFlow/3/start/1. You will see pathspecs in logs and use them to look things up, so get comfortable reading them.
An artifact is any value you store on self inside a step. If you write self.accuracy = 0.93, then accuracy is an artifact: Metaflow saves it when the task finishes, attaches it to that exact task in that exact run, and makes it available to every later step. Ordinary local variables, the ones without self., are not saved and vanish when the step ends. This is the single most important rule in the guide, and Section "Everyday commands" returns to it, because breaking it is the most common beginner mistake.
Two more nouns you will meet early. A parameter is a value you pass when starting a run, such as a learning rate, set once and then read-only. The datastore is where artifacts are kept. On your laptop that is a hidden folder called .metaflow in your working directory, and in a team setting it is a cloud bucket. Because the location is configuration rather than code, the flow file does not change when the storage does.
self. A step cannot see a local variable from the step before, in the same way a notebook cell cannot see variables from a kernel that was restarted. Once that feels natural, the rest of Metaflow follows.
- Without running anything, say what the pathspec
TrainFlow/12/evaluate/4identifies: the flow, the run, the step and the task. - Say which of these survive into the next step:
self.rows = 10,rows = 10. Explain why.
Installing Metaflow and checking the setup
Metaflow is a Python package and runs on macOS and Linux. Use a virtual environment so the install does not touch your system Python:
python -m venv .venv
source .venv/bin/activate
pip install metaflow
If you prefer uv, the documented route is:
uv init myflow
cd myflow
uv add pandas metaflow
The package depends on only two libraries, requests and boto3, so the install is quick. Check that it worked:
metaflow --version
python -c "import metaflow; print(metaflow.__version__)"
Either command should print a version number in the 2.19 line or later. To see which metadata provider your setup uses (more on that below), run metaflow status. On a fresh install it reports a local setup, which is exactly what you want for learning.
Windows is not supported natively. Metaflow's local runtime relies on a POSIX-only part of Python, so on a Windows machine you will see ImportError: No module named fcntl. Install WSL 2 with Ubuntu and do everything inside it, exactly as if you were on Linux. If you cannot install anything, the browser-based Metaflow Sandbox offered by Outerbounds, the company founded by Metaflow's creators, runs flows with no local setup. Either path teaches the same thing.
Metaflow can also pull in a set of tutorial flows that are useful for checking your install:
metaflow tutorials pull
metaflow tutorials list
The first command copies the tutorials into a folder called metaflow-tutorials in your current directory. Open 00-helloworld in it and you will recognise everything you are about to write yourself.
You do not need Docker, Kubernetes, a cloud account, or a database for anything in this guide. Everything runs locally and stores results under .metaflow. If your employer or university later gives you a configured cloud setup, a configuration file on your machine, not your code, will point Metaflow at it. That is covered near the end of the guide.
- Create a virtual environment, install Metaflow, and run
metaflow --version. - Run
metaflow tutorials pulland list the folder it creates. Open the hello world flow and read it without running it.
Your first flow, step by step
Create a file called hello_flow.py. Type it rather than pasting it; the shape matters.
from metaflow import FlowSpec, step
class HelloFlow(FlowSpec):
"""A flow that greets the world and keeps a message."""
@step
def start(self):
print("HelloFlow is starting")
self.next(self.hello)
@step
def hello(self):
self.message = "Hello, Metaflow"
print("Saying hello")
self.next(self.end)
@step
def end(self):
print("The message was:", self.message)
if __name__ == "__main__":
HelloFlow()
Read it top to bottom. HelloFlow inherits from FlowSpec, which gives it all of Metaflow's behaviour. Each method has @step above it. Each step except the last ends with self.next(self.something), passing the next method itself, not a string. The end step has no self.next because nothing follows it. The last two lines create the flow when the file is executed as a script; without them, running the file would do nothing.
Notice self.message = "Hello, Metaflow" in the hello step and the read in end. The value crossed a step boundary because it is an artifact.
Now use the commands Metaflow gives you, in this order. First check the file for mistakes:
python hello_flow.py check
Then draw the graph:
python hello_flow.py show
show prints each step with its docstring and what it leads to, which is a fast way to confirm the flow is wired the way you think. Finally run it:
python hello_flow.py run
You will see output shaped like this, though the IDs and timestamps will differ:
Workflow starting (run-id 1):
[1/start/1 (pid 40211)] Task is starting.
[1/start/1 (pid 40211)] HelloFlow is starting
[1/start/1 (pid 40211)] Task finished successfully.
[1/hello/2 (pid 40217)] Task is starting.
[1/hello/2 (pid 40217)] Saying hello
[1/hello/2 (pid 40217)] Task finished successfully.
[1/end/3 (pid 40222)] Task is starting.
[1/end/3 (pid 40222)] The message was: Hello, Metaflow
[1/end/3 (pid 40222)] Task finished successfully.
Done!
Read the bracketed prefix. 1/start/1 is run 1, step start, task 1. The number in parentheses is the operating system process that executed it, because Metaflow runs each task as its own process. Every print you wrote appears after its task's prefix, so when a flow is noisy you can still tell which task said what.
Two things happened that you did not code. A folder named .metaflow appeared next to your file, holding run 1's artifacts. And every task ran in a separate process, which is why steps cannot share local variables, and why a crash in one step cannot corrupt another.
Run it a second time. You get run 2, stored alongside run 1. Nothing was overwritten. This is the behaviour that makes Metaflow useful for experiments: every attempt is preserved and comparable.
python hello_flow.py run prints nothing or shows a usage message, check that the last lines are if __name__ == "__main__": HelloFlow(). Without it there is no flow to run. Also check that every step ends in self.next(...) except end; the linter will say so if not.
- Type in the flow, then run
check,showandrunin that order. - Run it twice. Look inside
.metaflowand confirm that run 1 was kept when run 2 happened. - Delete the
self.next(self.end)line inhelloand runcheckto read the error it gives. Put it back.
Artifacts: how data moves between steps
A flow is only useful if steps can hand data to each other, and the rule is simple: anything assigned to self is an artifact and is saved. Here is a flow that loads some numbers, computes statistics, and uses them later:
from metaflow import FlowSpec, step
class StatsFlow(FlowSpec):
@step
def start(self):
self.numbers = [4, 8, 15, 16, 23, 42]
self.next(self.compute)
@step
def compute(self):
total = sum(self.numbers)
self.mean = total / len(self.numbers)
self.biggest = max(self.numbers)
self.next(self.end)
@step
def end(self):
print(f"mean={self.mean:.2f} biggest={self.biggest}")
if __name__ == "__main__":
StatsFlow()
In compute, self.numbers was read, because it was saved by start. The variable total was a plain local, so it is gone after the step, while self.mean and self.biggest persist. By the end of the run, four artifacts exist: numbers, mean, biggest, and anything else you assigned.
There are three practical consequences worth knowing on day one.
Artifacts are saved when the task finishes, using Python's pickle. Most objects (numbers, strings, lists, dictionaries, pandas DataFrames, NumPy arrays, scikit-learn models) pickle fine. Open file handles, database connections and locks do not. If you store one, the step fails at the end with a pickling error, not where you created the object. The fix is to keep those objects as local variables inside a step and save only the results you need.
Artifacts are copied, not shared. A step that modifies self.numbers modifies its own task's version. Later steps see the new version; earlier steps are unchanged. That is a feature: old runs stay exactly as they were.
Size matters. Saving a five-gigabyte DataFrame as an artifact in five consecutive steps writes it several times. Metaflow deduplicates identical data within a flow, which helps, but you should still avoid storing large intermediate objects you will never read again. Keep them local if only one step needs them.
An everyday bug is worth naming now. If you write df = load_data() in one step and use df in the next, you will get AttributeError: 'StatsFlow' object has no attribute 'df', because the variable was never stored on self. The fix is always the same: assign it to self if a later step needs it.
self.something. If it is not needed, leave it local and save the space. Decide deliberately for each variable, because this one rule explains most "attribute not found" errors beginners report.
- Run
StatsFlow. Changeself.meanto a plainmeanand read the error whenendtries to use it. - Add a third step between
computeandendthat printsself.numbers, to confirm that an artifact set early in the flow is visible late in the flow.
Branches and joins: running work in parallel
Some work is naturally parallel, such as training two different models on the same data. Metaflow lets a step name more than one next step. Those steps, called branches, run at the same time, and you must bring them back together in a join step.
from metaflow import FlowSpec, step
class BranchFlow(FlowSpec):
@step
def start(self):
self.data = list(range(1, 11))
self.next(self.sum_it, self.max_it)
@step
def sum_it(self):
self.result = sum(self.data)
self.next(self.join)
@step
def max_it(self):
self.result = max(self.data)
self.next(self.join)
@step
def join(self, inputs):
self.sum_result = inputs.sum_it.result
self.max_result = inputs.max_it.result
self.merge_artifacts(inputs, exclude=["result"])
self.next(self.end)
@step
def end(self):
print(self.sum_result, self.max_result, len(self.data))
if __name__ == "__main__":
BranchFlow()
self.next(self.sum_it, self.max_it) with two arguments starts both branches. The join step is recognised by its second parameter, conventionally called inputs. It receives one entry per branch, and you can reach a branch by the name of its step: inputs.sum_it.result.
Both branches set an artifact called result, and they disagree. Metaflow cannot decide which is right, so it refuses to guess. Inside join you copy out the ones you want under distinct names (self.sum_result, self.max_result). Then self.merge_artifacts(inputs, exclude=["result"]) carries over the artifacts that are not in conflict, here data, and skips the one you excluded. If you forget to exclude or set the conflicting artifact, the run stops with "cannot merge the following artifacts due to them having conflicting values" and tells you to set them explicitly before calling merge_artifacts. Use either exclude or include on that call, never both.
The rules Metaflow enforces, and that check will report, are these. Every branch must be joined before end. A join step cannot itself start a new split: if it must, split it in two steps. And a step with an extra argument that is not preceded by a split is reported as a join with nothing to join.
- Run
BranchFlow. Then removeexclude=["result"]and read the conflict error. - Add a third branch that computes the minimum, and extend the join to collect it.
Foreach: one task for every item in a list
Branches are fixed in the code. Often the number of parallel jobs depends on your data: one per file, one per customer, one per hyperparameter setting. That is what foreach is for.
from metaflow import FlowSpec, step
class ForeachFlow(FlowSpec):
@step
def start(self):
self.learning_rates = [0.001, 0.01, 0.1]
self.next(self.train, foreach="learning_rates")
@step
def train(self):
lr = self.input
self.lr = lr
self.score = 1.0 - abs(0.01 - lr)
self.next(self.join)
@step
def join(self, inputs):
best = max(inputs, key=lambda i: i.score)
self.best_lr = best.lr
self.best_score = best.score
self.next(self.end)
@step
def end(self):
print("best learning rate:", self.best_lr)
if __name__ == "__main__":
ForeachFlow()
The phrase foreach="learning_rates" names an artifact that must hold a list. Metaflow creates one train task per item, running them in parallel, up to a local limit. Inside each task, self.input is that task's item. Note the name of the argument: it is a string naming the variable, not the variable itself.
The join step receives all the tasks. You can loop over them with for i in inputs, or use max(inputs, ...) as above, and each one exposes the artifacts that task saved. Here each train task saved lr and score, so the join reads i.score. Saving self.lr = lr was deliberate: it makes the input part of the record so the join can tell which result belongs to which setting.
The limits are worth memorising. Locally at most 16 tasks run concurrently by default, controllable with --max-workers. A foreach that produces more than 100 tasks is rejected with a message telling you to raise --max-num-splits; that safety cap exists to stop an accidental list of a million items from starting a million tasks. An empty list also fails with a "zero splits" message, which usually means an earlier step forgot to fill it.
python foreach_flow.py run --max-num-splits 1000, but consider grouping items so each task handles a batch. Every task has start-up cost.
Since version 2.18, flows can also choose between branches based on a value (a conditional transition) or loop back to an earlier step (recursion). They are useful and they work locally, but they are not needed to get started, so they are left to the Mid-level guide.
- Run
ForeachFlow, then add two more learning rates and watch the extra tasks appear in the output. - Run it with
--max-workers 1and notice that the tasks now run one after another.
Parameters: changing a run without editing code
A learning rate hard-coded inside a step means editing the file for every experiment, and the edit history becomes your only record of what you tried. A parameter moves that value to the command line. You declare it as a class attribute:
from metaflow import FlowSpec, Parameter, step
class ParamFlow(FlowSpec):
alpha = Parameter("alpha", default=0.01, help="Learning rate")
epochs = Parameter("epochs", default=5, help="Number of passes")
label = Parameter("label", default="baseline", help="Name for this run")
@step
def start(self):
print(f"alpha={self.alpha} epochs={self.epochs} label={self.label}")
self.loss = self.alpha * 100 / self.epochs
self.next(self.end)
@step
def end(self):
print("loss:", self.loss)
if __name__ == "__main__":
ParamFlow()
Pass values on the command line, after run:
python param_flow.py run --alpha 0.5 --epochs 10 --label bigger
python param_flow.py run --help
The second command lists every parameter with its help text and default, so a colleague can discover what your flow accepts without reading the code. Metaflow infers each parameter's type from its default: 0.01 makes a float, 5 makes an integer, and a string default makes a string. You can also set type=int explicitly, mark required=True so there is no default, or pass structured data with type=JSONType, where the command line supplies a JSON string.
Three facts keep you out of trouble. A parameter is read-only for the whole run, so you cannot reassign self.alpha in a step; copy it to another artifact if you need a variant. Parameters are stored as artifacts of the run, which means you can always look up later exactly what a run was given. And for that same reason, never pass a password or an API key as a parameter, because it would be saved in the run's record.
A fourth fact catches people the first time: the order of words matters. Options that belong to run, including your parameters, go after run. A small group of top-level options that change how Metaflow itself behaves, such as --environment, --with or --datastore, go before run:
python param_flow.py --with retry run --alpha 0.5
Write that shape as python flow.py [top-level options] command [command options] and you will rarely get it wrong. Metaflow also has a separate Config object, evaluated when a flow is deployed, used to set up decorators and production deployments. As a beginner you only need parameters; the Mid-level guide explains when to reach for Config instead.
You can also attach labels to a run with --tag, which you can repeat:
python param_flow.py run --tag experiment-7 --tag gpu-later
Tags help you find runs afterwards from the Client API, which a later section shows.
- Run
ParamFlowwith the defaults, then with three different values of--alpha. - Run
python param_flow.py run --helpand find your three parameters in the output. - Pass
--epochs abcand read what Metaflow says about the wrong type.
When things fail: resume, retry, catch and timeout
Real pipelines fail: a download times out, a file is missing, a cell of code has a typo. The question is how much work you lose. Metaflow's answer is that the work of every step that finished is saved, so a failure costs you only the failing step.
Here is a flow with a deliberate bug in its second step:
from metaflow import FlowSpec, step
class FragileFlow(FlowSpec):
@step
def start(self):
print("Doing the slow, expensive part")
self.data = list(range(1000))
self.next(self.fragile)
@step
def fragile(self):
self.ratio = len(self.data) / 0
self.next(self.end)
@step
def end(self):
print("ratio:", self.ratio)
if __name__ == "__main__":
FragileFlow()
Running it, the fragile step crashes with ZeroDivisionError, and Metaflow stops the run and prints the failure with the task's prefix. Now fix the bug (change / 0 to / 4) and run:
python fragile_flow.py resume
resume finds the most recent run, clones the results of the steps that succeeded (it does not re-run start), and re-executes from the step that failed. The slow, expensive part is skipped. By default it resumes the last run; to choose another, use --origin-run-id 3, and to restart from a step of your choosing, name it: python fragile_flow.py resume start.
Resume uses the original values of the parameters, not new ones. If you want to change a parameter, that is a new run, not a resume. The new code is used, though, which is exactly why resume is the tool for the edit-and-retry loop.
Some failures are not bugs but bad luck: a network blip, a rate limit. For those Metaflow has three decorators you place above @step:
from metaflow import FlowSpec, step, retry, catch, timeout
class RobustFlow(FlowSpec):
@retry(times=3, minutes_between_retries=1)
@timeout(seconds=30)
@step
def start(self):
import random
if random.random() < 0.5:
raise RuntimeError("flaky network")
self.value = 42
self.next(self.optional)
@catch(var="failure", print_exception=True)
@step
def optional(self):
raise ValueError("this part is nice to have")
self.next(self.end)
@step
def end(self):
print("value:", self.value)
print("optional step failed:", self.failure is not None)
if __name__ == "__main__":
RobustFlow()
@retry re-runs a failing task. Its defaults are three retries with two minutes between them, and the maximum you may ask for is times=4; asking for more gives "The maximum number of retries is @retry(times=4)". Each try is an attempt of the same task, and only the last attempt's artifacts and logs are kept. @timeout kills a step that runs too long; its seconds, minutes and hours arguments add up. @catch is different: it swallows the exception after all retries, stores it in the artifact whose name you give (failure), and lets the flow continue. Use it only when the flow can genuinely carry on without that step's output, and always check the artifact afterwards.
You can also add a decorator to every step at once from the command line, without touching the file, using --with, placed before run:
python robust_flow.py --with retry run
That is handy to make a whole flow tolerant of transient errors for one run.
- Run
FragileFlow, fix the bug, and useresume. Check that the output showsstartwas not executed again. - Run
RobustFlowseveral times and watch retries happen when the random check fails.
Logs, dump and cards: looking inside a run
While a flow runs, its print output streams to your terminal. Afterwards, Metaflow keeps it, so you can read the output of a past task:
python hello_flow.py logs 1/hello/2
The argument is the task's run/step/task identifier from the log prefix. Add --stderr for error output or --timestamps to see when each line was written. To see what a task saved, use dump:
python stats_flow.py dump 1/compute/2
This prints the artifacts of that task and their values. Use --include mean,biggest to show only some, and --file out.pkl to write them to a file.
For anything more readable than console text, Metaflow has cards: small HTML reports attached to a task. You add @card above a step, and Metaflow generates a default card showing the task's artifacts, parameters and metadata. You can add your own content through current.card:
from metaflow import FlowSpec, step, card, current
from metaflow.cards import Markdown, Table
class CardFlow(FlowSpec):
@step
def start(self):
self.scores = {"baseline": 0.81, "tuned": 0.87}
self.next(self.report)
@card(type="blank")
@step
def report(self):
current.card.append(Markdown("# Model comparison"))
rows = [[name, score] for name, score in self.scores.items()]
current.card.append(Table(rows, headers=["model", "accuracy"]))
self.next(self.end)
@step
def end(self):
print("report built")
if __name__ == "__main__":
CardFlow()
After a run, view the card in your browser:
python card_flow.py card view report
card view opens the newest card for that step. The blank type starts empty so you control everything, while the default type pre-fills artifacts for you. The components you will use most are Markdown for text, Table for rows, Image for pictures (including a matplotlib figure through Image.from_matplotlib) and Artifact for showing a stored value. Cards are the fastest way to share a result with a teammate who will not read your code. You can list a run's cards with card list, and card server starts a local viewer, by default on port 8324, that refreshes as the run progresses.
- Run
HelloFlow, then read the log of itshellotask withlogs, and dump the artifacts of itsendtask. - Build
CardFlowand open its card withcard view report. Add a second table of your own.
The Client API: reading past runs from Python
Everything Metaflow saved can be read back from ordinary Python, in a notebook or a script, without re-running anything. This is called the Client API, and it is where Metaflow's two halves meet: you author with FlowSpec, and you analyse with Flow, Run, Step and Task.
from metaflow import Flow, Run, Step, Task
flow = Flow("StatsFlow")
run = flow.latest_successful_run
print(run.id, run.successful, run.finished_at)
print(run.data.mean)
step = Step("StatsFlow/1/compute")
task = step.task
print(task.data.biggest)
print(task.stdout)
Flow("StatsFlow") is the collection of all runs of that flow. latest_run is the newest run, finished or not; latest_successful_run is the newest that completed. run.data exposes the artifacts of the end step of that run, so run.data.mean is the final value of mean. For artifacts from earlier steps, go through a Step or a Task and read .data there. Tasks also expose stdout, stderr, and exception if they failed.
You can iterate over history and filter it. A classic use is to find the best run of an experiment:
from metaflow import Flow
best = None
for run in Flow("ForeachFlow").runs():
if run.successful and (best is None or run.data.best_score > best.data.best_score):
best = run
print(best.id, best.data.best_lr, best.data.best_score)
Runs you tagged at the command line are easy to select with Flow("ParamFlow").runs("experiment-7").
One concept surprises nearly every beginner: namespaces. By default the Client API only shows runs that you created, filtered by a tag of the form user:<your username>. If you try to read someone else's run, or a run that was made by a scheduler, you get Object not in the current namespace. The fix is to widen the view:
from metaflow import namespace, Flow
namespace(None)
print(Flow("TeamFlow").latest_run)
namespace(None) selects the global namespace, showing every run. The documentation is clear that namespaces are a way to avoid confusion, not an access-control mechanism. They do not hide anything from anyone who can reach the datastore.
The other common error is Object not found with a message such as Flow('X') does not exist. On a local setup it almost always means you are in a different folder from the one where you ran the flow, because the local record lives in .metaflow in the working directory. Run metaflow status to see which metadata provider and working tree you are using, and change to the right folder.
- Open a Python shell in the folder where you ran
StatsFlowand printFlow("StatsFlow").latest_run.data.mean. - Write a loop over all runs of
ParamFlowthat prints each run's ID andloss. - Change folder to somewhere else and try again. Note the error, then return.
What happens during a local run
It helps to know what Metaflow does when you type python flow.py run, because it explains behaviour that otherwise looks like magic. The file is executed as a normal Python script. Metaflow reads your class, builds the graph from the self.next calls, and runs the linter. If the linter is happy, the local runtime starts the start step as a separate operating system process, waits for it to finish, reads the transitions, and launches the next step the same way. Anything that can run at the same time, such as the branches of a split or the tasks of a foreach, runs in parallel, up to the worker limit.
That design has four consequences that are worth saying plainly.
Each task starts with a clean interpreter. Module-level code at the top of your file runs once per task, not once per run. If you load a large file at the top of the script, every task loads it again, and if you print something there it will print many times. Put expensive work inside steps, where it runs only when needed.
Tasks communicate only through saved artifacts. When a task finishes, its artifacts are written to the datastore. When the next task starts, it reads the ones it needs from the datastore. Nothing is passed in memory. This is slower than a function call, and it is exactly what makes a step resumable and later movable to another computer: the next step needs nothing except what was saved.
The local runtime is for development. The documentation is clear that the local scheduler is not a production-grade scheduler. It does not restart itself after a reboot and it does not run at night on its own. That job belongs to a production orchestrator, which is the subject of the Senior guide. The important point for a beginner is that the flow file does not change when you move to one.
A crash in a step leaves the datastore consistent. A task either finishes and saves its artifacts or it does not. That is why resume can trust the steps that completed.
You can see the concurrency limit in action. By default up to 16 tasks run at once, and you can lower it when you want deterministic, sequential output or when you are hunting a bug:
python foreach_flow.py run --max-workers 1
With one worker, foreach tasks are processed one after another, so their log lines do not interleave. Another option that stops you from flooding a laptop with output is --max-log-size, which limits how many megabytes of a task's log are stored. The default is ten.
Good habits from the first day
A few habits are cheap to adopt at the beginning and painful to add later.
Keep each flow in its own small folder, with its own requirements file. Metaflow stores records in a .metaflow folder under the directory you run from. If you run flows from different folders, the Client API will only find the ones in the folder you are in, and Flow does not exist will confuse you. One folder per project removes the problem.
Name steps after what they do, not their order. Names like load_data, clean, train and evaluate make the show output and the log prefixes readable six months later. Step names can only contain lowercase letters, digits and underscores.
Make steps small enough to be worth resuming, and big enough to be worth saving. A step that loads a file and a step that trains a model are separate because you will want to re-run the second without the first. A step that adds one to a variable is not worth a step, because every step pays a small cost to save its artifacts.
Store results, not machinery. Save accuracy numbers, trained models, prediction tables. Do not save database connections, open files or thread pools. If you do, the save fails with a pickling error at the end of the step, and the message will not point at the line where you created the object.
Set random seeds as parameters. A run whose seed is a parameter can be reproduced from its record alone. A run whose seed is hidden in the code cannot, unless you also remember the version of the code.
Tag runs you care about. Tags cost nothing and turn "which run was that?" into a one-line query later.
Run check before run. It is instant and it catches the structural mistakes that otherwise cost you a run.
Commit your flow files to Git. Metaflow remembers the results of each run, but the record of which code produced them is your repository. A run that used code nobody committed cannot be reproduced.
--max-workers 1 and print self.input at the top of the step. Sequential output with clear labels finds most bugs faster than a debugger attached to many processes. If you do use an IDE debugger, configure it with the argument run and enable sub-process debugging, because each task is a child process.
- Add a
print("module loaded")at the top offoreach_flow.py, outside the class, and count how many times it prints in one run. Explain the number. - Rename the steps of one of your flows to describe what they do, and run
showto see the readable graph.
Dependencies: a short introduction
Your flow will soon need libraries beyond Metaflow, such as pandas or scikit-learn. On your laptop that is a simple matter: install them into your virtual environment and import them. Every step runs in a process started from that same environment, so they are all available.
Metaflow also has decorators, @pypi and @conda, that build a per-step environment with exact versions, which is essential when steps run on other machines. They only work if you launch with --environment=pypi (or conda), placed before run, and their packages must be imported inside the step. If you forget the flag, you get an "Incompatible environment" error that names the decorator and the flags it needs. The Mid-level guide covers this in depth. For now, know that the error means "tell Metaflow how to build environments", not "your code is wrong".
If you use uv, the documented pattern is uv run flow.py run for a flow inside a uv project. Keep a requirements.txt or lock file next to your flow; it is the easiest kind of reproducibility you can have, and it costs nothing.
Configuration and common errors
By default Metaflow stores everything locally. The three settings that matter are which datastore holds artifacts (local by default), which metadata provider records runs (local by default), and which environment builds dependencies (local by default). They live in a configuration file, config.json inside the folder ~/.metaflowconfig, and any setting can be overridden by an environment variable with the prefix METAFLOW_, such as METAFLOW_DEFAULT_DATASTORE. To use another profile, set METAFLOW_PROFILE=name, and Metaflow reads config_name.json instead.
You will rarely edit these by hand. If your organization has set up shared infrastructure, an administrator gives you a configuration, and the interactive metaflow configure commands (for example metaflow configure aws) write it for you; metaflow configure show prints the current values. As a learner you can ignore all of it, but knowing it exists explains two otherwise puzzling situations: flows that work on a colleague's machine because it points at the shared store, and flows that say Flow does not exist because your machine points somewhere else.
A note on where code lives. Metaflow packages the files of your working directory that end in .py, .R or .RDS when it needs to ship code to another machine. On a laptop that does not matter. It matters later, and it is one reason to keep each flow in a tidy folder of its own.
The errors you will actually meet, and how to read them, are below. Metaflow prints a short headline and then a message. Linter errors arrive under the headline "Validity checker found an issue", before anything runs.
Step *X* is missing a self.next() transition to the next step.A step does not end withself.next(...). Add it as the last line.Step *X* specifies a self.next() transition to an unknown step, *Y*.A typo in the step name. Compare it with the method definitions.Step *X* is unreachable from the entry step *start*.A step no one leads to. Wire it in or delete it.There is a loop in your flowA cycle in the transitions. Loops of this kind are not allowed; restructure, or use the recursion support covered in the Mid-level guide.The terminal step *end* was reached before a split started at step(s) ... were joined.You split with a branch or foreach and never joined. Add a join step, which takes aninputsargument.Step *X* is both a join step ... and a split step ...One step cannot join and also fan out. Make them two steps.Your flow must have exactly one start step.Name one stepstart, or mark exactly one as the entry step.Step *X* has an invalid name.Step names may contain only lowercase letters, digits and underscores.Foreach variable *self.x* in step *s* does not exist.The artifact named inforeach=was not set. Assign the list toselffirst.Step *join* cannot merge the following artifacts due to them having conflicting valuesSet those artifacts explicitly in the join, or useexclude.AttributeError: 'MyFlow' object has no attribute 'x'The variable was local in an earlier step. Save it toself.@catch is defined for the step *s* but @catch is not supported in foreach split steps.Move@catchto the inner step, not the one that fans out.ImportError: No module named fcntlYou are on native Windows. Use WSL 2.
check runs the same linter as run, so make it a reflex: if you change the shape of a flow, run python flow.py check first and save a failed run for real bugs.
- Run
metaflow configure showand read what it says about your current setup. - Break
BranchFlowon purpose in three different ways (remove a join, misspell a step, leave off aself.next) and runcheckeach time. Match the message to the list above.
Putting it all together
Here is one small project that uses almost everything above: train a few models in parallel with different settings, choose the best, report it on a card, and read the results afterward. It uses scikit-learn, so first pip install scikit-learn.
from metaflow import FlowSpec, Parameter, step, retry, card, current
from metaflow.cards import Markdown, Table
class TuneFlow(FlowSpec):
seed = Parameter("seed", default=42, help="Random seed")
depths = Parameter("depths", default="2,4,8", help="Tree depths, comma separated")
@step
def start(self):
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
self.split = train_test_split(X, y, test_size=0.3, random_state=self.seed)
self.depth_list = [int(d) for d in self.depths.split(",")]
self.next(self.train, foreach="depth_list")
@retry(times=2, minutes_between_retries=0)
@step
def train(self):
from sklearn.tree import DecisionTreeClassifier
X_train, X_test, y_train, y_test = self.split
self.depth = self.input
model = DecisionTreeClassifier(max_depth=self.depth, random_state=self.seed)
model.fit(X_train, y_train)
self.accuracy = float(model.score(X_test, y_test))
self.next(self.choose)
@step
def choose(self, inputs):
results = [(i.depth, i.accuracy) for i in inputs]
self.results = sorted(results)
self.best_depth, self.best_accuracy = max(results, key=lambda r: r[1])
self.next(self.report)
@card(type="blank")
@step
def report(self):
current.card.append(Markdown(f"# Best depth: {self.best_depth}"))
current.card.append(Table([[d, a] for d, a in self.results],
headers=["depth", "accuracy"]))
self.next(self.end)
@step
def end(self):
print(f"best depth {self.best_depth} with accuracy {self.best_accuracy:.3f}")
if __name__ == "__main__":
TuneFlow()
Notice the decisions in it. The data split is stored on self because several later tasks need it. Imports happen inside steps, a habit that pays off when environments get involved. The parameter depths is a string split into a list, because a comma-separated string is the simplest way to accept a list on the command line. Each train task saves depth and accuracy so the join can pair them. The join is named choose rather than join: any name works, since a join is defined by its extra argument, not by its name.
Run and inspect it:
python tune_flow.py check
python tune_flow.py run --depths 1,3,5,9 --tag first-try
python tune_flow.py card view report
Then, from Python:
from metaflow import Flow
run = Flow("TuneFlow").latest_successful_run
print(run.data.best_depth, run.data.best_accuracy)
print(run.data.results)
If a train task fails, fix the code and python tune_flow.py resume; the data loading will not repeat. If you want to try other depths, start a new run, because resume keeps the original parameters.
- Build
TuneFlowand run it with two different--depthslists and two different tags. - From Python, find the run with the higher best accuracy by looping over
Flow("TuneFlow").runs(). - Swap the model for another scikit-learn classifier and add its name as a parameter.
What you can now do, and what comes next
You can now take a script and structure it as a flow with start, end and steps between them, and you know that only values on self cross a step boundary. You can branch and join, fan out over a list with foreach and collect the results, and merge artifacts sensibly. You can pass parameters and tags at run time, recover from failures with resume, and make flaky steps tolerant with retry, catch and timeout. You can look back at logs, dump artifacts, build cards, and read any past run from Python with the Client API. You can read Metaflow's linter errors and fix the structure, not just the symptom.
| You want to | Reach for |
|---|---|
| Check the structure before running | python flow.py check and show |
| Keep a value for later steps | Assign it to self |
| Run things in parallel, fixed count | self.next(self.a, self.b) plus a join |
| Run one task per item in a list | self.next(self.step, foreach="items") |
| Change a value per run | Parameter, then run --name value |
| Continue after a failure | python flow.py resume |
| Survive flaky steps | @retry, @catch, @timeout |
| Read a past task's output | logs and dump |
| Share a result visually | @card and card view |
| Analyse runs in Python | Flow, Run, Step, Task |
Mid-level takes every one of those topics further: conditional transitions and recursion, running individual steps quickly with spin, per-step dependencies with @pypi and @conda, sending steps to remote compute on Kubernetes or AWS Batch with @resources, Config objects, secrets, the Runner and Deployer APIs, and debugging flows in an IDE.
Senior then covers running Metaflow as a platform for a team: the datastore and metadata service, deploying flows to a production orchestrator such as Argo Workflows, scheduling and event triggers, @project branches, security and access, cost, upgrades, and when another tool is the better choice.
If you are deciding what to learn alongside it, MLflow is the usual companion for tracking experiments and models, DVC covers versioning data files, and Docker and Kubernetes are what you will meet when your steps leave the laptop.