Skip to content
Back to student guides
PrefectMLOpsPipelines & orchestration3 levels108 sectionsCovers Prefect 3.8

The Complete Prefect Guide

Orchestrate Python data and ML workflows with Prefect flows, tasks and deployments. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
22sections
38examples

This is part one of three, written against Prefect 3.8 (tested on 3.8.7). It assumes you have never used Prefect, and it assumes you can read basic Python. By the end you will have written a workflow, watched it retry a flaky step, cached an expensive one, run it on a schedule, sent it to a worker, and read the dashboard to find out what went wrong when something failed. Mid-level and Senior take the same ideas into production; nothing you learn here is thrown away.

Each section ends with a Try it task. Do them in order. Orchestration is one of those subjects where reading is not enough: the concepts only make sense once you have watched your own flow fail, retry, and recover.

What Prefect is, and the problem it solves

Prefect is a Python library and a service for running your code reliably and seeing what it did. You write ordinary Python functions. You put a decorator on them. From then on, every call is recorded, can be retried when it fails, can be scheduled, and shows up in a web dashboard with its logs and its outcome.

That sounds modest, so it helps to see the problem it removes. Imagine a data team at a retailer in Riyadh. Every morning a script must download yesterday's orders from an API, clean them, load them into a warehouse, and refresh a model's features. The first version is a file called run_daily.py and a cron entry on someone's server. It works for a month. Then the orders API times out one night and the script dies halfway. Nobody knows until the dashboard is empty at nine in the morning. The fix is to log in, read a log file, and rerun the whole thing by hand, including the download that already succeeded. A week later the server is rebooted and cron is silently gone.

Every one of those failures has the same shape. The script did its job, but nothing around it did its job: nothing noticed the failure, nothing retried, nothing remembered which steps had finished, nothing told a human. Workflow orchestration is the name for the layer that does those things. Prefect is one of several orchestrators, and its particular design goal is to add that layer with as little change to your Python as possible.

YOUR PYTHONfunctions you already wrote
→
@flow / @tasktwo decorators
→
PREFECT APIstores state and history
→
DASHBOARDruns, logs, schedules

Read the diagram left to right. Your code stays your code. The decorators make it report to an API. The API keeps the record, and the dashboard is a window onto that record.

What people use it for:

🗓️

Scheduled data pipelines

Nightly loads, hourly syncs, weekly reports, each with retries and a visible history.

🤖

ML workflows

Train, evaluate and publish a model as separate steps, so a failed evaluation does not repeat the training.

🔁

Glue between systems

Call an API, transform the answer, write it somewhere, and alert someone if any step breaks.

👀

Observability for scripts

Even a script with no schedule gains logs, timings and a searchable record of every run.

Try it
  1. Think of one script you have written that must run repeatedly, even a homework one.
  2. Write down three things that could go wrong with it while you are asleep.
  3. Keep the list. By the end of this guide you will know which Prefect feature answers each item.
The features map to problems, not the other way round, which is the easiest way to remember them.

What came before, and where Prefect sits

The oldest answer is cron: a line in a file that starts a script at a chosen time. Cron is reliable at exactly one thing, starting a process. It does not know whether the process succeeded, it cannot retry, it cannot express "run B after A finished", and it keeps no history beyond whatever your script wrote to a file.

The next generation were workflow schedulers such as Apache Airflow (see the Airflow guide). Airflow asks you to describe your workflow as a graph, a DAG, in a file the scheduler parses. That graph is fixed in advance: the tasks and their dependencies are declared, and the scheduler runs them. This is a powerful model and it runs a huge share of the world's data pipelines, but it asks you to think in Airflow's vocabulary, and dynamic behaviour (loop over however many files arrived today) takes extra machinery.

Prefect's bet is different. You write normal Python, with normal loops, conditions and exceptions, and Prefect observes it while it runs instead of asking you to declare a graph first. If you want a step to repeat for each file, you write a for loop or call .map(). If you want to branch, you write an if. The graph of what happened is reconstructed from the run.

A fair comparison includes the other modern orchestrators. Dagster (see the Dagster guide) organises the world around the data assets your pipeline produces. Prefect organises it around the functions that run. Neither is universally better; the choice is usually about how your team prefers to think.

One more piece of history matters because it will save you hours. Prefect 3.0 became generally available in September 2024. Prefect 2.x is now legacy, yet a large amount of blog posts and AI-generated code still teaches the 2.x API: agents, Deployment.build_from_flow, infrastructure blocks. Those are gone in 3.x. This guide teaches only the 3.x way, and whenever an old name would trap you, it says so.

Old tutorials will mislead you If an example uses prefect agent start, prefect deployment build, prefect deployment apply or Deployment.build_from_flow, it targets Prefect 2. None of those exist in Prefect 3. Their replacements are prefect worker start, prefect deploy, and flow.serve() or flow.deploy(), all covered below.
Try it
  1. Search the web for "prefect tutorial" and open two results.
  2. Check the publication date and whether they mention agents. If they do, close the tab.
You are training a reflex you will use for every fast-moving tool.

The mental model: four nouns and a server

Prefect has a small vocabulary. Learn these words now and every later section is a variation on them.

A flow is a Python function decorated with @flow. It is the unit you run, schedule and deploy. A flow can call tasks, call other flows, and run any Python you like. Each time a flow is called, Prefect creates a flow run: a record with an ID, a name, parameters, a start time, logs and a final state.

A task is a Python function decorated with @task. It is a smaller unit of work with extra powers: it can be retried, cached, given a timeout, and run concurrently with its siblings. Each call creates a task run. Tasks are where you put the steps that can fail or are expensive.

A state describes where a run is in its life. The everyday ones are Scheduled, Pending, Running, Completed, Failed, Cancelled and Crashed. The two to learn early are the last two failures. Failed means your code raised an exception and retries (if any) are used up. Crashed means something outside your code killed the run: the machine ran out of memory, or a signal stopped the process. Keeping those two apart saves real debugging time, because one is a bug in your code and the other is a problem with the place it ran.

A deployment is a flow plus instructions for running it remotely: where its code lives, what parameters to use, what schedule to follow, and which infrastructure should execute it. You do not need a deployment to run a flow on your laptop. You need one when you want Prefect, rather than you, to start the run.

Behind these nouns sits one more thing: the Prefect API. It is a server, either Prefect Cloud (hosted by Prefect) or one you run yourself, that stores states, logs and schedules. The critical design point is the split between control and execution. The API stores metadata and decides what should run. Your code executes on machines you control, and only metadata goes to the API. In a region with data-residency rules, such as many Gulf and Egyptian employers have, that split matters: your data never has to leave your own infrastructure for Prefect to orchestrate it.

CONTROL PLANEPrefect Cloud or your own server: stores states, logs, schedules, deployments
EXECUTION PLANEyour laptop, worker, container or cluster: runs the Python and holds the data
Run it locally first You can run a flow with no server, no account and no deployment. Everything in the first half of this guide works that way. Prefect starts a temporary local server for you behind the scenes.
Try it
  1. Take the Riyadh retailer story above and label each step as a flow or a task: download, clean, load, refresh features.
  2. Decide which step you would most want to retry automatically, and why.
A good rule: anything that touches the network deserves to be a task.

Installing Prefect and checking the setup

Prefect 3.8 needs Python 3.10 to 3.14. Python 3.9 support was dropped in Prefect 3.5, so an old system Python is the most common reason an install fails. Check yours first.

BASH
python3 --version

Always use a virtual environment so Prefect's dependencies do not collide with other projects. On Linux and macOS:

BASH
python3 -m venv .venv && source .venv/bin/activate
pip install -U prefect

On Windows PowerShell the activation line differs:

POWERSHELL
py -m venv .venv
.venv\Scripts\Activate.ps1
pip install -U prefect

If you use uv, the equivalent is uv venv && source .venv/bin/activate && uv pip install prefect, or uv add prefect inside a project that has a pyproject.toml. Conda works too. On Windows the official documentation notes you may need to add Python's Scripts folder to your Path before the prefect command is found.

Check that the install worked:

BASH
prefect version

The output lists the Prefect version, the API version, the Python version, your operating system, the active profile, and the server type. A fresh install prints something like Server type: ephemeral, which means "no server is configured; Prefect will start a temporary one when needed". Do not worry about that word yet. It is explained in the next section.

There is also a much smaller package, prefect-client, meant for containers that only need to talk to an existing API and do not need the CLI or the server. As a beginner you want the full prefect package.

Pin the version in shared projects Write prefect==3.8.7 (or whichever version you tested) in requirements.txt. A newer client talking to an older server can produce confusing 422 errors, so teams keep client and server versions aligned on purpose.
Try it
  1. Create a folder prefect-lab, make a virtual environment inside it and install Prefect.
  2. Run prefect version and find the Python version and the Server type lines.
  3. Run prefect config view to see which profile is active.
If prefect is "command not found", your environment is not activated.

Where runs are recorded: ephemeral, local server, Cloud

Before writing a flow, settle where its record will go. There are three options and beginners are often confused because the first one is invisible.

Ephemeral mode is the default. With nothing configured, when you run a flow Prefect starts a temporary local API in the background, records the run, and shuts it down. It is convenient for learning because there is nothing to set up, but you cannot browse a dashboard, because the server disappears when your script ends. This is governed by the setting PREFECT_SERVER_EPHEMERAL_ENABLED, which the default ephemeral profile switches on.

A local server is a long-running API you start yourself. It gives you the dashboard and is the right choice for learning everything else in this guide.

BASH
prefect server start

Leave that terminal open. It serves the dashboard at http://127.0.0.1:4200 and the API at http://127.0.0.1:4200/api. The local server stores its data in a SQLite file under ~/.prefect by default. In a second terminal, with the same virtual environment active, tell your client where the API is:

BASH
prefect config set PREFECT_API_URL="http://127.0.0.1:4200/api"

Do not forget the /api ending. Leaving it off is one of the most common beginner mistakes, and it gives confusing errors.

Prefect Cloud is the hosted option. You run prefect cloud login, sign in through the browser, and choose a workspace. Runs are then recorded in Cloud and you use the Cloud dashboard. The free tier is enough to learn on. Code and data still stay on your machine; only metadata goes to Prefect.

BASH
prefect cloud login

The key that browser login creates expires after 30 days. For automation such as CI you would create a service account key and log in non-interactively. That belongs to later levels.

Local server

  • Free, private, no account
  • You start and stop it yourself
  • Good for learning and small teams

Cloud

  • Nothing to host or upgrade
  • Needs an account and outbound internet
  • Has plan limits on API rate and history

Either is fine for this guide. The rest of the text says "the dashboard" and means whichever you chose. Since Prefect 3.8, the self-hosted dashboard is the newer React-based interface by default, so screenshots in older blog posts may look different from what you see.

Try it
  1. In one terminal run prefect server start.
  2. Open http://127.0.0.1:4200 in your browser and find the empty Runs page.
  3. In a second terminal run the prefect config set command above, then prefect config view and confirm the URL.
Keep the server running for the rest of the guide.

Your first flow

Create a file named hello.py. The smallest useful Prefect program is one decorated function:

hello.py
from prefect import flow

@flow(log_prints=True)
def hello(name: str = "MENA"):
    print(f"Hello, {name}!")

if __name__ == "__main__":
    hello()

Run it the way you run any Python script:

BASH
python hello.py

You will see log lines in the terminal, roughly: the flow run was created with a random name such as amiable-falcon, the message printed, and the flow run finished in state Completed. Now open the dashboard and click Runs. The run is there, with its name, duration, and logs containing your greeting.

Three details deserve attention. First, log_prints=True tells Prefect to capture print() output as log entries, which is why the greeting appears in the dashboard and not only in your terminal. Without it, prints go to the terminal alone. Second, the flow is called like a normal function, including its parameter. There is no special runner. Third, the flow run got a generated name; you can control it with flow_run_name=, for example @flow(flow_run_name="greet-{name}").

The type hint on name: str is not decoration. Prefect validates the parameters of a flow against its hints, using Pydantic. Call hello(name=123) and Prefect will refuse the run with a validation error before any of your code executes. That early failure is a feature: bad input is caught at the door.

What happens if the function raises? Try it by changing the body to raise ValueError("boom"). The flow run ends in state Failed, the traceback appears in the logs, and the exception is re-raised to your script. A flow fails when an exception propagates out of it, which is the same rule you already know from Python.

Try it
  1. Create hello.py and run it twice with different names.
  2. Open the dashboard and find both runs.
  3. Make it raise an exception, run it again, and compare the Completed and Failed runs.
Click into the failed run and find the traceback in its logs.

Tasks: the steps inside a flow

A flow that only prints is not an orchestrator. The interesting part begins when you split the work into tasks. Here is a small pipeline that downloads data, cleans it and summarises it. It uses a public endpoint so you can run it yourself.

pipeline.py
import httpx
from prefect import flow, task

@task(retries=3, retry_delay_seconds=5)
def fetch_users() -> list[dict]:
    response = httpx.get("https://jsonplaceholder.typicode.com/users", timeout=10)
    response.raise_for_status()
    return response.json()

@task
def extract_cities(users: list[dict]) -> list[str]:
    return sorted({u["address"]["city"] for u in users})

@task
def summarise(cities: list[str]) -> str:
    return f"{len(cities)} distinct cities: " + ", ".join(cities)

@flow(log_prints=True)
def city_report():
    users = fetch_users()
    cities = extract_cities(users)
    print(summarise(cities))

if __name__ == "__main__":
    city_report()

Install httpx first with pip install httpx, then run the file. In the dashboard, open the run and look at the task runs: three of them, each with its own state, duration and logs. That is the payoff of decorating. Before, you had one opaque script. Now you can see that the fetch took two seconds and the summary took a millisecond.

The decorator option that matters most is retries. retries=3 means that if the task raises, Prefect runs it again, up to three more times. retry_delay_seconds=5 waits five seconds between attempts. Network calls fail transiently all the time, and this one line converts a three-in-the-morning page into a line in a log. You can also pass a list of delays, such as retry_delay_seconds=[1, 10, 60], for growing waits, or use exponential_backoff from prefect.tasks.

Tasks called inside a flow look like normal function calls, and they are: calling fetch_users() runs it right then and returns its result. Data passes between tasks as ordinary return values and arguments. There is no separate wiring language.

A few other task options appear early: timeout_seconds=30 fails a task that runs too long, name="fetch" and tags=["api"] label it in the dashboard, and log_prints=True works on tasks as well. Tasks are also callable outside a flow in Prefect 3, but as a beginner keep them inside flows.

How do you decide what is a task? Use tasks for units that can fail independently or are worth retrying, caching or timing on their own: network calls, database queries, slow computations, file writes. Do not turn every two-line helper into a task. Each task run is recorded with states and events, and thousands of tiny tasks add overhead without adding insight. A plain function is fine for trivial steps.

Retries re-run the whole task When a task is retried, the entire function runs again from its first line. A task that sends an email and then fails on the next line will send the email again on retry. Keep tasks small, and make side effects safe to repeat.
Try it
  1. Run pipeline.py and open the run in the dashboard.
  2. Change the URL to one that does not exist and run again.
  3. Watch the task go through retries, and read the retry entries in the logs.
After the last retry the task fails, and because the exception propagates, so does the flow.

Logging and reading states

When something goes wrong you will reach for logs, so set them up properly from the start. There are two ways to write logs that Prefect captures. The quick one is log_prints=True on the decorator, as you have used. The better one for real work is the run logger:

PYTHON
from prefect import flow, task, get_run_logger

@task
def load(rows: list[dict]) -> int:
    logger = get_run_logger()
    logger.info("Loading %d rows", len(rows))
    if not rows:
        logger.warning("Nothing to load")
    return len(rows)

@flow
def nightly():
    load([{"id": 1}, {"id": 2}])

get_run_logger() returns a standard Python logger that is attached to the current flow run or task run, so each message appears in the dashboard against the right run and carries a level: debug, info, warning or error. It only works inside a flow or task. Calling it from plain module code raises MissingContextError; outside of runs, use from prefect.logging import get_logger.

Logs are sent to the API unless you disable that. The default level is INFO, so debug lines are hidden until you change PREFECT_LOGGING_LEVEL.

Now that you have seen states in the dashboard, here is how they fit together. A scheduled run begins as Scheduled. When a process picks it up it becomes Pending, then Running. If a retry is pending it shows AwaitingRetry and then Retrying. It ends in one of the terminal states: Completed, Failed, Cancelled or Crashed. Two states worth recognising: Late marks a run that was supposed to start and did not, typically because nothing was available to run it, and Cached marks a task that did not execute because a saved result was reused. The next section explains where that comes from.

You can also ask for a state rather than an exception. Calling a task with return_state=True returns the State object instead of raising, which lets a flow decide what to do with a failure:

PYTHON
@flow
def tolerant():
    state = risky_task(return_state=True)
    if state.is_failed():
        print("risky_task failed, continuing with a fallback")
Try it
  1. Add get_run_logger() to one task from the last section and write an info line and a warning line.
  2. Find both lines in the dashboard and note how the level is shown.
  3. Write an on_failure hook that prints the final state name, attach it to a flow that raises, and run it.
The hook receives the flow, the flow run and the state, in that order.

Running tasks concurrently with submit and map

So far tasks ran one after another. Calling fetch_users() blocks until it finishes. When you have many independent items, such as a hundred URLs, waiting for each in turn wastes time. Prefect tasks can run concurrently.

The two tools are .submit() and .map(). Both return futures, which are placeholders for results that may not be ready yet.

concurrent.py
from prefect import flow, task

@task
def process_customer(customer_id: str) -> str:
    return f"Processed {customer_id}"

@flow(log_prints=True)
def many():
    ids = [f"customer{n}" for n in range(10)]
    futures = process_customer.map(ids)
    results = [f.result() for f in futures]
    print(results[:3])

if __name__ == "__main__":
    many()

.map(ids) launches one task run per item and returns a list of futures. Calling .result() on a future waits for that task and gives you its return value. You can also pass a future straight into another task, and Prefect waits for it automatically before running the downstream task. That is how you build a pipeline where step two starts as soon as its own input is ready.

.submit(x) does the same for a single call. Use it when you want to start a task and keep going, then collect the result later:

PYTHON
future = process_customer.submit("customer1")
other_work = do_something_else()
print(future.result())

What runs these tasks side by side? A task runner. The default is ThreadPoolTaskRunner, which uses threads, good for work that mostly waits on the network. For CPU-heavy Python there is ProcessPoolTaskRunner, and dedicated runners for Dask and Ray live in separate packages. As a beginner, the default is right.

Two small rules prevent the usual stumbles. If a task takes several arguments and only one varies per item, wrap the constant ones in unmapped() from prefect, as in process.map(ids, config=unmapped(cfg)). Without it, .map() will try to iterate the constant too and complain with MappingLengthMismatch or MappingMissingIterable. And in Prefect 3, futures are synchronous: do not await a future from .submit().

Flow outcome and task outcome In Prefect 3, a flow fails if an exception propagates out of it. A task that fails inside .submit() or .map() only affects the flow when you read its result (for instance with .result()), because that is what re-raises the error.
Try it
  1. Run concurrent.py and open the run; count the ten task runs.
  2. Add time.sleep(2) inside the task and compare the total duration with and without .map().
  3. Change the task to raise for one specific customer and see what the flow reports.
Ten two-second tasks finishing in about two seconds, not twenty, is the point of concurrency.

Results, persistence and caching

A result is what a task or flow returns. By default Prefect keeps results in memory for the duration of the run, so passing data between tasks works, but it does not save them anywhere after the run is over. That default is deliberate: results can be large or sensitive, and Prefect does not assume you want them stored.

Saving results is called persistence. You turn it on per task or flow with persist_result=True, or globally with PREFECT_RESULTS_PERSIST_BY_DEFAULT=true. Persisted results are written to local storage, which defaults to ~/.prefect/storage in current versions, or to a storage location you choose with result_storage. Teams running work on several machines need shared storage, such as a cloud bucket, because a result saved on one machine's disk cannot be read by another. If you ever see MissingResult, it means you asked for a result that was never persisted.

Caching builds on persistence. If a task has already run with the same inputs, you may not want to run it again. Prefect 3 makes this easy, and gives a surprise to beginners: tasks use a default cache policy made of the task's inputs, its source code, and the flow run ID. Within one flow run, calling the same task with the same arguments reuses the saved result. Across different runs, the run ID changes, so nothing is reused unless you say so.

To reuse results across runs, choose a policy that leaves the run ID out:

cache_demo.py
from datetime import timedelta
from prefect import flow, task
from prefect.cache_policies import INPUTS, TASK_SOURCE

@task(cache_policy=INPUTS + TASK_SOURCE, cache_expiration=timedelta(hours=1))
def expensive_lookup(country: str) -> dict:
    print(f"computing for {country}")
    return {"country": country, "score": len(country)}

@flow(log_prints=True)
def report():
    print(expensive_lookup("Egypt"))
    print(expensive_lookup("Egypt"))

if __name__ == "__main__":
    report()

Run it twice. The first run prints "computing for Egypt" once and shows the second call as Cached. The second run, within the hour, prints no computing line at all, because the saved result is reused. The key is built from the inputs plus the task's own source code, so if you edit the function, the cache is invalidated on purpose. Setting a cache policy turns persistence on for that task automatically.

The policies you meet first are INPUTS, TASK_SOURCE, FLOW_PARAMETERS and NO_CACHE, and you combine them with +. To switch caching off for a task, write cache_policy=NO_CACHE. To force a fresh run while keeping the policy, pass refresh_cache=True.

Why is my task not re-running? If a task shows Cached and you expected it to run, the default policy or one you set found a matching saved result. Set cache_policy=NO_CACHE or refresh_cache=True. If instead you see a HashError, an input could not be hashed, for example a database connection; exclude it with something like INPUTS - "client".
Try it
  1. Run cache_demo.py twice and compare which runs print the computing line.
  2. Change the country in one call to "Saudi Arabia" and see a fresh computation.
  3. Edit the function body slightly and run again; explain why the cache missed.
The cache key includes the source code, so any edit to the task is a miss.

Parameters and subflows

A flow is a function, so its arguments are its parameters. They are how you reuse one flow for many cases: a date, a country, a dry-run switch. Give them type hints and defaults, and Prefect validates them.

params.py
from datetime import date
from pydantic import BaseModel
from prefect import flow

class ReportConfig(BaseModel):
    country: str = "Egypt"
    top_n: int = 5

@flow(log_prints=True)
def daily_report(day: date, config: ReportConfig = ReportConfig()):
    print(f"Report for {day} in {config.country}, top {config.top_n}")

if __name__ == "__main__":
    daily_report(day=date(2026, 9, 30), config=ReportConfig(country="Saudi Arabia"))

Pydantic models as parameters are a good habit, because the deployment UI later renders them as a form and rejects malformed input up front. If the parameters fail validation, the run goes from Pending straight to Failed without ever reaching Running, and you get ParameterTypeError.

A flow can also call other flows. The inner flow is a subflow: it gets its own flow run, linked to the parent in the dashboard. Use subflows when a section of your pipeline deserves its own record, parameters and final state, for instance one subflow per country.

PYTHON
@flow
def load_country(country: str):
    ...

@flow
def load_all():
    for country in ["Egypt", "Saudi Arabia", "UAE"]:
        load_country(country)

The parent waits for each subflow to finish. Prefer tasks for small steps and subflows for meaningful units; a rule of thumb is that if you would want to rerun it by itself from the dashboard, it may be a flow.

You can also set a time limit and retries at the flow level. @flow(retries=2, retry_delay_seconds=30, timeout_seconds=600) reruns the whole flow on failure and fails it if it runs past ten minutes. Flow retries rerun the flow function from the start, so combine them with caching, so the tasks that already succeeded are not repeated.

Try it
  1. Write a flow with a date parameter and call it with a string that is not a date.
  2. Read the error, then look for the run in the dashboard and note its state.
  3. Add a subflow and find its parent link in the run page.
A run that fails validation never reaches Running, which tells you the failure was in the input.

Serving a flow: your first scheduled run

Everything so far ran when you typed python file.py. To have Prefect start runs on its own, you create a deployment. The quickest way is .serve().

serve_demo.py
from prefect import flow

@flow(log_prints=True)
def morning_report(country: str = "Egypt"):
    print(f"Good morning from {country}")

if __name__ == "__main__":
    morning_report.serve(
        name="morning-report",
        cron="0 8 * * *",
        parameters={"country": "Egypt"},
    )

Start it with python serve_demo.py while your local server is running. Unlike before, the process does not exit. It stays alive, registers the deployment with the API, and polls for runs. The dashboard now has a Deployments page listing morning-report/morning-report, with a next run scheduled for 08:00.

You do not have to wait until morning. Open a second terminal and start a run by hand:

BASH
prefect deployment run "morning-report/morning-report" --param country="Saudi Arabia"

The name format is always flow-name/deployment-name. You will see the serving process pick the run up and execute it, and the run appears in the dashboard with the parameter you passed. You can do the same from the dashboard's Run button, which offers a form for the parameters.

The crucial limitation to understand: a served flow only runs while the serving process is alive. Close the terminal, and the deployment still exists but nothing is listening, so scheduled runs become Late. .serve() is perfect for learning, for a single machine, and for containers you keep running with a process manager. For anything that must survive a reboot or scale across machines, the next sections introduce work pools and workers.

Keep the serving process alive On a real machine run it under a supervisor (systemd, a container with a restart policy, or a process manager) so it restarts after failure. That supervisor is the piece that replaces "someone remembers to start it".
Try it
  1. Run serve_demo.py and keep it running.
  2. In a second terminal, trigger a run with prefect deployment run.
  3. Stop the serving process with Ctrl+C, trigger another run, and watch it sit in a waiting state until you start serving again.
This is the clearest demonstration of why workers exist.

Schedules

A schedule tells Prefect when to create runs for a deployment. Prefect supports three kinds. Cron uses the familiar five-field expression. Interval repeats every N seconds. RRule follows the iCalendar recurrence format and suits irregular patterns such as "first Monday of each month".

With .serve() the shortcuts are cron=, interval= and rrule=. A few examples:

PYTHON
morning_report.serve(name="every-morning", cron="0 8 * * *")
morning_report.serve(name="every-ten-min", interval=600)

For more control, including a timezone, use the schedules argument with the classes from prefect.schedules:

PYTHON
from prefect.schedules import Cron

morning_report.serve(
    name="riyadh-morning",
    schedules=[Cron("0 9 * * MON-FRI", timezone="Asia/Riyadh")],
)

The timezone matters more than beginners expect. Without one, a cron expression is interpreted in UTC, so "8 a.m." would land at 11 a.m. in Riyadh. State the timezone you mean. Prefect keeps cron and RRule schedules at the same local wall-clock time across daylight-saving changes, which is irrelevant in Riyadh or Cairo for most of the year but useful to know if your team spans regions.

A deployment can hold several schedules, each with its own parameters. The scheduler service generates upcoming runs ahead of time, so the dashboard shows future runs as Scheduled. By default it plans up to 100 runs, at most 100 days ahead, for each deployment. You manage schedules from the dashboard, from your code, or from the CLI:

BASH
prefect deployment schedule ls "morning-report/morning-report"
prefect deployment schedule pause "morning-report/morning-report" --all
Late runs A run is marked Late if it is still Scheduled 15 seconds after its start time. Late usually means nothing was available to pick it up. It is a symptom, and the cause is almost always a missing or paused worker or serving process.
Try it
  1. Serve a flow with an interval of 60 seconds and watch three runs appear.
  2. Change it to a Cron schedule with the timezone Asia/Riyadh or your own.
  3. List the schedule from the CLI and pause it, then confirm in the dashboard that no new runs appear.
Pausing a schedule stops future runs; it does not cancel ones already running.

Work pools and workers

.serve() ties execution to one process you started. Real teams want execution to be decoupled from deployment: developers register what should run, and separate infrastructure decides where and how. Prefect does this with work pools and workers.

A work pool is a named bridge between deployments and infrastructure. It has a type (process, docker, kubernetes, and several cloud types) and a template describing how to start a run of that type. A deployment targets a pool. A worker is a small long-running process, started with the CLI, that polls one pool for runs and, for each, creates the infrastructure: a subprocess for a process pool, a container for a docker pool, a Job for Kubernetes. The worker's type must match the pool's type. Workers only make outbound connections to the API, so the machine running them does not need to accept inbound traffic.

DEPLOYMENTflow + schedule + pool
→
WORK POOLtyped queue of runs
→
WORKERpolls and launches
→
INFRASTRUCTUREprocess, container, Job

The simplest pool type is process: runs execute as plain subprocesses on the machine where the worker runs. Create one and start a worker:

BASH
prefect work-pool create my-pool --type process
prefect worker start --pool my-pool

Leave the worker running. Now deploy a flow to that pool. Instead of .serve(), use .from_source(...).deploy(...), which says where the code lives and which pool should run it:

deploy_demo.py
from prefect import flow

if __name__ == "__main__":
    flow.from_source(
        source="https://github.com/your-org/your-repo.git",
        entrypoint="flows/report.py:morning_report",
    ).deploy(
        name="morning-report-pool",
        work_pool_name="my-pool",
        cron="0 8 * * *",
    )

Replace the repository and entrypoint with your own. The entrypoint has the form path/to/file.py:function_name. The worker, when it gets a run, clones the repository, loads the flow from that file and runs it. That is the big difference from serving: the worker fetches code at run time, so pushing new code to the repository changes what tomorrow's run executes, with no restart.

The docker and kubernetes pool types are the route to containers and clusters; they need pip install "prefect[docker]" or "prefect[kubernetes]" and a little more configuration. See the Docker guide and Kubernetes guide for the underlying tools. A worker will offer to install the matching integration package when it is missing, and you can set the behaviour with --install-policy.

Check on things with:

BASH
prefect work-pool ls
prefect work-pool inspect my-pool

Workers send a heartbeat every 30 seconds and are shown offline after missing a few, so the dashboard's Work Pools page tells you at a glance whether anyone is listening. This is the first place to look when runs go Late.

Dependencies are your job now Since Prefect 3.7.3, code fetched by a worker no longer installs its requirements automatically. Either add a pip_install_requirements pull step (next section), bake the dependencies into an image, or set PREFECT_RUNNER_AUTO_INSTALL_DEPENDENCIES=true on the worker. A flow that imports a package the worker lacks fails at startup.
Try it
  1. Create a process work pool and start a worker for it.
  2. Push a tiny flow to a Git repository you own and deploy it to the pool with from_source.
  3. Trigger it with prefect deployment run and watch the worker log the clone and the run.
If the run stays Scheduled, check that the worker's terminal is alive and the pool is not paused.

Deploying from a prefect.yaml file

Writing deployments as Python is fine for one flow. Teams with several usually prefer a declarative file committed beside the code: prefect.yaml. You create a starter with:

BASH
prefect init

and then deploy with the prefect deploy command, which reads the file. A minimal version:

prefect.yaml
name: mena-pipelines
prefect-version: 3.8.7

pull:
  - prefect.deployments.steps.git_clone:
      repository: https://github.com/your-org/your-repo.git
      branch: main
  - prefect.deployments.steps.pip_install_requirements:
      requirements_file: requirements.txt
      directory: "{{ git_clone.directory }}"

deployments:
  - name: morning-report
    entrypoint: flows/report.py:morning_report
    parameters:
      country: Egypt
    schedules:
      - cron: "0 8 * * *"
        timezone: Asia/Riyadh
    work_pool:
      name: my-pool

Read the file as three ideas. The pull section lists steps a worker runs to get your code onto the machine before the flow starts: clone a repository, then install its requirements. The deployments list names each deployment, its entrypoint, parameters, schedules and work pool. Double-curly templating such as {{ git_clone.directory }} refers to the output of an earlier step, and {{ $ENV_VAR }} reads an environment variable at deploy time.

Then deploy:

BASH
prefect deploy --name morning-report
prefect deploy --all

Without flags, prefect deploy walks you through an interactive wizard that asks for the flow, the pool and a schedule and offers to save the answers to prefect.yaml. That is a gentle way to produce the first file. Keep the file in version control so a deployment is reviewable like any other change.

Re-deploy after changing deployment settings Edits to a flow's code are picked up at run time because the worker pulls fresh code. Edits to schedules, parameters or the pool live in the deployment, so run prefect deploy again after changing them.
Try it
  1. Run prefect init in your repository and read the generated file.
  2. Fill in a deployment for one of your flows and run prefect deploy --name for it.
  3. Trigger it from the dashboard and confirm the worker ran it.
Change a parameter in prefect.yaml and redeploy to see the difference in the dashboard.

Variables, secrets and blocks

Real flows need configuration: a bucket name, an API endpoint, and, worse, credentials. Do not hard-code them, and do not commit them. Prefect gives you three places to keep them outside your code.

Variables are named values stored on the server, meant for non-sensitive configuration such as an environment name. They are not encrypted. Set and read them from the CLI or Python:

BASH
prefect variable set env_name prod
prefect variable get env_name
PYTHON
from prefect.variables import Variable

Variable.set("env_name", "prod", overwrite=True)
env = Variable.get("env_name", default="dev")

Secrets hold sensitive strings. The Secret block stores its value encrypted on the server and hides it in the UI:

PYTHON
from prefect.blocks.system import Secret

Secret(value="s3cr3t").save("db-password", overwrite=True)
password = Secret.load("db-password").get()

The first line is usually run once, by a person, from a trusted machine. The second is what the flow runs. The password lives on the server, and the flow code only contains its name.

Blocks are the general mechanism: a typed, saved configuration object, such as credentials for a cloud bucket, that you load by name. Integration packages supply block types; for instance prefect-aws provides an S3Bucket block. You register a block type and create documents for it, then in code call .load("name"). Blocks are how a flow gets access to a storage service without knowing the key.

A Variable is not a secret Variables are stored unencrypted and visible to anyone who can open the workspace. Never put a password or token in one. Use a Secret block or your platform's secret store.
Try it
  1. Create a Variable called greeting and read it inside a flow.
  2. Save a Secret block and load it in a flow, printing only its length, never its value.
  3. Find both in the dashboard and compare how each is displayed.
Printing a secret to logs defeats its purpose; log that it loaded, not what it contains.

The everyday CLI, grouped by intent

Most of what you will do after setup happens in a handful of commands. Grouped by what you are trying to accomplish:

Check where you are.

BASH
prefect version
prefect config view
prefect profile ls

prefect config view shows the active profile's settings and --show-defaults adds the ones you have not changed. A profile is a named bundle of settings stored in ~/.prefect/profiles.toml. The defaults are ephemeral, local, test and cloud. Create one per backend and switch with prefect profile use local.

Find and inspect runs.

BASH
prefect flow-run ls
prefect flow-run ls --state Failed
prefect flow-run inspect <ID>
prefect flow-run logs <ID>

Act on runs.

BASH
prefect flow-run cancel <ID>
prefect flow-run retry <ID_OR_NAME>
prefect flow-run watch <ID>

Work with deployments.

BASH
prefect deployment ls
prefect deployment inspect "flow/deployment"
prefect deployment run "flow/deployment" --param key=value --watch

The --watch flag makes the command wait and stream the run's progress, which is handy in scripts. Since Prefect 3.6.21, many commands also accept --output json, useful when you need to feed the answer to another program.

Manage infrastructure.

BASH
prefect work-pool ls
prefect worker start --pool my-pool
Try it
  1. Run a flow that fails, then list failed runs with prefect flow-run ls --state Failed.
  2. Use prefect flow-run logs with the ID to read the traceback without opening the browser.
  3. Run prefect --help and scan the command groups so you know what exists.
Being able to debug from the terminal is what you will need over SSH on a server.

Configuration, profiles and where settings come from

Prefect is configured by settings. Each has a name like PREFECT_API_URL and a default. You can change them in several places, and when the same setting appears in more than one place, a fixed order decides the winner.

From strongest to weakest: environment variables, a .env file, a prefect.toml file in your project, the [tool.prefect] table in pyproject.toml, the active profile, and finally the built-in default. The practical rule is that an environment variable always wins, which is what makes containers and CI easy to configure.

The command prefect config set writes to the active profile, and prefect config view --show-sources tells you where each value came from, which is the fastest way to answer "why is it using that URL?".

prefect.toml
[api]
url = "http://127.0.0.1:4200/api"

[logging]
level = "DEBUG"

The settings you will meet first:

Setting What it does Default
PREFECT_API_URL Which API your client talks to none (ephemeral mode if allowed)
PREFECT_LOGGING_LEVEL How chatty the logs are INFO
PREFECT_RESULTS_PERSIST_BY_DEFAULT Save results after runs false
PREFECT_RESULTS_LOCAL_STORAGE_PATH Where local results are saved ~/.prefect/storage
PREFECT_SERVER_UI_V2_ENABLED New UI for self-hosted servers true since 3.8.0
PREFECT_HOME Where Prefect keeps its files ~/.prefect
Try it
  1. Run prefect config view --show-sources and find the source of PREFECT_API_URL.
  2. Override it for one command with an environment variable and confirm the change.
  3. Create a profile called cloud-test and switch between it and local.
Environment variables override everything, so check them first when a setting will not stick.

Reading the common errors

Most early problems fall into a few families. Here is how to read them.

"No Prefect API URL provided. Please set PREFECT_API_URL to the address of a running Prefect server." Your client has no backend and ephemeral mode is off, typically because you are in a container or a custom profile. Run prefect config set PREFECT_API_URL=http://host:4200/api or prefect cloud login, and remember the /api suffix.

"Found incompatible versions: client: X, server: Y. Major versions must match." A Prefect 2 client is talking to a Prefect 3 server, or the other way round. Align the major versions. A client newer than its server can also earn 422 responses, so keep the client at or below the server's version.

401 "Invalid authentication credentials" on Cloud. The API key is missing, wrong or expired. Browser-login keys last 30 days, so a laptop that worked last month can stop working. Log in again.

Runs stuck in Scheduled, then Late. Nothing is picking them up. Check that a worker is running for that pool, that its type matches the pool, that the pool and queue are not paused, and that no concurrency limit is full. Worker logs state plainly when a pool is paused. Work through prefect work-pool inspect and the Work Pools page.

"Unable to start worker. Please ensure you have the necessary dependencies installed..." The worker type needs an integration package. Install it, for example pip install prefect-kubernetes, or use --install-policy always.

Crashed. Not your code's exception. Suspect memory, a killed container or a lost machine. Read the worker logs and the container or pod logs, and raise memory limits if needed.

Script errors at startup such as MissingFlowError or "Script at ... encountered an exception". The entrypoint is wrong (check path/file.py:function spelling), the code was not pulled, or a dependency is missing on the worker. Since 3.7.3 this last one is the most common cause.

ParameterTypeError. The parameters do not match the flow's type hints. Fix the payload. If the deployment's saved parameters no longer match the flow's signature, you will see SignatureMismatchError; redeploy.

A reliable general method: open the failed run, read the last error in its logs, note which state it ended in, and ask whether the cause is your code (Failed), the infrastructure (Crashed), or the plumbing (Late, never started).

Try it
  1. Deliberately break one thing at a time: a wrong entrypoint, a stopped worker, a bad parameter.
  2. For each, write the exact message and the state the run ended in.
  3. Keep the list as your personal troubleshooting sheet.
Having seen each error once on purpose makes the real one far less frightening.

Putting it all together

Here is a small end-to-end project that uses most of what you have learned: fetch data with retries, cache an expensive step, run tasks concurrently, log properly, and schedule it. It is a nightly report that fetches posts for several users and counts words.

nightly_posts.py
import httpx
from datetime import timedelta
from prefect import flow, task, get_run_logger
from prefect.cache_policies import INPUTS, TASK_SOURCE

@task(retries=3, retry_delay_seconds=[2, 10, 30], timeout_seconds=30)
def fetch_posts(user_id: int) -> list[dict]:
    logger = get_run_logger()
    url = f"https://jsonplaceholder.typicode.com/posts?userId={user_id}"
    response = httpx.get(url, timeout=10)
    response.raise_for_status()
    logger.info("user %s returned %d posts", user_id, len(response.json()))
    return response.json()

@task(cache_policy=INPUTS + TASK_SOURCE, cache_expiration=timedelta(hours=6))
def count_words(posts: list[dict]) -> int:
    return sum(len(p["body"].split()) for p in posts)

@flow(log_prints=True, retries=1, retry_delay_seconds=30)
def nightly_posts(user_ids: list[int] = [1, 2, 3, 4, 5]):
    fetches = fetch_posts.map(user_ids)
    counts = count_words.map(fetches)
    totals = dict(zip(user_ids, [c.result() for c in counts]))
    print(f"Word counts by user: {totals}")
    return totals

if __name__ == "__main__":
    nightly_posts.serve(
        name="nightly-posts",
        cron="0 2 * * *",
        parameters={"user_ids": [1, 2, 3, 4, 5]},
    )

Walk through it, because each line is a lesson from earlier. fetch_posts is a task because it touches the network, so it has retries with growing delays and a timeout. count_words is cached on its inputs and source, so rerunning the flow within six hours skips recomputation. .map() runs the five fetches concurrently, and because count_words.map(fetches) receives futures, each count starts as soon as its own fetch finishes, with no manual waiting. The flow has its own retry for a bad night. The flow is served on a cron schedule for 02:00, and you can also trigger it by hand with prefect deployment run "nightly-posts/nightly-posts".

To make it survive reboots, you would move from .serve() to a work pool and a worker, a prefect.yaml in the repository, and a pull step that installs httpx. That is the same code with a different deployment path, and it is the natural bridge to the Mid-level guide.

Run it, let it complete once, and open the dashboard. You should be able to explain, for this one flow, every row: the flow run, five fetch task runs, five count task runs, and their states, durations and logs.

Try it
  1. Run the project and trigger it by hand twice in a row.
  2. Identify which task runs were Cached on the second run and explain why.
  3. Make one user ID invalid so that a fetch fails; observe the retries, then the flow's final state.
Being able to narrate a run from the dashboard is the skill interviewers and colleagues notice first.

What you can now do, and what comes next

You can now install Prefect and check the setup, choose between ephemeral mode, a local server and Cloud, turn Python functions into flows and tasks, and watch them in the dashboard. You can retry flaky steps, time out slow ones, cache expensive results, run tasks concurrently with .submit() and .map(), and pass validated parameters. You can serve a flow on a schedule, deploy one to a work pool with a worker, describe deployments in prefect.yaml, store configuration as Variables and Secrets, and read the common error messages well enough to tell a code bug from an infrastructure failure from a missing worker.

What comes next, in the Mid-level guide: Docker and Kubernetes work pools, base job templates and job variables, automations and events, concurrency limits, transactions, artifacts and testing your flows. Senior covers running Prefect as a platform: high-availability servers, security, multi-team setups and cost.

For neighbouring tools in this catalogue, the Airflow guide and Dagster guide show the same problem solved with different philosophies, the MLflow guide pairs naturally with Prefect for tracking the models your flows train, and the Docker guide is the prerequisite for container-based pools.

Sources