This is part one of three. It covers everything you need to put a machine learning model behind a web API with FastAPI, written for someone who has never built an API before. By the end you can install FastAPI, write endpoints that accept and validate data, load a model once when the server starts, return predictions as JSON, read the error messages the framework gives you, test the whole thing without starting a server, and package it for a colleague. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and an API only becomes real once you have sent it a bad request and watched it answer politely.
The guide was checked against FastAPI 0.142.2, released on 30 September 2026. FastAPI is still a 0.x library, so minor versions can change behaviour. We will point out the places where that matters.
What FastAPI is, and the problem it solves
A trained model is a file and a function. It sits in a notebook, and the only person who can use it is the one who owns the notebook. To make it useful to a mobile app, a website, a colleague's script, or another service, you need a way for other programs to send it inputs and receive outputs over a network. The standard way to do that is an HTTP API: a program that listens on a port, receives requests such as "predict the price for this house", runs your code, and sends back a response, usually in a text format called JSON.
FastAPI is a Python framework for building exactly that. You write ordinary Python functions, add type hints, and FastAPI turns them into web endpoints. It does three jobs for you that you would otherwise write by hand:
- It reads the request and converts the raw text into Python values.
- It validates those values against the types you declared and rejects bad input with a clear error before your code ever runs.
- It writes documentation for your API automatically, including a page where you can try each endpoint from your browser.
The diagram is the whole idea. The middle box is where FastAPI earns its keep, and the rest of this guide fills in each arrow.
It helps to know what came before. Python has had web frameworks for a long time. Flask is small and flexible, but it leaves validation and documentation to you, so every project grows its own handwritten checks. Django is large and includes a database layer and an admin interface, which is more than a model-serving service needs. FastAPI sits in between: small like Flask, but with validation and documentation built in because it reads your type hints.
Under the hood FastAPI is two libraries glued together by a thin layer of its own. Starlette handles the web part: routing, requests, responses, middleware. Pydantic handles the data part: checking that values have the right type and shape. FastAPI runs on ASGI, a standard interface between Python web applications and the servers that run them. ASGI is what lets one server process handle many waiting requests at once, which matters later when we discuss async.
Why does this matter for machine learning in particular? Because a model is only valuable once something else can call it, and because the failures in model serving are mostly boring input failures: a missing field, a string where a number was expected, a list of the wrong length. A framework that rejects those at the door, with an error that tells the caller exactly what was wrong, saves you days of debugging. It also gives you a clear contract. The documentation page that FastAPI generates is the contract between your model and everyone who uses it.
One note for readers in the Gulf and Egypt. If your employer handles personal or regulated data, where the API runs matters as much as how it is written. A prediction endpoint that receives customer records is subject to the same data-residency rules as the database those records came from. Keep that in mind when you reach the deployment sections, because a cloud region is a decision, not a default.
What you need to follow along: Python installed, a terminal, and a text editor. No web experience is assumed. We define each term as it appears.
- Pick a model you have trained or imagine one, for example a house-price predictor.
- Write down the inputs it needs (names and types) and the output it returns.
- Imagine three callers: a website, a mobile app, a colleague's script. Ask what each would have to do to use your notebook directly.
How the web conversation works
Before writing any FastAPI, you need five terms that every web API shares. They are not specific to FastAPI, and you will use them in every interview and every debugging session.
An HTTP request is a message a client sends to a server. It has a method (the verb), a path (the address of the thing on the server), optional headers (labelled metadata), and an optional body (the payload). The server answers with an HTTP response: a status code, headers, and usually a body.
The methods you need are few. GET asks for something and should not change anything on the server. POST sends data to be processed or stored. There are others (PUT, PATCH, DELETE) but a first model API needs just these two. A prediction can be a GET when the inputs are one or two numbers, and a POST when the inputs are a structured record.
The path may include a path parameter, a variable piece such as /items/42, where 42 identifies one item. It may also carry a query string, the part after the question mark, such as /predict?x=2. Query parameters are for optional or simple settings. Large or structured inputs belong in the body, which for APIs is almost always JSON: a text format made of objects (curly braces with named fields), lists (square brackets), strings, numbers, booleans and null.
Status codes are three-digit numbers that tell the client what happened. You will meet these constantly:
| Code | Meaning | When you see it |
|---|---|---|
| 200 | OK | The request worked |
| 400 | Bad request | You rejected something on purpose |
| 404 | Not found | The path or the resource does not exist |
| 422 | Unprocessable content | FastAPI's validation rejected the input |
| 500 | Internal server error | Your code crashed |
The first digit is the useful part. Codes starting with 2 mean success, 4 mean the caller made a mistake, and 5 mean the server did. When a request fails, that first digit tells you whose problem it is before you read anything else.
Finally, a server is a long-running program that listens on a port, a numbered door on a machine. When you run an API locally you will see an address like http://127.0.0.1:8000. The first part, 127.0.0.1, means "this computer only", and 8000 is the port. Nobody else on the network can reach it, which is exactly what you want while learning.
- Open any website, press F12 to open developer tools, and go to the Network tab.
- Reload the page and click one request.
- Find its method, its status code, and its response headers.
The mental model: four nouns
FastAPI has a large documentation site, but the core fits in four nouns. Learn these and every example you read will make sense.
Path operation. A function bound to a method and a path through a decorator. The line @app.get("/predict") means: when a GET request arrives at /predict, call the function underneath. The function itself is the path operation function. "Path operation" is FastAPI's name for the pairing of the two. (A decorator is a line starting with @ that wraps the function below it. You do not need to know how decorators work to use this one.)
Parameters. What the request carries in. FastAPI looks at your function's arguments and their type hints and decides where each comes from. An argument that matches a name in the path is a path parameter. A simple-typed argument that does not is a query parameter. An argument whose type is a Pydantic model is read from the request body. Header and cookie parameters exist too, and you declare them explicitly.
Pydantic model. A class that describes the shape of a piece of data: its field names, their types, and which are optional. Models are used for request bodies and for responses. When data arrives, Pydantic checks it against the class, converts what it can (the string "3" becomes the integer 3 where the declared type is int), and raises an error describing anything that does not fit.
Dependency. A function that FastAPI calls for you before your path operation function and whose result is handed in as an argument. Dependencies are how you share things like a loaded model, a database connection, or a login check across many endpoints without copying code. They appear in a later section. For now just remember that you declare them with Depends().
A fifth idea ties the first four together: the application object. You create it once with app = FastAPI(), and it collects every path operation you register. The server is told to run app, and everything else hangs off it.
The order matters for debugging. If the request path does not match any path operation, you get 404. If it matches but the parameters do not validate, you get 422 and your function never runs. Only if both pass does your code execute. When something fails, ask which stage it failed at.
- Take the model from the previous section.
- Name the path operation (method and path) you would use.
- List its parameters and say whether each is a path, query, or body parameter.
POST /predict with a body of four named fields. That is your whole API, on paper.
Installing FastAPI and checking the setup
The FastAPI documentation recommends the uv package manager, and plain pip works everywhere. Both approaches are shown. Use whichever you already have.
With uv:
uv init awesome-project --bare
cd awesome-project
uv add "fastapi[standard]"
With pip, inside a virtual environment (a private folder of packages for one project, so projects do not interfere):
python -m venv .venv
source .venv/bin/activate # Linux and macOS
pip install "fastapi[standard]"
On Windows the activation line is .venv\Scripts\activate in Command Prompt and .venv\Scripts\Activate.ps1 in PowerShell.
Two details trip up almost everyone, so read them slowly.
First, the quotes around fastapi[standard] are not optional on macOS. The default shell there, zsh, treats square brackets as a pattern to match file names. Without the quotes you get zsh: no matches found: fastapi[standard] and nothing installs. The quotes tell the shell to pass the text through untouched.
Second, what [standard] means. It is an extra: a named bundle of optional packages. The plain fastapi package is minimal. fastapi[standard] adds the command-line tool (fastapi dev, fastapi run), the server that runs your app, and other defaults you will want. For learning, always install the standard extra. There is also fastapi[standard-no-fastapi-cloud-cli], the same without the command-line client for the official hosting product, for teams that do not want it.
To check that everything works:
fastapi --version
On a fresh install this prints the version of the command-line tool and of the cloud client. Those are separate numbers from the framework's own, so do not be alarmed that they look different from 0.142.2. To see the framework version, ask Python:
python -c "import fastapi; print(fastapi.__version__)"
If the shell says the fastapi command is not found, there are two usual causes. Either you installed plain fastapi without the standard extra, or your virtual environment is not active. Re-run the activation line and check that your prompt shows the environment name.
Python version: current FastAPI documentation uses modern syntax such as str | None, which needs a reasonably recent Python 3. If you are on an old interpreter, upgrade before you debug anything else. The documentation's Docker example uses Python 3.14, and that is what this guide was run on.
pip install fastapi[standard] without quotes fails on zsh. It is the most common first-day failure on a Mac, and it has nothing to do with FastAPI.- Create a folder, a virtual environment, and install
"fastapi[standard]". - Run
fastapi --version. - Deliberately deactivate the environment and run it again.
Your first API, step by step
Create a file named main.py in your project folder. This is the whole program:
from fastapi import FastAPI
app = FastAPI()
@app.get("/")
async def root():
return {"message": "hello"}
Read it line by line, because every part is a pattern you will reuse.
from fastapi import FastAPI brings in the class. app = FastAPI() creates the application object. The name app is a convention, and the tools look for it. @app.get("/") registers the function below as the handler for GET requests at the root path. The function returns a Python dictionary, and FastAPI converts it to JSON automatically. You never call a JSON function yourself.
Now run it:
fastapi dev
The command finds main.py in the current folder, imports app from it, and starts a development server. It prints something like this:
Starting FastAPI in development mode
Server started at http://127.0.0.1:8000
Documentation at http://127.0.0.1:8000/docs
INFO: Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)
INFO: Started server process [67552]
INFO: Waiting for application startup.
INFO: Application startup complete.
The exact decoration varies between versions, but three things are stable. The address is http://127.0.0.1:8000. The documentation lives at /docs. And "Application startup complete" tells you the app loaded without errors. Uvicorn is the server program that FastAPI starts for you.
fastapi dev runs in development mode, which watches your files and reloads automatically when you save a change. This is the reason to use it while learning: edit, save, retry, with no restart. It is not for production. For that there is a separate fastapi run, covered later.
Open a second terminal and call the API:
curl http://127.0.0.1:8000/
{"message":"hello"}
curl is a command-line tool that sends HTTP requests. You can also just paste the address into a browser. If neither works, check that the server terminal still shows the "Uvicorn running" line and that you used the same port.
Now the feature that makes FastAPI popular. Open http://127.0.0.1:8000/docs in a browser. You will see an interactive page listing your endpoint. Click it, press Try it out, press Execute, and the page sends a real request and shows you the response. You wrote no documentation. FastAPI built it from your code and its OpenAPI schema, a standard machine-readable description of an API, which you can see raw at /openapi.json. Whenever you add or change an endpoint, this page changes with it, and it can never drift out of date because it is generated.
If fastapi dev complains that it cannot find the app, check three things: that your file is called main.py, that you are in the same folder, and that the variable is called app. You can also pass the path explicitly with fastapi dev main.py.
/docs page lets you call every endpoint without writing a client. Use it constantly while you develop.- Create
main.pyas above and runfastapi dev. - Call it with
curl, then in the browser at/docs. - Change the message text, save, and call it again without restarting.
Inputs: path, query and body parameters
A prediction API takes inputs. FastAPI decides where an input comes from by looking at how you declare the function argument, so this section is really about reading your own function signatures.
Query parameters. Any simple-typed argument that is not in the path becomes a query parameter:
@app.get("/predict")
def predict(x: float, scale: float = 1.0):
return {"result": x * 42 * scale}
Here x has no default, so it is required. scale has a default, so it is optional. A call to /predict?x=2 returns {"result":84.0}. A call to /predict?x=2&scale=0.5 returns 42.0. The type hint float does real work: FastAPI converts the text "2" to the number 2.0 before your function runs.
Notice the word def rather than async def. We will explain the difference shortly. For now, this is the safe choice for code that calls a model.
Path parameters. A variable piece of the path is declared in curly braces and matched by name:
@app.get("/models/{model_name}")
def get_model(model_name: str):
return {"model": model_name}
A request to /models/churn calls the function with model_name="churn". Use path parameters to identify which thing, such as which model or which record, and query parameters for settings.
Body parameters. When you declare an argument whose type is a Pydantic model, FastAPI reads it from the JSON body. We build one next section.
What happens when a caller gets it wrong? This is the best part. Call /predict with no x:
{"detail":[{"type":"missing","loc":["query","x"],"msg":"Field required","input":null}]}
with HTTP status 422. And call /predict?x=abc:
{"detail":[{"type":"float_parsing","loc":["query","x"],"msg":"Input should be a valid number, unable to parse string as a number","input":"abc"}]}
Read these as a trained eye does. detail is always a list, because several things can be wrong at once. Each entry has a type (a machine-readable category), a loc (where the problem is: first the source such as query or body, then the field name), a msg (a human sentence), and the input that was received. Entries for rules with limits also carry a ctx with the limit values. You wrote no error handling for any of this. Your function never ran.
Parameters can carry extra rules through Query(), Path() and the like, using the Annotated type form that the current documentation teaches:
from typing import Annotated
from fastapi import Query
@app.get("/search")
def search(q: Annotated[str, Query(min_length=3, max_length=50)]):
return {"q": q}
Annotated[type, extra] means "this type, with this extra information attached". Older tutorials write q: str = Query(min_length=3) instead. That still appears around the internet, but the documentation now prefers Annotated, so learn it first.
None, means optional. Beginners add a default just to silence a 422 and then wonder why the model receives nonsense.- Add the
/predictendpoint above. - Call it with a valid
x, with nox, and withx=abc. - For each failure, find the
locandmsgin the response.
Request bodies and Pydantic models
Real model inputs are rarely one number. A house has an area, a number of rooms, a district. These belong in a JSON body, described by a Pydantic model.
from pydantic import BaseModel, Field
class House(BaseModel):
area_m2: float = Field(gt=0, description="Living area in square metres")
rooms: int = Field(ge=1, le=20)
district: str
has_parking: bool = False
class Prediction(BaseModel):
price: float
model_version: str
@app.post("/predict-house")
def predict_house(house: House) -> Prediction:
price = 1200 * house.area_m2 + 5000 * house.rooms
return Prediction(price=price, model_version="demo-1")
A model is a class that inherits from BaseModel, with one annotated attribute per field. area_m2: float says the field is a float and is required. has_parking: bool = False says it is optional with a default. Field(gt=0) adds a rule: greater than zero. ge means greater than or equal and le means less than or equal. The description shows up in the generated docs.
Send a request:
curl -X POST http://127.0.0.1:8000/predict-house \
-H "Content-Type: application/json" \
-d '{"area_m2": 120, "rooms": 3, "district": "Maadi"}'
{"price":159000.0,"model_version":"demo-1"}
The -X POST sets the method, -H sets a header telling the server the body is JSON, and -d supplies the body. Because this command has several parts, a mistake in any one produces a confusing error, so type it carefully.
The -> Prediction return annotation does two jobs. It tells FastAPI what shape the response has, so the docs show it, and it makes FastAPI validate and filter the output. If your function accidentally returns extra fields, such as an internal ID or a database password hash, the ones not in Prediction are dropped. This is a cheap safeguard, and it is a good habit to always declare what you return. You may also pass response_model=Prediction in the decorator, which does the same thing and is what you will see in older code.
Now send bad data:
curl -X POST http://127.0.0.1:8000/predict-house \
-H "Content-Type: application/json" \
-d '{"area_m2": -5, "rooms": 3, "district": "Maadi"}'
{"detail":[{"type":"greater_than","loc":["body","area_m2"],"msg":"Input should be greater than 0","input":-5,"ctx":{"gt":0.0}}]}
Status 422 again, with loc pointing into the body. A body that is not valid JSON at all, or not an object, is rejected the same way, with a message such as "Input should be a valid dictionary or object to extract fields from".
Three habits make models safer.
- Put the rules in the model, not the function. A rule written once as
Field(gt=0)is enforced for every caller and shown in the docs. A rule buried in anifstatement is hidden and easy to forget. - Use Pydantic v2 names. FastAPI moved to Pydantic v2 in version 0.100.0. To turn a model into a dictionary use
house.model_dump(), and for a JSON stringhouse.model_dump_json(). Older code that calls.dict()or.json()is v1 style and shows deprecation warnings. - Optional means
| None. Writedescription: str | None = Nonefor a field that may be absent. The| Noneis the type; the= Noneis the default. You need both.
The 400-range errors that you raise yourself use HTTPException, covered in the section on errors below.
House class can be imported by your training code, your tests and your API. One definition of the input means they cannot disagree.- Define a Pydantic model for the inputs of your own model, with at least one numeric rule.
- Add the
POSTendpoint and call it withcurl. - Send a body with a wrong type, then one missing a field, and read each error.
loc beginning with body, and the same endpoint working again with good data.
Serving a real model: load once with lifespan
This is the section that matters most for machine learning, and the place beginners most often go wrong.
A model must be loaded from disk into memory before it can predict. Loading is slow, often seconds, and for a large model, minutes. If you load it inside the endpoint function, you pay that cost on every request, and a hundred simultaneous callers load a hundred copies. The correct pattern is to load once when the server starts, keep it in memory, and have every request reuse it.
FastAPI gives you a hook for this called lifespan: an asynchronous context manager passed to the application. Code before the yield runs at startup. Code after it runs at shutdown. This is the official documentation's own example, which uses a stand-in function in place of a real model:
from contextlib import asynccontextmanager
from fastapi import FastAPI
def fake_answer_to_everything_ml_model(x: float):
return x * 42
ml_models = {}
@asynccontextmanager
async def lifespan(app: FastAPI):
# Load the ML model
ml_models["answer_to_everything"] = fake_answer_to_everything_ml_model
yield
# Clean up the ML models and release the resources
ml_models.clear()
app = FastAPI(lifespan=lifespan)
@app.get("/predict")
async def predict(x: float):
return {"result": ml_models["answer_to_everything"](x)}
Walk through what happens. When the server starts, FastAPI runs lifespan up to the yield. The model goes into the ml_models dictionary. The server then begins accepting requests, and each call to /predict looks the model up in the dictionary: no loading, no waiting. When the server stops, the code after yield runs, which clears the dictionary and lets you close files or connections.
To use a real scikit-learn model that you saved earlier with joblib, replace the fake line:
import joblib
@asynccontextmanager
async def lifespan(app: FastAPI):
ml_models["house_price"] = joblib.load("models/house_price.joblib")
yield
ml_models.clear()
@app.post("/predict-house")
def predict_house(house: House) -> Prediction:
model = ml_models["house_price"]
features = [[house.area_m2, house.rooms]]
price = float(model.predict(features)[0])
return Prediction(price=price, model_version="2026-10-01")
Two small details deserve attention. The call is wrapped in float(...) because many ML libraries return NumPy numbers, which are not plain Python types, and the response validation can reject or mishandle them. And the features are a list of lists, because scikit-learn expects a batch of rows even when you predict one.
Why a dictionary at module level? It is the simplest thing that works: a place both the lifespan function and the endpoints can reach. In the next sections we see a cleaner way to hand the model to endpoints, and a later guide covers application state.
Watch your startup log. A failing model load now fails at startup, loudly, before any request arrives. That is a feature: a server that starts successfully is a server that has its model. If the file is missing you will see the FileNotFoundError in the terminal and the app will not start, instead of the first user receiving a 500.
A deprecation to know about: older tutorials load models with @app.on_event("startup") and @app.on_event("shutdown"). Those are deprecated in favour of lifespan. Worse, if you pass lifespan= to FastAPI(...), any on_event handlers are silently not called. If you copy old code into a lifespan-based app and your model dictionary is empty, this is the reason.
lifespan, use in endpoints.- Train a tiny scikit-learn model (even
LinearRegressionon four rows) and save it withjoblib.dump. - Load it in
lifespanand serve it from aPOSTendpoint. - Rename the model file and restart the server.
def or async def: the question every beginner asks
You have now seen both def and async def on endpoints, and the documentation is relaxed about which to use. For model serving the choice has a real consequence, so here is the rule in full.
async def defines a coroutine: a function that can pause while it waits for something slow, such as a network call or a database reply, and let the server work on other requests in the meantime. The pause points are marked with await. The server's single worker handles many requests by switching between them whenever one is waiting.
Plain def defines an ordinary blocking function. FastAPI notices this and runs it in a threadpool: a set of background threads, so the main loop is not held up.
The official rules are short:
- Use
async defwhen the library you call is designed to be awaited, so you writeawait something(). - Use plain
defwhen the library is not awaitable. Most database drivers and most machine learning libraries are not. - When unsure, use
def.
Now the trap. Calling model.predict(...) is blocking, CPU-heavy work. If you put it inside an async def endpoint without awaiting anything, the server's event loop is occupied for the entire time the prediction runs. Every other request, even a trivial health check, waits. One slow prediction freezes the whole server.
Safe for a model call
def predict(...)with a blockingmodel.predict- FastAPI runs it in the threadpool
- Other requests keep being served
Freezes the server
async def predict(...)with a blockingmodel.predictinside- Nothing to await, so the loop is stuck
- Every other request waits
The async def example in the official lifespan snippet is fine because its fake model returns instantly. A real model does not. When you replace the fake with a real one, change the endpoint to def.
A threadpool is not a cure for slow models, though. Python threads share one process, and heavy computation in pure Python code is still limited by a lock called the Global Interpreter Lock. Many numerical libraries release that lock while they compute, which helps, but it is not guaranteed. The documentation's advice for CPU-bound work such as inference is to use multiprocessing. In practice that means running several worker processes or several copies of the service, which a later section covers. For a beginner: use def, measure, and scale out when one process is not enough.
await, write def.- Make an endpoint that does
time.sleep(5)insideasync def. - Make a second one that does the same inside plain
def. - Call the slow one in one terminal and a fast endpoint in another.
async def version and answering instantly with def.
Dependencies: handing things to endpoints
You now have a model in a module-level dictionary and endpoints that reach into it. That works, but it couples every endpoint to a global, and it makes tests awkward. FastAPI's answer is dependency injection: instead of reaching out for what you need, you declare it as a parameter, and FastAPI supplies it.
A dependency is any function. You attach it to an endpoint argument with Depends:
from typing import Annotated
from fastapi import Depends
def get_model():
return ml_models["house_price"]
Model = Annotated[object, Depends(get_model)]
@app.post("/predict-house")
def predict_house(house: House, model: Model) -> Prediction:
price = float(model.predict([[house.area_m2, house.rooms]])[0])
return Prediction(price=price, model_version="2026-10-01")
Read Depends(get_model) as "before running this endpoint, call get_model and pass me whatever it returns as model". Pass the function itself, without parentheses: writing Depends(get_model()) calls it too early and is a classic slip.
Model = Annotated[object, Depends(get_model)] is a type alias: a short name for a long annotation, so every endpoint can write model: Model instead of repeating the whole thing. It is the pattern the documentation uses for shared dependencies.
Why bother? Three reasons that you will appreciate as the project grows.
- One place to change. If you later load the model from a registry instead of a file, or add a check that the model is ready, you edit
get_modeland no endpoint changes. - Easy testing. Tests can replace
get_modelwith a fake, so they do not need the real model file. That is the subject of the testing section. - Sharing logic. Anything several endpoints need, such as checking an API key, parsing pagination parameters, or opening a database session, can be a dependency.
Dependencies can have their own dependencies, forming a small tree, and FastAPI resolves it for you. They can be def or async def in any mix. A dependency can also take query or header parameters, which then appear in the documentation for every endpoint that uses it.
An example that shows the sharing idea is a simple API-key check:
from fastapi import Header, HTTPException
def require_api_key(x_api_key: Annotated[str, Header()]):
if x_api_key != "change-me":
raise HTTPException(status_code=401, detail="Invalid API key")
@app.post("/predict-house", dependencies=[Depends(require_api_key)])
def predict_house(house: House, model: Model) -> Prediction:
...
The dependencies=[...] list in the decorator is for dependencies that must run but whose return value you do not need. Two things to note. FastAPI converts the argument name x_api_key into the header name x-api-key for you, replacing underscores with hyphens. And the hard-coded key here is only for learning. Real keys come from environment variables or a secret manager, never from source code.
- Replace direct dictionary access in your endpoint with a
get_modeldependency. - Add the API-key dependency and call the endpoint with and without the header.
- Look at the
/docspage: what changed for the endpoint?
x-api-key field that now appears in the docs.
Errors: raising them and reading them
An API that only handles good input is half an API. You will deal with errors in two directions: the ones you raise on purpose, and the ones that reach you uninvited.
Raising errors on purpose. Use HTTPException:
from fastapi import HTTPException
@app.get("/models/{model_name}")
def get_model_info(model_name: str):
if model_name not in ml_models:
raise HTTPException(status_code=404, detail="Model not found")
return {"model": model_name}
raise stops the function immediately, and FastAPI converts the exception into a response with the status you chose and a body of {"detail": "Model not found"}. Pick the status code that matches the situation: 404 when a thing does not exist, 400 for a request that is well-formed but unacceptable (for example a combination of values your model cannot handle), 409 for a conflict such as creating something that already exists.
Reserve error responses for the caller's mistakes. If your code has a bug, you will get a 500 automatically, and you should fix the bug rather than catching it. A common beginner mistake is wrapping everything in try/except and returning 200 with an "error" field. That hides failures from monitoring tools and from callers, which rely on the status code.
Reading errors that arrive. Here is a short field guide, based on what you will actually see.
| Symptom | What it means | What to do |
|---|---|---|
422 with loc: ["body", ...] |
The JSON body failed validation | Read loc and msg, fix the request |
422 with loc: ["query", ...] |
A query parameter is missing or wrong | Same, and check spelling |
404 {"detail":"Not Found"} |
No path operation matches | Check path and method; a POST to a GET route gives 405 |
| 405 Method Not Allowed | Path exists, method does not | Use the declared method |
| 500 Internal Server Error | Your code raised an exception | Read the traceback in the server terminal |
command not found: fastapi |
Plain install or inactive venv | Install the standard extra, activate |
Address already in use |
Another process owns the port | Stop it, or use --port 8001 |
The most important line is the 500 row. The browser or curl only receives the words "Internal Server Error". The reason is in the terminal where the server runs. Beginners stare at the client and miss the traceback printed a few inches away. The last line of a traceback names the exception, and the lines above it point to your file and line number.
Validation also covers responses. If you declared -> Prediction and your function returns something that cannot be turned into a Prediction, FastAPI raises an error on the server and the caller gets a 500, with the details in the server log. That is a bug in your code, not in the request.
You can also control how a Pydantic model treats unexpected extra fields. By default, extra fields in the body are ignored. Whether to forbid them is a design choice: forbidding catches typos in field names (such as rooom), at the cost of being strict with clients who send harmless extras.
- Add an endpoint that raises
HTTPException(status_code=404, detail="Model not found"). - Add another with a deliberate bug, such as dividing by zero.
- Call both, and compare what the client sees with what the server terminal prints.
Testing without starting a server
Testing an API by hand with curl stops scaling around the fifth endpoint. FastAPI ships a test client that calls your app directly in Python, with no server and no network.
Install the two packages you need (the test client uses httpx, an HTTP library):
uv add httpx pytest
# or: pip install httpx pytest
Create test_main.py next to main.py:
from fastapi.testclient import TestClient
from main import app
client = TestClient(app)
def test_predict_house_rejects_negative_area():
response = client.post(
"/predict-house",
json={"area_m2": -5, "rooms": 3, "district": "Maadi"},
)
assert response.status_code == 422
assert response.json()["detail"][0]["loc"] == ["body", "area_m2"]
Run it:
pytest
Note that the test function is a plain def, not async, and you do not use await. The client handles that. client.post(...) takes a json= argument and does the encoding and the header for you, which is less error-prone than the long curl command. The response object offers status_code and .json().
One complication specific to this guide's pattern. A bare TestClient(app) does not run the lifespan code, so the model never loads and ml_models["house_price"] raises a KeyError. To run startup and shutdown, use the client as a context manager:
def test_predict_house_works():
with TestClient(app) as client:
response = client.post(
"/predict-house",
json={"area_m2": 120, "rooms": 3, "district": "Maadi"},
)
assert response.status_code == 200
assert "price" in response.json()
Inside the with block, the lifespan runs on entry and its shutdown half runs on exit.
This is where the dependency from the previous section pays off. The test above needs the real model file. A faster test replaces the dependency with a stub using app.dependency_overrides, a dictionary mapping a dependency function to a replacement:
class StubModel:
def predict(self, rows):
return [100.0]
def test_predict_with_stub_model():
app.dependency_overrides[get_model] = lambda: StubModel()
try:
client = TestClient(app)
response = client.post(
"/predict-house",
json={"area_m2": 120, "rooms": 3, "district": "Maadi"},
)
assert response.json()["price"] == 100.0
finally:
app.dependency_overrides.clear()
Here from main import get_model is also needed at the top. The try/finally guarantees the override is removed even if the assertion fails, so one test cannot leak into the next. Stub tests check your API logic (validation, error codes, response shape) in milliseconds, while a few tests with the real model check that the model loads and predicts. You want both kinds.
Depending on your versions of Starlette and httpx, you may see a deprecation warning about httpx when the test client is imported. In our run with FastAPI 0.142.2 the warning appeared and the tests still passed. A warning is not a failure, but read the warning text when you upgrade, because the underlying test dependency may change.
- Write one test that sends a negative area and expects 422.
- Write one test using
with TestClient(app) as clientand a real model. - Write one test with a stub model through
dependency_overrides.
Running it properly: dev, run, and configuration
So far you have used fastapi dev. It is for development only. When the API runs for other people, use the production command:
fastapi run
fastapi run starts the same app without reload and bound for real use. You can point it at a file and set the port:
fastapi run app/main.py --port 80
It also accepts --workers to start several worker processes and --proxy-headers, used when the app sits behind a proxy that terminates HTTPS. Why would you want several workers? One Python process runs one thing at a time for CPU-heavy work, so on a server with several cores you start several processes. The cost is memory: each worker loads its own copy of the model. Four workers with a 2 GB model use roughly 8 GB. Check your memory before you raise the count.
In a managed environment such as Kubernetes (covered in the Kubernetes guide), the usual practice is the opposite: one process per container, and let the platform run several containers. You do not need that decision yet, only to know that it exists.
Configuration. Your API will need settings: the path to a model file, an API key, a log level. Do not write them into the code. Read them from environment variables, which are named values the operating system gives to a program:
import os
MODEL_PATH = os.environ.get("MODEL_PATH", "models/house_price.joblib")
MODEL_PATH=/data/models/v2.joblib fastapi run
The first argument is the variable name and the second is a default for local development. The same code then works on your laptop, in CI and in production, with only the environment changing. FastAPI's documentation has a page on settings and environment variables built around Pydantic's settings support, which gives you typed, validated settings. It is worth reading once your configuration has more than three values.
Secrets such as API keys never go in the code or the repository. Pass them through the environment or a secret manager. Once a key has been committed to a public repository, treat it as leaked, even after you delete the commit.
A security note that is easy to forget. The interactive docs at /docs are served by default, and they list every endpoint. That is wonderful in development. For a public service, decide deliberately whether to expose them. You can turn them off when creating the app by passing docs_url=None, redoc_url=None and openapi_url=None to FastAPI(...). The parameter names are in the documentation's metadata and URLs page, and this is a mid-level topic, but know that the choice exists.
Finally, a quick word on networking. The address 127.0.0.1 accepts connections only from your own machine. Inside a container you must listen on all interfaces so that the outside world can reach the port. fastapi run handles this for you, which is another reason to use it rather than a raw server command inside a Dockerfile.
- Make the model path configurable through an environment variable with a default.
- Run
fastapi runon a different port, using--port 8001. - Point
MODEL_PATHat a file that does not exist.
Packaging it in a container
Your API works on your machine. To hand it to a colleague, or to run it on a server, the standard tool is a container image: your code, Python and every dependency in one portable package. If containers are new to you, read the Docker guide first, then return here.
First, list your dependencies in requirements.txt. Pin the framework to a minor version, because FastAPI is still 0.x and minor releases may include breaking changes:
fastapi[standard]>=0.142.0,<0.143.0
scikit-learn
joblib
(Pin your ML libraries to the exact versions you trained with. A model saved with one scikit-learn version may not load in another.)
Then a Dockerfile, following the shape in the official deployment documentation:
FROM python:3.14
WORKDIR /code
COPY ./requirements.txt /code/requirements.txt
RUN pip install --no-cache-dir --upgrade -r /code/requirements.txt
COPY ./app /code/app
CMD ["fastapi", "run", "app/main.py", "--port", "80"]
Three decisions in these few lines are deliberate.
The requirements.txt is copied and installed before the application code. Docker caches each step, and dependencies change rarely while code changes constantly. With this order, editing your code rebuilds only the last copy step, not the slow install.
The CMD uses the exec form, a JSON list of strings, and not the shell form (CMD fastapi run ... as a bare string). With the shell form, the signals Docker sends to stop the container do not reach your app, which breaks graceful shutdown, so your lifespan cleanup does not run, and in Docker Compose it causes a delay of about ten seconds on every stop. The documentation warns against it explicitly.
And the base image is python. You may find older tutorials that use a tiangolo/uvicorn-gunicorn-fastapi image. That image is deprecated, so build your own as above.
Your project layout for this Dockerfile looks like this:
.
├── Dockerfile
├── requirements.txt
└── app
├── main.py
└── models
└── house_price.joblib
Build and run it:
docker build -t house-api .
docker run -p 8000:80 house-api
The -p 8000:80 flag maps port 8000 on your machine to port 80 inside the container, which is where the CMD told the app to listen. Then http://127.0.0.1:8000/docs works exactly as before. If the container exits immediately, run docker logs on it and read the last lines: a model path that is wrong inside the container is the usual culprit, because relative paths are resolved from the working directory, /code.
Model files raise a question: should the model be inside the image or fetched at startup? For a small model, copying it in keeps the image self-contained and the deployment simple. For a large one, fetching from storage at startup keeps the image small but makes startup depend on the network. Both are valid, and you will meet this trade-off again in the mid-level guide.
For regional readers: if your data may not leave the country, choose the cloud region your container runs in deliberately, and prefer a region inside the Gulf or Egypt where your provider offers one. The image is portable, but the data it processes is not free to roam.
- Move your app into an
appfolder and write the Dockerfile above. - Build the image and run it with
-p 8000:80. - Open
/docsand send a prediction from the browser.
Putting it all together
Here is one small project that uses everything: a house-price service with validation, a model loaded once, a dependency, an error, and a test. Create the folder house-api with this layout:
house-api/
├── main.py
├── test_main.py
├── train.py
└── models/
First, train.py creates the model file. It uses scikit-learn, which you install with pip install scikit-learn:
import joblib
from sklearn.linear_model import LinearRegression
X = [[50, 1], [80, 2], [120, 3], [200, 5]]
y = [60000, 100000, 159000, 270000]
model = LinearRegression().fit(X, y)
joblib.dump(model, "models/house_price.joblib")
print("saved")
Run mkdir models && python train.py. Now the application:
import os
from contextlib import asynccontextmanager
from typing import Annotated
import joblib
from fastapi import Depends, FastAPI, HTTPException
from pydantic import BaseModel, Field
MODEL_PATH = os.environ.get("MODEL_PATH", "models/house_price.joblib")
ml_models = {}
@asynccontextmanager
async def lifespan(app: FastAPI):
ml_models["house_price"] = joblib.load(MODEL_PATH)
yield
ml_models.clear()
app = FastAPI(title="House price API", lifespan=lifespan)
class House(BaseModel):
area_m2: float = Field(gt=0, description="Living area in square metres")
rooms: int = Field(ge=1, le=20)
district: str
class Prediction(BaseModel):
price: float
def get_model():
return ml_models["house_price"]
Model = Annotated[object, Depends(get_model)]
@app.get("/health")
def health():
return {"status": "ok", "model_loaded": "house_price" in ml_models}
@app.post("/predict-house")
def predict_house(house: House, model: Model) -> Prediction:
if house.area_m2 > 2000:
raise HTTPException(status_code=400, detail="Area outside the trained range")
price = float(model.predict([[house.area_m2, house.rooms]])[0])
return Prediction(price=price)
Note the /health endpoint. It costs four lines and every deployment platform you will meet uses something like it to ask "is this service alive?". The 400 check shows a business rule that Pydantic's simple limits cannot express, for example refusing inputs far outside what the model saw in training. A model asked about a 50,000 square metre house will happily return a number, and the number will be wrong.
Now the tests:
from fastapi.testclient import TestClient
from main import app
def test_health():
with TestClient(app) as client:
assert client.get("/health").json()["model_loaded"] is True
def test_prediction_is_a_number():
with TestClient(app) as client:
r = client.post("/predict-house", json={"area_m2": 100, "rooms": 3, "district": "Maadi"})
assert r.status_code == 200
assert isinstance(r.json()["price"], float)
def test_validation_rejects_bad_rooms():
with TestClient(app) as client:
r = client.post("/predict-house", json={"area_m2": 100, "rooms": 0, "district": "Maadi"})
assert r.status_code == 422
def test_out_of_range_area_is_400():
with TestClient(app) as client:
r = client.post("/predict-house", json={"area_m2": 5000, "rooms": 3, "district": "Maadi"})
assert r.status_code == 400
Run pytest, then fastapi dev, open /docs, and try the endpoint by hand. You have built a validated, tested, documented model API in about eighty lines.
- Build the project exactly as shown and make all four tests pass.
- Add a new field,
has_parking: bool = False, toHouse. - Check
/docsand re-run the tests.
What you can now do, and what comes next
You can now explain what an HTTP API is and read status codes by their first digit. You can install FastAPI with the right extra and quote it correctly on a Mac. You can write path operations that take query, path and body parameters, describe inputs with Pydantic models, and let the framework reject bad requests with a precise 422. You can load a model once in lifespan, hand it to endpoints through a dependency, and choose between def and async def with a reason. You can raise your own errors, read a 500 by finding the traceback in the server terminal, test with TestClient including startup code, replace the model with a stub, and package the whole thing in a container with an exec-form CMD.
Three cautions to carry forward. FastAPI is a 0.x library, so pin the minor version and read the release notes' breaking-changes entries before upgrading; version 0.137.0, for example, changed how routers include other routers, which affects code that inspects router.routes. Prefer the current forms: Annotated over default-value Depends, lifespan over on_event, Pydantic v2 names such as model_dump. And remember that FastAPI makes the web layer easy but does not make your model fast or safe on its own. Slow predictions need more processes or a dedicated serving system, and personal data needs the same care it had before it reached an API.
What comes next. The mid-level guide covers how requests are actually processed, routers for organising larger apps, background tasks, middleware, CORS, settings, structured logging, proper async patterns, and running several workers or replicas. The senior guide covers the platform view: scaling, security, observability including the native OpenTelemetry support added in 0.142.0, upgrades and when to choose something else.
If you want to go sideways instead, these guides are natural neighbours. MLflow tracks the experiments and models that you would load into this API. BentoML and Triton are dedicated model-serving tools you can compare with a hand-written FastAPI service, and KServe runs model servers on Kubernetes. Knowing where a plain FastAPI service is enough, and where a specialised server is better, is one of the clearest signs of experience.
Sources
- FastAPI home and documentation: https://fastapi.tiangolo.com/
- Tutorial index (installation, first steps, running): https://fastapi.tiangolo.com/tutorial/
- Release notes: https://fastapi.tiangolo.com/release-notes/
- Lifespan events and the machine learning model example: https://fastapi.tiangolo.com/advanced/events/
- Concurrency and async/await (
defversusasync def): https://fastapi.tiangolo.com/async/ - Dependencies: https://fastapi.tiangolo.com/tutorial/dependencies/
- Testing: https://fastapi.tiangolo.com/tutorial/testing/
- Deployment overview: https://fastapi.tiangolo.com/deployment/
- Deploy with Docker: https://fastapi.tiangolo.com/deployment/docker/
- About FastAPI versions: https://fastapi.tiangolo.com/deployment/versions/
- FastAPI 0.137.0 release notes on GitHub: https://github.com/fastapi/fastapi/releases/tag/0.137.0