This is part one of three. It covers everything you need to take a model that works in a notebook and turn it into an HTTP service someone else can call, packaged as an artifact you can hand to a colleague or a cluster. By the end you can write a BentoML service, serve it locally, call it from curl and from Python, save a model into BentoML's store, build a versioned Bento, turn that Bento into a container image, and read the errors that come up along the way. Mid-level and Senior take the same topics further — batching, concurrency, multi-service graphs, autoscaling, security — and nothing here is thrown away when you get there.
Each section ends with a Try it task. Do them as you go. Serving is one of those subjects that looks obvious on the page and teaches you something the first time your own service refuses to start.
Everything below is written against BentoML 1.4.x and the current service API. That qualifier matters more than usual here, and the section on the old API explains why.
What BentoML is, and the problem it solves
BentoML is a Python framework for turning model inference code into a running HTTP service, and for packaging that service into a single versioned artifact. You write a normal Python class, decorate it, and BentoML gives you a web server, request validation, an OpenAPI description, Prometheus metrics, a build command that produces a reproducible bundle, and a command that turns that bundle into a container image.
That list is deliberately boring. The interesting claim BentoML makes is not that it can serve a model — anything can serve a model — but that serving and packaging are the same problem, and that solving them separately is what goes wrong in practice.
Read that left to right as the arc of this guide. A trained model on its own is a file. Wrapped in a service it becomes something callable. Built into a Bento it becomes something storable, listable and shippable. Containerized it becomes something a platform team can run without knowing any Python at all.
To see why that arc is worth a framework, look at what people build without one. The usual first attempt is a small web application — a Flask or FastAPI file that loads a pickle at import time and exposes one POST route that pulls JSON out of the request body, calls model.predict, and returns JSON. It takes twenty minutes and it works. It is also where a long list of problems starts.
The model loads once per worker process, or once per request if you got it wrong, and nothing in the code tells you which. The input is parsed by hand, so a client sending a string where a float was expected produces a five-hundred error and a stack trace in the logs rather than a clear four-hundred response. There is no description of the endpoint, so the team consuming it reads your source code or guesses. The requirements.txt next to the file drifts from the environment you actually trained in. The weights live somewhere outside version control — a shared drive, an S3 bucket, someone's home directory — so "which model is in production" is answered by asking around. And when a second model arrives, the whole thing is copied and the copies diverge.
None of those are hard problems individually. Together they are the reason a model that worked in March is unreproducible in June. BentoML's answer is to make the service definition carry all of it: the code, the dependency specification, the model references and the runtime configuration travel inside one artifact with a version string on it.
Two consequences of that design shape everything that follows.
A Bento is an immutable, self-describing bundle. When you run bentoml build, BentoML collects your source files, records the Python packages, notes which service class to start and which models to include, and writes the whole thing into the local Bento store under a name and an auto-generated version. You do not edit a Bento; you build a new one. That is what makes "roll back to the version that worked" a real option rather than a hopeful git checkout.
Your service is ordinary Python. There is no configuration language for describing inputs, no graph to declare, no model-server plugin to write. You load the model in __init__ and you write methods. Type hints on those methods become the request and response schema. This is the part people find surprisingly pleasant: the thing you test in a unit test is the thing that serves traffic.
What teams actually use it for:
Model as an API
A data scientist's model becomes an endpoint the backend team can call, with a schema and a generated OpenAPI description.
Reproducible artifacts
One versioned Bento per release, built in CI, exportable as a file and importable anywhere.
Container without Dockerfile work
bentoml containerize generates the image, so platform teams get an ordinary OCI image to deploy.
Pipelines of services
CPU preprocessing and GPU inference as separate services that scale independently, called like local objects.
You need very little to follow along: Python 3.9 or newer, a terminal, and — for the container sections only — a container runtime such as Docker or Podman. A model is optional at first; the first service we write does not need one, which is on purpose, because the point of the first exercise is the serving machinery and not the machine learning.
- Find a model you or a colleague has deployed, however informally, and write down where its weights live and what pins its dependency versions.
- Write down how a caller discovers the input format.
- Write down what you would do to put last month's version back.
Before you read any tutorial: the old API
This section is out of order on purpose. BentoML changed its public API in version 1.2, and the overwhelming majority of blog posts, Stack Overflow answers, course videos and generated code still show the older style. If you do not know what the old style looks like, you will copy it, and it will not run.
The old pattern, from 1.1 and earlier, looked like this. It is shown here only so you can recognise it:
import bentoml
from bentoml.io import NumpyNdarray
runner = bentoml.sklearn.get("iris_clf:latest").to_runner()
svc = bentoml.Service("iris_classifier", runners=[runner])
@svc.api(input=NumpyNdarray(), output=NumpyNdarray())
def classify(input_series):
return runner.predict.run(input_series)
Three tells identify it. There is a module-level svc = bentoml.Service(...) object. There are IO descriptors imported from bentoml.io — NumpyNdarray, JSON, Image — passed as input= and output=. And there are Runners, created with .to_runner() and called with runner.predict.run(...).
The current style, from 1.2 onward, replaces all three. The service is a class decorated with @bentoml.service. The endpoints are methods decorated with @bentoml.api. The input and output types are expressed as ordinary Python type hints instead of descriptor objects. And where you previously used a Runner to put model inference in its own process, you now use bentoml.depends() to depend on another service — a mid-level topic, but worth knowing the name of now.
import bentoml
@bentoml.service
class IrisClassifier:
def __init__(self) -> None:
...
@bentoml.api
def classify(self, sepal_length: float) -> str:
...
bentoml.io or to_runner, the example is out of date
Runners and IO descriptors are the legacy pattern. Code written against them raises import or attribute errors on a current install — typically ModuleNotFoundError: No module named 'bentoml.io' or an AttributeError about bentoml having no attribute Service. The fix is never to downgrade; it is to rewrite the handful of lines in the current style. Treat the 1.1 to 1.2 move as a rewrite of your service file, not a dependency bump.
Why spend a whole section on this? Because it is the single most common way beginners lose an afternoon with BentoML, and because knowing the history makes the current design easier to read. The old API asked you to describe your inputs in BentoML's vocabulary. The new one reads your Python type hints instead. Everything that looks elegant about the current API is the removal of a translation layer.
- Search the web for "bentoml example" and open the first three results.
- For each, look for
bentoml.io,bentoml.Service(orto_runner. - Note which ones you would have copied.
The four nouns
BentoML has a small vocabulary, and nearly every error message and documentation page uses it precisely. Four nouns carry most of the weight.
| Noun | What it is | Where it lives | How you make one |
|---|---|---|---|
| Service | A Python class decorated with @bentoml.service |
Your source file | Write it |
| API | A method decorated with @bentoml.api, exposed as an HTTP endpoint |
Inside a service | Write it |
| Model | Saved weights under a name and version | The local model store | bentoml.<framework>.save_model or a reference |
| Bento | The built bundle: code, dependency spec, models, config | The local Bento store | bentoml build |
A service is the unit of deployment. One service becomes one set of worker processes with its own resource requests and its own scaling behaviour. Its __init__ runs once per worker, which is where model loading belongs: load the weights there and every request after the first finds them already in memory. Anything expensive and per-process goes in __init__.
An API is one HTTP endpoint. By default the route is the method name, so a method called summarize answers at /summarize. You can override that with @bentoml.api(route="/custom/path") when the endpoint must match a contract you do not control. The parameters and return annotation of the method define what the endpoint accepts and returns, and BentoML validates incoming requests against them before your code runs.
A model in BentoML's sense is an entry in the local model store: weights plus metadata, addressed as name:version, with name:latest resolving to the newest. This matters because it separates "the model" from "the code that serves it". Two services can reference the same stored model; a new model version can be saved without touching the service; and the Bento records which model version it was built with. Models can also be references to external sources rather than local copies — a Hugging Face model reference, for instance — which keeps large weights out of your artifact.
A Bento is everything above, frozen. bentoml build produces one, tagged name:version, where the name comes from your service class and the version is generated for you. It sits in the local Bento store until you export it to a file, push it to BentoCloud, or containerize it.
Two more words you will meet early, worth defining now even though they belong to later sections. A deployment is a running instance of a Bento on BentoCloud, the hosted platform. Adaptive batching is BentoML grouping concurrent requests into a single model call, switched on per API with batchable=True; it is the main throughput lever and the subject of the mid-level guide.
- Say out loud, without looking: what is the difference between a model and a Bento?
- Which of the two has a version you choose, and which has one generated for you?
- Where does model loading belong, and why not in the API method?
__init__ happens once per worker; loading inside the API method happens once per request, which is the classic beginner performance bug.
Installing BentoML and checking the setup
BentoML is a Python package. The only hard requirement is Python 3.9 or newer. Install it into a virtual environment — not because BentoML is unusually fragile, but because BentoML reads your environment to decide what to put in the Bento, so a clean environment is a clean artifact.
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install bentoml
bentoml --version
On Windows, use PowerShell:
py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install -U pip
pip install bentoml
bentoml --version
If PowerShell refuses to run the activation script, Set-ExecutionPolicy -Scope CurrentUser RemoteSigned once will allow it. For the container sections later on, Windows users are better served by working inside WSL2, where the container runtime and the Linux filesystem behave the way the documentation assumes.
The second command to run is the one nobody runs until they need help:
bentoml env
This prints your system information and the versions of the relevant packages. Run it now so you know what healthy looks like, and include its output whenever you ask a colleague or a maintainer for help — most "it does not work on my machine" threads are answered by one line of it.
For production, pin the version rather than floating:
pip install "bentoml==1.4.39"
Pinning matters here more than for most libraries, for two reasons. BentoML's API changed meaningfully at 1.2, so an unpinned install in a long-lived environment can pick up a version your code does not match. And the 1.4 patch series has carried several security fixes — file path resolution on user input, symlink validation when extracting archives, hardening of the template engine that generates Dockerfiles — which is an argument for pinning to a specific patch and then deliberately moving that pin forward, rather than never pinning and never thinking about it.
pip install, model downloads and any call to a hosted platform fail with SSL: CERTIFICATE_VERIFY_FAILED or SELF_SIGNED_CERT_IN_CHAIN. The correct fix is to export the proxy's CA certificate and point REQUESTS_CA_BUNDLE, SSL_CERT_FILE and PIP_CERT at the bundle. Do not disable verification; you will forget, and the setting will follow you into something that matters.
Telemetry is worth a word because it surprises people. BentoML's CLI reports anonymous usage by default. Every command accepts --do-not-track, and the behaviour can be switched off for a whole environment with the BENTOML_DO_NOT_TRACK environment variable. If you work somewhere that reviews outbound traffic, set it before your first command rather than explaining it afterwards.
- Create a fresh virtual environment and install BentoML into it.
- Run
bentoml --versionand thenbentoml env, and read the output properly. - Deactivate the environment and run
bentoml --versionagain.
Your first service
We will build a text summarization service, because it needs no training step and exercises every part of the machinery: a model loaded once, an endpoint with a typed input, a real dependency, and something interesting to look at in the response.
Install the two extra packages:
pip install torch transformers
Create a file called service.py:
import bentoml
from transformers import pipeline
@bentoml.service
class Summarization:
def __init__(self) -> None:
self.pipeline = pipeline("summarization")
@bentoml.api
def summarize(self, text: str) -> str:
result = self.pipeline(text)
return result[0]["summary_text"]
Eleven lines, and every one of them is doing something worth naming.
@bentoml.service on the class is what makes this more than a class. It tells BentoML this is a deployable unit, and it is where configuration goes later — resources, worker count, timeouts, metrics settings all become arguments to this decorator.
__init__ runs once per worker process when the server starts. Downloading and constructing a Hugging Face pipeline takes seconds; doing it here means those seconds are paid at startup, once, and not on every request. Note that there is nothing BentoML-specific inside it. It is the code you would write in a notebook.
@bentoml.api on summarize turns the method into an endpoint. The route defaults to the method name, so this one answers at /summarize.
The annotations text: str and -> str are the contract. BentoML reads them, builds a request schema from them, validates incoming JSON against that schema, and publishes the result as an OpenAPI description. A caller who sends a number instead of a string is rejected before your pipeline is ever invoked, with a response that says what was wrong.
Start it:
bentoml serve service:Summarization
The argument is module:ClassName — the Python module (service, from service.py) and the class inside it. Get either wrong and the server will tell you it could not import the service. Run it from the directory that contains service.py, because that is what makes the module importable.
The server listens on port 3000 by default. The first start is slow, because the model downloads; later starts are fast. When it is up, the logs tell you the address, and BentoML serves an interactive API page where you can see every endpoint it found and try one from the browser. That page is generated from your type hints, which is the first moment the design pays off visibly: you wrote two annotations and got documentation.
Call it from another terminal:
curl -X POST http://localhost:3000/summarize \
-H "Content-Type: application/json" \
-d '{"text": "BentoML is a Python framework for serving and packaging machine learning models. It turns inference code into an HTTP service and bundles the code, its dependencies and its model references into a single versioned artifact that can be containerized and deployed."}'
Three details in that command are worth fixing in your memory, because every one of them is a question people ask. The method is POST, not GET. The content type is JSON. And the JSON keys are the parameter names of your method — text here, because the parameter is called text. Rename the parameter and the request body changes with it. That coupling is intentional: there is one definition of the interface and it is the Python signature.
curl. Almost every confusing serving problem is explained by a line in the server log that you never saw because the log was behind the window you were typing in.
- Build and serve the summarization service, then call it with
curl. - Open
http://localhost:3000in a browser and find thesummarizeendpoint. - Send
{"text": 42}and read the response and the status code. - Rename the parameter to
document, restart, and send the old body.
Calling your service three ways
A service that only answers curl is not finished. Three callers matter in practice, and each teaches something different.
The interactive page. BentoML serves a browser UI generated from the OpenAPI description of your service. Use it to check what a client will see: the endpoint list, the field names, the types, whether a field is required. When a frontend developer asks you for the API contract, this page is the answer, and because it comes from your type hints it cannot drift from the implementation.
curl or any HTTP client. The endpoint is an ordinary JSON POST. Nothing about calling it requires BentoML on the client side, which is the point: a Java backend, a Node service or a cron job can call it with whatever HTTP library it already has. This is also how you will test from CI.
The Python client. When the caller is Python, BentoML ships a client that makes the remote endpoints look like methods:
import bentoml
with bentoml.SyncHTTPClient("http://localhost:3000") as client:
summary = client.summarize(
text="BentoML turns inference code into an HTTP service and packages it."
)
print(summary)
There is an asynchronous equivalent for callers that already run an event loop. Use the client for integration tests and for one service calling another over HTTP during development. Note that the method name and keyword argument match your service exactly — the same contract again, from a third angle.
Which to reach for depends on what you are checking. If you are asking "is the server up and the model loaded", curl is fastest. If you are asking "what will the consuming team see", use the browser page. If you are writing a test, use the Python client, because a failure gives you a Python exception instead of a string you have to parse.
One endpoint you get for free and should know about: BentoML exposes Prometheus metrics on /metrics, on the same port, enabled by default. Open it now. The counters and histograms there — request totals, in-progress requests, request duration — are what monitoring will scrape later, and seeing them exist from the first service is worth more than reading about them in the senior guide.
- Call
summarizefrom the browser page, fromcurl, and from the Python client. - Open
http://localhost:3000/metricsand find a counter whose value went up by three. - Make one call, then immediately reload
/metricstwice and watch what changes.
Types are the interface
The single idea that most repays attention at this level is that your type hints are the API specification. Getting comfortable with that turns a lot of later work into type design rather than framework configuration.
Simple scalars work as you would expect: str, int, float, bool. Lists of them work. For anything with structure, use a Pydantic model, and your endpoint gains a documented, validated object schema:
from typing import List
import bentoml
from pydantic import BaseModel
class SummarizeRequest(BaseModel):
text: str
max_words: int = 60
class SummarizeResponse(BaseModel):
summary: str
input_words: int
@bentoml.service
class Summarization:
def __init__(self) -> None:
from transformers import pipeline
self.pipeline = pipeline("summarization")
@bentoml.api
def summarize(self, request: SummarizeRequest) -> SummarizeResponse:
result = self.pipeline(request.text)
return SummarizeResponse(
summary=result[0]["summary_text"],
input_words=len(request.text.split()),
)
Now the request body is a nested object, max_words has a default so clients may omit it, and the response has named fields instead of a bare string. The browser page updates itself. No descriptor classes, no schema file, no duplication between the validation layer and the implementation.
Two more type choices come up constantly in model serving.
Arrays. A numpy array annotation is accepted directly, which is what you want for a classical model whose predict takes a matrix. A list of floats is often the friendlier choice for a beginner API, because JSON clients do not have to think about shapes.
Files. For images, audio or documents, annotate with pathlib.Path. BentoML handles the upload and hands your method a path to a file on disk, which means your code opens a file like any other Python code rather than decoding a multipart body by hand.
from pathlib import Path
import bentoml
@bentoml.service
class ImageInfo:
@bentoml.api
def describe(self, image: Path) -> dict:
return {"name": image.name, "bytes": image.stat().st_size}
A word of advice about the return type. dict works and is tempting, but an untyped dictionary gives the caller no schema and gives you no protection against renaming a key. Prefer a Pydantic response model for anything another team consumes. The extra five lines are the cheapest documentation you will ever write.
Typed and documented
- Pydantic request and response models
- Invalid input rejected before the model runs
- Browser page shows every field and default
- Renaming a field is a visible change
Untyped and implicit
dictin,dictout- Bad input becomes a server-side exception
- Callers read your source to learn the fields
- Renaming a key breaks clients silently
- Convert your
summarizeendpoint to the Pydantic request and response models above. - Call it without
max_wordsand confirm the default applies. - Send
{"request": {"text": "hi", "max_words": "sixty"}}and read the error.
Saving and loading a model
So far the model has been downloaded by transformers at startup. That is fine for a demo and wrong for anything you need to reproduce, because "the model" is then whatever the internet served that morning. BentoML's model store exists to fix that.
The store is a local, versioned registry of saved models, managed through bentoml models. You save a trained object into it under a name; BentoML assigns a version and records metadata. From then on the model is addressable as name:version, or name:latest for the newest.
Train and save a small scikit-learn model so the whole loop is visible:
import bentoml
from sklearn import datasets, svm
X, y = datasets.load_iris(return_X_y=True)
clf = svm.SVC(gamma="scale")
clf.fit(X, y)
saved = bentoml.sklearn.save_model("iris_clf", clf)
print(saved.tag)
pip install scikit-learn
python train.py
bentoml models list
The printed tag looks like iris_clf:xxxxxxxxxxxx, with a generated version. Run train.py twice and bentoml models list shows two versions; iris_clf:latest points at the second. That is the whole mechanism, and it is enough to answer "which model is this" for the rest of your career.
Now serve it. The model is loaded in __init__, from the store, by tag:
import bentoml
import numpy as np
@bentoml.service
class IrisClassifier:
def __init__(self) -> None:
self.model = bentoml.sklearn.load_model("iris_clf:latest")
@bentoml.api
def classify(self, features: list[float]) -> int:
prediction = self.model.predict(np.array([features]))
return int(prediction[0])
bentoml serve service:IrisClassifier
curl -X POST http://localhost:3000/classify \
-H "Content-Type: application/json" \
-d '{"features": [5.1, 3.5, 1.4, 0.2]}'
Each supported framework has its own module with the same shape — save_model and load_model — so the pattern transfers. For deep learning work where the weights come from Hugging Face, BentoML can reference the external model instead of copying gigabytes into your store and your artifact, which keeps builds fast.
latest is convenient and dangerous
iris_clf:latest resolves at load time, so a service pinned to latest can change behaviour because somebody retrained, with no change to your code and nothing in your git history to explain it. Use latest while you are iterating locally. Pin an explicit version in anything you build for a shared environment, so the Bento records exactly which weights it serves.
Loading models from untrusted sources deserves one sentence now rather than a surprise later: many model formats are pickles, and loading a pickle executes code. The same applies to Bento archives you import from elsewhere. Treat both like running a script someone sent you, because that is what it is.
- Run
train.pytwice and list the model store. - Serve the classifier against
iris_clf:latestand classify a flower. - Change the service to pin the first version explicitly and restart.
Building a Bento
Your service runs. It runs on your laptop, in your virtual environment, from your working directory. bentoml build is the step that removes all three of those dependencies.
Build configuration lives in bentofile.yaml next to your code, or in pyproject.toml under [tool.bentoml.build] if you prefer one file. The minimum is which service to start and what to include:
service: "service:IrisClassifier"
include:
- "*.py"
python:
packages:
- scikit-learn
- numpy
bentoml build
bentoml list
The build prints a tag like iris_classifier:xxxxxxxxxxxx and stores the Bento locally. bentoml list shows what you have; bentoml get iris_classifier:latest shows the detail of one, including the path on disk — open that directory once, because seeing the structure makes the concept concrete. Inside you will find your source files, the recorded Python dependencies, the service configuration and the model references.
The fields of bentofile.yaml worth knowing on day one:
| Field | What it does |
|---|---|
service |
Required. The module:ClassName to serve. |
include |
Glob patterns of files to put in the Bento. |
exclude |
Patterns to leave out; a .bentoignore file works the same way, gitignore-style. |
python.packages |
Packages to install in the Bento's environment. |
python.requirements_txt |
Use an existing requirements file instead of listing packages. |
python.lock_packages |
On by default: resolves and pins the whole dependency set at build time. |
docker.distro |
Base distribution for the generated image, e.g. debian, alpine, ubi8, amazonlinux. |
docker.python_version |
Python version in the image. |
docker.system_packages |
OS packages to install, for libraries that need them. |
description, labels |
Documentation and metadata that travel with the Bento. |
Be deliberate with include. The default temptation is to include everything, which quietly puts your notebooks, your test data, your .env and your local virtual environment into a shareable artifact. List what the service needs. A Bento is something you hand to other people.
python.lock_packages is the field that earns its keep. With it on — which is the default — the build resolves your dependency specification into a fully pinned set and records that. Two builds from the same source therefore install the same versions, which is the difference between an artifact and a wish.
You can build from a different directory or config file when your repository layout calls for it:
bentoml build -f ./deploy/bentofile.yaml ./
And you can move Bentos between machines without any platform at all:
bentoml export iris_classifier:latest ./iris_classifier.bento
bentoml import ./iris_classifier.bento
bentoml delete iris_classifier:latest
That export file is the simplest possible answer to "how do I give this to the ops team". It is also, as noted above, executable content: import only what you trust.
- Write
bentofile.yamland runbentoml build. - Run
bentoml get iris_classifier:latest, find the path, and explore the directory. - Export the Bento to a file and check its size.
- Add
excludefor anything that should not have been there, rebuild, and compare sizes.
Serving a Bento, and containerizing it
A built Bento can be served directly by tag, which is how you check that the artifact — not your working directory — actually works:
bentoml serve iris_classifier:latest
This is a genuinely important test and people skip it. Serving from source uses your environment; serving the Bento uses what the Bento recorded. If the Bento is missing a file or a package, this is where you find out, in two seconds, instead of in a container build ten minutes later.
Then the step that makes the artifact portable:
bentoml containerize iris_classifier:latest
BentoML generates a Dockerfile from the Bento's configuration and builds an image through your local container backend. Docker is the common choice; Podman and Buildah also work. The image is an ordinary OCI image with no BentoML-specific runtime requirement on the host, which is exactly what a platform team wants to receive.
Run it:
docker run --rm -p 3000:3000 iris_classifier:xxxxxxxxxxxx
Use the tag the containerize command printed. Then call http://localhost:3000/classify exactly as before. The response should be identical to the one from bentoml serve, and confirming that identity is the point of the exercise: the thing you tested is the thing that ships.
Two practical notes. First, the container build installs your dependencies, so the first one is slow — expect minutes, not seconds, for anything with PyTorch in it. Second, if you are on an Apple Silicon Mac and the target is an x86 cloud instance, the image you build locally is ARM by default and will not run there; building for the target architecture is a container-level concern covered in the Docker guide. The honest advice for a beginner is to build deployment images in CI, on Linux, and keep local containerize runs for verification.
- Serve the Bento by tag and confirm the endpoint still answers.
- Run
bentoml containerizeand then run the image withdocker run. - Compare the
curlresponse from the container with the one frombentoml serve.
The commands you will actually use
Grouped by intent rather than alphabetically, because that is how you will reach for them.
While writing a service
bentoml serve service:MyService # run from source
bentoml env # environment report for bug reports
bentoml --version
Working with models
bentoml models list
bentoml models get iris_clf:latest
bentoml models delete iris_clf:xxxxxxxxxxxx
Building and inspecting Bentos
bentoml build
bentoml build -f ./deploy/bentofile.yaml ./
bentoml list
bentoml get iris_classifier:latest
bentoml serve iris_classifier:latest
bentoml delete iris_classifier:xxxxxxxxxxxx
Moving Bentos around
bentoml containerize iris_classifier:latest
bentoml export iris_classifier:latest ./out.bento
bentoml import ./out.bento
bentoml push iris_classifier:latest # to BentoCloud
bentoml pull iris_classifier:latest
Four options are shared by every command and are worth knowing as a set: --verbose when you need to see what BentoML is doing, -q or --quiet in scripts, --do-not-track to suppress telemetry for one invocation, and --context to select which configured remote a command talks to.
A habit worth forming now: after any build, run list and then serve the tag. Three commands, ten seconds, and it catches the two most common packaging mistakes — a missing file and a missing package — at the cheapest possible moment.
- Run each command in the first three groups once against your project.
- Delete an old Bento version and an old model version.
- Re-run a command with
--verboseand skim the extra output.
Configuring the service
Configuration in current BentoML is mostly arguments to the @bentoml.service decorator, which means it lives next to the code it configures and travels inside the Bento. If you find a tutorial editing a global YAML configuration file, you are reading about the older style again.
import bentoml
@bentoml.service(
resources={"cpu": "2"},
traffic={"timeout": 30},
workers=2,
path_prefix="/v1",
)
class IrisClassifier:
...
resources declares what one replica of this service needs — CPU, memory, and for model work a GPU. traffic controls request handling, most usefully timeout, in seconds, after which a request is abandoned. workers sets how many worker processes serve this service; remember that __init__ runs once per worker, so four workers means four copies of your model in memory. path_prefix puts every endpoint of the service under a prefix, which is how you version an API or fit into an existing URL scheme. There is also a metrics argument for tuning what Prometheus sees, including the histogram buckets for request duration.
The one you should think about on day one is workers, and the reason is memory arithmetic. A large model with several workers will exhaust the memory of a modest machine, and the symptom is a process being killed with no Python traceback at all. Start with one worker, measure, then raise it. Concurrency and batching — the other two throughput levers — are mid-level topics for a good reason: they interact, and tuning them without measurement makes things worse.
Two lifecycle hooks are worth knowing by name. @bentoml.on_deployment runs once before any worker starts, which is where one-time preparation such as downloading weights belongs; doing it in __init__ would repeat the work per worker. @bentoml.on_shutdown runs on the way out, for cleanup. Use the first sparingly and the second rarely, but recognise them when you read someone else's service.
A last word on secrets, because it is a mistake that is easy to make and hard to undo. The build configuration accepts an envs list, and it is tempting to put an API key there. Do not. The values are baked into the Bento, and a Bento is an artifact you export, push and share. Secrets belong in whatever your runtime provides — Kubernetes Secrets, Vault, a cloud secrets manager, or the hosted platform's secret store — injected as environment variables when the container starts.
- Add
path_prefix="/v1"to your service and restart. - Call the old URL, then the new one.
- Set
traffic={"timeout": 1}, add a deliberatetime.sleep(5)to the API, and call it.
Reading the errors
Serving errors are mostly import errors, port errors and version errors wearing different clothes. Learning the five shapes below covers nearly everything you will hit in your first month.
The command is not found. bentoml: command not found, or on Windows a message about an unrecognised command. Your virtual environment is not active, or the scripts directory is not on PATH. Activate the environment. If you need to confirm which interpreter owns the install, python -m bentoml --version bypasses PATH entirely.
The service cannot be imported. An import failure when you run bentoml serve service:MyService almost always means one of three things: the module name is wrong (it is the filename without .py), the class name is wrong or misspelled, or you are in the wrong directory so Python cannot find the module. Check in that order. Read the traceback to the bottom, too — sometimes the real failure is inside your own module, a missing dependency or a syntax error, and the service import message is just the outer wrapper.
Legacy API errors. A ModuleNotFoundError for bentoml.io, or an AttributeError saying bentoml has no attribute Service, means you are running 1.1-era code on a current install. Rewrite the file in the class-based style. The instinct to pin an old BentoML instead is understandable and wrong: you will be writing the migration eventually, and the old series does not get the security patches.
The port is taken. An address-already-in-use error on port 3000 means something else is listening — nine times out of ten a bentoml serve you started in another terminal and forgot. Stop it and start again.
The Bento or model is not found. A not-found error naming a tag means the store does not have what you asked for. Run bentoml list or bentoml models list and compare the tag carefully, including the version. This bites hardest in containers and CI, where the store is not the one on your laptop: a model saved locally is not inside a Bento unless the build included it.
Two more that are worth recognising because the cause is not in your code. GPU code that reports no device available in a container is usually a mismatch between the image's CUDA version and the host driver rather than a BentoML problem; upgrading to a current patch release is the first thing to try, since CUDA base image corrections have shipped in the 1.4 series. And failures during pip install, model download or any call to a hosted service that mention certificate verification are the TLS-interception problem from the installation section, not a BentoML bug.
The general technique matters more than the list. When a service misbehaves, the order is: read the server log in the terminal where it runs; reproduce with curl so you see the status code and body; re-run the command with --verbose; and if the artifact is involved, bentoml serve <tag> to separate "my code is broken" from "my Bento is broken". Collect bentoml env output before asking anyone for help.
- Cause each of the first four errors deliberately: deactivate the venv, misspell the class, write a
bentoml.ioimport, and start two servers. - Write the exact message you got next to each cause in your notes.
- Fix each one.
Putting it all together
One small end-to-end project, using everything above. A sentiment classifier, trained, saved, served with a typed interface, built, containerized and called from a container.
import bentoml
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
texts = [
"excellent service, very fast delivery",
"the product arrived broken and late",
"great quality for the price",
"terrible experience, would not buy again",
"friendly support and a quick refund",
"completely useless, waste of money",
]
labels = [1, 0, 1, 0, 1, 0]
model = make_pipeline(TfidfVectorizer(), LogisticRegression())
model.fit(texts, labels)
saved = bentoml.sklearn.save_model("sentiment", model)
print("saved:", saved.tag)
import bentoml
from pydantic import BaseModel
class Review(BaseModel):
text: str
class Sentiment(BaseModel):
label: str
score: float
@bentoml.service(resources={"cpu": "1"}, traffic={"timeout": 15})
class SentimentClassifier:
def __init__(self) -> None:
self.model = bentoml.sklearn.load_model("sentiment:latest")
@bentoml.api
def classify(self, review: Review) -> Sentiment:
probabilities = self.model.predict_proba([review.text])[0]
positive = float(probabilities[1])
return Sentiment(
label="positive" if positive >= 0.5 else "negative",
score=round(max(positive, 1 - positive), 4),
)
service: "service:SentimentClassifier"
description: "TF-IDF and logistic regression sentiment classifier"
labels:
owner: mlops-course
include:
- "service.py"
python:
packages:
- scikit-learn
models:
- "sentiment:latest"
Then the whole loop, in order:
python train.py
bentoml models list
bentoml serve service:SentimentClassifier
# in a second terminal:
curl -X POST http://localhost:3000/classify \
-H "Content-Type: application/json" \
-d '{"review": {"text": "the delivery was quick and the product is great"}}'
bentoml build
bentoml list
bentoml serve sentiment_classifier:latest
bentoml containerize sentiment_classifier:latest
docker run --rm -p 3000:3000 sentiment_classifier:xxxxxxxxxxxx
Walk through what each stage proved. train.py put weights in the model store with a version, so the model is now a named thing rather than a file. Serving from source proved the code works. The typed interface means a client sending the wrong shape gets a validation error, and the browser page documents the fields without you writing documentation. bentoml build froze code, pinned dependencies and the model reference into one tagged artifact — note the models entry in the build file, which is what puts the stored model inside the Bento so the container does not need your laptop's store. Serving the Bento proved the artifact is complete. Containerizing produced an image that any platform can run. Running it and getting the same answer proved the chain end to end.
Six commands, and the gap between "it works in my notebook" and "another team can deploy this" is closed. That gap is where most model projects die, so closing it reliably is a more valuable skill than it sounds.
- Run the whole project, start to finish, on your own machine.
- Add a second endpoint,
classify_batch, taking a list of reviews and returning a list of sentiments. - Rebuild, containerize, and call the new endpoint from the container.
- Then delete the model from the store and run the container again.
models entry, it would not, and understanding that difference is understanding what a Bento is.
What you can now do, and what comes next
You can write a BentoML service in the current API and recognise the legacy one on sight. You can serve it, call it from three kinds of client, and read the generated API page. You can express an interface in type hints and Pydantic models and let BentoML validate it. You can save a model into the store with a version, load it by tag, and understand why pinning matters. You can write a build file, produce a Bento, inspect it, serve it, export it and containerize it. You can set resources, timeouts, workers and a path prefix, and you know where secrets must not go. And you can diagnose the five errors that account for most first-month frustration.
That is enough to ship a real internal service. It is not yet enough to run one under load, and the honest list of what you do not yet know is short and specific.
Throughput. Adaptive batching groups concurrent requests into one model call and is the main lever for GPU utilisation. It interacts with traffic.concurrency and with workers, and tuning it without load testing makes latency worse. That is the centre of the mid-level guide.
Multi-service graphs. bentoml.depends() lets one service call another as if it were a local object, so CPU preprocessing and GPU inference can be separate services with separate resources and separate scaling. This replaces the old Runner concept.
Async, streaming and tasks. API methods can be async def; generators give streaming responses, which matters for anything token-by-token; and @bentoml.task handles long-running work through a queue the client polls.
Operations. Metrics beyond the defaults, tracing, autoscaling with minimum and maximum replicas and stabilization windows, secrets management, cost control on GPU replicas, and the question of self-hosting on Kubernetes versus the hosted platform. That is the senior guide.
Where to go next in this catalogue depends on what you are missing. If model versioning and experiment tracking are the weak link upstream of serving, read MLflow — the pairing of MLflow for tracking and BentoML for serving is a common and sensible stack. If the serving target is Kubernetes and you want to compare BentoML's packaging-first approach with a Kubernetes-native one, read KServe and Seldon Core. If you want the container fundamentals under bentoml containerize — layers, platforms, image size, running as a non-root user — read Docker, and then Kubernetes for where the image goes. If what you are serving is a language model and you care about tokens per second, vLLM covers the inference engine that BentoML would wrap. And for watching a deployed service, Prometheus and Grafana are what consume that /metrics endpoint you opened earlier.
Before you move on, do one thing: take a model that already exists in your work or your coursework and put it through the loop in the previous section. Not the iris dataset — something of yours, with a real input shape and a real dependency list. Every rough edge you hit doing that is a thing the next two guides will explain, and you will read them very differently having felt the problem first.
Then read the Interview Prep and Tips files. The first turns this into answers you can give under pressure; the second is the list of mistakes that cost other people their afternoons.