This is part one of three. It teaches you TorchServe from zero: what it is, how a PyTorch model becomes a web service, and how to run, call, scale and debug that service on your own machine. By the end you will have taken a trained image model, packaged it into a single file, started a server, sent it a photo over HTTP, watched it handle a burst of requests with batching, and read its logs when something went wrong. Mid-level and Senior take the same ground further.
Each section ends with a Try it task. Do them as you go. Serving concepts feel abstract until you have watched a worker process die and restart, and that takes only a few minutes.
What TorchServe is, and the problem it solves
You have trained a model in PyTorch. It lives in a Python file or a notebook, and it answers questions when you call it. Nobody else can use it. A mobile app, a website backend or a colleague's script cannot reach into your notebook and call model(image).
Serving means turning that model into a long-running program that other programs can talk to over the network. A client sends a request (an image, a sentence, a row of numbers), the server runs the model, and the client receives the prediction. TorchServe is a ready-made server for exactly this, built for PyTorch models by the PyTorch and AWS teams.
The diagram is the whole workflow. You will do each arrow by hand in this guide.
Before tools like this, teams wrote their own serving code. The usual version was a small Flask or FastAPI application that loaded the model once at startup and exposed a /predict route. That works, and for a prototype it is a fine choice (the FastAPI guide at /student-guides/fastapi shows the pattern). But a hand-rolled server leaves you to solve a list of problems yourself, and each one is harder than it looks:
- Concurrency. One Python process runs one thing at a time for CPU-heavy work. To use several CPU cores or several GPUs you need several processes, and you need to route requests between them.
- Batching. A GPU is far more efficient when it processes eight images together than when it processes one image eight times. Collecting requests that arrive close together into one batch, without making anyone wait too long, is fiddly to write correctly.
- Many models and versions. Real systems serve more than one model, and need to roll out version 2 while version 1 still answers traffic.
- Operations. Health checks, logs, metrics, restarting a crashed worker, loading models at startup, changing the number of workers without a restart.
TorchServe packages answers to all of these. You bring a model and a small piece of Python that describes how to turn a request into a tensor and a tensor into a response; TorchServe supplies everything around it.
The maintenance status, in plain words
"Limited maintenance" is not the same as "broken". Version 0.12.0 still installs, still runs, and still serves models. What it means is that when a bug or a security problem is found, nobody at the project will fix it, and the official security page warns that vulnerabilities may not be addressed. For a learner that is fine: you are on your own laptop, with no public traffic. For an employer that handles customer data, it is a real risk, and a manager will want to know you understand it. Being able to say "this tool is archived, here is what I would pick instead and why" in an interview is worth more than pretending the status does not exist. We will return to it in the last section and in the interview file.
- Open
https://github.com/pytorch/servein your browser. - Find the banner at the top of the page and read the date it was archived.
- Open the Releases page and find the newest version number.
The mental model: six nouns
TorchServe has a small vocabulary. Learn these six words and every error message and every documentation page becomes readable.
Model. Your trained PyTorch network: its weights (the numbers learned during training) and, depending on how you saved it, the Python class that defines its layers.
Model archive (.mar file). A single file that bundles everything TorchServe needs to serve one model: the weights, an optional file defining the model class, the handler, any extra files such as a list of class names, and a small MANIFEST.json describing the contents. You create it with a command-line tool called torch-model-archiver. Think of it as a zip file with a contract. One thing to remember from the start: the official docs say TorchServe "executes the arbitrary python code packaged in the mar file", so you must only load archives from sources you trust.
Model store. A plain folder holding .mar files. You tell the server where it is with --model-store.
Handler. The Python code that defines what happens to a request. It has four stages: initialize loads the model once, preprocess turns the raw request into model input, inference runs the model, and postprocess turns the output into a response. TorchServe ships built-in handlers for common tasks, so for an image classifier you often write no handler code at all.
Worker. A Python process that holds one copy of a model in memory and runs the handler. A model can have several workers. Workers are what give you parallelism.
Frontend and backend. The frontend is a Java server (built on Netty) that accepts HTTP and gRPC connections, queues requests, and manages workers. The backend is the Python worker processes. This split is why TorchServe needs Java installed even though you never write Java: the part that talks to the network is a Java program.
Between the frontend and the workers sits a job queue for each model. Requests wait there until a worker is free. By default the queue holds 100 requests (job_queue_size); when it is full, the next request is refused with HTTP 503 rather than being allowed to pile up forever. That refusal is deliberate and healthy: a server that says "busy" quickly is easier to build on than one that silently takes minutes to reply.
Three APIs on three ports
TorchServe exposes three separate HTTP APIs, each on its own port, and the separation is a security feature:
| API | Default port | Used for |
|---|---|---|
| Inference | 8080 | Sending data and getting predictions; the health check at /ping |
| Management | 8081 | Listing, registering, scaling and deleting models |
| Metrics | 8082 | Numbers about the server's behaviour |
There are also gRPC versions of the inference and management APIs on ports 7070 and 7071. All of them listen on localhost only by default, which means other machines cannot reach them. That default protects you; later we will see how easy it is to undo it by accident.
- Without looking back, draw the path of one request on paper: client, frontend, queue, worker, handler stages, response.
- Label which part is Java and which is Python.
- Mark which port the client would use to ask for a prediction, and which port an administrator would use to add a model.
Installing TorchServe and checking the setup
TorchServe 0.12.0 expects a specific environment, and most beginner trouble is a mismatch here. The official requirements for this release are:
- Python 3.8 to 3.11. Newer Python versions are not listed as supported, and since the project is archived they never will be.
- JDK 17. Java 17 specifically. A missing or different Java is the single most common cause of a server that will not start, and it shows up as a
java.lang.NoSuchMethodError. - PyTorch 2.4, matched to your hardware (CPU, or a CUDA GPU). CUDA 11.8 is the stable GPU target for this release.
- A supported OS: Ubuntu 20.04, macOS 10.14 or newer, Windows 10 Pro, Windows Server 2019, or WSL.
Linux and macOS
Create an isolated Python environment first, so a TorchServe install cannot disturb other projects. Use Python 3.11 or older:
python3.11 -m venv ts-env
source ts-env/bin/activate
pip install torch torchvision
pip install torchserve torch-model-archiver torch-workflow-archiver
Three packages come from that last line. torchserve is the server. torch-model-archiver makes .mar files. torch-workflow-archiver is for chaining several models together, which this guide does not use. The same three packages are available through conda if you prefer it: conda install torchserve torch-model-archiver torch-workflow-archiver -c pytorch.
Now install Java 17 using your system's package manager. On macOS with Homebrew, brew install openjdk@17; on Ubuntu, sudo apt install openjdk-17-jdk. Whichever you use, confirm the result and make sure the java that your shell finds is version 17:
java -version
torchserve --version
torch-model-archiver --version
The first command should print a line containing 17. The other two print version numbers, which should read 0.12.0 for the current release.
java -version reports 11 or 21, TorchServe may fail even though JDK 17 is installed somewhere. Set JAVA_HOME to the 17 installation and put its bin directory first on your PATH, then open a fresh terminal and check again.
Windows
The official Windows route has more steps: conda packages are not supported on Windows, and the instructions need administrator rights, Git, OpenJDK 17, Node.js and the Visual C++ Redistributable, and then a clone of the repository and its install_dependencies.py script. The 0.12.0 release notes list WSL as a supported platform, and for a student the least painful path is to install WSL2 with Ubuntu and follow the Linux instructions inside it. This is practical advice rather than an official recommendation, but it avoids a long list of Windows-specific setup steps.
Installing with Docker instead
If installing Java and matching versions sounds like a chore, Docker (see /student-guides/docker) removes it. The project publishes images pytorch/torchserve (CPU) and a GPU variant. Because the project is archived, the latest tag will never change again, but it is still better practice to pin a specific tag so that what you run today is what you run next year. Check the tag list on Docker Hub before choosing one. A later section shows how to run it.
- Create a virtual environment and install the three packages above.
- Run
java -version,torchserve --versionandtorch-model-archiver --version. - If the Java version is not 17, fix it before moving on.
Preparing a model to serve
TorchServe does not train anything. It needs a saved model, and PyTorch gives you two common ways to save one. The difference decides which archiver flags you use, so it is worth understanding.
Eager mode (a state dictionary). You save only the learned numbers with torch.save(model.state_dict(), "model.pth"). At serving time you need the Python class that defines the network, so you must also ship a file containing that class. This is flexible and common, and it is what the archiver calls an eager-mode model.
TorchScript. You convert the model into a self-contained, serialisable form with torch.jit.trace or torch.jit.script. The saved file includes the computation itself, so no separate model-definition file is needed. It is simpler to package.
For a first project we will use the TorchScript route with a pretrained ResNet-18 from torchvision, because it needs no custom code and the built-in image_classifier handler fits it. Create a working folder and a script:
import torch
from torchvision import models
# Load a ResNet-18 pretrained on ImageNet and switch to inference mode.
weights = models.ResNet18_Weights.DEFAULT
model = models.resnet18(weights=weights)
model.eval()
# Trace the model with a dummy 224x224 RGB image to produce TorchScript.
example = torch.rand(1, 3, 224, 224)
traced = torch.jit.trace(model, example)
traced.save("resnet18.pt")
print("saved resnet18.pt")
Run it with python export_model.py. The first run downloads the pretrained weights (roughly 45 MB), then writes resnet18.pt. Two details matter. model.eval() turns off training behaviours such as dropout and batch-norm updates; forgetting it gives subtly wrong predictions. And torch.jit.trace records what the model does for one example input, so the dummy input must have the same shape the real images will have after preprocessing.
You also need a test image. Any small JPEG will do; a photo of a cat or a dog works well because ImageNet has many categories for them. Save it as kitten.jpg in the same folder.
Class names, so the output is readable
A classifier outputs a number for each of 1,000 ImageNet classes, and the useful answer is "tabby cat", not "281". The built-in image_classifier handler can translate indices into names if you supply a JSON file called index_to_name.json through the archiver's --extra-files option. The torchvision package ships the names inside the weights' metadata, so you can write the file yourself:
import json
from torchvision import models
names = models.ResNet18_Weights.DEFAULT.meta["categories"]
mapping = {str(i): name for i, name in enumerate(names)}
with open("index_to_name.json", "w") as f:
json.dump(mapping, f)
print(len(mapping), "classes written")
This prints 1000 classes written. The file maps the string "0" to the first class name and so on. If you skip this file, the handler still works; the response will simply identify classes by index rather than by name, so keep the file when you want readable answers.
- Run
export_model.pyandmake_labels.pyin an empty folder. - Put a JPEG named
kitten.jpgbeside them. - Run
ls -lhand look at the sizes.
resnet18.pt of roughly 45 MB and a small index_to_name.json. These are the only ingredients you need for the next step.
Packaging the model with torch-model-archiver
The archiver is the tool that produces the .mar file. Its job is to bring the weights, the handler and the extras together under one name and version, so that the server can load the result without knowing anything else about your project.
For our TorchScript model the command is:
torch-model-archiver \
--model-name resnet18 \
--version 1.0 \
--serialized-file resnet18.pt \
--handler image_classifier \
--extra-files index_to_name.json \
--export-path model_store \
--force
Read it flag by flag, because these are the flags you will use every time:
--model-nameis the name clients will use in the URL. It becomes/predictions/resnet18. Choose something short, without spaces.--versionlabels this build of the model. TorchServe supports several versions of the same name side by side, and you will use that to roll out updates safely.--serialized-filepoints to the weights file: the.ptTorchScript file here, or a.pthstate dictionary for eager mode.--handlernames the handler.image_classifieris one of the built-in ones; you can instead give the path to your own Python file.--extra-fileslists additional files, separated by commas, to include in the archive.--export-pathis the folder where the.maris written. We point it atmodel_store, which will be our model store.--force(short form-f) overwrites an existing archive of the same name. Without it, re-running the command fails when the file already exists.
If a folder named model_store does not exist yet, create it first with mkdir model_store. After success you will find model_store/resnet18.mar.
The eager-mode variant
If you saved a state dictionary instead of TorchScript, you must additionally provide the file that defines the model class, with --model-file, and usually a --requirements-file for extra Python packages the handler needs:
torch-model-archiver \
--model-name mynet --version 1.0 \
--model-file model.py \
--serialized-file mynet_weights.pth \
--handler image_classifier \
--requirements-file requirements.txt \
--export-path model_store
The rule: --model-file is mandatory in eager mode and unnecessary for TorchScript. If you forget it in eager mode, the worker will fail when it tries to build the model, and the error appears in the logs rather than at archive time, which is a classic way to lose twenty minutes.
Other options exist: --runtime to choose the Python runtime, --config-file to attach a YAML configuration, and --archive-format to choose between the default .mar, a tgz file, a zip without compression, or a plain folder (no-archive). You can leave all of those at their defaults for now.
What is inside
A .mar is a zip-like file, so you can inspect it. Run unzip -l model_store/resnet18.mar and you will see your weights, the labels file and a MANIFEST.json inside a MAR-INF folder. The manifest records the model name, version and handler. Looking inside once removes the mystery: the archive is just your files plus a description.
- Create
model_storeand run the archiver command above. - Run
unzip -l model_store/resnet18.mar. - Open the manifest inside and find the handler name.
index_to_name.json and a manifest, and a manifest that records the handler. Nothing more is hidden in there.
Starting the server and sending your first prediction
Now the payoff. With model_store/resnet18.mar in place, start TorchServe:
torchserve --start --ncs --model-store model_store --models resnet18=resnet18.mar --disable-token-auth
Each flag has a reason:
--startlaunches the server in the background and returns your prompt. The server keeps running after the command ends.--ncs(long form--no-config-snapshots) turns off TorchServe's habit of saving its state to snapshot files. We will explain snapshots shortly; for learning, switching them off avoids surprising behaviour after a restart.--model-storeis the folder of.marfiles. It is the one option that is mandatory.--modelslists which models to load at startup, asname=file.mar. The name before the equals sign is the name clients use in the URL. You can list several, separated by spaces.--disable-token-authturns off token authentication. More on this in a moment; it is for local learning only.
After a few seconds, check that the server is alive:
curl http://localhost:8080/ping
A healthy server answers:
{
"status": "Healthy"
}
/ping returns success when the number of active workers is at least the model's minimum. If you call it too quickly after starting, or a worker failed to load, you get an error status instead. That makes it useful as a readiness check later in Kubernetes (/student-guides/kubernetes).
Now ask for a prediction. Send the image file as the request body with curl's -T option:
curl http://localhost:8080/predictions/resnet18 -T kitten.jpg
The reply is a JSON object mapping class names to probabilities, best guess first. With the labels file included, the keys are readable names such as a cat breed; the values are numbers between 0 and 1 that sum to about 1 across the top few classes. The built-in handler returns the five most likely classes. You have just served a PyTorch model over HTTP.
Stop the server when you are done:
torchserve --stop
Why you pass --disable-token-auth
Since version 0.11.1, TorchServe enforces token authorization by default. When the server starts it writes a file called key_file.json into the directory you launched it from, containing an inference key, a management key and an API key used to mint new tokens. Every request must then carry a header of the form Authorization: Bearer <key>. Tokens expire after 60 minutes by default.
This is a good default, and in production you should keep it on. But most tutorials on the internet were written for older versions and omit the header, so beginners follow a tutorial, get a 401-style refusal and conclude TorchServe is broken. There are two honest ways forward. For local learning, start with --disable-token-auth, as above. Or keep authentication on and send the key:
curl http://127.0.0.1:8080/predictions/resnet18 \
-H "Authorization: Bearer <INFERENCE_KEY>" -T kitten.jpg
where the inference key is the value found in key_file.json. Never commit that file to Git and never copy it into a container image: it contains live credentials.
curl -X POST "localhost:8081/models?url=..." as the first step. On current versions that is refused unless the model API is explicitly enabled and a management token is supplied. The simple, reliable beginner approach is to pre-load models with --models at startup, exactly as we did.
Snapshots, briefly
TorchServe can record which models and how many workers were running into snapshot files, so that after a restart it comes back in the same state. That is helpful for servers, and confusing for learners, because an old snapshot can restore a state you forgot about, or fail to load and produce an InvalidSnapshotException. Using --ncs while learning sidesteps both.
- Start the server with the command above, wait a few seconds, then call
/ping. - Send
kitten.jpgto/predictions/resnet18. - Send a different image, ideally something that is not an animal, and read the result.
- Run
torchserve --stop.
Handlers: how a request becomes a prediction
So far we used a built-in handler and wrote no code. To understand what happened, and to serve any model that is not a plain image classifier, you need to know what a handler does.
A request arrives at the frontend, which passes the raw bytes to a worker. The worker's handler runs a fixed pipeline of four stages:
initialize(context)runs once, when the worker starts. It loads the model weights and moves them to the right device (CPU or a specific GPU). Anything expensive belongs here, because it must not be repeated for every request.preprocess(data)runs for each request (or each batch). It converts the raw input, such as JPEG bytes, into a tensor in the shape the model expects.inference(data)runs the model on that tensor and returns the raw output.postprocess(output)turns the raw output into something that can be returned to the client, usually a list with one entry per request.
A function called handle(data, context) orchestrates the four. Most handlers do not implement all of it themselves; they inherit from BaseHandler, which provides sensible defaults for model loading and inference, and override only the stages they need to change.
The built-in handlers
TorchServe includes ready-made handlers for common jobs, selected by name in the --handler flag:
| Handler | Purpose |
|---|---|
image_classifier |
Classify an image into labels |
image_segmenter |
Label each pixel of an image |
object_detector |
Find and box objects in an image |
text_classifier |
Classify a piece of text |
Each lives in the ts.torch_handler package, so a custom handler can reuse one with an import such as from ts.torch_handler.image_classifier import ImageClassifier. Note that the official 0.12.0 release notes deprecate TorchText support, so treat NLP handlers that depend on it with caution.
Writing a small custom handler
Suppose your model is a text or tabular model, or an image model with unusual preprocessing. You write a Python file with a class that inherits from BaseHandler. Here is a handler for a model that takes a list of numbers and returns one score:
import json
import torch
from ts.torch_handler.base_handler import BaseHandler
class NumericHandler(BaseHandler):
"""Serve a model that takes a JSON list of floats and returns a score."""
def preprocess(self, data):
rows = []
for item in data:
payload = item.get("data") or item.get("body")
if isinstance(payload, (bytes, bytearray)):
payload = payload.decode("utf-8")
rows.append(json.loads(payload) if isinstance(payload, str) else payload)
return torch.tensor(rows, dtype=torch.float32)
def postprocess(self, output):
# Must return a list with exactly one element per request in the batch.
return output.squeeze(-1).tolist()
Several things in this small file are worth stopping on:
- The handler is a class that inherits
BaseHandler. If a file contains several classes, the entry class has to be the first one. datais always a list, one element per request in the current batch. When batching is off, the list has one element, but your code should still loop over it. Writing handlers as if there were always exactly one request is the root of many bugs once batching is switched on.- The request content arrives under a key such as
dataorbody, depending on how it was sent, which is why the code checks both. The raw bytes may need decoding. postprocessmust return a list the same length as the input list. If it returns a bare number or a list of the wrong length, the responses are matched to the wrong clients or the request fails.
You package it by passing the file path as the handler:
torch-model-archiver --model-name scorer --version 1.0 \
--model-file model.py --serialized-file scorer.pth \
--handler numeric_handler.py --export-path model_store
When you write code that will run inside a worker, remember that every print goes into a log file rather than your terminal, and an exception during initialize kills the worker. The log section below shows where to look.
- Find the file for the built-in
image_classifierhandler in your virtual environment'ssite-packages/ts/torch_handlerfolder. - Read it and locate which of the four stages it overrides.
- Compare it with
base_handler.pyin the same folder.
Managing a running server
Once a server is running you will want to look at it: which models are loaded, how many workers each has, whether they are healthy. The management API on port 8081 answers those questions. Because it can change the server, it is protected more strictly than the inference API, and on current versions the same token rules apply (with a management key rather than the inference key). If you started with --disable-token-auth, you can call it without headers.
List the models:
curl http://localhost:8081/models
Describe one model:
curl http://localhost:8081/models/resnet18
The description includes the model name and version, the minimum and maximum worker counts, the batch size and batch delay, a list of workers with their status, GPU use and memory, and a jobQueueStatus showing how many requests are waiting and how much capacity remains. When someone asks "is the model keeping up?", the queue status is the first place to look: a queue that is regularly full means requests arrive faster than workers can handle them.
Workers: scaling a single model
A worker is one process holding one copy of the model. By default the number of workers per model follows your hardware (the number of GPUs, or logical CPUs). You can change the number for a specific model at runtime:
curl -X PUT "http://localhost:8081/models/resnet18?min_worker=2&synchronous=true"
min_worker=2 asks for at least two workers. With synchronous=true the call waits until the workers are ready and answers with status 200; without it, the call returns 202 immediately and the workers start up in the background. Each worker loads its own copy of the model, so two workers use roughly twice the memory. On a CPU-only laptop, more workers than cores brings no benefit and can slow everything down, because they compete for the same cores.
You can see the effect by running the describe command again and counting the entries under workers. Scaling a model is the cheapest experiment in this guide and the one that teaches most about how serving capacity works.
Versions
TorchServe supports multiple versions of one model name. Archive the model again with --version 2.0, and the server can hold both. Requests to /predictions/resnet18 go to the default version, requests to /predictions/resnet18/2.0 go to that specific version, and an administrator can switch the default with PUT /models/resnet18/2.0/set-default. This lets you deploy a new model next to the old one, test it with a version-specific URL, and only then promote it. The principle is the same as blue-green deployments elsewhere in operations.
0.0.0.0 to make it reachable, you have given strangers a way to run code on your machine. Keep management on a private network and keep token authorization enabled.
- Start the server again and describe the
resnet18model; count the workers. - Scale it to two workers with
synchronous=true, then describe it again. - Watch memory use in your system monitor before and after.
workers list and a visible rise in memory. That is the cost of parallelism: each worker owns a full copy of the model.
Dynamic batching: trading a little latency for throughput
A GPU, and to a lesser degree a CPU, handles many inputs at once much more efficiently than one at a time. If a thousand clients each send one image, running a thousand separate forward passes wastes most of the hardware. Dynamic batching groups requests that arrive close together into one forward pass.
TorchServe controls it with two numbers per model:
batchSize, the maximum number of requests to combine into one batch.maxBatchDelay, the longest time in milliseconds the server will wait for a batch to fill before running with whatever it has.
The rule is: the worker collects requests until it has batchSize of them, or until maxBatchDelay has passed since the first one, whichever comes first. A quiet period therefore costs each request at most the delay, while a busy period fills batches quickly and gets the efficiency gain.
Choosing values is a tuning exercise. A large batchSize with a long delay maximises throughput and hurts the latency of each individual request. For interactive applications, a short delay such as 10 to 50 ms keeps responses snappy; for offline scoring, a long delay and a large batch are fine. There is no universal right answer, which is why you measure with a load test rather than guessing.
You set these values when registering a model. With the model API disabled by default, the straightforward beginner route is the configuration file, shown in the next section. The registration call's parameters are named batch_size and max_batch_delay, for example POST /models?url=resnet-152.mar&batch_size=8&max_batch_delay=50, if you enable the model API in your setup.
Your handler must cooperate
Batching only works if the handler expects a list. As described earlier, the data argument is a list with one entry per request, and postprocess must return a list of the same length. The built-in handlers already do this. A custom handler that assumes a single request will fail or mix up answers the moment batching is turned on. Write for the list from the beginning.
- Write a small Python loop that sends 50 requests to the server one after another and prints the total time.
- Repeat it using a thread pool of 8 to send requests at once.
- Compare the two totals and look at the queue status in the describe output while the second run is going.
Configuration: config.properties and command-line options
Typing long commands every time is error-prone, and settings that are not in a file cannot be reviewed or shared. TorchServe reads settings from a config.properties file, which you point at with --ts-config:
torchserve --start --ncs --ts-config config.properties
A beginner-friendly file looks like this:
# Where the three APIs listen. These are the defaults; keep them local.
inference_address=http://127.0.0.1:8080
management_address=http://127.0.0.1:8081
metrics_address=http://127.0.0.1:8082
# Where model archives live and which to load at startup.
model_store=model_store
load_models=resnet18.mar
# Capacity: how many requests may wait for a worker.
job_queue_size=100
# Larger uploads than the default need this raised.
max_request_size=20000000
max_response_size=20000000
Each line is a key=value pair; lines starting with # are comments. The settings above are the ones a beginner meets most:
inference_address,management_addressandmetrics_addressset where the three APIs listen. The defaults bind to127.0.0.1, meaning "only this machine".model_storeandload_modelsreplace the--model-storeand--modelsflags.load_modelsaccepts a list of.marfiles, orallto load everything in the store.default_workers_per_modelsets how many workers each model gets; if you leave it out, TorchServe uses the number of GPUs, or the number of logical CPUs on a CPU-only machine.job_queue_sizeis the per-model queue length (default 100).max_request_sizeandmax_response_sizelimit message sizes in bytes. The Troubleshooting page notes that uploads of more than about 6.5 MB fail by default, and raising these values is the fix.install_py_dep_per_model=truetells TorchServe to install the packages from a model'srequirements.txtwhen it loads the model. Without this setting, the file you passed with--requirements-fileis not used and imports fail inside the worker.allowed_urlsis a list of patterns (regular expressions) that restricts where models may be downloaded from. The default permitsfile://andhttp(s)://sources; in a real deployment you narrow it to your own model registry.
The configuration page also describes how to define models, with their versions, worker counts and batch settings, inside the file using a models= entry in JSON form. That is the usual way to set batchSize and maxBatchDelay without the management API. The exact structure is documented on the configuration page of the official docs. Read it before copying an example, because a small JSON mistake makes the file fail to parse.
Environment variables
With enable_envvars_config=true in the file, you can override properties with environment variables named TS_ followed by the property in capitals, for example TS_INFERENCE_ADDRESS. This is how container deployments change settings without editing files. Be careful when mixing the three places (file, environment, command line): the docs describe a precedence order, but different pages word it differently, so avoid depending on subtle overrides. Pick one mechanism for each setting and stay with it.
config.properties contains no secrets unless you add TLS keystore passwords to it, so it belongs in version control next to your archiving script. A teammate can reproduce your server from the repository, which was the whole point of serving the model this way.
- Stop the server, write the
config.propertiesabove and start TorchServe with--ts-config. - Call
/pingand a prediction to confirm that it works. - Change
inference_addressto port 8090, restart, and confirm that port 8080 no longer answers.
Running TorchServe in Docker
Containers solve the "which Java, which Python, which PyTorch" problem by shipping them together. The project publishes the images pytorch/torchserve for CPU and a -gpu variant. To start the CPU image, publish the ports, mount your model store and tell it which model to load:
docker run --rm -it \
-p 127.0.0.1:8080:8080 -p 127.0.0.1:8081:8081 -p 127.0.0.1:8082:8082 \
--mount type=bind,source="$(pwd)/model_store",target=/tmp/models \
pytorch/torchserve:latest \
torchserve --model-store=/tmp/models --models resnet18=resnet18.mar --disable-token-auth
The details matter:
- The
-p 127.0.0.1:8080:8080form publishes the container's port 8080 on your loopback interface only. Writing-p 8080:8080instead publishes it on all interfaces, which can expose the port to your whole network. The project's own Docker instructions bind to127.0.0.1for this reason. --mount type=bindmakes yourmodel_storefolder appear inside the container at/tmp/models.- The command at the end is the same
torchservecommand you already know. The container's job is just to run it in a controlled environment. In a container the server runs in the foreground, so you do not use--start.
For a more production-like run, the official Docker instructions add --shm-size=1g --ulimit memlock=-1 --ulimit stack=67108864, which give the worker processes enough shared memory. A common surprise for beginners is that PyTorch workers need more shared memory than the container default, and the symptom is a worker that crashes while handling a request.
latest will not move again, but pinning is still the right habit: choose an explicit version tag from the Docker Hub tag list, write it in your notes, and use it everywhere. An unpinned tag is a surprise waiting for the day somebody rebuilds the image on a different machine.
- Run the container with your
model_storemounted. - From another terminal, call
/pingand sendkitten.jpg. - Press Ctrl+C in the container terminal and confirm the port stops answering.
Logs, metrics and reading errors
When a model server misbehaves, the answer is almost always in its logs. TorchServe writes them to a logs folder in the directory where you started it. The two to know first are ts_log.log, which has the server's own messages including worker start-up and crashes, and model_log.log, which carries output from your handler code, including anything you print. There is also an access log recording each request. If you are unsure of exact names in your installation, run ls logs after starting the server.
A good habit: when something fails, run tail -n 100 logs/ts_log.log before changing any code. The log usually tells you the specific exception, and reading it first is quicker than guessing.
Metrics
The metrics API on port 8082 and the log files report what the server is doing: requests by status class, latency, queue time, CPU and memory use, and, on a GPU machine, GPU utilisation. By default metrics are written to log files (ts_metrics.log and model_metrics.log). The docs also describe a mode that exposes them in Prometheus format, which pairs with Prometheus and Grafana (/student-guides/prometheus, /student-guides/grafana); enabling it is a configuration choice for the mid-level guide.
The errors you will actually meet
These come from the official Troubleshooting page and from the behaviour described above. Learn to recognise them.
| What you see | What it means | What to do |
|---|---|---|
Failed to bind to address: http://127.0.0.1:8080 |
Something else already uses that port, often an old TorchServe | Run torchserve --stop, find the process, or change the port in config.properties |
java.lang.NoSuchMethodError at startup |
The wrong Java version | Install and select JDK 17 |
HTTP 404, ModelNotFoundException |
The model name in the URL does not match, or the .mar is not in the store |
Check GET /models and the file name in model_store |
HTTP 503, ServiceUnavailableException |
No worker is ready, or the queue is full | Check worker status with the describe call, scale workers, or raise job_queue_size |
Backend worker monitoring thread interrupted |
A worker died while loading, often an import error or missing file in the handler | Read ts_log.log and model_log.log for the Python traceback |
HTTP 409, ConflictStatusException |
That model version is already registered | Use a different version or name |
HTTP 400, DownloadModelException |
The model URL cannot be reached or is not in allowed_urls |
Fix the path or the allow-list |
| Large uploads fail | Request-size limit around 6.5 MB | Raise max_request_size and max_response_size |
InvalidSnapshotException after a restart |
A stale or broken snapshot | Start with --ncs, or delete the snapshot files |
| Unauthorized errors on every call | Token authorization is on and the request has no valid token | Send the Authorization: Bearer header with a key from key_file.json, or use --disable-token-auth locally. Tokens expire after 60 minutes |
- Deliberately break something: archive a model with a handler name that does not exist, or remove the labels file and archive again.
- Start the server and call the describe endpoint.
- Read
logs/ts_log.logand find the line that explains the failure.
Safe defaults and the security you need from day one
Serving a model means running a network service, and a network service is a target. The good news is that TorchServe's defaults are cautious. The habits below keep them that way.
Keep the ports local. All the APIs listen on 127.0.0.1 by default. In Docker, publish ports with the 127.0.0.1: prefix. If you must reach the server from other machines, put a reverse proxy or load balancer in front for the inference port and keep management and gRPC management ports (8081 and 7071) private.
Keep token authorization on outside your laptop. --disable-token-auth is a convenience for learning. Any shared or deployed environment should use the keys from key_file.json, with the inference key for clients and the management key only for administrators.
Trust only archives you made or verified. A .mar file contains Python code that the worker executes. Downloading a model archive from an unknown website and loading it is equivalent to running an unknown script. Restrict allowed_urls to your own storage.
Do not commit key_file.json. Add it to .gitignore the moment you start the server in a project folder.
Know that the project will not patch itself. Because TorchServe is no longer maintained, any vulnerability found in it or in its Java components stays unfixed in this project. If you ever run it where it matters, isolate it: a private network, a non-root user, nothing sensitive reachable from it. In the Gulf and Egypt, where many employers have data-residency rules and security reviews, expect a reviewer to ask exactly this question, and to prefer an actively maintained server.
- Start the server without
--disable-token-authand call the prediction URL without a header. - Open
key_file.json, copy the inference key, and call again with the header. - Add
key_file.jsonto your.gitignore.
Putting it all together
Here is a small end-to-end project that uses everything above. The goal: a reproducible folder that anyone can clone, run one script, and get an image classifier on port 8080.
image-service/
export_model.py
make_labels.py
build.sh
config.properties
.gitignore
README.md
The build.sh script performs the packaging steps in order, so that the .mar is always built the same way:
#!/usr/bin/env bash
set -euo pipefail
python export_model.py # writes resnet18.pt
python make_labels.py # writes index_to_name.json
mkdir -p model_store
torch-model-archiver \
--model-name resnet18 --version 1.0 \
--serialized-file resnet18.pt \
--handler image_classifier \
--extra-files index_to_name.json \
--export-path model_store --force
echo "built model_store/resnet18.mar"
And a run.sh that starts the server, waits for it to be healthy, and sends a test image:
#!/usr/bin/env bash
set -euo pipefail
torchserve --start --ncs --ts-config config.properties --disable-token-auth
# Wait up to 30 seconds for the model's workers to become ready.
for i in $(seq 1 30); do
if curl -fs http://127.0.0.1:8080/ping > /dev/null; then break; fi
sleep 1
done
curl http://127.0.0.1:8080/predictions/resnet18 -T kitten.jpg
Add key_file.json, logs/, model_store/ and *.pt to .gitignore, because they are generated artifacts, not source. Your README should state the pinned versions: Python 3.11, JDK 17, PyTorch 2.4, TorchServe 0.12.0.
Work through this checklist as the final exercise:
- Run
build.shin a clean folder and confirm the.marappears. - Run
run.shand confirm the prediction. - Describe the model, scale to two workers, and confirm the change.
- Look into the logs and find the line recording the model's load.
- Stop the server and run the same project in Docker with the bind mount.
- Change the model name to
resnet18-v2by rebuilding with--version 2.0, and think about how you would roll it out beside version 1.
If all six work, you have gone through the full lifecycle of a served model: export, package, serve, call, scale, observe, containerise and version.
- Delete the generated files and rebuild everything from the scripts alone.
- Give the folder to a friend, or to a second machine, and ask them to run it with no help.
What you can now do, and what comes next
You can now take a PyTorch model and make it a service. You know the vocabulary (archive, model store, handler, worker, queue, frontend, backend), the three ports and what each is for, how to archive a model, how to start the server with the right flags for the current version, how to call it, how to scale workers and read the describe output, how batching trades latency for throughput, how a configuration file replaces command-line flags, how to run it in Docker without exposing it, and how to read a log when a worker fails.
That is the whole beginner toolkit, and it is enough to understand most conversations about model serving.
What comes next in this series
The mid-level guide goes into custom handlers in depth, batching and worker tuning, metrics in Prometheus mode, and deployment on Kubernetes. The senior guide covers failure modes, security hardening, and what to do about the project's maintenance status in a real organisation.
Where to look if you need an actively maintained server
Because TorchServe has stopped receiving updates, a team starting a new project today would normally weigh these alternatives:
- Triton Inference Server (/student-guides/triton) serves models from several frameworks, with strong batching and GPU support.
- BentoML (/student-guides/bentoml) packages models as Python services with a friendlier developer experience.
- Ray Serve (/student-guides/ray-serve) scales Python serving across a cluster.
- KServe (/student-guides/kserve) is a Kubernetes-native layer for model serving; TorchServe can be one of its runtimes.
- vLLM (/student-guides/vllm) is the usual choice for serving large language models.
The concepts you learned here (an archive of model and code, a handler with a lifecycle, workers, queues, batching, health checks) transfer directly to all of them. Learning one serving system thoroughly makes the next one a matter of reading its documentation.
Sources
- TorchServe documentation home (redirected from https://pytorch.org/serve/)
- Getting started
- Serving models and the command line
- Configuration
- Management API
- Inference API
- Token authorization
- Metrics
- Security
- Custom service (handlers)
- Batch inference
- Troubleshooting
- TorchServe on GitHub (archived)
- Release notes
- Docker instructions
- Model archiver