Skip to content
Back to student guides
TorchServeMLOpsModel serving3 levels91 sectionsCovers TorchServe 0.12

The Complete TorchServe Guide

Serve PyTorch models in production with TorchServe. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
16sections
23examples

This is part one of three. It teaches you TorchServe from zero: what it is, how a PyTorch model becomes a web service, and how to run, call, scale and debug that service on your own machine. By the end you will have taken a trained image model, packaged it into a single file, started a server, sent it a photo over HTTP, watched it handle a burst of requests with batching, and read its logs when something went wrong. Mid-level and Senior take the same ground further.

Each section ends with a Try it task. Do them as you go. Serving concepts feel abstract until you have watched a worker process die and restart, and that takes only a few minutes.

Read this before you invest a month in TorchServe TorchServe is in limited maintenance. The official documentation says the project is no longer actively maintained, with no planned updates, bug fixes, new features or security patches. The GitHub repository was archived and made read-only on 7 August 2025. The last release is 0.12.0, from 30 September 2024. Nothing in this guide is invented for a newer version, because there is no newer version. Learn it for the ideas (model archives, handlers, workers, dynamic batching) because they appear in almost every serving system, and think hard before choosing it for a brand-new production system. The closing sections name the alternatives.

What TorchServe is, and the problem it solves

You have trained a model in PyTorch. It lives in a Python file or a notebook, and it answers questions when you call it. Nobody else can use it. A mobile app, a website backend or a colleague's script cannot reach into your notebook and call model(image).

Serving means turning that model into a long-running program that other programs can talk to over the network. A client sends a request (an image, a sentence, a row of numbers), the server runs the model, and the client receives the prediction. TorchServe is a ready-made server for exactly this, built for PyTorch models by the PyTorch and AWS teams.

YOUR MODELweights + code
→
.MAR FILEone archive
→
TORCHSERVEHTTP / gRPC
→
CLIENTSapps, scripts, curl

The diagram is the whole workflow. You will do each arrow by hand in this guide.

Before tools like this, teams wrote their own serving code. The usual version was a small Flask or FastAPI application that loaded the model once at startup and exposed a /predict route. That works, and for a prototype it is a fine choice (the FastAPI guide at /student-guides/fastapi shows the pattern). But a hand-rolled server leaves you to solve a list of problems yourself, and each one is harder than it looks:

  • Concurrency. One Python process runs one thing at a time for CPU-heavy work. To use several CPU cores or several GPUs you need several processes, and you need to route requests between them.
  • Batching. A GPU is far more efficient when it processes eight images together than when it processes one image eight times. Collecting requests that arrive close together into one batch, without making anyone wait too long, is fiddly to write correctly.
  • Many models and versions. Real systems serve more than one model, and need to roll out version 2 while version 1 still answers traffic.
  • Operations. Health checks, logs, metrics, restarting a crashed worker, loading models at startup, changing the number of workers without a restart.

TorchServe packages answers to all of these. You bring a model and a small piece of Python that describes how to turn a request into a tensor and a tensor into a response; TorchServe supplies everything around it.

The maintenance status, in plain words

"Limited maintenance" is not the same as "broken". Version 0.12.0 still installs, still runs, and still serves models. What it means is that when a bug or a security problem is found, nobody at the project will fix it, and the official security page warns that vulnerabilities may not be addressed. For a learner that is fine: you are on your own laptop, with no public traffic. For an employer that handles customer data, it is a real risk, and a manager will want to know you understand it. Being able to say "this tool is archived, here is what I would pick instead and why" in an interview is worth more than pretending the status does not exist. We will return to it in the last section and in the interview file.

Try it
  1. Open https://github.com/pytorch/serve in your browser.
  2. Find the banner at the top of the page and read the date it was archived.
  3. Open the Releases page and find the newest version number.
an archived-repository banner dated 7 August 2025, and v0.12.0 as the newest release. You now know the exact state of the tool you are about to learn.

The mental model: six nouns

TorchServe has a small vocabulary. Learn these six words and every error message and every documentation page becomes readable.

Model. Your trained PyTorch network: its weights (the numbers learned during training) and, depending on how you saved it, the Python class that defines its layers.

Model archive (.mar file). A single file that bundles everything TorchServe needs to serve one model: the weights, an optional file defining the model class, the handler, any extra files such as a list of class names, and a small MANIFEST.json describing the contents. You create it with a command-line tool called torch-model-archiver. Think of it as a zip file with a contract. One thing to remember from the start: the official docs say TorchServe "executes the arbitrary python code packaged in the mar file", so you must only load archives from sources you trust.

Model store. A plain folder holding .mar files. You tell the server where it is with --model-store.

Handler. The Python code that defines what happens to a request. It has four stages: initialize loads the model once, preprocess turns the raw request into model input, inference runs the model, and postprocess turns the output into a response. TorchServe ships built-in handlers for common tasks, so for an image classifier you often write no handler code at all.

Worker. A Python process that holds one copy of a model in memory and runs the handler. A model can have several workers. Workers are what give you parallelism.

Frontend and backend. The frontend is a Java server (built on Netty) that accepts HTTP and gRPC connections, queues requests, and manages workers. The backend is the Python worker processes. This split is why TorchServe needs Java installed even though you never write Java: the part that talks to the network is a Java program.

CLIENTcurl, app
→
FRONTENDJava: HTTP, queue
→
WORKERSPython: handler + model
→
RESPONSE

Between the frontend and the workers sits a job queue for each model. Requests wait there until a worker is free. By default the queue holds 100 requests (job_queue_size); when it is full, the next request is refused with HTTP 503 rather than being allowed to pile up forever. That refusal is deliberate and healthy: a server that says "busy" quickly is easier to build on than one that silently takes minutes to reply.

Three APIs on three ports

TorchServe exposes three separate HTTP APIs, each on its own port, and the separation is a security feature:

API Default port Used for
Inference 8080 Sending data and getting predictions; the health check at /ping
Management 8081 Listing, registering, scaling and deleting models
Metrics 8082 Numbers about the server's behaviour

There are also gRPC versions of the inference and management APIs on ports 7070 and 7071. All of them listen on localhost only by default, which means other machines cannot reach them. That default protects you; later we will see how easy it is to undo it by accident.

Try it
  1. Without looking back, draw the path of one request on paper: client, frontend, queue, worker, handler stages, response.
  2. Label which part is Java and which is Python.
  3. Mark which port the client would use to ask for a prediction, and which port an administrator would use to add a model.
a drawing with the queue between frontend and worker, prediction on 8080, administration on 8081. If you can draw it, you can debug most beginner problems.

Installing TorchServe and checking the setup

TorchServe 0.12.0 expects a specific environment, and most beginner trouble is a mismatch here. The official requirements for this release are:

  • Python 3.8 to 3.11. Newer Python versions are not listed as supported, and since the project is archived they never will be.
  • JDK 17. Java 17 specifically. A missing or different Java is the single most common cause of a server that will not start, and it shows up as a java.lang.NoSuchMethodError.
  • PyTorch 2.4, matched to your hardware (CPU, or a CUDA GPU). CUDA 11.8 is the stable GPU target for this release.
  • A supported OS: Ubuntu 20.04, macOS 10.14 or newer, Windows 10 Pro, Windows Server 2019, or WSL.

Linux and macOS

Create an isolated Python environment first, so a TorchServe install cannot disturb other projects. Use Python 3.11 or older:

BASH
python3.11 -m venv ts-env
source ts-env/bin/activate
pip install torch torchvision
pip install torchserve torch-model-archiver torch-workflow-archiver

Three packages come from that last line. torchserve is the server. torch-model-archiver makes .mar files. torch-workflow-archiver is for chaining several models together, which this guide does not use. The same three packages are available through conda if you prefer it: conda install torchserve torch-model-archiver torch-workflow-archiver -c pytorch.

Now install Java 17 using your system's package manager. On macOS with Homebrew, brew install openjdk@17; on Ubuntu, sudo apt install openjdk-17-jdk. Whichever you use, confirm the result and make sure the java that your shell finds is version 17:

BASH
java -version
torchserve --version
torch-model-archiver --version

The first command should print a line containing 17. The other two print version numbers, which should read 0.12.0 for the current release.

Several Javas on one machine If java -version reports 11 or 21, TorchServe may fail even though JDK 17 is installed somewhere. Set JAVA_HOME to the 17 installation and put its bin directory first on your PATH, then open a fresh terminal and check again.

Windows

The official Windows route has more steps: conda packages are not supported on Windows, and the instructions need administrator rights, Git, OpenJDK 17, Node.js and the Visual C++ Redistributable, and then a clone of the repository and its install_dependencies.py script. The 0.12.0 release notes list WSL as a supported platform, and for a student the least painful path is to install WSL2 with Ubuntu and follow the Linux instructions inside it. This is practical advice rather than an official recommendation, but it avoids a long list of Windows-specific setup steps.

Installing with Docker instead

If installing Java and matching versions sounds like a chore, Docker (see /student-guides/docker) removes it. The project publishes images pytorch/torchserve (CPU) and a GPU variant. Because the project is archived, the latest tag will never change again, but it is still better practice to pin a specific tag so that what you run today is what you run next year. Check the tag list on Docker Hub before choosing one. A later section shows how to run it.

Try it
  1. Create a virtual environment and install the three packages above.
  2. Run java -version, torchserve --version and torch-model-archiver --version.
  3. If the Java version is not 17, fix it before moving on.
Java 17 and TorchServe 0.12.0 printed with no errors. Every later section depends on this.

Preparing a model to serve

TorchServe does not train anything. It needs a saved model, and PyTorch gives you two common ways to save one. The difference decides which archiver flags you use, so it is worth understanding.

Eager mode (a state dictionary). You save only the learned numbers with torch.save(model.state_dict(), "model.pth"). At serving time you need the Python class that defines the network, so you must also ship a file containing that class. This is flexible and common, and it is what the archiver calls an eager-mode model.

TorchScript. You convert the model into a self-contained, serialisable form with torch.jit.trace or torch.jit.script. The saved file includes the computation itself, so no separate model-definition file is needed. It is simpler to package.

For a first project we will use the TorchScript route with a pretrained ResNet-18 from torchvision, because it needs no custom code and the built-in image_classifier handler fits it. Create a working folder and a script:

export_model.py
import torch
from torchvision import models

# Load a ResNet-18 pretrained on ImageNet and switch to inference mode.
weights = models.ResNet18_Weights.DEFAULT
model = models.resnet18(weights=weights)
model.eval()

# Trace the model with a dummy 224x224 RGB image to produce TorchScript.
example = torch.rand(1, 3, 224, 224)
traced = torch.jit.trace(model, example)
traced.save("resnet18.pt")
print("saved resnet18.pt")

Run it with python export_model.py. The first run downloads the pretrained weights (roughly 45 MB), then writes resnet18.pt. Two details matter. model.eval() turns off training behaviours such as dropout and batch-norm updates; forgetting it gives subtly wrong predictions. And torch.jit.trace records what the model does for one example input, so the dummy input must have the same shape the real images will have after preprocessing.

You also need a test image. Any small JPEG will do; a photo of a cat or a dog works well because ImageNet has many categories for them. Save it as kitten.jpg in the same folder.

Class names, so the output is readable

A classifier outputs a number for each of 1,000 ImageNet classes, and the useful answer is "tabby cat", not "281". The built-in image_classifier handler can translate indices into names if you supply a JSON file called index_to_name.json through the archiver's --extra-files option. The torchvision package ships the names inside the weights' metadata, so you can write the file yourself:

make_labels.py
import json
from torchvision import models

names = models.ResNet18_Weights.DEFAULT.meta["categories"]
mapping = {str(i): name for i, name in enumerate(names)}
with open("index_to_name.json", "w") as f:
    json.dump(mapping, f)
print(len(mapping), "classes written")

This prints 1000 classes written. The file maps the string "0" to the first class name and so on. If you skip this file, the handler still works; the response will simply identify classes by index rather than by name, so keep the file when you want readable answers.

Match preprocessing to training The built-in image handler resizes and normalises images the standard ImageNet way. That suits a stock torchvision model. If you trained your own model with different image sizes or different normalisation numbers, the built-in handler will quietly feed it the wrong input and the predictions will be poor with no error at all. That is the moment to write a custom handler, covered later in this guide.
Try it
  1. Run export_model.py and make_labels.py in an empty folder.
  2. Put a JPEG named kitten.jpg beside them.
  3. Run ls -lh and look at the sizes.
a resnet18.pt of roughly 45 MB and a small index_to_name.json. These are the only ingredients you need for the next step.

Packaging the model with torch-model-archiver

The archiver is the tool that produces the .mar file. Its job is to bring the weights, the handler and the extras together under one name and version, so that the server can load the result without knowing anything else about your project.

For our TorchScript model the command is:

BASH
torch-model-archiver \
  --model-name resnet18 \
  --version 1.0 \
  --serialized-file resnet18.pt \
  --handler image_classifier \
  --extra-files index_to_name.json \
  --export-path model_store \
  --force

Read it flag by flag, because these are the flags you will use every time:

  • --model-name is the name clients will use in the URL. It becomes /predictions/resnet18. Choose something short, without spaces.
  • --version labels this build of the model. TorchServe supports several versions of the same name side by side, and you will use that to roll out updates safely.
  • --serialized-file points to the weights file: the .pt TorchScript file here, or a .pth state dictionary for eager mode.
  • --handler names the handler. image_classifier is one of the built-in ones; you can instead give the path to your own Python file.
  • --extra-files lists additional files, separated by commas, to include in the archive.
  • --export-path is the folder where the .mar is written. We point it at model_store, which will be our model store.
  • --force (short form -f) overwrites an existing archive of the same name. Without it, re-running the command fails when the file already exists.

If a folder named model_store does not exist yet, create it first with mkdir model_store. After success you will find model_store/resnet18.mar.

The eager-mode variant

If you saved a state dictionary instead of TorchScript, you must additionally provide the file that defines the model class, with --model-file, and usually a --requirements-file for extra Python packages the handler needs:

BASH
torch-model-archiver \
  --model-name mynet --version 1.0 \
  --model-file model.py \
  --serialized-file mynet_weights.pth \
  --handler image_classifier \
  --requirements-file requirements.txt \
  --export-path model_store

The rule: --model-file is mandatory in eager mode and unnecessary for TorchScript. If you forget it in eager mode, the worker will fail when it tries to build the model, and the error appears in the logs rather than at archive time, which is a classic way to lose twenty minutes.

Other options exist: --runtime to choose the Python runtime, --config-file to attach a YAML configuration, and --archive-format to choose between the default .mar, a tgz file, a zip without compression, or a plain folder (no-archive). You can leave all of those at their defaults for now.

What is inside

A .mar is a zip-like file, so you can inspect it. Run unzip -l model_store/resnet18.mar and you will see your weights, the labels file and a MANIFEST.json inside a MAR-INF folder. The manifest records the model name, version and handler. Looking inside once removes the mystery: the archive is just your files plus a description.

Try it
  1. Create model_store and run the archiver command above.
  2. Run unzip -l model_store/resnet18.mar.
  3. Open the manifest inside and find the handler name.
a listing with your weights, index_to_name.json and a manifest, and a manifest that records the handler. Nothing more is hidden in there.

Starting the server and sending your first prediction

Now the payoff. With model_store/resnet18.mar in place, start TorchServe:

BASH
torchserve --start --ncs --model-store model_store --models resnet18=resnet18.mar --disable-token-auth

Each flag has a reason:

  • --start launches the server in the background and returns your prompt. The server keeps running after the command ends.
  • --ncs (long form --no-config-snapshots) turns off TorchServe's habit of saving its state to snapshot files. We will explain snapshots shortly; for learning, switching them off avoids surprising behaviour after a restart.
  • --model-store is the folder of .mar files. It is the one option that is mandatory.
  • --models lists which models to load at startup, as name=file.mar. The name before the equals sign is the name clients use in the URL. You can list several, separated by spaces.
  • --disable-token-auth turns off token authentication. More on this in a moment; it is for local learning only.

After a few seconds, check that the server is alive:

BASH
curl http://localhost:8080/ping

A healthy server answers:

JSON
{
  "status": "Healthy"
}

/ping returns success when the number of active workers is at least the model's minimum. If you call it too quickly after starting, or a worker failed to load, you get an error status instead. That makes it useful as a readiness check later in Kubernetes (/student-guides/kubernetes).

Now ask for a prediction. Send the image file as the request body with curl's -T option:

BASH
curl http://localhost:8080/predictions/resnet18 -T kitten.jpg

The reply is a JSON object mapping class names to probabilities, best guess first. With the labels file included, the keys are readable names such as a cat breed; the values are numbers between 0 and 1 that sum to about 1 across the top few classes. The built-in handler returns the five most likely classes. You have just served a PyTorch model over HTTP.

Stop the server when you are done:

BASH
torchserve --stop

Why you pass --disable-token-auth

Since version 0.11.1, TorchServe enforces token authorization by default. When the server starts it writes a file called key_file.json into the directory you launched it from, containing an inference key, a management key and an API key used to mint new tokens. Every request must then carry a header of the form Authorization: Bearer <key>. Tokens expire after 60 minutes by default.

This is a good default, and in production you should keep it on. But most tutorials on the internet were written for older versions and omit the header, so beginners follow a tutorial, get a 401-style refusal and conclude TorchServe is broken. There are two honest ways forward. For local learning, start with --disable-token-auth, as above. Or keep authentication on and send the key:

BASH
curl http://127.0.0.1:8080/predictions/resnet18 \
  -H "Authorization: Bearer <INFERENCE_KEY>" -T kitten.jpg

where the inference key is the value found in key_file.json. Never commit that file to Git and never copy it into a container image: it contains live credentials.

The version-0.11.1 change that breaks old tutorials The same release that turned on token authorization also disabled registering and deleting models over the HTTP and gRPC management APIs by default. Older guides show curl -X POST "localhost:8081/models?url=..." as the first step. On current versions that is refused unless the model API is explicitly enabled and a management token is supplied. The simple, reliable beginner approach is to pre-load models with --models at startup, exactly as we did.

Snapshots, briefly

TorchServe can record which models and how many workers were running into snapshot files, so that after a restart it comes back in the same state. That is helpful for servers, and confusing for learners, because an old snapshot can restore a state you forgot about, or fail to load and produce an InvalidSnapshotException. Using --ncs while learning sidesteps both.

Try it
  1. Start the server with the command above, wait a few seconds, then call /ping.
  2. Send kitten.jpg to /predictions/resnet18.
  3. Send a different image, ideally something that is not an animal, and read the result.
  4. Run torchserve --stop.
a "Healthy" status and a JSON list of class guesses. The non-animal image still gets confident-looking guesses from the 1,000 classes the model knows; a classifier always answers, even when the answer is nonsense.

Handlers: how a request becomes a prediction

So far we used a built-in handler and wrote no code. To understand what happened, and to serve any model that is not a plain image classifier, you need to know what a handler does.

A request arrives at the frontend, which passes the raw bytes to a worker. The worker's handler runs a fixed pipeline of four stages:

  1. initialize(context) runs once, when the worker starts. It loads the model weights and moves them to the right device (CPU or a specific GPU). Anything expensive belongs here, because it must not be repeated for every request.
  2. preprocess(data) runs for each request (or each batch). It converts the raw input, such as JPEG bytes, into a tensor in the shape the model expects.
  3. inference(data) runs the model on that tensor and returns the raw output.
  4. postprocess(output) turns the raw output into something that can be returned to the client, usually a list with one entry per request.

A function called handle(data, context) orchestrates the four. Most handlers do not implement all of it themselves; they inherit from BaseHandler, which provides sensible defaults for model loading and inference, and override only the stages they need to change.

The built-in handlers

TorchServe includes ready-made handlers for common jobs, selected by name in the --handler flag:

Handler Purpose
image_classifier Classify an image into labels
image_segmenter Label each pixel of an image
object_detector Find and box objects in an image
text_classifier Classify a piece of text

Each lives in the ts.torch_handler package, so a custom handler can reuse one with an import such as from ts.torch_handler.image_classifier import ImageClassifier. Note that the official 0.12.0 release notes deprecate TorchText support, so treat NLP handlers that depend on it with caution.

Writing a small custom handler

Suppose your model is a text or tabular model, or an image model with unusual preprocessing. You write a Python file with a class that inherits from BaseHandler. Here is a handler for a model that takes a list of numbers and returns one score:

numeric_handler.py
import json
import torch
from ts.torch_handler.base_handler import BaseHandler


class NumericHandler(BaseHandler):
    """Serve a model that takes a JSON list of floats and returns a score."""

    def preprocess(self, data):
        rows = []
        for item in data:
            payload = item.get("data") or item.get("body")
            if isinstance(payload, (bytes, bytearray)):
                payload = payload.decode("utf-8")
            rows.append(json.loads(payload) if isinstance(payload, str) else payload)
        return torch.tensor(rows, dtype=torch.float32)

    def postprocess(self, output):
        # Must return a list with exactly one element per request in the batch.
        return output.squeeze(-1).tolist()

Several things in this small file are worth stopping on:

  • The handler is a class that inherits BaseHandler. If a file contains several classes, the entry class has to be the first one.
  • data is always a list, one element per request in the current batch. When batching is off, the list has one element, but your code should still loop over it. Writing handlers as if there were always exactly one request is the root of many bugs once batching is switched on.
  • The request content arrives under a key such as data or body, depending on how it was sent, which is why the code checks both. The raw bytes may need decoding.
  • postprocess must return a list the same length as the input list. If it returns a bare number or a list of the wrong length, the responses are matched to the wrong clients or the request fails.

You package it by passing the file path as the handler:

BASH
torch-model-archiver --model-name scorer --version 1.0 \
  --model-file model.py --serialized-file scorer.pth \
  --handler numeric_handler.py --export-path model_store

When you write code that will run inside a worker, remember that every print goes into a log file rather than your terminal, and an exception during initialize kills the worker. The log section below shows where to look.

Test handler logic without a server Python errors inside a running worker are slow to debug, because you must archive, restart and send a request each time. Keep your pre- and post-processing logic in plain functions, run them in a normal Python shell on a sample input, and only then put them into the handler class. A thirty-second test in a shell saves a ten-minute archive-and-restart loop.
Try it
  1. Find the file for the built-in image_classifier handler in your virtual environment's site-packages/ts/torch_handler folder.
  2. Read it and locate which of the four stages it overrides.
  3. Compare it with base_handler.py in the same folder.
a short subclass that overrides preprocessing and postprocessing while inheriting the rest. Reading real handlers is the fastest way to learn the conventions.

Managing a running server

Once a server is running you will want to look at it: which models are loaded, how many workers each has, whether they are healthy. The management API on port 8081 answers those questions. Because it can change the server, it is protected more strictly than the inference API, and on current versions the same token rules apply (with a management key rather than the inference key). If you started with --disable-token-auth, you can call it without headers.

List the models:

BASH
curl http://localhost:8081/models

Describe one model:

BASH
curl http://localhost:8081/models/resnet18

The description includes the model name and version, the minimum and maximum worker counts, the batch size and batch delay, a list of workers with their status, GPU use and memory, and a jobQueueStatus showing how many requests are waiting and how much capacity remains. When someone asks "is the model keeping up?", the queue status is the first place to look: a queue that is regularly full means requests arrive faster than workers can handle them.

Workers: scaling a single model

A worker is one process holding one copy of the model. By default the number of workers per model follows your hardware (the number of GPUs, or logical CPUs). You can change the number for a specific model at runtime:

BASH
curl -X PUT "http://localhost:8081/models/resnet18?min_worker=2&synchronous=true"

min_worker=2 asks for at least two workers. With synchronous=true the call waits until the workers are ready and answers with status 200; without it, the call returns 202 immediately and the workers start up in the background. Each worker loads its own copy of the model, so two workers use roughly twice the memory. On a CPU-only laptop, more workers than cores brings no benefit and can slow everything down, because they compete for the same cores.

You can see the effect by running the describe command again and counting the entries under workers. Scaling a model is the cheapest experiment in this guide and the one that teaches most about how serving capacity works.

Versions

TorchServe supports multiple versions of one model name. Archive the model again with --version 2.0, and the server can hold both. Requests to /predictions/resnet18 go to the default version, requests to /predictions/resnet18/2.0 go to that specific version, and an administrator can switch the default with PUT /models/resnet18/2.0/set-default. This lets you deploy a new model next to the old one, test it with a version-specific URL, and only then promote it. The principle is the same as blue-green deployments elsewhere in operations.

Never expose the management port Whoever can reach port 8081 can load models, and a model archive runs arbitrary Python. By default the management API listens on localhost only, which is safe. If you ever change the address to 0.0.0.0 to make it reachable, you have given strangers a way to run code on your machine. Keep management on a private network and keep token authorization enabled.
Try it
  1. Start the server again and describe the resnet18 model; count the workers.
  2. Scale it to two workers with synchronous=true, then describe it again.
  3. Watch memory use in your system monitor before and after.
one more entry in the workers list and a visible rise in memory. That is the cost of parallelism: each worker owns a full copy of the model.

Dynamic batching: trading a little latency for throughput

A GPU, and to a lesser degree a CPU, handles many inputs at once much more efficiently than one at a time. If a thousand clients each send one image, running a thousand separate forward passes wastes most of the hardware. Dynamic batching groups requests that arrive close together into one forward pass.

TorchServe controls it with two numbers per model:

  • batchSize, the maximum number of requests to combine into one batch.
  • maxBatchDelay, the longest time in milliseconds the server will wait for a batch to fill before running with whatever it has.

The rule is: the worker collects requests until it has batchSize of them, or until maxBatchDelay has passed since the first one, whichever comes first. A quiet period therefore costs each request at most the delay, while a busy period fills batches quickly and gets the efficiency gain.

REQ 1t = 0 ms
+
REQ 2t = 8 ms
+
REQ 3t = 20 ms
→
ONE BATCHdelay expired, run 3

Choosing values is a tuning exercise. A large batchSize with a long delay maximises throughput and hurts the latency of each individual request. For interactive applications, a short delay such as 10 to 50 ms keeps responses snappy; for offline scoring, a long delay and a large batch are fine. There is no universal right answer, which is why you measure with a load test rather than guessing.

You set these values when registering a model. With the model API disabled by default, the straightforward beginner route is the configuration file, shown in the next section. The registration call's parameters are named batch_size and max_batch_delay, for example POST /models?url=resnet-152.mar&batch_size=8&max_batch_delay=50, if you enable the model API in your setup.

Your handler must cooperate

Batching only works if the handler expects a list. As described earlier, the data argument is a list with one entry per request, and postprocess must return a list of the same length. The built-in handlers already do this. A custom handler that assumes a single request will fail or mix up answers the moment batching is turned on. Write for the list from the beginning.

Try it
  1. Write a small Python loop that sends 50 requests to the server one after another and prints the total time.
  2. Repeat it using a thread pool of 8 to send requests at once.
  3. Compare the two totals and look at the queue status in the describe output while the second run is going.
the concurrent run finishes in far less than 8 times the sequential time per request, and the queue briefly shows pending requests. That is the pressure that workers and batching exist to absorb.

Configuration: config.properties and command-line options

Typing long commands every time is error-prone, and settings that are not in a file cannot be reviewed or shared. TorchServe reads settings from a config.properties file, which you point at with --ts-config:

BASH
torchserve --start --ncs --ts-config config.properties

A beginner-friendly file looks like this:

config.properties
# Where the three APIs listen. These are the defaults; keep them local.
inference_address=http://127.0.0.1:8080
management_address=http://127.0.0.1:8081
metrics_address=http://127.0.0.1:8082

# Where model archives live and which to load at startup.
model_store=model_store
load_models=resnet18.mar

# Capacity: how many requests may wait for a worker.
job_queue_size=100

# Larger uploads than the default need this raised.
max_request_size=20000000
max_response_size=20000000

Each line is a key=value pair; lines starting with # are comments. The settings above are the ones a beginner meets most:

  • inference_address, management_address and metrics_address set where the three APIs listen. The defaults bind to 127.0.0.1, meaning "only this machine".
  • model_store and load_models replace the --model-store and --models flags. load_models accepts a list of .mar files, or all to load everything in the store.
  • default_workers_per_model sets how many workers each model gets; if you leave it out, TorchServe uses the number of GPUs, or the number of logical CPUs on a CPU-only machine.
  • job_queue_size is the per-model queue length (default 100).
  • max_request_size and max_response_size limit message sizes in bytes. The Troubleshooting page notes that uploads of more than about 6.5 MB fail by default, and raising these values is the fix.
  • install_py_dep_per_model=true tells TorchServe to install the packages from a model's requirements.txt when it loads the model. Without this setting, the file you passed with --requirements-file is not used and imports fail inside the worker.
  • allowed_urls is a list of patterns (regular expressions) that restricts where models may be downloaded from. The default permits file:// and http(s):// sources; in a real deployment you narrow it to your own model registry.

The configuration page also describes how to define models, with their versions, worker counts and batch settings, inside the file using a models= entry in JSON form. That is the usual way to set batchSize and maxBatchDelay without the management API. The exact structure is documented on the configuration page of the official docs. Read it before copying an example, because a small JSON mistake makes the file fail to parse.

Environment variables

With enable_envvars_config=true in the file, you can override properties with environment variables named TS_ followed by the property in capitals, for example TS_INFERENCE_ADDRESS. This is how container deployments change settings without editing files. Be careful when mixing the three places (file, environment, command line): the docs describe a precedence order, but different pages word it differently, so avoid depending on subtle overrides. Pick one mechanism for each setting and stay with it.

Keep the file in Git config.properties contains no secrets unless you add TLS keystore passwords to it, so it belongs in version control next to your archiving script. A teammate can reproduce your server from the repository, which was the whole point of serving the model this way.
Try it
  1. Stop the server, write the config.properties above and start TorchServe with --ts-config.
  2. Call /ping and a prediction to confirm that it works.
  3. Change inference_address to port 8090, restart, and confirm that port 8080 no longer answers.
everything works with no model flags on the command line, and the port change moves the API. Settings in a file are repeatable settings.

Running TorchServe in Docker

Containers solve the "which Java, which Python, which PyTorch" problem by shipping them together. The project publishes the images pytorch/torchserve for CPU and a -gpu variant. To start the CPU image, publish the ports, mount your model store and tell it which model to load:

BASH
docker run --rm -it \
  -p 127.0.0.1:8080:8080 -p 127.0.0.1:8081:8081 -p 127.0.0.1:8082:8082 \
  --mount type=bind,source="$(pwd)/model_store",target=/tmp/models \
  pytorch/torchserve:latest \
  torchserve --model-store=/tmp/models --models resnet18=resnet18.mar --disable-token-auth

The details matter:

  • The -p 127.0.0.1:8080:8080 form publishes the container's port 8080 on your loopback interface only. Writing -p 8080:8080 instead publishes it on all interfaces, which can expose the port to your whole network. The project's own Docker instructions bind to 127.0.0.1 for this reason.
  • --mount type=bind makes your model_store folder appear inside the container at /tmp/models.
  • The command at the end is the same torchserve command you already know. The container's job is just to run it in a controlled environment. In a container the server runs in the foreground, so you do not use --start.

For a more production-like run, the official Docker instructions add --shm-size=1g --ulimit memlock=-1 --ulimit stack=67108864, which give the worker processes enough shared memory. A common surprise for beginners is that PyTorch workers need more shared memory than the container default, and the symptom is a worker that crashes while handling a request.

Pin the image tag Because the project is archived, latest will not move again, but pinning is still the right habit: choose an explicit version tag from the Docker Hub tag list, write it in your notes, and use it everywhere. An unpinned tag is a surprise waiting for the day somebody rebuilds the image on a different machine.
Try it
  1. Run the container with your model_store mounted.
  2. From another terminal, call /ping and send kitten.jpg.
  3. Press Ctrl+C in the container terminal and confirm the port stops answering.
the same results as the local install, without needing Java or TorchServe on your machine. If it did not work, check that the port mapping and the bind-mount path are right.

Logs, metrics and reading errors

When a model server misbehaves, the answer is almost always in its logs. TorchServe writes them to a logs folder in the directory where you started it. The two to know first are ts_log.log, which has the server's own messages including worker start-up and crashes, and model_log.log, which carries output from your handler code, including anything you print. There is also an access log recording each request. If you are unsure of exact names in your installation, run ls logs after starting the server.

A good habit: when something fails, run tail -n 100 logs/ts_log.log before changing any code. The log usually tells you the specific exception, and reading it first is quicker than guessing.

Metrics

The metrics API on port 8082 and the log files report what the server is doing: requests by status class, latency, queue time, CPU and memory use, and, on a GPU machine, GPU utilisation. By default metrics are written to log files (ts_metrics.log and model_metrics.log). The docs also describe a mode that exposes them in Prometheus format, which pairs with Prometheus and Grafana (/student-guides/prometheus, /student-guides/grafana); enabling it is a configuration choice for the mid-level guide.

The errors you will actually meet

These come from the official Troubleshooting page and from the behaviour described above. Learn to recognise them.

What you see What it means What to do
Failed to bind to address: http://127.0.0.1:8080 Something else already uses that port, often an old TorchServe Run torchserve --stop, find the process, or change the port in config.properties
java.lang.NoSuchMethodError at startup The wrong Java version Install and select JDK 17
HTTP 404, ModelNotFoundException The model name in the URL does not match, or the .mar is not in the store Check GET /models and the file name in model_store
HTTP 503, ServiceUnavailableException No worker is ready, or the queue is full Check worker status with the describe call, scale workers, or raise job_queue_size
Backend worker monitoring thread interrupted A worker died while loading, often an import error or missing file in the handler Read ts_log.log and model_log.log for the Python traceback
HTTP 409, ConflictStatusException That model version is already registered Use a different version or name
HTTP 400, DownloadModelException The model URL cannot be reached or is not in allowed_urls Fix the path or the allow-list
Large uploads fail Request-size limit around 6.5 MB Raise max_request_size and max_response_size
InvalidSnapshotException after a restart A stale or broken snapshot Start with --ncs, or delete the snapshot files
Unauthorized errors on every call Token authorization is on and the request has no valid token Send the Authorization: Bearer header with a key from key_file.json, or use --disable-token-auth locally. Tokens expire after 60 minutes
Try it
  1. Deliberately break something: archive a model with a handler name that does not exist, or remove the labels file and archive again.
  2. Start the server and call the describe endpoint.
  3. Read logs/ts_log.log and find the line that explains the failure.
a model with no healthy workers and a traceback in the log. Having caused and read a failure once makes the next real one far less frightening.

Safe defaults and the security you need from day one

Serving a model means running a network service, and a network service is a target. The good news is that TorchServe's defaults are cautious. The habits below keep them that way.

Keep the ports local. All the APIs listen on 127.0.0.1 by default. In Docker, publish ports with the 127.0.0.1: prefix. If you must reach the server from other machines, put a reverse proxy or load balancer in front for the inference port and keep management and gRPC management ports (8081 and 7071) private.

Keep token authorization on outside your laptop. --disable-token-auth is a convenience for learning. Any shared or deployed environment should use the keys from key_file.json, with the inference key for clients and the management key only for administrators.

Trust only archives you made or verified. A .mar file contains Python code that the worker executes. Downloading a model archive from an unknown website and loading it is equivalent to running an unknown script. Restrict allowed_urls to your own storage.

Do not commit key_file.json. Add it to .gitignore the moment you start the server in a project folder.

Know that the project will not patch itself. Because TorchServe is no longer maintained, any vulnerability found in it or in its Java components stays unfixed in this project. If you ever run it where it matters, isolate it: a private network, a non-root user, nothing sensitive reachable from it. In the Gulf and Egypt, where many employers have data-residency rules and security reviews, expect a reviewer to ask exactly this question, and to prefer an actively maintained server.

Try it
  1. Start the server without --disable-token-auth and call the prediction URL without a header.
  2. Open key_file.json, copy the inference key, and call again with the header.
  3. Add key_file.json to your .gitignore.
a refusal first and a prediction second. You have now seen the production-style behaviour and the reason the default changed.

Putting it all together

Here is a small end-to-end project that uses everything above. The goal: a reproducible folder that anyone can clone, run one script, and get an image classifier on port 8080.

TEXT
image-service/
  export_model.py
  make_labels.py
  build.sh
  config.properties
  .gitignore
  README.md

The build.sh script performs the packaging steps in order, so that the .mar is always built the same way:

build.sh
#!/usr/bin/env bash
set -euo pipefail

python export_model.py          # writes resnet18.pt
python make_labels.py           # writes index_to_name.json

mkdir -p model_store
torch-model-archiver \
  --model-name resnet18 --version 1.0 \
  --serialized-file resnet18.pt \
  --handler image_classifier \
  --extra-files index_to_name.json \
  --export-path model_store --force

echo "built model_store/resnet18.mar"

And a run.sh that starts the server, waits for it to be healthy, and sends a test image:

run.sh
#!/usr/bin/env bash
set -euo pipefail

torchserve --start --ncs --ts-config config.properties --disable-token-auth

# Wait up to 30 seconds for the model's workers to become ready.
for i in $(seq 1 30); do
  if curl -fs http://127.0.0.1:8080/ping > /dev/null; then break; fi
  sleep 1
done

curl http://127.0.0.1:8080/predictions/resnet18 -T kitten.jpg

Add key_file.json, logs/, model_store/ and *.pt to .gitignore, because they are generated artifacts, not source. Your README should state the pinned versions: Python 3.11, JDK 17, PyTorch 2.4, TorchServe 0.12.0.

Work through this checklist as the final exercise:

  1. Run build.sh in a clean folder and confirm the .mar appears.
  2. Run run.sh and confirm the prediction.
  3. Describe the model, scale to two workers, and confirm the change.
  4. Look into the logs and find the line recording the model's load.
  5. Stop the server and run the same project in Docker with the bind mount.
  6. Change the model name to resnet18-v2 by rebuilding with --version 2.0, and think about how you would roll it out beside version 1.

If all six work, you have gone through the full lifecycle of a served model: export, package, serve, call, scale, observe, containerise and version.

Try it
  1. Delete the generated files and rebuild everything from the scripts alone.
  2. Give the folder to a friend, or to a second machine, and ask them to run it with no help.
it works the first time from the README. If it does not, the missing step belongs in the README or the scripts.

What you can now do, and what comes next

You can now take a PyTorch model and make it a service. You know the vocabulary (archive, model store, handler, worker, queue, frontend, backend), the three ports and what each is for, how to archive a model, how to start the server with the right flags for the current version, how to call it, how to scale workers and read the describe output, how batching trades latency for throughput, how a configuration file replaces command-line flags, how to run it in Docker without exposing it, and how to read a log when a worker fails.

That is the whole beginner toolkit, and it is enough to understand most conversations about model serving.

What comes next in this series

The mid-level guide goes into custom handlers in depth, batching and worker tuning, metrics in Prometheus mode, and deployment on Kubernetes. The senior guide covers failure modes, security hardening, and what to do about the project's maintenance status in a real organisation.

Where to look if you need an actively maintained server

Because TorchServe has stopped receiving updates, a team starting a new project today would normally weigh these alternatives:

  • Triton Inference Server (/student-guides/triton) serves models from several frameworks, with strong batching and GPU support.
  • BentoML (/student-guides/bentoml) packages models as Python services with a friendlier developer experience.
  • Ray Serve (/student-guides/ray-serve) scales Python serving across a cluster.
  • KServe (/student-guides/kserve) is a Kubernetes-native layer for model serving; TorchServe can be one of its runtimes.
  • vLLM (/student-guides/vllm) is the usual choice for serving large language models.

The concepts you learned here (an archive of model and code, a handler with a lifecycle, workers, queues, batching, health checks) transfer directly to all of them. Learning one serving system thoroughly makes the next one a matter of reading its documentation.

Sources

Try it
  1. Start a model with two workers, then run ps aux | grep -i -E "java|ts_" and count the processes. One JVM, two workers.
  2. Watch memory with top while you scale workers up with the management API, and compute the per-worker footprint.

Following one request through queue, batch and worker

A prediction request does not go straight to your handler. Knowing the path lets you diagnose latency by stage instead of guessing.

First the frontend authenticates the request if token authorization is on, which is the default since 0.11.1. Then it resolves the model name and version. If you did not specify a version, it uses the default version. The request is turned into a job and placed on the job queue for that model. The queue holds up to job_queue_size entries, 100 by default. A worker pulls jobs from the queue, calls the handler, and the response is sent back to the waiting connection.

Two failure modes come from this design, and they look similar from the client.

When there are no workers at all for the model, the frontend answers 503 ServiceUnavailableException. When there are workers but the queue is full, the answer is the same 503. The official troubleshooting page gives the same two fixes: scale the workers, or raise job_queue_size. Note the ordering of the thinking, though. Raising the queue size when workers cannot keep up only adds latency, because requests wait longer before they fail or succeed. A bigger queue is appropriate for absorbing short bursts. A sustained overload needs more workers, a faster handler, or load shedding upstream.

The queue is also where queue latency is born. The metric ts_queue_latency_microseconds measures how long a request waited before a worker picked it up, and it is the single best autoscaling signal you have. Inference time measured inside the handler tells you how fast the model is. Queue time tells you whether you have enough workers.

Dynamic batching

For models that run on GPUs, running one request at a time wastes the hardware. Dynamic batching lets TorchServe collect several requests and hand them to the handler as a list. You set two numbers per model: batchSize, the maximum number of requests per batch, and maxBatchDelay, the maximum number of milliseconds the worker will wait for the batch to fill.

The rule is: the worker takes whatever is in the queue, and if it has fewer than batchSize jobs, it waits up to maxBatchDelay milliseconds for more. Whichever limit is reached first triggers the batch. This yields the tuning model you need:

  • Under heavy load the batch fills immediately and maxBatchDelay never matters. You get the throughput benefit at no latency cost.
  • Under light load a single request waits the full maxBatchDelay for company that never arrives, so you add up to that delay to every request's latency. Set it too high for a low-traffic service and you have made it slower for nothing.
  • The right maxBatchDelay is a fraction of your latency budget, often a few tens of milliseconds, and the right batchSize is the largest batch that still fits in GPU memory and meets the budget at p99.

Batching is configured at registration time, either in the models JSON in config.properties, in a model config file, or through the management API:

BASH
curl -X POST "http://localhost:8081/models?url=resnet-152.mar&batch_size=8&max_batch_delay=50" \
  -H "Authorization: Bearer $MGMT"

The catch, which breaks more first attempts than any other, is that your handler must be written for batches. The data argument is a list with one entry per request, and your return value must be a list of the same length, in the same order. A handler that assumes data[0] is the only request will silently return an answer for the first caller and drop the rest, or crash. We cover this in the handler section.

Measure before you batch Batching helps when the GPU is underused per request. For a tiny CPU model, a batch of one is already fast, and a batch delay only adds latency. Benchmark at batch sizes 1, 4, 8 and 16 under your real concurrency before settling.
Try it
  1. Register a model with batch_size=1, send 50 concurrent requests, and record the median and p99 latency.
  2. Repeat with batch_size=8&max_batch_delay=20. Compare throughput and tail latency, and note what happens when you send only one request at a time.

Handlers that survive production

The built-in handlers cover standard image and text cases. Real services need custom behaviour: a different input format, business logic after inference, a model that is not a single nn.Module. A custom handler is how you get that, and it is also where most production bugs live, because it is arbitrary Python executing inside a worker.

The lifecycle

The pipeline from the facts above is initialize, then for every request or batch preprocess, inference, postprocess, orchestrated by handle(data, context). BaseHandler implements all of it, so the usual pattern is to subclass it and override only what you need:

handler.py
import json
import torch
from ts.torch_handler.base_handler import BaseHandler


class ScoringHandler(BaseHandler):
    def initialize(self, context):
        super().initialize(context)          # loads the model, picks the device
        self.threshold = 0.5

    def preprocess(self, data):
        batch = []
        for row in data:                     # one entry per request in the batch
            payload = row.get("data") or row.get("body")
            if isinstance(payload, (bytes, bytearray)):
                payload = json.loads(payload)
            batch.append(torch.tensor(payload["features"], dtype=torch.float32))
        return torch.stack(batch).to(self.device)

    def postprocess(self, output):
        probs = torch.sigmoid(output).squeeze(-1).tolist()
        return [{"score": p, "label": p >= self.threshold} for p in probs]

If the file contains several classes, the entry class must come first, because TorchServe looks for the first one. A handler may instead be a module-level function handle(data, context), which is fine for small cases but harder to test.

What belongs in initialize

initialize runs once per worker, at startup, and it is where expensive, shared state belongs: loading weights, building tokenizers, warming the model, reading side files shipped with the archive. The context object tells you where things are. The model directory is available in context.system_properties, and the GPU id for this worker is there too, which is how BaseHandler chooses the device. Never do that work in preprocess, because it would repeat for every request.

A subtle point: if initialize raises, the worker dies, the frontend restarts it, and the loop repeats. The log line you see is Backend worker monitoring thread interrupted, and the real exception is in the backend log, so read that rather than the frontend output. A missing import and an out-of-memory error during weight loading look identical from outside.

A good habit is to warm the model inside initialize by running one dummy forward pass. This forces CUDA context creation, kernel selection and lazy allocations to happen before the worker reports ready, so that the first real request does not pay a multi-second penalty and trigger a timeout.

Writing for batches

Rules for batch-safe handlers, which apply even when you configure batchSize as 1, so that you can change it later without a rewrite:

  1. Treat data as a list. Loop over it. Never index [0].
  2. Return a list whose length equals len(data), in the same order. If one request is malformed, you must still return something at its position, usually an error object, rather than raising and failing the whole batch.
  3. Keep per-request failure isolated. A single bad image in a batch of eight should produce one error entry and seven predictions.
  4. Pad or stack carefully. Variable-length text needs padding to a common length before stacking, and the padding must be removed or masked again in postprocessing.

A defensive preprocess wraps each element:

PYTHON
def preprocess(self, data):
    tensors, errors = [], {}
    for i, row in enumerate(data):
        try:
            tensors.append(self._decode(row))
        except Exception as exc:             # one bad request must not poison the batch
            errors[i] = str(exc)
            tensors.append(self._placeholder())
    self._errors = errors
    return torch.stack(tensors).to(self.device)

Storing self._errors on the handler instance is safe here because a worker processes one batch at a time. Postprocess then replaces the placeholder predictions at those indices with error objects.

Testing a handler without a server

The best news about handlers is that they are plain Python classes. You can unit test them without starting TorchServe by building a small fake context:

test_handler.py
from types import SimpleNamespace
import torch
from handler import ScoringHandler


def make_handler(tmp_path):
    h = ScoringHandler()
    h.model = torch.nn.Linear(4, 1)
    h.device = torch.device("cpu")
    h.threshold = 0.5
    h.initialized = True
    return h


def test_batch_returns_one_result_per_request(tmp_path):
    h = make_handler(tmp_path)
    body = b'{"features": [0.1, 0.2, 0.3, 0.4]}'
    requests = [{"body": body}, {"body": body}, {"body": body}]
    out = h.postprocess(h.inference(h.preprocess(requests)))
    assert len(out) == 3
    assert all("score" in r for r in out)

This test runs in milliseconds in CI and catches the single most common handler bug, the batch of N returning fewer than N answers. The inference method of BaseHandler expects self.model and self.device to be set, which the test does by hand. If your base-class version names those attributes differently, read the installed ts/torch_handler/base_handler.py, which is the authority.

A handler is code that the server trusts completely Everything in a .mar, including your handler, executes inside the worker with the worker's privileges. Handler code that shells out, opens sockets or reads environment variables does so as the service. Review handlers like application code, and never load archives whose origin you cannot establish.
Try it
  1. Take a handler you use and add a unit test that sends three requests and asserts three results.
  2. Add one deliberately malformed request in the middle and make the test pass by returning an error object at that position only.

Archives you can rebuild and trust

Beginner showed one torch-model-archiver command. At work, the archive is a release artefact, and the questions change: can I rebuild this exact file, what is inside it, and how do I know nobody swapped it?

What goes in

A .mar is a packaged directory holding the serialised weights (--serialized-file), the handler, an optional model definition file (--model-file, mandatory in eager mode and unnecessary for TorchScript), any extra files you list with --extra-files as a comma-separated list, a generated MANIFEST.json, an optional requirements file and an optional model config YAML. You can choose the archive format with --archive-format: the default, tgz, no-archive which writes a plain directory, or zip-store. The no-archive form is handy in development because you can edit the handler in place, and the default is the one to ship.

Pin the version in the archive with --version, and treat it as part of the model's identity. Two archives with the same name and version but different weights are a recipe for confusion, because a registration of the second returns a 409 ConflictStatusException while the first is loaded. Use -f (--force) only for local overwriting of the output file, not as a release process.

BASH
torch-model-archiver \
  --model-name fraud \
  --version 2025.03.1 \
  --serialized-file build/fraud.pt \
  --model-file src/model.py \
  --handler src/handler.py \
  --extra-files src/labels.json \
  --requirements-file requirements.txt \
  --config-file model-config.yaml \
  --export-path model_store

Dependencies

Python packages that your handler needs go in requirements.txt, passed with --requirements-file. TorchServe installs them into the worker environment only when install_py_dep_per_model=true is set. If it is not set, the requirements file is silently not used, and you see ModuleNotFoundError in the backend log at worker start. This is the official troubleshooting entry for missing Python packages, and it is the most common reason a model that worked on your laptop fails in the container.

Installing at startup has a cost and a risk. It slows every start, it needs network access to a package index, and it means the running service is not the thing you tested. For production, the better practice is the opposite: bake the dependencies into the container image, leave install_py_dep_per_model at its default of false, and treat the requirements file as documentation of what the image must contain.

What the archive trusts

The documentation is blunt: TorchServe executes the arbitrary Python code packaged in the .mar file. Loading an archive is the same act as running a stranger's script. Combined with the management API's ability to register a model from a URL, this is the main attack surface of the project, and it is why the next sections lock registration down.

Keep weights and code releasable on their own When archives are large, rebuilding one for a one-line handler fix is slow. Keep the handler small and test it as ordinary Python so that a fix is a quick rebuild, and keep large reference data in --extra-files only if it genuinely belongs to a specific model version.
Try it
  1. Script your archiver command in a Makefile target or CI step that takes the version from a git tag.
  2. Unpack your .mar with unzip -l if your archive is zip-compatible, or tar -tzf for the default format and read MANIFEST.json.

Configuration without surprises

TorchServe is configured by a config.properties file, by command-line flags and by environment variables, and mixing them is where confusion starts.

The file

You point the server at a file with --ts-config. The important keys, with their documented defaults where they have one:

config.properties
inference_address=http://127.0.0.1:8080
management_address=http://127.0.0.1:8081
metrics_address=http://127.0.0.1:8082
job_queue_size=100
default_workers_per_model=2
load_models=all
model_store=/models
install_py_dep_per_model=false
enable_metrics_api=true
metrics_mode=prometheus
allowed_urls=file://.*
max_request_size=65535000
max_response_size=65535000

The values for max_request_size and max_response_size here are examples chosen for larger payloads. The official troubleshooting page says uploads of around 6.5 MB fail with the defaults, and the fix is to raise both. Raise them deliberately and to the size you actually need, since they are also a cheap defence against oversized requests.

Per-model settings in the file

Instead of registering models through the API, you can declare them in a models property holding JSON. This is the production-friendly path because the whole serving state is in one reviewable file:

PROPERTIES
load_models=standalone
models={\
  "fraud": {\
    "2025.03.1": {\
      "defaultVersion": true,\
      "marName": "fraud.mar",\
      "minWorkers": 2,\
      "maxWorkers": 2,\
      "batchSize": 8,\
      "maxBatchDelay": 20,\
      "responseTimeout": 60\
    }\
  }\
}

The shape is model name, then version, then settings. minWorkers and maxWorkers set the worker count, and setting them equal gives you a fixed pool with no scaling surprises. The same file can set batchSize, maxBatchDelay and responseTimeout. Because this is an ordinary properties file, the trailing backslashes continue the line.

Precedence

Here the documentation is not perfectly consistent, and you should know that rather than trust a rule of thumb. The server documentation states that command-line arguments override the config file. The configuration page describes environment variables as having higher priority than the command line or the file. The token authorization page states the order command line, then environment variable, then config file. The safe approach is to avoid depending on the ordering at all: set each key in exactly one place. Use the file for everything static, use environment variables only for values injected at deploy time such as secrets, and never set the same key two ways. If a setting seems ignored, search every layer for a second definition before suspecting a bug.

Environment variables

Setting enable_envvars_config=true allows properties to be overridden by environment variables named TS_ followed by the upper-case property name, for example TS_INFERENCE_ADDRESS. This is how containers get per-environment configuration without baking different files into images, and it matters for secrets, which should not sit in a file in the image. A few properties have their own documented variable names too, such as TS_METRICS_MODE and TS_DISABLE_TOKEN_AUTHORIZATION.

Snapshots

TorchServe records the models and worker counts it has registered into snapshot files so it can restore them after a restart, and by default it does so. This is convenient and also a trap, because the state you get at restart is the last snapshot, not your config file. The documented symptom is an InvalidSnapshotException after a restart, with a fix of removing the snapshot files (under the log location's config directory, falling back to ./log/config), starting with --ts-config, or disabling snapshots with --ncs. For containers, which should be disposable and declarative, pass --ncs so that every start reads only what you configured.

Stale snapshots hide config changes If you edit your model settings and the server behaves as before, check whether it restored a snapshot. Declarative deployments should start with --ncs so the file is the only source of truth.
Try it
  1. Move one model from an API registration call into the models property, start with --ncs, and confirm with GET /models/<name> that workers and batch settings match.
  2. Set TS_INFERENCE_ADDRESS with enable_envvars_config=true and verify it takes effect.

The management API, and why you should rarely call it

The management API on 8081 is how you register, scale, inspect and remove models at runtime. Beginner may have used it freely. Since version 0.11.1 two defaults changed that matter greatly: token authorization is enforced, and registering or deleting models through HTTP or gRPC calls is disabled. Old tutorials that run curl -X POST localhost:8081/models?url=... fail on a current server, and the usual bad advice online is to switch both protections off.

A mid-level engineer keeps the protections on and reads the management API for what remains useful:

  • GET /models lists models, with pagination through limit and next_page_token.
  • GET /models/{name} describes a model: its version, runtime, minWorkers, maxWorkers, batchSize, maxBatchDelay, a workers array with each worker's id, start time, status, GPU and memory usage, and jobQueueStatus with remainingCapacity and pendingRequests.
  • GET /models/{name}/all returns every version, and ?customized=true includes custom metadata.
  • PUT /models/{name}?min_worker=3&synchronous=true scales workers. Without synchronous it returns 202 and does the work in the background, and with it the call blocks and returns 200 once done.
  • PUT /models/{name}/{version}/set-default changes which version unversioned requests reach.

The describe call is your best live diagnostic. pendingRequests against remainingCapacity tells you whether the queue is filling, and the worker list tells you which workers are alive and how much memory each holds, without needing shell access.

Preload instead of register

In production the clean pattern is to treat the model set as immutable at start. List the models in the models property or with --models, set load_models, disable the model API, and roll out a new model by rolling out a new container. That gives you what runtime registration never will: the thing that ran in staging is byte for byte the thing in production, there is no state to restore, and no network path exists that can load new code.

If you really do need runtime registration, for a research cluster or a notebook environment, enable it on purpose using the model API setting, using the model API flag listed by torchserve --help for your version, keep the management address on localhost or a private interface, and require the management token.

Try it
  1. Call GET /models/<name> while load testing and watch pendingRequests change.
  2. Run torchserve --help and write down the exact flag names for your installed version in your runbook.

Securing a server that will not be patched

Security deserves its own section here because the project says plainly that no security patches are coming, and its security page lists only 0.11.1 as supported. You cannot count on upstream to close a hole, so your controls must assume the software has holes. The aim is layers: reduce what the server can do, reduce who can reach it, and limit what it can touch if something goes wrong.

Bind addresses

By default all three HTTP ports and both gRPC ports listen on localhost. That is a safe default that people undo within the first week. Setting an address to 0.0.0.0 exposes that API on every interface. The inference port may need to be reachable from other pods or services, so bind it to the interface it needs. The management port and gRPC management port 7071 should stay on localhost or a private management network. In Docker, remember that the image publishes 8080 to 8082 and 7070 to 7071, and that -p 8081:8081 publishes to all host interfaces. Bind explicitly, as in -p 127.0.0.1:8081:8081.

Token authorization

Since 0.11.1 the server enforces token authorization on its APIs. When it starts, it creates a key_file.json in the working directory containing a management key, an inference key and an API key. The two scoped keys are used as bearer tokens, for example -H "Authorization: Bearer [INFERENCE_KEY]". The API key is for minting new tokens through GET /token?type=management or type=inference, also as a bearer. Tokens expire after 60 minutes by default, controlled by token_expiration_min.

Read what this gives you and what it does not. It gives you two scopes, so a client holding an inference token cannot call management. It does not give you per-user identity, per-tenant quotas or revocation lists. It is a shared-secret scheme, closer to an API key than to OAuth. For real client authentication, put an API gateway or service mesh in front, let it authenticate callers, and let it hold the inference token to forward on. A frequent bug is a client that works for an hour and then fails with an unauthorised error, because it cached a token and never re-minted it. Make your calling code refresh on a 401.

key_file.json is a live credential file. Do not commit it, do not bake it into an image, and be careful about the working directory it lands in when the container runs. Mount a volume you control, or generate and consume the keys as part of your startup, and keep them out of logs.

Switching token authorization off is possible with disable_token_authorization=true, the --disable-token-auth flag, or the environment variable TS_DISABLE_TOKEN_AUTHORIZATION when environment configuration is enabled. The documentation positions this as a development convenience. If a pod must run without it, the pod must sit behind something that authenticates, with a network policy that allows nothing else in.

Controlling where models come from

The allowed_urls property is a list of regular expressions that restrict which URLs the management API may load models from. The default allows file://.* and http(s)?://.*, which is to say anywhere. A failed fetch produces DownloadModelException with HTTP 400, which is also what you see when your own URL is blocked. Tighten it to your own storage:

PROPERTIES
allowed_urls=https://models\.example\.com/.*

Even if you preload and keep the model API disabled, setting this is cheap defence in depth, since it limits the damage if someone later turns registration on.

Containing the blast radius

Because the handler and the archive are trusted code and the engine itself will not be patched, harden the runtime:

  • Run as a non-root user and drop capabilities.
  • Use a read-only root filesystem, with writable emptyDir mounts for the log and temporary directories only.
  • Apply a network policy that allows ingress to the inference port from known callers and egress only to what the model really needs, and nothing at all to the metadata service on clouds.
  • Scan the image regularly with a tool such as Trivy, and keep a written record of accepted findings. You will see findings you cannot fix, since no upstream fix exists, and the record is how you show they were considered.
  • Pull secrets from your secret manager, for instance Vault, at runtime.

Data residency matters here too. If your employer in the Gulf or Egypt processes personal data under a national regulation, an inference service that logs request payloads is a data store. Check what your handler and your access logs write, and keep the service in a region and a network zone that match the data's obligations.

The two ports that matter most Management on 8081 and gRPC management on 7071 can change what code runs. A cloud security group that exposes either to the internet is a serious incident waiting for a scanner to find it.
Try it
  1. From a second machine, try to reach your server's management port. It should time out or be refused.
  2. Call an inference endpoint without a token, then with an expired one, and check your client code handles the failure by re-minting.

Metrics and observability

A model server that you cannot see into is one you will debug by guessing. TorchServe has a metrics system with three modes, set by metrics_mode: log, the default, which writes metrics to log files, prometheus, which exposes them for scraping, and legacy, which keeps the pre-0.8 format and which you should avoid for new work.

Log mode and Prometheus mode

In log mode, frontend metrics go to ts_metrics.log and backend metrics to model_metrics.log under the log directory. That is workable if you run a log shipper and extract values from lines, and it is how many existing deployments work. In Prometheus mode the metrics are served from the metrics API on port 8082, conventionally at /metrics (confirm the path against your running server with curl). The setting can also come from TS_METRICS_MODE, and a metrics_config property points to a metrics.yaml that defines which metrics exist and their Prometheus types: counter, gauge or histogram. Setting model_metrics_auto_detect brings backend model metrics in without declaring each one, and it is off by default.

For a service that runs on Kubernetes, use Prometheus mode and scrape it with Prometheus, then chart it in Grafana.

Which metrics to watch

The documentation lists a long set of defaults. A small subset answers most questions:

Question Metric
Are requests failing? Requests2XX, Requests4XX, Requests5XX
How much traffic? ts_inference_requests_total
End-to-end latency inside the server? ts_inference_latency_microseconds
Do I have enough workers? ts_queue_latency_microseconds, QueueTime
How slow is my model itself? HandlerTime, PredictionTime (backend)
Are workers healthy or churning? WorkerLoadTime, WorkerThreadTime
Is the host under pressure? CPUUtilization, MemoryUtilization, GPUUtilization, GPUMemoryUsed

The difference between inference latency and queue latency is the diagnostic that matters. If ts_queue_latency_microseconds grows while PredictionTime stays flat, the model is fine and you are short of workers or batch capacity. If PredictionTime rises, something about the model, the inputs or the hardware changed.

Custom metrics

Anything you record through context.metrics in a handler can show up here. Good candidates are input size, number of items per request, a counter of fallback paths and a measure of confidence. With Prometheus mode and a metrics.yaml, define them there with a type so dashboards behave. Keep dimensions to a few low-cardinality labels such as model name and a coarse category. Putting a user ID or request ID in a label creates one time series per value and will hurt your Prometheus server.

Try it
  1. Start the server with metrics_mode=prometheus and run curl localhost:8082/metrics. Find the queue latency metric.
  2. Add one custom metric from a handler and confirm it appears.

Containers and Kubernetes

The reference path for running TorchServe in production is a container. The maintained images are pytorch/torchserve with a -gpu variant. Since the project is archived, latest is a moving label that will stop moving, and you should pin an explicit tag or, better, a digest, as the Docker guide recommends for any dependency. Tag names should be confirmed on Docker Hub, and the safest practice is to build your own image from the repository's build_image.sh script and push it to your registry, so that you control what the base contains.

Production run flags

The Docker documentation's production-style invocation adds --shm-size=1g, --ulimit memlock=-1 and --ulimit stack=67108864. The shared memory flag matters because PyTorch worker processes exchange tensors through /dev/shm, and Docker's default of 64 MB is too small, which produces odd crashes under load rather than a clear error. Models are mounted from a bind mount or a volume and the server is started with --model-store pointing at it:

BASH
docker run --rm --shm-size=1g --ulimit memlock=-1 --ulimit stack=67108864 \
  -p 127.0.0.1:8080:8080 -p 127.0.0.1:8081:8081 -p 127.0.0.1:8082:8082 \
  --mount type=bind,source=/srv/models,target=/tmp/models \
  my-registry/torchserve:0.12.0-cpu \
  torchserve --model-store=/tmp/models --ts-config /tmp/models/config.properties --ncs

A Kubernetes deployment

On Kubernetes you deploy this as a standard Deployment behind a Service. A sketch of the parts that are specific to TorchServe:

deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: fraud-ts
spec:
  replicas: 3
  selector:
    matchLabels: { app: fraud-ts }
  template:
    metadata:
      labels: { app: fraud-ts }
    spec:
      containers:
        - name: torchserve
          image: my-registry/torchserve@sha256:REPLACE_WITH_DIGEST
          args: ["torchserve", "--model-store", "/models", "--ts-config", "/config/config.properties", "--ncs"]
          ports:
            - { name: inference, containerPort: 8080 }
            - { name: metrics, containerPort: 8082 }
          readinessProbe:
            httpGet: { path: /ping, port: inference }
            initialDelaySeconds: 20
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /ping, port: inference }
            initialDelaySeconds: 120
            periodSeconds: 20
            failureThreshold: 3
          resources:
            requests: { cpu: "2", memory: 6Gi }
            limits: { memory: 6Gi }
          volumeMounts:
            - { name: shm, mountPath: /dev/shm }
            - { name: config, mountPath: /config }
      volumes:
        - name: shm
          emptyDir: { medium: Memory, sizeLimit: 1Gi }
        - name: config
          configMap: { name: fraud-ts-config }

Points that deserve explanation:

  • Probes on /ping. The documented behaviour is that /ping returns 200 when the active workers are at least minWorkers, and 500 otherwise. That makes it a good readiness probe, because a pod does not receive traffic until the model is actually loaded. It makes it a risky liveness probe if the delay is short, because a slow model load will look like a dead container and be killed in a loop. Give liveness a generous initial delay, longer than your worst load time.
  • Memory. Size the request and limit from the workers x per-worker footprint calculation earlier, with headroom. The kernel kills a container that exceeds its limit and the worker death looks like a mysterious restart.
  • Shared memory. A memory-backed emptyDir at /dev/shm replaces the 64 MB default, the Kubernetes equivalent of --shm-size.
  • Model delivery. Either bake the .mar into the image, which gives the most reproducible deployment, or pull it with an init container from object storage into an emptyDir. Baking is simpler and matches the immutable-model approach above.
  • Scaling. TorchServe has no built-in autoscaling of replicas. Scale with a Horizontal Pod Autoscaler driven by a custom metric, such as queue latency exposed through Prometheus, or by request rate. Scaling on CPU alone is misleading for GPU-bound models.
  • Graceful rollout. Make sure rolling updates respect readiness, set a terminationGracePeriodSeconds longer than your longest request, and consider a preStop pause so the Service stops sending traffic before the process receives its stop signal.

For GPU nodes, request nvidia.com/gpu in limits and set workers per model to match the GPUs you assign, since the default worker count follows the GPU count.

Try it
  1. Run your container locally with the production flags and kill it with docker stop. Time how long it takes and check the logs for an orderly shutdown.
  2. Set a readiness probe on /ping in a test cluster and watch the pod stay unready until the model loads.

Testing and load testing

A serving stack needs three layers of tests, and the cheapest layer catches most defects.

Handler unit tests run the handler's methods directly, with no server, as shown earlier. They check shapes, batch lengths, error isolation and postprocessing. They run in milliseconds and belong in every pull request.

Archive and server smoke tests build the .mar in CI, start TorchServe with the exact production configuration, wait for /ping to return healthy, send a known input and compare the output with a stored expected value within a tolerance. This catches a missing --extra-files entry, a requirements problem or a wrong handler path, which unit tests cannot. Run it in the same container image you deploy, so that the Java, Python and PyTorch versions are the real ones. A short script is enough:

smoke.sh
#!/usr/bin/env bash
set -euo pipefail
torchserve --start --ncs --model-store model_store --models fraud=fraud.mar \
  --ts-config config.properties
for i in $(seq 1 60); do
  curl -sf http://127.0.0.1:8080/ping >/dev/null && break
  sleep 2
done
curl -sf -X POST http://127.0.0.1:8080/predictions/fraud \
  -H "Authorization: Bearer ${INFERENCE_KEY}" \
  -H "Content-Type: application/json" \
  -d @tests/sample.json | tee out.json
python tests/compare.py out.json tests/expected.json
torchserve --stop

If token authorization is on in CI, read the inference key from the generated key_file.json rather than printing it in the log.

Load tests answer the questions unit tests cannot: how many requests per second at what latency, where the queue starts to grow, and what the batch settings really do. Use any HTTP load tool you know, such as hey, wrk, k6 or Locust, and be disciplined about how you measure.

  1. Warm up first. The first requests pay for lazy initialisation. Discard them.
  2. Test at your real concurrency and payload size. A load test with a tiny payload hides decode and transfer cost.
  3. Ramp load in steps and watch queue latency and pendingRequests. The knee of the curve, where latency climbs while throughput is flat, is your capacity.
  4. Record p50, p95 and p99, not the mean. Batching trades the median for throughput, and the tail tells you whether it was worth it.
  5. Change one thing at a time: worker count, then batch size, then delay.

The official performance guide exists and is worth reading for the measurement tooling it describes, and large-model topics have their own page. Treat published numbers as shape, not promises, because your model and hardware decide the result.

The most useful single experiment for a GPU service is to sweep the number of workers per GPU. Because workers are separate processes, a second worker on the same GPU can fill gaps while the first does CPU-side preprocessing, and often raises throughput, until GPU memory or compute saturates. Past that point you only add memory pressure and context switching. Find the number by measurement.

Put the load test in the pipeline, small A one-minute test at fixed concurrency against a staging container, with a threshold on p99, catches a regression that a functional test never will, such as a handler change that quietly doubled preprocessing time.
Try it
  1. Add the smoke script to CI, using the production image and config.
  2. Sweep workers 1, 2 and 4 at one fixed concurrency and graph throughput and p99 for each.

Debugging beyond the basics

When something is wrong, resist the urge to restart. A fixed order of questions finds causes quickly.

1. Is the process up, and are the workers? Check GET /ping. A 500 means fewer active workers than minWorkers. Then GET /models/<name> and read the workers array. An empty list, or workers cycling in status, means they are dying during startup.

2. If workers die, read the backend log, not the frontend output. The message Backend worker monitoring thread interrupted is a symptom. The cause, an import error, a wrong path, a missing file from --extra-files, a CUDA out-of-memory error, is in the worker's log under the logs/ directory. File names such as ts_log.log, model_log.log and access_log.log are the ones the project's layout uses, so check your own log directory to confirm.

3. Map the HTTP status to a stage.

Status or exception Meaning First action
404 ModelNotFoundException The .mar is not in the store, or the name is wrong List the model store and check GET /models
409 ConflictStatusException That name and version is already registered Change the version or unregister the old one
400 DownloadModelException URL unreachable or not allowed Check allowed_urls and network access
503 ServiceUnavailableException No workers, or the queue is full Check workers, then pendingRequests, then scale
401 or 403 Token missing, wrong scope or expired Re-mint, and check which key you used
Uploads over about 6.5 MB fail Default request size limit Raise max_request_size and max_response_size

4. Port conflicts. Failed to bind to address: http://127.0.0.1:8080 means another process holds the port, often an earlier TorchServe that was never stopped. Use torchserve --stop, check with lsof -i :8080, or move the address in the config.

7. Slow only under load. Compare queue latency and prediction time. High queue latency means too few workers or a batch delay that is too long. High prediction time under load but not alone suggests contention: workers competing for a GPU, memory swapping, or CPU throttling from a low Kubernetes CPU limit. A container with a CPU limit far below the number of workers will stall in bursts even though the average looks fine.

8. Memory growth. Watch the memoryUsage in the describe output per worker over hours. A steady climb in one worker points to a leak in your handler, such as accumulating tensors in a list on self, or holding the autograd graph by forgetting torch.no_grad() or inference mode. BaseHandler runs inference without gradients, so a leak is more often in code you added.

To reproduce a failure on a single input outside the server, instantiate the handler in a Python shell with a fake context, as in the unit test, and call the stages one at a time. Most handler bugs show up in the first stage that touches the data.

Try it
  1. Break a handler on purpose with a bad import, start the server, and find the real error in the backend log.
  2. Register the same model and version twice and read the 409.

Planning your way off TorchServe

Because there is no future release, a responsible mid-level engineer keeps a migration plan, even if it is a short document, and does the preparation now, while nothing is on fire.

Decouple the contract from the engine. Callers should depend on a stable internal URL and request schema, not on :8080/predictions/fraud. Put a thin gateway or a service name in front, so that a replacement can be introduced behind it. If your callers already use the KServe-compatible routes, you can move to a platform that speaks that protocol with less client change.

Pick a destination by need:

  • You want Python-first packaging and an API you control: BentoML or a FastAPI service for small cases.
  • You need composition of several models, or autoscaling with Python: Ray Serve.
  • You run on Kubernetes and want a standard serving abstraction: KServe.
  • You need high GPU utilisation across frameworks: Triton.
  • You serve language models: vLLM, which TorchServe 0.12.0 itself integrated for that reason.
Try it
  1. List every caller of your service and the exact URL each uses. That list is your migration scope.
  2. Move the preprocessing code from your handler into a module with its own tests, and import it back.

Putting it all together

Here is a small but complete production-shaped deployment, assembled from the pieces above. The scenario is a binary scoring model served as fraud.

  1. Code. The handler is a BaseHandler subclass whose preprocessing and postprocessing live in a separate tested module. Unit tests assert that N requests give N results and that a bad request yields one error entry.
  2. Build. CI builds the .mar from a tagged commit with the archiver, versioned by the tag, and stores it with a checksum.
  3. Image. A Dockerfile starts from a pinned base, installs the dependencies directly, copies the .mar into /models and the config.properties into /config, and runs as a non-root user.
  4. Config. The properties file preloads the model through the models property, with explicit minWorkers and maxWorkers, a batchSize and maxBatchDelay chosen from load tests, and a responseTimeout and startup timeout that match reality. Metrics are in Prometheus mode, allowed_urls is restricted, the model API stays off, token authorization stays on and addresses are bound deliberately. The server starts with --ncs.
  5. Smoke test. The pipeline starts the image, waits for /ping, scores a known sample and checks the answer, then runs a one-minute load test with a p99 threshold.
  6. Deploy. A Kubernetes Deployment pins the image digest, mounts a memory-backed /dev/shm, uses /ping for readiness with a patient liveness probe, and has requests and limits derived from the worker arithmetic. A network policy keeps management traffic out, and a gateway holds the inference token.
  7. Observe. Prometheus scrapes port 8082. Dashboards show request rate, 5XX rate, queue latency and prediction time per model. Alerts fire on sustained queue latency and worker restarts.
  8. Plan. The runbook lists the pinned versions, the accepted vulnerabilities, the controls that compensate for them and a dated plan to migrate, with the contract decoupled behind a gateway.

Nothing in that list is exotic. The skill at this level is doing every item deliberately, and knowing which of them breaks first when load, a dependency or a person changes.

Where you are now

You can now explain how TorchServe divides work between a Java frontend and Python workers, and use that to size memory and predict failures. You can follow a request through the queue and tune batching against a latency budget. You write handlers that are safe for batches, test them without a server, and package them in archives that you can rebuild. You configure the server from one source of truth, keep the management API closed, preload your models, and secure the service with token authorization, restricted URLs, private ports and a hardened container. You can observe it with Prometheus, load test it, debug it by a fixed order of questions, and plan the move to something that is still maintained.

The senior level takes the next step: custom handler internals and model configuration for large models, running TorchServe as a platform for several teams, the security and supply chain questions in depth, and the judgement about when to stop running it at all. Until then, the most valuable thing you can do is write your migration plan while the service is calm.

Sources