Skip to content
Back to student guides
Triton Inference ServerMLOpsModel serving3 levels88 sectionsCovers Triton Inference Server 2.72 (container 26.08)

The Complete Triton Inference Server Guide

Serve models from any framework on GPUs and CPUs with NVIDIA Triton. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
13sections
33examples

This is part one of three. It takes you from never having heard of Triton Inference Server to serving a model over HTTP, sending it requests from Python, reading its metrics, and fixing the errors that stop most newcomers. By the end you will have a working model repository on your laptop, you will know what every line of its configuration means, and you will be able to explain to a colleague why a team would put Triton between their trained model and their users. Mid-level and Senior take the same ground further; nothing in this part is a teaser.

The guide was checked against Triton Inference Server 2.72.0, shipped as container release 26.08. Triton releases monthly and tags its containers by year and month, so by the time you read this a newer tag may exist. Everything taught here is stable across releases, but always pin the tag you tested with rather than guessing.

Each section ends with a Try it task. Do them as you go. The ideas here stick only once you have watched your own server start, reject a bad request, and start again properly.

What Triton is, and the problem it solves

You have trained a model. On your laptop it is a file, perhaps model.onnx or model.pt, plus a notebook cell that loads it and calls it on some numbers. That is fine for you. It is not fine for a mobile app, a website backend, or another team's batch job, because none of them can import your notebook. They need a network service: something they can send a request to and get a prediction back from, reliably, many times a second.

The obvious first attempt is to write that service yourself. You wrap the model in a small web application, load the file at startup, and call it inside a request handler. This works, and for a lot of projects it is the right answer; the FastAPI guide on this site (/student-guides/fastapi) covers exactly that. But as soon as the traffic grows, the same set of problems appears every time:

  • The GPU sits idle. A GPU is fast when it processes many inputs at once and wasteful when it processes one. A hand-written handler takes one request, runs it, returns it, and takes the next. The expensive hardware spends most of its time waiting.
  • Every framework needs its own glue. One team has a PyTorch model, another has ONNX, a third has a TensorRT engine, a fourth has a gradient-boosted tree. Four formats mean four services with four sets of conventions for health checks, metrics, and error handling.
  • Operations are rebuilt every time. Loading and unloading models, running two versions side by side, exposing metrics, limiting concurrency: each of these is code someone has to write and then maintain.

Triton Inference Server is NVIDIA's answer to those three problems. It is a ready-made server program, not a library. You do not write a web application around your model. You describe your model in a small configuration file, put it in a folder with a particular layout, and start Triton pointing at that folder. Triton loads the model, opens network ports, and answers requests. Along the way it does the things you would otherwise build by hand: it can group requests from different clients into one batch for the GPU, run several copies of a model at once, serve many different model formats from the same process, and publish metrics that a monitoring system can scrape.

YOUR MODELa trained file
→
CONFIGconfig.pbtxt
→
TRITONloads and serves
→
CLIENTSHTTP or gRPC

The diagram is the whole idea. Everything else in this guide is detail about one of those four boxes.

A word on the name, because it confuses people. Triton was originally called the TensorRT Inference Server, and you will still see old blog posts using that name. It now serves many kinds of models, not only TensorRT ones, and the name was changed accordingly. Do not confuse it with the Triton language and compiler for writing GPU kernels, which is an unrelated project that happens to share the name.

It also helps to know where Triton sits among its neighbours. If you want a general-purpose Python service where you control every line, FastAPI is the lighter choice. If your whole team lives in PyTorch and wants a PyTorch-specific server, there is TorchServe (/student-guides/torchserve). If you are deploying large language models and want that single job done as well as possible, vLLM (/student-guides/vllm) and TensorRT-LLM (/student-guides/tensorrt-llm) are specialised engines, and Triton can host both of them as backends. Triton's particular strength is being the one server that speaks to all of these formats at once, which is why managed cloud services such as SageMaker (/student-guides/sagemaker) and Vertex AI (/student-guides/vertex-ai) use it as a serving runtime.

For readers in the Gulf and Egypt, one practical point deserves mention early. Triton runs wherever you can run a container, which means it runs inside your own data centre or in a cloud region of your choosing. Employers in finance, government, and healthcare often require that model inputs, which can be customer data, never leave the country. A self-hosted inference server satisfies that requirement in a way a third-party prediction API cannot.

Try it
  1. Think of one model you have trained or used, in any framework.
  2. Write down who would need to call it, from what kind of program, and roughly how many times per second.
  3. Decide whether a single process handling one request at a time could keep up. That gap is the problem Triton exists for.

The mental model: four nouns

Triton has a lot of features, but a beginner needs only four ideas. Learn these and the documentation stops looking like a wall.

The model repository. This is a folder, on local disk or in cloud storage, that holds everything Triton should serve. Each model gets its own subfolder, named after the model. Inside that subfolder sits one configuration file and one or more numbered version folders containing the actual model file. The folder layout is not a suggestion; Triton finds your model by reading the structure, so spelling and nesting matter.

The backend. A backend is the piece of code inside Triton that knows how to run one kind of model. There is a backend for ONNX Runtime, one for TensorRT, one for PyTorch's TorchScript format, one for OpenVINO, one for tree-based models such as XGBoost, and one that runs arbitrary Python code that you write. When you tell Triton "this model uses the onnxruntime backend", you are telling it which engine to hand the model file to. Triton itself does not understand your neural network; the backend does.

The model configuration. Each model has a file called config.pbtxt. It is written in a text format called protobuf text, which looks a little like JSON without commas between fields. It declares the model's name, which backend runs it, what inputs it expects (names, data types, shapes), what outputs it returns, and how Triton should run it (how many copies, whether to batch). The configuration is the contract between your model and the world: clients must send inputs exactly as it describes.

The instance. A running copy of a model is called an instance. By default Triton makes one instance per model on each GPU, or one on the CPU if there is no GPU. You can ask for more. Because each instance can work on a request independently, two instances can process two requests at the same moment.

ClientsHTTP on 8000, gRPC on 8001
Triton serverchecks, queues, batches
BackendONNX Runtime, TensorRT, PyTorch, Python
Model repositoryconfig.pbtxt and model files

Two more terms will come up. A protocol is the language clients use to talk to the server. Triton speaks HTTP with JSON bodies and gRPC, a faster binary protocol, and both follow a community specification called the KServe v2 inference protocol. This matters because it means a client written for one KServe-compatible server largely works against another. A tensor is simply a multi-dimensional array of numbers with a fixed data type, and it is the unit Triton sends and receives. Your image is a tensor of shape three by 224 by 224. Your sentence, after tokenising, is a tensor of integers. When Triton documentation says "input", it means a named tensor.

Finally, remember three port numbers, because you will type them constantly. Port 8000 serves HTTP, port 8001 serves gRPC, and port 8002 serves metrics in a format the Prometheus monitoring system understands.

Tip. When something goes wrong, ask which of the four nouns is at fault. A folder in the wrong place is a repository problem. A wrong file name or a missing engine is a backend problem. A mismatch between what the client sent and what the model expects is a configuration problem. Performance trouble is usually an instance or batching problem.
Try it
  1. Without looking back, name the four nouns and give a one-line definition of each.
  2. Say which noun owns each of these facts: "this model takes a 3 by 224 by 224 float input", "this model file is ONNX", "two copies of this model run at once".

Installing and checking your setup

Triton's supported way to run is as a container image published on NVIDIA's NGC registry, so installing Triton mostly means installing Docker. If containers are new to you, read the Docker guide (/student-guides/docker) first; you only need images, docker run, port mapping, and volume mounts, all of which it teaches.

On Linux with an NVIDIA GPU, install Docker and the NVIDIA Container Toolkit, which lets containers see the GPU. Your host also needs an NVIDIA driver recent enough for the CUDA version inside the container. The 26.08 container is built on CUDA 13.4.1, so check the release notes for the minimum driver, and expect that very old GPUs are no longer supported by current containers. Turing-generation cards (compute capability 7.5) and newer are the safe baseline.

On a Mac or a laptop without an NVIDIA GPU, there is no official native Triton server for macOS. You can still learn everything in this guide, because Triton runs on the CPU too. Run the same container under Docker Desktop or Colima, which gives you a Linux virtual machine. You will not get GPU acceleration, and the tag you pull must have a build for your processor architecture, which is worth confirming on the NGC page when you pull. The Python client library installs and runs natively on a Mac.

On Windows, the practical route is WSL2 (Windows Subsystem for Linux) with Docker, then following the Linux steps. Official native Windows support is limited, so do not plan around it.

Pull the image:

BASH
docker pull nvcr.io/nvidia/tritonserver:26.08-py3

The image is large, several gigabytes, because it bundles CUDA libraries and several inference engines. Expect the first pull to take a while. The -py3 suffix means the full image with all common backends; NVIDIA also publishes slimmer variants for specific uses, and a separate -py3-sdk image that contains client programs and the Perf Analyzer benchmarking tool.

If you are on a machine with a company network that inspects TLS traffic, a pull may fail with a certificate error. That is a trust problem with your network, not with Triton; add your organisation's certificate to your container runtime rather than disabling verification.

Install the Python client library on your own machine too, so you can send requests without entering a container:

BASH
python -m venv .venv
source .venv/bin/activate
pip install "tritonclient[http]" numpy

The square-bracket part chooses which protocols to install. [http] is enough for this guide, [grpc] adds gRPC, and [all] installs both.

Checking the install is easiest if you start the server once with an empty repository, to prove the container runs:

BASH
mkdir -p models
docker run --rm -p 8000:8000 -p 8001:8001 -p 8002:8002 \
  -v "$PWD/models:/models" \
  nvcr.io/nvidia/tritonserver:26.08-py3 \
  tritonserver --model-repository=/models

Break that command down, because you will reuse it all guide. --rm deletes the container when it stops. Each -p publishes a port from the container to your machine. -v "$PWD/models:/models" mounts your local models folder inside the container at /models, so Triton sees your files and you can edit them from outside. The image name and tag come next, and then the command to run inside the container: tritonserver with the flag --model-repository=/models telling it where to look. On a machine with a working GPU setup you would add --gpus=1 right after docker run to give the container one GPU; leave it off to run on the CPU.

With an empty folder, Triton should start, print a table of loaded models that is empty, and keep running. In a second terminal, ask whether it is ready:

BASH
curl -v localhost:8000/v2/health/ready

Look for HTTP/1.1 200 OK in the output. A 200 status means the server is up and every model it was asked to load is ready. There is also /v2/health/live, which only says that the process is running. The distinction matters later: "live" means do not restart it, "ready" means it is safe to send traffic. Stop the server with Ctrl+C in its terminal.

Trap. Using an absolute path for the volume is not optional in older Docker setups, and a relative path such as -v models:/models creates a named Docker volume instead of mounting your folder. Always use "$PWD/models" or a full path so Docker treats it as a directory on your machine.
Try it
  1. Pull the image and start the server against an empty models folder.
  2. Run the health check with curl and confirm you see a 200.
  3. Stop the container, then run curl again and observe the connection failure. Knowing what "down" looks like makes later debugging faster.

Your first model, built step by step

To keep this first project small and free of downloads, we will serve a model written in Python: a function that adds two vectors of four numbers. It is deliberately trivial. The point is not the arithmetic; it is the shape of everything around it, which is identical for a real neural network. Later you will swap the Python logic for an ONNX file and nothing else about the workflow changes.

Create the folder structure. The model is named add_vectors, and it has one version, numbered 1:

BASH
mkdir -p models/add_vectors/1

The result must look like this:

TEXT
models/
  add_vectors/
    config.pbtxt
    1/
      model.py

Three rules about this layout. The folder name under models/ is the model's name, and clients will use it in URLs. The version folder must be an integer; a folder named latest or v1 is silently ignored. And the Python backend expects the file inside the version folder to be called model.py.

Now write the configuration:

models/add_vectors/config.pbtxt
name: "add_vectors"
backend: "python"
max_batch_size: 8

input [
  {
    name: "INPUT0"
    data_type: TYPE_FP32
    dims: [ 4 ]
  },
  {
    name: "INPUT1"
    data_type: TYPE_FP32
    dims: [ 4 ]
  }
]

output [
  {
    name: "OUTPUT0"
    data_type: TYPE_FP32
    dims: [ 4 ]
  }
]

instance_group [
  {
    count: 1
    kind: KIND_CPU
  }
]

Read it line by line. name must match the folder name. backend: "python" selects the backend that runs your model.py. max_batch_size: 8 says Triton may hand this model up to eight requests stacked together; we will explain the consequences of this number in a moment, because it changes how dims is read. Each entry in input gives a tensor name, a data type (TYPE_FP32 is a 32-bit float), and dims, the shape of one example. The output block does the same for what the model returns. Finally instance_group asks for one copy of the model running on the CPU.

That max_batch_size detail is the single most common point of confusion for newcomers, so here it is plainly. When max_batch_size is greater than zero, Triton adds an invisible leading dimension for the batch. You write dims: [ 4 ] meaning "one example is four numbers", and a client sends a tensor of shape [1, 4] for one example or [5, 4] for five. When max_batch_size is zero, batching is off, Triton adds nothing, and your dims must describe the complete shape the client sends, batch dimension included. Mixing these up gives shape errors, covered in the errors section.

Now the Python code. A Python-backend model is a class called TritonPythonModel with up to three methods Triton knows how to call:

models/add_vectors/1/model.py
import numpy as np
import triton_python_backend_utils as pb_utils


class TritonPythonModel:
    def initialize(self, args):
        # Runs once when the model loads. Load weights or open
        # resources here, not inside execute().
        self.model_name = args["model_name"]

    def execute(self, requests):
        # Runs for each group of requests Triton hands us.
        # We must return exactly one response per request, in order.
        responses = []
        for request in requests:
            a = pb_utils.get_input_tensor_by_name(request, "INPUT0").as_numpy()
            b = pb_utils.get_input_tensor_by_name(request, "INPUT1").as_numpy()
            out = pb_utils.Tensor("OUTPUT0", (a + b).astype(np.float32))
            responses.append(pb_utils.InferenceResponse(output_tensors=[out]))
        return responses

    def finalize(self):
        # Runs once when the model unloads. Free anything you opened.
        pass

The module triton_python_backend_utils exists only inside the Triton container; you cannot import it on your laptop, so do not worry if your editor underlines it. initialize is where expensive setup belongs, because Triton calls it once, not per request. execute receives a list of requests, and the rule that trips people up is that you must return a list of responses with the same length and order. finalize is cleanup.

Notice what the code does not contain: no web framework, no ports, no JSON parsing, no health endpoint. Triton supplies all of that. Your code is only the model.

Start the server, this time with the folder populated:

BASH
docker run --rm -p 8000:8000 -p 8001:8001 -p 8002:8002 \
  -v "$PWD/models:/models" \
  nvcr.io/nvidia/tritonserver:26.08-py3 \
  tritonserver --model-repository=/models

Watch the startup output. Near the end Triton prints a table with columns for model, version, and status. You are looking for this line:

TEXT
| add_vectors | 1       | READY  |

READY means the model loaded and can serve. If you see UNAVAILABLE with a reason in the last column, the reason is the thing to read; later sections explain the usual ones. Triton also prints the HTTP and gRPC endpoints it started and the port the metrics endpoint listens on. In the second terminal check the model itself:

BASH
curl localhost:8000/v2/models/add_vectors/ready
curl localhost:8000/v2/models/add_vectors

The first returns an empty 200 response when the model is ready. The second returns a JSON description, including the name, the versions, and the input and output tensors as Triton understands them. Compare it with your config.pbtxt; seeing your own configuration echoed back is a good confirmation that Triton read what you meant.

Note. Triton can fill in some of a configuration automatically. For ONNX, TensorRT, and OpenVINO models it can read the input and output definitions from the model file itself, so a very small config is enough. For Python and PyTorch models there is no such file metadata, which is why we wrote everything by hand. Even when auto-complete works, writing the config out explicitly is a good habit, because the file then documents your service.
Try it
  1. Build the folder, config.pbtxt, and model.py exactly as shown and start the server.
  2. Find the add_vectors row in the startup table and confirm it says READY.
  3. Run both curl commands above and compare the JSON against your config.

Sending your first request

A server is only useful once someone calls it. Triton's HTTP interface is plain JSON, so you can use curl. A request to run inference is a POST to /v2/models/<name>/infer, with a body listing each input by name, shape, data type, and flattened data:

BASH
curl -s -X POST localhost:8000/v2/models/add_vectors/infer \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": [
      {"name": "INPUT0", "shape": [1, 4], "datatype": "FP32", "data": [1, 2, 3, 4]},
      {"name": "INPUT1", "shape": [1, 4], "datatype": "FP32", "data": [10, 20, 30, 40]}
    ]
  }'

Notice the details. The shape is [1, 4], one example of four numbers, because batching is enabled and the first dimension is the batch. The data type is written FP32, without the TYPE_ prefix used in config.pbtxt; the config uses the protobuf enum names and the wire protocol uses the short names. The data is a flat list in row-major order, and its length must equal the product of the shape.

The response looks like this, with some fields trimmed:

JSON
{
  "model_name": "add_vectors",
  "model_version": "1",
  "outputs": [
    {"name": "OUTPUT0", "datatype": "FP32", "shape": [1, 4], "data": [11.0, 22.0, 33.0, 44.0]}
  ]
}

That is your model, answering over the network. You wrote about twenty lines of Python and a config file.

Sending JSON from the shell is awkward for anything real, particularly images or large arrays, so most programs use the client library. Here is the same call in Python:

client.py
import numpy as np
import tritonclient.http as httpclient

client = httpclient.InferenceServerClient(url="localhost:8000")

a = np.array([[1, 2, 3, 4]], dtype=np.float32)
b = np.array([[10, 20, 30, 40]], dtype=np.float32)

inputs = [
    httpclient.InferInput("INPUT0", a.shape, "FP32"),
    httpclient.InferInput("INPUT1", b.shape, "FP32"),
]
inputs[0].set_data_from_numpy(a)
inputs[1].set_data_from_numpy(b)

result = client.infer("add_vectors", inputs)
print(result.as_numpy("OUTPUT0"))

Run it with python client.py and you should see [[11. 22. 33. 44.]]. Look at how the code maps to the concepts. An InferInput is one named input tensor, built from a name, a shape, and a data type string. set_data_from_numpy loads your array into it, and it checks that the array's dtype agrees with the declared data type, which is a helpful early error. client.infer sends the request, and as_numpy("OUTPUT0") extracts one output as a NumPy array by name.

The most important point here is a mistake worth making deliberately. Change INPUT1 to be shaped [1, 5] and rerun. The server rejects the request with a message saying the shape of the input does not match what the model expects. Triton validates every request against the configuration before your code ever runs. That is a feature: bad requests never reach your model, and the error tells the caller exactly what was expected.

There is also a gRPC client, tritonclient.grpc, with almost the same API, which connects to port 8001. gRPC is generally a better choice for high request rates and for large tensors, because it is a compact binary protocol. HTTP is easier to debug with curl. Start with HTTP, switch if profiling says to.

Try it
  1. Send the curl request and confirm the output is the element-wise sum.
  2. Run client.py and confirm it prints the same numbers.
  3. Send a batch of three examples by changing the shape to [3, 4] and supplying twelve numbers per input. One request, three answers.
  4. Send nine examples. This exceeds max_batch_size of 8; read the error.

Understanding the model configuration

You have written one config.pbtxt. Now it is worth knowing the full vocabulary, because the configuration file is the control panel for everything Triton does.

Names and types. name identifies the model. Data types are written TYPE_BOOL, TYPE_UINT8, TYPE_INT32, TYPE_INT64, TYPE_FP16, TYPE_FP32, TYPE_FP64, and TYPE_STRING for text and byte strings. Choose the type your model was exported with; a type mismatch between client and config is rejected, not silently converted.

Dimensions and variable sizes. In dims, a value of -1 means "any size on this axis". A text model might declare dims: [ -1 ] for a sequence of token ids of unknown length. You can use -1 on any axis, but be careful: the underlying model must genuinely support variable sizes there, which many exported ONNX graphs only do if they were exported with dynamic axes.

Batching. Batching is how a GPU earns its money, so understand it properly. A GPU runs one large computation much more efficiently than many small ones, so processing eight examples together can cost little more than processing one. max_batch_size declares the largest batch the model accepts. Setting it above zero does two things: it adds the invisible batch dimension we met earlier, and it permits Triton to do the stacking for you. By itself it does not make Triton wait to collect requests; if clients already send batches, they are passed through. To have Triton combine separate requests from separate clients into one execution, you add the dynamic batcher:

PROTOBUF
dynamic_batching { }

With that empty block Triton will form batches opportunistically from whatever requests are queued while the model is busy. You can tune it with a list of preferred_batch_size values and a max_queue_delay_microseconds, the longest Triton will wait for more requests to fill a batch before running what it has. That delay is a deliberate trade of a little latency for more throughput. The documentation's own example shows eight concurrent clients going from roughly 73 to roughly 272 inferences per second after enabling it, though your numbers depend entirely on your model and hardware. Treat that as a hint that it is worth trying, not a promise. For a beginner, the lesson is only that the knob exists and where it lives; you will tune it properly in the mid-level guide.

If your model cannot be batched at all, set max_batch_size: 0. Then dims must include the batch axis and Triton will never stack or split requests.

Instance groups. The instance_group block says how many copies of the model to run and where. Two examples:

PROTOBUF
instance_group [ { count: 2 kind: KIND_GPU } ]
PROTOBUF
instance_group [ { count: 1 kind: KIND_CPU } ]

KIND_GPU places copies on GPUs, KIND_CPU on the CPU. If you omit the block entirely, Triton uses one instance per available GPU, or one CPU instance when there is none. Two instances on the same GPU let it overlap the work of two requests, which helps when a single instance cannot keep the card busy, but each copy consumes GPU memory, so more is not automatically better.

Version policy. By default Triton serves only the latest version it finds (the highest-numbered folder). The version_policy field changes this:

PROTOBUF
version_policy: { latest: { num_versions: 2 } }
PROTOBUF
version_policy: { specific: { versions: [ 1, 3 ] } }
PROTOBUF
version_policy: { all: { } }

This lets you keep an old version serving next to a new one so that you can roll back without redeploying, and clients can address a version in the URL with /v2/models/<name>/versions/<n>/infer.

Warmup. The first request to a freshly loaded model is often slow, because engines allocate memory and compile kernels lazily. A model_warmup block makes Triton run sample requests at load time, so the model is genuinely ready before it reports ready. You will want this once you serve real models; for now it is enough to know it exists.

Inspecting the real configuration. Triton may modify your configuration while loading it, for example by filling in fields it can infer. To see what is actually in effect, ask the server:

BASH
curl localhost:8000/v2/models/add_vectors/config

When behaviour surprises you, compare this output to the file you wrote. A difference between them is often the whole explanation.

With max_batch_size above zero

  • dims describe one example
  • Client shape is [batch, ...dims]
  • Dynamic batching is available
  • Use for models whose first axis is a batch

With max_batch_size 0

  • dims describe the full tensor
  • Client shape equals dims exactly
  • Triton never stacks requests
  • Use for models that cannot batch
Try it
  1. Add a dynamic_batching { } block to your config and restart. Confirm the model still loads.
  2. Change max_batch_size to 0 and set dims to [ 1, 4 ] for all three tensors. Restart and send the same request. What changed about the shape you must send?
  3. Fetch /v2/models/add_vectors/config and find the fields you set.

Serving real model files: backends and file names

The Python backend is excellent for learning and for glue code, but most of what Triton serves in practice are exported model files executed by an optimised engine. The mechanics are the same as before; the version folder holds a model file instead of model.py, and the config names a different backend.

Triton looks for a default file name in the version folder, depending on the backend:

  • ONNX Runtime: model.onnx
  • TensorRT: model.plan
  • TorchScript (PyTorch): model.pt
  • OpenVINO: model.xml together with model.bin
  • Python: model.py
  • DALI (a data-preprocessing library): model.dali

If your file has a different name, set default_model_filename in the config. If you export from PyTorch, remember that Triton's PyTorch backend runs TorchScript files, produced by tracing or scripting a model, rather than a plain saved state dictionary.

Here is what a complete ONNX model looks like, using the image classifier from NVIDIA's own quickstart, called densenet_onnx:

TEXT
models/
  densenet_onnx/
    config.pbtxt
    1/
      model.onnx
models/densenet_onnx/config.pbtxt
name: "densenet_onnx"
backend: "onnxruntime"
max_batch_size: 8
input [
  { name: "data_0" data_type: TYPE_FP32 dims: [ 3, 224, 224 ] }
]
output [
  { name: "fc6_1" data_type: TYPE_FP32 dims: [ 1000 ] }
]

Read the dimensions as you now can: one example is a three-channel image 224 pixels square, and the output is a vector of 1,000 class scores. The tensor names data_0 and fc6_1 are not arbitrary. They are the names baked into the ONNX file, and the config must use them exactly. If you are not sure what names your own file contains, open it in a viewer such as Netron, or inspect it with the onnx Python package, before writing the config. For ONNX the configuration can in fact be auto-completed from the file, but writing it out documents the contract.

NVIDIA's quickstart provides example models, including this classifier, through a script named fetch_models.sh in the docs/examples folder of the Triton server repository on GitHub. Cloning that repository and running the script is the quickest way to get a real model to experiment with; check the repository for the current location of the script, since folders occasionally move. Mount the resulting model_repository folder at /models exactly as before, and the server will load it.

Then, from the SDK container, you can classify an image without writing any client code:

BASH
docker run -it --rm --net=host nvcr.io/nvidia/tritonserver:26.08-py3-sdk
/workspace/install/bin/image_client -m densenet_onnx -c 3 -s INCEPTION /workspace/images/mug.jpg

The flags select the model, ask for the top three classes, and choose an image-scaling convention. The output lists class scores with labels. --net=host lets the client container reach the server on localhost; this works as written on Linux, while on Docker Desktop you would instead point the client at the host address.

One backend deserves a warning at this stage: TensorRT. A TensorRT model.plan file is not a portable model format like ONNX. It is an engine compiled for a specific TensorRT version and a specific GPU architecture. A plan built on one machine will often refuse to load on another, or in a container with a different TensorRT version. If you use TensorRT, build plans inside the same container release you will serve with, on the same kind of GPU you will serve on. Triton upgrades its container monthly and bumps TensorRT along with CUDA, so plans must be rebuilt as part of upgrading. ONNX and TorchScript files do not have this limitation, which makes them the friendlier starting point.

One more feature is worth knowing exists: ensembles and business logic scripting. Real systems rarely call a lone model. They resize an image, run a classifier, and decode the result. Triton lets you declare such a pipeline in configuration (an ensemble) or code it in the Python backend (business logic scripting), so that the client makes one call and Triton wires the steps together on the server. You do not need either yet. When your client starts doing preprocessing that belongs next to the model, you will recognise the need.

Trap. The names in input and output must match the model file, not your preferences. A renamed tensor produces a load failure or an "unexpected inference input" error. Look at the model file first, then write the config.
Try it
  1. Export any small model from your framework of choice to ONNX, or download one.
  2. Open it in Netron and write down the input and output names, types, and shapes.
  3. Place it at models/mymodel/1/model.onnx, write a config using those details, and start the server.

Running the server: flags, logs, and model control

Until now we have used a single flag, --model-repository. A few more are worth knowing on day one, because they change how Triton behaves when things go wrong.

Where the models come from. --model-repository can be passed more than once to combine several repositories, as long as model names are unique across them. It also accepts cloud paths such as s3://bucket/path for Amazon S3, gs://bucket/path for Google Cloud Storage, and as://account/container/path for Azure. Cloud credentials come from the standard environment variables of each provider, for example the AWS access key variables or GOOGLE_APPLICATION_CREDENTIALS. Where you can, prefer a role or workload identity attached to the machine over long-lived keys pasted into environment variables.

What happens at startup failure. By default Triton exits if any model fails to load; this is the --exit-on-error flag, which defaults to true. That is a safe default for production, because a half-broken service should not accept traffic. While developing it can be annoying, since one broken model stops all the others. Set --exit-on-error=false and Triton will keep running, mark the failed model UNAVAILABLE in its table with a reason, and serve the rest. Related is --strict-readiness, true by default: the readiness endpoint reports not ready while any model has failed.

Model control mode. This setting decides how Triton reacts to changes in the repository after it starts. There are three values:

  • --model-control-mode=none is the default. Triton loads everything at startup and ignores later changes. To change a model, you restart the server.
  • --model-control-mode=poll makes Triton check the repository periodically, controlled by --repository-poll-secs, and reload what changed. The documentation advises against it for production use.
  • --model-control-mode=explicit loads nothing unless told to, either with --load-model=<name> at startup or by calling the repository API later.

For explicit mode, the API is a pair of POST requests:

BASH
curl -X POST localhost:8000/v2/repository/models/add_vectors/unload
curl -X POST localhost:8000/v2/repository/models/add_vectors/load

and POST /v2/repository/index lists every model Triton can see, with its state. A model passes through the states UNAVAILABLE, LOADING, READY, and UNLOADING. Explicit mode is a handy way to experiment, since you can change a file and reload a single model without restarting the whole server. The security-minded advice in the official deployment guide is to run production servers in the default mode, because letting clients load models is letting them run code. We return to that in the senior guide.

Logging. Triton logs to the terminal by default. --log-verbose=1 turns on verbose logging, which is noisy but invaluable when a model refuses to load and the normal message is not enough. --log-file writes to a file. When a problem is mysterious, add --log-verbose=1, reproduce it once, and read the lines around the failure.

Switching protocols off. --allow-http=false and --allow-grpc=false disable the corresponding endpoints. Reducing what is exposed is good hygiene once you know which protocol your clients use.

A complete invocation using several of these might look like this:

BASH
docker run --rm -p 8000:8000 -p 8001:8001 -p 8002:8002 \
  -v "$PWD/models:/models" \
  nvcr.io/nvidia/tritonserver:26.08-py3 \
  tritonserver --model-repository=/models \
               --model-control-mode=explicit \
               --load-model=add_vectors \
               --exit-on-error=false \
               --log-verbose=1

Another Docker detail will matter as soon as you use the Python backend with larger tensors: containers get a small shared-memory area by default, and the Python backend uses shared memory to pass data between processes. If you see crashes or "Bus error" messages under load, give the container more with --shm-size=1g (or more) on the docker run line.

Try it
  1. Add a second, deliberately broken model (an empty config.pbtxt in a new folder). Start with the defaults and watch the whole server exit.
  2. Start again with --exit-on-error=false. Confirm add_vectors still serves and that the broken one is UNAVAILABLE.
  3. Restart in explicit mode, load and unload add_vectors with the API, and request the repository index.

Watching it work: health, metadata, and metrics

A service you cannot observe is a service you cannot trust. Triton ships with three observation tools, all reachable without writing code.

Health endpoints. You met /v2/health/live and /v2/health/ready. They return empty bodies with a status code, which is exactly what orchestration systems want: Kubernetes (/student-guides/kubernetes) readiness and liveness probes can point straight at them. For a particular model, /v2/models/<name>/ready answers the same question.

Metadata. GET /v2 describes the server (name, version, extensions it supports), and GET /v2/models/<name> describes a model's tensors as Triton understands them. A client can fetch the metadata at startup and build its requests from it, which is how generic tooling talks to arbitrary models.

Metrics. Port 8002 serves metrics in the Prometheus text format:

BASH
curl localhost:8002/metrics

You will get a long list. Do not read it all; learn a handful. Send a few requests first, then look for these:

  • nv_inference_request_success and nv_inference_request_failure count requests that worked and requests that failed, per model.
  • nv_inference_count counts individual inferences (each example in a batch counts), and nv_inference_exec_count counts model executions. If requests are batched well, the first is noticeably larger than the second. Dividing them gives the average batch size, which is a quick test of whether batching is doing anything.
  • nv_inference_request_duration_us is the total time spent on requests, and nv_inference_queue_duration_us is the part spent waiting in line. If queue time dominates, the model is the bottleneck, not the network.
  • nv_gpu_utilization and nv_gpu_memory_used_bytes describe the GPU when there is one.

The durations are cumulative counters in microseconds, so a single reading means little. A monitoring system such as Prometheus (/student-guides/prometheus), with dashboards in Grafana (/student-guides/grafana), turns them into rates over time. For a beginner the point is simply to know that the data is there, free, from the first moment you run the server.

Triton also includes a benchmarking tool called Perf Analyzer, shipped in the SDK image. It sends synthetic load at a model and reports throughput and latency percentiles:

BASH
perf_analyzer -m add_vectors --concurrency-range 1:4

Run it from the SDK container with host networking, as above. It is the honest way to answer "how fast is it?", because it measures the server under load rather than a single request. Compare the numbers before and after you change a configuration setting such as dynamic batching, and you have a method for improving performance that rests on measurement instead of guessing. Keep in mind that the trivial adder will be limited by the Python backend and the network, not by a GPU, so use a real model when you care about the result.

Tip. After any change to a model's config, record a Perf Analyzer run before and after. A change that you cannot show helped is a change you should be suspicious of.
Try it
  1. Send twenty requests with a short Python loop.
  2. Fetch /metrics and find nv_inference_request_success for add_vectors. Does it read twenty?
  3. Fetch it again after one more request and confirm the counter only goes up.

Reading errors and fixing the common ones

Most of your first hours with Triton will be spent reading error messages. They are quite good once you know where to look. The first rule is to read the startup output from the top, not the bottom: the last line of a failed start is usually a generic summary, and the cause is several lines above it.

"creating server: Internal - failed to load all models." This is the generic summary printed when any model failed to load and --exit-on-error is true. Scroll up. Immediately before it Triton prints a line per failing model saying which model and why. Fix that one, or use --exit-on-error=false while you investigate.

The model shows UNAVAILABLE. The table's last column gives the reason. Common reasons are a missing model file (the file name does not match the backend's default, such as model.onnx.bak or a file placed in the model folder instead of the version folder), a version folder with a non-numeric name, a tensor name in the config that does not exist in the model, or a Python import error inside initialize. The debugging guide's advice is worth following as a habit: before blaming Triton, confirm that the model runs in its own framework outside Triton. If it fails to run in ONNX Runtime directly, no config will fix it.

A batch-size message at load time, such as "unexpected configuration maximum batch size ... model maximum is 1". This happens most with TensorRT engines and some ONNX files that were exported with a fixed batch dimension. The model itself cannot batch, but your config claimed it could. Either set max_batch_size: 0 and put the full shape in dims, or re-export the model with a dynamic batch axis.

A request error saying the shape does not match. The message names the input, the expected shape, and the shape you sent. With batching enabled, remember that the first dimension is the batch and is not part of dims. Nearly every shape error from beginners is the max_batch_size confusion from the first-project section.

A request error naming an unexpected input. The tensor name in your client does not match the config; the message lists the allowed names. Names are case-sensitive.

A batch-size limit error. You sent more examples in one request than max_batch_size. Either send fewer or raise the limit if the model supports it.

"Request for unknown model." Either the name is misspelled in the URL or client call, or the model failed to load. Check /v2/repository/index or the startup table.

CPU-only complaints when loading a GPU model. If you start the container without --gpus but the model's config places instances on KIND_GPU, or the model uses a GPU-only backend like TensorRT, it cannot load. Either give the container a GPU or set instance_group to KIND_CPU and use a CPU-capable backend.

An engine incompatible with this TensorRT version. A TensorRT plan built elsewhere or by another release will not deserialise. Rebuild it inside the container you serve with.

"NVIDIA driver too old" style failures. The host driver must meet the CUDA requirement of the container. Updating the driver, or pulling an older container tag that matches it, fixes this. The release notes list the requirement.

Health returning a non-200 while the server runs. The server is alive but not ready: a model is still loading, or a model has failed and strict readiness is on. Large models can take minutes to load, which is why production systems use a generous startup probe.

Bus error or shared-memory failures with the Python backend. Increase --shm-size on docker run.

Permission denied reading the model folder. The container user cannot read files that were created with restrictive permissions on your machine. Make the repository readable.

All the exact message wordings can differ slightly between releases, so match on the meaning, not the characters. When the message is unclear, run with --log-verbose=1 and reproduce once.

Trap. Do not edit files inside a model folder while Triton is loading or unloading that model, and do not expect edits to take effect in the default control mode. Triton read the folder at startup. Restart the server, or use explicit mode and call the reload endpoints.
Try it
  1. Rename 1 to v1 and start the server. Read the result and explain why the model is not found.
  2. Rename it back, then misspell the output name in config.pbtxt. Read the error carefully.
  3. Send a request with the wrong shape and note exactly what the message says to fix.

Putting it all together

Let us build one small project end to end, using everything above. The goal: a text-length model served by Triton, a client that calls it, a metrics check, and a script that starts it all, so a teammate can reproduce your work with two commands. We will reuse the Python backend so that nothing needs downloading. The model takes a list of strings and returns each string's length.

Create the layout:

TEXT
triton-demo/
  models/
    text_length/
      config.pbtxt
      1/
        model.py
  client.py
  run.sh

The configuration declares a string input of variable length and an integer output. Because strings are variable-sized byte sequences, the type is TYPE_STRING:

triton-demo/models/text_length/config.pbtxt
name: "text_length"
backend: "python"
max_batch_size: 16

input [
  { name: "TEXT" data_type: TYPE_STRING dims: [ 1 ] }
]
output [
  { name: "LENGTH" data_type: TYPE_INT32 dims: [ 1 ] }
]

dynamic_batching { }

instance_group [ { count: 1 kind: KIND_CPU } ]

Each example is one string, so dims is [ 1 ], and the client sends shape [batch, 1]. The dynamic batcher is on, so concurrent requests can be combined.

triton-demo/models/text_length/1/model.py
import numpy as np
import triton_python_backend_utils as pb_utils


class TritonPythonModel:
    def execute(self, requests):
        responses = []
        for request in requests:
            text = pb_utils.get_input_tensor_by_name(request, "TEXT").as_numpy()
            lengths = np.vectorize(len)(text).astype(np.int32)
            out = pb_utils.Tensor("LENGTH", lengths)
            responses.append(pb_utils.InferenceResponse(output_tensors=[out]))
        return responses

A string tensor arrives in the Python backend as a NumPy array of bytes objects, and len of a bytes value is the byte count. For non-ASCII text such as Arabic, byte length differs from character count, since Arabic letters take two bytes in UTF-8. If you wanted characters you would decode first; the example is a good reminder that data types have consequences.

The client uses the object type for strings:

triton-demo/client.py
import numpy as np
import tritonclient.http as httpclient

client = httpclient.InferenceServerClient(url="localhost:8000")

texts = ["hello", "triton", "مرحبا"]
data = np.array(texts, dtype=object).reshape(-1, 1)

inp = httpclient.InferInput("TEXT", data.shape, "BYTES")
inp.set_data_from_numpy(data)

result = client.infer("text_length", [inp])
for text, length in zip(texts, result.as_numpy("LENGTH").flatten()):
    print(f"{text!r}: {length} bytes")

The wire type for strings is BYTES, not STRING; remember it as another case where the config and the protocol differ in spelling. The script prints each word with its byte length, and you will see the Arabic word reports ten bytes for five characters.

Wrap it so a teammate does not need to remember the long docker run:

triton-demo/run.sh
#!/usr/bin/env bash
set -euo pipefail

docker run --rm \
  -p 8000:8000 -p 8001:8001 -p 8002:8002 \
  -v "$PWD/models:/models" \
  nvcr.io/nvidia/tritonserver:26.08-py3 \
  tritonserver --model-repository=/models

Now exercise the whole thing. In one terminal run bash run.sh. Wait for the READY row in the startup table. In another, run curl localhost:8000/v2/health/ready, then python client.py, then curl localhost:8002/metrics | grep text_length to find the counters for your model. Every step you took in this guide appears here: layout, configuration, a Python-backend model, a client, health, metrics, and a pinned container tag. Commit the folder to Git, and anyone on your team can reproduce it, which is precisely the benefit of describing a service as configuration. If you later containerise your own serving image or deploy to Kubernetes, the model repository folder you just wrote moves unchanged.

Try it
  1. Build the triton-demo project and run it with the script.
  2. Change the model so it returns character counts instead of byte counts, using .decode("utf-8"). Restart and compare the Arabic word.
  3. Add a second version folder 2 with the change, set version_policy: { all: { } }, and call each version through its own URL.

What you can now do, and what comes next

You started this guide with a model file and no way for anyone else to use it. You can now explain what Triton is for, in terms of idle GPUs, format sprawl, and rebuilt operations. You can name the four nouns, the repository, the backend, the configuration, and the instance, and use them to locate a problem. You can pull and run the container, lay out a model repository, write a config.pbtxt with the right dimension convention, serve both a Python model and an exported file, call it from curl and from the Python client, watch its health and metrics, and read the startup errors that stop most first attempts.

What this guide did not do is make you an operator. You have not tuned batching against real traffic, built an ensemble pipeline, served a large language model, deployed to Kubernetes, protected the server on a network, or planned upgrades across the monthly release train. Those are the next two levels.

The mid-level guide covers how the scheduler and dynamic batcher really work and how to choose instance counts, how to accelerate ONNX models with TensorRT, how to build ensembles and business logic scripts, and how to benchmark with Perf Analyzer and Model Analyzer. The senior guide covers security (Triton has no built-in authentication and must sit behind a gateway), multi-tenancy, scaling, upgrades, cost, and incident response.

Beyond this series, natural neighbours on this site are the Docker guide (/student-guides/docker) for the container skills Triton leans on, the KServe guide (/student-guides/kserve) for running Triton on Kubernetes through a standard serving layer, the BentoML guide (/student-guides/bentoml) and the Ray Serve guide (/student-guides/ray-serve) for alternative serving approaches that you can compare against, the TorchServe guide (/student-guides/torchserve) for a PyTorch-specific alternative, and the MLflow guide (/student-guides/mlflow) for tracking the models you eventually put in the repository.

One last piece of advice. Keep a small, working example like the one in this guide somewhere you can run in five minutes. When a colleague asks "how would we serve this?", answering with a running server and a twenty-line client is far more persuasive than any slide.

Sources