Skip to content
Back to student guides
Text Generation InferenceLLMsInference & gateways3 levels100 sectionsCovers TGI 3.3.7 (final release, maintenance mode)

The Complete Text Generation Inference Guide

Serve open LLMs in production with Hugging Face Text Generation Inference. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
15sections
19examples

This is part one of three. It covers everything you need to start a Text Generation Inference server, talk to it, and understand what it prints when something goes wrong. By the end you can launch a model in a container, check that it is healthy, send it prompts through three different APIs, stream tokens back as they are produced, read the startup logs and the common errors, and hand a colleague one script that brings the whole thing up. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. A language-model server rewards hands-on work more than most tools, because the interesting behaviour (the wait while weights load, the memory errors, the way tokens stream) only becomes real once you have watched it happen on your own machine or a rented GPU.

One piece of context comes before any command, because it changes how you should read the rest of this guide.

What TGI is, and what you must know before you start

Text Generation Inference, usually shortened to TGI, is a server for large language models. You point it at a model from the Hugging Face Hub, it loads the weights onto one or more GPUs, and it exposes an HTTP API. Your application sends a prompt to that API and gets generated text back, either all at once or token by token as the model produces it. It is built by Hugging Face, written mostly in Rust (the web layer) and Python (the model layer), and licensed under Apache 2.0.

Now the important caveat. TGI is in maintenance mode, and its repository is archived. Hugging Face moved the project into maintenance mode on 11 December 2025, accepting only minor bug fixes and documentation fixes from then on, and archived the GitHub repository (made it read-only) on 21 March 2026. The final release is v3.3.7, published on 19 December 2025. No further releases are expected, which means no new model architectures and no new security patches.

That does not make TGI a waste of your time, and the reasons matter:

  • It is still widely deployed. Plenty of production systems, tutorials, cloud templates and job descriptions mention it, so you will meet it at work whether or not you would start a new project on it.
  • It is frozen, which is a kind of stability. The behaviour in v3.3.7 will not shift under you. A guide written today stays true.
  • Its ideas transfer directly. Continuous batching, a paged KV cache, prefix caching, chunked prefill, and an OpenAI-compatible chat endpoint are the core concepts of every modern inference engine. Hugging Face itself now recommends vLLM and SGLang, as well as local engines such as llama.cpp and MLX, for new work. Learning TGI properly means learning the vocabulary those tools share. If you want to compare, the catalogue has guides for vLLM and SGLang.
Plan for the exit from day one. If you are choosing an engine for a brand-new project, start with vLLM or SGLang. Use this guide when you inherit a TGI deployment, when a platform you rely on still offers it, or when you want to understand how an inference server works. Every part of this series ends with a note on what carries over.

The rest of this guide treats TGI as what it is: a stable, finished product at version 3.3.7. Where a tutorial you find online disagrees with this guide, check its date; many were written for versions 1.x or 2.x, and some defaults and APIs have changed since.

Try it
  1. Open the repository page for huggingface/text-generation-inference in a browser.
  2. Find the banner or notice about maintenance mode and the archived status.
  3. Open the Releases page and find v3.3.7.
a read-only repository whose newest release is 3.3.7. That version number is the one every command in this guide uses.

The problem TGI solves, and what came before it

To see why a dedicated server exists, start with the simplest way to run a language model: load it in a Python script with the transformers library and call generate(). That works well for one person in a notebook. It falls apart as soon as several users share the model, for three reasons.

First, a language model produces text one token at a time. A token is a small chunk of text, often a word or part of a word. To write a 200-token answer, the model runs a full forward pass 200 times, each time using everything generated so far. A naive script processes one request at a time, so the second user waits for the first to finish completely.

Second, a naive batch wastes the GPU. The obvious fix is to group several requests into a batch and run them together, since a GPU is far more efficient doing many things at once. But requests are different lengths. In a static batch, everyone waits for the longest answer, and the GPU sits idle on the short ones that finished early. The user whose reply needed five tokens waits as long as the one who asked for an essay.

Third, the memory is the real limit. While generating, the model keeps the attention keys and values of every previous token so it does not recompute them. This is called the KV cache, and it grows with every token of every active request. Managing it badly means the GPU runs out of memory long before the compute is exhausted.

TGI solves these problems in the following ways, each of which you will meet properly later:

  • Continuous batching. New requests join the running batch between steps, and finished requests leave it immediately. Nobody waits for a slow neighbour.
  • Paged KV cache. Memory is split into fixed-size blocks that need not be contiguous, so there is little waste and blocks can be shared between requests with a common prefix.
  • Optimised kernels and tensor parallelism. Fast attention implementations, and the ability to split one large model across several GPUs.
  • Streaming and a real HTTP API. Tokens flow back to the client as they are produced, over Server-Sent Events, with the same API shape you already know from hosted providers.
  • Production plumbing. Health checks, Prometheus metrics, OpenTelemetry tracing, and request validation.

Before servers like this existed, teams wrote their own Flask or FastAPI wrapper around transformers, handled concurrency with locks or a queue, and discovered the problems above in production. TGI was Hugging Face's answer, first built to power its own hosted inference and the HuggingChat product, and later released for everyone.

A note on terms you will see constantly. Inference means using a trained model to produce output, as opposed to training it. Serving means running a model behind an API so other programs can call it. Latency is how long one request takes; throughput is how many tokens per second the server produces across all users. The two pull against each other, and most settings in an inference server are a trade between them.

Try it
  1. Think of a chatbot with ten simultaneous users, where nine want a one-line answer and one wants a long report.
  2. Write down what a one-request-at-a-time server does to the nine short requests.
  3. Write down what continuous batching changes about that.
with one-at-a-time, the nine wait behind the report. With continuous batching, they join the running batch, finish quickly, and leave while the long reply keeps going.

The mental model: three processes and a handful of nouns

Almost every failure in TGI makes sense once you know that it is three cooperating programs, not one. The Docker image starts them for you, but when something breaks, the logs are labelled by which one spoke.

CLIENTcurl, Python, your app
→
ROUTERHTTP, queue, batching
→
MODEL SERVERone shard per GPU
→
GPUweights + KV cache

The launcher (text-generation-launcher, written in Rust) is the program the container runs when it starts. It reads your flags, reads the model's config.json, downloads and converts the weights, starts one model-server process per GPU, then starts the router, and finally supervises all of them. If any child dies, the launcher shuts everything down. That is why a crash in one place shows up as "Shard crashed" or "Webserver crashed" followed by the container exiting.

The router, also called the web server (text-generation-router, Rust), is the part you talk to. It accepts HTTP requests, tokenizes the text, validates it against the limits, puts it in a queue, and builds batches. It talks to the model server over gRPC on a Unix socket, and streams tokens back to you. It has no configuration file; everything is a flag or an environment variable.

The model server (text-generation-server, Python and PyTorch) loads the weights, owns the KV cache, and runs the forward passes on the GPU. With several GPUs, there is one process per GPU, called a shard, and the shards coordinate through NCCL, NVIDIA's library for GPU-to-GPU communication.

The vocabulary you need on day one:

  • Prefill. The first pass over your whole prompt. It fills the KV cache and produces the first token. It is compute-heavy and the most memory-hungry step.
  • Decode. Every later step: one new token per active request. It is limited by memory bandwidth rather than compute.
  • KV cache. The stored attention state of past tokens. Its size decides how many requests the server can hold at once.
  • Continuous batching. New requests join the running batch between decode steps.
  • Shard. One slice of the model on one GPU, when the model is split across GPUs (tensor parallelism).
  • Warmup. At startup, the server runs a worst-case prefill to measure memory, then reserves the KV cache and prepares its kernels. Most out-of-memory failures happen here, which is good: you learn about them before users do.
  • Token. The unit of text the model reads and writes. Limits in TGI are all in tokens, not characters or words. A rough rule for English is about three-quarters of a word per token, but code and non-English text, Arabic included, can use noticeably more tokens for the same meaning.

One more concept: TGI serves one model per server. There is no switching between models inside a running instance. If you need two models, you run two containers. The one exception is multi-LoRA serving (small fine-tuned adapters on one base model), which later guides cover.

Try it
  1. Draw the diagram above on paper and add a box labelled "launcher" around the router and the model server.
  2. Mark which of the three processes would log a message about a prompt being too long.
  3. Mark which would log a CUDA out-of-memory error.
the router rejects prompts that are too long (it validates), while the model server, which holds the GPU, reports memory errors.

What you need before you install

TGI is a Linux and NVIDIA GPU server. That sentence saves a lot of confusion, so here is the full picture.

The supported path is a Linux machine with an NVIDIA GPU, a driver that supports CUDA 12.2 or newer, Docker, and the NVIDIA Container Toolkit (the piece that lets containers see the GPU). The images for 3.3.1 and later are built on Torch 2.7 and CUDA 12.8, so a recent driver is the safest choice. The optimised kernels are officially tested on H100, A100, A10G and T4 cards. On other NVIDIA GPUs TGI still runs and still batches continuously, but some operations such as flash attention and paged attention may not be used, so it can be slower. Flash attention itself needs an Ampere-generation GPU or newer; older cards fall back to a different path.

Other accelerators exist but are a separate topic: AMD ROCm images, Intel GPU and CPU images, Intel Gaudi, AWS Inferentia2 and Trainium, and TensorRT-LLM and llama.cpp backends. The docs say plainly that not all variants offer the same features. This guide uses the NVIDIA path throughout.

On a Mac, there is no accelerated path. None of the images support Apple Silicon GPUs, and Docker Desktop on a Mac cannot pass a GPU through to a container. The only documented macOS steps are for building from source on the CPU, which the docs themselves say is not the recommended usage. If you are on a Mac, do your hands-on work on a rented Linux GPU machine (any cloud provider offers an hourly one, including regions in the Gulf and Europe), or use Hugging Face Inference Endpoints. For running models locally on a Mac, the sensible tools are MLX or llama.cpp directly, and the catalogue's llama.cpp guide and Ollama guide cover that ground.

On Windows, the TGI docs say nothing. The practical route, which comes from Docker's and NVIDIA's own documentation rather than TGI's, is WSL2 with Docker Desktop's WSL2 backend and an NVIDIA driver that supports CUDA in WSL. Then the Linux commands work unchanged. Treat it as community-supported at best. Native Windows builds are not supported.

How big a GPU do you need? A rule of thumb: the weights alone take roughly two bytes per parameter in half precision. A 7-billion-parameter model needs about 14 GB for weights, and then more for the KV cache, so a 24 GB card is comfortable and a 16 GB card is tight. Smaller models, such as those with one to three billion parameters, run on modest cards and are the best choice for learning, because they download and start in minutes rather than tens of minutes.

Learn on a small model. Pick a model of a few billion parameters or less for every exercise here. The behaviour (logs, API, errors) is identical to a big model, and your mistakes cost a two-minute restart instead of a twenty-minute one.

Check your setup before pulling a multi-gigabyte image. Run these two commands on the Linux machine:

BASH
nvidia-smi
docker run --rm --gpus all ubuntu nvidia-smi

The first shows your GPU, driver version and the highest CUDA version it supports. The second proves that Docker can pass the GPU into a container. If the second fails with an error about a missing runtime or driver, the NVIDIA Container Toolkit is not installed or Docker was not restarted after installing it. Fix that first; nothing else in this guide will work without it.

Try it
  1. Run nvidia-smi and note the GPU model and its total memory.
  2. Run the second command above and confirm the same table appears from inside a container.
  3. Work out roughly how many billion parameters fit in half precision, at two bytes each, leaving a third of the memory free.
the same GPU table both times, and a number such as about 8 billion parameters for a 24 GB card.

Running your first server

The official way to run TGI is a single docker run command. Here it is, with a small model and a version pinned to the final release:

BASH
model=HuggingFaceH4/zephyr-7b-beta
volume=$PWD/data

docker run --gpus all --shm-size 1g -p 8080:80 -v $volume:/data \
    ghcr.io/huggingface/text-generation-inference:3.3.7 \
    --model-id $model

That command is dense, so take it apart. Everything before the image name is a Docker option; everything after it is passed to TGI's launcher.

  • --gpus all gives the container access to every GPU. TGI uses all visible GPUs by default, one shard each.
  • --shm-size 1g gives the container 1 GB of shared memory. Docker's default is only 64 MB, and the GPU shards use shared memory to talk to each other through NCCL. With one GPU you may get away without it; with several, the lack of it causes hangs or errors. Always include it.
  • -p 8080:80 publishes the container's port 80 on your machine's port 8080. The image sets its listening port to 80 inside the container, so the right-hand number is 80; the left-hand number is your choice. (Outside Docker, the launcher's own default port is 3000, which is a common source of confusion.)
  • -v $volume:/data mounts a folder from your machine at /data in the container. The image stores downloaded weights there (it sets the Hugging Face home directory to /data). Without this mount, the weights live inside the container and are downloaded again every time you start a fresh one. For a multi-gigabyte model that is the difference between a thirty-second start and a twenty-minute one.
  • ghcr.io/huggingface/text-generation-inference:3.3.7 is the image from GitHub's container registry. Pin the version. The latest tag exists, but because the project is frozen, a pinned tag is the only way you can state, a year from now, exactly what you ran. Older tutorials pull from other registries, such as an Azure one; those images are no longer pushed, so use ghcr.io.
  • --model-id $model is the first TGI flag: which model to serve. It is a Hugging Face Hub identifier of the form organisation/name. If you leave it out, the launcher defaults to a tiny placeholder model, bigscience/bloom-560m, which is fine for a smoke test but useless for real work.

What happens next takes a while, and it is worth knowing the stages so you do not think it has hung. The launcher downloads the weights from the Hub into /data (minutes for a large model, and only the first time), converts them if needed, starts the shard, loads the weights onto the GPU, runs warmup, and finally starts the router. Only then does the server answer requests. The first start of a 7B model on a good connection commonly takes several minutes.

Run it in the foreground at first, so you can watch the logs. Once you are comfortable, add -d to detach and --name tgi to name the container, then follow logs with docker logs -f tgi.

Do not delete the volume mount to "simplify" the command. A server without -v is the single most common reason a beginner's restarts take twenty minutes. The weights are downloaded into the container's writable layer, and they vanish when the container is removed.
Try it
  1. Pick a small model from the Hub whose licence lets you download it without accepting terms.
  2. Run the command above with that model and a data folder in your current directory.
  3. While it runs, open a second terminal and run du -sh data every minute.
the folder grows as weights download, then stops growing while the model loads onto the GPU.

Reading the startup logs and confirming the server is healthy

Do not treat the startup output as noise. It tells you which attention method was chosen, how the memory was divided, and where it failed if it did. Here is the sequence you should see, in order, with what each line means.

You will see the launcher print its parsed arguments, then lines about downloading. A line saying Successfully downloaded weights for followed by the model name means the files are on disk. Then:

TEXT
Waiting for shard to be ready...
Shard ready in 31.2s

This means the model server has loaded the weights onto the GPU. Next, a line similar to Using attention flashinfer - Prefix caching true tells you the launcher's choices. Flashinfer is the default attention implementation. Prefix caching being on means that if two requests begin with the same text, such as a shared system prompt, the work for that shared part is reused. You do not need to understand more yet; just notice that these decisions are visible in the log.

Then comes Default max_batch_prefill_tokens to ..., and warmup lines from the router. When the router prints that it is connected and listening, the server is ready. The exact wording varies slightly between versions, so look for the sequence rather than an exact string: downloaded, shard ready, attention chosen, warmup, connected.

Now confirm from outside. Run this in a second terminal:

BASH
curl -s localhost:8080/health -o /dev/null -w '%{http_code}\n'

The -s silences the progress meter, -o /dev/null discards the body, and -w prints just the HTTP status code. You should see 200. Before the server is ready you will see a connection error, or 503 with a body of {"error":"unhealthy","error_type":"healthcheck"}. Keep re-running it, or put it in a loop, until you get 200.

Next, ask the server what it decided:

BASH
curl -s localhost:8080/info

This returns a JSON document describing the running server. The fields to look at first:

  • model_id and model_sha: which model and which exact revision you are serving.
  • max_input_tokens and max_total_tokens: the longest prompt accepted and the longest prompt-plus-answer allowed. In TGI version 3, you normally do not set these; the server computes them from your hardware and the model.
  • version: the TGI version, which should say 3.3.7.

Those token limits matter more than any other numbers on this page, and we return to them in the configuration section. Looking at /info right after startup is the best habit a beginner can build, because it shows you what the server actually chose rather than what you assumed.

There is also an interactive API reference. Open http://localhost:8080/docs in a browser and you get a Swagger page listing every route, with a "Try it out" button for each. It is a good way to see which fields exist before writing code.

Poll, do not guess. Write a small loop such as until curl -sf localhost:8080/health; do sleep 5; done in your scripts. A fixed sleep 60 is wrong for a cold start and wasteful for a warm one.
Try it
  1. With your server running, call /health and confirm a 200.
  2. Call /info and write down max_input_tokens and max_total_tokens.
  3. Open /docs in a browser and find the /generate route.
a 200, two token limits that satisfy max_input_tokens being smaller than max_total_tokens, and a Swagger page.

Your first generation with /generate

TGI has its own native endpoint, POST /generate. You send a JSON body with the prompt in inputs and the options in parameters, and you get JSON back. It is the simplest possible interaction, so start here.

BASH
curl localhost:8080/generate -X POST \
    -H 'Content-Type: application/json' \
    -d '{"inputs":"What is Deep Learning?","parameters":{"max_new_tokens":20}}'

The reply looks like this:

JSON
{"generated_text":" Deep learning is a subset of machine learning that uses neural networks with many layers to"}

Notice three things. First, the answer starts mid-sentence and stops abruptly: max_new_tokens is 20, so the model was cut off after twenty tokens. That is a limit you set, not an error. Second, the text begins with a space; the model continues from your prompt, token by token, so what you get is the continuation, not a tidy reply. Third, there was no chat structure: /generate hands your string to the model as-is. For a base model that is exactly right; for an instruction-tuned model, which expects a conversation in a specific format, you usually want the chat endpoint described in the next section.

The parameters object controls how the text is produced. The ones to know now:

  • max_new_tokens: the most tokens to generate. In TGI 3, if you omit it, the server picks a value automatically, up to what fits in max_total_tokens minus your prompt length. Setting it explicitly is still good practice, because it bounds your cost and latency.
  • temperature: how random the output is. Higher values give more varied text. It must be strictly greater than zero.
  • top_p and top_k: restrict sampling to the most likely tokens. top_p takes a value between 0 and 1, exclusive; top_k takes a positive integer.
  • do_sample: whether to sample at all. With false, the model always takes its most likely token, which is deterministic ("greedy") and what you want for a repeatable test.
  • stop: a list of strings that end generation when they appear. The server allows up to four by default.
  • seed: a number that makes sampled output repeatable.
  • repetition_penalty: values above 1 discourage repeating text.
  • details: when true, the response also includes details such as the finish reason and the token count.

Try the details flag, because it teaches you how to see why generation ended:

BASH
curl localhost:8080/generate -X POST \
    -H 'Content-Type: application/json' \
    -d '{"inputs":"Name three colours:","parameters":{"max_new_tokens":30,"details":true,"do_sample":false}}'

The response now includes a details object with a finish_reason. You will see length when the token limit stopped it, eos_token when the model decided it was done, and stop_sequence when one of your stop strings matched. Reading the finish reason is the quickest way to tell "the model finished" from "I cut it off".

Temperature zero is an error, not "deterministic". Sending "temperature": 0 returns a 422 validation error saying the temperature must be strictly positive. For repeatable output, send "do_sample": false instead.
Try it
  1. Send the same prompt three times with do_sample set to true and a temperature of 1.0. Compare the outputs.
  2. Send it three times with do_sample set to false. Compare again.
  3. Add details and read the finish reason each time.
three different texts when sampling, three identical texts when greedy, and a finish reason of length when your token limit was the reason it stopped.

The chat endpoint and the OpenAI-compatible API

Most models you will serve are instruction-tuned: trained to follow a conversation made of messages with roles, such as a system message, user messages and assistant replies. Each model family expects those messages wrapped in its own special format, called a chat template, which lives in the model's tokenizer_config.json. Writing that by hand is error-prone.

TGI does it for you. Since version 1.4 it has served the Messages API, an endpoint compatible with the OpenAI Chat Completions format, at POST /v1/chat/completions. In version 3 it is always on; you no longer need a flag. The server applies the model's own chat template to your messages and then generates.

BASH
curl localhost:8080/v1/chat/completions \
    -X POST \
    -H 'Content-Type: application/json' \
    -d '{
      "model": "tgi",
      "messages": [
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain what a GPU does in one sentence."}
      ],
      "max_tokens": 60
    }'

The model field is informational: TGI serves one model, so the value is not used to choose anything, and the docs use the placeholder "tgi". The reply has the familiar OpenAI shape, with the text at choices[0].message.content and token counts under usage.

Because the shape matches OpenAI's, you can use OpenAI's own Python library by pointing it at your server. Install it with pip install openai, then:

PYTHON
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8080/v1/", api_key="-")

reply = client.chat.completions.create(
    model="tgi",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain what a GPU does in one sentence."},
    ],
    max_tokens=60,
)
print(reply.choices[0].message.content)

The api_key="-" is a dummy: the library insists on a value, and TGI ignores it unless you started the server with --api-key. Everything else, including the base_url ending in /v1/, is how you redirect any OpenAI-compatible tool to your own server. This is a major practical benefit: code written against a hosted provider can be switched to your TGI server by changing one URL, and switched again later when you move to another engine.

Hugging Face's own client, huggingface_hub, also works and is the one TGI's documentation recommends:

PYTHON
from huggingface_hub import InferenceClient

client = InferenceClient(base_url="http://localhost:8080/v1/")
out = client.chat.completions.create(
    model="tgi",
    messages=[{"role": "user", "content": "Say hello in Arabic."}],
    max_tokens=40,
)
print(out.choices[0].message.content)

You will see an older package, text-generation (installed with pip install text-generation), in many tutorials. Its own README now carries a legacy warning and recommends huggingface_hub instead. Avoid it in new code.

Two chat-specific traps are worth knowing now. Some models' templates do not allow a system role, and sending one produces an error starting Template error:; the fix is to fold your instructions into the first user message. And if a model has no chat template at all (typically a base model rather than an instruct model), the chat endpoint cannot work, and you should use /generate instead.

There is also POST /v1/completions, the older OpenAI completions shape, and GET /v1/models, which lists the single served model. You will rarely need either.

Try it
  1. Send the curl chat request above, then read the usage block in the response.
  2. Run the same request through the OpenAI Python client.
  3. Change the system message to ask for answers in French, and compare the two replies.
matching replies from both routes, a usage block with prompt, completion and total token counts, and tone that clearly follows the system message.

Streaming tokens as they are produced

A model that takes eight seconds to write a long answer feels broken if the user sees nothing for eight seconds, and fine if words appear immediately. Streaming is how every chat product achieves that. The server sends each token the moment it is generated, over Server-Sent Events (SSE), a simple web standard where one HTTP response stays open and delivers a series of small data: messages.

With the chat endpoint, add "stream": true:

BASH
curl -N localhost:8080/v1/chat/completions \
    -X POST \
    -H 'Content-Type: application/json' \
    -d '{
      "model": "tgi",
      "messages": [{"role": "user", "content": "Write a haiku about the desert."}],
      "max_tokens": 60,
      "stream": true
    }'

The -N flag turns off curl's output buffering; without it you would see nothing until the end, which defeats the purpose. The output is a run of lines each beginning with data: followed by a JSON chunk. In every chunk, the new piece of text sits at choices[0].delta.content, and the final chunk carries a finish reason. The stream ends with a closing message, and the connection closes.

In Python, the OpenAI client makes this a loop:

PYTHON
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8080/v1/", api_key="-")

stream = client.chat.completions.create(
    model="tgi",
    messages=[{"role": "user", "content": "Write a haiku about the desert."}],
    max_tokens=60,
    stream=True,
)
for chunk in stream:
    piece = chunk.choices[0].delta.content
    if piece:
        print(piece, end="", flush=True)
print()

The if piece: guard matters: some chunks, such as the first (which carries the role) and the last, have no text, and printing None would be a bug. The flush=True makes the characters appear immediately instead of waiting for the terminal's buffer.

On the native API, streaming has its own route, POST /generate_stream, with the same request body as /generate. Each event carries one token and, on the last event, the full generated text. Alternatively POST / accepts the same body and chooses between streaming and not based on a stream field.

Why streaming matters beyond looks: it changes what you measure. Time to first token is how long until the first piece appears, and it is dominated by queueing and the prefill step. Inter-token latency is the gap between later tokens, set by the decode step. Users judge a chat by the first and feel the second as reading speed. When you start tuning a server, these are the two numbers to watch, and streaming is the only way to see them from the client.

One limitation: the best_of parameter, which generates several candidates and returns the best, is not allowed with streaming. The server rejects it with a 422 error saying best_of is not supported when streaming tokens.

Try it
  1. Run the streaming curl command with -N, then run it again without -N.
  2. Run the Python streaming loop and time how long until the first character appears.
  3. Ask for a long answer and notice whether the words arrive at a steady pace.
words trickling in with -N and arriving in one lump without it, and a first-character delay of well under a second on a warm server.

Gated models, tokens, local models and the weights cache

Not every model downloads freely. Many popular ones, such as the Llama and Gemma families, are gated: you must log in to the Hugging Face Hub, read and accept the model's licence on its page, and then authenticate when you download. If you launch such a model without credentials, the download fails and the launcher exits with Error: DownloadError, with a message about a 401, 403 or gated repository a few lines above.

The fix has two parts, and both are required. On the Hub, visit the model page while logged in and accept its terms. Then create an access token in your account settings, choosing a read token (the least powerful kind, which is all you need), and pass it to the container:

BASH
export HF_TOKEN=hf_your_read_token_here

docker run --gpus all --shm-size 1g -p 8080:80 -v $PWD/data:/data \
    -e HF_TOKEN=$HF_TOKEN \
    ghcr.io/huggingface/text-generation-inference:3.3.7 \
    --model-id meta-llama/Llama-3.2-1B-Instruct

The environment variable is HF_TOKEN. You will see an older name, HUGGING_FACE_HUB_TOKEN, in older material; the router still falls back to it, but HF_TOKEN is current. Never write the token into a script you commit, and never bake it into an image. Read it from your shell, a secrets manager, or a file that is in your .gitignore.

Three related ideas make downloads manageable:

  • Pin a revision. The --revision flag takes a commit hash, branch or tag from the model's repository. Without it you get the default branch, which the model's authors can change at any time. For anything you want to reproduce, pin the commit hash.
  • Use a local model. If you have a model directory on disk (for example one saved with save_pretrained), mount it under /data and pass its in-container path: --model-id /data/my-model. The folder must contain the weights, the tokenizer files and the config.
  • Understand the cache layout. Since version 2.3, downloaded models live under /data/hub/models--organisation--name inside the container, which is ./data/hub/... on your side. If you reuse a cache volume from an old tutorial that stored models directly under /data, TGI will not find them and will download again.

A note on file formats, since it causes real confusion. Model weights come as safetensors files (a safe, pickle-free format) or as older .bin files, which are Python pickles. Pickles can execute arbitrary code when loaded, so since TGI 2.0 it does not convert them unless you pass --trust-remote-code. Always prefer a repository that has safetensors. Tensor parallelism across several GPUs also requires safetensors.

For teams in the Gulf or Egypt with data-residency rules, one more property is useful: once the weights are in your volume, you can run the container with HF_HUB_OFFLINE=1 and no token at all, and it never contacts the Hub. Prompts never leave your infrastructure because the model runs on your own GPU, which is often the reason an organisation chooses to self-host.

Gated model error is not a GPU problem. When the log ends in Error: DownloadError, look a few lines up for a 401, 403 or "gated" message. The cause is almost always a missing token or licence you have not accepted on the Hub, and no amount of restarting fixes it.
Try it
  1. Choose a gated model, accept its terms on the Hub, and create a read token.
  2. Launch it once without HF_TOKEN and read the error.
  3. Launch it again with the token and the same volume.
a DownloadError the first time with a gated-repository message above it, then a normal startup the second time.

Configuration: flags, environment variables and the token budget

Every launcher flag has an environment-variable twin: the flag --max-total-tokens is the variable MAX_TOTAL_TOKENS. In Kubernetes and managed platforms you will set variables; in a hand-run command you will use flags. They do the same thing. To see every flag for your version, run:

BASH
docker run --rm ghcr.io/huggingface/text-generation-inference:3.3.7 --help

The list is long, and a beginner needs only a few of them. Start with these:

Flag What it does
--model-id Which model to serve: a Hub id or a local path under /data
--revision A commit hash, branch or tag to pin
--port The listening port (80 in the Docker image, 3000 otherwise)
--num-shard How many GPUs to split the model across; defaults to all visible GPUs
--quantize Load weights at lower precision to save memory (for example eetq or fp8)
--max-total-tokens The most tokens one request may use, prompt plus answer
--max-input-tokens The longest prompt accepted
--api-key Require a bearer token on the main endpoints
--json-output Emit logs as JSON, for log pipelines

The most useful idea in this whole section is the token budget. Four numbers govern memory, and they stand in a fixed relationship:

  • max_input_tokens: the largest prompt.
  • max_total_tokens: the prompt plus the generated answer, per request. The docs call it the most important value to set, because it is the memory budget of one request.
  • max_batch_prefill_tokens: the total prompt tokens the server will process in one prefill step.
  • max_batch_total_tokens: all the tokens held in the KV cache across the whole batch.

The rules: the input limit must be smaller than the total limit, the prefill budget must be at least the input limit, and the total limit must not exceed the batch total.

Here is the good news. Since TGI 3.0, you should leave all four unset. The release notes call it "zero config": the server inspects your GPU and your model and chooses these numbers itself, and the official advice is to remove your flags and let it decide. You can see what it chose in /info. Many old tutorials set --max-input-length, --max-batch-prefill-tokens and friends by hand; copying them into version 3 often makes things worse, and the old name --max-input-length is deprecated in favour of --max-input-tokens (setting both is an error).

When do you intervene? When you want to trade context for concurrency. The KV cache is a fixed pool. If you allow every request up to 32,000 tokens, fewer requests fit at once than if you allow 4,000. If your application never sends more than 4,000-token prompts, lowering --max-total-tokens frees memory for more simultaneous users. That is the first real tuning decision you will make, and Mid-level covers it properly.

Quantization deserves a short mention. Quantization stores weights with fewer bits, shrinking memory use so a bigger model fits on a smaller card, at some cost to quality. For a beginner, the safe rule is: if a model repository is already quantized (for example AWQ or GPTQ), TGI reads that from the model's config.json and you pass nothing. If you need TGI to shrink a normal model on the fly, --quantize eetq is the drop-in 8-bit choice, and --quantize fp8 suits H100-class cards. The older --quantize bitsandbytes is deprecated and also disables CUDA graphs, so avoid it.

Finally, two operational flags. --api-key makes the main endpoints require an Authorization: Bearer <key> header, but it does not protect /health, /info, /metrics or /docs, and it provides one shared key, not per-user keys. It is a basic gate, not an authentication system; put TGI behind a gateway if the server is reachable by anyone you do not trust. And TGI sends anonymous usage statistics by default when it runs inside Docker; if your organisation forbids that, add --usage-stats off.

Try it
  1. Run --help against the image and find the entries for --max-total-tokens and --usage-stats.
  2. Restart your server with --max-total-tokens 2048 and read /info.
  3. Restart without it and read /info again; compare.
a total limit of 2048 in the first case and a much larger automatic value in the second.

Common errors and how to read them

When TGI fails at startup, the last line is usually unhelpful: Error: ShardCannotStart, Error: WebserverFailed or Error: DownloadError. These are only the launcher's name for which stage failed. The real cause is a few lines above, usually under a heading such as "Shard complete standard error output". Always scroll up before searching the web.

Here are the errors beginners meet most, with their meaning.

Out of memory during warmup. The message reads Not enough memory to handle {N} prefill tokens. You need to decrease --max-batch-prefill-tokens, followed by Unable to warmup the Python model shards. The weights loaded, but the worst-case prefill did not fit in the remaining memory. Causes: a manually set --max-batch-prefill-tokens or --max-input-tokens that is too large, another process holding the GPU (check nvidia-smi), or a model that is simply too big for the card. Fixes, in order: remove manual limits, lower the prefill budget, use a quantized model, use a smaller model, or use more GPUs.

A second variant. Not enough memory to handle max_total_tokens={N} means one request's full context does not fit in the KV cache left over. Lower --max-total-tokens, or add memory.

Invalid token settings. Messages such as max_input_tokens (X) must be < max_total_tokens (Y) tell you your numbers break the rules from the previous section. Fix the numbers, or better, delete them.

GPUs not found. sharded is true but only found 1 CUDA devices or an error about not having enough CUDA devices means the container cannot see the GPUs you asked for. Check --gpus all and --num-shard. If the second check from the prerequisites section fails, fix Docker's GPU access first.

NCCL or hang with several GPUs. An error mentioning NCCL, or startup that never finishes, usually means /dev/shm is too small. Add --shm-size 1g. As a last resort, NCCL_SHM_DISABLE=1 avoids shared memory at a performance cost.

Flash attention on an old GPU. FlashAttention only supports Ampere GPUs or newer. means your card is too old for that kernel. Older cards generally fall back automatically; if not, use --disable-custom-kernels or newer hardware.

Download failures. Error: DownloadError with a 401, 403 or gated message: a token or licence problem, as above. Without those words, check network access and free disk space in the volume.

A model with only .bin files. Pickle weights are not converted without --trust-remote-code. Prefer a safetensors version of the model. Trust remote code only for repositories you have vetted, and pin a revision when you do, because it runs Python from the repository.

Then there are the HTTP errors your client sees once the server is up. They arrive as JSON with an error message and an error_type.

  • 422, validation. Your request broke a rule. The message tells you which: inputs must have less than {max} tokens, inputs tokens + max_new_tokens must be <= {max_total}, temperature must be strictly positive, top_p must be > 0.0 and < 1.0, or stop supports up to 4 stop sequences. Shorten the prompt, lower max_new_tokens, or fix the parameter.
  • 429, "Model is overloaded". More requests are in flight than --max-concurrent-requests allows (128 by default). Retry with backoff, or add replicas.
  • 424, generation failed. Something broke inside a shard mid-generation, often a CUDA error. Read the server logs.
  • 401 with an empty body. The server runs with --api-key and your header is missing or wrong.
  • 413, payload too large. Your request body exceeds --payload-limit, 2 MB by default. This mostly happens with large base64 images.

Last, a problem that is not an error: the first request is slow when it uses grammar, response_format or tools. TGI compiles the constraint on first use and caches it. Send one warm-up request after a deploy.

The three-step triage. One: find which stage failed from the last Error: line. Two: scroll up to the first real error above it. Three: compare that message against the list here before changing anything.
Try it
  1. Send a request with "temperature": 0 and read the 422 body.
  2. Send one with a max_new_tokens far larger than your max_total_tokens.
  3. Launch a server with --max-total-tokens 100 --max-input-tokens 200 and read the startup error.
two precise 422 messages naming the limit you broke, and a launcher refusing to start because the input limit must be below the total.

Putting it all together

Here is one small end-to-end project: a script that starts a server, waits until it is ready, runs a short chat with streaming, and cleans up. It uses only what this guide has covered.

First, a start script. Save it as start-tgi.sh:

start-tgi.sh
#!/usr/bin/env bash
set -euo pipefail

MODEL="${MODEL:-HuggingFaceTB/SmolLM2-1.7B-Instruct}"
IMAGE="ghcr.io/huggingface/text-generation-inference:3.3.7"

mkdir -p "$PWD/data"

docker run -d --name tgi --gpus all --shm-size 1g \
    -p 8080:80 -v "$PWD/data:/data" \
    -e HF_TOKEN="${HF_TOKEN:-}" \
    "$IMAGE" --model-id "$MODEL" --usage-stats off

echo "Waiting for the server (first start downloads weights)..."
until curl -sf localhost:8080/health > /dev/null; do
    if ! docker ps --format '{{.Names}}' | grep -q '^tgi$'; then
        echo "Container exited. Last logs:"; docker logs --tail 40 tgi; exit 1
    fi
    sleep 5
done
echo "Ready:"; curl -s localhost:8080/info | head -c 400; echo

Notice the loop does two jobs: it waits for health, and it notices if the container died, so a failed start prints the last logs instead of waiting forever. The model default is a small instruction-tuned model; change it with the MODEL variable.

Second, a client script. Save it as chat.py:

chat.py
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8080/v1/", api_key="-")

history = [{"role": "system", "content": "You are a concise assistant."}]

while True:
    question = input("you> ").strip()
    if not question:
        break
    history.append({"role": "user", "content": question})

    answer = ""
    stream = client.chat.completions.create(
        model="tgi", messages=history, max_tokens=300, stream=True
    )
    for chunk in stream:
        piece = chunk.choices[0].delta.content
        if piece:
            answer += piece
            print(piece, end="", flush=True)
    print()
    history.append({"role": "assistant", "content": answer})

This keeps the conversation in a list and resends it on every turn. That is how every chat client works with a server like this: the server remembers nothing between requests. The history is yours. Every turn sends a longer prompt, which is why prefill grows, why long chats get slower, and why prefix caching (reusing the shared start of each prompt) helps so much for exactly this pattern.

Run it:

BASH
chmod +x start-tgi.sh
./start-tgi.sh
python chat.py

When you are done, stop and remove the container; the weights stay in ./data for next time:

BASH
docker rm -f tgi

Now extend it. Add a --max-total-tokens value to the start script and watch /info change. Point chat.py at a different model by changing MODEL. Time the first token. Break something on purpose, such as an invalid token limit, and practise reading the failure from the logs. A beginner who can start a server, diagnose a failed start, and write a streaming client has done the whole job at this level.

Try it
  1. Create both files and run the project end to end with a small model.
  2. Hold a five-turn conversation, and notice whether replies slow down.
  3. Run docker rm -f tgi, run the start script again, and compare the start time.
a working streamed conversation, and a much faster second start because the weights are cached in the volume.

What you can now do, and what comes next

You can now do the whole beginner loop. You know that TGI is a Linux and GPU server made of a launcher, a router and a model server, and that it is a frozen, archived product at version 3.3.7. You can start it with a pinned image and a persistent volume, wait for it with a health check, and read /info to see the token limits it chose. You can call it three ways: the native /generate endpoint, the OpenAI-compatible chat endpoint, and streaming. You can supply a Hub token for gated models, choose a revision, and use a local model. And you can read a failed startup, telling a memory problem from a download problem from a GPU-visibility problem.

What the other levels add:

  • Mid-level goes under the hood: how continuous batching and the paged KV cache decide capacity, quantization, multi-GPU tensor parallelism, structured output with grammars, tool calling, monitoring with Prometheus, and running TGI on Kubernetes.
  • Senior treats the server as a platform: capacity planning, the security model (including what --api-key does and does not protect), upgrade policy for a frozen project, and a migration plan to vLLM or SGLang.

Where to look next in this catalogue: vLLM and SGLang are the engines Hugging Face recommends going forward, and the concepts here map across directly. If you want a gentler local-first path, try Ollama or llama.cpp. To package or orchestrate the container you now run by hand, the Docker and Kubernetes guides are the natural follow-ups, and LiteLLM is a common gateway to put in front of a server like this.

One habit to carry away: write down the version, the image tag and the model revision of anything you deploy. With a frozen project that is easy to do and very valuable, because it is the information you will need on the day you migrate.

Sources