Skip to content
Back to student guides
vLLMLLMsInference & gateways3 levels101 sectionsCovers vLLM 0.30

The Complete vLLM Guide

Serve LLMs with high throughput using vLLM’s PagedAttention and OpenAI-compatible server. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
15sections
18examples

This is part one of three, and it is written for someone who has never run a large language model on their own hardware. By the end you will be able to install vLLM, serve an open-weight model behind an HTTP endpoint that any OpenAI client can talk to, call it from curl and Python, run the same model offline from a script, size its memory so it starts reliably, and read the errors it throws when it does not. The guide is checked against vLLM 0.30.0, released in September 2026. vLLM ships a minor release roughly every two weeks and removes flags as it goes, so a lot of tutorials you will find online show options that no longer exist. Where that matters, this guide says so.

Each section ends with a Try it task. Do them. A server that loads a model onto a GPU teaches you more in one failed start than a page of reading, and the memory numbers in this guide only mean something once you have watched your own machine produce them.

What vLLM is, and the problem it solves

A large language model is a file of numbers (the weights) plus a program that, given some text, predicts the next piece of text one token at a time. A token is a chunk of a word, roughly three to four characters of English. To answer a question, the model predicts a token, appends it to the input, predicts the next, and repeats until it decides to stop. Producing a 300-token answer means running the model 300 times in sequence.

You can run that loop yourself with the Hugging Face transformers library, and for one user in a notebook it works fine. The trouble starts when several people use it at once. A GPU is an expensive machine that is only efficient when it has a lot of work to do in parallel. Serving one request at a time leaves most of it idle. Naive batching, where you wait for eight requests, run them together, and return them all when the slowest one finishes, wastes time on the short answers that finished early. And the model needs a large working memory for every request in flight, which a naive server reserves in big contiguous slabs sized for the worst case, so most of that memory sits empty.

vLLM is an inference and serving engine built to fix exactly those three problems. It is an open-source project that started at UC Berkeley and is now developed by a large community. You give it a model, it loads the weights onto your GPU or GPUs, and it serves many requests at once while keeping the hardware busy. It does this with two main ideas that you will meet in the next section: continuous batching and PagedAttention.

What came before it was a mix of approaches. You could write your own serving loop around transformers, which is flexible but slow under load. You could use a dedicated server such as Text Generation Inference, or a model-specific runtime. Or you could call a hosted API and never think about GPUs at all. vLLM sits in the middle: it is a general-purpose server for open-weight models that you run yourself, and it speaks the same HTTP dialect as the OpenAI API, so code written for a hosted API can be pointed at your own machine by changing one URL.

Why would you run a model yourself? Cost at volume is one reason. Control is another: you pick the model, the version, and the behaviour. A third reason is data residency. Many employers in the Gulf and in Egypt need prompts and documents to stay inside a particular country or a private network, and a model served from your own infrastructure never sends data to a third party. vLLM does not make that decision for you, but it makes the self-hosted option practical.

YOUR APPOpenAI client, curl
→
vllm serveHTTP on port 8000
→
SCHEDULERbatches every step
→
GPUweights + KV cache

The diagram is the whole shape of the tool. Requests arrive over HTTP, a scheduler decides which requests share the next step on the GPU, and the GPU holds both the model and a pool of working memory. Everything else in this guide is detail about one of those four boxes.

Try it
  1. Write down one project of yours that would call a language model, and who would be unhappy if its prompts left your network.
  2. Decide how many users might ask something at the same moment. That number is what makes serving hard.

The mental model: four ideas you must know

vLLM has a large surface, with more than two hundred command-line flags, but the behaviour you need to predict as a beginner comes from four ideas. If you hold these, almost every error message will make sense.

First, the KV cache. When a model reads your prompt, it computes for every token two vectors called a key and a value, at every layer. The next token's prediction needs to look back at all of them. Recomputing them for every new token would be hopelessly slow, so the engine stores them. That store is the KV cache. It grows by one token's worth of keys and values for every token the request has seen, and it lives in GPU memory. For most deployments the size of the KV cache, not the size of the weights, decides how many requests you can serve at once. Remember that sentence; it explains most memory errors.

Second, PagedAttention. Older servers reserved one big contiguous block of memory per request, big enough for the longest possible answer, and most of it went unused. vLLM instead cuts the KV cache into small fixed-size blocks and hands them out on demand, the way an operating system hands out pages of RAM. A request holds a list of blocks (a block table) that need not be next to each other. Nothing is reserved in advance, nothing is stranded between requests, and memory is returned the moment a request finishes. This is the idea the project is named for and the reason it fits many more concurrent requests on the same GPU.

Third, continuous batching. A server that batches whole requests waits for the slowest one. vLLM's scheduler instead decides, at every single token step, which requests run next. A request that finishes leaves immediately and a waiting one takes its place on the very next step. The GPU is therefore always working on a full batch. The scheduler works with a token budget per step (the flag --max-num-batched-tokens) and a cap on concurrent requests (--max-num-seqs).

Fourth, prefill and decode. Every request has two phases. Prefill reads your whole prompt in one go; it is heavy on computation and determines how long you wait for the first token (time to first token, or TTFT). Decode then produces the answer one token per step; it is limited by memory bandwidth and determines how fast tokens stream out. Long prompts are split into chunks so they do not block the decoding of other users, a feature called chunked prefill that is on by default. Another default is automatic prefix caching: if two requests start with the same text, such as the same system prompt, the engine reuses the KV blocks it already computed for it and skips that part of the prefill.

Why your first token is slower than the rest If you notice a pause before text starts and then a fast stream, that is prefill followed by decode. A long prompt makes the pause longer. A repeated prompt prefix makes it shorter, because of prefix caching.

There are also a few nouns for the software itself. The LLM class is the Python interface for running a model inside your own script with no server, called offline inference. The command vllm serve starts an OpenAI-compatible server and is called online serving. Behind it, an API server process handles HTTP and tokenization, an engine core process runs the scheduler and manages the KV cache, and one worker process per GPU holds the weights and runs the model. You will see these names in logs. You do not need to manage them.

Try it
  1. In your own words, explain why a server that serves ten users at once needs more memory than one that serves a single user, even though the model weights are the same size.
  2. Explain what continuous batching does that "wait for eight requests, then run them together" does not.

What hardware you need

The first question every beginner asks is whether they can try this on a laptop. The honest answer depends on the machine, and the official platform matrix is worth knowing before you spend an afternoon on a setup that cannot work.

The primary, best-supported platform is Linux with an NVIDIA GPU of compute capability 7.5 or higher. That covers the T4, RTX 20-series and newer, L4, A100, H100 and B200. AMD GPUs on Linux (the MI200, MI300 and MI350 families and recent Radeon cards) are supported through ROCm. Intel GPUs and Google TPUs have their own, more specialised paths. CPU-only Linux is supported, with prebuilt wheels, and is useful for learning the API with a tiny model, though it is slow.

Windows is not supported natively. The documentation is direct about this. The supported route is WSL2, the Windows Subsystem for Linux, where you install the NVIDIA driver on the Windows side and then follow the Linux steps inside an Ubuntu shell. Docker Desktop with the WSL2 backend and an NVIDIA GPU also works, since it runs the Linux image.

macOS on Apple silicon is experimental. vLLM runs on the CPU backend there and only from a source build, with 32-bit and 16-bit floating point. It is fine for learning the API with a tiny model and not a path to good performance. For Apple GPU acceleration the docs point to a community plugin called vllm-metal, which is separate from the core project. If your goal is simply to chat with a model on a MacBook, a tool such as Ollama or llama.cpp is a better fit; come back to vLLM when you have a Linux GPU box or a cloud instance.

How much GPU memory do you need? A rough rule: a model stored in 16-bit precision needs about two bytes per parameter for the weights, so a 7-billion-parameter model needs about 14 GB, plus room for the KV cache and some overhead. A 24 GB card runs a 7B or 8B model comfortably with a moderate context. A 0.6-billion-parameter model such as Qwen3-0.6B, which this guide uses for every example, needs little more than a gigabyte and a half for weights and runs on almost any NVIDIA card with a few GB free. That is why it is the official smoke-test model: you prove the whole pipeline works before you download a model that takes half an hour.

No GPU at home? Rent one for an hour An hour on a small cloud GPU instance costs little and is enough to finish this guide. Pick a Linux image with the NVIDIA driver preinstalled and check with nvidia-smi before you do anything else. If your employer needs data to stay in a given country, choose a region there for real work; for these exercises with a public model and no private data, any region will do.
Try it
  1. Decide which row of the platform list you are on: Linux with NVIDIA, WSL2, a Mac, or a rented machine.
  2. If you have an NVIDIA GPU, run nvidia-smi and write down the GPU name, total memory and driver version.

Installing vLLM and checking the setup

vLLM depends on one specific version of PyTorch, and its wheels contain compiled GPU kernels built against that exact version. That is why the install guidance is strict: use a clean virtual environment, and do not install or upgrade PyTorch yourself afterwards. The documented path uses uv, a fast Python package manager, and Python 3.12. The package metadata allows 3.10 to 3.14 and the docs mention 3.10 to 3.13, but 3.12 is what the docs' own example uses and what the AMD and Intel builds require, so teach yourself 3.12 and avoid the question.

On Linux with an NVIDIA GPU:

BASH
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

The --torch-backend=auto flag tells uv to look at your installed driver and pick the matching PyTorch build. The default vLLM wheel is built for CUDA 12.9. If the install fails with an index error, update uv itself with uv self update and retry. If you do not have uv, install it first from its official site, or use plain pip with the extra index URL documented on the installation page. In a course, pin the version so that your results match this guide:

BASH
uv pip install vllm==0.30.0 --torch-backend=auto

Do not install the CUDA Toolkit expecting it to be needed. You need an NVIDIA driver that is recent enough for the CUDA runtime in the wheel; the toolkit with the nvcc compiler is only for building vLLM from source.

Now verify, step by step, from the cheapest check to the most expensive:

BASH
nvidia-smi
python -c "import vllm; print(vllm.__version__)"
vllm --version
vllm collect-env

nvidia-smi shows the driver sees your GPU. The next two print the installed version, which should be 0.30.0. vllm collect-env prints a full report of your environment, including the PyTorch, CUDA and driver versions, and it is exactly what maintainers ask you to paste into a bug report, so it is worth running once to see what it contains.

If you are on Windows, open PowerShell and run wsl --install, restart, install the NVIDIA Windows driver (do not install a Linux driver inside WSL, since WSL passes the GPU through), open the Ubuntu shell, and follow the Linux steps above inside it. If a source build runs out of memory under WSL, export MAX_JOBS=1 slows the build but lowers the peak. On a CPU-only Linux machine, the project publishes CPU wheels and CPU Docker images; the CPU installation page gives the exact wheel URL, and you set VLLM_CPU_KVCACHE_SPACE to the number of gigabytes you want for the KV cache.

Two small habits belong here. First, gated models such as Meta's Llama need you to accept a licence on Hugging Face and then supply a token; export it as HF_TOKEN in the shell that runs vLLM. The model we use, Qwen3-0.6B, is not gated and needs no token. Second, vLLM collects anonymous usage statistics by default. If you or your employer want that off, set VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1, or create the file ~/.config/vllm/do_not_track.

Do not "fix" the install by upgrading torch If you run pip install torch or a library that drags in a different torch after installing vLLM, the compiled kernels no longer match and importing vLLM fails with undefined-symbol errors mentioning vllm._C. The cure is a clean virtual environment and a fresh uv pip install vllm --torch-backend=auto. Treat the environment as owned by vLLM.
Try it
  1. Create the virtual environment and install vLLM as shown, pinned to 0.30.0.
  2. Run all four verification commands and confirm the version prints 0.30.0.
  3. Open the vllm collect-env output and find your PyTorch version and your CUDA version.

Your first program: offline inference

The smallest working vLLM program does not need a server at all. It loads a model into your own Python process, generates text for a list of prompts, and exits. This is called offline inference and it is ideal for batch jobs, such as classifying ten thousand support tickets overnight, and for first experiments.

Save this as first.py:

first.py
from vllm import LLM, SamplingParams

if __name__ == "__main__":
    llm = LLM(model="facebook/opt-125m")
    params = SamplingParams(temperature=0.8, top_p=0.95, max_tokens=40)
    outputs = llm.generate(["The capital of France is"], params)
    print(outputs[0].outputs[0].text)

Run it with python first.py. The first run downloads the model from the Hugging Face Hub into a cache under your home directory, so it takes a moment. You will see a long stream of log lines as vLLM loads the weights, compiles the model, captures CUDA graphs, and measures how much memory is left for the KV cache. Then a short continuation of the sentence prints. The model facebook/opt-125m is tiny and not instruction-tuned, so do not expect a clever answer; the point is that the pipeline works.

Read the code from the top. LLM(model=...) is the engine: constructing it is the slow step, because it is where the weights load and the GPU memory is planned. Construct it once and reuse it. SamplingParams holds the decoding settings for a request, covered below. llm.generate takes a list of prompts and returns a list of result objects in the same order. Each result has an outputs list (one entry per requested completion), and each entry has .text. You pass many prompts in a single call and vLLM batches them for you; looping over generate with one prompt at a time throws away the very thing the library is good at.

The if __name__ == "__main__": guard is not optional. vLLM starts worker processes, and on platforms where child processes re-import your script, an unguarded LLM(...) at the top of the file starts the engine recursively. The symptom is RuntimeError: An attempt has been made to start a new process before the current process has finished its bootstrapping phase, sometimes preceded by a warning that CUDA was previously initialized and the spawn method is being forced. Put everything that creates the LLM object inside the guard and the error disappears.

For an instruction-tuned model, you usually want to give it a conversation rather than raw text, because the model was trained on a specific chat format. The chat method applies the model's own chat template for you:

chat_offline.py
from vllm import LLM, SamplingParams

if __name__ == "__main__":
    llm = LLM(model="Qwen/Qwen3-0.6B")
    messages = [
        {"role": "system", "content": "You answer in one short sentence."},
        {"role": "user", "content": "What is a GPU used for?"},
    ]
    out = llm.chat(messages, SamplingParams(temperature=0.7, max_tokens=200))
    print(out[0].outputs[0].text)

A chat template is a small recipe, shipped inside the model's tokenizer files, that turns a list of role-tagged messages into the exact string of special tokens the model was trained to expect. Using the wrong format makes a good model answer badly, which is why chat and the server's chat endpoint exist. Qwen3 is a model that can print its reasoning inside think tags before answering, so you may see some of that in the text; that is the model, not vLLM.

Always guard the entry point A script that builds an LLM must keep that line under if __name__ == "__main__":. Missing it is the most common reason a first offline script prints a bootstrapping error instead of text.
Try it
  1. Run first.py. Find, in the log, the line or lines that report the memory available for the KV cache.
  2. Change the prompt list to three prompts and print every result. Confirm the order matches.
  3. Run chat_offline.py, then change the system message and watch the answer change.

Your first server

For most real use you want vLLM as a service that other programs call. The command is vllm serve followed by the model name:

BASH
vllm serve Qwen/Qwen3-0.6B

The older form python -m vllm.entrypoints.openai.api_server was deprecated in 0.29; if a tutorial tells you to use it, use vllm serve instead and treat the rest of the tutorial with a little suspicion.

The server prints a lot at startup. Roughly, it downloads the model if it is not cached, loads the weights, compiles the model and captures CUDA graphs (a one-time cost that makes later steps faster), measures free memory, and sizes the KV cache. When it is ready, you will see lines saying the application startup is complete and that the server is listening on port 8000. Startup for a small model takes tens of seconds to a couple of minutes. For a large model it can take many minutes, which surprises people who expect a web server to be instant.

While it runs, open a second terminal and check three endpoints:

BASH
curl http://localhost:8000/health
curl http://localhost:8000/v1/models
curl http://localhost:8000/version

/health returns an HTTP 200 with an empty body once the engine is ready, and is the check a load balancer or a Kubernetes probe would use. /v1/models lists the model names the server will answer to, which is important in the next section. /version returns the vLLM version.

Notice what the server does not do. By default it binds all network interfaces, since --host defaults to nothing, which means any machine that can reach yours can reach port 8000. On a laptop on a café network, or a cloud machine with an open security group, that is an unauthenticated model endpoint open to the world. For anything beyond a throwaway test, set the host explicitly:

BASH
vllm serve Qwen/Qwen3-0.6B --host 127.0.0.1 --port 8000

With 127.0.0.1 only processes on the same machine can connect. We return to authentication and safe exposure later.

Default binding is wider than you think Without --host, vLLM listens on every interface. Use --host 127.0.0.1 while learning, and put a gateway in front of it, rather than the bare server, when anything else needs to reach it.
Try it
  1. Start vllm serve Qwen/Qwen3-0.6B --host 127.0.0.1 and watch the log until the server is ready.
  2. Time how long that took. Stop the server with Ctrl+C and start it again; compare the two times.
  3. Run the three curl checks and read the response from /v1/models.

Talking to the server

The server speaks the OpenAI API dialect, so you can use the official openai Python package, plain curl, or any tool that accepts a custom base URL. Start with curl, because it shows exactly what crosses the wire:

BASH
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-0.6B",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}]
  }'

The reply is a JSON object. The answer lives at choices[0].message.content. There is also a usage object with prompt_tokens, completion_tokens and total_tokens, which is how you learn what a request cost in tokens, and a finish_reason: stop means the model decided it was done, while length means it hit your max_tokens limit and was cut off.

The same call from Python with the OpenAI client, installed alongside vLLM:

client.py
from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

response = client.chat.completions.create(
    model="Qwen/Qwen3-0.6B",
    messages=[{"role": "user", "content": "Tell me a joke about GPUs."}],
    max_tokens=200,
)
print(response.choices[0].message.content)

Two details trip people up. The base_url must end in /v1, because that is the path prefix of the OpenAI-style routes. And the api_key is required by the client library even though the server is not checking one yet, so any placeholder string works.

The model field must match a name the server knows. By default that is the exact model path you gave vllm serve. If you send a different name, such as gpt-4, the server replies with HTTP 404 and a message that the model does not exist. You can choose a friendlier public name with --served-model-name:

BASH
vllm serve Qwen/Qwen3-0.6B --served-model-name tiny-chat --host 127.0.0.1

Clients now send "model": "tiny-chat". This is useful because it lets you swap the model behind the name without touching client code, which is the same trick a hosted API plays.

To stream tokens as they are produced, which is what makes chat interfaces feel fast, add stream=True and iterate over the chunks:

stream.py
from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

stream = client.chat.completions.create(
    model="Qwen/Qwen3-0.6B",
    messages=[{"role": "user", "content": "Explain KV cache in two sentences."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()

Besides /v1/chat/completions, the server also offers /v1/completions, the older raw-text endpoint that takes a plain prompt and applies no chat template, and /v1/embeddings for models built to produce embeddings. For a chat model, use the chat endpoint. vLLM also has convenience clients in the terminal: with a server running, vllm chat --quick "hi" sends a one-shot message to it.

Because the interface is standard, the rest of the ecosystem plugs in. Frameworks that accept an OpenAI base URL work unchanged, and a gateway such as LiteLLM can sit in front of one or many vLLM servers to give you keys, budgets and routing. If you have used the hosted API described in the OpenAI API guide, the only changes are the base URL and the model name.

Try it
  1. With the server running, make the curl request and find finish_reason and the token counts in the reply.
  2. Run client.py, then stream.py, and notice how the second prints incrementally.
  3. Restart with --served-model-name tiny-chat and make the old model name fail with a 404. Read the message.

Controlling the output: sampling parameters

The model does not output a word. It outputs a probability for every token in its vocabulary, and a sampling step chooses one. The settings that govern the choice are the main way you control how an answer reads. In the offline API they live in SamplingParams; through the server you pass them in the request body, with the vLLM-specific ones inside extra_body when you use the OpenAI client.

Temperature reshapes the probabilities. Near zero, the model almost always picks the single most likely token, so answers are focused and repeatable. At 1.0 it samples more freely, and above that answers get erratic. top_p (nucleus sampling) keeps only the smallest set of tokens whose probabilities add up to that fraction, so top_p=0.9 ignores the long tail of unlikely words. top_k keeps only the k most likely tokens; vLLM accepts it, though it is not part of the standard OpenAI API, so with the OpenAI client you send it as extra_body={"top_k": 20}. max_tokens caps the length of the answer, and stop gives strings that end generation when produced. n asks for several independent completions of one prompt. seed makes sampling reproducible for a given setup.

params.py
from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

r = client.chat.completions.create(
    model="Qwen/Qwen3-0.6B",
    messages=[{"role": "user", "content": "Name three colours."}],
    temperature=0.2,
    top_p=0.9,
    max_tokens=60,
    extra_body={"top_k": 20},
)
print(r.choices[0].message.content)

A subtle and important default: if you do not send a setting, vLLM may fill it in from the model's own generation_config.json. Model authors ship recommended values for temperature, top_p and similar, and since version 0.8 vLLM applies them as the server's defaults. This is helpful, since you get the author's intended behaviour, but it explains the confusing report "I changed nothing and the answers changed after an upgrade" or "the same prompt behaves differently from another server". If you want vLLM's own neutral defaults instead, start the server with --generation-config vllm. In offline code the equivalent is LLM(model=..., generation_config="vllm"). Being explicit by always sending the parameters you care about is the safest habit.

Reproducible answers For tests, send temperature=0 and a fixed seed. Expect the same answer on the same setup; do not promise identical text across different GPUs, library versions or batch sizes, because floating-point arithmetic can differ in small ways that occasionally flip a close token choice.
Try it
  1. Ask the same creative question five times at temperature 0 and five times at 1.0. Compare the variety.
  2. Set max_tokens=5 and look at the finish_reason.
  3. Restart the server with --generation-config vllm and see whether the unset defaults change the behaviour.

The flags you will actually use

vllm serve accepts hundreds of flags, which is intimidating. A beginner needs about ten. The server also lets you explore the rest yourself: vllm serve --help=listgroup lists the flag groups, vllm serve --help=<GroupName> prints one group, and vllm serve --help=<flag> explains a single flag.

Here is a realistic first configuration, and then each flag explained:

BASH
vllm serve Qwen/Qwen3-0.6B \
  --host 127.0.0.1 --port 8000 \
  --served-model-name tiny-chat \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.85 \
  --api-key "$VLLM_API_KEY"

--host and --port choose where the server listens, as above. --served-model-name sets the name clients use. --api-key makes the server require an Authorization: Bearer <key> header on its main API routes; you can also set the VLLM_API_KEY environment variable instead of putting the key on the command line, which keeps it out of the process list that other users on a shared machine can see.

--max-model-len is the maximum number of tokens in one request, counting both the prompt and the answer. If you omit it, vLLM uses the context length from the model's configuration, which for modern models can be tens of thousands of tokens or more. A long maximum is not free: the engine must be able to fit at least one full-length request in the KV cache. Setting it to what you actually need, such as 8192 here, is the single most common fix for memory startup errors. The flag accepts friendly forms like 8k.

--gpu-memory-utilization is the fraction of the GPU's total memory that this vLLM instance may use for weights, working space and KV cache combined. The default is 0.92. Many older guides say 0.9, which was the previous default. It is a cap on the instance, not "extra" memory on top of the model, and it is a fraction of the total, not of what is currently free. If something else already holds memory on the GPU, a high value can fail at startup. Whatever is left after the weights and working memory becomes the KV cache, so raising the number gives you more concurrency and lowering it gives you headroom for other programs.

--dtype sets the numeric precision of the weights. The default, auto, uses the precision recorded in the model, normally bfloat16. Older GPUs such as the T4 (compute capability 7.5) do not support bfloat16; there you pass --dtype half. --trust-remote-code allows a model that ships its own Python code to run that code. It is off by default for a good reason: a model repository can contain arbitrary code, so only enable it for repositories you trust. --tensor-parallel-size splits one model across several GPUs of one machine when it does not fit on one; it is a topic for the next level. --download-dir and --revision control where weights are cached and which version of the repository is used.

--enforce-eager turns off CUDA graph capture. Startup is faster and memory use is slightly lower, at the cost of slower generation. It is useful for debugging and quick experiments, not for production.

Set deliberately

  • --host 127.0.0.1 while learning
  • --max-model-len sized to your real prompts
  • an explicit --served-model-name
  • the key read from VLLM_API_KEY

Left to chance

  • binding every interface by default
  • a context length of tens of thousands you never use
  • clients depending on a long model path
  • a secret typed straight on the command line

If you find yourself repeating the same long command, move it into a YAML file whose keys are the flag names without the dashes, and run vllm serve --config config.yaml. Anything you also type on the command line wins over the file.

Try it
  1. Run vllm serve --help=listgroup and pick one group to print with --help=<GroupName>.
  2. Start the server with the full configuration above, setting VLLM_API_KEY in your shell first.
  3. Send a request without the key and then with -H "Authorization: Bearer $VLLM_API_KEY". Compare the status codes.

Memory, the KV cache and reading the startup log

Because memory causes most beginner failures, it is worth understanding how vLLM budgets it. At startup the engine does three things in order. It loads the weights. It runs a dummy pass through the model to measure how much temporary working memory the computation needs. Then it takes the budget (the total memory times --gpu-memory-utilization), subtracts the weights and the working memory, and gives everything left to the KV cache, which it carves into blocks.

The startup log reports this. You will see a line giving the memory the model weights took, and lines reporting the GPU KV cache size in tokens and the maximum concurrency for requests of your maximum length. Read the KV cache size as the total number of tokens that can be in flight across all requests at once. If it says the cache holds 50,000 tokens and each request uses about 1,000 tokens, you can serve around fifty requests together. If requests are long, the number of simultaneous requests drops; if they are short, it rises. This is the quantity you tune.

A tiny numeric example makes it concrete. Imagine a GPU with 24 GB, a utilization of 0.9 (21.6 GB budget), weights of 16 GB and working memory of 2 GB. That leaves about 3.6 GB for the KV cache. Raise utilization to 0.95 and the budget grows to 22.8 GB, so the KV cache grows to 4.8 GB, a third more room for requests. Cut --max-model-len and nothing about that pool changes, but you stop the engine rejecting startup because it could not fit one huge request. The numbers here are illustrative, not measurements from a specific model.

What happens when the cache fills mid-run? The scheduler preempts a request: it evicts its blocks and later recomputes them when there is room again. In the current engine this is by recomputation, not by swapping memory to the CPU, so the old --swap-space flag no longer exists. Preemption is not an error, but it costs speed and shows up as a warning. The cure is more KV memory: raise the utilization, shorten the maximum length, quantize the model, or use more GPUs.

Two other things appear in the log and are worth recognising. One is the compilation and CUDA graph capture step, which can take a while the first time and is cached afterwards under ~/.cache/vllm, the directory named by VLLM_CACHE_ROOT. The other is the periodic stats line that reports throughput and the number of running and waiting requests, a handy live view of load.

The cache is a directory, not magic Downloaded weights live in the Hugging Face cache (by default under ~/.cache/huggingface) and compile artifacts live in ~/.cache/vllm. When you run in a container, mount both on a volume, or every start downloads and compiles again.
Try it
  1. Start the server and find the log lines about the weights memory and the KV cache size.
  2. Restart with --gpu-memory-utilization 0.5 and compare the KV cache size.
  3. Restart with a very small --max-model-len 512 and send a long prompt. Read the 400 error.

Common errors and how to read them

Most vLLM failures fall into a small set. When a start fails, the last line is often a generic wrapper, RuntimeError: Engine core initialization failed. See root cause above. Read that literally: scroll up a few lines. The real exception is above it, and that is the one to search for.

Not enough memory for one full request. The message begins To serve at least one request with the model's max seq len (131072), (16.00 GiB KV cache is needed, which is larger than the available KV cache memory (9.23 GiB). Based on the available memory, the estimated maximum model length is 75584. Your numbers will differ. It means the model's default context is longer than the leftover memory can hold. It even tells you a length that fits. Pass that as --max-model-len, or a smaller figure, raise --gpu-memory-utilization, or choose a smaller or quantized model.

No room for the cache at all. No available memory for the cache blocks. Try increasing gpu_memory_utilization when initializing the engine. The weights and working memory already exceed the budget. A bigger GPU, more GPUs, or a smaller model are the real fixes; raising utilization helps only if you were just under the limit.

The GPU is already in use. Free memory on device cuda:0 (10.5/79.2 GiB) on startup is less than desired GPU memory utilization (0.92, 72.9 GiB). Another process holds memory: a second vLLM, a notebook, or a zombie from a crashed run. Run nvidia-smi, find the process, and stop it, or lower the utilization. Two instances on one GPU need roughly 0.45 each.

Wrong model name, HTTP 404. The model `gpt-4` does not exist. The model in the request does not match the served name. Check /v1/models.

Prompt too long, HTTP 400. This model's maximum context length is 4096 tokens. However, you requested 1000 output tokens and your prompt contains 3500 input tokens, for a total of 4500 tokens. Prompt plus requested output must fit within the maximum length. Shorten one of them or serve with a larger --max-model-len if memory allows.

bfloat16 on an old GPU. Bfloat16 is only supported on GPUs with compute capability of at least 8.0. Your Tesla T4 GPU has compute capability 7.5. You can use float16 instead by explicitly setting the dtype flag. Add --dtype half.

Custom code needed. Failed to load the tokenizer. If the tokenizer is a custom tokenizer not yet available in the HuggingFace transformers library, consider setting trust_remote_code=True. Either upgrade the libraries or add --trust-remote-code, and only after reading what the repository's code does.

Unguarded offline script. The bootstrapping RuntimeError described earlier. Add the __main__ guard.

Device cannot be found. RuntimeError: Failed to infer device type usually means the wrong wheel (a CPU build on a GPU machine, or the reverse) or no visible GPU. Run VLLM_LOGGING_LEVEL=DEBUG vllm serve ... for more detail.

A hang with no error. Usually a download or a slow disk. Pre-download the model with the Hugging Face CLI, and keep weights on a local disk rather than a network share.

For anything unclear, VLLM_LOGGING_LEVEL=DEBUG gives much more output, and vllm collect-env gives the facts a maintainer will ask for.

Use the error's own suggestion first vLLM's memory errors are written to be actionable. They often name the flag to change and print an estimate of a length that fits. Try that before changing anything else.
Try it
  1. Cause the 404 by sending the wrong model name, then fix it with /v1/models.
  2. Cause the 400 by setting a tiny --max-model-len and sending a long prompt.
  3. Start a second server on the same GPU and read the free-memory error. Fix it by lowering the utilization of both.

Running vLLM in Docker

Once the server works on your machine, the natural next step is a container, so the environment is the same everywhere. If Docker is new to you, read the Docker guide first. The project publishes an image named vllm/vllm-openai, and everything you put after the image name is passed to vllm serve:

BASH
docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HF_TOKEN=$HF_TOKEN" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:v0.30.0 \
  --model Qwen/Qwen3-0.6B

Each piece has a reason. --runtime nvidia --gpus all gives the container access to the GPU through the NVIDIA container toolkit, which must be installed on the host. The -v mount shares your Hugging Face cache so the model downloads once, not at every start. HF_TOKEN passes your token for gated models. -p 8000:8000 publishes the port. --ipc=host (or a generous --shm-size) matters because PyTorch and the GPU communication library use shared memory, and Docker's small default makes multi-GPU runs fail with obscure errors. Finally, a pinned tag such as v0.30.0 is better than latest, because latest will move under you and a restart might quietly give you a new version with removed flags.

Add -v vllm-cache:/root/.cache/vllm to persist the compile cache so restarts do not recompile. If you publish the port to a network, remember that the container binds all interfaces; publish as -p 127.0.0.1:8000:8000 to keep it local, and place a gateway in front for anything shared. For the Kubernetes version of this deployment, see the Kubernetes guide, and note the official warning to avoid naming a Kubernetes Service vllm, since Kubernetes injects environment variables named VLLM_* that collide with vLLM's own settings.

Try it
  1. Run the container on a machine with the NVIDIA container toolkit and call /v1/models.
  2. Stop and start it again, and notice the second start reuses the downloaded model.
  3. Republish the port on 127.0.0.1 only and check you can still reach it locally.

Putting it all together

Here is one small project that uses everything above: a local, key-protected summarisation service and a client that summarises text files. The goal is a repeatable setup you can hand to a colleague.

First write a configuration file, so the command line stays short:

config.yaml
model: Qwen/Qwen3-0.6B
host: "127.0.0.1"
port: 8000
served-model-name: summariser
max-model-len: 8192
gpu-memory-utilization: 0.85

Start it with the key in the environment, not in the command:

BASH
export VLLM_API_KEY="change-me-for-real-use"
vllm serve --config config.yaml

Then write the client. It reads the key from the same environment variable, so no secret is in the code:

summarise.py
import os
import sys
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["VLLM_API_KEY"],
    base_url="http://127.0.0.1:8000/v1",
)

def summarise(text: str) -> str:
    r = client.chat.completions.create(
        model="summariser",
        messages=[
            {"role": "system", "content": "Summarise the user's text in three bullet points."},
            {"role": "user", "content": text},
        ],
        temperature=0.2,
        max_tokens=200,
    )
    return r.choices[0].message.content

if __name__ == "__main__":
    for path in sys.argv[1:]:
        with open(path, encoding="utf-8") as f:
            print(f"== {path}")
            print(summarise(f.read()))

Run python summarise.py notes.txt report.txt. Because the system prompt is the same for every call, prefix caching reuses its blocks after the first request, a small free speed-up. Check the server is healthy with curl http://127.0.0.1:8000/health before you start a batch. Note that the health route is not behind the key, which is a hint of something bigger: the API key only protects the main API routes, not every endpoint, so a bare vLLM server is not a safe thing to expose to the internet. Put it behind a reverse proxy that authenticates callers and only forwards the paths you intend.

Finally, if you want a single place for monitoring, the server exposes Prometheus metrics at /metrics, including running and waiting request counts and KV cache usage. Scrape them with Prometheus and chart them in Grafana. Like the health route, this endpoint is unauthenticated, so keep it on a private network.

Try it
  1. Build the project exactly as shown and summarise two text files of your own.
  2. Run curl http://127.0.0.1:8000/metrics and find the metric for waiting requests.
  3. Launch three copies of the client at once and watch the running-request count in the server's stats line.

What you can now do, and what comes next

You can now explain what vLLM is for and why it fits many more requests on one GPU than a naive loop: PagedAttention for memory, continuous batching for scheduling, prefix caching for repeated prompts. You can install a pinned version into a clean environment and verify it. You can run a model offline from a guarded script, and serve one with vllm serve, call it from curl and the OpenAI client, stream the answer, set sampling parameters, and understand where the unset defaults come from. You can budget memory with --max-model-len and --gpu-memory-utilization, read the startup log, run the same thing in Docker, and recognise the most common errors.

Some habits to carry forward: pin versions, bind to localhost until you have a gateway, size the context to your real needs, and read the error above the error.

The mid-level guide takes the same tool into real deployments: configuration files, tool calling, structured outputs, LoRA adapters, quantization, tuning the scheduler and benchmarking with vllm bench. The senior guide covers multi-GPU and multi-node scaling, security and the trust model, upgrades, cost and running vLLM as a platform. Neighbouring guides worth reading are Text Generation Inference and SGLang as alternative servers, LiteLLM as a gateway, KServe for running models on Kubernetes, and Langfuse for tracing the applications that call your model.

Sources