Skip to content
Back to student guides
DeepSpeedLLMsFine-tuning & training3 levels115 sectionsCovers DeepSpeed 0.19

The Complete DeepSpeed Guide

Train and serve large models across many GPUs with DeepSpeed’s ZeRO and parallelism. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
15sections
20examples

This is part one of three. It takes you from never having used DeepSpeed to running real multi-GPU training with it, and explaining what it did and why. By the end you can install it and check the install, write a training script that uses it, read and write its JSON configuration, choose a ZeRO stage and a precision mode, offload to CPU when memory runs out, launch on one machine, save and recover checkpoints, plug it into Hugging Face Trainer, and read its most common error messages without panic. Mid-level and Senior take the same ground further; nothing here is thrown away.

This guide was checked against DeepSpeed 0.19.7, released on 16 September 2026. DeepSpeed is still a 0.x project, which means minor and even patch releases can remove features, and 0.19.7 did exactly that. Where something changed recently, the guide says so and shows the current form.

Each section ends with a Try it task. You do not need a cluster for most of them. A single NVIDIA GPU is enough to run every example, and where an example needs two GPUs the text says so.

What DeepSpeed is, and the problem it solves

Training a neural network on one GPU is easy to picture. The model's weights sit in GPU memory, a batch of data goes in, the loss comes out, gradients flow backwards, and an optimizer updates the weights. The trouble starts when the model grows. A model with several billion parameters does not fit in the memory of one GPU once you count everything training needs, and even when it fits, one GPU is slow. So you want many GPUs working together, and you want the work split so that adding GPUs lets you train models that were impossible before, not merely the same model a bit faster.

DeepSpeed is a PyTorch library, plus a command-line launcher, that handles this. You hand it your ordinary torch.nn.Module and a JSON configuration file. It hands back an object called the engine that wraps your model. The engine takes over the distributed setup, mixed-precision arithmetic, splitting of memory across GPUs, offloading to CPU memory or disk, gradient accumulation, gradient clipping, checkpointing and the optimizer step. Your training loop shrinks to three lines: run the model forward, call backward, call step.

YOUR MODELa normal nn.Module
→
ds_config.jsonwhat to do
→
ENGINEdeepspeed.initialize
→
MANY GPUsstarted by the launcher

The diagram is the whole idea. Everything else in this guide is detail about one of those four boxes.

It helps to know what came before. The standard PyTorch answer to multi-GPU training is DistributedDataParallel, usually shortened to DDP. DDP starts one process per GPU, and every process holds a complete copy of the model, a complete copy of the gradients and a complete copy of the optimizer state. Each process trains on different data, and the gradients are averaged across processes after every backward pass. DDP is simple and fast, but it is wasteful in one specific way: if you have eight GPUs, you hold eight identical copies of everything. Adding GPUs gives you speed, not room. The biggest model you can train is still the biggest model that fits on a single GPU.

DeepSpeed's headline contribution, called ZeRO (Zero Redundancy Optimizer), removes that redundancy. Instead of every GPU holding every copy, ZeRO splits the training state across the GPUs, so with eight GPUs each one holds roughly an eighth of it. Suddenly adding GPUs does buy you room. The same project added offload to CPU memory and to NVMe disks, ways to split a model by layers or by weight matrices, long-sequence support, and tools for inference. The beginner path needs only a small slice: the engine, the config, the launcher, ZeRO stages 1 to 3, mixed precision and checkpoints.

Try it
  1. Pick a model you would like to fine-tune and find its parameter count on its model card.
  2. Look up the memory of the GPU you have access to.
  3. Multiply the parameter count by 16 to estimate the bytes training needs, and compare it with the GPU memory.
for anything beyond a few billion parameters the estimate is far larger than one GPU. That gap is the problem DeepSpeed exists to close, and the next section explains where the 16 comes from.

Where the memory goes

Before touching any configuration, you need a feel for why training uses so much more memory than the model's file size suggests. Without that, ZeRO's stages look like arbitrary numbers. With it, they are obvious.

When you train with the Adam optimizer in mixed precision, four kinds of data live in GPU memory for every parameter of the model.

The first is the parameters themselves, the weights, stored in a 16-bit format such as bf16 or fp16 for the forward and backward passes. That costs 2 bytes per parameter. The second is the gradients, one number per parameter, also in 16-bit, another 2 bytes. The third and fourth are the optimizer's own state. Adam keeps two running averages per parameter (the momentum and the variance), and mixed-precision training also keeps a full-precision 32-bit master copy of the weights, because tiny updates would vanish if they were added to 16-bit numbers. Those three 32-bit values cost 12 bytes per parameter. The DeepSpeed team's own figure for fp32 Adam is about 12 bytes per parameter on top of the weights, which matches. Add it up: 2 plus 2 plus 12 is roughly 16 bytes per parameter. That is the number from the previous Try it task. Treat it as an estimate, because activations and buffers are extra.

A model with 7 billion parameters therefore needs about 112 billion bytes, roughly 112 GB, for model state alone, before a single activation is stored. The weights file for that model in 16-bit is only 14 GB, which is why people are surprised. Most of the cost is not the model but the optimizer's bookkeeping.

Plain data parallel (DDP)

  • Every GPU holds all 16 bytes per parameter
  • More GPUs means more speed, not more room
  • The model must fit on one GPU, state and all

ZeRO

  • The 16 bytes are split across the GPUs
  • More GPUs means more room as well as speed
  • Stage 3 can train models larger than any single GPU

Now the three ZeRO stages make sense, because each one decides how much of the three big items to split across the GPUs. Stage 1 splits the optimizer state, the 12-byte part. Stage 2 splits the optimizer state and the gradients. Stage 3 splits all three, including the parameters themselves. The more you split, the less each GPU holds, and the more the GPUs have to talk to each other. That trade, memory against communication, is the central trade of the whole library, and later sections return to it.

Tip DeepSpeed ships estimators that print the memory a model needs under each ZeRO configuration without training anything. They are covered under the memory commands later, and they are the fastest way to answer "will this fit?" before you reserve expensive GPUs.
Try it
  1. For a 1.3 billion parameter model, compute the model-state memory at 16 bytes per parameter.
  2. Divide it by 4, as if ZeRO stage 3 were splitting everything across four GPUs.
  3. Compare the result with what a 24 GB GPU can hold.
about 21 GB without splitting and about 5 GB per GPU with it. The same model that barely fits on one card has comfortable headroom once the state is shared, which is the whole payoff.

The mental model: four nouns

You need only a handful of terms to read DeepSpeed code and documentation. Learn these and the rest follows.

The engine. The object returned by deepspeed.initialize is a DeepSpeedEngine. It wraps your model, and you call it just like the model: model_engine(batch) runs the forward pass. It also has the methods that replace the usual training boilerplate, notably backward, step, save_checkpoint and load_checkpoint. If you remember one fact, remember that after initialization you stop using your original model object for training and use the engine instead.

The config. A JSON file, or a Python dictionary, describing everything the engine should do: batch sizes, the optimizer and its learning rate, the learning-rate scheduler, the precision mode, the ZeRO stage, offload, clipping and monitoring. It is parsed when the engine is created, and mistakes are reported then, which is good: you find out in the first seconds, not after an hour.

Ranks. Distributed training starts several copies of your script, one process per GPU. Each process is called a rank and has an integer ID. The world size is the total number of ranks. The local rank is a process's index on its own machine, which is how it picks its GPU. The launcher sets the environment variables RANK, LOCAL_RANK, WORLD_SIZE, MASTER_ADDR and MASTER_PORT for you, so your script does not compute any of them.

ZeRO. The memory-splitting scheme described above, with stages 0 to 3. Stage 0 means no splitting, which behaves like ordinary data parallel training. You select the stage in the config with "zero_optimization": {"stage": 2}, and nothing else in your training code changes. That is the quiet superpower: moving from stage 1 to stage 3 is a one-number edit.

There is a fifth idea worth carrying even though it is not a noun: data parallelism. Every rank runs the same model on a different slice of the data, and the gradients are averaged. Even with ZeRO stage 3, DeepSpeed is doing data parallelism. It merely stores the state in pieces and gathers what it needs just in time.

One machine, four GPUs
rank 0
GPU 0
rank 1
GPU 1
rank 2
GPU 2
rank 3
GPU 3
Each rank runs your script on different data. With ZeRO stage 3, each also holds only one quarter of the parameters, gradients and optimizer state.

One piece of arithmetic appears in almost every beginner error, so learn it now. DeepSpeed relates three batch numbers:

train_batch_size = train_micro_batch_size_per_gpu x gradient_accumulation_steps x number of GPUs

The micro-batch is how many samples one GPU processes in a single forward and backward pass; it is what has to fit in memory. Gradient accumulation means running several micro-batches and summing their gradients before updating the weights, which lets you simulate a big batch on a small GPU. The train batch size is the effective batch per weight update across all GPUs. You set any two of the three and DeepSpeed infers the third. If you set all three and they disagree, you get the first error message in the troubleshooting section.

Try it
  1. Suppose you want an effective batch of 64, you have 4 GPUs, and each GPU can fit 4 samples.
  2. Work out the gradient accumulation steps.
  3. Then redo it for 2 GPUs and see which number changes.
4 accumulation steps on 4 GPUs and 8 on 2 GPUs. The effective batch stays 64 while the work per weight update is spread differently, and that is why batch keys must be rechecked whenever your GPU count changes.

Installing DeepSpeed and checking the setup

The golden rule: install PyTorch first. DeepSpeed's setup.py imports torch while building, so installing DeepSpeed into an environment with no PyTorch fails. DeepSpeed requires PyTorch 2.0 or newer, and the documentation recommends a current stable PyTorch 2.x. Pick the PyTorch build that matches your GPU driver, using the selector on pytorch.org.

On Linux, the primary platform, use a virtual environment:

BASH
python -m venv .venv && source .venv/bin/activate
pip install torch            # pick the CUDA or ROCm build matching your driver
pip install deepspeed        # source distribution; ops compile on first use
ds_report                    # or: python -m deepspeed.env_report

Python 3.10 to 3.12 is a safe choice. PyPI publishes DeepSpeed only as a source distribution, with no pre-built wheels; that is normal, and it is why the next idea matters.

Ops and JIT compilation. DeepSpeed includes fast C++ and CUDA extensions, called ops: a fused Adam optimizer, a CPU Adam for offload, an asynchronous file-I/O library for NVMe and others. By default these are compiled just in time the first time something needs them. The first run of a job that uses a fused optimizer therefore pauses for a minute or two while ninja and your compiler build the extension. The result is cached, by default in ~/.cache/torch_extensions, so later runs start fast. To compile these ops at install time instead, which is smart for containers and clusters where every node would otherwise compile the same thing, use build flags:

BASH
DS_BUILD_OPS=1 pip install deepspeed                       # all compatible ops
DS_BUILD_CPU_ADAM=1 DS_BUILD_AIO=1 pip install deepspeed  # selected ops only

Compiling CUDA ops needs a CUDA compiler (nvcc) whose version matches the CUDA version your PyTorch was built with, plus ninja and a C++ compiler. A pure-Python install without nvcc works for many things, but any op that needs compiling then fails when first used, with MissingCUDAException: CUDA_HOME does not exist, unable to compile CUDA op(s). For NVMe offload you also need the system library libaio, installed with apt install libaio-dev on Debian and Ubuntu.

On Apple Silicon, DeepSpeed 0.19.6 added support for the mps accelerator, covering ZeRO stages 0 to 3. The documented route is:

BASH
pip install torch
DS_ACCELERATOR=mps pip install --no-build-isolation deepspeed
ds_report

That is a useful way to learn the API on a laptop, with one important limit: a Mac has a single MPS device, so you cannot do multi-GPU data parallelism on it. Treat it as a place to test that your script and config are valid, not to train seriously.

Checking the install. The tool for this is ds_report. It prints two things. First, a table of every op with an installed column and a compatible column. [YES] under installed means the op was pre-built. [OKAY] under compatible means it can be compiled on demand. [NO] under compatible means your machine cannot build it, for example because libaio is missing. Second, general environment information: the torch version and location, the DeepSpeed version and location, the CUDA version torch was built with, the nvcc version and the shared-memory size. Read this output before debugging anything else, because many mysterious failures turn out to be version mismatches visible right there.

Trap If you upgrade PyTorch after pre-building DeepSpeed ops, you will meet PyTorch version mismatch! DeepSpeed ops were compiled and installed with a different version than what is being used at runtime. The ops are tied to the torch and CUDA they were compiled against. Reinstall DeepSpeed against the new torch, and rebuild any pre-built ops.
Try it
  1. Create a fresh virtual environment, install torch, then install deepspeed.
  2. Run ds_report and find the DeepSpeed version, the torch version and the nvcc line.
  3. Find the row for cpu_adam and note whether it says YES or OKAY under its columns.
the versions match what you installed. Most ops show OKAY rather than YES, meaning they will compile on first use, and that is expected for a plain pip install.

Your first project, step by step

We will train a deliberately tiny model on random data. The point is not the model; it is seeing every moving part of DeepSpeed work end to end with the least possible code. Create a folder with three files: train.py, ds_config.json and, later, nothing else.

Start with the config. It is the contract between you and the engine:

ds_config.json
{
  "train_micro_batch_size_per_gpu": 8,
  "gradient_accumulation_steps": 2,
  "optimizer": {
    "type": "AdamW",
    "params": { "lr": 0.0003, "weight_decay": 0.01 }
  },
  "scheduler": {
    "type": "WarmupLR",
    "params": { "warmup_min_lr": 0, "warmup_max_lr": 0.0003, "warmup_num_steps": 10 }
  },
  "bf16": { "enabled": true },
  "zero_optimization": { "stage": 2 },
  "gradient_clipping": 1.0,
  "steps_per_print": 10
}

Read it as a list of decisions. We say each GPU processes 8 samples per micro-step, and we accumulate 2 micro-steps per update. We do not give train_batch_size, because DeepSpeed computes it: 8 x 2 x the number of GPUs. We choose AdamW and a warmup schedule that ramps the learning rate from 0 over 10 steps. We turn on bf16, which needs a GPU of the Ampere generation or newer; on older cards use fp16 instead, covered below. We choose ZeRO stage 2, and we state clipping at 1.0 explicitly even though that is now the default.

Now the script:

train.py
import argparse

import deepspeed
import torch
import torch.nn as nn
from torch.utils.data import Dataset


class RandomData(Dataset):
    """1,024 random regression examples: 64 features in, 1 number out."""

    def __init__(self, n=1024):
        self.x = torch.randn(n, 64)
        self.y = self.x.sum(dim=1, keepdim=True)

    def __len__(self):
        return len(self.x)

    def __getitem__(self, i):
        return self.x[i], self.y[i]


class TinyNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.body = nn.Sequential(nn.Linear(64, 256), nn.ReLU(), nn.Linear(256, 1))
        self.loss_fn = nn.MSELoss()

    def forward(self, x, y):
        return self.loss_fn(self.body(x), y)


parser = argparse.ArgumentParser()
parser.add_argument("--local_rank", type=int, default=-1)  # the launcher passes it
parser = deepspeed.add_config_arguments(parser)             # adds --deepspeed, --deepspeed_config
args = parser.parse_args()

model = TinyNet()
model_engine, optimizer, train_loader, lr_scheduler = deepspeed.initialize(
    args=args,
    model=model,
    model_parameters=model.parameters(),
    training_data=RandomData(),
)

for epoch in range(3):
    for step, (x, y) in enumerate(train_loader):
        x = x.to(model_engine.device, dtype=torch.bfloat16)
        y = y.to(model_engine.device, dtype=torch.bfloat16)
        loss = model_engine(x, y)
        model_engine.backward(loss)
        model_engine.step()
    print(f"epoch {epoch} done, last loss {loss.item():.4f}")

Walk through it from the top, because every line has a reason.

deepspeed.add_config_arguments(parser) adds two command-line flags to your argument parser: --deepspeed and --deepspeed_config. That is how the path of your JSON file reaches the script. Alternatively you can skip the flags and pass config="ds_config.json" to initialize, but you must not do both; if you do, DeepSpeed stops with a message that it was given configs in two places. The --local_rank argument is there because the launcher passes it to your script, and an argument parser that does not declare it would reject the unknown flag.

deepspeed.initialize is the heart of the program. It takes your model, the list of parameters to optimize, and optionally a dataset. It returns four things: the engine, the optimizer, a data loader and a learning-rate scheduler. The last three can be None if you did not supply or configure them. Here, because we passed training_data and put the optimizer and scheduler in the JSON, we get all four. The data loader it builds is a DeepSpeedDataLoader, sized to the micro-batch in the config and sharded across ranks so each GPU sees different samples. Note that initialize also sets up the distributed process group for you; you do not call init_process_group yourself.

Inside the loop, we move each batch to model_engine.device, which is the right GPU for this rank. Because bf16 is enabled, the engine converts the model to bf16, so we cast the inputs to match. Then comes the three-line core: model_engine(x, y) runs forward and returns our loss, model_engine.backward(loss) runs backpropagation, and model_engine.step() does everything that follows. That step call is more than an optimizer step. It applies gradient clipping, runs the optimizer, advances the learning-rate scheduler and zeroes the gradients, but only at the gradient accumulation boundary. With gradient_accumulation_steps at 2, the first call to step accumulates and the second updates the weights. You never write if step % 2 == 0, and you never call optimizer.zero_grad(). If you do those by hand you will fight the engine.

Launch it with the DeepSpeed launcher:

BASH
deepspeed --num_gpus 2 train.py --deepspeed --deepspeed_config ds_config.json

The launcher starts two processes, one per GPU. Rank 0 prints a banner that includes DeepSpeed info: version=0.19.7, git-hash=..., git-branch=..., then configuration details, and then your epoch lines. With one GPU, use --num_gpus 1. On a machine with no --num_gpus flag, the launcher uses every GPU it finds. Because the script prints from every rank, you will see each message twice; real scripts guard prints with a rank check, which later sections show.

Note On the first run you may see a pause while ops compile. That is the JIT build described in the install section, and it will not happen on the second run.
Try it
  1. Create the two files above and run them with one GPU.
  2. Change gradient_accumulation_steps to 4 and run again.
  3. Find the line in the startup output that reports the effective train batch size.
with one GPU the train batch size goes from 16 to 32. The engine derived it from the keys you gave, which confirms you now control batch size through the config and not through code.

The configuration file, key by key

A DeepSpeed config looks long until you realise it is a set of independent blocks. You add a block only when you want that feature. This section covers the ones a beginner needs.

Batch size. Covered above: train_batch_size, train_micro_batch_size_per_gpu and gradient_accumulation_steps. Provide at least train_batch_size or train_micro_batch_size_per_gpu, because there are no numeric defaults. If you omit both you get Either train_batch_size or train_micro_batch_size_per_gpu needs to be provided.

Optimizer. The optimizer block has a type and a params dictionary. DeepSpeed natively understands Adam, AdamW, Lamb, Lion, Adagrad and Muon, matched case-insensitively. Any other name is looked up in torch.optim, so "SGD" works. For Adam and AdamW, DeepSpeed uses its fused implementation by default, or its CPU implementation when offloading, which is faster than the stock PyTorch loop. One important note on removed features: the 1-bit optimizers OneBitAdam, OneBitLamb and ZeroOneAdam were deleted in 0.19.7. Older tutorials mention them; do not use them.

If you pass an optimizer object to deepspeed.initialize in code, it overrides the optimizer block in the JSON. Use one or the other so you know which is in charge. And if you use a custom, non-DeepSpeed optimizer with ZeRO, DeepSpeed will refuse unless you set "zero_allow_untested_optimizer": true.

Scheduler. The scheduler block takes LRRangeTest, OneCycle, WarmupLR, WarmupDecayLR or WarmupCosineLR, each with parameters. They step once per engine.step() call, which means once per weight update, not once per micro-batch. Remember that when you reason about warmup_num_steps.

Precision. This is where beginners spend the most time, so slow down. You choose how numbers are stored and computed:

  • "bf16": {"enabled": true} uses bfloat16, a 16-bit format with the same range as 32-bit floats but fewer digits of precision. It is stable and needs no loss scaling. It requires hardware that supports it, meaning NVIDIA Ampere and newer, or recent AMD and Intel accelerators. This is the default recommendation for new work. Since 0.19.3, bf16 also checks gradients for infinities and NaNs and skips such steps, matching fp16.
  • "fp16": {"enabled": true} uses half precision. It has a narrower range, so it needs loss scaling: the loss is multiplied by a large number before backward so small gradients do not underflow to zero, and the scale adjusts dynamically. Setting "loss_scale": 0 means dynamic scaling. You will see log lines such as OVERFLOW! Rank 0 Skipping step when the scale is too big; DeepSpeed skips that step and lowers the scale. A few of these at the start are normal. A constant stream means instability.
  • "torch_autocast" drives PyTorch's native autocast from the config.
  • The amp block (Apex) is deprecated. Do not start anything new with it.

You cannot enable fp16 and bf16 together; DeepSpeed stops with bfloat16 and fp16 modes cannot be simultaneously enabled. If neither is enabled, training runs in fp32, which is the safest and slowest choice and a decent way to debug.

Gradient clipping. "gradient_clipping": 1.0 rescales gradients so their overall norm never exceeds 1.0, which protects against sudden spikes. Pay attention to the history here, because old tutorials contradict current behaviour: until 0.19.3 the default was 0.0 (disabled). Since 0.19.3 the default is 1.0. A config that omits the key now clips. If you really want no clipping, write "gradient_clipping": 0.0.

Trap Hugging Face users often leave their precision settings in two places, the training arguments and the DeepSpeed JSON, and they disagree. The cure is the "auto" value covered in the Trainer section: let the JSON borrow the value from the training arguments.
Try it
  1. In your first project, change bf16 to fp16 and cast the inputs to torch.float16.
  2. Run it and watch for any overflow lines.
  3. Now enable both blocks at once and read the error.
the fp16 run works, possibly with a couple of skipped steps while the loss scale settles, and enabling both gives the assertion about modes that cannot be enabled simultaneously.

ZeRO stages in practice

The previous sections introduced ZeRO as a memory-splitting idea. This section turns it into a decision you can make.

STAGE 0nothing split
→
STAGE 1optimizer state
→
STAGE 2+ gradients
→
STAGE 3+ parameters

Stage 0 disables ZeRO. The engine behaves like plain data parallelism. Use it when the model and its full training state fit comfortably, or to compare behaviour.

Stage 1 partitions the optimizer state, the 12-byte part, across the ranks. It needs almost no extra communication, so it is nearly free. With eight GPUs, the optimizer cost per GPU drops to about an eighth. Because the optimizer state is the largest of the three pieces, stage 1 alone recovers a large share of the memory.

Stage 2 also partitions gradients. Each rank keeps only the gradients for the slice of parameters whose optimizer state it owns, and the averaging step becomes a reduce-scatter, which sends each slice to its owner. Communication volume stays similar to plain data parallelism, which is why stage 2 is the most common choice for fine-tuning when the 16-bit weights themselves fit on each GPU. If you are unsure where to start, start here.

Stage 3 also partitions the parameters. No GPU holds the full model at rest. During the forward pass, right before a layer runs, its weights are gathered from all ranks; right after, the gathered copy is dropped. The same happens in backward. This is what lets you train models larger than a single GPU's memory. The price is communication: the volume is about 1.5 times that of plain data parallelism, so a fast interconnect (NVLink inside a machine, InfiniBand or RoCE between machines) matters. On slow networks stage 3 can crawl.

A beginner-friendly rule of thumb. Use stage 2 first. Move to stage 3 only when the model weights alone, or stage 2 with offload, still do not fit. Move down to stage 1 or 0 if you have memory to spare and want maximum speed.

A realistic stage 3 block looks like this:

ds_config.json
{
  "train_micro_batch_size_per_gpu": 4,
  "gradient_accumulation_steps": 8,
  "bf16": { "enabled": true },
  "zero_optimization": {
    "stage": 3,
    "overlap_comm": true,
    "contiguous_gradients": true,
    "stage3_gather_16bit_weights_on_model_save": true
  }
}

The keys beyond stage are tuning knobs. overlap_comm runs communication alongside computation to hide latency, and in the current code it resolves to true on its own at stage 3, though writing it is harmless. contiguous_gradients copies gradients into a contiguous buffer to reduce fragmentation and is true by default. stage3_gather_16bit_weights_on_model_save matters when you want to export a normal 16-bit weights file from a stage 3 run, as the checkpoint section explains. The other bucket-size and prefetch settings have sane defaults. Leave them alone until Mid-level.

What changes in your code at stage 3? For most scripts nothing. But two things surprise people. First, parameters are partitioned, so if you inspect param.shape on a stage 3 model during or after initialization you may see torch.Size([0]), because that rank holds only a shard. Second, if the model is so large that even building it on one GPU runs out of memory, you must construct it inside a deepspeed.zero.Init context, which partitions the parameters as they are created:

PYTHON
with deepspeed.zero.Init(config_dict_or_path="ds_config.json"):
    model = BigModel()

Hugging Face's from_pretrained does this for you when Trainer is configured for stage 3, as the integration section explains.

Tip You can switch stages without touching Python. Keep stage as the only difference between two config files and compare memory and speed with the same script. That is the cleanest experiment a beginner can run.
Try it
  1. Run your first project with stage 0, then 1, then 2, then 3, changing only the number.
  2. Each time, watch peak memory with nvidia-smi in another terminal.
  3. Write down the four numbers.
on a tiny model the differences are small and stage 3 may even look worse, because overhead dominates. The pattern only becomes dramatic as models grow, which is why you test real sizes before trusting a stage.

Offloading to CPU and NVMe

Sometimes even stage 3 across all your GPUs is not enough, or you only have one GPU and a large model. DeepSpeed can push parts of the training state out of GPU memory into the much larger CPU RAM, or even onto NVMe disk, and bring it back when needed. This is called offload.

The cheapest and most popular form is ZeRO-Offload: keep the optimizer state in CPU memory and run the optimizer update on the CPU. The 12-byte-per-parameter piece that dominated GPU memory moves off the GPU entirely. It works with stages 1, 2 and 3:

ds_config.json
{
  "train_micro_batch_size_per_gpu": 2,
  "gradient_accumulation_steps": 16,
  "bf16": { "enabled": true },
  "zero_optimization": {
    "stage": 2,
    "offload_optimizer": { "device": "cpu" }
  }
}

With this setting DeepSpeed uses its DeepSpeedCPUAdam implementation, a vectorized CPU version of Adam. The reward is a model that trains on a GPU that could not otherwise hold it. The cost is speed: every step moves gradients over the PCIe bus to the CPU and the updated weights back, and the CPU does the optimizer maths. Expect a slowdown, and expect it to depend heavily on how many CPU cores and how much memory bandwidth you have.

For stage 3 you can additionally offload the parameters with "offload_param": {"device": "cpu"}. And you can offload to NVMe by setting "device": "nvme" together with an nvme_path; that combination is ZeRO-Infinity, which is stage 3 only and requires libaio and an NVMe drive fast enough to be worth it. Offloading parameters is only valid at stage 3. A beginner should try optimizer-to-CPU first, then parameters to CPU, and only consider NVMe if the model really demands it.

The pinned memory change. Pinned (page-locked) CPU memory makes transfers between CPU and GPU much faster. Since 0.19.5, pin_memory inside the offload blocks defaults to true; before, it was false. That helps speed but pinned memory counts against the operating system's locked-memory limit. If you upgrade and suddenly see out-of-memory errors in offload runs, particularly on hosts or containers with a low ulimit -l, either raise the limit (in Docker, --ulimit memlock=-1) or set "pin_memory": false explicitly in both offload blocks.

Trap Offloading with a non-DeepSpeed optimizer is rejected on purpose. If you create an optimizer in code and enable ZeRO-Offload, you get a ZeRORuntimeException telling you to use DeepSpeedCPUAdam or define the optimizer in the config, because a stock PyTorch optimizer would crawl on the CPU. The fix is to put the optimizer in the JSON.
Try it
  1. Take your first-project config and add the offload_optimizer block with device cpu.
  2. Run it and compare GPU memory in nvidia-smi with and without offload.
  3. Compare the time taken for the same number of steps.
lower GPU memory use and a longer step time. That is the trade in miniature, and you now know the lever to pull when a job will not fit.

The launcher: running on one machine and beyond

The deepspeed command is more than a convenience wrapper; it is how ranks, environment variables and, across machines, SSH connections get set up.

On one machine the common forms are:

BASH
deepspeed train.py --deepspeed --deepspeed_config ds_config.json         # all local GPUs
deepspeed --num_gpus 2 train.py --deepspeed --deepspeed_config ds_config.json
deepspeed --include localhost:0,1 train.py --deepspeed --deepspeed_config ds_config.json
CUDA_VISIBLE_DEVICES=0,1 deepspeed train.py --deepspeed --deepspeed_config ds_config.json

The first uses every GPU on the machine. The second limits it to a number. The third chooses specific GPU indices; here GPUs 0 and 1. The fourth uses the familiar CUDA environment variable and works too. Everything after the script name goes to your script untouched, which is how --deepspeed and --deepspeed_config reach add_config_arguments.

The launcher uses port 29500 by default for the rendezvous where ranks find each other. If you run two jobs on the same machine at once, give the second a different one with --master_port, or the second fails with an address-already-in-use error. You may also see a warning that says Unable to find hostfile, will proceed with training with local resources only. On a single machine this is expected and harmless; it only means you did not provide a multi-machine host list.

You do not have to use the deepspeed command at all. Because DeepSpeed reads the standard PyTorch distributed environment variables, torchrun --nproc_per_node 2 train.py --deepspeed_config ds_config.json works, and so does accelerate launch. Hugging Face Trainer works under all of these.

Multiple machines. The default method is pdsh, which logs in to every machine over SSH without a password and starts the job. You describe the machines in a hostfile, an MPI-style text file:

hostfile
worker-1 slots=8
worker-2 slots=8

Each line names a machine and how many GPUs (slots) it offers. Then:

BASH
deepspeed --hostfile=hostfile train.py --deepspeed --deepspeed_config ds_config.json

Multi-node runs need passwordless SSH between machines and pdsh installed, otherwise you meet Cannot find pdsh, please install via 'apt-get install -y pdsh'. Alternatives exist: --launcher openmpi, slurm and others, and on Kubernetes or any system without SSH you start the launcher on each node yourself with --no_ssh --node_rank=<n> --master_addr=<ip>. All of those belong to Mid-level and Senior. As a beginner, get everything right on one machine first. The vast majority of fine-tuning jobs people do never need more than a single node of 8 GPUs, and multi-node adds failure modes you do not want while learning.

Try it
  1. Launch your project with --include localhost:0, then with --num_gpus 2 if you have two GPUs.
  2. Launch it twice at the same time and read the second job's error.
  3. Fix it with a different --master_port.
the second job fails because port 29500 is taken, and a new port number lets both run. You have now met the most common launcher collision.

Saving, loading and exporting checkpoints

Training jobs die: hardware fails, a scheduler preempts you, you press the wrong key. Checkpoints let you resume. With ZeRO, they work differently from a plain torch.save(model.state_dict()), and the difference trips up nearly everyone once.

The reason is sharding. Under ZeRO each rank holds only a slice of the optimizer state, and at stage 3 only a slice of the parameters. There is no single place that has the whole thing. So DeepSpeed saves in pieces: every rank writes its own shard.

PYTHON
model_engine.save_checkpoint("checkpoints", tag=None, client_state={"epoch": epoch})

The critical rule: every rank must call save_checkpoint, not just rank 0. If you wrap it in if rank == 0, the other ranks never write their shards, and the program may hang waiting on a collective operation that rank 0 started alone. With tag=None, the tag becomes global_step<N>, a folder name, and a small text file called latest records which tag is newest. The client_state dictionary is for anything of your own that you want stored, such as the epoch number. Do not put secrets in it.

The result looks like this on disk: a checkpoints/ directory holding a folder per tag, with a model-states file named something like mp_rank_00_model_states.pt, optimizer-states files per rank for ZeRO, the latest file, and a copy of a script called zero_to_fp32.py.

To resume:

PYTHON
load_path, client_state = model_engine.load_checkpoint("checkpoints")
if load_path is not None:
    start_epoch = client_state["epoch"] + 1

load_checkpoint returns the path it loaded from, or None if nothing was found, plus your client_state. By default it restores the module weights, the optimizer state and the scheduler state. Call it after deepspeed.initialize, on every rank. One documented caveat for stage 3: you cannot call load_checkpoint right after save_checkpoint within the same engine; re-create the engine first. Normal resume, which starts a fresh process, is unaffected.

Getting a normal weights file out. Sooner or later you want a plain file you can load with from_pretrained or torch.load outside DeepSpeed, and the sharded folder is not that. There are two routes.

The first is for ZeRO-3 only, and has to be planned: set stage3_gather_16bit_weights_on_model_save: true in the config, and then call model_engine.save_16bit_model("export_dir", save_filename="pytorch_model.bin") on all ranks. The engine gathers the full 16-bit weights and rank 0 writes the file. This is why that flag appeared in the stage 3 example.

The second route needs no GPU and works after the fact. The checkpoint folder contains the script zero_to_fp32.py, which merges the shards into a full-precision file:

BASH
python checkpoints/zero_to_fp32.py checkpoints exported_fp32 --safe_serialization

In 0.19.x the second argument is an output directory. Older tutorials, including some in the Hugging Face documentation, show the form python zero_to_fp32.py . pytorch_model.bin, which does not match the current behaviour. The optional --safe_serialization flag writes safetensors instead of a pickle file, and -t global_step1000 selects a specific tag. Expect it to use roughly twice the checkpoint size in CPU RAM.

A safety point for later: DeepSpeed checkpoints use PyTorch's pickle-based format, so loading a checkpoint from someone you do not trust can run arbitrary code. Load only checkpoints you or your team produced, and prefer safetensors when sharing.

Trap The most common checkpoint bug is saving on rank 0 only. Symptoms are a hang at the save step or a checkpoint that cannot be loaded because shards are missing. Always call save_checkpoint on every rank, and guard only your own file writes and prints with a rank check.
Try it
  1. Add save_checkpoint("checkpoints") at the end of every epoch in your first project.
  2. List the checkpoints folder and read the latest file.
  3. Run the project again with load_checkpoint("checkpoints") before the loop and confirm it resumes.
a folder named global_step followed by a number, a latest file containing that name, and a second run that starts from the saved state instead of from scratch.

Using DeepSpeed through Hugging Face Trainer

Many people never write deepspeed.initialize themselves. They fine-tune with the Hugging Face Trainer, which has built-in DeepSpeed support, and hand it the same JSON config. This is the most common way DeepSpeed is used in practice, so learn it.

Install the integration extras with pip install transformers[deepspeed], then pass the config file to the training arguments:

train_hf.py
from transformers import AutoModelForCausalLM, Trainer, TrainingArguments

args = TrainingArguments(
    output_dir="out",
    per_device_train_batch_size=4,
    gradient_accumulation_steps=8,
    learning_rate=2e-5,
    bf16=True,
    deepspeed="ds_config.json",
)
model = AutoModelForCausalLM.from_pretrained("your-model-id")
trainer = Trainer(model=model, args=args, train_dataset=my_dataset)
trainer.train()

Two details. First, create TrainingArguments before loading the model. For ZeRO-3, the arguments object is what lets from_pretrained load the model through the partitioned zero.Init path, so a model too large for one GPU never fully materializes on it. If you reverse the order at stage 3, loading can run out of memory. Second, launch this script with deepspeed train_hf.py, with torchrun, or with accelerate launch; a bare python starts only one process.

The slick part is the "auto" value. In the JSON, instead of repeating numbers that already live in TrainingArguments, write "auto" and Trainer fills them in:

ds_config.json
{
  "train_micro_batch_size_per_gpu": "auto",
  "gradient_accumulation_steps": "auto",
  "train_batch_size": "auto",
  "gradient_clipping": "auto",
  "bf16": { "enabled": "auto" },
  "optimizer": {
    "type": "AdamW",
    "params": { "lr": "auto", "weight_decay": "auto" }
  },
  "scheduler": {
    "type": "WarmupLR",
    "params": { "warmup_min_lr": "auto", "warmup_max_lr": "auto", "warmup_num_steps": "auto" }
  },
  "zero_optimization": { "stage": 2 }
}

Now per_device_train_batch_size flows into the micro-batch, learning_rate into the optimizer, and bf16=True into the bf16 block. One source of truth means no mismatch between the two systems. This also removes the batch-identity assertion for most Trainer users, because Trainer computes a consistent set.

If you use Hugging Face Accelerate instead, you run accelerate config, choose DeepSpeed, and either answer questions about ZeRO stage and offload or point it at your JSON file with deepspeed_config_file. Do not do both: if the file is given, Accelerate warns that its own variables will be ignored, and an inconsistent mix produces a ValueError listing them. The Hugging Face guide and the PEFT and TRL guide show these training libraries in full.

Try it
  1. Fine-tune a small model for 20 steps with Trainer and no DeepSpeed.
  2. Add deepspeed="ds_config.json" with the all-auto config above and launch with deepspeed.
  3. Compare memory use per GPU.
similar loss behaviour but lower memory per GPU on two or more GPUs, achieved without changing any model code, which is exactly the promise of the config-driven design.

Common errors and how to read them

DeepSpeed's errors are mostly assertions raised while the config is parsed or the engine is built. They look alarming and are usually one-line fixes. Here are the ones you will meet, with the cause and the repair.

Check batch related parameters. train_batch_size is not equal to micro_batch_per_gpu * gradient_acc_step * world_size 32 != 4 * 1 * 4. The three batch keys disagree with the real GPU count. In the example, you asked for a batch of 32 but 4 x 1 x 4 GPUs gives 16. This typically appears after you change the number of GPUs. Give only two of the three keys, or use "auto" in Hugging Face.

DeepSpeed requires --deepspeed_config to specify configuration file. No config reached initialize. Either you forgot deepspeed.add_config_arguments, or you did not pass --deepspeed_config path on the command line, or you did not give config=. The sibling message, Not sure how to proceed, we were given deepspeed configs in the deepspeed arguments and deepspeed.initialize() function call, means you passed it both ways.

bfloat16 and fp16 modes cannot be simultaneously enabled. Enable only one precision block.

DeepSpeedConfigError: DeepSpeed Sparse Attention has been removed..., or the same for Nebula, the compression library, MoQ and MiCS. You are using an old config on 0.19.7 or later. Delete the named block. For MiCS the replacement is ZeRO-3 with ZeRO++ hierarchical partitioning (zero_hpz_partition_size). The message ends with the link to the tracking issue that explains why.

You are using an untested ZeRO Optimizer. Please add "zero_allow_untested_optimizer": true. You passed an optimizer DeepSpeed has not validated with ZeRO. Either set that flag or use a native optimizer through the config.

MissingCUDAException: CUDA_HOME does not exist, unable to compile CUDA op(s). An op needs compiling and there is no CUDA toolkit. Common inside runtime-only containers. Use a -devel CUDA image, set CUDA_HOME, or pre-build ops at install time with DS_BUILD_OPS=1.

CUDAMismatchException: Installed CUDA version 12.4 does not match the version torch was compiled with 12.8. The system nvcc and PyTorch's CUDA disagree. Install the matching toolkit, point CUDA_HOME at the right one, or match the torch build to the toolkit. DS_SKIP_CUDA_CHECK=1 forces acceptance, but treat that as a last resort.

Watchdog timeout: Watchdog caught collective operation timeout ... Timeout(ms)=600000. One rank stopped participating in a collective operation: a slow data loader, a crashed rank, an uneven forward pass or a very slow checkpoint write. Since 0.19.6 the default timeout is 10 minutes, down from 30. Find the stuck rank with NCCL_DEBUG=INFO before raising the limit with DEEPSPEED_TIMEOUT=30, which fits only genuinely long operations.

OVERFLOW! Rank 0 Skipping step. Attempted loss scale: 65536, reducing to 32768. Not an error. It is fp16 loss scaling doing its job. A few at the start are normal. Constant skipping with a loss that never moves means instability: switch to bf16 or lower the learning rate. NaN loss from the very first step means a data or learning-rate problem, not DeepSpeed.

Out-of-memory after upgrading to 0.19.5 or later with offload. pin_memory now defaults to true and counts against the locked-memory limit. Raise ulimit -l or set "pin_memory": false.

Tip When a multi-GPU run fails confusingly, reproduce it with --num_gpus 1 first. Config and build errors show up on one GPU just as well, and a single process gives one clean traceback instead of eight interleaved ones.
Try it
  1. Break your config on purpose three ways: set all three batch keys inconsistently, enable bf16 and fp16 together, and delete --deepspeed_config from the command.
  2. For each, read the error and fix it.
  3. Keep a note of which source file each message mentions.
three distinct assertion messages that each point straight at the mistake, which shows how quickly this class of error resolves once you can read it.

Putting it all together

Here is one small end-to-end project that uses most of what you learned. We fine-tune a small causal language model with Trainer, ZeRO stage 2, bf16, optional CPU offload, and a clean export at the end. It runs on a single GPU or several.

The config, with everything driven by Trainer through auto:

ds_config.json
{
  "train_micro_batch_size_per_gpu": "auto",
  "gradient_accumulation_steps": "auto",
  "train_batch_size": "auto",
  "gradient_clipping": "auto",
  "bf16": { "enabled": "auto" },
  "optimizer": { "type": "AdamW", "params": { "lr": "auto", "weight_decay": "auto" } },
  "scheduler": {
    "type": "WarmupLR",
    "params": { "warmup_min_lr": "auto", "warmup_max_lr": "auto", "warmup_num_steps": "auto" }
  },
  "zero_optimization": { "stage": 2, "overlap_comm": true }
}

The script:

train_hf.py
from datasets import load_dataset
from transformers import (
    AutoModelForCausalLM, AutoTokenizer, DataCollatorForLanguageModeling,
    Trainer, TrainingArguments,
)

MODEL = "your-small-model-id"          # a small causal LM from the Hugging Face Hub

tok = AutoTokenizer.from_pretrained(MODEL)
if tok.pad_token is None:
    tok.pad_token = tok.eos_token

raw = load_dataset("your-text-dataset", split="train[:2000]")
data = raw.map(lambda b: tok(b["text"], truncation=True, max_length=256),
               batched=True, remove_columns=raw.column_names)

args = TrainingArguments(                  # created BEFORE the model
    output_dir="out",
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    learning_rate=2e-5,
    num_train_epochs=1,
    bf16=True,
    logging_steps=10,
    save_strategy="epoch",
    deepspeed="ds_config.json",
)

model = AutoModelForCausalLM.from_pretrained(MODEL)
trainer = Trainer(
    model=model, args=args, train_dataset=data,
    data_collator=DataCollatorForLanguageModeling(tok, mlm=False),
)
trainer.train()
trainer.save_model("final")

Launch it:

BASH
deepspeed --num_gpus 2 train_hf.py

The launcher starts two ranks, and Trainer calls deepspeed.initialize for you, filling the auto values from the training arguments: 4 per GPU, 4 accumulation steps, so 32 samples per update on two GPUs. ZeRO stage 2 splits the optimizer state and gradients between the GPUs. If memory is still tight, add "offload_optimizer": {"device": "cpu"} inside zero_optimization. For stage 3, change the stage to 3 and add "stage3_gather_16bit_weights_on_model_save": true; nothing else changes.

Try it
  1. Choose a small causal language model and a small text dataset, fill in the two placeholder ids, and run the project on your hardware.
  2. Run it again with stage 3 and note the change in memory and step time.
  3. Load the saved final folder with from_pretrained and generate a few tokens.
a model that loads and generates text. Stage 3 should use less GPU memory per card and run slower per step on a small model, which confirms both halves of the trade.

What you can now do, and what comes next

You can explain what DeepSpeed solves and why training needs about 16 bytes per parameter. You can name the engine, the config, ranks and ZeRO. You can install DeepSpeed in the right order and verify it with ds_report, write a training loop around deepspeed.initialize, backward and step, and keep the three batch keys consistent. You can choose bf16 or fp16, pick a ZeRO stage, turn on CPU offload, launch on one machine, save and resume with every rank participating, export a plain weights file, drive all of it from Hugging Face Trainer, and read the common errors.

Can you...
Say what DeepSpeed replaces in a training loop? Distributed setup, precision, memory sharding, clipping, the optimizer step
Explain the three ZeRO stages? Optimizer state, then plus gradients, then plus parameters
Compute the batch identity? micro-batch x accumulation x GPUs
Say why step is not just the optimizer? It also clips, schedules and zeroes, at the accumulation boundary
Choose between bf16 and fp16? bf16 where hardware supports it; fp16 needs loss scaling
Say who calls save_checkpoint? Every rank
Export a normal weights file? save_16bit_model or zero_to_fp32.py
Say what changed in 0.19.7? Nebula, compression, 1-bit optimizers, MiCS and Sparse Attention were removed

Mid-level takes each topic further. It explains how ZeRO works under the hood precisely enough to predict its behaviour, tunes the bucket and prefetch settings, introduces ZeRO++ and NVMe offload in depth, covers multi-node launching with hostfiles, SLURM and Kubernetes, and builds production patterns around checkpoints, monitoring and CI. Senior covers parallelism composition such as tensor and pipeline parallelism, security and the trust model of the launcher, cost, upgrades under a 0.x release cadence, and where DeepSpeed stops and PyTorch FSDP or a serving engine such as vLLM takes over. For the cluster side, the Kubernetes guide is the natural next stop, and MLflow is a good home for the runs you create.

Sources