This is part one of three. It covers everything you need to do real work with Weights & Biases, not a teaser. By the end you can log a training run from a script, read the charts it produces, compare it against other runs, version a dataset and a model as artifacts, launch a hyperparameter search, and recover when something breaks. Mid-level and Senior take the same topics further, and nothing you learn here is thrown away.
This guide is written against the Python SDK 0.30.0, released in September 2026. The SDK is still a 0.x library, which means the minor version can remove things, so if you copy code from an older blog post and it fails with an AttributeError, the reason is often a removal and not your mistake. We point out the important removals as we meet them.
Each section ends with a Try it task. Do them as you go. Experiment tracking only makes sense once you have watched your own numbers appear on a chart.
What Weights & Biases is, and what you can do by the end
Weights & Biases, usually shortened to W&B and spoken as "wand-B", is a service and a Python library for recording what happens while you train machine learning models. Your training script calls a few functions. Those functions send the numbers (loss, accuracy, learning rate), the settings you used, the machine's GPU usage and even the code version to a web dashboard. You then open the dashboard, see your runs drawn as charts, and compare them side by side.
That sounds modest, and it is also the thing that separates a person who is guessing from a person who is learning. Consider what happens without it. You train a model on Monday and it reaches 91 percent accuracy. On Wednesday you change the learning rate, train again and get 93 percent. On Friday your manager asks which settings produced the 93 and whether it is better than last week's model. If the answer lives in terminal scrollback and file names like model_final_v2_REAL.pt, you cannot give it.
By the end of this guide you will be able to:
- Install the library, create an account, and authenticate safely.
- Write a training script that records its settings and metrics in a run.
- Read the dashboard: line charts, the run table, the system metrics and the overview page.
- Store a dataset or a model as a versioned artifact and fetch it again from another script.
- Describe a hyperparameter search in a YAML file and let W&B drive it.
- Work offline, sync later, and read the most common error messages.
The diagram is the whole product in one line. Everything else in this guide is a detail of one of those four boxes.
A note on where the service runs. The free, hosted version is called the Multi-tenant Cloud. The documentation has recently begun branding it CoreWeave Forge, after the company that now operates W&B, and some pages still say "W&B Multi-tenant Cloud" and link to wandb.ai. They are the same service for your purposes. The library still talks to https://api.wandb.ai by default, so you do not need to change any URL to use it. Companies with strict data rules can also run a Dedicated Cloud in a region they choose, or host the server themselves. That matters for employers in the Gulf and Egypt who must keep data in a particular country, and Mid-level and Senior explain the options. For this guide, use the hosted service.
- Think of the last model you trained. Write down which settings you used and the final metric. Could you reproduce that result from your notes alone?
- Write one question about your training that you could not answer today, for example "which run used the smaller batch size?". Keep it, because you will answer it in the dashboard by the end.
The problem it solves, and what came before
Training a model is an experiment, and experiments need a lab notebook. Before tools like W&B, people used three approaches, each with a clear weakness.
The first was print statements and terminal scrollback. You watch the loss scroll by. When the run ends, the numbers are gone unless you redirected the output to a file, and even then you cannot draw a chart without writing more code.
The second was spreadsheets and naming conventions. You copy final numbers into a sheet by hand and encode settings into folder names such as run_lr0.01_bs64. This works for three runs and collapses at thirty, because people forget to fill in the row, misspell a setting, or never record the code version.
The third was TensorBoard, a viewer that reads log files written to your disk. It draws good charts, but the files live on whichever machine trained the model. Sharing them with a teammate means copying directories around, and comparing a run on your laptop with one on a cloud GPU means moving files first.
W&B's approach is to make the record automatic and central. The library sends data to a server while your code runs, so the run exists in one place the moment it starts, visible to anyone on your team. It also records things you did not ask for: the Python version, the operating system, the command that started the script, the GPU, and the git commit. When a teammate asks "how did you get that number?", the run page answers.
Two products in the same family are worth a name, so you are not confused later. Weave is a separate W&B product for tracing and evaluating applications built on large language models, and it is outside this guide. Registry, which we touch on in the artifacts section, is the organisation-wide shelf where finished models and datasets are published. Everything in this guide concerns the experiment-tracking product, which W&B calls Models.
Other tools solve the same problem. MLflow does experiment tracking and you can self-host it for free, and our MLflow guide is a good companion. Comet and ClearML are commercial and open alternatives. DVC focuses on versioning data and pipelines, and there is a DVC guide as well. W&B's strength is the dashboard, the speed of getting started and the depth of its team features. Its cost is that the hosted service is a third-party dependency, which you should weigh if your employer has data-residency rules.
- Pick one project where you currently track results by hand. List the three pieces of information you would most want recorded automatically.
- Compare your list against the next section. Which of your three items is a config, which is a metric, and which is a file?
The core mental model: entity, project, run, config, history, summary
W&B has a small vocabulary. Learn these words now, because every error message and every doc page uses them.
A run is one unit of computation that you log. In practice, it is one execution of your training script. You create a run by calling wandb.init(), and it ends when the script finishes or you call finish. Everything else hangs off a run. Each run has a unique ID (eight characters, such as 3x7k9a2q, generated for you) and a human-friendly name, which W&B invents as two random words unless you choose one.
A project is a folder of runs that belong together. You might have a project called churn-model holding every run you trained for that task. Project names can be up to 128 characters and cannot contain /, \, #, ?, % or :. If you forget to name one, W&B guesses from your git folder or script name, and if it cannot guess it uses uncategorized. Give projects names on purpose, because a pile of runs in uncategorized is hard to use.
An entity is the owner of a project. It is either your personal username or a team. A team is a group of people who share projects, and it is what you use when several colleagues need to see the same runs. The entity must exist already, and if you leave it out, your default entity is used. A detail that will matter later: artifacts logged under a personal entity cannot be linked to the organisation Registry, so work you intend to share or promote should live under a team.
Config is the set of inputs to a run, meaning the hyperparameters and settings that define the experiment: learning rate, batch size, number of epochs, model name, dataset version. You pass it as a dictionary, and it is stored in run.config. Config answers the question "what did I choose?".
History is the time series of everything you log during the run. Each call to run.log() adds a row of numbers, and by default each call moves the run forward by one step. History answers "what happened, and when?", and it is what draws the line charts.
Summary holds one final value per metric. By default it is the last value you logged for each name, so a run that ended with val_acc of 0.93 has a summary of 0.93. You can override it, for instance to keep the best value instead of the last. Summary answers "how did it end?", and it is what the run table sorts on.
An artifact is a versioned bundle of files that a run consumes or produces: a dataset, a model checkpoint, an evaluation set. Artifacts are the answer to "which exact data and which exact weights?", and we cover them in their own section.
A sweep is a hyperparameter search. You describe the space of settings to explore, and W&B hands different combinations to one or more worker processes called agents.
There is one more piece of machinery to know about, because it explains behaviour that would otherwise seem magical. When you call wandb.init(), the library starts a helper process called wandb-core, written in Go. Your Python code hands data to it, and it does the uploading, the streaming of metrics and the monitoring of CPU, memory and GPU. It also writes everything to a local file named run-<id>.wandb inside a folder called wandb/ in your working directory. This has a comforting consequence: if the network drops or the helper crashes, your training keeps going, and the local file lets you upload the data later with wandb sync.
- For a model you have trained, write one example each of config (an input you chose), history (a number that changed during training), summary (a final number) and artifact (a file you would want to keep).
- Decide whether your project would live under your personal entity or a team, and why.
Installing and checking your setup
W&B needs Python 3.10 or newer. Older releases of Python are no longer supported by the current library. On macOS you need version 13 (Ventura) or newer, and on Windows you need a 64-bit Python, because 32-bit wheels are no longer built. Linux works on both x86-64 and ARM machines.
Create a virtual environment so the library does not collide with other projects, then install it:
python -m venv .venv
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install --upgrade wandb
Inside a notebook, use !pip install wandb in a cell. Confirm the version from both the command line and Python:
wandb --version
python -c "import wandb; print(wandb.__version__)"
Both should print 0.30.0 or newer. If they disagree, you have two Python installations and the wandb command belongs to a different one than the python you are using. Activating the virtual environment normally fixes that.
Next, create an account at the W&B website and generate an API key. The key is the password your scripts use. Open your profile menu, choose User settings, then API keys, then New key. Copy it straight away: W&B shows the full key only once, and afterwards the page shows only a short prefix. Newer keys are around 86 characters long. Very old library versions (before 0.22.3) reject them with API key must be 40 characters long, yours was 86, and the cure is to upgrade the library.
Now log in:
wandb login
It prompts you to paste the key. The library then stores it in a file called .netrc in your home directory, so you only do this once per machine. You can also pass the key directly, as in wandb login <KEY>, or set an environment variable, which is the normal approach on servers and in automation:
export WANDB_API_KEY=your-key-here # Linux or macOS
$env:WANDB_API_KEY="your-key-here" # Windows PowerShell
train.py ends up in version control the first time you commit, and anyone who can read the repository can log runs as you. Keep the key in wandb login's stored credentials or in an environment variable. If you do leak one, delete it on the API keys page and create a new one. A key is revoked the moment you delete it.
Verify the whole setup with two small checks:
wandb login --verify
wandb status
The first confirms that the stored credentials work and shows which source they came from and your default team. The second prints your current settings. As of 0.30.0 it no longer prints the API key itself, which is sensible because people paste this output into support chats. You can also test from Python, which is handy in scripts and returns a boolean without prompting:
import wandb
print(wandb.login(prompt=False)) # True if credentials are configured
Finally, a smoke test. This tiny script should print a line ending in View run at ... and produce a run marked Finished in your dashboard:
import wandb
with wandb.init(project="smoke-test", config={"lr": 0.01}) as run:
run.log({"loss": 0.1})
- Create a virtual environment, install
wandb, and runwandb --version. - Generate an API key, run
wandb login, thenwandb login --verify. - Run the smoke-test script, open the printed link, and find the run named with two random words.
Your first real run, built step by step
The smoke test logged one number. Now build something that looks like real training. To keep the focus on W&B and not on a machine-learning framework, we will fit a straight line to noisy data with gradient descent in plain Python. The structure is identical to a neural network: a loop, a loss that shrinks, and numbers worth recording.
Start with the data and the settings. Save this as train.py:
import random
import wandb
config = {
"learning_rate": 0.05,
"epochs": 30,
"noise": 0.3,
"seed": 42,
}
random.seed(config["seed"])
xs = [i / 10 for i in range(50)]
ys = [2.0 * x + 1.0 + random.gauss(0, config["noise"]) for x in xs]
The true relationship is y = 2x + 1, with noise added. Our model has two unknowns, a slope w and an intercept b, which training should recover. Now add the training loop inside a run:
with wandb.init(project="line-fit", name="baseline", config=config,
tags=["baseline"], notes="First attempt, plain gradient descent") as run:
w, b = 0.0, 0.0
lr = run.config["learning_rate"]
for epoch in range(run.config["epochs"]):
grad_w = grad_b = loss = 0.0
for x, y in zip(xs, ys):
err = (w * x + b) - y
loss += err * err
grad_w += 2 * err * x
grad_b += 2 * err
n = len(xs)
w -= lr * grad_w / n
b -= lr * grad_b / n
run.log({"train/loss": loss / n, "weights/w": w, "weights/b": b})
run.summary["final_w"] = w
run.summary["final_b"] = b
Run it with python train.py. The terminal prints a link to the run. Open it, and you will see three line charts appear: the loss falling toward the noise level, and the two weights climbing toward 2 and 1. You wrote no chart code at all.
Walk through what each part did, because the details are where beginners go wrong.
The with block. Using wandb.init(...) as a context manager means the run is finished for you when the block ends. If your code raises an exception, the run is marked Failed, which is honest. Without the with block, a crash often leaves the run marked Crashed, which W&B decides only after it stops receiving heartbeats, and that is slower and less informative. Prefer the context manager in every script.
config=config. The dictionary becomes run.config. We then read learning_rate back from run.config and not from our original dictionary. That habit matters because in a sweep, W&B overwrites the config with the values it chose, and reading from run.config picks them up with no code change.
name, tags and notes. The name labels the run in the table. Tags are short words you can filter by, and notes are free text. All three are optional, and all three save you later.
run.log({...}). Each call records one step: a row of numbers. Put everything you want to see at the same step into one dictionary and one call. Calling run.log separately for each metric advances the step each time, so your charts end up with misaligned x-axes and far more steps than you intended.
Metric names with a slash. Names such as train/loss and weights/w make the dashboard group charts into collapsible sections called train and weights. That is the standard convention in the official examples. The documentation recommends names built from letters, digits and underscores, with / used for grouping. Avoid spaces, commas and hyphens in metric names, because they can cause trouble when you query the data later.
run.summary. We set two extra values by hand. Anything you assign there appears as a column in the run table.
Now run the script a second time with a different learning rate. Edit config to use "learning_rate": 0.01, change the name to "slow-lr", and run it again. Go to the project page. You now have two runs in the same project, and this is the moment the tool begins to earn its keep. The charts overlay both runs, and the run table has a column for learning_rate, so you can see immediately that the lower rate converges more slowly.
Two habits are worth forming from the start. First, log inside the with block and never afterwards, because logging to a finished run produces the warning Run (<id>) is finished. The call to log will be ignored. Second, prefer run.log(...) over the module-level wandb.log(...). The module-level version works on a global variable that can point at the wrong run if you ever create more than one, and the documentation steers learners toward the method on the run object.
- Type in
train.pyand run it. Open the run page and find thetrainandweightschart sections. - Run it again with
learning_rateset to0.01and a new name. Find the project page and compare the two loss curves. - Add a third run with
epochsset to 60. Which setting changed the final loss more?
Logging in depth: metrics, config, summary and media
The first project used plain numbers. Real projects log more kinds of data, and knowing the options saves you from reinventing them.
Scalars are single numbers like loss and accuracy. You have seen them. Log training metrics every step or every few steps, and validation metrics once per epoch. There is no need to log every step of a long run. The documentation suggests staying below roughly 1,000 log calls per minute, and logging every batch of a fast loop can exceed that.
Config values are set when the run starts. You may add to them later with run.config.update({"dataset_size": 5000}). One rule catches beginners: by default, changing a config key that is already set raises an error such as Attempted to change value of key "lr" from 0.01 to 0.02. The protection exists because config is meant to be a fixed record. If you really need to change it, pass allow_val_change=True. Better still, if a value changes during training, it was never a config value. Log it as a metric. Config keys should also avoid a dot in their names, and the whole config should stay under 10 MB, which is generous.
Summary defaults to the last value of each metric, and that is often wrong. For a validation loss, the last epoch is usually not the best one. You can tell W&B how to summarise a metric with define_metric:
run.define_metric("val/loss", summary="min")
run.define_metric("val/acc", summary="max")
Call these once, right after init. The accepted summaries include min, max, mean and last. Two older spellings are gone: summary="best" and the goal= argument were removed in 0.29.0, so if you find them in an old tutorial, use min or max instead.
define_metric has a second job, which is choosing the x-axis. By default the x-axis is the step. If you would rather plot against epochs, log an epoch number and declare it:
run.define_metric("epoch")
run.define_metric("val/*", step_metric="epoch")
The star matches every metric that starts with val/, so one line covers them all.
Images are logged with wandb.Image. This is essential for computer-vision work, because a wrong label is easier to spot in a picture than in a number. Pass a PIL image, a NumPy array or a path:
run.log({"examples": wandb.Image(img_array, caption="input 17")})
There is a change to know about. Since 0.29.0, wandb.Image no longer rescales pixel values for you, and its normalize argument was removed. An array must already hold values in the range 0 to 255. If you pass floats between 0 and 1, the image appears almost black, and the fix is to multiply by 255 and convert to an unsigned 8-bit integer first.
Tables are the tool for examining predictions. A wandb.Table holds rows and columns that you can sort and filter in the browser, and it can contain images too:
table = wandb.Table(columns=["id", "text", "label", "prediction"],
data=[[1, "great film", "pos", "pos"],
[2, "dull plot", "neg", "pos"]])
run.log({"val_predictions": table})
Open the run and you can sort by the rows where label differs from prediction, which is often the fastest way to find out what your model misunderstands. Other media types exist, including audio, video, HTML, histograms and 3D objects. Note that wandb.Video built from an array now requires a format= argument.
System metrics need no code at all. While a run is active, W&B samples the machine every 15 seconds and records CPU, memory, disk and network use, plus GPU utilisation, memory and temperature for NVIDIA cards. It also supports AMD GPUs, Apple Silicon and TPUs. These charts live in the run's System tab, and they answer a common question: is my GPU actually busy, or is my data loader the bottleneck? A GPU hovering at 20 percent usually points at the input pipeline and not the model.
Code and git. W&B records the git commit and any uncommitted difference for runs started inside a git repository, so the run page can tell you which version of your code produced it. If you also want a copy of the script itself, pass save_code=True to wandb.init. To switch git and code capture off, there are the settings disable_git and disable_code, which you would use if the repository contains something you may not upload.
log call under 25 MB, and fewer than 1,000 files per run. Another cause of a sluggish dashboard is a project with a huge number of distinct metric names, for example a name that embeds a step number such as loss_step_417. Keep the set of names small and let the step be the x-axis.
- Extend
train.pywith a held-out validation split and logval/losseach epoch. - Use
define_metricso the run table shows the minimumval/loss, not the last one. - Log a
wandb.Tableof five predictions with their true values and open it in the browser.
Reading the dashboard
Logging is half the skill. The other half is reading what you logged, and the dashboard rewards a few minutes of orientation.
The project page has a workspace, which is a grid of charts across all the project's runs, and a runs table along the side or in its own tab. The table is the place to sort, filter and group. Click a column header such as a summary value to sort the runs by it. Use the filter to show only runs with a given tag, and the group option to collapse runs that share a setting into one averaged line. Colours are assigned per run, and clicking the eye icon beside a run hides or reveals it on every chart.
A workspace can be automatic, where W&B makes a panel for every metric you log, or manual, where you choose the panels. Automatic is perfect while you are learning. On a project with hundreds of metric names, switch to manual and add only the panels you care about, because automatic workspaces become slow.
Charts have useful controls. Hover to see exact values for every run at a step. Drag across a region to zoom. Change the x-axis from step to another logged metric, or to wall-clock time. Smoothing is on by default to make noisy curves readable, and it can hide a real spike, so lower it when you are hunting for instability.
A run page shows one run in detail, and its tabs are worth knowing:
| Tab | What you find there |
|---|---|
| Overview | Name, ID, state, tags, notes, the command that started the run, git commit, Python and OS versions, hardware, the config and the summary |
| Charts | The charts for just this run |
| System | CPU, memory, GPU and network charts sampled every 15 seconds |
| Logs | The terminal output your script printed, captured as output.log |
| Files | Files the run saved, including requirements.txt, config.yaml and wandb-summary.json |
| Artifacts | Which artifacts this run used and produced |
Learn the run states, because they explain what you see in the table. Running means data is still arriving. Finished means the script ended normally. Failed means the script ended with a non-zero exit code or an exception. Crashed means W&B stopped receiving heartbeats from the process, typically because the machine ran out of memory, the job was killed, or the network was lost for a long time. Killed means someone stopped the run from outside. A sweep run can also be Pending while it waits in a queue.
When a run is Crashed and you know the process actually kept going, the data is not lost. It sits in the local wandb/ folder, and wandb sync wandb/run-<timestamp>-<id> uploads it.
Beyond the table and charts, a report is a document that mixes text with live panels. It is how teams write up an experiment, with a chart, a paragraph of interpretation and a link back to the runs. Creating one is a matter of clicking Create report from a project, and it stays connected to the data underneath.
- Open your
line-fitproject. Sort the runs table by final loss and hide the worst run. - Open the System tab of one run. Note the CPU and memory charts.
- Open the Files tab and read
config.yamlandrequirements.txt. Could a teammate rebuild the environment from them?
Artifacts: versioning data and models
Runs record numbers, and artifacts record files. An artifact is a versioned bundle of files, and it solves a problem every team meets: a model is the product of its code, its settings and its data, and you cannot reproduce it without all three. Artifacts give the data and the trained weights a permanent, versioned identity.
The best way to see the value is through an example. Imagine a training script that reads data/train.csv. Someone fixes a labelling bug in that file. A week later, nobody remembers which version a given model saw. If the CSV had been an artifact, each version would have its own name, train-data:v0, train-data:v1, and each run would list exactly which version it used.
Here is the producing side. One script creates a dataset and logs it:
import wandb
with wandb.init(project="line-fit", job_type="prepare-data") as run:
art = wandb.Artifact(name="line-data", type="dataset",
description="Noisy points around y = 2x + 1",
metadata={"noise": 0.3, "rows": 50})
art.add_file("data.csv")
run.log_artifact(art)
Three nouns appear here. The name identifies the artifact within the project. The type is a free-form label, conventionally dataset or model, used to group things in the UI. The metadata is a dictionary of facts about the contents. The job_type on the run is a label that describes what the run does, such as prepare-data or train, and helps the lineage graph read well.
W&B gives each new version a number starting at zero: v0, v1, v2. It also keeps movable labels called aliases. The alias latest is added automatically and always points at the newest version. You may add your own, such as best or production, when logging:
run.log_artifact(art, aliases=["best"])
If you log identical contents twice, W&B notices that nothing changed and does not create a new version. Newer library versions print a warning when this happens, which surprises people the first time.
Now the consuming side. A training script fetches the data by name and alias:
import wandb
with wandb.init(project="line-fit", job_type="train") as run:
data_art = run.use_artifact("line-data:latest")
folder = data_art.download() # downloads into ./artifacts/...
print("data is in", folder)
use_artifact records that this run used that version. download copies the files to a local folder, by default under ./artifacts. Because both the log and the use are recorded, W&B can draw a lineage graph: a picture showing dataset v1 feeding run baseline, which produced model v0. The graph is in the Artifacts tab and is one of the most persuasive features for explaining a result to a colleague.
Saving a trained model follows the same pattern. After training, write the weights to a file, then log it:
model_art = wandb.Artifact(name="line-model", type="model",
metadata={"final_loss": 0.09})
model_art.add_file("model.json")
run.log_artifact(model_art, aliases=["best"])
Two more ideas complete the picture. A reference artifact records a pointer to files that live elsewhere, such as an s3:// or gs:// location, using add_reference, without copying the bytes into W&B. Choose it when the data is large or must stay in your own cloud storage, which is a common requirement for regional data rules. And the Registry is an organisation-wide catalogue. When a model is good enough to share, you link a version from a project into a registry collection. Linking creates a pointer and not a copy. Remember that artifacts logged under a personal entity cannot be linked there, which is why teams log to a team.
You can also reach artifacts from the command line. wandb artifact ls <project> lists them, wandb artifact get <path> downloads one, and wandb artifact put <path> uploads a file or folder. These are useful in shell scripts and CI jobs where writing Python is overkill.
- Save the points from
train.pyintodata.csvand log it with the producing script above. - Log it again unchanged and note that no
v1appears. Change one row, log again, and confirm thatv1is created. - Run the consuming script, then open the Artifacts tab and find the lineage graph.
Hyperparameter sweeps
Once you can track runs, the next question is how to choose settings. Trying learning rates by editing the file and re-running gets slow. A sweep automates it: you describe the space of values to explore, and W&B hands out combinations.
A sweep has two parts. The controller lives on the W&B server and decides which configuration to try next. Agents are processes on your machines that ask the controller for a configuration, run your script with it, and report back. You can start several agents on different machines and they share the work.
Describe the search in a YAML file:
program: train.py
method: random
metric:
name: val/loss
goal: minimize
parameters:
learning_rate:
distribution: log_uniform_values
min: 0.001
max: 0.5
epochs:
values: [20, 40, 60]
Read it line by line. program is the script to run. method is the search strategy, and the choices are grid (every combination), random (random draws) and bayes (Bayesian optimisation, which uses the results so far to choose promising values). metric tells W&B what to optimise, and it must match a name your script logs exactly, so val/loss here. Under parameters, a list of values is a choice, and a distribution with min and max is a range. The distribution log_uniform_values draws numbers evenly on a logarithmic scale between the two values, which is what you want for learning rates, since 0.001 and 0.01 are as different as 0.1 and 1.
Your script needs to read from run.config, which we recommended earlier. Change the hard-coded lr line to read run.config["learning_rate"] and keep epochs likewise. Give config a default for each key so the script still works when run alone.
Create the sweep and start an agent from the terminal:
wandb sweep --project line-fit sweep.yaml
This prints a sweep ID and a ready-made command of the form wandb agent <entity>/line-fit/<sweep_id>. Run that command to start an agent, and cap it so it does not run forever:
wandb agent --count 10 your-entity/line-fit/SWEEP_ID
The --count 10 flag makes this agent run ten configurations and stop. Open the sweep page in the dashboard and you will see a parallel coordinates chart and a parameter importance panel. The first draws one line per run through each parameter, so you can spot which ranges produce low loss. The second estimates which parameters mattered most.
Two cautions. Grid search over a continuous range never ends, so use random or bayes with a --count or a run_cap line in the YAML. And a sweep can be controlled while it runs: wandb sweep --pause, --resume, --stop (let active runs finish) and --cancel (kill them) each take the sweep ID.
You can start a sweep from Python as well, which is convenient in notebooks:
sweep_id = wandb.sweep(sweep=sweep_config, project="line-fit")
wandb.agent(sweep_id, function=train, count=10)
Here sweep_config is the same structure as a dictionary, and train is a function that calls wandb.init and logs as usual.
- Make
train.pyreadlearning_rateandepochsfromrun.config, and logval/loss. - Save
sweep.yaml, create the sweep, and start an agent with--count 8. - Open the sweep page. Which learning-rate range gave the lowest loss, and does the parameter importance panel agree?
Offline mode, environment variables and configuration
Not every machine has a network, and not every script should call the cloud. W&B handles both through modes and environment variables.
The mode of a run is online by default. Set it to offline and the run writes only to your disk, then you upload later. Set it to disabled and every W&B call becomes a harmless no-op, which is ideal for unit tests and quick debugging where you do not want clutter on the dashboard.
export WANDB_MODE=offline
python train.py
wandb sync wandb/offline-run-20260930_101500-abc123xy
The offline run lives in a folder beginning offline-run-. Running wandb sync with that path, or with no argument to sync everything pending, uploads it. The command has a --dry-run flag for checking first, and --yes for scripts that must not ask for confirmation. Since 0.30.0, several older wandb sync options were removed. In particular --sync-all now errors, so plain wandb sync is the current form. Once runs are safely uploaded, the new command wandb clean deletes local data for runs that have already been synced.
The command line also has toggles. wandb offline and wandb online switch the mode for the current directory, and wandb disabled and wandb enabled turn W&B off and on.
Environment variables let you configure without touching code, which is how you keep a script identical between your laptop and a server. The ones a beginner uses are:
| Variable | What it does |
|---|---|
WANDB_API_KEY |
Your credentials, for machines where you cannot run wandb login interactively |
WANDB_ENTITY |
The team or user that owns the runs |
WANDB_PROJECT |
The project name, used when init is not given one |
WANDB_MODE |
online, offline or disabled |
WANDB_NAME |
A name for the run |
WANDB_TAGS |
Comma-separated tags |
WANDB_DIR |
Where the wandb/ run folder is created |
WANDB_SILENT |
Set to true to reduce console output |
A script can therefore call wandb.init() with no arguments, and the environment decides where the run goes. In CI or on a cluster node, set WANDB_API_KEY, WANDB_ENTITY and WANDB_PROJECT and you are done.
W&B keeps a few settings files as well. The system-wide one lives at ~/.config/wandb/settings, and a per-directory one lives at ./wandb/settings, which wandb init writes. One quirk from recent versions: if any key in a settings file is invalid, the library ignores all settings files and prints an error. Fix the typo, or reset with wandb init --reset.
Other places W&B writes are worth knowing so your disk does not surprise you. Run files go in ./wandb, downloaded artifacts in ./artifacts, and a cache in ~/.cache/wandb. The cache can grow. The command wandb artifact cache cleanup and the newer wandb purge-cache tidy it.
Add wandb/ to your .gitignore, because the run folders are local working data and should not be committed.
WANDB_MODE=disabled for the test run. The code path stays the same, but nothing is sent and no dashboard run is created, so your test suite does not litter the project or need network access.
- Set
WANDB_MODE=offline, runtrain.py, and confirm a run appears only in the localwandb/folder. - Upload it with
wandb sync --dry-runfirst, thenwandb sync. - Set
WANDB_MODE=disabledand run again. What changed in your output?
Working in notebooks and with your framework
Many beginners meet W&B in a Jupyter or Colab notebook. The rules are the same, with a few differences worth knowing.
Install with !pip install wandb in a cell, then call wandb.login(). In a notebook the login prompts you inline to paste the key. In init, the notebook default differs from scripts in one respect: changing a config value that is already set is allowed, which is why you may see no error in a notebook that would fail in a script. Use the with wandb.init(...) as run: form here too, so the run is closed when the cell finishes. If you forget, re-running a cell can start a second run while the first is still open, and the console warns you about it.
Most frameworks have a hook. You will meet these in real projects:
- PyTorch Lightning and Hugging Face Transformers can log to W&B with a single setting, such as choosing W&B as the logger or reporting destination. Both send the loss and learning-rate curves for you.
- Keras has callbacks. The old
WandbCallbackwas removed in 0.27.1. The current ones areWandbMetricsLogger,WandbModelCheckpointandWandbEvalCallback, imported fromwandb.integration.keras. - scikit-learn, XGBoost and LightGBM work simply: train as usual, then call
run.logwith the scores.
When a framework integration logs for you, you still call wandb.init first, and you still choose the project and config. The integration is only a convenience for the run.log calls.
If you use a framework that watches gradients, run.watch(model) logs gradient and parameter histograms during training for a PyTorch model. It is useful for diagnosing vanishing gradients, and it also adds overhead, so switch it off once the model behaves.
- Open a notebook, log in, and run a short training cell inside a
with wandb.init(...)block. - Re-run the cell twice. Confirm each execution creates exactly one finished run.
Reading errors, and fixing the common ones
Most first-week problems fall into a handful of messages. Learn to read them by their first line.
UsageError: No API key configured. Use wandb login to log in. The script has no credentials and no terminal to ask for them, as happens on servers, in cron jobs and in containers. Set WANDB_API_KEY, or run wandb login beforehand. A quick test is wandb.login(prompt=False), which returns False when nothing is configured.
Invalid API key or a key-length message. Usually the key was truncated when pasted, or the variable includes quote marks or trailing whitespace. Generate a fresh one. If you see API key must be 40 characters long, yours was 86, your library is older than 0.22.3. Run pip install -U wandb, then wandb login --relogin.
AuthenticationError: Both WANDB_API_KEY and WANDB_IDENTITY_TOKEN_FILE are set. These two are alternatives. Unset one.
Failed to verify credentials. Your key is fine but the destination is wrong, for example a key from the hosted service used against a company's own server, or the reverse. Run wandb status and check which host is configured.
Timed out initializing run followed by advice about init_timeout. The script could not reach W&B within 90 seconds, often behind a firewall, a proxy or a bad connection. Fix the network, or run with mode="offline" and sync later. Raising the timeout is a last resort:
wandb.init(settings=wandb.Settings(init_timeout=180))
ConfigError: Attempted to change value of key "lr". You assigned a config key that was already set. Decide whether it is really config. If it changes, log it as a metric. If you truly mean it, pass allow_val_change=True.
Run (<id>) is finished. The call to log will be ignored. You logged after the run ended. Move the run.log call inside the with block.
Invalid project name. The name contained a forbidden character, such as a slash, or exceeded 128 characters.
429 Rate limit exceeded. You are logging faster than the service allows. Metric uploads are retried with backoff and do not stop your training, but they can delay the end of the run. Log less often, put several metrics in one log call, and stagger many simultaneous job starts.
402 Payment Required. You have hit a plan or storage limit. Since 0.30.0 this is reported at once and not retried. Delete old artifacts or raise the plan.
A run shows Crashed. Open the Logs tab and the files debug.log and debug-internal.log in the Files tab. Out-of-memory kills are the usual cause. If training kept running, recover with wandb sync.
AttributeError after an upgrade. Something was removed. Commonly hit examples are wandb.PublicApi, run.get_url(), run.project_name() and the length attribute on the results of api.runs. The replacements are wandb.Api(), the properties run.url and run.project, and len(runs).
wandb.Image(normalize=...), define_metric(summary="best") and wandb.finish(quiet=True) were all removed in 0.29.0. The replacements are pre-scaled pixel arrays, summary="min" or "max", and wandb.Settings(quiet=True).
The last habit is to read the debug logs when the error text is not enough. Each run folder contains logs/debug.log and logs/debug-internal.log, and support staff will ask for them.
- Temporarily unset
WANDB_API_KEYand move your.netrc, then run a script in a fresh shell without a terminal prompt. Read the error. - Set a config key twice in a script without
allow_val_changeand read theConfigError. - Restore your credentials afterwards.
Putting it all together
Here is one small end-to-end project that uses every idea above. It prepares a dataset as an artifact, trains from it with config and metrics, logs the trained model as an artifact, and can be swept. Save it as project.py.
import json
import random
import wandb
DEFAULTS = {"learning_rate": 0.05, "epochs": 30, "noise": 0.3, "seed": 42}
def make_points(noise, seed):
random.seed(seed)
xs = [i / 10 for i in range(50)]
ys = [2.0 * x + 1.0 + random.gauss(0, noise) for x in xs]
return xs, ys
def prepare():
with wandb.init(project="line-fit", job_type="prepare-data", config=DEFAULTS) as run:
xs, ys = make_points(run.config["noise"], run.config["seed"])
with open("points.json", "w") as f:
json.dump({"xs": xs, "ys": ys}, f)
art = wandb.Artifact("line-data", type="dataset", metadata={"rows": len(xs)})
art.add_file("points.json")
run.log_artifact(art)
def train():
with wandb.init(project="line-fit", job_type="train", config=DEFAULTS) as run:
run.define_metric("train/loss", summary="min")
folder = run.use_artifact("line-data:latest").download()
with open(f"{folder}/points.json") as f:
data = json.load(f)
xs, ys = data["xs"], data["ys"]
w = b = 0.0
lr = run.config["learning_rate"]
for epoch in range(run.config["epochs"]):
loss = gw = gb = 0.0
for x, y in zip(xs, ys):
err = (w * x + b) - y
loss += err * err
gw += 2 * err * x
gb += 2 * err
n = len(xs)
w -= lr * gw / n
b -= lr * gb / n
run.log({"epoch": epoch, "train/loss": loss / n,
"weights/w": w, "weights/b": b})
with open("model.json", "w") as f:
json.dump({"w": w, "b": b}, f)
model = wandb.Artifact("line-model", type="model",
metadata={"final_loss": loss / n})
model.add_file("model.json")
run.log_artifact(model, aliases=["best"])
if __name__ == "__main__":
prepare()
train()
Run python project.py. Afterwards, in the dashboard you should find one prepare-data run and one train run, a dataset artifact, a model artifact, and a lineage graph connecting them in order: data, to training run, to model. Open the train run and confirm that learning_rate and epochs appear in its config and that the summary shows the minimum loss. For a sweep, point program in sweep.yaml at a small wrapper that calls only train(), since the data needs preparing once.
Notice what this gives a teammate. They can open the model artifact, follow its lineage to the exact data version and the exact run, read the settings, and reproduce the result.
- Run
project.pyend to end and find the lineage graph. - Change
noiseto0.6, rerun, and confirm the dataset artifact gains av1and the new model is linked to it. - Compare the two training runs in the project workspace and explain the difference in final loss.
What you can now do, and what comes next
You can install W&B, authenticate without leaking a key, write a training script that records its settings and metrics, and read the resulting dashboard. You can version a dataset and a model as artifacts and follow their lineage, describe a sweep in YAML, work offline and sync later, and read the common errors by their first line.
Here is the quick reference for the everyday tasks.
| You want to... | Use |
|---|---|
| Start recording a run | with wandb.init(project=..., config=...) as run: |
| Record a number each step | run.log({"train/loss": value}) |
| Keep the best value as the summary | run.define_metric("val/loss", summary="min") |
| Show predictions as a sortable table | wandb.Table logged with run.log |
| Version a file or folder | wandb.Artifact, then run.log_artifact |
| Fetch a version by name | run.use_artifact("name:latest").download() |
| Search hyperparameters | wandb sweep sweep.yaml, then wandb agent |
| Work without a network | WANDB_MODE=offline, then wandb sync |
| Turn W&B off in tests | WANDB_MODE=disabled |
| Check that login works | wandb login --verify |
Mid-level takes each topic further. It covers how the client and the wandb-core process actually move data, resuming and forking runs that were interrupted, artifact workflows with aliases and the Registry, sweeps with early termination, the public API for pulling runs back into pandas, grouping runs for cross-validation, and using W&B from CI. It also connects W&B to its neighbours, including MLflow and Docker, so the tracking fits the rest of your pipeline.
Senior covers what you own when W&B is a platform for a team: deployment choices between the hosted service, Dedicated Cloud and self-managed Kubernetes, scaling limits and what slows down first, security through single sign-on, service accounts and identity federation, bringing your own storage bucket for data-residency needs, rate limits, upgrades, automations, and when to choose a different tool.
The best next step is a real one. Take a model you already train, add five lines of W&B to it, run it three times with different settings, and answer the question you wrote in the first section.
Sources
- W&B quickstart
- Create an experiment
- Configure experiments
- Log data from experiments
- Environment variables
- Limits and performance guidance
- Run states
- Command-line reference
- wandb login
- wandb sync
- wandb sweep
- wandb agent
- wandb artifact
- Python reference: init
- Python reference: Run
- Python reference: Artifact
- SDK release notes
- SDK changelog at v0.30.0
- Breaking-change policy
- wandb on PyPI