This is part one of three. It covers everything you need to do real work with Comet's experiment tracking, not a teaser. By the end you can install the Python SDK, log in, record a training run with its settings, its metrics and its results, open the dashboard and read it, compare runs, version a dataset, register a model, and recover when a run fails to upload. Mid-level and Senior take the same topics further. Nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and tracking only makes sense once you have watched your own numbers appear in a browser tab and then compared them with a second run.
One naming note before we start. Comet is the company, and it sells two families of products on one platform. The first is the MLOps family: experiment management, artifacts, the model registry, the hyperparameter optimizer, and production monitoring. The second is Opik, an open-source product for tracing and evaluating LLM applications. This guide is about the first family and the Python package comet_ml. If you want to trace prompts and agents, go to the Opik guide after this one.
What Comet is, and what you can do with it
Comet is a place to keep a record of every machine-learning run you ever do. You add a few lines to your training script. Each time the script runs, the SDK streams the run's hyperparameters, metrics, code, environment, and any files you choose to a server. You then open a web dashboard and see the run as a page: loss curves drawn live, a table of settings, the exact code and git state that produced it, the Python packages installed, even CPU and GPU usage.
The value is not the charts. The value is that the record exists. Three weeks from now, when someone asks which learning rate produced the model in production, you do not search your terminal history. You open the run and read it.
Here is what a beginner can do by the end of this guide:
- Run a training script and see its metrics update live in a browser.
- Record the settings of each run so that two runs can be compared in a table.
- Attach a confusion matrix, a chart, a dataset and a trained model to a run.
- Give runs readable names and tags so you can find them later.
- Promote a good model into the model registry with a version number.
- Fetch your runs back from Python to analyse them in a notebook.
- Keep working when the network drops, by logging offline and uploading later.
The diagram is the whole idea. Everything else in this guide is detail about the four boxes.
https://www.comet.com in a browser and create a free account. Do not install anything yet. Just click around the empty workspace so the screens feel familiar when your first run arrives.
The problem Comet solves, and what came before
Machine learning has a problem that ordinary software mostly does not. In ordinary software, the code is the thing. If two developers have the same commit, they have the same program. In machine learning, the same code can produce a different result each time, depending on the data version, the random seed, the hyperparameters, the library versions, the GPU, and the amount of patience you had when you pressed Ctrl+C. The result you care about is a function of everything, and most of that everything is not in git.
Before experiment trackers, people coped in a few ways, and you have probably used at least one of them.
The notebook with a graveyard of cells keeps results in cell outputs. It works until you restart the kernel, and then the numbers are gone. Nobody can tell which cell ran in which order.
The spreadsheet holds one row per run: learning rate, batch size, accuracy. It works until you forget to fill a row, mistype a number, or need the loss curve rather than the final value. The spreadsheet is also never linked to the code that produced the row.
The folder naming convention gives you directories like run_final_v2_lr0.001_REALLY_final. It works for a week.
The TensorBoard log directory was a real step forward, because it records curves automatically. But logs live on the machine that trained, they are awkward to share, they do not capture the code and environment, and comparing forty runs from three laptops and a cluster means copying files around.
A hosted experiment tracker replaces these with one habit: every run logs itself to a shared server, and the server stores the settings, the curves, the code, and the files together. Comet is one such tracker. Others are covered elsewhere in this series, for example MLflow, Weights and Biases and ClearML. They differ in details, but the core idea is the same, and what you learn here transfers.
What Comet adds on top of plain logging is a set of neighbours around the experiment. Artifacts version your datasets. A model registry holds the models you promoted, with version numbers and approval states. An optimizer runs hyperparameter searches and logs each trial as a run. Panels and reports let you turn charts into shared documents. A beginner needs the first two or three. The rest wait for later levels.
The mental model: four nouns and a hierarchy
Comet's documentation describes a hierarchy of containers. From the outside in:
- An Organization is your company or school. There is normally one.
- A Workspace sits inside it, typically one per team. Every user also has a personal default workspace.
- A Project sits inside a workspace, typically one per machine-learning problem, such as "churn-prediction" or "arabic-sentiment".
- An Experiment sits inside a project. It is one execution of your code.
As a beginner you will live almost entirely at the bottom two levels. You will pick a project name, and every run you start becomes an experiment in it. If you do not pick a project name, runs land in a project called Uncategorized. That works but fills up with clutter, so name your projects from the start.
Now the nouns you need to be precise about.
An experiment is the unit of measurement. The docs define it as a single execution of code with some associated data. One training run with one set of hyperparameters is one experiment. Each experiment has an experiment key, a unique identifier. By default Comet generates a random one. If you supply your own, it must be alphanumeric and between 32 and 50 characters long. You will rarely need to.
A parameter is a setting of the run: learning rate, batch size, number of trees, the name of the dataset. A parameter is set once, or rarely changes during the run. You log parameters so that runs can be compared by their settings.
A metric is a measured value that usually changes over time: training loss at each step, validation accuracy at each epoch. A metric is a time series. Each value carries a step (how many batches so far) and optionally an epoch, plus a timestamp. Comet draws metrics as line charts. One detail to know early: beyond 15,000 values for a single metric in a single experiment, Comet down-samples what it displays, so a very dense curve may look slightly thinned. That is almost never a problem for beginners.
An asset is a file attached to one experiment: an image, a chart, a plot, a saved model, a piece of source code. Assets belong to the run that logged them and are not versioned.
An artifact is different, and the difference trips beginners. An artifact is a versioned object that lives at the workspace level, not inside one run. Each version is an immutable snapshot of files. The typical use is a dataset: train-data version 1, then version 2 after you clean it. Comet records which experiment produced each artifact version and which experiments used it, so you can trace a model back to the exact dataset it learned from. If an asset is "a file stuck to this run", an artifact is "a named, versioned thing many runs share".
Two more terms appear in this guide. A model logged with log_model() is a set of files attached to an experiment, shown in the Assets and Artifacts tab. A registered model is what you get after promoting one into the model registry, a workspace-level catalogue where every version has a number, tags, and a status. We come to that later.
Installing the SDK and checking the setup
The SDK is a pure Python package named comet_ml. The steps are the same on Linux, macOS and Windows. It needs Python 3.8 or newer. Create a virtual environment first, so that this course does not disturb other projects:
python -m venv .venv
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install comet_ml
You can also write pip install comet-ml; PyPI treats the two spellings as the same package. In a notebook such as Jupyter or Colab, use %pip install comet_ml in a cell. If you prefer conda, the docs give conda install -c anaconda -c conda-forge -c comet_ml comet_ml.
The install adds one command-line tool named comet. Check that it works and see which SDK version you got:
comet --version
At the time of writing, the current release is 3.58.7, so the command prints 3.58.7. If your shell says comet: command not found, the scripts directory of your environment is not on your PATH. On Windows this is common. The fix is to activate your virtual environment again or call the module through Python; the simplest check is pip show comet_ml.
NumPy is optional for the SDK. If it is missing you will see a line saying numpy not installed; some functionality will be unavailable. Since nearly every machine-learning script installs NumPy anyway, that message usually never appears.
Now the most useful diagnostic you will meet in this guide:
comet check
This command tests whether your machine can reach the three Comet services the SDK talks to: the server that receives logged data, the REST API, and the optimizer service. It also prints your Python version, any proxy settings, whether TLS certificate checking is on, and the path of the certificate bundle in use. It ends with a summary line such as Server connectivity True. Run it once right now, before you have any problem, so you know what healthy output looks like. Add --debug when something is wrong, to see the raw HTTP exchange.
If you have no API key yet, comet check can still show whether the network path is open. A corporate proxy or a security gateway that inspects encrypted traffic will show up here as a failed connection or a certificate complaint. Section thirteen covers what to do then.
pip install comet_ml, then comet --version and comet check. Copy the final connectivity lines into a note. If any line says False, stop and read the section on common errors before continuing.
Logging in: the API key and where it lives
Comet needs to know who you are. It identifies you with an API key, a long secret string tied to your account. To find it, open the Comet website, click your avatar, choose Account settings, open the API keys tab, and copy the key.
Treat the key like a password. Anyone who has it can log runs into your account and read your projects. Never paste it into code that you commit. This is more important with Comet than with many tools, because Comet automatically records your source code, so a key typed inside the script would be uploaded along with it.
The safe and simple way to store it is the login command:
comet login
It prompts for the key and writes it into a configuration file named .comet.config in your home directory. That is ~/.comet.config on Linux and macOS, and %USERPROFILE%\.comet.config on Windows. From then on, every script on this machine finds the key automatically.
From Python, the equivalent is:
import comet_ml
comet_ml.login()
This call asks for the key only if none is configured already. If a key exists, it does nothing visible. There are two variants worth knowing. comet_ml.login(anonymous=True) logs to a public, claimable project without an account; it is fine for trying the tool, but never for real or sensitive data, because anyone can see what you log. And comet_ml.login(project_name="my-project") saves extra settings such as the project name into the config file as well.
Older tutorials tell you to call comet_ml.init() or to run comet init --api-key. Both are deprecated. Using them prints a warning such as comet_ml.init() is deprecated and will be removed soon. Please use comet_ml.login(). Use login everywhere and ignore any tutorial that uses init.
You can also provide the key through an environment variable, which is the standard approach on servers and in automated jobs:
export COMET_API_KEY="paste-your-key-here" # macOS and Linux
$env:COMET_API_KEY="paste-your-key-here" # Windows PowerShell
Which source wins when several are present? Comet reads its settings in a fixed order of priority. Arguments written directly in code come first. Then operating-system environment variables. Then a .env file in the current folder. Then a file named by COMET_CONFIG. Then a .comet.config in the current folder, and finally the one in your home directory. The practical rule is that the more specific and closer to the run a setting is, the more power it has. A value in code beats a value in the environment, which beats a value in a file.
comet_ml.login(api_key="abc...") puts the secret into a file that Comet uploads and that you may push to GitHub. The docs specifically warn against it. Use comet login, an environment variable, or a secrets store. In Colab, read it with userdata.get("COMET_API_KEY") instead of pasting it.
The config file uses the INI format. Open yours after logging in and you will see something like this:
[comet]
api_key = your-key-here
workspace = my-team
project_name = my-project
You can edit it by hand to set a default workspace and project. Each line has an environment-variable equivalent: COMET_API_KEY, COMET_WORKSPACE, COMET_PROJECT_NAME. The pattern is COMET_ plus the setting name in capitals.
comet login, paste your key, then open ~/.comet.config and confirm the key is there. Add a line project_name = first-steps under [comet]. Run comet check again and confirm the server connectivity line still says True.
Your first experiment, step by step
Let us log a run that does no real machine learning, so that nothing distracts from the Comet part. Create a file called hello_comet.py:
import comet_ml
comet_ml.login()
exp = comet_ml.start(project_name="first-steps")
exp.log_parameters({"batch_size": 32, "learning_rate": 1e-4})
for step in range(100):
exp.log_metrics({"loss": 1 / (step + 1)}, step=step)
exp.end()
Run it with python hello_comet.py. Read the console output as it goes. Near the top you will see a line like COMET INFO: Experiment is live on comet.com followed by a web address. At the end you will see a summary of what was logged. The summary level is one by default and you can ask for more detail later.
Now walk through the script line by line, because each line teaches a concept.
import comet_ml brings in the SDK. Import order matters, and we will return to it in a moment.
comet_ml.login() makes sure the key is available. If you already ran comet login, this line is harmless.
comet_ml.start(project_name="first-steps") is the heart of it. It creates a new experiment in the project named first-steps, creating the project if it does not exist. It returns an object, here called exp, which you use for all logging. Since version 3.47.1, start() is the recommended way to begin a run. You will see older code that writes comet_ml.Experiment(...). That still works, but the current docs recommend start, and this guide uses it throughout.
exp.log_parameters({...}) records settings. You pass a dictionary. Nested dictionaries are supported. You can call log_parameter("name", value) for a single one.
exp.log_metrics({...}, step=step) records a measured value at a given step. Here the loss falls as one over the step number, which produces a nice decaying curve. You can log several metrics in one call by putting several keys in the dictionary. Calling it once with a dictionary is cheaper than calling log_metric many times, which matters when you hit rate limits later.
exp.end() tells Comet the run is finished. It waits until all queued data has been sent and then closes the experiment. Always call it, especially in notebooks. In a notebook, end() also records the cells you executed.
Open the web address printed in the console. You will land on the experiment page. Look at these tabs, which are the ones you will use daily:
- Charts or Panels: your
losscurve. Hover over it to read values. - Parameters: a table with
batch_sizeandlearning_rate. - Metrics: the final, minimum and maximum of each metric.
- Code: the exact script that ran, stored automatically.
- System metrics: CPU, memory and, if present, GPU usage over time.
- Installed packages: every package and version in your environment.
- Output: what the script printed to the console.
You wrote three lines of logging and got an archive of code, environment and results. That trade is the reason people use the tool.
Now the import-order rule. If your script uses PyTorch, TensorFlow, Keras, fastai or similar, put import comet_ml before importing them. Comet hooks into those libraries to log metrics automatically, and it can only do so if it is loaded first. If you get the order wrong, the current SDK prints a warning: To get all data logged automatically, import comet_ml before the following modules. Older documentation describes this as an ImportError, but the behaviour today is a warning and partial logging. There are two escape hatches: run your script with comet python script.py, which prepares the SDK before your code starts, or set COMET_DISABLE_AUTO_LOGGING=1 to turn the automatic logging off and log everything by hand.
Experiment(). Tutorials from before 2024 do. It still works, but comet_ml.start() supersedes it and is where new features land. The mapping is simple: Experiment() becomes start(), and OfflineExperiment() becomes start(online=False).
hello_comet.py. Open the link. Find the loss curve, the two parameters, and the Code tab. Then change learning_rate to 1e-3, change the loss formula to 2 / (step + 1), and run it again. You now have two experiments in the same project. Leave them for the comparison section.
A real training run with scikit-learn
The toy script showed the mechanics. A real run has data, a model, a split, training, and evaluation. We will use scikit-learn because it is installed on most machines and needs no GPU. Install it with pip install scikit-learn if you need to. The example trains a random forest on the built-in breast-cancer dataset and logs what a beginner should always log.
import comet_ml
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, f1_score
from sklearn.model_selection import train_test_split
import joblib
comet_ml.login()
params = {
"n_estimators": 200,
"max_depth": 6,
"random_state": 42,
"test_size": 0.2,
}
exp = comet_ml.start(project_name="breast-cancer")
exp.set_name("rf-200-trees-depth6")
exp.add_tags(["baseline", "random-forest"])
exp.log_parameters(params)
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=params["test_size"], random_state=params["random_state"]
)
exp.log_parameter("n_train_rows", len(X_train))
model = RandomForestClassifier(
n_estimators=params["n_estimators"],
max_depth=params["max_depth"],
random_state=params["random_state"],
)
model.fit(X_train, y_train)
with exp.train():
exp.log_metric("accuracy", accuracy_score(y_train, model.predict(X_train)))
with exp.test():
preds = model.predict(X_test)
exp.log_metric("accuracy", accuracy_score(y_test, preds))
exp.log_metric("f1", f1_score(y_test, preds))
exp.log_confusion_matrix(y_true=y_test, y_predicted=preds)
joblib.dump(model, "rf_model.joblib")
exp.log_model("breast-cancer-rf", "rf_model.joblib")
exp.end()
Run it and open the experiment. Here is what each new idea does.
exp.set_name(...) gives the run a readable name. Without it, Comet assigns a random name. Names are for humans. A good name states what is different about this run, such as rf-200-trees-depth6. You can change it later in the interface.
exp.add_tags([...]) attaches short labels. Tags are for filtering: show me everything tagged baseline. You can add a single tag with add_tag.
The with exp.train(): and with exp.test(): blocks are contexts. Inside a context, Comet prefixes the metric names, so the same call log_metric("accuracy", ...) becomes train_accuracy in one block and test_accuracy in the other. This is how you keep training and evaluation numbers apart without writing different names by hand. There is also exp.validate() for a validation context.
exp.log_confusion_matrix(y_true=..., y_predicted=...) takes the true labels and the predicted labels and draws a confusion matrix in the dashboard. A confusion matrix is a table that shows, for each true class, how many examples were predicted as each class. It tells you not just how often the model is wrong but how it is wrong, which a single accuracy number hides.
exp.log_model("breast-cancer-rf", "rf_model.joblib") uploads the model file and attaches it to the run under the name you chose. The first argument is the name, the second is the path of a file or folder. The docs recommend logging models with log_model on the experiment, rather than as artifacts, because logged models can later be promoted to the registry.
Look at the Assets and Artifacts tab after the run. The model file is there, downloadable.
A word about what you did not have to write. Comet also recorded the script text, the git commit if you are in a repository, the operating system, the Python version, the installed packages, and the console output. These happen automatically. They are controlled by settings you can turn off, which Mid-level covers. If your code or data is sensitive, consider these defaults. For example, if your employer has data-residency rules, be aware that the source code and console output are uploaded to Comet's servers, or to your own servers if your organisation self-hosts.
random_state is in the parameters on purpose. A run you cannot reproduce is half a record. Every random seed, split ratio and dataset name belongs in log_parameters.
train_rf.py. In the dashboard open the run, find test_accuracy, test_f1 and the confusion matrix, then download the model from the Assets tab. Then change max_depth to 3, change the name and run again. Two runs, one project. Keep them.
Reading and using the dashboard
You now have several runs in a project. Click the project name to open the project view. This is a table with one row per experiment, plus a set of charts above or beside it. The table is the most useful part for a beginner.
The columns show run names, status, creation time, the parameters you logged, and the last value of each metric. Click a column header to sort. Click the column chooser to show or hide columns. If you logged max_depth as a parameter, it is now a column you can sort by. That is the payoff for logging parameters properly: the question "which setting gave the best F1?" is answered by sorting.
You can filter the table. Type in the search box or build a filter on a parameter, a metric, or a tag. Filtering by the baseline tag shows only runs you tagged that way. This is why tags deserve the thirty seconds it takes to add them.
Select two or more rows and choose Compare. The compare view overlays the curves and shows parameter differences side by side, highlighting what changed. In the current cloud release the compare view handles up to ten experiments at once, while the project view can chart many more. Comparing is the thing you do a hundred times, so get comfortable with it early.
Panels are the charts. A project view and an experiment view are made of panels, and you can add built-in ones such as line charts, scatter plots, bar charts and parallel-coordinates plots. A scatter plot of learning rate against accuracy, or parallel coordinates across several hyperparameters, shows patterns a table cannot. You can also write custom Python panels, but that belongs to later levels.
If you work in a notebook, you can even embed the dashboard inside a cell:
exp.display(tab="panels")
The tab argument accepts names such as code, parameters, metrics, output, system-metrics, installed-packages, assets and artifacts. This is handy in Colab or Jupyter when you do not want to switch windows.
breast-cancer project, sort the table by test_f1, filter by the tag baseline, then select your two runs and open Compare. Write down in one sentence what changed between them and whether it helped.
Logging more: images, figures, text, code and notes
Metrics and parameters cover most of what you log. A few more methods earn their place, and they all follow the same pattern: call a method on exp, and something appears in the dashboard.
To attach a chart you made with Matplotlib, use log_figure:
import matplotlib.pyplot as plt
fig, ax = plt.subplots()
ax.bar(["a", "b", "c"], [3, 7, 5])
exp.log_figure(figure_name="example-bar", figure=fig)
The figure appears in the graphics tab. From version 3.51.0, the format argument defaults to automatic selection, so you rarely set it.
To attach an image from a file or an array, use log_image:
exp.log_image("confusion.png", name="confusion-plot")
To attach any file, such as a CSV of predictions or a text report, use log_asset:
exp.log_asset("predictions.csv", file_name="predictions.csv")
To store free-form key-value data that does not fit parameters or metrics, such as the name of the GPU node or the ticket number, use log_other("ticket", "ML-142").
To record the source code explicitly, for example when the automatic capture cannot find your script, call exp.log_code(file_name="train_rf.py"). You will see this needed when your entry point is wrapped by a launcher. If the console prints COMET ERROR: Failed to set run source code, this is the fix.
Steps and epochs deserve a word, because people confuse them. A step is usually one batch or one update. An epoch is one full pass through the training data. log_metric(name, value, step=..., epoch=...) accepts both. If you always log at the end of an epoch, pass epoch. If you log per batch, pass step. You can set defaults with exp.set_step(n) and exp.set_epoch(n), and later calls use them. Parameters accept step only.
On rates: Comet limits how fast the cloud accepts data, per data type. Metrics are accepted at 12,000 per minute, parameters at 10,000 per minute, and console output at 12,000 per minute. If you exceed them you see Experiment has been throttled. Some data (like experiment metrics) might be missing. The cure is to log per epoch rather than per batch, or to send several metrics in one log_metrics call. Artifacts are never throttled.
train_rf.py. Add a Matplotlib bar chart of model.feature_importances_ for the top ten features with log_figure, add log_other("dataset", "sklearn-breast-cancer"), rerun, and find both in the dashboard.
Organising runs: projects, names, tags and contexts
After twenty runs, a project without structure becomes hard to read. Four habits keep it usable.
One project per problem. Do not put every experiment of the course into Uncategorized. Name projects after the question they answer: breast-cancer, arabic-sentiment. Set the name in start(project_name=...) or in the config file, and the project is created automatically. Choose the name deliberately, because a typo creates a new project silently. Once it exists, you can move or rename runs in the interface.
Names that describe the difference. run-7 tells you nothing. rf-200-trees-depth6 tells you what the run is. Setting a name costs one line. If you forget, you can rename in the dashboard.
Tags for categories that cut across runs. Use a short, consistent vocabulary: baseline, tuned, bug, final, ablation. Agree on it with your team, because Baseline and baseline are two tags.
Contexts for train, validate and test. As shown earlier, the with exp.train(): style keeps metric names consistent.
breast-cancer with different n_estimators, each with a name describing the change and the tag tuned. In the project table, filter by tuned and sort by test_f1.
Datasets as artifacts
Models are the output of a run. Data is the input, and data changes. If you clean your dataset on Tuesday and retrain on Wednesday, a score change might come from the cleaning rather than the model. Artifacts exist to make that visible.
An artifact is created locally, given files, and logged from an experiment:
from comet_ml import Artifact, start
exp = start(project_name="breast-cancer")
art = Artifact(name="cancer-train-data", artifact_type="dataset")
art.add("data/train.csv")
exp.log_artifact(art)
exp.end()
The constructor takes a name and an optional type, such as dataset or model. add attaches a local file or folder. log_artifact uploads it and creates version 1.0.0. If you log an artifact with the same name again and the content differs, Comet creates a new version. Nothing is overwritten; each version is an immutable snapshot. Artifact names are limited to 100 characters.
Later, in another run or another script, you use the dataset by name:
exp = start(project_name="breast-cancer")
logged = exp.get_artifact("cancer-train-data")
logged.download("./data")
The signature is get_artifact(artifact_name, workspace=None, version_or_alias=None). Without a version it fetches the latest. Because you fetched it through the experiment, Comet records that this run consumed that artifact version. The artifact page then shows its lineage: which run produced it and which runs used it. When a result looks suspicious, lineage tells you which dataset version was in play.
If your data is large or already lives in cloud storage, you do not have to upload it. A remote artifact records a reference instead:
art = Artifact(name="big-images", artifact_type="dataset")
art.add_remote("s3://my-bucket/images/")
exp.log_artifact(art)
For S3 or Google Cloud Storage addresses, the SDK lists the objects under the prefix and records their addresses and checksums, and their version identifiers if bucket versioning is on. Your machine needs credentials configured the provider's normal way. Without them, only the address string is recorded. This suits Gulf and Egyptian teams who keep data in a regional bucket for residency reasons: the data stays in your region, and Comet stores only a pointer and metadata.
If you already know DVC, think of artifacts as a hosted, UI-visible alternative for small and medium datasets. They serve a similar goal with a different workflow.
log_model on the experiment. The docs recommend this so models can be registered. Artifacts are for datasets and other inputs that many runs share.
data/train.csv, log it as an artifact, then change one row, log it again under the same name, and open the artifact page. Confirm that two versions exist and that the experiment is listed as the producer.
The model registry: from a good run to a named model
Logging a model on a run keeps it attached to that run. The model registry is the catalogue where you keep the models you actually intend to use. Each registered model has a name, and each version has a number, so the question "which model is in production?" has an answer.
The flow has two steps. First log the model on the experiment, then register it:
exp.log_model("Breast Cancer RF", "rf_model.joblib")
exp.register_model("Breast Cancer RF")
Two details. The registry name becomes lowercase with hyphens: "Breast Cancer RF" appears as breast-cancer-rf. And the first version gets the default version string 1.0.0. Versions follow semantic versioning, which means three numbers separated by dots, such as 1.0.0 or 2.1.5. Register again from a better run and the new version is created under the same model name, which you can also set explicitly with the version argument.
Open the registry in the dashboard: look for the model registry in the left navigation, click your model, and you see its versions, each with the run it came from, tags, notes, and a history of changes.
Older tutorials talk about model stages: Staging, Production. Stages are deprecated. Their data became tags, and the official lifecycle mechanism is now status: a model version has a status such as Development, Staging or Production, which you set per version. In many setups, a regular member changes the status by sending an approval request that a workspace owner approves. This is deliberate: it puts a human between "a run finished" and "a model is labelled production". Using the stage-based API methods prints deprecation warnings, so skip them.
From code, after a run has ended, you can manage versions through the API object:
api = comet_ml.API()
model = api.get_model(workspace="my-team", model_name="breast-cancer-rf")
model.add_tag(version="1.0.0", tag="reviewed")
model.set_status(version="1.0.0", status="Staging")
model.download("1.0.0", output_folder="./downloaded-model")
get_model takes the workspace and the registry name. add_tag and set_status take the version first. download fetches the files of one version into a folder you name. From the command line there is a matching helper:
comet models list -w my-team
comet models download -w my-team --model-name breast-cancer-rf --model-version 1.0.0 --output ./model
The built-in help text of the command contains example flags that do not exist, written with underscores. The real flags are --model-name and --model-version, with hyphens, as above. If you paste an example from the help and it is rejected, that is why.
Registered models are also the point where serving tools and deployment pipelines connect: a deployment job asks the registry for "version 1.0.0 of this model" instead of copying a file from someone's laptop.
breast-cancer runs, log the model as Breast Cancer RF and register it. Open the registry. Confirm the name is lowercase, the version is 1.0.0, and add a note saying which run and why it was chosen.
Getting your data back: the API and querying
The SDK has two halves. comet_ml.start() gives you an object for logging during a run. comet_ml.API is a thin client over Comet's REST interface and is for reading afterwards: listing experiments, searching, downloading. You use it in an analysis notebook after your runs are done.
import comet_ml
api = comet_ml.API()
experiments = api.get_experiments("my-team", project_name="breast-cancer")
for e in experiments:
print(e)
The call takes a workspace and a project name, and returns a list of experiment objects. You can also fetch one by its path, api.get("my-team/breast-cancer/<experiment-key>"), or by key alone with api.get_experiment_by_key(key). The old get_experiment_by_id is deprecated.
To filter by content, use the query helpers:
from comet_ml.query import Tag, Metric, Parameter
baseline = api.query("my-team", "breast-cancer", Tag() == "baseline")
You can build conditions on tags, metrics and parameters. From version 3.54.0 there is also api.search(workspace, project_name, query), which supports combining conditions with AND and OR and handles paging through large result sets.
To bring metric curves into a pandas table for your own plotting:
df = api.get_metrics_df(
experiment_keys=[e.id for e in experiments],
metrics=["test_accuracy"],
)
print(df.head())
The function returns a data frame with one row per logged value. This is how you build your own comparison plot, or a table for a report, without using the web interface.
You can also append to a finished run using APIExperiment, for example to attach a late-arriving evaluation score. comet_ml.APIExperiment(previous_experiment="<key>") gives you a handle on an existing run for one-off reads and writes through the REST API. To continue a training run that is still going, rather than patch one afterwards, use comet_ml.start(mode="get", experiment_key="<key>").
breast-cancer project with api.get_experiments, then use get_metrics_df to load test_accuracy for all of them and print the best run's name.
Working offline and when the network fails
Not every training job has a good internet connection. You might train on a cluster without outbound access, on a train, or behind a firewall that blocks Comet. Comet handles this with offline experiments, which write everything to a local archive that you upload later.
exp = comet_ml.start(online=False, project_name="breast-cancer")
Log as normal. When you call exp.end(), the SDK writes a .zip file into a folder named .cometml-runs by default, and prints a message telling you the exact command to upload it:
comet upload .cometml-runs/*.zip
You can point comet upload at individual archives, and pass --workspace and --project-name to override where they go. Offline experiments have no rate limits. The only limits are the hard limits on counts, such as the number of images.
You do not need to choose offline in advance. Comet has a fallback to offline: if the connection drops during a run and the SDK cannot re-establish it after several retries, it switches to logging locally. If the connection returns, it uploads what it held. If it does not, you get a zip. The console tells you so: Could not send live data to Comet during experiment runtime. An offline experiment will be available for upload, followed by a command. The printed command includes --force-reupload. That flag is deprecated; use --force-upload instead. If comet upload complains experiment was not found, you can upload it as a new experiment by using the --force-upload flag, it is telling you the same thing.
What matters for a beginner is to not panic. A failed upload is not a lost run. Look for a zip file, read the last lines of the console, and run the command.
Finally, sometimes you want a script to run with Comet code in it but send nothing, for example during a dry run or a unit test. Set COMET_AUTO_LOG_DISABLE=1. This disables all network communication. Since version 3.37.2 it does not even need an API key.
start(online=False). Run hello_comet.py offline, find the zip in .cometml-runs, then reconnect and run comet upload on it. Check that the run appears in the dashboard.
Common errors and how to read them
Comet's messages start with a prefix: COMET INFO, COMET WARNING or COMET ERROR. Read the whole line; it usually names the cause. Here are the ones you will meet.
The given API key '...' is invalid ... Your experiment will not be logged. The key is wrong, or it belongs to a different installation, for example a key from the cloud used against a self-hosted server. Copy it again from Account settings, then API keys. Then run comet check.
API key is not set, or a ValueError from start(). No key was found anywhere. Run comet login or set COMET_API_KEY. To run code without logging, set COMET_AUTO_LOG_DISABLE=1.
Run will not be logged. The first handshake with the server failed. This points to the network, a proxy, a certificate problem, or an outage. Run comet check --debug. You can also test the server directly with curl -i https://www.comet.com/clientlib/isAlive/ping, which should return a short JSON body containing Healthy Server.
Failed to create Comet experiment, reason: .... Something went wrong inside start(), and the reason after the colon says what. Authentication, workspace and mode mistakes are the usual ones.
Workspace ... doesn't exist. A typo in the workspace name, or you are not a member. Check COMET_WORKSPACE.
To get all data logged automatically, import comet_ml before the following modules. You imported Comet after a machine-learning library. Move the import to the top of the file, or use comet python script.py.
Experiment has been throttled. You log too fast. Log per epoch, batch your metrics, or go offline.
Maximum number of experiments reached for project. A project holds at most 25,000 experiments. Archive old ones or use a new project. You will not meet this in the first months.
The SSL certificate message. There's seem to be an issue with your system's SSL certificate bundle. The message is Comet's wording. It means Python cannot verify the server's certificate. Three causes are common: a Python installed from python.org on macOS without its certificates installed, a company or security gateway that inspects encrypted traffic with its own certificate, or a self-signed certificate on a self-hosted server. The fix is to repair the certificate store, or to point the REQUESTS_CA_BUNDLE environment variable at your organisation's certificate file. Turning verification off with COMET_INTERNAL_CHECK_TLS_CERTIFICATE=0 exists, but it removes the protection that keeps your key safe. Use it only as a last resort.
Comet failed to send all the data back. Shown at the end of a run, this means end() ran out of patience waiting for the upload, or the process was killed. Always call exp.end(), and check your network. The waiting time is controlled by COMET_TIMEOUT_CLEANING.
A deprecation warning about init or COMET_INI. Not an error. Replace init() with login(), and COMET_INI with COMET_CONFIG.
One more class of problem is not an error message at all: the dashboard looks wrong. If curves look thinned, you probably logged more than 15,000 values of one metric. If a metric is missing, check the name and whether a context added a prefix. If a run is stuck on "running", your script crashed before end(); it will time out, and you can archive it.
COMET_API_KEY=wrong and run hello_comet.py; read the message. Set the workspace to a nonsense name; read that message. Then fix both. Knowing what failure looks like is half of debugging.
Configuration you will actually touch
Most beginners need only a handful of settings, and the simplest place for them is the [comet] section of .comet.config or the matching environment variables.
| Setting | Environment variable | What it does |
|---|---|---|
api_key |
COMET_API_KEY |
Your secret key |
workspace |
COMET_WORKSPACE |
Default workspace for new runs |
project_name |
COMET_PROJECT_NAME |
Default project; otherwise Uncategorized |
offline_directory |
COMET_OFFLINE_DIRECTORY |
Where offline zips go; default .cometml-runs |
display_summary_level |
COMET_DISPLAY_SUMMARY_LEVEL |
How much the end-of-run summary shows; 0, 1 or 2 |
Options that control what is logged automatically go into comet_ml.ExperimentConfig, which you pass to start:
config = comet_ml.ExperimentConfig(
name="rf-baseline",
tags=["baseline"],
log_code=True,
auto_output_logging="simple",
)
exp = comet_ml.start(project_name="breast-cancer", experiment_config=config)
ExperimentConfig accepts the name and tags, as shown, plus switches for logging code, the git state, environment details, and framework parameters and metrics. You can set disabled=True to switch Comet off entirely, which is useful for unit tests.
train_rf.py to use ExperimentConfig for the name and tags, and to read the project name from COMET_PROJECT_NAME by removing it from start. Confirm the run lands in the right project.
Putting it all together
Here is one small end-to-end project that uses everything so far. The goal: train three random forests with different depths, log them properly, register the best, and read the results back.
import comet_ml
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import f1_score
from sklearn.model_selection import train_test_split
import joblib
comet_ml.login()
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
best_f1, best_key = 0.0, None
for depth in [2, 4, 8]:
params = {"n_estimators": 200, "max_depth": depth, "random_state": 42}
exp = comet_ml.start(project_name="breast-cancer-final")
exp.set_name(f"rf-depth-{depth}")
exp.add_tags(["final-comparison"])
exp.log_parameters(params)
model = RandomForestClassifier(**params).fit(X_train, y_train)
preds = model.predict(X_test)
score = f1_score(y_test, preds)
with exp.test():
exp.log_metric("f1", score)
exp.log_confusion_matrix(y_true=y_test, y_predicted=preds)
joblib.dump(model, f"rf_depth_{depth}.joblib")
exp.log_model(f"rf-depth-{depth}", f"rf_depth_{depth}.joblib")
if score > best_f1:
best_f1, best_key = score, depth
exp.end()
print("best f1:", best_f1, "at depth", best_key)
Run it. The three runs appear in breast-cancer-final. In the project view, filter by the tag and sort by test_f1. Open Compare on all three and read the parameter difference next to the curves. Open the best run, go to its Assets tab, and register its model. Then in a notebook, use comet_ml.API() to list the runs and pull test_f1 into a table.
Notice what happened without any extra effort: each run has its code, its environment, its settings and its confusion matrix stored. If a teammate asks next month why depth eight was chosen, you send a link.
Also notice what this small project does not do. It does not version the dataset; adding an artifact would. It does not automate the search; Comet's Optimizer would. It does not run on a schedule or in a pipeline. Those are the next steps. They belong to later levels, and to neighbouring tools such as Kubernetes for running jobs, and Evidently for watching a model after deployment.
What you can now do, and what comes next
You can install the SDK, verify it with comet --version and comet check, and log in without putting a secret in code. You can start an experiment with comet_ml.start(), log parameters and metrics, use contexts to separate training from testing, attach figures, confusion matrices, files and models, and end the run cleanly. You can read the dashboard: sort, filter, compare, and chart. You can version a dataset as an artifact and trace which run used it. You can register a model, tag it, and set its status. You can fetch runs back with the API, work offline, and read the common error messages.
What to practise next, in order of value:
- Use Comet for every training run in your next project, even tiny ones. The habit matters more than any feature.
- Always log the seed, the dataset name and the data size.
- Keep a tag vocabulary and use it.
- Move to the Mid-level guide, which explains how the SDK works under the hood: what it captures automatically, how configuration resolves, how to resume runs, how artifacts and the registry behave in a team, the hyperparameter Optimizer, and testing your logging code.
For the LLM side, where you trace prompts, tool calls and agents instead of training curves, go to the Opik guide. For a different tracker with the same core idea, compare MLflow. When you are ready to put a registered model behind an API, see BentoML.
Sources
- Comet documentation home
- Quickstart
- Python SDK overview
- Start an experiment
- comet_ml.start reference
- ExperimentConfig reference
- Experiment reference
- API reference
- Running offline
- Fallback to offline
- Command line overview
- Configure the SDK
- Log metrics and parameters
- Log models
- Using artifacts
- Remote artifacts
- Troubleshooting and FAQ
- Limits and performance
- Changelog