Skip to content
Back to student guides
Azure Machine LearningMLOpsCloud ML platforms3 levels115 sectionsCovers azure-ai-ml 1.35

The Complete Azure Machine Learning Guide

Build, train and deploy models on Azure with Azure Machine Learning. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
15sections
36examples

This is part one of three. It covers everything you need to do real work with Azure Machine Learning: create a workspace, give yourself compute that costs nothing while idle, run a training script in the cloud, read the results, register the model, and put it behind a web endpoint that returns predictions. Mid-level and Senior take the same topics further. Nothing here gets thrown away.

Each section ends with a Try it task. Do them as you go. Cloud tooling is the kind of subject you only understand after you have watched a job fail with a real error message and fixed it yourself.

A note on the shape of the product before we start. Azure Machine Learning (Azure ML from now on) is a managed service. There is no server for you to install and no "Azure ML 4.2" to download. What you install is a client: the Azure CLI with its ml extension, or the Python package azure-ai-ml. The service itself lives in Microsoft's cloud. This guide was checked against the Python SDK azure-ai-ml 1.35.0 (released September 2026) and the current CLI v2, which is what the catalogue entry means by its version.

What Azure ML is, and the problem it solves

A machine learning model starts life as a script on a laptop. It reads a CSV, trains something, prints an accuracy, and the person who wrote it is happy. Then the real questions arrive. The dataset is now forty gigabytes and does not fit. Training takes six hours and the laptop is needed for other things. A colleague wants to reproduce the result and cannot, because nobody wrote down which library versions were used. The model has to be called by a web application, which means somebody has to wrap it in a server, package it, host it, secure it, and keep it running.

Azure ML exists to take those problems off your hands. It gives you rented computers on demand, a place to keep versioned data and models, a record of every training run, and a way to publish a trained model as an HTTPS endpoint. It does this by organising everything around a single top-level resource called a workspace, and by letting you describe your work in small YAML files or short Python programs.

DATAa versioned reference
→
JOByour script on rented compute
→
MODELregistered, versioned
→
ENDPOINTa URL that returns predictions

That diagram is the spine of the whole guide. Every section below is one arrow in it, or the plumbing that makes an arrow possible.

What came before? Teams used to rent a virtual machine, install Python on it by hand, copy files over with scp, and leave the machine running over the weekend because nobody remembered to switch it off. Some wrote their own scheduling scripts. Azure ML replaces that with managed compute that scales up when you submit work and back down when it finishes, and with a record of what ran, with which code, on which data, producing which result.

There is also an older generation of Azure ML itself that you will meet in search results. Azure ML has two client generations, called v1 and v2. The v1 command-line extension (azure-cli-ml) lost support on 30 September 2025, and the v1 Python SDK (azureml-core and its relatives) lost support on 30 June 2026. Microsoft says existing v1 workflows continue to operate, but they may be exposed to security risks or breaking changes in architecture. Everything in this guide uses v2. If a tutorial you find imports from azureml.core import Workspace, it is v1: close the tab. The v1 and v2 labels describe the client, not the service, so there is no workspace upgrade to perform. One workspace can serve both.

Do not mix v1 and v2 in one script Class names such as Workspace and Model exist in both generations and clash with each other. A script that imports from azureml.core and from azure.ai.ml together will fail in confusing ways. Pick v2 and stay there. The v2 package is imported as azure.ai.ml; v1 was azureml.core.

Who uses Azure ML? Anyone whose employer already lives in the Microsoft cloud: banks, telecoms, government bodies and large enterprises across the Gulf and Egypt often standardise on Azure, and for them the practical questions are where the data may live and who may see it. Azure ML workspaces are created in a specific Azure region, and you choose that region when you create the workspace. Pick a region that satisfies your employer's data-residency rules before you create anything, because the workspace and its storage live there.

Try it
  1. Think of one model you have trained, or would like to. Write down where the data lives, where training ran, and how anyone else could call the finished model.
  2. Mark each answer as "my laptop" or "somewhere shared and repeatable".
  3. Count the "my laptop" answers.
a list dominated by "my laptop". Each of those lines is something Azure ML moves into a shared, recorded, repeatable place, which is the point of the next sections.

The mental model: six nouns

Azure ML has a long list of features, but a beginner needs only a handful of nouns. Learn these and the documentation becomes readable.

Subscription and resource group. These are Azure's general scoping tools, not Azure ML ideas. A subscription is the billing container. A resource group is a folder inside it that holds related resources so you can see them together and delete them together. Every command you run needs to know the subscription, the resource group, and the workspace.

Workspace. The top-level Azure ML resource. It holds your jobs, models, data references, environments, compute and endpoints. When you create one, Azure also creates or links four supporting resources: a storage account (where your files and logs go), a Key Vault (where secrets are kept), a container registry (where Docker images are stored), and Application Insights (monitoring). You rarely touch those directly at first, but you will see them in your resource group and wonder what they are. Now you know.

Compute. The machines that run your work. There are four kinds you should recognise. A compute instance is one virtual machine used as a personal development workstation. A compute cluster (the type is called AmlCompute) is a pool that scales between a minimum and a maximum number of machines. Serverless compute means you do not create anything: you submit a job without naming any compute and Azure allocates it for you. Kubernetes compute lets you attach an existing Kubernetes cluster, which is an advanced topic.

Data asset. A named, versioned reference to data. It does not copy the data; it records where the data is. The three types are uri_file (one file), uri_folder (a folder) and mltable (a table described by a small definition file). A datastore is the related idea of a saved connection to a storage service such as Blob storage.

Environment. A Docker image plus a conda file describing which Python packages to install. It answers "what software does this job need?". Microsoft maintains curated environments whose names start with AzureML-, and you can build your own.

Job. One unit of execution: "run this command, with this code, on this environment, on this compute, with these inputs." The common kind is a command job. Other kinds exist (sweep for hyperparameter search, pipeline for chained steps, Spark, and AutoML), but they are all variations on the same idea. Jobs are grouped under an experiment name, which is just a label that keeps related runs together in the studio.

After training, two more nouns appear. A model is a registered, versioned artifact: the thing your job produced. An endpoint is a stable URL and an authentication setting, and a deployment is the model, code and environment actually running behind it. One endpoint can hold several deployments with traffic split between them. We return to this in the deployment section.

Subscription
billing
→
Resource group
a folder
→
Workspace
jobs, models, data, compute, endpoints
Storage
Key Vault
Container registry
Application Insights
created alongside the workspace

There is one more habit of mind worth adopting now. Azure ML is YAML-first. Almost every entity (a job, a compute cluster, an endpoint) can be written as a YAML file, checked into git, and created with one command. The Python SDK does the same things in code. The CLI and the SDK are two front doors to the same house, so you can learn the concepts once and switch between them. This guide shows both, because you will meet both in the wild.

Try it
  1. Without looking back, write the six nouns in the order a training request meets them: where it is billed, where it is grouped, where it is managed, what runs it, what it needs installed, what it reads.
  2. Then check yourself against this section.
subscription, resource group, workspace, compute, environment, data. If you can say it in one breath, the rest of the guide is detail.

Installing the tools and signing in

There is nothing to install for the service. You install a client. This guide uses both clients, because the CLI is best for creating infrastructure and running YAML jobs, while the Python SDK is best when your logic needs loops or conditions.

You also need an Azure account with a subscription where you are allowed to create resources. If you work for a company, ask your administrator which subscription to use. If you are studying alone, Microsoft offers a free account for new users; check the current terms on the Azure website, since offers change.

The Azure CLI and its ml extension. The Azure CLI is the az command. Azure ML adds its commands through an extension named ml. The official page says you need Azure CLI version 2.38.0 or newer. On Debian or Ubuntu, Microsoft's documented install is:

BASH
curl -sL https://aka.ms/InstallAzureCLIDeb | sudo bash
az extension add -n ml -y

On macOS, Homebrew is the usual route, and on Windows the winget package manager or Microsoft's MSI installer works. Check the Azure CLI install page for the exact current instructions for your platform, then add the extension:

BASH
brew update && brew install azure-cli
az extension add -n ml

Check that everything is in place:

BASH
az version
az ml -h

az version prints the CLI version and the extensions installed; confirm that ml appears. az ml -h prints the list of Azure ML command groups, and seeing job, model, compute and online-endpoint among them tells you the extension loaded. Upgrade it later with az extension update -n ml.

If you ever used the old v1 extension, remove it first, because it uses the same az ml prefix and the two conflict:

BASH
az extension list
az extension remove -n azure-cli-ml
az extension remove -n ml
az extension add -n ml

The Python SDK. Install the package that provides the azure.ai.ml module, plus the identity library that handles sign-in:

BASH
pip install azure-ai-ml azure-identity

Do this inside a virtual environment (python -m venv .venv, then activate it), so that these packages do not mingle with other projects. Microsoft links its install instructions at https://aka.ms/sdk-v2-install; check them for the Python versions currently supported.

Signing in. Both clients use your Azure identity. Sign in once from a terminal:

BASH
az login
az account set -s "<SUBSCRIPTION_NAME_OR_ID>"

az login opens a browser window. az account set chooses which subscription later commands target, which matters if your account can see more than one. The Python SDK then reuses that sign-in through a class called DefaultAzureCredential, which tries several sources of credentials in turn, including your az login session. In automated systems the same class picks up a service principal from environment variables, which is how a CI pipeline signs in without a person.

Behind a corporate TLS-inspecting proxy? If your laptop sits behind a security gateway that inspects encrypted traffic (Cloudflare Gateway is one example), az login may fail with CERTIFICATE_VERIFY_FAILED while talking to login.microsoftonline.com. The correct fix is to export your organisation's root certificate and tell the tool to trust it. Do not turn certificate verification off to make the error go away. This is local experience rather than Microsoft documentation, so ask your IT team which certificate to trust.
Try it
  1. Install the CLI and the extension, then run az version and az ml -h.
  2. Run az login and complete the browser sign-in.
  3. Run az account show and confirm the subscription named is the one you intend to use.
a JSON block with your subscription name and id. If it names the wrong subscription, run az account set -s now, before you create anything that costs money.

Creating a workspace and setting defaults

The workspace is the home for everything else, so we create it first. Two environment variables keep the commands short. Set them in your shell (the docs use Bash syntax; PowerShell users must adapt the variable lines):

BASH
GROUP=mlops-mena-rg
LOCATION=uaenorth
WORKSPACE=mlops-mena-ws

The location is the Azure region. uaenorth is a Middle East region, used here as an example; list the regions your subscription can use with az account list-locations -o table and choose the one that matches your data-residency requirements. Not every service or machine size is available in every region, so if a later step complains about capacity, a different region is a legitimate remedy.

Create the resource group and the workspace:

BASH
az group create -n $GROUP -l $LOCATION
az ml workspace create -n $WORKSPACE -g $GROUP -l $LOCATION

The workspace command takes a minute or two, because it creates the storage account, Key Vault, container registry and Application Insights alongside it. When it finishes it prints a JSON description. Now save defaults so you stop typing the group and workspace on every command:

BASH
az configure --defaults group=$GROUP workspace=$WORKSPACE location=$LOCATION
az configure -l -o table

Every az ml command needs a workspace and a resource group, and these defaults supply them. The second command shows what is currently set, which is the first thing to check when a command complains it does not know your workspace.

Open the studio, the web interface, at https://ml.azure.com, sign in, and choose your workspace. You will use it to watch jobs, view metrics and browse assets. Everything it shows is the same data the CLI sees; the studio is a window, not a separate system.

From Python, connect with a client object. Every SDK call goes through it:

PYTHON
from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential

subscription_id = "<SUBSCRIPTION_ID>"
resource_group = "mlops-mena-rg"
workspace = "mlops-mena-ws"

ml_client = MLClient(DefaultAzureCredential(), subscription_id, resource_group, workspace)
print(ml_client.workspace_name)

If this prints your workspace name, authentication, subscription and workspace are all correct. If it raises DefaultAzureCredential failed, you have not run az login (or have not set service principal variables). If it raises ImportError: No module named 'azure.identity', run pip install azure-identity.

Keep the ids out of your code Put the subscription id, group and workspace in environment variables, or in a small config.json that is listed in .gitignore, and read them at startup. Hard-coded ids in a public repository are not secrets as such, but they tell strangers exactly which resources to probe.
Try it
  1. Create the resource group and workspace with the commands above, using your own names.
  2. Run az configure -l -o table and confirm the defaults.
  3. Open https://ml.azure.com and find your workspace. Then run the Python snippet and confirm it prints the same name.
the same workspace name in the terminal, the studio and Python. That is three front doors to one house.

Compute: renting machines and not paying for idle ones

Your laptop is not where training should happen. In Azure ML you rent machines in two shapes, and the difference matters both for convenience and for your bill.

A compute cluster is the workhorse for jobs. You define a name, a virtual machine size, a minimum and a maximum number of nodes. When you submit a job, the cluster scales up from the minimum to run it, and when the work is done, it scales back down after an idle timeout. Set the minimum to 0 and you hold no running machines between jobs. Create one with the CLI:

BASH
az ml compute create -n cpu-cluster --type amlcompute --min-instances 0 --max-instances 4

The same cluster from Python:

PYTHON
from azure.ai.ml.entities import AmlCompute

compute = AmlCompute(
    name="cpu-cluster",
    size="STANDARD_D2_V2",
    min_instances=0,
    max_instances=4,
)
ml_client.compute.begin_create_or_update(compute).result()

Two details deserve explanation. First, the begin_ prefix on SDK methods marks a long-running operation: the call returns immediately with a poller object, and .result() waits for it to finish. You will see this pattern everywhere in the SDK. Second, size names an Azure virtual machine type. The right choice depends on your workload and on what your subscription and region have available. If you pick a size you have no quota for, you will meet QuotaExceeded, covered in the errors section.

Serverless compute is the zero-setup alternative. If you omit the compute setting from a job, Azure allocates compute for that job automatically. For a beginner's first jobs this is attractive, because there is nothing to create and nothing to forget about. The cluster is better when you want control over the machine size, want to reuse the same pool repeatedly, or want to cap the number of nodes.

A compute instance is a single VM meant to be your cloud workstation, with notebooks and a terminal available through the studio. It is convenient for exploration, but it has a property that catches beginners: it is a machine that stays on until you stop it, and a running machine is billed. Treat it like a laptop you must remember to close.

Cost discipline from day one The fastest way to a surprising Azure bill is a forgotten machine. Use a minimum of 0 on clusters, stop compute instances when you finish for the day, and delete what you no longer need. The documentation's own clean-up section recommends deleting the cluster when you are done with a tutorial, and says a cluster continues to bill while it exists. Check the Azure ML pricing page for how that applies to your setup, since this guide does not quote prices.

Inspect and remove compute with the same grammar as everything else. The CLI is always az ml <noun> <verb> <options>:

BASH
az ml compute list -o table
az ml compute show --name cpu-cluster
az ml compute delete -n cpu-cluster --yes
Try it
  1. Create cpu-cluster with a minimum of 0 and a maximum of 2 nodes.
  2. Run az ml compute list -o table and find it.
  3. Open the studio, go to Compute, and look at the cluster's node count.
a cluster with zero running nodes. It exists, it is ready, and nothing is running on it until a job arrives.

Your first training job

We now run a script in the cloud. A command job has five ingredients: the code to upload, the command to run, the environment, the compute, and the inputs. Start with a project folder:

TEXT
azureml-first/
  src/
    main.py
  job.yml

Here is a small training script. It reads a CSV of the classic iris flower dataset, trains a LightGBM classifier, and records what it did with MLflow, the open-source tracking library that Azure ML recommends for logging:

src/main.py
import argparse

import lightgbm as lgb
import mlflow
import mlflow.lightgbm
import pandas as pd
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--iris-csv", type=str, required=True)
    parser.add_argument("--learning-rate", type=float, default=0.1)
    args = parser.parse_args()

    df = pd.read_csv(args.iris_csv)
    target = df.columns[-1]
    X = df.drop(columns=[target])
    y = df[target].astype("category").cat.codes

    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)

    model = lgb.LGBMClassifier(learning_rate=args.learning_rate, n_estimators=50)
    model.fit(X_train, y_train)

    accuracy = accuracy_score(y_test, model.predict(X_test))
    print(f"accuracy={accuracy:.3f}")

    mlflow.log_param("learning_rate", args.learning_rate)
    mlflow.log_metric("accuracy", accuracy)
    mlflow.lightgbm.log_model(model, "model")


if __name__ == "__main__":
    main()

Read the script as an ordinary Python program, because that is what it is. Azure ML does not require a special framework; your script only has to be runnable from a command line and read its settings from arguments. The only Azure-aware parts are the mlflow calls, and even those work on your laptop.

Now the job definition. This is the YAML that tells Azure ML what to do:

job.yml
$schema: https://azuremlschemas.azureedge.net/latest/commandJob.schema.json
code: src
command: >-
  python main.py
  --iris-csv ${{inputs.iris_csv}}
  --learning-rate ${{inputs.learning_rate}}
inputs:
  iris_csv:
    type: uri_file
    path: https://azuremlexamples.blob.core.windows.net/datasets/iris.csv
  learning_rate: 0.9
environment: azureml:AzureML-lightgbm-3.3@latest
compute: azureml:cpu-cluster
display_name: lightgbm-iris-first-run
experiment_name: lightgbm-iris-first-run

Go through it line by line. $schema points to a published definition of the file format, which lets editors such as VS Code validate and autocomplete it. code: src says which local folder to upload. command is what runs inside the machine, and the ${{inputs.iris_csv}} placeholders are replaced with the actual input values at run time. inputs declares two inputs: a data file and a plain number. Note how the file input is given a type (uri_file) and a location, while the number is just written down. environment names a curated environment. The @latest suffix means "the newest version of that environment". compute names the cluster from the previous section; delete that line to use serverless compute instead.

Curated environment names rotate The documentation's own CLI and SDK examples use different curated environment names, and Microsoft retires old ones. If your job fails with EnvironmentNotFound, list what exists with ml_client.environments.list() in Python or az ml environment list on the command line, and pick a current one. Also check that the environment you choose contains mlflow; if your script fails on import mlflow, you need an environment that does, or your own.

Submit the job:

BASH
run_id=$(az ml job create -f job.yml --query name -o tsv)
az ml job show -n $run_id --web
az ml job stream -n $run_id

The first line submits the job and captures its generated name in a shell variable, using --query name -o tsv to extract just that field. The second opens the job's page in the studio. The third streams the job's log to your terminal until it finishes. The job's status moves through Starting, Preparing, Running and finally Completed. The first run is slow, since Azure may need to build the environment image and the cluster has to scale from zero. Later runs reuse the cached image and are much faster. Check the status at any time with:

BASH
az ml job show -n $run_id --query status -o tsv

The same job through the Python SDK reads like this:

PYTHON
from azure.ai.ml import command, Input

job = command(
    code="./src",
    command="python main.py --iris-csv ${{inputs.iris_csv}} --learning-rate ${{inputs.learning_rate}}",
    environment="AzureML-lightgbm-3.3@latest",
    inputs={
        "iris_csv": Input(
            type="uri_file",
            path="https://azuremlexamples.blob.core.windows.net/datasets/iris.csv",
        ),
        "learning_rate": 0.9,
    },
    compute="cpu-cluster",
    display_name="lightgbm-iris-first-run",
    experiment_name="lightgbm-iris-first-run",
)

returned_job = ml_client.jobs.create_or_update(job)
print(returned_job.studio_url)
ml_client.jobs.stream(returned_job.name)

create_or_update submits the job and returns right away; studio_url is a link to its page; and stream follows the logs. Notice that the YAML and the Python are the same job: code, command, environment, inputs and compute appear in both.

Try it
  1. Create the folder and the two files above. If the curated environment name fails, list environments and substitute one.
  2. Submit with az ml job create -f job.yml and stream the logs.
  3. When it finishes, open the studio link and look at the Overview tab.
status Completed, the experiment name you chose, and your command echoed back. The first run takes minutes; if it seems stuck on Preparing, that is the image being built, not a hang.

Reading results: logs, metrics and outputs

A job that ran is only useful if you can see what it did. Azure ML records three kinds of evidence, and knowing where each lives will save you hours.

The studio job page is the main view. The Overview tab shows status, timing, inputs, the compute used and the command. The Metrics tab shows the values you logged with mlflow.log_metric, and for repeated values over time it draws a chart. The Outputs + logs tab holds the files: the standard output of your script (what print wrote) is in a log file there, alongside system logs from the environment preparation. The Code tab shows the exact snapshot of your src folder that was uploaded, which is how you answer "what code produced this result?" months later.

When a job fails, read logs in this order. First the user log, the one that contains your script's output, because most failures are your own code: a missing import, a wrong file path, a typo in an argument. Only if your script never started do you look at the system logs, which describe the environment build and container start. A job that dies in Preparing almost always has an environment problem; one that dies in Running almost always has a script problem. That single distinction solves a large share of beginner failures.

This is also the place to understand why logging with MLflow matters. A printed accuracy is gone once you close the terminal. A logged metric is attached to the job, stored in the workspace, and comparable with other runs. Run the job a second time with a different learning_rate, and the studio lets you tick both runs and compare their metrics side by side. That comparison is the daily work of model development.

From the command line, the same information is available without opening a browser:

BASH
az ml job list -o table
az ml job show -n $run_id
az ml job stream -n $run_id

job list shows recent jobs with their status, job show prints the full definition and state, and stream replays the log. You can also download everything a job produced with az ml job download -n $run_id, which is useful when you want the model files locally.

Name your runs Set display_name and experiment_name in every job. Without them the studio shows machine-generated names, and a list of forty runs called things like jolly_basket_x7q2 is unreadable. A name such as iris-lr0.9-baseline costs five seconds and pays back the first time you compare runs.
Try it
  1. Change learning_rate in job.yml to 0.05 and change display_name to something that mentions it. Submit again.
  2. In the studio, open the experiment, select both runs and choose to compare them.
  3. Find where your print output appears under Outputs + logs.
two runs in one experiment, each with its own accuracy metric and parameter. This side-by-side view is what MLflow logging buys you.

Data: assets, datastores and inputs

So far the data was a public URL, which is convenient for a tutorial and unrealistic for work. Real data lives in your own storage, and Azure ML lets you refer to it by name and version rather than by a path someone might change.

A data asset is that named reference. Recall the three types: uri_file for a single file, uri_folder for a directory, mltable for tabular data described by a definition file. The word URI tells you that the asset stores a location, not a copy. Creating a data asset from a local file also uploads that file to the workspace's default storage, so the asset has somewhere stable to point.

A YAML file describes an asset:

data.yml
$schema: https://azuremlschemas.azureedge.net/latest/data.schema.json
name: iris-data
version: 1
type: uri_file
path: ./data/iris.csv
description: Iris flower measurements used in the first-job tutorial.
BASH
az ml data create -f data.yml
az ml data list -o table

The Python equivalent uses the Data entity and AssetTypes constants:

PYTHON
from azure.ai.ml.entities import Data
from azure.ai.ml.constants import AssetTypes

iris = Data(
    name="iris-data",
    version="1",
    type=AssetTypes.URI_FILE,
    path="./data/iris.csv",
    description="Iris flower measurements used in the first-job tutorial.",
)
ml_client.data.create_or_update(iris)

Once registered, a job refers to it as azureml:iris-data:1 in place of the URL. The path in the job YAML becomes:

YAML
inputs:
  iris_csv:
    type: uri_file
    path: azureml:iris-data:1

Why bother? Because versions make experiments reproducible. If somebody replaces the CSV next week, version 1 still points at what you trained on, and a new file becomes version 2. When a model's accuracy drops, you can ask which data version it was trained on and get an answer. A bare path cannot tell you that. If you already know DVC, the idea will feel familiar; see DVC for versioning data with git.

A datastore is the saved connection that makes this possible. The workspace creates a default one for its storage account; you can register more to point at other Blob containers or file shares, and the credentials stay in the workspace rather than in your scripts. One limitation to know about: database datastores are not supported in v2, so data in a database has to be exported to object storage such as Blob first.

What is mltable? An mltable asset is a folder containing your data plus a small MLTable definition file that says how to read it: which files, what delimiter, which columns and types. It is useful when "read this CSV" needs more detail than a single path can carry. As a beginner you can do most things with uri_file and uri_folder.
Try it
  1. Download iris.csv into a local data folder and register it as shown.
  2. Run az ml data list -o table, then open Data in the studio and find version 1.
  3. Edit job.yml to use azureml:iris-data:1 and run it again.
the job succeeds exactly as before, but its inputs now show a named, versioned data asset rather than a web address.

Registering the model

A finished job leaves behind a trained model, but it is buried in that job's outputs. Registering a model gives it a name and a version in the workspace, which is what later steps (deployment, sharing, comparison) refer to.

Because our script saved the model with mlflow.lightgbm.log_model(model, "model"), the job's output contains an MLflow-format model at the path model. Azure ML supports three model types: mlflow_model, custom_model and triton_model. MLflow models are the easy path, since MLflow records how to load and call the model. Register from the job on the command line:

BASH
az ml model create -n sklearn-iris-example -v 1 -p runs:/$run_id/model --type mlflow_model

Here -n is the model name, -v the version, -p the path (runs:/<job name>/model means "the folder named model inside that job's output"), and --type says it is an MLflow model. In Python the same step is:

PYTHON
from azure.ai.ml.entities import Model
from azure.ai.ml.constants import AssetTypes

model = Model(
    path=f"azureml://jobs/{returned_job.name}/outputs/artifacts/paths/model/",
    name="run-model-example",
    type=AssetTypes.MLFLOW_MODEL,
)
ml_client.models.create_or_update(model)

Check that it registered:

BASH
az ml model list -o table

The mental picture is a library shelf. Each registered model has a name; each registration adds a version; and the studio's Models page lists them with the job that produced each. When you deploy, you say "name and version", never "that folder somewhere".

Try it
  1. Register the model from your successful run with az ml model create.
  2. Run az ml model list -o table.
  3. Open Models in the studio and follow the link from the model back to its job.
a model at version 1 whose page links back to the run that produced it. That chain from model to job to code to data is the audit trail Azure ML exists to keep.

Deploying a model as an online endpoint

A model that nobody can call is a file. An online endpoint turns it into an HTTPS service that takes a request and returns a prediction in real time. A batch endpoint, a different feature, scores large datasets asynchronously; here we cover the real-time kind.

Two words must be clear. The endpoint is the stable front: a name, a URL and an authentication mode. The deployment is what actually runs behind it: a model, the code that loads and calls it, an environment, and a machine size with an instance count. An endpoint can hold more than one deployment, conventionally called blue and green, with traffic split between them by percentage. That is how you ship a new model version safely: add green, send it ten percent of traffic, and shift more as confidence grows.

A managed online endpoint means Azure runs the underlying infrastructure for you. You choose the instance type and count and do not manage servers.

The endpoint is a short YAML file:

endpoint.yml
$schema: https://azuremlschemas.azureedge.net/latest/managedOnlineEndpoint.schema.json
name: iris-endpoint-demo
auth_mode: key

auth_mode is key (a secret key you send with each request) or aml_token (a short-lived Microsoft Entra token). Keys are simplest for a first test.

The deployment YAML ties the pieces together. For a non-MLflow model you also supply a scoring script; for an MLflow model Azure ML can generate one, but you should still know what a scoring script looks like, because you will meet it in other people's projects. It has two functions:

score.py
import json
import os

import mlflow


def init():
    """Runs once when the container starts. Load the model here."""
    global model
    model_dir = os.getenv("AZUREML_MODEL_DIR")
    model = mlflow.pyfunc.load_model(os.path.join(model_dir, "model"))


def run(raw_data):
    """Runs for every request. raw_data is the request body as a string."""
    data = json.loads(raw_data)["data"]
    predictions = model.predict(data)
    return predictions.tolist()

init() runs once when the container starts, which is where you load the model so that you do not reload it for every request. run() runs for every call. If either raises an exception, callers see HTTP 502, as the errors section explains.

The deployment file references a registered model, the scoring script, an environment, and a machine size:

blue-deployment.yml
$schema: https://azuremlschemas.azureedge.net/latest/managedOnlineDeployment.schema.json
name: blue
endpoint_name: iris-endpoint-demo
model: azureml:sklearn-iris-example:1
code_configuration:
  code: ./onlinescoring
  scoring_script: score.py
environment: azureml:AzureML-lightgbm-3.3@latest
instance_type: Standard_DS3_v2
instance_count: 1

The environment must contain everything score.py imports, including the azureml-inference-server-http package that Azure ML uses to wrap your script in a web server. A curated environment aimed at training may not include it, so a deployment-specific environment is common. Check the deployment documentation for the environment you intend to use.

Test locally first. Deploying to the cloud can take many minutes, and a typo costs you the whole wait. Local deployment runs the same container on your machine and needs Docker installed (see Docker):

BASH
az ml online-endpoint create --local -n iris-endpoint-demo -f endpoint.yml
az ml online-deployment create --local -n blue --endpoint iris-endpoint-demo -f blue-deployment.yml
az ml online-endpoint invoke --local --name iris-endpoint-demo --request-file sample-request.json
az ml online-deployment get-logs --local -n blue --endpoint iris-endpoint-demo

Here sample-request.json is a file you write containing the input your run() function expects, for example {"data": [[5.1, 3.5, 1.4, 0.2]]}. The --local flag on each command keeps everything on your machine, and get-logs is the first place to look when something is wrong.

When local testing works, drop --local and deploy for real:

BASH
az ml online-endpoint create --name iris-endpoint-demo -f endpoint.yml
az ml online-deployment create --name blue --endpoint iris-endpoint-demo -f blue-deployment.yml --all-traffic
az ml online-endpoint show -n iris-endpoint-demo
az ml online-endpoint invoke --name iris-endpoint-demo --request-file sample-request.json

--all-traffic sends one hundred percent of requests to the new deployment; without it the deployment exists but receives none. To call the endpoint from outside, fetch its address and key:

BASH
ENDPOINT_KEY=$(az ml online-endpoint get-credentials -n iris-endpoint-demo -o tsv --query primaryKey)
SCORING_URI=$(az ml online-endpoint show -n iris-endpoint-demo -o tsv --query scoring_uri)
curl -X POST "$SCORING_URI" \
  -H "Authorization: Bearer $ENDPOINT_KEY" \
  -H "Content-Type: application/json" \
  -d @sample-request.json

The SDK version of the endpoint and deployment looks like this:

PYTHON
from azure.ai.ml.entities import ManagedOnlineEndpoint, ManagedOnlineDeployment, CodeConfiguration

endpoint = ManagedOnlineEndpoint(name="iris-endpoint-demo", auth_mode="key")

blue = ManagedOnlineDeployment(
    name="blue",
    endpoint_name="iris-endpoint-demo",
    model="azureml:sklearn-iris-example:1",
    code_configuration=CodeConfiguration(code="./onlinescoring/", scoring_script="score.py"),
    environment="AzureML-lightgbm-3.3@latest",
    instance_type="Standard_DS3_v2",
    instance_count=1,
)

ml_client.begin_create_or_update(endpoint).result()
ml_client.begin_create_or_update(blue).result()

endpoint.traffic = {"blue": 100}
ml_client.begin_create_or_update(endpoint).result()
An endpoint bills while it runs A deployment holds real virtual machines for as long as it exists, whether or not anyone calls it. Delete endpoints you are finished with: az ml online-endpoint delete --name iris-endpoint-demo --yes --no-wait. Make it the last step of every tutorial. Confirm the billing details on the pricing page.
Try it
  1. Write score.py and sample-request.json, then deploy locally with the four --local commands.
  2. Invoke the endpoint and read the response, then read get-logs.
  3. Only when local works, deploy to the cloud, invoke it once, and then delete the endpoint.
a list of predictions returned for your sample request, first from your machine and then from the cloud. Deleting the endpoint at the end is part of the exercise, not an afterthought.

The everyday commands, grouped by what you are trying to do

The CLI has one regular grammar: az ml <noun> <verb> <options>. Learn the verbs once and every noun behaves the same way. Here is a working reference grouped by intent.

Run something. az ml job create -f job.yml submits a job. az ml job stream -n <name> follows its log. az ml job show -n <name> --web opens it in the studio. az ml job cancel -n <name> stops a running job.

Create or change infrastructure. az ml compute create, az ml data create, az ml environment create, az ml model create and az ml online-endpoint create all take -f file.yml, so the file in git is the source of truth. Changing a thing is usually update with a modified file, as in az ml online-deployment update -n blue --endpoint <name> -f blue-deployment.yml.

Look at what exists. list and show exist for every noun. Add -o table to list for a readable summary, and --query with a JMESPath expression to pull one field out of the JSON, as in --query name -o tsv.

Export an existing thing as YAML. az ml <noun> show --output yaml prints the definition of something already created, which is a handy way to learn the schema. Delete the system-generated properties (ids, timestamps) before reusing it as a file.

Clean up. az ml compute delete -n <name> --yes, az ml online-endpoint delete --name <name> --yes --no-wait. The --yes flag skips the confirmation prompt, and --no-wait returns immediately rather than waiting for deletion to finish.

In the SDK, the same nouns appear as attributes of ml_client: ml_client.jobs, ml_client.compute, ml_client.data, ml_client.environments, ml_client.models. Each offers list, get and create_or_update, and the ones that take time use the begin_ prefix with .result(). For example:

PYTHON
for job in ml_client.jobs.list():
    print(job.name, job.display_name, job.status)

for env in ml_client.environments.list():
    print(env.name)

The second loop is the one to run when EnvironmentNotFound appears. A good habit is to explore with list and show, because they cost nothing and teach the vocabulary.

When should you choose the CLI and when the SDK? Microsoft's own guidance is that the CLI is well suited to automation and repeatable pipelines, and the SDK suits programmatic work with logic in it. In practice, many teams write YAML, run it with the CLI from a CI workflow (see GitHub Actions), and reach for Python when they need loops, conditions or tight integration with notebooks.

Try it
  1. Run az ml job list -o table, then az ml model list -o table, then az ml compute list -o table.
  2. Pick one job and run az ml job show -n <name> --output yaml to see its full definition.
  3. Use --query name -o tsv on one of the lists to print only names.
three tidy tables and one full YAML definition. Seeing a real job as YAML is the quickest way to learn which fields you can set.

Configuration and the errors you will actually meet

Most of Azure ML's configuration is the YAML you already wrote, but a few settings live elsewhere. Defaults (group, workspace, location) live in the Azure CLI configuration and are managed with az configure. Credentials come from az login or from service principal environment variables. Project settings such as the environment and compute belong in the job file, and that is where you should keep them.

Errors fall into predictable families. Here are the ones beginners meet, with how to read each.

Sign-in and setup errors.

  • ImportError: No module named 'azure.identity': the identity package is missing. Run pip install azure-identity.
  • DefaultAzureCredential failed: the SDK found no credentials. Run az login, or set service principal variables in an automated environment.
  • ComputeNotFound: the compute name in your job does not match a cluster that exists, or the cluster was deleted. Run az ml compute list -o table and compare the name.
  • EnvironmentNotFound: the curated environment was retired or misspelled. List with ml_client.environments.list().
  • QuotaExceeded: your subscription does not have enough virtual CPU quota for the machine size and node count. Request a quota increase or choose a smaller size.
  • No such file or directory: endpoint.yml: you ran the command from the wrong folder. Change into the folder that contains the YAML.

Deployment errors. These come from the online endpoint troubleshooting page and appear as ERROR: <Code>.

  • ImageBuildFailure: Azure could not build the container image. The build log is in the workspace's default storage. A message mentioning container registry authorization failure is fixed with az ml workspace sync-keys.
  • OutOfQuota: you ran out of something: CPU quota, disk, memory, endpoint count, or the region simply lacks capacity for that VM size. Choose another size or region, or request quota.
  • BadArgument: a family of causes, including a role assignment problem at startup or a model that cannot be downloaded. Read the full message; it names the cause.
  • ResourceNotReady: the container crashed while running your score.py. Typical causes are a missing import, a syntax error, or an exception in init(). Use az ml online-deployment get-logs, or reproduce it with a local deployment.
  • ResourceNotFound, OperationCanceled and InternalServerError: look at the detail in the message. A canceled operation usually means another operation took priority.

HTTP status codes when you call the endpoint.

  • 409: another operation is in progress on the endpoint, for example you tried to delete while it was still updating. Wait and retry.
  • 502: an exception or crash inside run() or init() of your scoring script. Read the logs.
  • 503: a spike in requests that the instances could not absorb. Add instances.
  • 504: the request timed out.
  • 500: a failure on Azure's side.

A mistake worth naming: when a job fails, read the first error in the log, not the last. Later messages are usually consequences. The line where the traceback begins is the real cause.

Private networks change the rules If your workspace sits behind a private endpoint, a legacy-mode setting called v1_legacy_mode is turned on automatically and blocks v2 APIs and managed online endpoints until it is set to false. Check with your network team first. This is rare for a personal learning workspace, and common in corporate ones, so know the term when you meet it.
Try it
  1. Break your job on purpose: set compute: azureml:no-such-cluster in job.yml and submit it.
  2. Read the error text and match it to the list above.
  3. Fix it and resubmit.
an error that names the missing compute, which you can fix in one line. Seeing a failure on purpose makes the real ones less frightening.

Putting it all together

Here is one small end-to-end project that uses every piece. Treat it as a script for a rehearsal, and do it once from an empty directory.

  1. Log in and pick the subscription. az login, then az account set -s "<SUBSCRIPTION>".
  2. Create the group and workspace. az group create, then az ml workspace create, then az configure --defaults.
  3. Create a cluster that scales to zero. az ml compute create -n cpu-cluster --type amlcompute --min-instances 0 --max-instances 2.
  4. Register the data. Write data.yml and run az ml data create -f data.yml.
  5. Write the training script and job file. Put main.py in src, write job.yml that uses azureml:iris-data:1, and set display_name and experiment_name.
  6. Run it. az ml job create -f job.yml, stream the logs, and read the metrics in the studio.
  7. Compare. Change the learning rate, run again, and compare the two runs.
  8. Register the best model. az ml model create -n iris-model -v 1 -p runs:/<run>/model --type mlflow_model.
  9. Deploy locally, then to the cloud. Test with --local first, then deploy with --all-traffic, and call the endpoint with curl.
  10. Clean up. Delete the endpoint, delete the cluster, and delete the resource group if this was only practice.

That last step deserves a command of its own, because deleting the resource group removes everything inside it, including the workspace and its supporting storage, in one go:

BASH
az group delete -n $GROUP --yes --no-wait

Do this only when you are sure. It is irreversible, and it is the cleanest way to be certain nothing is left running and billing.

What did you just build? A full loop: data in, training on rented machines, results recorded, a model with a name and version, and a live endpoint. Every production machine learning platform is a larger version of that loop with more safeguards: tests, approvals, scheduled retraining, monitoring and access control. The next level of this guide covers those.

To keep the project reproducible, commit the YAML files, the scripts and a requirements file to git, and write the commands in a short README. A teammate with an Azure subscription can then recreate your result from scratch, which is the real test of whether you have understood the system.

Try it
  1. Run the ten steps in order from a clean folder, timing each one.
  2. Note which step took longest and which produced the most confusion.
  3. Delete everything with the clean-up commands.
a log of your own session. The longest step is usually the first job (image build and cluster scale-up) or the cloud deployment. Knowing where time goes is the beginning of knowing how to speed it up.

What you can now do, and what comes next

By now you can explain what a workspace, a compute cluster, an environment, a data asset, a job, a model and an endpoint are, and how they connect. You can install the CLI and SDK, sign in, and create a workspace with sensible defaults. You can run a training script in the cloud, read its logs and metrics, register the resulting model, test a deployment locally and publish it, and you know how to read the common errors and how to turn everything off to stop the meter.

Some things were deliberately left for later. Pipelines chain several jobs into a graph, and components are reusable steps within them. Sweep jobs search over hyperparameters. Registries share models and environments across workspaces, so a model proven in a development workspace can be promoted to production without being rebuilt. Autoscaling, network isolation, private endpoints, managed identity, role-based access, monitoring, cost controls and infrastructure as code are the topics that turn a working demo into something a team can rely on. The Mid-level guide starts there.

Neighbouring guides in the catalogue fit naturally around this one. MLflow explains the tracking and model-format layer that Azure ML builds on. Docker makes environments and local deployment testing much clearer. Terraform is one way to create workspaces and compute as code, and Kubernetes is where Kubernetes compute and endpoints lead. If you are working with large language models rather than classical ones, look at Azure OpenAI, which is a separate Azure service with its own model.

Finally, keep the migration lesson in mind. Cloud products change names, retire environments and replace whole client generations, as v1 to v2 showed. Pin the version of azure-ai-ml in a requirements file, re-test when you upgrade, and always check the date on a tutorial before copying commands from it.

Try it
  1. Pick one idea from "left for later" and read its page in the official documentation for twenty minutes.
  2. Write three sentences about what problem it solves, in your own words.
  3. Decide whether your own project needs it yet.
a short note you can reuse in an interview or a design discussion. Knowing what you do not need yet is part of using a large platform well.

Sources