تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
Azure OpenAILLMsModel platforms & hubs3 مستويات96 قسمًايغطّي Azure OpenAI v1 APIدليل بالإنجليزية

The Complete Azure OpenAI Guide

Run OpenAI models in Azure: deployments, quotas, private networking and enterprise controls. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
16sections
31examples

This is part one of three. It covers everything you need to do real work with Azure OpenAI as a first-time user, not a teaser. By the end you can create the Azure resources that host a model, deploy a model under a name of your choosing, call it from curl and from Python, sign in without a password in your code, stream an answer, hold a conversation, turn text into embeddings, and read the errors the service sends back. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task where it makes sense. Do them as you go. They take a few minutes each, and the service only starts to feel predictable once you have watched your own deployment answer, fail, and answer again.

One honest note before we start. Azure OpenAI is a managed cloud service, not software you install, so "the version" does not mean what it does for a command-line tool. This guide teaches the v1 API (generally available since August 2025), the openai Python package 3.x, and Azure CLI 2.90. Where Microsoft's own pages still show an older pattern, we say so and show the current form.

What Azure OpenAI is, and why people choose it

Azure OpenAI is Microsoft's hosting of OpenAI's models (the GPT family, the embedding models, image and audio models) inside your own Azure subscription. You send text to an HTTPS endpoint that belongs to you, a model reads it, and you get text back. Microsoft's current name for it is "Azure OpenAI in Microsoft Foundry Models". People still say "Azure OpenAI", and so does this guide.

If you have called the OpenAI API directly, the experience is almost the same. The difference is everything around the model. Your requests go to an endpoint that sits in your Azure subscription, you pay on your Azure bill, you control who may call it with the same identity system you use for the rest of Azure, and Microsoft's data commitments apply. Those commitments, in Microsoft's own words, are that your prompts, completions, embeddings and training data are not available to other customers or to OpenAI, and are not used to train the foundation models. For an employer in the Gulf or Egypt that has a security team, a procurement process and a rule about where data may live, that wrapper is often the reason Azure OpenAI is chosen over the public API.

YOUR APPPython, curl, .NET
→
YOUR ENDPOINTresource-name.openai.azure.com
→
YOUR DEPLOYMENTa named model
→
ANSWERtext, vectors, images

The diagram is the whole idea, and the rest of this guide fills in each box. What you can do once the path works:

💬

Generate and transform text

Summaries, drafts, classification, extraction, translation, and answers to questions, from one API call.

🧭

Embed text as vectors

Turn sentences into lists of numbers so that you can search by meaning, the first half of retrieval-augmented generation.

🔐

Keep access under control

Use Azure roles instead of shared passwords, so you can say exactly who and what may call the model.

📍

Choose where it runs

Pick global, data-zone or regional processing depending on your residency rules.

You need little to follow along: an Azure subscription you are allowed to create resources in, a terminal, and Python 3.10 or newer. If your employer manages your subscription, ask for permission to create a resource group and an Azure OpenAI resource in it; the rest of the guide depends on that.

Try it
  1. Open https://portal.azure.com and sign in.
  2. Find your subscription under Subscriptions and note its name.
  3. Write down whether your organisation requires a particular Azure region for data.
a subscription name and a one-line residency rule. You will need both in the next sections.

The problem it solves, and what came before

Large language models are expensive to run. A serious model needs specialised hardware, careful capacity planning, and a team that keeps it patched and fast. For most companies, running one themselves is out of the question, so the first way to use a frontier model was a public web API: create an account, get a key, send requests.

That worked for experiments and became awkward for companies. A key pasted into a script has full power and no owner. Data leaves for a service your security team has not reviewed. There is no way to say "only this application may call the model" in the language the organisation already uses for permissions. Billing sits on a separate card. And the questions auditors ask, such as where the data is processed and who can see it, have answers that live outside your cloud contract.

Azure OpenAI answers those questions by putting the same family of models behind the controls Azure already has. The model is a resource like a database or a storage account. You create it in a region, give it a name, restrict its network if you need to, and grant roles to people and applications. Microsoft's billing, monitoring and policy tooling all see it.

Two consequences shape everything that follows, so notice them now.

You do not call a model, you call your deployment of a model. In the public API you write model="gpt-5.1". In Azure OpenAI you first create a deployment of that model, give the deployment a name you choose, and write that name in the model field. Mix the two up and you get a 404. This single difference causes more first-day failures than anything else, and we will return to it more than once.

Capacity is something you are assigned, not something you assume. Every deployment has a rate limit measured in tokens per minute, drawn from a quota that your subscription holds per model and region. Hit it and the service answers with HTTP 429. Nothing is broken when that happens; you asked for more than you were allotted. We will look at the numbers when we get there.

The mental model: five nouns

Azure OpenAI has more settings than you will ever need on day one, but it rests on a few nouns. Learn these five and every screen and error message makes sense.

SUBSCRIPTIONthe bill and the quota
→
RESOURCEthe endpoint and keys
→
DEPLOYMENTa named model
→
REQUESTtokens in, tokens out

The resource. An Azure OpenAI resource is an object in your subscription (technically of type Microsoft.CognitiveServices/accounts) that lives in one region. It owns an endpoint, two access keys, network settings and an optional managed identity. There are two flavours. The classic one has kind: OpenAI. The newer one, a Foundry resource, has kind: AIServices; it is a superset that adds other providers' models, agents and projects. Upgrading from the first to the second keeps your name, endpoint, keys and settings, and Microsoft is gradually upgrading eligible resources on its own. For a beginner the difference is small, and the commands in this guide work against both. The one thing to remember is that a resource has a custom subdomain, a name that becomes part of its URL. Passwordless sign-in requires one.

The endpoint. Your endpoint is a URL of the form https://YOUR-RESOURCE-NAME.openai.azure.com. Every call goes to a path under it. The current API lives under /openai/v1/, so the base URL you will use throughout is https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/.

The model and its version. A model is an entry in Microsoft's catalogue, such as gpt-5.1. A model version is a dated release of it, such as 2025-11-13. Models also come in tiers: a family like GPT-5 has big and small variants, with names like mini and nano for the smaller, cheaper ones.

The deployment. A deployment is a named instance of one model version on one resource, with a deployment type and a capacity. The name is yours: my-gpt, chat-prod, support-bot. That name is what API calls address. A resource can hold up to 32 standard deployments, and you can run the same model under several deployment names with different capacities or guardrails.

The token. Models do not read characters or words; they read tokens, which are fragments of text, often a short word or part of a longer one. You are billed per token, limits are expressed in tokens, and a request has input tokens (what you send) and output tokens (what the model writes). For ordinary English, a token is roughly three quarters of a word, so a hundred words is a bit over a hundred tokens. Arabic and other non-Latin scripts often use more tokens per word, which matters when you estimate cost for Arabic content.

Two planes also matter, and you will see both in the docs. The control plane is Azure Resource Manager at management.azure.com: creating resources, deployments, quota and guardrail policies. The data plane is the model API at your endpoint: sending prompts, uploading files, running batches. The Azure CLI commands in this guide use the control plane. Your application code uses the data plane. Permissions are separate for each, which is why someone can be an Owner of the resource and still be refused when calling the model. We will see that exact trap in the errors section.

Trap Model name versus deployment name. gpt-5.1 is a model. my-gpt is a deployment of it. The model field in your code takes the deployment name, and it is case-sensitive. Throughout this guide my-gpt is a placeholder for whatever name you chose.
Try it
  1. Draw the five nouns on paper in order: subscription, resource, deployment, request, token.
  2. Next to each, write one sentence about what it holds or controls.
  3. Mark the one that your code will name in the model field.
deployment. If your drawing says model, redraw it before you continue.

Deployment types, quota and where your data is processed

When you create a deployment you choose a deployment type. The names look intimidating, but they answer two simple questions: where may the request be processed, and how is capacity billed?

On the first question there are three scopes. Global deployments let Microsoft process the request in any Azure region, which gives the newest models, the highest default quota and the lowest price. Data Zone deployments keep processing inside a zone: the United States, the European Union, or Asia-Pacific. Standard (regional) deployments process within the geography of your resource. In every case, your stored data at rest, such as uploaded files, stays in the geography of the resource. That last sentence is worth reading twice if your employer has a residency rule, because "processing location" and "storage location" are different questions with different answers.

On the second question, most deployments are pay-as-you-go: you pay per token and draw on a quota. The alternative is provisioned capacity, where you reserve throughput units and pay for them by the hour whether you use them or not. Provisioned capacity suits steady, high-volume production workloads and is a mid-level topic. A third type, batch, runs large asynchronous jobs at a discount. As a beginner you want one thing: Global Standard. In code or in the CLI it is the SKU name GlobalStandard. The other names you will see in the portal are DataZoneStandard, Standard, GlobalBatch, DataZoneBatch and the provisioned ones, and you can leave them alone for now.

Tip New models arrive first as Global Standard, then Global Provisioned, then Data Zone, and regional Standard comes last, with no guaranteed date. Starting with Global Standard means the model you want will almost always be available. Switch to Data Zone or regional only when a residency rule forces you to, and expect to wait for newer models.

Now the numbers. Quota is an allowance of tokens per minute (TPM) that your subscription holds for a given model and deployment type. When you create a deployment you assign part of that quota to it, and that slice becomes the deployment's rate limit. For standard deployments, one unit of capacity equals 1,000 TPM, so a capacity of 10 means 10,000 tokens per minute. A requests-per-minute limit is derived from it: older chat models get six requests per minute for each 1,000 TPM, and most newer models get one request per minute for each 1,000 TPM. The service also looks at short windows, so a deployment limited to 600 requests per minute can be throttled if it receives more than ten in a single second.

Microsoft has been changing how quota works. Quota now comes in tiers: a free tier (Tier 0) and Tiers 1 to 6, which move up automatically based on usage and payment history. Starting after 7 May 2026, quota has also been moving to a subscription-level pool, where Global Standard shares one pool per model and version across all regions. Exact numbers differ by model and change over time, so do not memorise them. Look at the Quota page in the portal for your own subscription, and treat any figure in a tutorial as an example. As an example only, the free tier lists 200,000 TPM for gpt-4.1-mini and 500,000 TPM for gpt-5-mini.

The practical beginner lesson is modest. Start with a small capacity (10 units is plenty for learning), expect to see a 429 when you loop too fast, and remember that you can raise a deployment's capacity from the portal as long as unassigned quota remains in the pool. Changes can take about 15 minutes to apply.

Try it
  1. In the Foundry portal, open the Quota page for your subscription.
  2. Find a model you would like to use and note the tokens-per-minute figure for GlobalStandard.
  3. Convert it to a rough number of requests per minute at 1,000 tokens per request.
a number you can compare with how many requests your first application will make. If the two are far apart, quota will be your first bottleneck.

Setting up: an account, the Azure CLI and a Python environment

Azure OpenAI is not installed, so setup means three things: an Azure subscription, the command-line tool for managing it, and a language library for calling the model. Many people do the first part in the portal. Doing it from the command line instead gives you something you can paste into a note, repeat next month, and eventually turn into infrastructure-as-code.

Install the Azure CLI, called az. On macOS use Homebrew:

BASH
brew update && brew install azure-cli

On Windows:

BASH
winget install --exact --id Microsoft.AzureCLI

On Debian or Ubuntu Linux:

BASH
curl -sL https://aka.ms/InstallAzureCLIDeb | sudo bash

Check that it works:

BASH
az version

You should see a JSON block listing azure-cli with a version number; this guide was prepared against 2.90. Then sign in:

BASH
az login

A browser window opens, you sign in with your work or personal Azure account, and the terminal prints the subscriptions you can use. If you belong to several organisations, name the one you want with az login --tenant <TENANT_ID>. Confirm where you landed:

BASH
az account show --query "{subscription:name, tenant:tenantId}" --output table

If the subscription is not the one you intended, switch with az account set --subscription "<name or id>" before creating anything. Creating a resource in the wrong subscription is an easy mistake and an annoying one to undo.

Now create a Python environment and install the two libraries you need. The openai package is the client, and azure-identity handles passwordless sign-in:

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install openai azure-identity

On Windows PowerShell, activate with .venv\Scripts\Activate.ps1. The current openai release at the time of writing is in the 3.x line. A virtual environment keeps these packages away from the rest of your system, and it is the habit to build from day one.

Note If you work behind a corporate proxy or a security gateway that inspects TLS traffic, az login or Python may fail with certificate errors. Do not turn certificate verification off. Ask your IT team for the organisation's root certificate and point the tools at it.
Try it
  1. Run az version and az login.
  2. Run the az account show command above.
  3. Create the virtual environment and run pip show openai.
a subscription name, a tenant ID and an installed openai package. If az login opens a browser and returns an error, read the message before retrying; it usually names the missing permission.

Creating a resource and your first deployment

You will now create the two things that make the endpoint exist: a resource and a deployment. The commands below are the ones in Microsoft's documentation, adjusted to a current model.

First a resource group, a folder-like container that makes cleanup easy because deleting it deletes everything inside:

BASH
az group create --name OAIResourceGroup --location eastus

Pick the region with care. A resource lives in one region, and the models and quota available differ from region to region. Check Microsoft's model availability pages for the region closest to you that also satisfies your residency rule. If you are in the Middle East, remember that not every model is offered in every nearby region, and that the processing location depends on the deployment type you pick, as described earlier.

Now the resource itself:

BASH
az cognitiveservices account create \
  --name MyOpenAIResource \
  --resource-group OAIResourceGroup \
  --location eastus \
  --kind OpenAI \
  --sku s0 \
  --custom-domain MyOpenAIResource \
  --yes

Resource names must be unique across Azure because the name becomes part of the URL, so choose your own and use the same string for --custom-domain. Always pass --custom-domain. Without it, passwordless sign-in does not work, and Microsoft documents a known problem where resources created from the CLI without a proper subdomain misbehave with batch jobs. The --kind OpenAI flag makes a classic Azure OpenAI resource. If you want a Foundry resource, the kind is AIServices; for learning, OpenAI is simpler.

Read back the two values you will use in code:

BASH
az cognitiveservices account show \
  --name MyOpenAIResource --resource-group OAIResourceGroup \
  --query properties.endpoint --output tsv

az cognitiveservices account keys list \
  --name MyOpenAIResource --resource-group OAIResourceGroup \
  --query key1 --output tsv

The first prints the endpoint, something like https://MyOpenAIResource.openai.azure.com/. The second prints a secret key. Treat that key like a password; we will discuss storing it shortly.

Now deploy a model:

BASH
az cognitiveservices account deployment create \
  --name MyOpenAIResource \
  --resource-group OAIResourceGroup \
  --deployment-name my-gpt \
  --model-name gpt-5.1 \
  --model-version "2025-11-13" \
  --model-format OpenAI \
  --sku-name GlobalStandard \
  --sku-capacity 10

Read it as a sentence. On the resource MyOpenAIResource, create a deployment called my-gpt, which is the model gpt-5.1 at version 2025-11-13, as Global Standard, with a capacity of 10 units (10,000 tokens per minute). If it fails with a message about quota or availability, either the region lacks that model or your subscription has no quota for it. List what you have and try a different region or model:

BASH
az cognitiveservices account deployment list \
  --name MyOpenAIResource --resource-group OAIResourceGroup --output table

The deployment you just created should appear in the table with its model and capacity. Remember that a newly created deployment can answer with a "deployment does not exist" error for a few minutes; the service says so in the message.

The same steps work in the portal. Open Microsoft Foundry at ai.azure.com, select your resource, go to the model catalogue, choose the model, and click Deploy. The portal shows a default deployment name equal to the model name, and many beginners keep it. That is fine, as long as you remember that the name in the box is now the name your code must use. There are two portal experiences, the new Foundry and "Foundry (classic)", switched with a toggle, and screenshots in tutorials may match either. The CLI does not change, which is another reason to learn it.

When you finish experimenting, clean up. Deleting a resource group removes everything in it:

BASH
az group delete --name OAIResourceGroup --yes

One caution. A deleted Azure OpenAI resource is soft-deleted. If you delete a resource through the API or CLI while deployments still exist, its quota stays locked for 48 hours unless you purge it. Delete the deployments first, then the resource, or purge the deleted resource.

Trap Forgetting that you are billed. A pay-as-you-go deployment costs nothing while idle, but a loop that calls the model costs money per token. Keep experiments short and delete the resource group when you are done.
Try it
  1. Create the resource group and the resource with your own unique name.
  2. Create the deployment my-gpt and list deployments.
  3. Save the endpoint and key into environment variables, as shown in the next section.
one deployment shown as succeeded. If creation was refused for quota, try another region or a smaller model rather than fighting it.

Your first call

You now have an endpoint, a key and a deployment. Put the key and endpoint into environment variables so they never appear in your code. On macOS or Linux:

BASH
export AZURE_OPENAI_API_KEY="paste-your-key-here"
export AZURE_OPENAI_ENDPOINT="https://MyOpenAIResource.openai.azure.com"

Start with curl, because it shows you the raw request without any library in the way:

BASH
curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/responses" \
  -H "Content-Type: application/json" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -d '{"model": "my-gpt", "input": "This is a test."}'

Look at its parts. The URL is your endpoint plus /openai/v1/responses. The key travels in an api-key header. The body is JSON with two fields: model, which holds your deployment name, and input, which holds the prompt. The reply is a larger JSON document containing, among other things, an id, a status of completed, token counts under usage, and an output list whose first message holds the text.

Now the same call in Python, using the plain OpenAI client pointed at your endpoint. This is the current pattern for the v1 API:

hello.py
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.getenv("AZURE_OPENAI_API_KEY"),
    base_url=os.getenv("AZURE_OPENAI_ENDPOINT") + "/openai/v1/",
)

response = client.responses.create(
    model="my-gpt",  # the DEPLOYMENT name, not the model name
    input="Explain what a token is in one sentence.",
)

print(response.output_text)

Run it with python hello.py. You should see one sentence about tokens. If you do, you have completed the hard part, and everything else in this guide is variations on this call.

Notice also the API we called. responses.create is the Responses API, which is now the recommended way to talk to Azure OpenAI models. The older Chat Completions API still works, and you will see it everywhere in tutorials, so we cover it below. The shape of the answer differs slightly: the Responses API gives you response.output_text as a convenience, while Chat Completions nests the text under choices[0].message.content.

Tip If you set OPENAI_API_KEY and OPENAI_BASE_URL in your environment, the client can be created with no arguments at all: client = OpenAI(). That is convenient, but it also means a stray variable from another project can silently redirect your calls, so be explicit while you learn.
Try it
  1. Run the curl call and find output_text equivalents in the JSON by eye.
  2. Run hello.py.
  3. Change the model value to a name you did not create and run it again.
a 404 saying the deployment does not exist. Seeing that error on purpose now makes it easy to recognise later.

Signing in without a password: keys versus Entra ID

You have been using an API key. A key is simple, and it is also the weakest option. It grants full access to the data plane with no role checks, anyone who holds it can use your quota and your bill, and it is easy to leak into a repository or a screenshot. Azure gives you a better way: Microsoft Entra ID, Azure's identity system, where the caller proves who it is and Azure checks a role.

With Entra ID, your code holds no secret. Locally, it borrows the identity you used for az login. In Azure, it uses a managed identity attached to the service that runs your code. The library that does this is azure-identity, and its workhorse is DefaultAzureCredential, which tries several sources in order (environment variables, a managed identity, your Azure CLI login, and more) until one works.

The code changes by only a few lines:

entra.py
from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider

token_provider = get_bearer_token_provider(
    DefaultAzureCredential(), "https://ai.azure.com/.default"
)

client = OpenAI(
    base_url="https://MyOpenAIResource.openai.azure.com/openai/v1/",
    api_key=token_provider,
)

response = client.responses.create(model="my-gpt", input="Say hello.")
print(response.output_text)

The only odd line is api_key=token_provider. The api_key parameter accepts a function as well as a string. The client calls the function before each request, gets a fresh short-lived token, and sends it as an Authorization: Bearer header. You never handle tokens yourself, and they refresh automatically. The string https://ai.azure.com/.default is the scope, the audience the token is meant for. Older documentation uses https://cognitiveservices.azure.com/.default; current docs use the Foundry one shown above, and a token for one audience does not work on another, so copy the scope from current documentation.

Two things must be true for this to work, and each is a classic first-week failure.

First, your resource needs the custom subdomain you set with --custom-domain. Calls to a regional endpoint instead of your own subdomain return 401.

Second, your identity needs a role on the resource. Being the Owner of the subscription is not enough. Owner and Contributor are control-plane roles: they let you manage the resource but not call the model through Entra. The role for inference is Cognitive Services OpenAI User, and on a Foundry project the equivalent is Foundry User. Assign it to yourself with the CLI:

BASH
az role assignment create \
  --role "Cognitive Services OpenAI User" \
  --assignee "$(az ad signed-in-user show --query id --output tsv)" \
  --scope "$(az cognitiveservices account show \
      --name MyOpenAIResource --resource-group OAIResourceGroup \
      --query id --output tsv)"

Role assignments can take up to five minutes to propagate, so a 403 right after assigning one is normal. Wait a few minutes before debugging. Note that the Cognitive Services Contributor role, despite its name, cannot make Entra inference calls either; it manages the resource and its keys.

You can check that Azure will hand you a token, which separates "I am not logged in" from "I lack a role":

BASH
az account get-access-token --resource https://ai.azure.com --query expiresOn --output tsv

If that prints a date, your login works and any remaining failure is about roles or the endpoint.

So which should you use? For learning on your own laptop, a key is acceptable if you keep it in an environment variable and never commit it. For anything that a team or a company depends on, use Entra ID, and eventually turn keys off entirely. That hardening step is a senior topic, and the point for now is to build the habit: your first real application should use DefaultAzureCredential.

Entra ID

  • No secret in your code
  • Access checked against a role
  • Short-lived tokens, refreshed for you
  • Needs a custom subdomain and a role

API key

  • A secret that can leak
  • Full data-plane access, no role checks
  • Same key for everyone who holds it
  • Works immediately, which is its only advantage
Try it
  1. Assign yourself the Cognitive Services OpenAI User role on your resource.
  2. Wait five minutes, then run entra.py.
  3. Unset AZURE_OPENAI_API_KEY and run it again.
the same answer with no key in your environment. That is the proof your code holds no secret.

Everyday calls: instructions, conversations, streaming and Chat Completions

With the connection solved, most of your work is choosing what to send. This section covers the calls you will make daily.

Giving the model instructions. A model behaves according to its input. The Responses API lets you pass a separate instructions string that sets the role and rules, and an input that holds the user's request:

instructions.py
response = client.responses.create(
    model="my-gpt",
    instructions="You are a support assistant for a bank. Answer in two sentences, in plain language.",
    input="How do I reset my online banking password?",
)
print(response.output_text)

Instructions are the main lever you have. Be specific about audience, length, format and what to do when the model is unsure. "Answer only from the text below; if the answer is not there, say you do not know" is worth more than any amount of hoping.

Holding a conversation. Models are stateless, so a second request knows nothing about the first unless you tell it. The Responses API gives you two ways to carry context. The first is to chain by response ID:

chain.py
first = client.responses.create(
    model="my-gpt",
    input="My name is Aya and I work in Cairo.",
)

second = client.responses.create(
    model="my-gpt",
    input="Where do I work?",
    previous_response_id=first.id,
)
print(second.output_text)

Here the service remembers the first exchange on your behalf. By default it stores responses for 30 days, which is how previous_response_id can work. If you must not have the service store anything, pass store=False and use the second approach, which is to send the history yourself: keep a Python list of messages, append the model's output and your new question, and send the whole list each time. You can also fetch a stored response with client.responses.retrieve(first.id) and remove it with client.responses.delete(first.id). Either way, remember that each turn re-reads the entire history, and you are billed for those input tokens every time. Long chats get more expensive per message as they grow.

Streaming. A long answer can take many seconds to finish. Streaming sends the text as it is generated, so a user sees words appear immediately:

stream.py
stream = client.responses.create(
    model="my-gpt",
    input="Write a short paragraph about the Nile.",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
print()

The stream is a series of events. You print the delta of each event whose type is response.output_text.delta, and ignore the rest. One subtlety worth knowing: if a failure happens in the middle of a stream, the HTTP status was already 200, so the error arrives as an event inside the stream rather than as a status code. Code that only checks the status will miss it.

Chat Completions, the older API. Most tutorials you will find use Chat Completions, and it is still supported, so recognise it:

chat.py
completion = client.chat.completions.create(
    model="my-gpt",
    messages=[
        {"role": "developer", "content": "You are a concise assistant."},
        {"role": "user", "content": "Name three Azure regions in the Middle East."},
    ],
    max_completion_tokens=500,
)
print(completion.choices[0].message.content)

Here the conversation is a list of messages, each with a role. The instruction message takes the role developer, which behaves the same as the older system role; send one or the other, not both. The reply text is at completion.choices[0].message.content. Note max_completion_tokens rather than max_tokens: reasoning models reject max_tokens. Use Chat Completions when you must follow an existing tutorial or codebase, and Responses for new work.

Embeddings. An embedding turns text into a list of numbers so that texts with similar meaning end up close together. You use them for search, recommendations and retrieval. You need a separate deployment of an embedding model, such as text-embedding-3-small, and you call it by that deployment name:

embed.py
result = client.embeddings.create(
    model="my-embedding",  # the name of YOUR embedding deployment
    input=["How do I reset my password?", "Forgot login credentials"],
)
print(len(result.data), len(result.data[0].embedding))

You get one vector per input. For text-embedding-3-small each vector has 1,536 numbers, and text-embedding-3-large produces 3,072. A single call accepts up to 2,048 inputs and 300,000 tokens in total, and each input may be up to 8,192 tokens. The two vectors above would be close to each other, which is the point. One rule matters for planning: you cannot mix vectors from different embedding models, and you cannot upgrade in place. If you change the model, you must embed everything again, so choose deliberately before you embed a million documents. When you are ready to store and search vectors, the guides on pgvector, Qdrant and FAISS pick up where this section stops.

Trap The Assistants API appears in many older Azure tutorials. Microsoft retired it on 26 August 2026, so any tutorial that creates an "assistant" or a "thread" will no longer work. The Responses API, with conversations and tools, replaced it.
Try it
  1. Run chain.py and confirm the model remembers your city.
  2. Change the second call to omit previous_response_id.
  3. Run stream.py and watch the text arrive.
with the chain, the right city; without it, the model says it does not know. That difference is what "stateless" means.

Reasoning models and the parameters that change

Some deployments are reasoning models: they think through a problem internally before answering, which improves results on logic, maths and code at the cost of time and tokens. Models in the o-series and the GPT-5 family behave this way. For a beginner they matter because they accept different parameters from classic chat models, and sending the old ones produces a 400 error.

Three rules cover most of the trouble. First, use max_output_tokens on the Responses API, or max_completion_tokens on Chat Completions, never max_tokens. Second, do not send sampling controls such as temperature, top_p, presence_penalty, frequency_penalty, logprobs or logit_bias to the o-series and older GPT-5 models; they are rejected. The newest GPT-6 models are an exception and do accept logprobs, temperature and top_p. Third, the output budget includes the model's hidden reasoning tokens, which are billed as output tokens. A small max_output_tokens can be consumed entirely by thinking, leaving a response that is cut off or empty.

Reasoning models also take a reasoning effort, which says how hard to think:

effort.py
response = client.responses.create(
    model="my-gpt",
    input="A train leaves at 09:10 and arrives at 11:45. How long is the trip?",
    reasoning={"effort": "low"},
    max_output_tokens=2000,
)
print(response.output_text)

The accepted effort values depend on the model: none, minimal, low, medium, high, and for some newer models xhigh and max. Support differs, and one detail catches people who upgrade: gpt-5.1 defaults to none, so code written for an older reasoning model that relied on a default will behave differently. If quality matters, pass the effort explicitly. Higher effort costs more tokens and takes longer, so match it to the task: low for simple lookups and formatting, higher for puzzles and code.

When a Responses call returns status: "incomplete" with an incomplete_details.reason of max_output_tokens, the budget ran out before the answer finished. Raise max_output_tokens, shorten the prompt, or lower the effort. Reading status is a habit worth forming: a response with HTTP 200 is not always a finished response.

Tip If you are unsure whether a deployment is a reasoning model, send a minimal request first with no optional parameters. Add parameters one at a time. The error message names the parameter that was rejected, so you learn the model's rules from the service itself.
Try it
  1. Add temperature=0.2 to a call against a reasoning deployment.
  2. Read the error and note which parameter it names.
  3. Set max_output_tokens=20 on a harder question and check the status field.
a 400 for the first and an incomplete response for the second, which together explain most reasoning-model surprises.

Safety filters, now called guardrails

Every Azure OpenAI deployment has a safety layer in front of the model and behind it. Microsoft's classifiers read the prompt before it reaches the model and read the completion before it reaches you. They score text in four categories (hate, sexual content, violence, and self-harm) at levels from safe through low and medium to high. Out of the box, content at medium severity or above is blocked. These used to be called content filters; the current name is guardrails, and in the portal they live under "Guardrails + controls". Optional extras exist, such as prompt-injection detection (Prompt Shields), protected-material detection, and PII detection.

The filters are attached to a deployment, and as a beginner you will meet them as behaviour rather than configuration. There are two ways they show up in your results. If the prompt is blocked, the call fails with HTTP 400:

JSON
{"error": {"message": "The response was filtered", "param": "prompt", "code": "content_filter", "status": 400}}

In Python this raises openai.BadRequestError, and you can detect it by checking the code:

filtered.py
import openai

try:
    response = client.responses.create(model="my-gpt", input=user_text)
except openai.BadRequestError as err:
    if getattr(err, "code", None) == "content_filter":
        print("That request was blocked by the content guardrails. Please rephrase.")
    else:
        raise

If the completion is filtered, the call succeeds with HTTP 200, but the output is empty or cut short and the finish reason is content_filter. On Chat Completions you read completion.choices[0].finish_reason. That second case is the dangerous one, because nothing raised an error. An application that does not check will show the user a blank answer.

So build two habits. Catch the 400 and tell the user kindly to rephrase. And always look at the finish reason or the response status rather than assuming that a 200 contains a usable answer. You cannot switch the filters off on your own; relaxing them requires an application to Microsoft for modified content filtering, and that is for specific approved use cases. For a beginner project, the defaults are what you will use, and they are a feature: they are part of why a company can put the model in front of customers.

Try it
  1. Wrap one of your calls in the try block above.
  2. Print response.status and the usage numbers after each successful call.
  3. Write one sentence in your notes about what your app should show if a completion comes back filtered.
a wrapper that handles both a rejected prompt and an unfinished answer. Do not test the filters with harmful content; reading the documentation of the response format is enough.

Configuration, cost and good housekeeping

A handful of settings and habits keep a first project cheap and tidy.

Keep configuration in the environment. Your code needs three values: the endpoint, the deployment name, and (if you use a key) the key. Read them from environment variables so the same code runs on your laptop, in a test environment and in production. A small settings block at the top of your file is enough:

settings.py
import os

ENDPOINT = os.environ["AZURE_OPENAI_ENDPOINT"].rstrip("/")
DEPLOYMENT = os.environ.get("AZURE_OPENAI_DEPLOYMENT", "my-gpt")
BASE_URL = f"{ENDPOINT}/openai/v1/"

Using os.environ[...] for the endpoint makes the program fail loudly on a missing variable instead of sending a call to nowhere. Put local values in a file named .env or a shell profile, and add that file to .gitignore before your first commit. A key committed to a repository must be treated as leaked: rotate it. The resource has two keys, key1 and key2, precisely so that you can switch your applications to one and regenerate the other without downtime.

Understand what you are billed for. Standard deployments are billed per token, with separate prices for input tokens, cached input tokens and output tokens, and the price depends on the model. Reasoning tokens count as output. Every response tells you what it consumed in response.usage, so you can calculate the cost of a call and watch it. Two rules reduce a bill quickly. Send less: trim long histories and unnecessary context. And set a sensible max_output_tokens so one runaway answer cannot cost much. Current prices change, so use the official pricing page rather than any figure from a guide.

Watch your limits. The service returns headers on each response, including x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-limit-tokens and x-ratelimit-remaining-tokens. On a 429 it adds retry-after-ms, the delay it wants you to wait. The OpenAI Python client already retries rate-limit errors and transient failures twice by default. You can raise that:

retries.py
client = OpenAI(
    base_url=BASE_URL,
    api_key=token_provider,
    max_retries=5,
)

If you add your own retry library such as tenacity on top, set max_retries=0 on the client, or the two layers multiply and you send far more traffic than you expect, which makes the rate limit worse.

Keep an eye on the lifecycle. Models retire. Check the retirement schedule for your model when you start a project and again every few months. A deployment can be set to upgrade automatically to a newer default version of the same model, or to stay put; with the second setting, the deployment stops working on the retirement date. Subscribe to Azure Service Health advisories so that the notice reaches a person and not only an unread mailbox.

Try it
  1. Move your endpoint and deployment name into environment variables.
  2. Print response.usage after a call and compute the cost with the pricing page.
  3. Open the Metrics page of your resource and chart processed tokens.
a code file with no hard-coded endpoint, a cost per call you can state, and a chart that moves when you run your script.

Common errors and how to read them

Azure OpenAI errors are specific once you know how to read them. The error body is JSON with a code and a message, and the HTTP status tells you which side to look at: 400 means your request was wrong, 401 and 403 mean identity or permission, 404 means something named does not exist, 429 means a limit, and 5xx means the service. The most common ones follow.

404 DeploymentNotFound: The API deployment for this resource does not exist. Three causes, in order of likelihood. You wrote the model name where the deployment name belongs. You spelled the deployment name with different capitalisation. Or the endpoint belongs to a different resource from the one where you created the deployment. The message also suggests waiting five minutes if you created the deployment very recently, and that is real. Check with az cognitiveservices account deployment list and compare the name character by character.

401 Unauthorized with an Entra token. The resource has no custom subdomain, or you called a regional endpoint instead of https://<name>.openai.azure.com. Recreate with --custom-domain, or add a subdomain, and use your own endpoint.

403 Forbidden or PermissionDenied. You authenticated, but lack a data-plane role. Assign Cognitive Services OpenAI User (or Foundry User on a project) on the resource, and wait up to five minutes. Being an Owner or Contributor does not count, because those roles are control-plane only.

DefaultAzureCredential failed to retrieve a token. The credential chain found nothing it could use. Run az login again, with --tenant if you belong to several tenants, then retry the az account get-access-token check from earlier. If the code works on your laptop and fails in Azure, the cause is almost always that the managed identity is not enabled, or that the role was given to your user account and not to the managed identity. With a user-assigned managed identity you must also tell the library which one by setting AZURE_CLIENT_ID.

429 rate limit. The messages read "Requests to the ... have exceeded call rate limit" or "Rate limit is exceeded". You used more tokens or requests per minute than the deployment's allocation. Back off using retry-after-ms, lower max_output_tokens, raise the deployment's capacity if you have spare quota, and smooth bursts, since the requests-per-minute limit is checked over windows of a second or ten. Two surprises: the service estimates tokens from your prompt plus max_output_tokens and not the tokens you will really use, so a very large output cap can trip the limit early; and failed requests still count toward the limit. A different 429 reads "The service is temporarily unable to process your request" or mentions high demand. That one is capacity pressure on Microsoft's side; wait and retry.

400 with content_filter. Covered earlier: a guardrail blocked the prompt. Rephrase it.

400 on a reasoning model naming temperature or max_tokens. You sent a parameter the model rejects. Remove the sampling parameters, and use max_output_tokens or max_completion_tokens. On GPT-5.6 models there is one more case: using tools with reasoning on Chat Completions fails with "Function tools with reasoning_effort are not supported ... To use function tools, use /v1/responses or set reasoning_effort to 'none'." The fix is the Responses API.

410 Gone. The model version is retired. Create a deployment of a current model and change the deployment name or its model.

status: "incomplete" with HTTP 200. The response ran out of max_output_tokens. Raise the limit, lower reasoning effort, or trim the prompt.

Try it
  1. Make a deliberate error of each kind you can safely cause: a wrong deployment name, a wrong key, an unsupported parameter.
  2. For each, write down the status, the code, and your fix in one line.
  3. Keep the list; you will use it again.
a three-line cheat sheet in your own words, which is worth more than any table because you earned it.

Putting it all together: a small command-line assistant

Here is one small project that uses what you have learned. It is a command-line assistant for a fictional help desk. It reads its settings from the environment, signs in with Entra ID, keeps a conversation going, streams the answer, handles filtered prompts and rate limits, and shows what each turn cost in tokens.

assistant.py
import os
import sys

import openai
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
from openai import OpenAI

ENDPOINT = os.environ["AZURE_OPENAI_ENDPOINT"].rstrip("/")
DEPLOYMENT = os.environ.get("AZURE_OPENAI_DEPLOYMENT", "my-gpt")

token_provider = get_bearer_token_provider(
    DefaultAzureCredential(), "https://ai.azure.com/.default"
)
client = OpenAI(
    base_url=f"{ENDPOINT}/openai/v1/",
    api_key=token_provider,
    max_retries=4,
)

INSTRUCTIONS = (
    "You are a friendly help-desk assistant for an IT team. "
    "Answer in at most four sentences. If you do not know, say so."
)

previous_id = None
print("Ask a question, or type 'quit'.")

while True:
    question = input("\nYou: ").strip()
    if question.lower() in {"quit", "exit"}:
        break
    if not question:
        continue

    try:
        stream = client.responses.create(
            model=DEPLOYMENT,
            instructions=INSTRUCTIONS,
            input=question,
            previous_response_id=previous_id,
            max_output_tokens=800,
            stream=True,
        )
        print("Assistant: ", end="", flush=True)
        for event in stream:
            if event.type == "response.output_text.delta":
                print(event.delta, end="", flush=True)
            elif event.type == "response.completed":
                previous_id = event.response.id
                usage = event.response.usage
                print(f"\n[tokens in={usage.input_tokens} out={usage.output_tokens}]")
    except openai.BadRequestError as err:
        if getattr(err, "code", None) == "content_filter":
            print("Assistant: I cannot help with that request. Please rephrase it.")
        else:
            print(f"Request problem: {err}")
    except openai.RateLimitError:
        print("Assistant: The service is busy. Please try again in a moment.")
    except openai.AuthenticationError:
        sys.exit("Sign-in failed. Run 'az login' and check your role assignment.")
    except openai.NotFoundError:
        sys.exit("Deployment not found. Check AZURE_OPENAI_DEPLOYMENT and your endpoint.")

Read it as a checklist of what you now understand. The settings come from the environment. The client signs in with DefaultAzureCredential, so there is no key in the file. The conversation continues through previous_response_id, which the service honours because responses are stored by default. The answer streams, and the final event gives you the usage numbers. Each exception class corresponds to an error from the previous section, and each produces a message a human can act on. The max_retries=4 setting means the client already retries transient failures before you ever see a RateLimitError.

Run it with:

BASH
export AZURE_OPENAI_ENDPOINT="https://MyOpenAIResource.openai.azure.com"
export AZURE_OPENAI_DEPLOYMENT="my-gpt"
python assistant.py

Ask it a few questions, then ask one that depends on the previous answer ("and how long does that take?") to confirm the conversation chaining works. Stop with quit. When you are finished with the project, remove everything with az group delete --name OAIResourceGroup --yes.

Try it
  1. Build and run assistant.py against your own deployment.
  2. Add a question that triggers a follow-up and confirm it uses the earlier answer.
  3. Temporarily rename the deployment in the environment variable and watch the not-found message.
a working assistant, plus the not-found message arriving in your own words rather than a stack trace.

What you can now do, and what comes next

You can now explain, in plain terms, what Azure OpenAI is: a model hosted in your own Azure subscription, reached through an endpoint that you control and secure with Azure roles. You can name its nouns (resource, deployment, quota, token), and you know the difference that trips everyone up, that code addresses the deployment name and not the model name. You can create a resource group, a resource and a deployment from the command line, call the deployment with curl and with the Python openai client against the v1 API, sign in with Entra ID instead of a key, stream and chain responses, create embeddings, and read the common errors by status code and message.

You also know what you do not yet know. You have not sized capacity for real traffic or compared provisioned throughput with pay-as-you-go. You have not used structured outputs, tools, vision or batch jobs. You have not locked the resource to a private network, turned off keys, written infrastructure as code, or built dashboards and alerts. You have not planned a model upgrade across a retirement date. Those are the mid-level and senior topics, and the mid-level guide picks up exactly where you now stand.

Some neighbouring guides in this catalogue complete the picture. If you want to compare Azure's hosting with the original, read the OpenAI API guide; the calls are nearly identical and the differences are the ones you have just learned. If your team uses Azure for training and model management, the Azure ML guide covers that side of the platform. To put several model providers behind a single interface, see LiteLLM. And once your application is running, Langfuse shows how to trace what it actually sends and receives.

Sources