This is part one of three. It covers what you need to do real work with Amazon SageMaker AI, from an empty AWS account to a trained model, a live endpoint you can call, and a clean shutdown that leaves nothing billing in the background. By the end you will understand the handful of concepts the whole service is built from, you will have run the commands that matter, and you will be able to read an error message and know which of three or four causes produced it. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task where it makes sense. Do them as you go. SageMaker is a cloud service, so these tasks cost a little real money, but the guide tells you which resources bill and how to delete them, and the cost of the whole beginner path done carefully is small.
What SageMaker AI is, and what you can do by the end
Amazon SageMaker AI is a managed service for the machine learning lifecycle on AWS. You bring data and code. SageMaker rents you the computers to train a model on that data, stores the result, and then hosts the trained model behind a web address so that other programs can send it inputs and receive predictions. You do not install or patch a single server. You describe what you want in a few lines of Python or a few CLI commands, and AWS starts machines, runs your work on them, and shuts them down.
That diagram is the spine of the service. Almost everything else in SageMaker is a variation, an automation, or a safety net around one of those four boxes.
By the end of this guide you will be able to:
- Explain the difference between a notebook, a training job, a model, and an endpoint, and say which of them bills you while idle.
- Set up an AWS account safely, configure the command line, and confirm that your credentials work.
- Create the IAM role SageMaker needs, and explain why it is separate from your own permissions.
- Start a training job from Python with the current SDK, find its logs, and read why it failed when it fails.
- Deploy a model to a real-time endpoint, call it, and delete it.
- Read the most common SageMaker error messages and name the fix.
Two warnings before the first command. First, SageMaker bills by the second or hour for machines that are running, so the discipline of cleaning up is part of the skill, not an afterthought. Second, most tutorials you will find on the web were written for version 2 of the Python SDK, and version 3, released in November 2025, removed much of what they teach. This guide teaches the current version and tells you when you are looking at the old one.
- Open the AWS pricing page for SageMaker AI (linked in Sources) and find the hourly price of an
ml.m5.xlargeinstance in the region you plan to use. Multiply it by 24 and by 30. That number is what a forgotten endpoint costs in a month, and it is the reason for the cleanup habits later in this guide.
The problem it solves, and what came before
Training a machine learning model sounds like a software task, but it behaves like an infrastructure task. A model that learns from a few thousand rows runs fine on a laptop. A model that learns from millions of images, or a language model you want to adapt to Arabic text, needs a GPU, and a GPU machine is expensive to own and mostly idle if you only train a few times a week. Then, once the model exists, somebody has to serve it: keep a process running, answer requests quickly, survive a machine failing, and handle ten times the traffic on a busy day.
Before managed services, a team solved this by hand. Someone rented virtual machines, installed GPU drivers (a famously fragile step), copied data over, ran a training script inside a terminal session, copied the resulting file somewhere, and wrote a small web server to load it. Each of those steps was a place for the environment to drift: a driver mismatch here, a forgotten package there, a model file nobody could find six months later. Machines were left running over weekends. Nobody could say which data and which code had produced the model currently in production.
SageMaker AI packages the repeatable parts. A training job is a container that runs on machines AWS starts for you and tears down when the script finishes, so you pay only for the minutes of training. Data comes from, and the model goes to, Amazon S3, so inputs and outputs live in durable, addressable places. The hosting side is the same idea: you say "run this model on this instance type", and AWS keeps it running behind an HTTPS address, replaces it if it fails, and can add copies when traffic grows. Around those core pieces sit services for tracking experiments, orchestrating multi-step workflows, registering approved model versions, and watching a deployed model for drift. The Mid-level and Senior guides cover those.
It is worth being honest about the trade-off. SageMaker is not the only way to do this, and it is not free of friction. You are working inside AWS's identity, networking, and billing systems, so a good part of learning SageMaker is learning IAM roles and S3 permissions. In exchange you avoid running GPU fleets yourself. If you later work on another cloud, the same ideas appear under different names: Google's equivalent is covered in the Vertex AI guide, and if you prefer to run serving on your own Kubernetes cluster, KServe and Kubeflow are the open-source counterparts.
For learners in the Gulf and Egypt, region matters. AWS operates regions in the Middle East, including Bahrain (me-south-1) and the UAE (me-central-1), and an employer subject to data-residency rules may require that training data stay in one of them. Not every instance type is offered in every region, and the regional list changes over time, so check the regional endpoints and quotas page (linked in Sources) before you promise a project to a client.
A naming note: SageMaker, SageMaker AI, and the two SDK generations
Two naming changes confuse almost every newcomer, so it is better to meet them on purpose.
The product was renamed. On 3 December 2024, "Amazon SageMaker" became "Amazon SageMaker AI". The name "Amazon SageMaker" was then reused for a broader umbrella platform that includes SageMaker AI alongside data and analytics services such as the lakehouse, the data catalog, and SageMaker Unified Studio. The machine learning service this guide covers is SageMaker AI. Crucially, nothing in code was renamed. The API namespace is still sagemaker, the CLI command is still aws sagemaker, the managed policies are still called AmazonSageMaker..., and the Python package is still sagemaker. When the docs say "SageMaker AI" and your terminal says sagemaker, they mean the same service.
The Python SDK had a breaking release. The SDK is the library you pip install to drive SageMaker from Python. Version 2 had classes called Estimator, Model, and Predictor, plus per-framework subclasses such as sagemaker.pytorch.PyTorch. Version 3.0.0, released on 19 November 2025, states in its changelog that these older interfaces are not supported in V3. Their replacements are ModelTrainer for training and ModelBuilder for building and deploying models. The SDK is also split into modular pieces (sagemaker-core, sagemaker-train, sagemaker-serve, sagemaker-mlops), though pip install sagemaker installs the lot. At the time of writing the current stable release is 3.23.0, which needs Python 3.10 or newer.
The practical consequence is simple and important. If you copy a tutorial and see from sagemaker.estimator import Estimator or from sagemaker.pytorch import PyTorch, you are reading version 2 code. On a fresh install of version 3 it fails with an import error. You have two options: translate it to the v3 form, or create a separate virtual environment and install sagemaker==2.* for that tutorial. This guide uses v3 throughout, and when a v2 name appears it is labelled as legacy.
pip show sagemaker and look at the version line. Half of the "SageMaker is broken" questions beginners ask turn out to be v2 code running on v3, or the reverse. Knowing which generation you are on removes a whole category of confusion.
Underneath either SDK there is also boto3, the general AWS SDK for Python. It exposes every SageMaker API call directly through two clients: sagemaker for the control plane (create, describe, delete things) and sagemaker-runtime for calling a deployed endpoint. The high-level SageMaker SDK is built on top of the same API, which is why the AWS CLI, boto3, and the SDK all agree on the underlying names. You will use boto3 in this guide for the endpoint lifecycle because it shows exactly what is created, one resource at a time.
The mental model: the nouns you need
SageMaker has hundreds of features but a small vocabulary at its core. Learn these nouns well and the documentation becomes readable.
Account, region, and IAM. Everything in AWS lives in an account and, for SageMaker, in a region. A training job in eu-west-1 is invisible from me-central-1. Many beginner errors (an endpoint "not found", an S3 bucket that cannot be read) are simply the wrong region. Access to everything is governed by IAM, AWS's identity and permissions system.
S3. Amazon S3 is object storage: buckets containing files addressed by URIs like s3://my-bucket/train/data.csv. SageMaker reads training data from S3 and writes model files back to S3. It is the shared hard drive of the whole service.
Execution role. SageMaker does work on your behalf, such as reading your S3 bucket, pulling a container image from the registry, and writing logs. To do that it assumes an IAM role that you provide, called the execution role. This is not your own login. It is a separate identity that the service borrows for the duration of the job. Keeping the two apart is the single most common stumbling block, and it gets its own section below.
Container image. Every training job and every endpoint runs a Docker container. If containers are new to you, read the Docker guide first or alongside. AWS publishes prebuilt images for common frameworks, and you can build your own and store it in Amazon ECR, AWS's container registry.
Training job. A one-off run of a container on one or more managed instances. It receives data in channels, runs your code, and saves the result. When the code exits, the instances are released. A training job has a name, a status (InProgress, Completed, Failed, Stopped), and a permanent record you can inspect afterwards.
Model artifact. The output of training, usually a file named model.tar.gz in S3. It is just an archive of whatever your training code saved into the directory /opt/ml/model.
Model. In SageMaker's API, a "model" is a small resource that pairs a container image with the location of a model artifact. It does not run anything by itself. It is a recipe for how to serve.
Endpoint configuration and endpoint. The endpoint configuration says which model(s) to run, on which instance type, and how many. The endpoint is the live thing created from it: instances running your container behind an HTTPS address. Endpoints bill for every hour they exist, whether or not anyone calls them.
Instance types. Names like ml.m5.xlarge or ml.g5.xlarge describe the machine. The ml. prefix means a SageMaker-managed instance. The letter and number give the family and generation (m is general purpose, g and p are GPU families), and the suffix is the size. Bigger and GPU instances cost more per hour and, importantly, often have a default quota of zero in a new account.
Keep this stack in mind. When something fails, the useful question is always "which layer?". A permissions error lives in the control plane or the role. A Python traceback lives in the instance layer and shows up in the logs. A "file not found" for data lives in the storage layer.
Beyond these nouns there are many features you will meet later: Studio (the browser IDE), Pipelines (multi-step workflows), the Model Registry (versioned, approved models), Feature Store, Model Monitor, JumpStart (a catalogue of pretrained models), and HyperPod (large-scale training clusters). They are all built from the nouns above. Do not try to learn them yet.
- Without looking back, draw the four-box flow from the first section and label which boxes are temporary and which are permanent. Check yourself against the text: training instances are temporary, S3 objects are permanent, and an endpoint is temporary only if you delete it.
Setting up your AWS account safely
SageMaker needs an AWS account. If you already have one through an employer, ask whether you may use a sandbox account, because training jobs and endpoints can create meaningful bills and security teams care about who runs what.
For a personal account, the order of operations matters.
- Create the account and secure the root user. The root user is the email address you registered with. Turn on multi-factor authentication for it immediately, then stop using it. Never create access keys for the root user.
- Create a day-to-day identity. The modern recommendation is IAM Identity Center, which gives you a user that signs in through a browser and receives short-lived credentials. A plain IAM user with an access key also works for learning, but long-lived keys are the kind of secret that ends up in a public repository. For learning, give the identity broad rights, and tighten them once you understand what you need.
- Set a budget alert. In the Billing console, create a budget with an email alert at a small amount. This is a safety net for the day you forget an endpoint.
- Choose one region and stay in it. Pick the region closest to you or required by your employer, and use it for everything in this guide.
- Check quotas early. New accounts commonly have a quota of zero for GPU instance types (
ml.p*andml.g*families) for training and endpoints. Raising a quota is a request in the Service Quotas console and can take time, so file it before you need it. Everything in this guide runs on small CPU instances, which are usually available without a request, but verify in your own account.
Installing the tools and checking the setup
There is no SageMaker server to install. You need the AWS CLI, Python, and the SDK on your own machine.
The AWS CLI (version 2). On macOS, brew install awscli or the official installer package works. On Linux, download the zip for your architecture from AWS, unzip it, and run sudo ./aws/install. On Windows, use the MSI installer or winget install Amazon.AWSCLI. Confirm with:
aws --version
You should see a line starting with aws-cli/2.. If you see aws-cli/1., you have the older generation, and the configuration commands below differ.
Credentials. The preferred route is single sign-on:
aws configure sso
aws sts get-caller-identity
The first command walks you through naming a profile and signing in with your browser. The second asks AWS "who am I?" and prints your account number and identity ARN (an ARN, Amazon Resource Name, is the unique string identifying any AWS resource). If you chose a plain IAM user instead, run aws configure and paste the access key, secret key, and default region. If the identity call prints an account and ARN, your credentials work. If it prints an error about expired or missing credentials, sign in again before doing anything else.
If your laptop sits behind a corporate proxy or a security gateway that inspects TLS, the CLI may fail with a certificate error even though the browser works. That is a trust-store issue with the gateway's certificate authority, and the fix is to point the CLI at the exported CA certificate (aws configure set ca_bundle /path/to/ca.pem), not to disable verification.
Python and the SDK. Use Python 3.10, 3.11, or 3.12. Always work in a virtual environment so SageMaker's dependencies do not collide with other projects:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install --upgrade sagemaker boto3
pip show sagemaker
The last line should print version 3.something. To protect yourself against a future major release changing the API under you, pin it in real projects:
pip install "sagemaker>=3,<4"
If pip answers Could not find a version that satisfies the requirement sagemaker>=3, your Python is older than 3.10. Install a newer Python and recreate the virtual environment.
A final check that the pieces talk to each other:
import boto3
session = boto3.Session() # add profile_name="your-profile" if you used SSO
print(session.region_name)
print(session.client("sts").get_caller_identity()["Account"])
The first line should print the region you expect. If it prints None, no default region is configured, so set one with aws configure set region <your-region> or pass region_name= explicitly.
- Run
aws sts get-caller-identityand write down the account number. Then runaws sagemaker list-training-jobs --max-results 5. An empty list is a success: it proves your credentials and region work and that you have not run anything yet. - Run
pip show sagemakerand confirm the major version is 3.
The execution role, the idea that trips everyone
Here is the part of SageMaker that surprises beginners most. There are two different identities in every job. The first is you: the user or SSO identity whose credentials you configured. The second is the execution role, an IAM role that SageMaker assumes while running your job. You, the caller, hand the role to SageMaker when you start the job. SageMaker then uses that role to read the S3 data, pull the container from ECR, and write logs to CloudWatch.
This design is deliberate. The job runs on machines you do not control, long after you closed your laptop, and it needs its own limited permissions rather than your personal ones. But it means two sets of permissions can each be wrong, and the errors look similar.
Confused
My user is an administrator, so the training job can read my S3 bucket.
Correct
My user can start the job. The execution role is what reads S3, so the role needs access to the bucket.
Two permissions have to line up.
- The role must trust SageMaker. Its trust policy must allow the principal
sagemaker.amazonaws.comto assume it. If this is wrong you will see errors about not being able to assume the role. - You must be allowed to pass the role. The action
iam:PassRoleis the permission to hand a role to a service. If your identity lacks it for that role, the start-job call fails with an access-denied message mentioningiam:PassRole.
The easiest way to create a role in a learning account is the IAM console: choose Roles, create a role, pick SageMaker as the trusted use case, and attach the managed policy AmazonSageMakerFullAccess. Name it something obvious such as SageMakerLearningRole. Note the ARN it gives you, which looks like arn:aws:iam::123456789012:role/SageMakerLearningRole.
There is a catch hiding in that managed policy, documented on the AWS roles page: AmazonSageMakerFullAccess grants S3 access only to buckets and objects whose names contain SageMaker, Sagemaker, sagemaker, or aws-glue. Name your learning bucket something like sagemaker-yourname-learning-<random digits> and the default policy just works. If you use a bucket with a different name, the job fails with an AccessDenied when it tries to read, and the fix is to add your own policy granting that bucket to the role.
If you work inside SageMaker Studio, which you will meet shortly, the role attached to your space is available automatically. When you run code on your laptop there is no attached role, so you pass the ARN explicitly, as the examples below do.
Your first project: data in S3, a training job, a model artifact
Time to build the spine of the service. You will create a bucket, upload a small dataset, and start a training job.
Step 1: a bucket. Bucket names are global across all AWS accounts, so pick a unique one, and keep sagemaker in it for the managed policy reason above:
aws s3 mb s3://sagemaker-yourname-learning-4821 --region <your-region>
aws s3 cp train.csv s3://sagemaker-yourname-learning-4821/train/train.csv
aws s3 ls s3://sagemaker-yourname-learning-4821/train/
Replace the name and region with your own. The last command should list your file. Use any small CSV you have. The content is not the point; the plumbing is.
Step 2: understand the training contract. A training job runs a container, and the container must follow SageMaker's conventions for where things live. Data you declare as a channel named training appears inside the container under /opt/ml/input/data/training. Whatever your code writes into /opt/ml/model is packaged and uploaded to S3 as model.tar.gz when the job finishes successfully. Hyperparameters and job metadata are provided to your code as configuration, and anything else you write to /opt/ml/output is saved on failure for debugging. Your training script reads from the first path and writes to the second, and does not need to know anything about S3.
Step 3: start the job with ModelTrainer. In version 3 the class for training is ModelTrainer. The following is the minimal pattern from the SDK's own README:
from sagemaker.train import ModelTrainer
from sagemaker.train.configs import InputData
trainer = ModelTrainer(
training_image="<account>.dkr.ecr.<region>.amazonaws.com/my-training-image:latest",
role="arn:aws:iam::123456789012:role/SageMakerLearningRole",
)
train_data = InputData(
channel_name="training",
data_source="s3://sagemaker-yourname-learning-4821/train",
)
trainer.train(input_data_config=[train_data])
Read it from the top. training_image is the container to run, which must be an image your execution role can pull: either an AWS-provided framework image or one you built and pushed to ECR. role is the execution role ARN. InputData describes one channel: its name, which determines the directory under /opt/ml/input/data/, and where its files live in S3. trainer.train(...) creates the training job and, by default, streams its logs to your terminal while it runs.
ModelTrainer also accepts further configuration, such as pointing at a local source directory and entry script instead of baking code into the image, choosing the compute instance type and count, and setting hyperparameters and stopping conditions. The SDK documentation on the readthedocs site lists the configuration classes for these. Check it for the exact names in your installed version, because this part of the API is still evolving in the 3.x line, and a name that is right today may gain companions tomorrow.
Step 4: watch it. While the job runs, and afterwards, you can ask AWS about it from any terminal:
aws sagemaker list-training-jobs --max-results 5
aws sagemaker describe-training-job --training-job-name <job-name>
list-training-jobs shows recent jobs and their statuses. describe-training-job returns a large JSON record: its status, the instance type, the S3 location of the output model (ModelArtifacts), how long it ran, and, if it failed, a FailureReason string. Always read FailureReason before anything else. It is often a one-line statement of exactly what went wrong.
A job moves through statuses you can name: InProgress while it runs (with a secondary status showing starting, downloading, training, and uploading), then Completed, Failed, or Stopped. A fast failure within a minute or two usually means permissions, a missing image, or a missing quota. A failure after several minutes of training usually means your own code crashed.
- Create the bucket and upload any small file. Confirm with
aws s3 ls. - Start a training job with the pattern above, using an image and script you control. Run
aws sagemaker describe-training-jobon it and find the three fieldsTrainingJobStatus,ModelArtifacts, and, if it failed,FailureReason. - When it completes, download the artifact with
aws s3 cpfrom theS3ModelArtifactsURI and unpack it withtar -xzf model.tar.gz. You are looking at exactly what your code saved to/opt/ml/model.
Reading logs and failures
When a job fails, the answer is almost always in the logs, and SageMaker sends them to Amazon CloudWatch Logs. Training jobs write to the log group /aws/sagemaker/TrainingJobs, with one log stream per job and instance. Processing jobs use /aws/sagemaker/ProcessingJobs, and endpoints use /aws/sagemaker/Endpoints/<endpoint-name>. You can read them in the CloudWatch console, or tail them from the CLI:
aws logs tail /aws/sagemaker/TrainingJobs --follow
You will see your own print output and any Python traceback. The classic failure is the line AlgorithmError with ExecuteUserScriptError, which means that SageMaker started your container correctly and your script then crashed. The real explanation is the traceback a few lines above it in the log.
Use a simple triage order, because it saves hours.
- Did the job start at all? If creation fails instantly with an error in your terminal and no job appears in
list-training-jobs, it is a permission, quota, or parameter problem raised by the API call. Read the error text; it names the cause. - Did it fail before your code ran?
describe-training-jobshowsFailureReason. Image pull failures and S3 download failures appear here. - Did your code run and crash? Read the CloudWatch log from the bottom upward until you find the first traceback.
The SDK wraps job failures in an exception called UnexpectedStatusException, with text like Error for Training job X: Failed. Reason: .... That is only the messenger. The reason it prints, plus the CloudWatch log, is what you act on.
Deploying the model: from artifact to endpoint
A model artifact is a file. To get predictions you need a running process that loads it and answers requests. In SageMaker that is the endpoint. The lifecycle has three resources, created in order, and understanding them one at a time is worth the extra lines. The high-level SDK class in version 3, ModelBuilder, wraps these steps, but the underlying calls are the same, and the exact ModelBuilder signature is best checked in the SDK documentation for your installed version. The boto3 version below is stable and shows exactly what exists.
import time
import boto3
region = "<your-region>"
sm = boto3.client("sagemaker", region_name=region)
name = f"demo-{int(time.time())}"
role = "arn:aws:iam::123456789012:role/SageMakerLearningRole"
sm.create_model(
ModelName=name,
ExecutionRoleArn=role,
PrimaryContainer={
"Image": "<account>.dkr.ecr.<region>.amazonaws.com/my-inference-image:latest",
"ModelDataUrl": "s3://sagemaker-yourname-learning-4821/output/model.tar.gz",
},
)
sm.create_endpoint_config(
EndpointConfigName=name,
ProductionVariants=[{
"VariantName": "AllTraffic",
"ModelName": name,
"InstanceType": "ml.m5.xlarge",
"InitialInstanceCount": 1,
}],
)
sm.create_endpoint(EndpointName=name, EndpointConfigName=name)
waiter = sm.get_waiter("endpoint_in_service")
waiter.wait(EndpointName=name)
print("endpoint ready:", name)
Walk through it. create_model registers the pairing of an inference image and your artifact; the image here is the container that will serve, which must follow the serving contract below. create_endpoint_config says "run that model on one ml.m5.xlarge instance" and names the variant AllTraffic. A variant is a unit of traffic, and later you can split traffic across several. create_endpoint starts the real machines. The waiter polls until the status is InService, which typically takes several minutes, since an instance must start and your container must load the model.
The serving contract is the other half of the container rules. An inference container must run a web server on port 8080 that answers GET /ping with a healthy status and POST /invocations with predictions. If the server does not answer /ping in time, endpoint creation fails with the message that the primary container for production variant AllTraffic did not pass the ping health check, and the endpoint log group says why. AWS's prebuilt framework images already implement this contract; if you build your own container, you implement it yourself.
Once the status is InService, call it through the runtime client, which is a separate client from the control-plane one:
import boto3
rt = boto3.client("sagemaker-runtime", region_name="<your-region>")
response = rt.invoke_endpoint(
EndpointName="<endpoint-name>",
ContentType="text/csv",
Body="5.1,3.5,1.4,0.2",
)
print(response["Body"].read().decode("utf-8"))
ContentType tells your container how to parse the body, so it must match what your serving code expects. From the CLI the equivalent is:
aws sagemaker-runtime invoke-endpoint --endpoint-name <name> \
--content-type application/json --body fileb://in.json out.json
cat out.json
The fileb:// prefix sends the file as raw bytes, which the CLI otherwise would try to interpret as base64 text.
There are several endpoint types, and it helps to know the list even though you will start with the first. Real-time endpoints are the always-on kind described above, built for steady, low-latency traffic. Serverless inference starts capacity on demand and suits spiky, tolerant traffic. Asynchronous inference takes requests through a queue and writes results to S3, suited to large payloads and long processing. Batch transform runs a model over a whole dataset offline, without a persistent endpoint. For learning, use real-time, and delete it when you are done.
delete-endpoint stops the charge.
Cleaning up
Cleanup is three deletions, mirroring the three creations, and the order does not matter much, but the endpoint is the one that costs money.
aws sagemaker delete-endpoint --endpoint-name <name>
aws sagemaker delete-endpoint-config --endpoint-config-name <name>
aws sagemaker delete-model --model-name <name>
Verify nothing is left running:
aws sagemaker list-endpoints
The list should be empty, or show only endpoints you intend to keep. Training jobs do not need deletion; they stop billing when they finish, and their records are kept for reference. S3 objects bill for storage at a low rate, so delete the training and output files in your learning bucket when you no longer need them with aws s3 rm --recursive. Studio apps and notebook instances, covered next, also bill while running.
Make this a reflex: whenever you create an endpoint, write its deletion command at the same moment, ideally in the same script wrapped in a try ... finally.
- Create an endpoint with the script above, call it once, and delete it with the three commands. Confirm
list-endpointsis empty. - In the Billing console, look at your current month's cost by service. Find SageMaker in the list. This is where a forgotten endpoint would show up first.
Where you write code: Studio, notebook instances, and your laptop
So far you have run commands from your own machine. SageMaker also provides hosted places to write code, and you should know what they are and what they cost.
SageMaker Studio is the browser-based environment. It is organised around a domain: a per-account, per-region environment holding user profiles, shared storage, and network and IAM settings. Inside a domain, each person has a user profile, and each profile runs its tools in spaces. A space can host JupyterLab, a Code Editor based on Code-OSS (the open-source core of VS Code), or RStudio. SageMaker Canvas is a separate no-code interface for analysts. Studio has had a long-standing classic version and a newer experience, and AWS documentation steers new domains to the newer one, so if a tutorial's screenshots do not match your console, that is why. A newer, broader environment called SageMaker Unified Studio combines data and AI work and belongs to the umbrella platform mentioned earlier.
Notebook instances are the older option: a single managed virtual machine running Jupyter. They are still supported but AWS steers new users to Studio.
What matters for a beginner is the cost model. A running JupyterLab or Code Editor space is backed by an instance, and that instance bills for every second it is running, including overnight if you close the browser tab without stopping it. Closing the tab does not stop the app. Use the stop control for the app, or configure an idle-shutdown lifecycle script so unattended instances shut themselves down.
Another important property: inside Studio, credentials and the execution role come from the space automatically, so you do not paste ARNs the way you do on a laptop. That is convenient, but it also means code that works in Studio may fail on your laptop for want of an explicit role, which is an easy mistake when moving a script into CI.
The everyday commands, grouped by what you are trying to do
You now have the structure. This section collects the commands you will use constantly, grouped by intent. Replace the placeholders with your own names.
Find out what exists.
aws sagemaker list-training-jobs --max-results 5
aws sagemaker list-endpoints
aws sagemaker list-models
aws sagemaker list-endpoint-configs
Inspect one thing in detail.
aws sagemaker describe-training-job --training-job-name <name>
aws sagemaker describe-endpoint --endpoint-name <name>
For an endpoint, the fields to read are EndpointStatus (Creating, InService, Updating, Failed, Deleting) and, on failure, FailureReason. Add --query to pull out one field, for instance --query EndpointStatus --output text.
Stop something that is running.
aws sagemaker stop-training-job --training-job-name <name>
Call a model.
aws sagemaker-runtime invoke-endpoint --endpoint-name <name> \
--content-type application/json --body fileb://in.json out.json
Clean up.
aws sagemaker delete-endpoint --endpoint-name <name>
aws sagemaker delete-endpoint-config --endpoint-config-name <name>
aws sagemaker delete-model --model-name <name>
Move data.
aws s3 cp local.csv s3://<bucket>/train/local.csv
aws s3 sync ./data s3://<bucket>/data
aws s3 ls s3://<bucket>/ --recursive
Read logs.
aws logs tail /aws/sagemaker/TrainingJobs --follow
aws logs tail /aws/sagemaker/Endpoints/<name> --follow
Two habits make all of these safer. Pass --region explicitly or set a default, so you never look in the wrong region and conclude a resource vanished. And use --profile if you keep multiple accounts, so a command intended for your sandbox never reaches a work account.
When you move from the terminal into Python, two clients cover almost everything. boto3.client("sagemaker") mirrors the aws sagemaker commands one to one, with method names in snake case (list_training_jobs, describe_endpoint). boto3.client("sagemaker-runtime") mirrors aws sagemaker-runtime, with invoke_endpoint for ordinary calls and invoke_endpoint_with_response_stream for streaming responses. Learning one makes the other easy, which is why the AWS CLI reference (linked in Sources) is worth bookmarking: any field name you see there is the same field name in Python.
Choosing instances and keeping cost under control
Beginners over-think instance choice and under-think duration. For a first model on a small tabular dataset, an ml.m5.xlarge (a general-purpose CPU machine) is plenty. Move to GPU families, ml.g5.xlarge being a common starting point, only when you train deep learning models that need them, and expect to request a quota first.
Billing works differently per resource, and knowing the shape of each tells you where to look.
- Training and processing jobs bill per instance-second while running. A ten-minute job on one instance costs ten minutes. They stop billing when they finish.
- Notebook instances and Studio apps bill per instance while running, whether or not you are using them.
- Real-time endpoints bill per instance-hour for as long as they exist.
- Serverless endpoints bill for the compute used by requests, not for idle time, which is why they suit occasional traffic.
- Storage and data transfer bill separately: S3 storage, volumes attached to instances, and traffic leaving AWS or crossing zones.
AWS offers managed spot training, which runs your job on spare capacity at a discount, with AWS stating savings of up to 90 percent, in exchange for the possibility of interruption. It needs checkpointing so a job can resume, which makes it a Mid-level topic. Savings Plans for SageMaker AI offer lower rates against a one- or three-year commitment, which only makes sense once you have steady usage.
Your beginner safeguards are simple: a budget alert, a rule that every endpoint has a deletion command written beside its creation command, stopped Studio apps at the end of a session, and a monthly look at the Billing console. The most common bill shocks in this service are a forgotten endpoint, a GPU Studio app left running over a weekend, and a NAT gateway that processes large data transfers for jobs running in a private network. The last of these is why, as you move toward production, cost reviews include networking and not only SageMaker line items.
- Create a budget of a small amount in the Billing console with an email alert at 80 percent.
- List your Studio apps and notebook instances in the console and confirm none are running when you are not using them.
Configuration and common errors, and how to read them
Most beginner time is lost to a short list of errors. Each is worth recognising on sight. The exact wording of messages can shift between releases, so match on the key phrases.
ResourceLimitExceeded ... The account-level service limit 'ml.p3.2xlarge for training job usage' is 0 Instances. Your account has no quota for that instance type. This is a quota, not a bug. Open Service Quotas and request an increase, or use a different instance type or region. New accounts commonly hit this on GPU instances.
AccessDeniedException ... is not authorized to perform: iam:PassRole on resource: ...role/.... You, the caller, are not allowed to hand that role to SageMaker. Add iam:PassRole for that role to your identity, and make sure the role's trust policy lists sagemaker.amazonaws.com.
ValidationException: Could not assume role or Could not find role. The role ARN is mistyped, in the wrong account, or its trust policy does not trust SageMaker. Copy the ARN from the console again and check the trust relationship.
CapacityError: Unable to provision requested ML compute capacity. AWS had no machines of that type free right now. Retry later, try a different instance type, or try another region. This is not about your account.
AlgorithmError with ExecuteUserScriptError. Your script crashed. Read the traceback in CloudWatch. Out-of-memory messages (including CUDA out of memory) mean a smaller batch size or a larger instance.
UnexpectedStatusException: Error for Training job X: Failed. Reason: .... The SDK's wrapper around any failed job. Read the reason after the colon, and FailureReason from describe-training-job.
AccessDenied when reading S3. The execution role cannot read that bucket. If your bucket name lacks sagemaker, the managed policy does not cover it. Attach a policy for your bucket to the role.
Endpoint creation fails with did not pass the ping health check. Your serving container did not answer /ping on port 8080 in time. Read the endpoint's CloudWatch log group. Common causes are a missing or wrong model artifact, a container built for the wrong CPU architecture, or a server that fails on startup. The default startup timeout is 600 seconds.
ModelError ... Received server error (500) from primary. The endpoint is up but your inference code raised an error on that request. Read the endpoint log group.
ValidationError ... Endpoint NAME of account ACCOUNT not found. Wrong region, a typo, or the endpoint was deleted.
ImportError: cannot import name 'Estimator' from 'sagemaker.estimator'. Version 2 code on a version 3 install. Migrate to ModelTrainer, or create a separate environment with pip install "sagemaker==2.*".
An image pull that hangs. If you run jobs inside a private network without a route to ECR and S3, the container cannot download its image. This is a networking setup problem; VPC endpoints or a NAT route resolve it.
A final configuration note: the lowest-friction way to keep your settings consistent is environment variables or a named profile, for example export AWS_PROFILE=sandbox and export AWS_DEFAULT_REGION=<your-region>. Both the CLI and boto3 honour them, so scripts do not need hard-coded values.
Do not put secrets such as API keys in hyperparameters or environment variables of a training job. The job description, including those values, is visible to anyone who can run describe-training-job. Fetch secrets at runtime from AWS Secrets Manager using the execution role instead.
Putting it all together
Here is one small end-to-end project that exercises everything above. It trains a model, deploys it, calls it once, and cleans up. Treat it as a checklist you run through in one sitting of about an hour, being careful with the deletions.
- Prepare the account. Confirm
aws sts get-caller-identityworks, your region is set, a budget alert exists, and your execution role hasAmazonSageMakerFullAccesswith a trust policy forsagemaker.amazonaws.com. - Create the bucket and upload data. A bucket with
sagemakerin its name, and a small CSV undertrain/. - Run local sanity tests. Run your training script on your laptop against a few rows and confirm it writes a model file where your code expects.
- Build and push the images. A training image and an inference image, or the prebuilt AWS framework images that already honour both contracts, available to your role in ECR.
- Train. Use
ModelTrainerwith anInputDatachannel pointing ats3://.../train. Watch the log stream. If it fails, follow the triage order: did it start, did it fail before your code, did your code crash. - Find the artifact.
describe-training-job, thenaws s3 lson the model location, and download and unpackmodel.tar.gzonce to see what is inside. - Deploy. Run
create_model,create_endpoint_config, andcreate_endpoint, then wait for theendpoint_in_servicewaiter. - Invoke. Send one request through
sagemaker-runtime, and read the response. - Read the endpoint logs. Open
/aws/sagemaker/Endpoints/<name>in CloudWatch and find your request. - Clean up. Delete the endpoint, the endpoint config, and the model, confirm
list-endpointsis empty, stop any Studio apps, and remove the S3 objects you no longer need.
The deletion step is the one that marks the difference between a beginner who has run SageMaker and one who can be trusted with an account. Write the deletion commands first, put them in the same script as the creation, and run them in a finally block so a crash in step eight still removes the endpoint.
import boto3
sm = boto3.client("sagemaker", region_name="<your-region>")
name = "demo-endpoint"
try:
# create_model, create_endpoint_config, create_endpoint, wait, invoke ...
pass
finally:
for call, kwargs in [
(sm.delete_endpoint, {"EndpointName": name}),
(sm.delete_endpoint_config, {"EndpointConfigName": name}),
(sm.delete_model, {"ModelName": name}),
]:
try:
call(**kwargs)
except sm.exceptions.ClientError as err:
print("cleanup skipped:", err)
The try inside the loop matters: if the failure happened before the endpoint was created, its deletion will report that it does not exist, and you still want the remaining deletions to run.
- Run the whole checklist once, end to end, then repeat it with a deliberate mistake: use a bucket name without
sagemakerin it, and watch theAccessDeniedappear. Fix it by attaching a bucket policy to the role. Reproducing one failure on purpose teaches more than ten successes.
What you can now do, and what comes next
You can now describe SageMaker AI in two sentences, name its core nouns, and say which of them bill while idle. You can set up an account safely and verify the CLI and SDK. You understand why the execution role exists and what the two permission checks, trust and iam:PassRole, protect. You can train with ModelTrainer, find the result in S3, deploy through the three-step lifecycle, call the endpoint, and delete everything. You can read the main error messages and know which layer to look in. And you know that older tutorials use version 2 names which no longer exist in version 3.
The Mid-level guide takes these same pieces and makes them production-shaped: script-mode training with your own code, managed spot training with checkpoints, hyperparameter tuning, serverless and asynchronous endpoints, Pipelines, the Model Registry, and experiment tracking with MLflow. The Senior guide covers the platform view: auto scaling and scale to zero, blue/green deployments with rollback alarms, network isolation and encryption, multi-account design, and cost governance.
Natural next steps in this catalogue are experiment tracking with MLflow, packaging models for serving with BentoML, the Kubernetes-native serving route with KServe, using foundation models through Amazon Bedrock, and pretrained models from Hugging Face, which SageMaker supports through prebuilt containers and its JumpStart catalogue. If your team already schedules work with a general orchestrator, the Airflow guide shows the alternative to SageMaker Pipelines.
Sources
- What is Amazon SageMaker AI
- Amazon SageMaker AI documentation root
- How to use SageMaker AI execution roles
- AWS managed policies for SageMaker AI
- SageMaker AI API reference
- AWS CLI reference for sagemaker
- SageMaker AI endpoints and quotas
- Amazon SageMaker AI pricing
- SageMaker Python SDK on GitHub (README and changelog)
- sagemaker on PyPI
- SageMaker Python SDK documentation