Skip to content
Back to student guides
Vertex AIMLOpsCloud ML platforms3 levels103 sectionsCovers google-cloud-aiplatform 2.3

The Complete Vertex AI Guide

Build, train and deploy models on Google Cloud with Vertex AI. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
17sections
20examples

This is part one of three. It covers everything you need to do real work with Vertex AI: set up a Google Cloud project safely, train a model on managed hardware, register it, put it behind an endpoint, ask it for predictions, run a batch job, call a Gemini model, and clean up so you are not billed for things you forgot. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. Cloud platforms are learned by clicking, failing, reading the error, and trying again, and every error message in this guide is one you will probably meet.

The product has two names In 2026 Google began folding Vertex AI into a broader product called the Gemini Enterprise Agent Platform. The official documentation, the Python package, the gcloud ai commands and the resource names you will see below are still the ones this guide uses, and tutorials, Stack Overflow answers and job posts will use either name for the same platform. When you read "Vertex AI" here, and "Agent Platform" in a newer page, think of one thing. Where the rename changed code you would actually write, this guide says so.

What Vertex AI is, and the problem it solves

Vertex AI is Google Cloud's managed platform for machine learning. It gives you hardware to train models on, a place to keep the models you produced, a way to serve them over HTTPS, a way to orchestrate multi-step workflows, and access to Google's Gemini family of generative models. You use it through a web console, a command-line tool called gcloud, a Python SDK, or plain REST calls. They all drive the same underlying service.

To see why it exists, picture what a team does without a platform. A data scientist trains a model on a laptop and saves a model.pkl file. Someone must now rent a server, install the right Python version and library versions, write a small web server around the model, open a port, add authentication, make it restart when it crashes, add more servers when traffic grows, and keep track of which file is the current model. Each of those steps is a small engineering project, and each is done differently on every team. Then the data scientist wants to retrain on a GPU for an hour, and someone has to rent one, install drivers, and remember to switch it off.

Vertex AI takes over the repetitive plumbing. You hand it your training script and say "run this on one machine of this size". You hand it a trained model and say "serve this, between one and three replicas". Google provisions the machines, runs your code, collects the logs, and bills you per unit of compute used. You still decide what the model is and whether it is any good. The platform decides where it runs.

DATACloud Storage, BigQuery
→
TRAININGyour code, managed machines
→
MODEL REGISTRYversioned models
→
ENDPOINTlive predictions

The diagram is the spine of the platform. Almost every feature you will meet later, pipelines, experiments, monitoring, feature stores, is a way of automating, recording, or watching one of those four stages.

What people use it for:

🏋️

Training without owning hardware

Run a script on a CPU or GPU machine for exactly as long as it takes, then let the machine disappear.

🚀

Serving models

Turn a saved model into an HTTPS endpoint that scales between a minimum and maximum number of replicas.

🔁

Repeatable pipelines

Describe a workflow as steps, run it on a schedule, and have every run recorded.

✨

Generative AI

Call Gemini and other hosted models through the same project, billing and permissions as everything else.

Who it is not for is worth saying early. If you want a model that already exists to answer questions, you may not need to train anything at all, and the generative section near the end of this guide is your shortest path. If you have a notebook experiment and no intention of serving it, the managed notebook feature may be all you need. And if your employer is on another cloud, you will meet the same ideas under other names; the SageMaker guide covers the Amazon equivalent, and nearly every concept transfers.

Try it
  1. Open the Google Cloud console in a browser and find the project selector at the top of the page. Note whether you already have a project.
  2. Search the console's top bar for "Vertex AI" and then for "Agent Platform". See which name your console shows, and write down the left-hand menu entries you recognise from the diagram above.

The mental model: five nouns

Vertex AI has hundreds of pages in its documentation, but a beginner needs only five ideas. Learn them properly and the rest of the platform becomes variations.

Project. Every Google Cloud resource lives inside a project. A project is the unit that holds your billing account link, your enabled services, your permissions, and your quotas. When you delete a project, everything inside it is deleted, which makes projects a convenient sandbox. Every project has a human-readable name, a globally unique project ID such as mlops-lab-123456, and a numeric project number. Code and commands nearly always want the ID.

Location, also called region. Most Vertex AI resources are regional: a model uploaded to us-central1 exists only there, and an endpoint in europe-west4 serves from there. The region is where your data is processed and stored, which matters for latency, for price, and for legal reasons. Not every feature exists in every region, and the official locations page is the authority for what is available where. This is the single most common source of "it worked yesterday in a different region" confusion, so get in the habit of treating the region as part of a resource's name.

Training job. A training job is a request that says "run this code on these machines". The platform starts the machines, runs your code, streams its logs, and shuts the machines down. There are two broad styles. AutoML needs no code: you supply data and the platform picks the model. Custom training runs code you wrote, either in a container Google provides for popular frameworks or in a container you build yourself. This guide teaches custom training because it is the one you will use at work.

Model. After training, you register the result as a Model in the Model Registry. A Model is not a file; it is a record that points to your saved artifact in Cloud Storage and to a serving container, the program that knows how to load that artifact and answer requests. A model can have versions, and versions can carry aliases such as default or champion. Think of the registry as a labelled shelf.

Endpoint. An endpoint is a stable address that serves predictions. You deploy a model to an endpoint, which means placing it on one or more machine replicas that sit behind the endpoint, and you can split traffic between several deployed models. An endpoint with a model on it is running machines that bill by the hour, whether or not anyone is sending traffic. Hold on to that sentence; it is the most expensive thing in this guide.

Say it in one line A training job produces files; a model says what those files are and how to serve them; an endpoint runs that model for callers. If you can explain that triangle, you understand the platform.

Two supporting nouns will appear constantly. Cloud Storage (GCS) is Google's object store, and its containers are called buckets, addressed with gs://bucket-name/path. Vertex AI uses a bucket as a staging area for your code and as the place training jobs write their outputs. Service accounts are non-human identities. When a training job runs, it runs as a service account, and the job can touch only what that account is allowed to touch. Later sections show where this bites.

Try it
  1. Without looking back, draw the four-stage diagram and label the noun that belongs on each arrow: job, model, endpoint.
  2. Pick the region closest to your employer's users or data-protection rules. Write it down; you will use it throughout.

Getting set up: account, project, billing, API

Vertex AI is a service, not a program, so "installing" it means preparing your Google Cloud account. Doing the setup carefully once saves an afternoon of confusing permission errors.

Create or choose a project. In the console, open the project selector and create a new project for this course. Use a fresh project rather than an existing work project: if something goes wrong or you forget to clean up, you can delete the whole project and be certain nothing is still billing.

Link a billing account. Vertex AI is a paid service. New Google Cloud accounts have commonly received free trial credits, but the offer changes, so read the current terms on the sign-up page rather than assuming. Billing must be linked to the project before the API will work. Whatever credits you have, set a budget alert in the Billing section so an email arrives when spend crosses an amount you pick. A budget alert does not stop spending by itself, but it is your early warning.

Enable the API. Google Cloud services are off by default in a new project. Vertex AI's is called aiplatform.googleapis.com. You can enable it from the console by searching for Vertex AI and pressing the enable button, or from a terminal after you install the command-line tool in the next section.

Authenticate. There are two separate kinds of login, and mixing them up is the classic beginner stumble. One is for the gcloud command-line tool itself. The other, called Application Default Credentials (ADC), is what the Python SDK and other libraries look for when your code calls Google. You run a command for each, and the next section does so.

Do not paste a key file into a repository Older tutorials tell you to download a service-account JSON key and point GOOGLE_APPLICATION_CREDENTIALS at it. A key file is a password that never expires. If it lands in a public repository, automated scanners find it within minutes. On your own laptop, use gcloud auth application-default login instead, which keeps short-lived credentials in your home directory and never gives you a file to leak. This repository is public; treat every project the same way.

Your own user account, when you created the project, is its owner and can do everything. That is convenient for learning and dangerous at work. In a company project you will usually receive a narrower role, such as Vertex AI User (roles/aiplatform.user), which allows you to run jobs and call endpoints but not to administer the project. If you see a 403 permission error, the question to ask is always "which role does the identity running this need?", and we return to that in the errors section.

Try it
  1. Create a new project named for this course, link billing, and create a budget alert at an amount you are comfortable losing.
  2. In the console, open APIs & Services and enable the Vertex AI API for the project. Note how long it takes.

Installing the tools and checking the setup

You need two things on your machine: the gcloud command-line tool, and the Python SDK. Neither is large.

Install gcloud. The official installation page covers every platform. On macOS with Homebrew, brew install --cask google-cloud-sdk works; on Linux the page offers apt and yum repositories and a tarball; on Windows it offers an installer. Open a new terminal afterwards so gcloud is on your path, then run the initial configuration:

BASH
gcloud init
gcloud auth login
gcloud auth application-default login
gcloud config set project PROJECT_ID
gcloud services enable aiplatform.googleapis.com

Read what each line does, because they are not interchangeable. gcloud init walks you through choosing an account and a default project. gcloud auth login authorizes the command-line tool. gcloud auth application-default login creates the Application Default Credentials that your Python code will find automatically. gcloud config set project sets which project later commands act on; replace PROJECT_ID with the ID, not the display name. The last line enables the Vertex AI API if you have not already.

Install the Python SDK. The current SDK needs Python 3.10 or newer; Python 3.9 and older are deprecated. Always use a virtual environment so the SDK's many dependencies stay out of your system Python.

BASH
python -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
pip install google-cloud-aiplatform google-genai

google-cloud-aiplatform is the classic machine-learning SDK: it holds CustomTrainingJob, Model, Endpoint and the other nouns from the previous section. google-genai is the Google Gen AI SDK, used for Gemini and other generative models. They are separate packages on purpose, and the reason matters, so read the next callout.

The SDK changed shape in 2026 Version 2.0 of google-cloud-aiplatform was a breaking release. The generative modules that used to live inside it, vertexai.generative_models and its siblings, were deprecated in June 2025 and scheduled for removal in June 2026, and generative code now belongs in google-genai. The agent tooling moved into a separate package, google-cloud-agentplatform. Classic ML code, the aiplatform.init(), CustomTrainingJob and Model.deploy() style used in this guide, stays in google-cloud-aiplatform. If an old tutorial imports from vertexai.generative_models import GenerativeModel and fails, that is why. The package releases roughly weekly, so pin a version in any project you intend to keep.

Check the setup. Run these in order. Each checks a different layer, so when one fails you know where to look.

BASH
gcloud config list
gcloud auth list
gcloud ai models list --region=us-central1

The first shows your active project. The second shows which account is active. The third calls Vertex AI itself; on a new project it should print Listed 0 items. and not an error. Then check the SDK from Python:

PYTHON
from google.cloud import aiplatform

aiplatform.init(project="PROJECT_ID", location="us-central1")
print(aiplatform.Model.list())

An empty list [] means everything works. If you see DefaultCredentialsError, you skipped gcloud auth application-default login. If you see a 403 that mentions the API being disabled, run the enable command from the setup section.

Try it
  1. Run the five setup commands, then the three check commands. Fix any error before moving on.
  2. In a virtual environment, run pip show google-cloud-aiplatform and note the version. Write it in a requirements.txt with an exact pin.

Cloud Storage: the bucket every job needs

Before the first training job, make a bucket. Vertex AI uses Cloud Storage as the shared disk between your laptop and its machines: you upload data there, the platform stages your code there, and trained model files are written back there.

Bucket names are globally unique across all of Google Cloud, so choose something with your project in it. Create the bucket in the same region as your Vertex AI resources. A bucket in a different location from your job usually works, but it adds latency and egress charges, and some operations complain.

BASH
export REGION=us-central1
export BUCKET=gs://PROJECT_ID-vertex-lab
gcloud storage buckets create $BUCKET --location=$REGION
gcloud storage ls

Upload a small file to prove access, then read it back:

BASH
echo "hello" > hello.txt
gcloud storage cp hello.txt $BUCKET/test/hello.txt
gcloud storage ls $BUCKET/test/

The gs:// prefix is how every Vertex AI argument refers to storage. Folders in a bucket are a convention, since object stores have just names with slashes in them, but the console and ls display them as folders. You will pass these paths as the staging bucket, as training data locations, and as model artifact locations.

For the data itself, Vertex AI can also read straight from BigQuery, Google's warehouse, which is often where the company's tabular data already lives. For this guide we will generate data inside the training script so that nothing depends on a dataset you do not have.

Try it
  1. Create your bucket in your chosen region, upload a text file, and list it.
  2. In the console's Cloud Storage page, find the bucket's location and confirm it matches the region you wrote down earlier.

Your first training job

This is the section the rest of the guide leans on. We will train a very small model, a scikit-learn classifier, on managed hardware. The model is deliberately boring. What matters is that you see the whole shape of a Vertex AI job: a script, a container, a machine, and an artifact.

Step 1. The training script

Create train.py. It builds a toy dataset, trains a classifier, evaluates it, and saves the model to a location Vertex AI gives it through an environment variable.

train.py
import os
import joblib
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=2000, n_features=3, n_informative=3,
                           n_redundant=0, random_state=7)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=7)

model = LogisticRegression().fit(X_train, y_train)
print("accuracy:", model.score(X_test, y_test))

model_dir = os.environ.get("AIP_MODEL_DIR", "./model")
if model_dir.startswith("gs://"):
    local_dir = "/tmp/model"
else:
    local_dir = model_dir
os.makedirs(local_dir, exist_ok=True)
joblib.dump(model, os.path.join(local_dir, "model.joblib"))
print("saved to", local_dir)

Two details deserve explanation. The script prints the accuracy because anything your script prints to standard output becomes a log line you can read in the console, and a training script that prints nothing is a training script you cannot debug. And it saves to a directory taken from AIP_MODEL_DIR, an environment variable that Vertex AI sets for jobs created through the SDK's training-job classes. The platform decides where outputs go and tells your code, rather than your code hardcoding a path.

The exact way the saved file reaches Cloud Storage varies with the container and with how you submit the job. The simplest reliable approach for a beginner is the SDK's CustomTrainingJob, which handles that copy for you when you use a prebuilt training container, so we use it below. If you instead write your own container, you are responsible for uploading the model file to the storage location the job gives you, and the Mid-level guide covers it.

Run the script locally first. This catches every bug that has nothing to do with the cloud, and it is free.

BASH
pip install scikit-learn joblib numpy
python train.py

You should see an accuracy line and a "saved to" line. If this fails, do not submit it to the cloud; the cloud will only fail slower and charge you for it.

Step 2. Submit it to Vertex AI

submit_train.py
from google.cloud import aiplatform

PROJECT = "PROJECT_ID"
REGION = "us-central1"
BUCKET = "gs://PROJECT_ID-vertex-lab"

aiplatform.init(project=PROJECT, location=REGION, staging_bucket=BUCKET)

job = aiplatform.CustomTrainingJob(
    display_name="train-sklearn-toy",
    script_path="train.py",
    container_uri="us-docker.pkg.dev/vertex-ai/training/sklearn-cpu.1-0:latest",
    requirements=["joblib"],
    model_serving_container_image_uri="us-docker.pkg.dev/vertex-ai/prediction/sklearn-cpu.1-0:latest",
)

model = job.run(replica_count=1, machine_type="n1-standard-4")
print(model.resource_name)

Read it top to bottom. aiplatform.init sets the defaults for every later call: which project, which region, and the staging bucket where the SDK uploads your script. If you leave out the staging bucket, the SDK will refuse to create the job, because it has nowhere to put your code.

CustomTrainingJob describes a job whose code is a local Python script. display_name is the label you will see in the console. script_path is the file to run. container_uri names the container image the script runs inside: Google publishes prebuilt training images for scikit-learn, TensorFlow, PyTorch and XGBoost, with the framework already installed, so you do not build your own. requirements lists extra pip packages to install at start-up. model_serving_container_image_uri names a prebuilt prediction image, the program that will later load your model for serving. Training and serving images are different images with different jobs.

Check the image tags before you run The container image paths above follow the pattern in the official prebuilt-container pages, but framework versions in those images change and old ones are retired. Open the "prebuilt containers for custom training" and "prebuilt containers for prediction" pages in the documentation, copy a current scikit-learn image URI for each, and use that. A container_uri that no longer exists fails at job start with an image-pull error, after you have waited for the machine.

job.run(...) is where the cloud work happens. replica_count=1 asks for a single machine, and machine_type="n1-standard-4" picks four virtual CPUs and a standard amount of memory. By default the call blocks: your terminal waits while Vertex AI provisions the machine, pulls the image, runs the script, and finishes. That takes several minutes even for a tiny job, because starting a machine is not instant. Expect machine start-up time to dominate a job this small.

When it finishes, run returns a Model object, because we gave the job a serving container. The platform took the files your script saved, registered them as a model, and attached the serving image. model.resource_name prints something like projects/123456789/locations/us-central1/models/987654321. That long string is the model's identity, and it shows the region is part of the name.

Step 3. Watch it in the console

While job.run blocks, open the console, go to the training section, and find the custom job. You will see its state move through pending, running, and succeeded, and a logs link that opens Cloud Logging filtered to this job. Open the logs. You should see your accuracy: line among the platform's own messages. This is how you debug a failing job, so make friends with it now.

Try it
  1. Run train.py locally, then submit it with submit_train.py, after replacing the image tags with current ones from the documentation.
  2. In the console, open the job's logs and find the accuracy line. Note how long the job took from submission to completion, and how much of that was start-up.
  3. Run gcloud ai custom-jobs list --region=us-central1 and find your job in the output.

The Model Registry: versions and aliases

When the job finished, a model appeared. Open the Model Registry page in the console and find it. This is the shelf we mentioned earlier, and it earns its place the moment you train a second time.

A registered model has a name, a region, a pointer to the artifact in Cloud Storage, the serving container to use, and labels you can add, which are key-value tags used for search and for cost reports. The registry also records lineage: which training job produced the model and from what. When someone asks in six months which job produced the model now in production, the registry is where you look.

Versions solve the "retrained, but did it get better?" problem. Rather than creating an unrelated model each time, you upload the new artifact as a new version of the same model. Versions are addressed as MODEL_ID@VERSION. Aliases are movable labels on versions: default points to the version served when nobody says which one, and teams add their own such as champion or candidate. Promoting a model then means moving an alias rather than copying files. The Mid-level guide shows promotion in CI.

From Python you can look at what you have:

PYTHON
from google.cloud import aiplatform

aiplatform.init(project="PROJECT_ID", location="us-central1")

for m in aiplatform.Model.list():
    print(m.display_name, m.resource_name)

And the same thing from the terminal:

BASH
gcloud ai models list --region=us-central1

You can also register a model you trained somewhere else, such as on your laptop, without any Vertex AI training job. You upload the artifact to a bucket and call Model.upload, which we use in the next section's alternative. This matters because many real teams train in a notebook and use Vertex AI only for serving.

Name things on purpose The display_name you pick shows up in the console, in lists, and on invoices. train-sklearn-toy is findable; test2 is not. Add a label such as owner and purpose on anything you create, and you will be able to filter the console and the bill by it later.
Try it
  1. Find your model in the Model Registry page. Note its version number, its artifact location and its serving container.
  2. Run the training job a second time with a different random_state and see whether a new model or a new version appears. Read the documentation on uploading a model version and explain the difference to a friend.

Deploying to an endpoint and asking for predictions

A registered model does nothing until you deploy it. Deploying creates, or reuses, an endpoint and puts the model on machines behind it.

deploy.py
from google.cloud import aiplatform

aiplatform.init(project="PROJECT_ID", location="us-central1")

model = aiplatform.Model.list(filter='display_name="train-sklearn-toy"')[0]

endpoint = model.deploy(
    machine_type="n1-standard-2",
    min_replica_count=1,
    max_replica_count=1,
)
print(endpoint.resource_name)

model.deploy creates an endpoint for you when you do not pass an existing one. It then provisions a machine of the type you asked for and loads your model onto it using the serving container. Expect this to take several minutes, and it blocks while it works. We set both replica counts to one for a lab: min is the number of replicas that are always running, and max is how far autoscaling may grow under load. For anything real you would set min to at least two for availability, but in a lab one is enough.

The meter is running An endpoint with a model deployed is billed for the machines the whole time it is up, including every night and weekend when nobody calls it. This is the most common cause of surprise bills on this platform. Before you close your laptop, run the teardown at the end of this section.

Asking for a prediction

Once deployed, call the endpoint with a list of instances. For our model, an instance is a row of three numbers, because the training data had three features:

predict.py
from google.cloud import aiplatform

aiplatform.init(project="PROJECT_ID", location="us-central1")

endpoint = aiplatform.Endpoint.list(filter='display_name="train-sklearn-toy-endpoint"')[0]
response = endpoint.predict(instances=[[0.5, -1.2, 0.3], [1.1, 0.4, -0.7]])
print(response.predictions)

The result is a list with one entry per instance, here the predicted class, such as [1, 0]. The format of instances and predictions depends on the serving container: the scikit-learn image expects a list of feature rows, and the TensorFlow and PyTorch images have their own conventions. If you send the wrong shape, the error comes back as a 400 with a message from the model server, and reading that message is how you learn the expected format.

Note the filter on display_name. When we let deploy create the endpoint, its name was generated for us; in your own code you will usually create the endpoint explicitly so that you control the name, as in the next block. Use whichever you like, but know which one you used when you look for it later.

PYTHON
endpoint = aiplatform.Endpoint.create(display_name="train-sklearn-toy-endpoint")
model.deploy(endpoint=endpoint, machine_type="n1-standard-2",
             min_replica_count=1, max_replica_count=1)

The terminal has equivalents for the same operations, which are valuable in scripts and in debugging:

BASH
gcloud ai endpoints list --region=us-central1
gcloud ai endpoints create --region=us-central1 --display-name=my-ep
gcloud ai endpoints deploy-model ENDPOINT_ID --region=us-central1 \
  --model=MODEL_ID --display-name=DEPLOYMENT_NAME \
  --machine-type=n1-standard-4 --min-replica-count=1 --max-replica-count=10

Endpoints are also reachable over REST, which is what other services will call in production. The URL has the form https://REGION-aiplatform.googleapis.com/v1/projects/PROJECT/locations/REGION/endpoints/ENDPOINT_ID:predict, with a bearer token in the Authorization header. The SDK builds that request for you; knowing the shape helps when you debug from curl.

Tearing it down

This is not optional. Deleting an endpoint that still has a model deployed fails, so undeploy first:

teardown.py
endpoint.undeploy_all()
endpoint.delete()

Verify with gcloud ai endpoints list --region=us-central1, which should show nothing. The registered model costs almost nothing to keep, since it is just a pointer and a file in a bucket; the running replicas are what cost money.

Try it
  1. Deploy your model with one replica, send it two instances, and read the predictions.
  2. Send it an instance with the wrong number of features. Read the error message carefully and note which layer produced it.
  3. Tear down the endpoint, then confirm with gcloud ai endpoints list that it is gone.

Batch prediction: when you do not need an endpoint

Many problems do not need a live service. If you score a million customers every night, or label a folder of images once, an always-on endpoint is wasteful. Batch prediction runs a one-off job: it reads inputs from Cloud Storage or BigQuery, runs your model on temporary machines, writes the outputs, and disappears. You pay only while it runs.

The input for our scikit-learn model is a file of JSON lines, one instance per line, in a bucket. Make one:

BASH
cat > batch_in.jsonl <<'EOF'
[0.5, -1.2, 0.3]
[1.1, 0.4, -0.7]
[0.0, 0.0, 0.0]
EOF
gcloud storage cp batch_in.jsonl gs://PROJECT_ID-vertex-lab/batch/in.jsonl

Then submit the job:

batch.py
from google.cloud import aiplatform

aiplatform.init(project="PROJECT_ID", location="us-central1")
model = aiplatform.Model.list(filter='display_name="train-sklearn-toy"')[0]

batch_job = model.batch_predict(
    job_display_name="batch-toy",
    gcs_source="gs://PROJECT_ID-vertex-lab/batch/in.jsonl",
    gcs_destination_prefix="gs://PROJECT_ID-vertex-lab/batch/out/",
    machine_type="n1-standard-4",
)
print(batch_job.state)

The job writes result files under the destination prefix, in a subfolder it creates for the run. List them with gcloud storage ls --recursive gs://PROJECT_ID-vertex-lab/batch/out/ and read one with gcloud storage cat. The format of the input file must match what the serving container expects, which is the same rule as for online prediction.

The decision between the two styles is a trade-off you should be able to explain.

Batch prediction fits when

Results are needed on a schedule, volume is large, nobody is waiting for the answer, and you would rather not pay for idle machines. Examples: nightly scoring, monthly reports, one-off relabelling.

An endpoint fits when

A person or another service needs an answer in milliseconds to seconds, requests arrive one at a time, and the cost of an always-running replica is justified by the traffic. Examples: a recommendation widget, a fraud check at payment time.

Try it
  1. Create the three-line input file, upload it, and run the batch job against your model. The model must still be registered even if the endpoint is gone.
  2. Read the output files and check that you got one prediction per input line.

Calling Gemini: generative AI in the same project

Not every task needs a model you trained. Vertex AI also gives you hosted generative models, notably Gemini, which you call over an API without any training or deployment step. You pay per use, measured in tokens (small chunks of text), rather than for machine hours.

Use the Google Gen AI SDK, which you installed as google-genai. This is the current way to call Gemini on Vertex AI; the older vertexai.generative_models classes belong to the pre-2.0 SDK and are on their way out.

gemini.py
from google import genai

MODEL_ID = "REPLACE_WITH_A_CURRENT_MODEL_ID"

client = genai.Client(vertexai=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
    model=MODEL_ID,
    contents="Explain a Vertex AI endpoint to a new engineer in two sentences.",
)
print(response.text)

The vertexai=True argument is what routes the call to Vertex AI, with your Google Cloud project, billing and permissions, instead of to the separate Gemini Developer API with an API key. The same client can be configured by environment variables instead: GOOGLE_GENAI_USE_VERTEXAI=true, together with variables for the project and location. Check the Gen AI SDK documentation for the exact names before relying on that form.

Notice that the model name is a placeholder. Model identifiers change quickly: Google retires older Gemini versions on a published schedule, and a hardcoded ID in a tutorial can be dead within months. Gemini 1.5 access, for example, was discontinued in September 2025. Open the "model versions" page in the documentation, pick a current stable model, and put its ID in one constant at the top of your file so that updating it is a one-line change. If you see 404 Publisher Model ... was not found or your project does not have access to it, the ID has been retired or is unavailable in your location.

The location for generative calls is often global or one of several supported regions, and not every model is offered in every region. If data residency matters for your employer, check the models and locations pages for the combination you need before building anything on it. For companies in the Gulf or Egypt, this is a real design constraint, not a footnote: some regulated data may not be allowed to leave a particular country, and the answer lives in the location support tables and in your compliance team's rules.

A chat-style interaction, where the model remembers earlier turns, uses a chat object from the same SDK; the Gen AI SDK documentation shows the current pattern. If you want to build agents that call tools, that is a larger topic with its own package, and the Google ADK guide covers the toolkit Google built for it. The Gemini API guide covers the same models through the developer-facing route.

Retry transient errors Generative calls can return 429 when you exceed a per-model, per-region quota. Wrap calls in a retry with exponential backoff (wait 1 second, then 2, then 4) rather than looping instantly, which only makes the throttling worse. Quota increases are requested through the console and are not instant.
Try it
  1. Find a current Gemini model ID on the model versions page, put it in MODEL_ID, and run gemini.py.
  2. Change the location to a regional one that the documentation says supports your model, and see whether the call still works.
  3. Cause a 404 on purpose by using a nonsense model name, and read the message.

Notebooks and the console: where beginners actually start

Plenty of people meet Vertex AI through a notebook rather than a terminal, and it is a fine way to learn. Vertex AI Workbench gives you a managed JupyterLab instance running on a Google Cloud virtual machine, with gcloud and the SDK already installed and authenticated as the instance's service account. Colab Enterprise gives you notebooks inside the console without managing a machine. Older "managed notebooks" and "user-managed notebooks" products reached end of support in 2025 and 2026, so when a tutorial tells you to create one of those, use a Workbench instance instead.

A Workbench instance is a machine that bills while it is running. Stop it when you finish for the day, and delete it when the course is over. The console shows a stop button on the instance list.

The console is also where you inspect everything this guide created. The sidebar groups the platform by stage. You will find training jobs, the Model Registry, endpoints, batch predictions, pipelines, and Model Garden, a catalogue of Google, open and partner models you can deploy or try. Learn to read the Logs tab on jobs and the Monitoring graphs on endpoints; they answer most "why did this fail" and "is it being used" questions without any code.

Pipelines, experiments tracking, feature stores and model monitoring are the next layer of the platform. They matter in production, and they build on exactly what you have done here: a pipeline runs a training job, registers the model and deploys it automatically. The Mid-level guide covers them, and the Kubeflow guide explains the pipeline format Vertex AI Pipelines runs. If you are weighing it against open-source tracking tools, the MLflow guide is the natural comparison.

Try it
  1. Open the Model Garden page and read the description of one model. Note which region its deployment options support.
  2. If you create a Workbench instance, write down the exact time you started it and set a calendar reminder to stop it.

Permissions, regions and quotas

Most failures on a cloud platform are not bugs in your code. They are one of three things: you are not allowed, you are in the wrong place, or you ran out. Learn to recognise them.

Not allowed: IAM. Google Cloud's permission system is called IAM. A principal (a user or service account) is granted a role (a bundle of permissions) on a resource. When your call fails with 403 and a message such as Permission 'aiplatform.endpoints.predict' denied, the message names the missing permission. The fix is usually to grant a role that contains it, such as roles/aiplatform.user, to the identity that made the call. Remember that identity might not be you: a training job runs as a service account, and that account needs access to your bucket and, for private images, to Artifact Registry. Granting yourself roles does nothing for the job's account. Prefer narrow roles; the project Owner role works and is a poor habit.

Wrong place: regions. Create your bucket, your jobs, your models and your endpoints in the same region unless you have a reason not to. A model in us-central1 cannot be deployed to an endpoint in europe-west4. A feature may exist in one region and not another. When NotFound appears for something you are sure you created, check the region in your aiplatform.init call and in the --region flag first. For employers in the Middle East, check the locations page for what each Google Cloud region in the Gulf region offers before you design around it; the supported set of features differs by region and changes.

Ran out: quotas. Google limits how much of each resource a project can use, per region. GPUs are the usual surprise: a new project often has little or no GPU quota, and a job that requests one fails with a 429 RESOURCE_EXHAUSTED message naming the quota metric. The fix is to request a quota increase in the console (which can take time), pick another region, or use a smaller machine. Plan for this before a deadline, not on the day.

Reading a Google error Google API errors have a numeric code (403, 404, 429), a status word (PERMISSION_DENIED, NOT_FOUND, RESOURCE_EXHAUSTED) and a human sentence that usually names the exact thing that went wrong. Read the whole sentence before you search the internet. It frequently contains the permission, the API, or the resource you need.
Try it
  1. In the console, open IAM and find which roles your own account has on the project.
  2. Open the quotas page and search for a custom model training GPU metric in your region. Note its current limit.

Cost control: the habits that protect your wallet

Cloud bills are made of a few line items. Knowing which they are makes the bill predictable.

Endpoint replicas bill per node-hour while deployed, idle or not. Training jobs bill per machine-hour while running, and a job that hangs keeps billing until it times out or you cancel it. Notebook instances bill while running. Cloud Storage bills per gigabyte stored per month, which is small for toy data. Generative calls bill per token. Batch predictions bill for the machines while the job runs.

The habits that follow from this are simple.

  1. Work in a disposable project and delete the project when done. Deleting a project begins a grace period before everything is purged, but billing for its resources stops.
  2. Set a budget alert the day you create the project, with several thresholds.
  3. Undeploy the moment you finish. Make endpoint.undeploy_all() a reflex.
  4. Prefer batch prediction when nobody is waiting for the result.
  5. Start small. A job on n1-standard-4 for a toy model is plenty; a GPU is for models that need one.
  6. Use labels on everything you create so the billing report can be filtered by owner or purpose.
  7. Check the console daily during the course for endpoints and instances you forgot.

Prices change and differ by region, so use the official pricing page and the pricing calculator for estimates, and do not trust a number you remember from a blog.

Try it
  1. Open Billing, then Reports, and group the last day's costs by service. Identify which service your lab produced.
  2. Add a label purpose=vertex-lab to your model from the Model Registry page.

Reading the common errors

These are the failures a beginner meets, with what they mean and what to do. The exact wording can vary between SDK versions and services, so match on the meaning and the status word rather than on every character.

DefaultCredentialsError: Your default credentials were not found. Your code asked for Application Default Credentials and none exist. Run gcloud auth application-default login. Remember that gcloud auth login is a different credential and does not satisfy this.

403 ... Vertex AI API has not been used in project ... before or it is disabled. The API is not enabled. Run gcloud services enable aiplatform.googleapis.com, wait a minute, and retry.

403 Permission ... denied. The identity lacks a permission. Read the permission named in the message, find which identity ran the call, and grant a role containing it. Do not guess which identity: for a training job, it is the job's service account.

404 ... not found. Either the resource does not exist, it exists in another region, or, for a Gemini model, the model was retired. Check the region and the exact ID before assuming deletion.

429 RESOURCE_EXHAUSTED or a quota message. You hit a limit. For a GPU, request quota or change region. For generative calls, add retries with backoff, or ask for more quota.

400 ... staging bucket errors. aiplatform.init needs staging_bucket="gs://..." for custom training, and the bucket must exist and be reachable.

FailedPrecondition: Model server never became ready. Deployment started the serving container and it never answered its health check. Open Cloud Logging for the endpoint, read the container's output, and look for a Python exception while loading the model. A mismatch between the library version you trained with and the one in the serving image is a typical cause: a scikit-learn model pickled with one version may not load in an image with another. Train and serve with matching versions.

Image pull failures at job start. The container_uri is wrong or retired, or the job's service account cannot read a private repository. Copy a current URI from the docs, and for private images grant the right Artifact Registry reader role.

ImportError or ModuleNotFoundError for vertexai.generative_models. Your code is from before the 2.0 SDK. Move to google-genai, as in the Gemini section.

A job that "succeeds" but produces nothing. Your script wrote the model somewhere other than where the platform looks. Print the output path in your script, read it in the logs, and check it against the model directory the job reports.

Do not debug by resubmitting Each resubmission costs minutes of start-up and some money. Reproduce locally when you can, and read the job's logs in full before you change anything. Most cloud failures are named in the first red line of the log.
Try it
  1. Run gcloud auth application-default revoke, then run the Python check from the setup section and read the error. Log in again to fix it.
  2. Temporarily change location in aiplatform.init to another region and run the model listing. Explain why the result changed.

Putting it all together

Here is one small end-to-end project that uses every piece. You will build it from scratch, tear it down, and write down what each step cost you in time and money. Treat it as a rehearsal for a real interview task.

Goal. Train a classifier on managed hardware, register it, serve it for five minutes, score a batch file, call Gemini to explain the result in plain language, and delete everything.

  1. Project. Create a new project, link billing, set a budget alert, and enable the Vertex AI API.
  2. Tools. Install gcloud, run the login and ADC commands, create a virtual environment, and install google-cloud-aiplatform and google-genai. Pin the versions.
  3. Bucket. Create a bucket in your chosen region.
  4. Train. Run train.py locally, then submit it with CustomTrainingJob using current container image URIs.
  5. Register. Confirm the model in the Model Registry and add a label.
  6. Serve. Deploy to an endpoint with one replica, send predictions from Python, then from curl using the REST URL and a token from gcloud auth print-access-token.
  7. Batch. Upload a JSON-lines file and run batch_predict.
  8. Explain. Call Gemini with a prompt that includes the predictions and asks for a one-paragraph summary for a non-technical manager.
  9. Clean up. Undeploy and delete the endpoint, delete the bucket's test objects, and stop any notebook instance. Verify with gcloud ai endpoints list and the console.
  10. Write down. In a short README, record the region, the SDK version, the model ID you used, the time each step took, and anything that surprised you.
BASH
# the REST call from step 6, after you substitute your own values
curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://us-central1-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/endpoints/ENDPOINT_ID:predict \
  -d '{"instances": [[0.5, -1.2, 0.3]]}'

The call returns JSON with a predictions array. If you get 401, your token expired or you are logged into the wrong account; if 404, check the endpoint ID and region in the URL.

Try it
  1. Complete the ten steps above in one sitting and keep the README.
  2. Ask a colleague to follow your README on their own project and note where they get stuck. Fix the README.

What you can now do, and what comes next

If you worked through the Try it tasks, you can do things many people who list Vertex AI on a resume cannot. You can set up a project safely, with a budget alert and no leaked keys. You can authenticate the command-line tool and the SDK and know which login is which. You can train on managed hardware from a script, read the logs of a job, and register the result. You can deploy to an endpoint, call it from Python and from REST, and tear it down so it stops billing. You can run a batch job when an endpoint is overkill. You can call a Gemini model through the current SDK and know that model IDs rotate. And you can read the platform's errors by their code and their sentence rather than by panic.

You also know what you have not yet touched. Pipelines automate the whole sequence you just did by hand and record each run. Experiments track parameters and metrics across runs. Custom containers let you serve anything, not just frameworks Google prebuilt for. Hyperparameter tuning jobs search for better settings. Model monitoring watches for data drift after deployment. Private networking and encryption keys matter in regulated industries. Feature stores share engineered features between teams.

What comes next. The Mid-level track picks up exactly here, with pipelines, experiment tracking, custom serving containers, autoscaling behaviour, service accounts per job, and CI/CD that promotes a model by moving an alias. The Senior track covers private endpoints, perimeter controls, multi-project design, upgrade strategy, and when to choose a different platform entirely. Alongside it, read the neighbouring guides in this catalogue: Kubeflow for the pipeline model, MLflow for experiment tracking, BentoML and KServe for serving outside a managed platform, and SageMaker to see how Amazon solved the same problems. Seeing two clouds side by side is the fastest way to separate platform ideas from vendor vocabulary.

Keep the habit this guide started: when you build something on a managed platform, write down what it costs to leave running, and then switch it off.

Sources