What is Vertex AI? Platform, modules, and pricing (now the Gemini Enterprise Agent Platform)

Jerzy Kopaczewski 29 September 2026 9 min read
Contents

What is Vertex AI? Platform, modules, and pricing (now the Gemini Enterprise Agent Platform)

Vertex AI is Google Cloud's managed platform for the full machine learning lifecycle - data prep, training, deployment, serving, and increasingly agents and generative AI. One thing to know up front, because it changes how you read every doc and tutorial: at Google Cloud Next in April 2026, Google retired "Vertex AI" as a standalone brand and renamed it the Gemini Enterprise Agent Platform. The platform underneath is the same; this guide walks through what it does, its modules, and how pricing works - under the name you are still searching for.
We are an AWS Partner, and for qualifying workloads a proof of concept can be funded through AWS Partner programmes. If you are weighing a managed ML platform across clouds, talk to us about whether your PoC qualifies before you commit spend.

If you are specifically comparing Vertex AI against AWS’s platform, that is a separate piece: Vertex AI vs SageMaker. Here we focus on what Vertex AI is on its own.

 

The rebrand, and why this guide still says “Vertex AI”

At Google Cloud Next in April 2026, Google retired the “Vertex AI” brand and folded the platform into the Gemini Enterprise Agent Platform. This is a positioning change, not a teardown. Everything that was Vertex AI - Model Garden, custom training, AutoML, the model registry, endpoints, feature store, and pipelines - still exists and still works. It is now presented under the Agent Platform umbrella, with the roadmap leaning towards agent workloads.

We keep saying “Vertex AI” throughout because that is the term in near-universal use: it is what people search, what the existing documentation and tutorials call it, and what shows up in a year of blog posts and Stack Overflow answers. When you open the Google Cloud console today you will see the new name, so the practical mapping is: Vertex AI (what you searched) = Gemini Enterprise Agent Platform (what the console shows). For the official product page, see Gemini Enterprise Agent Platform (formerly Vertex AI).

 

What Vertex AI actually is

Vertex AI launched in 2021 as Google’s unified ML platform, bringing what had been scattered services (AI Platform, AutoML, and others) under one roof. The pitch is a single managed environment where a team can go from raw data to a deployed model and a served prediction without stitching together separate tools or running its own training and serving infrastructure.

In practice it sits on top of Google Cloud’s data and compute stack. Its natural home is a team already on Google Cloud, especially one with data in BigQuery, because the integrations there are the tightest. If that describes you, the equivalent AWS pattern is a useful contrast for understanding what “unified data and AI” looks like on the other major cloud.

 

The modules that matter

Vertex AI is best understood as a set of modules you can adopt independently rather than one monolith. These are the ones that come up on real projects.

  • Model Garden. A catalogue of models you can deploy or call: Google’s own Gemini family, open-weight models, and third-party models including Claude. This is the entry point for most generative AI work on the platform, and - as we will see - it is where the pricing gets layered.
  • Custom training. Managed training jobs on CPU, GPU, or TPU, so you bring your training code and Google runs the infrastructure. You pay for the compute while the job runs.
  • AutoML. Train models on tabular, image, or text data without writing the model code, trading control for speed.
  • Model registry and endpoints. Register model versions, then deploy them to endpoints for online or batch prediction. Endpoints are the piece most likely to surprise you on the bill - more on that below.
  • Pipelines. Orchestrate the steps of an ML workflow (data prep, training, evaluation, deployment) as a repeatable, versioned process rather than a pile of scripts.
  • Feature store and metadata. Shared infrastructure for serving features consistently between training and inference, and for tracking lineage.

The agent and generative-AI tooling that the 2026 rebrand foregrounds builds on top of these same primitives - Model Garden for the models, endpoints for serving, pipelines for orchestration.

 

Planning a Vertex AI build and want a second opinion on the architecture?

Book a free 30-min call

 

How Vertex AI pricing works

There is no single “Vertex AI price”. The bill is the sum of several independently metered pieces, and the two things that catch teams out are endpoints and Model Garden.

  • Endpoints bill while deployed, not while used. You pay for a model deployed to an endpoint for as long as it is deployed, even when it serves zero predictions. To stop the charge you have to undeploy the model. A forgotten endpoint quietly running overnight and over a weekend is the single most common bill surprise on the platform.
  • Model Garden is several price models at once. Third-party models such as Claude bill at the provider’s rate; open-weight models bill per token at a Google-set rate; and models you deploy yourself bill by the GPU-hour of the machine they run on. Choosing a model therefore picks a billing model, not just a number.
  • Training bills for the compute while the job runs. CPU, GPU, or TPU time for the duration of the training job, which is bounded and easy to reason about - the opposite of the always-on endpoint.
  • Committed-use discounts move the bill the most. For predictable, steady workloads, Google Cloud committed-use discounts can meaningfully cut the effective rate. On a stable platform the discount strategy usually matters more than the per-unit list price.
  • Egress and supporting services (storage, networking, the feature store) are the quiet additions that make the effective cost diverge from the headline compute rate.

New Google Cloud accounts come with free credits (US$300 at the time of writing) that are enough to explore the platform, but not enough to hide a long-running endpoint for long.

For the discount mechanics in detail, see GCP cost optimisation - committed-use discounts, spot VMs, FinOps.

A note on pricing: Google changes Vertex AI prices regularly, and the figures depend on region, model, machine type, and configuration. Treat everything above as how the bill is shaped, not as a quote. Confirm current rates on the Vertex AI pricing page and model your own token and endpoint volume before committing to anything.

 

When Vertex AI is the right call

  • Your data already lives in Google Cloud. The BigQuery integration is the strongest single reason to standardise on Vertex AI; running it against data in another cloud adds an integration tax that usually outweighs any feature advantage.
  • You want Gemini and a multi-provider Model Garden. If access to Google’s models alongside open-weight and third-party options matters, Model Garden is the draw.
  • You are building around agents. The post-rebrand direction is explicitly agent-first, so new agent tooling will land here first.
  • Your team already runs on GKE and Google’s open-source ML stack. The platform meets you where you already are.

If instead your infrastructure and data are on AWS, the honest answer is usually to stay there - the Vertex AI vs SageMaker comparison works through that trade-off in full.

 

When it is worth bringing in an external partner

  • Platform selection across clouds. If your data is in one cloud and your team knows another, the integration cost is real; cost it out before committing rather than after.
  • Cost modelling before rollout. Endpoint, token, and egress costs are easy to underestimate. A realistic model up front avoids the idle-endpoint surprise and makes committed-use decisions defensible.
  • The rebrand transition. Existing scripts, tooling, and internal docs that reference “Vertex AI” are worth mapping onto the Gemini Enterprise Agent Platform deliberately, rather than discovering the renames one broken link at a time.
  • Agent and generative-AI builds. The agent-first direction opens new patterns; getting the architecture right early is cheaper than refactoring a production agent later.

 

Summary

Vertex AI is Google Cloud’s full-lifecycle managed ML platform - Model Garden, training, endpoints, pipelines, and the supporting infrastructure around them - and since April 2026 it is delivered as the Gemini Enterprise Agent Platform.

Three things to remember:

  1. The name changed, the platform did not. “Vertex AI” is now the Gemini Enterprise Agent Platform, repositioned around agents. Search still uses the old name, so that is what you will find in docs and tutorials - and in the console you will see the new one.
  2. The modules are adoptable independently. Model Garden, custom training, AutoML, endpoints, and pipelines are pieces you can take as needed, not an all-or-nothing commitment.
  3. The bill is shaped, not fixed. Idle endpoints bill anyway, Model Garden layers several price models, and committed-use discounts move the effective cost more than headline rates - so model your own usage before you commit.
Jerzy Kopaczewski

Building on Vertex AI?

Book a free 30-minute call. We help teams decide whether Vertex AI (now the Gemini Enterprise Agent Platform) fits their cloud, model the real endpoint and token costs before rollout, and plan the architecture for ML and agent workloads.

Book a call
GCP Vertex AI machine learning MLOps GenAI Model Garden

Read also:

Previous post