Skip to content

Google Vertex AI

Description

Google Vertex AI is the enterprise AI platform provided by Google Cloud. The gateway integrates with Vertex AI for Gemini-family models. The wire format, request, response, and SSE streaming behaviour are identical to the Google Gemini AI Studio integration; only the base URL and the authentication differ.

Vertex AI is a separate integration from Google Gemini AI Studio. Vertex AI runs inside a Google Cloud project and authenticates with a service account bound to that project; Gemini AI Studio uses an API key bound to a Google account. Choose Vertex AI for enterprise deployments that require Google Cloud billing, IAM, and audit controls.

Feature Limitation
Chat completions for Gemini-family models —
Streaming responses (SSE) —
Vision input Not forwarded. Image and document blocks are replaced with a text placeholder; the model receives no visual or file input. Native image support is a follow-up.
Tool use Not supported — the gateway does not forward client tool definitions to Gemini/Vertex, so the tool-calling loop never fires.
Service-account OAuth2 authentication Supported. The gateway signs a JWT with the service account's private key and exchanges it for a short-lived OAuth2 access token (RFC 7523 JWT-bearer grant), sent as Authorization: Bearer.
Embeddings Not exposed by the gateway.

Vertex AI supports native web search through Google Search grounding, identical to the Google Gemini AI Studio integration. When a request carries a web_search tool (set the x-aig-web-search: 1 header, or configure the gateway with mode: "always"), the gateway enables grounding by sending tools: [{ "googleSearch": {} }] in a single upstream leg; no separate search round-trip is performed. The grounded answer is returned in the normal response.

Grounding runs from the gateway's resolved vertex_region, so on a gateway with eu_region_routing: true the grounded search stays within the configured EU region — the same data-residency posture as every other Vertex AI request (see the note under Required gateway configuration).


Required gateway configuration

The gateway resolves the Google Cloud project and region from its configuration before Vertex AI requests can be routed:

Key Required Default Description
vertex_project yes* — The Google Cloud project ID. If omitted, the gateway falls back to a deployment-wide default project; if neither is set the request fails closed* with a configuration_error (the gateway never dials an empty-project URL).
vertex_region no us-central1 The Google Cloud region. Falls back to a deployment-wide default region when unset.

These keys live in the gateway configuration JSON, not in the Add Model dialog. The deployment-wide default project / region fallbacks let your operator enable Vertex fleet-wide without setting the project on every gateway.

EU data residency. On a gateway with eu_region_routing: true, vertex_region must be an EU-member region (europe-west1/3/4/8/9/10/12, europe-north1, europe-north2, europe-central2, europe-southwest1) — otherwise the request is rejected 403 data_residency_blocked and never reaches a US endpoint. europe-west2 (London) and europe-west6 (Zurich) are not EU-member and are rejected. See Data residency.


BYOK key format

The BYOK value for the vertex provider is a Google Cloud service-account JSON key (the file downloaded from the Google Cloud console when you create a key for a service account). The gateway reads three fields from it — client_email, private_key (a PKCS#8 PEM), and token_uri — and rejects the credential (fails closed, no request sent) if any is missing, if the private key does not load, or if token_uri is not an https:// URL.

On each request the gateway mints an OAuth2 access token from the service account:

  1. It builds a JWT assertion (iss = client_email, scope = https://www.googleapis.com/auth/cloud-platform, aud = token_uri, iat/exp) and signs it RS256 with the service account's private key.
  2. It exchanges the assertion at the service account's token_uri for a short-lived access token (RFC 7523 JWT-bearer grant).
  3. The token is sent as Authorization: Bearer <token> and cached per service account until shortly before it expires, so the exchange runs about once an hour, not per request.

If the credential is invalid or the service account is disabled, the request fails closed with a configuration_error and the gateway falls through to any configured fallback provider — it never sends an unauthenticated request.

EU data residency — the token exchange. The OAuth2 token exchange in step 2 goes to the service account's token_uri (normally oauth2.googleapis.com, which is Google-global, not region-pinned). This is an authentication handshake only: the request carries just the signed JWT assertion (the service account's own identity — iss/scope/aud/iat/exp) and returns a service-account-scoped access token. No user prompt, conversation, or personal data crosses it. The inference call — the only leg that carries user content — always goes to the resolved regional endpoint ({vertex_region}-aiplatform.googleapis.com), so on a gateway with eu_region_routing: true user data stays in the configured EU region.


Adding the key

The procedure for storing a BYOK key is the same for every provider. See Provider keys (BYOK) for the steps. Select vertex in the Provider drop-down list and paste the entire service-account JSON (including the -----BEGIN PRIVATE KEY----- block) into the key field.