Google Vertex AI
Description
Google Vertex AI is the enterprise AI platform provided by Google Cloud. The gateway integrates with Vertex AI for Gemini-family models. The wire format, request, response, and SSE streaming behaviour are identical to the Google Gemini AI Studio integration; only the base URL and the authentication differ.
Vertex AI is a separate integration from Google Gemini AI Studio. Vertex AI runs inside a Google Cloud project and authenticates with a service account bound to that project; Gemini AI Studio uses an API key bound to a Google account. Choose Vertex AI for enterprise deployments that require Google Cloud billing, IAM, and audit controls.
| Feature | Limitation |
|---|---|
| Chat completions for Gemini-family models | — |
| Streaming responses (SSE) | — |
| Vision input | Not forwarded. Image and document blocks are replaced with a text placeholder; the model receives no visual or file input. Native image support is a follow-up. |
| Tool use | Not supported — the gateway does not forward client tool definitions to Gemini/Vertex, so the tool-calling loop never fires. |
| Service-account OAuth2 authentication | Supported. The gateway signs a JWT with the service account's private key and exchanges it for a short-lived OAuth2 access token (RFC 7523 JWT-bearer grant), sent as Authorization: Bearer. |
| Embeddings | Not exposed by the gateway. |
Web search
Vertex AI supports native web search through Google Search grounding, identical to the Google Gemini AI Studio integration. When a request carries a web_search tool (set the x-aig-web-search: 1 header, or configure the gateway with mode: "always"), the gateway enables grounding by sending tools: [{ "googleSearch": {} }] in a single upstream leg; no separate search round-trip is performed. The grounded answer is returned in the normal response.
Grounding runs from the gateway's resolved vertex_region, so on a gateway with eu_region_routing: true the grounded search stays within the configured EU region — the same data-residency posture as every other Vertex AI request (see the note under Required gateway configuration).
Required gateway configuration
The gateway resolves the Google Cloud project and region from its configuration before Vertex AI requests can be routed:
| Key | Required | Default | Description |
|---|---|---|---|
vertex_project |
yes* | — | The Google Cloud project ID. If omitted, the gateway falls back to a deployment-wide default project; if neither is set the request fails closed* with a configuration_error (the gateway never dials an empty-project URL). |
vertex_region |
no | us-central1 |
The Google Cloud region. Falls back to a deployment-wide default region when unset. |
These keys live in the gateway configuration JSON, not in the Add Model dialog. The deployment-wide default project / region fallbacks let your operator enable Vertex fleet-wide without setting the project on every gateway.
EU data residency. On a gateway with
eu_region_routing: true,vertex_regionmust be an EU-member region (europe-west1/3/4/8/9/10/12,europe-north1,europe-north2,europe-central2,europe-southwest1) — otherwise the request is rejected403data_residency_blockedand never reaches a US endpoint.europe-west2(London) andeurope-west6(Zurich) are not EU-member and are rejected. See Data residency.
BYOK key format
The BYOK value for the vertex provider is a Google Cloud service-account JSON key (the file downloaded from the Google Cloud console when you create a key for a service account). The gateway reads three fields from it — client_email, private_key (a PKCS#8 PEM), and token_uri — and rejects the credential (fails closed, no request sent) if any is missing, if the private key does not load, or if token_uri is not an https:// URL.
On each request the gateway mints an OAuth2 access token from the service account:
- It builds a JWT assertion (
iss=client_email,scope=https://www.googleapis.com/auth/cloud-platform,aud=token_uri,iat/exp) and signs it RS256 with the service account's private key. - It exchanges the assertion at the service account's
token_urifor a short-lived access token (RFC 7523 JWT-bearer grant). - The token is sent as
Authorization: Bearer <token>and cached per service account until shortly before it expires, so the exchange runs about once an hour, not per request.
If the credential is invalid or the service account is disabled, the request fails closed with a configuration_error and the gateway falls through to any configured fallback provider — it never sends an unauthenticated request.
EU data residency — the token exchange. The OAuth2 token exchange in step 2 goes to the service account's
token_uri(normallyoauth2.googleapis.com, which is Google-global, not region-pinned). This is an authentication handshake only: the request carries just the signed JWT assertion (the service account's own identity —iss/scope/aud/iat/exp) and returns a service-account-scoped access token. No user prompt, conversation, or personal data crosses it. The inference call — the only leg that carries user content — always goes to the resolved regional endpoint ({vertex_region}-aiplatform.googleapis.com), so on a gateway witheu_region_routing: trueuser data stays in the configured EU region.
Adding the key
The procedure for storing a BYOK key is the same for every provider. See Provider keys (BYOK) for the steps. Select vertex in the Provider drop-down list and paste the entire service-account JSON (including the -----BEGIN PRIVATE KEY----- block) into the key field.