Skip to content

API quickstart

This guide takes an API-first developer from a gateway token to a working inference request. The gateway exposes an OpenAI-compatible surface, so most existing OpenAI clients work by changing two settings: the base URL and the API key.

💡 Note: This is the developer on-ramp. For the complete request, streaming, header, and error contract, see Inference API (/v1). To set up the tenant, gateway, and provider key in the first place, see Setting up a production tenant.

Before you begin, ensure the following conditions are met:

  • ☑ You have a tenant slug and a gateway slug (for example myapp / production).
  • ☑ The gateway has a provider key stored (Provider keys / BYOK).
  • ☑ You have a gateway token, or admin access to create one (see below).

Getting a token

Every inference request carries a gateway token. Create one against the admin API:

curl -X POST "https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/tokens" \
  -H "Content-Type: application/json" \
  -d '{"label": "quickstart", "scopes": ["inference"]}'

The plaintext token is returned once in the response token field — copy it immediately; it is never shown again. For expiry, per-token rate limits, and budgets, see Authentication.

💡 Note: While exploring, an administrator can disable authentication on a gateway so you can send unauthenticated requests. Never do this in production — see Authentication — Disabling authentication.


Your first request

The fastest path is the unified compat endpoint: it infers the provider from the model field and always returns an OpenAI-shaped response.

curl -s -X POST \
  "https://<your-gateway-host>/v1/myapp/production/compat/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <GATEWAY_TOKEN>" \
  -d '{
    "model": "claude-opus-4-6",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

A successful response:

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "claude-opus-4-6",
  "choices": [{"message": {"role": "assistant", "content": "Hello! How can I help?"}}],
  "usage": {"prompt_tokens": 10, "completion_tokens": 9, "total_tokens": 19}
}

-> The gateway resolves the provider from the model name, forwards the request, and returns the response. Change the model to switch provider — gpt-4o routes to OpenAI, claude-opus-4-6 to Anthropic, grok-3 to xAI — without changing anything else. See Providers overview for the full resolution rules.

💡 Note: To pin one exact provider with no model-name guessing, replace compat in the path with the provider name, for example /v1/myapp/production/openai/chat/completions.


Authenticating

The gateway accepts the token in any of three headers, checked in this order — the first one present wins:

Order Header Form Compatible with
1 x-aig-token myra_xxxx Gateway-native
2 Authorization Bearer myra_xxxx OpenAI SDKs
3 x-api-key myra_xxxx Anthropic SDKs

This is why the gateway drops in as a replacement for the OpenAI or Anthropic API without changing how your client sends credentials.


Using the OpenAI SDK

Point the OpenAI SDK at the gateway by setting the base URL to the compat endpoint and the API key to your gateway token:

from openai import OpenAI

client = OpenAI(
    base_url="https://<your-gateway-host>/v1/myapp/production/compat",
    api_key="<GATEWAY_TOKEN>",
)

resp = client.chat.completions.create(
    model="claude-opus-4-6",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

The SDK appends /chat/completions to the base URL and sends the key as Authorization: Bearer, both of which the gateway accepts unchanged.

💡 Note: To route Claude Code or another Anthropic-native client through the gateway, see Claude Code.


Streaming

Set "stream": true to receive a Server-Sent Events stream of OpenAI-shaped chunks, terminated by a data: [DONE] sentinel:

curl -sN -X POST \
  "https://<your-gateway-host>/v1/myapp/production/compat/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <GATEWAY_TOKEN>" \
  -d '{
    "model": "claude-opus-4-6",
    "messages": [{"role": "user", "content": "Count to three."}],
    "stream": true
  }'
data: {"id":"...","choices":[{"delta":{"content":"One"},"finish_reason":null}]}

data: [DONE]

See Inference API — Streaming for the finish_reason normalisation and the streaming guardrail-block behaviour.


Handling errors

Errors return a structured JSON envelope and set the X-AIG-Error response header to the same code:

{
  "error": {
    "code": "invalid_request",
    "message": "Human-readable detail about what went wrong"
  }
}

Branch your client logic on the stable code, never on the human-readable message. A few you will meet early:

Code HTTP Cause
unauthorized 401 Missing, invalid, or wrong-gateway token.
token_expired 401 The token passed its expiry — mint or refresh a token.
token_revoked 401 The token was revoked — mint a new token.
rate_limited 429 The sliding-window rate limit was exceeded.
quota_exceeded 429 A budget cap was reached.
guardrail_blocked 400 A guardrail blocked the request (or, in streaming mode, a 200 with a synthetic SSE error chunk).

The full list is in Error codes.


Next steps


See also