API quickstart
This guide takes an API-first developer from a gateway token to a working inference request. The gateway exposes an OpenAI-compatible surface, so most existing OpenAI clients work by changing two settings: the base URL and the API key.
💡 Note: This is the developer on-ramp. For the complete request, streaming, header, and error contract, see Inference API (/v1). To set up the tenant, gateway, and provider key in the first place, see Setting up a production tenant.
Before you begin, ensure the following conditions are met:
- ☑ You have a tenant slug and a gateway slug (for example
myapp/production). - ☑ The gateway has a provider key stored (Provider keys / BYOK).
- ☑ You have a gateway token, or admin access to create one (see below).
Getting a token
Every inference request carries a gateway token. Create one against the admin API:
curl -X POST "https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/tokens" \
-H "Content-Type: application/json" \
-d '{"label": "quickstart", "scopes": ["inference"]}'
The plaintext token is returned once in the response token field — copy it
immediately; it is never shown again. For expiry, per-token rate limits, and
budgets, see Authentication.
💡 Note: While exploring, an administrator can disable authentication on a gateway so you can send unauthenticated requests. Never do this in production — see Authentication — Disabling authentication.
Your first request
The fastest path is the unified compat endpoint: it infers the provider from the
model field and always returns an OpenAI-shaped response.
curl -s -X POST \
"https://<your-gateway-host>/v1/myapp/production/compat/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <GATEWAY_TOKEN>" \
-d '{
"model": "claude-opus-4-6",
"messages": [{"role": "user", "content": "Hello!"}]
}'
A successful response:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "claude-opus-4-6",
"choices": [{"message": {"role": "assistant", "content": "Hello! How can I help?"}}],
"usage": {"prompt_tokens": 10, "completion_tokens": 9, "total_tokens": 19}
}
-> The gateway resolves the provider from the model name, forwards the request, and
returns the response. Change the model to switch provider — gpt-4o routes to
OpenAI, claude-opus-4-6 to Anthropic, grok-3 to xAI — without changing anything
else. See Providers overview for the full resolution
rules.
💡 Note: To pin one exact provider with no model-name guessing, replace
compatin the path with the provider name, for example/v1/myapp/production/openai/chat/completions.
Authenticating
The gateway accepts the token in any of three headers, checked in this order — the first one present wins:
| Order | Header | Form | Compatible with |
|---|---|---|---|
| 1 | x-aig-token |
myra_xxxx |
Gateway-native |
| 2 | Authorization |
Bearer myra_xxxx |
OpenAI SDKs |
| 3 | x-api-key |
myra_xxxx |
Anthropic SDKs |
This is why the gateway drops in as a replacement for the OpenAI or Anthropic API without changing how your client sends credentials.
Using the OpenAI SDK
Point the OpenAI SDK at the gateway by setting the base URL to the compat
endpoint and the API key to your gateway token:
from openai import OpenAI
client = OpenAI(
base_url="https://<your-gateway-host>/v1/myapp/production/compat",
api_key="<GATEWAY_TOKEN>",
)
resp = client.chat.completions.create(
model="claude-opus-4-6",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
The SDK appends /chat/completions to the base URL and sends the key as
Authorization: Bearer, both of which the gateway accepts unchanged.
💡 Note: To route Claude Code or another Anthropic-native client through the gateway, see Claude Code.
Streaming
Set "stream": true to receive a Server-Sent Events stream of OpenAI-shaped chunks,
terminated by a data: [DONE] sentinel:
curl -sN -X POST \
"https://<your-gateway-host>/v1/myapp/production/compat/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <GATEWAY_TOKEN>" \
-d '{
"model": "claude-opus-4-6",
"messages": [{"role": "user", "content": "Count to three."}],
"stream": true
}'
See Inference API — Streaming for the finish_reason
normalisation and the streaming guardrail-block behaviour.
Handling errors
Errors return a structured JSON envelope and set the X-AIG-Error response header
to the same code:
{
"error": {
"code": "invalid_request",
"message": "Human-readable detail about what went wrong"
}
}
Branch your client logic on the stable code, never on the human-readable
message. A few you will meet early:
| Code | HTTP | Cause |
|---|---|---|
unauthorized |
401 | Missing, invalid, or wrong-gateway token. |
token_expired |
401 | The token passed its expiry — mint or refresh a token. |
token_revoked |
401 | The token was revoked — mint a new token. |
rate_limited |
429 | The sliding-window rate limit was exceeded. |
quota_exceeded |
429 | A budget cap was reached. |
guardrail_blocked |
400 | A guardrail blocked the request (or, in streaming mode, a 200 with a synthetic SSE error chunk). |
The full list is in Error codes.
Next steps
- Inference API (/v1) — the complete request body, streaming, and header contract
- Providers overview — supported providers and how
compatresolves the model - Authentication — token lifecycle, scopes, and the security model
- Error codes — every error code, cause, and HTTP status
- Playground — try requests from the admin UI before wiring up code
See also
- Quick start — the UI-first setup walkthrough
- Setting up a production tenant — the full tenant-to-request checklist