Docs chatbot (POST /docs/v1/chat)
A public, unauthenticated endpoint that answers questions about the AI Gateway product from the official documentation, grounded in the docs corpus with inline citations. It backs both the in-app "Ask the docs" panel and the public documentation-site widget. No token or session is required.
The endpoint is a dedicated front door: it can only reach the documentation assistant. It cannot be used for general-purpose inference, and it cannot reach any tenant, gateway, or agent other than the documentation one — the target is pinned server-side, so nothing in the request selects it.
Request
POST /docs/v1/chat · Content-Type: application/json · no Authorization
messages— required, a non-empty array of turns (the client keeps the short conversation history and resends it each call).- each turn:
{ "role": "user" | "assistant", "content": "<string>" }. rolemust be exactlyuserorassistant;system/developerturns are rejected (the assistant's grounding instructions cannot be overridden by the caller).contentmust be a string.stream— optional boolean.truestreams the answer token-by-token as Server-Sent Events (Content-Type: text/event-stream): OpenAI-styledata: {chat.completion.chunk}frames whosechoices[0].delta.contentcarries each token, adata: [DONE]terminator, and: hbheartbeat comments to ignore. Absent/falsereturns one buffered JSON completion. To stop mid-stream, the client closes the connection (e.g.AbortController) — the server honours the abort.
CORS is open (Access-Control-Allow-Origin: *); the endpoint reads no cookie and carries no
credentials, so it may be called from any web origin (e.g. an embedded docs-site widget).
What is rejected (HTTP 400)
The body is validated at the trust boundary and fails closed — an absent and a malformed value are both rejected, never defaulted through:
- a missing/empty body, or a body that is not a JSON object;
messagesabsent, not an array, or empty;- a message that is not an object, whose
roleis notuser/assistant, or whosecontentis not a string (a missingcontent, a number, or an object all reject); - a request with no
userturn; - more than 20 turns, a single message over 8 000 bytes, or a conversation over 24 000 bytes total.
Response
200 OK with an OpenAI-compatible completion (or an SSE stream when stream:true); the answer text
carries inline [file | page | section] citations to the documentation — clients may turn these into
links to the live docs page. If the docs do not cover the question, the assistant says so rather than
guessing.
Client rendering (how the answer is displayed)
The answer content is Markdown. The in-app panel and the public docs-site widget both render a
safe Markdown subset — headings, paragraphs, bold/italic, inline and fenced code, ordered/unordered
lists, links, and soft line breaks. The model answer is treated as untrusted at the render boundary:
- it is turned into DOM nodes (never injected as HTML), so any
<tag>or<script>in the answer shows as literal text and can never execute; - link targets are scheme-checked —
http(s),mailto:, and relative/anchor links render as links (external ones open in a new tab,rel="noopener noreferrer"); any other scheme (javascript:,data:,vbscript:, protocol-relative//host, or control-char-obfuscated variants) renders as inert text, never a clickable link; - anything outside the subset (tables, images, raw HTML) falls back to plain text.
A custom client should apply the same policy; do not render the answer as raw HTML.
Feedback
POST /docs/v1/feedback — also public, anonymous, credential-less (CORS *). Records a thumbs rating on
an answer:
rating— required, the number1(helpful) or-1(not helpful). A string/boolean/absent value is rejected (400) — fail-closed.question/answer— optional strings (length-capped server-side).
Returns 204 on success, 400 on an invalid body, 429 when the per-IP rate limit is exceeded. The IP
is recorded only as a coarse bucket (never a raw address); the timestamp is server-set.
Rate limits & errors
| Status | Meaning |
|---|---|
429 |
Rate limit exceeded — a per-IP cap (15 requests / 60 s) and a global cap (100 requests / 60 s) protect the shared assistant. Honour Retry-After. |
400 |
The request body was invalid (see above). |
503 |
ask_docs_unconfigured — the documentation assistant is not provisioned in this environment. Retry later. |
The endpoint runs the same billing, logging, and safety pipeline as authenticated inference; every call is recorded. It has no web search, file fetch, code execution, or external-tool access — it only retrieves from the documentation corpus.