Conversations API
The conversations API manages chat conversations, messages, attachments, presets, and stored memories. All endpoints require an authenticated admin session.
Conversations are user-scoped: a non-admin caller can only access conversations they own or that are shared with them through a project.
While impersonating ("View as User"), these endpoints are refused with
403impersonation_forbidden— a user's private chats, messages, attachments, images, knowledge-file downloads, comments, live sessions and memories are not readable or writable act-as. See Admin impersonation.
Listing conversations
GET /admin/v1/conversations
Returns the conversations of the calling user.
Optional query parameters:
| Parameter | Description |
|---|---|
limit |
Maximum rows. |
offset |
Paging offset. |
archived |
1 to return archived conversations only; otherwise active conversations are returned. |
Each row includes shared_in_project (0/1) — whether the conversation is shared into its project's feed (see Sharing a conversation with a project). It is 0 for conversations that are not in a project.
Searching conversations
GET /admin/v1/conversations/search?q=<TERM>
Full-text search over the calling user's conversations. Rows carry the same fields as the list endpoint, including shared_in_project (0/1).
Where the search term goes. By default the term is matched against conversation titles inside the deployment — it does not leave. A tenant that has configured a semantic-cache embedding endpoint on one of its own interfaces additionally gets meaning-based matching, which means the term is sent to that endpoint. Two rules bound that, and neither is configurable:
- The endpoint is only ever one configured on your own tenant's interfaces. Another tenant's endpoint and API key are never used, even if yours has none configured (in that case you simply get title matching).
- If any of your conversations belongs to a project on a restricted access tier (Local only or PII mandatory), the search term is not sent anywhere: the search falls back to title matching, which runs entirely inside the deployment. The same applies if that check cannot be completed — the term is withheld rather than sent on an unproven assumption.
The title of a conversation is treated the same way: a conversation in a Local only or PII mandatory project never has its title sent to an embedding endpoint, and a conversation whose project binding cannot be resolved is treated as restricted.
Creating a conversation
POST /admin/v1/conversations
The client posts routing hints, not a resolved gateway and model. The server is the sole authority on which gateway the conversation lands on.
How the server picks a model for Auto
When no concrete model is pinned (hint_tier: "auto" and no hint_model_override), the server resolves the model itself, from the tenant's own catalogue. The rules, in order:
- Only production gateways are candidates. A gateway with
purposeother thanproductionnever receives traffic. - A mandated request — or a cloud pick that wants protection — picks a protecting gateway. If the project's access tier is PII mandatory or the tenant enforces masking, the server restricts the candidate gateways to those that actually scrub personal data on the request — using the same test the inference pipeline enforces with, so it can never select a gateway the request would then be refused on. If none qualifies, it widens to gateways that have any personal-data protection active (for example a keyword list); only if neither exists does it fall back to the full set. When the pick comes from that narrowed set the usual PII-twin swap is skipped. A plain
hint_pii_preferencewithout a mandate restricts the candidates only when the resolved model is a cloud provider; a local first-party model (the usual Auto pick) is served on the non-masking gateway by default — even when it uses web search — because masking a first-party leg only corrupts your own data. Enforcement is the sole thing that forces a local pick onto a scrubbing gateway. - Local models are always available. The Myra-hosted local fleet needs no provider key and is routable from every gateway, so it is part of every gateway's catalogue even though it has no stored provider credential.
- Proven tool calling comes first. Chat has web search and file tools enabled by default, so a model that cannot call tools is not a usable default. The server therefore looks for a verified tool-capable model across the preferred providers — local fleet first, then Anthropic — before it considers a non-verified one. A tenant whose local fleet has no verified model but who also uses Anthropic will therefore get an Anthropic model: a working tool-capable default beats a local one that stalls mid-answer.
- Then provider preference, then cost. If no preferred provider offers a verified model, the server takes any chat model from a preferred provider, then a verified model from any provider, then any chat model at all. Gateways are always walked in name order, and within one gateway the cheapest model wins; equal prices prefer the newer model of a family (so a new release is not silently undone by its predecessor once their prices converge), then break remaining ties by model name and then by provider name, so the same catalogue always yields the same choice.
- Plan limits apply. On a subscription plan with a model list, only models included in that plan can be chosen — both for Auto and for a project's preferred model.
- Non-chat models (embedding, reranking, image, speech) are never chosen, and a deprecated model is never chosen.
In the current fleet this resolves to the local Qwen 3.6 model. A project's preferred model (see Projects) overrides the choice when the pinned model is routable, including when it is a local model.
Access-tier enforcement during a database fault. The tier that governs a conversation is resolved server-side on every request. If that lookup fails transiently, the gateway fails closed — it applies the most restrictive treatment rather than serving the request unrestricted — but only for a conversation it cannot prove is project-free. A chat that is not in a project is not silently treated as Local only, so a database blip cannot turn ordinary chats into "this project only allows local models" refusals.
A Local only project constrains the choice. For a project on the Tier-1 Local only access tier the server considers only local models at every step above — it never proposes a cloud model that the inference guard would then refuse. A preferred model that is not local is ignored rather than adopted (the local choice stands). If the project is Local only and no local model is currently routable — none served, or all of them excluded by the gateway's residency/allowlist or by the subscription plan — the request fails with 409 no_local_route (the body carries the machine string in both error and a sibling code), a distinct code from no_runnable_route, so the client can say that this project only allows local models rather than reporting a missing gateway. The tier is read server-side from the project row and is never taken from the client; it binds a conversation even for a caller who is no longer a member of the project.
| Field | Type | Required |
|---|---|---|
hint_tier |
string | no — routing tier; legacy fast/balanced/deep map to auto. |
hint_task_slug |
string | no — chip slug from the hero; steers the system prompt only, does not pick a model. |
hint_model_override |
object | no — { gateway_id, model } to pin a concrete model. The gateway must be one that routes (purpose: "production"): a caller who is not a platform administrator naming a test / benchmark / archived gateway is refused 400 model_override_gateway_not_visible and nothing is created — such a pin used to be accepted and produced a conversation that failed on every send. The same rule applies to the raw gateway_id on the conversation PATCH — but there only when the request MOVES the conversation to a different gateway, so a request naming the gateway it already sits on is never refused — and on a share-continue create (which has no committed gateway, so any supplied id is a move); both answer 403 gateway_not_allowed_for_project. A platform administrator may still pin a non-production gateway deliberately (triage), the same exemption that lets them pin across tenants. |
hint_pii_preference |
boolean | no |
project_id |
string | no |
web_search |
integer | no — 0 or 1. |
title |
string | no |
ghost |
boolean | no |
memory_disabled |
boolean | no |
personal_style_disabled |
boolean | no |
source_share_token |
string | no — when provided, the gateway forks from the referenced shared conversation; the response describes the new fork. Must be a string when present: a truthy non-string value (a number, an object, or true) is rejected with 400, while a falsy value (false, null, or "") is treated as absent and the request becomes a normal create. |
gateway_id |
string | conditional — required only on the fork path (when source_share_token is present). Must be a string (400 otherwise), must exist, and — for a non-admin caller — must belong to the caller's own workspace AND be a gateway that routes (purpose: "production"); anything else is refused with 403 gateway_not_allowed_for_project before anything is written. |
Continuing a shared conversation
POST /admin/v1/conversations with source_share_token creates a new conversation in the caller's account holding a copy of the shared transcript.
The continuation inherits the source conversation's project — or it is refused. A conversation that lives in a project carries that project's access tier, which is what guarantees where its data may go. Copying its transcript into an unbound conversation would strip that guarantee, so the server resolves the source's binding first and decides:
| Source conversation | Continuing user | Result |
|---|---|---|
| not in a project | anyone | created, unbound |
| in a project, and the user is authorized on it (member, group grant, org share, or platform admin) | created and bound to that project | |
| in a Configurable (Tier‑3) project | not authorized | created, unbound — that tier imposes no residency or PII constraint |
| in a Local only (Tier‑1) or PII mandatory (Tier‑2) project | not authorized | 403 share_source_project_forbidden |
| binding cannot be confirmed — the project belongs to another workspace, was deleted, or its tier is unrecognized | anyone | 403 share_source_project_forbidden |
| the binding lookup fails transiently | anyone | 503 project_tier_unresolved — retryable, and deliberately not the permanent refusal |
| the access check fails transiently | 503 project_access_unresolved — retryable; a database fault is never treated as a denial |
The refusal names no project, workspace, tier or owner, and "not authorized" and "could not confirm" share one code on purpose, so the response is not an oracle over a stranger's project. It does disclose one bit: that the source is in a project you are not authorized on.
The binding is read under the source conversation's own workspace, derived server‑side from the share record — never from anything the client sends — so two people holding the same link always get the same answer. A continuation is only bound when that workspace is also the workspace of the gateway it will run on, which is the workspace the inference path enforces the tier under; that keeps the decision made here and the decision made at send time identical by construction.
Also on this path:
- The continuation is created on the validated
gateway_id, not on the gateway recorded in the frozen snapshot (which belongs to the sharer and can be in another workspace). - It carries the snapshot's last assistant model and sets
model_unlocked = 1, the same one-shot picker waiver the in-account fork uses — the copied messages would otherwise freeze the picker on a model the recipient never chose. The first real turn clears it. - A
project_idin the request body is ignored; the binding is the server's answer about the source. - The imported messages are a frozen copy. The binding governs the conversation's future turns; the original project of the copied text is not recorded, and project files referenced by the transcript are not carried (an authorized continuer regains access to project knowledge as a consequence of the binding, nothing more).
- A continuation is all or nothing: if any message of the snapshot fails to import, the partially-built conversation is removed and the request returns
500(could not import the shared conversation) rather than handing the recipient a transcript with holes it is told is complete. - Only the message text crosses. Attachments, generated documents and images, conversation-owned files and tool results are not copied — the share token is a messages-only credential by design. The continuation is therefore flagged
share_forked: 1in the conversation object (server-derived, read-only, never accepted from a request body) — set to1only when the frozen snapshot actually imported at least one user/assistant message; an empty or system-only snapshot stays0and gets no notice — and the model is told plainly, in its instructions, that the file artifacts the transcript refers to are not available to it and that it must not restate their contents or present a reconstruction as the original. Without that notice a model reads a transcript saying "here is the v2 I generated" and answers as if it had the file — which is exactly how a production document lineage came to rest on a reconstruction. The flag rides along when the recipient forks the continuation ("start a new chat from here"), because the fork copies the same file-less messages. A conversation that carries a genuinely readable file — one listed in its project knowledge, or attached after the continuation — is unaffected: the notice says so explicitly, soread_filekeeps working on the very path an authorized binding restores. - Retention: a bound continuation in the same workspace inherits the project's deletion period. An unbound continuation does not — it falls back to the user/workspace period.
Limits of this control. It guards the in-product continue path against accidental egress; it is not a data-loss-prevention wall.
Two of the ways around it are now closed. Minting a link for a conversation in such a project is refused, and the public read is refused as well — so links that already exist, and links whose project is tightened after the link was sent, stop resolving. And releasing the conversation from the project — detaching it, or moving it elsewhere — requires owner rank on that project, so the person who published the link cannot simply detach the source and let every outstanding link continue unbound.
What remains, deliberately: deleting the project does detach every conversation that belonged to it, after which those conversations carry no restriction. Deleting requires owner rank — the same rank that could set the project's tier to user configurable and release everything anyway — so this is an authorised act, not a bypass; it is recorded in the audit log (project.conversations_released) and the delete dialog states the consequence. A transcript's text can also always be copied out by anyone who can read it; the tier is a guardrail against accidental egress, not a wall against a determined reader.
Prompt-example chips
GET /admin/v1/easy/chips
Returns the suggestion chips shown on the /easy hero. Each entry carries a slug, a label, an icon, and one or more example prompts; the response is scoped to the caller's tenant, with any per-persona overrides applied. Only enabled chips are returned, and the server-side steering system prompt attached to a chip is never included in the response. A tenant with no configured examples receives the curated defaults.
The chip a user picks is sent back on conversation creation as hint_task_slug, which steers the system prompt only — it does not select a model.
Getting a conversation
GET /admin/v1/conversations/<ID>
Returns the conversation with its messages (always a JSON array). Read access is the owner or a
member of the project the conversation is shared into.
| Status | Body | Meaning |
|---|---|---|
200 |
the conversation object | — |
404 |
{ "error": "not_found", "code": "conversation_not_found" } |
The conversation does not exist, was deleted (in another tab, on another device, or by retention), or is not visible to the caller. One answer for all three — no existence oracle. |
500 |
{ "error": "<db error>" } |
A database fault; retryable. Never folded into the 404. |
code: "conversation_not_found" is the typed signal the app keys on: a conversation it still has
open that answers this way was deleted elsewhere, so the app clears the thread, drops it from the
list, shows a non-error "deleted in another tab" notice and returns to the start page — the same
path a send, a regenerate, an edit or the post-turn reconcile takes when their own writes answer
with the same code (see Messages). The read first waits (up to a few seconds) for any
turn that is still being committed for this conversation so it never races ahead of the gateway's
persist; a conversation that was deleted while a turn was streaming does not wait — its orphaned
claim row is ignored and the 404 is immediate.
Updating a conversation
PATCH /admin/v1/conversations/<ID>
Patchable fields: title, model, system_prompt, max_tokens, gateway_id, starred (0/1), memory_disabled (0/1), personal_style_disabled (0/1), web_search (0/1), picked_tier (or null to clear an explicit mid-chat model pick), archived_at (timestamp or null), and project_id.
Patching project_id moves the conversation to another project (or detaches it with null). project_id must be a non-empty string or null — any other type (a number, an object, "") is rejected with 400, on the platform-admin path too.
Attaching requires membership: a non-admin caller must already be a member of the destination project, otherwise the request returns 403.
Releasing — detaching the conversation, or moving it to a different project — is gated when the conversation's current project restricts data residency (Local only or PII mandatory). Attaching is gated because the conversation would gain the project's knowledge; releasing is the direction that drops a residency guarantee, so it requires owner rank on that project (platform admins bypass, as everywhere). A project member who is not an owner gets 403 project_release_forbidden and the committed binding is unchanged. An unrestricted project detaches freely, exactly as before.
If the conversation's current project cannot be confirmed at all (it was deleted, or it belongs to another workspace), every project_id change is refused with the same 403 for non-admins: nobody holds owner rank on a project that no longer resolves, and moving the conversation anywhere would drop the restriction that project carried. If any of the three authoritative reads in that chain — the conversation's gateway, its project's tier, or your access to that project — fails transiently, the request returns 503 (project_tier_unresolved / project_access_unresolved) — never a silent detach.
Re-sending the project the conversation is already in is not a release and is never gated: the gate keys on an actual change, not on the field being present.
An authorised release is written to the audit log as conversation.residency_released, with the project it left and that project's tier.
temperature is no longer accepted and is silently ignored.
Route policy validation
A patch that changes the conversation's route target — model, gateway_id, a routing hint_* field, or project_id — is validated server-side against the post-patch project and gateway before it is persisted. The client picker is a convenience, never the authorization boundary: the server evaluates the resulting (model, gateway_id) with the same tier/PII predicates the inference path enforces one step later, so a pin can never be saved that would make every subsequent message fail. Validation covers a project-move-plus-model in a single request (the post-patch project is used), and runs on the resolved values (a routing hint is re-resolved first, then the result is validated).
Rejected (the patch is refused and the committed route is left unchanged):
400 model must be a string— a presentmodel(top‑level, or insidehint_model_override) that is not a string and not JSONnullis malformed and is rejected at the boundary, before routing or persistence, on any tier. A string pins;nullor an absentmodelis a no‑op (the pin is left unchanged —null(not an empty string) is the deferral signal, and it does not clear the pin). An empty string resolves to no provider: on alocal_onlyproject it is refused (403, below); on other tiers it defers to Auto.403 model_not_allowed_for_project— the resulting model resolves to a non-local provider on alocal_only(Tier‑1) project. Only locally‑hosted models are permitted; pick a local model. A first‑party (Myra/local) pick is always permitted here.403 gateway_not_allowed_for_project— the resulting gateway cannot satisfy the conversation's PII policy (apii_mandatoryproject, or a tenant with mandatory PII masking, pinned to a gateway that does not scrub request‑phase PII), or thegateway_idis unknown (any caller), or — for a non-admin caller — the request MOVES the conversation to a gateway that does not route (purposeother thanproduction; the conversation's already-committed gateway is never re-refused, so an unrelated edit of an existing conversation is unaffected), or belongs to another tenant (non‑admin callers only — platform admins may route through another tenant's gateway).503— a transient database/config fault while validating the pin (retryable); the patch fails closed rather than skipping the check.
Not validated (unaffected): a title/starred/web_search/memory_disabled/personal_style_disabled/archived_at-only patch (it does not change the route target), and a non-owner patch (owner-scoped no-op). A patch that does not pin a concrete gateway defers the check to send time.
Deleting a conversation
DELETE /admin/v1/conversations/<ID>
Soft-deletes the conversation (owner-scoped) and best-effort reaps its file
space (a reap failure is logged, and the deleted-state filter keeps orphans
unreadable) and any persistent code-interpreter kernel. The <ID> path segment is untrusted input;
ownership is enforced server-side in the delete itself — the client is never the
authz boundary.
200 {ok:true}— idempotent: returned for an actual delete AND for any owner-scoped miss (a foreign user's id, an unknown id, or an already-deleted id). A non-owner calling this endpoint gets the same 200 shape while the row — and the owner's live code-interpreter kernel — is untouched; there is deliberately no existence oracle.500 {error}— a storage fault. The delete did NOT happen; retry. (Previously a fault was swallowed as a dishonest200 {ok:true}.)
Forking a conversation ("new chat from here")
POST /admin/v1/conversations/<ID>/fork
Creates a new conversation that copies every committed message of <ID> up to
and including a chosen branch-point message, re-references the files those messages
carry, and continues in the same project with the source's model preserved. This
backs the "New chat from here" action on an assistant reply: the user branches the
thread and keeps the full prior context and the uploaded files.
Request body:
branch_message_id— required, non-empty string. The id of the message (the clicked assistant reply) that is the last message copied into the fork.
Behaviour:
- Copies the messages
≤the branch point (in conversation read order), the source'sproject_id,model,gateway_id,picked_tierand system/toggle settings. - Sets
model_unlocked = 1on the new conversation: the model picker is unlocked for the fork even though it carries messages (the normal mid-thread model freeze is waived), so the user can switch models. This is one-shot — the first real turn persisted afterwards clears it and the fork re-freezes like any other conversation. - Project-owned files are re-referenced (the same file rows — no byte copy), so they stay authorized in the same project. Conversation-owned files (a project-less chat's own attachments / generated files / write_file artifacts) are duplicated under the new conversation so the copied messages keep valid, authorized, downloadable, inference-visible file links.
- Returns
201with the new conversation (same shape as create/get, messages included).
Authorization and validation (the client is never the authz boundary):
- The caller may fork a conversation they own, or a conversation shared into a
project's feed that they have access to (the same access that lets them view it —
owner ∪ (shared_in_project ∧ project-access)). Project access is resolved through the one project-access resolver, so it counts a direct member, a member of a group granted on the project, and — while the tenant's org-wide sharing is on — any user reached by the project's organisation share (the same three channels that let them open the project and read its files). The fork is always a private copy owned by the caller (not shared into any feed). A conversation that is not owned by the caller and not a feed conversation of a project they can access — or an unknown<ID>— returns404 not_found(no existence leak). Access is same-tenant, so no cross-tenant copy is possible. branch_message_idis validated to be a committed message of<ID>. A missing or empty value returns400; an unknown, foreign, deleted, or still-streaming id returns404.- The copy runs in one transaction that takes the source conversation's row lock first — the
same order every message writer (edit, correction set and clear, delete) uses — and copies
each message's large text columns on the
database side, so a message the gateway could store can always be forked. A transient
serialization conflict (a deadlock or lock wait against a concurrent write) is retried
transparently; if it persists, the fork answers
503(retry later) rather than a500.
Sharing a conversation publicly
POST /admin/v1/conversations/<ID>/share
Generates a public read-only token and returns { "token": "<TOKEN>", "url": "<URL>" }. The url is the absolute link a recipient opens: /shared/<TOKEN> on this deployment's own app host (the same origin every other product link uses), so a link minted on one environment always resolves against that environment's data.
Refused for a restricted project. A share link is a persistent, unauthenticated read credential. Minting one for a conversation whose project restricts data residency (Local only or PII mandatory) — or whose project cannot be confirmed — returns 403 share_restricted_project and writes no share row; a transient failure to read the project's settings returns 503. Unlike the release gate on PATCH, this has no owner-rank exception: an owner who genuinely wants a public link can move the conversation out of the project first, which is gated and audited, whereas a mint that succeeded from inside the project would be indistinguishable from an accidental one.
GET /admin/v1/conversations/<ID>/share — returns the active share token, if any. It applies the same check, so a token minted before the project was tightened is not handed back as a live-looking URL for a link the public read now refuses.
DELETE /admin/v1/conversations/<ID>/share — revokes the share.
Reading a shared snapshot (public, unauthenticated)
GET /share/<TOKEN> — the public read-only snapshot the /shared/<TOKEN> page renders. No authentication. Served same-origin on the app host (the SPA fetches it relative, so a recipient's browser never needs a second origin), with permissive CORS (Access-Control-Allow-Origin: *) so it is also readable cross-origin.
The token is the only input; it is matched exactly (a 256-bit random hex string — not enumerable) and resolved fail-closed through the owning tenant's live state, so a revoked share or a decommissioned tenant is indistinguishable from an unknown token.
Responses:
| Status | Meaning | Body |
|---|---|---|
200 |
Snapshot found | { "title": string, "messages": [ { id, role, content, created_at, … } ] } — messages is always a JSON array (an empty conversation returns [], never {}); system turns are excluded. |
404 |
Unknown / revoked / expired token, the conversation was deleted, or the owning tenant is decommissioned | { "error": "not found" } |
403 |
The conversation's project restricts data residency, or its binding cannot be confirmed — permanent | { "error": "share_source_project_forbidden", "message": … } (the machine code, unlike the human string the other error bodies carry) |
503 |
Service starting (schema migration in progress) — retryable | { "error": … } (carries Retry-After + CORS) |
400 |
Missing token | { "error": "missing token" } |
500 |
Corrupt stored snapshot | { "error": "invalid snapshot" } |
A recipient client MUST treat a 404 as a genuine "no longer shared", a 403 as a permanent policy refusal (never offer a retry), and any other non-200 (a 503, a 5xx, a non-JSON body, or a network error) as a transient, retryable failure — never as "not found".
The residency check runs on every read, not only at mint time: that is what closes links which were already sent, and links whose project is tightened afterwards. The refusal deliberately names no project, tier or owner — the reader here is unauthenticated, so a finer answer would be an oracle over a stranger's project — and it shares one code with "cannot confirm" for the same reason.
What a token holder can still infer, stated plainly. The body says nothing, but the status does: someone holding a valid token can tell 403 (restricted, or unconfirmable) from 404 (unknown / revoked / deleted) and from 200, and can poll it. So they can observe that a project was tightened, or that the conversation moved, without learning which project, which tier, or whose. That is an accepted trade: collapsing the refusal into 404 would tell a legitimate recipient that the owner revoked their link when they did not, and would deny the SPA the honest "not available" state. The token itself is a 256-bit random string, so none of this is reachable without one. Every response carries Cache-Control: no-store, because the decision is taken per request and must never be replayed from a CDN.
Deleting the conversation now also takes its public link down; previously the snapshot stayed readable at /share/<TOKEN> after the owner deleted the thread. The same resolver backs continuing a shared conversation, so continuing a share whose source was deleted returns 404 as well.
Sharing a conversation with a project
POST /admin/v1/conversations/<ID>/share-project — shares a conversation the caller owns into its project's feed, so everyone with access to the project can read it. The conversation must already belong to a project; otherwise the request returns 400.
Who may share. The caller must have access to the project through any channel — a direct invite, a group grant, or the organisation share — resolved by the one project-access resolver (the same access that lets them open the project and chat in it). A caller with no project access via any channel is refused 403 "not a project member"; a transient failure to resolve access returns 503 (retryable), never a silent 403. Because the conversation read above is owner-scoped, a caller can only share a conversation they own.
DELETE /admin/v1/conversations/<ID>/share-project — removes a conversation from the feed. This reverts it to private to its author (shared_in_project = 0); the conversation is not deleted, and any files it generated revert to Only you visibility (the generated-file share gate keys purely off the source conversation's shared state). Un-sharing is permitted for the conversation's author, any effective project owner (a direct owner or an owner-group grant), and platform admins — so the person responsible for a project can remove another member's conversation from the feed they own (for example after a name-invited member leaves). The refusal is deliberately split so it cannot become an existence oracle: a non-author who cannot see the conversation at all — it has no project, it is not on the feed (shared_in_project ≠ 1), or they have no access to its project via any channel — gets the same 404 not_found as an unknown id. Only a caller who can see it on the feed but is a plain editor/viewer (neither author nor owner) gets the explanatory 403; a transient failure to resolve access returns 503. Un-sharing a conversation that is already private is a graceful no-op.
Neither endpoint attaches the conversation to a project — patch project_id to move it.
Live shared session (room)
Turns a conversation that is shared into a project (see above) into a live room: project members with read access watch new messages appear in real time, see who else is present, and take turns writing through a single-driver write lock.
Only the current driver may write. Turn-taking is enforced server-side — the client is never the authorization boundary. The lock is user-keyed with a lease: a driver who disconnects without releasing is reclaimable once the lease lapses.
In the app, any participant who may drive (the owner, or a project member with the owner/editor role) can take the wheel when the seat is free and hand it over when done; the composer is enabled only for whoever currently holds the seat. Taking the wheel only succeeds when the seat is free or its lease has lapsed — it never forcibly steals an active driver's turn. A viewer-role member can watch but never drive.
Start / stop
POST /admin/v1/conversations/<ID>/live — start the session (owner only). The conversation must already be shared into a project; otherwise the request returns 409. The starter becomes the initial driver. Response: { session, roster }.
DELETE /admin/v1/conversations/<ID>/live — stop the session (owner only). Clears the roster and driver lock.
Snapshot
GET /admin/v1/conversations/<ID>/live — current { session, roster, can_drive } for the caller. Read access required (owner or project member); a non-member receives 404 (no id-enumeration surface). can_drive is true for the owner or a project member whose effective role is owner/editor.
Live stream (SSE)
GET /admin/v1/conversations/<ID>/live/stream — a long-lived text/event-stream. Read access required (non-member → 404, before any bytes stream). Each event is a JSON object with a type:
type |
payload |
|---|---|
hello |
{ you, can_drive, session, roster } — sent once on connect |
presence |
{ session, roster } — the current roster + driver, each tick |
message |
{ message } — a message committed after connect (history is loaded via the normal conversation GET) |
session_ended |
the session was stopped |
The stream DB-tails the conversation (poll interval ~2s); there is no cross-container push signal by design (the shared database is the source of truth across the split multi-container deployment). The subscriber's open connection is its presence — a dropped connection ages out of the roster within ~12s. Message delivery is message-level (a committed message appears within one poll); token-level streaming to watchers is not part of Phase 1.
Each connection re-seeds its message-tail cursor from the current maximum message on hello (it streams only messages committed after connect). Clients therefore reconcile on every (re)connect: a watcher whose stream drops (laptop sleep, network blip) and reconnects re-loads the conversation via the normal GET, so no message committed during the gap is lost. A watcher who is already viewing an inactive conversation re-polls the snapshot while no session is live, so it observes the owner starting a session without a manual reload. When the owner leaves (tab close / network drop) the session stays active until an explicit stop or the 30-minute cap, but the owner ages out of the roster; watchers surface an honest "owner away / paused" state rather than sitting on a stale "no one is driving" until the cap.
Driver (write) token
POST /admin/v1/conversations/<ID>/live/driver — request the turn. can_drive-gated (a viewer-role member receives 403). 200 { acquired: true } on success; 409 { acquired: false, driver_user_id } when another participant holds a live lease; 409 when no session is active.
DELETE /admin/v1/conversations/<ID>/live/driver — release the turn (only the current holder has effect; a stale caller is a harmless no-op → 200).
Writing while a session is live
When a session is active, an inference turn on the conversation (/v1 chat path) and an appended message (POST /admin/v1/conversations/<ID>/messages) are rejected for anyone who is not the current driver — error code live_not_driver (403). A transient failure to resolve the session state also blocks the write (fail-closed). Content shown in the room is the stored conversation content already visible to project members via the conversation read path — the live room adds no new content exposure and honors the same PII masking and retention as the persisted conversation.
Accepted / rejected input: the conversation id and driver token are resolved server-side against the caller's session and project membership; a forged or borrowed id resolves to 404/403. A non-owner cannot start or stop a session; a viewer-role member cannot acquire the driver; a non-driver cannot write.
Submitting and reading conversation feedback
PUT /admin/v1/conversations/<ID>/feedback — submit a 1–5 rating (5 = best, 1 = worst) with an optional comment.
Accepted / rejected input: rating must be a number (or a numeric string) in 1..5;
anything else — missing, null, a non-numeric string, NaN, Infinity, or out of range
— is refused with 400. comment is optional and coerced rather than rejected: a
non-string (null, number, boolean, object) becomes empty, and a long comment is
truncated to the column limit, so a malformed comment can never discard the rating.
GET /admin/v1/conversations/<ID>/feedback — read the submitted feedback.
Listing all chat feedback (admin)
GET /admin/v1/feedback
Required permission: FEEDBACK_TRIAGE (a platform-admin permission).
PATCH /admin/v1/feedback/<ID> — update the status of a feedback entry. Required permission: FEEDBACK_TRIAGE (a platform-admin permission).
Conversation summaries
GET /admin/v1/conversations/<ID>/summaries — list stored summaries. Readable by
anyone who can view the conversation — the owner or a member of the project it is
shared into (the same read access as GET /admin/v1/conversations/<ID>). Returns 404
not_found only when the conversation truly does not exist or the caller cannot see it,
never a misleading 404 for a transient fault: a database error on the access check
returns 500 (retryable), while a database error reading the summaries list itself
returns 503 (retryable). Both fail closed rather than mislabelling an outage as
not-found.
POST /admin/v1/conversations/<ID>/summaries — store a new summary entry.
Accepted / rejected input: summary_text (up to 65535 bytes), first_message_id and
last_message_id (up to 36 characters each) are required and must be non-empty strings. A
missing field, a JSON null, a number, a boolean, an object/array, an empty string, or a
value longer than its limit is refused with 400 before anything is written. The two
optional fields are coerced rather than rejected: message_count becomes a whole number
in range (a non-numeric value, NaN, Infinity, a negative, or one beyond the column's
range becomes 0; a fraction is truncated) and model_used becomes a string truncated to
128 characters (anything that is not a string becomes empty).
supersedes_ids is optional and, when present, must be an array of non-empty strings (an
empty array is legal — it asserts "no active summaries"); anything else is refused with
400. Supplying it turns on the concurrent-write check: a stale view is answered with
409 conflict_supersedes_stale instead of silently superseding a summary another client
wrote a moment earlier. Omitting it keeps the previous behaviour, in which this route
supersedes every active summary unconditionally — prefer to send it.
POST /admin/v1/conversations/<ID>/summarize — request the gateway to produce a summary now.
The summary call resolves its target provider and model through the same authority as a
normal inference dispatch: gateway-aware provider inference, then the gateway's routing
rules (the inferred provider is seeded into rule evaluation, and rules are evaluated on the
bare model), then the gateway configuration (Azure endpoint, custom provider base URLs). A
summary is therefore generated by the provider and model an inference turn on the same
gateway and model would use — a routing rule that remaps the provider or model applies here
too. A summary against a provider whose per-request key is unavailable returns 422, and a
provider/upstream failure is surfaced with the provider's own error shape (never a bare
500).
The conversation's access tier and the privacy filters apply here too. A summary is generated from the conversation's own text, so it is treated as an outbound request like any other turn — this was not always true, and the gap is worth stating plainly because it was invisible: compaction runs automatically after every turn and can carry tens of thousands of tokens of verbatim conversation per call.
- A Local only (Tier-1) project's conversation is never summarised by a non-local model.
The attempt is refused with
403 model_not_allowed_for_projectbefore anything is sent. In practice this changes nothing for a working conversation: such a project's ordinary turns are already restricted to local models, so the summary uses one too. - If the conversation's project binding cannot be confirmed, the summary is refused with a
retryable
503 project_tier_unresolvedrather than being sent unrestricted — and deliberately not with the permanent "only local models" message, which would be untrue for a project that may not be restricted at all. - The gateway's request-phase privacy filters run on the summary text. Personal data is
masked before the text reaches the provider and restored afterwards, so the stored summary
reads normally while the provider only ever saw the masked form. A
scrub-style filter, which replaces a match with a fixed placeholder rather than a reversible token, is irreversible by design: the placeholder stays in the stored summary, because the model never saw anything else. - If the project mandates PII protection and the gateway has no filter configured, the
summary is refused (
403 gateway_not_allowed_for_project) rather than sent unprotected. - If a filter is temporarily unavailable and the gateway is set to refuse rather than skip
it, the summary is refused with a retryable
503 guardrail_unavailable— distinct from a content block, because nothing about the content was wrong. - If a content filter blocks the text outright, the summary is refused with
400 guardrail_blocked, the same verdict the equivalent chat turn would get.
The stored summary matters more than a single request: it replaces the messages it covers and is sent as context on every later turn, so it is the conversation's only remaining record of that history.
Accepted / rejected input. The request body is validated at the trust boundary before
any provider call, key fetch or database write. Every shape rejection below is a 400
with an explanatory error, never a 500 (authorization is separate: a gateway_id
outside the caller's tenant is 403, an unknown one 404, and a database fault while
resolving it 503):
| Field | Accepted | Rejected |
|---|---|---|
gateway_id |
non-empty string, and must belong to the caller's tenant | missing, null, number, boolean, object/array, empty string |
model |
non-empty string (a provider prefix such as openai/ is stripped); stored truncated to 128 characters |
missing, null, number, boolean, object/array, empty string |
messages |
a non-empty array of objects (still required for back-compatibility) — but its content is no longer used: the text summarized is read from the stored transcript span (see below) | missing, null, number, string, boolean, [], or any entry that is not an object |
first_message_id, last_message_id |
non-empty strings, at most 36 characters each, and must both resolve to messages in this conversation, in order | missing, null, number, boolean, object, empty string, longer than 36 characters; an id not present in the conversation (404 span_not_found); first positioned after last (400 span_inverted) |
supersedes_ids |
absent, or an array of non-empty strings (an empty array is legal — it asserts "no active summaries") | a value with any non-sequential key (an object with keys, a sparse or mixed table), or any element that is not a non-empty string (including NaN, numbers, null, objects, ""). An empty JSON object {} is indistinguishable from [] here and is treated as an empty array |
target_tokens |
any number; clamped to [100, 8000] |
— (a non-numeric value, NaN, or null falls back to the default; Infinity clamps to the maximum) |
previous_summary |
string | — (any other type is treated as absent) |
Because a stored summary supersedes the history it covers, these are fail-closed rejections rather than best-effort coercions: a body the gateway cannot interpret exactly must not produce a summary that silently omits part of the conversation.
The summarized text is read from the stored transcript, not from the request. The gateway
resolves first_message_id..last_message_id to a contiguous span of the stored conversation
and summarizes that — the messages array in the body is accepted for backward compatibility
but its content is ignored. This removes an unbounded, untrusted input: the request body can no
longer decide what text is sent to the provider. (A message whose stored content is not text is
refused with 400 span_content_invalid.)
The request is bounded, rate-limited and metered — before any provider call, key fetch, or privacy scan:
- Size ceiling (
413 corpus_too_large). The assembled corpus (the resolved transcript span plus anyprevious_summary) is refused if it exceeds either an absolute raw-byte cap (a worker-protection bound, independent of the model) or the resolved model's context window. This is a byte ceiling, not a message-count limit — a 500-message conversation whose text fits the window still summarizes. - Per-user rate limit (
429 rate_limited, withRetry-After). Each user's summarize requests are charged in token-equivalents against a per-minute budget; a tight loop is throttled after a handful of full-size compactions while ordinary automatic compaction is unaffected. - Budget enforcement (
429 quota_exceeded, withRetry-After). A tenant or gateway at itsbudget_usdlimit is refused before the provider is called — the route re-runs the same spend predicates the/v1path uses (the ledger read, the prepaid-wallet funds check, and the self-serve degrade window), so a funded wallet or an allowance still inside its degrade window carries the request through just as on/v1. - Account state (
402 trial_expired/402 subscription_inactive). A workspace whose trial has ended, or whose self-serve subscription is inactive past its grace period, is refused before the provider is called — the same predicate the/v1path uses, with the same rules: a served subscription (active or in grace) supersedes a stale trial date, grace warns rather than blocks, and a manual/enterprise workspace has no subscription gate at all. Decided before the budget gate, so a lapsed account reads "reactivate" rather than "over quota". These refusals carry noRetry-After— they do not clear with time, only with a billing action. One residual gap: per-user trial-credit exhaustion (trial_budget_exhausted) is still enforced only on/v1, not here. A trial user who has spent their personal credit while the workspace stays under its own cap is refused on the inference path but can still run (and be billed for) a summarize. Its spend remains bounded by the tenant/gateway budget caps and the rate limit above. - Real cost. A successful summarize records its actual
cost_usd(priced from the provider's reported usage) inrequest_logand debits the same wallet/quota scopes an inference turn does. A provider call that returned no usable output (an empty completion, or a response that could not be parsed) is not billed — that cost is absorbed rather than charged as an estimate.
Because the summary is a real provider dispatch, it enforces the same fail-closed
data-residency and provider-allowlist gates as an inference request, evaluated
against the gateway's inheritance-folded configuration (the tenant-level EU-residency /
allowlist floor applies even when the gateway itself does not set the flag). If the
gateway requires EU model hosting and the model resolves to a non-EU-hosted provider,
the request is refused with 403 data_residency_blocked; if the gateway enforces a
provider allowlist and the model resolves to a provider not on the approved list, it is
refused with 403 provider_not_allowed. Both gates run before any provider key is
fetched or decrypted, so a blocked provider never touches a stored credential.
Streaming variant (Accept: text/event-stream)
By default summarize returns the finished summary as a buffered 201 JSON row. When
the request carries Accept: text/event-stream, the gateway instead streams the
summary as it is produced (Server-Sent Events), so a long compaction shows live progress
instead of blocking. All of the gates above (residency, provider-allowlist, key
availability) still run before the stream is opened — a blocked or misconfigured
provider is returned as the same JSON error (403 / 422 / 404) with no stream
opened, exactly like the buffered path. Absent the header, the buffered 201 is
returned unchanged (old clients are unaffected).
Each event is a data: line whose JSON payload carries a type:
{"type":"delta","text":"…"}— an incremental fragment of the summary text.{"type":"retry","reason":"stream_truncated"}— non-terminal: the gateway discarded the attempt streamed so far (see One automatic retry below) and is re-running the summarize call. Everydeltasent before this frame is dead text; a client that reconstructs the summary fromdeltafragments MUST discard what it has accumulated and start again from the nextdelta. Onlydone.rowis authoritative.{"type":"done","row":{…}}— terminal success;rowis the storedconversation_summary(same shape as the buffered201body).{"type":"error","code":"…","message":"…"}— terminal failure.codeis one ofstream_incomplete(the upstream stream ended before a completion marker — e.g. a mid-stream reset),empty_summary(the model produced no content),sse_line_overflow(a single upstream line exceeded the 1 MiB cap),stream_error(an upstream read error),stream_truncated(the model hit an output cap and the summary was cut off mid-sentence),content_filtered(the model refused to summarise the segment),stream_aborted(the provider ended the stream on a terminal that is not a clean finish — e.g. a mid-stream providererrorevent, or a reason the gateway does not recognise),summary_too_long(the summary exceeded the 65535-byte stored limit),conflict_supersedes_stale(a concurrent compaction won the supersede race — the client refetches/summariesand continues, not a hard failure), ordb_error.
The buffered variant returns the four terminal-reason codes — stream_truncated,
content_filtered, stream_aborted, summary_too_long — as
502 {"error":"…","code":"…"}, so a failed compaction is classified identically in both
shapes. Three buffered terminals differ by design: empty_summary is a 502 without a
code, conflict_supersedes_stale is a 409 whose body nests {"error":{"code","message"}},
and a storage failure is a 500 with code: "db_error". The three transport codes
(stream_incomplete, stream_error, sse_line_overflow) exist only on the streaming
variant.
When a summary is stored
A stored summary supersedes every earlier summary for the conversation, and the client then sends only that summary plus the messages after its cut — so storing an incomplete one silently destroys the history behind it. The row is therefore persisted only when all of the following hold:
- the upstream stream reached a completion marker and no read error occurred (streaming variant only — a mid-stream reset looks like an ordinary end of body);
- the provider reported a clean terminal — an allowlist, not a denylist:
stop,end_turn,pause_turn,tool_calls/tool_use,stop_sequence, or no terminal at all. A length cap (length/max_tokens), a content filter, a providererrorevent, and any unrecognised reason are all refused; - the resulting text is non-empty after trimming;
- the text is at most 65535 bytes (the storage limit).
Anything else stores nothing and reports the matching code above; the conversation's existing summaries and full history are left exactly as they were, so the client shows a "could not compact" notice and retries later rather than losing context.
Drop a span from the assistant's context
POST /admin/v1/conversations/<ID>/context-drop — body
{ "first_message_id": "<ID>", "last_message_id": "<ID>" }.
Some spans can never be summarised: the model truncates every attempt, refuses the content, or produces more than the storage limit. Because the summary chain always starts at the oldest un-summarised message, one such span stops compaction for the whole conversation — the history then grows until the request hits the model's context limit. This route is the user's way out: it drops that span from what the assistant is SENT, without deleting anything.
Nothing is deleted. The messages stay in the conversation and stay visible; they simply stop being included in future requests, and the stored summary gains a server-authored note telling the model that a gap exists, so it cannot invent what was there.
The client names a span and nothing else. The stored summary text, the message counts and
the note are all composed server-side — a caller cannot write summary text through this route
(unlike POST /summaries, which exists for a client-authored summary and is kept separate for
exactly that reason). supersedes_ids is likewise derived server-side from the same read used
to compose the text, so the concurrent-write check cannot be disarmed by omitting it.
| Input | Rejected when | Status |
|---|---|---|
<ID> (conversation) |
not owned by the caller, or unknown | 404 |
<ID> (conversation) |
the lookup itself failed (DB fault) | 503 — retryable, never reported as "gone" |
<ID> (conversation) |
the conversation has no messages | 409 empty_conversation |
first_message_id, last_message_id |
not a non-empty string of at most 36 characters | 400 |
| the span | either id is not a message of THIS conversation | 404 span_not_found |
| the span | first_message_id comes after last_message_id |
400 span_inverted |
| the span | last_message_id is the newest message (would drop the entire conversation) |
409 span_too_recent |
| the span | it lies at or behind the active summary's cut (already summarised) | 409 span_behind_cut |
| the composed summary | would exceed 65535 bytes | 409 summary_too_long |
| — | a concurrent summarisation moved the active set | 409 conflict_supersedes_stale — refetch /summaries and retry |
On success: 201 with the new summary row. It supersedes the active summaries, keeps their
text verbatim (the real memory is never discarded to make room for the note), records
model_used: "context-drop" — no model produced it — and carries
dropped_message_count, which accumulates across drops.
dropped_message_count is the durable record of the gap. The note itself cannot live in
the summary text alone: every later compaction sends the current summary back to the model as
previous_summary with an instruction to fold it into the new one, so an LLM would decide
each pass whether the note survived. The gateway therefore re-renders the sentence from this
counter after every summary write (stripping any copy the model echoed back), which is also
what lets the UI say "N earlier messages are no longer available to the assistant" instead of
claiming they were summarised.
Output-cap clamp
The first provider call's max_tokens (derived from target_tokens) is clamped to the
model's own output ceiling before the request is sent — the value the catalog records for
that model, or, for a self-hosted model, whatever leaves room for the prompt within the
context window. Without the clamp an over-target request could ask for more output tokens
than the provider's model allows, which a provider such as OpenAI rejects outright (a failure
before any content), so the conversation would never compact. The retry below omits
max_tokens entirely, so the clamp applies to the first attempt only.
One automatic retry on truncation
A cut-off summary is usually caused by the gateway's own output cap (max_tokens,
derived from target_tokens), so stream_truncated on the first provider call
triggers exactly one retry — the same request with max_tokens omitted entirely, so
the model may use its full output budget. At most two provider calls are made per
request; a second truncation is reported, never retried again. Both attempts are recorded
in request_log (see below) and are distinguished by meta.attempt (1 or 2).
The retry covers Claude too: the gateway fills an omitted max_tokens with the model's real
output ceiling before the request leaves the gateway, so a retry that drops the field still
reaches Anthropic with a valid, full-budget value rather than its native max_tokens: Field
required rejection.
The retry is skipped — and the failure reported immediately — when it could not help: when
the client has already disconnected (streaming variant only, and only once the attempt has
emitted visible text — the buffered variant cannot observe a departed client at all), or when
the resolved provider substitutes a smaller default of its own so a no-cap retry would ask
for less and truncate again (currently bedrock).
Upstream SSE bytes are treated as untrusted throughout: over-long lines are rejected
(sse_line_overflow), a malformed frame is skipped rather than aborting the stream, and
the terminal reason is validated against the clean-finish allowlist before it can allow a
write (a non-string terminal such as JSON null is treated as "no terminal reported").
Provider coverage. The terminal-reason gate can only act on a reason the provider actually reports. The OpenAI-compatible family (including self-hosted models) and Anthropic report one; Gemini, Vertex, Cohere, and Bedrock currently do not, so a truncated summary from those providers is still treated as clean.
Logging (meta.kind = "compaction")
Every summarize call — buffered and streaming, success and failure — is now recorded in
request_log via the same emitter the inference dispatch path uses, so compaction
latency, token counts, and provider status are queryable. These rows are identified by
meta.kind = "compaction" (with the conversation id under meta.conversation_id and the
pre-compaction estimated input tokens under meta.tokens_before); they do not set the
compaction_triggered / compaction_tokens_* columns, which remain reserved for a
provider's native in-turn compaction. A logging failure never fails the summary.
Viewers cannot spend credit or reshape a conversation
A user with the viewer role may read their own chat and do owner-housekeeping on it, but may
not spend inference credit or reshape a conversation's content/context. Every route in the first
category below is refused with a typed 403 { "error": <prose>, "code": "viewer_read_only" } on
the admin session plane; the routes in the second category stay open. This is not a blanket
"read-only" — a viewer can still rename, delete, rate and detach their own data (none of which spend
credit or change what the model sees). The check runs before any conversation lookup or body
validation, so a viewer is refused identically for a real and a non-existent id (no existence
oracle), and the app disables the composer and hides the per-message mutate affordances for a viewer
as a matter of UX — the server, not the client, is the boundary. (The viewer role is separately
barred from all /v1 inference — see
Inference is refused for viewers; this admin-plane
gate closes the second door, the chat writes that reach a model through the session cookie.)
Gated (a viewer is refused viewer_read_only) — spends inference credit OR reshapes content/context:
| Route | Why it is refused |
|---|---|
POST /admin/v1/conversations |
creates a conversation (a fresh chat, or a share-continue) |
POST /admin/v1/conversations/<ID>/fork |
creates a conversation ("new chat from here") |
POST /admin/v1/conversations/<CID>/messages |
appends a message and triggers inference |
PATCH /admin/v1/conversations/<CID>/messages/<MID> |
edits a message or sets the display-only corrected_content overlay — no carve-out; a viewer writes nothing |
DELETE /admin/v1/conversations/<CID>/messages/<MID> |
trims a message (regenerate) |
POST /admin/v1/conversations/<CID>/attachments |
uploads a file, which runs image-analysis inference on an image |
POST /admin/v1/conversations/<CID>/summarize |
makes a real provider call (inference credit) to compact the thread — fired automatically by the app on open |
POST /admin/v1/conversations/<CID>/summaries |
persists a summary that supersedes memory (changes what the model sees on every later turn) |
POST /admin/v1/conversations/<CID>/context-drop |
drops a message span from the model's context and writes a server-composed summary — a context reshape |
PUT /admin/v1/conversations/<CID>/knowledge/<KID> |
rewrites a conversation artifact's extracted_text, which is re-sent to the model as knowledge |
POST /admin/v1/chat/files |
uploads a chat file, which runs qwen OCR and a provider Files-API upload (inference credit) |
POST /admin/v1/chat/transcribe |
speech-to-text: a real provider (Whisper) inference call on the fleet from raw audio (inference credit) |
Allowed (reads + owner-housekeeping — no chat inference, no context reshape): listing and
searching conversations, getting a conversation, reading messages and generated files/images
(GET/list/search); renaming a conversation (PATCH /admin/v1/conversations/<ID>); deleting
a conversation (DELETE /admin/v1/conversations/<ID>); rating it (PUT /admin/v1/conversations/<ID>/feedback);
Caveat: on a tenant with a semantic-cache embedding endpoint, search and rename
currently trigger a small embedding call (a cheap provider cost, with a local keyword fallback when
no endpoint is configured) — these are allowed to a viewer today; skipping the embedding for viewers
is tracked as a follow-up. They do not run chat inference or reshape a conversation's content.
deleting an attachment (DELETE /admin/v1/attachments/<AID>); minting and reading a public share
link; and conversation comments. Read-aloud (text-to-speech) is also intentionally allowed
(passive accessibility) — see Text-to-speech. A pure viewer never
owns a conversation (they can no longer create or fork one), so the owner-scoped housekeeping routes
are unreachable to them in practice; they stay ungated because the policy is content/credit, not a
blanket write ban, and a member later downgraded to viewer must still tidy their own history.
Messages
POST /admin/v1/conversations/<CID>/messages — append a message and trigger inference.
The write gate (POST, PATCH, DELETE)
All three message writes run the same server-side gate before touching a row, and every
refusal is a typed body the app acts on without reading prose. The <CID> / <MID> path segments
are untrusted input; ownership, project membership and the live-session driver lock are all
resolved server-side — the client is never the authz boundary.
| Status | Body | When |
|---|---|---|
403 |
{ "error": "…", "code": "viewer_read_only" } |
The caller holds the viewer role — refused every chat write that spends credit or reshapes context. Decided first, before any conversation lookup, so a viewer is refused whether or not the id exists (no existence oracle). See Viewers cannot spend credit or reshape a conversation. |
404 |
{ "error": "conversation not found", "code": "conversation_not_found" } |
The caller cannot see the conversation at all: unknown id, deleted (another tab, another device, retention), or never shared with them. Returned only for a conversation the caller has no read access to — never for one they can still read — so the app may safely treat it as "this conversation no longer exists" (clear the thread, drop it from the list, non-error notice, back to the start page, draft kept). |
403 |
{ "error": "…", "code": "live_not_driver" } |
The caller can read the conversation, a live shared session is active, and they are not the current driver (or their lease lapsed). The owner is not exempt. A refusal, not an error — the app shows it as a notice and does not report it. |
403 |
{ "error": "…", "code": "forbidden" } |
The caller can read the conversation (a project member it is shared with) but no session is active — off-session only the owner writes. |
500 |
{ "error": "internal error" } |
A database fault while deciding, or (DELETE) while deleting. Retryable; never disguised as a 403/404 that would make the app drop the user's message or treat a live conversation as deleted. The decision runs under one read retry on a fresh connection, so a pooled-connection blip does not surface. |
A not_found is decided from visibility alone, before any session state is consulted, so an
invisible caller gets the same answer whether or not a session is live (no oracle). The
attachment and knowledge routes under a conversation keep their own owner-scoped
"conversation not found" bodies without the code — there a miss can also mean "visible but not
the owner", and the app must not read that as a deletion.
PATCH /admin/v1/conversations/<CID>/messages/<MID> — edit an existing message. The body carries
one of two mutually distinct operations. Both pass the write gate above first; the row update
itself is then owner-scoped in SQL, so a driving non-owner's PATCH matches 0 rows and is a no-op
(200), never a cross-user write. The corrected_content operation is gated on visibility
only (it never reaches model history, so the driver lock does not apply — the owner can fix a
typo while someone else drives); the content operation is driver-locked like POST, because the
app trims and regenerates behind it.
| Field | Type | Effect |
|---|---|---|
content |
string | Rewrites the message body (the user-message edit path). This IS what the model sees as history on later turns. |
corrected_content |
string | null | A user's display-only correction of an assistant answer (fix a typo without re-prompting). |
corrected_content semantics. It is a display OVERLAY, not a rewrite: the app renders / copies /
exports / reads-aloud corrected_content when present, but the message's content (the model's
original output) is never touched — and content, not the correction, is what is ever sent back to
the model as conversation history. A correction can therefore never feed the model's context.
Moderation (report-message) and FULLTEXT search likewise operate on the original content by
design. Accepted: a string that is valid UTF-8 with no NUL bytes and within the shared text size
ceiling (~7.5 MB, the same validate_client_text limit the knowledge upload uses). An
empty/whitespace string or JSON null clears the correction (reverts the display to the
model original). Rejected with 400: a non-string value (number, boolean, object, array),
invalid UTF-8, NUL bytes, or an over-ceiling value. Rejected with 413: a correction that,
together with the message's other large columns (content, and the sources / search-steps
lists a web-search turn stores), would make the row larger than the database can read back in
one packet (~15.9 MB for all of them together — a correction on an answer that is itself near
the commit ceiling); the row is left unchanged and the body carries code: "row_too_large"
(see error codes). The same two rules apply to an edit of content (400
for invalid text, 413 for a row the edit would make unreadable) and the 413 to the legacy
POST …/messages commit. A row the writers admit can always be forked (the fork copies the
large columns server-side). A body carrying both content and
corrected_content is treated as the correction. If neither field is present → 400.
Known limitation: a correction is an in-place update with no new sequence number, so it is not broadcast to a live co-chat viewer in real time — it appears on their next conversation load. A public share link minted after a correction freezes the corrected text in its snapshot.
DELETE /admin/v1/conversations/<CID>/messages/<MID> — delete a message (the app's
edit-and-resend trims the replies after the edited message with it; a regenerate deletes
nothing — the gateway supersedes the prior answer atomically when the replacement commits, see
x-aig-regen-of-turn in the inference reference). Passes the write gate above, then
soft-deletes the row owner-scoped. A message id that does not exist (or is already deleted)
inside a conversation the caller may write to is an idempotent 200 {ok:true} — the app
relies on that after a turn that was never persisted. A conversation that no longer exists is
the typed 404 conversation_not_found, so an edit-resend never streams a "ghost" answer onto a
dead thread; an owner who is not the current driver is refused (403 live_not_driver)
before the row is touched, so a reply is never deleted only for the follow-up turn to be
blocked. A database error on the delete is a 500, never a 200 that would let a second
answer be generated on top of a still-live row.
The message body does not repoint the conversation's gateway
The legacy POST …/messages commit body may carry gateway_id and model recording which
gateway and model produced the reply. model is propagated to the conversation as live state
(so a reload reflects the model that answered last, unless the conversation is on Auto, whose
row model is the create-time resolution and is never overwritten by a per-turn upgrade). The
gateway_id is recorded on the message row only — it does not change the conversation's
gateway. The conversation's gateway determines whether request-phase PII masking and per-gateway
malware scanning apply, so it is a security setting. It is set at creation, and the only
client-driven way to change it is the guarded PATCH /admin/v1/conversations/<ID> (see Route
policy validation), which authorizes the target gateway (unknown /
cross-tenant / non-production-on-move refused); the server additionally keeps it in sync as it
records the gateway that produced each turn (always the caller's own tenant). A gateway_id in the
message-append body — of any gateway, in
any tenant, under any role — is ignored for the conversation's gateway and leaves it unchanged;
an attachment uploaded afterward is still scanned according to the conversation's original gateway.
This route is not a gateway-authorization boundary, so it accepts no gateway repoint at all.
Regenerate-reason provenance (regen_reason)
When a user regenerates an assistant answer with a steering nudge (for example
"too long" or "too formal"), the client sends the chosen preset key and the
gateway stores it on the assistant message as regen_reason, so the app can show
a subtle "Adjusted · …" provenance chip that survives reload.
Accepted values — exactly one of the fixed set:
tooLong, tooShort, tooFormal, tooCasual, offTopic, inaccurate.
Anything else (an unknown key, an oversize string, injection) is rejected at
the trust boundary and stored as NULL (no chip) — the value is never echoed
back or used unvalidated. A plain "try again" regenerate carries no value and
resets the column to NULL (single-writer, last-nudge-wins).
Two ingress paths, one allowlist:
- streaming inference — the
x-aig-regen-reasonrequest header (read per turn, never forwarded to the model provider); - the legacy
POST …/messagesbody — an optionalregen_reasonfield.
The value appears on each assistant message in the conversation-GET response; the field is omitted when unset (a plain turn), so the app treats absent and empty identically. It is enum-like text, capped at 32 characters on the wire as defense-in-depth; the app-side allowlist is the authoritative validator.
Web-search steps (search_steps)
When an answer used web search, the gateway records the per-query search
process — each query plus its result list {title, url} — so the collapsed
"Searched the web" disclosure rehydrates identically on reload. The server-owned
commit stores this itself; on the legacy commit path (an older gateway that
did not acknowledge server-side turn ownership, or the Playground) the client is
the writer, so the optional search_steps field may accompany the assistant
message on POST …/messages.
Accepted shape — an array of steps, each { "query": <string>, "results":
[{ "title": <string>, "url": <string> }] }. Every field is re-validated at the
trust boundary through the same sanitizer applied on read (it is untrusted at
rest as well as on ingress): the list is capped (32 steps, 10 results per step),
query/title are length-clamped, and a url must be http(s):// — a
javascript:/data: or over-long url is dropped, a step whose query is missing
or non-string is dropped, and a wholly unusable list stores NULL. Absent on a
non-search turn → the column stays NULL.
The value appears as search_steps on each assistant message in the
conversation-GET response and is omitted when unset; the raw stored blob is
never echoed back.
Guardrail-block class (guardrail_block_class)
When a request is blocked by a request-phase guardrail (content policy,
mandatory PII protection, an unscannable attachment, or a fail-closed
"guardrail temporarily unavailable"), the gateway persists a distinct
policy-block turn and stamps the block class slug on that assistant message
as guardrail_block_class. This lets a reloaded conversation rehydrate the
distinct, localized policy-block bubble instead of reverting the block to a
plain-looking assistant reply.
Accepted values — a fixed, gateway-authored class slug:
guardrail_blocked, guardrail_pii_mandatory, guardrail_pii_media,
guardrail_unavailable, guardrail_role_model. Only the gateway writes this field (there is no client
ingress path); it never carries user or model text and never the tripped
safety category. It is enum-like text, capped at 32 characters on the wire as
defense-in-depth.
On load the value is untrusted (a DB value could be malformed by a future
bug or manual edit): the app renders it only as a localized i18n key with a
fail-closed fallback, so an unrecognized non-empty slug degrades to a generic
"Request blocked" bubble (never a raw key, never markup, never a crash). The
field is null/omitted on every normal (non-blocked) turn, which the app
renders as an ordinary message.
Hard-failure error class (error_class)
When a turn hard-fails — the request is too large / oversized attachments
(413), the prompt exceeds the model's context window, every configured provider
fails, a request times out, or a provider returns a client error the gateway
could not recover — the gateway now persists a distinct failure turn and
stamps the gateway error code on that assistant message as error_class.
This lets a reloaded conversation show an honest, assistant-visible failure
bubble ("this turn failed: <reason>") instead of a dangling user
message with no response. It is the sibling of guardrail_block_class above.
The failure row is written even when the client disconnects while the turn is
failing (tab close, navigation, network drop): the gateway finishes the
failing turn server-side, so the reloaded conversation is always honest. A
disconnect never fabricates a failure for work that did not fail: a
disconnected turn whose dispatch succeeds simply ends as an empty turn (the
same behavior as pressing Stop before the first token), and only a real
failure persists.
Accepted values — a fixed, gateway-authored, server-sanitized code:
all_providers_failed, request_too_large, request_timeout,
context_overflow, context_length_exceeded, model_capability_mismatch,
provider_error, web_search_not_supported. The gateway is the only writer
(there is no client ingress path). Where the code originates in a provider's own
4xx error body, it is validated against this allowlist at the trust boundary and
any other value — an arbitrary, oversized, or non-[a-z0-9_] string — is
coerced to provider_error before it is stored, so the raw provider text
never lands at rest. The accompanying message content is a gateway-authored
sentence (never the raw provider prose), capped defensively; the field is
enum-like text, capped at 48 characters on the wire.
On load the value is untrusted and rendered only as a localized i18n key
with a fail-closed fallback (an unrecognized non-empty code degrades to a generic
failure bubble — never a raw key, never markup, never a crash). The field is
null/omitted on every normal (successful) turn.
Attachments
POST /admin/v1/conversations/<CID>/attachments — upload a file. The body is JSON.
| Field | Type | Required | Description |
|---|---|---|---|
message_id |
string | yes | The message the attachment belongs to. Must belong to the conversation in the path; a message_id from another conversation returns 404. |
filename |
string | yes | Original file name. |
mime_type |
string | yes | MIME type of the file. |
data |
string | yes | File bytes, base64-encoded (standard alphabet, no line breaks). Validated fail-closed before anything is read or stored: a non-string value (null, object, boolean, number) or an empty string is rejected 400, code: "invalid_body"; a string that does not decode as base64 (an embedded newline is enough) is rejected 400, code: "bad_base64". Nothing that fails to decode is ever persisted — there is no "opaque text" attachment (such a body used to skip the malware scan and be stored verbatim). Both rejections are answered before the conversation lookup. |
Size — oversized uploads are refused with 413, never a 500. Two limits guard this inline-stored endpoint, and whichever trips first answers:
- A request body over the admin host's 16 MB cap (client_max_body_size on the ^/admin/ location) is rejected by the edge as 413 payload_too_large before the handler runs. Because the body is the base64 encoding of the file, this is the dominant answer for any clearly-oversized upload — a raw file of ~12 MB already exceeds 16 MB once base64-encoded.
- A body under that cap whose inline stored row would still exceed the database single-packet limit (16 MB) is rejected by the handler as 413, code: "attachment_too_large" — the same packet ceiling as row_too_large, classified distinctly because it is a stored attachment rather than a message-row write. No partial row is written. This covers the narrow band between the two 16 MB limits (it replaced a raw 500 that leaked the internal packet error).
Support for much larger attachments — chunked/resumable upload, so the 16 MB body cap no longer applies — is tracked internally. (This is a different endpoint from the /chat/files staging upload under File staging, which travels to a 140 MB location and has its own, larger limits.)
GET /admin/v1/attachments/<AID> — download an attachment.
DELETE /admin/v1/attachments/<AID> — delete an attachment.
GET /admin/v1/conversations/<CID>/images/<IID> — fetch the bytes of a generated image in the conversation. Returns { "id", "mime_type", "size", "data" }, where data is the image bytes base64-encoded (the SPA builds a data: URL). Authorization scopes on both the conversation (owned by the caller) and the image id within it; an image id that does not belong to the conversation returns 404.
Generated files (code interpreter)
When the code-interpreter tool produces a downloadable file, the gateway persists it into
the conversation's (or project's) file space and the SPA renders it as a file card. Text
files (csv/txt/json) carry their content inline; a binary document — Excel .xlsx,
Word .docx, PowerPoint .pptx, .pdf, or OpenDocument .odt/.ods/.odp — stores its
bytes and is fetched from a dedicated download route.
GET /admin/v1/conversations/<CID>/knowledge/<KID>
Returns the artifact's stored row as JSON (the metadata the SPA renders as a file card
— for a text file, the inline content). Authorization is scoped to the caller's own
conversation: a <CID> the caller does not own returns 404/403, a storage fault
returns 500, and a <KID> that does not belong to <CID> returns 404.
On a PII-masking gateway, a generated text file (csv/txt/json) has its reversible
masked values restored for the authorized reader on this metadata read — the restore is applied
when the artifact is persisted, so this route returns the stored, already-unmasked text, matching
the same unmask the gateway applies to a chat message it shows that user. Irreversible
custom-keyword redactions ([MYRA-CUSTOM:…]) stay masked, and binary artifacts
(.xlsx/.docx/.pptx/.pdf/.od*) are never rewritten — the /download route below serves
their raw bytes unchanged (and a text-only entry, which has no stored blob, 404s there).
GET /admin/v1/conversations/<CID>/knowledge/<KID>/download
Returns the raw bytes of a generated binary file (an .xlsx/.docx/.pptx/.pdf/
.odt/.ods/.odp produced by the code interpreter) with the gateway-pinned
content_type as the Content-Type, falling back to application/octet-stream when the
entry has none, a sanitized Content-Disposition attachment filename (quotes, semicolons,
CR/LF and path separators are stripped), and X-Content-Type-Options: nosniff (the bytes
are never MIME-sniffed). Authorization is scoped to the caller's own conversation: a
<CID> the caller does not own returns 403/404, and a <KID> that does not belong to
<CID> returns 404 (no cross-conversation or cross-tenant access). The gate is blob
existence — an entry with no stored bytes returns 404; a storage fault returns 500.
What is accepted as a generated file (untrusted runner output, validated fail-closed).
The runner's artifacts[] are hostile until validated; the declared MIME is advisory and
never trusted — the served MIME is pinned by the gateway from the validated kind/extension.
An artifact is dropped (its bytes never surface, counted in artifacts_rejected) unless it
is one of:
- a real PNG (admitted by magic, filename-independent);
- a valid text data file —
csv/txt(valid UTF-8, no NUL) orjson(parses); - an OOXML document —
.xlsx/.docx/.pptx— whose bytes are a ZIP (PK\x03\x04) with a[Content_Types].xmlpart and the extension's REQUIRED subtree (xl/for xlsx,word/for docx,ppt/for pptx), keyed off the sanitized extension; - a PDF (
.pdf) whose bytes begin with%PDF-at offset 0; - an OpenDocument —
.odt/.ods/.odp— whose bytes are a ZIP whose first member ismimetype(name at offset 30) and which carriesMETA-INF/manifest.xml(the manifest excludes an EPUB, which also hasmimetype@30); the gate is compression-independent and the served MIME is keyed off the extension.
The gateway never parses or inflates the bytes (no server-side XXE/zip-bomb; a generated blob
skips async text extraction). A wrong-magic file, a mislabeled container (e.g. a .docx zip
named .xlsx, an OOXML zip named .odt, an EPUB named .odt), a %PDF- not at offset 0, an
oversized file (over the tenant's per-artifact cap — 4 MiB by default, platform-admin-raisable
to 25 MiB via code_interpreter_artifact_max_bytes, see tenants), an
over-aggregate (the default 8 MiB scales with the cap to at most 25 MiB) or an over-count (> 4)
set is rejected.
Re-running the same generating turn creates a numbered copy (report-2.docx) rather than
overwriting, and never fails with a duplicate-key error.
Chat commands
GET /admin/v1/chat-commands
POST /admin/v1/chat-commands
PATCH /admin/v1/chat-commands/<ID>
DELETE /admin/v1/chat-commands/<ID>
A chat command is a saved prompt template that the user can invoke from the Chat view with a slash command. Commands are per-user: every endpoint is scoped to the calling user, so one user can never read, edit, or search another user's commands.
Request body (JSON) for POST / PATCH:
| Field | Type | Required | Description |
|---|---|---|---|
name |
string | yes (POST) | The slash-command name. |
template |
string | yes (POST) | The prompt template; {{variable}} placeholders are substituted at invoke time. |
description |
string | no | Short description shown in the command picker. |
category |
string | null | no | Optional user-chosen category for grouping/filtering. Accepted: a string of at most 64 characters (Unicode codepoints, not bytes — trimmed of surrounding whitespace). An empty/whitespace string or JSON null clears the category ("uncategorized"). Rejected with 400: a non-string value (number, boolean, object, array), a value longer than 64 characters, or malformed UTF-8. On PATCH, omitting category leaves it unchanged; sending null clears it. |
Full-text search — GET /admin/v1/chat-commands?q=<query>
When q is a non-empty string, the endpoint returns only the caller's own commands whose name, description, or template match the query (MariaDB InnoDB FULLTEXT, BOOLEAN mode). The query is escaped at the trust boundary: it is split into word tokens, each ≥ 3 bytes is quoted as a literal phrase (neutralizing every BOOLEAN-mode operator: + - * " ( ) ~ < > @), and at most 16 tokens are used. A non-empty query that tokenizes to nothing (only whitespace or sub-3-character tokens) or matches nothing returns an empty array [] (never an error). A literal empty value (?q=), a repeated q param, or a bare ?q (no value) is treated as absent and returns the full list.
Version history and restore
GET /admin/v1/chat-commands/<ID>/versions
GET /admin/v1/chat-commands/<ID>/versions/<VERSION>
POST /admin/v1/chat-commands/<ID>/versions/<VERSION>/restore
Every create or edit stores an immutable snapshot (name, description, category, template) as a new version, and the command carries a monotonic version counter. The list endpoint returns the caller's own version snapshots newest-first; the single-version endpoint returns one snapshot. These endpoints are owner-only: a command that is not the caller's own reads as 404 (a 404, not a 403, so one user cannot probe another user's command ids), and an empty history likewise returns 404 (every command has at least one version).
The <VERSION> path segment is validated at the trust boundary: it must be a positive, finite integer in [1, 2147483647]. Rejected with 400 invalid version: a non-numeric, fractional, NaN/±inf, zero/negative, or out-of-range value (so a hostile value never reaches SQL). A well-formed version that does not exist for the command returns 404 version not found. A transient database fault returns 503.
Restore is non-destructive: it re-applies the chosen snapshot through the normal update path, minting a new version whose content equals the restored one — no history is overwritten. A null snapshot category is faithfully restored to SQL NULL (not left at the current value).
Sharing with selected users
PUT /admin/v1/chat-commands/<ID>/shares
GET /admin/v1/chat-commands/<ID>/shares
An owner can share a personal prompt with individually selected users in their own tenant. A shared prompt appears in each target's GET /chat-commands (and search) list read-only: the response carries an owned field (1 = the caller's own → editable/restorable/shareable/deletable; 0 = shared with them → read-only). Sharing never grants edit, restore, delete, or re-share to the recipient.
Both endpoints are owner-only. GET .../shares returns the current target users (id, email, name); a command that is not the caller's own returns 404. PUT .../shares replaces the whole share set atomically.
Request body (JSON) for PUT:
| Field | Type | Required | Description |
|---|---|---|---|
user_ids |
array of strings | yes | The users to share with. Rejected with 400: a non-array value, or any non-string element, or more than 200 entries. |
Targets are resolved server-side against the owner's own tenant: an id that is unknown, belongs to another tenant, is soft-deleted, or is the owner themselves is silently dropped (not an error), and duplicates are de-duplicated. An empty user_ids array un-shares the prompt from everyone. A non-owner (or unknown command id) is rejected with 403 forbidden. The client is never the authz boundary — the server derives the owner's tenant and pins every target to it.
Tenant shared commands
GET /admin/v1/me/shared-commands
Returns the slash commands that an administrator has configured for the calling user's tenant (see Shared commands) as a JSON array. These are read-only for the member and are separate from the user's own chat commands above. A caller with no tenant, a tenant with no shared commands, or an empty configuration receives an empty array [].
File staging
POST /admin/v1/chat/files — stage an uploaded file before it is attached to a conversation. The body carries the raw bytes base64-encoded; the response is either the extracted text or an Anthropic Files API file id, depending on the MIME type.
Request body (JSON):
| Field | Type | Required | Description |
|---|---|---|---|
data |
string | yes | File bytes, base64-encoded. |
filename |
string | yes | Original file name. |
mime_type |
string | no | A hint only. The server dispatches on the filename extension (and, for a name with no decisive extension, the file's magic bytes); a generic or wrong client MIME (application/octet-stream for a .xlsx, application/vnd.ms-excel for a .csv, text/plain for a .csv/.tsv — a common Linux browser labeling) is reconciled to the correct path, so a mislabeled CSV still gets the charset ladder and the data-file handling. |
extract_text |
boolean | no | Ask for server-side text extraction of images, PDFs, and plain text. It is a hint, not a control: the server may extract regardless (see the privacy gate below), and can never be used to force a file upload. Images are OCR'd by the on-prem vision model (see Image OCR below), never a raster egress to the customer's provider. |
user_prompt |
string | no | The user's turn text for this send (e.g. "traduci esto"). Used only to make image OCR request-aware — it is appended inside a delimited context block, explicitly marked as the reader's request and not an instruction to the model, so the model extracts what the user needs without ever dropping text. Untrusted: the server codepoint-bounds it to 2000 (an oversized value is truncated; a non-string is ignored), collapses whitespace, and neutralizes the block delimiter. Absent/empty ⇒ plain verbatim OCR. Ignored for non-image uploads. |
gateway_id |
string | yes | The gateway this upload is staged against. It must belong to the caller's tenant — a gateway_id owned by another tenant (or one that does not exist) is rejected with 403/404 before any extraction or the privacy gate runs, so the endpoint cannot be used to probe a foreign gateway's privacy state or existence. The privacy gate below is evaluated from this gateway, and the provider key is resolved from it. |
Response: { "text": ... } on an extraction path (with an optional warning when an image yields little text, and — for a .docx or a text-extracted PDF on a non-privacy gateway — an images array of its embedded images; see below), or { "file_id": ... } on the provider Files API upload path. A caller that omits extract_text for a spreadsheet must handle either shape: the request returns file_id on a normal gateway and text on a privacy-protected one (below). The in-app composer always sends extract_text: true and consumes only the text shape — Anthropic's document content blocks accept only PDF and plaintext files, so a spreadsheet file_id cannot be referenced from a chat message and would fail the turn with 400. The upload path remains available to API callers with their own consumer for the staged file.
When the inline text is a data file (.csv, .tsv, .xlsx, .xlsm, .ods — never .pptx/.docx/plain text), the text leads with an authoritative facts directive — the exact row/column/sheet counts computed by the extractor over the complete file (never estimated from the inline preview), framed as an instruction the model must narrate when asked about the file's structure. The example rows follow, explicitly labelled a partial sample. This lets a weak local, tool-incapable, or private-mode model answer "how many rows?" correctly without eyeballing a truncated preview.
Source-code and plain-text files are accepted by EXTENSION, not by content. Roughly 55 code/config extensions — .py, .pyw, .ts, .tsx, .jsx, .js, .go, .rs, .rb, .php, .c/.h/.cpp, .java, .cs, .sh, .sql, .toml, .ini, .conf, .env, .log, and the like — route to the text/plain inline branch regardless of the client MIME. That branch performs no binary/NUL/UTF-8 check, so a genuine binary given one of these names (or a .txt name) inlines as garbage rather than being rejected — the same long-standing risk profile as .txt; a content-sniff gate that fail-closes on binary bytes is tracked as future hardening. (The .json branch, by contrast, validates UTF-8, and project-knowledge ingestion byte-sniffs its input — those paths differ, so the same file can be accepted here yet rejected as project knowledge.) Files with a dedicated branch — .csv/.tsv/.json/.md/.html/.xml/.yaml/.docx/.xlsx/… — keep their own handling.
A staged .py/.pyw is additionally a runnable input: when code interpreter is enabled, the model may execute the attached script in the sandbox (read /input, write /output), under the same sandbox limits as any code it writes itself.
A Word .docx upload is extracted to structure-preserving GitHub-Flavored Markdown (headings, bullet/numbered lists, tables as pipe tables, bold/italic, hyperlinks, footnotes) rather than flattened to one line per paragraph — so a model reading the file sees which cell belongs to which row and column. The docx bytes are untrusted (accepted shape: a valid Word OPC zip; rejected: anything the extractor cannot read). Extraction runs the bytes through pandoc under a network-disabled sandbox, an address-space (memory) limit, and a wall-clock timeout; an external XML entity is never fetched (no file read / SSRF) and an entity-expansion bomb is capped — a document that trips these fails the rich path. Dangerous link schemes are neutralized: only http, https, and mailto links survive as links (a javascript:/data:/file: target is reduced to plain text), matching the renderer's own allowlist. An enormous document (beyond a size/complexity threshold, or one that exhausts the memory/time limit) gracefully degrades to flat text — complete content, unstructured — instead of failing; the request only errors if the document is unreadable by both paths. (Project-knowledge ingestion of the same .docx uses a separate segmenting extractor for citations, so its stored/preview text may differ from this one-shot chat/workflow read.)
A .docx upload additionally surfaces its embedded images (word/media/*) so a multimodal model can see the letterhead logo, signature, or figures — not only the text. On a non-privacy gateway the response carries an images array; each entry is { "mime", "b64" }, and the in-app composer attaches them to the user turn as ordinary inline image_url blocks (identical to a pasted image — they then inherit request-phase PII blocking, the vision-degrade for a non-vision target, and the per-turn image cap). Two count fields accompany it:
| Field | Type | Meaning |
|---|---|---|
images |
array | The embedded images surfaced, each { "mime": "image/png"\|"image/jpeg", "b64": <base64> }. Always present on a non-privacy gateway (an empty array [] when the doc carries no usable image), and absent on a privacy gateway (below). |
images_total |
number | How many embedded image parts the document carries. |
images_shown |
number | How many are in images (i.e. survived validation and the caps). When images_shown < images_total the composer shows a "N of M images attached" notice. |
The embedded images are untrusted attacker-controlled bytes (the docx is a ZIP), validated fail-closed: the ZIP is read under a bounded, streaming reader with a memory limit and timeout (never a naive unzip — zip-bomb safe); each part's magic bytes are re-sniffed server-side (never a filename or declared type) and only PNG and JPEG are kept — an oversize-dimension image (over the safe-decode pixel ceiling) or a part whose bytes are not really an image is rejected; a vector SVG logo is rasterized to PNG server-side (via resvg, at a bounded width) and then re-sniffed like any other image, but only when it is fully self-contained — an SVG that references an external resource (<image>/<use>/CSS url() pointing at a file path or URL) is dropped, so it can never make the renderer read a host file; vector metafiles (EMF/WMF) and every other raster kind (GIF, TIFF, BMP, WebP) are dropped. The count of surfaced images is capped, and the aggregate is bounded to the inference host's wire budget (oldest dropped first to fit) so the resulting chat turn never exceeds the edge. A dropped image is reflected in images_total > images_shown (the notice), never a failed upload — the text extraction is independent and unaffected. On a privacy (PII-masking) gateway no images are produced at all (the response is text-only, no images field): a raster image cannot be masked, so it must never egress; this is fail-closed (an unreadable gateway config is treated as PII-active). block-level image metadata (the source filename) is used only for de-duplication and chip display — it is stripped before the request reaches the provider.
The facts also carry a column dictionary (schema card): for each column, its 1-indexed position, the header name, and its inferred type (number, boolean, or text); low-cardinality text columns additionally list their distinct values exactly as stored in the file. Type inference matches default pandas.read_csv — so a column of German-decimal amounts (1.234,56) or grouped numbers is reported text, not number, and blank/NA cells do not demote an otherwise-numeric column; number/boolean columns carry no value list. This exists so a model writing code filters on the value as stored (e.g. "Maennlich") rather than the literal from the request (e.g. "Männlich") — a mismatch that silently matches nothing. Column names and values are untrusted file content: they are stripped of control characters and the summary's own structural markers, length-clamped, and the whole card is byte-bounded, so a crafted header or cell cannot inject a second facts line. The card covers .csv/.tsv/.ods and single-sheet .xlsx/.xlsm (multi-sheet workbooks keep counts only); .xlsx stores dates and booleans as numeric serials, so such columns report number.
The response additionally carries four honesty fields describing what a model will actually see:
| Field | Type | Meaning |
|---|---|---|
truncated |
boolean | The extracted text was cut at the inline size cap — only the beginning of the file is in text. For a data file it is set only on the fallback path when no structural facts were available (a data file with facts is sampled instead); a PDF extraction whose inline text hit the 512 KB cap also sets it. |
sampled |
boolean | The file exceeded the sample cap, so text carries the exact facts directive plus a bounded sample of head rows (not the whole file, and not cut mid-table). The counts are exact; answers about specific rows beyond the sample may be estimates. Mutually exclusive with truncated per file. |
masked |
boolean | The gateway is PII-active, so the inline text will additionally pass request-phase masking before a model sees it. |
gateway_compute_path |
boolean | Some tool-capable model on this gateway could compute exactly over the raw file via the code interpreter's staged-input channel (sandbox configured, the gateway has not opted out of the code interpreter — it is on by default, only an explicit code_interpreter.enabled: false disables it — and, on a PII gateway, a fail-closed tool-result mask). This is a gateway-level fact: per-request gates (provider tool support, client-supplied tools, a project's local-only egress policy) can still deny a specific turn, so it is a "possible", never a "will run". |
All four are computed server-side, fail-closed: an unreadable gateway config yields masked: true and gateway_compute_path: false. The facts directive is built only from a real, non-empty extractor summary (an absent or malformed summary is ignored, never surfaced as a bogus count), and rides the inline text only — never the bytes uploaded on the PII-off code-execution path (that would make the CSV reader ingest the directive as a data row). The in-app composer uses these fields for a send-time notice; API callers may ignore them (additive fields, no shape change for non-data files).
PDF extraction — page bound and truncation honesty
A PDF on the extraction path (extract_text, or any upload on a privacy gateway) is read up to a 200-page bound (raised from a silent 20-page cap that made models confidently deny later-page content existed). Pages with a real text layer are extracted deterministically; scanned pages are OCR'd by the on-prem vision model under a wall-clock budget that degrades gracefully — transcribed pages are kept and the cut is reported, never a whole-document failure.
When the extraction is incomplete, that fact is surfaced twice — to the model and to the caller:
- The inline
textends with a model-visible marker (e.g.[... note: only the first 20 of 43 pages were extracted; content on the remaining pages is missing ...], or a "some pages were too dense to fully transcribe" variant), appended after the 512 KB inline cap so it is never cut off. This is what stops a model from asserting that dropped pages do not exist. - The response carries additive fields (absent on a complete extraction — same convention as
images_shown/images_total):
| Field | Type | Meaning |
|---|---|---|
pdf_pages_shown |
number | Pages actually extracted. Sent only when pages were dropped (pdf_pages_shown < pdf_pages_total). |
pdf_pages_total |
number | The document's true page count, recorded before the page bound. |
pdf_transcription_truncated |
boolean | Some transcribed pages were cut mid-transcription (a page too dense for the OCR output budget, or the OCR wall clock). |
The bound is exact: a document of exactly 200 pages is complete — the page fields and the page-drop marker appear only from 201 pages on, because pdf_pages_shown is sent only when it is strictly below pdf_pages_total (never as "200 of 200"); the "too dense to fully transcribe" variant is independent of the page count. The shared truncated field (above) additionally reports the 512 KB inline byte cap on this path. The in-app composer localizes these into a send-time notice ("Only the first 200 of 201 pages of
The extractor subprocess's control lines (usage/billing accounting interleaved with the extracted text) are authenticated with a per-invocation random nonce, so a crafted document whose text contains a forged accounting line cannot alter billing or forge/suppress the truncation report — the forged line is treated as ordinary document content. The subprocess also runs under the same memory ceiling and wall-clock timeout as the Office-document extractors (below); a decompression-bomb PDF is rejected 422, never a memory exhaustion.
A text-extracted PDF additionally surfaces its embedded images — the figures, charts, diagrams, scanned signatures or letterhead logo the text/OCR extraction discards — so a multimodal model can see them, not only the transcription. This is the PDF mirror of the .docx embedded-image leg (above) and uses the same response fields, caps, validation and privacy gate:
| Field | Type | Meaning |
|---|---|---|
images |
array | The embedded images surfaced, each { "mime": "image/png"\|"image/jpeg", "b64": <base64> }. Present (an empty array [] when the PDF carries no usable image) on a non-privacy gateway whenever the PDF took the text-extraction path; absent on a privacy gateway, and on the native-inline path below. |
images_total |
number | How many real embedded figures the document carries. |
images_shown |
number | How many are in images (survived validation and the caps). When images_shown < images_total the composer shows an "N of M images attached" notice. |
The images are untrusted attacker-controlled bytes, validated fail-closed exactly like the docx leg. The PDF is parsed (PyMuPDF) under the same memory ceiling and wall-clock timeout; its embedded image XObjects are enumerated and each image's encoded stream is read without rendering the page. Each kept image's magic bytes are re-sniffed server-side (never the PDF's declared encoding) and only PNG and JPEG are kept — any other encoding (JPEG 2000, JBIG2, CCITT fax, …), an over-the-pixel-ceiling raster, an oversize image, or a part whose bytes are not really an image is dropped; a dropped figure is reflected in images_total > images_shown (the notice), never a failed upload. Degenerate slivers (sub-16 px rules, 1 px spacers, mask fragments) and bomb-dimension rasters are skipped entirely and are not counted (they are not figures, so counting them would make the notice lie). An image reused on several pages is counted once. The surfaced count is capped and the aggregate is bounded to the inference host's wire budget (oldest dropped first to fit). A password-protected (encrypted) PDF yields no images. On a privacy (PII-masking) gateway no images are produced (text-only — a raster image cannot be masked, so it must never egress; fail-closed, an unreadable gateway config is treated as PII-active). The image leg is independent: if it degrades, the text extraction is unaffected (images are additive). It applies only to the text-extraction path — a document-capable model on a non-privacy gateway receives the PDF as a native inline document block and already sees the pages (figures and all), so no separate image leg runs there.
Spreadsheets are a code-execution data path
Analysing a spreadsheet ("how many rows…", "total revenue…") is a computation, not a text-extraction task. In chat, exact computation is provided by the gateway code interpreter (when enabled on the gateway): the staged data file is placed into the sandbox and the model runs read_excel()/read_csv() over the raw bytes for exact results. Without a code-execution path, the model answers from the bounded inline preview, whose leading facts directive carries the exact structure counts (below). The provider Files API upload (file_id shape, PII-off gateways only) stages the raw file — the .xlsx workbook, or a UTF-8-transcoded .csv/.ods — for callers with their own consumer; note that Anthropic chat document blocks accept only PDF and plaintext files, so a spreadsheet file_id is not usable there. A .pptx is not a data file and is always returned as extracted text, never uploaded.
Character sets are normalized to UTF-8 at the trust boundary via a fail-safe ladder (UTF-16 BOM → UTF-8 → Windows-1252 → Latin-1), so a German-Excel CSV export (CP1252/Latin-1) is read correctly instead of failing to decode or arriving as mojibake. Inline extracted text is capped, with an explicit truncation notice, so a very large file cannot overflow the model's context window.
When a data file is returned as inline text, the extracted text is returned as-is. (A computed [File structure] row/column/sheet-count prefix was previously prepended to help a model that cannot run code answer "how many rows…" without eyeball-counting; it is currently disabled — models are instead steered to run the data through the code interpreter and compute exact figures over the raw file, which supersedes the text summary. The extractor still computes the summary and the prepend code is retained (commented out), so it can be re-enabled if needed.)
Privacy gate (fail-closed)
The caller's ownership of gateway_id is verified first, before the privacy gate or any extraction, so the privacy state of a gateway is only ever evaluated for a gateway the caller's tenant owns.
When the target gateway has an active PII guardrail (a pii_protector, custom_pii, or Presidio detector whose action scrubs or blocks), attachment content is never uploaded to a provider Files API — it would bypass the request-phase masking that runs on the inference body. Instead the file is extracted to text server-side and returned inline, so it passes through masking before it can leave the gateway. This is enforced on the server, fail-closed: the decision is made from the gateway's guardrail config (not from any client field), an indeterminate config read is treated as PII-active, and a request that reaches the upload step on a PII-active gateway is rejected with 422. A client cannot cause a raw upload on a privacy-protected gateway by omitting or altering extract_text.
Scope: this gate covers the spreadsheet / Files-API path. Native inline image and PDF blocks (sent in the inference body, where guardrails mask them) are a separate path.
Office-document extraction (.docx, .xlsx/.xlsm, .pptx) — and, PDF extraction too — runs in a subprocess bounded by a memory ceiling and a wall-clock timeout. A crafted file that inflates far beyond its on-disk size (for example a small archive whose many internal parts decompress to gigabytes of text, or a PDF whose compressed text layer does the same) is rejected with 422 when the subprocess hits the ceiling — it can no longer exhaust host memory. Legitimate documents are unaffected.
Image OCR
An uploaded image is turned into text by the gateway's on-prem vision model (an OCR/verbatim-transcription prompt — transcribe all text, never a UI/design description). The image bytes never leave the gateway trust domain for this: on a privacy (PII-masking) gateway the image cannot be sent to the model as pixels (a raster cannot be masked), so this on-prem OCR is the only way its content reaches a model — the transcription then passes request-phase masking before any customer-provider egress. Any Pillow-decodable format is normalized to PNG first, so JPEG/PNG/GIF/WebP/TIFF/BMP are all supported.
The OCR is request-aware: when the composer sends user_prompt (the user's turn), it is appended to the transcription prompt inside a delimited block explicitly marked as the reader's request (not a model instruction, with the delimiter neutralized and whitespace collapsed) so the model surfaces what the user needs (e.g. an emailed screenshot sent with "translate this" is transcribed so it can be translated). The turn is the user's own text and reaches the on-prem model before request-phase masking; this is an accepted, documented pre-mask exposure confined to on-prem infrastructure (the image bytes already go there), and the resulting text is still masked before it can leave the gateway.
Model transcription is untrusted output: a repetition-loop / degenerate transcription (the model looping on one token) is detected and retried once; if it persists, the response is the no-text placeholder (warning set) — garbage is never stored or shown. A genuinely text-poor image yields the same placeholder; a short but real transcription is returned with a "limited text" warning.
Undecodable / unsupported image. An image is accepted by MIME prefix (image/*), but the server can only decode the formats its image library supports: JPEG, PNG, GIF, WebP, TIFF, BMP, HEIC/HEIF, AVIF, and DNG (Apple ProRAW and other camera raws — the server extracts the raw's embedded full-resolution JPEG preview; a raw with no embedded preview cannot be decoded). A genuinely unsupported format, or a truncated / corrupt / zero-byte file, is rejected at the server (in a local decode preflight, before any model call) with 422 and a stable code: "unsupported_image_format" plus a short, safe message — never the raw decoder exception (which previously leaked a Python-internal string, including a memory address, to the user). The in-app composer maps this code to a localized "unsupported or damaged image" message; API callers can branch on the code. A transient extraction failure (a model/network blip) keeps the generic 422 without that code.
Document extraction failure (PDF/DOCX/DOC/ODS/XLSX/PPTX/CSV/image-normalize). When the extractor subprocess for one of these formats fails (a corrupt/unreadable file, an unsupported internal structure, the subprocess crashing or timing out), the response is 422 with a stable code: "extract_failed" and a short, filename-agnostic message naming only the failed format (e.g. "Could not extract text from this .docx file.") — never the extractor's raw diagnostic (a python traceback line, an OS exit code, a subprocess spawn-failure reason). That raw detail is written to the operator-facing server log only; a leak was fixed where it was previously concatenated into this same client-facing error field. The in-app composer maps extract_failed to a localized "couldn't be read, try attaching again or continue without it" message; the failed attachment is restored to the composer so the user can retry or remove it.
Error codes: 400 (missing or invalid data/filename, or missing gateway_id), 403/404 (gateway_id is not owned by the caller's tenant, or does not exist — returned before any extraction), 422 (extraction failed — code: "extract_failed" for a PDF/DOCX/DOC/ODS/XLSX/PPTX/CSV — an undecodable or unsupported image — code: "unsupported_image_format" — a document that exceeds the extraction memory/time limit, an attempted upload on a privacy-protected gateway, or — on a malware-scanning gateway — an infected file, code: "malware_detected"), 502 (provider Files API error), 503 (provider key not configured for the gateway, a transient gateway lookup failure, the shared upload pool saturated — code: "ingest_busy", with Retry-After — or, on a malware-scanning gateway, the scanner unreachable, code: "scan_unavailable", with Retry-After).
Client-side upload size limit
The in-app composer enforces a client-side per-file guard sized to this host's edge upload cap (140 MB, giving a raw-file limit of about 110 MB) before it POSTs to this endpoint (raised from 56 MB so a ~100 MB Office document — the >100-page reference case — uploads). The composer's own attachment-staging cap is derived from that same budget (no second number). A file over the guard is rejected in the browser with a specific "…is too large to upload…" message and the request is never sent.
The client guard is not the trust boundary. The server independently rejects a decoded upload over a 100 MiB ceiling with a 413 JSON error (a clear message, never a bare 413) before any extraction or provider upload, and admits a bounded number of concurrent uploads — an upload arriving while the file-processing pool is saturated gets a 503 with code: "ingest_busy" and a Retry-After header (a transient state — retry shortly) rather than exhausting worker memory (each large upload peaks at the base64 body plus the decoded bytes plus its extractor subprocess). This is the same shared upload-pool budget, and the same machine-readable code, as the project knowledge-upload endpoint. The web chat branches on the exact string code: "ingest_busy" (a 503 without exactly this code falls to the generic extraction-failure message): it shows a transient "busy — please try sending again" banner naming the file and keeps the attachment in the composer; it does not retry on its own. In a multi-attachment send the other files' finished extractions are kept with their composer chips (the failed file's chip is marked), so the next Send re-submits only the failed file — never the whole batch again into the pool that just refused one.
Malware scanning (fail-closed, opt-in per gateway). Malware scanning is enabled per gateway (av_scan.enabled in the gateway config) and is off by default — a gateway that has not opted in skips this step and the upload proceeds normally. When the target gateway has opted in: after the size/authorization checks and before any text extraction or provider Files-API upload, the decoded bytes are streamed to a self-hosted ClamAV daemon (clamd, INSTREAM). An infected verdict rejects the upload with 422 and code: "malware_detected" ("This file was blocked by malware scanning."); the ClamAV signature is logged operator-side only and never returned. If clamd is unreachable, times out, or returns an error, the upload is rejected — not accepted unscanned — with 503, code: "scan_unavailable" and a short Retry-After ("Malware scanning is temporarily unavailable. Please try again in a few moments."). The same opt-in scan, with the same codes, guards the conversation attachment endpoint (POST /conversations/{id}/attachments, gated on the conversation's gateway) — attachments are persisted and served back raw. The virus scanner is configured for your deployment by Myra and is consulted only for an opted-in gateway.
Note this is larger than the inline image/PDF limit on the inference host (~7.4 MB): the two travel to different hosts with different edge caps (admin /chat/files at 140 MB; inference compat/chat/completions at 11 MB), so a document uploaded here for text extraction is bounded independently from a native-inline attachment. See the inference reference for the inline guard. (A document beyond the 200-page PDF extraction bound — or one whose text exceeds the 512 KB inline context ceiling — is best added as project knowledge, where it is chunk-stored and retrieved by RAG, rather than inlined into a single chat turn; an inline extraction past either bound says so explicitly, both to the model and in the response fields.)
This is a UX guard, not part of the endpoint contract — the endpoint itself has no application-layer per-file size cap and returns no 413. The guard exists because a body over the host's cap is dropped with a 413; behind the CDN that 413 historically carried no CORS headers, so the browser blinded fetch to an opaque network error indistinguishable from a genuine outage. Rejecting the oversized file up front turns that "the gateway is unreachable, keep retrying" dead end into a clear, actionable message. Direct API callers are not bound by this browser guard, but a body above the edge cap will still be dropped at the edge.
Conversation export
Render markdown content — one assistant message or a whole conversation — into a downloadable file. Nine formats share one request and response contract.
| Endpoint | Output |
|---|---|
POST /admin/v1/chat/export-pdf |
|
POST /admin/v1/chat/export-docx |
Word document |
POST /admin/v1/chat/export-odt |
OpenDocument Text |
POST /admin/v1/chat/export-txt |
Plain text — the message rendered without markup |
POST /admin/v1/chat/export-xlsx |
Excel — one sheet per markdown table |
POST /admin/v1/chat/export-ods |
OpenDocument Spreadsheet |
POST /admin/v1/chat/export-csv |
CSV — every markdown table, rows written consecutively (same table source as Excel) |
POST /admin/v1/chat/export-pptx |
PowerPoint |
POST /admin/v1/chat/export-odp |
OpenDocument Presentation |
The CSV and spreadsheet (xlsx/ods) exporters share one table-extraction engine, so a CSV agrees cell-for-cell with the Excel file for the same content: every GFM pipe table in the markdown is exported, numbers are detected locale-free (both 1,234.56 and 1.234,56 become the value 1234.56), and identifier-like values (leading zeros, >15 digits) stay text. CSV is a flat format with no sheets or table-boundary concept, so when a message has more than one table their rows are written consecutively (each table keeps its own header row); no blank-line separator is inserted (a data format should not carry a non-data structural line). A CSV is UTF-8 with a byte-order mark (so umlauts render correctly in Excel on Windows). To defend against CSV formula injection, any text cell (including a header cell) that begins with =, +, - or @ is prefixed with an apostrophe so a spreadsheet cannot execute it as a formula or DDE command; a genuine negative number such as -5 is written as a number and is not prefixed.
The plain-text (txt) exporter renders the message with all markup removed (headings, emphasis and links flattened; tables laid out as aligned text). Unlike the spreadsheet formats it never returns 422, because any content — prose or tables — has a plain-text form.
Request body (JSON):
| Field | Type | Required | Description |
|---|---|---|---|
markdown |
string | yes | Markdown source to render. Maximum 8 MiB. |
filename |
string | no | Download base name. The format extension is appended. Defaults to conversation. |
reference_attachment_id |
string | no | export-docx only. The id of a .docx the caller previously uploaded in one of their own conversations, used as the Word letterhead template: the produced document carries that file's header/footer, embedded logo, brand colours and fonts while the body is the rendered markdown. The attachment must belong to a conversation the caller owns (foreign or unknown ids are not distinguished — both are treated as "no reference"); it must be a .docx under 25 MiB. It is sanitized at the trust boundary before use — external-target relationships (which would make the produced file fetch a remote URL when opened), OLE objects, macros/VBA and external-data (mail-merge) settings are stripped; the embedded logo (an internal relationship) is preserved. Ignored for every non-docx export. |
Response: the rendered file is streamed back (HTTP 200, Content-Type per the table, Content-Disposition: attachment). The body is the raw file bytes, not JSON. When reference_attachment_id was supplied, the response carries an X-Aig-Reference-Applied header (true/false, exposed cross-origin) so the client can tell the user whether the letterhead was actually applied: if the reference is missing, not owned, not a valid .docx, or fails to render, the document is produced without the letterhead (never a 500, never silently branded) and the header is false.
Error codes: 400 (markdown missing or whitespace-only), 413 (markdown exceeds 8 MiB), 422 (the CSV, spreadsheet or presentation renderer found no usable content — e.g. a CSV or Excel export of a message with no table), 500 (renderer failed), 503 (at capacity, maximum four concurrent renders).
Images and untrusted media
The markdown is untrusted input handed to the renderers (pandoc / WeasyPrint), which would otherwise fetch any URL or read any local file a reference points at. It is sanitized at the trust boundary before rendering, fail-closed:
- Kept: a self-contained inline image whose source is a
data:image/pngordata:image/jpegURI. The payload is base64-decoded and validated by its real magic bytes (the declared MIME is not trusted), and rejected if its pixel dimensions exceed 25 megapixels. This is how embedded and server-generated images survive the export. - Rejected (the image is dropped, its alt text kept): any remote (
http(s)://, protocol-relative//host), local or relative path,file:URL,data:image/svg+xmlor any non-imagedata:URI, reference-style/shortcut image, and raw HTML (<img>,<style>,<div style="…">, etc.). Raw HTML in the source is not rendered at all — the pandoc reader is restricted to a safe set of markdown extensions, so no HTML/CSS/metadata construct can trigger a network fetch.
No image reference in the exported document can cause the server to make an outbound request or read a local file.
Corporate-design branding (PDF, Word, PowerPoint)
The PDF, Word (docx) and PowerPoint (pptx) exports apply the caller's organization (tenant) corporate design when it is configured: brand_primary_color recolours headings (and, in the PDF, the top-heading rule and links); brand_font_family sets the document font; and the organization logo — the same image uploaded for display in the app — is placed on the generated document. Branding is resolved server-side from the authenticated caller's tenant — it is never sent in the export request. When nothing is configured the platform default is used, and each element degrades independently.
How each element is applied per format:
Word (docx) |
PowerPoint (pptx) |
||
|---|---|---|---|
| Color | headings, heading rule, links | headings + title (document theme) | slide titles and section headings |
| Font | body font | document body + headings (theme) | all text (code stays monospaced) |
| Logo | top of the first page | top of the first page | top-right of every slide |
The logo is the same asset uploaded for the in-app logo (one upload, reused here — there is no separate export-logo upload). It is size-normalized to a small PNG before embedding; a logo with transparency is placed on a white background. The other export formats (odt, ods, odp, xlsx, csv, txt) use the platform default in this release.
The color/font are set on the tenant (see the tenant admin API) and are treated as untrusted at every boundary — validated on write, and re-validated where the stylesheet / Word theme / slide is generated (defense in depth) — so a stored value can never inject into the generated document. Accepted shapes:
brand_primary_color— exactly a#RRGGBBhex color (case-insensitive, stored lower-cased). Anything else (named colors,#RGB, values containing;/}/quotes) is rejected. The exports always use the stored value exactly; only the app's dark theme renders an automatically lightened variant of it (to stay WCAG-AA readable on dark surfaces) — the stored value itself never changes.brand_font_family— 1–120 characters of letters, digits, spaces and hyphens only. Quotes, semicolons, braces and angle brackets are rejected so the value cannot break out of the generated CSS/theme. (The font must also be installed on the server for the PDF to render in it; an unavailable font falls back to the default stack — the Word/PowerPoint files name the font regardless, and the viewer substitutes if absent.)
Every branding element is fail-closed: an invalid stored color/font is dropped per-field, and a missing or corrupt logo is omitted — the export still renders, unbranded for that element, never erroring.
An invalid value is a 400 on the tenant PATCH; a value that is somehow invalid at render time is dropped per-field (fail-closed to the default) rather than emitted.
Audio transcription
POST /admin/v1/chat/transcribe — transcribe a recorded audio message to text. The web app sends the browser's native recording as-is (WebM/Opus on Chrome/Firefox/Android, MP4/AAC on Safari); the gateway re-encodes it to Opus server-side and retries on an empty transcript. It does not decode or re-encode client-side — sending the compressed recording (rather than a WAV) is what keeps a multi-minute recording under the per-request cap. The common browser codecs are handled by the image's stable ffmpeg; a rarer non-browser codec that only a native upload could carry — notably xHE-AAC / USAC — is decoded best-effort by a modern fallback ffmpeg the server retries with when the primary decode produces nothing, so such an upload transcribes instead of failing closed. Genuinely corrupt or truly unsupported audio still fails both and returns 422.
Request body (JSON):
| Field | Type | Required | Description |
|---|---|---|---|
audio |
string | yes | Audio bytes, base64-encoded. Maximum 10 MiB of decoded bytes per request — for a compressed recording (~64 kbps) that is roughly 20 minutes of speech. Any browser-recorded container is accepted (the bytes are re-encoded server-side, so mime_type is advisory). |
mime_type |
string | no | Advisory container hint (e.g. audio/webm). Ignored for routing — the server sniffs the bytes. |
language |
string | no | ISO 639-1 language hint. Whisper auto-detects when omitted. |
Response: { "text": "<transcript>" }. An empty string is a valid result for silence.
Input validation. audio is required and must be a base64 string — any non-string value (a number, boolean, JSON null, or object) is rejected with a 400. mime_type and language are advisory and untrusted: a non-string value is silently ignored (treated as absent), never a 400 — mime_type only labels a debug capture, and a non-two-letter or non-string language falls back to Whisper's auto-detect.
Error codes:
| Status | Meaning | Retry? |
|---|---|---|
200 |
Success — { "text": "<transcript>" }. An empty string is a valid result for genuine silence. |
— |
400 |
audio missing, not a base64 string, or not valid base64 (the client-side envelope check). |
Fix the request |
413 |
audio exceeds the 10 MiB decoded-bytes per-request limit — the web app caps recording length, so this is only reached by an unusually long recording. |
Record shorter |
422 |
The recording could not be turned into text — either a local decode failure, or the local minimum source-audio-duration floor (a recording whose ffprobe container/stream duration is below ~0.2 s is treated as an accidental tap, not speech, and short-circuited to the same "no speech detected" response before any transcode or backend call — a recording with no reported duration is never rejected on this basis), or the speech-to-text backend returning a 400 that carries Whisper's audio-reject signature ("Invalid or unsupported audio file"), which means the audio itself was unusable — overwhelmingly a too-short / empty recording. The web app surfaces a friendly "no speech detected — try again" and the user re-records. |
Re-record |
429 |
Genuine saturation ("busy"). Either the gateway's local admission cap of two concurrent transcriptions, or the speech-to-text backend signalling a 429 rate-limit. Carries Retry-After. |
Yes, shortly |
502 |
A permanent upstream failure the client cannot recover by retrying — a contract break (non-JSON / missing text), a non-audio non-400 upstream 4xx such as an auth error (401/403) or a not-found (404), or a misconfiguration 400 whose body does not carry the audio-reject signature (an unrouted model, a bad language code). A misconfig 400 is a hard 502 (not the friendly 422) and is additionally logged at ERR for operators, since a systemic misconfig fails 100% of transcriptions and shows up as an operator-visible ERR flood, not a per-user event. Every such failure is logged server-side. |
No |
503 |
The upstream is down or broken ("temporarily unavailable") — a transport failure (the backend is unreachable; a 5 s connect deadline turns a black-holed host into a fast failure) or an upstream 5xx. Carries Retry-After. This is deliberately distinct from 429: an overload-503 and a hard-down-503 are indistinguishable, so a 5xx is never reported as "at capacity" / "busy". |
Yes, shortly |
A transient failure is reported as 429 (busy) or 503 (unavailable), never 502, so a client can tell "busy/unavailable, try again shortly" from "broken". The gateway does not auto-retry a 429/503 — the retry loop only re-attempts an empty transcript (its known intermittent-empty result), because retrying while holding a scarce transcription slot would amplify load on a saturated backend.
Text-to-speech (read aloud)
POST /admin/v1/chat/speak — synthesize text to speech ("read aloud"). The synthesis runs off the streaming critical path; it is a separate request the client makes after the answer has finished streaming.
Speech is produced by a self-hosted EU-hosted Piper model (piper-de, voice de_DE-thorsten-medium) on the sovereign fleet — the request never leaves the EU boundary, matching the residency guarantee of every other model call. The voice, model, and audio format are fixed server-side; the client supplies only the text.
One request carries one segment, not a whole answer. Synthesis costs roughly 3.6 seconds per 1000 characters, so a long answer sent as a single request takes almost a minute before the first byte of audio exists. The web app therefore splits the answer at sentence boundaries and requests the segments one at a time — playing segment 1 while segment 2 is still being synthesized — so audio starts after about 1.5 seconds regardless of how long the answer is. In practice each request carries roughly 400 characters (the first) or 1200 characters (the rest). Integrators are strongly encouraged to do the same: a single 16 KiB request is accepted but will keep you waiting for the whole synthesis, and it occupies the shared engine for the entire time.
Request body (JSON):
| Field | Type | Required | Description |
|---|---|---|---|
text |
string | yes | The text to read aloud. Must be valid UTF-8 and at most 16384 bytes. |
language |
string | no | ISO-639-1 hint for which voice to synthesize with: de, en, fr, or nl. Any other value — an unsupported code, a non-string, or the field omitted — falls back to German (de). The web app detects the answer's language once over the whole answer and sends the same value for every segment. |
message_id |
string | no | The id of the committed assistant message this text was taken from. When present and the message belongs to one of your own conversations, the gateway serves the pronunciation-normalization step from a precomputed cache (see below) and populates it on a miss. A missing, unknown, or another user's id is silently ignored — the request simply takes the live normalization path — so it is never an authorization surface. |
Response: the raw audio bytes with Content-Type: audio/mpeg (MP3), for inline playback.
Error codes:
| Status | Meaning | Retry? |
|---|---|---|
400 |
text is missing, empty, or not valid UTF-8 |
no — fix the request |
401 / 403 |
the session has expired or lacks access | no — sign in again |
413 |
text exceeds 16384 bytes |
no — split the text and send it as segments |
429 |
per-user rate limit | yes, after 60 seconds |
502 |
synthesis genuinely unavailable (the upstream is down or unreachable) | no |
503 |
at capacity — the gateway admits two concurrent syntheses, and the shared engine synthesizes one clip at a time | yes, after a few seconds |
429 and 503 both carry a Retry-After header, and both mean "come back shortly". 502 means the service is not working and retrying will not help. Capacity pressure anywhere in the chain — this gateway, the fleet's admission limits, the engine's own queue, or a synthesis that runs past the gateway's deadline — is reported as 503, never as 502, precisely so a client can tell a temporary condition from a broken one.
Two practical notes for clients:
Retry-Afteris not a CORS-safelisted response header, so browser JavaScript cannot read it. Browser clients should key their retry policy on the status code alone, and should treat429as "wait about a minute" rather than retrying immediately.- The rate limit is charged by size, not by request count. The budget is 15 segment-equivalents per 60 seconds per user, where one segment-equivalent is 1200 codepoints (characters, counted independently of UTF-8 byte length); a request is charged
ceil(codepoints / 1200). Sending segments of the recommended size therefore costs one unit each, while a single maximum-size request costs 14. This exists because the shared engine synthesizes one clip at a time and the cost of a request varies roughly 40x across the accepted range.
The gateway also scales its own upstream deadline to the size of the request, so a small segment fails fast when the engine is wedged while a maximum-size request still gets the time it legitimately needs.
Voices and language routing: the deployed engine serves five voices — German (
de_DE-thorsten-medium), English (en_US-lessac-medium), French (fr_FR-siwis-medium), Dutch (nl_NL-mls-medium) and Spanish (es_ES-davefx-medium). The read-aloud request'slanguagehint selects the matching voice; German is the default when the hint is absent or unsupported. A single answer is synthesized in one language (the detected/hinted one) — a mixed-language answer is read in its dominant language.Pronunciation normalization: before synthesis, the
textis rewritten into fully spoken form so the phonemizer reads it correctly. First, a small deterministic pass expands the few language-universal approximation symbols that must never be left to a model — a~or≈immediately before a number (at a word boundary) becomes "circa" (e.g.~64→ "circa 64"); a tilde glued to a token like a git ref (HEAD~3) or a path (~/dir) is left untouched. Then the primary layer, an internal EU language model, expands numbers, dates, times, currency amounts, numeric ranges, percentages, mathematical symbols, acronyms, and abbreviations into spoken words in the answer's own language — this is language-agnostic, with no per-language rule set (for example500–1000→ "five hundred to one thousand"/"fünfhundert bis eintausend",31.07.2026→ the spoken ordinal date). It is a pure rewrite that never changes the meaning of the answer, and it is fenced so that text inside the answer is treated strictly as content to rewrite, never as instructions. The model layer is fail-open: if it is unavailable, times out, or returns a truncated reply, the request falls back to the synthesis engine's own deterministicnum2wordsnormalizer (the deterministic symbol pass having already run), so read-aloud never breaks — it only reads a little less fluently. No layer ever routes text outside the EU boundary.Normalization cache: the model normalization above is a pure rewrite — for a fixed
(language, exact text)it always produces the same spoken form. When a request carries amessage_idyou own, the gateway caches that(text → spoken)result and reuses it on any later read of the same text, so a replay skips the language-model call entirely (its ~1–3 s of latency and its cost). The cache is keyed by a hash of the exact text sent — a single changed byte is a different key, so it can only ever return the normalization of precisely what was asked for, never stale or mismatched audio. It is a pure optimization: it does not change what is synthesized, only how fast the normalization step returns, and it never affects the fleet-serialized synthesis itself. A miss (a first read, an edited answer, or any request without a usablemessage_id) transparently takes the live path. The cache is scoped to your own messages and is deleted together with the message under retention and account-deletion.
Memories
GET /admin/v1/memories — list stored memories for the calling user. The response is bounded: it returns at most 50 memories per scope (the per-scope cap, see below), in the stored order (sort_order, then created_at, then id). A user who somehow holds more than the cap (rows created before the cap existed) sees the first 50 in that order.
POST /admin/v1/memories — create a new memory entry. Each scope holds at most 50 memories; a create that would exceed the cap is rejected with 409 and a machine-readable body { "error": "memory limit reached for this scope", "code": "memory_quota_exceeded" } (the count is not silently clamped — the create simply does not happen). Delete a memory to make room.
PATCH /admin/v1/memories/<ID> — update a memory entry.
DELETE /admin/v1/memories/<ID> — delete a memory entry.
POST /admin/v1/memories/reorder — set the ranking of memories within one scope (drag-and-drop). Body: { "ids": ["<memory-id>", ...], "project_id": "<id>" }. ids is the desired top-to-bottom order and must be a non-empty JSON array of strings — a missing ids, a non-array ids, an empty array, or any non-string / empty-string element is rejected with 400. project_id is optional: omit it for the caller's user-scope pool, or supply it for a project scope (a non-member, non-admin caller gets 403). Each id is applied only to a row the caller owns in that scope; ids the caller does not own are silently ignored (no cross-user or cross-scope write). Returns the re-listed memories in their new order. The ranking is applied atomically — a failure mid-list rolls back so the order never lands half-applied.
A memory is a piece of long-lived context that the chat engine may inject into future requests. Memories are injected into the model's system prompt in this stored order, so a higher-ranked memory takes precedence; new memories are appended at the bottom of their scope.
Per-scope cap. Each scope — the caller's user scope, and each project scope independently — holds at most 50 memories. This is a security ceiling, not only a tidiness limit: every stored memory is encrypted at rest, and list_memories runs on every chat turn and decrypts each returned row (one key derivation per row on a cold cache), so an unbounded pool would make each turn arbitrarily expensive. The cap bounds that per-turn cost. It applies to every create path — the POST /memories above and the accepted/edited items of the suggestion-feedback batch below both reject with the same 409 memory_quota_exceeded when the target scope is full.
Context-layer suggestion feedback
POST /admin/v1/memories/feedback — record the user's decisions on the model-proposed <memory> suggestions and, for the accepted/edited ones, create the corresponding memory. The whole batch runs in a single transaction: either every decision and every created memory is persisted, or nothing is (a mid-batch failure rolls back, so a client retry never duplicates a memory or double-counts a metric).
Request body: { "items": [ { "decision": "accepted"|"rejected"|"edited", "proposed_content": "<the model's original suggestion>", "final_content": "<edited text>", "type": "fact"|"preference"|"instruction", "project_id": "<id>" } ] }.
itemsmust be a non-empty JSON array of at most 100 objects — a missing, non-array, empty, or over-capitems, or a non-object element, is rejected with400.decisionis required and must be one ofaccepted,rejected,edited— any other value is rejected with400.proposed_contentis required, must be a non-empty string, and is capped at 4000 characters (over-length →400).final_contentis required for anediteddecision (non-empty, same 4000-char cap; empty/missing →400). It is ignored foraccepted(the proposal is saved verbatim) and forrejected(which creates no memory).typeis best-effort metadata: an unknown or missing value is clamped tofactrather than rejected, so a stray model-emitted type never blocks the batch.- Scope is per item: each item's optional
project_idselects its scope (absent/null = the caller's user scope). Every distinctproject_idin the batch is checked for existence (404if it does not exist) and membership (403if the caller is neither a member nor a platform admin) before anything is written.tenant_idanduser_idare taken from the session, never from the body.
On success returns 200 with { "created": [ <memory>, ... ] } — the memories created for the accepted/edited items (an all-rejected batch returns { "created": [] }). Rejected suggestions are also fed back into the memory system: the model stops re-proposing a fact the user explicitly rejected (bounded to the user's 200 most recent rejections per scope).
GET /admin/v1/memories/feedback/metrics — the S6.2 quality metrics for the caller's tenant, tenant-isolated: { "accepted": N, "rejected": N, "edited": N, "total": N, "accept_rate": 0..1, "reject_rate": 0..1, "edit_rate": 0..1 }. The counts never include another tenant's rows; a tenant-less caller sees only their own rows. accept_rate counts only verbatim accepts (edited is its own bucket, so a "kept" rate is accept_rate + edit_rate); rates are over total decisions and a zero-decision scope returns all zeros. Each decision is an append-only event — re-deciding a proposal records another event by design.
How memories are injected (fenced private context)
Stored memories are never mixed into a user message or into the text of an attached file. They are injected as their own clearly-delimited block inside the system prompt, fenced between a [Stored memory …] opening line and a [End of stored memory] closing line, and placed directly after the user's profile block, before the task and project blocks — user context, never the trailing block of the prompt (see System prompt block order below). On the one route that has no system role — Mistral with web search enabled (the Mistral Agents API, /v1/conversations) — the gateway folds the whole system prompt into the first user message; there the memory block sits inside the folded system text, separated from the user's own message by the prompt's block rule (---), never running straight into it. The fence labels the block as private operating context: the model may use the facts to inform its answers and state a specific fact when the user asks for it, but must not quote or reproduce the block (or its markers) to the user, and must never treat it as an attached document to evaluate, summarise, or review. This keeps a request like "evaluate the attached text" scoped to the attachment only, and keeps the internal <memory> proposal envelope out of visible output.
System prompt block order
The gateway composes one system message per turn, in this fixed order (a block is omitted when it has nothing to say; blocks 1-8 are joined by a --- rule; the memory-capture instruction, item 9, is appended after a blank line):
- the default prompt (identity, date, product self-help, the reply-language rule with its explicit-instruction precedence);
- the tenant's organisation policy;
- the user's profile block (preferred name, work category, the user's own instructions);
- the stored-memory block (fenced, see above);
- the curator task;
- the project block — the project's instructions, then the reply-language anchor that qualifies them, then the project-files tool rules;
- the inline file-write fallback (providers without native tools) and the imported-transcript notice, when applicable;
- the response-style block;
- the memory-capture instruction (capability-gated per model).
The per-request reply-language directive ("If the project instructions or the user's own instructions explicitly name a reply language, use that language … Otherwise: the user's most recent message is written in German. Write your entire reply in German …") is appended by the same middleware after these blocks; later content middlewares (the web-search notice, the context-compaction guard) may merge their own text after it. Client-sent system messages are dropped wholesale.
The rendered memory block is also byte-bounded: it emits at most 16 KiB into the system prompt regardless of how many memories the scope holds, so the per-turn input-token cost can never grow without limit. When the memories would exceed that budget they are included in stored order until the budget is reached and the boundary entry is truncated; truncation is deterministic (the same memories in the same order always produce the same block). With the 50-per-scope cap and the per-memory length limit this budget is not reached by ordinary use — it is a backstop against pathological content.
Memory content is treated as untrusted input at the composition boundary: any value that tries to forge the fence's [Stored memory / [End of stored memory] markers, or to embed a live <memory>…</memory> envelope, is neutralised (its leading bracket/angle is broken) so it cannot break the framing or ride into the prompt as a live tag. Empty or whitespace-only content is skipped. This is spotlighting-grade hardening — it reduces, it does not eliminate, a model's own confusion — layered over the gateway's output strip, which remains the backstop that removes any <memory> tag the model emits before it reaches the user.