Agents API
A saved agent is a named, gateway-scoped inference config (instructions +
model + tools + optional knowledge project), invoked head-lessly at
/v1/{tenant}/{gateway}/agents/{slug}/invoke. Agents are user-owned
(created_by); admin/tenant_admin govern all agents in their scope.
Admin endpoints require an authenticated admin session (aig_admin cookie). The
base URL is https://ai-api-admin.myra.eu/admin/v1.
Per-tenant feature gate (agents_enabled)
Agents is a per-tenant feature (tenant.agents_enabled, default 0 = off for every newly-created tenant / migration 0293 — a premium, not-yet-public feature; a platform admin enables it per-tenant; flip via
PATCH /admin/v1/tenants/{id} — see Tenants). Enforced server-side,
not just by hiding the SPA nav (invariant 11 — the client is never the authz boundary):
- Admin routes — every
…/gateways/{gw}/agents…route (list, get, create, update, delete, sharing, curation, rating, versions, schedules, webhook-triggers, runs) and the org agent-catalog return403 { "error": "feature_disabled" }when the tenant is disabled. - Invoke — a human agent run — the interactive
/v1/{tenant}/{gw}/agents/{slug}/invoke(user session / user-bound token), the run-as-self preview, and the in-chat agent bridge — returns403 { "error": { "code": "agents_disabled" } }(a distinct code so the SPA can localize "ask your workspace admin to enable Agents"). - Scheduled / webhook agent runs are withheld at the claim layer (no fire, no failure
email, no auto-pause); the inbound agent-webhook receiver fail-fasts
503without enqueuing. They resume automatically when Agents is re-enabled.
Exempt (still reachable while disabled, by design): the governance review decisions
(approve/flag/disable/reactivate) on agents that already exist — so a pending review can be
resolved (mirrors the workflow-approval precedent) — the gateway compliance audit read, and an
agent-step that runs inside an enabled Workflow (governed by workflows_enabled, not
agents_enabled).
Four-eyes config approval (202 pending_approval)
When the tenant enables four-eyes configuration approval, a
mutating call — create (POST …/agents), update (PATCH …/agents/{id}),
restore (POST …/agents/{id}/versions/{version}/restore), delete
(DELETE …/agents/{id}), schedule create (POST …/agents/{agent}/schedules),
schedule update (PATCH …/agents/{agent}/schedules/{id}), and schedule
delete (DELETE …/agents/{agent}/schedules/{id}) — made by a tenant admin
does not apply immediately. Instead it is held for a second admin and the
route returns 202 { "status": "pending_approval", "approval_id": "<id>" }. The
change applies only after a different admin approves it, re-validated against the
then-current state (never a stale replay); a delete and a schedule update/delete
are captured and replayed against the current row (a schedule update whose row has
since changed incompatibly fails closed at apply rather than reverting it). If the
approval gate itself is unavailable the call fails closed with 503. See
Configuration approvals for the approval lifecycle.
Knowledge areas
An agent can be linked to one or more knowledge areas (Wissensbereiche —
chat_project spaces, each with its own access control). At retrieval time the
agent's search_knowledge tool searches the union of the bound areas' documents
and returns passages with their [file | page | section] citation unchanged.
Create / update body
| Field | Shape | Notes |
|---|---|---|
project_ids |
string[] |
The authoritative set of bound areas. Replaces the binding on PATCH. Max 16 ids. Each must be a project in the gateway's tenant that the agent owner can access — otherwise 400 (fail closed). Duplicates collapse. [] clears all bindings. |
project_id |
string |
Legacy, single-area. On create it seeds a one-element set. On PATCH it is honored only for a true legacy agent (no project_ids binding yet) — once an agent has area rows, a stray project_id is ignored so it can't silently collapse a multi-area agent. |
A deleted knowledge area restricts the run (fail closed). A knowledge area is soft-deleted, and deleting one does not unbind the agents that use it — so an agent can outlive one of its areas. When the gateway resolves an agent's areas at invoke time and finds that a bound area no longer exists (deleted, or not a live project of this tenant), it cannot tell what policy that area carried. Because the missing area may have been the one that made the agent local only or PII mandatory, the run is forced to the most restrictive setting on both axes: local models only, PII masking on, and every externally-egressing tool (web search, URL fetch, external MCP, sub-agent delegation, image generation, code interpreter) disabled.
What you will see. If the agent is pinned to a cloud model, the run is refused with
403 agent_area_unresolved — "A knowledge area bound to this agent no longer exists, so the
run is restricted to local models. Open the agent and save its knowledge areas to fix it."
That is permanent, not retried: a scheduled agent in this state fails every run until it is
fixed. If the agent runs on a local model it still completes, silently restricted; the
agent_area_unresolved step in the run's trace is where that shows up.
How to see it. The agent list badges such an agent Knowledge area missing (a warning,
not an info tag — it is not a working agent), and opening it shows the same explanation directly
above the knowledge-area picker. On the API this is areas_unresolved: true on the owner view
of GET-one and the list; the field is present only when true, so its presence is the signal.
How to fix it. Open the agent and save its knowledge areas. The save writes the live set
and prunes the dead binding. A PATCH that omits project_ids does not clear it — the
binding is only rewritten when the field is present.
This is deliberate: before it, a deleted area simply vanished from the agent's area set, and an agent whose only area was local only silently became unrestricted — its scheduled runs kept working, on cloud models, with no residency guarantee.
A purged area fails closed too. The guarantee covers a hard delete — a right-to-erasure
purge that removes the project row itself, not just its deleted_at — as well as the ordinary
soft delete. (A whole-workspace purge deletes the agents with the areas, so nothing survives to
be restricted there.) The binding row is deliberately left behind by such a purge
(it is an asserted area id, and its liveness is resolved at read time), so the agent still
asserts an area the gateway cannot classify and is restricted exactly as above. Before this,
the purge removed the assertion with the project and the agent silently ran unrestricted.
Scope
This restricts what the RUN may reach. It does not change what the scheduler does with
the finished output: delivery (email / webhook / chat) is gated on PII activity, not on
residency, exactly as it is for a genuinely local_only area.
Responses. GET-one, create (201), update (200) and the list return
project_ids (the area set) alongside project_id (the derived primary = the
first area, kept for backward compat), plus areas_unresolved: true when at least one
asserted area no longer resolves (absent otherwise). For a non-owner (shared/catalog) view both
project_id and project_ids are redacted. On create/update/restore/curation
the write is authoritative and commits before the areas are read back for the
response, so a rare transient failure of that read-back does not turn a
succeeded write into a 500 (which would desync the list from the detail): the
committed agent is returned with project_ids as [], and the true set reloads
cleanly on the next fetch. (The read-only GET keeps its honest 500 — nothing is
committed there.)
Rejected input (each → 400): a non-array project_ids, a non-string element,
more than 16 ids, or an id the owner cannot access / that is not in the tenant. An id
the agent is already bound to is the exception — it is kept without re-validation,
so an owner who has since lost access to an area (or whose area was deleted or purged)
can still save the agent, and the area's restriction is not silently dropped. Only
newly-added ids are access-checked.
Retrieval & ACL. The union is resolved per invocation against the run-as-owner
identity: an area the owner cannot access is dropped from the union (never
searched, no citation) — the union never widens access. The corpus is capped at
50 documents, allocated fairly (round-robin) across the areas, so
one large area cannot starve the others; a document-count cap, not a byte cap. Egress
tiers are unioned per axis: if any bound area is local_only the whole invoke is
egress-blocked, and if any is pii_mandatory PII scrubbing is forced (the two are
independent — binding one of each keeps both protections).
Model pin offer validation
An agent pins a model (and an optional provider) that every invocation of that agent
dispatches to on the agent's gateway. When a create (POST) or edit (PATCH) sets that pin, the
server validates the resolved (provider, model) the agent would dispatch to against the gateway's
dispatch gates — the same EU data-residency and provider-allowlist checks the inference path
enforces — evaluated on the gateway's resolved (tenant-floor-folded) config. This is the same
authority the project default pin is validated against (providers.pin_gate). A pin
the gateway would refuse at dispatch is rejected at save time rather than accepted and then
failing every conversation with a 403:
- The model resolves to a provider the gateway's EU data-residency enforcement blocks →
400withcode: "data_residency_blocked". - The model resolves to a provider not on the gateway's approved provider allowlist →
400withcode: "provider_not_allowed".
Details:
- Dispatch-faithful resolution. The provider checked is the one the invoke path would actually use: a provider-remapping routing rule wins over the pin, and every load-balance fan-out target is checked (a single blocked target blocks the save).
- Touch-guarded (
PATCH). The pin is re-validated only when the request touchesmodelorprovider. An unrelated edit (rename, new instructions) on an agent that already carries a pin — including a legacy pin that predates this validation — is never re-validated, so it stays editable. - Routing-default agents (no
modelpinned) have nothing to validate and are unaffected. - Fail closed. If the gateway cannot be resolved to a folded config at save time (a transient
database error), the save fails with a retryable
5xxrather than silently storing an unvalidated pin.
Provider availability (the explicit provider field)
provider is an optional, client-supplied pin. It is validated at the server as an untrusted
input — the client is never the authority. The value is accepted when the provider can actually
be served on the agent's gateway, which is the same credential-sourcing decision the inference
(chat) path makes — so a provider that works for chat on this gateway is pinnable, and save ==
invoke by construction. Accepted when the provider is:
- platform-managed (the Myra-hosted
myraprovider) — always routable, no key, no opt-in; - a BYO provider the gateway opted into (a provider-config key → the gateway's
configured_providers); - available via the managed Anthropic pool — either a granted managed model
(
managed_models) or a model entitled by the gateway's self-serve plan. On a managed/wallet gatewayconfigured_providersis empty, yet Claude models are served keylessly on Myra's pool; this is exactly the case AGF-3129 fixed (create previously rejectedprovider: "anthropic"here with"not configured on this gateway"even though Haiku chat worked).
Rejected (400), each a distinct answer — a malformed value never degrades to the inferred
path:
- an unknown provider string →
unknown provider: <value>; - a provider that does not serve the pinned model →
provider '<p>' does not serve model '<m>'; - a provider not available on this gateway (not platform-managed, not configured, not on the
managed pool) →
provider '<p>' is not configured on this gateway. This is a fail-closed control — it is NOT dropped for the managed case, so a genuinely unavailable provider (e.g. a BYO provider with no key on this gateway) is still refused rather than accepted and then failing at the credential read; - a present but non-string
provider(123,[...],{}) →provider must be a string. A present-but-malformed value is rejected; it is not silently treated as absent.
Absent (null / "" / omitted) → the provider is inferred from the model: the agent is
stored unpinned when the inferred provider is not runnable here (routing re-infers at each invoke),
or pinned to the resolved provider when it is. A transient failure to verify availability fails the
save with a retryable 5xx (never a spurious 400).
The invoke / preview path validates a stored pin with the same predicate
(core.provider_runnable.on_gateway), so an agent pinned to a managed-pool provider runs; a pin to
a provider no longer available fails with pinned provider '<p>' is not configured on this gateway.
Known limitation. A provider-prefixed model id (e.g.
anthropic/claude-haiku-4-5) combined with an explicitprovideris not normalized before the availability check, so a managed-pool match on the bare id is missed and the request reports the generic… is not configured on this gateway. Pass the bare model id (claude-haiku-4-5) with the explicitprovider— which is what the Agents UI sends.
Tool-capability guard
A model curated as tool-incapable (e.g. Perplexity Sonar — see supports_function_calling in
Models) breaks the moment the gateway injects any function tool. So a create (POST)
or edit (PATCH) whose effective model is tool-incapable is rejected 400 with
code: "MODEL_CAPABILITY_MISMATCH" when it would leave any function-tool affordance enabled — web
search, URL fetch (agentic_fetch), the file tool, connectors (mcp / mcp_allow), sub-agents
(agents), or a bound knowledge area (which arms search_knowledge). This is the config-time
prevention counterpart to the runtime MODEL_CAPABILITY_MISMATCH net, refused consistently across
every affordance rather than one leaky toggle.
- Permissive on the unknown space. Only a curated tool-incapable model is blocked; an
uncatalogued model (and a routing-default agent with no
model) is never blocked here — the runtime net owns those, so a tool-capable model whose flag is unknown is never wrongly refused. - Effective (patch-merged) evaluation. A
tool_config-onlyPATCHkeeps the existingprovider/model; amodel-onlyPATCHto a tool-incapable model still catches tools already enabled on the stored row. The guard runs only when the request touchesmodel,tool_config, or the knowledge areas — a benign rename of a legacy tool-incapable+tools agent is not blocked.
Sharing — admin-plane only
Sharing controls who discovers, reads and clones an agent on the admin plane.
It does NOT change the data-plane invoke gate. Invoke is run-as-owner: a
user-bound caller may invoke only an agent they own, so a sharee cannot invoke
the owner's slug (AGENT_NOT_FOUND) — that would run with the owner's MCP
credentials and bound knowledge. Therefore agent sharing means see + clone-then-own,
not direct use of the owner's agent. (A private agent is still invokable by the
gateway service token, unchanged.)
Access resolves uniformly from rows + a flag: a user sees an agent if they own
it, or it is shared with them directly, via a group they belong to, or it
is org-shared (and the tenant's org_share_enabled is on). visibility
(private | users | group | org) is a derived display value.
List agents
GET /admin/v1/gateways/{gw}/agents
Returns the caller's own + shared agents (admins: all in the gateway).
Non-owned (shared) rows are redacted: tool_config is filtered to an allow-list of
non-sensitive display flags (web_search, file, agentic_fetch, sources_footer) — every
other key is dropped, so the owner's bound MCP connector ids (tool_config.mcp / mcp_allow),
sub-agent bindings (tool_config.agents), and any future credential-bearing key never reach
a non-owner. The knowledge project_id / project_ids and
source_metadata are removed too.
Query params: ?q=<name> (substring search), ?shared=1 ("shared with
me" only), ?usage=1 (annotate each owner row with usage_count, see below).
The management UI additionally applies, client-side over the returned page (no
extra request per keystroke), a tool facet (web search / knowledge files /
fetch / MCP connectors), a tag facet (the curation category), a usage
facet (used / never used), and a sort order (newest / recently updated /
name / most used). Because facets/sort run over the returned page and the list is
capped at 500 (created_at DESC), an agent outside the newest 500 is not surfaced
by "most used" — the cap is shared by all sorts.
usage_count (?usage=1 only). Invocation count for the agent over the
last 90 days on this gateway, counted from request_log (every agent-channel
row, including blocked/failed — "how often invoked"). Emitted only for owner
rows (it is the owner's own activity, like schedule_count; sharees don't get
it) and only when the caller passes ?usage=1. The aggregate scans the
high-volume request_log, so it is opt-in — the UI requests it lazily, only when
the user engages the popularity sort or usage facet. Absent = treat as 0.
Read an agent
GET /admin/v1/gateways/{gw}/agents/{id}
Owner/admin get the full config; a sharee gets the redacted view; everyone
else 404.
Deleting an agent
DELETE /admin/v1/gateways/{gw}/agents/{id}
Deletes the agent. Restricted to the agent's owner or a tenant_admin/admin
(a non-owned / cross-tenant / missing id → 404, no existence leak — the same gate
as the other agent routes). On success the response is 200 { "ok": true }, and the
deletion is recorded in the audit trail as agent.deleted.
Version history & restore
Every create/edit snapshots the agent's full config into an append-only
agent_version history (agent.version is a monotonic counter). These endpoints
browse that history and restore an old version. All three are owner or
tenant_admin/admin only — a non-owned / cross-tenant / missing agent id →
404 (no existence leak), same gate as the other agent routes. The publish model is
always-latest: the org catalog keeps cloning the agent's current live config,
so history is read-only and restore is non-destructive (there is no pinned
"published" version).
A run records which agent_version produced it, so run-detail can reconstruct
the exact config afterwards. Because invoke resolves the agent live (always
latest), the version is stamped from the value the invoke actually ran — every agent
invoke returns an X-AIG-Agent-Version response header, and the scheduler writes it
onto the run at completion. This closes a window where a scheduled run pinned the
version at poll time but a newer edit ran at invoke time, which would have made
run-detail show a config that never ran. When the header is absent (an
older gateway), the run keeps its poll-time snapshot.
GET /admin/v1/gateways/{gw}/agents/{id}/versions — the history, newest first
(bounded to the most recent 200). Each entry: id, version, name, model, provider,
created_by, created_at. Empty history encodes as [].
GET /admin/v1/gateways/{gw}/agents/{id}/versions/{version} — one version's full,
decoded config. version must be a positive, finite, whole integer in
[1, 2147483647]; anything else (non-numeric, 0, negative, fractional, inf,
over-range) → 400. A syntactically-valid but non-existent version → 404
(distinct from a 500 on a storage error).
POST /admin/v1/gateways/{gw}/agents/{id}/versions/{version}/restore — restore
the chosen version. Not a mutation of history: it re-applies that version's config
through the normal update path, bumping the agent to a new version whose body
equals the old one (auditable, agent.version_restored). The request body is
ignored — the config comes entirely from the stored snapshot keyed by the validated
version. Blocked for viewer/demouser (403). The restored knowledge-area
set is re-validated against the owner's current access, with one exception: an area
the agent already asserts is carried through unvalidated, including one that was
deleted or purged since — that is what keeps such an agent fail-closed instead of
making its whole version history unrestorable. Only an id the agent does not
currently assert is access-checked, and a failure there is a 400. An explicit
empty snapshot ([]) clears the areas, while
a pre-history (NULL-snapshot) version leaves the current area binding untouched. The
tool_config (incl. MCP connector ids) is restored verbatim — faithful to the
old config, no re-resolution. A tool_config.mcp connector that was deleted since
the snapshot is not FK-backed: it is simply skipped at invoke (fail-closed) and
pruned on the next edit, exactly as for any live agent whose connector was later
removed — so it needs no restore-time re-validation.
The provider pin is re-validated for runnability (mirroring the edit path):
restoring a historical pin whose provider is no longer configured or no longer
serves the model on this gateway would otherwise land a silently non-invokable agent
(the next invoke 502s as AGENT_MISCONFIGURED). If the historical pin is still
runnable it is kept; if it is genuinely not runnable here it is cleared (the
agent infers its provider at invoke, self-healing) and the response carries
provider_reset: true so the UI can warn; a transient verification failure (a DB
read error) fails the restore closed (500) rather than silently unpinning a
valid agent. A version that was historically unpinned stays unpinned.
source_metadata (import provenance, not run config) is not versioned, so restore
leaves the current value. Restoring the current version is allowed and is a
harmless no-op-content bump to a new version.
The restored pin is also run through the model pin offer validation:
if the gateway's residency/allowlist has tightened since the version was saved (or the pin was never
offerable), restoring an old US-only model onto a now-EU-enforced gateway is refused with 400
(data_residency_blocked / provider_not_allowed) rather than restored into an agent that then
403s every run.
Setting sharing
PUT /admin/v1/gateways/{gw}/agents/{id}/sharing
Atomically replaces the agent's sharing state. Owner (or admin/tenant_admin) only. Body:
visibility |
Extra body | Gate |
|---|---|---|
private |
— | owner |
users |
user_ids[] (≤200, each same-tenant) |
owner — member self-service |
group |
group_ids[] (≤200, each same-tenant) |
ki_manager + owner |
org |
— | ki_manager + owner + tenant org_share_enabled=1 (else 403 org_share_disabled) |
Cross-tenant user_ids/group_ids → 404. Audited
(agent.sharing_changed, before/after subjects). The current state is read via
GET …/agents/{id}/shares (owner/admin only).
Org agent catalog
GET /admin/v1/tenants/{id}/agent-catalog
Tenant-wide list of org-shared agents (redacted), gated on
require_tenant_access and the tenant org_share_enabled (read at query time, so
disabling the capability empties the catalog). demouser is excluded. Reuse a
catalog agent by cloning it (a fresh create validated as the cloner — the
owner's private connector/project bindings are stripped/rejected, never copied).
Curation. Each row carries its curation metadata: category (a
free-text section tag, or null) and featured (0|1). The list is ordered
featured-first, then newest-first within each band — so pinned agents surface at
the top (the catalog "featured section"). The admin UI renders a Featured badge,
a per-row category chip, and a category filter over the returned page.
Rating signal. Each row also carries the community rating: rating_avg
(mean of all 1–5 ratings, rounded to 1 dp, 0 when unrated), rating_count, the
caller's own my_rating (1..5, or null), and can_rate (server-decided —
false for the caller's own agent). These are computed with two grouped queries
over the whole tenant page (no N+1); my_rating/can_rate make the response
per-caller. The same four fields also appear on the shared (non-owner) Read an
agent response.
Curate a catalog agent
PATCH /admin/v1/gateways/{gw}/agents/{id}/curation
Sets the org-catalog presentation of an agent. tenant_admin/admin only — an
ordinary owner must not self-feature their agent, so this is admin-gated even though
the owner can otherwise edit the agent (403 curation requires tenant_admin
otherwise). It is a separate endpoint from the generic agent PATCH because
curation is presentation, not run-affecting config: it does not bump
agent.version or write an agent_version snapshot.
| Field | Shape | Notes |
|---|---|---|
category |
string | null |
Section tag, max 64 bytes. null clears it (uncategorized). Absent = left unchanged. |
featured |
boolean |
true pins the agent to the top of the catalog. Absent (or null) = left unchanged. |
Rejected input (each → 400): a non-string category, a category over 64
bytes, a non-boolean featured, or a body with neither field. Featuring a
non-org-shared agent is accepted but inert (it only surfaces once the agent is
org-shared). Audited (agent.curated, before/after).
In the admin UI this endpoint backs the Agent catalog, opened from the
Agents page via its Agent catalog button — a modal with a searchable list
of the org-shared agents and a per-row Clone action. The button is hidden when
the tenant's org_share_enabled is off. (There is no separate catalog page; the
old /agents/catalog route redirects to /agents.) Cloning targets the source
agent's own gateway, so its pinned (provider, model) stays runnable; the clone
becomes the caller's own editable agent with the owner's connectors/knowledge
omitted (rebind your own after cloning).
Rate a catalog agent
PUT /admin/v1/gateways/{gw}/agents/{id}/rating — body { "rating": 1..5 }
DELETE /admin/v1/gateways/{gw}/agents/{id}/rating — retract the caller's rating
A rating is the caller's own feedback (like conversation feedback), not agent
config, so these use the catalog's visibility gate — require_gateway_access
(tenant membership) plus a demouser block — not the author gate, so a viewer
who can see the catalog can also rate. You may rate any agent you can see
(org-shared / shared-to-you) except your own.
There is exactly one rating per (agent, user): the write is an upsert, so
re-rating updates in place and can never inflate rating_count. Both routes
return the fresh aggregate: { rating_avg, rating_count, my_rating } (my_rating is
null after a DELETE).
Rejected input:
| Condition | Status |
|---|---|
rating missing / null / not a JSON number / not a whole number / outside 1..5 (e.g. 0, 6, 3.5, "4") |
400 (fail closed — not tonumber-coerced) |
| agent not visible to the caller (private, not shared, or another tenant) | 404 |
the caller is the agent's owner (created_by) — anti-inflation self-rate |
403 cannot rate your own agent |
demouser |
403 |
| unauthenticated | 401 |
Backing table agent_rating (migration 0133): PRIMARY KEY (agent_id, user_id)
enforces the one-per-user rule; agent_id CASCADEs on agent hard-delete. In the
admin UI the Agent catalog modal renders each row's aggregate plus a 1–5 star
picker (shown only when can_rate).
Clone lineage (source_agent_id) — for per-agent usage stats
The create endpoint (POST /admin/v1/gateways/:id/agents) accepts an optional
source_agent_id — the id of the agent being cloned. The clone UI sends it
automatically. It powers the per-agent distinct-users and conversations
roll-up in /admin/v1/stats/analytics (each clone's usage counts toward the shared
original; see the stats reference).
- Accepted: a string id of an agent the caller can see (own, org-shared, or shared to them/their group), on the clone's target gateway. The server re-validates through that visibility gate and flattens the value to the source's own lineage root, so the lineage tree stays one level deep.
- Rejected / ignored (fail-open to NULL): an absent, non-string, unknown, or
not-visible
source_agent_idis dropped toNULL— the new agent becomes its own lineage root and the create still succeeds. It is never an error (a share revoked between opening the catalog and cloning must not break a valid create), and a source the caller cannot see can never be attributed — so lineage cannot be forged across tenants or onto another member's private agent. Immutable after create (not editable viaPATCH). Redacted from the catalog/sharee view.
System prompt on invoke
When an agent is invoked, the gateway assembles the agent's system message from these parts, in order:
- Today's date — a
Today's date is <Month D, YYYY>.line is prepended to every invocation (from the same source the/easychat surface uses). Without it the model has no notion of the current date and treats "today" as unknown or future — a "Daily briefings" agent was observed refusing to run a web search, claiming the current date "lies in the future" and citing its training cutoff. The date is a plain fact and is injected unconditionally, for every provider and every agent (including structured-output agents). - Organization policy — the tenant's
system_instruction(if set), as a steering preface. - Anti-fabrication guardrail — a tool-agnostic rule that the agent must not
claim it searched the web, looked something up, read a document, or ran a tool
unless a tool actually returned it (that turn or earlier in the conversation);
answering from the model's own knowledge is still allowed, just not dressed up as
a lookup. This is the same rule the
/easychat surface carries — a single shared source — applied on every agent completion (invoke, draft preview, scheduled task) so smaller models don't announce-then-stop or invent a lookup's results. (The companion "emit the search call itself" nudge is added separately, only when theweb_searchtool is offered.) - The agent's
instructions— the agent owns its persona; caller-suppliedsystem/developerturns are dropped and never override it. - Structured-output directive — appended when the agent has a
response_schema(see below).
Callers cannot influence any of this beyond the request input/messages: the
agent is the sole source of truth for its system prompt.
Web-search sources footer (tool_config.sources_footer)
When web search is enabled, the gateway appends a "Quellen" (sources) footer
to the answer — a numbered list of the web result URLs the turn used, matching the
inline [N] citation markers. This footer is added by the gateway after the model
finishes, so a prompt instruction like "don't list sources" cannot remove it.
sources_footer(boolean, defaulttrue): setfalseon the agent to suppress the footer entirely. Editable in the agent editor as "Include sources list in the answer", shown under the Web Search toggle. Only affects web-search output; an agent without web search never emits the footer regardless.- Cap: even with the footer on, at most 100 sources render. Beyond that a single
line reports how many further distinct sources were hidden (
… (N weitere Quellen ausgeblendet)) — this bounds the runaway 300+-link footers that used to flood scheduled-agent email / Mattermost delivery. A citation whose number exceeds the cap degrades to plain[N]text in the web UI rather than a link. - Trade-off when off:
sources_footer: falseremoves all web-source visibility, including the inline[N]links — web-search sources have no separate structured panel (that panel is for knowledge-area / file citations), so with the footer gone the[N]markers stay plain text.
Streamed interactive invoke (stream)
The invoke and draft-preview request bodies accept an optional stream
field:
- Accepted: a JSON boolean.
falseor absent → the historical buffered JSON response (the default; the scheduler and API integrations are unchanged).true→ the response is SSE (text/event-stream). - Rejected: any other shape —
"true",1,null, an object or array — is refused with400 INVALID_REQUEST(never coerced). This deliberately tightens the previous behaviour, which silently ignored the field.
stream:true exists for the interactive surfaces (Try-It, draft Preview,
the workflow step-test): a buffered invoke writes zero bytes until the tool
loop finishes, so an idle CDN connection dies at the CDN's idle cap while a
long agent is still working. A streamed invoke opens the wire immediately and
the gateway emits an SSE comment heartbeat (: hb) roughly every 15 seconds,
so the connection is provably live for the whole run.
Semantics of a streamed invoke:
- Policy still wins. Runs that must be validated as a whole are executed
internally buffered and the finished, validated body is then re-emitted onto
the open SSE wire as ordinary chunks: agents with a
response_schema(egress validation +STRUCTURED_OUTPUT_FAILEDare enforced exactly as documented above), and gateways with a response-phase guardrail detector the live stream cannot satisfy (a content-safety block, or a regex/presidio scrub). A gateway whose only response detector is an inline PII token-restore masker (pii_protector/custom_pii) streams incrementally — tokens are restored on the live wire, no internal buffering.stream:trueis a wire format, never a policy bypass. - Headers: the
X-AIG-Trace-Id,X-AIG-Agent-Version,X-AIG-Tool-Warnings,X-AIG-PII-ActiveandX-AIG-Cost-Microsresponse headers are absent on a streamed invoke (they are committed before the run finishes). A caller that needs them — the scheduler's fail-closed PII-egress gate readsX-AIG-PII-Active— must call buffered; the scheduler always does. - Errors after the wire opens arrive as a terminal SSE event followed by
data: [DONE]. A client that opted into the gateway'saig_*side channel (x-aig-turn-id, orx-aig-extensions: 1) getsaig_status:"provider_error"(error_class,message,user_message.body); a plain OpenAI-compatible client gets the OpenAI-shaped{"error":{"message":"…","type":"gateway_error","code":"<error_class>"}}instead, so its schema validation does not abort the stream. See Inference — what a plain client receives. - No response caching: a streamed invoke bypasses the response cache
(parity with streamed chat). In fact no agent invoke — streamed or buffered
— is ever served from or written to the gateway response cache, even when
cache_ttlis set on the gateway: agent runs must return fresh output every run, and every buffered agent invoke must reach the full response path so its fail-closedX-AIG-PII-Activeegress header is emitted for the scheduler (including when a request-phase guardrail blocks the run). See Response caching.
Unattended runs — extended tool-loop budget
Every agent turn is bounded by a wall-clock budget on its server-side tool loop (web search, file tools, MCP): 360 s by default. That cap exists to protect interactive clients from silence timeouts. An unattended caller — the built-in scheduler, or your own automation that owns its HTTP read timeout — can opt into a 600 s budget per invoke:
Accepted shape and rejection (the gate fails closed to 360 s):
- The value must be exactly the string
1. Any other value (true,0, empty, whitespace-padded), a repeated header, or a non-string is ignored. - The request must authenticate with a non-user-bound gateway service token. A user-bound API token, an SPA session, or an auth-less gateway never gets the extended budget — the header is silently ignored, never an error.
- Applies to
POST …/agents/{slug}/invokeand the scheduled-task invoke. The scheduler sends it on every run it fires.
The header extends wall-clock only: the per-turn tool-round cap and the token's spend budget are unchanged. If your client opts in, raise your own read timeout above ~800 s (budget + one in-flight tool round); the gateway sends the response as a single JSON body at the end of the run.
Budget exhaustion delivers a partial answer
When the wall-clock budget runs out after tool results were already gathered, the run is not discarded: the gateway makes one final model call without tools (bounded to at most ~60 s and 1024 output tokens) over the results collected so far and delivers that partial answer, followed by an honest note — "Stopped early: the time budget ran out — this answer is based on partial results." (German for users whose profile language is German; runs invoked with a service token have no user profile and receive the English note). The note is part of the answer text, so it persists with the turn.
This applies to interactive chats and unattended runs alike. If the budget
runs out before any tool result exists, or the final call itself fails or
produces nothing, the run ends with the unchanged
Stopped: tool-loop time budget exhausted (360s/600s) note instead — the
gateway never invents an answer from tool calls that did not complete.
Human-in-the-loop tool gate — suspend / review / resume
An agent can require a human to approve specific tool calls before they run in an
unattended (scheduled / webhook) run. Configure it on the agent's tool_config:
approval.tools is a list of tool wire-names whose calls must be human-approved. Scope: the gate matches any tool that dispatches this run by its wire name — a gateway
builtin (write_file, fetch_url, web_search, agentic_fetch, read_file, search_knowledge,
code_interpreter, generate_image), an MCP-connector tool (e.g. send_email), or a
sub-agent tool (agent__<slug>). Tool wire-names are unique per run, so a name in the list
matches exactly the tool the model called; a listed name that dispatches to no tool this run
never matches (it could not run anyway). Absent / empty approval.tools = OFF (today's behaviour,
byte-for-byte). The list is validated at agent create/update (see the four-eyes/create contract):
each entry is a non-empty, bounded, de-duplicated string; a malformed approval is rejected
400, never silently dropped (a dropped gate would run a tool the admin meant to hold).
Approval authorizes the tool NAME/action — not a bypass of live controls. When a held MCP / sub-agent / egress-builtin call is approved and resumes, it still passes every runtime control at its execution chokepoint: the fail-closed egress gate (an approval granted before the project became
local_only/ no-egress does not egress after — the held tool is blocked on resume), the MCP-argument PII block (approved args carrying protected personal data are still refused), and connector tenant-isolation + sub-agent same-owner authz.
Behaviour. The gate applies only to unattended runs (the scheduler; a stream:false
service-token invoke) — an interactive human invoke is never suspended (the human is present).
When the server-side tool loop reaches a round whose executing batch contains a gated tool
(builtin, MCP-connector, or sub-agent), the run suspends between legs: it persists enough
state to resume — including the dispatch identity of any held MCP / sub-agent tool so it can run
after approval — records a review_item
(status pending), and stops without running any tool in that batch (whole-batch hold — a
non-gated sibling in the same round is held too, so no side effect fires while a gated call
waits). The run's status becomes suspended; no output is delivered.
A tenant admin decides from the review inbox (see the decision API below):
- Approve → the scheduler's resume sweep re-drives the run: the held tool(s) execute and the
loop continues to completion (then delivers per the schedule, subject to the unchanged
fail-closed PII egress gate + the optional delivery-approval hold). The decided batch executes
exactly once — a retried or duplicate resume runs the held tools at most once (see
X-AIG-Resume-Run). - Deny → the run resumes with the gated tool(s) returning a denial result
(
[denied] This action was declined by a human reviewer.); non-gated siblings still execute and the agent continues, adapting to the denial. The same exactly-once fence applies, so the non-gated siblings also run at most once under a retried resume. - Undecided for 24 hours → the review expires and the run ends
failed(review not decided); the gated tool never runs (fail-closed). Expiry counts as a denial with no side effect.
🔒 Only a tool the gateway actually offered can run. The review queue shows the gated names, so a non-gated sibling in the same batch executes on resume with no separate approval. Whether each held call named a tool the gateway offered is therefore decided when the batch is suspended — the only moment that is knowable — and recorded with it; on resume a call marked unoffered is refused and the model receives
[error] tool '<name>' is not available on this turn …in its place, while the batch keeps its original shape so every other call defers and executes exactly as it would have. The same check gates dispatch on the live path, so a model that invents a call for a tool that was never sent cannot execute it.The parallel-tool cap in force when the batch was held is recorded with it for the same reason: a call that was beyond the cap — and so was never shown to the reviewer — still defers on resume even if
max_parallel_toolsis raised while the review is pending. It re-enters the executing window on a later round, and is reviewed then, exactly when it would run.Reviews created before this shipped resume unchanged.
A resumed run that reaches another gated call in a later round suspends again (a fresh
review). Personal-data masking is latched across the approval: if the pre-suspend legs
masked PII, the resumed run's X-AIG-PII-Active stays 1, so restored PII can never egress
because a run happened to pause.
Correlation headers (service-token only)
The built-in scheduler carries two headers on its unattended invoke. Both are honored only
for an unattended service-token invoke (the same gate as X-AIG-Unattended); a user-bound
or browser caller's headers are ignored (fail-closed). They are re-validated server-side.
| Header | Value | Meaning |
|---|---|---|
X-AIG-Run-Id |
the agent_run id |
Attaches a suspension to this run. The run must belong to the URL's (gateway, agent) — a mismatched / unknown id is refused (no cross-run / cross-tenant resume). |
X-AIG-Resume-Run |
exactly 1 |
Re-drive a suspended run: the gateway rehydrates the persisted conversation and executes the decided batch. Requires the run to be claimed (running) with persisted suspend state; otherwise INVALID_REQUEST. The decided batch executes exactly once per decided review: a per-review execution fence (an atomic compare-and-swap that marks the review executed before any held tool fires) refuses a second/duplicate resume with INVALID_REQUEST once the first has claimed execution — so a side-effecting gated tool can never double-fire, even across a scheduler retry. (A crash after the fence is claimed but before the run finishes forfeits the remainder of the run rather than re-running the batch — at-most-once is the safe direction for side effects.) If the fence cannot be evaluated the resume refuses fail-closed and runs nothing — INTERNAL when the persisted review reference is missing/corrupt or the fence's datastore is unavailable. |
Review decision API
Base URL https://ai-api-admin.myra.eu/admin/v1. All three require the EGRESS_APPROVE
permission (admin / tenant_admin governance — the same permission as the agent delivery-approval
inbox) and are tenant-scoped server-side (the client is never the authz boundary). Rows
carry only the gated tool name(s) — never the tool arguments or any personal data.
| Method | Path | Purpose |
|---|---|---|
GET |
/admin/v1/review-items |
This tenant's reviews. ?status= (pending|approved|denied|expired|closed|all) and ?kind= (tool_gate|flagged_prompt|flagged_output|all) filter — any other value is rejected 400 (fail closed, never reaches the query). Returns { "items": [...] } ([] when none). |
POST |
/admin/v1/review-items/{id}/approve |
Approve. Optional { "note": "<reviewer annotation>" } (length-capped). |
POST |
/admin/v1/review-items/{id}/deny |
Deny. Optional { "note": "<reviewer annotation>" } (length-capped). |
The optional note (both decisions) is stored as review_item.decision_note — the reviewer
annotation.
Decision semantics (both): a POST-after-auth tenant-scoped compare-and-swap. A
cross-tenant / unknown id → 404 (no existence leak). A second decision on an
already-decided review → 409 already_decided; an approve after the 24h TTL lapsed →
409 expired (fail-closed — an approve can never beat the expiry sweep and run a tool that
was never approved in time). A storage fault → 503 (never a 4xx). The decision is
recorded in the audit trail as review_item.approved / review_item.denied.
The review_item table is a generic review queue (shared with the flagged-prompt /
flagged-output moderation queue below): its run_id is nullable and un-foreign-keyed, so a review
need not be tied to an agent run. The held tool-loop resume state lives on
agent_run.suspend_state, not on the review row.
Flagged-prompt / flagged-output moderation queue
The same inbox and decision API moderate guardrail-flagged interactions. When a guardrail
detector's action is flag (observe-only — the request is not blocked; see
Guardrails) on a top-level turn, the gateway records a review_item:
kind=flagged_prompt(a request-phase flag) orflagged_output(a response-phase flag);run_idisnull(no agent run).subject_ref= the interaction's trace reference (request_log.trace_idwhen the request is traced, else itsrequest_id) — a display-only pointer. The prompt / output text and the detected entities are NEVER copied into the review row — a security reviewer correlates thesubject_refin the Request Logs view, which enforces its own access controls.summary= the flag detector name(s) (admin-authored gateway config), never the matched content.
A security reviewer allows (approve), denies (deny), or annotates (note) each
flagged interaction through the SAME /admin/v1/review-items decision API and inbox above
(permission EGRESS_APPROVE; tenant-scoped; the same 404 / 409 / 503 semantics).
Scope + lifecycle:
- Top-level turns only. Inner tool-loop / continuation legs, agent-as-tool delegated
sub-agent legs, the scrub-preview draft path, and a
count_tokensestimate (a local, non-egressing call a client issues right before the real completion) do not create moderation rows — no rows for un-sent drafts, internal fan-out, or duplicates of the completion. A top-level agent invoke (interactive or unattended scheduled) is moderated — a flagged unattended prompt/output is exactly what a reviewer wants. - One row per flagged top-level turn per phase (≤ 2 per turn); multiple flag detectors in a
phase are joined into that row's
summary. - Best-effort + fire-and-forget. The row is written from a background timer, so a moderation write never adds latency to the request and can never break it — the flag verdict already applied. A write failure is logged, not surfaced.
- Retention. A still-
pendingflagged review moves toexpiredafter a 30-day active window (it gates nothing, so nothing is failed on expiry — unlike atool_gate). Terminal flagged rows (approved/denied/expired) are then pruned by the shared agent-history retention sweep.
Draft preview — run-as-self
POST /v1/{tenant}/{gateway}/agents/preview dry-runs an unsaved agent draft so
the create/edit UI can show exactly how the agent will behave before it is saved.
Unlike /agents/{slug}/invoke (which runs as the agent's owner), preview runs
as the current user (the identity bound to the playground/personal token that
calls it) — so it faithfully exercises the SAME server-side assembly a saved invoke
uses (today's date, organization policy, instructions, web search, agentic fetch,
MCP connector discovery, and knowledge-area/RAG binding) without needing a saved row.
The request body is an inline draft config and is treated as untrusted input, validated at the trust boundary against the caller's identity (fail closed):
| Field | Accepted | Rejected → 400 |
|---|---|---|
model |
non-empty string (required) | missing/empty/non-string |
instructions |
string ≤ 40000 chars (optional) | non-string / too long |
provider |
optional pin: platform-managed (myra), a provider configured on this gateway, or one available via the managed pool (granted managed model / self-serve plan); absent → inferred from the model (see Provider availability) |
unknown provider; a provider that does not serve the model; a provider not available on this gateway; a present non-string value |
max_tokens |
positive integer ≤ 200000 (optional) | non-integer / ≤ 0 / too large |
tool_config |
object with the known agent keys (web_search, agentic_fetch, sources_footer, mcp, mcp_allow, file, agents, approval) |
unknown key / malformed / oversized |
tool_config.sources_footer |
boolean; default true (footer shown). A non-boolean is coerced to the default (shown), not rejected |
— (never rejected; coerced) |
tool_config.mcp |
connector ids the caller may use — in-tenant, and (if private) owned by the caller; at most 10 | a foreign/private, cross-tenant, or non-existent connector; more than 10 connectors |
project_ids |
knowledge-area ids the caller can access (member/admin), in the gateway's tenant, ≤ 16 | a foreign / inaccessible / cross-tenant area |
input / messages |
a non-empty input string or a messages array |
neither present, or a malformed turn |
Caller-supplied system/developer turns are dropped (the draft owns its system
prompt), and tool_config.agents (sub-agent delegation) is ignored in preview —
a dry-run never spawns owner-scoped sub-agent runs.
Authorization: preview requires a user-bound token holding the AGENTS_AUTHOR
permission — the same permission the agent create/edit routes require. For the built-in
roles that is member, ki_manager, tenant_admin, and admin; a tenant custom role
granting AGENTS_AUTHOR is admitted regardless of its base role. An unbound (gateway
service) token, and any caller lacking AGENTS_AUTHOR (e.g. viewer / demouser) —
none of which may create or own an agent — are refused (403). POST-only (405
otherwise). Because it runs as the caller with the caller's own connectors/areas,
preview grants no capability the caller doesn't already have. Preview runs are marked
as agent-shaped traffic but carry no agent_id, so they never affect per-agent
usage statistics.
Why direct invoke isn't shared
A sharee who can see + clone an agent still cannot invoke the owner's slug
with their own token. The agent runs as its owner (created_by), so granting a
sharee invoke would expose the owner's MCP credentials and bound knowledge files
— an escalation the run-as-owner boundary exists to prevent. Reuse is via
clone-then-own (the clone re-resolves the model/provider for the cloner's
gateway and binds the cloner's own connectors/knowledge).
Agent-as-tool — delegation
A supervisor agent can call another saved agent as a tool. Declare the
callable sub-agents in the supervisor's tool_config:
Each entry is a sub-agent slug on the same gateway. At invoke time the
gateway offers the supervisor's model one agent__<slug> tool per resolvable,
same-owner sub-agent (input schema: { "input": "<task>" }). Calling it runs
the sub-agent's full configured pipeline (its own model, tools, MCP
connectors and knowledge) via an internal in-process subrequest, and returns its
final answer to the supervisor as a tool result. The sub-agent runs as the
supervisor's owner — the same run-as-owner boundary as a direct invoke.
Accepted / rejected (tool_config.agents): an array of ≤ 8 valid slugs
(^[a-z0-9][a-z0-9_-]*$, ≤128 bytes), deduplicated. A non-array value, a
malformed slug, a duplicate, or more than 8 entries is rejected with 400 at
create/update. A declared slug that no longer resolves, or is owned by a
different user, is silently not offered (fail-closed) — it never becomes a
callable tool.
Guards (non-negotiable, fail-closed). Every delegation in a run draws from one run-wide budget and is refused (returned to the model as a tool-error, never a crash) when any limit is hit:
| Guard | Limit | Refusal |
|---|---|---|
| Nesting depth | 3 | delegation_depth_exceeded |
| Cycle (A→B→A) | active call-path repeat | delegation_cycle_detected |
| Total sub-invokes (fan-out) | 8 per run | delegation_fanout_exceeded |
| Token budget | 300 000 tokens per run | delegation_budget_exceeded |
| Wall-clock (run) | 240 s per run | delegation_time_exceeded |
| Wall-clock (per child) | 180 s per sub-agent | delegation_deadline_exceeded |
The depth, fan-out, token and run wall-clock guards are checked at the entry gate
before each sub-invoke. The per-child wall-clock (180 s, or the remaining run
budget if that is sooner) is additionally enforced mid-flight, so a single sub-agent
that keeps running cannot overshoot: its provider stream read is cut when the deadline
passes, and its own tool loop stops cleanly between rounds with a
delegation_deadline_exceeded note. The per-child cap is deliberately tighter than a
direct (non-delegated) turn's tool-loop budget so one runaway sub-agent can't consume the
whole run and starve its siblings; a legitimate multi-round sub-agent finishes well inside
it. (A sub-agent leg already producing output but stalling between chunks is interrupted
when its current read returns — the same granularity as the client-disconnect kill; a leg
that goes silent right after a late chunk can therefore overshoot the per-child cap by up to
one inter-chunk read budget before it is cut.)
A diamond (A→B→C then A→C) is allowed — only a true cycle on the active path is
refused. This includes a supervisor that lists its own slug (S→S) and a
deep call back to the run's originating supervisor (S→…→S): both are refused as
cycles (a self-listed slug is additionally never even offered as a tool). A
malicious or looping agent graph therefore always terminates.
Security. The agent__ tool-name prefix is reserved: an MCP connector that
exposes a tool named agent__* is rejected, so it can never shadow or hijack a
sub-agent tool. Cross-owner delegation is refused unconditionally — a
sub-agent's owner must equal the (already-resolved human) owner of the run, with
no service-token / unattended-automation exception, so a scheduled supervisor run
can never borrow another owner's agent (and its credentials). The model-supplied
input is treated as untrusted (byte-capped; the sub-agent runs its own
guardrails/PII on it), and the sub-agent's returned text is fenced as untrusted
before the supervisor model sees it.
Billing + lineage. Each sub-agent invocation is a separate request billed
under its own request id, attributed to the same owner; the parent's
agent_delegate trace step records parent_run_id, child_run_id and depth,
and the child's request_log.meta carries parent_request_id. Nested delegate
runs are metered and logged like any other run — a build guard ensures a
configuration change can never leave nested runs silently unbilled.
Parallel fan-out (agent_fanout)
Calling sub-agents one at a time costs the sum of their wall-clock times. When a
supervisor wants several independent sub-agent tasks done at once (e.g. "run 3
review agents on this diff"), it can call the agent_fanout tool instead —
offered automatically alongside the per-slug agent__<slug> tools whenever at
least one is offered — with an array of {agent, input} pairs:
{
"agents": [
{ "agent": "research-helper", "input": "Check the API section for accuracy." },
{ "agent": "summarizer", "input": "Summarize the intro in 2 sentences." }
]
}
Every named sub-agent must already be one of the supervisor's declared,
resolvable, same-owner tool_config.agents — agent_fanout does not grant
access to anything a plain agent__<slug> call couldn't already reach; it only
changes HOW MANY calls run per model turn and WHEN they run. The children run
concurrently (within the same request) and their answers are merged
into one tool result inside ONE outer untrusted-content frame around the whole
batch — each child's answer is individually labeled (not its own separately
closed frame — only the outer batch frame is opened/closed, execute_tool
closes it once at the end) and defanged (reusing the same marker-forgery
defenses agent__<slug> fencing uses — see Security above); every refusal
string in the merge (including a raw, unresolved model-supplied agent
value) is defanged too, so a crafted agent field can't forge a fake
close/open marker pair inside its own refusal text.
Accepted / rejected (agents):
| Condition | Outcome |
|---|---|
| Not an array, empty, or more than 5 entries | whole-batch refusal, no sub-agent is run |
An entry isn't an object, or agent/input is missing/non-string/empty |
that entry's own refusal; the rest of the batch still runs |
input exceeds the byte limit (same cap as agent__<slug>) |
that entry's own refusal |
agent is not one of this supervisor's configured, resolvable sub-agents |
that entry's own refusal |
| Depth / tokens / deadline (single early check) or fan-out (N-aware early check) obviously over budget for the WHOLE batch | a fast, whole-batch pre-flight refusal before any sub-agent runs |
| Fan-out budget exhausted specifically by concurrent siblings (invokes race) | an individual entry's own refusal from its race-safe late re-check, even though the early batch check passed |
| One child's sub-request crashes, or returns an error / empty answer | that entry's own error string; siblings are unaffected — this is a partial merge, not an all-or-nothing failure |
The agents array shares the SAME run-wide guards as sequential delegation
(depth 3, 8 total sub-invokes per run, 300 000-token budget, 240 s run
deadline, 180 s per-child deadline) — a fan-out batch cannot raise the ceiling,
only spend it faster. Every fanned-out child's invocation is counted exactly
once against the fan-out guard, and every child's real token usage is counted
exactly once against the token budget, regardless of the order the concurrent
children finish in.
Trade-offs, named explicitly (not silent):
- Shared output budget. The merged batch shares ONE 80 KB tool-result cap
across up to 5 children (a sequential agent__<slug> call gets the full 80 KB
to itself) — a fan-out of several verbose sub-agents divides that budget
between them.
- Token-ceiling overshoot window. The 300 000-token run-wide ceiling is a
pre-flight check against already-completed usage, not a hard per-call
limiter (true sequentially too) — under concurrency, up to 5 children can
each independently pass that check before any of them is reconciled, so the
worst-case transient overshoot widens from "one child's response" to "up to
5 children's worth" before the NEXT call sees the corrected total.
- Client-abort is not retroactive. If the client disconnects while a
fan-out batch is already running, the already-spawned children are not
killed (the underlying mechanism cannot interrupt an in-flight sub-agent
invocation) — they run to completion or their own deadline, bounded by
5 × 180 s in the worst case, same as an in-flight agent__<slug> call today.
Privacy. When agent names a sub-agent that isn't actually one of the
supervisor's configured, resolvable sub-agents, the refusal identifies the
attempted (model-supplied, unbounded) value the same way every other
delegation refusal does — but because that value never resolved to a real,
gateway-known identifier (unlike every other refusal reason, where the named
sub-agent IS a small, bounded, already-configured slug), the value persisted
to the queryable trace record for this ONE refusal reason follows the same
opt-in gate every other piece of raw request content does: omitted for a
private (ghost) turn or when the gateway hasn't opted into
tracing.include_bodies, present otherwise. The refusal reason itself
(agent_not_configured) is always recorded regardless.
Structured output (response_schema)
An agent can require its answer to be a JSON value conforming to a schema, so a
downstream consumer (a scheduled webhook, another agent in a chain) can parse it
reliably instead of scraping prose. Set response_schema on the agent to a JSON
Schema object:
{
"response_schema": {
"type": "object",
"required": ["summary", "sentiment"],
"properties": {
"summary": { "type": "string" },
"sentiment": { "enum": ["pos", "neg", "neutral"] },
"items": { "type": "array", "items": {
"type": "object", "required": ["id"],
"properties": { "id": { "type": "integer" } } } }
}
}
}
Accepted / rejected (response_schema, each → 400 at create/update). The value
must be a JSON object (a top-level array or scalar is rejected — a schema root is
an object), must fit the 60 000-byte column, and must not be excessively nested (at most
24 levels of nested objects — deeper is rejected so egress validation can descend
it in full without a silent depth cut-off). An absent / null value means the agent has no structured-output
requirement (unchanged behaviour, no overhead).
How it is enforced (provider-strict where possible, always validated on egress)
Enforcement is two-layered, and the egress layer is the hard guarantee:
- Native provider-strict (a belt, where the provider supports it). For a single-shot agent (no tools/web-search/MCP/knowledge/delegation) whose model is served by a provider confirmed to enforce it, the schema is sent natively so the provider constrains generation to it. Tool-using agents are excluded — constraining every leg to the schema would starve the model of tool calls. Two native shapes are supported today:
- OpenAI-style
response_format: json_schema— OpenAIgpt-4o/gpt-4o-mini. - Gemini
generationConfig.responseSchema(Googlegemini-2.5-flash/-2.5-pro/-2.0-flash/-1.5-flash/-1.5-pro, and the same models via Vertex). The JSON schema is translated to Gemini's OpenAPI-subset (types uppercased,enum→STRING+format:enum, a["T","null"]union →nullable). The translation is fail-closed and faithful: any construct it cannot represent exactly —oneOf/anyOf/allOf/$ref/const/patternProperties,additionalProperties:false, a property-less object or item-less array, a multi-type union — is not sent natively; that agent falls back to the prompt directive + egress validation below (never a lossy native schema). - Egress validation (the universal guarantee, every provider). An agent invoke with
a
response_schemaalways executes internally buffered — a caller'sstream:trueonly changes the delivery (the validated body is re-emitted as SSE, see Streamed interactive invoke below) — and every provider's answer is normalised to a single message, then the gateway extracts the JSON from that answer (tolerating a```jsonfence or a leading citation bracket), validates it against the schema, and replaces the delivered content with the canonical re-serialised JSON. This runs for all providers — including those with no native belt (Anthropic, Gemini, the default Myra/vLLM models) — so structured output is enforced uniformly, not best-effort.
Validation scope. Recursive: the top-level type; object required fields and
each declared property (recursively); array items (each element, recursively);
enum membership; and integer (a fractional number is rejected). An array is
distinguished from an object, so an object where an array is required (or vice-versa)
is rejected. Not enforced (validation only checks what is declared):
minItems/maxItems/minLength/pattern/format/oneOf/anyOf/$ref, tuple
items (an array of per-position schemas), and non-primitive enum members. In
particular additionalProperties is not applied — fields the model adds beyond the
schema are validated-around, not stripped; a consumer that needs a closed object
must not assume unknown keys are absent.
On failure — never silent prose. If the model's answer cannot be extracted,
parsed, or validated against the schema, the invoke fails closed with
structured_output_failed (HTTP 422) — the caller never receives non-conformant
output dressed as success. (The main inference is still billed, since it already ran.)
Schedule endpoints
An agent's schedules are managed through these routes. All require access to the
gateway; the mutating routes additionally require the AGENTS_AUTHOR permission
(the same gate as agent create/edit — the built-in member/ki_manager/tenant_admin/
admin roles, or a custom role granting it) and are ownership-checked against the
schedule's parent agent (a non-owned / missing schedule → 404).
| Method | Path | Purpose |
|---|---|---|
GET |
/admin/v1/gateways/{gw}/agents/{agent}/schedules |
List the agent's schedules as a JSON array ([] when none). Each row's output_action is redacted to { kind } — the webhook URL/token, Mattermost channel, and email recipient are never returned on this gateway-access route. |
GET |
/admin/v1/gateways/{gw}/agents/{agent}/schedules/{id} |
Fetch one schedule with its output_action returned whole (kind plus the to / url / channel_id), so the owner's edit form can prefill the current delivery target. This is why the read is owner-gated (AGENTS_AUTHOR + parent-agent ownership) unlike the redacted list; a non-owner gets 404 (no recipient/URL leak), a caller lacking AGENTS_AUTHOR (e.g. viewer/demouser) 403. |
POST |
/admin/v1/gateways/{gw}/agents/{agent}/schedules |
Create a schedule (201; see the input contract below). |
PATCH |
/admin/v1/gateways/{gw}/agents/{agent}/schedules/{id} |
Update a schedule (partial). Accepts output_action to change the delivery target; send { "kind": "none" } to clear an existing target (a JSON null is treated as "not provided" and leaves it unchanged). |
DELETE |
/admin/v1/gateways/{gw}/agents/{agent}/schedules/{id} |
Delete a schedule (200 { "ok": true }). |
POST |
/admin/v1/gateways/{gw}/agents/{agent}/schedules/{id}/run-now |
Mark the schedule due so the scheduler claims and runs it on its next tick (within about a minute), executing invoke, delivery, and the run record exactly as a normal run. A disabled schedule is refused with 409 — enable it first. |
The run-now route loads the schedule by id and gateway and checks ownership on the
schedule's parent agent (not the path {agent}), so it cannot be used to trigger
another agent's schedule.
Schedules — input contract for enabled
POST /admin/v1/gateways/{gw}/agents/{agent}/schedules and
PATCH …/schedules/{id} accept an optional enabled field:
- Accepted shape: a strict JSON boolean (
true/false), or JSONnull/ absent meaning "not provided". Anything else — a string, a number, an object — is rejected400 enabled must be a booleanbefore any storage access. (Same strictness asrequire_approvalbelow and the workflow-schedule route.) - On create,
enabledabsent ornulldefaults to enabled — a fresh schedule runs. (This differs from the member scheduled-tasks API, whereenabledis PATCH-only.) - Re-enabling re-arms the next run. Flipping a disabled schedule back to
enabled: truerecomputesnext_run_atexactly like a cadence change: an interval schedule becomes due promptly (within about a minute); a daily/weekly schedule arms to the next occurrence of its configured time. A stale run slot from before the pause is never fired — re-enabling can not trigger an immediate catch-up run. - Storage errors are not client errors. The only
400from the update's storage layer isno fields to update(an empty PATCH). A database fault answers an opaque500 { "error": "db" }; raw driver/SQL text never appears in a client-error (4xx) body — every 4xx on these routes is a fixed validator message. 404means genuinely absent, not "the lookup failed". OnPATCH,DELETE, and…/run-now, a404 schedule not foundis returned only when the row does not exist; a transient database fault on the lookup answers an opaque500 { "error": "db" }, never a404.POST(create) is the exception: because the row is already inserted, a failure to read it back for the response is not an error — the route still returns201with a minimal{ "id": … }body rather than a500(a500there would make a client retry and create a duplicate schedule).- Integer cadence fields are canonicalized.
interval_secand, for weekly agent schedules,daily_doware stored as canonical integers, so a syntactically-integer value never reaches the database in a form that would surface as a500; a value the shared cadence validator does not accept as an in-range integer is a400. - Cadence fields belong to exactly one kind. A schedule's cadence is
described by the field(s) its
schedule_kinduses —intervalusesinterval_sec,dailyusesdaily_at,weeklyusesdaily_atanddaily_dow— and nothing else. On this route that also meansdaily_dowis rejected outright: agent schedules areintervalordailyonly, so no kind here uses it. A cadence field the kind does not use is rejected400on create and onPATCH(it used to be accepted and then dropped), and the fields the new kind does not use are cleared when the kind changes, so a schedule never reports a cadence it does not run on. APATCHmay change the kind alone when the stored schedule already carries what the new kind needs; it is rejected400when the new kind needs a field the schedule's previous kind never used, rather than silently reusing a leftover value the request never named. The same rule applies when a queued four-eyes schedule-create is applied — a payload carrying a field its kind cannot use fails at apply with that reason recorded, instead of persisting it. -
output_actionis validated at the trust boundary. The optional delivery target is an object whosekindis one ofnone/email/webhook/mattermost(any other value →400). Foremail,tomust be a syntactically valid address within the length cap; forwebhook,urlmust be anhttps://URL within the length cap; formattermost,channel_idmust be a 26-character lower-case-alphanumeric id. A malformed value is rejected400before any storage write, on bothPOSTandPATCH— the client-side form check is a convenience, never the authority. Set{ "kind": "none" }onPATCHto clear an existing target; a JSONnullis read as "not provided" and leaves the stored target unchanged. The delivery destination's real egress controls (the fail-closed PII gate, SSRF/private-address refusal, and the per-tenant Mattermost-delivery flag — on by default, a platform admin may disable it per tenant) are enforced at delivery time, not here. A Mattermost delivery uses the tenant's own bot token and requires that bot to be a member of the target channel; if it cannot deliver (no per-tenant bot token, an undecryptable token, or the post is refused — e.g. the bot is not in the channel), the result is emailed to the schedule owner instead (never-fail), never silently dropped. -
A value that cannot be stored as JSON is rejected
400. JSON parsers accept the non-finite number literalsnan,Infinityand overflowing exponents such as1e999, but those values cannot be re-serialised, so they can never be persisted. A body such as{"output_action": {"kind": "webhook", "url": "https://x/y", "retries": 1e999}}now returns400 output_action is not JSON-encodablebefore any storage write. Previously the size guard was skipped on such a value and it reached the write, where the unencodable column silently shifted the remaining SQL parameters and the request answered200for an update that never happened. The same rule holds wherever the admin API stores a caller-supplied JSON object (gatewayconfig, routing-ruleconditions/actions, workflowgraph_json, governance-templateobligation_refs/target_controls, tokenscopes): the write is refused with an error naming the field, never partially applied.
Scheduled email delivery blocked by PII policy — the user is notified
A scheduled run's outbound delivery always passes a fail-closed PII gate: when the
run masked personal data (or the PII signal is missing), the result must not leave the
platform — the run records delivery_status: blocked_pii and no content egresses.
This holds on a local-only gateway of an enforcing tenant too: an agent-shaped run there is
masked under the PII mandate like on any other enforcing gateway (the first-party exemption is
for chat turns only).
That block stays. What changed: a blocked email delivery is no longer silent.
- The scheduler sends the recipient a notification email stating that the run completed and that delivery was blocked by the organization's PII policy. The notification never contains the run output or any personal data — only the agent and schedule name and where to log in. It is localized (English/German) by the recipient's user locale when the address belongs to a user of the tenant; otherwise English.
- Exactly one notification per blocked run — no retries, no notification loops. If
the notification itself fails, the failure is recorded on the run
(
delivery_error: "…; notification email failed: …") and nothing is retried; a successful notification records"…; recipient notified by email". - The run detail (Agents → Runs) shows the same honest status: the
Blocked (PII)badge plus a hint explaining that the block is the tenant's PII policy, not a delivery failure. - Webhook and Mattermost deliveries blocked by the PII gate keep their status-only behaviour (there is no address to notify).
Input contract (POST /v1/{tenant}/{gateway}/agents/{slug}/deliver-email, service
token only): the optional notice field requests this content-free notification and
accepts exactly one value, "blocked_pii" — any other value is rejected 400.
notice and text are mutually exclusive (400 if both are present); without
notice, a non-empty text remains required. All other deliver-email protections
(service-token-only auth, disabled/flagged-agent refusal, recipient re-derived from
the stored schedule, expected_to_hash recipient pin) apply unchanged in notice mode.
Scheduled run FAILED — the owner is notified
A scheduled run used to deliver only on success: if a run failed (a quota 429, a
provider error, a timeout), the configured destination received nothing and the
owner had no signal. Now a failed scheduled run whose schedule has a delivery
destination (output_action.kind of email / webhook / mattermost) sends a short
failure notice through that same channel.
- The notice is content-free operational metadata — the agent name, an actionable
reason, and a UTC timestamp. It never contains the prompt or any model output (a
failed run has none). The email subject reads
Agent run failed: <agent>; a webhook payload carriesevent: "run_failed"so an automation can tell an alert from a result. - The reason is the gateway's typed error message when it returns one, not a bare
status code. A run stopped by a budget cap — for example the shared
scheduler-service token reaching its monthly
budget_usd— reports which cap tripped and that it resets at the start of the next period, so the owner can act (raise the budget, or wait for the reset) instead of guessing at an opaquehttp 429. The same reason is stored onagent_run.errorand shown in the run history. Any bearer token in the body is redacted, and a response with no usable message degrades to the barehttp <status>. - At most one notice per schedule per 24 hours (anti-flap) — a schedule failing every
minute alerts once, not 1 440 times. This 24 h window is now shared with the owner
account-email channel (see below): both the destination notice and the owner email
fire at most once per window, off the one watermark. If the schedule uses the
delivery-approval gate (
require_approval), the failure notice is still sent: that gate reviews run output, and a content-free alert has none. A schedule with the no-egress profile sends no destination notice (but the owner account email — not an external egress — still fires). - If the notice's own delivery fails, it is not retried within the period and the failed run stays visible in the run history (Agents → Runs) — no notification loops.
Failing scheduled runs are surfaced, not silently disabled
Failure state was already persisted, but the only proactive owner channel was delivery to
the schedule's own external destination — skipped for output_action.kind: none, for member
prompt-tasks, and for scheduled workflows, and unreliable when the destination itself was the
fault. And after enough consecutive failures a schedule was silently auto-disabled, so an
owner who never opened the runs page never learned their automation had stopped. Both are fixed.
- Owner account email, destination-independent. On a failed scheduled run (agent schedule,
member prompt-task, or scheduled workflow) the schedule's creator is emailed at their
account address — resolved server-side from
created_by, never fromoutput_action.to. This fires even whenkind: none(no external destination). The email carries the schedule name, the same bounded, bearer-redacted typed reason stored on the run, and a UTC timestamp; the reason is HTML-escaped. It is localized (English/German) by the owner's locale. A schedule whose creator has no account email (an SSO/service identity) or is GDPR-erased (created_byNULL) gets no email — the in-app alert below still covers it. - A delivery-failed-on-a-successful-run is a distinct event. When a run succeeds but
delivering its output to the destination is an ongoing failure, the owner gets a
separate email that says results were produced but delivery failed. This covers a transport
failure, a destination that now resolves to a private IP (SSRF-blocked — a misconfig or
hijack the owner must fix), and a disabled tenant integration (e.g. Mattermost delivery
turned off). A PII-policy block (
blocked_pii, working as designed) and an empty result (nothing to deliver) are not alerted. Because a broken destination fails on every run, this email has its own 24 h window — separate from the run-failure window, so neither masks the other — bounding a schedule that succeeds every minute against a dead destination to one alert per window instead of up to 1 440. - In-app alert on
/me.GET/PATCH /admin/auth/menow returns afailing_runsarray — the caller's OWN failing scheduled runs — rendered in the always-on notifications center for every user (unlikebudget_alerts, which is admin/tenant_admin only). Each entry is{ kind: "agent"|"prompt"|"workflow", schedule_id, name, reason, failed_at?, delivery }(delivery: true= the delivery-failed-on-success state). The projection is owner-scoped in SQL (created_by= the caller); a tenant admin additionally sees rows whose owner is NULL (an orphaned/erased schedule that would otherwise be invisible) — never another member's rows. The array is always present ([]when none); a projection read error degrades to[](logged) and never fails/me. The center is non-dismissible — a broken automation cannot be swiped away. - No more silent auto-disable. A repeatedly-failing schedule stays enabled. Instead, an
interval schedule's next run is pushed later by a bounded exponential backoff
(
interval_sec × 2^min(cf, 6), capped at 24 h, only ever forward) so a fast-interval schedule cannot bill every minute forever; the backoff resets to the base cadence on the next success. daily/weekly schedules keep their exact chosen cadence (widening them would skip a chosen fire). Scheduled workflows whose recent runs all fail are likewise widened (interval) or left at exact cadence (daily/weekly) — never disabled. (An owner can still disable a schedule manually; that is unchanged, and a manually-disabled schedule produces nofailing_runsalert.)
Proactive token pre-expiry warning
The single most damaging silent failure is an auth/service token expiring unnoticed: a scheduler service token that lapsed took six scheduled agents down for days with no warning. To make that impossible, a daily sweep warns a token's owner before it expires:
- When. Once per day the scheduler triggers a server-side sweep of every
auth_tokenwhoseexpires_atfalls within the next 7 days and is not yet lapsed (a pre-expiry warning, not a post-mortem). Ephemeral playground tokens (30-minute UI probes) are excluded — they expire by design. The 7-day window matches the "expires soon" badge in the gateway's Auth Tokens table, so the two never disagree. - Who is warned — resolved server-side (never from a request). A user token warns its
owner (the
auth_token.user_idaccount). A service token (user_idNULL — a class which has no user owner) warns the tenant-admins of the token's gateway's tenant. The recipient set is derived entirely server-side from the token id inside the internal sweep route; a client can never redirect the warning to another address, and the admin fan-out is tenant-fenced in SQL — never cross-tenant (invariant 11). - Two channels. (1) An account email — for a service token it explicitly warns that
scheduled automations will stop when it lapses, and reminds the reader that a token issued with
no expiry date never lapses. (2) An in-app alert on
/me(see below). - At most once per token per window (dedup). Each warned token records the time it was warned
(
auth_token.last_expiry_warned_at, unix seconds); the sweep re-warns only when that watermark is unset or predates the current warning window (expires_at − 7 days). So repeated sweeps inside one window send nothing after the first, while extending a token's expiry (which opens a fresh window) correctly re-arms exactly one new warning. - Recipient-less tokens. A token whose owner was GDPR-erased, or a service token on a tenant
with no admins, sends no email — but the in-app
/mealert still surfaces it (invariant 4), and the watermark is still stamped so a daily no-op email is not re-attempted. - In-app alert on
/me.GET/PATCH /admin/auth/mereturns anexpiring_tokensarray — the caller's OWN soon-expiring tokens (a tenant admin additionally sees the tenant's service tokens) — rendered in the same always-on notifications center asfailing_runs, for every user. Each entry is{ token_id, gateway_id?, label, expires_at?, service }(service: true= a no-owner gateway service token, shown with the "automations will stop" wording). The projection is owner-scoped in SQL (auth_token.user_id= the caller; a tenant admin additionally seesuser_idNULL) and tenant-fenced. The array is always present ([]when none); a read error degrades to[](logged) and never fails/me. Likefailing_runs, it is non-dismissible. - Safe default at mint time. The token create/edit form defaults to no expiry (leave the field blank), which is the right choice for a service or scheduler token — a token that never lapses cannot die silently. The scheduler-token provisioning script mints non-expiring tokens and refuses to set an expiry.
Delivery approval — human-in-the-loop (HITL) gate
An unattended (scheduled) run whose result is delivered to an external destination (email / webhook / Mattermost) can require explicit human approval before it is sent. This is opt-in per schedule and defaults OFF — existing schedules keep auto-delivering exactly as before.
Opt-in field: require_approval
Create or update a schedule with require_approval: true (a strict JSON boolean — any
other type is rejected 400). When set, a succeeded run whose output_action.kind
is side-effecting (email / webhook / mattermost) does not auto-deliver: the run
executes and produces output, but delivery is held — the run's delivery_status
becomes await_approval and a pending approval appears in the inbox below. A schedule
with output_action.kind: none is unaffected (nothing is delivered, nothing is held).
If nobody decides within 24 hours the hold expires and is denied — the output is never delivered (fail-closed; an unapproved run never silently egresses).
The held output is not copied: it is referenced by the run and re-fetched at
delivery time. This requires payload logging to be enabled on the gateway and the
request_log retention to exceed the 24h approval window (both hold by default). If the
stored output cannot be re-fetched at delivery, the approved delivery fails closed to
skipped_empty (never a blank or wrong send).
List approvals (tenant governance)
Authz: tenant admin (admin / tenant_admin) — approving an egress is a tenant
governance action. Tenant-scoped: you only ever see your own tenant's approvals; a
cross-tenant id is never revealed. status (optional) must be one of
pending|approved|denied|expired|closed|all — any other value is rejected 400
(omitting it, like all, returns every status).
Each row carries the agent name/slug, the output_kind (e.g. email) and a short
action_summary (e.g. email / 1234 chars), the timestamps, and the run's real
delivery_status. It never returns the recipient or the answer.
Approve / deny
Authz: tenant admin, tenant-scoped. Both are POST-after-auth. The decision is a single tenant-scoped compare-and-swap:
- approve flips a
pendinghold toapproved; the scheduler delivers it on its next tick via the existing egress path (all three gates — empty-output, PII-active, SSRF / tenant-allowlist — re-run, and the destination is re-checked against what was approved; a schedule edited after approval fails closed toblocked_changed). Approving an expired hold is refused (409) — the TTL can never be bypassed. - deny flips it to
denied; it is never delivered.
Responses: 200 on success; 404 if the id is not in your tenant; 409 if it was
already decided (or, for approve, already expired). Each decision is recorded in the
control-plane audit log (agent_approval.approved / .denied) with the deciding
admin (best-effort write: a failed audit insert is logged server-side and never
blocks the decision).
Publish-time review governance (KI-Governance)
Complements the runtime HITL delivery gate: this governs the agent's own
publish lifecycle. Every agent carries a review_status — draft | in_review | approved
| flagged | disabled — surfaced on the agent GET (the decision provenance
reviewed_by/reviewed_at is owner-only; redacted for sharees). Only an approved
agent is reachable in the org catalog (and, for a review-required tenant, org-publish is
refused for a non-approved agent). This is regression-safe: every existing agent is
backfilled approved.
Per-tenant opt-in
PATCH /admin/v1/tenants/{id} accepts agent_review_required (0|1, default 0), and
GET /admin/v1/tenants/{id} (and the tenant list) return the current value so a client can
read the setting back. When 1:
new agents are created draft (must be approved before org-publish), and the org-publish
gate requires approved. Default 0 ⇒ new agents born approved, publish works as before.
The flag/disable kill-switch and the approved-gated catalog visibility apply to ALL
tenants (additive governance — they only take effect when someone explicitly acts).
Transitions
POST /admin/v1/gateways/{gw}/agents/{id}/review/submit (owner or admin)
POST /admin/v1/gateways/{gw}/agents/{id}/review/approve (ki_manager governance)
POST /admin/v1/gateways/{gw}/agents/{id}/review/flag (ki_manager governance)
POST /admin/v1/gateways/{gw}/agents/{id}/review/disable (ki_manager governance)
POST /admin/v1/gateways/{gw}/agents/{id}/review/reactivate (ki_manager governance)
Valid transitions (any other → 409 invalid_transition, no state change, no audit):
| decision | from | to | who |
|---|---|---|---|
| submit | draft, flagged | in_review | owner (require_author + owns the agent) |
| approve | in_review | approved | governance (require_gateway_access + require_ki_manager) |
| flag | in_review, approved | flagged | governance |
| disable | draft, in_review, approved, flagged | disabled | governance |
| reactivate | disabled | draft | governance |
Governance decisions require BOTH require_gateway_access (the tenant fence — a
ki_manager can only govern their own tenant's agents) AND the ki_manager role; they
stamp reviewed_by/reviewed_at and audit-log agent.review.<decision>. submit is an
owner action (no stamp). A flag/disable immediately drops the agent from the org
catalog and every non-owner org path; the owner keeps full access. Beyond the built-in
ki_manager/tenant_admin/admin roles, a custom tenant role granted the agent-review
capability can also make these governance decisions.
Re-check on update
In a review-required tenant, editing an already-approved agent (a PATCH
.../agents/{id} or a version restore that changes run-affecting config) invalidates
that approval: the agent is downgraded approved → in_review atomically in the same
transaction as the config write, so it immediately leaves the org catalog until a
ki_manager re-approves the new config. Only approved downgrades — a flagged/disabled
agent stays in its blocked state (an owner cannot un-flag/un-disable by editing), and
draft/in_review are already unapproved. Default-OFF tenants are unaffected. The reset
is visible in the agent.updated audit's review_status before/after.
Invoke-time enforcement (runtime teeth)
A flagged or disabled agent is not directly invokable — the data-plane
/agents/{slug}/invoke fails closed for every caller, including the agent's owner and
the scheduler service token, regardless of the tenant's review switch (a governance STOP
is absolute, not merely a catalog-visibility change). The rejection mirrors the
run-as-owner ACL: it returns AGENT_NOT_FOUND so a blocked slug cannot be probed.
draft/in_review/approved stay invokable (an owner may test a not-yet-approved draft;
approved is the normal live state).
Run history and audit trail (transparency)
These read-only endpoints reconstruct what an agent did, from storage alone — no live re-run.
List agent runs
GET /admin/v1/gateways/{gateway_id}/agents/{agent_id}/runs
Returns the recent runs of the agent as a JSON array. Required role: the agent's owner, or tenant_admin/admin (a non-owned or unknown agent returns 404 — the same gate as Read one agent run). The array is empty when the agent has never run.
Read one agent run
GET /admin/v1/gateways/{gateway_id}/agents/{agent_id}/runs/{run_id}
Returns one run fully reconstructed: { "run": { … }, "steps": [ … ] } — the durable run record plus its per-step trace (tool rounds, MCP discovery, and the termination reason) and the run's cost and token totals. The steps are metadata-only and carry no response content (the §12 transparency contract). Returns 404 when the run does not exist for that agent and gateway — and, so one owner cannot probe another's run ids, also 404 when the caller is a plain member who does not own the agent.
result — the run output, owner-only
A single additional field, result, carries the run's model output — the de-tokenized text the agent produced (the real body, exactly as the client received it, with any PII restored). It exists so a scheduled run whose external delivery was blocked by PII policy is no longer empty to the person who owns it.
result is returned to the run's owner only — the user who created the agent (agents run as their owner). This is enforced server-side: the field is joined from the request log and added to the response only when the caller's own user id equals the agent's creator. The distinction matters because the route itself is also reachable by governance roles:
- Owner (the agent's creator): receives
result. - Tenant administrator / administrator who is not the creator: can read the run (status, steps, usage) for governance, but the response never includes
result. - Any other member:
404(cannot read the run at all).
result is absent (the key is simply omitted, never a masked or empty placeholder) when there is no stored body — an aborted run with no request reference, a gateway with payload logging turned off, or a row whose response was PII-scrubbed. The masked projection (response_raw) is never returned by this endpoint.
Read the gateway audit trail
GET /admin/v1/gateways/{gateway_id}/audit
Returns the control-plane audit events for all agents and schedules of the gateway — who created, changed, enabled, disabled, or deleted what, with redacted before/after values. Required role: tenant admin or admin (the trail spans every agent, so it is governance-scoped, not per-owner). Query parameters limit (default 100) and offset (default 0) page the result.