Workflows API
A workflow is a versioned, gateway-scoped DAG of typed steps — Trigger,
Agent-Step, Condition, Human-Approval, Deliver — authored in the visual builder,
rendered as a vertical chain/tree and stored as a directed acyclic graph
(graph_json). A workflow is edited as a draft, then published to an
immutable, semver-tagged snapshot (workflow_version); a run is bound to the exact
version it started on, so a later publish never mutates a running or suspended run.
A manual run always executes the latest published version, never the unsaved
draft. So the caller can tell when the draft has diverged, GET /admin/v1/gateways/{gateway_id}/workflows/{id}
returns has_unpublished_changes (boolean): true when the draft graph_json differs
from the published version a run would execute — publish first to run the edits. It is
false when the draft is in sync or the workflow was never published.
Runs are executed by a dedicated runner process (scripts/workflow_runner.py)
that claims pending runs via a compare-and-swap lease and persists progress after
every step, so a resumed/replayed run in a fresh process resolves data-flow
references from durable state.
Admin endpoints require an authenticated admin session. Base URL:
https://ai-api-admin.myra.eu/admin/v1. The whole feature is gated per workspace by
tenant.workflows_enabled (default off, fail-closed): every admin workflow route
returns 403 { "error": "feature_disabled" } until an Admin enables it.
That workspace flag is flipped through its own route, which is platform-admin only
(role == "admin"; a tenant admin receives 403 — stricter than the
require_author gate the workflow routes themselves use):
POST /admin/v1/gateways/{gateway_id}/workflows-feature with { "enabled": <boolean> }
-> 200 { "enabled": <boolean> }. A non-boolean enabled is rejected 400.
Authorization matrix (server-side; the client is never the authz boundary)
| Action | Who |
|---|---|
| create / edit / publish / enable / disable / seed example / triggers + schedule config | any authoring user — member, ki_manager, tenant_admin, admin (require_author; viewer/demouser → 403). The builder is open to all authoring users; workflows are gateway-shared (per-user sharing is on the roadmap) |
read a workflow definition — list, detail (graph_json), versions, versions-diff |
any authoring user (require_author; viewer/demouser → 403). The definition is authoring IP (prompts, agent references, step wiring), gateway-shared among authors, not owner-scoped. These four reads were previously gateway-access-only (no role gate) — the one gap left when the rest of the surface was author-gated |
| trigger / run | any authoring user (require_author; viewer/demouser → 403) |
| view runs (list + detail) | any authoring user, owner-scoped — run detail exposes per-step outputs (run_state, possibly PII). A non-manager sees only their own runs (created_by); a manager (admin/tenant_admin/ki_manager) sees all. A non-owner requesting another user's run → 404 (never 403 — no run-existence oracle); a viewer/demouser → 403 |
enable / disable the feature for a tenant (workflows-feature) |
Platform admin only (role == "admin") |
| cancel a run | the run's owner, an Admin, or a Tenant-Admin |
| approve / reject a suspended run | only the node's named approver user (checked server-side); any authenticated user can be a named approver — the approver inbox is not KI-Manager-gated. Group approvers are a fast-follow. |
| referenced agents | the workflow owner (created_by) must own every agent an Agent-Step references (see below) |
Throughout this page, an "authoring" session / "any authoring user" (the
require_author gate) means a caller holding the WORKFLOWS_AUTHOR permission. For the
built-in roles that is member, ki_manager, tenant_admin, and admin; a tenant custom
role granting WORKFLOWS_AUTHOR is admitted regardless of its base role. A viewer/demouser
(or any role without the permission) is refused 403.
run-as-owner + agent ownership. A workflow runs as workflow.created_by; an
Agent-Step invokes the referenced agent over the gateway service-token path, so the
agent executes in the responsibility of the agent's owner. To prevent a
confused-deputy, an Agent-Step may reference only an agent the workflow owner owns
(ownership, not share-visibility). This is enforced at publish (a reference to a
non-owned agent is a publish error) and re-checked at run-claim (ownership can
change after publish → the run fails closed with agent_not_owned).
Workflow CRUD
GET /admin/v1/gateways/{gateway_id}/workflows → { workflows: [ … ] }
POST /admin/v1/gateways/{gateway_id}/workflows → create a draft.
Both routes require an authoring session (require_author — viewer/demouser → 403);
listing definitions is not open to read-only roles.
Each list row carries the workflow's identity/status/version/timestamps plus three server-derived list metrics:
steps— count of actionable step nodes in the draft graph (excludes the singletontriggerentry).agents— the distinct agent slugs the draft graph's Agent-Steps reference (a JSON array;[]when none).runs_30d— number of runs of this workflow in the last 30 days.
The list row does not include graph_json — the full definition is returned only by the
detail GET …/{id}. steps/agents derive from the draft graph, so a published
workflow whose draft has diverged shows its current edited structure.
Create body (accepted shape; anything else is rejected 400):
GET /admin/v1/gateways/{gateway_id}/workflows/{id} → the workflow + current draft graph.
PATCH /admin/v1/gateways/{gateway_id}/workflows/{id} → edit the draft (autosave).
DELETE /admin/v1/gateways/{gateway_id}/workflows/{id} → delete.
All three require an authoring session (require_author; viewer/demouser → 403) — the
detail response returns the full graph_json definition.
Draft edit is optimistic-concurrency guarded (D12): the body MUST carry the
updated_at the client last read. A mismatch (another tab saved first) returns
409 { "error": "stale", "updated_at": <current> }; the client reloads and retries.
On success the response echoes the freshly stamped token —
200 { "ok": true, "updated_at": <new> } — which is the value the next save must send
as its precondition. The client reads it straight from this response and needs no follow-up
GET. (A GET read-back was previously required to learn the new token; it was a second
failure surface that could misreport a persisted save and then self-409 the next save into
a false "another tab saved" state — removed.)
{
"name": "…",
"run_budget_micros": 5000000, // optional € budget (micros); null = no cap
"graph_json": { "steps": [ … ], "edges": [ … ] },
"updated_at": 1752000000 // REQUIRED precondition
}
POST /admin/v1/gateways/{gateway_id}/workflows/{id}/status { "status": "disabled" }
— enable/disable (published ⇄ disabled).
Seed the example workflows
POST /admin/v1/gateways/{gateway_id}/workflows/seed-example → create two publish-clean
example drafts owned by the caller and return their ids. This backs the empty-state
CTA ("Beispiel-Workflow erstellen") and the demo-tenant provisioning. Same authorization
as create: any authoring user (require_author) + the workspace feature flag.
Presseschau mit Freigabe— Trigger → Agent-Step → Human-Approval → Condition(true/else) → Deliver / Deliver (the full node vocabulary; the §demo flagship).PII-Testlauf— Trigger → Agent-Step → Deliver (its output carries test personal data so the Deliver egress trips the fail-closed PII gate →blocked_pii).
Both Agent-Steps reference an agent owned by the caller. When the caller owns none, a
per-user example agent (beispiel-agent-<uid8>, model left to the gateway default) is
auto-provisioned so the example publishes clean on a fresh tenant.
Request body — all fields optional (accepted shape; anything else is rejected 400):
{
"agent_slug": "presseschau-agent", // pin an OWNED agent (else auto-provisioned)
"approver_email": "redaktion@example.org" // a distinct in-tenant approver (else the caller)
}
Rejections (fail-closed): agent_slug not owned by the caller → 400; approver_email
not a live user in this workflow's tenant → 400; feature disabled → 403; not
an author (viewer/demouser) → 403.
Response 201:
Idempotent per the caller's own workflows by name: a repeat returns the existing ids (never duplicates); it never returns another user's same-named workflow. The examples are drafts — not auto-published — so the operator can adjust before publishing.
Demo-tenant provisioning (§10): seed, then set each Deliver channel_id to a real
channel and (for the two-person Freigabe beat) seed with the approver persona's
approver_email, then POST …/{id}/publish each. A blank channel_id publishes but does
not deliver — set it before the live run (and before the PII-Testlauf blocked_pii
beat, which needs a reached egress).
graph_json shape
{
"steps": [
{ "key": "trigger", "node_type": "trigger", "config": { "trigger_kind": "manual", "input_schema": { … } } },
{ "key": "classify", "node_type": "agent", "config": { "agent_slug": "klassifizierer",
"input": "Klassifiziere: {{trigger.output.text}}" } },
{ "key": "notify", "node_type": "deliver", "config": { "kind": "mattermost", "channel_id": "…",
"text": "Ergebnis: {{classify.output.summary}}" } }
],
"edges": [ { "from": "trigger", "to": "classify" }, { "from": "classify", "to": "notify" } ]
}
node_type ∈ trigger | agent | condition | loop | approval | deliver | fetch | code | connector. Edges may carry
a branch slot ("true"/"else" on a Condition, "body" on a Loop, "error" on an
Agent/Deliver node whose on_error is error_branch — see Error handling). key is a stable,
immutable step id; display labels rename freely. At publish the referenced agents'
response_schema + version are snapshotted into the version (drift guard).
Per-node config keys the runner consumes:
| Node | config keys |
|---|---|
| trigger | (none read at run time — the trigger's output is the run's trigger input; trigger_kind and input_schema are builder/validation metadata). A run's trigger_kind records how it started: manual | webhook | schedule | form | email (an inbound email — see Email-ingest triggers) |
| agent | agent_slug (an OWNED agent), input — the per-step task/prompt (a {{step.output.path}} template). input is required and must resolve to a non-empty string; an empty or missing input fails the step at run time (agent_failed). It is distinct from the agent's own saved system prompt. |
| condition | left_ref (a single {{…}} ref), op (eq\|ne\|contains\|gt\|lt\|is_empty), right_value (ignored for is_empty) |
| loop | items_ref (a single {{…}} ref to a JSON array, ≤ 10 items) |
| approval | approver_kind (user), approver_id, timeout_policy (skip\|end), timeout_hours (1…8760) |
| deliver | kind (mattermost\|webhook\|email), the per-kind target (channel_id | url | subject), text (a {{…}} template), and (email only) optional to — an EXTERNAL recipient, allowlist-gated (see below) |
| fetch | url (a literal or a {{…}} template) — the document to download; allowed_content_types (optional, a non-empty array of MIME strings; a malformed value is rejected at publish, never treated as "any"); max_bytes (optional positive int, clamped to the ~6 MiB artifact cap); filename (optional — else derived from the URL path basename). Emits an artifact reference {{fetch_key.output.artifact}} ({handle, filename, mime_type, size_bytes, sha256}) for a later node. A network fetch is a RETRY node (retry/on_error accepted) — a transient failure retries; a deterministic one takes the error_branch or fails the run. See the Fetch node section below. |
| connector | connector_id (an existing MCP/API connector of the tenant), tool (the tools/call tool name), input (a JSON object literal of arguments, or a single {{…}} ref that resolves to one). Calls the connector mid-run as the workflow owner (the owner's credential / private connector; a tenant-shared connector ignores the user). Emits the scrubbed string result {{connector_key.output.result}}. Side-effecting + ambiguous-outcome → NOT a retry node: retry/on_error/an error-edge are rejected at publish, and any failure is terminal (never auto-retried — see the Connector node section). |
| code | code (Python source, ≤ 64 KiB) — run in the network-isolated code-interpreter sandbox. The source is opaque: it is never template-resolved ({{…}} in the code is literal Python — an f-string f"{{x}}" or a set literal publishes clean), so run data reaches the program only as data, never spliced into source. Data channels: input_refs (optional array of ≤ 7 artifact-handle refs, typically {{prior_step.output.artifact.handle}}, staged as /input/<name>) and the always-present /input/_context.json (the run's trigger input as a JSON object — see the Code node section). output_filename (optional — selects the /output file by name and stores under it; else the sole output's own name, synthesized for a nameless image). Emits an artifact reference {{code_key.output.artifact}} for a later node. A RETRY node, but only a transient broker failure retries — a deterministic user-code failure does not. See the Code node section below. |
Publish
POST /admin/v1/gateways/{gateway_id}/workflows/{id}/publish →
201 { "version_id": "…", "semver": "1.1.0", "updated_at": 1742551200 }. The
updated_at is the optimistic-concurrency token the next draft PATCH must echo.
Publish runs graph validation (below) and, on success, writes an immutable
workflow_version and auto minor-bumps the semver. On validation failure the response
is 422:
{ "error": "graph_invalid",
"errors": [ { "step_key": "notify", "code": "template_unresolved" } ],
"warnings": [] }
Errors block publish (rendered as per-node red badges); warnings do not. (The validation model is block-only today — references that could be unavailable at runtime are rejected at publish, never downgraded to warnings.)
Validation rules (also re-run at run-claim as a drift guard): the graph must be a
DAG (a cycle is rejected); every edge must reference existing steps with
type-compatible ports; every {{step.output.path}} reference must resolve to an
upstream step's declared output schema (a dangling reference is
template_unresolved); a reference into a Condition/Loop branch member from outside
that branch is a hard publish block (dangling_ref — only ancestors on the same
chain resolve, so a maybe-skipped value can never be referenced); every referenced agent must be owned by the workflow owner (agent_not_owned);
malformed / oversize graph_json is rejected fail-closed. A {{}} reference to an
output-less node (condition / deliver / approval — they route, egress, or gate
but expose no output) is a publish block (ref_to_output_less); only trigger / agent
/ loop produce a referenceable value. An approval node must name a valid approver
(approver_kind: "user" + an approver_id that is a live user in the workflow's tenant,
else approver_missing / approver_not_in_tenant) and a valid timeout
(timeout_policy ∈ {skip, end}, timeout_hours a finite number in [1, 8760], else
bad_timeout).
GET /admin/v1/gateways/{gateway_id}/workflows/{id}/versions → the published versions
(newest first, capped at the newest 100), each enriched with run_count (how many runs
executed on that version — the run→version pin is permanent), is_current
(semver == workflow.version), and published_by_name (the publisher's email, or the
raw id if the user was deleted). Requires an authoring session (require_author;
viewer/demouser → 403), then gateway access + the workspace feature flag. It exposes
historical version config — the same config class as the current draft, never run-output —
so it is author-gated exactly like the definition reads.
Diff two versions
GET /admin/v1/gateways/{gateway_id}/workflows/{id}/versions/diff?from={version_id}&to={version_id}
-> 200 { "diff": { … }, "from": …, "to": … }. The diff is computed server-side from
the two immutable graph snapshots; the response lists, per step (keyed by the immutable
step key): added, removed, changed, unchanged, plus edge edges_added /
edges_removed. Each changed step carries a fields list of
{ field, kind, before, after } where kind ∈ {added_field, removed_field,
changed_field} (the absent side is JSON null). Every list is always a JSON array.
Requires an authoring session (require_author; viewer/demouser → 403) — same gate
as the other definition reads.
Accepted / rejected: both from and to are required and non-blank (else 400);
each must be a version of this workflow — a version id belonging to another workflow
or tenant, or a nonexistent id, is 404 (the workflow itself is gateway/tenant-fenced
first). A stored snapshot that fails to parse is 422 { "error": "diff_failed" } (fail
closed, never a 500). Diffing a version against itself returns an all-unchanged diff.
Roll back to a version
POST /admin/v1/gateways/{gateway_id}/workflows/{id}/versions/{version_id}/rollback
-> 201 { "version_id": …, "semver": …, "updated_at": …, "rolled_back_from": … }.
Rollback republishes the chosen version's graph as a NEW version — the history is
append-only, old versions are never mutated, and runs stay pinned to the version that
executed them. It runs the same validation as a normal publish, so a rollback to a
version that now references a deleted agent (agent_not_owned) or a
cross-tenant/deleted approver (approver_not_in_tenant) is rejected 422. The
workflow's current enabled/disabled status is preserved — rolling back a disabled
workflow does not silently re-enable it. Rolling back to the already-current version is
409 { "error": "already_current" }. The editable draft is not touched; rollback
changes what is published (future runs use the rolled-back graph), not your draft.
Requires an authoring session (require_author) — the publish authorization.
Test a single step (builder, no persisted run)
The builder's agent-step config panel offers "Diesen Schritt testen" — run just that
one agent step in isolation, without executing the whole workflow. It reuses the agent
invoke path (agents/{slug}/invoke via a short-lived playground token, run as the
agent owner), so it does not create a workflow_run and does not appear in the run
history. Because a mid-chain step has no upstream output at design time, the panel lists
each {{step.output.path}} the step references and lets the author supply a sample
value; those are substituted into the step input before the invoke. Fail-closed: a
deleted/unavailable agent, an unresolved (dangling) reference, or an empty resolved input
blocks the test with no invoke. The test-run cost is surfaced from the synchronous
X-AIG-Cost-Micros response header and labelled as a test. Scope: agent steps only
(Trigger / Condition / Loop / Approval / Deliver are not testable in isolation).
NL copilot — generate / fix / explain a draft graph
A natural-language assistant for the builder. Three modes: generate a draft graph from a process description, fix a graph from its validation errors, explain a graph in plain language. It never publishes — a human reviews and publishes, so the governance story is intact. The proposed graph is untrusted model output and is validated through the FULL publish validator before it may be applied (and again at publish).
The copilot is split into two calls, by concern:
1. Inference (metered) — /v1 pseudo-provider
POST /v1/{tenant}/{gateway}/workflows/copilot (inference host; Bearer a short-lived
playground token, exactly like "Test a single step"). This rides the normal inference
pipeline, so the run draws the tenant's normal metering (spend_ledger), PII
scrubbing, EU-residency and model-allowlist clamps, and returns the cost in
X-AIG-Cost-Micros. It is a passthrough (buffered, stream:false); for
generate/fix the answer is constrained to the graph JSON schema and validated
fail-closed (422 structured_output_failed on non-conforming output, billed first so a
caller can't harvest free runs).
Authorization: the token must be user-bound and hold the WORKFLOWS_AUTHOR
permission — the SAME permission the admin build/publish routes require, so the copilot
admits exactly the users who may author a workflow, never a stale subset. For the built-in
roles that is member, ki_manager, tenant_admin, and admin; a tenant custom role
that grants WORKFLOWS_AUTHOR is admitted too, regardless of its base role. An unbound
(service) token or a caller lacking WORKFLOWS_AUTHOR (e.g. viewer/demouser) is
refused 403 {error.code: "forbidden"}. The
feature must also be enabled for the workspace, else 403 {error.code: "workflows_disabled"}
— a distinct code from the role refusal, so the client can prompt the user to ask an admin
to enable Workflows rather than showing a generic permission error.
Accepted body (anything else is rejected 400 — validated at the trust boundary,
Invariant 11):
{ "mode": "generate|fix|explain",
"model": "<a runnable model id>",
"description": "describe the process (required for generate, optional for fix)",
"graph": { "steps": [ … ], "edges": [ … ] },
"validation_errors": [ { "step_key": "…", "code": "…" } ] }
mode— must be one of the three; missing/unknown →400.model— required, non-empty; absent/empty →400(no model call). The pipeline clamps it to the gateway's allowlist/residency.description— required forgenerate; capped at 4 KB; larger →400.graph— required forfixandexplain; the wire shape (see graph_json shape); capped at 64 KB for the prompt; malformed/oversized →400.validation_errors— optional, only used to phrase thefixprompt.
For generate/fix the response's assistant content is the proposed graph as JSON; for
explain it is plain-language text. To keep proposals publishable, the prompt is
grounded in the caller's owned agent slugs (agent nodes must reference one) and the
caller as the default approver.
2. Server-side validation + diff (before applying)
POST /admin/v1/gateways/{gateway_id}/workflows/{id}/copilot-validate (admin host;
require_author; feature-gated). Runs the same "publishable" validator as publish
(core.wf_validate + owner-agent fence + approver-in-tenant fence) on the untrusted
proposed graph, and computes the changed-nodes diff vs the current draft.
Accepted body:
graph— required object; not-an-object / not-encodable / >512 KB →400.base_graph— optional; defaults to an empty graph{"steps":[],"edges":[]}when absent (a fresh generate), so the diff never sees an absent side.
Response 200:
{ "ok": true,
"errors": [],
"warnings": [],
"diff": { "added": [ … ], "removed": [ … ], "changed": [ … ], "unchanged": [ … ],
"edges_added": [ … ], "edges_removed": [ … ] } }
ok:false returns the validation errors (step_key + code) and the client surfaces them
without applying — a graph that fails validation is never applied as a partial graph.
errors reflects wf_validate's phased short-circuit (up to the first failing phase, not
exhaustive), so the fix UX re-runs validate → fix → validate. A DB failure during
validation returns 503 (never a fake validation error). Applying an ok proposal is a
normal draft PATCH (the client is never the boundary — publish re-validates regardless).
Trigger a run (manual)
POST /admin/v1/gateways/{gateway_id}/workflows/{id}/runs →
202 { "run_id": "…", "status": "pending" }.
The workflow must be published — an unpublished workflow returns 409 { "error":
"workflow not published" }, and one whose versions have gone returns 409 { "error":
"no published version" }. The run starts as pending; the runner claims it within
~2 s. A manual run consumes only idempotency_key; it does not accept a trigger
input payload (the run begins from the Trigger node with no supplied input, and no
input_schema validation is applied on this route — an input field in the body is
ignored). To pass a payload, fire the workflow through its webhook or email trigger.
Idempotency (fail-safe against double-clicks/retries): if idempotency_key is
supplied, a second POST with the same key for this workflow returns the same
run_id (and "was_existing": true) — never a second run, never double spend or
double egress. Omitting the key (or sending null / "") means every POST is a
distinct run.
idempotency_key, when present, must be an opaque token: a string of at most 64
characters drawn from A-Z a-z 0-9 - _ . :. Anything else — a longer key, other
characters, or a non-string value — is rejected with 400 { "error": "invalid
idempotency key" }. This is the same shape the webhook fire path enforces on
X-AIG-Idempotency-Key.
Schedule trigger (time triggers)
A workflow can fire on a schedule — a fixed interval, once a day at a set time
(UTC), or once a week on a chosen weekday at a set time (UTC) — in addition to
manual / webhook triggers. There is at most one schedule
per workflow. It reuses the existing agent scheduler (no second scheduler): the
schedule is a polymorphic agent_schedule row with workflow_id set; the scheduler's
per-minute tick claims a due row via the same at-most-once compare-and-swap and, for a
workflow row, asks the gateway to enqueue a pending workflow_run (trigger_kind =
schedule, run-as-owner) instead of invoking an agent. The runner then executes it
exactly like a manual/webhook run.
Saving a daily or weekly schedule arms its next run to the next occurrence of the chosen time (for weekly: on the chosen weekday), not to the moment of saving — so "daily at 10:00" created at 08:36 first fires at 10:00, never a spurious extra run right after save. An interval schedule starts promptly (its first run is due immediately). Editing only metadata (e.g. renaming, or switching a still-enabled schedule off) preserves the already-armed run time; the next run is recomputed only when the cadence itself changes (including changing only the weekday of a weekly schedule) or a disabled schedule is switched back on.
Authz: any authoring user (require_author; setting a schedule is a create/edit function), plus
the per-workspace feature gate. All three routes are gateway-scoped.
| Method / path | Behaviour |
|---|---|
GET …/workflows/{id}/schedule |
200 { "schedule": false } when none is set, else { "schedule": { enabled, schedule_kind, interval_sec, daily_at, daily_dow, next_run_at, last_run_at, consecutive_failures, last_status, last_error } }. daily_dow is the weekly schedule's weekday, integer 0..6 (0 = Sunday … 6 = Saturday, UTC); absent/null for interval/daily rows. The three health fields let the UI show an auto-disabled (enabled: 0) or failing schedule instead of rendering it as healthy. last_status is ok / retrying / failed / paused / null; for a workflow schedule the reachable set is paused (the backstop switched it off) or null, which is exactly what lets the UI distinguish a system stop from a human switching the schedule off. last_error is fixed operator-facing English, never model output or PII: either a dispatch string such as no gateway service token, or the auto-disable sentence, which embeds the failing run's error_code accepted only as [A-Za-z0-9_.-]{1,48} — anything else (empty, over-length, markup, quotes, newlines, multibyte) is rejected and rendered as unknown. consecutive_failures is the streak toward the agent auto-pause threshold; on a workflow schedule the scheduler zeroes it on every fire, so it stays 0 and the count is carried in last_error instead. |
PUT …/workflows/{id}/schedule |
Create-or-replace (race-safe upsert). Returns the stored cadence. Disabling cancels the workflow's still-pending schedule runs. |
DELETE …/workflows/{id}/schedule |
Remove the schedule + cancel its pending schedule runs → 200 { "ok": true }. |
PUT body — accepted shape (trust boundary, Invariant 11; validated by
core.schedule_cadence, the one shared cadence validator):
{ "schedule_kind": "interval", "interval_sec": 3600, "enabled": true }
{ "schedule_kind": "daily", "daily_at": "06:00", "enabled": true }
{ "schedule_kind": "weekly", "daily_at": "06:00", "daily_dow": 3, "enabled": true }
schedule_kind—interval|daily|weekly(required; anything else →400).interval_sec— integer 60 … 366 days (required forinterval). Below 60, above the cap, non-integer, or non-numeric →400, no write.daily_at—HH:MMin UTC, 24-hour (required fordailyandweekly).24:00,12:60, a single-digit hour, an empty string, or a non-string →400, no write.daily_dow— integer 0 … 6 (0 = Sunday … 6 = Saturday, UTC; required forweekly— same contract as scheduled tasks). Out of range, fractional, non-numeric, boolean, or JSONnullwithweekly→400, no write. Forinterval/dailythe server storesNULLregardless of what was sent (the client never picks what persists).enabled— strict JSON boolean (a governance toggle is never coerced from a truthy string). Absent → enabled. Any other type →400.- All cadence fields are range/format-checked whenever present, even those the
kinddoes not use (a garbage value in an unused field is rejected here rather than reaching the DB).
Fail-closed fire path (never a surprise run or double spend): the scheduler enqueues
only when, at fire time, the workflow is published, its workspace feature is on, it has
a published version, and no run is already in flight (any manual/webhook/schedule run
in pending/running/suspended → the fire is skipped, at-most-once), the schedule still
exists, and the schedule is still enabled (a fire is refused for a schedule switched off
after the tick claimed it — this narrows, but cannot fully close, that race: a PUT landing
between the check and the insert still yields one run). All identity (tenant / gateway /
owner / version) is derived server-side from the workflow row — the scheduler is
authenticated but never trusted for identity.
Auto-disable backstop, and how to recover from it. A schedule whose last 5 schedule runs
since it was last enabled all failed is automatically switched off, to bound a broken
schedule's spend. The window starts when the schedule was last saved with enabled: true,
which is what makes the stop recoverable: fix the cause, switch the schedule back on and
save — that PUT re-arms the failure window, and the next fire produces a real run. Five
fresh failures after that re-arm switch it off again, so the backstop keeps its teeth. The
stop records why, so it never looks like a human turning the schedule off:
last_status: "paused" plus a last_error naming the failure. Both fields are cleared by the
same re-arming save, so a schedule an operator later switches off by hand does not inherit an
old system cause.
Two operational notes:
- Never re-enable a stopped schedule with raw SQL.
UPDATE agent_schedule SET enabled=1does not move the failure-window watermark, so the next tick switches it straight back off. Use thePUTabove. - A run that ended
cancelledinside the window breaks the all-failed streak and defers the backstop until five newer failures accrue — disabling or deleting a schedule cancels its pending schedule runs, so this is the expected behaviour after a stop/start cycle.
Timezone: daily_at is UTC in the MVP (cron expressions and per-workflow timezones
are a later roadmap item).
Runs
Both read routes require an authoring session (require_author — viewer/demouser
get 403) plus the workspace feature flag, and are owner-scoped: run detail exposes
per-step model outputs (run_state), which may carry personal data, so a non-manager
may read only their OWN runs. Scoping is by workflow_run.created_by — the triggerer for a
manual run, the workflow owner for a schedule/webhook/form run. The runs list returns
only the caller's own runs; run detail for a run the caller does not own returns 404
(never 403 — a non-owner must not learn the run id exists). An admin, tenant_admin, or
ki_manager sees every run as the audit view. The client is never the authz boundary.
GET /admin/v1/gateways/{gateway_id}/workflows/{id}/runs → { "runs": [...] } with
status, trigger_kind, cost_micros, tokens_used, pii_active, suspended_reason,
approver_ref, started_at, finished_at, error_code. status ∈ pending | running |
succeeded | failed | suspended | cancelled.
GET /admin/v1/gateways/{gateway_id}/workflows/runs/{run_id} → { "run": {...},
"steps": [...] } (404 if the run is not on this gateway). run adds the technical
error string and run_state — a JSON string (a step_key → output object,
bounded D15; absent on a run that has produced no output yet — NULL columns are
dropped, so clients must treat it as optional). Per-step steps rows carry
step_key, node_type, node_label, status, attempt, request_ref, cost_micros, tokens, error,
delivery_status, delivery_error, started_at, finished_at. node_label is the
step's friendly builder name — the run's pinned version-graph node label, resolved on read;
it is absent when the node has no label (or the version cannot be resolved), and the UI then
falls back to the translated node-type name (never the raw step_key). Step status ∈ pending |
running | succeeded | failed | skipped | delivering | continued. continued marks a
step whose final failure was absorbed by its on_error policy (continue or
error_branch) — the run proceeded; the row keeps the underlying error (and for
Deliver steps the delivery_status, e.g. an absorbed blocked_pii). When a step is
retried, EVERY attempt persists its own row with the same step_key and an
incrementing attempt (1-based) — each attempt row carries that attempt's
cost_micros/tokens, so retry cost is attributed per attempt; the step's terminal
disposition is the highest-attempt row. A Deliver step is written
status: "delivering" before egress, then flipped to a terminal succeeded/failed (or continued when the failure is absorbed by the step's on_error policy)
once the outcome is known; its detailed outcome is in delivery_status, so
render Deliver rows from that field. A row still showing delivering means the egress
outcome is unknown (a crash mid-send) — such a run fails closed on resume and is never
re-delivered. pii_active is an integer (1/0), written by the runner at the
run's terminal/suspend boundary. 1 means the run handled PII fail-closed:
it is set when any agent step did not explicitly report x-aig-pii-active: 0 — which
includes agent errors and timeouts ("PII could not be ruled out this run"), matching
the Deliver egress gate exactly; it does not mean "PII was provably masked". The flag is
monotonic within a run (latched via GREATEST — once 1, a later terminal write can
never lower it to 0, so a resume/late-fail never loses the signal). Known limitation: a
run reaped mid-execution before any terminal/suspend write reads 0 (the transient signal
died with the crashed tick; no egress occurred). The approver provenance (approver_ref,
rendered "freigegeben von X am Y") and the typed error_code (see Error codes) are on
run.
Run-detail display. The runs UI derives the shown status from
error_code, not the raw status: a deliver_blocked_pii run reads "Blockiert —
Datenschutz" and a budget_exceeded run reads "Gestoppt — Budget erreicht", both
in a protective amber tone — these guardrail terminal states are the platform working
as designed, so alarm-red is reserved for a genuine failed. Because the runner stores
the trigger node's output as the raw, unmasked trigger input, the run-detail view
never renders that step's output in cleartext — it shows a redaction notice instead
(agent-step outputs already arrive masked and stay visible). The display never mutates
the stored status/error_code; it is a read-time projection.
POST /admin/v1/gateways/{gateway_id}/workflows/runs/{run_id}/cancel →
200 { "cancelled": true }; 409 (already terminal — the CAS loser); 403 (not the
run's owner, Admin, or Tenant-Admin); 404 (unknown/other-gateway run). The runner observes the cancel at
its next step boundary and aborts before any further egress; every runner write is
lease-fenced, so a step that completes after the cancel cannot write a phantom row.
Cost & budget (D10/D14). Per-step cost_micros (€ micros) is taken from the
Agent-Step invoke's synchronous X-AIG-Cost-Micros response header (server-side
pricing; the runner performs no pricing math). The run's cost_micros is the sum;
request_log is the independent drill-down. If the run exceeds the workflow's
run_budget_micros, the run aborts with error_code: "budget_exceeded" ("Budget
erreicht").
Approval: the "Meine Freigaben" inbox
When a run reaches a Human-Approval node it is suspended (status: suspended,
suspended_reason: awaiting_approval), its lease cleared (so the lease-reaper — which
also excludes suspended runs — can never reap it), and the node's approver + timeout
policy + the suspended step key are snapshotted onto the run. No resume token is minted
(the token-link / notification path is a roadmap fast-follow); the approver decides in
their in-app inbox.
The inbox routes are reachable by any authenticated user (a named approver need not be a KI-Manager) and are fenced server-side by the caller's tenant and identity — you see and decide only the runs where you are the named approver.
GET /admin/v1/me/approvals → { "approvals": [ { id, workflow_id, workflow_name,
approval_step_key, approval_node_label, approval_deadline, cost_micros, created_at,
requested_by_email, requested_by_name } ] } — my pending approvals (an empty inbox is
{ "approvals": [] }). requested_by_email / requested_by_name name the run-as-owner
(workflow.created_by); both are always strings — "" when the owner is unknown or a deleted
user (the UI shows "—"). approval_node_label is the approval step's friendly
builder name (its version-graph node label, resolved on read); absent when the node has no
label, and the UI then falls back to the translated "approval" type name — never the raw
approval_step_key.
GET /admin/v1/me/approvals/{run_id} → { "approval": { … approval_node_label, run_state,
requested_by_email, requested_by_name … }, "steps": [ … ] } — the confirm-page detail. Its
approval carries the same approval_node_label as the inbox, and each steps row carries the
per-step node_label (see Runs) so the confirm page names every produced output by its
friendly step name. It is
identity-scoped: a caller who is not the named approver (or the run is not in their
tenant / no longer suspended) gets 404 — run_state (which may carry PII) never leaks to
a non-approver.
The decision is a POST after authentication (never a GET with a side-effect):
POST /admin/v1/me/approvals/{run_id}/decision
decision ∈ {approve, reject} (any other value → 400). approve resumes the run
(status: suspended → pending; the runner then carries it past the approval node and
continues); reject fails it closed (status: failed, error_code:
approval_rejected).
Reject requires an audit reason (public-sector audit trail). Accepted shape: a
string of 1..1000 characters (Unicode codepoints) after trimming surrounding
whitespace. It is validated at the trust boundary — after the 403/404 authz gate,
before the state change — and anything else is rejected fail-closed with the run left
suspended (never a partial reject): a missing/non-string or empty/whitespace-only
reason → 400 reason required; invalid UTF-8 → 400 reason invalid; an over-1000-char
reason → 400 reason too long.
On approve, reason is ignored. The reason is stored (parameterized) in the run's
error field alongside error_code: approval_rejected, so the run owner sees WHY it was
rejected in the run detail.
Both approve and reject are a single atomic compare-and-swap gated on
status = 'suspended' (and the caller's tenant_id + approver_id), so a second decision
— or a timeout that fires concurrently — loses and receives 409 { "error":
"already_decided" }. A caller who is not the named approver gets 403; a run absent from
their tenant, 404. The approver provenance (approver_ref, the approver's email) is
recorded on the run and rendered "freigegeben von X".
Approval timeout policy (skip | end, configured on the node; default end) is enforced
by an expired-approval scan in the runner: past the deadline the run either auto-proceeds
(skip — a Freigabe bypass when the approver is gone) or fails closed (end), per
policy — never hangs indefinitely. The timeout apply and a human decision race on the same
status = 'suspended' CAS, so exactly one wins.
Deliver node — egress is fail-closed
A Deliver node reuses the existing run-delivery egress controls (one mechanism). Its
outcome is the full delivery_status enum: delivered | failed | skipped_empty |
blocked_pii | blocked_ssrf | blocked_tenant. PII gate (the §1 claim): if any
Agent-Step in the run masked PII (X-AIG-PII-Active: 1, or an absent header — treated
as active, fail-closed), external egress is refused and the Deliver step shows
blocked_pii. A Deliver that does not reach delivered hard-fails the run
(error_code: deliver_blocked_pii etc.); skipped_empty (nothing to send) is a
success. The delivering step row is persisted before egress, so a replay refuses
to re-fire a non-idempotent delivery with an unknown outcome (no doubled egress).
Email attachment
An email-kind Deliver node may attach a workflow_artifact (e.g. the filled PDF a
code node produced) via one optional config key:
attachment_ref— a single{{step.output.artifact.handle}}reference to a prior node's artifact handle. It must be a single template ref; a literal handle is rejected at publish (deliver_bad_attachment_ref) because a handle is a run-scoped value minted at execution time. Absent → a plain email (unchanged behaviour).
At delivery the runner resolves the ref to a handle and the loopback /deliver-email
endpoint loads that artifact scoped to the current run (a node can only attach an
artifact of its own run) and composes a multipart/mixed message. Accepted / rejected
(fail-closed, Inv. 11):
- the handle is validated (present-but-empty / malformed →
400 bad_handle); an unknown or foreign handle →404 attachment_not_found; a corrupt / sha-mismatched artifact →422 attachment_corrupt; an over-cap artifact (the ~6 MiB artifact-store cap) →413 attachment_too_large. Any load failure hard-fails the send — a requested attachment is never silently dropped. - the attachment filename and content-type are sanitized before they enter the
MIME headers (CR/LF/control/quote/backslash stripped) so a crafted name can never
inject a header; a NULL content-type falls back to
application/octet-stream. - a
4xxfrom the endpoint is a deterministic terminal outcome (failed_input, not retried);5xx/transport stays retryable. - PII: on the PII-internal delivery path (a masked run delivering to the internal
owner) the unmaskable binary attachment is omitted; the delivered row records
attachment_omitted_pii(observable, never a silent drop). Only the message body (subject to the existing PII gate) is delivered there.
External recipient via the per-tenant allowlist
By default a deliver-email node sends to the workflow owner (resolved server-side; never
from the request). An email deliver node may instead target an external recipient via one
optional config key:
to— the external recipient: a single{{…}}template ref (resolved at runtime, e.g.{{trigger.output.recipient}}) OR a literal address. Malformed at publish →deliver_bad_recipient. Absent → the owner recipient (unchanged behaviour).
The external send is permitted ONLY if the recipient is on the tenant's
workflow_recipient_allowlist — a tenant-admin-curated JSON array of allowlist entries, each
an exact address (vipwing@munich-airport.de) or a domain (@munich-airport.de). It is
set on the tenant via the admin tenant PATCH (a member cannot curate it), DB-backed config. Domain entries match by exact domain equality — @munich-airport.de matches
x@munich-airport.de only, never a subdomain (x@evil.munich-airport.de) or suffix
(x@munich-airport.de.evil.com). IDN/non-ASCII addresses are not allowlistable (fail-closed).
Fences at the loopback /deliver-email (fail-closed, Inv. 11):
- a present-but-malformed
to→400 bad_recipient; - an external
toon a PII-masked run →403 recipient_pii_blocked, no send — an unmaskable body/attachment must never egress to a non-tenant address (the anti-exfiltration control is preserved; only the internal-owner relaxation applies on masked runs); - no allowlist configured, an unreadable/empty allowlist, or a recipient not on it →
403 recipient_not_allowed, no send (never a silent fall-back to the owner); - a matched external send is audited (
workflow.deliver_external, with the recipient + matched entry). The subject stays server-built; a notice (blocked-PII) always goes to the owner.
An invalid allowlist entry at the admin write door → 400 recipient_allowlist_error.
Condition node — one comparator, an always-present else (D6)
A Condition routes the chain down one of two indented sub-chains. Its config is
{ "left_ref": "{{step.output.path}}", "op": "eq|ne|contains|gt|lt|is_empty",
"right_value": "…" }. In the graph a Condition has three out-edges: a
branch:"true" head, a branch:"else" head (the else branch is always present),
and an optional plain "after" edge where the main chain resumes. Branch nesting is
depth 1 (a branch may not contain another Condition/Loop), and the branch sub-chains
never rejoin — no fan-in.
At run time the runner resolves left_ref and applies the comparator: eq/ne/
contains compare the value's string form against right_value; gt/lt are numeric
(a non-numeric value or operand is a hard fail, condition_type_error); is_empty is
true for "", [], or {} (the right_value is ignored). The taken branch executes;
every step in the not-taken branch is marked skipped (no execution, no egress).
A cjson null or a missing value on the compared path is a documented hard
fail (template_null_path / template_unresolved) — never a silent else. A
Condition produces no referenceable output ({{condition.output.…}} is a publish
block). Replay-safe: the not-taken skipped rows are persisted before the Condition's
own row, so a reaped run never re-runs a not-taken branch.
Loop-lite node — sequential, bounded, collect-all (D7)
A Loop iterates a JSON array. Its config is { "items_ref": "{{step.output.array}}" }
and it has a branch:"body" sub-chain head plus a plain "after" edge. The array is
resolved once and bounded before iteration one: a non-array is loop_not_array and
more than 10 items is loop_too_many (rejected, never clamped). For each item (0-
based) the body runs with the current item bound as {{loop.output.item}} and
{{loop.output.index}}; each body step is persisted as an iteration row keyed
step_key#i. The loop is collect-all: after the last item, loop.output is
{ "items": [ …one entry per iteration… ], "count": N }, referenceable by the after-
chain. First failure fails the loop (fixed policy — the per-node on_error key
described under Error handling below is not available inside loop bodies, and a
body step always uses the error policy; per-step retry IS available in bodies).
A loop body must be side-effect-free: deliver and approval nodes are rejected in
a loop body at publish AND run-claim (loop_body_side_effect) — because a mid-loop crash
restarts the whole loop, which is only safe when the body just transforms data (a re-run
re-invokes body agents, re-charging tokens, but delivers nothing twice — that egress class is designed out). Loop body outputs are transient (not persisted to
run_state); only the collected {items,count} is.
Structural validation (blocks publish, re-checked at run-claim): a graph must have
exactly one Trigger with no in-edges and every other node exactly one in-edge (no fan-
in); at most one plain out-edge and one edge per branch slot per node; branch slots must
match the container (condition → true/else, loop → body); branch depth ≤ 1;
every node reachable from the Trigger; step keys [A-Za-z_][A-Za-z0-9_]*, ≤ 60 chars
(so a #i iteration suffix fits). A {{}} ref to an output-less node (Condition,
Deliver) is a block.
Error handling — per-step retry and on_error
Agent and Deliver nodes accept two optional config keys (absent = today's
behavior, so existing graphs are untouched):
retry.max_attempts— integer1..5(default1= no retry).retry.backoff_seconds— number0..30(default5), a fixed delay between attempts; additionally(max_attempts - 1) * backoff_secondsmust be<= 120. Out-of-range values are a publish block (bad_retry_config).- Only known-transient failures retry, fail-closed: an Agent step retries on a
transport error or HTTP
5xx/429; a Deliver step retries only adelivery_status: "failed"outcome. Policy blocks (blocked_pii/blocked_ssrf/blocked_tenant) and other 4xx are permanent and never retried; every Deliver retry re-runs the full fail-closed gate chain (PII, SSRF pinning, tenant allowlist). Deliver retry is at-least-once: afailedoutcome can be a timeout after the request was already received, so a retry can deliver twice — make webhook receivers idempotent. (failedalso covers permanent causes, e.g. an invalid channel id, which retry pointlessly — bounded bymax_attempts.) - Each attempt persists its own step row (
attempt1-based) with that attempt's cost/tokens; the lease is re-extended around every attempt and backoff sleep. Retries count against the run budget, and the run deadline/cancel/budget are checked before every attempt — under deadline pressure absorption is best-effort (arun_deadline/budget_exceededabort wins over the policy).
on_error — what happens when the FINAL attempt fails ("error" default):
| Policy | Behavior |
|---|---|
error |
the run fails with the step's error code — exactly today. |
continue |
the step's final row becomes continued (failure absorbed, underlying error/delivery_status preserved); the run proceeds. The step produces no output, so {{}} references to it are publish-blocked (ref_to_fallible_step). Not allowed inside loop bodies (on_error_in_loop_body). |
error_branch |
the node carries a second out-edge (branch: "error" — the catch chain). On final failure the entire after-chain is persisted skipped, the step becomes continued, and the catch chain runs; on success the catch chain is persisted skipped and never runs. Catch chains are linear (no Condition/Loop/nested error edges — branch_depth_exceeded); references to the failing step from its own catch chain are publish-blocked (ref_to_failed_source), while after-chain references stay legal (the after-chain only runs on success). Requires exactly one error edge (error_branch_missing_edge / error_edge_without_policy); not allowed inside containers (error_branch_in_branch). |
A run whose failures were all absorbed finishes succeeded — the step rows carry
the truth (the runs UI shows per-attempt rows, the continued disposition, and any
preserved block outcome; a blocked_pii absorbed by continue still renders the
PII block on its step row).
Publish-time validation codes added by this feature: bad_retry_config,
bad_on_error, error_branch_missing_edge, error_edge_without_policy,
error_branch_in_branch, on_error_in_loop_body, ref_to_fallible_step,
ref_to_failed_source (all re-checked at run-claim like every other rule).
Error codes
The run's typed error_code drives a translated message in the runs UI via the
workflow.runErrors.* i18n namespace (kept separate from the builder's validation
workflow.errors.* codes to avoid a key collision), with a generic fallback for an
unrecognised code; the raw English error string is shown only in an expandable
technical-details view (never verbatim on the main surface).
error_code |
Meaning |
|---|---|
lease_expired |
the runner died mid-run; the lease reaper marked it failed (visible, never eternal "running"). Manual replay from the run detail. |
budget_exceeded |
the run hit run_budget_micros; aborted at a step boundary. |
graph_invalid |
the bound version failed structural validation at run-claim (e.g. an agent's schema drifted, or a cycle). Re-validation happens at claim, before any step runs. |
agent_not_owned |
an Agent-Step references an agent the workflow owner no longer owns (or that was deleted). Enforced at publish AND re-checked at run-claim — the run fails closed before any invoke, because the per-gateway service token could otherwise invoke any agent. |
agent_failed |
an Agent-Step invoke returned a non-2xx at runtime (the agent exists and is owned, but the model/provider call failed) and the step's retry/on_error policy did not absorb it. (Independent of configured retries, a single in-attempt retry fires on a 422 structured_output_failed response — stochastic and idempotent.) |
run_deadline |
the run exceeded its wall-clock cap at a step boundary. |
template_unresolved |
a {{step.output.path}} reference could not be resolved (missing/unfinished step) — a hard step error, never a silent empty string. |
template_null_path |
a resolved path traversed a JSON null — documented hard fail (a Condition's compared path or any {{}} ref). |
condition_type_error |
a Condition gt/lt comparison hit a non-numeric value or operand — fail-closed, never a silent else. |
loop_not_array |
a Loop's items_ref did not resolve to a JSON array — rejected before iteration one. |
loop_too_many |
a Loop's array exceeded the 10-iteration cap — rejected before iteration one (no silent clamp). |
step_output_too_large |
a step's output exceeded the run_state cap (256 KB/run, 32 KB/step). |
deliver_blocked_pii / deliver_blocked_ssrf / deliver_blocked_tenant / deliver_failed |
a Deliver node did not reach delivered. |
schedule_deleted / schedule_disabled |
a still-pending schedule-triggered run was cancelled because its schedule was removed or disabled (kill-switch parity — no orphaned pending fire). |
webhook_trigger_deleted |
a still-pending webhook-triggered run was cancelled because its webhook trigger was deleted (the same kill-switch parity as the schedule codes above). |
fetch_failed |
a Fetch node's download failed transiently (network / timeout / redirect-loop / upstream 5xx / 429) and the step's retry/on_error policy did not absorb it. Retryable (mapped status 502). |
fetch_blocked |
a Fetch node's URL was refused by the SSRF pin (private/link-local/loopback/metadata target, or an unsafe redirect). Deterministic — never retried. |
fetch_http_<status> |
a Fetch node's upstream returned a non-2xx. A 4xx (except 429) is deterministic (not retried); a 5xx/429 is retryable. |
fetch_too_large |
a Fetch node's body exceeded the fetch/artifact cap (~6 MiB). Deterministic. |
fetch_empty |
a Fetch node got a 2xx with an empty body. Deterministic. |
content_type_rejected / content_type_not_allowed |
a Fetch node's bytes failed the magic-authoritative content sniff, or the sniffed canonical MIME was not in the node's allowed_content_types. Deterministic. |
egress_blocked_residency |
a Fetch or Connector node was refused because the gateway is residency-restricted — EU-residency enforced (per-gateway flag, tenant floor, or the deployment env) or all providers local. The node cannot egress to a third party at all. Deterministic (never retried); no call/download occurs. |
url_pii_blocked |
a Fetch node's resolved URL carried protected structured PII (a checksum-valid national ID, IBAN, card number, or a configured keyword) on a PII-active / PII-masking-enforced gateway. A URL cannot be tokenized and stay fetchable, so the policy is refuse, not scrub. Deterministic; no download occurs. |
gateway_unresolved |
a Fetch node's gateway could not be resolved (soft-deleted or corrupt config). Deterministic fail-closed (no egress). A transient config read error instead returns a retryable db. |
code_disabled |
a Code node ran on a gateway where the code-interpreter is not fleet-configured, or has been explicitly disabled (code_interpreter.enabled: false). Deterministic (501) — absorbable by on_error, never a transient tick-retry. |
code_unavailable |
a Code node's sandbox call reached the broker but it was unavailable/saturated (no-isolation / busy / transport). Retryable (mapped status 502). |
code_too_large |
a Code node's job exceeded the sandbox pre-flight size cap (the encoded {code,stdin,limits} body). Deterministic. |
code_failed / code_timeout |
a Code node's program exited non-zero (or with an unknown exit) / hit the wall-clock limit. Deterministic. The response carries the exit_code only — never stderr (it would persist in workflow_step_run.error). |
code_no_output / code_no_usable_output / code_ambiguous_output |
a Code node produced no /output artifact / only harvest-rejected outputs / more than one with no output_filename to disambiguate. Deterministic. |
code_duplicate_input / input_not_found |
two input_refs resolved to the same /input filename / an input_refs handle is not an artifact of this run. Deterministic (422 / 404). |
code_reserved_input |
an input_refs artifact's stored filename is the reserved _context.json (the always-staged trigger-context slot). Deterministic (422) — rejected before the run so a step cannot spoof/clobber the context. |
code_output_too_large |
a Code node's /output exceeded the ~6 MiB artifact store cap. Deterministic. |
connector_missing_field |
a Connector node's connector_id, tool, or input was absent or empty (publish-time validation; also the runtime bad_* guard). Deterministic. |
connector_bad_input |
a Connector node's resolved input did not decode to a JSON object — a scalar, a non-empty array, or undecodable JSON (ABSENT≠MALFORMED: a value that arrives wrong is rejected, never degraded to "no args"). Deterministic; no call occurs. |
connector_pii_blocked |
a Connector node's arguments carried protected structured PII (checksum ID / IBAN / card / configured keyword) on a PII-active / PII-masking-enforced gateway and the connector is not a trusted_subprocessor. Refuse, not scrub. Deterministic; no call occurs. |
connector_gateway_unresolved |
a Connector node's gateway could not be resolved (soft-deleted or corrupt config). Deterministic fail-closed (no call). A transient config read error instead returns a retryable db. |
connector_call_failed |
a Connector node's tools/call failed — a definitive error (not_found / tool_blocked / credential_required / credential_expired / init_error) or an AMBIGUOUS upstream outcome (timeout / reset / 5xx-after-send). Always terminal, NEVER auto-retried (D1 — a mutating call whose outcome is unknown must not re-fire). A crash mid-call instead surfaces the shared deliver_unknown_outcome on resume (the resume reap is node-agnostic). |
Fetch node — SSRF-pinned download
A Fetch node downloads a URL to a binary workflow artifact (the /artifact store)
and emits its handle ({{fetch_key.output.artifact}}) for a later node to consume. Because the
Python runner has no HTTP-fetch and no SSRF pin, the download runs server-side in the gateway
(the loopback /fetch endpoint, below) — making the gateway the single, audited egress
authority for workflow downloads. The trust-boundary contract (Invariant 11):
urlis untrusted: the SSRF pin (connection_plan/resolve_pinned_ip) fails closed — private/link-local/loopback/cloud-metadata targets, userinfo, control bytes, DNS-rebinding, and unsafe redirects are all blocked (fetch_blocked). The pin is re-validated on every redirect hop.- The fetched bytes are untrusted third-party content: the artifact's MIME is decided by a
magic-byte sniff of the bytes (
upload_sniff), never the responseContent-Typeheader — a spoofed or absent header cannot smuggle a disallowed type past the gate. A node's optionalallowed_content_typesis enforced against the canonical (sniffed) MIME; a malformedallowed_content_typesis rejected, never treated as "allow any". - Caps: the body is capped at
max_bytes(clamped to the ~6 MiB artifact cap); an oversize body isfetch_too_large. A 2xx empty body isfetch_empty. Both deterministic. - Retry: transient failures (
fetch_failed, upstream5xx/429) retry per the node'sretryconfig or take itserror_branch; deterministic failures do not. - Audit: every egress attempt (blocked, failed, rejected, or stored) writes an
audit_logrow (action=workflow.fetch) with the authority only (scheme://host— never the path or query, which can carry a token) plus a SHA-256 of the full URL for correlation, the byte size, and the outcome. Best-effort (an audit failure never fails the fetch). - Residency (fail-closed): on a residency-restricted gateway the Fetch node is refused
before any download (
egress_blocked_residency). "Restricted" means the gateway is EU-residency enforced (the per-gatewayeu_region_routingflag, the tenant residency floor, or the deployment-wide enforcement) or every configured model provider is local — and, fail-closed, any gateway whose residency posture cannot be positively determined. A workflow run has no conversation/project context, so this binds to the gateway's own residency posture. (Code nodes are unaffected: their sandbox is network-isolated — no third-party egress on any gateway.) - Structured-PII in the URL (fail-closed): on a PII-active or PII-masking-enforced gateway, a
resolved
urlcarrying protected structured PII (national ID, IBAN, card number, or a configured keyword) is refused (url_pii_blocked) — mirroring the chat fetch tool and the Deliver node. A non-UTF-8 URL is rejected up front (bad_url). Free-text names pass. - Egress reach: the fetch shares the process-wide operator egress allowlist (empty in int/prod → fully fail-closed). An allowlisted host short-circuits the private-IP block and is honored on redirect hops — set it only for operator-vouched hosts. A dedicated per-tenant workflow-egress allowlist is a planned hardening.
Code node — headless sandbox run
A Code node runs author-supplied Python in the network-isolated code-interpreter sandbox
(--network none KVM micro-VM; pypdf/reportlab/openpyxl are available) over prior nodes' artifacts,
and stores the program's /output as a new workflow artifact ({{code_key.output.artifact}}). The
sandbox call runs server-side in the gateway (the loopback /code endpoint, below), reusing the
same ctx-free broker the chat code-interpreter uses. The trust-boundary contract (Invariant 11):
- Enablement (opt-out; AGF-2950): the code-interpreter is on by default, so the node runs
whenever it is fleet-configured and the run's gateway has not opted out
(
code_interpreter.enabled != false; absent /null→ on). Only an explicitcode_interpreter.enabled: falseyieldscode_disabled(501, no sandbox call). codeis OPAQUE: the source is untrusted but runs only in the isolated sandbox (never in the gateway) — isolation is the control; it is size-bounded (≤ 64 KiB; the broker re-checks the encoded job against its 128 KiB cap →code_too_large). It is never template-resolved: a{{…}}-shaped token in the code is literal Python, so run data can never be spliced into executed source (a code-injection seam) and legit Python containing{{(an f-stringf"{{x}}", a set literal{{1,2}}) publishes clean. Run data reaches the program only as data, via/input/_context.jsonandinput_refs./input/_context.json— the trigger context: every Code node's sandbox is staged with a/input/_context.jsonfile carrying the run's consumed trigger output as a JSON object. Read it withjson.load(open("/input/_context.json")). For a form run it is the validated{field_key: value}submission — and a file field carries the extracted, PII-masked text (identical to{{trigger.output.<field>}}), never the upload sentinel or its single-usesubmission_token. For a webhook run it is the webhook JSON payload. Every submitted trigger field lands in the sandbox (not only referenced ones). Values are data, not source — quotes, newlines, backslashes and non-ASCII round-trip because they are JSON-encoded. The file is always present and always a JSON object: a manual run with no input, or a trigger output that is absent / not an object / a JSON array, degrades fail-closed to{}(sojson.load(...).get(...)never raises) — a data-only empty object, never a control bypass. A pathologically deep trigger payload that cannot be re-encoded also degrades to{}(fail-closed)._context.jsonis a RESERVED name: aninput_refsartifact whose stored filename is_context.jsonis rejected before the run (422 code_reserved_input) — a prior or malicious step cannot spoof or clobber the trigger context.input_refsare artifact handles of this run: each is loaded run-scoped + lease-fenced + SHA-verified (a foreign/unknown handle →input_not_found, never a cross-run/tenant read), capped at 7 files (the sandbox holds 8; one slot is reserved for/input/_context.json) — enforced at both publish (code_bad_input_refs) and runtime (400 bad_input_refs), so there is no publish-clean/runtime-fail drift. They are staged as/input/<stored-filename>. Duplicate/inputnames are rejected (code_duplicate_input) before the run. The context file counts toward the shared 16 MiB input-files total, so a code node already staging near 16 MiB ofinput_refscan newly trip the total-size cap (code_bad_input).- Output: the
/outputartifact is magic-gated + capped by the sandbox harvest, then re-capped at the ~6 MiB artifact store. Selected byoutput_filename(or the sole output's own name); bytes come from the artifact'stext(csv/txt/json) orbytes(pdf/xlsx/png) by kind. Zero/ambiguous output is a deterministiccode_no_output/code_ambiguous_output. - Classification: a broker call that reaches the wire but fails (no-isolation / busy / transport)
is the only retryable class (
code_unavailable, 502). A pre-flight reject (never reached the sandbox — oversized job, unsafe input) and a program failure (code_failed/code_timeout, from a successful broker call with a non-zero/timeout result) are terminal. - Audit: every run attempt writes an
audit_logrow (action=workflow.code) with the input handle count, output size, and outcome — never the code source or stderr (they would persist and can embed secrets). Best-effort.
Connector node — call an MCP/API connector mid-run
A Connector node calls an existing MCP/API connector's tools/call mid-run and maps the
(PII-scrubbed, size-capped) result into run_state as {{connector_key.output.result}} for a later
step. The Python runner has no MCP/credential/SSRF capability, so the call runs server-side in the
gateway (the loopback /connector endpoint) — the same single, audited egress authority as Fetch.
The trust-boundary contract (Invariant 11):
- Run-as-owner. The call is made as the workflow owner (the workflow's
created_by), using the owner's credential and resolving the owner's private connector. A cross-user private connector is invisible (not_found→connector_call_failed); a tenant-shared connector (created_byNULL) is tenant-fenced, not owner-fenced, so any author in the tenant may bind it. Trigger-supplied data (e.g. an inbound email) that flows throughinputinto the owner's connector + credential is an accepted run-as-owner trust boundary, governed by the PII/residency gates below. inputis untrusted. The resolvedinputmust decode to a JSON object (or empty). A scalar, a non-empty array, or undecodable JSON is rejected (connector_bad_input) — ABSENT and MALFORMED do not share a permissive path; a value that arrives wrong is rejected, never degraded to "no args". A non-UTF-8inputis rejected up front (bad_input).- Residency (fail-closed). On a residency-restricted gateway (EU-residency enforced, all
providers local, or an indeterminate posture) the node is refused before any call
(
egress_blocked_residency). A workflow run has no conversation/project context, so this binds to the gateway's own residency posture. - Structured-PII in the arguments (fail-closed). On a PII-active or PII-masking-enforced gateway
the
argumentsare scanned; protected structured PII (national ID, IBAN, card number, or a configured keyword) is refused (connector_pii_blocked), not scrubbed (a token would break the tool) — unless the connector is marked atrusted_subprocessor(a vetted EU/DPA subprocessor; the per-connector exemption set via the four-eyes setter — see MCP connectors). Free-text names pass. The samemcp_egressauthority governs the chat agent tool path. - Result is scrubbed + capped. The result is stringified, pre-capped to 20 KiB, then run through
the strict persistence-leg PII fence (the same masker the form-file consume uses): on a PII-active
gateway without a fail-closed masker, or under an untokenizable mandate, the content is withheld
behind a safe placeholder — the raw result is never persisted into
run_state. - No retry — ambiguity fails closed (D1). The connector node is not a retry node: it accepts no
retry/on_error/error-edge (rejected at publish —bad_retry_config/bad_on_error/invalid_branch_slot) and may not sit inside a loop body (loop_body_side_effect). Because the MCP client collapses timeout / connection-reset / 5xx-after-send into one ambiguousupstreamoutcome with no pre-send-vs-post-send signal, a mutating call that ends ambiguous must not re-fire: every failure (definitive or ambiguous) is terminal (connector_call_failed). A crash mid-call leaves adeliveringstep row, which the node-agnostic resume reap fails closed asdeliver_unknown_outcome(never a second call). - Audit. Every call attempt (every outcome) writes an
audit_logrow (action=workflow.connector) with theconnector_id,tool, and outcome only — never the arguments or the result (both can carry the tenant's PII). Best-effort (an audit failure never fails the call).
Worked example — download, fill, and email a PDF form
The Fetch, Code, and Deliver nodes compose into the document-fill workflow the epic exists to enable: download a blank PDF form, fill it with the requester's data in the sandbox, and email the completed form as an attachment to an external recipient — behind a human approval. The builder exposes Fetch and Code in the node palette, so this is authored end to end in the UI, not only via the API.
The graph is trigger → fetch → code → approval → deliver:
- Fetch downloads the blank form (
url,allowed_content_types: ["application/pdf"]) →{{f1.output.artifact}}. - Code stages that artifact at
/input/<name>, runs Python (pypdf) that fills the AcroForm text, checkbox, and radio fields from an embedded/parameterised profile, and writes/output/<output_filename>→{{fill.output.artifact}}. The fill sets both each field's value and every widget's/ASappearance state across all pages — a checkbox or radio widget can live on any page, and a stale/ASon a second widget silently defeats the fill. - Approval suspends the run for a named approver to review before anything leaves the tenant.
- Deliver (
kind: email) attaches the filled form viaattachment_ref: "{{fill.output.artifact.handle}}"and sends it to an externaltoaddress — permitted only because the recipient's domain is on the tenant'sworkflow_recipient_allowlist(see External recipient). On a PII-masked run the external send is blocked fail-closed (recipient_pii_blocked).
A committed fixture of this workflow (a small AcroForm PDF, the fill program, and the graph) lives at
tests/fixtures/vipwing/ and is driven end to end by case_vipwing_reservation in
scripts/workflow_e2e.py (bash dev-instance.sh workflow-e2e).
Template resolution ({{ }}) — trust boundary (Invariant 11)
Data flows between steps via {{step_key.output.path}} references, inserted by the
builder's picker (chips, not free-typed handlebars). Resolution is single-pass over
graph-authored templates only: step outputs and trigger payloads are opaque data and
are never re-scanned — a {{ }} sequence appearing inside a webhook payload or a
model response is not expanded (closes the second-pass injection class). A missing
or unfinished referenced step is a hard step error (template_unresolved), never a
silent empty substitution; a cjson.null anywhere on the path is a documented hard
fail (template_null_path). References into Condition/Loop branch members from
outside the branch are impossible by construction — the publish-time ancestor rule
(dangling_ref) rejects them, so no runtime "skipped reference" case exists.
References to a step whose on_error is continue are likewise publish-blocked
(ref_to_fallible_step); references to an error_branch step are legal from its
after-chain (which only runs when the step succeeded) and publish-blocked from its
own catch chain (ref_to_failed_source).
Internal runner seam (not public)
The runner (scripts/workflow_runner.py) holds no DB credentials; it drives a
loopback-only internal API (127.0.0.1:8083, location ~ ^/internal/workflow/),
authenticated by the shared X-AIG-Scheduler-Secret. Endpoints: /reap, /due-runs,
/claim, /load-run, /extend-lease, /save-step, /step-state, /step-cost,
/finish-run, /suspend, /due-approvals, /apply-approval-timeout, /set-delivery,
/artifact, /artifact-get, /fetch, /code. Every runner-side
mutation is lease-token fenced (claimed_by = token AND status = 'running'); a
409 { "error": "lease_lost" } tells the runner it was reaped and must abort with no
further egress or overwrite. This surface is unreachable off-container and never called
from a browser.
Although loopback-only and secret-authenticated, the runner is still treated as an
untrusted-shape client (Invariant 11): every field is validated before any storage call.
The pii_active field on /finish-run and /suspend is accepted only as an
integer 1/0; any other shape is coerced fail-closed to 0 — a missing field, a
JSON boolean (tonumber(true) is nil), or a non-numeric string all become 0. The
value is never trusted to lower the stored flag: storage latches it with
GREATEST(pii_active, ?), so a 0 can only ever be a no-op.
/save-step persists a workflow_step_run row, whose tenant_id scopes GDPR export
(list_workflow_step_runs_for_export) and the erasure scrub. That tenant_id is
DERIVED from the lease-fenced parent run, never from the body: the fence
SELECT returns the run's tenant, and the INSERT stamps that value. A body
tenant_id that is present and different from the run's tenant is a runner bug and
is rejected 400 { "error": "tenant_mismatch" } with no write (fail-closed, loud);
an absent or JSON-null tenant_id is treated as absent and derives silently; a
run row whose own tenant is empty/NULL (schema-impossible corruption) is rejected
500 { "error": "db" } (run_tenant_missing), also with no write. Every field on the
row-writing endpoints — /save-step, /set-delivery, /finish-run, /suspend — is
validated column-true before the write: integers must be finite and inside the real
column range (tokens [0, 2³¹); cost_micros/tokens_used/started_at/finished_at
[0, 2⁵³) — a present-but-invalid value is 400, never a silent floor/zero), identifier
strings must be valid utf8mb4 within the column width (step_key/request_ref/
approval_step_key ≤ 64, node_type/status/delivery_status ≤ 16, suspended_reason
≤ 32, approver_id ≤ 36 codepoints, else 400), and free-text error/delivery_error
are clamped to 16000 codepoints while error_code clamps to 48 (invalid utf8mb4 → 400
on all of them). On /save-step and /set-delivery an empty-string or JSON-null
optional field is stored as SQL NULL, not ''. So no client shape can mint a 5xx
from the write.
/artifact and /artifact-get — binary artifact store
A workflow node produces a binary artifact (e.g. a fetched or filled document) that a later
node consumes by handle, without ever putting bytes in run_state (JSON, 256 KB cap). Bytes
live in exactly one row (workflow_artifact.data, LONGBLOB, FK → workflow_run ON DELETE
CASCADE); run_state holds only a reference {"artifact": {"handle", "filename", "mime_type",
"size_bytes", "sha256"}}. Both endpoints are lease-token fenced like every other runner mutation.
POST /artifact — {run_id, token, step_key, filename, mime_type, data_b64} → {handle,
size_bytes, sha256}. Accepted/rejected shape (Invariant 11):
- data_b64 is read raw and type()-checked, then base64-decoded, with the emptiness check on
the decoded bytes (not the raw string, because a base64 decode of a padding-only string like "====" yields "", not
nil): a non-string field → 400 missing_artifact_data; a decode == nil → 400
bad_artifact_b64; a decoded length of 0 (covering "", "=", "==", "====") → 400
empty_artifact. No 0-byte or wrong-size row is ever stored; size_bytes is the decoded length.
- filename is basename-stripped (path separators removed) then re-checked empty ("../",
"/", ".", ".." → 400), and filename/mime_type/step_key are utf8mb4-validated within
their column widths (else 400, never a store-time 5xx). Note for the consuming node:
basename does NOT strip NUL/control bytes — the artifact is stored opaque here, so the node that
stages it at /input/<name> MUST sanitise NUL/control at that boundary before it reaches a
C-based path handler.
- tenant_id is DERIVED from the lease-fenced parent run, never from the body. The fence is an
INSERT … SELECT … FROM workflow_run WHERE id = ? AND claimed_by = ? AND status = 'running';
rowcount 0 (reaped lease, or unknown run — no existence oracle) → 409 lease_lost with no write.
- Per-artifact and per-run aggregate size caps are enforced before the insert (413
artifact_too_large / artifact_run_total_too_large); a MySQL max-packet refusal is mapped to the
same 413, never a 500.
POST /artifact-get — {run_id, token, handle} → {filename, mime_type, size_bytes, sha256,
data_b64}. Lease-fenced (lost → 409); the row lookup is scoped WHERE id = ? AND run_id = ?, so a
handle belonging to another run/tenant → 404 not_found (a valid-lease holder cannot distinguish
foreign-exists from not-exists). content_sha256 is recomputed over the stored bytes and compared
on read — a mismatch fails closed (sha_mismatch, row withheld).
/fetch — SSRF-pinned server-side download
POST /fetch — {run_id, token, step_key, url, allowed_content_types?, max_bytes?, filename?} →
{handle, size_bytes, sha256, content_type, filename}. The Fetch node's server-side worker (see the Fetch node section
above for the full contract). Order, all
fail-closed: validate shape (allowed_content_types an explicit array-of-strings — a malformed value
→ 400 bad_allowed_types, never allow-all; max_bytes clamped to the artifact cap; filename
basename-sanitised, malformed → 400 bad_filename) → lease-fence the run FIRST (load_workflow_run;
lost → 409, no egress) → fetch_bytes (SSRF-pinned) → audit the egress attempt (authority-only,
no path/query) → empty-body → 422 → magic content sniff (415 on reject / not-allowed) → store via
create_workflow_artifact (409 lost lease, 413 cap). Status map: SSRF 400; network/timeout/
redirect/upstream-5xx/429 502 (retryable); upstream other-4xx & empty 422; oversize 413. The
runner's HttpWorkflowDb.fetch treats a 502 as retryable and every other non-2xx as a terminal
FetchError; a 409 is a lost lease (abort).
/code — headless sandbox run
POST /code — {run_id, token, step_key, code, input_refs?, output_filename?, context?} →
{handle, size_bytes, sha256, mime_type, filename}. The Code node's server-side worker (see the Code
node section above for the full contract). context is an internal loopback field (not a persisted
contract): the runner always sends the run's consumed trigger output (state[trigger_key]) as a
nested JSON object — code is passed verbatim (never template-resolved). Order, all fail-closed:
validate shape (code non-empty ≤ 64 KiB; input_refs an explicit array-of-strings capped at 7 —
#input_refs + 1 > 8 reserves the /input/_context.json slot → 400 bad_input_refs; output_filename
basename-sanitised) → lease-fence the run FIRST (load_workflow_run; lost → 409)
→ enablement gate (is_configured() + resolve_gateway_config_by_id + config_code_interpreter_enabled;
disabled → 501 code_disabled, no sandbox call) → load each input_ref (get_workflow_artifact,
run-scoped; unknown → 404 input_not_found) + reject the reserved name _context.json
(422 code_reserved_input) + reject duplicate /input names (422 code_duplicate_input) → stage
/input/_context.json from context (object-or-empty guard: a non-object / JSON array / un-encodable
value → the literal {}, fail-closed) → run_code (ctx-free broker; artifact_max_bytes = the store cap) → classify by return shape:
(nil,err,outcome≠nil) broker failure → 502 code_unavailable (retryable), (nil,err,nil-outcome)
pre-flight → terminal 413 code_too_large / 422 code_bad_input, a RunResult with timed_out →
422 code_timeout / exit_code≠0|nil → 422 code_failed (exit_code only, no stderr) → select the output
by art.name (bytes by kind) → store via create_workflow_artifact (409 lost lease, 413
code_output_too_large) → audit (metadata only). The runner's HttpWorkflowDb.code maps only 500
→ TransportError (transient tick-retry); everything else — including 501 code_disabled — is a
terminal CodeError (so on_error/error_branch can absorb it); 502 is node-retryable; 409 is a
lost lease (abort).