Skip to content

Workflows API

A workflow is a versioned, gateway-scoped DAG of typed steps — Trigger, Agent-Step, Condition, Human-Approval, Deliver — authored in the visual builder, rendered as a vertical chain/tree and stored as a directed acyclic graph (graph_json). A workflow is edited as a draft, then published to an immutable, semver-tagged snapshot (workflow_version); a run is bound to the exact version it started on, so a later publish never mutates a running or suspended run.

A manual run always executes the latest published version, never the unsaved draft. So the caller can tell when the draft has diverged, GET /admin/v1/gateways/{gateway_id}/workflows/{id} returns has_unpublished_changes (boolean): true when the draft graph_json differs from the published version a run would execute — publish first to run the edits. It is false when the draft is in sync or the workflow was never published.

Runs are executed by a dedicated runner process (scripts/workflow_runner.py) that claims pending runs via a compare-and-swap lease and persists progress after every step, so a resumed/replayed run in a fresh process resolves data-flow references from durable state.

Admin endpoints require an authenticated admin session. Base URL: https://ai-api-admin.myra.eu/admin/v1. The whole feature is gated per workspace by tenant.workflows_enabled (default off, fail-closed): every admin workflow route returns 403 { "error": "feature_disabled" } until an Admin enables it.

That workspace flag is flipped through its own route, which is platform-admin only (role == "admin"; a tenant admin receives 403 — stricter than the require_author gate the workflow routes themselves use):

POST /admin/v1/gateways/{gateway_id}/workflows-feature with { "enabled": <boolean> } -> 200 { "enabled": <boolean> }. A non-boolean enabled is rejected 400.


Authorization matrix (server-side; the client is never the authz boundary)

Action Who
create / edit / publish / enable / disable / seed example / triggers + schedule config any authoring user — member, ki_manager, tenant_admin, admin (require_author; viewer/demouser → 403). The builder is open to all authoring users; workflows are gateway-shared (per-user sharing is on the roadmap)
read a workflow definition — list, detail (graph_json), versions, versions-diff any authoring user (require_author; viewer/demouser → 403). The definition is authoring IP (prompts, agent references, step wiring), gateway-shared among authors, not owner-scoped. These four reads were previously gateway-access-only (no role gate) — the one gap left when the rest of the surface was author-gated
trigger / run any authoring user (require_author; viewer/demouser → 403)
view runs (list + detail) any authoring user, owner-scoped — run detail exposes per-step outputs (run_state, possibly PII). A non-manager sees only their own runs (created_by); a manager (admin/tenant_admin/ki_manager) sees all. A non-owner requesting another user's run → 404 (never 403 — no run-existence oracle); a viewer/demouser → 403
enable / disable the feature for a tenant (workflows-feature) Platform admin only (role == "admin")
cancel a run the run's owner, an Admin, or a Tenant-Admin
approve / reject a suspended run only the node's named approver user (checked server-side); any authenticated user can be a named approver — the approver inbox is not KI-Manager-gated. Group approvers are a fast-follow.
referenced agents the workflow owner (created_by) must own every agent an Agent-Step references (see below)

Throughout this page, an "authoring" session / "any authoring user" (the require_author gate) means a caller holding the WORKFLOWS_AUTHOR permission. For the built-in roles that is member, ki_manager, tenant_admin, and admin; a tenant custom role granting WORKFLOWS_AUTHOR is admitted regardless of its base role. A viewer/demouser (or any role without the permission) is refused 403.

run-as-owner + agent ownership. A workflow runs as workflow.created_by; an Agent-Step invokes the referenced agent over the gateway service-token path, so the agent executes in the responsibility of the agent's owner. To prevent a confused-deputy, an Agent-Step may reference only an agent the workflow owner owns (ownership, not share-visibility). This is enforced at publish (a reference to a non-owned agent is a publish error) and re-checked at run-claim (ownership can change after publish → the run fails closed with agent_not_owned).


Workflow CRUD

GET /admin/v1/gateways/{gateway_id}/workflows → { workflows: [ … ] } POST /admin/v1/gateways/{gateway_id}/workflows → create a draft.

Both routes require an authoring session (require_author — viewer/demouser → 403); listing definitions is not open to read-only roles.

Each list row carries the workflow's identity/status/version/timestamps plus three server-derived list metrics:

  • steps — count of actionable step nodes in the draft graph (excludes the singleton trigger entry).
  • agents — the distinct agent slugs the draft graph's Agent-Steps reference (a JSON array; [] when none).
  • runs_30d — number of runs of this workflow in the last 30 days.

The list row does not include graph_json — the full definition is returned only by the detail GET …/{id}. steps/agents derive from the draft graph, so a published workflow whose draft has diverged shows its current edited structure.

Create body (accepted shape; anything else is rejected 400):

{ "name": "Presseschau mit Freigabe" }

GET /admin/v1/gateways/{gateway_id}/workflows/{id} → the workflow + current draft graph. PATCH /admin/v1/gateways/{gateway_id}/workflows/{id} → edit the draft (autosave). DELETE /admin/v1/gateways/{gateway_id}/workflows/{id} → delete.

All three require an authoring session (require_author; viewer/demouser → 403) — the detail response returns the full graph_json definition.

Draft edit is optimistic-concurrency guarded (D12): the body MUST carry the updated_at the client last read. A mismatch (another tab saved first) returns 409 { "error": "stale", "updated_at": <current> }; the client reloads and retries.

On success the response echoes the freshly stamped token — 200 { "ok": true, "updated_at": <new> } — which is the value the next save must send as its precondition. The client reads it straight from this response and needs no follow-up GET. (A GET read-back was previously required to learn the new token; it was a second failure surface that could misreport a persisted save and then self-409 the next save into a false "another tab saved" state — removed.)

{
  "name": "…",
  "run_budget_micros": 5000000,       // optional € budget (micros); null = no cap
  "graph_json": { "steps": [ … ], "edges": [ … ] },
  "updated_at": 1752000000            // REQUIRED precondition
}

POST /admin/v1/gateways/{gateway_id}/workflows/{id}/status { "status": "disabled" } — enable/disable (published ⇄ disabled).

Seed the example workflows

POST /admin/v1/gateways/{gateway_id}/workflows/seed-example → create two publish-clean example drafts owned by the caller and return their ids. This backs the empty-state CTA ("Beispiel-Workflow erstellen") and the demo-tenant provisioning. Same authorization as create: any authoring user (require_author) + the workspace feature flag.

  • Presseschau mit Freigabe — Trigger → Agent-Step → Human-Approval → Condition(true/else) → Deliver / Deliver (the full node vocabulary; the §demo flagship).
  • PII-Testlauf — Trigger → Agent-Step → Deliver (its output carries test personal data so the Deliver egress trips the fail-closed PII gate → blocked_pii).

Both Agent-Steps reference an agent owned by the caller. When the caller owns none, a per-user example agent (beispiel-agent-<uid8>, model left to the gateway default) is auto-provisioned so the example publishes clean on a fresh tenant.

Request body — all fields optional (accepted shape; anything else is rejected 400):

{
  "agent_slug": "presseschau-agent",        // pin an OWNED agent (else auto-provisioned)
  "approver_email": "redaktion@example.org" // a distinct in-tenant approver (else the caller)
}

Rejections (fail-closed): agent_slug not owned by the caller → 400; approver_email not a live user in this workflow's tenant → 400; feature disabled → 403; not an author (viewer/demouser) → 403.

Response 201:

{ "primary_id": "<Presseschau id>", "workflows": [ { "id": "…", "name": "…" }, … ] }

Idempotent per the caller's own workflows by name: a repeat returns the existing ids (never duplicates); it never returns another user's same-named workflow. The examples are drafts — not auto-published — so the operator can adjust before publishing.

Demo-tenant provisioning (§10): seed, then set each Deliver channel_id to a real channel and (for the two-person Freigabe beat) seed with the approver persona's approver_email, then POST …/{id}/publish each. A blank channel_id publishes but does not deliver — set it before the live run (and before the PII-Testlauf blocked_pii beat, which needs a reached egress).

graph_json shape

{
  "steps": [
    { "key": "trigger",  "node_type": "trigger",  "config": { "trigger_kind": "manual", "input_schema": { … } } },
    { "key": "classify", "node_type": "agent",    "config": { "agent_slug": "klassifizierer",
                                                              "input": "Klassifiziere: {{trigger.output.text}}" } },
    { "key": "notify",   "node_type": "deliver",  "config": { "kind": "mattermost", "channel_id": "…",
                                                              "text": "Ergebnis: {{classify.output.summary}}" } }
  ],
  "edges": [ { "from": "trigger", "to": "classify" }, { "from": "classify", "to": "notify" } ]
}

node_type ∈ trigger | agent | condition | loop | approval | deliver | fetch | code | connector. Edges may carry a branch slot ("true"/"else" on a Condition, "body" on a Loop, "error" on an Agent/Deliver node whose on_error is error_branch — see Error handling). key is a stable, immutable step id; display labels rename freely. At publish the referenced agents' response_schema + version are snapshotted into the version (drift guard).

Per-node config keys the runner consumes:

Node config keys
trigger (none read at run time — the trigger's output is the run's trigger input; trigger_kind and input_schema are builder/validation metadata). A run's trigger_kind records how it started: manual | webhook | schedule | form | email (an inbound email — see Email-ingest triggers)
agent agent_slug (an OWNED agent), input — the per-step task/prompt (a {{step.output.path}} template). input is required and must resolve to a non-empty string; an empty or missing input fails the step at run time (agent_failed). It is distinct from the agent's own saved system prompt.
condition left_ref (a single {{…}} ref), op (eq\|ne\|contains\|gt\|lt\|is_empty), right_value (ignored for is_empty)
loop items_ref (a single {{…}} ref to a JSON array, ≤ 10 items)
approval approver_kind (user), approver_id, timeout_policy (skip\|end), timeout_hours (1…8760)
deliver kind (mattermost\|webhook\|email), the per-kind target (channel_id | url | subject), text (a {{…}} template), and (email only) optional to — an EXTERNAL recipient, allowlist-gated (see below)
fetch url (a literal or a {{…}} template) — the document to download; allowed_content_types (optional, a non-empty array of MIME strings; a malformed value is rejected at publish, never treated as "any"); max_bytes (optional positive int, clamped to the ~6 MiB artifact cap); filename (optional — else derived from the URL path basename). Emits an artifact reference {{fetch_key.output.artifact}} ({handle, filename, mime_type, size_bytes, sha256}) for a later node. A network fetch is a RETRY node (retry/on_error accepted) — a transient failure retries; a deterministic one takes the error_branch or fails the run. See the Fetch node section below.
connector connector_id (an existing MCP/API connector of the tenant), tool (the tools/call tool name), input (a JSON object literal of arguments, or a single {{…}} ref that resolves to one). Calls the connector mid-run as the workflow owner (the owner's credential / private connector; a tenant-shared connector ignores the user). Emits the scrubbed string result {{connector_key.output.result}}. Side-effecting + ambiguous-outcome → NOT a retry node: retry/on_error/an error-edge are rejected at publish, and any failure is terminal (never auto-retried — see the Connector node section).
code code (Python source, ≤ 64 KiB) — run in the network-isolated code-interpreter sandbox. The source is opaque: it is never template-resolved ({{…}} in the code is literal Python — an f-string f"{{x}}" or a set literal publishes clean), so run data reaches the program only as data, never spliced into source. Data channels: input_refs (optional array of ≤ 7 artifact-handle refs, typically {{prior_step.output.artifact.handle}}, staged as /input/<name>) and the always-present /input/_context.json (the run's trigger input as a JSON object — see the Code node section). output_filename (optional — selects the /output file by name and stores under it; else the sole output's own name, synthesized for a nameless image). Emits an artifact reference {{code_key.output.artifact}} for a later node. A RETRY node, but only a transient broker failure retries — a deterministic user-code failure does not. See the Code node section below.

Publish

POST /admin/v1/gateways/{gateway_id}/workflows/{id}/publish → 201 { "version_id": "…", "semver": "1.1.0", "updated_at": 1742551200 }. The updated_at is the optimistic-concurrency token the next draft PATCH must echo.

Publish runs graph validation (below) and, on success, writes an immutable workflow_version and auto minor-bumps the semver. On validation failure the response is 422:

{ "error": "graph_invalid",
  "errors":   [ { "step_key": "notify", "code": "template_unresolved" } ],
  "warnings": [] }

Errors block publish (rendered as per-node red badges); warnings do not. (The validation model is block-only today — references that could be unavailable at runtime are rejected at publish, never downgraded to warnings.)

Validation rules (also re-run at run-claim as a drift guard): the graph must be a DAG (a cycle is rejected); every edge must reference existing steps with type-compatible ports; every {{step.output.path}} reference must resolve to an upstream step's declared output schema (a dangling reference is template_unresolved); a reference into a Condition/Loop branch member from outside that branch is a hard publish block (dangling_ref — only ancestors on the same chain resolve, so a maybe-skipped value can never be referenced); every referenced agent must be owned by the workflow owner (agent_not_owned); malformed / oversize graph_json is rejected fail-closed. A {{}} reference to an output-less node (condition / deliver / approval — they route, egress, or gate but expose no output) is a publish block (ref_to_output_less); only trigger / agent / loop produce a referenceable value. An approval node must name a valid approver (approver_kind: "user" + an approver_id that is a live user in the workflow's tenant, else approver_missing / approver_not_in_tenant) and a valid timeout (timeout_policy ∈ {skip, end}, timeout_hours a finite number in [1, 8760], else bad_timeout).

GET /admin/v1/gateways/{gateway_id}/workflows/{id}/versions → the published versions (newest first, capped at the newest 100), each enriched with run_count (how many runs executed on that version — the run→version pin is permanent), is_current (semver == workflow.version), and published_by_name (the publisher's email, or the raw id if the user was deleted). Requires an authoring session (require_author; viewer/demouser → 403), then gateway access + the workspace feature flag. It exposes historical version config — the same config class as the current draft, never run-output — so it is author-gated exactly like the definition reads.

Diff two versions

GET /admin/v1/gateways/{gateway_id}/workflows/{id}/versions/diff?from={version_id}&to={version_id} -> 200 { "diff": { … }, "from": …, "to": … }. The diff is computed server-side from the two immutable graph snapshots; the response lists, per step (keyed by the immutable step key): added, removed, changed, unchanged, plus edge edges_added / edges_removed. Each changed step carries a fields list of { field, kind, before, after } where kind ∈ {added_field, removed_field, changed_field} (the absent side is JSON null). Every list is always a JSON array.

Requires an authoring session (require_author; viewer/demouser → 403) — same gate as the other definition reads.

Accepted / rejected: both from and to are required and non-blank (else 400); each must be a version of this workflow — a version id belonging to another workflow or tenant, or a nonexistent id, is 404 (the workflow itself is gateway/tenant-fenced first). A stored snapshot that fails to parse is 422 { "error": "diff_failed" } (fail closed, never a 500). Diffing a version against itself returns an all-unchanged diff.

Roll back to a version

POST /admin/v1/gateways/{gateway_id}/workflows/{id}/versions/{version_id}/rollback -> 201 { "version_id": …, "semver": …, "updated_at": …, "rolled_back_from": … }. Rollback republishes the chosen version's graph as a NEW version — the history is append-only, old versions are never mutated, and runs stay pinned to the version that executed them. It runs the same validation as a normal publish, so a rollback to a version that now references a deleted agent (agent_not_owned) or a cross-tenant/deleted approver (approver_not_in_tenant) is rejected 422. The workflow's current enabled/disabled status is preserved — rolling back a disabled workflow does not silently re-enable it. Rolling back to the already-current version is 409 { "error": "already_current" }. The editable draft is not touched; rollback changes what is published (future runs use the rolled-back graph), not your draft. Requires an authoring session (require_author) — the publish authorization.


Test a single step (builder, no persisted run)

The builder's agent-step config panel offers "Diesen Schritt testen" — run just that one agent step in isolation, without executing the whole workflow. It reuses the agent invoke path (agents/{slug}/invoke via a short-lived playground token, run as the agent owner), so it does not create a workflow_run and does not appear in the run history. Because a mid-chain step has no upstream output at design time, the panel lists each {{step.output.path}} the step references and lets the author supply a sample value; those are substituted into the step input before the invoke. Fail-closed: a deleted/unavailable agent, an unresolved (dangling) reference, or an empty resolved input blocks the test with no invoke. The test-run cost is surfaced from the synchronous X-AIG-Cost-Micros response header and labelled as a test. Scope: agent steps only (Trigger / Condition / Loop / Approval / Deliver are not testable in isolation).


NL copilot — generate / fix / explain a draft graph

A natural-language assistant for the builder. Three modes: generate a draft graph from a process description, fix a graph from its validation errors, explain a graph in plain language. It never publishes — a human reviews and publishes, so the governance story is intact. The proposed graph is untrusted model output and is validated through the FULL publish validator before it may be applied (and again at publish).

The copilot is split into two calls, by concern:

1. Inference (metered) — /v1 pseudo-provider

POST /v1/{tenant}/{gateway}/workflows/copilot (inference host; Bearer a short-lived playground token, exactly like "Test a single step"). This rides the normal inference pipeline, so the run draws the tenant's normal metering (spend_ledger), PII scrubbing, EU-residency and model-allowlist clamps, and returns the cost in X-AIG-Cost-Micros. It is a passthrough (buffered, stream:false); for generate/fix the answer is constrained to the graph JSON schema and validated fail-closed (422 structured_output_failed on non-conforming output, billed first so a caller can't harvest free runs).

Authorization: the token must be user-bound and hold the WORKFLOWS_AUTHOR permission — the SAME permission the admin build/publish routes require, so the copilot admits exactly the users who may author a workflow, never a stale subset. For the built-in roles that is member, ki_manager, tenant_admin, and admin; a tenant custom role that grants WORKFLOWS_AUTHOR is admitted too, regardless of its base role. An unbound (service) token or a caller lacking WORKFLOWS_AUTHOR (e.g. viewer/demouser) is refused 403 {error.code: "forbidden"}. The feature must also be enabled for the workspace, else 403 {error.code: "workflows_disabled"} — a distinct code from the role refusal, so the client can prompt the user to ask an admin to enable Workflows rather than showing a generic permission error.

Accepted body (anything else is rejected 400 — validated at the trust boundary, Invariant 11):

{ "mode": "generate|fix|explain",
  "model": "<a runnable model id>",
  "description": "describe the process (required for generate, optional for fix)",
  "graph": { "steps": [ … ], "edges": [ … ] },
  "validation_errors": [ { "step_key": "…", "code": "…" } ] }
  • mode — must be one of the three; missing/unknown → 400.
  • model — required, non-empty; absent/empty → 400 (no model call). The pipeline clamps it to the gateway's allowlist/residency.
  • description — required for generate; capped at 4 KB; larger → 400.
  • graph — required for fix and explain; the wire shape (see graph_json shape); capped at 64 KB for the prompt; malformed/oversized → 400.
  • validation_errors — optional, only used to phrase the fix prompt.

For generate/fix the response's assistant content is the proposed graph as JSON; for explain it is plain-language text. To keep proposals publishable, the prompt is grounded in the caller's owned agent slugs (agent nodes must reference one) and the caller as the default approver.

2. Server-side validation + diff (before applying)

POST /admin/v1/gateways/{gateway_id}/workflows/{id}/copilot-validate (admin host; require_author; feature-gated). Runs the same "publishable" validator as publish (core.wf_validate + owner-agent fence + approver-in-tenant fence) on the untrusted proposed graph, and computes the changed-nodes diff vs the current draft.

Accepted body:

{ "graph": { "steps": [ … ], "edges": [ … ] },
  "base_graph": { "steps": [ … ], "edges": [ … ] } }
  • graph — required object; not-an-object / not-encodable / >512 KB → 400.
  • base_graph — optional; defaults to an empty graph {"steps":[],"edges":[]} when absent (a fresh generate), so the diff never sees an absent side.

Response 200:

{ "ok": true,
  "errors": [],
  "warnings": [],
  "diff": { "added": [ … ], "removed": [ … ], "changed": [ … ], "unchanged": [ … ],
            "edges_added": [ … ], "edges_removed": [ … ] } }

ok:false returns the validation errors (step_key + code) and the client surfaces them without applying — a graph that fails validation is never applied as a partial graph. errors reflects wf_validate's phased short-circuit (up to the first failing phase, not exhaustive), so the fix UX re-runs validate → fix → validate. A DB failure during validation returns 503 (never a fake validation error). Applying an ok proposal is a normal draft PATCH (the client is never the boundary — publish re-validates regardless).


Trigger a run (manual)

POST /admin/v1/gateways/{gateway_id}/workflows/{id}/runs → 202 { "run_id": "…", "status": "pending" }.

The workflow must be published — an unpublished workflow returns 409 { "error": "workflow not published" }, and one whose versions have gone returns 409 { "error": "no published version" }. The run starts as pending; the runner claims it within ~2 s. A manual run consumes only idempotency_key; it does not accept a trigger input payload (the run begins from the Trigger node with no supplied input, and no input_schema validation is applied on this route — an input field in the body is ignored). To pass a payload, fire the workflow through its webhook or email trigger.

{ "idempotency_key": "optional-client-uuid" }

Idempotency (fail-safe against double-clicks/retries): if idempotency_key is supplied, a second POST with the same key for this workflow returns the same run_id (and "was_existing": true) — never a second run, never double spend or double egress. Omitting the key (or sending null / "") means every POST is a distinct run.

idempotency_key, when present, must be an opaque token: a string of at most 64 characters drawn from A-Z a-z 0-9 - _ . :. Anything else — a longer key, other characters, or a non-string value — is rejected with 400 { "error": "invalid idempotency key" }. This is the same shape the webhook fire path enforces on X-AIG-Idempotency-Key.


Schedule trigger (time triggers)

A workflow can fire on a schedule — a fixed interval, once a day at a set time (UTC), or once a week on a chosen weekday at a set time (UTC) — in addition to manual / webhook triggers. There is at most one schedule per workflow. It reuses the existing agent scheduler (no second scheduler): the schedule is a polymorphic agent_schedule row with workflow_id set; the scheduler's per-minute tick claims a due row via the same at-most-once compare-and-swap and, for a workflow row, asks the gateway to enqueue a pending workflow_run (trigger_kind = schedule, run-as-owner) instead of invoking an agent. The runner then executes it exactly like a manual/webhook run.

Saving a daily or weekly schedule arms its next run to the next occurrence of the chosen time (for weekly: on the chosen weekday), not to the moment of saving — so "daily at 10:00" created at 08:36 first fires at 10:00, never a spurious extra run right after save. An interval schedule starts promptly (its first run is due immediately). Editing only metadata (e.g. renaming, or switching a still-enabled schedule off) preserves the already-armed run time; the next run is recomputed only when the cadence itself changes (including changing only the weekday of a weekly schedule) or a disabled schedule is switched back on.

Authz: any authoring user (require_author; setting a schedule is a create/edit function), plus the per-workspace feature gate. All three routes are gateway-scoped.

Method / path Behaviour
GET …/workflows/{id}/schedule 200 { "schedule": false } when none is set, else { "schedule": { enabled, schedule_kind, interval_sec, daily_at, daily_dow, next_run_at, last_run_at, consecutive_failures, last_status, last_error } }. daily_dow is the weekly schedule's weekday, integer 0..6 (0 = Sunday … 6 = Saturday, UTC); absent/null for interval/daily rows. The three health fields let the UI show an auto-disabled (enabled: 0) or failing schedule instead of rendering it as healthy. last_status is ok / retrying / failed / paused / null; for a workflow schedule the reachable set is paused (the backstop switched it off) or null, which is exactly what lets the UI distinguish a system stop from a human switching the schedule off. last_error is fixed operator-facing English, never model output or PII: either a dispatch string such as no gateway service token, or the auto-disable sentence, which embeds the failing run's error_code accepted only as [A-Za-z0-9_.-]{1,48} — anything else (empty, over-length, markup, quotes, newlines, multibyte) is rejected and rendered as unknown. consecutive_failures is the streak toward the agent auto-pause threshold; on a workflow schedule the scheduler zeroes it on every fire, so it stays 0 and the count is carried in last_error instead.
PUT …/workflows/{id}/schedule Create-or-replace (race-safe upsert). Returns the stored cadence. Disabling cancels the workflow's still-pending schedule runs.
DELETE …/workflows/{id}/schedule Remove the schedule + cancel its pending schedule runs → 200 { "ok": true }.

PUT body — accepted shape (trust boundary, Invariant 11; validated by core.schedule_cadence, the one shared cadence validator):

{ "schedule_kind": "interval", "interval_sec": 3600, "enabled": true }
{ "schedule_kind": "daily",    "daily_at": "06:00",  "enabled": true }
{ "schedule_kind": "weekly",   "daily_at": "06:00",  "daily_dow": 3, "enabled": true }
  • schedule_kind — interval | daily | weekly (required; anything else → 400).
  • interval_sec — integer 60 … 366 days (required for interval). Below 60, above the cap, non-integer, or non-numeric → 400, no write.
  • daily_at — HH:MM in UTC, 24-hour (required for daily and weekly). 24:00, 12:60, a single-digit hour, an empty string, or a non-string → 400, no write.
  • daily_dow — integer 0 … 6 (0 = Sunday … 6 = Saturday, UTC; required for weekly — same contract as scheduled tasks). Out of range, fractional, non-numeric, boolean, or JSON null with weekly → 400, no write. For interval/daily the server stores NULL regardless of what was sent (the client never picks what persists).
  • enabled — strict JSON boolean (a governance toggle is never coerced from a truthy string). Absent → enabled. Any other type → 400.
  • All cadence fields are range/format-checked whenever present, even those the kind does not use (a garbage value in an unused field is rejected here rather than reaching the DB).

Fail-closed fire path (never a surprise run or double spend): the scheduler enqueues only when, at fire time, the workflow is published, its workspace feature is on, it has a published version, and no run is already in flight (any manual/webhook/schedule run in pending/running/suspended → the fire is skipped, at-most-once), the schedule still exists, and the schedule is still enabled (a fire is refused for a schedule switched off after the tick claimed it — this narrows, but cannot fully close, that race: a PUT landing between the check and the insert still yields one run). All identity (tenant / gateway / owner / version) is derived server-side from the workflow row — the scheduler is authenticated but never trusted for identity.

Auto-disable backstop, and how to recover from it. A schedule whose last 5 schedule runs since it was last enabled all failed is automatically switched off, to bound a broken schedule's spend. The window starts when the schedule was last saved with enabled: true, which is what makes the stop recoverable: fix the cause, switch the schedule back on and save — that PUT re-arms the failure window, and the next fire produces a real run. Five fresh failures after that re-arm switch it off again, so the backstop keeps its teeth. The stop records why, so it never looks like a human turning the schedule off: last_status: "paused" plus a last_error naming the failure. Both fields are cleared by the same re-arming save, so a schedule an operator later switches off by hand does not inherit an old system cause.

Two operational notes:

  • Never re-enable a stopped schedule with raw SQL. UPDATE agent_schedule SET enabled=1 does not move the failure-window watermark, so the next tick switches it straight back off. Use the PUT above.
  • A run that ended cancelled inside the window breaks the all-failed streak and defers the backstop until five newer failures accrue — disabling or deleting a schedule cancels its pending schedule runs, so this is the expected behaviour after a stop/start cycle.

Timezone: daily_at is UTC in the MVP (cron expressions and per-workflow timezones are a later roadmap item).


Runs

Both read routes require an authoring session (require_author — viewer/demouser get 403) plus the workspace feature flag, and are owner-scoped: run detail exposes per-step model outputs (run_state), which may carry personal data, so a non-manager may read only their OWN runs. Scoping is by workflow_run.created_by — the triggerer for a manual run, the workflow owner for a schedule/webhook/form run. The runs list returns only the caller's own runs; run detail for a run the caller does not own returns 404 (never 403 — a non-owner must not learn the run id exists). An admin, tenant_admin, or ki_manager sees every run as the audit view. The client is never the authz boundary.

GET /admin/v1/gateways/{gateway_id}/workflows/{id}/runs → { "runs": [...] } with status, trigger_kind, cost_micros, tokens_used, pii_active, suspended_reason, approver_ref, started_at, finished_at, error_code. status ∈ pending | running | succeeded | failed | suspended | cancelled.

GET /admin/v1/gateways/{gateway_id}/workflows/runs/{run_id} → { "run": {...}, "steps": [...] } (404 if the run is not on this gateway). run adds the technical error string and run_state — a JSON string (a step_key → output object, bounded D15; absent on a run that has produced no output yet — NULL columns are dropped, so clients must treat it as optional). Per-step steps rows carry step_key, node_type, node_label, status, attempt, request_ref, cost_micros, tokens, error, delivery_status, delivery_error, started_at, finished_at. node_label is the step's friendly builder name — the run's pinned version-graph node label, resolved on read; it is absent when the node has no label (or the version cannot be resolved), and the UI then falls back to the translated node-type name (never the raw step_key). Step status ∈ pending | running | succeeded | failed | skipped | delivering | continued. continued marks a step whose final failure was absorbed by its on_error policy (continue or error_branch) — the run proceeded; the row keeps the underlying error (and for Deliver steps the delivery_status, e.g. an absorbed blocked_pii). When a step is retried, EVERY attempt persists its own row with the same step_key and an incrementing attempt (1-based) — each attempt row carries that attempt's cost_micros/tokens, so retry cost is attributed per attempt; the step's terminal disposition is the highest-attempt row. A Deliver step is written status: "delivering" before egress, then flipped to a terminal succeeded/failed (or continued when the failure is absorbed by the step's on_error policy) once the outcome is known; its detailed outcome is in delivery_status, so render Deliver rows from that field. A row still showing delivering means the egress outcome is unknown (a crash mid-send) — such a run fails closed on resume and is never re-delivered. pii_active is an integer (1/0), written by the runner at the run's terminal/suspend boundary. 1 means the run handled PII fail-closed: it is set when any agent step did not explicitly report x-aig-pii-active: 0 — which includes agent errors and timeouts ("PII could not be ruled out this run"), matching the Deliver egress gate exactly; it does not mean "PII was provably masked". The flag is monotonic within a run (latched via GREATEST — once 1, a later terminal write can never lower it to 0, so a resume/late-fail never loses the signal). Known limitation: a run reaped mid-execution before any terminal/suspend write reads 0 (the transient signal died with the crashed tick; no egress occurred). The approver provenance (approver_ref, rendered "freigegeben von X am Y") and the typed error_code (see Error codes) are on run.

Run-detail display. The runs UI derives the shown status from error_code, not the raw status: a deliver_blocked_pii run reads "Blockiert — Datenschutz" and a budget_exceeded run reads "Gestoppt — Budget erreicht", both in a protective amber tone — these guardrail terminal states are the platform working as designed, so alarm-red is reserved for a genuine failed. Because the runner stores the trigger node's output as the raw, unmasked trigger input, the run-detail view never renders that step's output in cleartext — it shows a redaction notice instead (agent-step outputs already arrive masked and stay visible). The display never mutates the stored status/error_code; it is a read-time projection.

POST /admin/v1/gateways/{gateway_id}/workflows/runs/{run_id}/cancel → 200 { "cancelled": true }; 409 (already terminal — the CAS loser); 403 (not the run's owner, Admin, or Tenant-Admin); 404 (unknown/other-gateway run). The runner observes the cancel at its next step boundary and aborts before any further egress; every runner write is lease-fenced, so a step that completes after the cancel cannot write a phantom row.

Cost & budget (D10/D14). Per-step cost_micros (€ micros) is taken from the Agent-Step invoke's synchronous X-AIG-Cost-Micros response header (server-side pricing; the runner performs no pricing math). The run's cost_micros is the sum; request_log is the independent drill-down. If the run exceeds the workflow's run_budget_micros, the run aborts with error_code: "budget_exceeded" ("Budget erreicht").


Approval: the "Meine Freigaben" inbox

When a run reaches a Human-Approval node it is suspended (status: suspended, suspended_reason: awaiting_approval), its lease cleared (so the lease-reaper — which also excludes suspended runs — can never reap it), and the node's approver + timeout policy + the suspended step key are snapshotted onto the run. No resume token is minted (the token-link / notification path is a roadmap fast-follow); the approver decides in their in-app inbox.

The inbox routes are reachable by any authenticated user (a named approver need not be a KI-Manager) and are fenced server-side by the caller's tenant and identity — you see and decide only the runs where you are the named approver.

GET /admin/v1/me/approvals → { "approvals": [ { id, workflow_id, workflow_name, approval_step_key, approval_node_label, approval_deadline, cost_micros, created_at, requested_by_email, requested_by_name } ] } — my pending approvals (an empty inbox is { "approvals": [] }). requested_by_email / requested_by_name name the run-as-owner (workflow.created_by); both are always strings — "" when the owner is unknown or a deleted user (the UI shows "—"). approval_node_label is the approval step's friendly builder name (its version-graph node label, resolved on read); absent when the node has no label, and the UI then falls back to the translated "approval" type name — never the raw approval_step_key.

GET /admin/v1/me/approvals/{run_id} → { "approval": { … approval_node_label, run_state, requested_by_email, requested_by_name … }, "steps": [ … ] } — the confirm-page detail. Its approval carries the same approval_node_label as the inbox, and each steps row carries the per-step node_label (see Runs) so the confirm page names every produced output by its friendly step name. It is identity-scoped: a caller who is not the named approver (or the run is not in their tenant / no longer suspended) gets 404 — run_state (which may carry PII) never leaks to a non-approver.

The decision is a POST after authentication (never a GET with a side-effect):

POST /admin/v1/me/approvals/{run_id}/decision

{ "decision": "approve" }
{ "decision": "reject", "reason": "Zahlen nicht plausibel — bitte prüfen." }

decision ∈ {approve, reject} (any other value → 400). approve resumes the run (status: suspended → pending; the runner then carries it past the approval node and continues); reject fails it closed (status: failed, error_code: approval_rejected).

Reject requires an audit reason (public-sector audit trail). Accepted shape: a string of 1..1000 characters (Unicode codepoints) after trimming surrounding whitespace. It is validated at the trust boundary — after the 403/404 authz gate, before the state change — and anything else is rejected fail-closed with the run left suspended (never a partial reject): a missing/non-string or empty/whitespace-only reason → 400 reason required; invalid UTF-8 → 400 reason invalid; an over-1000-char reason → 400 reason too long. On approve, reason is ignored. The reason is stored (parameterized) in the run's error field alongside error_code: approval_rejected, so the run owner sees WHY it was rejected in the run detail.

Both approve and reject are a single atomic compare-and-swap gated on status = 'suspended' (and the caller's tenant_id + approver_id), so a second decision — or a timeout that fires concurrently — loses and receives 409 { "error": "already_decided" }. A caller who is not the named approver gets 403; a run absent from their tenant, 404. The approver provenance (approver_ref, the approver's email) is recorded on the run and rendered "freigegeben von X".

Approval timeout policy (skip | end, configured on the node; default end) is enforced by an expired-approval scan in the runner: past the deadline the run either auto-proceeds (skip — a Freigabe bypass when the approver is gone) or fails closed (end), per policy — never hangs indefinitely. The timeout apply and a human decision race on the same status = 'suspended' CAS, so exactly one wins.


Deliver node — egress is fail-closed

A Deliver node reuses the existing run-delivery egress controls (one mechanism). Its outcome is the full delivery_status enum: delivered | failed | skipped_empty | blocked_pii | blocked_ssrf | blocked_tenant. PII gate (the §1 claim): if any Agent-Step in the run masked PII (X-AIG-PII-Active: 1, or an absent header — treated as active, fail-closed), external egress is refused and the Deliver step shows blocked_pii. A Deliver that does not reach delivered hard-fails the run (error_code: deliver_blocked_pii etc.); skipped_empty (nothing to send) is a success. The delivering step row is persisted before egress, so a replay refuses to re-fire a non-idempotent delivery with an unknown outcome (no doubled egress).

Email attachment

An email-kind Deliver node may attach a workflow_artifact (e.g. the filled PDF a code node produced) via one optional config key:

  • attachment_ref — a single {{step.output.artifact.handle}} reference to a prior node's artifact handle. It must be a single template ref; a literal handle is rejected at publish (deliver_bad_attachment_ref) because a handle is a run-scoped value minted at execution time. Absent → a plain email (unchanged behaviour).

At delivery the runner resolves the ref to a handle and the loopback /deliver-email endpoint loads that artifact scoped to the current run (a node can only attach an artifact of its own run) and composes a multipart/mixed message. Accepted / rejected (fail-closed, Inv. 11):

  • the handle is validated (present-but-empty / malformed → 400 bad_handle); an unknown or foreign handle → 404 attachment_not_found; a corrupt / sha-mismatched artifact → 422 attachment_corrupt; an over-cap artifact (the ~6 MiB artifact-store cap) → 413 attachment_too_large. Any load failure hard-fails the send — a requested attachment is never silently dropped.
  • the attachment filename and content-type are sanitized before they enter the MIME headers (CR/LF/control/quote/backslash stripped) so a crafted name can never inject a header; a NULL content-type falls back to application/octet-stream.
  • a 4xx from the endpoint is a deterministic terminal outcome (failed_input, not retried); 5xx/transport stays retryable.
  • PII: on the PII-internal delivery path (a masked run delivering to the internal owner) the unmaskable binary attachment is omitted; the delivered row records attachment_omitted_pii (observable, never a silent drop). Only the message body (subject to the existing PII gate) is delivered there.

External recipient via the per-tenant allowlist

By default a deliver-email node sends to the workflow owner (resolved server-side; never from the request). An email deliver node may instead target an external recipient via one optional config key:

  • to — the external recipient: a single {{…}} template ref (resolved at runtime, e.g. {{trigger.output.recipient}}) OR a literal address. Malformed at publish → deliver_bad_recipient. Absent → the owner recipient (unchanged behaviour).

The external send is permitted ONLY if the recipient is on the tenant's workflow_recipient_allowlist — a tenant-admin-curated JSON array of allowlist entries, each an exact address (vipwing@munich-airport.de) or a domain (@munich-airport.de). It is set on the tenant via the admin tenant PATCH (a member cannot curate it), DB-backed config. Domain entries match by exact domain equality — @munich-airport.de matches x@munich-airport.de only, never a subdomain (x@evil.munich-airport.de) or suffix (x@munich-airport.de.evil.com). IDN/non-ASCII addresses are not allowlistable (fail-closed).

Fences at the loopback /deliver-email (fail-closed, Inv. 11):

  • a present-but-malformed to → 400 bad_recipient;
  • an external to on a PII-masked run → 403 recipient_pii_blocked, no send — an unmaskable body/attachment must never egress to a non-tenant address (the anti-exfiltration control is preserved; only the internal-owner relaxation applies on masked runs);
  • no allowlist configured, an unreadable/empty allowlist, or a recipient not on it → 403 recipient_not_allowed, no send (never a silent fall-back to the owner);
  • a matched external send is audited (workflow.deliver_external, with the recipient + matched entry). The subject stays server-built; a notice (blocked-PII) always goes to the owner.

An invalid allowlist entry at the admin write door → 400 recipient_allowlist_error.


Condition node — one comparator, an always-present else (D6)

A Condition routes the chain down one of two indented sub-chains. Its config is { "left_ref": "{{step.output.path}}", "op": "eq|ne|contains|gt|lt|is_empty", "right_value": "…" }. In the graph a Condition has three out-edges: a branch:"true" head, a branch:"else" head (the else branch is always present), and an optional plain "after" edge where the main chain resumes. Branch nesting is depth 1 (a branch may not contain another Condition/Loop), and the branch sub-chains never rejoin — no fan-in.

At run time the runner resolves left_ref and applies the comparator: eq/ne/ contains compare the value's string form against right_value; gt/lt are numeric (a non-numeric value or operand is a hard fail, condition_type_error); is_empty is true for "", [], or {} (the right_value is ignored). The taken branch executes; every step in the not-taken branch is marked skipped (no execution, no egress). A cjson null or a missing value on the compared path is a documented hard fail (template_null_path / template_unresolved) — never a silent else. A Condition produces no referenceable output ({{condition.output.…}} is a publish block). Replay-safe: the not-taken skipped rows are persisted before the Condition's own row, so a reaped run never re-runs a not-taken branch.

Loop-lite node — sequential, bounded, collect-all (D7)

A Loop iterates a JSON array. Its config is { "items_ref": "{{step.output.array}}" } and it has a branch:"body" sub-chain head plus a plain "after" edge. The array is resolved once and bounded before iteration one: a non-array is loop_not_array and more than 10 items is loop_too_many (rejected, never clamped). For each item (0- based) the body runs with the current item bound as {{loop.output.item}} and {{loop.output.index}}; each body step is persisted as an iteration row keyed step_key#i. The loop is collect-all: after the last item, loop.output is { "items": [ …one entry per iteration… ], "count": N }, referenceable by the after- chain. First failure fails the loop (fixed policy — the per-node on_error key described under Error handling below is not available inside loop bodies, and a body step always uses the error policy; per-step retry IS available in bodies).

A loop body must be side-effect-free: deliver and approval nodes are rejected in a loop body at publish AND run-claim (loop_body_side_effect) — because a mid-loop crash restarts the whole loop, which is only safe when the body just transforms data (a re-run re-invokes body agents, re-charging tokens, but delivers nothing twice — that egress class is designed out). Loop body outputs are transient (not persisted to run_state); only the collected {items,count} is.

Structural validation (blocks publish, re-checked at run-claim): a graph must have exactly one Trigger with no in-edges and every other node exactly one in-edge (no fan- in); at most one plain out-edge and one edge per branch slot per node; branch slots must match the container (condition → true/else, loop → body); branch depth ≤ 1; every node reachable from the Trigger; step keys [A-Za-z_][A-Za-z0-9_]*, ≤ 60 chars (so a #i iteration suffix fits). A {{}} ref to an output-less node (Condition, Deliver) is a block.


Error handling — per-step retry and on_error

Agent and Deliver nodes accept two optional config keys (absent = today's behavior, so existing graphs are untouched):

{ "retry":    { "max_attempts": 3, "backoff_seconds": 10 },
  "on_error": "error_branch" }
  • retry.max_attempts — integer 1..5 (default 1 = no retry). retry.backoff_seconds — number 0..30 (default 5), a fixed delay between attempts; additionally (max_attempts - 1) * backoff_seconds must be <= 120. Out-of-range values are a publish block (bad_retry_config).
  • Only known-transient failures retry, fail-closed: an Agent step retries on a transport error or HTTP 5xx/429; a Deliver step retries only a delivery_status: "failed" outcome. Policy blocks (blocked_pii / blocked_ssrf / blocked_tenant) and other 4xx are permanent and never retried; every Deliver retry re-runs the full fail-closed gate chain (PII, SSRF pinning, tenant allowlist). Deliver retry is at-least-once: a failed outcome can be a timeout after the request was already received, so a retry can deliver twice — make webhook receivers idempotent. (failed also covers permanent causes, e.g. an invalid channel id, which retry pointlessly — bounded by max_attempts.)
  • Each attempt persists its own step row (attempt 1-based) with that attempt's cost/tokens; the lease is re-extended around every attempt and backoff sleep. Retries count against the run budget, and the run deadline/cancel/budget are checked before every attempt — under deadline pressure absorption is best-effort (a run_deadline/budget_exceeded abort wins over the policy).

on_error — what happens when the FINAL attempt fails ("error" default):

Policy Behavior
error the run fails with the step's error code — exactly today.
continue the step's final row becomes continued (failure absorbed, underlying error/delivery_status preserved); the run proceeds. The step produces no output, so {{}} references to it are publish-blocked (ref_to_fallible_step). Not allowed inside loop bodies (on_error_in_loop_body).
error_branch the node carries a second out-edge (branch: "error" — the catch chain). On final failure the entire after-chain is persisted skipped, the step becomes continued, and the catch chain runs; on success the catch chain is persisted skipped and never runs. Catch chains are linear (no Condition/Loop/nested error edges — branch_depth_exceeded); references to the failing step from its own catch chain are publish-blocked (ref_to_failed_source), while after-chain references stay legal (the after-chain only runs on success). Requires exactly one error edge (error_branch_missing_edge / error_edge_without_policy); not allowed inside containers (error_branch_in_branch).

A run whose failures were all absorbed finishes succeeded — the step rows carry the truth (the runs UI shows per-attempt rows, the continued disposition, and any preserved block outcome; a blocked_pii absorbed by continue still renders the PII block on its step row).

Publish-time validation codes added by this feature: bad_retry_config, bad_on_error, error_branch_missing_edge, error_edge_without_policy, error_branch_in_branch, on_error_in_loop_body, ref_to_fallible_step, ref_to_failed_source (all re-checked at run-claim like every other rule).

Error codes

The run's typed error_code drives a translated message in the runs UI via the workflow.runErrors.* i18n namespace (kept separate from the builder's validation workflow.errors.* codes to avoid a key collision), with a generic fallback for an unrecognised code; the raw English error string is shown only in an expandable technical-details view (never verbatim on the main surface).

error_code Meaning
lease_expired the runner died mid-run; the lease reaper marked it failed (visible, never eternal "running"). Manual replay from the run detail.
budget_exceeded the run hit run_budget_micros; aborted at a step boundary.
graph_invalid the bound version failed structural validation at run-claim (e.g. an agent's schema drifted, or a cycle). Re-validation happens at claim, before any step runs.
agent_not_owned an Agent-Step references an agent the workflow owner no longer owns (or that was deleted). Enforced at publish AND re-checked at run-claim — the run fails closed before any invoke, because the per-gateway service token could otherwise invoke any agent.
agent_failed an Agent-Step invoke returned a non-2xx at runtime (the agent exists and is owned, but the model/provider call failed) and the step's retry/on_error policy did not absorb it. (Independent of configured retries, a single in-attempt retry fires on a 422 structured_output_failed response — stochastic and idempotent.)
run_deadline the run exceeded its wall-clock cap at a step boundary.
template_unresolved a {{step.output.path}} reference could not be resolved (missing/unfinished step) — a hard step error, never a silent empty string.
template_null_path a resolved path traversed a JSON null — documented hard fail (a Condition's compared path or any {{}} ref).
condition_type_error a Condition gt/lt comparison hit a non-numeric value or operand — fail-closed, never a silent else.
loop_not_array a Loop's items_ref did not resolve to a JSON array — rejected before iteration one.
loop_too_many a Loop's array exceeded the 10-iteration cap — rejected before iteration one (no silent clamp).
step_output_too_large a step's output exceeded the run_state cap (256 KB/run, 32 KB/step).
deliver_blocked_pii / deliver_blocked_ssrf / deliver_blocked_tenant / deliver_failed a Deliver node did not reach delivered.
schedule_deleted / schedule_disabled a still-pending schedule-triggered run was cancelled because its schedule was removed or disabled (kill-switch parity — no orphaned pending fire).
webhook_trigger_deleted a still-pending webhook-triggered run was cancelled because its webhook trigger was deleted (the same kill-switch parity as the schedule codes above).
fetch_failed a Fetch node's download failed transiently (network / timeout / redirect-loop / upstream 5xx / 429) and the step's retry/on_error policy did not absorb it. Retryable (mapped status 502).
fetch_blocked a Fetch node's URL was refused by the SSRF pin (private/link-local/loopback/metadata target, or an unsafe redirect). Deterministic — never retried.
fetch_http_<status> a Fetch node's upstream returned a non-2xx. A 4xx (except 429) is deterministic (not retried); a 5xx/429 is retryable.
fetch_too_large a Fetch node's body exceeded the fetch/artifact cap (~6 MiB). Deterministic.
fetch_empty a Fetch node got a 2xx with an empty body. Deterministic.
content_type_rejected / content_type_not_allowed a Fetch node's bytes failed the magic-authoritative content sniff, or the sniffed canonical MIME was not in the node's allowed_content_types. Deterministic.
egress_blocked_residency a Fetch or Connector node was refused because the gateway is residency-restricted — EU-residency enforced (per-gateway flag, tenant floor, or the deployment env) or all providers local. The node cannot egress to a third party at all. Deterministic (never retried); no call/download occurs.
url_pii_blocked a Fetch node's resolved URL carried protected structured PII (a checksum-valid national ID, IBAN, card number, or a configured keyword) on a PII-active / PII-masking-enforced gateway. A URL cannot be tokenized and stay fetchable, so the policy is refuse, not scrub. Deterministic; no download occurs.
gateway_unresolved a Fetch node's gateway could not be resolved (soft-deleted or corrupt config). Deterministic fail-closed (no egress). A transient config read error instead returns a retryable db.
code_disabled a Code node ran on a gateway where the code-interpreter is not fleet-configured, or has been explicitly disabled (code_interpreter.enabled: false). Deterministic (501) — absorbable by on_error, never a transient tick-retry.
code_unavailable a Code node's sandbox call reached the broker but it was unavailable/saturated (no-isolation / busy / transport). Retryable (mapped status 502).
code_too_large a Code node's job exceeded the sandbox pre-flight size cap (the encoded {code,stdin,limits} body). Deterministic.
code_failed / code_timeout a Code node's program exited non-zero (or with an unknown exit) / hit the wall-clock limit. Deterministic. The response carries the exit_code only — never stderr (it would persist in workflow_step_run.error).
code_no_output / code_no_usable_output / code_ambiguous_output a Code node produced no /output artifact / only harvest-rejected outputs / more than one with no output_filename to disambiguate. Deterministic.
code_duplicate_input / input_not_found two input_refs resolved to the same /input filename / an input_refs handle is not an artifact of this run. Deterministic (422 / 404).
code_reserved_input an input_refs artifact's stored filename is the reserved _context.json (the always-staged trigger-context slot). Deterministic (422) — rejected before the run so a step cannot spoof/clobber the context.
code_output_too_large a Code node's /output exceeded the ~6 MiB artifact store cap. Deterministic.
connector_missing_field a Connector node's connector_id, tool, or input was absent or empty (publish-time validation; also the runtime bad_* guard). Deterministic.
connector_bad_input a Connector node's resolved input did not decode to a JSON object — a scalar, a non-empty array, or undecodable JSON (ABSENT≠MALFORMED: a value that arrives wrong is rejected, never degraded to "no args"). Deterministic; no call occurs.
connector_pii_blocked a Connector node's arguments carried protected structured PII (checksum ID / IBAN / card / configured keyword) on a PII-active / PII-masking-enforced gateway and the connector is not a trusted_subprocessor. Refuse, not scrub. Deterministic; no call occurs.
connector_gateway_unresolved a Connector node's gateway could not be resolved (soft-deleted or corrupt config). Deterministic fail-closed (no call). A transient config read error instead returns a retryable db.
connector_call_failed a Connector node's tools/call failed — a definitive error (not_found / tool_blocked / credential_required / credential_expired / init_error) or an AMBIGUOUS upstream outcome (timeout / reset / 5xx-after-send). Always terminal, NEVER auto-retried (D1 — a mutating call whose outcome is unknown must not re-fire). A crash mid-call instead surfaces the shared deliver_unknown_outcome on resume (the resume reap is node-agnostic).

Fetch node — SSRF-pinned download

A Fetch node downloads a URL to a binary workflow artifact (the /artifact store) and emits its handle ({{fetch_key.output.artifact}}) for a later node to consume. Because the Python runner has no HTTP-fetch and no SSRF pin, the download runs server-side in the gateway (the loopback /fetch endpoint, below) — making the gateway the single, audited egress authority for workflow downloads. The trust-boundary contract (Invariant 11):

  • url is untrusted: the SSRF pin (connection_plan/resolve_pinned_ip) fails closed — private/link-local/loopback/cloud-metadata targets, userinfo, control bytes, DNS-rebinding, and unsafe redirects are all blocked (fetch_blocked). The pin is re-validated on every redirect hop.
  • The fetched bytes are untrusted third-party content: the artifact's MIME is decided by a magic-byte sniff of the bytes (upload_sniff), never the response Content-Type header — a spoofed or absent header cannot smuggle a disallowed type past the gate. A node's optional allowed_content_types is enforced against the canonical (sniffed) MIME; a malformed allowed_content_types is rejected, never treated as "allow any".
  • Caps: the body is capped at max_bytes (clamped to the ~6 MiB artifact cap); an oversize body is fetch_too_large. A 2xx empty body is fetch_empty. Both deterministic.
  • Retry: transient failures (fetch_failed, upstream 5xx/429) retry per the node's retry config or take its error_branch; deterministic failures do not.
  • Audit: every egress attempt (blocked, failed, rejected, or stored) writes an audit_log row (action=workflow.fetch) with the authority only (scheme://host — never the path or query, which can carry a token) plus a SHA-256 of the full URL for correlation, the byte size, and the outcome. Best-effort (an audit failure never fails the fetch).
  • Residency (fail-closed): on a residency-restricted gateway the Fetch node is refused before any download (egress_blocked_residency). "Restricted" means the gateway is EU-residency enforced (the per-gateway eu_region_routing flag, the tenant residency floor, or the deployment-wide enforcement) or every configured model provider is local — and, fail-closed, any gateway whose residency posture cannot be positively determined. A workflow run has no conversation/project context, so this binds to the gateway's own residency posture. (Code nodes are unaffected: their sandbox is network-isolated — no third-party egress on any gateway.)
  • Structured-PII in the URL (fail-closed): on a PII-active or PII-masking-enforced gateway, a resolved url carrying protected structured PII (national ID, IBAN, card number, or a configured keyword) is refused (url_pii_blocked) — mirroring the chat fetch tool and the Deliver node. A non-UTF-8 URL is rejected up front (bad_url). Free-text names pass.
  • Egress reach: the fetch shares the process-wide operator egress allowlist (empty in int/prod → fully fail-closed). An allowlisted host short-circuits the private-IP block and is honored on redirect hops — set it only for operator-vouched hosts. A dedicated per-tenant workflow-egress allowlist is a planned hardening.

Code node — headless sandbox run

A Code node runs author-supplied Python in the network-isolated code-interpreter sandbox (--network none KVM micro-VM; pypdf/reportlab/openpyxl are available) over prior nodes' artifacts, and stores the program's /output as a new workflow artifact ({{code_key.output.artifact}}). The sandbox call runs server-side in the gateway (the loopback /code endpoint, below), reusing the same ctx-free broker the chat code-interpreter uses. The trust-boundary contract (Invariant 11):

  • Enablement (opt-out; AGF-2950): the code-interpreter is on by default, so the node runs whenever it is fleet-configured and the run's gateway has not opted out (code_interpreter.enabled != false; absent / null → on). Only an explicit code_interpreter.enabled: false yields code_disabled (501, no sandbox call).
  • code is OPAQUE: the source is untrusted but runs only in the isolated sandbox (never in the gateway) — isolation is the control; it is size-bounded (≤ 64 KiB; the broker re-checks the encoded job against its 128 KiB cap → code_too_large). It is never template-resolved: a {{…}}-shaped token in the code is literal Python, so run data can never be spliced into executed source (a code-injection seam) and legit Python containing {{ (an f-string f"{{x}}", a set literal {{1,2}}) publishes clean. Run data reaches the program only as data, via /input/_context.json and input_refs.
  • /input/_context.json — the trigger context: every Code node's sandbox is staged with a /input/_context.json file carrying the run's consumed trigger output as a JSON object. Read it with json.load(open("/input/_context.json")). For a form run it is the validated {field_key: value} submission — and a file field carries the extracted, PII-masked text (identical to {{trigger.output.<field>}}), never the upload sentinel or its single-use submission_token. For a webhook run it is the webhook JSON payload. Every submitted trigger field lands in the sandbox (not only referenced ones). Values are data, not source — quotes, newlines, backslashes and non-ASCII round-trip because they are JSON-encoded. The file is always present and always a JSON object: a manual run with no input, or a trigger output that is absent / not an object / a JSON array, degrades fail-closed to {} (so json.load(...).get(...) never raises) — a data-only empty object, never a control bypass. A pathologically deep trigger payload that cannot be re-encoded also degrades to {} (fail-closed).
  • _context.json is a RESERVED name: an input_refs artifact whose stored filename is _context.json is rejected before the run (422 code_reserved_input) — a prior or malicious step cannot spoof or clobber the trigger context.
  • input_refs are artifact handles of this run: each is loaded run-scoped + lease-fenced + SHA-verified (a foreign/unknown handle → input_not_found, never a cross-run/tenant read), capped at 7 files (the sandbox holds 8; one slot is reserved for /input/_context.json) — enforced at both publish (code_bad_input_refs) and runtime (400 bad_input_refs), so there is no publish-clean/runtime-fail drift. They are staged as /input/<stored-filename>. Duplicate /input names are rejected (code_duplicate_input) before the run. The context file counts toward the shared 16 MiB input-files total, so a code node already staging near 16 MiB of input_refs can newly trip the total-size cap (code_bad_input).
  • Output: the /output artifact is magic-gated + capped by the sandbox harvest, then re-capped at the ~6 MiB artifact store. Selected by output_filename (or the sole output's own name); bytes come from the artifact's text (csv/txt/json) or bytes (pdf/xlsx/png) by kind. Zero/ambiguous output is a deterministic code_no_output / code_ambiguous_output.
  • Classification: a broker call that reaches the wire but fails (no-isolation / busy / transport) is the only retryable class (code_unavailable, 502). A pre-flight reject (never reached the sandbox — oversized job, unsafe input) and a program failure (code_failed / code_timeout, from a successful broker call with a non-zero/timeout result) are terminal.
  • Audit: every run attempt writes an audit_log row (action=workflow.code) with the input handle count, output size, and outcome — never the code source or stderr (they would persist and can embed secrets). Best-effort.

Connector node — call an MCP/API connector mid-run

A Connector node calls an existing MCP/API connector's tools/call mid-run and maps the (PII-scrubbed, size-capped) result into run_state as {{connector_key.output.result}} for a later step. The Python runner has no MCP/credential/SSRF capability, so the call runs server-side in the gateway (the loopback /connector endpoint) — the same single, audited egress authority as Fetch. The trust-boundary contract (Invariant 11):

  • Run-as-owner. The call is made as the workflow owner (the workflow's created_by), using the owner's credential and resolving the owner's private connector. A cross-user private connector is invisible (not_found → connector_call_failed); a tenant-shared connector (created_by NULL) is tenant-fenced, not owner-fenced, so any author in the tenant may bind it. Trigger-supplied data (e.g. an inbound email) that flows through input into the owner's connector + credential is an accepted run-as-owner trust boundary, governed by the PII/residency gates below.
  • input is untrusted. The resolved input must decode to a JSON object (or empty). A scalar, a non-empty array, or undecodable JSON is rejected (connector_bad_input) — ABSENT and MALFORMED do not share a permissive path; a value that arrives wrong is rejected, never degraded to "no args". A non-UTF-8 input is rejected up front (bad_input).
  • Residency (fail-closed). On a residency-restricted gateway (EU-residency enforced, all providers local, or an indeterminate posture) the node is refused before any call (egress_blocked_residency). A workflow run has no conversation/project context, so this binds to the gateway's own residency posture.
  • Structured-PII in the arguments (fail-closed). On a PII-active or PII-masking-enforced gateway the arguments are scanned; protected structured PII (national ID, IBAN, card number, or a configured keyword) is refused (connector_pii_blocked), not scrubbed (a token would break the tool) — unless the connector is marked a trusted_subprocessor (a vetted EU/DPA subprocessor; the per-connector exemption set via the four-eyes setter — see MCP connectors). Free-text names pass. The same mcp_egress authority governs the chat agent tool path.
  • Result is scrubbed + capped. The result is stringified, pre-capped to 20 KiB, then run through the strict persistence-leg PII fence (the same masker the form-file consume uses): on a PII-active gateway without a fail-closed masker, or under an untokenizable mandate, the content is withheld behind a safe placeholder — the raw result is never persisted into run_state.
  • No retry — ambiguity fails closed (D1). The connector node is not a retry node: it accepts no retry/on_error/error-edge (rejected at publish — bad_retry_config / bad_on_error / invalid_branch_slot) and may not sit inside a loop body (loop_body_side_effect). Because the MCP client collapses timeout / connection-reset / 5xx-after-send into one ambiguous upstream outcome with no pre-send-vs-post-send signal, a mutating call that ends ambiguous must not re-fire: every failure (definitive or ambiguous) is terminal (connector_call_failed). A crash mid-call leaves a delivering step row, which the node-agnostic resume reap fails closed as deliver_unknown_outcome (never a second call).
  • Audit. Every call attempt (every outcome) writes an audit_log row (action=workflow.connector) with the connector_id, tool, and outcome only — never the arguments or the result (both can carry the tenant's PII). Best-effort (an audit failure never fails the call).

Worked example — download, fill, and email a PDF form

The Fetch, Code, and Deliver nodes compose into the document-fill workflow the epic exists to enable: download a blank PDF form, fill it with the requester's data in the sandbox, and email the completed form as an attachment to an external recipient — behind a human approval. The builder exposes Fetch and Code in the node palette, so this is authored end to end in the UI, not only via the API.

The graph is trigger → fetch → code → approval → deliver:

  1. Fetch downloads the blank form (url, allowed_content_types: ["application/pdf"]) → {{f1.output.artifact}}.
  2. Code stages that artifact at /input/<name>, runs Python (pypdf) that fills the AcroForm text, checkbox, and radio fields from an embedded/parameterised profile, and writes /output/<output_filename> → {{fill.output.artifact}}. The fill sets both each field's value and every widget's /AS appearance state across all pages — a checkbox or radio widget can live on any page, and a stale /AS on a second widget silently defeats the fill.
  3. Approval suspends the run for a named approver to review before anything leaves the tenant.
  4. Deliver (kind: email) attaches the filled form via attachment_ref: "{{fill.output.artifact.handle}}" and sends it to an external to address — permitted only because the recipient's domain is on the tenant's workflow_recipient_allowlist (see External recipient). On a PII-masked run the external send is blocked fail-closed (recipient_pii_blocked).

A committed fixture of this workflow (a small AcroForm PDF, the fill program, and the graph) lives at tests/fixtures/vipwing/ and is driven end to end by case_vipwing_reservation in scripts/workflow_e2e.py (bash dev-instance.sh workflow-e2e).


Template resolution ({{ }}) — trust boundary (Invariant 11)

Data flows between steps via {{step_key.output.path}} references, inserted by the builder's picker (chips, not free-typed handlebars). Resolution is single-pass over graph-authored templates only: step outputs and trigger payloads are opaque data and are never re-scanned — a {{ }} sequence appearing inside a webhook payload or a model response is not expanded (closes the second-pass injection class). A missing or unfinished referenced step is a hard step error (template_unresolved), never a silent empty substitution; a cjson.null anywhere on the path is a documented hard fail (template_null_path). References into Condition/Loop branch members from outside the branch are impossible by construction — the publish-time ancestor rule (dangling_ref) rejects them, so no runtime "skipped reference" case exists. References to a step whose on_error is continue are likewise publish-blocked (ref_to_fallible_step); references to an error_branch step are legal from its after-chain (which only runs when the step succeeded) and publish-blocked from its own catch chain (ref_to_failed_source).


Internal runner seam (not public)

The runner (scripts/workflow_runner.py) holds no DB credentials; it drives a loopback-only internal API (127.0.0.1:8083, location ~ ^/internal/workflow/), authenticated by the shared X-AIG-Scheduler-Secret. Endpoints: /reap, /due-runs, /claim, /load-run, /extend-lease, /save-step, /step-state, /step-cost, /finish-run, /suspend, /due-approvals, /apply-approval-timeout, /set-delivery, /artifact, /artifact-get, /fetch, /code. Every runner-side mutation is lease-token fenced (claimed_by = token AND status = 'running'); a 409 { "error": "lease_lost" } tells the runner it was reaped and must abort with no further egress or overwrite. This surface is unreachable off-container and never called from a browser.

Although loopback-only and secret-authenticated, the runner is still treated as an untrusted-shape client (Invariant 11): every field is validated before any storage call. The pii_active field on /finish-run and /suspend is accepted only as an integer 1/0; any other shape is coerced fail-closed to 0 — a missing field, a JSON boolean (tonumber(true) is nil), or a non-numeric string all become 0. The value is never trusted to lower the stored flag: storage latches it with GREATEST(pii_active, ?), so a 0 can only ever be a no-op.

/save-step persists a workflow_step_run row, whose tenant_id scopes GDPR export (list_workflow_step_runs_for_export) and the erasure scrub. That tenant_id is DERIVED from the lease-fenced parent run, never from the body: the fence SELECT returns the run's tenant, and the INSERT stamps that value. A body tenant_id that is present and different from the run's tenant is a runner bug and is rejected 400 { "error": "tenant_mismatch" } with no write (fail-closed, loud); an absent or JSON-null tenant_id is treated as absent and derives silently; a run row whose own tenant is empty/NULL (schema-impossible corruption) is rejected 500 { "error": "db" } (run_tenant_missing), also with no write. Every field on the row-writing endpoints — /save-step, /set-delivery, /finish-run, /suspend — is validated column-true before the write: integers must be finite and inside the real column range (tokens [0, 2³¹); cost_micros/tokens_used/started_at/finished_at [0, 2⁵³) — a present-but-invalid value is 400, never a silent floor/zero), identifier strings must be valid utf8mb4 within the column width (step_key/request_ref/ approval_step_key ≤ 64, node_type/status/delivery_status ≤ 16, suspended_reason ≤ 32, approver_id ≤ 36 codepoints, else 400), and free-text error/delivery_error are clamped to 16000 codepoints while error_code clamps to 48 (invalid utf8mb4 → 400 on all of them). On /save-step and /set-delivery an empty-string or JSON-null optional field is stored as SQL NULL, not ''. So no client shape can mint a 5xx from the write.

/artifact and /artifact-get — binary artifact store

A workflow node produces a binary artifact (e.g. a fetched or filled document) that a later node consumes by handle, without ever putting bytes in run_state (JSON, 256 KB cap). Bytes live in exactly one row (workflow_artifact.data, LONGBLOB, FK → workflow_run ON DELETE CASCADE); run_state holds only a reference {"artifact": {"handle", "filename", "mime_type", "size_bytes", "sha256"}}. Both endpoints are lease-token fenced like every other runner mutation.

POST /artifact — {run_id, token, step_key, filename, mime_type, data_b64} → {handle, size_bytes, sha256}. Accepted/rejected shape (Invariant 11): - data_b64 is read raw and type()-checked, then base64-decoded, with the emptiness check on the decoded bytes (not the raw string, because a base64 decode of a padding-only string like "====" yields "", not nil): a non-string field → 400 missing_artifact_data; a decode == nil → 400 bad_artifact_b64; a decoded length of 0 (covering "", "=", "==", "====") → 400 empty_artifact. No 0-byte or wrong-size row is ever stored; size_bytes is the decoded length. - filename is basename-stripped (path separators removed) then re-checked empty ("../", "/", ".", ".." → 400), and filename/mime_type/step_key are utf8mb4-validated within their column widths (else 400, never a store-time 5xx). Note for the consuming node: basename does NOT strip NUL/control bytes — the artifact is stored opaque here, so the node that stages it at /input/<name> MUST sanitise NUL/control at that boundary before it reaches a C-based path handler. - tenant_id is DERIVED from the lease-fenced parent run, never from the body. The fence is an INSERT … SELECT … FROM workflow_run WHERE id = ? AND claimed_by = ? AND status = 'running'; rowcount 0 (reaped lease, or unknown run — no existence oracle) → 409 lease_lost with no write. - Per-artifact and per-run aggregate size caps are enforced before the insert (413 artifact_too_large / artifact_run_total_too_large); a MySQL max-packet refusal is mapped to the same 413, never a 500.

POST /artifact-get — {run_id, token, handle} → {filename, mime_type, size_bytes, sha256, data_b64}. Lease-fenced (lost → 409); the row lookup is scoped WHERE id = ? AND run_id = ?, so a handle belonging to another run/tenant → 404 not_found (a valid-lease holder cannot distinguish foreign-exists from not-exists). content_sha256 is recomputed over the stored bytes and compared on read — a mismatch fails closed (sha_mismatch, row withheld).

/fetch — SSRF-pinned server-side download

POST /fetch — {run_id, token, step_key, url, allowed_content_types?, max_bytes?, filename?} → {handle, size_bytes, sha256, content_type, filename}. The Fetch node's server-side worker (see the Fetch node section above for the full contract). Order, all fail-closed: validate shape (allowed_content_types an explicit array-of-strings — a malformed value → 400 bad_allowed_types, never allow-all; max_bytes clamped to the artifact cap; filename basename-sanitised, malformed → 400 bad_filename) → lease-fence the run FIRST (load_workflow_run; lost → 409, no egress) → fetch_bytes (SSRF-pinned) → audit the egress attempt (authority-only, no path/query) → empty-body → 422 → magic content sniff (415 on reject / not-allowed) → store via create_workflow_artifact (409 lost lease, 413 cap). Status map: SSRF 400; network/timeout/ redirect/upstream-5xx/429 502 (retryable); upstream other-4xx & empty 422; oversize 413. The runner's HttpWorkflowDb.fetch treats a 502 as retryable and every other non-2xx as a terminal FetchError; a 409 is a lost lease (abort).

/code — headless sandbox run

POST /code — {run_id, token, step_key, code, input_refs?, output_filename?, context?} → {handle, size_bytes, sha256, mime_type, filename}. The Code node's server-side worker (see the Code node section above for the full contract). context is an internal loopback field (not a persisted contract): the runner always sends the run's consumed trigger output (state[trigger_key]) as a nested JSON object — code is passed verbatim (never template-resolved). Order, all fail-closed: validate shape (code non-empty ≤ 64 KiB; input_refs an explicit array-of-strings capped at 7 — #input_refs + 1 > 8 reserves the /input/_context.json slot → 400 bad_input_refs; output_filename basename-sanitised) → lease-fence the run FIRST (load_workflow_run; lost → 409) → enablement gate (is_configured() + resolve_gateway_config_by_id + config_code_interpreter_enabled; disabled → 501 code_disabled, no sandbox call) → load each input_ref (get_workflow_artifact, run-scoped; unknown → 404 input_not_found) + reject the reserved name _context.json (422 code_reserved_input) + reject duplicate /input names (422 code_duplicate_input) → stage /input/_context.json from context (object-or-empty guard: a non-object / JSON array / un-encodable value → the literal {}, fail-closed) → run_code (ctx-free broker; artifact_max_bytes = the store cap) → classify by return shape: (nil,err,outcome≠nil) broker failure → 502 code_unavailable (retryable), (nil,err,nil-outcome) pre-flight → terminal 413 code_too_large / 422 code_bad_input, a RunResult with timed_out → 422 code_timeout / exit_code≠0|nil → 422 code_failed (exit_code only, no stderr) → select the output by art.name (bytes by kind) → store via create_workflow_artifact (409 lost lease, 413 code_output_too_large) → audit (metadata only). The runner's HttpWorkflowDb.code maps only 500 → TransportError (transient tick-retry); everything else — including 501 code_disabled — is a terminal CodeError (so on_error/error_branch can absorb it); 502 is node-retryable; 409 is a lost lease (abort).