Skip to content

Projects API

The projects API manages chat projects, project membership, and project knowledge attachments.

All endpoints require an authenticated admin session (aig_admin cookie). The base URL is https://ai-api-admin.myra.eu/admin/v1.

Project-level access is governed by per-project membership: owner, editor, viewer.


Listing projects

GET /admin/v1/projects

Returns the projects the calling user can access. A non-admin caller is scoped to the projects in their own tenant. An admin-role caller must pass ?tenant_id=<ID> to select the tenant; the request returns 400 if it is omitted.

Each project card's knowledge_count is source-ACL filtered for the calling user — it counts only the knowledge files that user is permitted to read at the source, so it always matches the (also filtered) knowledge list on the project detail. A member of a project that holds creator-only connector-synced documents therefore never sees a card count larger than the files they can actually open. (Admins are filtered too, with no bypass, consistent with the detail list.)

Required role: authenticated.


Creating a project

POST /admin/v1/projects

Required role: any authenticated role except viewer. The caller is auto-added as a project member with the owner role.

Four-eyes config approval. When the organization has config approval enabled and the caller is a Fachadmin (tenant_admin / ki_manager), a project create/update/delete does not apply immediately — it returns 202 Accepted with { "status": "pending_approval", "approval_id": "…" } and is held until a second, different administrator approves it. At approval the change is re-validated against current state (the same offer validation below, and the requester's project role — including this update's owner-only tier/retention gate) — never a stale replay. With config approval off (the default) these calls apply immediately, exactly as documented here.

Field Type Required
name string yes
description string no
instructions string no
icon string no
color string no — the project accent shown on the project banner, welcome tile, and folder icon. Accepted shape: a #RRGGBB hex string exactly (e.g. #2563eb); send JSON null to reset it to the default #2563eb (the column always holds a colour). Any other value — CSS keywords, 3-digit hex, url(...) strings, or a non-string type — returns 400 invalid color (the value feeds CSS custom properties in the app, so it is validated fail-closed at the write boundary). In dark mode the app renders an automatically lightened variant of the accent where it acts as a link/border colour, so it stays readable on dark surfaces; the stored value is unchanged.
permission_tier user_configurable | pii_mandatory | local_only no — defaults to user_configurable. An unknown value returns 400.
default_gateway_id string no — optional gateway pin. Usually left null: the preferred model (below) routes on a gateway the server derives, so callers need not choose one.
default_model string | null no — the project's preferred/default model. A soft preselect: new conversations in the project open with this model chosen; a member can still switch model per chat. Send null to clear. Set stand-alone (model only, default_gateway_id omitted or null) — the routing gateway is derived from the model. When both default_gateway_id and default_model are set, every project conversation routes to that pair; see the offer validation below.
tenant_id string required for admin callers placing the project in a different tenant; ignored for non-admin callers (the caller's tenant is used).

The permission_tier is the project's access tier: local_only restricts the project to local Myra models (no cloud egress), pii_mandatory forces PII masking on — and, on a gateway with no reversible masker, withholds the third-party tool egress the masking cannot cover (the web search; a fetched document bound for a non-local inner agent model) and refuses a URL or MCP argument carrying protected structured PII — and user_configurable (the default) leaves the PII choice per chat with all models available.

Project default pin offer validation. When a request sets both default_gateway_id and default_model, the server validates that pair against the pinned gateway's dispatch gates — the same EU data-residency and provider-allowlist checks the inference path enforces — evaluated on the gateway's resolved (tenant-floor-folded) config. A pin the gateway would refuse at dispatch is rejected here rather than accepted and then failing every conversation:

  • The model resolves to a provider the gateway's EU data-residency enforcement blocks → 400 with code: "data_residency_blocked".
  • The model resolves to a provider not on the gateway's approved provider allowlist → 400 with code: "provider_not_allowed".

Local-only access tier — the default must be a local model. A project whose permission_tier is local_only may only route to local Myra models, so its stored default must resolve to a local provider. This is checked at save time for both pin shapes (unlike the residency / allowlist check above, which is both-pin only):

  • A both-pinned default whose resolved dispatch provider is not local → 400 with code: "model_not_allowed_for_project". Without this the pin would be honoured verbatim and then refused on every conversation with no Auto fallback — a hard dead-end.
  • A model-only default whose model is served only by cloud providers (no local model_price provider) → 400 with the same code. The model would otherwise be stored but never honoured (the resolver falls back to a local Auto pick).
  • Changing a project's access tier to local_only while it still holds a non-local default is rejected the same way — clear or replace the default first.

An unknown model (no model_price row) is tolerated on a local_only project: it is not provably cloud, and the resolver simply falls back to the local Auto pick. A model served by at least one local provider is accepted even if a cloud provider also serves it (it can run locally). pii_mandatory and user_configurable projects are not constrained by this check — they may pin a cloud model.

A pin with only one of the two fields set is not offer-validated (a lone pin falls back to the normal Auto pick, which is itself offer-filtered). If the pinned gateway cannot be resolved to a folded config at save time (a transient database error), the save fails closed with a retryable 5xx rather than silently storing an unvalidated pin.

Model-only preferred model (soft default). Setting default_model alone is the common case (the UI stores the model only). On a user_configurable or pii_mandatory project it is accepted without offer-validation — an unroutable value is not rejected; it simply has no effect (a conversation whose preferred model cannot be routed falls back to the normal Auto pick). On a local_only project it IS validated (see Local-only access tier above): a model-only default served only by cloud providers is rejected 400. On a self-serve workspace a created or changed model-only default is also judged against the plan and the EU-Gov add-on: a model outside the plan → 400 plan_model_not_allowed; a Myra-hosted model without EU-Gov → 400 eu_gov_model_not_allowed (an unchanged stored default stays editable). Accepted shape: any string (subject to the local-only rule), or null to clear. When the preferred model can be routed, the server derives a gateway that serves it (offer-filtered, PII-twin-aware) both for the composer's preselect and for conversation routing, so callers never have to understand gateways. Saving a preferred model normalises the row to model-only (any previously stored default_gateway_id is cleared).

A local preferred model is routable like any other: the local fleet needs no provider key and is served by every gateway, so pinning one now takes effect (it was previously accepted and then silently ignored). When the request wants or requires personal-data protection, the server derives a gateway that actually masks personal data and serves the pinned model; if no such gateway serves it, the pin is still honoured on another gateway and the normal PII-twin selection applies. On a subscription plan with a model list, a preferred model outside that list is not adopted.


Getting a project

GET /admin/v1/projects/<ID>

Required role: project member or admin.

Who can see the project. The response carries three parallel access views so "who has access?" is complete: members (direct chat_project_member rows), groups (user groups granted access), and org_access — the organization-wide grant. org_access is { "role": "editor" | "viewer" } only when an org grant resolves to real access (a grant row exists and the tenant capability org_share_enabled is on); it is absent otherwise, including when a grant row exists but the capability is off (fail closed — see the org-share section in groups, which also documents the raw org_grant toggle field). member_count (and the members array length) counts named (direct) members only — it deliberately does not fold in an org-wide grant, since org size is not a project fact; the org grant is surfaced explicitly via org_access instead.

The response's knowledge list is source-ACL filtered, identical to retrieval: a connector-synced document is included only if the requesting user is permitted to read it at the source (SharePoint per-item ACL, or a Confluence source the user created or a tenant admin published org-wide). A member — or an admin — never sees the filename of a creator-only synced document they cannot open, matching the file browser, item preview, download, and chat retrieval. Manual uploads are always listed to project members. See connector sync — access control.

write_file filenames — accepted and rejected shapes. The chat write_file tool stores text only. Accepted: any validated basename (no path separators, no control bytes, ≤ 255 chars, no leading dot) whose extension is a text/code format or a renderable document format (.docx, .pdf, .pptx, .xlsx, .odt, .odp, .ods) — document-named text is rendered into the real binary at download time via the /chat/export-* pipeline, so the name's promise is kept. Rejected (tool error steering the model to a renderable or plain-text name): binary extensions with no renderer — legacy Office (.doc/.xls/.ppt), archives (.zip/.gz/.tar/.7z/.rar), images (.png/.jpg/.gif/…), executables/libraries (.exe/.dll/.so/.wasm/.jar), audio/video, and fonts. A rejected write persists nothing.

Chat-generated files follow the source chat's share gate. A file produced by a chat tool (write_file / the code interpreter) inside a project chat is stored as project knowledge, but it is not auto-published to the other members. Like the conversation that produced it, it is visible only to its creator until that conversation is shared into the project (shared_in_project); once the chat is shared, its generated files become visible to all members (and un-sharing hides them again). This applies to every project-knowledge surface — the file list, item preview, download, the card count, and chat/inference retrieval. Manual uploads and admin-added notes are unaffected (always member-visible). Note: files generated before this control shipped that cannot be traced to a source conversation remain member-visible (a deliberate backward-compatibility grandfather — existing shared files are not retroactively removed).

This gate is now surfaced so a member is never left with a silently shorter list. Each knowledge row carries an advisory, read-only member_visible boolean — false marks a chat-generated file the caller can see only because they are its creator (its chat isn't shared yet), which the UI badges "Only you". A separate endpoint reports how many files are hidden from the caller by this gate (see Hidden knowledge count below). Both are display hints only — enforcement remains the server-side gate; the hidden count is a deliberate, bounded cross-user disclosure (a number of other members' private generated files, never their names or content).

In addition to the stored fields, the response includes the writable permission_tier enum and a server-derived, read-only permission object:

{
  "permission_tier": "local_only",
  "permission": {
    "tier": "local_only",
    "allows_cloud": false,
    "pii_mandatory": false,
    "user_configurable": false
  }
}

The permission object is computed from permission_tier on every read; it is never persisted and is never accepted as input. allows_cloud is false only for local_only, pii_mandatory is true only for pii_mandatory, and user_configurable is true only for user_configurable.


Updating a project

PATCH /admin/v1/projects/<ID>

Required role: owner or editor. Changing permission_tier (the project's PII/residency policy) or conversation_retention_days requires the project owner or a platform admin — an editor may edit every other field, but a 403 is returned if an editor attempts to change either of those to a new value.

Patchable fields: name, description, instructions, icon, color, permission_tier, default_gateway_id, default_model, conversation_retention_days (an integer 30–3650, or null/0 to clear; a bad value returns 400).

When a PATCH touches default_gateway_id or default_model, the resulting effective pin (the new value where supplied, else the project's stored value; an explicit null clears it) is run through the same offer validation as create — so pinning a model the gateway would residency/allowlist-block returns 400 with the corresponding code. A PATCH that does not touch either pin field (e.g. a rename) is never offer-validated, so an unrelated edit on a project that already carries a pin is unaffected. The project UI saves the preferred model as default_model with default_gateway_id: null (model-only normalisation); patching only non-pin fields leaves any stored pin untouched.

permission_tier may be changed but never cleared: it is NOT NULL, so sending an explicit null (or an unknown value) returns 400. Omit the field to leave the tier unchanged.


Deleting a project

DELETE /admin/v1/projects/<ID>

Required role: owner.

The conversations attached to the project are detached but not deleted. Any of the project's documents still queued for ingestion (status uploading or pending) are cancelled, so no further processing is started, and no embedding runs, for the deleted project.

Detaching lifts a residency restriction. If the project's access tier was Local only or PII mandatory, the conversations it held carried that restriction only through their binding to it — so once they are detached, their future turns may route to a cloud model, or egress without a masking mandate, like any unbound conversation. This includes conversations owned by other members. The delete is authorised (it requires owner rank, the same rank that could set the tier to user configurable and release everything anyway) and is therefore not blocked, but it is recorded: the audit log gets a project.conversations_released row carrying the project's tier and the number of conversations released, and the delete dialog states the consequence before the click. Releasing a single conversation is gated separately — see Updating a conversation.


Listing add-member candidates

GET /admin/v1/projects/<ID>/members/candidates

Returns the users of the project's tenant who are not already members — the data source for the invite dialog's browsable email typeahead. Each entry is { id, email, name } (no role/login/status metadata). The server excludes, and the client cannot broaden: soft-deleted users and existing members. Unlike the groups variant, platform admins are included — the add-member mutation accepts any same-tenant user, so an admin colleague is a valid, offerable target. The result is capped at 1000 rows (ordered by email); anyone beyond the cap is still addable through the exact-email path.

Authorized like the add-member mutation: project owner (or platform admin). An under-rank member (editor/viewer) → 403; an unknown project, a foreign tenant, or a non-member → 404 (no existence oracle). A database fault while reading the roster → 503 (retryable; the 403/404 authz stage precedes it).

Exposure note (deliberate): this is the first tenant-roster enumeration surface below the groups (GROUPS_MANAGE) tier — any project owner reads the tenant's user emails and names. In this pure-B2B product a tenant is one organisation, so the roster is the inviter's own colleagues; every other picker (agent/prompt sharing) stays exact-email only.


Adding a project member

POST /admin/v1/projects/<ID>/members

Required role: project owner (or platform admin).

Field Type Required
user_id string yes — an existing, non-deleted user of the project's tenant
role owner | editor | viewer no — defaults to viewer
share_existing boolean no — defaults to false

The new member receives a push notification on the project.

share_existing (bulk-share at invite). When set to boolean true, every conversation the calling user currently owns in this project that is still private is shared into the project Feed (shared_in_project), so the newly-added member can see them. This is a convenience for the existing per-conversation share to project action, not a new capability:

  • Authorization is server-side and fail-closed. The endpoint already requires the caller to be the project owner (or a platform admin); the bulk-share is additionally scoped to the caller's own conversations (WHERE user_id = <caller>), so it can never expose another member's private conversation — only that conversation's own creator may share it. The client cannot widen this scope.
  • Input validation. Only a real JSON boolean true triggers the share. Any other value — "true" (string), 1, null, an object/array, or the field's absence — is treated as false (no share). The share_existing flag is never the authorization boundary; the role gate and the owner-scoped query are.
  • Feed-wide, not per-invitee. Sharing to the Feed makes a conversation visible to every current and future project member, exactly like the per-conversation share action — not only to the invited member. Un-sharing is per-conversation (DELETE …/conversations/<ID>/share-project).
  • Idempotent. Already-shared conversations are skipped, so re-inviting (or a repeated call) shares nothing further.
  • Audited. The membership grant (project.member_added) is always written to the audit log; a successful bulk-share additionally writes project.member.share_existing (with the target user and the number of conversations shared). If the bulk-share hits a database fault nothing is shared and no share audit row is written — the fault is logged server-side and reported to the caller (see the response below).

Response. 201 with { "ok": true, "shared_count": <N> } — the number of conversations newly shared (0 when share_existing was not set or nothing was eligible). If the membership was added but the bulk-share hit a database fault, the response is still 201 (the member was added) with "shared_count": null and "share_failed": true, so the outcome is never silently swallowed.

Rejected (validated server-side; the client is never the authz boundary):

  • 409 — user_id is already a member. The add is a plain insert; it never rewrites an existing member's role — role changes go through the PATCH below, where the last-owner guard applies. (Previously a re-add silently overwrote the role, which could demote the last owner.)
  • 404 — unknown user, a user of another tenant, or a soft-deleted user (no deletion oracle: same status as unknown).
  • 400 — missing user_id or an invalid role.
  • 500 — a storage fault (the member was NOT added).

A rejected request (409/404/400) performs no membership audit entry, no push notification, and no share_existing bulk-share — those run only after a successful insert.


Updating a member role

PATCH /admin/v1/projects/<ID>/members/<USER_ID>

Required role: project owner (or platform admin).

Field Type Required
role owner | editor | viewer yes

Returns 409 conflict when the change would demote the last owner of the project.


Removing a member

DELETE /admin/v1/projects/<ID>/members/<USER_ID>

A member may remove themselves (USER_ID equals the caller). Removing other members requires the project owner role (or platform admin). The last owner cannot be removed.


Listing knowledge attachments

GET /admin/v1/projects/<ID>/knowledge

Required role: project member.

Returns a JSON array of knowledge rows (source-ACL and share-gate filtered for the caller). Each row carries the stored fields plus a read-only member_visible boolean: false means this is a chat-generated file the caller sees only as its creator (its chat isn't shared into the project yet). A chat-generated row also carries source_conversation_id — the id of the conversation that produced the file (absent/null for an uploaded file, which has no source chat). It is derived from the same chat_message_file → chat_conversation link the member_visible gate uses, so the two never disagree on which conversation governs a file. The UI uses it to offer a "Share this chat" action on an "Only you" row: sharing that conversation (POST /conversations/<id>/share-project) un-hides the file to project members. Rejected: a non-member is refused by the access gate (403/404). A DB error is a 500, never a null or false-empty 200.

Hidden knowledge count

GET /admin/v1/projects/<ID>/knowledge-hidden-count

Required role: project member.

Returns { "hidden_count": N } — how many of the project's files are hidden from the calling user by the chat-generated share gate (chat-generated, unshared, not their own). Backs the UI hint "N files are only visible to their creator". The count is a deliberate, bounded cross-user disclosure: it reveals only the number of other members' private generated files, never their filenames or content. A DB error is a 500 (fail-closed — never a misleading 0). The path is hyphenated (not /knowledge/hidden-count) so it can never be mistaken for a single knowledge item id.


Uploading knowledge

POST /admin/v1/projects/<ID>/knowledge

POST /admin/v1/projects/<ID>/knowledge/upload

Required role: owner or editor.

The first form accepts a JSON body for plain-text knowledge entries. The second form also accepts a JSON body for file uploads:

Field Type Required Description
filename string yes Original file name; a leading path is the file's folder (reports/q1/report.pdf). See the filename rules below.
data string yes File bytes, base64-encoded.
mime_type string no Defaults to application/octet-stream (null and "" count as absent). Valid UTF-8, at most 128 characters; a non-string, malformed or longer value → 400. A known extension overrides it.

Filename rules (both forms, and the PUT …/knowledge/<ENCODED_FILENAME> path segment). The name is validated once, at the trust boundary, before any processing or storage call. Accepted: a non-empty string of at most 255 characters (code points — a 255-umlaut name is fine), valid UTF-8, where every /-separated segment is non-empty and is not . or ... Rejected with 400 and a reason: a missing / non-string / empty name, more than 255 characters, malformed UTF-8, a control character (bytes 0x00–0x1F, 0x7F) or a backslash, a leading or trailing /, an empty segment (a//b.txt), or a . / .. segment. Names are never rewritten — what the client sends is what is stored.

Name conflicts (409, code: "knowledge_name_conflict"). A file name is unique per project, and "the same name" is the database's notion (collation utf8mb4_uca1400_ai_ci): the comparison ignores case (Notes.txt = notes.txt), accents (Café.txt = Cafe.txt), ß/ss and ligature folds, and trailing spaces. A write whose name already exists — on either form, or an upload against a pasted-text row and vice versa — is rejected with 409 and a body naming the request's spelling: {"error": "A document named \"notes.txt\" already exists in this project. Delete it first.", "code": "knowledge_name_conflict"}. Branch on code; the prose may change. The existing row is untouched, sibling uploads in the same batch are unaffected, and no second row is ever created. (The upload form's extension gate runs first, so notes.txt with a trailing space is a 422 Unsupported file type there, not a 409.) Any other insert failure is a 500 with the fixed text insert failed — the database error is logged, never returned.

The decoded file size is capped per format at the trust boundary (the browser-supplied mime_type is advisory — the cap is keyed on the canonical format resolved from the file extension, so a mislabelled file cannot dodge it). A file over its cap is rejected with 413 and a message that names the limit, before any processing:

Format Decoded size cap
Office + PDF (.pdf, .docx/.pptx/.xlsx/.xlsm/.ods/.odt) 100 MB
XML (.xml), JSON (.json), HTML (.html/.htm) 64 MB
Text (.txt/.md/.markdown/.csv/.yaml/.yml/.rst/.log) 7.5 MB (must fit the extracted-text storage limit)
Images (.png/.jpg/.jpeg/.webp) and everything else 20 MB

Large uploads are stored chunked internally (a single file may exceed the database packet size); this is transparent and does not change the request shape.

Malware scanning (fail-closed, opt-in per gateway). Malware scanning is enabled per gateway (av_scan.enabled in a gateway config) and is off by default. A project has no single gateway and its knowledge blob is persisted and served back raw, so the upload is scanned when the project's tenant has malware scanning enabled on any of its gateways; otherwise this step is skipped. When it applies, the decoded bytes are streamed to a self-hosted ClamAV daemon (clamd, INSTREAM) after the size caps and before the document is staged or stored. An infected verdict rejects the upload with 422 and code: "malware_detected" (the ClamAV signature is logged operator-side only, never returned). If clamd is unreachable, times out, or errors, the upload is rejected — not accepted unscanned — with 503, code: "scan_unavailable" and a short Retry-After. The virus scanner itself is configured for your deployment by Myra.

PDF, text (.txt/.md/.markdown/.csv/.yaml/.yml/.rst/.log), and Word (.docx) uploads are processed asynchronously and segmented. (This ingestion path is deliberately distinct from the synchronous chat/workflow read of the same .docx: here the file is segmented into char-offset units with a heading trail for retrieval and citations — a contract the pandoc-to-Markdown chat reader cannot produce — so the two paths use different extractors and their stored/preview text may differ.) The endpoint validates the file synchronously at upload (a corrupt or password-protected PDF is rejected up front with 422 "Could not read the PDF (it may be corrupt or password-protected)" — no row, blob, or ingest slot is consumed, so it never stages pending only to terminal-fail later; a PDF encrypted with an owner password but no user password is readable and still ingests; a text file over 7.5 MB returns 413; a .docx that is not a valid Word zip, or whose XML declares a DTD, is later marked failed by the worker with a clear reason), stores it, and returns 202 Accepted with the row at ingest_status: "pending" — extraction + segmentation then run in the background and the status advances to segmented (or partial for a truncated large PDF). If a PDF or Word document's extracted text exceeds the 7.5 MB storage limit (the extracted text can be far larger than the uploaded file), the document is not failed: its text is truncated to the limit on a character boundary and stored as a doc-level extracted entry (readable in chat up to the doc-level size cap, without a page/section index), and the truncation is recorded on the document's ingest trace. This degrade is distinct from the text-file case above, where an over-7.5 MB .txt/.md/.markdown/.csv/.yaml/.yml/.rst/.log upload is rejected with 413 before processing. A text upload whose content is binary rather than text — a non-text file renamed to a text extension, detected by a NUL byte in the decoded content — is marked failed by the worker (it is not silently stored as mojibake); the persist boundary rejects any text-file extractor output carrying a NUL, so a masquerading binary never reaches project knowledge. PDF and image (OCR) extraction is different: those formats have an OCR fallback, so their extractor scrubs NUL/control bytes (rather than failing the document) and, for a PDF, routes a garbage or broken-font text layer to OCR — see the note on the PDF text-quality gate under Managing knowledge files. A stray control byte in a PDF's text layer therefore ingests cleanly instead of losing the document. The file extension determines the format (the browser-supplied MIME is overridden — .md/.csv/.yaml/.yml/.rst/.log/.docx are canonicalized regardless of what the browser sends; .yaml/.yml/.rst/.log are treated as plain text and line-segmented like .txt). Note: .yaml/.yml/.rst/.log files uploaded before this became searchable are stored doc-level and are not retroactively indexed — re-upload them to make them searchable. Poll GET …/knowledge (or the item endpoint) until the status leaves pending/extracting. If too many of the tenant's documents are still processing, a further upload returns 429 (retry once some finish). Text files gain a line segment index (L.n locators, or L.a–b for a range; Markdown carries the heading trail, CSV carries the header as section context). A very long text file — more paragraphs, headings or row groups than the index holds — is indexed with each entry covering a RANGE of lines rather than a single unit, so the index still reaches the document's last line and every part of it stays citable; .docx gains a para index (¶N locators + the Word heading trail as section context). All become page/section-citable and RAG-narrowable like PDF. PowerPoint (.pptx) gains a slide index (slide N locators + the slide title as section context) and Excel (.xlsx/.xlsm) gains a sheet index (sheet N locators + the sheet name as section context), also async. Each sheet's name also heads its own block in the extracted text (a plain header line, the same way a heading heads a .docx paragraph block), so the model can name the workbook's tabs and attribute rows to the right one — not just read anonymous cells. Images (.png/.jpg/.jpeg/.webp) are OCR'd asynchronously (on-prem qwen vision model) into a region-segmented doc with ocr confidence — the transcribed text is retrievable and RAG-searchable like any other document, but its segments are non-strict-citable (OCR offsets are best-effort, not verbatim page citations), and figures/numbers should be verified against the original. Same async staging as PDF (202 → pending → extracted/segmented). An image or PDF with no readable text (a photo, a blank scan, or an OCR that returned nothing usable) is stored non-terminally with a short "no readable text" marker and is served (never failed) — the row lands segmented with an empty index, not extracted. Unsupported image types (.gif/.bmp/.tiff/.heic/ .svg) return 422. An uploaded image — and a scanned PDF page, which is OCR'd the same way — is untrusted input (the OCR model's response is untrusted too): the transcription prompt carries the data-boundary clause described in the threat model, which instructs the model to treat text inside the file as data to transcribe rather than as instructions to follow. As a further guard, if the OCR model echoes the task prompt back verbatim instead of transcribing (a known failure mode), that echoed output is detected and dropped rather than stored as document content — a page/image with only an echoed prompt becomes the "no readable text" marker. When an OCR transcription is truncated because a dense page exceeded the model's output window, the document is stored partial (its extracted text is real but incomplete) and, when read in chat, carries an honesty marker so the model does not treat it as complete. Like the other spotlighting markers it reduces, but does not eliminate, the risk. The clause applies at ingest time going forward; documents ingested earlier keep the transcription they were stored with. OpenDocument spreadsheets (.ods) are extracted asynchronously (202 → pending → segmented): the sheets are flattened to CSV and given a line (L.n) index like a .csv — a multi-sheet .ods is concatenated into one table (no per-sheet names or sheet N locators, unlike .xlsx). This matches how a connector-synced .ods is ingested. Legacy .xls is unsupported (422). The JSON POST …/knowledge form (a JSON body {"filename", "extracted_text", "content_type"?}, no file) stays for pasted-text entries, which are stored doc-level (no index). Its filename follows the filename rules above (400 otherwise) and the same per-project uniqueness (409 + knowledge_name_conflict); its content_type is optional (null / "" = absent → text/plain), and must be a valid-UTF-8 string of at most 128 characters (400 otherwise — a malformed value is never read as absent). The extracted_text must be valid UTF-8 (utf8mb4) text of at most 7.5 MB with no NUL byte; a body that is oversize, not valid UTF-8, or contains a NUL byte (the signature of binary content decoded as text — e.g. an image read client-side) is rejected with 400. On the project routes an image/* content_type is rejected with 400 except image/svg+xml (SVG is text) — images belong on the …/knowledge/upload route, not the pasted-text form. size_bytes is derived server-side from the text; any client-supplied size_bytes is ignored. The same extracted_text validation (UTF-8 / ≤ 7.5 MB / no NUL) applies to the PUT …/knowledge/<ENCODED_FILENAME> and the conversation PUT …/conversations/<ID>/knowledge/<KID> write-back forms. The content_type checks differ, though: the conversation write-back inherits content_type from the existing row (the file was created by write_file) and only nullable-checks a supplied value — it does not apply the image/* reject or a length cap there.

For an uploaded .docx the extractor accepts a valid OOXML zip containing word/document.xml and rejects (→ failed, never crashing) a non-zip, a missing/oversized document.xml, or any XML that declares a DTD/entity (an entity-expansion or external-entity attack is refused at the parser). Only the document body and its tables are indexed; text in headers/footers, footnotes, and text boxes is not (the whole document text is still stored and served). A .docx attached in a chat (rather than a project) is still stored as flat text without the para index.

XML (.xml), JSON (.json), and HTML (.html/.htm) uploads are extracted asynchronously (202 → pending → segmented), up to the 64 MB per-format cap above. XML is parsed with a DTD-hardened SAX pass — any DOCTYPE, internal/external entity, or parameter-entity declaration is refused (billion-laughs / XXE / SSRF-via-DTD are rejected, not expanded); a malformed or empty document is marked failed, never crashing. JSON is validated and pretty-printed for retrieval; NaN/Infinity, an overflowing exponent, a huge integer literal, deeply nested input, and non-UTF-8 bytes are all rejected. HTML has its script/style/noscript content stripped — including the self-closing <script/>/<style/>/<noscript/> forms — so text hidden from a JS-enabled reader (notably <noscript>) never reaches the searchable index; its visible text is line-segmented. (<head> is not depth-suppressed — a <title> is human-visible in the browser tab, and suppressing an unclosed <head> would blank a real page.) For any format, if the extracted text is very large the emitted result is bounded so the worker never overflows its output pipe — the text is truncated (on a codepoint boundary) and stored as a doc-level extracted entry with the truncation recorded on the ingest trace, exactly like the PDF/Word over-7.5 MB degrade above. A file up to 64 MB therefore always ingests (its text truncated to the indexed ceiling if needed); 64 MB + 1 byte is rejected with 413 at upload.

The upload endpoint returns two distinct 503s, told apart by a stable machine-readable code in the JSON body ({"error": "…", "code": "…"}). A client must branch on code (the error prose is informational and may change) — never on the message text:

code Meaning Nature Retry-After
ingest_busy The shared document-processing pool is momentarily saturated (this endpoint and POST /admin/v1/chat/files share one bounded slot budget). Transient — retry the same upload shortly. 10 (seconds)
ingest_unavailable Async ingestion is not provisioned on this deployment (no ingest pipeline configured). Permanent / config — an operator must enable ingestion. —
knowledge_name_conflict (409) A knowledge file with this name (case-, accent- and trailing-space-insensitive) already exists in the project — see Name conflicts above. Also returned by the pasted-text form and the move endpoint. Permanent for this name — delete the existing file or choose another name. —

Async ingestion (PDF + text + Office + XML/JSON/HTML + images) requires the deployment to be provisioned for it. When the ingest pipeline is not configured the upload returns 503 with code: "ingest_unavailable" and a clear error ("Document ingestion is not currently available. Contact your administrator.") instead of accepting a document that could never be processed — an upload never silently sticks in pending. When ingestion is provisioned but the shared upload pool is momentarily at capacity, the upload instead returns 503 with code: "ingest_busy" and a Retry-After: 10 header — a transient state the caller resolves by retrying, not by contacting an administrator. A pending row that is never picked up by a worker (for example after an outage) is automatically failed by the background reaper so it surfaces as failed rather than processing indefinitely.

Semantic retrieval for large documents

When the assistant reads a large project document (one whose text exceeds the per-read size limit and would otherwise be truncated to its opening pages), the gateway returns the passages most relevant to the user's question — selected by semantic similarity from anywhere in the document — instead of just the beginning. Each served passage keeps its [file | page | section] citation anchor, so grounded citations still reference only pages the model actually received. This engages only once a document has been semantically indexed in the background. Until then the document is served from its beginning up to the per-read size limit and the remainder is not returned by a read — a very large document is therefore only partly readable while its index is still being built, or if it is too large to index at all. A document that fits within the limit is served in full as before. A scanned (OCR) page is still served when relevant, but without a citation anchor (its page location is not strictly verifiable). The user's question is used solely as a similarity query and never affects which documents are authorized — the per-request x-aig-knowledge-ids selection and project scope are unchanged.

Verified excerpts in the sources panel

When the assistant supports a point with a verbatim quote from a cited page, the gateway checks deterministically that the quote actually appears on that page and, if so, attaches it to the source as a verified excerpt shown under the page in the sources panel. This is purely additive — the deterministic list of pages the assistant was given is never filtered or removed; a paraphrased or ungrounded quote simply gets no excerpt. "Verified" means the quote is genuinely present on the named page (open the document there to see it in context) — it is a traceability aid, not a guarantee that the surrounding claim is correct. A verified excerpt is real document text, so it is shown only to viewers authorized for that document; a shared-conversation reader without document access sees the page reference without the excerpt.


Getting a knowledge entry

GET /admin/v1/projects/<ID>/knowledge/<KID>

GET /admin/v1/projects/<ID>/knowledge/<KID>/download

Required role: project member.

The /download sub-path returns the entry's original bytes with a sanitized Content-Disposition attachment filename. It serves any blob-backed entry — an uploaded file (source: "upload") or a code-interpreter generated file (source: "generated", e.g. an .xlsx workbook produced in a project chat). A text-only entry that has no stored bytes (a pasted-text or write_file row) returns 404; a storage fault returns 500.

The knowledge entry includes, for paged documents:

Field Type Description
page_count integer | null True total number of pages in the source document. null for non-paged formats.
pages_extracted integer | null Pages actually extracted into the stored text (capped at the extractor's page limit). null for non-paged formats.
ingest_status string Extraction lifecycle: extracted/segmented = ready; uploading/pending/extracting = still processing; failed = could not be extracted.

When pages_extracted < page_count the document was truncated at the page limit; the later pages are not yet searchable or citable. Clients should surface this (the web UI shows a "showing N of M pages" notice).

A document is only readable in chat once its ingest_status is extracted or segmented. While it is still processing (uploading/pending/extracting) the read_file tool returns a "still being processed" message instead of its content, and the document is listed as not-yet-readable in the assistant's file list — so an in-progress document is never answered over as if it were empty.

Page-anchored retrieval

For a document that has a page/section index (ingest_status segmented or partial — PDF by page, .txt/.md/.markdown/.csv/.yaml/.yml/.rst/.log by line, .docx by paragraph, .pptx by slide, .xlsx/.xlsm by sheet), the read_file tool returns the document text annotated with anchors. Each boundary is marked inline with an anchor label of the form [<filename> | <locator> | <section>], for example [report.pdf | p.12 | Introduction], [notes.md | L.40–58 | Overview], or [memo.docx | ¶12 | Methods] (the section part is omitted when the extractor found no heading/header). The full document text is always returned — the anchors are inserted between the text, never in place of it, so no content is dropped. This gives the assistant an explicit, line- or page-level reference it can cite when answering from your own documents. A document with no index (a doc-level extracted row — a pasted-text entry, a document whose segmentation degraded, or one whose oversize extracted text was truncated to the storage limit) is returned as-is, with no anchors.

Getting the ingest trace

GET /admin/v1/projects/<ID>/knowledge/<KID>/trace

Required role: project member (viewer or above). The <KID> must belong to <ID>; a knowledge id from another project returns 404 (it is never resolved outside the project you can access).

Returns the append-only timeline of a document's ingest lifecycle — one event per state transition — to diagnose a document that is stuck or degraded:

{
  "knowledge_id": "abc123",
  "events": [
    { "id": 1, "ts": 1751700000, "stage": "pending",     "detail": null },
    { "id": 2, "ts": 1751700003, "stage": "extracting",  "detail": "application/pdf" },
    { "id": 3, "ts": 1751700009, "stage": "segmented",   "detail": "pdf-qwen-1" }
  ]
}
Field Type Description
id integer Monotonic event id; events are returned oldest-first.
ts integer Unix timestamp (seconds) the event was recorded.
stage string Lifecycle stage: pending, extracting, segmented, partial, extracted, failed.
detail string | null Free-text context: the content type at claim, the extractor version at completion, or the failure reason.

events is always a JSON array (empty [] for a document with no recorded events — e.g. one attached before tracing existed). The trace is a diagnostic aid, not the source of truth: the authoritative state is the entry's ingest_status field. A missing or partial trace never implies a document failed.


Updating a knowledge entry

PUT /admin/v1/projects/<ID>/knowledge/<ENCODED_FILENAME>

Required role: owner or editor.

<ENCODED_FILENAME> is the stored name, percent-encoded once by the client (it is decoded once, by the web server — a name containing a literal % is sent as %25 and stored as %, and a + is a literal plus, never a space); it is validated by the same filename rules as the upload forms (400 for an over-long name, malformed UTF-8, a control character or a backslash such as %5C). This form reaches top-level names only: a folder separator is decoded from %2F into a real / before routing, so reports%2Fq1%2Fx.md matches no knowledge route (404) — a foldered file is written through the upload form and re-filed with the move endpoint. Likewise a .. or // in the URL is normalized by the web server before routing and a trailing / matches no knowledge route, so those shapes never reach this handler (they answer 404, or whatever unrelated route the normalized path lands on). The body follows the pasted-text form (extracted_text required, optional content_type with the same rules). An existing row of that name is updated in place — this form never raises a name conflict. Any storage failure on this form, or on DELETE …/knowledge/<KID>, is a 500 with a fixed message (the database error is logged, never returned).


Deleting a knowledge entry

DELETE /admin/v1/projects/<ID>/knowledge/<KID>

Required role: owner or editor.


Moving a knowledge entry into a folder

PATCH /admin/v1/projects/<ID>/knowledge/<KID>/move

Required role: owner or editor.

Knowledge attachments are organized into a folder hierarchy. A file's folder is the leading path of its name — reports/2024/summary.pdf lives in folder reports/2024 — so a folder exists exactly while at least one file references it (there is no separate empty folder to create or delete). This endpoint changes a file's folder while keeping its own name.

Request body:

{ "folder": "reports/2024" }
Field Accepted shape
folder A /-separated relative path, or "" / omitted for the top level. Segments are trimmed of surrounding spaces; leading, trailing and duplicated / are normalized away.

The folder is sanitized at the trust boundary and the request is rejected with 400 when it contains:

  • malformed UTF-8;
  • a control character (bytes 0x00–0x1F, 0x7F) or a backslash;
  • a . or .. path segment (no traversal);
  • a segment longer than 96 characters, more than 12 nesting levels, or a total path longer than 200 characters.

Responses:

  • 200 with the refreshed knowledge list on success (a no-op move — same target folder — also returns 200).
  • 400 when the folder is invalid (see above), or when the joined folder/name would exceed 255 characters (counted in characters, not bytes). Only the width is checked here — the file's own name is the stored one and is not re-validated by a move.
  • 404 when the knowledge id does not belong to this project.
  • 409 with code: "knowledge_name_conflict" when a file with the same name (per the collation rules under Name conflicts) already exists in the target folder.
  • 500 with a fixed message on any other database failure (the error is logged, never returned).

Listing project conversations

GET /admin/v1/projects/<ID>/conversations

Returns only the calling user's own conversations in the project, not every member's (a privacy boundary). Required role: project member.


Listing the project feed

GET /admin/v1/projects/<ID>/feed

Returns the project's shared conversations, newest first (the rows where shared_in_project = 1). Despite the name it does not include knowledge or membership events. Required role: project member.


Reading the combined knowledge text

GET /admin/v1/projects/<ID>/knowledge-text

Returns a JSON array of per-entry objects — each { id, filename, token_count, extracted_text, ingest_status } — one per knowledge entry (not a single concatenated blob). Used by the chat UI for context injection.