Direct document-AI actions (/direct/v1)
The /direct/v1 surface lets a browser-based Office add-in run document-AI actions through a Myra
gateway. Three add-ins consume it: the Office.js Outlook task pane (AGF-2812, over the open
message), the Office.js Excel task pane (AGF-2823, over the selected cell range) and the Office.js
Word task pane (AGF-2687, over the selection/document, inserting as tracked changes). It is the
direct-adapter sibling of the Microsoft 365 Copilot lane (/copilot/v1):
the action shaping, server-pinned prompts, model pinning, CONTENT pipeline, PII/redact behaviour,
billing, and one-request_log-row semantics are identical — read that page for the request/response
shape and per-action params. This page documents only what differs.
What differs from /copilot/v1
/copilot/v1 (Copilot) |
/direct/v1 (Office add-in) |
|
|---|---|---|
| Caller | Microsoft's servers (server-to-server) | the user's browser (Office task pane) |
| CORS | none (server-side call) | enabled for the SPA origin (a deployment setting) |
| Entra app | the Copilot app (copilot_entra_*) |
the add-in's own app (direct_entra_*) |
generate action |
rejected 400 (never widens) |
available (direct-only) — a generic generative verb (AGF-2823) |
ground action |
permanently excluded (Art.44) | reserved direct-only (not yet built) |
Everything else is shared code (one ENTRA_ACCESS_PIPELINE, one entra_auth middleware, one action
shaper); the lane is selected by a server-set tag, never by the request.
Authentication
Send a Microsoft Entra OAuth2 access token minted by Office SSO
(OfficeRuntime.auth.getAccessToken) as a bearer:
It is validated exactly as the Copilot token (signature against the tenant JWKS, issuer pin,
aud/azp allowlist, mandatory exp, a delegated-token scp gate) — but against the
direct_entra_* settings, which describe the add-in's own Entra app registration, not the
Copilot app. Key consequences:
- The token's
audis the add-in's App ID URI (api://<host>/{add-in-client-id}). This value must be listed indirect_entra_aud_allowand must not overlapcopilot_entra_aud_allow— the disjointaudis the fence that keeps a Copilot server-to-server token off this browser lane (the write API rejects an overlapping value400). - The token's
azp/appidis a Microsoft Office host client id (host/platform-dependent), not the add-in's app id — list the relevant host id(s) indirect_entra_azp_allow.azpvalues may overlap the Copilot lane (they are Microsoft's), so no overlap check is applied to them.
The Myra tenant (tid) and user (oid) resolution and the document-AI entitlement gate are
identical to the Copilot lane. Unset direct_entra_aud_allow/direct_entra_azp_allow → the lane is
dormant (every token 401) — the feature ships dark.
CORS
/direct/v1 is CORS-enabled for the single configured SPA origin (set by your operator, e.g.
https://ai.myra.eu — the task panes are hosted there under /outlook/ and /excel/):
OPTIONSpreflight →204withAccess-Control-Allow-Methods: POST, OPTIONSandAccess-Control-Allow-Headers: Authorization, Content-Type.- The
POST/error response carriesAccess-Control-Allow-Originon every status (so the pane can read the body of a401/403/413/429, not an opaque CORS error) andAccess-Control-Expose-Headers: X-Request-Id, Retry-After, X-AIG-PII-Active, X-AIG-Model, X-AIG-Doc-AI-Action. - No
Access-Control-Allow-Credentials— auth is a Bearer token, never a cookie.
Endpoints, request/response, and errors
Identical to /copilot/v1 with the path prefix /direct/v1:
POST /direct/v1/summarize
POST /direct/v1/rewrite
POST /direct/v1/translate
POST /direct/v1/redact
POST /direct/v1/generate # direct-only (AGF-2823)
Request body, per-action params, the 200 chat-completions shape, redact fail-closed behaviour,
and the error table are the same. Lane-specific notes:
- A URL whose prefix does not match the authenticated lane (e.g.
/copilot/v1/...presented with an add-in token, or vice versa) is rejected400 invalid_request— the action is resolved from the lane, not trusted from the URL. generate(direct-only). A generic generative action: "using the suppliedtextas context, produce exactly whatparams.instructionasks for; return ONLY that." The Excel add-in uses it for its formula / clean / build-a-table actions (the fixedrewriteprompt — "improve clarity, preserve meaning" — is a poor base for those).params.instructionis required (400if absent — a generate with no instruction is meaningless) and is sanitised into one bounded slot exactly likerewrite's. It is rejected400on/copilot/v1(identical to an unknown action — the copilot lane never widens); it resolves only on the direct lane.groundis not a valid action on either lane yet (400); when built it will be direct-only.
Untrusted model output written into cells (Excel add-in)
The /direct/v1 responses are OpenAI chat-completions text — untrusted (a model response is
hostile until validated). The Excel add-in writes some of that text back into spreadsheet cells, which
is a spreadsheet formula / CSV-injection trust boundary handled client-side, in the add-in (the
server returns text; the add-in decides how to write it):
- Accepted / neutralised: a value written as data (the clean / build-a-table actions) that begins
with
=,+,-,@(or after leading whitespace/tab/CR/LF) is written as literal text — the add-in apostrophe-prefixes it AND forces the target cells to text (numberFormat = "@"), so Excel never evaluates it. Exception: a leading+/-on a pure numeric literal (e.g.-5,+3.14,-1.5e-3) is preserved as a computable number (Excel stores it as the number, not a formula) — only non-numeric trigger values are neutralised, so a cleaned negative column stays summable. Embedded tab/newline characters in a source cell are collapsed to a space so a hostile sheet cannot forge extra columns/rows in the prompt. - Rejected: a formula (the write-a-formula action, written live) that uses a network/exfil/DDE or
code-execution function (
WEBSERVICE,FILTERXML,ENCODEURL,RTD,IMAGE,HYPERLINK,CALL,REGISTER,EXEC,IMPORT*,PY,SQL.REQUEST, DDEcmd|…!) is not written and the user is told why — even though the user asked for a formula (a denylist; the residual risk of a novel function is mitigated by the mandatory preview-before-write).IMAGEis the worst exfil vector — it auto-fetches an arbitrary URL on write with no click;PY(Python in Excel) andSQL.REQUEST(ODBC) are the code-execution / data-connection vectors. Leading-trigger disguises (+WEBSERVICE,@WEBSERVICE,\t=WEBSERVICE,=+PY) are canonicalised before matching.
Nothing is ever written to the sheet automatically — every write shows a preview the user must confirm.
The Excel add-in also offers a read-only "Ask a question" action (AGF-3123): it sends the user's
question as generate's params.instruction with the selected range (TSV) as context and renders the
answer in the task pane only (as textContent, never innerHTML). It has no write path, so it
adds no new cell-injection surface — the untrusted answer is display text.
The pane's tone / length / format quick-controls (AGF-3121) are composed client-side onto the
action's existing params slot by the shared office-addins/action-controls.ts helper — they never
introduce a new slot: generate/rewrite get the control phrases appended to instruction, and
summarize gets the length slot; translate/redact take none. The write-a-formula action is
deliberately exempt (controls would corrupt its "return ONLY a formula" instruction, so its instruction
stays authoritative and the controls are hidden for it).
Untrusted model output inserted into the document (Word add-in)
The Word add-in (AGF-2687) inserts /direct/v1 response text into the user's document — the same
untrusted model output boundary as Excel, handled client-side, in the add-in:
- Accepted: markdown headings, paragraphs and list items are parsed (client-side, pure
md.ts) into a flat list of inert text blocks and inserted via the Word API'sinsertText/insertParagraph+ a built-in heading style ONLY. - Neutralised: markdown links
[text](url)and imagesare rendered as inert plain text (the link becomestext (url), the image becomes its alt text) — never a live hyperlink or a remote image. So a[x](javascript:…)orin a hostile response cannot execute a URL or fire a remote-fetch beacon. Emphasis/code markers are flattened. - Rejected (by construction + enforced): the add-in never calls a Word sink that could render
a live hyperlink / image / HTML / field code —
insertHtml,insertOoxml,insertFileFromBase64,insertInlinePictureFromBase64,insertField,range.hyperlinksare all forbidden in the add-in source by a lint grep-guard (scripts/lint_office_addins.sh) and the inert-block behaviour is unit-tested + fuzzed (md.test.ts/md.fuzz.test.ts). This is the Word analogue of Excel's formula denylist above.
Every AI edit lands as a tracked change (when the Word host supports WordApi 1.4), so the user reviews and accepts/rejects it — nothing is applied silently.
Enabling document AI for a tenant
Provisioning is identical to the Copilot lane — see
Enabling document AI for a tenant. The same
designated gateway + model serve both lanes; note that this gateway's ip_allowlist must stay empty
if used by the browser lane (arbitrary end-user IPs). Add-in provisioning + settings:
docs/internal/outlook-addin-lane.md (Outlook),
docs/internal/excel-addin-lane.md (Excel).