Skip to content

Connector sync API

The connector sync API pulls documents from an external content source — a SharePoint drive or a Confluence space — into one project's knowledge base. A source is registered once against a project and a connector, then kept current in two ways: an explicit "Sync now" request at any time, and an automatic scheduled re-sync on a per-source cadence (new sources default to every 24 hours; see Automatic re-sync below).

All endpoints require an authenticated admin session (aig_admin cookie). The base URL is https://ai-api-admin.myra.eu/admin/v1. Access is governed by the target project's per-project membership: owner, editor, viewer.


Overview

A sync source binds one project to one selection in an external system:

  • a SharePoint drive, reached through a Microsoft Graph-scoped OAuth connector (delegated, see the setup note below), or
  • a Confluence space, reached through the Atlassian MCP connector.

Running a source performs a pull crawl of that one selection into the target project's knowledge base — the same knowledge store used by uploaded files, so synced documents are read, cited, and retrieved exactly like any other project knowledge. The sync is incremental and replacing: a re-sync brings changed and new items up to date and tombstones items that have been removed at the source, so the project's copy tracks the source rather than accumulating stale documents.

Updating a document does not take it away while the update runs. When a document has changed at the source, the new version is prepared alongside the existing one and swapped in only once its content is stored. If the download or the storage step fails, the previous version stays exactly where it is and remains readable, citable and retrievable; the crawl reports the item as skipped and retries it on the next pass. A document you filed into a folder also keeps that folder when the source updates it.

Two limits are worth knowing. If the source replaces a document with a file type the extractor cannot read, there is nothing to index and the document is removed (counted as skipped, like any unsupported file). And if the swap is interrupted after the new version is stored, the document stays available but may briefly show a placeholder name beginning .aig-restage-; the next sync restores its proper name.

A synced password-protected PDF (encrypted with a user password) cannot be read: it is staged and then terminal-fails extraction, and its ingest trace records a password-protected reason (rather than a generic extraction failure) so the cause is diagnosable. An owner-password-only PDF (readable without a password) ingests normally.

For a source that carries its own access permissions (SharePoint), those permissions are re-read for the updated document. If that lookup fails, the update still goes through, and the document's permissions are re-read on a later sync — the same fail-closed behaviour as any newly synced document: while the permissions are unknown, the document is not shown.

Tombstoning is fail-safe against transient errors. A document is tombstoned only when the crawl positively confirms it is gone at the source — never merely because a transient fault (a database blip, a per-item download error, or a malformed source response) prevented the crawl from confirming it this pass. If any such error occurs during a crawl, the end-of-crawl removal step is suppressed for that whole crawl (including across a paused/resumed crawl): every already-indexed document is kept, and removed items are reaped on the next fully clean crawl instead. The trade-off is deliberate — the gateway always errs toward keeping your documents. The count of items a crawl could not process is surfaced as items_skipped (the "could not be imported" notice), so a persistently failing item that keeps deferring removal is visible. (A benign skip — an unsupported file type the extractor can't read — is a positive outcome, not an error, and does not suppress removal.)

A crawl is resumable. It records a cursor and pauses cleanly, then continues from where it stopped on the next run, when either of two limits is hit:

  • Rate limit — the upstream returns 429; the gateway honours Retry-After and does not run the source again until that backoff has elapsed. The source status becomes paused_rate_limited.
  • Per-tenant ingest cap — at most 50 documents are ingested per tenant in one pass; when the cap is reached the crawl pauses with status paused_cap and is resumable on the next run. The cap counts only in-flight documents in live projects — documents left processing in a soft-deleted project do not consume it.

A paused crawl (either kind) is picked up again automatically on the next scheduled re-sync tick, or immediately by a manual "Sync now" — a large drive fills in over several passes rather than in one.

Source status

Status Meaning
idle Registered, not currently running; ready to sync.
running A crawl is in progress.
completed The last crawl finished with no items left to fetch.
error The last crawl failed; see last_error.
paused_cap Paused after reaching the per-tenant ingest cap (50); resumable.
paused_rate_limited Paused on an upstream 429; resumable on the next run once the Retry-After backoff has elapsed.

provider is one of sharepoint, confluence, or url.

The url provider (a public website) has no credential holder — it takes no connector_id, and its source_ref is the page URL itself. The gateway fetches the page over its SSRF-guarded HTTP client, extracts the readable text, and indexes it into the project knowledge base exactly like an uploaded document; a re-sync re-fetches the same URL and re-indexes only if the page changed.


Access-control stance

Read this before syncing anything.

SharePoint per-item source permissions ARE mirrored. At sync time the gateway captures each SharePoint document's per-item ACL (Microsoft Graph /permissions) and, at retrieval, filters synced documents per requesting user by that source ACL, fail-closed. A document a user cannot open in SharePoint is not findable, readable, or citable to them here either — across every surface (chat retrieval / search_knowledge, the file browser, item preview, and download). A synced document whose source ACL could not be captured is hidden, never shown. A gateway admin does not bypass the source ACL.

  • Accepted principal shapes (SharePoint): a Graph grant to a user (matched by the requester's Entra object id — their SCIM externalId, or their OIDC subject when the tenant maps subject_claim=oid); a grant to a group (matched via the SCIM/IdP group mapping); an organization-wide / "Everyone" grant (any member of the same tenant); a "specific people" sharing link (expanded to the individual users). Rejected / dropped (fail-closed): anonymous or external sharing links, unrecognised principal shapes, and — for a user whose Entra object id is not known to the gateway (no SCIM externalId and not an oid-claim tenant) — a per-user grant simply does not match (they under-serve, never over-serve).
  • Freshness: SharePoint permissions are re-captured on every "Sync now", including for unchanged documents. A revoked user loses access on the next sync run that completes past that document.

Confluence defaults to the sync creator only, with an admin org-wide opt-in. The Confluence connector does not read per-page restrictions, so by default a Confluence-synced document is visible only to the user who created the sync source — no other project member can read or cite it (fail-closed).

A tenant admin may instead publish a Confluence source org-wide (share_scope: "org", via the share endpoint below). When published, every page that source synced becomes readable and citable to everyone who can open the project — including pages that are restricted in Confluence. This is deliberate admin curation, not permission mirroring: the tenant admin takes responsibility for exposing those pages tenant-wide. It never crosses tenant boundaries (a source is only ever reachable within its own tenant's project). Reverting to creator restores the creator-only default instantly (no re-sync needed). Per-user Confluence-parity sharing (a per-page restriction read plus an Atlassian↔gateway identity map) remains future work.

Manual uploads are unchanged: a document uploaded directly to a project (not via a connector sync) remains visible to the project's members at their project role, as before.

  • Run-as-invoker. A sync runs under the OAuth/MCP credentials of the user who created the source (its created_by), not the person who clicked "Sync now". This is why the run endpoint additionally requires the caller to be the source's creator. If that user is deprovisioned, their credentials stop working — delete and recreate the source under an active user to restore syncing.

Registering a sync source

POST /admin/v1/connector-sync

Registers a new source against a project and connector. Required role: project editor (owner or editor) on the target project.

Field Type Required
project_id string yes
connector_id string yes for sharepoint/confluence; must be omitted for url (a website has no connector — supplying one returns 400).
provider sharepoint | confluence | url yes
source_ref string yes — for sharepoint/confluence: the source selector (a SharePoint drive id or a Confluence space key), charset ^[A-Za-z0-9_.!-]+$. For url: the website URL (http/https). Max 512 characters.
source_label string no — a human-readable label for the source.
auto_sync_interval_secs integer no — the automatic re-sync cadence in seconds. Omitted → 86400 (24h, ON). Accepted values: 0 (off) or an integer in 3600–604800 (1 hour to 7 days). Any other value is rejected with 400.

On success the response is 201 Created with the new source at status: "idle":

{
  "id": "src_abc123",
  "project_id": "prj_1",
  "connector_id": "mcp_1",
  "provider": "sharepoint",
  "source_ref": "b!drive-id",
  "source_label": "Team drive",
  "status": "idle",
  "auto_sync_interval_secs": 86400
}

Accepted shape / What is rejected (fail-closed):

  • A missing or wrongly-typed field, or an out-of-range value, returns 400.
  • provider must be exactly sharepoint, confluence, or url; any other value returns 400.
  • auto_sync_interval_secs, when present, must be an integer that is 0 or in 3600–604800. A non-integer (e.g. 3600.5), a value below the 1-hour floor (e.g. 60), a value above the 7-day ceiling, or a non-numeric junk value returns 400. When omitted it defaults to 86400 (24h ON) — a newly registered source auto-syncs daily unless you opt out.
  • source_ref must match ^[A-Za-z0-9_.!-]+$ and be at most 512 characters — a value containing spaces, /, :, ?, #, %, \, control characters, or any other character outside that set is rejected with 400. The source_ref is a bare selector, never a URL or a path.
  • A caller who is not a project editor on project_id returns 403.
  • A connector_id that belongs to another tenant returns 403 — a source can never bind a project to a connector outside its own tenant.
  • An unknown connector_id or project_id returns 404.
  • A project may hold at most 50 sync sources (all providers); registering beyond that returns 400. This bounds the recurring outbound-fetch fan-out from the gateway's egress.

The url provider (websites) — accepted shape and what is rejected

  • No connector_id. Supplying one with provider: "url" returns 400.
  • The source_ref URL is canonicalised (scheme + host lower-cased; path/query preserved) and stored in that form; a missing scheme defaults to https://. Two URLs that differ only by host/scheme case are the same source; http vs https, a trailing slash, or a www. prefix are treated as distinct sources.
  • The URL is SSRF-checked at registration and again on every fetch (including each redirect hop): a private/internal/loopback/link-local address (e.g. 127.0.0.1, 10.0.0.0/8, 169.254.169.254), a DNS name that resolves to one, or a non-http(s) scheme is rejected with 400 at create (fixed body "URL not allowed" — no internal-reachability oracle) and blocked at fetch time.
  • Duplicate URL (same canonical form already a url source in the project) → 409.
  • At fetch time (during a sync, not at registration) the page is fetched fail-closed: the response must be text/html or application/xhtml+xml (an absent or other content-type is treated as non-HTML and the page is not indexed), at most 10 MiB (a larger body is aborted), and reachable. A dead 404/410 removes the previously-indexed page; any transient failure (DNS, timeout, 5xx, 401/403, non-HTML, oversize, blocked redirect) surfaces a sync error and retains the last successfully-indexed copy — a blip never empties the knowledge base. Only <script>/<style>/<noscript> content is dropped (including the self-closing <.../> forms — so text hidden from a JS-enabled reader never reaches the index); the visible page text, including any <title>, is indexed.
  • Recurring egress: like the other providers a url source defaults to 24 h automatic re-sync — the gateway will re-fetch the tenant-supplied public URL on that cadence until auto-sync is set to 0. The per-project source cap above bounds the total fan-out.

Listing sync sources

GET /admin/v1/connector-sync?project_id=<id>

Returns every sync source registered on the project. Required role: project viewer (any project member).

The response is a JSON array — [] when the project has no sources. Each element is the full source object:

Field Type Description
id string Source id.
tenant_id string Owning tenant.
project_id string Target project.
connector_id string Connector used to reach the source.
provider sharepoint | confluence | url Source system.
source_ref string The source selector.
source_label string | null Human-readable label.
share_scope creator | org Whether documents imported from this source are visible only to their creator or shared with the whole organization (set via the source's share endpoint).
status string See the status table above.
crawl_started_at unix seconds | null When the current/last crawl began.
running_since unix seconds | null When the current run started (null when not running).
last_synced_at unix seconds | null When the last successful pass completed.
last_error string | null Reason the last crawl failed, when status is error.
next_earliest_run unix seconds | null Earliest time a run may start again (set by a rate-limit or auto-sync backoff).
auto_sync_interval_secs integer The automatic re-sync cadence in seconds. 0 = off (manual "Sync now" only); a value in 3600–604800 = auto re-sync every N seconds.
items_synced integer Documents currently held from this source.
items_skipped integer Documents the last run could not import (unsupported type or a transient fetch/stage error). Persisted so the "could not be imported" warning survives a page reload; overwritten each run (last-run count, 0 after a clean pass).
created_by string The user the sync runs as.
created_at unix seconds Registration time.
updated_at unix seconds Last update time.

Accepted shape / What is rejected: project_id is required; a request without it returns 400. The listing is scoped to the caller's tenant and to a project they are a member of.


Running a sync ("Sync now")

POST /admin/v1/connector-sync/{id}/run

Runs one crawl pass of the source. Required role: the caller must be both a project editor and the source's creator (created_by) — this is the run-as-invoker boundary described above. A caller who is a project editor but not the creator (or vice versa) returns 403.

On success the response is 200 with the pass result:

{
  "status": "completed",
  "items_synced": 12,
  "skipped": 3
}
Field Type Description
status string The source status after this pass (completed, paused_cap, paused_rate_limited, or error).
items_synced integer Documents newly ingested, updated, or re-confirmed unchanged this pass.
skipped integer Items that could not be ingested this pass: an unsupported file type (always skipped), or a transient fetch/stage error (a later run may pick it up).
error string Present only when the pass ended in error.

Supported file types. A synced document is ingested only when the async extraction pipeline can extract it: PDF, DOCX, PPTX, XLSX/XLSM, OpenDocument spreadsheets (.ods) and text documents (.odt), the text formats (.txt, .md, .markdown, .csv, .yaml, .yml, .rst, .log), structured text (.xml, .json, .html, .htm), and images (.png, .jpg, .jpeg, .webp, OCR-extracted) — the same set the project knowledge upload form accepts. Other types — including legacy .xls (no extractor) — are skipped and counted in skipped, never reported as failed. A synced document is ingested identically to one uploaded directly in the knowledge UI (both take the async path).

Structured formats are validated at the trust boundary and a hostile file is rejected (skipped), not ingested: an .xml that declares a DOCTYPE or any entity (billion-laughs / XXE / an external SYSTEM DTD) is refused before parsing, and a .json that is not valid UTF-8 JSON — or contains NaN/Infinity — is refused. Accepted .xml is reduced to its text content; accepted .json is pretty-printed so each key/value segments for retrieval.

When a run cannot start, the response is 409 with { "conflict": true }. This happens when a sync is already running for the source, or when the rate-limit backoff has not yet elapsed (next_earliest_run is in the future). Retry once the running pass finishes or the backoff expires.

Accepted shape / What is rejected (fail-closed):

  • A caller who is not both a project editor and the source's creator returns 403.
  • A source id that is not in the caller's tenant returns 404 — it is never resolved outside the tenant you can access.
  • A concurrent run or an un-elapsed backoff returns 409 { "conflict": true }, never a partial double-crawl.
  • Third-party responses are validated at the trust boundary and treated as hostile. Every item the Microsoft Graph or Confluence API returns is checked before ingestion: an item with a null or missing required field is rejected, an oversize payload is rejected, and a document download URL that is not https is refused. The connector's bearer token is sent only to graph.microsoft.com (the Graph API host) — it is never attached to a document downloadUrl, so a malicious or redirected download target can never capture the caller's token. A source that returns malformed data fails closed (status error) rather than ingesting unvalidated content.

Setting the sharing scope of a Confluence source

PATCH /admin/v1/connector-sync/{id}/share

Publishes a Confluence sync source org-wide, or reverts it to the creator-only default. Required role: tenant admin (a project editor/member is not sufficient — org-wide publishing is a tenant-governance action). This is the admin org-wide opt-in described above: publishing makes every page the source synced visible to everyone who can open the project, including pages restricted in Confluence (admin curation, not permission mirroring).

Request body:

{ "scope": "org" }
Field Type Notes
scope "org" | "creator" org publishes the source org-wide; creator reverts to the creator-only default (visible only to the sync creator). Any other value is rejected.

The response is 200 with { "id": "<source id>", "share_scope": "<scope>" }. Reverting to creator takes effect immediately — no re-sync is needed (the underlying creator grant was never removed).

Accepted shape / What is rejected (fail-closed):

  • A caller who is not a tenant admin (or platform admin) returns 403; an unauthenticated caller returns 401. The client is never the authz boundary.
  • A source id that is not in the caller's tenant returns 404 — it is never resolved outside the tenant you can access, so a tenant admin can never share another tenant's source.
  • A non-Confluence source (e.g. SharePoint) returns 400: org-wide sharing deliberately bypasses per-item source ACLs, which is valid only for Confluence. SharePoint keeps its real per-item ACL (see above) and cannot be org-shared here.
  • A scope that is absent, non-string, or not exactly "org"/"creator" returns 400. Nothing is written.

Automatic re-sync

Every source can re-crawl itself on a schedule, so connector-sourced knowledge stays current without anyone clicking "Sync now".

  • New sources default to every 24 hours (auto_sync_interval_secs: 86400). Registering a source with POST /admin/v1/connector-sync and no auto_sync_interval_secs opts it into daily re-sync.
  • Existing sources (registered before this feature) start OFF (auto_sync_interval_secs: 0) and re-sync only when you set a cadence — no source silently begins re-crawling on upgrade.
  • Per-source opt-out / retune at any time with the schedule endpoint below: 0 turns automatic re-sync off (manual "Sync now" still works); 3600–604800 sets a cadence of 1 hour to 7 days.

How the cadence is honoured. Due-ness is derived from last_synced_at: a source becomes due again auto_sync_interval_secs seconds after its last successful crawl. A manual "Sync now" therefore also pushes the next automatic run out by a full interval (no redundant crawl right after a manual one). The scheduler re-crawls due sources a few at a time; a run that pauses (ingest cap or rate limit) simply stays due and resumes on the next tick, so the cadence is "at least every N seconds", drifting by at most one crawl duration — deliberately spread out to level load, not a to-the-second guarantee. A source whose creator has lost project access, or whose connector was removed, is skipped and backed off rather than run.

Setting the auto-sync cadence of a source

PATCH /admin/v1/connector-sync/{id}/schedule

Enables, disables, or retunes a source's automatic re-sync. Required role: project editor (owner or editor) on the source's project — the same gate as "Sync now" and delete.

Request body:

{ "auto_sync_interval_secs": 86400 }
Field Type Notes
auto_sync_interval_secs integer 0 disables automatic re-sync (manual "Sync now" still works); an integer in 3600–604800 sets the cadence (1 hour to 7 days).

The response is 200 with { "id": "<source id>", "auto_sync_interval_secs": <secs> }.

Accepted shape / What is rejected (fail-closed):

  • auto_sync_interval_secs is required; a body without it returns 400.
  • The value must be an integer that is 0 or in 3600–604800. A non-integer, a value below the 1-hour floor (a per-minute connector crawl is deliberately not allowed), a value above the 7-day ceiling, or non-numeric junk returns 400 and nothing is written.
  • A caller who is not a project editor returns 403; an unauthenticated caller returns 401.
  • A source id not in the caller's tenant returns 404 — it is never resolved outside the tenant you can access.

Deleting a sync source

DELETE /admin/v1/connector-sync/{id}

Removes the source registration and all knowledge rows this source produced in the project. Required role: project editor.

The response is 200 with { "deleted": true }. Deleting a source detaches its synced documents from the project knowledge base (they are removed, not orphaned); uploaded files and other sources are untouched. A source id outside the caller's tenant is not resolved (404).


Setup: the SharePoint Graph-scoped OAuth connector

SharePoint sync uses a separate connector from the SharePoint MCP tool connector. The MCP tool connector lets the assistant call SharePoint tools during a chat; connector sync uses a dedicated Graph-scoped OAuth connector that reads drive contents through the Microsoft Graph API. The two are distinct registrations and are not interchangeable.

The Graph-scoped OAuth connector is a delegated OAuth connection requesting the scopes:

  • Sites.Read.All
  • Files.Read.All

To set up SharePoint sync:

  1. In Microsoft Entra, register an enterprise application and grant it the delegated Graph scopes Sites.Read.All and Files.Read.All.
  2. In the gateway, create and connect the Graph-scoped OAuth connector using that Entra app (supply your Directory (tenant) GUID and the enterprise-app client id — see the tenant-scoped connector setup). Completing the delegated OAuth sign-in stores the per-user credential the sync will run under.
  3. Add a sync source with POST /admin/v1/connector-sync, referencing that connector_id, provider: "sharepoint", and the drive's source_ref.

Confluence sync needs no separate OAuth connector: it uses the existing Atlassian MCP connector, with provider: "confluence" and the space key as source_ref.