Skip to content

Web search

Myra AI Workspace augments model requests with live web search results. When web search is active, the gateway intercepts the outbound request, retrieves relevant search results, and injects them into the conversation context before forwarding to the model. No changes to client code are required beyond an optional request header.

How it works

Two-leg agentic flow (most providers)

For the majority of supported providers, web search uses a two-leg agentic loop:

  1. Leg 1 (buffered): The gateway injects a web_search tool into the request and makes a non-streaming, buffered call to the provider.
  2. Tool decision: If the model answers directly without calling the tool, the Leg 1 response is returned to the client as-is. No search is performed.
  3. Search: If the model calls the web_search tool, the gateway runs parallel Brave Search queries.
  4. URL fetch: The gateway fetches up to two of the top result URLs in parallel and appends the page content to the search results. The same fetch is used when a message references a page directly (for example a bare domain such as example.com/page). The fetch follows up to five HTTP redirects — so a bare domain that redirects to its www host (or any Location redirect) still yields the real page rather than a redirect stub — and re-applies the SSRF safety check to every redirect target before connecting, so a redirect to a private, loopback, or link-local address is rejected. The URL itself is validated at that same gate before the first connection: a URL carrying a control character (a CR/LF, a tab, a NUL — the shape of a request-line or header injection) is refused outright as an unsupported address, never sent; percent-encoded bytes are sent as they are. A page that cannot be read returns a clear "could not read the page" signal to the model instead of empty content — and that signal names the website's condition, never a gateway fault: a login wall (401), a site that blocks automated access (403), a paywall (402), a missing page (404), a site rate-limiting requests (429), a server error (any 5xx other than 504 — "may work later"), any other refusal (400/405/410/422/451…, described as a refusal, not an outage or a content problem), a request the site rejects as too large (413), a redirect loop (over the five-hop limit, with the reason), a site that does not respond within the fetch deadline — or answers 408 or 504 — ("did not respond in time"), a connection or TLS failure on the remote side ("could not be reached"), a blocked address, non-text content, or a page larger than the 20 MiB fetch cap (reported as unreadable content, like a non-text page — the site answered, it was not unreachable). In every case the model is told what happened, told to relay it to the user, and told not to invent results. When a page is reduced to readable text, the content of <script>, <style>, and <noscript> elements is dropped — not merely their tags — so text a JS-enabled reader never sees (inline script source, CSS, <noscript> fallbacks) cannot reach the model as page content; an unclosed such element is dropped through to the end of the page (fail-closed). The same rule is applied to search-result titles and snippets.

Documents (PDF, DOCX, XLSX). The top-two auto-fetch described here now also reads a linked PDF, DOCX, or XLSX among the search results, extracting its embedded text layer so the model can answer over the document — a PDF surfaced by the search is no longer skipped. To keep an interactive turn fast and off the inference fleet, the auto-fetch is text-layer only: it does not OCR. A scanned-image PDF (no text layer) contributes no text and is disclosed to the model (N of M pages are scanned images…); a fully-scanned document returns a short "no extractable text layer — fetch it directly to OCR" marker. When the model instead fetches a URL directly (the fetch_url / agentic_fetch tools — a link it found, or one the user pasted), it additionally OCRs scanned pages. See Document fetch below.

What the SSRF guard accepts and rejects. The fetch only performs an HTTP(S) GET to a public address, and the web-search auto-fetch returns readable text/*/HTML content plus the extracted text layer of a magic-sniffed PDF/DOCX/XLSX (see the Documents note above); every other binary is unreadable. It rejects, fail-closed (the request is never sent), any of: a private, loopback, link-local, unique-local, or cloud-metadata address (127.0.0.0/8, 10/8, 172.16/12, 192.168/16, 169.254/16, ::1, fc00::/7, fe80::/10, IPv4-mapped forms, and the 0.0.0.0/unspecified address); any numeric IP encoding that resolves to one of those (decimal, octal, hex, short a.b/a.b.c forms, and expanded/compressed IPv6); a URL carrying userinfo (user@host, which HTTP clients treat as the connect host); a hostname whose DNS answer includes any such address; and any malformed authority. DNS names are resolved once and the fetch connects to that exact validated IP — it does not re-resolve the name at connect time, so a rebinding DNS answer that returns a public address to the check and a private address to the connect is defeated (the TLS SNI and Host header still carry the original hostname, so certificate verification and virtual-host routing are unchanged). Each redirect hop repeats this resolve-and-pin. An operator allowlist (configured by your operator: exact host:port entries, port required) is the only way to permit an internal address, and it never widens which port is reached. 5. SSE status event: For streaming requests from a client that opted into the aig_* side channel, an aig_status event is emitted before the URL fetch so the client can display a loading indicator. See SSE status event. 6. Leg 2 (streaming): The enriched results are injected into the conversation as tool results, and the final request is forwarded to the provider. Streaming is used if the original request requested it.

Google Gemini — native grounding

For Google Gemini, web search uses the built-in googleSearch grounding feature of Gemini. No Brave Search API call is made, and the api_key in the gateway config is not used for Gemini requests. The gateway automatically converts the web search instruction to the Gemini grounding format. This is a single-leg request with no tool injection.

Model-driven streaming (Myra / vLLM)

For Myra-hosted models (Qwen3, Gemma) served over vLLM, web search is handled inside the gateway's streaming tool loop rather than the buffered two-leg flow above. The gateway offers the model a web_search tool on the live streaming request; when the model calls it, the gateway runs the same Brave Search queries and URL fetch and feeds the results back into the same turn. There is no separate buffered Leg 1 — the tool call and the final answer stream in one loop.

Per-turn cap and within-turn de-duplication

The streaming tool loop lets the model call web_search across many rounds in one turn, and smaller local models (Qwen3, Gemma) tend to over-call it — re-issuing near-identical low-value queries every iteration, which multiplies the paid search-provider (e.g. linkup) egress cost. Two controls bound this, on the model-driven path only:

  • Within-turn de-duplication — repeat and near-identical queries in the same turn (compared after lower-casing, collapsing whitespace, and trimming) are served from the first query's result. No second search is sent to the provider, so there is no duplicate cost, search card, or citation. Sub-second "freshness" between rounds is illusory, and the repeat is exactly the waste being removed.
  • Per-turn cap (max_searches_per_turn, default 8) — a hard ceiling on the number of distinct searches per turn. It sits well above typical legitimate research (roughly 2–6 distinct searches, since de-duplication already removes the near-duplicates) and well below the pathological fanout. On reaching the cap the model receives a short notice instead of a search result and answers from what it has already gathered — never a silent truncation. Raise max_searches_per_turn for research-heavy workloads.

These controls apply only to the model-driven (Myra / vLLM) tool-loop path; the two-leg agentic flow issues its searches in a single leg and is unaffected.

Document fetch (PDF, DOCX, XLSX)

The gateway reads not only web pages but also PDF, DOCX, and XLSX documents, extracting their text so the model can answer over the document's content. This applies on two paths, which differ only in whether scanned pages are OCR'd:

  • Direct fetch — the fetch_url and agentic_fetch tools, offered when a fetchable URL is present in the conversation (or co-armed with web search): reads web pages and documents, and OCRs scanned pages (up to the full 200-page / 300-second budget below).
  • Web-search top-two auto-fetch (step 4): now also extracts a linked PDF/DOCX/XLSX, but text-layer only — no OCR, on a tight per-document budget (default 50 pages / 20 s, tenant-overridable via url_fetch.auto_doc_max_pages / url_fetch.auto_doc_timeout_ms). A scanned page is skipped and disclosed (N of M pages are scanned images…); an all-scanned document returns a "fetch it directly to OCR" marker. This keeps every web search fast and off the inference fleet while still reading the common text-layer report/PDF; the model can always fetch a scanned document directly for OCR. The auto-fetch does not consume the model's per-turn fetch_url budget below.

  • Type detection is by content, not by claim. The document type is determined by the file's magic bytes, never by the server's Content-Type header or the URL's extension (both are untrusted and only used to reject spoofs). A URL ending in .pdf that actually serves an HTML paywall is read as a web page; a text/html-labelled response whose bytes are a real PDF is read as a PDF. Only PDF, DOCX, and XLSX are extracted; images, presentations (PPTX), OpenDocument (ODT), archives, and other binaries are not, and return a clear "could not read" signal.

  • Bounded download. The document body is capped at 20 MB and fetched through the same SSRF guard (see step 4) with a fetch deadline. Larger files are refused rather than read.
  • Scanned PDFs are read via OCR, and labelled low-confidence. A PDF with no embedded text layer (a scan or an image-only export) is transcribed page-by-page by an image-capable vision model, which reads the full page in one pass — the model is instructed to transcribe every table row verbatim rather than summarize or omit. Because OCR can still misread a figure or a date, OCR-derived text reaches the model labelled ("Read via OCR — may contain transcription errors; verify specific figures/dates against the source"), and the model is told to qualify figures it quotes from it rather than present them as authoritative. If a page was too dense to fully transcribe within the output budget, the label also says so.
  • Long documents are read up to 200 pages, and truncation is disclosed. A fetched PDF is read up to a 200-page bound (raised from a previous silent 20-page limit); if the document has more pages than were read, the model is told explicitly ("only the first N of M pages were extracted; content on the remaining pages is missing") so it never confidently denies content that lies past the cut. The real ceiling on very large text is still a byte cap on the extracted text (also disclosed), so page and byte truncation are each surfaced rather than silently dropped.
  • Query-directed page range (pages). When the answer lies on a later page of a long PDF (past the truncation cut), the model can call fetch_url/agentic_fetch again with the same URL and a pages argument — a 1-based range "START-END" (e.g. "110-125") or a single page "117" — to read only that window instead of always the first pages. The model picks the range from the document's table of contents (which is usually within the first pages it already read). The honored range is disclosed factually ("pages X–Y of Z shown", no "missing" clause). pages applies to PDFs only (DOCX/XLSX/web pages ignore it). A range that starts beyond the document, or contains no readable text, returns a served notice ("Requested pages start beyond … the N-page document") — not an error. pages is untrusted model input: a non-string, reversed, zero, or malformed value is ignored and the read falls back to the first pages (see the API reference).
  • Per-turn limit. To bound extraction cost (a scanned PDF is OCR'd), the number of documents extracted in a single turn is capped (default 8, tenant-overridable via the gateway config url_fetch.max_doc_fetches_per_turn — API-managed, no UI control; an invalid or non-positive value falls back to the default). On reaching the cap the model receives a short notice and answers from what it has, or asks the user to narrow the request — never a silent truncation. Note that a document-heavy turn is additionally bounded by the tool loop's own wall-clock budget: each document costs a fetch plus an extraction (up to minutes for a scanned PDF), so raising this cap far above the default mostly moves the bound to the wall clock.
  • Personal-data handling. Extracted document text is treated as untrusted bulk content and passes the same masking/injection-scan boundary as every other tool result before it reaches the model. The agentic_fetch inner-model leg has no such central scrub, so it is guarded explicitly: under a PII mandate (a pii_mandatory project, a forced agent invoke, or a tenant that enforces masking) on a gateway with no reversible masker configured, the fetched document is withheld from the inner model (fail-closed) — unless the inner model is a first-party local model, where nothing leaves the estate and the leg is not a mandate case (inside an agent-shaped run the inner leg is never local); otherwise the extracted text is masked/scanned before the inner model sees it. If that scan cannot run (PII engine outage) the document is withheld under a PII mandate — except when the inner model is a first-party local model, where the leg is not a mandate case — and otherwise the detector's fail_open decides; the inner leg's own request phase then scans the document again with its own locality and records any outage on the request — for every detector that phase runs (an enabled request-phase masker, an enabled injection classifier); a disabled or response-only one is re-run by nobody, so its fail-open forward at this seam is recorded as the gap it is.

The top-two auto-fetch inside a web search opens only the two highest-ranked result URLs (now including a PDF/DOCX/XLSX there, text-layer only); when the answer sits behind a link on one of those pages — a sub-page, a deeper document, or a scanned PDF that needs OCR — the model needs to fetch that link itself. To make that possible within the same turn, whenever the gateway web search tool is offered, the gateway also offers fetch_url / agentic_fetch — even when no URL was present in the conversation. The fetch tool is therefore already armed when the search results land, and the model can open a link it found in them.

  • Only for the gateway search path. The co-arm keys off the gateway's own web_search tool being offered — the case where results come back as a tool result the model can read a URL from. It does not arm fetch for providers that run search server-side and fetch their own pages (Groq compound models, or a native-search provider such as Anthropic that is not forced onto the gateway two-leg). Under the EU residency floor those native providers are forced onto the gateway two-leg, so there fetch is co-armed.
  • Per-turn fetch limit. To bound fan-out, the total number of URL fetches in a single turn (web pages and documents combined) is capped (default 24, raised from 8 so a multi-source research turn — e.g. 14-20 pasted source URLs — completes without punting mid-task); the document-extraction sub-cap (default 8, above) still applies within that budget. Both caps are tenant-overridable via the gateway config url_fetch section (max_fetches_per_turn / max_doc_fetches_per_turn — API-managed, no UI control; invalid, zero, negative, or non-finite values fall back to the defaults, so a config typo can never uncap or brick the tool). On reaching either cap the model receives a short notice naming the effective limit and answers from what it has — the two caps produce distinct messages so a page fetch tripping the total limit never misreports a document limit.
  • Guards unchanged. Every co-armed fetch still passes the full execution boundary per hop: the SSRF/redirect guard, the outbound personal-data/secret exfiltration check (fail-closed), the residency re-check, and the untrusted-framing + injection scan on the fetched content. A local_only (no-egress) workspace still blocks fetch_url entirely, even with web search on.

Search-first behavior

When web search is enabled for a turn, the assistant is instructed to run the search directly rather than asking the user for permission first — enabling the feature is itself the go-ahead. The model still uses judgement: it skips the search for requests that plainly need no external information (greetings, rewriting or translating text you provided, or a follow-up answerable from the conversation so far). The tool is only offered (never forced), so a turn that genuinely needs no search is never made to run one.

The instruction also guards against a smaller-model failure mode where the assistant announces a search ("Let me look that up…") and then stops without actually calling the tool, or claims it checked a source and invents the result. When external information is needed, the model is told to emit the search call itself on the same turn rather than describe the search it would run; and it must not present anything as looked-up, checked, or verified unless a tool actually returned it (that turn or earlier in the same conversation). Answering from the model's own stable knowledge — without dressing it up as a lookup — remains allowed.

Personal data in the search query

On a gateway that runs a PII layer, the model-generated search query is scrubbed before it egresses to the search provider. Every detected personal-data entity is replaced with a reversible token, so the search backend (Brave or linkup) — and the X-Web-Search-Query response header, the logs, and the trace — only ever see the tokenised query, never the raw personal data. The token map is shared with the request phase, so an entity gets the same token throughout the turn, and the final answer is restored to the real values before it reaches the user.

This applies to the gateway two-leg flow. Provider-native search (Anthropic, Gemini, and others that run the search on their own server) is covered by this scrub only when it is forced onto the gateway two-leg under the EU residency floor (see EU data residency); otherwise the query reaches the provider search backend without the gateway scrub.

The scrub activates only when the gateway has a PII detector. On a gateway with no PII layer, the query is not tokenised.

Under a PII mandate (a pii_mandatory project, or a tenant with masking enforced) the query scrub applies the German PERSON / LOCATION floor whatever the detector's language says — including on a gateway whose every model is first-party (local-only), where the model leg itself is exempt from the mandate: the query's bytes leave for the search provider, so the exemption stops at the model leg. Such a request records the pii_egress_floor_forced_by_mandate marker in its log (unless the detector is German-calibrated, where the floor was already on); the cost is that a place name in an English query may reach the search provider as a redaction token.

⚠️ Caution: Under a PII mandate — a pii_mandatory project, a forced agent invoke, or a tenant that enforces masking (an unresolvable project tier blocks every egress tool outright, one step earlier) — a gateway that cannot tokenise the query (no pii_protector) withholds the search rather than sending it, and the model is told the workspace requires personal-data masking (the same wording on both search engines — the older engine used to claim a "PII scan failed"). When the PII scan is unavailable, the outcome depends on whether a PII mandate is in force: under a mandate (a pii_mandatory project, or a tenant with masking enforced — including a local-only gateway, whose model leg is exempt but whose query still leaves) the search is withheld whatever the detector's fail_open says, and the request records pii_mandate_masker_unavailable. Outside a mandate the detector's fail_open decides: false (the seam's default) withholds the search and the reply reports that web search is unavailable; true forwards the query with at most the offline German-ID masking applied — and that forward is recorded on the request as a guardrail gap (guardrail_degraded, pii_scan_degraded, guardrail_error) once the search actually leaves: on a multi-query turn where a later query is withheld, every query is withheld and nothing is recorded as a gap.

Supported providers

The web search mechanism depends on the provider.

Native search:

  • Anthropic (native endpoint) — the provider runs the search server-side.
  • Google Gemini — native grounding, a single leg with no tool injection.
  • Google Vertex AI — native grounding, the same single-leg mechanism as Google Gemini (Vertex hosts the same Gemini models).

Two-leg agentic loop (tool injection):

  • OpenAI
  • DeepSeek
  • Cerebras
  • Together AI
  • Fireworks
  • xAI
  • HuggingFace
  • SambaNova
  • NVIDIA
  • Azure OpenAI
  • Cloudflare Workers AI

Mistral Agents path:

  • Mistral — web search is routed to the Mistral Agents endpoint (/v1/conversations).

Server-side plugin:

  • OpenRouter — web search runs server-side via OpenRouter's default web plugin. OpenRouter uses the underlying model's native search when it is available and otherwise falls back to its own web search, so every routed model gets working web search without a gateway two-leg loop.

Model-driven streaming:

Not supported:

  • AWS Bedrock, Perplexity, Cohere, and Groq models.

💡 Note: For Cohere and Groq, a request with web search enabled returns the web_search_not_supported error. For the other unsupported providers, the request proceeds without web search.

💡 Note: The search backend for the two-leg and streaming flows above is pluggable via web_search.provider: Brave Search (the default) or linkup.so (EU-based; SOC 2 II, Art. 28 DPA, zero-retention — the EU-sovereign option for a residency-enforced gateway). Only linkup queries are metered into the tenant's spend and budget cap (each successfully-answered query draws down the same allowance as model usage); Brave queries are not currently metered. The linkup per-query price is a model_price catalog row (provider='linkup', model='web_search', input_per_1k = 10.0 = $0.01/query), editable on the Provider Costs page (Settings › Costs); a deployment default applies as the fallback when that row is absent. See web-search cost.

Result language

The gateway detects the language of the user's own words in the latest message server-side (the same detector that pins the reply language — text attached or pasted into the turn is set aside first, so an English question about a German document searches in English) and, for Brave searches, sends Brave's language/region targeting so results come back in the asked language rather than the language of the ambient context:

  • search_lang — the content language (ISO 639-1: de, en, fr, nl).
  • ui_lang — the response-metadata language (de-DE, en-US, fr-FR, nl-NL).
  • country — a 2-char region, sent only where it is derivable from the language (DE/FR/NL). English is left without a country (it spans many regions; Brave defaults to US) — search_lang=en already pins the result language.

Detection is precision-biased and fails open: if the message language is ambiguous or unsupported, no language parameter is sent and Brave uses its default behaviour (exactly as before this feature). Only the four supported languages steer the search.

linkup has no request-level language or country parameter, so its results cannot be pinned this way; the lever for linkup is the reply-language directive, which makes the model answer in the question's language regardless of the source language of the results.

The gateway's own web search (the two-leg agentic loop) can run against an EU-sovereign search provider — linkup.so (France; DPA, zero-retention). Provider-native search (Anthropic's server-side search, OpenRouter's web plugin, Gemini/Vertex grounding, Mistral Agents) instead egresses the model-generated search query to the provider's own (non-EU) search backend, which breaks an EU-residency commitment. On a non-enforced gateway this native egress is not blocked but it is no longer invisible: it records a non_eu residency-evidence leg for the data-localization report — see EU data residency — search-egress jurisdiction.

To keep the query inside the EU, the gateway forces gateway search — it disables provider-native search tools and knobs and routes web search through the gateway's own (linkup) two-leg — when either:

  • the gateway's EU residency floor is armed (eu_region_routing, or the inherited tenant floor / deployment default); or
  • the gateway is configured with an EU search provider (web_search.provider: "linkup").

A gateway configured with a non-EU provider (brave) and not enforced keeps provider-native search unchanged — forced gateway search has zero effect on non-sovereign deployments.

Under forced gateway search:

  • Anthropic, OpenRouter, and Mistral models run web search through the gateway's linkup two-leg (they support the gateway tool loop). Any client-supplied native search tool (web_search_20250305, the OpenRouter web plugin or :online model suffix, OpenAI web_search_options) is stripped from the outbound request.
  • Gemini / Vertex and a raw /v1/messages (native Anthropic SDK / Claude Code) client have no gateway two-leg path, so web search fails closed — it is disabled rather than egressing to a non-EU backend. No search runs; the answer is produced without live web results.

The stripped native paths are the per-request search tools and knobs: Anthropic's web_search_20250305, the OpenRouter web plugin and :online model suffix, OpenAI's web_search_options, and Gemini/Vertex grounding. Search that is intrinsic to a model slug (for example gpt-*-search-preview, Groq compound, Perplexity sonar) runs on the model's own server regardless of tools — such a model is guaranteed off only under the enforced residency floor, where the model itself is dispatch-blocked as a non-EU provider. On a gateway that only configures an EU search provider without arming the floor, those models still reach their (non-EU) backend for the whole request, so pair the EU search provider with the residency floor for full sovereignty.

Configuration

Web search is configured at the gateway level under config.web_search.

Screenshot: the Web Search section of the Edit Gateway dialog The Web Search section of the Edit Gateway dialog

{
  "config": {
    "web_search": {
      "enabled": true,
      "api_key": "BSA...",
      "max_results": 5,
      "mode": "opt-in",
      "default_freshness": "pw"
    }
  }
}
Field Type Default Description
enabled boolean false Enables web search on this gateway. A null value is treated as disabled (fail-closed).
provider string "brave" Gateway search backend: "brave" (US) or "linkup" (EU-sovereign). Configuring "linkup" forces gateway search over provider-native search (see EU data residency). Omit it to inherit the tenant default (web_search_provider); see below.
api_key string — Search provider API key (Brave or linkup). Required for all providers except Google Gemini.
max_results integer 5 Maximum number of search results to retrieve per query.
max_searches_per_turn integer 8 Maximum number of distinct web searches the model may run in ONE turn on the model-driven (Myra/vLLM) tool-loop path. Repeated/near-identical queries within a turn are de-duplicated (served from the first result, not re-fetched). On reaching the cap the model receives a notice and answers from the results already gathered. Raise it for research-heavy workloads.
mode string "opt-in" Controls when web search is triggered. See Modes.
default_freshness string — (unset → unfiltered) Default recency window applied server-side when a request sends no x-aig-web-search-freshness header. Accepts the same values as that header: pd / pw / pm / py (past day/week/month/year) or an explicit inclusive range YYYY-MM-DDtoYYYY-MM-DD. Any other value is ignored (fail-closed → no narrowing from this source). Unset by default, so older/historical results are never silently excluded; set it only for a gateway whose queries are predominantly time-sensitive. A valid per-request header always overrides it. Brave only — the linkup adapter does not apply a recency window.

Tenant provider default

A tenant carries a web_search_provider default (brave | linkup, set via PATCH /admin/v1/tenants/{id} — see Tenants API). A gateway that has a web_search block but sets no provider of its own inherits that tenant default; a per-gateway provider overrides it. Under an armed EU residency floor (eu_region_routing), a non-EU per-gateway provider is clamped up to an EU tenant default (so a sovereign tenant's gateways cannot silently sit on Brave), while a gateway may still choose another EU provider. The tenant default never enables web search on a gateway that has no web_search block, and the gateway's api_key must match the effective provider.

Default on every new gateway (platform Linkup key)

Every new production gateway — created at self-service signup (trial or paid) or by an administrator — starts with web search enabled via Linkup and no key of its own. You do not need to supply an api_key: when a platform admin has configured the shared platform key (the setting trial_linkup_api_key — the name is historical, it applies to every plan), the gateway serves web search through the EU-sovereign Linkup backend. The key is injected into the gateway's web_search config at read time, is never stored on the gateway, and is never returned to any client. A gateway that already carries its own web_search block when it is created (your own Linkup or Brave key, another provider, or an explicit enabled: false) keeps it exactly as supplied — the default is never forced over a deliberate choice.

If the platform key is unset, such gateways show web search as set up but not available (the composer shows a notice), the Health dashboard flags Platform web-search key: MISSING, and the platform is alerted — nothing is silently off. Setting the key from the admin console takes effect within about a minute for every affected gateway, with no redeploy and no backfill. It was formerly a deployment setting and formerly applied to trials only.

A trial that converts to a paid plan keeps its web search unchanged. To use your own Linkup or Brave account instead of the platform key, set the gateway's web_search.api_key (and provider) — an explicit key always takes precedence.

Modes

Mode Behaviour
"opt-in" Web search triggers only when the client includes the X-AIG-Web-Search: 1 request header.
"always" Web search is attempted on every request through this gateway, regardless of client headers.

Getting a Brave Search API key

Obtain a Brave Search API key from the Brave Search API developer portal. The Free tier provides 2,000 queries per month; paid plans offer higher limits.

💡 Note: For Google Gemini, web search uses the built-in googleSearch grounding feature of Gemini rather than Brave Search. The api_key in the gateway config is not used for Gemini requests. The gateway automatically converts the web_search tool to the Gemini grounding format.

Request and response headers

Client opt-in header (required when mode is "opt-in"):

X-AIG-Web-Search: 1

Client recency header (optional):

X-AIG-Web-Search-Freshness: pw

Narrows results to a recency window: pd / pw / pm / py (past day/week/month/year) or an explicit inclusive range YYYY-MM-DDtoYYYY-MM-DD. Any other value — including an empty string — is ignored (fail-closed → unfiltered). When the header is absent or invalid, the gateway falls back to the per-gateway default_freshness knob (itself unset by default → unfiltered). A valid header always wins over the gateway default. Brave only — the linkup adapter does not apply a recency window.

Response header (always set when a search is performed):

X-Web-Search-Query: <search query used>

Result dates

When the search backend reports a result's publication / last-seen date (Brave's page_age / age), the gateway surfaces it to the model on an as of <date> line in the search-results listing, and labels the corresponding fetched page in the ## Page Content section as ### Source: <url> (as of <date>). This lets the model judge how current each source is instead of guessing. The date string is third-party content, so it is treated as untrusted: any markup, newlines, and control characters are stripped and the value is length-clamped before it reaches the model. linkup results carry no per-result date, so no as of line is shown for them.

SSE status event

For streaming requests from a client that opted into the gateway's aig_* side channel (x-aig-turn-id, as the chat UI always sends, or x-aig-extensions: 1), the gateway emits an aig_status server-sent event (SSE) before fetching URL content. Clients use this to display a loading indicator. A plain OpenAI-compatible client receives a pure OpenAI stream and none of these events — see Inference — what a plain client receives.

data: {"aig_status": "fetching", "count": 2}

The count field indicates the number of URLs being fetched.

Live search results (web_search_results)

For streaming requests from a client that opted into the aig_* side channel, the gateway emits a web_search_results event immediately after each successful web_search — both on the tool-loop path (Myra-hosted models) and on the Anthropic-native path (Claude models run the search server-side and stream a web_search_tool_result block) — so the client can render the live "Searched the web" card (query line + a source list) identically on either provider:

data: {"aig_status": "web_search_results", "query": "bitcoin price today USD", "count": 3, "results": [
  {"title": "CoinMarketCap", "url": "https://coinmarketcap.com/"},
  {"title": "Coinbase", "url": "https://www.coinbase.com/"},
  {"title": "Yahoo Finance", "url": "https://finance.yahoo.com/"}
]}
  • query is the original search query (the card headline).
  • results is capped at 10 entries; each carries a length-clamped title and an http(s) url (the client derives the display domain from the URL). Consumers must treat every field as untrusted input and ignore or clamp malformed entries. Length clamping is UTF-8-aware — a title or query that exceeds its byte budget is truncated on a character boundary, never mid-codepoint, so the persisted utf8mb4 blob is always valid (a byte-split multibyte character would otherwise reject the whole turn's write).
  • The event is emitted only on streaming turns to a client that opted into the aig_* side channel; a buffered turn (agent invoke, playground preview) omits it, and so does a plain OpenAI-compatible client's stream. A search that returned no results emits no event.
  • The per-query steps are also persisted on the assistant message, so the collapsed "Searched the web" disclosure rehydrates identically on conversation reload.

Enabling web search on a gateway

Proceed as follows to enable web search on a gateway:

  1. Send a PATCH request to /admin/v1/gateways/{id} with the web_search configuration block.
curl -X PATCH "https://<your-gateway-host>/admin/v1/gateways/{id}" \
  -H "Cookie: aig_admin=<SESSION>" \
  -H "Content-Type: application/json" \
  -d '{
    "config": {
      "web_search": {
        "enabled": true,
        "api_key": "BSA_your_brave_key",
        "max_results": 5,
        "mode": "opt-in"
      }
    }
  }'
  • The gateway configuration is updated immediately.

-> Web search is active for all subsequent requests through this gateway.

Making a search-augmented request in opt-in mode

Proceed as follows to send a search-augmented request when mode is "opt-in":

  1. Include the X-AIG-Web-Search: 1 header in the client request.
curl -X POST "https://<your-gateway-host>/v1/myapp/prod/openai/chat/completions" \
  -H "x-aig-token: <token>" \
  -H "X-AIG-Web-Search: 1" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "What is the latest news about AI?"}],
    "stream": true
  }'
  • The gateway runs the two-leg agentic flow and injects search results before the final model call.

-> The model response includes context from live web search results.

⚠️ Caution: Not all models support tool use. Web search requires the model to support tool/function calling. If the model does not call the tool, the Leg 1 response is returned directly and no search is performed.

💡 Note: Web search is not supported when using the OpenAI-compatible endpoint with Anthropic as the provider. Use the native Anthropic endpoint instead.

See also