Skip to content

Response caching

Myra AI Workspace implements an exact-match response cache that short-circuits the entire upstream call when a matching request is found. Cache hits return the stored response immediately, saving both cost and latency.


When to enable caching

Enable caching when your workload includes repeated identical prompts — for example, a support chatbot that frequently receives the same questions, or a pipeline that runs the same classification prompt over many inputs.

Not useful for: conversational flows where each message is unique, or prompts that vary by temperature or other parameters (any change to the request produces a different cache key and a cache miss).

💡 Note: The cache is purely exact-match. Two requests with identical prompts but different temperature or max_tokens values do not share a cache entry. Only byte-for-byte identical requests hit the cache.

⭐ Example: How the cache key is constructed — The cache key is computed from the tenant id, the gateway id, the provider name, the model name as requested (a retired Myra-fleet id and its successor are two different keys, and a hit on the old id serves the stored entry without the upgrade headers — see Retired model ids), the request body, and three request-scoping headers: x-project-id, x-aig-web-search, and x-aig-knowledge-ids (normalised/sorted). So two requests with identical bodies but a different project, web-search setting, or bound knowledge set do not share a cache entry. Only the stream and user fields are excluded from the body before hashing — these are delivery preferences that do not affect the response of the model. All other fields (messages, temperature, max_tokens, system, tools, metadata, etc.) are included. Field order within the JSON object does not matter. If either of the two single-valued scoping headers (x-project-id or x-aig-web-search) is sent more than once, the request is not cached at all (neither read nor written): a duplicated single-value header names no single project or setting, and a guessed one would let a differently-scoped answer share a key with a normal request. (A repeated x-aig-knowledge-ids is the exception — because a knowledge selection is inherently a list, its duplicate values are merged and the request is still cached normally.) A request that carries x-aig-knowledge-exclude-ids is likewise not cached at all (neither the exact-match nor the semantic cache), for a different reason: an exclusion describes an open-ended set — "every file except these" — whose membership changes the moment a file is added to the project, so a cached answer for it could not be invalidated correctly.

Because the key is scoped per tenant and per gateway, two gateways — even within the same tenant — never share cache entries. This keeps differing guardrail or PII policies isolated: a gateway that scrubs PII can never be served a raw completion cached by a gateway that does not. The cache fails closed: if the gateway identity is absent, caching is disabled rather than falling back to a shared namespace.


TTL configuration

Cache TTL (time to live) is configured per gateway in the gateway config JSON:

{
  "cache_ttl": 300
}
Value Behaviour
0 (default) Caching disabled
> 0 Cache entries live for this many seconds

💡 Note: There is no way to set a different TTL per model or per token — the TTL applies to all cache entries for the gateway.


Cache hit behaviour

On a cache hit:

  1. The stored response body is returned immediately with HTTP 200.
  2. The X-AIG-Cache: HIT response header is set.
  3. The provider is not called.
  4. The log entry records cached = true, saved_cost_usd, and saved_latency_ms.
# Verify a cache hit
curl -i -X POST https://<your-gateway-host>/v1/myapp/prod/openai/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"gpt-4o","messages":[{"role":"user","content":"What is 2+2?"}]}'

# Second identical request
HTTP/1.1 200 OK
X-AIG-Cache: HIT

Savings tracking

For every cache hit, the gateway logs the following fields:

Field Description
saved_cost_usd The cost_usd value stored when the entry was written (the cost that would have been incurred)
saved_latency_ms Estimated upstream latency saved, based on the average upstream latency for the provider/model

These fields are visible in the request logs and aggregated in the Stats API response under the today, hour, and other period windows.


What is not cached

The following responses are never written to the cache:

  • Streaming responses — "stream": true requests are passed through as SSE and cannot be buffered for caching.
  • Non-200 responses — provider errors, gateway blocks, and rate limit responses are not cached.
  • Requests when cache_ttl is 0 — the default; caching must be explicitly enabled per gateway.
  • Agent invokes — requests to /v1/{tenant}/{gateway}/agents/{slug}/invoke are never served from or written to the response cache, even when cache_ttl is set on the gateway. Agent runs (especially scheduled or tool-using agents such as web search) must return fresh output on every run, and an agent invoke must always flow through the full response path so its fail-closed PII egress header is emitted for the scheduler. This applies to every agent-invoke entry point (scheduled runs, the admin "Try it" preview, agent-as-tool delegation, and workflow steps).
  • Requests carrying MCP tools — a request that references an MCP connector (tools: [{ type: "mcp", … }]) is never cached, so a per-user connector result can't leak across users.
  • Web-search turns — a turn that runs a live web search (the gateway's web_search is configured and either its mode is "always" or the request opts in with the x-aig-web-search header set to a plain value other than 0/false) is never served from or written to the response cache, even when cache_ttl is set. The answer is grounded in results fetched for that specific request, so a cached completion for an identical body (e.g. "weather now?") would otherwise be replayed from a prior session for up to cache_ttl. A globe-off turn (header 0/false, not always mode) is a normal turn and still caches. This is decided on exactly the same condition the search engine runs on, so caching and searching cannot disagree.

Disabling cache per request

There is no per-request cache bypass header. To force a cache miss:

  • Change a field that is included in the cache key (e.g. add a unique value to the messages).
  • Temporarily set cache_ttl: 0 on the gateway config.

Semantic caching

Semantic caching is currently disabled (security)

Semantic caching is hard-disabled at the code level and ignores the semantic_cache configuration below. It embedded the full prompt (including personal data) to an external embedding endpoint and stored responses keyed only by gateway+model with no per-user scope — so it could send PII outside the EU region and serve one user's cached answer to another. It will remain off until it can run with an EU-hosted embedder, a per-user/tenant cache key, encryption at rest, and PII masking before embedding. Exact-match caching (above) is unaffected.

Exact-match caching only helps when two requests are byte-for-byte identical. Semantic caching catches near-identical prompts that differ in phrasing — for example, "What is the capital of France?" and "Tell me the capital city of France?" map to essentially the same response.

When enabled, the gateway computes an embedding vector for each incoming prompt and compares it against stored embeddings using cosine similarity. If the best match exceeds the configured threshold, the stored response is returned immediately without calling the provider.

When to use semantic caching

Semantic caching is most effective for:

  • FAQ / support bots — users rephrase the same small set of questions in slightly different ways.
  • Classification pipelines — similar inputs map to identical labels.
  • Content generation with stable topics — "Write a short bio for Einstein" vs "Give me a brief biography of Albert Einstein".

Not recommended for:

  • Conversational flows where context changes with every message.
  • Prompts where small phrasing differences meaningfully change the correct answer.
  • Low-latency requirements (the embedding call adds ~50 ms per miss; see note below).

Configuration

{
  "semantic_cache": {
    "enabled": true,
    "threshold": 0.95,
    "embedding_url": "https://api.openai.com/v1/embeddings",
    "embedding_api_key": "sk-...",
    "embedding_model": "text-embedding-3-small",
    "max_candidates": 100,
    "ttl": 86400
  }
}
Field Type Default Description
enabled boolean false Activates semantic caching
threshold number 0.95 Cosine similarity cutoff. Hits require similarity ≥ threshold.
embedding_url string — OpenAI-compatible embeddings endpoint
embedding_api_key string — Bearer token for the embedding endpoint
embedding_model string text-embedding-3-small Embedding model name
max_candidates integer 100 Maximum stored embeddings to compare per query
ttl integer 86400 Seconds before a stored embedding expires

Threshold guidance

Threshold Behaviour
0.97–1.00 Very strict — only near-identical rephrasing hits
0.95 (default) Balanced — catches common reformulations, avoids false positives
0.92–0.94 Loose — higher hit rate; risk of semantically adjacent but distinct prompts sharing a cached response

Embedding model choice

Any OpenAI-compatible embeddings endpoint works:

  • text-embedding-3-small (OpenAI) — recommended; small, fast, 1536 dimensions
  • text-embedding-3-large (OpenAI) — higher quality; 3072 dimensions, more storage

How a semantic hit is served

On a semantic cache hit:

  1. The stored response body is returned immediately with HTTP 200.
  2. The X-AIG-Cache: SEMANTIC_HIT response header is set.
  3. The X-AIG-Similarity: 0.97 header indicates the cosine similarity score.
  4. The provider is not called.

Latency note

The embedding call adds approximately 50 ms on a cache miss. LLM inference calls typically take between 500 ms and 3 s, so the overhead is small relative to the savings on hits. The embedding storage step on writes is fully asynchronous and never adds latency to the response path.

💡 Note: Requests with "stream": true are not stored in or served from the semantic cache. Exact-match caching also does not apply to streaming requests.


See also