Skip to content

Glossary

Definitions for all key terms used across the Myra AI Workspace documentation, listed alphabetically.


Artifact card

A card in the chat message thread that represents a file the model generated (for example with the write_file tool). The card shows the file name and format and opens the file in a split-screen panel. A Markdown artifact can be exported to PDF, Word, PowerPoint, Excel, or OpenDocument formats. See Easy chat — Generated files.

Auth token

An access token issued by the admin UI or API that authenticates inference requests. The plaintext token is returned once at creation time and cannot be recovered afterwards. Tokens are scoped to a single gateway. See Authentication.

Block

A guardrail or policy action that rejects a request entirely and returns an error response to the caller. Compare with Scrub (redact and continue) and Flag (log and continue). See Guardrail Pipeline.

Budget

A cumulative spend cap in USD applied at either the gateway level (config.budget_usd) or the per-token level (auth_token.budget_usd). Once the cap is reached all requests return 429 quota_exceeded. Budgets are reset via the admin API. See Budget & Quota Enforcement.

BYOK

Bring Your Own Key — the practice of supplying your own provider API keys to the gateway rather than using any keys the gateway operator holds. Keys are stored encrypted at rest and decrypted in-process at request time. The encryption key is managed by the Myra Security platform. See also: BYOK alias.

BYOK alias

A named label for a stored provider key, e.g. default or team-a. Multiple keys for the same provider can be stored under different aliases. The alias is selected per-request via the x-aig-byok-alias header. See Gateway Configuration Reference.

Cache TTL

The time-to-live in seconds for cached inference responses. Set via config.cache_ttl. When 0, caching is disabled. See also: Exact-match cache.

Compat endpoint

The unified POST /v1/{tenant}/{gateway}/compat/chat/completions endpoint that accepts any model name and automatically resolves the correct provider using prefix matching and an OpenRouter fallback. Always returns an OpenAI-shaped response. See Providers Overview.

Connector (MCP)

An external tool server, registered with the gateway through the Model Context Protocol, whose tools the model can call during chat. A tenant-scoped connector is configured by an admin in the MCP Connectors view and shares one credential. A user-scoped connector is connected by each user with a personal credential in the My MCP Connectors view. See MCP Connectors and My MCP Connectors.

Conversation file space

The set of files that belong to a conversation that is not part of a project. Files the model generates during the conversation are stored with the conversation and appear as artifact cards. A conversation attached to a project shares the file space of the project instead. See Conversation file space.

cost_usd

The estimated cost of an inference request in US dollars, calculated from token counts multiplied by the per-token pricing of the model in the pricing table of the gateway. Stored as micro-dollars internally. Appears in log entries and in the Stats API. See Models & Pricing API.

DLP

Data Loss Prevention — the practice of scanning request and response content for sensitive patterns (PII, credentials, etc.) before they leave or enter the system. In Myra AI Workspace, DLP is implemented through the guardrail pipeline using regex and keyword guardrails, or the NLP PII Detector sidecar for NER-based PII detection.

Exact-match cache

The response caching mechanism of the gateway. Responses are keyed on a SHA-256 hash of the tenant, gateway, provider, and model, plus the canonicalised request body (only the stream and user fields are excluded — metadata is part of the key) and the request-scoping headers. A cache hit returns the stored response without calling the upstream provider, saving cost and latency. Controlled by config.cache_ttl.

Fallback

An alternative provider and model that the gateway tries when the primary provider fails all retries. Fallbacks are defined in routing rule actions.fallbacks as an ordered array. Each fallback in the chain is attempted once. Only the primary provider uses retry_count. See Routing Rules API.

Flag

A detector action that records a match in the log entry (detector_flags) and continues the request without modifying the content. Compare with Block and Scrub.

Gateway

The second-level entity in the tenant hierarchy. A gateway belongs to a tenant and holds a configuration object, a set of provider keys, routing rules, and auth tokens. Each gateway corresponds to a unique path prefix in inference endpoint URLs: /v1/{tenant}/{gateway}/.... See Tenants & Gateways API.

Guardrail

A configurable content inspection component in the two-tier pipeline of the gateway. Tier 1 guardrails (regex, keyword, jailbreak, language, gibberish, JSON schema, contains-code, custom PII) run in-process in sub-milliseconds. Tier 2 guardrails (presidio, prompt_guard, prompt_injection, pii_protector) call external HTTP sidecars. Each guardrail has an action: block, scrub, or flag. See Guardrail Pipeline.

Llama Guard

An open-weight safety classification model (Meta's Llama Guard 3) that inspects prompt and response content for harm across 14 categories. In Myra AI Workspace it is used as the prompt_guard Tier-2 guardrail sidecar. See Prompt Guard.

Micro-dollars

The internal unit for cost storage: cost_usd × 1,000,000 stored as an integer. This avoids floating-point precision loss when accumulating small per-request costs in an atomic counter. The API always returns values in USD (as a float), not micro-dollars.

Myra

The Myra-hosted inference provider. Myra serves Qwen3 and Gemma models on Myra infrastructure through a vLLM backend using an OpenAI-compatible wire format. Myra uses an internal token rather than a BYOK key. Web search is supported for Myra models through the gateway's streaming tool loop. See Providers Overview.

NLP PII Detector

An NLP-based named entity recognition engine for PII (personally identifiable information) detection. Runs as a locally hosted sidecar within the Myra infrastructure — data never leaves the Myra perimeter. More accurate than the in-process regex guardrails for unstructured text but adds network round-trip latency. See NLP PII Detector.

Personal style

The writing style from the user's profile, applied to the model's replies. The personal style is toggled per conversation in the chat composer and is shown by the Personal style chip. The toggle appears only when the profile contains writing-style notes. See Easy chat — Applying a personal style.

Playground

The interactive model testing interface in the admin UI. Supports multi-model comparisons, streaming responses, web search, and session persistence. Playground sessions use short-lived tokens issued by POST /admin/v1/playground/token. See Admin API.

Playground token

A short-lived token issued for the Playground UI. Expires after 30 minutes and is managed automatically by the admin UI. Cannot be used for production inference. See Authentication.

Provider

An upstream AI inference service (e.g. OpenAI, Anthropic, Google Gemini, AWS Bedrock). The gateway supports 21 providers. Each provider has a native endpoint path and a registered adapter that translates the OpenAI-compatible request format to the native wire format of the provider. See Providers Overview.

Quota

The exhaustion state when a budget cap is reached. Requests are blocked with 429 quota_exceeded until the budget is reset. See Budget.

Rate limit

A sliding-window request frequency cap, applied at the gateway level (config.rate_limit) or per token (auth_token.rate_limit). Exceeded limits return 429 rate_limited. See Rate Limiting.

Retry

An automatic re-attempt of a failed upstream provider request on 5xx errors. The number of retries is configured via config.retry_count (default: 2). Each retry is a separate upstream call to the same provider. If all retries fail, the gateway walks the fallback chain.

Routing rule

A priority-ordered conditional rule that rewrites the provider and model for matching requests. Rules are evaluated in descending priority order (highest priority number first); the first match wins. See Routing Rules API.

saved_cost_usd

For cache-hit requests: the cost that would have been incurred if the response had been fetched from the provider rather than served from cache. Reported in log entries and aggregated in stats as a savings metric. See Stats API.

Scrub

A detector action that replaces matched content with a placeholder (e.g. [REDACTED]) and allows the (modified) request to continue. Compare with Block and Flag.

SigV4

The request signing scheme used by AWS services including Bedrock. The gateway signs Bedrock requests automatically using credentials stored via BYOK — you do not need to handle signing yourself. See AWS Bedrock.

Slug

A URL-safe lowercase string identifier (e.g. myapp, production). Slugs are unique within their scope (tenant slugs are globally unique; gateway slugs are unique within a tenant). They appear directly in inference endpoint URLs: /v1/{tenant_slug}/{gateway_slug}/....

SSE

Server-Sent Events — the streaming transport used for "stream": true inference requests. The gateway proxies SSE chunks from the provider to the client as they arrive. On the compat endpoint, provider-native SSE formats are re-encoded into OpenAI chat.completion.chunk format.

Streaming

The mode of operation when a request includes "stream": true. The provider sends tokens as they are generated; the gateway proxies each SSE chunk to the client without buffering the full response first. Usage data (token counts) is emitted in a final chunk just before data: [DONE].

Tenant

The top-level organisational unit. A tenant corresponds to one application or team. Each tenant has a unique slug that appears in all inference endpoint URLs. Tenants contain gateways, users, and tokens. See Tenants & Gateways API.

Tier 1 guardrail

An in-process guardrail (regex, keyword, jailbreak, language, gibberish, JSON schema, contains-code, custom PII) that runs entirely inside the gateway process with no external calls. Execution time is in the sub-millisecond range. Tier 1 guardrails always run before Tier 2. See Guardrail Pipeline.

Tier 2 guardrail

An HTTP-sidecar-based guardrail (presidio, prompt_guard, pii_protector, prompt_injection) that sends content to an external service for analysis. Execution time is in the millisecond range due to the network round trip. Tier 2 guardrails run only after all Tier 1 guardrails pass without blocking. See Guardrail Pipeline.

Triage queue

The unified queue of operational issues that the gateway collects from several sources — model errors, client errors, content reports, and chat-session feedback — for review and resolution. See Reports and Feedback review.

Trace

A structured step-by-step record of the execution of a single request through the gateway pipeline. Each trace captures per-phase timing and data (auth, guardrail, upstream call, log) and is stored in the playground_trace table. Gateway traces are linked to log entries via the trace_id field. See Traces API.


See also