Skip to content

Rate limiting

Rate limiting caps the request rate at the gateway. The gateway uses a sliding-window dual-bucket approximation: instead of counting a literal trailing window, it interpolates between the previous and current fixed windows. A request is admitted only when the resulting effective count is below the configured limit.

Levels

Rate limits are configured at two independent levels:

  • Gateway-level — caps the aggregate request rate across every token of the gateway.
  • Token-level — caps the request rate of a single authentication token.

The most restrictive applicable limit wins. A request is admitted only when every applicable level has capacity.

Configuration

A rate limit is configured as an object with two fields:

{
  "rate_limit": {
    "requests": 100,
    "window_sec": 60
  }
}
  • requests — the maximum number of requests admitted per window. When a rate_limit object is present but omits this field, it falls back to 100.
  • window_sec — the trailing window length in seconds. When present but omitted, it falls back to 60.

A gateway has no rate limit until a rate_limit object is set: an absent object means unlimited, not 100/60. The 100/60 fallbacks apply only to a missing field within a rate_limit object that is otherwise present.

Window semantics

The sliding window is approximated with two fixed buckets — the previous window and the current window. The effective count weights the previous bucket by the fraction of the current window not yet elapsed and adds the current bucket's count; new requests are rejected when that effective count reaches the configured limit. This smooths bursts at window boundaries without storing individual request timestamps.

Every response — both admitted requests and rejections — carries the current rate-limit state in its headers:

  • X-RateLimit-Limit — the configured request ceiling for the applicable window.
  • X-RateLimit-Remaining — the number of requests still available in the current window (0 on a rejection).

A rejection is returned as HTTP 429 and additionally carries a Retry-After header set to the window length in seconds.

💡 Note: Each 429 block is also counted cross-worker. Every rejection bumps a lock-free per-worker tally that worker 0 flushes on a timer (default 60 s, configured by your operator) to a persisted per-local-calendar-day aggregate. The gateway or auth token being throttled then surfaces as a rate_limit-scope alert — carrying that day's throttle count — in the notifications center (not the dismissible dashboard banner). This block-event tally is distinct from the sliding-window admission counters above, which are ephemeral and bump only on admitted requests.

Difference from budgets

A rate limit caps the request rate (requests per minute or per hour). A budget caps spend (currency units per period). Both are enforced independently; a request must satisfy both to be admitted.

See also