Rate limiting
Rate limiting caps the request rate at the gateway. The gateway uses a sliding-window dual-bucket approximation: instead of counting a literal trailing window, it interpolates between the previous and current fixed windows. A request is admitted only when the resulting effective count is below the configured limit.
Levels
Rate limits are configured at two independent levels:
- Gateway-level — caps the aggregate request rate across every token of the gateway.
- Token-level — caps the request rate of a single authentication token.
The most restrictive applicable limit wins. A request is admitted only when every applicable level has capacity.
Configuration
A rate limit is configured as an object with two fields:
requests— the maximum number of requests admitted per window. When arate_limitobject is present but omits this field, it falls back to100.window_sec— the trailing window length in seconds. When present but omitted, it falls back to60.
A gateway has no rate limit until a rate_limit object is set: an absent object means unlimited, not 100/60. The 100/60 fallbacks apply only to a missing field within a rate_limit object that is otherwise present.
Window semantics
The sliding window is approximated with two fixed buckets — the previous window and the current window. The effective count weights the previous bucket by the fraction of the current window not yet elapsed and adds the current bucket's count; new requests are rejected when that effective count reaches the configured limit. This smooths bursts at window boundaries without storing individual request timestamps.
Every response — both admitted requests and rejections — carries the current rate-limit state in its headers:
X-RateLimit-Limit— the configured request ceiling for the applicable window.X-RateLimit-Remaining— the number of requests still available in the current window (0on a rejection).
A rejection is returned as HTTP 429 and additionally carries a Retry-After header set to the window length in seconds.
💡 Note: Each
429block is also counted cross-worker. Every rejection bumps a lock-free per-worker tally that worker 0 flushes on a timer (default 60 s, configured by your operator) to a persisted per-local-calendar-day aggregate. The gateway or auth token being throttled then surfaces as arate_limit-scope alert — carrying that day's throttle count — in the notifications center (not the dismissible dashboard banner). This block-event tally is distinct from the sliding-window admission counters above, which are ephemeral and bump only on admitted requests.
Difference from budgets
A rate limit caps the request rate (requests per minute or per hour). A budget caps spend (currency units per period). Both are enforced independently; a request must satisfy both to be admitted.
See also
- Rate limiting — how to configure gateway and token limits