Rate Limits

How requests are throttled and how to handle the limit

The API enforces rate limits to maintain stable operations and ensure fair usage across tenants. Limits are applied per tenant, one tenant's traffic cannot exhaust another's budget.

Limits

The limit is a token bucket. The bucket holds 100 requests and refills at 50 requests per second. You can send short bursts of up to 100 requests, but over time you can't exceed 50 requests per second.

📘

Need a higher limit?

Reach out to [email protected] with details about your expected traffic.

Response headers

Every response includes the current rate-limit state:

HeaderDescription
ratelimit-limitThe sustained per-second limit applied to your tenant.
ratelimit-remainingRequests left at the sustained rate. Doesn't include the burst.
retry-afterOn 429 responses only. Seconds to wait before retrying.

ratelimit-remaining: 0 doesn't mean your next request will be rejected. It means you are using your burst and should slow down.

When the limit is exceeded

When both the sustained rate and the burst are used up, the API responds with 429 Too Many Requests and a Retry-After header that gives the number of seconds to wait.

HTTP/1.1 429 Too Many Requests
ratelimit-limit: 50
ratelimit-remaining: 0
retry-after: 1

Did this page help you?