Rate limits
The limits that apply to your requests, the headers that describe them, and how to handle a 429.
Limits apply in layers. Every 429 says when to retry in retry_after_ms, and the request-rate
limits name themselves in error.details.layer (ip, burst, recovery).
| Layer | Applies to | Limit | On limit |
|---|---|---|---|
ip | Anonymous routes (catalog, /openapi.json, /llms.txt), per IP | 60 per 60 s | 429 rate_limited |
burst | Each API key | 150 per 10 s | 429 rate_limited, Retry-After: 10 |
| Concurrency | Calls in flight per workspace | 200 at once | 429 rate_limited, retry after 250 ms |
recovery | Recovery-code redemption, per IP | 30 per 60 s, and a workspace lockout after 5 failed codes in 15 min | 429 rate_limited |
| Credit limits and budgets | Keys, agents, apps and runs with a spend cap | as configured | 403 budget_blocked |
| Vendor capacity | Each vendor and credential, from the vendor's published limits | per vendor | 503 provider_capacity_unavailable |
| Platform | Everyone, under extreme load | — | 503 grid_saturated with Retry-After |
Dashboard sign-in, sign-up and password reset have their own per-IP limits.
Headers
Responses carry the IETF RateLimit-Policy header for the limiter that applied:
RateLimit-Policy: "burst";q=150;w=10A 429 also carries RateLimit with nothing remaining and Retry-After:
HTTP/1.1 429 Too Many Requests
RateLimit: "burst";r=0;t=10
Retry-After: 10These headers are exposed to browsers through CORS.
Vendor limits
Each vendor has a token bucket, a concurrency limit and a circuit breaker per credential, set from the
vendor's published limits and your plan with them. When a vendor answers 429, GridRouter honors
its Retry-After. A routed /v1/run moves on to the next vendor instead of waiting, and returns
503 provider_capacity_unavailable only when no vendor has capacity. Adding a
fallback vendor key gives the governor a second bucket.
Handling a 429
- Wait
retry_after_ms(orRetry-Afterseconds), with jitter. - For bulk work, use lists or batch waterfall runs, which pace rows per vendor instead of failing them.
- Spread high-volume traffic across several keys only if they belong to separate services; the workspace concurrency limit still applies.
Next