# Rate limits (/docs/rate-limits)



Limits apply in layers. Every `429` says when to retry in `retry_after_ms`, and the request-rate
limits name themselves in `error.details.layer` (`ip`, `burst`, `recovery`).

| Layer                     | Applies to                                                       | Limit                                                               | On limit                                |
| ------------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------- | --------------------------------------- |
| `ip`                      | Anonymous routes (catalog, `/openapi.json`, `/llms.txt`), per IP | 60 per 60 s                                                         | `429 rate_limited`                      |
| `burst`                   | Each API key                                                     | 150 per 10 s                                                        | `429 rate_limited`, `Retry-After: 10`   |
| Concurrency               | Calls in flight per workspace                                    | 200 at once                                                         | `429 rate_limited`, retry after 250 ms  |
| `recovery`                | Recovery-code redemption, per IP                                 | 30 per 60 s, and a workspace lockout after 5 failed codes in 15 min | `429 rate_limited`                      |
| Credit limits and budgets | Keys, agents, apps and runs with a spend cap                     | as configured                                                       | `403 budget_blocked`                    |
| Vendor capacity           | Each vendor and credential, from the vendor's published limits   | per vendor                                                          | `503 provider_capacity_unavailable`     |
| Platform                  | Everyone, under extreme load                                     | —                                                                   | `503 grid_saturated` with `Retry-After` |

Dashboard sign-in, sign-up and password reset have their own per-IP limits.

## Headers [#headers]

Responses carry the IETF `RateLimit-Policy` header for the limiter that applied:

```http
RateLimit-Policy: "burst";q=150;w=10
```

A `429` also carries `RateLimit` with nothing remaining and `Retry-After`:

```http
HTTP/1.1 429 Too Many Requests
RateLimit: "burst";r=0;t=10
Retry-After: 10
```

These headers are exposed to browsers through CORS.

## Vendor limits [#vendor-limits]

Each vendor has a token bucket, a concurrency limit and a circuit breaker per credential, set from the
vendor's published limits and your plan with them. When a vendor answers `429`, GridRouter honors
its `Retry-After`. A routed `/v1/run` moves on to the next vendor instead of waiting, and returns
`503 provider_capacity_unavailable` only when no vendor has capacity. Adding a
[fallback vendor key](/docs/getting-started/vendor-keys) gives the governor a second bucket.

## Handling a 429 [#handling-a-429]

1. Wait `retry_after_ms` (or `Retry-After` seconds), with jitter.
2. For bulk work, use [lists](/docs/api/async/list_create) or
   [batch waterfall runs](/docs/api/waterfalls/waterfall_batch_run), which pace rows per vendor
   instead of failing them.
3. Spread high-volume traffic across several keys only if they belong to separate services;
   the workspace concurrency limit still applies.

