# Alerts (/docs/concepts/alerts)



An **alert rule** watches one metric in your workspace: the error rate of your calls, what they
cost, how much of a budget is used, a vendor rejecting your own key, and more. When the metric
crosses the rule's threshold, GridRouter opens an **alert**, shows it as a toast in the dashboard
and notifies the rule's [destinations](/docs/concepts/integrations): email, your own
[webhook endpoints](/docs/concepts/webhooks), Slack, Discord, Microsoft Teams, PagerDuty or
Opsgenie. When the metric recovers, the alert resolves on its own and destinations that take
resolve notifications are told.

Manage rules, alerts and destinations in the dashboard under Settings → **Alerts & integrations**,
through the API or with the MCP tools below.

## Metrics [#metrics]

Every rule measures one of these. The table is generated from the same definitions the API
validates against.

| Metric | Measures | Threshold | Fires | Evaluated | Scopes | One alert per | Default |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `error_rate` | Error rate: The share of calls that failed in the window. | Share, 0–1 (`0.5` is 50%) | At or above | Live, within seconds (windows up to an hour) | `workspace`, `key`, `vendor`, `capability` | Rule | 50% over 5 minutes, at least 20 events |
| `p95_latency` | p95 latency: How long the slowest 5% of calls took, end to end. | Milliseconds | At or above | Every 5 minutes | `workspace`, `key`, `vendor`, `capability` | Rule | 5 s over 15 minutes, at least 20 events |
| `spend` | Spend: What calls cost in the window, in USD. | Micro-USD (`1000000` is $1) | At or above | Live, within seconds (windows up to an hour) | `workspace`, `key`, `vendor`, `capability` | Rule | $5.00 over 5 minutes |
| `budget_used` | Budget used: How much of a workspace or key budget this period has used. | Percent of the limit (`80`) | At or above | Every 5 minutes | `workspace`, `key` | Budget | 80% over 5 minutes |
| `vendor_failure_spike` | Vendor failure spike: One vendor's failed share, alerted separately for each vendor. | Share, 0–1 (`0.5` is 50%) | At or above | Live, within seconds (windows up to an hour) | `workspace`, `vendor`, `capability` | Vendor | 50% over 5 minutes, at least 10 events |
| `vendor_key_failing` | Vendor key failing: A vendor rejected one of your own vendor keys (HTTP 401 or 403): expired or revoked. | Count | At or above | Live, within seconds (windows up to an hour) | `workspace`, `vendor` | Vendor | 1 over 5 minutes |
| `billing_mismatch` | Billing mismatches: Calls a vendor billed against its own published rule (charged for a miss, twice, …). | Count | At or above | Live, within seconds (windows up to an hour) | `workspace`, `vendor`, `key` | Rule | 3 over an hour |
| `cache_hit_rate` | Cache hit rate: The share of cacheable lookups answered from your private cache. | Share, 0–1 (`0.5` is 50%) | At or below | Every 5 minutes | `workspace`, `key`, `vendor`, `capability` | Rule | 20% over an hour, at least 50 events |
| `delivery_backlog` | Delivery backlog: Webhook deliveries and log-drain batches that failed or were dead-lettered. | Count | At or above | Every 5 minutes | `workspace` | Rule | 10 over an hour |
| `security_event` | Security events: Invalid-key bursts, a key used from a new country, admin actions on the workspace. | Count | At or above | When the signal arrives | `workspace`, `key` | Signal kind | 1 over an hour |
| `new_ip` | Key used from a new IP: A key was used from an IP address it has never used before. | Count | At or above | When the signal arrives | `workspace`, `key` | Key | 1 over an hour |

`threshold` is in the metric's unit. Money is always integer micro-USD, so `5000000` is $5. Rates
(`error_rate`, `vendor_failure_spike`, `p95_latency`, `cache_hit_rate`) wait for `min_events` calls
or lookups in the window before they can fire, so a single failed call in a quiet minute doesn't
page anyone.

## Default rules [#default-rules]

Every new workspace starts with these rules. They're ordinary rules: edit, disable or delete any
of them.

| Rule | Severity | Fires when | Evaluated | Cooldown | Notifies |
| --- | --- | --- | --- | --- | --- |
| Vendor key failing | critical | Alerts when a vendor rejects your own vendor key (HTTP 401 or 403) within 5 minutes: usually an expired or revoked key. | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Budget at 80% | warning | Alerts when a workspace or key budget reaches 80% of its limit for the period. | Every 5 minutes | a day | Every destination, and a toast |
| Budget reached | critical | Alerts when a workspace or key budget reaches 100% of its limit for the period. | Every 5 minutes | a day | Every destination, and a toast |
| Billing mismatch spike | warning | Alerts when 3 or more calls are billed against the vendor's published rule within an hour. | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Error-rate spike | warning | Alerts when 50% or more of calls fail over 5 minutes (once there are at least 20 calls). | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Spend velocity | warning | Alerts when calls cost $5.00 or more within 5 minutes. | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Key used from a new IP | info | Alerts when a key is used from an IP address it hasn't used before. | When the signal arrives | 15 minutes | Dashboard toast only |
| Deliveries failing | warning | Alerts when 10 or more webhook deliveries or log-drain batches fail within an hour. | Every 5 minutes | 15 minutes | Every destination, and a toast |
| Security events | warning | Alerts on every security event. | When the signal arrives | 15 minutes | Every destination, and a toast |

"Key used from a new IP" is informational: it shows as a toast but goes to no destination until
you route it to one. A key's first IP is its baseline and never alerts.

## Rule settings [#rule-settings]

| Field                                 | Meaning                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `metric`, `threshold`                 | What to measure, and the value that fires the rule (at or above, or at or below for `cache_hit_rate`)                                                                  |
| `window_s`                            | The window the metric is measured over, 60 seconds to 7 days                                                                                                           |
| `min_events`                          | Rates only: calls or lookups the window needs before the rule can fire (default 1)                                                                                     |
| `severity`                            | `info`, `warning` (default) or `critical`                                                                                                                              |
| `scope`                               | `{ "type": "workspace" }` (default), or `key`, `vendor` or `capability` with a `value` (a key id, vendor slug or capability id). Each metric lists the scopes it takes |
| `kinds`                               | `security_event` only: the signal kinds that count (`invalid_key_burst`, `new_geo`, `admin_action`, `key_ip_denied`, `scope_denied_burst`); empty means every kind     |
| `cooldown_s`                          | After an alert resolves, the same rule and subject can't open a new one for this long (default 900, at most 86,400)                                                    |
| `all_destinations`, `destination_ids` | Notify every enabled destination (default), or only the listed ones (up to 20)                                                                                         |
| `toast`                               | Also show the alert as a dashboard toast (default on)                                                                                                                  |
| `muted_until`                         | Snooze until this time; see [Mute](#mute)                                                                                                                              |

Each destination has its own `min_severity` (alerts below it are not sent there) and
`notify_on_resolve` (default on).

## Lifecycle [#lifecycle]

An alert is `triggered`, then optionally `acknowledged`, then `resolved`. Its timeline records
every step: when it triggered, each notification sent or failed, acknowledgement and resolution
(the newest 60 entries, always keeping the trigger).

* **Dedupe.** While an alert is open, further breaches of the same rule and subject don't open
  another one. They raise the alert's `count` and update its value and `last_seen_at`.
* **Grouping.** Some metrics open one alert per subject instead of one per rule:
  `vendor_failure_spike` and `vendor_key_failing` per vendor, `budget_used` per budget (the
  workspace's or one key's), `new_ip` per key and `security_event` per signal kind. The subject is
  on the alert as `subject`.
* **Cooldown.** After an alert resolves, the same rule and subject can't open a new alert for
  `cooldown_s`. The budget rules default to a day, so a budget that hovers around 80% alerts once
  a day at most.
* **Auto-resolve.** A live rule reports that it has cleared once it has stayed on the right side of
  its threshold for a minute. A scheduled rule resolves on the next five-minute run that's no longer
  breached. An alert that nothing has confirmed for a while resolves too: 5 to 30 minutes for live
  rules (the window, within that range), 15 minutes for scheduled rules, and the rule's window for
  event rules.
* **Acknowledge.** `POST /v1/alerts/{id}/acknowledge` marks a triggered alert as being handled.
  PagerDuty and Opsgenie are acknowledged too; chat and email destinations only hear about
  triggers and resolutions.
* **Resolve.** `POST /v1/alerts/{id}/resolve` closes an alert now, and the rule's cooldown starts
  from there. Changing what a rule measures (metric, threshold, window, `min_events`, scope or
  kinds), disabling it or deleting it resolves its open alerts.

### Mute [#mute]

`POST /v1/alerts/rules/{id}/mute` with `{ "minutes": 60 }` (or `{ "until": "<ISO time>" }`) snoozes
a rule (`minutes` goes up to 30 days). It keeps evaluating: its alerts still open and resolve, show in history
marked `muted` and still show as toasts, but nothing is sent to destinations or webhook
endpoints. `{ "minutes": 0 }` unmutes it.

## Live and scheduled evaluation [#live-and-scheduled-evaluation]

Where a rule runs depends on its metric and window:

* **Live.** `error_rate`, `spend`, `vendor_failure_spike`, `vendor_key_failing` and
  `billing_mismatch` with windows up to an hour run in your workspace's
  [live tail](/docs/concepts/logs-and-drains#live-tail) hub on every call, so they fire within
  seconds.
* **Scheduled.** `p95_latency`, `cache_hit_rate`, `budget_used`, `delivery_backlog` and any live
  metric with a window over an hour are evaluated every five minutes, against your call log (hourly
  rollups for whole-hour windows), your budgets and the webhook delivery log.
* **Event.** `new_ip` and `security_event` fire when the signal arrives.

Every rule shows its state: `ok`, `firing`, `muted`, `disabled` or `no_data`. `no_data` means the
rule couldn't be evaluated, with the reason in `state.note`; for example, in a deployment
without the call-log store, scheduled call metrics report `no_data` while every other rule keeps
working.

## Toasts [#toasts]

Alerts with `toast` on appear as toasts in the dashboard when they trigger and when they resolve.
They arrive over the live tail as `alert` events, so any client that tails `kinds=alert` sees them
too:

```bash
curl -N "https://api.gridrouter.io/v1/logs/live?format=sse&kinds=alert&replay=0" \
  -H "Authorization: Bearer $GRID_API_KEY"
```

## Create a rule [#create-a-rule]

Creating and changing rules needs a key with `keys:write`. This rule pages on-call when a third of
email-finding calls fail over ten minutes, and only notifies one destination:

```bash
curl https://api.gridrouter.io/v1/alerts/rules \
  -H "Authorization: Bearer $GRID_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Email finding failing",
    "metric": "vendor_failure_spike",
    "threshold": 0.33,
    "window_s": 600,
    "min_events": 25,
    "severity": "critical",
    "scope": { "type": "capability", "value": "people.email.find" },
    "all_destinations": false,
    "destination_ids": ["dst_4f8k2m9q4w8x1c5v"]
  }'
```

The response is the rule with its `id`, its `evaluator` and its `state`.
`POST /v1/alerts/rules/{id}/test` sends a test alert to every destination the rule notifies and
reports each result; nothing is recorded as a real alert.

## API and MCP [#api-and-mcp]

Every operation is a REST route and an [MCP tool](/docs/mcp) of the same name. Reads need
`logs:read`, rule and alert changes need `keys:write`, and destinations need `credentials:write`
because they hold secrets.

Reads, with `logs:read`:

* `alerts_summary`: `GET /v1/alerts/summary`
* `alerts_rules_list` and `alerts_rules_get`: `GET /v1/alerts/rules` and
  `GET /v1/alerts/rules/{id}`
* `alerts_list` and `alerts_get`: `GET /v1/alerts?status=open` and `GET /v1/alerts/{id}`
* `alerts_destinations_list` and `alerts_destinations_get`: `GET /v1/alerts/destinations` and
  `GET /v1/alerts/destinations/{id}`

Rule and alert changes, with `keys:write`:

* `alerts_rules_create`: `POST /v1/alerts/rules`
* `alerts_rules_update` and `alerts_rules_delete`: `PUT` and `DELETE /v1/alerts/rules/{id}`
* `alerts_rules_mute` and `alerts_rules_test`: `POST /v1/alerts/rules/{id}/mute` and
  `POST /v1/alerts/rules/{id}/test`
* `alerts_acknowledge` and `alerts_resolve`: `POST /v1/alerts/{id}/acknowledge` and
  `POST /v1/alerts/{id}/resolve`

Destination changes, with `credentials:write`:

* `alerts_destinations_create`: `POST /v1/alerts/destinations`
* `alerts_destinations_update` and `alerts_destinations_delete`: `PUT` and
  `DELETE /v1/alerts/destinations/{id}`
* `alerts_destinations_test`: `POST /v1/alerts/destinations/{id}/test`

`alerts_list` takes `status` (`open`, `triggered`, `acknowledged`, `resolved` or `all`),
`rule_id` and `limit` (up to 200). A `PUT` changes only the fields you send.

## Limits [#limits]

| Limit                      | Value            |
| -------------------------- | ---------------- |
| Alert rules per workspace  | 50               |
| Destinations per workspace | 20               |
| Resolved alerts kept       | The newest 2,000 |
| Timeline entries per alert | 60               |

