Alerts
Rules over your calls, budgets, vendor keys, billing checks and deliveries. Alerts open, repeat, get acknowledged and resolve, and notify the destinations you choose.
An alert rule watches one metric in your workspace: the error rate of your calls, what they cost, how much of a budget is used, a vendor rejecting your own key, and more. When the metric crosses the rule's threshold, GridRouter opens an alert, shows it as a toast in the dashboard and notifies the rule's destinations: email, your own webhook endpoints, Slack, Discord, Microsoft Teams, PagerDuty or Opsgenie. When the metric recovers, the alert resolves on its own and destinations that take resolve notifications are told.
Manage rules, alerts and destinations in the dashboard under Settings → Alerts & integrations, through the API or with the MCP tools below.
Metrics
Every rule measures one of these. The table is generated from the same definitions the API validates against.
| Metric | Measures | Threshold | Fires | Evaluated | Scopes | One alert per | Default |
|---|---|---|---|---|---|---|---|
error_rate | Error rate: The share of calls that failed in the window. | Share, 0–1 (0.5 is 50%) | At or above | Live, within seconds (windows up to an hour) | workspace, key, vendor, capability | Rule | 50% over 5 minutes, at least 20 events |
p95_latency | p95 latency: How long the slowest 5% of calls took, end to end. | Milliseconds | At or above | Every 5 minutes | workspace, key, vendor, capability | Rule | 5 s over 15 minutes, at least 20 events |
spend | Spend: What calls cost in the window, in USD. | Micro-USD (1000000 is $1) | At or above | Live, within seconds (windows up to an hour) | workspace, key, vendor, capability | Rule | $5.00 over 5 minutes |
budget_used | Budget used: How much of a workspace or key budget this period has used. | Percent of the limit (80) | At or above | Every 5 minutes | workspace, key | Budget | 80% over 5 minutes |
vendor_failure_spike | Vendor failure spike: One vendor's failed share, alerted separately for each vendor. | Share, 0–1 (0.5 is 50%) | At or above | Live, within seconds (windows up to an hour) | workspace, vendor, capability | Vendor | 50% over 5 minutes, at least 10 events |
vendor_key_failing | Vendor key failing: A vendor rejected one of your own vendor keys (HTTP 401 or 403): expired or revoked. | Count | At or above | Live, within seconds (windows up to an hour) | workspace, vendor | Vendor | 1 over 5 minutes |
billing_mismatch | Billing mismatches: Calls a vendor billed against its own published rule (charged for a miss, twice, …). | Count | At or above | Live, within seconds (windows up to an hour) | workspace, vendor, key | Rule | 3 over an hour |
cache_hit_rate | Cache hit rate: The share of cacheable lookups answered from your private cache. | Share, 0–1 (0.5 is 50%) | At or below | Every 5 minutes | workspace, key, vendor, capability | Rule | 20% over an hour, at least 50 events |
delivery_backlog | Delivery backlog: Webhook deliveries and log-drain batches that failed or were dead-lettered. | Count | At or above | Every 5 minutes | workspace | Rule | 10 over an hour |
security_event | Security events: Invalid-key bursts, a key used from a new country, admin actions on the workspace. | Count | At or above | When the signal arrives | workspace, key | Signal kind | 1 over an hour |
new_ip | Key used from a new IP: A key was used from an IP address it has never used before. | Count | At or above | When the signal arrives | workspace, key | Key | 1 over an hour |
threshold is in the metric's unit. Money is always integer micro-USD, so 5000000 is $5. Rates
(error_rate, vendor_failure_spike, p95_latency, cache_hit_rate) wait for min_events calls
or lookups in the window before they can fire, so a single failed call in a quiet minute doesn't
page anyone.
Default rules
Every new workspace starts with these rules. They're ordinary rules: edit, disable or delete any of them.
| Rule | Severity | Fires when | Evaluated | Cooldown | Notifies |
|---|---|---|---|---|---|
| Vendor key failing | critical | Alerts when a vendor rejects your own vendor key (HTTP 401 or 403) within 5 minutes: usually an expired or revoked key. | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Budget at 80% | warning | Alerts when a workspace or key budget reaches 80% of its limit for the period. | Every 5 minutes | a day | Every destination, and a toast |
| Budget reached | critical | Alerts when a workspace or key budget reaches 100% of its limit for the period. | Every 5 minutes | a day | Every destination, and a toast |
| Billing mismatch spike | warning | Alerts when 3 or more calls are billed against the vendor's published rule within an hour. | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Error-rate spike | warning | Alerts when 50% or more of calls fail over 5 minutes (once there are at least 20 calls). | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Spend velocity | warning | Alerts when calls cost $5.00 or more within 5 minutes. | Live, within seconds (windows up to an hour) | 15 minutes | Every destination, and a toast |
| Key used from a new IP | info | Alerts when a key is used from an IP address it hasn't used before. | When the signal arrives | 15 minutes | Dashboard toast only |
| Deliveries failing | warning | Alerts when 10 or more webhook deliveries or log-drain batches fail within an hour. | Every 5 minutes | 15 minutes | Every destination, and a toast |
| Security events | warning | Alerts on every security event. | When the signal arrives | 15 minutes | Every destination, and a toast |
"Key used from a new IP" is informational: it shows as a toast but goes to no destination until you route it to one. A key's first IP is its baseline and never alerts.
Rule settings
| Field | Meaning |
|---|---|
metric, threshold | What to measure, and the value that fires the rule (at or above, or at or below for cache_hit_rate) |
window_s | The window the metric is measured over, 60 seconds to 7 days |
min_events | Rates only: calls or lookups the window needs before the rule can fire (default 1) |
severity | info, warning (default) or critical |
scope | { "type": "workspace" } (default), or key, vendor or capability with a value (a key id, vendor slug or capability id). Each metric lists the scopes it takes |
kinds | security_event only: the signal kinds that count (invalid_key_burst, new_geo, admin_action, key_ip_denied, scope_denied_burst); empty means every kind |
cooldown_s | After an alert resolves, the same rule and subject can't open a new one for this long (default 900, at most 86,400) |
all_destinations, destination_ids | Notify every enabled destination (default), or only the listed ones (up to 20) |
toast | Also show the alert as a dashboard toast (default on) |
muted_until | Snooze until this time; see Mute |
Each destination has its own min_severity (alerts below it are not sent there) and
notify_on_resolve (default on).
Lifecycle
An alert is triggered, then optionally acknowledged, then resolved. Its timeline records
every step: when it triggered, each notification sent or failed, acknowledgement and resolution
(the newest 60 entries, always keeping the trigger).
- Dedupe. While an alert is open, further breaches of the same rule and subject don't open
another one. They raise the alert's
countand update its value andlast_seen_at. - Grouping. Some metrics open one alert per subject instead of one per rule:
vendor_failure_spikeandvendor_key_failingper vendor,budget_usedper budget (the workspace's or one key's),new_ipper key andsecurity_eventper signal kind. The subject is on the alert assubject. - Cooldown. After an alert resolves, the same rule and subject can't open a new alert for
cooldown_s. The budget rules default to a day, so a budget that hovers around 80% alerts once a day at most. - Auto-resolve. A live rule reports that it has cleared once it has stayed on the right side of its threshold for a minute. A scheduled rule resolves on the next five-minute run that's no longer breached. An alert that nothing has confirmed for a while resolves too: 5 to 30 minutes for live rules (the window, within that range), 15 minutes for scheduled rules, and the rule's window for event rules.
- Acknowledge.
POST /v1/alerts/{id}/acknowledgemarks a triggered alert as being handled. PagerDuty and Opsgenie are acknowledged too; chat and email destinations only hear about triggers and resolutions. - Resolve.
POST /v1/alerts/{id}/resolvecloses an alert now, and the rule's cooldown starts from there. Changing what a rule measures (metric, threshold, window,min_events, scope or kinds), disabling it or deleting it resolves its open alerts.
Mute
POST /v1/alerts/rules/{id}/mute with { "minutes": 60 } (or { "until": "<ISO time>" }) snoozes
a rule (minutes goes up to 30 days). It keeps evaluating: its alerts still open and resolve, show in history
marked muted and still show as toasts, but nothing is sent to destinations or webhook
endpoints. { "minutes": 0 } unmutes it.
Live and scheduled evaluation
Where a rule runs depends on its metric and window:
- Live.
error_rate,spend,vendor_failure_spike,vendor_key_failingandbilling_mismatchwith windows up to an hour run in your workspace's live tail hub on every call, so they fire within seconds. - Scheduled.
p95_latency,cache_hit_rate,budget_used,delivery_backlogand any live metric with a window over an hour are evaluated every five minutes, against your call log (hourly rollups for whole-hour windows), your budgets and the webhook delivery log. - Event.
new_ipandsecurity_eventfire when the signal arrives.
Every rule shows its state: ok, firing, muted, disabled or no_data. no_data means the
rule couldn't be evaluated, with the reason in state.note; for example, in a deployment
without the call-log store, scheduled call metrics report no_data while every other rule keeps
working.
Toasts
Alerts with toast on appear as toasts in the dashboard when they trigger and when they resolve.
They arrive over the live tail as alert events, so any client that tails kinds=alert sees them
too:
curl -N "https://api.gridrouter.io/v1/logs/live?format=sse&kinds=alert&replay=0" \
-H "Authorization: Bearer $GRID_API_KEY"Create a rule
Creating and changing rules needs a key with keys:write. This rule pages on-call when a third of
email-finding calls fail over ten minutes, and only notifies one destination:
curl https://api.gridrouter.io/v1/alerts/rules \
-H "Authorization: Bearer $GRID_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Email finding failing",
"metric": "vendor_failure_spike",
"threshold": 0.33,
"window_s": 600,
"min_events": 25,
"severity": "critical",
"scope": { "type": "capability", "value": "people.email.find" },
"all_destinations": false,
"destination_ids": ["dst_4f8k2m9q4w8x1c5v"]
}'The response is the rule with its id, its evaluator and its state.
POST /v1/alerts/rules/{id}/test sends a test alert to every destination the rule notifies and
reports each result; nothing is recorded as a real alert.
API and MCP
Every operation is a REST route and an MCP tool of the same name. Reads need
logs:read, rule and alert changes need keys:write, and destinations need credentials:write
because they hold secrets.
Reads, with logs:read:
alerts_summary:GET /v1/alerts/summaryalerts_rules_listandalerts_rules_get:GET /v1/alerts/rulesandGET /v1/alerts/rules/{id}alerts_listandalerts_get:GET /v1/alerts?status=openandGET /v1/alerts/{id}alerts_destinations_listandalerts_destinations_get:GET /v1/alerts/destinationsandGET /v1/alerts/destinations/{id}
Rule and alert changes, with keys:write:
alerts_rules_create:POST /v1/alerts/rulesalerts_rules_updateandalerts_rules_delete:PUTandDELETE /v1/alerts/rules/{id}alerts_rules_muteandalerts_rules_test:POST /v1/alerts/rules/{id}/muteandPOST /v1/alerts/rules/{id}/testalerts_acknowledgeandalerts_resolve:POST /v1/alerts/{id}/acknowledgeandPOST /v1/alerts/{id}/resolve
Destination changes, with credentials:write:
alerts_destinations_create:POST /v1/alerts/destinationsalerts_destinations_updateandalerts_destinations_delete:PUTandDELETE /v1/alerts/destinations/{id}alerts_destinations_test:POST /v1/alerts/destinations/{id}/test
alerts_list takes status (open, triggered, acknowledged, resolved or all),
rule_id and limit (up to 200). A PUT changes only the fields you send.
Limits
| Limit | Value |
|---|---|
| Alert rules per workspace | 50 |
| Destinations per workspace | 20 |
| Resolved alerts kept | The newest 2,000 |
| Timeline entries per alert | 60 |
Logs, live tail and drains
Every call is logged and streamed as it happens. Search it, tail it from the CLI, SDK or MCP, and drain it to your own stack.
Webhooks
Signed HTTPS events for calls, jobs, lists, waterfall runs, alerts, vendor keys, billing checks, budgets, keys and members. Standard Webhooks signatures, retries for about a day, a delivery log and replay.