Skip to content
GridRouterhome

Search

Search providers, capabilities and pages

Docs
Concepts

Alerts

Rules over your calls, budgets, vendor keys, billing checks and deliveries. Alerts open, repeat, get acknowledged and resolve, and notify the destinations you choose.

An alert rule watches one metric in your workspace: the error rate of your calls, what they cost, how much of a budget is used, a vendor rejecting your own key, and more. When the metric crosses the rule's threshold, GridRouter opens an alert, shows it as a toast in the dashboard and notifies the rule's destinations: email, your own webhook endpoints, Slack, Discord, Microsoft Teams, PagerDuty or Opsgenie. When the metric recovers, the alert resolves on its own and destinations that take resolve notifications are told.

Manage rules, alerts and destinations in the dashboard under Settings → Alerts & integrations, through the API or with the MCP tools below.

Metrics

Every rule measures one of these. The table is generated from the same definitions the API validates against.

MetricMeasuresThresholdFiresEvaluatedScopesOne alert perDefault
error_rateError rate: The share of calls that failed in the window.Share, 0–1 (0.5 is 50%)At or aboveLive, within seconds (windows up to an hour)workspace, key, vendor, capabilityRule50% over 5 minutes, at least 20 events
p95_latencyp95 latency: How long the slowest 5% of calls took, end to end.MillisecondsAt or aboveEvery 5 minutesworkspace, key, vendor, capabilityRule5 s over 15 minutes, at least 20 events
spendSpend: What calls cost in the window, in USD.Micro-USD (1000000 is $1)At or aboveLive, within seconds (windows up to an hour)workspace, key, vendor, capabilityRule$5.00 over 5 minutes
budget_usedBudget used: How much of a workspace or key budget this period has used.Percent of the limit (80)At or aboveEvery 5 minutesworkspace, keyBudget80% over 5 minutes
vendor_failure_spikeVendor failure spike: One vendor's failed share, alerted separately for each vendor.Share, 0–1 (0.5 is 50%)At or aboveLive, within seconds (windows up to an hour)workspace, vendor, capabilityVendor50% over 5 minutes, at least 10 events
vendor_key_failingVendor key failing: A vendor rejected one of your own vendor keys (HTTP 401 or 403): expired or revoked.CountAt or aboveLive, within seconds (windows up to an hour)workspace, vendorVendor1 over 5 minutes
billing_mismatchBilling mismatches: Calls a vendor billed against its own published rule (charged for a miss, twice, …).CountAt or aboveLive, within seconds (windows up to an hour)workspace, vendor, keyRule3 over an hour
cache_hit_rateCache hit rate: The share of cacheable lookups answered from your private cache.Share, 0–1 (0.5 is 50%)At or belowEvery 5 minutesworkspace, key, vendor, capabilityRule20% over an hour, at least 50 events
delivery_backlogDelivery backlog: Webhook deliveries and log-drain batches that failed or were dead-lettered.CountAt or aboveEvery 5 minutesworkspaceRule10 over an hour
security_eventSecurity events: Invalid-key bursts, a key used from a new country, admin actions on the workspace.CountAt or aboveWhen the signal arrivesworkspace, keySignal kind1 over an hour
new_ipKey used from a new IP: A key was used from an IP address it has never used before.CountAt or aboveWhen the signal arrivesworkspace, keyKey1 over an hour

threshold is in the metric's unit. Money is always integer micro-USD, so 5000000 is $5. Rates (error_rate, vendor_failure_spike, p95_latency, cache_hit_rate) wait for min_events calls or lookups in the window before they can fire, so a single failed call in a quiet minute doesn't page anyone.

Default rules

Every new workspace starts with these rules. They're ordinary rules: edit, disable or delete any of them.

RuleSeverityFires whenEvaluatedCooldownNotifies
Vendor key failingcriticalAlerts when a vendor rejects your own vendor key (HTTP 401 or 403) within 5 minutes: usually an expired or revoked key.Live, within seconds (windows up to an hour)15 minutesEvery destination, and a toast
Budget at 80%warningAlerts when a workspace or key budget reaches 80% of its limit for the period.Every 5 minutesa dayEvery destination, and a toast
Budget reachedcriticalAlerts when a workspace or key budget reaches 100% of its limit for the period.Every 5 minutesa dayEvery destination, and a toast
Billing mismatch spikewarningAlerts when 3 or more calls are billed against the vendor's published rule within an hour.Live, within seconds (windows up to an hour)15 minutesEvery destination, and a toast
Error-rate spikewarningAlerts when 50% or more of calls fail over 5 minutes (once there are at least 20 calls).Live, within seconds (windows up to an hour)15 minutesEvery destination, and a toast
Spend velocitywarningAlerts when calls cost $5.00 or more within 5 minutes.Live, within seconds (windows up to an hour)15 minutesEvery destination, and a toast
Key used from a new IPinfoAlerts when a key is used from an IP address it hasn't used before.When the signal arrives15 minutesDashboard toast only
Deliveries failingwarningAlerts when 10 or more webhook deliveries or log-drain batches fail within an hour.Every 5 minutes15 minutesEvery destination, and a toast
Security eventswarningAlerts on every security event.When the signal arrives15 minutesEvery destination, and a toast

"Key used from a new IP" is informational: it shows as a toast but goes to no destination until you route it to one. A key's first IP is its baseline and never alerts.

Rule settings

FieldMeaning
metric, thresholdWhat to measure, and the value that fires the rule (at or above, or at or below for cache_hit_rate)
window_sThe window the metric is measured over, 60 seconds to 7 days
min_eventsRates only: calls or lookups the window needs before the rule can fire (default 1)
severityinfo, warning (default) or critical
scope{ "type": "workspace" } (default), or key, vendor or capability with a value (a key id, vendor slug or capability id). Each metric lists the scopes it takes
kindssecurity_event only: the signal kinds that count (invalid_key_burst, new_geo, admin_action, key_ip_denied, scope_denied_burst); empty means every kind
cooldown_sAfter an alert resolves, the same rule and subject can't open a new one for this long (default 900, at most 86,400)
all_destinations, destination_idsNotify every enabled destination (default), or only the listed ones (up to 20)
toastAlso show the alert as a dashboard toast (default on)
muted_untilSnooze until this time; see Mute

Each destination has its own min_severity (alerts below it are not sent there) and notify_on_resolve (default on).

Lifecycle

An alert is triggered, then optionally acknowledged, then resolved. Its timeline records every step: when it triggered, each notification sent or failed, acknowledgement and resolution (the newest 60 entries, always keeping the trigger).

  • Dedupe. While an alert is open, further breaches of the same rule and subject don't open another one. They raise the alert's count and update its value and last_seen_at.
  • Grouping. Some metrics open one alert per subject instead of one per rule: vendor_failure_spike and vendor_key_failing per vendor, budget_used per budget (the workspace's or one key's), new_ip per key and security_event per signal kind. The subject is on the alert as subject.
  • Cooldown. After an alert resolves, the same rule and subject can't open a new alert for cooldown_s. The budget rules default to a day, so a budget that hovers around 80% alerts once a day at most.
  • Auto-resolve. A live rule reports that it has cleared once it has stayed on the right side of its threshold for a minute. A scheduled rule resolves on the next five-minute run that's no longer breached. An alert that nothing has confirmed for a while resolves too: 5 to 30 minutes for live rules (the window, within that range), 15 minutes for scheduled rules, and the rule's window for event rules.
  • Acknowledge. POST /v1/alerts/{id}/acknowledge marks a triggered alert as being handled. PagerDuty and Opsgenie are acknowledged too; chat and email destinations only hear about triggers and resolutions.
  • Resolve. POST /v1/alerts/{id}/resolve closes an alert now, and the rule's cooldown starts from there. Changing what a rule measures (metric, threshold, window, min_events, scope or kinds), disabling it or deleting it resolves its open alerts.

Mute

POST /v1/alerts/rules/{id}/mute with { "minutes": 60 } (or { "until": "<ISO time>" }) snoozes a rule (minutes goes up to 30 days). It keeps evaluating: its alerts still open and resolve, show in history marked muted and still show as toasts, but nothing is sent to destinations or webhook endpoints. { "minutes": 0 } unmutes it.

Live and scheduled evaluation

Where a rule runs depends on its metric and window:

  • Live. error_rate, spend, vendor_failure_spike, vendor_key_failing and billing_mismatch with windows up to an hour run in your workspace's live tail hub on every call, so they fire within seconds.
  • Scheduled. p95_latency, cache_hit_rate, budget_used, delivery_backlog and any live metric with a window over an hour are evaluated every five minutes, against your call log (hourly rollups for whole-hour windows), your budgets and the webhook delivery log.
  • Event. new_ip and security_event fire when the signal arrives.

Every rule shows its state: ok, firing, muted, disabled or no_data. no_data means the rule couldn't be evaluated, with the reason in state.note; for example, in a deployment without the call-log store, scheduled call metrics report no_data while every other rule keeps working.

Toasts

Alerts with toast on appear as toasts in the dashboard when they trigger and when they resolve. They arrive over the live tail as alert events, so any client that tails kinds=alert sees them too:

curl -N "https://api.gridrouter.io/v1/logs/live?format=sse&kinds=alert&replay=0" \
  -H "Authorization: Bearer $GRID_API_KEY"

Create a rule

Creating and changing rules needs a key with keys:write. This rule pages on-call when a third of email-finding calls fail over ten minutes, and only notifies one destination:

curl https://api.gridrouter.io/v1/alerts/rules \
  -H "Authorization: Bearer $GRID_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Email finding failing",
    "metric": "vendor_failure_spike",
    "threshold": 0.33,
    "window_s": 600,
    "min_events": 25,
    "severity": "critical",
    "scope": { "type": "capability", "value": "people.email.find" },
    "all_destinations": false,
    "destination_ids": ["dst_4f8k2m9q4w8x1c5v"]
  }'

The response is the rule with its id, its evaluator and its state. POST /v1/alerts/rules/{id}/test sends a test alert to every destination the rule notifies and reports each result; nothing is recorded as a real alert.

API and MCP

Every operation is a REST route and an MCP tool of the same name. Reads need logs:read, rule and alert changes need keys:write, and destinations need credentials:write because they hold secrets.

Reads, with logs:read:

  • alerts_summary: GET /v1/alerts/summary
  • alerts_rules_list and alerts_rules_get: GET /v1/alerts/rules and GET /v1/alerts/rules/{id}
  • alerts_list and alerts_get: GET /v1/alerts?status=open and GET /v1/alerts/{id}
  • alerts_destinations_list and alerts_destinations_get: GET /v1/alerts/destinations and GET /v1/alerts/destinations/{id}

Rule and alert changes, with keys:write:

  • alerts_rules_create: POST /v1/alerts/rules
  • alerts_rules_update and alerts_rules_delete: PUT and DELETE /v1/alerts/rules/{id}
  • alerts_rules_mute and alerts_rules_test: POST /v1/alerts/rules/{id}/mute and POST /v1/alerts/rules/{id}/test
  • alerts_acknowledge and alerts_resolve: POST /v1/alerts/{id}/acknowledge and POST /v1/alerts/{id}/resolve

Destination changes, with credentials:write:

  • alerts_destinations_create: POST /v1/alerts/destinations
  • alerts_destinations_update and alerts_destinations_delete: PUT and DELETE /v1/alerts/destinations/{id}
  • alerts_destinations_test: POST /v1/alerts/destinations/{id}/test

alerts_list takes status (open, triggered, acknowledged, resolved or all), rule_id and limit (up to 200). A PUT changes only the fields you send.

Limits

LimitValue
Alert rules per workspace50
Destinations per workspace20
Resolved alerts keptThe newest 2,000
Timeline entries per alert60