Logs, live tail and drains
Every call is logged and streamed as it happens. Search it, tail it from the CLI, SDK or MCP, and drain it to your own stack.
Every call, waterfall step, job and list update is written to your workspace's log and published to its live tail. The dashboard, the CLI, the SDK and MCP clients subscribe to the same stream, and log drains batch it out to your webhook, Axiom, Datadog or object storage.
Search the log
/logs and GET /v1/logs share one query language:
provider:hunter status:>=400 -capability:email.find latency:>2000 cost:<0.05 error:"rate limited"field:a,bincludes values and-field:aexcludes them. Fields:outcome,status,provider,capability,endpoint,billing,client,client_mode,mode,credential,key,error,countryandcache.latency(ms),cost(USD) andstatustake>,>=,<,<=anda..b.status:4xxis 400–499.- A bare word searches call, run and parent ids, the endpoint and the error code.
Export streams CSV or NDJSON for the current filter, up to 10,000 rows.
Live tail
GET /v1/logs/live streams over WebSocket, or SSE with ?format=sse. It needs logs:read.
curl -N "https://api.gridrouter.io/v1/logs/live?format=sse&kinds=call&status=error" \
-H "Authorization: Bearer $GRID_API_KEY"| Filter | Meaning |
|---|---|
kinds | call, run, job, list, alert (default all) |
provider, capability, endpoint_id, key_id, app_id, credential, http_status, error_code | Call dimensions |
status | hit, miss, failed (error is an alias) |
run_id, waterfall_id | One run and its attempts, or one waterfall |
latency_min, cost_min | Minimum latency (ms) and cost (micro-USD) |
query | The log query language above (calls only) |
sample | 0–1, deterministic: every subscriber keeps the same events |
fields | Project call rows to these columns |
replay | Buffered events sent first, 0–1,000 (default 100) |
max_rate | Events per second for this subscriber (default 500, max 5,000) |
Each frame is JSON with seq, ts and shard; (shard, seq) is unique, so de-duplicate on it
after a reconnect. Control frames share the stream: hello, dropped (events this subscriber
missed because it was too slow) and heartbeat (every 15 s). A stream is closed after 30 minutes;
reconnect and replay.
GET /v1/logs/live/recent?limit=50 returns the newest events without streaming. A call made with
observability.tail: false is logged but not published to the tail.
Alerts
Fast alert rules (error rate, spend, vendor failures, vendor keys and
billing mismatches over windows up to an hour) run on the tail as calls arrive, and alerts reach
the dashboard as alert events (kinds=alert). Rules, destinations and
webhooks are documented on their own pages.
Drains
Plan
Log drains are a Pro feature (2 drains; Team and above unlimited). On Free, creating one fails
with plan_limit_reached.
| Type | Delivery |
|---|---|
webhook | POST {drain_id, batch_id, events} as JSON, signed with Grid-Signature |
axiom | NDJSON to your dataset's ingest endpoint (token) |
datadog | JSON to the Logs intake, ddsource: gridrouter (api_key) |
s3, gcs, r2 | SigV4-signed PUT of prefix/YYYY/MM/DD/HH/<batch>.ndjson.gz |
Batches go out at 200 events or after 5 s (configurable up to 1,000 events and 60 s). batch_id
is stable, so de-duplicate on it. 408, 429, 5xx and network errors retry 8 times with backoff from
10 s up to 30 minutes. Each drain reports its health: idle, healthy, degraded, failing or
disabled.
A webhook drain's whsec_… secret is shown once. Verify each batch against
Grid-Signature: t=<unix seconds>,v1=<hex HMAC-SHA256 of "t.body"> with a 5-minute tolerance.
Drain credentials are write-only, and batches carry only public call columns, never bodies.
Private cache
Your workspace's encrypted cache of what it already fetched. Repeats are free, and partly known records only fetch the missing fields.
Alerts
Rules over your calls, budgets, vendor keys, billing checks and deliveries. Alerts open, repeat, get acknowledged and resolve, and notify the destinations you choose.