# Logs, live tail and drains (/docs/concepts/logs-and-drains)



Every call, waterfall step, job and list update is written to your workspace's log and published
to its **live tail**. The dashboard, the [CLI](/docs/cli), the [SDK](/docs/sdk) and
[MCP clients](/docs/mcp) subscribe to the same stream, and **log drains** batch it out to your
webhook, Axiom, Datadog or object storage.

## Search the log [#search-the-log]

`/logs` and [`GET /v1/logs`](/docs/api/logs/logs_list) share one query language:

```text
provider:hunter status:>=400 -capability:email.find latency:>2000 cost:<0.05 error:"rate limited"
```

* `field:a,b` includes values and `-field:a` excludes them. Fields: `outcome`, `status`,
  `provider`, `capability`, `endpoint`, `billing`, `client`, `client_mode`, `mode`, `credential`,
  `key`, `error`, `country` and `cache`.
* `latency` (ms), `cost` (USD) and `status` take `>`, `>=`, `<`, `<=` and `a..b`. `status:4xx` is
  400–499.
* A bare word searches call, run and parent ids, the endpoint and the error code.

Export streams CSV or NDJSON for the current filter, up to 10,000 rows.

## Live tail [#live-tail]

`GET /v1/logs/live` streams over WebSocket, or SSE with `?format=sse`. It needs `logs:read`.

```bash
curl -N "https://api.gridrouter.io/v1/logs/live?format=sse&kinds=call&status=error" \
  -H "Authorization: Bearer $GRID_API_KEY"
```

| Filter                                                                                                 | Meaning                                                        |
| ------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------- |
| `kinds`                                                                                                | `call`, `run`, `job`, `list`, `alert` (default all)            |
| `provider`, `capability`, `endpoint_id`, `key_id`, `app_id`, `credential`, `http_status`, `error_code` | Call dimensions                                                |
| `status`                                                                                               | `hit`, `miss`, `failed` (`error` is an alias)                  |
| `run_id`, `waterfall_id`                                                                               | One run and its attempts, or one waterfall                     |
| `latency_min`, `cost_min`                                                                              | Minimum latency (ms) and cost (micro-USD)                      |
| `query`                                                                                                | The log query language above (calls only)                      |
| `sample`                                                                                               | 0–1, deterministic: every subscriber keeps the same events     |
| `fields`                                                                                               | Project call rows to these columns                             |
| `replay`                                                                                               | Buffered events sent first, 0–1,000 (default 100)              |
| `max_rate`                                                                                             | Events per second for this subscriber (default 500, max 5,000) |

Each frame is JSON with `seq`, `ts` and `shard`; `(shard, seq)` is unique, so de-duplicate on it
after a reconnect. Control frames share the stream: `hello`, `dropped` (events this subscriber
missed because it was too slow) and `heartbeat` (every 15 s). A stream is closed after 30 minutes;
reconnect and replay.

`GET /v1/logs/live/recent?limit=50` returns the newest events without streaming. A call made with
`observability.tail: false` is logged but not published to the tail.

## Alerts [#alerts]

Fast [alert rules](/docs/concepts/alerts) (error rate, spend, vendor failures, vendor keys and
billing mismatches over windows up to an hour) run on the tail as calls arrive, and alerts reach
the dashboard as `alert` events (`kinds=alert`). Rules, destinations and
[webhooks](/docs/concepts/webhooks) are documented on their own pages.

## Drains [#drains]

<Callout title="Plan">
  Log drains are a Pro feature (2 drains; Team and above unlimited). On Free, creating one fails
  with `plan_limit_reached`.
</Callout>

| Type              | Delivery                                                                  |
| ----------------- | ------------------------------------------------------------------------- |
| `webhook`         | `POST {drain_id, batch_id, events}` as JSON, signed with `Grid-Signature` |
| `axiom`           | NDJSON to your dataset's ingest endpoint (`token`)                        |
| `datadog`         | JSON to the Logs intake, `ddsource: gridrouter` (`api_key`)               |
| `s3`, `gcs`, `r2` | SigV4-signed `PUT` of `prefix/YYYY/MM/DD/HH/<batch>.ndjson.gz`            |

Batches go out at 200 events or after 5 s (configurable up to 1,000 events and 60 s). `batch_id`
is stable, so de-duplicate on it. 408, 429, 5xx and network errors retry 8 times with backoff from
10 s up to 30 minutes. Each drain reports its health: `idle`, `healthy`, `degraded`, `failing` or
`disabled`.

A webhook drain's `whsec_…` secret is shown once. Verify each batch against
`Grid-Signature: t=<unix seconds>,v1=<hex HMAC-SHA256 of "t.body">` with a 5-minute tolerance.
Drain credentials are write-only, and batches carry only public call columns, never bodies.

