# Waterfalls (/docs/concepts/waterfalls)



A **waterfall** is an ordered set of catalog endpoints for one capability, plus a speed profile,
stop rules and a merge strategy per output field. A **pipeline** chains up to 10 waterfalls as
stages, and a later stage can read an earlier one's output (`stages.<stage_id>.<field>`). Build them
in the dashboard at `/waterfalls`, or with the [waterfall API](/docs/api/waterfalls/waterfall_create).
Most start from one of the [waterfall templates](/docs/concepts/waterfall-templates) (browse them
with each step explained at [/waterfalls/templates](/waterfalls/templates)); see
[build a waterfall](/docs/concepts/build-a-waterfall) and [merge rules](/docs/concepts/merge-rules).
Most start from one of the [waterfall templates](/docs/concepts/waterfall-templates) (browse them
with each step explained at [/waterfalls/templates](/waterfalls/templates)); see
[build a waterfall](/docs/concepts/build-a-waterfall) and [merge rules](/docs/concepts/merge-rules).

## Steps [#steps]

A step is one endpoint (`hunter/email.find`) or a **routed** pool: the capability's endpoints,
filtered with the same [provider preferences](/docs/concepts/routing#provider-preferences) as
`/v1/run`. Each step can set a timeout, retries, a max cost, required fields, a minimum confidence,
verified-only, what a vendor 4xx does, and a `when` guard (field conditions such as `eq`, `gte`,
`exists`, `in`, combined with `all`, `any` and `not`).

## Speed profiles [#speed-profiles]

| Profile    | Behavior                                                                | Default deadline |
| ---------- | ----------------------------------------------------------------------- | ---------------- |
| `fastest`  | Starts 3 steps at once, hedges the next every 750 ms, aborts the losers | 15 s             |
| `balanced` | One at a time; hedges a slow step after 3 s                             | 30 s             |
| `cheapest` | Strictly sequential, lowest quote first                                 | 60 s             |
| `custom`   | Your concurrency, hedge delay, step timeout, deadline and sort          | 30 s             |

## Stop rules and merge [#stop-rules-and-merge]

Stop rules combine with `stop_mode: any | all`: `first_hit`, `first_verified`,
`until_fields_filled`, `all_then_merge` and `condition`.

Each output field has its own merge strategy: `highest_quality`, `consensus`, `most_recent`,
`prefer_verified`, `first_non_null` or `union`. Values that disagree are kept in
`_provenance.<field>.conflicts`.

## Running [#running]

| Mode      | How                                                             | Response                                                                    |
| --------- | --------------------------------------------------------------- | --------------------------------------------------------------------------- |
| Sync      | `POST /v1/waterfalls/{id}/run`                                  | `200` with the run (deadline capped at 50 s)                                |
| Wait      | `Prefer: wait=N`                                                | `200` if done within N s, else `202` with a `Location` to poll              |
| Async     | `Prefer: respond-async`                                         | `202`; poll the run, stream `/events` (SSE) or `/ws`, or pass `webhook_url` |
| Scheduled | `"delay_seconds": N` (≤ 12 h)                                   | `202`                                                                       |
| Batch     | `POST /v1/waterfalls/batch-run`, up to 1,000 JSON rows or a CSV | a list id; rows are paced per vendor                                        |

Durable runs execute each attempt as its own step, so a crash resumes without paying for an attempt
twice. `POST …/runs/{run_id}/cancel` stops a run between attempts.

A run returns `outcome` (`hit`, `partial`, `miss`, `failed`), `stop_reason`, the merged record in
`data`, `_provenance` per field (vendor, call id, confidence, verified) and `_attempts`. Every attempt
is its own row in the [call log](/docs/concepts/logs-and-drains) with the run id as its parent.

## Publish [#publish]

Publishing freezes a semver version. A published waterfall is your own endpoint at
`POST /v1/x/{workspace}/{name}`, an MCP tool call (`x_run`) and an OpenAPI document
(`GET /v1/waterfalls/{id}/openapi`). Roll back with `POST /v1/waterfalls/{id}/rollback`.

## Errors [#errors]

| Case                           | Behavior                                                                               |
| ------------------------------ | -------------------------------------------------------------------------------------- |
| Invalid definition or input    | `422 validation_failed` with `details.issues[].path`                                   |
| Vendor 4xx                     | Not retried; next step                                                                 |
| Vendor 429                     | Next step now; back to this one after `Retry-After`, within the deadline               |
| Vendor 5xx or timeout          | Retried with backoff up to the step's `retries`, then the next step                    |
| Run budget or deadline reached | Returns the partial record with `stop_reason` `budget_exceeded` or `deadline_exceeded` |
| Every step misses              | `200` with `outcome: miss`                                                             |
| Account errors                 | Stop everything; the partial run is in `details.run`                                   |

