Waterfalls
Save an ordered set of vendors, a speed profile, stop rules and a per-field merge as your own versioned endpoint.
Before you start
A waterfall is an ordered set of catalog endpoints for one capability, plus a speed profile,
stop rules and a merge strategy per output field. A pipeline chains up to 10 waterfalls as
stages, and a later stage can read an earlier one's output (stages.<stage_id>.<field>). Build them
in the dashboard at /waterfalls, or with the waterfall API.
Steps
A step is one endpoint (hunter/email.find) or a routed pool: the capability's endpoints,
filtered with the same provider preferences as
/v1/run. Each step can set a timeout, retries, a max cost, required fields, a minimum confidence,
verified-only, what a vendor 4xx does, and a when guard (field conditions such as eq, gte,
exists, in, combined with all, any and not).
Speed profiles
| Profile | Behavior | Default deadline |
|---|---|---|
fastest | Starts 3 steps at once, hedges the next every 750 ms, aborts the losers | 15 s |
balanced | One at a time; hedges a slow step after 3 s | 30 s |
cheapest | Strictly sequential, lowest quote first | 60 s |
custom | Your concurrency, hedge delay, step timeout, deadline and sort | 30 s |
Stop rules and merge
Stop rules combine with stop_mode: any | all: first_hit, first_verified,
until_fields_filled, all_then_merge and condition.
Each output field has its own merge strategy: highest_quality, consensus, most_recent,
prefer_verified, first_non_null or union. Values that disagree are kept in
_provenance.<field>.conflicts.
Running
| Mode | How | Response |
|---|---|---|
| Sync | POST /v1/waterfalls/{id}/run | 200 with the run (deadline capped at 50 s) |
| Wait | Prefer: wait=N | 200 if done within N s, else 202 with a Location to poll |
| Async | Prefer: respond-async | 202; poll the run, stream /events (SSE) or /ws, or pass webhook_url |
| Scheduled | "delay_seconds": N (≤ 12 h) | 202 |
| Batch | POST /v1/waterfalls/batch-run, up to 1,000 JSON rows or a CSV | a list id; rows are paced per vendor |
Durable runs execute each attempt as its own step, so a crash resumes without paying for an attempt
twice. POST …/runs/{run_id}/cancel stops a run between attempts.
A run returns outcome (hit, partial, miss, failed), stop_reason, the merged record in
data, _provenance per field (vendor, call id, confidence, verified) and _attempts. Every attempt
is its own row in the call log with the run id as its parent.
Publish
Publishing freezes a semver version. A published waterfall is your own endpoint at
POST /v1/x/{workspace}/{name}, an MCP tool call (x_run) and an OpenAPI document
(GET /v1/waterfalls/{id}/openapi). Roll back with POST /v1/waterfalls/{id}/rollback.
Errors
| Case | Behavior |
|---|---|
| Invalid definition or input | 422 validation_failed with details.issues[].path |
| Vendor 4xx | Not retried; next step |
| Vendor 429 | Next step now; back to this one after Retry-After, within the deadline |
| Vendor 5xx or timeout | Retried with backoff up to the step's retries, then the next step |
| Run budget or deadline reached | Returns the partial record with stop_reason budget_exceeded or deadline_exceeded |
| Every step misses | 200 with outcome: miss |
| Account errors | Stop everything; the partial run is in details.run |