Routing preferences
By default the proxy picks an operator the same way the chat app does, favoring a fast, cheap, available operator for the model you asked for.
For finer control, from your own code or a tool that exposes it, add an
optional top-level provider object to the request body (OpenRouter-style).
The proxy reads it to pick an operator, then strips it before the request leaves
your machine, so it never reaches the operator or the model. A missing or
malformed field is ignored.
{
"model": "qwen3-coder",
"messages": [/* … */],
"provider": {
"only": ["1234", "5678:2"], // allowlist of operator refs
"ignore": ["9999:1"], // denylist (deny wins over allow)
"order": ["5678:2", "1234"], // explicit preference order
"allow_fallbacks": false, // false ⇒ pin to exactly `order`/`only`
"max_price": { "input": 2.0, "output": 6.0 }, // USD per 1M tokens, ceilings
"require_tools": true, // only route to operators with built-in tools
"sort": "latency", // price | throughput | latency
"relay": "off" // per-request transport-privacy override
}
}
A ref is either "<operatorId>" (any node run by that operator) or
"<operatorId>:<nodeId>" (one exact node). Operator and node ids are listed by
GET /v1/zs/operators and in the chat app's model picker.
A node flagged staging isn't listed there unless you turned
on allow_staging — its operator still is, if it also runs a production node, but
the operatorId:nodeId row for the staging node is absent. You can still write
such a ref by hand, because refs aren't checked against the directory, but the
request fails: the ref filter runs first and lets the pin through, then the staging
step drops it, and you get 400 no_production_operator rather than the
400 no_pinned_operator an unmatched pin normally gives.
A node whose signing account has less than 5 ALGO spendable (that account pays
the network fees of every request the node serves) stays listed in
GET /v1/zs/operators, with signer_underfunded: true
and its balance in signer_spendable_microalgos, so you can see why it isn't
being picked. It still won't be routed to: a pin that only matches it returns
503 no_funded_operator, and it comes back on its own within about five
minutes of the operator topping it up.
Fields
| Field | What it does |
|---|---|
only | Restrict routing to these operator refs. |
ignore | Exclude these operator refs — wins over only if both match the same ref. |
order | Try these operators in this order before falling back to the rest. |
allow_fallbacks | false pins the request to exactly only/order — no fallback to anything else if they're all unavailable. |
max_price | USD-per-1M-token ceilings (input/output); operators above either are excluded. |
require_tools | Only route to operators that advertise the built-in tools the request needs. |
sort | Re-weight ordering among the remaining candidates — see below. |
relay | Override transport privacy for this one request — see below. |
A previous_response_id continuation stays pinned to the operator whose account
holds the conversation, and that pin always wins over provider. If your
preferences exclude that operator, the request is not re-routed; you get
400 pin_conflicts_with_continuation.
When preferences can't be satisfied
| Situation | What you get back |
|---|---|
only/order with allow_fallbacks: false matched no available operator | 400 no_pinned_operator |
A previous_response_id continuation is excluded by only/ignore/order | 400 pin_conflicts_with_continuation |
max_price excluded every operator | 400 price_ceiling_exceeded |
require_tools and no operator advertises the requested tools | 503 no_tool_capacity |
| TEE was required and no operator serving the model is verified-attested | 503 no_tee_capacity |
Every operator serving the model is a staging node and you haven't enabled allow_staging | 400 no_production_operator |
Every operator left for the model — or the one a pin or a previous_response_id continuation names — has less than 5 ALGO spendable in its signing account | 503 no_funded_operator |
sort — what to optimize for
By default the proxy bands candidates by a combined expected-response-time
score (network latency + time-to-first-token + decode speed) and picks the
cheapest within the fastest band. sort changes that:
| Value | Optimizes for |
|---|---|
latency | Time to first token only (ignores decode speed). |
throughput | Decode speed only (ignores time to first token). |
price | Price only: the cheapest available operator. |
| (absent) | The combined default. |
relay — transport privacy, per request
By default the proxy routes each request through another operator acting as a
relay, so the operator running your model never sees your IP — see
Privacy & security. relay overrides that for one request:
| Value | Effect |
|---|---|
"off" | Send directly to the target — lower latency, but it sees your IP. Only honored if zs.allow_privacy_override is enabled (the default). |
"required" | Force a relay hop, failing with 503 no_relay_available rather than send directly. |
"auto" / absent | Use the configured default (zs.privacy). |
The X-Zs-Relay request header sets the same thing; if both are present, the
body wins.
X-Zs-Require-TEE — attested hardware, per request
With this header, the proxy routes the request only to an operator running in
confidential compute whose attestation
it verified itself: it fetches the node's hardware evidence, checks the
quote to Intel's root, replays the measurement log, and confirms the measured
software is a zs-node build we published.
curl http://localhost:9376/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'X-Zs-Require-TEE: 1' \
-d '{"model":"...","messages":[{"role":"user","content":"..."}]}'
- Any non-empty value turns it on.
1,true,yes— the proxy checks only that the header is present and non-blank. There is no way to spell "off"; omit the header instead. - It does not enter the sealed envelope. Only your own proxy reads it, to choose an operator. It is not sent to the node, and a node cannot see it.
- It never falls back. If every operator that could otherwise serve the
request is unattested, you get
503 no_tee_capacity, a distinct code from the genericno_operator. - The proxy's own check decides. An operator's advertised
tee.modemakes the proxy run the verification; it does not pass the check by itself.
To require it for certain models on every request instead of per call, use
zs.tee.require_for_models.
Either that setting or the header is enough to require attestation.
What's next
- Configuration — the config-file defaults these fields override per request.
- Connecting AI tools — the rest of the request API.