Skip to main content

Routing preferences

By default the proxy picks an operator the same way the chat app does, favoring a fast, cheap, available operator for the model you asked for.

For finer control, from your own code or a tool that exposes it, add an optional top-level provider object to the request body (OpenRouter-style). The proxy reads it to pick an operator, then strips it before the request leaves your machine, so it never reaches the operator or the model. A missing or malformed field is ignored.

{
"model": "qwen3-coder",
"messages": [/* … */],
"provider": {
"only": ["1234", "5678:2"], // allowlist of operator refs
"ignore": ["9999:1"], // denylist (deny wins over allow)
"order": ["5678:2", "1234"], // explicit preference order
"allow_fallbacks": false, // false ⇒ pin to exactly `order`/`only`
"max_price": { "input": 2.0, "output": 6.0 }, // USD per 1M tokens, ceilings
"require_tools": true, // only route to operators with built-in tools
"sort": "latency", // price | throughput | latency
"relay": "off" // per-request transport-privacy override
}
}

A ref is either "<operatorId>" (any node run by that operator) or "<operatorId>:<nodeId>" (one exact node). Operator and node ids are listed by GET /v1/zs/operators and in the chat app's model picker.

A node flagged staging isn't listed there unless you turned on allow_staging — its operator still is, if it also runs a production node, but the operatorId:nodeId row for the staging node is absent. You can still write such a ref by hand, because refs aren't checked against the directory, but the request fails: the ref filter runs first and lets the pin through, then the staging step drops it, and you get 400 no_production_operator rather than the 400 no_pinned_operator an unmatched pin normally gives.

A node whose signing account has less than 5 ALGO spendable (that account pays the network fees of every request the node serves) stays listed in GET /v1/zs/operators, with signer_underfunded: true and its balance in signer_spendable_microalgos, so you can see why it isn't being picked. It still won't be routed to: a pin that only matches it returns 503 no_funded_operator, and it comes back on its own within about five minutes of the operator topping it up.

Fields​

FieldWhat it does
onlyRestrict routing to these operator refs.
ignoreExclude these operator refs — wins over only if both match the same ref.
orderTry these operators in this order before falling back to the rest.
allow_fallbacksfalse pins the request to exactly only/order — no fallback to anything else if they're all unavailable.
max_priceUSD-per-1M-token ceilings (input/output); operators above either are excluded.
require_toolsOnly route to operators that advertise the built-in tools the request needs.
sortRe-weight ordering among the remaining candidates — see below.
relayOverride transport privacy for this one request — see below.

A previous_response_id continuation stays pinned to the operator whose account holds the conversation, and that pin always wins over provider. If your preferences exclude that operator, the request is not re-routed; you get 400 pin_conflicts_with_continuation.

When preferences can't be satisfied​

SituationWhat you get back
only/order with allow_fallbacks: false matched no available operator400 no_pinned_operator
A previous_response_id continuation is excluded by only/ignore/order400 pin_conflicts_with_continuation
max_price excluded every operator400 price_ceiling_exceeded
require_tools and no operator advertises the requested tools503 no_tool_capacity
TEE was required and no operator serving the model is verified-attested503 no_tee_capacity
Every operator serving the model is a staging node and you haven't enabled allow_staging400 no_production_operator
Every operator left for the model — or the one a pin or a previous_response_id continuation names — has less than 5 ALGO spendable in its signing account503 no_funded_operator

sort — what to optimize for​

By default the proxy bands candidates by a combined expected-response-time score (network latency + time-to-first-token + decode speed) and picks the cheapest within the fastest band. sort changes that:

ValueOptimizes for
latencyTime to first token only (ignores decode speed).
throughputDecode speed only (ignores time to first token).
pricePrice only: the cheapest available operator.
(absent)The combined default.

relay — transport privacy, per request​

By default the proxy routes each request through another operator acting as a relay, so the operator running your model never sees your IP — see Privacy & security. relay overrides that for one request:

ValueEffect
"off"Send directly to the target — lower latency, but it sees your IP. Only honored if zs.allow_privacy_override is enabled (the default).
"required"Force a relay hop, failing with 503 no_relay_available rather than send directly.
"auto" / absentUse the configured default (zs.privacy).

The X-Zs-Relay request header sets the same thing; if both are present, the body wins.

X-Zs-Require-TEE — attested hardware, per request​

With this header, the proxy routes the request only to an operator running in confidential compute whose attestation it verified itself: it fetches the node's hardware evidence, checks the quote to Intel's root, replays the measurement log, and confirms the measured software is a zs-node build we published.

curl http://localhost:9376/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'X-Zs-Require-TEE: 1' \
-d '{"model":"...","messages":[{"role":"user","content":"..."}]}'
  • Any non-empty value turns it on. 1, true, yes — the proxy checks only that the header is present and non-blank. There is no way to spell "off"; omit the header instead.
  • It does not enter the sealed envelope. Only your own proxy reads it, to choose an operator. It is not sent to the node, and a node cannot see it.
  • It never falls back. If every operator that could otherwise serve the request is unattested, you get 503 no_tee_capacity, a distinct code from the generic no_operator.
  • The proxy's own check decides. An operator's advertised tee.mode makes the proxy run the verification; it does not pass the check by itself.

To require it for certain models on every request instead of per call, use zs.tee.require_for_models. Either that setting or the header is enough to require attestation.

What's next​