Configuration
The proxy runs with no config file at all out of the box. A config-less first run boots on mainnet, using embedded defaults for the escrow contract id and algod endpoint, and walks you through wallet setup (see Quick start).
How config is loaded
-
File. An explicit
--config <path>wins; otherwise the proxy checks thePROXY_CONFIGenvironment variable, thenconfig.yamlin its state directory (~/Library/Application Support/zerosignal/on macOS,~/.config/zerosignal/on Linux,%AppData%\zerosignal\on Windows), then./config.yamlin the current directory. -
Embedded defaults. File and environment values win where you set them; the embedded defaults fill the gaps.
-
Environment overrides. Settings can be overridden by an environment variable named after the YAML path, with
_separators and aPROXY_prefix:PROXY_SERVER_LISTENoverridesserver.listen, andPROXY_ZS_ESCROW_APP_IDoverrideszs.escrow_app_id. -
To inspect what's active:
zs-proxy config path # the file path that would be loadedzs-proxy config print-effective # the merged result: defaults + file + env + flags
server — where the proxy listens
| Key | Default | Notes |
|---|---|---|
listen | 127.0.0.1:9376 | Loopback by default; the proxy is meant to be reached only from your own machine. |
read_timeout | 30s | Request-read deadline. |
write_timeout | 0s | Keep at 0 (disabled); streaming responses require it. |
idle_timeout | 120s | Keep-alive idle deadline. |
algod — the Algorand connection
| Key | Default | Notes |
|---|---|---|
network | mainnet | The Algorand network to use. It selects the algod endpoint and token for you. |
endpoint | "" | Optional override for a private or paid algod endpoint. |
token | "" | Optional override. Prefer the PROXY_ALGOD_TOKEN env var so it stays out of a committed file. |
zs — operator selection and behavior
| Key | Default | Notes |
|---|---|---|
escrow_app_id | — | The deployed escrow contract id. On a public network you can usually omit it; the proxy falls back to the canonical app id for that network. |
operator_ids | all registered | Optional allowlist of on-chain operator ids to restrict routing to. Omit it to use every registered operator. |
concurrent_slots | 10 | How many requests can be in flight at once, each holding a ticket backed by the prepaid pool. Raise it if you run many requests in parallel (e.g. several agents at once). Change it with zs-proxy slots <n>, which also funds the pool to match; the limit only really moves when both change. |
fallback_max_output_tokens | 32768 | Output ceiling for a request that omits max_tokens, since the escrow ticket needs a bound to price against. The proxy sizes each reserve to the chosen operator: the model's declared output limit, or else a quarter of its context window capped by this value, so a small-context operator isn't skipped. Leave it at the default unless you know why you're changing it: the chat app uses the same value, and changing it makes the two disagree on which operators can serve a request. Set 0 to refuse such requests (max_output_required). default_max_output_tokens overrides it per model. A client that does send max_tokens is bounded by what operators advertise instead (Response length). |
privacy | true | Transport privacy: routes every request through a relay so the operator you talk to never sees your IP. Set false to talk directly to operators (lower latency, less private). |
allow_privacy_override | true | Whether a single request may opt out of privacy mode via the provider.relay: "off" field — see Routing preferences. |
allow_staging | false | Route to and list operator nodes flagged on-chain as staging. While it's off, no staging node appears in /v1/zs/operators or contributes to the pricing and context sizes on /v1/zs/details, and a model served only by staging nodes is absent from /v1/models and 404s on /v1/models/{id} — so it isn't advertised and then refused at dispatch. Turn it on and those nodes appear, each marked "staging": true in the operator directory. Leave it off unless you're testing one. Also settable per launch with --allow-staging. |
affinity_policy | prefer | How strictly multi-turn tool calls are routed back to the operator that handled the previous turn. prefer falls back if needed, strict refuses to fall back, none disables it. |
affinity_ttl | 10m | How long a tool-call continuation stays pinned to the operator that served the previous turn. |
affinity_max | 1024 | How many such pins the proxy remembers at once (a memory bound). |
default_max_output_tokens | (empty) | Per-model map of output ceilings for requests that omit max_tokens. A model without an entry falls through to fallback_max_output_tokens. File only; there is no environment form. |
reserve_timeout | 5s | Time limit for one price reservation with a candidate operator. On a timeout the proxy moves to the next candidate, so raising it (10s–30s, for slow operators across the open internet) only delays failover; it never blocks a request indefinitely. |
settlement_grace | — | Informational only; the contract enforces the real grace window. The proxy uses this value for the timeout hints it shows a client when a response lags past the refund window. |
proxy_payer_addr | derived | The Algorand address the wallet signs with. Optional when exactly one mnemonic is loaded, since the proxy derives it. Required when two or more are loaded, to say which one pays. Must be in the keystore either way. |
mbr_deposit_slots was renamed to concurrent_slots. The old key is no
longer read, so a config file still using it silently falls back to the
default of 10. If you had raised it, set the new key (or run
zs-proxy slots <n>). The environment variable moved too:
PROXY_ZS_MBR_DEPOSIT_SLOTS is now PROXY_ZS_CONCURRENT_SLOTS.
allow_private_operators lets the proxy dial private addresses, for a local
development deployment. Keep it off (the default) everywhere else: it guards
against the proxy being tricked into dialing an address on your own network.
zs.tee — confidential compute
Controls when the proxy insists on an operator running in
confidential compute, and how it
decides one qualifies. You can leave this whole block unset: the trust
anchors are compiled into the binary, so verification already works. Most
people set only require_for_models.
zs:
tee:
require_for_models:
- some-sensitive-model
| Key | Default | Notes |
|---|---|---|
require_for_models | (empty) | Model ids that may only be served by a verified-attested operator, on every request, without the caller asking. Matched exactly and case-sensitively against the model id on the wire. When no attested operator serves one, the request fails 503 no_tee_capacity rather than falling back — see Routing preferences. |
allow_stub | false | Accept a node's stub evidence, which is deterministic, well-formed, and proves nothing. Development only. In production it makes "attested" meaningless, since any node can emit a stub bundle. |
verifier | builtin | Which verifier implementation to use. builtin is the only one; nras is accepted as a legacy alias and normalized away. |
The remaining keys are trust anchors, and they ship populated: the compose
skeleton digest and image reference of every published zs-node release are
compiled in and updated automatically with each release. Setting them by hand
narrows what you accept; it never widens it:
| Key | Notes |
|---|---|
trusted_os_images | Guest OS image digests. Without this a node could present a valid quote taken on a dstack development image, where every other check passes identically but the platform's hardening does not hold. |
trusted_compose_skeletons / trusted_node_images | The published-release match. It checks that the node runs software from a published release, a stronger test than finding nothing structurally disqualifying. |
enforce_compose_skeleton | Whether that match is a gate. On by default. |
trusted_compose_hashes | Pin specific whole compose documents. Applied in addition to the release match, never instead of it. |
trusted_pre_launch_digests | Pre-launch measurement values, for platforms that publish them. |
An empty list means nothing verifies. It is never read as "accept anything."
Environment overrides follow the usual shape: PROXY_ZS_TEE_REQUIRE_FOR_MODELS
(comma-separated; empty resets the list), PROXY_ZS_TEE_ALLOW_STUB,
PROXY_ZS_TEE_VERIFIER.
To see what the verifier makes of a specific node, including which of the two
trust values it would need on the lists, run
zs-proxy tee inspect <node-url> --chain.
Discovery tuning
The proxy keeps a live picture of which operator nodes exist and what each one
can serve. It refreshes that picture in the background, reading the operator
set from the chain and then asking nodes directly what models, tools, and
prices they currently offer. Each round probes up to probe_budget nodes; the
rest keep their last-known state. You shouldn't need to touch these keys
unless you run against a large or unusually slow network.
| Key | Default | Notes |
|---|---|---|
refresh_interval | 5m | How often the background refresh runs. |
refresh_timeout | 30s | Time limit for one whole round. A round that overruns is abandoned and the previous picture is kept. |
max_cache_age | 15m | The oldest the saved picture may be before a request forces a refresh on the spot instead of waiting for the next round. Keep it comfortably above refresh_interval so that rarely happens. |
details_timeout | 2s | Time limit for asking a single node. Short, so one slow node can't starve the round. |
probe_budget | 96 | How many nodes one refresh round asks directly. A network at or under this size is checked in full every round; past it, rounds take turns. |
max_carry_age | 90m | How long a node that hasn't been asked recently keeps counting as available, based on what it last reported. |
target_topk | 8 | When you request a model, how many of its fastest known servers get re-checked that round, so the ones you're likely to use stay fresh. |
details_concurrency | 24 | How many nodes are asked at once within a round. Lower it if the proxy opens more sockets than your network likes. |
catalog_path | state dir | Where the last-known picture of the network is saved, so a restart resumes from it instead of rediscovering everything. Set "off" to disable saving. |
latency_state_max | 4096 | Upper bound on how many measured routes the proxy remembers timings for (a memory ceiling). |
If you change probe_budget, check max_carry_age too. Together with
refresh_interval they decide how long one full pass over the network takes
(roughly nodes ÷ probe_budget × refresh_interval). If max_carry_age is
shorter than that pass, nodes go stale before their turn comes round again and
visibly flicker in and out of /v1/models. Raising probe_budget also makes
each round do more work, so keep it in proportion to details_concurrency and
details_timeout so a round still finishes inside refresh_timeout.
Each of these can also be set as an environment variable:
PROXY_ZS_PROBE_BUDGET, PROXY_ZS_MAX_CARRY_AGE, and so on.
spend — your own safety cap
This is the one section worth configuring yourself, especially if you point an autonomous agent or a loop at the proxy. These caps sit on top of the network's own per-request price ceiling, so a misbehaving agent can't spend past the caps you set.
| Key | Default | Notes |
|---|---|---|
daily_cap_usdc | 0 (unlimited) | Maximum total USDC the proxy will commit per calendar day (UTC); the counter resets at midnight UTC. Persisted across restarts. |
per_request_cap_usdc | 0 (unlimited) | Maximum USDC a single request may cost. A request that would exceed it is refused before it's sent. |
spend:
daily_cap_usdc: 5.00
per_request_cap_usdc: 0.50
Both are whole USDC (e.g. 5.00), and either can also be set via
PROXY_SPEND_DAILY_CAP_USDC / PROXY_SPEND_PER_REQUEST_CAP_USDC. The daily
counter tracks your actual settled spend: what each request really cost
after the operator's signed receipt reconciles, usually far below the price
ceiling reserved on-chain, so refunded headroom returns to the cap. When
deciding whether to admit a new request, it also counts requests still in
flight, so the worst-case commitment can never exceed the cap.
zs-proxy status shows today's spend against it.
logging
| Key | Default | Values |
|---|---|---|
level | info | debug | info | warn | error |
format | text | text | json |
The proxy never logs prompts, responses, or tool-call content, only metadata such as token counts, ticket/operator ids, model names, and cost. See Privacy & security.
A minimal config.yaml
To pin mainnet with a spend cap, this is enough:
algod:
network: "mainnet"
spend:
daily_cap_usdc: 10.00
per_request_cap_usdc: 1.00
What's next
- Wallet & funding — funding, the prepaid pool, and
wallet_unfunded. - Routing preferences — per-request operator selection, pricing ceilings, and the relay override.
- Running as a service — keep the proxy up across reboots.