Configuration
Operator and escrow settings live under zs:, with NODE_ZS_* environment
overrides. For the reasoning behind pricing and the model catalog, see
Serving models & pricing; for the signing mnemonic, see
Encryption & keys.
Generate it: zs-node init
Before writing this file by hand, try the wizard. Point it at your inference
backend and it produces a complete, commented config.yaml:
zs-node init --base-url=http://127.0.0.1:8080/v1
It queries the backend:
- It identifies the runtime (vLLM, SGLang, llama.cpp, LM Studio, Ollama, or a
hosted gateway), sets
llm.providerto match, and finds the real API root if you gave it the wrong one. - It checks whether
/v1/responsesexists, which decidesllm.openai.translate_responses_to_chat, and whether streaming reports token usage. A backend that doesn't is rejected, because the node bills streaming work from that usage and refuses to start without it. - It probes each model for tool calling, reasoning, image input, per-prompt image limits, and its output ceiling.
- It looks up list pricing in the models.dev catalog, applies your margin, and checks your on-chain operator and node records, your keystore, your USDC opt-in, and your signing balance.
The probe sends real requests, one to sixteen output tokens each, so it learns
what your deployment does. It catches deployment gaps, such as a vLLM started
without --enable-auto-tool-choice, which cannot make tool calls whatever the
model card says. Against a local backend it just runs;
against a metered endpoint it asks first. Use --probe=passive to read metadata
only.
The wizard loads the generated file exactly as the daemon would and refuses to write anything that wouldn't start. Every value is annotated with where it came from, and anything surprising becomes a note in the file header.
# Re-run against an existing config; your own rates and settings are kept.
zs-node init --from=config.yaml --dry-run
# A paid gateway, two models, resold at a 25% margin.
zs-node init --base-url=https://api.z.ai/api/paas/v4 --models=glm-4.6 --margin=25
# Unattended.
zs-node init --base-url=http://vllm:8000/v1 --network=mainnet \
--operator-id=7 --node-id=1 --non-interactive --yes
It never writes an API key or a mnemonic into the file, never submits a
transaction, and always keeps a .bak copy before overwriting. Run
zs-node init -h for every flag.
How config is loaded
- File. The node reads
./config.yamlby default. Point it elsewhere with the--config <path>flag or theNODE_CONFIGenvironment variable. - Environment overrides. A curated set of settings can be overridden by an
environment variable with the
NODE_prefix following the YAML path — e.g.NODE_SERVER_LISTENoverridesserver.listen,NODE_ZS_SELF_EVICTION_ENABLEDoverrideszs.self_eviction.enabled. Not every key has one: several timeouts, the oracle,min_charge, and the rate-limit buckets are file-only. Where an override exists, the environment wins over the file. - Secrets come from the environment, never the file. API keys
(
NODE_LLM_OPENAI_API_KEY), the algod token (NODE_ALGOD_TOKEN), and the signing mnemonic (OPERATOR_SIGNING_MNEMONIC/ZS_MNEMONIC_URLS) are read from the environment or a secret manager. Don't commit them toconfig.yaml.
Three timeouts must stay 0 for streaming. server.write_timeout,
llm.openai.timeout, and llm.vertexai.timeout all default to 0s (disabled).
Inference responses are long-lived SSE streams, and a non-zero value cuts a
generation off mid-stream.
server — listeners, timeouts, TLS
| Key | Default | What it controls |
|---|---|---|
listen | 127.0.0.1:9090 | Public listener (prompts, reserve, discovery). A serving node binds all interfaces with :9090 (dual-stack; 0.0.0.0 is IPv4-only), which zs-node init writes; see The two listeners. Override with NODE_SERVER_LISTEN. |
private_listen | 127.0.0.1:9091 | Ops listener for /healthz, /livez and /metrics, off the public port and exempt from rate limiting. Set "" to colocate them on listen. Override with NODE_SERVER_PRIVATE_LISTEN. |
read_timeout | 30s | Request-read deadline. |
write_timeout | 0s | Keep at 0 — disabled, required for streams. |
idle_timeout | 120s | Keep-alive idle deadline. |
sse_keepalive_interval | 15s | Sends a : zs-keepalive SSE comment whenever a streaming response produces no bytes (a long prefill before the first token, or a gap between tokens), so a reverse proxy or relay in front of you doesn't reset the idle stream at its own timeout. SSE clients ignore the comment, and it never touches billing. Set it well below your front door's idle timeout; 0 disables. See Clients get 502 or 504 errors. Override with NODE_SERVER_SSE_KEEPALIVE_INTERVAL. |
drain_grace | 20s | On SIGTERM, how long reservations issued before the shutdown are still honored, so a request already in flight lands. New reservations are refused immediately regardless. Override with NODE_SERVER_DRAIN_GRACE. |
drain_timeout | 5m | How long shutdown waits for inference already running to finish. Size it for your worst-case single response, but no longer than any per-request timeout in front of the node, which would cut the stream anyway. 0 opts out of draining entirely and also disables drain_grace (the grace hold runs inside the same wait). Override with NODE_SERVER_DRAIN_TIMEOUT. |
shutdown_timeout | 30s | Final connection-close budget after inference has drained. 0 is treated as 30s, not an opt-out, so a connection that never closes can't hang the process past your supervisor's deadline. Override with NODE_SERVER_SHUTDOWN_TIMEOUT. |
Your supervisor's kill deadline must exceed drain_grace + drain_timeout + shutdown_timeout (5m50s with the defaults) or it force-kills the node mid-drain.
See Graceful shutdown.
The private listener is never TLS-wrapped. In YAML, quote a leading-colon
value like ":9091". The routes each listener serves are in the
Endpoint reference.
server.tls
TLS for the public listener. Three modes:
| Mode | What it does |
|---|---|
off (default) | Plain HTTP (loopback-only default). Prefer acme to serve HTTPS directly; or terminate TLS at a single reverse proxy in front of one node. |
manual | Load cert_path + key_path from disk at startup; rotate by replacing the files and restarting. |
acme | Provision and renew a Let's Encrypt certificate automatically (detailed below). |
cert_path and key_path are both required when mode: manual, and are read
once at startup — there is no hot reload.
The acme mode provisions and renews a Let's Encrypt certificate in the
background via DNS-01, solved against your operator's NFD u.dns (the only
supported DNS provider). Prerequisites, all enforced at boot:
algod.network∈{testnet, mainnet}— localnet has no public DNS suffix.- A non-zero NFD app id, from
zs.nfd_app_idor gap-filled from the operator box whenzs.escrow_app_idis set. Otherwise the node exits withtls.mode=acme requires the operator's NFD on chain — set zs.escrow_app_id and ensure the operator is registered with an NFD, or set tls.mode=off. - The keystore holds the key for whichever account currently owns that NFD. That is not automatically your node's signing address, and getting it wrong fails at runtime rather than at boot — see Which account signs the DNS writes.
| Key | Default | Notes |
|---|---|---|
cache_dir | ./tls-cache | Relative to the process working directory. Holds the Let's Encrypt account key as well as issued certs, so it must be writable and durable — a container without a volume or a shifting WorkingDirectory re-registers the account every boot and eventually hits LE rate limits. Use an absolute path. |
email | operator_<id>@algo.xyz (mainnet) / @dotalgo.io (testnet) | Let's Encrypt account contact. Need not be deliverable; set it only to receive expiry notices. |
domains | auto-derived | A single name, from your NFD plus zs.nfd_record_name. An explicit list bypasses the label and is not cross-checked — a name outside your NFD's suffix fails at challenge time, not at startup. |
directory_url | LE production | Point at https://acme-staging-v02.api.letsencrypt.org/directory to rehearse the wiring without spending quota. |
propagation_delay | 30s | How long the solver waits after writing the TXT record before it starts polling. Don't shorten it — the record has to travel chain confirm → NFD indexer → authoritative zone, and polling early caches an NXDOMAIN against it. |
propagation_timeout | 6m | Overall deadline for the challenge to become visible. |
When TLS is on, bind listen to a public (non-loopback) address. See
Endpoints for the public-URL requirements, and
Reaching your node for the end-to-end walkthrough: name,
DNS record, certificate, base URL, and what each one costs.
server.tls.ip_sync
Opt-in background loop (default off) that detects the node's public IP and
rewrites the A/AAAA record on the operator's NFD when it changes, at the apex
or under zs.nfd_record_name. Steady state costs one HTTPS probe per active
family per tick and zero on-chain transactions; only real drift writes. See
Reaching your node for what it's for
and why it pairs with acme.
| Key | Default | Notes |
|---|---|---|
enabled | false | Master switch. |
ipv4 / ipv6 | auto-detect | Tri-state: omitted = auto, false = never publish, true = require (a sustained detect failure then WARNs). Setting both to false while enabled is a startup error. |
interval | 1m | Detection cadence. 1m is also the minimum. |
ttl | 1m | Published record TTL. Matches interval by default so resolvers don't cache a moved IP longer than the loop takes to notice. |
url.enabled | true when ip_sync.enabled | Also keeps this node's on-chain NodeRecord.baseUrl in sync, via updateNodeUrl signed by the node's signing key. The URL is derived once at startup as https://<nfd-derived host>:<port> — the port is always appended and comes only from server.listen, so the default publishes :9090. Set false when a CDN or reverse proxy fronts the node on a different port. |
These preconditions refuse to start the node rather than warn: ip_sync
requires tls.mode to be manual or acme, and algod.network to be testnet
or mainnet. url.enabled additionally requires tls.mode: acme plus
zs.escrow_app_id. Since it defaults to on, a manual-TLS node with
ip_sync enabled won't boot until you set url.enabled: false.
Once the node is running, every failure in the loop is a WARN, nothing gates serving, and the next tick retries.
Env: NODE_SERVER_TLS_IP_SYNC_{ENABLED,IPV4,IPV6,INTERVAL,TTL,URL_ENABLED}.
algod — Algorand connection
Selects the chain the node verifies payments and submits settlements against.
Setting network alone is enough: it picks a default endpoint (nodely.dev on a
public network, http://localhost:4001 for localnet) and token.
| Key | Default | Notes |
|---|---|---|
network | mainnet | mainnet | testnet | localnet. Omitting the whole block boots on mainnet with its canonical escrow app id. |
endpoint | "" | Optional override for a private/paid algod. Setting a raw endpoint with no network stays fully manual: no network default is stamped, so nothing gap-fills the escrow app id. |
token | "" | Optional override — prefer NODE_ALGOD_TOKEN. |
Your algod must be able to observe the mempool. Admission verifies the payer's
escrow open() via a pending-transaction lookup before it confirms. Public
RPC (Nodely / AlgoNode) works out of the box; a self-hosted algod must be
participating in consensus or have ForceFetchTransactions: true set.
logging
| Key | Default | Values |
|---|---|---|
level | info | debug | info | warn | error |
format | text | text | json |
llm — text inference backend
Picks the upstream that runs your text models. One provider per node. Omit
the whole llm: block for an image-only node.
llm.provider is one of:
| Provider | What it does |
|---|---|
openai_passthrough | Any OpenAI-compatible base URL (real OpenAI, vLLM, Ollama, aggregators such as OpenRouter, Anthropic's OpenAI-compat endpoint). Configure llm.openai.base_url + api_key. Aggregators need care when pricing — see Aggregator backends. |
lmstudio | Local LM Studio. base_url defaults to http://localhost:1234/v1; queries /api/v1/models for metadata. |
llamacpp | An externally-running llama-server. base_url defaults to http://127.0.0.1:8080/v1; queries /props for metadata. |
kronk | An externally-running Kronk server — a multi-model llama.cpp pool. base_url defaults to http://127.0.0.1:11435/v1; queries /v1/kronk/* for metadata and each model's provenance source, which its model ids can't supply: since Kronk 1.31.4 both model lists report <org>/<file stem> (a bare stem gets a 400), and neither form names the HuggingFace repo. |
local | The node supervises a llama-server child over loopback (recommended for self-hosted GPUs). See llm.local below. |
vertexai | Google Vertex AI's OpenAI-compatible API, authenticated via Application Default Credentials (no static key). See llm.vertexai. |
"" (empty) | No text inference. Combined with image_llm this is image-only mode; combined with zs.relay_only: true it's relay-only mode. |
Two route-level keys sit directly under llm:, not under llm.openai:. Each has
an image_llm.<key> twin that you set independently, since the two routes can
point at different upstreams.
| Key | Default | Notes |
|---|---|---|
upstream_zdr_declared | false | Assert a zero-retention agreement the node has no way to check (a first-party OpenAI or Azure contract, hardware you own). Advertises retention: operator_declared, the weakest tier. Ignored wherever the node already derives a tier. See Privacy & retention. |
allow_upstream_retention | false | Release the OpenRouter privacy pin so a model with no zero-retention endpoint becomes servable. It also drops data_collection: "deny", re-opening OpenRouter's Data Training settings, and the route then advertises no upstream_enforced tier. No effect on any other upstream, and it does not touch the xAI gate. Read the four consequences before setting it: Releasing the pin. |
llm.openai
Connection settings shared by openai_passthrough, lmstudio, llamacpp, and
kronk.
| Key | Default | Notes |
|---|---|---|
base_url | https://api.openai.com/v1 | Upstream API root. Include the version segment (usually /v1) yourself. For lmstudio/llamacpp/kronk, leave empty to take their loopback default. |
api_key | "" | Prefer NODE_LLM_OPENAI_API_KEY env. Optional for lmstudio/llamacpp/kronk — but for kronk, required once KRONK_AUTHORIZATION_MODE is anything but open, and it must be an admin token for the node to read model provenance. |
timeout | 0s | Keep at 0 for long streams. |
translate_responses_to_chat | provider-dependent | Emulate /v1/responses on top of /v1/chat/completions. Defaults true for llamacpp (llama-server has no Responses route), false for lmstudio, kronk, and openai_passthrough. Set true only for an upstream that lacks /v1/responses. Serving that route is mandatory: leaving this false against an upstream that 404s it is a startup failure — see Serving models. |
log_raw_usage | false | Emits a raw upstream usage INFO line carrying the upstream's usage object verbatim for every response, including cached-token fields the node may not yet interpret. Privacy-safe (counts only, no content), so it is not gated by tee.mode. Turn it on for a capture run to confirm whether your provider reports cached tokens, then off again. Env NODE_LLM_OPENAI_LOG_RAW_USAGE. |
debug_dump_errors | false | Dev only. On any non-2xx from the upstream, prints the request body the node sent plus the upstream's response body to stderr, both pretty-printed. It writes decrypted content, logs a loud startup WARN whenever on, and is a hard startup error under tee.mode != none. Never enable it on a node serving real payers — see Privacy & non-retention. Env NODE_LLM_OPENAI_DEBUG_DUMP_ERRORS. |
Aggregator backends
An aggregator like OpenRouter (base_url: https://openrouter.ai/api/v1) is not one upstream. It serves a single model id
from several provider endpoints at different prices and chooses one per
request, so two identical requests can cost you very different amounts.
Measured on google/gemini-3.7-flash, August 2026: six live endpoints spanning
$0.1875–$1.35 per 1M input tokens — a 7.2x spread — with two identical calls
landing on different endpoints at $0.00186 and $0.00924.
A ticket carries one signed rate, so if you price from a model's headline
figure, you are paid that rate but billed the endpoint's. Both public catalogs
quote the cheap end (models.dev and OpenRouter's own /v1/models both report
$0.375/M for that model), 3.6x below its dearest endpoint.
zs-node init reads each selected model's per-endpoint rate cards, prices from
the worst endpoint, and records the spread in the generated config. In a
hand-written config, price from the worst endpoint yourself or accept the loss.
Endpoint rotation has two more consequences:
- Cached-read discounts are intermittent. A prompt cache lives on the
endpoint that built it, so a request routed elsewhere reports
cached_tokens: 0and bills at your full input rate. Setcache_read_rateregardless; omitting it resolves to yourinput_rate, so cached reads bill at full price with nothing reported. - Some catalog entries can't be served and
zs-node initexcludes them::batchvariants (asynchronous batch endpoints with no live route — 60 of OpenRouter's 416 models) and meta-routers such asopenrouter/auto, which publish a negative price because their cost depends on what they route to.
llm.local
When provider: local, the node starts and restarts a llama-server child
process. Install llama-server once; the node supervises it (5 restart attempts
in 60s, then permanent failure). There is one model per node, so more than
one entry is a config error. There is no Hugging Face pull yet; point
model_path at a local GGUF.
| Key | Default | Notes |
|---|---|---|
binary_path | /usr/local/bin/llama-server | macOS dev: /opt/homebrew/bin/llama-server (brew install llama.cpp). |
startup_timeout | 60s | Wait for the child's health check; big GGUFs on cold storage need this. |
host / port | 127.0.0.1 / 0 | Loopback only; 0 = OS-assigned port. |
parallel_slots | 0 | Concurrent decode sessions; 0 derives from zs.max_active_tickets. |
models[] | — | Exactly one: id, model_path (GGUF), context_window, gpu_layers (0–99), threads, extra_args[]. |
KV-cache memory scales with context_window × parallel_slots, so raising
zs.max_active_tickets raises memory use. Do not put --parallel, -np,
--cont-batching, or -c in extra_args — those are derived and rejected at
startup.
llm.vertexai
When provider: vertexai, authenticate via Application Default Credentials
(Cloud Run / GKE Workload Identity / GCE SA, or gcloud auth application-default login for dev) — there is no api_key. Required:
project, location (e.g. us-central1), and models[] (each an
publisher-prefixed id like google/gemini-2.5-flash plus context_window).
timeout defaults to 0s — keep it there for streams. /v1/responses is
emulated automatically.
llm.health_check
Always on; this block only tunes cadence. The node probes the provider's
discovery endpoint. While the backend is up, /v1/zs/details advertises every
text zs.models entry and, when default_pricing is set, every model the
backend lists, subject to the
identity gate. A model served
through default_pricing alone drops out when the backend stops listing it.
While the backend is down, none of its text models are advertised and a reserve
for one returns 503 provider_unavailable. Image models follow the image
backend's own monitor, which runs on the same cadence.
| Key | Default | Notes |
|---|---|---|
interval | 30s | Probe cadence. |
timeout | 5s | Per-probe deadline (must be < interval). |
failure_threshold | 2 | Consecutive failures before the backend is marked down. |
local and vertexai derive their model list from config and are always
reported available.
image_llm — image backend (optional)
Independent of llm — run both, either, or just images. When unset, the node
404s the image routes and no model may declare an image: block.
| Key | Notes |
|---|---|
provider | "" | comfyui | comfyui_cloud | openai_passthrough (a hosted OpenAI-compatible image API such as xAI; see Image generation). |
comfyui.base_url | e.g. http://127.0.0.1:8000. |
comfyui.data_dir | Optional absolute path to the ComfyUI base dir so the node cleans up generated files. Override with NODE_IMAGE_LLM_COMFYUI_DATA_DIR. |
Each image-capable model in zs.models[] declares its own image: block
(backend, image_rate, max_n, defaults, ComfyUI template). See
Serving models & pricing for the
pricing model.
zs — operator identity, pricing, and behavior
Identity and concurrency
| Key | Default | Notes |
|---|---|---|
operator_id | 0 | Your on-chain operator id from registration (0 is reserved). Owner address + NFD are read from the operator box at startup. |
node_id | 0 | This node's per-operator id (0 reserved). Its signing address is read from the node box. |
max_active_tickets | 16 | Default per-model concurrency cap; reserve returns 429 no_capacity when a model hits its cap. Each model has its own independent pool — override per model under zs.models[].max_active_tickets. |
payer_slot_cap | 4 | How many of a model's max_active_tickets slots a single payer may hold at once, so one funded account can't take a model's whole pool. Over it, reserve returns 429 payer_slot_cap, a distinct code from no_capacity. Per-model override at zs.models[].payer_slot_cap; 0 disables the per-payer cap while leaving max_active_tickets in force. At the default of 4, one payer can hold half a pool of 8. File-only (no env override). |
inject_safety_identifier | false | Opt-in. Fills OpenAI's safety_identifier field from the verified payer, so abuse signals at the upstream attach to an account rather than to your whole key. Env NODE_ZS_INJECT_SAFETY_IDENTIFIER. |
owner_addr | "" | Optional override. Normally read from the operator box on chain; pin it only if you need to bypass the chain lookup. Env NODE_ZS_OWNER_ADDR. |
signing_addr | "" | Optional override. Normally read from the node box on chain; the node still refuses to start unless the matching mnemonic is loaded via the keystore. Env NODE_ZS_SIGNING_ADDR. |
The signing mnemonic itself is not a config key — it's provisioned through the keystore (see below and Encryption & keys).
NFD identity
Which NFD the node publishes DNS under, and where inside it. These keys only
matter when tls.mode is acme or ip_sync is on, but nfd_record_name is
validated on every boot regardless.
| Key | Default | Notes |
|---|---|---|
nfd_app_id | 0 | The NFD whose u.dns this node writes. Gap-filled from the operator box only when it is 0, so an explicit value always wins; set one to give a node its own NFD instead of sharing the operator's. Env NODE_ZS_NFD_APP_ID. |
nfd.api_url | per-network default | Override the NFDomains API endpoint. Only needed for tests or an unusual network path. Env NODE_ZS_NFD_API_URL. |
nfd_record_name | "" (apex) | Bare DNS label the node publishes under, inside nfd_app_id's u.dns. Lowercase letters, digits and internal hyphens; dotted labels (api.node1) allowed. Uppercase is rejected, not lowercased. A literal @ is rejected (leave it empty for the apex), and a value ending in .algo is rejected — that's an NFD name and belongs in nfd_app_id. Keep it short: it eats into the 248-byte base-URL ceiling. Env NODE_ZS_NFD_RECORD_NAME. |
Never let two nodes publish at the same record. Two nodes of one operator
both left at the apex overwrite each other's A record on every tick — a
permanent flap costing an on-chain transaction per tick per node, with clients
reaching whichever wrote last. Because each write rewrites the whole u.dns
document, racing writers can also drop each other's unrelated records, including
a live _acme-challenge TXT mid-issuance. Nothing detects this. Give every node
its own nfd_record_name, or its own nfd_app_id — see
One label per node.
Ticket lifetimes
| Key | Default | Notes |
|---|---|---|
ticket_ttl | 5s | Window between reserve and the inference POST — not inference duration. A POST after expiry gets 402 ticket_invalid. A long stream on a 5s-TTL ticket still completes normally. Kept tight so an abandoned reserve frees its slot fast. |
default_expires_after | 5m | Fallback settlement-complete deadline gating the contract's refund_inactive. Per-model override: expires_after. Keep it ≳ your p99 inference duration plus the ~30s settlement watchdog. |
Pricing
Rates are USD per 1,000,000 tokens (paste directly from a vendor price card) and
are net of the protocol fee — what you receive. The node grosses up the
escrowed max_price to cover the fee, so don't inflate your published rates.
default_pricing—{input_rate, output_rate}applied to any model without its own entry. An optionalcache_read_ratediscounts the cached-read input subset (see below).models[]— per-model entries.pricingoverrides the default; an omitted or emptypricing: {}inherits it; explicit{input_rate: 0, output_rate: 0}is a free model (never inherits).contextdeclares the model's capabilities and size limits (see below).max_active_ticketsoverrides the global pool.
models[].context
What the node advertises about a model, and the reserve-time ceiling it enforces. Every field is optional.
| Field | What it does |
|---|---|
context_window | Total tokens the model accepts, input plus output. Reserves are sized against it. Most backends publish it and the node reads it automatically — vLLM and SGLang via max_model_len, llama.cpp and LM Studio via their own metadata, Vertex from its model list. Declare it only for gateways that publish no metadata (OpenAI, xAI), or to enforce something tighter. On a self-hosted runtime, prefer leaving it out: a hand-declared value overrides the real --max-model-len your server is running with. |
max_output_tokens | The model's output ceiling. Must be below context_window. See Set an output ceiling — leaving it out caps long answers at a quarter of the context window (up to 32,768), and lets over-large requests through to your backend. |
input_modalities | Content types the model accepts: text, image, audio, video. Omitting implies text-only, which also suppresses feeding generated images back to the model. |
output_modalities | Content types it produces. Usually ["text"]. |
tool_use | Whether the model can call tools. Overrides discovery. |
reasoning | {supported, allowed_efforts, default_effort}. Overrides discovery; gates reasoning replay in the tool loop. |
max_input_images | Per-prompt cap on input images. The node trims the oldest rather than letting the request fail. Enforced locally, not advertised. |
tags | Free-form labels, merged with any from the model's HuggingFace card. |
Run zs-node doctor after editing this block — it re-probes the backend and
reports where your declarations and its actual behavior disagree.
Two more per-model blocks sit beside context and pricing:
source(effectively required) — the model's checkable identity, e.g.source: "hf:org/model". A model that satisfies neither this, anorg/model-shaped id, a backend-supplied source, nor the frontier whitelist is dropped from your catalog, with only a WARN in the log. A malformed value is a hard startup error. See Model identity.weights(optional) —{files, root, digest}, used to advertise a cryptographic digest of the bytes you serve for cross-operator comparison. You usually need none of it:provider: localandprovider: kronkproduce a stronger digest with no config at all. Declarefilesonly when fronting an external engine. See Model integrity — and note that a digest your peers disagree with costs you routing placement.
And two rate modifiers:
cache_read_rate(optional, inside anypricingblock) — the USD/1M rate for the cached-read input subset. Omit it and cached reads bill atinput_rate(no change);0makes them free. It only takes effect when the upstream reports a cached count — vLLM needs--enable-prompt-tokens-details; LM Studio / Ollama don't report one, so it's a no-op there. See Serving models & pricing.long_context(optional, inside anypricingblock) — a surcharge tier for upstreams (xAI/Grok) that charge a higher rate once a prompt reaches a size threshold.{threshold_tokens, input_rate, output_rate}are required (the high rates must be ≥ their base counterparts), plus an optional highcache_read_ratethat defaults to your base discount scaled by the input step-up. A request whose prompt reachesthreshold_tokensbills at the high rates for the whole request. With a tier set, raisecontext_windowto the model's true max instead of capping it at the threshold — the node re-checks the tier against the prompt actually sent when it bills, so an over-sized reserve doesn't over-charge. See Serving models & pricing.
Set either default_pricing or at least one models[] entry (or both).
Full detail — modalities, capability badges, image routes — is in
Serving models & pricing.
zs.coordinates
Model-taxonomy discovery, on by default. With
discover_huggingface: true the node fetches each served model's source repo
once near startup and uses it to enrich what clients display: the author's
curated card tags are unioned into the model's advertised tags, and
params / family / quantization are surfaced on the per-model drill-in.
zs:
coordinates:
discover_huggingface: true # default
The fetch is cached per repo for the process lifetime and fails open, so it
never blocks a model from being served. It reveals your served set to
HuggingFace, a set already public on /v1/zs/details. Set it to false only
for a node that must make no outbound calls to huggingface.co; that also turns
off the registry cross-check described in Model
integrity.
Confirm it is working with zs_hf_coordinates_total{result}. File-only (no env
override).
zs.min_charge
Per-request minimum charge. min_price is the max of two components; settlement
clamps amount_charged up to it on non-zero usage. Free models bypass it.
| Key | Default | Notes |
|---|---|---|
output_tokens | 1000 | Bill as if at least this many output tokens were produced, so a tiny response never settles ~0. 0 disables. |
algo_txns | 0 | µALGO network fees (in minTxnFee units of 1,000 µALGO) to recover via the oracle. Recommended 7 — the ~7,000 µALGO the operator absorbs per paid request on a live deployment (2 at open() + 5 at atomic settle()). The settle fee is sized to the payments that actually happen, so a free model costs 2 and a failed request 3. Auto-disabled on free models and while the oracle is down. |
zs.reserve
How much of a payer's USDC is locked per request, and whether your node checks that the prompt it receives matches what was reserved for it.
The reserve is sized from the actual request body plus a tool-loop headroom
term, not from the model's full context window, so a short question against a
1M-context model locks cents rather than dollars. The client computes it; your
node caps it at context_window − max_output, never raises it, and
re-measures the decrypted body against it at inference time.
| Key | Default | Notes |
|---|---|---|
enforce_input_budget | true | Reconcile the decrypted prompt against the reserved input_count. Over budget, a plain request is refused before any upstream call (sealed, zero-cost, refunded in full) and a tool loop is cut off at the boundary — still returning a final answer synthesized from the results already gathered. This is your protection against a payer that under-declares. A prompt whose reserve is already at the context_window − max_output cap is forwarded instead, as long as its size estimate is within three times the budget, and your backend decides whether it fits (see How the ceiling gets sized). false runs monitor mode: over-budget requests are metered and logged but still served, so you can size the real rate before enforcing. The node logs a startup WARN while it's off. |
input_budget_tolerance | 0.10 | Slack the measured input may exceed the reserve by before it counts as over budget — absorbs estimator jitter between the client's sizing and your re-measurement. Read clamped to [0, 1]. |
tool_headroom_per_iteration | 4000 | Per-iteration input-token headroom you advertise on /v1/zs/details. Advisory — your node never adds it and never recomputes it; clients read it and size their own reserve. |
Env overrides: NODE_ZS_RESERVE_ENFORCE_INPUT_BUDGET,
NODE_ZS_RESERVE_INPUT_BUDGET_TOLERANCE,
NODE_ZS_RESERVE_TOOL_HEADROOM_PER_ITERATION.
A plain turn is never falsely rejected: the client sizes from the same bound
your node re-measures with, so the two agree within tolerance. Clients add the
headroom term only when the request's tools can grow the context your node
measures — your zs_* built-in tools, or a frontier
model's server-side tools. Tools the caller executes itself (a coding agent's
shell and edit tools, client-side MCP) add none, so reserves from agent-style
clients are legitimately much smaller.
Sizing tool_headroom_per_iteration. Derive it from what your own tools
actually return. zs.builtin_tools.web_read.max_bytes defaults to 64 KiB of
markdown — roughly 16,000 tokens for one read, about 4× the default
per-iteration headroom. At the shipped defaults the aggregate
(max_iterations × tool_headroom_per_iteration) absorbs that, but if you lower
zs.builtin_tools.max_iterations without raising this value, tool loops get cut
short. Watch zs_reserve_input_budget_over_total{outcome="tool_cutoff"} — see
Monitoring & metrics.
zs.oracle
ALGO/USD price feed. Used only so Reserve can express ALGO network fees in
USD — it is not in the inference-pricing path, and Reserve tolerates an
unhealthy oracle (it just omits algo_usd_price).
| Key | Default | Notes |
|---|---|---|
source | coingecko | Only CoinGecko is supported today. |
refresh_interval | 30s | Fetch cadence. |
max_staleness | 5m | Past this, Reserve omits the price. |
min_algo_usd | 0.01 | Reject quotes below this as unreliable. |
http_timeout | 10s | Per-fetch deadline. |
zs.escrow_app_id and mempool polling
The deployed ZeroSignalEscrow app id for your network.
- On a public network you may omit it — the node falls back to the canonical
embedded app id: testnet
765860477, mainnet3628061142. An explicit value always wins. Localnet has no embedded default — set it there. escrow_app_id: 0is a test-isolation hook only (short-circuits payment verification); a0-mode node is undiscoverable in production and can't claim USDC.mempool_poll_timeout(5s) andmempool_poll_interval(100ms) bound how long the node waits for theopen()group to appear in algod's pending pool. The interval is latency your callers feel: the first check is immediate, so whenever the payment hasn't propagated yet, a full interval elapses before inference starts. Raising it only reduces algod calls.
zs.rate_limits
Public-facing throttling, on by default with the values below. A partial
block keeps the defaults for the buckets you don't mention. Each bucket is a
token bucket; setting its rps or burst to 0 disables that one bucket, and
enabled: false switches the whole subsystem off (for private deployments
behind their own edge). /healthz, /livez and /metrics are always exempt.
| Key | Default | Scope |
|---|---|---|
enabled | true | Master switch. false leaves max_active_tickets as the only admission throttle. |
| Bucket | Default rps / burst | Scope |
|---|---|---|
reserve_per_ip | 2 / 5 | /v1/zs/reserve per IP. |
reserve_per_account | 5 / 10 | Per Algorand payer address (escrow enabled). |
llm_per_ip | 5 / 20 | Chat + responses per IP. |
discovery_per_ip | 20 / 40 | /v1/models, /v1/zs/details, etc. |
relay_per_ip | 10 / 20 | /v1/zs/relay forwards per IP (the transport-privacy hops you carry for others). |
attestation_nonce | 1 / 5 | /v1/zs/attestation?nonce=…, one bucket for the whole node. Confidential mode only. |
Plus idle_eviction (30m) and max_keys (100000). Client IP is read from
CF-Connecting-IP > X-Real-IP > rightmost X-Forwarded-For > TCP peer.
attestation_nonce limits freshness challenges,
each of which makes the enclave mint a new quote. It is not per IP because
challenges arrive through relays, so the IP you see is a relay's. Separately,
the node mints at most two challenged quotes at once, and that cap stays in
force with rate limits switched off. Over either limit, the caller gets
429 attestation_rate_limited. All requests to /v1/zs/attestation also count
against discovery_per_ip; requests without a nonce count only against that.
zs.builtin_tools
In-loop tools the node advertises on /v1/zs/details and runs locally,
aggregating all rounds into one receipt. See Built-in tools
for the tool catalog and how they run.
| Key | Default | Notes |
|---|---|---|
enabled | true | Master switch. false disables the whole subsystem (including zs_get_time). |
max_iterations | 20 | Absolute ceiling on chat→tool→chat rounds (clamped [1,20]). Applied when the key is omitted; the shipped example sets it to 5. Advertised to clients as max_tool_iterations. |
max_stalled_iterations | 3 | Consecutive repeated-call rounds tolerated before the loop stops (clamped [1,10]). Node-internal; not advertised. |
web_search.enabled | true | Per-tool toggle for zs_web_search. |
web_search.max_results | 10 | DuckDuckGo HTML; 1..25. |
web_search.safe_search | moderate | strict | moderate | off — an unknown value is a hard startup error. off returns unfiltered results; set it only if that's your intent. |
web_search.timeout | 15s | Per-search deadline. |
web_read.enabled | true | Per-tool toggle for zs_web_read. |
web_read.max_bytes | 65536 | Truncation ceiling on the markdown handed to the model (keeps fetched pages within its context). |
web_read.max_download | 4194304 | Ceiling on the raw page fetched off the wire, before conversion. Over this, the read is refused, not truncated. Must be >= max_bytes (compared on effective values, so an explicit 0 still means "the default"); a smaller value is a hard startup error. Not a memory budget: converting a page peaks at roughly 20–70× the page's size, so raising this raises RAM use by a large multiple. 4 MiB is ~2× the largest real-world article observed. |
web_read.max_concurrent | 4 | How many pages the node converts at the same time. Peak memory ≈ max_download × 20–70 × this, so it's the setting to lower under a memory limit. Lowering it adds latency, because reads queue; lowering max_download instead would make large pages permanently unreadable. 4 suits a node that owns its box; 1 is reasonable in a small container. See Sizing the node process. |
web_read.timeout | 15s | Per-fetch deadline. |
web_read.allow_private_targets | false | Localnet/dev only — leave off in production. Disables the guard that refuses to fetch a URL resolving to a private / loopback / link-local / carrier-NAT address. Your node does its own fetching, so with this on, a crafted URL can probe your LAN or cloud metadata endpoint. Logs a startup WARN when enabled. The web_read counterpart of zs.allow_private_relay_targets. |
favicons.enabled | true | Fetches each search-result site's favicon and inlines the bytes on the round's status frame, so chat clients show real site marks on the source chips. Fires only on the streaming Responses path. Turn it off on a node with metered or locked-down egress; clients then show a generic link glyph. |
favicons.timeout | 1.2s | Bounds one tool call's whole fetch batch, not one site and not one turn: each call with hosts not already cached pays it again. A host already sent on the stream is skipped. |
favicons.max_bytes | 16384 | Caps one icon. Oversize is dropped, not truncated. The default is the wire-format maximum; larger values are clamped to it. |
favicons.cache_max_bytes | 4194304 | In-memory host→icon cache size. |
favicons.cache_ttl | 24h | How long a resolved icon stays fresh. A miss is re-probed after 1/24th of it, and a fetch that merely timed out isn't cached at all. |
favicons.allow_private_targets | false | Localnet/dev only. Same guard and same warning as web_read.allow_private_targets. |
image_search.* | — | Same shape (enabled, max_results 1..100, safe_search, timeout) but force-disabled — see below. |
Why the node fetches favicons instead of the user's browser, and what egress that adds, is in Source favicons.
image_search is force-disabled in code — its DuckDuckGo image backend is
broken, so the node won't advertise or run it regardless of what you set here
until a working provider is wired up.
zs.relay behavior
Every node ships the /v1/zs/relay route, registered whenever a relay
directory is available (escrow + algod configured). See Relays.
relay_only: true(orNODE_ZS_RELAY_ONLY=true) — pure relay: text and image providers empty, nozs.models, a non-zeroescrow_app_idrequired. Prompt routes 404; only relay and discovery GETs are served.allow_private_relay_targets(defaultfalse) — SSRF guard. The relay refuses to dial private/loopback/link-local/CGNAT/ULA targets so a malicious operator can't turn your relay into a LAN prober. Settrueonly for localnet/dev.
zs.self_eviction
Watchdog (on by default) that re-reads this node's own operator box and exits
non-zero if the operator is removed on-chain (admin-evicted or unregistered) so a
supervisor surfaces it. Requires escrow_app_id + operator_id.
| Key | Default | Notes |
|---|---|---|
enabled | true | Disable with false. |
interval | 5m | Re-check cadence. |
threshold | 3 | Consecutive missing reads before exiting (rides out a reorg). |
Env: NODE_ZS_SELF_EVICTION_{ENABLED,INTERVAL,THRESHOLD}.
Evicting a misbehaving operator is an admin moderation action
(evictOperator); there is no permissionless eviction sweep, and nodes don't run
one. self_eviction only notices that this node was removed and shuts it down
cleanly.
zs.signing_balance
Background poller that reads your hot signing address's ALGO balance so you can
alert on it before on-chain transactions start failing. That account fee-pays
every settlement and, when URL sync is on, the updateNodeUrl call. The poller
publishes the spendable balance as zs_signing_balance_algos on /metrics,
logs a WARN when the spendable balance (balance minus Algorand's
minimum balance) drops below the threshold, and logs an ERROR when it falls
under 5 ALGO — the floor below which proxies and the web client stop sending
requests to your node; see Monitoring.
The NFD u.dns writes that ACME and IP sync make are fee-paid by whichever
account owns the NFD, a different account unless you made them the same —
see Which account signs the DNS writes.
| Key | Default | Notes |
|---|---|---|
poll_interval | 5m | How often to read the balance. Set 0 to disable the poller, which also turns off both log lines; the node warns at startup if you do. |
warn_below_algos | 10 | Emit a low-balance WARN when the spendable balance drops below this many ALGO. Keep it above 5, the routing floor, or the warning arrives only once traffic has stopped — the node warns at startup if it is 5 or lower. Advisory — independent of the hard 1 ALGO boot floor the node enforces at startup. |
The gap between the default 10 and the 5 ALGO floor covers about 700 paid
requests. On a busy node that can run out between two polls, so the WARN and
the ERROR land together. To get the WARN at least one poll ahead of the floor,
set warn_below_algos to at least
5 + (paid requests per second × poll_interval in seconds × 0.007) — about
5 + 2.1 × requests per second at the default 5m — plus enough to cover the
time it takes you to refill.
The poller runs for relay-only nodes too, since they still fee-pay
updateNodeUrl when URL sync is on. A relay-only node never receives inference
requests, so the 5 ALGO floor doesn't apply to it and the node doesn't log the
routing-floor ERROR; you can lower warn_below_algos to whatever float its
ACME and IP-sync writes need.
zs.settlement_db_path
Where the node records served tickets while the background settlement driver
drives each through on-chain escrow.settle. On startup, leftover settling
entries are reconciled against algod.
| Value | Behavior |
|---|---|
<path> | Durable SQLite (WAL). Recommended for production. |
:memory: | Ephemeral SQLite (dev only). |
"" | In-memory store, non-durable (dev only; the contract's refund_inactive backstop still protects payer funds). |
Related:
settlement_watchdog_seconds(30) — how long the driver waits after a request completes before considering a standalone settle. Normally the client broadcasts the atomic two-transaction group the node pre-signed, which finalizes the ticket. Set this too short and the driver pre-empts that group: payers are unaffected, but you pay network fees for a standalone settle that wasn't needed. Raise it if your clients' acknowledgement latency is consistently above the window; lower it only if a metric says so.-1fires immediately (tests only).settlement_lapse_grace_seconds(300) — node-side wait before it force-finalizes a claim the client never acknowledged; must be ≥ the contract's grace default.settlement_retention(720h, i.e. 30 days) andsettlement_retention_interval— how long finalized ledger rows are kept before the node auto-purges them, and how often it sweeps.
See The payment flow for the full settlement lifecycle.
tee — confidential mode (opt-in)
A non-none tee.mode advertises the node as TEE-capable on /v1/zs/details,
exposes /v1/zs/attestation, and makes the proxy verify attestation before
routing prompts here. tee.dataflow declares where plaintext comes to rest.
The trust model, hardware requirements, what each posture requires of your
backends, and the in-CVM identity bootstrap are in
Confidential compute (TEE).
| Key | Default | Notes |
|---|---|---|
mode | none | none | stub | dstack-tdx | nvidia-cc-tdx | nvidia-cc-snp. dstack-tdx is the mode that runs in production; stub is laptop-only untrusted evidence; the nvidia-cc-* modes fail the startup check with a "not yet implemented" error until GPU attestation lands. |
dataflow | sealed_local | sealed_local | attested_passthrough. Decides the retention tier payers see; see Two postures. No env override, because it has to be part of the measured config. |
attestation.nras_url | https://nras.attestation.nvidia.com | NVIDIA Remote Attestation Service — nvidia-cc-* only. dstack-tdx mints its quote from the local guest agent and never calls it. |
attestation.refresh_interval | 1h | /v1/zs/attestation serves 503 once stale past 2× this. Evidence is also re-minted on every ephemeral key rotation, independent of this cadence. |
attestation.evidence_cache_path | — | Disk path on a CVM-encrypted volume so a brief bounce doesn't drop you from the verified set. |
attestation.pccs_url | https://pccs.phala.network | Where a dstack-tdx node fetches Intel's TCB and revocation collateral to carry alongside its quote. Change it only for a PCCS inside your own network; see Collateral your node carries. |
Env: NODE_TEE_MODE, NODE_TEE_ATTESTATION_{NRAS_URL,REFRESH_INTERVAL,EVIDENCE_CACHE_PATH,PCCS_URL}.
A minimal config.yaml
A working passthrough node: one model, an OpenAI-compatible upstream, and the canonical escrow app id for its network.
server:
listen: ":9090" # bind all interfaces (dual-stack); override of the loopback default
private_listen: "127.0.0.1:9091"
write_timeout: "0s" # keep 0 for streams (also the default)
# No algod block: mainnet is the default, and it carries the canonical
# escrow app id. Set algod.network only to target a different network.
logging:
level: "info"
format: "text"
llm:
provider: "openai_passthrough"
openai:
base_url: "https://api.openai.com/v1"
# api_key comes from NODE_LLM_OPENAI_API_KEY — never commit it
timeout: "0s" # keep 0 for streams
zs:
operator_id: 42 # your id from registration
node_id: 1 # this node's id
models:
gpt-5.4-mini:
# No `source:` — a frontier id clears the identity gate on its own. A
# self-hosted model would need `source: "hf:org/model"`; see below.
pricing:
input_rate: 0.31 # USD per 1M input tokens, net of the protocol fee
output_rate: 2.50 # USD per 1M output tokens
# cache_read_rate: 0.08 # optional cached-read discount (no-op unless
# the upstream reports a cached count)
min_charge:
output_tokens: 1000
algo_txns: 7 # recover the ~7,000 µALGO network fees per paid request (live deployment)
settlement_db_path: "./settlement.db" # dev value; for the systemd unit use /var/lib/zs-node/settlement.db (see installation.md)
Every model needs a checkable identity, and a model without one is dropped
with only a WARN. An always-on gate drops any model that has neither a source
nor a frontier id from the advertised catalog. The node starts fine and /v1/models just comes back
short. A frontier id (gpt-5.4-mini, claude-*, gemini-*, grok-*) clears the gate on the id
alone. Recognized provider-qualified spellings clear it too — for example,
x-ai/grok-4.5 and google/gemini-2.5-pro fold onto their curated bare-id
entries. This does not whitelist arbitrary org/model strings. A declared
org/model id derives its own source, and Kronk supplies one per model (see
Model identity). Anything else,
such as a bare GGUF stem or a custom id, needs an explicit
source: "hf:org/model". It must be the hf:-prefixed repo ref, and a
malformed one is a startup error, not a silent drop.
default_pricing on its own is not a substitute: upstream-discovered ids
carry no source unless the backend supplies one, as Kronk does, so from any
other backend only frontier-whitelisted ones survive. Declare what you
intend to serve, or let zs-node init write the
catalog for you — then zs-node doctor reports models pass the provenance gate when it's right.
Provide secrets through the environment, not the file:
export OPERATOR_SIGNING_MNEMONIC="word1 word2 ... word25"
export NODE_LLM_OPENAI_API_KEY="sk-..."
The signing mnemonic must reach the node at startup — it refuses to start
without the mnemonic for its node's signing address. Any environment variable
ending in _MNEMONIC is picked up (the label is informational; lookup is by
derived address), or use a cloud secret manager via ZS_MNEMONIC_URLS
(comma-separated name=url pairs). See
Encryption & keys for the full
keystore options.
For the exact install paths, Docker run lines, and reverse-proxy setup, see
Installation. For registering the operator and getting your
operator_id / node_id, see Registering on-chain.