Skip to main content

Configuration

Operator and escrow settings live under zs:, with NODE_ZS_* environment overrides. For the reasoning behind pricing and the model catalog, see Serving models & pricing; for the signing mnemonic, see Encryption & keys.

Generate it: zs-node init​

Before writing this file by hand, try the wizard. Point it at your inference backend and it produces a complete, commented config.yaml:

zs-node init --base-url=http://127.0.0.1:8080/v1

It queries the backend:

  • It identifies the runtime (vLLM, SGLang, llama.cpp, LM Studio, Ollama, or a hosted gateway), sets llm.provider to match, and finds the real API root if you gave it the wrong one.
  • It checks whether /v1/responses exists, which decides llm.openai.translate_responses_to_chat, and whether streaming reports token usage. A backend that doesn't is rejected, because the node bills streaming work from that usage and refuses to start without it.
  • It probes each model for tool calling, reasoning, image input, per-prompt image limits, and its output ceiling.
  • It looks up list pricing in the models.dev catalog, applies your margin, and checks your on-chain operator and node records, your keystore, your USDC opt-in, and your signing balance.
info

The probe sends real requests, one to sixteen output tokens each, so it learns what your deployment does. It catches deployment gaps, such as a vLLM started without --enable-auto-tool-choice, which cannot make tool calls whatever the model card says. Against a local backend it just runs; against a metered endpoint it asks first. Use --probe=passive to read metadata only.

The wizard loads the generated file exactly as the daemon would and refuses to write anything that wouldn't start. Every value is annotated with where it came from, and anything surprising becomes a note in the file header.

# Re-run against an existing config; your own rates and settings are kept.
zs-node init --from=config.yaml --dry-run

# A paid gateway, two models, resold at a 25% margin.
zs-node init --base-url=https://api.z.ai/api/paas/v4 --models=glm-4.6 --margin=25

# Unattended.
zs-node init --base-url=http://vllm:8000/v1 --network=mainnet \
--operator-id=7 --node-id=1 --non-interactive --yes

It never writes an API key or a mnemonic into the file, never submits a transaction, and always keeps a .bak copy before overwriting. Run zs-node init -h for every flag.

How config is loaded​

  • File. The node reads ./config.yaml by default. Point it elsewhere with the --config <path> flag or the NODE_CONFIG environment variable.
  • Environment overrides. A curated set of settings can be overridden by an environment variable with the NODE_ prefix following the YAML path — e.g. NODE_SERVER_LISTEN overrides server.listen, NODE_ZS_SELF_EVICTION_ENABLED overrides zs.self_eviction.enabled. Not every key has one: several timeouts, the oracle, min_charge, and the rate-limit buckets are file-only. Where an override exists, the environment wins over the file.
  • Secrets come from the environment, never the file. API keys (NODE_LLM_OPENAI_API_KEY), the algod token (NODE_ALGOD_TOKEN), and the signing mnemonic (OPERATOR_SIGNING_MNEMONIC / ZS_MNEMONIC_URLS) are read from the environment or a secret manager. Don't commit them to config.yaml.
warning

Three timeouts must stay 0 for streaming. server.write_timeout, llm.openai.timeout, and llm.vertexai.timeout all default to 0s (disabled). Inference responses are long-lived SSE streams, and a non-zero value cuts a generation off mid-stream.

server — listeners, timeouts, TLS​

KeyDefaultWhat it controls
listen127.0.0.1:9090Public listener (prompts, reserve, discovery). A serving node binds all interfaces with :9090 (dual-stack; 0.0.0.0 is IPv4-only), which zs-node init writes; see The two listeners. Override with NODE_SERVER_LISTEN.
private_listen127.0.0.1:9091Ops listener for /healthz, /livez and /metrics, off the public port and exempt from rate limiting. Set "" to colocate them on listen. Override with NODE_SERVER_PRIVATE_LISTEN.
read_timeout30sRequest-read deadline.
write_timeout0sKeep at 0 — disabled, required for streams.
idle_timeout120sKeep-alive idle deadline.
sse_keepalive_interval15sSends a : zs-keepalive SSE comment whenever a streaming response produces no bytes (a long prefill before the first token, or a gap between tokens), so a reverse proxy or relay in front of you doesn't reset the idle stream at its own timeout. SSE clients ignore the comment, and it never touches billing. Set it well below your front door's idle timeout; 0 disables. See Clients get 502 or 504 errors. Override with NODE_SERVER_SSE_KEEPALIVE_INTERVAL.
drain_grace20sOn SIGTERM, how long reservations issued before the shutdown are still honored, so a request already in flight lands. New reservations are refused immediately regardless. Override with NODE_SERVER_DRAIN_GRACE.
drain_timeout5mHow long shutdown waits for inference already running to finish. Size it for your worst-case single response, but no longer than any per-request timeout in front of the node, which would cut the stream anyway. 0 opts out of draining entirely and also disables drain_grace (the grace hold runs inside the same wait). Override with NODE_SERVER_DRAIN_TIMEOUT.
shutdown_timeout30sFinal connection-close budget after inference has drained. 0 is treated as 30s, not an opt-out, so a connection that never closes can't hang the process past your supervisor's deadline. Override with NODE_SERVER_SHUTDOWN_TIMEOUT.
warning

Your supervisor's kill deadline must exceed drain_grace + drain_timeout + shutdown_timeout (5m50s with the defaults) or it force-kills the node mid-drain. See Graceful shutdown.

The private listener is never TLS-wrapped. In YAML, quote a leading-colon value like ":9091". The routes each listener serves are in the Endpoint reference.

server.tls​

TLS for the public listener. Three modes:

ModeWhat it does
off (default)Plain HTTP (loopback-only default). Prefer acme to serve HTTPS directly; or terminate TLS at a single reverse proxy in front of one node.
manualLoad cert_path + key_path from disk at startup; rotate by replacing the files and restarting.
acmeProvision and renew a Let's Encrypt certificate automatically (detailed below).

cert_path and key_path are both required when mode: manual, and are read once at startup — there is no hot reload.

The acme mode provisions and renews a Let's Encrypt certificate in the background via DNS-01, solved against your operator's NFD u.dns (the only supported DNS provider). Prerequisites, all enforced at boot:

  • algod.network ∈ {testnet, mainnet} — localnet has no public DNS suffix.
  • A non-zero NFD app id, from zs.nfd_app_id or gap-filled from the operator box when zs.escrow_app_id is set. Otherwise the node exits with tls.mode=acme requires the operator's NFD on chain — set zs.escrow_app_id and ensure the operator is registered with an NFD, or set tls.mode=off.
  • The keystore holds the key for whichever account currently owns that NFD. That is not automatically your node's signing address, and getting it wrong fails at runtime rather than at boot — see Which account signs the DNS writes.
KeyDefaultNotes
cache_dir./tls-cacheRelative to the process working directory. Holds the Let's Encrypt account key as well as issued certs, so it must be writable and durable — a container without a volume or a shifting WorkingDirectory re-registers the account every boot and eventually hits LE rate limits. Use an absolute path.
emailoperator_<id>@algo.xyz (mainnet) / @dotalgo.io (testnet)Let's Encrypt account contact. Need not be deliverable; set it only to receive expiry notices.
domainsauto-derivedA single name, from your NFD plus zs.nfd_record_name. An explicit list bypasses the label and is not cross-checked — a name outside your NFD's suffix fails at challenge time, not at startup.
directory_urlLE productionPoint at https://acme-staging-v02.api.letsencrypt.org/directory to rehearse the wiring without spending quota.
propagation_delay30sHow long the solver waits after writing the TXT record before it starts polling. Don't shorten it — the record has to travel chain confirm → NFD indexer → authoritative zone, and polling early caches an NXDOMAIN against it.
propagation_timeout6mOverall deadline for the challenge to become visible.

When TLS is on, bind listen to a public (non-loopback) address. See Endpoints for the public-URL requirements, and Reaching your node for the end-to-end walkthrough: name, DNS record, certificate, base URL, and what each one costs.

server.tls.ip_sync​

Opt-in background loop (default off) that detects the node's public IP and rewrites the A/AAAA record on the operator's NFD when it changes, at the apex or under zs.nfd_record_name. Steady state costs one HTTPS probe per active family per tick and zero on-chain transactions; only real drift writes. See Reaching your node for what it's for and why it pairs with acme.

KeyDefaultNotes
enabledfalseMaster switch.
ipv4 / ipv6auto-detectTri-state: omitted = auto, false = never publish, true = require (a sustained detect failure then WARNs). Setting both to false while enabled is a startup error.
interval1mDetection cadence. 1m is also the minimum.
ttl1mPublished record TTL. Matches interval by default so resolvers don't cache a moved IP longer than the loop takes to notice.
url.enabledtrue when ip_sync.enabledAlso keeps this node's on-chain NodeRecord.baseUrl in sync, via updateNodeUrl signed by the node's signing key. The URL is derived once at startup as https://<nfd-derived host>:<port> — the port is always appended and comes only from server.listen, so the default publishes :9090. Set false when a CDN or reverse proxy fronts the node on a different port.
warning

These preconditions refuse to start the node rather than warn: ip_sync requires tls.mode to be manual or acme, and algod.network to be testnet or mainnet. url.enabled additionally requires tls.mode: acme plus zs.escrow_app_id. Since it defaults to on, a manual-TLS node with ip_sync enabled won't boot until you set url.enabled: false.

Once the node is running, every failure in the loop is a WARN, nothing gates serving, and the next tick retries.

Env: NODE_SERVER_TLS_IP_SYNC_{ENABLED,IPV4,IPV6,INTERVAL,TTL,URL_ENABLED}.

algod — Algorand connection​

Selects the chain the node verifies payments and submits settlements against. Setting network alone is enough: it picks a default endpoint (nodely.dev on a public network, http://localhost:4001 for localnet) and token.

KeyDefaultNotes
networkmainnetmainnet | testnet | localnet. Omitting the whole block boots on mainnet with its canonical escrow app id.
endpoint""Optional override for a private/paid algod. Setting a raw endpoint with no network stays fully manual: no network default is stamped, so nothing gap-fills the escrow app id.
token""Optional override — prefer NODE_ALGOD_TOKEN.
info

Your algod must be able to observe the mempool. Admission verifies the payer's escrow open() via a pending-transaction lookup before it confirms. Public RPC (Nodely / AlgoNode) works out of the box; a self-hosted algod must be participating in consensus or have ForceFetchTransactions: true set.

logging​

KeyDefaultValues
levelinfodebug | info | warn | error
formattexttext | json

llm — text inference backend​

Picks the upstream that runs your text models. One provider per node. Omit the whole llm: block for an image-only node.

llm.provider is one of:

ProviderWhat it does
openai_passthroughAny OpenAI-compatible base URL (real OpenAI, vLLM, Ollama, aggregators such as OpenRouter, Anthropic's OpenAI-compat endpoint). Configure llm.openai.base_url + api_key. Aggregators need care when pricing — see Aggregator backends.
lmstudioLocal LM Studio. base_url defaults to http://localhost:1234/v1; queries /api/v1/models for metadata.
llamacppAn externally-running llama-server. base_url defaults to http://127.0.0.1:8080/v1; queries /props for metadata.
kronkAn externally-running Kronk server — a multi-model llama.cpp pool. base_url defaults to http://127.0.0.1:11435/v1; queries /v1/kronk/* for metadata and each model's provenance source, which its model ids can't supply: since Kronk 1.31.4 both model lists report <org>/<file stem> (a bare stem gets a 400), and neither form names the HuggingFace repo.
localThe node supervises a llama-server child over loopback (recommended for self-hosted GPUs). See llm.local below.
vertexaiGoogle Vertex AI's OpenAI-compatible API, authenticated via Application Default Credentials (no static key). See llm.vertexai.
"" (empty)No text inference. Combined with image_llm this is image-only mode; combined with zs.relay_only: true it's relay-only mode.

Two route-level keys sit directly under llm:, not under llm.openai:. Each has an image_llm.<key> twin that you set independently, since the two routes can point at different upstreams.

KeyDefaultNotes
upstream_zdr_declaredfalseAssert a zero-retention agreement the node has no way to check (a first-party OpenAI or Azure contract, hardware you own). Advertises retention: operator_declared, the weakest tier. Ignored wherever the node already derives a tier. See Privacy & retention.
allow_upstream_retentionfalseRelease the OpenRouter privacy pin so a model with no zero-retention endpoint becomes servable. It also drops data_collection: "deny", re-opening OpenRouter's Data Training settings, and the route then advertises no upstream_enforced tier. No effect on any other upstream, and it does not touch the xAI gate. Read the four consequences before setting it: Releasing the pin.

llm.openai​

Connection settings shared by openai_passthrough, lmstudio, llamacpp, and kronk.

KeyDefaultNotes
base_urlhttps://api.openai.com/v1Upstream API root. Include the version segment (usually /v1) yourself. For lmstudio/llamacpp/kronk, leave empty to take their loopback default.
api_key""Prefer NODE_LLM_OPENAI_API_KEY env. Optional for lmstudio/llamacpp/kronk — but for kronk, required once KRONK_AUTHORIZATION_MODE is anything but open, and it must be an admin token for the node to read model provenance.
timeout0sKeep at 0 for long streams.
translate_responses_to_chatprovider-dependentEmulate /v1/responses on top of /v1/chat/completions. Defaults true for llamacpp (llama-server has no Responses route), false for lmstudio, kronk, and openai_passthrough. Set true only for an upstream that lacks /v1/responses. Serving that route is mandatory: leaving this false against an upstream that 404s it is a startup failure — see Serving models.
log_raw_usagefalseEmits a raw upstream usage INFO line carrying the upstream's usage object verbatim for every response, including cached-token fields the node may not yet interpret. Privacy-safe (counts only, no content), so it is not gated by tee.mode. Turn it on for a capture run to confirm whether your provider reports cached tokens, then off again. Env NODE_LLM_OPENAI_LOG_RAW_USAGE.
debug_dump_errorsfalseDev only. On any non-2xx from the upstream, prints the request body the node sent plus the upstream's response body to stderr, both pretty-printed. It writes decrypted content, logs a loud startup WARN whenever on, and is a hard startup error under tee.mode != none. Never enable it on a node serving real payers — see Privacy & non-retention. Env NODE_LLM_OPENAI_DEBUG_DUMP_ERRORS.

Aggregator backends​

An aggregator like OpenRouter (base_url: https://openrouter.ai/api/v1) is not one upstream. It serves a single model id from several provider endpoints at different prices and chooses one per request, so two identical requests can cost you very different amounts.

Measured on google/gemini-3.7-flash, August 2026: six live endpoints spanning $0.1875–$1.35 per 1M input tokens — a 7.2x spread — with two identical calls landing on different endpoints at $0.00186 and $0.00924.

A ticket carries one signed rate, so if you price from a model's headline figure, you are paid that rate but billed the endpoint's. Both public catalogs quote the cheap end (models.dev and OpenRouter's own /v1/models both report $0.375/M for that model), 3.6x below its dearest endpoint.

zs-node init reads each selected model's per-endpoint rate cards, prices from the worst endpoint, and records the spread in the generated config. In a hand-written config, price from the worst endpoint yourself or accept the loss.

Endpoint rotation has two more consequences:

  • Cached-read discounts are intermittent. A prompt cache lives on the endpoint that built it, so a request routed elsewhere reports cached_tokens: 0 and bills at your full input rate. Set cache_read_rate regardless; omitting it resolves to your input_rate, so cached reads bill at full price with nothing reported.
  • Some catalog entries can't be served and zs-node init excludes them: :batch variants (asynchronous batch endpoints with no live route — 60 of OpenRouter's 416 models) and meta-routers such as openrouter/auto, which publish a negative price because their cost depends on what they route to.

llm.local​

When provider: local, the node starts and restarts a llama-server child process. Install llama-server once; the node supervises it (5 restart attempts in 60s, then permanent failure). There is one model per node, so more than one entry is a config error. There is no Hugging Face pull yet; point model_path at a local GGUF.

KeyDefaultNotes
binary_path/usr/local/bin/llama-servermacOS dev: /opt/homebrew/bin/llama-server (brew install llama.cpp).
startup_timeout60sWait for the child's health check; big GGUFs on cold storage need this.
host / port127.0.0.1 / 0Loopback only; 0 = OS-assigned port.
parallel_slots0Concurrent decode sessions; 0 derives from zs.max_active_tickets.
models[]—Exactly one: id, model_path (GGUF), context_window, gpu_layers (0–99), threads, extra_args[].

KV-cache memory scales with context_window × parallel_slots, so raising zs.max_active_tickets raises memory use. Do not put --parallel, -np, --cont-batching, or -c in extra_args — those are derived and rejected at startup.

llm.vertexai​

When provider: vertexai, authenticate via Application Default Credentials (Cloud Run / GKE Workload Identity / GCE SA, or gcloud auth application-default login for dev) — there is no api_key. Required: project, location (e.g. us-central1), and models[] (each an publisher-prefixed id like google/gemini-2.5-flash plus context_window). timeout defaults to 0s — keep it there for streams. /v1/responses is emulated automatically.

llm.health_check​

Always on; this block only tunes cadence. The node probes the provider's discovery endpoint. While the backend is up, /v1/zs/details advertises every text zs.models entry and, when default_pricing is set, every model the backend lists, subject to the identity gate. A model served through default_pricing alone drops out when the backend stops listing it. While the backend is down, none of its text models are advertised and a reserve for one returns 503 provider_unavailable. Image models follow the image backend's own monitor, which runs on the same cadence.

KeyDefaultNotes
interval30sProbe cadence.
timeout5sPer-probe deadline (must be < interval).
failure_threshold2Consecutive failures before the backend is marked down.

local and vertexai derive their model list from config and are always reported available.

image_llm — image backend (optional)​

Independent of llm — run both, either, or just images. When unset, the node 404s the image routes and no model may declare an image: block.

KeyNotes
provider"" | comfyui | comfyui_cloud | openai_passthrough (a hosted OpenAI-compatible image API such as xAI; see Image generation).
comfyui.base_urle.g. http://127.0.0.1:8000.
comfyui.data_dirOptional absolute path to the ComfyUI base dir so the node cleans up generated files. Override with NODE_IMAGE_LLM_COMFYUI_DATA_DIR.

Each image-capable model in zs.models[] declares its own image: block (backend, image_rate, max_n, defaults, ComfyUI template). See Serving models & pricing for the pricing model.

zs — operator identity, pricing, and behavior​

Identity and concurrency​

KeyDefaultNotes
operator_id0Your on-chain operator id from registration (0 is reserved). Owner address + NFD are read from the operator box at startup.
node_id0This node's per-operator id (0 reserved). Its signing address is read from the node box.
max_active_tickets16Default per-model concurrency cap; reserve returns 429 no_capacity when a model hits its cap. Each model has its own independent pool — override per model under zs.models[].max_active_tickets.
payer_slot_cap4How many of a model's max_active_tickets slots a single payer may hold at once, so one funded account can't take a model's whole pool. Over it, reserve returns 429 payer_slot_cap, a distinct code from no_capacity. Per-model override at zs.models[].payer_slot_cap; 0 disables the per-payer cap while leaving max_active_tickets in force. At the default of 4, one payer can hold half a pool of 8. File-only (no env override).
inject_safety_identifierfalseOpt-in. Fills OpenAI's safety_identifier field from the verified payer, so abuse signals at the upstream attach to an account rather than to your whole key. Env NODE_ZS_INJECT_SAFETY_IDENTIFIER.
owner_addr""Optional override. Normally read from the operator box on chain; pin it only if you need to bypass the chain lookup. Env NODE_ZS_OWNER_ADDR.
signing_addr""Optional override. Normally read from the node box on chain; the node still refuses to start unless the matching mnemonic is loaded via the keystore. Env NODE_ZS_SIGNING_ADDR.

The signing mnemonic itself is not a config key — it's provisioned through the keystore (see below and Encryption & keys).

NFD identity​

Which NFD the node publishes DNS under, and where inside it. These keys only matter when tls.mode is acme or ip_sync is on, but nfd_record_name is validated on every boot regardless.

KeyDefaultNotes
nfd_app_id0The NFD whose u.dns this node writes. Gap-filled from the operator box only when it is 0, so an explicit value always wins; set one to give a node its own NFD instead of sharing the operator's. Env NODE_ZS_NFD_APP_ID.
nfd.api_urlper-network defaultOverride the NFDomains API endpoint. Only needed for tests or an unusual network path. Env NODE_ZS_NFD_API_URL.
nfd_record_name"" (apex)Bare DNS label the node publishes under, inside nfd_app_id's u.dns. Lowercase letters, digits and internal hyphens; dotted labels (api.node1) allowed. Uppercase is rejected, not lowercased. A literal @ is rejected (leave it empty for the apex), and a value ending in .algo is rejected — that's an NFD name and belongs in nfd_app_id. Keep it short: it eats into the 248-byte base-URL ceiling. Env NODE_ZS_NFD_RECORD_NAME.
danger

Never let two nodes publish at the same record. Two nodes of one operator both left at the apex overwrite each other's A record on every tick — a permanent flap costing an on-chain transaction per tick per node, with clients reaching whichever wrote last. Because each write rewrites the whole u.dns document, racing writers can also drop each other's unrelated records, including a live _acme-challenge TXT mid-issuance. Nothing detects this. Give every node its own nfd_record_name, or its own nfd_app_id — see One label per node.

Ticket lifetimes​

KeyDefaultNotes
ticket_ttl5sWindow between reserve and the inference POST — not inference duration. A POST after expiry gets 402 ticket_invalid. A long stream on a 5s-TTL ticket still completes normally. Kept tight so an abandoned reserve frees its slot fast.
default_expires_after5mFallback settlement-complete deadline gating the contract's refund_inactive. Per-model override: expires_after. Keep it ≳ your p99 inference duration plus the ~30s settlement watchdog.

Pricing​

Rates are USD per 1,000,000 tokens (paste directly from a vendor price card) and are net of the protocol fee — what you receive. The node grosses up the escrowed max_price to cover the fee, so don't inflate your published rates.

  • default_pricing — {input_rate, output_rate} applied to any model without its own entry. An optional cache_read_rate discounts the cached-read input subset (see below).
  • models[] — per-model entries. pricing overrides the default; an omitted or empty pricing: {} inherits it; explicit {input_rate: 0, output_rate: 0} is a free model (never inherits). context declares the model's capabilities and size limits (see below). max_active_tickets overrides the global pool.

models[].context​

What the node advertises about a model, and the reserve-time ceiling it enforces. Every field is optional.

FieldWhat it does
context_windowTotal tokens the model accepts, input plus output. Reserves are sized against it. Most backends publish it and the node reads it automatically — vLLM and SGLang via max_model_len, llama.cpp and LM Studio via their own metadata, Vertex from its model list. Declare it only for gateways that publish no metadata (OpenAI, xAI), or to enforce something tighter. On a self-hosted runtime, prefer leaving it out: a hand-declared value overrides the real --max-model-len your server is running with.
max_output_tokensThe model's output ceiling. Must be below context_window. See Set an output ceiling — leaving it out caps long answers at a quarter of the context window (up to 32,768), and lets over-large requests through to your backend.
input_modalitiesContent types the model accepts: text, image, audio, video. Omitting implies text-only, which also suppresses feeding generated images back to the model.
output_modalitiesContent types it produces. Usually ["text"].
tool_useWhether the model can call tools. Overrides discovery.
reasoning{supported, allowed_efforts, default_effort}. Overrides discovery; gates reasoning replay in the tool loop.
max_input_imagesPer-prompt cap on input images. The node trims the oldest rather than letting the request fail. Enforced locally, not advertised.
tagsFree-form labels, merged with any from the model's HuggingFace card.

Run zs-node doctor after editing this block — it re-probes the backend and reports where your declarations and its actual behavior disagree.

Two more per-model blocks sit beside context and pricing:

  • source (effectively required) — the model's checkable identity, e.g. source: "hf:org/model". A model that satisfies neither this, an org/model-shaped id, a backend-supplied source, nor the frontier whitelist is dropped from your catalog, with only a WARN in the log. A malformed value is a hard startup error. See Model identity.
  • weights (optional) — {files, root, digest}, used to advertise a cryptographic digest of the bytes you serve for cross-operator comparison. You usually need none of it: provider: local and provider: kronk produce a stronger digest with no config at all. Declare files only when fronting an external engine. See Model integrity — and note that a digest your peers disagree with costs you routing placement.

And two rate modifiers:

  • cache_read_rate (optional, inside any pricing block) — the USD/1M rate for the cached-read input subset. Omit it and cached reads bill at input_rate (no change); 0 makes them free. It only takes effect when the upstream reports a cached count — vLLM needs --enable-prompt-tokens-details; LM Studio / Ollama don't report one, so it's a no-op there. See Serving models & pricing.
  • long_context (optional, inside any pricing block) — a surcharge tier for upstreams (xAI/Grok) that charge a higher rate once a prompt reaches a size threshold. {threshold_tokens, input_rate, output_rate} are required (the high rates must be ≥ their base counterparts), plus an optional high cache_read_rate that defaults to your base discount scaled by the input step-up. A request whose prompt reaches threshold_tokens bills at the high rates for the whole request. With a tier set, raise context_window to the model's true max instead of capping it at the threshold — the node re-checks the tier against the prompt actually sent when it bills, so an over-sized reserve doesn't over-charge. See Serving models & pricing.

Set either default_pricing or at least one models[] entry (or both). Full detail — modalities, capability badges, image routes — is in Serving models & pricing.

zs.coordinates​

Model-taxonomy discovery, on by default. With discover_huggingface: true the node fetches each served model's source repo once near startup and uses it to enrich what clients display: the author's curated card tags are unioned into the model's advertised tags, and params / family / quantization are surfaced on the per-model drill-in.

zs:
coordinates:
discover_huggingface: true # default

The fetch is cached per repo for the process lifetime and fails open, so it never blocks a model from being served. It reveals your served set to HuggingFace, a set already public on /v1/zs/details. Set it to false only for a node that must make no outbound calls to huggingface.co; that also turns off the registry cross-check described in Model integrity. Confirm it is working with zs_hf_coordinates_total{result}. File-only (no env override).

zs.min_charge​

Per-request minimum charge. min_price is the max of two components; settlement clamps amount_charged up to it on non-zero usage. Free models bypass it.

KeyDefaultNotes
output_tokens1000Bill as if at least this many output tokens were produced, so a tiny response never settles ~0. 0 disables.
algo_txns0µALGO network fees (in minTxnFee units of 1,000 µALGO) to recover via the oracle. Recommended 7 — the ~7,000 µALGO the operator absorbs per paid request on a live deployment (2 at open() + 5 at atomic settle()). The settle fee is sized to the payments that actually happen, so a free model costs 2 and a failed request 3. Auto-disabled on free models and while the oracle is down.

zs.reserve​

How much of a payer's USDC is locked per request, and whether your node checks that the prompt it receives matches what was reserved for it.

The reserve is sized from the actual request body plus a tool-loop headroom term, not from the model's full context window, so a short question against a 1M-context model locks cents rather than dollars. The client computes it; your node caps it at context_window − max_output, never raises it, and re-measures the decrypted body against it at inference time.

KeyDefaultNotes
enforce_input_budgettrueReconcile the decrypted prompt against the reserved input_count. Over budget, a plain request is refused before any upstream call (sealed, zero-cost, refunded in full) and a tool loop is cut off at the boundary — still returning a final answer synthesized from the results already gathered. This is your protection against a payer that under-declares. A prompt whose reserve is already at the context_window − max_output cap is forwarded instead, as long as its size estimate is within three times the budget, and your backend decides whether it fits (see How the ceiling gets sized). false runs monitor mode: over-budget requests are metered and logged but still served, so you can size the real rate before enforcing. The node logs a startup WARN while it's off.
input_budget_tolerance0.10Slack the measured input may exceed the reserve by before it counts as over budget — absorbs estimator jitter between the client's sizing and your re-measurement. Read clamped to [0, 1].
tool_headroom_per_iteration4000Per-iteration input-token headroom you advertise on /v1/zs/details. Advisory — your node never adds it and never recomputes it; clients read it and size their own reserve.

Env overrides: NODE_ZS_RESERVE_ENFORCE_INPUT_BUDGET, NODE_ZS_RESERVE_INPUT_BUDGET_TOLERANCE, NODE_ZS_RESERVE_TOOL_HEADROOM_PER_ITERATION.

A plain turn is never falsely rejected: the client sizes from the same bound your node re-measures with, so the two agree within tolerance. Clients add the headroom term only when the request's tools can grow the context your node measures — your zs_* built-in tools, or a frontier model's server-side tools. Tools the caller executes itself (a coding agent's shell and edit tools, client-side MCP) add none, so reserves from agent-style clients are legitimately much smaller.

info

Sizing tool_headroom_per_iteration. Derive it from what your own tools actually return. zs.builtin_tools.web_read.max_bytes defaults to 64 KiB of markdown — roughly 16,000 tokens for one read, about 4× the default per-iteration headroom. At the shipped defaults the aggregate (max_iterations × tool_headroom_per_iteration) absorbs that, but if you lower zs.builtin_tools.max_iterations without raising this value, tool loops get cut short. Watch zs_reserve_input_budget_over_total{outcome="tool_cutoff"} — see Monitoring & metrics.

zs.oracle​

ALGO/USD price feed. Used only so Reserve can express ALGO network fees in USD — it is not in the inference-pricing path, and Reserve tolerates an unhealthy oracle (it just omits algo_usd_price).

KeyDefaultNotes
sourcecoingeckoOnly CoinGecko is supported today.
refresh_interval30sFetch cadence.
max_staleness5mPast this, Reserve omits the price.
min_algo_usd0.01Reject quotes below this as unreliable.
http_timeout10sPer-fetch deadline.

zs.escrow_app_id and mempool polling​

The deployed ZeroSignalEscrow app id for your network.

  • On a public network you may omit it — the node falls back to the canonical embedded app id: testnet 765860477, mainnet 3628061142. An explicit value always wins. Localnet has no embedded default — set it there.
  • escrow_app_id: 0 is a test-isolation hook only (short-circuits payment verification); a 0-mode node is undiscoverable in production and can't claim USDC.
  • mempool_poll_timeout (5s) and mempool_poll_interval (100ms) bound how long the node waits for the open() group to appear in algod's pending pool. The interval is latency your callers feel: the first check is immediate, so whenever the payment hasn't propagated yet, a full interval elapses before inference starts. Raising it only reduces algod calls.

zs.rate_limits​

Public-facing throttling, on by default with the values below. A partial block keeps the defaults for the buckets you don't mention. Each bucket is a token bucket; setting its rps or burst to 0 disables that one bucket, and enabled: false switches the whole subsystem off (for private deployments behind their own edge). /healthz, /livez and /metrics are always exempt.

KeyDefaultScope
enabledtrueMaster switch. false leaves max_active_tickets as the only admission throttle.
BucketDefault rps / burstScope
reserve_per_ip2 / 5/v1/zs/reserve per IP.
reserve_per_account5 / 10Per Algorand payer address (escrow enabled).
llm_per_ip5 / 20Chat + responses per IP.
discovery_per_ip20 / 40/v1/models, /v1/zs/details, etc.
relay_per_ip10 / 20/v1/zs/relay forwards per IP (the transport-privacy hops you carry for others).
attestation_nonce1 / 5/v1/zs/attestation?nonce=…, one bucket for the whole node. Confidential mode only.

Plus idle_eviction (30m) and max_keys (100000). Client IP is read from CF-Connecting-IP > X-Real-IP > rightmost X-Forwarded-For > TCP peer.

attestation_nonce limits freshness challenges, each of which makes the enclave mint a new quote. It is not per IP because challenges arrive through relays, so the IP you see is a relay's. Separately, the node mints at most two challenged quotes at once, and that cap stays in force with rate limits switched off. Over either limit, the caller gets 429 attestation_rate_limited. All requests to /v1/zs/attestation also count against discovery_per_ip; requests without a nonce count only against that.

zs.builtin_tools​

In-loop tools the node advertises on /v1/zs/details and runs locally, aggregating all rounds into one receipt. See Built-in tools for the tool catalog and how they run.

KeyDefaultNotes
enabledtrueMaster switch. false disables the whole subsystem (including zs_get_time).
max_iterations20Absolute ceiling on chat→tool→chat rounds (clamped [1,20]). Applied when the key is omitted; the shipped example sets it to 5. Advertised to clients as max_tool_iterations.
max_stalled_iterations3Consecutive repeated-call rounds tolerated before the loop stops (clamped [1,10]). Node-internal; not advertised.
web_search.enabledtruePer-tool toggle for zs_web_search.
web_search.max_results10DuckDuckGo HTML; 1..25.
web_search.safe_searchmoderatestrict | moderate | off — an unknown value is a hard startup error. off returns unfiltered results; set it only if that's your intent.
web_search.timeout15sPer-search deadline.
web_read.enabledtruePer-tool toggle for zs_web_read.
web_read.max_bytes65536Truncation ceiling on the markdown handed to the model (keeps fetched pages within its context).
web_read.max_download4194304Ceiling on the raw page fetched off the wire, before conversion. Over this, the read is refused, not truncated. Must be >= max_bytes (compared on effective values, so an explicit 0 still means "the default"); a smaller value is a hard startup error. Not a memory budget: converting a page peaks at roughly 20–70× the page's size, so raising this raises RAM use by a large multiple. 4 MiB is ~2× the largest real-world article observed.
web_read.max_concurrent4How many pages the node converts at the same time. Peak memory ≈ max_download × 20–70 × this, so it's the setting to lower under a memory limit. Lowering it adds latency, because reads queue; lowering max_download instead would make large pages permanently unreadable. 4 suits a node that owns its box; 1 is reasonable in a small container. See Sizing the node process.
web_read.timeout15sPer-fetch deadline.
web_read.allow_private_targetsfalseLocalnet/dev only — leave off in production. Disables the guard that refuses to fetch a URL resolving to a private / loopback / link-local / carrier-NAT address. Your node does its own fetching, so with this on, a crafted URL can probe your LAN or cloud metadata endpoint. Logs a startup WARN when enabled. The web_read counterpart of zs.allow_private_relay_targets.
favicons.enabledtrueFetches each search-result site's favicon and inlines the bytes on the round's status frame, so chat clients show real site marks on the source chips. Fires only on the streaming Responses path. Turn it off on a node with metered or locked-down egress; clients then show a generic link glyph.
favicons.timeout1.2sBounds one tool call's whole fetch batch, not one site and not one turn: each call with hosts not already cached pays it again. A host already sent on the stream is skipped.
favicons.max_bytes16384Caps one icon. Oversize is dropped, not truncated. The default is the wire-format maximum; larger values are clamped to it.
favicons.cache_max_bytes4194304In-memory host→icon cache size.
favicons.cache_ttl24hHow long a resolved icon stays fresh. A miss is re-probed after 1/24th of it, and a fetch that merely timed out isn't cached at all.
favicons.allow_private_targetsfalseLocalnet/dev only. Same guard and same warning as web_read.allow_private_targets.
image_search.*—Same shape (enabled, max_results 1..100, safe_search, timeout) but force-disabled — see below.

Why the node fetches favicons instead of the user's browser, and what egress that adds, is in Source favicons.

info

image_search is force-disabled in code — its DuckDuckGo image backend is broken, so the node won't advertise or run it regardless of what you set here until a working provider is wired up.

zs.relay behavior​

Every node ships the /v1/zs/relay route, registered whenever a relay directory is available (escrow + algod configured). See Relays.

  • relay_only: true (or NODE_ZS_RELAY_ONLY=true) — pure relay: text and image providers empty, no zs.models, a non-zero escrow_app_id required. Prompt routes 404; only relay and discovery GETs are served.
  • allow_private_relay_targets (default false) — SSRF guard. The relay refuses to dial private/loopback/link-local/CGNAT/ULA targets so a malicious operator can't turn your relay into a LAN prober. Set true only for localnet/dev.

zs.self_eviction​

Watchdog (on by default) that re-reads this node's own operator box and exits non-zero if the operator is removed on-chain (admin-evicted or unregistered) so a supervisor surfaces it. Requires escrow_app_id + operator_id.

KeyDefaultNotes
enabledtrueDisable with false.
interval5mRe-check cadence.
threshold3Consecutive missing reads before exiting (rides out a reorg).

Env: NODE_ZS_SELF_EVICTION_{ENABLED,INTERVAL,THRESHOLD}.

info

Evicting a misbehaving operator is an admin moderation action (evictOperator); there is no permissionless eviction sweep, and nodes don't run one. self_eviction only notices that this node was removed and shuts it down cleanly.

zs.signing_balance​

Background poller that reads your hot signing address's ALGO balance so you can alert on it before on-chain transactions start failing. That account fee-pays every settlement and, when URL sync is on, the updateNodeUrl call. The poller publishes the spendable balance as zs_signing_balance_algos on /metrics, logs a WARN when the spendable balance (balance minus Algorand's minimum balance) drops below the threshold, and logs an ERROR when it falls under 5 ALGO — the floor below which proxies and the web client stop sending requests to your node; see Monitoring.

The NFD u.dns writes that ACME and IP sync make are fee-paid by whichever account owns the NFD, a different account unless you made them the same — see Which account signs the DNS writes.

KeyDefaultNotes
poll_interval5mHow often to read the balance. Set 0 to disable the poller, which also turns off both log lines; the node warns at startup if you do.
warn_below_algos10Emit a low-balance WARN when the spendable balance drops below this many ALGO. Keep it above 5, the routing floor, or the warning arrives only once traffic has stopped — the node warns at startup if it is 5 or lower. Advisory — independent of the hard 1 ALGO boot floor the node enforces at startup.

The gap between the default 10 and the 5 ALGO floor covers about 700 paid requests. On a busy node that can run out between two polls, so the WARN and the ERROR land together. To get the WARN at least one poll ahead of the floor, set warn_below_algos to at least 5 + (paid requests per second × poll_interval in seconds × 0.007) — about 5 + 2.1 × requests per second at the default 5m — plus enough to cover the time it takes you to refill.

The poller runs for relay-only nodes too, since they still fee-pay updateNodeUrl when URL sync is on. A relay-only node never receives inference requests, so the 5 ALGO floor doesn't apply to it and the node doesn't log the routing-floor ERROR; you can lower warn_below_algos to whatever float its ACME and IP-sync writes need.

zs.settlement_db_path​

Where the node records served tickets while the background settlement driver drives each through on-chain escrow.settle. On startup, leftover settling entries are reconciled against algod.

ValueBehavior
<path>Durable SQLite (WAL). Recommended for production.
:memory:Ephemeral SQLite (dev only).
""In-memory store, non-durable (dev only; the contract's refund_inactive backstop still protects payer funds).

Related:

  • settlement_watchdog_seconds (30) — how long the driver waits after a request completes before considering a standalone settle. Normally the client broadcasts the atomic two-transaction group the node pre-signed, which finalizes the ticket. Set this too short and the driver pre-empts that group: payers are unaffected, but you pay network fees for a standalone settle that wasn't needed. Raise it if your clients' acknowledgement latency is consistently above the window; lower it only if a metric says so. -1 fires immediately (tests only).
  • settlement_lapse_grace_seconds (300) — node-side wait before it force-finalizes a claim the client never acknowledged; must be ≥ the contract's grace default.
  • settlement_retention (720h, i.e. 30 days) and settlement_retention_interval — how long finalized ledger rows are kept before the node auto-purges them, and how often it sweeps.

See The payment flow for the full settlement lifecycle.

tee — confidential mode (opt-in)​

A non-none tee.mode advertises the node as TEE-capable on /v1/zs/details, exposes /v1/zs/attestation, and makes the proxy verify attestation before routing prompts here. tee.dataflow declares where plaintext comes to rest. The trust model, hardware requirements, what each posture requires of your backends, and the in-CVM identity bootstrap are in Confidential compute (TEE).

KeyDefaultNotes
modenonenone | stub | dstack-tdx | nvidia-cc-tdx | nvidia-cc-snp. dstack-tdx is the mode that runs in production; stub is laptop-only untrusted evidence; the nvidia-cc-* modes fail the startup check with a "not yet implemented" error until GPU attestation lands.
dataflowsealed_localsealed_local | attested_passthrough. Decides the retention tier payers see; see Two postures. No env override, because it has to be part of the measured config.
attestation.nras_urlhttps://nras.attestation.nvidia.comNVIDIA Remote Attestation Service — nvidia-cc-* only. dstack-tdx mints its quote from the local guest agent and never calls it.
attestation.refresh_interval1h/v1/zs/attestation serves 503 once stale past 2× this. Evidence is also re-minted on every ephemeral key rotation, independent of this cadence.
attestation.evidence_cache_path—Disk path on a CVM-encrypted volume so a brief bounce doesn't drop you from the verified set.
attestation.pccs_urlhttps://pccs.phala.networkWhere a dstack-tdx node fetches Intel's TCB and revocation collateral to carry alongside its quote. Change it only for a PCCS inside your own network; see Collateral your node carries.

Env: NODE_TEE_MODE, NODE_TEE_ATTESTATION_{NRAS_URL,REFRESH_INTERVAL,EVIDENCE_CACHE_PATH,PCCS_URL}.

A minimal config.yaml​

A working passthrough node: one model, an OpenAI-compatible upstream, and the canonical escrow app id for its network.

server:
listen: ":9090" # bind all interfaces (dual-stack); override of the loopback default
private_listen: "127.0.0.1:9091"
write_timeout: "0s" # keep 0 for streams (also the default)

# No algod block: mainnet is the default, and it carries the canonical
# escrow app id. Set algod.network only to target a different network.

logging:
level: "info"
format: "text"

llm:
provider: "openai_passthrough"
openai:
base_url: "https://api.openai.com/v1"
# api_key comes from NODE_LLM_OPENAI_API_KEY — never commit it
timeout: "0s" # keep 0 for streams

zs:
operator_id: 42 # your id from registration
node_id: 1 # this node's id
models:
gpt-5.4-mini:
# No `source:` — a frontier id clears the identity gate on its own. A
# self-hosted model would need `source: "hf:org/model"`; see below.
pricing:
input_rate: 0.31 # USD per 1M input tokens, net of the protocol fee
output_rate: 2.50 # USD per 1M output tokens
# cache_read_rate: 0.08 # optional cached-read discount (no-op unless
# the upstream reports a cached count)
min_charge:
output_tokens: 1000
algo_txns: 7 # recover the ~7,000 µALGO network fees per paid request (live deployment)
settlement_db_path: "./settlement.db" # dev value; for the systemd unit use /var/lib/zs-node/settlement.db (see installation.md)
warning

Every model needs a checkable identity, and a model without one is dropped with only a WARN. An always-on gate drops any model that has neither a source nor a frontier id from the advertised catalog. The node starts fine and /v1/models just comes back short. A frontier id (gpt-5.4-mini, claude-*, gemini-*, grok-*) clears the gate on the id alone. Recognized provider-qualified spellings clear it too — for example, x-ai/grok-4.5 and google/gemini-2.5-pro fold onto their curated bare-id entries. This does not whitelist arbitrary org/model strings. A declared org/model id derives its own source, and Kronk supplies one per model (see Model identity). Anything else, such as a bare GGUF stem or a custom id, needs an explicit source: "hf:org/model". It must be the hf:-prefixed repo ref, and a malformed one is a startup error, not a silent drop.

default_pricing on its own is not a substitute: upstream-discovered ids carry no source unless the backend supplies one, as Kronk does, so from any other backend only frontier-whitelisted ones survive. Declare what you intend to serve, or let zs-node init write the catalog for you — then zs-node doctor reports models pass the provenance gate when it's right.

Provide secrets through the environment, not the file:

export OPERATOR_SIGNING_MNEMONIC="word1 word2 ... word25"
export NODE_LLM_OPENAI_API_KEY="sk-..."

The signing mnemonic must reach the node at startup — it refuses to start without the mnemonic for its node's signing address. Any environment variable ending in _MNEMONIC is picked up (the label is informational; lookup is by derived address), or use a cloud secret manager via ZS_MNEMONIC_URLS (comma-separated name=url pairs). See Encryption & keys for the full keystore options.

info

For the exact install paths, Docker run lines, and reverse-proxy setup, see Installation. For registering the operator and getting your operator_id / node_id, see Registering on-chain.