Skip to main content

Installation

The node software sits in front of your inference backend, advertises your catalog, and speaks the network's wire protocol. It's distributed as prebuilt binaries from the zs-node repository under the Functional Source License (FSL-1.1-ALv2), which converts to Apache 2.0 two years after each release. The source is not public yet; it will be published soon.

For what the node does once it's running, see How the network works and the concept pages.

Supported platforms & hardware​

The node does not run models, so size it and the backend separately:

What it isWhat it costs
The node processThe zs-node binary — envelopes, admission, settlement, built-in toolsSmall and bounded. 1–2 vCPU, 2 GB RAM. Measured; see Sizing the node process.
The inference backendWhatever actually runs the modelEverything else. Zero if the model runs somewhere you don't pay for.

The backend's cost depends on where it runs, not which provider you name:

DeploymentConfigThis host needs
Backend runs elsewhere — a hosted gateway, OpenAI, Vertexopenai_passthrough with a remote base_url, or vertexaiNo GPU. 1–2 vCPU, 2 GB RAM.
Backend runs on this box — vLLM, SGLang, Ollamaopenai_passthrough with a local base_url, or llamacpp / lmstudio / kronk pointed at a sibling processThe node's 2 GB plus what the backend needs.
Backend runs on this box, node-supervisedlocal (llama.cpp)Same, plus the node derives the backend's concurrency flags — see KV cache.
Image generationimage_llm.provider: comfyuiIts own VRAM, on top of any text model — see Image models.
Relay-onlyzs.relay_only: true1 vCPU, 256 MB. No models, no tools, no settlement.
warning

openai_passthrough appears on both sides of that table, and it's the most common way to run your own GPU. vLLM and SGLang speak the OpenAI API, so a node in front of your own vLLM is openai_passthrough with base_url: http://127.0.0.1:8000/v1, not provider: local, which means specifically "node-supervised llama.cpp".

Only local is a child process the node supervises. lmstudio, llamacpp, kronk, a same-box vLLM, and ComfyUI are separate processes with their own memory; the node's GOMEMLIMIT and container limits do not restrain them.

If your backend runs elsewhere, the node never loads weights: read Sizing the node process, then skip to Install paths.

What the inference backend needs​

VRAM ≈ model weights + KV cache

Weights depend on the model and the quantization you pull. Model names usually carry the format (…-nvfp4, …-fp8-dynamic, …-Q4_K_M), and it changes the answer by up to 4×:

Format≈ GB per 1B paramsRuntimeGPU floor
BF16 / FP162.0anythingany
FP8 / fp8-dynamic1.0vLLM, SGLangHopper or newer (H100/H200/Blackwell)
NVFP4~0.6vLLMBlackwell (B100/B200, RTX 50-series)
MXFP4~0.6vLLM, llama.cppwider than NVFP4
GGUF Q4_K_M~0.6llama.cpp, LM Studio, Ollamaany

For a mixture-of-experts model the multiplier applies to total parameters, not active ones, because every expert stays resident. Qwen3.6 35B-A3B routes 3B active parameters through 256 experts, and you still pay for all 35B of weights.

warning

FP8 and NVFP4 have a hardware floor. The A100 has no FP8 tensor cores: vLLM either refuses --dtype fp8 or silently falls back to FP16, doubling the memory you budgeted. NVFP4 needs Blackwell. Check the format against your card before you pull 60 GB of weights.

KV cache depends on your config, and it scales with concurrency:

KV bytes/token = 2 × layers × kv_heads × head_dim × bytes_per_element
KV total = context_window × parallel_slots × KV bytes/token

parallel_slots: 0, the default, resolves to zs.max_active_tickets, so raising your ticket concurrency multiplies your VRAM requirement. A config that boots at max_active_tickets: 2 can fail to allocate at 16 with nothing else changed.

info

layers means KV-bearing layers, not the model's layer count. Hybrid stacks (Qwen3.6's Gated DeltaNet, and linear-attention families generally) keep no conventional cache on their linear layers: Qwen3.6 27B has three Gated DeltaNet layers for every Gated Attention layer, so only 16 of its 64 layers hold a KV cache. Sliding-window layers cap at their window, so they add a flat cost per sequence rather than a per-token one. Plugging the full layer count into the formula on either architecture overestimates by ~4×.

Real numbers​

KV figures come from each model's published config.json at f16 cache; small marks sliding-window or compressed attention, where the cache is far cheaper than the parameter count suggests. Total is weights + KV at context_window: 32768 with 4 slots and nothing else on the card. These are planning figures: your runtime, format and context all move them.

ModelParamsFormatWeightsKV / tokenKV @ 32k × 4TotalRealistic card
Gemma 4 E4B4.5B eff.Q4_K_M~3 GBsmall (SWA)~1 GB~4 GB8 GB, or CPU
Gemma 4 12B12BQ4_K_M~7.5 GBsmall (SWA)~2 GB~10 GB16 GB
Gemma 4 26B-A4B26B MoE (3.8B active)NVFP4~16 GBsmall (SWA)~2 GB~18 GB24 GB Blackwell
Qwen3.6 35B-A3B35B MoE (3B active)Q4_K_M~21 GB20 KB (10 of 40 layers)~2.5 GB~24 GB32 GB / 2× 24 GB
Qwen3.6 27B27B denseQ4_K_M~17 GB64 KB (16 of 64 layers)~8 GB~25 GB32 GB / 2× 24 GB
Qwen3.6 27B27B denseFP8~27 GB64 KB~8 GB~35 GB40 GB Hopper+
Gemma 4 31B31B denseQ4_K_M~19 GB40 KB global + flat local~8 GB~27 GB32 GB
gpt-oss-120b120B MoE (5B active)MXFP4~61 GB~36 KB~5 GB~66 GB1× 80 GB (H100)
DeepSeek V4 Flash284B MoE (13B active)FP8~284 GBsmall (compressed)—~290 GB4× 80 GB
GLM-5.2744B MoE (40B active)FP8~744 GB——~800 GB8× 141 GB — one node
Qwen3.8 / Kimi K32.4T / 2.8T MoEMXFP4~1.2–1.5 TB———multi-node cluster
  • Parameter count doesn't predict KV cost. gpt-oss-120b has over 4× the parameters of Qwen3.6 27B and roughly half the per-token KV, because half its 36 layers are sliding-window and its head dimension is 64.
  • KV cache catches the weights when you widen the window, not when you add slots. Qwen3.6 27B advertises a 262K context; one slot at the full window is ~16 GB of cache against ~17 GB of weights. Gemma 4 31B lands in the same place at 256K. Four slots at 32k is the cheap end of that trade. Sizing a card off the weights column while advertising a six-figure context window leads to Out of VRAM at startup.
  • Serve the top of the range through passthrough. GLM-5.2 needs an entire 8-GPU node before it serves one token; Qwen3.8 and Kimi K3 are multi-node at any quantization. Serving those means openai_passthrough to someone who runs the cluster.

If it doesn't fit​

In rough order of what you give up:

  1. Lower context_window. Linear on KV. On a sliding-window model it moves only the global layers, since the local ones are already capped at their window. It's also the per-session budget you advertise, so it's a product decision.
  2. Lower zs.max_active_tickets (or pin llm.local.parallel_slots). Linear on KV. You serve fewer requests at once.
  3. Quantize the KV cache. extra_args: ["--cache-type-k", "q8_0", "--cache-type-v", "q8_0"] roughly halves it, and llama.cpp's newer TurboQuant path goes further (3-bit KV via a randomized Hadamard transform) if your build has it. The node derives -c, --parallel and --cont-batching itself and rejects them in extra_args, but the cache-type flags are yours to set.
  4. Lower gpu_layers to spill layers to CPU. It fits, and it's slow.
  5. Smaller quantization, or a smaller model.
warning

Nothing checks this for you. The node reads no GGUF metadata and never looks at your VRAM; gpu_layers is passed to llama-server verbatim. The one guard is a sanity ceiling of 4M tokens on context_window × parallel_slots, which counts tokens and so can't tell a 24 GB card from a 192 GB one.

The failure therefore comes at startup, in the backend: llama-server fails to allocate, the supervisor restarts it five times in sixty seconds, then gives up permanently. Do the arithmetic before you deploy.

Image models​

Image generation is a separate provider (image_llm.provider), and its VRAM adds to the text model's if you serve both on one box. There is no KV cache: the cost is the transformer plus the VAE, text encoder and activations, so it's near-constant per model rather than scaling with concurrency.

The Template column is the node's built-in workflow name, which you put in image.comfyui.template_internal. Anything else needs a template_path to a workflow JSON you supply.

ModelFormatVRAMTemplate
Z-Image Turbo (6B)BF1614–16 GBzimage, comfy_zimage
Z-Image Turbo (6B)FP8~8 GB (GGUF ~6 GB)same
FLUX.2 klein (4B / 9B)FP8 / GGUF12–16 GBnone — bring a template_path
FLUX.2 dev (32B)FP8~32 GB + text encoderflux2_edit (edit only)
HiDream E1.1BF1624+ GBhidream_edit (edit only)

Text encoders dominate two of these:

  • FLUX.2's text encoder is itself a 24B model. The shipped flux2_edit workflow pairs an fp8-mixed transformer with a BF16 Mistral-3-Small encoder because the all-BF16 pairing is ~64 GB and doesn't fit an 80 GB card in practice.
  • hidream_edit loads four text encoders (CLIP-G, CLIP-L, T5-XXL fp8 and Llama-3.1-8B fp8) on top of the transformer and VAE. Budget VRAM for the whole stack.

ComfyUI is operator-installed, not packaged with the node. On Linux + AMD, install the ROCm extras: the smaller models work and FLUX is hit-or-miss. Setup is in Serving models & pricing.

GPU and platform notes​

  • NVIDIA driver ≥ 535 (CUDA 12.x runtime) on any GPU host.
  • Multi-GPU works without NVLink. llama.cpp splits by layer across cards, and layer-split is usually faster than row-split on non-NVLink pairs. NVLink helps tensor-parallel runtimes (vLLM) far more than it helps llama.cpp.
  • Apple Silicon shares one memory pool between weights, KV cache and the OS. Apply the same arithmetic to unified memory and leave several GB for macOS. Fine for development, not production (see the OS table below).
  • Confidential compute has its own hardware rules. The shipping dstack-tdx mode attests the CPU (Intel TDX) only: it needs no GPU, and an attached GPU is outside the attestation. The forthcoming GPU modes need H100, H200 or Blackwell (A100 is not supported). See Confidential compute (TEE).

Sizing the node process​

Skip this on a dedicated machine, where the inference engine dwarfs the node. It matters when the node gets a memory limit of its own: a container, a cgroup, or a VM sized to the pod rather than the box.

The node idles in the low hundreds of MB. The one path that spikes is web read: converting a page to markdown peaks at roughly 20–70× the page's size, because the readability pass builds and scores a document tree of the whole page. The multiple depends on how the page is built; a hostile page sits at the top, and the caller chooses the page.

zs.builtin_tools.web_read.max_concurrent (default 4) caps how many pages the node converts at once, so a burst of large reads queues instead of multiplying:

peak ≈ max_download × (20–70) × max_concurrent

Measured on the shipped defaults with 16 concurrent readers: ~270 MB on ~2 MB pages (the largest real article we found), and ~560 MB when every reader fetches a page that nearly fills the 4 MiB ceiling. Those are ordinary page structures; a pathological page costs several times more per byte, up to the top of the 20–70× range. Budget for it.

DeploymentRequestLimitGOMEMLIMITConfig
Serving node (inference elsewhere)256Mi2Gi1700MiBdefaults
Serving node, tighter256Mi1Gi850MiBmax_concurrent: 1
Serving node, small container128Mi512Mi430MiBmax_concurrent: 1, max_download: 1048576
Relay-only128Mi256Mi200MiBn/a — no tools

On a tight box, lower max_concurrent. Reads queue instead of failing, so it adds latency but every page stays readable. Measured peak live heap, same load, other settings at defaults:

max_concurrent1248
peak live heap153 MB220 MB345 MB513 MB

Lowering max_download also cuts memory, but makes large pages permanently unreadable; the 512Mi row above accepts that trade. web_read.enabled: false removes the spike entirely. See Built-in tools.

A relay-only node only forwards traffic and answers discovery. It runs no built-in tools, settlement, or oracle, so it has no spike to budget for.

warning

Set GOMEMLIMIT, not just a memory limit. A limits.memory on its own does not restrain the node; it only decides when the kernel kills it.

Go's collector lets the heap grow well past what's live before it runs, so a conversion burst can get the node OOM-killed for memory it wasn't really using. GOMEMLIMIT is a soft limit: near it, the collector runs more often to stay under, at the cost of CPU. Set it to roughly 85% of your limit to leave room for stacks and allocator overhead the Go heap doesn't count.

Go reads the cgroup CPU limit but not the memory limit (golang/go#75164 tracks adding it), so set GOMEMLIMIT. Leave GOMAXPROCS unset: the runtime derives it from the CPU limit, and setting it disables that.

Kubernetes, serving node:

resources:
requests:
memory: 256Mi
cpu: 500m
limits:
memory: 2Gi
env:
- name: GOMEMLIMIT
value: "1700MiB"

Operating systems:

OSNotes
Linux (recommended for production)Ubuntu 22.04+ or any modern x86_64 distro. The node's SQLite is pure-Go (modernc.org/sqlite), so there's no libsqlite to install.
WindowsRuns natively or under WSL2; both work, using the prebuilt CUDA llama.cpp binary.
macOSSupported as a dev environment (Metal). Not recommended for production because of lower concurrent throughput.

To build from source, Go 1.26+. GPU hosts also need the driver listed in GPU and platform notes.

Network:

  • Outbound HTTPS to your algod provider, to CoinGecko (for ALGO/USD pricing), and to your upstream LLM if you use OpenAI-compatible passthrough or Vertex AI.
  • Outbound HTTPS to the NFDomains API (api.nf.domains / api.testnet.nf.domains) and to api64.ipify.org, if you use tls.mode: acme or server.tls.ip_sync. An egress ACL that omits these breaks certificate renewal and IP sync at runtime, not at startup.
  • Inbound HTTPS to the node's public listener, through your reverse proxy or the node's own TLS; see Exposing the endpoint over HTTPS.
info

You should already have an operator_id before you install. Register with your owner wallet through the "Register Operator" form on the operator dashboard at operator.zerosignal.ai; it returns the operator_id you'll put in the node config. See Registering on-chain, and Encryption & keys for which keys the node holds.

Install paths​

Pick one. A release binary (install script / Homebrew / Scoop / direct download) is the quickest; Docker is simplest for production. The bare-metal binary and building inside Docker build from source, so they apply once the source is published.

macOS builds are Developer ID–signed and notarized, so there's no Gatekeeper prompt on first run.

Install script (Linux / macOS):

curl -fsSL https://zerosignal.ai/install.sh | sh -s -- zs-node

The script detects your OS and architecture, verifies the SHA-256 against the release's checksums.txt, and installs to /usr/local/bin if writable, else ~/.local/bin. It never uses sudo and never starts anything. Read it first with curl -fsSL https://zerosignal.ai/install.sh | less. Pin a release with ZS_VERSION=X.Y.Z, or choose the directory with ZS_INSTALL_DIR. Re-run it to upgrade; it reinstalls into the same directory, which matters because the service unit records the binary's path.

Homebrew (macOS):

brew install txnlab/tap/zs-node

Scoop (Windows):

scoop bucket add txnlab https://github.com/txnlab/scoop-bucket
scoop install zs-node

Direct download (Linux / any): download the archive for your platform from the latest release, verify its checksum against checksums.txt, and put the binary on your PATH:

PlatformArchive
Linux x86_64zs-node_<version>_linux_amd64.tar.gz
Linux arm64zs-node_<version>_linux_arm64.tar.gz
macOS (universal Intel + Apple Silicon)zs-node_<version>_darwin_all.tar.gz
Windows x86_64zs-node_<version>_windows_amd64.zip
Windows arm64zs-node_<version>_windows_arm64.zip
sha256sum -c checksums.txt --ignore-missing # verify
tar -xzf zs-node_<version>_linux_amd64.tar.gz
sudo install -m 755 zs-node /usr/local/bin/zs-node

Then run it as a service: zs-node install-service generates and enables the unit for you. See Running as a service. (The Homebrew and Scoop installs already put zs-node on your PATH.)

Wiring the node to an inference backend​

The node forwards decrypted prompts to the inference backend you select with llm.provider (one provider per node):

ProviderWhat it does
openai_passthroughAny OpenAI-compatible base URL (vLLM, Ollama, LM Studio, hosted gateways, or a real OpenAI-compatible endpoint).
localThe node supervises a llama-server child process over loopback. The recommended path for a self-hosted GPU.
lmstudio, llamacppPoint at a separately-running LM Studio or llama-server.
kronkPoint at a separately-running Kronk server — a multi-model llama.cpp pool in one process.
vertexaiGoogle Vertex AI's OpenAI-compatible API (no static API key; uses Application Default Credentials).

Per-provider configuration (base URLs, API keys, GGUF model paths, gpu_layers, slots, image generation, and multi-model layouts) is in Serving models & pricing and Configuration.

tip

You usually don't have to pick by hand. Point the wizard at your running backend and it works out which provider to use, the real API root, which endpoints exist, and what each model can do:

zs-node init --base-url=http://127.0.0.1:8080/v1

See Generate it: zs-node init.

llama.cpp setup (the local provider)​

With provider: "local", install llama-server once; the node starts, supervises, and restarts it.

warning

Pin a recent llama.cpp, not just a known-good one. The 2026 hybrid models use new layer operators (Qwen3.6's Gated DeltaNet), and a llama-server older than the build that added them refuses the GGUF outright instead of falling back. The node treats a backend that won't start as a crash loop, so an old pin looks like a broken node, not an unsupported model. Pin a tag, and re-check it when you change model generation.

# 1. Driver — aim for >= 535 (CUDA 12.x runtime).
sudo apt install -y nvidia-driver-550
sudo reboot
nvidia-smi # confirm GPU(s) visible

# 2. Build llama-server with CUDA. There is NO prebuilt Linux CUDA archive on
# the releases page — the ubuntu-* assets are CPU, Vulkan, ROCm and SYCL
# only, so CUDA on Linux means compiling. Pin a tag you've smoke-tested
# against the model you serve; hybrid-attention GGUFs need a current build.
LLAMA_TAG=b10628
sudo apt install -y build-essential cmake git libcurl4-openssl-dev
git clone --depth 1 --branch ${LLAMA_TAG} https://github.com/ggml-org/llama.cpp /tmp/llama
cmake -S /tmp/llama -B /tmp/llama/build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build /tmp/llama/build --config Release -j --target llama-server
sudo install -m 755 /tmp/llama/build/bin/llama-server /usr/local/bin/llama-server
sudo cp /tmp/llama/build/bin/lib*.so /usr/local/lib/
sudo ldconfig
llama-server --version

# 3. Pull a GGUF that fits weights AND KV cache — see "What the inference
# backend needs" above. Budget ~0.6 GB per billion params for weights, then
# add context_window x KV-bearing-layer KV-bytes-per-token x parallel_slots.
mkdir -p /srv/models
huggingface-cli download unsloth/Qwen3.6-27B-GGUF \
Qwen3.6-27B-Q4_K_M.gguf --local-dir /srv/models
tip

If you'd rather not compile, run llama.cpp from its official CUDA image (ghcr.io/ggml-org/llama.cpp:server-cuda) as a separate service and point the node at it with provider: "llamacpp" instead of local. You give up the node's supervision and derived concurrency flags (you set -c, --parallel and --cont-batching yourself), and you need no toolchain on the box.

llm:
provider: "local"
local:
binary_path: "/usr/local/bin/llama-server"
parallel_slots: 0 # 0 = derive from max_active_tickets
models:
- id: "qwen3.6-27b"
model_path: "/srv/models/Qwen3.6-27B-Q4_K_M.gguf"
context_window: 32768 # x parallel_slots = your KV budget
gpu_layers: 99
warning

context_window: 32768 with the default parallel_slots: 0 sizes the KV cache at 32768 × zs.max_active_tickets tokens. For Qwen3.6 27B that's ~2 GB per slot, so four slots need ~8 GB on top of the ~17 GB of weights. The ~25 GB total fits a 32 GB card but not a 24 GB one. Check the arithmetic against your card before you start the node, or lower one of the two. See What the inference backend needs.

info

The local provider supervises one llama-server instance, so it serves a single model per node; a multi-entry models[] is a config error. To serve several local models on one host, run one node per model, each with its own config.yaml, public listen port, and settlement DB. They can share one operator identity and signing mnemonic. See Serving models & pricing.

Exposing the endpoint over HTTPS​

Clients and relays reach you at the public HTTPS base URL you registered on-chain (e.g. https://node.example.com). The node faces the internet directly. Bind server.listen to all interfaces with :9090, which is dual-stack (IPv4 + IPv6); 0.0.0.0 would be IPv4-only. Then give it HTTPS one of two ways:

  • The node's own TLS (simplest). With server.tls.mode: acme, the node provisions and renews a Let's Encrypt certificate through your operator's NFD DNS, and keeps the DNS record and the on-chain base URL current. With manual, you supply cert_path + key_path. Nothing sits in front of the node. acme needs an NFD whose owner key the node holds; read Reaching your node before you pick a wallet.
  • A TLS-terminating reverse proxy. If you already run one, terminate TLS there and forward to the node. It is a single TLS front door for a single node, not a load balancer.
warning

Do not put the node behind a load balancer. A node id is one process with one signing key and per-process admission and settlement state; multiple instances of the same node id corrupt accounting. To add capacity, register additional nodes (each with its own node id, key, and URL) under your operator and let clients load-balance across them; see Registration.

A reverse proxy in front of a node must forward streamed responses without buffering the whole reply, and its connection-draining window must outlive the node's shutdown drain, or it cuts responses mid-stream on every restart. See Graceful shutdown.

Reaching your node walks through the hostname, NFD record, certificate, on-chain base URL, and which account signs what. Configuration is the knob-by-knob reference, and Endpoint reference states the transport requirements.

Public vs private listeners​

The node binds a public listener (server.listen) for the protocol routes and a private one (server.private_listen, default 127.0.0.1:9091) for /healthz, /livez and /metrics. The private listener is never TLS-wrapped and is meant to stay on loopback or a private network you control. The port server.listen binds is the port published in your base URL on chain, and zs-node init writes this bind for you. Defaults, the single-port option, and YAML quoting are in The two listeners.

Running as a service​

systemd (bare-metal install). Let the node write and enable the unit:

sudo zs-node install-service

On first run it writes a 0600 secrets template to /etc/zerosignal/secrets.env and stops so you can add your mnemonic; run it again to install, enable, and start the service. sudo zs-node uninstall-service removes it. It generates exactly this unit, which you can also write by hand:

# /etc/systemd/system/zs-node.service
# Generated by `zs-node install-service`. Re-run to regenerate.
[Unit]
Description=ZeroSignal node
Documentation=https://docs.zerosignal.ai/operators
After=network-online.target
Wants=network-online.target
# Bound a crash loop (e.g. a bad config): 5 restarts / 60s, then give up.
StartLimitIntervalSec=60
StartLimitBurst=5

[Service]
Type=simple
DynamicUser=yes
StateDirectory=zs-node
EnvironmentFile=/etc/zerosignal/secrets.env
ExecStart=/usr/local/bin/zs-node --config /etc/zerosignal/config.yaml
# Graceful stop: on SIGTERM the node stops accepting new work, then finishes the
# inference it already accepted before exiting. TimeoutStopSec MUST exceed the
# node's whole drain budget (server.drain_grace + drain_timeout +
# shutdown_timeout — defaults 20s + 5m + 30s = 350s) or systemd SIGKILLs it
# mid-drain, truncating paid inference and stranding payer escrow. Raise this if
# you raise drain_timeout. See docs.zerosignal.ai/operators/operations.
KillSignal=SIGTERM
TimeoutStopSec=400
Restart=on-failure
RestartSec=5
LimitNOFILE=65536

# Hardening
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
PrivateTmp=true

[Install]
WantedBy=multi-user.target

DynamicUser=yes runs the node as a transient system user that systemd allocates at start, so there's no useradd and no user to manage. StateDirectory=zs-node creates and owns /var/lib/zs-node for it, and ProtectSystem=strict makes that the one writable path. Point your settlement DB there. With tls.mode: acme, point server.tls.acme.cache_dir there too (/var/lib/zs-node/tls-cache): its default, ./tls-cache, is relative to the working directory, which under this unit is neither writable nor durable, and losing it re-registers your Let's Encrypt account on every restart.

EnvironmentFile=/etc/zerosignal/secrets.env (mode 0600) holds your secrets:

OPERATOR_SIGNING_MNEMONIC=word1 word2 ... word25
NODE_LLM_OPENAI_API_KEY=sk-...

If you wrote the unit by hand, enable it (install-service does this for you):

sudo systemctl daemon-reload
sudo systemctl enable --now zs-node
sudo journalctl -u zs-node -f

Docker: the --restart unless-stopped flag in the Docker install gives the same always-up behavior.

info

The signing mnemonic is the only secret the node must have at startup; it's the on-chain identity that signs your tickets and receipts. Any env var ending in _MNEMONIC (such as OPERATOR_SIGNING_MNEMONIC) is picked up automatically, or you can load it from a cloud secret manager via ZS_MNEMONIC_URLS. See Encryption & keys for the full keystore options.

Your algod must observe the mempool​

The node admits a paid request by checking the payer's escrow open() transaction while it's still pending, before it confirms, via algod.PendingTransactionInformation. The algod you point the node at must see incoming pending transactions in the mempool, or admission never succeeds.

  • A public RPC (Nodely / AlgoNode) satisfies this with no configuration.
  • A self-hosted algod must either be participating in consensus or have ForceFetchTransactions: true set in its config.json. A non-participating archival node with default settings does not observe the mempool and silently fails admission.
info

Privacy: your algod provider learns payer addresses. To admit a request, the node also looks up the payer's account (an escrow opt-in check and a free-tier allowance simulate), and both send the payer's Algorand address to your algod. A shared public RPC therefore sees which payer is reserving on your node. That address is a stable pseudonym, never the prompt, and there's no protocol fix for it today. To keep it from a third party, run your own algod.

You configure algod with the algod block (network, plus optional endpoint and token; prefer NODE_ALGOD_TOKEN for the token). Full Algorand wiring, including the escrow app id defaults, is in Configuration.