Installation
The node software sits in front of your inference backend, advertises your catalog, and speaks the network's wire protocol. It's distributed as prebuilt binaries from the zs-node repository under the Functional Source License (FSL-1.1-ALv2), which converts to Apache 2.0 two years after each release. The source is not public yet; it will be published soon.
For what the node does once it's running, see How the network works and the concept pages.
Supported platforms & hardware
The node does not run models, so size it and the backend separately:
| What it is | What it costs | |
|---|---|---|
| The node process | The zs-node binary — envelopes, admission, settlement, built-in tools | Small and bounded. 1–2 vCPU, 2 GB RAM. Measured; see Sizing the node process. |
| The inference backend | Whatever actually runs the model | Everything else. Zero if the model runs somewhere you don't pay for. |
The backend's cost depends on where it runs, not which provider you name:
| Deployment | Config | This host needs |
|---|---|---|
| Backend runs elsewhere — a hosted gateway, OpenAI, Vertex | openai_passthrough with a remote base_url, or vertexai | No GPU. 1–2 vCPU, 2 GB RAM. |
| Backend runs on this box — vLLM, SGLang, Ollama | openai_passthrough with a local base_url, or llamacpp / lmstudio / kronk pointed at a sibling process | The node's 2 GB plus what the backend needs. |
| Backend runs on this box, node-supervised | local (llama.cpp) | Same, plus the node derives the backend's concurrency flags — see KV cache. |
| Image generation | image_llm.provider: comfyui | Its own VRAM, on top of any text model — see Image models. |
| Relay-only | zs.relay_only: true | 1 vCPU, 256 MB. No models, no tools, no settlement. |
openai_passthrough appears on both sides of that table, and it's the most
common way to run your own GPU. vLLM and SGLang speak the OpenAI API, so a node
in front of your own vLLM is openai_passthrough with base_url: http://127.0.0.1:8000/v1, not provider: local, which means specifically
"node-supervised llama.cpp".
Only local is a child process the node supervises. lmstudio, llamacpp,
kronk, a same-box vLLM, and ComfyUI are separate processes with their own
memory; the node's GOMEMLIMIT and container limits do not restrain them.
If your backend runs elsewhere, the node never loads weights: read Sizing the node process, then skip to Install paths.
What the inference backend needs
VRAM ≈ model weights + KV cache
Weights depend on the model and the quantization you pull. Model names
usually carry the format (…-nvfp4, …-fp8-dynamic, …-Q4_K_M), and it
changes the answer by up to 4×:
| Format | ≈ GB per 1B params | Runtime | GPU floor |
|---|---|---|---|
BF16 / FP16 | 2.0 | anything | any |
FP8 / fp8-dynamic | 1.0 | vLLM, SGLang | Hopper or newer (H100/H200/Blackwell) |
NVFP4 | ~0.6 | vLLM | Blackwell (B100/B200, RTX 50-series) |
MXFP4 | ~0.6 | vLLM, llama.cpp | wider than NVFP4 |
GGUF Q4_K_M | ~0.6 | llama.cpp, LM Studio, Ollama | any |
For a mixture-of-experts model the multiplier applies to total parameters, not active ones, because every expert stays resident. Qwen3.6 35B-A3B routes 3B active parameters through 256 experts, and you still pay for all 35B of weights.
FP8 and NVFP4 have a hardware floor. The A100 has no FP8 tensor cores: vLLM
either refuses --dtype fp8 or silently falls back to FP16, doubling the
memory you budgeted. NVFP4 needs Blackwell. Check the format against your card
before you pull 60 GB of weights.
KV cache depends on your config, and it scales with concurrency:
KV bytes/token = 2 × layers × kv_heads × head_dim × bytes_per_element
KV total = context_window × parallel_slots × KV bytes/token
parallel_slots: 0, the default, resolves to zs.max_active_tickets, so
raising your ticket concurrency multiplies your VRAM requirement. A config
that boots at max_active_tickets: 2 can fail to allocate at 16 with nothing
else changed.
layers means KV-bearing layers, not the model's layer count. Hybrid
stacks (Qwen3.6's Gated DeltaNet, and linear-attention families generally) keep
no conventional cache on their linear layers: Qwen3.6 27B has three Gated
DeltaNet layers for every Gated Attention layer, so only 16 of its 64 layers
hold a KV cache. Sliding-window layers cap at their window, so they add a flat
cost per sequence rather than a per-token one. Plugging the full layer count
into the formula on either architecture overestimates by ~4×.
Real numbers
KV figures come from each model's published config.json at f16 cache;
small marks sliding-window or compressed attention, where the cache is far
cheaper than the parameter count suggests. Total is weights + KV at
context_window: 32768 with 4 slots and nothing else on the card. These are
planning figures: your runtime, format and context all move them.
| Model | Params | Format | Weights | KV / token | KV @ 32k × 4 | Total | Realistic card |
|---|---|---|---|---|---|---|---|
| Gemma 4 E4B | 4.5B eff. | Q4_K_M | ~3 GB | small (SWA) | ~1 GB | ~4 GB | 8 GB, or CPU |
| Gemma 4 12B | 12B | Q4_K_M | ~7.5 GB | small (SWA) | ~2 GB | ~10 GB | 16 GB |
| Gemma 4 26B-A4B | 26B MoE (3.8B active) | NVFP4 | ~16 GB | small (SWA) | ~2 GB | ~18 GB | 24 GB Blackwell |
| Qwen3.6 35B-A3B | 35B MoE (3B active) | Q4_K_M | ~21 GB | 20 KB (10 of 40 layers) | ~2.5 GB | ~24 GB | 32 GB / 2× 24 GB |
| Qwen3.6 27B | 27B dense | Q4_K_M | ~17 GB | 64 KB (16 of 64 layers) | ~8 GB | ~25 GB | 32 GB / 2× 24 GB |
| Qwen3.6 27B | 27B dense | FP8 | ~27 GB | 64 KB | ~8 GB | ~35 GB | 40 GB Hopper+ |
| Gemma 4 31B | 31B dense | Q4_K_M | ~19 GB | 40 KB global + flat local | ~8 GB | ~27 GB | 32 GB |
| gpt-oss-120b | 120B MoE (5B active) | MXFP4 | ~61 GB | ~36 KB | ~5 GB | ~66 GB | 1× 80 GB (H100) |
| DeepSeek V4 Flash | 284B MoE (13B active) | FP8 | ~284 GB | small (compressed) | — | ~290 GB | 4× 80 GB |
| GLM-5.2 | 744B MoE (40B active) | FP8 | ~744 GB | — | — | ~800 GB | 8× 141 GB — one node |
| Qwen3.8 / Kimi K3 | 2.4T / 2.8T MoE | MXFP4 | ~1.2–1.5 TB | — | — | — | multi-node cluster |
- Parameter count doesn't predict KV cost.
gpt-oss-120bhas over 4× the parameters of Qwen3.6 27B and roughly half the per-token KV, because half its 36 layers are sliding-window and its head dimension is 64. - KV cache catches the weights when you widen the window, not when you add slots. Qwen3.6 27B advertises a 262K context; one slot at the full window is ~16 GB of cache against ~17 GB of weights. Gemma 4 31B lands in the same place at 256K. Four slots at 32k is the cheap end of that trade. Sizing a card off the weights column while advertising a six-figure context window leads to Out of VRAM at startup.
- Serve the top of the range through passthrough. GLM-5.2 needs an entire
8-GPU node before it serves one token; Qwen3.8 and Kimi K3 are multi-node at
any quantization. Serving those means
openai_passthroughto someone who runs the cluster.
If it doesn't fit
In rough order of what you give up:
- Lower
context_window. Linear on KV. On a sliding-window model it moves only the global layers, since the local ones are already capped at their window. It's also the per-session budget you advertise, so it's a product decision. - Lower
zs.max_active_tickets(or pinllm.local.parallel_slots). Linear on KV. You serve fewer requests at once. - Quantize the KV cache.
extra_args: ["--cache-type-k", "q8_0", "--cache-type-v", "q8_0"]roughly halves it, and llama.cpp's newer TurboQuant path goes further (3-bit KV via a randomized Hadamard transform) if your build has it. The node derives-c,--paralleland--cont-batchingitself and rejects them inextra_args, but the cache-type flags are yours to set. - Lower
gpu_layersto spill layers to CPU. It fits, and it's slow. - Smaller quantization, or a smaller model.
Nothing checks this for you. The node reads no GGUF metadata and never looks
at your VRAM; gpu_layers is passed to llama-server verbatim. The one guard
is a sanity ceiling of 4M tokens on context_window × parallel_slots, which
counts tokens and so can't tell a 24 GB card from a 192 GB one.
The failure therefore comes at startup, in the backend: llama-server fails
to allocate, the supervisor restarts it five times in sixty seconds, then gives
up permanently. Do the arithmetic before you deploy.
Image models
Image generation is a separate provider (image_llm.provider), and its VRAM
adds to the text model's if you serve both on one box. There is no KV cache: the
cost is the transformer plus the VAE, text encoder and activations, so it's
near-constant per model rather than scaling with concurrency.
The Template column is the node's built-in workflow name, which you put in
image.comfyui.template_internal. Anything else needs a template_path to a
workflow JSON you supply.
| Model | Format | VRAM | Template |
|---|---|---|---|
| Z-Image Turbo (6B) | BF16 | 14–16 GB | zimage, comfy_zimage |
| Z-Image Turbo (6B) | FP8 | ~8 GB (GGUF ~6 GB) | same |
| FLUX.2 klein (4B / 9B) | FP8 / GGUF | 12–16 GB | none — bring a template_path |
| FLUX.2 dev (32B) | FP8 | ~32 GB + text encoder | flux2_edit (edit only) |
| HiDream E1.1 | BF16 | 24+ GB | hidream_edit (edit only) |
Text encoders dominate two of these:
- FLUX.2's text encoder is itself a 24B model. The shipped
flux2_editworkflow pairs an fp8-mixed transformer with a BF16 Mistral-3-Small encoder because the all-BF16 pairing is ~64 GB and doesn't fit an 80 GB card in practice. hidream_editloads four text encoders (CLIP-G, CLIP-L, T5-XXL fp8 and Llama-3.1-8B fp8) on top of the transformer and VAE. Budget VRAM for the whole stack.
ComfyUI is operator-installed, not packaged with the node. On Linux + AMD, install the ROCm extras: the smaller models work and FLUX is hit-or-miss. Setup is in Serving models & pricing.
GPU and platform notes
- NVIDIA driver ≥ 535 (CUDA 12.x runtime) on any GPU host.
- Multi-GPU works without NVLink. llama.cpp splits by layer across cards, and layer-split is usually faster than row-split on non-NVLink pairs. NVLink helps tensor-parallel runtimes (vLLM) far more than it helps llama.cpp.
- Apple Silicon shares one memory pool between weights, KV cache and the OS. Apply the same arithmetic to unified memory and leave several GB for macOS. Fine for development, not production (see the OS table below).
- Confidential compute has its own hardware rules. The shipping
dstack-tdxmode attests the CPU (Intel TDX) only: it needs no GPU, and an attached GPU is outside the attestation. The forthcoming GPU modes need H100, H200 or Blackwell (A100 is not supported). See Confidential compute (TEE).
Sizing the node process
Skip this on a dedicated machine, where the inference engine dwarfs the node. It matters when the node gets a memory limit of its own: a container, a cgroup, or a VM sized to the pod rather than the box.
The node idles in the low hundreds of MB. The one path that spikes is web read: converting a page to markdown peaks at roughly 20–70× the page's size, because the readability pass builds and scores a document tree of the whole page. The multiple depends on how the page is built; a hostile page sits at the top, and the caller chooses the page.
zs.builtin_tools.web_read.max_concurrent (default 4) caps how many pages the
node converts at once, so a burst of large reads queues instead of multiplying:
peak ≈ max_download × (20–70) × max_concurrent
Measured on the shipped defaults with 16 concurrent readers: ~270 MB on ~2 MB pages (the largest real article we found), and ~560 MB when every reader fetches a page that nearly fills the 4 MiB ceiling. Those are ordinary page structures; a pathological page costs several times more per byte, up to the top of the 20–70× range. Budget for it.
| Deployment | Request | Limit | GOMEMLIMIT | Config |
|---|---|---|---|---|
| Serving node (inference elsewhere) | 256Mi | 2Gi | 1700MiB | defaults |
| Serving node, tighter | 256Mi | 1Gi | 850MiB | max_concurrent: 1 |
| Serving node, small container | 128Mi | 512Mi | 430MiB | max_concurrent: 1, max_download: 1048576 |
| Relay-only | 128Mi | 256Mi | 200MiB | n/a — no tools |
On a tight box, lower max_concurrent. Reads queue instead of failing, so
it adds latency but every page stays readable. Measured peak live heap, same
load, other settings at defaults:
max_concurrent | 1 | 2 | 4 | 8 |
|---|---|---|---|---|
| peak live heap | 153 MB | 220 MB | 345 MB | 513 MB |
Lowering max_download also cuts memory, but makes large pages permanently
unreadable; the 512Mi row above accepts that trade. web_read.enabled: false
removes the spike entirely. See Built-in tools.
A relay-only node only forwards traffic and answers discovery. It runs no built-in tools, settlement, or oracle, so it has no spike to budget for.
Set GOMEMLIMIT, not just a memory limit. A limits.memory on its own does
not restrain the node; it only decides when the kernel kills it.
Go's collector lets the heap grow well past what's live before it runs, so a
conversion burst can get the node OOM-killed for memory it wasn't really using.
GOMEMLIMIT is a soft limit: near it, the collector runs more often to stay
under, at the cost of CPU. Set it to roughly 85% of your limit to
leave room for stacks and allocator overhead the Go heap doesn't count.
Go reads the cgroup CPU limit but not the memory limit
(golang/go#75164 tracks adding it),
so set GOMEMLIMIT. Leave GOMAXPROCS unset: the runtime derives it from the
CPU limit, and setting it disables that.
Kubernetes, serving node:
resources:
requests:
memory: 256Mi
cpu: 500m
limits:
memory: 2Gi
env:
- name: GOMEMLIMIT
value: "1700MiB"
Operating systems:
| OS | Notes |
|---|---|
| Linux (recommended for production) | Ubuntu 22.04+ or any modern x86_64 distro. The node's SQLite is pure-Go (modernc.org/sqlite), so there's no libsqlite to install. |
| Windows | Runs natively or under WSL2; both work, using the prebuilt CUDA llama.cpp binary. |
| macOS | Supported as a dev environment (Metal). Not recommended for production because of lower concurrent throughput. |
To build from source, Go 1.26+. GPU hosts also need the driver listed in GPU and platform notes.
Network:
- Outbound HTTPS to your algod provider, to CoinGecko (for ALGO/USD pricing), and to your upstream LLM if you use OpenAI-compatible passthrough or Vertex AI.
- Outbound HTTPS to the NFDomains API (
api.nf.domains/api.testnet.nf.domains) and toapi64.ipify.org, if you usetls.mode: acmeorserver.tls.ip_sync. An egress ACL that omits these breaks certificate renewal and IP sync at runtime, not at startup. - Inbound HTTPS to the node's public listener, through your reverse proxy or the node's own TLS; see Exposing the endpoint over HTTPS.
You should already have an operator_id before you install. Register with your
owner wallet through the "Register Operator" form on the operator dashboard
at operator.zerosignal.ai; it returns the
operator_id you'll put in the node config. See
Registering on-chain, and
Encryption & keys for which keys the node holds.
Install paths
Pick one. A release binary (install script / Homebrew / Scoop / direct download) is the quickest; Docker is simplest for production. The bare-metal binary and building inside Docker build from source, so they apply once the source is published.
- Release binary
- Docker
- Bare-metal binary
- Build inside Docker
macOS builds are Developer ID–signed and notarized, so there's no Gatekeeper prompt on first run.
Install script (Linux / macOS):
curl -fsSL https://zerosignal.ai/install.sh | sh -s -- zs-node
The script detects your OS and architecture, verifies the SHA-256 against the
release's checksums.txt, and installs to /usr/local/bin if writable, else
~/.local/bin. It never uses sudo and never starts anything. Read it first
with curl -fsSL https://zerosignal.ai/install.sh | less. Pin a release with
ZS_VERSION=X.Y.Z, or choose the directory with ZS_INSTALL_DIR. Re-run it to
upgrade; it reinstalls into the same directory, which matters because the
service unit records the binary's path.
Homebrew (macOS):
brew install txnlab/tap/zs-node
Scoop (Windows):
scoop bucket add txnlab https://github.com/txnlab/scoop-bucket
scoop install zs-node
Direct download (Linux / any): download the archive for your platform from
the latest release, verify
its checksum against checksums.txt, and put the binary on your PATH:
| Platform | Archive |
|---|---|
| Linux x86_64 | zs-node_<version>_linux_amd64.tar.gz |
| Linux arm64 | zs-node_<version>_linux_arm64.tar.gz |
| macOS (universal Intel + Apple Silicon) | zs-node_<version>_darwin_all.tar.gz |
| Windows x86_64 | zs-node_<version>_windows_amd64.zip |
| Windows arm64 | zs-node_<version>_windows_arm64.zip |
sha256sum -c checksums.txt --ignore-missing # verify
tar -xzf zs-node_<version>_linux_amd64.tar.gz
sudo install -m 755 zs-node /usr/local/bin/zs-node
Then run it as a service: zs-node install-service generates and enables the
unit for you. See Running as a service. (The Homebrew
and Scoop installs already put zs-node on your PATH.)
The image is public and multi-arch (linux/amd64 + linux/arm64) on GHCR, and
docker run pulls it on first use. Mount your config file, and keep secrets
(the API key, the signing mnemonic) in an --env-file rather than inline -e
flags so they stay out of your shell history:
# /etc/zerosignal/secrets.env (chmod 600)
# NODE_LLM_OPENAI_API_KEY=sk-...
# OPERATOR_SIGNING_MNEMONIC=word1 word2 ... word25
docker run -d \
--name zs-node \
--restart unless-stopped \
--stop-timeout 400 \
-p 9090:9090 \
-p 127.0.0.1:9091:9091 \
--env-file /etc/zerosignal/secrets.env \
-e NODE_SERVER_LISTEN=:9090 \
-e NODE_SERVER_PRIVATE_LISTEN=:9091 \
-e NODE_ZS_SETTLEMENT_DB_PATH=/app/data/settlement.db \
-v /etc/zerosignal/config.yaml:/app/config.yaml:ro \
-v zs-node-data:/app/data \
ghcr.io/txnlab/zs-node:latest --config /app/config.yaml
--stop-timeout 400 gives the node its full drain budget to finish the
inference it already accepted; Docker's default is 10 seconds. See
Graceful shutdown.
Put the settlement DB inside the mounted volume. The code default for
settlement_db_path is empty (""), which keeps a non-durable in-memory
ledger that a restart wipes: fine for local dev, not for production. The example
config's ./settlement.db would resolve to /app/settlement.db here, outside
the zs-node-data volume. Set settlement_db_path: /app/data/settlement.db (or
the NODE_ZS_SETTLEMENT_DB_PATH env shown above). Losing the ledger strands any
in-flight settlements your node hadn't finalized; the contract's
refund_inactive backstop still protects payer funds.
Both listeners default to loopback (server.listen 127.0.0.1:9090,
server.private_listen 127.0.0.1:9091), which nothing outside the container
can reach. Set NODE_SERVER_LISTEN=:9090 so the published port serves traffic.
Set NODE_SERVER_PRIVATE_LISTEN=:9091 if your probes or Prometheus scraper run
in another container, and publish it on the host's loopback only
(-p 127.0.0.1:9091:9091) so /healthz, /livez and /metrics never leave the
host. For dual-stack binding and quoting ":9091" in YAML or compose, see
The two listeners.
Requires Go 1.26+. The node uses the experimental JSON v2 encoder, so the build tag is required:
go build -tags=goexperiment.jsonv2 -trimpath -ldflags "-s -w" -o zs-node .
sudo install -m 755 zs-node /usr/local/bin/zs-node
Then run it as a service — see Running as a service.
To produce a custom image, use the multi-stage Dockerfile as a base and build
with the same Go toolchain and build tag as the bare-metal binary. If your build
fetches private modules, mount the token secret at build time (see the comments
in the repo's Dockerfile).
Wiring the node to an inference backend
The node forwards decrypted prompts to the inference backend you select with
llm.provider (one provider per node):
| Provider | What it does |
|---|---|
openai_passthrough | Any OpenAI-compatible base URL (vLLM, Ollama, LM Studio, hosted gateways, or a real OpenAI-compatible endpoint). |
local | The node supervises a llama-server child process over loopback. The recommended path for a self-hosted GPU. |
lmstudio, llamacpp | Point at a separately-running LM Studio or llama-server. |
kronk | Point at a separately-running Kronk server — a multi-model llama.cpp pool in one process. |
vertexai | Google Vertex AI's OpenAI-compatible API (no static API key; uses Application Default Credentials). |
Per-provider configuration (base URLs, API keys, GGUF model paths,
gpu_layers, slots, image generation, and multi-model layouts) is in
Serving models & pricing and
Configuration.
You usually don't have to pick by hand. Point the wizard at your running backend and it works out which provider to use, the real API root, which endpoints exist, and what each model can do:
zs-node init --base-url=http://127.0.0.1:8080/v1
llama.cpp setup (the local provider)
With provider: "local", install llama-server once; the node starts,
supervises, and restarts it.
Pin a recent llama.cpp, not just a known-good one. The 2026 hybrid models
use new layer operators (Qwen3.6's Gated DeltaNet), and a llama-server older
than the build that added them refuses the GGUF outright instead of falling
back. The node treats a backend that won't start as a crash loop, so an old pin
looks like a broken node, not an unsupported model. Pin a tag, and re-check it
when you change model generation.
- Linux + NVIDIA CUDA
- Windows + NVIDIA CUDA
- macOS + Metal
# 1. Driver — aim for >= 535 (CUDA 12.x runtime).
sudo apt install -y nvidia-driver-550
sudo reboot
nvidia-smi # confirm GPU(s) visible
# 2. Build llama-server with CUDA. There is NO prebuilt Linux CUDA archive on
# the releases page — the ubuntu-* assets are CPU, Vulkan, ROCm and SYCL
# only, so CUDA on Linux means compiling. Pin a tag you've smoke-tested
# against the model you serve; hybrid-attention GGUFs need a current build.
LLAMA_TAG=b10628
sudo apt install -y build-essential cmake git libcurl4-openssl-dev
git clone --depth 1 --branch ${LLAMA_TAG} https://github.com/ggml-org/llama.cpp /tmp/llama
cmake -S /tmp/llama -B /tmp/llama/build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build /tmp/llama/build --config Release -j --target llama-server
sudo install -m 755 /tmp/llama/build/bin/llama-server /usr/local/bin/llama-server
sudo cp /tmp/llama/build/bin/lib*.so /usr/local/lib/
sudo ldconfig
llama-server --version
# 3. Pull a GGUF that fits weights AND KV cache — see "What the inference
# backend needs" above. Budget ~0.6 GB per billion params for weights, then
# add context_window x KV-bearing-layer KV-bytes-per-token x parallel_slots.
mkdir -p /srv/models
huggingface-cli download unsloth/Qwen3.6-27B-GGUF \
Qwen3.6-27B-Q4_K_M.gguf --local-dir /srv/models
If you'd rather not compile, run llama.cpp from its official CUDA image
(ghcr.io/ggml-org/llama.cpp:server-cuda) as a separate service and point the
node at it with provider: "llamacpp" instead of local. You give up the
node's supervision and derived concurrency flags (you set -c, --parallel and
--cont-batching yourself), and you need no toolchain on the box.
llm:
provider: "local"
local:
binary_path: "/usr/local/bin/llama-server"
parallel_slots: 0 # 0 = derive from max_active_tickets
models:
- id: "qwen3.6-27b"
model_path: "/srv/models/Qwen3.6-27B-Q4_K_M.gguf"
context_window: 32768 # x parallel_slots = your KV budget
gpu_layers: 99
Windows 11 or Server 2022. This is the native path; WSL2 also works.
# 1. NVIDIA Studio or Game Ready driver (latest). The CUDA toolkit is NOT
# required at runtime — the prebuilt binary ships the CUDA runtime libs.
# 2. Download and extract the Windows CUDA build — Windows is the one platform
# with a prebuilt CUDA archive. Pin a tag you've tested, and match the CUDA
# version in the asset name to your driver.
$tag = "b10628"
$url = "https://github.com/ggml-org/llama.cpp/releases/download/$tag/llama-$tag-bin-win-cuda-12.4-x64.zip"
Invoke-WebRequest -Uri $url -OutFile "$env:TEMP\llama.zip"
Expand-Archive "$env:TEMP\llama.zip" -DestinationPath "C:\llama"
# 3. GGUF model (huggingface-cli, or a direct download).
mkdir C:\models
Use forward slashes in the config, or escape the backslashes:
llm:
provider: "local"
local:
binary_path: "C:/llama/llama-server.exe"
models:
- id: "qwen3.6-27b"
model_path: "C:/models/Qwen3.6-27B-Q4_K_M.gguf"
context_window: 16384
gpu_layers: 99
Fine for development; not recommended for production.
brew install llama.cpp
llm:
provider: "local"
local:
binary_path: "/opt/homebrew/bin/llama-server"
models:
- id: "gemma-4-e4b"
model_path: "/Users/you/models/gemma-4-e4b-it-Q4_K_M.gguf"
context_window: 32768
gpu_layers: 99
context_window: 32768 with the default parallel_slots: 0 sizes the KV cache
at 32768 × zs.max_active_tickets tokens. For Qwen3.6 27B that's ~2 GB per
slot, so four slots need ~8 GB on top of the ~17 GB of weights. The ~25 GB total
fits a 32 GB card but not a 24 GB one. Check the arithmetic against your card
before you start the node, or lower one of the two. See What the inference backend needs.
The local provider supervises one llama-server instance, so it serves a
single model per node; a multi-entry models[] is a config error. To serve
several local models on one host, run one node per model, each with its own
config.yaml, public listen port, and settlement DB. They can share one
operator identity and signing mnemonic. See
Serving models & pricing.
Exposing the endpoint over HTTPS
Clients and relays reach you at the public HTTPS base URL you registered
on-chain (e.g. https://node.example.com). The node faces the internet
directly. Bind server.listen to all interfaces with :9090, which is
dual-stack (IPv4 + IPv6); 0.0.0.0 would be IPv4-only. Then give it HTTPS one
of two ways:
- The node's own TLS (simplest). With
server.tls.mode: acme, the node provisions and renews a Let's Encrypt certificate through your operator's NFD DNS, and keeps the DNS record and the on-chain base URL current. Withmanual, you supplycert_path+key_path. Nothing sits in front of the node.acmeneeds an NFD whose owner key the node holds; read Reaching your node before you pick a wallet. - A TLS-terminating reverse proxy. If you already run one, terminate TLS there and forward to the node. It is a single TLS front door for a single node, not a load balancer.
Do not put the node behind a load balancer. A node id is one process with one signing key and per-process admission and settlement state; multiple instances of the same node id corrupt accounting. To add capacity, register additional nodes (each with its own node id, key, and URL) under your operator and let clients load-balance across them; see Registration.
A reverse proxy in front of a node must forward streamed responses without buffering the whole reply, and its connection-draining window must outlive the node's shutdown drain, or it cuts responses mid-stream on every restart. See Graceful shutdown.
Reaching your node walks through the hostname, NFD record, certificate, on-chain base URL, and which account signs what. Configuration is the knob-by-knob reference, and Endpoint reference states the transport requirements.
Public vs private listeners
The node binds a public listener (server.listen) for the protocol routes
and a private one (server.private_listen, default 127.0.0.1:9091) for
/healthz, /livez and /metrics. The private listener is never TLS-wrapped
and is meant to stay on loopback or a private network you control. The port
server.listen binds is the port published in your base URL on chain, and
zs-node init writes this bind for you. Defaults, the single-port option, and
YAML quoting are in The two listeners.
Running as a service
systemd (bare-metal install). Let the node write and enable the unit:
sudo zs-node install-service
On first run it writes a 0600 secrets template to
/etc/zerosignal/secrets.env and stops so you can add your mnemonic; run it
again to install, enable, and start the service. sudo zs-node uninstall-service removes it. It generates exactly this unit, which you can
also write by hand:
# /etc/systemd/system/zs-node.service
# Generated by `zs-node install-service`. Re-run to regenerate.
[Unit]
Description=ZeroSignal node
Documentation=https://docs.zerosignal.ai/operators
After=network-online.target
Wants=network-online.target
# Bound a crash loop (e.g. a bad config): 5 restarts / 60s, then give up.
StartLimitIntervalSec=60
StartLimitBurst=5
[Service]
Type=simple
DynamicUser=yes
StateDirectory=zs-node
EnvironmentFile=/etc/zerosignal/secrets.env
ExecStart=/usr/local/bin/zs-node --config /etc/zerosignal/config.yaml
# Graceful stop: on SIGTERM the node stops accepting new work, then finishes the
# inference it already accepted before exiting. TimeoutStopSec MUST exceed the
# node's whole drain budget (server.drain_grace + drain_timeout +
# shutdown_timeout — defaults 20s + 5m + 30s = 350s) or systemd SIGKILLs it
# mid-drain, truncating paid inference and stranding payer escrow. Raise this if
# you raise drain_timeout. See docs.zerosignal.ai/operators/operations.
KillSignal=SIGTERM
TimeoutStopSec=400
Restart=on-failure
RestartSec=5
LimitNOFILE=65536
# Hardening
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
PrivateTmp=true
[Install]
WantedBy=multi-user.target
DynamicUser=yes runs the node as a transient system user that systemd
allocates at start, so there's no useradd and no user to manage.
StateDirectory=zs-node creates and owns /var/lib/zs-node for it, and
ProtectSystem=strict makes that the one writable path. Point your settlement
DB there. With tls.mode: acme, point server.tls.acme.cache_dir there too
(/var/lib/zs-node/tls-cache): its default, ./tls-cache, is relative to the
working directory, which under this unit is neither writable nor durable, and
losing it re-registers your Let's Encrypt account on every restart.
EnvironmentFile=/etc/zerosignal/secrets.env (mode 0600) holds your secrets:
OPERATOR_SIGNING_MNEMONIC=word1 word2 ... word25
NODE_LLM_OPENAI_API_KEY=sk-...
If you wrote the unit by hand, enable it (install-service does this for you):
sudo systemctl daemon-reload
sudo systemctl enable --now zs-node
sudo journalctl -u zs-node -f
Docker: the --restart unless-stopped flag in the
Docker install gives the same always-up behavior.
The signing mnemonic is the only secret the node must have at startup; it's
the on-chain identity that signs your tickets and receipts. Any env var ending
in _MNEMONIC (such as OPERATOR_SIGNING_MNEMONIC) is picked up automatically,
or you can load it from a cloud secret manager via ZS_MNEMONIC_URLS. See
Encryption & keys for the full keystore options.
Your algod must observe the mempool
The node admits a paid request by checking the payer's escrow open()
transaction while it's still pending, before it confirms, via
algod.PendingTransactionInformation. The algod you point the node at must
see incoming pending transactions in the mempool, or admission never succeeds.
- A public RPC (Nodely / AlgoNode) satisfies this with no configuration.
- A self-hosted algod must either be participating in consensus or
have
ForceFetchTransactions: trueset in itsconfig.json. A non-participating archival node with default settings does not observe the mempool and silently fails admission.
Privacy: your algod provider learns payer addresses. To admit a request, the node also looks up the payer's account (an escrow opt-in check and a free-tier allowance simulate), and both send the payer's Algorand address to your algod. A shared public RPC therefore sees which payer is reserving on your node. That address is a stable pseudonym, never the prompt, and there's no protocol fix for it today. To keep it from a third party, run your own algod.
You configure algod with the algod block (network, plus optional endpoint
and token; prefer NODE_ALGOD_TOKEN for the token). Full Algorand wiring,
including the escrow app id defaults, is in Configuration.