Operations
Most config changes take effect on restart. In-flight tickets always carry
the rate and terms they were quoted at, so a restart never changes the price of
a request already under way — only new reserves see new settings. Run the node
under a supervisor (systemd, Docker --restart, an orchestrator) so "restart"
is a one-liner; see Installation.
Top up the signing account
Your hot signing account fee-pays every on-chain action the node signs:
every settlement, plus the base-URL update if you've enabled ip_sync's URL
sync. It needs an ALGO float on top of the account's own minimum balance (0.1
ALGO for a plain account). Budget roughly
5 ALGO + minimum balance + (daily request volume × ~7,000 µALGO) + a few days of buffer;
Staking & economics
breaks down the per-request cost.
- Keep at least 5 ALGO spendable. Spendable means the balance minus Algorand's minimum balance. Proxies and the web client check it every few minutes and stop sending requests to a node below 5 ALGO spendable — the node keeps running, it just gets no traffic until you top it up.
- Watch it. The node polls the balance and publishes the spendable amount
as
zs_signing_balance_algos, warns when it drops below your threshold (default10ALGO), logs an error once it falls under the 5 ALGO routing floor, and refuses to start below1ALGO. Alert on that metric or thesigning-address ALGO balance is lowlog line — see Monitoring. - Raise the warning on a busy node. The gap between 10 and 5 ALGO covers
about 700 paid requests. A node that serves several requests a second can use
that up between two balance checks, so the warning and the error arrive
together. Set
warn_below_algosto at least5 + (paid requests per second × 2.1)at the default 5-minute check, plus enough to cover the time it takes you to refill — seezs.signing_balance. - Refill from the owner address. USDC income lands at your owner address, not the signing address, so topping up means an off-chain hop: swap USDC→ALGO and transfer, or send ALGO manually.
- Fund the NFD's owner too, if it's a different account. ACME and IP sync write DNS records into your NFD, fee-paid by whoever owns the NFD: roughly one transaction pair per certificate renewal and one per IP change. That account needs its own small float, and the metric above doesn't watch it. See Which account signs the DNS writes.
Update pricing
Edit the rates
Change zs.default_pricing and/or the pricing block on any
zs.models[<id>] entry in your config. Rates are quoted in USD per million
tokens (see Serving models & pricing).
Restart
In-flight tickets keep their pinned rate; only new reserves see the new price.
Confirm
Check your live catalog on /v1/zs/details (or in the model picker) to
verify the new rate is being advertised.
Add or change a model
Declare it
Add a zs.models[<id>] entry. With zs.default_pricing set you don't need
a pricing block unless you want a per-model rate; otherwise add explicit rates
(input_rate: 0 / output_rate: 0 makes it free). Add a context
block too — it's recommended, and required for a local model.
Wire the backend
For a local model, add the corresponding llm.local.models entry (a node
hosts one local model; see
Serving models & pricing for multi-model hosts). For an
OpenAI-compatible or Vertex backend, availability is governed by the
backend's own model list; here you're just declaring pricing and capacity.
Restart
The change is picked up on restart. Clients see the new model on their next
discovery probe: about once a minute in the web app, and every 5 minutes by
default for zs-proxy (its
refresh_interval). See
Health & compatibility.
To stop serving a model, remove its entry and restart, so clients stop routing to it.
Update context windows and concurrency
All three are per-model edits that take effect on restart:
- Context window — edit
context_windowin the model'scontextblock. The proxy reads the new capacity from your details document on its next refresh. - Output ceiling — edit
max_output_tokensin the same block, keeping it belowcontext_window. If you raise the window, revisit this too: it does not scale with it, and a stale low ceiling silently truncates long answers. See Set an output ceiling. - Concurrency cap — set
max_active_ticketson the model to give it its own admission pool, or omit it to inherit the node default. Each model's in-flight count is tracked independently, so raising one model's cap doesn't touch the others. Raise it alongside real backend/VRAM headroom, or you'll just move a429 no_capacityinto an upstream error.
Rotate the signing key
The signing key is a hot key — rotate it whenever you'd rotate any production credential.
Create the new key
Generate a new mnemonic and address in your wallet.
Add it to the keystore
Add the new mnemonic to your keystore source (env var or secret manager) — keep the old one for now so in-flight tickets can still be settled.
Update the node record
From the operator dashboard, connect your owner wallet and update the node record's signing address to the new one. Only the operator owner can authorize this. See Registering on-chain.
Restart and verify
The node picks up the new on-chain signing address, matches it to the new mnemonic, and starts signing with it. Once old in-flight tickets have settled, remove the old mnemonic from the keystore.
Rotate the encryption key — nothing to do
There is no operator-managed encryption key. The node generates, signs, and rotates a short-lived recipient in memory on its own, with no file and no on-chain transaction. See Encryption & keys.
Rotate the owner account
From the operator dashboard, connect the
current owner wallet and update the operator record's owner address. Only
the current owner can authorize it, and every later rotation must come from the
new owner account. The node signs with the hot signing key, not the owner key,
so a running node needs no restart for an owner change, unless you've pinned
zs.owner_addr in config; then update it to match and restart.
Upgrade the node
Pull the new version
Fetch the new binary or container image.
Restart
Send SIGTERM (what systemd, Docker, and Kubernetes all do) and the node drains
gracefully; see Graceful shutdown. The settlement
driver reconciles any in-flight settlements against the chain at startup, so a
planned restart never loses ledger state.
Keep your node software current — a node on an incompatible protocol version is
still up and serving, but clients drop it from selection (see
Health & compatibility). After
upgrading, confirm your monitoring targets still point at the private
listener for /healthz, /livez, and /metrics (see Monitoring).
Graceful shutdown
On SIGTERM the node stops accepting new work and waits a bounded time for the
inference it already accepted to finish. It does that in three stages:
- Immediately it stops advertising models, refuses new reservations, sheds
new relay circuits, and reports unhealthy on
/healthz. Clients re-checking during your restart route to a different operator instead of paying to reserve against a node that is about to exit. - For
server.drain_grace(default20s) it still honors reservations issued before the shutdown began, so a request already on its way to you lands instead of losing the payment it has already committed. - For up to
server.drain_timeout(default5m) it waits for the inference already running to finish.
Then server.shutdown_timeout (default 30s) closes the remaining connections.
If you sent the SIGTERM, a second one skips the wait when you need the
process gone immediately. That shortcut applies only to a shutdown you
signalled. When the node starts draining on its own, because your operator was
evicted or unregistered on-chain or a listener failed, it ignores SIGTERM for
the rest of the drain so a routine stop can't truncate inference that's still
running, and it runs to drain_timeout. Use SIGKILL if you have to cut that
short.
Four things must line up or the drain is defeated:
- Your supervisor's kill deadline must be longer than the total drain —
drain_grace + drain_timeout + shutdown_timeout, 5m50s with the defaults. If you installed the service withzs-node install-service, this is already set for you (systemdTimeoutStopSec=400, launchdExitTimeOut=400). You only need to set it by hand for a hand-written unit, a container, or Kubernetes:terminationGracePeriodSeconds: 400,docker run --stop-timeout 400, composestop_grace_period: 400s. Watch the Docker defaults especially — both--stop-timeoutandstop_grace_perioddefault to 10 seconds, and Kubernetes to 30. Raise it everywhere if you raisedrain_timeout. - Whatever fronts the node needs a drain window too. A reverse proxy or
managed load balancer has its own connection-draining / deregistration delay,
and it will cut a response mid-stream while the node is still draining. Size it
to outlive
drain_grace + drain_timeout + shutdown_timeout, and never stop or reload the front door before the node has finished draining. Some platforms disable connection draining by default, which drops everything in flight the instant the node leaves the pool; check yours. The order you want is: node drain < front-door drain < platform kill deadline. - Your readiness probe must target the private listener, where
/healthzlives (see Monitoring). That's what pulls you out of the load balancer in stage 1. Without it, your own front door keeps forwarding requests to a node that is shutting down and answers them with a502or504. - Point your liveness probe at
/livez, not/healthz./healthzreports unhealthy for the whole drain, so a liveness check on it tells your orchestrator to restart the node mid-drain, killing the paid work the drain exists to finish. An ordinary restart hides this, but when the node drains on its own (because your operator was evicted or unregistered on-chain), nothing is terminating the container, so a failing liveness probe restarts it. If your node build has no/livezyet (it returns404), drop the liveness probe instead.
On managed Kubernetes, the platform caps your grace period. It enforces its
own ceiling during node upgrades and scale-downs regardless of what your Pod spec
asks for, so the whole drain budget has to fit underneath it; that ceiling is the
real limit on drain_timeout. Preemptible / Spot instances get a drastically
shorter ceiling (tens of seconds), too short for even the default drain, which
makes them a poor fit for a serving node. Check your platform's published limits
before raising drain_timeout, and if it offers an annotation to exempt
long-running Pods from autoscaler eviction, use it.
A node that dies seconds after issuing a reservation leaves that payer's USDC locked in escrow until the inactivity refund window elapses, for work you never performed. Rolling a fleet without draining does this to every payer who reserved in the last seconds of each node's life.
Disputed tickets need you
A payer who disagrees with a charge can freeze the ticket on-chain. Frozen is terminal: no contract path leaves that state, and the funds stay in escrow until off-chain arbitration produces a resolution. It is the one settlement outcome the node cannot resolve on its own.
When the settlement driver meets one — either on its pre-flight read, or because the settle transaction reverts — it logs a single ERROR and stops retrying:
settlement: ticket FROZEN (payer disputed); off-chain arbitration required
ticket_id=… node_claimed_microalgos=… onchain_pending_amount=…
onchain_pending_by=… onchain_max_price=…
The matching ledger row is marked failed. The same line is emitted when a
force-finalization reverts for the same reason.
Alert on this message. Both signed claims are now on-chain for whoever
arbitrates. The node_claimed_microalgos and onchain_* fields in the line
record what each side claimed; an arbitrator will want them.
Back up what matters
-
The signing mnemonic — back it up like any wallet seed (secret manager, paper, hardware safe). There's no encryption-key file to back up; that key is ephemeral and disposable.
-
The settlement ledger — the local SQLite database at
zs.settlement_db_path. The contract's inactivity backstop protects payer funds even without it, but the ledger is your only record of in-flight obligations. A daily copy is plenty:sqlite3 /var/lib/zs-node/settlement.db \".backup /backups/settlement-$(date +%F).db"