Skip to main content

Operations

Most config changes take effect on restart. In-flight tickets always carry the rate and terms they were quoted at, so a restart never changes the price of a request already under way — only new reserves see new settings. Run the node under a supervisor (systemd, Docker --restart, an orchestrator) so "restart" is a one-liner; see Installation.

Top up the signing account​

Your hot signing account fee-pays every on-chain action the node signs: every settlement, plus the base-URL update if you've enabled ip_sync's URL sync. It needs an ALGO float on top of the account's own minimum balance (0.1 ALGO for a plain account). Budget roughly 5 ALGO + minimum balance + (daily request volume × ~7,000 µALGO) + a few days of buffer; Staking & economics breaks down the per-request cost.

  • Keep at least 5 ALGO spendable. Spendable means the balance minus Algorand's minimum balance. Proxies and the web client check it every few minutes and stop sending requests to a node below 5 ALGO spendable — the node keeps running, it just gets no traffic until you top it up.
  • Watch it. The node polls the balance and publishes the spendable amount as zs_signing_balance_algos, warns when it drops below your threshold (default 10 ALGO), logs an error once it falls under the 5 ALGO routing floor, and refuses to start below 1 ALGO. Alert on that metric or the signing-address ALGO balance is low log line — see Monitoring.
  • Raise the warning on a busy node. The gap between 10 and 5 ALGO covers about 700 paid requests. A node that serves several requests a second can use that up between two balance checks, so the warning and the error arrive together. Set warn_below_algos to at least 5 + (paid requests per second × 2.1) at the default 5-minute check, plus enough to cover the time it takes you to refill — see zs.signing_balance.
  • Refill from the owner address. USDC income lands at your owner address, not the signing address, so topping up means an off-chain hop: swap USDC→ALGO and transfer, or send ALGO manually.
  • Fund the NFD's owner too, if it's a different account. ACME and IP sync write DNS records into your NFD, fee-paid by whoever owns the NFD: roughly one transaction pair per certificate renewal and one per IP change. That account needs its own small float, and the metric above doesn't watch it. See Which account signs the DNS writes.

Update pricing​

Edit the rates​

Change zs.default_pricing and/or the pricing block on any zs.models[<id>] entry in your config. Rates are quoted in USD per million tokens (see Serving models & pricing).

Restart​

In-flight tickets keep their pinned rate; only new reserves see the new price.

Confirm​

Check your live catalog on /v1/zs/details (or in the model picker) to verify the new rate is being advertised.

Add or change a model​

Declare it​

Add a zs.models[<id>] entry. With zs.default_pricing set you don't need a pricing block unless you want a per-model rate; otherwise add explicit rates (input_rate: 0 / output_rate: 0 makes it free). Add a context block too — it's recommended, and required for a local model.

Wire the backend​

For a local model, add the corresponding llm.local.models entry (a node hosts one local model; see Serving models & pricing for multi-model hosts). For an OpenAI-compatible or Vertex backend, availability is governed by the backend's own model list; here you're just declaring pricing and capacity.

Restart​

The change is picked up on restart. Clients see the new model on their next discovery probe: about once a minute in the web app, and every 5 minutes by default for zs-proxy (its refresh_interval). See Health & compatibility.

To stop serving a model, remove its entry and restart, so clients stop routing to it.

Update context windows and concurrency​

All three are per-model edits that take effect on restart:

  • Context window — edit context_window in the model's context block. The proxy reads the new capacity from your details document on its next refresh.
  • Output ceiling — edit max_output_tokens in the same block, keeping it below context_window. If you raise the window, revisit this too: it does not scale with it, and a stale low ceiling silently truncates long answers. See Set an output ceiling.
  • Concurrency cap — set max_active_tickets on the model to give it its own admission pool, or omit it to inherit the node default. Each model's in-flight count is tracked independently, so raising one model's cap doesn't touch the others. Raise it alongside real backend/VRAM headroom, or you'll just move a 429 no_capacity into an upstream error.

Rotate the signing key​

The signing key is a hot key — rotate it whenever you'd rotate any production credential.

Create the new key​

Generate a new mnemonic and address in your wallet.

Add it to the keystore​

Add the new mnemonic to your keystore source (env var or secret manager) — keep the old one for now so in-flight tickets can still be settled.

Update the node record​

From the operator dashboard, connect your owner wallet and update the node record's signing address to the new one. Only the operator owner can authorize this. See Registering on-chain.

Restart and verify​

The node picks up the new on-chain signing address, matches it to the new mnemonic, and starts signing with it. Once old in-flight tickets have settled, remove the old mnemonic from the keystore.

Rotate the encryption key — nothing to do​

There is no operator-managed encryption key. The node generates, signs, and rotates a short-lived recipient in memory on its own, with no file and no on-chain transaction. See Encryption & keys.

Rotate the owner account​

From the operator dashboard, connect the current owner wallet and update the operator record's owner address. Only the current owner can authorize it, and every later rotation must come from the new owner account. The node signs with the hot signing key, not the owner key, so a running node needs no restart for an owner change, unless you've pinned zs.owner_addr in config; then update it to match and restart.

Upgrade the node​

Pull the new version​

Fetch the new binary or container image.

Restart​

Send SIGTERM (what systemd, Docker, and Kubernetes all do) and the node drains gracefully; see Graceful shutdown. The settlement driver reconciles any in-flight settlements against the chain at startup, so a planned restart never loses ledger state.

Keep your node software current — a node on an incompatible protocol version is still up and serving, but clients drop it from selection (see Health & compatibility). After upgrading, confirm your monitoring targets still point at the private listener for /healthz, /livez, and /metrics (see Monitoring).

Graceful shutdown​

On SIGTERM the node stops accepting new work and waits a bounded time for the inference it already accepted to finish. It does that in three stages:

  1. Immediately it stops advertising models, refuses new reservations, sheds new relay circuits, and reports unhealthy on /healthz. Clients re-checking during your restart route to a different operator instead of paying to reserve against a node that is about to exit.
  2. For server.drain_grace (default 20s) it still honors reservations issued before the shutdown began, so a request already on its way to you lands instead of losing the payment it has already committed.
  3. For up to server.drain_timeout (default 5m) it waits for the inference already running to finish.

Then server.shutdown_timeout (default 30s) closes the remaining connections.

If you sent the SIGTERM, a second one skips the wait when you need the process gone immediately. That shortcut applies only to a shutdown you signalled. When the node starts draining on its own, because your operator was evicted or unregistered on-chain or a listener failed, it ignores SIGTERM for the rest of the drain so a routine stop can't truncate inference that's still running, and it runs to drain_timeout. Use SIGKILL if you have to cut that short.

warning

Four things must line up or the drain is defeated:

  • Your supervisor's kill deadline must be longer than the total drain — drain_grace + drain_timeout + shutdown_timeout, 5m50s with the defaults. If you installed the service with zs-node install-service, this is already set for you (systemd TimeoutStopSec=400, launchd ExitTimeOut=400). You only need to set it by hand for a hand-written unit, a container, or Kubernetes: terminationGracePeriodSeconds: 400, docker run --stop-timeout 400, compose stop_grace_period: 400s. Watch the Docker defaults especially — both --stop-timeout and stop_grace_period default to 10 seconds, and Kubernetes to 30. Raise it everywhere if you raise drain_timeout.
  • Whatever fronts the node needs a drain window too. A reverse proxy or managed load balancer has its own connection-draining / deregistration delay, and it will cut a response mid-stream while the node is still draining. Size it to outlive drain_grace + drain_timeout + shutdown_timeout, and never stop or reload the front door before the node has finished draining. Some platforms disable connection draining by default, which drops everything in flight the instant the node leaves the pool; check yours. The order you want is: node drain < front-door drain < platform kill deadline.
  • Your readiness probe must target the private listener, where /healthz lives (see Monitoring). That's what pulls you out of the load balancer in stage 1. Without it, your own front door keeps forwarding requests to a node that is shutting down and answers them with a 502 or 504.
  • Point your liveness probe at /livez, not /healthz. /healthz reports unhealthy for the whole drain, so a liveness check on it tells your orchestrator to restart the node mid-drain, killing the paid work the drain exists to finish. An ordinary restart hides this, but when the node drains on its own (because your operator was evicted or unregistered on-chain), nothing is terminating the container, so a failing liveness probe restarts it. If your node build has no /livez yet (it returns 404), drop the liveness probe instead.
info

On managed Kubernetes, the platform caps your grace period. It enforces its own ceiling during node upgrades and scale-downs regardless of what your Pod spec asks for, so the whole drain budget has to fit underneath it; that ceiling is the real limit on drain_timeout. Preemptible / Spot instances get a drastically shorter ceiling (tens of seconds), too short for even the default drain, which makes them a poor fit for a serving node. Check your platform's published limits before raising drain_timeout, and if it offers an annotation to exempt long-running Pods from autoscaler eviction, use it.

A node that dies seconds after issuing a reservation leaves that payer's USDC locked in escrow until the inactivity refund window elapses, for work you never performed. Rolling a fleet without draining does this to every payer who reserved in the last seconds of each node's life.

Disputed tickets need you​

A payer who disagrees with a charge can freeze the ticket on-chain. Frozen is terminal: no contract path leaves that state, and the funds stay in escrow until off-chain arbitration produces a resolution. It is the one settlement outcome the node cannot resolve on its own.

When the settlement driver meets one — either on its pre-flight read, or because the settle transaction reverts — it logs a single ERROR and stops retrying:

settlement: ticket FROZEN (payer disputed); off-chain arbitration required
ticket_id=… node_claimed_microalgos=… onchain_pending_amount=…
onchain_pending_by=… onchain_max_price=…

The matching ledger row is marked failed. The same line is emitted when a force-finalization reverts for the same reason.

Alert on this message. Both signed claims are now on-chain for whoever arbitrates. The node_claimed_microalgos and onchain_* fields in the line record what each side claimed; an arbitrator will want them.

Back up what matters​

  • The signing mnemonic — back it up like any wallet seed (secret manager, paper, hardware safe). There's no encryption-key file to back up; that key is ephemeral and disposable.

  • The settlement ledger — the local SQLite database at zs.settlement_db_path. The contract's inactivity backstop protects payer funds even without it, but the ledger is your only record of in-flight obligations. A daily copy is plenty:

    sqlite3 /var/lib/zs-node/settlement.db \
    ".backup /backups/settlement-$(date +%F).db"