Skip to main content

Troubleshooting

info

Before working through these by hand, run zs-node doctor. It checks your config against your live backend and the chain in one pass, and it catches the whole class of problem where your config advertises something your backend won't actually do, which produces client-visible failures your own logs never mention. Add --trace=/tmp/doctor.jsonl for a raw dump of every request and response. See the Node CLI reference.

Node won't start: missing signing address​

The node refuses to start because the signing address it expects doesn't match any mnemonic it loaded. The address it wants comes from zs.signing_addr (or, if you didn't set that, your on-chain operator record); the mnemonics come from the environment or your secret manager. Check, in order:

  • The mnemonic actually reached the process. A variable set in your shell isn't set in the service's environment. For systemd, confirm the EnvironmentFile= points at a readable file that defines it.
  • The mnemonic derives the right address. Decode it locally (any wallet that shows the derived address will do) and compare.
  • The on-chain record names the address you think it does. Open the operator dashboard and check the node record's signing address.

See Encryption & keys for how the signing mnemonic is provisioned.

Node won't start: ACME requires the operator's NFD on chain​

tls.mode=acme requires the operator's NFD on chain — set zs.escrow_app_id and
ensure the operator is registered with an NFD, or set tls.mode=off

The node resolved an NFD app id of 0, and ACME has no name to get a certificate for. Three ways to satisfy it:

  • Set zs.escrow_app_id so the node can read your operator box, and make sure that operator record actually has an NFD linked (the field is optional at registration; see Your name and avatar).
  • Set zs.nfd_app_id explicitly. This always wins over the value read from chain, and is how you point a node at its own NFD rather than the operator's. See NFD identity.
  • Set tls.mode: off and terminate TLS somewhere else.

algod.network must also be testnet or mainnet; localnet has no public DNS suffix, so ACME is refused there outright.

Node won't start with ip_sync enabled and manual certificates​

server.tls.ip_sync.url.enabled requires server.tls.mode="acme"

ip_sync itself works under tls.mode: manual, but its URL sync defaults to on, and that half requires acme. Set it off explicitly:

server:
tls:
mode: manual
ip_sync:
enabled: true
url:
enabled: false

You then manage the base URL yourself from the operator dashboard.

Certificate never issues: keystore has no key for owner​

nfd update "yourname.algo": keystore has no key for owner ABCD…XYZ

An NFD's DNS records can only be written by the account that owns the NFD, and your node's keystore doesn't hold that key. This is the common case: registering your operator pins the NFD to your cold owner wallet, while the node holds only its hot signing key. The node starts and serves normally; the failure only surfaces on the first certificate renewal or IP change, as a warning that's easy to scroll past.

Your options (a second NFD held by the signing address, transferring the existing one, or not using NFD DNS at all) are laid out in Which account signs the DNS writes.

HTTPS fails right after start, then works​

Certificate provisioning runs in the background. The node logs listener up scheme=https and begins accepting connections before the first certificate exists, so TLS handshakes fail until it lands, while /healthz on the plaintext private listener reads green the whole time.

Give it a few minutes: the DNS-01 challenge has to travel from an on-chain confirmation through the NFD indexer to the authoritative zone, which is why propagation_delay waits 30s before polling even starts. Watch the node's log rather than the health endpoint.

If it never succeeds and the logs mention Let's Encrypt rate limits, check server.tls.acme.cache_dir. It defaults to ./tls-cache, relative to the working directory, and holds your ACME account key. A container without a volume, or a systemd unit whose working directory moves, discards it and registers a fresh account on every boot until the rate limiter cuts you off. Point it at an absolute, durable path.

Two nodes keep overwriting each other's DNS record​

Symptoms: clients reach a different one of your nodes each time; your NFD shows an on-chain update roughly every minute, per node; certificates occasionally fail to renew for no obvious reason.

Two nodes under one operator are both publishing at the same record, almost always both left at the NFD apex with zs.nfd_record_name unset. Each believes it owns that A record and rewrites it, forever. Because each write replaces the whole u.dns document, they can also delete each other's unrelated records, including a live _acme-challenge TXT mid-issuance.

Give each node its own zs.nfd_record_name (or its own NFD via zs.nfd_app_id), restart both, then delete the stale apex record they were fighting over:

zs-node nfd-dns delete --name=@ --type=a

See One label per node.

Reserve fails: payment not verified​

Reserves are rejected with a payment-verification error, and the node's logs show it can't find the payer's transaction. This almost always means the algod your node talks to isn't observing pending transactions, so between the payer broadcasting the escrow-open and the network confirming it, the node looks for that transaction and comes up empty.

Give the node an algod that sees the mempool. If you run your own non-participation node, set ForceFetchTransactions: true in its config.json and restart it. See Installation → your algod must observe the mempool for the details and the valid algod configurations.

warning

Tightening zs.mempool_poll_timeout / mempool_poll_interval only changes how the node spaces its lookups; it can't create visibility that isn't there.

The oracle is unavailable and prices show 0​

/v1/zs/details reports oracle_healthy: false and algo_usd_price reads 0. This is non-fatal: your node still issues tickets and serves requests normally. For that pricing cycle, the ALGO-network-fee figure clients display in USD is dropped, and the (optional) per-request minimum ALGO charge falls to zero. Token pricing and in-flight tickets are unaffected, since those rates were pinned at reserve time.

Common causes:

  • The ALGO/USD price feed is rate-limiting or down.
  • Outbound network access to the feed is blocked.
  • The node host's clock has drifted (the oracle compares timestamps).

If you don't charge the optional ALGO minimum and your clients don't need the fee-in-USD figure, an unavailable oracle is harmless. Otherwise, widen zs.oracle.max_staleness cautiously, or point zs.oracle.endpoint at a mirror you control (see Configuration).

Reserve returns a no-capacity 429​

A model's admission pool is full: every slot is held by a reserved or in-flight request, so the node refuses new reserves with 429 no_capacity and a short Retry-After. Confirm it on metrics: zs_reserve_inflight_slots{model} has reached zs_reserve_capacity_slots{model} while zs_reserve_capacity_denied_total{model} climbs. Then either:

  • Raise the model's max_active_tickets, and verify your backend and VRAM can actually handle the added concurrency (see Serving models).
  • Lower zs.ticket_ttl so unused reservations recycle faster.
  • Add another node serving the model so clients can spread the load.

This is separate from the rate limiter's 429 rate_limited, which throttles per-IP or per-account (see Configuration), and from 429 payer_slot_cap below.

Reserve returns a payer-slot-cap 429​

A single payer is already holding its zs.payer_slot_cap share of a model's pool and asked for another concurrent slot. The node refuses with 429 payer_slot_cap, a distinct code from no_capacity, which means the whole pool is full across all payers.

Tell them apart on metrics: zs_reserve_payer_slot_cap_denied_total{model} climbs while zs_reserve_inflight_slots{model} is still below the model's capacity. Then:

  • If it's abuse (one funded account trying to monopolize a model), the cap is working as intended. No action needed.
  • If it's a legitimate high-parallelism caller being pinched, raise zs.payer_slot_cap globally or per-model, or set it to 0 to disable the per-payer cap and rely on max_active_tickets plus the reserve_per_account rate limit alone.
warning

The cap only resists a determined attacker while payer-signature verification is on (the default). Without it the payer address is spoofable, so an attacker simply rotates it and sidesteps the cap entirely.

Reserve returns a 403 free-quota error​

Only reachable on a free model: the payer has used its daily allowance of free tickets and is refused with 403 free_quota_exhausted until the window resets. It is a 403, not a 429, so clients don't retry or fail over to another operator; neither would help. Nothing on your node is wrong, and there is no knob to raise; see Free models are rationed by the contract. Confirm on zs_reserve_free_quota_total{outcome="exhausted"}.

What to check instead: if that same counter's cold_miss_budget or chain_error labels are climbing, those refusals are 503s, not 403s. The allowance read itself is failing (algod unreachable, or a rotating-address flood exhausting the lookup budget), and legitimate free reserves are being turned away for an infrastructure reason. Treat it as an algod-connectivity alert.

Tool loops end early, or reserves are refused​

Answers arrive after one tool round when they should have taken several, or prompts come back refused with input_budget_exceeded. Both mean a request's measured input exceeded what its ticket reserved; see zs.reserve. Split them apart on the metric, which labels every occurrence by what happened:

curl -s localhost:9091/metrics | grep zs_reserve_input_budget_over_total
  • outcome="tool_cutoff" — loops are hitting the ceiling. Your tool_headroom_per_iteration (default 4000) is undersized for the tools you serve: one web_read at the default 64 KiB max_bytes is roughly 16,000 tokens. Raise it, or lower zs.builtin_tools.web_read.max_bytes so results fit the budget you advertise. Lowering zs.builtin_tools.max_iterations without raising the headroom makes this worse, because the reserve clients compute is the product of the two.
  • outcome="rejected" — plain requests are being refused before any upstream call, at zero cost to the payer. The paired WARN line reserve input budget exceeded on initial body carries a code. context_length_exceeded is a prompt whose size estimate is far past the model's window, which honest users send; at the context cap that means more than three times the budget, which whitespace-heavy or \u-escaped text can reach before the real window. input_budget_exceeded is a caller that reserved less than its prompt: an honest client sizes with the same bound your node measures with, so isolated hits are one caller sizing badly, and a broad, sustained rate across many payers is more likely a client older than your node's protocol version.
  • outcome="forwarded" — not a problem. These were long prompts whose reserve was already at the context cap, so your node let your backend decide whether they fit instead of refusing them on the size estimate. The ones your backend refuses show up on zs_provider_upstream_rejection_total as code="context_length_exceeded", and the payer gets that code at zero cost. With enforcement on, a forwarded request is not offered your built-in tools, since one round alone can fill the window. Forwarding happens with enforcement off too, so these never count as monitored.
  • outcome="monitored" — you have enforce_input_budget: false. Nothing is being refused or cut off; these requests were served at your expense.

To serve while you tune, set enforce_input_budget: false (monitor mode), but only temporarily. Enforcement is what stops a payer under-declaring input_count and having you serve the difference, and the node logs a startup WARN for as long as it's off.

Clients say tools never fire, but the model supports them​

Three different faults look identical from the outside. zs-node doctor tells them apart: it reports tools accepted and tool calls emitted as separate lines, because they have different fixes.

  • The backend rejects tools[] outright. On vLLM or SGLang that is a serving flag, not a model limitation: you need --enable-auto-tool-choice together with a --tool-call-parser matching the model's chat template. llama-server needs --jinja with a tool-declaring template. If your config also says context.tool_use: true, doctor reports it as a failure: the node is advertising a capability every client then gets a 4xx for.
  • The backend accepts tools[] and never returns a call. Almost always a --tool-call-parser that doesn't match the template: the request is valid, the model emits a call in its own format, and the runtime doesn't recognise it.
  • The backend accepts tools[] but rejects tool_choice: "required". Normal tool calling works; only clients that force a specific call get a 4xx. Doctor reports this as a warning rather than a failure, and says which form worked.
zs-node doctor --models=<your-model-id> --trace=/tmp/tools.jsonl
jq -c 'select(.step|startswith("tool_use")) | {step,status,response_body}' /tmp/tools.jsonl

Reasoning shows up in the backend but not through the node​

Your backend visibly produces reasoning, but clients don't see it, or the model loses its thread across a tool call. Two independent settings under context.reasoning are involved.

context.reasoning.supported gates whether the node replays reasoning back to the model between tool-loop iterations. A reasoning model declared as supported: false still streams its reasoning to the client, but stops having it replayed, so it re-plans from scratch each iteration. This most often affects OpenAI's o-series and GPT-5 behind openai_passthrough, where nothing in the model list announces that the model reasons.

context.reasoning.allowed_efforts is what clients are offered, and the node forwards a client's chosen effort to your backend as reasoning_effort. Plenty of runtimes emit reasoning natively while rejecting that field as an unknown argument, so advertising an effort your backend doesn't take turns into a 400 for the client that picks it. If your backend is in that group, declare supported: true with no allowed_efforts.

zs-node doctor reports reasoning emitted and reasoning_effort accepted separately, and flags either one that disagrees with your config.

Doctor says a model rejects image input, but it doesn't​

zs-node doctor establishes image support by sending a small test image. If your backend refuses that particular image (a minimum-dimension rule, a format restriction, a decode failure), the exam quotes the upstream's own message alongside the finding. Read it. Only a complaint about the model ("this model does not support image input") means the model lacks image support. A complaint about the image ("invalid image", "dimensions below the minimum") does not.

The probe tries two images and refuses to record a negative from wording it can't attribute to the capability, so this should be rare. If you see it anyway for a model you know takes images, --trace has the exact exchange:

zs-node doctor --models=<your-model-id> --trace=/tmp/vision.jsonl
jq -c 'select(.step|startswith("vision")) | {step,status,response_body}' /tmp/vision.jsonl

Report the upstream message; the classifier is a list of known wordings. In the meantime the node advertises what context.input_modalities says, and doctor's finding is advisory.

The local model keeps crashing​

When you run a local model, its supervisor restarts a crashed engine up to five times in a minute, then gives up. Check:

  • The model file path is correct and readable by the user the node runs as.
  • gpu_layers doesn't exceed your VRAM. Drop it, or use a smaller quantization.
  • context_window × parallel_slots fits in VRAM. That product is your KV cache, and it is routinely larger than the weights. Do the arithmetic in What the inference backend needs.
  • Your extra_args don't re-set values the node derives itself (context size, parallelism, continuous batching). Those are rejected at validation, and if one slips through it silently breaks the per-session contract.
  • On dual-GPU without NVLink, layer-split mode is usually faster than row-split.

The node is healthy but gets no traffic​

Probes succeed, /v1/zs/details looks right, and yet no requests arrive. Check the signing account first: proxies and the web client skip any node whose signing account has less than 5 ALGO spendable (its balance minus Algorand's minimum balance), because that account pays the fees of every request.

  • Look at the zs_signing_balance_algos metric, or for a signing-address spendable ALGO is below the routing floor error in the log. zs-node doctor reports it too.
  • Top up the signing account (see Operations). Clients pick up the new balance within about five minutes.

If the float is fine, work through the rest of Keeping your node routable.

Settlement entries stuck in settling​

A settle transaction was submitted but algod didn't confirm it in time, so the entry flips back to be retried automatically. Occasional churn here is normal. Entries that stay stuck point at one of:

  • algod is unreachable — the node can't submit or confirm.
  • The signing account is out of ALGO — no fee budget to settle. Top it up (see Operations).
  • Network congestion — confirmations are slow but eventual.

A late confirmation isn't lost: the node reconciles in-flight settlements against the chain when it restarts, so a planned restart never drops ledger state. If you need to inspect or hand-repair a specific ticket, the settlement admin subcommand is the escape hatch.

Out of VRAM at startup​

Your local engine won't allocate. If it doesn't fit lists what to lower, in order of what you give up: the model's context_window, zs.max_active_tickets or llm.local.parallel_slots, the KV cache precision, and gpu_layers, before falling back to a smaller quantization or model.

Clients get 502 or 504 errors, but the node looks fine​

Anything you put in front of the node (CDN, ingress controller, reverse proxy, load balancer) can answer for it, and when it does the error comes from that layer, not from the node. Clients can tell the difference: the node's own refusals always carry a structured error code, and a gateway's don't.

While a request is still being set up, the client retries and reroutes automatically, so brief blips never reach the user. They still cost you the traffic, and a client that keeps failing against you there demotes you in its routing for several minutes. Once a request is paid for and sent, there is no retry: a gateway error during inference goes straight back to the caller as a failed request. Repeated timeouts on requests forwarded through you also demote you if your peers carry the same nodes without them. That takes more failures than an ordinary error, and it fades once the timeouts stop.

Two causes, in order of likelihood:

Restart gaps. If your front door serves errors while the node is down, use graceful shutdown and point your readiness probe at the private listener; see Graceful shutdown. A draining node stops advertising before the gap opens, so clients route elsewhere and never meet the error.

Origin timeouts shorter than your model's thinking time. A reasoning-heavy request can send nothing for a long stretch before the first token. A gateway that gives up during that silence reports 504 (Cloudflare: 524) even though the node was working normally. How long a silence you have to survive depends on how the caller asked:

RequestKeepalive during silenceYour origin read/idle timeout must exceed
StreamedYes: an invisible keepalive every server.sse_keepalive_interval (15s by default), starting as soon as the response headers are sent, before the model is called.server.sse_keepalive_interval.
Buffered (non-streaming)No. Nothing is sent until generation completes.Your slowest complete generation, a far bigger number. This is the usual cause of a 504 that lands almost exactly on a round minute.

You can't control which mode callers use, so size for the buffered case if you want to serve them reliably. Defaults that are commonly too low:

GatewayDefault
nginx proxy_read_timeout60s
AWS ALB idle timeout60s
Cloudflare~100s, and not raisable except on Enterprise. If your slowest generation exceeds it, that's a hard ceiling to design around.

Errors in the 520–527 range are Cloudflare-specific and always mean the problem is between Cloudflare and your node: 521 origin down, 522 connection timed out, 523 origin unreachable, 524 origin timed out.

The node can't reach algod, the price feed, or the backend​

Almost always an egress firewall. The node needs outbound HTTPS to:

  • your algod endpoint,
  • the ALGO/USD price feed (unless you've overridden zs.oracle.endpoint),
  • your inference backend's URL,
  • and, if you use tls.mode: acme or server.tls.ip_sync, the NFDomains API (api.nf.domains / api.testnet.nf.domains) and api64.ipify.org.

From inside the container or host, curl -v against each is the fastest way to find which one is blocked.

The node only listens on localhost​

server.listen defaults to loopback (127.0.0.1:9090), so the node isn't exposed yet. Set NODE_SERVER_LISTEN=:9090 to bind all interfaces, and serve HTTPS with tls.mode: acme or behind a single reverse proxy, never a load balancer. See The two listeners, Installation, and Reaching your node. The port you bind is the port published in your base URL on chain, so changing it means updating that record too. The private listener (server.private_listen, default 127.0.0.1:9091) is separate: override it if your probes or scraper live in another container.