Troubleshooting
Before working through these by hand, run zs-node doctor. It checks your
config against your live backend and the chain in one pass, and it catches the
whole class of problem where your config advertises something your backend won't
actually do, which produces client-visible failures your own logs never mention.
Add --trace=/tmp/doctor.jsonl for a raw dump of every request and response.
See the Node CLI reference.
Node won't start: missing signing address
The node refuses to start because the signing address it expects doesn't match
any mnemonic it loaded. The address it wants comes from zs.signing_addr
(or, if you didn't set that, your on-chain operator record); the mnemonics come
from the environment or your secret manager. Check, in order:
- The mnemonic actually reached the process. A variable set in your shell
isn't set in the service's environment. For systemd, confirm the
EnvironmentFile=points at a readable file that defines it. - The mnemonic derives the right address. Decode it locally (any wallet that shows the derived address will do) and compare.
- The on-chain record names the address you think it does. Open the operator dashboard and check the node record's signing address.
See Encryption & keys for how the signing mnemonic is provisioned.
Node won't start: ACME requires the operator's NFD on chain
tls.mode=acme requires the operator's NFD on chain — set zs.escrow_app_id and
ensure the operator is registered with an NFD, or set tls.mode=off
The node resolved an NFD app id of 0, and ACME has no name to get a certificate
for. Three ways to satisfy it:
- Set
zs.escrow_app_idso the node can read your operator box, and make sure that operator record actually has an NFD linked (the field is optional at registration; see Your name and avatar). - Set
zs.nfd_app_idexplicitly. This always wins over the value read from chain, and is how you point a node at its own NFD rather than the operator's. See NFD identity. - Set
tls.mode: offand terminate TLS somewhere else.
algod.network must also be testnet or mainnet; localnet has no public DNS
suffix, so ACME is refused there outright.
Node won't start with ip_sync enabled and manual certificates
server.tls.ip_sync.url.enabled requires server.tls.mode="acme"
ip_sync itself works under tls.mode: manual, but its URL sync defaults to
on, and that half requires acme. Set it off explicitly:
server:
tls:
mode: manual
ip_sync:
enabled: true
url:
enabled: false
You then manage the base URL yourself from the operator dashboard.
Certificate never issues: keystore has no key for owner
nfd update "yourname.algo": keystore has no key for owner ABCD…XYZ
An NFD's DNS records can only be written by the account that owns the NFD, and your node's keystore doesn't hold that key. This is the common case: registering your operator pins the NFD to your cold owner wallet, while the node holds only its hot signing key. The node starts and serves normally; the failure only surfaces on the first certificate renewal or IP change, as a warning that's easy to scroll past.
Your options (a second NFD held by the signing address, transferring the existing one, or not using NFD DNS at all) are laid out in Which account signs the DNS writes.
HTTPS fails right after start, then works
Certificate provisioning runs in the background. The node logs
listener up scheme=https and begins accepting connections before the first
certificate exists, so TLS handshakes fail until it lands, while /healthz on
the plaintext private listener reads green the whole time.
Give it a few minutes: the DNS-01 challenge has to travel from an on-chain
confirmation through the NFD indexer to the authoritative zone, which is why
propagation_delay waits 30s before polling even starts. Watch the node's log
rather than the health endpoint.
If it never succeeds and the logs mention Let's Encrypt rate limits, check
server.tls.acme.cache_dir. It defaults to ./tls-cache, relative to the
working directory, and holds your ACME account key. A container without a
volume, or a systemd unit whose working directory moves, discards it and
registers a fresh account on every boot until the rate limiter cuts you off.
Point it at an absolute, durable path.
Two nodes keep overwriting each other's DNS record
Symptoms: clients reach a different one of your nodes each time; your NFD shows an on-chain update roughly every minute, per node; certificates occasionally fail to renew for no obvious reason.
Two nodes under one operator are both publishing at the same record, almost
always both left at the NFD apex with zs.nfd_record_name unset. Each believes
it owns that A record and rewrites it, forever. Because each write replaces the
whole u.dns document, they can also delete each other's unrelated records,
including a live _acme-challenge TXT mid-issuance.
Give each node its own zs.nfd_record_name (or its own NFD via
zs.nfd_app_id), restart both, then delete the stale apex record they were
fighting over:
zs-node nfd-dns delete --name=@ --type=a
See One label per node.
Reserve fails: payment not verified
Reserves are rejected with a payment-verification error, and the node's logs show it can't find the payer's transaction. This almost always means the algod your node talks to isn't observing pending transactions, so between the payer broadcasting the escrow-open and the network confirming it, the node looks for that transaction and comes up empty.
Give the node an algod that sees the mempool. If you run your own
non-participation node, set ForceFetchTransactions: true in its config.json
and restart it. See
Installation → your algod must observe the mempool
for the details and the valid algod configurations.
Tightening zs.mempool_poll_timeout / mempool_poll_interval only changes
how the node spaces its lookups; it can't create visibility that isn't there.
The oracle is unavailable and prices show 0
/v1/zs/details reports oracle_healthy: false and algo_usd_price reads
0. This is non-fatal: your node still issues tickets and serves requests
normally. For that pricing cycle, the ALGO-network-fee figure clients display in
USD is dropped, and the (optional) per-request minimum ALGO charge falls to
zero. Token pricing and in-flight tickets are unaffected, since those rates were
pinned at reserve time.
Common causes:
- The ALGO/USD price feed is rate-limiting or down.
- Outbound network access to the feed is blocked.
- The node host's clock has drifted (the oracle compares timestamps).
If you don't charge the optional ALGO minimum and your clients don't need the
fee-in-USD figure, an unavailable oracle is harmless. Otherwise, widen
zs.oracle.max_staleness cautiously, or point zs.oracle.endpoint at a
mirror you control (see Configuration).
Reserve returns a no-capacity 429
A model's admission pool is full: every slot is held by a reserved or in-flight
request, so the node refuses new reserves with 429 no_capacity and a short
Retry-After. Confirm it on metrics:
zs_reserve_inflight_slots{model} has reached
zs_reserve_capacity_slots{model} while
zs_reserve_capacity_denied_total{model} climbs. Then either:
- Raise the model's
max_active_tickets, and verify your backend and VRAM can actually handle the added concurrency (see Serving models). - Lower
zs.ticket_ttlso unused reservations recycle faster. - Add another node serving the model so clients can spread the load.
This is separate from the rate limiter's 429 rate_limited, which throttles
per-IP or per-account (see Configuration), and from
429 payer_slot_cap below.
Reserve returns a payer-slot-cap 429
A single payer is already holding its
zs.payer_slot_cap share of a
model's pool and asked for another concurrent slot. The node refuses with
429 payer_slot_cap, a distinct code from no_capacity, which means the
whole pool is full across all payers.
Tell them apart on metrics:
zs_reserve_payer_slot_cap_denied_total{model} climbs while
zs_reserve_inflight_slots{model} is still below the model's capacity. Then:
- If it's abuse (one funded account trying to monopolize a model), the cap is working as intended. No action needed.
- If it's a legitimate high-parallelism caller being pinched, raise
zs.payer_slot_capglobally or per-model, or set it to0to disable the per-payer cap and rely onmax_active_ticketsplus thereserve_per_accountrate limit alone.
The cap only resists a determined attacker while payer-signature verification is on (the default). Without it the payer address is spoofable, so an attacker simply rotates it and sidesteps the cap entirely.
Reserve returns a 403 free-quota error
Only reachable on a free model: the payer has used its daily allowance of
free tickets and is refused with 403 free_quota_exhausted until the window
resets. It is a 403, not a 429, so clients don't retry or fail over to
another operator; neither would help. Nothing on your node is wrong, and there
is no knob to raise; see
Free models are rationed by the contract.
Confirm on zs_reserve_free_quota_total{outcome="exhausted"}.
What to check instead: if that same counter's cold_miss_budget or
chain_error labels are climbing, those refusals are 503s, not 403s. The
allowance read itself is failing (algod unreachable, or a rotating-address flood
exhausting the lookup budget), and legitimate free reserves are being turned
away for an infrastructure reason. Treat it as an algod-connectivity alert.
Tool loops end early, or reserves are refused
Answers arrive after one tool round when they should have taken several, or
prompts come back refused with input_budget_exceeded. Both mean a request's
measured input exceeded what its ticket reserved; see
zs.reserve. Split them apart on the metric,
which labels every occurrence by what happened:
curl -s localhost:9091/metrics | grep zs_reserve_input_budget_over_total
outcome="tool_cutoff"— loops are hitting the ceiling. Yourtool_headroom_per_iteration(default4000) is undersized for the tools you serve: oneweb_readat the default 64 KiBmax_bytesis roughly 16,000 tokens. Raise it, or lowerzs.builtin_tools.web_read.max_bytesso results fit the budget you advertise. Loweringzs.builtin_tools.max_iterationswithout raising the headroom makes this worse, because the reserve clients compute is the product of the two.outcome="rejected"— plain requests are being refused before any upstream call, at zero cost to the payer. The paired WARN linereserve input budget exceeded on initial bodycarries acode.context_length_exceededis a prompt whose size estimate is far past the model's window, which honest users send; at the context cap that means more than three times the budget, which whitespace-heavy or\u-escaped text can reach before the real window.input_budget_exceededis a caller that reserved less than its prompt: an honest client sizes with the same bound your node measures with, so isolated hits are one caller sizing badly, and a broad, sustained rate across many payers is more likely a client older than your node's protocol version.outcome="forwarded"— not a problem. These were long prompts whose reserve was already at the context cap, so your node let your backend decide whether they fit instead of refusing them on the size estimate. The ones your backend refuses show up onzs_provider_upstream_rejection_totalascode="context_length_exceeded", and the payer gets that code at zero cost. With enforcement on, a forwarded request is not offered your built-in tools, since one round alone can fill the window. Forwarding happens with enforcement off too, so these never count asmonitored.outcome="monitored"— you haveenforce_input_budget: false. Nothing is being refused or cut off; these requests were served at your expense.
To serve while you tune, set enforce_input_budget: false (monitor mode), but
only temporarily. Enforcement is what stops a payer under-declaring
input_count and having you serve the difference, and the node logs a startup
WARN for as long as it's off.
Clients say tools never fire, but the model supports them
Three different faults look identical from the outside. zs-node doctor tells
them apart: it reports tools accepted and tool calls emitted as separate
lines, because they have different fixes.
- The backend rejects
tools[]outright. On vLLM or SGLang that is a serving flag, not a model limitation: you need--enable-auto-tool-choicetogether with a--tool-call-parsermatching the model's chat template. llama-server needs--jinjawith a tool-declaring template. If your config also sayscontext.tool_use: true, doctor reports it as a failure: the node is advertising a capability every client then gets a 4xx for. - The backend accepts
tools[]and never returns a call. Almost always a--tool-call-parserthat doesn't match the template: the request is valid, the model emits a call in its own format, and the runtime doesn't recognise it. - The backend accepts
tools[]but rejectstool_choice: "required". Normal tool calling works; only clients that force a specific call get a 4xx. Doctor reports this as a warning rather than a failure, and says which form worked.
zs-node doctor --models=<your-model-id> --trace=/tmp/tools.jsonl
jq -c 'select(.step|startswith("tool_use")) | {step,status,response_body}' /tmp/tools.jsonl
Reasoning shows up in the backend but not through the node
Your backend visibly produces reasoning, but clients don't see it, or the model
loses its thread across a tool call. Two independent settings under
context.reasoning are involved.
context.reasoning.supported gates whether the node replays reasoning back to
the model between tool-loop iterations. A reasoning model declared as
supported: false still streams its reasoning to the client, but stops having it
replayed, so it re-plans from scratch each iteration. This most often affects
OpenAI's o-series and GPT-5 behind openai_passthrough, where nothing in the
model list announces that the model reasons.
context.reasoning.allowed_efforts is what clients are offered, and the node
forwards a client's chosen effort to your backend as reasoning_effort. Plenty
of runtimes emit reasoning natively while rejecting that field as an unknown
argument, so advertising an effort your backend doesn't take turns into a 400
for the client that picks it. If your backend is in that group, declare
supported: true with no allowed_efforts.
zs-node doctor reports reasoning emitted and reasoning_effort accepted
separately, and flags either one that disagrees with your config.
Doctor says a model rejects image input, but it doesn't
zs-node doctor establishes image support by sending a small test image. If your
backend refuses that particular image (a minimum-dimension rule, a format
restriction, a decode failure), the exam quotes the upstream's own message
alongside the finding. Read it. Only a complaint about the model ("this model
does not support image input") means the model lacks image support. A complaint
about the image ("invalid image", "dimensions below the minimum") does not.
The probe tries two images and refuses to record a negative from wording it
can't attribute to the capability, so this should be rare. If you see it anyway
for a model you know takes images, --trace has the exact exchange:
zs-node doctor --models=<your-model-id> --trace=/tmp/vision.jsonl
jq -c 'select(.step|startswith("vision")) | {step,status,response_body}' /tmp/vision.jsonl
Report the upstream message; the classifier is a list of known wordings. In the
meantime the node advertises what context.input_modalities says, and
doctor's finding is advisory.
The local model keeps crashing
When you run a local model, its supervisor restarts a crashed engine up to five times in a minute, then gives up. Check:
- The model file path is correct and readable by the user the node runs as.
gpu_layersdoesn't exceed your VRAM. Drop it, or use a smaller quantization.context_window × parallel_slotsfits in VRAM. That product is your KV cache, and it is routinely larger than the weights. Do the arithmetic in What the inference backend needs.- Your
extra_argsdon't re-set values the node derives itself (context size, parallelism, continuous batching). Those are rejected at validation, and if one slips through it silently breaks the per-session contract. - On dual-GPU without NVLink, layer-split mode is usually faster than row-split.
The node is healthy but gets no traffic
Probes succeed, /v1/zs/details looks right, and yet no requests arrive. Check
the signing account first: proxies and the web client skip any node whose
signing account has less than 5 ALGO spendable (its balance minus Algorand's
minimum balance), because that account pays the fees of every request.
- Look at the
zs_signing_balance_algosmetric, or for asigning-address spendable ALGO is below the routing floorerror in the log.zs-node doctorreports it too. - Top up the signing account (see Operations). Clients pick up the new balance within about five minutes.
If the float is fine, work through the rest of Keeping your node routable.
Settlement entries stuck in settling
A settle transaction was submitted but algod didn't confirm it in time, so the entry flips back to be retried automatically. Occasional churn here is normal. Entries that stay stuck point at one of:
- algod is unreachable — the node can't submit or confirm.
- The signing account is out of ALGO — no fee budget to settle. Top it up (see Operations).
- Network congestion — confirmations are slow but eventual.
A late confirmation isn't lost: the node reconciles in-flight settlements
against the chain when it restarts, so a planned restart never drops ledger
state. If you need to inspect or hand-repair a specific ticket, the
settlement admin subcommand is the escape hatch.
Out of VRAM at startup
Your local engine won't allocate. If it doesn't fit
lists what to lower, in order of what you give up: the model's
context_window, zs.max_active_tickets or llm.local.parallel_slots, the KV
cache precision, and gpu_layers, before falling back to a smaller quantization
or model.
Clients get 502 or 504 errors, but the node looks fine
Anything you put in front of the node (CDN, ingress controller, reverse proxy, load balancer) can answer for it, and when it does the error comes from that layer, not from the node. Clients can tell the difference: the node's own refusals always carry a structured error code, and a gateway's don't.
While a request is still being set up, the client retries and reroutes automatically, so brief blips never reach the user. They still cost you the traffic, and a client that keeps failing against you there demotes you in its routing for several minutes. Once a request is paid for and sent, there is no retry: a gateway error during inference goes straight back to the caller as a failed request. Repeated timeouts on requests forwarded through you also demote you if your peers carry the same nodes without them. That takes more failures than an ordinary error, and it fades once the timeouts stop.
Two causes, in order of likelihood:
Restart gaps. If your front door serves errors while the node is down, use graceful shutdown and point your readiness probe at the private listener; see Graceful shutdown. A draining node stops advertising before the gap opens, so clients route elsewhere and never meet the error.
Origin timeouts shorter than your model's thinking time. A reasoning-heavy
request can send nothing for a long stretch before the first token. A gateway
that gives up during that silence reports 504 (Cloudflare: 524) even though
the node was working normally. How long a silence you have to survive depends on
how the caller asked:
| Request | Keepalive during silence | Your origin read/idle timeout must exceed |
|---|---|---|
| Streamed | Yes: an invisible keepalive every server.sse_keepalive_interval (15s by default), starting as soon as the response headers are sent, before the model is called. | server.sse_keepalive_interval. |
| Buffered (non-streaming) | No. Nothing is sent until generation completes. | Your slowest complete generation, a far bigger number. This is the usual cause of a 504 that lands almost exactly on a round minute. |
You can't control which mode callers use, so size for the buffered case if you want to serve them reliably. Defaults that are commonly too low:
| Gateway | Default |
|---|---|
nginx proxy_read_timeout | 60s |
| AWS ALB idle timeout | 60s |
| Cloudflare | ~100s, and not raisable except on Enterprise. If your slowest generation exceeds it, that's a hard ceiling to design around. |
Errors in the 520–527 range are Cloudflare-specific and always mean the
problem is between Cloudflare and your node: 521 origin down, 522 connection
timed out, 523 origin unreachable, 524 origin timed out.
The node can't reach algod, the price feed, or the backend
Almost always an egress firewall. The node needs outbound HTTPS to:
- your algod endpoint,
- the ALGO/USD price feed (unless you've overridden
zs.oracle.endpoint), - your inference backend's URL,
- and, if you use
tls.mode: acmeorserver.tls.ip_sync, the NFDomains API (api.nf.domains/api.testnet.nf.domains) andapi64.ipify.org.
From inside the container or host, curl -v against each is the fastest way to
find which one is blocked.
The node only listens on localhost
server.listen defaults to loopback (127.0.0.1:9090), so the node isn't
exposed yet. Set NODE_SERVER_LISTEN=:9090 to bind all interfaces, and serve
HTTPS with tls.mode: acme or behind a single reverse proxy, never a load
balancer. See The two listeners,
Installation, and Reaching your node.
The port you bind is the port published in your base URL on chain, so changing
it means updating that record too. The private listener
(server.private_listen, default 127.0.0.1:9091) is separate: override it if
your probes or scraper live in another container.