Monitoring & metrics
/healthz, /livez and /metrics are served on the private listener
(server.private_listen, default 127.0.0.1:9091), not the public port, so
point your scrape config and uptime checks there. If probes or a scraper run in
another container, bind the private listener to a reachable address (e.g.
NODE_SERVER_PRIVATE_LISTEN=:9091) rather than moving these routes to the
public port. See the
endpoint reference.
Health — /healthz and /livez
Both report process health only, not whether your inference backend is reachable or how many models you serve. For those, watch the backend metrics.
| Endpoint | Question it answers | On a serving node | While shutting down |
|---|---|---|---|
GET /healthz | Readiness — should traffic go here? | 200 {"status":"ok"} | 503 {"status":"draining","inflight_inference":N} |
GET /livez | Liveness — is the process wedged? | 200 {"status":"ok"} | 200 {"status":"draining",…} |
/healthz returns 503 for the whole
graceful shutdown, removing the node from
rotation. Use it for readiness probes and for uptime monitors that treat
draining as normal.
Point liveness probes at /livez, never /healthz. See
Graceful shutdown for why, and for node builds
without /livez.
Metrics — /metrics
GET /metrics returns a Prometheus exposition. These are the signals worth
scraping and alerting on:
Backend health & catalog
| Metric | Type | What it tells you |
|---|---|---|
zs_provider_healthy | Gauge | 1 when your text backend is reachable, 0 when it's down (after llm.health_check.failure_threshold failed probes) or not yet probed. While 0, the node advertises zero text models and returns 503 on text reserves. Alert on a sustained 0. |
zs_advertised_models | Gauge | How many text models you're actually serving right now — the number clients see. |
zs_provider_discovered_models | Gauge | Raw count of models your backend's own listing reported on the last probe. A diagnostic, not a serving count — third-party OpenAI-compatible backends often ship an incomplete list, so this can differ from zs_advertised_models. |
zs_image_provider_healthy | Gauge | Image-backend (ComfyUI / Comfy Cloud) twin of zs_provider_healthy. Only present when an image backend is configured, and independent of it — a mixed node can have one backend down while the other serves. |
zs_image_provider_discovered_models | Gauge | Image-backend twin of zs_provider_discovered_models — the image backend's raw upstream probe count, not your served image catalog. |
zs_image_advertised_models | Gauge | Image-route models you're currently advertising. |
zs_hf_coordinates_total{result} | Counter | HuggingFace model-info lookups — the startup tag warm and the advisory coordinates drill-in — by outcome: hit, negative (repo didn't resolve), error (timeout / 429 / 5xx / gated), cached. Zero when discover_huggingface is off. A hit rate confirms discovery is working; a sustained error rate means HF is unreachable, rate-limiting, or blocked. Advisory — never affects routing, pricing, or settlement. |
Admission & capacity
| Metric | Type | What it tells you |
|---|---|---|
zs_reserve_capacity_slots{model} | Gauge | Per-model concurrent-admission ceiling — the model's max_active_tickets (or the node default). The denominator for utilization. |
zs_reserve_inflight_slots{model} | Gauge | Slots currently held by reserved / in-flight tickets. capacity − inflight = free slots; 0 free means the next reserve for that model returns 429 no_capacity. |
zs_reserve_capacity_denied_total{model} | Counter | Reserves turned away with 429 no_capacity because the model was at its ceiling — the actionable "we're refusing paying callers" signal. Alert on a sustained rate() > 0. Distinct from the per-IP rate limiter below. |
zs_reserve_payer_slot_cap_denied_total{model} | Counter | Reserves turned away with 429 payer_slot_cap because a single payer already held its per-model zs.payer_slot_cap share. Unlike zs_reserve_capacity_denied_total (the whole pool is full), a non-zero rate here is one payer repeatedly reaching past its share, usually the cap doing its job. Raise the cap if a legitimate high-parallelism caller is being pinched. |
zs_reserve_free_quota_total{outcome} | Counter | Reserves for a free model refused by the contract's per-payer daily free-ticket quota. exhausted = 403, the payer spent its 24h allowance — a low background rate is normal on a popular free model, not a fault. cold_miss_budget / chain_error = 503 fail-closed, meaning the allowance read was budget-denied or algod failed; alert on a sustained rate of those two, since free reserves are then being refused for an infrastructure reason rather than a policy one. Paid reserves never touch this counter. |
zs_reserve_funds_gate_total{outcome} | Counter | Reserves refused by the payer prepaid-pool opt-in gate. not_opted_in = 403: the payer has never deposited into its prepaid ticket pool, so it cannot fund a ticket box — a client problem, not yours. cold_miss_budget = 503 fail-closed: the budget for checking a never-seen payer on chain was exhausted, which is what denies new payers under a flood of rotating addresses. Alert on a sustained rate here: it means the gate is under pressure. chain_error = the check itself failed (transient algod). |
zs_reserve_abuse_banned_total | Counter | Reserves rejected with 429 reserve_temporarily_banned because the (signature-verified) payer is in the penalty box — an escalating, time-boxed ban triggered by persistent reserve-and-abandon or repeated slot-cap trips. Node-local, never on-chain. A sustained rate means a payer is being actively throttled for abuse. It has no labels, so that the payer address is not exported. |
zs_reserve_input_budget_over_total{model,reason,outcome} | Counter | Requests whose measured input exceeded the reserved input_count (see zs.reserve). reason is where it was caught — initial_prompt, tool_loop, or chain. outcome is what happened — rejected (plain request refused zero-cost), tool_cutoff (loop ended early with a final answer), forwarded (the reserve was already at the context cap, so your backend decided whether the prompt fits), or monitored (enforcement off, served anyway). A sustained rate under monitored is the signal to set enforce_input_budget: true; a sustained tool_cutoff means your tool_headroom_per_iteration is too low for the tools you serve. |
zs_provider_upstream_rejection_total{model,code} | Counter | Non-2xx responses from your inference backend, labelled by model and error code — a rising rate points at a backend problem, not the network. An HTTP rejection that says the prompt is too long is labelled code="context_length_exceeded" whatever shape your backend sent it in. Some of those are expected: a long prompt whose reserve was at the context cap goes to your backend to decide whether it fits, and some genuinely don't. A rate well above zs_reserve_input_budget_over_total{outcome="forwarded"} means something else: your backend serves a smaller window than the context_window you advertise. |
zs_ratelimit_denied_total{bucket} | Counter | Requests dropped by the public-facing rate limiter, by bucket (per-IP / per-account). Distinct from no_capacity. |
Signing-account balance
| Metric | Type | What it tells you |
|---|---|---|
zs_signing_balance_algos | Gauge | Your hot signing address's spendable ALGO (balance minus Algorand's minimum balance) from the last successful poll. This address fee-pays every on-chain action (settlement, and ACME / IP-sync updates), and it's the number proxies and the web client check: below 5 they stop sending requests to your node. Alert above that (e.g. < 10). Reads 0 before the first poll — cross-check the timestamp below. |
zs_signing_balance_last_success_timestamp_seconds | Gauge | Unix time of the last successful balance poll. If it stops advancing, the balance gauge is stale (algod polling is failing) — alert on staleness so a stuck poller isn't mistaken for a healthy balance. |
Keeping this account funded is covered in Staking & economics.
Tool loops
Exposed when you enable built-in tools, alongside the standard go_* /
process_* runtime metrics.
| Metric | Type | What it tells you |
|---|---|---|
zs_builtin_tool_invocations_total{tool} | Counter | Tool dispatches by name. |
zs_builtin_tool_errors_total{tool} | Counter | Tool errors by name. |
zs_builtin_tool_duration_seconds{tool} | Histogram | Tool latency. |
zs_tool_loop_iterations | Histogram | Iterations per loop. |
zs_tool_loop_duration_seconds | Histogram | Wall time of a loop. |
zs_tool_loop_finished_total{outcome} | Counter | How loops ended. See the outcome vocabulary below. |
zs_tool_call_leak_total{syntax,disposition} | Counter | Rounds where the backend returned the model's native tool-call markup as visible content instead of parsing it into a call. Any sustained rate means the runtime has no tool-call parser configured for that model's chat template — on vLLM/SGLang that's --enable-auto-tool-choice plus a --tool-call-parser matching the template. disposition="recovered" means the node ran the call anyway and the turn completed; "ignored" means the markup sat beside a real answer and was left alone; "suppressed" means it happened on a round the input budget had already forced final, so the turn ended zero-cost with no answer; "void_round" means the round was cut off mid-call by the backend's output cap (see the Kronk max_tokens warning). Run zs-node doctor to confirm the cause. |
zs_tool_loop_finished_total outcomes: ok (the model produced a final
answer), cap_exceeded (hit the absolute iteration cap), stalled (repeated
tool calls without progress), provider_error (the upstream failed the round),
append_error (a tool result couldn't be appended to the conversation),
leak_suppressed (unparsed tool-call markup stood in for the answer on a
budget-forced round — zero-cost, no answer shipped), open_unconfirmed (the
payer's escrow transaction never confirmed and the watchdog cancelled), and
policy_refusal (a node-side policy gate refused the upstream response — today
that is the xAI zero-data-retention
check). All of
them are pre-materialized at zero from boot, so rate() has a flat line to
start from.
Log lines worth alerting on
upstream LLM unreachable— a backend stopped answering health probes; the node now advertises none of that backend's models and returns503on reserves for them. Thecomponentfield names the backend:llm-healthfor text,image-llm-healthfor image. The other backend keeps serving. Paired withupstream LLM reachableon recovery.signing-address ALGO balance is low— the signing account's spendable balance dropped below yourwarn_below_algosthreshold (default10). Refill before it reaches the 5 ALGO routing floor. Paired with a recovery line once topped up.signing-address balance poll failedmeans algod is unreachable and the balance metrics are stale.signing-address spendable ALGO is below the routing floor— an ERROR: spendable is under 5 ALGO, so proxies and the web client have stopped sending requests to this node. Top up the signing account. The node logs the matching… back above the routing floorline on its next balance poll, and proxies and the web client resume sending requests within about five minutes, when they next read the balance.reserve input budget exceeded on initial body— a decrypted prompt was larger than what its ticket reserved for it. Withenforce_input_budget: true(the default) the request was refused before any upstream call, at zero cost to the payer. Isolated lines are a caller sizing badly; a sustained stream from many payers points at your own config — see Tool loops end early, or reserves are refused.input bound over budget at the context ceiling; forwarding for the upstream to judge fit— INFO, not an error: a long prompt's reserve was already at the context cap, so the node sent it to your backend rather than refusing it on the size estimate. If the backend refuses it as too long, the payer seescontext_length_exceededat zero cost.tool loop reached reserved input budget— INFO, not an error: a tool loop hit its reserved ceiling and was asked for a final answer instead of another round. The user still got a reply. A steady rate meanstool_headroom_per_iterationis undersized for the tools you serve.oracle fetch failed— the ALGO/USD price feed is unreachable. Non-fatal: reserves still succeed, but the ALGO-fee-in-USD figure is dropped for that cycle. See Troubleshooting.settlement: confirmation failed, will retry— a settle transaction was submitted but not confirmed in time; the driver retries automatically. Occasional lines are normal; a sustained stream means the driver isn't draining (see Troubleshooting).settlement: terminal settle error; giving up— a ticket needs manual reconciliation (the contract's inactivity backstop still protects the payer's funds).settlement: ticket FROZEN (payer disputed)— a payer has disputed a charge and the ticket can no longer settle on-chain. Terminal, and the only one on this list that needs a human: see Disputed tickets. Alert on it.
What to dashboard
- Liveness — the
/livez200rate. Don't alert on/healthz: it returns503for the whole drain on every restart, so it would page you on routine upgrades. Dashboard it as the rotation signal instead; a node at503far longer than your drain budget is stuck, not draining. - Backend up —
zs_provider_healthy(andzs_image_provider_healthyif you serve images); alert on sustained0. - Admission saturation — per-model utilization
zs_reserve_inflight_slots / zs_reserve_capacity_slots(1.0= full, the next reserve429s); page onrate(zs_reserve_capacity_denied_total[5m]) > 0. Briefly hitting full is healthy use of capacity; sustained denials mean you're turning callers away — raisemax_active_tickets(with matching backend capacity / VRAM) or add a node. - Signing float —
zs_signing_balance_algos, alerting well above the 5 ALGO routing floor, plus staleness on the timestamp gauge. - Backend errors —
rate(zs_provider_upstream_rejection_total{code!="context_length_exceeded"}[5m])by code. Too-long prompts are expected, so watch them separately: alert whenrate(zs_provider_upstream_rejection_total{code="context_length_exceeded"}[1h])runs well aboverate(zs_reserve_input_budget_over_total{outcome="forwarded"}[1h]), which points at an over-declaredcontext_window.
The troubleshooting guide maps common symptoms to fixes.