Skip to main content

Health & compatibility

Clients continuously probe every registered node to learn whether it's reachable, what it currently serves, and whether it speaks a compatible protocol version. They also read each node's signing-account balance from the chain. A node that probes cleanly and has a funded signing account is routable; one that doesn't is flagged or skipped.

What clients probe​

Probing reads your node's public details document, an unauthenticated, unencrypted endpoint that returns your live catalog and capabilities (see Serving models & pricing). From it, and from your on-chain record, a client determines:

  • Reachability: does the endpoint respond?
  • Protocol version: is your wire-protocol version compatible with the client's?
  • Your catalog: which models, at which prices, with which capabilities.

When both reachability and a compatible protocol version hold, the client marks your node OK and ready to route. Separately, the client verifies your encryption recipient (your current sealing key, unexpired and signed by your on-chain key; see Encryption & keys) at probe time and again just before sealing. If the recipient is missing, expired, or unsigned, the node stays OK but gets no traffic: at send time the client fails over to another node rather than downgrade.

Probe states​

A client classifies each node into one of a few states:

StateMeaningRoutable?
ProbingThe client hasn't finished its first probe yet.Not yet
OKReachable and on a compatible protocol version.Yes, if its signing account is funded and its recipient passes the send-time check
IncompatibleReachable, but the protocol version doesn't match.No — flagged in the picker, never chosen by Auto
FailingThe probe errored, or (in privacy mode) no relay could reach it.No

The decision flow:

An incompatible node is the most common avoidable problem: it's up and serving, but on a protocol version the client can't speak, so it's dropped from selection. Keep your node software current.

How often you're probed​

Clients re-discover registered nodes on a short background cadence:

  • Web clients pick up changes to your details document (a price change, a newly served model) within about a minute, and on-chain changes (a new node, a cleared staging flag) within about five minutes.
  • The proxy re-reads the chain and re-probes nodes on a five-minute cycle by default.
  • A node found unreachable is re-checked less often, every few minutes, until it recovers, so a node that's down isn't spammed.

No announcement is needed. The intervals are client-side and may change.

Performance is measured, not claimed​

Beyond up/down, the network measures two performance signals per node and keeps them on-chain as rolling averages:

SignalWhat it measures
Time to first tokenHow quickly you start responding (a latency and hardware-readiness signal).
Decode throughputTokens per second once you're streaming.

These feed target selection, including users' Auto mode, so faster nodes get picked to serve more often. Relay selection is weighted by network round-trip time only. You can't game these signals by advertising; they come from real served requests.

Signing-account float​

Your node's signing account pays the network fees of every request it serves (see Staking & economics). Clients and proxies read its balance from the chain on the same background cadence as the node list, and skip a node with less than 5 ALGO spendable: the balance minus Algorand's minimum balance. The node isn't flagged as failing and keeps serving as a relay (relaying costs it no fees). It just isn't chosen to serve requests, and a model that only underfunded nodes serve drops out of the catalog, until you top the account up. Routing resumes within about five minutes.

The node tracks this for you: it warns when the spendable balance drops below zs.signing_balance.warn_below_algos (10 ALGO by default) and logs an error when it crosses the 5 ALGO floor. See Monitoring.

Staging nodes​

You can set a registered node's staging flag from the operator dashboard. Ordinary clients won't dispatch to a staging node, even though it's registered and reachable. They also won't see it: it's absent from the operator directory, it doesn't contribute to the published capability and pricing summary, and a model only your staging nodes serve doesn't show up in the model catalog or the chat app's picker at all. Use it to:

  • Bring up a new machine or a new model and exercise it end to end before it takes production traffic.
  • Test a software upgrade against real protocol behavior without risking paying users.

Users can opt in to staging nodes with an Allow staging toggle, so you can point a cooperating tester at your staging node. That's also how you test it yourself: a default zs-proxy won't list your staged model, so run it with --allow-staging (or set zs.allow_staging: true). Clear the flag when the node is ready, and it joins normal rotation within about five minutes.

Staging does not take your node out of service as a relay. Relays only forward encrypted bytes and run no inference, so a staging node keeps carrying other people's traffic. That also exercises the network path without touching your model.

Keeping your node routable — checklist​

  • Stay reachable at your advertised endpoint over HTTPS.
  • Keep your node software current, so your protocol version stays compatible.
  • Keep a valid, signed, unexpired encryption recipient advertised.
  • Keep at least 5 ALGO spendable on the signing account, and preferably well above it.
  • Serve what your details document advertises. If you stop serving a model, drop it from the catalog so clients don't route to a model you no longer have.