Serving models & pricing
Your node advertises what it offers in a single public details document, the same one clients probe for health (see Health & compatibility). Clients read it directly to build the live catalog users see in the model picker.
The details document
The details document is served unauthenticated and unencrypted at
/v1/zs/details. It carries:
| Field | What it is |
|---|---|
| Protocol version | The wire version your node speaks. Clients drop nodes on an incompatible major version (see Health & compatibility). |
| Models | The catalog. See Per-model advertisement. |
| Built-in tools | Any in-loop tools your node offers, such as web search or image generation, with their names and descriptions. |
| Encryption recipient | Your current sealing key, signed and with an expiry (see Encryption & keys). |
| Config hash | A sha256: fingerprint of your node's declared policy. See The config hash. |
The config hash
config_hash fingerprints your rate card and tool roster: the model catalog and
its capabilities, built-in tools, rate limits, minimum charge, oracle source, and
relay mode. It ignores deployment and identity details (listen addresses,
file paths, operator_id, addresses, secrets) and knobs that fit a node to its
hardware, such as web_read.max_download. Nodes running the same policy
therefore advertise the same hash wherever they run, on a small VM or a large
one.
Use it to confirm a fleet is uniformly configured (jq -r .config_hash across
your nodes should match) or to notice that a node's policy drifted after an
edit. It is advisory; nothing on the wire is gated on it. It has two limits:
- It covers your declared policy, not what's reachable. If your upstream
goes down, the models endpoint advertises nothing while
config_hashstays the same. - Compare only across nodes on the same release. A release that changes which fields are fingerprinted changes every node's value, and nothing on the wire distinguishes that from an edit you made.
Per-model advertisement
Each model in your catalog declares what it is and what it costs:
| Field | What it is |
|---|---|
| Model id | The identifier users select and the one carried in requests. |
| Source | The public artifact the model is (hf:org/model). Effectively required: a model with no checkable identity is dropped from your catalog, with only a WARN in the log. See Model identity & integrity. |
| Input and output modalities | E.g. text in / text out, or text+image in. Drives the Vision, Generation, and Edit capability badges. |
| Context window and maximum output length | The model's size limits. Set the output length explicitly; see Set an output ceiling. |
| Reasoning support | Whether the model supports reasoning, the effort levels it allows, and a default. Drives the Reasoning badge. |
| Tool support | Whether the model can call tools. Drives the Tools badge. |
| Tags | Descriptive and policy tags (e.g. content-policy labels). Auto-populated from the model's HuggingFace card when you declare a source; anything you add is merged in. They appear in the picker and can affect how Auto mode prioritizes the model. |
| Rates | See Pricing the routes. |
Advertise only what you actually serve. Clients route to anything you list, so
a model you no longer have means failed requests: remove it from zs.models
when you stop serving it. The
health check withdraws models on
its own only when the whole backend is down, or when a model you serve through
default_pricing alone disappears from the backend's list.
Set an output ceiling
Set max_output_tokens for every model you serve. It is the model's output
ceiling: advertised to clients, enforced when a request reserves payment, and,
because clients take a declared ceiling at face value, the number a client that
doesn't ask for a specific one receives.
zs:
models:
"Qwen/Qwen2.5-32B-Instruct":
context:
context_window: 131072
max_output_tokens: 32768
Leave it out and clients fall back to a quarter of the context window, capped at 32,768 tokens. Then:
- Long answers get cut short. Even a 256K-context model stops at 32,768 tokens.
- Large explicit requests aren't caught. A client that does ask for more meets no ceiling on your node, so the request reaches your inference server and fails there, after the user's payment has been reserved.
When every operator serving a model declares a ceiling, the proxy lowers an over-large request to the largest one advertised. Operators who declared that much stay eligible and those who declared less drop out; price and responsiveness then decide who gets the request. If you declare none, the proxy applies no operator's ceiling, and the request reaches your backend at full size and fails there.
max_output_tokens must be below context_window. A reservation covers
input plus output, so a ceiling at or above the window leaves no room for a
prompt and your node is never selected for that model. The node still loads and
advertises the model normally, so nothing looks wrong.
The rule applies to the window your node advertises: with a ceiling but no
context_window, that is the window discovered from your backend.
A quarter of the context window is a reasonable starting point. zs-node init
suggests it and asks about every model, and zs-node doctor flags a model with
a missing or impossible ceiling.
Reasoning models
The Reasoning badge and the per-request effort selector appear only when you
declare the capability. LM Studio publishes it automatically; vLLM, llama.cpp,
and any OpenAI-compatible passthrough (including OpenAI's own o-series / gpt-5)
publish none, so declare it under the model's context:
zs:
models:
google/gemma-4-31b-it:
context:
reasoning:
supported: true
allowed_efforts: ["low", "medium", "high"]
default_effort: "medium"
Advertised effort tiers are also what let the client's toggle turn thinking on.
A gemma-style template doesn't think until the request carries an effort. vLLM
0.25+ turns its enable_thinking flag on for any effort except "none", on
both the Responses reasoning.effort and the Chat reasoning_effort, so the
client's selection turns thinking on with no server-side change.
On a vLLM older than 0.25, or a template with a non-standard kwarg name, enable
thinking on your inference server instead, e.g. vLLM's
--default-chat-template-kwargs '{"enable_thinking": true}'. The node reads
both vLLM's current reasoning field and the older reasoning_content, and a
backend's native /v1/responses reasoning items pass through untouched.
A thinking model spends part of its output-token budget on the chain-of-thought before the visible answer, so on a hard prompt it can use the whole budget and return an empty, cut-off answer. The surest fix is an explicit output ceiling for the model.
Choosing a backend
Your node doesn't run inference itself. It sits in front of an inference engine,
on loopback or over the network, chosen with one setting: llm.provider. One
provider per node. Every provider and its defaults are listed under
llm. An empty provider ("")
disables text inference, for image-only or
relay-only mode.
- openai_passthrough
- lmstudio / llamacpp
- kronk
- local
- vertexai
The most flexible option. Set the upstream's API root in llm.openai.base_url,
and supply the key through the NODE_LLM_OPENAI_API_KEY env var rather than
committing it:
llm:
provider: "openai_passthrough"
openai:
base_url: "https://api.openai.com/v1" # include the version segment yourself
api_key: "" # prefer NODE_LLM_OPENAI_API_KEY env var
timeout: "0s" # 0 = disabled; required for long streams
translate_responses_to_chat: false # true ONLY for upstreams without /v1/responses
/v1/responsesPayers assume the Responses API is there. At startup the node sends one
zero-token probe to your upstream's /v1/responses and exits on a 404, 405,
or 501 rather than advertising models it cannot serve. Any other outcome (a
4xx about the request itself, a timeout, a connection error) says nothing about
whether the route exists, so the node logs a warning and boots. With
translate_responses_to_chat: true the node serves the route itself and skips
the probe.
base_url is appended to directly (/chat/completions, /responses,
/models), so include the version segment, usually /v1, yourself. For a
non-/v1-rooted upstream, give the full prefix (e.g.
https://api.z.ai/api/paas/v4). A local engine works the same way, e.g. vLLM at
http://127.0.0.1:8000/v1 or Ollama at http://127.0.0.1:11434/v1.
Set translate_responses_to_chat: true when the upstream has no native
/v1/responses (llama.cpp's own server and Ollama, and vLLM/TGI unless you've
enabled its Responses endpoint); the node then emulates the route on top of
/v1/chat/completions. Leave it false for a backend that implements Responses
natively (OpenAI, Azure OpenAI, LM Studio). If you leave it false against a
backend with no Responses route, the node refuses to start and names the route.
If you set it true against a backend that has one, the node boots but emulates
the route and loses native Responses behavior. zs-node doctor settles it
empirically; see
doctor.
These share the passthrough chat path and also query the runtime's native
metadata API, so /v1/zs/details carries context windows and modalities
without you hand-authoring every zs.models[].context:
lmstudio:base_urldefaults tohttp://localhost:1234/v1,api_keyoptional.translate_responses_to_chatdefaults to false (LM Studio implements/v1/responsesnatively).llamacpp:base_urldefaults tohttp://127.0.0.1:8080/v1,api_keyoptional.translate_responses_to_chatdefaults to true (llama-serverhas no/v1/responses).
Kronk (source)
is an OpenAI-compatible server that links llama.cpp in-process and adds its own
serving layer. Unlike
llama-server, it runs a real multi-model pool in one process, with a
memory budget, eviction, and an idle TTL, so it is the way to serve several
local models from a single node.
llm:
provider: "kronk"
openai:
base_url: "http://127.0.0.1:11435/v1" # the default
api_key: "" # optional; authorization mode defaults to open
translate_responses_to_chat defaults to false; Kronk implements
/v1/responses natively.
Why this provider exists. The node's always-on model identity
gate advertises a model only
if it resolves to a canonical public artifact: a zs.models[<id>].source you
declare (hf:org/repo), or a recognized frontier model. A Kronk model id
never names the repo it came from, so under a plain openai_passthrough:
| Kronk | Model id | Result |
|---|---|---|
| 1.31.3 and earlier | Bare GGUF file stem (Qwen3-8B-Q8_0) | Every model is dropped from your advertised catalog, with only a WARN in the log. |
| 1.31.4 and later | <org>/<file stem> (unsloth/Qwen3-8B-Q8_0) | Looks like a HuggingFace org/repo, but the repo is really the file's parent directory, typically …-GGUF. The node points at a repo that does not exist, and your models are grouped apart from every other operator serving the same weights. |
The kronk provider reads Kronk's native endpoints and rebuilds each model's
HuggingFace ref from its models-directory layout
(<org>/<repo>/<file>.gguf → hf:<org>/<repo>@<revision>), so provenance,
context windows, and capabilities resolve on their own. A source: you declare
in zs.models still wins. zs-node doctor reports a non-kronk provider
pointed at a Kronk as a failure, listing which of your models the backend
does and doesn't supply a source for.
Upgrading to Kronk 1.31.4 or later means renaming your model keys in two
places. 1.31.4 rejects a bare model id with a 400, and the node forwards
your configured id upstream verbatim, so a zs.models key left at the old stem
loads, advertises, and reserves, then fails every inference. Rename each key
to the id GET /v1/models now reports:
zs:
models:
# before (Kronk <= 1.31.3)
# Qwen3-8B-Q8_0:
# after (Kronk >= 1.31.4)
unsloth/Qwen3-8B-Q8_0:
pricing: { input_rate: 0.05, output_rate: 0.20 }
Kronk's own ~/.kronk/models/model_config.yaml keys must be canonical too, or
its model pool will not initialize. Change both, then restart both. zs-node doctor reports a stale key as a failure and names the exact id to use.
Back up model_config.yaml before the first start. If the file has no
version: key, 1.31.4 rewrites it in place on first run, wrapping your
settings under models: and stamping version: 1. Comments and formatting are
lost, and no backup is kept.
Other Kronk upgrades that need action:
| Kronk | Change | What to do |
|---|---|---|
| 1.30.3 | Both listeners default to loopback; earlier versions bound 0.0.0.0. | Verify the bind addresses on an upgraded box. |
| 1.30.3 | KRONK_AUTHORIZATION_MODE replaces KRONK_AUTH_ADMIN_ENABLED / KRONK_AUTH_LOCAL_ENABLED. | Set KRONK_AUTHORIZATION_MODE instead. |
| 1.30.3 | Reasoning, control, and buffered tool-call tokens count in completion_tokens, even when they never surface as assistant text. Charges for the same work read higher than on earlier versions. | Nothing. The node bills what the backend reports; this is more complete counting, not a pricing change. |
| 1.30.4–1.30.8 | A tool call cut off by the output cap reaches a streaming caller as raw tool syntax. | Upgrade to 1.30.9 or later. |
| 1.31.6 | Requires Go 1.27 and macOS 13 or later. | With an older toolchain or Mac, stay on 1.31.5. |
| 1.32.0 | Fallback sampling defaults change, with no error and no release-note mention: temperature 0.8 → 1.0, top_k 40 → 20, top_p 0.9 → 0.95. They apply only when neither model_config.yaml nor the model's GGUF metadata sets a value. | Set the values explicitly for any model that relied on the old fallbacks. |
| 1.32.4 | The default prompt-cache working set shrinks from nseq-max × max(3, queue-depth) to nseq-max × queue-depth — with the nseq-max: 2 the bundled model_config.yaml uses, four warm sessions instead of six. More than that many active conversation branches now re-read their prompt more often. | Nothing, unless you see cache hit rates fall. Set imc-session-capacity explicitly to keep the old size. |
| 1.32.5 | Library verification is on by default: Kronk checks the native libraries against digests published for the release and refuses to start if any file disagrees. (It still starts, with a warning, for a release that publishes no per-file digests.) | Leave KRONK_LIB_VERSION unset (see below). If you pinned a version from before 1.32.5, remove the pin: that pin now stops Kronk starting. If it still refuses to start, delete the libraries directory under its base path and let it reinstall; deleting models is not necessary. |
| 1.32.6 | A tool call whose arguments fail to parse no longer passes through untouched. The bad calls are dropped; if that leaves none, the reply ends with "The model produced an invalid tool call. Please try again." While streaming, if the call was already announced, the request errors instead. | Nothing on the node side. Keep it in mind when reading logs for a model that emits malformed calls. |
| 1.32.6 | In the official container image, the default KRONK_POOL_MODEL_CONFIG_FILE moves from /etc/kronk/model_config.yaml to /kronk/models/model_config.yaml. Kronk's own default was always the latter; only the image changed. | If you bind-mount a config at the old path, move the mount or set KRONK_POOL_MODEL_CONFIG_FILE yourself. A named volume on /kronk that already has data will not be seeded with the new default file. |
Check the bind addresses before exposing the host. The API listens on
127.0.0.1:11435 and debug on 127.0.0.1:11445 by default. The debug server
serves /metrics and /debug/pprof/* with no auth and no TLS, and an MCP
listener (localhost:9000) and a web admin UI are both on by default.
To reach Kronk off-box (a container, a LAN, a Kubernetes service), set
KRONK_WEB_API_HOST=0.0.0.0:11435 explicitly; publishing a container port is
not enough when the process is bound to the container's own loopback. Leave
KRONK_WEB_DEBUG_HOST on loopback, or firewall it.
-
Never set
KRONK_INSECURE_LOGGING=trueon a serving node. It logs prompts, breaking the node's non-retention guarantee on your own box. -
A reasoning model with no output cap can fail the node's startup check. Before serving, the node runs one real streaming completion, with no
max_tokensand a 60-second limit, to confirm your backend reports the token counts it bills streaming work by. Withenable_thinkingon andmax_tokens: 0, a one-word prompt can produce several hundred reasoning tokens, and the first request after Kronk starts is by far the slowest (we measured roughly a twelfth of later throughput), so a large model on modest hardware can run out of time. If the node exits with a probe timeout, capmax_tokensor setenable_thinking: falsefor that model in Kronk'smodel_config.yaml; there is no node-side setting. Check the resolved values atGET /v1/kronk/models/<id>undermodel_config.sampling-parameters. -
Idle expiry is off by default (1.30.8 and later): the effective
KRONK_POOL_TTLdefault is0. Set a positive duration to unload idle models automatically. At0, count- and memory-pressure eviction still operate. -
Leave
KRONK_LIB_VERSIONunset on 1.32.5 and later. Kronk downloads its native llama.cpp libraries at runtime. From 1.32.5 the binary contains the llama.cpp build tag it asks for and a checksum of the digest list published for that build. The checksum pins the list, and the list pins every file. With the variable unset, verification runs against those digests.Setting a value replaces the baked-in tag and checksum with a bare build tag such as
b11017, and Kronk then trusts whatever digests the download host serves next to the files. That catches a corrupted download but not a substituted one. A tag from before the digest lists existed has no digests at all, so Kronk fails to start. This is what happens if you carry an old pin across the 1.32.5 upgrade.On earlier Kronk nothing is verified either way. Unset does not float: the build tag is baked into whichever Kronk binary you run, so it changes only when you upgrade Kronk. Pinning there records a version you chose; it adds no integrity check. Leave
--allow-upgradeat itsfalsedefault on any version. -
Kronk's server settings can live in
model_config.yamlunder akms:block (1.31.4 and later). They override the built-in defaults but lose toKRONK_*environment variables andkronk server startflags. -
Mind
KRONK_AUTHORIZATION_MODE./v1/kronk/*is Kronk's management API and needs an admin token in every mode exceptopen; the token that authenticates inference is not enough. Without one the node cannot read model provenance: chat keeps working, but you must declare asource:for every model. Runopen, or give the node an admin token.Mode /v1/modelsInference /v1/kronk/*openpublic public public managementpublic public admin token authenticatedany valid token any valid token admin token full-protectedany valid token token + endpoint grant admin token Under
authenticatedandfull-protected,/v1/modelsitself needs a token, and the node uses that list to decide whether the backend is up. With no workingapi_keyit advertises zero models and refuses every reserve with a503while Kronk is healthy; look for a401(missing or invalid token) or403(valid, but not admin) on the discovery probe in your node logs. Underfull-protected, inference tokens also need the matching endpoint grant (chat-completions,responses).
Size your ticket cap from Kronk's admission capacity. Kronk's per-model
model_config publishes nseq-max (execution slots), queue-depth (the
admission multiplier), and admission-capacity (the total it admits at once,
running plus queued). Kronk guarantees that an admitted request runs, so
max_active_tickets should match
admission-capacity. zs-node init writes it for you: one global
zs.max_active_tickets, plus a per-model max_active_tickets for each model
whose capacity differs. zs-node doctor reports drift:
a failure when you admit more than the backend will, a warning when you
admit fewer. An older Kronk that publishes only nseq-max yields that more
conservative number. Nothing recomputes this at runtime, so pin it by hand if
you disagree.
Kronk reports real cached-token counts, so
cache_read_rate works on this
backend, but the cache is per conversation, not pool-wide. Each conversation
thread gets its own session, so a new conversation's first request reports
cached_tokens: 0 even when an identical system preamble is warm in another
conversation. From the thread's second turn on, reuse is ~99%. (Measured at
1.30.3: a 1219-token prompt went 0 → 1214 cached on repeat, while a sibling with
the same preamble started at 0 and reached 1216 on its own second call.) Don't
size the discount as though a shared preamble were free across users.
Vision needs the projector file pulled as well as the model. A multimodal
GGUF whose mmproj companion isn't on disk does not reject an image; it
silently answers from the text alone. The node advertises image input only when
Kronk reports has_projection for the model (check GET /v1/kronk/models). If
a vision model advertises text-only, re-pull it so the projection lands.
The node never sends these parameters, but a caller that writes its own request body may:
| Parameter | Kronk's handling |
|---|---|
stop | 400 on /v1/responses, the endpoint this provider defaults to. Accepted on /v1/chat/completions since 1.30.3 (a string, or up to four; the matched sequence is omitted from the response). |
n | Since 1.30.4, must be 1, null, or omitted. Kronk generates one choice per request. |
Invalid grammar or response_format | 400 since 1.30.4. Earlier versions ignore it and generate unconstrained output. |
Kronk cannot be made to force a tool call. It accepts every tool_choice
form but acts on two: "none" drops the tools, and a named function narrows the
offered list to that one. Neither "required" nor a named function constrains
generation, so a client that depends on forcing a call can get a prose answer
with no error. zs-node doctor probes both forms per model; trust its result
over this paragraph if Kronk changes.
Check max_tokens in model_config.yaml if you serve tool-using models.
When the output cap lands mid tool call, Kronk returns no tool calls and
finishes with finish_reason: "length". From 1.30.9 the answer is one sentence,
Response truncated before completion., streaming or not. On 1.30.4–1.30.8 it
is the raw tool syntax: the node strips it from non-streaming responses, but a
streaming caller has already received it.
Either way the round yields no tool result and the caller is still charged. The
node never executes a call the model did not finish, including an
announced-but-incomplete call in a truncated stream. Kronk's shipped AGENT
profiles cap max_tokens (8192 from 1.30.4, 16384 in 1.31.x), so requests that
would otherwise run to the full context window can hit this. Raise the cap for
models that make long tool calls; the void_round disposition on
zs_tool_call_leak_total counts these rounds.
The recommended path for a self-hosted GPU. The node starts, health-checks, and
restarts a llama-server child over loopback (5 attempts in 60s, then permanent
failure). There's no Docker image for this path yet, so install the
llama-server binary yourself.
llm:
provider: "local"
local:
binary_path: "/usr/local/bin/llama-server" # macOS dev: /opt/homebrew/bin/llama-server
startup_timeout: "60s" # big GGUFs on cold storage may need more
host: "127.0.0.1" # loopback only; the node is the public face
port: 0 # 0 = OS picks an ephemeral port
parallel_slots: 0 # 0 = derive from zs.max_active_tickets
models:
- id: "qwen-2.5-0.5b"
model_path: "/models/Qwen2.5-0.5B-Instruct-Q4_K_M.gguf" # local GGUF
context_window: 32768 # per-session token budget
gpu_layers: 99 # 0 = CPU only; 99 = all layers on GPU (CUDA/Metal)
threads: 0 # 0 = llama-server default
extra_args: []
Rules this path enforces (key reference:
llm.local):
- One model per node, from a local GGUF. A second
models[]entry is a config error, and there's no Hugging Face pull yet. To serve several local GGUFs, run one node per model. - Keep
--parallel,-np,--cont-batching,-c, and--ctx-sizeout ofextra_args. They are derived fromparallel_slotsandcontext_windowand rejected at startup. The supervisor passes--cont-batchingitself when slots > 1. - KV cache scales with concurrency, roughly
context_window × parallel_slots × per_token_bytes, and often exceeds the weights. Raisingmax_active_tickets(whichparallel_slots: 0follows) raises memory and VRAM use. Worked numbers for real models are in What the inference backend needs.
Targets Vertex AI's OpenAI-compatible API. There is no api_key field:
the node mints short-lived OAuth tokens from Application Default Credentials
(Cloud Run service identity, GKE Workload Identity, a GCE service account, or
gcloud auth application-default login for dev) and refreshes them
automatically. You declare the catalog yourself with the publisher-prefixed ids
Vertex expects, and /v1/responses is emulated for you.
llm:
provider: "vertexai"
vertexai:
project: "my-gcp-project"
location: "us-central1" # us-central1 has the broadest model availability
timeout: "0s" # 0 = disabled; required for long SSE streams
models:
- id: "google/gemini-2.5-flash"
context_window: 1048576
- id: "google/gemini-2.5-pro"
context_window: 2097152
Serving many models
A local node serves one model, but a host can serve many in two ways:
- Passthrough to a multi-model engine. Run Kronk, LM Studio, Ollama, or vLLM
with several models loaded and point one node at it via
kronk,lmstudio, oropenai_passthrough; discovery picks up everything the engine serves. For llama.cpp specifically, Kronk is the closest fit: it pools many GGUFs in one process with a memory budget and eviction, wherellama-serverloads exactly one. - One node per local model. Run several
localnodes, each with its ownconfig.yaml,server.listenport, andzs.settlement_db_path. They can share the same operator identity and signing mnemonic.
Either way, the proxy routes each request by its model field across the
operator directory, so users see one merged catalog.
Image generation
Image generation runs on its own image_llm backend, independent of text, so a
node can serve text only, images only, or both. The provider is comfyui
(self-hosted), comfyui_cloud, or openai_passthrough (a hosted
OpenAI-compatible image API such as xAI), and each image model declares an
image: block in zs.models[] with its backend, per-image rate, and max_n
ceiling. Omit llm: while image_llm is set and the node runs image-only.
For a hosted backend, the zs-node init wizard does the setup: point
it at https://api.x.ai/v1 and it discovers each image model from /v1/models
(via the image_price field), derives your per-image rate, and writes both the
image_llm block and the per-model image: blocks. xAI image backends must
confirm zero data retention,
as on the text side.
Installing ComfyUI, workflow templates, edits, Comfy Cloud, output formats, in-chat image tools, image billing, and disk cleanup are covered in Image generation.
The health check gates what you advertise
The node continuously probes its backend's discovery endpoint: GET /v1/models
for openai_passthrough, llamacpp, and kronk, and LM Studio's native
/api/v1/models for lmstudio. The probe's cadence, timeout, and failure
threshold are set under llm.health_check.
While the backend is up, /v1/zs/details advertises your zs.models entries
and, when default_pricing is set, every model the backend lists. After
failure_threshold consecutive failures the backend is marked down:
/v1/zs/detailsadvertises zero models./v1/zs/reservereturns 503 (provider_unavailable) until a probe succeeds again.
The local and vertexai providers derive their model list from config and are
always reported available. Two metrics on the private listener expose the
state: zs_provider_healthy (1/0) and zs_provider_discovered_models.
xAI backends must confirm zero data retention
If your base_url points at xAI (api.x.ai), the node also requires Zero
Data Retention, automatically and with no opt-out. xAI retains API traffic
for 30 days by default; the only way to disable that is the org-level ZDR toggle
in the xAI console (Team Settings → Zero Data Retention), after which xAI
confirms it on every response. Enable ZDR for your xAI team before serving.
The node's startup usage probe is a real inference, so the ZDR guard sees its
response: a node whose xAI team has ZDR switched off does not start, and
exits with startup usage-reporting probe failed. There is no periodic ZDR
probe and no "zero models until confirmed" state. If ZDR is switched off under a
running node, each prompt-carrying request is refused before any stream frame is
emitted, at no charge to the payer. (store is separate and does not affect
xAI's audit retention.)
The openai_passthrough image backend has the same requirement when its
base_url is xAI, checked per request: an image response without the ZDR
confirmation is refused at no charge (403). See
Image generation.
OpenRouter backends have ZDR on by default
If your base_url points at OpenRouter (openrouter.ai), the node sets
provider.zdr: true and provider.data_collection: "deny" on every
prompt-carrying request, on both the text and image routes. The one way to turn
it off is Releasing the pin.
Your OpenRouter dashboard settings don't affect this, in either direction.
OpenRouter's privacy controls only tighten: a request-level zdr is OR'd with
your account and guardrail settings, so the node's request decides, whatever
you set at Settings → Privacy. The Data Training block there is separate from
Zero Data Retention and ships with "Allow free endpoints that train on request
data" switched on; data_collection: "deny" is what closes it.
A model with no ZDR endpoint on OpenRouter cannot be served; startup fails and names it. The node checks OpenRouter's public ZDR endpoint listing at startup:
startup zero-data-retention coverage check failed
text route: OpenRouter publishes no zero-data-retention endpoint for this
model ...: meta/llama-guard
For such a model, drop it, serve it from another upstream, or release the pin.
:free variants are the common case: a free endpoint that trains on requests
satisfies neither constraint. If the listing can't be read (OpenRouter is down,
or the response is partial), the node logs a warning and starts anyway; every
request still pins the constraints.
Because the node enforces the constraints, it advertises
retention: upstream_enforced for these models on /v1/zs/details, which
clients can show to payers. See
Privacy and retention.
Releasing the pin
To serve a model OpenRouter has no zero-retention route for, turn the pin off. It is set per route, since text and images can point at different upstreams:
llm:
allow_upstream_retention: true
image_llm:
allow_upstream_retention: true # independent of the text route
The node then starts, serves that model, and sends an unconstrained body. Four things change:
- Zero data retention is no longer required. A retaining endpoint can serve your payers' prompts.
data_collection: "deny"goes with it. That re-opens the Data Training block, which ships with "Allow free endpoints that train on request data" switched on. Operators most often miss this consequence.- This route advertises no
upstream_enforcedtier. Payers comparing operators see no retention guarantee from you, and yourconfig_hashchanges to match. - Both keys become caller-writable. Only the proxy strips a caller's
providerobject, so a client talking straight to your node can set them itself.
It does not touch the xAI gate, which only reads a header xAI already sends
and stays unconditional, and it does nothing on an upstream that isn't
OpenRouter. zs-node doctor warns both when it's live and when it's set where
it has no effect. Remove the key to restore the pin.
With no derived tier left, an upstream_zdr_declared that doctor previously
reported as having no effect here becomes live, and the route advertises
operator_declared instead. This is intended: you can still declare a
zero-retention agreement you hold, at the lowest tier.
Pricing the routes
A model can offer any combination of three priced request types, each with its own rate.
A request can also run tool calls that cost you, such as a frontier
vendor's server-side web search or your own zs_ built-ins. Charge for those
with Per-call tool pricing.
Text (token-based)
The primary route. You set an input rate and an output rate, each in
USD per million tokens, the unit vendors quote, so you can paste values
straight off a price card. At reserve the node converts each to microUSDC as
ceil(usd_per_1m × 1_000_000) (USDC has 6 decimals and is dollar-pegged) and
pins the result into the signed ticket.
Set rates fleet-wide in zs.default_pricing, per model under
zs.models[id].pricing, or both. An omitted or empty pricing: {}
inherits the default; an explicit {input_rate: 0, output_rate: 0} is a
free model that does not inherit. Which rate applies:
zs:
default_pricing:
input_rate: 0.15 # USD per 1M input tokens
output_rate: 0.60 # USD per 1M output tokens
models:
gpt-4o-mini:
pricing:
input_rate: 0.15
output_rate: 0.60
context:
context_window: 128000
max_output_tokens: 16384
input_modalities: ["text"]
output_modalities: ["text"]
gpt-4o:
pricing: {} # inherits default_pricing
context:
context_window: 128000
max_output_tokens: 16384
input_modalities: ["text", "image"] # vision-capable
output_modalities: ["text"]
max_active_tickets: 4 # tighter per-model concurrency cap
Your published rates are net of the protocol fee: set them to what you want to
keep. The protocol fee (the escrow contract's protocolFeeBps) is added on
top and paid by the payer; the node grosses up the escrowed max_price at
reserve to cover it. Do not inflate your rates to absorb the fee. See
Staking & economics.
Discounting cached input (cache_read_rate)
Many upstreams serve part of a repeated prompt from a prefix/prompt cache and
report that part as a "cached" token count. Pass the saving on to the user with
an optional third rate on any pricing block:
zs:
default_pricing:
input_rate: 0.15
output_rate: 0.60
cache_read_rate: 0.0375 # ~25% of input_rate — the cached-read discount
cache_read_rate (USD per 1M tokens) bills the cached-read part of the input in
place of input_rate; the rest of the input still bills at input_rate. Omit
it and cached reads bill at input_rate (no discount); 0 makes them free. Any
value up to input_rate only ever lowers the final charge, and the escrowed
maximum price is unaffected.
A real discount is also advertised on /v1/zs/details as
cache_read_rate_usd_per_1m, so clients can show the cached rate before a
request; the chat app shows it as a "Cached input" row in the model's pricing
detail. A model you don't discount advertises no cache rate, and a 0
advertises as free cached reads.
The discount takes effect only when your upstream reports a cached count.
OpenAI, z.ai/GLM, SGLang, DeepSeek, Vertex/Gemini, and a local llama.cpp all
report one; vLLM reports it only when started with
--enable-prompt-tokens-details (off by default); LM Studio and Ollama don't
report it, so the rate is a no-op there. To confirm what your backend reports,
set llm.openai.log_raw_usage: true for a capture run and watch the
raw upstream usage log line.
Long-context surcharge (long_context)
Some upstreams, xAI/Grok most notably, price with a context-size cliff:
below a prompt-token threshold you pay one rate, and at or above it a higher
rate for every token in the request (input, cached, and output all step up).
Declare it with a long_context block on any pricing (or default_pricing)
so a large-context request bills correctly instead of being turned away:
zs:
models:
grok-4.5:
pricing:
input_rate: 2.0 # base (below-threshold) tier
output_rate: 6.0
cache_read_rate: 0.30
long_context:
threshold_tokens: 200000 # prompt tokens at/above which the high tier applies
input_rate: 4.0 # high rates — each must be >= its base counterpart
output_rate: 12.0
cache_read_rate: 0.60 # optional — omitted, the base discount is carried forward
context:
context_window: 256000 # the model's TRUE max — see below
- The prompt (input) tokens alone decide the tier. A request whose reserved
input reaches
threshold_tokensbills at the high rates for the whole request; output tokens bill at the high output rate but never decide the tier. The node resolves the tier at reserve and pins its rates onto the ticket, so the user's app agrees on the price up front. - The high rates can't be lower than the base rates. The high
input_rateandoutput_rateare required and must each be at least the base rate, and the highcache_read_ratemust be at or below the highinput_rate, as in the base tier. The node rejects an inconsistent tier, or a tier on a fully free model, at startup. - A base cached-read discount carries into the high tier. Leave the high
cache_read_rateunset and the node scales the base discount by the input rate's step-up:0.30 × (4.0 / 2.0) = 0.60, exactly what xAI charges. Set it explicitly if your upstream differs. When the base tier has a discount, an explicit highcache_read_rateat or above the highinput_rateis rejected at startup. - Raise
context_windowto the model's true maximum rather than capping it at the threshold; withlong_contextset, a request above the threshold is admitted and billed at the high tier.threshold_tokensmust be strictly belowcontext_windowor startup fails, since otherwise the surcharge could never apply and you would pay the higher upstream price while billing the lower one. (Omitcontext_windowand the check doesn't apply; the backend's own limit is used at runtime.) - A large reserve doesn't mean a large bill. A reserve is a worst case: apps
add headroom for tool loops, and a multi-turn session that keeps its history
on the server must reserve the whole window, because your node can't see that
history. With a big
context_window, such requests reserve above the threshold even for a short prompt. So the node re-checks the tier when it bills, against the prompt actually sent, and charges the base rates if it never reached the threshold. It only adjusts downward: a request that did cross the threshold on a base-priced ticket still bills at the base rate. - It's advertised. The threshold and high rates ride
/v1/zs/details, so the chat app previews the higher price for a long prompt and shows an "Above<threshold>" row in the model's pricing detail. The authoritative price is still the one pinned into the ticket.
Dedicated image routes
Models that generate or edit images directly are priced per image:
image_rate for generation and image_edit_rate for edits, each in USD per
1024²-standard image. The rate is size-aware: the charge is
rate × (w × h) / 1024² × quality_multiplier, where the quality multiplier is
×0.25 (low), ×1.0 (medium / standard), or ×4.0 (high / hd). Each route is
declared per model, so a model can offer one without the other. A dedicated
image model (one with an image: block) needs no pricing entry and never
inherits default_pricing. See
How billing works.
In-loop image tools
A text model can produce images mid-response through a built-in image tool. These are priced per image as tool output, separately from the dedicated image routes: a dedicated route is model-to-image, while the tool rate covers images produced inside a text turn. A positive tool rate is what marks the tool as served.
How rates become the price a user sees
Users see and agree to an all-in price: your charge plus the protocol fee. When a request runs:
- The reserve ticket's maximum is sized from your rates and the request's
budget to cover the worst case. For text,
max_output_tokensis rounded up to the next 1K (minimum 1K) when sizingmax_price; input tokens are sized exactly. - The receipt's actual charge reflects what was really consumed, within that ceiling. Input is billed exactly per token; output is billed per token with a 1000-token floor (see the minimum charge).
- Unused microUSDC is refunded to the user automatically by the escrow contract at settlement.
Minimum charge, context, and concurrency
zs.min_chargeis a per-request floor, so a tiny successful response never settles for ~0. It is the larger of a token floor (output_tokens, default 1000: bill as if at least that many output tokens were produced, at the ticket's output rate) and a network-fee floor (algo_txns, in units of 1000 µALGO, default 0, converted through the ALGO/USD oracle). 7 is recommended foralgo_txns, to recover the ~7,000 µALGO of network fees you absorb per paid request on a live deployment (2 atopen()+ 5 atsettle()). Free models bypass the floor, and the µALGO component is skipped while the oracle is unavailable. An image model counts as paid whenever its image rate is above zero, but only thealgo_txnscomponent applies to it, since it produces no tokens. Seezs.min_charge.zs.models[id].contextis a defense-in-depth ceiling enforced at reserve time. The proxy is the primary sizing gatekeeper; this catches a misconfigured proxy or a client that bypasses it. Omit it and the node uses the context window your backend reports; with neither, no ceiling is enforced for that model.zs.max_active_tickets(default 16) is the per-model concurrency cap. Each model has its own independent pool, with no shared cross-model ceiling, andzs.models[id].max_active_ticketsoverrides it per model.reservereturns 429 when a model hits its cap.
The two ticket timers bound different things:
zs.ticket_ttl(default 5s) is the window betweenreserveand the inference POST: the client/proxy round trip only. It does not bound how long inference or streaming runs; a 5-minute stream on a 5s-TTL ticket completes normally. A POST that arrives after expiry gets HTTP 402ticket_invalid. Keeping it tight frees an abandoned reservation's slot in seconds.zs.default_expires_after(default 5m, per-model overrideexpires_after) is the settlement-complete deadline the contract reads to gaterefund_inactive. The co-signed settle still fires seconds after delivery; this only widens the window before an abandoned ticket becomes refundable, so slower or reasoning models settle without per-model tuning.
Free models are rationed by the contract
If you serve a model at 0 / 0, free tickets are rationed per payer by the
escrow contract: at most 20 per fixed 24-hour window, starting at that payer's
first free ticket. This is protocol-wide, not an operator setting: the cap lives
in the contract and applies across every node in the network. The count is spent
when the escrow opens and never given back, even if the free ticket is refunded
or lapses.
At reserve your node checks the payer's remaining allowance and refuses an
over-quota caller with 403 free_quota_exhausted, whose message and
Retry-After name the reset, so the caller gets a clear answer instead of an
escrow transaction that reverts on chain. A popular free model normally sees a
low background rate of these. Watch zs_reserve_free_quota_total. Paid
reserves never consult the quota.
A free model still costs the caller the usual Algorand network fees on the escrow open and settle, and a small minimum-balance amount is locked for the ticket's lifetime and refunded at settlement.
When a failed request still bills
A request that ends in an error almost always settles at 0 and refunds the payer in full. Billing depends on who caused the failure, not on what the client saw:
| Outcome | Billed | Why |
|---|---|---|
Your infrastructure failed: the backend returned 429/502/503/504, or the connection died with no HTTP status at all | 0, full refund | The caller did nothing wrong and could not have avoided it. A backend at its admission capacity lands here, so an over-large max_active_tickets costs a caller a request but never money. |
The request itself was refused: a non-transient 4xx such as context_length_exceeded, or a tool schema the backend rejected | the work actually measured | The request is what failed, and you really ran the iterations that got there. |
| The client hung up mid-stream | the work done | You produced output they walked away from; the force-finalization path pays you after the grace window. |
The difference shows in tool loops. A single-shot request that is refused has done no billable work yet, so it settles at 0 either way. A tool loop that completed two rounds and is refused on the third has two rounds of real, model-reported usage, and that is what gets billed. A capacity failure at the same point still bills 0.
You are never paid an estimate for a refusal. Only usage the model itself reported is billable here; the node will not synthesize a token count from frames it delivered. If your backend refuses a request without reporting usage, you absorb the cost.
A 500 counts as "the request was refused" and is billable, because a
500 is often a deterministic, caller-reproducible provider bug rather than a
transient outage. If your backend returns 500 on genuine crashes, you will be
billing callers for them; tell us and we'll move it.
Billing what you serve
Your receipt reports the actual counts you bill on: input and output token counts for text, and images produced for image routes and in-loop image tools. The receipt is signed and bound to the response body the user received (see The payment flow), so your billed counts have to match what you delivered. A receipt the client can't reconcile against the response it got won't settle.
The exact JSON shape of the details document and the receipt (field names, encodings, and units) is part of the wire protocol your node software implements, and a full protocol specification is coming soon.