Skip to main content

Serving models & pricing

Your node advertises what it offers in a single public details document, the same one clients probe for health (see Health & compatibility). Clients read it directly to build the live catalog users see in the model picker.

The details document​

The details document is served unauthenticated and unencrypted at /v1/zs/details. It carries:

FieldWhat it is
Protocol versionThe wire version your node speaks. Clients drop nodes on an incompatible major version (see Health & compatibility).
ModelsThe catalog. See Per-model advertisement.
Built-in toolsAny in-loop tools your node offers, such as web search or image generation, with their names and descriptions.
Encryption recipientYour current sealing key, signed and with an expiry (see Encryption & keys).
Config hashA sha256: fingerprint of your node's declared policy. See The config hash.

The config hash​

config_hash fingerprints your rate card and tool roster: the model catalog and its capabilities, built-in tools, rate limits, minimum charge, oracle source, and relay mode. It ignores deployment and identity details (listen addresses, file paths, operator_id, addresses, secrets) and knobs that fit a node to its hardware, such as web_read.max_download. Nodes running the same policy therefore advertise the same hash wherever they run, on a small VM or a large one.

Use it to confirm a fleet is uniformly configured (jq -r .config_hash across your nodes should match) or to notice that a node's policy drifted after an edit. It is advisory; nothing on the wire is gated on it. It has two limits:

  • It covers your declared policy, not what's reachable. If your upstream goes down, the models endpoint advertises nothing while config_hash stays the same.
  • Compare only across nodes on the same release. A release that changes which fields are fingerprinted changes every node's value, and nothing on the wire distinguishes that from an edit you made.

Per-model advertisement​

Each model in your catalog declares what it is and what it costs:

FieldWhat it is
Model idThe identifier users select and the one carried in requests.
SourceThe public artifact the model is (hf:org/model). Effectively required: a model with no checkable identity is dropped from your catalog, with only a WARN in the log. See Model identity & integrity.
Input and output modalitiesE.g. text in / text out, or text+image in. Drives the Vision, Generation, and Edit capability badges.
Context window and maximum output lengthThe model's size limits. Set the output length explicitly; see Set an output ceiling.
Reasoning supportWhether the model supports reasoning, the effort levels it allows, and a default. Drives the Reasoning badge.
Tool supportWhether the model can call tools. Drives the Tools badge.
TagsDescriptive and policy tags (e.g. content-policy labels). Auto-populated from the model's HuggingFace card when you declare a source; anything you add is merged in. They appear in the picker and can affect how Auto mode prioritizes the model.
RatesSee Pricing the routes.

Advertise only what you actually serve. Clients route to anything you list, so a model you no longer have means failed requests: remove it from zs.models when you stop serving it. The health check withdraws models on its own only when the whole backend is down, or when a model you serve through default_pricing alone disappears from the backend's list.

Set an output ceiling​

Set max_output_tokens for every model you serve. It is the model's output ceiling: advertised to clients, enforced when a request reserves payment, and, because clients take a declared ceiling at face value, the number a client that doesn't ask for a specific one receives.

zs:
models:
"Qwen/Qwen2.5-32B-Instruct":
context:
context_window: 131072
max_output_tokens: 32768

Leave it out and clients fall back to a quarter of the context window, capped at 32,768 tokens. Then:

  • Long answers get cut short. Even a 256K-context model stops at 32,768 tokens.
  • Large explicit requests aren't caught. A client that does ask for more meets no ceiling on your node, so the request reaches your inference server and fails there, after the user's payment has been reserved.

When every operator serving a model declares a ceiling, the proxy lowers an over-large request to the largest one advertised. Operators who declared that much stay eligible and those who declared less drop out; price and responsiveness then decide who gets the request. If you declare none, the proxy applies no operator's ceiling, and the request reaches your backend at full size and fails there.

danger

max_output_tokens must be below context_window. A reservation covers input plus output, so a ceiling at or above the window leaves no room for a prompt and your node is never selected for that model. The node still loads and advertises the model normally, so nothing looks wrong.

The rule applies to the window your node advertises: with a ceiling but no context_window, that is the window discovered from your backend.

A quarter of the context window is a reasonable starting point. zs-node init suggests it and asks about every model, and zs-node doctor flags a model with a missing or impossible ceiling.

Reasoning models​

The Reasoning badge and the per-request effort selector appear only when you declare the capability. LM Studio publishes it automatically; vLLM, llama.cpp, and any OpenAI-compatible passthrough (including OpenAI's own o-series / gpt-5) publish none, so declare it under the model's context:

zs:
models:
google/gemma-4-31b-it:
context:
reasoning:
supported: true
allowed_efforts: ["low", "medium", "high"]
default_effort: "medium"

Advertised effort tiers are also what let the client's toggle turn thinking on. A gemma-style template doesn't think until the request carries an effort. vLLM 0.25+ turns its enable_thinking flag on for any effort except "none", on both the Responses reasoning.effort and the Chat reasoning_effort, so the client's selection turns thinking on with no server-side change.

info

On a vLLM older than 0.25, or a template with a non-standard kwarg name, enable thinking on your inference server instead, e.g. vLLM's --default-chat-template-kwargs '{"enable_thinking": true}'. The node reads both vLLM's current reasoning field and the older reasoning_content, and a backend's native /v1/responses reasoning items pass through untouched.

warning

A thinking model spends part of its output-token budget on the chain-of-thought before the visible answer, so on a hard prompt it can use the whole budget and return an empty, cut-off answer. The surest fix is an explicit output ceiling for the model.

Choosing a backend​

Your node doesn't run inference itself. It sits in front of an inference engine, on loopback or over the network, chosen with one setting: llm.provider. One provider per node. Every provider and its defaults are listed under llm. An empty provider ("") disables text inference, for image-only or relay-only mode.

The most flexible option. Set the upstream's API root in llm.openai.base_url, and supply the key through the NODE_LLM_OPENAI_API_KEY env var rather than committing it:

llm:
provider: "openai_passthrough"
openai:
base_url: "https://api.openai.com/v1" # include the version segment yourself
api_key: "" # prefer NODE_LLM_OPENAI_API_KEY env var
timeout: "0s" # 0 = disabled; required for long streams
translate_responses_to_chat: false # true ONLY for upstreams without /v1/responses
Every node must serve /v1/responses

Payers assume the Responses API is there. At startup the node sends one zero-token probe to your upstream's /v1/responses and exits on a 404, 405, or 501 rather than advertising models it cannot serve. Any other outcome (a 4xx about the request itself, a timeout, a connection error) says nothing about whether the route exists, so the node logs a warning and boots. With translate_responses_to_chat: true the node serves the route itself and skips the probe.

base_url is appended to directly (/chat/completions, /responses, /models), so include the version segment, usually /v1, yourself. For a non-/v1-rooted upstream, give the full prefix (e.g. https://api.z.ai/api/paas/v4). A local engine works the same way, e.g. vLLM at http://127.0.0.1:8000/v1 or Ollama at http://127.0.0.1:11434/v1.

Set translate_responses_to_chat: true when the upstream has no native /v1/responses (llama.cpp's own server and Ollama, and vLLM/TGI unless you've enabled its Responses endpoint); the node then emulates the route on top of /v1/chat/completions. Leave it false for a backend that implements Responses natively (OpenAI, Azure OpenAI, LM Studio). If you leave it false against a backend with no Responses route, the node refuses to start and names the route. If you set it true against a backend that has one, the node boots but emulates the route and loses native Responses behavior. zs-node doctor settles it empirically; see doctor.

Serving many models​

A local node serves one model, but a host can serve many in two ways:

  • Passthrough to a multi-model engine. Run Kronk, LM Studio, Ollama, or vLLM with several models loaded and point one node at it via kronk, lmstudio, or openai_passthrough; discovery picks up everything the engine serves. For llama.cpp specifically, Kronk is the closest fit: it pools many GGUFs in one process with a memory budget and eviction, where llama-server loads exactly one.
  • One node per local model. Run several local nodes, each with its own config.yaml, server.listen port, and zs.settlement_db_path. They can share the same operator identity and signing mnemonic.

Either way, the proxy routes each request by its model field across the operator directory, so users see one merged catalog.

Image generation​

Image generation runs on its own image_llm backend, independent of text, so a node can serve text only, images only, or both. The provider is comfyui (self-hosted), comfyui_cloud, or openai_passthrough (a hosted OpenAI-compatible image API such as xAI), and each image model declares an image: block in zs.models[] with its backend, per-image rate, and max_n ceiling. Omit llm: while image_llm is set and the node runs image-only.

For a hosted backend, the zs-node init wizard does the setup: point it at https://api.x.ai/v1 and it discovers each image model from /v1/models (via the image_price field), derives your per-image rate, and writes both the image_llm block and the per-model image: blocks. xAI image backends must confirm zero data retention, as on the text side.

Installing ComfyUI, workflow templates, edits, Comfy Cloud, output formats, in-chat image tools, image billing, and disk cleanup are covered in Image generation.

The health check gates what you advertise​

The node continuously probes its backend's discovery endpoint: GET /v1/models for openai_passthrough, llamacpp, and kronk, and LM Studio's native /api/v1/models for lmstudio. The probe's cadence, timeout, and failure threshold are set under llm.health_check.

While the backend is up, /v1/zs/details advertises your zs.models entries and, when default_pricing is set, every model the backend lists. After failure_threshold consecutive failures the backend is marked down:

  • /v1/zs/details advertises zero models.
  • /v1/zs/reserve returns 503 (provider_unavailable) until a probe succeeds again.

The local and vertexai providers derive their model list from config and are always reported available. Two metrics on the private listener expose the state: zs_provider_healthy (1/0) and zs_provider_discovered_models.

xAI backends must confirm zero data retention​

If your base_url points at xAI (api.x.ai), the node also requires Zero Data Retention, automatically and with no opt-out. xAI retains API traffic for 30 days by default; the only way to disable that is the org-level ZDR toggle in the xAI console (Team Settings → Zero Data Retention), after which xAI confirms it on every response. Enable ZDR for your xAI team before serving.

The node's startup usage probe is a real inference, so the ZDR guard sees its response: a node whose xAI team has ZDR switched off does not start, and exits with startup usage-reporting probe failed. There is no periodic ZDR probe and no "zero models until confirmed" state. If ZDR is switched off under a running node, each prompt-carrying request is refused before any stream frame is emitted, at no charge to the payer. (store is separate and does not affect xAI's audit retention.)

The openai_passthrough image backend has the same requirement when its base_url is xAI, checked per request: an image response without the ZDR confirmation is refused at no charge (403). See Image generation.

OpenRouter backends have ZDR on by default​

If your base_url points at OpenRouter (openrouter.ai), the node sets provider.zdr: true and provider.data_collection: "deny" on every prompt-carrying request, on both the text and image routes. The one way to turn it off is Releasing the pin.

Your OpenRouter dashboard settings don't affect this, in either direction. OpenRouter's privacy controls only tighten: a request-level zdr is OR'd with your account and guardrail settings, so the node's request decides, whatever you set at Settings → Privacy. The Data Training block there is separate from Zero Data Retention and ships with "Allow free endpoints that train on request data" switched on; data_collection: "deny" is what closes it.

A model with no ZDR endpoint on OpenRouter cannot be served; startup fails and names it. The node checks OpenRouter's public ZDR endpoint listing at startup:

startup zero-data-retention coverage check failed
text route: OpenRouter publishes no zero-data-retention endpoint for this
model ...: meta/llama-guard

For such a model, drop it, serve it from another upstream, or release the pin. :free variants are the common case: a free endpoint that trains on requests satisfies neither constraint. If the listing can't be read (OpenRouter is down, or the response is partial), the node logs a warning and starts anyway; every request still pins the constraints.

info

Because the node enforces the constraints, it advertises retention: upstream_enforced for these models on /v1/zs/details, which clients can show to payers. See Privacy and retention.

Releasing the pin​

To serve a model OpenRouter has no zero-retention route for, turn the pin off. It is set per route, since text and images can point at different upstreams:

llm:
allow_upstream_retention: true
image_llm:
allow_upstream_retention: true # independent of the text route

The node then starts, serves that model, and sends an unconstrained body. Four things change:

  1. Zero data retention is no longer required. A retaining endpoint can serve your payers' prompts.
  2. data_collection: "deny" goes with it. That re-opens the Data Training block, which ships with "Allow free endpoints that train on request data" switched on. Operators most often miss this consequence.
  3. This route advertises no upstream_enforced tier. Payers comparing operators see no retention guarantee from you, and your config_hash changes to match.
  4. Both keys become caller-writable. Only the proxy strips a caller's provider object, so a client talking straight to your node can set them itself.

It does not touch the xAI gate, which only reads a header xAI already sends and stays unconditional, and it does nothing on an upstream that isn't OpenRouter. zs-node doctor warns both when it's live and when it's set where it has no effect. Remove the key to restore the pin.

note

With no derived tier left, an upstream_zdr_declared that doctor previously reported as having no effect here becomes live, and the route advertises operator_declared instead. This is intended: you can still declare a zero-retention agreement you hold, at the lowest tier.

Pricing the routes​

A model can offer any combination of three priced request types, each with its own rate.

info

A request can also run tool calls that cost you, such as a frontier vendor's server-side web search or your own zs_ built-ins. Charge for those with Per-call tool pricing.

Text (token-based)​

The primary route. You set an input rate and an output rate, each in USD per million tokens, the unit vendors quote, so you can paste values straight off a price card. At reserve the node converts each to microUSDC as ceil(usd_per_1m × 1_000_000) (USDC has 6 decimals and is dollar-pegged) and pins the result into the signed ticket.

Set rates fleet-wide in zs.default_pricing, per model under zs.models[id].pricing, or both. An omitted or empty pricing: {} inherits the default; an explicit {input_rate: 0, output_rate: 0} is a free model that does not inherit. Which rate applies:

zs:
default_pricing:
input_rate: 0.15 # USD per 1M input tokens
output_rate: 0.60 # USD per 1M output tokens

models:
gpt-4o-mini:
pricing:
input_rate: 0.15
output_rate: 0.60
context:
context_window: 128000
max_output_tokens: 16384
input_modalities: ["text"]
output_modalities: ["text"]
gpt-4o:
pricing: {} # inherits default_pricing
context:
context_window: 128000
max_output_tokens: 16384
input_modalities: ["text", "image"] # vision-capable
output_modalities: ["text"]
max_active_tickets: 4 # tighter per-model concurrency cap
info

Your published rates are net of the protocol fee: set them to what you want to keep. The protocol fee (the escrow contract's protocolFeeBps) is added on top and paid by the payer; the node grosses up the escrowed max_price at reserve to cover it. Do not inflate your rates to absorb the fee. See Staking & economics.

Discounting cached input (cache_read_rate)​

Many upstreams serve part of a repeated prompt from a prefix/prompt cache and report that part as a "cached" token count. Pass the saving on to the user with an optional third rate on any pricing block:

zs:
default_pricing:
input_rate: 0.15
output_rate: 0.60
cache_read_rate: 0.0375 # ~25% of input_rate — the cached-read discount

cache_read_rate (USD per 1M tokens) bills the cached-read part of the input in place of input_rate; the rest of the input still bills at input_rate. Omit it and cached reads bill at input_rate (no discount); 0 makes them free. Any value up to input_rate only ever lowers the final charge, and the escrowed maximum price is unaffected.

A real discount is also advertised on /v1/zs/details as cache_read_rate_usd_per_1m, so clients can show the cached rate before a request; the chat app shows it as a "Cached input" row in the model's pricing detail. A model you don't discount advertises no cache rate, and a 0 advertises as free cached reads.

The discount takes effect only when your upstream reports a cached count. OpenAI, z.ai/GLM, SGLang, DeepSeek, Vertex/Gemini, and a local llama.cpp all report one; vLLM reports it only when started with --enable-prompt-tokens-details (off by default); LM Studio and Ollama don't report it, so the rate is a no-op there. To confirm what your backend reports, set llm.openai.log_raw_usage: true for a capture run and watch the raw upstream usage log line.

Long-context surcharge (long_context)​

Some upstreams, xAI/Grok most notably, price with a context-size cliff: below a prompt-token threshold you pay one rate, and at or above it a higher rate for every token in the request (input, cached, and output all step up). Declare it with a long_context block on any pricing (or default_pricing) so a large-context request bills correctly instead of being turned away:

zs:
models:
grok-4.5:
pricing:
input_rate: 2.0 # base (below-threshold) tier
output_rate: 6.0
cache_read_rate: 0.30
long_context:
threshold_tokens: 200000 # prompt tokens at/above which the high tier applies
input_rate: 4.0 # high rates — each must be >= its base counterpart
output_rate: 12.0
cache_read_rate: 0.60 # optional — omitted, the base discount is carried forward
context:
context_window: 256000 # the model's TRUE max — see below
  • The prompt (input) tokens alone decide the tier. A request whose reserved input reaches threshold_tokens bills at the high rates for the whole request; output tokens bill at the high output rate but never decide the tier. The node resolves the tier at reserve and pins its rates onto the ticket, so the user's app agrees on the price up front.
  • The high rates can't be lower than the base rates. The high input_rate and output_rate are required and must each be at least the base rate, and the high cache_read_rate must be at or below the high input_rate, as in the base tier. The node rejects an inconsistent tier, or a tier on a fully free model, at startup.
  • A base cached-read discount carries into the high tier. Leave the high cache_read_rate unset and the node scales the base discount by the input rate's step-up: 0.30 × (4.0 / 2.0) = 0.60, exactly what xAI charges. Set it explicitly if your upstream differs. When the base tier has a discount, an explicit high cache_read_rate at or above the high input_rate is rejected at startup.
  • Raise context_window to the model's true maximum rather than capping it at the threshold; with long_context set, a request above the threshold is admitted and billed at the high tier. threshold_tokens must be strictly below context_window or startup fails, since otherwise the surcharge could never apply and you would pay the higher upstream price while billing the lower one. (Omit context_window and the check doesn't apply; the backend's own limit is used at runtime.)
  • A large reserve doesn't mean a large bill. A reserve is a worst case: apps add headroom for tool loops, and a multi-turn session that keeps its history on the server must reserve the whole window, because your node can't see that history. With a big context_window, such requests reserve above the threshold even for a short prompt. So the node re-checks the tier when it bills, against the prompt actually sent, and charges the base rates if it never reached the threshold. It only adjusts downward: a request that did cross the threshold on a base-priced ticket still bills at the base rate.
  • It's advertised. The threshold and high rates ride /v1/zs/details, so the chat app previews the higher price for a long prompt and shows an "Above <threshold>" row in the model's pricing detail. The authoritative price is still the one pinned into the ticket.

Dedicated image routes​

Models that generate or edit images directly are priced per image: image_rate for generation and image_edit_rate for edits, each in USD per 1024²-standard image. The rate is size-aware: the charge is rate × (w × h) / 1024² × quality_multiplier, where the quality multiplier is ×0.25 (low), ×1.0 (medium / standard), or ×4.0 (high / hd). Each route is declared per model, so a model can offer one without the other. A dedicated image model (one with an image: block) needs no pricing entry and never inherits default_pricing. See How billing works.

In-loop image tools​

A text model can produce images mid-response through a built-in image tool. These are priced per image as tool output, separately from the dedicated image routes: a dedicated route is model-to-image, while the tool rate covers images produced inside a text turn. A positive tool rate is what marks the tool as served.

How rates become the price a user sees​

Users see and agree to an all-in price: your charge plus the protocol fee. When a request runs:

  • The reserve ticket's maximum is sized from your rates and the request's budget to cover the worst case. For text, max_output_tokens is rounded up to the next 1K (minimum 1K) when sizing max_price; input tokens are sized exactly.
  • The receipt's actual charge reflects what was really consumed, within that ceiling. Input is billed exactly per token; output is billed per token with a 1000-token floor (see the minimum charge).
  • Unused microUSDC is refunded to the user automatically by the escrow contract at settlement.

Minimum charge, context, and concurrency​

  • zs.min_charge is a per-request floor, so a tiny successful response never settles for ~0. It is the larger of a token floor (output_tokens, default 1000: bill as if at least that many output tokens were produced, at the ticket's output rate) and a network-fee floor (algo_txns, in units of 1000 µALGO, default 0, converted through the ALGO/USD oracle). 7 is recommended for algo_txns, to recover the ~7,000 µALGO of network fees you absorb per paid request on a live deployment (2 at open() + 5 at settle()). Free models bypass the floor, and the µALGO component is skipped while the oracle is unavailable. An image model counts as paid whenever its image rate is above zero, but only the algo_txns component applies to it, since it produces no tokens. See zs.min_charge.
  • zs.models[id].context is a defense-in-depth ceiling enforced at reserve time. The proxy is the primary sizing gatekeeper; this catches a misconfigured proxy or a client that bypasses it. Omit it and the node uses the context window your backend reports; with neither, no ceiling is enforced for that model.
  • zs.max_active_tickets (default 16) is the per-model concurrency cap. Each model has its own independent pool, with no shared cross-model ceiling, and zs.models[id].max_active_tickets overrides it per model. reserve returns 429 when a model hits its cap.

The two ticket timers bound different things:

  • zs.ticket_ttl (default 5s) is the window between reserve and the inference POST: the client/proxy round trip only. It does not bound how long inference or streaming runs; a 5-minute stream on a 5s-TTL ticket completes normally. A POST that arrives after expiry gets HTTP 402 ticket_invalid. Keeping it tight frees an abandoned reservation's slot in seconds.
  • zs.default_expires_after (default 5m, per-model override expires_after) is the settlement-complete deadline the contract reads to gate refund_inactive. The co-signed settle still fires seconds after delivery; this only widens the window before an abandoned ticket becomes refundable, so slower or reasoning models settle without per-model tuning.

Free models are rationed by the contract​

If you serve a model at 0 / 0, free tickets are rationed per payer by the escrow contract: at most 20 per fixed 24-hour window, starting at that payer's first free ticket. This is protocol-wide, not an operator setting: the cap lives in the contract and applies across every node in the network. The count is spent when the escrow opens and never given back, even if the free ticket is refunded or lapses.

At reserve your node checks the payer's remaining allowance and refuses an over-quota caller with 403 free_quota_exhausted, whose message and Retry-After name the reset, so the caller gets a clear answer instead of an escrow transaction that reverts on chain. A popular free model normally sees a low background rate of these. Watch zs_reserve_free_quota_total. Paid reserves never consult the quota.

A free model still costs the caller the usual Algorand network fees on the escrow open and settle, and a small minimum-balance amount is locked for the ticket's lifetime and refunded at settlement.

When a failed request still bills​

A request that ends in an error almost always settles at 0 and refunds the payer in full. Billing depends on who caused the failure, not on what the client saw:

OutcomeBilledWhy
Your infrastructure failed: the backend returned 429/502/503/504, or the connection died with no HTTP status at all0, full refundThe caller did nothing wrong and could not have avoided it. A backend at its admission capacity lands here, so an over-large max_active_tickets costs a caller a request but never money.
The request itself was refused: a non-transient 4xx such as context_length_exceeded, or a tool schema the backend rejectedthe work actually measuredThe request is what failed, and you really ran the iterations that got there.
The client hung up mid-streamthe work doneYou produced output they walked away from; the force-finalization path pays you after the grace window.

The difference shows in tool loops. A single-shot request that is refused has done no billable work yet, so it settles at 0 either way. A tool loop that completed two rounds and is refused on the third has two rounds of real, model-reported usage, and that is what gets billed. A capacity failure at the same point still bills 0.

You are never paid an estimate for a refusal. Only usage the model itself reported is billable here; the node will not synthesize a token count from frames it delivered. If your backend refuses a request without reporting usage, you absorb the cost.

info

A 500 counts as "the request was refused" and is billable, because a 500 is often a deterministic, caller-reproducible provider bug rather than a transient outage. If your backend returns 500 on genuine crashes, you will be billing callers for them; tell us and we'll move it.

Billing what you serve​

Your receipt reports the actual counts you bill on: input and output token counts for text, and images produced for image routes and in-loop image tools. The receipt is signed and bound to the response body the user received (see The payment flow), so your billed counts have to match what you delivered. A receipt the client can't reconcile against the response it got won't settle.

info

The exact JSON shape of the details document and the receipt (field names, encodings, and units) is part of the wire protocol your node software implements, and a full protocol specification is coming soon.