Connecting AI tools
The proxy is a drop-in OpenAI endpoint: point any tool that speaks the OpenAI
API at http://localhost:9376 (or your configured listen address) instead of
api.openai.com. Start the proxy first (zs-proxy proxy start). connect
reads the live endpoint and model list, and tells you to start the proxy if it
isn't up.
The connect command
zs-proxy connect # detect installed tools → confirm → configure
zs-proxy connect opencode # configure just one tool
zs-proxy connect opencode --model qwen3-coder # ...with a specific default model
zs-proxy connect --all # configure every detected tool, no prompts
zs-proxy connect --list # show what would be detected/configured — no writes
zs-proxy connect opencode --print # render the merged config to stdout — no writes
For each tool, connect merges a zerosignal provider into the tool's
existing config. Your other settings are kept (and comments too, for JSONC
configs like opencode's), the original file is backed up to <file>.bak, and
the write is atomic. Re-running it is safe and idempotent. Hermes' .env is the
exception: it is spliced line by line and gets no backup — see
its guide.
Every flag, including --plugin (Hermes only) and --url, is listed in the
Proxy CLI reference.
Seeded model prices
In opencode, OpenClaw, and Pi, connect seeds each priced model's cost (USD
per 1M tokens), so the tool's spend display shows an estimate instead of $0.
For Aider it writes the same rates per token (input_cost_per_token,
output_cost_per_token, cache_read_input_token_cost) into its
model metadata file, with no cache-write rate. Each rate is
the cheapest serving operator's advertised rate, grossed up by the
protocol fee — what you actually pay. In the other three tools, the
cache-write rate equals the input rate, since ZeroSignal has no cache-write
premium. Free or unpriced
models get a zeroed cost in OpenClaw and Pi, no cost block in opencode, and
no cost keys in Aider.
These are estimates. The authoritative charge is the amount the chosen operator
signs on-chain for each request, which connect can't know in advance, and
routing may pick a different-priced operator. See Pricing for
the full model.
Supported tools
| Tool | connect name |
|---|---|
| opencode | opencode |
| OpenClaw | openclaw |
| Pi | pi |
| Hermes | hermes |
| Aider | aider |
| Continue | continue |
| Codex CLI | codex |
| Anything else | generic — prints the OPENAI_BASE_URL / OPENAI_API_KEY env vars to set yourself |
No API key is needed, since your wallet's on-chain seal is the credential.
connect writes a placeholder key wherever a tool requires a non-empty one.
Manual setup
For a tool that isn't in the list, or any OpenAI SDK / script, point it at:
- Base URL:
http://localhost:9376/v1(or your configuredserver.listen— see Configuration) - API key: anything non-empty; the proxy ignores it
export OPENAI_BASE_URL="http://localhost:9376/v1"
export OPENAI_API_KEY="not-checked"
For popular GUI apps — SillyTavern, Open WebUI, LibreChat, Chatbox, Cherry Studio — there are screen-by-screen walkthroughs in How-to guides.
What the proxy serves
| Endpoint | Notes |
|---|---|
POST /v1/chat/completions | Standard chat completions, streaming or buffered. Prefer streaming — see below. |
POST /v1/responses | The Responses API, streaming or buffered. Prefer streaming — see below. |
POST /v1/completions | Legacy text-completion shape, translated to a chat request under the hood. The prompt must be a single string (a one-element string array works too). |
POST /v1/images/generations | Image generation. |
POST /v1/images/edits | Image edits (multipart upload). |
GET /v1/models / GET /v1/models/{id} | The live, on-chain model catalog — the same one the chat app's model picker reads. |
GET /v1/zs/operators | The operator directory — per-operator ids, models, and pricing. Where the ids for routing preferences come from. |
GET /v1/zs/details | The network's aggregate capability and pricing summary (union across operators, no per-operator ids). |
GET /healthz | Liveness probe — handy for scripts and supervisors. |
All three discovery rows cover production nodes only unless you set
allow_staging.
Requests must name a concrete model id from /v1/models; the chat app's
"Auto" model lane is an app feature, not a proxy one. The proxy picks the
operator serving that model — see Models & operators.
Response length
Most tools send a max_tokens (or max_completion_tokens, or
max_output_tokens) taken from their own built-in model card. Operators
advertise their own output ceiling per model, and the two often disagree: a
tool's card may say 65,536 for a model every operator serving it caps at 32,768.
The proxy lowers an over-large max_tokens to the highest ceiling advertised
for that model, so the operators offering it stay eligible. You still get at
most what you asked for, and the escrow reserve shrinks to match. If any
operator serving the model advertises no ceiling at all, your value passes
through unchanged, because that operator can honor it.
Only operators you could be routed to count, so a staging node or one your
routing preferences exclude never sets your ceiling.
The exception is a provider.max_price ceiling: price filtering happens after
this step, so an operator it rules out still counts.
When the proxy lowers your value, the response carries
X-Zs-Max-Output-Clamped with the ceiling it applied. Check it if a reply stops
earlier than your tool expected: without it, a reply cut off at the lowered
ceiling looks the same as the model choosing to stop.
Send nothing and the proxy picks a ceiling sized to whichever operator it routes
you to (see fallback_max_output_tokens in Configuration).
Setting a smaller value than you need is still the best way to keep replies and
cost tight.
Prefer streaming
Set stream: true on chat and responses requests whenever your tool lets you.
A buffered (non-streaming) request sends nothing back over the connection
until the model has finished the entire reply. Any layer between the proxy and
the operator — a CDN, a load balancer, a reverse proxy — can have its own idle
timeout, and 60 seconds is a common default. If the reply takes longer than
that, the layer gives up on the silent connection and answers 504, even
though the operator was working normally.
Streaming sends tokens as they're produced, plus an invisible keepalive during quiet stretches such as a long reasoning pass. Keepalives prevent intermediary idle timeouts, and you see the reply as it arrives.
Buffered requests are fine for short replies. Streaming matters more as the expected output grows: big contexts, reasoning models, and agent turns that plan before answering.
Choosing where a request goes
By default the proxy routes the same way the chat app does. For finer control
over which operator handles a request — pinning to one, setting a price
ceiling, requiring tool support, or trading latency for cost — use the
provider field on the request body. See
Proxy routing preferences.
Seeing what you paid
For a non-streaming response, the cost is in the X-Zs-Inference-Amount and
related X-Zs-* response headers. For a streaming response, set the header
X-Zs-Usage-Frame: 1 on your request and the proxy adds an event: zs.usage
frame to the stream with the same breakdown as JSON. The frame is opt-in
because its payload is not a valid OpenAI chunk: strict OpenAI clients that
parse every data: line fail on it. Set the header only from a client that
handles the frame. See Pricing for what the numbers mean.
Troubleshooting
context_length_exceeded on a conversation that isn't long
A 400 like this one:
no operator serves <model> with input_tokens=36576 max_output_tokens=65536
operators advertising this model:
- op 1: max_output_tokens=32768 (over by 32768)
names the limit each candidate operator exceeded. Read the per-operator lines rather than the error code: despite its name, this error fires whenever no operator fits the request, and the two causes look different:
-
context_window=N (request needs M)— your conversation plus the reserved output doesn't fit the model's window by the proxy's size estimate, which runs about twice a real token count on prose, so this can fire well before the window is actually full. Start a new conversation, trim the history, or pick a model with a larger window. -
max_output_tokens=N (over by M)— your tool asked for more output than that operator caps at. The proxy normally lowers this for you (see Response length), so if this is the only reason every operator was dropped, the lowering didn't apply. Set a smaller max response length in your tool, or check the two things that prevent the lowering:- an operator serving this model advertises no ceiling, so your value
passes through unchanged; that operator's entry has no
max_output_tokens=line; - you set a
provider.max_priceceiling (Routing preferences) that rules out the operators with the larger ceilings.
If neither applies, check
zs-proxy version: releases before 0.14.3 don't lowermax_tokens, so upgrade. - an operator serving this model advertises no ceiling, so your value
passes through unchanged; that operator's entry has no
The same code without any per-operator lines comes from the operator after the request was sent: its backend refused the conversation as too long, usually because it serves a smaller window than the operator advertises. There is no inference charge. Retry, which may reach a different operator, or trim the history.
504 errors, usually after about a minute
An error like this one:
504 operator returned HTTP 504 with a body this proxy could not parse as an
OpenAI error (content-type "text/html")
comes from a CDN or load balancer between the proxy and the operator, not from the operator's software, which is why the body is an HTML page rather than an API error. These arrive at a round number, typically around 60 seconds.
Almost always this is a buffered request that ran long. Set
stream: true first — see Prefer streaming.
| Stage | Retried by the proxy? |
|---|---|
| Setting up the request | Yes. The proxy retries over different network paths and reroutes to other operators serving the same model, so most transient trouble never reaches you. |
| Paid for and sent | No. The request goes out exactly once, with no retry or failover, because the payment is already committed on-chain. A 504 here reaches your tool as-is. |
Repeated attempts in your logs are therefore your own tool's retry logic. Each of those attempts is a separate paid request.
Every error the proxy returns carries an X-Should-Retry header, which the
official OpenAI SDKs (Python, Node, Go) follow instead of retrying every 5xx.
It says false for a request that was already paid for, and for an error a
retry won't fix, so tools built on those SDKs don't resend those requests. A tool
with its own retry logic may ignore the header and resend anyway.
The unused part of the escrow comes back at settlement, and a ticket nobody claims is refunded automatically after it expires. If the node had already generated output when the connection dropped, it can still bill for that work.
If you're already streaming and still see 504s from the same operator, that
operator's front door is unhealthy. A retry normally lands on a different one;
to rule it out for good, add its id to ignore in
routing preferences.
What's next
- How-to guides — per-tool setup for every supported tool, plus the GUI chat apps (SillyTavern, Open WebUI, LibreChat, Chatbox, Cherry Studio).
- Wallet & funding — keep the proxy funded.
- Configuration — change the listen address, network, or spend caps.
- Proxy CLI reference — every
connectflag and the rest of the commands.