Skip to main content

Connecting AI tools

The proxy is a drop-in OpenAI endpoint: point any tool that speaks the OpenAI API at http://localhost:9376 (or your configured listen address) instead of api.openai.com. Start the proxy first (zs-proxy proxy start). connect reads the live endpoint and model list, and tells you to start the proxy if it isn't up.

The connect command​

zs-proxy connect # detect installed tools → confirm → configure
zs-proxy connect opencode # configure just one tool
zs-proxy connect opencode --model qwen3-coder # ...with a specific default model
zs-proxy connect --all # configure every detected tool, no prompts
zs-proxy connect --list # show what would be detected/configured — no writes
zs-proxy connect opencode --print # render the merged config to stdout — no writes

For each tool, connect merges a zerosignal provider into the tool's existing config. Your other settings are kept (and comments too, for JSONC configs like opencode's), the original file is backed up to <file>.bak, and the write is atomic. Re-running it is safe and idempotent. Hermes' .env is the exception: it is spliced line by line and gets no backup — see its guide.

Every flag, including --plugin (Hermes only) and --url, is listed in the Proxy CLI reference.

Seeded model prices​

In opencode, OpenClaw, and Pi, connect seeds each priced model's cost (USD per 1M tokens), so the tool's spend display shows an estimate instead of $0. For Aider it writes the same rates per token (input_cost_per_token, output_cost_per_token, cache_read_input_token_cost) into its model metadata file, with no cache-write rate. Each rate is the cheapest serving operator's advertised rate, grossed up by the protocol fee — what you actually pay. In the other three tools, the cache-write rate equals the input rate, since ZeroSignal has no cache-write premium. Free or unpriced models get a zeroed cost in OpenClaw and Pi, no cost block in opencode, and no cost keys in Aider.

These are estimates. The authoritative charge is the amount the chosen operator signs on-chain for each request, which connect can't know in advance, and routing may pick a different-priced operator. See Pricing for the full model.

Supported tools​

Toolconnect name
opencodeopencode
OpenClawopenclaw
Pipi
Hermeshermes
Aideraider
Continuecontinue
Codex CLIcodex
Anything elsegeneric — prints the OPENAI_BASE_URL / OPENAI_API_KEY env vars to set yourself

No API key is needed, since your wallet's on-chain seal is the credential. connect writes a placeholder key wherever a tool requires a non-empty one.

Manual setup​

For a tool that isn't in the list, or any OpenAI SDK / script, point it at:

  • Base URL: http://localhost:9376/v1 (or your configured server.listen — see Configuration)
  • API key: anything non-empty; the proxy ignores it
export OPENAI_BASE_URL="http://localhost:9376/v1"
export OPENAI_API_KEY="not-checked"
info

For popular GUI apps — SillyTavern, Open WebUI, LibreChat, Chatbox, Cherry Studio — there are screen-by-screen walkthroughs in How-to guides.

What the proxy serves​

EndpointNotes
POST /v1/chat/completionsStandard chat completions, streaming or buffered. Prefer streaming — see below.
POST /v1/responsesThe Responses API, streaming or buffered. Prefer streaming — see below.
POST /v1/completionsLegacy text-completion shape, translated to a chat request under the hood. The prompt must be a single string (a one-element string array works too).
POST /v1/images/generationsImage generation.
POST /v1/images/editsImage edits (multipart upload).
GET /v1/models / GET /v1/models/{id}The live, on-chain model catalog — the same one the chat app's model picker reads.
GET /v1/zs/operatorsThe operator directory — per-operator ids, models, and pricing. Where the ids for routing preferences come from.
GET /v1/zs/detailsThe network's aggregate capability and pricing summary (union across operators, no per-operator ids).
GET /healthzLiveness probe — handy for scripts and supervisors.

All three discovery rows cover production nodes only unless you set allow_staging.

Requests must name a concrete model id from /v1/models; the chat app's "Auto" model lane is an app feature, not a proxy one. The proxy picks the operator serving that model — see Models & operators.

Response length​

Most tools send a max_tokens (or max_completion_tokens, or max_output_tokens) taken from their own built-in model card. Operators advertise their own output ceiling per model, and the two often disagree: a tool's card may say 65,536 for a model every operator serving it caps at 32,768.

The proxy lowers an over-large max_tokens to the highest ceiling advertised for that model, so the operators offering it stay eligible. You still get at most what you asked for, and the escrow reserve shrinks to match. If any operator serving the model advertises no ceiling at all, your value passes through unchanged, because that operator can honor it.

Only operators you could be routed to count, so a staging node or one your routing preferences exclude never sets your ceiling. The exception is a provider.max_price ceiling: price filtering happens after this step, so an operator it rules out still counts.

When the proxy lowers your value, the response carries X-Zs-Max-Output-Clamped with the ceiling it applied. Check it if a reply stops earlier than your tool expected: without it, a reply cut off at the lowered ceiling looks the same as the model choosing to stop.

Send nothing and the proxy picks a ceiling sized to whichever operator it routes you to (see fallback_max_output_tokens in Configuration). Setting a smaller value than you need is still the best way to keep replies and cost tight.

Prefer streaming​

Set stream: true on chat and responses requests whenever your tool lets you.

A buffered (non-streaming) request sends nothing back over the connection until the model has finished the entire reply. Any layer between the proxy and the operator — a CDN, a load balancer, a reverse proxy — can have its own idle timeout, and 60 seconds is a common default. If the reply takes longer than that, the layer gives up on the silent connection and answers 504, even though the operator was working normally.

Streaming sends tokens as they're produced, plus an invisible keepalive during quiet stretches such as a long reasoning pass. Keepalives prevent intermediary idle timeouts, and you see the reply as it arrives.

Buffered requests are fine for short replies. Streaming matters more as the expected output grows: big contexts, reasoning models, and agent turns that plan before answering.

Choosing where a request goes​

By default the proxy routes the same way the chat app does. For finer control over which operator handles a request — pinning to one, setting a price ceiling, requiring tool support, or trading latency for cost — use the provider field on the request body. See Proxy routing preferences.

Seeing what you paid​

For a non-streaming response, the cost is in the X-Zs-Inference-Amount and related X-Zs-* response headers. For a streaming response, set the header X-Zs-Usage-Frame: 1 on your request and the proxy adds an event: zs.usage frame to the stream with the same breakdown as JSON. The frame is opt-in because its payload is not a valid OpenAI chunk: strict OpenAI clients that parse every data: line fail on it. Set the header only from a client that handles the frame. See Pricing for what the numbers mean.

Troubleshooting​

context_length_exceeded on a conversation that isn't long​

A 400 like this one:

no operator serves <model> with input_tokens=36576 max_output_tokens=65536
operators advertising this model:
- op 1: max_output_tokens=32768 (over by 32768)

names the limit each candidate operator exceeded. Read the per-operator lines rather than the error code: despite its name, this error fires whenever no operator fits the request, and the two causes look different:

  • context_window=N (request needs M) — your conversation plus the reserved output doesn't fit the model's window by the proxy's size estimate, which runs about twice a real token count on prose, so this can fire well before the window is actually full. Start a new conversation, trim the history, or pick a model with a larger window.

  • max_output_tokens=N (over by M) — your tool asked for more output than that operator caps at. The proxy normally lowers this for you (see Response length), so if this is the only reason every operator was dropped, the lowering didn't apply. Set a smaller max response length in your tool, or check the two things that prevent the lowering:

    • an operator serving this model advertises no ceiling, so your value passes through unchanged; that operator's entry has no max_output_tokens= line;
    • you set a provider.max_price ceiling (Routing preferences) that rules out the operators with the larger ceilings.

    If neither applies, check zs-proxy version: releases before 0.14.3 don't lower max_tokens, so upgrade.

The same code without any per-operator lines comes from the operator after the request was sent: its backend refused the conversation as too long, usually because it serves a smaller window than the operator advertises. There is no inference charge. Retry, which may reach a different operator, or trim the history.

504 errors, usually after about a minute​

An error like this one:

504 operator returned HTTP 504 with a body this proxy could not parse as an
OpenAI error (content-type "text/html")

comes from a CDN or load balancer between the proxy and the operator, not from the operator's software, which is why the body is an HTML page rather than an API error. These arrive at a round number, typically around 60 seconds.

Almost always this is a buffered request that ran long. Set stream: true first — see Prefer streaming.

StageRetried by the proxy?
Setting up the requestYes. The proxy retries over different network paths and reroutes to other operators serving the same model, so most transient trouble never reaches you.
Paid for and sentNo. The request goes out exactly once, with no retry or failover, because the payment is already committed on-chain. A 504 here reaches your tool as-is.

Repeated attempts in your logs are therefore your own tool's retry logic. Each of those attempts is a separate paid request.

Every error the proxy returns carries an X-Should-Retry header, which the official OpenAI SDKs (Python, Node, Go) follow instead of retrying every 5xx. It says false for a request that was already paid for, and for an error a retry won't fix, so tools built on those SDKs don't resend those requests. A tool with its own retry logic may ignore the header and resend anyway.

The unused part of the escrow comes back at settlement, and a ticket nobody claims is refunded automatically after it expires. If the node had already generated output when the connection dropped, it can still bill for that work.

If you're already streaming and still see 504s from the same operator, that operator's front door is unhealthy. A retry normally lands on a different one; to rule it out for good, add its id to ignore in routing preferences.

What's next​

  • How-to guides — per-tool setup for every supported tool, plus the GUI chat apps (SillyTavern, Open WebUI, LibreChat, Chatbox, Cherry Studio).
  • Wallet & funding — keep the proxy funded.
  • Configuration — change the listen address, network, or spend caps.
  • Proxy CLI reference — every connect flag and the rest of the commands.