SillyTavern
SillyTavern is a popular front end for character-driven chat and roleplay. It speaks the OpenAI Chat Completions API, so it points at the proxy like any other custom endpoint. Every message goes over the same paid, encrypted network the chat app uses.
Before you start
The proxy must be running (zs-proxy proxy start) and funded
— see How-to guides.
SillyTavern runs natively on your machine, so the default base URL
http://localhost:9376/v1 works as-is. (If you run SillyTavern in Docker, see
the Docker note.)
Connect it
Open the API connections panel
In SillyTavern's top bar, click the API Connections icon (the plug).
Choose Chat Completion → Custom
- Set API to Chat Completion.
- Set Chat Completion Source to Custom (OpenAI-compatible).
Fill in the endpoint
- Custom Endpoint (Base URL):
http://localhost:9376/v1 - Custom API Key: anything non-empty (e.g.
zerosignal) — the proxy ignores it.
Leave off any /chat/completions suffix; SillyTavern adds it.
Connect and pick a model
Click Connect. SillyTavern reads the proxy's /v1/models and fills the
model dropdown with the live catalog — pick whichever model you want to
chat with. (If the dropdown stays empty, type a model id by hand; see
Picking a model.)
Send a test message
Use Test Message to fire a quick prompt, or just start chatting with a character. A reply confirms the setup works.
Notes for roleplay
- Streaming works normally — leave it on for token-by-token replies.
- SillyTavern's samplers (temperature, top-p, penalties) pass straight through to the operator.
- Response length. If you don't set a max response length, the proxy picks a
sensible output ceiling for the operator you're routed to (up to
fallback_max_output_tokens, 32768 by default) so it can price the request — plenty for chat. You can set a smaller "Max Response Length" in SillyTavern to keep replies (and cost) tight. - Context size is up to the model you pick — larger contexts cost more per turn because more tokens are sent. See Pricing.
Troubleshooting
| Symptom | Fix |
|---|---|
| "API returned an error" / can't connect | Confirm the proxy is up: curl http://localhost:9376/healthz. Start it with zs-proxy proxy start. |
| Red status light, but chats work | Tick Bypass API status check — SillyTavern's probe is stricter than the endpoint needs. |
| Empty model dropdown | The proxy wasn't running when you clicked Connect, or /v1/models returned nothing. Start the proxy, then reconnect — or type a model id from curl http://localhost:9376/v1/models. |
wallet_unfunded / payment errors | Add funds in the chat app, or with zs-proxy fund for a proxy-only wallet — see Wallet & funding. |
| Works from Docker? Nothing connects | See the Docker note — use host.docker.internal and bind the proxy off loopback. |
What's next
- Pricing — what a chat turn costs.
- Configuration — spend caps and the output ceiling.
- How-to guides — connect another app.