Per-call tool pricing
Tokens aren't the only thing a request can cost you. Two kinds of tool call run up a bill beyond the prompt and completion:
- Vendor server-side tools. Frontier models like xAI Grok, OpenAI GPT, and Kimi can run their own hosted tools inside a single completion: web search, X search, code execution, file search. The vendor executes them and bills you, the operator, per call (xAI charges around $5 per 1,000 web searches, for example), on top of the tokens.
- Your node's own
zs_built-ins.zs_web_search,zs_web_read, andzs_image_searchcost you real compute and bandwidth to run;zs_web_readin particular can spike memory converting a large page.
Per-call tool pricing lets you charge for both. You set a price per tool call; the node reserves for it, caps it, and bills the payer, so a harness that wants Grok's web search keeps working and you get paid for it.
This is a pricing layer on top of the token rates in
Serving models & pricing. Your zs_ built-in tools
themselves are set up in Built-in tools; this page is only
about charging for tool calls.
Two kinds of tool call
| Vendor server-side tool | Node zs_ built-in | |
|---|---|---|
| Who executes it | The upstream vendor, inside one completion | Your node |
| Examples | web_search, x_search, code_interpreter, file_search | zs_web_search, zs_web_read, zs_image_search |
| Who bills you | The vendor (per call) | Nobody directly — it's your compute |
| Default when unpriced | Stripped from the request | Free (offered at no charge) |
| Cap it stays under | The vendor's own cap knob + your call_cap | The tool-loop iteration limit |
A third group costs nothing per call and is left alone; see Tools that are never billed.
zs_image_generation / zs_image_edit are not priced here; they already
bill per produced image (see Built-in tools → Image tools).
Default-deny: you're never billed for a tool you didn't price
By default:
- An unpriced vendor server-side tool in a request is stripped before the node forwards it upstream, so the vendor can't run it and can't bill you.
- Your
zs_tools stay free unless you name a price.
A client's web_search is stripped before it reaches your upstream unless
you price the tool (below) or opt out with vendor.enabled: false, which passes
vendor tools through unpriced, at your own cost. If you rely on server-side
tools reaching the upstream, price them.
Default-deny does not apply to the
never-billed types; those keep reaching your
upstream while they stay unpriced. Giving mcp an explicit rate promotes it to
the billable tier, which also subjects it to chat_server_tools on chat.
Tools that are never billed
These tool types carry no per-call fee at any vendor. They cost tokens, like any other part of the prompt, so the node passes them through untouched and never meters them:
| Tool type | What it is | Why it's free |
|---|---|---|
mcp | A remote MCP connector the vendor calls out to | Token cost only (OpenAI, Anthropic) |
custom | OpenAI freeform tools | Your caller executes it |
local_shell | Shell commands run on the caller's machine | Your caller executes it |
apply_patch | Structured file edits the caller applies | Your caller executes it |
computer_use | Lets a computer-use model drive a screen: it emits click / type / scroll actions and reads back screenshots | Your caller executes it — and the model refuses to run unless the tool is declared |
shell with environment.type: "local" | The successor to the deprecated local_shell — the command runs on the caller's machine | Your caller executes it; no container is allocated |
bash, text_editor, str_replace_editor, str_replace_based_edit_tool | Anthropic's caller-executed tools (the _YYYYMMDD version suffix is normalized away) | Your caller executes them |
Stripping these would break a caller's harness and save you nothing, which is why
they're exempt from default-deny. Everything else with an unrecognized type is
still treated as billable and stripped unless priced.
Read the shell row carefully: shell is a different tool from
local_shell. In its container_auto / container_reference modes OpenAI
allocates a managed container and bills you per session, so those modes stay
default-denied; only environment.type: "local" is exempt. OpenAI has
deprecated local_shell in favour of shell + environment.type: "local",
and without the exemption a Codex harness that has migrated would have its shell
tool silently stripped.
If your upstream does charge per call for a vendor-hosted one (in practice,
mcp), name it in tools with a rate; an explicit rate wins over the
never-billed default. The default_usd_per_1k_calls catch-all does not
reach them, so a catch-all meant to cover Grok's search can't start
charging your users for an MCP connector that cost you nothing.
The caller-executed ones can't be priced at all; the node rejects the config at startup. Your caller runs those tools, so no vendor ever reports a call for the node to meter, and a rate would inflate every reserve for the model against a charge that can never land.
A surviving hosted tool still counts against call_cap, even at a rate of
zero, because every hosted call re-feeds its results into the context and that
token growth needs a bound. Caller-executed tools don't count against it.
Pricing tools (zs.tool_pricing)
One block prices both kinds of tool. Rates are USD per 1,000 calls, the unit vendors quote, so you can copy a number off a price card and add your markup.
zs:
tool_pricing:
tools: # named per-call rates
web_search: { usd_per_1k_calls: 6.00 } # your vendor cost + markup
x_search: { usd_per_1k_calls: 6.00 }
code_interpreter:{ usd_per_1k_calls: 6.00 }
zs_web_search: { usd_per_1k_calls: 4.00 } # also prices YOUR own tools
zs_web_read: { usd_per_1k_calls: 2.00 }
default_usd_per_1k_calls: 6.00 # catch-all for any UNNAMED vendor tool
vendor:
dialect: auto # auto | openai | xai | zai | moonshot | openrouter | generic
call_cap: 16 # per-request cap on vendor tool calls
tokens_per_call: 3000 # est. context inflation per vendor call
chat_server_tools: strip # strip (default) | allow — see below
Named rates and the vendor catch-all
Every entry in tools is keyed by the tool's canonical name. A key beginning
zs_ prices one of your built-ins; any other key prices a vendor tool. Each entry
must set usd_per_1k_calls: an explicit 0 means free, and a missing rate is a
config error.
default_usd_per_1k_calls is the catch-all: any vendor tool a client
sends that you didn't name is billed at this rate instead of being stripped, so a
harness can use a vendor tool you didn't think to enumerate. The catch-all
never applies to your zs_ tools, so setting it to cover Grok's search can't
start charging users for your own search. An unnamed zs_ tool is free.
default_usd_per_1k_calls: 0 — pass everything through, bill nothing
Setting the catch-all to zero says "my upstream doesn't charge me per call." It is not the same as leaving the block out: a zero catch-all makes every vendor server-side tool survive and bill nothing instead of being stripped.
zs:
tool_pricing:
default_usd_per_1k_calls: 0 # nothing here bills per call — pass it all through
vendor:
call_cap: 64 # a catch-all still caps calls; raise it if that bites
Use this when you serve an open-weight or self-hosted upstream (vLLM,
SGLang, llama.cpp, LM Studio) that has no hosted tools to bill for, or a
provider that folds tool cost into its token rates. The same works per tool:
web_search: { usd_per_1k_calls: 0 } offers exactly that one for free.
A zero rate removes the per-call fee, not the token cost. Server-side tools still re-feed their results into the context, and those input tokens bill at your normal rate. Size your token rates knowing the context can grow.
The rules in one line each:
- Never-billed tool → passed through,
never billed, unless you name a rate for a hosted one such as
mcp(caller-executed types reject a rate at startup). - Vendor tool → named rate, else the catch-all, else stripped.
zs_tool → named rate, else free.- No
tool_pricingblock at all → billable vendor tools stripped,zs_tools free.
Vendor knobs
Under vendor:
dialect— how the node reads your upstream's tool conventions.autoinfers it from the host ofllm.openai.base_url:x.ai→ xai,openai.com→ openai,moonshot.ai/moonshot.cn/kimi.ai/kimi.com→ moonshot,z.ai/bigmodel.cn→ zai,openrouter.ai→ openrouter, each including its subdomains. A proxy or gateway in front of the vendor is on a different host, so it needsdialectset explicitly. An unrecognized host falls back to a generic reader. Set it explicitly if auto-detection guesses wrong.call_cap— the most vendor server-side calls you'll reserve (and pay) for in one request. On the Responses API the node injects this as the vendor's own cap knob (max_tool_calls/max_turns), so the model can't exceed it. Defaults to 16 once you've priced anything (a named rate or a catch-all). With notool_pricingblock at all there is no cap, so a never-billed hosted tool likemcppasses through unbounded. A default cap would truncate a legitimate remote-MCP session, and the payer pays up tomax_priceeither way. Setvendor.call_capif you want the bound anyway.tokens_per_call— your estimate of how many extra input tokens each vendor call adds. Server-side tools re-feed their results into the context, so a search-heavy request can bill several times its original prompt in input tokens; this term reserves for that.chat_server_tools—strip(default) orallow: what to do with a priced vendor tool on/v1/chat/completions, where its call count can't be capped upstream. See Chat has no cap knob.enabled: false— opt out of vendor handling entirely: pass server-side tools through unpriced, at your own cost. For an operator with a separate billing arrangement with the upstream. It doesn't affectzs_pricing.
Per-model overrides
zs.tool_pricing is the fleet-wide default. Override it for one model under
zs.models[<id>].tool_pricing, for example when only some models point at a
vendor that charges for tools:
zs:
models:
grok-4.5:
tool_pricing:
vendor: { dialect: xai, call_cap: 12 }
Per-model entries and knobs are merged over the default, entry by entry.
Examples by vendor
The rates in the comments are the vendor's list price (your cost) at the
time of writing. Set usd_per_1k_calls to that plus your margin, and always
check the current price card. These are for OpenAI-compatible upstreams;
Gemini and Anthropic are covered under the caveats below.
- xAI Grok
- OpenAI GPT
- Kimi (Moonshot)
Grok's server-side tools are Responses-API tools, so call_cap is enforced.
Point llm.openai.base_url at https://api.x.ai/v1.
zs:
tool_pricing:
tools:
web_search: { usd_per_1k_calls: 6.00 } # xAI $5/1k + markup
x_search: { usd_per_1k_calls: 6.00 } # xAI $5/1k
code_interpreter: { usd_per_1k_calls: 6.00 } # xAI $5/1k (code execution)
file_search: { usd_per_1k_calls: 3.00 } # xAI $2.50/1k (Collections Search)
attachment_search: { usd_per_1k_calls: 12.00 } # xAI $10/1k (File Attachments)
vendor:
dialect: xai
call_cap: 16
tokens_per_call: 3000
Watch the names. file_search is xAI's alias for Collections Search at
$2.50/1k; the $10/1k tool is the separate attachment_search. Don't assume a
tool name means what it means at OpenAI. xAI's primary names are code_execution,
collections_search and attachment_search; the aliases above are accepted on
the REST API but reportedly not on the gRPC SDK. document_search has no
published price, so leave it unpriced (stripped) rather than guess.
GPT's built-in tools are Responses-API tools (capped by call_cap). Point
base_url at https://api.openai.com/v1. Web search may also bill
search-content input tokens, which ride your normal token rate, so size
tokens_per_call to cover them. OpenAI's flat "8k tokens per call" figure
applies only to gpt-4o-mini and gpt-4.1-mini; other models bill content
tokens as actually retrieved, and web_search_preview on non-reasoning models
bills none. Measure yours before committing to a number.
zs:
tool_pricing:
tools:
web_search: { usd_per_1k_calls: 12.00 } # OpenAI $10/1k + markup (but see below)
file_search: { usd_per_1k_calls: 3.00 } # OpenAI $2.50/1k queries
vendor:
dialect: openai
call_cap: 12
tokens_per_call: 9000 # measure yours; see the note below
web_search_preview on a non-reasoning model is $25/1k, not $10/1k. It
normalizes to the same web_search billing key, so an operator pricing at 12.00
is underwater on that path. Price for the highest tier you actually serve.
code_interpreter is billed per 20-minute container session (memory-tiered:
$0.03 / $0.12 / $0.48 / $1.92 for 1 / 4 / 16 / 64 GB, minute-billed with a
5-minute minimum), not per call, so it doesn't map cleanly to a per-call rate.
Leave it unpriced (stripped) unless you have a reason to offer it. The hosted
shell tool shares that container pool and price line, so don't price it per
call either. (shell with environment.type: "local" runs on the caller's
machine, costs you nothing, and is passed through automatically.) On
/v1/chat/completions, GPT web search is the dedicated *-search-preview
models via web_search_options (stripped by default under strict chat). Those
search inherently, so price the model instead.
Kimi's published price is $0.005 per successful call = $5/1k (CN platform ¥0.03/call). Moonshot currently flags its web-search docs as outdated and advises against relying on the feature near-term; check before you enable it.
Kimi's $web_search runs on chat completions, so it's stripped by default.
Set chat_server_tools: allow to offer it (uncapped; size call_cap for a
realistic worst case). Key the rate as web_search; the $ is normalized away.
zs:
tool_pricing:
tools:
web_search: { usd_per_1k_calls: 6.00 } # Kimi $5/1k ($0.005/successful call)
vendor:
dialect: moonshot
call_cap: 8
chat_server_tools: allow
Gemini and Anthropic don't fit yet
- Google Gemini grounding (Google Search) is around $14/1k on the Gemini
3.x family (after a monthly free tier) and $35/1k on 2.5, but the unit
differs: 3.x bills per search query the model runs, while 2.5 and older
bill per grounded prompt, however many queries fan out. So $14/1k is not a
discount; an agentic prompt averaging 3+ searches costs more on 3.x. It doesn't
fit
tool_pricingeither: Gemini's grounding tool uses a non-OpenAI shape ({"google_search": {}}, with notypefield), so the node doesn't classify it as a priced tool, and on thevertexaiprovider non-function tools are dropped in translation anyway. Fold grounding into the model's token rates, or runlog_raw_usageto see how your Gemini upstream exposes and counts it before trying atool_pricingentry. - Anthropic Claude server-side tools (
web_search,code_execution) run only on the native Messages API, which the node doesn't serve yet, so there's nothing to price.
How it's billed
A tool fee rides on top of the token charge, on the same receipt, and is
capped at the ticket's escrowed max_price. The node counts the calls that
actually ran (a vendor's usage report for server-side tools; the node's own count
for zs_ tools) and bills Σ calls × your rate. Only successful calls are
billed. A request that fails before completing bills $0, even if a tool
already ran.
What it does to the reserve
To cover the fees, the node folds a worst-case tool-fee headroom into
max_price for any model that prices tools:
max_tool_iterations × (priciest zs_ rate)
+ call_cap × (priciest vendor rate)
+ the input-token cost of call_cap × tokens_per_call
It's per-model, not per-request. Every reserve against a tool-pricing model
locks this headroom up front, even a plain chat turn that uses no tools, and the
unused part is refunded at settlement. Size call_cap and tokens_per_call to a
realistic worst case. Larger values make plain requests lock more of a user's
USDC than they need to.
Chat has no cap knob. call_cap is enforced upstream only on the Responses
API (/v1/responses), where max_tool_calls / max_turns live. On
/v1/chat/completions the count can't be capped, so vendor.chat_server_tools
decides what happens to priced vendor tools there:
chat_server_tools | Priced vendor tools on chat |
|---|---|
strip (default) | Removed, even when priced. Offer them on the Responses API instead, where call_cap is enforced. |
allow | Kept, uncapped. A request may run more than call_cap calls, and you eat any overage beyond the reserve. |
A tool priced at exactly 0 survives strip: the rule exists to stop an
uncapped per-call overage, and at a rate of zero there isn't one.
strip also removes the top-level chat search fields that aren't tool entries:
xAI's legacy search_parameters and OpenAI's web_search_options. Those two
follow the web_search rate specifically; pricing some other tool at zero
won't retain them. One caveat strip can't fix: a dedicated search model
(OpenAI's gpt-4o-search-preview, Perplexity sonar, ...) searches on every
request as part of the model itself, not as a strippable tool. Removing
web_search_options only stops a client dialing the search up, not the
model's baseline search. Price those models to cover their search cost, or don't
serve them.
Why did my tool disappear?
When a user reports that a tool they sent didn't run, check the node log for:
tool pricing stripped a vendor server-side tool from a request
The line names the canonical tool and the model, and is emitted once per tool name for the life of the process, so it tells you what to price without flooding a busy node. It logs tool names only, never arguments or request content.
To make it stop, in rough order of preference:
- Price the tool —
tools: { web_search: { usd_per_1k_calls: 6.00 } }. - Pass it through free if your upstream doesn't charge —
usd_per_1k_calls: 0, ordefault_usd_per_1k_calls: 0for all of them. - Opt out of vendor handling entirely —
vendor.enabled: false(unpriced, at your own cost).
If the tool is on chat rather than the Responses API, also check
chat_server_tools: strict chat strips positively-priced vendor tools there.
Confirm your upstream's counts before going live
The node bills vendor tools from the call counts your upstream reports in its
usage object (or, for OpenAI-style responses, from the tool-call items in the
response). Providers report these differently. Turn on
llm.openai.log_raw_usage: true for a capture run and watch the
raw upstream usage log line to confirm the shape before you rely on it. The
log is counts-only, never prompt or response content.
See Examples by vendor above for the current per-vendor list prices and the tool names to key your rates by.
See also
- Serving models & pricing — the token rates a tool fee sits on top of.
- Built-in tools — enabling and tuning your
zs_tools. - The payment flow — how a reserve, receipt, and settlement fit together.