Skip to main content

Per-call tool pricing

Tokens aren't the only thing a request can cost you. Two kinds of tool call run up a bill beyond the prompt and completion:

  • Vendor server-side tools. Frontier models like xAI Grok, OpenAI GPT, and Kimi can run their own hosted tools inside a single completion: web search, X search, code execution, file search. The vendor executes them and bills you, the operator, per call (xAI charges around $5 per 1,000 web searches, for example), on top of the tokens.
  • Your node's own zs_ built-ins. zs_web_search, zs_web_read, and zs_image_search cost you real compute and bandwidth to run; zs_web_read in particular can spike memory converting a large page.

Per-call tool pricing lets you charge for both. You set a price per tool call; the node reserves for it, caps it, and bills the payer, so a harness that wants Grok's web search keeps working and you get paid for it.

info

This is a pricing layer on top of the token rates in Serving models & pricing. Your zs_ built-in tools themselves are set up in Built-in tools; this page is only about charging for tool calls.

Two kinds of tool call​

Vendor server-side toolNode zs_ built-in
Who executes itThe upstream vendor, inside one completionYour node
Examplesweb_search, x_search, code_interpreter, file_searchzs_web_search, zs_web_read, zs_image_search
Who bills youThe vendor (per call)Nobody directly — it's your compute
Default when unpricedStripped from the requestFree (offered at no charge)
Cap it stays underThe vendor's own cap knob + your call_capThe tool-loop iteration limit

A third group costs nothing per call and is left alone; see Tools that are never billed.

zs_image_generation / zs_image_edit are not priced here; they already bill per produced image (see Built-in tools → Image tools).

Default-deny: you're never billed for a tool you didn't price​

By default:

  • An unpriced vendor server-side tool in a request is stripped before the node forwards it upstream, so the vendor can't run it and can't bill you.
  • Your zs_ tools stay free unless you name a price.
warning

A client's web_search is stripped before it reaches your upstream unless you price the tool (below) or opt out with vendor.enabled: false, which passes vendor tools through unpriced, at your own cost. If you rely on server-side tools reaching the upstream, price them.

Default-deny does not apply to the never-billed types; those keep reaching your upstream while they stay unpriced. Giving mcp an explicit rate promotes it to the billable tier, which also subjects it to chat_server_tools on chat.

Tools that are never billed​

These tool types carry no per-call fee at any vendor. They cost tokens, like any other part of the prompt, so the node passes them through untouched and never meters them:

Tool typeWhat it isWhy it's free
mcpA remote MCP connector the vendor calls out toToken cost only (OpenAI, Anthropic)
customOpenAI freeform toolsYour caller executes it
local_shellShell commands run on the caller's machineYour caller executes it
apply_patchStructured file edits the caller appliesYour caller executes it
computer_useLets a computer-use model drive a screen: it emits click / type / scroll actions and reads back screenshotsYour caller executes it — and the model refuses to run unless the tool is declared
shell with environment.type: "local"The successor to the deprecated local_shell — the command runs on the caller's machineYour caller executes it; no container is allocated
bash, text_editor, str_replace_editor, str_replace_based_edit_toolAnthropic's caller-executed tools (the _YYYYMMDD version suffix is normalized away)Your caller executes them

Stripping these would break a caller's harness and save you nothing, which is why they're exempt from default-deny. Everything else with an unrecognized type is still treated as billable and stripped unless priced.

Read the shell row carefully: shell is a different tool from local_shell. In its container_auto / container_reference modes OpenAI allocates a managed container and bills you per session, so those modes stay default-denied; only environment.type: "local" is exempt. OpenAI has deprecated local_shell in favour of shell + environment.type: "local", and without the exemption a Codex harness that has migrated would have its shell tool silently stripped.

If your upstream does charge per call for a vendor-hosted one (in practice, mcp), name it in tools with a rate; an explicit rate wins over the never-billed default. The default_usd_per_1k_calls catch-all does not reach them, so a catch-all meant to cover Grok's search can't start charging your users for an MCP connector that cost you nothing.

The caller-executed ones can't be priced at all; the node rejects the config at startup. Your caller runs those tools, so no vendor ever reports a call for the node to meter, and a rate would inflate every reserve for the model against a charge that can never land.

info

A surviving hosted tool still counts against call_cap, even at a rate of zero, because every hosted call re-feeds its results into the context and that token growth needs a bound. Caller-executed tools don't count against it.

Pricing tools (zs.tool_pricing)​

One block prices both kinds of tool. Rates are USD per 1,000 calls, the unit vendors quote, so you can copy a number off a price card and add your markup.

zs:
tool_pricing:
tools: # named per-call rates
web_search: { usd_per_1k_calls: 6.00 } # your vendor cost + markup
x_search: { usd_per_1k_calls: 6.00 }
code_interpreter:{ usd_per_1k_calls: 6.00 }
zs_web_search: { usd_per_1k_calls: 4.00 } # also prices YOUR own tools
zs_web_read: { usd_per_1k_calls: 2.00 }
default_usd_per_1k_calls: 6.00 # catch-all for any UNNAMED vendor tool
vendor:
dialect: auto # auto | openai | xai | zai | moonshot | openrouter | generic
call_cap: 16 # per-request cap on vendor tool calls
tokens_per_call: 3000 # est. context inflation per vendor call
chat_server_tools: strip # strip (default) | allow — see below

Named rates and the vendor catch-all​

Every entry in tools is keyed by the tool's canonical name. A key beginning zs_ prices one of your built-ins; any other key prices a vendor tool. Each entry must set usd_per_1k_calls: an explicit 0 means free, and a missing rate is a config error.

default_usd_per_1k_calls is the catch-all: any vendor tool a client sends that you didn't name is billed at this rate instead of being stripped, so a harness can use a vendor tool you didn't think to enumerate. The catch-all never applies to your zs_ tools, so setting it to cover Grok's search can't start charging users for your own search. An unnamed zs_ tool is free.

default_usd_per_1k_calls: 0 — pass everything through, bill nothing​

Setting the catch-all to zero says "my upstream doesn't charge me per call." It is not the same as leaving the block out: a zero catch-all makes every vendor server-side tool survive and bill nothing instead of being stripped.

zs:
tool_pricing:
default_usd_per_1k_calls: 0 # nothing here bills per call — pass it all through
vendor:
call_cap: 64 # a catch-all still caps calls; raise it if that bites

Use this when you serve an open-weight or self-hosted upstream (vLLM, SGLang, llama.cpp, LM Studio) that has no hosted tools to bill for, or a provider that folds tool cost into its token rates. The same works per tool: web_search: { usd_per_1k_calls: 0 } offers exactly that one for free.

warning

A zero rate removes the per-call fee, not the token cost. Server-side tools still re-feed their results into the context, and those input tokens bill at your normal rate. Size your token rates knowing the context can grow.

The rules in one line each:

  • Never-billed tool → passed through, never billed, unless you name a rate for a hosted one such as mcp (caller-executed types reject a rate at startup).
  • Vendor tool → named rate, else the catch-all, else stripped.
  • zs_ tool → named rate, else free.
  • No tool_pricing block at all → billable vendor tools stripped, zs_ tools free.

Vendor knobs​

Under vendor:

  • dialect — how the node reads your upstream's tool conventions. auto infers it from the host of llm.openai.base_url: x.ai → xai, openai.com → openai, moonshot.ai/moonshot.cn/kimi.ai/kimi.com → moonshot, z.ai/bigmodel.cn → zai, openrouter.ai → openrouter, each including its subdomains. A proxy or gateway in front of the vendor is on a different host, so it needs dialect set explicitly. An unrecognized host falls back to a generic reader. Set it explicitly if auto-detection guesses wrong.
  • call_cap — the most vendor server-side calls you'll reserve (and pay) for in one request. On the Responses API the node injects this as the vendor's own cap knob (max_tool_calls / max_turns), so the model can't exceed it. Defaults to 16 once you've priced anything (a named rate or a catch-all). With no tool_pricing block at all there is no cap, so a never-billed hosted tool like mcp passes through unbounded. A default cap would truncate a legitimate remote-MCP session, and the payer pays up to max_price either way. Set vendor.call_cap if you want the bound anyway.
  • tokens_per_call — your estimate of how many extra input tokens each vendor call adds. Server-side tools re-feed their results into the context, so a search-heavy request can bill several times its original prompt in input tokens; this term reserves for that.
  • chat_server_tools — strip (default) or allow: what to do with a priced vendor tool on /v1/chat/completions, where its call count can't be capped upstream. See Chat has no cap knob.
  • enabled: false — opt out of vendor handling entirely: pass server-side tools through unpriced, at your own cost. For an operator with a separate billing arrangement with the upstream. It doesn't affect zs_ pricing.

Per-model overrides​

zs.tool_pricing is the fleet-wide default. Override it for one model under zs.models[<id>].tool_pricing, for example when only some models point at a vendor that charges for tools:

zs:
models:
grok-4.5:
tool_pricing:
vendor: { dialect: xai, call_cap: 12 }

Per-model entries and knobs are merged over the default, entry by entry.

Examples by vendor​

The rates in the comments are the vendor's list price (your cost) at the time of writing. Set usd_per_1k_calls to that plus your margin, and always check the current price card. These are for OpenAI-compatible upstreams; Gemini and Anthropic are covered under the caveats below.

Grok's server-side tools are Responses-API tools, so call_cap is enforced. Point llm.openai.base_url at https://api.x.ai/v1.

zs:
tool_pricing:
tools:
web_search: { usd_per_1k_calls: 6.00 } # xAI $5/1k + markup
x_search: { usd_per_1k_calls: 6.00 } # xAI $5/1k
code_interpreter: { usd_per_1k_calls: 6.00 } # xAI $5/1k (code execution)
file_search: { usd_per_1k_calls: 3.00 } # xAI $2.50/1k (Collections Search)
attachment_search: { usd_per_1k_calls: 12.00 } # xAI $10/1k (File Attachments)
vendor:
dialect: xai
call_cap: 16
tokens_per_call: 3000

Watch the names. file_search is xAI's alias for Collections Search at $2.50/1k; the $10/1k tool is the separate attachment_search. Don't assume a tool name means what it means at OpenAI. xAI's primary names are code_execution, collections_search and attachment_search; the aliases above are accepted on the REST API but reportedly not on the gRPC SDK. document_search has no published price, so leave it unpriced (stripped) rather than guess.

Gemini and Anthropic don't fit yet​

  • Google Gemini grounding (Google Search) is around $14/1k on the Gemini 3.x family (after a monthly free tier) and $35/1k on 2.5, but the unit differs: 3.x bills per search query the model runs, while 2.5 and older bill per grounded prompt, however many queries fan out. So $14/1k is not a discount; an agentic prompt averaging 3+ searches costs more on 3.x. It doesn't fit tool_pricing either: Gemini's grounding tool uses a non-OpenAI shape ({"google_search": {}}, with no type field), so the node doesn't classify it as a priced tool, and on the vertexai provider non-function tools are dropped in translation anyway. Fold grounding into the model's token rates, or run log_raw_usage to see how your Gemini upstream exposes and counts it before trying a tool_pricing entry.
  • Anthropic Claude server-side tools (web_search, code_execution) run only on the native Messages API, which the node doesn't serve yet, so there's nothing to price.

How it's billed​

A tool fee rides on top of the token charge, on the same receipt, and is capped at the ticket's escrowed max_price. The node counts the calls that actually ran (a vendor's usage report for server-side tools; the node's own count for zs_ tools) and bills Σ calls × your rate. Only successful calls are billed. A request that fails before completing bills $0, even if a tool already ran.

What it does to the reserve​

To cover the fees, the node folds a worst-case tool-fee headroom into max_price for any model that prices tools:

max_tool_iterations × (priciest zs_ rate)
+ call_cap × (priciest vendor rate)
+ the input-token cost of call_cap × tokens_per_call

It's per-model, not per-request. Every reserve against a tool-pricing model locks this headroom up front, even a plain chat turn that uses no tools, and the unused part is refunded at settlement. Size call_cap and tokens_per_call to a realistic worst case. Larger values make plain requests lock more of a user's USDC than they need to.

Chat has no cap knob. call_cap is enforced upstream only on the Responses API (/v1/responses), where max_tool_calls / max_turns live. On /v1/chat/completions the count can't be capped, so vendor.chat_server_tools decides what happens to priced vendor tools there:

chat_server_toolsPriced vendor tools on chat
strip (default)Removed, even when priced. Offer them on the Responses API instead, where call_cap is enforced.
allowKept, uncapped. A request may run more than call_cap calls, and you eat any overage beyond the reserve.

A tool priced at exactly 0 survives strip: the rule exists to stop an uncapped per-call overage, and at a rate of zero there isn't one.

strip also removes the top-level chat search fields that aren't tool entries: xAI's legacy search_parameters and OpenAI's web_search_options. Those two follow the web_search rate specifically; pricing some other tool at zero won't retain them. One caveat strip can't fix: a dedicated search model (OpenAI's gpt-4o-search-preview, Perplexity sonar, ...) searches on every request as part of the model itself, not as a strippable tool. Removing web_search_options only stops a client dialing the search up, not the model's baseline search. Price those models to cover their search cost, or don't serve them.

Why did my tool disappear?​

When a user reports that a tool they sent didn't run, check the node log for:

tool pricing stripped a vendor server-side tool from a request

The line names the canonical tool and the model, and is emitted once per tool name for the life of the process, so it tells you what to price without flooding a busy node. It logs tool names only, never arguments or request content.

To make it stop, in rough order of preference:

  1. Price the tool — tools: { web_search: { usd_per_1k_calls: 6.00 } }.
  2. Pass it through free if your upstream doesn't charge — usd_per_1k_calls: 0, or default_usd_per_1k_calls: 0 for all of them.
  3. Opt out of vendor handling entirely — vendor.enabled: false (unpriced, at your own cost).

If the tool is on chat rather than the Responses API, also check chat_server_tools: strict chat strips positively-priced vendor tools there.

Confirm your upstream's counts before going live​

The node bills vendor tools from the call counts your upstream reports in its usage object (or, for OpenAI-style responses, from the tool-call items in the response). Providers report these differently. Turn on llm.openai.log_raw_usage: true for a capture run and watch the raw upstream usage log line to confirm the shape before you rely on it. The log is counts-only, never prompt or response content.

See Examples by vendor above for the current per-vendor list prices and the tool names to key your rates by.

See also​