Skip to main content

Built-in tools

Your node can run a small set of built-in tools for the models it serves: web search, web read, and, when you wire an image backend, image generation and editing. They're in-loop tools: when a model calls one mid-turn, your node executes it locally, feeds the result back to the model, and continues the same turn. Every round of a turn settles into one receipt.

Your node advertises the tools it offers on /v1/zs/details by type only. Names, descriptions, and parameters are identical fleet-wide, so clients fill them in. A model only gets tools if it declares tool support and zs.builtin_tools.enabled is on; a relay-only node runs none.

info

Built-in tools are free by default. To charge per search or read, or for a frontier model's own server-side tools, see Per-call tool pricing.

The tools​

ToolWhat it doesRuns against
Web searchQuery the web, return numbered results the model cites.DuckDuckGo HTML — no API key.
Web readFetch a URL and hand the model clean markdown.The page itself, fetched by your node — no third party, no API key.
Get timeReport the node's clock, optionally in a given IANA timezone.The node itself — no network.
Image generationCreate an image from a prompt, inline in the reply.Your image_llm backend.
Image editingModify an image — the model's or the conversation's — from a prompt.Your image_llm backend.
info

A web image search tool exists in the code but is force-disabled: its DuckDuckGo image backend is broken, so the node won't advertise or run it, regardless of config, until a working provider is wired up.

Where the requests come from​

Web search and web read make outbound requests from your node's network: a client's "search the web" becomes an HTTP request from your host. Tool activity is never logged as content. Prompts, fetched page bodies, tool arguments, and image bytes stay out of logs, traces, and metrics; only counts, durations, and IDs are recorded. Errors surfaced back to the model are sanitized so upstream URLs don't leak. For the user-facing view, see Privacy & security; for what you bill, see The payment flow.

Web read fetches the page itself. Your node makes the request, converts the HTML to markdown in-process, and hands the result to the model, so no intermediary service learns what your users asked to read. The site and your DNS resolver still see the request.

  • The site sees your node's IP. Your node identifies itself as zs-node/web_read, so a site that wants to exclude it can.
  • Your node will not fetch its own network. Web read refuses any URL that resolves to a private, loopback, link-local, or carrier-NAT address. The check runs against the address actually being dialled, so a public hostname pointing inward (or rebinding mid-request) is refused too. This stops a crafted URL from turning a read into a probe of your LAN or your cloud provider's metadata endpoint. It is on by default; leave it on. web_read.allow_private_targets exists for localnet development only.

Web read is also gated on provenance: the model may only read a URL the user supplied or one a prior search returned, never one it invented from a page it just read. This stops a hostile page from talking the model into encoding your user's conversation into a URL and fetching it.

info

Web search still uses a third party. Search queries go to DuckDuckGo, which sees the query. web_search.enabled: false turns search off and leaves web read working.

A confidential node can serve these too. Running confidential compute does not mean turning the web tools off. A tool runs only because the caller listed it in that request's tools[], so the user decides whether a query or a URL leaves the enclave. A user who leaves web search off gets the confidential-posture claim from the encrypted request they sent, not from anything you advertise. Your node logs one warning at startup naming the web tools it offers. The image backend has its own constraints; see Confidential compute.

Configuration​

The web tools and the loop are configured under zs.builtin_tools. See Configuration → zs.builtin_tools for the full knob table (per-tool enabled, max_results / max_bytes / max_download, timeout, safe_search, and the loop guards).

Two ceilings on a read:

KnobDefaultCaps
web_read.max_bytes64 KiBThe markdown the model receives. This is the context-window guard.
web_read.max_download4 MiBThe raw page your node downloads before converting it. A page can be far larger than the markdown it distills to.

A page over the download ceiling is refused, not truncated, because half an HTML document converts to nonsense. Setting max_bytes above max_download is rejected at startup, since no page could then fill the markdown budget.

warning

max_download is not a memory budget. Converting a page costs roughly 20–70× the page's own size in peak memory: the readability pass builds and scores a document tree of the whole page, and the multiple depends on the page's structure. A 4 MiB page can mean a few hundred MB while it converts.

What bounds the total is web_read.max_concurrent (default 4): peak memory is roughly max_download × 20–70 × max_concurrent. Lower it when the node has a memory limit. The cost is latency, because reads queue. Lowering max_download instead would make large pages permanently unreadable.

Web read is the only path that makes the node's memory use spike, so it determines the limit to set if you run the node under a container or cgroup limit. See Sizing the node process.

Content filtering (web_search.safe_search). SafeSearch defaults to moderate; strict and off are the other choices. off returns unfiltered results, so set it deliberately. An unrecognized value is a hard startup error, not a silent fallback. image_search has the same knob, but that tool is force-disabled, so it has no effect today.

Loop guardWhat it does
max_iterationsThe absolute ceiling on chat→tool→chat rounds in a single turn. Advertised to clients as max_tool_iterations.
max_stalled_iterationsAn internal safety heuristic that stops a model spinning on the same repeated call. Not advertised.

Offering tools raises what callers reserve. A tool result is re-fed to the model, so a tool-carrying request needs more input budget than its prompt alone. Clients size that from two values you advertise: max_iterations and zs.reserve.tool_headroom_per_iteration. Keep the second in step with what your tools actually return, or loops end early at the budget boundary. A single web_read at the default 64 KiB max_bytes is roughly 16,000 tokens, well above the 4,000-per-iteration default. Tools the client runs itself (a coding agent's own shell and edit tools) add no headroom, because they never grow the context your node measures.

Source favicons​

After a search round on a streaming request, the node fetches each result site's favicon and inlines the bytes into the round's status frame, so chat clients can show real site marks on the source chips instead of a generic link glyph. It's on by default and configured under zs.builtin_tools.favicons.

The node fetches them so the client doesn't. A browser loading https://thatsite.com/favicon.ico itself would tell every host in a search result the user's IP address and the moment they looked, and that set of hosts is chosen by a third-party search backend steered by a model-generated query. Fetching server-side puts the bytes inside the sealed response, where the relay can't read them either.

It only fires on streaming /v1/responses. Neither /v1/chat/completions nor a non-streaming request has a status frame to carry the icons, so a chat-heavy node shows no favicon egress at all.

This is the one path that makes your node contact hosts it was not otherwise going to fetch: a search result the model never reads. On a node with metered or locked-down egress, set favicons.enabled: false; clients then show the same generic glyph they show for any site without a favicon. The timeout, per-icon cap, and cache knobs are in Configuration.

Image tools​

Image generation and editing inside a text turn are enabled per model, not globally. A chat model opts in under zs.models[<id>].image_tools.{generation,edit} with a positive per-image rate, and the images are produced by the same image_llm backend that powers the dedicated /v1/images/* routes. Set up the backend and price these in Image generation.

Observability​

When built-in tools are enabled, the node exposes tool and loop metrics on its private listener, among them zs_builtin_tool_invocations_total, zs_builtin_tool_errors_total, zs_builtin_tool_duration_seconds, zs_tool_loop_finished_total{outcome}, and zs_tool_call_leak_total. See Monitoring & metrics.