Skip to main content

Confidential compute (TEE)

In standard mode, prompts are end-to-end encrypted to your node (see Encryption & keys), and your node decrypts them to run the model, so you could read the plaintext. In confidential mode, your node runs inside a hardware-attested confidential VM (CVM), and the decryption key is generated inside the attested hardware. Payers get hardware-signed proof that you cannot read prompts, and a payer who requires attested hardware checks it before sending you anything.

This page is the operator's side: hardware, configuration, deployment, and keeping your node verifiable. What the attestation proves, and how a payer checks it, is in Verifying an attestation yourself.

Status by tee.mode:

ModeStatusCPU TEEGPUNotes
dstack-tdxWorking end to endIntel TDX, in a CVM managed by dstack (Phala Cloud or self-hosted)Not required; an attached GPU is not attestedThe node mints a TDX quote bound to its sealing key and publishes it with the runtime measurement log. A proxy verifies the quote against Intel's certificate chain, replays the log, and checks the result against its allowlist. Never calls NRAS.
nvidia-cc-tdxForthcomingIntel TDXNVIDIA GPU in CC modeThe node side is not wired. A proxy tags this evidence verifier_not_implemented and will not route Require-TEE traffic to it.
nvidia-cc-snpForthcomingAMD SEV-SNPNVIDIA GPU in CC modeSame as nvidia-cc-tdx.
stubDevelopment onlyNoneNoneDeterministic, well-formed but untrusted evidence for a laptop. Production proxies reject it unless explicitly opted in.

What confidential mode means for you​

  • There is no encryption key to manage. The node generates a fresh age (X25519) ephemeral identity from the in-enclave RNG, on boot and on every ~20-minute rotation. The private half never exists outside CVM-encrypted memory, so there is no host-side age-keygen step and nothing to copy in.
  • You run a published release, unmodified. The attestation measures the node, its container images and your configuration. A modified build, an extra container or a root-login setting still boots, but payers no longer count it as attested, and it drops out of routing for anyone who requires attested hardware. The full list is what the operator cannot change.
  • You declare where plaintext ends up. Either the node runs the weights inside the CVM, or it forwards to one named zero-retention upstream; see Two postures. The node publishes that choice, and the validator refuses a configuration that contradicts it.

Confidential mode seals content only. You and the payer still see metadata such as ticket IDs, model names and token counts, and attestation does not defend against side channels, so keep your CPU microcode patched. What this does not prove has the full list.

Hardware and platform requirements​

dstack-tdx needs an Intel TDX host and nothing else: no GPU, no NVIDIA driver, no nvtrust. Both postures run on it: sealed_local serves a CPU-sized model from inside the CVM, and attested_passthrough runs no model locally. On Phala Cloud, pick a TDX instance.

The nvidia-cc-* modes need both a CPU TEE and an NVIDIA GPU in CC mode. The CPU TEE protects host memory, but inference happens in GPU HBM, which only GPU CC protects.

CPU TEE​

Either of:

PlatformMinimumNotes
Intel TDXSapphire Rapids (4th-gen Xeon) or newerMost cloud confidential offerings (Azure NCC, GCP CC) target TDX.
AMD SEV-SNPEPYC Genoa (Zen 4) or Milan (Zen 3)More common for self-hosted bare metal.

Earlier generations do not qualify. Intel SGX (enclave-only, not whole-VM encryption) and AMD SEV-ES (no attested measurement of the full image) are not supported. The threat model requires whole-VM memory encryption and an attested measurement of the entire OS image.

GPU TEE — nvidia-cc-* only​

For the forthcoming nvidia-cc-* modes, NVIDIA Confidential Computing requires H100, H200, or Blackwell (B100/B200). A100 is not supported; it has no on-die security processor.

You also need:

  • NVIDIA driver ≥ 535 with CC support compiled in.
  • The nvtrust SDK, accessible to the node, to generate the GPU EAT. On Phala/dstack it's bundled; on Azure NCC it ships in the confidential-VM image.
  • A GPU bootstrapped in CC mode at provisioning time. CC mode is a host-level setting made before the CVM starts; it is not toggleable per request. Confirm it from inside the CVM with nvidia-smi conf-compute -f (must report ON).

Inference engine inside the CVM — sealed_local only​

Under attested_passthrough there is no local engine; the node forwards to the one named upstream.

Under sealed_local, run a local inference engine in the same CVM as the node, one of:

  • llama-server from llama.cpp, supervised by the node's local provider (the engine documented in Serving models & pricing). Set llm.provider: local.
  • A digest-pinned engine container in the same compose file, reached over loopback: kronk, lmstudio, or llamacpp. Pinning it by digest puts it inside the measurement. Available under tee.mode: dstack-tdx and stub only.

vLLM does not work under sealed_local today. The node reaches vLLM only through llm.provider: openai_passthrough, which sealed_local refuses even on 127.0.0.1 (see below). A dedicated provider for an in-CVM vLLM sidecar is on the roadmap. When it ships, vLLM will need --enforce-eager under GPU CC, because its CUDA-graph compilation does not yet work cleanly there (~3–5% throughput hit).

The node and its inference engine must run inside the same CVM, talking over loopback or a Unix socket. Splitting them across VMs, even confidential ones, would put the plaintext prompt on a network the host can see.

Running the engine on a GPU: what is and isn't proven

We have run this setup on a dstack CVM with an H200, and it attests cleanly. It is a dstack-tdx CPU-TEE deployment with a GPU attached, not the nvidia-cc-* GPU confidential-computing modes. There is no GPU attestation in it (POST /v1/AttestGpu is not available on the dstack NVIDIA 0.5.x line), so nothing here says anything about what is in GPU memory. It does prove which weights are loaded, to the byte.

On zs-node 0.23.0 or later, that is the only gap in the evidence a payer checks. Kronk keeps its native llama.cpp library on a volume that is not part of the measurement, and verifies it at startup: the Kronk binary carries the digest of the manifest for the llama.cpp build it pins, and checks every file that manifest lists. That includes the library Docker copies from the engine's image onto a new volume; the published image ships one, so nothing is downloaded. The binary is inside the measurement, so the digest is too, and a listed file that is missing or altered stops the engine from starting.

This check needs Kronk 1.32.5 or later, which the GPU deployment of zs-node 0.23.0 and later pins. Payers do not count the GPU deployments of earlier releases as attested, because their engine does not verify the library. A GPU node on one of them keeps serving ordinary traffic, and must upgrade to be routed requests that require attested hardware.

The digest check does not change who builds those libraries. They are built and published by github.com/hybridgroup/llama-cpp-builder, a third-party rebuild of llama.cpp, not by llama.cpp's own releases. The baked digest lets you verify you got that project's artifact; it does not make the artifact an upstream llama.cpp release.

Two things follow for your deploy. zs-node tee compose --gpu (0.23.0 or later) handles the first: it pins an engine that verifies, and leaves KRONK_LIB_VERSION unset, because the version baked into the engine is the one verification runs against, and setting it by hand weakens or breaks the check.

The second is yours, and applies in one case only. Upgrading an existing CVM keeps its volume. Normally that is fine: a library older than the engine's baked version is replaced and verified as it is installed. It fails when the library on disk is newer, which an engine that ran unpinned could have left behind. Kronk then keeps that library, has no record to verify it against, and refuses to start. There is no shell into the CVM to delete only the library, so the fix is a new volume and a re-pull of your models. Only this case needs a new volume.

Generate this deployment's compose file with zs-node tee compose --gpu. It is a published release like the CPU one, so payers verify it with no special configuration. The engine's image is inside the measurement, so changing it is a release change, not a configuration change. This lets a payer verify which engine served them.

Two postures, and why you have to pick one​

A hardware quote proves which software ran. It does not say where your users' prompts come to rest, so you declare that in tee.dataflow, and the node publishes it to payers.

sealed_local (default)attested_passthrough
Where plaintext ends upInside the CVM. The node runs the weights, and nothing leaves the machine to answer the request.One named upstream, under its zero-retention terms.
Text backendAn engine inside the CVM: llm.provider: local, or lmstudio / llamacpp / kronk on a loopback base_urlA passthrough to an upstream the node collects evidence about: xAI, whose zero retention it confirms on every response, or an ACI/1 gateway whose enclave it verifies
Image backend, if anycomfyui on this machinecomfyui on this machine, or a hosted backend with confirmed zero retention
Egress built-in toolsNot constrained — caller opt-in, see belowNot constrained — caller opt-in, see below
Advertised retention tiertee_attestedupstream_confirmed

The validator refuses to start a node whose backends don't match its declared posture.

The built-in tools are not part of that check. zs_web_search sends a model-authored query derived from the prompt to a search backend, zs_web_read fetches URLs taken verbatim from it, and zs_image_search does both. That can look like a reason to refuse them on a confidential node.

The node does not refuse them, because the caller makes that choice. A built-in tool runs only because the caller asked for it in that request, by listing it in the request's tools[] array. A user who leaves web search off gets the posture's claim without qualification, and that follows from the encrypted request they sent, not from anything you advertise. Refusing the tools would take that choice away from them, and the only way to offer web search would be a node with no tee.mode at all.

The posture therefore describes the inference route: where your node sends the prompt in order to answer it. Leave the tools on and they show up in your node's builtin_tools list, where a user's client reads them and offers the switch. The node logs one warning at startup saying so.

You decide whether to offer them at all. To withhold them from every caller, set:

zs:
builtin_tools:
web_search: {enabled: false}
image_search: {enabled: false}
web_read: {enabled: false}

zs_get_time is unaffected either way; it runs locally and makes no network request.

The image backend row in the table above is different. You choose where image_llm points, a user can't see or choose it, and a dedicated image model reaches it without any tool call. That row is a requirement, and the validator enforces it.

sealed_local — the default​

llm.provider: local is accepted under every mode. A terminal runtime on a loopback base_url must run as a digest-pinned service in the same measured compose file, and is accepted under dstack-tdx and stub only, because those are the modes where something measures a second container. Under nvidia-cc-tdx / nvidia-cc-snp it is refused, with an error saying so.

openai_passthrough is refused even on 127.0.0.1: the node cannot see whether that endpoint forwards requests onward, and to the node a local relay to a hosted API looks the same as a local engine.

tee.mode="dstack-tdx" with tee.dataflow="sealed_local" requires an engine
inside the CVM: llm.provider=local (the supervised LocalLlamaProvider,
node/internal/llm/local.go) or a terminal runtime on a loopback base_url
(lmstudio, llamacpp, kronk), pinned by digest as a service in this same
measured compose. Got llm.provider="openai_passthrough" at
"http://127.0.0.1:8080/v1", and forwarding plaintext there would leak the
very prompts the TEE is meant to seal. […]

What the validator can and cannot check here. It checks the provider name and that base_url is loopback. It cannot check that your engine is digest-pinned or in the compose file; the payer verifies that against the measured compose document your node publishes. A node pointing at an unpinned engine boots and advertises tee_attested, but clients reject the compose, so it gets no traffic from payers who require attested hardware. You are responsible for pinning the engine.

attested_passthrough​

The node decrypts inside the CVM and forwards to exactly one named upstream. A payer is told:

Your prompt is readable by exactly one party — the named upstream — under their zero-retention terms. The node operator is not that party and cannot become it.

This guarantee is weaker than sealed_local, so you must declare it explicitly. The default is sealed_local, so setting tee.mode while pointed at a hosted API fails startup. The node does not fall back to the weaker claim under the same badge.

There are two ways an upstream qualifies, and both are evidence the node collects itself:

  • xAI, whose zero retention the node confirms from the x-zero-data-retention: true header on every reply. There is no knob to disable that check. The node decides that an upstream is xAI from the URL's host, which must be x.ai or a subdomain of it. A host such as api.x.ai.example.com, or a URL with api.x.ai in its path, is not xAI.
  • An upstream that attests its own enclave, verified by the node — see Verifying your upstream's enclave below.

Either way, llm.openai.base_url must be an https URL with a plain host name and no user@ part, and so must image_llm.openai.base_url. Over plain http, anyone who controls the network or DNS of the machine hosting your CVM could answer as the upstream. Payers check the upstream your attestation names against the same rule.

llm.upstream_zdr_declared does not qualify an upstream: it is an assertion the node cannot check, and attesting it does not make it verifiable. Supporting a third upstream requires building an enforcement mechanism for it; adding its name to a list is not enough.

The image backend must meet the first of these requirements, confirmed zero retention. The enclave check covers the text route only, and an image_llm on the gateway's host is refused. image_llm receives the prompt of every zs_image_generation and zs_image_edit call and of every request to a dedicated image model, and you choose where it points, so the validator holds it to the posture (the image backend row in the table above). zs_image_search is one of the built-in web tools and stays a caller opt-in. comfyui_cloud never runs on this machine, so it is refused under both postures.

caution

previous_response_id continuation does not work on an attested_passthrough node. The node forces store: false even when a client explicitly asks for store: true. You own the upstream account, so upstream retention would otherwise let you read prompts back while every attestation claim you make stays literally true.

Verifying your upstream's enclave​

Some inference providers run inside a TEE of their own and publish evidence for it. If yours speaks ACI/1 (Attested Confidential Inference v1 — Phala's https://inference.phala.com/v1 is the one deployment today), your node can check that evidence itself instead of taking the upstream's word:

tee:
mode: dstack-tdx
dataflow: attested_passthrough
upstream_attestation:
protocol: aci/1
require_upstream_providers: ["phala"]
allow_root_backdoor_env: ["DSTACK_ROOT_PUBLIC_KEY"]
refresh_interval: 15m

At startup the node fetches the gateway's hardware attestation with a nonce it chose, replays the boot log against the quote, checks the measured configuration, follows the certificate chain to Intel, and pins the TLS key the quote commits to so nothing can answer for that hostname afterwards. Then it checks every model you declared. If any of those checks fails, the node exits instead of starting with a claim it cannot back.

require_upstream_providers is the field that matters most and it has no safe default, so the node refuses to start with it empty. An ACI gateway is a router: it verifies each backend's enclave and serves several independent operators from one address. Verifying the gateway establishes which code read your payer's prompt first, not who ran the model. On 2026-09-24 that endpoint fanned 272 live sessions out to four operators — 129 to Phala's own GPU CVMs, 126 to Chutes, 16 to NEAR AI, 1 to Tinfoil — so without the list, prompts you sealed for one party reach three others. Every request also carries a pin that the gateway refuses to serve unless it verified the backend first, and every response carries a signed receipt naming the enclave that answered.

allow_root_backdoor_env waives part of the upstream check. dstack can be told to install an operator-held root key inside a CVM, and Phala's gateway declares that channel. Our own nodes are refused for it outright; an upstream may declare it, and if you accept it here, the exact list is republished on your evidence as upstream_attestation.carve_outs for a payer to read. Only real dstack channel names are accepted, so the list cannot overstate what you waived.

It is also honoured only on a guest image whose filesystem has been enumerated and found to contain no SSH daemon. With no SSH daemon, nothing in the image can use the root key, which is why accepting it is defensible. Today that is the CPU image, dstack-0.5.9. On any other image, including the GPU line's dstack-nvidia-0.5.9, the waiver is withheld and the node refuses the upstream at boot. If that happens the error names the image it saw. Adding entries to this list does not fix it: the image has not been scanned, and no configuration change can supply that.

caution

Tell payers how many enclaves can read a prompt; don't say "nobody". With this in place a prompt is readable inside three measured enclaves on three codebases (your CVM, the gateway's, and the model's GPU CVM) and on no operator's disk. That is stronger than any hosted API, but it is not the sealed_local claim. Four things stay open: which model answered is not measured; the inference engine is pinned by digest but nobody attests the build was vetted; the GPU attestation proves a genuine confidential GPU, not that it is the one bound to the serving CPU enclave; and any root-shell channel you accepted weakens "the measurement pins what runs" for that upstream.

For a stream, the pin is applied before any byte is sent, but the receipt exists only once the response is finished. A stream is therefore pinned before it starts and verified after it ends. A receipt that fails verification de-routes the upstream and fails the next request; it cannot withhold a response already delivered.

None of this changes the retention tier your models advertise today. It adds evidence of where the prompt went.

Why not OpenRouter?​

Unless you have released it with llm.allow_upstream_retention (see Releasing the pin), your node pins zero data retention on every OpenRouter request and refuses to start if a model you declared has no zero-retention endpoint. That pin takes effect earlier than xAI's check, but it produces less evidence, and this posture promises evidence:

AxisOpenRouter pinxAI header
Prevention timingApplied before routing, so a retaining endpoint is never selected.Read after the prompt has reached xAI.
ConfirmationNothing comes back. A broker that ignored the pin looks identical to one that obeyed it.Arrives with every response, so a lapse is caught on the next request.
The upstream's own retentionOpenRouter itself reads the prompt on the way through, and its prompt logging is a separate account setting that you own and no API exposes for reading. Turning it on would let you read every payer's prompt back. Nothing like the forced store: false closes it.The node forces store: false.
What "zero data retention" meansIncludes implicit KV caching to on-premises SSDs, unlike Google's reading of the same words.—
Who reads the promptThe posture would name https://openrouter.ai/api/v1, but the reader is whichever provider OpenRouter picked for that request, never published. The measured image fixes where the node sends a prompt, not who reads it.xAI, the named upstream.

Pinning a single sub-provider in your config answers only the last row.

What to do instead: run your OpenRouter upstream on a node with no tee.mode. As long as you have not released the pin, it stays in force and your models advertise retention: upstream_enforced. You lose the confidential-compute badge, which never applied to this setup.

The tee: config block​

The whole opt-in lives under one top-level YAML block, alongside the zs: and llm: sections (see Configuration):

tee:
# mode: which confidential-computing platform this node runs on.
# none - default; standard trust model, not a TEE node.
# stub - dev only; deterministic, well-formed-but-untrusted
# evidence. Exercises the whole pipeline on a laptop with
# no special hardware. Proxies reject stub evidence in
# production unless explicitly opted in.
# dstack-tdx - CPU-only Intel TDX CVM managed by dstack (Phala Cloud
# and self-hosted). The mode with a working verifier.
# nvidia-cc-tdx - Intel TDX host + NVIDIA GPU CC. (forthcoming)
# nvidia-cc-snp - AMD SEV-SNP host + NVIDIA GPU CC. (forthcoming)
mode: dstack-tdx

# dataflow: where plaintext comes to rest. sealed_local (default) runs the
# weights in the CVM and requires an engine there; attested_passthrough
# forwards to one named zero-retention upstream. See "Two postures" above —
# this is the gate, and forgetting it fails the boot rather than quietly
# weakening the claim.
dataflow: sealed_local

attestation:
# NVIDIA Remote Attestation Service endpoint. Required for the nvidia-cc-*
# modes only — dstack-tdx mints its quote from the local guest agent and
# never calls NRAS. Leave at the default unless you run a private
# attestation deployment.
nras_url: https://nras.attestation.nvidia.com

# How often the node mints fresh attestation evidence. /v1/zs/attestation
# serves the cached bundle and re-mints on this cadence. Evidence older than
# 2x this interval is served as 503 (stale), and the proxy demotes you out of
# the TEE-eligible set until fresh evidence is available.
refresh_interval: 1h

# Where to persist the latest evidence bundle so a brief restart doesn't drop
# you out of the verified set. MUST live on a CVM-encrypted volume — the
# bundle binds to your current ephemeral recipient and the host operator must
# not be able to read past evidence. Unset = mint fresh on every boot.
evidence_cache_path: /var/lib/zs-node/attestation/last.json

# Where this node reads Intel's revocation and TCB collateral from, to carry
# alongside its quote. Unset = Phala's CORS-enabled mirror. See "Collateral
# your node carries" below before changing it — pointing this at Intel's own
# PCS is the tempting wrong answer.
pccs_url: https://pccs.phala.network

# upstream_attestation: OMIT THIS unless your upstream speaks ACI/1. It is
# what turns attested_passthrough from "one upstream I trust" into "one
# upstream I check", and it requires llm.provider: openai_passthrough — a
# locally-hosted runtime publishes no attestation to appraise. See
# "Verifying your upstream's enclave" above for what each field does and
# what accepting a carve-out costs you.
#
# upstream_attestation:
# protocol: aci/1
# require_upstream_providers: ["phala"]
# allow_root_backdoor_env: ["DSTACK_ROOT_PUBLIC_KEY"]
# refresh_interval: 15m

None of the upstream_attestation fields has a NODE_TEE_* environment override, for the same reason tee.dataflow has none: on dstack your config is literal text inside the measured document, while environment values are substituted after that document is hashed. An override would let a node publish one carve-out list in the measurement and appraise against another. The field a payer reads to see what you waived would then describe a policy your node never applied.

Collateral your node carries​

Verifying a quote also takes Intel's current revocation lists and TCB status for your platform. A verifier normally fetches those itself. Your node also fetches them at mint time and publishes them in the bundle, so a payer whose network cannot reach Intel can still verify you.

  • It is never fatal. A PCCS outage costs the bundle its collateral and nothing else. Your node mints, boots, and serves normally, and a verifier with no carried set fetches its own.
  • It is a last resort for the verifier. A verifier uses your copy only when it has neither live collateral nor a usable cached copy of its own, and never to overturn a refusal. Where the Intel collateral comes from explains why.
  • Keep the default; don't point it at Intel's own PCS. A browser payer cannot reach api.trustedservices.intel.com at all, because it sends no CORS headers, so a node pointed there ships documents its clients could not have obtained independently. Change pccs_url only for a PCCS inside your own network.

Validation at startup​

The config validator refuses to start when:

  • The backends don't match tee.dataflow. See Two postures for what each posture requires. The built-in tools are not part of this — they are a caller opt-in, and the same section explains why.
  • tee.mode is not one of none, stub, dstack-tdx, nvidia-cc-tdx, nvidia-cc-snp. An unknown tee.dataflow is rejected too; neither is silently defaulted.
  • nras_url is missing under an nvidia-cc-* mode. stub and dstack-tdx don't need it.
  • refresh_interval is not positive (not checked under stub). Keep the default 1h. On dstack-tdx the node also re-mints on every ~20-minute key rotation, whatever this is set to (see Key rotation), so a shorter interval adds mints and little freshness. The interval also sets how old evidence may get before the node stops serving it (503 stale_attestation) and payers refuse it: 2× the interval, which payers cap at 80 minutes.
  • llm.openai.debug_dump_errors is true under any non-none mode. It writes decrypted prompt content to stderr.

Environment variables​

Most fields have a NODE_TEE_* override (env wins over YAML):

VariableMaps to
NODE_TEE_MODEtee.mode
NODE_TEE_ATTESTATION_NRAS_URLtee.attestation.nras_url
NODE_TEE_ATTESTATION_REFRESH_INTERVALtee.attestation.refresh_interval (Go duration: 1h, 30m, 15m)
NODE_TEE_ATTESTATION_EVIDENCE_CACHE_PATHtee.attestation.evidence_cache_path
NODE_TEE_ATTESTATION_PCCS_URLtee.attestation.pccs_url

tee.dataflow has no env override, on purpose. On dstack your config is part of the measured compose file, while environment values are delivered encrypted and separately, so they are not measured. An override would let a node run attested_passthrough behind an image a verifier had confirmed as sealed_local.

The tee: block holds no secrets, so the whole thing is safe to commit to YAML.

How the in-CVM key binds to your attestation​

The node binds the in-CVM key to the attestation automatically. It reuses a standard node's ephemeral key handling (see Encryption & keys) and adds a hardware proof on top.

The encryption key is never published on chain. Your on-chain record anchors only your identity, the Ed25519 signing key, and a signature binds the ephemeral recipient to it. There is no placeholder pubkey to register, no owner-signed pubkey rotation, and no updateOperator call. The flow:

  1. Register normally. Operator and node registration take your owner and signing addresses, base URL, and stake, with no encryption key (see Registering on-chain). Bake the issued operator_id into the CVM config (zs.operator_id in YAML, or NODE_ZS_OPERATOR_ID in env). The attestation binds to this id, so the CVM must know it at boot.

  2. Boot the CVM with tee.mode != none. Inside the enclave, the node generates a fresh age (X25519) ephemeral recipient from the in-enclave RNG, signs it under your on-chain signing key, and advertises it on /v1/zs/details (ephemeral_age_pubkey + ephemeral_sig), as a standard node does. The private half lives only in CVM-encrypted RAM.

  3. The node mints attestation evidence, served from GET /v1/zs/attestation. Its report_data field is a hash of the current ephemeral recipient and your operator_id, so the node re-mints on every rotation. The field holds the first 32 bytes of the hardware's 64-byte REPORT_DATA; the last 32 are the model binding.

  4. Payers verify the binding from your on-chain record, not from the bundle: the ephemeral must carry a valid signature from your registered signing address, and the hash recomputed from it must equal the signed report. The exact steps, and why the evidence names the ephemeral rather than your signing key, are in The key binding.

To sanity-check a deployment, read the live values from outside:

NODE_URL=https://your-node.example.com
curl -s "$NODE_URL/v1/zs/details" | jq '{ephemeral_age_pubkey, ephemeral_sig, tee}'
curl -s "$NODE_URL/v1/zs/attestation" | jq '{report_data, generated_at, app_models}'
info

No persistence required. The ephemeral private key is never written to disk, not even to a CVM-encrypted volume, and is zeroed on each ~20-minute rotation. A fresh key on every boot is normal; the node attests the new one. The optional evidence_cache_path persists only the attestation evidence bundle, never the key, so a brief restart doesn't drop you from the verified set while fresh evidence is minted.

Key rotation​

The node rotates its in-CVM ephemeral on the same fixed cadence as a standard node. The evidence binds the current ephemeral, so every rotation triggers a re-mint, not only the refresh_interval cadence. The rotation hook requests the re-mint, and the refresh loop independently checks that the cached bundle still names the live key, so a rotation that slipped past the hook is caught rather than left advertising evidence for a key the node no longer holds. A node that has rotated but not yet re-minted is refused until the next probe picks up the fresh bundle.

The model binding​

The last 32 bytes of REPORT_DATA are a hash of your node id, your model catalog (each model's id, source, digest and how the digest was produced) and, since protocol 9.10, your node's posture: where plaintext ends up, and for a named upstream, its URL, whether zero retention is enforced, and whether you configured tee.upstream_attestation. Prices and tools are not included. The node publishes the hashed list as app_models and the posture as posture on /v1/zs/attestation so anyone can recompute it. Confidential mode turns this on; there is nothing to configure. How a verifier checks it is in The model binding.

When the signed list changes​

The node re-mints on key rotation (every 20 minutes), when it finishes hashing model weights, and when its catalog changes, for example when an upstream goes unhealthy. It checks the catalog every 15 seconds but re-mints for a change at most once a minute, so /v1/zs/details and app_models can disagree for up to about 75 seconds.

After a restart the node hashes model weights in the background, which can take several minutes. Until it finishes, every model is listed as unverifiable. The first evidence after boot lists no models; verifiers accept it.

If the node builds a model list that cannot be hashed, for example two entries with the same id, it does not mint. It logs tee attestation refresh failed with cannot measure the model set this node serves in err. This is a node bug; report it with that log line. Meanwhile the node keeps serving its previous evidence, which payers stop accepting at the next key rotation, within 20 minutes, because it binds a key the node no longer uses.

Freshness challenges​

A caller can request /v1/zs/attestation?nonce=HEX with a random value of its own, and the enclave mints a new quote over that value (details). The reply is not cached, and your cached evidence never carries a caller's nonce. Challenges are rate-limited across the whole node; see Configuration.

When the model binding fails​

A node refused for one of the reasons below loses its attested verdict. It keeps serving ordinary traffic, is shown to payers as unattested with the reason attached, and drops out of routing for anyone who requires attested hardware. Your node mints evidence normally, so its log never shows why that traffic stopped. To see the reason yourself, open the card for a model your node serves in the chat app, page to your node, and hover the attestation badge.

ReasonWhat it meansWhat to do
aux_binding_absentThe last 32 bytes are all zero, which is what a node older than 9.9 mints. The chat app shows this as Node software out of date.Upgrade zs-node and redeploy the CVM.
aux_binding_mismatchThe signed hash does not match your node id, app_models, posture and nonce.Look for something that rewrites responses between your node and the payer: a reverse proxy in front of your node, or the relay the payer fetched through. If the chat app shows your node as attested, the problem is on that payer's path. A node and a payer on opposite sides of protocol 9.10 also refuse each other this way, so upgrade whichever is older.
happ_preimage_missingapp_models is absent or empty, but the quote does not commit to an empty catalog.Check that nothing between your node and the payer rewrites the response body. A reverse proxy that drops unknown JSON fields will do this.
happ_preimage_invalidapp_models cannot be hashed: a malformed entry, an empty or duplicate model id, an unknown weights_state, or a state that contradicts its digest.Your node refuses to mint over such a list, so the likely cause is a rewritten response body.
posture_preimage_invalidposture cannot be hashed: it is not an object, or a field has the wrong type.Your node never publishes one like that, so look for a rewritten response body.
posture_upstream_unverifiableThe signed posture names an upstream the payer will not accept. The URL must be https with a plain host, and then either be xAI with zero-retention enforcement, or the node must be configured with tee.upstream_attestation and the response must carry its aci/1 upstream_attestation block.A current node refuses to start unless one of those holds, so on an attested upstream gateway the block is usually missing. First check your node's logs for a failed upstream appraisal: the node drops the block while its last check of the upstream's enclave has failed or expired, and requests fail with upstream_attestation_required meanwhile. Otherwise check that nothing between your node and the payer strips the block from the response.

A payer on a zs-proxy older than 9.9 refuses a 9.9 node with key_binding_mismatch, because that release requires the last 32 bytes to be zero. The payer has to upgrade; nothing on your node fixes it.

A wrong zs.node_id fails before the model binding is checked. Boot stops with node X/Y not found on chain or keystore: signing address not loaded. If the id belongs to a sibling node that shares this node's hot key, the node boots, but its signed ephemeral key names the wrong node. Payers then reject it and cannot seal to the node at all, confidential or not. They report key_binding_mismatch.

Validate the whole pipeline with stub mode​

Before you pay for CVM time, exercise the entire confidential-mode path on your laptop with tee.mode: stub. Stub mode produces a deterministic, well-formed but untrusted bundle, so every component on both sides runs the code it runs in production: the node's config gate, the /v1/zs/details advertisement, the /v1/zs/attestation handler, the proxy verifier, the routing filter, and the dashboard badge. Only the hardware-rooted evidence is missing. A stub bundle has no quote, so it has no model binding either.

Add the block to the same config.yaml a non-TEE deployment already uses:

tee:
mode: stub
attestation:
refresh_interval: 1m
evidence_cache_path: /tmp/zerosignal-tee-stub.json

The validator still requires an engine inside the CVM, so point llm.local at your llama-server binary with one model configured (see Serving models & pricing). stub also accepts the loopback-sidecar shape, so you can rehearse a compose-file deployment off-platform; stub evidence is untrusted anyway. Start the node and check what it advertises on the public port (9090 by default):

NODE_URL=http://localhost:9090

curl -s "$NODE_URL/v1/zs/details" | jq .tee
# { "mode": "stub", "attested_at": "...", "evidence_url": "/v1/zs/attestation" }

curl -s "$NODE_URL/v1/zs/attestation" | jq .
# { "mode": "stub", "node_pubkey": "...", "operator_id": 42, "report_data": "...", "stub": true, ... }

Both should return 200. A 404 on /v1/zs/attestation means the node didn't pick up tee.mode: stub; check the startup logs for a validator error.

Production proxies reject stub evidence. A real proxy marks a stub-mode operator unverified and drops it only when a request requires TEE: a sensitive route carrying the X-Zs-Require-TEE header, or a model on that proxy's zs.tee.require_for_models list. Plain non-TEE traffic to the same operator still routes normally. To exercise the gate end to end against your own dev proxy, send that header or opt the dev proxy into accepting stub bundles. That opt-in is dev-only; never set it in production.

Stub mode does not validate the real TDX/SEV-SNP signature chains, the NVIDIA EAT / NRAS round-trip, the CVM memory-encryption guarantee, or the encrypted-volume requirement for evidence_cache_path (the recipe above writes it to /tmp).

Deployment recipes​

info

One recipe works today: dstack (tee.mode: dstack-tdx), CPU-only. The GPU paths below wait on the nvidia-cc-* modes.

Your zs-node binary already has the compose file. Run zs-node tee compose, which writes the measured artifact for your release with the published image digest resolved. Re-run it when you finish editing: with the file already present it verifies instead of overwriting, and tells you if an edit moved the measurement.

  • Your configuration goes in the compose file, as literal text. dstack hashes the compose before substituting ${VAR} references, so a NODE_CONFIG_YAML block written out in full is inside the measurement, while your API keys, which stay ${VAR} references, are not. The block is also published: the node serves the whole document unauthenticated, because it is the preimage a payer hashes. Never inline a key there.
  • The compose file is a published artifact, not a template. Change the image reference and your config block (and, on the GPU shape, the engine's config block under zs-engine-config:) and nothing else. A verifier lifts exactly those spans out and checks the rest, comments included, against the published release. Reformatting anything outside them fails silently: your node deploys, boots, and attests, and payers who require attested hardware route elsewhere. zs-proxy tee inspect <your-node-url> --chain shows you that verdict. The engine block is checked for content rather than matched byte for byte: adapters and template are refused, since either changes what the measured weights produce.
  • Inside your config block, edit freely: the whole block, not just the two EDIT ME markers. The lift covers the entire NODE_CONFIG_YAML body, so max_active_tickets, ticket_ttl, your models and rates, and any setting the example does not mention are yours to change; redeploy and nobody has to be told. The EDIT ME markers show what you must fill in (startup fails if you leave them); you may change other lines too. A few lines need care (the dataflow, the upstream base_url, the drain timings), and a comment on each says so.

Always deploy with --no-dev-os. Without it Phala may install an SSH key as an authorized root key inside the CVM. That undoes the enclave's protection, and payers detect and reject it anyway.

Decide on Phala's --listed flag before you deploy. It is off by default, and adding it later means building a new CVM. It publishes your deployment to Phala's trust center, at https://trust.phala.com/app/<your-app-id>, where Phala re-verifies the same quote and event log your node serves, independently of us. The chat client links payers there from the attestation badge. Without the flag your node attests just as well, but that link leads nowhere, and a payer has only our verification instead of two independent ones.

Phala cannot change the flag in place, so a later phala deploy --cvm-id will not turn it on. Turning it on means a fresh CVM with a new app id and gateway hostname, and a baseUrl to re-register on chain. The node repo's dstack deployment guide has the full command and a one-line check for the state you are in. The related --public-sysinfo / --public-tcbinfo / --public-logs flags default on, can be changed later, and serve Phala's node-information page. Leaving them on costs nothing and gives payers one more independent view.

Once stub validation passes, the real deployment swaps tee.mode: stub for the platform string and removes any dev-only stub opt-in from the proxy. The paths still to come:

  • Phala Cloud / dstack with GPU confidential computing — the same CVM setup as the shipped CPU recipe, on a TDX host with an NVIDIA H200. Closely modeled on Phala's open private-ml-sdk.
  • Azure NCC H100 v5 — TDX hosts with H100 GPUs in CC mode, via a confidential VM (--security-type ConfidentialVM). Best when you want managed Kubernetes or region-specific data residency.
  • Bare-metal SEV-SNP — EPYC Genoa/Milan with SEV-SNP in BIOS (use tee.mode: nvidia-cc-snp) and a CVM-aware hypervisor.
  • GCP H100 confidential — TDX-based; identical config to the Azure recipe.

Pitfalls, and which paths each applies to:

  • Pin every image by digest, never by tag (every path). The compose / VM image hash is what the CPU report measures. A floating tag produces a different measurement on every pull, and verifiers refuse to seal to a measurement that keeps changing.
  • Keep evidence_cache_path on a CVM-encrypted volume (every path). A plain host bind-mount is readable and writable by the host operator, who must not be able to inspect or tamper with past evidence.
  • Allow outbound 443 to NRAS (nras.attestation.nvidia.com) from the CVM (nvidia-cc-* only), or GPU-EAT verification fails.
  • Pin the driver and nvtrust versions together inside the CVM image (nvidia-cc-* only). The combination is part of the attested measurement.

When the GPU modes land, only the mode string you set changes; the tee: config and the verification steps below stay the same.

Verifying a deployment from outside the CVM​

Run these checks from a workstation that is not the CVM host, to see what a third party sees.

  1. Capability advertisement. GET /v1/zs/details → the tee field should report your configured mode and an attested_at within the last 2 × refresh_interval. Anything older means the refresh loop is wedged.

  2. Evidence is fetchable. GET /v1/zs/attestation → 200 with a well-formed bundle. 503 means stale or cold-start; 404 means the node isn't in TEE mode.

  3. The advertised ephemeral is signed. The ephemeral_sig on /v1/zs/details verifies under your on-chain signing address.

  4. report_data binding is consistent. It matches the ephemeral currently advertised, computed as in The key binding. A mismatch right after a rotation is normal for one probe; a persistent one means the rotation hook is not re-minting.

    That field is the first 32 bytes. The last 32 are the model binding, which hashes the bundle's app_models and posture. Compare the list against what you advertise:

    curl -s "$NODE_URL/v1/zs/attestation" | jq -r '(.app_models // [])[].model_id' | sort
    curl -s "$NODE_URL/v1/zs/details" | jq -r '.models[].id' | sort

    After a catalog change or a restart the two can differ for about 75 seconds. A difference lasting longer means the evidence is not being re-minted; check attested_at as in step 1.

  5. Run the verifier against yourself. The steps above show only that your node is internally consistent, which is true of a node nobody routes to. This one runs the hardware, image and compose checks a payer runs. It does not check the two REPORT_DATA bindings, which need the node's on-chain record (its node id and signing address). For those, open the card for a model your node serves in the chat app, page to your node, and hover the attestation badge.

    zs-proxy tee inspect https://your-node.example.com --chain

    Always pass --chain. Without it the quote's signature is never examined, so the report only checks values your node supplied against each other. Read three lines:

    • quote chain — PASS, and which collateral it used. Plain "verified against Intel's PCK chain" means live collateral. It can also say the pass used a stale cache or the collateral your node supplied. Both are genuine passes, and both mean the checking party could not reach Intel just then.
    • compose-hash gate — PASS means the document you published is the one your CVM measured.
    • violations and release match — ENFORCED with violations: none is the routable state. A skeleton or image that is not on the lists fails silently: your node deploys, boots, attests, and payers who require attested hardware route elsewhere, with nothing in your logs.
  6. An end-to-end Require-TEE request lands. A request carrying X-Zs-Require-TEE is routed to your node and returns a normal completion.

info

Your release has to be on the lists; you don't add it yourself. The two values tee inspect prints, the compose skeleton digest and the image reference, ship compiled into zs-proxy and the browser client, and every zs-node release adds its own through an automated pull request into both. The only requirement is to run a published release unmodified. If you build your own image, no payer counts it as attested, however valid the quote.

Monitoring​

The standard monitoring on the private listener (/healthz, /livez, /metrics on 9091) still applies. Confidential mode adds two things to watch:

EndpointWhat to watch
GET /v1/zs/attestationShould respond 200 with generated_at newer than 2 × refresh_interval. A 503 (stale) means the refresh loop is failing — check node logs for the underlying mint error (on dstack-tdx, the call to the local dstack guest agent; on nvidia-cc-*, NRAS / nvtrust).
GET /v1/zs/details (tee field)Should report your configured mode and an attested_at matching the most recent successful refresh.

Make sure /v1/zs/attestation is reachable on your public HTTPS URL; if a TLS reverse proxy fronts the node, expose it there. Proxies fetch it during their per-operator details poll, and the verifier needs it to keep you in the TEE-eligible set.

Where this fits​

  • Verifying an attestation yourself covers what the attestation establishes: every value that gets measured, what each one proves, how a payer confirms it, and what it leaves open. This page covers what you do to pass those checks.
  • Encryption & keys describes the standard trust model this mode upgrades.
  • The sealed_local engine rule ties into provider selection in Serving models & pricing; the tee: block sits alongside the rest of your Configuration.
  • Confidential nodes still relay and are still relayed for, unchanged (see Relays). Relaying never touches plaintext on either side.