Confidential compute (TEE)
In standard mode, prompts are end-to-end encrypted to your node (see Encryption & keys), and your node decrypts them to run the model, so you could read the plaintext. In confidential mode, your node runs inside a hardware-attested confidential VM (CVM), and the decryption key is generated inside the attested hardware. Payers get hardware-signed proof that you cannot read prompts, and a payer who requires attested hardware checks it before sending you anything.
This page is the operator's side: hardware, configuration, deployment, and keeping your node verifiable. What the attestation proves, and how a payer checks it, is in Verifying an attestation yourself.
Status by tee.mode:
| Mode | Status | CPU TEE | GPU | Notes |
|---|---|---|---|---|
dstack-tdx | Working end to end | Intel TDX, in a CVM managed by dstack (Phala Cloud or self-hosted) | Not required; an attached GPU is not attested | The node mints a TDX quote bound to its sealing key and publishes it with the runtime measurement log. A proxy verifies the quote against Intel's certificate chain, replays the log, and checks the result against its allowlist. Never calls NRAS. |
nvidia-cc-tdx | Forthcoming | Intel TDX | NVIDIA GPU in CC mode | The node side is not wired. A proxy tags this evidence verifier_not_implemented and will not route Require-TEE traffic to it. |
nvidia-cc-snp | Forthcoming | AMD SEV-SNP | NVIDIA GPU in CC mode | Same as nvidia-cc-tdx. |
stub | Development only | None | None | Deterministic, well-formed but untrusted evidence for a laptop. Production proxies reject it unless explicitly opted in. |
What confidential mode means for you
- There is no encryption key to manage. The node generates a fresh
age (X25519) ephemeral identity from the
in-enclave RNG, on boot and on every ~20-minute rotation. The private half
never exists outside CVM-encrypted memory, so there is no host-side
age-keygenstep and nothing to copy in. - You run a published release, unmodified. The attestation measures the node, its container images and your configuration. A modified build, an extra container or a root-login setting still boots, but payers no longer count it as attested, and it drops out of routing for anyone who requires attested hardware. The full list is what the operator cannot change.
- You declare where plaintext ends up. Either the node runs the weights inside the CVM, or it forwards to one named zero-retention upstream; see Two postures. The node publishes that choice, and the validator refuses a configuration that contradicts it.
Confidential mode seals content only. You and the payer still see metadata such as ticket IDs, model names and token counts, and attestation does not defend against side channels, so keep your CPU microcode patched. What this does not prove has the full list.
Hardware and platform requirements
dstack-tdx needs an Intel TDX host and nothing else: no GPU, no NVIDIA driver,
no nvtrust. Both postures run on it: sealed_local serves a CPU-sized model
from inside the CVM, and attested_passthrough runs no model locally. On Phala
Cloud, pick a TDX instance.
The nvidia-cc-* modes need both a CPU TEE and an NVIDIA GPU in CC mode.
The CPU TEE protects host memory, but inference happens in GPU HBM, which only
GPU CC protects.
CPU TEE
Either of:
| Platform | Minimum | Notes |
|---|---|---|
| Intel TDX | Sapphire Rapids (4th-gen Xeon) or newer | Most cloud confidential offerings (Azure NCC, GCP CC) target TDX. |
| AMD SEV-SNP | EPYC Genoa (Zen 4) or Milan (Zen 3) | More common for self-hosted bare metal. |
Earlier generations do not qualify. Intel SGX (enclave-only, not whole-VM encryption) and AMD SEV-ES (no attested measurement of the full image) are not supported. The threat model requires whole-VM memory encryption and an attested measurement of the entire OS image.
GPU TEE — nvidia-cc-* only
For the forthcoming nvidia-cc-* modes, NVIDIA Confidential Computing requires
H100, H200, or Blackwell (B100/B200).
A100 is not supported; it has no on-die security processor.
You also need:
- NVIDIA driver ≥ 535 with CC support compiled in.
- The
nvtrustSDK, accessible to the node, to generate the GPU EAT. On Phala/dstack it's bundled; on Azure NCC it ships in the confidential-VM image. - A GPU bootstrapped in CC mode at provisioning time. CC mode is a host-level
setting made before the CVM starts; it is not toggleable per request.
Confirm it from inside the CVM with
nvidia-smi conf-compute -f(must reportON).
Inference engine inside the CVM — sealed_local only
Under attested_passthrough there is no local engine; the node forwards to the
one named upstream.
Under sealed_local, run a local inference engine in the same CVM as the
node, one of:
llama-serverfrom llama.cpp, supervised by the node'slocalprovider (the engine documented in Serving models & pricing). Setllm.provider: local.- A digest-pinned engine container in the same compose file, reached over
loopback:
kronk,lmstudio, orllamacpp. Pinning it by digest puts it inside the measurement. Available undertee.mode: dstack-tdxandstubonly.
vLLM does not work under sealed_local today. The node reaches vLLM only
through llm.provider: openai_passthrough, which sealed_local refuses even on
127.0.0.1 (see below). A dedicated provider for
an in-CVM vLLM sidecar is on the roadmap. When it ships, vLLM will need
--enforce-eager under GPU CC, because its CUDA-graph compilation does not yet
work cleanly there (~3–5% throughput hit).
The node and its inference engine must run inside the same CVM, talking over loopback or a Unix socket. Splitting them across VMs, even confidential ones, would put the plaintext prompt on a network the host can see.
We have run this setup on a dstack CVM with an H200, and it attests cleanly. It
is a dstack-tdx CPU-TEE deployment with a GPU attached, not the
nvidia-cc-* GPU confidential-computing modes. There is no GPU attestation in it
(POST /v1/AttestGpu is not available on the dstack NVIDIA 0.5.x line), so
nothing here says anything about what is in GPU memory. It does prove which
weights are loaded, to the byte.
On zs-node 0.23.0 or later, that is the only gap in the evidence a payer checks. Kronk keeps its native llama.cpp library on a volume that is not part of the measurement, and verifies it at startup: the Kronk binary carries the digest of the manifest for the llama.cpp build it pins, and checks every file that manifest lists. That includes the library Docker copies from the engine's image onto a new volume; the published image ships one, so nothing is downloaded. The binary is inside the measurement, so the digest is too, and a listed file that is missing or altered stops the engine from starting.
This check needs Kronk 1.32.5 or later, which the GPU deployment of zs-node 0.23.0 and later pins. Payers do not count the GPU deployments of earlier releases as attested, because their engine does not verify the library. A GPU node on one of them keeps serving ordinary traffic, and must upgrade to be routed requests that require attested hardware.
The digest check does not change who builds those libraries. They are built and
published by github.com/hybridgroup/llama-cpp-builder, a third-party rebuild of
llama.cpp, not by llama.cpp's own releases. The baked digest lets you verify you
got that project's artifact; it does not make the artifact an upstream llama.cpp
release.
Two things follow for your deploy. zs-node tee compose --gpu (0.23.0 or later)
handles the first: it pins an engine that verifies, and leaves
KRONK_LIB_VERSION unset, because the version baked into the engine is the one
verification runs against, and setting it by hand weakens or breaks the check.
The second is yours, and applies in one case only. Upgrading an existing CVM keeps its volume. Normally that is fine: a library older than the engine's baked version is replaced and verified as it is installed. It fails when the library on disk is newer, which an engine that ran unpinned could have left behind. Kronk then keeps that library, has no record to verify it against, and refuses to start. There is no shell into the CVM to delete only the library, so the fix is a new volume and a re-pull of your models. Only this case needs a new volume.
Generate this deployment's compose file with zs-node tee compose --gpu. It is
a published release like the CPU one, so payers verify it with no special
configuration. The engine's image is inside the measurement, so changing it
is a release change, not a configuration change. This lets a payer verify which
engine served them.
Two postures, and why you have to pick one
A hardware quote proves which software ran. It does not say where your users'
prompts come to rest, so you declare that in tee.dataflow, and the node
publishes it to payers.
sealed_local (default) | attested_passthrough | |
|---|---|---|
| Where plaintext ends up | Inside the CVM. The node runs the weights, and nothing leaves the machine to answer the request. | One named upstream, under its zero-retention terms. |
| Text backend | An engine inside the CVM: llm.provider: local, or lmstudio / llamacpp / kronk on a loopback base_url | A passthrough to an upstream the node collects evidence about: xAI, whose zero retention it confirms on every response, or an ACI/1 gateway whose enclave it verifies |
| Image backend, if any | comfyui on this machine | comfyui on this machine, or a hosted backend with confirmed zero retention |
| Egress built-in tools | Not constrained — caller opt-in, see below | Not constrained — caller opt-in, see below |
| Advertised retention tier | tee_attested | upstream_confirmed |
The validator refuses to start a node whose backends don't match its declared posture.
The built-in tools are not part of that check. zs_web_search sends a
model-authored query derived from the prompt to a search backend, zs_web_read
fetches URLs taken verbatim from it, and zs_image_search does both. That can
look like a reason to refuse them on a confidential node.
The node does not refuse them, because the caller makes that choice. A built-in
tool runs only because the caller asked for it in that request, by listing
it in the request's tools[] array. A user who leaves web search off gets the
posture's claim without qualification, and that follows from the encrypted
request they sent, not from anything you advertise. Refusing the tools would
take that choice away from them, and the only way to offer web search would be
a node with no tee.mode at all.
The posture therefore describes the inference route: where your node sends
the prompt in order to answer it. Leave the tools on and they show up in your
node's builtin_tools list, where a user's client reads them and offers the
switch. The node logs one warning at startup saying so.
You decide whether to offer them at all. To withhold them from every caller, set:
zs:
builtin_tools:
web_search: {enabled: false}
image_search: {enabled: false}
web_read: {enabled: false}
zs_get_time is unaffected either way; it runs locally and makes no network
request.
The image backend row in the table above is different. You choose where
image_llm points, a user can't see or choose it, and a dedicated image model
reaches it without any tool call. That row is a requirement, and the validator
enforces it.
sealed_local — the default
llm.provider: local is accepted under every mode. A terminal runtime on a
loopback base_url must run as a digest-pinned service in the same measured
compose file, and is accepted under dstack-tdx and stub only, because
those are the modes where something measures a second container. Under
nvidia-cc-tdx / nvidia-cc-snp it is refused, with an error saying so.
openai_passthrough is refused even on 127.0.0.1: the node cannot see
whether that endpoint forwards requests onward, and to the node a local relay to
a hosted API looks the same as a local engine.
tee.mode="dstack-tdx" with tee.dataflow="sealed_local" requires an engine
inside the CVM: llm.provider=local (the supervised LocalLlamaProvider,
node/internal/llm/local.go) or a terminal runtime on a loopback base_url
(lmstudio, llamacpp, kronk), pinned by digest as a service in this same
measured compose. Got llm.provider="openai_passthrough" at
"http://127.0.0.1:8080/v1", and forwarding plaintext there would leak the
very prompts the TEE is meant to seal. […]
What the validator can and cannot check here. It checks the provider name
and that base_url is loopback. It cannot check that your engine is
digest-pinned or in the compose file; the payer verifies that against the
measured compose document your node publishes. A node pointing at an unpinned
engine boots and advertises tee_attested, but clients reject the compose, so
it gets no traffic from payers who require attested hardware. You are responsible for pinning the engine.
attested_passthrough
The node decrypts inside the CVM and forwards to exactly one named upstream. A payer is told:
Your prompt is readable by exactly one party — the named upstream — under their zero-retention terms. The node operator is not that party and cannot become it.
This guarantee is weaker than sealed_local, so you must declare it
explicitly. The default is sealed_local, so setting tee.mode while pointed
at a hosted API fails startup. The node does not fall back to the weaker claim
under the same badge.
There are two ways an upstream qualifies, and both are evidence the node collects itself:
- xAI, whose zero retention the node confirms from the
x-zero-data-retention: trueheader on every reply. There is no knob to disable that check. The node decides that an upstream is xAI from the URL's host, which must bex.aior a subdomain of it. A host such asapi.x.ai.example.com, or a URL withapi.x.aiin its path, is not xAI. - An upstream that attests its own enclave, verified by the node — see Verifying your upstream's enclave below.
Either way, llm.openai.base_url must be an https URL with a plain host
name and no user@ part, and so must image_llm.openai.base_url. Over plain
http, anyone who controls the network or DNS of the machine hosting your CVM
could answer as the upstream. Payers check the upstream your attestation
names against the same rule.
llm.upstream_zdr_declared does not qualify an upstream: it is an
assertion the node cannot check, and attesting it does not make it verifiable.
Supporting a third upstream requires building an enforcement mechanism for it;
adding its name to a list is not enough.
The image backend must meet the first of these requirements, confirmed zero
retention. The enclave check covers the text route only, and an image_llm on
the gateway's host is refused. image_llm receives the
prompt of every zs_image_generation and zs_image_edit call and of every
request to a dedicated image model, and you choose where it points, so the
validator holds it to the posture (the image backend row in the table above).
zs_image_search is one of the built-in web tools and stays a caller opt-in.
comfyui_cloud never runs on this machine, so it is refused under both
postures.
previous_response_id continuation does not work on an
attested_passthrough node. The node forces store: false even when a
client explicitly asks for store: true. You own the upstream account, so
upstream retention would otherwise let you read prompts back while every
attestation claim you make stays literally true.
Verifying your upstream's enclave
Some inference providers run inside a TEE of their own and publish evidence for
it. If yours speaks ACI/1 (Attested Confidential Inference v1 — Phala's
https://inference.phala.com/v1 is the one deployment today), your node can
check that evidence itself instead of taking the upstream's word:
tee:
mode: dstack-tdx
dataflow: attested_passthrough
upstream_attestation:
protocol: aci/1
require_upstream_providers: ["phala"]
allow_root_backdoor_env: ["DSTACK_ROOT_PUBLIC_KEY"]
refresh_interval: 15m
At startup the node fetches the gateway's hardware attestation with a nonce it chose, replays the boot log against the quote, checks the measured configuration, follows the certificate chain to Intel, and pins the TLS key the quote commits to so nothing can answer for that hostname afterwards. Then it checks every model you declared. If any of those checks fails, the node exits instead of starting with a claim it cannot back.
require_upstream_providers is the field that matters most and it has no safe
default, so the node refuses to start with it empty. An ACI gateway is a
router: it verifies each backend's enclave and serves several independent
operators from one address. Verifying the gateway establishes which code read
your payer's prompt first, not who ran the model. On 2026-09-24 that endpoint
fanned 272 live sessions out to four operators — 129 to Phala's own GPU CVMs,
126 to Chutes, 16 to NEAR AI, 1 to Tinfoil — so without the list, prompts you
sealed for one party reach three others. Every request also carries a pin that
the gateway refuses to serve unless it verified the backend first, and every
response carries a signed receipt naming the enclave that answered.
allow_root_backdoor_env waives part of the upstream check. dstack can be told
to install an operator-held root key inside a CVM, and Phala's gateway declares
that channel. Our own nodes are refused for it outright; an upstream may declare
it, and if you accept it here, the exact list is republished on your evidence as
upstream_attestation.carve_outs for a payer to read. Only real dstack channel
names are accepted, so the list cannot overstate what you waived.
It is also honoured only on a guest image whose filesystem has been enumerated
and found to contain no SSH daemon. With no SSH daemon, nothing in the image can
use the root key, which is why accepting it is defensible. Today that is the CPU
image, dstack-0.5.9. On any other image, including the GPU line's
dstack-nvidia-0.5.9, the waiver is withheld and the node refuses the upstream
at boot. If that happens the error names the image it saw. Adding entries to
this list does not fix it: the image has not been scanned, and no configuration
change can supply that.
Tell payers how many enclaves can read a prompt; don't say "nobody". With
this in place a prompt is readable inside three measured enclaves on three
codebases (your CVM, the gateway's, and the model's GPU CVM) and on no
operator's disk. That is stronger than any hosted API, but it is not the
sealed_local claim. Four things stay open: which model answered is not
measured; the inference engine is pinned by digest but nobody attests the build
was vetted; the GPU attestation proves a
genuine confidential GPU, not that it is the one bound to the serving CPU
enclave; and any root-shell channel you accepted weakens "the measurement pins
what runs" for that upstream.
For a stream, the pin is applied before any byte is sent, but the receipt exists only once the response is finished. A stream is therefore pinned before it starts and verified after it ends. A receipt that fails verification de-routes the upstream and fails the next request; it cannot withhold a response already delivered.
None of this changes the retention tier your models advertise today. It adds evidence of where the prompt went.
Why not OpenRouter?
Unless you have released it with llm.allow_upstream_retention (see
Releasing the pin), your node pins zero
data retention on every OpenRouter request and refuses to start if a model
you declared has no zero-retention endpoint. That pin takes effect earlier than
xAI's check, but it produces less evidence, and this posture promises evidence:
| Axis | OpenRouter pin | xAI header |
|---|---|---|
| Prevention timing | Applied before routing, so a retaining endpoint is never selected. | Read after the prompt has reached xAI. |
| Confirmation | Nothing comes back. A broker that ignored the pin looks identical to one that obeyed it. | Arrives with every response, so a lapse is caught on the next request. |
| The upstream's own retention | OpenRouter itself reads the prompt on the way through, and its prompt logging is a separate account setting that you own and no API exposes for reading. Turning it on would let you read every payer's prompt back. Nothing like the forced store: false closes it. | The node forces store: false. |
| What "zero data retention" means | Includes implicit KV caching to on-premises SSDs, unlike Google's reading of the same words. | — |
| Who reads the prompt | The posture would name https://openrouter.ai/api/v1, but the reader is whichever provider OpenRouter picked for that request, never published. The measured image fixes where the node sends a prompt, not who reads it. | xAI, the named upstream. |
Pinning a single sub-provider in your config answers only the last row.
What to do instead: run your OpenRouter upstream on a node with no
tee.mode. As long as you have not released the pin, it stays in force and
your models advertise
retention: upstream_enforced.
You lose the confidential-compute badge, which never applied to this setup.
The tee: config block
The whole opt-in lives under one top-level YAML block, alongside the zs: and
llm: sections (see Configuration):
tee:
# mode: which confidential-computing platform this node runs on.
# none - default; standard trust model, not a TEE node.
# stub - dev only; deterministic, well-formed-but-untrusted
# evidence. Exercises the whole pipeline on a laptop with
# no special hardware. Proxies reject stub evidence in
# production unless explicitly opted in.
# dstack-tdx - CPU-only Intel TDX CVM managed by dstack (Phala Cloud
# and self-hosted). The mode with a working verifier.
# nvidia-cc-tdx - Intel TDX host + NVIDIA GPU CC. (forthcoming)
# nvidia-cc-snp - AMD SEV-SNP host + NVIDIA GPU CC. (forthcoming)
mode: dstack-tdx
# dataflow: where plaintext comes to rest. sealed_local (default) runs the
# weights in the CVM and requires an engine there; attested_passthrough
# forwards to one named zero-retention upstream. See "Two postures" above —
# this is the gate, and forgetting it fails the boot rather than quietly
# weakening the claim.
dataflow: sealed_local
attestation:
# NVIDIA Remote Attestation Service endpoint. Required for the nvidia-cc-*
# modes only — dstack-tdx mints its quote from the local guest agent and
# never calls NRAS. Leave at the default unless you run a private
# attestation deployment.
nras_url: https://nras.attestation.nvidia.com
# How often the node mints fresh attestation evidence. /v1/zs/attestation
# serves the cached bundle and re-mints on this cadence. Evidence older than
# 2x this interval is served as 503 (stale), and the proxy demotes you out of
# the TEE-eligible set until fresh evidence is available.
refresh_interval: 1h
# Where to persist the latest evidence bundle so a brief restart doesn't drop
# you out of the verified set. MUST live on a CVM-encrypted volume — the
# bundle binds to your current ephemeral recipient and the host operator must
# not be able to read past evidence. Unset = mint fresh on every boot.
evidence_cache_path: /var/lib/zs-node/attestation/last.json
# Where this node reads Intel's revocation and TCB collateral from, to carry
# alongside its quote. Unset = Phala's CORS-enabled mirror. See "Collateral
# your node carries" below before changing it — pointing this at Intel's own
# PCS is the tempting wrong answer.
pccs_url: https://pccs.phala.network
# upstream_attestation: OMIT THIS unless your upstream speaks ACI/1. It is
# what turns attested_passthrough from "one upstream I trust" into "one
# upstream I check", and it requires llm.provider: openai_passthrough — a
# locally-hosted runtime publishes no attestation to appraise. See
# "Verifying your upstream's enclave" above for what each field does and
# what accepting a carve-out costs you.
#
# upstream_attestation:
# protocol: aci/1
# require_upstream_providers: ["phala"]
# allow_root_backdoor_env: ["DSTACK_ROOT_PUBLIC_KEY"]
# refresh_interval: 15m
None of the upstream_attestation fields has a NODE_TEE_* environment
override, for the same reason tee.dataflow has none: on dstack your config is
literal text inside the measured document, while environment values are
substituted after that document is hashed. An override would let a node publish
one carve-out list in the measurement and appraise against another. The field a
payer reads to see what you waived would then describe a policy your node never
applied.
Collateral your node carries
Verifying a quote also takes Intel's current revocation lists and TCB status for your platform. A verifier normally fetches those itself. Your node also fetches them at mint time and publishes them in the bundle, so a payer whose network cannot reach Intel can still verify you.
- It is never fatal. A PCCS outage costs the bundle its collateral and nothing else. Your node mints, boots, and serves normally, and a verifier with no carried set fetches its own.
- It is a last resort for the verifier. A verifier uses your copy only when it has neither live collateral nor a usable cached copy of its own, and never to overturn a refusal. Where the Intel collateral comes from explains why.
- Keep the default; don't point it at Intel's own PCS. A browser payer
cannot reach
api.trustedservices.intel.comat all, because it sends no CORS headers, so a node pointed there ships documents its clients could not have obtained independently. Changepccs_urlonly for a PCCS inside your own network.
Validation at startup
The config validator refuses to start when:
- The backends don't match
tee.dataflow. See Two postures for what each posture requires. The built-in tools are not part of this — they are a caller opt-in, and the same section explains why. tee.modeis not one ofnone,stub,dstack-tdx,nvidia-cc-tdx,nvidia-cc-snp. An unknowntee.dataflowis rejected too; neither is silently defaulted.nras_urlis missing under annvidia-cc-*mode.stubanddstack-tdxdon't need it.refresh_intervalis not positive (not checked understub). Keep the default1h. Ondstack-tdxthe node also re-mints on every ~20-minute key rotation, whatever this is set to (see Key rotation), so a shorter interval adds mints and little freshness. The interval also sets how old evidence may get before the node stops serving it (503 stale_attestation) and payers refuse it: 2× the interval, which payers cap at 80 minutes.llm.openai.debug_dump_errorsis true under any non-nonemode. It writes decrypted prompt content to stderr.
Environment variables
Most fields have a NODE_TEE_* override (env wins over YAML):
| Variable | Maps to |
|---|---|
NODE_TEE_MODE | tee.mode |
NODE_TEE_ATTESTATION_NRAS_URL | tee.attestation.nras_url |
NODE_TEE_ATTESTATION_REFRESH_INTERVAL | tee.attestation.refresh_interval (Go duration: 1h, 30m, 15m) |
NODE_TEE_ATTESTATION_EVIDENCE_CACHE_PATH | tee.attestation.evidence_cache_path |
NODE_TEE_ATTESTATION_PCCS_URL | tee.attestation.pccs_url |
tee.dataflow has no env override, on purpose. On dstack your config is
part of the measured compose file, while environment values are delivered
encrypted and separately, so they are not measured. An override would let a
node run attested_passthrough behind an image a verifier had confirmed as
sealed_local.
The tee: block holds no secrets, so the whole thing is safe to commit to YAML.
How the in-CVM key binds to your attestation
The node binds the in-CVM key to the attestation automatically. It reuses a standard node's ephemeral key handling (see Encryption & keys) and adds a hardware proof on top.
The encryption key is never published on chain. Your on-chain record anchors
only your identity, the Ed25519 signing key, and a signature binds the
ephemeral recipient to it. There is no placeholder pubkey to register, no
owner-signed pubkey rotation, and no updateOperator call. The flow:
-
Register normally. Operator and node registration take your owner and signing addresses, base URL, and stake, with no encryption key (see Registering on-chain). Bake the issued
operator_idinto the CVM config (zs.operator_idin YAML, orNODE_ZS_OPERATOR_IDin env). The attestation binds to this id, so the CVM must know it at boot. -
Boot the CVM with
tee.mode != none. Inside the enclave, the node generates a freshage(X25519) ephemeral recipient from the in-enclave RNG, signs it under your on-chain signing key, and advertises it on/v1/zs/details(ephemeral_age_pubkey+ephemeral_sig), as a standard node does. The private half lives only in CVM-encrypted RAM. -
The node mints attestation evidence, served from
GET /v1/zs/attestation. Itsreport_datafield is a hash of the current ephemeral recipient and youroperator_id, so the node re-mints on every rotation. The field holds the first 32 bytes of the hardware's 64-byteREPORT_DATA; the last 32 are the model binding. -
Payers verify the binding from your on-chain record, not from the bundle: the ephemeral must carry a valid signature from your registered signing address, and the hash recomputed from it must equal the signed report. The exact steps, and why the evidence names the ephemeral rather than your signing key, are in The key binding.
To sanity-check a deployment, read the live values from outside:
NODE_URL=https://your-node.example.com
curl -s "$NODE_URL/v1/zs/details" | jq '{ephemeral_age_pubkey, ephemeral_sig, tee}'
curl -s "$NODE_URL/v1/zs/attestation" | jq '{report_data, generated_at, app_models}'
No persistence required. The ephemeral private key is never written to
disk, not even to a CVM-encrypted volume, and is zeroed on each ~20-minute
rotation. A fresh key on every boot is normal; the node attests the new one. The
optional
evidence_cache_path persists only the attestation evidence bundle, never the
key, so a brief restart doesn't drop you from the verified set while fresh
evidence is minted.
Key rotation
The node rotates its in-CVM ephemeral on the same fixed cadence as a standard
node. The evidence binds the current ephemeral, so every rotation triggers a
re-mint, not only the refresh_interval cadence. The rotation hook requests the
re-mint, and the refresh loop independently checks that the cached bundle still
names the live key, so a rotation that slipped past the hook is caught rather
than left advertising evidence for a key the node no longer holds. A node that
has rotated but not yet re-minted is refused until the next probe picks up the
fresh bundle.
The model binding
The last 32 bytes of REPORT_DATA are a hash of your node id, your model
catalog (each model's id, source, digest and how the digest was produced) and,
since protocol 9.10, your node's posture: where plaintext ends up, and for a
named upstream, its URL, whether zero retention is enforced, and whether you
configured tee.upstream_attestation. Prices and tools are not included. The node publishes the hashed list as app_models and the
posture as posture on /v1/zs/attestation so anyone can recompute it. Confidential mode turns this on; there is nothing to configure. How a
verifier checks it is in
The model binding.
When the signed list changes
The node re-mints on key rotation (every 20 minutes), when it finishes hashing
model weights, and when its catalog changes, for example when an upstream goes
unhealthy. It checks the catalog every 15 seconds but re-mints for a change at
most once a minute, so /v1/zs/details and app_models can disagree for up to
about 75 seconds.
After a restart the node hashes model weights in the background, which can take
several minutes. Until it finishes, every model is listed as unverifiable. The
first evidence after boot lists no models; verifiers accept it.
If the node builds a model list that cannot be hashed, for example two entries
with the same id, it does not mint. It logs tee attestation refresh failed
with cannot measure the model set this node serves in err. This is a node
bug; report it with that log line. Meanwhile the node keeps serving its previous
evidence, which payers stop accepting at the next key rotation, within 20
minutes, because it binds a key the node no longer uses.
Freshness challenges
A caller can request /v1/zs/attestation?nonce=HEX with a random value of its
own, and the enclave mints a new quote over that value
(details). The reply is not
cached, and your cached evidence never carries a caller's nonce. Challenges are
rate-limited across the whole node; see
Configuration.
When the model binding fails
A node refused for one of the reasons below loses its attested verdict. It keeps serving ordinary traffic, is shown to payers as unattested with the reason attached, and drops out of routing for anyone who requires attested hardware. Your node mints evidence normally, so its log never shows why that traffic stopped. To see the reason yourself, open the card for a model your node serves in the chat app, page to your node, and hover the attestation badge.
| Reason | What it means | What to do |
|---|---|---|
aux_binding_absent | The last 32 bytes are all zero, which is what a node older than 9.9 mints. The chat app shows this as Node software out of date. | Upgrade zs-node and redeploy the CVM. |
aux_binding_mismatch | The signed hash does not match your node id, app_models, posture and nonce. | Look for something that rewrites responses between your node and the payer: a reverse proxy in front of your node, or the relay the payer fetched through. If the chat app shows your node as attested, the problem is on that payer's path. A node and a payer on opposite sides of protocol 9.10 also refuse each other this way, so upgrade whichever is older. |
happ_preimage_missing | app_models is absent or empty, but the quote does not commit to an empty catalog. | Check that nothing between your node and the payer rewrites the response body. A reverse proxy that drops unknown JSON fields will do this. |
happ_preimage_invalid | app_models cannot be hashed: a malformed entry, an empty or duplicate model id, an unknown weights_state, or a state that contradicts its digest. | Your node refuses to mint over such a list, so the likely cause is a rewritten response body. |
posture_preimage_invalid | posture cannot be hashed: it is not an object, or a field has the wrong type. | Your node never publishes one like that, so look for a rewritten response body. |
posture_upstream_unverifiable | The signed posture names an upstream the payer will not accept. The URL must be https with a plain host, and then either be xAI with zero-retention enforcement, or the node must be configured with tee.upstream_attestation and the response must carry its aci/1 upstream_attestation block. | A current node refuses to start unless one of those holds, so on an attested upstream gateway the block is usually missing. First check your node's logs for a failed upstream appraisal: the node drops the block while its last check of the upstream's enclave has failed or expired, and requests fail with upstream_attestation_required meanwhile. Otherwise check that nothing between your node and the payer strips the block from the response. |
A payer on a zs-proxy older than 9.9 refuses a 9.9 node with
key_binding_mismatch, because that release requires the last 32 bytes to be
zero. The payer has to upgrade; nothing on your node fixes it.
A wrong zs.node_id fails before the model binding is checked. Boot stops with
node X/Y not found on chain or keystore: signing address not loaded. If the
id belongs to a sibling node that shares this node's hot key, the node boots,
but its signed ephemeral key names the wrong node. Payers then reject it and
cannot seal to the node at all, confidential or not. They report
key_binding_mismatch.
Validate the whole pipeline with stub mode
Before you pay for CVM time, exercise the entire confidential-mode path on
your laptop with tee.mode: stub. Stub mode produces a deterministic,
well-formed but untrusted bundle, so every component on both sides runs the code
it runs in production: the node's config gate, the /v1/zs/details
advertisement, the /v1/zs/attestation handler, the proxy verifier, the routing
filter, and the dashboard badge. Only the hardware-rooted evidence is missing.
A stub bundle has no quote, so it has no model binding either.
Add the block to the same config.yaml a non-TEE deployment already uses:
tee:
mode: stub
attestation:
refresh_interval: 1m
evidence_cache_path: /tmp/zerosignal-tee-stub.json
The validator still requires an engine inside the CVM, so point llm.local at
your llama-server binary with one model configured (see
Serving models & pricing). stub also accepts the
loopback-sidecar shape, so you can rehearse a compose-file deployment
off-platform; stub evidence is untrusted anyway. Start the node and check what
it advertises on the public port (9090 by default):
NODE_URL=http://localhost:9090
curl -s "$NODE_URL/v1/zs/details" | jq .tee
# { "mode": "stub", "attested_at": "...", "evidence_url": "/v1/zs/attestation" }
curl -s "$NODE_URL/v1/zs/attestation" | jq .
# { "mode": "stub", "node_pubkey": "...", "operator_id": 42, "report_data": "...", "stub": true, ... }
Both should return 200. A 404 on /v1/zs/attestation means the node
didn't pick up tee.mode: stub; check the startup logs for a validator error.
Production proxies reject stub evidence. A real proxy marks a stub-mode
operator unverified and drops it only when a request requires TEE: a sensitive
route carrying the
X-Zs-Require-TEE
header, or a model on that proxy's zs.tee.require_for_models list. Plain
non-TEE traffic to the same operator still routes normally. To exercise the gate
end to end against your own dev proxy, send that header or opt the dev proxy
into accepting stub bundles. That opt-in is dev-only; never set it in
production.
Stub mode does not validate the real TDX/SEV-SNP signature chains, the
NVIDIA EAT / NRAS round-trip, the CVM memory-encryption guarantee, or the
encrypted-volume requirement for evidence_cache_path (the recipe above writes
it to /tmp).
Deployment recipes
One recipe works today: dstack (tee.mode: dstack-tdx), CPU-only. The GPU
paths below wait on the nvidia-cc-* modes.
Your zs-node binary already has the compose file. Run
zs-node tee compose, which writes the measured artifact for your release with
the published image digest resolved. Re-run it when you finish editing: with the
file already present it verifies instead of overwriting, and tells you if an
edit moved the measurement.
- Your configuration goes in the compose file, as literal text. dstack
hashes the compose before substituting
${VAR}references, so aNODE_CONFIG_YAMLblock written out in full is inside the measurement, while your API keys, which stay${VAR}references, are not. The block is also published: the node serves the whole document unauthenticated, because it is the preimage a payer hashes. Never inline a key there. - The compose file is a published artifact, not a template. Change the image
reference and your config block (and, on the GPU shape, the engine's config
block under
zs-engine-config:) and nothing else. A verifier lifts exactly those spans out and checks the rest, comments included, against the published release. Reformatting anything outside them fails silently: your node deploys, boots, and attests, and payers who require attested hardware route elsewhere.zs-proxy tee inspect <your-node-url> --chainshows you that verdict. The engine block is checked for content rather than matched byte for byte:adaptersandtemplateare refused, since either changes what the measured weights produce. - Inside your config block, edit freely: the whole block, not just the two
EDIT MEmarkers. The lift covers the entireNODE_CONFIG_YAMLbody, somax_active_tickets,ticket_ttl, your models and rates, and any setting the example does not mention are yours to change; redeploy and nobody has to be told. TheEDIT MEmarkers show what you must fill in (startup fails if you leave them); you may change other lines too. A few lines need care (the dataflow, the upstreambase_url, the drain timings), and a comment on each says so.
Always deploy with --no-dev-os. Without it Phala may install an SSH key as
an authorized root key inside the CVM. That undoes the enclave's protection, and
payers
detect and reject
it anyway.
Decide on Phala's --listed flag before you deploy. It is off by default,
and adding it later means building a new CVM. It publishes your deployment to
Phala's trust center, at https://trust.phala.com/app/<your-app-id>, where
Phala re-verifies the same quote and event log your node serves, independently
of us. The chat client links payers there from the attestation badge. Without
the flag your node attests just as well, but that link leads nowhere, and a
payer has only our verification instead of two independent ones.
Phala cannot change the flag in place, so a later phala deploy --cvm-id will
not turn it on. Turning it on means a fresh CVM with a new app id and gateway
hostname, and a baseUrl to re-register on chain. The node repo's
dstack deployment guide
has the full command and a one-line check for the state you are in. The related
--public-sysinfo / --public-tcbinfo / --public-logs flags default on, can
be changed later, and serve Phala's node-information page. Leaving them on costs
nothing and gives payers one more independent view.
Once stub validation passes, the real deployment swaps tee.mode: stub for the
platform string and removes any dev-only stub opt-in from the proxy. The paths
still to come:
- Phala Cloud / dstack with GPU confidential computing — the same CVM setup
as the shipped CPU recipe, on a TDX host with an NVIDIA H200. Closely modeled
on Phala's open
private-ml-sdk. - Azure NCC H100 v5 — TDX hosts with H100 GPUs in CC mode, via a confidential
VM (
--security-type ConfidentialVM). Best when you want managed Kubernetes or region-specific data residency. - Bare-metal SEV-SNP — EPYC Genoa/Milan with SEV-SNP in BIOS (use
tee.mode: nvidia-cc-snp) and a CVM-aware hypervisor. - GCP H100 confidential — TDX-based; identical config to the Azure recipe.
Pitfalls, and which paths each applies to:
- Pin every image by digest, never by tag (every path). The compose / VM image hash is what the CPU report measures. A floating tag produces a different measurement on every pull, and verifiers refuse to seal to a measurement that keeps changing.
- Keep
evidence_cache_pathon a CVM-encrypted volume (every path). A plain host bind-mount is readable and writable by the host operator, who must not be able to inspect or tamper with past evidence. - Allow outbound 443 to NRAS (
nras.attestation.nvidia.com) from the CVM (nvidia-cc-*only), or GPU-EAT verification fails. - Pin the driver and
nvtrustversions together inside the CVM image (nvidia-cc-*only). The combination is part of the attested measurement.
When the GPU modes land, only the mode string you set changes; the tee: config
and the verification steps below stay the same.
Verifying a deployment from outside the CVM
Run these checks from a workstation that is not the CVM host, to see what a third party sees.
-
Capability advertisement.
GET /v1/zs/details→ theteefield should report your configuredmodeand anattested_atwithin the last2 × refresh_interval. Anything older means the refresh loop is wedged. -
Evidence is fetchable.
GET /v1/zs/attestation→200with a well-formed bundle.503means stale or cold-start;404means the node isn't in TEE mode. -
The advertised ephemeral is signed. The
ephemeral_sigon/v1/zs/detailsverifies under your on-chain signing address. -
report_databinding is consistent. It matches the ephemeral currently advertised, computed as in The key binding. A mismatch right after a rotation is normal for one probe; a persistent one means the rotation hook is not re-minting.That field is the first 32 bytes. The last 32 are the model binding, which hashes the bundle's
app_modelsandposture. Compare the list against what you advertise:curl -s "$NODE_URL/v1/zs/attestation" | jq -r '(.app_models // [])[].model_id' | sortcurl -s "$NODE_URL/v1/zs/details" | jq -r '.models[].id' | sortAfter a catalog change or a restart the two can differ for about 75 seconds. A difference lasting longer means the evidence is not being re-minted; check
attested_atas in step 1. -
Run the verifier against yourself. The steps above show only that your node is internally consistent, which is true of a node nobody routes to. This one runs the hardware, image and compose checks a payer runs. It does not check the two
REPORT_DATAbindings, which need the node's on-chain record (its node id and signing address). For those, open the card for a model your node serves in the chat app, page to your node, and hover the attestation badge.zs-proxy tee inspect https://your-node.example.com --chainAlways pass
--chain. Without it the quote's signature is never examined, so the report only checks values your node supplied against each other. Read three lines:- quote chain —
PASS, and which collateral it used. Plain "verified against Intel's PCK chain" means live collateral. It can also say the pass used a stale cache or the collateral your node supplied. Both are genuine passes, and both mean the checking party could not reach Intel just then. - compose-hash gate —
PASSmeans the document you published is the one your CVM measured. - violations and release match —
ENFORCEDwithviolations: noneis the routable state. A skeleton or image that is not on the lists fails silently: your node deploys, boots, attests, and payers who require attested hardware route elsewhere, with nothing in your logs.
- quote chain —
-
An end-to-end Require-TEE request lands. A request carrying
X-Zs-Require-TEEis routed to your node and returns a normal completion.
Your release has to be on the lists; you don't add it yourself.
The two values tee inspect prints, the compose skeleton digest and the image
reference, ship compiled into zs-proxy and the browser client, and every
zs-node release adds its own through an automated pull request into both.
The only requirement is to run a published release unmodified. If you
build your own image, no payer counts it as attested, however valid the quote.
Monitoring
The standard monitoring on the private listener (/healthz, /livez, /metrics on
9091) still applies. Confidential mode adds two things to watch:
| Endpoint | What to watch |
|---|---|
GET /v1/zs/attestation | Should respond 200 with generated_at newer than 2 × refresh_interval. A 503 (stale) means the refresh loop is failing — check node logs for the underlying mint error (on dstack-tdx, the call to the local dstack guest agent; on nvidia-cc-*, NRAS / nvtrust). |
GET /v1/zs/details (tee field) | Should report your configured mode and an attested_at matching the most recent successful refresh. |
Make sure /v1/zs/attestation is reachable on your public HTTPS URL; if a TLS
reverse proxy fronts the node, expose it there. Proxies fetch it during their
per-operator details poll, and the verifier needs it to keep you in the
TEE-eligible set.
Where this fits
- Verifying an attestation yourself covers what the attestation establishes: every value that gets measured, what each one proves, how a payer confirms it, and what it leaves open. This page covers what you do to pass those checks.
- Encryption & keys describes the standard trust model this mode upgrades.
- The
sealed_localengine rule ties into provider selection in Serving models & pricing; thetee:block sits alongside the rest of your Configuration. - Confidential nodes still relay and are still relayed for, unchanged (see Relays). Relaying never touches plaintext on either side.