Model identity & integrity
Two different questions travel with every model you advertise:
- Identity — what model is this? Answered by the
sourcepointer, which the always-on identity gate requires before a model is advertised at all. - Integrity — are you and your competitors serving the same bytes under
that name? Answered by an optional
weights_digest.
The first is a hard gate: get it wrong and the model vanishes from your catalog. The second is advisory, except that a contested digest demotes you in routing (see What a contradicted digest costs you).
Identity: the provenance gate
A model is advertised (and therefore reservable) only if it carries a checkable identity. Four ways to satisfy that, in the order the node resolves them:
- You declare a
source— a scheme-tagged pointer to the public artifact the model is:source: "hf:org/model", or"hf:org/model@revision". Your declaration always wins. - The id is itself the repo. If the wire id is an
org/modelpair (Qwen/Qwen2.5-7B-Instruct), the node deriveshf:<id>at load and logs it, so you don't restate the id as a source. - The backend supplies it. Kronk
knows which artifact each GGUF came from and hands the node
hf:<org>/<repo>@<rev>. This fills a gap; it never overrides asource:you declared, and it outranks the syntactic guess in (2), since a slash in a wire id isn't always an org/repo boundary. - The id is on the frontier whitelist of vendor-hosted models (
gpt-4o,claude-opus-4-5,gemini-2.5-pro,grok-4.5,glm-5.2,kimi-k2.6, …). The whitelist entry is their provenance, so declare nosource. Provider-qualified spellings of a whitelisted model clear the same gate when the prefix is a known routing vendor (x-ai/grok-4.5,google/gemini-2.5-pro,z-ai/glm-4.6).
zs:
models:
qwen2.5-7b-instruct: # an open-weights id you renamed
source: "hf:Qwen/Qwen2.5-7B-Instruct" # required — links it to the repo
pricing: { input_rate: 0.1, output_rate: 0.3 }
Qwen/Qwen2.5-7B-Instruct: # id IS the repo → derived for you
pricing: { input_rate: 0.1, output_rate: 0.3 }
gpt-4o-mini: # frontier-whitelisted → no source
pricing: { input_rate: 0.15, output_rate: 0.60 }
A model that satisfies none of these is dropped, and the node still starts.
It disappears from /v1/models and /v1/zs/details and becomes unreservable.
The node logs a WARN once at startup naming the dropped ids, and
zs-node doctor flags them in its provenance check. There is no knob to turn
the gate off.
The case that catches people is default_pricing: it advertises every id your
upstream discovers, but a discovered id has no zs.models entry and so can
declare no source. Under the gate, every non-frontier discovered model is
dropped. Because it isn't in your config, the startup WARN can't name it; you
get a one-shot runtime WARN the first time it would have been advertised
instead. To expose a fleet of open-weights models, give each one a zs.models
entry with a source:.
- A malformed
source("not-a-real-ref") is a hard startup error, independent of the gate. Fix it or remove it. - Image models are exempt. A generation workflow maps to no single
HuggingFace repo, so a model with an
image:block never needs asource. See Image generation. - The derivation is syntactic, not verified. The gate is offline; it never
confirms the repo exists. A derived
hf:ref is operator-attested exactly like one you typed. If anorg/model-shaped id isn't really a public repo, fix the wire id or declare the realsourcerather than leaning on the guess. - Whitelisted open-weights models get their repo attached for you. For the
GLM and Kimi families the node adds the real repo (
glm-5.2→hf:zai-org/GLM-5.2, logged at startup) so your hosted-API spelling clusters with operators serving the same weights self-hosted. Leave it to the node: declaring it by hand pins today's mapping and opts you out of later corrections.
What source buys you beyond the gate
Declaring a source also lets the node enrich the model from HuggingFace for
clients to display. This is on by default
(zs.coordinates.discover_huggingface: true) and fetches each served model's
repo once near startup:
- Tags — the model author's curated card tags (
medical,code) are exposed as the model'stags, unioned with anycontext.tagsyou declare, so you don't hand-maintain them. The union can only add labels. The HF-discovered ones are not part of the signed policy fingerprint, so your advertisedtagsare a superset of the hash-covered ones. Content gating still works, becausensfwis declared incontext.tags. - Coordinates — params, family, quantization, surfaced on the lazy per-model
drill-in (
?expand=coordinates).
Set discover_huggingface: false only for a node that must make no outbound
calls to huggingface.co. With it on, the fetch reveals your served set to
HuggingFace, a set that is already public on /v1/zs/details. To confirm it's
working, watch zs_hf_coordinates_total{result} (a hit rate means it's
resolving; error means HF is unreachable or blocked), or check that a model's
tags include its HuggingFace labels.
Integrity: the weights digest
The node can hash the weight bytes it serves and advertise the result as
weights_digest on /v1/zs/details. Two operators claiming the same model id
can then be compared: same digest, same bytes.
zs:
models:
qwen2.5-7b-instruct:
source: "hf:Qwen/Qwen2.5-7B-Instruct"
weights:
files: ["Qwen2.5-7B-Instruct-Q4_K_M.gguf"] # resolved against root, or CWD
# root: "/models/qwen" # optional
# digest: "sha256:..." # optional precomputed pin
Read this before treating a digest as proof of anything. On a standard
node the digest is self-reported by the operator's own node, in every
configuration below. Nothing stops an operator from pasting a well-known model's
genuine public checksum into weights.digest while serving something else:
the node performs zero verification of a pin beyond checking that the string is
shaped like sha256:<64 hex>. Cross-operator comparison doesn't catch that
either. It flags disagreement, and a liar who copies the correct hash produces
none.
Treat weights_digest as a tool for catching accidental drift (a stale
file, a wrong quant, a sloppy mismatch), not as evidence against a deliberately
dishonest operator. The exception is a
confidential node: there the engine runs inside the
hardware measurement, so the digest it reports for the artifact it loaded is
evidence rather than a claim, and a match against the source repository's
published digest is a binding the operator cannot forge. That is what
runtime_attested plus matched means on an attested node, with one disclosed
gap: the engine's own downloaded runtime library sits outside the measurement.
On a confidential node, the hardware-signed REPORT_DATA also covers the whole
advertised model list
(details).
You usually don't need to configure anything. On provider: "local" the
node already knows the GGUF path and hashes it. On provider: "kronk" it asks
Kronk which artifact it loaded and takes the digest Kronk attests for it.
Declare weights.files yourself only for lmstudio / llamacpp /
openai_passthrough / vertexai, or for a Kronk whose management API you've
locked down.
Three states, all advisory:
| State | Meaning |
|---|---|
weights_digest: "sha256:<hex>" | The node hashed local files, the runtime attested the digest, or you supplied a pin. |
weights_unverifiable: true | No usable digest — a remote provider, an image model, a hashing failure, or a runtime that won't vouch for its own artifact. The reason appears only on the deep ?expand=digest&model=<id> path. |
| Both absent | No weights: block and no provider fallback applies, or the background hash hasn't published its first snapshot yet. |
One file (a GGUF) hashes to that file's raw sha256, byte-identical to
sha256sum model.gguf and to HuggingFace's LFS oid. Several files
(safetensors shards plus tokenizer and config) hash to a fold of their sorted
per-file content digests; paths and sizes don't enter it. It runs once at
startup, in the background, so a multi-gigabyte hash never blocks boot or adds
request latency. A file swapped in after startup isn't reflected until you
restart.
How much to trust a digest: weights_digest_trust
A digest string looks the same however it was produced. The deep
?expand=digest&model=<id> path reports which of four tiers produced it. The
tier is self-reported, not verified. The tiers are ranked, strongest first:
| Tier | How it's produced | What it tells you |
|---|---|---|
node_supervised | provider: "local" with no weights: block — the node execs llama-server --model <path> itself and hashes that exact path. | The strongest. No separate declared file list can diverge undetected from what's loaded, because the node built the launch command. Still not proof: host access can swap the file after the startup hash. |
runtime_attested | provider: "kronk" with no weights: block — the node asks Kronk which artifact it loaded and takes the digest Kronk verified against it. Automatic. | The node hashed nothing itself and is relaying a claim, so it ranks below node_supervised. It ranks above the declared tiers because the claim comes from the process actually serving your requests. Kronk doesn't re-hash per request, so this means "the bytes matched when last checked". |
operator_declared | A weights.files path you supplied. | The node hashes real bytes at that path, but has no supervisory tie to whatever process is answering requests. The declared path and the serving engine can diverge without the node knowing. |
operator_pinned | A weights.digest pin. | No hashing happens at all. A bare assertion, no more trustworthy than source. |
To reach a higher tier, declare less. Run local or kronk without a
weights: block and you get the strongest tier available to your setup for
free. If you must declare one because you front an
external engine (vLLM, LM Studio, a llama-server outside the node's
supervision), you're asking clients to trust your word a little more
(operator_declared) or a lot more (operator_pinned) than they'd need to
otherwise.
Kronk specifics. The integrity API shares the admin grant that gates
/v1/kronk/*, so under any authorization mode except open the node can't read
it and those models fall back to "no local weights" (or to whatever weights:
you declare); one INFO line at startup names the failure. An artifact Kronk
hasn't verified (or has marked stale after the file changed) also comes back
without a digest by design; declare weights.files yourself if you need one.
Multi-part (split) GGUFs are covered. Every model above HuggingFace's 50 GB per-file cap ships as several parts, and Kronk reports a content digest for each one. The node combines those into a single composite and advertises it as "runtime attested", so a frontier-scale model gets the same tier a small one does.
It is all parts or none: if Kronk hasn't verified every part, the model reports "runtime unverified" rather than a digest over the verified subset, which would describe weights nobody is serving.
Declaring weights.files for those same parts agrees with the runtime, since
both fold the artifacts' content digests with no filename, path or size
involved. A declaration matching what the runtime loaded is corroborated by it
and promotes to "runtime attested". You don't need to declare anything; if you
do, it's a free upgrade.
If you declare a multi-file weights.files set and upgrade from a release
whose fold included each file's path and size, the digest you advertise
changes. Until peers serving the same model upgrade too, you and they report
different digests for identical weights. Single-file models are unaffected;
their digest is still plain sha256sum.
Does the source repo publish it: weights_registry_match
Whenever the node holds a digest, however produced, it also asks the model's
source repo whether it publishes that exact content digest, and reports the
answer on the same deep path:
| Verdict | Meaning |
|---|---|
matched | The repo publishes that digest. |
mismatch | The repo's artifact list was read in full and does not contain it. |
unchecked | No conclusive lookup — no source, discover_huggingface: false, a gated or missing repo, a network failure, or a listing too large to rule it out. |
For a multi-artifact model (a split GGUF, or a declared file set) the
question is asked about each artifact rather than the folded digest, because no
repository publishes a fold. matched there means the set is one model the
repo publishes: every artifact is present, and for each one the repo names as
a shard (…-00003-of-00009.gguf) your set holds that whole shard group, not
parts from two quantizations and not a subset of one. A repo like
DeepSeek-V3-GGUF carries seven quantizations side by side, so "every file is
in the repo" alone would not prove much. Files the repo does not name as shards
(a projection alongside the weights) are judged by presence alone, and a set may
hold more than one whole group, as a sharded pipeline does. The path field is
left empty: no single page prints the value being advertised.
Read it together with the trust tier, never alone. node_supervised +
matched is your node hashing the file it launched and an independent registry
agreeing, the strongest evidence available short of hardware attestation.
operator_pinned + matched corroborates nothing, because the pin could have
been copied off the very page being consulted.
The comparison is by content, not filename, because runtimes rename
artifacts on download. A real Kronk install serves HuggingFace's
Qwen3.6-35B-A3B-Q8_0.gguf as mtp-Qwen3.6-35B-A3B-Q8_0.gguf, and stores that
repo's mmproj-F16.gguf under the weights file's own stem. Both match fine.
A mismatch only logs a WARN. The digest is still advertised, still
compared across operators, and admission, pricing and settlement are untouched.
Innocent explanations are common: you serve a locally requantized or repacked
copy, or the repo re-uploaded the file after you pulled it. A payer's router never reads this verdict at all. If
neither explanation fits, check that source: names the repo you're actually
serving.
When your own runtime contradicts you
If you declare a digest and the serving runtime independently reports
loading a different artifact, the node advertises the runtime's digest,
reports runtime_attested, sets weights_declaration_conflict: true, and logs a
WARN naming both values. A copied pin can pass the registry check, but the
runtime's digest still replaces it.
The common cause is innocent: pinning the digest of a model's original repo
while serving a GGUF quantization of it produces a conflict, as does declaring a
file set that is a sibling of what the backend loaded. To clear the warning in
that case, drop the declaration; a Kronk-backed model needs no
weights: block and reports runtime_attested on its own.
It never triggers on a runtime digest the runtime hasn't verified (that is not
evidence, so your declaration stands), or on provider: "local", where the node
hashes the path it launched and there is no second opinion.
On a confidential node, the model list is in the attestation
On a confidential node, the catalog on
/v1/zs/details is also hashed into the CPU-signed REPORT_DATA. That covers
every model id, its source, its digest and how the digest was produced. A
verifier recomputes the hash from the list the node publishes alongside its
evidence and refuses the node when the two disagree. See
The model binding.
A confidential node has two integrity signals; a standard node has only the first.
| Signal | Who produces it | What it decides |
|---|---|---|
| Cross-operator agreement | Your peers, by advertising a digest for the same model id. | Placement only. A contradicted digest moves you to the end of the candidate list. |
| The model binding | The CPU, over bytes chosen inside the measured VM. | Pass or fail. A node whose published catalog does not reproduce the signed value loses its attested verdict. |
Consensus applies only when another operator serves the same model id. The model
binding applies to every model but shows only that your node reported these
digests. Checking the weights against the repo is
weights_registry_match,
which is advisory. The signature records one of
three coarse states, not the
trust tier. The tier on
/v1/zs/details is unsigned.
What a contradicted digest costs you
The rule: when two or more distinct operator owners advertise the same digest for a model, an operator advertising a different digest for that model drops to the tail of the payer's candidate list for that request. That operator stays routable as a last resort, but every agreeing peer is tried first.
What costs you nothing:
- Publishing no digest. Every hosted-passthrough model is permanently in this state, since there is no local file to hash for grok, gpt or claude.
- A
mismatchoruncheckedregistry verdict on its own. Payer routers don't read those fields. - Your self-reported trust tier.
node_superviseddoes not move you up. Nothing you say about your own digest can raise you; only a peer's digest can lower you. - Being outnumbered by one operator's fleet. The count is over distinct owners, so a competitor spinning up nodes cannot vote you down.
- A two-way split. If the operators serving a model divide evenly, nobody is demoted.
The honest way to land here is to serve a requantized or repacked copy under
the same model id everyone else uses. Your bytes really do differ, and the router
can't tell your Q4 repack from a substitution. Say so: point source: at the
quantization's own repo, or advertise it under its own model id. Both are more
accurate anyway, and both move you into your own identity group, where the
comparison no longer applies.