Skip to main content

Model identity & integrity

Two different questions travel with every model you advertise:

  • Identity — what model is this? Answered by the source pointer, which the always-on identity gate requires before a model is advertised at all.
  • Integrity — are you and your competitors serving the same bytes under that name? Answered by an optional weights_digest.

The first is a hard gate: get it wrong and the model vanishes from your catalog. The second is advisory, except that a contested digest demotes you in routing (see What a contradicted digest costs you).

Identity: the provenance gate​

A model is advertised (and therefore reservable) only if it carries a checkable identity. Four ways to satisfy that, in the order the node resolves them:

  1. You declare a source — a scheme-tagged pointer to the public artifact the model is: source: "hf:org/model", or "hf:org/model@revision". Your declaration always wins.
  2. The id is itself the repo. If the wire id is an org/model pair (Qwen/Qwen2.5-7B-Instruct), the node derives hf:<id> at load and logs it, so you don't restate the id as a source.
  3. The backend supplies it. Kronk knows which artifact each GGUF came from and hands the node hf:<org>/<repo>@<rev>. This fills a gap; it never overrides a source: you declared, and it outranks the syntactic guess in (2), since a slash in a wire id isn't always an org/repo boundary.
  4. The id is on the frontier whitelist of vendor-hosted models (gpt-4o, claude-opus-4-5, gemini-2.5-pro, grok-4.5, glm-5.2, kimi-k2.6, …). The whitelist entry is their provenance, so declare no source. Provider-qualified spellings of a whitelisted model clear the same gate when the prefix is a known routing vendor (x-ai/grok-4.5, google/gemini-2.5-pro, z-ai/glm-4.6).
zs:
models:
qwen2.5-7b-instruct: # an open-weights id you renamed
source: "hf:Qwen/Qwen2.5-7B-Instruct" # required — links it to the repo
pricing: { input_rate: 0.1, output_rate: 0.3 }
Qwen/Qwen2.5-7B-Instruct: # id IS the repo → derived for you
pricing: { input_rate: 0.1, output_rate: 0.3 }
gpt-4o-mini: # frontier-whitelisted → no source
pricing: { input_rate: 0.15, output_rate: 0.60 }
warning

A model that satisfies none of these is dropped, and the node still starts. It disappears from /v1/models and /v1/zs/details and becomes unreservable. The node logs a WARN once at startup naming the dropped ids, and zs-node doctor flags them in its provenance check. There is no knob to turn the gate off.

The case that catches people is default_pricing: it advertises every id your upstream discovers, but a discovered id has no zs.models entry and so can declare no source. Under the gate, every non-frontier discovered model is dropped. Because it isn't in your config, the startup WARN can't name it; you get a one-shot runtime WARN the first time it would have been advertised instead. To expose a fleet of open-weights models, give each one a zs.models entry with a source:.

  • A malformed source ("not-a-real-ref") is a hard startup error, independent of the gate. Fix it or remove it.
  • Image models are exempt. A generation workflow maps to no single HuggingFace repo, so a model with an image: block never needs a source. See Image generation.
  • The derivation is syntactic, not verified. The gate is offline; it never confirms the repo exists. A derived hf: ref is operator-attested exactly like one you typed. If an org/model-shaped id isn't really a public repo, fix the wire id or declare the real source rather than leaning on the guess.
  • Whitelisted open-weights models get their repo attached for you. For the GLM and Kimi families the node adds the real repo (glm-5.2 → hf:zai-org/GLM-5.2, logged at startup) so your hosted-API spelling clusters with operators serving the same weights self-hosted. Leave it to the node: declaring it by hand pins today's mapping and opts you out of later corrections.

What source buys you beyond the gate​

Declaring a source also lets the node enrich the model from HuggingFace for clients to display. This is on by default (zs.coordinates.discover_huggingface: true) and fetches each served model's repo once near startup:

  • Tags — the model author's curated card tags (medical, code) are exposed as the model's tags, unioned with any context.tags you declare, so you don't hand-maintain them. The union can only add labels. The HF-discovered ones are not part of the signed policy fingerprint, so your advertised tags are a superset of the hash-covered ones. Content gating still works, because nsfw is declared in context.tags.
  • Coordinates — params, family, quantization, surfaced on the lazy per-model drill-in (?expand=coordinates).

Set discover_huggingface: false only for a node that must make no outbound calls to huggingface.co. With it on, the fetch reveals your served set to HuggingFace, a set that is already public on /v1/zs/details. To confirm it's working, watch zs_hf_coordinates_total{result} (a hit rate means it's resolving; error means HF is unreachable or blocked), or check that a model's tags include its HuggingFace labels.

Integrity: the weights digest​

The node can hash the weight bytes it serves and advertise the result as weights_digest on /v1/zs/details. Two operators claiming the same model id can then be compared: same digest, same bytes.

zs:
models:
qwen2.5-7b-instruct:
source: "hf:Qwen/Qwen2.5-7B-Instruct"
weights:
files: ["Qwen2.5-7B-Instruct-Q4_K_M.gguf"] # resolved against root, or CWD
# root: "/models/qwen" # optional
# digest: "sha256:..." # optional precomputed pin
warning

Read this before treating a digest as proof of anything. On a standard node the digest is self-reported by the operator's own node, in every configuration below. Nothing stops an operator from pasting a well-known model's genuine public checksum into weights.digest while serving something else: the node performs zero verification of a pin beyond checking that the string is shaped like sha256:<64 hex>. Cross-operator comparison doesn't catch that either. It flags disagreement, and a liar who copies the correct hash produces none.

Treat weights_digest as a tool for catching accidental drift (a stale file, a wrong quant, a sloppy mismatch), not as evidence against a deliberately dishonest operator. The exception is a confidential node: there the engine runs inside the hardware measurement, so the digest it reports for the artifact it loaded is evidence rather than a claim, and a match against the source repository's published digest is a binding the operator cannot forge. That is what runtime_attested plus matched means on an attested node, with one disclosed gap: the engine's own downloaded runtime library sits outside the measurement. On a confidential node, the hardware-signed REPORT_DATA also covers the whole advertised model list (details).

You usually don't need to configure anything. On provider: "local" the node already knows the GGUF path and hashes it. On provider: "kronk" it asks Kronk which artifact it loaded and takes the digest Kronk attests for it. Declare weights.files yourself only for lmstudio / llamacpp / openai_passthrough / vertexai, or for a Kronk whose management API you've locked down.

Three states, all advisory:

StateMeaning
weights_digest: "sha256:<hex>"The node hashed local files, the runtime attested the digest, or you supplied a pin.
weights_unverifiable: trueNo usable digest — a remote provider, an image model, a hashing failure, or a runtime that won't vouch for its own artifact. The reason appears only on the deep ?expand=digest&model=<id> path.
Both absentNo weights: block and no provider fallback applies, or the background hash hasn't published its first snapshot yet.

One file (a GGUF) hashes to that file's raw sha256, byte-identical to sha256sum model.gguf and to HuggingFace's LFS oid. Several files (safetensors shards plus tokenizer and config) hash to a fold of their sorted per-file content digests; paths and sizes don't enter it. It runs once at startup, in the background, so a multi-gigabyte hash never blocks boot or adds request latency. A file swapped in after startup isn't reflected until you restart.

How much to trust a digest: weights_digest_trust​

A digest string looks the same however it was produced. The deep ?expand=digest&model=<id> path reports which of four tiers produced it. The tier is self-reported, not verified. The tiers are ranked, strongest first:

TierHow it's producedWhat it tells you
node_supervisedprovider: "local" with no weights: block — the node execs llama-server --model <path> itself and hashes that exact path.The strongest. No separate declared file list can diverge undetected from what's loaded, because the node built the launch command. Still not proof: host access can swap the file after the startup hash.
runtime_attestedprovider: "kronk" with no weights: block — the node asks Kronk which artifact it loaded and takes the digest Kronk verified against it. Automatic.The node hashed nothing itself and is relaying a claim, so it ranks below node_supervised. It ranks above the declared tiers because the claim comes from the process actually serving your requests. Kronk doesn't re-hash per request, so this means "the bytes matched when last checked".
operator_declaredA weights.files path you supplied.The node hashes real bytes at that path, but has no supervisory tie to whatever process is answering requests. The declared path and the serving engine can diverge without the node knowing.
operator_pinnedA weights.digest pin.No hashing happens at all. A bare assertion, no more trustworthy than source.

To reach a higher tier, declare less. Run local or kronk without a weights: block and you get the strongest tier available to your setup for free. If you must declare one because you front an external engine (vLLM, LM Studio, a llama-server outside the node's supervision), you're asking clients to trust your word a little more (operator_declared) or a lot more (operator_pinned) than they'd need to otherwise.

info

Kronk specifics. The integrity API shares the admin grant that gates /v1/kronk/*, so under any authorization mode except open the node can't read it and those models fall back to "no local weights" (or to whatever weights: you declare); one INFO line at startup names the failure. An artifact Kronk hasn't verified (or has marked stale after the file changed) also comes back without a digest by design; declare weights.files yourself if you need one.

tip

Multi-part (split) GGUFs are covered. Every model above HuggingFace's 50 GB per-file cap ships as several parts, and Kronk reports a content digest for each one. The node combines those into a single composite and advertises it as "runtime attested", so a frontier-scale model gets the same tier a small one does.

It is all parts or none: if Kronk hasn't verified every part, the model reports "runtime unverified" rather than a digest over the verified subset, which would describe weights nobody is serving.

Declaring weights.files for those same parts agrees with the runtime, since both fold the artifacts' content digests with no filename, path or size involved. A declaration matching what the runtime loaded is corroborated by it and promotes to "runtime attested". You don't need to declare anything; if you do, it's a free upgrade.

Upgrading changes multi-file digests

If you declare a multi-file weights.files set and upgrade from a release whose fold included each file's path and size, the digest you advertise changes. Until peers serving the same model upgrade too, you and they report different digests for identical weights. Single-file models are unaffected; their digest is still plain sha256sum.

Does the source repo publish it: weights_registry_match​

Whenever the node holds a digest, however produced, it also asks the model's source repo whether it publishes that exact content digest, and reports the answer on the same deep path:

VerdictMeaning
matchedThe repo publishes that digest.
mismatchThe repo's artifact list was read in full and does not contain it.
uncheckedNo conclusive lookup — no source, discover_huggingface: false, a gated or missing repo, a network failure, or a listing too large to rule it out.

For a multi-artifact model (a split GGUF, or a declared file set) the question is asked about each artifact rather than the folded digest, because no repository publishes a fold. matched there means the set is one model the repo publishes: every artifact is present, and for each one the repo names as a shard (…-00003-of-00009.gguf) your set holds that whole shard group, not parts from two quantizations and not a subset of one. A repo like DeepSeek-V3-GGUF carries seven quantizations side by side, so "every file is in the repo" alone would not prove much. Files the repo does not name as shards (a projection alongside the weights) are judged by presence alone, and a set may hold more than one whole group, as a sharded pipeline does. The path field is left empty: no single page prints the value being advertised.

Read it together with the trust tier, never alone. node_supervised + matched is your node hashing the file it launched and an independent registry agreeing, the strongest evidence available short of hardware attestation. operator_pinned + matched corroborates nothing, because the pin could have been copied off the very page being consulted.

The comparison is by content, not filename, because runtimes rename artifacts on download. A real Kronk install serves HuggingFace's Qwen3.6-35B-A3B-Q8_0.gguf as mtp-Qwen3.6-35B-A3B-Q8_0.gguf, and stores that repo's mmproj-F16.gguf under the weights file's own stem. Both match fine.

info

A mismatch only logs a WARN. The digest is still advertised, still compared across operators, and admission, pricing and settlement are untouched. Innocent explanations are common: you serve a locally requantized or repacked copy, or the repo re-uploaded the file after you pulled it. A payer's router never reads this verdict at all. If neither explanation fits, check that source: names the repo you're actually serving.

When your own runtime contradicts you​

If you declare a digest and the serving runtime independently reports loading a different artifact, the node advertises the runtime's digest, reports runtime_attested, sets weights_declaration_conflict: true, and logs a WARN naming both values. A copied pin can pass the registry check, but the runtime's digest still replaces it.

The common cause is innocent: pinning the digest of a model's original repo while serving a GGUF quantization of it produces a conflict, as does declaring a file set that is a sibling of what the backend loaded. To clear the warning in that case, drop the declaration; a Kronk-backed model needs no weights: block and reports runtime_attested on its own.

It never triggers on a runtime digest the runtime hasn't verified (that is not evidence, so your declaration stands), or on provider: "local", where the node hashes the path it launched and there is no second opinion.

On a confidential node, the model list is in the attestation​

On a confidential node, the catalog on /v1/zs/details is also hashed into the CPU-signed REPORT_DATA. That covers every model id, its source, its digest and how the digest was produced. A verifier recomputes the hash from the list the node publishes alongside its evidence and refuses the node when the two disagree. See The model binding.

A confidential node has two integrity signals; a standard node has only the first.

SignalWho produces itWhat it decides
Cross-operator agreementYour peers, by advertising a digest for the same model id.Placement only. A contradicted digest moves you to the end of the candidate list.
The model bindingThe CPU, over bytes chosen inside the measured VM.Pass or fail. A node whose published catalog does not reproduce the signed value loses its attested verdict.

Consensus applies only when another operator serves the same model id. The model binding applies to every model but shows only that your node reported these digests. Checking the weights against the repo is weights_registry_match, which is advisory. The signature records one of three coarse states, not the trust tier. The tier on /v1/zs/details is unsigned.

What a contradicted digest costs you​

The rule: when two or more distinct operator owners advertise the same digest for a model, an operator advertising a different digest for that model drops to the tail of the payer's candidate list for that request. That operator stays routable as a last resort, but every agreeing peer is tried first.

What costs you nothing:

  • Publishing no digest. Every hosted-passthrough model is permanently in this state, since there is no local file to hash for grok, gpt or claude.
  • A mismatch or unchecked registry verdict on its own. Payer routers don't read those fields.
  • Your self-reported trust tier. node_supervised does not move you up. Nothing you say about your own digest can raise you; only a peer's digest can lower you.
  • Being outnumbered by one operator's fleet. The count is over distinct owners, so a competitor spinning up nodes cannot vote you down.
  • A two-way split. If the operators serving a model divide evenly, nobody is demoted.

The honest way to land here is to serve a requantized or repacked copy under the same model id everyone else uses. Your bytes really do differ, and the router can't tell your Q4 repack from a substitution. Say so: point source: at the quantization's own repo, or advertise it under its own model id. Both are more accurate anyway, and both move you into your own identity group, where the comparison no longer applies.