Image generation
Image generation is configured independently of text, so a node can serve
text only, images only, or both. When it is configured, the node serves
POST /v1/images/generations and POST /v1/images/edits. When it isn't, those
routes return 404 image_generation_not_supported.
Two backend families, selected by image_llm.provider:
- ComfyUI (
comfyuiself-hosted, orcomfyui_cloud) — renders through workflow templates you author. openai_passthrough— forwards to a hosted OpenAI-compatible image API such as xAI.
Either way, images are priced in USD per image, with image_rate for
generations and image_edit_rate for edits, scaled by size and quality. Image
tools on a chat model carry their own image_rate. See
How billing works.
Pick a backend
- comfyui (self-hosted)
- comfyui_cloud
- openai_passthrough
ComfyUI is operator-installed, not packaged with the node. Its VRAM adds to any text model on the same box; see Image models for per-model figures.
git clone https://github.com/comfyanonymous/ComfyUI ~/ComfyUI
cd ~/ComfyUI
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Drop your checkpoint into models/checkpoints/
# (e.g. sd_xl_base_1.0.safetensors from HuggingFace, ~6.5 GB)
python main.py --listen 127.0.0.1 --port 8000
Platform notes: Linux + NVIDIA needs CUDA 12.x; macOS Apple Silicon uses MPS automatically (16+ GB unified memory for Z-Image Turbo at BF16, 32+ GB for the FLUX.2 family); Linux + AMD needs the ROCm extras on a ROCm 6.x kernel, where the smaller models work and FLUX is hit-or-miss. For a steady-state install put ComfyUI under systemd or launchd so it and the node restart together.
image_llm:
provider: comfyui
comfyui:
base_url: "http://127.0.0.1:8000"
# data_dir: "/home/comfy/ComfyUI" # absolute; see "Disk cleanup" below
Comfy Cloud runs the same workflow JSON and event protocol on Comfy Org's hardware. Mint an API key at platform.comfy.org. It is shown once, on creation, and rotated from the same dashboard.
image_llm:
provider: comfyui_cloud
comfyui_cloud:
base_url: "https://cloud.comfy.org/api" # the default; rarely overridden
zs:
models:
sdxl-cloud:
expires_after: 15m
max_active_tickets: 3 # match your tier's concurrent-job cap
image:
backend: comfyui_cloud
image_rate: 0.05
max_n: 4
defaults: {width: 1024, height: 1024, steps: 20, cfg: 7.0}
comfyui_cloud:
template_path: "~/.config/zs-node/templates/sdxl.json"
export NODE_LLM_COMFYUI_CLOUD_API_KEY=<your-key>
Startup validation refuses an empty key. Differences from self-hosted:
- The key is sent on every call as
X-API-Key, and appended to the WebSocket handshake as&token=…. - Tier concurrency caps are yours to enforce. Standard has no API access;
Creator allows 3 concurrent jobs, Pro 5, Enterprise custom. Set the per-model
max_active_ticketsto match, so the ticket store refuses admission before a burst of reserves meets an upstream429. - Outputs arrive as signed URLs. The node follows the view endpoint's
redirect to short-lived storage, fetches the bytes, and base64-encodes them
into the response; the signed URL never reaches the client.
X-API-Keyis stripped on cross-host redirects, so the CDN never sees your key.
A workflow declares its backend in exactly one block — image.comfyui or
image.comfyui_cloud, never both — and the block must match the active
image_llm.provider. To migrate a model between them, swap the block name and
the backend value together.
Forwards /v1/images/generations and /v1/images/edits to a hosted upstream.
There are no workflow templates and no sampler defaults: the ComfyUI
defaults: knobs (width/height/steps/cfg) do nothing here.
image_llm:
provider: openai_passthrough
openai:
base_url: "https://api.x.ai/v1" # the API ROOT — /images/* is appended
api_key: "" # prefer NODE_IMAGE_LLM_OPENAI_API_KEY
zs:
models:
"grok-imagine-image-quality":
image:
backend: "openai_passthrough"
image_rate: 0.021 # USD per 1024²-STANDARD image
image_edit_rate: 0.021 # omit to not serve /v1/images/edits
default_size: "2048x2048" # fill in when the caller omits size
max_size: "2048x2048" # clamp DOWN
max_quality: "medium" # blocks the pointless ×4
The key is independent of the text NODE_LLM_OPENAI_API_KEY, so text and images
can point at different upstreams. A model serves generation when image_rate is
set and edits when image_edit_rate is set; declare at least one.
Response format is forced to b64_json. The node hosts no URLs and won't
fetch an upstream-controlled one, so an upstream that returns a URL anyway is
rejected. Transient upstream errors (429/5xx, or a pre-response
transport error) are retried in place with bounded backoff honoring
Retry-After; non-transient errors are forwarded verbatim.
xAI request shaping and zero data retention are both automatic, inferred from the URL, with no config knob.
When base_url resolves to xAI, the node rebuilds requests for xAI's image API,
which is not OpenAI-shaped:
- Edits go as
application/jsonwith the source image as a data-URI object; xAI rejects the multipart edit form. - On both routes, the OpenAI-only
size/quality/style/output_formatparams are dropped (xAI answers400 Argument not supported: sizeotherwise), and a requestedsizeis translated into xAI'saspect_ratio+resolution. - An edit
maskhas no xAI analog and is dropped.
xAI also requires org-level Zero Data Retention (Team Settings → Zero Data
Retention), confirmed by an x-zero-data-retention: true header on every image
response. Without it, each image request is refused with
403 zero_data_retention_required. store: false does not provide this.
The check runs per request; there is no startup probe, because it would cost a
real image. If you also run an xAI text backend, its startup probe already
fails boot on a non-ZDR account.
Sizing xAI fairly. xAI renders only its 1k / 2k tiers, and an omitted
resolution renders 1k; on an edit it does not follow the source image.
Because image_rate prices a 1024²-standard image, a 2k square bills ×4.
The tiers are also not area-preserving across aspect ratios. Measured against
grok-imagine-image-quality:
| Request | Delivered | Factor |
|---|---|---|
1:1 at 2k | 2048×2048 | ×4.00 |
16:9 at 2k | 2816×1584 | ×4.25 |
16:9, no resolution | 1280×720 | ×0.88 |
Billing uses the delivered dimensions, clamped to max_size. Set
default_size and max_size to "2048x2048", max_quality to "medium",
and price image_rate at your intended 2k price ÷ 4.
Image-only mode is automatic. Omit the llm: block (or set
llm.provider: "") while image_llm.provider is set. /v1/chat/completions
and /v1/responses then return 404 text_inference_not_supported, and
/v1/models and /v1/zs/details advertise only your image models. Startup
refuses image-only with no image model declared, image-only with a text model
declared, and an image-flagged model with no image backend.
Image models are exempt from the provenance
gate: a generation workflow
maps to no single HuggingFace repo, so an image:-flagged model is never
dropped for lacking a source. Declare one only if the checkpoint really is on
HF and you want the cross-operator grouping.
Set expires_after on every image model
Image generation takes far longer than the 5-minute zs.default_expires_after,
which is the signed deadline the contract uses to gate the payer's inactivity
refund. Without an explicit larger value, your settle transaction races the
payer's refund:
zs:
models:
sdxl-base:
expires_after: 15m # SDXL on a real GPU: 10–15m. flux-dev / heavier: 20–30m.
Pad for queueing plus the settlement driver's lapse-grace buffer (5 minutes by
default). This is independent of zs.ticket_ttl, which stays at its 5s
default and bounds only how long a client has to POST after reserving.
Authoring a workflow template
The node drives ComfyUI's graph API through a workflow JSON that you author. In the ComfyUI UI:
- Settings → enable Dev mode Options.
- Build the workflow you want to serve (checkpoint loader, sampler, text encode, latent, save image — the standard SDXL graph works).
- Click Save (API Format) to download the JSON.
Then edit the saved file to insert Go text/template placeholders for the
fields the node fills at request time:
| Placeholder | Where it goes | Source |
|---|---|---|
{{ .Prompt | toJSON }} | positive CLIPTextEncode text | request prompt |
{{ .NegPrompt | toJSON }} | negative CLIPTextEncode text | request negative_prompt |
{{ .Width }} / {{ .Height }} | EmptyLatentImage width / height | parsed from request size |
{{ .BatchSize }} | EmptyLatentImage batch_size | request n |
{{ .Seed }} | KSampler seed | request seed, or random |
{{ .Steps }} | KSampler steps | per-model default |
{{ .CFG }} | KSampler cfg | per-model default |
{{ .Checkpoint | toJSON }} | CheckpointLoaderSimple ckpt_name | per-model default |
{{ .InputImageName | toJSON }} | LoadImage image (edit workflows) | filename ComfyUI returns after the node uploads the request's image part; empty on the generation path |
{{ .MaskImageName | toJSON }} | LoadImage image for a mask node (edit workflows) | same, for the optional mask part; empty when no mask was supplied |
Keep the | toJSON filter on every string placeholder; it makes filenames with
spaces or unicode safe.
The built-in workflows
Four workflows are compiled into the binary and referenced by bare name:
template_internal | What it is |
|---|---|
zimage | Text-to-image on z-image-turbo. |
comfy_zimage | The Comfy-native variant of the same. |
flux2_edit | Image-to-image / inpainting on Flux2 (Flux2-dev fp8mixed + Mistral-3-small CLIP). |
hidream_edit | Image-to-image on HiDream-E1.1 (bf16, four text encoders: CLIP-G, CLIP-L, T5-XXL fp8, Llama-3.1-8B fp8) — a smaller VRAM footprint than Flux2. |
The name must match exactly, with no .json suffix and no aliasing.
template_internal: zimage resolves to zimage.json; anything not in the table
above — sdxl, sdxl-1024 — is a startup failure. If you need a different
model, drop your own API-format workflow on disk and point template_path: at
it.
All four hardcode their UNET / CLIP / VAE filenames, so they ignore
{{ .Checkpoint }} and the defaults.checkpoint you set. A template you author
around CheckpointLoaderSimple should keep the placeholder and let config
supply the filename.
The two edit built-ins expose different knobs to defaults:
defaults knob | flux2_edit | hidream_edit |
|---|---|---|
width / height | Ignored — derived from the uploaded image. | Ignored — derived from the uploaded image. |
steps | Used (the full-quality branch; the Turbo lora branch stays at 8). | Used (denoise fixed at 1.0). Typical 20–28. |
cfg | Used as Flux2's FluxGuidance, a different scale from KSampler's cfg: use roughly 3–5, not 7–9. cfg: 7 overshoots guidance and degrades output. | Not used. Fork the workflow and use edit_template_path to tune guidance. |
checkpoint | Not used. | Not used. |
For zimage-style KSampler workflows, defaults.cfg keeps its conventional
meaning (5–9).
Declaring the model
Each model that serves images gets an image: block. Exactly one of
template_internal or template_path must be set for the generation workflow;
edit workflows are optional, and at most one of edit_template_internal or
edit_template_path may be set.
zs:
models:
flux2:
expires_after: 15m
image:
backend: comfyui
image_rate: 0.05 # $0.05 per 1024²-standard image (text → image)
image_edit_rate: 0.10 # $0.10 per 1024²-standard edit (image → image)
max_n: 4 # 0 = no clamp
defaults: {width: 1024, height: 1024, steps: 20, cfg: 7.0}
edit_defaults: {steps: 28, cfg: 3.0} # gen keeps 7.0, edit uses 3.0
comfyui:
template_internal: zimage
edit_template_internal: flux2_edit # or: hidream_edit
One model id serves both routes; the request route picks the workflow. A field
left at zero in edit_defaults inherits from defaults verbatim. width and
height usually stay unset, since edit templates derive output dimensions from
the input image. Negative values are rejected at startup.
Edit requests bill at image_edit_rate, typically higher because
image-to-image encodes the input through the VAE before sampling.
What happens on an edit request
- The node parses the multipart body for
image, optionalmask,model,prompt,n,size. - It POSTs the input image to ComfyUI's
/upload/imageand captures the assigned filename (and asset UUID, which Comfy Cloud always returns and self-hosted does not). - It renders the edit template with that filename substituted into
.InputImageName(and.MaskImageNamewhen a mask is present), then submits the job exactly like the generation path. - It cleans up the files it created — see Disk cleanup.
Edit-only models
Omit the generation template and its image_rate entirely to serve only
/v1/images/edits:
zs:
models:
hidream-edit:
expires_after: 15m
image:
backend: comfyui
image_edit_rate: 0.10 # no gen-side image_rate
max_n: 2
edit_defaults: {steps: 28, cfg: 3.0}
comfyui:
edit_template_internal: hidream_edit
The model advertises input_modalities: [text, image] and
output_modalities: [image] with no image_rate, so payers skip you for
generation and route only edits here. No defaults block is needed; the
validator only requires the effective edit-path steps and cfg to be > 0.
Discovery includes image in input_modalities only when the edit template
and image_edit_rate are both declared, so SDKs that auto-detect from
modalities won't try the edits route against a generation-only model. An edit
POSTed to a generation-only model gets 400 image_edit_not_supported.
Output format (PNG / WebP / JPEG)
Images default to PNG. Callers can request png, webp, or jpeg with an
output_format field on /v1/images/generations, on /v1/images/edits (as a
form field), and on the built-in zs_image_generation / zs_image_edit chat
tools. WebP is typically 2–4× smaller at similar quality, which makes it a good
default for chat.
The node asks ComfyUI to re-encode through its standard preview path, so stock
ComfyUI supports it with no workflow change. If your build ignores the
request and returns PNG anyway, the node detects that and labels the image
correctly. An unrecognized output_format falls back to PNG.
WebP is also accepted as an input to /v1/images/edits. Its dimensions are
read automatically, so a WebP edit needs no explicit size.
Image tools on /v1/responses
A chat model can also call generation and editing as built-in tools
mid-conversation, so "draw a cat, now put a hat on it" runs inside one request.
Declare image_tools on the chat model, naming a configured image model to
dispatch into:
zs:
models:
sdxl: # the image-route model the tools dispatch into
pricing: { input_rate: 0, output_rate: 0 }
image:
backend: comfyui
image_rate: 0.05
image_edit_rate: 0.10
defaults: { width: 1024, height: 1024, steps: 20, cfg: 7.0 }
comfyui:
template_internal: zimage
edit_template_internal: flux2_edit
gpt-4o: # the chat model that OFFERS the tools
pricing: { input_rate: 2.5, output_rate: 10.0 }
image_tools:
generation:
enabled: true
model: sdxl
image_rate: 0.05 # required, must be > 0 — no free-rate tools
max_size: 1024x1024 # default cap
max_quality: medium # default cap
default_size: 1024x1024
max_n: 4
edit:
enabled: true
model: sdxl
image_rate: 0.10
max_size: 1024x1024
- The model picks
sizeandqualitymid-loop, so the caps matter. The node clamps the model's choice down to the cap (the request still succeeds) and reservesmax_n × image_rate × factor(max_size, max_quality). The charge is measured off the images produced, as described in How billing works. max_size/max_qualitydefault to1024x1024/mediumwhen unset.default_size/default_qualityfill in a value the model omits.default_sizeis also a floor: a smaller request is raised to it by area, keeping its aspect ratio, and a bare ratio like16:9is resolved to pixels at that area. Without the floor, a model could ask for1024x1024against a 2k-pinned tool and pay ¼ for an image the backend renders at 2k anyway. The default also sets the price clients display.max_quality/default_qualitydo nothing on anedittool —zs_image_editexposes no quality argument and always renders the standard tier. Startup logs a WARN if you set them.max_n: 0(or absent) inherits the referenced image model'smax_n; if that is also0, there is no per-call clamp.- A model id is either a chat model (may carry
image_tools) or an image-route model (carriesimage:) — never both. Startup rejects a model that declares both, an enabled tool with no positiveimage_rate, or a referenced model that doesn't serve the matching route. - Clients opt in per request with a
tool_budgetsobject; omitted means the tool can't produce an image even if it is listed intools[].
Whether the produced image is fed back to the model depends on its declared
modalities. The image always reaches the client, as a marker item carrying
the base64, so generation works on any chat model. It is fed back to the model
as an input_image, so a follow-up "now put a hat on it" can edit what it just
made, only when context.input_modalities includes "image". Runtimes that
publish modalities (LM Studio, llama.cpp) are detected automatically. A
passthrough upstream that publishes nothing (OpenAI, z.ai) is text-only by
default, so declare it on a vision-capable chat model:
zs:
models:
glm-4.5v:
context:
input_modalities: ["text", "image"]
A text-only model gets the tool's text status instead; an image fed back to it
would fail the whole turn with an opaque 400.
How billing works
An image bills image_rate × factor(size, quality) in microUSDC, so the price
scales with image area. Declare each rate as the USD price of one
1024²-standard-quality image:
factor(1024², standard) = 1.0— that reference image is exactly whatimage_rateprices.- Area scales linearly:
factor = (width × height) / (1024 × 1024). So 2048² ≈ 4× and 512² ≈ 0.25×. A request with nosize, or a bare aspect ratio like16:9, prices at the 1024² reference. - Quality tiers:
low×0.25; empty /medium/standard×1.0;high/hd×4.0.
So image_rate: 0.05 charges $0.05 for a 1024²-standard image, $0.20 for a
2048², and $0.003125 for a 512²-low. The node computes the exact figure from the
request, signs it into the receipt tagged Images (not tokens), and sizes the
reserve to match.
For openai_passthrough the request resolves through the size and quality knobs
in this order: default_size / default_quality fill in an omitted value; an
aspect-ratio size is resolved to concrete pixels covering default_size's area;
default_size acts as a floor; then max_size / max_quality clamp down. The
reserve resolves the same knobs, so the escrowed ceiling always covers the
charge.
The final charge is measured off the image you actually deliver, shrunk to
max_size's area if it overshoots, not off the size resolved above. That keeps
the price right when a backend ignores the size it was handed or renders coarse
tiers. A smaller image bills less, and the charge never exceeds the ceiling
reserved for it. Measurement covers size only: quality leaves no observable trace in
the returned image, so max_quality is the only bound on it.
If your backend renders more pixels than the request was priced for, the node
bills the reserved amount, you absorb the difference, and the node logs a
WARN with the delivered dimensions. A WARN on every request means your
default_size / max_size don't describe what your backend produces.
Failures cost the payer nothing. A template rejection, a ComfyUI outage, or a generation error signs a zero-cost receipt, so the escrow refunds in full.
Rules the validator enforces at startup:
image_rateis required on a served generation route, andimage_edit_rateon a served edit route. Images have no other pricing basis.- A rate without its matching template, or a template without its rate, is rejected.
- A dedicated image model needs no
pricingentry, since it produces no tokens, and does not inheritdefault_pricing. Apricingblock is still accepted and validated if you write one; its token rates go unused.
/v1/zs/details advertises image_rate as the per-1024²-standard base, which
reserve sizing and cross-operator price ordering key off, plus two display
fields:
| Field | Shows |
|---|---|
image_default_micro_usdc | The price a caller who sends no size pays. |
image_cap_micro_usdc | The ceiling your caps impose. |
Pinning default_size == max_size therefore displays one price rather than a
÷4 unit rate. Both fields are omitted when you set none of these knobs.
Operational notes
- ComfyUI processes one job at a time per instance. Concurrency comes from
the existing
max_active_tickets(global or per-model) — there is no separate image queue. - The image backend has its own health gate, running in parallel with the
text one on the same
llm.health_checkcadence. When it's unreachable, image models drop from/v1/modelsand/v1/zs/detailsand image reserves return503 provider_unavailable; the text backend keeps serving. Watchzs_image_provider_healthyand alert on a sustained0. - Template changes need a node restart. Workflows are loaded at startup and do not hot-reload.
- Comfy Cloud cost variance is yours. The cloud bills by GPU-second while
your receipt charges per image regardless of how long the workflow ran. Pick
an
image_ratethat covers your worst-case runtime at your tier.
Disk cleanup
Every request leaves files on the ComfyUI side: uploaded inputs and masks in
input/ (named zs-<random>.<ext>), generated images in output/, and
previews in temp/. Left alone these grow without bound.
- Comfy Cloud cleans up automatically. The node fires a best-effort asset delete for the input, mask, and output after each request.
- Self-hosted ComfyUI has no delete API. Two options:
- Let the node do it (recommended when co-located). Set
image_llm.comfyui.data_dirto ComfyUI's base directory (the one holdinginput/,output/,temp/), and the node deletes the files it created right after each request, by path. The node needs write access there, which holds when ComfyUI runs on the same host or shares a volume. The path must be absolute. A not-yet-mounted path doesn't block startup; cleanup no-ops with a WARN until it appears. Prefer a local filesystem: the delete runs synchronously on the request path, so a slow network mount adds its unlink latency to the response. - Run your own cron. Leave
data_dirunset and sweep<comfyui>/input/zs-*.*,<comfyui>/output/, and<comfyui>/temp/on whatever cadence matches your throughput. This is the right choice for a remote or networked ComfyUI.
- Let the node do it (recommended when co-located). Set
Environment variables
| Variable | Effect |
|---|---|
NODE_IMAGE_LLM_PROVIDER | Override image_llm.provider (comfyui, comfyui_cloud, openai_passthrough, or empty to disable) |
NODE_IMAGE_LLM_COMFYUI_BASE_URL | Override image_llm.comfyui.base_url |
NODE_IMAGE_LLM_COMFYUI_DATA_DIR | Override image_llm.comfyui.data_dir |
NODE_IMAGE_LLM_COMFYUI_CLOUD_BASE_URL | Override image_llm.comfyui_cloud.base_url |
NODE_LLM_COMFYUI_CLOUD_API_KEY | Comfy Cloud API key — required for comfyui_cloud. Never in YAML. |
NODE_IMAGE_LLM_OPENAI_API_KEY | Image-backend upstream key for openai_passthrough. Never in YAML. |