Skip to main content

Image generation

Image generation is configured independently of text, so a node can serve text only, images only, or both. When it is configured, the node serves POST /v1/images/generations and POST /v1/images/edits. When it isn't, those routes return 404 image_generation_not_supported.

Two backend families, selected by image_llm.provider:

  • ComfyUI (comfyui self-hosted, or comfyui_cloud) — renders through workflow templates you author.
  • openai_passthrough — forwards to a hosted OpenAI-compatible image API such as xAI.

Either way, images are priced in USD per image, with image_rate for generations and image_edit_rate for edits, scaled by size and quality. Image tools on a chat model carry their own image_rate. See How billing works.

Pick a backend​

ComfyUI is operator-installed, not packaged with the node. Its VRAM adds to any text model on the same box; see Image models for per-model figures.

git clone https://github.com/comfyanonymous/ComfyUI ~/ComfyUI
cd ~/ComfyUI
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

# Drop your checkpoint into models/checkpoints/
# (e.g. sd_xl_base_1.0.safetensors from HuggingFace, ~6.5 GB)

python main.py --listen 127.0.0.1 --port 8000

Platform notes: Linux + NVIDIA needs CUDA 12.x; macOS Apple Silicon uses MPS automatically (16+ GB unified memory for Z-Image Turbo at BF16, 32+ GB for the FLUX.2 family); Linux + AMD needs the ROCm extras on a ROCm 6.x kernel, where the smaller models work and FLUX is hit-or-miss. For a steady-state install put ComfyUI under systemd or launchd so it and the node restart together.

image_llm:
provider: comfyui
comfyui:
base_url: "http://127.0.0.1:8000"
# data_dir: "/home/comfy/ComfyUI" # absolute; see "Disk cleanup" below

Image-only mode is automatic. Omit the llm: block (or set llm.provider: "") while image_llm.provider is set. /v1/chat/completions and /v1/responses then return 404 text_inference_not_supported, and /v1/models and /v1/zs/details advertise only your image models. Startup refuses image-only with no image model declared, image-only with a text model declared, and an image-flagged model with no image backend.

info

Image models are exempt from the provenance gate: a generation workflow maps to no single HuggingFace repo, so an image:-flagged model is never dropped for lacking a source. Declare one only if the checkpoint really is on HF and you want the cross-operator grouping.

Set expires_after on every image model​

Image generation takes far longer than the 5-minute zs.default_expires_after, which is the signed deadline the contract uses to gate the payer's inactivity refund. Without an explicit larger value, your settle transaction races the payer's refund:

zs:
models:
sdxl-base:
expires_after: 15m # SDXL on a real GPU: 10–15m. flux-dev / heavier: 20–30m.

Pad for queueing plus the settlement driver's lapse-grace buffer (5 minutes by default). This is independent of zs.ticket_ttl, which stays at its 5s default and bounds only how long a client has to POST after reserving.

Authoring a workflow template​

The node drives ComfyUI's graph API through a workflow JSON that you author. In the ComfyUI UI:

  1. Settings → enable Dev mode Options.
  2. Build the workflow you want to serve (checkpoint loader, sampler, text encode, latent, save image — the standard SDXL graph works).
  3. Click Save (API Format) to download the JSON.

Then edit the saved file to insert Go text/template placeholders for the fields the node fills at request time:

PlaceholderWhere it goesSource
{{ .Prompt | toJSON }}positive CLIPTextEncode textrequest prompt
{{ .NegPrompt | toJSON }}negative CLIPTextEncode textrequest negative_prompt
{{ .Width }} / {{ .Height }}EmptyLatentImage width / heightparsed from request size
{{ .BatchSize }}EmptyLatentImage batch_sizerequest n
{{ .Seed }}KSampler seedrequest seed, or random
{{ .Steps }}KSampler stepsper-model default
{{ .CFG }}KSampler cfgper-model default
{{ .Checkpoint | toJSON }}CheckpointLoaderSimple ckpt_nameper-model default
{{ .InputImageName | toJSON }}LoadImage image (edit workflows)filename ComfyUI returns after the node uploads the request's image part; empty on the generation path
{{ .MaskImageName | toJSON }}LoadImage image for a mask node (edit workflows)same, for the optional mask part; empty when no mask was supplied

Keep the | toJSON filter on every string placeholder; it makes filenames with spaces or unicode safe.

The built-in workflows​

Four workflows are compiled into the binary and referenced by bare name:

template_internalWhat it is
zimageText-to-image on z-image-turbo.
comfy_zimageThe Comfy-native variant of the same.
flux2_editImage-to-image / inpainting on Flux2 (Flux2-dev fp8mixed + Mistral-3-small CLIP).
hidream_editImage-to-image on HiDream-E1.1 (bf16, four text encoders: CLIP-G, CLIP-L, T5-XXL fp8, Llama-3.1-8B fp8) — a smaller VRAM footprint than Flux2.
warning

The name must match exactly, with no .json suffix and no aliasing. template_internal: zimage resolves to zimage.json; anything not in the table above — sdxl, sdxl-1024 — is a startup failure. If you need a different model, drop your own API-format workflow on disk and point template_path: at it.

All four hardcode their UNET / CLIP / VAE filenames, so they ignore {{ .Checkpoint }} and the defaults.checkpoint you set. A template you author around CheckpointLoaderSimple should keep the placeholder and let config supply the filename.

The two edit built-ins expose different knobs to defaults:

defaults knobflux2_edithidream_edit
width / heightIgnored — derived from the uploaded image.Ignored — derived from the uploaded image.
stepsUsed (the full-quality branch; the Turbo lora branch stays at 8).Used (denoise fixed at 1.0). Typical 20–28.
cfgUsed as Flux2's FluxGuidance, a different scale from KSampler's cfg: use roughly 3–5, not 7–9. cfg: 7 overshoots guidance and degrades output.Not used. Fork the workflow and use edit_template_path to tune guidance.
checkpointNot used.Not used.

For zimage-style KSampler workflows, defaults.cfg keeps its conventional meaning (5–9).

Declaring the model​

Each model that serves images gets an image: block. Exactly one of template_internal or template_path must be set for the generation workflow; edit workflows are optional, and at most one of edit_template_internal or edit_template_path may be set.

zs:
models:
flux2:
expires_after: 15m
image:
backend: comfyui
image_rate: 0.05 # $0.05 per 1024²-standard image (text → image)
image_edit_rate: 0.10 # $0.10 per 1024²-standard edit (image → image)
max_n: 4 # 0 = no clamp
defaults: {width: 1024, height: 1024, steps: 20, cfg: 7.0}
edit_defaults: {steps: 28, cfg: 3.0} # gen keeps 7.0, edit uses 3.0
comfyui:
template_internal: zimage
edit_template_internal: flux2_edit # or: hidream_edit

One model id serves both routes; the request route picks the workflow. A field left at zero in edit_defaults inherits from defaults verbatim. width and height usually stay unset, since edit templates derive output dimensions from the input image. Negative values are rejected at startup.

Edit requests bill at image_edit_rate, typically higher because image-to-image encodes the input through the VAE before sampling.

What happens on an edit request​

  1. The node parses the multipart body for image, optional mask, model, prompt, n, size.
  2. It POSTs the input image to ComfyUI's /upload/image and captures the assigned filename (and asset UUID, which Comfy Cloud always returns and self-hosted does not).
  3. It renders the edit template with that filename substituted into .InputImageName (and .MaskImageName when a mask is present), then submits the job exactly like the generation path.
  4. It cleans up the files it created — see Disk cleanup.

Edit-only models​

Omit the generation template and its image_rate entirely to serve only /v1/images/edits:

zs:
models:
hidream-edit:
expires_after: 15m
image:
backend: comfyui
image_edit_rate: 0.10 # no gen-side image_rate
max_n: 2
edit_defaults: {steps: 28, cfg: 3.0}
comfyui:
edit_template_internal: hidream_edit

The model advertises input_modalities: [text, image] and output_modalities: [image] with no image_rate, so payers skip you for generation and route only edits here. No defaults block is needed; the validator only requires the effective edit-path steps and cfg to be > 0. Discovery includes image in input_modalities only when the edit template and image_edit_rate are both declared, so SDKs that auto-detect from modalities won't try the edits route against a generation-only model. An edit POSTed to a generation-only model gets 400 image_edit_not_supported.

Output format (PNG / WebP / JPEG)​

Images default to PNG. Callers can request png, webp, or jpeg with an output_format field on /v1/images/generations, on /v1/images/edits (as a form field), and on the built-in zs_image_generation / zs_image_edit chat tools. WebP is typically 2–4× smaller at similar quality, which makes it a good default for chat.

The node asks ComfyUI to re-encode through its standard preview path, so stock ComfyUI supports it with no workflow change. If your build ignores the request and returns PNG anyway, the node detects that and labels the image correctly. An unrecognized output_format falls back to PNG.

WebP is also accepted as an input to /v1/images/edits. Its dimensions are read automatically, so a WebP edit needs no explicit size.

Image tools on /v1/responses​

A chat model can also call generation and editing as built-in tools mid-conversation, so "draw a cat, now put a hat on it" runs inside one request. Declare image_tools on the chat model, naming a configured image model to dispatch into:

zs:
models:
sdxl: # the image-route model the tools dispatch into
pricing: { input_rate: 0, output_rate: 0 }
image:
backend: comfyui
image_rate: 0.05
image_edit_rate: 0.10
defaults: { width: 1024, height: 1024, steps: 20, cfg: 7.0 }
comfyui:
template_internal: zimage
edit_template_internal: flux2_edit
gpt-4o: # the chat model that OFFERS the tools
pricing: { input_rate: 2.5, output_rate: 10.0 }
image_tools:
generation:
enabled: true
model: sdxl
image_rate: 0.05 # required, must be > 0 — no free-rate tools
max_size: 1024x1024 # default cap
max_quality: medium # default cap
default_size: 1024x1024
max_n: 4
edit:
enabled: true
model: sdxl
image_rate: 0.10
max_size: 1024x1024
  • The model picks size and quality mid-loop, so the caps matter. The node clamps the model's choice down to the cap (the request still succeeds) and reserves max_n × image_rate × factor(max_size, max_quality). The charge is measured off the images produced, as described in How billing works.
  • max_size / max_quality default to 1024x1024 / medium when unset.
  • default_size / default_quality fill in a value the model omits. default_size is also a floor: a smaller request is raised to it by area, keeping its aspect ratio, and a bare ratio like 16:9 is resolved to pixels at that area. Without the floor, a model could ask for 1024x1024 against a 2k-pinned tool and pay ¼ for an image the backend renders at 2k anyway. The default also sets the price clients display.
  • max_quality / default_quality do nothing on an edit tool — zs_image_edit exposes no quality argument and always renders the standard tier. Startup logs a WARN if you set them.
  • max_n: 0 (or absent) inherits the referenced image model's max_n; if that is also 0, there is no per-call clamp.
  • A model id is either a chat model (may carry image_tools) or an image-route model (carries image:) — never both. Startup rejects a model that declares both, an enabled tool with no positive image_rate, or a referenced model that doesn't serve the matching route.
  • Clients opt in per request with a tool_budgets object; omitted means the tool can't produce an image even if it is listed in tools[].
info

Whether the produced image is fed back to the model depends on its declared modalities. The image always reaches the client, as a marker item carrying the base64, so generation works on any chat model. It is fed back to the model as an input_image, so a follow-up "now put a hat on it" can edit what it just made, only when context.input_modalities includes "image". Runtimes that publish modalities (LM Studio, llama.cpp) are detected automatically. A passthrough upstream that publishes nothing (OpenAI, z.ai) is text-only by default, so declare it on a vision-capable chat model:

zs:
models:
glm-4.5v:
context:
input_modalities: ["text", "image"]

A text-only model gets the tool's text status instead; an image fed back to it would fail the whole turn with an opaque 400.

How billing works​

An image bills image_rate × factor(size, quality) in microUSDC, so the price scales with image area. Declare each rate as the USD price of one 1024²-standard-quality image:

  • factor(1024², standard) = 1.0 — that reference image is exactly what image_rate prices.
  • Area scales linearly: factor = (width × height) / (1024 × 1024). So 2048² ≈ 4× and 512² ≈ 0.25×. A request with no size, or a bare aspect ratio like 16:9, prices at the 1024² reference.
  • Quality tiers: low ×0.25; empty / medium / standard ×1.0; high / hd ×4.0.

So image_rate: 0.05 charges $0.05 for a 1024²-standard image, $0.20 for a 2048², and $0.003125 for a 512²-low. The node computes the exact figure from the request, signs it into the receipt tagged Images (not tokens), and sizes the reserve to match.

For openai_passthrough the request resolves through the size and quality knobs in this order: default_size / default_quality fill in an omitted value; an aspect-ratio size is resolved to concrete pixels covering default_size's area; default_size acts as a floor; then max_size / max_quality clamp down. The reserve resolves the same knobs, so the escrowed ceiling always covers the charge.

info

The final charge is measured off the image you actually deliver, shrunk to max_size's area if it overshoots, not off the size resolved above. That keeps the price right when a backend ignores the size it was handed or renders coarse tiers. A smaller image bills less, and the charge never exceeds the ceiling reserved for it. Measurement covers size only: quality leaves no observable trace in the returned image, so max_quality is the only bound on it.

If your backend renders more pixels than the request was priced for, the node bills the reserved amount, you absorb the difference, and the node logs a WARN with the delivered dimensions. A WARN on every request means your default_size / max_size don't describe what your backend produces.

Failures cost the payer nothing. A template rejection, a ComfyUI outage, or a generation error signs a zero-cost receipt, so the escrow refunds in full.

Rules the validator enforces at startup:

  • image_rate is required on a served generation route, and image_edit_rate on a served edit route. Images have no other pricing basis.
  • A rate without its matching template, or a template without its rate, is rejected.
  • A dedicated image model needs no pricing entry, since it produces no tokens, and does not inherit default_pricing. A pricing block is still accepted and validated if you write one; its token rates go unused.

/v1/zs/details advertises image_rate as the per-1024²-standard base, which reserve sizing and cross-operator price ordering key off, plus two display fields:

FieldShows
image_default_micro_usdcThe price a caller who sends no size pays.
image_cap_micro_usdcThe ceiling your caps impose.

Pinning default_size == max_size therefore displays one price rather than a ÷4 unit rate. Both fields are omitted when you set none of these knobs.

Operational notes​

  • ComfyUI processes one job at a time per instance. Concurrency comes from the existing max_active_tickets (global or per-model) — there is no separate image queue.
  • The image backend has its own health gate, running in parallel with the text one on the same llm.health_check cadence. When it's unreachable, image models drop from /v1/models and /v1/zs/details and image reserves return 503 provider_unavailable; the text backend keeps serving. Watch zs_image_provider_healthy and alert on a sustained 0.
  • Template changes need a node restart. Workflows are loaded at startup and do not hot-reload.
  • Comfy Cloud cost variance is yours. The cloud bills by GPU-second while your receipt charges per image regardless of how long the workflow ran. Pick an image_rate that covers your worst-case runtime at your tier.

Disk cleanup​

Every request leaves files on the ComfyUI side: uploaded inputs and masks in input/ (named zs-<random>.<ext>), generated images in output/, and previews in temp/. Left alone these grow without bound.

  • Comfy Cloud cleans up automatically. The node fires a best-effort asset delete for the input, mask, and output after each request.
  • Self-hosted ComfyUI has no delete API. Two options:
    • Let the node do it (recommended when co-located). Set image_llm.comfyui.data_dir to ComfyUI's base directory (the one holding input/, output/, temp/), and the node deletes the files it created right after each request, by path. The node needs write access there, which holds when ComfyUI runs on the same host or shares a volume. The path must be absolute. A not-yet-mounted path doesn't block startup; cleanup no-ops with a WARN until it appears. Prefer a local filesystem: the delete runs synchronously on the request path, so a slow network mount adds its unlink latency to the response.
    • Run your own cron. Leave data_dir unset and sweep <comfyui>/input/zs-*.*, <comfyui>/output/, and <comfyui>/temp/ on whatever cadence matches your throughput. This is the right choice for a remote or networked ComfyUI.

Environment variables​

VariableEffect
NODE_IMAGE_LLM_PROVIDEROverride image_llm.provider (comfyui, comfyui_cloud, openai_passthrough, or empty to disable)
NODE_IMAGE_LLM_COMFYUI_BASE_URLOverride image_llm.comfyui.base_url
NODE_IMAGE_LLM_COMFYUI_DATA_DIROverride image_llm.comfyui.data_dir
NODE_IMAGE_LLM_COMFYUI_CLOUD_BASE_URLOverride image_llm.comfyui_cloud.base_url
NODE_LLM_COMFYUI_CLOUD_API_KEYComfy Cloud API key — required for comfyui_cloud. Never in YAML.
NODE_IMAGE_LLM_OPENAI_API_KEYImage-backend upstream key for openai_passthrough. Never in YAML.