Say what actually failed, and tell the model how to use the thing

Two problems, both found by looking rather than by guessing.

ComfyUI writes its history entry in task_done and nowhere else, so the entry
appearing IS "finished" -- but it sets completed=e.success, which means an
out-of-memory, a cancelled job and a broken node all stay completed:false for
ever. await_images waited on that flag. So every failure sat for the full 600s
timeout and then reported a timeout, when ComfyUI had known within one second and
written down the node, the exception type and the message. Proved by causing both
against the real instance: an OOM now raises in 1.0s and an interrupt in 4.0s,
each naming the node.

The terminal condition is a record with a status, and status.messages is read for
the last execution_error or execution_interrupted. OutOfMemory and Interrupted
are their own classes because they are the two failures with an obvious next
move: the first tells the model to retry at a named smaller size -- worked out
from what it actually asked for, since "use a lower resolution" against a request
that was already 512x512 is advice nobody can follow -- or with a lighter
checkpoint; the second says somebody pressed stop, so do not simply start again.
Everything else gets the reason and no advice, because a model told to try again
after a broken workflow tries the identical thing.

The OOM message is cut to its first sentence. The rest is allocator advice --
PYTORCH_CUDA_ALLOC_CONF, fragmentation notes -- addressed to whoever runs the box
and meaningless to a model, in a tool result that is already a failure.

Second: the parameters were described in the register of a reference table, and
"cfg: prompt adherence, default 8" tells a model nothing it can act on. Measured
on a 4B model, same request, same everything else: with the old wording it sent
prompt and template and nothing more -- so 512x512 on an SDXL checkpoint, which
is exactly the duplicated-limbs failure the width description now warns about.
With descriptions that say what each value does to the picture and when to move
it, the same model sent a portrait 1024x1536 and a deliberate sampler. ~3KB of
schema per request in a chat that can draw, and the difference between having ten
parameters and having one.

docs/image-generation-instructions.md is the long version for the admin
instructions box, for models that need more than the harness can afford to carry
on every request in every chat.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jaroslav Beneš
2026-08-05 14:55:18 +02:00
parent b2a05e0351
commit 178742501d
6 changed files with 447 additions and 40 deletions
+83 -14
View File
@@ -93,6 +93,29 @@ class ComfyError(LLMError):
"""Anything that stopped a generation, in words worth showing somebody."""
class OutOfMemory(ComfyError):
"""The far side ran out of VRAM.
Its own class because it is the one failure with an obvious next move --
a smaller picture, or a smaller checkpoint -- and the model is told to make
it. Everything else is reported and stopped at.
"""
class Interrupted(ComfyError):
"""Somebody cancelled it from ComfyUI's own interface, or it was stopped.
Distinct because it is not a fault: retrying is reasonable, and "the
workflow failed" would be describing a decision as a breakage.
"""
# What `exception_type` looks like when a GPU has run out. Matched on the type
# rather than on the message, which is a paragraph of allocator advice written
# for whoever is running the box and not for a model.
_OOM_TYPES = ("outofmemory", "out_of_memory", "cuda error: out of memory")
def _transport_error(exc: httpx.RequestError, config: Config) -> ComfyError:
"""The `wrap_transport_error` shape, said about ComfyUI rather than an LLM.
@@ -133,9 +156,7 @@ async def submit(config: Config, workflow: dict[str, Any]) -> str:
body = {"prompt": workflow, "client_id": uuid.uuid4().hex}
try:
async with httpx.AsyncClient(timeout=60.0) as client:
response = await client.post(
config.url("prompt"), headers=config.headers(), json=body
)
response = await client.post(config.url("prompt"), headers=config.headers(), json=body)
if response.status_code >= 400:
raise ComfyError(_refusal(response))
data = response.json()
@@ -185,19 +206,24 @@ def _describe_nodes(errors: dict[str, Any]) -> str:
async def await_images(config: Config, prompt_id: str) -> list[Ref]:
"""Wait for one queued workflow and answer with what it saved.
`/history/{id}` is empty while the job is queued or running and gains the
whole record when it ends, so an empty answer is "not yet" rather than
"nothing" -- which is why the deadline is the only thing that ends this.
**The record existing is what "finished" means, not `status.completed`.**
ComfyUI writes the history entry in `task_done` and nowhere else, so it
appears exactly once the job is over -- but it sets `completed=e.success`,
so a run that failed is `completed: false` for ever. Waiting on that flag
means every out-of-memory, every cancelled job and every broken node hangs
the reply for the whole timeout and then reports a timeout, when ComfyUI
knew what was wrong within seconds and said so.
So: no record means not yet, a record means done, and `status_str` says
which kind of done.
"""
deadline = time.monotonic() + config.timeout
while True:
record = (await _get_json(config, f"history/{prompt_id}")).get(prompt_id)
if isinstance(record, dict) and (record.get("status") or {}).get("completed"):
if isinstance(record, dict) and record.get("status") is not None:
status = record.get("status") or {}
if status.get("status_str") not in (None, "success"):
raise ComfyError(
f"ComfyUI could not finish the workflow ({status.get('status_str')})."
)
if status.get("status_str") != "success":
raise _failure(status)
return _refs_in(record.get("outputs") or {})
if time.monotonic() > deadline:
raise ComfyError(
@@ -207,6 +233,51 @@ async def await_images(config: Config, prompt_id: str) -> list[Ref]:
await asyncio.sleep(POLL_INTERVAL)
def _failure(status: dict[str, Any]) -> ComfyError:
"""Why a workflow stopped, out of the messages ComfyUI recorded against it.
`status.messages` is a list of `[name, payload]` pairs -- the lifecycle of
the run. The last `execution_error` or `execution_interrupted` in it is the
thing that ended it, and carries the node and the exception. Without reading
these the only thing that could be said is "error", which is what ComfyUI's
own status string amounts to.
"""
event, payload = "", {}
for entry in status.get("messages") or []:
if isinstance(entry, list | tuple) and len(entry) == 2:
name, body = entry
if name in ("execution_error", "execution_interrupted"):
event, payload = str(name), body if isinstance(body, dict) else {}
node = str(payload.get("node_type") or "").strip()
where = f" in {node}" if node else ""
if event == "execution_interrupted":
return Interrupted(f"The image was cancelled on the ComfyUI side{where}.")
kind = str(payload.get("exception_type") or "")
detail = _first_sentence(str(payload.get("exception_message") or ""))
if any(marker in kind.lower() for marker in _OOM_TYPES) or "out of memory" in detail.lower():
return OutOfMemory(f"ComfyUI ran out of video memory{where}. {detail}".strip())
if not detail and not kind:
return ComfyError(f"ComfyUI could not finish the workflow{where}.")
return ComfyError(f"ComfyUI could not finish the workflow{where}: {detail or kind}")
def _first_sentence(message: str) -> str:
"""Enough of an exception to act on, and no more.
A torch OOM runs to several lines of allocator advice -- environment
variables to set, fragmentation notes -- addressed to whoever runs the box.
None of it means anything to a model, and all of it costs tokens in a tool
result that is already a failure.
"""
first = message.strip().split("\n", 1)[0].strip()
if len(first) > 200:
first = first[:200].rsplit(" ", 1)[0] + ""
return first
def _refs_in(outputs: dict[str, Any]) -> list[Ref]:
"""Every image any node saved, in node order.
@@ -233,9 +304,7 @@ async def fetch_image(config: Config, ref: Ref) -> bytes:
params = {"filename": ref.filename, "subfolder": ref.subfolder, "type": ref.kind}
try:
async with httpx.AsyncClient(timeout=120.0) as client:
response = await client.get(
config.url("view"), headers=config.headers(), params=params
)
response = await client.get(config.url("view"), headers=config.headers(), params=params)
response.raise_for_status()
payload = response.content
except httpx.HTTPStatusError as exc:
+114 -11
View File
@@ -61,9 +61,19 @@ SCHEMA: dict[str, Any] = {
"type": "string",
"description": "What to draw. Describe the subject, the setting and the style.",
},
# Every description below says what the value *does to the picture* and
# when to move it, not what it is called. A model that is told "cfg:
# prompt adherence, default 8" has been told nothing it can act on, and
# the observable result is a model that sends the prompt alone and
# leaves ten parameters at their defaults for ever.
"negative": {
"type": "string",
"description": "What to keep out of the picture. Defaults to 'text, watermark'.",
"description": (
"Comma-separated things to keep OUT of the picture, as plain nouns and "
"adjectives: 'blurry, extra fingers, text, watermark'. Not a sentence, "
"and never phrased as an instruction — 'do not add text' puts *text* in "
"the picture. Defaults to 'text, watermark'."
),
},
"template": {
"type": "string",
@@ -71,22 +81,80 @@ SCHEMA: dict[str, Any] = {
},
"model": {
"type": "string",
"description": "Which checkpoint to draw with. Omit to use this chat's usual one.",
"description": (
"Which checkpoint to draw with. Pick by what it is good at; omit to use "
"this chat's usual one."
),
},
"seed": {
"type": "integer",
"description": (
"Omit it, or pass -1, for a new random image. Repeat a seed you were "
"told about to get the same image again."
"told about to get that same image again — which is how you change one "
"thing about a picture and keep the rest."
),
},
"steps": {
"type": "integer",
"description": (
"How long to refine, 1-150. Default 20. Around 20-30 for most things; "
"8-12 for a quick draft or when several are wanted; 40+ only for fine "
"detail, and past about 50 it stops improving and only costs time."
),
},
"cfg": {
"type": "number",
"description": (
"How literally to follow the prompt, 0-30. Default 8. 3-6 gives the "
"model room and looks more natural; 7-9 is the usual range; 12+ forces "
"the words through and starts to look burnt and over-saturated. Lower "
"it if the picture looks harsh, raise it if the subject is being "
"ignored."
),
},
"width": {
"type": "integer",
"description": (
"Pixels, 64-2048, a multiple of 8. Default 512. Use the size the "
"checkpoint was trained for — about 512 for SD1.5, about 1024 for SDXL "
"— and change the ratio rather than the total: 512x768 for a portrait, "
"768x512 for a landscape. Going far above what the checkpoint expects "
"produces duplicated limbs and repeated horizons, not more detail."
),
},
"height": {
"type": "integer",
"description": (
"Pixels, 64-2048, a multiple of 8. Default 512. See width: the aspect "
"ratio is the thing to choose, and taller than wide suits a person, "
"wider than tall suits a place."
),
},
"sampler": {
"type": "string",
"description": (
"How the image is solved. Default euler. 'euler' is safe and fast; "
"'dpmpp_2m' is a good general improvement; 'dpmpp_2m_sde' for more "
"texture; 'ddim' for a clean flat look. Leave it out unless you have a "
"reason."
),
},
"scheduler": {
"type": "string",
"description": (
"How the steps are spaced. Default normal. 'karras' pairs well with the "
"dpmpp samplers and usually helps at low step counts; 'normal' "
"otherwise. Leave it out unless you are also setting the sampler."
),
},
"denoise": {
"type": "number",
"description": (
"How much of the starting noise to replace, 0-1. Default 1, which is "
"what you want for a picture drawn from nothing. Lower values only mean "
"something for a workflow that starts from an existing image."
),
},
"steps": {"type": "integer", "description": "Sampling steps. Default 20."},
"cfg": {"type": "number", "description": "Prompt adherence. Default 8."},
"width": {"type": "integer", "description": "Pixels. Default 512."},
"height": {"type": "integer", "description": "Pixels. Default 512."},
"sampler": {"type": "string", "description": "Sampler name. Default euler."},
"scheduler": {"type": "string", "description": "Scheduler name. Default normal."},
"denoise": {"type": "number", "description": "0 to 1. Default 1."},
},
"required": ["prompt"],
}
@@ -371,6 +439,9 @@ async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
attempts: list[Attempt] = []
kept: tuple[bytes, dict[str, Any]] | None = None
# What the last attempt actually asked for, so a failure can name concrete
# numbers back at the model rather than saying "try something smaller".
params_used: dict[str, Any] = workflow.resolve(given)
try:
for number in range(1, tries + 1):
@@ -378,6 +449,7 @@ async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
await _unload_llm(context)
params = workflow.resolve({**given, "seed": args.get("seed") if number == 1 else None})
params_used = params
refs = await comfy.await_images(
config, await comfy.submit(config, workflow.fill(template, params))
)
@@ -402,9 +474,12 @@ async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
break
except comfy.ComfyError as exc:
if preserve:
# It failed *inside* the far side, so its models are still resident
# and the language model is still unloaded. Freeing here is what
# lets the reply carry on and say what happened.
await comfy.free(config)
return ToolOutcome(
f"The image could not be generated: {exc.message}",
f"The image could not be generated: {exc.message}{_advice(exc, params_used)}",
{**event, "status": "error", "error": exc.message},
)
@@ -454,6 +529,34 @@ async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
)
def _advice(exc: comfy.ComfyError, params: dict[str, Any]) -> str:
"""What to do about a failure, when there is something to do about it.
Only for the two that have an obvious next move. Everything else gets the
reason and nothing else -- a model told to "try again" after a broken
workflow will try the identical thing, and a suggestion invented for a
failure nobody understands is a guess wearing the application's authority.
The numbers are concrete on purpose. "Use a lower resolution" against a
request that was already 512x512 is advice that cannot be followed, so the
halved size is worked out here where the request is known.
"""
if isinstance(exc, comfy.Interrupted):
return (
" Somebody stopped it deliberately, so do not simply start it again — say so and ask."
)
if not isinstance(exc, comfy.OutOfMemory):
return ""
width, height = int(params.get("width") or 512), int(params.get("height") or 512)
smaller = f"{max(256, width // 2)}x{max(256, height // 2)}"
return (
f" Try once more at a smaller size — {smaller} instead of {width}x{height}"
"or with a lighter checkpoint if one is offered. Do not repeat the same "
"request unchanged; it will run out of memory again."
)
def _pick(rows: list[Any], wanted: str, chat_default: str, values: dict[str, Any]) -> Any:
"""The workflow to use: asked for, then the chat's, then the instance's."""
by_slug = {row.slug: row for row in rows}
+31 -14
View File
@@ -1025,21 +1025,38 @@ BUILTIN: tuple[Fragment, ...] = (
group=GROUP_TOOLS,
order=243,
families=("image",),
hint="Appears when image generation is offered. The sentence about the "
"picture already being on screen is the one that earns its place: "
"without it the commonest thing a model does next is offer to show you "
"the image, which it has no way of doing and which has already "
"happened.",
hint="Appears when image generation is offered. Two sentences here earn "
"their place against the tool's own descriptions. The picture already "
"being on screen, because without it the commonest thing a model does "
"next is offer to show you the image which it cannot do and which has "
"already happened. And the shape of a prompt: a small model left to "
"itself passes the request through verbatim, which is why so many "
"generations look like nobody thought about them.",
default=(
"- You can draw a picture with image_generate. Describe what you want in "
"the prompt as fully as you can — subject, setting, lighting, style — "
"because the prompt is the whole of what the picture is made from.\n"
"- The picture appears in the conversation as soon as the tool returns. "
"It is already on screen: do not offer to show it, link to it or "
"describe how to open it.\n"
"- Only the prompt is required. Everything else has a sensible default, "
"so set a parameter when you have a reason to and leave it out "
"otherwise. Repeat a seed to get the same picture again."
"- You can draw a picture with image_generate. Only `prompt` is required.\n"
"- Write the prompt as a description, not as the request you were given. "
"Comma-separated phrases work better than a sentence, and the order matters "
"— subject first, then what it is doing, then the setting, then the light, "
"then the style and medium. \"a red bicycle\" is a worse prompt than \"a red "
"bicycle leaning on a whitewashed wall, morning light, long shadows, 35mm "
"photograph, shallow depth of field\". Expand what you were asked for into "
"one of these; do not ask the person to write it for you.\n"
"- Use `negative` for what must not appear, as plain nouns: \"blurry, extra "
"fingers, text, watermark\". Never phrase it as an instruction — \"no text\" "
"puts text in the picture.\n"
"- Set `width` and `height` to suit the subject rather than leaving both at "
"the default: taller than wide for a person, wider than tall for a place. "
"Match the size the checkpoint expects; far above it produces duplicated "
"limbs rather than more detail.\n"
"- The other parameters have sensible defaults. Change one when you have a "
"reason — fewer steps for a quick draft, lower cfg when a picture looks "
"harsh — and leave it out otherwise.\n"
"- The picture appears in the conversation as soon as the tool returns. It "
"is already on screen: do not offer to show it, link to it, or describe how "
"to open it. Say what you made and what you would change.\n"
"- If it fails because the machine ran out of video memory, try once more at "
"a smaller size or with a lighter checkpoint. Do not repeat the same request "
"unchanged."
),
),
Fragment(