Say what actually failed, and tell the model how to use the thing
Two problems, both found by looking rather than by guessing. ComfyUI writes its history entry in task_done and nowhere else, so the entry appearing IS "finished" -- but it sets completed=e.success, which means an out-of-memory, a cancelled job and a broken node all stay completed:false for ever. await_images waited on that flag. So every failure sat for the full 600s timeout and then reported a timeout, when ComfyUI had known within one second and written down the node, the exception type and the message. Proved by causing both against the real instance: an OOM now raises in 1.0s and an interrupt in 4.0s, each naming the node. The terminal condition is a record with a status, and status.messages is read for the last execution_error or execution_interrupted. OutOfMemory and Interrupted are their own classes because they are the two failures with an obvious next move: the first tells the model to retry at a named smaller size -- worked out from what it actually asked for, since "use a lower resolution" against a request that was already 512x512 is advice nobody can follow -- or with a lighter checkpoint; the second says somebody pressed stop, so do not simply start again. Everything else gets the reason and no advice, because a model told to try again after a broken workflow tries the identical thing. The OOM message is cut to its first sentence. The rest is allocator advice -- PYTORCH_CUDA_ALLOC_CONF, fragmentation notes -- addressed to whoever runs the box and meaningless to a model, in a tool result that is already a failure. Second: the parameters were described in the register of a reference table, and "cfg: prompt adherence, default 8" tells a model nothing it can act on. Measured on a 4B model, same request, same everything else: with the old wording it sent prompt and template and nothing more -- so 512x512 on an SDXL checkpoint, which is exactly the duplicated-limbs failure the width description now warns about. With descriptions that say what each value does to the picture and when to move it, the same model sent a portrait 1024x1536 and a deliberate sampler. ~3KB of schema per request in a chat that can draw, and the difference between having ten parameters and having one. docs/image-generation-instructions.md is the long version for the admin instructions box, for models that need more than the harness can afford to carry on every request in every chat. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
b2a05e0351
commit
178742501d
@@ -93,6 +93,29 @@ class ComfyError(LLMError):
|
||||
"""Anything that stopped a generation, in words worth showing somebody."""
|
||||
|
||||
|
||||
class OutOfMemory(ComfyError):
|
||||
"""The far side ran out of VRAM.
|
||||
|
||||
Its own class because it is the one failure with an obvious next move --
|
||||
a smaller picture, or a smaller checkpoint -- and the model is told to make
|
||||
it. Everything else is reported and stopped at.
|
||||
"""
|
||||
|
||||
|
||||
class Interrupted(ComfyError):
|
||||
"""Somebody cancelled it from ComfyUI's own interface, or it was stopped.
|
||||
|
||||
Distinct because it is not a fault: retrying is reasonable, and "the
|
||||
workflow failed" would be describing a decision as a breakage.
|
||||
"""
|
||||
|
||||
|
||||
# What `exception_type` looks like when a GPU has run out. Matched on the type
|
||||
# rather than on the message, which is a paragraph of allocator advice written
|
||||
# for whoever is running the box and not for a model.
|
||||
_OOM_TYPES = ("outofmemory", "out_of_memory", "cuda error: out of memory")
|
||||
|
||||
|
||||
def _transport_error(exc: httpx.RequestError, config: Config) -> ComfyError:
|
||||
"""The `wrap_transport_error` shape, said about ComfyUI rather than an LLM.
|
||||
|
||||
@@ -133,9 +156,7 @@ async def submit(config: Config, workflow: dict[str, Any]) -> str:
|
||||
body = {"prompt": workflow, "client_id": uuid.uuid4().hex}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=60.0) as client:
|
||||
response = await client.post(
|
||||
config.url("prompt"), headers=config.headers(), json=body
|
||||
)
|
||||
response = await client.post(config.url("prompt"), headers=config.headers(), json=body)
|
||||
if response.status_code >= 400:
|
||||
raise ComfyError(_refusal(response))
|
||||
data = response.json()
|
||||
@@ -185,19 +206,24 @@ def _describe_nodes(errors: dict[str, Any]) -> str:
|
||||
async def await_images(config: Config, prompt_id: str) -> list[Ref]:
|
||||
"""Wait for one queued workflow and answer with what it saved.
|
||||
|
||||
`/history/{id}` is empty while the job is queued or running and gains the
|
||||
whole record when it ends, so an empty answer is "not yet" rather than
|
||||
"nothing" -- which is why the deadline is the only thing that ends this.
|
||||
**The record existing is what "finished" means, not `status.completed`.**
|
||||
ComfyUI writes the history entry in `task_done` and nowhere else, so it
|
||||
appears exactly once the job is over -- but it sets `completed=e.success`,
|
||||
so a run that failed is `completed: false` for ever. Waiting on that flag
|
||||
means every out-of-memory, every cancelled job and every broken node hangs
|
||||
the reply for the whole timeout and then reports a timeout, when ComfyUI
|
||||
knew what was wrong within seconds and said so.
|
||||
|
||||
So: no record means not yet, a record means done, and `status_str` says
|
||||
which kind of done.
|
||||
"""
|
||||
deadline = time.monotonic() + config.timeout
|
||||
while True:
|
||||
record = (await _get_json(config, f"history/{prompt_id}")).get(prompt_id)
|
||||
if isinstance(record, dict) and (record.get("status") or {}).get("completed"):
|
||||
if isinstance(record, dict) and record.get("status") is not None:
|
||||
status = record.get("status") or {}
|
||||
if status.get("status_str") not in (None, "success"):
|
||||
raise ComfyError(
|
||||
f"ComfyUI could not finish the workflow ({status.get('status_str')})."
|
||||
)
|
||||
if status.get("status_str") != "success":
|
||||
raise _failure(status)
|
||||
return _refs_in(record.get("outputs") or {})
|
||||
if time.monotonic() > deadline:
|
||||
raise ComfyError(
|
||||
@@ -207,6 +233,51 @@ async def await_images(config: Config, prompt_id: str) -> list[Ref]:
|
||||
await asyncio.sleep(POLL_INTERVAL)
|
||||
|
||||
|
||||
def _failure(status: dict[str, Any]) -> ComfyError:
|
||||
"""Why a workflow stopped, out of the messages ComfyUI recorded against it.
|
||||
|
||||
`status.messages` is a list of `[name, payload]` pairs -- the lifecycle of
|
||||
the run. The last `execution_error` or `execution_interrupted` in it is the
|
||||
thing that ended it, and carries the node and the exception. Without reading
|
||||
these the only thing that could be said is "error", which is what ComfyUI's
|
||||
own status string amounts to.
|
||||
"""
|
||||
event, payload = "", {}
|
||||
for entry in status.get("messages") or []:
|
||||
if isinstance(entry, list | tuple) and len(entry) == 2:
|
||||
name, body = entry
|
||||
if name in ("execution_error", "execution_interrupted"):
|
||||
event, payload = str(name), body if isinstance(body, dict) else {}
|
||||
|
||||
node = str(payload.get("node_type") or "").strip()
|
||||
where = f" in {node}" if node else ""
|
||||
|
||||
if event == "execution_interrupted":
|
||||
return Interrupted(f"The image was cancelled on the ComfyUI side{where}.")
|
||||
|
||||
kind = str(payload.get("exception_type") or "")
|
||||
detail = _first_sentence(str(payload.get("exception_message") or ""))
|
||||
if any(marker in kind.lower() for marker in _OOM_TYPES) or "out of memory" in detail.lower():
|
||||
return OutOfMemory(f"ComfyUI ran out of video memory{where}. {detail}".strip())
|
||||
if not detail and not kind:
|
||||
return ComfyError(f"ComfyUI could not finish the workflow{where}.")
|
||||
return ComfyError(f"ComfyUI could not finish the workflow{where}: {detail or kind}")
|
||||
|
||||
|
||||
def _first_sentence(message: str) -> str:
|
||||
"""Enough of an exception to act on, and no more.
|
||||
|
||||
A torch OOM runs to several lines of allocator advice -- environment
|
||||
variables to set, fragmentation notes -- addressed to whoever runs the box.
|
||||
None of it means anything to a model, and all of it costs tokens in a tool
|
||||
result that is already a failure.
|
||||
"""
|
||||
first = message.strip().split("\n", 1)[0].strip()
|
||||
if len(first) > 200:
|
||||
first = first[:200].rsplit(" ", 1)[0] + "…"
|
||||
return first
|
||||
|
||||
|
||||
def _refs_in(outputs: dict[str, Any]) -> list[Ref]:
|
||||
"""Every image any node saved, in node order.
|
||||
|
||||
@@ -233,9 +304,7 @@ async def fetch_image(config: Config, ref: Ref) -> bytes:
|
||||
params = {"filename": ref.filename, "subfolder": ref.subfolder, "type": ref.kind}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=120.0) as client:
|
||||
response = await client.get(
|
||||
config.url("view"), headers=config.headers(), params=params
|
||||
)
|
||||
response = await client.get(config.url("view"), headers=config.headers(), params=params)
|
||||
response.raise_for_status()
|
||||
payload = response.content
|
||||
except httpx.HTTPStatusError as exc:
|
||||
|
||||
Reference in New Issue
Block a user