The last unbuilt capability, and built the way CLAUDE.md said it had to be: a
ToolDef reaching resolve_tools plus a permission and a capability flag, not a new
code path. The only genuinely new UI is one branch in the transcript.
services/images/ is three modules. comfy.py speaks HTTP -- submit, poll /history,
fetch the PNG, /free, and an /object_info discovery for the admin page only.
Polled and not socketed, because holding a connection open for the length of a
generation is the live-connection state the whole ssh.py design forbids, and the
thing being waited for takes tens of seconds anyway. The base URL is exempt from
the SSRF guard by construction, exactly as Connection.base_url and the audio
endpoints are -- said out loud in the docstring, because a default of
127.0.0.1:8188 is precisely the shape that guard exists to refuse and therefore
reads as a hole rather than a decision.
workflow.py fills a template, and the one thing that matters is that it walks the
parsed JSON rather than the text of it. A value that is exactly "{{steps}}"
becomes the number 20; ComfyUI validates types and refuses the string. A
placeholder inside a longer string is still text, which is what makes
"{{prompt}}, masterpiece" work -- and text substitution would additionally mean a
prompt containing a quotation mark produced a document that no longer parses, on
the one input guaranteed to hold arbitrary text. Which node holds the prompt is
the administrator's statement rather than a guess from node types: sniffing for
the first CLIPTextEncode works on the shipped workflow and on nothing else, and
swaps positive for negative the first time somebody reorders them. seed has no
fixed default, because one would make every unspecified generation identical and
make the retry loop redraw the same rejected picture four times.
tool.py is one call, one finished image. Returning every attempt to the
conversation would cost a round each, make the ceiling advisory rather than
enforced, and walk the reader past every reject -- so the reviewer lives inside
the tool and is asked about *bytes*: an attempt about to be discarded should not
leave an Attachment behind, so it sees a downscaled preview built in memory and
only the kept image is written. Anything that goes wrong in review is a keep;
losing a picture because a judging request timed out would be the check
destroying the thing it was checking. The last attempt is kept whatever the
verdict, so a request always produces something. Rejects are recorded, not
stored.
Preserve VRAM unloads the chat's own connection and nothing else, because the
memory being freed belongs to one machine: local llama-swap answers GET /unload,
and a box on the network has no reason to be unloaded when ComfyUI wants memory
here. The swap goes round the review rather than round the tool, which costs two
model loads per retry -- so the two settings are independent and the page warns
when both are on. Nothing loads the LLM back: the reply's next request does, and
that step exists in the description and not in the code, so the code says so.
Two rules elsewhere had to be drawn for the first time. message_payload sends
images only on user turns -- no assistant message had ever carried one, and the
moment one does the multimodal list form on an assistant turn is rejected by
OpenAI and most local runners, breaking every later turn in the chat. And
files.store gained keep_original, because _process_image turns anything without
alpha into JPEG q85 at 1400px: right for a phone photo, a visible loss on the one
output this feature exists to produce.
/image sends the ordinary message with force_tool, which becomes tool_choice for
the first round only -- left in place the reply would draw a picture, be asked
again, and draw another. FORCEABLE_TOOLS is an allow list because the name is
read off a form.
ToolContext gained chat_id, and that fixed a tool nobody had ever successfully
run: _run_scratch_write read context.chat_id on a dataclass with no such field,
so every call raised AttributeError, swallowed by run_tool's blanket except into
"the scratch_write tool failed" -- indistinguishable from a model calling it
wrongly. The test that existed asserted the family and the risk, which are
properties of the declaration rather than of the code.
Verified against the real ComfyUI 0.27.0 on this machine rather than against
documentation: every endpoint shape here was read off it, a generation ran end to
end through the client, the reviewer was shown a matching and a mismatched prompt
and answered KEEP and RETRY correctly, and the unload hook fired for the local
llama-swap and not for the remote box.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
160 lines
6.6 KiB
Python
160 lines
6.6 KiB
Python
"""OpenAI-compatible endpoint connections and their discovered models."""
|
|
|
|
from __future__ import annotations
|
|
|
|
from datetime import datetime
|
|
from typing import TYPE_CHECKING, Any
|
|
|
|
from sqlalchemy import (
|
|
Boolean,
|
|
Column,
|
|
DateTime,
|
|
ForeignKey,
|
|
Integer,
|
|
String,
|
|
Table,
|
|
Text,
|
|
UniqueConstraint,
|
|
)
|
|
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
|
|
|
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
|
from lembas.db.types import JSONDict
|
|
|
|
if TYPE_CHECKING:
|
|
# Import only for the annotation; at runtime SQLAlchemy resolves the
|
|
# name through its own class registry, so there is no import cycle.
|
|
from lembas.db.models.user import Group
|
|
|
|
# Which groups may use a given model. A model with no rows here is reachable
|
|
# only by administrators unless it is marked public.
|
|
model_groups = Table(
|
|
"model_groups",
|
|
Base.metadata,
|
|
Column("model_id", String(32), ForeignKey("models.id", ondelete="CASCADE"), primary_key=True),
|
|
Column("group_id", String(32), ForeignKey("groups.id", ondelete="CASCADE"), primary_key=True),
|
|
)
|
|
|
|
|
|
class Connection(UUIDPrimaryKey, Timestamps, Base):
|
|
"""A configured upstream endpoint speaking the OpenAI HTTP API.
|
|
|
|
Works for api.openai.com as well as LM Studio, vLLM, llama.cpp, Ollama's
|
|
compatibility layer, OpenRouter, and anything else exposing /v1.
|
|
"""
|
|
|
|
__tablename__ = "connections"
|
|
|
|
name: Mapped[str] = mapped_column(String(120), nullable=False)
|
|
base_url: Mapped[str] = mapped_column(String(500), nullable=False)
|
|
|
|
# Fernet ciphertext, never the raw key. See lembas.services.crypto.
|
|
# Empty string is legitimate: local endpoints often need no auth at all.
|
|
api_key_encrypted: Mapped[str] = mapped_column(Text, default="")
|
|
|
|
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
|
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
|
|
|
# Extra headers merged into every request (e.g. OpenRouter's HTTP-Referer).
|
|
extra_headers_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
|
|
|
# How to ask this endpoint to drop its model from memory, for the Preserve
|
|
# VRAM option in image generation. Per connection and not instance-wide,
|
|
# because the VRAM being freed is a particular machine's: llama-swap on this
|
|
# host answers `GET /unload`, while a remote vLLM has no such call and no
|
|
# reason to be unloaded when ComfyUI needs memory *here*.
|
|
#
|
|
# Empty means "this connection cannot be unloaded", which is the honest
|
|
# default -- there is no call that works everywhere, and guessing one would
|
|
# send an unexplained request to somebody's endpoint.
|
|
unload_url: Mapped[str] = mapped_column(String(500), default="")
|
|
unload_method: Mapped[str] = mapped_column(String(8), default="POST")
|
|
|
|
# Result of the most recent "Test & refresh", surfaced in the admin list.
|
|
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
|
last_error: Mapped[str] = mapped_column(Text, default="")
|
|
|
|
models: Mapped[list[Model]] = relationship(
|
|
back_populates="connection",
|
|
cascade="all, delete-orphan",
|
|
order_by="Model.model_id",
|
|
)
|
|
|
|
def __repr__(self) -> str:
|
|
return f"<Connection {self.name} {self.base_url}>"
|
|
|
|
|
|
class Model(UUIDPrimaryKey, Timestamps, Base):
|
|
"""A model advertised by a connection, cached locally.
|
|
|
|
Cached rather than fetched live so the chat UI stays responsive and keeps
|
|
working when an endpoint is briefly unreachable. Refreshed on demand from
|
|
the admin screen.
|
|
"""
|
|
|
|
__tablename__ = "models"
|
|
__table_args__ = (UniqueConstraint("connection_id", "model_id"),)
|
|
|
|
connection_id: Mapped[str] = mapped_column(
|
|
String(32), ForeignKey("connections.id", ondelete="CASCADE"), nullable=False, index=True
|
|
)
|
|
model_id: Mapped[str] = mapped_column(String(300), nullable=False)
|
|
display_name: Mapped[str] = mapped_column(String(300), default="")
|
|
description: Mapped[str] = mapped_column(Text, default="")
|
|
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
|
|
|
# Sort order in every picker. Ties fall back to model_id so the order is
|
|
# stable rather than whatever SQLite feels like today.
|
|
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
|
# Pinned models are offered first, before the full list.
|
|
pinned: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
|
|
|
# Public models are usable by anyone; otherwise access comes from `groups`.
|
|
public: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
|
|
|
# Filename under <data>/uploads/models. Stored rather than a URL so the
|
|
# image cannot become a request to a third party on every page render.
|
|
image_path: Mapped[str] = mapped_column(String(300), default="")
|
|
|
|
# Applied to chats using this model when the chat has none of its own.
|
|
# See services.chat.effective_system_prompt for the precedence.
|
|
system_prompt: Mapped[str] = mapped_column(Text, default="")
|
|
|
|
# Endpoints do not reliably advertise capabilities, so these are admin
|
|
# overrides. Recognised keys: vision, tools, reasoning.
|
|
capabilities_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
|
# Default sampling params applied to new chats using this model.
|
|
params_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
|
|
|
# How many tokens this model can hold. 0 means unknown, which is what an
|
|
# endpoint that does not advertise it leaves behind -- and unknown has to
|
|
# stay tellable from "small", because the context percentage and automatic
|
|
# compaction both refuse to act on a number nobody supplied.
|
|
#
|
|
# A column rather than a key in capabilities_json: that dict is rebuilt
|
|
# wholesale from the submitted checkboxes on every save (api/admin_models.py),
|
|
# so a number living in it would be destroyed the next time an administrator
|
|
# ticked anything.
|
|
context_length: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
|
|
|
connection: Mapped[Connection] = relationship(back_populates="models")
|
|
groups: Mapped[list[Group]] = relationship(
|
|
"Group", secondary=model_groups, back_populates="models"
|
|
)
|
|
|
|
@property
|
|
def label(self) -> str:
|
|
return self.display_name or self.model_id
|
|
|
|
@property
|
|
def supports_reasoning(self) -> bool:
|
|
return bool((self.capabilities_json or {}).get("reasoning"))
|
|
|
|
@property
|
|
def initial(self) -> str:
|
|
"""First character of the label, for the fallback avatar."""
|
|
return (self.label.strip() or "?")[0].upper()
|
|
|
|
def __repr__(self) -> str:
|
|
return f"<Model {self.model_id}>"
|