A crowd in one chat

The chat's own model answers, then each other member in order, then the order runs
backwards asking each whether it disagrees, ending at the main model, which either
closes or sends them round again. Design and reasoning: LLeMbas.wiki/Crowd-chats.

THE SPEAKER SEAM, WHICH IS ALSO A BUG FIX

`chat_service.speaker_for` makes the *message* name the answering model and the
chat only the default. That closes a live half-wired feature -- `wake_chat` takes a
model override and `schedule/runner` passes one, and it reached the row and never
the request, so a schedule naming another model got the chat's model wearing the
other one's name.

The seam is wider than `build_request`: `{{model_name}}`, the authored prompt's
model layer, `vision` (where a wrong answer makes the endpoint reject the whole
request), the effort vocabulary (which raises inside the model's own chat template,
and whose refusal narrows every Model row sharing the id), `resolve_tools`,
`context_length` -> `_too_big`, and `ToolContext.model_id`. `resolve_endpoint` may
now only write back `chat.connection_id` when the speaker *is* the chat's model.

WHY N CHAINED REPLIES

`Generation` is one reply's state and `_follow` streams per message, so one
generation cannot stream into nine bubbles and `ensure` would not know which of the
nine it was after a restart. A subagent per speaker cannot work either: its answer
comes back as a tool result and tool results are never replayed, so speaker 3 could
not see speaker 2 -- which is the whole point. Chained, exactly one incomplete row
exists at a time, and `tests/test_crowd_chain.py` asserts that at every
observation.

The round lives on `Message.crowd_json`, not on the chat: the row is the authority,
and chat-level state would describe turns a rewind or a restart had removed.
`crowd.next_turn` is pure, so all eight refusals are tested with no endpoint.

THREE RULES, EACH A BUG WRITTEN THE OTHER WAY ROUND

- `if not _advance_crowd(g): _drain(g)` -- advancing must *suppress* draining, or a
  queued human turn puts a second incomplete row beside the next speaker's.
- `_advance_crowd` refuses unless the finishing row is the newest, or regenerating
  member 2 creates a second member 3 and two chains race down one turn.
- an error skips one speaker and two in a row end the round: the usual failure is a
  small member's window overflowing, and `_drain`'s stop-on-error would kill every
  crowd at whichever member is smallest.

Each other speaker's turn is relabelled as attributed user content, which is both
how a model can disagree with words it did not write and how the history keeps
alternating. The per-speaker instruction is payload-only -- as a row it could be
dropped from the request by a `created_at` tie, and every later speaker would answer
it. Compaction, titling and the notification are gated to once per turn; `_inject`
is off during a round; the way back gets no tools and a member is treated as
unattended.

Membership stores the model as text with no foreign key: "Test & refresh" deletes
and recreates Model rows, and a cascade would empty the crowd out of every chat.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-26 13:38:52 +00:00
co-authored by Claude Opus 5
parent ac51dd46cc
commit da0797ccad
27 changed files with 3018 additions and 59 deletions
+85 -6
View File
@@ -162,6 +162,15 @@ FAMILY_FRIEND = "friend"
# a model up: either it may form and keep opinions of this kind or it may not.
FAMILY_PERSONA = "persona"
# Sending a crowd round again. Its own family so `harness._families` can map the
# name back to one, and deliberately **not in `FAMILIES`**: that tuple is the list
# of things an administrator switches on, and this is mechanism. Being in it would
# mint a `tool_crowd` capability checkbox and demand a `tools.crowd` permission
# that does not exist -- which, because `_family_allowed` falls through to
# `allowed.get(...)`, would mean the tool could never be offered at all. Its real
# gate is `resolve_tools(crowd_again=…)`: one turn of one round.
FAMILY_CROWD = "crowd"
# The built-in families, in the order they are offered.
FAMILIES = (
FAMILY_SEARCH,
@@ -1635,6 +1644,13 @@ def _family_allowed(
# separate switch would be a second door to the cost with nothing
# naming it. `Helpers` on /admin/agents is where both are bounded.
return bool(allowed.get("tools.friend") and subagents)
if gate == FAMILY_CROWD:
# Always allowed, because whether it is *offered* is decided before this:
# `resolve_tools` puts it in the book only on the main model's closing turn
# with a round still left. A permission here would be a second switch for
# one already-enabled feature, and an absent one would silently make the
# crowd a single round for ever.
return True
if gate in (
FAMILY_CUSTOM,
FAMILY_MCP,
@@ -1721,6 +1737,13 @@ def _friend_defs() -> list[ToolDef]:
return subagent_service.friend_tool_defs()
def _crowd_defs() -> list[ToolDef]:
"""The go-round-again tool. Imported inside the call for the reason above."""
from lembas.services import crowd as crowd_service
return crowd_service.tool_defs()
def _image_defs(db: DBSession, values: dict | None = None) -> list[ToolDef]:
"""The image tool, whose schema carries this instance's own choices.
@@ -1776,6 +1799,7 @@ def registry(db: DBSession) -> dict[str, ToolDef]:
*_schedule_defs(),
*_subagent_defs(),
*_friend_defs(),
*_crowd_defs(),
]
)
@@ -1786,13 +1810,31 @@ def families(db: DBSession) -> tuple[str, ...]:
return (*FAMILIES, *rows)
def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
"""Every tool this chat may call right now, with its runner attached."""
def resolve_tools(
db: DBSession,
chat: Chat,
user: User | None,
speaker=None,
*,
crowd_turn=None,
crowd_again: bool = False,
) -> ToolSet:
"""Every tool this chat may call right now, with its runner attached.
The capabilities are the **answering** model's. `tools` being off is the first
gate and returns nothing at all, so handing a crowd member the main model's
switches would offer a tool list to an endpoint that rejects the request for
carrying one.
"""
from lembas.security import permissions
from lembas.services import chat as chat_service
capabilities = {}
model = chat_service.model_for(db, chat)
model = (
chat_service.model_row(db, speaker)
if speaker is not None
else chat_service.model_for(db, chat)
)
if model is not None:
capabilities = model.capabilities_json or {}
@@ -1819,6 +1861,12 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
*(_schedule_defs() if schedules_on else []),
*(_subagent_defs() if subagents_on else []),
*(_friend_defs() if subagents_on else []),
# Only on the closing turn, and only with a round left. Not gated on a
# capability or a permission: a tool that exists on exactly one turn of
# one feature is mechanism, and an administrator switching it off would
# be switching off the main model's ability to use the feature it
# already enabled.
*(_crowd_defs() if crowd_again else []),
]
)
@@ -1832,6 +1880,26 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
off = scoped_off(chat)
empty_library = not skills_service.count_enabled(db, user, exclude=scoped_skills_off(chat))
# What a crowd speaker may do, which is narrower than what the chat may.
if crowd_turn is not None:
from lembas.services import crowd as crowd_service
if crowd_turn.phase == crowd_service.PHASE_BACK:
# The way back is "do you disagree with any of this", which needs
# nothing looked up: everything it is about is already in front of it.
# An empty toolset also guarantees the turn ends in words, which is the
# shape `_wrap_up` relies on.
return ToolSet()
if not crowd_turn.is_main:
# A member answers a machine-composed instruction with several models'
# words quoted into it, and nobody is waiting on *it* in particular.
# So: it cannot stop the round for an approval or a question -- one
# card would park every remaining speaker for `approval_timeout` -- it
# cannot fan out, and it cannot rewrite a personality under wording it
# did not choose. The same set `unattended` withdraws, for the same
# reasons, applied for a different one.
off = off | {FAMILY_ASK, FAMILY_SUBAGENT, FAMILY_FRIEND, FAMILY_PERSONA}
# A scheduled task runs with nobody present, so `ask_user` cannot work here:
# it pauses the reply and waits for a POST that will never come, until
# `approval_timeout` expires -- a run that silently does nothing for fifteen
@@ -2017,10 +2085,21 @@ def context_for(
chat: Chat | None = None,
*,
tools: ToolSet | None = None,
speaker=None,
) -> ToolContext:
"""The snapshot a running tool needs, taken while the session is open."""
"""The snapshot a running tool needs, taken while the session is open.
`speaker` is the model answering, and it decides which model a tool acts *as*:
which personality `persona_write` rewrites, and whose endpoint the image
reviewer and the Preserve-VRAM unload reach for. It defaults to the chat's own
model.
"""
from lembas.services import chat as chat_service
from lembas.services.agent import session as agent_session
if chat is not None and speaker is None:
speaker = chat_service.speaker_for(db, chat)
return ToolContext(
agent=agent_session.resolve(db, chat, user) if chat is not None else None,
owner_id=user.id if user else "",
@@ -2029,8 +2108,8 @@ def context_for(
image_config=settings_store.images(db),
image_workflow_id=(chat.image_workflow_id or "") if chat is not None else "",
image_checkpoint=(chat.image_checkpoint or "") if chat is not None else "",
model_id=(chat.model_id or "") if chat is not None else "",
connection_id=(chat.connection_id or "") if chat is not None else "",
model_id=(speaker.model_id or "") if speaker is not None else "",
connection_id=(speaker.connection_id or "") if speaker is not None else "",
base_ids=[base.id for base in chat.knowledge_bases] if chat is not None else [],
skills_off=scoped_skills_off(chat),
tools=tools.by_name if tools is not None else None,