Work handed to a second model, which may not ask

subagent_run gives a self-contained piece of work to a helper carrying the
parent's connection, directory, model and effort, and hands its answer back as
the tool result. The mechanism is the one scheduled runs already use -- a hidden
chat, one turn, wake_chat, and a poll -- so tools, rounds, budgets, metrics and
steps all work with no second implementation. The two alternatives were
rejected where they had already been rejected once: a nested Generation is two
replies writing one transcript, and a one-shot complete() has no tools, which
schedule/runner.py records as useless for exactly this case.

Every restriction is a property of the child's row, applied by resolve_tools
after the gates, because a rule that lives in a system message is one a page the
model just read can argue with. No questions, no recursion, nothing that writes
unless the call asked for it and the parent's own mode would not have stopped
first, and commands only from a fixed read-only list -- in every mode including
Auto, because the task text can have come from a page.

Withdrawing ask_user turned out to be half of "nobody is watching". An approval
still built a card nobody could see and parked the reply until approval_timeout,
which from every screen is the feature not working. Chat.unattended is the
question now, and not the kind: _authorise answers with a refusal instead. A
scheduled task's chat had the same hole and is covered by the same flag.

Three bounds, counted where each is knowable: per reply on the parent's
Generation, instance-wide in a set a restart clears, and per helper in settings
of its own so one runs out of room long before the reply that asked. Past the
clock the helper is stopped rather than abandoned, so a partial answer comes
back with a sentence saying so.

Also: four gates had shipped into the scope menu with no name, taking the first
tool's label instead -- the canvas switch read "Canvas written". There is a test
that refuses a family without one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jaroslav Beneš
2026-08-06 15:05:30 +02:00
co-authored by Claude Opus 5
parent 0fa05c88b2
commit 46066150d9
17 changed files with 1850 additions and 13 deletions
+96 -4
View File
@@ -137,6 +137,13 @@ FAMILY_REPORT = "report"
# note, said it had scheduled something, and nothing anywhere disagreed.
FAMILY_SCHEDULE = "schedule"
# Handing a self-contained piece of work to a second model that runs on its own
# and reports back. Its own family because it is the one tool whose cost is
# another whole reply -- an instance may reasonably offer everything else and
# not this, and on a single local endpoint four helpers at once is four times
# the queue rather than four times the speed.
FAMILY_SUBAGENT = "subagent"
# The built-in families, in the order they are offered.
FAMILIES = (
FAMILY_SEARCH,
@@ -150,6 +157,7 @@ FAMILIES = (
FAMILY_IMAGE,
FAMILY_REPORT,
FAMILY_SCHEDULE,
FAMILY_SUBAGENT,
FAMILY_AGENT,
)
@@ -209,6 +217,14 @@ class ToolContext:
# something. Read from the instance settings while the session was open,
# like everything else here.
interaction_timeout: float = 900.0
# Whether there is anybody who could answer. False for an ordinary chat;
# true for a scheduled task's and a subagent's. `ask_user` is already
# withdrawn when it is set, so what this reaches is `_authorise`, which
# answers an approval with a refusal instead of pausing on a card nobody
# can see. Without it the reply stalls for `interaction_timeout` and then
# gives up having done nothing -- which is the failure the withdrawal was
# added to prevent, arriving by the other door.
unattended: bool = False
# Set only for an agent chat: the machine to act on, the mode in force, and
# the decrypted credential. None everywhere else, which is what every agent
# runner checks first. `generation` clears it when the reply ends.
@@ -1301,6 +1317,7 @@ def _family_allowed(
allowed: dict,
images: bool = False,
schedules: bool = False,
subagents: bool = False,
) -> bool:
"""Whether one family is on for this chat.
@@ -1342,6 +1359,14 @@ def _family_allowed(
# offer this at all, or a model spends a round being told the tool it
# was handed does not work.
return bool(allowed.get("schedule.use") and schedules)
if gate == FAMILY_SUBAGENT:
# Its own permission and its own instance switch, for the reason the
# image tool has both: what this costs is a second reply, which is not
# a cost the tools around it have, and an instance whose endpoint is one
# local card has a real reason to say no. `subagents` is passed in
# rather than read here so that the whole gate is answered from the
# snapshot `resolve_tools` already took.
return bool(allowed.get("tools.subagent") and subagents)
if gate in (
FAMILY_CUSTOM,
FAMILY_MCP,
@@ -1412,6 +1437,13 @@ def _schedule_defs() -> list[ToolDef]:
return schedule_tool.tool_defs()
def _subagent_defs() -> list[ToolDef]:
"""The subagent tool. Imported inside the call for the reason above."""
from lembas.services import subagent as subagent_service
return subagent_service.tool_defs()
def _image_defs(db: DBSession, values: dict | None = None) -> list[ToolDef]:
"""The image tool, whose schema carries this instance's own choices.
@@ -1454,7 +1486,6 @@ def registry(db: DBSession) -> dict[str, ToolDef]:
working on. The same omission cost custom tools their guidance once already.
"""
from lembas.services.agent import tools as agent_tools
from lembas.services.schedule import tool as schedule_tool
return _book(
[
@@ -1465,7 +1496,8 @@ def registry(db: DBSession) -> dict[str, ToolDef]:
# `schedule_create` back to a family and the guidance never
# reaches the model. That omission has cost two features their
# instructions already.
*schedule_tool.tool_defs(),
*_schedule_defs(),
*_subagent_defs(),
]
)
@@ -1494,6 +1526,7 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
image_values = settings_store.images(db)
images_ready = settings_store.images_ready(db)
schedules_on = bool(settings_store.schedules(db).get("enabled"))
subagents_on = bool(settings_store.subagents(db).get("enabled"))
# Resolved against what this reader may see, not against everything that
# exists: a tool restricted to a group is not offered outside it. The image
@@ -1506,6 +1539,7 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
*_agent_defs(db, chat, user),
*(_image_defs(db, image_values) if images_ready else []),
*(_schedule_defs() if schedules_on else []),
*(_subagent_defs() if subagents_on else []),
]
)
@@ -1526,8 +1560,25 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
# merely discouraged in `core.unattended`, because a rule living only in a
# system message is one a page the model just read can argue with. The
# fragment is the half that stops it *planning* around a tool it has not got.
if chat is not None and chat.kind == KIND_TASK:
off = off | {FAMILY_ASK}
#
# A subagent's chat is unattended for a different reason and arrives at the
# same place, which is why the question asked is `unattended` and not the
# kind: it is also where the *recursion* stops. A helper that could spawn a
# helper is a fan-out with no bound anybody set.
if unattended(chat):
off = off | {FAMILY_ASK, FAMILY_SUBAGENT}
# Everything that changes something, withheld. Set by `services/subagent.py`
# on the chat it creates and by nothing else, so absent means on exactly as
# every other key here does. Keyed on the tool's declared **risk** rather
# than on a list of names, because a list is a thing that goes out of date
# silently: a tool added next year would default into a read-only helper's
# set unless somebody remembered.
#
# `RISK_EXECUTE` is deliberately not included. In an agent chat it is
# governed by the mode and the allow list instead, which is a finer
# instrument -- `git log` is a read whatever its risk class says.
writes_off = scoped_writes_off(chat)
return ToolSet(
tuple(
@@ -1540,8 +1591,10 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
allowed=allowed,
images=images_ready,
schedules=schedules_on,
subagents=subagents_on,
)
and gate_of(tool.family) not in off
and not (writes_off and tool.risk == RISK_WRITE)
# Nothing to read and nothing to improve. Offering `skill_get` with
# no skills is what makes a model spend a round looking one up and
# being told it does not exist -- and `context.skills` already
@@ -1557,6 +1610,38 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
_NEEDS_A_SKILL = frozenset({"skill_get", "skill_edit"})
def unattended(chat: Chat | None) -> bool:
"""Whether there is anybody who could answer a question in this chat.
Two things make a chat unattended and they are not the same fact. A
scheduled task's chat is one because of what starts it; a subagent's is one
because of what it *is*. `Chat.unattended` is the column both now set, and
the kind is still consulted beside it because the column was added to a
table that already had task chats in it -- `sync_schema` backfills a new
NOT NULL column with its type default, so every task chat written before
this reads back as attended. Dropping the kind check would silently give
every existing scheduled task a tool that stalls it for fifteen minutes.
"""
if chat is None:
return False
return bool(getattr(chat, "unattended", False)) or chat.kind == KIND_TASK
def scoped_writes_off(chat: Chat | None) -> bool:
"""Whether this chat has had everything that changes something withdrawn.
One key rather than a family list, because "may not write" is a property of
the *conversation* and not of any one gate: a read-only helper must not
write a note, file a report, save a memory or edit a file, and those are
four gates it would otherwise have to name — and a fifth would arrive
unnamed. Absent means writes are on, the same convention as everything else
under `scope_json`.
"""
if chat is None:
return False
return (getattr(chat, "scope_json", None) or {}).get("write") is False
def scoped_off(chat: Chat | None) -> frozenset[str]:
"""Gates this chat has switched off. **Absent means on**, always.
@@ -1595,6 +1680,12 @@ def scoped_allow(chat: Chat | None) -> tuple[str, ...]:
compared, and it refuses to produce anything at all for a command line
carrying a shell metacharacter. "Always" can therefore only ever mean "this
exact thing again".
There is a second server-side writer now: `services/subagent.py` puts its
fixed safe list here when it creates a helper's chat. That does not weaken
the property above -- the list is a constant in this codebase, the chat is
created here and never by a request, and the model asking for the helper
chooses none of it.
"""
if chat is None:
return ()
@@ -1637,6 +1728,7 @@ def context_for(
skills_off=scoped_skills_off(chat),
tools=tools.by_name if tools is not None else None,
interaction_timeout=float(settings_store.agents(db)["approval_timeout"]),
unattended=unattended(chat),
)