Three features sharing one idea: a model here started from nothing every
conversation and had no notion that anything else existed.
THE ROSTER. `chat.roster_block` builds one line per model this *person* can
reach -- through `permissions.models_visible_to`, never the table -- and
`{{model_roster}}` carries it, gated on the `friend` family for the reason the
memories block is gated on `memory`: a list of peers a model cannot talk to is
context spent on nothing, and one checkbox is then the whole switch. New
`Model.notes` column, a column and not a `capabilities_json` key for the reason
`context_length` and `reasoning_efforts` both carry.
ASKING A FRIEND. A second entry point in `services/subagent.py` rather than a
second module, so one place still owns the bounds and the lifecycle. `_create_
child` takes the friend's (model_id, connection_id) *pair*, because Model is
unique on both and an id alone does not say which endpoint. Three things differ
from a helper: the effort is the friend's own default and never the parent's (the
1.3.0 bug by another door -- the vocabularies differ and a level a model does not
take raises inside its chat template), the chat is ordinary even when the asker's
is an agent chat, and `scope_json["role"]` marks it so `core.friend` speaks
instead of `core.subagent`. `friend` joins the unattended withdrawal set: a
friend that could ask a friend is the same unbounded fan-out in politer clothes.
Budget, concurrency and quota are shared with helpers, so one reply cannot spend
the allowance twice.
PERSONALITY. One table, two roles, `owner_id IS NULL` the discriminator: the
model's own persona, and its read of one person. Keyed on the model's *text* id
with no foreign key, because "Test & refresh" deletes a model the endpoint has
stopped listing and a personality must not be collateral. `PersonaRevision`
copies SkillRevision, and so does the argument: the safety story for a model
rewriting itself is a record and a way back, not a gate. The reflection is shown
to the person it is about, in their own settings, which is the whole of why
keeping one is acceptable. `persona` is withdrawn from any unattended chat --
a helper's task, a friend's question and a schedule's instruction are all words
nobody watched being written.
Two bugs found while reading for this, both silent:
`review_model_id` stored a `Model` primary key, so a refresh taken while an
endpoint was not listing that model unset the administrator's choice -- and
`_reviewer` then fell back to the chat's own model, so pictures were judged by
a model nobody chose. Now the text id, with the primary key still accepted.
`_messages_after` used a bare `>` on `created_at`, so a row sharing the edited
turn's microsecond survived a rewind -- and `_send` writes a user turn and its
placeholder back to back, which is exactly that tie. Deliberately NOT
`thread_tail`'s `(created_at, id)` tiebreak: ids are random UUIDs, so that
settles a tie by coin toss. A tie now reads as "later", which is the safe
direction for an operation whose purpose is to discard what follows.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
876 lines
37 KiB
Python
876 lines
37 KiB
Python
"""Handing a piece of a reply to a second model that runs on its own.
|
|
|
|
## What it is
|
|
|
|
`subagent_run` creates a hidden chat, puts one self-contained task into it, lets
|
|
the ordinary generation loop answer it, and gives the answer back as the tool
|
|
result. That is the whole mechanism. Nothing about streaming, rounds, budgets,
|
|
metrics, steps or tools is re-implemented here, because a second implementation
|
|
of any of those is a second thing to keep correct.
|
|
|
|
Two shapes were rejected on the way.
|
|
|
|
A **nested `Generation` in the parent's chat** would mean two replies writing one
|
|
transcript, which `services/wake.py` exists to make impossible: a chat has one
|
|
generation at a time, and the Stop button points at whichever bubble comes first
|
|
in the document.
|
|
|
|
A **one-shot `complete()`** — the shape `generate_title` uses — has no tools and
|
|
no rounds, which `schedule/runner.py` already records as useless for exactly this
|
|
case. A helper that cannot search is not a helper.
|
|
|
|
So the pattern is `runner.fire`'s: `wake_chat`, then poll `running_for` until it
|
|
stops. The child chat is `temporary`, so it is in no listing, and it is deleted
|
|
when its answer has been handed over unless an administrator asked to keep it.
|
|
|
|
## What makes it safe with nobody watching
|
|
|
|
The rule this codebase holds is that restriction happens **at tool resolution,
|
|
never in the prompt** — a model can be talked out of a system message by a page
|
|
it just read, and a subagent's task text is written by a model that has been
|
|
reading pages. So every restriction below is a property of the child's row,
|
|
applied by `resolve_tools` after every gate:
|
|
|
|
- **Nothing may ask.** `Chat.unattended` withdraws `ask_user`, and
|
|
`generation._authorise` answers an approval with a refusal instead of pausing.
|
|
Without the second half a helper in Manual mode would sit on a card nobody can
|
|
see until `approval_timeout`, which is fifteen minutes of doing nothing.
|
|
- **No recursion.** The same flag withdraws the `subagent` family from the child.
|
|
- **No writing, by default.** `scope_json["write"] = False` drops every tool
|
|
whose declared risk is `RISK_WRITE` — notes, memories, reports, skills, the
|
|
canvas, file writes, schedules, images. A *writing* helper is a per-call
|
|
parameter and is refused outright in a chat whose own mode would have stopped
|
|
to ask before writing, because a subagent that writes where its parent had to
|
|
ask is the mode being laundered through a tool call.
|
|
- **Commands from a fixed list only, in every mode including Auto.** The child
|
|
runs in Plan or Edit mode, both of which resolve `RISK_EXECUTE` to ASK, and
|
|
ASK in an unattended chat is a refusal. What runs is what matches
|
|
`SAFE_COMMANDS` — and `policy.subject` refuses to match any line carrying a
|
|
shell metacharacter, so `git log; curl … | sh` matches nothing. Auto is
|
|
deliberately **not** inherited: the task text can have come from a page, and
|
|
that is the injection path this list exists to close.
|
|
- **The credential is copied**, never referenced, for the reason
|
|
`agent/jobs.py:start_watch` gives: the parent's `spec` is cleared when its
|
|
reply ends and a helper outlives nothing but is not entitled to assume so.
|
|
|
|
## What bounds it
|
|
|
|
`settings_store.subagents` — how many one reply may spawn, how many run at once
|
|
across the instance, and what one of them may spend. The per-reply count lives on
|
|
the parent's `Generation`, which is the only object that knows what "this reply"
|
|
means; the instance count is a set here, because a counter that a restart clears
|
|
is the correct shape for a thing that cannot survive a restart anyway.
|
|
|
|
The wall clock is the honest bound. When it runs out the child is **stopped**,
|
|
not abandoned: `request_stop` keeps whatever it had written and marks the message
|
|
`stopped`, so the parent gets a partial answer and a sentence saying it is one,
|
|
rather than silence or a wait that outlives the reply that asked for it.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import logging
|
|
import time
|
|
from typing import TYPE_CHECKING, Any
|
|
|
|
from lembas.db.models import KIND_AGENT, KIND_CHAT, Chat, Model, User
|
|
from lembas.db.session import session_scope
|
|
from lembas.security import permissions
|
|
from lembas.services import settings_store
|
|
from lembas.services.agent import policy as agent_policy
|
|
|
|
if TYPE_CHECKING: # pragma: no cover - typing only
|
|
from lembas.services.tools import ToolContext, ToolDef, ToolOutcome
|
|
|
|
log = logging.getLogger(__name__)
|
|
|
|
# How often the parent looks to see whether its helper has finished. Coarser
|
|
# than a stream and finer than the schedule runner's three seconds: a caller is
|
|
# blocked on this, so a second of latency at the end is worth avoiding, and
|
|
# anything finer is a poll per hundred milliseconds for a minute of work.
|
|
POLL_SECONDS = 1.0
|
|
|
|
# How long to keep asking after `request_stop`, before giving up on the row and
|
|
# reading whatever is there. The producer checks `cancel` between streamed
|
|
# chunks, so a stop lands within one chunk unless the far side has stalled.
|
|
STOP_GRACE = 20.0
|
|
|
|
# What a helper may run on a machine, whatever mode its parent is in.
|
|
#
|
|
# Every entry is a command that reads. There is no `find -delete`, no `git
|
|
# checkout`, no package manager: the list is short because the argument for
|
|
# adding to it is always "this one is fine", and the sum of those is a shell.
|
|
# `policy.subject` normalises whitespace and refuses to match anything holding a
|
|
# shell metacharacter, so none of these can be extended with a `;` or a pipe.
|
|
#
|
|
# `file_read` and `file_list` are here as tool names rather than commands, which
|
|
# is what `subject` returns for anything that is not `shell_run` -- the same
|
|
# entries `agents.allow_default` ships with.
|
|
SAFE_COMMANDS: tuple[str, ...] = (
|
|
"file_read",
|
|
"file_list",
|
|
"ls",
|
|
"ls *",
|
|
"pwd",
|
|
"cat *",
|
|
"head *",
|
|
"tail *",
|
|
"wc *",
|
|
"file *",
|
|
"stat *",
|
|
"du *",
|
|
"df *",
|
|
"tree *",
|
|
"find *",
|
|
"grep *",
|
|
"rg *",
|
|
"git status",
|
|
"git log*",
|
|
"git show*",
|
|
"git diff*",
|
|
"git branch",
|
|
"git remote -v",
|
|
)
|
|
|
|
# Modes a helper may be given, and nothing else. Plan reads; Edit reads and
|
|
# writes files. Neither allows a command outside the list above, because both
|
|
# resolve RISK_EXECUTE to ASK and an unattended chat cannot ask.
|
|
MODE_READING = agent_policy.MODE_PLAN
|
|
MODE_WRITING = agent_policy.MODE_EDIT
|
|
|
|
# The parent modes from which a *writing* helper may be asked for. In Manual and
|
|
# Plan the reader is stopped before anything is written, and a helper that wrote
|
|
# on the model's own authority would be that rule going through a side door.
|
|
WRITING_ALLOWED_FROM = (agent_policy.MODE_EDIT, agent_policy.MODE_AUTO)
|
|
|
|
# What `scope_json["role"]` says on the chat of a model that has been asked a
|
|
# question rather than given a job. A key on the scope and not a column: it is
|
|
# read in one place, to pick which of two sentences the child's own system
|
|
# prompt carries, and `Chat.unattended` already carries every *behavioural*
|
|
# consequence of being somebody's child.
|
|
ROLE_FRIEND = "friend"
|
|
|
|
# Helpers running right now, across the instance, by child chat id. In-process
|
|
# and cleared by a restart, which is correct: a restart abandons replies in
|
|
# flight, so there is nothing for a durable count to describe.
|
|
_LIVE: set[str] = set()
|
|
|
|
|
|
def live_count() -> int:
|
|
return len(_LIVE)
|
|
|
|
|
|
def clear() -> None:
|
|
"""For tests. The set is the only state this module holds."""
|
|
_LIVE.clear()
|
|
|
|
|
|
# --- Building the child --------------------------------------------------------
|
|
def _child_scope(parent: Chat, *, write: bool) -> dict[str, Any]:
|
|
"""What the helper's chat is narrowed to.
|
|
|
|
`families` names the two withdrawals that are absolute; `write` is the
|
|
risk-class narrowing; `allow` is the command list. Everything else the
|
|
parent had, the child has -- searching, fetching, reading the library --
|
|
because a helper that cannot look things up is a slower way of asking the
|
|
same model the same question.
|
|
|
|
The parent's own narrowing is carried across whole. A chat with web search
|
|
switched off must not be able to reach it by delegating.
|
|
"""
|
|
inherited = dict((parent.scope_json or {}).get("families") or {})
|
|
inherited.update({"ask": False, "subagent": False, "friend": False})
|
|
return {
|
|
"families": inherited,
|
|
"skills": dict((parent.scope_json or {}).get("skills") or {}),
|
|
"write": bool(write),
|
|
"allow": list(SAFE_COMMANDS),
|
|
}
|
|
|
|
|
|
def _create_child(
|
|
db,
|
|
parent: Chat,
|
|
*,
|
|
title: str,
|
|
write: bool,
|
|
friend: Model | None = None,
|
|
) -> Chat:
|
|
"""The hidden chat one helper or one friend runs in.
|
|
|
|
A helper inherits the parent's model, connection, directory and reasoning
|
|
effort, and nothing else. The effort has to be **seeded onto the row** rather
|
|
than left to be inherited at request time: `chat_service.resolved_effort`
|
|
reads the chat's own `params_json` and deliberately consults no fallback, so
|
|
a helper of a high-effort reply would otherwise quietly run at none.
|
|
|
|
`friend` makes it somebody else's chat instead, and changes three things.
|
|
|
|
**The model and the connection are the friend's**, as a pair rather than an
|
|
id: `Model` is unique on `(connection_id, model_id)`, so the same name can
|
|
live behind two endpoints and an id alone does not say which.
|
|
|
|
**The effort is the friend's own default, never the parent's.** Inheriting it
|
|
across models is the 1.3.0 bug with a new door: the vocabularies differ, and
|
|
`high` handed to a Bonsai raises inside its chat template rather than being
|
|
ignored. A level the friend does not take is simply not sent.
|
|
|
|
**It is not put to work on a machine.** A friend is asked what it thinks, so
|
|
it gets no SSH profile, no project directory and no agent mode even when the
|
|
asking chat has all three -- and `scope_json["role"]` marks it so its own
|
|
system prompt can say it is answering a peer rather than running an errand.
|
|
"""
|
|
from lembas.services import chat as chat_service
|
|
|
|
peer = friend is not None
|
|
child = Chat(
|
|
user_id=parent.user_id,
|
|
# An ordinary chat for a friend even when the asking one is an agent
|
|
# chat: KIND_AGENT brings a harness about the machine it is working on,
|
|
# and a peer being asked a question is not working on one.
|
|
kind=KIND_CHAT if peer else parent.kind,
|
|
title=title[:200] or ("Question" if peer else "Helper"),
|
|
model_id=friend.model_id if peer else parent.model_id,
|
|
connection_id=friend.connection_id if peer else parent.connection_id,
|
|
# Never in a listing, and swept a day later even if it is kept.
|
|
temporary=True,
|
|
parent_chat_id=parent.id,
|
|
unattended=True,
|
|
scope_json=_child_scope(parent, write=write),
|
|
)
|
|
if not peer and parent.kind == KIND_AGENT:
|
|
child.ssh_profile_id = parent.ssh_profile_id
|
|
child.project_dir = parent.project_dir
|
|
child.agent_mode = MODE_WRITING if write else MODE_READING
|
|
if peer:
|
|
child.scope_json = {**(child.scope_json or {}), "role": ROLE_FRIEND}
|
|
effort = str((friend.params_json or {}).get("reasoning_effort") or "")
|
|
if effort not in chat_service.efforts_for(friend):
|
|
effort = ""
|
|
else:
|
|
effort = chat_service.resolved_effort(parent)
|
|
if effort:
|
|
child.params_json = {"reasoning_effort": effort}
|
|
# The bases the parent is scoped to, or the helper searches everything its
|
|
# owner can see and answers from documents the parent was not looking at.
|
|
child.knowledge_bases = list(parent.knowledge_bases)
|
|
db.add(child)
|
|
db.commit()
|
|
return child
|
|
|
|
|
|
def _task_turn(task: str, context: str) -> str:
|
|
"""The one turn a helper is given.
|
|
|
|
Named as a delegation in *words*, for the reason `wake.py` sets out: the
|
|
role stays `user` because `build_messages` requires one there, so the
|
|
framing cannot live in the role. `core.subagent` is the other half — this
|
|
says what the job is, the fragment says what being a helper means.
|
|
"""
|
|
lines = [
|
|
"You are answering a request from another model, not from a person. "
|
|
"Nobody is reading this conversation; your reply is handed back whole "
|
|
"as the result of one tool call.",
|
|
"",
|
|
"## The task",
|
|
task.strip(),
|
|
]
|
|
if context.strip():
|
|
lines += ["", "## What you have been told about it", context.strip()]
|
|
return "\n".join(lines)
|
|
|
|
|
|
# --- Running one ---------------------------------------------------------------
|
|
async def _await_reply(chat_id: str, message_id: str, deadline: float) -> bool:
|
|
"""Wait for the helper's reply. True if it finished on its own.
|
|
|
|
Polled rather than awaited on the task, the same reasoning
|
|
`schedule/runner._await_reply` writes down: `generation` owns its registry
|
|
and reaching into it from here would couple this to internals whose whole
|
|
job is to be replaceable.
|
|
"""
|
|
from lembas.services import generation as generation_service
|
|
|
|
while time.monotonic() < deadline:
|
|
running = generation_service.running_for(chat_id)
|
|
# The id check is what stops this waiting on some *later* reply in the
|
|
# same chat -- there is nothing else to produce one here, but the same
|
|
# loop in `runner` needed it and the cost of keeping it is nothing.
|
|
if running is None or running.message_id != message_id:
|
|
return True
|
|
await asyncio.sleep(POLL_SECONDS)
|
|
return False
|
|
|
|
|
|
async def _stop(chat_id: str, message_id: str) -> None:
|
|
"""End a helper that has run out of clock, keeping what it wrote.
|
|
|
|
`request_stop` rather than cancelling the task: it sets the flag the
|
|
producer checks between chunks, so the partial reply is persisted and marked
|
|
`stopped` rather than `error`. An abandoned generation would go on spending
|
|
the endpoint after the parent had stopped caring.
|
|
"""
|
|
from lembas.services import generation as generation_service
|
|
|
|
generation_service.request_stop(message_id)
|
|
deadline = time.monotonic() + STOP_GRACE
|
|
while time.monotonic() < deadline:
|
|
running = generation_service.running_for(chat_id)
|
|
if running is None or running.message_id != message_id:
|
|
return
|
|
await asyncio.sleep(POLL_SECONDS)
|
|
log.warning("subagent %s did not stop within the grace period", chat_id)
|
|
|
|
|
|
def _harvest(db, chat_id: str, message_id: str) -> tuple[str, str]:
|
|
"""The helper's answer, and what went wrong if anything did.
|
|
|
|
Read from the row rather than from the `Generation`, because `_persist` is
|
|
the single writer and the row is authoritative the moment `done` is set --
|
|
the same order `_follow` depends on.
|
|
"""
|
|
from lembas.db.models import Message
|
|
|
|
message = db.get(Message, message_id)
|
|
if message is None or message.chat_id != chat_id:
|
|
return "", "The helper's reply could not be found."
|
|
text = (message.content or "").strip()
|
|
if message.error:
|
|
return text, str(message.error)
|
|
if not text:
|
|
return "", "The helper produced no answer."
|
|
return text, ""
|
|
|
|
|
|
def _tool_names(db, chat_id: str) -> list[str]:
|
|
"""What the helper actually did, for the transcript.
|
|
|
|
The tool names off its messages, in order, deduplicated by run. It is what
|
|
makes a collapsed block worth expanding: an answer alone cannot say whether
|
|
it was researched or recalled.
|
|
"""
|
|
from sqlalchemy import select
|
|
|
|
from lembas.db.models import Message
|
|
|
|
names: list[str] = []
|
|
rows = db.scalars(
|
|
select(Message).where(Message.chat_id == chat_id).order_by(Message.created_at)
|
|
)
|
|
for row in rows:
|
|
for event in row.tool_calls_json or []:
|
|
name = str((event or {}).get("name") or "")
|
|
if name and (not names or names[-1] != name):
|
|
names.append(name)
|
|
return names
|
|
|
|
|
|
def _cleanup(chat_id: str, *, keep: bool) -> None:
|
|
"""Delete the helper's chat unless an administrator asked to keep it.
|
|
|
|
Best-effort and outside every other session: a helper whose answer has been
|
|
handed back has done its job, and failing to tidy up must not turn a good
|
|
result into an error. Kept chats are `temporary`, so the day-old sweep gets
|
|
them either way.
|
|
"""
|
|
if keep:
|
|
return
|
|
try:
|
|
with session_scope() as db:
|
|
from lembas.services import chat as chat_service
|
|
|
|
chat = db.get(Chat, chat_id)
|
|
if chat is not None:
|
|
# A writing helper can generate an image or attach a file, and
|
|
# `db.delete` would leave both on disk with the row that named
|
|
# them gone.
|
|
chat_service.delete_chats(db, [chat])
|
|
except Exception: # noqa: BLE001 - tidying up is not the result
|
|
log.debug("could not remove subagent chat %s", chat_id, exc_info=True)
|
|
|
|
|
|
# --- The tool ------------------------------------------------------------------
|
|
def _outcome(text: str, event: dict[str, Any]) -> ToolOutcome:
|
|
"""Imported inside the call: `services/tools.py` imports this module to build
|
|
the definition, so a top-level import back is a cycle."""
|
|
from lembas.services.tools import ToolOutcome
|
|
|
|
return ToolOutcome(text, event)
|
|
|
|
|
|
def _error(message: str, *, task: str = "") -> ToolOutcome:
|
|
return _outcome(
|
|
message,
|
|
{"name": "subagent_run", "status": "error", "query": task[:120], "error": message},
|
|
)
|
|
|
|
|
|
def _budget(generation, values: dict[str, Any]) -> str:
|
|
"""Whether this reply may spawn another helper, and why not if it may not.
|
|
|
|
Counted on the parent's `Generation` because that is the only object that
|
|
knows what "this reply" is: a chat-keyed counter would have to be reset by
|
|
something, and every candidate for that something is a place to forget.
|
|
Read and incremented with no `await` between, which is what makes it safe
|
|
against the four calls a round runs together.
|
|
|
|
`values` is passed in rather than read here, so this can be called from
|
|
inside the caller's session -- opening a second one underneath an open one
|
|
is a shape this codebase does not have anywhere else and is not worth
|
|
introducing for a settings lookup.
|
|
"""
|
|
if generation is None:
|
|
# No generation means no reply to bound. It happens if a runner is ever
|
|
# reached outside the loop; refusing is the answer that cannot be wrong.
|
|
return "This reply cannot delegate."
|
|
if generation.subagents >= int(values["max_per_reply"]):
|
|
return (
|
|
f"This reply has already used its {values['max_per_reply']} helpers. "
|
|
"Do the rest yourself, or answer with what you have."
|
|
)
|
|
if len(_LIVE) >= int(values["max_concurrent"]):
|
|
return (
|
|
"Too many helpers are running on this instance right now. "
|
|
"Do this part yourself rather than waiting."
|
|
)
|
|
generation.subagents += 1
|
|
return ""
|
|
|
|
|
|
async def _run_subagent(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
|
|
from lembas.services import generation as generation_service
|
|
from lembas.services import wake as wake_service
|
|
|
|
task = str(args.get("task") or "").strip()
|
|
title = str(args.get("title") or "").strip() or task[:60]
|
|
briefing = str(args.get("context") or "")
|
|
want_write = bool(args.get("write"))
|
|
|
|
if not task:
|
|
return _error(
|
|
"A helper needs a task: what to find out or do, written out in full. "
|
|
"It starts with none of this conversation, so say everything it needs."
|
|
)
|
|
|
|
parent_id = context.chat_id
|
|
if not parent_id:
|
|
return _error("There is no conversation to delegate from.", task=task)
|
|
|
|
with session_scope() as db:
|
|
parent = db.get(Chat, parent_id)
|
|
if parent is None:
|
|
return _error("That conversation no longer exists.", task=task)
|
|
# Enforced here as well as by the withdrawn family, because this is the
|
|
# cheaper half to get right and the two failures look different: the
|
|
# family withdrawal means the model is never offered the tool, and this
|
|
# means a call that arrived by some other path is refused rather than
|
|
# opening a third level.
|
|
if parent.parent_chat_id or parent.unattended:
|
|
return _error("A helper cannot ask for a helper of its own.", task=task)
|
|
|
|
write = want_write
|
|
if write and parent.kind == KIND_AGENT and parent.agent_mode not in WRITING_ALLOWED_FROM:
|
|
return _error(
|
|
"This chat is in "
|
|
f"{agent_policy.MODE_LABELS.get(parent.agent_mode, parent.agent_mode)} mode, "
|
|
"where you are stopped before anything is written — so a helper "
|
|
"cannot write either, since nobody can be stopped to ask. Send a "
|
|
"reading helper and make the changes yourself, or ask the reader "
|
|
"to switch mode.",
|
|
task=task,
|
|
)
|
|
owner = db.get(User, parent.user_id)
|
|
if owner is None: # pragma: no cover - a chat outliving its owner
|
|
return _error("That account no longer exists.", task=task)
|
|
|
|
# After the refusals above and before anything is created. The order is
|
|
# the design: a call that could never have worked should be told *why*
|
|
# rather than told it has run out of helpers, and the counter should
|
|
# only move for a call that is about to spend one.
|
|
values = settings_store.subagents(db)
|
|
# A group's ceiling narrows the instance's, never widens it. Zero on
|
|
# either side means "no opinion", so the two cannot be folded with
|
|
# `min` -- see generation._narrower for the same arithmetic.
|
|
allowance = permissions.limit(db, owner, "helpers_per_reply")
|
|
if allowance:
|
|
values = {**values, "max_per_reply": min(int(values["max_per_reply"]), allowance)}
|
|
refusal = _budget(generation_service.running_for(parent_id), values)
|
|
if refusal:
|
|
return _error(refusal, task=task)
|
|
|
|
child = _create_child(db, parent, title=title, write=write)
|
|
child_id = child.id
|
|
|
|
_LIVE.add(child_id)
|
|
started = time.monotonic()
|
|
try:
|
|
message_id = await wake_service.wake_chat(child_id, _task_turn(task, briefing))
|
|
if not message_id:
|
|
_cleanup(child_id, keep=False)
|
|
return _error("The helper could not be started.", task=task)
|
|
|
|
finished = await _await_reply(
|
|
child_id, message_id, started + float(values["wall_seconds"])
|
|
)
|
|
if not finished:
|
|
await _stop(child_id, message_id)
|
|
|
|
with session_scope() as db:
|
|
answer, problem = _harvest(db, child_id, message_id)
|
|
used = _tool_names(db, child_id)
|
|
finally:
|
|
_LIVE.discard(child_id)
|
|
|
|
elapsed = time.monotonic() - started
|
|
keep = bool(values.get("keep_transcript"))
|
|
_cleanup(child_id, keep=keep)
|
|
|
|
if not answer:
|
|
return _error(problem or "The helper produced no answer.", task=task)
|
|
|
|
# The account of what it did goes in the *event*, where the transcript shows
|
|
# it; the answer goes to the model. Putting the tool list in front of the
|
|
# model as well would be spending its window on our own bookkeeping.
|
|
note = "" if finished else "\n\n(It ran out of time; this is as far as it got.)"
|
|
return _outcome(
|
|
f"The helper answered:\n\n{answer}{note}\n\n"
|
|
"This is another model's work, not yours and not the reader's. Check it "
|
|
"against what you know before relying on it, and say what came from it.",
|
|
{
|
|
"name": "subagent_run",
|
|
"status": "ok" if finished else "error" if not answer else "ok",
|
|
"query": title,
|
|
"detail": (
|
|
f"{len(used)} tool call(s), {elapsed:.0f}s"
|
|
+ ("" if finished else ", stopped at the time limit")
|
|
),
|
|
"text": answer,
|
|
"why": ", ".join(used[:8]) if used else "",
|
|
},
|
|
)
|
|
|
|
|
|
# --- Asking a friend -----------------------------------------------------------
|
|
def _friend_error(message: str, *, question: str = "") -> ToolOutcome:
|
|
return _outcome(
|
|
message,
|
|
{"name": "ask_friend", "status": "error", "query": question[:120], "error": message},
|
|
)
|
|
|
|
|
|
def _resolve_friend(db, owner: User, wanted: str, *, asking: str) -> tuple[Model | None, str]:
|
|
"""The model a call named, or a refusal that says what it could have named.
|
|
|
|
The name arrives in a tool call, which is to say it was written by a model
|
|
that may have been reading a web page, so it is matched against what **this
|
|
account** can reach rather than against the table. `roster_models` is the
|
|
same list the prompt was built from, so a refusal here cannot disagree with
|
|
what the model was told.
|
|
|
|
Matched on `model_id` first and on the label second, because the roster
|
|
prints both and a model will sometimes type back the pretty one.
|
|
"""
|
|
from lembas.services import chat as chat_service
|
|
|
|
question_for = wanted.strip()
|
|
candidates = chat_service.roster_models(db, owner, exclude=asking)
|
|
if not candidates:
|
|
return None, (
|
|
"There is no other model here to ask. Answer from what you know."
|
|
)
|
|
if not question_for:
|
|
return None, (
|
|
"Name the model to ask, exactly as it is written in brackets in the "
|
|
"list you were given:\n"
|
|
+ chat_service.roster_block(db, owner, exclude=asking)
|
|
)
|
|
|
|
lowered = question_for.lower()
|
|
for model in candidates:
|
|
if model.model_id.lower() == lowered:
|
|
return model, ""
|
|
for model in candidates:
|
|
if model.label.lower() == lowered:
|
|
return model, ""
|
|
|
|
# `candidates` already excludes the asker, so its own name would otherwise
|
|
# fall through to "there is no model called that", which is both untrue and
|
|
# unhelpful.
|
|
if lowered == asking.lower():
|
|
return None, "That is you. Ask somebody else, or answer it yourself."
|
|
|
|
return None, (
|
|
f"There is no model called {question_for!r} that you can reach. "
|
|
"These are the ones you can:\n"
|
|
+ chat_service.roster_block(db, owner, exclude=asking)
|
|
)
|
|
|
|
|
|
def _question_turn(question: str, context: str, asker: str) -> str:
|
|
"""The one turn a friend is given.
|
|
|
|
Deliberately not `_task_turn`. A helper is told it is doing a job nobody is
|
|
reading; a friend is told another model wants its opinion, which is a
|
|
different thing to be and produces a different answer -- a helper reports,
|
|
a peer disagrees. The framing lives in words for the reason `wake.py` sets
|
|
out: the role has to stay `user`, because `build_messages` requires a user
|
|
turn there.
|
|
"""
|
|
lines = [
|
|
f"Another model ({asker}) is asking you a question, on behalf of the "
|
|
"person it is talking to. Nobody is reading this conversation directly: "
|
|
"your reply is handed back whole as the answer.",
|
|
"",
|
|
"Answer it as yourself. If you think the question rests on something "
|
|
"wrong, say so — that is usually why you were asked. If you do not know, "
|
|
"say that rather than guessing; a confident wrong answer is worse than "
|
|
"no answer, because it will be relied on.",
|
|
"",
|
|
"## The question",
|
|
question.strip(),
|
|
]
|
|
if context.strip():
|
|
lines += ["", "## What you have been told about it", context.strip()]
|
|
return "\n".join(lines)
|
|
|
|
|
|
async def _run_ask_friend(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
|
|
from lembas.services import generation as generation_service
|
|
from lembas.services import wake as wake_service
|
|
|
|
question = str(args.get("question") or "").strip()
|
|
wanted = str(args.get("model") or "")
|
|
briefing = str(args.get("context") or "")
|
|
|
|
if not question:
|
|
return _friend_error(
|
|
"Ask something. The model you are asking sees none of this "
|
|
"conversation, so the question has to stand on its own."
|
|
)
|
|
|
|
parent_id = context.chat_id
|
|
if not parent_id:
|
|
return _friend_error("There is no conversation to ask from.", question=question)
|
|
|
|
with session_scope() as db:
|
|
parent = db.get(Chat, parent_id)
|
|
if parent is None:
|
|
return _friend_error("That conversation no longer exists.", question=question)
|
|
# The same belt-and-braces as `_run_subagent`: the family is withdrawn
|
|
# from an unattended chat, and a call arriving by any other route is
|
|
# refused here rather than opening a third level.
|
|
if parent.parent_chat_id or parent.unattended:
|
|
return _friend_error(
|
|
"You are answering a question yourself. Answer it, or say you "
|
|
"cannot — you may not pass it on.",
|
|
question=question,
|
|
)
|
|
owner = db.get(User, parent.user_id)
|
|
if owner is None: # pragma: no cover - a chat outliving its owner
|
|
return _friend_error("That account no longer exists.", question=question)
|
|
|
|
friend, refusal = _resolve_friend(db, owner, wanted, asking=parent.model_id)
|
|
if friend is None:
|
|
return _friend_error(refusal, question=question)
|
|
|
|
# Bounded by the same allowance as a helper, and counted on the same
|
|
# counter: both spend one reply to get another, and two separate budgets
|
|
# would let one reply spend both.
|
|
values = settings_store.subagents(db)
|
|
allowance = permissions.limit(db, owner, "helpers_per_reply")
|
|
if allowance:
|
|
values = {**values, "max_per_reply": min(int(values["max_per_reply"]), allowance)}
|
|
refusal = _budget(generation_service.running_for(parent_id), values)
|
|
if refusal:
|
|
return _friend_error(refusal, question=question)
|
|
|
|
asker = parent.model_id
|
|
label = friend.label
|
|
child = _create_child(
|
|
db, parent, title=f"Asking {label}"[:200], write=False, friend=friend
|
|
)
|
|
child_id = child.id
|
|
|
|
_LIVE.add(child_id)
|
|
started = time.monotonic()
|
|
try:
|
|
message_id = await wake_service.wake_chat(
|
|
child_id, _question_turn(question, briefing, asker)
|
|
)
|
|
if not message_id:
|
|
_cleanup(child_id, keep=False)
|
|
return _friend_error(f"{label} could not be reached.", question=question)
|
|
|
|
finished = await _await_reply(
|
|
child_id, message_id, started + float(values["wall_seconds"])
|
|
)
|
|
if not finished:
|
|
await _stop(child_id, message_id)
|
|
|
|
with session_scope() as db:
|
|
answer, problem = _harvest(db, child_id, message_id)
|
|
finally:
|
|
_LIVE.discard(child_id)
|
|
|
|
elapsed = time.monotonic() - started
|
|
_cleanup(child_id, keep=bool(values.get("keep_transcript")))
|
|
|
|
if not answer:
|
|
return _friend_error(problem or f"{label} did not answer.", question=question)
|
|
|
|
note = "" if finished else "\n\n(It ran out of time; this is as far as it got.)"
|
|
return _outcome(
|
|
f"{label} answered:\n\n{answer}{note}\n\n"
|
|
"That is another model's opinion, not a fact and not the reader's. Say "
|
|
"whose it is when you use it, and say so too if you disagree with it.",
|
|
{
|
|
"name": "ask_friend",
|
|
"status": "ok" if finished else "error",
|
|
"query": f"{label}: {question}"[:160],
|
|
"detail": f"{elapsed:.0f}s" + ("" if finished else ", stopped at the time limit"),
|
|
"text": answer,
|
|
"why": label,
|
|
},
|
|
)
|
|
|
|
|
|
def friend_tool_defs() -> list[ToolDef]:
|
|
"""The ask-a-friend tool. Its own family; see `services/tools.py`."""
|
|
from lembas.services.tools import FAMILY_FRIEND, RISK_READ, ToolDef
|
|
|
|
return [
|
|
ToolDef(
|
|
name="ask_friend",
|
|
family=FAMILY_FRIEND,
|
|
description=(
|
|
"Put one question to another model here and get its answer. Use "
|
|
"it for a second opinion, for something outside what you are good "
|
|
"at, or to have your own reasoning checked by something that "
|
|
"thinks differently — the list of models you can ask, and what "
|
|
"each is for, is in your instructions. It answers as itself and "
|
|
"sees none of this conversation, so the question must stand on "
|
|
"its own. Its answer is an opinion: say whose it is, and say so "
|
|
"if you disagree. Do not ask for something you can work out "
|
|
"yourself, and do not ask the same thing of several models hoping "
|
|
"one agrees with you."
|
|
),
|
|
parameters={
|
|
"type": "object",
|
|
"properties": {
|
|
"model": {
|
|
"type": "string",
|
|
"description": (
|
|
"Which model to ask, written exactly as the id in "
|
|
"brackets in the list you were given."
|
|
),
|
|
},
|
|
"question": {
|
|
"type": "string",
|
|
"description": (
|
|
"The question, written out in full. It is read on its "
|
|
"own, with none of this conversation around it."
|
|
),
|
|
},
|
|
"context": {
|
|
"type": "string",
|
|
"description": (
|
|
"Anything it needs to answer — the code in question, "
|
|
"the constraint, what has already been tried. Not a "
|
|
"summary of the conversation."
|
|
),
|
|
},
|
|
},
|
|
"required": ["model", "question"],
|
|
},
|
|
run=_run_ask_friend,
|
|
# A read, for the reason `subagent_run` is one: what the answer costs
|
|
# is another reply, and nothing in this instance is changed by it.
|
|
risk=RISK_READ,
|
|
),
|
|
]
|
|
|
|
|
|
def tool_defs() -> list[ToolDef]:
|
|
"""The one tool, built here so `services/tools.py` need not know the wording."""
|
|
from lembas.services.tools import FAMILY_SUBAGENT, RISK_READ, ToolDef
|
|
|
|
return [
|
|
ToolDef(
|
|
name="subagent_run",
|
|
family=FAMILY_SUBAGENT,
|
|
description=(
|
|
"Hand one self-contained piece of work to a helper — a second "
|
|
"model that works on its own and gives you its answer. Use it "
|
|
"to cover several independent areas at once: call it several "
|
|
"times in one turn and each runs in parallel. It gets a "
|
|
"narrower set of tools than you: it reads, it cannot ask you or "
|
|
"the reader anything, it cannot delegate further, it changes "
|
|
"nothing unless you set write, and on a machine it may run only "
|
|
"a fixed list of read-only commands. It starts knowing nothing "
|
|
"about this conversation, so the task must say everything it "
|
|
"needs. Do not use it for something you could do in one call "
|
|
"yourself, or for anything needing a decision only the reader "
|
|
"can make."
|
|
),
|
|
parameters={
|
|
"type": "object",
|
|
"properties": {
|
|
"task": {
|
|
"type": "string",
|
|
"description": (
|
|
"What the helper is to do, written out in full and as "
|
|
"an instruction. Say what a good answer contains and "
|
|
"how long it should be. It is read on its own, with "
|
|
"none of this conversation around it."
|
|
),
|
|
},
|
|
"title": {
|
|
"type": "string",
|
|
"description": "A few words naming this piece of work.",
|
|
},
|
|
"context": {
|
|
"type": "string",
|
|
"description": (
|
|
"Facts the helper needs that it cannot look up — what "
|
|
"the reader asked for, decisions already made, names "
|
|
"and paths. Not a summary of the conversation."
|
|
),
|
|
},
|
|
"write": {
|
|
"type": "boolean",
|
|
"description": (
|
|
"True if the helper must change something: write a "
|
|
"file, keep a note, file a report. Leave it out for "
|
|
"anything that only reads, which is nearly always. A "
|
|
"writing helper is refused where you would have been "
|
|
"stopped for approval yourself."
|
|
),
|
|
},
|
|
},
|
|
"required": ["task"],
|
|
},
|
|
run=_run_subagent,
|
|
# It reads, from the parent's side: what it changes, it changes
|
|
# through tools that carry their own risk class inside the helper's
|
|
# own chat, where the mode and the scope decide. Classing the spawn
|
|
# itself as a write would put an approval card in front of every
|
|
# research fan-out in Edit mode, which is the mode that permits
|
|
# writing anyway.
|
|
risk=RISK_READ,
|
|
),
|
|
]
|
|
|
|
|
|
__all__ = [
|
|
"MODE_READING",
|
|
"ROLE_FRIEND",
|
|
"MODE_WRITING",
|
|
"SAFE_COMMANDS",
|
|
"WRITING_ALLOWED_FROM",
|
|
"clear",
|
|
"friend_tool_defs",
|
|
"live_count",
|
|
"tool_defs",
|
|
]
|