Agent chats run commands, and stop to ask first
The four tools an agent chat has -- shell_run, file_read, file_write, file_list -- and the mode table wired into the loop that decides which of them stop for approval. Verified end to end against a real Kali container over SSH: the card shows the command, allowing it runs it there, and the file it writes is visible from outside. The mode is enforced in `_authorise`, in the generation loop, server-side, keyed on each tool's declared risk. Not in the prompt: a model is told which mode it is in so it behaves sensibly, but everything it reads -- a web page, a README, the output of the last command -- is untrusted, and a rule written only into a system message is one a poisoned file can argue with. Within an agent chat every call goes through the table, including the built-in ones, because notes_edit writes and Plan mode meaning "look but do not touch" has to mean that too. Two things this turned up. The runners re-check the mode as a backstop, and that backstop refused the very thing a person had just approved -- the mode says "ask", and asking was exactly what happened. Approval is now threaded per call, on a copy of the context, because a round runs its calls together and only some of them were allowed. And the harness said nothing at all, because `registry` maps an offered tool *name* back to a family and did not know the agent tools existed. So shell_run resolved to no family and the fragment naming the machine, the directory and the mode was never admitted. The same omission cost custom tools their guidance once already; there is a test for it now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
6a849dc1ec
commit
b6aab8de55
@@ -73,6 +73,11 @@ FAMILY_MCP = "mcp"
|
||||
# is the only tool the model cannot resolve by itself.
|
||||
FAMILY_ASK = "ask"
|
||||
|
||||
# Acting on the machine an agent chat is pointed at. Offered only when the chat
|
||||
# is one, has a usable connection, and the feature is switched on -- see
|
||||
# services/agent/session.py:resolve, which answers all three at once.
|
||||
FAMILY_AGENT = "agent"
|
||||
|
||||
# The built-in families, in the order they are offered.
|
||||
FAMILIES = (
|
||||
FAMILY_SEARCH,
|
||||
@@ -81,6 +86,7 @@ FAMILIES = (
|
||||
FAMILY_MEMORY,
|
||||
FAMILY_SKILLS,
|
||||
FAMILY_ASK,
|
||||
FAMILY_AGENT,
|
||||
)
|
||||
|
||||
GATES = (*FAMILIES, FAMILY_CUSTOM, FAMILY_MCP)
|
||||
@@ -128,6 +134,10 @@ class ToolContext:
|
||||
# something. Read from the instance settings while the session was open,
|
||||
# like everything else here.
|
||||
interaction_timeout: float = 900.0
|
||||
# Set only for an agent chat: the machine to act on, the mode in force, and
|
||||
# the decrypted credential. None everywhere else, which is what every agent
|
||||
# runner checks first. `generation` clears it when the reply ends.
|
||||
agent: Any = None
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -825,7 +835,7 @@ def _family_allowed(
|
||||
and config.get("enabled")
|
||||
and not search_service.availability(str(config.get("provider") or "ddgs"))
|
||||
)
|
||||
if gate in (FAMILY_CUSTOM, FAMILY_MCP, FAMILY_ASK):
|
||||
if gate in (FAMILY_CUSTOM, FAMILY_MCP, FAMILY_ASK, FAMILY_AGENT):
|
||||
# Deliberately without `library.use`: an HTTP endpoint an administrator
|
||||
# wrote has nothing to do with this person's own documents and notes,
|
||||
# and requiring the library permission for it would be a coincidence of
|
||||
@@ -852,6 +862,23 @@ def _row_defs(db: DBSession, user: User | None, *, everything: bool = False) ->
|
||||
return [*custom, *mcp_registry.tool_defs(db, user, everything=everything, taken=taken)]
|
||||
|
||||
|
||||
def _agent_defs(db: DBSession, chat: Chat | None, user: User | None) -> list[ToolDef]:
|
||||
"""The agent tools, when this chat is pointed at a machine it can use.
|
||||
|
||||
Everything that would make them useless -- not an agent chat, the feature
|
||||
switched off, the connection deleted or disabled, SSH not installed -- comes
|
||||
back as an empty list, because offering a tool that fails on its first call
|
||||
is worse than not offering it.
|
||||
"""
|
||||
from lembas.services.agent import session as agent_session
|
||||
from lembas.services.agent import tools as agent_tools
|
||||
|
||||
context = agent_session.resolve(db, chat, user) if chat is not None else None
|
||||
if context is None:
|
||||
return []
|
||||
return agent_tools.tool_defs(context)
|
||||
|
||||
|
||||
def _book(defs: list[ToolDef]) -> dict[str, ToolDef]:
|
||||
"""Keyed by name, first claim winning.
|
||||
|
||||
@@ -872,8 +899,16 @@ def registry(db: DBSession) -> dict[str, ToolDef]:
|
||||
an administrator-defined tool is a row. Callers that only need to map a name
|
||||
back to a family use this; callers deciding what to *offer* use
|
||||
`resolve_tools`, which applies the gates as well.
|
||||
|
||||
The agent tools are listed here **unbound to any chat**. Mapping a name back
|
||||
to its family is exactly what the harness does to decide whether a
|
||||
fragment applies, and without them `shell_run` would resolve to no family at
|
||||
all -- so an agent chat would be told nothing about the machine it is
|
||||
working on. The same omission cost custom tools their guidance once already.
|
||||
"""
|
||||
return _book(_row_defs(db, None, everything=True))
|
||||
from lembas.services.agent import tools as agent_tools
|
||||
|
||||
return _book([*_row_defs(db, None, everything=True), *agent_tools.tool_defs()])
|
||||
|
||||
|
||||
def families(db: DBSession) -> tuple[str, ...]:
|
||||
@@ -900,7 +935,7 @@ def resolve_tools(db: DBSession, chat: Chat, user: User | None) -> ToolSet:
|
||||
|
||||
# Resolved against what this reader may see, not against everything that
|
||||
# exists: a tool restricted to a group is not offered outside it.
|
||||
book = _book(_row_defs(db, user))
|
||||
book = _book([*_row_defs(db, user), *_agent_defs(db, chat, user)])
|
||||
return ToolSet(
|
||||
tuple(
|
||||
tool
|
||||
@@ -929,7 +964,10 @@ def context_for(
|
||||
tools: ToolSet | None = None,
|
||||
) -> ToolContext:
|
||||
"""The snapshot a running tool needs, taken while the session is open."""
|
||||
from lembas.services.agent import session as agent_session
|
||||
|
||||
return ToolContext(
|
||||
agent=agent_session.resolve(db, chat, user) if chat is not None else None,
|
||||
owner_id=user.id if user else "",
|
||||
search_config=settings_store.search(db),
|
||||
base_ids=[base.id for base in chat.knowledge_bases] if chat is not None else [],
|
||||
@@ -1105,6 +1143,7 @@ def _row_source(db: DBSession):
|
||||
|
||||
__all__ = [
|
||||
"FAMILIES",
|
||||
"FAMILY_AGENT",
|
||||
"MAX_ROUNDS",
|
||||
"REGISTRY",
|
||||
"ToolCallAccumulator",
|
||||
|
||||
Reference in New Issue
Block a user