Agent chats run commands, and stop to ask first

The four tools an agent chat has -- shell_run, file_read, file_write,
file_list -- and the mode table wired into the loop that decides which of
them stop for approval. Verified end to end against a real Kali container
over SSH: the card shows the command, allowing it runs it there, and the
file it writes is visible from outside.

The mode is enforced in `_authorise`, in the generation loop, server-side,
keyed on each tool's declared risk. Not in the prompt: a model is told
which mode it is in so it behaves sensibly, but everything it reads -- a
web page, a README, the output of the last command -- is untrusted, and a
rule written only into a system message is one a poisoned file can argue
with. Within an agent chat every call goes through the table, including
the built-in ones, because notes_edit writes and Plan mode meaning "look
but do not touch" has to mean that too.

Two things this turned up.

The runners re-check the mode as a backstop, and that backstop refused the
very thing a person had just approved -- the mode says "ask", and asking
was exactly what happened. Approval is now threaded per call, on a copy of
the context, because a round runs its calls together and only some of them
were allowed.

And the harness said nothing at all, because `registry` maps an offered
tool *name* back to a family and did not know the agent tools existed. So
shell_run resolved to no family and the fragment naming the machine, the
directory and the mode was never admitted. The same omission cost custom
tools their guidance once already; there is a test for it now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jaroslav Beneš
2026-08-02 00:08:48 +02:00
parent 6a849dc1ec
commit b6aab8de55
14 changed files with 1579 additions and 29 deletions
+30
View File
@@ -41,6 +41,35 @@ MODE_HINTS = {
MODE_PLAN: "Reads freely, changes nothing, and finishes by proposing a plan.",
}
# What the *model* is told about the mode it is in. Different words from
# MODE_HINTS, which describes it to a person: this is about how to behave, and
# says the one thing that changes what a competent model does -- that being
# stopped for approval is normal and worth batching for.
MODE_GUIDANCE = {
MODE_MANUAL: (
"You are in **Manual** mode: everything you do is shown to them for "
"approval first. Expect to be interrupted, and say what you are about "
"to do before you do it."
),
MODE_EDIT: (
"You are in **Edit** mode: you may read and write files freely, but "
"every command is shown to them for approval first. Prefer reading and "
"writing files over shelling out where both would work."
),
MODE_AUTO: (
"You are in **Auto** mode: nothing is shown to them first. That is trust "
"rather than permission — be as careful as you would be if each step "
"were being watched, and stop to say so if you find yourself about to "
"do something you could not undo."
),
MODE_PLAN: (
"You are in **Plan** mode: read and explore freely, but change nothing. "
"Anything that writes or runs will be stopped for approval, so do not "
"rely on it. Finish by setting out what you would do, as steps, so it "
"can be carried out afterwards."
),
}
ALLOW = "allow"
ASK = "ask"
@@ -172,6 +201,7 @@ __all__ = [
"MODES",
"MODE_AUTO",
"MODE_EDIT",
"MODE_GUIDANCE",
"MODE_HINTS",
"MODE_LABELS",
"MODE_MANUAL",