Agent chats run commands, and stop to ask first
The four tools an agent chat has -- shell_run, file_read, file_write, file_list -- and the mode table wired into the loop that decides which of them stop for approval. Verified end to end against a real Kali container over SSH: the card shows the command, allowing it runs it there, and the file it writes is visible from outside. The mode is enforced in `_authorise`, in the generation loop, server-side, keyed on each tool's declared risk. Not in the prompt: a model is told which mode it is in so it behaves sensibly, but everything it reads -- a web page, a README, the output of the last command -- is untrusted, and a rule written only into a system message is one a poisoned file can argue with. Within an agent chat every call goes through the table, including the built-in ones, because notes_edit writes and Plan mode meaning "look but do not touch" has to mean that too. Two things this turned up. The runners re-check the mode as a backstop, and that backstop refused the very thing a person had just approved -- the mode says "ask", and asking was exactly what happened. Approval is now threaded per call, on a copy of the context, because a round runs its calls together and only some of them were allowed. And the harness said nothing at all, because `registry` maps an offered tool *name* back to a family and did not know the agent tools existed. So shell_run resolved to no family and the fragment naming the machine, the directory and the mode was never admitted. The same omission cost custom tools their guidance once already; there is a test for it now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -41,6 +41,35 @@ MODE_HINTS = {
|
||||
MODE_PLAN: "Reads freely, changes nothing, and finishes by proposing a plan.",
|
||||
}
|
||||
|
||||
# What the *model* is told about the mode it is in. Different words from
|
||||
# MODE_HINTS, which describes it to a person: this is about how to behave, and
|
||||
# says the one thing that changes what a competent model does -- that being
|
||||
# stopped for approval is normal and worth batching for.
|
||||
MODE_GUIDANCE = {
|
||||
MODE_MANUAL: (
|
||||
"You are in **Manual** mode: everything you do is shown to them for "
|
||||
"approval first. Expect to be interrupted, and say what you are about "
|
||||
"to do before you do it."
|
||||
),
|
||||
MODE_EDIT: (
|
||||
"You are in **Edit** mode: you may read and write files freely, but "
|
||||
"every command is shown to them for approval first. Prefer reading and "
|
||||
"writing files over shelling out where both would work."
|
||||
),
|
||||
MODE_AUTO: (
|
||||
"You are in **Auto** mode: nothing is shown to them first. That is trust "
|
||||
"rather than permission — be as careful as you would be if each step "
|
||||
"were being watched, and stop to say so if you find yourself about to "
|
||||
"do something you could not undo."
|
||||
),
|
||||
MODE_PLAN: (
|
||||
"You are in **Plan** mode: read and explore freely, but change nothing. "
|
||||
"Anything that writes or runs will be stopped for approval, so do not "
|
||||
"rely on it. Finish by setting out what you would do, as steps, so it "
|
||||
"can be carried out afterwards."
|
||||
),
|
||||
}
|
||||
|
||||
ALLOW = "allow"
|
||||
ASK = "ask"
|
||||
|
||||
@@ -172,6 +201,7 @@ __all__ = [
|
||||
"MODES",
|
||||
"MODE_AUTO",
|
||||
"MODE_EDIT",
|
||||
"MODE_GUIDANCE",
|
||||
"MODE_HINTS",
|
||||
"MODE_LABELS",
|
||||
"MODE_MANUAL",
|
||||
|
||||
Reference in New Issue
Block a user