3fc3449726
An agent command is one blocking conn.run over a per-call connection, killed the
moment it hits its timeout -- so a ten-minute apt install is impossible, which is
exactly what a user hit. This is the substrate for running it detached instead:
the model can ask for background=true, or a command that outlasts its timeout is
kept running rather than killed, and either way the model gets tools to read and
stop it. Opt-in, off by default, under Admin -> Agents; off is byte-for-byte the
old behaviour.
The mechanism has to survive the connection closing (that is the whole premise
of the per-call model), so a job is a setsid-detached process on the far side,
redirected to a remote logfile and an exit-file; LLeMbas reconnects, as always,
to read it later. services/agent/jobs.py holds the wrappers.
Three things in those wrappers are load-bearing and each was got wrong in the
first sketch:
- The command never touches a quoted shell context. sh -c '<cmd>' shatters the
instant the command contains a quote -- git commit -m 'fix', awk '{…}', sed
's/…/…/' are the common case, and it is an injection hole besides. So the
command is base64-encoded in Python and decoded on the far side into a script
file; it is bytes, never shell syntax.
- The child records its own pid via $$ as its first act, under setsid where it
is the session leader, so job_stop can kill the whole process group. echo $!
from the launcher captures the wrong pid.
- The command's exit status comes from the exit-file, never the wrapper's own
status -- which is ~0 from its trailing rm. Reading the wrapper's status would
mark every job a success.
A command that finishes in time is indistinguishable from a foreground one --
same output, same wording; the difference shows only when it does not, where
instead of "stopped after Ns" it becomes a job id. Auto-convert is its own
sub-switch: with it off, a timeout stays a hard stop and nothing is left
running, because routing the plain case through the detached wrapper would leave
an orphan running past a stop an administrator asked for.
New agent tools job_output/job_list/job_stop, offered only when the feature is
on (the plan_submit gating pattern); job_stop is RISK_EXECUTE since it kills a
process. A job's files are namespaced by the calling chat's id and the wrappers
are always built from it, so a model in one chat cannot even name another's job.
Tested against a real local /bin/sh rather than the fake echo-the-command sshd
fixture, because the shell logic -- setsid, base64, the wait loop, the child
surviving the wait being cut off -- is the whole of the risk. The auto-wake that
prompts the model back when a job finishes is the next commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
217 lines
7.1 KiB
Python
217 lines
7.1 KiB
Python
"""What a tool is called in the interface, and what it looks like.
|
|
|
|
Four places have to agree about one tool, and for the whole life of the feature
|
|
they did not:
|
|
|
|
- the transcript (`chat/_tool_activity.html`) showed the SSH profile's name for
|
|
an agent tool -- "homeserver · ls -la", naming the machine rather than the
|
|
thing that was done -- and the raw function name for everything else, so a
|
|
saved memory read `memory_add`;
|
|
- the status line while a round runs said "Running shell_run…";
|
|
- the approval card had its own hand-written if-chain;
|
|
- and nothing checked that any of the three matched.
|
|
|
|
So the table lives here and each of them reads it.
|
|
|
|
`LABELS` and `ACTIONS` are deliberately different words for the same tool, the
|
|
same way `policy.MODE_HINTS` and `policy.MODE_GUIDANCE` are. A label is a noun
|
|
phrase in a list of things that happened; an approval card is a sentence
|
|
somebody is agreeing to, and "Bash" is not one.
|
|
|
|
**The precedence is inverted on purpose, and that is the whole design.**
|
|
Tool events are persisted in `Message.tool_calls_json`, so every row written
|
|
before today already carries `label: "homeserver"`. A resolver that preferred
|
|
the stored label would fix nothing for any transcript that already exists. So
|
|
a name this module knows about resolves from the static table and the stored
|
|
label is ignored; a name it does not know -- a custom HTTP tool, an MCP tool,
|
|
whose labels are per row and cannot be tabulated -- keeps its own. One rule,
|
|
both cases correct.
|
|
|
|
Resolved **without a database**. It is called once per rendered event, and
|
|
reaching for `tools.registry(db)` from a Jinja global would be two table scans
|
|
per bubble.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from typing import Any
|
|
|
|
# What the transcript calls each tool. Kept in the same order as the families
|
|
# in `services/tools.py` so that adding one has an obvious home.
|
|
LABELS: dict[str, str] = {
|
|
# Acting on the machine an agent chat is pointed at.
|
|
"shell_run": "Bash",
|
|
"file_read": "Read",
|
|
"file_write": "Write",
|
|
"file_edit": "Update",
|
|
"file_list": "List",
|
|
"plan_submit": "Plan",
|
|
"plan_update": "Plan updated",
|
|
"job_output": "Job output",
|
|
"job_list": "Jobs",
|
|
"job_stop": "Job stopped",
|
|
# The web.
|
|
"web_search": "Web search",
|
|
"fetch": "Fetch",
|
|
# The library.
|
|
"knowledge_search": "Knowledge searched",
|
|
"knowledge_get": "Document read",
|
|
"notes_search": "Notes searched",
|
|
"notes_get": "Note read",
|
|
"notes_create": "Note written",
|
|
"notes_edit": "Note updated",
|
|
"notes_delete": "Note deleted",
|
|
"memory_add": "Memory saved",
|
|
"memory_forget": "Memory removed",
|
|
"skill_get": "Skill read",
|
|
"skill_create": "Skill written",
|
|
"skill_edit": "Skill updated",
|
|
# Stopping to ask.
|
|
"ask_user": "Asked you",
|
|
}
|
|
|
|
# A symbol id from templates/partials/icons.html. Everything used to be the
|
|
# sparkle, which said only "a model did something".
|
|
ICONS: dict[str, str] = {
|
|
"shell_run": "terminal",
|
|
"file_read": "file-text",
|
|
"file_write": "pencil",
|
|
"file_edit": "diff",
|
|
"file_list": "folder",
|
|
"plan_submit": "check",
|
|
"plan_update": "check",
|
|
"job_output": "clock",
|
|
"job_list": "dots",
|
|
"job_stop": "stop-circle",
|
|
"web_search": "globe",
|
|
"fetch": "link",
|
|
"knowledge_search": "archive",
|
|
"knowledge_get": "file-text",
|
|
"notes_search": "search",
|
|
"notes_get": "file-text",
|
|
"notes_create": "pencil",
|
|
"notes_edit": "pencil",
|
|
"notes_delete": "trash",
|
|
"memory_add": "star",
|
|
"memory_forget": "trash",
|
|
"skill_get": "sparkle",
|
|
"skill_create": "sparkle",
|
|
"skill_edit": "sparkle",
|
|
"ask_user": "chat",
|
|
}
|
|
|
|
# The icon for an event whose tool is not in the table -- a custom HTTP tool, an
|
|
# MCP tool, or a row written before `kind` existed.
|
|
KIND_ICONS: dict[str, str] = {
|
|
"search": "globe",
|
|
"fetch": "link",
|
|
"custom": "link",
|
|
"mcp": "server",
|
|
}
|
|
FALLBACK_ICON = "sparkle"
|
|
|
|
# What an approval card is headed. A sentence somebody agrees to, in the
|
|
# imperative, because that is what pressing the button does.
|
|
ACTIONS: dict[str, str] = {
|
|
"shell_run": "Run a command",
|
|
"file_read": "Read a file",
|
|
"file_write": "Write a file",
|
|
"file_edit": "Update a file",
|
|
"file_list": "List a directory",
|
|
"web_search": "Search the web",
|
|
"fetch": "Fetch a page",
|
|
"knowledge_search": "Search the library",
|
|
"knowledge_get": "Read a document",
|
|
"notes_search": "Search notes",
|
|
"notes_get": "Read a note",
|
|
"notes_create": "Write a note",
|
|
"notes_edit": "Change a note",
|
|
"notes_delete": "Delete a note",
|
|
"memory_add": "Remember something",
|
|
"memory_forget": "Forget something",
|
|
"skill_get": "Read a skill",
|
|
"skill_create": "Write a skill",
|
|
"skill_edit": "Change a skill",
|
|
"job_stop": "Stop a background job",
|
|
}
|
|
|
|
# Which argument is the thing being agreed to. Shown verbatim and escaped on the
|
|
# card: a summary that paraphrased it would be a card approving something other
|
|
# than what runs.
|
|
DETAIL_KEYS: dict[str, str] = {
|
|
"shell_run": "command",
|
|
"file_read": "path",
|
|
"file_write": "path",
|
|
"file_edit": "path",
|
|
"file_list": "path",
|
|
"fetch": "url",
|
|
"web_search": "query",
|
|
"knowledge_search": "query",
|
|
"notes_search": "query",
|
|
"job_stop": "id",
|
|
}
|
|
|
|
|
|
def _name_of(event: dict[str, Any] | str) -> str:
|
|
if isinstance(event, str):
|
|
return event
|
|
return str(event.get("name") or "")
|
|
|
|
|
|
def label_for(event: dict[str, Any] | str) -> str:
|
|
"""What to call this tool in the transcript.
|
|
|
|
The static table wins over anything stored on the event. See the module
|
|
docstring: rows already on disk carry the wrong label, and deferring to them
|
|
would leave every existing transcript naming a machine.
|
|
"""
|
|
name = _name_of(event)
|
|
if name in LABELS:
|
|
return LABELS[name]
|
|
if isinstance(event, dict):
|
|
stored = str(event.get("label") or "").strip()
|
|
if stored:
|
|
return stored
|
|
return name
|
|
|
|
|
|
def icon_for(event: dict[str, Any] | str) -> str:
|
|
"""A symbol id for this event, never empty.
|
|
|
|
Falls through the tool's own icon, then the event's `kind`, then the
|
|
generic one -- so a custom tool still gets a link and an MCP tool a server,
|
|
which is what the template used to decide for itself.
|
|
"""
|
|
name = _name_of(event)
|
|
if name in ICONS:
|
|
return ICONS[name]
|
|
kind = str(event.get("kind") or "") if isinstance(event, dict) else ""
|
|
if not kind and name == "web_search":
|
|
# Rows written before `kind` existed. The template made the same
|
|
# allowance for the same reason.
|
|
kind = "search"
|
|
return KIND_ICONS.get(kind, FALLBACK_ICON)
|
|
|
|
|
|
def describe(name: str, args: dict[str, Any]) -> tuple[str, str]:
|
|
"""What an approval card says about one call: a title, and the detail."""
|
|
title = ACTIONS.get(name)
|
|
key = DETAIL_KEYS.get(name)
|
|
if title is not None:
|
|
return title, str(args.get(key) or "") if key else ""
|
|
detail = ", ".join(f"{k}={v!r}" for k, v in list(args.items())[:4])
|
|
return f"Use {label_for(name)}", detail[:400]
|
|
|
|
|
|
__all__ = [
|
|
"ACTIONS",
|
|
"DETAIL_KEYS",
|
|
"FALLBACK_ICON",
|
|
"ICONS",
|
|
"KIND_ICONS",
|
|
"LABELS",
|
|
"describe",
|
|
"icon_for",
|
|
"label_for",
|
|
]
|