Files
LLeMbas/src/lembas/services/agent/session.py
T
Jaroslav Beneš 3fc3449726 Let a command run in the background instead of being killed
An agent command is one blocking conn.run over a per-call connection, killed the
moment it hits its timeout -- so a ten-minute apt install is impossible, which is
exactly what a user hit. This is the substrate for running it detached instead:
the model can ask for background=true, or a command that outlasts its timeout is
kept running rather than killed, and either way the model gets tools to read and
stop it. Opt-in, off by default, under Admin -> Agents; off is byte-for-byte the
old behaviour.

The mechanism has to survive the connection closing (that is the whole premise
of the per-call model), so a job is a setsid-detached process on the far side,
redirected to a remote logfile and an exit-file; LLeMbas reconnects, as always,
to read it later. services/agent/jobs.py holds the wrappers.

Three things in those wrappers are load-bearing and each was got wrong in the
first sketch:

- The command never touches a quoted shell context. sh -c '<cmd>' shatters the
  instant the command contains a quote -- git commit -m 'fix', awk '{…}', sed
  's/…/…/' are the common case, and it is an injection hole besides. So the
  command is base64-encoded in Python and decoded on the far side into a script
  file; it is bytes, never shell syntax.
- The child records its own pid via $$ as its first act, under setsid where it
  is the session leader, so job_stop can kill the whole process group. echo $!
  from the launcher captures the wrong pid.
- The command's exit status comes from the exit-file, never the wrapper's own
  status -- which is ~0 from its trailing rm. Reading the wrapper's status would
  mark every job a success.

A command that finishes in time is indistinguishable from a foreground one --
same output, same wording; the difference shows only when it does not, where
instead of "stopped after Ns" it becomes a job id. Auto-convert is its own
sub-switch: with it off, a timeout stays a hard stop and nothing is left
running, because routing the plain case through the detached wrapper would leave
an orphan running past a stop an administrator asked for.

New agent tools job_output/job_list/job_stop, offered only when the feature is
on (the plan_submit gating pattern); job_stop is RISK_EXECUTE since it kills a
process. A job's files are namespaced by the calling chat's id and the wrappers
are always built from it, so a model in one chat cannot even name another's job.

Tested against a real local /bin/sh rather than the fake echo-the-command sshd
fixture, because the shell logic -- setsid, base64, the wait loop, the child
surviving the wait being cut off -- is the whole of the risk. The auto-wake that
prompts the model back when a job finishes is the next commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:15:49 +02:00

200 lines
8.9 KiB
Python

"""What one agent chat is pointed at, resolved while a session is open.
Everything a runner needs travels in `AgentContext`: the machine, the decrypted
credential, the mode in force, and the two lists that adjust it. Nothing is
looked up later, for the reason `Endpoint` is a frozen copy of a `Connection`
and `ToolContext` carries an owner id rather than a `User` -- a generation
outlives the request that started it, and a detached instance is a trap.
The mode is read **once, at the start of the reply**, and deliberately does not
change under a reply already in flight. Somebody switching to Auto halfway
through must not retroactively approve what is already queued.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass, field, replace
from typing import Any
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import KIND_AGENT, Chat, SshProfile, User
from lembas.services import settings_store
from lembas.services.agent import policy
from lembas.services.agent import ssh as ssh_service
from lembas.services.agent.base import Executor
from lembas.services.agent.policy import Limits
log = logging.getLogger(__name__)
@dataclass
class AgentContext:
"""The machine an agent chat acts on, and what it may do there."""
chat_id: str
label: str
project_dir: str
# The connection's id, carried so a runner can drop the project listing it
# has just invalidated. `index` is keyed on the connection and the
# directory, not on the chat -- two chats on one tree share a listing.
profile_id: str = ""
mode: str = policy.MODE_MANUAL
allow: tuple[str, ...] = ()
deny: tuple[str, ...] = ()
limits: Limits = field(default_factory=Limits)
# Per-command bounds, from the instance settings.
timeout: float = 60.0
max_timeout: float = 600.0
max_output: int = 64 * 1024
# The decrypted credential. Held here and nowhere else, and cleared by
# `generation` when the reply ends -- a finished Generation lingers five
# minutes so late followers get the final frames, and a private key should
# not linger with it.
spec: dict[str, Any] = field(default_factory=dict)
# Set only on the per-call copy handed to a runner whose call a person has
# just allowed. The runners re-check the mode as a backstop, and without
# this they would refuse the very thing that was approved -- the mode says
# "ask", and asking is exactly what happened.
approved: bool = False
# Absolute paths this reply has read. `file_edit` refuses a file that is not
# in here, because a patch written from memory against a file the model has
# not looked at is how a rewrite silently loses somebody's work.
#
# Here rather than on `Generation` for two reasons. Runners never see a
# Generation -- they get a `ToolContext`, which is a session-free snapshot
# precisely so nothing in a tool holds live state -- and a read path is a
# fact about the machine, which is what this class is.
#
# It is **shared with the approved copy**: `as_approved` is
# `dataclasses.replace`, which copies field references, so a path read
# through an approved call is visible here. That is wanted and is not
# obvious, so there is a test for it.
#
# It resets each reply, and that is correct rather than a limitation.
# `Message.tool_calls_json` is deliberately never replayed as context, so on
# the next turn the model does not have the file's contents either --
# requiring a re-read in the reply that edits is asking for something it
# needs anyway.
read_paths: set[str] = field(default_factory=set)
# The plan currently in force, seeded from `chat.plan_message_id` when this
# is resolved. Mutable and read/written in place by `plan_update`, for a
# reason that is not obvious: a runner cannot write the message row --
# `_persist` is the single writer -- so it returns the merged plan on its
# event and the loop carries it. Two updates in one reply would then both
# read the same stale plan from the database and the second would lose the
# first. This snapshot is what they actually merge into.
plan: dict[str, Any] = field(default_factory=dict)
# Whether commands may run detached. When off, `shell_run` is byte-for-byte
# what it always was and the `job_*` tools are not offered -- a command that
# times out is killed, as before. When on, a command can be launched in the
# background (or converted to one when it times out) and the model gets the
# tools to check on it. `on_timeout` is the sub-switch for the auto-convert.
background: bool = False
background_on_timeout: bool = True
# Whether a finished job wakes the model on its own, rather than only being
# seen when it next runs. Read by the wording here and by the watcher.
background_notify: bool = True
def executor(self) -> Executor:
return ssh_service.SshExecutor(self.spec, self.project_dir)
def clear(self) -> None:
self.spec = {}
def as_approved(self) -> AgentContext:
"""A copy of this context for one call a person has allowed."""
return replace(self, approved=True)
def _plan_of(db: DBSession, chat: Chat) -> dict[str, Any]:
"""The plan this chat is working to, or an empty dict.
One `db.get` by primary key -- the column exists to avoid a scan for "the
newest message carrying a plan", because this runs while a request is
waiting. The id is validated here rather than constrained in the schema, for
the reason the column's comment gives.
"""
from lembas.db.models import Message
from lembas.services import plans
if not chat.plan_message_id:
return {}
message = db.get(Message, chat.plan_message_id)
if message is None or message.chat_id != chat.id:
return {}
return plans.normalise(message.plan_json)
def profile_for(db: DBSession, chat: Chat, user: User | None) -> SshProfile | None:
"""The connection this chat is pointed at, if it is still usable.
Ownership is re-checked here rather than trusted from when the chat was
created: a profile can be deleted, disabled, or moved to a host whose key
has not been confirmed since, and any of those should stop the chat acting
rather than be discovered at the first command.
"""
if chat is None or chat.kind != KIND_AGENT or not chat.ssh_profile_id:
return None
profile = db.get(SshProfile, chat.ssh_profile_id)
if profile is None or not profile.enabled:
return None
if user is not None and profile.owner_id != user.id:
return None
return profile
def resolve(db: DBSession, chat: Chat, user: User | None) -> AgentContext | None:
"""This chat's agent setup, or None if it has none it can use.
None is the answer to every "no": not an agent chat, the feature switched
off, the connection gone or disabled, SSH not installed. Each of those means
the agent tools are not offered at all, which is better than offering a tool
that fails on its first call.
A profile whose host key has never been confirmed is deliberately *not* one
of them. The tools are offered and the failure is explicit, because "check
the connection and accept its fingerprint" is a thing the reader can act on,
while a silently missing tool is not.
"""
profile = profile_for(db, chat, user)
if profile is None:
return None
values = settings_store.agents(db)
if not values.get("enabled"):
return None
if ssh_service.available():
return None
return AgentContext(
chat_id=chat.id,
label=profile.label,
plan=_plan_of(db, chat),
project_dir=chat.project_dir or profile.default_dir or "",
profile_id=profile.id,
mode=chat.agent_mode if chat.agent_mode in policy.MODES else policy.MODE_MANUAL,
allow=tuple(values.get("allow_default") or ()),
deny=tuple(values.get("deny_default") or ()),
limits=Limits(
steps=int(values.get("max_steps") or 200),
wall_seconds=float(values.get("max_wall_seconds") or 900),
output_bytes=int(values.get("max_total_output_bytes") or 1024 * 1024),
# `or 0` would turn a deliberate 0 into the default, and 0 is how an
# administrator says "no ceiling". `agents()` has already clamped it.
completion_tokens=int(values.get("max_completion_tokens", 200_000) or 0),
),
timeout=float(values.get("default_timeout") or 60),
max_timeout=float(values.get("max_timeout") or 600),
max_output=int(values.get("max_output_bytes") or 64 * 1024),
background=bool(values.get("background_enabled")),
background_on_timeout=bool(values.get("background_on_timeout", True)),
background_notify=bool(values.get("background_notify", True)),
spec=ssh_service.spec_from(profile),
)
__all__ = ["AgentContext", "profile_for", "resolve"]