Files
LLeMbas/CLAUDE.md
T
Jaroslav Beneš 131a4083f8 The terminal learns where one command ends, and can be dragged wider
"The last command and its output" was not something the panel could honestly
offer. sendToChat took the last forty rows of the screen buffer, hard-wrapped at
the terminal's width with no way to tell a wrap from a newline -- its own comment
said so. So bash and zsh are given the OSC 133 markers VS Code and WezTerm use,
and Copy, Send and an Auto toggle are built on those.

The integration is written by the PTY command string itself, with printf. sshd
runs that string through $SHELL -c, so it can case on the shell's own name and
needs no probe, no second channel and no writable home. Passing it through the
environment does not work -- every distribution ships AcceptEnv LANG LC_*, so
anything else is dropped silently -- and feeding `source ...` in as keystrokes
races a slow .zshrc, echoes into the scrollback and lands in shell history.

Nothing needs hiding, which is the point of choosing it: the setup runs before
the shell exists and never writes to the PTY's input side, so there is nothing
to echo and no fan-out gate to build.

Two things were wrong in the first version and both were found by running it
against real shells rather than the fake one. bash: the DEBUG trap fires before
every simple command *including each one inside PROMPT_COMMAND*, so $? read from
there is whatever ran a moment ago -- every command reported success. The status
is captured in the trap now, which also removes the two-entry PROMPT_COMMAND
dance entirely. zsh: $ZDOTDIR is already ours by the time .zshenv runs, so the
shims were sourcing themselves and none of the user's configuration loaded; the
original is passed on the exec line.

Parsing is server-side. The `behind` path resets the terminal and replays a
truncated scrollback, so a client parser routinely sees a finish with no start;
two tabs share one shell and can disagree; and what comes out of this ends up
inside a prompt, so deriving it here leaves nothing to disbelieve. The bytes are
fanned out unchanged -- xterm consumes an OSC it has no handler for.

Output is bounded head and tail, 48KB and 16KB: a build that fails ten megabytes
in has the invocation at the top and the error at the bottom. Carriage returns
collapse to the last state of each line, which is the difference between a
usable prompt and two megabytes of spinner. The fence is sized to its content,
because output containing three backticks would otherwise break out and read as
prose.

Any shell that is not bash or zsh starts exactly as it did before. The buttons
then scrape the screen and say so, and Auto is disabled rather than degraded:
forty arbitrary lines on every message is worse than nothing.

Also a generic [data-resize] handle, keyboard included, persisted the way the
theme is. The inspector and sidebar can have it whenever they want it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 17:28:38 +02:00

926 lines
54 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
Working notes for LLeMbas. Read this before changing anything.
## What it is
A self-hosted web UI for OpenAI-compatible LLM endpoints, written in Python and
themed after Middle-earth. Server-rendered FastAPI + Jinja + htmx; SQLite;
no JavaScript build step.
## Commands
```bash
. .venv/bin/activate
pip install -e ".[dev,search,ssh]" # `search` adds ddgs for DuckDuckGo,
# `ssh` adds asyncssh for agent chats
lembas serve # http://127.0.0.1:8080
lembas info # paths + counts, useful when confused
lembas secret-key # generate LEMBAS_SECRET_KEY
lembas create-admin # create or promote an admin
pytest # 891 tests, ~55s
# PLAN.md tracks what is and is not built
ruff check . # lint (line length 100)
python scripts/build_artwork.py # regenerate artwork (SVG + PWA icons;
# needs fonttools and cairosvg)
python scripts/fetch_vendor.py # verify vendored JS against the lockfile
```
## Hard rules
These are the constraints the project is built around. Breaking one is a
redesign, not a tweak.
1. **No Node, no npm, no build step.** Browser libraries are downloaded once by
`scripts/fetch_vendor.py`, hash-pinned in `scripts/vendor.lock.json`, and
committed under `web/static/vendor/`. Adding one is a `--update`, same as
bumping one: a name in `PACKAGES` with no lock entry has nothing to verify
against, and the script refuses rather than writing it unpinned. xterm is
the one heavy dependency — 280KB, more than everything else together — and
is loaded only on a chat that can open a terminal.
2. **Nothing loads from a CDN at runtime.** A self-hosted tool must work
offline and must not report page views to a third party.
3. **No hard-coded values outside `tokens.css`.** Every colour, space, radius
and control height resolves through a CSS variable. `--control-h` is why
buttons, inputs and selects line up: they all take their height from it, so
a mixed row is flush by construction rather than by nudging.
4. **Additive-only schema changes.** SQLite only, no Alembic. `init_db()` runs
`db/migrations.py:sync_schema()`, which creates missing tables *and* adds
missing columns by diffing the models against the database. Renames, drops
and retypes are still manual. See "Changing the schema" below.
5. **Secrets never reach the browser.** API keys are Fernet-encrypted at rest
and only ever rendered masked.
6. **Model output is untrusted.** Everything from an endpoint goes through
`services/markdown.py` (markdown-it → nh3) or `escape_text()`. Never
`|safe` on anything that has not.
## The flavour rule
Middle-earth lives in the **artwork, theme names, empty states, loading lines
and error pages**. It does not live in the functional UI.
Chats are called *Chats*, not *Tales*. Folders are *Folders*, not *Chapters*.
Buttons say what they do. Someone who has never read the books must be able to
use this without a glossary. The two themes are named `moria` and `shire`, and
the 404 says "Not all those who wander are lost. This page, however, is." —
that is the right amount.
## Layout
```
src/lembas/
main.py app factory, lifespan, error handlers
config.py pydantic-settings, all LEMBAS_* variables
cli.py typer entry points
api/
deps.py Db / CurrentUser / RequiredUser / AdminUser
auth.py register, login, logout
pages.py full-page routes (chat shell, settings)
chats.py messaging + the SSE stream
folders.py folder CRUD
admin.py connections + instance settings
admin_models.py model ordering, defaults, images, access
admin_users.py users, groups, permissions
admin_audio.py speech-to-text and text-to-speech endpoints
admin_search.py web search provider and credentials
admin_prompts.py the prompt fragment editor and its preview
admin_suggestions.py the cards offered on the new-chat screen
admin_tools.py custom HTTP tools and MCP servers
admin_agents.py whether agent chats exist, and what they may spend
agents.py SSH connections, kept by the people who own them
terminal.py the terminal panel's WebSocket, and its two locks
audio.py transcribe, speak, voice discovery
library.py knowledge, notes, skills pages; memory CRUD
files.py upload, serve, remove attachments
preferences.py per-user theme, default model, password, audio
db/
base.py Base, UUID/Timestamp mixins
session.py engine, SQLite pragmas, init_db, session_scope
migrations.py additive schema sync (tables + columns)
models/ user, chat, connection, setting
security/ passwords (argon2), sessions, permissions
services/
llm/openai_client.py httpx streaming + model discovery
search/ ddgs, SearXNG and Firecrawl behind one shape
library/ documents, notes, memories, skills, FTS
mcp/ remote MCP servers: framing, transport, rows to tools
agent/ agent chats: the mode table, SSH, the four tools,
terminal.py (shells held open behind the panel),
shell_marks.py + capture.py (where one command ends),
index.py (what is in the project directory)
audio.py OpenAI-shaped /v1/audio/* client
fetch.py URL retrieval, HTML to text, the SSRF guard
sharing.py one visibility rule for every library store
prompts.py every injected prompt fragment, and {{variables}}
metrics.py tokens, context percentage and tokens/second
tokens.py the chars/4 estimate, for endpoints that report none
compaction.py summarising the earlier turns of a long chat
suggestions.py new-chat starting points, seeded once
harness.py the operational prompt built from what a model has
tools.py tool registry, schemas, streamed-call reassembly
custom_tools.py the admin-defined HTTP tool runner
tool_access.py who may be offered which admin-defined tool
interaction.py pausing a reply to ask the reader something
chat.py request building, endpoint resolution, titles
markdown.py markdown-it + pygments + nh3
crypto.py Fernet encrypt/decrypt/mask
files.py attachment validation, images, PDF/text extraction
reasoning.py splits thinking from the answer
settings_store.py runtime instance settings
uploads.py validated image storage
sse.py event framing
web/
templating.py render() -- always use this, not TemplateResponse
templates/ Jinja
static/ css, js, vendor, img, sw.js
js/commands.js the / table, and the keyboard that does the same jobs
js/composer.js the menu that / and @ open, over the message box
assets/ SVG masters and PWA icons (generated)
deploy/ systemd unit, nginx vhost, install/update scripts
```
## Things that will bite you
**`render()`, not `TemplateResponse`.** `web/templating.py:render()` injects
`user`, `theme`, `layout`, `version` and `allow_signup`. Templates assume they
exist. If you must call `templates.TemplateResponse` directly (the SSE path does,
because
there is no `Request`), pass `user` explicitly — `chat/_message.html` renders
both roles and the user branch dereferences it.
**The message template is the state machine.** `chat/_message.html` renders an
incomplete assistant message as a streaming shell carrying `sse-connect`, and a
complete one as finished output. That is the *only* thing that starts a
generation. A consequence worth knowing: loading a page whose last reply is
unfinished restarts it, which is how a dropped connection recovers.
**SSE framing.** `services/sse.py:event()` splits payloads on newlines into
several `data:` lines. A raw newline in a single `data:` line truncates the
event — the failure shows up the first time a model emits a code block.
**Streaming opens its own database session.** `api/chats.py:_generate()` uses
`session_scope()`, not the request's session, because streaming outlives the
request handler.
**Escaping is chunk-safe on purpose.** `escape_text()` is `html.escape`, which
works character by character, so escaping stream chunks separately equals
escaping the whole string. `nh3.clean_text` would also be safe but escapes
spaces and slashes, tripling the size of every streamed token.
**The fence renderer is replaced, not configured.** markdown-it's `highlight`
option re-wraps output in `<pre><code>` unless the string starts with `<pre`,
which would nest a second `<pre>` inside our wrapper. `markdown.py` overrides
`renderer.rules["fence"]` instead. There is a regression test for this.
**SVG `<style>` is document-scoped.** Two text runs in one SVG sharing a class
name means the later rule recolours both. `build_artwork.py` takes class names
as parameters for exactly this reason.
**Gradient ids are document-global.** The `mark()` macro takes a `uid` because
two marks on one page with identical ids make the second silently reuse the
first one's gradients.
**Two kinds of settings.** `lembas.config` is deployment configuration read
from the environment at startup. `services/settings_store.py` is instance
settings an admin edits at runtime, stored in the `settings` table. Environment
variables seed the latter as an *initial* value only — once stored, the database
wins, or a toggle in the UI would silently revert on the next restart.
**Permissions are a union, and admins bypass them.** `security/permissions.py`
resolves a baseline (instance setting) widened by each group. A group grants;
it never denies — otherwise "why can this user not do X" needs a simulation of
every group to answer. Model *access* is separate: `models_visible_to()`.
**FastAPI cannot tell an empty form field from an absent one.** With
`x: str | None = Form(None)`, a submitted `x=` arrives as `None`, so "clear this
field" is indistinguishable from "leave it alone". `api/chats.py:update_chat`
reads `await request.form()` and checks key presence instead. Anything with a
clearable field must do the same.
**`Mapped[list]` without an element type is not a collection.** SQLAlchemy
treats a bare `Mapped[list]` as a scalar and hands back `None` instead of `[]`.
Always write `Mapped[list[Group]]`, with a `TYPE_CHECKING` import if the class
lives in another module.
**Reasoning arrives two ways.** A `reasoning_content` delta field (llama.cpp,
llama-swap, vLLM) or `<think>` tags inline in `content` (Ollama and friends).
`services/reasoning.py` handles the second with a streaming splitter, because
the tags arrive split across chunks. Reasoning is stored in `Message.reasoning`
and is deliberately **not** replayed as context on the next turn.
**Attachments are typed by their bytes, not their name.** `services/files.py`
sniffs magic numbers; a `.png` full of text is stored as text. Images are
downscaled and re-encoded (a phone photo is megabytes of base64), PDFs have
their text extracted **once at upload** — re-extracting per request would let a
reply change because a parser was upgraded.
**Images only go to models marked `vision`.** Sending content parts to an
endpoint without multimodal support is not graceful degradation; most reject
the whole request. `build_request()` checks the capability and falls back to a
plain string. A plain text turn must *stay* a plain string for the same reason.
**Attachments are served, never linked.** Images reach the model as base64 data
URIs: a local endpoint has no route back to LLeMbas and a hosted one has no
credentials. Non-images are served `Content-Disposition: attachment` with
`nosniff`, so an uploaded `.html` cannot execute in this origin.
**Uploads are unbound until the message is sent.** `Attachment.message_id` is
null in the composer; `files.claim()` binds them, and only unclaimed rows owned
by that user, so a forged id cannot pull in someone else's file. Abandoned ones
are swept at startup.
**Generation is a background task; the SSE endpoint only follows it.**
`services/generation.py` owns the work and the registry; `api/chats.py:_follow`
watches a `Generation` and streams what it sees. Closing the connection does
NOT stop the reply -- that was the old behaviour and it cut answers off when
the reader navigated away. Any route that creates an assistant placeholder must
also call `generation.ensure()`.
**`ensure` attaches, `restart` replaces.** The registry is keyed on message id
and finished generations linger `KEEP_FINISHED` so a follower arriving at the
last moment still gets the final frames. `ensure` is idempotent because a page
load finding an unfinished reply must attach rather than start a second one.
Regeneration is the only caller that reuses a `Message` row, and therefore the
only one for which idempotence is wrong -- it got the finished generation back,
made no request, and left the browser reconnecting to a stream with nothing to
say. It calls `restart`. `_persist` refuses to write when another generation
owns the message, because a cancelled predecessor's `finally:` still runs.
**The row is written before `done` is set.** `_follow` breaks out the instant it
sees that flag and re-renders the bubble from the database, so the row has to be
authoritative first. The other order silently showed the previous turn's stored
metrics.
**Stream frames carry whole blocks, not deltas.** `render`, `reasoning`,
`metrics` and `status` all send the complete value each time, and every one of
them is swapped with `innerHTML`. `reasoning` used `beforeend` and so repeated
everything already shown on every frame. That is what makes reattaching mid-reply
work: a follower arriving late has no earlier fragments to append to. It also
means Markdown is re-rendered whole, which is required anyway -- a list or code
fence is only correct once its context exists.
**Two frames must be able to blank themselves, and the rest must not.**
`reasoning`, `tools` and `render` are only sent when they have something in
them, so a frame can never wipe what is on screen. `metrics`, `status` and
`ask` are sent on every version bump *including empty*, because each has to be
able to clear: an approval card that survived being answered would be a button
you could press twice.
**Stopping sets a flag the producer checks -- except while it is paused.**
`generation.request_stop()`; whatever arrived is kept and the message is marked
`stopped`, distinct from `error`. In-process, so single-worker only. `cancel` is
read in exactly one place, between streamed chunks, and a reply waiting on an
approval produces no chunks -- so `request_stop` also resolves
`generation.pending`, and that is the wakeup. Without it the Stop button does
nothing at all while a card is on screen, silently.
**Asking a person is one primitive with three uses.**
`services/interaction.py`: a command waiting to be allowed, a question the model
asked, and "this reply is waiting for you" are all *pause, render a block in the
bubble, wait for a POST, resume*. It pauses a **round, not a call** -- a round's
calls run together under a semaphore, and parking four coroutines on four
separate answers inside that gather would queue them behind each other
invisibly, and hand the reader four cards for commands whose order matters. So
one card covers everything in the round, and `_authorise` returns pre-decided
outcomes keyed by call index, which is what keeps
`zip(calls, outcomes, strict=True)` aligned.
**One `ask_user` call may carry several questions, and they come back at once.**
Each becomes an `Item` with its own `key`; several items can share an `index`
because they belong to one call, and one tool turn answers them all with each
answer quoted beside its question. Asking one at a time would cost a round trip
and an interruption each, and answering the third would mean having forgotten
the first. `_questions_in` reads the singular form and bare strings too: a small
model sends something close to the schema rather than the schema, and getting it
wrong costs a whole round trip to show a card that says nothing.
**A paused reply is deliberately not `done`.** That is what lets a page reload
reattach to it. Its *timeout* is what stops it lingering, not `_prune`, which
only drops finished ones -- so `approval_timeout` is clamped to at least a
minute on read, and `_prune` resolves anything whose deadline is long past as a
backstop.
**Nothing is persisted while paused.** A restart abandons the pending question
along with the reply, and a reload starts the turn afresh -- the model asks
again. That is consistent with "a restart abandons replies in flight", but it
means an approval is not a durable record of consent.
**A chat's kind and connection are fixed at creation; only the mode moves.**
`Chat.kind`, `ssh_profile_id` and `project_dir` are chosen on the new-chat screen
and refused by `update_chat` thereafter with a 409 — a transcript whose earlier
turns ran somewhere else is not one conversation. `agent_mode` is the exception
and changes freely: it decides what gets asked about, not what the conversation
is. The mode is read **once per reply**, so switching to Auto mid-reply cannot
retroactively approve what is already queued.
**The mode is enforced in the loop, never in the prompt.** `_authorise` consults
`agent/policy.py:decide()` server-side, keyed on each `ToolDef.risk`. A model is
*told* which mode it is in so it behaves sensibly, but everything it reads — a
web page, a README, the output of the last command — is untrusted, and a rule
living only in a system message is one a poisoned file can argue with. Within an
agent chat **every** call goes through the table, including the built-ins:
`notes_edit` writes, and Plan mode meaning "look but do not touch" has to mean
that too.
**An approved call needs telling.** Every agent runner re-checks the mode as a
backstop, so a call arriving by a path that skipped `_authorise` cannot walk
past it. That backstop refused the very thing a person had just approved — the
mode says "ask", and asking is exactly what happened. `AgentContext.approved` is
threaded per call on a *copy* of the context, because a round runs its calls
together and only some of them were allowed.
**`registry(db)` must know every tool that can be offered, agent tools
included.** It maps an offered tool *name* back to a family, which is how the
harness decides that `tool.agent` applies. They are listed there unbound to any
chat. Without them `shell_run` resolves to no family, and an agent chat is told
nothing about the machine it is working on. The identical omission cost custom
tools their guidance once already; there is a test for it now.
**A tool description is schema; the harness is where "where" lives.**
Descriptions are sent verbatim and are deliberately not editable, so they state
facts about the runner. Which machine, which directory and which mode belong to
*this chat* and live in the `tool.agent` fragment, where they can change without
the schema shifting under a model mid-conversation.
**Each command is a fresh shell.** Connections are per call, so `cd build`
followed by `make` fails silently — `cwd` is a first-class parameter reaching the
executor, never spliced into the command string. This is the likeliest single
cause of "the agent seems stupid", and the harness says it out loud. So does the
other one: on a Debian-derived host `apt-get install` reports the package missing
until `apt-get update` has run.
**Files never go through a shell.** The SSH exec protocol carries one command
*string* that the far side parses, with no argv form at all, so a model-supplied
path in a command line is unavoidably a quoting problem. `file_read`/`file_write`
/`file_list` use SFTP, where a path is a path.
**asyncssh's defaults are wrong here, all four of them.** Every LLeMbas user
shares one unix account, so `known_hosts` unset reads a *shared* trust store
(and `None` disables checking entirely), `client_keys` unset loads whatever is in
`~/.ssh`, `config` unset lets a `ProxyCommand` redirect the connection, and
`agent_path` unset uses `$SSH_AUTH_SOCK`. All four are passed explicitly on every
connection, and the test that proves it needs no server.
**A pinned host key belongs to a host and a port.** Moving a profile forgets it
deliberately. `capture_host_key` completes the key exchange and stops, so a host
that has not been accepted is never offered a username, let alone a credential —
which is what makes accepting a fingerprint from a button safe.
**A plan ends the turn, but not mid-sentence.** `plan_submit` is offered in Plan
mode only, and the round after it runs with the tools withdrawn: the model gets
to say what it proposed, and cannot spend three more rounds changing its mind
about a plan somebody is being asked to approve. Carrying it out switches to
**Edit, never Auto**, and the plan goes back quoted and attributed rather than
stated — text that came out of a file the model read must not arrive wearing the
reader's authority.
**Rewind rewinds the transcript, not the machine.** Editing or regenerating in an
agent chat stamps `Chat.rewound_at` and the harness warns that files from steps
no longer in the transcript are still there. Nothing tries to undo them: the
project directory is somebody's real working tree, and deleting their work to
match would be far worse than the inconsistency.
**The project listing is read from a cache and never fetched.**
`harness.context_variables` runs synchronously on the request path, so
`agent/index.py:cached()` is all it may call — an SFTP round trip from there
would hold a request open while somebody's box thought about it. The walk
happens in `generation._warm_index`, which is async and already doing network
work, with a short wait. A chat whose first reply outruns its first walk simply
has no listing that turn, and the fragment's `requires` makes it vanish rather
than appear as an empty heading. Anything else wanting the listing gets the same
deal: the `@` picker offers no files until one exists, because a keystroke must
never wait on a machine.
**A listing is budgeted, not dumped.** A tree of a thousand files costs the
window on every request forever and buries the four names that mattered.
`index.render` collapses what will not fit to `src/vendor/ (412 files)` and says
so. Collapsing picks the **deepest and largest first**: by saving alone it would
take `src/` before `src/web/static/vendor/`, because it contains it, and lose
every name worth having. Watch the double-count — collapsing a parent subsumes a
child already collapsed, and adding both savings stops the loop early believing
it has made room it has not.
**A slash command must never swallow a message.** `static/js/commands.js`
intercepts only an exact match against its table; `//` escapes, and anything
unrecognised is sent as written. Eating somebody's message because it began with
a slash is a far worse failure than an unknown command, and it is the one the
implementation has to be arranged around rather than patched for afterwards.
**`@` inserts a reference *and* attaches the contents.** The token stays in the
sentence so "change the thing in @main.py" reads as one, and the file arrives as
an attachment chip — the same component every other attach path returns, so the
composer learns nothing new. `Attachment.source_path` and `source_label` carry
where it came from into `chat.document_context`'s tag, because a model handed
`main.py` cannot tell which of four it is looking at and cannot name it back when
asked to change something. Those two are attribute values in a tag we write, so
`_attr` strips quotes and angle brackets rather than escaping them.
**Unread is polled, not pushed.** A browser on another chat has no connection
to the one that finished. `/api/chats/unread` returns out-of-band dot spans and
an `HX-Trigger` for the toast; `unread_notified` stops the same arrival being
announced every tick. Re-rendering the whole sidebar instead would reset the
folder open/closed state every 10 seconds.
**Editing rewinds, it does not branch.** `POST .../messages/{id}/edit` rewrites
a user turn and **deletes everything after it**. Branching would need a UI for
choosing between versions; "go back and try again from here" is what was asked
for and what other clients do. `Message.parent_id` still exists unused.
**Dialogs and toasts are ours, not the browser's.** `static/js/ui.js` provides
`lembas.notify/confirm/prompt`, and intercepts htmx's `htmx:confirm` so every
existing `hx-confirm` gets the themed dialog with no change at the call site.
Plain forms opt in with `data-confirm`, lone submit buttons with
`data-confirm-button`. Never add a `window.confirm` back.
**The model picker is hand-built.** A `<select>` renders only text in an
`<option>` -- no avatar, no description, no badges. `chat/_model_picker.html`
plus the picker block in `ui.js`; the value lives in a hidden input so it still
behaves as a form field.
**Chats are created lazily.** There is no endpoint that makes an empty chat.
"New chat" is a link to `/chat`, which renders a composer with no row behind
it; `POST /api/chats/start` writes the chat together with its first message.
That is why an opened-and-abandoned chat never appears in the sidebar. Tests
that just need a chat use the `make_chat` fixture rather than the HTTP flow.
**Admin lists are list-plus-detail, never a form per row.** `/admin/models`
renders compact rows with search, filter tabs and pagination; the full form
lives at `/admin/models/{id}/edit`. A connection can advertise a hundred models,
and a page that renders a form for each is unusable. Any future admin list
(tools, agents) should follow the same shape.
**Route order matters for static path segments.** FastAPI matches in
registration order, so `/admin/models/bulk` must be registered *before*
`/admin/models/{model_id}` or "bulk" is parsed as a model id and 404s. This has
already been a bug once.
**Pinning is not ordering.** The model picker is always in the administrator's
`position` order. Pinned models get shortcuts in the chat sidebar and nothing
else -- a picker whose order differs from the admin screen is just confusing.
**System prompts are precedence, not concatenation.** chat > model > instance,
most specific wins outright (`services/chat.py:effective_system_prompt`).
Stacking them reads well in a settings screen and badly in practice: two layers
that disagree give the model contradictory instructions and nobody can tell
which is losing.
**JSON columns need reassignment.** `user.settings_json["theme"] = x` on a
plain dict is not detected. The columns use `MutableDict` (`db/types.py`), but
the safe habit is `obj.field = {**obj.field, "k": v}`.
**`[hidden]` needs `!important`.** The browser's rule is `[hidden] { display:
none }`, which any class setting `display` outranks — and `.btn` is
`display: inline-flex`. That is not theoretical: it is why the old Stop button,
created and then `hidden = true`, sat permanently beside Send. `app.css` forces
the attribute to win. Anything toggled with `hidden` depends on that line.
**Send and Stop are one button.** `chat/_composer.html` renders both icons and
`ui.js` flips `data-composer-action` plus `type` (`submit``button`) when a
message in the thread is still streaming. Do not add a second button back.
**The tool loop is inside one generation.** `services/generation.py:_run()` runs
up to `tools_service.MAX_ROUNDS` request rounds for a single reply: stream,
accumulate tool calls, run them, append the results, ask again. `Generation`
accumulates content across all of them, so text emitted before a tool call
survives. Tools are only offered when search is enabled, the user has
`tools.web_search`, **and** the model is flagged `tools` — sending a `tools`
array to an endpoint without support fails the whole request, exactly as images
do without `vision`.
**The registry is resolved per request, not imported.** `REGISTRY` holds the
built-ins; a custom tool or an MCP tool is a row. `tools.resolve_tools()` returns
a `ToolSet` carrying the schemas *and* the runners, and the runners travel to the
loop on `ToolContext.tools` — because a generation outlives the session that
could look them up. `run_tool` consults that map, so what may be *run* is what
was *offered*. `None` means nobody resolved a set and falls back to the built-ins;
an empty dict is authoritative. Reaching for the global registry instead is how a
model naming a tool its chat was gated out of used to get it run anyway.
**A row-backed tool's family is `gate:slug`.** `custom:weather`, `mcp:github`.
The capability flag and the permission are named after the *gate*
(`tool_custom` / `tools.custom`), so a server advertising forty tools does not
mean forty checkboxes on every model; the full family exists so each row can
carry its own prompt fragment, gated to appear exactly when its tool is offered.
`harness._families` and `admin_prompts` therefore take a `db`. Custom and MCP
deliberately do **not** require `library.use`: an endpoint an administrator wrote
has nothing to do with anyone's own notes.
**An argument may fill a hole; it may never move the target.** A custom tool's
URL is a template. The scheme and host must be literal — checked at save *and*
again at call time, since a row can predate a check — values are escaped for
where they land (`quote(safe="")` in a URL, JSON-escaped in a body, control
characters stripped in a header), and the filled URL's origin is compared with
the template's afterwards. An undeclared `{{name}}` becomes nothing rather than
passing through, which is the opposite of `prompts.substitute` and deliberately
so: a literal `{{x}}` in a URL is not a feature.
**Three places now follow redirects by hand.** `fetch.fetch`,
`custom_tools._send` and `mcp.client.Session._post`, each re-running
`check_url` on every hop. `fetch()` itself is not reusable — GET-only,
bodyless, and it *raises* on any content type that is not HTML or text, which is
every JSON API there is. The duplication is deliberate; bending a page fetcher
into a general HTTP client is not. A secret is dropped when a hop leaves the
origin it was issued for.
**MCP sessions are per call.** Initialize, `notifications/initialized`, the call,
then a best-effort `DELETE`. Caching one would need an owner, a TTL, eviction, a
lock (a round runs its tools concurrently) and a shutdown hook, and the server
can expire it underneath all of that anyway — `ToolContext` is a session-free
snapshot precisely so nothing in a tool holds live state. The cost is one POST in
front of a call that is already a network round trip. `307`/`308` are followed;
`301`/`302`/`303` turn a POST into a GET and are refused rather than guessed at.
**A discovered MCP tool is JSON, not a row.** `McpServer.tools_json` caches
`tools/list`. `Model` is a table because each row carries eight independent admin
decisions; a discovered tool carries one (offered or not, in
`tool_overrides_json`, where absent means on), credentials and guidance are per
server, and the list is replaced wholesale on every refresh — a table would mean
reconciling rows against a cache of somebody else's document.
**An MCP tool has two names.** The server's own, which `tools/call` needs, and
the offered one in the schema — `slug_tool`, lowercased into
`[a-z0-9_-]{1,64}` because endpoints accept less than MCP does. Built-ins claim
their names first and can never be shadowed; a custom tool whose slug collides is
*refused at save*, an MCP tool is renamed silently, since the one that can adapt
should be the one that has to. The rename never leaves `mcp/registry.py`.
**A server's tool metadata is untrusted input that becomes instructions.**
Names, descriptions and schemas from `tools/list` are bounded and sanitised in
`mcp/protocol.clean_tool` before anything reaches a model. What a tool *returns*
is untrusted too, and is rendered as escaped preformatted text — never through
`services/markdown.py`, which is the one path allowed to emit HTML.
**A round's tool calls run together.** `generation._run_calls` gathers them under
a semaphore of four and keeps the results **indexed, not appended as they
finish**: each tool turn must line up with the assistant turn's `tool_calls` or
an endpoint matching on `tool_call_id` pairs the right id with the wrong content.
Safe because `run_tool` never raises and every runner opens its own session.
`generation.status` names what is running, because a remote tool taking seconds
with nothing streaming is exactly what a hang looks like.
**Tool-call arguments arrive in fragments.** `delta.tool_calls` carries an
`index`, a name that appears once, and an `arguments` string split across
chunks. `tools.ToolCallAccumulator` rejoins them keyed on `index` — not on
name, which breaks the moment a model calls one tool twice in a turn.
**Four stores, four different reasons.** `services/library/``documents`
(uploaded by a person, searched by the model), `notes` (written by the model,
searched), `memories` (short, and *injected whole* every turn), `skills` (index
injected, body fetched by tool). The shape of each follows from how it reaches
the model: a memory is capped short because it costs tokens on every request
forever, a note is not injected because a dozen would fill the window.
**Documents live in knowledge bases, and the base is what is shared.** A
`Document` always belongs to a `KnowledgeBase`; visibility comes from the base,
never the document, which is why `Document` is absent from
`sharing.RESOURCE_TYPES` and `documents.visible()` filters on
`base_id IN (visible bases)`. Per-document grants would mean answering "who can
see this?" by checking every file. `Document.base_id` is nullable only because
the column had to be added to a table that already had rows;
`documents.sweep_unfiled()` runs at startup and files anything predating bases
into its owner's default.
**A chat attached to bases is scoped to them.** `Chat.knowledge_bases` is
many-to-many; empty means "everything the owner can see", not "nothing".
`tools.context_for(db, user, chat)` carries the ids and `knowledge_search`
filters on them — and the harness names the bases, because otherwise the model
cannot tell "there is nothing about this" from "I am only allowed to see the
contracts folder".
**Sharing goes through one helper, and admins do not bypass it.**
`services/sharing.py:visible_to()` is the only definition of who can see a
library item, and every listing and tool uses it. `permissions.resolve` gives an
admin everything, deliberately — but that is about configuration, which an admin
can grant themselves anyway. Reading someone's private notes is not the same
act, so `sharing` has no admin branch. Sharing grants **reading only**.
**FTS5 tables are outside the model-driven schema sync.** They are not
SQLAlchemy models, so `sync_schema()` cannot diff them; `db/migrations.py:
ensure_fts()` writes them out with `IF NOT EXISTS` and creates the triggers that
keep an external-content index correct. It runs at every startup and converges,
like the column sync beside it. `tests/conftest.py` calls `sync_schema` rather
than `create_all` so tests run against the same schema.
**A failed search rolls back.** One broken FTS statement otherwise leaves the
session unusable and every later query in the request fails too, which looks
nothing like a search problem.
**Knowledge attachments are copies.** Attaching a library document to a message
duplicates its text and its file (`files.copy_document`). Referencing it would
mean a conversation changing when a document is edited or deleted later — the
same reason PDF text is extracted once at upload.
**The link fetcher is an SSRF hole unless guarded.** `services/fetch.py` refuses
loopback, private and link-local addresses **after resolution** — a hostname
pointing at 127.0.0.1 walks past any check that only reads the URL — and follows
redirects by hand so every hop is checked. An admin can open it deliberately.
The URL can come from a model, which can be talked into things by a page it just
read.
**The harness is an exception to the prompt-precedence rule, on purpose.**
"System prompts are precedence, not concatenation" governs the three *authored*
layers, and it stands: exactly one still wins, and `effective_system_prompt`
still decides which. `services/harness.py` is a different axis — it describes
the machinery rather than the behaviour, nobody authored it, and there is
nothing for it to disagree with. It is prepended to whichever authored prompt
won, in one system message (several endpoints reject a second one), and
`build_request` is where the two meet.
**The harness holds no text.** Every piece of it is a `Fragment` in
`services/prompts.py`, edited on `/admin/prompts`. `harness.py` decides which
fragments apply and what their variables resolve to; `prompts.py` owns the
wording, the storage and the substitution, and knows nothing about chats or
tools. Four rules hold the whole thing up:
- **Defaults live in code, overrides live in the database**, and text equal to
its default is never stored. That is what lets a later release improve a
default and have it reach an instance whose administrator once pressed Save.
- **An empty override means off**, which is why there is no separate enable
flag: clearing the box in the admin page *is* the switch. A fragment that was
not submitted at all keeps whatever it had — it may be missing from the page
because the thing contributing it is switched off.
- **A fragment carries its gate as data** (`families`, `requires`,
`when_tools`), never as a callable, because a database row can carry the same
three fields. `requires` is why there is no longer a hand-written pair of
memory-guidance variants: the sentence that refers to a section lives *inside*
that section, so it cannot outlive it.
- **`{{name}}`, and anything unrecognised passes through verbatim.** Names are
lowercase letters, digits and underscores, so `{"total": 1}` and `${PATH}` are
never candidates. Substitution is one pass and never recursive — `{{memories}}`
carries text a model wrote, and a memory reading `{{skills}}` must not expand.
A model with no tools now gets the core fragments too, the date above all.
"An empty harness is worse than none" was about tokens that say nothing, and a
model with no clock being asked about the present is not that. Clearing those
fragments restores the old silence exactly.
**Tool descriptions are not fragments.** They are schema, sent verbatim in the
`tools` array, and they state facts about what a runner does — an administrator
editing `notes_edit`'s "omit a field to leave it alone" would make the text a
lie with nothing to catch it. The page lists them read-only so nothing injected
is hidden. A *custom* tool's description will be editable, because it is a row.
**A model's tool flags default to on when `tools` is on.** Rows configured
before the per-tool split have no `tool_*` keys. Reading absent as off would
silently take web search away from every model already set up for it, so
`tools.enabled_tools` treats absent as inherited.
**Tool results are not replayed.** Like reasoning, `Message.tool_calls_json` is
stored and rendered but never fed back as context. The answer already contains
what the model made of the results; replaying stale results and the schema into
every later request wastes the window and reliably sends a small model into a
search loop. The sources stay visible in the transcript.
**Search results are untrusted.** Hard rule 6 covers them as much as model
output. `chat/_tool_activity.html` escapes everything and only renders `http`
and `https` URLs as links — a result carrying a `javascript:` URL must never
become an anchor.
**A message bubble is rendered from four places.** `pages.py`,
`chats.post_message`, `chats.regenerate` and `chats._follow`. Each needs
`audio_service.template_flags(db, user)` or the speaker button's conditions are
undefined; the template uses `| default(false)` so a missed one degrades to no
button rather than an exception. `_follow` also passes `just_finished`, which is
what read-aloud-automatically keys off — without it, reopening a chat would
start reading its last reply out loud.
**Dictation audio never touches disk.** `api/audio.py` reads it into memory,
capped, and streams it upstream. It is not an attachment: it has no owner, no
row, and nothing would ever sweep it.
**The service worker must skip `/api/`.** A reply is an endless event stream and
passing one through a worker turns it into one delivery at the end, or nothing.
`static/js/sw.js` bails out on `/api/`, `/auth/`, `/admin/` and any request
accepting `text/event-stream`. It is served from `GET /sw.js` rather than the
static mount because a worker's scope is the path it came from.
**XSS is now a root shell, not a leaked chat.** `api/terminal.py` is the one
WebSocket here, it is same-origin, the cookie rides along automatically, and
what it opens is an interactive shell. Every other route a script could reach
gives up a conversation; this one gives up the machine. Nothing about hard rule
6 changes — it was already absolute — but the *price* of getting it wrong did,
and so did the price of a stray `|safe`. The two locks are: the session cookie
is SameSite Lax, so a foreign page's handshake carries no cookie, and the
endpoint additionally **requires** an Origin header matching Host rather than
checking one when it happens to be present.
**A WebSocket dependency must be typed `HTTPConnection`.** `api/deps.py:
get_current_user` used to take a `Request`; FastAPI injects a `WebSocket` on a
websocket route, so the annotation fails at *connect* time rather than at
import. That is a failure which passes every test that does not open a socket
and breaks in a browser. `HTTPConnection` is the shared base and carries both
the cookies and `.state`.
**Terminal sessions are keyed on the chat, and outlive the socket.** A reload is
indistinguishable from a second tab, so anything finer needs an id in the
browser's storage — and then an abandoned tab leaks a PTY nothing in the UI can
find. One chat, one shell; two tabs share it and the smaller window decides the
size. Closing the panel calls `detach`, never `close`: a build running behind a
shut panel is the case the whole lifetime exists for. What ends one is the idle
timeout (nobody attached *and* nothing typed), deleting the chat, disabling,
moving or deleting the connection, forgetting its host key, or a restart.
**Unlike generations, nothing here ends by itself.** `generation.ensure` can
prune inside itself because a reply finishes and something calls in again. A
shell sits at a prompt forever, so `agent/terminal.py` runs a reaper task
instead. Copying the generation shape would mean nothing was ever swept.
**A slow viewer is dropped, not buffered.** Each viewer has a bounded queue; one
that fills is disconnected and reconnects with the scrollback, which costs it
nothing because the scrollback *is* the state. Blocking the pump instead would
stall every other viewer and buffer without bound — and `yes` is one word to
type. The reflex fix is an unbounded queue; it is the wrong one.
**Terminal traffic is bytes in both directions, and nothing decodes it.** A read
on the far side lands mid-character often enough to matter. xterm's decoder is
stateful across `write()` calls, so passing raw bytes through is correct by
construction, while decoding each frame server-side would corrupt every
boundary. Only `resize`, `ready`, `closed` and `error` are text, and they are
JSON.
**The modes do not govern the keyboard, and now there are three exceptions, not
one.** `agent/policy.py` exists because a model reads pages, files and command
output it did not write and can be talked into things. A person typing into the
terminal panel holds the credential already and could open the same shell with
an ssh client, so nothing they type is checked against the mode or the two
lists. The directory browser (`GET /api/agents/{id}/browse`) and the project
listing (`agent/index.py`) are the same argument again: both are read-only, both
are LLeMbas acting on somebody's instruction rather than a model choosing to,
and both would be pointless if they asked. But it does mean **Manual** mode's
"everything is shown to you before it happens" is now true of the *model* and
not of the interface, and that is worth saying out loud rather than discovering.
There is a test named after the first one, because it reads like a bug next to
`policy.py` and "fixing" it would make the panel useless in the mode people
spend the most time in.
**A control wired to a method its route does not serve fails silently.** The
agent-mode select posted with `hx-post` against a route that only answers
`PATCH`, so every change returned 405 and the mode never moved — for the whole
life of the feature. htmx surfaces nothing on a failed request, so the select
stayed where it was put and the server ignored it, which looks exactly like
working. `tests/test_agent_mode.py` asserts the method is *refused* as well as
that the right one works, because only the second half would have passed
throughout. When adding a control that writes, check the verb against the route,
and assert on the row rather than on the response.
**Shell integration is best-effort, and the fallback is the point.**
`agent/shell_marks.py` gives bash and zsh hooks that emit OSC 133 around the
prompt, the command and its result, so the panel can say what "the last command
and its output" means. Three things about it:
- **It is written by the PTY command string itself**, with `printf`. sshd runs
that string through `$SHELL -c`, so it can `case` on the shell's own name and
needs no probe, no second channel and no writable `$HOME`. Environment
variables do not work — every distribution ships `AcceptEnv LANG LC_*`, so
anything else is dropped silently — and feeding `source …` in as keystrokes
races a slow `.zshrc`, echoes, and lands in shell history.
- **Nothing needs hiding.** The setup runs before the shell exists and never
writes to the PTY's *input* side, so there is nothing to echo and no fan-out
gate. That is why this mechanism was chosen over the one that looks obvious.
- **The exit status is captured in the `DEBUG` trap, not in `PROMPT_COMMAND`.**
DEBUG fires before every simple command *including each one inside
`PROMPT_COMMAND`*, so `$?` read from there is whatever ran a moment ago. This
was wrong in the first version and every command reported success. zsh has the
mirror-image trap: `$ZDOTDIR` is already ours by the time `.zshenv` runs, so
the user's own must be passed on the exec line or the shims source themselves
and none of somebody's configuration loads.
Any shell that is not bash or zsh gets exactly the command that ran before, and
therefore no markers — at which point Copy and Send fall back to scraping the
screen and say so, and the automatic toggle is **disabled rather than degraded**.
Forty arbitrary lines attached to every message is worse than nothing attached.
**The nginx vhost must pass upgrades through.** `deploy/nginx-vhost.conf` used
to set `Connection ""`, which is right for SSE and fails every WebSocket
handshake — and a failed handshake tells the browser nothing: no status, no
reason. It now uses `map $http_upgrade`, which yields the empty string when
nothing asked to upgrade, so one `location` serves both. `update.sh` has a drift
check for exactly this.
**`data-toggle` syncs every toggle, not the one that was clicked.** A panel can
be opened by the topbar button and closed by its own Close, and now also closed
by nothing at all: `data-toggle-group="side"` makes the terminal and the
inspector mutually exclusive, because at 1280px both plus the sidebar leave the
conversation about seventy pixels wide. `app.js:setPanel` applies the state and
then brings every `[data-toggle]` pointing at that panel in line, and fires
`lembas:toggle` — which is how `terminal.js` learns it is visible and may
measure itself. xterm's `fit()` reads `offsetWidth`, which is 0 inside a
`[hidden]` ancestor, so fitting early is a silent no-op that leaves an
80-column terminal in a 34rem panel.
**xterm holds colours as values, so the theme has to be pushed at it.**
`applyTheme` dispatches `lembas:theme`; without it, switching to `shire` leaves
a black rectangle in a light interface. Same reason a `ResizeObserver` is on the
panel: a window `resize` never fires when the sidebar is toggled beside it.
## Changing the schema
There is no Alembic, but there *is* `db/migrations.py`. It compares the declared
models against the live database and issues `ALTER TABLE ... ADD COLUMN` for
anything missing, so adding a column to a model is free: restart and it appears,
with existing rows backfilled from a type-derived default.
**A nullable column is added with no default**, so existing rows get NULL --
the value the model treats as absent. Only a NOT NULL column gets one, because
SQLite refuses to add one without. An earlier version defaulted every column by
type, which meant an added foreign key arrived as `""` on old rows and every
"is this set?" check downstream was wrong about them.
It cannot rename, drop or retype a column, or add a UNIQUE/PRIMARY KEY to an
existing table — SQLite mostly cannot do those with ALTER TABLE either. Those
need the create-copy-swap dance by hand; record them in `MANUAL_STEPS` so a
failure has somewhere to point.
Because the runner exists, forward-looking columns are cheap now. `Message.parent_id`
and `content_parts_json` (branching, multimodal) predate it and are still unread.
## Artwork
Do not hand-edit files in `assets/` — they are generated. Change
`scripts/build_artwork.py` and re-run it. It also copies the few files the app
serves into `web/static/img/`.
The leaf geometry is defined once (`LEAF_BLADE`, `LEAF_MIDRIB`, …) and reused by
the icon, favicon, lockup and banner. The 64×64 mark must stay legible at 16px:
the favicon variant drops the score lines, rim and veins because they turn to
mud at that size. The icon sprite is a **template partial**
(`templates/partials/icons.html`), not an asset, because same-document
`<use href="#id">` is universally supported and the cross-document form is not.
The `mark()` macro in `_macros.html` duplicates the mark geometry so it can be
inlined and themed. If the mark changes, update both.
## Deployment
`deploy/` holds a systemd unit template, an nginx vhost template, and
install/update scripts. Both templates are parameterised (`__PREFIX__`,
`__SITE_HOST__`, …) and substituted at install time, so nothing host-specific is
committed here. See `deploy/README.md`.
This repository is **public**. Keep deployment-specific hostnames, ports and
internal infrastructure detail out of it — those belong in whatever private
notes describe the machine.
## Not built yet
Image generation, and a nav entry marks where it goes. The tool loop in
`services/generation.py` is what a new capability plugs into — a tool is a
`ToolDef` reaching `tools.resolve_tools()` plus a permission and a capability
flag, not a new code path. Its guidance is the same shape: a
`prompts.register_source` yielding one `Fragment` per row puts it in the harness,
on the admin page and in the preview without touching the assembler, the save
handler or a template. Custom HTTP tools, MCP servers and agent chats are the
three worked examples.
**Nothing executes on this machine, and that is the design.** Agent chats run
their commands on a host reached over SSH. A local sandbox was designed in
detail — bubblewrap, a masked data directory, a curated bind list — and dropped,
because every hard problem in it came from running on the machine that holds the
database and the encryption key: the service user cannot traverse `/home`,
granting it needs ACLs, `RLIMIT_NPROC` is counted per *uid* so a fork bomb
starves the server too, `--size` applies only to tmpfs so there is no disk quota,
and a bind list is a standing invitation to widen until the sandbox is
decoration. Over SSH, isolation is somebody's considered choice of host, using
tools far better at it than anything that could be built here — and it is the
only version that is honestly multi-user.
Two consequences worth stating plainly. **The security of an agent chat is the
security of the host behind its profile**, and nothing here can tell a throwaway
container from a production server. And there is no equivalent of the
`network: False` switch the local sandbox would have had, because the network
belongs to the far side — so an instruction injected through something the model
read can, in principle, be carried out from there.
**Local MCP is absent for the same reason.** Only remote servers over streamable
HTTP. Spawning `npx` would be a subprocess on this machine, which is the thing
that is deliberately not done.
**Unknown is not zero.** `Model.context_length` of 0 means nobody has said how
big the window is, which is different from "small". The context percentage is
omitted rather than computed, and automatic compaction never fires. Token counts
fall back to `services/tokens.py` -- four characters to a token -- and anything
derived from an estimate is shown with a `~`. Compaction *does* act on an
estimate, because a premature compaction costs one turn of answer quality rather
than data: the messages are kept.
**Compaction hides turns, it does not delete them.** `Chat.compact_summary` plus
`compacted_through_id` say how far it reached; the messages stay in the
transcript behind a `<details>` divider and simply stop being part of the
request. The summary is carried by a **user turn and an assistant turn**, not
one: a leading `assistant` breaks templates that require the first non-system
message to be `user`, and a lone leading `user` produces `user, user` whenever
the kept history starts on a user turn -- which it always does, because the
cutoff lands on a finished reply. `compacted_through_id` is a plain id, not a
foreign key, because `migrations.py` compiles only the column type and a
`REFERENCES` clause would exist on a fresh database and not on an upgraded one;
`compaction.cutoff_message` validates it on every read instead.
**Compare message timestamps through `compaction.moment()`.** SQLite does not
store the offset, so a row loaded from disk is naive while one still in the
session's identity map keeps its tzinfo. Comparing the two raises.
No OCR: a scanned PDF is stored with an explanatory `extraction_error` rather
than silently contributing nothing.