The instructions LLeMbas puts in front of a model were hard-coded: six
strings in a GUIDANCE dict, two headings, and the title request inline in
chat.py. An operator could not see what was being sent, let alone change
it, and there was nowhere for a custom tool to contribute its own guidance
when custom tools land.
services/prompts.py now holds each piece as a Fragment, and /admin/prompts
edits them with a preview of the whole assembled system message including
unsaved edits. harness.py keeps only the decisions -- which fragments apply
to this request, and what their variables resolve to.
The design turns on one choice: a fragment carries its gate as data
(families, requires, when_tools) rather than as a callable, because a
database row can carry the same three fields. Custom tools will therefore
register a fragment source and change nothing else -- there is a test that
says exactly that, and it is the reason the rest of the shape is what it is.
Consequences worth knowing:
- Defaults live in code, overrides in the database, and text equal to its
default is never stored. Otherwise pressing Save once would freeze
today's wording forever and no later release could improve it.
- An empty override means off. A fragment that was not submitted at all
keeps what it had, because it may be missing from the page only because
whatever contributes it is currently switched off.
- requires= replaced the hand-written pair of memory guidance variants.
The sentence that refers to a section now lives inside that section, so
it cannot outlive it. That was the general problem the pair was a
special case of.
- {{name}}, with anything unrecognised passing through verbatim. The name
grammar is the guard: {"total": 1} and ${PATH} are not candidates.
Substitution is one pass and never recursive, because {{memories}}
carries text a model wrote.
The wording is also overhauled, and a model now gets the core fragments
even with no tools -- the date above all. "An empty harness is worse than
none" was about tokens that say nothing; a model with no clock being asked
about the present is not that. Clearing those boxes restores the old
silence exactly. New: today's date, who it is talking to, the three-round
tool budget, that tool results are not replayed, that anything a tool
returns is data rather than instruction, and what the <document> wrapper
around an attachment is. Extended: memory_forget, notes_edit/delete,
skill_create/edit, and reading a knowledge document in full rather than
answering from an extract.
Tool descriptions stay in code and are listed read-only. They are schema
and they state facts about what a runner does; an edit would make the text
a lie with nothing to catch it.
No schema change -- one JSON row in the settings table.
488 tests. Version 0.2.0, which also invalidates the service worker cache
so the green artwork appears without a hard reload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
26 KiB
CLAUDE.md
Working notes for LLeMbas. Read this before changing anything.
What it is
A self-hosted web UI for OpenAI-compatible LLM endpoints, written in Python and themed after Middle-earth. Server-rendered FastAPI + Jinja + htmx; SQLite; no JavaScript build step.
Commands
. .venv/bin/activate
pip install -e ".[dev,search]" # `search` adds ddgs for DuckDuckGo
lembas serve # http://127.0.0.1:8080
lembas info # paths + counts, useful when confused
lembas secret-key # generate LEMBAS_SECRET_KEY
lembas create-admin # create or promote an admin
pytest # 488 tests, ~29s
# PLAN.md tracks what is and is not built
ruff check . # lint (line length 100)
python scripts/build_artwork.py # regenerate artwork (SVG + PWA icons;
# needs fonttools and cairosvg)
python scripts/fetch_vendor.py # verify vendored JS against the lockfile
Hard rules
These are the constraints the project is built around. Breaking one is a redesign, not a tweak.
- No Node, no npm, no build step. Browser libraries are downloaded once by
scripts/fetch_vendor.py, hash-pinned inscripts/vendor.lock.json, and committed underweb/static/vendor/. - Nothing loads from a CDN at runtime. A self-hosted tool must work offline and must not report page views to a third party.
- No hard-coded values outside
tokens.css. Every colour, space, radius and control height resolves through a CSS variable.--control-his why buttons, inputs and selects line up: they all take their height from it, so a mixed row is flush by construction rather than by nudging. - Additive-only schema changes. SQLite only, no Alembic.
init_db()runsdb/migrations.py:sync_schema(), which creates missing tables and adds missing columns by diffing the models against the database. Renames, drops and retypes are still manual. See "Changing the schema" below. - Secrets never reach the browser. API keys are Fernet-encrypted at rest and only ever rendered masked.
- Model output is untrusted. Everything from an endpoint goes through
services/markdown.py(markdown-it → nh3) orescape_text(). Never|safeon anything that has not.
The flavour rule
Middle-earth lives in the artwork, theme names, empty states, loading lines and error pages. It does not live in the functional UI.
Chats are called Chats, not Tales. Folders are Folders, not Chapters.
Buttons say what they do. Someone who has never read the books must be able to
use this without a glossary. The two themes are named moria and shire, and
the 404 says "Not all those who wander are lost. This page, however, is." —
that is the right amount.
Layout
src/lembas/
main.py app factory, lifespan, error handlers
config.py pydantic-settings, all LEMBAS_* variables
cli.py typer entry points
api/
deps.py Db / CurrentUser / RequiredUser / AdminUser
auth.py register, login, logout
pages.py full-page routes (chat shell, settings)
chats.py messaging + the SSE stream
folders.py folder CRUD
admin.py connections + instance settings
admin_models.py model ordering, defaults, images, access
admin_users.py users, groups, permissions
admin_audio.py speech-to-text and text-to-speech endpoints
admin_search.py web search provider and credentials
admin_prompts.py the prompt fragment editor and its preview
audio.py transcribe, speak, voice discovery
library.py knowledge, notes, skills pages; memory CRUD
files.py upload, serve, remove attachments
preferences.py per-user theme, default model, password, audio
db/
base.py Base, UUID/Timestamp mixins
session.py engine, SQLite pragmas, init_db, session_scope
migrations.py additive schema sync (tables + columns)
models/ user, chat, connection, setting
security/ passwords (argon2), sessions, permissions
services/
llm/openai_client.py httpx streaming + model discovery
search/ ddgs, SearXNG and Firecrawl behind one shape
library/ documents, notes, memories, skills, FTS
audio.py OpenAI-shaped /v1/audio/* client
fetch.py URL retrieval, HTML to text, the SSRF guard
sharing.py one visibility rule for every library store
prompts.py every injected prompt fragment, and {{variables}}
harness.py the operational prompt built from what a model has
tools.py tool registry, schemas, streamed-call reassembly
chat.py request building, endpoint resolution, titles
markdown.py markdown-it + pygments + nh3
crypto.py Fernet encrypt/decrypt/mask
files.py attachment validation, images, PDF/text extraction
reasoning.py splits thinking from the answer
settings_store.py runtime instance settings
uploads.py validated image storage
sse.py event framing
web/
templating.py render() -- always use this, not TemplateResponse
templates/ Jinja
static/ css, js, vendor, img, sw.js
assets/ SVG masters and PWA icons (generated)
deploy/ systemd unit, nginx vhost, install/update scripts
Things that will bite you
render(), not TemplateResponse. web/templating.py:render() injects
user, theme, version and allow_signup. Templates assume they exist. If
you must call templates.TemplateResponse directly (the SSE path does, because
there is no Request), pass user explicitly — chat/_message.html renders
both roles and the user branch dereferences it.
The message template is the state machine. chat/_message.html renders an
incomplete assistant message as a streaming shell carrying sse-connect, and a
complete one as finished output. That is the only thing that starts a
generation. A consequence worth knowing: loading a page whose last reply is
unfinished restarts it, which is how a dropped connection recovers.
SSE framing. services/sse.py:event() splits payloads on newlines into
several data: lines. A raw newline in a single data: line truncates the
event — the failure shows up the first time a model emits a code block.
Streaming opens its own database session. api/chats.py:_generate() uses
session_scope(), not the request's session, because streaming outlives the
request handler.
Escaping is chunk-safe on purpose. escape_text() is html.escape, which
works character by character, so escaping stream chunks separately equals
escaping the whole string. nh3.clean_text would also be safe but escapes
spaces and slashes, tripling the size of every streamed token.
The fence renderer is replaced, not configured. markdown-it's highlight
option re-wraps output in <pre><code> unless the string starts with <pre,
which would nest a second <pre> inside our wrapper. markdown.py overrides
renderer.rules["fence"] instead. There is a regression test for this.
SVG <style> is document-scoped. Two text runs in one SVG sharing a class
name means the later rule recolours both. build_artwork.py takes class names
as parameters for exactly this reason.
Gradient ids are document-global. The mark() macro takes a uid because
two marks on one page with identical ids make the second silently reuse the
first one's gradients.
Two kinds of settings. lembas.config is deployment configuration read
from the environment at startup. services/settings_store.py is instance
settings an admin edits at runtime, stored in the settings table. Environment
variables seed the latter as an initial value only — once stored, the database
wins, or a toggle in the UI would silently revert on the next restart.
Permissions are a union, and admins bypass them. security/permissions.py
resolves a baseline (instance setting) widened by each group. A group grants;
it never denies — otherwise "why can this user not do X" needs a simulation of
every group to answer. Model access is separate: models_visible_to().
FastAPI cannot tell an empty form field from an absent one. With
x: str | None = Form(None), a submitted x= arrives as None, so "clear this
field" is indistinguishable from "leave it alone". api/chats.py:update_chat
reads await request.form() and checks key presence instead. Anything with a
clearable field must do the same.
Mapped[list] without an element type is not a collection. SQLAlchemy
treats a bare Mapped[list] as a scalar and hands back None instead of [].
Always write Mapped[list[Group]], with a TYPE_CHECKING import if the class
lives in another module.
Reasoning arrives two ways. A reasoning_content delta field (llama.cpp,
llama-swap, vLLM) or <think> tags inline in content (Ollama and friends).
services/reasoning.py handles the second with a streaming splitter, because
the tags arrive split across chunks. Reasoning is stored in Message.reasoning
and is deliberately not replayed as context on the next turn.
Attachments are typed by their bytes, not their name. services/files.py
sniffs magic numbers; a .png full of text is stored as text. Images are
downscaled and re-encoded (a phone photo is megabytes of base64), PDFs have
their text extracted once at upload — re-extracting per request would let a
reply change because a parser was upgraded.
Images only go to models marked vision. Sending content parts to an
endpoint without multimodal support is not graceful degradation; most reject
the whole request. build_request() checks the capability and falls back to a
plain string. A plain text turn must stay a plain string for the same reason.
Attachments are served, never linked. Images reach the model as base64 data
URIs: a local endpoint has no route back to LLeMbas and a hosted one has no
credentials. Non-images are served Content-Disposition: attachment with
nosniff, so an uploaded .html cannot execute in this origin.
Uploads are unbound until the message is sent. Attachment.message_id is
null in the composer; files.claim() binds them, and only unclaimed rows owned
by that user, so a forged id cannot pull in someone else's file. Abandoned ones
are swept at startup.
Generation is a background task; the SSE endpoint only follows it.
services/generation.py owns the work and the registry; api/chats.py:_follow
watches a Generation and streams what it sees. Closing the connection does
NOT stop the reply -- that was the old behaviour and it cut answers off when
the reader navigated away. Any route that creates an assistant placeholder must
also call generation.ensure().
Stream frames carry whole blocks, not deltas. Both render and reasoning
send the complete text each time. That is what makes reattaching mid-reply
work: a follower arriving late has no earlier fragments to append to. It also
means Markdown is re-rendered whole, which is required anyway -- a list or code
fence is only correct once its context exists.
Stopping sets a flag the producer checks. generation.request_stop();
whatever arrived is kept and the message is marked stopped, which is distinct
from error. In-process, so single-worker only.
Unread is polled, not pushed. A browser on another chat has no connection
to the one that finished. /api/chats/unread returns out-of-band dot spans and
an HX-Trigger for the toast; unread_notified stops the same arrival being
announced every tick. Re-rendering the whole sidebar instead would reset the
folder open/closed state every 10 seconds.
Editing rewinds, it does not branch. POST .../messages/{id}/edit rewrites
a user turn and deletes everything after it. Branching would need a UI for
choosing between versions; "go back and try again from here" is what was asked
for and what other clients do. Message.parent_id still exists unused.
Dialogs and toasts are ours, not the browser's. static/js/ui.js provides
lembas.notify/confirm/prompt, and intercepts htmx's htmx:confirm so every
existing hx-confirm gets the themed dialog with no change at the call site.
Plain forms opt in with data-confirm, lone submit buttons with
data-confirm-button. Never add a window.confirm back.
The model picker is hand-built. A <select> renders only text in an
<option> -- no avatar, no description, no badges. chat/_model_picker.html
plus the picker block in ui.js; the value lives in a hidden input so it still
behaves as a form field.
Chats are created lazily. There is no endpoint that makes an empty chat.
"New chat" is a link to /chat, which renders a composer with no row behind
it; POST /api/chats/start writes the chat together with its first message.
That is why an opened-and-abandoned chat never appears in the sidebar. Tests
that just need a chat use the make_chat fixture rather than the HTTP flow.
Admin lists are list-plus-detail, never a form per row. /admin/models
renders compact rows with search, filter tabs and pagination; the full form
lives at /admin/models/{id}/edit. A connection can advertise a hundred models,
and a page that renders a form for each is unusable. Any future admin list
(tools, agents) should follow the same shape.
Route order matters for static path segments. FastAPI matches in
registration order, so /admin/models/bulk must be registered before
/admin/models/{model_id} or "bulk" is parsed as a model id and 404s. This has
already been a bug once.
Pinning is not ordering. The model picker is always in the administrator's
position order. Pinned models get shortcuts in the chat sidebar and nothing
else -- a picker whose order differs from the admin screen is just confusing.
System prompts are precedence, not concatenation. chat > model > instance,
most specific wins outright (services/chat.py:effective_system_prompt).
Stacking them reads well in a settings screen and badly in practice: two layers
that disagree give the model contradictory instructions and nobody can tell
which is losing.
JSON columns need reassignment. user.settings_json["theme"] = x on a
plain dict is not detected. The columns use MutableDict (db/types.py), but
the safe habit is obj.field = {**obj.field, "k": v}.
[hidden] needs !important. The browser's rule is [hidden] { display: none }, which any class setting display outranks — and .btn is
display: inline-flex. That is not theoretical: it is why the old Stop button,
created and then hidden = true, sat permanently beside Send. app.css forces
the attribute to win. Anything toggled with hidden depends on that line.
Send and Stop are one button. chat/_composer.html renders both icons and
ui.js flips data-composer-action plus type (submit ↔ button) when a
message in the thread is still streaming. Do not add a second button back.
The tool loop is inside one generation. services/generation.py:_run() runs
up to tools_service.MAX_ROUNDS request rounds for a single reply: stream,
accumulate tool calls, run them, append the results, ask again. Generation
accumulates content across all of them, so text emitted before a tool call
survives. Tools are only offered when search is enabled, the user has
tools.web_search, and the model is flagged tools — sending a tools
array to an endpoint without support fails the whole request, exactly as images
do without vision.
Tool-call arguments arrive in fragments. delta.tool_calls carries an
index, a name that appears once, and an arguments string split across
chunks. tools.ToolCallAccumulator rejoins them keyed on index — not on
name, which breaks the moment a model calls one tool twice in a turn.
Four stores, four different reasons. services/library/ — documents
(uploaded by a person, searched by the model), notes (written by the model,
searched), memories (short, and injected whole every turn), skills (index
injected, body fetched by tool). The shape of each follows from how it reaches
the model: a memory is capped short because it costs tokens on every request
forever, a note is not injected because a dozen would fill the window.
Documents live in knowledge bases, and the base is what is shared. A
Document always belongs to a KnowledgeBase; visibility comes from the base,
never the document, which is why Document is absent from
sharing.RESOURCE_TYPES and documents.visible() filters on
base_id IN (visible bases). Per-document grants would mean answering "who can
see this?" by checking every file. Document.base_id is nullable only because
the column had to be added to a table that already had rows;
documents.sweep_unfiled() runs at startup and files anything predating bases
into its owner's default.
A chat attached to bases is scoped to them. Chat.knowledge_bases is
many-to-many; empty means "everything the owner can see", not "nothing".
tools.context_for(db, user, chat) carries the ids and knowledge_search
filters on them — and the harness names the bases, because otherwise the model
cannot tell "there is nothing about this" from "I am only allowed to see the
contracts folder".
Sharing goes through one helper, and admins do not bypass it.
services/sharing.py:visible_to() is the only definition of who can see a
library item, and every listing and tool uses it. permissions.resolve gives an
admin everything, deliberately — but that is about configuration, which an admin
can grant themselves anyway. Reading someone's private notes is not the same
act, so sharing has no admin branch. Sharing grants reading only.
FTS5 tables are outside the model-driven schema sync. They are not
SQLAlchemy models, so sync_schema() cannot diff them; db/migrations.py: ensure_fts() writes them out with IF NOT EXISTS and creates the triggers that
keep an external-content index correct. It runs at every startup and converges,
like the column sync beside it. tests/conftest.py calls sync_schema rather
than create_all so tests run against the same schema.
A failed search rolls back. One broken FTS statement otherwise leaves the session unusable and every later query in the request fails too, which looks nothing like a search problem.
Knowledge attachments are copies. Attaching a library document to a message
duplicates its text and its file (files.copy_document). Referencing it would
mean a conversation changing when a document is edited or deleted later — the
same reason PDF text is extracted once at upload.
The link fetcher is an SSRF hole unless guarded. services/fetch.py refuses
loopback, private and link-local addresses after resolution — a hostname
pointing at 127.0.0.1 walks past any check that only reads the URL — and follows
redirects by hand so every hop is checked. An admin can open it deliberately.
The URL can come from a model, which can be talked into things by a page it just
read.
The harness is an exception to the prompt-precedence rule, on purpose.
"System prompts are precedence, not concatenation" governs the three authored
layers, and it stands: exactly one still wins, and effective_system_prompt
still decides which. services/harness.py is a different axis — it describes
the machinery rather than the behaviour, nobody authored it, and there is
nothing for it to disagree with. It is prepended to whichever authored prompt
won, in one system message (several endpoints reject a second one), and
build_request is where the two meet.
The harness holds no text. Every piece of it is a Fragment in
services/prompts.py, edited on /admin/prompts. harness.py decides which
fragments apply and what their variables resolve to; prompts.py owns the
wording, the storage and the substitution, and knows nothing about chats or
tools. Four rules hold the whole thing up:
- Defaults live in code, overrides live in the database, and text equal to its default is never stored. That is what lets a later release improve a default and have it reach an instance whose administrator once pressed Save.
- An empty override means off, which is why there is no separate enable flag: clearing the box in the admin page is the switch. A fragment that was not submitted at all keeps whatever it had — it may be missing from the page because the thing contributing it is switched off.
- A fragment carries its gate as data (
families,requires,when_tools), never as a callable, because a database row can carry the same three fields.requiresis why there is no longer a hand-written pair of memory-guidance variants: the sentence that refers to a section lives inside that section, so it cannot outlive it. {{name}}, and anything unrecognised passes through verbatim. Names are lowercase letters, digits and underscores, so{"total": 1}and${PATH}are never candidates. Substitution is one pass and never recursive —{{memories}}carries text a model wrote, and a memory reading{{skills}}must not expand.
A model with no tools now gets the core fragments too, the date above all. "An empty harness is worse than none" was about tokens that say nothing, and a model with no clock being asked about the present is not that. Clearing those fragments restores the old silence exactly.
Tool descriptions are not fragments. They are schema, sent verbatim in the
tools array, and they state facts about what a runner does — an administrator
editing notes_edit's "omit a field to leave it alone" would make the text a
lie with nothing to catch it. The page lists them read-only so nothing injected
is hidden. A custom tool's description will be editable, because it is a row.
A model's tool flags default to on when tools is on. Rows configured
before the per-tool split have no tool_* keys. Reading absent as off would
silently take web search away from every model already set up for it, so
tools.enabled_tools treats absent as inherited.
Tool results are not replayed. Like reasoning, Message.tool_calls_json is
stored and rendered but never fed back as context. The answer already contains
what the model made of the results; replaying stale results and the schema into
every later request wastes the window and reliably sends a small model into a
search loop. The sources stay visible in the transcript.
Search results are untrusted. Hard rule 6 covers them as much as model
output. chat/_tool_activity.html escapes everything and only renders http
and https URLs as links — a result carrying a javascript: URL must never
become an anchor.
A message bubble is rendered from four places. pages.py,
chats.post_message, chats.regenerate and chats._follow. Each needs
audio_service.template_flags(db, user) or the speaker button's conditions are
undefined; the template uses | default(false) so a missed one degrades to no
button rather than an exception. _follow also passes just_finished, which is
what read-aloud-automatically keys off — without it, reopening a chat would
start reading its last reply out loud.
Dictation audio never touches disk. api/audio.py reads it into memory,
capped, and streams it upstream. It is not an attachment: it has no owner, no
row, and nothing would ever sweep it.
The service worker must skip /api/. A reply is an endless event stream and
passing one through a worker turns it into one delivery at the end, or nothing.
static/js/sw.js bails out on /api/, /auth/, /admin/ and any request
accepting text/event-stream. It is served from GET /sw.js rather than the
static mount because a worker's scope is the path it came from.
Changing the schema
There is no Alembic, but there is db/migrations.py. It compares the declared
models against the live database and issues ALTER TABLE ... ADD COLUMN for
anything missing, so adding a column to a model is free: restart and it appears,
with existing rows backfilled from a type-derived default.
A nullable column is added with no default, so existing rows get NULL --
the value the model treats as absent. Only a NOT NULL column gets one, because
SQLite refuses to add one without. An earlier version defaulted every column by
type, which meant an added foreign key arrived as "" on old rows and every
"is this set?" check downstream was wrong about them.
It cannot rename, drop or retype a column, or add a UNIQUE/PRIMARY KEY to an
existing table — SQLite mostly cannot do those with ALTER TABLE either. Those
need the create-copy-swap dance by hand; record them in MANUAL_STEPS so a
failure has somewhere to point.
Because the runner exists, forward-looking columns are cheap now. Message.parent_id
and content_parts_json (branching, multimodal) predate it and are still unread.
Artwork
Do not hand-edit files in assets/ — they are generated. Change
scripts/build_artwork.py and re-run it. It also copies the few files the app
serves into web/static/img/.
The leaf geometry is defined once (LEAF_BLADE, LEAF_MIDRIB, …) and reused by
the icon, favicon, lockup and banner. The 64×64 mark must stay legible at 16px:
the favicon variant drops the score lines, rim and veins because they turn to
mud at that size. The icon sprite is a template partial
(templates/partials/icons.html), not an asset, because same-document
<use href="#id"> is universally supported and the cross-document form is not.
The mark() macro in _macros.html duplicates the mark geometry so it can be
inlined and themed. If the mark changes, update both.
Deployment
deploy/ holds a systemd unit template, an nginx vhost template, and
install/update scripts. Both templates are parameterised (__PREFIX__,
__SITE_HOST__, …) and substituted at install time, so nothing host-specific is
committed here. See deploy/README.md.
This repository is public. Keep deployment-specific hostnames, ports and internal infrastructure detail out of it — those belong in whatever private notes describe the machine.
Not built yet
Custom tools and MCP, agentic execution (local subprocess and SSH connection
profiles), image generation. Nav entries mark where each one goes. The tool
loop in services/generation.py is what they plug into — a new tool is a
ToolDef in services/tools.py:REGISTRY plus a permission and a capability
flag, not a new code path. Its guidance is the same shape: a
prompts.register_source yielding one Fragment per tool row puts it in the
harness, on the admin page and in the preview without touching the assembler,
the save handler or a template.
No OCR: a scanned PDF is stored with an explanatory extraction_error rather
than silently contributing nothing.