# LLeMbas — plan and status Where the project is, what is deliberately not built yet, and the decisions that would be expensive to revisit. Kept current as work lands; the detail of *how* things work lives in [`CLAUDE.md`](CLAUDE.md). **Status:** usable daily. Streaming chat, attachments, reasoning, tool calling with web search, a knowledge library, notes, memory and skills, speech in and out, users and groups, model administration, installable as an app. 437 tests, `ruff` clean. --- ## The shape of it A self-hosted web UI for OpenAI-compatible endpoints, written in Python, themed after Middle-earth. | | | |---|---| | Stack | FastAPI + Jinja + htmx + a little Alpine | | Build step | none — no Node, no npm, no CDN at runtime | | Database | SQLite, schema synchronised additively at startup | | Deployment | systemd unit + nginx vhost, one worker | These are load-bearing. Dropping the no-build rule or moving off SQLite would be a different project, not a refactor. --- ## Done ### Chat - [x] Streaming replies over server-sent events - [x] **Markdown renders progressively** — re-rendered whole every 100ms rather than appending tokens, because a list or code fence is only correct once its context exists - [x] Syntax highlighting (Pygments), sanitised with nh3 - [x] **Generation runs in the background** — a task, not the request. Navigate away, open another chat, close the tab: the reply keeps being written and reattaching replays the whole state - [x] **Stop** — the send button becomes Stop while writing; what arrived is kept - [x] **Rewind** — edit one of your own turns and the conversation runs on from there. Truncates rather than branching - [x] Copy, regenerate, automatic chat titles - [x] Chats created on first message, so an abandoned composer leaves nothing - [x] **Unread indicator** — a green dot and a toast when a reply lands while you were elsewhere - [x] Folders, arbitrarily nested; deleting one keeps the chats inside it - [x] Per-reply metrics — tokens, context used as a percentage, tokens/second, live while streaming and kept afterwards. Estimated with a `~` when the endpoint reports no usage - [x] Compaction — a button, and automatically at a configurable percentage of the model's context. Summarised turns are kept and collapsed, not deleted - [x] Temporary chats — never listed, swept after a day, with a Keep button - [x] An admin-only request inspector beside the thread ### Tools - [x] **Tool calling** — one reply is a bounded loop of requests, not one request. Text produced before a call is kept - [x] **Web search** as the first tool: DuckDuckGo (no setup), SearXNG or Firecrawl, chosen in the admin area - [x] Only offered to models flagged `tools`, because an endpoint without support rejects the whole request rather than ignoring the array - [x] Sources stay in the transcript; results are **not** replayed as context on the next turn, for the same reasons reasoning is not - [x] A round's calls run together, and the reply says which tool is running — a remote tool taking seconds with nothing streaming looks like a hang - [x] **Custom HTTP tools** — an administrator describes one call: a JSON Schema, a URL template, headers, an encrypted secret and how to read the answer. Arguments may fill a hole but never move the target: the scheme and host are literal, values are escaped for where they land, and the origin is pinned afterwards - [x] **MCP servers** over streamable HTTP — a hand-written client, so that `check_url` runs on every hop rather than being bypassed by somebody else's transport. Tools are discovered and cached by a button, namespaced per server, and a server's own descriptions are bounded before they reach a model as instructions - [x] Both gated like the built-ins — a model capability, a permission — and restrictable to groups, with guidance of their own on `/admin/prompts` - [x] Local MCP over stdio is deliberately absent: spawning a subprocess is the agentic-execution feature and wants a confirmation model first ### The library - [x] **Knowledge bases** — documents, images and saved web pages, grouped into named collections and ingested through the same pipeline as chat attachments, searched with SQLite FTS5 - [x] A chat can be pointed at particular bases, so "answer from the contracts folder" is a different question from "answer from everything I have" - [x] **Notes** — longer things the model writes down and searches later; editable by hand, because they are yours - [x] **Memory** — short facts, injected on every turn to a budget rather than searched, and managed in your settings - [x] **Skills** — saved procedures. Only the name and description are injected; the body is fetched when the model decides it applies - [x] A model may write and revise its own notes, memories and skills. Every skill revision is kept, attributed and revertible — the safety story is a record and a way back, not a gate - [x] **Sharing** — a knowledge base, a note or a skill can be shared with a group or with named people, read-only. One visibility rule, and administrators do not bypass it. Documents are shared through their base - [x] **The harness** — an operational prompt assembled from what a model actually has, so the tools get used rather than ignored - [x] Attach menu: file, image, a web page fetched on the spot, or a document from the library ### Audio - [x] **Dictation** — record in the composer, transcribed by any OpenAI-shaped `/v1/audio/transcriptions` endpoint. The recording never touches disk - [x] **Read aloud** — any `/v1/audio/speech` endpoint, with the voice list discovered from the server where it offers one - [x] Instance defaults in Admin, per-reader overrides in Settings — voice, speed, dictation language, and whether replies play automatically ### Models and reasoning - [x] OpenAI-compatible connections with encrypted keys and model discovery - [x] **Reasoning display** — `reasoning_content` and inline `` tags, collapsed by default, labelled with how long it took, never replayed as context - [x] Model admin as a list plus a page per model; scales to hundreds - [x] Ordering, pinning (a sidebar shortcut, *not* a reordering), instance default, per-user default, images, capability flags - [x] Custom model picker showing avatars, descriptions and capabilities ### Attachments - [x] Drag, paste or pick images, PDFs and text files - [x] Images downscaled and sent to vision models as content parts - [x] PDF and text extracted at upload and placed in the prompt - [x] Type decided by inspecting bytes, random names on disk, non-images served as downloads with `nosniff` - [x] No OCR: a scanned PDF says so rather than silently contributing nothing ### People - [x] Accounts, argon2, revocable server-side sessions, self-service password change - [x] Users and groups with permissions that **union** rather than override - [x] Model access restricted to chosen groups - [x] Registration toggle, instance settings stored in the database ### Prompts - [x] Three layers — instance, model, chat — with the most specific winning **outright** rather than being concatenated - [x] Every injected fragment editable at `/admin/prompts`: the tool guidance, the memory and skill sections, the seam above the authored prompt, and the request that names a chat - [x] `{{variables}}` with a legend, values shown as they currently resolve, and pass-through for anything that is not one - [x] A preview of the whole assembled system message, including unsaved edits - [x] Defaults in code and overrides in the database, so improving a default still reaches an instance that never edited it ### Suggestions - [x] Admin-managed cards on the new-chat screen; three seeded once at startup ### Interface - [x] **Installable** — manifest, generated PWA icons, a service worker for the shell and a themed offline page. The worker deliberately never touches `/api/`: a reply is an event stream and caching one breaks it - [x] Two themes (`moria`, `shire`) from one set of design tokens - [x] Every control sized from `--control-h`, so rows line up by construction - [x] Toasts and dialogs of our own; no `window.confirm` anywhere - [x] Original SVG artwork generated from a single source ### Operations - [x] Additive schema sync — new tables and columns applied at startup - [x] `deploy/` — systemd unit and nginx templates, install and update scripts --- ## Not built yet In the order they are likely to be worth doing. ### Agentic execution Two modes, as originally specified: - **local** — subprocess on the machine LLeMbas runs on - **remote** — SSH connection profiles, with `shell.run` / `fs.read` / `fs.write` Needs a confirmation model before it does anything. Note that the systemd unit is deliberately only `ProtectSystem=full` rather than `strict` **because** of this — revisit the hardening when the real filesystem needs are known. ### Image generation Left until last from the start, as it needs heavy customisation. ComfyUI is already running on this machine and is the obvious first target. ### Smaller things - **OCR** for scanned PDFs - **Conversation branching** — `Message.parent_id` exists unused; needs a UI for choosing between versions, which is why rewind truncates for now - **Chat export** (Markdown, JSON) - **Semantic search** in the library — the retrieval service is one call, so an embedding backend can go behind it without touching the tools or the UI - **Archived chats** — the column exists, nothing surfaces it - **Per-user quotas** --- ## Known limits Worth knowing before they surprise someone. **One worker.** The generation registry and the stop mechanism are in-process. Running several workers needs that state in the database or a broker, because the request following a reply would not necessarily land in the process writing it. **A restart abandons replies in flight.** Shutdown cancels them and keeps what each had. There is no resume. **Schema changes are additive only.** New tables and columns apply themselves; renames, drops and retypes are manual against the SQLite file. `MANUAL_STEPS` in `db/migrations.py` is where such a step gets recorded. **Attachments live on disk, unreferenced files are swept at startup.** No deduplication, no size quota. **Unread is polled every 10 seconds.** A push channel would be more responsive but means an always-on connection per tab for the sake of a green dot. **Installing needs HTTPS or localhost.** Service workers are unavailable over plain HTTP, so a LAN install without TLS is a normal browser tab. The microphone is unavailable for the same reason. **Tool calling needs a model that supports it.** The `tools` flag is an administrator's assertion, not something endpoints reliably advertise. Set it on a model that cannot, and its replies fail rather than degrade. **Library search is keyword, not semantic.** FTS5 ranks well and needs no dependency or embedding endpoint, but "how do I get paid" will not find a document that says "invoicing". **A model can write its own skills, and they take effect at once.** Marked as model-authored and fully revertible, but a model that has just read a hostile page could save a skill that outlives the conversation. The mitigation is that it is visible and undoable, not that it was prevented. --- ## Deliberate decisions Recorded because each looks like an oversight until you know the reason. - **No JavaScript build step.** Browser libraries are hash-pinned and committed. A self-hosted tool should work offline and not report page views to a CDN. - **Permissions union, never deny.** With denies, "why can this user not do X" cannot be answered without simulating every group. - **System prompts replace, never stack.** Two layers that disagree give the model contradictory instructions and nobody can tell which is losing. - **Rewind truncates, does not branch.** Branching needs a UI for choosing between versions; "go back and try again from here" is what was asked for. - **Pinning is a shortcut, not an ordering.** A picker whose order silently differs from the admin screen is confusing. - **Images only reach models marked `vision`.** Not graceful degradation: most endpoints reject the entire request rather than ignoring an image part. Tools are gated the same way, for the same reason. - **Sharing grants reading, never writing.** Two people editing one note with no history and no merge is worse than the inconvenience of copying it. - **Memory is never shareable.** A record about a person is not content to hand round. - **Knowledge attached to a message is copied, not referenced.** History must not change under a conversation because a document was edited later. - **The harness is prepended to the authored prompt, not a fourth layer.** It describes the machinery; the authored layers describe the behaviour. Only one authored layer still wins. - **Tool results are not replayed.** Like reasoning: the answer already contains what the model made of them, and replaying stale results into every later request wastes the window and sends small models into search loops. - **The service worker caches the shell, never a page with a user in it.** A cached conversation would be a snapshot that silently went stale, belonging to whoever was signed in last. - **Markdown rendered server-side.** One code path produces the streamed and the stored view, so they cannot disagree. - **This repository is public.** Deployment hostnames, ports and paths stay out of it; `deploy/` is templates, and the real values live in private notes.