# LLeMbas — plan and status Where the project is, what is deliberately not built yet, and the decisions that would be expensive to revisit. Kept current as work lands; the detail of *how* things work lives in [`CLAUDE.md`](CLAUDE.md). **Status:** usable daily. Streaming chat, attachments, reasoning, tool calling with web search, a knowledge library, notes, memory and skills, speech in and out, users and groups, model administration, installable as an app. 428 tests, `ruff` clean. --- ## The shape of it A self-hosted web UI for OpenAI-compatible endpoints, written in Python, themed after Middle-earth. | | | |---|---| | Stack | FastAPI + Jinja + htmx + a little Alpine | | Build step | none — no Node, no npm, no CDN at runtime | | Database | SQLite, schema synchronised additively at startup | | Deployment | systemd unit + nginx vhost, one worker | These are load-bearing. Dropping the no-build rule or moving off SQLite would be a different project, not a refactor. --- ## Done ### Chat - [x] Streaming replies over server-sent events - [x] **Markdown renders progressively** — re-rendered whole every 100ms rather than appending tokens, because a list or code fence is only correct once its context exists - [x] Syntax highlighting (Pygments), sanitised with nh3 - [x] **Generation runs in the background** — a task, not the request. Navigate away, open another chat, close the tab: the reply keeps being written and reattaching replays the whole state - [x] **Stop** — the send button becomes Stop while writing; what arrived is kept - [x] **Rewind** — edit one of your own turns and the conversation runs on from there. Truncates rather than branching - [x] Copy, regenerate, automatic chat titles - [x] Chats created on first message, so an abandoned composer leaves nothing - [x] **Unread indicator** — a green dot and a toast when a reply lands while you were elsewhere - [x] Folders, arbitrarily nested; deleting one keeps the chats inside it ### Tools - [x] **Tool calling** — one reply is a bounded loop of requests, not one request. Text produced before a call is kept - [x] **Web search** as the first tool: DuckDuckGo (no setup), SearXNG or Firecrawl, chosen in the admin area - [x] Only offered to models flagged `tools`, because an endpoint without support rejects the whole request rather than ignoring the array - [x] Sources stay in the transcript; results are **not** replayed as context on the next turn, for the same reasons reasoning is not ### The library - [x] **Knowledge** — documents, images and saved web pages, ingested through the same pipeline as chat attachments, searched with SQLite FTS5 - [x] **Notes** — longer things the model writes down and searches later; editable by hand, because they are yours - [x] **Memory** — short facts, injected on every turn to a budget rather than searched, and managed in your settings - [x] **Skills** — saved procedures. Only the name and description are injected; the body is fetched when the model decides it applies - [x] A model may write and revise its own notes, memories and skills. Every skill revision is kept, attributed and revertible — the safety story is a record and a way back, not a gate - [x] **Sharing** — any of the three can be shared with a group or with named people, read-only. One visibility rule, and administrators do not bypass it - [x] **The harness** — an operational prompt assembled from what a model actually has, so the tools get used rather than ignored - [x] Attach menu: file, image, a web page fetched on the spot, or a document from the library ### Audio - [x] **Dictation** — record in the composer, transcribed by any OpenAI-shaped `/v1/audio/transcriptions` endpoint. The recording never touches disk - [x] **Read aloud** — any `/v1/audio/speech` endpoint, with the voice list discovered from the server where it offers one - [x] Instance defaults in Admin, per-reader overrides in Settings — voice, speed, dictation language, and whether replies play automatically ### Models and reasoning - [x] OpenAI-compatible connections with encrypted keys and model discovery - [x] **Reasoning display** — `reasoning_content` and inline `` tags, collapsed by default, labelled with how long it took, never replayed as context - [x] Model admin as a list plus a page per model; scales to hundreds - [x] Ordering, pinning (a sidebar shortcut, *not* a reordering), instance default, per-user default, images, capability flags - [x] Custom model picker showing avatars, descriptions and capabilities ### Attachments - [x] Drag, paste or pick images, PDFs and text files - [x] Images downscaled and sent to vision models as content parts - [x] PDF and text extracted at upload and placed in the prompt - [x] Type decided by inspecting bytes, random names on disk, non-images served as downloads with `nosniff` - [x] No OCR: a scanned PDF says so rather than silently contributing nothing ### People - [x] Accounts, argon2, revocable server-side sessions, self-service password change - [x] Users and groups with permissions that **union** rather than override - [x] Model access restricted to chosen groups - [x] Registration toggle, instance settings stored in the database ### Prompts - [x] Three layers — instance, model, chat — with the most specific winning **outright** rather than being concatenated ### Interface - [x] **Installable** — manifest, generated PWA icons, a service worker for the shell and a themed offline page. The worker deliberately never touches `/api/`: a reply is an event stream and caching one breaks it - [x] Two themes (`moria`, `shire`) from one set of design tokens - [x] Every control sized from `--control-h`, so rows line up by construction - [x] Toasts and dialogs of our own; no `window.confirm` anywhere - [x] Original SVG artwork generated from a single source ### Operations - [x] Additive schema sync — new tables and columns applied at startup - [x] `deploy/` — systemd unit and nginx templates, install and update scripts --- ## Not built yet In the order they are likely to be worth doing. ### Custom tools and MCP servers An MCP client managing configured servers, their tools surfaced alongside the built-in ones. The loop they plug into exists now — `services/tools.py` is a registry of thirteen tools and `services/generation.py` already runs bounded rounds — so this is a client and an admin screen rather than a change to how chat works. ### Agentic execution Two modes, as originally specified: - **local** — subprocess on the machine LLeMbas runs on - **remote** — SSH connection profiles, with `shell.run` / `fs.read` / `fs.write` Needs a confirmation model before it does anything. Note that the systemd unit is deliberately only `ProtectSystem=full` rather than `strict` **because** of this — revisit the hardening when the real filesystem needs are known. ### Image generation Left until last from the start, as it needs heavy customisation. ComfyUI is already running on this machine and is the obvious first target. ### Smaller things - **OCR** for scanned PDFs - **Conversation branching** — `Message.parent_id` exists unused; needs a UI for choosing between versions, which is why rewind truncates for now - **Chat export** (Markdown, JSON) - **Semantic search** in the library — the retrieval service is one call, so an embedding backend can go behind it without touching the tools or the UI - **Archived chats** — the column exists, nothing surfaces it - **Per-user quotas** --- ## Known limits Worth knowing before they surprise someone. **One worker.** The generation registry and the stop mechanism are in-process. Running several workers needs that state in the database or a broker, because the request following a reply would not necessarily land in the process writing it. **A restart abandons replies in flight.** Shutdown cancels them and keeps what each had. There is no resume. **Schema changes are additive only.** New tables and columns apply themselves; renames, drops and retypes are manual against the SQLite file. `MANUAL_STEPS` in `db/migrations.py` is where such a step gets recorded. **Attachments live on disk, unreferenced files are swept at startup.** No deduplication, no size quota. **Unread is polled every 10 seconds.** A push channel would be more responsive but means an always-on connection per tab for the sake of a green dot. **Installing needs HTTPS or localhost.** Service workers are unavailable over plain HTTP, so a LAN install without TLS is a normal browser tab. The microphone is unavailable for the same reason. **Tool calling needs a model that supports it.** The `tools` flag is an administrator's assertion, not something endpoints reliably advertise. Set it on a model that cannot, and its replies fail rather than degrade. **Library search is keyword, not semantic.** FTS5 ranks well and needs no dependency or embedding endpoint, but "how do I get paid" will not find a document that says "invoicing". **A model can write its own skills, and they take effect at once.** Marked as model-authored and fully revertible, but a model that has just read a hostile page could save a skill that outlives the conversation. The mitigation is that it is visible and undoable, not that it was prevented. --- ## Deliberate decisions Recorded because each looks like an oversight until you know the reason. - **No JavaScript build step.** Browser libraries are hash-pinned and committed. A self-hosted tool should work offline and not report page views to a CDN. - **Permissions union, never deny.** With denies, "why can this user not do X" cannot be answered without simulating every group. - **System prompts replace, never stack.** Two layers that disagree give the model contradictory instructions and nobody can tell which is losing. - **Rewind truncates, does not branch.** Branching needs a UI for choosing between versions; "go back and try again from here" is what was asked for. - **Pinning is a shortcut, not an ordering.** A picker whose order silently differs from the admin screen is confusing. - **Images only reach models marked `vision`.** Not graceful degradation: most endpoints reject the entire request rather than ignoring an image part. Tools are gated the same way, for the same reason. - **Sharing grants reading, never writing.** Two people editing one note with no history and no merge is worse than the inconvenience of copying it. - **Memory is never shareable.** A record about a person is not content to hand round. - **Knowledge attached to a message is copied, not referenced.** History must not change under a conversation because a document was edited later. - **The harness is prepended to the authored prompt, not a fourth layer.** It describes the machinery; the authored layers describe the behaviour. Only one authored layer still wins. - **Tool results are not replayed.** Like reasoning: the answer already contains what the model made of them, and replaying stale results into every later request wastes the window and sends small models into search loops. - **The service worker caches the shell, never a page with a user in it.** A cached conversation would be a snapshot that silently went stale, belonging to whoever was signed in last. - **Markdown rendered server-side.** One code path produces the streamed and the stored view, so they cannot disagree. - **This repository is public.** Deployment hostnames, ports and paths stay out of it; `deploy/` is templates, and the real values live in private notes.