Files
LLeMbas/PLAN.md
T
Jaroslav Beneš ecb52e9978 MCP servers, over streamable HTTP
A server is a row with a URL; its tools are discovered by a button and
cached, then offered beside the built-in ones. Written by hand rather than
taken from the reference SDK, because that SDK's transport does its own
connecting -- and the one thing that must not be bypassed is check_url on
every hop. Owning the transport is the point; the framing beside it is the
small part.

Sessions are per call: initialize, initialized, the call, a best-effort
DELETE. Caching one wants an owner, a TTL, eviction, a lock and a shutdown
hook, and the server may expire it under all of that anyway -- ToolContext
is a session-free snapshot precisely so nothing in a tool holds live state.

A server's names and descriptions reach the model as instructions and are
bounded before they do; what it returns is escaped preformatted text, never
markdown. Tools are namespaced per server, so two servers exposing "search"
do not collide and neither shadows a built-in.

Also: a round's calls now run together under a semaphore, results indexed
so each tool turn stays paired with its call, and generation.status names
what is running -- a remote tool is latency-bound, and a silent pause is
what a hang looks like.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 16:44:29 +02:00

276 lines
14 KiB
Markdown

# LLeMbas — plan and status
Where the project is, what is deliberately not built yet, and the decisions
that would be expensive to revisit. Kept current as work lands; the detail of
*how* things work lives in [`CLAUDE.md`](CLAUDE.md).
**Status:** usable daily. Streaming chat, attachments, reasoning, tool calling
with web search, a knowledge library, notes, memory and skills, speech in and
out, users and groups, model administration, installable as an app. 437 tests,
`ruff` clean.
---
## The shape of it
A self-hosted web UI for OpenAI-compatible endpoints, written in Python, themed
after Middle-earth.
| | |
|---|---|
| Stack | FastAPI + Jinja + htmx + a little Alpine |
| Build step | none — no Node, no npm, no CDN at runtime |
| Database | SQLite, schema synchronised additively at startup |
| Deployment | systemd unit + nginx vhost, one worker |
These are load-bearing. Dropping the no-build rule or moving off SQLite would
be a different project, not a refactor.
---
## Done
### Chat
- [x] Streaming replies over server-sent events
- [x] **Markdown renders progressively** — re-rendered whole every 100ms rather
than appending tokens, because a list or code fence is only correct once
its context exists
- [x] Syntax highlighting (Pygments), sanitised with nh3
- [x] **Generation runs in the background** — a task, not the request. Navigate
away, open another chat, close the tab: the reply keeps being written and
reattaching replays the whole state
- [x] **Stop** — the send button becomes Stop while writing; what arrived is kept
- [x] **Rewind** — edit one of your own turns and the conversation runs on from
there. Truncates rather than branching
- [x] Copy, regenerate, automatic chat titles
- [x] Chats created on first message, so an abandoned composer leaves nothing
- [x] **Unread indicator** — a green dot and a toast when a reply lands while
you were elsewhere
- [x] Folders, arbitrarily nested; deleting one keeps the chats inside it
- [x] Per-reply metrics — tokens, context used as a percentage, tokens/second,
live while streaming and kept afterwards. Estimated with a `~` when the
endpoint reports no usage
- [x] Compaction — a button, and automatically at a configurable percentage of
the model's context. Summarised turns are kept and collapsed, not deleted
- [x] Temporary chats — never listed, swept after a day, with a Keep button
- [x] An admin-only request inspector beside the thread
### Tools
- [x] **Tool calling** — one reply is a bounded loop of requests, not one
request. Text produced before a call is kept
- [x] **Web search** as the first tool: DuckDuckGo (no setup), SearXNG or
Firecrawl, chosen in the admin area
- [x] Only offered to models flagged `tools`, because an endpoint without
support rejects the whole request rather than ignoring the array
- [x] Sources stay in the transcript; results are **not** replayed as context on
the next turn, for the same reasons reasoning is not
- [x] A round's calls run together, and the reply says which tool is running —
a remote tool taking seconds with nothing streaming looks like a hang
- [x] **Custom HTTP tools** — an administrator describes one call: a JSON Schema,
a URL template, headers, an encrypted secret and how to read the answer.
Arguments may fill a hole but never move the target: the scheme and host
are literal, values are escaped for where they land, and the origin is
pinned afterwards
- [x] **MCP servers** over streamable HTTP — a hand-written client, so that
`check_url` runs on every hop rather than being bypassed by somebody
else's transport. Tools are discovered and cached by a button, namespaced
per server, and a server's own descriptions are bounded before they reach
a model as instructions
- [x] Both gated like the built-ins — a model capability, a permission — and
restrictable to groups, with guidance of their own on `/admin/prompts`
- [x] Local MCP over stdio is deliberately absent: spawning a subprocess is the
agentic-execution feature and wants a confirmation model first
### The library
- [x] **Knowledge bases** — documents, images and saved web pages, grouped into
named collections and ingested through the same pipeline as chat
attachments, searched with SQLite FTS5
- [x] A chat can be pointed at particular bases, so "answer from the contracts
folder" is a different question from "answer from everything I have"
- [x] **Notes** — longer things the model writes down and searches later;
editable by hand, because they are yours
- [x] **Memory** — short facts, injected on every turn to a budget rather than
searched, and managed in your settings
- [x] **Skills** — saved procedures. Only the name and description are injected;
the body is fetched when the model decides it applies
- [x] A model may write and revise its own notes, memories and skills. Every
skill revision is kept, attributed and revertible — the safety story is a
record and a way back, not a gate
- [x] **Sharing** — a knowledge base, a note or a skill can be shared with a
group or with named people, read-only. One visibility rule, and
administrators do not bypass it. Documents are shared through their base
- [x] **The harness** — an operational prompt assembled from what a model
actually has, so the tools get used rather than ignored
- [x] Attach menu: file, image, a web page fetched on the spot, or a document
from the library
### Audio
- [x] **Dictation** — record in the composer, transcribed by any OpenAI-shaped
`/v1/audio/transcriptions` endpoint. The recording never touches disk
- [x] **Read aloud** — any `/v1/audio/speech` endpoint, with the voice list
discovered from the server where it offers one
- [x] Instance defaults in Admin, per-reader overrides in Settings — voice,
speed, dictation language, and whether replies play automatically
### Models and reasoning
- [x] OpenAI-compatible connections with encrypted keys and model discovery
- [x] **Reasoning display**`reasoning_content` and inline `<think>` tags,
collapsed by default, labelled with how long it took, never replayed as
context
- [x] Model admin as a list plus a page per model; scales to hundreds
- [x] Ordering, pinning (a sidebar shortcut, *not* a reordering), instance
default, per-user default, images, capability flags
- [x] Custom model picker showing avatars, descriptions and capabilities
### Attachments
- [x] Drag, paste or pick images, PDFs and text files
- [x] Images downscaled and sent to vision models as content parts
- [x] PDF and text extracted at upload and placed in the prompt
- [x] Type decided by inspecting bytes, random names on disk, non-images served
as downloads with `nosniff`
- [x] No OCR: a scanned PDF says so rather than silently contributing nothing
### People
- [x] Accounts, argon2, revocable server-side sessions, self-service password
change
- [x] Users and groups with permissions that **union** rather than override
- [x] Model access restricted to chosen groups
- [x] Registration toggle, instance settings stored in the database
### Prompts
- [x] Three layers — instance, model, chat — with the most specific winning
**outright** rather than being concatenated
- [x] Every injected fragment editable at `/admin/prompts`: the tool guidance,
the memory and skill sections, the seam above the authored prompt, and the
request that names a chat
- [x] `{{variables}}` with a legend, values shown as they currently resolve, and
pass-through for anything that is not one
- [x] A preview of the whole assembled system message, including unsaved edits
- [x] Defaults in code and overrides in the database, so improving a default
still reaches an instance that never edited it
### Suggestions
- [x] Admin-managed cards on the new-chat screen; three seeded once at startup
### Interface
- [x] **Installable** — manifest, generated PWA icons, a service worker for the
shell and a themed offline page. The worker deliberately never touches
`/api/`: a reply is an event stream and caching one breaks it
- [x] Two themes (`moria`, `shire`) from one set of design tokens
- [x] Every control sized from `--control-h`, so rows line up by construction
- [x] Toasts and dialogs of our own; no `window.confirm` anywhere
- [x] Original SVG artwork generated from a single source
### Operations
- [x] Additive schema sync — new tables and columns applied at startup
- [x] `deploy/` — systemd unit and nginx templates, install and update scripts
---
## Not built yet
In the order they are likely to be worth doing.
### Agentic execution
Two modes, as originally specified:
- **local** — subprocess on the machine LLeMbas runs on
- **remote** — SSH connection profiles, with `shell.run` / `fs.read` / `fs.write`
Needs a confirmation model before it does anything. Note that the systemd unit
is deliberately only `ProtectSystem=full` rather than `strict` **because** of
this — revisit the hardening when the real filesystem needs are known.
### Image generation
Left until last from the start, as it needs heavy customisation. ComfyUI is
already running on this machine and is the obvious first target.
### Smaller things
- **OCR** for scanned PDFs
- **Conversation branching** — `Message.parent_id` exists unused; needs a UI for
choosing between versions, which is why rewind truncates for now
- **Chat export** (Markdown, JSON)
- **Semantic search** in the library — the retrieval service is one call, so an
embedding backend can go behind it without touching the tools or the UI
- **Archived chats** — the column exists, nothing surfaces it
- **Per-user quotas**
---
## Known limits
Worth knowing before they surprise someone.
**One worker.** The generation registry and the stop mechanism are in-process.
Running several workers needs that state in the database or a broker, because
the request following a reply would not necessarily land in the process writing
it.
**A restart abandons replies in flight.** Shutdown cancels them and keeps what
each had. There is no resume.
**Schema changes are additive only.** New tables and columns apply themselves;
renames, drops and retypes are manual against the SQLite file. `MANUAL_STEPS`
in `db/migrations.py` is where such a step gets recorded.
**Attachments live on disk, unreferenced files are swept at startup.** No
deduplication, no size quota.
**Unread is polled every 10 seconds.** A push channel would be more responsive
but means an always-on connection per tab for the sake of a green dot.
**Installing needs HTTPS or localhost.** Service workers are unavailable over
plain HTTP, so a LAN install without TLS is a normal browser tab. The
microphone is unavailable for the same reason.
**Tool calling needs a model that supports it.** The `tools` flag is an
administrator's assertion, not something endpoints reliably advertise. Set it on
a model that cannot, and its replies fail rather than degrade.
**Library search is keyword, not semantic.** FTS5 ranks well and needs no
dependency or embedding endpoint, but "how do I get paid" will not find a
document that says "invoicing".
**A model can write its own skills, and they take effect at once.** Marked as
model-authored and fully revertible, but a model that has just read a hostile
page could save a skill that outlives the conversation. The mitigation is that
it is visible and undoable, not that it was prevented.
---
## Deliberate decisions
Recorded because each looks like an oversight until you know the reason.
- **No JavaScript build step.** Browser libraries are hash-pinned and committed.
A self-hosted tool should work offline and not report page views to a CDN.
- **Permissions union, never deny.** With denies, "why can this user not do X"
cannot be answered without simulating every group.
- **System prompts replace, never stack.** Two layers that disagree give the
model contradictory instructions and nobody can tell which is losing.
- **Rewind truncates, does not branch.** Branching needs a UI for choosing
between versions; "go back and try again from here" is what was asked for.
- **Pinning is a shortcut, not an ordering.** A picker whose order silently
differs from the admin screen is confusing.
- **Images only reach models marked `vision`.** Not graceful degradation: most
endpoints reject the entire request rather than ignoring an image part. Tools
are gated the same way, for the same reason.
- **Sharing grants reading, never writing.** Two people editing one note with no
history and no merge is worse than the inconvenience of copying it.
- **Memory is never shareable.** A record about a person is not content to hand
round.
- **Knowledge attached to a message is copied, not referenced.** History must not
change under a conversation because a document was edited later.
- **The harness is prepended to the authored prompt, not a fourth layer.** It
describes the machinery; the authored layers describe the behaviour. Only one
authored layer still wins.
- **Tool results are not replayed.** Like reasoning: the answer already contains
what the model made of them, and replaying stale results into every later
request wastes the window and sends small models into search loops.
- **The service worker caches the shell, never a page with a user in it.** A
cached conversation would be a snapshot that silently went stale, belonging to
whoever was signed in last.
- **Markdown rendered server-side.** One code path produces the streamed and
the stored view, so they cannot disagree.
- **This repository is public.** Deployment hostnames, ports and paths stay out
of it; `deploy/` is templates, and the real values live in private notes.