Files
LLeMbas/README.md
T
Homer cdad9f0bc7 1.0.0
The version, the changelog entry, the plan and the README. Nothing else,
which is what makes this readable as a release rather than as work.

CHANGELOG.md's 1.0.0 entry is assembled from every version below it, as
that file has said it would be since it was written: those shipped as a
running deployment rather than as releases, and this is what they add up
to. It is also what an administrator reads -- /admin/updates takes release
notes out of the annotated tag, so the tag message is this entry.

It says what arrived, then the part worth reading: the nine things that
had shipped looking correct and were found by five audit passes. Then
where the edges are, because a first release should say what it does not
do before somebody finds out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:26:44 +02:00

503 lines
25 KiB
Markdown

<p align="center">
<img src="assets/banner.svg" alt="LLeMbas — waybread for the long road of thought" width="100%">
</p>
<p align="center">
<strong>A self-hosted web UI for your language models, written in Python.</strong><br>
Talks to anything that speaks the OpenAI API. Themed after Middle-earth.
</p>
<p align="center">
<img alt="Version 1.0.0" src="https://img.shields.io/badge/version-1.0.0-6B8E4E?style=flat-square">
<img alt="Python 3.11+" src="https://img.shields.io/badge/python-3.11%2B-3E6B7A?style=flat-square">
<img alt="License GPL-3.0" src="https://img.shields.io/badge/license-GPL--3.0-C9A227?style=flat-square">
<img alt="No Node required" src="https://img.shields.io/badge/build%20step-none-6B8E4E?style=flat-square">
</p>
---
*Lembas* is the Elvish waybread — one bite sustains a traveller for a day's
march. The capitals hide what it runs on: **LLeM**bas.
## Why this exists
Most self-hosted LLM front-ends are large JavaScript applications with a Python
API bolted underneath. LLeMbas is the other way round: **server-rendered
Python**, with htmx and a little Alpine for interactivity. There is no
`package.json`, no bundler, no build step, and nothing is fetched from a CDN at
runtime. Clone it, `pip install -e .`, run it.
## Features
**Working now**
- **Chats** — streaming replies, Markdown with server-side syntax highlighting,
copy and regenerate, automatic chat titles. Chats are created when you send
the first message, so an abandoned one never clutters the sidebar
- **System prompts** — instance-wide, per-model and per-chat, with the most
specific winning outright
- **Reasoning display** — thinking streams into its own collapsible block
(closed by default), labelled with how long it took, and is never replayed as
context
- **Live Markdown** — formatting appears as the model writes, not at the end
- **Stop and rewind** — cut a reply short and keep what arrived, or edit an
earlier message and run the conversation on from there
- **Replies keep running in the background** — navigate away, open another
chat, close the tab; a green dot and a notification tell you when it lands
- **Attachments** — drag, paste or pick images, PDFs and text files. Images are
downscaled and sent to vision models; PDF and text content is extracted and
put in the prompt
- **`@` to name something** — a document from your library, or in an agent chat
a file in the project directory. The reference stays in the sentence you are
writing and the contents come with it
- **`/` for commands** — `/compact`, `/usage`, `/mode plan`, `/effort high`,
`/model`, `/title`, `/terminal`, `/theme`. The list appears as you type and
filters as you go; `/help` shows all of them with the keyboard shortcuts
beside them. A message that merely starts with a slash is still sent as
written, and both `@` and a recognised command are marked in the box as you
type so you can see what will happen before you press Enter
- **Reasoning effort** — `/effort low`, `medium` or `high` on a model marked as
reasoning, with a per-model default in the admin area. Sent two ways at once,
because there is no single field every endpoint reads
- **Folders** — arbitrarily nested, delete a folder without losing the chats
inside it
- **Web search** — offered to the model as a tool it calls when a question needs
it. DuckDuckGo out of the box (no account, no key), or point it at your own
SearXNG, or Firecrawl. The sources stay in the transcript
- **Your own tools** — describe an HTTP call in the admin area (a schema, a URL
template, a secret) and a model can make it. Or add an **MCP server** by URL
and its tools appear beside the built-in ones. Both restrictable to groups,
and neither can be pointed at your own network unless you say so
- **Agent chats** — start a chat as an *Agent* instead, pointed at one of your
own SSH connections and a directory on it, and a model can read files, write
files and run commands **there**. Nothing ever runs on the machine LLeMbas
itself is on. What it may do without asking is a mode you set and can change
mid-conversation: *Manual* shows you everything first, *Edit* writes freely
but asks before commands, *Auto* asks about nothing, and *Plan* reads freely,
changes nothing, and finishes by proposing steps you can carry out with one
button. Adding a host shows you its fingerprint before anything is sent to it
- **A terminal beside the chat** — the same connection, a real shell, opened and
closed like any panel. It survives closing the panel and reloading the page,
so a build keeps running; the model cannot see it, and a button hands it the
output you choose
- **It can ask you things** — a model that needs a decision can stop and put a
few questions on one card, with answers to pick from and a box to write your
own. In any chat, not only an agent one
- **Speech in and out** — dictate a message and have replies read aloud, against
any OpenAI-compatible audio endpoint (whisper.cpp, Speaches, Kokoro…). Each
person picks their own voice
- **A library** — four places a model can reach for. **Knowledge**: documents,
images and web pages you collect, grouped into named bases so a chat can be
pointed at just the right one, searched before the web. **Notes**: longer
things it writes down and finds again later. **Memory**: short facts about you,
in front of it on every turn. **Skills**: saved procedures it can follow, and
write. All of it visible and editable by you, and shareable with a group or a
person, read-only
- **Installable** — add it to a phone home screen or a desktop launcher and it
runs in its own window
- **OpenAI connections** — point at OpenAI, LM Studio, vLLM, llama.cpp,
llama-swap, Ollama or OpenRouter; models are discovered and cached
- **Model settings** — searchable, filterable list with a page per model:
ordering, pinned models, an instance default and a per-user default, custom
names, descriptions and images. Scales to hundreds of models
- **Things that happen because time passed** — say "every Monday at nine" and a
model can set it up itself, against the same recurrence rule the manual form
uses. A run can file a **report** you read later, send you a message, or work
on in a chat of its own. The reply says the timing back in words, which is the
one moment anybody can check that Monday was understood as Monday
- **News that finds you** — a dot in the sidebar, a count in the tab title while
you are looking elsewhere, and **web push** so a schedule firing at seven in
the morning reaches a browser that is shut. Opt-in per device
- **Helpers** — a reply can hand a self-contained piece of work to another model
that runs on its own and reports back, several at once, so research fans out
instead of queueing. A helper cannot ask questions, cannot send helpers of its
own, and on a machine runs only a fixed list of read-only commands
- **Drawing** — point it at a ComfyUI and a model can make images, against
workflow templates and defaults you set: size, steps, sampler, scheduler,
checkpoint. It reviews its own result and can try again
- **Semantic search** — pick an embedding model and library search fuses keyword
and meaning, so *"how do I get paid"* finds a document that says *"invoicing"*.
Choosing none is not a degraded mode: it is byte-for-byte the keyword search
that was always there, with nothing written and no requests made
- **Users, groups & permissions** — per-group grants that union rather than
override, model access restricted to chosen groups, read and write split for
notes, memory and skills, and a screen that answers *"what can this account
actually do?"* by naming where each permission came from
- **Quotas** — monthly tokens, concurrent replies, agent wall clock, images a
day, helpers a reply. Resolved by maximum across a person's groups, with zero
meaning *no limit*
- **Sharing** — hand a document, a note, a skill or a report to a group or a
person, read-only, with a *Shared with me* filter in every listing
- **Make it yours** — name, tagline, logo, favicon and launcher icons; the
Middle-earth wording is editable data; custom **themes** defined as a set of
colours rather than a stylesheet, and global CSS overrides
- **Accounts** — first account becomes the administrator, argon2 password
hashing, revocable server-side sessions, self-service password change,
admin-managed accounts
- **Admin settings** — registration, upload and extraction limits, prompt
fragments, and an **Updates** page showing what is running, what is available
and what changed between
- **Two themes and your own** — *Moria* (dark), *Shire* (light), and as many
more as you care to define
**Planned**
OCR for scanned PDFs · conversation branching · chat export · archived chats.
See [PLAN.md](PLAN.md) for what is built, what is not, and why.
## Quick start
```bash
git clone https://git.houmeres.sk/Houmeres/LLeMbas.git
cd LLeMbas
python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev,search,ssh]" # search: DuckDuckGo. ssh: agent chats.
# Drop either if you do not want it
cp .env.example .env
lembas secret-key # paste the result into LEMBAS_SECRET_KEY
lembas serve # http://127.0.0.1:8080
```
Open the address and create the first account — it becomes the administrator.
Then go to **Admin → Connections** and add an endpoint. For a local runner that
is usually `http://localhost:1234/v1` with no API key. Press **Test & refresh**
and its models appear in the chat model picker.
> The vendored browser libraries (htmx, Alpine) are committed, so no network
> access is needed to run. To re-fetch or bump them:
> `python scripts/fetch_vendor.py --update`.
### Web search
**Admin → Web search.** DuckDuckGo needs nothing beyond the `search` extra
above. SearXNG needs its JSON format enabled — add `- json` under
`search.formats` in its `settings.yml`, or every search fails. Firecrawl needs
an API key.
Search is offered to the model as a *tool*, so it decides when a question needs
looking up. It is only offered to models marked **tools** under
**Admin → Models**: an endpoint without tool support rejects the whole request
rather than ignoring the extra field, so the flag is a real switch and not a
hint.
### Audio
**Admin → Audio.** Two endpoints, because they are usually two servers:
| | Speaks | Example |
|---|---|---|
| Dictation | `POST /v1/audio/transcriptions` | whisper.cpp's `whisper-server`, Speaches, faster-whisper-server |
| Read aloud | `POST /v1/audio/speech` | Kokoro-FastAPI, OpenAI |
If the speech endpoint also answers `GET /v1/audio/voices` the voice list is
read from it, and each person can pick their own under **Settings → Audio**.
Recorded audio is passed straight through and never written to disk.
> The microphone needs HTTPS or localhost. Browsers do not grant it over plain
> HTTP, so a LAN install without TLS will not offer dictation.
### Agent chats
**Admin → Agents** to turn the feature on, then **Connections** in the sidebar
to add a machine. Three things have to line up before an agent chat can start:
the feature enabled, the *Run commands* permission, and a model flagged **Agent
execution**. All three are off by default, on purpose.
Nothing an agent does runs on the machine LLeMbas is on. Commands go to a host
you name over SSH, which means **the containment is that host** — a container
built for the job is a very different thing from a key to a server you care
about, and LLeMbas cannot tell them apart. A throwaway container is the intended
shape:
```bash
docker run -d --name agent-box -p 127.0.0.1:2222:22 <an sshd image>
```
Adding a connection does not connect to it. **Check** shows you the host's
fingerprint with nothing sent — not your username, not your key — and only
accepting pins it. If that host later answers with a different key, it is
refused rather than quietly trusted.
Then start a chat with the **Agent** toggle, pick the connection, browse to a
directory, and choose a mode — all of it under the message box, before you send
anything. The connection and the directory are fixed once the chat exists; the
mode changes at any time and stays where you chose it:
| | Reads | Writes files | Runs commands |
|---|---|---|---|
| **Manual** | asks | asks | asks |
| **Edit** | free | free | asks |
| **Auto** | free | free | free |
| **Plan** | free | asks | asks |
The mode is enforced in the reply loop, not written into the prompt: everything
a model reads — a web page, a README, the last command's output — is untrusted,
and a rule that lives only in a system message is one a poisoned file can argue
with. In **Auto**, nothing stands between that and a command running.
*Plan* finishes by proposing steps, with a button that carries them out — which
switches to *Edit*, never *Auto*, because the plan was written under a mode
where every command still asked.
#### What the model knows about the directory
An agent chat starts by listing the project directory, so a reply does not spend
its first rounds finding out what is there. It is one read-only command —
`git ls-files` in a repository, so `.gitignore` is honoured for free, otherwise
`find` with the usual noise pruned — and it is cached and shared by every chat
pointed at the same place.
What reaches the model is budgeted rather than dumped: a directory that will not
fit is shown as `node_modules/ (4,102 files)` and the model is told to open it
itself if it needs to. **Admin → Agents** sets the budget, and `0` keeps the
listing for the `@` picker while putting none of it in the prompt.
Listing a directory and browsing one are things *you* asked for, not things a
model chose, so neither goes through the modes above. Worth knowing if you read
**Manual** as "nothing happens without me": it means nothing the *model* does.
#### The terminal
An agent chat has a **Terminal** button in its header, which opens a real shell
on that chat's connection, in its directory, beside the conversation. It needs
the *Open a terminal* permission, which is off by default.
The modes above do not apply to it. They exist because a model reads pages,
files and command output it did not write; you hold the credential and could
open the same shell with an ssh client, so nothing you type is queued for your
own approval. The model cannot see the panel either — three buttons in its
header decide what it sees: **Copy** takes the last command and its output to
the clipboard, **Send** puts the same into the message box, and **Auto**
collects every command you run into your next message. Nothing is ever sent on
its own; the box is where you read it first.
Knowing what "the last command" means takes a little help from the shell.
LLeMbas gives bash and zsh the same invisible markers VS Code and WezTerm use,
written into a temporary file the shell deletes itself, so it can tell one
command's output from the next and record the exit status and the directory.
Your own dotfiles are loaded first and nothing of yours is skipped. Any other
shell starts exactly as it would have; the two buttons then copy the last of the
screen as it appeared, say so, and Auto is switched off rather than guessing.
Drag the panel's left edge to make it wider — a terminal narrower than eighty
columns re-wraps everything a program prints — and the width follows you to
another browser.
The shell is not tied to the panel. Close it and a build carries on; come back,
or reload, and you reattach with the scrollback. Two tabs share one shell, and
the smaller window decides the size. It ends when nobody has watched it and
nothing has been typed for a while, when the chat is deleted, when the
connection is disabled or deleted, or when LLeMbas restarts — a deploy cuts off
whatever was running, and the panel says so rather than quietly opening a fresh
shell that has lost your working directory.
> Nothing typed here is in the transcript and nothing is logged but the opening
> and the closing. If you are running this over plain http, note that the
> session cookie is not marked `secure` so a LAN install works at all — with a
> terminal switched on, that is worth a certificate.
### The library
**Sidebar → Library**, and **Settings → Memory**. Nothing is on by default for a
model: give it the tools it should have under **Admin → Models**, where
`tools` decides whether a tool list may be sent at all and the built-in tools are
chosen one by one.
Knowledge is organised into **bases** — one per subject, project or client. A
chat with no base attached searches everything you have; tick some in the chat's
settings panel and it searches only those. Sharing happens at the base: share it
and everything in it comes too, read-only.
Search is SQLite's FTS5 — keyword matching with BM25 ranking, no embedding
service to run and nothing that stops working offline. It will not match a
paraphrase, so a line of description on a document is worth writing.
> Saving a **link** makes your server fetch a URL. Addresses on your own machine
> and network are refused unless an administrator opts in under
> **Admin → Web search**, because the address can come from a model and the
> server can reach things your browser cannot.
### Installing as an app
Open it in a browser and use *Install* (Chromium) or *Share → Add to Home
Screen* (iOS). This also needs HTTPS or localhost — service workers are
unavailable over plain HTTP, and without one there is nothing to install.
There is no offline mode beyond a page saying so. Everything is rendered by your
server, so a cached conversation would be a snapshot that silently went stale.
## Configuration
All variables are prefixed `LEMBAS_` and can live in `.env`. See
[`.env.example`](.env.example) for the annotated list.
| Variable | Default | Purpose |
|---|---|---|
| `LEMBAS_SECRET_KEY` | *generated* | Signs sessions and encrypts stored API keys. **Set this.** A generated key changes every restart, signing everyone out and making stored API keys unreadable. |
| `LEMBAS_DATA_DIR` | `./data` | SQLite database and uploads. |
| `LEMBAS_HOST` / `LEMBAS_PORT` | `127.0.0.1` / `8080` | Bind address. |
| `LEMBAS_ALLOW_SIGNUP` | `true` | Whether new users may register themselves — the *initial* value only. Once set under **Admin → General** the stored setting wins. The first account is always an admin regardless. |
| `LEMBAS_DEFAULT_THEME` | `moria` | `moria` (dark) or `shire` (light). |
| `LEMBAS_SESSION_TTL` | `2592000` | Session lifetime in seconds. |
| `LEMBAS_REQUEST_TIMEOUT` | `300` | Seconds to wait on an upstream model. |
### Commands
```bash
lembas serve # run the server
lembas info # where data lives, what is configured
lembas secret-key # generate a value for LEMBAS_SECRET_KEY
lembas create-admin # create or promote an administrator
```
## Running it somewhere
Three ways, all in this repository.
### Docker
```bash
export LEMBAS_SECRET_KEY="$(lembas secret-key)" # required; there is no default
docker compose up -d
```
One stage, no build step, non-root. The image bakes **no secret key, no data and
no `.git`** — a key inside an image is one every copy shares, and rotating it
makes stored API keys unreadable. Data lives in a named volume on `/data`.
`docker-compose.yml` publishes on `127.0.0.1` and expects a TLS proxy in front:
the service worker and the microphone both require HTTPS or localhost, so plain
http on a LAN address is a constraint rather than a preference. One replica, and
that is deliberate — the generation registry, the terminal sessions and the
schedule ticker are all in-process, so two would mean every schedule firing
twice.
**Updating a container is pulling a new image**, and `/admin/updates` says so
rather than offering a button:
```bash
docker compose pull && docker compose up -d
```
There is deliberately no in-container update helper. The one the other install
paths use restarts a systemd service; the equivalent here would be a process
inside the container reaching the Docker socket to replace the container it is
running in — which is root on the host, granted to anybody who can administer
the web interface. The image is the unit of deployment, and that is the whole
point of it.
### A machine of its own
`deploy/` holds a systemd unit, an nginx vhost, and install/update scripts. Every
template is parameterised and substituted at install time, so nothing
host-specific is committed here. See [deploy/README.md](deploy/README.md).
`deploy/lxc-install.sh` creates an unprivileged Proxmox container and runs that
same installer inside it — a wrapper around what already works rather than a
second install path:
```bash
CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
```
The container gets the **update helper by default**, unlike a bare
`install.sh`. The installer defaults it off because it cannot know what it is
installing onto; a container this script made thirty seconds ago to run one
thing, on a hypervisor you own, is not that host — and an appliance you cannot
update without a shell is one nobody updates. `INSTALL_UPDATE_HELPER=0` opts
out.
### Updating
**Admin → Updates** shows the version running, what is available on the channel
this host follows, and the commits between. `stable` is the newest `vX.Y.Z` tag;
`edge` is the branch tip, which is whatever was pushed most recently.
The button that applies an update is **opt-in**, and that is the design: the
service runs unprivileged and cannot restart itself, so the request is a file
that a systemd `.path` unit picks up and runs as root. It carries no ref and no
channel — pressing it is always "deploy the channel this host was configured
with", never "deploy something else". Install it with
`INSTALL_UPDATE_HELPER=1`; without it the page says so and prints the command to
run by hand.
Release notes come out of the annotated tag itself, so no forge API is involved
anywhere.
Root runs a **copy** of `deploy/update.sh` that the installer places outside the
checkout and root owns. It must not run the one in the checkout: that file
belongs to the unprivileged service account, so anything able to write as that
account could rewrite it and become root — and so could whoever controls the
branch, since a pull happens as that account and root would run whatever it
fetched. The cost is that changing `update.sh` needs the installer re-run, and
it tells you when your copy has fallen behind.
**If you installed the helper before this changed, re-run the installer.** The
old wiring points systemd at the checkout, and the update script now says so
loudly when it notices it is running from there.
## How it fits together
```
Browser ──form POST──▶ FastAPI ──▶ SQLite
▲ │
│ └──httpx──▶ any OpenAI-compatible endpoint
└──── server-sent events ◀───────────────┘ (streamed reply)
```
Sending a message stores the turn and returns two HTML fragments: the user's
bubble and an empty assistant bubble carrying an `sse-connect`. That opens a
server-sent event stream which appends tokens as they arrive, then replaces the
whole bubble with the finished, Markdown-rendered version. Rendering and
highlighting happen in Python, so the streamed and final views cannot disagree.
```
src/lembas/
api/ routes: auth, chats, folders, admin, pages
db/models/ SQLAlchemy schema
security/ password hashing, sessions
services/ llm client, chat orchestration, markdown, crypto, sse
web/ Jinja templates and static assets
assets/ SVG artwork masters
scripts/ artwork generator, vendored-JS fetcher
deploy/ systemd unit and nginx vhost for a real install
```
## Development
```bash
pytest # test suite
ruff check . # lint
python scripts/build_artwork.py # regenerate the SVG artwork
python scripts/fetch_vendor.py # verify vendored JS against the lockfile
```
There is no Alembic. The schema is SQLite-only and synchronised at startup:
missing tables and missing columns are added automatically, so adding a field to
a model needs nothing but a restart. Renames, drops and retypes are still manual
— see `CLAUDE.md`.
## Artwork
The logo, favicon and banner are original vector work, generated by
[`scripts/build_artwork.py`](scripts/build_artwork.py) so the mallorn leaf stays
identical across every size it appears at. The wordmark is
[Source Serif 4](https://github.com/adobe-fonts/source-serif) (SIL OFL 1.1)
converted to outlines — a README banner cannot load a webfont, and `<text>`
would render in whatever serif the reader happens to have.
## Licence
[GPL-3.0](LICENSE).
## A note on the theme
This is an independent hobby project, themed as an affectionate nod to
J.R.R. Tolkien's world. It is **not affiliated with, endorsed by, or connected
to** the Tolkien Estate, Middle-earth Enterprises, or any related rights
holder. All artwork here is original.