0bee366488
composer.js built its menu lazily inside show(), and refresh() wrote list.innerHTML before calling it. `list` is null until build() has run, so the first `/` or `@` ever typed threw a TypeError and took the handler with it. The menu has never appeared in any browser. That is why /compact "isn't there": nothing was. I shipped it having only run `node --check`, which parses the file happily. So this also brings the thing that catches it: a DOM stub driven under node -- not committed, hard rule 1 stands, it is an instrument like curl. It reproduced the crash in one run and immediately found two more: choosing a command from the menu left `/help` sitting in the box so the next Enter ran it again, and Tab completed nothing. Tab now completes and Enter runs, which is the split that matters for a command taking an argument. `.select--sm` was used three times and defined nowhere. I deleted the copy in chat.css and left a comment saying it "is defined once, in app.css", where it did not exist -- so those selects fell back to plain `.select`: width 100% in a flex row where four siblings wanted the same, all of them shrinking together until each was a few characters wide, and half a rem taller than everything beside them. That was the whole of "the connection switch needs to be wider". The connection and directory move to the topbar. They cannot change -- update_chat refuses both with a 409 -- so they are facts about the chat, of a kind with the Temporary badge, not controls on the message. The mode stays by the box. Compaction says it is working. It makes a model call that takes seconds and had no indicator anywhere: `hx-indicator` appears nowhere in this codebase, and the Generation.status channel that says "Summarising earlier messages…" for the automatic path cannot be borrowed, because it lives in the streaming bubble and this endpoint refuses to run while any message is unfinished. The overflow menu now runs the same code as /compact rather than posting for itself, so there is one implementation, one spinner, and one place the endpoint's four carefully written 409s finally reach somebody. /effort, low medium high, per chat with a per-model default. It goes out twice because there is no field that works everywhere: OpenAI and vLLM read reasoning_effort, llama.cpp's own docs say other values "have no effect" and its maintainer says the field "simply gets dropped without error or logging" -- what reaches gpt-oss behind it is chat_template_kwargs. Both are sent, and only once an effort has been chosen, so a provider strict about unknown parameters sees exactly the request it always did until somebody opts in. The control appears only on a model marked `reasoning`, a flag that has existed since the beginning with no reader at all. Mentions and recognised commands are marked as you type -- a mirror behind the textarea holding the same text with every character transparent, contributing nothing but a rounded rectangle, so a pixel of drift is a misplaced rectangle rather than a doubled glyph. A command is marked only when it resolves, so `/thoughts on this` visibly is not one before you send it. And again in the transcript, where user turns had no render step at all and now escape before they inject. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
385 lines
19 KiB
Markdown
385 lines
19 KiB
Markdown
<p align="center">
|
|
<img src="assets/banner.svg" alt="LLeMbas — waybread for the long road of thought" width="100%">
|
|
</p>
|
|
|
|
<p align="center">
|
|
<strong>A self-hosted web UI for your language models, written in Python.</strong><br>
|
|
Talks to anything that speaks the OpenAI API. Themed after Middle-earth.
|
|
</p>
|
|
|
|
<p align="center">
|
|
<img alt="Python 3.11+" src="https://img.shields.io/badge/python-3.11%2B-3E6B7A?style=flat-square">
|
|
<img alt="License GPL-3.0" src="https://img.shields.io/badge/license-GPL--3.0-C9A227?style=flat-square">
|
|
<img alt="No Node required" src="https://img.shields.io/badge/build%20step-none-6B8E4E?style=flat-square">
|
|
</p>
|
|
|
|
---
|
|
|
|
*Lembas* is the Elvish waybread — one bite sustains a traveller for a day's
|
|
march. The capitals hide what it runs on: **LLeM**bas.
|
|
|
|
## Why this exists
|
|
|
|
Most self-hosted LLM front-ends are large JavaScript applications with a Python
|
|
API bolted underneath. LLeMbas is the other way round: **server-rendered
|
|
Python**, with htmx and a little Alpine for interactivity. There is no
|
|
`package.json`, no bundler, no build step, and nothing is fetched from a CDN at
|
|
runtime. Clone it, `pip install -e .`, run it.
|
|
|
|
## Features
|
|
|
|
**Working now**
|
|
|
|
- **Chats** — streaming replies, Markdown with server-side syntax highlighting,
|
|
copy and regenerate, automatic chat titles. Chats are created when you send
|
|
the first message, so an abandoned one never clutters the sidebar
|
|
- **System prompts** — instance-wide, per-model and per-chat, with the most
|
|
specific winning outright
|
|
- **Reasoning display** — thinking streams into its own collapsible block
|
|
(closed by default), labelled with how long it took, and is never replayed as
|
|
context
|
|
- **Live Markdown** — formatting appears as the model writes, not at the end
|
|
- **Stop and rewind** — cut a reply short and keep what arrived, or edit an
|
|
earlier message and run the conversation on from there
|
|
- **Replies keep running in the background** — navigate away, open another
|
|
chat, close the tab; a green dot and a notification tell you when it lands
|
|
- **Attachments** — drag, paste or pick images, PDFs and text files. Images are
|
|
downscaled and sent to vision models; PDF and text content is extracted and
|
|
put in the prompt
|
|
- **`@` to name something** — a document from your library, or in an agent chat
|
|
a file in the project directory. The reference stays in the sentence you are
|
|
writing and the contents come with it
|
|
- **`/` for commands** — `/compact`, `/usage`, `/mode plan`, `/effort high`,
|
|
`/model`, `/title`, `/terminal`, `/theme`. The list appears as you type and
|
|
filters as you go; `/help` shows all of them with the keyboard shortcuts
|
|
beside them. A message that merely starts with a slash is still sent as
|
|
written, and both `@` and a recognised command are marked in the box as you
|
|
type so you can see what will happen before you press Enter
|
|
- **Reasoning effort** — `/effort low`, `medium` or `high` on a model marked as
|
|
reasoning, with a per-model default in the admin area. Sent two ways at once,
|
|
because there is no single field every endpoint reads
|
|
- **Folders** — arbitrarily nested, delete a folder without losing the chats
|
|
inside it
|
|
- **Web search** — offered to the model as a tool it calls when a question needs
|
|
it. DuckDuckGo out of the box (no account, no key), or point it at your own
|
|
SearXNG, or Firecrawl. The sources stay in the transcript
|
|
- **Your own tools** — describe an HTTP call in the admin area (a schema, a URL
|
|
template, a secret) and a model can make it. Or add an **MCP server** by URL
|
|
and its tools appear beside the built-in ones. Both restrictable to groups,
|
|
and neither can be pointed at your own network unless you say so
|
|
- **Agent chats** — start a chat as an *Agent* instead, pointed at one of your
|
|
own SSH connections and a directory on it, and a model can read files, write
|
|
files and run commands **there**. Nothing ever runs on the machine LLeMbas
|
|
itself is on. What it may do without asking is a mode you set and can change
|
|
mid-conversation: *Manual* shows you everything first, *Edit* writes freely
|
|
but asks before commands, *Auto* asks about nothing, and *Plan* reads freely,
|
|
changes nothing, and finishes by proposing steps you can carry out with one
|
|
button. Adding a host shows you its fingerprint before anything is sent to it
|
|
- **A terminal beside the chat** — the same connection, a real shell, opened and
|
|
closed like any panel. It survives closing the panel and reloading the page,
|
|
so a build keeps running; the model cannot see it, and a button hands it the
|
|
output you choose
|
|
- **It can ask you things** — a model that needs a decision can stop and put a
|
|
few questions on one card, with answers to pick from and a box to write your
|
|
own. In any chat, not only an agent one
|
|
- **Speech in and out** — dictate a message and have replies read aloud, against
|
|
any OpenAI-compatible audio endpoint (whisper.cpp, Speaches, Kokoro…). Each
|
|
person picks their own voice
|
|
- **A library** — four places a model can reach for. **Knowledge**: documents,
|
|
images and web pages you collect, grouped into named bases so a chat can be
|
|
pointed at just the right one, searched before the web. **Notes**: longer
|
|
things it writes down and finds again later. **Memory**: short facts about you,
|
|
in front of it on every turn. **Skills**: saved procedures it can follow, and
|
|
write. All of it visible and editable by you, and shareable with a group or a
|
|
person, read-only
|
|
- **Installable** — add it to a phone home screen or a desktop launcher and it
|
|
runs in its own window
|
|
- **OpenAI connections** — point at OpenAI, LM Studio, vLLM, llama.cpp,
|
|
llama-swap, Ollama or OpenRouter; models are discovered and cached
|
|
- **Model settings** — searchable, filterable list with a page per model:
|
|
ordering, pinned models, an instance default and a per-user default, custom
|
|
names, descriptions and images. Scales to hundreds of models
|
|
- **Users, groups & permissions** — per-group grants that union rather than
|
|
override, and model access restricted to chosen groups
|
|
- **Accounts** — first account becomes the administrator, argon2 password
|
|
hashing, revocable server-side sessions, self-service password change,
|
|
admin-managed accounts
|
|
- **Admin settings** — open or close registration from the UI, stored in the
|
|
database and effective immediately
|
|
- **Two themes** — *Moria* (dark) and *Shire* (light), switchable per user
|
|
|
|
**Planned**
|
|
|
|
Image generation · OCR for scanned PDFs · semantic search in the library.
|
|
|
|
See [PLAN.md](PLAN.md) for what is built, what is not, and why.
|
|
|
|
## Quick start
|
|
|
|
```bash
|
|
git clone https://git.houmeres.sk/Houmeres/LLeMbas.git
|
|
cd LLeMbas
|
|
|
|
python -m venv .venv && . .venv/bin/activate
|
|
pip install -e ".[dev,search,ssh]" # search: DuckDuckGo. ssh: agent chats.
|
|
# Drop either if you do not want it
|
|
|
|
cp .env.example .env
|
|
lembas secret-key # paste the result into LEMBAS_SECRET_KEY
|
|
|
|
lembas serve # http://127.0.0.1:8080
|
|
```
|
|
|
|
Open the address and create the first account — it becomes the administrator.
|
|
Then go to **Admin → Connections** and add an endpoint. For a local runner that
|
|
is usually `http://localhost:1234/v1` with no API key. Press **Test & refresh**
|
|
and its models appear in the chat model picker.
|
|
|
|
> The vendored browser libraries (htmx, Alpine) are committed, so no network
|
|
> access is needed to run. To re-fetch or bump them:
|
|
> `python scripts/fetch_vendor.py --update`.
|
|
|
|
### Web search
|
|
|
|
**Admin → Web search.** DuckDuckGo needs nothing beyond the `search` extra
|
|
above. SearXNG needs its JSON format enabled — add `- json` under
|
|
`search.formats` in its `settings.yml`, or every search fails. Firecrawl needs
|
|
an API key.
|
|
|
|
Search is offered to the model as a *tool*, so it decides when a question needs
|
|
looking up. It is only offered to models marked **tools** under
|
|
**Admin → Models**: an endpoint without tool support rejects the whole request
|
|
rather than ignoring the extra field, so the flag is a real switch and not a
|
|
hint.
|
|
|
|
### Audio
|
|
|
|
**Admin → Audio.** Two endpoints, because they are usually two servers:
|
|
|
|
| | Speaks | Example |
|
|
|---|---|---|
|
|
| Dictation | `POST /v1/audio/transcriptions` | whisper.cpp's `whisper-server`, Speaches, faster-whisper-server |
|
|
| Read aloud | `POST /v1/audio/speech` | Kokoro-FastAPI, OpenAI |
|
|
|
|
If the speech endpoint also answers `GET /v1/audio/voices` the voice list is
|
|
read from it, and each person can pick their own under **Settings → Audio**.
|
|
Recorded audio is passed straight through and never written to disk.
|
|
|
|
> The microphone needs HTTPS or localhost. Browsers do not grant it over plain
|
|
> HTTP, so a LAN install without TLS will not offer dictation.
|
|
|
|
### Agent chats
|
|
|
|
**Admin → Agents** to turn the feature on, then **Connections** in the sidebar
|
|
to add a machine. Three things have to line up before an agent chat can start:
|
|
the feature enabled, the *Run commands* permission, and a model flagged **Agent
|
|
execution**. All three are off by default, on purpose.
|
|
|
|
Nothing an agent does runs on the machine LLeMbas is on. Commands go to a host
|
|
you name over SSH, which means **the containment is that host** — a container
|
|
built for the job is a very different thing from a key to a server you care
|
|
about, and LLeMbas cannot tell them apart. A throwaway container is the intended
|
|
shape:
|
|
|
|
```bash
|
|
docker run -d --name agent-box -p 127.0.0.1:2222:22 <an sshd image>
|
|
```
|
|
|
|
Adding a connection does not connect to it. **Check** shows you the host's
|
|
fingerprint with nothing sent — not your username, not your key — and only
|
|
accepting pins it. If that host later answers with a different key, it is
|
|
refused rather than quietly trusted.
|
|
|
|
Then start a chat with the **Agent** toggle, pick the connection, browse to a
|
|
directory, and choose a mode — all of it under the message box, before you send
|
|
anything. The connection and the directory are fixed once the chat exists; the
|
|
mode changes at any time and stays where you chose it:
|
|
|
|
| | Reads | Writes files | Runs commands |
|
|
|---|---|---|---|
|
|
| **Manual** | asks | asks | asks |
|
|
| **Edit** | free | free | asks |
|
|
| **Auto** | free | free | free |
|
|
| **Plan** | free | asks | asks |
|
|
|
|
The mode is enforced in the reply loop, not written into the prompt: everything
|
|
a model reads — a web page, a README, the last command's output — is untrusted,
|
|
and a rule that lives only in a system message is one a poisoned file can argue
|
|
with. In **Auto**, nothing stands between that and a command running.
|
|
|
|
*Plan* finishes by proposing steps, with a button that carries them out — which
|
|
switches to *Edit*, never *Auto*, because the plan was written under a mode
|
|
where every command still asked.
|
|
|
|
#### What the model knows about the directory
|
|
|
|
An agent chat starts by listing the project directory, so a reply does not spend
|
|
its first rounds finding out what is there. It is one read-only command —
|
|
`git ls-files` in a repository, so `.gitignore` is honoured for free, otherwise
|
|
`find` with the usual noise pruned — and it is cached and shared by every chat
|
|
pointed at the same place.
|
|
|
|
What reaches the model is budgeted rather than dumped: a directory that will not
|
|
fit is shown as `node_modules/ (4,102 files)` and the model is told to open it
|
|
itself if it needs to. **Admin → Agents** sets the budget, and `0` keeps the
|
|
listing for the `@` picker while putting none of it in the prompt.
|
|
|
|
Listing a directory and browsing one are things *you* asked for, not things a
|
|
model chose, so neither goes through the modes above. Worth knowing if you read
|
|
**Manual** as "nothing happens without me": it means nothing the *model* does.
|
|
|
|
#### The terminal
|
|
|
|
An agent chat has a **Terminal** button in its header, which opens a real shell
|
|
on that chat's connection, in its directory, beside the conversation. It needs
|
|
the *Open a terminal* permission, which is off by default.
|
|
|
|
The modes above do not apply to it. They exist because a model reads pages,
|
|
files and command output it did not write; you hold the credential and could
|
|
open the same shell with an ssh client, so nothing you type is queued for your
|
|
own approval. The model cannot see the panel either — three buttons in its
|
|
header decide what it sees: **Copy** takes the last command and its output to
|
|
the clipboard, **Send** puts the same into the message box, and **Auto**
|
|
collects every command you run into your next message. Nothing is ever sent on
|
|
its own; the box is where you read it first.
|
|
|
|
Knowing what "the last command" means takes a little help from the shell.
|
|
LLeMbas gives bash and zsh the same invisible markers VS Code and WezTerm use,
|
|
written into a temporary file the shell deletes itself, so it can tell one
|
|
command's output from the next and record the exit status and the directory.
|
|
Your own dotfiles are loaded first and nothing of yours is skipped. Any other
|
|
shell starts exactly as it would have; the two buttons then copy the last of the
|
|
screen as it appeared, say so, and Auto is switched off rather than guessing.
|
|
|
|
Drag the panel's left edge to make it wider — a terminal narrower than eighty
|
|
columns re-wraps everything a program prints — and the width follows you to
|
|
another browser.
|
|
|
|
The shell is not tied to the panel. Close it and a build carries on; come back,
|
|
or reload, and you reattach with the scrollback. Two tabs share one shell, and
|
|
the smaller window decides the size. It ends when nobody has watched it and
|
|
nothing has been typed for a while, when the chat is deleted, when the
|
|
connection is disabled or deleted, or when LLeMbas restarts — a deploy cuts off
|
|
whatever was running, and the panel says so rather than quietly opening a fresh
|
|
shell that has lost your working directory.
|
|
|
|
> Nothing typed here is in the transcript and nothing is logged but the opening
|
|
> and the closing. If you are running this over plain http, note that the
|
|
> session cookie is not marked `secure` so a LAN install works at all — with a
|
|
> terminal switched on, that is worth a certificate.
|
|
|
|
### The library
|
|
|
|
**Sidebar → Library**, and **Settings → Memory**. Nothing is on by default for a
|
|
model: give it the tools it should have under **Admin → Models**, where
|
|
`tools` decides whether a tool list may be sent at all and the built-in tools are
|
|
chosen one by one.
|
|
|
|
Knowledge is organised into **bases** — one per subject, project or client. A
|
|
chat with no base attached searches everything you have; tick some in the chat's
|
|
settings panel and it searches only those. Sharing happens at the base: share it
|
|
and everything in it comes too, read-only.
|
|
|
|
Search is SQLite's FTS5 — keyword matching with BM25 ranking, no embedding
|
|
service to run and nothing that stops working offline. It will not match a
|
|
paraphrase, so a line of description on a document is worth writing.
|
|
|
|
> Saving a **link** makes your server fetch a URL. Addresses on your own machine
|
|
> and network are refused unless an administrator opts in under
|
|
> **Admin → Web search**, because the address can come from a model and the
|
|
> server can reach things your browser cannot.
|
|
|
|
### Installing as an app
|
|
|
|
Open it in a browser and use *Install* (Chromium) or *Share → Add to Home
|
|
Screen* (iOS). This also needs HTTPS or localhost — service workers are
|
|
unavailable over plain HTTP, and without one there is nothing to install.
|
|
|
|
There is no offline mode beyond a page saying so. Everything is rendered by your
|
|
server, so a cached conversation would be a snapshot that silently went stale.
|
|
|
|
## Configuration
|
|
|
|
All variables are prefixed `LEMBAS_` and can live in `.env`. See
|
|
[`.env.example`](.env.example) for the annotated list.
|
|
|
|
| Variable | Default | Purpose |
|
|
|---|---|---|
|
|
| `LEMBAS_SECRET_KEY` | *generated* | Signs sessions and encrypts stored API keys. **Set this.** A generated key changes every restart, signing everyone out and making stored API keys unreadable. |
|
|
| `LEMBAS_DATA_DIR` | `./data` | SQLite database and uploads. |
|
|
| `LEMBAS_HOST` / `LEMBAS_PORT` | `127.0.0.1` / `8080` | Bind address. |
|
|
| `LEMBAS_ALLOW_SIGNUP` | `true` | Whether new users may register themselves — the *initial* value only. Once set under **Admin → General** the stored setting wins. The first account is always an admin regardless. |
|
|
| `LEMBAS_DEFAULT_THEME` | `moria` | `moria` (dark) or `shire` (light). |
|
|
| `LEMBAS_SESSION_TTL` | `2592000` | Session lifetime in seconds. |
|
|
| `LEMBAS_REQUEST_TIMEOUT` | `300` | Seconds to wait on an upstream model. |
|
|
|
|
### Commands
|
|
|
|
```bash
|
|
lembas serve # run the server
|
|
lembas info # where data lives, what is configured
|
|
lembas secret-key # generate a value for LEMBAS_SECRET_KEY
|
|
lembas create-admin # create or promote an administrator
|
|
```
|
|
|
|
## How it fits together
|
|
|
|
```
|
|
Browser ──form POST──▶ FastAPI ──▶ SQLite
|
|
▲ │
|
|
│ └──httpx──▶ any OpenAI-compatible endpoint
|
|
└──── server-sent events ◀───────────────┘ (streamed reply)
|
|
```
|
|
|
|
Sending a message stores the turn and returns two HTML fragments: the user's
|
|
bubble and an empty assistant bubble carrying an `sse-connect`. That opens a
|
|
server-sent event stream which appends tokens as they arrive, then replaces the
|
|
whole bubble with the finished, Markdown-rendered version. Rendering and
|
|
highlighting happen in Python, so the streamed and final views cannot disagree.
|
|
|
|
```
|
|
src/lembas/
|
|
api/ routes: auth, chats, folders, admin, pages
|
|
db/models/ SQLAlchemy schema
|
|
security/ password hashing, sessions
|
|
services/ llm client, chat orchestration, markdown, crypto, sse
|
|
web/ Jinja templates and static assets
|
|
assets/ SVG artwork masters
|
|
scripts/ artwork generator, vendored-JS fetcher
|
|
deploy/ systemd unit and nginx vhost for a real install
|
|
```
|
|
|
|
## Development
|
|
|
|
```bash
|
|
pytest # test suite
|
|
ruff check . # lint
|
|
python scripts/build_artwork.py # regenerate the SVG artwork
|
|
python scripts/fetch_vendor.py # verify vendored JS against the lockfile
|
|
```
|
|
|
|
There is no Alembic. The schema is SQLite-only and synchronised at startup:
|
|
missing tables and missing columns are added automatically, so adding a field to
|
|
a model needs nothing but a restart. Renames, drops and retypes are still manual
|
|
— see `CLAUDE.md`.
|
|
|
|
## Artwork
|
|
|
|
The logo, favicon and banner are original vector work, generated by
|
|
[`scripts/build_artwork.py`](scripts/build_artwork.py) so the mallorn leaf stays
|
|
identical across every size it appears at. The wordmark is
|
|
[Source Serif 4](https://github.com/adobe-fonts/source-serif) (SIL OFL 1.1)
|
|
converted to outlines — a README banner cannot load a webfont, and `<text>`
|
|
would render in whatever serif the reader happens to have.
|
|
|
|
## Licence
|
|
|
|
[GPL-3.0](LICENSE).
|
|
|
|
## A note on the theme
|
|
|
|
This is an independent hobby project, themed as an affectionate nod to
|
|
J.R.R. Tolkien's world. It is **not affiliated with, endorsed by, or connected
|
|
to** the Tolkien Estate, Middle-earth Enterprises, or any related rights
|
|
holder. All artwork here is original.
|