Files
LLeMbas/README.md
T
Jaroslav Beneš 5f020ef33f Live Markdown, stop, rewind, custom picker, dialogs
Seven things.

**Reasoning starts closed.** The answer is what the reader is waiting
for; the thinking is one click away.

**Image borders.** .attachments__image was a block-level <a>, so its
border stretched the full column around a narrow picture. inline-block,
and the frame is the picture. Same fix for the composer thumbnail.

**Markdown now renders during the stream.** The generator re-renders the
answer so far and sends it as a `render` event at most every 100ms,
swapped with innerHTML, instead of appending escaped tokens and
formatting everything at the end. Re-rendering whole rather than
appending is the point: a list or a code fence is only correct once its
context exists, and partial syntax resolves itself as more arrives.
Measured against a live model: 29 render events, formatting visible from
the first content token.

**Stop button.** A stop request goes into an in-process set the
generator checks between chunks; whatever arrived is kept, because a
half-written answer the reader chose to cut short is still worth having.
Measured: stream ended 0.2s after the request, 1155 characters
preserved, message marked stopped rather than errored. Navigating away
does the same thing via CancelledError.

**Rewind and edit.** Edit one of your own turns and everything after it
is deleted, then the conversation runs on from there. Deliberately not
branching: that needs a UI for choosing between versions, and "go back
and try again from here" is what was asked for. The form states how many
messages will be discarded before you confirm.

**Custom model picker.** A <select> renders only text in an <option>, so
it can never show an avatar. Built from buttons and a hidden input, with
descriptions, capability tags, a filter box past eight models, and
arrow-key navigation written out by hand since there is no native widget
doing it.

**Notification system.** lembas.notify/confirm/prompt in ui.js, built on
<dialog> so focus trapping, Escape and page inertness come from the
browser. htmx:confirm is intercepted, so every existing hx-confirm gets
the themed dialog with no change at the call site; the browser's grey
confirm() is gone from every template.

230 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 14:33:04 +02:00

176 lines
7.4 KiB
Markdown

<p align="center">
<img src="assets/banner.svg" alt="LLeMbas — waybread for the long road of thought" width="100%">
</p>
<p align="center">
<strong>A self-hosted web UI for your language models, written in Python.</strong><br>
Talks to anything that speaks the OpenAI API. Themed after Middle-earth.
</p>
<p align="center">
<img alt="Python 3.11+" src="https://img.shields.io/badge/python-3.11%2B-3E6B7A?style=flat-square">
<img alt="License GPL-3.0" src="https://img.shields.io/badge/license-GPL--3.0-C9A227?style=flat-square">
<img alt="No Node required" src="https://img.shields.io/badge/build%20step-none-6B8E4E?style=flat-square">
</p>
---
*Lembas* is the Elvish waybread — one bite sustains a traveller for a day's
march. The capitals hide what it runs on: **LLeM**bas.
## Why this exists
Most self-hosted LLM front-ends are large JavaScript applications with a Python
API bolted underneath. LLeMbas is the other way round: **server-rendered
Python**, with htmx and a little Alpine for interactivity. There is no
`package.json`, no bundler, no build step, and nothing is fetched from a CDN at
runtime. Clone it, `pip install -e .`, run it.
## Features
**Working now**
- **Chats** — streaming replies, Markdown with server-side syntax highlighting,
copy and regenerate, automatic chat titles. Chats are created when you send
the first message, so an abandoned one never clutters the sidebar
- **System prompts** — instance-wide, per-model and per-chat, with the most
specific winning outright
- **Reasoning display** — thinking streams into its own collapsible block
(closed by default), labelled with how long it took, and is never replayed as
context
- **Live Markdown** — formatting appears as the model writes, not at the end
- **Stop and rewind** — cut a reply short and keep what arrived, or edit an
earlier message and run the conversation on from there
- **Attachments** — drag, paste or pick images, PDFs and text files. Images are
downscaled and sent to vision models; PDF and text content is extracted and
put in the prompt
- **Folders** — arbitrarily nested, delete a folder without losing the chats
inside it
- **OpenAI connections** — point at OpenAI, LM Studio, vLLM, llama.cpp,
llama-swap, Ollama or OpenRouter; models are discovered and cached
- **Model settings** — searchable, filterable list with a page per model:
ordering, pinned models, an instance default and a per-user default, custom
names, descriptions and images. Scales to hundreds of models
- **Users, groups & permissions** — per-group grants that union rather than
override, and model access restricted to chosen groups
- **Accounts** — first account becomes the administrator, argon2 password
hashing, revocable server-side sessions, self-service password change,
admin-managed accounts
- **Admin settings** — open or close registration from the UI, stored in the
database and effective immediately
- **Two themes** — *Moria* (dark) and *Shire* (light), switchable per user
**Planned**
Built-in tools with admin settings · custom tools and MCP servers · agentic
execution (local and over SSH) · image generation · OCR for scanned PDFs.
## Quick start
```bash
git clone https://git.houmeres.sk/Houmeres/LLeMbas.git
cd LLeMbas
python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"
cp .env.example .env
lembas secret-key # paste the result into LEMBAS_SECRET_KEY
lembas serve # http://127.0.0.1:8080
```
Open the address and create the first account — it becomes the administrator.
Then go to **Admin → Connections** and add an endpoint. For a local runner that
is usually `http://localhost:1234/v1` with no API key. Press **Test & refresh**
and its models appear in the chat model picker.
> The vendored browser libraries (htmx, Alpine) are committed, so no network
> access is needed to run. To re-fetch or bump them:
> `python scripts/fetch_vendor.py --update`.
## Configuration
All variables are prefixed `LEMBAS_` and can live in `.env`. See
[`.env.example`](.env.example) for the annotated list.
| Variable | Default | Purpose |
|---|---|---|
| `LEMBAS_SECRET_KEY` | *generated* | Signs sessions and encrypts stored API keys. **Set this.** A generated key changes every restart, signing everyone out and making stored API keys unreadable. |
| `LEMBAS_DATA_DIR` | `./data` | SQLite database and uploads. |
| `LEMBAS_HOST` / `LEMBAS_PORT` | `127.0.0.1` / `8080` | Bind address. |
| `LEMBAS_ALLOW_SIGNUP` | `true` | Whether new users may register themselves — the *initial* value only. Once set under **Admin → General** the stored setting wins. The first account is always an admin regardless. |
| `LEMBAS_DEFAULT_THEME` | `moria` | `moria` (dark) or `shire` (light). |
| `LEMBAS_SESSION_TTL` | `2592000` | Session lifetime in seconds. |
| `LEMBAS_REQUEST_TIMEOUT` | `300` | Seconds to wait on an upstream model. |
### Commands
```bash
lembas serve # run the server
lembas info # where data lives, what is configured
lembas secret-key # generate a value for LEMBAS_SECRET_KEY
lembas create-admin # create or promote an administrator
```
## How it fits together
```
Browser ──form POST──▶ FastAPI ──▶ SQLite
▲ │
│ └──httpx──▶ any OpenAI-compatible endpoint
└──── server-sent events ◀───────────────┘ (streamed reply)
```
Sending a message stores the turn and returns two HTML fragments: the user's
bubble and an empty assistant bubble carrying an `sse-connect`. That opens a
server-sent event stream which appends tokens as they arrive, then replaces the
whole bubble with the finished, Markdown-rendered version. Rendering and
highlighting happen in Python, so the streamed and final views cannot disagree.
```
src/lembas/
api/ routes: auth, chats, folders, admin, pages
db/models/ SQLAlchemy schema
security/ password hashing, sessions
services/ llm client, chat orchestration, markdown, crypto, sse
web/ Jinja templates and static assets
assets/ SVG artwork masters
scripts/ artwork generator, vendored-JS fetcher
deploy/ systemd unit and nginx vhost for a real install
```
## Development
```bash
pytest # test suite
ruff check . # lint
python scripts/build_artwork.py # regenerate the SVG artwork
python scripts/fetch_vendor.py # verify vendored JS against the lockfile
```
There is no Alembic. The schema is SQLite-only and synchronised at startup:
missing tables and missing columns are added automatically, so adding a field to
a model needs nothing but a restart. Renames, drops and retypes are still manual
— see `CLAUDE.md`.
## Artwork
The logo, favicon and banner are original vector work, generated by
[`scripts/build_artwork.py`](scripts/build_artwork.py) so the mallorn leaf stays
identical across every size it appears at. The wordmark is
[Source Serif 4](https://github.com/adobe-fonts/source-serif) (SIL OFL 1.1)
converted to outlines — a README banner cannot load a webfont, and `<text>`
would render in whatever serif the reader happens to have.
## Licence
[GPL-3.0](LICENSE).
## A note on the theme
This is an independent hobby project, themed as an affectionate nod to
J.R.R. Tolkien's world. It is **not affiliated with, endorsed by, or connected
to** the Tolkien Estate, Middle-earth Enterprises, or any related rights
holder. All artwork here is original.