Jaroslav Beneš ca3e4fd04f Background generation, unread replies, send/stop, PLAN.md
**Replies now run in the background.** Generation was driven by the SSE
request, so navigating away or opening another chat cut the answer off
mid-sentence. services/generation.py owns the work as its own task and
the SSE endpoint merely follows it. Verified: attached briefly, closed
the connection, went to another page -- the reply finished anyway, 832
characters, not marked stopped, auto-titled.

Reattaching works because both `render` and `reasoning` frames now carry
the whole block rather than a delta. A follower arriving late has no
earlier fragments to append to, so deltas would leave it permanently
missing the beginning. Verified: attached six seconds in and the first
frame already contained 517 characters written while nobody watched.

**Unread indicator.** A reply that lands with no follower attached marks
its chat unread; the sidebar polls every 10s for out-of-band dot spans
plus an HX-Trigger that raises a toast. Polled rather than pushed: a
browser sitting on another chat has no connection to the one that
finished, and an always-on channel per tab is a lot of machinery for a
green dot. `unread_notified` stops the same arrival being announced
every tick. Follower count is what decides "was anyone watching", so
reading it as it arrives does not mark it unread -- verified both ways.

**Stop is the send button.** While a reply is being written the send
button becomes a red stop square, found via a MutationObserver on the
thread since the composer and the streaming bubble are far apart in the
document. The in-bubble Stop is gone.

**Attachment border removed.** As asked -- an attachment is a picture,
and the frame only ever drew at the wrong width. The anchor now
shrink-wraps and the img's width/height attributes are overridden so a
small image shows at its own size.

Adds PLAN.md: what is built, what is not, known limits, and the
decisions that look like oversights until you know the reason.

239 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 14:52:28 +02:00
2026-07-21 08:02:48 +00:00

LLeMbas — waybread for the long road of thought

A self-hosted web UI for your language models, written in Python.
Talks to anything that speaks the OpenAI API. Themed after Middle-earth.

Python 3.11+ License GPL-3.0 No Node required


Lembas is the Elvish waybread — one bite sustains a traveller for a day's march. The capitals hide what it runs on: LLeMbas.

Why this exists

Most self-hosted LLM front-ends are large JavaScript applications with a Python API bolted underneath. LLeMbas is the other way round: server-rendered Python, with htmx and a little Alpine for interactivity. There is no package.json, no bundler, no build step, and nothing is fetched from a CDN at runtime. Clone it, pip install -e ., run it.

Features

Working now

  • Chats — streaming replies, Markdown with server-side syntax highlighting, copy and regenerate, automatic chat titles. Chats are created when you send the first message, so an abandoned one never clutters the sidebar
  • System prompts — instance-wide, per-model and per-chat, with the most specific winning outright
  • Reasoning display — thinking streams into its own collapsible block (closed by default), labelled with how long it took, and is never replayed as context
  • Live Markdown — formatting appears as the model writes, not at the end
  • Stop and rewind — cut a reply short and keep what arrived, or edit an earlier message and run the conversation on from there
  • Replies keep running in the background — navigate away, open another chat, close the tab; a green dot and a notification tell you when it lands
  • Attachments — drag, paste or pick images, PDFs and text files. Images are downscaled and sent to vision models; PDF and text content is extracted and put in the prompt
  • Folders — arbitrarily nested, delete a folder without losing the chats inside it
  • OpenAI connections — point at OpenAI, LM Studio, vLLM, llama.cpp, llama-swap, Ollama or OpenRouter; models are discovered and cached
  • Model settings — searchable, filterable list with a page per model: ordering, pinned models, an instance default and a per-user default, custom names, descriptions and images. Scales to hundreds of models
  • Users, groups & permissions — per-group grants that union rather than override, and model access restricted to chosen groups
  • Accounts — first account becomes the administrator, argon2 password hashing, revocable server-side sessions, self-service password change, admin-managed accounts
  • Admin settings — open or close registration from the UI, stored in the database and effective immediately
  • Two themesMoria (dark) and Shire (light), switchable per user

Planned

Built-in tools with admin settings · custom tools and MCP servers · agentic execution (local and over SSH) · image generation · OCR for scanned PDFs.

See PLAN.md for what is built, what is not, and why.

Quick start

git clone https://git.houmeres.sk/Houmeres/LLeMbas.git
cd LLeMbas

python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"

cp .env.example .env
lembas secret-key           # paste the result into LEMBAS_SECRET_KEY

lembas serve                # http://127.0.0.1:8080

Open the address and create the first account — it becomes the administrator. Then go to Admin → Connections and add an endpoint. For a local runner that is usually http://localhost:1234/v1 with no API key. Press Test & refresh and its models appear in the chat model picker.

The vendored browser libraries (htmx, Alpine) are committed, so no network access is needed to run. To re-fetch or bump them: python scripts/fetch_vendor.py --update.

Configuration

All variables are prefixed LEMBAS_ and can live in .env. See .env.example for the annotated list.

Variable Default Purpose
LEMBAS_SECRET_KEY generated Signs sessions and encrypts stored API keys. Set this. A generated key changes every restart, signing everyone out and making stored API keys unreadable.
LEMBAS_DATA_DIR ./data SQLite database and uploads.
LEMBAS_HOST / LEMBAS_PORT 127.0.0.1 / 8080 Bind address.
LEMBAS_ALLOW_SIGNUP true Whether new users may register themselves — the initial value only. Once set under Admin → General the stored setting wins. The first account is always an admin regardless.
LEMBAS_DEFAULT_THEME moria moria (dark) or shire (light).
LEMBAS_SESSION_TTL 2592000 Session lifetime in seconds.
LEMBAS_REQUEST_TIMEOUT 300 Seconds to wait on an upstream model.

Commands

lembas serve          # run the server
lembas info           # where data lives, what is configured
lembas secret-key     # generate a value for LEMBAS_SECRET_KEY
lembas create-admin   # create or promote an administrator

How it fits together

Browser  ──form POST──▶  FastAPI  ──▶  SQLite
   ▲                        │
   │                        └──httpx──▶  any OpenAI-compatible endpoint
   └──── server-sent events ◀───────────────┘   (streamed reply)

Sending a message stores the turn and returns two HTML fragments: the user's bubble and an empty assistant bubble carrying an sse-connect. That opens a server-sent event stream which appends tokens as they arrive, then replaces the whole bubble with the finished, Markdown-rendered version. Rendering and highlighting happen in Python, so the streamed and final views cannot disagree.

src/lembas/
  api/         routes: auth, chats, folders, admin, pages
  db/models/   SQLAlchemy schema
  security/    password hashing, sessions
  services/    llm client, chat orchestration, markdown, crypto, sse
  web/         Jinja templates and static assets
assets/        SVG artwork masters
scripts/       artwork generator, vendored-JS fetcher
deploy/        systemd unit and nginx vhost for a real install

Development

pytest                              # test suite
ruff check .                        # lint
python scripts/build_artwork.py     # regenerate the SVG artwork
python scripts/fetch_vendor.py      # verify vendored JS against the lockfile

There is no Alembic. The schema is SQLite-only and synchronised at startup: missing tables and missing columns are added automatically, so adding a field to a model needs nothing but a restart. Renames, drops and retypes are still manual — see CLAUDE.md.

Artwork

The logo, favicon and banner are original vector work, generated by scripts/build_artwork.py so the mallorn leaf stays identical across every size it appears at. The wordmark is Source Serif 4 (SIL OFL 1.1) converted to outlines — a README banner cannot load a webfont, and <text> would render in whatever serif the reader happens to have.

Licence

GPL-3.0.

A note on the theme

This is an independent hobby project, themed as an affectionate nod to J.R.R. Tolkien's world. It is not affiliated with, endorsed by, or connected to the Tolkien Estate, Middle-earth Enterprises, or any related rights holder. All artwork here is original.

S
Description
No description provided
Readme GPL-3.0 13 MiB
1.0.2 Latest
2026-08-07 23:38:39 +00:00
Languages
Python 79.1%
HTML 12.4%
JavaScript 4.5%
CSS 3.2%
Shell 0.7%