dd9e0e9440
LLeMbas now runs end to end. Register, add an OpenAI-compatible connection, and hold a real streaming conversation organised into folders. Verified against the local llama-swap instance. Streaming is the one genuinely tricky part. Sending a message returns two HTML fragments -- the user bubble and an empty assistant bubble carrying an sse-connect -- and that attribute is the ONLY thing that starts a generation. Rendering an incomplete assistant message as a streaming shell falls out of the same template, which means loading a page whose last reply never finished simply picks it up again. Details worth knowing about, each commented where it matters: - SSE payloads are split across several data: lines. A raw newline in one data: line truncates the event, which shows up the first time a model emits a code block. - Markdown is rendered server-side by the same helper for both the page and the final streamed frame, so the two cannot disagree. The fence renderer is replaced outright rather than using markdown-it's highlight option, which re-wraps output in a second <pre>. - escape_text is html.escape, not nh3.clean_text: it escapes character by character, so escaping stream chunks separately equals escaping the whole string. - The stream opens its own session via session_scope(); it outlives the request handler and the dependency-scoped session may be closed. - Deleting a folder keeps the chats inside it (FK is SET NULL). Losing a conversation to a mis-clicked folder delete is unforgivable. - Login failures use one message for "no such account" and "wrong password" so the form cannot enumerate registered addresses. Also adds deploy/ for the gamebox install at https://chat.lan: system unit, nginx vhost with buffering off (buffering on turns streaming into one lump at the end), and install/update scripts following the same service-user and /srv bind-mount conventions as llama-swap and comfyui. 70 tests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
50 lines
1.7 KiB
Desktop File
50 lines
1.7 KiB
Desktop File
# LLeMbas system service.
|
|
#
|
|
# Deployed to /etc/systemd/system/lembas.service by deploy/install.sh.
|
|
#
|
|
# A system unit, not a user unit, so it survives logout and comes up at boot
|
|
# without anyone signing in -- matching llama-swap and comfyui on this box.
|
|
#
|
|
# /srv/lembas is a bind mount of /home/lembas: the root LV is only 50 GB and
|
|
# the venv plus SQLite database belong on /home, same trick as /srv/llama.
|
|
|
|
[Unit]
|
|
Description=LLeMbas - web UI for language models
|
|
Documentation=https://git.houmeres.sk/Houmeres/LLeMbas
|
|
After=network-online.target
|
|
Wants=network-online.target
|
|
# The unit is useless without the bind mount: the venv and database live there.
|
|
RequiresMountsFor=/srv/lembas
|
|
# Not a hard dependency. LLeMbas starts fine with the endpoint down and shows a
|
|
# readable error in the admin UI, which is better than refusing to boot.
|
|
After=llama-swap.service
|
|
|
|
[Service]
|
|
Type=simple
|
|
User=lembas
|
|
Group=lembas
|
|
WorkingDirectory=/srv/lembas/app
|
|
EnvironmentFile=/srv/lembas/lembas.env
|
|
ExecStart=/srv/lembas/venv/bin/lembas serve
|
|
Restart=on-failure
|
|
RestartSec=5
|
|
|
|
# Reachable only through the nginx chat.lan vhost, never directly on the LAN.
|
|
# The bind address is set by LEMBAS_HOST in the environment file.
|
|
|
|
# --- Hardening -------------------------------------------------------------
|
|
# Modest rather than maximal: the agentic features planned for later will need
|
|
# to run commands, so ProtectSystem=strict would only be torn out again.
|
|
NoNewPrivileges=yes
|
|
PrivateTmp=yes
|
|
ProtectSystem=full
|
|
ProtectKernelTunables=yes
|
|
ProtectControlGroups=yes
|
|
RestrictSUIDSGID=yes
|
|
# Only /srv/lembas needs to be writable; /home/lembas is the same inode.
|
|
ReadWritePaths=/srv/lembas
|
|
LimitNOFILE=65535
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|