dd9e0e9440
LLeMbas now runs end to end. Register, add an OpenAI-compatible connection, and hold a real streaming conversation organised into folders. Verified against the local llama-swap instance. Streaming is the one genuinely tricky part. Sending a message returns two HTML fragments -- the user bubble and an empty assistant bubble carrying an sse-connect -- and that attribute is the ONLY thing that starts a generation. Rendering an incomplete assistant message as a streaming shell falls out of the same template, which means loading a page whose last reply never finished simply picks it up again. Details worth knowing about, each commented where it matters: - SSE payloads are split across several data: lines. A raw newline in one data: line truncates the event, which shows up the first time a model emits a code block. - Markdown is rendered server-side by the same helper for both the page and the final streamed frame, so the two cannot disagree. The fence renderer is replaced outright rather than using markdown-it's highlight option, which re-wraps output in a second <pre>. - escape_text is html.escape, not nh3.clean_text: it escapes character by character, so escaping stream chunks separately equals escaping the whole string. - The stream opens its own session via session_scope(); it outlives the request handler and the dependency-scoped session may be closed. - Deleting a folder keeps the chats inside it (FK is SET NULL). Losing a conversation to a mis-clicked folder delete is unforgivable. - Login failures use one message for "no such account" and "wrong password" so the form cannot enumerate registered addresses. Also adds deploy/ for the gamebox install at https://chat.lan: system unit, nginx vhost with buffering off (buffering on turns streaming into one lump at the end), and install/update scripts following the same service-user and /srv bind-mount conventions as llama-swap and comfyui. 70 tests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
51 lines
1.6 KiB
Bash
Executable File
51 lines
1.6 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Pull the latest LLeMbas and restart the service.
|
|
#
|
|
# This is what to run after pushing: it fetches, hard-resets the deployment
|
|
# checkout to the remote branch, reinstalls dependencies if they changed, and
|
|
# restarts. Nothing is ever edited in place at /srv/lembas/app, so a hard reset
|
|
# is safe and avoids merge conflicts from a dirty deployment tree.
|
|
set -euo pipefail
|
|
|
|
SERVICE_USER=lembas
|
|
PREFIX=/srv/lembas
|
|
APP="$PREFIX/app"
|
|
VENV="$PREFIX/venv"
|
|
BRANCH="${LEMBAS_BRANCH:-main}"
|
|
|
|
if [[ ! -d "$APP/.git" ]]; then
|
|
echo "No deployment at $APP. Run deploy/install.sh first." >&2
|
|
exit 1
|
|
fi
|
|
|
|
before=$(sudo -u "$SERVICE_USER" git -C "$APP" rev-parse HEAD)
|
|
|
|
echo "== fetching =="
|
|
sudo -u "$SERVICE_USER" git -C "$APP" fetch --quiet origin "$BRANCH"
|
|
sudo -u "$SERVICE_USER" git -C "$APP" reset --hard --quiet "origin/$BRANCH"
|
|
|
|
after=$(sudo -u "$SERVICE_USER" git -C "$APP" rev-parse HEAD)
|
|
|
|
if [[ "$before" == "$after" ]]; then
|
|
echo " already at $(git -C "$APP" rev-parse --short HEAD), nothing to pull"
|
|
else
|
|
echo " $(echo "$before" | cut -c1-7) -> $(echo "$after" | cut -c1-7)"
|
|
sudo -u "$SERVICE_USER" git -C "$APP" --no-pager log --oneline "$before..$after" | sed 's/^/ /'
|
|
fi
|
|
|
|
# Cheap and idempotent; catches a dependency added since the last deploy.
|
|
echo "== dependencies =="
|
|
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet -e "$APP"
|
|
|
|
echo "== restart =="
|
|
sudo systemctl restart lembas
|
|
sleep 2
|
|
|
|
if systemctl is-active --quiet lembas; then
|
|
echo " lembas is running at https://chat.lan"
|
|
else
|
|
echo " lembas FAILED to start:" >&2
|
|
sudo journalctl -u lembas -n 30 --no-pager >&2
|
|
exit 1
|
|
fi
|