Files
LLeMbas/src/lembas/services/sse.py
T
Jaroslav Beneš 0f44e8d24c Working chat: auth, connections, streaming, folders
LLeMbas now runs end to end. Register, add an OpenAI-compatible
connection, and hold a real streaming conversation organised into
folders. Verified against the local llama-swap instance.

Streaming is the one genuinely tricky part. Sending a message returns
two HTML fragments -- the user bubble and an empty assistant bubble
carrying an sse-connect -- and that attribute is the ONLY thing that
starts a generation. Rendering an incomplete assistant message as a
streaming shell falls out of the same template, which means loading a
page whose last reply never finished simply picks it up again.

Details worth knowing about, each commented where it matters:

- SSE payloads are split across several data: lines. A raw newline in
  one data: line truncates the event, which shows up the first time a
  model emits a code block.
- Markdown is rendered server-side by the same helper for both the page
  and the final streamed frame, so the two cannot disagree. The fence
  renderer is replaced outright rather than using markdown-it's
  highlight option, which re-wraps output in a second <pre>.
- escape_text is html.escape, not nh3.clean_text: it escapes character
  by character, so escaping stream chunks separately equals escaping
  the whole string.
- The stream opens its own session via session_scope(); it outlives the
  request handler and the dependency-scoped session may be closed.
- Deleting a folder keeps the chats inside it (FK is SET NULL). Losing
  a conversation to a mis-clicked folder delete is unforgivable.
- Login failures use one message for "no such account" and "wrong
  password" so the form cannot enumerate registered addresses.

Also adds deploy/ for the gamebox install at https://chat.lan: system
unit, nginx vhost with buffering off (buffering on turns streaming into
one lump at the end), and install/update scripts following the same
service-user and /srv bind-mount conventions as llama-swap and comfyui.

70 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:04:13 +02:00

25 lines
864 B
Python

"""Server-sent event framing.
Small, but worth isolating: getting the wire format subtly wrong is the usual
cause of a stream that "works" until a model emits a newline.
"""
from __future__ import annotations
# Every 15s of silence, so proxies that kill idle connections (nginx defaults
# to 60s) do not drop a stream while a model is still thinking.
KEEPALIVE = ": keepalive\n\n"
def event(name: str, data: str) -> str:
"""Frame one SSE event.
A payload containing newlines must be split across several `data:` lines;
the browser rejoins them with "\\n". Sending a raw newline inside a single
data line silently truncates the event, which is exactly what happens the
first time a model emits a code block.
"""
lines = data.split("\n")
body = "".join(f"data: {line}\n" for line in lines)
return f"event: {name}\n{body}\n"