Files
LLeMbas/deploy/chat.lan.nginx.conf
T
Jaroslav Beneš dd9e0e9440 Working chat: auth, connections, streaming, folders
LLeMbas now runs end to end. Register, add an OpenAI-compatible
connection, and hold a real streaming conversation organised into
folders. Verified against the local llama-swap instance.

Streaming is the one genuinely tricky part. Sending a message returns
two HTML fragments -- the user bubble and an empty assistant bubble
carrying an sse-connect -- and that attribute is the ONLY thing that
starts a generation. Rendering an incomplete assistant message as a
streaming shell falls out of the same template, which means loading a
page whose last reply never finished simply picks it up again.

Details worth knowing about, each commented where it matters:

- SSE payloads are split across several data: lines. A raw newline in
  one data: line truncates the event, which shows up the first time a
  model emits a code block.
- Markdown is rendered server-side by the same helper for both the page
  and the final streamed frame, so the two cannot disagree. The fence
  renderer is replaced outright rather than using markdown-it's
  highlight option, which re-wraps output in a second <pre>.
- escape_text is html.escape, not nh3.clean_text: it escapes character
  by character, so escaping stream chunks separately equals escaping
  the whole string.
- The stream opens its own session via session_scope(); it outlives the
  request handler and the dependency-scoped session may be closed.
- Deleting a folder keeps the chats inside it (FK is SET NULL). Losing
  a conversation to a mis-clicked folder delete is unforgivable.
- Login failures use one message for "no such account" and "wrong
  password" so the form cannot enumerate registered addresses.

Also adds deploy/ for the gamebox install at https://chat.lan: system
unit, nginx vhost with buffering off (buffering on turns streaming into
one lump at the end), and install/update scripts following the same
service-user and /srv bind-mount conventions as llama-swap and comfyui.

70 tests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:04:13 +02:00

58 lines
1.9 KiB
Plaintext

# chat.lan - HTTPS reverse proxy to LLeMbas (127.0.0.1:8080).
# Deployed to /etc/nginx/conf.d/chat.lan.conf. Self-signed cert (chat.lan).
#
# Mirrors the comfy.lan and llama.lan vhosts on this box.
server {
listen 80;
listen [::]:80;
server_name chat.lan;
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
listen [::]:443 ssl;
http2 on;
server_name chat.lan;
ssl_certificate /etc/nginx/ssl/chat.lan.crt;
ssl_certificate_key /etc/nginx/ssl/chat.lan.key;
ssl_protocols TLSv1.2 TLSv1.3;
# File uploads land here once that feature exists; 0 = no limit.
client_max_body_size 0;
location / {
proxy_pass http://127.0.0.1:8080;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# Streamed replies are server-sent events. Every one of these matters:
# with buffering on, nginx holds the whole reply and delivers it in one
# lump at the end, which looks exactly like streaming being broken.
proxy_buffering off;
proxy_request_buffering off;
proxy_cache off;
# SSE is plain HTTP/1.1 chunked, so the connection header must not be
# the websocket upgrade dance -- it must simply stay open.
proxy_set_header Connection "";
# A model can think for minutes before the first token. The default
# 60s read timeout would cut long generations off mid-sentence.
proxy_read_timeout 3600s;
proxy_send_timeout 3600s;
}
# Static assets are immutable per release and never need revalidating.
location /static/ {
proxy_pass http://127.0.0.1:8080;
proxy_set_header Host $host;
expires 1h;
add_header Cache-Control "public";
}
}