Compaction: a button, and automatically when the window fills

A long conversation eventually just stops working. Compaction summarises the
earlier turns and sends the summary in their place.

The messages are kept. They stay in the transcript behind a collapsed
divider and simply stop being part of the request, which is what makes the
button safe to press and automatic compaction safe to have at all: a summary
that came out badly is a bad turn, not a lost conversation.

Stored on the Chat, not as a synthetic Message. A synthetic row needs a
role -- `system` breaks the one-system-message rule the moment build_messages
emits it beside the harness, and user/assistant makes it a turn people can
edit, regenerate from and copy, indistinguishable from a real one in all
four places a bubble is rendered. Worse, "editing rewinds, it does not
branch" would silently delete it and leave no marker that compaction had
happened at all.

The summary goes out as a user turn and an assistant turn, not one. A
leading assistant breaks templates requiring the first non-system message to
be user; a lone leading user produces user, user whenever the kept history
starts on a user turn -- which it always does, because the cutoff lands on a
finished reply.

compacted_through_id is a plain id rather than a foreign key: migrations.py
compiles only the column type, so a REFERENCES clause would exist on a fresh
database and not on an upgraded one, and a constraint half the fleet has is
worse than none. cutoff_message validates it on every read instead, and a
rewind past the boundary clears it.

Compacting again summarises only the delta, with the previous summary
supplied to be subsumed. Re-summarising the whole chat each time grows
quadratically and eventually exceeds the window it is protecting.

Automatically at the top of _run, not in post_message: that route's contract
is to return immediately and leave the slow part to a resumable connection,
and it also means build_request is called once, after compaction, with no
second assembly path. The trigger is the last reply's recorded usage plus an
estimate of the new turn -- retrospective because true prompt_tokens are only
knowable after a response, plus the delta because otherwise fifty thousand
characters pasted into the composer overflow a window that read 90% last
turn. It never fires when the context length is unknown. It does fire on
estimated counts, which is safe here precisely because nothing is lost.

_maybe_compact never raises: a failure logs and sends the uncompacted
request. A `status` event says "Summarising earlier messages…" in the
meantime, because a silent multi-second pause before the first token is what
a hang looks like.

The wording is three fragments under Admin - Prompts. Clearing task.compact
turns compaction off entirely.

Also adds compaction.moment(): SQLite does not store the offset, so a row
loaded from disk is naive while one in the session's identity map keeps its
tzinfo, and comparing the two raises. Every comparison here is between
exactly those.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jaroslav Beneš
2026-08-01 01:02:02 +02:00
parent 314cc946d7
commit 17f3fa1946
17 changed files with 1068 additions and 12 deletions
+60
View File
@@ -174,11 +174,31 @@ def build_messages(
resolved, which is how the harness gets in front of the authored prompt
without this function knowing anything about tools.
"""
from lembas.services import compaction as compaction_service
from lembas.services import prompts as prompts_service
payload: list[dict[str, Any]] = []
system = effective_system_prompt(db, chat) if system_prompt is None else system_prompt
if system:
payload.append({"role": ROLE_SYSTEM, "content": system})
# Compacted turns are replaced by a summary carried in two turns rather than
# one. A leading `assistant` breaks templates that require the first
# non-system message to be `user`; a lone leading `user` produces user, user
# whenever the kept history starts on a user turn -- which it always does,
# because the cutoff lands on a finished reply. The pair alternates
# correctly in both directions and keeps exactly one system message.
cutoff = compaction_service.cutoff_message(db, chat)
if cutoff is not None:
lead = prompts_service.resolve(db, "task.compact_lead").strip()
ack = prompts_service.resolve(db, "task.compact_ack").strip()
summary = chat.compact_summary.strip()
payload.append(
{"role": ROLE_USER, "content": f"{lead}\n\n{summary}" if lead else summary}
)
if ack:
payload.append({"role": ROLE_ASSISTANT, "content": ack})
history = db.scalars(
select(Message).where(Message.chat_id == chat.id).order_by(Message.created_at)
).all()
@@ -186,6 +206,10 @@ def build_messages(
for message in history:
if upto is not None and message.id == upto.id:
break
if cutoff is not None and compaction_service.moment(
message
) <= compaction_service.moment(cutoff):
continue
# Skip turns that failed or produced nothing -- but a message carrying
# only an attachment has no text and must still be sent.
if message.error:
@@ -391,6 +415,42 @@ def create_message(
return message
async def summarise_for_compaction(
endpoint: Endpoint,
model_id: str,
*,
transcript: str,
previous_summary: str,
template: str,
) -> str:
"""Ask the model to summarise the earlier turns.
`template` is passed in for the same reason `generate_title`'s is: this runs
after the generation's session has closed, and opening another one there is
how you get a session that outlives its scope. An empty template means an
administrator cleared the fragment, and nothing is asked of anyone.
"""
from lembas.services import prompts as prompts_service
if not template.strip() or not transcript.strip():
return ""
prompt = prompts_service.substitute(
template, {"transcript": transcript, "previous_summary": previous_summary}
)
raw = await complete(
endpoint,
{
"model": model_id,
"messages": [{"role": ROLE_USER, "content": prompt}],
"max_tokens": 1200,
# Low, but not zero: this is recall, not invention.
"temperature": 0.3,
},
)
return raw.strip()
def sweep_temporary(db: DBSession, older_than: timedelta = TEMPORARY_LIFETIME) -> int:
"""Delete temporary chats nobody has touched for a day.
+192
View File
@@ -0,0 +1,192 @@
"""Carrying a long conversation forward without carrying all of it.
Past a certain length every chat stops working: the window fills, and the only
options are to lose the beginning or to start again. Compaction summarises the
earlier turns and sends the summary in their place.
**The messages are kept.** They stay in the transcript, collapsed behind a
divider, and simply stop being part of the request. A summary that turned out
badly is then a bad turn rather than a lost conversation, which is what makes
the button safe to press and automatic compaction safe to have at all.
**Stored on the Chat, not as a synthetic Message.** A synthetic row would need a
role: `system` breaks the one-system-message rule the moment `build_messages`
emits it beside the harness, and `user`/`assistant` makes it a turn people can
edit, regenerate from and copy, indistinguishable from a real one in all four
places a bubble is rendered. Worse, "editing rewinds, it does not branch" would
silently delete it and leave no marker that compaction had ever happened.
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import ROLE_ASSISTANT, Chat, Message
from lembas.services import metrics as metrics_service
from lembas.services import settings_store, tokens
log = logging.getLogger(__name__)
# What the summariser is shown. Past this the oldest turns are dropped with a
# marker: a transcript that does not fit the window it is protecting is no use.
MAX_TRANSCRIPT_CHARS = 24_000
# Settings key, in the GENERAL group. 0 turns automatic compaction off; the
# button still works, because a person asking for it does not need a threshold.
THRESHOLD_KEY = "compact_threshold"
DEFAULT_THRESHOLD = 95
def threshold(db: DBSession) -> int:
value = settings_store.get(db, THRESHOLD_KEY)
return int(value) if isinstance(value, (int, float)) else DEFAULT_THRESHOLD
def moment(message: Message) -> datetime:
"""A message's timestamp, always comparable.
SQLite does not store the offset, so a row loaded from disk comes back naive
while one still in the session's identity map keeps the tzinfo it was
created with. Comparing the two raises, and every comparison here is between
exactly those: a cutoff fetched by id against history loaded in bulk.
`files.sweep_orphans` already normalises for the same reason.
"""
created = message.created_at
return created if created.tzinfo is not None else created.replace(tzinfo=UTC)
def cutoff_message(db: DBSession, chat: Chat) -> Message | None:
"""The message compaction reached, or None if it never has.
There is no foreign key to null this out on an upgraded database, so the
check is load-bearing rather than defensive: an id pointing at a message
that has been deleted means the boundary no longer describes anything, and
the chat has to read as uncompacted.
"""
if not chat.compact_summary or not chat.compacted_through_id:
return None
message = db.get(Message, chat.compacted_through_id)
if message is None or message.chat_id != chat.id:
return None
return message
def reset(chat: Chat) -> None:
"""Forget that this chat was ever compacted."""
chat.compact_summary = ""
chat.compacted_through_id = None
chat.compacted_at = None
def apply(chat: Chat, *, summary: str, upto: Message) -> None:
"""Record a summary and move the boundary. Caller commits."""
chat.compact_summary = summary.strip()
chat.compacted_through_id = upto.id
chat.compacted_at = datetime.now(UTC)
def split(
db: DBSession, chat: Chat, messages: list[Message]
) -> tuple[list[Message], list[Message]]:
"""(summarised, live) -- what is behind the divider, and what is not."""
cutoff = cutoff_message(db, chat)
if cutoff is None:
return [], list(messages)
boundary = moment(cutoff)
return (
[m for m in messages if moment(m) <= boundary],
[m for m in messages if moment(m) > boundary],
)
def last_complete(db: DBSession, chat: Chat) -> Message | None:
"""The newest finished assistant turn: where compaction should stop.
Landing on a reply rather than a question means the kept history starts on a
user turn, which is what every chat template expects.
"""
return db.scalar(
select(Message)
.where(
Message.chat_id == chat.id,
Message.role == ROLE_ASSISTANT,
Message.complete.is_(True),
)
.order_by(Message.created_at.desc())
)
def transcript(db: DBSession, chat: Chat, *, upto: Message) -> str:
"""The turns to summarise, oldest first, as plain text.
Only the delta since the last compaction: the previous summary is supplied
separately, and the instruction asks for one record, so each summary
subsumes the one before it. Re-summarising the whole chat every time grows
quadratically and eventually exceeds the very window this protects.
"""
previous = cutoff_message(db, chat)
query = select(Message).where(
Message.chat_id == chat.id,
Message.created_at <= upto.created_at,
Message.error == "",
)
if previous is not None:
query = query.where(Message.created_at > previous.created_at)
lines: list[str] = []
for message in db.scalars(query.order_by(Message.created_at)):
body = message.content.strip()
if not body:
continue
lines.append(f"{message.role}: {body}")
text = "\n\n".join(lines)
if len(text) > MAX_TRANSCRIPT_CHARS:
# Keep the most recent part: the older it is, the more likely the
# previous summary already covers it.
text = "[earlier turns omitted]\n\n" + text[-MAX_TRANSCRIPT_CHARS:]
return text
def previous_summary_block(chat: Chat) -> str:
"""The earlier summary, headed, or "" on a first compaction.
Empty is fine to pass straight through: `prompts.substitute` drops a line
that held a known variable and expanded to nothing, so the prompt does not
end up with a hole where a heading was.
"""
if not chat.compact_summary.strip():
return ""
return "## Summary of even earlier turns\n\n" + chat.compact_summary.strip()
def should_compact(db: DBSession, chat: Chat, *, pending: str = "") -> bool:
"""Whether the next request should be summarised first.
Judged from the last reply's recorded usage plus an estimate of the new
turn. True prompt_tokens are only knowable after a response, so a
retrospective figure is the honest basis -- but on its own it is one turn
stale, and fifty thousand characters pasted into the composer would overflow
a window that measured 90% last time. The estimator covers only that delta.
Never fires when the model's context length is unknown. Acting on a number
nobody supplied is exactly what the 0-means-unknown rule exists to prevent.
"""
limit = threshold(db)
if limit <= 0:
return False
last = last_complete(db, chat)
if last is None:
return False
usage = metrics_service.from_message(last.usage_json)
if usage.context_limit <= 0 or usage.context_tokens <= 0:
return False
projected = usage.context_tokens + tokens.estimate(pending)
return projected >= usage.context_limit * limit / 100
+85
View File
@@ -22,9 +22,12 @@ import time
from dataclasses import dataclass, field
from datetime import UTC, datetime, timedelta
from sqlalchemy import select
from lembas.db.models import ROLE_ASSISTANT, ROLE_USER, Chat, Message, User
from lembas.db.session import session_scope
from lembas.services import chat as chat_service
from lembas.services import compaction as compaction_service
from lembas.services import metrics as metrics_service
from lembas.services import prompts as prompts_service
from lembas.services import tokens
@@ -95,6 +98,10 @@ class Generation:
# woken individually: with a 100ms cadence a short poll is simpler than
# future bookkeeping, and cannot drop a wakeup.
version: int = 0
# What the reply is doing when it is not producing tokens. Shown in the
# streaming bubble, because a silent multi-second pause before the first
# token is what a hang looks like.
status: str = ""
# Number of browsers currently watching. Decides whether a finished reply
# counts as unread.
followers: int = 0
@@ -214,6 +221,13 @@ async def _run(generation: Generation) -> None:
title_prompt = ""
try:
# Before the request is assembled, so build_request is called once and
# what goes out is the compacted conversation -- there is no second
# assembly path. Here rather than in post_message because that route's
# whole contract is to return immediately, and a three-second
# summarisation in front of it would break exactly that.
await _maybe_compact(generation)
with session_scope() as db:
chat = db.get(Chat, generation.chat_id)
message = db.get(Message, generation.message_id)
@@ -387,6 +401,77 @@ async def _run(generation: Generation) -> None:
generation.touch()
async def _maybe_compact(generation: Generation) -> None:
"""Summarise the earlier turns if the window is about to be full.
Never raises. A failed compaction logs and sends the uncompacted request,
which either works or fails upstream with a message that says what actually
happened -- refusing to answer because the summariser was unavailable would
be a worse trade.
The awaited call is deliberately outside any session, the same shape titling
uses: read everything needed, close, ask, reopen to write.
"""
try:
with session_scope() as db:
chat = db.get(Chat, generation.chat_id)
message = db.get(Message, generation.message_id)
if chat is None or message is None:
return
pending = _pending_text(db, message)
if not compaction_service.should_compact(db, chat, pending=pending):
return
template = prompts_service.resolve(db, "task.compact")
upto = compaction_service.last_complete(db, chat)
if not template.strip() or upto is None:
return
endpoint, model_id = chat_service.resolve_endpoint(db, chat)
transcript = compaction_service.transcript(db, chat, upto=upto)
previous = compaction_service.previous_summary_block(chat)
upto_id = upto.id
generation.status = "Summarising earlier messages…"
generation.touch()
summary = await chat_service.summarise_for_compaction(
endpoint,
model_id,
transcript=transcript,
previous_summary=previous,
template=template,
)
if not summary:
return
with session_scope() as db:
chat = db.get(Chat, generation.chat_id)
upto = db.get(Message, upto_id)
if chat is None or upto is None:
return
compaction_service.apply(chat, summary=summary, upto=upto)
db.commit()
log.info("chat %s compacted automatically through %s", chat.id, upto_id)
except Exception: # noqa: BLE001 - the reply matters more than the tidy-up
log.exception("automatic compaction failed for chat %s", generation.chat_id)
finally:
generation.status = ""
generation.touch()
def _pending_text(db, message: Message) -> str:
"""The user turn this reply is answering, for the size estimate."""
previous = db.scalars(
select(Message)
.where(Message.chat_id == message.chat_id, Message.created_at < message.created_at)
.order_by(Message.created_at.desc())
.limit(1)
).first()
return previous.content if previous is not None else ""
def _question_from(payload: dict) -> str:
"""The last thing the user said, for auto-titling."""
for entry in reversed(payload.get("messages", [])):
+81
View File
@@ -156,6 +156,16 @@ VARIABLES: tuple[Variable, ...] = (
),
Variable("question", "Question", "The first message. Chat title task only."),
Variable("answer", "Answer", "The first reply. Chat title task only."),
Variable(
"transcript",
"Transcript",
"The turns being summarised, oldest first. Compaction task only.",
),
Variable(
"previous_summary",
"Earlier summary",
"The summary from a previous compaction, if there was one. Compaction task only.",
),
)
VARIABLE_NAMES = frozenset(variable.name for variable in VARIABLES)
@@ -764,6 +774,77 @@ BUILTIN: tuple[Fragment, ...] = (
"Assistant: {{answer}}"
),
),
Fragment(
key="task.compact",
label="Compaction summary",
group=GROUP_TASKS,
order=410,
variables=("transcript", "previous_summary"),
hint="A separate one-message request, not part of any chat. Clear it to "
"turn compaction off entirely: the button says so and nothing is "
"summarised automatically.",
default=(
"Summarise the conversation below so it can be carried forward after the "
"earlier turns are dropped from your context. This is a working record, "
"not a report for a reader.\n"
"\n"
"Keep, under these headings and in this order:\n"
"\n"
"## What we are doing\n"
"The goal, and where we have got to.\n"
"\n"
"## Decisions\n"
"Anything settled, and why. A decision without its reason gets argued "
"again.\n"
"\n"
"## Facts established\n"
"Names, numbers, versions, file paths, URLs and identifiers, copied "
"exactly. Do not round them, paraphrase them or reconstruct one from "
"memory — if it is not in the transcript, leave it out.\n"
"\n"
"## Open threads\n"
"What is unfinished, and what was about to happen next.\n"
"\n"
"Leave out pleasantries, retracted ideas and anything already superseded. "
"Do not answer the conversation: you are recording it. Write in the "
"language of the conversation, and stay under 500 words.\n"
"\n"
"{{previous_summary}}\n"
"\n"
"## Transcript\n"
"\n"
"{{transcript}}"
),
),
Fragment(
key="task.compact_lead",
label="How a summary is introduced",
group=GROUP_TASKS,
order=420,
hint="Sits in front of the summary, in the turn that replaces the "
"messages no longer being sent. Without it a model reads the summary as "
"something the person has just typed.",
default=(
"Here is a summary of the earlier part of this conversation. Those "
"messages are no longer in your context. Treat this summary as an "
"accurate record of them and rely on it rather than on what you can no "
"longer see; if it does not cover something you need, say so instead of "
"filling the gap."
),
),
Fragment(
key="task.compact_ack",
label="The model's acknowledgement",
group=GROUP_TASKS,
order=430,
hint="One assistant turn after the summary, so the conversation still "
"alternates user, assistant, user. Several chat templates reject a "
"history that does not.",
default=(
"Understood. I have the summary of the earlier turns and will carry on "
"from there."
),
),
)
register_source(_builtin_source)
+6
View File
@@ -35,6 +35,12 @@ def _general_defaults() -> dict[str, Any]:
# Applied to every chat that has no model or chat prompt of its
# own. See services.chat.effective_system_prompt.
"system_prompt": "",
# Percentage of a model's context length at which the earlier turns are
# summarised automatically. 0 turns it off; the Compact button still
# works, because a person asking for it does not need a threshold.
# Never fires for a model whose context_length is 0, since that is
# "unknown" rather than "small". See services/compaction.py.
"compact_threshold": 95,
}