Narrow a chat before it starts, and find a file rather than spell it

Six things, all found by using the thing rather than by reading it.

The scope menu only appeared once a chat existed, on the reasoning that there was
no row to post to. True, and the wrong conclusion: the harness puts a tool's
guidance in front of the model the moment the tool is offered, so the menu could
not be reached until after the model had been told how to keep notes and handed
the tools to do it -- and switching it off then does not un-send that turn. It is
on the new-chat screen now and writes nothing: `_scope_context` builds a stand-in
Chat, which is `draft.as_chat`'s trick again, and the switches ride along with
the first message. Checked means on and a browser submits only the ticked boxes,
so every gate also renders a hidden input naming it and `start_chat` subtracts one
list from the other; inverting the control would read backwards under a menu that
says everything is on unless you say otherwise. Only the off ones are written,
because absent means on and one representation of it is what keeps "why is this
off?" to a single answer. Nothing is validated against the offered set, since
scope_json narrows after every gate -- naming a gate that was never offered
switches off something that was not on.

Then the scheduling instructions, audited against a 4B model on this machine
rather than against my own reading of them. Ten realistic requests, ten
compiled, twice over -- so the prompt is sound. What was not sound was
`describe`, which built a phrase by joining fragments and read "Every the 1st at
09:00" for the commonest monthly schedule there is, and "Every of January" for a
month with no day. That string is the whole of what somebody sees before
approving a schedule and the whole of what the model is told about its own chat,
so a phrase nobody can parse is a review step nobody performs. It reads as
English now, collapses Monday-to-Friday to "every weekday" and seven days to
"every day", and every case in the test is a rule that model actually produced.

The one mistake it made was naming Wednesday for "every other tuesday", so the
weekday numbering is spelled out rather than left as "0-6, Monday is 0": getting
that wrong is the error here that still looks like a working schedule. Roughly
one call in six also came back empty -- a local runner swapping models under the
request will do that -- so an unusable reply is asked for once more before giving
up. Not on an LLMError: an endpoint that refused will refuse again, and the
reader is better served by the form than by waiting twice for the same answer.

Canvas asked for a typed path, which was the last control in the application
expecting somebody to remember an absolute path on another machine -- the same
complaint the folder page's directory field answered with a picker. /browse takes
pick=file and the same fragment makes files buttons, because a second copy of
that listing is a second place for the path arithmetic to be got subtly
differently. The button carries data-canvas-open rather than an hx-post since the
path is not known until the dialog closes, and ui.js posts it through htmx.ajax
so the response lands in the panel exactly as every other canvas action's does.
The key is `agent:<path>`, so a file opened by hand and one opened by the model
are one tab rather than two spellings of it. The tabs already existed and already
closed; they now square off at the bottom and the active one takes the body's
background, so which is selected is structural rather than a tint nobody can see
in a theme they did not choose. Highlighting was already there for every language
named and is checked for fifteen of them.

Three smaller ones. Tabs kept their scroll position, so switching from a long
panel to a short one left the browser clamping to that panel's bottom: the end of
it above a screen of nothing, which reads as a page that failed to load. Nothing
in CSS can reset a scroll position. The sidebar's footer and the composer sit
either side of one vertical edge and were both content-sized, so their top
borders met it at different heights and read as one line that had been broken --
`--footer-height` is a calc of the pieces the footer is built from, applied as a
min-height to both, which is exactly what `--header-height` already does at the
top of the shell. And "Add a workflow" sat flush against the list it adds to,
stated as an adjacency because `.btn-row` is right to carry no margin everywhere
else it appears.

Both pieces of JavaScript were driven under a DOM stub before committing, which
is how the tab listener's delegation and the canvas button's six behaviours were
checked at all -- `node --check` parses a file that does nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jaroslav Beneš
2026-08-05 22:37:47 +02:00
parent 7ff4c2c0aa
commit 4c78215e31
23 changed files with 910 additions and 72 deletions
+53
View File
@@ -329,3 +329,56 @@ def test_reindexing_somebody_elses_chat_is_not_possible(
)
assert client.post(f"/api/chats/{chat.id}/index").status_code == 404
# --- Picking a file rather than a directory ---------------------------------------
def test_files_are_inert_when_a_directory_is_wanted(
client: TestClient, db, registered, served_tree
):
"""The default, and the older of the two. Hiding files would make a folder
of nothing but files look empty, which is worse than showing what is there
and not letting it be chosen."""
profile = _profile(db, served_tree)
body = client.get(
f"/api/agents/{profile.id}/browse", params={"path": served_tree["root"]}
).text
assert "README.md" in body
assert "data-file-open" not in body
def test_files_become_choices_when_a_file_is_wanted(
client: TestClient, db, registered, served_tree
):
"""Canvas asks for `pick=file`. One listing serves both, because a second
copy is a second place for the path arithmetic to be got subtly
differently -- and getting it differently means a file that opens to the
wrong path, or to nothing."""
profile = _profile(db, served_tree)
body = client.get(
f"/api/agents/{profile.id}/browse",
params={"path": served_tree["root"], "pick": "file"},
).text
assert "data-file-open" in body
assert f'data-file-open="{served_tree["root"]}/README.md"' in body
# Directories stay a step rather than becoming a choice.
assert "data-dir-open" in body
def test_an_unknown_pick_falls_back_to_directories(
client: TestClient, db, registered, served_tree
):
"""It arrives off a query string, so it is read as one of two things rather
than trusted -- the same shape every other value read off a request here
takes."""
profile = _profile(db, served_tree)
body = client.get(
f"/api/agents/{profile.id}/browse",
params={"path": served_tree["root"], "pick": "whatever"},
).text
assert "data-file-open" not in body
+73
View File
@@ -529,3 +529,76 @@ def test_the_tab_routes_refuse_the_wrong_method(
chat_id = make_chat()
assert getattr(client, verb)(f"/api/chats/{chat_id}/canvas/tabs").status_code == 405
assert getattr(client, verb)(f"/api/chats/{chat_id}/canvas/save").status_code == 405
# --- Opening a file by looking rather than by spelling --------------------------
def _agent_chat(db, make_chat) -> Chat:
"""An agent chat with a usable connection, so `agent_ready` says yes.
No SSH server is needed: the panel's head, tab strip and open row render
without reading anything off the far side, which is exactly the part under
test here.
"""
from lembas.db.models import SshProfile
_add_connection(db)
chat = db.get(Chat, make_chat())
profile = SshProfile(
owner_id=chat.user_id,
name="box",
host="127.0.0.1",
port=22,
username="tester",
auth="password",
password_encrypted=encrypt(""),
host_key="ssh-ed25519 AAAA",
default_dir="/srv/app",
enabled=True,
)
db.add(profile)
db.commit()
chat.kind = KIND_AGENT
chat.ssh_profile_id = profile.id
chat.project_dir = "/srv/app"
db.commit()
settings_store.update(db, {"enabled": True}, key=settings_store.AGENTS)
return chat
def test_the_panel_offers_a_dialog_rather_than_a_path_box(
client: TestClient, db, registered, make_chat
):
"""This was the one control left that asked somebody to remember an
absolute path on another machine -- the same complaint the folder page's
directory field answered with a picker.
The button cannot carry an `hx-post`: the path is not known until the dialog
closes, so ui.js posts it afterwards. What it must carry is the profile the
browse endpoint is hung off and the directory to open at, or the dialog
starts at the account's home and every path is a walk from there.
"""
chat = _agent_chat(db, make_chat)
body = client.get(f"/api/chats/{chat.id}/canvas").text
assert "data-canvas-open" in body
assert 'data-profile="' + chat.ssh_profile_id + '"' in body
assert 'data-dir="/srv/app"' in body
# The typed path box is gone, and with it the only field named `key` that a
# person was expected to fill in by hand.
assert 'name="key"' not in body
def test_a_chat_with_no_machine_is_offered_no_file_dialog(
client: TestClient, db, registered, make_chat
):
"""`canvas_agent` gates it, and it is re-derived server-side on every
request -- a button that opened an empty browser would be worse than none."""
_add_connection(db)
chat_id = make_chat()
body = client.get(f"/api/chats/{chat_id}/canvas").text
assert "data-canvas-open" not in body
# Scratch is still there: it needs no machine.
assert "scratch:" in body
+118
View File
@@ -12,6 +12,7 @@ from __future__ import annotations
import pytest
from fastapi.testclient import TestClient
from sqlalchemy import select
from lembas.db.models import Chat, Connection, Model, User
from lembas.services import settings_store
@@ -322,3 +323,120 @@ def test_with_nothing_to_narrow_there_is_no_button_at_all(
html = client.get(f"/chat/{chat.id}").text
assert "picker__menu--scope" not in html
# --- Before the chat exists -------------------------------------------------------
def test_the_menu_is_there_before_the_first_message(
client: TestClient, db, chat, registered
):
"""The bug this section exists for.
The harness puts a tool's guidance in front of the model the moment the tool
is offered — so a menu that only appeared once a chat existed was one you
could not reach until after the model had been told how to keep notes and
been handed the tools to do it. Switching it off then does not un-send that
turn.
"""
settings_store.update(db, {"enabled": True}, key=settings_store.SEARCH)
html = client.get("/chat").text
assert 'aria-label="Toggle"' in html
assert 'name="scope_all"' in html
assert 'name="scope_on"' in html
# And it posts nothing on its own: there is no row to post to yet.
assert "/scope" not in html
def test_the_prospective_switches_ride_with_the_first_message(
client: TestClient, db, chat, registered
):
"""A browser submits only the ticked boxes, so "which were unticked" needs
the hidden mirror. This asserts the pair exists per gate rather than that
the markup looks a certain way."""
from html.parser import HTMLParser
settings_store.update(db, {"enabled": True}, key=settings_store.SEARCH)
html = client.get("/chat").text
named: dict[str, list[str]] = {"scope_all": [], "scope_on": []}
class Finder(HTMLParser):
def handle_starttag(self, tag, attrs):
got = {key: (value or "") for key, value in attrs}
if got.get("name") in named:
named[got["name"]].append(got.get("value", ""))
Finder().feed(html)
assert named["scope_all"], "the prospective menu rendered no gates"
# Every gate offered has both halves, or one of them can never be turned off.
assert set(named["scope_all"]) == set(named["scope_on"])
def test_unticking_before_sending_writes_it_to_the_new_chat(
client: TestClient, db, chat, registered
):
"""End to end: what the menu was set to is what the row is created with, so
the very first request is already narrowed."""
settings_store.update(db, {"enabled": True}, key=settings_store.SEARCH)
client.post(
"/api/chats/start",
data={
"content": "hello",
"model_id": "m",
"scope_all": ["notes", "memory", "web_search"],
# `notes` left out: unticked.
"scope_on": ["memory", "web_search"],
},
)
fresh = db.scalars(
select(Chat).where(Chat.user_id == chat.user_id).order_by(Chat.created_at.desc())
).first()
assert tools_service.scoped_off(fresh) == frozenset({"notes"})
assert "notes_search" not in _names(db, fresh, db.get(User, chat.user_id))
def test_leaving_everything_ticked_writes_nothing(client: TestClient, db, chat, registered):
"""Absent means on, and there is one representation of it. A row full of
`True`s would be a second one, and "why is this off?" would have two
answers."""
client.post(
"/api/chats/start",
data={
"content": "hello",
"model_id": "m",
"scope_all": ["notes", "memory"],
"scope_on": ["notes", "memory"],
},
)
fresh = db.scalars(
select(Chat).where(Chat.user_id == chat.user_id).order_by(Chat.created_at.desc())
).first()
assert fresh.scope_json == {}
def test_starting_a_chat_cannot_widen_through_the_menu(
client: TestClient, db, chat, registered
):
"""The security-shaped half, from the new-chat side. `scope_json` narrows
inside `resolve_tools` *after* every gate, so naming a gate that was never
offered switches off something that was not on — which is nothing. A
crafted POST cannot turn anything on, because there is no representation
for "on" to send."""
client.post(
"/api/chats/start",
data={
"content": "hello",
"model_id": "m",
"scope_all": ["agent", "made_up"],
"scope_on": ["agent", "made_up"],
},
)
fresh = db.scalars(
select(Chat).where(Chat.user_id == chat.user_id).order_by(Chat.created_at.desc())
).first()
assert "shell_run" not in _names(db, fresh, db.get(User, chat.user_id))
+79
View File
@@ -337,3 +337,82 @@ def test_an_ordinary_chat_is_told_none_of_it(client: TestClient, db, registered)
block = harness.compose(db, _user(db), [], chat)
assert "nobody is necessarily reading" not in block
assert "This scheduled task" not in block
@pytest.mark.anyio
async def test_an_unusable_reply_is_asked_again_once(
client: TestClient, db, registered, monkeypatch
):
"""Measured against a 4B model: the prompt is sound — ten realistic requests
compiled ten times over, twice — but roughly one call in six came back empty
or truncated, which a local runner swapping models under the request will
do. One retry costs a second on a screen somebody is already waiting at."""
replies = ["", json.dumps({"title": "T", "instruction": "I",
"schedule": {"at": {"times": ["09:00"]}}})]
calls = 0
async def answer(endpoint, payload):
nonlocal calls
calls += 1
return replies.pop(0)
monkeypatch.setattr(compile_service, "complete", answer)
compiled = await compile_service.compile_request(
_endpoint(), "test-model", "daily at nine", template=_template(db), user=_user(db)
)
assert calls == 2
assert compiled.ok is True
@pytest.mark.anyio
async def test_it_gives_up_after_the_second_try(
client: TestClient, db, registered, monkeypatch
):
calls = 0
async def answer(endpoint, payload):
nonlocal calls
calls += 1
return "I'd suggest weekly."
monkeypatch.setattr(compile_service, "complete", answer)
compiled = await compile_service.compile_request(
_endpoint(), "test-model", "daily at nine", template=_template(db), user=_user(db)
)
assert calls == 2
assert compiled.ok is False
assert compiled.instruction == "daily at nine"
@pytest.mark.anyio
async def test_an_endpoint_that_refuses_is_not_asked_twice(
client: TestClient, db, registered, monkeypatch
):
"""It will refuse again, and the reader is better served by the form than by
waiting twice for the same answer."""
calls = 0
async def refuse(endpoint, payload):
nonlocal calls
calls += 1
raise LLMError("connection refused")
monkeypatch.setattr(compile_service, "complete", refuse)
await compile_service.compile_request(
_endpoint(), "test-model", "daily at nine", template=_template(db), user=_user(db)
)
assert calls == 1
def test_the_weekday_numbering_is_spelled_out(client: TestClient, db, registered):
"""The one mistake a small model actually made in the audit: "every other
tuesday" came back as Wednesday. Naming the wrong day is the error here that
still looks like a working schedule, so the mapping is written out rather
than left as "0-6, Monday is 0"."""
template = _template(db)
for day, number in (("Monday", 0), ("Wednesday", 2), ("Sunday", 6)):
assert f"{day}={number}" in template
+45 -1
View File
@@ -356,7 +356,51 @@ def test_describe_says_what_the_rule_actually_does():
{"start": "2026-08-05T12:00:00Z", "every": {"minutes": 10}, "count": 5},
"Every 10 minutes, 5 times",
),
({"at": {"days": [1], "times": ["09:00"]}}, "Every the 1st at 09:00"),
]
for raw, expected in cases:
assert rule_service.describe(rule_service.validate(raw), zone=PRAGUE) == expected
def test_describe_reads_as_english_for_what_a_model_actually_writes():
"""Every case below is a rule a 4B model produced from a plain request, and
the first two used to read "Every the 1st at 09:00" and "Every of January".
Worth pinning as *wording*, which is not a thing tests usually assert. This
string is the whole of what the reader sees before approving a schedule and
the whole of what the model is told about its own chat — so a phrase nobody
can parse is a review step nobody performs.
"""
cases = [
# "on the 1st of every month, write up what changed"
({"at": {"days": [1], "times": ["09:00"]}}, "On the 1st of each month at 09:00"),
(
{"at": {"days": [1, 15], "times": ["09:00"]}},
"On the 1st and 15th of each month at 09:00",
),
# "every weekday at 8am give me a briefing" — five names in a row is the
# commonest thing this produces otherwise.
({"at": {"weekdays": [0, 1, 2, 3, 4], "times": ["08:00"]}}, "Every weekday at 08:00"),
(
{"at": {"weekdays": [5, 6], "times": ["10:00"]}},
"Every Saturday and Sunday at 10:00",
),
# Naming all seven is no constraint at all, and saying so is how "every
# day" comes out of a rule that enumerated them.
(
{"at": {"weekdays": [0, 1, 2, 3, 4, 5, 6], "times": ["09:00"]}},
"Every day at 09:00",
),
({"at": {"months": [1, 7], "times": ["09:00"]}}, "Every day in January and July at 09:00"),
(
{"at": {"months": [1], "days": [1], "times": ["00:00"]}},
"On the 1st of January at 00:00",
),
# Both set is an AND, and rare. Said plainly rather than smoothed into
# something that reads like an OR.
(
{"at": {"weekdays": [0], "days": [1], "times": ["09:00"]}},
"On the 1st of each month, if it is a Monday at 09:00",
),
]
for raw, expected in cases:
assert rule_service.describe(rule_service.validate(raw), zone=PRAGUE) == expected
+70
View File
@@ -132,3 +132,73 @@ def test_the_finished_bubble_repeats_the_live_container_s_id():
message = (TEMPLATES / "chat/_message.html").read_text(encoding="utf-8")
assert message.count('id="steps-{{ message.id }}" data-steps') == 2
def test_switching_a_tab_puts_its_body_back_at_the_top():
"""A tab is a radio and a panel is shown by CSS, so switching one changes
nothing about `.tabs__body` -- the element that actually scrolls. Read half
way down a long panel, switch to a short one, and the browser clamps the
kept scrollTop to that panel's bottom: what lands on screen is the end of it
above a screen of nothing, which reads as a page that failed to load.
Nothing in CSS can reset a scroll position. Driven under a DOM stub before
committing; what is pinned here is that the listener is delegated and keyed
on the class rather than on one page's ids, because the next tabbed screen
would otherwise have the same bug and no sign of it.
"""
assert '.closest(".tabs__bar")' in SOURCE or 'closest(".tabs__bar")' in SOURCE
assert 'querySelector(".tabs__body")' in SOURCE
assert "scrollTop = 0" in SOURCE
def test_the_shell_has_one_line_along_its_bottom_edge():
"""The sidebar footer and the composer sit either side of the same vertical
edge and were both content-sized, so their top borders met it at different
heights and read as one line that had been broken.
Neither can match the other by accident: the footer's height depends on
which entries a reader's permissions allow, and the composer's on how much
has been typed. So both take a `min-height` from one token, the way
`--header-height` already does this at the top of the shell.
"""
tokens = (ROOT / "web/static/css/tokens.css").read_text(encoding="utf-8")
app = (ROOT / "web/static/css/app.css").read_text(encoding="utf-8")
chat = (ROOT / "web/static/css/chat.css").read_text(encoding="utf-8")
assert "--footer-height:" in tokens
# A calc of the pieces the footer is built from, not a measured constant --
# a number would stop being true the moment `--control-h` moved.
assert "var(--control-h)" in tokens.split("--footer-height:")[1].split(";")[0]
assert "min-height: var(--footer-height)" in app
assert "min-height: var(--footer-height)" in chat
def test_the_canvas_open_button_posts_through_htmx():
"""The button cannot carry an `hx-post`: the path is not known until the
dialog closes. So it posts afterwards — through htmx's own `ajax`, so the
response lands in the panel exactly as every other canvas action's does.
`fetch` would mean parsing and swapping the fragment by hand, and then there
would be two ways the canvas gets replaced. Driven under a DOM stub before
committing; what is pinned here is the shape that keeps them one.
"""
assert "window.htmx.ajax(" in SOURCE
assert '"#canvas-inner"' in SOURCE
# The key is `agent:<path>` -- the same key a tool call's read produces, so
# a file opened here and one opened by the model are one tab rather than two
# spellings of it. That is `canvas.path_key`'s whole job.
assert '"agent:" + path' in SOURCE
def test_the_file_picker_asks_for_files():
"""One listing serves the directory picker and the file picker, because a
second copy is a second place for the path arithmetic to be got subtly
differently — and getting it differently means a file that opens to the
wrong path, or to nothing."""
app = (ROOT / "web/static/js/app.js").read_text(encoding="utf-8")
assert "pick=file" in app
assert "data-file-open" in app or "dataset.fileOpen" in app
# Directories stay a step in file mode, or a file two folders down is
# unreachable.
assert app.count("[data-dir-open]") >= 2