What the audit is worth keeping, and where
Five passes produced a working document that said on its first line it was temporary. This is it being spent rather than abandoned. CLAUDE.md gains eleven paragraphs, each a thing that had shipped looking correct: a handler bound to a shared variable rather than its own socket, a script the base template already loads being loaded again, a control that stays clickable while it awaits permission, "is this name taken?" asked about visibility instead of ownership, root running a file the service account can write, sourcing anything under $PREFIX, a read-only command name that is not a read-only command, 0.0.0.0 being this machine, a folder that is not a label, a file that is not deleted by the row that named it, and a measuring harness that measured an unstyled page and reported a dramatic finding that was entirely an artefact. PLAN.md carries the seven things the audit found and deliberately did not fix, each with why: they change what something does rather than fix what it claims to do, which is not an audit's job. deploy/README.md says why root runs a copy, and that a host installed before this keeps the old wiring until the installer is re-run -- the button cannot fix it, because the button runs the old unit. docs/notes/audit-0.9.md is deleted, having been all three. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,416 +0,0 @@
|
||||
# The 0.9 audit
|
||||
|
||||
A working document, written during the pass over everything that landed between
|
||||
`0.8.1` and `0.9.8` -- 16 commits, 163 files, +16,795 / −1,608. It exists so the
|
||||
findings survive the session that produced them, and so the security and testing
|
||||
stages have something to work from rather than re-deriving it.
|
||||
|
||||
**This file is temporary.** Everything in it either becomes a `CHANGELOG.md`
|
||||
line, a paragraph in `CLAUDE.md`, or an entry under "After 1.0.0" in `PLAN.md`.
|
||||
It is deleted before the release.
|
||||
|
||||
Each finding carries a disposition:
|
||||
|
||||
- **fixed** -- landed, with a test
|
||||
- **later** -- real, but the fix changes what a feature does, so it is the
|
||||
user's call and it is written up rather than done
|
||||
- **stated** -- correct as it stands, and recorded so the next audit does not
|
||||
spend an hour rediscovering that it is deliberate
|
||||
|
||||
---
|
||||
|
||||
## The harness and the instruction prompts
|
||||
|
||||
### {{timezone}} shipped as a hole — *fixed*
|
||||
|
||||
`clock.name_for` returns `""` for anybody who has never chosen a zone, which is
|
||||
the default state of every account. `substitute` drops a line only when the
|
||||
whole line is **blank** after expansion, and this variable sits inside a
|
||||
sentence, so every such request carried:
|
||||
|
||||
```
|
||||
- Times the person gives you are in unless they say otherwise.
|
||||
```
|
||||
|
||||
Both the comment at `harness.py:191` and the fragment's own hint asserted the
|
||||
line disappeared. Neither had ever been true. The sibling call site
|
||||
(`schedule/compile.py:102`) has always had the fallback.
|
||||
|
||||
Worth noting *why* the fallback is right rather than merely convenient: `stamp`
|
||||
is `clock.now_for(user)`, which already falls back to the server's zone, and
|
||||
`{{now}}` already prints its offset. The times were already in that zone.
|
||||
Withholding the label from a value the model had been given was not restraint.
|
||||
|
||||
### The admin preview could not show most of what it previews — *fixed*
|
||||
|
||||
`admin_prompts.py` built its variables with `context_variables(db, user, [], None)`
|
||||
— no tools, no chat — and backfilled six names. Every other `requires` gate
|
||||
stayed empty, so **eleven fragments could never appear in a preview whatever an
|
||||
administrator ticked**: the whole agent surface, both scheduling fragments and
|
||||
the helper warning. Editing `tool.agent` and pressing preview showed a system
|
||||
message with `tool.agent` missing from it, silently.
|
||||
|
||||
Fixed with sample values, which is what `SAMPLE_DOCUMENTS` already did for
|
||||
attachments, plus two controls the "Tools offered" row could not reach — "This
|
||||
chat is" (ordinary / scheduled task / helper) and an agent-mode select. The
|
||||
samples are gated **exactly as `context_variables` gates the real values**,
|
||||
because a preview that admits a fragment the real request would not is worse
|
||||
than one that omits it.
|
||||
|
||||
### Fragments named tools that had been withdrawn — *fixed for Plan mode, later for the rest*
|
||||
|
||||
`resolve_tools` filters per **tool**; `harness._families` re-derives gates per
|
||||
**family**. So a family survives on its readers while its writers are gone, and
|
||||
the family-gated fragment still names them — under `core.tool_list` saying
|
||||
"Anything not named there does not exist here — calling it costs a round and
|
||||
returns nothing."
|
||||
|
||||
**Fixed:** Plan mode. `agent/tools.py:1185` withdraws `plan_update` there, and
|
||||
`tool.plan_update` was gated on `{{plan}}`, which is set whenever a plan exists
|
||||
in any mode. A new `plan_editable` variable — the plan text, blank in Plan mode
|
||||
— now gates it. Writing the test found a second instance the audit had missed:
|
||||
`context.plan` also said "Change it with plan_update", so that sentence moved to
|
||||
the fragment whose job it is.
|
||||
|
||||
**Later:** a read-only helper keeps seven fragments naming fifteen withdrawn
|
||||
write tools (`notes_create`, `report_write`, `file_write`, `skill_create`, the
|
||||
three schedule writers…). The principled fix is the split `tool.skills` /
|
||||
`tool.skills_write` already demonstrates, applied to `notes`, `report`,
|
||||
`schedule` and `agent_edits` — four more fragment pairs. That is a prompt
|
||||
restructure rather than a bug fix and it is proposed for after 1.0.0. The cost
|
||||
today is bounded: `{{tool_names}}` is authoritative and the model has it, so a
|
||||
helper wastes at most one round discovering a writer is absent.
|
||||
|
||||
### Three tool descriptions disagreed with their runners — *all fixed*
|
||||
|
||||
- **`subagent_run` opened "a second model with the same tools"** and contradicted
|
||||
itself two sentences later. The child is forced to `ask:False`,
|
||||
`subagent:False`, `write:False` with commands pinned to `SAFE_COMMANDS`. The
|
||||
leading clause is what a model reads when sizing a task.
|
||||
- **`knowledge_get` returned `extracted_text` whole** while every sibling reader
|
||||
capped and said so — `fetch` at 20k, `file_read`, the memories block, the
|
||||
skill index, the project listing. `MAX_EXTRACTED_CHARS` defaults to 120,000,
|
||||
so one call on a long PDF filled an ordinary window with nothing reporting it.
|
||||
Its description said "in full", which is why this read as correct: the tool did
|
||||
exactly what it claimed. Now `MAX_DOCUMENT_CHARS = 40_000`, cut with the model
|
||||
told, which is what `fetch` and `file_read` both do.
|
||||
- **`_run_ask_user` read `args["question"]`**, singular, against a schema
|
||||
declaring `questions` and a list — so its event always carried an empty
|
||||
`query`. Harmless only because the path is documented unreachable, which is
|
||||
exactly why nothing caught it.
|
||||
|
||||
### Fragment text that was false — *fixed*
|
||||
|
||||
`tool.subagent_agent` said a helper "reads and reports … and nothing else, in
|
||||
every mode" beside a `write` parameter that makes one write files, and named
|
||||
**seven** of the **twenty-three** commands in `SAFE_COMMANDS` — so a model
|
||||
avoided commands it was allowed, which costs nothing visible and is therefore
|
||||
never reported. Both halves are now pinned by tests against the real list and
|
||||
the real schema, because prose and a constant drift the moment one is edited
|
||||
alone.
|
||||
|
||||
The distinction the wording now has to carry, and did not before: **commands**
|
||||
are fixed to the read-only list in every mode, since an unattended chat cannot
|
||||
approve anything; **files** may be written, but only by a helper asked for with
|
||||
`write` and only from Edit or Auto.
|
||||
|
||||
### `tool.background` promises a notification that can be off — *later*
|
||||
|
||||
"When a background job finishes you are told in a new turn" has no `requires`
|
||||
for `agents.background_notify`, and the runner itself branches on that flag and
|
||||
says "Check on it with `job_output`" instead. One fragment, two behaviours. The
|
||||
fix is a variable and a gate, the same shape as `plan_editable`; deferred only
|
||||
because it wants the split above rather than a third one-off.
|
||||
|
||||
### `FAMILY_ASK` has no fragment at all — *later*
|
||||
|
||||
Every other family in `FAMILIES` contributes one. `ask_user` is offered in
|
||||
nearly every chat and **all** of its guidance lives in its schema description,
|
||||
which is the one thing an administrator cannot edit. Proposed rather than done:
|
||||
it is new prompt text, and new prompt text in an audit stage is how an audit
|
||||
turns into a feature.
|
||||
|
||||
### Smaller, all *stated*
|
||||
|
||||
- **`model_name` is resolved on every request and rendered by no shipped
|
||||
fragment.** It is documented in `VARIABLES` and available to an
|
||||
administrator's override, so it is an escape hatch rather than dead work. The
|
||||
cost is one small `select` on a table with tens of rows.
|
||||
- **Four pairs of fragments share an `order`**, which undercuts "sparse so a
|
||||
custom tool can be slotted between two of these". Ties break alphabetically
|
||||
and deterministically. One has a readable consequence: `context.schedule`
|
||||
sorts below `context.memories`, so a scheduled task's own instruction appears
|
||||
under the reader's remembered facts.
|
||||
- **`core.rounds` is dark on a stock install**, since `DEFAULT_CHAT_ROUNDS = 0`.
|
||||
Deliberate — `core.keep_working` is its complement and exactly one ever fires.
|
||||
- **A row-backed fragment with empty `guidance` can never be selected**, so for
|
||||
those "never configured" and "deliberately switched off" are the same state —
|
||||
the one distinction the absent-vs-empty convention exists to preserve.
|
||||
|
||||
---
|
||||
|
||||
## Controls
|
||||
|
||||
The four sweeps `CLAUDE.md` is shaped around were run by hand over every
|
||||
template and **all four came back clean**: 68 htmx verbs against 179 registered
|
||||
routes with zero method mismatches, 88 plain form actions and 30
|
||||
`fetch`/`htmx.ajax` call sites with zero, every verb-carrying element declaring
|
||||
its own `hx-target` or `hx-swap="none"` checked against every ancestor, no empty
|
||||
verb attributes, both `form="…"` sites correct, and every `hx-target`,
|
||||
`hx-include` and `from:` selector resolving to an id that exists.
|
||||
|
||||
Making each one a **test that sweeps** is Stage 2's work. A one-off audit that
|
||||
finds nothing is worth much less than the same audit made permanent.
|
||||
|
||||
## Capabilities that are built, documented and unreachable — *Stage 2*
|
||||
|
||||
The `Model.params_json` and `library.share` family: the code is right, the route
|
||||
validates, and nothing can reach it.
|
||||
|
||||
- **Folder nesting.** `PATCH /api/folders/{id}` handles `parent_id` with a cycle
|
||||
guard and `MAX_DEPTH = 8`, `partials/_folder.html` recurses, and `README.md`
|
||||
advertises "arbitrarily nested". No template or script submits `parent_id`.
|
||||
- **Moving a chat into a folder.** `api/chats.py` handles `folder_id` on PATCH;
|
||||
only the composer submits it, at creation. `chat/_composer.html` even carries
|
||||
a comment reasoning about "the only way into a folder was to make the chat
|
||||
elsewhere and move it".
|
||||
- **`Connection.extra_headers_json`** is read into every LLM request and written
|
||||
by no form, so its documented use — OpenRouter's `HTTP-Referer` — is
|
||||
unreachable. *later*: nothing advertises it, so nothing is currently untrue.
|
||||
|
||||
Written and never read: `Chat.compacted_at`, `User.last_login_at`,
|
||||
`Schedule.last_fire_at`, `Schedule.compiled_at`. Neither read nor written:
|
||||
`general.require_approval` (its comment says "reserved"), and the `back` form
|
||||
parameter on two `admin_models` routes.
|
||||
|
||||
**A third, found by asking where generated images go.** `db.delete(chat)`
|
||||
cascades to the message and attachment *rows* and leaves every file on disk;
|
||||
`files.remove_files_for_chats` was written for exactly that and was called from
|
||||
one place, the temporary-chat sweep. The delete button, a schedule's task chat,
|
||||
a helper's chat and account deletion all went straight to `db.delete`.
|
||||
|
||||
Underneath it, a second fault that made the first invisible: `files.claim`
|
||||
bound `message_id` and never `chat_id`, so anything uploaded on the *new-chat*
|
||||
screen kept an empty `chat_id` for the rest of its life. Six readers filter on
|
||||
that column, so for those files the harness did not name them among the attached
|
||||
documents, the canvas refused to open them, and the one caller the cleanup
|
||||
helper had found nothing to remove. Both fixed, and `chat_service.delete_chats`
|
||||
is now the only way a chat is deleted.
|
||||
|
||||
*Stated:* `Message.parent_id` and `content_parts_json` are deliberate forward
|
||||
-looking columns; `Chat.archived` is already recorded in `PLAN.md` as a column
|
||||
nothing surfaces.
|
||||
|
||||
## Verified by rendering, not by reading — *Stage 2*
|
||||
|
||||
There is no browser extension on this machine, so `/admin/prompts` and
|
||||
`/admin/customization` were checked by rendering them through `TestClient`,
|
||||
rewriting the asset URLs to `file://`, and driving headless Chromium against the
|
||||
result. Worth recording because the first run of that harness **measured an
|
||||
unstyled page**: the templates emit absolute `http://testserver/static/…` URLs
|
||||
via `url_for`, not `/static/…`, so the replacement matched nothing and no CSS
|
||||
loaded. It reported all five tab panels visible at once — a dramatic finding
|
||||
that was entirely an artefact. The harness now fails loudly on any unrewritten
|
||||
asset URL.
|
||||
|
||||
What it measured once fixed:
|
||||
|
||||
- `/admin/prompts`, tab bar's on-screen position across the five tabs:
|
||||
**385 → 642px before** (the page moving under the reader), **195px constant
|
||||
after**, with `scrollTop` staying at 0.
|
||||
- `/admin/customization`, document height: **5495px before, 2820px after**.
|
||||
|
||||
## Documentation — *Stage 2*
|
||||
|
||||
- `README.md` is stale by six phases. "Planned: Image generation · OCR ·
|
||||
semantic search" — two of the three shipped. Nothing under Features mentions
|
||||
schedules, reports, subagents, branding, quotas, web push, updates, Docker or
|
||||
LXC.
|
||||
- `CLAUDE.md` says "Three topics live in `docs/notes/`" and lists seven.
|
||||
- `CLAUDE.md` says 2088 tests; there are 2092 before this stage.
|
||||
- `tests/__pycache__/test_zz_{dump,live}*.pyc` are stale bytecode for two files
|
||||
that no longer exist.
|
||||
|
||||
## Security — *Stage 3*
|
||||
|
||||
**Two privilege escalations in the update helper, both root, both fixed.** The
|
||||
helper is the one place this application deliberately crosses a privilege
|
||||
boundary, and it crossed it twice more than intended. Neither is reachable
|
||||
from the web interface: both need code execution as the `lembas` service
|
||||
account first. That is precisely the boundary the unprivileged split exists to
|
||||
hold, so "you need a foothold" is the threat model, not a mitigation.
|
||||
|
||||
**1. Root ran a script the service account owns.**
|
||||
`ExecStart=/bin/bash __PREFIX__/app/deploy/update.sh` — inside the checkout,
|
||||
owned `lembas:lembas`, because `install.sh` clones as that user. So: write your
|
||||
payload into `deploy/update.sh`, `touch data/update-requested` (the service
|
||||
account owns that directory too), and systemd runs it as root. The web
|
||||
interface's `AdminUser` check is not the gate systemd honours.
|
||||
|
||||
There is a second path needing no compromise at all: an update pulls new code
|
||||
*as the service user*, and root then executes whatever `deploy/update.sh` that
|
||||
pull contained. **Control of the branch was control of root.**
|
||||
|
||||
Fixed by installing a root-owned copy at `/usr/local/lib/lembas/update.sh` and
|
||||
pointing the unit there. The cost — improving `update.sh` needs the installer
|
||||
re-run — is the right one: root should not execute a script that arrived over
|
||||
the network a moment ago. The script warns when its own copy has fallen behind.
|
||||
|
||||
**The old test asserted the vulnerable line**
|
||||
(`assert "ExecStart=/bin/bash __PREFIX__/app/deploy/update.sh" in unit`). It
|
||||
passed for the life of the feature and pinned the bug in place — the recurring
|
||||
failure of this codebase, applied to a privilege boundary: an assertion about
|
||||
the text rather than about the property the text was meant to have.
|
||||
|
||||
**2. Root sourced a file the service account can replace.**
|
||||
`. "$PREFIX/.deploy-env"`. The file is root-owned, having been written with
|
||||
`sudo tee` — but `$PREFIX` is the service account's own directory at mode 755,
|
||||
and write permission on a *directory* is all it takes to unlink a file and put
|
||||
another there. On the live host `.deploy-env` did not even exist, so it could
|
||||
simply be created. `.` runs its contents as root.
|
||||
|
||||
This one survived the first fix entirely, and the helper is what made it
|
||||
reachable: before the `.path` unit existed, `update.sh` only ran when an
|
||||
administrator invoked it from a shell. Fixed by parsing the two values it wants
|
||||
with strict patterns instead of sourcing. The test asserts that **nothing**
|
||||
under `$PREFIX` is sourced, rather than naming `.deploy-env`, because the next
|
||||
file read from there would have the same problem.
|
||||
|
||||
**Upgrade note:** a host that installed the helper before this keeps the old
|
||||
unit, and only re-running the installer moves it. `update.sh` now detects that
|
||||
it is running from inside the checkout and says so loudly — otherwise the
|
||||
vulnerable hosts are exactly the ones that never hear about it.
|
||||
|
||||
**Also fixed, sub-threshold as a vulnerability but a real bug:** the share
|
||||
panel built its `hx-vals` by pasting the search term into a JSON string. Jinja
|
||||
escapes the quote for HTML and the parser decodes it again before htmx parses
|
||||
the JSON, so a `"` in a search term ended the string and silently stopped every
|
||||
checkbox in the panel from submitting anything. `q` was the last key, so an
|
||||
injected one would also have won a duplicate-key parse. Built with `| tojson`
|
||||
over the whole object now.
|
||||
|
||||
**3. A read-only command that was not read-only.** `SAFE_COMMANDS` — the list a
|
||||
**subagent** is pinned to, in every mode, unattended, with no approval card
|
||||
possible — contained `find *`. GNU `find` writes files (`-fprintf`), runs
|
||||
programs (`-exec … +`) and deletes them (`-delete`), and none of those needs a
|
||||
character `policy._UNSAFE` refuses. `rg --pre` is the same shape.
|
||||
|
||||
So the chain was: a parent in **Plan** mode — which promises "reads freely,
|
||||
changes nothing" — spawns a helper on a `RISK_READ` tool with no card; the
|
||||
helper's `shell_run` survives because `writes_off` drops `RISK_WRITE` and
|
||||
deliberately keeps `RISK_EXECUTE`; `find . -maxdepth 0 -fprintf
|
||||
~/.ssh/authorized_keys 'ssh-ed25519 …'` matches `find *` and runs. Prompt
|
||||
injection from a page the model just read is enough to start it.
|
||||
|
||||
Fixed with `policy._ACTION`, refusing those flags in `subject()` rather than
|
||||
trimming the allow list — a pattern cannot express "and no dangerous flags",
|
||||
and "this one looks read-only" is exactly the reasoning that put `find *` there.
|
||||
It costs a false refusal on `grep -- -delete`, which is the right direction to
|
||||
be wrong in: a refusal asks, an allow does not.
|
||||
|
||||
**4. `0.0.0.0` walked past the loopback guard.** `_literal` answered from
|
||||
`is_loopback`, and `0.0.0.0`/`::` are `is_unspecified` — so it returned a
|
||||
*decided* `False`, which short-circuited `resolves_here` and skipped the DNS
|
||||
half too. `connect()` to either goes to loopback on Linux, so an SSH profile
|
||||
pointed at `0.0.0.0` reached this host's own sshd: the one spelling of "this
|
||||
machine" that walked past the guard whose whole job is that sentence.
|
||||
|
||||
**5. Push endpoints skipped the SSRF guard.** `POST /api/push/subscribe`
|
||||
checked `startswith("https://")` and nothing else, and `send_one` POSTed to it
|
||||
with no `check_url` — the only outbound client in the codebase not going
|
||||
through the guard. Delivery is triggered by the caller: send a message, close
|
||||
the tab, and `_persist` announces it because nobody is following. Checked now
|
||||
at subscribe **and** again before the POST, since the row outlives the first
|
||||
check.
|
||||
|
||||
**6. A chat could be put in somebody else's folder.** `effective_system_prompt`
|
||||
walks up from the chat through its folder and that folder's parents, so this
|
||||
reads another account's system prompt through a field that looks like a tag.
|
||||
Both paths had it, and `_new_chat`'s is the instructive one: it resolved the
|
||||
folder, discarded it when it was not the caller's, and then stored the **raw
|
||||
id** anyway — so the ownership check governed which *seeds* were applied and
|
||||
not where the chat actually went.
|
||||
|
||||
### Clean
|
||||
|
||||
Checked and found sound: the branding CSS and custom-theme generation (ids and
|
||||
colour values both validated on **read**, so a row written by hand still cannot
|
||||
emit a malformed rule; served as `text/css` rather than inline, so there is no
|
||||
HTML context to escape); the unauthenticated branding asset route (random
|
||||
names, traversal guarded twice, magic-number sniffing, SVG excluded); sharing
|
||||
authorisation on every route; the request file's contents reaching nothing;
|
||||
`updates._git`'s fixed argv; and the container (non-root, no secret baked, no
|
||||
docker socket, loopback only).
|
||||
|
||||
## Carried forward from earlier stages
|
||||
|
||||
Questions raised before the security stage, answered by it:
|
||||
|
||||
- the three hand-rolled redirect loops each re-run `check_url` per hop
|
||||
(confirmed); does each also drop the secret when a hop leaves its origin?
|
||||
- `admin_tools.py` calls `check_url(server.url, allow_private=True)`
|
||||
unconditionally at save — a syntax check, with the per-row flag governing at
|
||||
call time. Confirm that reading is right.
|
||||
- the subagent restriction is the one gate that fails open unattended.
|
||||
`_child_scope` forcing `ask:False`, `subagent:False`, `write:False` and
|
||||
pinning `allow` to `SAFE_COMMANDS` is confirmed present; the question is
|
||||
whether every path reaches it.
|
||||
- `pyproject.toml` pins no upper bounds and `deploy/update.sh` runs
|
||||
`pip install -e` on every update, so a breaking upstream release arrives on a
|
||||
button press. pip's `only-if-needed` default limits the blast radius, which is
|
||||
why this is a note rather than an emergency.
|
||||
|
||||
## Testing — *Stage 4*
|
||||
|
||||
2140 tests became **2283**, and writing them found four bugs that reading had
|
||||
not — three of them by driving the JavaScript under a Node DOM stub, which is
|
||||
the practice `CLAUDE.md` sets out and the reason it does.
|
||||
|
||||
- **The terminal dropped every keystroke after a reconnect.** `socket.onclose`
|
||||
closed over the module-level variable rather than its own connection, and
|
||||
`close()` queues its event: both reconnect paths close the old socket and
|
||||
immediately open a new one, so the old one's `close` arrived afterwards and
|
||||
nulled the *live* socket. Output kept coming, because `onmessage` is bound to
|
||||
the object; every send gates on `socket && readyState === OPEN`. It also said
|
||||
"Disconnected" about a shell that had just reconnected, because
|
||||
`closedOnPurpose` had been cleared for the new connection.
|
||||
- **Two scripts were loaded twice on `/messages`** — once by `base.html` and
|
||||
again by the page. Each is an IIFE with its own state, so four keyboard
|
||||
shortcuts toggled their panel twice and therefore did nothing, `/help` opened
|
||||
two dialogs, and an `@` mention attached its file twice. A sweep now refuses
|
||||
any template re-loading what `base.html` already has.
|
||||
- **The microphone had no guard while the permission prompt was up.** State only
|
||||
left `idle` after `getUserMedia` resolved, so every click in that window
|
||||
opened another stream; only the last was stopped.
|
||||
- **A shared skill took its name out of your own library.** `create` checked
|
||||
uniqueness with `by_name`, which is scoped to what is *visible* — owned or
|
||||
shared — against a `(owner_id, name)` constraint. `documents.create_base` next
|
||||
door asks it correctly. Found by writing HTTP tests for `api/library.py`.
|
||||
|
||||
Also: `--ink-faint` failed the 4.5:1 contrast minimum in **both** themes (3.85
|
||||
in Moria, 3.19 in Shire), measured by computing the ratios in a headless browser
|
||||
rather than by eye. Raised in both.
|
||||
|
||||
**Verified rather than assumed:** the suite now runs on Python **3.11, 3.12 and
|
||||
3.14**. It had only ever run on 3.14, while `Dockerfile` ships 3.12 and
|
||||
`requires-python` claimed 3.11 — the interpreter most people would actually run
|
||||
was the one nothing had tested. Both pass.
|
||||
|
||||
One measurement worth keeping honest: the `slow` marker splits 365 tests that
|
||||
stand up a real server, shell, PTY or git repository, but the fast subset is
|
||||
111s against 140s for everything. The split is for CI, not for a dramatic
|
||||
saving.
|
||||
|
||||
## Still open after the testing stage
|
||||
|
||||
- `api/library.py`: 28 routes, effectively zero HTTP coverage. Highest exposure
|
||||
— auth, ownership and permission gates all live at that boundary.
|
||||
- No real `sync_schema` upgrade test exists. `conftest.py` runs `create_all`
|
||||
then `sync_schema`, so the differ is only ever exercised as a no-op. Hard rule
|
||||
4 has no other safety net.
|
||||
- `cli.py`: zero tests, and `create-admin` is the documented bootstrap.
|
||||
- SSE has three end-to-end tests; `services/sse.py` has none.
|
||||
- `composer.js`, `terminal.js` and `audio.js` — 1,188 lines — have no assertion
|
||||
of any kind.
|
||||
- The suite has only ever run on Python 3.14. The Dockerfile ships 3.12 and
|
||||
`requires-python` claims 3.11.
|
||||
Reference in New Issue
Block a user