2 Commits
Author SHA1 Message Date
HomerandClaude Opus 5 32e2326d41 An effort the model had never heard of
Reported from a live instance, on Bonsai:

  Jinja Exception: Unexpected reasoning effort high. Supported types are
  xhigh (default), medium, and low.

Effort goes out two ways because no single field works, and the second --
chat_template_kwargs -- is not a parameter the server interprets. It is
rendered into the model's own chat template, which does not ignore a value it
does not know: it calls raise_exception, and the request dies before a token.
So a perfectly ordinary option, drawn by this application in its own menu, took
the whole reply with it.

The vocabulary is per model and nobody agrees. gpt-oss takes low/medium/high.
Bonsai takes low/medium/xhigh and refuses high. OpenAI has added minimal, xhigh
and max at different points, and which of them a given model accepts varies
again. One global tuple was going to be wrong for somebody whatever it held.

A model carries its own list now, and the picker, the slash command and the
request builder all read it. A column rather than a key in capabilities_json,
for the reason context_length is one: that dict is rebuilt wholesale from the
submitted checkboxes on every save.

And it corrects itself. A refusal retries the reply once without the effort
rather than losing it -- safe only because the template renders before any
token, so nothing has been emitted, and there is a guard that keeps it that way
-- then narrows the model's list. Bonsai's error states what it does take, so
that is what gets stored.

Note the parser bug, because it is a good one: "high" is a substring of
"xhigh", so reading the advertised list by substring learned `high` from a
sentence explaining that `high` is the problem. Whole words now, with a test
named after it.

/effort reads its levels off the picker instead of a second copy of the list
kept in the browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-25 21:20:06 +00:00
HomerandClaude Opus 5 b7bf7d728b An admin area a phone could reach and not navigate
Administration has a nav of its own rather than the chat sidebar, and 1.1.0
gave every `.sidebar` the drawer behaviour -- starts closed, slides in --
without giving that one any of the drawer's furniture. No id for the toggle to
resolve, no toggle, no close, no scrim: it sat at left:-280 with nothing in the
application able to open it. The close button and the scrim are partials now,
used by both, and the test that guards it *finds* sidebars by scanning the
templates rather than working from a list, which is exactly why this one was
missed.

The chat, measured at 390px, spent forty pixels of side padding and a
forty-four pixel avatar column before drawing a word -- close to a quarter of
the screen on margin, so anything that could not wrap had to be reached
sideways. Padding halved and the avatar moved above the turn; a code block
gained about sixty pixels.

Worse in the same row: `.topbar__actions` asked for 317px of a 390px bar,
because the control that used to give in that row is display:none below a
tablet width, so the group went rigid and the title -- flex: 1 -- was squeezed
to exactly zero. And `.btn--icon` sets a width with no `flex: none`, so the row
shrank the button instead of the text: the sidebar toggle measured eighteen
pixels across. The picker gives now, and shows its avatar rather than its name
on a phone.

Also the instrument, which lied twice more: it could not see horizontal
overflow at all, because `.shell` is overflow:hidden and its "is this
contained" test therefore answered yes for everything on the page; and run from
a copy it resolved `STATIC` to a directory that did not exist, rewrote every
asset URL to a dead file:// path and reported the whole application overflowing
by thirty thousand pixels. It resolves from the imported package now and
asserts that what it rewrote to is really there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-25 19:27:39 +00:00
20 changed files with 708 additions and 68 deletions
+54
View File
@@ -16,6 +16,60 @@ for 1.0.0 have something to be assembled from.
## Unreleased
## 1.2.0
- Fixed: **choosing a reasoning effort could kill the reply outright**, with a
Jinja traceback where the answer should have been. Reasoning effort is sent
two ways, and the second — `chat_template_kwargs` — is rendered into the
model's own chat template, which does not ignore a value it has never heard
of: it raises, and the whole request fails. The catch is that the vocabulary
is **not the same for every model**. gpt-oss takes `low/medium/high`; Bonsai
takes `low/medium/xhigh` and refuses `high`; OpenAI has added `minimal`,
`xhigh` and `max` at various points. This application offered the same three
to everything, so on some models the top setting was one the model would
throw for.
- **A model now has its own list of the efforts it accepts**, on its page under
Models, and the composer's picker and `/effort` offer only those. Tick none
and the familiar three are used, which is right for nearly everything.
- **And it corrects itself.** If an endpoint refuses an effort anyway — a model
swapped underneath a name, a runtime upgraded — that reply is retried once
without it instead of being lost, and the model's list is narrowed so the
menu stops offering something that does not work. Where the endpoint says
what it *does* take, that is what gets stored.
- `/effort` now reads the levels from the picker rather than from a second copy
of the list kept in the browser, so the two can no longer disagree about what
a valid effort is.
## 1.1.2
Two things a phone found that 1.1.0's phone pass had not.
- Fixed: **the administration area could not be navigated on a phone.** Admin
has a nav of its own rather than the chat sidebar, and 1.1.0 gave every
sidebar the drawer behaviour — starts closed, slides in — without giving that
one any of the drawer's furniture. So it sat off-screen with no button to open
it, no close, and nothing to tap beside it: every administration page was
reachable and then a dead end. It now opens, closes and dims the page like the
other one, and a test refuses any future sidebar that cannot be opened.
- Fixed: **the chat gave nearly a quarter of a phone screen to margins**, so
anything that could not wrap had to be scrolled to sideways. The thread's side
padding is halved, and the speaker's avatar moves above the turn instead of
sitting in a 44px column beside every line of it — a code block gained about
sixty pixels of readable width.
- Fixed: **the chat's title was squeezed to nothing.** The row's designated
shrinker is hidden below a tablet width, so on a phone the controls went rigid
and asked for 317 pixels of a 390 pixel bar; the heading was not truncated, it
simply stopped occupying space. The model picker gives now, and on a phone it
shows its avatar rather than its name — the name is one tap away and the
title is not.
- Tick boxes and the smaller buttons are big enough to hit on a phone. A
checkbox is drawn by the browser at about sixteen pixels whatever the type
around it, which made it the smallest target in the application by some way,
and the admin lists are mostly checkboxes.
- Fixed: **icon buttons could be squashed below their own size.** The sidebar
toggle measured eighteen pixels across on a phone, under half its target,
because a full row shrank the button rather than the text beside it.
## 1.1.1
One bug, and it is the one that made 1.1.0 look broken the moment you updated to
+35 -5
View File
@@ -29,10 +29,19 @@ import tempfile
from pathlib import Path
REPO = Path(__file__).resolve().parent.parent
SRC = REPO / "src"
sys.path.insert(0, str(SRC))
sys.path.insert(0, str(REPO / "src"))
STATIC = SRC / "lembas/web/static"
# Resolved from the package that actually got imported, not from where this
# file happens to sit. A copy of this script run from somewhere else silently
# pointed STATIC at a directory that did not exist, every asset URL was
# rewritten to a file:// path with nothing behind it, and the run measured an
# unstyled document -- reporting that every page in the application overflowed
# by thirty thousand pixels. The guard below only asked whether the URLs had
# been rewritten, which they had.
import lembas # noqa: E402
SRC = Path(lembas.__file__).resolve().parent.parent
STATIC = Path(lembas.__file__).resolve().parent / "web/static"
CHROMIUM = shutil.which("chromium") or shutil.which("chromium-browser")
# Routes that are served by the app rather than mounted, so the rewrite has to
@@ -131,8 +140,17 @@ window.__measure = function () {
tallCulprits: culprits('y'),
wideCulprits: culprits('x'),
/* The invariant: the application shell fills the window and the DOCUMENT
never scrolls. A document taller than the window is the /settings bug. */
documentScrolls: de.scrollHeight > window.innerHeight + 1,
never scrolls *for the reader*. A document taller than the window is the
/settings bug -- but only when the reader can actually move it. `overflow:
hidden` blocks a wheel and a finger while still permitting an assignment
to scrollTop, so a page whose shell clips a tall descendant reports a
scrollHeight of thousands and scrolls for nobody. /admin/prompts does
exactly that, and reading the raw height called it a bug four times. */
documentScrolls:
de.scrollHeight > window.innerHeight + 1 &&
["visible", "auto", "scroll"].indexOf(
getComputedStyle(document.documentElement).overflowY
) !== -1,
scrollsSideways: de.scrollWidth > window.innerWidth + 1,
smallTargets: small.slice(0, 40),
smallCount: small.length,
@@ -219,6 +237,18 @@ def rewrite(html: str, client, assets: Path) -> str:
f"{sorted(set(blocking))[:8]}"
)
# And that what they were rewritten *to* is really there. A rewrite that
# matches and produces a dead path is indistinguishable, from inside the
# browser, from no stylesheet at all -- and it is the failure that actually
# happened, twice.
missing = [
url
for url in re.findall(r'(?:href|src)="file://([^"?]+)"', html)
if not Path(url).exists()
]
if missing:
raise SystemExit(f"REWRITTEN TO NOTHING -- still an unstyled document: {missing[:5]}")
# The one-time notifications offer is a modal over the very page we came
# to measure, and it is gated on a localStorage key. Set it in the head, so
# it runs before the deferred script that reads it.
+1 -1
View File
@@ -1,3 +1,3 @@
"""LLeMbas - a Middle-earth themed web UI for OpenAI-compatible LLM endpoints."""
__version__ = "1.1.1"
__version__ = "1.2.0"
+16 -1
View File
@@ -177,7 +177,11 @@ async def model_detail(
"groups": list(db.scalars(select(Group).order_by(Group.name))),
"capabilities": PROTOCOL_CAPABILITIES,
"tool_capabilities": TOOL_CAPABILITIES,
# Every effort this application understands, so an administrator
# can tick the ones their model actually takes -- and the model's
# current answer, which is the common three until somebody says.
"efforts": chat_service.EFFORTS,
"model_efforts": chat_service.efforts_for(model),
# Rows predating the split have no tool_* keys at all. Showing them
# unticked would be a lie: tools.enabled_tools treats absent as on
# when `tools` is on, so that an upgrade does not silently take web
@@ -238,6 +242,7 @@ async def update_model(
position: str = Form(""),
context_length: str = Form(""),
default_effort: str = Form(""),
reasoning_efforts: list[str] = Form(default=[]),
group_ids: list[str] = Form(default=[]),
capability: list[str] = Form(default=[]),
) -> Response:
@@ -260,9 +265,19 @@ async def update_model(
# Merged rather than rebuilt, unlike the capabilities below: params_json
# holds whatever sampling defaults an administrator has set and this form
# only carries one of them.
# Which efforts this model takes at all. Submitted as a list of ticked
# values; empty means "nobody has said", and `chat.efforts_for` answers with
# the common three. Stored in the order `EFFORTS` declares rather than the
# order a browser happened to send.
chosen = [value for value in chat_service.EFFORTS if value in (reasoning_efforts or [])]
model.reasoning_efforts = chosen
params = dict(model.params_json or {})
wanted = default_effort.strip().lower()
if wanted in chat_service.EFFORTS:
# Checked against what this model takes, not against everything this
# application has heard of -- a default of `high` on a model whose template
# refuses it is a chat that fails on its first turn.
if wanted in chat_service.efforts_for(model):
params["reasoning_effort"] = wanted
else:
params.pop("reasoning_effort", None)
+6 -4
View File
@@ -82,10 +82,12 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
else []
),
"attached_base_ids": [base.id for base in chat.knowledge_bases] if chat else [],
# The three a reasoning model understands. From the service so the
# command, the control and the request builder cannot disagree about
# what is a valid effort.
"efforts": chat_service.EFFORTS,
# What *this* model takes, not the three every model used to be assumed
# to take. The vocabulary is per model -- gpt-oss has no `xhigh` and
# Bonsai has no `high`, and sending the wrong one does not degrade, it
# raises inside the chat template and fails the reply. From the service
# so the command, the control and the request builder cannot disagree.
"efforts": chat_service.efforts_for(current) if current else chat_service.DEFAULT_EFFORTS,
# What the picker shows, and what `build_request` will send. One
# resolver so the two cannot disagree.
"resolved_effort": chat_service.resolved_effort(chat) if chat else "",
+15 -1
View File
@@ -19,7 +19,7 @@ from sqlalchemy import (
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.types import JSONDict
from lembas.db.types import JSONDict, JSONList
if TYPE_CHECKING:
# Import only for the annotation; at runtime SQLAlchemy resolves the
@@ -137,6 +137,20 @@ class Model(UUIDPrimaryKey, Timestamps, Base):
# ticked anything.
context_length: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
# Which reasoning efforts this model actually accepts. Empty means "nobody
# has said", and `services/chat.efforts_for` answers with the common set.
#
# It has to be per model, because the vocabulary is: gpt-oss takes
# low/medium/high, Bonsai takes low/medium/xhigh and *raises* on high, and
# OpenAI's own list has grown minimal, xhigh and max at different times. A
# single global tuple is a guess that is wrong for somebody.
#
# ⚠ A column and not a key in `capabilities_json`, for exactly the reason
# `context_length` is one: that dict is rebuilt wholesale from the submitted
# checkboxes on every save, so anything in it that is not a checkbox is
# destroyed the next time an administrator ticks anything.
reasoning_efforts: Mapped[list[str]] = mapped_column(JSONList, default=list)
connection: Mapped[Connection] = relationship(back_populates="models")
groups: Mapped[list[Group]] = relationship(
"Group", secondary=model_groups, back_populates="models"
+58 -5
View File
@@ -387,7 +387,15 @@ def build_request(
):
body["tool_choice"] = {"type": "function", "function": {"name": force_tool}}
apply_effort(body, (chat.params_json or {}).get("reasoning_effort"))
# The model's own vocabulary, looked up here rather than passed in: every
# caller of `build_request` would otherwise have to remember, which is the
# trap `audio_service.template_flags` fell into.
chat_model = model_for(db, chat)
apply_effort(
body,
(chat.params_json or {}).get("reasoning_effort"),
efforts_for(chat_model) if chat_model is not None else None,
)
return body
@@ -405,7 +413,42 @@ def build_request(
# an effort on sends neither field and is byte-for-byte what it was. An endpoint
# strict about unknown parameters will refuse the extra one -- but on a chat
# somebody deliberately set an effort on, not on every chat in the instance.
EFFORTS = ("low", "medium", "high")
# Every reasoning effort this application understands, and the subset a model
# gets when nobody has said otherwise.
#
# 🚨 These are two different questions and conflating them is what broke a
# chat on Bonsai: `EFFORTS` was `("low", "medium", "high")` and was used both to
# validate what somebody chose *and* to decide what to offer, so a model whose
# vocabulary is low/medium/**xhigh** could not be given its own top setting,
# and the one it was given -- `high` -- made its chat template call
# `raise_exception` and took the whole reply with it.
#
# The known list is the union across providers, which have not agreed: OpenAI
# has added `minimal`, `xhigh` and `max` at different points; gpt-oss takes
# low/medium/high; Bonsai takes low/medium/xhigh and refuses high. `none` is
# deliberately absent -- this application already spells that `off`, and two
# spellings of off is the failure this codebase keeps cataloguing.
EFFORTS = ("minimal", "low", "medium", "high", "xhigh", "max")
# What a model is offered when its own list is empty. The three every reasoning
# model since the first one has understood.
DEFAULT_EFFORTS = ("low", "medium", "high")
def efforts_for(model) -> tuple[str, ...]:
"""The efforts this model accepts, in the order they should be offered.
A model's own list when an administrator has set one or the endpoint has
taught us one (see `generation._narrow_efforts`), and the common three
otherwise. Filtered against `EFFORTS` on the way out, so a value stored by
an older release -- or learned from an endpoint that advertised something
this application has never heard of -- cannot reach a request body.
"""
stored = list(getattr(model, "reasoning_efforts", None) or [])
chosen = [value for value in stored if value in EFFORTS]
if not chosen:
return DEFAULT_EFFORTS
return tuple(value for value in EFFORTS if value in chosen)
def resolved_effort(chat) -> str:
@@ -427,9 +470,19 @@ def resolved_effort(chat) -> str:
return value if value in EFFORTS else ""
def apply_effort(body: dict[str, Any], effort: str | None) -> None:
"""Put a chosen reasoning effort into a request body, in both forms."""
if not effort or effort not in EFFORTS:
def apply_effort(
body: dict[str, Any], effort: str | None, supported: tuple[str, ...] | None = None
) -> None:
"""Put a chosen reasoning effort into a request body, in both forms.
`supported` is the model's own vocabulary. An effort outside it is dropped
rather than sent, because the second form below is not advisory: it reaches
the model's Jinja chat template, and a template that does not know the value
raises rather than ignoring it -- which fails the whole request, not the
parameter.
"""
allowed = supported or DEFAULT_EFFORTS
if not effort or effort not in allowed:
return
body["reasoning_effort"] = effort
kwargs = dict(body.get("chat_template_kwargs") or {})
+118 -1
View File
@@ -19,6 +19,7 @@ import asyncio
import contextlib
import json
import logging
import re
import time
import uuid
from dataclasses import dataclass, field, replace
@@ -454,6 +455,122 @@ def _narrower(instance: float, quota: int) -> float:
return float(min(instance, quota))
# --- A reasoning effort the model will not take ------------------------------
#
# `chat_template_kwargs.reasoning_effort` is not advisory. It reaches the
# model's Jinja chat template, and a template that does not know the value does
# not ignore it -- gpt-oss and Bonsai both call `raise_exception`, which fails
# the whole request. The reader sees their reply die with a Jinja traceback in
# it, having chosen a perfectly ordinary-looking option from a menu this
# application drew.
#
# So the value is checked against the model's own vocabulary before it is sent
# (`chat.apply_effort`), and this is the second line: when it is refused anyway
# -- an endpoint upgraded underneath us, a model whose list nobody has set --
# the reply is retried once without it rather than lost, and the model's list is
# narrowed so the menu stops offering something that does not work.
def _effort_was_refused(message: str) -> bool:
"""Whether this error is the chat template refusing the effort we sent.
Deliberately narrow. Anything that merely mentions reasoning would also
match a model politely declining to think, and retrying *that* silently
would hide a real failure behind a second request.
"""
lowered = message.lower()
return "effort" in lowered and ("unexpected" in lowered or "supported" in lowered)
def _advertised_efforts(message: str) -> list[str]:
"""The efforts an error message says it will take, if it says.
Bonsai's is "Unexpected reasoning effort high. Supported types are xhigh
(default), medium, and low." -- which is the answer, written out, in the
failure. Read only from the part after "supported", so the *rejected* value
named in the first sentence is not collected as a supported one.
Best-effort by design: it only ever narrows what is offered, an
administrator can set the list by hand, and anything unrecognised is
dropped by `efforts_for` on the way out.
"""
lowered = message.lower()
if "supported" not in lowered:
return []
tail = lowered.split("supported", 1)[1]
# Whole words. `"high" in "xhigh"` is true, so a substring test reads
# Bonsai's "Supported types are xhigh (default), medium, and low" as
# advertising `high` -- the very value it has just refused -- and the list
# would learn the opposite of what the endpoint said.
words = set(re.findall(r"[a-z]+", tail))
return [effort for effort in chat_service.EFFORTS if effort in words]
def _learn_refused_effort(model_id: str, refused: str, message: str) -> None:
"""Write what the endpoint just taught us onto the model.
Its own session: this runs from inside a generation, which outlives the
request's session, and the whole point is that it survives to the next turn.
"""
from lembas.db.models import Model
if not model_id:
return
try:
with session_scope() as db:
models = list(db.scalars(select(Model).where(Model.model_id == model_id)))
for model in models:
advertised = _advertised_efforts(message)
current = list(model.reasoning_efforts or chat_service.DEFAULT_EFFORTS)
# What the endpoint advertised, when it did; otherwise simply
# the list it had, minus the one it has just refused.
wanted = advertised or [e for e in current if e != refused]
wanted = [e for e in wanted if e in chat_service.EFFORTS and e != refused]
if wanted and wanted != list(model.reasoning_efforts or []):
model.reasoning_efforts = wanted
log.info(
"model %s refused reasoning effort %r; efforts narrowed to %s",
model_id, refused, wanted,
)
except Exception: # noqa: BLE001 - never let bookkeeping fail a reply
log.exception("could not record the refused effort for model %s", model_id)
async def _stream_once(endpoint, payload, generation, model_id: str):
"""`stream_chat`, retried once without the reasoning effort if that is what
the endpoint objected to.
⚠ The retry is only safe because the template is rendered *before* any token
is produced, so a refusal arrives with nothing yet emitted. `sent` is the
guard that keeps it that way: once a single chunk has reached the caller,
the reply is under way and a second request would duplicate it.
"""
sent = False
try:
async for chunk in stream_chat(endpoint, payload):
sent = True
yield chunk
return
except LLMError as exc:
refused = str((payload.get("chat_template_kwargs") or {}).get("reasoning_effort") or "")
if sent or not refused or not _effort_was_refused(exc.message):
raise
log.info("retrying without reasoning effort %r: %s", refused, exc.message)
_learn_refused_effort(model_id, refused, exc.message)
retry = dict(payload)
retry.pop("reasoning_effort", None)
kwargs = dict(retry.get("chat_template_kwargs") or {})
kwargs.pop("reasoning_effort", None)
if kwargs:
retry["chat_template_kwargs"] = kwargs
else:
retry.pop("chat_template_kwargs", None)
async for chunk in stream_chat(endpoint, retry):
yield chunk
async def _run(generation: Generation) -> None:
"""Produce one reply, then persist it. Never raises into the task.
@@ -643,7 +760,7 @@ async def _run(generation: Generation) -> None:
# round thinks at all -- plenty of rounds do not.
round_thinking: tuple[float, float] | None = None
async for chunk in stream_chat(endpoint, payload):
async for chunk in _stream_once(endpoint, payload, generation, model_id):
counts = chunk_usage(chunk)
if counts is not None:
generation.reported_usage = True
+82 -3
View File
@@ -185,6 +185,12 @@ button, input, textarea, select {
/* Square, and the same height as everything beside it. */
.btn--icon {
width: var(--control-h);
/* Square, and it stays square. Without this a flex row that runs out of room
shrinks it instead of its neighbours -- the sidebar toggle measured 18px
across on a 390px chat, less than half the target it is supposed to be,
while the row beside it kept every pixel it had asked for. A control's
size is not the give in a layout; text is. */
flex: none;
padding: 0;
background: transparent;
border-color: transparent;
@@ -354,12 +360,28 @@ button, input, textarea, select {
}
.checkbox input {
accent-color: var(--accent);
width: 1rem;
height: 1rem;
width: var(--check-size);
height: var(--check-size);
flex: none;
cursor: pointer;
}
/* Every tick box, not only the ones inside a `.checkbox` label -- the admin
lists put bare ones in a row and those were 16px square on a phone. */
input[type="checkbox"],
input[type="radio"] {
accent-color: var(--accent);
width: var(--check-size);
height: var(--check-size);
}
/* Except the ones that are deliberately 1px: a visually-hidden radio is the
state behind a label, and the label is the target. */
input.visually-hidden[type="radio"],
input.visually-hidden[type="checkbox"] {
width: 1px;
height: 1px;
}
/* Multi-column form layout, one definition. */
.grid { display: grid; gap: var(--sp-4); }
.grid--2 { grid-template-columns: repeat(auto-fit, minmax(14rem, 1fr)); }
@@ -953,7 +975,48 @@ body.is-resizing .canvas__body { pointer-events: none; }
min-width: 0;
flex: 1;
}
.topbar__actions { display: flex; align-items: center; gap: var(--sp-2); flex: none; }
/*
The controls on the right of the topbar.
`flex: none` on the group with `min-width: 0` inside it: the group keeps the
width its controls need, and the one child whose width is a *name* rather
than a control -- the model picker -- is the thing allowed to give. Without
the second half the group asked for 317px of a 390px bar and the chat's
title, which is `flex: 1`, was squeezed to exactly zero: a heading that had
not been shortened or truncated but had simply ceased to occupy space.
*/
.topbar__actions {
display: flex;
align-items: center;
gap: var(--sp-2);
/* Allowed to give, which it was not. `--topbar__where` used to be the
designated shrinker in this row, and it is `display: none` below 64rem --
so on a phone the group became rigid, asked for 317px of a 390px bar, and
the title (`flex: 1`) was squeezed to exactly zero: a heading that had not
been truncated but had ceased to occupy space.
Nothing inside it shrinks except the model picker: every button here is
`flex: none` because a control's size is not the give in a layout. */
flex: 0 1 auto;
min-width: 0;
}
/* A title identifies the page, so it gets a floor and truncates rather than
disappearing. */
.topbar__title { min-width: 4rem; }
/* The one control in this row whose width is somebody else's decision -- a
model's label is whatever an administrator called it -- so it is the one
that gives, and it gives by truncating its name rather than its avatar or
its chevron. */
.topbar__actions .picker { min-width: 0; }
.topbar__actions .picker__button { max-width: 100%; }
.topbar__actions .picker__label {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.picker__avatar, .picker__chevron { flex: none; }
/*
Which machine an agent chat runs on, and where.
@@ -1243,6 +1306,22 @@ body.is-resizing .canvas__body { pointer-events: none; }
}
@media (max-width: 48rem) {
/* The bar is the densest row in the application and the one with the least
room: a toggle, a title, a model, and up to four panel buttons. Tighter
padding and a smaller gap buy back about 24px, which is the difference
between a title that truncates and one there is no room for at all. */
.topbar {
gap: var(--sp-2);
padding-right: max(var(--sp-2), var(--safe-right));
padding-left: max(var(--sp-2), var(--safe-left));
}
/* The model's name costs about a hundred pixels and its avatar does not,
and the picker opens onto a list of full names the moment it is touched.
So on a phone the avatar carries the identity and the chat's own title --
which nothing else on the screen tells you -- gets the room back. */
.topbar__actions .picker__label { display: none; }
.sidebar {
position: fixed;
inset: 0 auto 0 0;
+49
View File
@@ -1726,3 +1726,52 @@
.thread__intro > * { animation: intro-rise var(--dur-3) var(--ease-out) both; }
.thread__intro > *:nth-child(2) { animation-delay: 60ms; }
.thread__intro > *:nth-child(3) { animation-delay: 120ms; }
/*
--- A phone ----------------------------------------------------------------
The one width-aware block in this file, and the reason the blanket ban on
`@media` here was lifted: everything below is a *size*, and there is no
intrinsic-sizing trick that makes 24px of thread padding the right amount on
a 390px screen. The ban existed to stop the composer toolbar being "fixed"
with a breakpoint instead of by saying which child gives, and that guarantee
is asserted directly now (`tests/test_chat.py`) -- so this block may not touch
`.composer__toolbar` or `.composer__actions`, and a test refuses it if it
does.
What was wrong: a 390px screen spent 40px of its width on thread padding and
another 44 on the avatar gutter before a single word was drawn, which is
nearly a quarter of the screen given over to margin -- so anything that could
not wrap had to be scrolled to sideways.
*/
@media (max-width: 48rem) {
/* Half the horizontal padding. The vertical stays: it is what separates one
turn from the next, and turns are no closer together on a phone. */
.thread {
padding-left: var(--sp-3);
padding-right: var(--sp-3);
}
/* The avatar goes to the top of the turn rather than beside it, so the body
gets the whole width. The gutter is what identifies the speaker and it
still does; it simply stops costing 44px of every line. */
.msg {
grid-template-columns: 1fr;
gap: var(--sp-2);
}
.msg__gutter {
width: var(--control-h-sm);
height: var(--control-h-sm);
}
.msg__meta { gap: var(--sp-2); }
/* A bubble against the edge of the screen wants less inside it. */
.msg--user .msg__body--plain { padding: var(--sp-2) var(--sp-3); }
/* The composer is the other thing pressed against both edges. */
.composer { padding-left: var(--sp-2); padding-right: var(--sp-2); }
/* A hint that runs to four lines on a phone is a hint nobody reads, and it
sits directly under the thing a thumb is reaching for. */
.composer__hint { font-size: var(--text-xs); }
}
+11 -1
View File
@@ -145,6 +145,9 @@
Raising the token is the only version that reaches all of them, and it is
what `--control-h` exists for. */
--tap-min: 2.75rem;
/* A tick box, which does not take its size from `--control-h`: the browser
draws it and only `width`/`height` move it. */
--check-size: 1rem;
/* --- The window's own edges ---------------------------------------------
Installed on a phone, the page runs under the notch and the home
@@ -421,9 +424,16 @@
@media (pointer: coarse), (max-width: 48rem) {
:root {
--control-h: var(--tap-min);
--control-h-sm: 2.25rem;
/* 40px, not the 36 a comfortable pointer gets. A `.btn--sm` is a secondary
action, not an unimportant one -- Edit, Enable and Use default are all
`.btn--sm`, and on a phone they are the whole interaction. */
--control-h-sm: 2.5rem;
--control-px: var(--sp-4);
--control-px-sm: var(--sp-3);
/* A native checkbox is 13-16px whatever the surrounding type is, and no
amount of padding on its label changes the box itself. It is the
smallest target in the application on a phone by some margin. */
--check-size: 1.375rem;
}
}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 35 KiB

After

Width:  |  Height:  |  Size: 33 KiB

+25 -7
View File
@@ -271,8 +271,21 @@
/* --- Reasoning effort ---------------------------------------------------
The command drives the same select the composer shows, so there is one
piece of state and the control updates itself when the command is used. */
var EFFORTS = ["low", "medium", "high"];
piece of state and the control updates itself when the command is used.
Which efforts exist is read off that select's own options rather than
kept here. It used to be a second copy of `["low","medium","high"]`, which
was wrong the moment the vocabulary became per model: a Bonsai takes
`xhigh` and no `high`, so the list the server rendered and the list this
file believed in disagreed -- and the one that decides what `/effort xhigh`
does was this one. The select is the table; nothing else should hold it. */
function efforts() {
var select = el("[data-effort]");
if (!select) return [];
return Array.prototype.map
.call(select.options, function (option) { return option.value; })
.filter(function (value) { return value !== "off"; });
}
function setEffort(rest) {
var select = el("[data-effort]");
@@ -283,12 +296,14 @@
"error"
);
}
var available = efforts();
var listed = available.join(", ");
var wanted = (rest || "").trim().toLowerCase();
if (!wanted) {
return note(
EFFORTS.indexOf(select.value) === -1
? "No effort is being sent. Try low, medium or high."
: "Effort is " + select.value + ". /effort low, medium, high, or off."
available.indexOf(select.value) === -1
? "No effort is being sent. Try " + listed + "."
: "Effort is " + select.value + ". /effort " + listed + ", or off."
);
}
/* "off" is the option's real value, not an empty string: the new-chat form
@@ -296,8 +311,11 @@
sentinel and this has to match it. "default" and "none" still work,
because somebody's fingers will type them. */
if (wanted === "default" || wanted === "none") wanted = "off";
else if (wanted !== "off" && EFFORTS.indexOf(wanted) === -1) {
return note("“" + wanted + "” is not an effort. Try low, medium, high or off.", "error");
else if (wanted !== "off" && available.indexOf(wanted) === -1) {
return note(
"“" + wanted + "” is not an effort this model takes. Try " + listed + " or off.",
"error"
);
}
select.value = wanted;
select.dispatchEvent(new Event("change", { bubbles: true }));
+17 -4
View File
@@ -14,10 +14,20 @@
{% block body %}
<div class="shell">
<aside class="sidebar">
<div class="sidebar__header">
{{ brandlink(uid="admin") }}
</div>
{#
`id="sidebar"` and the drawer's furniture, because below the phone
breakpoint `.sidebar` is a fixed overlay that starts closed -- and this one
had neither an id for `data-toggle="#sidebar"` to find nor any control to
open it. The administration area was reachable on a phone and then
unnavigable once you arrived.
#}
<aside class="sidebar" id="sidebar">
<header class="sidebar__header">
<div class="sidebar__brand-slot">
{{ brandlink(uid="admin") }}
</div>
{% include "partials/_sidebar_close.html" %}
</header>
<nav class="sidebar__scroll" aria-label="Administration">
<div class="nav-group">
@@ -106,8 +116,11 @@
</div>
</aside>
{% include "partials/_sidebar_scrim.html" %}
<main class="main">
<header class="topbar">
{% include "partials/_sidebar_toggle.html" %}
<h1 class="topbar__title">{% block heading %}Administration{% endblock %}</h1>
<button class="btn btn--icon" type="button" data-theme-toggle aria-label="Switch theme">
<span class="theme-icon theme-icon--dark">{{ icon("moon") }}</span>
@@ -104,11 +104,39 @@
</p>
</div>
<div class="field">
<span class="field__label">Reasoning efforts this model accepts</span>
<div class="btn-row">
{% for value in efforts %}
<label class="checkbox">
<input type="checkbox" name="reasoning_efforts" value="{{ value }}"
{{ 'checked' if value in model_efforts }}>
<span class="mono">{{ value }}</span>
</label>
{% endfor %}
</div>
<p class="field__hint">
The vocabulary is <strong>not the same for every model</strong>, and
sending one a model does not know is not ignored — it is rendered into
the model's chat template, which raises and fails the whole reply.
gpt-oss takes <span class="mono">low/medium/high</span>; Bonsai takes
<span class="mono">low/medium/xhigh</span> and refuses
<span class="mono">high</span>; OpenAI has added
<span class="mono">minimal</span>, <span class="mono">xhigh</span> and
<span class="mono">max</span> at various points.
<br>
Tick none and the common three are offered, which is right for almost
everything. If an endpoint ever refuses one anyway, that reply is
retried without it and this list corrects itself — so this is worth
setting by hand only to save that one round trip.
</p>
</div>
<div class="field">
<label class="field__label" for="default-effort">Default reasoning effort</label>
<select class="select" id="default-effort" name="default_effort">
<option value="">None — send nothing</option>
{% for value in efforts %}
{% for value in model_efforts %}
<option value="{{ value }}"
{{ 'selected' if model.params_json.get('reasoning_effort') == value }}>
{{ value }}
@@ -0,0 +1,20 @@
{% from "_macros.html" import icon %}
{#
The way out of the drawer, and the reason it is *inside* it.
Below the phone breakpoint the sidebar is a fixed overlay and the toggle that
opens it is in the topbar underneath -- so once open, the control for closing
it is behind it. Its own partial because there are two sidebars in this
application, the chat one and the admin one, and the second was given the
drawer behaviour without the drawer's furniture: at a phone width it was
hidden off-screen with no toggle and no close anywhere, which is an admin area
that simply could not be navigated on a phone.
Hidden above that breakpoint, where the sidebar is an ordinary column.
#}
<div class="sidebar__actions-rail">
<button class="btn btn--icon sidebar__close" type="button"
aria-label="Close sidebar" data-toggle="#sidebar">
{{ icon("x") }}
</button>
</div>
@@ -0,0 +1,10 @@
{#
The scrim behind an open drawer. It carries the same `data-toggle` as every
other control that closes it, so tapping beside the drawer goes through one
code path rather than a second written for touch.
Rendered always and shown by CSS: it exists only below the breakpoint and only
while the drawer is open, which is a question about width and state that the
server cannot answer and the stylesheet can.
#}
<div class="sidebar-scrim" data-toggle="#sidebar" aria-hidden="true"></div>
+2 -31
View File
@@ -27,27 +27,7 @@
{{ brandlink(uid="side") }}
</div>
{#
The way out, and the reason it is *inside* the drawer.
Below the phone breakpoint this whole element is a fixed overlay, and the
toggle that opens it lives in the topbar underneath -- so once it was
open, the control for closing it was behind it. That was true on /chat,
where at least a toggle existed; on the seven other pages that carry this
sidebar there was no such control at all, and no way back.
Hidden above that breakpoint, where the sidebar is an ordinary column and
the topbar's toggle is perfectly visible. It shipped *visible* on the
desktop in 1.1.0 -- not because this rule was wrong, but because the
browser was still drawing the page with the previous release's
stylesheet; see `templating.asset`.
#}
<div class="sidebar__actions-rail">
<button class="btn btn--icon sidebar__close" type="button"
aria-label="Close sidebar" data-toggle="#sidebar">
{{ icon("x") }}
</button>
</div>
{% include "partials/_sidebar_close.html" %}
</header>
{% include "partials/_sidebar_actions.html" %}
@@ -126,13 +106,4 @@
</div>
</aside>
{#
The scrim behind the open drawer. It carries the same `data-toggle` as every
other control that closes it, so tapping beside the drawer closes it through
exactly one code path rather than a second one written for touch.
Rendered always and shown by CSS: it exists only below the breakpoint and only
while the drawer is open, which is a question about width and state that the
server cannot answer and the stylesheet can.
#}
<div class="sidebar-scrim" data-toggle="#sidebar" aria-hidden="true"></div>
{% include "partials/_sidebar_scrim.html" %}
+90
View File
@@ -366,3 +366,93 @@ def test_the_picker_never_says_default(client: TestClient, db, registered):
assert "Effort: default" not in html
assert "Effort: off" in html
assert '<option value="medium" selected>' in html.replace("\n", "").replace(" ", "")
# --- A vocabulary that is not the same for every model -----------------------
#
# Reported from a real instance, on a model called Bonsai:
#
# Jinja Exception: Unexpected reasoning effort high. Supported types are
# xhigh (default), medium, and low.
#
# `chat_template_kwargs.reasoning_effort` is rendered into the model's own chat
# template, and a template that does not know the value calls `raise_exception`
# rather than ignoring it -- so the whole reply died, from an option this
# application had drawn in a menu.
BONSAI_ERROR = (
"Jinja Exception: Unexpected reasoning effort high. "
"Supported types are xhigh (default), medium, and low."
)
class _FakeModel:
def __init__(self, efforts=None):
self.reasoning_efforts = efforts or []
def test_a_model_that_has_said_nothing_gets_the_common_three():
from lembas.services import chat as chat_service
assert chat_service.efforts_for(_FakeModel()) == ("low", "medium", "high")
def test_a_model_can_take_xhigh_and_not_high():
from lembas.services import chat as chat_service
bonsai = _FakeModel(["xhigh", "medium", "low"])
assert chat_service.efforts_for(bonsai) == ("low", "medium", "xhigh")
assert "high" not in chat_service.efforts_for(bonsai)
def test_an_effort_the_model_refuses_is_never_sent():
"""The check that stops the crash happening at all."""
from lembas.services import chat as chat_service
supported = chat_service.efforts_for(_FakeModel(["xhigh", "medium", "low"]))
body: dict = {}
chat_service.apply_effort(body, "high", supported)
assert body == {}
chat_service.apply_effort(body, "xhigh", supported)
assert body["reasoning_effort"] == "xhigh"
assert body["chat_template_kwargs"]["reasoning_effort"] == "xhigh"
def test_a_value_this_application_never_heard_of_cannot_reach_a_request():
from lembas.services import chat as chat_service
assert chat_service.efforts_for(_FakeModel(["ludicrous"])) == ("low", "medium", "high")
def test_the_refusal_is_recognised_and_the_supported_list_read_out_of_it():
from lembas.services import generation
assert generation._effort_was_refused(BONSAI_ERROR)
assert generation._advertised_efforts(BONSAI_ERROR) == ["low", "medium", "xhigh"]
def test_the_rejected_value_is_not_collected_as_a_supported_one():
"""The message names the refused effort first and the supported ones after,
so anything reading the whole string would learn `high` from a sentence
saying `high` is the problem."""
from lembas.services import generation
assert "high" not in generation._advertised_efforts(BONSAI_ERROR)
def test_an_ordinary_failure_is_not_retried_as_an_effort_problem():
"""Retrying a genuine failure would hide it behind a second request."""
from lembas.services import generation
for message in (
"Connection refused.",
"The model is still loading.",
"context length exceeded",
):
assert not generation._effort_was_refused(message)
def test_a_model_with_no_advertisement_simply_loses_the_refused_value():
from lembas.services import generation
assert generation._advertised_efforts("Unexpected reasoning effort high.") == []
+70 -3
View File
@@ -95,9 +95,12 @@ def test_inert_is_never_left_behind_on_a_widened_window():
def test_the_drawer_is_dismissible_without_finding_a_button():
sidebar = (TEMPLATES / "partials/sidebar.html").read_text(encoding="utf-8")
assert 'class="sidebar-scrim"' in sidebar
assert 'data-toggle="#sidebar"' in sidebar
"""The scrim is a partial because there are two sidebars, so the markup is
asserted where it is defined and its *inclusion* is asserted per sidebar by
`test_every_sidebar_carries_the_way_out_and_the_scrim`."""
scrim = (TEMPLATES / "partials/_sidebar_scrim.html").read_text(encoding="utf-8")
assert 'class="sidebar-scrim"' in scrim
assert 'data-toggle="#sidebar"' in scrim
assert ".sidebar-scrim" in APP_CSS
@@ -113,3 +116,67 @@ def test_the_toggle_is_a_real_target(client: TestClient, registered):
tokens = (ROOT / "web/static/css/tokens.css").read_text(encoding="utf-8")
assert "--tap-min: 2.75rem" in tokens
assert "--control-h: var(--tap-min)" in tokens
# --- Any sidebar, not only the one this was written for ----------------------
def _sidebar_templates() -> list[str]:
"""Every template that renders a sidebar of its own, found rather than
listed -- the admin one was missed precisely because it was not on a list."""
return [
str(p.relative_to(TEMPLATES))
for p in TEMPLATES.rglob("*.html")
if '<aside class="sidebar"' in p.read_text(encoding="utf-8")
]
def test_every_sidebar_is_one_the_toggle_can_find():
"""`data-toggle="#sidebar"` resolves by id, and below the phone breakpoint
`.sidebar` is a fixed overlay that starts closed. A sidebar without that id
is one nothing can open: the admin area shipped that way in 1.1.0 and 1.1.1
-- reachable on a phone, and unnavigable the moment you arrived."""
without = [
name
for name in _sidebar_templates()
if '<aside class="sidebar" id="sidebar"' not in (TEMPLATES / name).read_text(
encoding="utf-8"
)
]
assert not without, f"sidebar with no id, so nothing can open it: {without}"
def test_every_sidebar_carries_the_way_out_and_the_scrim():
missing = []
for name in _sidebar_templates():
text = (TEMPLATES / name).read_text(encoding="utf-8")
if "partials/_sidebar_close.html" not in text:
missing.append(f"{name}: no close button")
if "partials/_sidebar_scrim.html" not in text:
missing.append(f"{name}: no scrim")
assert not missing, missing
def test_the_admin_area_can_be_navigated_on_a_phone(client: TestClient, registered):
"""The whole of administration is in that nav and nowhere else."""
page = client.get("/admin/models").text
assert '<aside class="sidebar" id="sidebar"' in page
assert 'data-toggle="#sidebar"' in page
assert "sidebar-scrim" in page
# --- Controls do not shrink below their own size -----------------------------
def test_an_icon_button_keeps_its_size_in_a_tight_row():
"""`.btn--icon` sets a width and, without `flex: none`, a row that runs out
of room shrinks it instead of the text beside it -- the sidebar toggle
measured 18px across on a 390px chat, well under half its target."""
rule = APP_CSS.split(".btn--icon {", 1)[1].split("}", 1)[0]
assert "flex: none" in rule
def test_the_topbar_can_give_somewhere(client: TestClient, registered):
"""`.topbar__where` was the designated shrinker in that row and it is
`display: none` below 64rem, so on a phone the group went rigid and the
title -- which is `flex: 1` -- was squeezed to exactly zero width."""
actions = APP_CSS.split(".topbar__actions {", 1)[1].split("}", 1)[0]
assert "flex: 0 1 auto" in actions
assert "min-width: 0" in actions
assert "min-width" in APP_CSS.split(".topbar__title {", 1)[1].split("}", 1)[0]