An effort the model had never heard of

Reported from a live instance, on Bonsai:

  Jinja Exception: Unexpected reasoning effort high. Supported types are
  xhigh (default), medium, and low.

Effort goes out two ways because no single field works, and the second --
chat_template_kwargs -- is not a parameter the server interprets. It is
rendered into the model's own chat template, which does not ignore a value it
does not know: it calls raise_exception, and the request dies before a token.
So a perfectly ordinary option, drawn by this application in its own menu, took
the whole reply with it.

The vocabulary is per model and nobody agrees. gpt-oss takes low/medium/high.
Bonsai takes low/medium/xhigh and refuses high. OpenAI has added minimal, xhigh
and max at different points, and which of them a given model accepts varies
again. One global tuple was going to be wrong for somebody whatever it held.

A model carries its own list now, and the picker, the slash command and the
request builder all read it. A column rather than a key in capabilities_json,
for the reason context_length is one: that dict is rebuilt wholesale from the
submitted checkboxes on every save.

And it corrects itself. A refusal retries the reply once without the effort
rather than losing it -- safe only because the template renders before any
token, so nothing has been emitted, and there is a guard that keeps it that way
-- then narrows the model's list. Bonsai's error states what it does take, so
that is what gets stored.

Note the parser bug, because it is a good one: "high" is a substring of
"xhigh", so reading the advertised list by substring learned `high` from a
sentence explaining that `high` is the problem. Whole words now, with a test
named after it.

/effort reads its levels off the picker instead of a second copy of the list
kept in the browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 21:20:06 +00:00
co-authored by Claude Opus 5
parent b7bf7d728b
commit 32e2326d41
10 changed files with 382 additions and 21 deletions
+118 -1
View File
@@ -19,6 +19,7 @@ import asyncio
import contextlib
import json
import logging
import re
import time
import uuid
from dataclasses import dataclass, field, replace
@@ -454,6 +455,122 @@ def _narrower(instance: float, quota: int) -> float:
return float(min(instance, quota))
# --- A reasoning effort the model will not take ------------------------------
#
# `chat_template_kwargs.reasoning_effort` is not advisory. It reaches the
# model's Jinja chat template, and a template that does not know the value does
# not ignore it -- gpt-oss and Bonsai both call `raise_exception`, which fails
# the whole request. The reader sees their reply die with a Jinja traceback in
# it, having chosen a perfectly ordinary-looking option from a menu this
# application drew.
#
# So the value is checked against the model's own vocabulary before it is sent
# (`chat.apply_effort`), and this is the second line: when it is refused anyway
# -- an endpoint upgraded underneath us, a model whose list nobody has set --
# the reply is retried once without it rather than lost, and the model's list is
# narrowed so the menu stops offering something that does not work.
def _effort_was_refused(message: str) -> bool:
"""Whether this error is the chat template refusing the effort we sent.
Deliberately narrow. Anything that merely mentions reasoning would also
match a model politely declining to think, and retrying *that* silently
would hide a real failure behind a second request.
"""
lowered = message.lower()
return "effort" in lowered and ("unexpected" in lowered or "supported" in lowered)
def _advertised_efforts(message: str) -> list[str]:
"""The efforts an error message says it will take, if it says.
Bonsai's is "Unexpected reasoning effort high. Supported types are xhigh
(default), medium, and low." -- which is the answer, written out, in the
failure. Read only from the part after "supported", so the *rejected* value
named in the first sentence is not collected as a supported one.
Best-effort by design: it only ever narrows what is offered, an
administrator can set the list by hand, and anything unrecognised is
dropped by `efforts_for` on the way out.
"""
lowered = message.lower()
if "supported" not in lowered:
return []
tail = lowered.split("supported", 1)[1]
# Whole words. `"high" in "xhigh"` is true, so a substring test reads
# Bonsai's "Supported types are xhigh (default), medium, and low" as
# advertising `high` -- the very value it has just refused -- and the list
# would learn the opposite of what the endpoint said.
words = set(re.findall(r"[a-z]+", tail))
return [effort for effort in chat_service.EFFORTS if effort in words]
def _learn_refused_effort(model_id: str, refused: str, message: str) -> None:
"""Write what the endpoint just taught us onto the model.
Its own session: this runs from inside a generation, which outlives the
request's session, and the whole point is that it survives to the next turn.
"""
from lembas.db.models import Model
if not model_id:
return
try:
with session_scope() as db:
models = list(db.scalars(select(Model).where(Model.model_id == model_id)))
for model in models:
advertised = _advertised_efforts(message)
current = list(model.reasoning_efforts or chat_service.DEFAULT_EFFORTS)
# What the endpoint advertised, when it did; otherwise simply
# the list it had, minus the one it has just refused.
wanted = advertised or [e for e in current if e != refused]
wanted = [e for e in wanted if e in chat_service.EFFORTS and e != refused]
if wanted and wanted != list(model.reasoning_efforts or []):
model.reasoning_efforts = wanted
log.info(
"model %s refused reasoning effort %r; efforts narrowed to %s",
model_id, refused, wanted,
)
except Exception: # noqa: BLE001 - never let bookkeeping fail a reply
log.exception("could not record the refused effort for model %s", model_id)
async def _stream_once(endpoint, payload, generation, model_id: str):
"""`stream_chat`, retried once without the reasoning effort if that is what
the endpoint objected to.
⚠ The retry is only safe because the template is rendered *before* any token
is produced, so a refusal arrives with nothing yet emitted. `sent` is the
guard that keeps it that way: once a single chunk has reached the caller,
the reply is under way and a second request would duplicate it.
"""
sent = False
try:
async for chunk in stream_chat(endpoint, payload):
sent = True
yield chunk
return
except LLMError as exc:
refused = str((payload.get("chat_template_kwargs") or {}).get("reasoning_effort") or "")
if sent or not refused or not _effort_was_refused(exc.message):
raise
log.info("retrying without reasoning effort %r: %s", refused, exc.message)
_learn_refused_effort(model_id, refused, exc.message)
retry = dict(payload)
retry.pop("reasoning_effort", None)
kwargs = dict(retry.get("chat_template_kwargs") or {})
kwargs.pop("reasoning_effort", None)
if kwargs:
retry["chat_template_kwargs"] = kwargs
else:
retry.pop("chat_template_kwargs", None)
async for chunk in stream_chat(endpoint, retry):
yield chunk
async def _run(generation: Generation) -> None:
"""Produce one reply, then persist it. Never raises into the task.
@@ -643,7 +760,7 @@ async def _run(generation: Generation) -> None:
# round thinks at all -- plenty of rounds do not.
round_thinking: tuple[float, float] | None = None
async for chunk in stream_chat(endpoint, payload):
async for chunk in _stream_once(endpoint, payload, generation, model_id):
counts = chunk_usage(chunk)
if counts is not None:
generation.reported_usage = True