An effort the model had never heard of
Reported from a live instance, on Bonsai: Jinja Exception: Unexpected reasoning effort high. Supported types are xhigh (default), medium, and low. Effort goes out two ways because no single field works, and the second -- chat_template_kwargs -- is not a parameter the server interprets. It is rendered into the model's own chat template, which does not ignore a value it does not know: it calls raise_exception, and the request dies before a token. So a perfectly ordinary option, drawn by this application in its own menu, took the whole reply with it. The vocabulary is per model and nobody agrees. gpt-oss takes low/medium/high. Bonsai takes low/medium/xhigh and refuses high. OpenAI has added minimal, xhigh and max at different points, and which of them a given model accepts varies again. One global tuple was going to be wrong for somebody whatever it held. A model carries its own list now, and the picker, the slash command and the request builder all read it. A column rather than a key in capabilities_json, for the reason context_length is one: that dict is rebuilt wholesale from the submitted checkboxes on every save. And it corrects itself. A refusal retries the reply once without the effort rather than losing it -- safe only because the template renders before any token, so nothing has been emitted, and there is a guard that keeps it that way -- then narrows the model's list. Bonsai's error states what it does take, so that is what gets stored. Note the parser bug, because it is a good one: "high" is a substring of "xhigh", so reading the advertised list by substring learned `high` from a sentence explaining that `high` is the problem. Whole words now, with a test named after it. /effort reads its levels off the picker instead of a second copy of the list kept in the browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -387,7 +387,15 @@ def build_request(
|
||||
):
|
||||
body["tool_choice"] = {"type": "function", "function": {"name": force_tool}}
|
||||
|
||||
apply_effort(body, (chat.params_json or {}).get("reasoning_effort"))
|
||||
# The model's own vocabulary, looked up here rather than passed in: every
|
||||
# caller of `build_request` would otherwise have to remember, which is the
|
||||
# trap `audio_service.template_flags` fell into.
|
||||
chat_model = model_for(db, chat)
|
||||
apply_effort(
|
||||
body,
|
||||
(chat.params_json or {}).get("reasoning_effort"),
|
||||
efforts_for(chat_model) if chat_model is not None else None,
|
||||
)
|
||||
return body
|
||||
|
||||
|
||||
@@ -405,7 +413,42 @@ def build_request(
|
||||
# an effort on sends neither field and is byte-for-byte what it was. An endpoint
|
||||
# strict about unknown parameters will refuse the extra one -- but on a chat
|
||||
# somebody deliberately set an effort on, not on every chat in the instance.
|
||||
EFFORTS = ("low", "medium", "high")
|
||||
# Every reasoning effort this application understands, and the subset a model
|
||||
# gets when nobody has said otherwise.
|
||||
#
|
||||
# 🚨 These are two different questions and conflating them is what broke a
|
||||
# chat on Bonsai: `EFFORTS` was `("low", "medium", "high")` and was used both to
|
||||
# validate what somebody chose *and* to decide what to offer, so a model whose
|
||||
# vocabulary is low/medium/**xhigh** could not be given its own top setting,
|
||||
# and the one it was given -- `high` -- made its chat template call
|
||||
# `raise_exception` and took the whole reply with it.
|
||||
#
|
||||
# The known list is the union across providers, which have not agreed: OpenAI
|
||||
# has added `minimal`, `xhigh` and `max` at different points; gpt-oss takes
|
||||
# low/medium/high; Bonsai takes low/medium/xhigh and refuses high. `none` is
|
||||
# deliberately absent -- this application already spells that `off`, and two
|
||||
# spellings of off is the failure this codebase keeps cataloguing.
|
||||
EFFORTS = ("minimal", "low", "medium", "high", "xhigh", "max")
|
||||
|
||||
# What a model is offered when its own list is empty. The three every reasoning
|
||||
# model since the first one has understood.
|
||||
DEFAULT_EFFORTS = ("low", "medium", "high")
|
||||
|
||||
|
||||
def efforts_for(model) -> tuple[str, ...]:
|
||||
"""The efforts this model accepts, in the order they should be offered.
|
||||
|
||||
A model's own list when an administrator has set one or the endpoint has
|
||||
taught us one (see `generation._narrow_efforts`), and the common three
|
||||
otherwise. Filtered against `EFFORTS` on the way out, so a value stored by
|
||||
an older release -- or learned from an endpoint that advertised something
|
||||
this application has never heard of -- cannot reach a request body.
|
||||
"""
|
||||
stored = list(getattr(model, "reasoning_efforts", None) or [])
|
||||
chosen = [value for value in stored if value in EFFORTS]
|
||||
if not chosen:
|
||||
return DEFAULT_EFFORTS
|
||||
return tuple(value for value in EFFORTS if value in chosen)
|
||||
|
||||
|
||||
def resolved_effort(chat) -> str:
|
||||
@@ -427,9 +470,19 @@ def resolved_effort(chat) -> str:
|
||||
return value if value in EFFORTS else ""
|
||||
|
||||
|
||||
def apply_effort(body: dict[str, Any], effort: str | None) -> None:
|
||||
"""Put a chosen reasoning effort into a request body, in both forms."""
|
||||
if not effort or effort not in EFFORTS:
|
||||
def apply_effort(
|
||||
body: dict[str, Any], effort: str | None, supported: tuple[str, ...] | None = None
|
||||
) -> None:
|
||||
"""Put a chosen reasoning effort into a request body, in both forms.
|
||||
|
||||
`supported` is the model's own vocabulary. An effort outside it is dropped
|
||||
rather than sent, because the second form below is not advisory: it reaches
|
||||
the model's Jinja chat template, and a template that does not know the value
|
||||
raises rather than ignoring it -- which fails the whole request, not the
|
||||
parameter.
|
||||
"""
|
||||
allowed = supported or DEFAULT_EFFORTS
|
||||
if not effort or effort not in allowed:
|
||||
return
|
||||
body["reasoning_effort"] = effort
|
||||
kwargs = dict(body.get("chat_template_kwargs") or {})
|
||||
|
||||
@@ -19,6 +19,7 @@ import asyncio
|
||||
import contextlib
|
||||
import json
|
||||
import logging
|
||||
import re
|
||||
import time
|
||||
import uuid
|
||||
from dataclasses import dataclass, field, replace
|
||||
@@ -454,6 +455,122 @@ def _narrower(instance: float, quota: int) -> float:
|
||||
return float(min(instance, quota))
|
||||
|
||||
|
||||
# --- A reasoning effort the model will not take ------------------------------
|
||||
#
|
||||
# `chat_template_kwargs.reasoning_effort` is not advisory. It reaches the
|
||||
# model's Jinja chat template, and a template that does not know the value does
|
||||
# not ignore it -- gpt-oss and Bonsai both call `raise_exception`, which fails
|
||||
# the whole request. The reader sees their reply die with a Jinja traceback in
|
||||
# it, having chosen a perfectly ordinary-looking option from a menu this
|
||||
# application drew.
|
||||
#
|
||||
# So the value is checked against the model's own vocabulary before it is sent
|
||||
# (`chat.apply_effort`), and this is the second line: when it is refused anyway
|
||||
# -- an endpoint upgraded underneath us, a model whose list nobody has set --
|
||||
# the reply is retried once without it rather than lost, and the model's list is
|
||||
# narrowed so the menu stops offering something that does not work.
|
||||
|
||||
|
||||
def _effort_was_refused(message: str) -> bool:
|
||||
"""Whether this error is the chat template refusing the effort we sent.
|
||||
|
||||
Deliberately narrow. Anything that merely mentions reasoning would also
|
||||
match a model politely declining to think, and retrying *that* silently
|
||||
would hide a real failure behind a second request.
|
||||
"""
|
||||
lowered = message.lower()
|
||||
return "effort" in lowered and ("unexpected" in lowered or "supported" in lowered)
|
||||
|
||||
|
||||
def _advertised_efforts(message: str) -> list[str]:
|
||||
"""The efforts an error message says it will take, if it says.
|
||||
|
||||
Bonsai's is "Unexpected reasoning effort high. Supported types are xhigh
|
||||
(default), medium, and low." -- which is the answer, written out, in the
|
||||
failure. Read only from the part after "supported", so the *rejected* value
|
||||
named in the first sentence is not collected as a supported one.
|
||||
|
||||
Best-effort by design: it only ever narrows what is offered, an
|
||||
administrator can set the list by hand, and anything unrecognised is
|
||||
dropped by `efforts_for` on the way out.
|
||||
"""
|
||||
lowered = message.lower()
|
||||
if "supported" not in lowered:
|
||||
return []
|
||||
tail = lowered.split("supported", 1)[1]
|
||||
# Whole words. `"high" in "xhigh"` is true, so a substring test reads
|
||||
# Bonsai's "Supported types are xhigh (default), medium, and low" as
|
||||
# advertising `high` -- the very value it has just refused -- and the list
|
||||
# would learn the opposite of what the endpoint said.
|
||||
words = set(re.findall(r"[a-z]+", tail))
|
||||
return [effort for effort in chat_service.EFFORTS if effort in words]
|
||||
|
||||
|
||||
def _learn_refused_effort(model_id: str, refused: str, message: str) -> None:
|
||||
"""Write what the endpoint just taught us onto the model.
|
||||
|
||||
Its own session: this runs from inside a generation, which outlives the
|
||||
request's session, and the whole point is that it survives to the next turn.
|
||||
"""
|
||||
from lembas.db.models import Model
|
||||
|
||||
if not model_id:
|
||||
return
|
||||
try:
|
||||
with session_scope() as db:
|
||||
models = list(db.scalars(select(Model).where(Model.model_id == model_id)))
|
||||
for model in models:
|
||||
advertised = _advertised_efforts(message)
|
||||
current = list(model.reasoning_efforts or chat_service.DEFAULT_EFFORTS)
|
||||
# What the endpoint advertised, when it did; otherwise simply
|
||||
# the list it had, minus the one it has just refused.
|
||||
wanted = advertised or [e for e in current if e != refused]
|
||||
wanted = [e for e in wanted if e in chat_service.EFFORTS and e != refused]
|
||||
if wanted and wanted != list(model.reasoning_efforts or []):
|
||||
model.reasoning_efforts = wanted
|
||||
log.info(
|
||||
"model %s refused reasoning effort %r; efforts narrowed to %s",
|
||||
model_id, refused, wanted,
|
||||
)
|
||||
except Exception: # noqa: BLE001 - never let bookkeeping fail a reply
|
||||
log.exception("could not record the refused effort for model %s", model_id)
|
||||
|
||||
|
||||
async def _stream_once(endpoint, payload, generation, model_id: str):
|
||||
"""`stream_chat`, retried once without the reasoning effort if that is what
|
||||
the endpoint objected to.
|
||||
|
||||
⚠ The retry is only safe because the template is rendered *before* any token
|
||||
is produced, so a refusal arrives with nothing yet emitted. `sent` is the
|
||||
guard that keeps it that way: once a single chunk has reached the caller,
|
||||
the reply is under way and a second request would duplicate it.
|
||||
"""
|
||||
sent = False
|
||||
try:
|
||||
async for chunk in stream_chat(endpoint, payload):
|
||||
sent = True
|
||||
yield chunk
|
||||
return
|
||||
except LLMError as exc:
|
||||
refused = str((payload.get("chat_template_kwargs") or {}).get("reasoning_effort") or "")
|
||||
if sent or not refused or not _effort_was_refused(exc.message):
|
||||
raise
|
||||
log.info("retrying without reasoning effort %r: %s", refused, exc.message)
|
||||
_learn_refused_effort(model_id, refused, exc.message)
|
||||
|
||||
retry = dict(payload)
|
||||
retry.pop("reasoning_effort", None)
|
||||
kwargs = dict(retry.get("chat_template_kwargs") or {})
|
||||
kwargs.pop("reasoning_effort", None)
|
||||
if kwargs:
|
||||
retry["chat_template_kwargs"] = kwargs
|
||||
else:
|
||||
retry.pop("chat_template_kwargs", None)
|
||||
|
||||
async for chunk in stream_chat(endpoint, retry):
|
||||
yield chunk
|
||||
|
||||
|
||||
async def _run(generation: Generation) -> None:
|
||||
"""Produce one reply, then persist it. Never raises into the task.
|
||||
|
||||
@@ -643,7 +760,7 @@ async def _run(generation: Generation) -> None:
|
||||
# round thinks at all -- plenty of rounds do not.
|
||||
round_thinking: tuple[float, float] | None = None
|
||||
|
||||
async for chunk in stream_chat(endpoint, payload):
|
||||
async for chunk in _stream_once(endpoint, payload, generation, model_id):
|
||||
counts = chunk_usage(chunk)
|
||||
if counts is not None:
|
||||
generation.reported_usage = True
|
||||
|
||||
Reference in New Issue
Block a user