Reading the answer instead of asking somebody to know it

llama-server publishes the loaded model's Jinja chat template on /props, and
that template is the very thing that rejects a reasoning effort it does not
recognise -- so the accepted set is written down, authoritatively, in a place
this application can simply read. There is a button on the model's page that
does.

The parser handles both shapes a template uses: the values inline in the test
that rejects them (Bonsai), and a named list set elsewhere with nothing near
the mention spelling them out (gpt-oss). It is deliberately conservative,
because a wrong answer here silently removes a level somebody is entitled to:
only known efforts count, an unrelated list of quoted strings is ignored, and a
single match is read as a default -- `{%- set reasoning_effort = 'medium' %}`
-- rather than as a vocabulary of one.

An endpoint with no such route says so. OpenAI and vLLM do not publish a
template, and "this cannot tell us" must not be recorded as "this model accepts
nothing".

/props sits at the server root, beside the OpenAI-compatible surface rather
than inside it, so a base URL written as .../v1 needs the suffix stripped.
Getting that wrong is a silent 404 that looks like detection simply not
working, so there is a test on it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 22:23:22 +00:00
co-authored by Claude Opus 5
parent 32e2326d41
commit b1dbca7db6
7 changed files with 323 additions and 2 deletions
+61
View File
@@ -3,6 +3,7 @@
from __future__ import annotations
import logging
import re
from datetime import UTC, datetime, timedelta
from typing import Any
@@ -470,6 +471,66 @@ def resolved_effort(chat) -> str:
return value if value in EFFORTS else ""
def efforts_from_chat_template(template: str) -> list[str]:
"""Which efforts a model's Jinja chat template will actually accept.
The template is where the truth lives: the one on a Bonsai reads roughly
{%- if reasoning_effort not in ('xhigh', 'medium', 'low') %}
{{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ...
so the accepted set is written out beside the thing that rejects everything
else. `llama-server` hands the whole template over on `/props`, which makes
this readable rather than guessable.
Deliberately conservative, because a wrong answer here silently removes a
level somebody is entitled to:
- only quoted literals within a short window of a `reasoning_effort`
mention are considered, so an unrelated list elsewhere in a four-hundred
line template cannot contribute;
- the result is intersected with `EFFORTS`, so an unknown token is dropped
rather than stored;
- fewer than two survivors is treated as "the template did not say". One
match is far more likely to be a default assignment
(`{%- set reasoning_effort = 'medium' %}`) than a vocabulary.
Returns [] when nothing can be read, which every caller treats as "ask
somebody" rather than as "this model accepts nothing".
"""
if not template or "reasoning_effort" not in template:
return []
found: set[str] = set()
# Shape one: the values sit in the statement that tests them.
# {%- if reasoning_effort not in ('xhigh', 'medium', 'low') %}
for match in re.finditer(r"reasoning_effort", template):
window = template[match.start() : match.start() + 400]
# Stop at the end of the statement that mentions it, so a later,
# unrelated block cannot leak in.
window = window.split("%}")[0] if "%}" in window else window
for literal in re.findall(r"""['"]([a-z]{3,8})['"]""", window):
if literal in EFFORTS:
found.add(literal)
# Shape two: the values are a named list somewhere else, and the test says
# {%- if reasoning_effort not in valid_efforts %}
# so nothing near the mention names them. Any group of quoted literals in
# which *every* token is a known effort and there are at least two is taken
# -- that is a strong enough signal on its own, and a list of nothing but
# effort names that is not the effort vocabulary would be a strange thing
# for a chat template to contain.
for group in re.findall(r"[\[(]((?:\s*['\"][a-z]{3,8}['\"]\s*,?)+)[\])]", template):
literals = re.findall(r"""['"]([a-z]{3,8})['"]""", group)
if len(literals) >= 2 and all(value in EFFORTS for value in literals):
found.update(literals)
if len(found) < 2:
return []
return [effort for effort in EFFORTS if effort in found]
def apply_effort(
body: dict[str, Any], effort: str | None, supported: tuple[str, ...] | None = None
) -> None: