Reading the answer instead of asking somebody to know it

llama-server publishes the loaded model's Jinja chat template on /props, and
that template is the very thing that rejects a reasoning effort it does not
recognise -- so the accepted set is written down, authoritatively, in a place
this application can simply read. There is a button on the model's page that
does.

The parser handles both shapes a template uses: the values inline in the test
that rejects them (Bonsai), and a named list set elsewhere with nothing near
the mention spelling them out (gpt-oss). It is deliberately conservative,
because a wrong answer here silently removes a level somebody is entitled to:
only known efforts count, an unrelated list of quoted strings is ignored, and a
single match is read as a default -- `{%- set reasoning_effort = 'medium' %}`
-- rather than as a vocabulary of one.

An endpoint with no such route says so. OpenAI and vLLM do not publish a
template, and "this cannot tell us" must not be recorded as "this model accepts
nothing".

/props sits at the server root, beside the OpenAI-compatible surface rather
than inside it, so a base URL written as .../v1 needs the suffix stripped.
Getting that wrong is a silent 404 that looks like detection simply not
working, so there is a test on it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 22:23:22 +00:00
co-authored by Claude Opus 5
parent 32e2326d41
commit b1dbca7db6
7 changed files with 323 additions and 2 deletions
+62 -1
View File
@@ -4,6 +4,7 @@ from __future__ import annotations
import contextlib
import logging
from urllib.parse import quote
from fastapi import APIRouter, File, Form, HTTPException, Request, Response, UploadFile, status
from fastapi.responses import FileResponse, RedirectResponse
@@ -162,7 +163,13 @@ async def models_page(
@router.get("/admin/models/{model_id}/edit")
async def model_detail(
request: Request, db: Db, user: AdminUser, model_id: str, saved: str = ""
request: Request,
db: Db,
user: AdminUser,
model_id: str,
saved: str = "",
detected: str = "",
message: str = "",
):
"""Everything about one model, on its own page."""
model = _model(db, model_id)
@@ -182,6 +189,11 @@ async def model_detail(
# current answer, which is the common three until somebody says.
"efforts": chat_service.EFFORTS,
"model_efforts": chat_service.efforts_for(model),
# What `detect-efforts` found, if it has just run. Escaped by the
# template like every other value; it is prose the endpoint or this
# application wrote, not markup.
"detected": detected if detected in ("success", "warning") else "",
"detected_message": message[:400],
# Rows predating the split have no tool_* keys at all. Showing them
# unticked would be a lie: tools.enabled_tools treats absent as on
# when `tools` is on, so that an upgrade does not silently take web
@@ -342,6 +354,55 @@ async def move_model(
return RedirectResponse(back or "/admin/models", status_code=303)
@router.post("/admin/models/{model_id}/detect-efforts")
async def detect_efforts(db: Db, user: AdminUser, model_id: str) -> Response:
"""Ask the endpoint which reasoning efforts this model actually takes.
llama-server hands its loaded model's Jinja chat template over on `/props`,
and that template is the thing that rejects an effort it does not know -- so
the accepted set is written down in the one place that is authoritative,
rather than having to be guessed at or discovered by a failed reply.
Anything that is not a llama-server answers nothing here, and that is a
normal outcome: OpenAI and vLLM have no such route, and their models are
documented rather than introspectable. The result then says so instead of
claiming the model accepts nothing.
"""
from lembas.services.llm.openai_client import Endpoint, fetch_chat_template
model = _model(db, model_id)
connection = db.get(Connection, model.connection_id)
if connection is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That connection no longer exists.")
template = await fetch_chat_template(Endpoint.from_connection(connection))
found = chat_service.efforts_from_chat_template(template)
if found:
model.reasoning_efforts = found
db.commit()
message = "This model's template accepts: " + ", ".join(found) + "."
kind = "success"
elif template:
message = (
"The endpoint gave up its chat template, but nothing in it names a "
"set of reasoning efforts. Either this model does not take one, or "
"it accepts anything and never checks."
)
kind = "warning"
else:
message = (
"This endpoint does not publish its chat template, so there is "
"nothing to read. llama.cpp does; OpenAI and vLLM do not."
)
kind = "warning"
return RedirectResponse(
f"/admin/models/{model.id}/edit?detected={kind}&message={quote(message)}",
status_code=status.HTTP_303_SEE_OTHER,
)
@router.post("/admin/models/{model_id}/default")
async def set_default_model(
db: Db, user: AdminUser, model_id: str, back: str = Form("")