Files
LLeMbas/tests/test_effort.py
T
Jaroslav Beneš 0452e742e8 A menu for what a chat may use, and three keys
Six smaller things, all of them about the interface not saying what is true.

The @ button only ever inserted the character, which the @ key already does
without a button. It becomes the scope menu: what this chat may use, switched
off per chat. Chat.scope_json is filtered inside resolve_tools AFTER the
capability, permission and instance gates -- exactly as chat.knowledge_bases
narrows knowledge_search -- so a crafted POST turning something on reaches a
tool the gates already removed, and there is a test that writes the column
directly to prove it. Absent means on, for every key, so "why is this off?" has
one answer. It is keyed on the gate rather than the tool name, so notes is one
switch rather than five. The switches carry no role="menuitem", deliberately:
ui.js closes a picker when a menuitem is clicked, which is right for an action
menu and wrong for a list you want to set several of -- which is why the menu
needs no JavaScript at all. Typing @ is untouched.

With no skills, nothing should mention them. tool.skills was gated on the family
alone, so somebody with an empty library was told "the list below gives each
one's name" above no list, handed skill_get, and watched the model spend a round
finding out. It requires skills now; the writing half moved to
tool.skills_write, which is deliberately not gated, because saving the first one
is what somebody with none most needs. And core.tool_list finally reads
tool_names, which had been resolved and documented with no fragment using it.

The composer's toolbar is one row again. .composer__actions is last in the DOM
with margin-left:auto, so the moment an agent chat added a connection, a
directory and a mode, Send and the microphone dropped to a second line.
chat.css has no media queries by design and the fix is not to add one:
.composer__context is the single child allowed to shrink and scroll sideways.
There is a test asserting the file still contains no @media.

The effort picker shows the level in force. "Effort: default" named no level and
was true of nothing in particular; chat.resolved_effort is the chat's own value
and build_request reads the same field, so what is shown is what is sent. The
model's default is a seed, copied onto the row at creation and on a model
change, and never consulted at request time -- a fallback would resurrect it
underneath a cleared effort and make "off" silently do nothing. "off" is a
sentinel and not an empty value, because start_chat declares Form("") and cannot
tell absent from empty: with value="" the reader picks off and gets high.

Alt+M dictates, Alt+R reads the last reply aloud, Ctrl+Enter sends from
anywhere. All three click the button that already does the job, so audio.js
keeps its one delegated listener. Alt+M and not Alt+D, which is the address bar
in Chrome and Firefox. Ctrl+Enter never means Stop -- Send and Stop are the same
element, and Esc already stops. Driven under a DOM stub before committing, per
the rule in CLAUDE.md, and tests/test_commands_js.py pins that every key has a
row in SHORTCUTS, since /help reads that list.

And the memory tooling, which had seven defects. The worst: memory_forget was a
case-insensitive substring first-match delete with nothing warning about it, so
forgetting "coffee" against "Drinks coffee black" and "Allergic to coffee"
silently removed whichever was older -- a wrong deletion nobody would ever find
out about, from a tool whose description invited exactly the short fragment that
misfires. It matches exactly first, then by substring, and refuses an ambiguous
one while naming what it matched. add() refuses an exact duplicate. The
at-the-limit refusal no longer tells the model to delete one to make room: past
the block's budget it is not shown all of them and would be guessing, which
feeds straight back into the first defect. And context.memories no longer claims
the memories "still apply", which nothing checks and which taught a model to
trust a stale one over what the person had just said.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 11:22:03 +02:00

369 lines
13 KiB
Python

"""Reasoning effort: what goes out, and what does not.
The second half matters as much as the first. There is no field that works
everywhere -- OpenAI and vLLM read `reasoning_effort`, llama.cpp drops it
silently and reads only `chat_template_kwargs` -- so both are sent. That is only
safe because neither is sent at all until somebody chooses an effort, which is
what keeps an endpoint strict about unknown parameters working exactly as it
did.
"""
from __future__ import annotations
import pytest
from fastapi.testclient import TestClient
from sqlalchemy import select
from lembas.db.models import Chat, Connection, Model, User
from lembas.services import chat as chat_service
from .conftest import control_named
def _model(db, **capabilities) -> Model:
connection = Connection(name="c", base_url="http://127.0.0.1:1", api_key_encrypted="")
db.add(connection)
db.commit()
model = Model(
connection_id=connection.id,
model_id="m",
capabilities_json={"reasoning": True, **capabilities},
)
db.add(model)
db.commit()
return model
def _chat(db, effort: str | None = None) -> Chat:
model = _model(db)
user = db.scalars(select(User)).first()
chat = Chat(
user_id=user.id,
model_id=model.model_id,
connection_id=model.connection_id,
params_json={"reasoning_effort": effort} if effort is not None else {},
)
db.add(chat)
db.commit()
return chat
# --- What reaches the endpoint -----------------------------------------------
def test_an_effort_goes_out_in_both_forms(client: TestClient, db, registered):
"""One value, two fields. Neither endpoint family reads the other's."""
chat = _chat(db, "high")
body = chat_service.build_request(db, chat)
assert body["reasoning_effort"] == "high"
assert body["chat_template_kwargs"] == {"reasoning_effort": "high"}
def test_no_effort_means_neither_field(client: TestClient, db, registered):
"""The whole safety of sending both. A chat nobody has set an effort on is
byte-for-byte the request it was before this existed, so a provider that
refuses unknown parameters is untouched until somebody opts in."""
chat = _chat(db)
body = chat_service.build_request(db, chat)
assert "reasoning_effort" not in body
assert "chat_template_kwargs" not in body
def test_a_cleared_effort_means_neither_field(client: TestClient, db, registered):
"""Cleared is stored as None, like every other parameter here."""
chat = _chat(db, None)
body = chat_service.build_request(db, chat)
assert "reasoning_effort" not in body
@pytest.mark.parametrize("junk", ["sudo", "HIGH ", "maximum", "1"])
def test_a_value_that_is_not_an_effort_is_not_sent(client: TestClient, db, registered, junk):
"""Never trusted from the row: it could predate a change to the list."""
chat = _chat(db, junk)
assert "reasoning_effort" not in chat_service.build_request(db, chat)
def test_existing_chat_template_kwargs_are_kept(client: TestClient, db, registered):
"""Merged rather than replaced, so a future caller setting something else
there does not lose it."""
body: dict = {"chat_template_kwargs": {"enable_thinking": True}}
chat_service.apply_effort(body, "low")
assert body["chat_template_kwargs"] == {"enable_thinking": True, "reasoning_effort": "low"}
# --- Setting it ---------------------------------------------------------------
def test_patching_the_effort_stores_it(client: TestClient, db, registered):
chat = _chat(db)
assert client.patch(
f"/api/chats/{chat.id}", data={"reasoning_effort": "medium"}
).status_code == 204
db.refresh(chat)
assert chat.params_json["reasoning_effort"] == "medium"
def test_an_empty_effort_clears_it(client: TestClient, db, registered):
chat = _chat(db, "high")
client.patch(f"/api/chats/{chat.id}", data={"reasoning_effort": ""})
db.refresh(chat)
assert chat.params_json["reasoning_effort"] is None
def test_an_unknown_effort_leaves_the_old_one(client: TestClient, db, registered):
"""Ignored, not refused: a typo should not cost the setting you had."""
chat = _chat(db, "low")
client.patch(f"/api/chats/{chat.id}", data={"reasoning_effort": "extreme"})
db.refresh(chat)
assert chat.params_json["reasoning_effort"] == "low"
def test_the_control_only_appears_on_a_reasoning_model(client: TestClient, db, registered):
"""The flag has existed with no reader since the beginning; this is its
first job. Offering the control everywhere would offer a setting that does
nothing almost everywhere."""
chat = _chat(db)
assert "data-effort" in client.get(f"/chat/{chat.id}").text
model = db.scalars(select(Model)).one()
model.capabilities_json = {"reasoning": False}
db.commit()
assert "data-effort" not in client.get(f"/chat/{chat.id}").text
# --- The per-model default ----------------------------------------------------
def test_a_new_chat_starts_from_the_models_defaults(client: TestClient, db, registered):
"""`Model.params_json` has claimed to do this since it was added and did it
nowhere. It is empty on every existing row, so honouring it changes nothing
until an administrator sets something."""
model = _model(db)
model.params_json = {"reasoning_effort": "high"}
db.commit()
client.post("/api/chats/start", data={"content": "hello", "model_id": "m"})
chat = db.scalars(select(Chat)).one()
assert chat.params_json["reasoning_effort"] == "high"
def test_a_model_with_no_defaults_starts_a_plain_chat(client: TestClient, db, registered):
_model(db)
client.post("/api/chats/start", data={"content": "hello", "model_id": "m"})
assert db.scalars(select(Chat)).one().params_json == {}
# --- Choosable before the first prompt ---------------------------------------
def test_the_effort_select_carries_its_own_verb(client: TestClient, db, registered):
"""The same invariant the mode select needs, for the same reason.
Both were built on `form="…"` pointing at an empty sibling form holding the
`hx-patch`, and both therefore wrote nothing at all: `form=` scopes the
values a request carries, it does not route the event that starts one.
"""
chat = _chat(db)
select = control_named(client.get(f"/chat/{chat.id}").text, "reasoning_effort")
assert select["hx-patch"] == f"/api/chats/{chat.id}"
assert select["form"] == "chat-params-form"
def test_the_effort_is_offered_before_there_is_a_chat(client: TestClient, db, registered):
"""Otherwise it is a setting you can only reach once it is too late to use.
On the new-chat screen there is nothing to PATCH, so it is an ordinary field
of the composer's form and carries no verb -- `_new_chat` reads it.
"""
_model(db)
select = control_named(client.get("/chat").text, "reasoning_effort")
assert "hx-patch" not in select
assert "form" not in select
def test_starting_a_chat_with_an_effort_stores_it(client: TestClient, db, registered):
_model(db)
client.post(
"/api/chats/start",
data={"content": "hello", "model_id": "m", "reasoning_effort": "low"},
)
assert db.scalars(select(Chat)).one().params_json["reasoning_effort"] == "low"
def test_an_explicit_effort_beats_the_models_default(client: TestClient, db, registered):
"""An inherited value is a starting point, not a ceiling."""
model = _model(db)
model.params_json = {"reasoning_effort": "high"}
db.commit()
client.post(
"/api/chats/start",
data={"content": "hello", "model_id": "m", "reasoning_effort": "low"},
)
assert db.scalars(select(Chat)).one().params_json["reasoning_effort"] == "low"
def test_a_nonsense_effort_at_the_start_falls_back(client: TestClient, db, registered):
model = _model(db)
model.params_json = {"reasoning_effort": "high"}
db.commit()
client.post(
"/api/chats/start",
data={"content": "hello", "model_id": "m", "reasoning_effort": "extreme"},
)
assert db.scalars(select(Chat)).one().params_json["reasoning_effort"] == "high"
# --- What the picker says is what is sent ---------------------------------------
def test_the_resolver_is_what_the_request_carries(client: TestClient, db, registered):
"""One resolver, so the control and the request cannot disagree. That
disagreement is the whole reason the picker said "default": it named no
level, and was true of nothing in particular."""
chat = _chat(db, "high")
assert chat_service.resolved_effort(chat) == "high"
assert chat_service.build_request(db, chat)["reasoning_effort"] == "high"
def test_a_cleared_effort_is_not_resurrected_by_the_models_default(
client: TestClient, db, registered
):
"""What the no-fallback decision buys. If `build_request` fell back to the
model, `update_chat` storing None for a cleared effort would be undone
underneath it and the off option would silently do nothing."""
model = _model(db)
model.params_json = {"reasoning_effort": "high"}
user = db.scalars(select(User)).first()
chat = Chat(
user_id=user.id,
model_id=model.model_id,
connection_id=model.connection_id,
params_json={"reasoning_effort": None},
)
db.add(chat)
db.commit()
assert chat_service.resolved_effort(chat) == ""
body = chat_service.build_request(db, chat)
assert "reasoning_effort" not in body
assert "chat_template_kwargs" not in body
def test_choosing_off_before_the_chat_exists_sends_nothing(
client: TestClient, db, registered
):
"""The one that would otherwise ship broken. `start_chat` declares
`Form("")`, so an absent field and an empty one are the same thing there --
with `value=""` on the off option the reader picks off, the value falls out
of EFFORTS, and the model's default seeded onto the row stays. They get
"high"."""
model = _model(db)
model.params_json = {"reasoning_effort": "high"}
db.commit()
client.post(
"/api/chats/start",
data={"content": "hello", "model_id": "m", "reasoning_effort": "off"},
)
chat = db.scalars(select(Chat)).first()
assert not (chat.params_json or {}).get("reasoning_effort")
assert "reasoning_effort" not in chat_service.build_request(db, chat)
def test_patching_off_clears_it(client: TestClient, db, registered):
chat = _chat(db, "high")
client.patch(f"/api/chats/{chat.id}", data={"reasoning_effort": "off"})
db.expire_all()
assert chat_service.resolved_effort(db.get(Chat, chat.id)) == ""
def test_switching_model_seeds_an_effort_that_was_never_chosen(
client: TestClient, db, registered
):
"""So "what the picker shows is what is sent" stays true after a switch."""
chat = _chat(db)
second = Model(
connection_id=chat.connection_id,
model_id="m2",
capabilities_json={"reasoning": True},
params_json={"reasoning_effort": "medium"},
)
db.add(second)
db.commit()
client.patch(f"/api/chats/{chat.id}", data={"model_id": "m2"})
db.expire_all()
assert chat_service.resolved_effort(db.get(Chat, chat.id)) == "medium"
def test_switching_model_does_not_overwrite_a_chosen_effort(
client: TestClient, db, registered
):
chat = _chat(db, "low")
second = Model(
connection_id=chat.connection_id,
model_id="m2",
capabilities_json={"reasoning": True},
params_json={"reasoning_effort": "high"},
)
db.add(second)
db.commit()
client.patch(f"/api/chats/{chat.id}", data={"model_id": "m2"})
db.expire_all()
assert chat_service.resolved_effort(db.get(Chat, chat.id)) == "low"
def test_switching_model_does_not_resurrect_a_cleared_effort(
client: TestClient, db, registered
):
"""`None` means somebody cleared it deliberately. Only an ABSENT key is
seeded, or "off" would silently undo itself on the next model change."""
chat = _chat(db, "high")
client.patch(f"/api/chats/{chat.id}", data={"reasoning_effort": "off"})
second = Model(
connection_id=chat.connection_id,
model_id="m2",
capabilities_json={"reasoning": True},
params_json={"reasoning_effort": "high"},
)
db.add(second)
db.commit()
client.patch(f"/api/chats/{chat.id}", data={"model_id": "m2"})
db.expire_all()
assert chat_service.resolved_effort(db.get(Chat, chat.id)) == ""
def test_the_picker_never_says_default(client: TestClient, db, registered):
"""The one markup assertion. It named no level and was true of nothing."""
chat = _chat(db, "medium")
html = client.get(f"/chat/{chat.id}").text
assert "Effort: default" not in html
assert "Effort: off" in html
assert '<option value="medium" selected>' in html.replace("\n", "").replace(" ", "")