One round for a chat, as many as it takes for an agent
Two different jobs were sharing one number. A plain conversation asking a question is one round of looking things up and then an answer; the rounds after that were a small model that had decided searching was the answer searching until the context ran out, at a full request each. MAX_ROUNDS is 1 now. Several tools can still be called within that round, which is the thing worth telling the model. The trade is real and worth naming: a plain chat can no longer search and then read one of the results, because reading is a second round. That is what an agent chat is for. An agent chat is sized by Limits instead, where steps is now a runaway backstop and not a working budget. It was 40 and it was reached -- a step count low enough to be the thing that ends a reply is a count that ends it halfway. What bounds one now is the wall clock and a new completion-token ceiling, with zero meaning no ceiling, the same convention index_chars already uses. That ceiling would have been decorative. generation.completion_tokens is only populated when the endpoint sends a usage block, and llama.cpp, Ollama and friends never do; the fallback estimate is computed once, in _run's finally, long after the loop that needs it. So _written takes the larger of reported and estimated, and there is a test that runs the whole thing against a stream reporting no usage at all. A limit that works on OpenAI and silently does nothing everywhere else is the worst kind: one that looks configured. core.rounds could not stay one fragment. "You get at most N rounds" is not the same sentence with a different number in it -- a model told it has a budget rations it and stops early to report progress, which is exactly the behaviour that strands a long piece of work. So it splits: core.rounds keeps the one-round case and gates on a new round_budget variable that _agent_values blanks, and core.keep_working says the other thing to an agent chat. A queued message during a one-round reply is now never taken mid-reply -- there is no work under way to steer -- and falls through to _drain, which gives it a reply of its own. No code change went with that; it falls out of the guard, and there is a test so that "it happens to work" and "it is meant to work" stop looking the same. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
3345df5b38
commit
7977d4ef25
@@ -45,16 +45,31 @@ from lembas.services.search.base import SearchError
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# How many times a model may call tools before it has to answer with words.
|
||||
# Not a safety limit so much as a termination one: a small model that has
|
||||
# decided searching is the answer will otherwise search until the context runs
|
||||
# out, and each round costs a full request.
|
||||
MAX_ROUNDS = 3
|
||||
# How many times a model may call tools before it has to answer with words, in
|
||||
# an ORDINARY chat. An agent chat is sized by `agent/policy.py:Limits.steps`
|
||||
# instead, which is two orders of magnitude larger, because an agent reply is
|
||||
# meant to run until the work is done.
|
||||
#
|
||||
# One, deliberately. A plain conversation asking a question is one round of
|
||||
# looking things up and then an answer; the rounds after that were a small model
|
||||
# that had decided searching was the answer searching until the context ran out,
|
||||
# at a full request each. Several tools can still be called *within* that round,
|
||||
# which is the thing worth telling the model -- see `core.rounds`.
|
||||
#
|
||||
# The trade is real and worth naming: a chat can no longer search and then read
|
||||
# one of the results, because reading is a second round. That is what an agent
|
||||
# chat is for.
|
||||
MAX_ROUNDS = 1
|
||||
|
||||
# Tool families, matching the per-model capability flags and the permission
|
||||
# keys. The three names differ by prefix only, which is deliberate: adding a
|
||||
# family means adding one entry here and one permission.
|
||||
FAMILY_SEARCH = "web_search"
|
||||
# Reading one page, given its address. Its own family rather than part of
|
||||
# `web_search`: an administrator may reasonably want a model that can look
|
||||
# things up but not follow an arbitrary URL it read somewhere, and the SSRF
|
||||
# surface is entirely on this side.
|
||||
FAMILY_FETCH = "fetch"
|
||||
FAMILY_KNOWLEDGE = "knowledge"
|
||||
FAMILY_NOTES = "notes"
|
||||
FAMILY_MEMORY = "memory"
|
||||
|
||||
Reference in New Issue
Block a user