One round for a chat, as many as it takes for an agent
Two different jobs were sharing one number. A plain conversation asking a question is one round of looking things up and then an answer; the rounds after that were a small model that had decided searching was the answer searching until the context ran out, at a full request each. MAX_ROUNDS is 1 now. Several tools can still be called within that round, which is the thing worth telling the model. The trade is real and worth naming: a plain chat can no longer search and then read one of the results, because reading is a second round. That is what an agent chat is for. An agent chat is sized by Limits instead, where steps is now a runaway backstop and not a working budget. It was 40 and it was reached -- a step count low enough to be the thing that ends a reply is a count that ends it halfway. What bounds one now is the wall clock and a new completion-token ceiling, with zero meaning no ceiling, the same convention index_chars already uses. That ceiling would have been decorative. generation.completion_tokens is only populated when the endpoint sends a usage block, and llama.cpp, Ollama and friends never do; the fallback estimate is computed once, in _run's finally, long after the loop that needs it. So _written takes the larger of reported and estimated, and there is a test that runs the whole thing against a stream reporting no usage at all. A limit that works on OpenAI and silently does nothing everywhere else is the worst kind: one that looks configured. core.rounds could not stay one fragment. "You get at most N rounds" is not the same sentence with a different number in it -- a model told it has a budget rations it and stops early to report progress, which is exactly the behaviour that strands a long piece of work. So it splits: core.rounds keeps the one-round case and gates on a new round_budget variable that _agent_values blanks, and core.keep_working says the other thing to an agent chat. A queued message during a one-round reply is now never taken mid-reply -- there is no work under way to steer -- and falls through to _drain, which gives it a reply of its own. No code change went with that; it falls out of the guard, and there is a test so that "it happens to work" and "it is meant to work" stop looking the same. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -61,9 +61,10 @@ async def save_agents(
|
||||
default_timeout: int = Form(60),
|
||||
max_timeout: int = Form(600),
|
||||
max_output_bytes: int = Form(64 * 1024),
|
||||
max_steps: int = Form(40),
|
||||
max_steps: int = Form(200),
|
||||
max_wall_seconds: int = Form(900),
|
||||
max_total_output_bytes: int = Form(1024 * 1024),
|
||||
max_completion_tokens: int = Form(200_000),
|
||||
approval_timeout: int = Form(900),
|
||||
allow_default: str = Form(""),
|
||||
deny_default: str = Form(""),
|
||||
@@ -75,6 +76,8 @@ async def save_agents(
|
||||
terminal_integration: bool = Form(False),
|
||||
index_enabled: bool = Form(False),
|
||||
index_chars: int = Form(2000),
|
||||
instructions_enabled: bool = Form(False),
|
||||
instructions_chars: int = Form(4000),
|
||||
) -> Response:
|
||||
settings_store.update(
|
||||
db,
|
||||
@@ -86,9 +89,11 @@ async def save_agents(
|
||||
"default_timeout": min(max(default_timeout, 1), 3600),
|
||||
"max_timeout": min(max(max_timeout, 1), 3600),
|
||||
"max_output_bytes": min(max(max_output_bytes, 1024), 1024 * 1024),
|
||||
"max_steps": min(max(max_steps, 1), 200),
|
||||
"max_steps": min(max(max_steps, 1), 1000),
|
||||
"max_wall_seconds": min(max(max_wall_seconds, 30), 7200),
|
||||
"max_total_output_bytes": min(max(max_total_output_bytes, 4096), 8 * 1024 * 1024),
|
||||
# Floor of 0, not 1: zero is how "no ceiling" is said.
|
||||
"max_completion_tokens": min(max(max_completion_tokens, 0), 5_000_000),
|
||||
"approval_timeout": min(max(approval_timeout, 60), 3600),
|
||||
"allow_default": _lines(allow_default),
|
||||
"deny_default": _lines(deny_default),
|
||||
|
||||
Reference in New Issue
Block a user