MCP servers, over streamable HTTP
A server is a row with a URL; its tools are discovered by a button and cached, then offered beside the built-in ones. Written by hand rather than taken from the reference SDK, because that SDK's transport does its own connecting -- and the one thing that must not be bypassed is check_url on every hop. Owning the transport is the point; the framing beside it is the small part. Sessions are per call: initialize, initialized, the call, a best-effort DELETE. Caching one wants an owner, a TTL, eviction, a lock and a shutdown hook, and the server may expire it under all of that anyway -- ToolContext is a session-free snapshot precisely so nothing in a tool holds live state. A server's names and descriptions reach the model as instructions and are bounded before they do; what it returns is escaped preformatted text, never markdown. Tools are namespaced per server, so two servers exposing "search" do not collide and neither shadows a built-in. Also: a round's calls now run together under a semaphore, results indexed so each tool turn stays paired with its call, and generation.status names what is running -- a remote tool is latency-bound, and a silent pause is what a hang looks like. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -64,6 +64,22 @@ be a different project, not a refactor.
|
||||
support rejects the whole request rather than ignoring the array
|
||||
- [x] Sources stay in the transcript; results are **not** replayed as context on
|
||||
the next turn, for the same reasons reasoning is not
|
||||
- [x] A round's calls run together, and the reply says which tool is running —
|
||||
a remote tool taking seconds with nothing streaming looks like a hang
|
||||
- [x] **Custom HTTP tools** — an administrator describes one call: a JSON Schema,
|
||||
a URL template, headers, an encrypted secret and how to read the answer.
|
||||
Arguments may fill a hole but never move the target: the scheme and host
|
||||
are literal, values are escaped for where they land, and the origin is
|
||||
pinned afterwards
|
||||
- [x] **MCP servers** over streamable HTTP — a hand-written client, so that
|
||||
`check_url` runs on every hop rather than being bypassed by somebody
|
||||
else's transport. Tools are discovered and cached by a button, namespaced
|
||||
per server, and a server's own descriptions are bounded before they reach
|
||||
a model as instructions
|
||||
- [x] Both gated like the built-ins — a model capability, a permission — and
|
||||
restrictable to groups, with guidance of their own on `/admin/prompts`
|
||||
- [x] Local MCP over stdio is deliberately absent: spawning a subprocess is the
|
||||
agentic-execution feature and wants a confirmation model first
|
||||
|
||||
### The library
|
||||
- [x] **Knowledge bases** — documents, images and saved web pages, grouped into
|
||||
@@ -155,13 +171,6 @@ be a different project, not a refactor.
|
||||
|
||||
In the order they are likely to be worth doing.
|
||||
|
||||
### Custom tools and MCP servers
|
||||
An MCP client managing configured servers, their tools surfaced alongside the
|
||||
built-in ones. The loop they plug into exists now — `services/tools.py` is a
|
||||
registry of thirteen tools and `services/generation.py` already runs bounded
|
||||
rounds — so this is a client and an admin screen rather than a change to how
|
||||
chat works.
|
||||
|
||||
### Agentic execution
|
||||
Two modes, as originally specified:
|
||||
- **local** — subprocess on the machine LLeMbas runs on
|
||||
|
||||
Reference in New Issue
Block a user