Developer pages for 1.0.0

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HomerandClaude Opus 5.5 committed 2026-10-09 22:40:23 +00:00
1 parent a647fa620a
commit 9fd7f14e2a
4 files changed
+771

No files matched your search

+239
@@ -0,0 +1,239 @@
# Architecture
One process per front end, one event bus per session. `createApp` (in
`src/app.ts`) wires configuration, the project, the model and an `Engine`; the
TUI, `lembas run`, the ACP agent and the tests all start there and subscribe to
the same events.
## Layout of `src/`
```
cli.ts argv: the TUI, or run | models | sessions | config (check|schema|list|get|set)
| login | logout | serve --stdio | service | mcp | voice | kb | trust | update
| uninstall
app.ts wiring shared by every front end: config → project → model → Engine; spawn
(subagents), the live Settings, freeze() of memory and skills per session
headless.ts `lembas run "…"`: answer on stdout, tool lines on stderr, or --json lines
settings.ts a setting's value and source; set for session / global / project
reload.ts /reload: execv the binary on disk now, back on the same session
doctor.ts `lembas config check`: machine checks (another `lembas` first on PATH, …)
service.ts `lembas service`: the link as a systemd *user* unit, one copy at a time
uninstall.ts `lembas uninstall [--purge]`
harness.ts the harness spec embedded at build time: version, tool catalogue
version.ts replaced at compile time
bus/ the typed event bus, and the Asker interface (who answers approvals,
questions and plans)
config/ XDG paths (LEMBAS_HOME for tests), YAML loading and layering, {env:}/{file:}
substitution and key_cmd, the zod schema, the settings list, JSON Schemas
provider/ types.ts (Message, StreamEvent), common.ts (retry rule, part builder, stream
recorder), sse.ts, think.ts (inline <think> splitter), effort.ts, learned.ts,
discover.ts (context windows), http.ts (TLS, auth), tokens.ts, refs.ts (how a
model is named); one client per dialect: openai-chat.ts, responses.ts,
anthropic.ts, gemini.ts, ollama.ts
permission/ wildcard.ts, arity.ts (OpenCode), bash.ts (the small shell reader),
hardline.ts (the floor), evaluate.ts (rules, modes, paths)
tool/ tool.ts (the framework), registry.ts, names.ts (former names); read, write,
edit, multiedit, apply_patch (+ patch.ts, unidiff.ts, replace.ts), search.ts
(list, glob, grep), bash + jobs, web, question (ask_user), plan_exit
(plan_submit), todo, task, memory, skills, library, session_search,
settings, view_image, project_tools (tasks, decisions)
session/ engine.ts (the loop), store.ts (SQLite + FTS5), turns.ts (undo/redo),
title.ts, context.ts (/context), capacity.ts (the capacity rule),
commands.ts (slash commands shared by the TUI and ACP)
prompt/ assemble.ts (the system prompt), personality.ts
project/ root.ts (.agent, or .lembas), init.ts (first open: trust, git, skeleton),
safe.ts (no writes through a committed symlink), attach.ts (@ references),
files.ts (the @ index), commands.ts, agents.ts, board.ts, image.ts
git/ run.ts, repo.ts, snapshot.ts (the shadow store), branch.ts, commit.ts,
release.ts, worktree.ts
memory/ store.ts (MEMORY.md, USER.md), threats.ts (the injection scanner)
skill/ SKILL.md directories, the skills block of the prompt
library/ store.ts (notes, bases, documents, chunks; FTS5 + vectors), chunks.ts,
embed.ts, ingest.ts, cli.ts (`lembas kb`)
mcp/ index.ts (McpManager: stdio and HTTP servers), oauth.ts
search/ webui (a LLeMbas instance's), searxng, firecrawl, ddg; fetch.ts (pages)
voice/ audio.ts (recorders and players), speech.ts, text.ts, index.ts
lembas/ a LLeMbas instance: client.ts (discovery, device login), login.ts, webui.ts
(the `webui` connection), personal.ts (personalization, library, usage)
acp/ agent.ts, rpc.ts, link.ts, hub.ts, share.ts, prompt.ts, turns.ts,
pending.ts, shellrc.ts — see below
update/ index.ts, source.ts, sshsig.ts, version.ts — see below
tui/ OpenTUI + Solid — see below
```
Outside `src/`: `harness/` (the shared spec), `schema/` (JSON Schemas for the two
YAML files, from `bun run schema`), `scripts/` (build, dist, smoke, drive, ci,
licences, harness), `install.sh` and `get.sh`, `third_party/`.
## The engine loop (`src/session/engine.ts`)
1. Stream a reply from the provider; forward text, reasoning and tool-call deltas
to the bus.
2. No tool calls → done. Otherwise **decide every call first** (asking is
sequential): parse and validate the arguments, build the tool's permission
request, evaluate it, ask when needed. The same call three times running is
always asked about.
3. Run what was allowed: consecutive ordinary calls together, up to four; an
`exclusive` tool (such as `ask_user`) alone; two calls on the same file in
order. Errors become tool results, never exceptions.
4. Append the results and go round again.
Around that: **messages sent meanwhile** wait in the engine's inbox and enter
after the step's tool results as one user message marked `STEER_MARK`
(`busy_input: queue` holds them for the next prompt instead). **Stop** keeps what
streamed, so `/continue` picks up there. **Budgets** (`harness/loop.json`: wall
time, tool output bytes, written tokens) and the **step ceiling** withdraw the
tools and ask for an answer from what the model has. **Context**: before each
step, past `compaction.auto_at` of the window, old tool outputs are pruned, then
— if still over — the history is summarised and the turn's prompt put back
verbatim. **Capacity** (`session/capacity.ts`) refuses a subagent on a model its
server cannot hold beside the session's own.
## Permission evaluation (`src/permission/evaluate.ts`)
Every path is judged where it really is (`realPath`).
1. **Hardline**: the raw line, then each simple command in its plain spelling
(`plainCommands`: `VAR=` prefixes, wrappers such as `sudo -u x`, `env -i`,
`timeout 5`, `xargs` taken off; quotes and program paths removed; `sh -c "…"`
read as the line it runs); protected paths for writes. Deny.
2. **Rules**: defaults → global → project (trusted only) → session approvals.
Per pattern the last match wins, across patterns the strictest; an approval
never beats a written deny, and a deny on the plain spelling is a deny. A bash
line is split into simple commands; command substitution, `eval`, nested
shells or a redirect to a file turn an `allow` into `ask`. Files a command
names follow the `read` rules and the project boundary.
3. A path outside the project goes through `external_directory`.
4. **Mode**: `auto` allows; `plan` allows reads and writes only under the
project's `plans/`; `edit` allows file reads and writes inside the project
(not `.git/`, not the project's own config, agents, commands and skills);
`manual` is the rules as they are. A subagent can be stricter than its parent,
never looser.
"Always allow" covers the command's arity prefix (`git commit -m x` →
`git commit *`); a command that runs others is approved exactly, and a line where
one runs others offers no "always".
## Providers (`src/provider/`)
Every dialect turns the provider-neutral `Message[]` into its request and its
stream into the same `StreamEvent`s (`text`, `reasoning`, `tool_call_delta`,
`usage`, `progress`, `notice`, `finish`). What a provider needs back next turn
travels on the parts: Anthropic thinking signatures, a Responses reasoning item,
Gemini's thought signature. A part another provider cannot take is dropped on
conversion, never sent malformed. `common.retrying()` owns the retry rule: only
while nothing of the reply has been passed on, plus one attempt for a server
error. `effort.ts` and `learned.ts` remember an effort a server refused.
## The session store and turns
`session/store.ts` is `~/.local/share/lembas/sessions.db` (bun:sqlite), sessions
and messages with an FTS5 index for `session_search` and `/sessions`. Each prompt
is a **turn** (`session/turns.ts`): a snapshot before and after in a shadow git
repository (`git/snapshot.ts`, under `~/.local/share/lembas/snapshot/`, borrowing
the project's objects, skipping `.git`, `node_modules`, the project's `local/`
and files over 2 MB, off above 20 000 files). `/undo` restores exactly the paths
the turn changed and drops it from the conversation; `/redo` reverses that;
`/checkpoint` keeps a tree under a name with its objects packed so a `git gc` in
the project cannot orphan it. Titles are set from the first prompt at once and
improved by the model after the first reply.
## The system prompt (`src/prompt/assemble.ts`)
Identity → base and family overlay → environment, model roster, git state, task
board → `AGENTS.md` / `CLAUDE.md` and `instruction_files` → MCP server
instructions (scanned) → memory guidance and the memory snapshot → the skills
block → the mode → the person's custom instructions → the personality. The last
two use the same headings and order as the LLeMbas web UI. Memory and skills are
**frozen per session** (`freeze()` at start, `/new` and resume), so a save during
the session does not change the prompt under it and the provider's prompt cache
stays valid.
## The TUI (`src/tui/`)
A subscriber like any other. `state.ts` turns bus events into a list of items,
buffering streamed text and flushing about 30 times a second; the permission
asker is a store entry whose `resolve` a card calls. `index.tsx` builds the
renderer and gives the terminal back before any fatal error; `app.tsx` is `Root`
(layout, keys, dialogs); `components/` holds the prompt, messages, the
permission, question and plan cards, the spinner, the status bar, the todo list,
the first-run and setup screens; `commands.ts` the built-in slash commands;
`palette.ts` and `theme.ts` every colour; `clipboard.ts` copies by OSC 52 and the
system clipboard.
## ACP and the link (`src/acp/`)
- **`agent.ts`** — the CLI as an ACP agent: `lembas serve --stdio` for an editor,
and the far end of the LLeMbas link. A session is an ordinary session
(`createApp`): same tools, permissions, prompts and hardline as the TUI. A
client that says it is LLeMbas also gets every engine event as
`_lembas/event`. **Limits are the device's**: the global `remote:` block's
`roots`, `max_mode`, trust and approval timeout.
- **`rpc.ts`** — JSON-RPC 2.0, one message a line on stdio and one per WebSocket
frame on the link; written here so its few rules are what LLeMbas implements in
Python too.
- **`link.ts`** — the machine dials out to the instance it is logged in to; one
connection per device, re-dialled with backoff. Nothing listens.
- **`hub.ts`** — one session, seen from the terminal and the web alike. The
process holding the link (the service, or a terminal sharing with `/remote`) is
the hub, on a Unix socket in the state directory (mode 0600). A web-started
session runs in the hub; opening it in a terminal asks the hub to let go.
- **`share.ts`** — a terminal's side of the hub: its session shown in the web
UI, the web as a second keyboard; an approval asked on both sides, first answer
wins.
- **`prompt.ts`** — a web prompt made into what the TUI would have sent: `/name`
through `session/commands.ts`, `@path` through the TUI's expander, images and
files stored as attachments (10 MB a block, 25 MB a prompt).
- **`turns.ts`** — turn ids, the last 200 per session in a file shared by the
service and a terminal, so a prompt is never run twice across a handover or a
restart; `_lembas/session/status` reports running, queued and last finished.
- **`pending.ts`** — deletes the instance has not acknowledged, resent.
- **`shellrc.ts`** — shell integration for a terminal opened from the web: OSC
133 and OSC 7 marks for bash, zsh and fish.
The whole protocol, every `_lembas/*` method and capability, is
`harness/acp.md`. `lembas/client.ts` finds an instance
(`/.well-known/lembas.json`) and signs in by device code (RFC 8628);
`lembas/webui.ts` reads the account's models from `/v1/models` at every start.
## Updates and signing (`src/update/`)
`source.ts` lists releases from GitHub's or Gitea's API or a static directory
(default: the public GitHub repository). `version.ts` compares SemVer; a channel
is `stable` or `beta`. `index.ts` installs: fetch `SHA256SUMS` and
`SHA256SUMS.sig`; `sshsig.ts` checks the SSH signature itself (PROTOCOL.sshsig,
namespace `lembas-release`, a key from `RELEASE_KEYS` or `update.public_key`,
ed25519 over the hashed list), with no `ssh-keygen` needed; the list's first line
must name the release's version; the asset must match its checksum; the
unpacked binary must report that version. The running binary is hard-linked to
`.prev` and the new one renamed over it; the running process keeps its inode.
A build from source is never updated, and nothing older than what runs is
installed unless asked for by name. `RELEASE_KEYS` is a list so the key can be
rotated: a release signed by the old key ships the new one.
## The service (`src/service.ts`)
`lembas service install` writes `~/.config/systemd/user/lembas.service` and
starts it; `status`, `logs`, `uninstall`, and `run` (what the unit runs: the link
and the hub, in the foreground). It runs as the user, never root, listens on
nothing, and says — rather than does — that lingering is needed to keep it
running while nobody is logged in.
## Library, memory, skills, MCP, voice
- **Memory** is two small files (`MEMORY.md`, `USER.md`) under the config
directory, written by the `memory` tool and scanned by `memory/threats.ts`.
**Skills** are `SKILL.md` directories; only names and first lines enter the
prompt, and `skill_manage` is all-or-nothing across the files it touches.
- **The library** is one SQLite file: notes, knowledge bases, documents and
chunks (LLeMbas's chunker, shared by conformance), bm25 plus optional vectors,
fused by reciprocal rank. Logged in with `library: lembas`, the account's
library on the instance is used over `/mcp` instead.
- **MCP**: `McpManager` connects every server at once without blocking the TUI;
a server that connects mid-session is usable from the next prompt. Stdio
servers get a minimal environment; HTTP servers try streamable HTTP, then SSE;
OAuth is the SDK's flow with tokens in `mcp-auth.json` (mode 600).
- **Voice**: every recorder delivers raw 16 kHz mono s16le on stdout, so the
level meter and silence detection work the same; the recorder is its own
process group and is killed with it.
+84
@@ -0,0 +1,84 @@
# Configuration — for developers
Every key and what it does is documented for users at
<https://llembas.eu/docs/cli-configuration.html>. This page is about the
machinery: where configuration comes from, in what order, what a project may and
may not set, and what to touch when adding a key.
## The files
| File | Scope | Notes |
|---|---|---|
| `~/.config/lembas/connections.yaml` | global only | Endpoints, dialects, models, keys. Written `chmod 600`. A project can never define or override a connection. |
| `~/.config/lembas/config.yaml` | global | Everything else. |
| `<project>/.agent/config.yaml` | project | Read **only once the directory is trusted**. `.lembas/` is used instead when `.agent/` already belongs to another tool (`project/root.ts`). |
| `~/.config/lembas/{agents,commands,skills,memory}/` | global | Subagents, custom commands, skills, memory. The project's `.agent/agents`, `.agent/commands` and `.agent/skills` count only when trusted. |
| `~/.local/share/lembas/` | data | `sessions.db`, the library, snapshots, worktrees, `mcp-auth.json`. |
| `~/.local/state/lembas/` | state | `learned.json`, prompt history, the hub socket, the service's status. |
All three roots follow XDG; `LEMBAS_HOME` roots all three at once, which is what
the tests and the smoke test use (`config/paths.ts`).
## Loading (`src/config/load.ts`)
1. **Read the YAML**, then translate legacy keys and former names (a mode such as
`unrestricted`, a renamed permission key) with a note for the user.
2. **Drop a project's global-only keys before validation**, so their shape cannot
break the file: `GLOBAL_ONLY` = `hardline_extra`, `hardline_disable`, `voice`,
`update`, `settings_tool`, `embedding`, `remote`, `library`, `instructions`,
`personality_custom`. Each one dropped is a warning. A repository must not be
able to weaken the floor, redirect speech or updates, or write words into the
system prompt in the user's voice.
3. **Validate with zod, then substitute.** `{env:NAME}` and `{file:path}` are
filled after validation (`config/substitute.ts`); a missing variable or file is
an error, never an empty string. A URL field accepts a placeholder and is
checked again after substitution. `mcp`, `voice`, `search` and `update` are
filled one entry at a time, so one broken server or service turns off that
one thing, not the whole file. A connection whose substitution fails lands in
`broken` with the reason.
4. **Merge**: objects key by key, everything else replaced, project over global.
**Permission rules stack** instead (global first, then project), and a
project's rule cannot loosen a global one.
5. **`webui` connections are expanded** (`lembas/webui.ts`): one entry naming a
LLeMbas instance becomes the OpenAI-chat connection it is spoken to as, with
the account's models from `/v1/models` — refreshed at every start, kept from
the last good read when the instance cannot be reached. The account's
personalization replaces the local keys while logged in
(`lembas/personal.ts`). Models get refs (`provider/refs.ts`): an instance's are
`<provider>/<model>`.
## Settings while running (`src/config/settings.ts`, `src/settings.ts`)
One list, `SETTINGS`, is read three ways: `/settings`, `lembas config
get|set|list`, and the agent's `settings` tool. Each entry has a kind, the scopes
it may be set at (`session`, `global`, `project`), whether **the agent** may
change it (with approval; a less strict mode is asked every time), whether it
needs `/reload`, and what it is when unset. A value is resolved from the session,
then the project's file, then the global file, then the fallback. Writing uses
`yaml`'s `parseDocument … setIn`, so the user's comments and layout survive.
**Not in the list, on purpose:** permission rules, the hardline, MCP servers,
connections, the update source and key, and `settings_tool` itself. Those are
edited by hand, never by the agent.
## JSON Schemas
`schema/config.schema.json` and `schema/connections.schema.json` are generated
from the zod schemas (`config/jsonschema.ts`) by `bun run schema`, and written
beside a user's config by `lembas config schema`, so an editor with a YAML
language server completes and checks keys. Regenerate them in the same commit as
a schema change; a test compares them.
## Adding a key
1. Add it to the zod schema in `config/schema.ts` with a `.describe()` — the
description is what the JSON Schema and the editor show.
2. Decide whether a project may set it. If a cloned repository could use it to
weaken a check, redirect traffic or a key, or speak in the user's voice, add
it to `GLOBAL_ONLY`.
3. If it may be changed while running, add it to `SETTINGS` with its scopes, and
say whether the agent may change it. The default is no.
4. **Write its reader in the same change, with a test that reads the effect** —
a key the schema accepts and nothing reads is a promise the program does not
keep.
5. `bun run schema`, and document it on the website's configuration page.
+165
@@ -0,0 +1,165 @@
# Harness parity — LLeMbas CLI ⇄ LLeMbas
LLeMbas CLI (TypeScript) and [LLeMbas](https://github.com/LLeMbas/LLeMbas)
(Python) are two independent projects that share **one harness design**: the
same tools, the same permission floor, the same prompt texts, the same model
quirks and metadata, and one protocol between them. Each implements it natively;
neither ever executes the other. The specification is this repository's
`harness/` (see its `README.md`); the CLI embeds it at build time
(`src/harness.ts`), LLeMbas vendors a pinned, hash-checked copy
(`scripts/fetch_harness.py`, `scripts/harness.lock.json`, into
`src/lembas/harness_spec/`).
This page maps every shared concept to its file in both, and says where they
still differ. **Read it before touching either harness, and keep it true** — a
wrong row is how one side silently stops matching the other. Paths: CLI under
`src/` unless they start with `harness/`, `scripts/` or `tests/`; LLeMbas under
`src/lembas/` unless they start with `scripts/` or `tests/`.
## The rule
- A harness change (tools, permissions, prompts, quirks, the loop, the library,
the link) goes **into the spec first**, then into both implementations — or is
recorded as an open item against the other project.
- **A bug found in one is searched for in the other**, and the result —
including "not present" — is recorded with the fix.
- Every such fix brings a **conformance case** in `harness/conformance/`, which
both test suites run.
- An LLeMbas administrator's prompt overrides are a local setting, not a
difference.
## The spec itself
| Concept | LLeMbas CLI | LLeMbas |
|---|---|---|
| Where the spec lives | `harness/` (source of truth), embedded by `harness.ts` | `harness_spec/` (vendored), read by `services/harness_spec.py` |
| Pinning and updating | `harness/VERSION` | `scripts/fetch_harness.py --update vX.Y.Z` (a CLI tag), `scripts/harness.lock.json` |
| Spec version reported | `HARNESS_VERSION` (`harness.ts`), sent as `harness_spec` in `initialize` and the link's hello (`acp/agent.ts`, `acp/link.ts`) | `/.well-known/lembas.json` (`api/devices.py`), `/admin/updates` (`api/admin_updates.py`) |
| Conformance runner | `tests/conformance.ts` (areas → functions), `tests/conformance.test.ts`; `bun run harness record` (`scripts/harness.ts`) | `tests/test_harness_spec.py` (`AREAS`, and `PENDING` with reasons) |
| Tool definitions kept equal to the spec | `tests/harness.test.ts` | `tests/test_harness_spec.py` (`DEVICE_TOOLS`, `TOOLS_PENDING`) |
## Conformance areas
| Area | LLeMbas CLI | LLeMbas |
|---|---|---|
| `hardline` | `permission/hardline.ts` `hardlineCommand` | `services/agent/shell.py` `hardline_command` |
| `split` | `permission/bash.ts` `splitCommand` | `services/agent/shell.py` `split_command` |
| `plain` | `permission/hardline.ts` `plainCommands` | `services/agent/shell.py` `plain_commands` |
| `arity` | `permission/arity.ts` `prefix` | `services/agent/shell.py` `prefix` |
| `always` | `permission/evaluate.ts` `alwaysPatterns` | `services/agent/policy.py` `always_patterns` |
| `permission` | `permission/evaluate.ts` `evaluate` | **pending** — LLeMbas's modes are its own table (`services/agent/policy.py` `POLICY`) |
| `chunk` | `library/chunks.ts` `split` | `services/library/chunks.py` `split` |
| `think` | `provider/think.ts` `ThinkSplitter` | `services/reasoning.py` `ReasoningSplitter` |
| `effort` | `provider/effort.ts` `effortRefused`, `advertisedEfforts` | `services/generation.py` `_effort_was_refused`, `_advertised_efforts` |
| `edit` | `tool/replace.ts` | not run — the file tools are the CLI's |
| `unidiff` | `tool/unidiff.ts` | not run — the file tools are the CLI's |
| `envelope` | `tool/patch.ts` | not run — the file tools are the CLI's |
## Tools
Scopes come from each `harness/tools/<name>.json`: `shared` (both implement it),
`cli` (the CLI's alone). LLeMbas runs no agent of its own — an agent chat runs on
a linked CLI — so the execution tools are `cli`.
| Tool (spec name) | Scope | LLeMbas CLI | LLeMbas | Differs |
|---|---|---|---|---|
| `ask_user` | shared | `tool/question.ts`; card `tui/components/question.tsx` | `ask_user` in `services/tools.py`, `services/interaction.py` | LLeMbas's own schema and card; same name |
| `task` | shared | `tool/task.ts`, `app.ts` spawn, `project/agents.ts` | `subagent_run` in `services/subagent.py` | name; LLeMbas's helper always on the parent's model |
| `web_search` | shared | `tool/web.ts`, `search/` | `web_search` in `services/tools.py`, `services/search/` | LLeMbas's own description |
| `web_fetch` | shared | `tool/web.ts`, `search/fetch.ts` | `fetch` in `services/tools.py`, `services/fetch.py` | name |
| `memory` | shared | `tool/memory.ts`, `memory/store.ts` | `memory_add`, `memory_forget` in `services/tools.py`, `services/library/memories.py` | shape |
| `notes_search` · `note_view` · `note_manage` | shared | `tool/library.ts`, `library/store.ts` | `notes_search`, `notes_get`, `notes_create`, `notes_edit`, `notes_delete` in `services/tools.py`, `services/library/notes.py` | names |
| `skills_list` · `skill_view` · `skill_manage` | shared | `tool/skills.ts`, `skill/index.ts` | `skill_get`, `skill_create`, `skill_edit` in `services/tools.py`, `services/library/skills.py`; the index is in the prompt | names; no list tool |
| `knowledge_search` · `knowledge_get` | shared | `tool/library.ts`, `library/store.ts` | `services/tools.py`, `services/library/documents.py`, `retrieval.py` | LLeMbas's own descriptions; same names |
| `read` · `write` · `list` | cli | `tool/read.ts`, `tool/write.ts`, `tool/search.ts` | — | by design |
| `edit` · `multiedit` | cli | `tool/edit.ts`, `tool/multiedit.ts`, `tool/replace.ts` | — | by design |
| `apply_patch` | cli | `tool/apply_patch.ts`, `tool/patch.ts`, `tool/unidiff.ts` | — | by design |
| `glob` · `grep` | cli | `tool/search.ts` | — | by design |
| `bash` · `bash_output` · `bash_list` · `bash_kill` | cli | `tool/bash.ts`, `tool/jobs.ts` | — | by design |
| `todo` | cli | `tool/todo.ts`; `tui/components/todos.tsx` | drawn from the device's events above the composer | — |
| `plan_submit` | cli | `tool/plan_exit.ts`; `tui/components/plancard.tsx` | the plan card for a device chat (capability `plan`) | — |
| `tasks` · `decisions` | cli | `tool/project_tools.ts`, `project/board.ts` | — | by design |
| `session_search` · `settings` · `view_image` | cli | `tool/session_search.ts`, `tool/settings.ts`, `tool/view_image.ts` | — | by design |
| Former names | | `tool/names.ts` reads `aliases.cli` and `argument_aliases` | its tools still carry the names in `aliases.llembas` | see Open items |
| How a call is drawn (`block`) | | `tui/state.ts`, `tui/components/messages.tsx` | `services/tool_blocks.py` | — |
| LLeMbas only | | | `report_*`, `schedule_*`, `image_generate`, `scratch_write`, custom HTTP tools, remote MCP | web features, by design |
## Permissions
| Concept | LLeMbas CLI | LLeMbas |
|---|---|---|
| Modes `manual` `edit` `auto` `plan` | `harness/permission/modes.json`; `permission/evaluate.ts`; `config/schema.ts` (`unrestricted` read as `auto`) | `services/agent/policy.py` (the device's mode is clamped by the device) |
| Hardline floor, protected paths | `harness/permission/hardline.json`; `permission/hardline.ts` | `services/agent/shell.py`, `services/agent/policy.py` `refusal` |
| Judging a command line | `permission/bash.ts`, `permission/evaluate.ts` (each command and its plain spelling) | `services/agent/shell.py` |
| "Always allow" | `harness/permission/arity.json`; `permission/arity.ts`, `evaluate.ts` | `services/agent/policy.py` `always_patterns` |
| Default rules | `harness/permission/defaults.json`; `DEFAULT_RULES` in `evaluate.ts` | **pending** (see the `permission` area) |
| Approving | once, edit then once, session, project, deny, deny with a reason — `tui/components/permission.tsx`; over ACP `session/request_permission` | the approval card on a device chat; first answer on either side wins |
| Device limits | the global `remote:` block → `Limits` in `acp/agent.ts` | shown, never set |
## Loop, providers, prompts
| Concept | LLeMbas CLI | LLeMbas | Differs |
|---|---|---|---|
| The loop: step backstop, budgets, wrap-up, parallel calls | `harness/loop.json`; `session/engine.ts` | `services/generation.py` (`_run`, `_wrap_up`), `services/settings_store.py` | budgets unset by default in the CLI |
| Capacity rule | `session/capacity.ts` | `services/subagent.py`, the connection's flags (`db/models/connection.py`) | — |
| Dialects | `provider/openai-chat.ts`, `responses.ts`, `anthropic.ts`, `gemini.ts`, `ollama.ts` | `services/llm/openai_client.py`, `anthropic_client.py` | LLeMbas: OpenAI chat and Anthropic, by design |
| Quirks of real servers | `harness/quirks.json`; `provider/common.ts`, `sse.ts`, `openai-chat.ts` | `services/llm/openai_client.py`, `services/generation.py` | — |
| Reasoning effort | `provider/effort.ts`, `provider/learned.ts` | `services/chat.py`, `services/generation.py` | — |
| Model metadata | `harness/models/schema.json`; `config/schema.ts`, `lembas/webui.ts` reads it from `/v1/models` | served at `/v1/models` (`api/v1/openai.py`) from the model rows | — |
| Prompt texts | `harness/prompts/` (`index.json`: which are shared), `prompt/assemble.ts` | `services/prompts.py` fragments, `services/harness.py` | see below |
| `tasks/continue.md` | `session/commands.ts`, `session/engine.ts` | `harness_spec.prompt("tasks/continue.md")` in `api/chats.py`, `services/generation.py`, `web/templating.py` | identical |
| Personality presets | `harness/prompts/personality/`; `prompt/personality.ts` | `services/personalization.py` (reads the vendored presets) | identical |
| Compaction | prune, then summarise, prompt put back (`session/engine.ts`) | summary as a user + assistant turn, turns kept collapsed (`services/compaction.py`) | shape |
| Library: chunking, search, fusion | `library/chunks.ts`, `library/store.ts` | `services/library/chunks.py`, `retrieval.py` | — |
**Prompt texts.** `harness/prompts/index.json` records, for each shared text,
which LLeMbas fragment it is and whether LLeMbas's default is word for word the
same. Identical today: `tasks/continue.md` and the six `personality/*.md`.
Shared but still differing in LLeMbas: `system/default.md` (LLeMbas's `core.*`
fragments are separate, editable pieces), `family/*.md`, `modes/*.md`,
`blocks/memory.md`, `blocks/skills.md`, `tasks/compact.md`. The CLI's alone:
`system/identity.md`, `blocks/skills-empty.md`, `blocks/agents-*.md`,
`tasks/review.md`, `tasks/changelog.md`, `tasks/init.md`.
## Connecting
| Concept | LLeMbas CLI | LLeMbas |
|---|---|---|
| The protocol, written down | `harness/acp.md`, checked against the code by `tests/harness.test.ts` | vendored with the spec |
| ACP role | agent — `acp/agent.ts` | client — `services/device_link.py` |
| JSON-RPC framing | `acp/rpc.ts` | `device_link.Link` |
| The link | `acp/link.ts` (dials out), `service.ts` | `api/devices.py` `device_link_socket` (`/api/devices/link`) |
| `_lembas` protocol versions | `acp/agent.ts` `LEMBAS_PROTOCOL` = 2; says 2 only to an instance whose discovery lists it | `device_link.PROTOCOL` = 2, `PROTOCOLS` = (1, 2), `ACP_VERSION` = 1 |
| Discovery | `lembas/client.ts` reads `protocols`, at login and every dial | `/.well-known/lembas.json`: `protocol: 1`, `protocols: [1, 2]` |
| Capabilities (`files`, `commands`, `attachments`, `ask`, `plan`, `model`, `turns`, `shell`) | offered in `initialize`'s `_meta.lembas.capabilities` | used only when listed |
| One session, both sides | `acp/hub.ts`, `acp/share.ts` | announced sessions become chats; `Chat.device_session_id` |
| Turn ids | `acp/turns.ts` | messages keep the turn id; reattach asks `_lembas/session/status` |
| A web prompt made into a typed one | `acp/prompt.ts` (`/name` via `session/commands.ts`, `@path` via `project/attach.ts`) | the composer's `@` and `/` lists, uploads as `image` / `resource` blocks |
| Deletes not yet acknowledged | `acp/pending.ts` | kept and resent until answered |
| Terminal shell marks | `acp/shellrc.ts` (bash, zsh, fish) | `services/agent/terminal.py`, `services/agent/shell_marks.py` (OSC 133 / OSC 7) |
| Device login (RFC 8628) | `lembas/client.ts`, `lembas/login.ts` | `services/devices.py`, `api/devices.py` (`/device`) |
| The account's models | `lembas/webui.ts`, `provider/refs.ts` (`<provider>/<model>`) | `api/v1/openai.py` |
| Personalization and usage | `lembas/personal.ts` | `api/v1/me.py` (`/v1/me/personalization`, `/v1/usage`) |
| Library over MCP | `library: lembas` → `mcp/index.ts` `LIBRARY_SERVER` | `api/mcp_server.py` (`/mcp`) |
| Search, fetch and speech through the instance | `search/webui.ts`, `voice/speech.ts` | `api/v1/services.py` |
## Open items
1. **Shared tools still in LLeMbas's own shape** — `ask_user` (own schema),
`task` (`subagent_run`), `web_fetch` (`fetch`), and the library tools
(`memory_*`, `notes_*`, `skill_*`, no `skills_list`). `tests/test_harness_spec.py`
names each in `TOOLS_PENDING`.
2. **The `permission` conformance area** — LLeMbas's mode table is its own; the
spec's `modes.json` and default rules are to be adopted. Commands themselves
only ever run on a linked CLI, under the CLI's evaluator and the device's
limits, so this affects what LLeMbas shows, not what runs.
3. **Prompt texts** — beyond the seven identical ones, the shared texts differ
from LLeMbas's fragments by what only a coding agent needs; a shared base with
variables, as the tool descriptions have, is the next step.
## Known gaps, kept deliberately
- **Dialects in LLeMbas.** OpenAI chat and Anthropic cover its endpoints; the
other three are not required for parity.
- **Background jobs, snapshots and undo** are the CLI's: they happen on the
machine where the work is.
+283
@@ -0,0 +1,283 @@
# Working notes
The project's rules, how it is built, tested and released, and the lessons that
became rules. Each rule carries its reason, because a rule without one is the
first thing somebody "simplifies".
## Rules — breaking one is a redesign, not a tweak
1. **Nothing is predefined.** No connection, model, endpoint or key ships in the
binary. Every connection is one the user wrote into
`~/.config/lembas/connections.yaml`, or one a `/login` to a LLeMbas instance
wrote there. A model a server does not describe gets its context from the
config or from discovery, never from a table in the code.
2. **Connections are global only.** A project's `.agent/config.yaml` can never
define or override one — otherwise a cloned repository could point
`base_url` somewhere else and collect keys and prompts. Keys are
`{env:NAME}`, `{file:path}` or `key_cmd`, never pasted. The same holds for
every key in `config/load.ts:GLOBAL_ONLY` (`hardline_extra`,
`hardline_disable`, `voice`, `update`, `settings_tool`, `embedding`,
`remote`, `library`, `instructions`, `personality_custom`): a project that
sets one is warned and ignored.
3. **The hardline is a floor.** `src/permission/hardline.ts` (data:
`harness/permission/hardline.json`) is refused in every mode, `auto`
included. Only the global config may add (`hardline_extra`) or disable
(`hardline_disable`, by id) a rule.
4. **Model output is untrusted.** Tool results are data, not instructions. A
project's own config, commands, agents, skills and MCP servers count only once
the directory is trusted. Text from a third party — an MCP prompt, a skill, a
server's instructions — goes through none of the user's conveniences: `@path`
is expanded only in what the user typed.
5. **One thin native client per dialect, no AI SDK.** `openai-chat`,
`responses`, `anthropic`, `gemini`, `ollama` — one file each in
`src/provider/`, all producing the same `StreamEvent`s. The quirks of
llama.cpp, vLLM and friends are the point, and a generic SDK hides exactly
them.
6. **One process, one bus.** The engine emits events and knows nothing about who
listens. The TUI, `lembas run`, the ACP agent and the tests are subscribers.
7. **No colour outside the theme.** Every colour comes from `src/tui/palette.ts`
(the palettes) and `theme.ts`; no hex literal anywhere else under `src/tui/`.
8. **Lifted code keeps its header**, naming the upstream file and licence.
`src/tool/replace.ts` is kept **byte-for-byte** with OpenCode (hence its
`@ts-nocheck`) so it can be re-synced; a fix it needs lives beside it
(`replace-text.ts`).
9. **The binary is the product.** `bun test` cannot see what
`bun build --compile` breaks, so every tag is preceded by `bun run smoke`,
which drives the compiled binary — including the TUI in a real PTY at 120×40
and 80×24.
10. **The harness design is shared with LLeMbas.** Tools, permissions, prompt
texts, quirks, model metadata and the device link are one design, written as
data in `harness/` and implemented natively by both projects; LLeMbas vendors
a pinned copy. A change goes into the spec first, then into both. A bug found
here is searched for in LLeMbas, and the other way round. The two connect
(login, `/v1`, `/mcp`, the outbound link) but never depend on each other.
See [Harness parity](Harness-parity).
## The shared harness, in practice
- `harness/README.md` is the spec's own description. Bump `harness/VERSION` with
the change: patch for a text or a new conformance case, minor for a new tool,
rule or field, major for a rename or a removed tool.
- **A fix comes with a conformance case that fails without it.**
`bun run harness record` fills in the `expect` of a new case from this
implementation; `bun run harness record --all` shows every case whose answer
would change. `tests/conformance.ts` maps each area to the function that
answers it; LLeMbas runs the same files through its Python.
- **A test keeps the tool definitions honest**: the spec is the source of every
tool's name, description and parameters as the model sees them, and
`tests/harness.test.ts` keeps the zod schemas equal to it and `harness/acp.md`
naming every `_lembas` method and capability the code handles or sends.
- **A version field a released client compares exactly is frozen.** Discovery
says `protocol: 1` beside `protocols: [1, 2]` because older clients compared
`protocol` for equality. Add a list beside it; never raise it.
- Nothing in `harness/` names a host, a path on a server or anybody's setup —
both projects are public.
## Build, test, release
```sh
bun install # bunfig.toml's minimumReleaseAge refuses packages under 3 days old
bun test
bunx tsc --noEmit
bun run smoke # build dist/lembas, drive it
bun run licenses --check # the licences of everything the binary contains
bun run schema # regenerate schema/*.json from the zod schemas
scripts/ci.sh # all of it, plus shellcheck and gitleaks
```
`scripts/ci.sh` is the gate, here and on CI (`.github/workflows/ci.yml` runs it
with `CI=1`, so a missing shellcheck or gitleaks fails instead of being
skipped). `.gitleaksignore` holds the one false positive: the made-up
credential the memory threat scanner's test must catch.
**Versioning.** `package.json` is the single source; `src/version.ts` is
replaced at compile time. `main` is the stable channel, tagged
`vX.Y.Z`, and pre-releases are `vX.Y.Z-beta.N`. The updater compares versions as
SemVer (`src/update/version.ts`): git's own version sort would put
`v1.1.0-beta.1` above `v1.1.0` and hand a stable install a pre-release.
**A release, in order:**
1. `CHANGELOG.md`'s `[Unreleased]` becomes `[X.Y.Z] — date`; `package.json` is
bumped; one commit, files added by name.
2. `scripts/ci.sh` passes.
3. A **signed annotated tag** whose message is that changelog section:
`git tag -s vX.Y.Z --cleanup=verbatim -F notes.md`. Without
`--cleanup=verbatim` git deletes every line starting with `#` — every
`### Fixed` heading.
4. Push, then **from the tag**: `bun install --frozen-lockfile --os=linux
--cpu=arm64` (OpenTUI's arm64 library), then `bun run dist -- --sign <the
release key>`. `dist` builds the three binaries, writes `SHA256SUMS` — its
first line `# lembas X.Y.Z`, signed with the rest so an old release's files
cannot be served under a newer tag — signs it with `ssh-keygen -Y sign` in the
`lembas-release` namespace, verifies its own signature, and says whether the
key is the one built into the binary.
5. The Release is made from the tag **with the notes passed explicitly** and
every file `SHA256SUMS` lists — the three binaries, `install.sh`, `get.sh`,
`LICENSES.txt` — plus `SHA256SUMS.sig`. `get.sh` and the updater refuse a
release without a valid signature, so a Release missing its assets or its
signature breaks every install and every update from then on.
6. The public GitHub repository's Release is made by
`.github/workflows/release.yml`: it waits for the canonical release's files,
checks them exactly as `get.sh` would (signature, version line, every
checksum), and publishes the same files and notes; a `-beta.N` tag becomes a
prerelease. Nothing is built there.
**Never gate a commit on a pipeline.** `bun test | tail -3 && git commit` commits
when `tail` succeeds. Take the run's status on its own:
`bun test > log; rc=$?`.
## Defaults, and why
- **MCP tools ask**, whatever the server says about them, unless a `permission:`
rule allows them. A read-only hint only changes the class (plan mode then asks
instead of refusing).
- **A local MCP server gets a minimal environment**, not the shell's. A server
that needs a token names it under `env:`.
- **A server's instructions enter the system prompt only if the threat scan
finds nothing.** MCP prompts are commands the user runs, not tools.
- **`memory` is allowed without asking; `skill_manage` asks.** Memory is capped,
scanned and small; a skill is instructions later sessions will follow, so it is
approved like an edit.
- **A long skill description is warned about, not refused** — small models loop
on a refusal.
- **Subagents get the skills list, not memory**, custom instructions or the
personality: those are the main agent's business.
- **What you say is a draft by default** (`voice.submit: draft`), and `voice` is
global only, so a cloned repository cannot point speech and its key elsewhere.
- **The step ceiling is a runaway backstop, not a budget.** Budgets
(`harness/loop.json`) are unset by default.
- **Gemini and Ollama are untested against real servers.** Both were written from
the API references and checked against how Hermes Agent and OpenCode handle the
same APIs; `tests/gemini-ollama.test.ts` proves the shapes, not acceptance.
## Lessons that are rules
### Permissions: judge the thing that happens
- **A permission is about what the call does; when the program cannot see that,
it asks.** Every check that failed in review looked at a *description* of the
action: the command's arguments but not its input redirection (`cat < .env`); a
link's name but not its target; a URL's host but not its redirects or what the
name resolved to; the mode as typed, not as applied; a release's tag, not the
version inside its signed list.
- **Paths are judged as real paths.** `resolve` does not follow symlinks; a
repository shipping `docs -> ~/.ssh` would get reads there as in-project.
`realPath` (the longest existing prefix, links followed — dangling ones by hand)
is used for every path; protected paths are checked on both spellings.
- **A written deny is matched against the plain spelling too.** `sudo -u x git
push`, `env A=1 git push`, `timeout 5 …`, `sh -c "…"`, `/usr/bin/git push` all
walked past `git push *: deny` until each simple command was also looked up
with wrappers, quotes and program paths taken off. Only the deny is taken from
the plain spelling — an allow there would loosen (`sudo git status` is not
`git status`).
- **A rule right per command can be wrong per line.** "Always" on
`curl x | sh` stored `curl *` and `sh`, so the next `curl anything | sh` ran
unasked. No "always" is offered for a line where one command runs others
(`exact_only` in `harness/permission/arity.json`; conformance `always.json`).
- **Read an allowed command's options, not its purpose.** `rg --pre` runs a
program, `git diff --output` writes anywhere, `git branch -v -D` deletes,
`tree -o` writes. Ask rules follow the allows for those.
- **Edit mode's waiver needs `paths`.** A `write`-class call that names no file
(`skill_manage`, a read-only MCP tool) is not a file edit.
- **The project's rules come after the user's, and a project cannot loosen a
global rule** (last match would otherwise let a cloned repository do it).
- **Fail closed** where a check cannot be done — including `get.sh`, which never
falls back to an unsigned build.
### Providers and streams
- **An SSE stream's last line may have no newline**, and an error may be a bare
`{"error":…}` line inside a 200. `sse.ts` processes an unterminated last line
like any other and surfaces every error shape seen in the wild.
- **Never re-serialise a recorded stream.** Fixtures are sanitised by in-place
substitution only, byte-exact otherwise — a re-serialiser once added the very
newline whose absence was the bug. Streamed strings are reassembled across
frames before paths are replaced, then spread back over the same frames.
- **Only non-whitespace counts as output** for the retry-before-output rule.
- **A frame with `type: "server_error"` or an `internal…` code is a 500** and
gets the usual retry (a proxy unloading the model mid-reply says so this way).
- **A broken tool call must not end the session.** Arguments that are not a JSON
object go back into the history as `{}`; a call cut off at the output limit
(`finish: length`, kept by every dialect) is not run, and the model is told to
split it. Otherwise the server re-parses the broken arguments on every request.
- **The context meter leaves out thinking the dialect does not send back.**
- **Reasoning arrives as `reasoning_content`, `reasoning` or inline `<think>`
tags**; all are read. A thinking model that ends with everything in the
reasoning channel and no content is asked once for the reply.
- **A model told "carry on" retries a refused edit indefinitely.** Tell it up
front that nobody can approve, and withdraw the tool on a final refusal.
- **Small models drop the leading `/` of absolute paths** (recovered only inside
the project) and **search code text as a regex** (`grep` offers
`literal: true`).
- **Compaction puts the prompt that started the turn back verbatim** after an
automatic summary, at most twice a turn; a summary alone loses the task.
- **Bun's fetch does not use the OS certificate store** unless run with
`--use-system-ca`; the build passes it, and `tls.ca` covers the rest.
- **`Bun.spawn` without `env` passes the environment as it was at startup.** Pass
the live environment explicitly.
- **Bun decodes text dropping a UTF-8 BOM.** Files are read with
`ignoreBOM: true` so an edit does not strip one. Lifted code carries its
runtime's assumptions: Bun is not Node.
### The TUI
- **Solid props are live getters.** A dialog that cleared its own spec and then
called `spec.onSelect` threw inside OpenTUI's key handler, which logs and
carries on — a picker that closed without choosing. Take the callback before
closing; test dialogs inside the real `Root`.
- **One terminal-size subscription**, in `Root`, passed through context — one per
component leaked listeners until Node printed a warning across the screen.
- **Test the state a program leaves behind, not only what it shows.** A throw
after OpenTUI turned on mouse reporting left the shell receiving every mouse
move as input. `runTui` destroys the renderer before printing any fatal error,
and the smoke test reads the terminal's modes after it quits.
- **No model, or an unreadable config, is a screen with a retry, never an
exception.**
- **Markdown inside a `<scrollbox>` needs `scrollX={false}`**, and is parsed off
the main thread — wait for its text, not its box.
- **OpenTUI's `truncate` elides the middle** — right for a path, wrong for a
sentence.
- **Emoji with a variation selector are one or two columns** depending on the
terminal; plain emoji are two everywhere. `icons: plain` exists for fonts
without emoji.
- **Streamed text is flushed at about 30 fps**; re-parsing markdown per token
stutters on a small machine.
### Code and data
- **A field nothing reads is a promise.** A store method with no caller, a config
key accepted by the schema and read by nothing, a status-bar slot never filled
— grep for the reader before shipping it, and test the stored row, not a screen
with a fallback.
- **Validate after substitution.** Files are validated before `{env:}`/`{file:}`
are filled (so one connection's missing variable cannot stop the others); a URL
field therefore accepts a placeholder, and what it becomes is checked after,
turning off only the entry that is still wrong.
- **`String.replaceAll` reads `$&`, `$$`, `` $` `` and `$'` in the replacement.**
Escape `$`, or pass a function.
- **`\w` is ASCII in JavaScript.** Patterns ported from Python use `\p{L}` with
the `u` flag, or one accented word hides an injection.
- **zod 4's `z.record(z.enum(…))` is exhaustive**; a partial map is
`z.partialRecord`.
- **Messages stored in one millisecond tie on time**; order by id as well.
- **An rc wrapper replaces the shell's own startup order.** The shell
integration (`src/acp/shellrc.ts`) is tested by running real bash, zsh and
fish against the startup files people really have. In bash, `PS0` can set a
variable in the shell itself only through an arithmetic array subscript.
## Testing conventions
- **`bun test`** runs unit tests and end-to-end tests against
`tests/fake-provider.ts`, a scripted endpoint. Recorded real streams replay
through the same path. `tests/tui/` drives the real `Root`.
- **Bun's `toMatchObject` with asymmetric matchers writes the matchers into the
object** — use plain `toMatch` / `toContain` on an object that is looked at
again.
- **The smoke test** (`scripts/smoke.ts`, `scripts/drive.ts`) starts the
compiled binary on a fresh home: first run, prompt, approval, answer and quit
in a PTY at both sizes, nothing wider than the terminal, the status bar on the
last row, and the terminal's modes restored.
- **MCP and OAuth are tested against the SDK's own servers**, including its
example OAuth server, so the whole sign-in runs without a public service.
- **Several suites may share a small machine.** The per-test timeout is 20 s.