Clone
1
Working notes
Jaroslav Beneš edited this page 2026-10-09 22:40:23 +00:00
This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

Working notes

The project's rules, how it is built, tested and released, and the lessons that became rules. Each rule carries its reason, because a rule without one is the first thing somebody "simplifies".

Rules — breaking one is a redesign, not a tweak

  1. Nothing is predefined. No connection, model, endpoint or key ships in the binary. Every connection is one the user wrote into ~/.config/lembas/connections.yaml, or one a /login to a LLeMbas instance wrote there. A model a server does not describe gets its context from the config or from discovery, never from a table in the code.
  2. Connections are global only. A project's .agent/config.yaml can never define or override one — otherwise a cloned repository could point base_url somewhere else and collect keys and prompts. Keys are {env:NAME}, {file:path} or key_cmd, never pasted. The same holds for every key in config/load.ts:GLOBAL_ONLY (hardline_extra, hardline_disable, voice, update, settings_tool, embedding, remote, library, instructions, personality_custom): a project that sets one is warned and ignored.
  3. The hardline is a floor. src/permission/hardline.ts (data: harness/permission/hardline.json) is refused in every mode, auto included. Only the global config may add (hardline_extra) or disable (hardline_disable, by id) a rule.
  4. Model output is untrusted. Tool results are data, not instructions. A project's own config, commands, agents, skills and MCP servers count only once the directory is trusted. Text from a third party — an MCP prompt, a skill, a server's instructions — goes through none of the user's conveniences: @path is expanded only in what the user typed.
  5. One thin native client per dialect, no AI SDK. openai-chat, responses, anthropic, gemini, ollama — one file each in src/provider/, all producing the same StreamEvents. The quirks of llama.cpp, vLLM and friends are the point, and a generic SDK hides exactly them.
  6. One process, one bus. The engine emits events and knows nothing about who listens. The TUI, lembas run, the ACP agent and the tests are subscribers.
  7. No colour outside the theme. Every colour comes from src/tui/palette.ts (the palettes) and theme.ts; no hex literal anywhere else under src/tui/.
  8. Lifted code keeps its header, naming the upstream file and licence. src/tool/replace.ts is kept byte-for-byte with OpenCode (hence its @ts-nocheck) so it can be re-synced; a fix it needs lives beside it (replace-text.ts).
  9. The binary is the product. bun test cannot see what bun build --compile breaks, so every tag is preceded by bun run smoke, which drives the compiled binary — including the TUI in a real PTY at 120×40 and 80×24.
  10. The harness design is shared with LLeMbas. Tools, permissions, prompt texts, quirks, model metadata and the device link are one design, written as data in harness/ and implemented natively by both projects; LLeMbas vendors a pinned copy. A change goes into the spec first, then into both. A bug found here is searched for in LLeMbas, and the other way round. The two connect (login, /v1, /mcp, the outbound link) but never depend on each other. See Harness parity.

The shared harness, in practice

  • harness/README.md is the spec's own description. Bump harness/VERSION with the change: patch for a text or a new conformance case, minor for a new tool, rule or field, major for a rename or a removed tool.
  • A fix comes with a conformance case that fails without it. bun run harness record fills in the expect of a new case from this implementation; bun run harness record --all shows every case whose answer would change. tests/conformance.ts maps each area to the function that answers it; LLeMbas runs the same files through its Python.
  • A test keeps the tool definitions honest: the spec is the source of every tool's name, description and parameters as the model sees them, and tests/harness.test.ts keeps the zod schemas equal to it and harness/acp.md naming every _lembas method and capability the code handles or sends.
  • A version field a released client compares exactly is frozen. Discovery says protocol: 1 beside protocols: [1, 2] because older clients compared protocol for equality. Add a list beside it; never raise it.
  • Nothing in harness/ names a host, a path on a server or anybody's setup — both projects are public.

Build, test, release

bun install                 # bunfig.toml's minimumReleaseAge refuses packages under 3 days old
bun test
bunx tsc --noEmit
bun run smoke               # build dist/lembas, drive it
bun run licenses --check    # the licences of everything the binary contains
bun run schema              # regenerate schema/*.json from the zod schemas
scripts/ci.sh               # all of it, plus shellcheck and gitleaks

scripts/ci.sh is the gate, here and on CI (.github/workflows/ci.yml runs it with CI=1, so a missing shellcheck or gitleaks fails instead of being skipped). .gitleaksignore holds the one false positive: the made-up credential the memory threat scanner's test must catch.

Versioning. package.json is the single source; src/version.ts is replaced at compile time. main is the stable channel, tagged vX.Y.Z, and pre-releases are vX.Y.Z-beta.N. The updater compares versions as SemVer (src/update/version.ts): git's own version sort would put v1.1.0-beta.1 above v1.1.0 and hand a stable install a pre-release.

A release, in order:

  1. CHANGELOG.md's [Unreleased] becomes [X.Y.Z] — date; package.json is bumped; one commit, files added by name.
  2. scripts/ci.sh passes.
  3. A signed annotated tag whose message is that changelog section: git tag -s vX.Y.Z --cleanup=verbatim -F notes.md. Without --cleanup=verbatim git deletes every line starting with # — every ### Fixed heading.
  4. Push, then from the tag: bun install --frozen-lockfile --os=linux --cpu=arm64 (OpenTUI's arm64 library), then bun run dist -- --sign <the release key>. dist builds the three binaries, writes SHA256SUMS — its first line # lembas X.Y.Z, signed with the rest so an old release's files cannot be served under a newer tag — signs it with ssh-keygen -Y sign in the lembas-release namespace, verifies its own signature, and says whether the key is the one built into the binary.
  5. The Release is made from the tag with the notes passed explicitly and every file SHA256SUMS lists — the three binaries, install.sh, get.sh, LICENSES.txt — plus SHA256SUMS.sig. get.sh and the updater refuse a release without a valid signature, so a Release missing its assets or its signature breaks every install and every update from then on.
  6. The public GitHub repository's Release is made by .github/workflows/release.yml: it waits for the canonical release's files, checks them exactly as get.sh would (signature, version line, every checksum), and publishes the same files and notes; a -beta.N tag becomes a prerelease. Nothing is built there.

Never gate a commit on a pipeline. bun test | tail -3 && git commit commits when tail succeeds. Take the run's status on its own: bun test > log; rc=$?.

Defaults, and why

  • MCP tools ask, whatever the server says about them, unless a permission: rule allows them. A read-only hint only changes the class (plan mode then asks instead of refusing).
  • A local MCP server gets a minimal environment, not the shell's. A server that needs a token names it under env:.
  • A server's instructions enter the system prompt only if the threat scan finds nothing. MCP prompts are commands the user runs, not tools.
  • memory is allowed without asking; skill_manage asks. Memory is capped, scanned and small; a skill is instructions later sessions will follow, so it is approved like an edit.
  • A long skill description is warned about, not refused — small models loop on a refusal.
  • Subagents get the skills list, not memory, custom instructions or the personality: those are the main agent's business.
  • What you say is a draft by default (voice.submit: draft), and voice is global only, so a cloned repository cannot point speech and its key elsewhere.
  • The step ceiling is a runaway backstop, not a budget. Budgets (harness/loop.json) are unset by default.
  • Gemini and Ollama are untested against real servers. Both were written from the API references and checked against how Hermes Agent and OpenCode handle the same APIs; tests/gemini-ollama.test.ts proves the shapes, not acceptance.

Lessons that are rules

Permissions: judge the thing that happens

  • A permission is about what the call does; when the program cannot see that, it asks. Every check that failed in review looked at a description of the action: the command's arguments but not its input redirection (cat < .env); a link's name but not its target; a URL's host but not its redirects or what the name resolved to; the mode as typed, not as applied; a release's tag, not the version inside its signed list.
  • Paths are judged as real paths. resolve does not follow symlinks; a repository shipping docs -> ~/.ssh would get reads there as in-project. realPath (the longest existing prefix, links followed — dangling ones by hand) is used for every path; protected paths are checked on both spellings.
  • A written deny is matched against the plain spelling too. sudo -u x git push, env A=1 git push, timeout 5 …, sh -c "…", /usr/bin/git push all walked past git push *: deny until each simple command was also looked up with wrappers, quotes and program paths taken off. Only the deny is taken from the plain spelling — an allow there would loosen (sudo git status is not git status).
  • A rule right per command can be wrong per line. "Always" on curl x | sh stored curl * and sh, so the next curl anything | sh ran unasked. No "always" is offered for a line where one command runs others (exact_only in harness/permission/arity.json; conformance always.json).
  • Read an allowed command's options, not its purpose. rg --pre runs a program, git diff --output writes anywhere, git branch -v -D deletes, tree -o writes. Ask rules follow the allows for those.
  • Edit mode's waiver needs paths. A write-class call that names no file (skill_manage, a read-only MCP tool) is not a file edit.
  • The project's rules come after the user's, and a project cannot loosen a global rule (last match would otherwise let a cloned repository do it).
  • Fail closed where a check cannot be done — including get.sh, which never falls back to an unsigned build.

Providers and streams

  • An SSE stream's last line may have no newline, and an error may be a bare {"error":…} line inside a 200. sse.ts processes an unterminated last line like any other and surfaces every error shape seen in the wild.
  • Never re-serialise a recorded stream. Fixtures are sanitised by in-place substitution only, byte-exact otherwise — a re-serialiser once added the very newline whose absence was the bug. Streamed strings are reassembled across frames before paths are replaced, then spread back over the same frames.
  • Only non-whitespace counts as output for the retry-before-output rule.
  • A frame with type: "server_error" or an internal… code is a 500 and gets the usual retry (a proxy unloading the model mid-reply says so this way).
  • A broken tool call must not end the session. Arguments that are not a JSON object go back into the history as {}; a call cut off at the output limit (finish: length, kept by every dialect) is not run, and the model is told to split it. Otherwise the server re-parses the broken arguments on every request.
  • The context meter leaves out thinking the dialect does not send back.
  • Reasoning arrives as reasoning_content, reasoning or inline <think> tags; all are read. A thinking model that ends with everything in the reasoning channel and no content is asked once for the reply.
  • A model told "carry on" retries a refused edit indefinitely. Tell it up front that nobody can approve, and withdraw the tool on a final refusal.
  • Small models drop the leading / of absolute paths (recovered only inside the project) and search code text as a regex (grep offers literal: true).
  • Compaction puts the prompt that started the turn back verbatim after an automatic summary, at most twice a turn; a summary alone loses the task.
  • Bun's fetch does not use the OS certificate store unless run with --use-system-ca; the build passes it, and tls.ca covers the rest.
  • Bun.spawn without env passes the environment as it was at startup. Pass the live environment explicitly.
  • Bun decodes text dropping a UTF-8 BOM. Files are read with ignoreBOM: true so an edit does not strip one. Lifted code carries its runtime's assumptions: Bun is not Node.

The TUI

  • Solid props are live getters. A dialog that cleared its own spec and then called spec.onSelect threw inside OpenTUI's key handler, which logs and carries on — a picker that closed without choosing. Take the callback before closing; test dialogs inside the real Root.
  • One terminal-size subscription, in Root, passed through context — one per component leaked listeners until Node printed a warning across the screen.
  • Test the state a program leaves behind, not only what it shows. A throw after OpenTUI turned on mouse reporting left the shell receiving every mouse move as input. runTui destroys the renderer before printing any fatal error, and the smoke test reads the terminal's modes after it quits.
  • No model, or an unreadable config, is a screen with a retry, never an exception.
  • Markdown inside a <scrollbox> needs scrollX={false}, and is parsed off the main thread — wait for its text, not its box.
  • OpenTUI's truncate elides the middle — right for a path, wrong for a sentence.
  • Emoji with a variation selector are one or two columns depending on the terminal; plain emoji are two everywhere. icons: plain exists for fonts without emoji.
  • Streamed text is flushed at about 30 fps; re-parsing markdown per token stutters on a small machine.

Code and data

  • A field nothing reads is a promise. A store method with no caller, a config key accepted by the schema and read by nothing, a status-bar slot never filled — grep for the reader before shipping it, and test the stored row, not a screen with a fallback.
  • Validate after substitution. Files are validated before {env:}/{file:} are filled (so one connection's missing variable cannot stop the others); a URL field therefore accepts a placeholder, and what it becomes is checked after, turning off only the entry that is still wrong.
  • String.replaceAll reads $&, $$, $` and $' in the replacement. Escape $, or pass a function.
  • \w is ASCII in JavaScript. Patterns ported from Python use \p{L} with the u flag, or one accented word hides an injection.
  • zod 4's z.record(z.enum(…)) is exhaustive; a partial map is z.partialRecord.
  • Messages stored in one millisecond tie on time; order by id as well.
  • An rc wrapper replaces the shell's own startup order. The shell integration (src/acp/shellrc.ts) is tested by running real bash, zsh and fish against the startup files people really have. In bash, PS0 can set a variable in the shell itself only through an arithmetic array subscript.

Testing conventions

  • bun test runs unit tests and end-to-end tests against tests/fake-provider.ts, a scripted endpoint. Recorded real streams replay through the same path. tests/tui/ drives the real Root.
  • Bun's toMatchObject with asymmetric matchers writes the matchers into the object — use plain toMatch / toContain on an object that is looked at again.
  • The smoke test (scripts/smoke.ts, scripts/drive.ts) starts the compiled binary on a fresh home: first run, prompt, approval, answer and quit in a PTY at both sizes, nothing wider than the terminal, the status bar on the last row, and the terminal's modes restored.
  • MCP and OAuth are tested against the SDK's own servers, including its example OAuth server, so the whole sign-in runs without a public service.
  • Several suites may share a small machine. The per-test timeout is 20 s.