1 Commits
Author SHA1 Message Date
HomerandClaude Opus 5.5 9970bb43c6 Data groups: a provider's models read only their own group's data
Every connection is in a data group. Its models are handed, and can find,
only that group's memories, notes, skills, knowledge, reports and
personality -- by search and by id. A chat stays in the group it was
started in: switching its model, the endpoint fallback, the crowd, friends,
bases and the @ menu all stay inside it, and a chat whose model has moved
is refused rather than sent. A group may name its own embedder and image
reviewer. data.manage lets a person make personal groups, remap
connections for themselves and move their own records.

Also: a search no longer mixes two embedders of the same width.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 16:07:55 +00:00
61 changed files with 3595 additions and 200 deletions
+64
View File
@@ -16,6 +16,70 @@ for 1.0.0 have something to be assembled from.
## Unreleased
## 1.10.0
Data groups: a provider's models read only the data of the group their
connection is in. This is the first of three releases. The next two add rules
for which model may talk to which, and helpers on a different model.
- **Every connection is in a data group, and its models read only that group.**
This covers memories, notes, skills, knowledge bases and their documents,
reports, and the personality and impression a model keeps with you. It applies
both to what a model is handed at the start of a turn and to what its tools
can find. A tool can no longer open a record from another group by its id.
**Admin → Data groups** makes groups, puts connections in them, and shows
**where each group's data goes**, flagging a service whose connection sits in
a different group. Every instance starts with one group, called Default, with
every connection and every existing record in it. Until you make a second
group nothing changes, and nothing about groups is shown in the library.
- **A chat stays in the group it was started in.**
- Inside a chat, the model menu offers only models in the chat's group and
names the others underneath, with the reason.
- Switching to one is refused, because it would be sent the whole
conversation.
- The crowd picker, `ask_friend`, the list of other models, attaching a
knowledge base and the `@` menu all stay within the chat's group.
- If a chat's model is later moved into another group, the next reply is
refused with an explanation rather than sent.
- When a chat's own connection has gone, the fallback to "any connection
offering the same model" now only considers connections in the chat's
group. Before, it could silently move a conversation to another provider.
- **A group can have its own embedding model and image reviewer.** The embedder
is sent the full text of everything it indexes, and the reviewer every picture
with its prompt. A group that names neither uses the instance's.
- **Schedules run in their model's group.** A schedule's reports land in that
group. A schedule whose model is in another group than Messages cannot post
into Messages, and says so. The model that works out a schedule's timing from
plain words is now picked from the schedule's own group, not simply the first
pinned model.
- **Your own arrangement, with a new permission.** With *Manage their own data
groups* (off by default), a person can:
- make personal groups nobody else sees;
- choose which group each connection reads for them alone;
- move their own notes, skills and knowledge bases between groups.
Everybody can see which group each connection reads, on the new **Data** tab
in Settings, once there is more than one group. New notes, skills, knowledge
bases and memories can be put in any group you can use.
- **A search no longer mixes two embedding models of the same width.** It
checked only a vector's width, so two different 1024-wide models scored against
each other and returned confident nonsense. That used to need a model change
with a rebuild pending; with an embedder per group it would have been ordinary.
A query now carries the model that made it, and only that model's pieces are
scored.
- The crowd picker on the new-chat screen now leaves out the model the chat is
being started on. It was offering the chat's own model as a member.
- Four lines on the crowd page and in the crowd picker were still in English on
a Slovak instance. They are translated now.
- Known limits:
- A skill name and a knowledge base name are still unique per person across
all groups. The database constraint cannot be changed without rebuilding the
table.
- Speech to text and text to speech are not grouped: audio is sent to the
speech server and not kept.
- Moving a whole knowledge base to another group leaves its documents indexed
for the old group's embedder until the index is rebuilt.
## 1.9.1
- **A temporary chat can be started on any model.** On the new-chat screen,
+43
View File
@@ -261,6 +261,10 @@ def authored_sideways() -> tuple[str, ...]:
return tuple(selectors)
# The chat `build_client` seeds, so a run can name `/chat/<SHOOT_CHAT>`.
SHOOT_CHAT = "5" * 32
def build_client():
import lembas.config as config_mod
@@ -293,6 +297,45 @@ def build_client():
for name in ("gemma4-moe", "qwen3-coder"):
db.add(Model(connection_id=connection.id, model_id=name, display_name=name))
# A second data group with a hosted model in it, a note in each group and a
# chat with a fixed id (`SHOOT_CHAT`). Without them every grouped screen --
# the chips, the Data tab, the models named under the picker as being in
# another group -- is rendered with nothing in it, which is the "a screen the
# test never renders is unchecked" lesson again.
from lembas.db.models import Chat, DataGroup, Note, User
with session_scope() as db:
db.add(DataGroup(id="hosted", name="Hosted providers"))
hosted = Connection(
name="hosted",
base_url="http://127.0.0.1:2",
api_key_encrypted="",
data_group_id="hosted",
)
db.add(hosted)
db.flush()
db.add(
Model(
connection_id=hosted.id,
model_id="deepseek-flash",
display_name="DeepSeek Flash",
)
)
owner = db.query(User).first()
db.add(Note(owner_id=owner.id, title="A note at home", body="x", data_group_id="default"))
db.add(Note(owner_id=owner.id, title="A hosted note", body="x", data_group_id="hosted"))
local = db.query(Connection).filter_by(name="local").first()
db.add(
Chat(
id=SHOOT_CHAT,
user_id=owner.id,
model_id="gemma4-moe",
connection_id=local.id,
data_group_id="default",
title="A chat to measure",
)
)
# 🚨 The suggestion cards are seeded by the startup hook, and `TestClient(app)`
# runs a lifespan only inside a `with` block -- so every shot of the new-chat
# screen ever taken by this script was of a page with its cards missing. That
+1 -1
View File
@@ -1,3 +1,3 @@
"""LLeMbas - a Middle-earth themed web UI for OpenAI-compatible LLM endpoints."""
__version__ = "1.9.1"
__version__ = "1.10.0"
+12 -1
View File
@@ -13,7 +13,7 @@ from sqlalchemy.orm import Session as DBSession
from lembas.api.deps import AdminUser, Db
from lembas.db.models import Connection, Model, User
from lembas.services import settings_store
from lembas.services import data_groups, settings_store
from lembas.services.crypto import UNCHANGED_SENTINEL, decrypt, encrypt, mask
from lembas.services.llm.openai_client import Endpoint, LLMError, context_from, list_models
from lembas.web import i18n
@@ -109,6 +109,7 @@ async def connections_page(request: Request, db: Db, user: AdminUser, message: s
},
"message": message,
"unchanged": UNCHANGED_SENTINEL,
"data_group_choices": data_groups.instance_groups(db),
},
)
@@ -175,8 +176,17 @@ async def update_connection(
unload_url: str = Form(""),
unload_method: str = Form("POST"),
extra_headers: str = Form(""),
data_group_id: str = Form(""),
) -> Response:
connection = _connection(db, connection_id)
# Which of the instance's data groups this provider reads. Empty is "not
# submitted" -- an older page -- and leaves it alone; a personal group is
# somebody else's arrangement and cannot be chosen here.
chosen = data_group_id.strip()
if chosen:
group = data_groups.get(db, chosen)
if group is not None and group.owner_id is None:
connection.data_group_id = group.id
connection.name = name.strip()[:120] or connection.name
connection.base_url = base_url.strip().rstrip("/")
connection.enabled = enabled
@@ -227,6 +237,7 @@ async def test_connection(
"message": message,
"message_kind": "error" if error else "success",
"unchanged": UNCHANGED_SENTINEL,
"data_group_choices": data_groups.instance_groups(db),
},
)
+263
View File
@@ -0,0 +1,263 @@
"""Data groups: which provider may read which part of the people's data.
List plus detail, the shape every admin list here follows. The list says what
each group holds; the detail says which providers read it, which services send
its data somewhere else, and lets an administrator change both.
The one sentence this page exists to make answerable is "which provider has
seen this?". So the detail page lists every place a group's data can leave by --
its connections, its embedder and its reviewer -- and flags the ones whose
connection is in a *different* group, because those are the ones nobody would
think to check.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from fastapi import APIRouter, Form, HTTPException, Request, Response, status
from fastapi.responses import RedirectResponse
from sqlalchemy import func, select
from lembas.api.deps import AdminUser, Db
from lembas.db.models import Connection, DataGroup, Model, User
from lembas.services import data_groups, settings_store
from lembas.web.i18n import t
from lembas.web.templating import render
log = logging.getLogger(__name__)
router = APIRouter(prefix="/admin/data-groups", tags=["admin-data-groups"])
@dataclass(frozen=True)
class Exit:
"""One way a group's data leaves it: a provider, and whether it is outside."""
what: str
model: str
connection: str
elsewhere: str # the other group's name, or "" when it is this group's own
def labels() -> dict[str, str]:
"""The words `data_groups.COUNTED` and `exits` produce, in the reader's language.
Written out as literal `t()` calls because the templates look them up by
key, and a key that exists only as data is one the catalogue extractor never
finds -- so it would stay English forever, silently. Called per request, not
at import, because the language is the request's.
"""
return {
"chats": t("chats"),
"memories": t("memories"),
"notes": t("notes"),
"skills": t("skills"),
"knowledge bases": t("knowledge bases"),
"reports": t("reports"),
"connections": t("connections"),
"Chat models": t("Chat models"),
"Embedding": t("Embedding"),
"Image review": t("Image review"),
}
def _group(db, group_id: str) -> DataGroup:
group = data_groups.get(db, group_id)
if group is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "No such data group.")
return group
def _service_model(db, model_id: str, connection_id: str) -> Model | None:
if not model_id:
return None
return db.scalar(
select(Model)
.join(Connection)
.where(Model.model_id == model_id, Connection.enabled.is_(True))
.order_by(Model.connection_id != (connection_id or ""), Connection.position)
)
def exits(db, group: DataGroup) -> list[Exit]:
"""Every provider this group's data is sent to, as the instance has it set.
Personal remaps are not here: they belong to one person, change what that
person's providers read, and are listed on that person's own settings page.
"""
found: list[Exit] = []
for connection in db.scalars(select(Connection).order_by(Connection.position)):
if data_groups.for_connection(db, None, connection.id) == group.id:
found.append(Exit("Chat models", "", connection.name, ""))
extraction = settings_store.extraction(db)
images = settings_store.images(db)
services = (
(
"Embedding",
group.embedding_model_id or str(extraction.get("embedding_model_id") or ""),
group.embedding_connection_id if group.embedding_model_id else "",
),
(
"Image review",
group.review_model_id
or (str(images.get("review_model_id") or "") if images.get("review_enabled") else ""),
group.review_connection_id if group.review_model_id else "",
),
)
for what, model_id, connection_id in services:
model = _service_model(db, model_id, connection_id)
if model is None:
continue
other = data_groups.for_connection(db, None, model.connection_id)
found.append(
Exit(
what,
model.label,
model.connection.name if model.connection else "",
data_groups.name_of(db, other) if other != group.id else "",
)
)
return found
def _capable(db, capability: str) -> list[Model]:
rows = db.scalars(
select(Model)
.join(Connection)
.where(Model.enabled.is_(True), Connection.enabled.is_(True))
.order_by(Model.position, Model.model_id)
)
return [m for m in rows if (m.capabilities_json or {}).get(capability)]
@router.get("")
async def data_groups_page(request: Request, db: Db, user: AdminUser, saved: str = ""):
groups = data_groups.instance_groups(db)
personal = [g for g in data_groups.all_groups(db) if g.owner_id is not None]
owners = {
u.id: u
for u in db.scalars(select(User).where(User.id.in_([g.owner_id for g in personal])))
}
connections = {
group.id: db.scalar(
select(func.count())
.select_from(Connection)
.where(data_groups.condition(Connection, group.id))
)
or 0
for group in groups
}
return render(
request,
"admin/data_groups.html",
{
"groups": groups,
"personal": personal,
"owners": owners,
"connections": connections,
"counts": {g.id: data_groups.counts(db, g.id) for g in [*groups, *personal]},
"labels": labels(),
"saved": saved,
},
)
@router.post("")
async def create_group(db: Db, user: AdminUser, name: str = Form(...)) -> Response:
name = " ".join(name.split())[:120]
if not name:
raise HTTPException(status.HTTP_400_BAD_REQUEST, "A data group needs a name.")
position = db.scalar(select(func.coalesce(func.max(DataGroup.position), 0))) + 1
group = DataGroup(name=name, position=position)
db.add(group)
db.commit()
return RedirectResponse(f"/admin/data-groups/{group.id}", status_code=status.HTTP_303_SEE_OTHER)
@router.get("/{group_id}")
async def data_group_detail(
request: Request, db: Db, user: AdminUser, group_id: str, saved: str = "", error: str = ""
):
group = _group(db, group_id)
all_connections = list(db.scalars(select(Connection).order_by(Connection.position)))
return render(
request,
"admin/data_group_detail.html",
{
"group": group,
"owner": db.get(User, group.owner_id) if group.owner_id else None,
"connections": all_connections,
"member_ids": {
c.id
for c in all_connections
if data_groups.for_connection(db, None, c.id) == group.id
},
"embedders": _capable(db, "embeddings"),
"reviewers": _capable(db, "vision"),
"exits": exits(db, group),
"counts": data_groups.counts(db, group.id),
"in_use": data_groups.in_use(db, group.id),
"labels": labels(),
"saved": saved,
"error": error,
},
)
def _pair(value: str) -> tuple[str, str]:
"""`model_id|connection_id` from a select, or two blanks for the fallback."""
model_id, _, connection_id = (value or "").partition("|")
return model_id.strip()[:300], connection_id.strip()[:32]
@router.post("/{group_id}")
async def save_group(request: Request, db: Db, user: AdminUser, group_id: str) -> Response:
"""Name, description, services and -- for an instance group -- its connections.
Connections are read from a list that is always submitted, so unticking the
last one is a signal and not an absence. A connection taken out of a group
goes back to the default one, never to "no group": there is no such thing.
"""
group = _group(db, group_id)
form = await request.form()
name = " ".join(str(form.get("name") or "").split())[:120]
if name:
group.name = name
group.description = str(form.get("description") or "").strip()[:2000]
group.embedding_model_id, group.embedding_connection_id = _pair(
str(form.get("embedding") or "")
)
group.review_model_id, group.review_connection_id = _pair(str(form.get("reviewer") or ""))
if group.owner_id is None and "connections_sent" in form:
wanted = {str(v) for v in form.getlist("connection_ids") if v}
for connection in db.scalars(select(Connection)):
current = connection.data_group_id or data_groups.DEFAULT_GROUP
if connection.id in wanted:
connection.data_group_id = group.id
elif current == group.id and not group.is_default:
connection.data_group_id = data_groups.DEFAULT_GROUP
db.commit()
return RedirectResponse(
f"/admin/data-groups/{group.id}?saved=1", status_code=status.HTTP_303_SEE_OTHER
)
@router.post("/{group_id}/delete")
async def delete_group(db: Db, user: AdminUser, group_id: str) -> Response:
group = _group(db, group_id)
try:
data_groups.delete(db, group)
except ValueError as exc:
from urllib.parse import quote
return RedirectResponse(
f"/admin/data-groups/{group_id}?error={quote(str(exc))}",
status_code=status.HTTP_303_SEE_OTHER,
)
return RedirectResponse(
"/admin/data-groups?saved=deleted", status_code=status.HTTP_303_SEE_OTHER
)
+41 -7
View File
@@ -34,9 +34,9 @@ from lembas.security import permissions
from lembas.services import audio as audio_service
from lembas.services import chat as chat_service
from lembas.services import compaction as compaction_service
from lembas.services import data_groups, interaction, settings_store, sse
from lembas.services import files as files_service
from lembas.services import generation as generation_service
from lembas.services import interaction, settings_store, sse
from lembas.services import metrics as metrics_service
from lembas.services import prompts as prompts_service
from lembas.services import reports as reports_service
@@ -222,6 +222,14 @@ def _new_chat(
folder_id=folder.id if folder is not None else None,
model_id=chosen[0] if chosen else "",
connection_id=chosen[1] if chosen else None,
# The chat's data group is its model's, fixed now. It is what every
# later turn reads memories and notes from, and what decides which models
# this chat may be switched to -- see services/data_groups.py.
data_group_id=(
data_groups.for_pair(db, user, chosen[0], chosen[1])
if chosen
else data_groups.DEFAULT_GROUP
),
temporary=temporary,
kind=KIND_AGENT if profile is not None else KIND_CHAT,
ssh_profile_id=profile.id if profile is not None else None,
@@ -657,8 +665,12 @@ async def attach_base(
if not permissions.has(db, user, "library.use"):
raise HTTPException(status.HTTP_403_FORBIDDEN, "You may not use the library.")
# Only a base in the chat's own data group: its documents are what this
# chat's model would search, and another group's are not its to read.
base = db.scalar(
documents_service.visible_bases(db, user).where(KnowledgeBase.id == base_id)
documents_service.visible_bases(db, user, data_groups.for_chat(db, chat)).where(
KnowledgeBase.id == base_id
)
)
if base is None:
raise HTTPException(status.HTTP_404_NOT_FOUND, "That knowledge base is not available.")
@@ -1536,8 +1548,13 @@ def _apply_crowd(db: DBSession, chat: Chat, user: User, values: list[str]) -> No
from lembas.db.models import CrowdMember
settings = settings_store.crowd(db)
# Only models in the chat's own data group: a member is sent the whole
# conversation, so one from another group would carry it to that provider.
reachable = {
model.model_id: model for model in chat_service.available_models(db, user)
model.model_id: model
for model in chat_service.available_models(
db, user, data_groups.for_chat(db, chat)
)
}
wanted: list[str] = []
for value in values:
@@ -2125,11 +2142,28 @@ async def update_chat(request: Request, db: Db, user: RequiredUser, chat_id: str
)
# Checked against what this user can reach, not merely what exists --
# otherwise the picker is advisory and a crafted request bypasses it.
# And within the chat's own data group: the new model would be sent the
# whole history, which is exactly what a group keeps from its provider.
group = data_groups.for_chat(db, chat)
match = next(
(m for m in chat_service.available_models(db, user) if m.model_id == model_id),
(
m
for m in chat_service.available_models(db, user, group)
if m.model_id == model_id
),
None,
)
if match is None:
elsewhere = any(
m.model_id == model_id for m in chat_service.available_models(db, user)
)
if elsewhere:
raise HTTPException(
status.HTTP_409_CONFLICT,
f"That model is in another data group than this chat "
f"({data_groups.name_of(db, group)}), so it cannot be given this "
f"chat's history. Start a new chat with it instead.",
)
raise HTTPException(status.HTTP_403_FORBIDDEN, "That model is not available to you.")
chat.model_id = model_id
chat.connection_id = match.connection_id
@@ -2152,9 +2186,9 @@ async def update_chat(request: Request, db: Db, user: RequiredUser, chat_id: str
chat.knowledge_bases = (
list(
db.scalars(
documents_service.visible_bases(db, user).where(
KnowledgeBase.id.in_(wanted)
)
documents_service.visible_bases(
db, user, data_groups.for_chat(db, chat)
).where(KnowledgeBase.id.in_(wanted))
)
)
if wanted
+57 -18
View File
@@ -21,8 +21,8 @@ from sqlalchemy import select
from lembas.api.deps import Db, RequiredUser, require_permission
from lembas.db.models import Attachment, Chat, Document, KnowledgeBase, Note
from lembas.security import permissions
from lembas.services import data_groups, settings_store
from lembas.services import files as files_service
from lembas.services import settings_store
from lembas.services.fetch import FetchError, fetch
from lembas.services.library import documents as documents_service
from lembas.services.library import notes as notes_service
@@ -119,7 +119,7 @@ async def attach_link(
@router.post("/from-knowledge", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_knowledge(
request: Request, db: Db, user: RequiredUser, document_id: str = Form(""),
chat_id: str = Form(""),
chat_id: str = Form(""), model_id: str = Form(""),
) -> Response:
"""Attach a library document to the message being composed.
@@ -127,7 +127,8 @@ async def attach_from_knowledge(
conversation because a document was later edited or deleted -- the same
reason a PDF's text is extracted once at upload rather than per request.
"""
document = documents_service.get(db, document_id, user)
group = data_groups.for_composer(db, user, chat_id, model_id)
document = documents_service.get(db, document_id, user, group)
if document is None:
return templates.TemplateResponse(
request,
@@ -163,7 +164,12 @@ def _not_available(request: Request, what: str) -> Response:
@router.post("/from-note", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_note(
request: Request, db: Db, user: RequiredUser, note_id: str = Form(""), chat_id: str = Form("")
request: Request,
db: Db,
user: RequiredUser,
note_id: str = Form(""),
chat_id: str = Form(""),
model_id: str = Form(""),
) -> Response:
"""Attach a note the model wrote earlier.
@@ -171,7 +177,9 @@ async def attach_from_note(
document, and a transcript that changes underneath itself because somebody
tidied a note later is the thing all of this is arranged to prevent.
"""
note = notes_service.get(db, note_id, user)
note = notes_service.get(
db, note_id, user, data_groups.for_composer(db, user, chat_id, model_id)
)
if note is None:
return _not_available(request, "note")
@@ -226,7 +234,12 @@ async def attach_from_scratch(
@router.post("/from-skill", dependencies=[Depends(require_permission("files.upload"))])
async def attach_from_skill(
request: Request, db: Db, user: RequiredUser, skill_id: str = Form(""), chat_id: str = Form("")
request: Request,
db: Db,
user: RequiredUser,
skill_id: str = Form(""),
chat_id: str = Form(""),
model_id: str = Form(""),
) -> Response:
"""Hand a skill over directly, rather than hoping the model fetches it.
@@ -234,7 +247,9 @@ async def attach_from_skill(
a body on demand -- but only if the model decides to. `@` is the reader
saying "use this one", which is a different act and deserves a way to say it.
"""
skill = skills_service.get(db, skill_id, user)
skill = skills_service.get(
db, skill_id, user, data_groups.for_composer(db, user, chat_id, model_id)
)
if skill is None:
return _not_available(request, "skill")
@@ -259,6 +274,7 @@ async def attach_from_attachment(
user: RequiredUser,
attachment_id: str = Form(""),
chat_id: str = Form(""),
model_id: str = Form(""),
) -> Response:
"""Point at something already in this conversation, without uploading again.
@@ -269,6 +285,12 @@ async def attach_from_attachment(
original = db.get(Attachment, attachment_id)
if original is None or original.user_id != user.id:
return _not_available(request, "attachment")
# Another chat's attachment is that chat's data, in that chat's group.
source = db.get(Chat, original.chat_id) if original.chat_id else None
if source is not None and data_groups.for_chat(db, source) != data_groups.for_composer(
db, user, chat_id, model_id
):
return _not_available(request, "attachment")
return _chip(
request,
@@ -280,15 +302,24 @@ async def attach_from_attachment(
@router.get("/knowledge-picker", dependencies=[Depends(require_permission("files.upload"))])
async def knowledge_picker(
request: Request, db: Db, user: RequiredUser, q: str = "", chat_id: str = ""
request: Request,
db: Db,
user: RequiredUser,
q: str = "",
chat_id: str = "",
model_id: str = "",
) -> Response:
"""The list of documents shown by the composer's Knowledge option."""
"""The list of documents shown by the composer's Knowledge option.
Only the composer's own data group: see `data_groups.for_composer`.
"""
group = data_groups.for_composer(db, user, chat_id, model_id)
if q.strip():
found = documents_service.search(db, user, q, limit=20)
found = documents_service.search(db, user, q, limit=20, group=group)
else:
found = list(
db.scalars(
documents_service.visible(db, user)
documents_service.visible(db, user, group=group)
.order_by(Document.created_at.desc())
.limit(20)
)
@@ -311,6 +342,7 @@ async def mention_picker(
chat_id: str = "",
profile_id: str = "",
project_dir: str = "",
model_id: str = "",
) -> Response:
"""What `@` offers: files under the project directory, and the library.
@@ -348,24 +380,29 @@ async def mention_picker(
skills: list = []
bases: list = []
if permissions.has(db, user, "library.use"):
# The library half offers only the composer's own data group: whatever is
# picked becomes part of the conversation and goes to the chat's model.
group = data_groups.for_composer(db, user, chat_id, model_id)
if needle:
documents = documents_service.search(db, user, q, limit=10)
notes = notes_service.search(db, user, q, limit=5)
skills = skills_service.search(db, user, q, limit=5)
documents = documents_service.search(db, user, q, limit=10, group=group)
notes = notes_service.search(db, user, q, limit=5, group=group)
skills = skills_service.search(db, user, q, limit=5, group=group)
else:
documents = list(
db.scalars(
documents_service.visible(db, user)
documents_service.visible(db, user, group=group)
.order_by(Document.created_at.desc())
.limit(10)
)
)
notes = list(
db.scalars(
notes_service.visible(db, user).order_by(Note.updated_at.desc()).limit(5)
notes_service.visible(db, user, group)
.order_by(Note.updated_at.desc())
.limit(5)
)
)
skills = list(db.scalars(skills_service.visible(db, user).limit(5)))
skills = list(db.scalars(skills_service.visible(db, user, group).limit(5)))
# A whole base is a *reference*, not a copy: attaching one scopes the
# chat to it and the model searches inside it. Dumping the contents of
@@ -377,7 +414,9 @@ async def mention_picker(
bases = [
base
for base in db.scalars(
documents_service.visible_bases(db, user).order_by(KnowledgeBase.name)
documents_service.visible_bases(db, user, group).order_by(
KnowledgeBase.name
)
)
if not needle or needle in base.name.lower()
][:5]
+146 -16
View File
@@ -33,8 +33,8 @@ from lembas.db.models import (
User,
)
from lembas.security import permissions
from lembas.services import data_groups, settings_store, sharing
from lembas.services import files as files_service
from lembas.services import settings_store, sharing
from lembas.services.fetch import FetchError, fetch
from lembas.services.library import documents as documents_service
from lembas.services.library import memories as memories_service
@@ -60,6 +60,46 @@ def _page(db: DBSession, query, page: int):
return rows, {"page": page, "pages": pages, "total": total}
def _groups(db: DBSession, user: User) -> dict:
"""What `library/_group.html` needs on every library page.
`several_groups` false is the ordinary instance, and then nothing about
groups is rendered anywhere in the library.
"""
usable = data_groups.usable(db, user)
return {
"several_groups": len(usable) > 1,
"data_groups": usable,
"group_names": {group.id: group.name for group in data_groups.all_groups(db)},
"may_move_groups": data_groups.may_manage(db, user),
}
def _chosen_group(db: DBSession, user: User, value: str) -> str:
"""A submitted group, if this person may use it; otherwise the default."""
value = (value or "").strip()
if value and data_groups.may_use(db, user, value):
return value
return data_groups.DEFAULT_GROUP
def _move(db: DBSession, user: User, row, value) -> None:
"""Move a record into another group, when that was asked and is allowed.
`None` is a form that did not carry the field -- a single-group instance,
or somebody without `data.manage` -- and leaves the record where it is.
"""
if value is None or not data_groups.may_manage(db, user):
return
wanted = str(value).strip()
if wanted and data_groups.may_use(db, user, wanted):
row.data_group_id = wanted
def _group_filter(query, model, group: str):
return query.where(data_groups.condition(model, group)) if group else query
def _shared_context(db: DBSession, user: User, resource, kind: str) -> dict:
"""What the share placeholder needs, which is now three facts.
@@ -87,7 +127,12 @@ async def library_home(user: RequiredUser):
# FastAPI matches in registration order and this has bitten before.
@router.get("/library/knowledge")
async def knowledge_list(
request: Request, db: Db, user: RequiredUser, error: str = "", shared: bool = False
request: Request,
db: Db,
user: RequiredUser,
error: str = "",
shared: bool = False,
group: str = "",
):
"""The bases, not the documents. A library is a set of places first.
@@ -100,6 +145,7 @@ async def knowledge_list(
if shared
else documents_service.visible_bases(db, user)
)
query = _group_filter(query, KnowledgeBase, group)
bases = list(db.scalars(query.order_by(KnowledgeBase.name)))
counts = {
base.id: db.scalar(
@@ -117,6 +163,8 @@ async def knowledge_list(
"counts": counts,
"shared": shared,
"error": error,
"group": group,
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -124,11 +172,19 @@ async def knowledge_list(
@router.post("/api/library/bases")
async def create_base(
db: Db, user: RequiredUser, name: str = Form(""), description: str = Form("")
db: Db,
user: RequiredUser,
name: str = Form(""),
description: str = Form(""),
data_group_id: str = Form(""),
) -> Response:
try:
base = documents_service.create_base(
db, owner=user, name=name, description=description
db,
owner=user,
name=name,
description=description,
group=_chosen_group(db, user, data_group_id),
)
except ValueError as exc:
from urllib.parse import quote
@@ -163,6 +219,7 @@ async def knowledge_detail(request: Request, db: Db, user: RequiredUser, documen
.order_by(KnowledgeBase.name)
)
),
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -202,6 +259,7 @@ async def base_detail(
"q": q,
"pager": pager,
**_shared_context(db, user, base, "base"),
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -220,6 +278,8 @@ async def update_base(request: Request, db: Db, user: RequiredUser, base_id: str
if name:
base.name = name
base.description = str(form.get("description", "")).strip()[:2000]
# A base moves with every document in it: they have no group of their own.
_move(db, user, base, form.get("data_group_id"))
db.commit()
return RedirectResponse(
f"/library/knowledge/{base.id}", status_code=status.HTTP_303_SEE_OTHER
@@ -302,7 +362,17 @@ async def update_document(
wanted = str(form.get("base_id", "")).strip()
if wanted and wanted != document.base_id:
destination = documents_service.get_base(db, wanted, user)
if destination is not None and sharing.can_write(destination, user):
current = db.get(KnowledgeBase, document.base_id) if document.base_id else None
# Into a base in another data group is a move between groups, which
# changes which providers may read it -- `data.manage`, like any move.
crosses = current is not None and destination is not None and (
data_groups.group_of(current) != data_groups.group_of(destination)
)
if (
destination is not None
and sharing.can_write(destination, user)
and (not crosses or data_groups.may_manage(db, user))
):
document.base_id = destination.id
db.commit()
@@ -352,6 +422,7 @@ async def notes_list(
q: str = "",
page: int = 1,
shared: bool = False,
group: str = "",
):
"""`shared=1` narrows to what other people have given this reader.
@@ -363,7 +434,12 @@ async def notes_list(
"""
if q.strip():
rows = notes_service.search(
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
db,
user,
q,
limit=PAGE_SIZE,
vector=await retrieval.embed_query(db, q),
group=group or None,
)
pager = {"page": 1, "pages": 1, "total": len(rows)}
else:
@@ -372,6 +448,7 @@ async def notes_list(
if shared
else notes_service.visible(db, user)
)
query = _group_filter(query, Note, group)
rows, pager = _page(db, query.order_by(Note.updated_at.desc()), page)
return render(
request,
@@ -382,17 +459,25 @@ async def notes_list(
"q": q,
"shared": shared,
"pager": pager,
"group": group,
**_groups(db, user),
**sidebar_context(db, user),
},
)
@router.get("/library/notes/new")
async def new_note(request: Request, db: Db, user: RequiredUser):
async def new_note(request: Request, db: Db, user: RequiredUser, group: str = ""):
return render(
request,
"library/note_detail.html",
{"section": "notes", "note": None, **sidebar_context(db, user)},
{
"section": "notes",
"note": None,
"group": group,
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -409,6 +494,7 @@ async def note_detail(request: Request, db: Db, user: RequiredUser, note_id: str
"note": note,
"body_html": render_markdown(note.body),
**_shared_context(db, user, note, "note"),
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -416,9 +502,20 @@ async def note_detail(request: Request, db: Db, user: RequiredUser, note_id: str
@router.post("/api/library/notes")
async def create_note(
db: Db, user: RequiredUser, title: str = Form(""), body: str = Form("")
db: Db,
user: RequiredUser,
title: str = Form(""),
body: str = Form(""),
data_group_id: str = Form(""),
) -> Response:
note = notes_service.create(db, owner=user, title=title, body=body, author=AUTHOR_USER)
note = notes_service.create(
db,
owner=user,
title=title,
body=body,
author=AUTHOR_USER,
group=_chosen_group(db, user, data_group_id),
)
return RedirectResponse(f"/library/notes/{note.id}", status_code=status.HTTP_303_SEE_OTHER)
@@ -431,6 +528,7 @@ async def update_note(request: Request, db: Db, user: RequiredUser, note_id: str
raise HTTPException(status.HTTP_403_FORBIDDEN, "That note is not yours to change.")
form = await request.form()
_move(db, user, note, form.get("data_group_id"))
notes_service.update(db, note, title=str(form.get("title", "")), body=str(form.get("body", "")))
return RedirectResponse(f"/library/notes/{note.id}", status_code=status.HTTP_303_SEE_OTHER)
@@ -453,6 +551,7 @@ async def skills_list(
q: str = "",
page: int = 1,
shared: bool = False,
group: str = "",
):
"""`shared=1` narrows to what other people have given this reader.
@@ -464,7 +563,12 @@ async def skills_list(
"""
if q.strip():
rows = skills_service.search(
db, user, q, limit=PAGE_SIZE, vector=await retrieval.embed_query(db, q)
db,
user,
q,
limit=PAGE_SIZE,
vector=await retrieval.embed_query(db, q),
group=group or None,
)
pager = {"page": 1, "pages": 1, "total": len(rows)}
else:
@@ -473,6 +577,7 @@ async def skills_list(
if shared
else skills_service.visible(db, user)
)
query = _group_filter(query, Skill, group)
rows, pager = _page(db, query.order_by(Skill.name), page)
return render(
request,
@@ -483,17 +588,25 @@ async def skills_list(
"q": q,
"shared": shared,
"pager": pager,
"group": group,
**_groups(db, user),
**sidebar_context(db, user),
},
)
@router.get("/library/skills/new")
async def new_skill(request: Request, db: Db, user: RequiredUser):
async def new_skill(request: Request, db: Db, user: RequiredUser, group: str = ""):
return render(
request,
"library/skill_detail.html",
{"section": "skills", "skill": None, **sidebar_context(db, user)},
{
"section": "skills",
"skill": None,
"group": group,
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -510,6 +623,7 @@ async def skill_detail(request: Request, db: Db, user: RequiredUser, skill_id: s
"skill": skill,
"revisions": skill.revisions,
**_shared_context(db, user, skill, "skill"),
**_groups(db, user),
**sidebar_context(db, user),
},
)
@@ -522,10 +636,17 @@ async def create_skill(
name: str = Form(""),
description: str = Form(""),
body: str = Form(""),
data_group_id: str = Form(""),
) -> Response:
try:
skill = skills_service.create(
db, owner=user, name=name, description=description, body=body, author=AUTHOR_USER
db,
owner=user,
name=name,
description=description,
body=body,
author=AUTHOR_USER,
group=_chosen_group(db, user, data_group_id),
)
except skills_service.SkillError as exc:
raise HTTPException(status.HTTP_400_BAD_REQUEST, str(exc)) from exc
@@ -541,6 +662,7 @@ async def update_skill(request: Request, db: Db, user: RequiredUser, skill_id: s
raise HTTPException(status.HTTP_403_FORBIDDEN, "That skill is not yours to change.")
form = await request.form()
_move(db, user, skill, form.get("data_group_id"))
skills_service.update(
db,
skill,
@@ -581,9 +703,17 @@ async def delete_skill(db: Db, user: RequiredUser, skill_id: str) -> Response:
# Lives in Settings rather than in the library: it is a set of short facts about
# the reader, not content they collected.
@router.post("/api/library/memories")
async def add_memory(db: Db, user: RequiredUser, content: str = Form("")) -> Response:
async def add_memory(
db: Db, user: RequiredUser, content: str = Form(""), data_group_id: str = Form("")
) -> Response:
try:
memories_service.add(db, owner=user, content=content, author=AUTHOR_USER)
memories_service.add(
db,
owner=user,
content=content,
author=AUTHOR_USER,
group=_chosen_group(db, user, data_group_id),
)
except ValueError as exc:
from urllib.parse import quote
+80 -4
View File
@@ -18,6 +18,7 @@ from lembas.db.models import (
KIND_TASK,
KINDS,
Chat,
Connection,
Folder,
KnowledgeBase,
Message,
@@ -29,8 +30,8 @@ from lembas.services import branding as branding_service
from lembas.services import canvas as canvas_service
from lembas.services import chat as chat_service
from lembas.services import compaction as compaction_service
from lembas.services import data_groups, settings_store
from lembas.services import reports as reports_service
from lembas.services import settings_store
from lembas.services import suggestions as suggestions_service
from lembas.services.library import documents as documents_service
from lembas.services.schedule import clock
@@ -64,8 +65,18 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
"""
models = chat_service.available_models(db, user)
current = next((m for m in models if m.model_id == chat.model_id), None) if chat else None
# A chat that exists may only be switched to a model in its own data group:
# the new model would be sent the whole history. The rest are named below
# the picker rather than silently missing from it, so somebody looking for
# one learns where it went -- and that a new chat is how to reach it.
group = data_groups.for_chat(db, chat) if chat is not None else None
in_group = chat_service.available_models(db, user, group) if group is not None else models
in_group_ids = {m.id for m in in_group}
return {
"models": models,
"models": in_group,
"models_elsewhere": [m for m in models if m.id not in in_group_ids],
"chat_group_name": data_groups.name_of(db, group) if group is not None else "",
"several_groups": data_groups.several(db, user),
"current_model": current,
# Assistant bubbles show the avatar of the model that wrote them, which
# may not be the model the chat is set to now. Keyed by model_id, the
@@ -77,7 +88,9 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
"knowledge_bases": (
list(
db.scalars(
documents_service.visible_bases(db, user).order_by(KnowledgeBase.name)
documents_service.visible_bases(db, user, group).order_by(
KnowledgeBase.name
)
)
)
if permissions.has(db, user, "library.use")
@@ -85,7 +98,7 @@ def _chat_context(db: DBSession, user: User, chat: Chat | None) -> dict:
),
"attached_base_ids": [base.id for base in chat.knowledge_bases] if chat else [],
**_crowd_context(
db, user, chat, models, current.model_id if current is not None else ""
db, user, chat, in_group, current.model_id if current is not None else ""
),
# What *this* model takes, not the three every model used to be assumed
# to take. The vocabulary is per model -- gpt-oss has no `xhigh` and
@@ -200,6 +213,44 @@ def _scope_context(db: DBSession, user: User, chat: Chat | None) -> dict:
}
def _data_group_settings(db: DBSession, user: User) -> dict:
"""What the Data tab on /settings shows: which group each connection reads.
Every connection this person can reach a model on, with the instance's
choice beside their own. Their own is only in force while they hold
`data.manage`; without it the tab still says what applies to them, because
"which provider can read my notes?" is a question anybody may ask.
"""
from lembas.api import admin_data_groups
usable = data_groups.usable(db, user)
reachable = {m.connection_id for m in chat_service.available_models(db, user)}
connections = [
connection
for connection in db.scalars(select(Connection).order_by(Connection.position))
if connection.id in reachable
]
in_force = data_groups.connection_groups(db, user)
return {
"data_groups": usable,
"group_names": {group.id: group.name for group in data_groups.all_groups(db)},
"may_manage_groups": data_groups.may_manage(db, user),
"group_labels": admin_data_groups.labels(),
"group_counts": {
group.id: data_groups.counts(db, group.id, owner=user) for group in usable
},
"group_connections": [
{
"connection": connection,
"instance": connection.data_group_id or data_groups.DEFAULT_GROUP,
"chosen": data_groups.personal_map(user).get(connection.id, ""),
"in_force": in_force.get(connection.id, data_groups.DEFAULT_GROUP),
}
for connection in connections
],
}
def _crowd_context(
db: DBSession, user: User, chat: Chat | None, models: list, default_model_id: str = ""
) -> dict:
@@ -226,6 +277,17 @@ def _crowd_context(
# showing, so the list excludes it for the same reason it does in a chat:
# adding it would have it answer twice in a row.
own = chat.model_id if chat is not None else default_model_id
# Only the main model's data group: a member is sent the whole conversation.
# On the new-chat screen that is the group of the model the picker shows,
# which is the one the chat will be pinned to.
group = (
data_groups.for_chat(db, chat)
if chat is not None
else (data_groups.for_pair(db, user, own) if own else None)
)
if group is not None:
allowed = {m.id for m in chat_service.available_models(db, user, group)}
models = [model for model in models if model.id in allowed]
others = [model for model in models if model.model_id != own]
members = (
[
@@ -767,6 +829,16 @@ async def chat_index(
if preselected
else chat_service.DEFAULT_EFFORTS
),
# `_chat_context` had no chat and so no main model to build the crowd
# list from: the list offered here included the model it would be
# added to, and ignored which data group the new chat is going into.
**_crowd_context(
db,
user,
None,
context["models"],
preselected.model_id if preselected is not None else "",
),
"starting_temporary": temporary,
"temporary_toggle_url": new_chat_url(temporary="" if temporary else "1"),
"model_navigate_url": without_model + ("&" if "?" in without_model else "?") + "model=",
@@ -958,6 +1030,10 @@ async def settings_page(
# would leave no way to delete them.
"personalities": personas_service.personas_of(db, user),
"impressions": personas_service.impressions_for(db, user),
# A person's key carries the data group after the model id; the
# template shows the two apart rather than printing the raw key.
"split_key": personas_service.split_key,
**_data_group_settings(db, user),
# Sorted rather than left in set order, because a list of six
# hundred zones that is not alphabetical is one nobody can use.
"languages": i18n.LANGUAGES,
+78
View File
@@ -293,3 +293,81 @@ async def change_password(
path="/",
)
return response
# --- Data groups -------------------------------------------------------------
# A person's own arrangement of which provider may read which of their data.
# Everything here needs `data.manage`: it changes what a provider can see, and an
# instance that never granted it keeps the arrangement its administrator made.
def _refuse_without_manage(db, user) -> Response | None:
from lembas.services import data_groups
if data_groups.may_manage(db, user):
return None
return RedirectResponse(
"/settings?error=You+may+not+manage+your+own+data+groups.", status_code=303
)
@router.post("/data-groups")
async def set_data_groups(request: Request, db: Db, user: RequiredUser) -> Response:
"""Which group each connection reads, for this person.
One select per connection, named `group__<connection id>`; empty means
"follow the instance", which removes the entry rather than storing a copy of
the administrator's choice -- a copy would stop following it the day it
changed.
"""
from lembas.services import data_groups
refused = _refuse_without_manage(db, user)
if refused is not None:
return refused
form = await request.form()
chosen: dict[str, str] = {}
for key, value in form.items():
if not key.startswith("group__"):
continue
connection_id, group_id = key.removeprefix("group__"), str(value).strip()
if group_id and data_groups.may_use(db, user, group_id):
chosen[connection_id] = group_id
user.settings_json = {**(user.settings_json or {}), data_groups.SETTING_KEY: chosen}
db.commit()
return RedirectResponse("/settings?saved=Data+groups+updated.", status_code=303)
@router.post("/data-groups/new")
async def create_personal_group(
db: Db, user: RequiredUser, name: str = Form("")
) -> Response:
from lembas.db.models import DataGroup
refused = _refuse_without_manage(db, user)
if refused is not None:
return refused
name = " ".join(name.split())[:120]
if not name:
return RedirectResponse("/settings?error=A+data+group+needs+a+name.", status_code=303)
db.add(DataGroup(name=name, owner_id=user.id))
db.commit()
return RedirectResponse("/settings?saved=Data+group+created.", status_code=303)
@router.post("/data-groups/{group_id}/delete")
async def delete_personal_group(db: Db, user: RequiredUser, group_id: str) -> Response:
"""Remove one of this person's own groups -- only once nothing is in it."""
from urllib.parse import quote
from lembas.services import data_groups
refused = _refuse_without_manage(db, user)
if refused is not None:
return refused
group = data_groups.get(db, group_id)
if group is None or group.owner_id != user.id:
return RedirectResponse("/settings?error=No+such+data+group.", status_code=303)
try:
data_groups.delete(db, group)
except ValueError as exc:
return RedirectResponse(f"/settings?error={quote(str(exc))}", status_code=303)
return RedirectResponse("/settings?saved=Data+group+deleted.", status_code=303)
+1 -1
View File
@@ -237,7 +237,7 @@ async def describe_schedule(request: Request, db: Db, user: RequiredUser):
described = str(form.get("request") or "").strip()
template = prompts_service.resolve(db, "task.schedule_compile")
resolved = compile_service.endpoint_for(db, user)
resolved = compile_service.endpoint_for(db, user, str(form.get("model_id") or "").strip())
if resolved is None:
compiled = compile_service.Compiled(
instruction=described,
+4
View File
@@ -36,6 +36,7 @@ from lembas.db.models.chat import (
Message,
)
from lembas.db.models.connection import Connection, Model, model_groups
from lembas.db.models.data_group import DEFAULT_GROUP, DataGroup, InDataGroup
from lembas.db.models.image import ImageWorkflow
from lembas.db.models.library import (
AUTHOR_MODEL,
@@ -167,6 +168,9 @@ __all__ = [
"CrowdMember",
"Job",
"Connection",
"DEFAULT_GROUP",
"DataGroup",
"InDataGroup",
"CustomTool",
"CHUNK_DOCUMENT",
"CHUNK_KINDS",
+2 -1
View File
@@ -17,6 +17,7 @@ from sqlalchemy import (
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.models.data_group import InDataGroup
from lembas.db.types import JSONDict, JSONList
if TYPE_CHECKING:
@@ -173,7 +174,7 @@ class Folder(UUIDPrimaryKey, Timestamps, Base):
return f"<Folder {self.name}>"
class Chat(UUIDPrimaryKey, Timestamps, Base):
class Chat(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
__tablename__ = "chats"
user_id: Mapped[str] = mapped_column(
+2 -1
View File
@@ -19,6 +19,7 @@ from sqlalchemy import (
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.models.data_group import InDataGroup
from lembas.db.types import JSONDict, JSONList
if TYPE_CHECKING:
@@ -36,7 +37,7 @@ model_groups = Table(
)
class Connection(UUIDPrimaryKey, Timestamps, Base):
class Connection(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""A configured upstream endpoint speaking the OpenAI HTTP API.
Works for api.openai.com as well as LM Studio, vLLM, llama.cpp, Ollama's
+86
View File
@@ -0,0 +1,86 @@
"""Data groups: which provider may read which of a person's data.
A connection belongs to a data group, and so does everything a model can be
handed about a person -- memories, notes, skills, knowledge bases, reports, the
personality and impression a model keeps, and the chats themselves. A model
reads only the rows of the group its own connection is in. Two providers in one
group see the same data; two in different groups never see each other's.
**Data belongs to a group, not to a connection.** Moving a connection into
another group does not carry anything with it -- that provider simply starts
reading the other group. That is the only reading under which "which provider
has seen this?" has an answer that does not depend on history.
`id` is a short string rather than a generated UUID so the one group every
instance has can be called `"default"` in code and in the database alike, and
every row written before groups existed can be backfilled to it without a
lookup. See services/data_groups.py for how a group is resolved.
"""
from __future__ import annotations
from sqlalchemy import ForeignKey, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, new_id
# The group every instance has, every connection is in until somebody says
# otherwise, and every row written before 1.10.0 is backfilled to.
DEFAULT_GROUP = "default"
class DataGroup(Timestamps, Base):
"""One partition of the people's data, and the models that may read it.
`owner_id` NULL is an instance group, set up by an administrator and usable
by everybody. Set, it is somebody's personal group -- made by a person
holding `data.manage` to keep one provider away from the rest of their own
data, and invisible to everybody else.
The four model columns name the services that read a group's data without
being a chat's model: the embedder that indexes it and the model that
reviews generated images. Empty means the instance's own choice, which is
what every group starts with. They are text ids with a connection beside
them, never a `Model` primary key, for the reason `Chat.model_id` gives: a
"Test & refresh" recreates the row.
"""
__tablename__ = "data_groups"
id: Mapped[str] = mapped_column(String(32), primary_key=True, default=new_id)
name: Mapped[str] = mapped_column(String(120), nullable=False)
description: Mapped[str] = mapped_column(Text, default="")
position: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
owner_id: Mapped[str | None] = mapped_column(
String(32), ForeignKey("users.id", ondelete="CASCADE"), nullable=True, index=True
)
embedding_model_id: Mapped[str] = mapped_column(String(300), default="")
embedding_connection_id: Mapped[str] = mapped_column(String(32), default="")
review_model_id: Mapped[str] = mapped_column(String(300), default="")
review_connection_id: Mapped[str] = mapped_column(String(32), default="")
@property
def is_default(self) -> bool:
return self.id == DEFAULT_GROUP
@property
def personal(self) -> bool:
return self.owner_id is not None
def __repr__(self) -> str:
return f"<DataGroup {self.id} {self.name!r}>"
class InDataGroup:
"""Mixin: the data group a row belongs to.
Nullable, and NULL reads as the default group everywhere -- which is what a
row written before 1.10.0 holds until `data_groups.sweep_unassigned` reaches
it at startup. A plain string rather than a foreign key: `sync_schema` adds
a column with its type only, so a `REFERENCES` clause would exist on a fresh
database and not on an upgraded one, and the two would then disagree about
what deleting a group does.
"""
data_group_id: Mapped[str | None] = mapped_column(String(32), nullable=True)
+5 -4
View File
@@ -38,6 +38,7 @@ from sqlalchemy import (
from sqlalchemy.orm import Mapped, mapped_column, relationship
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.models.data_group import InDataGroup
# Who wrote a record. Not decoration: a skill the model wrote itself is the one
# worth looking at twice when its behaviour changes unexpectedly.
@@ -81,7 +82,7 @@ chat_knowledge_bases = Table(
)
class KnowledgeBase(UUIDPrimaryKey, Timestamps, Base):
class KnowledgeBase(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""A named collection of documents.
Sharing lives here rather than on the individual document: "this folder is
@@ -167,7 +168,7 @@ class Document(UUIDPrimaryKey, Timestamps, Base):
return f"<Document {self.title!r}>"
class Note(UUIDPrimaryKey, Timestamps, Base):
class Note(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""Something the model wrote down, or a person did.
Longer and more specific than a memory. Not injected: a handful of notes
@@ -188,7 +189,7 @@ class Note(UUIDPrimaryKey, Timestamps, Base):
return f"<Note {self.title!r}>"
class Memory(UUIDPrimaryKey, Timestamps, Base):
class Memory(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""One short fact, in front of the model on every turn.
Deliberately not shareable and deliberately small. The length cap is
@@ -208,7 +209,7 @@ class Memory(UUIDPrimaryKey, Timestamps, Base):
return f"<Memory {self.content[:40]!r}>"
class Skill(UUIDPrimaryKey, Timestamps, Base):
class Skill(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""A named set of instructions the model can choose to follow.
`description` is the load-bearing field: it is what gets injected, and it is
+2 -1
View File
@@ -6,6 +6,7 @@ from sqlalchemy import Boolean, ForeignKey, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.models.data_group import InDataGroup
# Where a report came from. Not a foreign key to anything -- see `source_id`.
SOURCE_SCHEDULE = "schedule"
@@ -14,7 +15,7 @@ SOURCE_MANUAL = "manual"
SOURCES = (SOURCE_SCHEDULE, SOURCE_CHAT, SOURCE_MANUAL)
class Report(UUIDPrimaryKey, Timestamps, Base):
class Report(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""A finished piece of work, filed.
Deliberately not a `Chat` with one `Message` in it. A report is read top to
+2 -1
View File
@@ -8,6 +8,7 @@ from sqlalchemy import Boolean, DateTime, ForeignKey, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from lembas.db.base import Base, Timestamps, UUIDPrimaryKey
from lembas.db.models.data_group import InDataGroup
from lembas.db.types import JSONDict
# Where a firing's result is delivered. Chosen per schedule rather than fixed by
@@ -26,7 +27,7 @@ ORIGIN_MODEL = "model"
ORIGINS = (ORIGIN_USER, ORIGIN_MODEL)
class Schedule(UUIDPrimaryKey, Timestamps, Base):
class Schedule(UUIDPrimaryKey, Timestamps, InDataGroup, Base):
"""One standing instruction and when it comes due.
The row carries no recurrence logic at all: `rule_json` is read by
+6
View File
@@ -18,6 +18,7 @@ from lembas.api import (
admin_audio,
admin_branding,
admin_crowd,
admin_data_groups,
admin_extraction,
admin_images,
admin_models,
@@ -84,6 +85,7 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
try:
from lembas.db.session import session_scope
from lembas.services.chat import sweep_temporary
from lembas.services.data_groups import sweep_unassigned
from lembas.services.files import sweep_orphans
from lembas.services.library.documents import sweep_unfiled
from lembas.services.library.indexing import sweep_orphans as sweep_chunks
@@ -94,6 +96,9 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
# Documents that predate knowledge bases have nowhere to live until
# this runs; see services/library/documents.py.
sweep_unfiled(db)
# Rows written before data groups existed, into the group they were
# read in -- see services/data_groups.py.
sweep_unassigned(db)
# Temporary chats older than a day. Startup only, like the sweeps
# above it -- see services/chat.py:sweep_temporary.
sweep_temporary(db)
@@ -225,6 +230,7 @@ def create_app() -> FastAPI:
app.include_router(admin_tools.router)
app.include_router(admin_agents.router)
app.include_router(admin_crowd.router)
app.include_router(admin_data_groups.router)
app.include_router(push.router)
app.include_router(branding.router)
+14
View File
@@ -328,6 +328,20 @@ PERMISSION_DEFS: tuple[PermissionDef, ...] = (
True,
"Library",
),
# Data groups keep one provider's models away from the data another's have
# been handed. The administrator's arrangement applies to everybody; this
# lets a person make groups of their own, put a connection into one for
# themselves, and move their own records between groups. Off by default,
# because it moves what a provider can read -- and an instance that never
# looks should keep the arrangement its administrator made.
PermissionDef(
"data.manage",
"Manage their own data groups",
"Make personal data groups, choose which of them each connection reads "
"for this person, and move their own records between groups.",
False,
"Library",
),
)
# Gates whose read and write halves are separate permissions. Keyed on the gate,
+41 -7
View File
@@ -20,6 +20,7 @@ from lembas.db.models import (
Connection,
Message,
Model,
User,
)
from lembas.services import files as files_service
from lembas.services.llm.openai_client import Endpoint, LLMError, complete
@@ -115,8 +116,14 @@ def resolve_endpoint(
if connection is None or not connection.enabled:
# The original connection is gone or disabled. Any enabled connection
# still offering this model id will do.
model = db.scalar(
# still offering this model id will do -- **in the chat's own data
# group**, for the chat's own model. Any connection at all would repoint
# the conversation onto whichever provider happened to serve the same
# id, and hand it the whole history on the way.
from lembas.services import data_groups
candidates = list(
db.scalars(
select(Model)
.join(Connection)
.where(
@@ -126,6 +133,17 @@ def resolve_endpoint(
)
.order_by(Connection.position)
)
)
if speaks_for_chat and candidates:
owner = db.get(User, chat.user_id) if chat.user_id else None
groups = data_groups.connection_groups(db, owner)
wanted = data_groups.for_chat(db, chat)
candidates = [
m
for m in candidates
if groups.get(m.connection_id, data_groups.DEFAULT_GROUP) == wanted
]
model = candidates[0] if candidates else None
if model is None:
raise LLMError(
f"No enabled connection currently offers the model "
@@ -841,16 +859,28 @@ def default_model(db: DBSession, user=None) -> tuple[str, str] | None:
return chosen.model_id, chosen.connection_id
def available_models(db: DBSession, user=None) -> list[Model]:
def available_models(db: DBSession, user=None, group: str | None = None) -> list[Model]:
"""Models this user may start a chat with, in the administrator's order.
Pinning does NOT hoist a model up this list: pinned models get their own
shortcuts in the sidebar, and a picker whose order silently differs from
the one configured in the admin screen is just confusing.
`group` narrows to the models whose connection is in one data group for this
person -- what a chat that already exists may switch to, and who may be
asked or added to it. `None` is the new-chat screen, where any model can
start a chat and the chat then takes that model's group.
"""
from lembas.security import permissions
reachable = permissions.models_visible_to(db, user)
if group is not None:
from lembas.services import data_groups
groups = data_groups.connection_groups(db, user)
reachable = [
m for m in reachable if groups.get(m.connection_id, data_groups.DEFAULT_GROUP) == group
]
return sorted(reachable, key=lambda m: (m.position, m.model_id))
@@ -865,7 +895,9 @@ MAX_ROSTER_CHARS = 2400
MAX_ROSTER_ENTRY = 300
def roster_models(db: DBSession, user=None, *, exclude: str = "") -> list[Model]:
def roster_models(
db: DBSession, user=None, *, exclude: str = "", group: str | None = None
) -> list[Model]:
"""The other models this person could reach, in the administrator's order.
`exclude` is a `model_id` and is normally the chat's own: a model does not
@@ -874,10 +906,12 @@ def roster_models(db: DBSession, user=None, *, exclude: str = "") -> list[Model]
would be both a leak and a dead end, since asking it anything is refused by
the same check.
"""
return [model for model in available_models(db, user) if model.model_id != exclude]
return [model for model in available_models(db, user, group) if model.model_id != exclude]
def roster_block(db: DBSession, user=None, *, exclude: str = "") -> str:
def roster_block(
db: DBSession, user=None, *, exclude: str = "", group: str | None = None
) -> str:
"""The roster as the models read it: one line each, name, id, what it is for.
The id is in brackets because it is what has to be typed back into
@@ -888,7 +922,7 @@ def roster_block(db: DBSession, user=None, *, exclude: str = "") -> str:
"""
lines: list[str] = []
budget = MAX_ROSTER_CHARS
for model in roster_models(db, user, exclude=exclude)[:MAX_ROSTER_MODELS]:
for model in roster_models(db, user, exclude=exclude, group=group)[:MAX_ROSTER_MODELS]:
parts = ((model.description or "").strip(), (model.notes or "").strip())
about = " ".join(part for part in parts if part)
about = " ".join(about.split())[:MAX_ROSTER_ENTRY]
+13 -2
View File
@@ -272,11 +272,18 @@ def member_speakers(db: DBSession, chat: Chat, user=None) -> list:
Deduplicated against the main model: adding the chat's own model to the crowd
would have it answer twice in a row, which is not what anybody meant by it.
And narrowed to the chat's data group: a member from another group would be
handed this conversation, which is exactly what groups exist to prevent.
"""
from lembas.services import chat as chat_service
from lembas.services import data_groups
reachable = {
model.model_id: model for model in chat_service.roster_models(db, user, exclude="")
model.model_id: model
for model in chat_service.roster_models(
db, user, exclude="", group=data_groups.for_chat(db, chat)
)
}
speakers = [chat_service.Speaker(chat.model_id, chat.connection_id)]
seen = {chat.model_id}
@@ -291,9 +298,13 @@ def member_speakers(db: DBSession, chat: Chat, user=None) -> list:
def unreachable_members(db: DBSession, chat: Chat, user=None) -> list[str]:
"""Members that will be skipped, so a screen can say so rather than lie."""
from lembas.services import chat as chat_service
from lembas.services import data_groups
reachable = {
model.model_id for model in chat_service.roster_models(db, user, exclude="")
model.model_id
for model in chat_service.roster_models(
db, user, exclude="", group=data_groups.for_chat(db, chat)
)
}
return [
member.model_id
+430
View File
@@ -0,0 +1,430 @@
"""Which data group a connection, a chat or a speaker is in.
A data group is the unit of isolation between providers: a model reads the
memories, notes, skills, knowledge, reports, personality and impression of
exactly one group -- the one its connection resolves to -- and a chat belongs
to the group it was started in. See db/models/data_group.py for what the group
itself is.
**The resolution order lives here and nowhere else.** For a connection:
1. the person's own mapping, in `settings_json["data_groups"]`, honoured only
while they hold `data.manage` and only to a group they may use -- so taking
the permission away puts them back on the instance's arrangement without
anybody having to find and clear what they set;
2. the administrator's, `Connection.data_group_id`;
3. the default group.
A mapping to a group that has since been deleted falls through to the next rung
rather than to nothing, for the same reason.
**A chat's group is stamped, not derived.** `for_chat` reads the row, and only
derives -- and stamps -- when the row predates the column. A chat whose model
has since been moved into another group therefore stays where it was, and
`refusal` is what stops that model being handed the chat's history. Deriving it
afresh every turn would instead carry the transcript into whichever group the
model happened to be in today.
"""
from __future__ import annotations
import logging
from typing import TYPE_CHECKING, Any
from sqlalchemy import func, or_, select, update
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import (
DEFAULT_GROUP,
Chat,
Connection,
DataGroup,
KnowledgeBase,
Memory,
Model,
Note,
Report,
Schedule,
Skill,
User,
)
if TYPE_CHECKING:
from lembas.services.chat import Speaker
log = logging.getLogger(__name__)
DEFAULT_NAME = "Default"
# Where a person's own connection-to-group choices live in `settings_json`.
SETTING_KEY = "data_groups"
# The permission that lets somebody make personal groups, move a connection
# into one for themselves, and move their own records between groups.
PERMISSION = "data.manage"
# Every table carrying `data_group_id` whose NULL means "written before groups
# existed" and therefore belongs in the default group. Chats are not on this
# list: a chat's group is derived from its model, see `for_chat`.
LIBRARY_TABLES: tuple[Any, ...] = (Memory, Note, Skill, KnowledgeBase, Report, Schedule)
# Rows of these, counted per group, on the admin and settings pages.
COUNTED: tuple[tuple[Any, str], ...] = (
(Chat, "chats"),
(Memory, "memories"),
(Note, "notes"),
(Skill, "skills"),
(KnowledgeBase, "knowledge bases"),
(Report, "reports"),
)
def group_of(row: Any) -> str:
"""The group a stored row belongs to. NULL is the default group."""
return getattr(row, "data_group_id", None) or DEFAULT_GROUP
def condition(model: Any, group: str):
"""A WHERE clause selecting the rows of `model` in `group`.
NULL counts as the default, so a row the startup sweep has not reached yet
is never lost from the group it belongs to.
"""
column = model.data_group_id
if group == DEFAULT_GROUP:
return or_(column.is_(None), column == DEFAULT_GROUP)
return column == group
# --- The groups themselves -----------------------------------------------------
def ensure_default(db: DBSession) -> DataGroup:
"""The default group, created the first time anything asks for it."""
group = db.get(DataGroup, DEFAULT_GROUP)
if group is None:
group = DataGroup(id=DEFAULT_GROUP, name=DEFAULT_NAME, position=0)
db.add(group)
db.commit()
return group
def get(db: DBSession, group_id: str | None) -> DataGroup | None:
if not group_id:
return None
if group_id == DEFAULT_GROUP:
return ensure_default(db)
return db.get(DataGroup, group_id)
def all_groups(db: DBSession) -> list[DataGroup]:
"""Every group on the instance, personal ones included. Administrators only."""
ensure_default(db)
return list(
db.scalars(
select(DataGroup).order_by(
DataGroup.owner_id.is_not(None), DataGroup.position, DataGroup.name
)
)
)
def instance_groups(db: DBSession) -> list[DataGroup]:
ensure_default(db)
return list(
db.scalars(
select(DataGroup)
.where(DataGroup.owner_id.is_(None))
.order_by(DataGroup.position, DataGroup.name)
)
)
def usable(db: DBSession, user: User | None) -> list[DataGroup]:
"""The groups this person's data may be in: the instance's, and their own."""
groups = instance_groups(db)
if user is not None:
groups += list(
db.scalars(
select(DataGroup).where(DataGroup.owner_id == user.id).order_by(DataGroup.name)
)
)
return groups
def may_use(db: DBSession, user: User | None, group_id: str) -> bool:
group = get(db, group_id)
if group is None:
return False
return group.owner_id is None or (user is not None and group.owner_id == user.id)
def several(db: DBSession, user: User | None) -> bool:
"""Whether there is any choice to show. One group means no chip, no select."""
return len(usable(db, user)) > 1
def name_of(db: DBSession, group_id: str | None) -> str:
group = get(db, group_id or DEFAULT_GROUP)
return group.name if group is not None else (group_id or DEFAULT_NAME)
def may_manage(db: DBSession, user: User | None) -> bool:
from lembas.security import permissions
return user is not None and permissions.has(db, user, PERMISSION)
# --- Connections ---------------------------------------------------------------
def personal_map(user: User | None) -> dict[str, str]:
"""The person's own connection -> group choices, as stored."""
if user is None:
return {}
stored = (user.settings_json or {}).get(SETTING_KEY) or {}
if not isinstance(stored, dict):
return {}
return {str(k): str(v) for k, v in stored.items() if k and v}
def connection_groups(db: DBSession, user: User | None) -> dict[str, str]:
"""Every connection's group for this person, resolved once.
One query for the lot, because the model lists call this for every model
they show and a query per model would be one per row of every picker.
"""
ensure_default(db)
known = {group.id for group in db.scalars(select(DataGroup))}
resolved: dict[str, str] = {}
for connection_id, group_id in db.execute(select(Connection.id, Connection.data_group_id)):
resolved[connection_id] = group_id if group_id in known else DEFAULT_GROUP
if user is not None and may_manage(db, user):
for connection_id, group_id in personal_map(user).items():
if connection_id in resolved and group_id in known and may_use(db, user, group_id):
resolved[connection_id] = group_id
return resolved
def for_connection(db: DBSession, user: User | None, connection_id: str | None) -> str:
if not connection_id:
return DEFAULT_GROUP
return connection_groups(db, user).get(connection_id, DEFAULT_GROUP)
def for_model(db: DBSession, user: User | None, model: Model) -> str:
return for_connection(db, user, model.connection_id)
def _connection_for(db: DBSession, model_id: str, connection_id: str | None) -> str | None:
"""The connection a (model id, connection) pair actually lands on.
The connection when one is named and still enabled; otherwise the first
enabled connection offering that id, which is exactly the one
`chat.resolve_endpoint` would fall back to.
"""
if connection_id:
connection = db.get(Connection, connection_id)
if connection is not None and connection.enabled:
return connection.id
return db.scalar(
select(Model.connection_id)
.join(Connection)
.where(
Model.model_id == model_id,
Model.enabled.is_(True),
Connection.enabled.is_(True),
)
.order_by(Connection.position)
)
def for_pair(
db: DBSession, user: User | None, model_id: str, connection_id: str | None = None
) -> str:
"""The group of a model named by its text id and, if known, its connection."""
return for_connection(db, user, _connection_for(db, model_id, connection_id))
# --- Chats and speakers --------------------------------------------------------
def for_chat(db: DBSession, chat: Chat | None) -> str:
"""The group a chat belongs to.
Read from the row. A row with none -- written before groups, or by a path
that creates a chat without going through `_new_chat` -- is given the group
of its model now, and keeps it.
"""
if chat is None:
return DEFAULT_GROUP
if chat.data_group_id:
return chat.data_group_id
owner = db.get(User, chat.user_id) if chat.user_id else None
group = DEFAULT_GROUP
if chat.model_id:
group = for_pair(db, owner, chat.model_id, chat.connection_id)
chat.data_group_id = group
return group
def for_speaker(
db: DBSession, user: User | None, chat: Chat | None, speaker: Speaker | None
) -> str:
"""The group whose data this speaker is handed.
The chat's own group for the chat's own model. For anybody else -- a crowd
member, a schedule's model -- the group of *their* connection: a model reads
its own group's stores and never the chat's, which is what keeps a crowd
member from another provider out of this group's memories even when a rule
has let it into the conversation.
"""
if chat is not None and (speaker is None or speaker.model_id == chat.model_id):
return for_chat(db, chat)
if speaker is None or not speaker.model_id:
return DEFAULT_GROUP
return for_pair(db, user, speaker.model_id, speaker.connection_id)
def for_composer(db: DBSession, user: User | None, chat_id: str, model_id: str = "") -> str:
"""The group a composer is writing into: its chat's, or its chosen model's.
What the `@` menu and the library picker filter on. A copy made from the
library becomes part of the conversation and is sent to the chat's model,
so offering another group's note there would be the boundary crossed by
hand. On the new-chat screen there is no chat yet, and the model the
composer has chosen decides -- it is the one the chat will be pinned to.
"""
chat = db.get(Chat, chat_id) if chat_id else None
if chat is not None and user is not None and chat.user_id == user.id:
return for_chat(db, chat)
if model_id:
return for_pair(db, user, model_id)
from lembas.services import chat as chat_service
chosen = chat_service.default_model(db, user)
return for_pair(db, user, *chosen) if chosen else DEFAULT_GROUP
def refusal(db: DBSession, user: User | None, chat: Chat, speaker: Speaker) -> str:
"""Why this speaker may not be sent this chat, or "" when it may.
Only the chat's own model is checked here: a crowd member reads its own
group's stores whatever the chat's group is, and whether it may join the
conversation at all is decided where the crowd is assembled.
Refused when the model's connection is now in a different group from the
chat -- an administrator moved it, or the person remapped it. Sending the
reply anyway would hand that provider the chat's whole history.
"""
if speaker.model_id != chat.model_id:
return ""
chat_group = for_chat(db, chat)
model_group = for_pair(db, user, speaker.model_id, speaker.connection_id)
if model_group == chat_group:
return ""
return (
f"This chat belongs to the data group {name_of(db, chat_group)!r}, and its "
f"model is now in {name_of(db, model_group)!r}, so it cannot be sent this "
f"chat's history. Pick a model in {name_of(db, chat_group)!r}, or start a "
f"new chat."
)
# --- Housekeeping ----------------------------------------------------------------
def sweep_unassigned(db: DBSession) -> int:
"""Put every row written before 1.10.0 into the group it belongs to.
Library rows go to the default group: before groups existed every
connection was in it, so that is where every one of them was read. Chats
get their model's group, which on an upgrade is the default too, and on a
later run is the right answer for a chat some path created without
stamping one. Runs at startup beside `documents.sweep_unfiled`.
"""
ensure_default(db)
moved = 0
for model in LIBRARY_TABLES:
result = db.execute(
update(model).where(model.data_group_id.is_(None)).values(data_group_id=DEFAULT_GROUP)
)
moved += result.rowcount or 0
for chat in db.scalars(select(Chat).where(Chat.data_group_id.is_(None))):
for_chat(db, chat)
moved += 1
db.commit()
if moved:
log.info("data groups: %d rows assigned", moved)
return moved
def counts(db: DBSession, group_id: str, *, owner: User | None = None) -> dict[str, int]:
"""How many of each kind of record are in a group, for one person or all."""
found: dict[str, int] = {}
for model, label in COUNTED:
query = select(func.count()).select_from(model).where(condition(model, group_id))
if owner is not None:
column = model.user_id if model is Chat else model.owner_id
query = query.where(column == owner.id)
found[label] = db.scalar(query) or 0
return found
def in_use(db: DBSession, group_id: str) -> dict[str, int]:
"""What still points at a group, which is what stops it being deleted."""
found = counts(db, group_id)
found["connections"] = (
db.scalar(
select(func.count())
.select_from(Connection)
.where(Connection.data_group_id == group_id)
)
or 0
)
return {label: n for label, n in found.items() if n}
def delete(db: DBSession, group: DataGroup) -> None:
"""Remove a group that nothing is in. Raises ValueError otherwise.
Refused rather than cascaded. Deleting a group's records along with it is
far too large a thing to do from one button, and moving them into another
group silently would hand them to that group's providers.
"""
if group.is_default:
raise ValueError("The default group cannot be deleted.")
busy = in_use(db, group.id)
if busy:
listed = ", ".join(f"{n} {label}" for label, n in busy.items())
raise ValueError(f"The group still holds {listed}. Move them out first.")
# Nobody may go on naming a group that is gone; their mapping falls through.
for user in db.scalars(select(User)):
mapped = personal_map(user)
if group.id in mapped.values():
kept = {k: v for k, v in mapped.items() if v != group.id}
user.settings_json = {**(user.settings_json or {}), SETTING_KEY: kept}
db.delete(group)
db.commit()
__all__ = [
"DEFAULT_GROUP",
"all_groups",
"condition",
"connection_groups",
"counts",
"delete",
"ensure_default",
"for_chat",
"for_composer",
"for_connection",
"for_model",
"for_pair",
"for_speaker",
"get",
"group_of",
"in_use",
"instance_groups",
"may_manage",
"may_use",
"name_of",
"personal_map",
"refusal",
"several",
"sweep_unassigned",
"usable",
]
+9 -2
View File
@@ -34,7 +34,7 @@ from lembas.services import canvas as canvas_service
from lembas.services import chat as chat_service
from lembas.services import compaction as compaction_service
from lembas.services import crowd as crowd_service
from lembas.services import interaction, settings_store, tokens, tool_labels
from lembas.services import data_groups, interaction, settings_store, tokens, tool_labels
from lembas.services import metrics as metrics_service
from lembas.services import prompts as prompts_service
from lembas.services import push as push_service
@@ -640,8 +640,15 @@ async def _run(generation: Generation) -> None:
# on has to survive that. It is also the only thing that can make the
# bubble's avatar and the model actually asked agree.
speaker = chat_service.speaker_for(db, chat, message)
endpoint, model_id = chat_service.resolve_endpoint(db, chat, speaker)
owner = db.get(User, chat.user_id)
# Before the endpoint is even resolved: a chat whose model has since
# been moved into another data group must not be sent to it, or that
# provider is handed the whole history the group was keeping from it.
moved = data_groups.refusal(db, owner, chat, speaker)
if moved:
generation.error = moved
return
endpoint, model_id = chat_service.resolve_endpoint(db, chat, speaker)
# Before the request is built, not while it streams. Every other
# budget here can only be noticed part way through and so ends with
+29 -6
View File
@@ -186,6 +186,20 @@ def context_variables(
# set one behaves exactly as it always did.
stamp = clock.now_for(user)
# Whose data this request may carry. The data group of the *answering*
# model -- the chat's own for the chat's model, the member's own for a crowd
# member -- so a model is handed exactly one group's memories, skills and
# personality and never the chat's merely for being in it. No chat is the
# admin preview, which has no speaker and shows every group.
group: str | None = None
if chat is not None:
from lembas.services import chat as chat_service
from lembas.services import data_groups
group = data_groups.for_speaker(
db, user, chat, speaker or chat_service.speaker_for(db, chat)
)
values: dict[str, str] = {
"today": stamp.strftime("%A %-d %B %Y"),
"now": stamp.strftime("%A %-d %B %Y, %H:%M (UTC%z)"),
@@ -227,9 +241,11 @@ def context_variables(
"unbounded": "" if settings_store.chat_rounds(db) else "yes",
"memory_limit": str(memories_service.MAX_MEMORY_CHARS),
"tool_names": _tool_names(offered),
"memories": memories_service.block(db, user) if "memory" in families else "",
"memories": memories_service.block(db, user, group) if "memory" in families else "",
"skills": (
skills_service.index_block(db, user, exclude=tools_service.scoped_skills_off(chat))
skills_service.index_block(
db, user, exclude=tools_service.scoped_skills_off(chat), group=group
)
if "skills" in families
else ""
),
@@ -304,8 +320,15 @@ def context_variables(
# Naming the bases a chat is scoped to matters: without it the model
# cannot tell "there is nothing about this" from "I am only allowed to
# see the contracts folder", and phrases a miss as the former.
if "knowledge" in families and chat.knowledge_bases:
values["knowledge_bases"] = ", ".join(base.name for base in chat.knowledge_bases)
# Only the attached bases in this speaker's group: a base from another
# group is not searchable here, and naming it would leak its name.
in_group = [
base
for base in chat.knowledge_bases
if (base.data_group_id or data_groups.DEFAULT_GROUP) == group
]
if "knowledge" in families and in_group:
values["knowledge_bases"] = ", ".join(base.name for base in in_group)
values["document_names"] = _document_names(db, chat)
# The one thing a tool description cannot carry, because a description
@@ -338,7 +361,7 @@ def context_variables(
# reason the roster and the tool are one checkbox rather than two.
if "friend" in families:
values["model_roster"] = chat_service.roster_block(
db, user, exclude=speaking.model_id
db, user, exclude=speaking.model_id, group=group
)
if "persona" in families:
@@ -346,7 +369,7 @@ def context_variables(
# administrator's default until the model has written one with them;
# and this model's impression of them, which has no default and never
# could.
key = speaking.model_id
key = personas_service.key_for(speaking.model_id, group)
values["persona"] = personas_service.block(db, key, user)
values["person_view"] = personas_service.view_block(db, key, user)
+17 -1
View File
@@ -330,7 +330,23 @@ def _reviewer(context: ToolContext) -> tuple[Endpoint, str] | None:
try:
with session_scope() as db:
model = None
if wanted:
# The chat's data group may name its own reviewer, because the
# reviewer is sent the picture and the prompt that described it --
# this group's data, going to whichever provider reviews. A group
# that names none uses the instance's choice below.
from lembas.services import data_groups
group = data_groups.get(db, context.data_group)
if group is not None and group.review_model_id:
model = db.scalar(
select(Model)
.where(Model.model_id == group.review_model_id)
.order_by(
Model.connection_id != (group.review_connection_id or ""),
Model.position,
)
)
if model is None and wanted:
# By the model's own id, and by primary key for anything stored
# before that was the rule -- a value written by an older release
# is a primary key and must keep working.
+82 -21
View File
@@ -20,6 +20,7 @@ from sqlalchemy.orm import Session as DBSession
from lembas.config import settings
from lembas.db.models import (
CHUNK_DOCUMENT,
DEFAULT_GROUP,
SOURCE_LINK,
SOURCE_UPLOAD,
Document,
@@ -69,19 +70,27 @@ def stored_path(stored_name: str) -> Path | None:
# --- Bases -------------------------------------------------------------------
def visible_bases(db: DBSession, user: User | None):
return select(KnowledgeBase).where(sharing.visible_to(KnowledgeBase, user))
def visible_bases(db: DBSession, user: User | None, group: str | None = None):
"""Bases this user may see; `group` narrows to one data group, for a model."""
return select(KnowledgeBase).where(sharing.visible_to(KnowledgeBase, user, group))
def get_base(db: DBSession, base_id: str, user: User | None) -> KnowledgeBase | None:
def get_base(
db: DBSession, base_id: str, user: User | None, group: str | None = None
) -> KnowledgeBase | None:
base = db.get(KnowledgeBase, base_id)
if base is None or not sharing.can_read(db, base, user):
if base is None or not sharing.can_read(db, base, user, group):
return None
return base
def create_base(
db: DBSession, *, owner: User, name: str, description: str = ""
db: DBSession,
*,
owner: User,
name: str,
description: str = "",
group: str = DEFAULT_GROUP,
) -> KnowledgeBase:
name = " ".join((name or "").split())[:200] or DEFAULT_BASE_NAME
existing = db.scalar(
@@ -90,28 +99,58 @@ def create_base(
)
)
if existing is not None:
# Unique per person across every data group: the constraint is
# `(owner_id, name)` and cannot be changed without rebuilding the table.
if (existing.data_group_id or DEFAULT_GROUP) != (group or DEFAULT_GROUP):
raise ValueError(
f"You already have a knowledge base called {name!r} in another "
f"data group. Names are unique across your groups."
)
raise ValueError(f"You already have a knowledge base called {name!r}.")
base = KnowledgeBase(owner_id=owner.id, name=name, description=description.strip()[:2000])
base = KnowledgeBase(
owner_id=owner.id,
data_group_id=group or DEFAULT_GROUP,
name=name,
description=description.strip()[:2000],
)
db.add(base)
db.commit()
return base
def default_base(db: DBSession, owner: User) -> KnowledgeBase:
"""The base a document goes into when none was chosen.
def default_base(db: DBSession, owner: User, group: str = DEFAULT_GROUP) -> KnowledgeBase:
"""The base a document goes into when none was chosen, in one data group.
Made on demand rather than at registration, so an account that never uses
the library never grows an empty one.
the library never grows an empty one. Outside the default group it is named
after the group, because a base name is unique per person across every group
and two called "My documents" cannot both exist.
"""
from lembas.services import data_groups
group = group or DEFAULT_GROUP
base = db.scalar(
select(KnowledgeBase)
.where(KnowledgeBase.owner_id == owner.id)
.where(
KnowledgeBase.owner_id == owner.id,
data_groups.condition(KnowledgeBase, group),
)
.order_by(KnowledgeBase.created_at)
)
if base is not None:
return base
base = KnowledgeBase(owner_id=owner.id, name=DEFAULT_BASE_NAME)
name = DEFAULT_BASE_NAME
if group != DEFAULT_GROUP:
name = f"{DEFAULT_BASE_NAME} — {data_groups.name_of(db, group)}"[:190]
taken = set(
db.scalars(select(KnowledgeBase.name).where(KnowledgeBase.owner_id == owner.id))
)
wanted, counter = name, 2
while wanted in taken:
wanted = f"{name} ({counter})"
counter += 1
base = KnowledgeBase(owner_id=owner.id, name=wanted, data_group_id=group)
db.add(base)
db.commit()
return base
@@ -172,10 +211,15 @@ def store_upload(
filename: str,
title: str = "",
base: KnowledgeBase | None = None,
group: str = DEFAULT_GROUP,
) -> Document:
"""Add an uploaded file to the library. Raises files.FileError if unusable."""
"""Add an uploaded file to the library. Raises files.FileError if unusable.
A document is in the data group of its base. `group` only chooses which
default base it lands in when no base was given.
"""
prepared = files_service.prepare(payload, filename)
base = base or default_base(db, owner)
base = base or default_base(db, owner, group)
stored_name = f"{secrets.token_hex(16)}{prepared.extension}"
(library_dir() / stored_name).write_bytes(prepared.payload)
@@ -205,14 +249,19 @@ def store_upload(
def store_page(
db: DBSession, *, owner: User, page: Fetched, base: KnowledgeBase | None = None
db: DBSession,
*,
owner: User,
page: Fetched,
base: KnowledgeBase | None = None,
group: str = DEFAULT_GROUP,
) -> Document:
"""Add a fetched web page to the library.
Saved as text rather than as the original HTML: the point of keeping it is
what it said, and the markup would have to be reduced again on every read.
"""
base = base or default_base(db, owner)
base = base or default_base(db, owner, group)
document = Document(
owner_id=owner.id,
base_id=base.id,
@@ -233,15 +282,22 @@ def store_page(
# --- Reading -----------------------------------------------------------------
def visible(db: DBSession, user: User | None, *, base_ids: list[str] | None = None):
def visible(
db: DBSession,
user: User | None,
*,
base_ids: list[str] | None = None,
group: str | None = None,
):
"""Documents this user may see, optionally narrowed to some bases.
Visibility comes from the base, not the document: a document is readable by
whoever can read the base it lives in. That is the whole reason bases are
shareable and documents are not.
shareable and documents are not -- and it is also why a document has no
data group of its own: it is in its base's.
"""
condition = Document.base_id.in_(
select(KnowledgeBase.id).where(sharing.visible_to(KnowledgeBase, user))
select(KnowledgeBase.id).where(sharing.visible_to(KnowledgeBase, user, group))
)
query = select(Document).where(condition)
if base_ids:
@@ -251,12 +307,14 @@ def visible(db: DBSession, user: User | None, *, base_ids: list[str] | None = No
return query
def get(db: DBSession, document_id: str, user: User | None) -> Document | None:
def get(
db: DBSession, document_id: str, user: User | None, group: str | None = None
) -> Document | None:
document = db.get(Document, document_id)
if document is None:
return None
base = db.get(KnowledgeBase, document.base_id) if document.base_id else None
if base is None or not sharing.can_read(db, base, user):
if base is None or not sharing.can_read(db, base, user, group):
return None
return document
@@ -310,6 +368,7 @@ def search(
limit: int = 10,
base_ids: list[str] | None = None,
vector: list[float] | None = None,
group: str | None = None,
) -> list[Document]:
"""Documents matching `needle` that this user may see, best match first.
@@ -331,7 +390,9 @@ def search(
order = {hit.id: position for position, hit in enumerate(hits)}
rows = list(
db.scalars(
visible(db, user, base_ids=base_ids).where(Document.id.in_(list(order)))
visible(db, user, base_ids=base_ids, group=group).where(
Document.id.in_(list(order))
)
)
)
rows.sort(key=lambda document: order.get(document.id, len(order)))
+53 -7
View File
@@ -55,6 +55,7 @@ from lembas.db.models import (
Chunk,
Connection,
Document,
KnowledgeBase,
Model,
Note,
Report,
@@ -92,19 +93,40 @@ class Embedder:
batch: int = 16
def embedder(db: DBSession) -> Embedder | None:
"""The configured embedding model, or None.
def embedder(db: DBSession, group: str | None = None) -> Embedder | None:
"""The embedding model for one data group, or the instance's, or None.
None is the answer to every "no" -- none chosen, the model row deleted, its
connection disabled -- and every caller reads it the same way: do nothing,
and let the keyword search stand. That is deliberately not an error. An
instance that never configured this is the common case, not a broken one.
**A group may name its own.** The embedder is sent the full text of every
record it indexes, so a group that keeps its data away from a provider has
to be able to keep it away from that provider's embedder too. A group that
names none uses the instance's, which is what every group starts with --
and the Data groups page says so, per group, beside the providers.
Mixing is impossible by construction rather than by care: a `Chunk` carries
the model that made its vector and `retrieval.semantic_ids` skips any other,
so a query embedded by one group's model never meets another's vectors.
"""
values = settings_store.extraction(db)
batch = int(values.get("embed_batch") or 16)
if group:
from lembas.services import data_groups
row = data_groups.get(db, group)
if row is not None and row.embedding_model_id:
return _resolve(db, row.embedding_model_id, row.embedding_connection_id, batch)
wanted = str(values.get("embedding_model_id") or "").strip()
if not wanted:
return None
model = db.scalar(
return _resolve(db, wanted, "", batch)
def _resolve(db: DBSession, wanted: str, connection_id: str, batch: int) -> Embedder | None:
query = (
select(Model)
.join(Connection)
.where(
@@ -112,8 +134,9 @@ def embedder(db: DBSession) -> Embedder | None:
Model.enabled.is_(True),
Connection.enabled.is_(True),
)
.order_by(Connection.position)
.order_by(Connection.id != (connection_id or ""), Connection.position)
)
model = db.scalar(query)
if model is None:
log.info("embedding model %r is configured but not available", wanted)
return None
@@ -123,7 +146,30 @@ def embedder(db: DBSession) -> Embedder | None:
return Embedder(
endpoint=Endpoint.from_connection(connection),
model_id=model.model_id,
batch=int(values.get("embed_batch") or 16),
batch=batch,
)
def group_of_row(db: DBSession, row) -> str:
"""The data group a record is in. A document is in its base's."""
from lembas.services import data_groups
if isinstance(row, Document):
base = db.get(KnowledgeBase, row.base_id) if row.base_id else None
return data_groups.group_of(base) if base is not None else data_groups.DEFAULT_GROUP
return data_groups.group_of(row)
def any_configured(db: DBSession) -> bool:
"""Whether any group -- or the instance -- has an embedder to index with."""
from lembas.services import data_groups
if embedder(db) is not None:
return True
return any(
embedder(db, group.id) is not None
for group in data_groups.all_groups(db)
if group.embedding_model_id
)
@@ -199,7 +245,7 @@ async def index_resource(kind: str, resource_id: str, *, force: bool = False) ->
if row is None:
forget_resource(db, kind, resource_id)
return 0
worker = embedder(db)
worker = embedder(db, group_of_row(db, row))
if worker is None:
return 0
body = text_of(row)
@@ -438,7 +484,7 @@ async def rebuild_all(*, force: bool = True) -> None:
_PROGRESS = Progress(running=True)
try:
with session_scope() as db:
if embedder(db) is None:
if not any_configured(db):
_PROGRESS.error = "No embedding model is configured."
return
work: list[tuple[str, str]] = []
+34 -12
View File
@@ -22,7 +22,7 @@ import logging
from sqlalchemy import func, select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, Memory, User
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, DEFAULT_GROUP, Memory, User
log = logging.getLogger(__name__)
@@ -39,14 +39,19 @@ MAX_TOTAL_CHARS = 4000
MAX_RECORDS = 200
def all_for(db: DBSession, user: User | None) -> list[Memory]:
def all_for(db: DBSession, user: User | None, group: str | None = None) -> list[Memory]:
"""This person's memories; `group` narrows to one data group, for a model.
`None` is the person's own list in their settings, which shows every group.
"""
if user is None:
return []
return list(
db.scalars(
select(Memory).where(Memory.owner_id == user.id).order_by(Memory.created_at)
)
)
query = select(Memory).where(Memory.owner_id == user.id)
if group is not None:
from lembas.services import data_groups
query = query.where(data_groups.condition(Memory, group))
return list(db.scalars(query.order_by(Memory.created_at)))
def get(db: DBSession, memory_id: str, user: User | None) -> Memory | None:
@@ -56,7 +61,14 @@ def get(db: DBSession, memory_id: str, user: User | None) -> Memory | None:
return memory
def add(db: DBSession, *, owner: User, content: str, author: str = AUTHOR_MODEL) -> Memory:
def add(
db: DBSession,
*,
owner: User,
content: str,
author: str = AUTHOR_MODEL,
group: str = DEFAULT_GROUP,
) -> Memory:
"""Record a fact. Raises ValueError when there is no room or nothing to say.
An exact repeat returns the record that already exists rather than making a
@@ -71,15 +83,24 @@ def add(db: DBSession, *, owner: User, content: str, author: str = AUTHOR_MODEL)
if not content:
raise ValueError("A memory cannot be empty.")
content = content[:MAX_MEMORY_CHARS]
group = group or DEFAULT_GROUP
# Both checks are per data group. A repeat of a fact another group already
# holds is a new memory *here* -- returning the other group's row would be
# telling this group's model that it had saved something it cannot see.
from lembas.services import data_groups
in_group = data_groups.condition(Memory, group)
existing = db.scalars(
select(Memory).where(Memory.owner_id == owner.id, Memory.content == content)
select(Memory).where(Memory.owner_id == owner.id, Memory.content == content, in_group)
).first()
if existing is not None:
return existing
count = db.scalar(
select(func.count()).select_from(Memory).where(Memory.owner_id == owner.id)
select(func.count())
.select_from(Memory)
.where(Memory.owner_id == owner.id, in_group)
)
if (count or 0) >= MAX_RECORDS:
# Deliberately does NOT say "remove one first". Past MAX_TOTAL_CHARS the
@@ -95,6 +116,7 @@ def add(db: DBSession, *, owner: User, content: str, author: str = AUTHOR_MODEL)
memory = Memory(
owner_id=owner.id,
data_group_id=group,
content=content,
author=author if author in (AUTHOR_USER, AUTHOR_MODEL) else AUTHOR_MODEL,
)
@@ -117,14 +139,14 @@ def delete(db: DBSession, memory: Memory) -> None:
db.commit()
def block(db: DBSession, user: User | None) -> str:
def block(db: DBSession, user: User | None, group: str | None = None) -> str:
"""The memories as they appear in the prompt, within the budget.
Oldest first, and truncation drops the *newest* -- a fact that has survived
a long time is more likely to be a standing preference than something said
once this morning.
"""
records = all_for(db, user)
records = all_for(db, user, group)
if not records:
return ""
+22 -9
View File
@@ -13,7 +13,7 @@ import logging
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, CHUNK_NOTE, Note, User
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, CHUNK_NOTE, DEFAULT_GROUP, Note, User
from lembas.services import sharing
from lembas.services.library import retrieval
@@ -26,20 +26,25 @@ MAX_BODY_CHARS = 40_000
SNIPPET_CHARS = 800
def visible(db: DBSession, user: User | None):
return select(Note).where(sharing.visible_to(Note, user))
def visible(db: DBSession, user: User | None, group: str | None = None):
"""Notes this user may see; `group` narrows to one data group, for a model."""
return select(Note).where(sharing.visible_to(Note, user, group))
def get(db: DBSession, note_id: str, user: User | None) -> Note | None:
def get(
db: DBSession, note_id: str, user: User | None, group: str | None = None
) -> Note | None:
note = db.get(Note, note_id)
if note is None or not sharing.can_read(db, note, user):
if note is None or not sharing.can_read(db, note, user, group):
return None
return note
def recent(db: DBSession, user: User | None, *, limit: int = 20) -> list[Note]:
def recent(
db: DBSession, user: User | None, *, limit: int = 20, group: str | None = None
) -> list[Note]:
return list(
db.scalars(visible(db, user).order_by(Note.updated_at.desc()).limit(limit))
db.scalars(visible(db, user, group).order_by(Note.updated_at.desc()).limit(limit))
)
@@ -50,6 +55,7 @@ def search(
*,
limit: int = 10,
vector: list[float] | None = None,
group: str | None = None,
) -> list[Note]:
"""Notes matching `needle` that this user may see, best match first.
@@ -62,16 +68,23 @@ def search(
if not hits:
return []
order = {hit.id: position for position, hit in enumerate(hits)}
rows = list(db.scalars(visible(db, user).where(Note.id.in_(list(order)))))
rows = list(db.scalars(visible(db, user, group).where(Note.id.in_(list(order)))))
rows.sort(key=lambda note: order.get(note.id, len(order)))
return rows[:limit]
def create(
db: DBSession, *, owner: User, title: str, body: str, author: str = AUTHOR_USER
db: DBSession,
*,
owner: User,
title: str,
body: str,
author: str = AUTHOR_USER,
group: str = DEFAULT_GROUP,
) -> Note:
note = Note(
owner_id=owner.id,
data_group_id=group or DEFAULT_GROUP,
title=(title.strip() or "Untitled")[:MAX_TITLE_CHARS],
body=body.strip()[:MAX_BODY_CHARS],
author=author if author in (AUTHOR_USER, AUTHOR_MODEL) else AUTHOR_USER,
+33 -6
View File
@@ -67,7 +67,7 @@ def embeddable(db: DBSession) -> bool:
return indexing.enabled(db)
def worker_for(db: DBSession):
def worker_for(db: DBSession, group: str | None = None):
"""The configured embedder, resolved while a session is open.
Split from the awaiting half deliberately. A caller that must not hold a
@@ -78,7 +78,21 @@ def worker_for(db: DBSession):
"""
from lembas.services.library import indexing
return indexing.embedder(db)
return indexing.embedder(db, group)
class QueryVector(list):
"""A query embedding that knows which model made it.
A plain list everywhere it is used as one. The extra attribute is what lets
`semantic_ids` skip chunks another model embedded: two models of the same
width produce vectors that score against each other perfectly happily and
mean nothing, and since 1.10.0 a data group may have its own embedder, so
two such models on one instance is an ordinary arrangement rather than a
rebuild left half done.
"""
model_id: str = ""
async def embed_with(worker, needle: str) -> list[float] | None:
@@ -100,7 +114,11 @@ async def embed_with(worker, needle: str) -> list[float] | None:
except LLMError as exc:
log.info("could not embed a query: %s", exc)
return None
return vectors[0] if vectors else None
if not vectors:
return None
vector = QueryVector(vectors[0])
vector.model_id = worker.model_id
return vector
async def embed_query(db: DBSession, needle: str) -> list[float] | None:
@@ -126,13 +144,22 @@ def semantic_ids(
Chunks whose width does not match the query's are skipped. That is a change
of embedding model with a rebuild still pending, and scoring across two
spaces produces a confident wrong answer rather than a missing one.
**And chunks another model made are skipped when the query says which model
it came from.** Width alone cannot tell two 1024-wide models apart, and with
an embedder per data group two of them on one instance is ordinary. A plain
list -- a caller that built its own vector -- keeps the width check alone.
"""
if not vector:
return []
width = len(vector)
rows = db.execute(
select(Chunk.resource_id, Chunk.vector, Chunk.dims).where(Chunk.resource_type == kind)
).all()
query = select(Chunk.resource_id, Chunk.vector, Chunk.dims).where(
Chunk.resource_type == kind
)
made_by = getattr(vector, "model_id", "")
if made_by:
query = query.where(Chunk.model_id == made_by)
rows = db.execute(query).all()
best: dict[str, float] = {}
for resource_id, blob, dims in rows:
+57 -15
View File
@@ -28,7 +28,15 @@ from collections.abc import Iterable
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import AUTHOR_MODEL, AUTHOR_USER, CHUNK_SKILL, Skill, SkillRevision, User
from lembas.db.models import (
AUTHOR_MODEL,
AUTHOR_USER,
CHUNK_SKILL,
DEFAULT_GROUP,
Skill,
SkillRevision,
User,
)
from lembas.services import sharing
from lembas.services.library import retrieval
@@ -55,18 +63,23 @@ def slugify(name: str) -> str:
return cleaned[:60]
def visible(db: DBSession, user: User | None):
return select(Skill).where(sharing.visible_to(Skill, user))
def visible(db: DBSession, user: User | None, group: str | None = None):
"""Skills this user may see; `group` narrows to one data group, for a model."""
return select(Skill).where(sharing.visible_to(Skill, user, group))
def get(db: DBSession, skill_id: str, user: User | None) -> Skill | None:
def get(
db: DBSession, skill_id: str, user: User | None, group: str | None = None
) -> Skill | None:
skill = db.get(Skill, skill_id)
if skill is None or not sharing.can_read(db, skill, user):
if skill is None or not sharing.can_read(db, skill, user, group):
return None
return skill
def by_name(db: DBSession, name: str, user: User | None) -> Skill | None:
def by_name(
db: DBSession, name: str, user: User | None, group: str | None = None
) -> Skill | None:
"""Look one up the way the model refers to it.
Scoped to what this person can **see**, which is theirs plus anything
@@ -77,7 +90,7 @@ def by_name(db: DBSession, name: str, user: User | None) -> Skill | None:
"""
if user is None:
return None
return db.scalar(visible(db, user).where(Skill.name == slugify(name)))
return db.scalar(visible(db, user, group).where(Skill.name == slugify(name)))
def owned_by_name(db: DBSession, name: str, owner: User) -> Skill | None:
@@ -101,7 +114,11 @@ def owned_by_name(db: DBSession, name: str, owner: User) -> Skill | None:
def enabled_for(
db: DBSession, user: User | None, *, exclude: Iterable[str] = ()
db: DBSession,
user: User | None,
*,
exclude: Iterable[str] = (),
group: str | None = None,
) -> list[Skill]:
"""Skills that should appear in the index, oldest first for a stable order.
@@ -112,7 +129,7 @@ def enabled_for(
return []
hidden = {slugify(name) for name in exclude}
rows = db.scalars(
visible(db, user)
visible(db, user, group)
.where(Skill.enabled.is_(True))
.order_by(Skill.name)
.limit(MAX_INDEX_SKILLS + len(hidden))
@@ -120,14 +137,20 @@ def enabled_for(
return [skill for skill in rows if skill.name not in hidden][:MAX_INDEX_SKILLS]
def count_enabled(db: DBSession, user: User | None, *, exclude: Iterable[str] = ()) -> int:
def count_enabled(
db: DBSession,
user: User | None,
*,
exclude: Iterable[str] = (),
group: str | None = None,
) -> int:
"""How many skills are available here at all.
Zero is what withdraws `skill_get` and `skill_edit`: reading and improving
are meaningless with nothing to read, and a model told to "read one with
skill_get" above a list that is not there spends a round finding out.
"""
return len(enabled_for(db, user, exclude=exclude))
return len(enabled_for(db, user, exclude=exclude, group=group))
def search(
@@ -137,6 +160,7 @@ def search(
*,
limit: int = 10,
vector: list[float] | None = None,
group: str | None = None,
) -> list[Skill]:
"""Skills matching `needle` that this user may see, best match first.
@@ -149,7 +173,7 @@ def search(
if not hits:
return []
order = {hit.id: position for position, hit in enumerate(hits)}
rows = list(db.scalars(visible(db, user).where(Skill.id.in_(list(order)))))
rows = list(db.scalars(visible(db, user, group).where(Skill.id.in_(list(order)))))
rows.sort(key=lambda skill: order.get(skill.id, len(order)))
return rows[:limit]
@@ -175,6 +199,7 @@ def create(
description: str,
body: str,
author: str = AUTHOR_USER,
group: str = DEFAULT_GROUP,
) -> Skill:
slug = slugify(name)
if not SKILL_NAME_PATTERN.match(slug):
@@ -182,7 +207,17 @@ def create(
"A skill name must be two or more letters, numbers or hyphens, "
"such as 'weekly-report'."
)
if owned_by_name(db, slug, owner) is not None:
taken = owned_by_name(db, slug, owner)
if taken is not None:
# A name is unique per person across every data group -- the table's
# constraint is `(owner_id, name)` and cannot be changed. "Edit it
# instead" would send a model in another group to a skill it cannot
# see, so that case gets its own sentence.
if (taken.data_group_id or DEFAULT_GROUP) != (group or DEFAULT_GROUP):
raise SkillError(
f"The name {slug!r} is already used by a skill in another data "
f"group. Choose a different name."
)
raise SkillError(f"A skill called {slug!r} already exists. Edit it instead.")
if not description.strip():
raise SkillError(
@@ -192,6 +227,7 @@ def create(
skill = Skill(
owner_id=owner.id,
data_group_id=group or DEFAULT_GROUP,
name=slug,
description=description.strip()[:MAX_DESCRIPTION_CHARS],
body=body.strip()[:MAX_BODY_CHARS],
@@ -257,9 +293,15 @@ def delete(db: DBSession, skill: Skill) -> None:
db.commit()
def index_block(db: DBSession, user: User | None, *, exclude: Iterable[str] = ()) -> str:
def index_block(
db: DBSession,
user: User | None,
*,
exclude: Iterable[str] = (),
group: str | None = None,
) -> str:
"""The one-line-per-skill listing that goes into the prompt."""
skills = enabled_for(db, user, exclude=exclude)
skills = enabled_for(db, user, exclude=exclude, group=group)
if not skills:
return ""
return "\n".join(f"- {skill.name}: {skill.description}" for skill in skills)
+37 -1
View File
@@ -31,6 +31,7 @@ from sqlalchemy.orm import Session as DBSession
from lembas.db.models import (
AUTHOR_MODEL,
AUTHOR_USER,
DEFAULT_GROUP,
Impression,
Persona,
PersonaRevision,
@@ -54,8 +55,39 @@ MAX_VIEW_CHARS = 800
# a model editing itself every turn cannot grow the table without limit.
MAX_REVISIONS = 20
# Between a model id and a data group in a person's key. Two characters, because
# one `@` is a character a model id could plausibly contain and this must never
# split one.
KEY_SEPARATOR = "@@"
def key_for(model_id: str, group: str | None) -> str:
"""The key a person's personality and impression are stored under.
**Namespaced by data group, and the reason is a constraint.** Both tables
are `UNIQUE(model_key, owner_id)`, SQLite cannot alter a constraint, and
this project's schema changes are additive only -- so a `data_group_id`
column could not let one person hold a personality for the same model id in
two groups, which is exactly what one model id served by two providers in
different groups needs.
The default group keeps the bare model id, which is what every row written
before groups existed already holds, so an upgrade moves nothing. The
administrator's default (`owner_id NULL`) is always bare: it is their text,
not a person's data, and every group falls back to it.
"""
if not model_id or not group or group == DEFAULT_GROUP:
return model_id
return f"{model_id}{KEY_SEPARATOR}{group}"
def split_key(model_key: str) -> tuple[str, str]:
"""(model id, data group) out of a stored key."""
model_id, separator, group = (model_key or "").rpartition(KEY_SEPARATOR)
if not separator:
return model_key or "", DEFAULT_GROUP
return model_id, group or DEFAULT_GROUP
def get(db: DBSession, model_key: str, owner: User | None) -> Persona | None:
"""One personality row, exactly as asked for and with no fallback.
@@ -86,7 +118,8 @@ def effective(db: DBSession, model_key: str, owner: User | None) -> Persona | No
own = get(db, model_key, owner)
if own is not None:
return own
return get(db, model_key, None) if owner is not None else None
# The default is keyed on the bare model id whatever group asked.
return get(db, split_key(model_key)[0], None) if owner is not None else None
def personas_of(db: DBSession, owner: User | None) -> list[Persona]:
@@ -302,6 +335,7 @@ def view_block(db: DBSession, model_key: str, owner: User | None) -> str:
__all__ = [
"KEY_SEPARATOR",
"MAX_PERSONA_CHARS",
"MAX_REVISIONS",
"MAX_VIEW_CHARS",
@@ -312,8 +346,10 @@ __all__ = [
"get",
"impression",
"impressions_for",
"key_for",
"personas_for",
"personas_of",
"split_key",
"view_block",
"revert",
"write",
+19 -8
View File
@@ -19,7 +19,7 @@ import logging
from sqlalchemy import func, select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import CHUNK_REPORT, SOURCE_MANUAL, SOURCES, Report, User
from lembas.db.models import CHUNK_REPORT, DEFAULT_GROUP, SOURCE_MANUAL, SOURCES, Report, User
from lembas.services import sharing
from lembas.services.library import retrieval
@@ -33,7 +33,7 @@ MAX_BODY_CHARS = 60_000
SNIPPET_CHARS = 400
def visible(user: User | None):
def visible(user: User | None, group: str | None = None):
"""Every report this person owns or has been shared.
Takes no session because it builds a query rather than running one, and
@@ -43,13 +43,17 @@ def visible(user: User | None):
It said "a later move to shared reports is a change of one line here", and
it was: `sharing.visible_to` is that line. Every listing, search and detail
page went through this already, which is what made the move safe.
`group` narrows to one data group, and is what a model's tools pass.
"""
return select(Report).where(sharing.visible_to(Report, user))
return select(Report).where(sharing.visible_to(Report, user, group))
def get(db: DBSession, report_id: str, user: User | None) -> Report | None:
def get(
db: DBSession, report_id: str, user: User | None, group: str | None = None
) -> Report | None:
report = db.get(Report, report_id)
if report is None or not sharing.can_read(db, report, user):
if report is None or not sharing.can_read(db, report, user, group):
return None
return report
@@ -67,8 +71,12 @@ def owned(db: DBSession, report_id: str, user: User | None) -> Report | None:
return report
def recent(db: DBSession, user: User | None, *, limit: int = 20) -> list[Report]:
return list(db.scalars(visible(user).order_by(Report.created_at.desc()).limit(limit)))
def recent(
db: DBSession, user: User | None, *, limit: int = 20, group: str | None = None
) -> list[Report]:
return list(
db.scalars(visible(user, group).order_by(Report.created_at.desc()).limit(limit))
)
def search(
@@ -78,6 +86,7 @@ def search(
*,
limit: int = 20,
vector: list[float] | None = None,
group: str | None = None,
) -> list[Report]:
"""Reports matching `needle`, best match first.
@@ -94,7 +103,7 @@ def search(
if not hits:
return []
order = {hit.id: position for position, hit in enumerate(hits)}
rows = list(db.scalars(visible(user).where(Report.id.in_(list(order)))))
rows = list(db.scalars(visible(user, group).where(Report.id.in_(list(order)))))
rows.sort(key=lambda report: order.get(report.id, len(order)))
return rows[:limit]
@@ -165,6 +174,7 @@ def create(
model_id: str = "",
error: str = "",
unread: bool = True,
group: str = DEFAULT_GROUP,
) -> Report:
"""File a report.
@@ -178,6 +188,7 @@ def create(
"""
report = Report(
owner_id=owner.id,
data_group_id=group or DEFAULT_GROUP,
title=(title.strip() or "Untitled report")[:MAX_TITLE_CHARS],
summary=(summary.strip() or _first_line(body))[:MAX_SUMMARY_CHARS],
body=(body or "").strip()[:MAX_BODY_CHARS],
+14 -2
View File
@@ -214,17 +214,29 @@ async def compile_request(
return Compiled(ok=True, title=title, instruction=instruction, target=target, rule=clean)
def endpoint_for(db, user: User) -> tuple[Endpoint, str] | None:
def endpoint_for(db, user: User, model_id: str = "") -> tuple[Endpoint, str] | None:
"""A connection and model to compile with, or None if there is none.
Built on a throwaway `Chat` that is never added to a session, exactly as
`agent/draft.py` does: `resolve_endpoint` reads `model_id` and
`connection_id` and nothing else, so it works unchanged and did not have to
learn what a compile is.
**From the schedule's own data group.** The request being compiled is the
person's words about their own work, and it goes to whichever model does the
compiling -- so that model is chosen among the ones that will run the
schedule, never merely the first one pinned. `model_id` is the model the
form has chosen; without one, the person's default model decides the group.
"""
from lembas.services import chat as chat_service
from lembas.services import data_groups
models = chat_service.available_models(db, user)
if model_id:
group = data_groups.for_pair(db, user, model_id)
else:
chosen = chat_service.default_model(db, user)
group = data_groups.for_pair(db, user, *chosen) if chosen else data_groups.DEFAULT_GROUP
models = chat_service.available_models(db, user, group)
if not models:
return None
chosen = next((m for m in models if m.pinned), models[0])
+14
View File
@@ -30,6 +30,7 @@ import logging
from datetime import UTC, datetime
from lembas.db.models import (
DEFAULT_GROUP,
ROLE_ASSISTANT,
TARGET_CHAT,
TARGET_MESSAGES,
@@ -187,6 +188,7 @@ async def deliver(schedule_id: str, message_id: str, *, since: datetime) -> None
source_id=chat_id,
schedule_id=schedule.id,
error="The run did not produce a reply.",
group=schedule.data_group_id or DEFAULT_GROUP,
)
return
reports_service.create(
@@ -198,6 +200,7 @@ async def deliver(schedule_id: str, message_id: str, *, since: datetime) -> None
source_id=chat_id,
schedule_id=schedule.id,
model_id=message.model_id or "",
group=schedule.data_group_id or DEFAULT_GROUP,
)
return
@@ -213,7 +216,18 @@ async def deliver(schedule_id: str, message_id: str, *, since: datetime) -> None
# not write it, and the bubble should not imply they did.
from lembas.services import chat as chat_service
from lembas.services import messages as messages_service
from lembas.services import schedules as schedules_service
# Checked when the schedule was saved, and again here: the person may
# have moved a connection or remapped a group since, and a turn in
# Messages is read by Messages' model on every later reply.
refused = schedules_service.messages_refusal(
db, owner, schedule.data_group_id or ""
)
if refused:
schedule.last_error = refused
db.commit()
return
conversation = messages_service.for_user(db, owner)
chat_service.create_message(
db,
+60
View File
@@ -14,10 +14,12 @@ from sqlalchemy import func, select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import (
KIND_MESSAGES,
KIND_TASK,
ORIGIN_USER,
ORIGINS,
TARGET_CHAT,
TARGET_MESSAGES,
TARGETS,
Chat,
Schedule,
@@ -55,6 +57,51 @@ def for_chat(db: DBSession, chat: Chat) -> Schedule | None:
return db.scalars(select(Schedule).where(Schedule.chat_id == chat.id)).first()
def group_for(db: DBSession, owner: User, model_id: str) -> str:
"""The data group a schedule runs in: its model's, or the person's default's.
Stamped on the schedule and on its task chat when it is made. A run reads
that group's memories and notes, and what it produces -- a report, a turn
in a chat -- lands in it.
"""
from lembas.services import chat as chat_service
from lembas.services import data_groups
if model_id:
return data_groups.for_pair(db, owner, model_id)
chosen = chat_service.default_model(db, owner)
return data_groups.for_pair(db, owner, *chosen) if chosen else data_groups.DEFAULT_GROUP
def messages_refusal(db: DBSession, owner: User, group: str) -> str:
"""Why a schedule in `group` may not post into Messages, or "".
Messages is one long conversation, pinned to one data group like any chat.
A run from another group posting its answer there would put that group's
output in front of this group's model on the next turn -- data crossing
between providers through the one channel nobody would think to check.
"""
from lembas.services import data_groups
conversation = db.scalars(
select(Chat)
.where(Chat.user_id == owner.id, Chat.kind == KIND_MESSAGES)
.order_by(Chat.created_at)
).first()
if conversation is not None:
home = data_groups.for_chat(db, conversation)
else:
home = group_for(db, owner, "")
if (group or data_groups.DEFAULT_GROUP) == home:
return ""
return (
f"This schedule's model is in the data group "
f"{data_groups.name_of(db, group)!r} and Messages is in "
f"{data_groups.name_of(db, home)!r}, so it cannot post there. File it as a "
f"report, or keep it in its own chat."
)
def count_for(db: DBSession, user: User) -> int:
return int(
db.scalar(
@@ -103,6 +150,13 @@ def create(
# on, so it is refused at the only moment somebody is present to be told.
raise ScheduleError("That schedule has no next run — its time has already passed.")
group = group_for(db, owner, model_id)
target = target if target in TARGETS else TARGET_CHAT
if target == TARGET_MESSAGES:
refused = messages_refusal(db, owner, group)
if refused:
raise ScheduleError(refused)
limit = int(settings_store.schedules(db).get("max_per_user") or 20)
if count_for(db, owner) >= limit:
raise ScheduleError(
@@ -115,6 +169,7 @@ def create(
kind=KIND_TASK,
title=(title.strip() or "Scheduled task")[:MAX_TITLE_CHARS],
model_id=model_id or "",
data_group_id=group,
# Said on the row as well as implied by the kind. `tools.unattended`
# reads both, because the column was added to a table that already held
# task chats and a backfill cannot know which they were -- but every one
@@ -134,6 +189,7 @@ def create(
target=target if target in TARGETS else TARGET_CHAT,
chat_id=chat.id,
model_id=model_id or "",
data_group_id=group,
origin=origin if origin in ORIGINS else ORIGIN_USER,
enabled=True,
next_fire_at=rule_service.next_after(
@@ -162,6 +218,10 @@ def update(
if instruction is not None:
schedule.instruction = instruction.strip()[:MAX_INSTRUCTION_CHARS]
if target is not None and target in TARGETS:
if target == TARGET_MESSAGES and schedule.target != TARGET_MESSAGES:
refused = messages_refusal(db, owner, schedule.data_group_id or "")
if refused:
raise ScheduleError(refused)
schedule.target = target
if rule is not None:
clean = rule_service.validate(rule)
+28 -3
View File
@@ -74,12 +74,19 @@ def principal_ids(user: User | None) -> tuple[list[str], list[str]]:
return [user.id], [group.id for group in user.groups]
def visible_to(model: Any, user: User | None) -> ColumnElement[bool]:
def visible_to(
model: Any, user: User | None, group: str | None = None
) -> ColumnElement[bool]:
"""A WHERE clause selecting the rows of `model` this user may see.
Returned as a condition rather than a query so callers can add their own
filtering, ordering and pagination without this module knowing about any of
it.
`group` narrows to one data group, and is what every path that hands rows to
a *model* passes -- a model reads only its own group's data. `None` is the
person's own view of their library, which shows every group they have: the
isolation is between providers, not between a person and their records.
"""
if user is None:
# Signed out sees nothing. Not an empty library -- no library.
@@ -93,7 +100,12 @@ def visible_to(model: Any, user: User | None) -> ColumnElement[bool]:
(Share.principal_type == PRINCIPAL_GROUP) & Share.principal_id.in_(groups or [""]),
),
)
return or_(model.owner_id == user.id, model.id.in_(shared))
seen = or_(model.owner_id == user.id, model.id.in_(shared))
if group is None:
return seen
from lembas.services import data_groups
return and_(seen, data_groups.condition(model, group))
def only_shared(model: Any, user: User | None) -> ColumnElement[bool]:
@@ -120,9 +132,22 @@ def owned_by(model: Any, user: User | None) -> ColumnElement[bool]:
return model.owner_id == user.id
def can_read(db: DBSession, resource: Any, user: User | None) -> bool:
def can_read(
db: DBSession, resource: Any, user: User | None, group: str | None = None
) -> bool:
"""Whether this user may read one row -- and, given `group`, whether it is in it.
The group half is what makes fetching a record *by id* obey the same
boundary as searching for it. Without it a model that learned an id from
another group's transcript could read the record straight past the filter.
"""
if user is None or resource is None:
return False
if group is not None:
from lembas.services import data_groups
if data_groups.group_of(resource) != group:
return False
if resource.owner_id == user.id:
return True
users, groups = principal_ids(user)
+27 -6
View File
@@ -77,7 +77,7 @@ from typing import TYPE_CHECKING, Any
from lembas.db.models import KIND_AGENT, KIND_CHAT, Chat, Model, User
from lembas.db.session import session_scope
from lembas.security import permissions
from lembas.services import settings_store
from lembas.services import data_groups, settings_store
from lembas.services.agent import policy as agent_policy
if TYPE_CHECKING: # pragma: no cover - typing only
@@ -224,6 +224,7 @@ def _create_child(
from lembas.services import chat as chat_service
peer = friend is not None
owner = db.get(User, parent.user_id) if parent.user_id else None
child = Chat(
user_id=parent.user_id,
# An ordinary chat for a friend even when the asking one is an agent
@@ -233,6 +234,14 @@ def _create_child(
title=title[:200] or ("Question" if peer else "Helper"),
model_id=friend.model_id if peer else parent.model_id,
connection_id=friend.connection_id if peer else parent.connection_id,
# The helper is in its parent's group, doing its parent's work. A friend
# is in its own model's, so it reads its own group's data and never the
# asker's.
data_group_id=(
data_groups.for_pair(db, owner, friend.model_id, friend.connection_id)
if peer
else data_groups.for_chat(db, parent)
),
# Never in a listing, and swept a day later even if it is kept.
temporary=True,
parent_chat_id=parent.id,
@@ -559,7 +568,9 @@ def _friend_error(message: str, *, question: str = "") -> ToolOutcome:
)
def _resolve_friend(db, owner: User, wanted: str, *, asking: str) -> tuple[Model | None, str]:
def _resolve_friend(
db, owner: User, wanted: str, *, asking: str, group: str | None = None
) -> tuple[Model | None, str]:
"""The model a call named, or a refusal that says what it could have named.
The name arrives in a tool call, which is to say it was written by a model
@@ -570,11 +581,15 @@ def _resolve_friend(db, owner: User, wanted: str, *, asking: str) -> tuple[Model
Matched on `model_id` first and on the label second, because the roster
prints both and a model will sometimes type back the pretty one.
`group` is the asking chat's data group. A friend is handed the question
and whatever context the asker wrote into it, so one in another group would
be carrying this group's data to another provider.
"""
from lembas.services import chat as chat_service
question_for = wanted.strip()
candidates = chat_service.roster_models(db, owner, exclude=asking)
candidates = chat_service.roster_models(db, owner, exclude=asking, group=group)
if not candidates:
return None, (
"There is no other model here to ask. Answer from what you know."
@@ -583,7 +598,7 @@ def _resolve_friend(db, owner: User, wanted: str, *, asking: str) -> tuple[Model
return None, (
"Name the model to ask, exactly as it is written in brackets in the "
"list you were given:\n"
+ chat_service.roster_block(db, owner, exclude=asking)
+ chat_service.roster_block(db, owner, exclude=asking, group=group)
)
lowered = question_for.lower()
@@ -603,7 +618,7 @@ def _resolve_friend(db, owner: User, wanted: str, *, asking: str) -> tuple[Model
return None, (
f"There is no model called {question_for!r} that you can reach. "
"These are the ones you can:\n"
+ chat_service.roster_block(db, owner, exclude=asking)
+ chat_service.roster_block(db, owner, exclude=asking, group=group)
)
@@ -670,7 +685,13 @@ async def _run_ask_friend(context: ToolContext, args: dict[str, Any]) -> ToolOut
if owner is None: # pragma: no cover - a chat outliving its owner
return _friend_error("That account no longer exists.", question=question)
friend, refusal = _resolve_friend(db, owner, wanted, asking=parent.model_id)
friend, refusal = _resolve_friend(
db,
owner,
wanted,
asking=parent.model_id,
group=data_groups.for_chat(db, parent),
)
if friend is None:
return _friend_error(refusal, question=question)
+66 -24
View File
@@ -32,7 +32,7 @@ from typing import Any
from sqlalchemy import select
from sqlalchemy.orm import Session as DBSession
from lembas.db.models import AUTHOR_MODEL, KIND_TASK, SOURCE_CHAT, Chat, User
from lembas.db.models import AUTHOR_MODEL, DEFAULT_GROUP, KIND_TASK, SOURCE_CHAT, Chat, User
from lembas.db.session import session_scope
from lembas.services import personas as personas_service
from lembas.services import prompts as prompts_service
@@ -289,6 +289,12 @@ class ToolContext:
image_checkpoint: str = ""
model_id: str = ""
connection_id: str = ""
# Which data group this call reads and writes, resolved from the answering
# model's connection. Every library runner passes it to its store and every
# write stamps it, so a model reaches exactly one group's records -- by
# search *and* by id, since a model that learned an id from somewhere else
# must not be able to fetch the record past the filter.
data_group: str = DEFAULT_GROUP
@dataclass
@@ -465,12 +471,18 @@ async def _run_knowledge_search(context: ToolContext, args: dict[str, Any]) -> T
# session held across one is the trade `_maybe_compact` already refuses.
# None for every "no" -- no model configured, endpoint down -- and the
# search is then exactly the keyword one it has always been.
vector = await _query_vector(query)
vector = await _query_vector(query, context.data_group)
with session_scope() as db:
user = db.get(User, context.owner_id)
found = documents_service.search(
db, user, query, limit=6, base_ids=context.base_ids, vector=vector
db,
user,
query,
limit=6,
base_ids=context.base_ids,
vector=vector,
group=context.data_group,
)
event = {
"name": "knowledge_search",
@@ -502,7 +514,7 @@ async def _run_knowledge_get(context: ToolContext, args: dict[str, Any]) -> Tool
document_id = str(args.get("id") or "").strip()
with session_scope() as db:
user = db.get(User, context.owner_id)
document = documents_service.get(db, document_id, user)
document = documents_service.get(db, document_id, user, context.data_group)
if document is None:
return ToolOutcome(
"There is no such document, or it is not available to you.",
@@ -526,7 +538,7 @@ async def _run_knowledge_get(context: ToolContext, args: dict[str, Any]) -> Tool
return ToolOutcome(f"{document.title}\n\n{body}", event)
async def _query_vector(query: str) -> list[float] | None:
async def _query_vector(query: str, group: str | None = None) -> list[float] | None:
"""The query as a vector, for the stores that can use one.
Its own session, opened and closed before the caller opens theirs: this is
@@ -538,20 +550,22 @@ async def _query_vector(query: str) -> list[float] | None:
from lembas.services.library import retrieval
with session_scope() as db:
worker = retrieval.worker_for(db)
worker = retrieval.worker_for(db, group)
return await retrieval.embed_with(worker, query)
# --- Notes -------------------------------------------------------------------
async def _run_notes_search(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
query = str(args.get("query") or "").strip()
vector = await _query_vector(query)
vector = await _query_vector(query, context.data_group)
with session_scope() as db:
user = db.get(User, context.owner_id)
found = (
notes_service.search(db, user, query, limit=8, vector=vector)
notes_service.search(
db, user, query, limit=8, vector=vector, group=context.data_group
)
if query
else notes_service.recent(db, user, limit=8)
else notes_service.recent(db, user, limit=8, group=context.data_group)
)
event = {
"name": "notes_search",
@@ -571,7 +585,7 @@ async def _run_notes_search(context: ToolContext, args: dict[str, Any]) -> ToolO
async def _run_notes_get(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
with session_scope() as db:
user = db.get(User, context.owner_id)
note = notes_service.get(db, str(args.get("id") or ""), user)
note = notes_service.get(db, str(args.get("id") or ""), user, context.data_group)
if note is None:
return ToolOutcome(
"There is no such note, or it is not available to you.",
@@ -599,7 +613,12 @@ async def _run_notes_create(context: ToolContext, args: dict[str, Any]) -> ToolO
with session_scope() as db:
user = db.get(User, context.owner_id)
note = notes_service.create(
db, owner=user, title=title, body=body, author=AUTHOR_MODEL
db,
owner=user,
title=title,
body=body,
author=AUTHOR_MODEL,
group=context.data_group,
)
return ToolOutcome(
f"Saved note {note.id} — {note.title!r}.",
@@ -615,7 +634,7 @@ async def _run_notes_create(context: ToolContext, args: dict[str, Any]) -> ToolO
async def _run_notes_edit(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
with session_scope() as db:
user = db.get(User, context.owner_id)
note = notes_service.get(db, str(args.get("id") or ""), user)
note = notes_service.get(db, str(args.get("id") or ""), user, context.data_group)
if note is None or note.owner_id != context.owner_id:
return ToolOutcome(
"There is no such note, or it belongs to someone else. A note "
@@ -642,7 +661,7 @@ async def _run_notes_edit(context: ToolContext, args: dict[str, Any]) -> ToolOut
async def _run_notes_delete(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
with session_scope() as db:
user = db.get(User, context.owner_id)
note = notes_service.get(db, str(args.get("id") or ""), user)
note = notes_service.get(db, str(args.get("id") or ""), user, context.data_group)
if note is None or note.owner_id != context.owner_id:
return ToolOutcome(
"There is no such note, or it belongs to someone else.",
@@ -735,7 +754,7 @@ async def _run_persona_write(context: ToolContext, args: dict[str, Any]) -> Tool
return _persona_error("persona_write", "There is nobody here to be this with.")
row = personas_service.write(
db,
model_key=context.model_id,
model_key=personas_service.key_for(context.model_id, context.data_group),
owner=user,
content=content,
author=AUTHOR_MODEL,
@@ -782,7 +801,9 @@ async def _run_impression_write(context: ToolContext, args: dict[str, Any]) -> T
if user is None:
return _persona_error("impression_write", "There is nobody here to describe.")
if not content:
row = personas_service.impression(db, context.model_id, user)
row = personas_service.impression(
db, personas_service.key_for(context.model_id, context.data_group), user
)
if row is not None:
personas_service.clear_impression(db, row)
return ToolOutcome(
@@ -791,7 +812,7 @@ async def _run_impression_write(context: ToolContext, args: dict[str, Any]) -> T
)
row = personas_service.write_impression(
db,
model_key=context.model_id,
model_key=personas_service.key_for(context.model_id, context.data_group),
owner=user,
content=content,
author=AUTHOR_MODEL,
@@ -819,7 +840,7 @@ async def _run_memory_add(context: ToolContext, args: dict[str, Any]) -> ToolOut
user = db.get(User, context.owner_id)
try:
memory = memories_service.add(
db, owner=user, content=content, author=AUTHOR_MODEL
db, owner=user, content=content, author=AUTHOR_MODEL, group=context.data_group
)
except ValueError as exc:
return ToolOutcome(
@@ -861,7 +882,7 @@ async def _run_memory_forget(context: ToolContext, args: dict[str, Any]) -> Tool
wanted = str(args.get("content") or "").strip().lower()
with session_scope() as db:
user = db.get(User, context.owner_id)
records = memories_service.all_for(db, user)
records = memories_service.all_for(db, user, context.data_group)
if not wanted:
return ToolOutcome(
"Say which memory to remove, quoting its text.",
@@ -918,6 +939,7 @@ async def _run_report_write(context: ToolContext, args: dict[str, Any]) -> ToolO
source=SOURCE_CHAT,
source_id=context.chat_id or "",
model_id=context.model_id or "",
group=context.data_group,
)
return ToolOutcome(
f"Filed report {report.id} — {report.title!r}. "
@@ -937,9 +959,11 @@ async def _run_report_search(context: ToolContext, args: dict[str, Any]) -> Tool
with session_scope() as db:
user = db.get(User, context.owner_id)
found = (
reports_service.search(db, user, query, limit=8, vector=vector)
reports_service.search(
db, user, query, limit=8, vector=vector, group=context.data_group
)
if query
else reports_service.recent(db, user, limit=8)
else reports_service.recent(db, user, limit=8, group=context.data_group)
)
event = {
"name": "report_search",
@@ -962,7 +986,7 @@ async def _run_report_search(context: ToolContext, args: dict[str, Any]) -> Tool
async def _run_report_get(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
with session_scope() as db:
user = db.get(User, context.owner_id)
report = reports_service.get(db, str(args.get("id") or ""), user)
report = reports_service.get(db, str(args.get("id") or ""), user, context.data_group)
if report is None:
return ToolOutcome(
"There is no such report.",
@@ -985,7 +1009,7 @@ async def _run_skill_get(context: ToolContext, args: dict[str, Any]) -> ToolOutc
name = str(args.get("name") or "").strip()
with session_scope() as db:
user = db.get(User, context.owner_id)
skill = skills_service.by_name(db, name, user)
skill = skills_service.by_name(db, name, user, context.data_group)
# Enforced here and not only in the listing. Without this the per-chat
# narrowing is advisory: a model can name a skill it was never shown --
# from an earlier turn, from a note -- and the runner would fetch it.
@@ -1020,6 +1044,7 @@ async def _run_skill_create(context: ToolContext, args: dict[str, Any]) -> ToolO
description=str(args.get("description") or ""),
body=str(args.get("body") or ""),
author=AUTHOR_MODEL,
group=context.data_group,
)
except skills_service.SkillError as exc:
return ToolOutcome(
@@ -1039,7 +1064,9 @@ async def _run_skill_create(context: ToolContext, args: dict[str, Any]) -> ToolO
async def _run_skill_edit(context: ToolContext, args: dict[str, Any]) -> ToolOutcome:
with session_scope() as db:
user = db.get(User, context.owner_id)
skill = skills_service.by_name(db, str(args.get("name") or ""), user)
skill = skills_service.by_name(
db, str(args.get("name") or ""), user, context.data_group
)
if skill is None or skill.owner_id != context.owner_id:
return ToolOutcome(
"There is no such skill, or it belongs to someone else.",
@@ -1878,7 +1905,16 @@ def resolve_tools(
# something on here would still be reaching for a tool the gates had
# already removed.
off = scoped_off(chat)
empty_library = not skills_service.count_enabled(db, user, exclude=scoped_skills_off(chat))
# Counted in the answering model's own data group: skills in another group
# are not readable here, so they must not keep `skill_get` on offer.
from lembas.services import data_groups
empty_library = not skills_service.count_enabled(
db,
user,
exclude=scoped_skills_off(chat),
group=data_groups.for_speaker(db, user, chat, speaker),
)
# What a crowd speaker may do, which is narrower than what the chat may.
if crowd_turn is not None:
@@ -2095,6 +2131,7 @@ def context_for(
model.
"""
from lembas.services import chat as chat_service
from lembas.services import data_groups
from lembas.services.agent import session as agent_session
if chat is not None and speaker is None:
@@ -2110,6 +2147,11 @@ def context_for(
image_checkpoint=(chat.image_checkpoint or "") if chat is not None else "",
model_id=(speaker.model_id or "") if speaker is not None else "",
connection_id=(speaker.connection_id or "") if speaker is not None else "",
data_group=(
data_groups.for_speaker(db, user, chat, speaker)
if chat is not None
else DEFAULT_GROUP
),
base_ids=[base.id for base in chat.knowledge_bases] if chat is not None else [],
skills_off=scoped_skills_off(chat),
tools=tools.by_name if tools is not None else None,
+69
View File
@@ -1790,3 +1790,72 @@ MESSAGES.update(
"https://mcp.example.com/mcp": "https://mcp.example.com/mcp",
}
)
# --- Data groups (1.10.0) -------------------------------------------------------
MESSAGES.update(
{
'Data group': 'Dátová oblasť',
'Data groups': 'Dátové oblasti',
'Data': 'Dáta',
"Its models read only this group's memories, notes, skills, knowledge and reports, and a chat started on one of them stays in it. Moving a connection does not move any data: it starts reading the other group.": 'Jeho modely čítajú len pamäť, poznámky, schopnosti, znalosti a správy tejto oblasti a konverzácia začatá na niektorom z nich v nej zostáva. Presunutie spojenia nepresúva žiadne dáta: začne čítať druhú oblasť.',
'Several models answering one turn, in any chat. Not an agent-chat feature: it works in an ordinary conversation, and the control is in the composer beside the tool switches.': 'Viacero modelov odpovedá v jednom ťahu, v ľubovoľnej konverzácii. Nie je to funkcia agentových konverzácií: funguje v obyčajnej konverzácii a ovládanie je v okne na písanie vedľa prepínačov nástrojov.',
'All data groups': 'Všetky dátové oblasti',
'A personal group of %(who)s.': 'Osobná oblasť používateľa %(who)s.',
'Connections in this group': 'Spojenia v tejto oblasti',
"Their models read this group's data and nothing else. A connection taken out of a group goes back to the default one. Moving a connection moves no data: it starts reading the other group's.": 'Ich modely čítajú dáta tejto oblasti a nič iné. Spojenie vybraté z oblasti sa vráti do predvolenej. Presunutie spojenia nepresúva žiadne dáta: začne čítať dáta druhej oblasti.',
'There are no connections yet.': 'Zatiaľ tu nie sú žiadne spojenia.',
'A connection leaves the default group by being put into another one.': 'Spojenie opustí predvolenú oblasť tak, že ho zaradíte do inej.',
'Services that read this group': 'Služby, ktoré čítajú túto oblasť',
"The embedder is sent the full text of every document, note, skill and report it indexes, and the reviewer is sent every generated picture with its prompt. A group can name its own, so its data does not have to go to the instance's.": 'Model pre vnorenia dostane celý text každého dokumentu, poznámky, schopnosti a správy, ktoré indexuje, a posudzovateľ dostane každý vygenerovaný obrázok aj s jeho zadaním. Oblasť si môže určiť vlastné, aby jej dáta nemuseli ísť k tým, ktoré zvolila inštancia.',
"The instance's choice": 'Voľba inštancie',
"Changing it leaves this group's search keyword-only until the index is rebuilt on the Extraction page.": 'Po zmene vyhľadáva táto oblasť len podľa kľúčových slov, kým sa index neprestavia na stránke Extrakcia.',
'Image reviewer': 'Posudzovateľ obrázkov',
'Only models that can see images are listed.': 'Zobrazujú sa len modely, ktoré vidia obrázky.',
'Delete the data group “%(name)s”?': 'Zmazať dátovú oblasť „%(name)s“?',
"Where this group's data goes": 'Kam idú dáta tejto oblasti',
"Every provider this group's data is sent to, as the instance has it set. A person holding “Manage their own data groups” may have arranged their own differently. Speech to text and text to speech are not grouped: audio is sent and not kept.": 'Každý poskytovateľ, ktorému sa dáta tejto oblasti posielajú, tak ako to má nastavené inštancia. Kto má oprávnenie „Spravovať vlastné dátové oblasti“, môže mať svoje usporiadané inak. Prevod reči na text a textu na reč sa do oblastí nedelí: zvuk sa posiela a neuchováva.',
'its connection is in %(group)s': 'jeho spojenie je v oblasti %(group)s',
'Nothing: no connection is in this group, and no service reads it.': 'Nikam: v tejto oblasti nie je žiadne spojenie a žiadna služba ju nečíta.',
'Holds': 'Obsahuje',
"A data group keeps one provider's models away from the data another provider's models have been given. Every connection is in one group. Its models read only that group's memories, notes, skills, knowledge, reports and personalities, and a chat stays in the group it was started in.": 'Dátová oblasť drží modely jedného poskytovateľa ďalej od dát, ktoré dostali modely iného poskytovateľa. Každé spojenie patrí do jednej oblasti. Jeho modely čítajú len pamäť, poznámky, schopnosti, znalosti, správy a osobnosti tejto oblasti a konverzácia zostáva v oblasti, v ktorej začala.',
'Data group deleted.': 'Dátová oblasť zmazaná.',
'Instance groups': 'Oblasti inštancie',
'%(n)s connections': 'spojenia: %(n)s',
'%(n)s chats': 'konverzácie: %(n)s',
'own embedder': 'vlastné vnorenia',
'own reviewer': 'vlastný posudzovateľ',
'New data group name': 'Názov novej dátovej oblasti',
'Create data group': 'Vytvoriť dátovú oblasť',
'A new group holds nothing and reads nothing until a connection is put into it. Nothing moves on its own.': 'Nová oblasť nič neobsahuje a nič nečíta, kým do nej nezaradíte spojenie. Nič sa nepresúva samo.',
'Personal groups': 'Osobné oblasti',
'Tick a model to have it answer after this one, then be asked whether it disagrees.': 'Označte model a odpovie po tomto, potom sa ho opýta, či s niečím nesúhlasí.',
'Also answering': 'Odpovedajú aj',
'Skipped, because you cannot reach them any more:': 'Vynechané, lebo k nim už nemáte prístup:',
'In another data group than this chat (%(group)s). Start a new chat to use them:': 'V inej dátovej oblasti než táto konverzácia (%(group)s). Ak ich chcete použiť, začnite novú konverzáciu:',
'Only models whose connection is in this group can read it.': 'Čítať to môžu len modely, ktorých spojenie je v tejto oblasti.',
"Moving it hands it to the other group's models, and takes it away from this group's.": 'Presunutím to dostanú modely druhej oblasti a modely tejto oblasti o to prídu.',
'Only models whose connection is in this group can read it. Moving records between groups needs the permission to manage your own data groups.': 'Čítať to môžu len modely, ktorých spojenie je v tejto oblasti. Na presúvanie záznamov medzi oblasťami potrebujete oprávnenie spravovať vlastné dátové oblasti.',
'Every group': 'Všetky oblasti',
'Show': 'Zobraziť',
'Your data groups': 'Vaše dátové oblasti',
"A model reads only the memories, notes, skills, knowledge, reports and personality of the data group its connection is in, and a chat stays in the group it was started in. Two providers in different groups never see each other's.": 'Model číta len pamäť, poznámky, schopnosti, znalosti, správy a osobnosť tej dátovej oblasti, v ktorej je jeho spojenie, a konverzácia zostáva v oblasti, v ktorej začala. Dvaja poskytovatelia v rôznych oblastiach nikdy nevidia dáta toho druhého.',
'yours': 'vaša',
'Create': 'Vytvoriť',
'A group of your own, for keeping one provider away from the rest of your data. Nobody else can see it.': 'Vlastná oblasť, aby ste jedného poskytovateľa udržali ďalej od zvyšku svojich dát. Nikto iný ju nevidí.',
'Which group each connection reads': 'Ktorú oblasť číta ktoré spojenie',
'Your choice applies to you alone. Following the instance keeps up with whatever the administrator sets. Moving a connection moves no data: from then on it reads the other group, and your chats on it that were started in the old group can no longer use it.': 'Vaša voľba platí len pre vás. Keď sa riadite inštanciou, sledujete to, čo nastaví správca. Presunutie spojenia nepresúva žiadne dáta: odvtedy číta druhú oblasť a vaše konverzácie na ňom, ktoré začali v pôvodnej oblasti, ho už nemôžu použiť.',
'Follow the instance (%(group)s)': 'Podľa inštancie (%(group)s)',
'Reads: %(group)s': 'Číta: %(group)s',
'Set by the administrator. Changing it for yourself needs the permission to manage your own data groups.': 'Nastavuje správca. Ak to chcete zmeniť pre seba, potrebujete oprávnenie spravovať vlastné dátové oblasti.',
'chats': 'konverzácie',
'memories': 'záznamy v pamäti',
'notes': 'poznámky',
'skills': 'schopnosti',
'knowledge bases': 'znalostné bázy',
'reports': 'správy',
'connections': 'spojenia',
'Chat models': 'Modely v konverzáciách',
'Embedding': 'Vnorenia',
'Image review': 'Posudzovanie obrázkov',
}
)
+20
View File
@@ -1811,6 +1811,26 @@ body.is-resizing .canvas__body { pointer-events: none; }
.model-vision { display: flex; color: var(--ink-faint); min-width: 1rem; }
.picker__list--models .picker__tick { display: flex; margin-top: 0; min-width: 1rem; visibility: hidden; }
.picker__list--models .picker__option.is-selected .picker__tick { visibility: visible; }
/* Models in another data group, named under the list inside a chat. Text, not
options: switching this chat to one would hand it the conversation. */
.picker__elsewhere {
border-top: 1px solid var(--border);
padding: var(--sp-2) var(--sp-3);
font-size: var(--text-sm);
color: var(--ink-muted);
}
.picker__elsewhere-head,
.picker__elsewhere-names {
margin: 0;
overflow-wrap: anywhere;
}
.picker__elsewhere-names {
margin-top: var(--sp-1);
color: var(--ink-faint);
}
.picker__empty {
padding: var(--sp-4);
margin: 0;
+11 -1
View File
@@ -279,6 +279,13 @@
return document.getElementById("attachments");
}
/* The composer's model, which on the new-chat screen decides which data
group the library picker may offer. See data_groups.for_composer. */
function modelId() {
var field = document.querySelector('.composer input[name="model_id"], [data-picker-input]');
return field ? field.value : "";
}
function chatId() {
var input = document.getElementById("file-input");
var url = (input && input.dataset.uploadUrl) || "";
@@ -335,7 +342,9 @@
var close = dialog.querySelector("button");
function load(query) {
fetch("/api/files/knowledge-picker?q=" + encodeURIComponent(query || ""), {
fetch("/api/files/knowledge-picker?q=" + encodeURIComponent(query || "") +
"&chat_id=" + encodeURIComponent(chatId()) +
"&model_id=" + encodeURIComponent(modelId()), {
credentials: "same-origin",
})
.then(function (response) { return response.text(); })
@@ -355,6 +364,7 @@
var body = new FormData();
body.append("document_id", option.dataset.attachKnowledge);
body.append("chat_id", chatId());
body.append("model_id", modelId());
postForChip("/api/files/from-knowledge", body);
finish();
});
+11 -2
View File
@@ -62,7 +62,14 @@
/* A chat under way carries its connection on the composer; a new one is
still choosing it, so the select and the hidden field are the truth. */
profileId: picker ? picker.value : (box && box.dataset.profileId) || "",
projectDir: dir ? dir.value : (box && box.dataset.projectDir) || ""
projectDir: dir ? dir.value : (box && box.dataset.projectDir) || "",
/* The model the composer is writing to. Only the new-chat screen needs
it -- a chat that exists answers "which data group" by itself -- but
there it decides which library items may be offered at all. */
modelId: (function () {
var field = document.querySelector('.composer input[name="model_id"], [data-picker-input]');
return field ? field.value : "";
})()
};
}
@@ -171,7 +178,8 @@
"/api/files/mention-picker?q=" + encodeURIComponent(query) +
"&chat_id=" + encodeURIComponent(where.chatId) +
"&profile_id=" + encodeURIComponent(where.profileId) +
"&project_dir=" + encodeURIComponent(where.projectDir);
"&project_dir=" + encodeURIComponent(where.projectDir) +
"&model_id=" + encodeURIComponent(where.modelId);
fetch(url, { credentials: "same-origin" })
.then(function (response) { return response.text(); })
@@ -266,6 +274,7 @@
var body = new FormData();
body.append("chat_id", where.chatId);
body.append("model_id", where.modelId);
if (option.dataset.mentionFile) {
body.append("profile_id", where.profileId);
body.append("path", option.dataset.mentionFile);
@@ -72,6 +72,22 @@
</p>
</div>
{# Which data group this provider reads. Only worth a control once there is a
second group to choose; until then every connection is in the default one
and the field would be a select with one option. #}
{% if data_group_choices is defined and data_group_choices|length > 1 %}
<div class="field">
<label class="field__label" for="group-{{ connection.id }}">{{ t("Data group") }}</label>
<select class="select" id="group-{{ connection.id }}" name="data_group_id">
{% for group in data_group_choices %}
<option value="{{ group.id }}"
{{ 'selected' if (connection.data_group_id or 'default') == group.id }}>{{ group.name }}</option>
{% endfor %}
</select>
<p class="field__hint">{{ t("Its models read only this group's memories, notes, skills, knowledge and reports, and a chat started on one of them stays in it. Moving a connection does not move any data: it starts reading the other group.") }}</p>
</div>
{% endif %}
<div class="field">
<label class="field__label" for="unload-{{ connection.id }}">{{ t("Unload URL") }}</label>
<div class="btn-row">
@@ -50,6 +50,12 @@
{{ icon("sliders", "icon--sm") }}
<span class="nav-item__label">{{ t("Models") }}</span>
</a>
{# Beside Connections and Models because it is about them: which
provider's models may read which part of the people's data. #}
<a class="nav-item {{ 'is-active' if section == 'data-groups' }}" href="/admin/data-groups">
{{ icon("shield", "icon--sm") }}
<span class="nav-item__label">{{ t("Data groups") }}</span>
</a>
{# Its own entry rather than a card on Agents, where it started. Sitting
there made it read as an agent-chat feature -- which is what the owner
took it for, reasonably, since that is what the page is called. #}
@@ -0,0 +1,143 @@
{% extends "admin/_layout.html" %}
{% from "_macros.html" import icon %}
{% set section = "data-groups" %}
{% block title %}{{ group.name }} - {{ t("Data groups") }} - {{ brand.name }}{% endblock %}
{% block heading %}{{ group.name }}{% endblock %}
{% block admin_content %}
<p class="admin-lede">
<a href="/admin/data-groups">{{ icon("chevron-left", "icon--sm") }} {{ t("All data groups") }}</a>
{% if owner %}· {{ t("A personal group of %(who)s.", who=owner.email) }}{% endif %}
</p>
{% if saved %}
<div class="alert alert--success">{{ icon("check", "icon--sm") }} <span>{{ t("Settings saved.") }}</span></div>
{% endif %}
{% if error %}
<div class="alert alert--error">{{ icon("warning", "alert__icon") }} <span>{{ error }}</span></div>
{% endif %}
{#
The delete form is declared before the main one and reached by `form=`, so it
can never end up nested inside it -- a nested <form> truncates its parent, the
1.3.0 bug that made a whole admin page unsaveable.
#}
<form id="delete-group" method="post" action="/admin/data-groups/{{ group.id }}/delete"></form>
<form method="post" action="/admin/data-groups/{{ group.id }}" class="form-grid">
<section class="card">
<h2 class="card__title">{{ t("Name") }}</h2>
<div class="field">
<label class="field__label" for="name">{{ t("Name") }}</label>
<input class="input" id="name" name="name" value="{{ group.name }}" required>
<p class="field__hint"></p>
</div>
<div class="field">
<label class="field__label" for="description">{{ t("What it is for") }}</label>
<input class="input" id="description" name="description"
value="{{ group.description }}" maxlength="2000">
<p class="field__hint"></p>
</div>
</section>
{% if not owner %}
<section class="card">
<h2 class="card__title">{{ t("Connections in this group") }}</h2>
<p class="card__lede">{{ t("Their models read this group's data and nothing else. A connection taken out of a group goes back to the default one. Moving a connection moves no data: it starts reading the other group's.") }}</p>
<input type="hidden" name="connections_sent" value="1">
<div class="field">
{% for connection in connections %}
<label class="checkbox">
<input type="checkbox" name="connection_ids" value="{{ connection.id }}"
{{ 'checked' if connection.id in member_ids }}
{{ 'disabled' if group.is_default and connection.id in member_ids }}>
<span>{{ connection.name }}{% if not connection.enabled %} <span class="badge">{{ t("disabled") }}</span>{% endif %}</span>
</label>
{% else %}
<p class="field__hint">{{ t("There are no connections yet.") }}</p>
{% endfor %}
{% if group.is_default %}
<p class="field__hint">{{ t("A connection leaves the default group by being put into another one.") }}</p>
{% endif %}
</div>
</section>
{% endif %}
<section class="card">
<h2 class="card__title">{{ t("Services that read this group") }}</h2>
<p class="card__lede">{{ t("The embedder is sent the full text of every document, note, skill and report it indexes, and the reviewer is sent every generated picture with its prompt. A group can name its own, so its data does not have to go to the instance's.") }}</p>
<div class="field">
<label class="field__label" for="embedding">{{ t("Embedding model") }}</label>
<select class="select" id="embedding" name="embedding">
<option value="">{{ t("The instance's choice") }}</option>
{% for model in embedders %}
<option value="{{ model.model_id }}|{{ model.connection_id }}"
{{ 'selected' if group.embedding_model_id == model.model_id and (not group.embedding_connection_id or group.embedding_connection_id == model.connection_id) }}>
{{ model.label }} · {{ model.connection.name }}
</option>
{% endfor %}
</select>
<p class="field__hint">{{ t("Changing it leaves this group's search keyword-only until the index is rebuilt on the Extraction page.") }}</p>
</div>
<div class="field">
<label class="field__label" for="reviewer">{{ t("Image reviewer") }}</label>
<select class="select" id="reviewer" name="reviewer">
<option value="">{{ t("The instance's choice") }}</option>
{% for model in reviewers %}
<option value="{{ model.model_id }}|{{ model.connection_id }}"
{{ 'selected' if group.review_model_id == model.model_id and (not group.review_connection_id or group.review_connection_id == model.connection_id) }}>
{{ model.label }} · {{ model.connection.name }}
</option>
{% endfor %}
</select>
<p class="field__hint">{{ t("Only models that can see images are listed.") }}</p>
</div>
</section>
<div class="btn-row">
<button class="btn btn--primary" type="submit">{{ t("Save") }}</button>
{% if not group.is_default %}
<button class="btn btn--danger" type="submit" form="delete-group"
data-confirm-button="{{ t('Delete the data group “%(name)s”?', name=group.name) }}">
{{ icon("trash", "icon--sm") }} {{ t("Delete") }}
</button>
{% endif %}
</div>
</form>
<section class="card" style="margin-top: var(--sp-6)">
<h2 class="card__title">{{ t("Where this group's data goes") }}</h2>
<p class="card__lede">{{ t("Every provider this group's data is sent to, as the instance has it set. A person holding “Manage their own data groups” may have arranged their own differently. Speech to text and text to speech are not grouped: audio is sent and not kept.") }}</p>
{% if exits %}
{# Rows rather than a table, so a phone gets a stack instead of a sideways
scroll -- the same list-row pattern as every other admin list. #}
<div class="model-rows">
{% for exit in exits %}
<div class="model-row">
<div class="model-row__main">
<span class="model-row__name">{{ labels.get(exit.what, exit.what) }}</span>
<span class="model-row__id">
{%- if exit.model %}{{ exit.model }} · {% endif %}{{ exit.connection -}}
</span>
</div>
<div class="model-row__meta">
{% if exit.elsewhere %}
<span class="badge badge--danger">{{ t("its connection is in %(group)s", group=exit.elsewhere) }}</span>
{% endif %}
</div>
</div>
{% endfor %}
</div>
{% else %}
<p class="field__hint">{{ t("Nothing: no connection is in this group, and no service reads it.") }}</p>
{% endif %}
<p class="field__hint">
{{ t("Holds") }}:
{% for label, n in counts.items() %}{{ n }} {{ labels.get(label, label) }}{% if not loop.last %}, {% endif %}{% endfor %}.
</p>
</section>
{% endblock %}
@@ -0,0 +1,73 @@
{% extends "admin/_layout.html" %}
{% from "_macros.html" import icon %}
{% set section = "data-groups" %}
{% block title %}{{ t("Data groups") }} - {{ brand.name }}{% endblock %}
{% block heading %}{{ t("Data groups") }}{% endblock %}
{% block admin_content %}
<p class="admin-lede">{{ t("A data group keeps one provider's models away from the data another provider's models have been given. Every connection is in one group. Its models read only that group's memories, notes, skills, knowledge, reports and personalities, and a chat stays in the group it was started in.") }}</p>
{% if saved %}
<div class="alert alert--success">{{ icon("check", "icon--sm") }} <span>{{ t("Data group deleted.") if saved == "deleted" else t("Settings saved.") }}</span></div>
{% endif %}
<h2 class="admin-section-title">
{{ t("Instance groups") }} <span class="badge">{{ groups|length }}</span>
</h2>
<div class="model-rows">
{% for group in groups %}
<a class="model-row" href="/admin/data-groups/{{ group.id }}">
<div class="model-row__main">
<span class="model-row__name">{{ group.name }}</span>
{% if group.description %}
<span class="model-row__id">{{ group.description }}</span>
{% endif %}
</div>
<div class="model-row__meta">
{% if group.is_default %}<span class="badge badge--leaf">{{ t("default") }}</span>{% endif %}
<span class="badge">{{ t("%(n)s connections", n=connections[group.id]) }}</span>
<span class="badge">{{ t("%(n)s chats", n=counts[group.id]["chats"]) }}</span>
{% if group.embedding_model_id %}<span class="badge">{{ t("own embedder") }}</span>{% endif %}
{% if group.review_model_id %}<span class="badge">{{ t("own reviewer") }}</span>{% endif %}
</div>
</a>
{% endfor %}
</div>
<section class="card" style="margin-top: var(--sp-6)">
<form method="post" action="/admin/data-groups" class="btn-row">
<input class="input" name="name" placeholder="{{ t('New data group name') }}" required
aria-label="{{ t('New data group name') }}">
<button class="btn btn--primary" type="submit">
{{ icon("plus", "icon--sm") }} {{ t("Create data group") }}
</button>
</form>
<p class="field__hint">{{ t("A new group holds nothing and reads nothing until a connection is put into it. Nothing moves on its own.") }}</p>
</section>
{#
Personal groups are somebody's own arrangement of their own data -- made by
a person holding "Manage their own data groups" -- and are listed so an
administrator can see they exist, not so they can be rearranged from here.
#}
{% if personal %}
<h2 class="admin-section-title">
{{ t("Personal groups") }} <span class="badge">{{ personal|length }}</span>
</h2>
<div class="model-rows">
{% for group in personal %}
<a class="model-row" href="/admin/data-groups/{{ group.id }}">
<div class="model-row__main">
<span class="model-row__name">{{ group.name }}</span>
<span class="model-row__id">{{ owners[group.owner_id].email if group.owner_id in owners else "" }}</span>
</div>
<div class="model-row__meta">
<span class="badge">{{ t("%(n)s chats", n=counts[group.id]["chats"]) }}</span>
</div>
</a>
{% endfor %}
</div>
{% endif %}
{% endblock %}
@@ -80,6 +80,22 @@
{% endfor %}
</div>
<p class="picker__empty" data-picker-empty hidden>{{ t("No model matches that.") }}</p>
{#
Models in another data group. Named rather than left out, so somebody
looking for one learns where it went; not options, because switching this
chat to one would hand it the whole conversation. Only inside a chat --
on /chat every model can start one.
#}
{% if chat and models_elsewhere %}
<div class="picker__elsewhere">
<p class="picker__elsewhere-head">
{{ t("In another data group than this chat (%(group)s). Start a new chat to use them:", group=chat_group_name) }}
</p>
<p class="picker__elsewhere-names">
{%- for model in models_elsewhere %}{{ model.label }}{% if not loop.last %}, {% endif %}{% endfor -%}
</p>
</div>
{% endif %}
</div>
{% if chat %}
@@ -0,0 +1,77 @@
{#
Data groups on the library pages. Import `with context`: both macros read
`several_groups`, `data_groups`, `group_names` and `may_move_groups`, which
`api/library.py:_groups` puts on every library page.
Nothing here appears on an instance with one group, so a library that never
uses groups looks exactly as it always did.
#}
{# Which group a row is in, as a badge. Only when there is a choice to see. #}
{% macro group_chip(row) -%}
{%- if several_groups -%}
<span class="badge" title="{{ t('Data group') }}">{{ group_names.get(row.data_group_id or 'default', '') }}</span>
{%- endif -%}
{%- endmacro %}
{#
The group a new record goes into, or the one an existing record is in.
A new record may go into any group this person can use: it is theirs, and
choosing where it lives is choosing which providers may read it. Moving one
that exists is `data.manage`, because that changes what a provider has
already been able to see. Every slot is emitted even when one is empty, so a
field row lines up by construction.
#}
{% macro group_field(row=None, chosen="") -%}
{%- if several_groups -%}
{%- set current = (row.data_group_id if row else chosen) or "default" -%}
<div class="field">
<label class="field__label" for="data_group_id">{{ t("Data group") }}</label>
{%- if not row or may_move_groups %}
<select class="select" id="data_group_id" name="data_group_id">
{%- for group in data_groups %}
<option value="{{ group.id }}" {{ 'selected' if group.id == current }}>{{ group.name }}</option>
{%- endfor %}
</select>
{%- else %}
<input class="input" id="data_group_id" value="{{ group_names.get(current, '') }}" disabled>
{%- endif %}
<p class="field__hint">
{%- if not row -%}
{{ t("Only models whose connection is in this group can read it.") }}
{%- elif may_move_groups -%}
{{ t("Moving it hands it to the other group's models, and takes it away from this group's.") }}
{%- else -%}
{{ t("Only models whose connection is in this group can read it. Moving records between groups needs the permission to manage your own data groups.") }}
{%- endif -%}
</p>
</div>
{%- endif -%}
{%- endmacro %}
{#
The list filter: a select that submits itself, as a GET so the filtered list
is a URL. The other filters on the page ride along as hidden fields.
#}
{% macro group_filter(action, current="", keep={}) -%}
{%- if several_groups -%}
{#- A compact control rather than a field: the label is for a screen reader,
because the options ("Every group", the group names) already say what the
select is for, and a full-width field here outweighed the list below it. #}
<form method="get" action="{{ action }}" class="btn-row" style="margin-bottom: var(--sp-4)">
{%- for name, value in keep.items() %}{% if value %}
<input type="hidden" name="{{ name }}" value="{{ value }}">
{%- endif %}{% endfor %}
<label class="visually-hidden" for="group-filter">{{ t("Data group") }}</label>
<select class="select" id="group-filter" name="group" style="flex: none; width: auto"
onchange="this.form.submit()">
<option value="">{{ t("Every group") }}</option>
{%- for group in data_groups %}
<option value="{{ group.id }}" {{ 'selected' if group.id == current }}>{{ group.name }}</option>
{%- endfor %}
</select>
<noscript><button class="btn btn--sm" type="submit">{{ t("Show") }}</button></noscript>
</form>
{%- endif -%}
{%- endmacro %}
@@ -1,5 +1,6 @@
{% extends "library/_layout.html" %}
{% from "_macros.html" import icon %}
{% from "library/_group.html" import group_chip, group_field, group_filter with context %}
{% set section = "knowledge" %}
{% block title %}{{ base.name }} - {{ brand.name }}{% endblock %}
@@ -102,6 +103,7 @@
value="{{ base.description }}" {{ 'disabled' if not is_owner }}>
</div>
</div>
{% if is_owner %}{{ group_field(base) }}{% endif %}
</section>
{% include "library/_share.html" %}
@@ -1,5 +1,6 @@
{% extends "library/_layout.html" %}
{% from "_macros.html" import icon %}
{% from "library/_group.html" import group_chip, group_field, group_filter with context %}
{% set section = "knowledge" %}
{% block title %}Knowledge - {{ brand.name }}{% endblock %}
@@ -14,6 +15,7 @@
<a class="filter-tab {{ 'is-active' if shared }}"
href="/library/knowledge?shared=1">Shared with me</a>
</div>
{{ group_filter("/library/knowledge", group, {"shared": "1" if shared else ""}) }}
{% if error %}
<div class="alert alert--error">{{ icon("warning", "alert__icon") }} <span>{{ error }}</span></div>
@@ -31,6 +33,7 @@
</div>
</div>
<div class="btn-row">
{{ group_chip(base) }}
{% if base.owner_id != user.id %}<span class="badge">{{ t("shared with you") }}</span>{% endif %}
</div>
</li>
@@ -59,6 +62,7 @@
placeholder="{{ t('What belongs in here.') }}">
</div>
</div>
{{ group_field(None, group) }}
<button class="btn btn--primary" type="submit">{{ icon("plus", "icon--sm") }} Create</button>
</form>
</section>
@@ -1,5 +1,6 @@
{% extends "library/_layout.html" %}
{% from "_macros.html" import icon %}
{% from "library/_group.html" import group_chip, group_field, group_filter with context %}
{% set section = "notes" %}
{% block title %}{{ note.title if note else "New note" }} - {{ brand.name }}{% endblock %}
@@ -27,6 +28,7 @@
{{ 'disabled' if note and not is_owner }}
placeholder="{{ t('Markdown.') }}">{{ note.body if note else '' }}</textarea>
</div>
{% if not note or is_owner %}{{ group_field(note, group|default("")) }}{% endif %}
{% if note and note.author == "model" %}
<p class="field__hint">
{{ icon("sparkle", "icon--sm") }}
@@ -1,5 +1,6 @@
{% extends "library/_layout.html" %}
{% from "_macros.html" import icon %}
{% from "library/_group.html" import group_chip, group_field, group_filter with context %}
{% set section = "notes" %}
{% block title %}Notes - {{ brand.name }}{% endblock %}
@@ -20,6 +21,7 @@
<a class="filter-tab {{ 'is-active' if not shared }}" href="/library/notes">All notes</a>
<a class="filter-tab {{ 'is-active' if shared }}" href="/library/notes?shared=1">Shared with me</a>
</div>
{{ group_filter("/library/notes", group, {"shared": "1" if shared else "", "q": q}) }}
<form method="get" action="/library/notes" class="btn-row" style="margin-bottom: var(--sp-5)">
<input class="input" type="search" name="q" value="{{ q }}" style="flex: 1"
placeholder="{{ t('Search notes…') }}">
@@ -49,6 +51,7 @@
<div class="text-xs faint">{{ note.body[:160] }}{{ "…" if note.body|length > 160 }}</div>
</div>
<div class="btn-row">
{{ group_chip(note) }}
{% if note.author == "model" %}<span class="badge badge--leaf">{{ t("written by a model") }}</span>{% endif %}
{% if note.owner_id != user.id %}<span class="badge">{{ t("shared") }}</span>{% endif %}
</div>
@@ -1,5 +1,6 @@
{% extends "library/_layout.html" %}
{% from "_macros.html" import icon %}
{% from "library/_group.html" import group_chip, group_field, group_filter with context %}
{% set section = "skills" %}
{% block title %}{{ skill.name if skill else "New skill" }} - {{ brand.name }}{% endblock %}
@@ -40,6 +41,7 @@
{{ 'disabled' if skill and not is_owner }}
placeholder="{{ t('Markdown. Steps, conventions, things to avoid.') }}">{{ skill.body if skill else '' }}</textarea>
</div>
{% if not skill or is_owner %}{{ group_field(skill, group|default("")) }}{% endif %}
{% if skill %}
<div class="field">
@@ -1,5 +1,6 @@
{% extends "library/_layout.html" %}
{% from "_macros.html" import icon %}
{% from "library/_group.html" import group_chip, group_field, group_filter with context %}
{% set section = "skills" %}
{% block title %}Skills - {{ brand.name }}{% endblock %}
@@ -19,6 +20,7 @@
<a class="filter-tab {{ 'is-active' if not shared }}" href="/library/skills">All skills</a>
<a class="filter-tab {{ 'is-active' if shared }}" href="/library/skills?shared=1">Shared with me</a>
</div>
{{ group_filter("/library/skills", group, {"shared": "1" if shared else "", "q": q}) }}
<form method="get" action="/library/skills" class="btn-row" style="margin-bottom: var(--sp-5)">
<input class="input" type="search" name="q" value="{{ q }}" style="flex: 1"
placeholder="{{ t('Search skills…') }}">
@@ -48,6 +50,7 @@
<div class="text-xs faint">{{ skill.description }}</div>
</div>
<div class="btn-row">
{{ group_chip(skill) }}
{% if not skill.enabled %}<span class="badge">{{ t("off") }}</span>{% endif %}
{% if skill.author == "model" %}<span class="badge badge--leaf">{{ t("written by a model") }}</span>{% endif %}
{% if skill.owner_id != user.id %}<span class="badge">{{ t("shared") }}</span>{% endif %}
+104 -2
View File
@@ -57,6 +57,11 @@
<label class="tabs__tab" for="tab-memory">{{ icon("sparkle", "icon--sm") }} Memory</label>
{% endif %}
{% if several_groups or may_manage_groups %}
<input class="visually-hidden" type="radio" name="settings-tab" id="tab-data">
<label class="tabs__tab" for="tab-data">{{ icon("shield", "icon--sm") }} {{ t("Data") }}</label>
{% endif %}
<input class="visually-hidden" type="radio" name="settings-tab" id="tab-security">
<label class="tabs__tab" for="tab-security">{{ icon("key", "icon--sm") }} Security</label>
</div>
@@ -388,6 +393,9 @@
<input class="input" name="content" value="{{ memory.content }}"
maxlength="{{ memory_limit }}" style="flex: 1"
aria-label="{{ t('What this memory says') }}">
{% if several_groups %}
<span class="badge" title="{{ t('Data group') }}">{{ group_names.get(memory.data_group_id or 'default', '') }}</span>
{% endif %}
<button class="btn btn--sm" type="submit">{{ t("Save") }}</button>
<button class="btn btn--sm btn--danger" type="submit"
formaction="/api/library/memories/{{ memory.id }}/delete"
@@ -416,6 +424,14 @@
maxlength="{{ memory_limit }}"
aria-label="{{ t('Something worth remembering') }}"
placeholder="{{ t('Prefers metric units and a 24-hour clock.') }}">
{% if several_groups %}
<select class="select" name="data_group_id" aria-label="{{ t('Data group') }}"
style="flex: none">
{% for group in data_groups %}
<option value="{{ group.id }}">{{ group.name }}</option>
{% endfor %}
</select>
{% endif %}
<button class="btn btn--primary" type="submit">{{ t("Remember") }}</button>
</form>
<p class="field__hint">
@@ -439,7 +455,9 @@
{% for personality in personalities %}
<li class="model-list__item">
<div style="min-width: 0">
<strong>{{ personality.model_key }}</strong>
{% set key = split_key(personality.model_key) %}
<strong>{{ key[0] }}</strong>
{% if several_groups %}<span class="badge">{{ group_names.get(key[1], key[1]) }}</span>{% endif %}
{% if personality.author == "model" %}
<span class="badge badge--leaf">{{ t("its own words") }}</span>
{% endif %}
@@ -484,7 +502,9 @@
{% for impression in impressions %}
<li class="model-list__item">
<div style="min-width: 0">
<strong>{{ impression.model_key }}</strong>
{% set key = split_key(impression.model_key) %}
<strong>{{ key[0] }}</strong>
{% if several_groups %}<span class="badge">{{ group_names.get(key[1], key[1]) }}</span>{% endif %}
<div class="text-sm">{{ impression.content }}</div>
</div>
<form method="post"
@@ -504,6 +524,88 @@
</section>
{% endif %}
{# --- Data groups --- #}
{#
Which provider may read which of this person's data. Shown to anybody
once there is more than one group, because "which provider can read
my notes?" is a question anybody may ask; editable only with
`data.manage`, because the answer changes what a provider can see.
#}
{% if several_groups or may_manage_groups %}
<section class="tabs__panel" data-tab="tab-data">
<div class="card">
<h2 class="card__title">{{ t("Your data groups") }}</h2>
<p class="card__lede">{{ t("A model reads only the memories, notes, skills, knowledge, reports and personality of the data group its connection is in, and a chat stays in the group it was started in. Two providers in different groups never see each other's.") }}</p>
<ul class="model-list">
{% for group in data_groups %}
<li class="model-list__item">
<div style="min-width: 0">
<strong>{{ group.name }}</strong>
{% if group.personal %}<span class="badge">{{ t("yours") }}</span>{% endif %}
<div class="text-xs faint">
{% for label, n in group_counts[group.id].items() %}{{ n }} {{ group_labels.get(label, label) }}{% if not loop.last %} · {% endif %}{% endfor %}
</div>
</div>
{% if group.personal and may_manage_groups %}
<form method="post" action="/api/preferences/data-groups/{{ group.id }}/delete">
<button class="btn btn--sm btn--danger" type="submit"
data-confirm-button="{{ t('Delete the data group “%(name)s”?', name=group.name) }}"
aria-label="{{ t('Delete this') }}" title="{{ t('Delete this') }}">
{{ icon("trash", "icon--sm") }}
</button>
</form>
{% endif %}
</li>
{% endfor %}
</ul>
{% if may_manage_groups %}
<form method="post" action="/api/preferences/data-groups/new" class="row"
style="gap: var(--sp-2); margin-top: var(--sp-4)">
<input class="input" name="name" required maxlength="120" style="flex: 1"
aria-label="{{ t('New data group name') }}"
placeholder="{{ t('New data group name') }}">
<button class="btn" type="submit">{{ icon("plus", "icon--sm") }} {{ t("Create") }}</button>
</form>
<p class="field__hint">{{ t("A group of your own, for keeping one provider away from the rest of your data. Nobody else can see it.") }}</p>
{% endif %}
</div>
<div class="card">
<h2 class="card__title">{{ t("Which group each connection reads") }}</h2>
{% if may_manage_groups %}
<p class="card__lede">{{ t("Your choice applies to you alone. Following the instance keeps up with whatever the administrator sets. Moving a connection moves no data: from then on it reads the other group, and your chats on it that were started in the old group can no longer use it.") }}</p>
<form method="post" action="/api/preferences/data-groups" class="form-grid">
{% for row in group_connections %}
<div class="field">
<label class="field__label" for="dg-{{ row.connection.id }}">{{ row.connection.name }}</label>
<select class="select" id="dg-{{ row.connection.id }}" name="group__{{ row.connection.id }}">
<option value="">{{ t("Follow the instance (%(group)s)", group=group_names.get(row.instance, '')) }}</option>
{% for group in data_groups %}
<option value="{{ group.id }}" {{ 'selected' if row.chosen == group.id }}>{{ group.name }}</option>
{% endfor %}
</select>
<p class="field__hint">{{ t("Reads: %(group)s", group=group_names.get(row.in_force, '')) }}</p>
</div>
{% endfor %}
<div class="btn-row">
<button class="btn btn--primary" type="submit">{{ t("Save") }}</button>
</div>
</form>
{% else %}
<p class="card__lede">{{ t("Set by the administrator. Changing it for yourself needs the permission to manage your own data groups.") }}</p>
<ul class="model-list">
{% for row in group_connections %}
<li class="model-list__item">
<strong>{{ row.connection.name }}</strong>
<span class="badge">{{ group_names.get(row.in_force, '') }}</span>
</li>
{% endfor %}
</ul>
{% endif %}
</div>
</section>
{% endif %}
{# --- Security --- #}
<section class="tabs__panel" data-tab="tab-security">
<div class="card">
+245
View File
@@ -0,0 +1,245 @@
"""A chat stays in the data group it was started in.
Its history is the group's data, so every way a chat could reach a model in
another group is closed here: switching its model, the endpoint fallback, a
crowd member, a friend, the roster, a base, the `@` menu, and Messages. And if
a chat's own model is moved into another group afterwards, the next reply is
refused rather than sent.
"""
from __future__ import annotations
import pytest
from sqlalchemy import select
from lembas.db.models import (
DEFAULT_GROUP,
TARGET_MESSAGES,
Chat,
Connection,
CrowdMember,
DataGroup,
Model,
User,
)
from lembas.services import chat as chat_service
from lembas.services import data_groups, schedules, settings_store
from lembas.services import subagent as subagent_service
from lembas.services.crypto import encrypt
from lembas.services.library import notes
HOSTED = "hosted"
@pytest.fixture
def owner(db, registered) -> User:
return db.scalars(select(User).order_by(User.created_at)).first()
@pytest.fixture
def setup(db, owner):
"""Two local models in the default group, one hosted model in its own."""
db.add(DataGroup(id=HOSTED, name="Hosted"))
local = Connection(name="Local", base_url="http://127.0.0.1:1", api_key_encrypted=encrypt(""))
cloud = Connection(
name="Cloud",
base_url="http://127.0.0.1:2",
api_key_encrypted=encrypt(""),
data_group_id=HOSTED,
)
db.add_all([local, cloud])
db.flush()
db.add_all(
[
Model(connection_id=local.id, model_id="local-a", position=0),
Model(connection_id=local.id, model_id="local-b", position=1),
Model(connection_id=cloud.id, model_id="cloud-model", position=2),
]
)
db.commit()
return local, cloud
def _chat(db, owner, model_id="local-a", connection=None, group=DEFAULT_GROUP) -> Chat:
chat = Chat(
user_id=owner.id,
model_id=model_id,
connection_id=connection.id if connection else None,
data_group_id=group,
title="t",
)
db.add(chat)
db.commit()
return chat
# --- Starting and switching ---------------------------------------------------------
def test_a_new_chat_takes_its_models_group(client, db, setup):
client.post("/api/chats/start", data={"content": "hi", "model_id": "cloud-model"})
chat = db.scalars(select(Chat).order_by(Chat.created_at.desc())).first()
assert chat.data_group_id == HOSTED
def test_a_chat_cannot_be_switched_to_another_groups_model(client, db, owner, setup):
local, _ = setup
chat = _chat(db, owner, connection=local)
response = client.patch(f"/api/chats/{chat.id}", data={"model_id": "cloud-model"})
assert response.status_code == 409
assert "data group" in response.text
db.refresh(chat)
assert chat.model_id == "local-a"
def test_a_chat_can_switch_within_its_group(client, db, owner, setup):
local, _ = setup
chat = _chat(db, owner, connection=local)
response = client.patch(f"/api/chats/{chat.id}", data={"model_id": "local-b"})
assert response.status_code in (200, 204)
db.refresh(chat)
assert chat.model_id == "local-b"
def test_the_picker_names_the_models_it_leaves_out(client, db, owner, setup):
"""Named, not silently missing -- and not offered as options either."""
local, _ = setup
chat = _chat(db, owner, connection=local)
page = client.get(f"/chat/{chat.id}").text
assert "In another data group" in page
assert 'data-picker-value="cloud-model"' not in page
def test_available_models_narrow_to_a_group(db, owner, setup):
ids = [m.model_id for m in chat_service.available_models(db, owner, HOSTED)]
assert ids == ["cloud-model"]
everything = [m.model_id for m in chat_service.available_models(db, owner)]
assert everything == ["local-a", "local-b", "cloud-model"]
# --- The endpoint fallback ----------------------------------------------------------
def test_the_fallback_never_repoints_a_chat_into_another_group(db, owner, setup):
"""A model id served by two connections: the chat's own going away must not
land it on the other provider, which would be handed the whole history."""
local, cloud = setup
db.add(Model(connection_id=cloud.id, model_id="local-a"))
local.enabled = False
db.commit()
chat = _chat(db, owner, connection=local)
with pytest.raises(chat_service.LLMError):
chat_service.resolve_endpoint(db, chat)
db.refresh(chat)
assert chat.connection_id == local.id
def test_a_chat_whose_model_moved_is_refused_rather_than_sent(db, owner, setup):
local, _ = setup
chat = _chat(db, owner, connection=local)
speaker = chat_service.speaker_for(db, chat)
assert data_groups.refusal(db, owner, chat, speaker) == ""
local.data_group_id = HOSTED
db.commit()
refusal = data_groups.refusal(db, owner, chat, speaker)
assert "Default" in refusal and "Hosted" in refusal
# --- Other models reaching the conversation ------------------------------------------
def test_a_crowd_member_from_another_group_is_refused(client, db, owner, setup):
settings_store.update(db, {"enabled": True}, key=settings_store.CROWD)
local, _ = setup
chat = _chat(db, owner, connection=local)
client.patch(f"/api/chats/{chat.id}", data={"crowd_model_ids": ["local-b", "cloud-model"]})
members = [row.model_id for row in db.scalars(select(CrowdMember))]
assert members == ["local-b"]
def test_a_crowd_member_that_left_the_group_does_not_speak(db, owner, setup):
local, cloud = setup
chat = _chat(db, owner, connection=local)
db.add(CrowdMember(chat_id=chat.id, model_id="cloud-model", connection_id=cloud.id))
db.commit()
from lembas.services import crowd
speakers = [s.model_id for s in crowd.member_speakers(db, chat, owner)]
assert speakers == ["local-a"]
def test_the_roster_and_the_friend_stay_in_the_group(db, owner, setup):
roster = chat_service.roster_block(db, owner, exclude="local-a", group=DEFAULT_GROUP)
assert "local-b" in roster and "cloud-model" not in roster
friend, refusal = subagent_service._resolve_friend(
db, owner, "cloud-model", asking="local-a", group=DEFAULT_GROUP
)
assert friend is None
# The name asked for is echoed back; the list of who *can* be asked is not
# allowed to carry it.
offered = refusal.split("These are the ones you can:")[-1]
assert "cloud-model" not in offered and "local-b" in offered
def test_a_friend_reads_its_own_group(db, owner, setup):
local, cloud = setup
parent = _chat(db, owner, connection=local)
friend = db.scalar(select(Model).where(Model.model_id == "cloud-model"))
child = subagent_service._create_child(db, parent, title="q", write=False, friend=friend)
assert child.data_group_id == HOSTED
# --- Bases and the @ menu -------------------------------------------------------------
def test_a_base_from_another_group_cannot_be_attached(client, db, owner, setup):
from lembas.services.library import documents
local, _ = setup
chat = _chat(db, owner, connection=local)
base = documents.create_base(db, owner=owner, name="Hosted base", group=HOSTED)
response = client.post(f"/api/chats/{chat.id}/bases", data={"base_id": base.id})
assert response.status_code == 404
def test_the_mention_menu_offers_only_the_chats_group(client, db, owner, setup):
local, _ = setup
chat = _chat(db, owner, connection=local)
notes.create(db, owner=owner, title="Home note", body="x")
notes.create(db, owner=owner, title="Hosted note", body="x", group=HOSTED)
page = client.get(f"/api/files/mention-picker?q=&chat_id={chat.id}").text
assert "Home note" in page and "Hosted note" not in page
def test_on_the_new_chat_screen_the_chosen_model_decides(client, db, owner, setup):
notes.create(db, owner=owner, title="Hosted note", body="x", group=HOSTED)
page = client.get("/api/files/mention-picker?q=&model_id=cloud-model").text
assert "Hosted note" in page
def test_a_note_from_another_group_cannot_be_attached(client, db, owner, setup):
local, _ = setup
chat = _chat(db, owner, connection=local)
hosted = notes.create(db, owner=owner, title="Hosted note", body="x", group=HOSTED)
response = client.post(
"/api/files/from-note", data={"note_id": hosted.id, "chat_id": chat.id}
)
assert "not available" in response.text
# --- Messages and schedules -------------------------------------------------------------
def test_a_schedule_from_another_group_cannot_post_to_messages(db, owner, setup):
rule = {"at": {"weekdays": [0], "times": ["15:00"]}}
with pytest.raises(schedules.ScheduleError, match="Messages"):
schedules.create(
db,
owner=owner,
title="t",
instruction="i",
rule=rule,
target=TARGET_MESSAGES,
model_id="cloud-model",
)
def test_a_schedule_is_stamped_with_its_models_group(db, owner, setup):
rule = {"at": {"weekdays": [0], "times": ["15:00"]}}
schedule = schedules.create(
db, owner=owner, title="t", instruction="i", rule=rule, model_id="cloud-model"
)
assert schedule.data_group_id == HOSTED
assert db.get(Chat, schedule.chat_id).data_group_id == HOSTED
+221
View File
@@ -0,0 +1,221 @@
"""A model reads one data group's data, and only that one -- from both sides.
Every store is checked twice: in the harness, where memories, skills and a
personality are *handed* to a model, and in the tool runners, where a model goes
looking. A test that only covered the search would miss the fetch by id, which is
the path a model takes after learning an id from somewhere it should not have.
"""
from __future__ import annotations
import json
import pytest
from sqlalchemy import select
from lembas.db.models import (
AUTHOR_MODEL,
DEFAULT_GROUP,
Chat,
Connection,
DataGroup,
Impression,
Model,
User,
)
from lembas.services import harness, personas, reports
from lembas.services import tools as tools_service
from lembas.services.crypto import encrypt
from lembas.services.library import documents, memories, notes, skills
HOSTED = "hosted"
TOOLS = {"tools": True}
@pytest.fixture
def owner(db, registered) -> User:
return db.scalars(select(User).order_by(User.created_at)).first()
@pytest.fixture
def chats(db, owner) -> tuple[Chat, Chat]:
"""One chat in the default group, one in the hosted group."""
db.add(DataGroup(id=HOSTED, name="Hosted"))
local = Connection(name="Local", base_url="http://127.0.0.1:1", api_key_encrypted=encrypt(""))
cloud = Connection(
name="Cloud",
base_url="http://127.0.0.1:2",
api_key_encrypted=encrypt(""),
data_group_id=HOSTED,
)
db.add_all([local, cloud])
db.flush()
db.add_all(
[
Model(connection_id=local.id, model_id="local-model", capabilities_json=TOOLS),
Model(connection_id=cloud.id, model_id="cloud-model", capabilities_json=TOOLS),
]
)
home = Chat(user_id=owner.id, model_id="local-model", connection_id=local.id,
data_group_id=DEFAULT_GROUP)
away = Chat(user_id=owner.id, model_id="cloud-model", connection_id=cloud.id,
data_group_id=HOSTED)
db.add_all([home, away])
db.commit()
return home, away
def _tools(*names):
return [tools_service.REGISTRY[name].schema for name in names]
def _context(db, owner, chat) -> tools_service.ToolContext:
return tools_service.context_for(db, owner, chat, tools=None)
async def _run(context, tool: str, **args):
return await tools_service.run_tool(context, tool, json.dumps(args))
# --- The harness: what a model is handed ------------------------------------------
def test_a_model_is_handed_only_its_own_groups_memories(db, owner, chats):
home, away = chats
memories.add(db, owner=owner, content="Home fact.", group=DEFAULT_GROUP)
memories.add(db, owner=owner, content="Hosted fact.", group=HOSTED)
at_home = harness.compose(db, owner, _tools("memory_add"), chat=home)
abroad = harness.compose(db, owner, _tools("memory_add"), chat=away)
assert "Home fact." in at_home and "Hosted fact." not in at_home
assert "Hosted fact." in abroad and "Home fact." not in abroad
def test_the_skill_index_is_one_groups(db, owner, chats):
home, away = chats
skills.create(db, owner=owner, name="home-skill", description="Home.", body="b")
skills.create(
db, owner=owner, name="away-skill", description="Away.", body="b", group=HOSTED
)
abroad = harness.compose(db, owner, _tools("skill_get"), chat=away)
assert "away-skill" in abroad
assert "home-skill" not in abroad
def test_a_personality_is_per_group(db, owner, chats):
home, away = chats
personas.write(
db,
model_key=personas.key_for("cloud-model", HOSTED),
owner=owner,
content="The hosted self.",
author=AUTHOR_MODEL,
)
abroad = harness.compose(db, owner, _tools("persona_write"), chat=away)
assert "The hosted self." in abroad
def test_a_base_in_another_group_is_not_named(db, owner, chats):
home, away = chats
base = documents.create_base(db, owner=owner, name="Home contracts")
away.knowledge_bases = [base]
db.commit()
abroad = harness.compose(db, owner, _tools("knowledge_search"), chat=away)
assert "Home contracts" not in abroad
# --- The tools: what a model can go and get ----------------------------------------
def test_the_tool_context_carries_the_chats_group(db, owner, chats):
home, away = chats
assert _context(db, owner, home).data_group == DEFAULT_GROUP
assert _context(db, owner, away).data_group == HOSTED
async def test_a_note_in_another_group_cannot_be_searched_or_fetched(db, owner, chats):
home, away = chats
secret = notes.create(db, owner=owner, title="Home only", body="mallorn", group=DEFAULT_GROUP)
context = _context(db, owner, away)
found = await _run(context, "notes_search", query="mallorn")
assert found.event["results"] == []
fetched = await _run(context, "notes_get", id=secret.id)
assert fetched.event["status"] == "error"
edited = await _run(context, "notes_edit", id=secret.id, body="gone")
assert edited.event["status"] == "error"
async def test_a_note_a_model_writes_lands_in_its_group(db, owner, chats):
home, away = chats
outcome = await _run(_context(db, owner, away), "notes_create", title="t", body="b")
assert outcome.event["status"] == "ok"
note = notes.get(db, outcome.event["results"][0]["id"], owner)
assert note.data_group_id == HOSTED
async def test_a_memory_is_recorded_in_the_group_and_forgotten_only_there(db, owner, chats):
home, away = chats
memories.add(db, owner=owner, content="Keep this at home.", group=DEFAULT_GROUP)
await _run(_context(db, owner, away), "memory_add", content="Hosted fact.")
assert [m.content for m in memories.all_for(db, owner, HOSTED)] == ["Hosted fact."]
await _run(_context(db, owner, away), "memory_forget", content="Keep this at home.")
assert [m.content for m in memories.all_for(db, owner, DEFAULT_GROUP)] == [
"Keep this at home."
]
# The same call from the memory's own group does forget it, so the refusal
# above is the group and not a mistyped argument.
await _run(_context(db, owner, home), "memory_forget", content="Keep this at home.")
db.expire_all()
assert memories.all_for(db, owner, DEFAULT_GROUP) == []
def test_the_same_fact_in_two_groups_is_two_memories(db, owner):
first = memories.add(db, owner=owner, content="Same.", group=DEFAULT_GROUP)
second = memories.add(db, owner=owner, content="Same.", group=HOSTED)
assert first.id != second.id
async def test_a_document_in_another_group_cannot_be_fetched_by_id(db, owner, chats):
home, away = chats
base = documents.create_base(db, owner=owner, name="Home")
document = documents.store_upload(
db, owner=owner, payload=b"The mallorn is golden.", filename="a.txt", base=base
)
context = _context(db, owner, away)
assert (await _run(context, "knowledge_search", query="mallorn")).event["results"] == []
assert (await _run(context, "knowledge_get", id=document.id)).event["status"] == "error"
async def test_a_report_in_another_group_cannot_be_read(db, owner, chats):
home, away = chats
report = reports.create(db, owner=owner, title="Home report", body="mallorn", unread=False)
context = _context(db, owner, away)
assert (await _run(context, "report_get", id=report.id)).event["status"] == "error"
async def test_a_skill_in_another_group_cannot_be_fetched(db, owner, chats):
home, away = chats
skills.create(db, owner=owner, name="home-skill", description="Home.", body="SECRET")
outcome = await _run(_context(db, owner, away), "skill_get", name="home-skill")
assert "SECRET" not in outcome.content
def test_a_skill_name_taken_in_another_group_says_so(db, owner):
skills.create(db, owner=owner, name="shared-name", description="d", body="b")
with pytest.raises(skills.SkillError, match="another data group"):
skills.create(
db, owner=owner, name="shared-name", description="d", body="b", group=HOSTED
)
async def test_an_impression_is_written_under_the_groups_key(db, owner, chats):
home, away = chats
await _run(_context(db, owner, away), "impression_write", content="Terse.")
row = db.scalar(select(Impression))
assert row.model_key == personas.key_for("cloud-model", HOSTED)
def test_each_group_gets_its_own_default_base(db, owner, chats):
home = documents.default_base(db, owner)
away = documents.default_base(db, owner, HOSTED)
assert home.id != away.id
assert away.data_group_id == HOSTED
assert home.name != away.name
+273
View File
@@ -0,0 +1,273 @@
"""The services that read a group's data, and the screens that arrange groups.
The embedder is sent the full text of everything it indexes and the reviewer is
sent every picture with its prompt, so a group can name its own of each and
falls back to the instance's when it names none. Then the three places groups
are arranged: the admin page, the person's own settings, and the library.
"""
from __future__ import annotations
import pytest
from sqlalchemy import select
from lembas.db.models import (
DEFAULT_GROUP,
Connection,
DataGroup,
Model,
Note,
User,
)
from lembas.services import data_groups, settings_store
from lembas.services import tools as tools_service
from lembas.services.crypto import encrypt
from lembas.services.images import tool as image_tool
from lembas.services.library import indexing, notes
HOSTED = "hosted"
@pytest.fixture
def owner(db, registered) -> User:
return db.scalars(select(User).order_by(User.created_at)).first()
@pytest.fixture
def setup(db, owner):
db.add(DataGroup(id=HOSTED, name="Hosted"))
local = Connection(name="Local", base_url="http://127.0.0.1:1", api_key_encrypted=encrypt(""))
cloud = Connection(
name="Cloud",
base_url="http://127.0.0.1:2",
api_key_encrypted=encrypt(""),
data_group_id=HOSTED,
)
db.add_all([local, cloud])
db.flush()
db.add_all(
[
Model(connection_id=local.id, model_id="local-embed",
capabilities_json={"embeddings": True}),
Model(connection_id=cloud.id, model_id="cloud-embed",
capabilities_json={"embeddings": True}),
Model(connection_id=local.id, model_id="local-eye",
capabilities_json={"vision": True}),
Model(connection_id=cloud.id, model_id="cloud-eye",
capabilities_json={"vision": True}),
]
)
db.commit()
return local, cloud
def _plain_user(db) -> User:
user = User(email="sam@shire.test", name="Sam", password_hash="x", role="user")
db.add(user)
db.commit()
return user
# --- The embedder -----------------------------------------------------------------
def test_a_group_without_its_own_embedder_uses_the_instances(db, setup):
settings_store.update(db, {"embedding_model_id": "local-embed"}, key=settings_store.EXTRACTION)
assert indexing.embedder(db, HOSTED).model_id == "local-embed"
def test_a_group_with_its_own_embedder_uses_that(db, setup):
settings_store.update(db, {"embedding_model_id": "local-embed"}, key=settings_store.EXTRACTION)
group = data_groups.get(db, HOSTED)
group.embedding_model_id = "cloud-embed"
db.commit()
assert indexing.embedder(db, HOSTED).model_id == "cloud-embed"
assert indexing.embedder(db, DEFAULT_GROUP).model_id == "local-embed"
def test_a_record_is_indexed_by_its_own_groups_embedder(db, owner, setup):
group = data_groups.get(db, HOSTED)
group.embedding_model_id = "cloud-embed"
db.commit()
note = notes.create(db, owner=owner, title="t", body="b", group=HOSTED)
assert indexing.group_of_row(db, note) == HOSTED
assert indexing.embedder(db, indexing.group_of_row(db, note)).model_id == "cloud-embed"
def test_any_configured_counts_a_groups_own_embedder(db, setup):
assert indexing.any_configured(db) is False
group = data_groups.get(db, HOSTED)
group.embedding_model_id = "cloud-embed"
db.commit()
assert indexing.any_configured(db) is True
# --- The reviewer -------------------------------------------------------------------
def test_the_reviewer_is_the_groups_own_when_it_names_one(db, setup):
group = data_groups.get(db, HOSTED)
group.review_model_id = "cloud-eye"
db.commit()
config = {"review_enabled": True, "review_model_id": "local-eye"}
hosted = tools_service.ToolContext(owner_id="x", image_config=config, data_group=HOSTED)
home = tools_service.ToolContext(owner_id="x", image_config=config)
assert image_tool._reviewer(hosted)[1] == "cloud-eye"
assert image_tool._reviewer(home)[1] == "local-eye"
# --- The admin page ---------------------------------------------------------------------
def test_the_admin_page_lists_groups_and_creates_one(client, db, setup):
assert "Hosted" in client.get("/admin/data-groups").text
response = client.post("/admin/data-groups", data={"name": "Work"}, follow_redirects=False)
assert response.status_code == 303
assert db.scalar(select(DataGroup).where(DataGroup.name == "Work")) is not None
def test_putting_a_connection_into_a_group_and_taking_it_out(client, db, setup):
local, cloud = setup
client.post(
f"/admin/data-groups/{HOSTED}",
data={"name": "Hosted", "connections_sent": "1", "connection_ids": [local.id]},
)
db.expire_all()
assert db.get(Connection, local.id).data_group_id == HOSTED
# Unticked: back to the default group, never to "no group".
assert db.get(Connection, cloud.id).data_group_id == DEFAULT_GROUP
def test_the_detail_page_flags_a_service_in_another_group(client, db, setup):
settings_store.update(db, {"embedding_model_id": "local-embed"}, key=settings_store.EXTRACTION)
page = client.get(f"/admin/data-groups/{HOSTED}").text
assert "its connection is in Default" in page
def test_deleting_a_group_in_use_is_refused_with_the_reason(client, db, setup):
response = client.post(f"/admin/data-groups/{HOSTED}/delete", follow_redirects=True)
assert "connections" in response.text
assert data_groups.get(db, HOSTED) is not None
def test_the_admin_page_is_for_administrators(client, db, setup, owner):
owner.role = "user"
db.commit()
assert client.get("/admin/data-groups").status_code in (303, 403, 404)
def test_the_connection_form_sets_the_group(client, db, setup):
local, _ = setup
client.post(
f"/admin/connections/{local.id}",
data={"name": "Local", "base_url": local.base_url, "data_group_id": HOSTED},
)
db.expire_all()
assert db.get(Connection, local.id).data_group_id == HOSTED
# --- The person's own settings -------------------------------------------------------------
def test_remapping_for_oneself_needs_the_permission(client, db, setup, owner):
local, _ = setup
owner.role = "user"
db.commit()
client.post("/api/preferences/data-groups", data={f"group__{local.id}": HOSTED})
db.expire_all()
assert data_groups.personal_map(db.get(User, owner.id)) == {}
settings_store.update(db, {"default_permissions": {data_groups.PERMISSION: True}})
client.post("/api/preferences/data-groups", data={f"group__{local.id}": HOSTED})
db.expire_all()
user = db.get(User, owner.id)
assert data_groups.personal_map(user) == {local.id: HOSTED}
assert data_groups.for_connection(db, user, local.id) == HOSTED
def test_a_personal_group_is_made_and_seen_only_by_its_owner(client, db, setup, owner):
client.post("/api/preferences/data-groups/new", data={"name": "Private"})
group = db.scalar(select(DataGroup).where(DataGroup.name == "Private"))
assert group.owner_id == owner.id
other = _plain_user(db)
assert group.id not in {g.id for g in data_groups.usable(db, other)}
def test_the_data_tab_appears_once_there_is_a_choice(client, db, setup):
page = client.get("/settings").text
assert 'id="tab-data"' in page
assert "Which group each connection reads" in page
# --- The library --------------------------------------------------------------------------------
def test_a_note_is_created_in_the_chosen_group(client, db, setup):
client.post("/api/library/notes", data={"title": "n", "body": "b", "data_group_id": HOSTED})
assert db.scalar(select(Note)).data_group_id == HOSTED
def test_a_group_the_person_may_not_use_falls_back_to_the_default(client, db, setup, owner):
other = _plain_user(db)
db.add(DataGroup(id="theirs", name="Theirs", owner_id=other.id))
db.commit()
client.post("/api/library/notes", data={"title": "n", "body": "b", "data_group_id": "theirs"})
assert db.scalar(select(Note)).data_group_id == DEFAULT_GROUP
def test_moving_a_note_needs_the_permission(client, db, setup, owner):
note = notes.create(db, owner=owner, title="n", body="b")
owner.role = "user"
db.commit()
client.post(f"/api/library/notes/{note.id}", data={"title": "n", "body": "b",
"data_group_id": HOSTED})
db.expire_all()
assert db.get(Note, note.id).data_group_id == DEFAULT_GROUP
settings_store.update(db, {"default_permissions": {data_groups.PERMISSION: True,
"library.use": True}})
client.post(f"/api/library/notes/{note.id}", data={"title": "n", "body": "b",
"data_group_id": HOSTED})
db.expire_all()
assert db.get(Note, note.id).data_group_id == HOSTED
def test_the_notes_list_filters_by_group(client, db, setup, owner):
notes.create(db, owner=owner, title="Home note", body="b")
notes.create(db, owner=owner, title="Hosted note", body="b", group=HOSTED)
page = client.get(f"/library/notes?group={HOSTED}").text
assert "Hosted note" in page and "Home note" not in page
def test_a_single_group_instance_shows_nothing_about_groups(client, db, owner):
"""The ordinary instance: no chip, no select, no tab -- for anybody who could
not make a personal group either. An administrator can, so the tab is theirs."""
notes.create(db, owner=owner, title="n", body="b")
owner.role = "user"
db.commit()
assert 'name="data_group_id"' not in client.get("/library/notes/new").text
assert 'id="tab-data"' not in client.get("/settings").text
# --- Two embedders of the same width -----------------------------------------------------
def test_a_query_skips_chunks_another_model_of_the_same_width_made(db, owner):
"""Width cannot tell two 1024-wide models apart, and with an embedder per
group two of them on one instance is ordinary. The query says which model
made it, and only that model's chunks are scored."""
from lembas.db.models import CHUNK_NOTE, Chunk
from lembas.services.library import chunks as chunk_service
from lembas.services.library import retrieval
ours = notes.create(db, owner=owner, title="Ours", body="x")
theirs = notes.create(db, owner=owner, title="Theirs", body="y")
for note, model_id, vector in ((ours, "embed-a", [0.6, 0.8]), (theirs, "embed-b", [1.0, 0.0])):
db.add(
Chunk(
owner_id=owner.id,
resource_type=CHUNK_NOTE,
resource_id=note.id,
ordinal=0,
text="t",
vector=chunk_service.pack(vector),
dims=2,
model_id=model_id,
)
)
db.commit()
query = retrieval.QueryVector([1.0, 0.0])
query.model_id = "embed-a"
assert [hit.id for hit in retrieval.semantic_ids(db, CHUNK_NOTE, query)] == [ours.id]
# A plain list keeps the old width-only behaviour, and the closer vector wins.
plain = retrieval.semantic_ids(db, CHUNK_NOTE, [1.0, 0.0])
assert [hit.id for hit in plain][0] == theirs.id
+293
View File
@@ -0,0 +1,293 @@
"""Data groups: how one is resolved, the startup sweep, deleting one, and upgrading.
What a group *isolates* is `test_data_group_isolation.py`; how a chat is pinned to
one is `test_data_group_chat_pin.py`. This file is the machinery underneath both:
the resolution order, which lives in one function and must keep living there, and
the upgrade from a database that has never heard of groups.
"""
from __future__ import annotations
import pytest
from sqlalchemy import inspect, select, text
from lembas.db.migrations import sync_schema
from lembas.db.models import (
DEFAULT_GROUP,
Chat,
Connection,
DataGroup,
Memory,
Model,
Note,
Persona,
User,
)
from lembas.db.session import get_engine
from lembas.services import data_groups, personas, settings_store
from lembas.services.crypto import encrypt
@pytest.fixture
def owner(db, registered) -> User:
return db.scalars(select(User).order_by(User.created_at)).first()
@pytest.fixture
def reader(db, registered) -> User:
"""A second account, not an administrator, so permissions actually apply."""
user = User(email="sam@shire.test", name="Sam", password_hash="x", role="user")
db.add(user)
db.commit()
return user
@pytest.fixture
def two(db) -> tuple[Connection, Connection]:
"""A local connection in the default group and a hosted one in its own."""
db.add(DataGroup(id="hosted", name="Hosted"))
local = Connection(name="Local", base_url="http://127.0.0.1:1", api_key_encrypted=encrypt(""))
cloud = Connection(
name="Cloud",
base_url="http://127.0.0.1:2",
api_key_encrypted=encrypt(""),
data_group_id="hosted",
)
db.add_all([local, cloud])
db.flush()
db.add_all(
[
Model(connection_id=local.id, model_id="local-model", position=0),
Model(connection_id=cloud.id, model_id="cloud-model", position=1),
]
)
db.commit()
return local, cloud
def _grant_manage(db, user: User) -> None:
settings_store.update(db, {"default_permissions": {data_groups.PERMISSION: True}})
# --- Resolution --------------------------------------------------------------------
def test_a_connection_with_no_group_is_in_the_default_one(db, two, reader):
local, _ = two
assert data_groups.for_connection(db, reader, local.id) == DEFAULT_GROUP
def test_the_administrators_choice_applies_to_everybody(db, two, reader):
_, cloud = two
assert data_groups.for_connection(db, reader, cloud.id) == "hosted"
assert data_groups.for_connection(db, None, cloud.id) == "hosted"
def test_a_personal_mapping_needs_the_permission(db, two, reader):
"""Stored and ignored without `data.manage`: taking the permission away puts a
person back on the instance's arrangement without anybody clearing anything."""
local, _ = two
reader.settings_json = {data_groups.SETTING_KEY: {local.id: "hosted"}}
db.commit()
assert data_groups.for_connection(db, reader, local.id) == DEFAULT_GROUP
_grant_manage(db, reader)
assert data_groups.for_connection(db, reader, local.id) == "hosted"
def test_a_mapping_to_somebody_elses_personal_group_is_ignored(db, two, reader, owner):
local, _ = two
db.add(DataGroup(id="theirs", name="Theirs", owner_id=owner.id))
reader.settings_json = {data_groups.SETTING_KEY: {local.id: "theirs"}}
db.commit()
_grant_manage(db, reader)
assert data_groups.for_connection(db, reader, local.id) == DEFAULT_GROUP
def test_a_connection_naming_a_deleted_group_falls_back_to_the_default(db, two, reader):
_, cloud = two
cloud.data_group_id = "gone"
db.commit()
assert data_groups.for_connection(db, reader, cloud.id) == DEFAULT_GROUP
def test_a_pair_without_a_connection_resolves_through_the_model(db, two, reader):
assert data_groups.for_pair(db, reader, "cloud-model") == "hosted"
assert data_groups.for_pair(db, reader, "local-model") == DEFAULT_GROUP
def test_a_chat_keeps_the_group_it_was_stamped_with(db, two, reader):
"""Derived once, then read. A model moved afterwards does not carry the chat."""
_, cloud = two
chat = Chat(user_id=reader.id, model_id="cloud-model", connection_id=cloud.id)
db.add(chat)
db.commit()
assert data_groups.for_chat(db, chat) == "hosted"
db.commit()
cloud.data_group_id = DEFAULT_GROUP
db.commit()
assert data_groups.for_chat(db, chat) == "hosted"
def test_usable_groups_are_the_instances_and_ones_own(db, two, reader, owner):
db.add_all(
[
DataGroup(id="mine", name="Mine", owner_id=reader.id),
DataGroup(id="not-mine", name="Not mine", owner_id=owner.id),
]
)
db.commit()
ids = {group.id for group in data_groups.usable(db, reader)}
assert ids == {DEFAULT_GROUP, "hosted", "mine"}
def test_one_group_means_there_is_nothing_to_choose(db, reader):
assert data_groups.several(db, reader) is False
db.add(DataGroup(id="second", name="Second"))
db.commit()
assert data_groups.several(db, reader) is True
# --- Deleting ------------------------------------------------------------------------
def test_the_default_group_cannot_be_deleted(db):
with pytest.raises(ValueError, match="default group"):
data_groups.delete(db, data_groups.ensure_default(db))
def test_a_group_with_records_in_it_cannot_be_deleted(db, reader):
group = DataGroup(id="busy", name="Busy")
db.add(group)
db.add(Note(owner_id=reader.id, title="n", body="b", data_group_id="busy"))
db.commit()
with pytest.raises(ValueError, match="1 notes"):
data_groups.delete(db, group)
def test_a_group_a_connection_is_in_cannot_be_deleted(db, two):
with pytest.raises(ValueError, match="connections"):
data_groups.delete(db, data_groups.get(db, "hosted"))
def test_deleting_an_empty_group_clears_every_mapping_to_it(db, reader):
group = DataGroup(id="empty", name="Empty")
db.add(group)
reader.settings_json = {data_groups.SETTING_KEY: {"some-connection": "empty"}}
db.commit()
data_groups.delete(db, group)
db.refresh(reader)
assert data_groups.personal_map(reader) == {}
assert data_groups.get(db, "empty") is None
# --- The sweep -------------------------------------------------------------------------
def test_the_sweep_puts_old_rows_in_the_default_group(db, reader):
db.add(Memory(owner_id=reader.id, content="old"))
db.commit()
assert db.scalar(select(Memory)).data_group_id is None
data_groups.sweep_unassigned(db)
assert db.scalar(select(Memory)).data_group_id == DEFAULT_GROUP
def test_the_sweep_gives_a_chat_its_models_group(db, two, reader):
"""Not simply the default: a chat some path created without stamping one
belongs where its model is, or its next turn is refused as a moved chat."""
chat = Chat(user_id=reader.id, model_id="cloud-model")
db.add(chat)
db.commit()
data_groups.sweep_unassigned(db)
db.refresh(chat)
assert chat.data_group_id == "hosted"
def test_a_row_the_sweep_has_not_reached_still_counts_as_default(db, reader):
db.add(Note(owner_id=reader.id, title="old", body=""))
db.commit()
found = db.scalars(select(Note).where(data_groups.condition(Note, DEFAULT_GROUP))).all()
assert [note.title for note in found] == ["old"]
# --- Personalities: namespaced, because the constraint cannot change ------------------
def test_the_default_group_keeps_the_bare_model_id():
assert personas.key_for("gpt-oss", DEFAULT_GROUP) == "gpt-oss"
assert personas.key_for("gpt-oss", None) == "gpt-oss"
def test_another_group_gets_its_own_key_and_splits_back():
key = personas.key_for("gpt-oss", "hosted")
assert key != "gpt-oss"
assert personas.split_key(key) == ("gpt-oss", "hosted")
assert personas.split_key("gpt-oss") == ("gpt-oss", DEFAULT_GROUP)
def test_a_group_falls_back_to_the_administrators_default(db, reader):
"""The admin default is keyed bare and reaches every group, until the model has
written one of its own with this person in that group."""
db.add(Persona(model_key="gpt-oss", owner_id=None, content="the default"))
db.commit()
key = personas.key_for("gpt-oss", "hosted")
assert personas.block(db, key, reader) == "the default"
personas.write(db, model_key=key, owner=reader, content="hosted self", author="model")
assert personas.block(db, key, reader) == "hosted self"
assert personas.block(db, "gpt-oss", reader) == "the default"
# --- Upgrading from 1.9.1 -----------------------------------------------------------------
# What 1.10.0 added: one table, and one column on each of these. Taken from the
# models rather than invented, and `test_the_recorded_shape_is_still_real`
# below is what stops the list rotting.
NEW_TABLE = "data_groups"
GROUPED_TABLES = (
"chats",
"connections",
"knowledge_bases",
"memories",
"notes",
"reports",
"schedules",
"skills",
)
def _rollback_to_1_9_1(engine) -> None:
with engine.begin() as connection:
connection.execute(text(f"DROP TABLE IF EXISTS {NEW_TABLE}"))
for table in GROUPED_TABLES:
connection.execute(text(f"ALTER TABLE {table} DROP COLUMN data_group_id"))
def test_the_recorded_shape_is_still_real():
from lembas.db.base import Base
grouped = {
table.name for table in Base.metadata.sorted_tables if "data_group_id" in table.columns
}
assert grouped == set(GROUPED_TABLES)
def test_a_1_9_1_database_with_data_upgrades(db, reader):
engine = get_engine()
db.add(Memory(owner_id=reader.id, content="kept"))
db.add(Persona(model_key="gpt-oss", owner_id=reader.id, content="mine"))
db.commit()
db.close()
_rollback_to_1_9_1(engine)
assert "data_group_id" not in {c["name"] for c in inspect(engine).get_columns("memories")}
changes = sync_schema(engine)
assert any(NEW_TABLE in change for change in changes)
for table in GROUPED_TABLES:
assert "data_group_id" in {c["name"] for c in inspect(engine).get_columns(table)}
# A nullable column arrives empty; the sweep is what files it.
memory = db.scalar(select(Memory))
assert memory.data_group_id is None
data_groups.sweep_unassigned(db)
db.refresh(memory)
assert memory.data_group_id == DEFAULT_GROUP
# The person's personality was keyed bare, and the default group still reads it.
user = db.get(User, reader.id)
assert personas.block(db, personas.key_for("gpt-oss", DEFAULT_GROUP), user) == "mine"
assert data_groups.get(db, DEFAULT_GROUP) is not None