chore: anchor fresh-start history to upstream
This commit is contained in:
@@ -0,0 +1,283 @@
|
||||
# 项目内历史对话召回与语义检索完整方案 V4
|
||||
|
||||
> **For Hermes:** Use subagent-driven-development when implementation is authorized. This document is a plan, not authorization to modify business code, migrate user data, commit, push, deploy, or send historical conversations to an external service.
|
||||
|
||||
**Goal:** 会话具有独立的项目归属属性;用户移动会话只修改归属;模型能够按需通过关键词与语义混合检索找回同项目历史原文。
|
||||
|
||||
**Architecture:** 分两阶段交付:A 完成 V3 的 project_id、创建/移动/列表与项目内关键词召回;B 在其上建立可重建的原文片段向量索引,并扩展现有 session_search 为项目内混合检索。原始消息和会话项目归属是事实来源,向量是派生数据。工作目录独立,不生成项目或会话摘要,不自动注入历史上下文。
|
||||
|
||||
**Tech Stack:** 现有 Python、SQLite/FTS5、SessionDB、TUI Gateway、Desktop TypeScript。语义层首选同项目片段的精确相似度计算;embedding 模型及必要依赖由小规模实测后确定,不预先引入独立向量服务。
|
||||
|
||||
**状态与版本:** 完整独立文件,包含 V3 基础能力和语义层,不需要同时阅读 V3 才能实施。替代此前版本作为后续设计依据,原文件保留。本文件是实施方案,不是功能已经完成或测试已经通过的证明。
|
||||
|
||||
---
|
||||
|
||||
## 1. 范围与交付目标
|
||||
|
||||
### 1.1 包含
|
||||
|
||||
1. sessions 保存可空 project_id,和 cwd/git_repo_root 分离。
|
||||
2. 新建、草稿落库、恢复、压缩延续、独立分支正确保存或继承归属。
|
||||
3. 显式移动会话改变整条逻辑对话的项目归属,不移动消息、不改变工作目录。
|
||||
4. 项目会话列表和 session_search 按持久化归属工作。
|
||||
5. 项目内关键词、语义及混合检索,返回原文和消息锚点。
|
||||
6. 原文片段向量增量更新、删除/修改处理、中断恢复、模型切换后的重建。
|
||||
7. 索引未配置、未完成或不可用时,项目内关键词检索仍可用。
|
||||
|
||||
### 1.2 不包含
|
||||
|
||||
- 项目摘要、会话摘要、结论蒸馏、自动每回合召回注入。
|
||||
- 新的核心工具、独立记忆管理平台、MemoryProvider 槽位占用。
|
||||
- 项目级权限撤销协议、项目版本平台、消息级项目授权、操作系统沙箱。
|
||||
- 全局 MEMORY/USER、技能、其他插件的项目隔离改造。
|
||||
- 旧会话自动按目录归类或复杂批量迁移审批系统。
|
||||
- 独立向量数据库服务、常驻队列平台或未验证必要性的近似索引。
|
||||
|
||||
## 2. 业务合同
|
||||
|
||||
| 操作 | 历史所属项目 | 工作目录 |
|
||||
|---|---|---|
|
||||
| 在项目中创建对话 | 保存选定 Project.id | 使用既有默认或用户选择 |
|
||||
| 未选择项目创建 | NULL | 按现有逻辑 |
|
||||
| 恢复对话 | 读取已存 project_id | 读取已存目录 |
|
||||
| 更改目录 | 不变 | 更新 cwd 与 git 元数据 |
|
||||
| 移动对话到项目 | 更新整条逻辑对话的 project_id | 不变 |
|
||||
| 移出项目 | 更新为 NULL | 不变 |
|
||||
| 修改项目名称或 folders | 已有归属不变 | 不自动修改 |
|
||||
| 搜索历史 | 使用当前会话持久归属 | 不参与过滤 |
|
||||
|
||||
项目归属是业务分类。移动前已经开始的检索可以完成,移动完成后发起的新查询按新归属执行;不取消在途工具、不要求先停止 Agent。已返回到其他会话的内容不会被追溯撤回。
|
||||
|
||||
无归属会话不是共同项目:不搜索其他 NULL 会话,但保留当前逻辑对话离开上下文后的压缩历史恢复能力。
|
||||
|
||||
同 profile 的显式项目 ID 才是归属值。相同路径不代表跨 profile 同项目,工具不通过任意 profile 参数或链接自动扩大检索范围。全局历史管理 UI 不因此变成权限系统。
|
||||
|
||||
## 3. 现有代码与改动落点
|
||||
|
||||
以下是已定位的现有文件;新字段、模块和测试均是拟新增设计项,不冒称已有 API。
|
||||
|
||||
| 职责 | 现有位置 | 预期改动 |
|
||||
|---|---|---|
|
||||
| 表结构及升级 | hermes_state_common.py、hermes_state_schema.py | project_id、索引、增量迁移 |
|
||||
| 会话创建/更新/浏览 | hermes_state_sessions.py | 保存归属、移动、继承、项目过滤 |
|
||||
| 搜索路径 | hermes_state_search.py | FTS/CJK/LIKE/回退的项目过滤 |
|
||||
| 工具 | tools/session_search_tool.py | 项目内所有形状、混合检索接线 |
|
||||
| Agent 实际入口 | agent/inline_tool_executors.py | 当前会话上下文及检索参数透传 |
|
||||
| 项目及目录 RPC | tui_gateway/methods_session.py、methods_projects.py、methods_config.py | 创建和移动归属、区分目录操作 |
|
||||
| 项目列表 | tui_gateway/project_tree.py | 会话成员按 project_id 分组 |
|
||||
| 桌面状态 | apps/desktop/src/store/projects.ts | 创建/移动请求、刷新与错误状态 |
|
||||
| 项目存储 | hermes_cli/projects_db.py | 复用真实 Project 对象和校验路径 |
|
||||
| 配置 | hermes_cli/config_defaults.py 及实际设置/加载路径 | embedding 配置及启用说明 |
|
||||
|
||||
会话插入/upsert、压缩延续、branch/new/reset 的具体调用者在实施前追踪后纳入。不得仅修改 facade 而漏掉实际 sibling 绑定;新逻辑过大时使用专题 sibling,不做无关拆分。
|
||||
|
||||
## 4. 阶段 A:项目归属与关键词召回
|
||||
|
||||
### 4.1 数据模型
|
||||
|
||||
sessions 增加 `project_id TEXT NULL`。在现有迁移机制中追加字段和必要索引,旧记录保留 NULL,原始消息不变。索引由实际 WHERE/排序和 EXPLAIN 决定,不预设多个冗余索引。
|
||||
|
||||
projects.db 与 state.db 是现有分库,不能直接声称有跨库 SQLite 外键。创建/移动入口校验目标项目存在且属于当前 profile;NULL 表示移出项目。归档保留归属;硬删除仍有绑定会话的项目时拒绝并提示先移动,避免孤儿归属。
|
||||
|
||||
upsert 必须区分“未提供 project_id”和“显式传入 NULL”,恢复或重复保存不得清空已有归属。
|
||||
|
||||
### 4.2 逻辑对话与移动
|
||||
|
||||
正常移动只更新属性。若用户可见的一条对话由多个历史压缩 session 记录组成,在同一个现有 SessionDB 写事务内更新其 project_id:
|
||||
- 原地压缩无需移动消息。
|
||||
- 压缩延续继承同一归属,和逻辑对话一起移动。
|
||||
- 独立 branch 创建时继承项目,之后独立移动。
|
||||
- reset/new 的新对话初始化项目,但不跟随旧对话后续移动。
|
||||
- 委派记录维持现有隐藏规则。
|
||||
|
||||
不能使用沿所有 parent 边走根的宽泛 helper 代替逻辑对话判定。现有代码区分 branch/reset/压缩,需复用其准确语义并测试。
|
||||
|
||||
移动与压缩延续创建采用现有数据库串行写事务;延续创建在写事务中读取最新来源归属,不从长期缓存复制旧 project_id。这是数据一致性要求,不新增项目锁服务或移动禁用机制。
|
||||
|
||||
### 4.3 旧会话与自动仓库入口
|
||||
|
||||
旧会话为 NULL,用户直接通过“移动到项目”归类即可。不做目录推断自动回填。
|
||||
|
||||
自动发现的 repo 路径不是 Project.id。该入口用于创建项目内对话或成为移动目标时,复用已有项目创建/查重能力取得稳定 ID。仅浏览仓库不必批量创建项目;不要把路径写进 project_id。
|
||||
|
||||
### 4.4 前后端接线
|
||||
|
||||
项目内新建与草稿第一次持久化都传入归属。恢复从 DB 读取,不按全局当前项目或 cwd 重算。
|
||||
|
||||
提供一个窄职责的设置会话项目 RPC,方法名在核实注册表后确定,入参使用已有会话标识和可空 project_id,返回实际写入结果。前端失败时不改成成功状态;成功后刷新源项目、目标项目和会话信息。
|
||||
|
||||
`session.cwd.set` 和 `session.workspace.move` 保留目录操作语义,不隐式改归属;“移动到项目”不能继续只调用 workspace.move。项目成员按 project_id 展示,repo/lane 和目录信息尽量沿用现有布局,不重写整个侧栏。
|
||||
|
||||
### 4.5 项目内检索
|
||||
|
||||
Agent 每次调用从可信当前 session/profile 读取当前归属,不长期缓存移动前的 ID。默认当前项目,无须添加只有 project 一个取值的 scope 参数。
|
||||
|
||||
查询时在 sessions 的 JOIN/WHERE 中施加 project_id,发生在候选 LIMIT 之前。不做全库先截再筛、不建立全量成员表、不将历史集合截为 900/2000 条。
|
||||
|
||||
关键词、标题、browse、read、scroll、首尾片段、窗口、lineage 展开和 profile fallback 均遵守项目范围;内部 SessionDB 可保留未过滤查询供其他可信调用者使用,但 Agent 不能把 NULL/读取错误当成不过滤。
|
||||
|
||||
保留既有角色、隐藏来源、压缩/撤销可见性和 bounded browse 行为。NULL 情况只恢复自身可靠压缩延续历史,不把 branch/reset/delegation 全树当作自身。
|
||||
|
||||
原文 session_id 与 message_id 必须配对,保留链接。不改系统提示词或已发送历史消息,只用工具结果进入上下文。
|
||||
|
||||
## 5. 阶段 B:语义与混合检索
|
||||
|
||||
### 5.1 技术路线与前置小实验
|
||||
|
||||
首版选用原文片段 embedding + 项目内精确余弦相似度 + 排名融合。先验证精确扫描的正确性、中文效果和实际规模耗时;不因“语义检索”直接引入远程向量服务或近似索引。
|
||||
|
||||
向量存储优先在现有 profile 的 state.db 中新增派生表,便于直接 JOIN sessions/messages 实现前置过滤。大 BLOB 对现有备份、数据库大小和读写锁的影响必须实测。若该选择不能通过规模验证,更新语义层存储选择,不绕过项目前置过滤,不阻塞阶段 A 的独立交付。
|
||||
|
||||
相似度可使用已验证可用的数值实现;仓库某些 optional extra 中出现 numpy 不表示基础安装已有该依赖。新增依赖必须明确声明、设置版本上界并更新锁文件。不要先写出项目里不存在的 embedding client/import。
|
||||
|
||||
### 5.2 模型配置与数据发送
|
||||
|
||||
正式回填前确定:embedding provider、模型及版本、向量维度、输入上限、批大小、超时,以及本地或远程运行方式。通过现有配置与 setup 路径设置,不增加非秘密 HERMES_* 环境变量,不读取或打印密钥。
|
||||
|
||||
建议拟用独立 `session_search.semantic` 配置节承载 enabled、模型路由与索引预算;最终字段在核实现有配置模式后确定。语义功能默认未配置/关闭时阶段 A 正常工作。
|
||||
|
||||
本方案保存不代表用户授权把历史发送给第三方。先使用合成测试对话验证候选模型;远程历史回填需确认服务方和项目范围,界面/设置说明后续新消息也会送该 embedding 服务。没有可用路由时不安装大型本地模型或猜用聊天模型代替。
|
||||
|
||||
### 5.3 原文切片与派生数据
|
||||
|
||||
按单条原始消息稳定切片,长消息分段并有有限重叠。保持消息 ID 和字符范围;命中后利用既有窗口读取相邻对话,不先总结再 embedding,也不把完整长会话当一个向量。
|
||||
|
||||
片段大小依据实际模型输入上限和 tokenizer 确定,明确 token/字符单位,不能混用。默认索引 user/assistant 文本;工具正文只有在用户启用相应索引范围后处理,其余显式 tool-role 查询仍走现有关键词路径并标明语义覆盖不足。索引未解析附件二进制,不把 base64 送入 embedding。
|
||||
|
||||
拟新增派生表使用清晰字段:session_id、message_id、片段起止位置、原文内容标识、切片策略版本、embedding 模型指纹/维度、向量、处理状态和必要时间戳。具体表名在实现时确定,原文仍由 messages 读取,不保存另一份权威消息。
|
||||
|
||||
模型指纹包含端点/提供方、模型身份、维度、归一化与切片版本等影响兼容性的非秘密信息,不能含 token。不同指纹的查询向量和索引向量不可混用。
|
||||
|
||||
不保存用于授权的 project_id 副本。向量通过 message/session 关联当前 sessions.project_id。片段标识稳定,唯一键保证重试不产生重复有效片段。
|
||||
|
||||
### 5.4 增量索引与中断恢复
|
||||
|
||||
初次索引按用户选择的项目分批执行;之后对已启用范围的新建/修改消息增量处理。先定位已有消息写入、更新、删除和可见性修改路径,复用实际持久化边界,不新造通用事件平台。
|
||||
|
||||
最低实现为派生行状态和内容标识:pending、ready、failed 可恢复。正文变化使旧片段失效并重建;仅记录最大 message_id 不能识别旧消息编辑。写入端不要做 embedding 网络调用。
|
||||
|
||||
使用嵌入宿主的有界后台批处理,明确启动、唤醒、关闭和每批预算;不建设独立常驻服务。进程退出后从持久化 pending/failed 恢复,不依赖 Agent 子进程存活。若多个进程同时执行,用现有 SQLite 事务做最小任务领取/幂等写入,不新增调度平台。
|
||||
|
||||
开始 embedding 前读取原文标识,结果落库前再检查原文仍存在且未变化;旧网络结果不能覆盖已编辑或已删除消息。网络期间不持有长写事务。首次回填可重试、可中断、可继续;失败记录原因,不无限重试阻塞用户回复。
|
||||
|
||||
查询端同时检查原消息当前可见性和内容标识,尚未物理清理的失效向量不能贡献结果。删除/撤销立即通过原文 JOIN 和可见性条件排除,后台清理派生行;压缩归档按既有历史规则保留,不能一律排除 active=0。分支副本与原消息各自保留正确来源。
|
||||
|
||||
移动会话只修改 project_id,不重算已有向量。若目标项目已启用语义但该会话尚未索引,正常增量批处理补齐缺片段;若目标未启用,不因为历史已有向量就绕过配置调用远程模型。
|
||||
|
||||
切换模型后逐步生成新指纹索引,旧指纹不可参与新模型检索。新索引部分覆盖时可返回部分语义结果并明确状态,关键词保持覆盖;按项目过滤后不会为了补足结果搜索其他项目。
|
||||
|
||||
### 5.5 查询与融合
|
||||
|
||||
仅 query discovery 启用语义;browse/read/scroll 继续直接读取,不产生无意义的 embedding 调用。
|
||||
|
||||
查询顺序:
|
||||
1. 从 DB 解析当前会话项目及现有角色/可见性限制。
|
||||
2. 在该范围执行既有关键词检索。
|
||||
3. 配置可用时生成查询向量;SQL 先 JOIN 当前 sessions/messages,过滤 project_id、有效原文、模型指纹和角色,再计算该集合的向量相似度。
|
||||
4. 精确扫描可分批读取整个允许集合,用固定大小候选堆保存最相关结果,不先按时间 LIMIT 一批向量,也不在全库 Top-K 之后才过滤项目。
|
||||
5. 关键词与语义按匹配消息锚点归并,以 Reciprocal Rank Fusion(基于排名而非原始分数相加)生成混合候选,再执行既有会话去重与结果展开。
|
||||
6. 对命中消息读取原文窗口,返回实际 session_id/message_id/link。窗口仍遵守阶段 A 的项目和可见性规则。
|
||||
|
||||
建议给现有 session_search 增加 `search_mode=hybrid|keyword|semantic`,默认 hybrid;函数、schema、registry 和 inline 映射同步。模式仅控制算法,不改变项目范围。keyword 便于精确符号检索和对照;semantic 用于用户明确的相似问题追溯及效果测试。
|
||||
|
||||
hybrid 在语义不可用时返回关键词结果并注明实际模式与原因。显式 semantic 失败应报告不可用,不冒称结果来自语义;用户或模型可改用 keyword。全程共用调用截止时间,不串行无限等待 embedding。
|
||||
|
||||
现有 sort=newest/oldest 保留为项目内结果的时间偏好,在两路候选融合后使用一致的已文档化顺序;不是承诺搜遍所有命中。确切排名参数通过固定测试集验证后在代码/配置中明确,不混用可变随机参数。
|
||||
|
||||
### 5.6 覆盖与降级状态
|
||||
|
||||
工具结果最少携带 requested_mode、used_mode、semantic_status(disabled/ready/partial/unavailable)、必要的截断/超时提示和匹配来源(keyword/semantic/both)。不要回传向量 BLOB。
|
||||
|
||||
“零命中”与“语义未索引完整”分开。已处理到最大 ID 不等于索引覆盖完整;覆盖统计依据已选范围的有效消息/片段状态,无法准确计算时报告 unknown/partial,不编造百分比。
|
||||
|
||||
失败不改变项目过滤,关键词 FTS/CJK/LIKE 的现有回退继续可用。原始历史与关键词检索不以语义服务成功为前提。
|
||||
|
||||
## 6. 实施阶段与退出条件
|
||||
|
||||
### A1:归属字段与数据行为
|
||||
|
||||
主要文件:hermes_state_common.py、hermes_state_schema.py、hermes_state_sessions.py 及实际创建/压缩调用点。
|
||||
|
||||
拟新增测试 `tests/test_session_project_assignment.py`:旧库升级、NULL 与省略区别、移动不改 cwd/消息、压缩延续共同移动、branch/reset 不级联、移动与延续创建的事务顺序。
|
||||
|
||||
先写失败测试,再实现最小字段和事务更新,再跑绿色回归。退出条件:数据属性正确,不触碰用户实际迁移。
|
||||
|
||||
### A2:界面与 RPC
|
||||
|
||||
修改第 3 节对应 TUI/桌面文件,接通项目内新建、草稿、恢复、移动、列表刷新及自动 repo 转显式项目。拟新增 `tests/tui_gateway/test_session_project_assignment.py`;扩展 projects.test.ts、现有项目树与 profile 测试。
|
||||
|
||||
退出条件:更改目录不换项目,移动项目不改目录;失败不误报成功,刷新恢复一致。
|
||||
|
||||
### A3:关键词项目召回
|
||||
|
||||
修改 tools/session_search_tool.py、agent/inline_tool_executors.py、hermes_state_search.py 和必要会话读取函数。拟新增 `tests/agent/test_session_search_project_assignment.py`,扩展 tests/tools/test_session_search.py。
|
||||
|
||||
退出条件:真实内联/registry 和全部搜索/读取形状按项目工作;其他项目同词不返回;移动后新查询使用新归属;原文链接可定位。
|
||||
|
||||
### A4:阶段 A 独立交付
|
||||
|
||||
真实可见 Desktop 建测试项目 P/Q:P 留历史,P 新对话能召回;cwd 改到 Q 目录仍属于 P;移动后 Q 可查、P 后续不可查;旧 NULL 会话手动移动后可查。相关 Python/TS 测试、独立代码复审通过后阶段 A 可单独交付,不等待 embedding 选型。
|
||||
|
||||
### B0:语义可行性实验
|
||||
|
||||
先检查当前仓库可复用模型配置、client、依赖和消息持久化钩子。可用则复用,不假定 memory 插件中的能力天然能用于 SessionDB。
|
||||
|
||||
用合成且明确标注的中文项目对话实验,验证候选模型的同义改写、符号查询、项目先过滤、精确扫描耗时/内存、向量表对 DB/备份影响。生成真实实验记录,包含模型、维度、实际规模、批大小、耗时、排名结果;不得用模拟 embedding 冒充真实效果。
|
||||
|
||||
阶段 A 不因 B0 网络/模型未配置而阻塞;但 B0 未通过,不宣称语义层可交付。正式历史回填前确认本地/远程服务及项目范围。
|
||||
|
||||
### B1:派生表、切片与增量任务
|
||||
|
||||
拟新增专题模块:`hermes_state_semantic.py`(派生存储/查询)、`agent/session_semantic_index.py`(切片与增量处理)。仅在不存在可复用同职责模块时创建,避免 facade 继续膨胀。
|
||||
|
||||
按第 5 节设计创建派生 schema、原文切片、模型指纹、状态恢复,接通实际消息写入/编辑/删除路径。拟新增 `tests/test_session_semantic_index.py`。
|
||||
|
||||
退出条件:重复处理幂等,中断可续,编辑/删除后旧结果不可见,网络晚到结果不会写回旧内容,原地压缩与旧式压缩历史符合可见性合同。单元测试可用确定性向量测试算法,不将其当模型实测。
|
||||
|
||||
### B2:混合检索与配置
|
||||
|
||||
拟新增 `tools/session_search_semantic.py` 承载算法辅助;现有工具入口调用它,不新增工具。修改 schema、inline/registry 参数传递、配置及用户文档。拟新增 `tests/tools/test_session_search_semantic.py`,补充真实内联测试。
|
||||
|
||||
退出条件:过滤先于向量候选截断,RRF 融合去重正确,关键词/语义模式含义明确,移动无需 embedding 重算,故障/部分覆盖状态准确,关键词独立可用。
|
||||
|
||||
### B3:效果验证与最终验收
|
||||
|
||||
使用事先固定且与调参样例分开的评估集:中文同义改写、精确函数名/错误码、多个会话相似话题、其他项目干扰、旧历史、无匹配查询。人工标注目标消息与可接受原文范围,不以模型自己说“相关”代替标准。
|
||||
|
||||
记录 keyword、semantic、hybrid 的 Top-K 目标命中、相关性和耗时;报告实际数据,不预先伪造指标。退出要求:同义改写能找回词面不同的目标;符号查询由关键词保底,不因融合明显退化;所有返回的原文属于当前项目;失败回退真实有效。B0 形成固定规模/延迟预算,B3 按该预算验收而非边跑边放宽。
|
||||
|
||||
可见 Desktop 用真实模型主动调用 hybrid,点击原文链接;移动已索引会话后再次检索,并通过索引记录/embedding 请求计数证明移动未触发已有片段重算。测试关闭语义服务时关键词仍能工作。
|
||||
|
||||
## 7. 测试命令与测试边界
|
||||
|
||||
Python 一律使用 scripts/run_tests.sh。基础文件存在后执行:
|
||||
|
||||
```bash
|
||||
scripts/run_tests.sh tests/test_session_project_assignment.py tests/tui_gateway/test_session_project_assignment.py tests/agent/test_session_search_project_assignment.py
|
||||
scripts/run_tests.sh tests/tools/test_session_search.py tests/tui_gateway/test_project_tree.py tests/hermes_cli/test_profiles_sidebar_scope.py
|
||||
scripts/run_tests.sh tests/test_session_semantic_index.py tests/tools/test_session_search_semantic.py
|
||||
scripts/run_tests.sh tests/tools/ tests/tui_gateway/ tests/hermes_cli/ tests/agent/ tests/test_hermes_state.py
|
||||
```
|
||||
|
||||
新测试路径是拟新增;实施时将实际发现的根目录 state/search/压缩专题测试加入,不将目录测试误当全部覆盖。TS 测试/typecheck/lint 使用实际 package.json 脚本。新增依赖后按仓库要求更新锁文件。
|
||||
|
||||
测试均使用临时 HERMES_HOME、真实 SessionDB 和隔离项目数据;禁止改用户其他 profile 或实际历史。行为单测与真实 embedding 网络实测分开标注。既有失败记录基线,候选造成的失败必须修复。
|
||||
|
||||
## 8. 风险、恢复和边界
|
||||
|
||||
- **本地/远程模型尚未选定:** B0 后冻结;无配置时保持关键词可用,不猜用聊天模型。
|
||||
- **远程隐私和费用:** 仅授权范围回填,设置批量/输入/超时预算;不自动发送所有 profile 历史。
|
||||
- **精确扫描规模:** SQL 项目过滤后流式扫描;实际超预算再评估索引,不提前引入外部平台。
|
||||
- **派生表增长:** 不复制完整正文,记录向量/锚点;实测 DB/备份开销,支持只清理派生索引再重建。
|
||||
- **模型切换:** 不混用指纹,新模型逐步重建,partial 期间关键词保底。
|
||||
- **停止或崩溃:** 持久化状态恢复;原文不依赖 worker,索引结果落库前校验原文标识。
|
||||
- **回滚:** 关闭 semantic 即恢复阶段 A,不删除历史或 project_id。退回不识别 project_id 的旧代码可能恢复旧宽检索/目录分组,不能当作满足新要求的安全回退。
|
||||
- **移动语义:** 后续查询按新归属,已开始的查询可以完成;不撤销已复制的历史,不建设版本撤销协议。
|
||||
|
||||
## 9. 评审结论与实施前确认
|
||||
|
||||
设计自审:V3 的简单项目属性模型完整保留;语义层只增加可重建的检索索引,不承担项目归属权威,不包含任何摘要。取消 V1/V2 的额外权限/归属扫描设计。
|
||||
|
||||
实施前确认项仅限语义技术落地:模型运行位置/服务、向量维度与输入限制、正式回填范围、实测性能预算。不是重新选择业务方案,也不阻断阶段 A。
|
||||
|
||||
本文件作为完整方案保存。架构方向已在会话中评审认可;向量性能与模型效果尚未实测,后续代码必须经过测试和独立复审。文档写入成功不代表实施通过。
|
||||
@@ -0,0 +1,322 @@
|
||||
# 项目内历史对话召回与语义检索完整方案 V4.1
|
||||
|
||||
> **For Hermes:** Use subagent-driven-development when implementation is authorized. This document is a plan, not authorization to modify business code, migrate user data, commit, push, deploy, or send historical conversations to an external service.
|
||||
|
||||
**Goal:** 会话具有独立的项目归属属性;用户移动会话只修改归属;模型能够按需通过关键词与语义混合检索找回同项目历史原文。
|
||||
|
||||
**Architecture:** 分两阶段交付:A 完成 V3 的 project_id、创建/移动/列表与项目内关键词召回;B 在其上建立可重建的原文片段向量索引,并扩展现有 session_search 为项目内混合检索。原始消息和会话项目归属是事实来源,向量是派生数据。工作目录独立,不生成项目或会话摘要,不自动注入历史上下文。
|
||||
|
||||
**Tech Stack:** 现有 Python、SQLite/FTS5、SessionDB、TUI Gateway、Desktop TypeScript。语义层首选同项目片段的精确相似度计算;embedding 模型及必要依赖由小规模实测后确定,不预先引入独立向量服务。
|
||||
|
||||
**状态与版本:** 完整独立文件,包含 V3 基础能力和语义层,不需要同时阅读旧版才能实施。V4.1 补齐对话驱动的授权、配置和自主检索流程,替代 V4 作为后续设计依据,原文件保留。本文件是实施方案,不是功能已经完成或测试已经通过的证明。
|
||||
|
||||
---
|
||||
|
||||
## 1. 范围与交付目标
|
||||
|
||||
### 1.1 包含
|
||||
|
||||
1. sessions 保存可空 project_id,和 cwd/git_repo_root 分离。
|
||||
2. 新建、草稿落库、恢复、压缩延续、独立分支正确保存或继承归属。
|
||||
3. 显式移动会话改变整条逻辑对话的项目归属,不移动消息、不改变工作目录。
|
||||
4. 项目会话列表和 session_search 按持久化归属工作。
|
||||
5. 项目内关键词、语义及混合检索,返回原文和消息锚点。
|
||||
6. 原文片段向量增量更新、删除/修改处理、中断恢复、模型切换后的重建。
|
||||
7. 索引未配置、未完成或不可用时,项目内关键词检索仍可用。
|
||||
8. Agent 在对话中完成首次远程服务说明、用户确认和配置;之后在已授权范围内自主检索和维护索引,不要求用户去设置页面启用功能。
|
||||
|
||||
### 1.2 不包含
|
||||
|
||||
- 项目摘要、会话摘要、结论蒸馏、自动每回合召回注入。
|
||||
- 新的核心工具、独立记忆管理平台、MemoryProvider 槽位占用。
|
||||
- 项目级权限撤销协议、项目版本平台、消息级项目授权、操作系统沙箱。
|
||||
- 全局 MEMORY/USER、技能、其他插件的项目隔离改造。
|
||||
- 旧会话自动按目录归类或复杂批量迁移审批系统。
|
||||
- 独立向量数据库服务、常驻队列平台或未验证必要性的近似索引。
|
||||
|
||||
## 2. 业务合同
|
||||
|
||||
| 操作 | 历史所属项目 | 工作目录 |
|
||||
|---|---|---|
|
||||
| 在项目中创建对话 | 保存选定 Project.id | 使用既有默认或用户选择 |
|
||||
| 未选择项目创建 | NULL | 按现有逻辑 |
|
||||
| 恢复对话 | 读取已存 project_id | 读取已存目录 |
|
||||
| 更改目录 | 不变 | 更新 cwd 与 git 元数据 |
|
||||
| 移动对话到项目 | 更新整条逻辑对话的 project_id | 不变 |
|
||||
| 移出项目 | 更新为 NULL | 不变 |
|
||||
| 修改项目名称或 folders | 已有归属不变 | 不自动修改 |
|
||||
| 搜索历史 | 使用当前会话持久归属 | 不参与过滤 |
|
||||
|
||||
项目归属是业务分类。移动前已经开始的检索可以完成,移动完成后发起的新查询按新归属执行;不取消在途工具、不要求先停止 Agent。已返回到其他会话的内容不会被追溯撤回。
|
||||
|
||||
无归属会话不是共同项目:不搜索其他 NULL 会话,但保留当前逻辑对话离开上下文后的压缩历史恢复能力。
|
||||
|
||||
同 profile 的显式项目 ID 才是归属值。相同路径不代表跨 profile 同项目,工具不通过任意 profile 参数或链接自动扩大检索范围。全局历史管理 UI 不因此变成权限系统。
|
||||
|
||||
## 3. 现有代码与改动落点
|
||||
|
||||
以下是已定位的现有文件;新字段、模块和测试均是拟新增设计项,不冒称已有 API。
|
||||
|
||||
| 职责 | 现有位置 | 预期改动 |
|
||||
|---|---|---|
|
||||
| 表结构及升级 | hermes_state_common.py、hermes_state_schema.py | project_id、索引、增量迁移 |
|
||||
| 会话创建/更新/浏览 | hermes_state_sessions.py | 保存归属、移动、继承、项目过滤 |
|
||||
| 搜索路径 | hermes_state_search.py | FTS/CJK/LIKE/回退的项目过滤 |
|
||||
| 工具 | tools/session_search_tool.py | 项目内所有形状、混合检索接线 |
|
||||
| Agent 实际入口 | agent/inline_tool_executors.py | 当前会话上下文及检索参数透传 |
|
||||
| 项目及目录 RPC | tui_gateway/methods_session.py、methods_projects.py、methods_config.py | 创建和移动归属、区分目录操作 |
|
||||
| 项目列表 | tui_gateway/project_tree.py | 会话成员按 project_id 分组 |
|
||||
| 桌面状态 | apps/desktop/src/store/projects.ts | 创建/移动请求、刷新与错误状态 |
|
||||
| 项目存储 | hermes_cli/projects_db.py | 复用真实 Project 对象和校验路径 |
|
||||
| 配置 | hermes_cli/config_defaults.py 及实际设置/加载路径 | embedding 配置及启用说明 |
|
||||
|
||||
会话插入/upsert、压缩延续、branch/new/reset 的具体调用者在实施前追踪后纳入。不得仅修改 facade 而漏掉实际 sibling 绑定;新逻辑过大时使用专题 sibling,不做无关拆分。
|
||||
|
||||
## 4. 阶段 A:项目归属与关键词召回
|
||||
|
||||
### 4.1 数据模型
|
||||
|
||||
sessions 增加 `project_id TEXT NULL`。在现有迁移机制中追加字段和必要索引,旧记录保留 NULL,原始消息不变。索引由实际 WHERE/排序和 EXPLAIN 决定,不预设多个冗余索引。
|
||||
|
||||
projects.db 与 state.db 是现有分库,不能直接声称有跨库 SQLite 外键。创建/移动入口校验目标项目存在且属于当前 profile;NULL 表示移出项目。归档保留归属;硬删除仍有绑定会话的项目时拒绝并提示先移动,避免孤儿归属。
|
||||
|
||||
upsert 必须区分“未提供 project_id”和“显式传入 NULL”,恢复或重复保存不得清空已有归属。
|
||||
|
||||
### 4.2 逻辑对话与移动
|
||||
|
||||
正常移动只更新属性。若用户可见的一条对话由多个历史压缩 session 记录组成,在同一个现有 SessionDB 写事务内更新其 project_id:
|
||||
- 原地压缩无需移动消息。
|
||||
- 压缩延续继承同一归属,和逻辑对话一起移动。
|
||||
- 独立 branch 创建时继承项目,之后独立移动。
|
||||
- reset/new 的新对话初始化项目,但不跟随旧对话后续移动。
|
||||
- 委派记录维持现有隐藏规则。
|
||||
|
||||
不能使用沿所有 parent 边走根的宽泛 helper 代替逻辑对话判定。现有代码区分 branch/reset/压缩,需复用其准确语义并测试。
|
||||
|
||||
移动与压缩延续创建采用现有数据库串行写事务;延续创建在写事务中读取最新来源归属,不从长期缓存复制旧 project_id。这是数据一致性要求,不新增项目锁服务或移动禁用机制。
|
||||
|
||||
### 4.3 旧会话与自动仓库入口
|
||||
|
||||
旧会话为 NULL,用户直接通过“移动到项目”归类即可。不做目录推断自动回填。
|
||||
|
||||
自动发现的 repo 路径不是 Project.id。该入口用于创建项目内对话或成为移动目标时,复用已有项目创建/查重能力取得稳定 ID。仅浏览仓库不必批量创建项目;不要把路径写进 project_id。
|
||||
|
||||
### 4.4 前后端接线
|
||||
|
||||
项目内新建与草稿第一次持久化都传入归属。恢复从 DB 读取,不按全局当前项目或 cwd 重算。
|
||||
|
||||
提供一个窄职责的设置会话项目 RPC,方法名在核实注册表后确定,入参使用已有会话标识和可空 project_id,返回实际写入结果。前端失败时不改成成功状态;成功后刷新源项目、目标项目和会话信息。
|
||||
|
||||
`session.cwd.set` 和 `session.workspace.move` 保留目录操作语义,不隐式改归属;“移动到项目”不能继续只调用 workspace.move。项目成员按 project_id 展示,repo/lane 和目录信息尽量沿用现有布局,不重写整个侧栏。
|
||||
|
||||
### 4.5 项目内检索
|
||||
|
||||
Agent 每次调用从可信当前 session/profile 读取当前归属,不长期缓存移动前的 ID。默认当前项目,无须添加只有 project 一个取值的 scope 参数。
|
||||
|
||||
查询时在 sessions 的 JOIN/WHERE 中施加 project_id,发生在候选 LIMIT 之前。不做全库先截再筛、不建立全量成员表、不将历史集合截为 900/2000 条。
|
||||
|
||||
关键词、标题、browse、read、scroll、首尾片段、窗口、lineage 展开和 profile fallback 均遵守项目范围;内部 SessionDB 可保留未过滤查询供其他可信调用者使用,但 Agent 不能把 NULL/读取错误当成不过滤。
|
||||
|
||||
保留既有角色、隐藏来源、压缩/撤销可见性和 bounded browse 行为。NULL 情况只恢复自身可靠压缩延续历史,不把 branch/reset/delegation 全树当作自身。
|
||||
|
||||
原文 session_id 与 message_id 必须配对,保留链接。不改系统提示词或已发送历史消息,只用工具结果进入上下文。
|
||||
|
||||
## 5. 阶段 B:语义与混合检索
|
||||
|
||||
### 5.1 技术路线与前置小实验
|
||||
|
||||
首版选用原文片段 embedding + 项目内精确余弦相似度 + 排名融合。先验证精确扫描的正确性、中文效果和实际规模耗时;不因“语义检索”直接引入远程向量服务或近似索引。
|
||||
|
||||
向量存储优先在现有 profile 的 state.db 中新增派生表,便于直接 JOIN sessions/messages 实现前置过滤。大 BLOB 对现有备份、数据库大小和读写锁的影响必须实测。若该选择不能通过规模验证,更新语义层存储选择,不绕过项目前置过滤,不阻塞阶段 A 的独立交付。
|
||||
|
||||
相似度可使用已验证可用的数值实现;仓库某些 optional extra 中出现 numpy 不表示基础安装已有该依赖。新增依赖必须明确声明、设置版本上界并更新锁文件。不要先写出项目里不存在的 embedding client/import。
|
||||
|
||||
### 5.2 模型配置与数据发送
|
||||
|
||||
正式回填前确定:embedding provider、模型及版本、向量维度、输入上限、批大小、超时,以及本地或远程运行方式。由 Agent 在对话中引导确认并调用现有配置与 setup 能力,不要求用户打开设置页。不增加非秘密 HERMES_* 环境变量;密钥通过已有安全凭据输入/存储能力处理,不进入普通聊天正文或工具日志。
|
||||
|
||||
建议拟用独立 `session_search.semantic` 配置节承载 enabled、模型路由与索引预算;最终字段在核实现有配置模式后确定。语义功能默认未配置/关闭时阶段 A 正常工作。
|
||||
|
||||
本方案保存不代表用户授权把历史发送给第三方。先使用合成测试对话验证候选模型;没有可用路由时不安装大型本地模型或猜用聊天模型代替。
|
||||
|
||||
### 5.2.1 对话内首次确认
|
||||
|
||||
首次需要远程语义检索且没有匹配的授权记录时,session_search 返回项目内关键词结果及结构化 `consent_required` 状态,不发送查询或历史正文给 embedding 服务。Agent 据此在当前对话说明:
|
||||
- 服务方及实际接收数据的 endpoint、拟用模型。
|
||||
- 发送的数据为当前检索文本与所选项目的历史原文片段;默认 user/assistant,不含未解析附件和未经批准的工具正文。
|
||||
- 授权项目范围,默认当前项目,明确包含首次历史索引及后续新增/修改消息的增量处理。
|
||||
- 已知计费方式及可确定的预算;未知费用明确说明,不编造估算。
|
||||
|
||||
使用现有对话确认能力让用户接受或拒绝,可修改范围。沉默、超时、普通继续执行指令和仅讨论设计均不算授权。拒绝后继续关键词检索,不重复强迫弹问;用户后来明确请求启用时再确认。
|
||||
|
||||
用户明确同意后,Agent 经现有配置/CLI 能力提交与确认内容一致的服务和范围。优先复用现有安全确认与配置记录机制;若不适用,仅增加本能力的最小配置字段和验证命令,不建设通用授权平台或新增核心工具。确认记录和实际配置写入后必须读回校验,再启用网络调用。不得仅凭模型传 `approved=true` 或临时变量放行。
|
||||
|
||||
需要凭据时使用现有安全输入能力在对话流程中取得;不猜测凭据,不要求明文贴到普通消息。凭据缺失、配置失败或用户未完成输入时,保留关键词路径,并明确语义尚未就绪。
|
||||
|
||||
### 5.2.2 授权范围与配置变更
|
||||
|
||||
最低持久记录包含当前 profile、明确的项目 ID 集、服务方/接收 endpoint、已说明的模型及数据类型、包含历史和后续增量的范围、确认来源/时间及启用状态。不记录密钥。授权只是远程 embedding 数据发送许可,不改变会话 project_id 或历史检索归属。
|
||||
|
||||
每次远程查询及后台发送批次开始前检查此记录与实际配置。更换服务方/接收 endpoint、增加项目或数据类型须再次在对话确认。仅密钥轮换不扩大授权;模型变化若改变接收方、数据使用条件或突破已确认费用预算也须确认。未获得适用授权时不将对话内容发出。
|
||||
|
||||
归属授权绑定项目 ID,不绑定目录:修改 cwd 不影响许可;把会话移动到未授权项目,不自动扩展许可。已授权项目的后续新增内容包括用户显式移入的历史会话,首次确认时应说明这一点;其缺失索引可自动补齐。跨 profile 不继承授权。
|
||||
|
||||
用户可以在对话中停用远程语义或撤销指定项目许可。停止新增远程请求,关键词继续可用;已经发送的请求和服务方已收到的数据不能追溯收回。保留本地派生索引不代表允许继续调用远程服务。
|
||||
|
||||
### 5.2.3 日常自主检索
|
||||
|
||||
配置和确认完成后,用户只需正常提问。Agent 判断是否需要项目历史,并调用现有 session_search;默认 hybrid,无需用户手选算法、索引批次或重试。
|
||||
|
||||
程序负责范围过滤、查询向量、增量索引、截止时间和状态。索引缺失时按已批准范围和批量预算自动补建;可在本次工具预算内完成则检索,否则先返回关键词结果及 `indexing/partial` 状态。Agent 判断关键词证据是否足够;不足时用现有工具再次检索查看进度,或清楚说明尚在索引,不能把未索引当作没有历史。
|
||||
|
||||
重查有界、遵守批处理退避和失败状态,不忙循环、不无限调用、不要求用户手动点重试。索引继续由宿主批处理推进,不依赖本轮模型持续占用。没有新授权需求时,不因每次调用、进程重启或会话切换重复询问。设置页可保留诊断信息,但不是主流程或交付前提。
|
||||
|
||||
### 5.3 原文切片与派生数据
|
||||
|
||||
按单条原始消息稳定切片,长消息分段并有有限重叠。保持消息 ID 和字符范围;命中后利用既有窗口读取相邻对话,不先总结再 embedding,也不把完整长会话当一个向量。
|
||||
|
||||
片段大小依据实际模型输入上限和 tokenizer 确定,明确 token/字符单位,不能混用。默认索引 user/assistant 文本;工具正文只有在用户启用相应索引范围后处理,其余显式 tool-role 查询仍走现有关键词路径并标明语义覆盖不足。索引未解析附件二进制,不把 base64 送入 embedding。
|
||||
|
||||
拟新增派生表使用清晰字段:session_id、message_id、片段起止位置、原文内容标识、切片策略版本、embedding 模型指纹/维度、向量、处理状态和必要时间戳。具体表名在实现时确定,原文仍由 messages 读取,不保存另一份权威消息。
|
||||
|
||||
模型指纹包含端点/提供方、模型身份、维度、归一化与切片版本等影响兼容性的非秘密信息,不能含 token。不同指纹的查询向量和索引向量不可混用。
|
||||
|
||||
不保存用于授权的 project_id 副本。向量通过 message/session 关联当前 sessions.project_id。片段标识稳定,唯一键保证重试不产生重复有效片段。
|
||||
|
||||
### 5.4 增量索引与中断恢复
|
||||
|
||||
初次索引在对话确认的项目范围内由程序自动分批执行;之后对已启用范围的新建/修改消息增量处理。Agent 无需逐批调度,用户无需手动启动回填。先定位已有消息写入、更新、删除和可见性修改路径,复用实际持久化边界,不新造通用事件平台。
|
||||
|
||||
最低实现为派生行状态和内容标识:pending、ready、failed 可恢复。正文变化使旧片段失效并重建;仅记录最大 message_id 不能识别旧消息编辑。写入端不要做 embedding 网络调用。
|
||||
|
||||
使用嵌入宿主的有界后台批处理,明确启动、唤醒、关闭和每批预算;不建设独立常驻服务。进程退出后从持久化 pending/failed 恢复,不依赖 Agent 子进程存活。若多个进程同时执行,用现有 SQLite 事务做最小任务领取/幂等写入,不新增调度平台。
|
||||
|
||||
开始 embedding 前读取原文标识并检查当前项目/服务发送许可,结果落库前再检查原文仍存在且未变化;旧网络结果不能覆盖已编辑或已删除消息。网络期间不持有长写事务。首次回填可重试、可中断、可继续;失败记录原因,不无限重试阻塞用户回复。查询与增量处理都必须经过相同许可检查,不能只检查前台工具而漏掉后台回填。
|
||||
|
||||
查询端同时检查原消息当前可见性和内容标识,尚未物理清理的失效向量不能贡献结果。删除/撤销立即通过原文 JOIN 和可见性条件排除,后台清理派生行;压缩归档按既有历史规则保留,不能一律排除 active=0。分支副本与原消息各自保留正确来源。
|
||||
|
||||
移动会话只修改 project_id,不重算已有向量。若目标项目已启用语义但该会话尚未索引,正常增量批处理补齐缺片段;若目标未启用,不因为历史已有向量就绕过配置调用远程模型。
|
||||
|
||||
切换模型后逐步生成新指纹索引,旧指纹不可参与新模型检索。新索引部分覆盖时可返回部分语义结果并明确状态,关键词保持覆盖;按项目过滤后不会为了补足结果搜索其他项目。
|
||||
|
||||
### 5.5 查询与融合
|
||||
|
||||
仅 query discovery 启用语义;browse/read/scroll 继续直接读取,不产生无意义的 embedding 调用。
|
||||
|
||||
查询顺序:
|
||||
1. 从 DB 解析当前会话项目及现有角色/可见性限制。
|
||||
2. 在该范围执行既有关键词检索。
|
||||
3. 配置可用时生成查询向量;SQL 先 JOIN 当前 sessions/messages,过滤 project_id、有效原文、模型指纹和角色,再计算该集合的向量相似度。
|
||||
4. 精确扫描可分批读取整个允许集合,用固定大小候选堆保存最相关结果,不先按时间 LIMIT 一批向量,也不在全库 Top-K 之后才过滤项目。
|
||||
5. 关键词与语义按匹配消息锚点归并,以 Reciprocal Rank Fusion(基于排名而非原始分数相加)生成混合候选,再执行既有会话去重与结果展开。
|
||||
6. 对命中消息读取原文窗口,返回实际 session_id/message_id/link。窗口仍遵守阶段 A 的项目和可见性规则。
|
||||
|
||||
建议给现有 session_search 增加 `search_mode=hybrid|keyword|semantic`,默认 hybrid;函数、schema、registry 和 inline 映射同步。模式仅控制算法,不改变项目范围。keyword 便于精确符号检索和对照;semantic 用于用户明确的相似问题追溯及效果测试。
|
||||
|
||||
hybrid 在语义不可用时返回关键词结果并注明实际模式与原因。显式 semantic 失败应报告不可用,不冒称结果来自语义;用户或模型可改用 keyword。全程共用调用截止时间,不串行无限等待 embedding。
|
||||
|
||||
现有 sort=newest/oldest 保留为项目内结果的时间偏好,在两路候选融合后使用一致的已文档化顺序;不是承诺搜遍所有命中。确切排名参数通过固定测试集验证后在代码/配置中明确,不混用可变随机参数。
|
||||
|
||||
### 5.6 覆盖与降级状态
|
||||
|
||||
工具结果最少携带 requested_mode、used_mode、semantic_status(disabled/consent_required/configuration_required/indexing/ready/partial/unavailable)、必要的截断/超时提示和匹配来源(keyword/semantic/both)。需要确认时给出非秘密的服务及范围说明,供 Agent 提问;需要配置时给出可执行的现有配置路径提示。不要回传向量 BLOB、凭据或模型可自行认可的授权令牌。
|
||||
|
||||
“零命中”与“语义未索引完整”分开。已处理到最大 ID 不等于索引覆盖完整;覆盖统计依据已选范围的有效消息/片段状态,无法准确计算时报告 unknown/partial,不编造百分比。
|
||||
|
||||
失败不改变项目过滤,关键词 FTS/CJK/LIKE 的现有回退继续可用。原始历史与关键词检索不以语义服务成功为前提。
|
||||
|
||||
## 6. 实施阶段与退出条件
|
||||
|
||||
### A1:归属字段与数据行为
|
||||
|
||||
主要文件:hermes_state_common.py、hermes_state_schema.py、hermes_state_sessions.py 及实际创建/压缩调用点。
|
||||
|
||||
拟新增测试 `tests/test_session_project_assignment.py`:旧库升级、NULL 与省略区别、移动不改 cwd/消息、压缩延续共同移动、branch/reset 不级联、移动与延续创建的事务顺序。
|
||||
|
||||
先写失败测试,再实现最小字段和事务更新,再跑绿色回归。退出条件:数据属性正确,不触碰用户实际迁移。
|
||||
|
||||
### A2:界面与 RPC
|
||||
|
||||
修改第 3 节对应 TUI/桌面文件,接通项目内新建、草稿、恢复、移动、列表刷新及自动 repo 转显式项目。拟新增 `tests/tui_gateway/test_session_project_assignment.py`;扩展 projects.test.ts、现有项目树与 profile 测试。
|
||||
|
||||
退出条件:更改目录不换项目,移动项目不改目录;失败不误报成功,刷新恢复一致。
|
||||
|
||||
### A3:关键词项目召回
|
||||
|
||||
修改 tools/session_search_tool.py、agent/inline_tool_executors.py、hermes_state_search.py 和必要会话读取函数。拟新增 `tests/agent/test_session_search_project_assignment.py`,扩展 tests/tools/test_session_search.py。
|
||||
|
||||
退出条件:真实内联/registry 和全部搜索/读取形状按项目工作;其他项目同词不返回;移动后新查询使用新归属;原文链接可定位。
|
||||
|
||||
### A4:阶段 A 独立交付
|
||||
|
||||
真实可见 Desktop 建测试项目 P/Q:P 留历史,P 新对话能召回;cwd 改到 Q 目录仍属于 P;移动后 Q 可查、P 后续不可查;旧 NULL 会话手动移动后可查。相关 Python/TS 测试、独立代码复审通过后阶段 A 可单独交付,不等待 embedding 选型。
|
||||
|
||||
### B0:语义可行性实验
|
||||
|
||||
先检查当前仓库可复用模型配置、client、依赖和消息持久化钩子。可用则复用,不假定 memory 插件中的能力天然能用于 SessionDB。
|
||||
|
||||
用合成且明确标注的中文项目对话实验,验证候选模型的同义改写、符号查询、项目先过滤、精确扫描耗时/内存、向量表对 DB/备份影响。生成真实实验记录,包含模型、维度、实际规模、批大小、耗时、排名结果;不得用模拟 embedding 冒充真实效果。
|
||||
|
||||
阶段 A 不因 B0 网络/模型未配置而阻塞;但 B0 未通过,不宣称语义层可交付。正式历史回填前确认本地/远程服务及项目范围。
|
||||
|
||||
### B1:派生表、切片与增量任务
|
||||
|
||||
拟新增专题模块:`hermes_state_semantic.py`(派生存储/查询)、`agent/session_semantic_index.py`(切片与增量处理)。仅在不存在可复用同职责模块时创建,避免 facade 继续膨胀。
|
||||
|
||||
按第 5 节设计创建派生 schema、原文切片、模型指纹、状态恢复,接通实际消息写入/编辑/删除路径。拟新增 `tests/test_session_semantic_index.py`。
|
||||
|
||||
退出条件:重复处理幂等,中断可续,编辑/删除后旧结果不可见,网络晚到结果不会写回旧内容,原地压缩与旧式压缩历史符合可见性合同。单元测试可用确定性向量测试算法,不将其当模型实测。
|
||||
|
||||
### B2:混合检索与配置
|
||||
|
||||
拟新增 `tools/session_search_semantic.py` 承载算法辅助;现有工具入口调用它,不新增工具。修改 schema、inline/registry 参数传递、配置及用户文档。拟新增 `tests/tools/test_session_search_semantic.py`,补充真实内联测试。
|
||||
|
||||
退出条件:过滤先于向量候选截断,RRF 融合去重正确,关键词/语义模式含义明确,移动无需 embedding 重算,故障/部分覆盖状态准确,关键词独立可用。
|
||||
|
||||
对话配置与自动运行必须作为 B2 的正式交付项:检索返回确认/配置状态 → Agent 调用现有确认能力 → 用户选择 → 配置及记录持久化/读回 → 自动索引 → 再次检索。按实际 setup/CLI 入口实现最小能力,不用新的设置页面替代此链路。
|
||||
|
||||
新增行为测试覆盖:未确认不发查询和历史;拒绝/超时继续关键词;确认后前后台可运行;重启和新对话复用同范围许可;换服务/加项目再次确认;撤销后不再启动远程请求;后台旧队列在许可变化后不能继续按旧配置发送。使用真实配置与临时 profile,HTTP 测试服务验证实际发送与未发送;真实模型效果另测。
|
||||
|
||||
### B3:效果验证与最终验收
|
||||
|
||||
使用事先固定且与调参样例分开的评估集:中文同义改写、精确函数名/错误码、多个会话相似话题、其他项目干扰、旧历史、无匹配查询。人工标注目标消息与可接受原文范围,不以模型自己说“相关”代替标准。
|
||||
|
||||
记录 keyword、semantic、hybrid 的 Top-K 目标命中、相关性和耗时;报告实际数据,不预先伪造指标。退出要求:同义改写能找回词面不同的目标;符号查询由关键词保底,不因融合明显退化;所有返回的原文属于当前项目;失败回退真实有效。B0 形成固定规模/延迟预算,B3 按该预算验收而非边跑边放宽。
|
||||
|
||||
可见 Desktop 从未配置状态开始,用户只通过聊天操作:提问历史 → Agent 在对话解释服务/范围并请求确认 → 用户同意并通过安全入口完成必要凭据 → 自动配置/索引 → hybrid 找回原文。全流程不要求用户进入设置页或手动启动索引。
|
||||
|
||||
再次开启同项目对话时正常提问即可检索,不重复授权;测试拒绝首次发送时仍返回关键词结果。移动已索引会话后再次检索,并通过索引记录/embedding 请求计数证明移动未触发已有片段重算。测试关闭语义服务时关键词仍能工作,增加未授权项目时只在对话请求新的范围确认。
|
||||
|
||||
## 7. 测试命令与测试边界
|
||||
|
||||
Python 一律使用 scripts/run_tests.sh。基础文件存在后执行:
|
||||
|
||||
```bash
|
||||
scripts/run_tests.sh tests/test_session_project_assignment.py tests/tui_gateway/test_session_project_assignment.py tests/agent/test_session_search_project_assignment.py
|
||||
scripts/run_tests.sh tests/tools/test_session_search.py tests/tui_gateway/test_project_tree.py tests/hermes_cli/test_profiles_sidebar_scope.py
|
||||
scripts/run_tests.sh tests/test_session_semantic_index.py tests/tools/test_session_search_semantic.py
|
||||
scripts/run_tests.sh tests/tools/ tests/tui_gateway/ tests/hermes_cli/ tests/agent/ tests/test_hermes_state.py
|
||||
```
|
||||
|
||||
新测试路径是拟新增;实施时将实际发现的根目录 state/search/压缩专题测试加入,不将目录测试误当全部覆盖。TS 测试/typecheck/lint 使用实际 package.json 脚本。新增依赖后按仓库要求更新锁文件。
|
||||
|
||||
测试均使用临时 HERMES_HOME、真实 SessionDB 和隔离项目数据;禁止改用户其他 profile 或实际历史。行为单测与真实 embedding 网络实测分开标注。既有失败记录基线,候选造成的失败必须修复。
|
||||
|
||||
## 8. 风险、恢复和边界
|
||||
|
||||
- **本地/远程模型尚未选定:** B0 后冻结;无配置时保持关键词可用,不猜用聊天模型。
|
||||
- **远程隐私和费用:** 对话首次确认后,在授权范围自动回填与增量索引,设置批量/输入/超时预算;不自动发送所有 profile 历史。查询文本同样属于出站数据,前后台统一检查许可。
|
||||
- **精确扫描规模:** SQL 项目过滤后流式扫描;实际超预算再评估索引,不提前引入外部平台。
|
||||
- **派生表增长:** 不复制完整正文,记录向量/锚点;实测 DB/备份开销,支持只清理派生索引再重建。
|
||||
- **模型切换:** 不混用指纹,新模型逐步重建,partial 期间关键词保底。
|
||||
- **停止或崩溃:** 持久化状态恢复;原文不依赖 worker,索引结果落库前校验原文标识。
|
||||
- **回滚:** 关闭 semantic 即恢复阶段 A,不删除历史或 project_id。退回不识别 project_id 的旧代码可能恢复旧宽检索/目录分组,不能当作满足新要求的安全回退。
|
||||
- **移动语义:** 后续查询按新归属,已开始的查询可以完成;不撤销已复制的历史,不建设版本撤销协议。
|
||||
|
||||
## 9. 评审结论与实施前确认
|
||||
|
||||
设计自审:V3 的简单项目属性模型完整保留;语义层只增加可重建的检索索引,不承担项目归属权威,不包含任何摘要。取消 V1/V2 的额外权限/归属扫描设计。
|
||||
|
||||
实施前确认项仅限语义技术落地:模型运行位置/服务、向量维度与输入限制、正式回填范围、实测性能预算。首次远程数据发送许可在对话中完成并持久化;之后同范围日常检索自主进行。不是重新选择业务方案,也不阻断阶段 A。用户对本交互规则的认可不等于已许可向某个尚未指定的远程服务发送历史。
|
||||
|
||||
本文件作为完整方案保存。架构方向已在会话中评审认可;向量性能与模型效果尚未实测,后续代码必须经过测试和独立复审。文档写入成功不代表实施通过。
|
||||
@@ -0,0 +1,351 @@
|
||||
# 项目历史召回插件设计方案 V2
|
||||
|
||||
> **状态:** 设计方案,非实施代码。V2 修正了 V1 中"插件叠加 workspace.move RPC"的错误前提,补齐执行所需的全部接口定义。
|
||||
|
||||
**Goal:** 一个插件解决两个需求:(1) 会话项目归属,(2) 项目内历史的按需关键词+语义检索。
|
||||
|
||||
**定位:** Footprint Ladder 第 4 级插件。装在 `~/.hermes/plugins/project-history-recall/`。核心零改动。
|
||||
|
||||
---
|
||||
|
||||
## 一、修正:workspace.move 不可用
|
||||
|
||||
V1 假设插件能调用 `session.workspace.move`(tui_gateway JSON-RPC 方法)。核实后这是错误的:
|
||||
|
||||
- 插件工具通过 `ctx.register_tool` 注册(`hermes_cli/plugins.py:449-491`),走 `handle_function_call` → `tools/registry.py` 路径。
|
||||
- `session.workspace.move` 是 tui_gateway 的 JSON-RPC 方法(`tui_gateway/methods_session.py:840`),插件工具没有它的句柄。
|
||||
- 两者是不同的调用通道,插件无法"叠加"到 workspace.move 上。
|
||||
|
||||
**修正后的移动路径:** 插件自己实现归属更新,不调用 workspace.move。归属与 cwd 一致移动的含义变为:**插件工具同时更新归属记录和会话的 cwd/git_repo_root**——通过直接调用 `SessionDB.update_session_cwd()`(`hermes_state_sessions.py:486`)写 state.db。这是唯一的例外写入,且复用核心已有的公共写接口,不绕过核心的持久化逻辑。
|
||||
|
||||
## 二、插件结构
|
||||
|
||||
```
|
||||
~/.hermes/plugins/project-history-recall/
|
||||
├── plugin.yaml
|
||||
├── __init__.py # register(ctx)
|
||||
├── attribution.py # 归属解析(纯函数,无 IO)
|
||||
├── search.py # 关键词+语义混合检索
|
||||
├── store.py # 归属记录读写(state.db 只读 + update_session_cwd 例外)
|
||||
└── config_schema.py # embedding 配置声明
|
||||
```
|
||||
|
||||
## 三、需求一:项目归属
|
||||
|
||||
### 3.1 归属解析(attribution.py)
|
||||
|
||||
纯函数,输入为数据不读文件:
|
||||
|
||||
```python
|
||||
def resolve_project(
|
||||
session_row: dict, # 当前会话行:{id, cwd, git_repo_root, ...}
|
||||
projects: list[dict], # projects_db.list_projects() 的结果(含 folders)
|
||||
) -> dict:
|
||||
"""返回归属身份。三态:
|
||||
- {"kind": "explicit", "id": "p_abc", "name": "Hermes"} — 显式项目
|
||||
- {"kind": "auto", "key": "path_key_of_repo_root"} — 自动 repo(Tier 2)
|
||||
- {"kind": "home"} — 无归属(Home 桶)
|
||||
"""
|
||||
```
|
||||
|
||||
规则(镜像桌面端 project_tree.py 的语义,插件内重新实现):
|
||||
1. 用 `session_row["cwd"]` 和 `session_row["git_repo_root"]` 匹配 projects 的 folders(最长祖先匹配,路径规范化处理 Windows casefold)。
|
||||
2. 显式命中 → `{"kind": "explicit", "id": ..., "name": ...}`。
|
||||
3. 无显式命中但有 git_repo_root → `{"kind": "auto", "key": path_key(git_repo_root)}`。
|
||||
4. 两者都无 → `{"kind": "home"}`。
|
||||
|
||||
### 3.2 归属记录存储
|
||||
|
||||
归属记录的权威存储是 **state.db 的 sessions 表**,但插件不加列。归属关系通过**运行时解析**得出:每次检索时,从当前会话行的 cwd/git_repo_root + projects.db 的 folders 计算归属。
|
||||
|
||||
这等价于"归属 = cwd/git_repo_root 的函数",不需要持久化 project_id 列。优点是核心零改动;缺点是归属随 cwd 变化——用户手动改 cwd 时归属跟着变。这是接受的行为(cwd 是用户选择的工作空间,归属跟随工作空间)。
|
||||
|
||||
### 3.3 移动会话(project_session_move 工具)
|
||||
|
||||
移动 = 改 cwd 到目标项目的目录。归属自然跟随(因为归属是 cwd 的函数)。
|
||||
|
||||
实现:调用 `SessionDB.update_session_cwd(session_id, new_cwd, git_branch, git_repo_root, replace_git_meta=True)`(`hermes_state_sessions.py:486-517`),同时做 git 探测(`git_probe.branch()` / `git_probe.common_repo_root()`)获取 branch 和 repo_root。
|
||||
|
||||
**只读约束的例外:** `update_session_cwd` 是核心已有的公共写接口,插件调用它写 state.db 是合法复用,不违背"不改核心 schema"的约束。插件绝不直接 INSERT/UPDATE/DELETE messages 表。
|
||||
|
||||
### 3.4 工具 schema
|
||||
|
||||
```python
|
||||
MOVE_SCHEMA = {
|
||||
"name": "project_session_move",
|
||||
"description": (
|
||||
"Move this conversation to a different project. The working directory "
|
||||
"switches to the target project's folder, and future history searches "
|
||||
"will look in the new project. Use when the user says 'move this to X' "
|
||||
"or 'this belongs to project Y'."
|
||||
),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"project_id": {
|
||||
"type": "string",
|
||||
"description": (
|
||||
"Target project ID (e.g. 'p_abc123'). Use '' to remove from "
|
||||
"any project (moves to Home). The conversation's working "
|
||||
"directory changes to the project's primary folder."
|
||||
),
|
||||
},
|
||||
},
|
||||
"required": ["project_id"],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Handler 逻辑:
|
||||
1. `project_id` 为空 → 移到 Home:调用 `update_session_cwd` 设为 `~`(或 `terminal.cwd` 配置)。
|
||||
2. `project_id` 非空 → 从 projects.db 读目标项目的 primary_path → git 探测该路径 → `update_session_cwd`。
|
||||
3. 返回 `{"success": true, "project": {...}, "cwd": "...", "message": "Moved to project X, cwd now ..."}`。
|
||||
4. 失败(项目不存在、路径不存在、git 探测失败)返回 `{"success": false, "error": "..."}`。
|
||||
|
||||
## 四、需求二:项目内历史检索
|
||||
|
||||
### 4.1 工具 schema
|
||||
|
||||
```python
|
||||
SEARCH_SCHEMA = {
|
||||
"name": "project_history_search",
|
||||
"description": (
|
||||
"Search past conversations in the SAME project as this one. "
|
||||
"Returns actual messages from other sessions, with links. "
|
||||
"Use when the user asks about prior work on this codebase: "
|
||||
"'what did we decide about X', 'where did we leave Y'. "
|
||||
"Modes: 'hybrid' (default, keyword + semantic), 'keyword' "
|
||||
"(exact terms), 'semantic' (similar meaning)."
|
||||
),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"query": {
|
||||
"type": "string",
|
||||
"description": "Search query. Keywords, phrases, or natural language.",
|
||||
},
|
||||
"mode": {
|
||||
"type": "string",
|
||||
"enum": ["hybrid", "keyword", "semantic"],
|
||||
"default": "hybrid",
|
||||
"description": (
|
||||
"'hybrid' (default): keyword + semantic combined. "
|
||||
"'keyword': exact terms only. 'semantic': similar meaning only. "
|
||||
"Use 'keyword' for code symbols/error codes; 'semantic' for "
|
||||
"'how did we handle X' when exact words differ."
|
||||
),
|
||||
},
|
||||
"limit": {
|
||||
"type": "integer",
|
||||
"default": 3,
|
||||
"description": "Max sessions to return (default 3, max 10).",
|
||||
},
|
||||
},
|
||||
"required": ["query"],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### 4.2 归属过滤的 SQL
|
||||
|
||||
归属 = cwd 的函数,所以过滤条件直接作用在 sessions 行的 cwd/git_repo_root 上:
|
||||
|
||||
```sql
|
||||
-- 显式项目:cwd 或 git_repo_root 在项目 folders 的子树内
|
||||
WHERE (
|
||||
s.cwd LIKE ? ESCAPE '\' -- folder_path + '/%'
|
||||
OR s.git_repo_root LIKE ? ESCAPE '\'
|
||||
OR s.cwd = ? -- folder_path 精确匹配
|
||||
OR s.git_repo_root = ?
|
||||
)
|
||||
|
||||
-- 自动 repo:git_repo_root 匹配
|
||||
WHERE s.git_repo_root = ?
|
||||
|
||||
-- Home 桶:无 cwd 且无 git_repo_root
|
||||
WHERE s.cwd IS NULL AND s.git_repo_root IS NULL
|
||||
```
|
||||
|
||||
LIKE 的 `%` 和 `_` 用 `ESCAPE '\'` 转义。路径比较前统一规范化(去尾部 `/`,Windows casefold)。
|
||||
|
||||
**当前会话排除:** 归属过滤天然包含当前会话(它的 cwd 在项目里)。关键词搜索时需要额外排除当前 session_id,避免返回自己:
|
||||
|
||||
```sql
|
||||
AND s.id != ? -- current_session_id
|
||||
```
|
||||
|
||||
### 4.3 关键词路径
|
||||
|
||||
```python
|
||||
def _keyword_search(db_path: str, query: str, scope_where: str, scope_params: tuple,
|
||||
current_session_id: str, limit: int) -> list[dict]:
|
||||
"""FTS5 search scoped to project sessions."""
|
||||
# 用 messages_fts FTS5 表 + sessions JOIN
|
||||
sql = f"""
|
||||
SELECT m.id, m.session_id, m.role, m.content, m.timestamp,
|
||||
s.title, s.started_at, s.cwd, s.git_repo_root,
|
||||
rank
|
||||
FROM messages_fts f
|
||||
JOIN messages m ON f.rowid = m.id
|
||||
JOIN sessions s ON m.session_id = s.id
|
||||
WHERE messages_fts MATCH ?
|
||||
AND m.role IN ('user', 'assistant')
|
||||
AND m.active = 1
|
||||
AND s.archived = 0 AND s.hidden = 0
|
||||
AND {scope_where}
|
||||
AND s.id != ?
|
||||
ORDER BY rank
|
||||
LIMIT ?
|
||||
"""
|
||||
params = (query, *scope_params, current_session_id, limit)
|
||||
# ... execute, return rows
|
||||
```
|
||||
|
||||
FTS5 查询语法沿用现有 `_sanitize_fts5_query` 的消毒规则。LIKE 回退(CJK 查询)用 `m.content LIKE ? ESCAPE '\'`。
|
||||
|
||||
### 4.4 语义路径
|
||||
|
||||
插件自有 `embeddings.db`(`~/.hermes/plugins/project-history-recall/embeddings.db`):
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS chunks (
|
||||
id INTEGER PRIMARY KEY,
|
||||
session_id TEXT NOT NULL,
|
||||
message_id INTEGER NOT NULL,
|
||||
chunk_index INTEGER NOT NULL,
|
||||
content_hash TEXT NOT NULL,
|
||||
model_fingerprint TEXT NOT NULL,
|
||||
vector BLOB NOT NULL,
|
||||
created_at REAL NOT NULL,
|
||||
UNIQUE(session_id, message_id, chunk_index, model_fingerprint)
|
||||
);
|
||||
```
|
||||
|
||||
**索引构建:**
|
||||
1. 归属解析出项目会话集合。
|
||||
2. 读这些会话的 user/assistant 消息(`active = 1`)。
|
||||
3. 按消息切片(chunk_size=512 tokens, overlap=64 tokens,起始值,P4 实测调整)。
|
||||
4. 计算 content_hash(SHA-256),跳过已有且未变的 chunk。
|
||||
5. 调用 embedding 模型,存向量。
|
||||
|
||||
**检索:**
|
||||
1. 归属解析 → 项目会话 ID 集合。
|
||||
2. 从 embeddings.db 读该集合内所有 chunk 的向量。
|
||||
3. 计算查询向量与每个 chunk 的余弦相似度。
|
||||
4. 取 Top-K(K=20,起始值),按 message_id 归并到消息级别。
|
||||
|
||||
**混合检索(RRF):**
|
||||
- 关键词结果按 rank 排序,语义结果按相似度排序。
|
||||
- RRF 分数:`1 / (k + rank_keyword) + 1 / (k + rank_semantic)`,k=60(标准值)。
|
||||
- 按 message_id 归并,取 Top-N。
|
||||
|
||||
### 4.5 结果格式
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"mode": "hybrid",
|
||||
"project": {"kind": "explicit", "id": "p_abc", "name": "Hermes"},
|
||||
"results": [{
|
||||
"session_id": "20260901_143022_a1b2",
|
||||
"link": "@session:20260901_143022_a1b2",
|
||||
"title": "项目归属设计讨论",
|
||||
"message_id": 47,
|
||||
"snippet": "……项目归属应该持久化……",
|
||||
"score_source": "both",
|
||||
"when": "2026-09-01 14:30"
|
||||
}],
|
||||
"sessions_in_scope": 12,
|
||||
"semantic_status": "ready"
|
||||
}
|
||||
```
|
||||
|
||||
`semantic_status`:`disabled / consent_required / indexing / ready / partial / unavailable`。
|
||||
|
||||
## 五、Embedding 配置与隐私
|
||||
|
||||
配置走 `config.yaml` 的 `plugins.project-history-recall` 节:
|
||||
|
||||
```yaml
|
||||
plugins:
|
||||
project-history-recall:
|
||||
embedding:
|
||||
provider: "openai" # 或 "local"、"voyage" 等
|
||||
model: "text-embedding-3-small"
|
||||
base_url: "" # 可选
|
||||
chunk_size: 512 # tokens
|
||||
chunk_overlap: 64
|
||||
top_k: 20
|
||||
enabled: false # 首次需对话确认后开启
|
||||
```
|
||||
|
||||
密钥进 `.env`(`PROJECT_HISTORY_EMBEDDING_API_KEY`)。
|
||||
|
||||
**隐私确认链路:**
|
||||
1. `enabled: false` 时 `mode=hybrid` 或 `mode=semantic` 返回 `consent_required`,Agent 在对话中说明。
|
||||
2. 用户确认后 Agent 调用配置命令写入 `enabled: true` + 服务信息。
|
||||
3. 授权记录存 `~/.hermes/plugins/project-history-recall/consent.json`:`{profile, projects, provider, endpoint, model, confirmed_at, scope: "history+incremental"}`。
|
||||
4. 换 provider/endpoint/扩大项目范围时重新确认。
|
||||
|
||||
## 六、check_fn
|
||||
|
||||
```python
|
||||
def check_available() -> bool:
|
||||
"""state.db 和 projects.db 至少一个存在即可注册(归属解析需要它们)。"""
|
||||
from hermes_state import _default_db_path
|
||||
return _default_db_path().parent.exists()
|
||||
```
|
||||
|
||||
工具注册后始终可见(`hermes tools` 里能看到),但调用时归属解析失败返回明确错误。
|
||||
|
||||
## 七、失败模式
|
||||
|
||||
| 场景 | 返回 |
|
||||
|---|---|
|
||||
| state.db 不存在 | `{"success": false, "error": "session database not found", "code": "no_db"}` |
|
||||
| projects.db 不存在 | 退化为 auto repo 归属(用 git_repo_root),不报错 |
|
||||
| 当前会话无 cwd 且无 git_repo_root | `{"success": true, "results": [], "message": "No project context (no working directory recorded)", "sessions_in_scope": 0}` |
|
||||
| 归属匹配到 0 个其他会话 | `{"success": true, "results": [], "message": "No other sessions in this project", "sessions_in_scope": 1}` |
|
||||
| FTS5 查询语法错误 | 回退到 LIKE 查询,不报错 |
|
||||
| embedding 未启用 | `semantic_status: "disabled"`,关键词结果正常返回 |
|
||||
| embedding 未确认 | `semantic_status: "consent_required"`,关键词结果正常返回 |
|
||||
| embedding 服务不可用 | `semantic_status: "unavailable"`,关键词结果正常返回 |
|
||||
| 移动目标项目不存在 | `{"success": false, "error": "project not found: p_xxx"}` |
|
||||
| 移动目标路径不存在 | `{"success": false, "error": "project folder not found: /path"}` |
|
||||
|
||||
## 八、与核心的交互边界
|
||||
|
||||
| 插件做的 | 插件不做的 |
|
||||
|---|---|
|
||||
| 读 state.db(SELECT only,除 update_session_cwd) | 写 messages 表、改 state.db schema |
|
||||
| 读 projects.db(SELECT only) | 写 projects.db |
|
||||
| 调用 SessionDB.update_session_cwd(核心公共接口) | 直接 UPDATE sessions 表 |
|
||||
| 自己的 embeddings.db | 改核心任何文件 |
|
||||
| 注册两个工具到 project toolset | 修改 session_search 或 inline executor |
|
||||
|
||||
## 九、实施顺序
|
||||
|
||||
| 步骤 | 内容 | 验证 |
|
||||
|---|---|---|
|
||||
| B0 | 验证 state.db 只读访问 + projects.db 读取 + update_session_cwd 可调用 | 临时 HERMES_HOME 下跑通 |
|
||||
| P1 | attribution.py 归属解析纯函数 + 单测 | 显式/自动/Home 三态 + Windows 路径 |
|
||||
| P2 | project_history_search 关键词路径 + 注册 | P/Q 隔离、Home 空、排除当前会话、溯源链接 |
|
||||
| P3 | project_session_move + update_session_cwd 集成 | 移动后新检索用新归属、cwd 已切换 |
|
||||
| P4 | embeddings.db + 语义路径 + RRF 混合 | 同义词命中、关键词保底、移动不重算向量 |
|
||||
| P5 | 隐私确认链路 | 未确认不发送、确认后自主检索 |
|
||||
|
||||
## 十、不做的事
|
||||
|
||||
- 不改核心 session_search、inline executor、SessionDB schema、project_tree.py。
|
||||
- 不做项目摘要、会话摘要、自动上下文注入。
|
||||
- 不做旧会话自动归类。
|
||||
- 不做权限撤销协议、版本表、跨 profile 联合检索。
|
||||
- 不动全局 MEMORY/USER。
|
||||
|
||||
## 十一、评审自审
|
||||
|
||||
- **V1 的 workspace.move 错误已修正:** 插件不调 JSON-RPC,改用核心公共接口 update_session_cwd。
|
||||
- **归属 = cwd 的函数:** 不加列、不改 schema,核心零改动。归属随 cwd 变化是接受的行为。
|
||||
- **只读约束:** 除 update_session_cwd 外,插件对 state.db 只有 SELECT。
|
||||
- **执行水准:** 工具 schema、归属输出形状、SQL WHERE 子句、失败模式、check_fn 全部定死。
|
||||
- **诚实边界:** 语义检索的 chunk_size/overlap/top_k/RRF-k 是起始值,P4 实测后可能调整;embedding 模型选型待 B0 确认。
|
||||
@@ -0,0 +1,375 @@
|
||||
# 项目历史召回插件设计方案 V3
|
||||
|
||||
> **状态:** 可执行方案。基于 V2 + 探针验证结果,所有接口、SQL、参数、失败模式已定死。
|
||||
|
||||
**Goal:** 一个插件解决两个需求:(1) 会话项目归属(cwd 跟随),(2) 项目内历史的按需关键词+语义检索。
|
||||
|
||||
**定位:** 独立插件 `~/.hermes/plugins/project-history-recall/`,核心零改动。
|
||||
|
||||
**探针验证结论(已跑通):**
|
||||
- 插件发现/注册/调用全链路 ✅
|
||||
- state.db 只读访问 ✅(1097 会话,656 有 cwd,441 Home 桶)
|
||||
- projects.db 读取 ✅(17 个显式项目)
|
||||
- SessionDB acquire + update_session_cwd ✅
|
||||
- FTS5 messages_fts 查询 ✅
|
||||
- messages 表含 `active`/`compacted` 列(可见性过滤用)
|
||||
|
||||
---
|
||||
|
||||
## 一、插件文件结构
|
||||
|
||||
```
|
||||
~/.hermes/plugins/project-history-recall/
|
||||
├── plugin.yaml
|
||||
├── __init__.py # register(ctx) — 注册两个工具
|
||||
├── attribution.py # 归属解析(纯函数)
|
||||
├── search.py # 关键词+语义混合检索
|
||||
├── store.py # state.db 只读查询 + update_session_cwd 封装
|
||||
├── embeddings.py # 向量索引(embeddings.db)
|
||||
└── consent.py # 隐私确认记录
|
||||
```
|
||||
|
||||
## 二、归属解析(attribution.py)
|
||||
|
||||
### 2.1 数据画像(探针实测)
|
||||
|
||||
| 类别 | 数量 | 占比 |
|
||||
|---|---|---|
|
||||
| 有 cwd | 656 | 60% |
|
||||
| 有 git_repo_root | 75 | 7% |
|
||||
| Home 桶(无 cwd) | 441 | 40% |
|
||||
|
||||
**结论:** Home 桶不是边缘情况,是主要路径之一。归属解析必须优雅处理。
|
||||
|
||||
### 2.2 归属解析函数
|
||||
|
||||
```python
|
||||
def resolve_project(session_row: dict, projects: list) -> dict:
|
||||
"""三态归属解析。
|
||||
|
||||
Returns:
|
||||
{"kind": "explicit", "id": "p_abc", "name": "...", "folders": [...]}
|
||||
{"kind": "auto", "key": "path_key", "root": "/path/to/repo"}
|
||||
{"kind": "home"}
|
||||
"""
|
||||
```
|
||||
|
||||
规则:
|
||||
1. **显式项目匹配:** 用 session 的 `cwd` 和 `git_repo_root` 分别与每个项目的 `folders` 做最长祖先匹配。路径规范化:去尾部 `/`,Windows casefold。取匹配深度最深的项目。
|
||||
2. **自动 repo:** 无显式命中但有 `git_repo_root` → `{"kind": "auto", "key": path_key(root), "root": root}`。
|
||||
3. **Home 桶:** 无 cwd 且无 git_repo_root → `{"kind": "home"}`。
|
||||
|
||||
### 2.3 归属过滤 SQL 生成
|
||||
|
||||
```python
|
||||
def scope_sql(identity: dict) -> tuple[str, tuple]:
|
||||
"""返回 (WHERE 片段, 参数)。用于 JOIN sessions 时的项目过滤。
|
||||
|
||||
explicit: 多 folder OR 条件
|
||||
auto: git_repo_root 精确匹配
|
||||
home: 无过滤(返回空结果,由调用方处理)
|
||||
"""
|
||||
```
|
||||
|
||||
显式项目的 WHERE(每个 folder 生成 4 个条件):
|
||||
|
||||
```sql
|
||||
-- 每个 folder 生成:
|
||||
(s.cwd LIKE ? ESCAPE '\' OR s.cwd = ? OR s.git_repo_root LIKE ? ESCAPE '\' OR s.git_repo_root = ?)
|
||||
-- 参数:(folder + '/%', folder, folder + '/%', folder)
|
||||
-- 多个 folder 用 OR 连接
|
||||
```
|
||||
|
||||
LIKE 转义:`\` → `\\`,`%` → `\%`,`_` → `\_`。
|
||||
|
||||
## 三、检索(search.py)
|
||||
|
||||
### 3.1 关键词路径
|
||||
|
||||
```sql
|
||||
SELECT m.id, m.session_id, m.role, m.content, m.timestamp,
|
||||
s.title, s.started_at, s.cwd, s.git_repo_root
|
||||
FROM messages_fts f
|
||||
JOIN messages m ON f.rowid = m.id
|
||||
JOIN sessions s ON m.session_id = s.id
|
||||
WHERE messages_fts MATCH ?
|
||||
AND m.role IN ('user', 'assistant')
|
||||
AND m.active = 1
|
||||
AND s.archived = 0 AND s.hidden = 0
|
||||
AND s.id != ? -- 排除当前会话
|
||||
AND {scope_where} -- 项目过滤
|
||||
ORDER BY rank
|
||||
LIMIT ?
|
||||
```
|
||||
|
||||
参数顺序:`(query, current_session_id, *scope_params, limit)`。
|
||||
|
||||
**CJK 回退:** FTS5 对 CJK 不友好。查询含 CJK 字符时回退到 LIKE:
|
||||
|
||||
```sql
|
||||
SELECT m.id, m.session_id, m.role, m.content, m.timestamp,
|
||||
s.title, s.started_at, s.cwd, s.git_repo_root
|
||||
FROM messages m
|
||||
JOIN sessions s ON m.session_id = s.id
|
||||
WHERE m.content LIKE ? ESCAPE '\'
|
||||
AND m.role IN ('user', 'assistant')
|
||||
AND m.active = 1
|
||||
AND s.archived = 0 AND s.hidden = 0
|
||||
AND s.id != ?
|
||||
AND {scope_where}
|
||||
ORDER BY m.timestamp DESC
|
||||
LIMIT ?
|
||||
```
|
||||
|
||||
**FTS5 查询消毒:** 用现有 `_sanitize_fts5_query` 的规则——引号包裹、OR/AND/NOT 保留、`*` 前缀匹配。语法错误时回退 LIKE。
|
||||
|
||||
### 3.2 归属会话集合
|
||||
|
||||
```python
|
||||
def get_project_sessions(db_conn, identity: dict, current_session_id: str) -> list[dict]:
|
||||
"""返回项目内所有会话的元数据(排除当前会话和隐藏来源)。"""
|
||||
```
|
||||
|
||||
用于语义索引时确定"需要索引哪些会话",也用于结果中的 `sessions_in_scope`。
|
||||
|
||||
### 3.3 语义路径(embeddings.py)
|
||||
|
||||
**存储:** `~/.hermes/plugins/project-history-recall/embeddings.db`
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS chunks (
|
||||
id INTEGER PRIMARY KEY,
|
||||
session_id TEXT NOT NULL,
|
||||
message_id INTEGER NOT NULL,
|
||||
chunk_index INTEGER NOT NULL,
|
||||
content_text TEXT NOT NULL, -- 原文片段(用于调试和 snippet)
|
||||
content_hash TEXT NOT NULL, -- SHA-256 hex
|
||||
model_fingerprint TEXT NOT NULL, -- provider:model:dimensions
|
||||
vector BLOB NOT NULL, -- float32 array
|
||||
created_at REAL NOT NULL,
|
||||
UNIQUE(session_id, message_id, chunk_index, model_fingerprint)
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS idx_chunks_session ON chunks(session_id);
|
||||
```
|
||||
|
||||
**切片参数(起始值,P4 实测调整):**
|
||||
|
||||
| 参数 | 起始值 | 说明 |
|
||||
|---|---|---|
|
||||
| chunk_size | 512 chars | 字符数,不是 token(简单可靠) |
|
||||
| chunk_overlap | 64 chars | 重叠 |
|
||||
| top_k | 20 | 语义检索候选数 |
|
||||
| rrf_k | 60 | RRF 标准常数 |
|
||||
| batch_size | 32 | embedding API 批量 |
|
||||
|
||||
**索引流程:**
|
||||
1. `get_project_sessions()` 得到项目会话集合。
|
||||
2. 对每个会话,读 `active=1` 的 user/assistant 消息。
|
||||
3. 每条消息按 chunk_size 切片,计算 content_hash。
|
||||
4. 跳过已有且 hash 未变的 chunk。
|
||||
5. 批量调用 embedding API,存向量。
|
||||
6. 返回索引统计(新增/跳过/失败数)。
|
||||
|
||||
**检索流程:**
|
||||
1. 归属解析 → 项目会话 ID 集合。
|
||||
2. 查询 embedding API 得到查询向量。
|
||||
3. 从 embeddings.db 读该集合内所有 chunk(按 session_id IN 过滤)。
|
||||
4. 计算余弦相似度(numpy 或纯 Python,B0 决定)。
|
||||
5. 取 top_k,按 message_id 归并。
|
||||
|
||||
**模型指纹:** `{provider}:{model}:{dimensions}`,如 `openai:text-embedding-3-small:1536`。不同指纹的向量不可混用。
|
||||
|
||||
### 3.4 混合检索(RRF)
|
||||
|
||||
```python
|
||||
def rrf_merge(keyword_results: list[dict], semantic_results: list[dict],
|
||||
k: int = 60) -> list[dict]:
|
||||
"""Reciprocal Rank Fusion. 按 message_id 归并两路结果。
|
||||
|
||||
score = 1/(k + rank_keyword) + 1/(k + rank_semantic)
|
||||
只出现在一路的结果,另一路 rank = infinity(贡献 0)。
|
||||
"""
|
||||
```
|
||||
|
||||
归并后按 RRF 分数降序,取 top `limit` 条。每条的 `score_source` 标注 `"keyword" / "semantic" / "both"`。
|
||||
|
||||
### 3.5 结果展开
|
||||
|
||||
对每个命中消息,读前后各 2 条消息作为上下文:
|
||||
|
||||
```sql
|
||||
SELECT id, role, content, timestamp FROM messages
|
||||
WHERE session_id = ? AND id BETWEEN ? - 2 AND ? + 2 AND active = 1
|
||||
ORDER BY id
|
||||
```
|
||||
|
||||
返回时包含 `session_id`、`message_id`、`link`、`title`、`snippet`(命中消息内容,截断 500 字符)、`context`(前后消息列表)、`when`、`score_source`。
|
||||
|
||||
## 四、移动(store.py)
|
||||
|
||||
```python
|
||||
def move_session(session_id: str, project_id: str, **kw) -> dict:
|
||||
"""移动会话到项目。cwd 跟随到项目 primary_path。"""
|
||||
```
|
||||
|
||||
逻辑:
|
||||
1. `project_id` 为空 → 移到 Home:cwd 设为 `~`。
|
||||
2. 非空 → 从 projects.db 读目标项目的 `primary_path`。
|
||||
3. git 探测:`git_probe.branch(path)` + `git_probe.common_repo_root(path)`。
|
||||
4. 调用 `SessionDB.update_session_cwd(session_id, path, branch, root, replace_git_meta=True)`。
|
||||
5. 返回 `{"success": true, "project": {...}, "cwd": path}`。
|
||||
|
||||
**获取 SessionDB:** `hermes_state_registry.acquire()`(探针已验证可用)。
|
||||
|
||||
## 五、隐私确认(consent.py)
|
||||
|
||||
**存储:** `~/.hermes/plugins/project-history-recall/consent.json`
|
||||
|
||||
```json
|
||||
{
|
||||
"version": 1,
|
||||
"grants": [{
|
||||
"profile": "default",
|
||||
"provider": "openai",
|
||||
"endpoint": "https://api.openai.com/v1",
|
||||
"model": "text-embedding-3-small",
|
||||
"projects": ["p_abc", "p_def"],
|
||||
"scope": "history+incremental",
|
||||
"confirmed_at": 1757486400.0
|
||||
}]
|
||||
}
|
||||
```
|
||||
|
||||
**确认流程:**
|
||||
1. `mode=semantic` 或 `mode=hybrid` 且 embedding 未配置 → 返回 `semantic_status: "consent_required"`。
|
||||
2. Agent 在对话中说明服务方、发送内容、项目范围。
|
||||
3. 用户确认后,Agent 写入 consent.json + 配置 embedding 参数。
|
||||
4. 后续同 provider+endpoint+项目范围不重复确认。
|
||||
5. 换 provider/endpoint/扩大项目范围 → 重新确认。
|
||||
|
||||
## 六、check_fn
|
||||
|
||||
```python
|
||||
def check_available() -> bool:
|
||||
"""state.db 存在即可注册。"""
|
||||
try:
|
||||
from hermes_state import _default_db_path
|
||||
return _default_db_path().exists()
|
||||
except Exception:
|
||||
return False
|
||||
```
|
||||
|
||||
## 七、工具注册(__init__.py)
|
||||
|
||||
```python
|
||||
def register(ctx) -> None:
|
||||
ctx.register_tool(
|
||||
name="project_history_search", toolset="project",
|
||||
schema=SEARCH_SCHEMA, handler=handle_search,
|
||||
check_fn=check_available, emoji="📂",
|
||||
)
|
||||
ctx.register_tool(
|
||||
name="project_session_move", toolset="project",
|
||||
schema=MOVE_SCHEMA, handler=handle_move,
|
||||
check_fn=check_available, emoji="📁",
|
||||
)
|
||||
```
|
||||
|
||||
## 八、完整工具 schema
|
||||
|
||||
### project_history_search
|
||||
|
||||
```python
|
||||
SEARCH_SCHEMA = {
|
||||
"name": "project_history_search",
|
||||
"description": (
|
||||
"Search past conversations in the SAME project as this one. "
|
||||
"Returns actual messages from other sessions with links. "
|
||||
"Modes: 'hybrid' (default, keyword+semantic), 'keyword' (exact terms), "
|
||||
"'semantic' (similar meaning). Use 'keyword' for code symbols/error codes; "
|
||||
"'semantic' for 'how did we handle X' when exact words differ."
|
||||
),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"query": {
|
||||
"type": "string",
|
||||
"description": "Search query. Keywords, phrases, or natural language.",
|
||||
},
|
||||
"mode": {
|
||||
"type": "string",
|
||||
"enum": ["hybrid", "keyword", "semantic"],
|
||||
"default": "hybrid",
|
||||
},
|
||||
"limit": {
|
||||
"type": "integer",
|
||||
"default": 3,
|
||||
"minimum": 1,
|
||||
"maximum": 10,
|
||||
},
|
||||
},
|
||||
"required": ["query"],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### project_session_move
|
||||
|
||||
```python
|
||||
MOVE_SCHEMA = {
|
||||
"name": "project_session_move",
|
||||
"description": (
|
||||
"Move this conversation to a different project. The working directory "
|
||||
"switches to the target project's folder. Use when the user says "
|
||||
"'move this to project X' or 'this belongs to project Y'."
|
||||
),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"project_id": {
|
||||
"type": "string",
|
||||
"description": (
|
||||
"Target project ID or slug (e.g. 'p_abc123' or 'hermes-agent'). "
|
||||
"Use '' to remove from any project."
|
||||
),
|
||||
},
|
||||
},
|
||||
"required": ["project_id"],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## 九、失败模式
|
||||
|
||||
| 场景 | 返回 |
|
||||
|---|---|
|
||||
| state.db 不存在 | `{"success": false, "error": "session database not found"}` |
|
||||
| projects.db 不存在 | 退化为 auto repo 归属,不报错 |
|
||||
| 当前会话无 cwd 且无 git_repo_root | `{"success": true, "results": [], "sessions_in_scope": 0, "message": "No project context"}` |
|
||||
| 项目内无其他会话 | `{"success": true, "results": [], "sessions_in_scope": 1, "message": "No other sessions in this project"}` |
|
||||
| FTS5 语法错误 | 回退 LIKE,不报错 |
|
||||
| embedding 未启用 | `semantic_status: "disabled"`,关键词正常返回 |
|
||||
| embedding 未确认 | `semantic_status: "consent_required"`,关键词正常返回 |
|
||||
| embedding 服务不可用 | `semantic_status: "unavailable"`,关键词正常返回 |
|
||||
| 移动目标项目不存在 | `{"success": false, "error": "project not found: {id}"}` |
|
||||
| 移动目标路径不存在 | `{"success": false, "error": "folder not found: {path}"}` |
|
||||
| SessionDB acquire 失败 | `{"success": false, "error": "session database unavailable"}` |
|
||||
|
||||
## 十、实施顺序
|
||||
|
||||
| 步骤 | 内容 | 验证命令 |
|
||||
|---|---|---|
|
||||
| B0 | ✅ 已完成(探针验证) | 已跑通 |
|
||||
| P1 | attribution.py 归属解析 + 单测 | 显式/自动/Home 三态 + Windows 路径 |
|
||||
| P2 | search.py 关键词路径 + 注册 + 归属过滤 | P/Q 隔离、Home 空、排除当前、溯源链接 |
|
||||
| P3 | store.py 移动 + update_session_cwd | 移动后新检索用新归属、cwd 已切换 |
|
||||
| P4 | embeddings.py 语义路径 + RRF 混合 | 同义词命中、关键词保底、移动不重算向量 |
|
||||
| P5 | consent.py 隐私确认 | 未确认不发送、确认后自主检索 |
|
||||
|
||||
## 十一、评审自审
|
||||
|
||||
- **V2 的 SessionDB 获取问题已解决:** `hermes_state_registry.acquire()` 探针验证可用。
|
||||
- **归属 = cwd 的函数:** 无持久化列,核心零改动。归属随 cwd 变化是接受的行为。
|
||||
- **只读约束:** 除 `update_session_cwd` 外,state.db 只有 SELECT。
|
||||
- **执行水准:** schema、SQL、参数、失败模式、check_fn 全部定死。
|
||||
- **诚实边界:** chunk_size/overlap/top_k/rrf_k 是起始值,P4 实测调整;embedding 模型选型待 B0 确认;40% 的 Home 桶会话无法检索是已知限制。
|
||||
@@ -0,0 +1,177 @@
|
||||
# 项目历史召回插件设计方案
|
||||
|
||||
> **状态:** 设计方案,非实施代码。本文件定义插件的能力边界与集成方式;代码实现需另行授权。
|
||||
|
||||
**Goal:** 用一个插件同时解决两个需求:(1) 会话项目归属(project_id),(2) 项目内历史的按需关键词+语义检索。
|
||||
|
||||
**定位:** Footprint Ladder 第 4 级 —— 独立插件,不动核心文件,不落 `plugins/memory/`(该目录已关闭新 provider)。装在 `~/.hermes/plugins/project-history-recall/` 或独立仓库 pip 包。
|
||||
|
||||
---
|
||||
|
||||
## 一、为什么用一个插件而不是两个
|
||||
|
||||
两个需求共享同一套底座:会话→项目的归属解析 + state.db 只读访问 + 检索结果格式化。拆成两个插件会重复这份底座并引入"哪个插件负责归属"的模糊边界。一个插件、两个工具,归属逻辑写一次,两个工具都吃它。
|
||||
|
||||
**与 V3/V4.1 的关系:** 本插件方案替代它们的"改核心 session_search"路径。V3/V4.1 的归属语义(project_id、移动只改归属、cwd 分离)和检索设计(项目过滤下推、混合检索、溯源)全部保留,但**落点从核心 session_search 改为插件的独立工具**,核心零改动。
|
||||
|
||||
## 二、插件结构与发现
|
||||
|
||||
```
|
||||
~/.hermes/plugins/project-history-recall/
|
||||
├── plugin.yaml # name/version/description,不动核心
|
||||
├── __init__.py # register(ctx):注册两个工具
|
||||
├── attribution.py # 归属解析(读 projects.db + sessions 行)
|
||||
├── search.py # 关键词+语义混合检索
|
||||
└── config_schema.py # embedding 配置项声明(如有)
|
||||
```
|
||||
|
||||
发现路径:`PluginManager`(`hermes_cli/plugins.py`)扫描 `~/.hermes/plugins/`,later-wins 发现 `register(ctx)`。
|
||||
|
||||
注册方式(对齐 `plugins/spotify/__init__.py:28-31` 的现有模式):
|
||||
|
||||
```python
|
||||
def register(ctx) -> None:
|
||||
ctx.register_tool(name="project_history_search", toolset="project",
|
||||
schema=SEARCH_SCHEMA, handler=handle_search,
|
||||
check_fn=check_available, emoji="📂")
|
||||
ctx.register_tool(name="project_session_move", toolset="project",
|
||||
schema=MOVE_SCHEMA, handler=handle_move,
|
||||
check_fn=check_available, emoji="📁")
|
||||
```
|
||||
|
||||
两个工具都在 `project` toolset(会话作用域,桌面端 GUI 平台自动折叠进可用集合,不依赖 `HERMES_DESKTOP` 环境变量)。
|
||||
|
||||
## 三、需求一:项目归属(project_id 语义)
|
||||
|
||||
**归属数据源:** 复用现有 `projects.db`(`hermes_cli/projects_db.py`)+ 会话行的 `cwd`/`git_repo_root`(`hermes_state_common.py:305-307`)。
|
||||
|
||||
**归属解析(attribution.py):** 纯函数,无 IO 写入。
|
||||
|
||||
1. 读 `projects.db` 的 `projects` + `project_folders` → 构建文件夹索引(最长祖先匹配,`_FolderIndex` 同款逻辑,插件内重新实现,不从 tui_gateway 复制——它是显示层,不是共享内核)。
|
||||
2. 当前会话行 → `cwd`/`git_repo_root` → 匹配显式项目文件夹 → `Project.id`。
|
||||
3. 无显式项目命中 → 用 `git_repo_root` 的 path_key 作为 auto 项目身份(Tier 2)。
|
||||
4. 无 cwd 且无 git_repo_root → **Home 桶**,返回明确空结果,不检索。
|
||||
|
||||
**移动(project_session_move 工具):** 更新会话的归属记录。归属与 cwd 一致移动——会话移到哪个项目,工作目录就跟到哪个项目的目录。通过现有 `session.workspace.move` RPC 完成(它已有持久化 cwd/git_repo_root 的完整路径,见 `tui_gateway/methods_session.py:840-876`),插件在其上叠加归属语义(校验目标项目存在、写入 project_id)。核心零改动。
|
||||
|
||||
**归属语义(保留 V3 合同,cwd 跟随移动):**
|
||||
- 移动会话 = 归属与 cwd 一起切到目标项目目录(workspace.move 的现有行为,插件叠加归属校验)。
|
||||
- 移动前开始的检索可以完成;移动后的新检索用新归属和新 cwd。
|
||||
- Home 桶(无 cwd)不检索,返回空并说明原因。
|
||||
- 跨 profile 归属不合并。
|
||||
|
||||
## 四、需求二:项目内历史检索(关键词 + 语义)
|
||||
|
||||
**工具:** `project_history_search`,两个模式由 `mode` 参数控制(默认 `hybrid`)。
|
||||
|
||||
### 4.1 关键词路径
|
||||
|
||||
插件直接只读查询 state.db(`SessionDB` 的只读实例,或 `sqlite3.connect(..., uri=True)` 只读模式):
|
||||
|
||||
```
|
||||
SELECT m.id, m.session_id, m.role, m.content, m.timestamp, s.title, s.started_at
|
||||
FROM messages m JOIN sessions s ON m.session_id = s.id
|
||||
WHERE s.<归属匹配条件> AND m.content LIKE / FTS5 MATCH ...
|
||||
ORDER BY rank / timestamp
|
||||
LIMIT ?
|
||||
```
|
||||
|
||||
归属匹配条件由 attribution.py 解析出的项目会话集合生成(`s.project_id = ?` 或 `s.cwd LIKE '...%'` 或 `s.git_repo_root = ?`,取决于归属是否已持久化)。
|
||||
|
||||
**只读约束是硬性的:** 插件对 state.db 只有 SELECT,绝不 INSERT/UPDATE/DELETE。state.db 是核心的权威存储,插件写入会破坏核心的一致性假设(FTS 索引、token 计数、lineage 状态)。
|
||||
|
||||
### 4.2 语义路径
|
||||
|
||||
向量索引为插件自有派生数据,存在插件自己的 SQLite 文件(`~/.hermes/plugins/project-history-recall/embeddings.db`)——不碰 state.db schema。
|
||||
|
||||
**索引构建:**
|
||||
- 按归属解析出项目会话集合 → 读 messages 的 user/assistant 文本 → 按消息切片(长消息分段)→ 调用 embedding 模型 → 存向量。
|
||||
- 首次索引按项目分批执行;后续增量(新消息)在每次检索前检查并补建。
|
||||
- 索引记录 `session_id + message_id + chunk_offset + content_hash + model_fingerprint + vector`。
|
||||
- 移动会话后向量不重算——归属解析在查询时重新执行,向量只关联 session_id。
|
||||
|
||||
**检索流程:**
|
||||
1. 归属解析 → 项目会话集合。
|
||||
2. 关键词:state.db FTS5 查询,限定在集合内。
|
||||
3. 语义:embeddings.db 向量查询,限定在集合内(先 JOIN 过滤再算相似度,不全库 Top-K 后过滤)。
|
||||
4. 混合:Reciprocal Rank Fusion 按消息锚点归并两路结果。
|
||||
5. 展开:对命中消息用 state.db 读原文窗口(前后各 N 条),返回 `session_id + message_id + @session:<id> 链接 + 原文片段`。
|
||||
|
||||
**模式:**
|
||||
- `hybrid`(默认):关键词 + 语义,语义不可用时退化为关键词并注明。
|
||||
- `keyword`:仅关键词,用于精确符号/错误码查询。
|
||||
- `semantic`:仅语义,用于不同措辞的相似问题追溯。
|
||||
|
||||
### 4.3 结果格式
|
||||
|
||||
JSON 字符串返回(对齐现有工具约定):
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"mode": "hybrid",
|
||||
"project": {"kind": "explicit", "id": "p_abc", "name": "Hermes"},
|
||||
"results": [{
|
||||
"session_id": "20260901_143022_a1b2",
|
||||
"link": "@session:20260901_143022_a1b2",
|
||||
"title": "session_search 项目过滤设计",
|
||||
"message_id": 47,
|
||||
"snippet": "……项目归属应该持久化在 sessions 表……",
|
||||
"score_source": "both",
|
||||
"when": "2026-09-01 14:30"
|
||||
}],
|
||||
"sessions_in_scope": 12,
|
||||
"semantic_status": "ready"
|
||||
}
|
||||
```
|
||||
|
||||
`semantic_status`:`disabled / consent_required / indexing / ready / partial / unavailable`。
|
||||
|
||||
## 五、Embedding 配置与隐私
|
||||
|
||||
embedding 模型配置通过插件自己的配置文件(`~/.hermes/plugins/project-history-recall/config.yaml`)或 `config.yaml` 的 `plugins.project-history-recall` 节,走 `hermes tools` 的插件配置 UX。密钥进 `.env`。
|
||||
|
||||
**隐私(保留 V4.1 的对话确认设计):**
|
||||
- 首次需要远程 embedding 且无授权记录时,工具返回 `consent_required` 状态,不发送任何内容。
|
||||
- Agent 在对话中说明服务方、发送内容、项目范围,用户确认后写入配置。
|
||||
- 授权持久化,同范围后续不重复询问;换服务/加项目重新确认。
|
||||
- 未确认时关键词路径正常工作。
|
||||
|
||||
## 六、与核心的交互边界
|
||||
|
||||
| 插件做的 | 插件不做的 |
|
||||
|---|---|
|
||||
| 读 projects.db、读 state.db(只读) | 写 state.db 的 sessions/messages 表(移动除外——workspace.move 有核心自己的写路径) |
|
||||
| 自己的 embeddings.db(向量存储) | 改 state.db schema |
|
||||
| 注册两个工具到 `project` toolset | 修改核心 session_search 或 inline executor |
|
||||
| 归属解析纯函数 | 修改 project_tree.py 或桌面端代码 |
|
||||
| embedding 调用(自有配置) | 新增核心 HERMES_* 环境变量 |
|
||||
|
||||
**核心配合需求:无。** 移动会话归属与 cwd 一致移动,复用现有 `session.workspace.move` RPC 的完整路径(持久化 + git 元数据更新),插件只在其上叠加项目校验和归属语义。核心零改动。
|
||||
|
||||
## 七、实施顺序
|
||||
|
||||
| 步骤 | 内容 | 验证 |
|
||||
|---|---|---|
|
||||
| B0 | 验证 state.db 只读访问 + projects.db 读取 + workspace.move 的归属叠加路径 | 真实临时 HERMES_HOME 下跑通 |
|
||||
| P1 | attribution.py 归属解析纯函数 + 单测 | 显式项目/自动 repo/Home 桶三态 |
|
||||
| P2 | project_history_search 关键词路径 + project toolset 注册 | P/Q 项目隔离、Home 桶空结果、溯源链接 |
|
||||
| P3 | project_session_move 移动 + 归属更新 | 移动后新检索用新归属 |
|
||||
| P4 | embeddings.db + 语义路径 + 混合检索 | 同义词命中、关键词保底、移动不重算向量 |
|
||||
| P5 | 对话确认链路 + 隐私控制 | 未确认不发送、确认后自主检索 |
|
||||
|
||||
## 八、不做的事
|
||||
|
||||
- 不改核心 session_search、inline executor、SessionDB schema。
|
||||
- 不做项目摘要、会话摘要、自动上下文注入。
|
||||
- 不做旧会话自动目录归类。
|
||||
- 不做权限撤销协议、版本表、跨 profile 联合检索。
|
||||
- 不动全局 MEMORY/USER 的项目隔离。
|
||||
|
||||
## 九、评审自审
|
||||
|
||||
- **Footprint Ladder:** 第 4 级插件,核心零改动,符合"能力在边缘"。
|
||||
- **一个插件两个工具:** 归属底座共享,不重复。
|
||||
- **只读 state.db:** 插件绝不写核心存储,向量索引自有。
|
||||
- **prompt cache:** 工具结果通道,不碰系统提示词。
|
||||
- **诚实边界:** 移动时 cwd 跟随归属是已确认的行为(workspace.move 现有路径),无悬而未决项;语义检索的性能包络待 P4 实测,不预先承诺。
|
||||
+77
@@ -0,0 +1,77 @@
|
||||
# Architecture Decision Records
|
||||
|
||||
## 2026-07-13: Scope plugin manager state by Hermes home/profile (keyed cache)
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Context:
|
||||
Hermes supports multiple profiles via different Hermes home directories.
|
||||
Homes are switched two ways in a running process: the `HERMES_HOME`
|
||||
environment variable (single-profile CLI/gateway processes), and the
|
||||
context-local `set_hermes_home_override()` (`hermes_constants.py`), which
|
||||
the multiplexed gateway worker (`gateway/run.py`'s `_profile_scope`) and
|
||||
subagent/embedded callers use to serve several profiles from one
|
||||
long-lived process. The override is a `ContextVar` and deliberately does
|
||||
**not** mutate `os.environ`, since that would leak one profile's home
|
||||
into every other concurrent task in the same process.
|
||||
|
||||
The plugin manager was a process-global single-slot singleton
|
||||
(`_plugin_manager`). User-installed plugins are discovered from
|
||||
`get_hermes_home() / "plugins"`, and context-engine plugins (e.g.
|
||||
`hermes-lcm`) capture profile-scoped state — such as the LCM database
|
||||
path — at registration time. A single-slot cache meant:
|
||||
|
||||
1. Switching homes via `set_hermes_home_override()` was invisible to a
|
||||
naive "did `HERMES_HOME` change" check, so the singleton silently kept
|
||||
serving the first profile's manager to every other profile in the
|
||||
process.
|
||||
2. Even when a fresh `PluginManager` *was* created for a new home, plugin
|
||||
modules are imported into `sys.modules` as `hermes_plugins.<slug>` by
|
||||
`_load_directory_module`, and only that top-level module was ever
|
||||
replaced. A same-slug plugin's *relative* imports
|
||||
(`from . import state`) are cached separately under
|
||||
`hermes_plugins.<slug>.<submodule>`, and Python's import machinery
|
||||
resolves those from `sys.modules` first — so a profile switch could
|
||||
silently keep serving a previous profile's already-imported submodule
|
||||
code/state instead of re-executing the new profile's plugin.
|
||||
|
||||
Decision:
|
||||
- Replace the single-slot singleton with a cache keyed on the *resolved*
|
||||
Hermes home path (`_plugin_managers_by_home: Dict[Path, PluginManager]`).
|
||||
`get_plugin_manager()` resolves the current home via `get_hermes_home()`
|
||||
(which itself already consults `get_hermes_home_override()` before
|
||||
`os.environ`), so both the env-var and context-local override paths are
|
||||
covered uniformly.
|
||||
- `_plugin_manager` (the old single-slot name) is kept as a thin "last
|
||||
manager returned" pointer purely for backward compatibility with
|
||||
existing test code that does
|
||||
`monkeypatch.setattr(plugins_mod, "_plugin_manager", some_manager)`.
|
||||
When that name is monkeypatched to a manager the keyed cache doesn't
|
||||
know about, `get_plugin_manager()` treats it as an explicit injection
|
||||
and adopts it into the cache under the *current* resolved home, rather
|
||||
than discarding it.
|
||||
- Both `PluginManager._load_directory_module` (initial/`force=True`
|
||||
reload within the same home) and the shared `_clear_plugin_submodules`
|
||||
helper (profile switch / test teardown) evict `sys.modules[module_name]`
|
||||
**and every name prefixed with `module_name + "."`** before a plugin
|
||||
slug is (re-)imported, so relative-import submodules can never survive
|
||||
a reload or a home switch.
|
||||
- Test isolation (`tests/conftest.py`'s `_hermetic_environment` fixture)
|
||||
calls a new `_reset_plugin_managers_for_tests()` helper that drops the
|
||||
entire keyed cache and purges every plugin submodule from `sys.modules`
|
||||
between tests, instead of only resetting the single-slot pointer.
|
||||
|
||||
Consequences:
|
||||
- Per-profile LCM instances (and any other context-engine plugin) use
|
||||
their own `{home}/lcm.db` regardless of whether the profile switch went
|
||||
through `HERMES_HOME` or `set_hermes_home_override()`.
|
||||
- Plugin discovery remains cached within a profile for normal
|
||||
performance, and re-entering a previously-seen profile reuses its
|
||||
cached manager instead of rebuilding from scratch.
|
||||
- Sequential *and* interleaved profile switching — in tests, the gateway
|
||||
multiplexer worker, or embedded callers using the context-local
|
||||
override — no longer leaks context-engine state, plugin module state,
|
||||
or stale relative-import submodules across profiles.
|
||||
- Regression coverage exercises the real production path
|
||||
(`set_hermes_home_override()`) rather than only the env-var path, and
|
||||
includes a dedicated relative-import leak test.
|
||||
@@ -0,0 +1,179 @@
|
||||
# Billing lifecycle: client-side state, errors, and recovery
|
||||
|
||||
This is the map from every `billing.*`/`subscription.*` state shape the gateway
|
||||
serves (from NAS) to what the terminal actually renders, and from every typed
|
||||
refusal/error code to its exact user-facing copy and recovery action. The
|
||||
guarantee: no NAS billing state and no typed refusal falls through to a
|
||||
generic toast — every case below is an explicit branch in
|
||||
`ui-tui/src/app/slash/commands/topup.ts`, `ui-tui/src/components/billingOverlay.tsx`,
|
||||
or `ui-tui/src/components/subscriptionOverlay.tsx`. An **unknown** code still
|
||||
degrades gracefully: it hits the `default` branch (a generic-but-real message
|
||||
pulled from the server payload, never a blank toast) rather than crashing or
|
||||
silently dropping the refusal.
|
||||
|
||||
## 1. `billing.state` shapes → render
|
||||
|
||||
Source: `ui-tui/src/components/billingOverlay.tsx` (`OverviewScreen`,
|
||||
`BuyScreen`, `AutoReloadScreen`), `ui-tui/src/app/slash/commands/topup.ts` (`/topup` run).
|
||||
|
||||
| State shape | Render |
|
||||
|---|---|
|
||||
| Logged out (`s.logged_in === false`) | Overlay never opens. `sys`: `💳 Not logged into Nous Portal — run /portal to log in, then /topup.` |
|
||||
| `billing.state` RPC fetch fails (transport/timeout) | **Fail-closed**: `.catch(ctx.guardedErr)` — overlay never opens, no state is assumed. `sys`: `error: <message or "request failed">`. Never renders "no card" or any other guessed state; user must retry `/topup`. |
|
||||
| `card: null` (no saved card), full menu (`is_admin && cli_billing_enabled`) | Overview shows `No saved card on file — "Add funds" walks you through adding one.` "Add funds" opens the **add-card path**: `Add a card on the portal` / `I've added it — check again` / `Back` (never an amount picker, which would 403 `no_payment_method`). |
|
||||
| `card` present, `resolved_via` set | `Card: {display}` (e.g. `Visa ····4242 — the card on your subscription`) using the provenance-aware `display` field. |
|
||||
| `card` present, `resolved_via` absent (older NAS) | Falls back to the generic `Card: {masked}`; Confirm screen adds `Your card saved on the portal will be charged.` |
|
||||
| `auto_reload: null` | No auto-reload line at all (`autoReloadLine` returns `null`) — the feature isn't surfaced. |
|
||||
| `auto_reload.card.kind: 'canonical'` | No distinct-card warning; card line falls back to the card on file. |
|
||||
| `auto_reload.card.kind: 'distinct'` | `⚠ Auto-refill is charging {brand} ••{last4} — not your card on file.` in the Auto-reload screen (the divergence notice). |
|
||||
| `auto_reload.card.kind: 'none'` | Same as `canonical` rendering-wise — no distinct-card warning shown. |
|
||||
| `monthly_cap` present, `limit_usd != null` | `{spent_display} of {limit_display} used this month` (+ ` (default ceiling)` iff `is_default_ceiling`). |
|
||||
| `monthly_cap` absent or `limit_usd == null` | `No monthly cap visible (managed on the portal).` |
|
||||
| Role without billing capability (`!is_admin`, menu collapses) | Note: `Billing actions need someone with billing permissions (owner, admin, or finance admin).` Menu collapses to `Manage on portal` / `Cancel`. |
|
||||
| Org kill-switch off (`is_admin` but `!cli_billing_enabled`) | Note: `Remote spending is off for this org — a billing admin can turn it on from the portal's Hermes Agent page.` Same collapsed menu. |
|
||||
|
||||
Note: `full = s.is_admin && s.cli_billing_enabled` gates the **org-level**
|
||||
switch, not the per-terminal `billing:manage` scope — that's discovered
|
||||
reactively (a charge 403s `insufficient_scope`) and routes to the resumable
|
||||
step-up screen instead of a preflight check.
|
||||
|
||||
## 2. Refusal codes (`renderBillingError`, in code order)
|
||||
|
||||
Source: `renderBillingError` in `ui-tui/src/app/slash/commands/topup.ts:37-149`.
|
||||
"Portal" row = `sys('Portal: {portal_url}')` is appended whenever `portal_url` is present, for every code (including default).
|
||||
|
||||
| `error` code | Copy | Portal URL | `retry_after` |
|
||||
|---|---|:-:|:-:|
|
||||
| `insufficient_scope` | `This needs Remote Spending allowed. Start a top-up to allow it, then retry.` | if present | — |
|
||||
| `remote_spending_revoked` (CF-4) | `{An admin stopped remote spending for this terminal. \| You stopped remote spending for this terminal.}` (by `actor`) `Reconnect to restore — run /portal to re-authorize this terminal.` Also clears `billing` overlay state immediately (doesn't wait for token refresh). | if present | — |
|
||||
| `session_revoked` | `Your session was logged out. Run /portal to log in again.` Also clears `billing` overlay state. | if present | — |
|
||||
| `cli_billing_disabled` / `remote_spending_disabled` (dual-emitted) | `Remote spending is off for this account — a billing admin can turn it on from the portal's Hermes Agent page.` | if present | — |
|
||||
| `role_required` | `Adding funds needs someone with billing permissions (owner, admin, or finance admin), or manage this on the portal.` | if present | — |
|
||||
| `consent_required` | `This action needs a one-time card confirmation and consent step on the portal before it can proceed.` | if present | — |
|
||||
| `org_access_denied` | `This token isn't bound to an org you can manage. Sign in with the right org, or manage this on the portal.` | if present | — |
|
||||
| `upgrade_cap_exceeded` | `🔴 Daily plan-change limit reached (5 per org) — try again tomorrow, or manage this on the portal.` | if present | — |
|
||||
| `auto_top_up_disabled_failures` | `Auto-reload was turned off after repeated charge failures. Fix the card issue, then re-enable it from /topup → Auto-reload.` | if present | — |
|
||||
| `idempotency_conflict` | `🔴 That charge key was already used for a different amount. Start a fresh top-up.` | if present | — |
|
||||
| `no_payment_method` | `💳 No saved card for terminal charges yet. Set one up on the portal (one-time credit buys don't save a reusable card).` | if present | — |
|
||||
| `monthly_cap_exceeded` | `🔴 Monthly spend cap reached — ${remainingUsd} headroom left.` if `payload.remainingUsd` present, else `🔴 Monthly spend cap reached.` | if present | — |
|
||||
| `rate_limited` / `temporarily_unavailable` | `🟡 Too many charges right now{ (try again in ~N min)}. This isn't a payment failure.` | if present | **yes** — minutes computed as `max(1, round(retry_after/60))` |
|
||||
| `stripe_unavailable` | `🟡 Stripe is having trouble right now — try again shortly{ (try again in ~N min)}.` | if present | **yes** (same formula) |
|
||||
| *default (unknown/other)* | `🔴 {message \|\| error \|\| 'Billing request failed.'}` — still surfaces whatever the server said, never a blank toast. | if present | — |
|
||||
|
||||
## 3. Charge settlement outcomes (`pollCharge` / `renderChargeFailed`)
|
||||
|
||||
Source: `pollCharge` (`ui-tui/src/app/slash/commands/topup.ts:170-258`) and
|
||||
`renderChargeFailed` (`:260-290`). Poll cadence: 2s interval, 5-minute cap
|
||||
(`POLL_INTERVAL_MS=2000`, `POLL_CAP_MS=5*60*1000`), applied on **every**
|
||||
non-terminal path (pending *and* throttled), so a sustained 429/503 can't
|
||||
keep the poll alive forever.
|
||||
|
||||
| Outcome | Copy | Notes |
|
||||
|---|---|---|
|
||||
| `status: 'settled'` | `✅ ${amount_usd} added.` (or `✅ Credits added.` if no amount) | Terminal success. |
|
||||
| `status: 'failed'`, `reason: 'authentication_required'` | `🔴 Your bank requires verification (3DS). Complete it on the portal to finish this purchase.` | + `Portal:` line if `portalUrl`. |
|
||||
| `status: 'failed'`, `reason: 'payment_method_expired'` | `🔴 Your card has expired. Update it on the portal.` | + `Portal:` line. |
|
||||
| `status: 'failed'`, `reason: 'card_declined'` | `🔴 Your card was declined. Try another card on the portal.` | + `Portal:` line. |
|
||||
| `status: 'failed'`, `reason: 'processing_error'` | `🔴 The charge didn't go through (processing_error).` | + `Portal:` line. |
|
||||
| `status: 'failed'`, unrecognized/missing `reason` | `🔴 The charge didn't go through ({reason \|\| 'processing_error'}).` | Same portal funnel — parity with `cli.py`'s `_billing_portal_hint`. |
|
||||
| Poll timeout (still `pending` past the 5-min cap) | `🟡 Still processing after 5 minutes — this is a timeout, not a failure. Check /topup or the portal shortly.` | + `Portal:` line if `portalUrl`. Explicitly NOT called a failure. |
|
||||
| Revocation mid-poll (`remote_spending_revoked` / `session_revoked` while polling) | Renders the matching §2 copy, **then** appends: `🟡 Your last charge's outcome is unconfirmed — check your balance/history before retrying.` | CF-7 rule 4: a post-revoke 403 while polling is ambiguous (the charge may have already settled) — never call it "failed". |
|
||||
| 429/503 while polling (`rate_limited`/`temporarily_unavailable`/`stripe_unavailable`) | No error shown; backs off using `retry_after` (default 5s, capped at 30s) and keeps polling until the 5-min cap, then reads as timeout. | Not a payment failure. |
|
||||
| Other `!ok` status-check error | `🔴 Could not check the charge: {message \|\| error \|\| 'error'}` | |
|
||||
| Transport loss (poll RPC throws/rejects) | `🟡 Your last charge's outcome is unconfirmed — check your balance/history before retrying.` (`UNCONFIRMED_CHARGE_MESSAGE`) | Same "unconfirmed, check balance" framing as revocation mid-poll — a dropped connection can never be read as "failed". |
|
||||
|
||||
## 4. Subscription preview / pending-change / upgrade outcomes
|
||||
|
||||
Source: `previewAndRoute`, `applyPendingAndRoute`, `upgradeResult`,
|
||||
`stepUpDenialResult` in `ui-tui/src/components/subscriptionOverlay.tsx`.
|
||||
|
||||
**Preview `effect` values** (drive the Confirm screen):
|
||||
|
||||
| `effect` | Confirm screen copy | Primary action |
|
||||
|---|---|---|
|
||||
| `charge_now` | `Upgrade to {target}. You will be charged {amount} now (prorated).` (+ monthly-credits delta, + which card if resolver confidently knows) | `Pay {amount} & upgrade now` |
|
||||
| `scheduled` | `Change to {target} — takes effect {date}. No charge now; you keep your current plan until then.` | `Schedule change to {target}` |
|
||||
| `no_op` | `You are already on {target} — nothing to change.` | none (Back only) |
|
||||
| `blocked` | `{preview.reason}` or fallback `That change cannot be made here — manage it on the portal.` | `Manage on portal` |
|
||||
| Preview RPC returns `null`/transport failure | routes straight to Result: `Could not preview that change.` | — |
|
||||
| Preview `!ok`, `insufficient_scope` | routes to `stepup` screen (`{kind:'preview', tierId}`) | — |
|
||||
| Preview `!ok`, other error | routes to Result with `errorResult(p)` (`message \|\| error \|\| 'Something went wrong. Try again, or manage on the portal.'`) | — |
|
||||
|
||||
**Pending-change apply outcomes** (`applyPendingAndRoute`):
|
||||
|
||||
| `pending.kind` | Success copy |
|
||||
|---|---|
|
||||
| `cancellation` | `Scheduled — your plan stays active until the end of the billing period, then it cancels. Nothing changes today.` |
|
||||
| `tier_change` (downgrade/schedule) | `Scheduled — your plan doesn't change today. You keep your current plan until the end of the billing period, then it switches.` |
|
||||
| `upgrade` | routed through `upgradeResult` (below) |
|
||||
| any kind, mutation `insufficient_scope` | routes to stepup (`{kind:'apply'}`) |
|
||||
|
||||
**Upgrade `status` × `reason` matrix** (`upgradeResult`, checked in this
|
||||
order — `reason` is checked *before* `status`):
|
||||
|
||||
| Condition | Result |
|
||||
|---|---|
|
||||
| `r === null` (transport failure on the charging route) | `Couldn't confirm the upgrade — your card may or may not have been charged. Re-run /subscription to check your plan before trying again.` — ambiguous, never a blind retry. |
|
||||
| `reason: 'authentication_required'` **or** `reason: 'subscription_payment_intent_requires_action'` | `Please verify your card in the portal to finish this upgrade.` → `recovery_url`. **Both reasons map to the same SCA copy** — the client branches on `reason`, not `status`, specifically so an SCA case that pre-#711 NAS mislabels with `status: 'payment_failed'` (no distinguishing reason yet) still routes to the correct "verify your card" copy instead of reading as a hard decline. |
|
||||
| `reason: 'card_declined'` | `Your card was declined — try a different card on the portal.` → `recovery_url`. |
|
||||
| `ok && status: 'already_on_tier'` | `You are already on {target_tier_name}.` (success) |
|
||||
| `ok && status: 'upgraded'` | `Upgraded to {target_tier_name}. Your new monthly credits land in a moment.` — starts the eventual-consistency apply-poll (below). |
|
||||
| `status: 'requires_action'` (no distinguishing reason) | `This upgrade needs extra verification (3DS). Finish it on the portal.` → `recovery_url`. |
|
||||
| `status: 'payment_failed'` (no distinguishing reason) | `Your card was declined. Update your payment method on the portal and try again.` → `recovery_url`. |
|
||||
| anything else | `errorResult(r)`: `message \|\| error \|\| 'Something went wrong. Try again, or manage on the portal.'` |
|
||||
|
||||
**Eventual-consistency apply-poll** (`ResultScreen`, only after `status:
|
||||
'upgraded'`): polls `billing`/subscription state every 2s
|
||||
(`UPGRADE_CONFIRM_INTERVAL_MS`) up to 15 attempts
|
||||
(`UPGRADE_CONFIRM_ATTEMPTS`, i.e. ~30s) until `current.tier_id` flips to the
|
||||
target. While waiting the screen reads `Applying…`; if it never flips inside
|
||||
the budget it reads `Still applying` / `Your upgrade succeeded and is still
|
||||
applying — refresh in a moment.` — the upgrade is never re-reported as failed
|
||||
just because NAS hasn't caught up yet.
|
||||
|
||||
**Step-up denial copy** (`stepUpDenialResult`, subscription flow):
|
||||
|
||||
| `error` | Copy |
|
||||
|---|---|
|
||||
| `session_revoked` | `Your session expired — run /portal to log in again, then retry the change.` |
|
||||
| `remote_spending_revoked` | `{message}` or `Remote spending was stopped for this terminal — reconnect from the portal, then retry.` |
|
||||
| `rate_limited` | `Too many attempts — wait a moment, then try again.` |
|
||||
| other/unknown | `{message}` or `Remote Spending was not allowed — someone with billing permissions (owner, admin, or finance admin) must approve it. You can also make this change on the portal.` |
|
||||
|
||||
A **repeat** scope denial during a post-grant replay never re-enters the
|
||||
step-up screen (it's already mounted there — re-patching would freeze it);
|
||||
`allowStepUp=false` instead surfaces a terminal result: `Remote Spending still
|
||||
isn’t active for this terminal — the authorization didn’t take. Retry, or make
|
||||
this change on the portal.`
|
||||
|
||||
## Text-mode (CLI) parity
|
||||
|
||||
`cli.py`'s `_show_billing` / `_billing_overview` and `_show_subscription` /
|
||||
`_subscription_overview` render the same state shapes (balance title, two-bar
|
||||
dollar usage, auto-reload line, card line, monthly cap) and share the
|
||||
"fail-open on logged-out/portal-hiccup, never crash" discipline. The CLI's
|
||||
`/subscription` gives a paid admin/owner in an interactive context the **full
|
||||
in-terminal change flow** (tier picker → preview → confirm → apply, parity
|
||||
with the TUI overlay); members and non-interactive contexts fall back to
|
||||
`_billing_portal_hint`'s deep-link to `subscription_manage_url`. `/topup`'s
|
||||
interactive modal (prompt_toolkit) mirrors the TUI overlay the same way, and
|
||||
non-interactive contexts fall back to the same text + portal-link rendering,
|
||||
never prompting.
|
||||
|
||||
| CLI surface / state | Behavior (parity with TUI / desktop) |
|
||||
| --- | --- |
|
||||
| `/subscription` on **Free** + admin/owner + interactive | `_subscription_free_catalog` prints the plan catalog from the same `tiers[]` data the TUI uses — one row per enabled paid tier, cheapest first, `name · $/mo · $credits/mo` (monthly credits are DOLLARS → `$22 credits/mo`, never a bare number). A numbered pick opens the `/manage-subscription` deep-link with `plan=<tier_id>` appended so the portal preselects the chosen plan. Starting a new subscription needs a fresh card, so the only action is the portal hand-off (the terminal never charges here). |
|
||||
| Any CLI-built manage/subscribe URL | `subscription_manage_url(state, tier_id=…)` appends `plan=<tier_id>` (the stable `tiers[]` id, never a name/slug) **only when a tier was picked** (the Free catalog). The portal validates it server-side and ignores an unknown tier, so the CLI appends unconditionally on a pick, mirroring the TUI's `?plan=`. `org_id` is emitted first, `plan` second. |
|
||||
| **Downgrades** in the CLI | Stay **native / in-app** for normal changes (chargeless scheduling via `put_subscription_pending_change`). A blocked downgrade may still print the generic manage URL, but it never carries `plan=<tier_id>` — selected-tier deep-links are reserved for new subscriptions and upgrades. |
|
||||
| `/topup` overview action copy | Splits one-time top-up from automatic refill, the distinction stated up front in each first sentence: `Add funds now — a single charge, added to your balance today.` vs `Refill when low — charges $X automatically when your balance falls below $Y.` ("credits" stays out of the dollars-only `/topup` surface — "Add funds now" carries the one-time meaning without it). When auto-reload is off, the automatic line omits concrete amounts. |
|
||||
|
||||
## Forward compatibility
|
||||
|
||||
Any `error`/`status`/`reason` code not in the tables above lands on the
|
||||
`default` branch in `renderBillingError` (§2) or `errorResult`/`upgradeResult`'s
|
||||
fallthrough (§4): it still renders the server's own `message` (never blank,
|
||||
never a crash), just without bespoke copy or a typed recovery affordance.
|
||||
NAS W3 introduces card-health codes (`card_paused`, `card_expired`,
|
||||
`card_mismatch`) that are not yet typed here — until a client update adds
|
||||
explicit branches, they will arrive as unknown codes and degrade to this
|
||||
default path.
|
||||
@@ -0,0 +1,222 @@
|
||||
# Chronos managed-cron — agent ↔ NAS wire contract
|
||||
|
||||
**Status:** authoritative wire spec for the Chronos cron provider.
|
||||
**Audience:** the NAS-side implementer of the `agent-cron` endpoints
|
||||
(`nous-account-service`) and anyone debugging the managed-cron path.
|
||||
|
||||
Chronos lets a hosted Hermes gateway **scale to zero** while idle and still
|
||||
fire cron jobs. Instead of an in-process 60-second ticker, the agent asks NAS
|
||||
to arm exactly **one external one-shot per job at that job's real next-fire
|
||||
time**. NAS calls the agent back at fire time over an authenticated webhook;
|
||||
the agent runs the job and re-arms the next one-shot. Between fires the agent
|
||||
process can be fully stopped — it wakes only on a genuine fire.
|
||||
|
||||
The external scheduler NAS uses to implement the one-shots is an **internal NAS
|
||||
implementation detail**. The agent never talks to it, never holds its
|
||||
credentials, and never names it. The agent only knows the three NAS endpoints
|
||||
below.
|
||||
|
||||
```
|
||||
create/update/pause/resume/remove a cron job (agent side)
|
||||
│
|
||||
▼
|
||||
ChronosCronScheduler.reconcile() ── agent computes next_run_at
|
||||
│ POST {portal}/api/agent-cron/provision (auth: agent's Nous access token)
|
||||
▼
|
||||
NAS arms a one-shot for fire_at ── NAS owns the scheduler + its creds
|
||||
│
|
||||
⏰ at fire_at
|
||||
▼
|
||||
scheduler → POST {portal}/api/agent-cron/relay (auth: scheduler signature, NAS-verified)
|
||||
│
|
||||
▼
|
||||
NAS mints a short-lived agent-audience JWT (purpose=cron_fire)
|
||||
│ POST {agent_callback_url}/api/cron/fire (auth: that JWT)
|
||||
▼
|
||||
agent verifies the NAS JWT → store CAS claim → run_one_job → re-arm next one-shot
|
||||
```
|
||||
|
||||
## Trust model (read this first)
|
||||
|
||||
| Hop | Who calls whom | Auth mechanism | Verified by |
|
||||
|---|---|---|---|
|
||||
| 1 | agent → NAS (`provision`/`cancel`/`list`) | the agent's existing **Nous Portal access token** (Bearer) — for a hosted agent this is the **bootstrap-session token** NAS planted in `auth.json` (client `hermes-cli-vps`), NOT an `agent:*` client token | NAS (its normal agent-token path) |
|
||||
| 2 | scheduler → NAS (`relay`) | the scheduler's request **signature** | NAS (the signature path it already has) |
|
||||
| 3 | NAS → agent (`/api/cron/fire`) | a **short-lived NAS-minted JWT** (`aud=agent:{instance_id}`, `purpose=cron_fire`) | agent (PyJWT against NAS JWKS) |
|
||||
|
||||
> **Which token, exactly (hop 1).** A hosted agent never holds an `agent:{instance_id}`
|
||||
> OAuth client credential — that shape is minted only by the interactive dashboard
|
||||
> auth-code grant (a browser user). For all of its own outbound portal calls the
|
||||
> agent uses the **bootstrap-session access token** (`resolve_nous_access_token`),
|
||||
> minted under the bootstrap-only client `hermes-cli-vps` and seeded into the
|
||||
> container on first boot. NAS therefore must resolve the calling agent's instance
|
||||
> id from EITHER an `agent:{id}` client (self-hosted/dashboard callers) OR — for the
|
||||
> bootstrap token — from `AgentInstance.bootstrapSessionId` matching the token's
|
||||
> session id (`sid`), org-scoped. The fire JWT minted at hop 3 still carries
|
||||
> `aud=agent:{instance_id}` regardless. (Gating hop 1 on an `agent:*` client alone
|
||||
> 403s every real hosted-agent provision — see `src/server/agent-cron/instance-auth.ts`.)
|
||||
|
||||
Why NAS-mediated rather than scheduler→agent direct: the scheduler signs with
|
||||
**NAS's** keys, which the agent does not (and should not) hold. The agent can
|
||||
only verify a **NAS-minted** token — a trust path it already has. This keeps
|
||||
all scheduler credentials inside NAS. (Full rationale: the plan's DQ-4.)
|
||||
|
||||
No new secret is introduced on the agent: hop 1 reuses the token the agent
|
||||
already uses for the portal, and hop 3 reuses the NAS-JWT verification the agent
|
||||
already performs.
|
||||
|
||||
---
|
||||
|
||||
## Endpoint 1 — `POST /api/agent-cron/provision` (agent → NAS)
|
||||
|
||||
Arm (or re-arm, idempotently) exactly one one-shot for a job.
|
||||
|
||||
- **Auth:** `Authorization: Bearer <agent Nous access token>`. NAS validates via
|
||||
its normal agent-token path and scopes the row to the calling agent/org.
|
||||
- **Request body:**
|
||||
```json
|
||||
{
|
||||
"job_id": "ab12cd34",
|
||||
"fire_at": "2026-06-18T12:34:56+00:00",
|
||||
"agent_callback_url": "https://agent-xyz.fly.dev",
|
||||
"dedup_key": "ab12cd34:2026-06-18T12:34:56+00:00"
|
||||
}
|
||||
```
|
||||
- `fire_at` — ISO 8601, **agent-computed**. May be sub-minute in the future;
|
||||
NAS must honor second-granularity (the agent owns the time, so there is no
|
||||
1-minute scheduler floor).
|
||||
- `agent_callback_url` — the agent's own publicly-reachable base URL. NAS
|
||||
POSTs `{agent_callback_url}/api/cron/fire` at fire time.
|
||||
- `dedup_key` — `"{job_id}:{fire_at}"`. NAS **upserts by `(agent_id, job_id)`**
|
||||
so re-arming the same fire is idempotent (no duplicate one-shots). A new
|
||||
`fire_at` for the same `job_id` replaces the prior arm.
|
||||
- **Action:** arm one one-shot to fire at `fire_at`, destined for the NAS
|
||||
**relay** route (Endpoint 3) — NOT the agent directly, so NAS stays in the
|
||||
loop to mint the agent JWT. Persist `(agent_id, job_id, schedule_id,
|
||||
agent_callback_url)`.
|
||||
- **Response:** `200 {"schedule_id": "<opaque>"}`.
|
||||
|
||||
## Endpoint 2 — `POST /api/agent-cron/cancel` (agent → NAS)
|
||||
|
||||
- **Auth:** same as Endpoint 1.
|
||||
- **Body:** `{"job_id": "ab12cd34"}`.
|
||||
- **Action:** cancel the armed one-shot for `(agent_id, job_id)` and delete the
|
||||
row. Idempotent — cancelling an unknown job is a 200 no-op.
|
||||
- **Response:** `200 {"ok": true}`.
|
||||
|
||||
## Endpoint 3 — `POST /api/agent-cron/relay` (scheduler → NAS, the fire relay)
|
||||
|
||||
- **Auth:** the scheduler's request **signature**, verified by NAS with the
|
||||
signature path it already has. This is the trust boundary for the fire — a
|
||||
forged relay call must be rejected here.
|
||||
- **Action:**
|
||||
1. Look up `(agent_id, job_id) → agent_callback_url` from the persisted row.
|
||||
2. Mint a **short-lived** JWT: `aud = "agent:{instance_id}"`,
|
||||
`iss = {portal_url}`, `purpose = "cron_fire"`, small `exp` (≈60–120s),
|
||||
signed with NAS's normal asymmetric signing key (published via JWKS).
|
||||
3. `POST {agent_callback_url}/api/cron/fire` with
|
||||
`Authorization: Bearer <that JWT>` and body `{"job_id": "...", "fire_at": "..."}`.
|
||||
4. Treat a non-2xx agent response as a **retryable** failure (let the
|
||||
scheduler retry the relay). The agent's store CAS de-dupes a double fire,
|
||||
so retries are safe.
|
||||
- **Response to the scheduler:** 2xx once the agent POST is accepted (202), so
|
||||
the scheduler does not retry a delivered fire.
|
||||
|
||||
---
|
||||
|
||||
## Inbound `POST /api/cron/fire` (NAS → agent) — agent side, already implemented
|
||||
|
||||
This is the agent endpoint NAS calls in Endpoint 3 step 3. Two hops on hosted
|
||||
deployments:
|
||||
|
||||
1. **Dashboard app** (`hermes_cli/web_server.py`) — the agent's only public
|
||||
HTTP surface (the Fly proxy exposes exactly one port, the dashboard's). It
|
||||
is in `PUBLIC_API_PATHS` so the dashboard cookie gate lets the bearer-JWT
|
||||
callback through to the verifier. The dashboard verifies the JWT, resolves
|
||||
the job's profile, then **forwards** the fire to hop 2 on loopback with the
|
||||
NAS bearer preserved — it does NOT execute the job itself.
|
||||
2. **Gateway `APIServerAdapter`** (`gateway/platforms/api_server.py`, loopback
|
||||
bind, default port 8642) — re-verifies the JWT (defense in depth) and runs
|
||||
the job with the gateway's **live platform adapters**, which is what makes
|
||||
delivery work for relay-fronted logical platforms and E2EE rooms (the
|
||||
standalone send path can serve neither). Self-host API-server deployments
|
||||
that expose the api_server directly hit hop 2 without hop 1.
|
||||
|
||||
Gateway unreachable from hop 1 (scale-to-zero wake still booting, restart
|
||||
window, api_server disabled) → the dashboard returns **503** and NAS retries
|
||||
(non-2xx = retryable, below); the store CAS de-dupes the eventual double fire.
|
||||
There is deliberately no in-dashboard execution fallback. The verifier is
|
||||
`plugins/cron/chronos/verify.py`.
|
||||
|
||||
- **Auth:** `Authorization: Bearer <NAS-minted JWT>`. The agent verifies:
|
||||
- signature against the NAS JWKS (`cron.chronos.nas_jwks_url`),
|
||||
- `aud` == `cron.chronos.expected_audience` (this agent's
|
||||
`agent:{instance_id}`),
|
||||
- `iss` == `cron.chronos.portal_url`,
|
||||
- `exp` / `nbf` (30s leeway),
|
||||
- `purpose == "cron_fire"` — a general agent JWT (no/other purpose) is
|
||||
rejected so it can't be replayed against this endpoint.
|
||||
- **Body:** `{"job_id": "ab12cd34", "fire_at": "..."}` (only `job_id` is used).
|
||||
- **Behavior:**
|
||||
- invalid/missing/forged/expired/wrong-aud/wrong-purpose token → **401**, no
|
||||
execution.
|
||||
- missing `job_id` → **400**.
|
||||
- valid → **202 `{"status": "accepted", "job_id": "..."}`** immediately, and
|
||||
the job runs in the background. 202-before-run means a long agent turn never
|
||||
trips the relay's HTTP timeout.
|
||||
- **At-most-once:** the agent claims the job with a store-level compare-and-set
|
||||
(`claim_job_for_fire`) before running. A relay/scheduler retry that arrives
|
||||
while the first fire is in flight (or after it completed) loses the claim and
|
||||
does not double-run.
|
||||
|
||||
---
|
||||
|
||||
## At-most-once & re-arm semantics
|
||||
|
||||
- **Recurring (cron/interval):** on fire, the agent advances `next_run_at`
|
||||
(under its store lock) as part of the claim, runs the job, then re-provisions
|
||||
a one-shot for the new `next_run_at`. A duplicate relay for the old `fire_at`
|
||||
finds the claim taken / time advanced and is dropped.
|
||||
- **One-shot (`30m`, `+90s`, etc.):** fires once; `mark_job_run` marks it
|
||||
completed. No re-arm.
|
||||
- **`repeat.times = N`:** `mark_job_run` deletes the job at the limit, so
|
||||
`get_job` returns `None` after the final fire → the agent does **not** re-arm
|
||||
→ the schedule stops cleanly with no orphaned one-shot.
|
||||
- **Multi-replica agents:** the store CAS makes the fire at-most-once across N
|
||||
gateway replicas sharing one `HERMES_HOME` — exactly one replica runs each
|
||||
fire.
|
||||
|
||||
## Reconcile (self-healing)
|
||||
|
||||
The agent reconciles desired (`jobs.json`) vs armed on:
|
||||
- `start()` (gateway boot / wake),
|
||||
- every successful job mutation (`on_jobs_changed`),
|
||||
- piggybacked after each fire (re-arm).
|
||||
|
||||
Reconcile arms missing/changed-time jobs and cancels orphans. A missed
|
||||
provision (transient NAS error) self-heals on the next reconcile. There is **no
|
||||
periodic wake** of a sleeping agent — that would negate scale-to-zero.
|
||||
|
||||
## Config (agent side)
|
||||
|
||||
All non-secret (`cron.chronos.*` in `config.yaml`); the agent holds no scheduler
|
||||
credentials. For hosted agents NAS sets these at provision time:
|
||||
|
||||
| key | meaning |
|
||||
|---|---|
|
||||
| `cron.provider` | `"chronos"` to activate (empty = built-in ticker) |
|
||||
| `cron.chronos.portal_url` | NAS base URL (also the expected JWT `iss`) |
|
||||
| `cron.chronos.callback_url` | the agent's own public base URL for NAS→agent fires |
|
||||
| `cron.chronos.expected_audience` | this agent's JWT `aud` (`agent:{instance_id}`) |
|
||||
| `cron.chronos.nas_jwks_url` | NAS JWKS for verifying the fire JWT |
|
||||
|
||||
If `callback_url` / `portal_url` is blank or the agent has no Nous login,
|
||||
`is_available()` returns False and the resolver falls back to the built-in
|
||||
in-process ticker — cron never loses its trigger.
|
||||
|
||||
## Escape hatch (not default)
|
||||
|
||||
The inbound `/api/cron/fire` verifier is pluggable (`get_fire_verifier()`). If
|
||||
relay volume through NAS ever saturates, a direct scheduler→agent mode with a
|
||||
per-job NAS-minted cron-key can replace the NAS-JWT verifier with **no change to
|
||||
the webhook handler**. NAS-mediated (this contract) is the default.
|
||||
@@ -0,0 +1,32 @@
|
||||
# Cron Doctor Spec
|
||||
|
||||
## Problem
|
||||
|
||||
Scheduled jobs can silently degrade when a script is moved, a workdir disappears,
|
||||
a provider run fails, or delivery starts failing. `hermes cron list` shows some of
|
||||
this inline, but there is no compact read-only health check that can be run from a
|
||||
terminal, cron job, or CI-style smoke check.
|
||||
|
||||
## Goal
|
||||
|
||||
Add `hermes cron doctor` as a read-only diagnostic command that summarizes cron
|
||||
job health and exits non-zero when actionable issues are found.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Do not mutate jobs or auto-repair state.
|
||||
- Do not start/stop the gateway.
|
||||
- Do not inspect secrets or print credentials.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `hermes cron doctor` returns `0` and prints a healthy message when active jobs
|
||||
have no detected issues.
|
||||
- It returns `1` and prints grouped job-level issues when any active job has:
|
||||
- last run failure (`last_status` not `ok`),
|
||||
- last delivery failure,
|
||||
- no `next_run_at` while still active,
|
||||
- `no_agent` enabled without a script,
|
||||
- script path missing/outside `HERMES_HOME/scripts`, or
|
||||
- configured workdir path missing.
|
||||
- Parser, command dispatch, and focused tests cover the new subcommand.
|
||||
@@ -0,0 +1,903 @@
|
||||
<!doctype html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8" />
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1" />
|
||||
<title>Hermes Kanban — Native Dialog Prototypes</title>
|
||||
<style>
|
||||
/* =========================================================== */
|
||||
/* Design tokens — host @nous-research/ui, oklch→sRGB approx. */
|
||||
/* =========================================================== */
|
||||
:root {
|
||||
--bg: #ffffff;
|
||||
--bg-subtle: #f8f9fa;
|
||||
--bg-muted: #f1f3f5;
|
||||
--border: #e5e7eb;
|
||||
--border-strong: #d1d5db;
|
||||
--fg: #0f172a;
|
||||
--fg-muted: #64748b;
|
||||
--fg-subtle: #94a3b8;
|
||||
--ring: #94a3b8;
|
||||
--accent: #2563eb;
|
||||
--accent-fg: #ffffff;
|
||||
--destructive: #dc2626;
|
||||
--destructive-fg: #ffffff;
|
||||
--success: #16a34a;
|
||||
--shadow-md: 0 4px 12px -2px rgba(15, 23, 42, 0.10),
|
||||
0 2px 4px -2px rgba(15, 23, 42, 0.06);
|
||||
--shadow-lg: 0 20px 25px -5px rgba(15, 23, 42, 0.10),
|
||||
0 8px 10px -6px rgba(15, 23, 42, 0.04);
|
||||
--radius: 10px;
|
||||
--radius-sm: 6px;
|
||||
--font: ui-sans-serif, system-ui, -apple-system, "Segoe UI",
|
||||
Roboto, "Helvetica Neue", Arial, sans-serif;
|
||||
--mono: ui-monospace, SFMono-Regular, "SF Mono", Menlo, monospace;
|
||||
}
|
||||
@media (prefers-color-scheme: dark) {
|
||||
:root {
|
||||
--bg: #0b0d10;
|
||||
--bg-subtle: #11141a;
|
||||
--bg-muted: #161a22;
|
||||
--border: #232833;
|
||||
--border-strong: #2c3140;
|
||||
--fg: #e6e8ec;
|
||||
--fg-muted: #9aa3b2;
|
||||
--fg-subtle: #6b7280;
|
||||
--ring: #4b5563;
|
||||
--accent: #3b82f6;
|
||||
--accent-fg: #ffffff;
|
||||
--destructive: #ef4444;
|
||||
--destructive-fg: #ffffff;
|
||||
--success: #22c55e;
|
||||
}
|
||||
}
|
||||
* { box-sizing: border-box; }
|
||||
html, body { margin: 0; padding: 0; }
|
||||
body {
|
||||
font-family: var(--font);
|
||||
color: var(--fg);
|
||||
background: var(--bg-subtle);
|
||||
line-height: 1.5;
|
||||
-webkit-font-smoothing: antialiased;
|
||||
}
|
||||
|
||||
/* =========================================================== */
|
||||
/* Page chrome */
|
||||
/* =========================================================== */
|
||||
header.page {
|
||||
padding: 24px 32px 16px;
|
||||
border-bottom: 1px solid var(--border);
|
||||
background: var(--bg);
|
||||
}
|
||||
header.page h1 {
|
||||
font-size: 18px; font-weight: 600; margin: 0 0 4px;
|
||||
letter-spacing: -0.01em;
|
||||
}
|
||||
header.page p {
|
||||
font-size: 13px; color: var(--fg-muted); margin: 0;
|
||||
max-width: 920px;
|
||||
}
|
||||
header.page .controls {
|
||||
margin-top: 14px; display: flex; gap: 8px; flex-wrap: wrap;
|
||||
}
|
||||
header.page .controls button {
|
||||
font: inherit; font-size: 12px; font-weight: 500;
|
||||
padding: 6px 12px; border-radius: var(--radius-sm);
|
||||
border: 1px solid var(--border); background: var(--bg);
|
||||
color: var(--fg); cursor: pointer;
|
||||
transition: background 120ms ease, border-color 120ms ease;
|
||||
}
|
||||
header.page .controls button:hover { background: var(--bg-muted); }
|
||||
header.page .controls button.active {
|
||||
background: var(--accent); border-color: var(--accent); color: var(--accent-fg);
|
||||
}
|
||||
|
||||
.board {
|
||||
padding: 24px 32px 80px;
|
||||
display: grid;
|
||||
grid-template-columns: repeat(4, minmax(0, 1fr));
|
||||
gap: 20px;
|
||||
max-width: 1700px;
|
||||
}
|
||||
@media (max-width: 1400px) { .board { grid-template-columns: repeat(2, minmax(0, 1fr)); } }
|
||||
@media (max-width: 720px) { .board { grid-template-columns: 1fr; } }
|
||||
|
||||
section.variant {
|
||||
background: var(--bg);
|
||||
border: 1px solid var(--border);
|
||||
border-radius: var(--radius);
|
||||
overflow: hidden;
|
||||
display: flex; flex-direction: column;
|
||||
min-width: 0;
|
||||
}
|
||||
section.variant header {
|
||||
padding: 14px 16px 12px;
|
||||
border-bottom: 1px solid var(--border);
|
||||
background: var(--bg-subtle);
|
||||
}
|
||||
section.variant header h2 {
|
||||
font-size: 11px; font-weight: 600; margin: 0 0 2px;
|
||||
letter-spacing: 0.08em; text-transform: uppercase;
|
||||
color: var(--fg-muted);
|
||||
}
|
||||
section.variant header .label {
|
||||
font-size: 15px; font-weight: 600; color: var(--fg);
|
||||
letter-spacing: -0.01em;
|
||||
}
|
||||
section.variant header p {
|
||||
font-size: 12px; color: var(--fg-muted); margin: 6px 0 0;
|
||||
line-height: 1.4;
|
||||
}
|
||||
section.variant .stage {
|
||||
flex: 1;
|
||||
padding: 24px 16px;
|
||||
display: flex; flex-direction: column; align-items: center;
|
||||
justify-content: flex-start;
|
||||
background: var(--bg);
|
||||
min-height: 540px;
|
||||
position: relative;
|
||||
}
|
||||
section.variant .actions {
|
||||
padding: 10px 16px;
|
||||
border-top: 1px solid var(--border);
|
||||
display: flex; flex-wrap: wrap; gap: 6px;
|
||||
background: var(--bg-subtle);
|
||||
}
|
||||
section.variant .actions button {
|
||||
font: inherit; font-size: 11px;
|
||||
padding: 5px 10px; border-radius: var(--radius-sm);
|
||||
border: 1px solid var(--border); background: var(--bg);
|
||||
color: var(--fg-muted); cursor: pointer;
|
||||
}
|
||||
section.variant .actions button:hover {
|
||||
border-color: var(--border-strong); color: var(--fg);
|
||||
}
|
||||
section.variant .actions .trash {
|
||||
color: var(--destructive); border-color: var(--destructive);
|
||||
background: color-mix(in srgb, var(--destructive) 6%, var(--bg));
|
||||
}
|
||||
section.variant .note {
|
||||
font-size: 11px; color: var(--fg-muted);
|
||||
padding: 8px 12px; background: var(--bg-muted);
|
||||
border-top: 1px solid var(--border);
|
||||
line-height: 1.45;
|
||||
}
|
||||
section.variant .note strong { color: var(--fg); }
|
||||
section.variant .note code {
|
||||
font-family: var(--mono); font-size: 10px;
|
||||
background: var(--bg); padding: 1px 4px; border-radius: 3px;
|
||||
border: 1px solid var(--border);
|
||||
}
|
||||
|
||||
/* =========================================================== */
|
||||
/* Backdrop + dialog (shared) */
|
||||
/* =========================================================== */
|
||||
.backdrop {
|
||||
position: absolute; inset: 0;
|
||||
background: rgba(15, 23, 42, 0.45);
|
||||
backdrop-filter: blur(2px);
|
||||
display: flex; align-items: center; justify-content: center;
|
||||
z-index: 10;
|
||||
animation: fade-in 160ms ease both;
|
||||
padding: 16px;
|
||||
}
|
||||
@keyframes fade-in { from { opacity: 0; } to { opacity: 1; } }
|
||||
@keyframes dialog-in {
|
||||
from { opacity: 0; transform: translateY(4px) scale(0.98); }
|
||||
to { opacity: 1; transform: none; }
|
||||
}
|
||||
.dialog {
|
||||
background: var(--bg);
|
||||
border: 1px solid var(--border);
|
||||
border-radius: var(--radius);
|
||||
box-shadow: var(--shadow-lg);
|
||||
width: 100%; max-width: 440px;
|
||||
padding: 20px;
|
||||
animation: dialog-in 180ms cubic-bezier(.2,.8,.2,1) both;
|
||||
font-size: 14px;
|
||||
}
|
||||
.dialog h3 {
|
||||
margin: 0; font-size: 16px; font-weight: 600;
|
||||
letter-spacing: -0.01em; display: flex; align-items: center; gap: 10px;
|
||||
color: var(--fg);
|
||||
}
|
||||
.dialog p.desc {
|
||||
margin: 8px 0 0; color: var(--fg-muted); font-size: 13px;
|
||||
line-height: 1.45;
|
||||
}
|
||||
.dialog .body { margin-top: 14px; }
|
||||
.dialog .actions {
|
||||
margin-top: 20px;
|
||||
display: flex; justify-content: flex-end; gap: 8px;
|
||||
}
|
||||
.dialog button {
|
||||
font: inherit; font-size: 13px; font-weight: 500;
|
||||
padding: 8px 14px; border-radius: var(--radius-sm);
|
||||
border: 1px solid var(--border); background: var(--bg);
|
||||
color: var(--fg); cursor: pointer;
|
||||
transition: background 120ms ease, border-color 120ms ease;
|
||||
}
|
||||
.dialog button:hover { background: var(--bg-muted); }
|
||||
.dialog button:focus-visible {
|
||||
outline: 2px solid var(--ring); outline-offset: 2px;
|
||||
}
|
||||
.dialog button.primary {
|
||||
background: var(--accent); border-color: var(--accent); color: var(--accent-fg);
|
||||
}
|
||||
.dialog button.primary:hover {
|
||||
background: color-mix(in srgb, var(--accent) 88%, black);
|
||||
}
|
||||
.dialog button.destructive {
|
||||
background: var(--destructive); border-color: var(--destructive); color: var(--destructive-fg);
|
||||
}
|
||||
.dialog button.destructive:hover {
|
||||
background: color-mix(in srgb, var(--destructive) 88%, black);
|
||||
}
|
||||
.dialog button:disabled { opacity: 0.5; cursor: not-allowed; }
|
||||
.dialog .spinner {
|
||||
display: inline-block;
|
||||
width: 12px; height: 12px;
|
||||
border: 2px solid currentColor;
|
||||
border-right-color: transparent;
|
||||
border-radius: 50%;
|
||||
animation: spin 0.7s linear infinite;
|
||||
vertical-align: -2px;
|
||||
margin-right: 6px;
|
||||
}
|
||||
@keyframes spin { to { transform: rotate(360deg); } }
|
||||
|
||||
.dialog textarea, .dialog input[type="text"] {
|
||||
font: inherit; font-size: 13px;
|
||||
width: 100%; padding: 8px 10px;
|
||||
border: 1px solid var(--border-strong); border-radius: var(--radius-sm);
|
||||
background: var(--bg); color: var(--fg);
|
||||
resize: vertical;
|
||||
}
|
||||
.dialog textarea:focus, .dialog input[type="text"]:focus {
|
||||
outline: 2px solid var(--ring); outline-offset: -1px; border-color: var(--ring);
|
||||
}
|
||||
.dialog textarea.invalid, .dialog input.invalid {
|
||||
border-color: var(--destructive);
|
||||
}
|
||||
.dialog .helper {
|
||||
font-size: 11px; color: var(--fg-muted);
|
||||
margin-top: 4px; min-height: 16px;
|
||||
}
|
||||
.dialog .helper.error { color: var(--destructive); }
|
||||
|
||||
/* =========================================================== */
|
||||
/* Icons (inline SVG, Lucide-style) */
|
||||
/* =========================================================== */
|
||||
.ico { width: 18px; height: 18px; stroke: currentColor; fill: none;
|
||||
stroke-width: 2; stroke-linecap: round; stroke-linejoin: round;
|
||||
flex-shrink: 0; }
|
||||
.ico.muted { color: var(--fg-muted); }
|
||||
.ico.destructive { color: var(--destructive); }
|
||||
.ico.success { color: var(--success); }
|
||||
|
||||
/* =========================================================== */
|
||||
/* Toast container (Variant C) */
|
||||
/* =========================================================== */
|
||||
.toast-container {
|
||||
position: absolute; right: 16px; bottom: 16px;
|
||||
display: flex; flex-direction: column-reverse; gap: 8px;
|
||||
z-index: 11; width: 320px;
|
||||
pointer-events: none;
|
||||
}
|
||||
.toast {
|
||||
background: var(--fg);
|
||||
color: var(--bg);
|
||||
border-radius: var(--radius-sm);
|
||||
padding: 10px 12px;
|
||||
box-shadow: var(--shadow-lg);
|
||||
display: flex; flex-direction: column; gap: 6px;
|
||||
pointer-events: auto;
|
||||
animation: toast-in 220ms cubic-bezier(.2,.8,.2,1) both;
|
||||
font-size: 13px;
|
||||
}
|
||||
.toast.error { background: var(--destructive); color: var(--destructive-fg); }
|
||||
@keyframes toast-in {
|
||||
from { opacity: 0; transform: translateY(8px); }
|
||||
to { opacity: 1; transform: none; }
|
||||
}
|
||||
.toast .title {
|
||||
display: flex; align-items: center; gap: 8px;
|
||||
font-weight: 600;
|
||||
}
|
||||
.toast .desc { font-size: 12px; opacity: 0.85; }
|
||||
.toast .row {
|
||||
display: flex; align-items: center; justify-content: space-between; gap: 8px;
|
||||
}
|
||||
.toast .undo {
|
||||
font: inherit; font-size: 11px; font-weight: 600;
|
||||
padding: 4px 10px; border-radius: var(--radius-sm);
|
||||
background: transparent; color: inherit;
|
||||
border: 1px solid currentColor;
|
||||
cursor: pointer;
|
||||
text-transform: uppercase; letter-spacing: 0.06em;
|
||||
}
|
||||
.toast .undo:hover { background: rgba(255,255,255,0.12); }
|
||||
|
||||
/* =========================================================== */
|
||||
/* Bulk-many list (Variant B-refined) */
|
||||
/* =========================================================== */
|
||||
.bulk-list {
|
||||
margin-top: 14px;
|
||||
border: 1px solid var(--border); border-radius: var(--radius-sm);
|
||||
background: var(--bg-subtle);
|
||||
max-height: 200px; overflow-y: auto;
|
||||
}
|
||||
.bulk-item {
|
||||
padding: 8px 10px;
|
||||
border-bottom: 1px solid var(--border);
|
||||
display: flex; align-items: flex-start; gap: 8px;
|
||||
font-size: 12px;
|
||||
}
|
||||
.bulk-item:last-child { border-bottom: none; }
|
||||
.bulk-item .meta { flex: 1; min-width: 0; }
|
||||
.bulk-item .meta .title {
|
||||
font-weight: 500; color: var(--fg);
|
||||
overflow: hidden; text-overflow: ellipsis; white-space: nowrap;
|
||||
}
|
||||
.bulk-item .meta .id {
|
||||
font-family: var(--mono); font-size: 10px; color: var(--fg-subtle);
|
||||
margin-top: 1px;
|
||||
}
|
||||
.bulk-item textarea {
|
||||
font: inherit; font-size: 11px;
|
||||
width: 100%; padding: 4px 6px;
|
||||
border: 1px solid var(--border); border-radius: 4px;
|
||||
background: var(--bg); color: var(--fg);
|
||||
margin-top: 4px;
|
||||
min-height: 22px; resize: vertical;
|
||||
}
|
||||
|
||||
/* =========================================================== */
|
||||
/* Decision legend */
|
||||
/* =========================================================== */
|
||||
.legend {
|
||||
padding: 16px 32px 24px;
|
||||
background: var(--bg);
|
||||
border-top: 1px solid var(--border);
|
||||
font-size: 12px; color: var(--fg-muted);
|
||||
max-width: 1700px;
|
||||
}
|
||||
.legend h3 {
|
||||
font-size: 11px; font-weight: 600; color: var(--fg);
|
||||
letter-spacing: 0.08em; text-transform: uppercase;
|
||||
margin: 0 0 8px;
|
||||
}
|
||||
.legend table { border-collapse: collapse; width: 100%; max-width: 1100px; }
|
||||
.legend th, .legend td {
|
||||
text-align: left; padding: 6px 10px;
|
||||
border-bottom: 1px solid var(--border);
|
||||
font-size: 12px;
|
||||
}
|
||||
.legend th { color: var(--fg); font-weight: 600; }
|
||||
.legend td.yes { color: var(--success); }
|
||||
.legend td.no { color: var(--fg-subtle); }
|
||||
.legend td.partial { color: var(--accent); }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
|
||||
<header class="page">
|
||||
<h1>Hermes Kanban — Native Dialog Prototypes</h1>
|
||||
<p>
|
||||
Four approaches to replacing <code>window.confirm()</code>,
|
||||
<code>window.prompt()</code>, and <code>window.alert()</code> in
|
||||
<code>plugins/kanban/dashboard/dist/index.js</code>. Click a page-level
|
||||
trigger to fire the same flow in every variant simultaneously, or use the
|
||||
per-variant buttons to fire flows unique to that variant.
|
||||
</p>
|
||||
<div class="controls">
|
||||
<button data-trigger="move-done">Trigger: Mark Done (with summary)</button>
|
||||
<button data-trigger="move-blocked">Trigger: Mark Blocked</button>
|
||||
<button data-trigger="bulk-delete">Trigger: Bulk delete (3 tasks)</button>
|
||||
<button data-trigger="archive-board">Trigger: Archive board</button>
|
||||
<button data-trigger="remove-attachment">Trigger: Remove attachment</button>
|
||||
<button data-trigger="error">Trigger: Error toast</button>
|
||||
<button data-trigger="clear">Clear all stages</button>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
<div class="board">
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- VARIANT A — Conservative -->
|
||||
<!-- ============================================================ -->
|
||||
<section class="variant" data-variant="A">
|
||||
<header>
|
||||
<h2>Variant A</h2>
|
||||
<div class="label">Conservative</div>
|
||||
<p>Direct 1:1 mapping to the host's <code>ConfirmDialog</code>. Single-line input. Minimal chrome.</p>
|
||||
</header>
|
||||
<div class="stage" data-stage="A">
|
||||
<div class="placeholder">Click a trigger to preview.</div>
|
||||
</div>
|
||||
<div class="note">
|
||||
<strong>Trade-off:</strong> Safest to ship — zero new components, zero new
|
||||
patterns. Fails the GPT-OSS review on three points: no multi-line
|
||||
summary, validation triggers a SECOND dialog instead of inline, no bulk
|
||||
affordance.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- VARIANT B — Strong-fit (Pro's pick) -->
|
||||
<!-- ============================================================ -->
|
||||
<section class="variant" data-variant="B">
|
||||
<header>
|
||||
<h2>Variant B</h2>
|
||||
<div class="label">Strong-fit (Pro's pick)</div>
|
||||
<p>Textarea + contextual SVG icon + inline validation + pluralized copy. What the design brief recommends.</p>
|
||||
</header>
|
||||
<div class="stage" data-stage="B">
|
||||
<div class="placeholder">Click a trigger to preview.</div>
|
||||
</div>
|
||||
<div class="note">
|
||||
<strong>Trade-off:</strong> Best baseline, but Pro suggested <em>either</em>
|
||||
a disabled-button <em>or</em> an inline error — GPT-OSS caught that both
|
||||
are needed (button disabled AND error visible) for screen-reader users.
|
||||
Bottom-sheet on mobile is the right move but still has keyboard edge cases.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- VARIANT B-refined (Pro + GPT-OSS synthesis) -->
|
||||
<!-- ============================================================ -->
|
||||
<section class="variant" data-variant="B2">
|
||||
<header>
|
||||
<h2>Variant B-refined</h2>
|
||||
<div class="label">Synthesis (recommended)</div>
|
||||
<p>All of B's improvements + GPT-OSS fixes: auto-focus, dual-validation, cancellable-spinner, per-task summaries in bulk.</p>
|
||||
</header>
|
||||
<div class="stage" data-stage="B2">
|
||||
<div class="placeholder">Click a trigger to preview.</div>
|
||||
</div>
|
||||
<div class="note">
|
||||
<strong>Why this is the recommendation:</strong> Single contextual icon
|
||||
(not four) keeps the title readable. Confirm button stays enabled until
|
||||
textarea has content; the inline error appears on submit-attempted-empty
|
||||
AND on blur if still empty. Cancel button stays clickable during the
|
||||
PATCH (only the confirm shows the spinner) so users can abort slow
|
||||
networks. Bulk-many shows an expandable list with per-task summary
|
||||
fields.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- VARIANT C — Divergent (undo toast, non-destructive only) -->
|
||||
<!-- ============================================================ -->
|
||||
<section class="variant" data-variant="C">
|
||||
<header>
|
||||
<h2>Variant C</h2>
|
||||
<div class="label">Divergent: undo toast</div>
|
||||
<p>Skip the modal entirely for non-destructive moves. Optimistic UI + 5s undo in a bottom-right toast.</p>
|
||||
</header>
|
||||
<div class="stage" data-stage="C">
|
||||
<div class="placeholder">Click a trigger to preview.</div>
|
||||
<div class="toast-container" data-toasts></div>
|
||||
</div>
|
||||
<div class="note">
|
||||
<strong>Trade-off:</strong> Radical speedup for routine moves, but breaks
|
||||
the required-summary flow (you can't optimistically "complete" a task
|
||||
that's missing required schema data). Best used as a
|
||||
<em>complement</em> to B-refined — undo toast for the safe moves,
|
||||
modal for <code>done</code>/<code>blocked</code>/<code>archive</code>.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
</div>
|
||||
|
||||
<div class="legend">
|
||||
<h3>Decision matrix — recommended pick: B-refined</h3>
|
||||
<table>
|
||||
<thead>
|
||||
<tr><th>Capability</th><th>A</th><th>B</th><th>B-refined</th><th>C</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr><td>Centered modal (Radix)</td> <td class="yes">✓</td> <td class="yes">✓</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Multi-line summary</td> <td class="no">✗</td> <td class="yes">✓</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Single contextual icon</td> <td class="no">✗</td> <td class="partial">4</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Inline validation (no second dialog)</td><td class="no">✗</td> <td class="partial">~</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Disabled button + persistent error</td> <td class="no">✗</td> <td class="partial">~</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Cancellable spinner during PATCH</td> <td class="no">✗</td> <td class="no">✗</td> <td class="yes">✓</td> <td class="yes">✓</td></tr>
|
||||
<tr><td>Bulk-many expandable list</td> <td class="no">✗</td> <td class="no">✗</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Auto-focus textarea + mobile scroll</td> <td class="no">✗</td> <td class="partial">~</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
<tr><td>Toast on success</td> <td class="no">✗</td> <td class="no">✗</td> <td class="no">✗</td> <td class="yes">✓</td></tr>
|
||||
<tr><td>Toast on error</td> <td class="no">✗</td> <td class="no">✗</td> <td class="partial">opt.</td> <td class="yes">✓</td></tr>
|
||||
<tr><td>Survives the required-summary flow</td> <td class="no">✗</td> <td class="yes">✓</td> <td class="yes">✓</td> <td class="no">✗</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
|
||||
<script>
|
||||
/* =========================================================== */
|
||||
/* Shared icon library (inline SVG) */
|
||||
/* =========================================================== */
|
||||
const ICONS = {
|
||||
check: '<svg class="ico success" viewBox="0 0 24 24"><path d="M20 6 9 17l-5-5"/></svg>',
|
||||
archive: '<svg class="ico muted" viewBox="0 0 24 24"><rect x="2" y="4" width="20" height="5" rx="1"/><path d="M4 9v9a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V9"/><path d="M10 13h4"/></svg>',
|
||||
pause: '<svg class="ico muted" viewBox="0 0 24 24"><circle cx="12" cy="12" r="10"/><path d="M10 9v6M14 9v6"/></svg>',
|
||||
trash: '<svg class="ico destructive" viewBox="0 0 24 24"><path d="M3 6h18"/><path d="M8 6V4a2 2 0 0 1 2-2h4a2 2 0 0 1 2 2v2"/><path d="M19 6v14a2 2 0 0 1-2 2H7a2 2 0 0 1-2-2V6"/></svg>',
|
||||
paperclip:'<svg class="ico muted" viewBox="0 0 24 24"><path d="m21.44 11.05-9.19 9.19a6 6 0 0 1-8.49-8.49l8.57-8.57A4 4 0 1 1 17.93 8.8l-8.59 8.57a2 2 0 0 1-2.83-2.83l8.49-8.48"/></svg>',
|
||||
info: '<svg class="ico muted" viewBox="0 0 24 24"><circle cx="12" cy="12" r="10"/><path d="M12 16v-4M12 8h.01"/></svg>',
|
||||
x: '<svg class="ico" viewBox="0 0 24 24"><path d="M18 6 6 18M6 6l12 12"/></svg>',
|
||||
undo: '<svg class="ico" viewBox="0 0 24 24"><path d="M3 7v6h6"/><path d="M21 17a9 9 0 0 0-15-6.7L3 13"/></svg>',
|
||||
};
|
||||
|
||||
/* =========================================================== */
|
||||
/* Dialog factories */
|
||||
/* Each factory: (stage, opts) => void; clears stage, mounts */
|
||||
/* backdrop + dialog. */
|
||||
/* =========================================================== */
|
||||
|
||||
function clearStage(stage) {
|
||||
// Remove everything except the toast container (for variant C)
|
||||
const toastContainer = stage.querySelector('.toast-container');
|
||||
stage.innerHTML = '';
|
||||
if (toastContainer) stage.appendChild(toastContainer);
|
||||
}
|
||||
|
||||
function mountBackdrop(stage, dialog) {
|
||||
const backdrop = document.createElement('div');
|
||||
backdrop.className = 'backdrop';
|
||||
backdrop.appendChild(dialog);
|
||||
stage.appendChild(backdrop);
|
||||
return backdrop;
|
||||
}
|
||||
|
||||
/* ----- Variant A: Conservative -------------------------------- */
|
||||
function A_moveDone(stage) {
|
||||
clearStage(stage);
|
||||
const dlg = document.createElement('div');
|
||||
dlg.className = 'dialog';
|
||||
dlg.innerHTML = `
|
||||
<h3>Mark this task as done?</h3>
|
||||
<p class="desc">The worker's claim is released and dependent children become ready.</p>
|
||||
<div class="body">
|
||||
<input type="text" placeholder="Completion summary" />
|
||||
</div>
|
||||
<div class="actions">
|
||||
<button data-act="cancel">Cancel</button>
|
||||
<button data-act="confirm" class="primary">Mark Done</button>
|
||||
</div>
|
||||
`;
|
||||
const backdrop = mountBackdrop(stage, dlg);
|
||||
dlg.querySelector('[data-act=cancel]').onclick = () => clearStage(stage);
|
||||
dlg.querySelector('[data-act=confirm]').onclick = () => {
|
||||
// No validation — just shows another alert if empty (current bug)
|
||||
const v = dlg.querySelector('input').value.trim();
|
||||
if (!v) { window.alert('Completion summary is required before marking a task done.'); return; }
|
||||
clearStage(stage);
|
||||
};
|
||||
backdrop.onclick = (e) => { if (e.target === backdrop) clearStage(stage); };
|
||||
}
|
||||
|
||||
/* ----- Variant B: Strong-fit ----------------------------------- */
|
||||
function B_moveDone(stage) {
|
||||
clearStage(stage);
|
||||
const dlg = document.createElement('div');
|
||||
dlg.className = 'dialog';
|
||||
dlg.innerHTML = `
|
||||
<h3>${ICONS.check} Mark this task as done?</h3>
|
||||
<p class="desc">The worker's claim is released and dependent children become ready.</p>
|
||||
<div class="body">
|
||||
<textarea rows="3" placeholder="Completion summary — this is stored as the task result."></textarea>
|
||||
<div class="helper" data-helper></div>
|
||||
</div>
|
||||
<div class="actions">
|
||||
<button data-act="cancel">Cancel</button>
|
||||
<button data-act="confirm" class="primary">Mark Done</button>
|
||||
</div>
|
||||
`;
|
||||
const backdrop = mountBackdrop(stage, dlg);
|
||||
const ta = dlg.querySelector('textarea');
|
||||
const helper = dlg.querySelector('[data-helper]');
|
||||
const confirm = dlg.querySelector('[data-act=confirm]');
|
||||
// Pro suggested either disabled OR error — let's show the error path
|
||||
confirm.onclick = () => {
|
||||
if (!ta.value.trim()) {
|
||||
helper.classList.add('error');
|
||||
helper.textContent = 'Completion summary is required before marking a task done.';
|
||||
ta.classList.add('invalid');
|
||||
ta.focus();
|
||||
return;
|
||||
}
|
||||
// Simulate PATCH
|
||||
confirm.disabled = true;
|
||||
confirm.innerHTML = '<span class="spinner"></span>Marking done…';
|
||||
setTimeout(() => clearStage(stage), 1200);
|
||||
};
|
||||
dlg.querySelector('[data-act=cancel]').onclick = () => clearStage(stage);
|
||||
backdrop.onclick = (e) => { if (e.target === backdrop) clearStage(stage); };
|
||||
}
|
||||
|
||||
/* ----- Variant B-refined: synthesis ---------------------------- */
|
||||
function B2_moveDone(stage, count) {
|
||||
clearStage(stage);
|
||||
const isBulk = count && count > 1;
|
||||
const label = isBulk ? `${count} selected tasks` : 'this task';
|
||||
const title = isBulk ? `Mark ${count} tasks as done?` : 'Mark this task as done?';
|
||||
const dlg = document.createElement('div');
|
||||
dlg.className = 'dialog';
|
||||
dlg.innerHTML = `
|
||||
<h3>${ICONS.check} ${title}</h3>
|
||||
<p class="desc">The worker's claim is released and dependent children become ready.</p>
|
||||
<div class="body">
|
||||
<textarea rows="3" autofocus placeholder="Completion summary for ${label}. This is stored as the task result."></textarea>
|
||||
<div class="helper" data-helper></div>
|
||||
${isBulk ? bulkList() : ''}
|
||||
</div>
|
||||
<div class="actions">
|
||||
<button data-act="cancel">Cancel</button>
|
||||
<button data-act="confirm" class="primary" disabled>Mark Done</button>
|
||||
</div>
|
||||
`;
|
||||
const backdrop = mountBackdrop(stage, dlg);
|
||||
const ta = dlg.querySelector('textarea');
|
||||
const helper = dlg.querySelector('[data-helper]');
|
||||
const confirm = dlg.querySelector('[data-act=confirm]');
|
||||
const cancel = dlg.querySelector('[data-act=cancel]');
|
||||
const REQUIRED = 'Completion summary is required before marking a task done.';
|
||||
|
||||
// Auto-focus + ensure visible on mobile keyboards (best-effort)
|
||||
setTimeout(() => {
|
||||
ta.focus();
|
||||
ta.scrollIntoView({ block: 'center', behavior: 'smooth' });
|
||||
}, 50);
|
||||
|
||||
ta.addEventListener('input', () => {
|
||||
const valid = ta.value.trim().length > 0;
|
||||
confirm.disabled = !valid;
|
||||
if (valid) {
|
||||
ta.classList.remove('invalid');
|
||||
helper.classList.remove('error');
|
||||
helper.textContent = '';
|
||||
}
|
||||
});
|
||||
ta.addEventListener('blur', () => {
|
||||
if (!ta.value.trim()) {
|
||||
ta.classList.add('invalid');
|
||||
helper.classList.add('error');
|
||||
helper.textContent = REQUIRED;
|
||||
}
|
||||
});
|
||||
confirm.onclick = () => {
|
||||
if (!ta.value.trim()) {
|
||||
ta.classList.add('invalid');
|
||||
helper.classList.add('error');
|
||||
helper.textContent = REQUIRED;
|
||||
ta.focus();
|
||||
return;
|
||||
}
|
||||
// Cancel stays enabled; only confirm gets the spinner (GPT-OSS fix)
|
||||
confirm.disabled = true;
|
||||
confirm.innerHTML = '<span class="spinner"></span>Marking done…';
|
||||
setTimeout(() => clearStage(stage), 1200);
|
||||
};
|
||||
cancel.onclick = () => clearStage(stage);
|
||||
backdrop.onclick = (e) => { if (e.target === backdrop) clearStage(stage); };
|
||||
// Cmd/Ctrl + Enter submits
|
||||
ta.addEventListener('keydown', (e) => {
|
||||
if ((e.metaKey || e.ctrlKey) && e.key === 'Enter' && !confirm.disabled) {
|
||||
confirm.click();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
/* Helper: bulk list markup for B-refined bulk-many dialogs */
|
||||
function bulkList() {
|
||||
const items = [
|
||||
{ id: 't_8c41', title: 'Refactor kanban dispatcher tick interval' },
|
||||
{ id: 't_8d12', title: 'Add batch delete endpoint smoke test' },
|
||||
{ id: 't_8e09', title: 'Document goal_mode lifecycle in user guide' },
|
||||
];
|
||||
return `
|
||||
<div class="bulk-list">
|
||||
${items.map(t => `
|
||||
<div class="bulk-item">
|
||||
<div class="meta">
|
||||
<div class="title">${t.title}</div>
|
||||
<div class="id">${t.id}</div>
|
||||
<textarea rows="1" placeholder="Completion summary for this task…"></textarea>
|
||||
</div>
|
||||
</div>
|
||||
`).join('')}
|
||||
</div>
|
||||
`;
|
||||
}
|
||||
|
||||
/* ----- Generic destructive confirm (used by A/B for blocked/archive/delete) -- */
|
||||
function makeConfirm(stage, opts) {
|
||||
const { icon, title, desc, confirmLabel, destructive, requireReason } = opts;
|
||||
const dlg = document.createElement('div');
|
||||
dlg.className = 'dialog';
|
||||
dlg.innerHTML = `
|
||||
<h3>${icon} ${title}</h3>
|
||||
<p class="desc">${desc}</p>
|
||||
${requireReason ? `
|
||||
<div class="body">
|
||||
<textarea rows="2" placeholder="${requireReason}"></textarea>
|
||||
<div class="helper" data-helper></div>
|
||||
</div>
|
||||
` : ''}
|
||||
<div class="actions">
|
||||
<button data-act="cancel">Cancel</button>
|
||||
<button data-act="confirm" class="${destructive ? 'destructive' : 'primary'}">${confirmLabel}</button>
|
||||
</div>
|
||||
`;
|
||||
const backdrop = mountBackdrop(stage, dlg);
|
||||
const confirm = dlg.querySelector('[data-act=confirm]');
|
||||
if (requireReason) {
|
||||
const ta = dlg.querySelector('textarea');
|
||||
confirm.disabled = true;
|
||||
ta.addEventListener('input', () => { confirm.disabled = ta.value.trim().length === 0; });
|
||||
}
|
||||
dlg.querySelector('[data-act=cancel]').onclick = () => clearStage(stage);
|
||||
confirm.onclick = () => {
|
||||
confirm.disabled = true;
|
||||
confirm.innerHTML = '<span class="spinner"></span>Working…';
|
||||
setTimeout(() => clearStage(stage), 800);
|
||||
};
|
||||
backdrop.onclick = (e) => { if (e.target === backdrop) clearStage(stage); };
|
||||
}
|
||||
|
||||
const makeA_confirm = (s, o) => makeConfirm(s, o);
|
||||
const makeB_confirm = (s, o) => makeConfirm(s, o);
|
||||
const makeB2_confirm = (s, o) => makeConfirm(s, o);
|
||||
|
||||
/* ----- Variant C: undo toast ----------------------------------- */
|
||||
function C_undoToast(stage, message, desc) {
|
||||
const container = stage.querySelector('.toast-container');
|
||||
const toast = document.createElement('div');
|
||||
toast.className = 'toast';
|
||||
toast.innerHTML = `
|
||||
<div class="title">${ICONS.check} ${message}</div>
|
||||
${desc ? `<div class="desc">${desc}</div>` : ''}
|
||||
<div class="row">
|
||||
<span style="opacity:0.6; font-size: 10px;">Auto-dismiss in 5s</span>
|
||||
<button class="undo" data-act="undo">Undo</button>
|
||||
</div>
|
||||
`;
|
||||
container.appendChild(toast);
|
||||
setTimeout(() => toast.remove(), 5200);
|
||||
toast.querySelector('[data-act=undo]').onclick = () => {
|
||||
toast.style.opacity = '0';
|
||||
toast.style.transition = 'opacity 200ms';
|
||||
setTimeout(() => toast.remove(), 220);
|
||||
};
|
||||
}
|
||||
function C_errorToast(stage, message) {
|
||||
const container = stage.querySelector('.toast-container');
|
||||
const toast = document.createElement('div');
|
||||
toast.className = 'toast error';
|
||||
toast.dataset.sticky = '1'; // visual marker; doesn't actually stop dismissal
|
||||
toast.innerHTML = `
|
||||
<div class="title">${ICONS.x} ${message}</div>
|
||||
`;
|
||||
container.appendChild(toast);
|
||||
// Longer dismiss for screenshot stability
|
||||
setTimeout(() => toast.remove(), 30000);
|
||||
}
|
||||
|
||||
/* =========================================================== */
|
||||
/* Dispatch table — fires the right handler per variant */
|
||||
/* =========================================================== */
|
||||
const HANDLERS = {
|
||||
'move-done': () => {
|
||||
A_moveDone(stage('A'));
|
||||
B_moveDone(stage('B'));
|
||||
B2_moveDone(stage('B2'), 1);
|
||||
C_undoToast(stage('C'), 'Task marked as done', 'The worker claim has been released.');
|
||||
},
|
||||
'move-blocked': () => {
|
||||
const opts = (variant) => ({
|
||||
icon: variant === 'A' ? '' : ICONS.pause,
|
||||
title: 'Mark this task as blocked?',
|
||||
desc: "The worker's claim is released.",
|
||||
confirmLabel: 'Mark Blocked',
|
||||
destructive: false,
|
||||
});
|
||||
makeA_confirm(stage('A'), opts('A'));
|
||||
makeB_confirm(stage('B'), opts('B'));
|
||||
makeB2_confirm(stage('B2'), opts('B2'));
|
||||
C_undoToast(stage('C'), 'Task marked as blocked', 'Worker claim released. Worker will re-prompt on unblock.');
|
||||
},
|
||||
'bulk-delete': () => {
|
||||
const isBulk = true;
|
||||
const count = 3;
|
||||
const titleAll = isBulk
|
||||
? `Permanently delete ${count} selected tasks?`
|
||||
: 'Permanently delete this task?';
|
||||
const opts = (variant) => ({
|
||||
icon: variant === 'A' ? '' : ICONS.trash,
|
||||
title: titleAll,
|
||||
desc: 'This cannot be undone.',
|
||||
confirmLabel: isBulk ? `Delete ${count} tasks` : 'Delete',
|
||||
destructive: true,
|
||||
requireReason: variant === 'B2' ? 'Optional: why are these being deleted?' : null,
|
||||
});
|
||||
makeA_confirm(stage('A'), opts('A'));
|
||||
makeB_confirm(stage('B'), opts('B'));
|
||||
// B2 gets the bulk list version
|
||||
clearStage(stage('B2'));
|
||||
const dlg = document.createElement('div');
|
||||
dlg.className = 'dialog';
|
||||
dlg.innerHTML = `
|
||||
<h3>${ICONS.trash} Permanently delete 3 selected tasks?</h3>
|
||||
<p class="desc">This cannot be undone.</p>
|
||||
<div class="body">${bulkList()}</div>
|
||||
<div class="actions">
|
||||
<button data-act="cancel">Cancel</button>
|
||||
<button data-act="confirm" class="destructive">Delete 3 tasks</button>
|
||||
</div>
|
||||
`;
|
||||
const backdrop = mountBackdrop(stage('B2'), dlg);
|
||||
dlg.querySelector('[data-act=cancel]').onclick = () => clearStage(stage('B2'));
|
||||
dlg.querySelector('[data-act=confirm]').onclick = () => {
|
||||
dlg.querySelector('[data-act=confirm]').disabled = true;
|
||||
dlg.querySelector('[data-act=confirm]').innerHTML = '<span class="spinner"></span>Deleting…';
|
||||
setTimeout(() => clearStage(stage('B2')), 1200);
|
||||
};
|
||||
backdrop.onclick = (e) => { if (e.target === backdrop) clearStage(stage('B2')); };
|
||||
C_errorToast(stage('C'), 'Bulk delete is destructive — undo toast not appropriate.');
|
||||
},
|
||||
'archive-board': () => {
|
||||
const msg = "Archive board 'atm10-server'? It will be moved to boards/_archived/ so you can recover it later. Tasks on this board will no longer appear anywhere in the UI.";
|
||||
const opts = (variant) => ({
|
||||
icon: variant === 'A' ? '' : ICONS.archive,
|
||||
title: "Archive board 'atm10-server'?",
|
||||
desc: msg,
|
||||
confirmLabel: 'Archive',
|
||||
destructive: true,
|
||||
});
|
||||
makeA_confirm(stage('A'), opts('A'));
|
||||
makeB_confirm(stage('B'), opts('B'));
|
||||
makeB2_confirm(stage('B2'), opts('B2'));
|
||||
C_errorToast(stage('C'), 'Archive is destructive — undo toast not appropriate.');
|
||||
},
|
||||
'remove-attachment': () => {
|
||||
const opts = (variant) => ({
|
||||
icon: variant === 'A' ? '' : ICONS.paperclip,
|
||||
title: 'Remove this attachment?',
|
||||
desc: 'The file will be unlinked from this task. Other references are unaffected.',
|
||||
confirmLabel: 'Remove',
|
||||
destructive: true,
|
||||
});
|
||||
makeA_confirm(stage('A'), opts('A'));
|
||||
makeB_confirm(stage('B'), opts('B'));
|
||||
makeB2_confirm(stage('B2'), opts('B2'));
|
||||
C_undoToast(stage('C'), 'Attachment removed', 'design-spec-v3.pdf');
|
||||
},
|
||||
'error': () => {
|
||||
// Variant C: error toast is the only place toasts really shine.
|
||||
[stage('A'), stage('B'), stage('B2')].forEach(s => {
|
||||
const banner = document.createElement('div');
|
||||
banner.style.cssText = 'position:absolute; top:12px; left:12px; right:12px; padding:10px 12px; background: color-mix(in srgb, var(--destructive) 10%, var(--bg)); border: 1px solid var(--destructive); color: var(--destructive); border-radius: var(--radius-sm); font-size: 12px;';
|
||||
banner.textContent = 'Move failed: HTTP 409 — task is in status "todo", only running/ready/blocked can be completed.';
|
||||
s.appendChild(banner);
|
||||
});
|
||||
C_errorToast(stage('C'), 'Move failed: HTTP 409');
|
||||
},
|
||||
'clear': () => {
|
||||
[stage('A'), stage('B'), stage('B2'), stage('C')].forEach(s => clearStage(s));
|
||||
},
|
||||
};
|
||||
|
||||
function stage(name) { return document.querySelector(`[data-stage="${name}"]`); }
|
||||
|
||||
/* =========================================================== */
|
||||
/* Wire up page-level triggers */
|
||||
/* =========================================================== */
|
||||
document.querySelectorAll('header.page .controls button').forEach(btn => {
|
||||
btn.addEventListener('click', () => {
|
||||
const handler = HANDLERS[btn.dataset.trigger];
|
||||
if (handler) handler();
|
||||
});
|
||||
});
|
||||
</script>
|
||||
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,208 @@
|
||||
# Multiplexing Gateway
|
||||
|
||||
One gateway process can serve every profile in the install. The mode is opt-in
|
||||
(`gateway.multiplex_profiles`, default `false`), and everything it changes
|
||||
reverts the moment the flag is off. This document is the design rationale
|
||||
referenced from `agent/secret_scope.py` ("Workstream A"): what is isolated per
|
||||
profile, the mechanism that isolates it, and what deliberately stays
|
||||
process-global.
|
||||
|
||||
## Overview
|
||||
|
||||
Without multiplexing, one gateway process serves exactly one profile — its
|
||||
`.env`, sessions, skills, and platform adapters — and multi-profile installs
|
||||
run one process per profile. Multiplexing collapses that into a single
|
||||
process: the default profile plus every served named profile get their own
|
||||
adapters, secrets, sessions, and cron ticks, while sharing one event loop, one
|
||||
HTTP listener, one process lock, and one status surface.
|
||||
|
||||
The design constraint that shapes everything below: **profile A's turns must
|
||||
never observe profile B's state**. Secrets, homes, sessions, and adapter lanes
|
||||
are isolated per profile; anything that cannot yet be isolated fails closed or
|
||||
is documented as a known limitation at the end of this document.
|
||||
|
||||
## The mode flag
|
||||
|
||||
- Config: `gateway.multiplex_profiles: true` (also accepted at top level).
|
||||
Parsed in `gateway/config.py` with precedence env > config > default.
|
||||
- Env override: `GATEWAY_MULTIPLEX_PROFILES` accepts explicit truthy/falsy
|
||||
tokens only; a blank or unrecognized value returns "no override" so an empty
|
||||
deployment secret cannot shadow a config opt-in.
|
||||
- At startup, `GatewayRunner.__init__` calls
|
||||
`agent.secret_scope.set_multiplex_active(...)` once. `_MULTIPLEX_ACTIVE` is
|
||||
a plain module global, not a contextvar: it describes the deployment mode,
|
||||
not a per-task value. Its only job is to arm the fail-closed behavior in
|
||||
`get_secret()`.
|
||||
|
||||
## Scope composition
|
||||
|
||||
Every inbound event composes the same two context-local scopes before any
|
||||
profile-owned code runs:
|
||||
|
||||
```
|
||||
platform event
|
||||
│
|
||||
▼
|
||||
profile_routes match ──► served-set check ──► SessionSource.profile stamped
|
||||
│ (gateway/profile_routing.py)
|
||||
▼
|
||||
_profile_runtime_scope(profile_home) (gateway/run.py)
|
||||
├── set_hermes_home_override(home) config / state.db / skills /
|
||||
│ memory / sessions resolve here
|
||||
└── set_secret_scope(profile .env + secret sources)
|
||||
│ provider keys, platform tokens
|
||||
▼
|
||||
agent turn (worker thread via copy_context())
|
||||
│
|
||||
▼
|
||||
scope unwound in finally
|
||||
```
|
||||
|
||||
`_profile_runtime_scope` wraps every seam where profile-owned code executes:
|
||||
secondary adapter startup, connect and reconnect, the primary platform event
|
||||
handler, inbound preprocessing, `/model` and session-info resolution,
|
||||
background tasks, and the agent turn itself. Config reloads run under the
|
||||
default profile's scope so global gateway settings (`#64674`) resolve
|
||||
consistently.
|
||||
|
||||
Both scopes are `contextvars`, so they propagate into executor worker threads
|
||||
via `copy_context()` and unwind deterministically — nothing is written to
|
||||
`os.environ`, ever.
|
||||
|
||||
## Workstream A: context-local secret scope
|
||||
|
||||
`agent/secret_scope.py` exists because the obvious implementation — union all
|
||||
profile `.env` files into `os.environ` — leaks profile A's keys into profile
|
||||
B's turns and into every subprocess spawned with `env=dict(os.environ)`.
|
||||
|
||||
- `build_profile_secret_scope(home)` merges the profile's `.env` with its
|
||||
configured secret sources, skipping globals.
|
||||
- `set_secret_scope(mapping)` installs it for the current task.
|
||||
- `get_secret(name)` resolves: global allowlist → active scope → fallback.
|
||||
The fallback is the load-bearing part:
|
||||
- multiplexing **off**: reads `os.environ`, so single-profile gateways and
|
||||
every non-gateway caller behave exactly as before;
|
||||
- multiplexing **on**, no scope installed: **raises `UnscopedSecretError`**
|
||||
rather than silently reading the process environment. An un-migrated call
|
||||
site fails loud at that exact line instead of leaking another profile's
|
||||
value.
|
||||
- A small allowlist (`HERMES_HOME`, `HERMES_PROFILE`, proxy settings,
|
||||
`API_SERVER_*` listener settings — but deliberately not `API_SERVER_KEY`)
|
||||
stays global because those describe the process, not a profile.
|
||||
|
||||
Because the per-turn `.env` reload is a no-op under multiplexing, rotated
|
||||
credentials are picked up through the profile scope on the next turn — never
|
||||
via `os.environ`. This holds at the loader boundary, not just the gateway's
|
||||
reload helper: `hermes_cli.env_loader.load_hermes_dotenv` skips the
|
||||
process-global load whenever multiplexing is active *and* a profile-home
|
||||
override is installed (import-time and cron callers hit it mid-turn), while
|
||||
still hydrating the profile's external secret sources into its private
|
||||
snapshot (`#77562`). The unscoped startup load is unchanged.
|
||||
|
||||
The same scope-authoritative rule covers the other `os.environ` seams a
|
||||
routed turn can reach: `${VAR}` / `${env:VAR}` references in a profile's
|
||||
`config.yaml` resolve through `get_secret` when a scope is installed
|
||||
(`#84079`), and `.env` writes made under a scope (`save_env_value`, e.g. a
|
||||
`/pair` grant mirror) update the installed scope mapping instead of the
|
||||
process environment (`#88441`).
|
||||
|
||||
## The HERMES_HOME override
|
||||
|
||||
`hermes_constants.py` holds a context-local override consulted by
|
||||
`get_hermes_home()` before the `HERMES_HOME` env var. Everything that resolves
|
||||
paths through it — config, `state.db`, skills, memory, SOUL, sessions, kanban,
|
||||
goals, plugin discovery, MCP startup — follows the active profile
|
||||
automatically. `get_process_hermes_home()` exists for the few machine-level
|
||||
assets that must not follow the override. `hermes_home_key()` gives
|
||||
per-home registries a stable scope key. A one-shot warning (`#18594`) fires if
|
||||
profile-scoped code runs without the override where one is expected.
|
||||
|
||||
## Inbound routing
|
||||
|
||||
`gateway.profile_routes` maps `(platform, guild_id, chat_id, thread_id)` to a
|
||||
profile; matching is conjunctive, most-specific-first, with parent-chain chat
|
||||
matching for threads. Routing only runs when multiplexing is active, and a
|
||||
matched route whose target is outside the served set is rejected (the event is
|
||||
dropped, not misdelivered). Full schema and matching rules:
|
||||
`docs/profile-routing.md`.
|
||||
|
||||
## Serving selected profiles
|
||||
|
||||
`profiles_to_serve(multiplex, profile_allowlist)` in `hermes_cli/profiles.py`
|
||||
is the single chokepoint for which profiles a multiplexer serves: default plus
|
||||
every valid profile directory, optionally filtered by allowlist. A malformed
|
||||
allowlist fails safe to default-only. The served set gates adapter startup,
|
||||
cron ticking (`#69377`), `/p/<profile>/` HTTP admission, route eligibility,
|
||||
and the runtime status surface. An excluded profile stays installed and can
|
||||
still run its own standalone gateway.
|
||||
|
||||
## Per-profile persistence
|
||||
|
||||
`SessionStore` binds no database handle at construction (`#88532`). Session
|
||||
DB handles are resolved at call time through the active HERMES_HOME override —
|
||||
one cached handle per resolved `profiles/<name>/state.db` — so sessions land
|
||||
in the owning profile's store even when the store object itself is shared.
|
||||
Pairing stores are constructed per served profile.
|
||||
|
||||
## Per-bot session lanes
|
||||
|
||||
Session keys are namespaced by profile (`agent:main` for default,
|
||||
`agent:<name>` for named profiles). Adapters carry `_owner_profile`
|
||||
(installed at adapter configuration time, before any inbound event) because
|
||||
adapter ingress runs before `SessionSource.profile` is stamped;
|
||||
`_session_key_profile` resolves source stamp → owner profile → store
|
||||
resolver. Text/media batching, active-session tracking, and the busy-session
|
||||
guard are all keyed per lane, so two bots sharing a chat do not share a
|
||||
session lane.
|
||||
|
||||
## Control plane
|
||||
|
||||
Desktop plugins reach the gateway only through the ws JSON-RPC door, so
|
||||
profile enumeration and configuration live in
|
||||
`tui_gateway/methods_profiles.py`: `profiles.list`, `profiles.create`,
|
||||
`profiles.describe`, `profiles.configure`, `profiles.set_asset`,
|
||||
`profiles.get_asset`. Reads and writes run under the target profile's
|
||||
HERMES_HOME override. Asset writes are atomic, type- and size-capped.
|
||||
|
||||
## Failure modes
|
||||
|
||||
- Fatal at startup: multiplex config errors and a secondary profile enabling a
|
||||
port-binding platform (`MultiplexConfigError`,
|
||||
`SecondaryPortBindingConfigError`) — one shared HTTP listener is owned by
|
||||
the default profile.
|
||||
- Skipped, not fatal: a single misconfigured secondary adapter is skipped with
|
||||
a warning rather than taking down the multiplexer.
|
||||
- Fail-closed: unscoped `get_secret()` under multiplexing raises; a routed
|
||||
event targeting an unserved profile is dropped; an unscoped `/p/` request
|
||||
enters the default profile's scope (`#61276`) rather than an undefined one.
|
||||
- Fallback: an external `cron.provider` does not support multiplexing and
|
||||
falls back to the built-in ticker with a warning.
|
||||
|
||||
## Known limitations
|
||||
|
||||
Process-global state that is not yet profile-scoped:
|
||||
|
||||
| Surface | State at time of writing |
|
||||
| --- | --- |
|
||||
| MCP discovery and tool registration | Process-global; the first profile to build an agent wins the discovery slot. Full per-profile MCP registries are tracked in `#67605`. |
|
||||
| Terminal / sandbox env (`TERMINAL_*`) | Global by allowlist; tools read it from the process environment. |
|
||||
| Built-in tool registry | Built-ins are process-global; plugin-registered tools are overlaid per profile via `hermes_home_key()`. |
|
||||
| Provider/capability registries | Same hybrid overlay pattern (browser, image-gen, TTS, transcription, video-gen, web-search, secret sources). |
|
||||
| HTTP listener, relay ingress, process lock | One per process, owned by the default/active profile. Per-profile `runtime_status.json` is still written. |
|
||||
|
||||
## Non-goals
|
||||
|
||||
Multiplexing isolates *profiles*; it does not authenticate or authorize *end
|
||||
users*. A profile is a configuration, not a person: the gateway trusts its
|
||||
transport and its routing table to decide which profile an event belongs to.
|
||||
Request-level identity and per-user authorization above the profile layer are
|
||||
out of scope for this document.
|
||||
|
||||
## Related
|
||||
|
||||
- `docs/profile-routing.md` — inbound routing schema and matching rules.
|
||||
- `website/docs/user-guide/multi-profile-gateways.md` — user-facing guide,
|
||||
including the standalone one-gateway-per-profile alternative.
|
||||
- `agent/secret_scope.py`, `hermes_constants.py`, `gateway/profile_routing.py`,
|
||||
`gateway/run.py` (`_profile_runtime_scope`), `hermes_cli/profiles.py`
|
||||
(`profiles_to_serve`), `gateway/session.py`, `tui_gateway/methods_profiles.py`.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Profile Builder — Dashboard-Native, Full-Featured Profile Creation
|
||||
|
||||
Status: design proposal (not yet implemented)
|
||||
Author: drafted for Teknium
|
||||
Supersedes: PR #31781 (prompt_toolkit `hermes profile wizard`)
|
||||
|
||||
## Why this, not the CLI wizard
|
||||
|
||||
PR #31781 added a keyboard-driven `hermes profile wizard` in the terminal.
|
||||
The decision is to **not** build the profile-creation experience in the CLI.
|
||||
The dashboard already owns mature, separate pages for every element a profile
|
||||
needs, and a profile is just a HERMES_HOME directory — so the dashboard is the
|
||||
right home for a full-featured builder, and it can reuse everything that
|
||||
already exists.
|
||||
|
||||
A profile = a full `~/.hermes/profiles/<name>/` directory with its own:
|
||||
- `config.yaml` — holds `model`/`provider`, `mcp_servers`, enabled skills
|
||||
- `skills/` — physical SKILL.md files (built-in seed + optional + hub installs)
|
||||
- `.env` — secrets
|
||||
- `SOUL.md` / `USER.md` — identity
|
||||
|
||||
So per-profile scoping of Model, MCPs, and Skills is **native** — no data-model
|
||||
change needed. The gap is purely UX: creation today is a thin modal
|
||||
(name + clone + model + description), and you can only compose skills/MCPs
|
||||
*after* the profile exists, by visiting other pages and remembering to scope
|
||||
them.
|
||||
|
||||
## What already exists (reuse, don't rebuild)
|
||||
|
||||
| Element | Existing page | Existing API | Profile-scopable? |
|
||||
|---|---|---|---|
|
||||
| Name / Description | ProfilesPage create modal | `POST /api/profiles` (`create_profile`) | yes (args) |
|
||||
| Model + Provider | ModelsPage | `_write_profile_model(profile_dir, …)` | yes — HERMES_HOME override, already wired into create endpoint |
|
||||
| MCPs | McpPage | `mcp_config._save_mcp_server` + `/api/mcp/catalog` | yes — wrap with HERMES_HOME override |
|
||||
| Skills (built-in/optional) | SkillsPage | `GET /api/skills`, `/api/skills/toggle` | yes — config write |
|
||||
| Skills (hub) | SkillsPage | `/api/skills/hub/search`, `/api/skills/hub/install` | **only via subprocess** — see seam #1 |
|
||||
|
||||
## Two architectural seams found while grounding this design
|
||||
|
||||
These are load-bearing — they change the implementation, not just the polish.
|
||||
|
||||
### Seam #1 — hub-skill install cannot use the HERMES_HOME override
|
||||
|
||||
`tools/skills_hub.py` binds `SKILLS_DIR = HERMES_HOME / "skills"` at **module
|
||||
import time**. The context-local `set_hermes_home_override()` swap (which makes
|
||||
`_write_profile_model` and the MCP write land in the target profile) does NOT
|
||||
retroactively rebind that already-imported module global. So a data-layer wrap
|
||||
of hub install would write into the dashboard's *own* active profile, not the
|
||||
new one.
|
||||
|
||||
The correct mechanism is the existing subprocess path: `_spawn_hermes_action`
|
||||
runs `python -m hermes_cli.main <subcommand>`, and `_apply_profile_override()`
|
||||
re-reads `sys.argv` at import in the fresh child. Prepend `-p <profile>`:
|
||||
|
||||
```python
|
||||
_spawn_hermes_action(["-p", profile, "skills", "install", identifier], "skills-install")
|
||||
```
|
||||
|
||||
A fresh subprocess re-imports `skills_hub` with the profile's HERMES_HOME bound
|
||||
from the start, so `SKILLS_DIR` resolves to `<profile>/skills/`. Correct by
|
||||
construction.
|
||||
|
||||
### Seam #2 — hub installs are async, so create cannot be fully atomic
|
||||
|
||||
Built-in/optional skill enabling and MCP writes are **synchronous config ops**
|
||||
and can be part of the create call. Hub installs are long-running git fetches
|
||||
spawned detached (`_spawn_hermes_action` returns a PID immediately). So the
|
||||
create flow is:
|
||||
|
||||
1. `create_profile()` — make the dir (synchronous)
|
||||
2. write model (synchronous, HERMES_HOME override)
|
||||
3. write selected MCP servers (synchronous, HERMES_HOME override)
|
||||
4. seed/enable selected built-in + optional skills (synchronous)
|
||||
5. spawn `hermes -p <profile> skills install <id>` per hub skill (async, returns PIDs)
|
||||
|
||||
Steps 1–4 commit before the response; step 5 returns a list of action PIDs the
|
||||
UI polls (same pattern as today's SkillsPage hub install). The builder's
|
||||
"Review → Create" returns `{ok, name, path, hub_installs: [{id, pid}]}` and the
|
||||
final screen shows live install progress for the hub skills.
|
||||
|
||||
## Proposed backend change (small, follows existing patterns)
|
||||
|
||||
Extend `ProfileCreate` and the create endpoint — no new endpoints, no rewrite:
|
||||
|
||||
```python
|
||||
class ProfileCreate(BaseModel):
|
||||
name: str
|
||||
clone_from: Optional[str] = None
|
||||
# Backward compatibility for older dashboard/desktop clients.
|
||||
clone_from_default: bool = False
|
||||
clone_all: bool = False
|
||||
no_skills: bool = False
|
||||
description: Optional[str] = None
|
||||
provider: Optional[str] = None
|
||||
model: Optional[str] = None
|
||||
# NEW — all optional, all best-effort post-create (profile already exists)
|
||||
mcp_servers: List[MCPServerCreate] = [] # synchronous, HERMES_HOME override
|
||||
builtin_skills: List[str] = [] # synchronous enable/seed
|
||||
hub_skills: List[str] = [] # async spawn, returns PIDs
|
||||
```
|
||||
|
||||
The endpoint already does best-effort post-create steps (`seed_profile_skills`,
|
||||
`_write_profile_model`). Add two more best-effort blocks (MCP write, hub-skill
|
||||
spawn) in the same style — a failure in any of them must not 500 the create,
|
||||
since the profile dir already exists and the user can fix it from the relevant
|
||||
page afterward. Mirror `_write_profile_model`'s HERMES_HOME-override helper for
|
||||
the MCP write (`_write_profile_mcp_servers(profile_dir, servers)`).
|
||||
|
||||
## Proposed frontend — dedicated builder page `/profiles/new`
|
||||
|
||||
A full page (not the cramped modal), stepped, each step reusing the existing
|
||||
page's component + API, targeted at the new profile:
|
||||
|
||||
```
|
||||
① Identity Name + Description (+ optional clone-from existing profile)
|
||||
② Model Provider + model picker (reuse ModelsPage picker)
|
||||
③ Skills Tabs: Built-in · Optional · Hub-search
|
||||
multi-select; "Start from default bundle" preset button
|
||||
④ MCPs Tabs: Catalog browse · Manual add (reuse McpPage form)
|
||||
⑤ Review Blueprint preview → Create
|
||||
→ progress screen for async hub installs
|
||||
```
|
||||
|
||||
Nothing writes to disk until ⑤.
|
||||
|
||||
## Open product decisions (need Teknium)
|
||||
|
||||
1. **Skills seeding default.** Fresh profiles auto-seed the default bundle
|
||||
today. In the builder, should the skill step **replace** the bundle (pick
|
||||
exactly what you want; offer a "start from default bundle" preset) or
|
||||
**augment** it? Recommendation: replace + preset button.
|
||||
|
||||
2. **Page vs richer modal.** Dedicated `/profiles/new` page (room to grow:
|
||||
SOUL editing, multi-agent fleets later) vs a bigger create modal on
|
||||
ProfilesPage. Recommendation: dedicated page — matches "full-featured / way
|
||||
more options."
|
||||
|
||||
## Verification plan (when built)
|
||||
|
||||
- Backend E2E with isolated HERMES_HOME: POST a full create body
|
||||
(name + model + 2 MCPs + 3 builtin skills + 1 hub skill), assert the new
|
||||
profile dir has the model in config.yaml, both MCP servers in config.yaml,
|
||||
the builtin skills enabled, and a spawned PID for the hub skill. Negative:
|
||||
a bad MCP entry must not 500 the create.
|
||||
- `cd web && npm run build` (no JS test suite in web/).
|
||||
- Targeted: `pytest tests/<web_server profile tests> -k profile_create`.
|
||||
Binary file not shown.
@@ -0,0 +1,47 @@
|
||||
# Multi-gateway deployment
|
||||
|
||||
Hermes supports multiple gateway processes running concurrently — one per profile
|
||||
(default, writer, admin, coder, researcher). Each gateway opens its own connection
|
||||
to platform APIs and delivers messages for its profile's subscribers.
|
||||
|
||||
Task subscriptions also cover review feedback. A `changes_requested` review
|
||||
event is delivered as an actionable review-BLOCK notification. Subscriptions
|
||||
using `notify+wake` additionally wake the exact originating chat/thread/session
|
||||
so the controller inspects the existing card and current run; `notify` remains
|
||||
passive-only and `wake` remains wake-only. Review feedback never creates,
|
||||
unblocks, requeues, or otherwise mutates a task.
|
||||
|
||||
## Single-dispatcher posture
|
||||
|
||||
Only one gateway owns the kanban dispatcher. The owning gateway keeps
|
||||
`kanban.dispatch_in_gateway: true` (the default); every other gateway sets it
|
||||
to `false`.
|
||||
|
||||
**Why this matters:** dispatching is single-owner so multiple gateways do not
|
||||
race to spawn the same work. Notification delivery is profile-owned instead:
|
||||
each gateway polls only subscriptions for profiles whose platform adapters it
|
||||
hosts. The atomic event claim prevents duplicate delivery across watcher
|
||||
processes.
|
||||
|
||||
## Configuration
|
||||
|
||||
On the dispatch-owning gateway (typically the `default` profile), no change is
|
||||
needed. On every other profile gateway, add to `~/.hermes/config.yaml`:
|
||||
|
||||
```yaml
|
||||
kanban:
|
||||
dispatch_in_gateway: false
|
||||
```
|
||||
|
||||
Or set the env var: `HERMES_KANBAN_DISPATCH_IN_GATEWAY=false`
|
||||
|
||||
## What each gateway does
|
||||
|
||||
| Gateway role | dispatch_in_gateway | Opens subscribed board DBs? | Dispatcher | Notifier |
|
||||
|---|---|---|---|---|
|
||||
| default (confirmed dispatch-lock owner) | true (default) | yes | yes | owned profiles + legacy unstamped subscriptions |
|
||||
| writer, admin, coder, etc. | false | yes, when the profile has subscriptions | no | that gateway's owned profiles |
|
||||
|
||||
Non-dispatch gateways still deliver messages for their own platform adapters
|
||||
(Telegram, Discord, etc.). They do not dispatch tasks, and they skip boards
|
||||
that have no subscriptions owned by their profiles.
|
||||
@@ -0,0 +1,401 @@
|
||||
# Micro-compaction
|
||||
|
||||
**A way to amortize the cost of compression.**
|
||||
|
||||
Long conversations eventually outgrow the model's context window, and something
|
||||
has to be thrown away or summarized. Hermes has always done this in one batch:
|
||||
when the transcript crosses a threshold, the session stops, a large chunk of the
|
||||
middle is summarized in a single call, and the conversation resumes. That works,
|
||||
but the whole bill comes due at once — one visible pause, one big summarization
|
||||
request, at whatever moment you happened to cross the line.
|
||||
|
||||
Micro-compaction pays the same bill in instalments. After each completed turn,
|
||||
Hermes folds the single oldest un-absorbed exchange into a running summary. The
|
||||
work is the same work; it just happens continuously, a piece at a time, instead
|
||||
of all at once in the middle of your session.
|
||||
|
||||
It is not free and it is not a magic bullet, and it is **off by default** —
|
||||
`compression.micro_compact: true` turns it on. Each pass is a real call to the
|
||||
compression model, and it runs at the end of a turn — your answer has already
|
||||
streamed, but the turn does not close until the pass finishes. Each pass also
|
||||
rewrites already-sent history, which breaks the provider prompt-cache prefix
|
||||
every turn; read [Prompt caching](#prompt-caching--the-cost-you-are-opting-into)
|
||||
before enabling it, because for some setups that cost exceeds the benefit.
|
||||
|
||||
What the feature gives you is a **tuning option**: you choose how the
|
||||
compression cost is distributed, and which model pays it. See
|
||||
[Choosing a compression model](#choosing-a-compression-model), because that
|
||||
choice matters more than anything else here.
|
||||
|
||||
**The tradeoff is that knowledge gets a little earlier than you may be used to.**
|
||||
Because compaction is always running, older parts of the conversation become
|
||||
summaries sooner than they would under batch compaction — which leaves
|
||||
everything verbatim until the window actually fills. Detail from earlier in the
|
||||
session turns second-hand faster. You trade some of that fidelity for never
|
||||
eating one long stall, and for a context window that stays consistently smaller
|
||||
rather than sawtoothing up to the threshold and back.
|
||||
|
||||
---
|
||||
|
||||
## What it does
|
||||
|
||||
After every turn that finishes normally, `finalize_turn` asks the context
|
||||
compressor to absorb **one** exchange:
|
||||
|
||||
1. Find the oldest exchange that hasn't been summarized yet.
|
||||
2. Send just that exchange, plus the current running summary, to the auxiliary
|
||||
summarization model.
|
||||
3. Replace those messages in the transcript with a single summary marker
|
||||
carrying the updated running summary.
|
||||
|
||||
One exchange per turn. The per-turn cost stays bounded no matter how long the
|
||||
conversation gets.
|
||||
|
||||
An **exchange** is one full agent turn: an assistant message together with its
|
||||
tool results and any follow-up assistant iterations, up to the next user
|
||||
message. In tool-heavy work that's where the bulk of the tokens live — a
|
||||
file read or a command's output dwarfs the surrounding prose — which is why
|
||||
absorbing one exchange at a time is worth doing at all. Taking the whole turn
|
||||
(rather than a single assistant+tools group) also keeps the transcript's role
|
||||
alternation strictly valid: the summary marker is an assistant-role message,
|
||||
and a full turn is always bounded by user messages on both sides.
|
||||
|
||||
## Your messages are never compacted
|
||||
|
||||
An exchange deliberately starts at the *assistant* message. Micro-compaction
|
||||
walks straight past user messages to get there, so **what you typed is never
|
||||
summarized** — your prompts stay verbatim for the entire session, no matter how
|
||||
long it runs or how many times compaction fires.
|
||||
|
||||
This is the most useful property of the whole design, and it's worth being
|
||||
explicit about why. What the assistant produces is largely an account of what it
|
||||
did: it read this file, it ran that command, it got this result. That kind of
|
||||
narration survives summarising with very little loss — "it did it this way" is
|
||||
about as informative compressed as it was in full. Your instructions are a
|
||||
different kind of thing. They're the intent everything else is derived from, and
|
||||
they cannot be reconstructed from the work that followed. Paraphrasing "use the
|
||||
existing retry helper, don't add a new one" into a summary is exactly how an
|
||||
agent ends up confidently doing the thing you told it not to, six turns later.
|
||||
|
||||
So the asymmetry is on purpose: compact the derived material, keep the source of
|
||||
truth. The cost is a floor on how small the middle can get, since user turns
|
||||
accumulate and are never absorbed. In practice that floor is low — a prompt is
|
||||
normally a tiny fraction of what a single tool result costs — but it is a real
|
||||
floor. If you routinely paste 10–20K-token prompts, that weight stays in context
|
||||
by design.
|
||||
|
||||
## What it never touches
|
||||
|
||||
Two more regions are protected and stay verbatim:
|
||||
|
||||
- **The head** — the system prompt and the opening messages, so the session's
|
||||
founding instructions are never paraphrased.
|
||||
- **The tail** — a token-budgeted window of the most recent messages, so
|
||||
everything that's immediately relevant is still there in full.
|
||||
|
||||
Micro-compaction only ever works in the middle, between those two.
|
||||
|
||||
## How it works
|
||||
|
||||
### The cursor
|
||||
|
||||
The compressor keeps a cursor: the index of the first message not yet absorbed.
|
||||
Each successful pass advances it past the exchange it just summarized.
|
||||
|
||||
If that in-memory cursor is missing or out of range — a fresh process, a resumed
|
||||
session — it's recovered by scanning the transcript for the last summary marker
|
||||
and resuming just after it. The transcript itself is the source of truth, so
|
||||
resuming a session doesn't re-summarize work already done.
|
||||
|
||||
### The rolling summary
|
||||
|
||||
Rather than keeping a pile of per-exchange summaries, there is exactly one
|
||||
running summary that each new exchange is merged into. The summarizer is asked
|
||||
to fold in the new material's decisions, requirements, file paths and open
|
||||
questions, drop details that are no longer relevant, and preserve the existing
|
||||
structure. It's also explicitly instructed to replace any credentials it
|
||||
encounters with `[REDACTED]`.
|
||||
|
||||
Because that summary is cumulative, only the newest marker is kept in the
|
||||
transcript. Earlier markers are strictly redundant — the current summary already
|
||||
contains everything they held — so they're dropped as they're superseded. This
|
||||
matters more than it sounds: leaving them in place stacks near-duplicate copies
|
||||
of the same text, each with its own heading and end-marker scaffolding, and the
|
||||
transcript grows on every turn instead of shrinking.
|
||||
|
||||
### Defrag
|
||||
|
||||
Merge into a summary often enough and it gets baggy — repetitive, and larger
|
||||
than the material justifies. When the running summary crosses a token threshold
|
||||
(2000 by default), the next pass **defrags**: one auxiliary call re-summarizes
|
||||
the running summary *itself* into a fresh compact version, and the summary
|
||||
marker in the transcript is rewritten in place.
|
||||
|
||||
Defrag never touches the transcript's structure — no messages are absorbed or
|
||||
spliced, the cursor doesn't move, and user turns are untouched. It processes
|
||||
only the accumulated summary text, never conversation messages, so the
|
||||
"your messages are never compacted" guarantee holds through it.
|
||||
|
||||
### Staying in step with the session database
|
||||
|
||||
The in-memory splice alone isn't enough. Hermes's normal session flush is
|
||||
append-only, so the original rows would stay marked active and a resume would
|
||||
load *both* the summary and the messages it replaced — putting the session
|
||||
straight over the context limit.
|
||||
|
||||
So each pass also calls `archive_and_compact`, which atomically soft-archives the
|
||||
active rows and inserts the compacted set. The messages are then stamped as
|
||||
already-persisted so the append-only flush that follows skips them. If that
|
||||
database step fails, it's logged and the session continues; the resume would
|
||||
double-load until the next batch compression cleans up.
|
||||
|
||||
### When the summarizer fails
|
||||
|
||||
A summarization call can fail — the auxiliary model is unreachable, out of quota,
|
||||
or the exchange itself is somehow unsummarizable. The transcript is left
|
||||
untouched and the failure is counted.
|
||||
|
||||
If the *same* exchange fails three times in a row, the cursor is advanced past it
|
||||
anyway. Without that, one bad exchange would be retried on every single turn
|
||||
forever. Those skipped messages stay in the transcript and get picked up by the
|
||||
next batch compaction.
|
||||
|
||||
## Interaction with batch compaction
|
||||
|
||||
Micro-compaction doesn't replace batch compaction — it defers it. Threshold-based
|
||||
compaction is still there and still fires if the window fills anyway, and its
|
||||
summary markers are the same format, so the two interoperate. In practice
|
||||
micro-compaction keeps the transcript far enough below the threshold that the
|
||||
batch path fires much less often.
|
||||
|
||||
## Configuration
|
||||
|
||||
Micro-compaction is **off by default**. Turn it on explicitly:
|
||||
|
||||
```yaml
|
||||
compression:
|
||||
micro_compact: true # default: false
|
||||
micro_compact_every_n_turns: 1 # cadence — how often a pass runs
|
||||
micro_compact_defrag_threshold_tokens: 2000
|
||||
```
|
||||
|
||||
With `micro_compact` unset or `false` Hermes behaves exactly as it always has:
|
||||
batch-only compaction. Everything else about compression is unchanged.
|
||||
|
||||
`micro_compact_every_n_turns` is the knob that matters most after the on/off
|
||||
switch, because it sets how often you pay the cache break described below. At
|
||||
`1` a pass runs after every completed turn: the most aggressive reclaim, and
|
||||
one broken prefix per turn. At `5` you get a fifth of the breaks and a fifth of
|
||||
the reclaim rate, which is the right direction if your sessions are long-lived
|
||||
and your provider's cache discount is deep. Values below `1` are clamped to `1`
|
||||
rather than silently disabling the feature. The counter advances per turn, not
|
||||
per committed pass, so a turn with nothing to absorb still moves the cadence
|
||||
along and cannot wedge it.
|
||||
|
||||
`micro_compact_defrag_threshold_tokens` is when the rolling summary gets
|
||||
re-summarized instead of growing forever — see [Defrag](#defrag).
|
||||
|
||||
It ships opt-in rather than on because of the prompt-cache cost described in
|
||||
the next section — that cost is real, it is not universally worth paying, and
|
||||
it should be a decision you make rather than one you inherit.
|
||||
|
||||
## Prompt caching — the cost you are opting into
|
||||
|
||||
Read this before enabling the feature. It is the strongest argument against it.
|
||||
|
||||
A long-lived conversation reuses a cached prompt prefix every turn, and cached
|
||||
input tokens are billed at a fraction of uncached ones. That discount survives
|
||||
only as long as the prefix does not change. **A micro-compaction pass rewrites
|
||||
already-sent history**, which invalidates the prefix from the rewrite point
|
||||
onward — so with micro-compaction on, you break the cache *every turn* instead
|
||||
of once per batch compaction.
|
||||
|
||||
This is the same cost the proactive prune deliberately avoids. That path gates
|
||||
itself behind `compression.proactive_prune_min_reclaim_tokens` (4096 by
|
||||
default) precisely so its rewrites stay, in the words of the config comment,
|
||||
"one big episodic break instead of a tiny break every tool iteration."
|
||||
|
||||
Micro-compaction has no equivalent *reclaim-size* gate — a pass commits
|
||||
whatever the one absorbed exchange happened to save, large or small. What it
|
||||
has instead is a *frequency* dial, `micro_compact_every_n_turns`. Raising it
|
||||
makes the breaks rarer and more episodic, which is the same end the prune's
|
||||
gate serves by a different route, though it gets there by absorbing less rather
|
||||
than by waiting for a bigger win. If you want the prune's exact semantics here,
|
||||
a reclaim threshold on micro-compaction is the obvious follow-up and does not
|
||||
exist yet.
|
||||
|
||||
So the honest framing is a trade of one cost for another, not a saving:
|
||||
|
||||
| | Batch-only (default) | Micro-compaction on |
|
||||
|---|---|---|
|
||||
| Compression stalls | One long stall at the threshold | Spread across turns |
|
||||
| Context occupancy | Sawtooths up to the threshold | Stays low and flat |
|
||||
| Cache prefix | Intact between compactions | Broken every turn |
|
||||
|
||||
Which side wins depends on numbers specific to you: how much your provider
|
||||
discounts cached input, how large your prefix is, how long your sessions run,
|
||||
and how much a mid-session stall actually costs you. On a provider with a deep
|
||||
cache discount and a big prefix, the per-turn invalidation can plausibly cost
|
||||
more than the stall it removes. Measure your own sessions — see
|
||||
[Measuring it](#measuring-it) — rather than assuming.
|
||||
|
||||
## Choosing a compression model
|
||||
|
||||
Micro-compaction uses the `auxiliary.compression` model:
|
||||
|
||||
```yaml
|
||||
auxiliary:
|
||||
compression:
|
||||
provider: openai-api
|
||||
model: <your choice>
|
||||
base_url: <endpoint>
|
||||
```
|
||||
|
||||
This is the single most important knob, and there is no universally right
|
||||
answer — it depends on your hardware and what you are willing to trade.
|
||||
|
||||
Each pass sends the running summary plus one exchange, so the prompt is small
|
||||
(a few thousand tokens) but the call happens **every turn**, at the end of the
|
||||
turn. Two properties matter:
|
||||
|
||||
- **Latency dominates.** Because a pass runs per turn, its wall-clock cost is
|
||||
felt repeatedly. A model that takes 30 seconds turns every turn into a turn
|
||||
plus 30 seconds.
|
||||
- **Reasoning models are a poor fit.** Merging one exchange into a summary is
|
||||
mechanical work. A thinking model will spend reasoning tokens on it and be
|
||||
substantially slower than a plain instruct model of similar size, for no
|
||||
benefit to the output.
|
||||
|
||||
Some measured points, on one particular setup — treat them as illustrations of
|
||||
the shape, not as recommendations:
|
||||
|
||||
| model | observed |
|
||||
|---|---|
|
||||
| 7B 4-bit instruct, local (MLX, Apple Silicon) | ~31s per pass; box also serving other work |
|
||||
| large MoE reasoning model, remote GPU | noticeably slower still — thinking tokens on a summarisation task |
|
||||
|
||||
The pattern is that a small, fast, non-reasoning instruct model is usually the
|
||||
right shape, and that a bigger or "smarter" model is often worse here rather
|
||||
than better. Where that lands for you depends on what you have to run it on.
|
||||
|
||||
If passes feel too slow, your options in rough order of effect are: pick a
|
||||
faster or smaller compression model; give it a less contended host; or turn
|
||||
micro-compaction off and go back to batch compaction.
|
||||
|
||||
## Measuring it
|
||||
|
||||
Micro-compaction is not primarily a token-saving or time-saving optimisation,
|
||||
and judging it on tokens saved will undersell it. The two things it actually
|
||||
buys you are:
|
||||
|
||||
1. **The long pause is amortized.** The same summarization work happens, but as
|
||||
small increments after turns instead of one stall in the middle of a session.
|
||||
2. **Your context lasts longer.** Because the middle is continuously reclaimed,
|
||||
occupancy stays low instead of sawtoothing up to the threshold. A session
|
||||
runs much further — often indefinitely — before it needs a hard compaction
|
||||
at all.
|
||||
|
||||
So the number that matters is **occupancy**: how full the window is being kept,
|
||||
as a percentage of the compaction threshold. A session that holds steady around
|
||||
40% has headroom to keep going; one climbing through 90% is about to stall. The
|
||||
second number is **how many batch compactions actually fired** — ideally none.
|
||||
|
||||
A session can save nothing on paper and still be a clear win on both counts.
|
||||
|
||||
Every pass emits one content-free JSON line, in the same style as the batch
|
||||
compaction telemetry:
|
||||
|
||||
```
|
||||
micro compaction telemetry: {"event":"micro_compaction","outcome":"absorbed",
|
||||
"tokens_before":12739,"tokens_after":12060,"tokens_delta":-679,
|
||||
"occupancy_pct":38.4,"threshold_tokens":34816,"context_limit":40960,
|
||||
"exchange_tokens":868,"rolling_summary_tokens":31,"passes_total":1,
|
||||
"tokens_saved_total":679,"duration_ms":14,...}
|
||||
```
|
||||
|
||||
`occupancy_pct` is `tokens_after` as a share of the compaction threshold -- the
|
||||
headroom figure. It is null when the model's window has not been resolved yet:
|
||||
the telemetry reads only the cached value, because resolving it can issue a
|
||||
synchronous `/models` probe and telemetry must never be what blocks a turn.
|
||||
|
||||
`tokens_delta` is negative when the pass shrank the transcript.
|
||||
`tokens_saved_total` and `passes_total` accumulate across the session, so a whole
|
||||
run can be summarised from its last line. No transcript content appears in the
|
||||
payload — only counts.
|
||||
|
||||
To turn a log into an answer:
|
||||
|
||||
```
|
||||
python scripts/micro_compaction_report.py [--per-session] [LOGFILE ...]
|
||||
```
|
||||
|
||||
Defaults to `$HERMES_HOME/logs/agent.log`. It reports passes, outcome mix, net
|
||||
tokens saved, mean absorbed-exchange size and pass durations.
|
||||
|
||||
### What it looks like when it is working
|
||||
|
||||
One real session — a 3.5 hour whole-project code review, ~75K tokens of
|
||||
transcript, 400K window, compaction threshold at 320K:
|
||||
|
||||
| pass | messages | tokens | delta | occupancy | duration |
|
||||
|---|---|---|---|---|---|
|
||||
| 1 | 40 -> 39 | 27,479 -> 27,778 | +299 | 8.7% | 2.2s |
|
||||
| 2 | 61 -> 59 | 48,676 -> 48,128 | -548 | 15.0% | 4.5s |
|
||||
| 3 | 70 -> 67 | 58,309 -> 55,915 | -2,394 | 17.5% | 9.1s |
|
||||
| 4 | 84 -> 80 | 75,251 -> 69,818 | -5,433 | 21.8% | 36.2s |
|
||||
| 5 | 84 -> 80 | 74,659 -> 70,264 | -4,395 | 22.0% | 31.2s |
|
||||
|
||||
Three things to read off it.
|
||||
|
||||
**Occupancy flattened.** It climbed to about 22% and stopped. The last two
|
||||
passes are identical (84 -> 80 messages); between them the conversation added
|
||||
4,841 tokens and micro-compaction reclaimed 4,395. That is equilibrium: the
|
||||
window holds steady instead of marching toward the threshold.
|
||||
|
||||
**No batch compaction fired.** Across the whole session the long pause never
|
||||
happened.
|
||||
|
||||
**Reclamation only ramps after the tail budget.** The first passes recovered
|
||||
almost nothing, because below the tail budget (here 64,000 tokens, 16% of the
|
||||
window) nearly the whole transcript is protected tail and there is very little
|
||||
that may be touched. Early sessions legitimately show no passes at all.
|
||||
|
||||
And the cost, stated plainly: passes ran 2 to 37 seconds, median around 31, on
|
||||
a small local model that was also serving other work. Roughly two minutes of
|
||||
summarisation spread across three and a half hours. Against one batch
|
||||
compaction of a 75K-token middle that is still the better trade, but a
|
||||
37-second increment is not a rounding error. See
|
||||
[Choosing a compression model](#choosing-a-compression-model).
|
||||
|
||||
### Reading the numbers honestly
|
||||
|
||||
**The first pass in a session usually costs tokens rather than saving them.**
|
||||
Inserting the summary marker carries a fixed ~400 tokens of scaffolding — the
|
||||
compaction preamble, the historical heading, the end marker — and on pass one
|
||||
that is paid against a single absorbed exchange. A first pass showing
|
||||
`tokens_delta: +330` is not a malfunction.
|
||||
|
||||
From the second pass on, the marker is *replaced* rather than added, so the
|
||||
scaffolding is already paid for and each absorbed exchange is close to pure
|
||||
saving. The break-even is normally the second or third pass. This is why the
|
||||
per-session view matters more than any single line: judge the feature on a
|
||||
session's trajectory, not on one turn.
|
||||
|
||||
The plainer human-readable lines are still there too:
|
||||
|
||||
```
|
||||
Micro-compaction: 37 -> 36 messages
|
||||
Micro-compaction defrag: rolling summary re-summarized (1843 chars)
|
||||
Micro-compaction: skipping exchange at cursor 12 after 3 consecutive failures
|
||||
```
|
||||
|
||||
Message counts move by small amounts — that's expected. The token count is where
|
||||
the effect shows: absorbing one tool-heavy exchange can drop hundreds of tokens
|
||||
while changing the message count by one or two.
|
||||
|
||||
## Failure behaviour
|
||||
|
||||
Micro-compaction is best-effort throughout. The call in `finalize_turn` is wrapped
|
||||
so that any exception is logged and swallowed — a failure returns the conversation
|
||||
unchanged and the turn completes normally. It can degrade, but it shouldn't be
|
||||
able to break a session.
|
||||
@@ -0,0 +1,261 @@
|
||||
# Hermes Middleware
|
||||
|
||||
Hermes middleware is the behavior-changing companion to observer hooks.
|
||||
Observer hooks report what happened. Middleware can change what happens by
|
||||
rewriting a request before execution or by wrapping the execution callback
|
||||
itself.
|
||||
|
||||
This contract is intentionally backend-neutral. A plugin can use it for local
|
||||
policy, request shaping, tracing, adaptive routing, cache control, sandbox
|
||||
selection, or handoff to runtimes such as NeMo Relay without changing Hermes'
|
||||
planner, model provider adapters, tool registry, memory, or CLI UX.
|
||||
|
||||
With middleware enabled, plugins can:
|
||||
|
||||
- Rewrite LLM provider request kwargs before Hermes calls the provider.
|
||||
- Rewrite tool arguments before guardrails, approval checks, hooks, and tool
|
||||
execution see them.
|
||||
- Wrap the actual LLM execution callback while preserving Hermes retry,
|
||||
streaming, interrupt, and hook behavior.
|
||||
- Wrap the actual tool execution callback while preserving Hermes guardrails,
|
||||
approval, post-tool hooks, and tool-result transformation.
|
||||
|
||||
## Contract
|
||||
|
||||
Plugins register middleware from `register(ctx)`:
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
ctx.register_middleware("llm_request", on_llm_request)
|
||||
ctx.register_middleware("llm_execution", on_llm_execution)
|
||||
ctx.register_middleware("tool_request", on_tool_request)
|
||||
ctx.register_middleware("tool_execution", on_tool_execution)
|
||||
```
|
||||
|
||||
Every middleware callback receives:
|
||||
|
||||
- `telemetry_schema_version`: currently `hermes.observer.v1`
|
||||
- `middleware_schema_version`: currently `hermes.middleware.v1`
|
||||
- Runtime context such as `session_id`, `task_id`, `turn_id`,
|
||||
`api_request_id`, `provider`, `model`, `api_mode`, `tool_name`, and
|
||||
`tool_call_id` when applicable.
|
||||
|
||||
Supported middleware kinds:
|
||||
|
||||
| Kind | Payload | Return shape | Purpose |
|
||||
| --- | --- | --- | --- |
|
||||
| `llm_request` | `request`, `original_request` | `{"request": {...}}` | Replace effective provider kwargs before provider execution. |
|
||||
| `tool_request` | `tool_name`, `args`, `original_args` | `{"args": {...}}` | Replace effective tool args before hooks, guardrails, approvals, and execution. |
|
||||
| `llm_execution` | `request`, `original_request`, `next_call` | Any provider response | Wrap or replace the actual provider call. |
|
||||
| `tool_execution` | `tool_name`, `args`, `original_args`, `next_call` | Any tool result | Wrap or replace the actual tool call. |
|
||||
|
||||
Request middleware can return optional trace fields:
|
||||
|
||||
```python
|
||||
return {
|
||||
"request": updated_request,
|
||||
"source": "my-plugin",
|
||||
"reason": "selected fallback model",
|
||||
}
|
||||
```
|
||||
|
||||
Hermes stores those trace entries in later observer hook payloads as
|
||||
`middleware_trace`.
|
||||
|
||||
Execution middleware receives a `next_call` callback. Call it to continue the
|
||||
chain:
|
||||
|
||||
```python
|
||||
def on_tool_execution(**kwargs):
|
||||
result = kwargs["next_call"](kwargs["args"])
|
||||
return result
|
||||
```
|
||||
|
||||
If multiple plugins register the same execution middleware kind, Hermes runs
|
||||
them as a nested chain in registration order. Middleware failures are fail-open:
|
||||
Hermes logs a warning and continues with the next middleware or the base
|
||||
runtime path.
|
||||
|
||||
## Execution Order
|
||||
|
||||
### LLM Calls
|
||||
|
||||
For each provider request, Hermes applies middleware in this order:
|
||||
|
||||
1. Build provider kwargs from the current conversation.
|
||||
2. Apply `llm_request` middleware.
|
||||
3. Emit `pre_api_request` observer hooks with the effective request.
|
||||
4. Run provider execution through `llm_execution` middleware.
|
||||
5. Emit `post_api_request` or `api_request_error` observer hooks.
|
||||
|
||||
Request middleware sees the full provider kwargs, including `messages` or
|
||||
Responses API `input`, model settings, tool definitions, stream options, and
|
||||
provider-specific options. Execution middleware receives the same effective
|
||||
request plus `next_call`.
|
||||
|
||||
### Tool Calls
|
||||
|
||||
For each tool call, Hermes applies middleware in this order:
|
||||
|
||||
1. Parse and coerce model-provided tool arguments.
|
||||
2. Apply `tool_request` middleware.
|
||||
3. Run the normal Hermes pre-execution path against the effective arguments:
|
||||
tool availability checks, observer block directives, guardrails, and
|
||||
approval checks.
|
||||
4. Run tool execution through `tool_execution` middleware.
|
||||
5. Emit `post_tool_call` observer hooks.
|
||||
6. Apply `transform_tool_result` hooks before the result is appended back into
|
||||
conversation context.
|
||||
|
||||
Tool request middleware runs before approval checks. Use it carefully: a
|
||||
rewritten path, command, or URL is the value downstream policy will evaluate.
|
||||
|
||||
## Enablement
|
||||
|
||||
Middleware only runs for enabled plugins. For a bundled plugin:
|
||||
|
||||
```bash
|
||||
hermes plugins enable <plugin-name>
|
||||
```
|
||||
|
||||
For isolated local testing, use one `HERMES_HOME` for plugin enablement and the
|
||||
agent run:
|
||||
|
||||
```bash
|
||||
export HERMES_HOME=/tmp/hermes-middleware-test
|
||||
mkdir -p "$HERMES_HOME"
|
||||
hermes plugins enable <plugin-name>
|
||||
hermes chat --query 'Reply exactly ok'
|
||||
```
|
||||
|
||||
For source checkouts, prefer the source command so the runtime sees plugins and
|
||||
middleware from the working tree:
|
||||
|
||||
```bash
|
||||
uv sync
|
||||
uv run hermes plugins enable <plugin-name>
|
||||
uv run hermes chat --query 'Reply exactly ok'
|
||||
```
|
||||
|
||||
## Generic Plugin Examples
|
||||
|
||||
The examples below are intentionally small. They show the middleware contract
|
||||
shape without depending on NeMo Relay.
|
||||
|
||||
### LLM Request Middleware
|
||||
|
||||
This plugin tags provider requests and records a middleware trace entry:
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
ctx.register_middleware("llm_request", tag_llm_request)
|
||||
|
||||
|
||||
def tag_llm_request(**kwargs):
|
||||
request = dict(kwargs["request"])
|
||||
extra_body = dict(request.get("extra_body") or {})
|
||||
extra_body.setdefault("metadata", {})["hermes_middleware_demo"] = True
|
||||
request["extra_body"] = extra_body
|
||||
return {
|
||||
"request": request,
|
||||
"source": "middleware-demo",
|
||||
"reason": "tagged provider request",
|
||||
}
|
||||
```
|
||||
|
||||
The effective request is passed to `pre_api_request`, provider execution, and
|
||||
`post_api_request`.
|
||||
|
||||
### Tool Request Middleware
|
||||
|
||||
This plugin constrains `terminal` calls to a known working directory:
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
ctx.register_middleware("tool_request", normalize_terminal_workdir)
|
||||
|
||||
|
||||
def normalize_terminal_workdir(**kwargs):
|
||||
if kwargs.get("tool_name") != "terminal":
|
||||
return None
|
||||
args = dict(kwargs["args"])
|
||||
args.setdefault("workdir", "/tmp/hermes-middleware-demo")
|
||||
return {
|
||||
"args": args,
|
||||
"source": "middleware-demo",
|
||||
"reason": "defaulted terminal workdir",
|
||||
}
|
||||
```
|
||||
|
||||
Because this runs before hooks and approvals, downstream telemetry and policy
|
||||
observe the rewritten `workdir`.
|
||||
|
||||
### LLM Execution Middleware
|
||||
|
||||
This plugin wraps the provider call and preserves the raw provider response:
|
||||
|
||||
```python
|
||||
import time
|
||||
|
||||
|
||||
def register(ctx):
|
||||
ctx.register_middleware("llm_execution", time_llm_execution)
|
||||
|
||||
|
||||
def time_llm_execution(**kwargs):
|
||||
started = time.monotonic()
|
||||
response = kwargs["next_call"](kwargs["request"])
|
||||
elapsed_ms = int((time.monotonic() - started) * 1000)
|
||||
print(f"llm_execution elapsed_ms={elapsed_ms}")
|
||||
return response
|
||||
```
|
||||
|
||||
Return the same response shape Hermes expects from the provider adapter. Do not
|
||||
wrap the response in a plugin-specific envelope unless the rest of the runtime
|
||||
expects that envelope.
|
||||
|
||||
### Tool Execution Middleware
|
||||
|
||||
This plugin wraps tool execution while preserving the tool result:
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
ctx.register_middleware("tool_execution", annotate_tool_execution)
|
||||
|
||||
|
||||
def annotate_tool_execution(**kwargs):
|
||||
result = kwargs["next_call"](kwargs["args"])
|
||||
# Metrics, logging, or external routing can happen here.
|
||||
return result
|
||||
```
|
||||
|
||||
Execution middleware may call `next_call(modified_args)` to pass a changed
|
||||
payload to later middleware and the base tool dispatcher.
|
||||
|
||||
Plugin-specific examples should live with the plugin that owns the behavior.
|
||||
NeMo Relay execution middleware is installed through an explicitly selected
|
||||
Relay `plugins.toml`; see
|
||||
[Relay shared metrics](../observability/relay-shared-metrics.md).
|
||||
|
||||
## Safety Notes
|
||||
|
||||
- Middleware should be deterministic for the same input unless it is explicitly
|
||||
routing to a dynamic external system.
|
||||
- Request middleware should return complete replacement payloads, not partial
|
||||
patches.
|
||||
- Execution middleware should call `next_call(...)` exactly once unless it is
|
||||
intentionally short-circuiting execution.
|
||||
- If execution middleware raises before calling `next_call(...)`, Hermes treats
|
||||
that as middleware failure and continues with the remaining middleware chain
|
||||
and base execution.
|
||||
- If execution middleware calls `next_call(...)` successfully and then raises
|
||||
during post-processing, Hermes preserves the downstream result and does not
|
||||
run the provider or tool a second time.
|
||||
- If downstream provider or tool execution fails, middleware may let that error
|
||||
propagate or translate it deliberately. Hermes does not convert downstream
|
||||
failure into a successful `None` result.
|
||||
- Tool request middleware runs before approvals. If it mutates file paths,
|
||||
commands, URLs, or arguments, the mutated values are what guardrails and
|
||||
approvals evaluate.
|
||||
- Observer hooks remain the right place for read-only telemetry. Use middleware
|
||||
only when a plugin needs to alter or wrap behavior.
|
||||
@@ -0,0 +1,325 @@
|
||||
# Hermes Observer Hooks
|
||||
|
||||
Hermes observer hooks are the read-only telemetry contract for plugins that
|
||||
need to reconstruct agent execution without changing runtime behavior. This
|
||||
contract supports trace, metrics, audit, replay, and export integrations such
|
||||
as Langfuse, OpenTelemetry-style collectors, and NeMo Relay.
|
||||
|
||||
Observer hooks are intentionally backend-neutral. They expose stable lifecycle
|
||||
events, correlation IDs, sanitized payloads, timing, status, and error fields.
|
||||
They do not replace Hermes' planner, model providers, memory, tool registry,
|
||||
approval UX, CLI, gateway behavior, or execution semantics.
|
||||
|
||||
Behavior-changing request or execution wrappers are outside this observer
|
||||
contract. Observer hooks should report what happened; they should not replace
|
||||
provider requests, tool arguments, or execution callbacks.
|
||||
|
||||
Hermes also has a first-party NeMo Relay shared-metrics path. It uses these
|
||||
lifecycle boundaries directly and does not require enabling an observability
|
||||
plugin. See [Relay shared metrics](relay-shared-metrics.md).
|
||||
|
||||
## Contract
|
||||
|
||||
Plugins register observer callbacks from `register(ctx)`:
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
ctx.register_hook("pre_api_request", on_pre_api_request)
|
||||
ctx.register_hook("post_api_request", on_post_api_request)
|
||||
ctx.register_hook("pre_tool_call", on_pre_tool_call)
|
||||
ctx.register_hook("post_tool_call", on_post_tool_call)
|
||||
```
|
||||
|
||||
Every hook callback receives keyword arguments. Plugins should accept
|
||||
`**kwargs` so additive fields remain backward-compatible:
|
||||
|
||||
```python
|
||||
def on_post_tool_call(**kwargs):
|
||||
tool_name = kwargs.get("tool_name")
|
||||
status = kwargs.get("status")
|
||||
result = kwargs.get("result")
|
||||
```
|
||||
|
||||
The plugin manager injects this field into every hook payload:
|
||||
|
||||
```text
|
||||
telemetry_schema_version = "hermes.observer.v1"
|
||||
```
|
||||
|
||||
Hook callbacks are fail-open. Hermes catches callback exceptions, logs a
|
||||
warning, and keeps the agent loop running.
|
||||
|
||||
Most observer hook return values are ignored. The exceptions are older
|
||||
behavior-affecting hooks:
|
||||
|
||||
| Hook | Return behavior |
|
||||
| --- | --- |
|
||||
| `pre_llm_call` | May return a string or `{"context": "..."}` to inject ephemeral context into the current user message. |
|
||||
| `pre_tool_call` | May return `{"action": "block", "message": "..."}` to block a tool before execution, or `{"action": "modify", "args": {...}}` to transform the tool's input arguments. |
|
||||
| `transform_tool_result` | May return a replacement tool result string after `post_tool_call`. |
|
||||
| `transform_llm_output` | May return a replacement final assistant text string. |
|
||||
|
||||
Telemetry plugins should treat these behavior-affecting returns as optional
|
||||
compatibility features, not as observability requirements.
|
||||
|
||||
## Correlation IDs
|
||||
|
||||
Observer payloads use stable IDs so plugins can join events without relying on
|
||||
callback order alone.
|
||||
|
||||
| Field | Meaning |
|
||||
| --- | --- |
|
||||
| `session_id` | Conversation/session identity. |
|
||||
| `task_id` | Task identity, especially useful for subagents and isolated execution. |
|
||||
| `turn_id` | User-turn identity shared by API attempts and tool calls in a turn. |
|
||||
| `api_request_id` | Opaque provider-attempt identity. Do not parse its string format. |
|
||||
| `api_call_count` | Numeric API attempt count within the agent loop. |
|
||||
| `tool_call_id` | Provider-supplied tool call ID when available. |
|
||||
| `parent_session_id` / `child_session_id` | Session link for delegated subagents. |
|
||||
| `parent_subagent_id` / `child_subagent_id` | Subagent link when available. |
|
||||
| `parent_turn_id` | Parent turn that spawned delegated work. |
|
||||
|
||||
Consumers should prefer explicit fields over parsing compound IDs. In
|
||||
particular, `api_request_id` is an opaque correlation value.
|
||||
|
||||
## Event Families
|
||||
|
||||
### Session Lifecycle
|
||||
|
||||
Session hooks describe conversation boundaries and resets:
|
||||
|
||||
| Hook | When it fires |
|
||||
| --- | --- |
|
||||
| `on_session_start` | A brand-new session starts after the system prompt is built. |
|
||||
| `on_session_end` | A `run_conversation` call ends, including interrupted or incomplete turns. |
|
||||
| `on_session_finalize` | CLI or gateway tears down an active session identity. |
|
||||
| `on_session_reset` | CLI or gateway moves from an old session identity to a new one. |
|
||||
|
||||
Common fields include `session_id`, `completed`, `interrupted`, `reason`,
|
||||
`old_session_id`, and `new_session_id` where available.
|
||||
|
||||
`on_session_end` is turn/run scoped. It is not necessarily the final lifetime
|
||||
boundary for a chat identity. Use `on_session_finalize` and `on_session_reset`
|
||||
for lifecycle cleanup that must happen once per session identity.
|
||||
|
||||
### Turn-Scoped LLM Hooks
|
||||
|
||||
These hooks frame the user turn, not individual provider API attempts:
|
||||
|
||||
| Hook | When it fires |
|
||||
| --- | --- |
|
||||
| `pre_llm_call` | Before the tool loop begins for a user turn. |
|
||||
| `post_llm_call` | After the turn completes with final assistant output. |
|
||||
|
||||
Common `pre_llm_call` fields include `session_id`, `turn_id`,
|
||||
`user_message`, `conversation_history`, `is_first_turn`, `model`, `platform`,
|
||||
and `sender_id`.
|
||||
|
||||
Common `post_llm_call` fields include `session_id`, `turn_id`,
|
||||
`user_message`, `assistant_response`, `conversation_history`, `model`, and
|
||||
`platform`.
|
||||
|
||||
Use request-scoped API hooks for LLM span telemetry. Use `pre_llm_call` and
|
||||
`post_llm_call` for turn-level context, compatibility, and final turn summary.
|
||||
|
||||
### Request-Scoped API Hooks
|
||||
|
||||
API hooks describe provider attempts inside the agent loop:
|
||||
|
||||
| Hook | When it fires |
|
||||
| --- | --- |
|
||||
| `pre_api_request` | Immediately before a provider API request. |
|
||||
| `post_api_request` | After a successful provider response. |
|
||||
| `api_request_error` | After a failed provider request or retryable error path. |
|
||||
|
||||
`pre_api_request` includes:
|
||||
|
||||
- identity: `session_id`, `task_id`, `turn_id`, `api_request_id`
|
||||
- runtime: `platform`, `model`, `provider`, `base_url`, `api_mode`
|
||||
- attempt metadata: `api_call_count`, `message_count`, `tool_count`,
|
||||
`approx_input_tokens`, `request_char_count`, `max_tokens`
|
||||
- timing: `started_at`
|
||||
- sanitized request payload: `request`
|
||||
|
||||
`post_api_request` includes the same identity/runtime fields plus:
|
||||
|
||||
- `api_duration`, `started_at`, `ended_at`
|
||||
- `finish_reason`, `message_count`, `response_model`
|
||||
- `usage`
|
||||
- `assistant_content_chars`, `assistant_tool_call_count`
|
||||
- sanitized response payload: `response`
|
||||
- compatibility object: `assistant_message`
|
||||
|
||||
`api_request_error` includes the same identity/runtime fields plus:
|
||||
|
||||
- `api_duration`, `started_at`, `ended_at`
|
||||
- `status_code`, `retry_count`, `max_retries`, `retryable`, `reason`
|
||||
- structured `error = {"type": ..., "message": ...}`
|
||||
- sanitized failed request payload: `request`
|
||||
|
||||
The sanitized `request`, `response`, and `error` fields are the canonical
|
||||
observer inputs for new consumers.
|
||||
|
||||
### Tool Lifecycle
|
||||
|
||||
Tool hooks describe individual tool calls:
|
||||
|
||||
| Hook | When it fires |
|
||||
| --- | --- |
|
||||
| `pre_tool_call` | Before guardrail-approved tool dispatch. |
|
||||
| `post_tool_call` | After tool dispatch, cancellation, block, or error completion. |
|
||||
| `transform_tool_result` | After `post_tool_call`, before the result is appended to model context. |
|
||||
|
||||
`pre_tool_call` includes `tool_name`, `args`, `task_id`, `session_id`,
|
||||
`tool_call_id`, `turn_id`, and `api_request_id`.
|
||||
|
||||
`post_tool_call` includes the same identity fields plus `result`,
|
||||
`duration_ms`, `status`, `error_type`, and `error_message`.
|
||||
|
||||
`status` is the observer-grade lifecycle outcome. Common values include:
|
||||
|
||||
| Status | Meaning |
|
||||
| --- | --- |
|
||||
| `ok` | Tool completed normally. |
|
||||
| `error` | Tool ran and returned or raised an error outcome. |
|
||||
| `blocked` | A `pre_tool_call` hook blocked execution. |
|
||||
| `cancelled` | Execution was cancelled before normal completion. |
|
||||
|
||||
`post_tool_call` is emitted for blocked and cancelled paths so telemetry
|
||||
plugins can close spans cleanly.
|
||||
|
||||
### Approval Lifecycle
|
||||
|
||||
Approval hooks describe dangerous-command approval prompts:
|
||||
|
||||
| Hook | When it fires |
|
||||
| --- | --- |
|
||||
| `pre_approval_request` | Before the approval request is shown or sent. |
|
||||
| `post_approval_response` | After the user responds or the request times out. |
|
||||
|
||||
Common fields include `command`, `description`, `pattern_key`,
|
||||
`pattern_keys`, `session_key`, and `surface`.
|
||||
|
||||
`post_approval_response` also includes `choice`, with values such as `once`,
|
||||
`session`, `always`, `deny`, and `timeout`.
|
||||
|
||||
Approval hooks are observer-only. Plugins cannot pre-answer or veto approvals
|
||||
from these hooks. To prevent a tool from reaching approval, use
|
||||
`pre_tool_call` blocking.
|
||||
|
||||
### Subagent Lifecycle
|
||||
|
||||
Subagent hooks describe delegated child-agent work:
|
||||
|
||||
| Hook | When it fires |
|
||||
| --- | --- |
|
||||
| `subagent_start` | A delegated child agent is created. |
|
||||
| `subagent_stop` | A delegated child agent returns or fails. |
|
||||
|
||||
`subagent_start` fields include `parent_session_id`, `parent_turn_id`,
|
||||
`parent_subagent_id`, `child_session_id`, `child_subagent_id`, `child_role`,
|
||||
and `child_goal`.
|
||||
|
||||
`subagent_stop` fields include parent/child session IDs, role/status fields,
|
||||
`child_summary`, `duration_ms`, and a metadata-only `tool_call_history`. Each
|
||||
history entry contains the tool name, argument names, bounded side-effect
|
||||
targets, input/output byte counts, and outcome. URL query strings and fragments
|
||||
are removed; raw arguments, prompts, commands, contents, headers, and results
|
||||
are intentionally excluded.
|
||||
|
||||
Observers can use these hooks to model nested trajectories while keeping child
|
||||
agent execution linked to the parent turn that spawned it.
|
||||
|
||||
## Payload Safety
|
||||
|
||||
Observer payloads are designed for telemetry consumers, not raw object access.
|
||||
New consumers should use the sanitized API payloads:
|
||||
|
||||
- `pre_api_request.request`
|
||||
- `post_api_request.response`
|
||||
- `api_request_error.request`
|
||||
- `api_request_error.error`
|
||||
|
||||
Sanitization converts provider objects to JSON-compatible structures, bounds
|
||||
large payloads, redacts sensitive keys, and avoids exposing raw response
|
||||
objects in sanitized fields.
|
||||
|
||||
Legacy compatibility fields such as `request_messages`, `conversation_history`,
|
||||
and `assistant_message` may still be present for existing plugins. New
|
||||
observability consumers should prefer the sanitized payloads.
|
||||
|
||||
## Performance
|
||||
|
||||
The default uninstrumented path should stay cheap. Expensive request/response
|
||||
payload construction is gated behind `has_hook(...)`, so Hermes only builds
|
||||
sanitized API telemetry payloads when at least one plugin registered the
|
||||
relevant hook.
|
||||
|
||||
Plugin authors should preserve this property:
|
||||
|
||||
- Register only hooks the plugin actually consumes.
|
||||
- Avoid deep-copying or re-sanitizing already sanitized payloads.
|
||||
- Keep hook callbacks fast and fail-open.
|
||||
- Offload network export or batch writes when practical.
|
||||
|
||||
## Writing An Observer Plugin
|
||||
|
||||
Minimal observer plugin:
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
ctx.register_hook("pre_api_request", on_pre_api_request)
|
||||
ctx.register_hook("post_api_request", on_post_api_request)
|
||||
ctx.register_hook("pre_tool_call", on_pre_tool_call)
|
||||
ctx.register_hook("post_tool_call", on_post_tool_call)
|
||||
|
||||
|
||||
def on_pre_api_request(**kwargs):
|
||||
start_llm_span(
|
||||
request_id=kwargs.get("api_request_id"),
|
||||
turn_id=kwargs.get("turn_id"),
|
||||
request=kwargs.get("request"),
|
||||
model=kwargs.get("model"),
|
||||
)
|
||||
|
||||
|
||||
def on_post_api_request(**kwargs):
|
||||
finish_llm_span(
|
||||
request_id=kwargs.get("api_request_id"),
|
||||
response=kwargs.get("response"),
|
||||
usage=kwargs.get("usage"),
|
||||
duration=kwargs.get("api_duration"),
|
||||
)
|
||||
|
||||
|
||||
def on_pre_tool_call(**kwargs):
|
||||
start_tool_span(
|
||||
call_id=kwargs.get("tool_call_id"),
|
||||
name=kwargs.get("tool_name"),
|
||||
args=kwargs.get("args"),
|
||||
)
|
||||
|
||||
|
||||
def on_post_tool_call(**kwargs):
|
||||
finish_tool_span(
|
||||
call_id=kwargs.get("tool_call_id"),
|
||||
result=kwargs.get("result"),
|
||||
status=kwargs.get("status"),
|
||||
duration_ms=kwargs.get("duration_ms"),
|
||||
)
|
||||
```
|
||||
|
||||
Use `session_id`, `turn_id`, `api_request_id`, and `tool_call_id` for span
|
||||
correlation. Use subagent and approval hooks when the export format supports
|
||||
nested agent work or security lifecycle events.
|
||||
|
||||
## Existing Consumers
|
||||
|
||||
The bundled Langfuse plugin demonstrates direct hook-based observability for
|
||||
turns, provider requests, and tool calls.
|
||||
|
||||
The native NeMo Relay SDK integration maps Hermes session, turn, LLM, and tool
|
||||
lifecycles to Relay. Explicit Relay plugin configuration can add
|
||||
[ATOF, ATIF, or OTEL](https://docs.nvidia.com/nemo/relay/configure-plugins/observability/about)
|
||||
exporters and execution middleware; see
|
||||
[Relay shared metrics](relay-shared-metrics.md).
|
||||
@@ -0,0 +1,304 @@
|
||||
# Gateway Monitoring
|
||||
|
||||
Service health monitoring plus structured operational diagnostics for the
|
||||
Hermes gateway daemon, exported over OTLP/HTTP to an operator-configured
|
||||
endpoint (OpenTelemetry Collector, DataDog, or any OTLP receiver).
|
||||
|
||||
This plane is content-free by construction. It exports gateway and cron
|
||||
lifecycle state, platform connector health, and content-free warning/error
|
||||
diagnostics. It never exports prompts, messages, tool arguments or results,
|
||||
job names, destinations, schedules, raw errors, session history, usage
|
||||
analytics, audit logs, or detailed execution traces. Run/model/tool trajectory
|
||||
capture is a separate plane served by Hermes's native NeMo Relay SDK
|
||||
integration and explicitly configured Relay subscribers or exporters.
|
||||
|
||||
## What gets exported
|
||||
|
||||
| Signal | OTLP route | Content |
|
||||
| --- | --- | --- |
|
||||
| Gateway gauges | `/v1/metrics` | `hermes.gateway.up/state/busy/drainable/active_agents/background_work/background_delegations/restart_requested`, `hermes.platform.up/degraded` with bounded `error_code` attributes |
|
||||
| Health/lifecycle events | `/v1/traces` | `gateway.lifecycle` state transitions (`starting -> running -> draining -> stopped`, `startup_failed`, exit), `gateway.health_snapshot`, platform state changes |
|
||||
| Diagnostics | `/v1/logs` | Warning/error gateway events with a constant body and bounded subsystem, severity, error class, and error code attributes; rendered log messages are never exported |
|
||||
| Cron scheduler gauges | `/v1/metrics` | Ticker heartbeat and last-success age (omitted when unavailable), a monotonic catch-up-occurrence count from the scheduler's stale-window branch, enabled/running job counts, and overdue count derived from persisted `next_run_at` plus the scheduler's existing grace rule |
|
||||
| Cron execution lifecycle | `/v1/traces` | Durable `claimed/running/completed/failed/unknown` states, bounded source and error class, opaque hashed job key, elapsed duration when timestamps exist, and delivery outcome when the scheduler knows it; terminal states make a fail-open flush attempt that can delay completion by up to one second |
|
||||
|
||||
Signals carry `service.name`, version, supervision mode, and a stable one-way
|
||||
hash of the install id so an operator can distinguish instances without
|
||||
exporting account/profile identity or the raw install identifier.
|
||||
|
||||
`hermes.gateway.active_agents`, `hermes.gateway.background_work`, and
|
||||
`hermes.gateway.background_delegations` are complementary. `active_agents`
|
||||
counts foreground message turns plus in-flight cron jobs plus API runs — the
|
||||
work the gateway drains on shutdown. `background_work` counts detached work that
|
||||
`active_agents` never includes: backgrounded `delegate_task` subagents,
|
||||
`terminal(background=true)` processes, and kanban workers; it is
|
||||
**task-granular** — a fan-out batch of N subagents counts as N — so it reflects
|
||||
real concurrent subagent load. `background_delegations` counts only async
|
||||
delegation **units** (each `delegate_task` dispatch is one, a fan-out batch is
|
||||
one), matching the async pool's capacity accounting; alert it against
|
||||
`delegation.max_concurrent_children` to see slot pressure. Sum `active_agents`
|
||||
and `background_work` for total live work per instance; use
|
||||
`background_delegations` for pool-saturation.
|
||||
|
||||
## Enabling
|
||||
|
||||
```yaml
|
||||
# config.yaml
|
||||
monitoring:
|
||||
gateway_health_export:
|
||||
enabled: true
|
||||
export:
|
||||
otlp:
|
||||
enabled: true
|
||||
endpoint: http://collector-host:4318/v1/traces # metrics/logs derive
|
||||
headers_env: {} # header name -> ENV VAR NAME (values never stored)
|
||||
```
|
||||
|
||||
Check the posture any time:
|
||||
|
||||
```bash
|
||||
hermes monitoring status
|
||||
```
|
||||
|
||||
The OpenTelemetry SDK is an optional extra (`pip install 'hermes-agent[otlp]'`),
|
||||
lazy-installed on first use. When the SDK is missing or the endpoint is down,
|
||||
the gateway runs unaffected: metric collection and ordinary event export stay
|
||||
off the hot path, while terminal cron events make one bounded fail-open flush
|
||||
attempt of up to one second so the final state is less likely to be lost.
|
||||
|
||||
Works identically under systemd/launchd/s6 supervision, containers, tmux, or
|
||||
a plain `hermes gateway run`: the exporter lives in the gateway process, so
|
||||
no sidecar, agent, or collector is required on the host.
|
||||
|
||||
## Collecting into DataDog
|
||||
|
||||
Run a customer-owned OpenTelemetry Collector and forward:
|
||||
|
||||
```yaml
|
||||
# otel-collector config
|
||||
receivers:
|
||||
otlp:
|
||||
protocols:
|
||||
http:
|
||||
exporters:
|
||||
datadog:
|
||||
api:
|
||||
key: ${env:DD_API_KEY}
|
||||
service:
|
||||
pipelines:
|
||||
metrics: {receivers: [otlp], exporters: [datadog]}
|
||||
traces: {receivers: [otlp], exporters: [datadog]}
|
||||
logs: {receivers: [otlp], exporters: [datadog]}
|
||||
```
|
||||
|
||||
Point `monitoring.export.otlp.endpoint` at the collector. Alerts belong on
|
||||
`hermes.gateway.up`, `hermes.platform.up`, and `hermes.platform.degraded`.
|
||||
|
||||
## Generic fleet queries and alerts
|
||||
|
||||
The exact syntax depends on the customer's observability backend. The examples
|
||||
below use PromQL-style expressions and intentionally avoid vendor-specific
|
||||
routing, destinations, or customer inventory.
|
||||
|
||||
Group fleet views by the opaque `service.instance.id` resource attribute. A
|
||||
process that has died cannot emit its own zero, so every deployment needs both
|
||||
explicit-state and missing-series detection.
|
||||
|
||||
```promql
|
||||
# Explicit gateway failure.
|
||||
hermes_gateway_up == 0
|
||||
|
||||
# Box disappeared or stopped exporting. Choose a window longer than the
|
||||
# configured export interval and collector retry allowance.
|
||||
absent_over_time(hermes_gateway_up[5m])
|
||||
|
||||
# Locally owned bridge is explicitly down.
|
||||
hermes_platform_up == 0
|
||||
|
||||
# Scheduler thread is stale even though the gateway may still be alive.
|
||||
hermes_cron_scheduler_heartbeat_age_seconds > 180
|
||||
|
||||
# Ticker loops but has not completed a successful tick recently.
|
||||
hermes_cron_scheduler_last_success_age_seconds > 300
|
||||
|
||||
# One or more jobs are beyond their existing scheduler grace window.
|
||||
hermes_cron_jobs_overdue > 0
|
||||
|
||||
# Catch-up counter increased, proving at least one stale occurrence was
|
||||
# collapsed and run once after a delay.
|
||||
increase(hermes_cron_scheduler_catch_up_occurrences[15m]) > 0
|
||||
```
|
||||
|
||||
Cron execution lifecycle records arrive as `hermes.cron_execution` spans.
|
||||
Alert or derive events from bounded attributes such as:
|
||||
|
||||
```text
|
||||
hermes.status = failed|unknown
|
||||
hermes.delivery_outcome = failed|not_configured
|
||||
hermes.error_class = auth_failed|rate_limited|timeout|network_error|
|
||||
dispatch_failed|interrupted|empty_response|
|
||||
invalid_config|unknown
|
||||
```
|
||||
|
||||
Recommended operator views:
|
||||
|
||||
1. one row per `service.instance.id` with gateway and configured local-platform
|
||||
state;
|
||||
2. scheduler heartbeat, last-success age, running count, overdue count, and
|
||||
catch-up increase;
|
||||
3. a cron lifecycle feed keyed only by opaque `hermes.job_key`;
|
||||
4. separate alerts for box absence, local bridge down, scheduler stale, cron
|
||||
failed/unknown, delivery failure, and overdue/catch-up activity.
|
||||
|
||||
Keep alert thresholds and routing in deployment-owned configuration. Do not add
|
||||
job names, prompts, outputs, schedules, destinations, raw errors, profile names,
|
||||
or account identity merely to make a dashboard easier to read.
|
||||
|
||||
## Release-validation scenarios
|
||||
|
||||
Before accepting a deployment, force and verify all five cases through the real
|
||||
collector and backend:
|
||||
|
||||
1. **Cron success:** observe `claimed -> running -> completed`, duration, and a
|
||||
truthful delivery outcome.
|
||||
2. **Cron failure:** observe `failed` plus a bounded error class, with no raw
|
||||
exception or content in the decoded OTLP payload.
|
||||
3. **Cron interruption:** stop the owning gateway during execution, restart it,
|
||||
and observe recovery to `unknown`.
|
||||
4. **Locally owned bridge outage:** break one native connector, observe its
|
||||
bounded down/retrying/fatal state and recovery, and verify unaffected boxes
|
||||
remain healthy.
|
||||
5. **Killed gateway:** terminate one canary, verify missing-series detection,
|
||||
restart it, and confirm the same opaque instance identity returns.
|
||||
|
||||
Hermes Agent-owned Relay transport health remains in scope. A separate gateway
|
||||
or connector service remains authoritative for any shared connected-platform
|
||||
state that it owns and should export that state through its own telemetry path.
|
||||
|
||||
For every scenario, verify the signal and alert clear on recovery, other boxes
|
||||
remain unaffected, collector failure stays fail-open, and decoded metrics,
|
||||
spans, logs, and resource attributes remain content-free.
|
||||
|
||||
## Local smoke test (no Docker)
|
||||
|
||||
```bash
|
||||
# terminal 1: capture collector on :4318
|
||||
python scripts/observability/otel_capture_collector.py \
|
||||
--host 127.0.0.1 --port 4318 --log /tmp/hermes_otel_capture.jsonl
|
||||
|
||||
# terminal 2: drive the real exporter through lifecycle transitions,
|
||||
# a fatal platform, and a structured warning event, then flush
|
||||
python scripts/observability/gateway_health_export_probe.py \
|
||||
--endpoint http://127.0.0.1:4318/v1/traces \
|
||||
--log /tmp/hermes_otel_capture.jsonl --wait 8
|
||||
# exit 0 prints: {"requests": 6, "paths": ["/v1/logs", "/v1/metrics", "/v1/traces"]}
|
||||
```
|
||||
|
||||
## Maintaining and extending this plane
|
||||
|
||||
This plane is a **fixed, enumerated, content-free vocabulary** by design. Adding
|
||||
a signal is not just "emit a new metric" — every new name and attribute must be
|
||||
declared in each layer that enforces the bounded vocabulary, or it is silently
|
||||
dropped downstream. Follow the checklist for the change you are making. The
|
||||
golden rule: **a new signal that is emitted but not declared in every layer
|
||||
looks like a code bug but is a vocabulary-registration bug — nothing errors, the
|
||||
signal just never arrives.**
|
||||
|
||||
### Content-free invariant (applies to every change)
|
||||
|
||||
Before adding anything, confirm it cannot carry content. Numbers, booleans,
|
||||
ages, durations, monotonic counts, and one-way hashes are safe. **Never** add an
|
||||
attribute that can hold a job name, prompt, output, schedule, destination, raw
|
||||
exception text, file path, profile name, account id, or free-form string. When
|
||||
you must key a record to a job/entity, hash it (`sha256(...)[:24]`, see
|
||||
`_job_key` in `agent/monitoring/cron_health.py`) — never emit the raw id. All
|
||||
string attributes that could touch user input must pass through
|
||||
`redaction.redact_for_export` and be truncated (see `_span_attrs` in
|
||||
`agent/monitoring/otlp_exporter.py`).
|
||||
|
||||
### Adding a new gauge/metric
|
||||
|
||||
1. Emit it in the snapshot builder (`agent/monitoring/gateway_health.py`
|
||||
`build_gateway_health_snapshot`, `cron_health.py` `build_cron_health_snapshot`,
|
||||
or a sibling reader wired into `_read_runtime_snapshot` in
|
||||
`gateway_health_export.py`). Best-effort: never let a reader raise into the
|
||||
collection loop — wrap it and log a **content-free WARNING with the exception
|
||||
TYPE name only** (the pattern the cron and background-work readers use), so a
|
||||
future regression is visible instead of silently dropping the signal.
|
||||
2. Register the dotted metric name in the observable-gauge `metric_names` list in
|
||||
`gateway_health_export.py::_start_metric_provider`. **A gauge that is emitted
|
||||
in the snapshot but not registered here is never observed.**
|
||||
3. Add the export-table row and an alert example in this file.
|
||||
4. If the deployment fronts the exporter with an OpenTelemetry Collector that
|
||||
uses a metric-name allowlist (a `filter/...` processor with `name != "..."`
|
||||
guards), add the new name there too — otherwise the collector drops it before
|
||||
the backend. This is not repo code, but it is the single most common reason a
|
||||
correctly-emitted new metric never appears; call it out in the PR so the
|
||||
deploying operator updates their collector config.
|
||||
|
||||
### Adding a new subsystem (a new family of signals)
|
||||
|
||||
Mirror the cron pattern (`cron_health.py` + its wiring): put the read/projection
|
||||
logic in its own module, expose one `build_<subsystem>_health_snapshot()` that
|
||||
returns bounded `GatewayMetric`s (and events if any), and extend it into
|
||||
`_read_runtime_snapshot` with the same best-effort try/except-WARNING guard.
|
||||
Then do the "adding a metric" checklist for each new name, and the "adding an
|
||||
attribute" checklist for each new event attribute. Add a release-validation
|
||||
scenario below for the subsystem's failure mode.
|
||||
|
||||
### Extending the error-class / status / source / state vocabularies
|
||||
|
||||
These are the closed enums that keep the plane bounded. Extend the SET, then the
|
||||
classifier, never one without the other:
|
||||
|
||||
- **Cron** (`agent/monitoring/cron_health.py`): `_KNOWN_STATUSES`,
|
||||
`_KNOWN_SOURCES`, `_KNOWN_DELIVERY_OUTCOMES`, and the `classify_cron_error`
|
||||
keyword buckets. Anything not in the set is coerced to `unknown` on the way
|
||||
out, so a new value that is not added to the set is invisible.
|
||||
- **Gateway/platform** (`agent/monitoring/gateway_health.py`):
|
||||
`_KNOWN_GATEWAY_STATES`, `_KNOWN_PLATFORM_STATES`, and `classify_gateway_error`.
|
||||
|
||||
Rules: keep the vocabulary SMALL and operationally meaningful (an error class
|
||||
should map to an operator action, not to an exception subclass); a new bucket
|
||||
must match on a stable keyword, not on message text that could vary; update the
|
||||
`hermes.error_class = ...` list in this file's alert section and the enum's unit
|
||||
test so the contract is asserted, not frozen as a count.
|
||||
|
||||
### Adding a content-free attribute to an existing event/span
|
||||
|
||||
Add the key to the emitter's per-kind `keep_by_kind` allowlist in
|
||||
`agent/monitoring/otlp_exporter.py::_span_attrs` (unlisted keys are dropped), run
|
||||
it through redaction if it is ever string-shaped, and — as with metrics — if the
|
||||
deployment's collector has a span-attribute `keep_keys(...)` allowlist, add the
|
||||
attribute there too or it is stripped in transit.
|
||||
|
||||
### Verify the whole chain, not just emission
|
||||
|
||||
Emitting is necessary but not sufficient. Confirm the signal survives all the
|
||||
way to the backend, because the enums, the `metric_names` registration, the
|
||||
emitter attribute allowlist, and any collector allowlist each drop unlisted
|
||||
values with no error:
|
||||
|
||||
```bash
|
||||
hermes monitoring status # posture
|
||||
python scripts/observability/gateway_health_export_probe.py \
|
||||
--endpoint http://127.0.0.1:4318/v1/traces \
|
||||
--log /tmp/cap.jsonl --wait 8 # drive the real exporter
|
||||
```
|
||||
|
||||
Decode the captured OTLP payload and assert the new name/attribute is present
|
||||
AND that no content leaked. When a real collector sits in front, add its
|
||||
allowlist entries and re-verify against the backend, not just the local capture.
|
||||
|
||||
## Boundaries and roadmap
|
||||
|
||||
The `hermes monitoring` CLI intentionally exposes `status` only. This first
|
||||
release covers only Hermes Agent-owned service-health and operational-diagnostic
|
||||
signals, including Hermes Agent-owned Relay transport health. Team Gateway's
|
||||
authoritative shared connector/platform state is explicitly out of scope, as
|
||||
are product analytics, audit/quality reporting, and detailed execution traces.
|
||||
Shared client usage metrics and enterprise trace telemetry are being designed on
|
||||
the NeMo Relay integration with their own consent, policy, and export
|
||||
boundaries; this monitoring plane stays narrow so an operator can enable it
|
||||
without touching any content-bearing signal. The telemetry surface may be
|
||||
reorganized as that lands.
|
||||
@@ -0,0 +1,482 @@
|
||||
# NeMo Relay Shared Metrics
|
||||
|
||||
Hermes includes NeMo Relay as a normal runtime dependency on platforms for
|
||||
which Relay publishes a native wheel. The shared-metrics integration is built
|
||||
into Hermes and does not require a Hermes observability plugin. Hermes remains
|
||||
importable without Relay on other native targets. Those targets use an
|
||||
explicit reduced-capability no-op host:
|
||||
Hermes execution remains available, while Relay scopes, middleware, plugins,
|
||||
and subscribers are unavailable. The `hermes-agent[nemo-relay]` extra remains
|
||||
as a no-op compatibility alias for existing installation commands.
|
||||
|
||||
> [!WARNING]
|
||||
> This removes the Hermes `observability/nemo_relay` plugin. Existing users
|
||||
> must remove `observability/nemo_relay` (or its legacy `nemo_relay` alias)
|
||||
> from `plugins.enabled` and move exporter configuration into a Relay
|
||||
> `plugins.toml` selected with `HERMES_NEMO_RELAY_PLUGINS_TOML`. The legacy
|
||||
> `HERMES_NEMO_RELAY_ATOF_*` and `HERMES_NEMO_RELAY_ATIF_*` variables no
|
||||
> longer activate exporters. Without the new variable, Hermes does not run
|
||||
> Relay plugin discovery, configuration layering, middleware, or exporters.
|
||||
|
||||
Hermes requires NeMo Relay 0.8.3 or later within the 0.8 release line. That
|
||||
line provides the provider-codec and canonical tool-result contracts Hermes
|
||||
uses for managed provider and tool calls.
|
||||
|
||||
## Runtime Dependency and Data Boundary
|
||||
|
||||
Hermes installs the platform-specific `nemo-relay` native wheel from the
|
||||
bounded `>=0.8.3,<0.9` dependency range. The published package is built from
|
||||
the [NVIDIA NeMo Relay repository](https://github.com/NVIDIA/NeMo-Relay).
|
||||
Unsupported platforms use the explicit no-op runtime described above rather
|
||||
than downloading a different implementation.
|
||||
|
||||
Operator-supplied typed native plugins must be rebuilt for Relay 0.8. `grpc-v1`
|
||||
workers must be regenerated and rebuilt when they use tool callbacks, tool
|
||||
execution intercepts, or manual tool-end APIs.
|
||||
|
||||
When Relay managed execution is active, the provider request and response pass
|
||||
through that native module in the Hermes process so configured interceptors can
|
||||
operate on the real call. This is separate from the shared-metrics data
|
||||
contract. Shared-metrics mode installs no rich-observability network exporter,
|
||||
and its subscriber
|
||||
accepts only the versioned, allowlisted projection described below. The
|
||||
opt-in package sender described in Appendix A is the only outbound path, it
|
||||
transmits nothing unless the user enables both `enabled` and `send`, and it
|
||||
sends whole packages rather than live spans. Enabling a
|
||||
separately configured rich-observability or dynamic plugin can create a
|
||||
different data path and requires its own policy review.
|
||||
|
||||
Collection remains off unless Hermes policy enables it:
|
||||
|
||||
```yaml
|
||||
telemetry:
|
||||
shared_metrics:
|
||||
enabled: true
|
||||
```
|
||||
|
||||
This choice is read from the profile's own `config.yaml`. A machine-managed
|
||||
configuration overlay cannot enable or disable shared metrics on the profile's
|
||||
behalf.
|
||||
|
||||
Relay plugin activation is owned by the native runtime and remains explicitly
|
||||
opt-in. Set `HERMES_NEMO_RELAY_PLUGINS_TOML` to a selected `plugins.toml` to
|
||||
activate configured middleware, exporters, or dynamic plugins. When the
|
||||
variable is unset, Hermes does not invoke Relay's plugin initializer, so Relay
|
||||
does not perform plugin configuration discovery or layering. When it is set
|
||||
and the selected file loads successfully, Relay discovers supported user and
|
||||
system `plugins.toml` files and layers the selected static configuration over
|
||||
them. Repository-local `.nemo-relay/plugins.toml` files are ignored. Dynamic
|
||||
`[[plugins.dynamic]]` records are loaded from the selected file only. If the
|
||||
selected file cannot be loaded, Hermes reports the error and does not invoke
|
||||
Relay initialization or fall back to ambient discovery.
|
||||
|
||||
## Session-Span Segmentation for Continuous Sessions
|
||||
|
||||
Relay exports a span when its scope closes. A continuous gateway session can
|
||||
remain open for days, so its session span remains open even though each turn
|
||||
span is exported normally. Optional segmentation rotates only the session
|
||||
scope at a turn boundary:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
telemetry:
|
||||
session_segments:
|
||||
on_compaction: false # rotate after context compaction
|
||||
max_turns: 0 # 0 = unlimited; N = turns per segment
|
||||
```
|
||||
|
||||
| Key | Default | Behavior |
|
||||
|---|---:|---|
|
||||
| `on_compaction` | `false` | Rotate after compaction completes, at the next turn boundary. |
|
||||
| `max_turns` | `0` | Rotate after every N completed turns; `0` disables the cap. |
|
||||
|
||||
Both defaults preserve one session scope for the full session. Rotated spans
|
||||
retain the same `session_id` and add `hermes.session.segment` plus
|
||||
`hermes.session.segment_reason` (`compaction` or `max_turns`).
|
||||
|
||||
## Process-Wide Plugin Policy and Profile Isolation
|
||||
|
||||
Relay plugin configuration is a process-level deployment choice, not a Hermes
|
||||
profile setting. The first hosted profile triggers lazy initialization, and
|
||||
every additional profile hosted by that Hermes process shares the resulting
|
||||
static middleware, dynamic plugins, subscribers, exporters, and guardrail
|
||||
policy. After initialization succeeds, Hermes logs:
|
||||
|
||||
```text
|
||||
Relay plugins are active process-wide and apply to all profiles hosted by this Hermes process.
|
||||
```
|
||||
|
||||
Profile scopes still preserve causal isolation inside that shared policy.
|
||||
ATIF groups events by their top-level Agent scope, so simultaneous profile
|
||||
sessions produce separate trajectories rather than one mixed trajectory.
|
||||
ATOF and other global subscribers observe events from every hosted profile.
|
||||
Static and dynamic middleware likewise runs for managed calls from every
|
||||
profile.
|
||||
|
||||
A worker plugin running in a separate worker process does not create a
|
||||
per-profile security boundary. One process-wide activation dispatches calls
|
||||
from all hosted profiles to that worker while preserving the invoking
|
||||
profile's Relay scope stack. Native dynamic plugins are loaded into the Hermes
|
||||
process and share the same policy boundary.
|
||||
|
||||
Run profiles in separate Hermes processes when they require different trust
|
||||
levels, plugin credentials, exporter destinations, or guardrail policies.
|
||||
This process-wide plugin contract does not change each profile's independent
|
||||
shared-metrics consent, local SQLite state, or ATIF trajectory grouping.
|
||||
|
||||
Hermes core owns one Relay host and one isolated Relay session scope per Hermes
|
||||
session. Core lifecycle producers use
|
||||
`agent.relay_runtime` to obtain the shared session handle or
|
||||
run Relay scope, LLM, tool, and mark APIs in that session context. New product
|
||||
marks do not require Hermes plugin registration. Shared-metrics marks must
|
||||
still contain only fields approved by the versioned allowlist; the hard
|
||||
dependency does not change the collection or privacy policy.
|
||||
|
||||
## Current Slices
|
||||
|
||||
The current vertical slices record pseudonymous profile activity, logical
|
||||
model calls, top-level task runs, tool and approval outcomes, and skill
|
||||
lifecycle and reuse:
|
||||
|
||||
```text
|
||||
Hermes turn, API, tool, and approval hooks
|
||||
-> Relay session, task, LLM, tool, and mark lifecycle
|
||||
-> Hermes shared-metrics subscriber
|
||||
-> SQLite counters
|
||||
-> immutable JSON delta package
|
||||
```
|
||||
|
||||
Hermes sends an empty `LLMRequest` into the metrics-owned lifecycle. This does
|
||||
not describe the separate managed-execution call through the native runtime
|
||||
documented above. The terminal metrics event contains the model identifier and
|
||||
provider route that Hermes used for the logical call, such as
|
||||
`nvidia/nemotron-3-ultra` through `openrouter`. These identifiers are
|
||||
lowercased and structurally bounded, but they are not normalized through a
|
||||
checked-in model catalog. Pricing and model-family classification belong to
|
||||
the metrics backend. Prompts, responses, endpoints, errors, session IDs, task
|
||||
IDs, and request IDs are not included in the metrics event or package.
|
||||
New calls use `hermes.model_route.count`. The previous
|
||||
`hermes.model_call.count` contract remains readable only so pending local
|
||||
counters created by older builds can be exported without losing data.
|
||||
|
||||
The first consented session start emits an empty `hermes.client.active` Relay
|
||||
mark. The profile-scoped subscriber creates a random UUID install identity and
|
||||
uses a transactional compare-and-set to record at most one client-active
|
||||
counter in any rolling 24-hour window. The metric has no dimensions; Hermes
|
||||
version, OS family, architecture, and install method remain bounded package
|
||||
resources. Concurrent Hermes processes share the SQLite latch, so simultaneous
|
||||
starts cannot double-count one install. A later session or task can attempt the
|
||||
mark again, but the subscriber suppresses it until the rolling window expires.
|
||||
|
||||
Each task run is a Relay `Function` scope named `hermes.task_run`, parented to
|
||||
the owning Hermes session. The start counter contains only bounded execution
|
||||
surface and entrypoint values. The terminal counter contains bounded outcome,
|
||||
end reason, termination status, duration, logical model-call count, terminal
|
||||
tool-call count, and provider-retry count buckets. Retries are additional
|
||||
provider attempts for the same Hermes API request ID; they do not inflate the
|
||||
logical model-call count. Tool calls are deduplicated by their Hermes tool-call
|
||||
ID after a terminal tool result is observed. The outer `AIAgent` execution
|
||||
boundary closes the task for normal returns, early returns, exceptions, and
|
||||
cancellations. Active task ownership follows the task ID if Hermes rotates its
|
||||
conversation session during context compression.
|
||||
|
||||
Each tool invocation is represented by a Relay tool lifecycle named
|
||||
`hermes.tool_call`. The terminal counter contains only bounded tool category,
|
||||
outcome, approval outcome, latency, and explicit retry-count buckets. Hermes
|
||||
derives the category from the toolset already declared in its runtime registry;
|
||||
custom and unrecognized toolsets collapse to `other` rather than exporting
|
||||
tool or plugin names. Hermes does not infer retries from repeated tool names or
|
||||
adjacent calls; when the
|
||||
hook does not provide an explicit retry relationship, the retry bucket is
|
||||
`unknown`. Approval decisions are emitted as `hermes.tool_approval` marks and
|
||||
recorded as attributed to a tool call or explicitly `unattributed`. Tool names,
|
||||
call IDs, arguments, results, commands, descriptions, and error text are not
|
||||
included in shared-metrics events or packages. A started tool that is still
|
||||
open when its task terminates is closed as failed, timed out, or cancelled and
|
||||
remains in the task's tool-count bucket.
|
||||
|
||||
Successful skill mutations emit `hermes.skill.lifecycle` marks with only a
|
||||
bounded action and provenance. Successful loads emit `hermes.skill.load`
|
||||
marks with bounded provenance, first-use or reuse state, reuse-after-patch
|
||||
state, and a use-count bucket. Hermes derives reuse and patch-generation
|
||||
continuity transactionally in its existing `skills/.usage.json` state; skill
|
||||
names and exact counts or generations never enter Relay metrics events,
|
||||
SQLite dimensions, or packages. A use after a new patch is counted once as
|
||||
`reused_after_patch`; later uses remain ordinary reuse until another patch.
|
||||
Task-outcome attribution after a patch remains deferred until its window and
|
||||
multi-skill semantics are defined.
|
||||
|
||||
Local state is written under:
|
||||
|
||||
```text
|
||||
$HERMES_HOME/telemetry/shared_metrics/metrics.sqlite3
|
||||
$HERMES_HOME/telemetry/shared_metrics/outbox/*.json
|
||||
```
|
||||
|
||||
The database keeps transactional aggregate and package-outbox state. Package
|
||||
files are immutable delta documents that conform to a closed JSON schema and
|
||||
are written with atomic replacement. Each package records the Hermes version,
|
||||
OS family, architecture, and install method as bounded client resources.
|
||||
Unrecognized platform or installation values are exported as `unknown`; raw
|
||||
platform strings, hostnames, and paths are never included. Fully packaged
|
||||
aggregate rows and successfully exported package rows and files are retained
|
||||
locally for 30 days. Pending package rows and counters with unexported deltas
|
||||
are never pruned.
|
||||
Package schema v1 remains unchanged for existing outbox files. New packages
|
||||
use v2, which accepts both the retired model-call contract and the current
|
||||
model-route contract so upgrades can drain pending counters safely.
|
||||
|
||||
Each package contains an `install_id` generated as a random UUID. Despite the
|
||||
schema field name, its current scope is one `HERMES_HOME`, so it is more
|
||||
precisely a persistent pseudonymous profile identifier. It is not derived from
|
||||
hardware, account, host, path, or credential data. It remains stable across
|
||||
packages from that profile and can therefore link those local packages.
|
||||
Deleting `$HERMES_HOME/telemetry/shared_metrics` resets the identifier together
|
||||
with all aggregates and package files.
|
||||
|
||||
Remote delivery is opt-in and off by default. Reusing the persistent local
|
||||
identifier remotely required a separate product and privacy decision covering
|
||||
consent, identity scope, reset behavior, retention, and deletion — that
|
||||
decision has been made.
|
||||
|
||||
> Those decisions are recorded in
|
||||
> [Appendix A](#appendix-a-remote-exporter-decisions-phase-2), and the exporter
|
||||
> implementing them has shipped. Collection alone still transmits nothing: the
|
||||
> sender runs only when `telemetry.shared_metrics.send` is also true. Each
|
||||
> transmitted package carries the stable `install_id` as-is (product decision,
|
||||
> 2026-08-27 — see A.2 for the record, including the superseded
|
||||
> HMAC-pseudonym design).
|
||||
|
||||
The install identity is scoped to one `HERMES_HOME`. To reset it, stop Hermes
|
||||
processes and remove `$HERMES_HOME/telemetry/shared_metrics`. This deliberately
|
||||
removes the old identity, aggregate database, and queued local packages
|
||||
together; the next consented session creates a new identity. Disabling shared
|
||||
metrics stops new collection but does not silently delete previously collected
|
||||
local state.
|
||||
|
||||
## Smoke Test
|
||||
|
||||
Run a real Hermes CLI turn against the deterministic local model server:
|
||||
|
||||
```bash
|
||||
./.venv/bin/python scripts/smoke_nemo_relay_shared_metrics.py
|
||||
```
|
||||
|
||||
The script uses the installed `nemo-relay` dependency by default. Pass
|
||||
`--relay-python ../nemo-relay/python` only when testing a locally built Relay
|
||||
binding.
|
||||
|
||||
The smoke has the local model request a real `read_file` tool call before its
|
||||
final response, then drives create, load, reuse, patch, edit, stale, archive,
|
||||
restore, and install skill transitions through the installed Relay binding. It
|
||||
verifies model, provider, task, tool, and skill counters in SQLite, validates
|
||||
all exported delta packages against the closed schema, verifies the
|
||||
pseudonymous client-active counter, and checks that prompt, response, tool-call
|
||||
ID, tool-result, and skill-name canaries are absent from the packages.
|
||||
|
||||
## Appendix A: Remote Exporter Decisions (Phase 2)
|
||||
|
||||
Status: **implemented.** This appendix answers the product and
|
||||
privacy questions that "Current Slices" defers to a future remote exporter. It
|
||||
records what was decided and why, so the reasoning survives the implementation.
|
||||
|
||||
Sending is off by default and requires both `telemetry.shared_metrics.enabled`
|
||||
and `telemetry.shared_metrics.send`.
|
||||
|
||||
The exporter sends the package files already written under
|
||||
`$HERMES_HOME/telemetry/shared_metrics/outbox/` to the Hermes telemetry ingest
|
||||
service. That service validates only the envelope (`schema_version` plus a UUID
|
||||
`package_id`) and stores the body verbatim in S3.
|
||||
|
||||
### A.1 Consent
|
||||
|
||||
Transmission is a **separate opt-in** from collection, under a new config key:
|
||||
|
||||
```yaml
|
||||
telemetry:
|
||||
shared_metrics:
|
||||
enabled: false # collect locally
|
||||
send: false # NEW: transmit to the Nous telemetry service
|
||||
```
|
||||
|
||||
- `send` defaults to **false**. Collection alone never transmits.
|
||||
- `send` requires `enabled`. It does **not** imply it: a transmission flag must
|
||||
not silently switch on collection. `send: true` with `enabled: false` warns
|
||||
and does nothing.
|
||||
- Like `enabled`, `send` is profile-owned and is not overridden by
|
||||
managed-scope configuration.
|
||||
|
||||
**A package is only sent when its whole period falls inside a recorded
|
||||
consent window.** Consent is stored as explicit intervals in the shared-
|
||||
metrics SQLite store (`send_consent_windows`): a window opens when `send:
|
||||
true` is first observed, is confirmed forward by every later observation,
|
||||
and closes — at the last *confirmed* moment, never at the wall clock — when
|
||||
`send: false` is observed. A single reconciler derives this table from the
|
||||
config on every process start, so wizard changes, hand-edits to
|
||||
`config.yaml`, and mid-pass revocations all take the same path, and no
|
||||
transition can be missed by any of them.
|
||||
|
||||
Any package whose period predates the first window, falls between windows,
|
||||
or runs past the newest confirmed moment is excluded — the gate fails
|
||||
closed. A fresh package therefore waits at most one process start after its
|
||||
period completes before becoming eligible.
|
||||
|
||||
The gate is on the **period**, not on the package's creation time. One period
|
||||
is split across several packages created on different days: a day's first
|
||||
package is written that day, and a tail package for the same period typically
|
||||
follows the next day. Gating on creation time would send a period's tail while
|
||||
dropping its head, reporting a **silently undercounted** day. Gating on the
|
||||
period keeps consent forward-only and every transmitted period complete.
|
||||
|
||||
Local history can be up to 30 days old, and that data was collected under a
|
||||
promise that nothing is uploaded. Honouring consent forward-only costs at most
|
||||
30 days of backlog we never had permission to send.
|
||||
|
||||
### A.2 Identity scope — the stable install_id is transmitted as-is
|
||||
|
||||
**Decision record.** The original design of this exporter (and revisions 1–8
|
||||
of this appendix) transmitted a keyed pseudonym instead of the identifier:
|
||||
`HMAC-SHA256(key = locally-held rotating salt, message = install_id)`, with
|
||||
the salt rotating every 30 days. On **2026-08-27**, before the feature
|
||||
shipped (zero consented users, zero production transmissions), the product
|
||||
owner decided the analytical need is a **stable cross-window identity** —
|
||||
retention curves, longitudinal install behaviour — which rotation by design
|
||||
destroys. The pseudonymization layer was removed in full rather than
|
||||
weakened in place.
|
||||
|
||||
What is transmitted now:
|
||||
|
||||
- Each package carries `install_id` verbatim: the persistent, profile-scoped
|
||||
random UUID described above.
|
||||
- It is generated locally (`uuid4`), contains no hardware, account, user, or
|
||||
machine-derived information, and identifies a *profile*, not a person.
|
||||
- It is stable until the user deletes the shared-metrics directory, which
|
||||
regenerates it (see A.4).
|
||||
|
||||
Consequences stated plainly rather than papered over:
|
||||
|
||||
- Packages from one profile correlate **indefinitely**, not per-window.
|
||||
Long-term linkability of one install's daily envelope sequence is now the
|
||||
designed behaviour, not a residue.
|
||||
- The A.3 residue analysis of the old design (stable `resource` tuple +
|
||||
contiguous periods bridging rotation windows) is moot — there is no window
|
||||
boundary left to bridge.
|
||||
- The setup wizard's consent language states this identity model explicitly;
|
||||
it was updated in the same change that removed the derivation, so no
|
||||
consent was ever collected under the old wording in any shipped build.
|
||||
|
||||
**Byte-identical resends still hold.** The transmitted id is recorded on the
|
||||
row (`sent_install_id`) when the package is first prepared, and the wire body
|
||||
is always rebuilt from that recorded value, so a retry rebuilds identical
|
||||
bytes. The contract requires this: resending a `package_id` with different
|
||||
content is undefined behaviour. (With a stable id the recorded copy is no
|
||||
longer load-bearing against rotation — it remains as the audit column and as
|
||||
cheap insurance against any future change to identity semantics.)
|
||||
|
||||
### A.3 Rotation — removed (decision record)
|
||||
|
||||
Salt rotation was deleted together with the derivation (product decision,
|
||||
2026-08-27). This section is retained as a record of what the earlier design
|
||||
did and why the removal was accepted:
|
||||
|
||||
- Rotation existed to bound long-term linkability: one identity per 30-day
|
||||
window, unrelated identities across windows.
|
||||
- The documented residue (see git history for the full analysis): the
|
||||
envelope's stable, low-entropy `resource` tuple plus contiguous daily
|
||||
periods could plausibly bridge windows for rare configurations anyway, so
|
||||
the boundary was a cost-raiser, not a wall.
|
||||
- The product need that killed it: cross-window continuity is precisely what
|
||||
retention analysis requires. A boundary that mostly inconveniences honest
|
||||
analysis while only raising costs for a determined correlator was judged
|
||||
the wrong trade once stable identity became a requirement.
|
||||
|
||||
There is no salt in the store, no rotation schedule, and no derived
|
||||
identifier anywhere in the pipeline.
|
||||
|
||||
### A.4 Reset behavior
|
||||
|
||||
Removing `$HERMES_HOME/telemetry/shared_metrics` still resets local identity,
|
||||
aggregates, and package files, exactly as documented above. Two honest
|
||||
qualifications now apply:
|
||||
|
||||
- Reset regenerates `install_id`, so subsequent packages transmit a **new**
|
||||
identity. Local reset does give a new remote identity.
|
||||
- Reset **cannot unsend**. Packages already transmitted remain in the ingest
|
||||
service's storage under the identifier they were sent with. There is no
|
||||
read-back or delete API in the v1 contract.
|
||||
|
||||
Setting `send: false` stops transmission immediately: consent is re-read
|
||||
before every package, so a pass already in flight stops after the package it
|
||||
is currently sending rather than draining its whole batch. It does not delete
|
||||
previously transmitted packages, and it does not stop local collection.
|
||||
|
||||
Turning sending off also **closes the consent window** — at the last moment
|
||||
consent was actually observed, not at the wall clock. Packages whose periods
|
||||
fall between one window and the next are never transmitted, even if sending
|
||||
is later re-enabled, and this holds for any number of on/off cycles, across
|
||||
hand-edits with no process running, and under a clock that jumps in either
|
||||
direction (window opens are clamped above every timestamp already in the
|
||||
store; observation marks advance by a bounded step per call, so one glitched
|
||||
forward sample cannot drag the confirmation horizon years ahead; a close
|
||||
never lands after the closing observation's own clock).
|
||||
Unlike the earlier single moving opt-in date, closing and reopening does NOT
|
||||
discard the still-undelivered backlog from a previous consented window —
|
||||
those packages stay inside their own interval and remain eligible.
|
||||
|
||||
One deliberate upgrade-path consequence: packages exported under the
|
||||
pre-interval consent model (before `send_consent_windows` existed) predate
|
||||
the first recorded window and are therefore never transmitted after an
|
||||
upgrade. This is the fail-closed direction — re-importing the old moving
|
||||
day-stamp to release them would re-import the semantics five review rounds
|
||||
showed to be unsound — and it costs at most the undelivered backlog, never
|
||||
collected data.
|
||||
|
||||
### A.5 Retention
|
||||
|
||||
- **Local:** unchanged — 30 days for successfully exported history, and pending
|
||||
deltas are kept until exported. Send state does **not** extend local
|
||||
retention: a package that could never be sent is still pruned at 30 days.
|
||||
Unbounded local growth against a permanently unreachable endpoint is a worse
|
||||
failure than losing metrics from an install that has been broken for a month.
|
||||
- **Remote:** raw packages are retained in S3 without expiry in production and
|
||||
for 30 days in staging.
|
||||
|
||||
### A.6 Deletion
|
||||
|
||||
There is no remote deletion path in the v1 contract, and this appendix does not
|
||||
invent one. What a user can do:
|
||||
|
||||
| Action | Effect |
|
||||
|---|---|
|
||||
| `send: false` | No further packages leave the machine |
|
||||
| `enabled: false` | Collection stops; existing local state remains |
|
||||
| Remove `.../shared_metrics` | Local identity, aggregates, and files reset; future sends use a new install_id |
|
||||
| Delete already-sent data | Not self-service — requires an operator acting on the S3 bucket |
|
||||
|
||||
If a deletion-on-request obligation is ever taken on, the lookup path is now
|
||||
direct: the user's `install_id` (readable from their local store) is the key
|
||||
their data is stored under. Building the service-side delete API remains a
|
||||
new product decision, not an implementation detail.
|
||||
|
||||
### A.7 What the outbox directory is
|
||||
|
||||
Recorded because it was misread once during Phase 2 planning, in a way that
|
||||
would have deleted user data.
|
||||
|
||||
The directory is **local history, not a send-queue**. `package_outbox` is the
|
||||
SQLite table; its `exported_at` column means "written to disk", not "sent".
|
||||
Files are immutable and pruned **by age alone**.
|
||||
|
||||
The ingest contract says senders should delete a package from their outbox on
|
||||
`202`. **The exporter does not do this.** Deleting on acknowledgement would
|
||||
repurpose the user's 30-day local history as a transmission queue and destroy
|
||||
state they were promised. Send state lives in new columns on the
|
||||
`package_outbox` table instead; the files are untouched by transmission.
|
||||
|
||||
### A.8 Scope note
|
||||
|
||||
The `install_id` field inside the package body is transmitted as the
|
||||
generator wrote it (rewritten from the row's frozen `sent_install_id`, which
|
||||
records the same value). No other payload field changes, nothing is added,
|
||||
and the service treats the whole body as opaque. Payload schema evolution
|
||||
therefore stays a sender-side concern, as before.
|
||||
@@ -0,0 +1,117 @@
|
||||
# Profile-Based Routing for Inbound Messages
|
||||
|
||||
> **Audience:** Gateway operators and contributors
|
||||
> **Source files:** `gateway/profile_routing.py`, `gateway/run.py` (`_profile_name_for_source`), `gateway/platforms/base.py` (`build_source`), `gateway/config.py`
|
||||
> **Related:** [Session Lifecycle](session-lifecycle.md), `docs/design/profile-builder.md`
|
||||
|
||||
## Overview
|
||||
|
||||
By default a single gateway run uses one profile (memory, persona, tools). **Profile-based
|
||||
routing** lets one gateway instance serve **multiple isolated profiles**, selecting which
|
||||
profile handles an inbound message based on *where the message came from* — the platform,
|
||||
server (`guild_id`), channel (`chat_id`), and/or thread (`thread_id`).
|
||||
|
||||
This is the inbound counterpart to multiplexing: instead of running N gateways, run one
|
||||
gateway and route per-community / per-channel / per-thread to a dedicated profile. Each
|
||||
profile keeps fully isolated state (`MEMORY.md`, `USER.md`, `SOUL.md`, sessions, tools).
|
||||
|
||||
Routing is **platform-generic**: it works for Discord, Telegram, Feishu, Slack, and every
|
||||
adapter — not just Discord.
|
||||
|
||||
## Configuring routes
|
||||
|
||||
Routes live under `profile_routes` in `config.yaml`. Both the top-level and the nested
|
||||
`gateway.profile_routes` forms are accepted (the nested form is what
|
||||
`hermes config set gateway.profile_routes ...` writes).
|
||||
|
||||
```yaml
|
||||
profile_routes:
|
||||
# Route an entire Discord server (guild) to one profile.
|
||||
- name: server-default
|
||||
platform: discord
|
||||
guild_id: "1234567890"
|
||||
profile: server-profile
|
||||
|
||||
# Override a specific channel within that server with a different profile.
|
||||
- name: support-channel
|
||||
platform: discord
|
||||
guild_id: "1234567890"
|
||||
chat_id: "9876543210"
|
||||
profile: support-profile
|
||||
|
||||
# Pin a Telegram group to a profile (Telegram has no guild_id — chat_id only).
|
||||
- name: tg-group
|
||||
platform: telegram
|
||||
chat_id: "-1001234567890"
|
||||
profile: tg-profile
|
||||
|
||||
# Route a single Discord thread.
|
||||
- name: standup-thread
|
||||
platform: discord
|
||||
guild_id: "1234567890"
|
||||
chat_id: "9876543210"
|
||||
thread_id: "1111111111"
|
||||
profile: standup
|
||||
```
|
||||
|
||||
### Fields
|
||||
|
||||
| Field | Required | Description |
|
||||
|---|---|---|
|
||||
| `name` | yes | Human-readable route identifier (used in logs). |
|
||||
| `platform` | yes | Adapter platform: `discord`, `telegram`, `feishu`, `slack`, … |
|
||||
| `profile` | yes | Target profile name (must exist under `~/.hermes/profiles/<name>`). |
|
||||
| `guild_id` | no | Server/guild (Discord). |
|
||||
| `chat_id` | no | Channel/group/DM id. |
|
||||
| `thread_id` | no | Thread id within a channel. |
|
||||
| `enabled` | no | Default `true`; set `false` to disable a route without removing it. |
|
||||
|
||||
## Matching rules
|
||||
|
||||
A route matches an inbound source when **every discriminator the route declares is satisfied**
|
||||
(conjunctive / AND). A field the route leaves unset is ignored.
|
||||
|
||||
- **`platform`** must equal the source platform exactly.
|
||||
- **`thread_id`** (if set) must equal the source thread id.
|
||||
- **`chat_id`** (if set) must match the source channel **or** its parent — a thread in a
|
||||
channel matches the channel's route (hierarchical match for Discord forums/threads).
|
||||
- **`guild_id`** (if set) must equal the source guild.
|
||||
|
||||
> A route declaring **both** `guild_id` and `chat_id` requires both to hold. A channel match
|
||||
> alone does not satisfy a guild constraint — this is intentional and tested.
|
||||
|
||||
When multiple routes match, the **most specific** one wins. Specificity is additive:
|
||||
|
||||
| Discriminator | Weight |
|
||||
|---|---|
|
||||
| `thread_id` | 8 |
|
||||
| `chat_id` | 4 |
|
||||
| `guild_id` | 2 |
|
||||
| (platform only) | 0 |
|
||||
|
||||
So a thread route (8) beats a channel route (4) beats a guild route (2) within the same server.
|
||||
If no route matches, the message uses the default/active profile.
|
||||
|
||||
## How it works at runtime
|
||||
|
||||
1. An inbound message arrives at a platform adapter.
|
||||
2. `BasePlatformAdapter.build_source` builds the `SessionSource` for the message. Every
|
||||
adapter carries a back-reference to the running `GatewayRunner`
|
||||
(`gateway_runner`, injected in `gateway/run.py`), so it asks the runner to resolve the
|
||||
target profile via `_profile_name_for_source`.
|
||||
3. `_profile_name_for_source` runs the configured routes through `match_profile_route` and
|
||||
stamps `source.profile` with the winning route's profile (or leaves it unset).
|
||||
4. Downstream, `_resolve_profile_home_for_source` chooses the profile home directory
|
||||
(`source.profile` → active profile → `default`) and the session is scoped per-profile, so
|
||||
each routed community gets isolated memory and conversation state.
|
||||
|
||||
Because `gateway_runner` is injected for **all** adapters (declared on `BasePlatformAdapter`),
|
||||
every platform goes through this path — not just Discord.
|
||||
|
||||
## Relationship to multiplexing
|
||||
|
||||
`profile_routes` requires `gateway.multiplex_profiles: true`. Multiplexing is what
|
||||
activates the per-profile runtime scope (per-profile `HERMES_HOME`, secret scope, and
|
||||
profile-namespaced session keys); routing is the decision layer that picks *which*
|
||||
profile a given guild/channel/thread lands in. With multiplexing off, `profile_routes`
|
||||
is ignored entirely — behavior is byte-identical to a single-profile gateway.
|
||||
@@ -0,0 +1,42 @@
|
||||
# Independent core project history recall
|
||||
|
||||
Status: core explicit retrieval implemented and targeted integration-tested; not the complete V4 automatic-memory rollout. The disabled external project-history-recall plugin is not imported, consulted, or modified.
|
||||
|
||||
## Implemented path
|
||||
|
||||
Agent sequential/inline session_search calls bind the live caller ID and current_project scope. Registry calls preserve the trusted session_id kwarg. The model cannot supply a broader scope. Internal Python callers retain the historical positional API, with new options keyword-only.
|
||||
|
||||
`hermes_project_scope.resolve_project_scope` derives a complete metadata-only allowed session set using shared Desktop ownership rules, rejects ambiguity and unrelated principals, and returns an optimistic metadata revision. No UI result limit determines access. Ordinary archives remain eligible; hidden/background rows do not.
|
||||
|
||||
`SessionDB.search_messages(allowed_session_ids=...)` applies a single JSON-bound session set before SQL ranking/LIMIT in FTS/CJK/trigram/LIKE and fallback paths. None retains legacy behavior; an empty set returns no matches. Compaction archives are retained, rewound rows excluded.
|
||||
|
||||
`tools/session_search_project.py` rereads source text before presentation, decodes multimodal storage, projects text only, excludes synthetic control messages and secret-bearing text, and rechecks source hashes plus scope before returning. Raw historical assistant replies are not confirmed facts.
|
||||
|
||||
## Continuation contracts
|
||||
|
||||
- Read: pass returned `next_after_message_id` as `after_message_id`.
|
||||
- Scroll: retain the real original `around_message_id`; pass returned forward/backward frontier. A frontier can be a filtered row and is not evidence.
|
||||
- Long text: use `content_offset` / `content_length` on a real message anchor; content hash covers the full safe decoded text.
|
||||
- Discovery: pass `next_search_cursor` with identical query/sort; cursor is bound to caller and scope revision. Live offsets may repeat sessions, and are not snapshot-consistent.
|
||||
- No project: do not browse unrelated Home sessions; only the caller's own compacted history is available for explicit recovery.
|
||||
- Agent project mode does not support role_filter; it returns unsupported instead of silently exposing tools.
|
||||
|
||||
## Verification receipt
|
||||
|
||||
The canonical runner completed 7 selected files with 121 tests passed, zero failed. Tests cover scope ownership, >5000 scope members, SQL fallback boundaries, source revocation, safe multiline/multimodal projection, pagination, trusted inline identity, and old session search behavior. The project-tool target has exhausted the explicitly authorized 10 starts; do not rerun or rename it to bypass that limit.
|
||||
|
||||
Existing Holographic legacy regressions separately passed 42 tests. The optional project Holographic mode is still under identity review; it is not enabled on the user's profile.
|
||||
|
||||
## Known boundaries
|
||||
|
||||
- Project/source checks are optimistic across multiple stores, not a distributed authorization transaction.
|
||||
- This is Recall subsystem isolation, not an OS sandbox against general file/terminal access.
|
||||
- Global MEMORY/USER and old unscoped Holographic data have not been migrated or claimed to be isolated.
|
||||
- Full source completion/origin policy, import-taint dependencies, automatic injection/revocation of prior api_content, confirmed/revised project facts, durable indexing/restore release gates and full Desktop journey are not complete.
|
||||
- Python callers passing a numeric-string around_message_id together with integer frontiers can hit a TypeError; schema-valid integer anchors are supported. Track as nonblocking compatibility debt before final release.
|
||||
- execute_code identity forwarding is fixed, but session_search remains disallowed by sandbox/agent-loop guards. No extra capability was opened.
|
||||
- No actual model history export/embedding request, user-config change, deployment, restart, or Git commit was performed.
|
||||
|
||||
## Rollout
|
||||
|
||||
Do not mark the whole feature complete or enable automatic retrieval based on these selected tests. Finish Holographic identity fixes, authenticate Desktop RPC, integrate the visible UI, run real user journeys, and close the V4 publication gates first. Source changes do not prove an already-running backend has loaded them.
|
||||
@@ -0,0 +1,54 @@
|
||||
# RCA: SSL CA cert bundle corruption after `hermes update`
|
||||
|
||||
**Status:** resolved by `fix(ssl): surface broken CA bundles before provider calls`
|
||||
**Severity:** P2 — degrades the agent into opaque provider/client failures until the user repairs deps or CA configuration.
|
||||
|
||||
## Summary
|
||||
|
||||
A partial `hermes update`, interrupted venv repair, or stale CA-bundle environment variable can leave Python TLS configuration pointing at a missing, empty, or unloadable CA bundle. The first outbound HTTPS client creation or request can then fail with a raw `FileNotFoundError: [Errno 2] No such file or directory` or a low-level SSL error that does not name the broken CA path.
|
||||
|
||||
## Root cause
|
||||
|
||||
Hermes uses OpenAI/httpx and requests-based clients for provider calls, model metadata, gateway delivery, and web tools. Those clients inherit CA bundle settings from:
|
||||
|
||||
- `HERMES_CA_BUNDLE`
|
||||
- `SSL_CERT_FILE`
|
||||
- `REQUESTS_CA_BUNDLE`
|
||||
- `CURL_CA_BUNDLE`
|
||||
- the bundled `certifi` package's `cacert.pem`
|
||||
|
||||
When the venv is partially refreshed, or when one of those env vars points at a file that no longer exists, provider client construction can fail before Hermes has enough context to produce a useful message.
|
||||
|
||||
## Fix
|
||||
|
||||
`agent/ssl_guard.py` validates CA bundle configuration before the OpenAI-compatible provider client is created in `agent/agent_init.py`. It:
|
||||
|
||||
1. Checks explicit CA bundle env vars and reports the exact broken variable/path,
|
||||
2. Verifies `certifi` is importable,
|
||||
3. Verifies `certifi.where()` points at an existing file of plausible size,
|
||||
4. Builds an `ssl.SSLContext` from each checked bundle,
|
||||
5. Raises a typed `SSLConfigurationError` with a repair hint before httpx/OpenAI can raise a raw low-level error.
|
||||
|
||||
`hermes_cli doctor` exposes the same check under `SSL / CA Certificates`, so users can diagnose the problem without starting a model session.
|
||||
|
||||
## Recovery
|
||||
|
||||
When the guard fires during agent init, the user sees a message like:
|
||||
|
||||
```text
|
||||
Failed to initialize OpenAI client: SSL_CERT_FILE points to a missing CA bundle: C:\path\to\missing\cacert.pem
|
||||
Repair: python -m pip install --force-reinstall certifi openai httpx
|
||||
If you configured a custom corporate CA bundle, fix or unset the broken CA bundle environment variable.
|
||||
```
|
||||
|
||||
For a normal corrupted Hermes venv, reinstall the affected client dependencies:
|
||||
|
||||
```bash
|
||||
python -m pip install --force-reinstall certifi openai httpx
|
||||
```
|
||||
|
||||
For a custom/corporate CA setup, fix the env var so it points at a real PEM bundle, or unset it if Hermes should use the bundled `certifi` store.
|
||||
|
||||
## Environment escape hatch
|
||||
|
||||
Set `HERMES_SKIP_SSL_GUARD=1` to bypass the preflight check. This is intended only for sandboxed or managed-trust environments where the Python CA path looks unusual but downstream clients are known to work.
|
||||
@@ -0,0 +1,786 @@
|
||||
# Relay ↔ Connector Contract (v1, EXPERIMENTAL)
|
||||
|
||||
> **Status:** EXPERIMENTAL. This contract MAY CHANGE without a deprecation
|
||||
> cycle until at least two real Class-1 platforms (Discord + Telegram) have
|
||||
> validated it. Evolution during the experimental phase is **additive-only**,
|
||||
> gated by `contract_version`. A breaking change updates both repos in lockstep.
|
||||
|
||||
This document is the formal interface between the **Hermes gateway** (Python,
|
||||
`gateway/relay/`) and the **connector** (Node/TypeScript,
|
||||
`NousResearch/gateway-gateway`). The connector implementer's first action is to
|
||||
read this file.
|
||||
|
||||
The gateway runs a generic `RelayAdapter` that dials **out** to the connector,
|
||||
receives a `CapabilityDescriptor` at handshake, then exchanges normalized
|
||||
`MessageEvent`s (inbound) and actions (outbound) over a per-turn bidirectional
|
||||
WebSocket. The gateway never learns which concrete platform is fronting it; the
|
||||
connector owns all platform-specific socket/identity logic.
|
||||
|
||||
---
|
||||
|
||||
## 1. Handshake
|
||||
|
||||
1. Gateway opens the transport (`connect`).
|
||||
2. Gateway calls `handshake()`; connector returns a `CapabilityDescriptor`
|
||||
(section 2) describing the platform this adapter instance fronts.
|
||||
3. Gateway configures the adapter from the descriptor (char limit, length unit,
|
||||
draft/edit/thread/markdown capabilities) and registers an inbound handler.
|
||||
4. Connector then streams inbound events and accepts outbound actions.
|
||||
|
||||
`contract_version` (currently `1`) is carried in the descriptor. The gateway
|
||||
ignores unknown descriptor fields (forward-compat) and fills missing optional
|
||||
fields from defaults.
|
||||
|
||||
---
|
||||
|
||||
## 2. CapabilityDescriptor (handshake payload)
|
||||
|
||||
JSON object. Source of truth: `gateway/relay/descriptor.py`.
|
||||
|
||||
| Field | Type | Required | Meaning |
|
||||
| --- | --- | --- | --- |
|
||||
| `contract_version` | int | yes | Contract version (additive-only within a version). |
|
||||
| `platform` | string | yes | Platform name (e.g. `"discord"`, `"telegram"`). |
|
||||
| `label` | string | yes | Human-readable label. |
|
||||
| `max_message_length` | int | yes | Char limit; gateway exposes as `MAX_MESSAGE_LENGTH`. 0 → treat as 4096. |
|
||||
| `supports_draft_streaming` | bool | yes | Native draft-streaming preview support. |
|
||||
| `supports_edit` | bool | yes | Edit-based streaming possible; if false, consumer degrades to one-message-per-segment. |
|
||||
| `supports_threads` | bool | yes | `create_handoff_thread` capability. |
|
||||
| `markdown_dialect` | string | yes | `"plain"`, `"markdown_v2"`, `"discord"`, … (drives `supports_code_blocks`). |
|
||||
| `len_unit` | string | yes | `"chars"` (builtin len) or `"utf16"` (Telegram UTF-16 code units). |
|
||||
| `emoji` | string | no | Display emoji (default 🔌). |
|
||||
| `platform_hint` | string | no | System-prompt platform hint. |
|
||||
| `pii_safe` | bool | no | Redact PII in session descriptions. |
|
||||
| `supports_context` | bool | no | Whether the connector can supply surrounding channel/group **context** for an addressed turn on this platform (Model A on-demand history fetch — Discord/Slack/Matrix; Model B passive buffer — Telegram/Signal/WhatsApp). Default false ⇒ no `context` is attached to inbound events. See §3. |
|
||||
| `supports_inchannel_continuable` | bool | no | Whether the platform can host a **flat continuable cron surface** (native Slack's `cron_continuable_surface: in_channel`): the brief posts top-level in the channel/DM and a plain reply continues the job via the flat `(platform, chat_id, None)` session. Default false ⇒ the gateway's scheduler fails safe to thread mode (D6 gate), so an older connector keeps today's thread behavior. |
|
||||
| `supports_block_formatting` | bool | no | Whether this platform's sender renders **block-level formatting** from raw markdown when the gateway stamps `metadata.format_hints` on `send`/`edit` frames (Slack: native `markdown` block for tables/lists/code, mrkdwn text kept as fallback). Default false ⇒ the gateway never stamps hints, so an older connector never receives the metadata. |
|
||||
| `supported_ops` | string[] | no | Op-level capability discovery: the outbound op names the connector's sender for this platform actually implements (e.g. `["send", "edit", "typing", "follow_up", "get_chat_info"]`). Absent/empty ⇒ the connector predates the field and the gateway assumes the legacy op set (`send`/`edit`/`typing`/`follow_up`); a NEW op is used only when explicitly advertised. |
|
||||
|
||||
Most fields are a projection of the gateway's existing `PlatformEntry`; the
|
||||
runtime-only fields (`len_unit`, `supports_*`, `markdown_dialect`) come from the
|
||||
live platform adapter's capability methods.
|
||||
|
||||
---
|
||||
|
||||
## 3. Inbound: `MessageEvent` envelope
|
||||
|
||||
The connector normalizes each platform wire event into a `MessageEvent`
|
||||
(`gateway/platforms/base.py`) and delivers it to the gateway. **Inbound is
|
||||
delivered over the gateway's OUTBOUND `/relay` WebSocket** (see the transport
|
||||
note below) — the connector pushes an `inbound` frame down the socket the
|
||||
gateway already dialed. The gateway keys the session via `build_session_key()`
|
||||
from the embedded `SessionSource` — so populating the right discriminators is
|
||||
the single highest-correctness responsibility of the connector.
|
||||
|
||||
### Inbound transport (WS back-channel, not HTTP)
|
||||
|
||||
The gateway dials **out** to the connector's `/relay` WebSocket for the
|
||||
handshake + outbound actions (§4) + its own `/stop` egress (§5). Inbound rides
|
||||
the **same socket** in the other direction: the connector pushes an `inbound`
|
||||
frame (and `interrupt_inbound` for §5) down the gateway's outbound WS. There is
|
||||
**no gateway-side inbound HTTP endpoint** — a gateway need not (and, when hosted,
|
||||
cannot) expose any inbound port; everything flows over the connection it
|
||||
initiated.
|
||||
|
||||
**Multi-instance routing.** The connector instance that owns a platform's socket
|
||||
(and thus produces inbound events) is generally **not** the instance the gateway
|
||||
dialed its outbound WS into. The producing instance therefore publishes the
|
||||
event on the connector's internal **relay bus** (Redis pub/sub; `RelayBus` in
|
||||
`src/core/relayBus.ts`) keyed by tenant. Every connector instance subscribes and
|
||||
routes each message to its **local** sessions for that tenant
|
||||
(`RelayServer.routeBusMessage`); the single instance that actually holds the
|
||||
gateway's socket delivers it, and instances with no local session for the tenant
|
||||
no-op. Cross-instance delivery is thus an in-cluster Redis hop, not a public
|
||||
HTTP call.
|
||||
|
||||
Frames (connector → gateway, over the WS):
|
||||
|
||||
- `{"type":"inbound", "event": <MessageEvent>, "bufferId"?}`
|
||||
- `{"type":"interrupt_inbound", "session_key", "chat_id"}` (§5)
|
||||
- `{"type":"passthrough_forward", "forward": <PassthroughForward>, "bufferId"?}` (§5.1)
|
||||
|
||||
**Channel context on inbound (design relay-channel-context).** When the source
|
||||
platform's descriptor advertised `supports_context` (§2) and the chat is
|
||||
multi-party (`chat_type` ∈ group/channel/thread/forum, never `dm`), the
|
||||
connector MAY attach two optional, additive fields to the inbound `MessageEvent`:
|
||||
|
||||
- `context`: an array of read-only surrounding messages (same channel, oldest→
|
||||
newest) — nearby non-addressed chatter the connector fetched (Model A) or
|
||||
buffered (Model B). REFERENCE ONLY: it never triggers the agent (the trigger
|
||||
decision was already made connector-side on the addressed event alone). The
|
||||
gateway renders it into `MessageEvent.channel_context` (the same read-only
|
||||
injection path history-backfill uses).
|
||||
- `context_error`: bool, true when the platform is context-capable but the
|
||||
fetch/buffer failed and the connector fail-opened to an empty `context`
|
||||
(observability marker; surfaced connector-side via the delivery span).
|
||||
|
||||
Both absent ⇒ byte-identical to today. A connector that never sends them, or a
|
||||
`dm`, or a no-context platform, yields no `channel_context`.
|
||||
|
||||
`PassthroughForward` is the wire form of a forwarded passthrough-plane request
|
||||
(Class-2/3 webhooks — Discord interactions, Twilio): `{platform, botId, method,
|
||||
path, headers: [[k,v],…], bodyB64, profile?}`. `profile` is optional — the
|
||||
connector stamps it when NAS resolves the target profile for a Team-Gateway
|
||||
interaction; omitting it (single-profile gateways) preserves legacy routing to
|
||||
the default `agent:main` session namespace, mirroring the `profile` field the
|
||||
`inbound` frame's `SessionSource` already carries (#60586). The body is
|
||||
base64-encoded so arbitrary bytes survive the newline-delimited-JSON transport;
|
||||
the gateway base64-decodes
|
||||
back to the exact bytes the connector forwarded (the connector already verified
|
||||
the provider signature and stripped any shared-identity credential at the edge —
|
||||
§6 — so the gateway re-processes a sanitized, token-free body and acts on it via
|
||||
the token-less `follow_up` path). See §3.1.
|
||||
|
||||
**Trust.** The WS upgrade is authenticated with the gateway's per-gateway secret
|
||||
(§6.1), so the channel is trusted end to end — inbound frames are not separately
|
||||
HMAC-signed (the authenticated socket subsumes the per-delivery origin proof the
|
||||
old HTTP path needed). The relay-bus hop is inside the connector trust domain
|
||||
(same as the lease/buffer/capability stores).
|
||||
|
||||
> Earlier drafts of this contract delivered inbound over a signed **HTTP POST**
|
||||
> to a `gatewayEndpoint` (`HttpGatewayDelivery` + a gateway-side
|
||||
> `inbound_receiver`), HMAC-signed with a per-tenant delivery key. That required
|
||||
> every gateway to expose a reachable inbound URL — impossible for hosted
|
||||
> gateways, which have no public IP. The WS back-channel above replaces it; the
|
||||
> per-tenant delivery key is retained at provision for forward-compat but is no
|
||||
> longer used for inbound. The **passthrough plane** (Class-2/3 webhooks like
|
||||
> Discord interactions / Twilio) historically still used `gatewayEndpoint` for
|
||||
> its post-ACK forward; Phase 5 §5.1 moves that forward onto the WS too (the
|
||||
> `passthrough_forward` frame above), so a hosted gateway needs zero public
|
||||
> inbound surface and `gatewayEndpoint` is retired once the cutover lands.
|
||||
|
||||
### 3.1 Passthrough-plane forward (§5.1)
|
||||
|
||||
The passthrough plane answers the provider's latency-critical ACK at the
|
||||
connector EDGE (e.g. Discord's deferred interaction response within ~3s), then
|
||||
does a **fire-and-forget** forward of the real request to the gateway. That
|
||||
forward needs no response back (the provider was already satisfied), so it rides
|
||||
the same outbound WS as `inbound` via a `passthrough_forward` frame rather than
|
||||
an HTTP POST. The gateway processes the decoded request through its normal agent
|
||||
path (a Discord interaction is decoded to a `MessageEvent` and handled like a
|
||||
message; the reply egresses over the outbound / `follow_up` path). `bufferId` is
|
||||
present when the forward was buffered (Phase 5 §5.3 buffered-only flip) and the
|
||||
gateway acks it after durable handoff.
|
||||
|
||||
|
||||
|
||||
### SessionSource fields (the wire surface)
|
||||
|
||||
Source of truth: `SessionSource.to_dict()` in `gateway/session.py`. These are
|
||||
every key the gateway accepts on the wire. `platform`, `chat_id`, `chat_type`,
|
||||
`user_id`, `user_name`, `thread_id`, `chat_name`, and `chat_topic` are always
|
||||
present (may be `null`); the rest are included only when set.
|
||||
|
||||
| Field | Type | Always sent | Meaning |
|
||||
| --- | --- | --- | --- |
|
||||
| `platform` | string | yes | Platform name (matches the descriptor's `platform`). |
|
||||
| `chat_id` | string | yes | Primary conversation id (channel/chat). Session-key discriminator. |
|
||||
| `chat_type` | string | yes | `dm` / `group` / `channel` / `thread` / `forum`. |
|
||||
| `chat_name` | string\|null | yes | Human-readable chat name. |
|
||||
| `user_id` | string\|null | yes | Message author id. Session-key discriminator. |
|
||||
| `user_name` | string\|null | yes | Author display name. |
|
||||
| `thread_id` | string\|null | yes | Thread/forum-topic id when in a thread. Session-key discriminator. |
|
||||
| `chat_topic` | string\|null | yes | Channel topic/description (Discord, Slack). |
|
||||
| `user_id_alt` | string | no | Platform-specific stable alt id (Signal UUID, Feishu union_id). |
|
||||
| `chat_id_alt` | string | no | Alternate chat id (e.g. Signal group internal id). |
|
||||
| `scope_id` | string | no | Platform-neutral **scope** discriminator: Discord guild / Slack workspace / Matrix server. **REQUIRED for Discord/Slack scope isolation.** Session-key discriminator. (Canonical name as of the D-Q2.5 wire migration.) |
|
||||
| `guild_id` | string | no | **Legacy alias, no longer read by the connector.** As of D-Q2.5c the connector reads and writes only `scope_id`; the gateway's agent-wide `SessionSource.to_dict()` still emits `guild_id` (mirrored to `scope_id`) for non-relay session persistence, so it may still appear on the wire but the connector ignores it. Do not depend on it. |
|
||||
| `parent_chat_id` | string | no | Parent channel when `chat_id` refers to a thread. |
|
||||
| `message_id` | string | no | Id of the triggering message (for pin/reply/react). |
|
||||
|
||||
> `is_bot` (author-is-a-bot/webhook classification) exists on the gateway-side
|
||||
> dataclass but is **intentionally NOT on the wire** in v1 — it is not part of
|
||||
> `to_dict()`. Do not add it to the connector's `SessionSource` until it is
|
||||
> first added here and to `to_dict()` (additive bump).
|
||||
|
||||
### SessionSource discriminators per platform
|
||||
|
||||
| Platform | chat_id | chat_type | user_id | thread_id | scope_id |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| **Discord** | channel id | `dm`/`group`/`thread` | author id | thread channel id (threads) | **guild id** (REQUIRED for server isolation) |
|
||||
| **Telegram** | chat id | `dm`/`group`/`forum` | from id | forum topic id (forums) | — |
|
||||
|
||||
**Get Discord's `guild_id` wrong and two servers collide into one session.**
|
||||
This is the #1 High-severity risk. The gateway's `build_session_key()` is the
|
||||
conformance oracle: for a given `SessionSource`, the connector's normalization
|
||||
must produce the same key the Python adapter would. (The Phase-1 stub tests
|
||||
assert known-input → known-key.)
|
||||
|
||||
### Bot identity vs tenant (single-bot consolidation, Appendix A)
|
||||
|
||||
The envelope carries the **originating bot identity** as a field **distinct from
|
||||
tenant**. Tenant is resolved from the event's own discriminator (Discord
|
||||
`guild_id`, Telegram `chat_id`, webhook path/subdomain) — **never** from which
|
||||
token/socket/process delivered it. This keeps one shared bot able to front many
|
||||
tenants (Phase 6) without overloading an existing field.
|
||||
|
||||
### Author-first resolution + the account-link (DM) path (Phase 7)
|
||||
|
||||
Phase 7 adds **self-serve, per-user onboarding to a shared bot**, which changes
|
||||
*which* discriminator resolves the instance for a routed inbound message — and
|
||||
adds a management path for users to bind their own account.
|
||||
|
||||
**Author-first resolution (the multi-tenant-guild rule, D-7.2).** A single
|
||||
Discord guild may hold **many** tenants — different members each linked to their
|
||||
own agent. So for delivery the connector resolves the destination instance from
|
||||
the **authenticated author binding** (`user_instance_binding`, keyed by
|
||||
`(tenant, platform, platform_user_id)` via `resolveByUser`), **NOT** by a
|
||||
guild→instance route. Concretely:
|
||||
|
||||
- A routed message authored by a **linked** user reaches **only that user's**
|
||||
instance — even when a second linked user in the **same guild** is served by a
|
||||
different instance (each reaches only their own).
|
||||
- A message authored by an **unlinked** user resolves to **no** instance and is
|
||||
dropped (**fail-closed** — never broadcast to the guild's other tenants).
|
||||
- The author id used is the **authentic `user_id` off the observed event**, the
|
||||
same `SessionSource.user_id` documented above — never a value asserted by a
|
||||
gateway or carried in a management frame.
|
||||
|
||||
This is the per-`user_id` owner-only routing the connector enforces in
|
||||
`WsGatewayDelivery` (the gateway-side multi-tenant-guild E2E driver
|
||||
`gateway_multitenant_guild_driver.py` is the cross-repo oracle).
|
||||
|
||||
**The account-link (DM) path.** A user binds their account to an instance with a
|
||||
one-time code, redeemed by DMing the shared bot:
|
||||
|
||||
1. The owner triggers a link from the Portal (or a self-hosted CLI). The
|
||||
connector mints a short-lived **link code** for the **authenticated**
|
||||
instance (`POST /manage/link`; instanceId comes from the caller's principal —
|
||||
a NAS-signed `aud=agent:{instanceId}` token or the instance's own per-gateway
|
||||
secret — **never** the request body).
|
||||
2. The user sends `/link <code>` as a **direct message** to the shared bot from
|
||||
the account they want to bind.
|
||||
3. The connector's inbound observer **consumes** that DM (it is not routed to any
|
||||
agent) and writes the `user_instance_binding` using the **authentic
|
||||
`user_id`** off the observed DM event. From then on, author-first resolution
|
||||
routes that user's messages to the bound instance.
|
||||
|
||||
**Opt-out is connector-authoritative.** Deprovisioning an instance
|
||||
(`POST /manage/deprovision`) drops its author bindings (so its users stop
|
||||
resolving to it) **and** revokes its per-gateway secret (so its socket can no
|
||||
longer authenticate — the next WS upgrade is closed **4401**). A gateway that
|
||||
sees a **4401 close after a previously-successful handshake** treats it as a
|
||||
terminal revocation: it stops reconnecting and reports the relay platform as
|
||||
**disabled** (not a retryable error). A 4401 *before* any successful handshake
|
||||
stays retryable (a cold-start / not-yet-provisioned race, not a revocation).
|
||||
|
||||
### 3.2 Going-idle / buffered-flip primitive (§5.3)
|
||||
|
||||
A scale-to-zero PRIMITIVE (not the behaviour — nothing here decides to sleep or
|
||||
suspends a machine; a later workstream consumes these frames). It lets a gateway
|
||||
enter a drain/idle transition without losing inbound that arrives while it is
|
||||
gone, by making the connector buffer for that instance and replay on reconnect.
|
||||
|
||||
Three frames (all keyed by the connection's **authenticated** per-instance id —
|
||||
read off the stored secret record at the WS upgrade, never asserted in a frame):
|
||||
|
||||
- `{"type":"going_idle"}` (gateway → connector) — emitted as part of the
|
||||
gateway's EXISTING drain transition (the adapter sends it before tearing down
|
||||
the socket). Asks the connector to flip this instance to **buffered-only**.
|
||||
- `{"type":"going_idle_ack"}` (connector → gateway) — the connector has flipped:
|
||||
live delivery has stopped and subsequent inbound for this instance buffers
|
||||
durably. The gateway **stays serving until this ack** (so an event landing in
|
||||
the flip window is delivered live, not lost — the same SUBSCRIBE-before-serve
|
||||
ordering discipline as the bus). Only after the ack is it safe to close.
|
||||
- `{"type":"inbound_ack", "bufferId"}` (gateway → connector) — durable receipt of
|
||||
a buffered `inbound` delivery (which carries its `bufferId`) replayed on
|
||||
reconnect. The connector acks the buffer entry only after this, giving
|
||||
drain-without-dup on the **delivery leg**: an instance that dies mid-drain
|
||||
redelivers exactly the unacked tail; an acked entry never redelivers.
|
||||
|
||||
**Buffer + drain.** While flipped, the connector appends inbound to a durable
|
||||
per-instance delivery-leg buffer (`delivery:<instanceId>`) instead of pushing it
|
||||
live. On the gateway's **reconnect** (a NET-NEW reconnect loop re-dials +
|
||||
re-handshakes after an unexpected close), the new handshake triggers the
|
||||
connector to drain that backlog over the new socket **in order, ack-gated**,
|
||||
then clear the flip so live delivery resumes. This reuses the same
|
||||
`drainWithoutDup` machinery as the Discord→connector ingest leg, applied to the
|
||||
connector→gateway delivery leg. Connector-authoritative throughout: a gateway can
|
||||
only flip/drain ITS OWN instance.
|
||||
|
||||
> NOT in scope (deferred behaviour): the autonomous idle timer that DECIDES to
|
||||
> drain, the actual machine suspend, and the NAS suspended-health model. The
|
||||
> primitive is "when the gateway drains, relay flips to buffered + replays on
|
||||
> reconnect, with no loss/dup"; WHAT triggers the drain is out of scope.
|
||||
|
||||
### 3.3 Wake poke (§5.2)
|
||||
|
||||
The other half of the sleep/wake loop: how a SUSPENDED gateway finds out it has
|
||||
buffered work waiting. A PRIMITIVE — nothing here suspends a machine; it wires
|
||||
the wake SIGNAL so a future scale-to-zero behaviour layer can rely on "buffered
|
||||
⇒ wake poked."
|
||||
|
||||
- **Registration.** The gateway registers a **wake URL** at enroll/provision —
|
||||
any reachable URL the connector can GET to wake it (a Fly autostart hostname,
|
||||
a dashboard host). Self-hosted: `hermes gateway enroll --wake-url <url>` (or
|
||||
`GATEWAY_RELAY_WAKE_URL` / `gateway.relay_wake_url`). Managed/NAS: stamped into
|
||||
the container env beside `GATEWAY_RELAY_URL`. Forwarded in the
|
||||
`/relay/provision` body as `wakeUrl` and stored per-instance on the connector's
|
||||
secret record (gateway-asserted but safely scoped — same posture as
|
||||
`instanceId`; the org/tenant stays token-verified, so a gateway can only
|
||||
register a wake target for ITS OWN instance). DISTINCT from the retired
|
||||
`gatewayEndpoint`: a **poke target**, not a delivery target.
|
||||
- **The poke.** When a buffered-only (going-idle) destination receives its FIRST
|
||||
buffered event, the connector issues a **payload-free, unsigned GET** to that
|
||||
instance's registered `wakeUrl`, **directly** (NOT NAS-mediated — relay stays
|
||||
NAS-independent). It carries no tenant data and no inbound: it only says "you
|
||||
have buffered work, reconnect." Tenant authority is re-established the normal
|
||||
way when the gateway re-dials (the authenticated WS upgrade), so a leaked/
|
||||
guessed wake URL can at worst cause a spurious reconnect of ITS OWN instance.
|
||||
Rate-limited per instance (one poke per cooldown window, not per event), and
|
||||
best-effort — a failed poke is swallowed; the gateway still drains whenever it
|
||||
next reconnects on its own. No new frame: the wake is an out-of-band HTTP GET,
|
||||
not a relay-WS message (the socket is down — that's the whole point).
|
||||
|
||||
> NOT in scope (deferred behaviour): the actual machine suspend (Fly
|
||||
> `autostop:"suspend"`) and the autonomous idle timer that decides to sleep. The
|
||||
> primitive is "buffered event for a sleeping instance ⇒ its wakeUrl gets poked";
|
||||
> WHAT makes the instance sleep (and wake-to-serve) is the behaviour layer.
|
||||
|
||||
### 3.4 Obligations on a future scale-to-zero behaviour layer
|
||||
|
||||
§3.2 and §3.3 ship the **primitives**; this section is the **contract a separate
|
||||
scale-to-zero behaviour workstream must honour to consume them safely.** It owns
|
||||
the *decision* to suspend, the actual machine suspend, and the platform/health
|
||||
model — none of which live here — but it MUST hold these guarantees, which the
|
||||
primitives assume:
|
||||
|
||||
1. **Register a `wakeUrl` before the instance can ever be suspended.** A
|
||||
suspended instance with no registered `wakeUrl` is a black hole — buffered
|
||||
inbound never triggers a poke, so it sleeps through its own traffic until
|
||||
something else reconnects it. The behaviour layer MUST ensure a reachable
|
||||
wake target is registered (self-hosted: `--wake-url`; managed: stamped) as a
|
||||
precondition of allowing suspend. A wake URL that is unreachable while the
|
||||
machine is suspended (e.g. points at the suspended machine itself with no
|
||||
platform autostart in front) is equivalent to none.
|
||||
2. **Drain through `going_idle` → await `going_idle_ack` BEFORE tearing down the
|
||||
socket or suspending.** Never suspend with an un-acked flip in flight. The
|
||||
ack is the connector's confirmation that delivery for this instance is now
|
||||
buffered-only; a machine that suspends after sending `going_idle` but before
|
||||
the ack can drop the inbound that races the flip. The gateway already gates
|
||||
socket teardown on the ack (Q-5.3c); the suspend step MUST sit *after* a
|
||||
clean drain completes, not race it.
|
||||
3. **Keep the NET-NEW reconnect loop live as a precondition of suspend.** The
|
||||
wake→drain contract is "poke ⇒ the gateway re-dials ⇒ the connector drains on
|
||||
the reconnect handshake." If the reconnect loop is disabled, a poke lands on a
|
||||
machine that never re-dials and the buffer strands. The behaviour layer must
|
||||
not suspend an instance whose relay transport won't reconnect on wake.
|
||||
4. **Treat suspended ≠ down in the health model (Q-5.3b).** A suspended instance
|
||||
is healthy-asleep, not failed. The health/monitoring layer MUST distinguish
|
||||
the two (e.g. via the platform machine-state) so a suspended instance is not
|
||||
restarted, alerted on, or reaped as unhealthy — that would defeat the suspend
|
||||
and can race the wake/drain.
|
||||
5. **The wake poke is best-effort and rate-limited — do not assume exactly-once
|
||||
or immediate wake.** At most one poke per cooldown window per instance, and a
|
||||
failed poke is swallowed. The behaviour layer must not rely on the poke as a
|
||||
guaranteed/prompt signal; correctness still rests on "the gateway drains
|
||||
whenever it next reconnects." A belt-and-suspenders wake (e.g. a scheduled
|
||||
job that also reconnects) is the behaviour layer's call, not the primitive's.
|
||||
6. **Suspend only when genuinely idle — and idle is connector-observable, not
|
||||
gateway-guessed.** WHAT counts as idle (no in-flight turn + no inbound for N
|
||||
min) is the behaviour layer's policy, but it must compose with the existing
|
||||
drain machinery (`gateway_state` running→draining) rather than introduce a
|
||||
parallel relay-only idle path — the same integration constraint §3.2 places
|
||||
on `going_idle`.
|
||||
|
||||
These are guarantees the behaviour layer OWES the primitives; the primitives owe
|
||||
the behaviour layer only what §3.2/§3.3 already specify (a flip-on-going_idle,
|
||||
a durable per-instance buffer + ack-gated reconnect drain, and a poke on the
|
||||
first buffered event for a flipped instance).
|
||||
|
||||
---
|
||||
|
||||
## 4. Outbound: action set
|
||||
|
||||
The gateway calls the transport with action dicts. Source of truth:
|
||||
`gateway/relay/transport.py` + `gateway/relay/adapter.py`.
|
||||
|
||||
| `op` | Fields | Result |
|
||||
| --- | --- | --- |
|
||||
| `send` | `chat_id`, `content`, `reply_to?`, `metadata?` | `{success: bool, message_id?, error?}` |
|
||||
| `edit` | `chat_id`, `message_id`, `content`, `metadata?` | `{success: bool, error?}` |
|
||||
| `typing` | `chat_id`, `content?`, `metadata?` | `{success: bool}` |
|
||||
| `follow_up` | `session_key`, `kind`, `content`, `metadata?` | `{success: bool, message_id?, error?}` |
|
||||
| `send_media` | `chat_id`, `media_kind`, `source_url`, `content?` (caption), `filename?`, `reply_to?`, `metadata?` | `{success: bool, message_id?, error?}` |
|
||||
| `prompt` | `chat_id`, `prompt_kind`, `prompt_id`, `content` (the question), `options[]{id,label,style?}`, `timeout_s?`, `reply_to?`, `metadata?` | `{success: bool, message_id?, error?}` |
|
||||
| `react` | `chat_id`, `message_id`, `emoji`, `remove?`, `metadata?` | `{success: bool, error?}` |
|
||||
| `thread_create` | `chat_id` (parent), `thread_name`, `message_id?` (anchor), `metadata?` | `{success: bool, thread_id?, error?}` |
|
||||
| `thread_rename` | `chat_id` (parent), `message_id` (the THREAD id), `thread_name`, `only_if_current_name?`, `metadata?` | `{success: bool, error?}` |
|
||||
|
||||
`get_chat_info(chat_id)` is a separate proxied call returning at least
|
||||
`{name, type}`.
|
||||
|
||||
**`send_media` (Phase 2 media egress).** Media crosses the wire BY REFERENCE:
|
||||
`source_url` is either (a) a **connector re-host** the gateway previously
|
||||
uploaded via `POST {connector}/relay/media` (raw bytes body, `Content-Type` +
|
||||
optional `X-Media-Filename` headers, per-gateway HMAC bearer — the same token
|
||||
scheme as the WS upgrade; response `{id, size}` → reference
|
||||
`{connector}/relay/media/{id}`), or (b) a **public http(s) URL** (e.g. a
|
||||
fal.media generation) the connector downloads directly. `media_kind` is one of
|
||||
`image` / `voice` / `audio` / `video` / `document` and selects the
|
||||
platform-native upload lane (Telegram `sendPhoto`/`sendVoice`/…, Discord
|
||||
multipart attachment, Slack external upload, WhatsApp media upload + media
|
||||
message). The caption rides `content` and renders through the platform's
|
||||
normal markdown lane; platforms without native captions get a follow-up text
|
||||
send (connector-side). Both routes and the op are gated on `supported_ops`
|
||||
advertising `send_media` — a legacy connector never sees the op (the gateway's
|
||||
media sends degrade to their pre-media text fallbacks). Size cap 25 MB
|
||||
(connector `mediaStore.ts` MEDIA_MAX_BYTES; uploads over it are rejected 413).
|
||||
|
||||
**Inbound media (Phase 2 media ingress).** An inbound event's `media_urls`
|
||||
carry fetchable references: platform-public URLs pass through (Discord CDN);
|
||||
auth-gated/expiring platform URLs (Telegram file API, Slack `url_private`,
|
||||
WhatsApp Graph media) are downloaded connector-side with the PLATFORM
|
||||
credential and re-hosted as `{connector}/relay/media/{id}` — the platform
|
||||
credential never crosses the wire. Re-host references are readable by any
|
||||
authenticated gateway (capability-URL semantics: the id is 128-bit random and
|
||||
was already delivered to every admitted recipient); the gateway downloads each
|
||||
reference with its per-gateway bearer and presents LOCAL file paths to the
|
||||
agent, mirroring native adapters. Re-hosts expire (TTL ~1h) — download on
|
||||
receipt, not lazily. A parallel `media` array (same order) adds `kind`, `mime`,
|
||||
`size`, `filename`, `caption` metadata; `message_type` reflects the first
|
||||
attachment's kind (`image`/`audio`/`document`).
|
||||
|
||||
**`prompt` (Phase 3 interactive).** One platform-abstract op renders the
|
||||
gateway's highest-stakes interactions (exec approvals, slash confirms,
|
||||
clarify pickers) with NATIVE controls: Discord button components, Telegram
|
||||
inline keyboards, Slack Block Kit actions, WhatsApp button messages (≤3
|
||||
options) / list messages (4–10; >10 degrades to the numbered-text fallback).
|
||||
`prompt_kind` (`approval`/`clarify`/`choice`) is a styling hint only.
|
||||
`prompt_id` is gateway-minted and opaque to the connector; each
|
||||
option's callback payload carries the token `hp1:<prompt_id>:<option_id>`
|
||||
(≤64 bytes — Telegram's `callback_data` cap binds every lane; option ids are
|
||||
`[A-Za-z0-9_.-]`, ≤32 chars). The gateway mints the prompt id as
|
||||
`<per-process nonce>.<8 hex>` within that same alphabet and budget: the
|
||||
connector fans a passthrough forward (a Discord press) out to EVERY live
|
||||
gateway session of the tenant, unlike a message, which it narrows to the
|
||||
admitted instance set, so the nonce is how a gateway tells its OWN prompt
|
||||
from a sibling's. `style` maps per-platform
|
||||
(primary/success/danger/secondary). `timeout_s` is advisory on the wire —
|
||||
expiry is enforced GATEWAY-side (the pending-prompt registry drops expired
|
||||
entries; the owning gateway then replies with a short "no longer waiting"
|
||||
notice).
|
||||
|
||||
**`prompt_response` (Phase 3 inbound).** The user's press crosses back as a
|
||||
normal inbound MessageEvent carrying
|
||||
`prompt_response: {prompt_id, option_id, label?, prompt_message_id?}` — never
|
||||
a bare platform `custom_id`. The event's `text` mirrors `/{option_id}` with
|
||||
`message_type: "command"` so a gateway predating the field routes the press
|
||||
as a typed reply instead of dropping it. A gateway that DOES understand the
|
||||
field always consumes the press instead: a prompt id it did not mint belongs
|
||||
to a sibling gateway that the same fan-out also reached, and letting the
|
||||
`/{option_id}` text reach the chat lane made every sibling answer
|
||||
"Unknown command" under the owner's single ack. The SOURCE is the authentic
|
||||
CLICKING user (connector-observed: Telegram `callback_query.from`, Slack
|
||||
`block_actions.user`, WhatsApp `messages[].from`, Discord interaction
|
||||
member/user), so gateway-side authorization gates apply to a button press
|
||||
exactly as to a typed `/approve`. Ingest lanes: Telegram `callback_query`
|
||||
(polled, `allowed_updates` widened; best-effort `answerCallbackQuery`
|
||||
spinner-stop), Slack `POST /slack/interactions` (raw-bytes HMAC + replay
|
||||
window, same posture as `/slack/commands`), WhatsApp interactive
|
||||
`button_reply`/`list_reply` (webhook normalize arm), Discord type-3
|
||||
component interactions (passthrough §5.1 sanitized forward; the type-3 edge
|
||||
ack is `DEFERRED_UPDATE` so no visible "thinking…" reply). Foreign
|
||||
callback payloads (another integration's buttons) never become prompt
|
||||
events: Telegram/Slack/WhatsApp drop them at the connector; Discord type-3
|
||||
forwards keep the legacy custom_id-as-text shape.
|
||||
|
||||
**`react` (Phase 3 ack lifecycle).** Adds/removes the bot's own `emoji`
|
||||
reaction on `message_id` — restoring the native adapters' 👀→✅/❌
|
||||
processing-lifecycle acks over the relay. Unicode emoji on the wire; the
|
||||
Slack sender maps to Slack's name vocabulary (`eyes`, `white_check_mark`, …)
|
||||
and treats `already_reacted`/`no_reaction` as success (idempotent). Telegram
|
||||
uses `setMessageReaction` (empty set = remove; Telegram's curated-emoji
|
||||
restriction can reject glyphs — the failure is structured and the gateway
|
||||
treats reactions as cosmetic). WhatsApp sends a reaction message (empty
|
||||
emoji = remove). Reactions are best-effort by contract: a `react` failure
|
||||
must never fail a turn.
|
||||
|
||||
**`thread_create` / `thread_rename` (Phase 4 thread lifecycle).** One
|
||||
platform-abstract pair covers handoff threads, Telegram DM/forum topics, and
|
||||
LLM-title semantic renames. `thread_create`: Discord posts a channel thread
|
||||
(type 11) or a message-anchored thread when `message_id` is set; Telegram
|
||||
`createForumTopic` (topic id returned); Slack posts a NAMED seed root
|
||||
message and returns its `ts` (threads there are message-anchored — an
|
||||
explicit `message_id` anchor is echoed back verbatim). The created id rides
|
||||
`SendResult.thread_id`. `thread_rename`: Discord PATCHes the thread channel;
|
||||
Telegram `editForumTopic`. The **`only_if_current_name` no-clobber guard**
|
||||
is the native adapters' human-rename-wins semantics, enforced
|
||||
CONNECTOR-side: Discord reads the current name first and no-ops (structured
|
||||
`success:false`) on mismatch; Telegram has no topic-name read, so a GUARDED
|
||||
rename is unsatisfiable and fails safe (unguarded renames proceed). Slack
|
||||
does not advertise `thread_rename` (a root message's text is content, not a
|
||||
name). WhatsApp advertises neither (no threads).
|
||||
|
||||
**Auto-thread markers + gateway-declared command manifest (Phase 4
|
||||
inbound/handshake).** When the connector's auto-thread egress policy creates
|
||||
a Discord thread, later inbound events from that thread carry
|
||||
`source.auto_thread_created: true` + `source.auto_thread_initial_name` — the
|
||||
connector-observed evidence that lights the gateway's semantic-rename lane
|
||||
(the LLM session title renames the thread via a GUARDED `thread_rename`;
|
||||
per-instance memory, so in an N>1 fleet a miss simply never lights the
|
||||
lane). The gateway may also declare its slash-command set on the Discord
|
||||
`hello` frame (`command_manifest: [{name, description, options?}]`); the
|
||||
connector reconciles Discord's GLOBAL application-command registration
|
||||
against it (GET → diff → bulk PUT overwrite; idempotent, debounced,
|
||||
best-effort — a registration failure never affects the handshake). Commands
|
||||
still dispatch through the passthrough plane as before; the manifest only
|
||||
keeps Discord's registry in sync with what the gateway's dispatcher handles.
|
||||
|
||||
**Inbound `reply_to` enrichment (Phase 4).** A platform reply may carry
|
||||
`reply_to: {text?, author?, is_own?}` alongside `reply_to_message_id` — what
|
||||
the user QUOTED, populated only from data the connector already had in hand
|
||||
(Discord's inline `referenced_message`, Telegram's inline
|
||||
`reply_to_message`, WhatsApp `context.from` + a bounded per-instance
|
||||
inbound-text cache for the text leg). Absent fields mean the platform didn't
|
||||
carry the data — never triggers an extra platform API call. `is_own` = the
|
||||
quoted message was authored by the fronted bot (same evidence as the
|
||||
`is_reply_to_bot` relevance marker). The gateway maps these onto the same
|
||||
MessageEvent reply-context fields native adapters populate.
|
||||
|
||||
**`typing` `content?` (Slack status clear).** A `typing` frame normally omits
|
||||
`content` — the connector renders its platform's active indicator ("is
|
||||
typing…" Assistant status on Slack, one-shot typing elsewhere). An **empty
|
||||
string** `content` is an explicit *clear* request: on Slack the connector sets
|
||||
the Assistant thread status to `""`, dismissing it. The gateway emits the
|
||||
clear only for Slack (persistent status); one-shot platforms never receive it.
|
||||
Additive within `contract_version` 1, but note the deploy order: a connector
|
||||
predating gateway-gateway #154 ignores `content` and would *set* "is typing…"
|
||||
on a clear frame — deploy the connector first.
|
||||
|
||||
**`follow_up` (A2 capability action).** Some inbound payloads carry a credential
|
||||
that acts on the **shared** bot identity (e.g. a Discord interaction follow-up
|
||||
token). Per §6 the connector strips that at the edge and binds it in its
|
||||
capability vault keyed by the session; it **never reaches the gateway**. To use
|
||||
it, the gateway issues `follow_up` naming the **session it is already in**
|
||||
(`session_key`) plus the capability `kind` (e.g. `discord.interaction_token`) —
|
||||
**never a token**. The connector resolves the real value from its vault,
|
||||
enforces the tenant match (tenant B can never wield tenant A's capability), and
|
||||
egresses. `success: false` when the capability is absent/expired or the tenant
|
||||
doesn't match — the gateway has nothing to retry with, by design (a leaked
|
||||
gateway holds zero capability material). Source of truth:
|
||||
`gateway/relay/transport.py` (`send_follow_up`) + `gateway/relay/adapter.py`.
|
||||
|
||||
---
|
||||
|
||||
## 5. Interrupt (`/stop`) routing
|
||||
|
||||
- **Gateway → connector:** `send_interrupt(session_key, reason?)` egresses a
|
||||
mid-turn `/stop` over the outbound WS. The connector MUST forward it to the
|
||||
gateway instance running that `session_key` (the routing invariant).
|
||||
- **Connector → gateway:** an inbound interrupt for a `session_key` is delivered
|
||||
as an `interrupt_inbound` frame down the gateway's outbound WS (§3 transport
|
||||
note) — routed cross-instance via the relay bus to whichever instance holds
|
||||
the socket — and bridged by the adapter's `on_interrupt(session_key, chat_id)`
|
||||
into the existing per-session interrupt mechanism, cancelling exactly that turn
|
||||
(siblings untouched).
|
||||
|
||||
Both directions ride the gateway's outbound WS: the gateway→connector `/stop`
|
||||
egresses over it, and the connector→gateway interrupt rides the same `inbound`
|
||||
back-channel as a normalized event.
|
||||
|
||||
---
|
||||
|
||||
## 6. Trust boundary & signed-body handling (A2)
|
||||
|
||||
**The connector is the sole crypto/identity boundary. The gateway re-validates
|
||||
nothing.**
|
||||
|
||||
Webhook signatures (Discord ed25519, Twilio HMAC, WeCom BizMsgCrypt) are
|
||||
computed over exact raw bytes, and some payloads are *encrypted* with a shared
|
||||
secret. The connector fronts a **shared** bot for many tenants and holds every
|
||||
tenant's platform secrets, so it:
|
||||
|
||||
- **verifies / decrypts at the edge** (the only place the secrets live),
|
||||
- **normalizes** the payload into a tenant-scoped `MessageEvent` (§3),
|
||||
- **strips any shared-identity capability** out of the payload and binds it in
|
||||
its capability vault, keyed by the session (see §4 `follow_up`),
|
||||
- **forwards only the sanitized `MessageEvent`** — never the raw signed body.
|
||||
|
||||
The gateway therefore performs **no** platform signature/crypto verification on
|
||||
the relay path; it trusts the normalized event. This is an enforced invariant on
|
||||
the gateway side (`tests/gateway/relay/test_relay_sheds_crypto.py`: the relay
|
||||
package imports/calls no platform-crypto).
|
||||
|
||||
**Why not "forward the signed body byte-for-byte so the gateway re-validates"?**
|
||||
That earlier model is incoherent under an untrusted, disposable tenant gateway:
|
||||
|
||||
- Re-validating Twilio HMAC / WeCom crypto would require handing the gateway the
|
||||
**shared signing secret** — which is itself the leak, and on a shared bot it's
|
||||
a *cross-tenant* leak.
|
||||
- WeCom payloads are encrypted with the shared secret; the connector must decrypt
|
||||
at the edge just to route, so forwarding ciphertext would again require giving
|
||||
the gateway the secret.
|
||||
- A Discord interaction token lives **inside** the signed JSON body — you cannot
|
||||
both preserve the bytes and strip the credential; they are the same bytes.
|
||||
|
||||
So byte-preservation is abandoned deliberately: the connector re-serializes the
|
||||
sanitized event and the gateway trusts it. This also unifies the passthrough and
|
||||
relay planes — both are "verify at the edge → emit a normalized event," differing
|
||||
only in transport. See `docs/capability-trust-boundary.md` (connector repo:
|
||||
`gateway-gateway`) for the full A2 rationale and the connector-side vault.
|
||||
|
||||
### 6.1 Channel authentication (the connector⇄gateway link itself)
|
||||
|
||||
A2 makes the connector the sole holder of platform secrets while the gateway may
|
||||
be **customer-managed and internet-exposed**, so the connector⇄gateway channel
|
||||
is itself authenticated. The gateway holds an enrollment- or provision-issued
|
||||
**per-gateway secret** (`hermes gateway enroll` → connector `/relay/enroll`, or
|
||||
managed self-provision → `/relay/provision`) that authenticates its outbound WS
|
||||
upgrade. It is an HMAC-SHA256 scheme with a multi-secret rotation verify list
|
||||
(gateway side: `gateway/relay/auth.py`; connector side:
|
||||
`src/core/relayAuthToken.ts`).
|
||||
|
||||
| Leg | Credential | Mechanism |
|
||||
|-----|-----------|-----------|
|
||||
| Gateway → connector WS upgrade | per-gateway secret | An `Authorization` bearer header on the `/relay` upgrade. The token is `base64url(payload:exp:sig)` where `payload = gatewayId` and `sig = HMAC(payload:exp, secret)`. Connector verifies and rejects the upgrade (**close 4401**) on mismatch/absence/revocation. The authenticated tenant comes from the connector's store, never the `hello` frame. |
|
||||
| Connector → gateway inbound (`inbound` / `interrupt_inbound` frames) | — (rides the authenticated WS) | Inbound is pushed down the gateway's already-authenticated outbound socket (§3), so no per-message signature is needed. A **per-tenant delivery key** is still issued at enroll/provision and retained for forward-compat, but is no longer used to sign inbound. |
|
||||
|
||||
This is the **channel** authenticator — distinct from platform crypto, which the
|
||||
relay path still sheds entirely (§6). The gateway holds zero platform secrets;
|
||||
the per-gateway secret authenticates only the connector link. Full threat model +
|
||||
enrollment/rotation/kill-switch design: `docs/connector-gateway-auth-design.md`
|
||||
(connector repo).
|
||||
|
||||
---
|
||||
|
||||
## 7. Per-instance delivery & the management plane (Phase 6)
|
||||
|
||||
Phases 1–5 treat the connector as a single-tenant front: inbound events for a
|
||||
tenant fan out to that tenant's gateway socket(s). **Phase 6 makes delivery
|
||||
per-INSTANCE** — a shared bot can front many users/agents in one tenant (one
|
||||
Discord guild, one Telegram bot) without cross-delivery — and adds a small
|
||||
**management plane** the agent (or a managed Portal) uses to declare who-sees-what
|
||||
and what's-relevant. All of this lives **connector-side**; the gateway's only new
|
||||
responsibility is to **declare its relevance policy** at boot (§7.3).
|
||||
|
||||
### 7.1 The delivery gate (connector-side, informational)
|
||||
|
||||
For each inbound event the connector decides which instances receive it by
|
||||
composing three AND-ed filters. The gateway does not implement these — they run
|
||||
in the connector — but they define the delivery semantics the gateway relies on:
|
||||
|
||||
| Layer | Question | Source of truth |
|
||||
| --- | --- | --- |
|
||||
| **owner / scope ∧ principal** | May this instance *see* this author here? | per-user `user_id → instance` bindings (the owner floor) + per-instance `(guild, channel)` scope grants + an `owner-only` / `allow-list` / `any` principal policy. |
|
||||
| **visibility floor** | Can the instance's bound owner actually `VIEW_CHANNEL` this in Discord? | live Discord ACL (effective permissions), fail-closed. Narrows an over-broad scope grant downward. |
|
||||
| **relevance** | *Given* it may see it, should the agent engage? | the relevance policy declared in §7.3 (address-gating / free-response / allow-bots). |
|
||||
|
||||
The composition only ever **narrows** delivery (`deliver ⇔ authorized ∧ visible
|
||||
∧ relevant`); the **owner floor bypasses the relevance layer** (an author's own
|
||||
message always reaches their own instance — you don't @mention your own agent).
|
||||
A message authored by an unbound user reaches no instance (fail-closed). The
|
||||
full design + invariants live in the connector repo
|
||||
(`NousResearch/gateway-gateway`); this section is the gateway-facing summary.
|
||||
|
||||
### 7.2 Management routes (connector-side, authenticated)
|
||||
|
||||
The connector mounts authenticated management routes. They share the **same
|
||||
dual-auth** as the WS upgrade: either a managed NAS-signed `aud=agent:{instanceId}`
|
||||
RS256 JWT, **or** the gateway's own per-gateway secret bearer (§6.1
|
||||
`make_upgrade_token`). In both cases the connector resolves the authoritative
|
||||
`{tenant, instanceId}` from its **stored** record — **never** from the request
|
||||
body (a body-asserted `instanceId` is ignored).
|
||||
|
||||
| Route | Purpose |
|
||||
| --- | --- |
|
||||
| `POST /manage/link` | Issue a short-lived code to bind a platform account to the authenticated instance (the `/link <code>` flow; the connector reads the authentic `user_id` off the inbound event). |
|
||||
| `POST /manage/scope`, `/manage/scope/release` | Claim / release a `(guild, channel)` scope for the authenticated instance. A channel is owned by at most one instance (non-overlap is a PK constraint). |
|
||||
| `POST /manage/principal` | Set the instance's principal policy (`owner-only` \| `allow-list` \| `any`). |
|
||||
| `POST /manage/dm-default` | Set the user's DM-default instance (DM tie-break when a user linked more than one). |
|
||||
| `POST /relay/policy` | Declare the instance's **relevance policy** (§7.3). |
|
||||
|
||||
These are connector-owned (the management plane is not part of the gateway's
|
||||
agent path); the gateway only calls `POST /relay/policy` (§7.3). The others are
|
||||
driven by the managed Portal / `hermes` CLI.
|
||||
|
||||
### 7.3 Relevance-policy declaration (the gateway's responsibility)
|
||||
|
||||
The relevance layer (§7.1) is the per-tenant parity for the gateway's own
|
||||
behaviour knobs (`require_mention`, `free_response_channels`,
|
||||
`{PLATFORM}_ALLOW_BOTS`). So the **same** behaviour governs relay delivery, the
|
||||
gateway projects those knobs into a **platform-agnostic** policy and POSTs it to
|
||||
`POST /relay/policy` at boot (after its per-gateway secret is resolved).
|
||||
|
||||
Body (`gateway/relay/__init__.py` `relay_relevance_policy()` → `send_relay_policy()`):
|
||||
|
||||
| Field | Type | Projected from | Meaning |
|
||||
| --- | --- | --- | --- |
|
||||
| `platform` | string | the fronted platform (`relay_platform_identity`) | which platform this policy applies to. |
|
||||
| `requireAddress` | bool | `require_mention` | a non-owner message must @mention / reply-to the bot to be relevant. |
|
||||
| `freeResponseScopes` | string[] | `free_response_channels` | scope (channel) ids where `requireAddress` is waived. Same scope vocabulary as §7.1's scope grants. |
|
||||
| `allowOtherBots` | bool | `{PLATFORM}_ALLOW_BOTS ∈ {mentions, all}` | admit bot-authored messages (default off). |
|
||||
|
||||
Auth is the per-gateway upgrade token (§6.1), so the connector attaches the
|
||||
policy to the authenticated instance. The gateway is the **source of truth** and
|
||||
re-declares **every boot** (a full replace, mirroring the `routeKeys` upsert at
|
||||
provision — self-healing). When the projected policy is all-default the gateway
|
||||
sends nothing (the connector's absent-row default already matches). The POST is
|
||||
**fail-soft**: a failure logs and boot proceeds — relevance is an optimization
|
||||
layered on the authorization gate (§7.1), never a boot dependency. There is **no
|
||||
new gateway inbound surface** and **no new credential** — it reuses the
|
||||
per-gateway secret and the same host as `/relay/provision`.
|
||||
|
||||
> A relevance drop happens **before** the connector wakes a scaled-to-zero agent
|
||||
> (Phase 5), so excluded chatter never spins an agent up — relevance is the
|
||||
> primary scale-to-zero lever as well as a correctness filter.
|
||||
|
||||
---
|
||||
|
||||
## 8. Gateway-side platform behavior controls (enterprise)
|
||||
|
||||
Enterprise deployments configure fronted-platform behavior on the GATEWAY
|
||||
side, under `platforms.relay.extra.<platform>` — a supported subset of that
|
||||
platform's native options. The native platform block (e.g. `platforms.slack`)
|
||||
is not read on the relay lane; the connector receives the *outcome* of these
|
||||
controls as frame metadata (§4) and executes mechanically — it holds no
|
||||
platform behavior policy of its own.
|
||||
|
||||
```yaml
|
||||
platforms:
|
||||
relay:
|
||||
extra:
|
||||
slack:
|
||||
reply_in_thread: true # default
|
||||
```
|
||||
|
||||
Resolution: nested `extra.<platform>` object wins → legacy flat key on
|
||||
`extra` honored as fallback → default. Source of truth:
|
||||
`RelayAdapter._effective_reply_in_thread` (`gateway/relay/adapter.py`).
|
||||
Values coerce exactly as the native Slack adapter's do — `1/true/yes/on`
|
||||
(case-insensitive, whitespace-trimmed) are ON, anything else is OFF — so a
|
||||
YAML-quoted `"false"` turns a knob off rather than being read as a truthy
|
||||
string.
|
||||
|
||||
Current controls (Slack):
|
||||
|
||||
| Key | Default | Effect |
|
||||
| --- | --- | --- |
|
||||
| `reply_in_thread` | `true` | `true`: thread-per-message — each top-level DM message anchors its own thread (status, progress, prompts, final reply all carry that `metadata.thread_id`). `false`: flat rolling DM — send-lane frames carry NO thread anchor (stripped, not omitted), one shared session per DM. |
|
||||
| `dm_top_level_threads_as_sessions` | `true` | Native-parity escape hatch (mirrors `platforms.slack.extra.dm_top_level_threads_as_sessions`). `true`: in thread-per-message mode each top-level DM message keys its own session, so concurrent messages run in parallel. `false`: threaded reply placement is kept but the session stamp is skipped — one rolling DM session (legacy steer/queue posture). No effect in flat mode, which always keeps the single rolling session. |
|
||||
|
||||
Typing/status frames always carry the triggering-ts anchor when one is known
|
||||
(liveliness is unconditional, both modes): Slack's status line is
|
||||
thread-scoped, and in flat mode the send-side anchor strip guarantees the
|
||||
status anchor can never leak into reply placement. Semantics of the native
|
||||
key: see `website/docs/user-guide/messaging/slack.md`.
|
||||
|
||||
Thread-anchor resolution applies to EVERY send lane — text (`send`) and media
|
||||
(`send_media`) alike — through one choke point
|
||||
(`RelayAdapter._apply_slack_thread_anchor`). Media frames egress via the same
|
||||
connector-side Slack sender, which threads on `metadata.thread_id` only, so an
|
||||
attachment resolves its anchor identically to a text reply: promoted into
|
||||
metadata in thread-per-message mode, stripped in flat mode.
|
||||
|
||||
Changes take effect on gateway restart; no connector involvement.
|
||||
|
||||
---
|
||||
|
||||
## 9. Versioning policy
|
||||
|
||||
- `contract_version` is an int; bump **only** for additive changes during the
|
||||
experimental phase (new optional fields, new `op`s).
|
||||
- A breaking change (renamed/removed field, changed semantics) requires a
|
||||
coordinated update of both repos and a version bump.
|
||||
- The connector's first PR references the commit SHA of this file it implements
|
||||
against.
|
||||
@@ -0,0 +1,134 @@
|
||||
# Research spike: plugin-architecture lessons from Pi and OpenCode
|
||||
|
||||
**Issue:** #64180 · **Informs:** #64164 (event bus), #64161 (streaming hooks), #64162 (pluggable approval), #64165 (manifest v2), #64229 (lifecycle/ledger), #64230 (Plugin Doctor)
|
||||
|
||||
**Method.** Both systems were read at source level (shallow clones pinned to a commit), not from docs sites alone: Pi (`badlogic/pi-mono`, now `earendil-works/pi`) at `eb79351` (v0.80.7, 2026-07-14) and OpenCode (`sst/opencode`, now `anomalyco/opencode`) at `c69abee` (v1.18.2, 2026-07-16). Claims below carry file:line references into those commits. Where something could not be found (ADRs, policies, timeouts), that absence was verified by search and is reported as a finding. Hermes ground rules from #64182 (additive-only, prompt-cache sacred, observer-first, fail-closed security) are treated as overriding constraints throughout — this report grades imported patterns against them, per the adapt-don't-copy rubric.
|
||||
|
||||
**Headline.** The two systems are near-perfect opposites on the four design axes Hermes currently has open, which makes them a natural controlled experiment:
|
||||
|
||||
| Axis | Pi | OpenCode | Hermes proposal on the table |
|
||||
|---|---|---|---|
|
||||
| Per-delta streaming hook | Yes — awaited inline, no timeout | Structurally absent (text-end only) | Observer per-delta + never-block contract (#64161) |
|
||||
| Veto semantics | Typed per-event result vocabulary (`{block}`, `{cancel}`, `"handled"`) | Veto-by-throw (bug and policy denial indistinguishable) | TBD in #64162 |
|
||||
| Guard-hook failure | Fail closed (`tool_call` only) | No runtime containment at all | Ground rule 4: fail closed on security-adjacent |
|
||||
| Plugin event bus | Yes — 33 lines, **un-namespaced** channels | None (plugins observe core bus, cannot emit) | Namespaced `ctx.emit`/`ctx.subscribe` (#64164) |
|
||||
|
||||
Neither system has hook timeouts. Both have shipped hang-class or drift-class failures because of it. That is the single strongest cross-cutting lesson for Hermes.
|
||||
|
||||
---
|
||||
|
||||
## 1. Pi (badlogic/pi-mono → earendil-works/pi)
|
||||
|
||||
Extensions are in-process TypeScript modules (loaded via jiti) receiving an `ExtensionAPI`; no separate process, no IPC, no manifest permissions. Pi explicitly rejected MCP as its extension mechanism (Zechner, ["What if you don't need MCP at all?"](https://mariozechner.at/posts/2025-11-02-what-if-you-dont-need-mcp/)). Notably, Pi's founding "minimal, no hooks" stance (Nov 2025) reversed into a 33-event extension system within ~7 months — extensibility demand won; the no-MCP stance held.
|
||||
|
||||
### 1.1 Hook/event taxonomy
|
||||
|
||||
33 event types (`src/core/extensions/types.ts:507-902`). The load-bearing design choice: **every mutating event gets its own typed emitter with its own result vocabulary**, rather than one generic middleware pipe —
|
||||
|
||||
- `tool_call` → `{ block: true, reason }`; argument mutation in place, explicitly "no re-validation after your mutation" (`docs/extensions.md:742-765`)
|
||||
- `session_before_*` → `{ cancel: true }`, first canceller wins, later handlers skipped
|
||||
- `input` → `"handled"` (short-circuit) vs `"transform"` (chain) (`runner.ts:1148-1188`)
|
||||
- `tool_result` → partial-patch accumulation across handlers (`runner.ts:835-883`)
|
||||
- `before_agent_start` → system-prompt chaining with a live `ctx.getSystemPrompt()` reflecting earlier handlers (`runner.ts:1034-1098`)
|
||||
- Observer events have **no return channel at all** — the observer/mutator split is enforced by which emitter the event flows through, not by convention.
|
||||
|
||||
Dispatch is fully sequential async (`runner.ts:759-791`): every handler awaited one at a time; the only ordering rule is extension load order (project → global → CLI, `loader.ts:660-708`) then registration order. No priorities, no phases — and no community demand for them in ~8 months. Name collisions are handled per registry with explicit deterministic policies instead of dependency resolution: first-wins tools, suffixed duplicate commands (`name:2`), last-wins shortcuts with a warning plus an 18-key reserved denylist (`runner.ts:421-604`).
|
||||
|
||||
Two ordering guarantees worth copying: extensions see events **before** the UI and **before** session persistence (`agent-session.ts:596-601`), and parallel sibling tool calls are "preflighted sequentially, then executed concurrently" (`docs/extensions.md:750`).
|
||||
|
||||
### 1.2 Plugin-to-plugin interaction
|
||||
|
||||
A 33-line shared event bus (`src/core/event-bus.ts`): `pi.events.emit/on` on **arbitrary string channels** — no namespacing, no declarations, no collision protection; per-handler try/catch and an unsubscribe closure. Everything richer (capability registries, dependencies) is deliberately absent; a community RFC to build richer coordination as an extension (pi#2715) was left to userland.
|
||||
|
||||
### 1.3 Compatibility strategy
|
||||
|
||||
No API versioning, no handshake, no deprecation annotations. In its place, three working practices: (a) loader-level alias shims that kept old imports resolving across the `@mariozechner`→`@earendil-works` package rename and a pi-ai API split, with pre-announced removal (`loader.ts:47-71`; `CHANGELOG.md:243`); (b) loud "Breaking Changes" changelog sections with **automatic migrations** (directory renames on startup, session format v2→v3 auto-migrated) for the big hooks+customTools→extensions unification (v0.35-0.37); (c) per-tool `prepareArguments()` shims for old argument shapes.
|
||||
|
||||
Real breakage on record: pi#2860 — an internal session-management refactor made `pi.sendUserMessage()` after `ctx.newSession()` silently drop messages. The remediation pattern is distinctive: **stale-context poisoning** — after session replacement, every captured context getter throws a paragraph-long teaching error pointing at the safe `withSession` pattern (`runner.ts:514-527`; `docs/extensions.md:1223-1265` documents the footgun with unsafe-pattern code). Compat by making misuse loud, not by never changing.
|
||||
|
||||
### 1.4 Failure isolation
|
||||
|
||||
Contain errors, don't contain time: every handler call is individually try/caught and surfaced (red stack trace in chat UI), load failures skip only the broken extension, and a `pi -ne` no-extensions escape hatch exists. The **one deliberate exception**: `tool_call` has no internal catch — a guard-hook crash blocks the tool ("Extension failed, blocking execution"), converted into an error tool result the LLM sees; the loop survives, the guard fails closed (`runner.ts:885-906`; `agent-session.ts:454-467`; `agent-loop.ts:657-665`).
|
||||
|
||||
**There are zero timeouts.** All handlers — including per-token `message_update` — are awaited inline in the agent event pipeline (`agent-session.ts:598, 728-734`). A hung extension freezes the agent; the changelog records hang-class fixes (pi#5687 background handles, pi#5115 shutdown drains). Mitigations offered are cooperative only (`ctx.signal` AbortSignal, Esc-abort).
|
||||
|
||||
No sandbox, documented as a decision: "a partial in-process sandbox would be easy to misunderstand as a security boundary" (`docs/security.md:33-35`); the only gate is project-trust on load.
|
||||
|
||||
### 1.5 Design-history record
|
||||
|
||||
**Pi has no ADRs** — the Discord recollection that prompted this spike is not substantiated for Pi. Rationale lives in four places: "footgun" sections inside user docs (effectively inline ADRs), a 5,000-line issue-linked changelog (which records abandoned designs: the hooks/customTools split, slash-commands→prompt-templates rename, a fetch-override proxy abandoned for undici dispatchers), Zechner's blog, and issues themselves.
|
||||
|
||||
### 1.6 Prompt/context construction
|
||||
|
||||
Pi is the only surveyed system that treats **prompt-cache stability as part of the extension API contract**: extensions receive structured prompt inputs (`systemPromptOptions` — the same decomposed inputs Pi itself uses) rather than a final string; per-request context transforms operate on a `structuredClone` so session history is never corrupted; and cache-friendly dynamic tool loading (v0.80.6, pi#6474) activates tools additively via native provider deferred loading (Anthropic `defer_loading`) explicitly to avoid invalidating the prefix cache — with documented warnings about second-order invalidation through prompt-metadata changes (`docs/extensions.md:2254-2290`).
|
||||
|
||||
---
|
||||
|
||||
## 2. OpenCode (sst/opencode → anomalyco/opencode)
|
||||
|
||||
Server (Bun/Effect) + clients (TUI/desktop/web); plugins are npm packages or local TS files loaded in-process, with **two plugin kinds in one package** — `server` and `tui` entrypoints, hard rule one-or-the-other (`shared.ts:103-114, 293-295`). A v2 plugin API ships side-by-side with v1 (`/v2/effect`, `/v2/promise` subpaths), designed in a candid 516-line `PLAN.md`.
|
||||
|
||||
### 2.1 Hook/event taxonomy
|
||||
|
||||
One `Hooks` bag returned by an async factory (`packages/plugin/src/index.ts:74, 222-335`): ~16 mutating `(input, output)` hooks (mutate `output` in place; host reads it back) plus declarative registries (tools, auth, providers) and a single firehose `event` observer. The whole dispatch engine is ~13 lines (`plugin/index.ts:280-293`): sequential, awaited, load-order, later hooks see earlier mutations — ordering documented as a spec ("global config → project config → global plugin dir → project plugin dir"; v2: "plugin registration order, then transform registration order").
|
||||
|
||||
**Veto is throw.** The documented idiom for denying a tool call is `throw new Error("Do not read .env files")` (`plugins.mdx:247-257`) — policy denial and plugin bug are indistinguishable in every consumer downstream.
|
||||
|
||||
**A typed-but-dead hook.** `permission.ask` — the only hook with real decision semantics (`output.status: "ask" | "deny" | "allow"`) — exists in the published types but has **no dispatch site anywhere in the tree**: a permission-subsystem rewrite orphaned it, types kept compiling, plugins silently no-op'd (oc#7006, open since Jan 2026). This is the sharpest single failure mode found in the whole spike.
|
||||
|
||||
### 2.2 Plugin-to-plugin interaction
|
||||
|
||||
No dependencies, no registry, no plugin-emitted events. Interaction is (a) blind composition through the shared mutable `output` (last-writer-wins per field) and (b) observing the core event bus. The 2026 core bus itself is heavyweight — event-sourced, durable (SQLite, per-aggregate sequences, idempotent replay), versioned wire types, and a **backpressure guard**: `allBounded` wraps a dropping queue and fails slow subscribers with `SubscriberOverflowError` (`packages/core/src/event.ts:152-164`). Plugins get the untyped tail of a three-tier bridge (Effect streams → global emitter → `event` hook), fired **fire-and-forget** — an async observer's rejection is an unhandled promise rejection invisible to host error routing (`plugin/index.ts:251-258`).
|
||||
|
||||
### 2.3 Compatibility strategy
|
||||
|
||||
Lockstep versioning (`opencode` = `@opencode-ai/plugin` = `@opencode-ai/sdk` = 1.18.2) plus one gate: npm plugins may declare `engines.opencode` semver ranges, checked at load (`shared.ts:194-205`) — but it is plugin-opt-in, and local file plugins skip it entirely. Three generations of module shape are loaded simultaneously; superseded packages are silently ignored via a hardcoded `DEPRECATED_PLUGIN_PACKAGES` list.
|
||||
|
||||
The community-experience record is instructive: v1.14.42 — a **patch release** — removed the whole `api.command.*` TUI namespace with no deprecation cycle (oc#26557); the aftermath is visible in-tree as a deprecated shim added back *after* the outcry ("Legacy `api.command` API kept so v1 plugins can initialize. Remove in v2", `tui.ts:87-120`). No written deprecation policy exists anywhere. Shims-after-outcry is the de facto process.
|
||||
|
||||
### 2.4 Failure isolation
|
||||
|
||||
Strong at the edges, absent in the middle. Load is staged (`install | entry | compatibility | missing | load`) with per-plugin, per-stage containment and user-visible toasts (`loader.ts:82-93`; `plugin/index.ts:215-249`). Runtime hooks have **no catch and no timeout**: a throw in a tool hook fails that tool call (the sanctioned veto), a throw in chat/transform hooks aborts the turn (session errors, server survives), a **hang hangs the turn forever** — the v2 PLAN explicitly lists "Transform timeouts" under *Deferred Decisions* (`PLAN.md:507-510`). Boot re-entrancy burned them: a plugin calling the SDK client during its own init deadlocked startup (oc#7741). The official troubleshooting page's first advice is "start by disabling plugins."
|
||||
|
||||
Two mature exceptions worth stealing: the streaming hot path is protected **structurally** — there is no per-delta hook at all; text hooks fire once at `text-end` (`processor.ts:512-524`) — and the TUI runtime has scope-tracked registrations (a `Proxy`-wrapped keymap API auto-records every registration per plugin, enabling clean live deactivate) plus a **hard 5s dispose budget** racing each cleanup against a timer (`runtime.ts:122-226, 388-468`).
|
||||
|
||||
### 2.5 Design-history record
|
||||
|
||||
No ADR system; the `PLAN.md` for v2 is the exception and is better than most ADR archives — it names v1's mistakes (the returned-hooks bag, finalizer-triggered special cases), specifies ordering as a contract, splits replayable *transforms* from live *hooks*, and — crucially — carries an honest **Deferred Decisions** section (typed error model, transform timeouts) rather than pretending closure.
|
||||
|
||||
### 2.6 Prompt/context construction
|
||||
|
||||
Plugins can touch every layer, but the deep layers are gated behind an `experimental.` prefix (`experimental.chat.system.transform`, `experimental.chat.messages.transform`, compaction prompt replacement) — a deliberate two-tier stability promise: interception at operation boundaries is stable, rewriting the context itself is not. No cache-stability discipline comparable to Pi's was found.
|
||||
|
||||
---
|
||||
|
||||
## 3. Adopt / Adapt / Avoid for Hermes
|
||||
|
||||
Grading against #64182's ground rules. "Validated" = the proposal already on the Hermes issue is independently confirmed by field evidence.
|
||||
|
||||
| # | Lesson | Verdict | Maps to | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| 1 | **Typed per-hook result vocabularies, not veto-by-throw.** Pi's `{block, reason}` / `{cancel}` / `"handled"` enums vs OpenCode's throw-idiom (bug ≡ policy denial) and its dead `permission.ask`. Approval/gate hooks need enumerated results dispatched from the policy engine itself. | **Adopt** | #64162 | Pi runner.ts:759-1188; oc plugins.mdx:247-257, oc#7006 |
|
||||
| 2 | **Guard hooks fail closed; observers fail open.** Pi contains every handler error except `tool_call`, whose crash blocks the tool with an LLM-visible error result. Independently confirms ground rule 4 — and refines it: the failure mode of a *crashed* security hook must also be closed, not just its config default. | **Adopt** (validated) | #64162, #64204 | Pi agent-session.ts:454-467, agent-loop.ts:657-665 |
|
||||
| 3 | **Hook wire-up drift is the killer bug class: CI-check that every declared hook has a live dispatch site.** OpenCode's `permission.ask` sat typed-but-dead for 6+ months after a subsystem rewrite. Hermes already stores unknown hook names for forward compat (`register_hook`, plugins.py ~L1158) — the same drift is possible. A `VALID_HOOKS` ↔ `invoke_hook(` cross-check is a one-file test. | **Adopt now** (cheap, standalone) | #64230 (Doctor), CI | oc index.ts:261 + absent trigger site, oc#7006 |
|
||||
| 4 | **Deadline budgets on plugin callbacks — be the first framework to have them.** Neither system times out runtime hooks; both shipped hang-class failures (pi#5687/#5115; oc#7741, "Transform timeouts" deferred twice). OpenCode's own TUI dispose path (hard 5s, per-cleanup timer race) proves the mechanism is practical. Observer hooks: enforce a budget and log-and-drop. Mutating/guard hooks: budget + fail per lesson 2. | **Adopt** (differentiator) | #64161, #64164, #64229 | Pi grep: zero timeout logic; oc PLAN.md:507-510, runtime.ts:122-226 |
|
||||
| 5 | **Per-delta streaming hooks are viable only if non-blocking is structural, not documentary.** The controlled experiment: Pi offers per-delta and awaits inline → slow observer throttles the visible stream, hangs freeze it; OpenCode offers nothing per-delta → hot path safe, TTS use case unserved. #64161's "never-block contract + buffered-queue helper" is the right middle — but make the bounded queue the *only* consumption path (drop/coalesce policy included, cf. OpenCode's `SubscriberOverflowError` dropping queue), not an optional convenience next to a raw sync callback. | **Adapt** | #64161 | Pi agent-session.ts:728-734; oc processor.ts:512-524, event.ts:152-164 |
|
||||
| 6 | **The namespaced bus proposal is ahead of both systems — proceed, with their two omissions fixed.** Pi's bus works but has arbitrary un-namespaced string channels and no discoverability; OpenCode has no plugin-emit at all. #64164's `<plugin_key>:` enforcement, reserved `hermes:` prefix, advisory declarations, recursion cap, and deterministic subscription order have no counterexample in the field. Carry over per-callback isolation (both systems do this right) and add lesson-4 budgets. Fire-and-forget with return-values-ignored matches both systems' stable practice. | **Adopt own design** (validated) | #64164 | Pi event-bus.ts (whole file); oc plugin/index.ts:251-258 |
|
||||
| 7 | **Load order as the only priority system; explicit per-registry collision policies.** Zero configuration, deterministic, and no field demand for priorities in either ecosystem. Document Hermes's ordering as a spec the way OpenCode's PLAN does; pick a collision rule per registry (Pi: first-wins tools / suffixed commands / reserved denylist) instead of building dependency resolution. | **Adopt** | #64164, #64229 | Pi loader.ts:660-708, runner.ts:421-604; oc PLAN.md:144-146 |
|
||||
| 8 | **Host-enforced compat gate + written deprecation window; migration tooling over semver ceremony.** OpenCode's `engines` gate is the right shape but plugin-opt-in only, and its patch-release API removal (oc#26557) shows lockstep versioning without policy is social, not mechanical. Pi shows the complement: loud breaking changes + automatic migrations + alias shims *before* removal. Manifest v2 should carry a host-checked `api_version` range; the repo should carry a one-paragraph deprecation policy. | **Adapt** | #64165, #64179 | oc shared.ts:194-205, tui.ts:87-120, oc#26557; Pi CHANGELOG:243, 3530-3620 |
|
||||
| 9 | **Scoped registrations with auto-tracked disposal; poison stale contexts with teaching errors.** OpenCode's Proxy-tracked per-plugin scopes (clean live disable) and Pi's post-replacement context poisoning (silent race pi#2860 → loud self-documenting error) are the two halves of a robust lifecycle story — exactly what the #64229 ownership ledger needs. | **Adopt** | #64229 | oc runtime.ts:143-160; Pi runner.ts:514-527, docs:1223-1265 |
|
||||
| 10 | **Prompt-cache stability as API contract is real and Pi proves it's implementable.** Structured prompt inputs instead of final strings, `structuredClone` for ephemeral transforms, additive-only tool activation with provider deferred loading, documented second-order invalidation warnings. Strongest possible validation of ground rule 2, with a concrete reference implementation for cache-safe injection. | **Adopt** (validated) | #64167 | Pi docs:2254-2290, CHANGELOG 0.80.6/pi#6474 |
|
||||
| 11 | **Half-sandboxes: both systems refuse, for the same stated reason.** Pi documents that a partial in-process sandbox "would be easy to misunderstand as a security boundary"; OpenCode runs plugins fully privileged with path-containment only. Viable while plugin authors ≈ users; Hermes's Skills-Hub-style trust/scan pipeline is the nearer-term marketplace answer than in-process isolation. | **Adapt with eyes open** | security posture | Pi docs/security.md:5-37; oc shared.ts:89-97 |
|
||||
| 12 | **Events to plugins before UI and before persistence; boot-stage the plugin-facing client.** Pi's explicit ordering guarantee removes a whole class of races; OpenCode's plugins-as-API-clients design is elegant but deadlocked startup when a plugin called the API mid-init (oc#7741) — if ctx ever grows client-like powers, stage them ("unavailable until ready"). | **Adapt** | #64178, #64229 | Pi agent-session.ts:596-601; oc plugin/index.ts:142-147, oc#7741 |
|
||||
| 13 | **Neither system has ADRs — and both paid for it.** Pi's rationale is scattered across changelog/blog/footguns; OpenCode broke APIs in patch releases partly because no decision record said not to. Hermes's per-sub-issue design sketches (#64182 style) are already ahead of both; add an explicit **Deferred Decisions** section per design (OpenCode's PLAN.md's best feature) so open questions stay visible instead of silently unresolved. | **Keep + adapt** | process | verified ADR absence in both repos |
|
||||
|
||||
## 4. Verified absences (findings, not gaps in the spike)
|
||||
|
||||
- **No ADRs in either repo** (repo-wide searches for `adr`/`decision` artifacts). The Discord recollection of "ADRs describing failure modes" is unsubstantiated for both; the nearest equivalents are Pi's docs footgun sections and OpenCode's single v2 PLAN.md.
|
||||
- **No runtime hook timeouts in either system** (grep-verified in Pi's `src/core/extensions/`; OpenCode's own plan defers them).
|
||||
- **No written deprecation or plugin-API stability policy in either repo.**
|
||||
- Caveats: Pi's #2860→remediation causality is inferred from matching failure/fix, not a maintainer statement; OpenCode pre-rewrite history was not diffed (shallow clones); maintainer replies inside cited issues were not visible in fetched content.
|
||||
|
||||
---
|
||||
|
||||
*Spike time-boxed per #64180. Primary sources: `earendil-works/pi` @ `eb79351` — `packages/coding-agent/src/core/extensions/{runner,loader,types}.ts`, `src/core/{agent-session,event-bus}.ts`, `packages/agent/src/agent-loop.ts`, `docs/{extensions,security}.md`, `CHANGELOG.md`; `anomalyco/opencode` @ `c69abee` — `packages/plugin/src/{index,tui}.ts`, `packages/plugin/src/v2/effect/PLAN.md`, `packages/opencode/src/plugin/{index,shared,loader}.ts`, `packages/opencode/src/plugin/tui/runtime.ts`, `packages/core/src/event.ts`, `packages/web/src/content/docs/plugins.mdx`; issues pi#2860/#2715/#5080/#5687, oc#7006/#26557/#7741/#4850/#12222; mariozechner.at posts (2025-11-02, 2025-11-30).*
|
||||
@@ -0,0 +1,144 @@
|
||||
# Plugin Config & State Bridge
|
||||
|
||||
**Status:** config + state slice implemented by #64227
|
||||
|
||||
**Original design:** Topher Ross (@thebizfixer), RFC PR #58542
|
||||
|
||||
**Concrete consumer:** kanban-advanced
|
||||
|
||||
## Scope
|
||||
|
||||
This slice adds two native `PluginContext` capabilities:
|
||||
|
||||
- typed, namespace-jailed settings via `ctx.get_config()` and `ctx.set_config()`;
|
||||
- atomic, profile-scoped runtime data via `ctx.state`.
|
||||
|
||||
Config schema registration, config defaults, and the cron facade from the
|
||||
original RFC remain separate follow-up work. No core model tool is added.
|
||||
|
||||
## Config API
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
endpoint = ctx.get_config("api_url", default="https://example.invalid")
|
||||
retries = ctx.get_config("retry.attempts", default=3)
|
||||
|
||||
ctx.set_config("api_url", "https://api.example.com")
|
||||
ctx.set_config("retry.attempts", 5)
|
||||
```
|
||||
|
||||
Keys are **relative to the calling plugin**. The example above reads and writes:
|
||||
|
||||
```yaml
|
||||
plugins:
|
||||
entries:
|
||||
<effective-plugin-id>:
|
||||
settings:
|
||||
api_url: https://api.example.com
|
||||
retry:
|
||||
attempts: 5
|
||||
```
|
||||
|
||||
`<effective-plugin-id>` is `manifest.key` when present, otherwise
|
||||
`manifest.name`. `settings` is the canonical namespace chosen after the issue
|
||||
discussion in #64227/#67531. For migration safety, reads fall back to the former
|
||||
`plugins.entries.<id>.config.*` subtree only when the canonical value is absent.
|
||||
Writes always target `settings`; they do not rewrite or delete legacy values.
|
||||
|
||||
### Namespace jail
|
||||
|
||||
The API does not accept full config paths. A plugin can never use it to inspect
|
||||
or change arbitrary Hermes configuration.
|
||||
|
||||
Accepted:
|
||||
|
||||
```python
|
||||
ctx.get_config("endpoint")
|
||||
ctx.set_config("retry.policy", {"attempts": 3})
|
||||
```
|
||||
|
||||
Rejected with `ValueError` and a warning log:
|
||||
|
||||
```python
|
||||
ctx.get_config("security.approval_mode")
|
||||
ctx.set_config("model.provider", "attacker-proxy")
|
||||
ctx.set_config("plugins.entries.other.settings.token", "...")
|
||||
ctx.set_config("../../security.approval_mode", "always_allow")
|
||||
ctx.set_config(r"..\..\model.provider", "attacker-proxy")
|
||||
```
|
||||
|
||||
There is no global read allowlist: `ctx.profile_name` already exposes the only
|
||||
small host fact requested by the RFC. Settings writes use Hermes'
|
||||
profile-aware config loader/saver and atomic YAML replacement. The bridge
|
||||
validates the existing YAML before writing so malformed config is never
|
||||
silently replaced. Every operation resolves the active context-local
|
||||
`HERMES_HOME`, so one globally loaded plugin context follows multiplexed
|
||||
profile turns without crossing profile data.
|
||||
|
||||
## Durable state API
|
||||
|
||||
Use state for plugin-owned runtime data such as cursors, dedupe sets, and
|
||||
caches. Do not put those values in user-owned config.
|
||||
|
||||
```python
|
||||
def register(ctx):
|
||||
cursor = ctx.state.get("cursor", default={"page": 0})
|
||||
ctx.state.set("cursor", {"page": cursor["page"] + 1})
|
||||
```
|
||||
|
||||
The facade stores one JSON object at:
|
||||
|
||||
```text
|
||||
<HERMES_HOME>/plugin-data/<plugin-data-namespace>/state.json
|
||||
```
|
||||
|
||||
Portable Agent Plugins use their existing `PLUGIN_DATA` namespace exactly.
|
||||
Native and nested plugin ids use the same collision-resistant, Windows-safe
|
||||
namespace algorithm. `ctx.state.data_dir` exposes the directory and
|
||||
`ctx.state.path` exposes the JSON file when a plugin needs to inspect its own
|
||||
location.
|
||||
|
||||
### State guarantees
|
||||
|
||||
- **Profile isolation:** the data root resolves from the active context-local
|
||||
Hermes home on every operation.
|
||||
- **Atomic replacement:** state writes use temp-file + `fsync` + `os.replace`.
|
||||
- **Concurrent updates:** a sibling lock file serializes read-modify-write across
|
||||
threads and processes (`fcntl` on POSIX, `msvcrt` on Windows).
|
||||
- **Quota:** the complete serialized state is limited to 10 MiB per plugin. A
|
||||
rejected update leaves the previous file untouched.
|
||||
- **Fail closed:** malformed/non-object JSON is reported and never overwritten.
|
||||
- **Typed values:** values must be JSON-serializable.
|
||||
|
||||
State keys are 1–128 characters and may contain letters, numbers, `_`, `-`,
|
||||
`.`, or `:`. Path separators and `..` are rejected.
|
||||
|
||||
## State vs. config
|
||||
|
||||
| Data | API | Ownership | Example |
|
||||
|---|---|---|---|
|
||||
| User-visible behavior | `ctx.get_config` / `ctx.set_config` | User/plugin settings in `config.yaml` | endpoint, timeout, feature mode |
|
||||
| Runtime bookkeeping | `ctx.state.get` / `ctx.state.set` | Plugin data under `plugin-data/` | cursor, cache, dedupe ids |
|
||||
|
||||
Both APIs are additive. Existing plugins that perform their own file I/O keep
|
||||
working, but new plugins should use this bridge for stable profile and Windows
|
||||
semantics.
|
||||
|
||||
## Verification contract
|
||||
|
||||
The implementation is covered with real temporary-Hermes-home tests for:
|
||||
|
||||
- fixture-plugin discovery and config/state round trips;
|
||||
- canonical `settings` writes and legacy `config` read fallback;
|
||||
- direct global, cross-plugin, POSIX traversal, and Windows traversal rejection;
|
||||
- concurrent settings writes without lost siblings;
|
||||
- cross-thread and cross-process state updates;
|
||||
- atomic quota rejection and malformed-state/config preservation;
|
||||
- two-profile isolation after the ambient profile changes;
|
||||
- Unicode and Windows-style path values.
|
||||
|
||||
## Related
|
||||
|
||||
- [Issue #64227](https://github.com/NousResearch/hermes-agent/issues/64227)
|
||||
- [RFC PR #58542](https://github.com/NousResearch/hermes-agent/pull/58542) by Topher Ross
|
||||
- #67531 — standalone plugin settings namespace discussion
|
||||
@@ -0,0 +1,195 @@
|
||||
# Network Egress Isolation for Docker Deployments
|
||||
|
||||
When running Hermes inside Docker, the default `network_mode: host` gives the
|
||||
agent process unrestricted outbound network access. This guide shows how to
|
||||
segment traffic so the agent core can only reach the services it needs, while
|
||||
blocking arbitrary outbound connections.
|
||||
|
||||
This is primarily a defense against prompt injection attacks that attempt to
|
||||
exfiltrate data via `curl`, `wget`, or raw HTTP from tool-generated shell
|
||||
commands.
|
||||
|
||||
## Threat Model
|
||||
|
||||
The Hermes [SECURITY.md](../../SECURITY.md) §2 defines the trust model. The
|
||||
terminal backend is the primary execution boundary. However, when running with
|
||||
`network_mode: host`, any command the agent executes can reach any endpoint on
|
||||
the network, including external ones.
|
||||
|
||||
Network egress isolation adds a second layer: even if a malicious command
|
||||
executes inside the container, it cannot reach endpoints outside the
|
||||
explicitly allowlisted set.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ Docker Network: internal (no internet) │
|
||||
│ │
|
||||
│ ┌──────────────┐ ┌──────────────────┐ │
|
||||
│ │ hermes-agent │ │ hermes-dashboard │ │
|
||||
│ └──────┬───────┘ └────────┬─────────┘ │
|
||||
│ │ │ │
|
||||
│ ▼ │ │
|
||||
│ ┌──────────────┐ │ │
|
||||
│ │ hermes-gtw │◄───────────┘ │
|
||||
│ └──────┬───────┘ │
|
||||
│ │ │
|
||||
└──────────┼───────────────────────────────────┘
|
||||
│
|
||||
┌──────────┼───────────────────────────────────┐
|
||||
│ Docker Network: egress (internet-capable) │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ ┌─────────────────┐ │
|
||||
│ │ egress-proxy │──► allowlisted hosts │
|
||||
│ │ (squid / envoy) │ │
|
||||
│ └─────────────────┘ │
|
||||
└──────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
Two Docker networks:
|
||||
|
||||
- **`internal`** — no default route, no internet access. The agent, dashboard,
|
||||
and gateway run here.
|
||||
- **`egress`** — has internet access. Only services that need to reach external
|
||||
APIs are attached to this network.
|
||||
|
||||
The gateway service is dual-homed (attached to both networks) so it can
|
||||
receive inbound messages from Telegram/Slack/etc. and forward them to the
|
||||
agent on the internal network.
|
||||
|
||||
## Compose Configuration
|
||||
|
||||
Override the default `docker-compose.yml` with a
|
||||
`docker-compose.override.yml`:
|
||||
|
||||
```yaml
|
||||
# docker-compose.override.yml
|
||||
# Network egress isolation for production deployments.
|
||||
#
|
||||
# Usage:
|
||||
# HERMES_UID=$(id -u) HERMES_GID=$(id -g) docker compose up -d
|
||||
#
|
||||
# This overrides network_mode: host with isolated Docker networks.
|
||||
|
||||
networks:
|
||||
internal:
|
||||
driver: bridge
|
||||
internal: true # no default route, no internet
|
||||
egress:
|
||||
driver: bridge
|
||||
|
||||
services:
|
||||
gateway:
|
||||
network_mode: "" # clear the host-mode default
|
||||
networks:
|
||||
- internal
|
||||
- egress # needs outbound for Telegram, LLM APIs
|
||||
ports:
|
||||
- "127.0.0.1:9119:9119" # dashboard proxy, localhost only
|
||||
|
||||
dashboard:
|
||||
network_mode: ""
|
||||
networks:
|
||||
- internal # internal only, no egress needed
|
||||
```
|
||||
|
||||
### With an Egress Proxy (Recommended)
|
||||
|
||||
For tighter control, route all outbound traffic through an HTTP proxy with
|
||||
an explicit allowlist:
|
||||
|
||||
```yaml
|
||||
# docker-compose.override.yml (with egress proxy)
|
||||
|
||||
networks:
|
||||
internal:
|
||||
driver: bridge
|
||||
internal: true
|
||||
egress:
|
||||
driver: bridge
|
||||
|
||||
services:
|
||||
gateway:
|
||||
network_mode: ""
|
||||
networks:
|
||||
- internal
|
||||
- egress
|
||||
environment:
|
||||
- HTTP_PROXY=http://egress-proxy:3128
|
||||
- HTTPS_PROXY=http://egress-proxy:3128
|
||||
- NO_PROXY=hermes,hermes-dashboard,localhost
|
||||
|
||||
dashboard:
|
||||
network_mode: ""
|
||||
networks:
|
||||
- internal
|
||||
|
||||
egress-proxy:
|
||||
image: ubuntu/squid:6.10-24.04_edge
|
||||
networks:
|
||||
- egress
|
||||
volumes:
|
||||
- ./config/squid-allowlist.conf:/etc/squid/conf.d/allowlist.conf:ro
|
||||
restart: unless-stopped
|
||||
```
|
||||
|
||||
Example `config/squid-allowlist.conf`:
|
||||
|
||||
```
|
||||
# Only allow HTTPS CONNECT to these hosts
|
||||
acl allowed_hosts dstdomain api.openai.com
|
||||
acl allowed_hosts dstdomain api.anthropic.com
|
||||
acl allowed_hosts dstdomain openrouter.ai
|
||||
acl allowed_hosts dstdomain generativelanguage.googleapis.com
|
||||
acl allowed_hosts dstdomain api.telegram.org
|
||||
acl allowed_hosts dstdomain api.github.com
|
||||
acl allowed_hosts dstdomain discord.com
|
||||
|
||||
http_access allow CONNECT allowed_hosts
|
||||
http_access deny all
|
||||
```
|
||||
|
||||
Adjust the allowlist to match your LLM provider and messaging platform.
|
||||
|
||||
## Validating the Setup
|
||||
|
||||
After bringing up the stack, verify isolation:
|
||||
|
||||
```bash
|
||||
# From the agent container: this should FAIL (no egress)
|
||||
docker compose exec gateway \
|
||||
curl -sf --max-time 5 https://example.com && echo "FAIL: egress not blocked" || echo "OK: egress blocked"
|
||||
|
||||
# From the agent container: this should SUCCEED (internal network)
|
||||
docker compose exec gateway \
|
||||
curl -sf --max-time 5 http://hermes-dashboard:9119/health && echo "OK: internal reachable" || echo "FAIL"
|
||||
|
||||
# If using egress proxy: this should SUCCEED (allowlisted)
|
||||
docker compose exec gateway \
|
||||
curl -sf --max-time 5 --proxy http://egress-proxy:3128 https://api.openai.com/v1/models && echo "OK" || echo "FAIL"
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- **DNS resolution:** The `internal` network can still resolve external DNS
|
||||
names unless you also run a local DNS resolver that blocks external queries.
|
||||
For most threat models this is acceptable since DNS resolution alone does not
|
||||
exfiltrate meaningful data.
|
||||
|
||||
- **Not a substitute for sandbox backends:** This guide isolates the agent
|
||||
*container's* network. If you use the default local terminal backend, tool
|
||||
commands execute inside the same container. For stronger isolation, combine
|
||||
network segmentation with a sandboxed terminal backend (Docker, Modal,
|
||||
Daytona).
|
||||
|
||||
- **Platform adapters need egress:** The gateway service needs outbound access
|
||||
to reach messaging platform APIs. If you add new platform adapters, add their
|
||||
API endpoints to the proxy allowlist.
|
||||
|
||||
## Related
|
||||
|
||||
- [SECURITY.md](../../SECURITY.md) — Hermes trust model and vulnerability reporting
|
||||
- [Terminal backends](../../README.md) — sandboxed execution targets
|
||||
- [docker-compose.yml](../../docker-compose.yml) — default compose configuration
|
||||
@@ -0,0 +1,614 @@
|
||||
# Session Lifecycle
|
||||
|
||||
> **Audience:** Gateway developers and maintainers
|
||||
> **Source files:** `gateway/session.py` (~1200 lines + `session_*.py` siblings), `gateway/run.py` (~5500 lines facade + `run_*.py` phases), `gateway/config.py`
|
||||
> **Last updated:** 2026-06-16
|
||||
|
||||
## Overview
|
||||
|
||||
A **session** represents a continuous conversation between the agent and one or more users on a
|
||||
messaging platform. The session lifecycle governs when conversations persist, when they reset,
|
||||
how they survive gateway restarts, and how messages queue during concurrent operations.
|
||||
|
||||
The session system lives primarily in two modules:
|
||||
|
||||
- `gateway/session.py` — Data model (`SessionSource`, `SessionEntry`, `SessionContext`),
|
||||
key generation (`build_session_key`), and the main store (`SessionStore`).
|
||||
- `gateway/run.py` — Gateway runner (`GatewayRunner`) facade that wires sessions into the message
|
||||
processing pipeline; the phases live in `run_*.py` siblings: session housekeeping
|
||||
(`run_watchers.py`), agent caching (`run_agent_cache.py`), restart recovery
|
||||
(`session_recovery.py`), and message queuing (`run_busy.py`).
|
||||
|
||||
---
|
||||
|
||||
## 1. SessionSource — Message Origin Descriptor
|
||||
|
||||
`SessionSource` is a frozen record of *where a message came from*. It is attached to every
|
||||
incoming `MessageEvent` and used for routing, isolation, and context injection.
|
||||
|
||||
### Fields
|
||||
|
||||
| Field | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `platform` | `Platform` | *(required)* | Enum identifying the messaging platform (telegram, discord, slack, signal, whatsapp, matrix, local, etc.). |
|
||||
| `chat_id` | `str` | *(required)* | Platform-level chat/group/channel identifier. Routed through the adapter's `chat_id_key` transform. |
|
||||
| `chat_name` | `Optional[str]` | `None` | Human-readable name of the chat or group. |
|
||||
| `chat_type` | `str` | `"dm"` | One of `"dm"`, `"group"`, `"channel"`, `"thread"`. Controls session key generation and isolation. |
|
||||
| `user_id` | `Optional[str]` | `None` | Platform-specific user identifier. Used for authorization and per-user session isolation. |
|
||||
| `user_name` | `Optional[str]` | `None` | Display name of the message author. Injected into system prompt. |
|
||||
| `thread_id` | `Optional[str]` | `None` | Forum topic / Discord thread / Slack thread identifier. Differentiates threaded conversations. |
|
||||
| `chat_topic` | `Optional[str]` | `None` | Channel topic or description (Discord channel topic, Slack channel purpose). |
|
||||
| `user_id_alt` | `Optional[str]` | `None` | Platform-specific stable alternative ID (Signal UUID, Feishu union_id). Used when `user_id` is ephemeral. |
|
||||
| `chat_id_alt` | `Optional[str]` | `None` | Signal group internal ID — maps a Signal group V2 identifier to its canonical form. |
|
||||
| `is_bot` | `bool` | `False` | True when the message author is a bot or webhook (Discord bots). |
|
||||
| `guild_id` | `Optional[str]` | `None` | Discord guild / Slack workspace / Matrix server scope identifier. |
|
||||
| `parent_chat_id` | `Optional[str]` | `None` | Parent channel when `chat_id` refers to a thread. |
|
||||
| `message_id` | `Optional[str]` | `None` | ID of the triggering message. Used for pin/reply/react operations and Discord ID injection. |
|
||||
| `role_authorized` | `bool` | `False` | True when adapter granted access via a platform role (not individual user ID). |
|
||||
|
||||
### Key Methods
|
||||
|
||||
- **`description`** (property: `str`) — Human-readable summary e.g. `"DM with Alice"`,
|
||||
`"group: My Group, thread: 12345"`.
|
||||
- **`to_dict()` / `from_dict()`** — Serialization round-trip for persistence in `sessions.json`.
|
||||
|
||||
---
|
||||
|
||||
## 2. SessionEntry — Active Session Record
|
||||
|
||||
`SessionEntry` is the per-session metadata record stored in memory and persisted to
|
||||
`{sessions_dir}/sessions.json`. Each entry maps a `session_key` to its current `session_id`.
|
||||
|
||||
### Fields
|
||||
|
||||
| Field | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `session_key` | `str` | *(required)* | Deterministic key identifying the conversation lane (see §4). |
|
||||
| `session_id` | `str` | *(required)* | Unique identifier for this specific conversation incarnation. Format: `YYYYMMDD_HHMMSS_<8hex>`. |
|
||||
| `created_at` | `datetime` | *(required)* | When this session incarnation was created. |
|
||||
| `updated_at` | `datetime` | *(required)* | Last activity timestamp used for resource housekeeping. |
|
||||
| `origin` | `Optional[SessionSource]` | `None` | The source that created this session, used for delivery routing. |
|
||||
| `display_name` | `Optional[str]` | `None` | Chat display name (sourced from `SessionSource.chat_name`). |
|
||||
| `platform` | `Optional[Platform]` | `None` | Platform enum persisted for routing across restarts. |
|
||||
| `chat_type` | `str` | `"dm"` | Chat type, also persisted for policy lookup. |
|
||||
| `input_tokens` | `int` | `0` | Cumulative LLM input (prompt) tokens consumed. |
|
||||
| `output_tokens` | `int` | `0` | Cumulative LLM output (completion) tokens consumed. |
|
||||
| `cache_read_tokens` | `int` | `0` | Cumulative prompt cache read tokens. |
|
||||
| `cache_write_tokens` | `int` | `0` | Cumulative prompt cache write tokens. |
|
||||
| `total_tokens` | `int` | `0` | Total token count across all turns. |
|
||||
| `estimated_cost_usd` | `float` | `0.0` | Estimated cumulative USD cost. |
|
||||
| `cost_status` | `str` | `"unknown"` | Cost tracking status label. |
|
||||
| `last_prompt_tokens` | `int` | `0` | Last API-reported prompt token count. Used for accurate compression pre-check. |
|
||||
|
||||
### Boolean Flags (State Machine)
|
||||
|
||||
SessionEntry has several boolean flags that form a simple state machine governing session
|
||||
behavior on the next access.
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `was_auto_reset` | `bool` | `False` | Set when explicit suspension causes a replacement session. Also retained for historical records. |
|
||||
| `auto_reset_reason` | `Optional[str]` | `None` | `"suspended"` for explicit suspension; older rows may retain historical reset reasons. |
|
||||
| `reset_had_activity` | `bool` | `False` | Whether the replaced session had prior activity. |
|
||||
| `is_fresh_reset` | `bool` | `False` | Set by explicit `/new` or `/reset`. Triggers topic/channel skill re-injection on first message. Distinguished from `was_auto_reset` to avoid misleading "session expired" notices. |
|
||||
| `expiry_finalized` | `bool` | `False` | Historical finalization fence retained for recovery; no timer writes it. |
|
||||
| `suspended` | `bool` | `False` | Hard force-wipe signal. Set by `/stop` or stuck-loop escalation (3+ consecutive restart failures). On next `get_or_create_session()`, forces a new `session_id` regardless of `resume_pending`. |
|
||||
| `resume_pending` | `bool` | `False` | Soft recovery marker. Set by `suspend_recently_active()` (crash recovery) or drain timeout. On next access, preserves the existing `session_id` — the user continues on the same transcript. Cleared after the next successful turn completes. |
|
||||
| `resume_reason` | `Optional[str]` | `None` | Why resume was marked: `"restart_timeout"`, `"shutdown_timeout"`, `"restart_interrupted"`. |
|
||||
| `last_resume_marked_at` | `Optional[datetime]` | `None` | Timestamp of the last resume-pending marking. |
|
||||
|
||||
### State Transition Logic (get_or_create_session)
|
||||
|
||||
```
|
||||
┌──────────┐
|
||||
│ Incoming │
|
||||
│ Message │
|
||||
└────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ session_key exists │──── No ──► Create fresh SessionEntry
|
||||
│ AND !force_new │
|
||||
└──────────┬───────────┘
|
||||
│ Yes
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ entry.suspended? │──── Yes ──► Auto-reset: new session_id
|
||||
└──────────┬───────────┘ (reason="suspended")
|
||||
│ No
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ entry.resume_pending?│──── Yes ──► Return existing entry
|
||||
└──────────┬───────────┘ (preserve session_id)
|
||||
│ No Clear flag on next successful turn
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ Policy says reset? │──── Yes ──► Auto-reset: new session_id
|
||||
└──────────┬───────────┘ (reason="idle"/"daily")
|
||||
│ No
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ Return existing │
|
||||
│ entry, bump │
|
||||
│ updated_at │
|
||||
└──────────────────────┘
|
||||
```
|
||||
|
||||
**Priority order in `get_or_create_session()`:**
|
||||
1. `suspended=True` → always force-reset (hard wipe)
|
||||
2. `resume_pending=True` → preserve session_id (soft recovery)
|
||||
3. No trigger → return existing entry (bump `updated_at`)
|
||||
|
||||
---
|
||||
|
||||
## 3. SessionStore — Storage and Operations
|
||||
|
||||
`SessionStore` is the main storage layer. It maintains an in-memory dict (`_entries`) persisted
|
||||
to `sessions.json`, with SQLite (`SessionDB`) as the canonical store for session metadata and
|
||||
message transcripts.
|
||||
|
||||
### Constructor
|
||||
|
||||
```python
|
||||
SessionStore(sessions_dir: Path, config: GatewayConfig, has_active_processes_fn=None)
|
||||
```
|
||||
|
||||
- `sessions_dir` — Directory where `sessions.json` lives.
|
||||
- `config` — `GatewayConfig` instance for routing and housekeeping settings.
|
||||
- `has_active_processes_fn` — Optional callback keyed by `session_key` to check for running
|
||||
background processes. Sessions with active processes are protected from routing-entry pruning.
|
||||
|
||||
### Operations (Methods)
|
||||
|
||||
| Method | Description |
|
||||
|---|---|
|
||||
| `get_or_create_session(source, force_new=False)` | Core entry point. Returns existing or creates new `SessionEntry`. Evaluates explicit suspension and restart recovery state. Creates/ends SQLite records. |
|
||||
| `update_session(session_key, last_prompt_tokens=None)` | Lightweight metadata update after an interaction. Bumps `updated_at`, optionally records `last_prompt_tokens`. |
|
||||
| `reset_session(session_key, display_name=None)` | Explicit reset (from `/new` or `/reset`). Creates new `session_id`, sets `is_fresh_reset=True`. Ends old SQLite session, creates new one. |
|
||||
| `switch_session(session_key, target_session_id)` | Switch to a different existing session ID (from `/resume`). Ends current SQLite session, reopens target. |
|
||||
| `suspend_session(session_key)` | Mark session as `suspended=True` (from `/stop`). Forces auto-reset on next access. |
|
||||
| `mark_resume_pending(session_key, reason)` | Mark session as `resume_pending=True` (from drain timeout). Preserves session_id on next access. Will NOT override `suspended=True`. |
|
||||
| `clear_resume_pending(session_key)` | Clear `resume_pending` after a successful resumed turn. Called from gateway after `run_conversation()` returns. |
|
||||
| `suspend_recently_active(max_age_seconds=120)` | Crash recovery: mark recently-active sessions as `resume_pending=True`. Skips already-pending and already-suspended entries. Called on startup after unclean shutdown. |
|
||||
| `prune_old_entries(max_age_days)` | Drop entries older than `max_age_days` (based on `updated_at`). Skips `suspended` entries and sessions with active processes. |
|
||||
| `list_sessions(active_minutes=None)` | Return all sessions, optionally filtered by recent activity. Sorted by `updated_at` descending. |
|
||||
| `lookup_by_session_id(session_id)` | Find the active `SessionEntry` for a persisted session ID. |
|
||||
| `has_any_sessions()` | Check if any sessions have ever been created (uses SQLite for history, not just in-memory dict). |
|
||||
| `append_to_transcript(session_id, message, skip_db=False)` | Append a message to SQLite transcript. `skip_db=True` prevents duplicate writes when the agent already persisted. |
|
||||
| `rewrite_transcript(session_id, messages)` | Full replacement of session transcript (used by `/retry`, `/undo`, `/compress`). |
|
||||
| `load_transcript(session_id)` | Load all messages from a session's SQLite transcript. |
|
||||
| `rewind_session(session_id, n=1)` | Back up `n` user turns via soft-delete (keeps audit trail). Returns `{rewound_count, turns_undone, target_text}`. |
|
||||
|
||||
### Internal Helpers
|
||||
|
||||
- `_ensure_loaded()` / `_ensure_loaded_locked()` — Load `sessions.json` into `_entries` dict.
|
||||
- `_save()` — Atomic write to `sessions.json` via temp file + `atomic_replace`.
|
||||
- `_generate_session_key(source)` — Delegates to `build_session_key()` with config params.
|
||||
|
||||
### Storage Layout
|
||||
|
||||
```
|
||||
{sessions_dir}/
|
||||
sessions.json # In-memory _entries dict, persisted as JSON
|
||||
Maps session_key → SessionEntry (metadata only)
|
||||
{session_id}.jsonl # (Legacy, removed in spec 002)
|
||||
```
|
||||
|
||||
The canonical transcript store is SQLite via `SessionDB` (from `hermes_state`). The
|
||||
`sessions.json` file persists the `session_key → session_id` mapping and entry metadata
|
||||
(flags, timestamps, token counts). If SQLite is unavailable, the store falls back to
|
||||
JSONL, but this is a degradation path.
|
||||
|
||||
---
|
||||
|
||||
## 4. SessionKey Generation Rules
|
||||
|
||||
Session keys are deterministic strings that identify a conversation lane. They are generated
|
||||
by `build_session_key(source, group_sessions_per_user, thread_sessions_per_user)`.
|
||||
|
||||
### Key Format
|
||||
|
||||
```
|
||||
agent:main:{platform}:{chat_type}[:{chat_id}][:{thread_id}][:{participant_id}]
|
||||
```
|
||||
|
||||
### DM Rules
|
||||
|
||||
| Scenario | Key |
|
||||
|---|---|
|
||||
| DM with chat_id | `agent:main:telegram:dm:12345` |
|
||||
| DM with chat_id + thread | `agent:main:telegram:dm:12345:thread_678` |
|
||||
| DM without chat_id, with participant_id | `agent:main:signal:dm:user_abc` |
|
||||
| DM without chat_id or participant_id | `agent:main:telegram:dm` |
|
||||
| WhatsApp DM (canonicalized) | `agent:main:whatsapp:dm:{canonical_number}` |
|
||||
|
||||
- DMs always include `chat_id` when present, isolating each private conversation.
|
||||
- `thread_id` further differentiates threaded DMs within the same DM chat.
|
||||
- Without `chat_id`, falls back to `user_id_alt` or `user_id` as participant_id.
|
||||
- Without any identifier, all DMs on that platform collapse to one shared session.
|
||||
|
||||
### Group/Channel Rules
|
||||
|
||||
| Scenario | Key |
|
||||
|---|---|
|
||||
| Group chat | `agent:main:telegram:group:-10012345` |
|
||||
| Group chat, per-user isolation | `agent:main:telegram:group:-10012345:user_abc` |
|
||||
| Thread in group, shared | `agent:main:discord:group:12345:thread_678` |
|
||||
| Thread in group, per-user | `agent:main:discord:group:12345:thread_678:user_abc` |
|
||||
| Channel | `agent:main:slack:channel:C12345` |
|
||||
| WhatsApp group (canonicalized) | `agent:main:whatsapp:group:{canonical_id}:{participant}` |
|
||||
|
||||
- `chat_id` identifies the parent group/channel.
|
||||
- `thread_id` differentiates threads within that parent.
|
||||
- **Per-user isolation** (append `participant_id`) is controlled by:
|
||||
- `group_sessions_per_user` (default: `True`) — group/channel sessions are isolated.
|
||||
- `thread_sessions_per_user` (default: `False`) — threads are **shared** by default
|
||||
(Telegram forum topics, Discord threads, Slack threads all share one session per thread).
|
||||
- `participant_id` = `user_id_alt` or `user_id` (in that priority).
|
||||
- WhatsApp identifiers are canonicalized to handle JID/LID alias flips.
|
||||
|
||||
### Special Case: WhatApp
|
||||
|
||||
WhatsApp phone numbers go through `canonical_whatsapp_identifier()` which strips the
|
||||
`@s.whatsapp.net` suffix and normalizes to E.164 format. This prevents session fragmentation
|
||||
when the bridge returns different alias forms of the same phone number.
|
||||
|
||||
---
|
||||
|
||||
## 5. Multi-User Isolation Strategy
|
||||
|
||||
Multi-user isolation determines whether multiple users in the same chat share a conversation
|
||||
or each get their own private session.
|
||||
|
||||
### Decision Logic (`is_shared_multi_user_session`)
|
||||
|
||||
```python
|
||||
def is_shared_multi_user_session(source, *, group_sessions_per_user, thread_sessions_per_user):
|
||||
if source.chat_type == "dm":
|
||||
return False # DMs are always private
|
||||
if source.thread_id:
|
||||
return not thread_sessions_per_user # Threads: shared unless per-user
|
||||
return not group_sessions_per_user # Groups: isolated unless shared
|
||||
```
|
||||
|
||||
### Summary
|
||||
|
||||
| Chat Type | Default | Config Control |
|
||||
|---|---|---|
|
||||
| DM | Private (never shared) | N/A |
|
||||
| Group/Channel | Per-user isolation | `group_sessions_per_user` (default: True) |
|
||||
| Thread (forum, discord) | Shared (all participants see same context) | `thread_sessions_per_user` (default: False) |
|
||||
|
||||
### Impact on System Prompt
|
||||
|
||||
When `shared_multi_user_session=True`, the system prompt omits a fixed user name and instead
|
||||
states: *"Multi-user {thread|session} — messages are prefixed with [sender name]. Multiple
|
||||
users may participate."* Individual sender names are prefixed on each user message by the
|
||||
gateway at runtime, preserving prompt caching (the system prompt doesn't change per-turn).
|
||||
|
||||
---
|
||||
|
||||
## 6. Explicit Conversation Boundaries
|
||||
|
||||
Inactivity and wall-clock time never rotate a conversation. `/new` and `/reset`
|
||||
create an explicit boundary; context compression continues to manage long histories.
|
||||
Legacy timer configuration is ignored. The existing `SessionResetPolicy` datatype
|
||||
is inert compatibility data, not a runtime policy.
|
||||
|
||||
Explicit suspension still creates a boundary on the next inbound turn. Recovery
|
||||
respects explicit and historical finalized boundaries rather than reopening them.
|
||||
Resource-only eviction and WebSocket orphan reaping leave conversations resumable.
|
||||
|
||||
---
|
||||
|
||||
## 7. Restart Recovery Flow
|
||||
|
||||
The restart recovery system ensures that in-flight sessions are preserved across gateway
|
||||
restarts, crashes, and drain timeouts. It is the solution to issue #7536.
|
||||
|
||||
### Startup Recovery Sequence
|
||||
|
||||
```
|
||||
Gateway starts
|
||||
│
|
||||
▼
|
||||
┌───────────────────────────────┐
|
||||
│ Check for .clean_shutdown │── Exists? ──► Skip suspension (clean exit)
|
||||
│ marker │
|
||||
└───────────────────────────────┘
|
||||
│ Missing
|
||||
▼
|
||||
┌───────────────────────────────┐
|
||||
│ session_store │── Marks sessions updated within
|
||||
│ .suspend_recently_active() │ last 120 seconds as resume_pending
|
||||
└───────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌───────────────────────────────┐
|
||||
│ _suspend_stuck_loop_sessions()│── Suspends sessions that have been
|
||||
│ │ active across 3+ restarts
|
||||
└───────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌───────────────────────────────┐
|
||||
│ Queue inbound messages while │
|
||||
│ startup restore runs │
|
||||
│ (_startup_restore_in_progress)│
|
||||
└───────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌───────────────────────────────┐
|
||||
│ For each adapter, find │
|
||||
│ resume_pending sessions → │
|
||||
│ synthesize MessageEvent and │
|
||||
│ run _handle_message to let │
|
||||
│ the agent auto-continue │
|
||||
└───────────────────────────────┘
|
||||
```
|
||||
|
||||
### suspend_recently_active(max_age_seconds=120)
|
||||
|
||||
Called on gateway startup when no `.clean_shutdown` marker exists (indicating a crash or
|
||||
unexpected exit). For each session updated within the last 120 seconds:
|
||||
|
||||
- Sets `resume_pending=True`, `resume_reason="restart_interrupted"`,
|
||||
`last_resume_marked_at=now`.
|
||||
- Skips entries already `resume_pending=True` (no double-mark).
|
||||
- Skips entries explicitly `suspended=True` (hard wipe should stay).
|
||||
|
||||
### Stuck-Loop Detection (`_suspend_stuck_loop_sessions`)
|
||||
|
||||
Counts consecutive restarts via a JSON file (`{HERMES_HOME}/restart_counts.json`). If a
|
||||
session has been active across 3+ consecutive restarts, it's auto-suspended so the user
|
||||
gets a clean slate.
|
||||
|
||||
### Drain-Timeout Marking
|
||||
|
||||
On graceful shutdown/restart, the drain system calls `mark_resume_pending()` for any
|
||||
session that was mid-turn when the drain timeout fired. Reasons:
|
||||
|
||||
- `"restart_timeout"` — killed during restart drain
|
||||
- `"shutdown_timeout"` — killed during shutdown drain
|
||||
- `"restart_interrupted"` — crash recovery (from `suspend_recently_active`)
|
||||
|
||||
All three reasons are in `_AUTO_RESUME_REASONS` and eligible for startup auto-resume.
|
||||
|
||||
### Auto-Resume on Next Access
|
||||
|
||||
When `get_or_create_session()` encounters `resume_pending=True`:
|
||||
|
||||
1. It returns the existing entry **without** creating a new `session_id`.
|
||||
2. The existing transcript is loaded intact.
|
||||
3. The marking is not cleared here — it survives until the next successful turn
|
||||
completes (`clear_resume_pending()` is called from the gateway after
|
||||
`run_conversation()` returns a real response).
|
||||
4. If the resumed turn is interrupted again, the `resume_pending` flag remains set,
|
||||
and the next restart will retry. The stuck-loop counter handles terminal escalation
|
||||
(3 retries → suspended).
|
||||
|
||||
### Clean Shutdown Marker (`.clean_shutdown`)
|
||||
|
||||
Written at the end of a graceful shutdown. On next startup:
|
||||
|
||||
- If present: skip `suspend_recently_active()` entirely. Active agents were already
|
||||
drained, so no sessions are stuck.
|
||||
- Then delete the marker.
|
||||
|
||||
This prevents unwanted auto-resets after `hermes update`, `hermes gateway restart`,
|
||||
or `/restart`.
|
||||
|
||||
---
|
||||
|
||||
## 8. Message Queuing Flow
|
||||
|
||||
The message queuing system handles two scenarios:
|
||||
|
||||
1. **Interrupt follow-ups** — When a user sends multiple messages while the agent is
|
||||
processing, subsequent messages are queued as single-slot pending messages.
|
||||
2. **`/queue` FIFO** — Explicit `/queue` commands that must each produce their own full
|
||||
agent turn, in order, without merging.
|
||||
|
||||
### Data Structures
|
||||
|
||||
```
|
||||
adapter._pending_messages: Dict[session_key, MessageEvent]
|
||||
└── Single "next-up" slot per session. Overwritten on repeat sends
|
||||
(burst collapse). Shared with photo-burst follow-ups.
|
||||
|
||||
self._queued_events: Dict[session_key, List[MessageEvent]]
|
||||
└── Overflow buffer. Each /queue invocation appends here when the
|
||||
slot is occupied. Promoted one-at-a-time after each drain.
|
||||
```
|
||||
|
||||
### Enqueue (`_enqueue_fifo`)
|
||||
|
||||
```
|
||||
_enqueue_fifo(session_key, event, adapter)
|
||||
│
|
||||
▼
|
||||
┌───────────────────────────────────────┐
|
||||
│ Is slot free? │
|
||||
│ (session_key NOT in _pending_messages)│── Yes ──► Place event in slot
|
||||
└───────────────────────────────────────┘
|
||||
│ No
|
||||
▼
|
||||
Append to _queued_events[session_key] (overflow tail)
|
||||
```
|
||||
|
||||
### Dequeue / Promotion (`_promote_queued_event`)
|
||||
|
||||
Called at the drain site after the slot was consumed. If there's an overflow item:
|
||||
|
||||
- When `pending_event is None` (slot was empty), return overflow head as the new event.
|
||||
- When `pending_event` exists, stage overflow head in the slot for the next recursion.
|
||||
- If no adapter available, push back to `_queued_events` (don't silently drop).
|
||||
|
||||
### Queue Depth
|
||||
|
||||
`_queue_depth(session_key, adapter)` returns `len(overflow) + (1 if slot occupied else 0)`.
|
||||
|
||||
### Clearing
|
||||
|
||||
Queued events for a session are cleared on `/new` and `/reset` (via `_handle_reset_command`).
|
||||
|
||||
### FIFO Invariant
|
||||
|
||||
Each `/queue` invocation produces exactly one full agent turn, in FIFO order, with no
|
||||
merging. The single-slot `_pending_messages` + overflow `_queued_events` design ensures
|
||||
that repeated sends during an active turn don't cause out-of-order processing.
|
||||
|
||||
---
|
||||
|
||||
## 9. Session Context Injection
|
||||
|
||||
`SessionContext` is built from a `SessionSource` and `GatewayConfig` and injected into the
|
||||
agent's system prompt. It tells the agent:
|
||||
|
||||
- Where the current message came from
|
||||
- What platforms are connected
|
||||
- Where it can deliver scheduled task outputs
|
||||
- Whether this is a shared multi-user session
|
||||
|
||||
### Construction (`build_session_context`)
|
||||
|
||||
```python
|
||||
def build_session_context(source, config, session_entry=None) -> SessionContext
|
||||
```
|
||||
|
||||
1. Collects connected platforms from config.
|
||||
2. Collects home channels for each platform.
|
||||
3. Determines `shared_multi_user_session` via `is_shared_multi_user_session()`.
|
||||
4. Attaches session metadata (key, id, timestamps) if `session_entry` is provided.
|
||||
|
||||
### PII Redaction (`build_session_context_prompt`)
|
||||
|
||||
The dynamic system prompt section (`## Current Session Context`) can optionally redact
|
||||
personally identifiable information before sending to the LLM:
|
||||
|
||||
- User IDs → `user_<12hex>` (SHA-256 prefix)
|
||||
- Chat IDs → `<platform>:<12hex>` or just `<12hex>`
|
||||
- Platforms excluded from redaction: Discord (needs raw IDs for `@mentions`),
|
||||
and any plugin-registered platform not marked `pii_safe`.
|
||||
|
||||
Redaction applies only to the system prompt text. Routing, session keys, and adapter
|
||||
operations always use the original values.
|
||||
|
||||
---
|
||||
|
||||
## 10. Background Housekeeping
|
||||
|
||||
The `_session_housekeeping_watcher` periodically sweeps idle cached agents, sheds
|
||||
cache entries under memory pressure, and prunes old routing entries hourly.
|
||||
It never ends a transcript because of inactivity or the time of day.
|
||||
|
||||
TTL, LRU and pressure eviction commit the live transcript to memory providers before
|
||||
soft-releasing clients. Active turns remain protected; terminal, browser and background
|
||||
process resources survive soft release. Routing-entry pruning preserves the canonical
|
||||
SQLite transcript, and live processes protect their routing entries from pruning.
|
||||
Historical `expiry_finalized` flags remain recovery fences but are no longer written
|
||||
by a timer watcher.
|
||||
|
||||
---
|
||||
|
||||
## 11. Agent Cache
|
||||
|
||||
The gateway maintains an LRU cache of `AIAgent` instances keyed by `session_key` to
|
||||
preserve prompt caching across turns.
|
||||
|
||||
### Cache Properties
|
||||
|
||||
- **Max size:** 128 entries (`agent.agent_cache.max_size`, default `_AGENT_CACHE_MAX_SIZE`).
|
||||
- **Eviction policy:** Least-recently-used (LRU via `OrderedDict`).
|
||||
- **Idle TTL:** 3600s (1h) — `agent.agent_cache.idle_ttl_secs`, enforced by
|
||||
`_session_housekeeping_watcher`.
|
||||
- **Memory budget:** `agent.agent_cache.memory_high_mb` (default `auto`) — see below.
|
||||
- **Lock:** `_agent_cache_lock` (threading) for thread safety.
|
||||
|
||||
### Memory-Pressure Eviction
|
||||
|
||||
A cached agent pins `_session_messages`, the full live transcript including tool
|
||||
outputs — tens of MB on a session with 100+ tool calls. The entry cap and the idle
|
||||
TTL are both blind to that: a gateway serving many chats keeps every warm transcript
|
||||
resident (agents that took a turn within the TTL are never idle-swept), so RSS climbs until
|
||||
the cgroup throttles and SIGTERM can no longer flush inside systemd's stop timeout
|
||||
(#80764).
|
||||
|
||||
`_sweep_agent_cache_under_pressure()` is the valve. Each watcher tick it compares the
|
||||
process's anonymous RSS against `memory_high_mb`; over budget, it evicts LRU agents
|
||||
through the same soft path the cap enforcer uses (`_commit_then_release_soft`), then
|
||||
runs `malloc_trim` so the freed arenas actually return to the OS. Evicted sessions
|
||||
rebuild their transcript from the persisted session on the next turn.
|
||||
|
||||
Three classes of session are never shed:
|
||||
|
||||
- agents currently mid-turn (their clients and sandboxes are in use);
|
||||
- the `protect_recent` most-recently-used sessions (their prompt cache is worth the
|
||||
most);
|
||||
- any session whose live transcript has not finished reaching disk —
|
||||
`transcript_persistence_caught_up()` compares `_last_flushed_db_idx` against
|
||||
`len(_session_messages)`, the same divergence the FTS write-corruption guard reacts
|
||||
to when it preserves live history over a lagging transcript.
|
||||
|
||||
`memory_high_mb: auto` derives the budget from the cgroup limit the gateway runs under
|
||||
(`memory.high`, then `memory.max`, then cgroup v1), falling back to total RAM when
|
||||
uncapped. Set a number to pin it, or `0`/`off` to disable the pass entirely. Helpers
|
||||
live in `gateway/agent_cache_pressure.py`.
|
||||
|
||||
### Cache Lifecycle
|
||||
|
||||
```
|
||||
Message arrives
|
||||
│
|
||||
▼
|
||||
get_or_create_session() → session_key obtained
|
||||
│
|
||||
▼
|
||||
Lookup _agent_cache[session_key]
|
||||
│
|
||||
├── Hit → move_to_end(), reuse AIAgent (preserves prompt cache)
|
||||
│
|
||||
└── Miss → create new AIAgent, store in cache
|
||||
(if at capacity, popitem(last=False) evicts LRU entry)
|
||||
│
|
||||
▼
|
||||
run_conversation() → agent processes message
|
||||
│
|
||||
▼
|
||||
Housekeeping soft-releases idle agents without ending transcripts
|
||||
```
|
||||
|
||||
### Cleanup Flow
|
||||
|
||||
Resource eviction removes the cached agent and commits memory before soft-releasing
|
||||
clients. Full `_cleanup_agent_resources(agent)` teardown is reserved for actual
|
||||
conversation boundaries and shutdown.
|
||||
|
||||
---
|
||||
|
||||
## Appendix: Key Configuration
|
||||
|
||||
| Config Key | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `group_sessions_per_user` | `bool` | `true` | Isolate group/channel sessions per user |
|
||||
| `thread_sessions_per_user` | `bool` | `false` | Isolate thread sessions per user |
|
||||
| `session_store_max_age_days` | `int` | `0` | Prune sessions older than N days (0=disabled) |
|
||||
| `agent.gateway_auto_continue_freshness` | `int` | `3600` | Seconds for resume freshness window |
|
||||
| `agent.gateway_timeout` | `int` | `1800` | Agent turn timeout (30 min default) |
|
||||
| `agent.agent_cache.max_size` | `int` | `128` | LRU entry cap on cached AIAgents |
|
||||
| `agent.agent_cache.idle_ttl_secs` | `int` | `3600` | Evict agents idle this long |
|
||||
| `agent.agent_cache.memory_high_mb` | `int`/`str` | `auto` | Anon-RSS budget above which LRU transcripts are shed |
|
||||
| `agent.agent_cache.max_evictions_per_pass` | `int` | `16` | Cap on sessions shed per pressure pass |
|
||||
| `agent.agent_cache.protect_recent` | `int` | `8` | MRU sessions the pressure pass never touches |
|
||||
|
||||
## State database and FTS recovery
|
||||
|
||||
The canonical transcript lives in the `sessions` and `messages` tables. FTS5
|
||||
tables and their sync triggers are derived indexes that can be detached and
|
||||
rebuilt without deleting canonical messages. See
|
||||
[`docs/state-db-recovery.md`](state-db-recovery.md) for the bounded live failure
|
||||
mode and the explicit repair procedure.
|
||||
|
||||
### Conversation lifetime
|
||||
|
||||
No idle or daily reset settings are supported. Explicit `/new` and `/reset`,
|
||||
compaction, suspension and crash recovery retain their separate lifecycle roles.
|
||||
@@ -0,0 +1,102 @@
|
||||
# State database and FTS recovery
|
||||
|
||||
`state.db` stores two different data classes:
|
||||
|
||||
- `sessions` and `messages` are the canonical transcript.
|
||||
- `messages_fts*` tables and their sync triggers are derived search indexes.
|
||||
|
||||
The derived indexes may be detached temporarily. They must not turn a live
|
||||
message write or search into an unbounded full-transcript rebuild.
|
||||
|
||||
## Live behavior when FTS is corrupt
|
||||
|
||||
If an FTS write or search reports the corruption error class, `SessionDB`:
|
||||
|
||||
1. records the durable `fts_stale` marker;
|
||||
2. removes the FTS sync triggers in the same transaction;
|
||||
3. retries canonical writes without the derived-index sinks; and
|
||||
4. serves searches from canonical rows through the `LIKE` fallback.
|
||||
|
||||
The failing live operation never runs `FTS5('rebuild')`. Existing recovery
|
||||
ownership remains unchanged: a later `SessionDB` open may rebuild under the
|
||||
cross-process admission lock and foreign-holder guard. If that guarded rebuild
|
||||
cannot run, FTS remains detached, canonical writes stay available, and
|
||||
`hermes doctor` reports the explicit repair command.
|
||||
|
||||
## Live behavior when the file itself is corrupt
|
||||
|
||||
If a live write reports bare `SQLITE_CORRUPT` / `SQLITE_NOTADB` (`database
|
||||
disk image is malformed`, `file is not a database`) with no FTS provenance,
|
||||
the damage is in a canonical B-tree, the schema, or the freelist. `SessionDB`
|
||||
then quarantines that handle (`StateDbCorruptError`):
|
||||
|
||||
1. the failing write propagates the typed error and nothing is retried;
|
||||
2. later writes on the handle fail immediately without touching the file;
|
||||
3. the handle never reopens its connection after `close()`; and
|
||||
4. `close()` skips its explicit WAL checkpoint.
|
||||
|
||||
Stopping the writes is the protection. In the field, a handle that kept
|
||||
writing for ~50 minutes after the first structural error checkpointed 15
|
||||
pages under the wrong page numbers on shutdown (page 1 received a
|
||||
`messages_fts_trigram_data` leaf) and turned a damaged-but-readable file into
|
||||
one that no longer opened at all. Skipping the explicit checkpoint is the
|
||||
second line of defence; on Python 3.12+ the quarantine also disables
|
||||
SQLite's own last-connection checkpoint (`SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE`),
|
||||
so the `-wal` sidecar survives `close()` for forensics. On Python 3.11 that
|
||||
switch is unavailable and SQLite may still checkpoint once on close, so copy
|
||||
`state.db`, `state.db-wal` and `state.db-shm` together before restarting
|
||||
anything.
|
||||
|
||||
The gateway and the agent flush path treat the quarantine like a replaced
|
||||
file: pending transcripts go to `sessions/<id>.jsonl` and the gateway
|
||||
`pending_messages/` spool instead of the retry queue, and the FTS one-shot
|
||||
rebuild never runs on the damaged file. The quarantine is per process — the
|
||||
shared handle stays poisoned for every holder until the process restarts on a
|
||||
repaired or restored file. Do not run `hermes doctor --fix` while the gateway
|
||||
is still up. Next steps:
|
||||
|
||||
```bash
|
||||
hermes gateway stop
|
||||
HERMES_HOME="$HOME/.hermes" hermes sessions recover --source "$HOME/.hermes/state.db" --inspect-only
|
||||
# if recoverable:
|
||||
HERMES_HOME="$HOME/.hermes" hermes sessions recover --source "$HOME/.hermes/state.db" --output "$HOME/recovered-state.db"
|
||||
```
|
||||
|
||||
or restore the newest snapshot from `state-snapshots/`.
|
||||
|
||||
## Explicit repair
|
||||
|
||||
Stop every process that can open the profile database before repairing it.
|
||||
Keep them stopped for the complete repair and verification window.
|
||||
|
||||
```bash
|
||||
hermes gateway stop
|
||||
HERMES_HOME="$HOME/.hermes" hermes sessions repair --check-only
|
||||
HERMES_HOME="$HOME/.hermes" hermes sessions repair
|
||||
```
|
||||
|
||||
`sessions repair` creates a SQLite backup by default and performs structural
|
||||
work through the repository's guarded snapshot-and-promotion path. Do not copy
|
||||
`state.db`, `state.db-wal`, and `state.db-shm` independently with `cp`; those
|
||||
files are one live SQLite image.
|
||||
|
||||
After repair, verify the health probe, stale marker, trigger set, and canonical
|
||||
row counts before restarting the gateway:
|
||||
|
||||
```bash
|
||||
HERMES_HOME="$HOME/.hermes" hermes sessions repair --check-only
|
||||
sqlite3 "$HOME/.hermes/state.db" \
|
||||
"SELECT key, value FROM state_meta WHERE key = 'fts_stale';"
|
||||
sqlite3 "$HOME/.hermes/state.db" \
|
||||
"SELECT type, name FROM sqlite_master WHERE name IN
|
||||
('messages_fts_insert','messages_fts_update','messages_fts_delete')
|
||||
ORDER BY name;"
|
||||
sqlite3 "$HOME/.hermes/state.db" \
|
||||
"SELECT 'sessions', COUNT(*) FROM sessions
|
||||
UNION ALL SELECT 'messages', COUNT(*) FROM messages;"
|
||||
```
|
||||
|
||||
The marker query should return no row, the expected FTS triggers should be
|
||||
present, and canonical row counts must not decrease. If repair fails, preserve
|
||||
both the live database and the reported backup; never delete canonical rows to
|
||||
make a derived-index error disappear.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Streaming TTS
|
||||
|
||||
Hermes can stream TTS audio as it arrives from the provider, instead of waiting
|
||||
for the full audio before playing. This is used by voice mode (CLI/TUI live
|
||||
conversation), the dashboard speak-stream WebSocket, and — via the gateway
|
||||
`StreamingTTSConsumer` — any platform adapter that opts into streaming audio.
|
||||
Voice replies start speaking after the first clause instead of after full
|
||||
generation + synthesis.
|
||||
|
||||
## Architecture
|
||||
|
||||
The streaming pipeline has four parts:
|
||||
|
||||
1. **Producer** — the LLM emits text deltas as it generates a response
|
||||
2. **Sentence chunker** — `tools.tts_streaming.SentenceChunker` accumulates
|
||||
deltas, strips `<think>` blocks (even split across deltas), and flushes
|
||||
complete sentences
|
||||
3. **TTS provider** — a registered `StreamingTTSProvider` turns each sentence
|
||||
into raw PCM chunks (int16 mono at the provider's declared `sample_rate`)
|
||||
4. **Audio sink** — `sounddevice.OutputStream` for local playback
|
||||
(`tools.tts_tool_speaker.stream_tts_to_speaker`), or a gateway platform adapter's
|
||||
`write_streaming_tts` seam (`gateway/streaming_tts_consumer.py`)
|
||||
|
||||
Providers with no chunked API still get per-*sentence* playback via the proven
|
||||
sync `text_to_speech_tool` path, so edge (the default) is conversational too.
|
||||
All spoken text is cleaned by `tools.tts_text_normalize.prepare_spoken_text`
|
||||
(one cleaner, all paths).
|
||||
|
||||
## How to pick a provider
|
||||
|
||||
By default the dispatcher streams with the provider you already configured
|
||||
(`tts.provider`) when that provider has a chunked API — it never silently
|
||||
swaps your voice for a different provider just to get streaming.
|
||||
|
||||
To override, set `tts.streaming.provider` in your `config.yaml`:
|
||||
|
||||
- a provider name (`elevenlabs`, `gemini`, `openai`, `xai`) pins that streamer
|
||||
- `auto` walks the priority list `elevenlabs → gemini → openai → xai` and uses
|
||||
the first one whose credentials resolve — an explicit opt-in to "best
|
||||
chunked voice available"
|
||||
|
||||
```yaml
|
||||
tts:
|
||||
provider: gemini
|
||||
streaming:
|
||||
provider: gemini # or "auto"
|
||||
gemini:
|
||||
model: gemini-2.5-flash-preview-tts
|
||||
voice: Kore
|
||||
```
|
||||
|
||||
## Capability matrix
|
||||
|
||||
| Provider | Transport | Chunked PCM | Credentials |
|
||||
|-------------|---------------------------------------|-------------|-------------|
|
||||
| elevenlabs | chunked HTTP (`pcm_24000`) | yes | `ELEVENLABS_API_KEY` / `tts.elevenlabs` |
|
||||
| openai | chunked HTTP (`with_streaming_response`, `pcm`) | yes | `tts.openai.api_key` → env → managed gateway |
|
||||
| gemini | SSE (`streamGenerateContent?alt=sse`) | yes | `GEMINI_API_KEY` / `GOOGLE_API_KEY` |
|
||||
| xai | WebSocket (`wss://api.x.ai/v1/tts`) | yes | xAI OAuth or `XAI_API_KEY` |
|
||||
| edge, piper, kitten, neutts, mistral, minimax, deepinfra, … | — | no (per-sentence sync fallback) | as usual |
|
||||
|
||||
All credential lookups go through `resolve_provider_secret()`
|
||||
(config > env/.env > credential pool) — never bare env reads. Streamed bodies
|
||||
are capped at 16 MiB per sentence, mirroring the sync providers' bounded
|
||||
upstream-body invariant.
|
||||
|
||||
## Adding a new streaming provider
|
||||
|
||||
1. Subclass `StreamingTTSProvider` in `tools/tts_streaming.py`
|
||||
2. Set `sample_rate` (and `channels` / `sample_width` if not int16 mono)
|
||||
3. Implement `available()` (a pure probe — never install anything) and
|
||||
`stream(self, text) -> Iterator[bytes]` yielding raw PCM chunks
|
||||
4. Decorate with `@register("yourname")`
|
||||
5. Add tests in `tests/tools/test_tts_streaming.py`
|
||||
|
||||
The ABC enforces the contract; the registry makes the provider discoverable;
|
||||
the dispatcher (`stream_tts_to_speaker`) and the gateway consumer handle the
|
||||
sentence buffer, stop events, and audio sink for free.
|
||||
|
||||
## Gateway streaming (platform adapters)
|
||||
|
||||
`gateway/streaming_tts_consumer.py` bridges agent deltas to an adapter's
|
||||
streaming-audio seam. Adapters opt in by overriding, on
|
||||
`BasePlatformAdapter`:
|
||||
|
||||
- `supports_streaming_tts(chat_id, audio_format) -> bool`
|
||||
- `begin_streaming_tts / write_streaming_tts / finish_streaming_tts /
|
||||
abort_streaming_tts`
|
||||
|
||||
All default to unsupported/no-op, so existing adapters are untouched. When a
|
||||
turn's streaming audio completes, the whole-file auto-TTS reply for that turn
|
||||
is suppressed (no double playback); when streaming fails before any audio was
|
||||
audible, the gateway falls back to the legacy whole-file voice reply.
|
||||
Reference in New Issue
Block a user