Commit Graph

27247 Commits

Author SHA1 Message Date
Cyber-Yichen a2fea79de6 fix(cron): isolate multiplex profile failures per profile (#74878)
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:

- startup recovery loop: a per-profile exception (e.g. an unreadable
  executions.db raising sqlite3.DatabaseError) was uncaught and killed the
  ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
  exception escaped to the cycle-wide handler, skipping every remaining
  profile that cycle and marking all of them failed.

Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).

Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).

Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
이민재 e48bb828d4 test(cron): preserve explicit suggestions path override 2026-09-02 06:27:24 -07:00
wanliqin 9da8842585 fix(cron): resolve profile store paths per call
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.

Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium 0fd9218e5a docs: ${VAR} config refs resolve per-profile under a multiplexed gateway 2026-09-02 06:19:03 -07:00
Teknium 638c1f3204 chore(contributors): map patryk.kopycinski@elastic.co -> patrykkopycinski 2026-09-02 06:19:03 -07:00
Teknium b6f0106602 docs(design): note loader-boundary dotenv guard, scoped ${VAR} expansion and scoped .env publish in multiplexing design doc
Accuracy pass on the #89950 doc for the fixes landing alongside it (#77562, #84079, #88441).
2026-09-02 06:19:03 -07:00
Eva 011a60b7cc docs(design): add the multiplexing-gateway design doc referenced by secret_scope
agent/secret_scope.py has pointed at docs/design/multiplexing-gateway.md
("Workstream A") since the fail-closed secret scope landed, but the file was
never added. This writes the missing doc from the code as it stands today:
the mode flag, scope composition (_profile_runtime_scope seams), the
context-local secret scope and HERMES_HOME override, routing/serving/
persistence/session-lane isolation, the control-plane RPCs, failure modes,
and an honest table of what is still process-global (per-profile MCP
registries are tracked in #67605).

Doc-only change; no code touched. Style follows docs/profile-routing.md.
2026-09-02 06:19:03 -07:00
Teknium 2475335443 fix(config): keep .env publishes inside the routed profile scope under multiplex (#88441)
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.

Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
2026-09-02 06:19:03 -07:00
Teknium b53c50cf8a fix(config): resolve ${VAR} config refs through the profile secret scope (#84079)
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).

Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
2026-09-02 06:19:03 -07:00
webtecnica d7dc75ccff fix(gateway): multiplex must not apply default-profile creds to unconfigured profiles (#84079)
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.

Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.

Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Patryk Kopycinski 1eb71756af toolchain-self-improve: route API_SERVER_KEY to profile env 2026-09-02 06:19:03 -07:00
Teknium 0a6aa7cce1 fix(env_loader): log routed-scope dotenv skip once per home; port single-profile control test
Follow-up to the #77592 salvage: emit a once-per-home debug line where the
multiplex guard skips the process-global dotenv load (requested on #77562),
and port the single-profile control test from #77970 so the guard is pinned
to the multiplex flag rather than the home override alone.

Co-authored-by: DonShelly <25538402+DonShelly@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Lester Liang 1aa62ceb45 fix(security): isolate multiplex dotenv reloads 2026-09-02 06:19:03 -07:00
Teknium 0058bde251 fix(bot-relay): a DM into a live Bot Chat queues as the next turn, never interrupts the one in flight 2026-09-02 06:18:55 -07:00
Teknium d29a7936e4 fix(bot-mode): DMs to a Desktop-owned Bot Chat land in the live session instead of being dropped (#100523)
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.

bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.

On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.

Closes #100523
Supersedes #100544, #100542

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 06:18:55 -07:00
Teknium 16370ae539 fix(tui_gateway): setup.status / setup.runtime_check answer for the requested profile
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.

`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.

The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.

Supersedes #94147 (python half)

Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
aeonsong 8076c78c87 fix(desktop): scope handoff config to session profile 2026-09-02 06:17:47 -07:00
Teknium 7a86397a46 fix(api_server): fail closed on unstamped runs; claim session-chat-stream run owner (#93689)
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:

- `_request_owns_run` no longer admits run state that exists without an
  owner stamp. Under gateway.multiplex_profiles every served profile holds
  a valid key, so the "backward compatibility" branch made the boundary
  allow-all whenever provenance was missing. Unstamped state now fails
  closed; only an in-memory owner match or a durable idempotency record
  under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
  inside the request's profile scope, so its run is confined to the
  creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
  (`_release_run_owner_if_forgotten`) and runs at every retirement point
  (task finally, SSE stream close, both sweep loops, chat-stream finally),
  not only the terminal-status sweep — no stranded entries, no stateful id
  ever left unowned.

Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).

Fixes #93689
Fixes #90415
Supersedes #93747, #93704, #92822

Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
Teknium af86ad0479 fix(doctor): reuse resolved memory config at Memory Provider section
Follow-up to salvaged #100677: the file checks read the memory section via the
run's hermes_home while the Memory Provider section re-read config with no
argument (module-global HERMES_HOME). Resolve once and reuse so both sections
report against the same config.
2026-09-02 06:17:25 -07:00
fangliquanflq 21cebbfd68 fix(doctor): honor disabled built-in memory stores
Salvaged from #100677. Fixes #100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
2026-09-02 06:17:25 -07:00
Teknium 70dc1606c6 test(cli): pin OSC 9 / Warp OSC 777 bell emitters; docs + contributor mappings
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
  only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
  of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
2026-09-02 06:17:10 -07:00
Teknium 632078bca7 feat(cli): OSC 9 + Warp OSC 777 notifications ride on the bell flags
Extend _ring_bell() so display.bell_on_prompt / bell_on_complete also
emit terminal-native desktop notifications from the same six call sites
(clarify, clarify batch, approval incl. computer_use, sudo password,
secret capture, turn complete). No new config keys.

- OSC 9 (ESC ] 9 ; body BEL): Ghostty / iTerm2 / Kitty / WezTerm raise an
  OS notification; unknown terminals drop it. Body is "Hermes: <context>"
  with C0 controls and DEL stripped. Written to /dev/tty (prompt_toolkit's
  stdout wrapper can buffer/strip raw escapes) with a sys.stdout fallback.
- Warp OSC 777 warp://cli-agent (agent "hermes", event permission_request
  / stop, compact JSON mirroring build-payload.sh). Gated on
  TERM_PROGRAM=WarpTerminal + WARP_CLI_AGENT_PROTOCOL_VERSION + the
  should-use-structured.sh broken-build floor (stable/preview builds at or
  before v0.2026.03.25.08.24.*_05 rejected). Never raises.

Salvages #58957 and #100805.

Co-authored-by: glitchbunny0 <glitchbunny0@proton.me>
Co-authored-by: harsha-usethread <harsha@usethread.io>
2026-09-02 06:17:10 -07:00
Teknium c2954c8934 feat(model-catalog): picker catalogs refresh every 20 minutes, gateway keeps them warm
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.

- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
  an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
  disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
  it off-thread every TTL window, so every surface on the machine reads
  a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
2026-09-02 06:16:54 -07:00
Teknium 11f932c935 fix(slack): interactive-caller and pre-fetch auth prefer the injected profile check; gate reads never fall through to os.environ
`SlackAdapter._is_interactive_user_authorized` (approval / slash-confirm /
clarify Block Kit clicks) and the early pre-fetch gate in the message
handler recovered the runner via `_message_handler.__self__`, which is
None on a multiplexed adapter (closure handler) — so both fell to env-only
auth. The fallback read `SLACK_ALLOW_ALL_USERS` raw from `os.environ` and
its `_env` helper fell through to `os.environ` on a scoped miss: the
DEFAULT profile's allow-all flag / allowlist authorized callers on every
other profile's bot.

- Prefer the wired `set_authorization_check` callback (profile-bound
  `_make_adapter_auth_check`) at both sites; keep `__self__` introspection
  only for adapters wired without one.
- Env-only fallback reads go through `authz_mixin._platform_gate_env`
  (scoped miss under multiplex → "", never os.environ); drop the raw
  `os.getenv("SLACK_ALLOW_ALL_USERS")` pre-read.

Reapplies #72657 onto current main (original commit carried a bot
co-author trailer). Same class as Telegram #86296 / #65589.

Co-authored-by: MilaArtyNew <261982280+MilaArtyNew@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
Teknium bbb087f3d1 fix(gateway): egress adapter and channel directory no longer follow the per-turn active profile
`_authorization_adapter` compared a stamped profile against
`_active_profile_name()`, which reads the per-turn HERMES_HOME override.
Inside a secondary profile's `_profile_runtime_scope` (cron, restored or
hand-built sources without transport provenance) that reported the
secondary itself, so it was handed the DEFAULT bot for egress instead of
the fail-closed None. Capture the launch identity once in `__init__`
(`_primary_profile_name`) and compare against that; the
`_active_profile_name()` fallback remains for partial fixtures.

`gateway/channel_directory.py` resolved `DIRECTORY_PATH` /
`CHANNEL_ALIASES_PATH` at import time, pinning every multiplexed profile's
directory to whichever home imported the module first. Resolve lazily
from the current home; the module attributes stay as explicit overrides
(tests patch them) and default to None.

Extracted from #87240 (topic-table half handled separately via #76487).

Co-authored-by: cherryb16 <166878179+cherryb16@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
Teknium 74775df53f fix(gateway): route-stamp primary callback auth and carry is_bot through the adapter auth check
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.

- `_make_adapter_auth_check`: for the shared primary adapter under
  multiplex, mirror the inbound message path exactly — stamp the
  `profile_routes` match so the routed profile's pairing store is
  consulted, and authorize under the TRANSPORT home via
  `_is_user_authorized_for_source` (same split as
  `_make_default_profile_message_handler`, 2afed50863). A rejected route
  fails closed like the ingress gate. Retain the receiving adapter as
  `_transport_adapter_ref` so config.yaml policy reads stay on it.
  Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
  `thread_id` as keywords only when set, so legacy 3-positional callbacks
  keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
  the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
  honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
  introspection class — fall back to the injected `gateway_runner` and
  the adapter's owner profile.

Fixes #86296
Fixes #92840

Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
elphamale 4346721117 fix(telegram): resolve button-caller authorization via the injected auth check, not handler introspection
_is_callback_user_authorized resolved the gateway's auth chain through
_message_handler.__self__. For a secondary multiplexed adapter the
message handler is a per-profile closure with no __self__, so the
introspection silently fell through to the env-only fallback -- which
knows nothing about config allowlists or the pairing store, denying
every button caller on that profile (fail-closed, but wrong).

Prefer the auth callback GatewayRunner already injects at connection
time via set_authorization_check (registered for primary and multiplexed
adapters alike, delegating to the full _is_user_authorized chain), and
keep the introspection plus env fallback for adapters wired without it.
Same resolution pattern the admin-tier gate uses.
2026-09-02 06:08:09 -07:00
Teknium 4afbecb429 test(cli): trim fast-serve coverage to the parity + dispatch invariants 2026-09-02 06:06:47 -07:00
Gille 5180601a6a perf(cli): dispatch serve without the full parser tree 2026-09-02 06:06:47 -07:00
Teknium e9dd0bf5d5 feat(desktop): polish bot roster sections — dialog rename, Undo delete, Esc-cancel drag, nested under gateways (salvage #100745)
Follow-up on @fortun8te's user-made roster sections:

- Sections start empty: no seeded General/Workforce/Clients. With no
  sections created the roster renders exactly as before.
- New section and Rename go through one Dialog + Input + Cancel/Save
  (the app's session-rename shape) instead of an inline caret; the row
  menu's "New section…" files the bot as it creates.
- Delete needs no confirmation: bots return to Unassigned and the toast
  offers Undo (restores the section in its slot and refiles its bots).
- Drag: single-row drag under a private MIME type, every valid target
  shows a faint outline while a drag is live, the hovered target lights
  up, the source section refuses the drop, Escape cancels, and the moved
  row no longer stays faded after it remounts under its new section.
- Multi-select (cmd/shift-click, querySelectorAll shift-range) dropped:
  the roster has no selection model. Per-bot saveBotMeta writes run in
  sequence, one per profile (membership IS a field on each profile).
- Section heading reuses RosterSectionHeader (gains `action` /
  `onDoubleClick`), so user sections fold and look like the gateway
  headings; ⋯ menu and right-click drive the same Rename / Move up /
  Move down / Delete. Empty sections show a dashed "Drag bots here" slot.
- Composes with gateway buckets: sections nest INSIDE each connection
  bucket, indented under a hairline rail (membership lives in the bot's
  profile on that gateway); empty sections repeat there only mid-drag.
- Full i18n parity (en / ja / zh / zh-hant) for every new string; icon
  toggle and the storage-async plumbing removed.
- Tests trimmed to the three invariants (membership persists through
  saveBotMeta + reload, remainder = Unassigned, delete returns bots +
  undo) plus a live Electron e2e covering the whole flow.
- Docs: "Organize bots into sections" in user-guide/bot-mode.md.
2026-09-02 06:06:32 -07:00
Michael Knaap bd9955d529 fix(desktop): section rename — Escape cancels, Enter commits once
Closing the rename field unmounts the input, and the unmount can still
fire onBlur, which committed the draft the user had just asked to throw
away with Escape. Enter also called commit() directly and then again from
blur. Route both through blur with a cancelled flag so the commit runs
exactly once and Escape never renames.
2026-09-02 06:06:32 -07:00
Michael Knaap 3d0ac691af feat(desktop): user-made sections in the bot roster, with drag-and-drop filing
The roster already has sections, but only automatic ones: one per gateway
connection plus the group-chat bucket. Those answer "where does this bot
run", which is not the question being asked when someone wants two client
bots filed together under "Clients" and the internal ones under "Team".

This adds a second axis that composes with the first: gateway sections keep
the top level whenever more than one connection is showing, and user
sections group the flat list underneath.

Design choices, each deliberate:

- Membership lives on the BOT (`ui_meta.sectionId`), not as a member list on
  the section. A bot can only be in one place, deleting a section cannot
  orphan anybody, and the assignment rides the same profile.yaml sync every
  other bot setting already uses, so it follows the profile to another
  machine. Section records (id, name, icon) live in plugin storage.
- "Unassigned" is not a section. It is whatever is left, always drawn last,
  and it is where members of a deleted section land. No record, so nothing
  to keep in sync.
- Three gestures, one rule: drag a row onto a section heading; cmd/ctrl-click
  and shift-click build a multi-selection (shift ranges in DOCUMENT order,
  anchored Finder-style); and the row's context menu gets "Move to section…"
  with the same targets the drag would use. Dragging a row that is part of
  the selection drags the whole selection.
- The drag uses a private MIME type, so a bot dropped on the composer or the
  transcript is simply not a valid payload there instead of pasting its key
  as text.
- Section headings rename inline (double-click / menu), reorder, hide their
  glyph, and delete (keeping their bots). Right-click and the ⋯ button open
  the same menu so neither can drift.

With no sections created the roster renders exactly as before.

Tests: user-sections.test.ts covers the pure model (normalisation, grouping
with unknown/deleted sections falling to Unassigned, drag payload
round-trip). The existing hermes-bots suite passes; tsc and eslint clean.
2026-09-02 06:06:32 -07:00
Teknium bb7838c57c chore: map contributor crdesign8@hotmail.com -> @crdesign8 2026-09-02 05:59:24 -07:00
Teknium 458e2ef1d2 refactor(state): collapse telegram topic v3 migration into one table-driven rebuild
Same behavior as the salvaged #76487 migration (fresh installs get the v3
shape; v1/v2 tables rebuild with profile_name leading the PK, legacy rows
into 'default' only, CASCADE FK supplied on the way), with the per-table
DDL written once instead of three times and the now-redundant v1->v2
CASCADE-only rebuild dropped (the v3 rebuild subsumes it).

Co-authored-by: Celio Monteiro <crdesign8@hotmail.com>
2026-09-02 05:59:24 -07:00
Celio Monteiro d55d9d128a fix(gateway): route profile into topic prune, cooldowns, and docs
Address hermes-sweeper review on #76487:

- Prefer hermes_profile from send metadata when pruning stale topic
  bindings so profile_routes cannot delete the transport adapter's
  namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
  (profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
2026-09-02 05:59:24 -07:00
Celio Monteiro 62be7043ff fix(gateway): pass routed source.profile into telegram topic state
Issue #76423 follow-up: wire SessionDB profile_name through gateway paths.

- Resolve profile from source.profile (never process-global active profile)
- Stamp adapter._hermes_profile_name for prune under multiplex
- /topic enable/status and binding record/recover/disable/restore paths
2026-09-02 05:59:24 -07:00
Celio Monteiro 351e4c0067 fix(state): namespace telegram topic tables by profile_name
Issue #76423: under multiplex_profiles a shared state.db keyed topic mode
and bindings only by Telegram chat_id/thread_id, so private-chat ids
collided across bots/profiles.

- Add profile_name to telegram_dm_topic_mode and telegram_dm_topic_bindings
- Schema v2→v3 rebuild; legacy rows migrate into the "default" namespace
- Keyword-only profile_name="default" on SessionDB topic APIs (compat)
2026-09-02 05:59:24 -07:00
Teknium 648c664eb3 feat(desktop): gate cold-start restore on display.resume_last_session
- use-desktop-integrations: hold the restore latch until the config
  record answers; when false, stay on the fresh chat (route and session
  restore alike) while still remembering the open chat for next launch.
- wiring: read the shared config-record query; undefined while pending,
  fetch failure falls back to the historical behavior (resume).
- appearance-settings: ToggleRow writing through the shared config
  cache with rollback + notifyError on a failed save.
- ar/ru strings, docs line in user-guide/desktop.md, two hook tests.
2026-09-02 05:56:54 -07:00
jinglun010 7aff724e56 feat(desktop): add display.resume_last_session config toggle (#60812)
Config default (true) plus the Appearance-settings strings for a
"Reopen Last Chat on Launch" switch. Salvaged from PR #60816 onto
current main (defaults moved to config_defaults.py since the PR).
2026-09-02 05:56:54 -07:00
Teknium 0437fe66f7 fix(discord): native slash commands honor guild/channel profile_routes
_build_slash_event and _dispatch_thread_session built their SessionSource
without guild_id/parent_chat_id, while on_message passes both. Route
matching in build_source keys off exactly those fields, so under
gateway.multiplex_profiles a guild- or channel-routed profile never matched
a native slash command: /new, /reset, /model, /profile, /status ... all ran
against the default profile and reset the wrong session (#69178, #91633).

Pass guild_id (interaction.guild_id, falling back to channel.guild like the
message path) and the thread's parent channel id into build_source at both
sites. One test pins channel + thread routing parity with messages.

Fixes #69178
Fixes #91633
Co-authored-by: Sora-bluesky <179361977+Sora-bluesky@users.noreply.github.com>
Co-authored-by: jondgilbert <42873618+jondgilbert@users.noreply.github.com>
Co-authored-by: tensorbit89-netizen <257030052+tensorbit89-netizen@users.noreply.github.com>
2026-09-02 05:55:36 -07:00
Teknium 17f86ab36f fix(gateway): enter per-message profile scope off the event loop too
Same class as the handoff watcher (#100014): the secondary/primary message
handlers and the scoped inbound-preprocess hop entered
_profile_runtime_scope synchronously on the loop thread, so a slow
profile .env read or external secret-source hydration stalled every other
adapter's traffic. Route those three async sites through
_async_profile_runtime_scope (asyncio.to_thread load, then the existing
sync scope with prepared_secret_scope=). _format_session_info_scoped is
already called via asyncio.to_thread and stays sync.
2026-09-02 05:55:36 -07:00
fangliquanflq f069ffd471 fix(gateway): offload handoff secret scope loading
The handoff watcher entered _profile_runtime_scope synchronously on the
event loop each tick; hydrate_profile_secret_sources + build_profile_secret_scope
do blocking file/secret-source IO, so a slow profile secret read stalled every
adapter (#100014). Load the secret scope via asyncio.to_thread, then enter
the existing sync scope with prepared_secret_scope=.

Fixes #100014
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-02 05:55:36 -07:00
Adolanium 55e1698979 fix(gateway): a secondary profile handoff fails closed when its config cannot load
_process_handoff caught a config load failure for a secondary profile,
logged a warning, and kept going with self.config, which is the primary
profile's config. The handoff then went out through the right bot to the
primary's home channel and the row was reported completed. That is the
exact wrong delivery the multi-profile handoff work exists to prevent,
and the same fail closed posture the no-live-adapters branch already
takes.

A load failure now logs an error and raises, which marks the row failed
so the CLI can report and retry it. The default profile path is
untouched, it never reloaded config.
2026-09-02 05:55:36 -07:00
kshitijk4poor e9fa7bc05e fix(state): keep the #94736 teardown self-heal alive under the deleted-WAL guard
Follow-ups on the salvaged #101081 guard:

- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
  guard treated that as a lost generation and permanently halted the
  handle, so the #94736 late-write self-heal reopen dropped transcript
  tails (4 existing tests failed). close() now clears the recorded
  sidecar generation, and _wal_generation_was_lost() re-adopts the
  current sidecars after a clean /proc/self probe instead of relying on
  a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
  generation is recorded, the stat-based inode check alone detects an
  unlink/replace. The fd probe only runs in the empty-identity state
  (fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
  gateway retry queue and run_agent flush divert transcripts to the
  JSONL fallback exactly as they do for a replaced store, instead of
  retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
  open, halving the system-wide /proc scan; dropped the dead
  include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
  non-linux test now patches sys.platform (the real gate) instead of
  _IS_WINDOWS.
2026-09-02 18:24:54 +05:30
Cursor Agent 7f7df1ce44 fix(state): refuse SessionDB open and writes on a deleted WAL generation
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.

Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
2026-09-02 18:24:54 +05:30
liguoyu 1592e48ac9 fix(desktop): a split-dragged Bot tab keeps its Bot workspace scope
Dragging a session tab to split in the Bot workspace committed through
openSessionTile with no scope, whose default { workspaceMode: 'sessions' }
was written onto the already-open tile before the pane move — so the tile
re-bucketed into the Sessions workspace and vanished from the Bot strip.

A scope-less open of an already-open tile is a MOVE, not a re-scope:
default the scope to the tile's current workspaceMode and only rewrite
scope when the caller passed one explicitly (sidebar/bot openers still
win). Brand-new tiles keep the 'sessions' default.

Closes #96865
Supersedes #96998

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-02 05:54:10 -07:00
Teknium d1f8c2ea11 test(desktop): group chat picker row label opts out of min-width:auto so names truncate
Renders CreateGroupChatDialog with a long-named bot and asserts the row
label carries min-w-0 (and the inner text column keeps min-w-0 flex-1 +
truncate). Sabotage-verified: fails against the pre-#90624 markup with
'expected [...] to include min-w-0'.
2026-09-02 05:42:14 -07:00
Jay c65c79a4a8 fix(bot-mode): group chat picker rows overflow and scroll names out of view
The New Group Chat picker's rows are `label` flex containers holding a
`min-w-0 flex-1` text column whose two lines are `truncate`. The label
itself has no `min-w-0`, so as a flex item it keeps its `auto` minimum
width and cannot shrink below its content. `truncate` therefore never
fires: the row grows to fit the longest secondary line instead, which is
`@handle · in "Group A", "Group B", …` and so scales with how many groups
a bot already belongs to.

The row is a grid item inside a Radix ScrollArea, whose viewport wraps
children in a `display: table` div that sizes to content, so the whole
list widens rather than clipping.

Measured on a 414px viewport with one bot in three groups:

  wrapper width  593  (viewport 414)
  widest row     585
  overflows      yes

Nothing is visibly wrong until the first click. The checkbox now sits
past the right edge, so focusing it scrolls it into view: `scrollLeft`
jumps 0 → 178.38 (= 593 − 414) and every row shifts to `left: -162px`,
clipping the bot names from the left — the user clicks a name and the
names disappear. The scroll offset persists after unchecking, until the
dialog is remounted.

Adding `min-w-0` to the label lets it shrink, so `truncate` engages as
the markup already intended. Same viewport, same data:

  wrapper width  414  (unchanged display: table)
  widest row     406
  overflows      no
  scrollLeft after clicking a row  0 → 0

`display: table` on the ScrollArea wrapper is untouched; the fix works
with it rather than around it. Verified against a packaged build via CDP,
before and after, on identical roster data.

No test. The rule for this is "extract the logic into a small
pure/DI-testable function and call it for real", but there is no logic
here — `min-w-0` is a class name, and the behaviour under test belongs to
the layout engine. jsdom does not lay out, so a unit test cannot observe
the overflow; the Playwright suite could, but has no bots-roster fixture,
which is a large scaffold to hang off a one-class change. A source-regex
assertion would pass without ever laying anything out, which is precisely
the false confidence AGENTS.md describes. The before/after measurements
above are offered as the evidence instead — happy to add a Playwright
case if you'd rather have the fixture.
2026-09-02 05:42:14 -07:00
Teknium 8d6a286fe8 fix(desktop): restored background tabs resolve their session title without a click
A restored session tile has no runtimeId and never mounts its pane until first
activation, so the by-id resolution effect inside SessionTilePane never runs.
When the row is also outside the recents page and project tree, tileTitle()
falls back to "New session" until the user clicks the tab (#94167).

Add a one-shot backfill, wired next to watchSessionTiles(): once the gateway
is open, look each unrestored, untitled, unlisted tile up via
resolveStoredSession(id, tile.ownerRoute). That call already upserts the row
into $sessions, which the tab strip watches, so the tab renames itself —
nothing new is persisted and workspaceTabTitle stays the Bot Chat marker.

Live repro (Electron e2e, target session pushed off the 50-row recents page
by 60 newer sessions, restored as a stacked background tab): main showed
"New session" after boot with no click; with this fix the tab reads the real
title while the pane is still unmounted.

Closes #94167
Supersedes #94212

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 05:42:00 -07:00