Commit Graph

28271 Commits

Author SHA1 Message Date
joaomarcos 11942f6de5 fix(gateway): fail closed when a shutdown worker survives the quiesce budget
andrexibiza's review on #101118 pointed out that the timeout branch
still ran the SessionDB close/checkpoint even when _shutdown_executor()
reported a live worker -- the exact sequence that produces the
wrong-page-number corruption in #101093. The close block now only runs
when _exec_live == 0; a surviving worker skips it entirely and leaves
the handle open for SQLite to recover from its own WAL on next open.

Adds test_stuck_worker_skips_the_session_db_close to prove the converse
of the existing ordering test: a worker that outlives the budget must
never be raced by close().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019JujDvCo2vfpEiizidAS2U
2026-09-02 20:43:06 +05:30
joaomarcos af52474aba fix(gateway): quiesce the thread pool before closing state.db at shutdown
`_shutdown_executor()` ran *after* the SessionDB close block in `_stop_impl`,
and it never waited. That left two ways for blocking DB work to outlive
`SessionDB.close()`:

  (a) `_executor_closing` was still False during the close, so a coroutine
      reaching `_run_in_executor_with_context` minted a brand-new pool and ran
      more blocking DB work against handles that had just been closed;
  (b) `cancel_futures` only drops work that has not started, and cancelling
      `self._background_tasks` does not stop the worker thread behind a
      `run_in_executor` future that is already running.

`SessionDB.close()` checkpoints the WAL and lets SQLite unlink the sidecar. A
write that lands after it silently reopens the handle (#94736) and mints a
fresh WAL generation behind that checkpoint, so teardown checkpoints the same
file a second time from a connection the shutdown log never accounts for --
the close-time page-write damage in #101093 and the split WAL generation in
#101064.

The quiesce now runs before the close and waits for the running workers. The
wait is bounded by `_EXECUTOR_QUIESCE_TIMEOUT` (2s) and clamped to what is left
of the shutdown watchdog leash minus a second for the close itself, so a stuck
worker can never cost the post-close cleanup window (#82161). Workers still
alive after the budget are logged as a warning instead of being waited on.

`_shutdown_executor()` keeps its no-argument fire-and-forget contract and now
returns the number of workers still running.

Refs #101093

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DGpGPnvz5FFFH999i59Xeb
2026-09-02 20:43:06 +05:30
kshitijk4poor 95f62ca3bf fix(cron): validate failure_deliver at preflight and dashboard update lanes
Follow-up to the failure_deliver salvage (#100375):

- _preflight_check_delivery also checks the failure lane, so a typo'd
  failure_deliver platform blocks at config-validation time instead of
  surfacing only when a failure occurs — exactly when the notice must
  not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
  deliver (text normalization; empty clears the optional override
  instead of coalescing), closing the one update path that could write
  an unnormalized value into jobs.json.

4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
2026-09-02 20:16:14 +05:30
Victor Kyriazakos fd35e1ec5a fix(cron): delivery bookkeeping reads the failure lane it actually routed through
Review findings (Salt, NS-788):

B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.

S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.

S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.

Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
2026-09-02 20:16:14 +05:30
Victor Kyriazakos c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium 73f68362b3 fix(sessions): auto-prune state.db by default (90d) and gate VACUUM on freelist ratio (#54189)
Flip the state.db retention defaults per Teknium's decision on #54189:

- sessions.auto_prune: false -> true. A stock install now prunes ENDED
  sessions inactive for retention_days at CLI/gateway/cron startup
  (at most once per min_interval_hours). Open, pinned and mid-turn
  sessions are never deleted; the only open rows touched are stale
  automation sessions (#100903 sweep), which are closed, not deleted,
  and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
  the file: PRAGMA freelist_count / page_count must exceed 25%
  (AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
  min_vacuum_interval_days throttle. Pruning a few small sessions on a
  dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
  Unknown ratio (pragma read failure) falls back to the time throttle.

Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.

Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
2026-09-02 07:26:52 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
joaomarcos 65b0f00002 fix(agent): stop the between-turns tool refresh from forking the cached prefix
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:

* a tool whose `check_fn` merely flapped (headless browser probe, expired
  credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.

Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.

`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.

Refs #100336
2026-09-02 07:22:59 -07:00
Teknium 2e25b47210 chore: map contributor email for #80825 salvage 2026-09-02 07:01:23 -07:00
Teknium febd2af391 fix(gateway): skip credential-less WhatsApp on secondary multiplex profiles
_start_one_profile_adapters skipped only Platform.RELAY as shared
process-level ingress. WhatsApp is the same shape: the bridge is one
authenticated session tied to a single phone number, so a secondary
profile has no credential of its own to bring; constructing an adapter
for it only produced a connect/retry loop that stalled startup for every
profile queued behind it. Treat WhatsApp like Relay -- the active profile
owns the connection and route-stamped source.profile fans inbound turns
out to secondary profiles.

Salvage of #69042 (narrowed by its author to this one behavioral line);
test re-expressed on the current secondary-startup fixtures.

Co-authored-by: sshawn <28279366+lsshawn@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Teknium 4fa1d5498e docs(multiplex): list teams among port-binding platforms 2026-09-02 07:01:23 -07:00
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
Teknium 001b8abbd4 fix(matrix): pin the E2EE crypto store per profile at connect(), not import
The multiplex gateway imports plugins/platforms/matrix/adapter.py once, so
the module-level _STORE_DIR/_CRYPTO_DB_PATH resolved against the root
HERMES_HOME for every profile: all bots' Olm identities landed in one
crypto.db and inbound E2EE failed with "no session found" (#89168).

connect() runs inside _profile_runtime_scope, so resolve the store dir
there via get_hermes_dir (honors the context-local HERMES_HOME) and cache
it on the instance -- diagnostics and error-log paths read outside the
scope then still report the store actually in use. Mirrors the
pairing-store fix (a6397c379).

Salvage of #89169 (per-call resolvers collapsed into one cached resolve;
dead `_CRYPTO_DB_PATH = None` alias dropped -- no external importers).
Also routes the last raw MATRIX_HOMESERVER read in check_matrix_requirements
through _startup_env_secret like its token/password neighbours (#69943).

Fixes #89168

Co-authored-by: Michael Short <18595461+mjshorty@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Teknium 9be1168cd3 fix(wecom): scope WECOM_WEBSOCKET_URL like its neighbours
Same class as #100627's WECOM_BOT_ID: the one remaining raw os.getenv in
WeComAdapter.__init__ let a secondary multiplex profile pick up the
default profile's bridged websocket URL. Route it through
_get_scoped_secret; folded into the existing scoped-miss test.
2026-09-02 07:01:23 -07:00
Teknium 2bcbdb61a7 fix(simplex): scope SIMPLEX_* reads to the active profile under multiplexing
SimplexAdapter.__init__ (auto_accept, group_allowed), the registry gates
check_requirements/validate_config/is_connected, _env_enablement and
_standalone_send all read SIMPLEX_* via raw os.getenv. Under
gateway.multiplex_profiles those paths run inside a secondary profile's
scope where os.environ holds the DEFAULT profile's YAML-to-env bridge
output -- so a secondary profile that never configured SimpleX was
auto-enabled on the default's daemon URL and inherited its group
allowlist / auto-accept setting.

Route every read through the module-local `_get_scoped_secret` wrapper
(get_secret; UnscopedSecretError -> os.getenv for the default profile,
which constructs unscoped) -- the same helper the IRC/ntfy/Photon/
Mattermost siblings use. Unlike the extra-only `_scoped_platform_setting`
shape proposed in #100241, this honors BOTH the secondary profile's own
.env (the scope) and its config.yaml extra, and needs no config.yaml
re-read in check_requirements.

Rewrite of #100241.

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
nftpoetrist 07ee457a21 fix(wecom): scope WECOM_BOT_ID reads to the active profile under multiplexing
WeComAdapter.__init__ read WECOM_BOT_ID via a raw os.getenv() call, while
the immediately adjacent line for WECOM_SECRET already used the module's
_get_scoped_secret() helper. Under gateway.multiplex_profiles, a secondary
profile's adapter is constructed inside a scoped context where os.environ
still holds the DEFAULT profile's env-bridge output -- so a secondary
profile's bot would silently connect using the default profile's bot_id
while (correctly) using its own secret, or vice versa on a scope miss.

Switch the bot_id read to _get_scoped_secret(), matching the sibling
_secret/_dm_policy/_group_policy/allow_from reads in the same __init__
that were already migrated in #76664/#93545. _standalone_send's
out-of-process fallback branch constructs a fresh WeComAdapter(pconfig)
and therefore inherits this fix automatically -- no separate change
needed there.

Adds two regression tests to the existing TestWeComAdapterAuthzScope
class (already covering dm_policy/allow_from scoping per #93522),
mirroring its established fixture/assertion style. Mutation-verified:
both fail against the pre-fix code (asserting the default profile's
bot_id leaks into a secondary profile's scope) and pass with the fix.
2026-09-02 07:01:23 -07:00
nftpoetrist 56d869d2d9 fix(mattermost): scope url/reply_mode/require_mention/free_response_channels/allowed_channels to the active profile under multiplexing
MattermostAdapter.__init__, validate_mattermost_config, _standalone_send,
and _handle_ws_event's mention-gating block all read MATTERMOST_URL/
MATTERMOST_REPLY_MODE/MATTERMOST_REQUIRE_MENTION/
MATTERMOST_FREE_RESPONSE_CHANNELS/MATTERMOST_ALLOWED_CHANNELS via raw
os.getenv -- only MATTERMOST_TOKEN was already scoped via
_get_scoped_secret. _apply_yaml_config additionally wrote
MATTERMOST_REQUIRE_MENTION/MATTERMOST_FREE_RESPONSE_CHANNELS/
MATTERMOST_ALLOWED_CHANNELS into the process-global os.environ
unconditionally (guarded only by `not os.getenv(...)`, first-writer-wins),
the same apply_yaml_config_fn bug class already fixed for the
Discord/Telegram/WhatsApp/DingTalk adapters in this series.

Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's
env-bridge output. A secondary profile with its own (or no) Mattermost
config could silently connect to the default profile's server, thread
its replies per the default profile's reply_mode, or -- since
_handle_ws_event's mention-gating block runs on every LIVE inbound
message, not just at construction -- have its require_mention/
free_response_channels/allowed_channels decisions driven by the default
profile's settings for the adapter's entire runtime lifetime.

Fix, mirroring the WhatsApp/DingTalk apply_yaml_config_fn pattern:
- Add _profile_scoped_config_load() (same helper as DingTalk).
- Rewrite _apply_yaml_config to skip the env-bridge write under a
  multiplexed secondary profile's scope, and instead return the YAML
  values as a dict merged into this profile's own PlatformConfig.extra.
- Make require_mention/free_response_channels read extra first (matching
  the existing allowed_channels precedent), falling back to
  _get_scoped_secret() instead of raw os.getenv when extra is absent --
  fixing a residual gap the DingTalk fix (#100615, this series' item 6)
  left in its own analogous extra-first-with-raw-fallback read sites
  (_dingtalk_require_mention et al. still fall back to bare os.getenv).
- Switch __init__'s url/reply_mode, validate_mattermost_config's url, and
  _standalone_send's url to _get_scoped_secret().
- Leave check_mattermost_requirements() (no longer reads any MATTERMOST_*
  var on current main -- just an aiohttp-importability probe) and
  _is_connected() (already scope-aware via hermes_cli.gateway.get_env_value,
  which itself routes through agent.secret_scope.get_secret) untouched.

Adds a new TestMultiplexProfileScope class to tests/gateway/test_mattermost.py
(7 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope, plus two
tests exercising _apply_yaml_config's new seeded-dict return directly.
Mutation-verified: stashed the production fix and confirmed 5 of 7 new
tests fail against pre-fix code (the other 2 are non-differentiating
regression guards -- extra-wins-over-env and unscoped-default-profile-
precedence -- which correctly pass either way). Restored the fix;
all 30 tests in the file, the plugin-setup test, and the full 75-test
tests/gateway/test_adapter_startup_secret_scope.py suite pass.
2026-09-02 07:01:23 -07:00
nftpoetrist c36def6aea fix(photon): scope project_id/node_bin/require_mention/reactions/sidecar config to the active profile under multiplexing
PhotonAdapter.__init__, check_requirements, validate_config,
_env_enablement, _markdown_enabled, _reactions_enabled, and
_standalone_send in adapter.py, plus load_project_credentials and
load_dashboard_project_id in auth.py, all read PHOTON_PROJECT_ID/
PHOTON_NODE_BIN/PHOTON_SIDECAR_PORT/PHOTON_SIDECAR_AUTOSTART/
PHOTON_PROBE_*/PHOTON_REQUIRE_MENTION/PHOTON_MENTION_PATTERNS/
PHOTON_REACTIONS/PHOTON_MARKDOWN/PHOTON_HOME_CHANNEL(_NAME)/
PHOTON_DASHBOARD_PROJECT_ID via raw os.getenv -- only
PHOTON_PROJECT_SECRET and PHOTON_SIDECAR_TOKEN were already scoped via
_get_scoped_secret.

Notably __init__'s project_id read was a stronger variant of the bug
(like the IRC fix in this series, item 11): the original
`os.getenv("PHOTON_PROJECT_ID") or extra.get("project_id") or stored_id`
ordering let a raw env read override even an explicitly configured
config.yaml extra -- a secondary profile that set its own project_id via
extra would still silently authenticate against the default profile's
Spectrum project, because the default profile's project id is always
bridged to os.environ under multiplex and env was checked first.

_reactions_enabled() and the require_mention/mention_patterns reads in
__init__ are exercised on every live inbound message / tapback, not just
at construction, so a secondary profile's reaction/mention-gating
behavior would be driven by the default profile's settings for the
adapter's entire runtime lifetime.

Switch every raw PHOTON_* read (except the two already scoped) to
_get_scoped_secret(), matching the module's existing helper (already
defined identically in both adapter.py and auth.py). Left
_dashboard_host()/_spectrum_host() and the interactive device-login flow
functions in auth.py untouched -- these are CLI-only management-plane
calls (`hermes photon login`/`setup`), not part of the gateway's
per-profile adapter construction/connection lifecycle, so they are not
reachable under a multiplexed secondary profile's scope; noted as a
"Scope note" in the PR body rather than silently expanding scope to
unreachable call sites.

Adds a new tests/plugins/platforms/photon/test_multiplex_profile_scope.py
(9 tests, two classes covering auth.py and adapter.py separately)
mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope, reusing
test_auth.py's tmp_hermes_home isolation pattern so tests don't depend on
the real ~/.hermes/auth.json fallback. Mutation-verified: stashed the
production fix and confirmed 7 of 9 new tests fail against pre-fix code
(the other 2 are non-differentiating regression guards -- unscoped-
default-profile-precedence, one per class -- which correctly pass either
way). Restored the fix; all 140 tests in tests/plugins/platforms/photon/,
the 10 photon-related parametrized tests in
test_adapter_startup_secret_scope.py, and the broader
test_multiplex_adapter_registry.py / test_adapter_connect_classification.py
suites (45 tests) pass.
2026-09-02 07:01:23 -07:00
nftpoetrist 327e9043e9 fix(ntfy): scope server/topic/publish_topic reads to the active profile under multiplexing
NtfyAdapter.__init__, _env_enablement, check_requirements, validate_config,
is_connected, and _standalone_send all read NTFY_SERVER_URL/NTFY_TOPIC/
NTFY_PUBLISH_TOPIC/NTFY_MARKDOWN/NTFY_HOME_CHANNEL(_NAME) via raw os.getenv
-- only NTFY_TOKEN already went through the module's _get_scoped_secret
helper. Under gateway.multiplex_profiles, env_enablement_fn/check_fn/
is_connected all run inside the registry-enablement loop in
load_gateway_config() (confirmed in gateway/config.py, lines ~2704-2820,
inside _profile_runtime_scope for secondary profiles), and adapter
construction likewise runs scoped -- so os.environ there still holds the
DEFAULT profile's env-bridge output. A secondary profile with its own (or
no) ntfy topic configured could silently:

- get auto-enabled via _env_enablement()/is_connected() using the default
  profile's topic, even though it never configured ntfy itself
- have its adapter subscribe to / publish on the default profile's topic
  and server instead of (or in addition to) its own
- deliver cron/send_message_tool messages via _standalone_send to the
  wrong topic

Switch every raw NTFY_* read (except the two secret-material fields
already scoped: NTFY_TOKEN) to _get_scoped_secret(), matching the
established helper already defined in this module and used for
NTFY_TOKEN, and the same pattern applied to the sibling LINE/DingTalk/
Teams/SMS/WeCom adapters in this series.

Adds a new TestMultiplexProfileScope class to tests/gateway/test_ntfy_plugin.py
(7 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope. Mutation-
verified: stashed the production fix and confirmed 5 of the 7 new tests
fail against pre-fix code (the other 2 are non-differentiating regression
guards -- extra-wins-over-env and unscoped-default-profile-precedence --
which correctly pass either way); restored the fix and confirmed all 37
tests in the file, plus the file's 5 parametrized _get_scoped_secret tests
in test_adapter_startup_secret_scope.py, pass.
2026-09-02 07:01:23 -07:00
nftpoetrist 05aa749977 fix(irc): scope server/port/nickname/channel/use_tls reads to the active profile under multiplexing
IRCAdapter.__init__, check_requirements, validate_config, is_connected,
_env_enablement, and _standalone_send all read IRC_SERVER/IRC_PORT/
IRC_NICKNAME/IRC_CHANNEL/IRC_USE_TLS via raw os.getenv -- only
IRC_SERVER_PASSWORD/IRC_NICKSERV_PASSWORD already went through the
module's _get_scoped_secret helper. Under gateway.multiplex_profiles,
env_enablement_fn/check_fn/is_connected all run inside the registry-
enablement loop in load_gateway_config() (gateway/config.py, ~lines
2704-2820), scoped for secondary profiles via _profile_runtime_scope,
and adapter construction runs scoped the same way -- so os.environ there
still holds the DEFAULT profile's env-bridge output.

Notably, __init__'s original `os.getenv("IRC_SERVER") or extra.get(...)`
ordering let a raw env read override even an explicitly configured
config.yaml extra -- a secondary profile that set its own server/channel
via config.yaml extra would still silently connect to the default
profile's IRC server/channel/nick if the default profile bridged its own
config to env (which it always does under multiplex). This is a stronger
variant of the same bug fixed for the sibling LINE/DingTalk/Teams/SMS/
WeCom/ntfy adapters in this series -- there, extra already won because
of the `extra.get(...) or os.getenv(...)` order.

Switch every raw IRC_* read (except IRC_SERVER_PASSWORD/
IRC_NICKSERV_PASSWORD, already scoped) to _get_scoped_secret(), matching
the module's existing helper. Also collapses a double os.getenv("IRC_USE_TLS")
read in __init__ into a single _get_scoped_secret() call (same behavior,
one scope lookup instead of two).

Adds a new TestMultiplexProfileScope class to tests/gateway/test_irc_adapter.py
(6 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope. Mutation-
verified: stashed the production fix and confirmed 5 of 6 new tests fail
against pre-fix code -- including the "extra wins" test, since IRC's
original env-first ordering meant even an explicit extra config was not
a safe differentiator boundary before the fix (only the DEFAULT-profile-
unscoped-precedence test is a non-differentiating regression guard that
correctly passes either way). Restored the fix; all 23 tests in the file
pass, plus the file's 5 parametrized _get_scoped_secret tests in
test_adapter_startup_secret_scope.py.
2026-09-02 07:01:23 -07:00
Teknium bd81bf0327 chore(contributors): map tky.juani@gmail.com -> JuaniLezcano 2026-09-02 07:00:13 -07:00
Teknium ee0e234a2c fix(gateway): discover and reload MCP servers per profile under multiplex
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.

- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
  served profile inside `_profile_runtime_scope`, carried into the
  executor via `copy_context()` (same shape as
  `_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
  caller (e.g. button-confirm callback) did not; shut down / rediscover /
  report only that profile's servers; refresh only that profile's cached
  agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
  `_server_scope_keys` ownership map; leaves the shared MCP loop running
  while other profiles' servers are live. Unscoped call keeps the full
  historical behavior.
- MCP tools register into the owning profile's registry overlay
  (`registry.register(scope=...)`), and `registry.deregister()` gains a
  matching `scope=` kwarg. Plugin callers still cannot name another
  profile's scope; the plugin-vs-global guard is unchanged for them.

Fixes #95518

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium dc7e1b7ab9 fix(webhook): load URL-resolved profile's skills under multiplex
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.

- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
  when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
  otherwise, same helper the runner uses) and wrap the script / render /
  skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
  `SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
  listed default's skills; the #88023 home-keyed cache alone could not fix
  that. Use the call-time `_skills_dir()` there and at the two other
  SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
  root, so a routed profile's absolute skill_dir was rejected by
  `skill_view` ("must be a relative path within the skills directory").
  Resolve against `_skills_dir()` — the root `skill_view` itself enforces.

Fixes #67277

Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium cd7811a7a7 fix(memory/hindsight): propagate profile scope into background threads under multiplex
Under multiplex_profiles the Hindsight provider's writer, daemon-start and
prefetch threads were spawned as bare threading.Thread, so they started with
an empty contextvars Context: no profile secret scope and no HERMES_HOME
override. get_secret() fails closed there, so the local_embedded daemon never
booted and every retain raised UnscopedSecretError, even though the spawning
thread (initialize()/sync_turn() inside the gateway's copy_context'd turn) had
the scope all along.

Spawn each thread with contextvars.copy_context().run so the child inherits
the spawner's scope + home override. No environ fallback, no re-parsed .env.
The shared hindsight-loop thread needs no wrap: coroutines submitted via
run_coroutine_threadsafe already run in the submitter's context per call.

Fixes #92608
Fixes #94933

Co-authored-by: KIAgent01 <297567825+KIAgent01@users.noreply.github.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium 4e7aa48716 fix(browser): reap idle multiplexed sessions under their owner profile scope
The inactivity janitor is one process-global thread started by whichever
profile first opens a browser, so under `gateway.multiplex_profiles` it runs
with no secret scope: `cleanup_browser` -> `is_camofox_mode` ->
`get_secret("CAMOFOX_URL")` raises UnscopedSecretError, the session entry is
never removed, and the same failure repeats every 30s while the Chromium
daemon leaks.

- `_update_session_activity` records the owning Hermes home per session;
  `_cleanup_inactive_browser_sessions` re-enters that owner's
  `set_hermes_home_override` + `build_profile_secret_scope` around each
  teardown (`_session_owner_scope`, mirroring `_profile_runtime_scope`).
  copy_context at thread spawn would pin the first profile's secrets onto
  every other profile's teardown; there is no os.environ fallthrough.
- 3 consecutive failures -> `_force_reap_browser_session`, which skips the
  failing `close` round-trips but still closes the cloud provider session
  and kills the local daemon via the shared `_release_session_resources`
  tail (extracted from `_cleanup_single_browser_session`, unchanged).
  An activity touch does not reset the failure budget.

Fixes #86402
Fixes #100738

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-02 07:00:13 -07:00
Teknium 98c27aa25c feat(webhooks): stamp the emitting profile on outbound webhook payloads
Under a multiplexed gateway every profile's outbound webhooks share one
delivery worker, so receivers could not tell which profile fired an
event. Add a top-level `profile` field to the payload, resolved at fire
time from the bound Hermes home via get_active_profile_name() ("default"
outside profiles). Documents the field in the wire-format section.

Reported by @vszgdcn8cj-ctrl.

Fixes #92674
2026-09-02 07:00:13 -07:00
chelsealong 600b9c7e7e fix(gateway): scope force-reload hook re-registration to its own home
re_register_config_hooks() cleared the entire process-global idempotence
set on every force-reload, so a profile-local plugin force-reload dropped
another live profile's ledger key without touching its still-registered
callback — the next registration call for that profile then appended a
duplicate. Scope the clear to the reloading profile's own home, and give
outbound webhooks the same force-reload restoration shell hooks already
had, since unload() wipes both from the shared _hooks dict.
2026-09-02 07:00:13 -07:00
chelsealong 29640e3b5a fix(gateway): register secondary profiles' shell hooks and outbound webhooks
Multiplex gateway startup only ever calls agent.shell_hooks/
outbound_webhooks register_from_config() once, against the
root/default profile's config, before any profile scope exists.
_start_one_profile_adapters() discovers Python plugins per profile
but never registered that profile's own declarative `hooks:` block,
so a secondary profile's shell hooks (e.g. a deny-writes gate) and
outbound webhooks silently never fire.

Load and register each profile's own config inside its
_profile_runtime_scope, and key the module-level idempotence sets in
shell_hooks.py/outbound_webhooks.py by resolved Hermes home so two
profiles configuring an identical hook/webhook both register on
their own plugin manager instead of the second being dropped as a
duplicate of the first.

Fixes #92672
2026-09-02 07:00:13 -07:00
Teknium ad8b9e04c8 chore(contributors): map gobeumsu@gmail.com -> GoBeromsu 2026-09-02 06:59:11 -07:00
Teknium 7463cd1202 fix(session): fence multiplex peer-fallback recovery and profile inheritance by owner (#74285, #88381)
The per-profile store partition (17ba992108, 5ffaed6e45, 5cc3da6827)
already keeps fresh rows apart, but legacy rows written to root state.db
before the partition still sat where the default profile's peer-tuple
fallback could adopt them: a Telegram DM's tuple (chat_id == user_id, no
thread) is identical for every bot. Three residual holes, closed with the
smallest predicate that fits main's design:

- hermes_state.find_latest_gateway_session_for_peer: the fallback query
  now requires COALESCE(s.profile_name, <store owner>) = <store owner>
  (owner via SessionDB._own_profile_name). Handles NULL legacy rows and
  the single→multiplex migration case a key-namespace fence would break;
  stores outside the profile tree (no derivable owner) are unchanged.
- gateway/session._recovered_row_allowed_for_active_profile: under
  multiplexing no longer `return True` — the recovered row's agent:<ns>:
  must match the REQUESTED key's namespace (the active profile is
  meaningless when several profiles serve concurrently). Single-profile
  behavior unchanged; keyless/unnamespaced rows stay adoptable.
- hermes_state create_session parent COALESCE: profile_name inherits only
  when parent and child agree on agent:<ns>: (or either is keyless), so a
  default child forked from a sibling row is not durably mislabelled.

Co-authored-by: pcaruba <31041167+pcaruba@users.noreply.github.com>
Co-authored-by: jiangtaoliu-source <308256854+jiangtaoliu-source@users.noreply.github.com>
Co-authored-by: 69k4xmdfm2-blip <275826864+69k4xmdfm2-blip@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium bb5a1b7b80 fix(plugins): platform_actions fails closed when the active profile cannot be resolved
Follow-up to #85245: on a profile-resolution exception the facade fell
through to `_authorization_adapter(platform, None)` — the DEFAULT bot —
which is exactly the fallthrough the profile-aware lookup exists to
prevent. Return `adapter_not_registered` instead. Tests now drive the
real `GatewayAuthorizationMixin` ladder (no hand-written stand-in) and
are trimmed to one positive + one parametrized fail-closed case.

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium bedebf5f7d fix(gateway): only coerce int route ids; warn when a float/bool id can never match (#86470)
Folds the #86815 nuance into the #70815 helper: `bool` is an int subclass
and floats stringify to "123.0", both of which silently recreate the
no-match #70815 fixes — pass them through with a load-time warning.
Trims the route-id tests to one positive (int/negative-int/omitted) and
one negative (float/bool warn).

Co-authored-by: pittosporum-seu <117899760+pittosporum-seu@users.noreply.github.com>
Co-authored-by: fangliquanflq <280272527+fangliquanflq@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium ca42d7a034 fix(gateway): isolate /voice state and voice-channel input per multiplexed profile (#84872)
Voice state was keyed `<platform>:<chat_id>` with no profile namespace,
so two bots in one Discord channel shared one /voice mode; every
`_voice_input_callback` was the bare `_handle_voice_channel_input`, which
(like `_handle_voice_timeout_cleanup` and the /voice slash handler) always
picked `self.adapters[DISCORD]` — a secondary profile's voice transcripts
were dispatched through the default profile's bot.

- `_voice_key(platform, chat_id, profile=None)`: named profiles get a
  `<profile>:` prefix; default keeps the legacy shape (persisted state valid).
- `_voice_key_for_source` keys by the transport-OWNING profile
  (`_adapter_profile_for_source`), matching what `_sync_voice_mode_state_to_adapter`
  now restores per adapter via `_owner_profile`.
- `_bind_voice_input_callback` binds the capturing adapter into the
  transcript handler (functools.partial); used at primary connect,
  primary reconnect, /voice channel join, and `_configure_profile_adapter`.
- `_handle_voice_timeout_cleanup` takes the adapter it was bound to.
- /voice, join, leave and `_should_send_voice_reply` resolve the adapter via
  `_adapter_for_source` (fail-closed) instead of `self.adapters[platform]`.
- #84872: `_start_one_profile_adapters` now calls
  `_sync_voice_mode_state_to_adapter` on secondary INITIAL connect, as the
  primary path and both reconnect paths already did.

Co-authored-by: davidxyuan <124700534+davidxyuan@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium e8231f01da fix(gateway): hydrate secondary-profile secrets off-loop at startup too (#99519 class)
_run_secondary_profile_reconnect now pre-hydrates external secret sources
in a worker thread (PR #99519); _start_one_profile_adapters entered
_profile_runtime_scope on the event loop three times per profile, each
running the same synchronous network-bound hydration under
_SECRET_SOURCE_CACHE_LOCK. Hydrate once via asyncio.to_thread and enter
every scope in that method with hydrate_secrets=False.

The reconnect test is parametrized over both entry points; the startup
reconnect handoff waits are deadline-based since the runner now hops to a
worker thread before publishing the replacement adapter.

Co-authored-by: GoBeromsu <37897508+GoBeromsu@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
pierrenode 956f493cf1 fix(plugins): route platform_actions through profile-aware adapter resolution
PlatformActions._resolve_adapter() unconditionally read
runner.adapters — the default profile's adapter registry — with no
awareness of multiplex/Team-Gateway secondary profiles. Every other
adapter-resolution path in this codebase (GatewayAuthorizationMixin.
_authorization_adapter, and the plugin message-injection path that
uses it) is careful about this split: a secondary profile's adapters
live in runner._profile_adapters[profile], and a stamped secondary
profile with no registry entry must fail closed rather than fall back
to the default profile's adapter of the same platform — falling back
would send replies out the wrong bot.

ctx.platform_actions had none of that: a plugin scoped to one profile
calling add_reaction/set_thread_title in a Team-Gateway deployment
would silently act through the DEFAULT profile's adapter/bot identity
instead of its own profile's, whenever the default profile also ran
that platform.

Route _resolve_adapter() through the same _authorization_adapter
lookup, resolving the calling profile via
hermes_cli.profiles.get_active_profile_name() (which reflects the
per-task HERMES_HOME override multiplex profiles already propagate).
Falls back to the old bare runner.adapters lookup only when the
runner predates _authorization_adapter (defensive, not expected in
practice).
2026-09-02 06:59:11 -07:00
Fangliquan df50b5287e fix(gateway): coerce profile route IDs to str for YAML ints 2026-09-02 06:59:11 -07:00
gobeumsu 04836886fa fix(gateway): run secondary-profile adapter auth setup off the event loop 2026-09-02 06:59:11 -07:00
Teknium b1193b27e2 test(desktop): corrupt-exe integrity test follows stage-and-swap — corrupt pack lands in staging, live app untouched
The Windows-only test still modelled the old in-place pack (corrupt exe already
at release/win-unpacked + a .bak to roll back to). Under stage-and-swap the pack
writes into -c.directories.output=<staging>; the integrity gate runs there and a
failure discards staging while the live tree is never touched. Verified on Linux
with sys.platform patched to win32 inside the test; sabotage (skip the discard)
fails it.
2026-09-02 06:53:20 -07:00
Teknium 8ecf19b328 chore: map contributor deathxdefeat@users.noreply.github.com -> @deathxdefeat 2026-09-02 06:53:20 -07:00
Teknium 1398c0f5ca fix(update): stage-and-swap the Desktop rebuild so a failed pack never removes the working app
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).

Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.

- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
  `_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
  retry purge and the integrity self-heal only ever clear the staging tree,
  never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.

Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.

Closes #86443

Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
2026-09-02 06:53:20 -07:00
Teknium 87ad7aa0b0 fix(gateway): route multiplexed profile logs to their own logs/ dir
setup_logging(mode="gateway") binds agent.log / errors.log / gateway.log
to the launch home, so under multiplex_profiles every secondary profile's
records (emitted inside _profile_runtime_scope) landed in the DEFAULT
profile's files. Main already has the routing primitives from #99440
(record.hermes_home factory + _ProfileRoutingFileHandler +
enable_profile_log_routing) but only the Desktop cron ticker used them.
Enable them at gateway startup for the served profile set; single-profile
gateways are untouched (routing is a no-op below two homes).

Supersedes #84954, which introduced a parallel routing mechanism; the
cross-profile isolation test is translated from its suite.

Co-authored-by: Michał Dziwisz <michal@dziwisz.net>
2026-09-02 06:48:31 -07:00
fangliquanflq 04c640a183 fix(gateway): honor secondary adapters under profile runtime scope
Multiplex turns enter _profile_runtime_scope, so get_active_profile_name()
equals the secondary profile and the old authz lookup consulted empty
self.adapters. Prefer _profile_adapters before the active-profile shortcut.
2026-09-02 06:48:31 -07:00
fangliquanflq 8167cfa4b6 fix(a2a): scope multiplexed peer authorization
ThreadingHTTPServer request threads do not inherit the gateway's profile
ContextVars, so security.authenticate()/is_trusted_peer() read the
process-global A2A_* env on every request — every secondary profile's
listener authenticated against the default profile's tokens. Capture an
immutable A2ASecurityContext at adapter construction (which runs inside
_profile_runtime_scope for secondary profiles) and have the request
handler consult it instead of re-reading env per request.

Salvage note: the original `get_secret() / except UnscopedSecretError:
os.getenv()` fallthrough in _startup_env was replaced by _profile_scoped()
gating (Buzz/Raft pattern) — inside a secondary profile's scope the scope
is authoritative and a miss never falls through to os.environ.

Salvaged from #80956.
2026-09-02 06:48:31 -07:00
Teknium 64bbbaa896 fix(gateway): honor explicit msgraph_webhook disable under env override
Sibling of the webhook branch fixed in the previous commit (#85637): the
MSGRAPH_WEBHOOK env-override branch force-set enabled=True whenever
MSGRAPH_WEBHOOK_ENABLED was truthy, ignoring an explicit
`platforms.msgraph_webhook.enabled: false` in config.yaml. Apply the same
_enabled_explicit guard; port/client_state extras still wire through.
2026-09-02 06:48:31 -07:00
webtecnica dd666d2418 fix(gateway): honor explicit webhook disable in _apply_env_overrides
The webhook-specific env block set config.platforms[Platform.WEBHOOK].enabled
= True directly whenever WEBHOOK_ENABLED was truthy, without honoring the
_enabled_explicit marker that the YAML merge sets for profiles that pin
platforms.webhook.enabled: false. In multiplex mode a secondary profile that
explicitly disables webhook still inherits the process-level WEBHOOK_ENABLED
(get_secret() falls back to the default profile's .env), so the env var
force-enabled the listener and tripped the MultiplexConfigError check.

Apply the same guard used by the api_server env block (and the generic
_enable_from_env path): pop the _enabled_explicit marker and only set
enabled = True when the platform was not explicitly disabled.

Closes #85637
2026-09-02 06:48:31 -07:00
liuhao1024 7d509657a8 fix(tts): resolve default output dir from the active profile
DEFAULT_OUTPUT_DIR was resolved once at import time, so long-lived
multi-profile runtimes (dashboard console, TUI/Desktop backend, cron,
kanban workers) kept writing synthesized audio into the launch
profile's cache/audio instead of the requesting profile's (#98749).

Same bug class and fix as skills_tool (f8723c478) and skills_sync
(#65828): keep the legacy module attribute for tests and external
patchers, but re-resolve from the live profile-scoped HERMES_HOME on
every synthesis call.
2026-09-02 06:48:31 -07:00
liuhao1024 39fca697ae fix(tools): log boot-time UnscopedSecretError probes at debug, not warning
With multiplexing on, check_fns evaluated before any profile secret scope
exists fail closed by design: get_secret raises UnscopedSecretError and the
tool re-probes on the first scoped turn. _run_check_fn_uncached logged that
expected signal like a crashed check_fn (WARNING + exc_info), so every
multiplexed gateway start printed three full tracebacks that drowned real
check_fn failures.

Split the handler: an unscoped read reported while the profile cache scope
was unresolved logs one debug line without a traceback; the same error with
the scope resolved is a genuinely lost scope and keeps the loud
warning + traceback.

Fixes #100697
2026-09-02 06:48:31 -07:00
nftpoetrist 79732c6452 fix(a2a): scope multiplex secondary-profile construction, not shared env
A2AAdapter.__init__ / _default_agent_name / _load_served_agents read
A2A_PORT, A2A_AGENT_NAME and A2A_AGENT_DESCRIPTION from raw os.environ,
so a secondary multiplex profile borrowed the default profile's port and
Agent Card identity. Skip the env read when constructed inside a
secondary profile's scope (_profile_scoped(), #98738 pattern) and fall
to config.extra / module defaults instead. The default profile keeps its
unscoped env precedence.

Salvaged from #100382 (tests trimmed to two).
2026-09-02 06:48:31 -07:00
nftpoetrist 504dd1f24f fix(raft): scope RAFT_PROFILE resolution to the active multiplex profile
_spawn_bridge, _env_enablement and register()'s platform_hint read
RAFT_PROFILE via raw os.environ. Under a multiplexed secondary profile
os.environ holds the DEFAULT profile's bridged value, so the bridge
subprocess / CLI hint pointed at another profile's Raft identity.
Resolve through get_secret() only when running inside a secondary
profile's scope (Buzz/SimpleX _profile_scoped() pattern, #98738); the
default profile keeps its unscoped os.environ read. No fallthrough to
os.environ after a scoped miss.

Salvaged from #100392 (tests trimmed to two).
2026-09-02 06:48:31 -07:00
Teknium d9b051a9d6 fix(desktop): pin the project '+' new chat to the profile its tree is shown under
workspace-session-target.ts never set $newChatProfile, so a new session
started from a project's "+" reached desktopSessionCreateParams with no
intent and fell back to $activeGatewayProfile — which an in-flight profile
swap can move between the click and Send, landing session.create on the
wrong backend (#79005 flaw 3, second path). Pin the tree's profile
(projectProfile()) via the same intent write newSessionInProfile uses.

Fixes #79005
2026-09-02 06:47:55 -07:00