Commit Graph

3089 Commits

Author SHA1 Message Date
Teknium 92d60de4ae refactor(gateway): unify hosted-room coercers/JSON/clock/connect into hosted_rooms_common; fold event insert + row helpers 2026-09-02 18:23:47 -07:00
Teknium 6360d3d9c4 Merge branch 'simp/r2-gw-rest' into simp/integration2
# Conflicts:
#	tests/gateway/test_48031_model_switch_after_auto_reset.py
2026-09-02 17:06:54 -07:00
Teknium d48a0e604f Merge branch 'simp/r2-gw-rest-config' into simp/integration2 2026-09-02 17:05:22 -07:00
Teknium b80dc57a09 Merge branch 'simp/r2-gwrun' into simp/integration2 2026-09-02 17:05:04 -07:00
Teknium ba7599c953 Merge branch 'simp/r2-gw-rest-config' into simp/r2-gw-rest 2026-09-02 16:55:01 -07:00
Teknium 4ad1c6c1be refactor(gateway): drop dead GatewayRunner._has_setup_skill (no production caller; tests only stub it) 2026-09-02 16:26:11 -07:00
Teknium 6bbb13d625 refactor(gateway/slash): extract introspection commands into gateway/slash_commands_status.py (GatewayStatusCommandsMixin) 2026-09-02 16:25:35 -07:00
Teknium c40645ea2f test(gateway): repoint two source-inspection/logger-patch tests to run_turn_runner / run_voice 2026-09-02 16:24:25 -07:00
Teknium 66366d3dab refactor(gateway): split GatewayRunner into 13 run_* mixins + TurnRunner module (run.py 31332 -> 6957)
AST-driven, body-identical move of 359 GatewayRunner methods into cohesive
mixin modules (gateway/run_{voice,adapters,topics,turn,shutdown,busy,
config_loaders,startup,watchers,notifications,inbound,goals,agent_cache}.py)
plus TurnRunner -> gateway/run_turn_runner.py. run.py-internal symbols are
imported lazily inside method bodies so patch('gateway.run.X') keeps
intercepting; neutral deps are top-level; logger name stays 'gateway.run'.
_UNSET moved to leaf gateway/run_common.py (def-time default-arg sentinel).
Whole-module inspect.getsource(gateway_run) AST-walker tests repointed to
the module that now holds the walked code.
2026-09-02 16:11:50 -07:00
Teknium 8ba0350b31 refactor(gateway/slash): extract model-route commands into gateway/slash_commands_model.py (GatewayModelCommandsMixin) 2026-09-02 15:59:06 -07:00
Teknium eef25a70be fix(integration): repoint telegram webhook-secret polling test at _start_webhook_mode/_start_polling_mode
The adapters branch extracted connect()'s webhook/polling branches into methods;
the guard is intact in _start_webhook_mode. The source-inspection test now checks
the two method bodies by AST instead of the old inline if/else regex.
2026-09-02 15:41:11 -07:00
Teknium 99164fca16 refactor(gateway): split GatewayRunner._handle_message into _hm_* stage helpers
1361 -> 186 LOC orchestrator; 11 contiguous helpers (ingress/auth admit, estop gate,
update-prompt / clarify / slash-confirm replies, stale-agent eviction, running-session
fast-path, command resolve+hooks, canonical dispatch, quick+plugin commands, skill
slash rewrite). Bodies are verbatim; early returns become Optional / (handled, result)
tuples the orchestrator returns as-is. Largest function in the region: 195 LOC.
2026-09-02 15:40:08 -07:00
Teknium ff67ee47ff refactor(gateway/config): extract env overrides to gateway/config_env.py as a _Cred/_ENV_STEPS table 2026-09-02 15:38:08 -07:00
Teknium 09adb5e2bd refactor(gateway/slash): unify cached-agent lookup, session-db reply, approval delivery, checkpoint mgr, approval setters; table-drive /busy /voice /rollback /fast 2026-09-02 15:32:49 -07:00
Teknium 01b523ee96 Merge branch 'simp/adapters' into simp/integration 2026-09-02 14:19:41 -07:00
Teknium 1593bee0c3 Merge branch 'simp/cron' into simp/integration 2026-09-02 14:19:40 -07:00
Teknium a07dceb01f refactor(adapters/small_group): 5147->3565; line/email/ntfy/homeassistant/sms send-family and standalone-send dedupe, god-method extraction, dead find_pending_for_chat removed 2026-09-02 14:06:48 -07:00
Teknium 6e5b084b8b refactor(adapters/discord): 11250->7733; _HermesView base for 6 component views (1074-line factory -> 634), _send_prompt/_resolve_channel/_send_url_media/_send_local_file unification, opus loading and setup helpers extracted, dead _discord_allow_any_attachment removed 2026-09-02 14:06:32 -07:00
Teknium cc11c6bfdd refactor(adapters/telegram): 11957->8819; unify send/edit retry classification, media caching, auth gates, model-picker branches, polling recovery; compact docs 2026-09-02 14:06:32 -07:00
Teknium 581d97e545 refactor(adapters/api_server): 10589->8871; extract OpenAI-compatible routes into api_server_openai_routes.py, dedupe SSE/response builders, drop dead idempotency acknowledge path 2026-09-02 14:06:32 -07:00
Teknium 659c6b9fff refactor(gateway): simplify session, status, stream_consumer, kanban_watchers, relay adapter, hosted-room driver/discussion/replicas, shutdown/lifecycle, pairing, channel_directory, control_socket and small modules 2026-09-02 13:31:54 -07:00
Teknium 75bafc197a refactor(plugins): unify provider registration and unload paths in hermes_cli/plugins.py
hermes_cli/plugins.py (7193 -> 4961):
- _register_scoped_provider: one body for the 8 scope-keyed register_*_provider methods;
  _track_callback/_track_mapping_entry unify manager-mapping lease tracking (4 sites);
  _track_scoped_registration for the ownership ledger. Public method names, signatures,
  return values and warning strings unchanged.
- Manifest v2 type checks table-driven (_manifest_field_of_type); _unload_scoped split into
  _unload_target_keys + _reset_after_unload_all; _gate_manifest/_record_placeholder out of
  _discover_and_load_inner; _track_tool_override_policy/_attribute_registrations out of
  _load_plugin_scoped; bounded hook worker lifted into _run_hook_callback_bounded;
  prompt-section rendering extracted; resolve_pre_tool_block delegates to
  _dispatch_pre_tool_call_hooks; _evict_modules (3 sites); _remove_name_if_unowned and
  _nowait_plugin_set unify unowned-name cleanup and *_nowait probes.
- Dead (zero references outside their own test): unload_plugins, has_portable_mcp_servers,
  _classify_entrypoint_kind, _reset_event_bus, get_plugin_subscriptions,
  get_telegram_handler_factories, pass-through _restore_* helpers; _env_enabled is now an alias
  of utils.env_var_enabled (plugins/memory still imports it).
- Docstrings/comments compacted; contract sentences and rationale kept (multi-profile ledger
  keying, persistent auth-provider registration, capability declaration is not a grant).
2026-09-02 13:31:44 -07:00
Teknium 3c7069bdcb refactor(gateway): dead-code removal, helper unification, defensive-layer collapse and rationale-preserving comment compaction across 40 modules
authz_mixin, browser_control_broker, delivery, delivery_ledger, display_config, drain_control,
hosted_room_links/peer/policy_checkpoint, hosted_rooms, platform_registry, relay/__init__,
relay/ws_transport, run.py and slash_commands.py (comments), session_context, session_state,
streaming_tts_consumer, turn_lease and small modules.

- HostedRoomPolicyCheckpoint._apply_event -> per-kind handler table
- WebSocketRelayTransport._handle_frame -> frame-handler table
- GatewayAuthorizationMixin: unified adapter setting/flag/extra readers
- dead symbols removed (verified zero references): RoomLinkProbe/select_room_link,
  relay_bot_username, is_restart_loop_tripped, debug_rows, DeadTargetRegistry.all_dead,
  BrowserControlBroker.detach_owner/_prune_tickets, StreamingTTSConsumer.started/_enqueue_done/
  _iter_stream_chunks/_next_stream_chunk, RecoverableHandleCache.status_for, _auth_env,
  _copy_default_catalog, _parse_timestamp_prefix, _present_* helpers, _send_result_error_kind,
  _truthy_env, SessionFieldView/TurnLeaseTokenView dunder shims, and their orphaned tests.
- lost WHY/invariant text from the earlier compaction restored compactly (541 hunks audited)
2026-09-02 13:30:50 -07:00
Teknium 1cb3ab6173 fix(gateway): /goal gate add now requires an explicitly configured admin (salvage #91677)
A goal quality gate is a shell command persisted by `/goal gate add` and
later executed with `subprocess.run(shell=True)` at every goal turn
boundary (run_gate), with no approval prompt. On the messaging gateway,
slash access is backward-compatible: with no `allow_admin_from` list
configured (the default), every *allowed* chat user is treated as
unrestricted. So an allowed but non-admin remote sender could add an
arbitrary shell command and get authenticated RCE as the Hermes process
account.

Gate ONLY the shell-creating operation (`gate add`) behind the existing
fail-closed explicit-admin check (`_resume_caller_is_admin`, the same one
that guards cross-origin `/resume`). `gate list` / `remove` / `clear`
stay open so a non-admin can still inspect and recover. CLI/TUI/Desktop
`gate add` is local (the user already has host access) and is unchanged.

Salvaged from #91677 by @unsupportedpastels — narrowed to the single
shell-creating op and reusing the existing admin helper rather than
renaming it. Contributor's regression tests preserved.

Co-authored-by: unsupportedpastels <theoldwizard123@pm.me>
2026-09-02 10:55:57 -07:00
Ayush Nangia 02103a2104 fix(gateway): drain detached hygiene workers (#98973) 2026-09-02 21:54:01 +05:30
joaomarcos 11942f6de5 fix(gateway): fail closed when a shutdown worker survives the quiesce budget
andrexibiza's review on #101118 pointed out that the timeout branch
still ran the SessionDB close/checkpoint even when _shutdown_executor()
reported a live worker -- the exact sequence that produces the
wrong-page-number corruption in #101093. The close block now only runs
when _exec_live == 0; a surviving worker skips it entirely and leaves
the handle open for SQLite to recover from its own WAL on next open.

Adds test_stuck_worker_skips_the_session_db_close to prove the converse
of the existing ordering test: a worker that outlives the budget must
never be raced by close().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019JujDvCo2vfpEiizidAS2U
2026-09-02 20:43:06 +05:30
joaomarcos af52474aba fix(gateway): quiesce the thread pool before closing state.db at shutdown
`_shutdown_executor()` ran *after* the SessionDB close block in `_stop_impl`,
and it never waited. That left two ways for blocking DB work to outlive
`SessionDB.close()`:

  (a) `_executor_closing` was still False during the close, so a coroutine
      reaching `_run_in_executor_with_context` minted a brand-new pool and ran
      more blocking DB work against handles that had just been closed;
  (b) `cancel_futures` only drops work that has not started, and cancelling
      `self._background_tasks` does not stop the worker thread behind a
      `run_in_executor` future that is already running.

`SessionDB.close()` checkpoints the WAL and lets SQLite unlink the sidecar. A
write that lands after it silently reopens the handle (#94736) and mints a
fresh WAL generation behind that checkpoint, so teardown checkpoints the same
file a second time from a connection the shutdown log never accounts for --
the close-time page-write damage in #101093 and the split WAL generation in
#101064.

The quiesce now runs before the close and waits for the running workers. The
wait is bounded by `_EXECUTOR_QUIESCE_TIMEOUT` (2s) and clamped to what is left
of the shutdown watchdog leash minus a second for the close itself, so a stuck
worker can never cost the post-close cleanup window (#82161). Workers still
alive after the budget are logged as a warning instead of being waited on.

`_shutdown_executor()` keeps its no-argument fire-and-forget contract and now
returns the number of workers still running.

Refs #101093

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DGpGPnvz5FFFH999i59Xeb
2026-09-02 20:43:06 +05:30
Victor Kyriazakos c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium febd2af391 fix(gateway): skip credential-less WhatsApp on secondary multiplex profiles
_start_one_profile_adapters skipped only Platform.RELAY as shared
process-level ingress. WhatsApp is the same shape: the bridge is one
authenticated session tied to a single phone number, so a secondary
profile has no credential of its own to bring; constructing an adapter
for it only produced a connect/retry loop that stalled startup for every
profile queued behind it. Treat WhatsApp like Relay -- the active profile
owns the connection and route-stamped source.profile fans inbound turns
out to secondary profiles.

Salvage of #69042 (narrowed by its author to this one behavioral line);
test re-expressed on the current secondary-startup fixtures.

Co-authored-by: sshawn <28279366+lsshawn@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
Teknium 001b8abbd4 fix(matrix): pin the E2EE crypto store per profile at connect(), not import
The multiplex gateway imports plugins/platforms/matrix/adapter.py once, so
the module-level _STORE_DIR/_CRYPTO_DB_PATH resolved against the root
HERMES_HOME for every profile: all bots' Olm identities landed in one
crypto.db and inbound E2EE failed with "no session found" (#89168).

connect() runs inside _profile_runtime_scope, so resolve the store dir
there via get_hermes_dir (honors the context-local HERMES_HOME) and cache
it on the instance -- diagnostics and error-log paths read outside the
scope then still report the store actually in use. Mirrors the
pairing-store fix (a6397c379).

Salvage of #89169 (per-call resolvers collapsed into one cached resolve;
dead `_CRYPTO_DB_PATH = None` alias dropped -- no external importers).
Also routes the last raw MATRIX_HOMESERVER read in check_matrix_requirements
through _startup_env_secret like its token/password neighbours (#69943).

Fixes #89168

Co-authored-by: Michael Short <18595461+mjshorty@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Teknium 9be1168cd3 fix(wecom): scope WECOM_WEBSOCKET_URL like its neighbours
Same class as #100627's WECOM_BOT_ID: the one remaining raw os.getenv in
WeComAdapter.__init__ let a secondary multiplex profile pick up the
default profile's bridged websocket URL. Route it through
_get_scoped_secret; folded into the existing scoped-miss test.
2026-09-02 07:01:23 -07:00
Teknium 2bcbdb61a7 fix(simplex): scope SIMPLEX_* reads to the active profile under multiplexing
SimplexAdapter.__init__ (auto_accept, group_allowed), the registry gates
check_requirements/validate_config/is_connected, _env_enablement and
_standalone_send all read SIMPLEX_* via raw os.getenv. Under
gateway.multiplex_profiles those paths run inside a secondary profile's
scope where os.environ holds the DEFAULT profile's YAML-to-env bridge
output -- so a secondary profile that never configured SimpleX was
auto-enabled on the default's daemon URL and inherited its group
allowlist / auto-accept setting.

Route every read through the module-local `_get_scoped_secret` wrapper
(get_secret; UnscopedSecretError -> os.getenv for the default profile,
which constructs unscoped) -- the same helper the IRC/ntfy/Photon/
Mattermost siblings use. Unlike the extra-only `_scoped_platform_setting`
shape proposed in #100241, this honors BOTH the secondary profile's own
.env (the scope) and its config.yaml extra, and needs no config.yaml
re-read in check_requirements.

Rewrite of #100241.

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
nftpoetrist 07ee457a21 fix(wecom): scope WECOM_BOT_ID reads to the active profile under multiplexing
WeComAdapter.__init__ read WECOM_BOT_ID via a raw os.getenv() call, while
the immediately adjacent line for WECOM_SECRET already used the module's
_get_scoped_secret() helper. Under gateway.multiplex_profiles, a secondary
profile's adapter is constructed inside a scoped context where os.environ
still holds the DEFAULT profile's env-bridge output -- so a secondary
profile's bot would silently connect using the default profile's bot_id
while (correctly) using its own secret, or vice versa on a scope miss.

Switch the bot_id read to _get_scoped_secret(), matching the sibling
_secret/_dm_policy/_group_policy/allow_from reads in the same __init__
that were already migrated in #76664/#93545. _standalone_send's
out-of-process fallback branch constructs a fresh WeComAdapter(pconfig)
and therefore inherits this fix automatically -- no separate change
needed there.

Adds two regression tests to the existing TestWeComAdapterAuthzScope
class (already covering dm_policy/allow_from scoping per #93522),
mirroring its established fixture/assertion style. Mutation-verified:
both fail against the pre-fix code (asserting the default profile's
bot_id leaks into a secondary profile's scope) and pass with the fix.
2026-09-02 07:01:23 -07:00
nftpoetrist 56d869d2d9 fix(mattermost): scope url/reply_mode/require_mention/free_response_channels/allowed_channels to the active profile under multiplexing
MattermostAdapter.__init__, validate_mattermost_config, _standalone_send,
and _handle_ws_event's mention-gating block all read MATTERMOST_URL/
MATTERMOST_REPLY_MODE/MATTERMOST_REQUIRE_MENTION/
MATTERMOST_FREE_RESPONSE_CHANNELS/MATTERMOST_ALLOWED_CHANNELS via raw
os.getenv -- only MATTERMOST_TOKEN was already scoped via
_get_scoped_secret. _apply_yaml_config additionally wrote
MATTERMOST_REQUIRE_MENTION/MATTERMOST_FREE_RESPONSE_CHANNELS/
MATTERMOST_ALLOWED_CHANNELS into the process-global os.environ
unconditionally (guarded only by `not os.getenv(...)`, first-writer-wins),
the same apply_yaml_config_fn bug class already fixed for the
Discord/Telegram/WhatsApp/DingTalk adapters in this series.

Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's
env-bridge output. A secondary profile with its own (or no) Mattermost
config could silently connect to the default profile's server, thread
its replies per the default profile's reply_mode, or -- since
_handle_ws_event's mention-gating block runs on every LIVE inbound
message, not just at construction -- have its require_mention/
free_response_channels/allowed_channels decisions driven by the default
profile's settings for the adapter's entire runtime lifetime.

Fix, mirroring the WhatsApp/DingTalk apply_yaml_config_fn pattern:
- Add _profile_scoped_config_load() (same helper as DingTalk).
- Rewrite _apply_yaml_config to skip the env-bridge write under a
  multiplexed secondary profile's scope, and instead return the YAML
  values as a dict merged into this profile's own PlatformConfig.extra.
- Make require_mention/free_response_channels read extra first (matching
  the existing allowed_channels precedent), falling back to
  _get_scoped_secret() instead of raw os.getenv when extra is absent --
  fixing a residual gap the DingTalk fix (#100615, this series' item 6)
  left in its own analogous extra-first-with-raw-fallback read sites
  (_dingtalk_require_mention et al. still fall back to bare os.getenv).
- Switch __init__'s url/reply_mode, validate_mattermost_config's url, and
  _standalone_send's url to _get_scoped_secret().
- Leave check_mattermost_requirements() (no longer reads any MATTERMOST_*
  var on current main -- just an aiohttp-importability probe) and
  _is_connected() (already scope-aware via hermes_cli.gateway.get_env_value,
  which itself routes through agent.secret_scope.get_secret) untouched.

Adds a new TestMultiplexProfileScope class to tests/gateway/test_mattermost.py
(7 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope, plus two
tests exercising _apply_yaml_config's new seeded-dict return directly.
Mutation-verified: stashed the production fix and confirmed 5 of 7 new
tests fail against pre-fix code (the other 2 are non-differentiating
regression guards -- extra-wins-over-env and unscoped-default-profile-
precedence -- which correctly pass either way). Restored the fix;
all 30 tests in the file, the plugin-setup test, and the full 75-test
tests/gateway/test_adapter_startup_secret_scope.py suite pass.
2026-09-02 07:01:23 -07:00
nftpoetrist 327e9043e9 fix(ntfy): scope server/topic/publish_topic reads to the active profile under multiplexing
NtfyAdapter.__init__, _env_enablement, check_requirements, validate_config,
is_connected, and _standalone_send all read NTFY_SERVER_URL/NTFY_TOPIC/
NTFY_PUBLISH_TOPIC/NTFY_MARKDOWN/NTFY_HOME_CHANNEL(_NAME) via raw os.getenv
-- only NTFY_TOKEN already went through the module's _get_scoped_secret
helper. Under gateway.multiplex_profiles, env_enablement_fn/check_fn/
is_connected all run inside the registry-enablement loop in
load_gateway_config() (confirmed in gateway/config.py, lines ~2704-2820,
inside _profile_runtime_scope for secondary profiles), and adapter
construction likewise runs scoped -- so os.environ there still holds the
DEFAULT profile's env-bridge output. A secondary profile with its own (or
no) ntfy topic configured could silently:

- get auto-enabled via _env_enablement()/is_connected() using the default
  profile's topic, even though it never configured ntfy itself
- have its adapter subscribe to / publish on the default profile's topic
  and server instead of (or in addition to) its own
- deliver cron/send_message_tool messages via _standalone_send to the
  wrong topic

Switch every raw NTFY_* read (except the two secret-material fields
already scoped: NTFY_TOKEN) to _get_scoped_secret(), matching the
established helper already defined in this module and used for
NTFY_TOKEN, and the same pattern applied to the sibling LINE/DingTalk/
Teams/SMS/WeCom adapters in this series.

Adds a new TestMultiplexProfileScope class to tests/gateway/test_ntfy_plugin.py
(7 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope. Mutation-
verified: stashed the production fix and confirmed 5 of the 7 new tests
fail against pre-fix code (the other 2 are non-differentiating regression
guards -- extra-wins-over-env and unscoped-default-profile-precedence --
which correctly pass either way); restored the fix and confirmed all 37
tests in the file, plus the file's 5 parametrized _get_scoped_secret tests
in test_adapter_startup_secret_scope.py, pass.
2026-09-02 07:01:23 -07:00
nftpoetrist 05aa749977 fix(irc): scope server/port/nickname/channel/use_tls reads to the active profile under multiplexing
IRCAdapter.__init__, check_requirements, validate_config, is_connected,
_env_enablement, and _standalone_send all read IRC_SERVER/IRC_PORT/
IRC_NICKNAME/IRC_CHANNEL/IRC_USE_TLS via raw os.getenv -- only
IRC_SERVER_PASSWORD/IRC_NICKSERV_PASSWORD already went through the
module's _get_scoped_secret helper. Under gateway.multiplex_profiles,
env_enablement_fn/check_fn/is_connected all run inside the registry-
enablement loop in load_gateway_config() (gateway/config.py, ~lines
2704-2820), scoped for secondary profiles via _profile_runtime_scope,
and adapter construction runs scoped the same way -- so os.environ there
still holds the DEFAULT profile's env-bridge output.

Notably, __init__'s original `os.getenv("IRC_SERVER") or extra.get(...)`
ordering let a raw env read override even an explicitly configured
config.yaml extra -- a secondary profile that set its own server/channel
via config.yaml extra would still silently connect to the default
profile's IRC server/channel/nick if the default profile bridged its own
config to env (which it always does under multiplex). This is a stronger
variant of the same bug fixed for the sibling LINE/DingTalk/Teams/SMS/
WeCom/ntfy adapters in this series -- there, extra already won because
of the `extra.get(...) or os.getenv(...)` order.

Switch every raw IRC_* read (except IRC_SERVER_PASSWORD/
IRC_NICKSERV_PASSWORD, already scoped) to _get_scoped_secret(), matching
the module's existing helper. Also collapses a double os.getenv("IRC_USE_TLS")
read in __init__ into a single _get_scoped_secret() call (same behavior,
one scope lookup instead of two).

Adds a new TestMultiplexProfileScope class to tests/gateway/test_irc_adapter.py
(6 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope. Mutation-
verified: stashed the production fix and confirmed 5 of 6 new tests fail
against pre-fix code -- including the "extra wins" test, since IRC's
original env-first ordering meant even an explicit extra config was not
a safe differentiator boundary before the fix (only the DEFAULT-profile-
unscoped-precedence test is a non-differentiating regression guard that
correctly passes either way). Restored the fix; all 23 tests in the file
pass, plus the file's 5 parametrized _get_scoped_secret tests in
test_adapter_startup_secret_scope.py.
2026-09-02 07:01:23 -07:00
Teknium ee0e234a2c fix(gateway): discover and reload MCP servers per profile under multiplex
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.

- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
  served profile inside `_profile_runtime_scope`, carried into the
  executor via `copy_context()` (same shape as
  `_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
  caller (e.g. button-confirm callback) did not; shut down / rediscover /
  report only that profile's servers; refresh only that profile's cached
  agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
  `_server_scope_keys` ownership map; leaves the shared MCP loop running
  while other profiles' servers are live. Unscoped call keeps the full
  historical behavior.
- MCP tools register into the owning profile's registry overlay
  (`registry.register(scope=...)`), and `registry.deregister()` gains a
  matching `scope=` kwarg. Plugin callers still cannot name another
  profile's scope; the plugin-vs-global guard is unchanged for them.

Fixes #95518

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium dc7e1b7ab9 fix(webhook): load URL-resolved profile's skills under multiplex
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.

- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
  when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
  otherwise, same helper the runner uses) and wrap the script / render /
  skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
  `SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
  listed default's skills; the #88023 home-keyed cache alone could not fix
  that. Use the call-time `_skills_dir()` there and at the two other
  SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
  root, so a routed profile's absolute skill_dir was rejected by
  `skill_view` ("must be a relative path within the skills directory").
  Resolve against `_skills_dir()` — the root `skill_view` itself enforces.

Fixes #67277

Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
chelsealong 29640e3b5a fix(gateway): register secondary profiles' shell hooks and outbound webhooks
Multiplex gateway startup only ever calls agent.shell_hooks/
outbound_webhooks register_from_config() once, against the
root/default profile's config, before any profile scope exists.
_start_one_profile_adapters() discovers Python plugins per profile
but never registered that profile's own declarative `hooks:` block,
so a secondary profile's shell hooks (e.g. a deny-writes gate) and
outbound webhooks silently never fire.

Load and register each profile's own config inside its
_profile_runtime_scope, and key the module-level idempotence sets in
shell_hooks.py/outbound_webhooks.py by resolved Hermes home so two
profiles configuring an identical hook/webhook both register on
their own plugin manager instead of the second being dropped as a
duplicate of the first.

Fixes #92672
2026-09-02 07:00:13 -07:00
Teknium 7463cd1202 fix(session): fence multiplex peer-fallback recovery and profile inheritance by owner (#74285, #88381)
The per-profile store partition (17ba992108, 5ffaed6e45, 5cc3da6827)
already keeps fresh rows apart, but legacy rows written to root state.db
before the partition still sat where the default profile's peer-tuple
fallback could adopt them: a Telegram DM's tuple (chat_id == user_id, no
thread) is identical for every bot. Three residual holes, closed with the
smallest predicate that fits main's design:

- hermes_state.find_latest_gateway_session_for_peer: the fallback query
  now requires COALESCE(s.profile_name, <store owner>) = <store owner>
  (owner via SessionDB._own_profile_name). Handles NULL legacy rows and
  the single→multiplex migration case a key-namespace fence would break;
  stores outside the profile tree (no derivable owner) are unchanged.
- gateway/session._recovered_row_allowed_for_active_profile: under
  multiplexing no longer `return True` — the recovered row's agent:<ns>:
  must match the REQUESTED key's namespace (the active profile is
  meaningless when several profiles serve concurrently). Single-profile
  behavior unchanged; keyless/unnamespaced rows stay adoptable.
- hermes_state create_session parent COALESCE: profile_name inherits only
  when parent and child agree on agent:<ns>: (or either is keyless), so a
  default child forked from a sibling row is not durably mislabelled.

Co-authored-by: pcaruba <31041167+pcaruba@users.noreply.github.com>
Co-authored-by: jiangtaoliu-source <308256854+jiangtaoliu-source@users.noreply.github.com>
Co-authored-by: 69k4xmdfm2-blip <275826864+69k4xmdfm2-blip@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium bedebf5f7d fix(gateway): only coerce int route ids; warn when a float/bool id can never match (#86470)
Folds the #86815 nuance into the #70815 helper: `bool` is an int subclass
and floats stringify to "123.0", both of which silently recreate the
no-match #70815 fixes — pass them through with a load-time warning.
Trims the route-id tests to one positive (int/negative-int/omitted) and
one negative (float/bool warn).

Co-authored-by: pittosporum-seu <117899760+pittosporum-seu@users.noreply.github.com>
Co-authored-by: fangliquanflq <280272527+fangliquanflq@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium ca42d7a034 fix(gateway): isolate /voice state and voice-channel input per multiplexed profile (#84872)
Voice state was keyed `<platform>:<chat_id>` with no profile namespace,
so two bots in one Discord channel shared one /voice mode; every
`_voice_input_callback` was the bare `_handle_voice_channel_input`, which
(like `_handle_voice_timeout_cleanup` and the /voice slash handler) always
picked `self.adapters[DISCORD]` — a secondary profile's voice transcripts
were dispatched through the default profile's bot.

- `_voice_key(platform, chat_id, profile=None)`: named profiles get a
  `<profile>:` prefix; default keeps the legacy shape (persisted state valid).
- `_voice_key_for_source` keys by the transport-OWNING profile
  (`_adapter_profile_for_source`), matching what `_sync_voice_mode_state_to_adapter`
  now restores per adapter via `_owner_profile`.
- `_bind_voice_input_callback` binds the capturing adapter into the
  transcript handler (functools.partial); used at primary connect,
  primary reconnect, /voice channel join, and `_configure_profile_adapter`.
- `_handle_voice_timeout_cleanup` takes the adapter it was bound to.
- /voice, join, leave and `_should_send_voice_reply` resolve the adapter via
  `_adapter_for_source` (fail-closed) instead of `self.adapters[platform]`.
- #84872: `_start_one_profile_adapters` now calls
  `_sync_voice_mode_state_to_adapter` on secondary INITIAL connect, as the
  primary path and both reconnect paths already did.

Co-authored-by: davidxyuan <124700534+davidxyuan@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium e8231f01da fix(gateway): hydrate secondary-profile secrets off-loop at startup too (#99519 class)
_run_secondary_profile_reconnect now pre-hydrates external secret sources
in a worker thread (PR #99519); _start_one_profile_adapters entered
_profile_runtime_scope on the event loop three times per profile, each
running the same synchronous network-bound hydration under
_SECRET_SOURCE_CACHE_LOCK. Hydrate once via asyncio.to_thread and enter
every scope in that method with hydrate_secrets=False.

The reconnect test is parametrized over both entry points; the startup
reconnect handoff waits are deadline-based since the runner now hops to a
worker thread before publishing the replacement adapter.

Co-authored-by: GoBeromsu <37897508+GoBeromsu@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Fangliquan df50b5287e fix(gateway): coerce profile route IDs to str for YAML ints 2026-09-02 06:59:11 -07:00
gobeumsu 04836886fa fix(gateway): run secondary-profile adapter auth setup off the event loop 2026-09-02 06:59:11 -07:00
Teknium 87ad7aa0b0 fix(gateway): route multiplexed profile logs to their own logs/ dir
setup_logging(mode="gateway") binds agent.log / errors.log / gateway.log
to the launch home, so under multiplex_profiles every secondary profile's
records (emitted inside _profile_runtime_scope) landed in the DEFAULT
profile's files. Main already has the routing primitives from #99440
(record.hermes_home factory + _ProfileRoutingFileHandler +
enable_profile_log_routing) but only the Desktop cron ticker used them.
Enable them at gateway startup for the served profile set; single-profile
gateways are untouched (routing is a no-op below two homes).

Supersedes #84954, which introduced a parallel routing mechanism; the
cross-profile isolation test is translated from its suite.

Co-authored-by: Michał Dziwisz <michal@dziwisz.net>
2026-09-02 06:48:31 -07:00
fangliquanflq 04c640a183 fix(gateway): honor secondary adapters under profile runtime scope
Multiplex turns enter _profile_runtime_scope, so get_active_profile_name()
equals the secondary profile and the old authz lookup consulted empty
self.adapters. Prefer _profile_adapters before the active-profile shortcut.
2026-09-02 06:48:31 -07:00
Teknium 64bbbaa896 fix(gateway): honor explicit msgraph_webhook disable under env override
Sibling of the webhook branch fixed in the previous commit (#85637): the
MSGRAPH_WEBHOOK env-override branch force-set enabled=True whenever
MSGRAPH_WEBHOOK_ENABLED was truthy, ignoring an explicit
`platforms.msgraph_webhook.enabled: false` in config.yaml. Apply the same
_enabled_explicit guard; port/client_state extras still wire through.
2026-09-02 06:48:31 -07:00
webtecnica dd666d2418 fix(gateway): honor explicit webhook disable in _apply_env_overrides
The webhook-specific env block set config.platforms[Platform.WEBHOOK].enabled
= True directly whenever WEBHOOK_ENABLED was truthy, without honoring the
_enabled_explicit marker that the YAML merge sets for profiles that pin
platforms.webhook.enabled: false. In multiplex mode a secondary profile that
explicitly disables webhook still inherits the process-level WEBHOOK_ENABLED
(get_secret() falls back to the default profile's .env), so the env var
force-enabled the listener and tripped the MultiplexConfigError check.

Apply the same guard used by the api_server env block (and the generic
_enable_from_env path): pop the _enabled_explicit marker and only set
enabled = True when the platform was not explicitly disabled.

Closes #85637
2026-09-02 06:48:31 -07:00