Commit Graph

223 Commits

Author SHA1 Message Date
Teknium 5820d0b0d5 refactor(gateway/config): extract the config.yaml phase to gateway/config_loader.py; table-driven from_dict; defensive collapse
- load_gateway_config (536 LOC) is now a thin orchestrator: legacy gateway.json
  -> config_loader.load_yaml_layer -> GatewayConfig.from_dict -> env overrides
  -> validation. The yaml phase lives in gateway/config_loader.py as small
  functions driven by tables: _TOPLEVEL_BRIDGE (23 top-level/nested
  gateway.<key> bridges with 5 fallback modes), _SHARED_KEYS (28 per-platform
  keys copied into extra, with per-platform restrictions and transforms) and
  _PORT_BRIDGE_KEYS. Logger name kept as "gateway.config".
- GatewayConfig.from_dict: shared pick()/key_label() helpers replace the
  repeated "top-level key present else nested gateway.<key>" blocks; warning
  order preserved.
- _normalize_unauthorized_dm_behavior / _normalize_notice_delivery unified into
  _normalize_choice; _ensure_platform_extra_dict -> _dict_slot (also used by
  persist_home_channel); _getenv_int removed (dead since the env pass moved to
  config_env, which has its own _int_or).
- Platform._missing_: one _add_pseudo_member helper for both branches.
- Single warning sites in _coerce_optional_positive_int and
  coerce_systemd_watchdog_seconds; _validate_gateway_config placeholder pass
  flattened; small to_dict/getter collapses. Ruff F401/SIM102 clean.
- tests/hermes_cli/test_config_read_guard.py: allowlist gateway/config_loader.py
  (same owner as gateway/config.py — the extracted load_gateway_config phase).

gateway/config.py 2684 -> 1319 LOC (-50.9%); largest function now
GatewayConfig.from_dict at 119 LOC. Resolved-config parity: 160 cells
(16 yaml fixtures x 10 env sets) byte-identical to the integration base,
including captured log records and stderr.
2026-09-02 16:48:36 -07:00
Teknium ff67ee47ff refactor(gateway/config): extract env overrides to gateway/config_env.py as a _Cred/_ENV_STEPS table 2026-09-02 15:38:08 -07:00
Teknium aed6720dab refactor(gateway/run, slash_commands): dispatch tables, helper unification and hand-reviewed comment compaction
run.py:
- built-in adapter creation: 9-branch if/elif -> _BUILTIN_ADAPTERS table
- idle slash-command routing: 35 `if canonical == ...` branches -> _gateway_idle_command_handlers()
- shared helpers: _send_command_ack (4 sites), _command_origin_for_source (2), _session_entry_for_manager
  (goal/heartbeat), _toggle_adapter_auto_tts_set (2), _load_env_or_agent_cfg_timeout (2), _float_env reuse (2),
  _resolve_session_key_or_none (3), _running_agent_ids (4), _schedule_rename_from_title_thread (2),
  _write_runtime_status_quiet (5), _AUTO_RESET_CONTEXT_NOTES/_auto_reset_reason_text
- ruff SIM102/SIM103/SIM105/SIM108/SIM118 + F401 across gateway/ (semantics re-reviewed; sqlite Row
  `.keys()` and side-effecting assignments kept)
- two hand-reviewed comment/docstring compaction passes (AST-identical, rationale kept)

slash_commands.py:
- /model: typed path and picker callback shared one 200-line commit block -> _perform_model_switch +
  _commit_model_switch
- comment/docstring compaction (AST-identical)
2026-09-02 13:31:53 -07:00
Teknium ffd628c3ee refactor(gateway/config): table-drive _apply_env_overrides via env-extra/home-channel helpers; extract plugin enable pass 2026-09-02 13:30:50 -07:00
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
Teknium 64bbbaa896 fix(gateway): honor explicit msgraph_webhook disable under env override
Sibling of the webhook branch fixed in the previous commit (#85637): the
MSGRAPH_WEBHOOK env-override branch force-set enabled=True whenever
MSGRAPH_WEBHOOK_ENABLED was truthy, ignoring an explicit
`platforms.msgraph_webhook.enabled: false` in config.yaml. Apply the same
_enabled_explicit guard; port/client_state extras still wire through.
2026-09-02 06:48:31 -07:00
webtecnica dd666d2418 fix(gateway): honor explicit webhook disable in _apply_env_overrides
The webhook-specific env block set config.platforms[Platform.WEBHOOK].enabled
= True directly whenever WEBHOOK_ENABLED was truthy, without honoring the
_enabled_explicit marker that the YAML merge sets for profiles that pin
platforms.webhook.enabled: false. In multiplex mode a secondary profile that
explicitly disables webhook still inherits the process-level WEBHOOK_ENABLED
(get_secret() falls back to the default profile's .env), so the env var
force-enabled the listener and tripped the MultiplexConfigError check.

Apply the same guard used by the api_server env block (and the generic
_enable_from_env path): pop the _enabled_explicit marker and only set
enabled = True when the platform was not explicitly disabled.

Closes #85637
2026-09-02 06:48:31 -07:00
Teknium 7ff8aae8fb gateway: warn once when an explicit platforms.<x>.enabled: false overrides env credentials
De-risking for the #48820 behaviour change: before this branch, credentials
in the environment force-enabled twelve platforms regardless of an explicit
enabled: false in config.yaml. Now that the explicit disable wins, users who
relied on the old override would see the platform go dark with no trace.

_enable_from_env (and Slack's inline copy) now emit ONE WARNING per platform
per process when the platform is explicitly disabled AND its env credentials
are present, naming the platform, the winning key
(platforms.<x>.enabled: false), the env var(s) being ignored, and the remedy.
A plain disable with no credentials, an enabled platform, and the env-only
(no YAML opinion) path stay silent; repeated config reloads do not repeat it.
_ENV_ENABLE_CREDENTIALS maps every _enable_from_env platform to its
triggering env var(s); a test pins that the map covers every routed branch.

Docs: messaging/index.md gains a 'Disabling a platform whose credentials are
still in .env' section with the exact warning text.

Live repro (real load_gateway_config on a temp HERMES_HOME with
platforms.weixin/telegram.enabled: false + WEIXIN_TOKEN/TELEGRAM_BOT_TOKEN in
env): before — both stayed disabled with zero log output; after — one
WARNING each ('Platform 'weixin' is explicitly disabled by
platforms.weixin.enabled: false ... (WEIXIN_TOKEN, WEIXIN_ACCOUNT_ID) will
NOT start its adapter ...'), none for the enabled homeassistant, none on the
second load.
2026-09-02 01:50:45 -07:00
Teknium 867e4158f0 fix(gateway): honor explicit platforms.<x>.enabled: false over env credentials (#48820)
Twelve credential-presence branches in _apply_env_overrides (weixin,
whatsapp_cloud, homeassistant, email, sms, dingtalk, feishu, wecom,
wecom_callback, bluebubbles, qqbot, yuanbao) force-set enabled = True
unconditionally, so a user's explicit `platforms.<x>.enabled: false` in
config.yaml was silently overridden whenever the platform's token/secret
lived in .env. Telegram/Discord/Slack/Signal/Matrix already routed through
_enable_from_env, which honors the `_enabled_explicit` marker written by
load_gateway_config.

Route all twelve sites through the same helper. Credentials are still wired
into the (disabled) PlatformConfig so send-only tooling keeps working —
the same contract Slack and api_server already follow.

Live repro (real load_gateway_config against a temp HERMES_HOME, yaml
`enabled: false` + creds in env): 12/13 platforms flipped to enabled=True
on main; 0/13 after the fix (telegram control unchanged).

Bug 2 of #48820. Fix direction from @JoaoMarcos44 in #48852 (surgically
reapplied on current main — the June branch no longer applies).

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-02 01:50:45 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
Kshitij Kapoor 8ee0103ea2 fix(gateway): finite-bounded watchdog knob validation + wire keys through load_gateway_config
Addresses both review findings from @egilewski on #89134:

- Non-finite values: _coerce_int now degrades int(inf) (OverflowError
  previously ABORTED gateway config loading); the clamp requires
  math.isfinite plus sane upper bounds (interval <=3600s, timeout
  <=600s, strikes <=1000), falling back to the shutdown_watchdog
  constants.
- Loader wiring: load_gateway_config builds gw_data FLAT and never
  forwarded the yaml gateway: section, so loop_watchdog* keys —
  including the PRE-EXISTING loop_watchdog bool documented in
  config_defaults — were silently ignored on the real startup path.
  Bridged with the established top-level-wins/nested-fallback pattern.

E2E: config.yaml with loop_watchdog:false + strikes:12 + interval:.inf
now yields False/12/30.0 through the real loader.
2026-08-22 20:40:23 +05:30
Kshitij Kapoor 3616145723 fix(gateway): keep loop-watchdog default at 3 strikes; dedupe constants; register knobs in config defaults
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:

- default max_strikes back to 3 everywhere (constant, dataclass,
  from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
  shutdown_watchdog DEFAULT_* constants instead of duplicating literals
  in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
  sibling gateway.loop_watchdog bool
2026-08-22 20:40:23 +05:30
devops aa08cb8ccb fix(gateway): make loop-liveness watchdog tolerant of transient reconnect stalls
The event-loop liveness watchdog (gateway.shutdown_watchdog) hard-exited with
code 75 after 3 consecutive missed probes (probe_interval=30s, timeout=10s,
max_strikes=3), i.e. ~90-120s of loop block. Telegram/Discord reconnect during
a network blip does synchronous socket I/O on the loop and can block it for
60-90s; these stalls self-recover (recurring fleet incidents on 2026-08-17
stalled cron dispatch ~21h via restart churn, kanban t_0f76430f).

Raise the default max_strikes 3->8 so a transient reconnect stall is tolerated
while a genuine multi-minute wedge still escalates, and expose the three
tolerance knobs via config.yaml (gateway.loop_watchdog_probe_interval_s /
_probe_timeout_s / _max_strikes) so operators can tune per deployment.

Refs: kanban t_70483f23
2026-08-22 20:40:23 +05:30
Ben Barclay 2ef095ce46 fix(gateway): source-neutral log wording for relay-exclusive sweep
sol-reviewer MINOR: the _enabled_explicit marker is set by config.yaml,
gateway.json, and dashboard PUTs alike, so the WARNING no longer claims
config.yaml specifically. Also names the opt-out env var in the message
so an operator seeing the WARNING knows the escape hatch.
2026-08-21 12:42:13 +10:00
Ben Barclay 477a0222f9 fix(gateway): read relay-exclusive env vars through profile secret scope
sol-reviewer IMPORTANT: the relay trigger and opt-out read os.getenv()
directly, bypassing the profile secret scope that _apply_env_overrides
uses everywhere else. Under a multiplexed gateway a profile-scoped
GATEWAY_RELAY_URL was invisible, and a process-global one leaked into
every profile, disabling direct platforms in profiles that are not
relay-fronted. Both reads now go through the scope-aware getenv.

Also folds the NIT: opt-out truthiness now uses the shared
is_truthy_value helper instead of a local truthy tuple.
2026-08-21 12:42:06 +10:00
Ben Barclay 2d3f5c1554 feat(gateway): GATEWAY_RELAY_ALLOW_DIRECT_PLATFORMS opt-out for relay-exclusive mode
Deployments that intentionally mix connector-fronted and direct ingress
can set GATEWAY_RELAY_ALLOW_DIRECT_PLATFORMS=true to keep directly-
connected messaging adapters enabled beside the relay. Unset, the
GATEWAY_RELAY_URL env stamp keeps its exclusive behavior. Like the
trigger, the opt-out is a deploy-stamp env var, not config.yaml.
2026-08-21 12:26:41 +10:00
Ben Barclay 67292ec5b8 fix(gateway): GATEWAY_RELAY_URL env stamp disables direct messaging platforms
A GATEWAY_RELAY_URL set in the process environment marks a
connector-fronted deployment where the connector owns every platform
connection. A directly-connected messaging adapter in the same process
is a second, unmanaged ingress path: it causes duplicate deliveries and
split sessions, and its live socket disarms scale-to-zero.

At the end of _apply_env_overrides, after all enablement passes, the
env stamp now disables every other enabled messaging platform:

- Explicitly-enabled platforms (config.yaml enabled: true) are disabled
  with a WARNING that names the platform.
- Credential-auto-enabled platforms are disabled with an INFO line.
- Non-messaging surfaces (local, api_server, webhook) are untouched --
  the same exclusion set as the scale-to-zero arm gate.
- gateway.relay_url in config.yaml alone (no env stamp) keeps the old
  additive behavior: relay runs beside direct adapters.
2026-08-21 12:16:06 +10:00
Trevin Chow c8f235a106 feat(gateway): allow selective multiplex profile serving 2026-08-10 22:48:24 -07:00
kshitij a658dfe509 fix: address self-review findings on the check_fn/ensure_deps_fn split
- gateway/config.py: rewrite the stale enablement-pass header comment that
  still described check_fn as 'the single source of truth for are-my-env-
  vars-set' / 'lazy-installs it' — both false under the new contract.
- teams: check_requirements docstring wrongly claimed credential checks
  (body checks only SDK/aiohttp presence); derive install_hint from the
  canonical LAZY_DEPS pins + sys.executable instead of hardcoding
  '~/.hermes/hermes-agent/venv/bin/pip' and version pins (wrong under
  HERMES_HOME overrides / profile installs; pins go stale on CVE bumps);
  connect() fatal-error hints now point at the venv pip instead of bare
  system pip (the PEP 668 trap the docs warn about).
- teams docs: drop exact version pins from the two manual-install commands
  (LAZY_DEPS is the source of truth; unpinned installs still work and the
  text can't go stale).
- hermes_cli/status.py: per-entry exception guard around check_fn so one
  raising probe can't abort the listing of all remaining plugin platforms
  (aligns with the other three call sites).
- tests: rename test_register_check_fn_is_active_lazy_installer ->
  test_register_splits_passive_probe_from_active_installer (name said the
  opposite of what it verifies).
2026-08-07 13:28:43 +05:30
kshitij 0d32607c62 fix(gateway): split check_fn (passive probe) from ensure_deps_fn (active installer)
PlatformEntry.check_fn served three contradictory roles: adapter-creation
gate, config auto-enablement gate, and status display. Plugins had to pick
one function for all three:

- Active installer as check_fn (discord/slack/telegram/matrix/dingtalk/
  feishu): every status display could pip-install SDKs as a side effect
  (the desktop 94% boot-loop class).
- Passive probe as check_fn (teams, wecom_callback): create_adapter()
  returned None before connect() could lazy-install, so the SDK never
  installed (#79812 deadlock; wecom_callback's platform.wecom_callback
  LAZY_DEPS entry was dead code).

The split makes both call sites correct by construction:

- check_fn is now contractually PASSIVE (probe only, never installs).
- New optional PlatformEntry.ensure_deps_fn is the ACTIVE installer;
  create_adapter() runs it exactly when check_fn is False — the platform
  is enabled+configured and the gateway is about to connect it.
- Config enablement keeps a configured platform whose deps are missing
  but installable; the install itself is deferred to create_adapter().
- Status surfaces (_platform_status, hermes status) read only the
  passive probe and can never trigger pip.

Migrated all lazy-installable platform plugins to the split; platforms
with no optional deps (irc/ntfy/buzz/simplex/line/a2a/...) are unchanged
— no ensure_deps_fn means a False check_fn stays a hard block.
wecom_callback gains a working installer for the first time.

Builds on @xxxigm's #79812 (both commits cherry-picked with authorship
preserved), reworking the check_fn swap into the two-field split so the
Teams fix doesn't reintroduce install-on-status.
2026-08-07 13:28:43 +05:30
teknium1 5b751dc0ad chore: remove unused imports and dead locals (ruff F401/F841 sweep)
Cleans F401 unused imports and F841 dead local assignments across
root *.py, agent/, hermes_cli/, tools/, gateway/, cron/, tui_gateway/
(tests/, plugins/, skills/ excluded).

Intentionally KEPT (false positives / test-patch surfaces):
- agent/transports/__init__.py package re-exports
- cli.py browser_connect re-exports (DEFAULT_BROWSER_CDP_URL area,
  used by tests/cli/test_cli_browser_connect.py)
- hermes_cli/main.py _prompt_auth_credentials_choice /
  _model_flow_bedrock_api_key (accessed via main_mod attr in tests)
- gateway/run.py aliased replay_cleanup + whatsapp_identity re-exports
  and _PORT_BINDING_PLATFORM_VALUES (test-referenced)
- hermes_cli/web_server.py get_running_pid (tests monkeypatch it) and
  _OAUTH_TOKEN_URL availability probe
- hermes_cli/config.py get_process_hermes_home re-export (noqa'd F811
  chain) and yaml availability-probe import
- hermes_cli/nous_subscription.py managed_nous_tools_enabled
  (tests patch hermes_cli.nous_subscription.managed_nous_tools_enabled)
- try/except ImportError availability probes (env_loader, tts_tool,
  mcp_tool, web_server anthropic OAuth block)
- tools/web_tools.py noqa F401 re-exports
- hermes_cli/setup_whatsapp_cloud.py:263 'proceed' skipped: possible
  missing-guard bug, flagged for separate review
- unused function parameters (signature changes out of scope)

Side-effect RHS calls preserved where only the binding was dead
(e.g. web_server proc = _spawn_hermes_action -> bare call).
2026-07-29 11:53:39 -07:00
Teknium 2c771be406 fix(gateway): dual-stack webhook bind for wecom/msgraph/whatsapp_cloud/teams/telegram siblings
Same class of bug as the LINE adapter (NS-603): defaulting the webhook
bind to "0.0.0.0" (or hardcoding it) binds IPv4 ONLY, so the listener
is unreachable over IPv6-only private networks such as Fly.io 6PN.

- wecom callback_adapter: DEFAULT_HOST None; config.py env seed no
  longer forces 0.0.0.0 when WECOM_CALLBACK_HOST is unset.
- msgraph_webhook: DEFAULT_HOST None; the allowed_source_cidrs
  requirement still fires for the all-interfaces default (host=None is
  treated as network-accessible).
- whatsapp_cloud: DEFAULT_WEBHOOK_HOST None.
- teams: hardcoded 0.0.0.0 TCPSite bind → _DEFAULT_HOST=None with new
  TEAMS_HOST / extra.host override (mirrors LINE_HOST pattern).
- telegram: hardcoded listen="0.0.0.0" → default "" (tornado
  bind_sockets opens one socket per address family; verified against
  PTB 22.6/tornado) with new TELEGRAM_WEBHOOK_HOST / extra.webhook_host
  override.

Explicit host overrides everywhere are preserved; empty/unset collapses
to the dual-stack default. "::" remains a bad substitute on
bindv6only=1 hosts (see LINE adapter comment).
2026-07-28 22:42:41 -07:00
Cyrus 35afa8ce06 feat(whatsapp): support inbound read receipts 2026-07-28 18:02:51 +05:30
Hermes Agent 9cd7296849 refactor(gateway): gate loop-liveness watchdog via config.yaml, drop HERMES_* env knobs
Follow-up to the salvaged #69164 commits: policy forbids introducing new
HERMES_* environment variables, so the four watchdog env knobs
(HERMES_GATEWAY_LOOP_WATCHDOG / _INTERVAL / _TIMEOUT / _STRIKES) are
replaced with a single config.yaml boolean:

  gateway:
    loop_watchdog: true   # default; false disables both guards

- gateway/config.py: new GatewayConfig.loop_watchdog field (default True),
  parsed from top-level or nested gateway: form, round-trips via
  to_dict/from_dict.
- gateway/run.py: _start_loop_liveness_guards() checks config.loop_watchdog
  before arming the floor timer + watchdog (getattr-guarded for bare
  object.__new__ runners).
- gateway/shutdown_watchdog.py: start_loop_liveness_watchdog() no longer
  reads the environment; probe interval/timeout/strikes are module
  constants (30s/10s/3 — ~90s to restart, matching the systemd watchdog
  layer's posture).
- hermes_cli/config.py: documented gateway.loop_watchdog default so
  'hermes config set gateway.loop_watchdog false' validates.
- tests: env-knob tests replaced with config-gate + round-trip tests;
  the final-strike boundary test injects its probe via max_strikes
  directly instead of patching the removed env helper.
2026-07-24 16:03:42 -07:00
Victor Kyriazakos 45a408f41a fix(gateway): deliver relay-backed homes after restart 2026-07-24 10:45:13 -07:00
kyssta-exe 53e6435908 fix(gateway): parse gateway.api_server YAML config section (#66630)
The gateway ignored  config when defined through
YAML (config.yaml). The API server only started when environment
variables like  and  were set,
even though other gateway subsections like  were
processed correctly.

Two changes in gateway/config.py's load_gateway_config():
1. Merge platform configs placed directly under gateway.*
   (e.g. gateway.api_server) via _merge_platform_map, matching
   the existing gateway.platforms.* and platforms.* merge
   paths.  Only keys matching known Platform values are picked up;
   non-platform keys like streaming are safely ignored.
2. Bridge api_server-specific keys (port, key, host, cors_origins,
   model_name) from the top-level config block into the extra
   dict so PlatformConfig.from_dict preserves them — matching what
   _apply_env_overrides already does for env var values.
2026-07-23 11:54:32 -07:00
arimu1 9e4b89857a fix(gateway): require usable API_SERVER_KEY to enroll the api_server platform at load time
Salvaged from PR #36180 (commits 68dfeb4b16 and 86f437509b by arimu1),
re-applied onto current main with the incidental black-reformat churn
stripped out (~1,700 lines -> the semantic change + tests).

Previously gateway/config.py enrolled the api_server platform on
`api_server_enabled or api_server_key`, so API_SERVER_ENABLED=true with
no key (or a weak/placeholder key) still loaded the platform: the
adapter is instantiated (ResponseStore/SQLite opened in __init__), the
reconnect watcher spins, and the startup guard refuses at connect() —
logging errors forever. Now the platform is enrolled only when
API_SERVER_KEY passes the same strength bar as the adapter's startup
guard (has_usable_secret, min_length=16), via a shared
_has_usable_api_server_key() helper.

The no-op `lambda cfg: True` connected-checker for API_SERVER is also
replaced with the same key check, so get_connected_platforms() only
reports the platform "up" when it could actually start.

Known limitation (intentionally out of scope): a YAML config with
`platforms.api_server.enabled: true` and no key still loads the
platform; this gate covers the env-override path only.

Dropped from the original PR: EMAIL/SMS checker additions (scope creep
beyond the PR title; absent on current main) and the wholesale black
reformat of gateway/config.py and tests.

Fixes #36111

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-23 11:54:05 -07:00
James Huang 74db4bfe68 fix(gateway): bridge top-level port/host into extra for webhook and api_server
WebhookAdapter and ApiServerAdapter read port/host from config.extra, but
PlatformConfig.from_dict only populates extra from the 'extra:' sub-key in
the YAML platform section. Top-level keys like port and host are silently
ignored, causing the adapter to fall back to DEFAULT_PORT (8644).

This causes silent port conflicts in multi-profile setups: a profile that
configures 'platforms.webhook.port: 8649' still binds 8644, colliding with
the default profile's webhook on the same port.

Fix: extend the shared-key bridging loop in load_gateway_config() to bridge
top-level port/host/secret into extra for WEBHOOK, MSGRAPH_WEBHOOK, and
API_SERVER platforms, following the same pattern already used for
dm_policy, allow_from, gateway_restart_notification, and other keys.

The extra dict takes precedence: if port is already under 'extra:', the
top-level value does not clobber it.
2026-07-23 11:31:46 -07:00
davidgut1982 c7fd3eb377 fix(gateway): honor explicit api_server enabled:false under env key
_apply_env_overrides() force-set ``api_server.enabled = True`` whenever
API_SERVER_KEY (or API_SERVER_ENABLED) was present in the environment.

In multiplex mode, a secondary profile pins
``platforms.api_server.enabled: false`` in its config.yaml so that it
shares the default profile's API-server listener instead of binding its
own port. That profile still inherits the process-level env, including
API_SERVER_KEY, so the unconditional re-enable flipped api_server back on
and tripped the MultiplexConfigError check.

Honor an explicit disable, flagged by ``_enabled_explicit`` in the
platform's extra. Use ``extra.pop("_enabled_explicit", False)``: the
api_server branch is terminal (unlike the migrated plugin platforms, no
later registry pass re-enables api_server), so popping consumes the flag
in a single read and avoids the double-read hazard, while the final
per-platform cleanup remains a no-op.

Adds a regression test asserting that with API_SERVER_KEY set, a config
with api_server explicitly enabled:false + _enabled_explicit:true survives
_apply_env_overrides() as enabled=False (fails without the fix), while the
key is still wired through for the shared listener.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 11:21:06 -07:00
And 4cbceae9f4 fix(gateway): normalize YAML boolean streaming mode and keep enabled a mode-only alias
Address PR #62873 review:

- Bare YAML `mode: off`/`on` parse to Python False/True (YAML 1.1). Stringifying
  False yielded "false" (not "off"), so `mode: off` wrongly enabled streaming.
  Add _normalize_transport_token() to map booleans to canonical off/auto tokens,
  mirroring the normalization documented in gateway/display_config.py.
- Only the `mode` alias infers `enabled`; a bare `transport` no longer enables
  streaming, preserving `streaming.enabled` as the documented master switch
  (website/docs/user-guide/configuration.md).
- Update tests to the corrected contract and add YAML-boolean coverage plus
  loader-level regressions for unquoted `mode: off` and nested mode enable.
2026-07-20 05:39:09 -07:00
And 62cdb3e1be Enable streaming when only streaming.mode is set
- StreamingConfig.from_dict now treats `mode` as an alias for `transport`
  that also implies `enabled`, so `streaming: {mode: auto}` turns streaming
  on instead of being silently ignored (enabled defaulted to False, which
  buffered the whole reply and sent it in one message)
- `mode: off` disables streaming; an explicit `enabled` key still wins; an
  explicit `transport` takes precedence over `mode`
- Add regression tests covering mode/transport/enabled precedence and the
  real-world `mode + preloader_frames` block
2026-07-20 05:39:09 -07:00
Teknium 73543744bc fix(gateway): key-presence precedence for session_reset/stt nested fallback
Follow-up for salvaged PR #59779: the session_reset and stt fallbacks
used truthiness/type checks, so a present-but-empty top-level value was
silently replaced by the nested gateway.* form — inconsistent with the
key-presence precedence every other key in the block uses. Switch both
to 'key not in yaml_cfg' gating and add precedence regression tests.
2026-07-20 03:34:58 -07:00
pierrenode e9bd3b6eeb fix(gateway): honor nested gateway.* form for 9 more top-level keys
load_gateway_config() already accepted both the top-level key and the
nested gateway.<key> form (written by `hermes config set gateway.<key>
...`) for multiplex_profiles, max_concurrent_sessions, streaming, and
write_sessions_json — each fixed one at a time as users hit it (most
recently #59320 for multiplex_profiles). Nine sibling top-level keys
never got the same nested fallback: session_reset, quick_commands, stt,
stt_echo_transcripts, group_sessions_per_user, thread_sessions_per_user,
reset_triggers, always_log_local, and unauthorized_dm_behavior.

`hermes config set gateway.<any-of-these> ...` builds exactly this nested
shape (hermes_cli/config.py's _set_nested has no schema, so it accepts
any dotted path), so a user following the same pattern that legitimately
works for gateway.multiplex_profiles/gateway.streaming gets a silent
no-op for these nine keys instead.

Fix: read `gateway: {...}` into a single `gateway_section` variable once
(consolidating three separate `yaml_cfg.get("gateway")` calls already in
the function) and add the same top-level-wins/nested-fallback check for
each of the nine keys, mirroring the existing write_sessions_json
precedent exactly.

Note: because every fallback here is guarded by
`isinstance(gateway_section, dict)`, this also makes the streaming
fallback tolerate a scalar `gateway:` block (e.g. `gateway: disabled`)
without crashing — the same crash #40837 (open) targets specifically for
streaming. This change doesn't set out to fix that PR's issue, but the
consolidated guard covers it as a side effect; flagging it for the
reviewer rather than leaving it to be found in review.
2026-07-20 03:34:58 -07:00
Soju06 c0c76a4715 perf(gateway): byte-stable session context prompts
The per-message ephemeral context prompt re-renders every turn, and
any byte change (Discord auto-thread rename, reset notes, voice
channel state) both breaks the provider prompt-cache prefix at the
head of every request and changes the gateway agent-cache signature,
forcing a full agent rebuild per message. Pin the rendered block per
session keyed by a hash of exactly the fields it renders, so only a
real input change (rename, topic edit, /sethome, redact_pii flip)
re-renders; deliver one-shot per-turn facts (auto-reset note,
first-contact intro, voice-channel changes) on the current user
message via the api_content sidecar instead of the system prompt; sort
get_connected_platforms for byte-stable ordering.
2026-07-19 14:58:59 +05:30
teknium1 d4396797c3 feat(gateway): live per-tool status line on Slack
Builds on the salvaged typing_status_text plumbing (PR #62007): instead
of a static 'is thinking...', Slack's assistant status line now updates
live as the agent works — 'is running pytest tests/…', 'is reading
docs/api.md…' — and reverts to the static text between tool calls.

Mechanics:
- agent/display.py: build_status_phrase() derives a <=49-char present-
  tense phrase from the existing _TOOL_VERBS table (+ 'is using <name>'
  for plugin/MCP tools; None for _thinking).
- base adapter: supports_status_text capability flag + set_status_text()
  per-chat store, cleared when the typing loop winds down.
- Slack adapter: send_typing() renders the live phrase when set, falling
  back to typing_status_text then 'is thinking...'.
- gateway/run.py: progress_callback stashes the phrase on tool.started
  and clears on tool.completed. Rendering rides the existing
  _keep_typing refresh cadence — zero additional Slack API calls, no
  rate-limit exposure. Works with tool_progress: off (Slack default);
  the callback is now armed whenever the adapter supports status text.
- display.live_status config (full|verb|off, default full): 'verb' hides
  argument previews for shared/customer-facing channels.

Also fixes a latent crash in the cherry-picked from_dict: malformed
non-dict 'extra' sections broke typing_status_text resolution (uses the
already-coerced extra dict).

Design notes: status text is a side-effect display channel only — never
enters the transcript, no prompt-cache impact. Lifecycle guarantees from
the stuck-status fix family are preserved (per-thread tracking,
clear-on-finish via existing stop_typing paths). Related: #45109
(closed; same direction via lifecycle states), #59010/#51363 (native
task cards — complementary, larger scope).
2026-07-18 12:28:59 -07:00
George Drury dc0c778b22 feat(gateway): make the working-state status text configurable
Adds PlatformConfig.typing_status_text for the two platforms that render
text for the working-state line: Slack's assistant.threads.setStatus
status (hardcoded 'is thinking...') and Google Chat's visible marker
message (hardcoded 'Hermes is thinking…'). None keeps each platform's
built-in default; to_dict omits the field when unset so existing configs
serialize unchanged. Plumbing mirrors typing_indicator exactly (typed
field, from_dict extra fallback, shared-key bridge).

Also documents that Slack's status line requires the assistant:write
scope — without it setStatus fails silently and Slack shows its own
generic placeholder, which previously made the behaviour undiagnosable
from config alone.
2026-07-18 12:28:59 -07:00
StellarisW f57157a128 fix(gateway): recover Discord websocket and event-loop stalls
Replace REST-based Discord liveness probe with local WebSocket/heartbeat
state detection. REST success doesn't prove Gateway event delivery — a
half-closed WebSocket can leave Bot.start() alive while REST returns 200.
Now samples ready/open/ACK state and heartbeat latency; consecutive
unhealthy samples emit one retryable fatal code so GatewayRunner rebuilds
the adapter through the existing reconnect path.

Also fixes three lifecycle gaps in the recovery path:
1. asyncio.wait_for() can remain blocked if adapter cleanup swallows
   cancellation — now uses bounded asyncio.wait() with task detachment.
2. Multiplexed secondary-profile adapters had no profile-scoped reconnect
   owner — now uses one runner-owned reconnect slot per profile.
3. An in-flight turn could send its final text through the disconnected
   adapter after a replacement was registered — now resolves the live
   same-profile replacement for unsent final responses only (message IDs
   never migrate, edits/deletes stay on the old transport).

Adds an opt-in Linux/systemd event-loop watchdog (gateway.systemd_watchdog_seconds,
default 0) for the failure mode where the whole asyncio loop stops making
progress and no in-process liveness task can run. stdlib-only sd_notify,
Type=notify/WatchdogSec generation, READY/STOPPING lifecycle.

Co-authored-by: 王鑫 <wx.xw@bytedance.com>
2026-07-18 20:01:55 +05:30
liuhao1024 9cb3569e97 fix(gateway): allow Feishu websocket mode in multiplex profiles
Feishu was unconditionally listed in _PORT_BINDING_PLATFORM_VALUES,
causing the multiplexer to reject ALL Feishu secondary profiles. But
Feishu in websocket mode (the default) uses an outbound WebSocket
connection and does NOT bind an HTTP port — only webhook/callback mode
needs a listener.

Add _platform_binds_port() helper that checks connection-dependent
platforms (currently only Feishu) against their actual config before
raising MultiplexConfigError. Feishu websocket profiles are now allowed;
Feishu webhook profiles still raise as before.

Fixes #52563
2026-07-16 07:17:55 -07:00
PRATHAMESH75 e984a61306 fix(dashboard): reject port-binding channels on secondary multiplexed profiles
The Channels API (PUT /api/messaging/platforms/{id}) accepted and persisted
enabling a port-binding platform on a secondary profile while
gateway.multiplex_profiles is on — a config the gateway only rejects on its
next start, aborting startup with MultiplexConfigError for every multiplexed
profile.

Validate before any .env/config.yaml write and return 409 for the enable
attempt. Disabling and clearing env stay allowed so an already-invalid
profile can be repaired. The port-binding platform set moves to
gateway/config.py (PORT_BINDING_PLATFORM_VALUES) as the single source of
truth shared by gateway startup validation and the dashboard, so the two
policies cannot drift. Platform config mutations now get a names-only audit
log line.

Fixes #62791
2026-07-16 07:17:55 -07:00
Teknium b4e4b5a43e fix(gateway): harden multiplex primary token gate — canonical platform map + unserved-platform warning
Follow-ups to @SAMBAS123's #64986 salvage:

- Replace the hardcoded token-platform set in _platform_has_bot_credential
  with PLATFORM_TOKEN_ENV_NAMES, a shared canonical map in gateway/config.py
  also used by the empty-token validation warning — one source of truth, so
  future token platforms can't silently bypass the gate or drift between
  the two sites.
- After secondary-profile startup, warn loudly for any platform skipped on
  the primary that no secondary profile ended up serving: an enabled
  platform with no credential anywhere is a config error, not a silent
  no-op.
- AUTHOR_MAP entry for the salvaged commit's author email.
2026-07-16 04:26:06 -07:00
Teknium 647520f83e fix(gateway): gate profile routing on multiplex_profiles + widen batch-key routing to all adapters
Follow-ups on the salvaged #20096 profile-routing feature:

- _profile_name_for_source now returns None unless gateway.multiplex_profiles
  is on. Routing stamps source.profile, which namespaces session/batch keys,
  but the profile-scoped agent run only activates under multiplexing — without
  the gate, configured routes with multiplexing off split batch/session keys
  into agent:<profile> while the agent still ran from agent:main.
- Widen the profile-aware _text_batch_key fix from Discord to every adapter
  that builds batch keys via build_session_key (telegram, whatsapp, matrix,
  feishu, wecom, weixin) — routing is platform-generic, so the batch-key
  namespace fix must be too.
- Downgrade the no-route-matched log from INFO to DEBUG (fired on every
  unrouted inbound message).
- GatewayConfig.to_dict(): serialize profile_routes as plain dicts
  (ProfileRoute dataclasses are not JSON-safe).
- Docs: correct the 'independent of multiplexing' claim in
  docs/profile-routing.md (routing requires multiplexing), fix the
  platform-only specificity row (0, not 1), and document profile_routes in
  website/docs/user-guide/multi-profile-gateways.md.
- Tests: pin the multiplex gate (routes ignored when off, active when on,
  build_source end-to-end stays in agent:main when off).
2026-07-15 09:50:05 -07:00
Burgunthy 5e65f6d79f feat(gateway): add profile-based routing for inbound messages
Adds gateway.profile_routes config that routes specific Discord
guilds/channels/threads (and other platforms) to different profiles.
The routing engine uses hierarchical specificity matching
(thread > channel > guild) with bounded LRU caching for forum post
resolution.

Routing result is stamped on source.profile by BasePlatformAdapter
.build_source() at inbound time. When gateway.multiplex_profiles is on,
the existing _profile_runtime_scope machinery picks up source.profile
and runs the whole turn inside the profile's HERMES_HOME — so memory,
skills, config, and secrets all resolve to that profile automatically.
No new isolation code is added; this PR only adds the routing decision
layer on top of the existing multiplexing infrastructure.

Configuration:

    gateway:
      multiplex_profiles: true
      profile_routes:
        - name: server-default
          platform: discord
          guild_id: "GUILD_ID"
          profile: server-profile
        - name: special-channel
          platform: discord
          guild_id: "GUILD_ID"
          chat_id: "CHANNEL_ID"
          profile: channel-profile

When multiplex_profiles is off, profile_routes is ignored (no behavior
change for single-profile gateways).

Tests: 29 unit tests covering specificity scoring, hierarchical
matching, path-traversal validation, and config parsing.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-15 09:50:05 -07:00
sharziki a7f65e3bcd fix(gateway): tolerate scalar gateway config block
The streaming fallback path read yaml_cfg.get("gateway", {}).get("streaming") when top-level streaming was absent or malformed. If a user accidentally set gateway to a scalar value, config loading crashed with AttributeError instead of ignoring the malformed block and using defaults.

Read the gateway block once, verify it is a mapping before accessing nested streaming, and keep the existing gateway.platforms fallback using the same checked value.

Adds a regression test for config.yaml containing gateway: disabled.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-09 18:23:27 -07:00
sharziki 50c66b2f8e fix(gateway): ignore malformed config sections
GatewayConfig.from_dict(), PlatformConfig.from_dict(), SessionResetPolicy.from_dict(), and StreamingConfig.from_dict() assumed their input sections were mappings. A malformed scalar from legacy gateway.json or an internal caller could crash config loading with AttributeError before env overrides/defaults had a chance to recover.

Coerce non-mapping sections to empty dicts, skip malformed platform entries, and keep valid sibling platform configs loading normally.

Tests cover scalar platform blocks, scalar nested reset/streaming sections, and malformed PlatformConfig home_channel/extra values.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-09 18:23:27 -07:00
Ben Barclay 75de0057bc feat(gateway): GATEWAY_MULTIPLEX_PROFILES env override for multiplex flag (#60589)
The connector now depends on the single multiplexed gateway for per-profile
relay routing, so hosted deployments need to FORCE multiplexing on regardless
of the image's config.yaml. gateway.multiplex_profiles was config.yaml-only,
which a user could leave unset or flip off.

Add GATEWAY_MULTIPLEX_PROFILES as a standard operator override on top of the
existing config key — the same 'config.yaml is canonical, env is the operator
override' pattern the Telegram/Signal require_mention bridges use:

  env (recognized token) > config.yaml (top-level or nested gateway.*) > False

- gateway/config.py: _env_multiplex_profiles_override() resolves the env var
  tri-state — recognized truthy/falsy token → bool; unset/blank/unrecognized
  → None (fall through to config). Blank is deliberately None, not False, so a
  provisioned-but-unpopulated Fly secret ('') can't shadow a config.yaml opt-in
  (the empty-secret trap). Wired into GatewayConfig.from_dict so every consumer
  (run.py, session.py via self.config) sees the resolved value.
- hermes_cli/gateway.py: the named-profile-start guard
  (_guard_named_profile_under_multiplexer) reads config.yaml directly, so it
  gets the SAME env precedence — otherwise env-forced multiplex would leave the
  guard blind and someone could start a conflicting per-profile gateway that
  double-binds a bot token. Env-forced-on trips the guard even with no
  config.yaml key; env-forced-off disables it over a config opt-in.

Tests: full 3-tier precedence in test_config.py (incl. the discriminating
env-overrides-config cases + the empty/whitespace/unrecognized fall-through
trap + resolver tri-state), mutation-verified (flipping precedence fails
exactly the two env-wins tests); guard env cases in test_multiplex_lifecycle.py.

Force-on is safe on a single-profile instance: session keys stay byte-identical
(agent:main) and the _run_agent wrapper installs the per-turn secret scope, so
the fail-closed get_secret() path is satisfied.
2026-07-08 00:34:34 +00:00
Teknium 9c272a306e feat(gateway): default session auto-reset to off (mode: none) (#60194)
Sessions no longer auto-reset by default. SessionResetPolicy.mode now
defaults to "none" (was "both": 24h idle + daily 4am), matching the
setup wizard's existing no-reset default and community feedback that
surprise context loss hurts more than it helps.

- gateway/config.py: dataclass default + from_dict fallback -> "none";
  installs whose config.yaml lacks a session_reset section stop
  auto-resetting
- hermes_cli/setup.py: "Never auto-reset" is now the recommended/default
  choice in hermes setup agent; stale comment updated
- docs (en + zh-Hans): default is no auto-reset, opt in via
  session_reset in config.yaml

Users who explicitly configured idle/daily/both resets keep them.
2026-07-07 05:11:10 -07:00
davidgut1982 d3602e6308 fix(gateway): read multiplex_profiles from nested gateway section
load_gateway_config() only surfaced the top-level `multiplex_profiles`
key into gw_data before calling GatewayConfig.from_dict(). A config.yaml
that pinned the flag under the nested `gateway:` section -- the form
written by `hermes config set gateway.multiplex_profiles true` -- was
silently ignored, so the gateway loaded with multiplex_profiles=False.

from_dict() already honors the nested fallback, but load_gateway_config()
builds gw_data from top-level keys first, so the nested value never
reached it.

Read gateway.multiplex_profiles into gw_data when the top-level key is
absent, mirroring the existing nested fallback for max_concurrent_sessions.

Adds a load_gateway_config() regression test that writes a config.yaml
with `gateway.multiplex_profiles: true` and asserts the loaded config has
multiplex_profiles=True (fails without the fix).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 22:11:03 -07:00
izumi0uu 0f154e780e fix(gateway): isolate multiplex profile config env reads
Fixes #50051 by preserving nested gateway.multiplex_profiles and routing gateway config env reads through the active profile secret scope when present.

This keeps secondary profile adapter startup from inheriting default-profile platform tokens or port-binding enables while preserving legacy single-profile behavior outside a scope.

Constraint: latest upstream main f57ff7aef1 still reproduced both nested-config loss and cross-profile env leakage
Rejected: special-casing API_SERVER_* only | left other profile-scoped tokens vulnerable to the same leak
Confidence: high
Scope-risk: moderate
Directive: keep future gateway/config env reads on the scoped helper path unless a variable is explicitly process-global
Tested: pytest -q tests/gateway/test_multiplex_phase0.py tests/gateway/test_multiplex_credential_isolation.py tests/gateway/test_config.py -k 'multiplex or scope or getenv or api_server or relay'
Not-tested: full gateway startup across live platform adapters
2026-07-05 22:00:25 -07:00
Teknium 94205a1139 refactor(gateway): move routing index to state.db, make sessions.json an optional legacy mirror (#59203)
Follow-up to #9006/#58899. The gateway routing index (session_key ->
SessionEntry) now lives in a new gateway_routing table in state.db as the
primary store; sessions.json is demoted to an optional legacy mirror.

- hermes_state.py: schema v19 — gateway_routing table (scope + session_key
  PK; scope = resolved sessions_dir so multiple stores sharing one state.db
  never cross-contaminate) with save/replace/load/delete methods
- gateway/session.py: _save() writes the whole index atomically to the DB
  (mirrors the old full-file JSON rewrite semantics) and only falls back to
  JSON when the DB write fails; _ensure_loaded reads the DB first and folds
  in legacy sessions.json entries for keys the DB lacks (pre-migration
  import; DB entries win over stale JSON)
- gateway/config.py + hermes_cli/config.py: new write_sessions_json flag
  (default true for compat/downgrade safety); gateway.write_sessions_json:
  false stops producing the file entirely
- sessions.json _README updated to say it's a legacy mirror + how to
  disable it

Rehydration is now lossless across restarts even with sessions.json deleted:
suspended/resume_pending/model_override/token state all round-trip through
the DB (the old sessions-table recovery only rebuilt the bare key mapping).
2026-07-05 19:25:51 -07:00
devatnull 4be749d151 fix: honor top-level STT transcript echo config 2026-07-05 06:12:49 -07:00