Fence owner discovery by profile, lease and loopback endpoint, and hand the authenticated URL to the existing Ink transport without acquiring a competing lease. Distinguish lease age from turn activity in unsupported-owner recovery.
Client slice only: requires the integration runtime to advertise shared_runtime_url and provide the session-attach handshake. Classic CLI attachment remains an integration gap.
`key_cmd` (#86891) authenticates a provider with a SHORT-LIVED bearer minted
by a command — SSO/OIDC brokers, cloud IAM, internal auth proxies. The
request path has honoured it since it landed, but the picker resolved probe
credentials from `api_key`/`key_env` ONLY, so a key_cmd provider probed
`/v1/models` with an EMPTY key.
Against an authenticated endpoint the probe 401s, discovery returns nothing,
and the provider falls back to its single configured default model. The
picker shows ONE model, indistinguishable from an endpoint that genuinely
serves one — while inference keeps working, because that path mints
correctly. Reproduced against a LiteLLM gateway behind Entra OIDC: 0 models
discovered with an empty key, 26 with the minted token.
Both picker probe sites already funnel through `_entry_credentials()`, so
the fix lands in one place: it now reports a `cmd:<key_cmd>` identity, and
each site falls back to `resolve_probe_token()` after api_key/key_env. An
explicit static key still wins, so existing configs are unaffected.
The identity is keyed on the COMMAND, never the minted token: the token
rotates on every refresh, so keying on its value would change the group
fingerprint constantly and force a re-probe on every open. Two entries on
one URL with different helpers still get distinct rows.
`resolve_probe_token()` lives in agent.command_token_source, which already
owns key_cmd minting, and shares the CommandTokenSource cache with the
request path — a cache read, not a fresh sign-in. Fail-closed: a helper
needing an interactive sign-in degrades to today's empty-key behaviour
rather than taking down every other provider's row.
`_model_flow_named_custom` (the `hermes model` setup flow) is the sibling
path — it builds its own `Authorization: Bearer` from the same incomplete
resolution — and is fixed the same way, with one ordering constraint: the
value persisted to config.yaml is computed BEFORE the mint, so a short-lived
bearer can never be written back to shadow the key_cmd meant to re-mint it.
Tests drive the real code paths and assert on the credential each probe
receives rather than on function source, so a semantics-preserving refactor
does not fail them. Verified they fail with the fix reverted.
Carry the owning task's notification subscriptions independently of dependency
edges, within the creation transaction. Prefer its durable session over worker
and request-local sessions while preserving explicit overrides. Cover worker
CLI create and built-in decomposition, and retain conversation route anchors.
Auto-subscribe no longer upgrades an inherited passive subscription.
Slim adaptation of Christopher-Schulze's session-precedence fix in #85687,
expanded to durable subscription provenance and sibling creation paths.
Related: #85575, #85687
Validation: strict RED/GREEN (7 failing cases before; 7 passing after), then
58 Kanban test files: 383 passed, 2 skipped. Real dispatcher-spawn subprocess
probe covers direct, linked, unlinked, explicit-session, worker CLI, built-in
children and a plain CLI negative control, with recording transport only.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
Drives the real run_gateway() boot with the service-manager environment (no bus
vars, INVOCATION_ID set) and asserts the bus address reaches the worker env the
dispatch paths build — including surviving the secret scrubber, without which the
fix would be a silent no-op. The companion test pins the other half: with no user
manager present nothing is fabricated, so the scope probe's refusal stays honest.
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
Sweep of `ruff check . --select F821 --target-version py311`: 2,234 hits. 2,201 are left
alone on purpose: tui_gateway (2,169; bind_module rebinds bodies onto server.py globals,
all names verified to resolve there), the Feishu adapter (27; globals().update() SDK
binding) and the godmode script (5; dead standalone script). The other 33 were all
genuine defects. No lint config change; no TYPE_CHECKING escape hatches — every
annotation names a real, imported type; ty on the touched files: 0 new diagnostics.
- gateway/slash_commands.py: HISTORY_UNREADABLE never imported after #102117
→ NameError on the /btw error branch (same one-liner as #102952).
- gateway/platforms/whatsapp_common.py: `-> Path` return annotation with no Path
import (the body uses `_Path`). Never raised at runtime thanks to
`from __future__ import annotations`, but `typing.get_type_hints()` and ty
both fail on it.
- gateway/run.py: ActivityProvenance imported at module level
(agent.session_activity has no gateway deps); stringly annotation and the
lazy in-function import are gone.
- tools/patch_parser.py: PatchResult imported at module level; real return
annotation. The "avoid circular import" lazy import guarded a cycle that
does not exist (file_operations_common never imports patch_parser).
- gateway/platforms/helpers.py: base.py imports helpers at module level, so
MessageEvent cannot be named here; TextBatchAggregator only reads .text and
.source, so it is typed by a BatchableEvent Protocol that MessageEvent
satisfies structurally.
- tools/mcp_tool_sampling.py: mcp_tool imports this module, so MCPServerTask
cannot be named here; ElicitationHandler only reads
owner._pending_call_context, typed by an ElicitationOwner Protocol.
- plugins/platforms/sms/adapter.py: aiohttp is an optional dep ([messaging] extra) →
module-level try/except ImportError binding `aiohttp = web = None`, the pattern the
homeassistant / webhook / whatsapp_cloud adapters already use. Retires three lazy
in-function imports and the `_aiohttp_available()` wrapper; `_handle_webhook` typed
`web.Request -> web.Response`.
- plugins/platforms/teams/summary_writer.py: plain module-level `import httpx` — httpx is a
hard core dependency (pyproject `httpx[socks]==0.28.1`), so the lazy import and the
"imported on every CLI start" docstring premise were both wrong (plugin discovery never
imports this module; it is reached only via the Teams adapter / meeting pipeline).
Tests:
- tests/hermes_cli/test_config.py: a test body orphaned by the wave-1 prune
(6b81590c55) sat inside the class as dead code with self/tmp_path unbound
— header restored, so the v11→12 custom_providers migration is covered.
- tests/tools/test_mcp_tool.py: @staticmethod recursing on `self` in the
win32 branch; call portalocker directly.
- tests/test_background_review_list_shapes.py: main() still ran 3 pruned tests.
- tests/agent/test_cursor_optimizations_parity.py: bench() used names only
imported inside a sibling test.
- GatewayRunner / FeishuAdapter / Dict / Optional: missing imports.
The slug pair {"gpt-6-astra", "gpt-6-astra-900k"} was spelled out five times
(reasoning_effort, transports/codex, models.py x2, codex_models) with five hand-rolled
``.strip().lower().rsplit("/", 1)[-1]`` normalisations. ``is_astra_model`` in
agent/reasoning_effort.py (the module the other four already import) is now the only home, so a
new Astra alias is one edit.
``_finalize_codex_models(..., allow_astra=)`` was a control-coupling flag set True by the two
live callers and left False by the two static ones. The filter now lives where the untrusted
inputs are — ``_drop_undiscovered_astra`` over config.toml default + models_cache.json in
``get_codex_model_ids`` — and ``_finalize_codex_models`` is back to its one-line original.
DEFAULT_CODEX_MODELS never contains Astra, so the static catalog needs no gate.
``_openai_catalog`` appends whatever Astra ids live discovery returned instead of a name-keyed
``if "gpt-6-astra" in live_lower`` branch with a literal fallback that could never fire.
The regression test builds its scratch repo under GIT_CONFIG_GLOBAL/
GIT_CONFIG_SYSTEM = /dev/null, so the init commit has no configured
identity. On CI runners the auto-detected ident is rejected
(user@<bare-hostname>.(none)) and 'git commit' exits 128, failing the
test that passes locally. Pin user.email/user.name on the commit, the
same pattern the install-script tests already use.
ssh bypasses stdin=DEVNULL and GIT_TERMINAL_PROMPT: when a git child
dials an SSH remote whose host key is unknown, ssh opens /dev/tty
directly and its yes/no prompt steals the caller's terminal — exactly
what noninteractive_git_env exists to prevent. Pin core.sshCommand to
"ssh -o BatchMode=yes" at the config-injection layer so the ssh child
fails instead of prompting; an agent-authenticated ssh still succeeds,
and an explicit user GIT_SSH_COMMAND env var still takes precedence
(#104591).
The startup update check resolved `git remote get-url origin` with the
user's global config in scope while the subsequent fetch runs under
noninteractive_git_env (GIT_CONFIG_GLOBAL=/dev/null). A global
url.<https>.insteadOf rewrite therefore made an SSH origin masquerade as
HTTPS, the SSH-avoiding fast path was skipped, and the fetch dialed the
raw SSH origin — whose host-key prompt opens /dev/tty directly and
steals the CLI's keystrokes (#104591).
Probe the origin URL under the same isolated env so both sides observe
the URL the fetch will actually dial.
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.
Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.
Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
Main now snapshots systemd unit listings before stopping old processes
and passes them into _restart_systemd_gateway_units_best_effort; the
catch-up budget test and eval call the new two-argument shape while
still asserting the unit's stop+start budget reaches the client timeout.
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.
Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Salvaged from #104272. Preserve restart and fatal-exit policy while classifying the planned restart code as success. Earlier analysis in #13604 by Justin Kausel.
Adapt the config-only portion of #104347; omit its environment flag and unrelated docs. Explicit update commands remain independent.
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.
Slim redo informed by #104274, #104283, and #104285.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Slim adaptation of anombyte93/hermes-agent@d06d2a49c5; use the canonical board resolver instead of inferring the slug from a path. Live isolated CLI probe confirms current-file, env and explicit board banners; event delivery remains live.
Co-authored-by: Hayden (Atlas agents) <212644172+anombyte93@users.noreply.github.com>
check_for_skill_updates() fetched every lock-file entry remotely, even
when the entry's install directory no longer existed, and each fetch had
no wall-clock bound — a few dead sources turned a routine
`hermes skills update` into a multi-minute stall (#104291).
- Entries whose recorded install_path resolves but does not exist are
reported as "orphaned" and skipped without a remote fetch;
unresolvable paths keep the previous fetch behavior.
- Each fetch now runs under a daemon helper thread with a hard timeout
(default 30 s) and degrades to "unavailable" when abandoned.
- `hermes skills check` prints a removal hint for orphaned entries.
Fixes#104291
Retain partial-line output, UTF-8 decoding, failure output and cancellation cleanup. Based on streaming investigations by Artemonim (#101850) and lEWFkRAD (#104843); gateway tee adapted from fangliquanflq (#97402). Live Linux child/tee probe: withheld or dropped on base, visible in 0.02 seconds after. Campaign-locked tests and native Windows proof are pending.
Preserve the two contributor fixes, slim them to two behavioral invariants, and enter explicitly requested homes even inside a nested scope. Real native remote Desktop changes Disabled to gateway_stopped for default and named profiles; direct API controls preserve explicit disable and empty-profile isolation. Unit A/B and regression suites remain queued under the shared campaign lock.
The desktop app always sends profile=default on GET /api/messaging/platforms.
_is_current_profile() recognized only None/""/"current" as the dashboard's
own profile — NOT the string "default" — so a single-profile install (the
standard `hermes gateway setup` flow: token in .env, no platforms: section in
config.yaml) entered the profile-scoped branch of _config_profile_scope().
That branch derives platform enablement from config.yaml only and never calls
load_gateway_config()'s env-override pass (which enables the platform when the
token is in the environment). Result: a platform connected via .env reported
enabled=false, state="disabled" while it was actually running. The unscoped
GET (no profile param) correctly reported enabled=true, state="connected".
Fix: classify by resolved path, not by string. After _is_current_profile()
fails, _config_profile_scope() now resolves the requested profile dir and
compares it against get_process_hermes_home().resolve() — the same comparison
_is_other_profile() already uses. When they match (profile=default on a
default-home process), yield None (no override), taking the unscoped path that
calls load_gateway_config(). A named-profile process (`-p worker`) has a
different HERMES_HOME, so its profile=default resolves to a different directory
and still scopes correctly — cross-profile secret isolation is preserved.
The scoped branch's config.yaml-only enablement is DELIBERATE:
load_gateway_config()'s env pass reads os.environ and would leak the root
install's tokens into a genuinely different profile's state. This fix only
reclassifies requests that name the process's OWN home; it does not touch the
scoped branch or gateway/config_env.py.
Refs #104614
all() over an empty tuple evaluates True, so the scoped credential
fallback reported platforms with required_env == () (whatsapp, yuanbao,
api_server, webhook, a2a, msgraph_webhook, relay, whatsapp_cloud) as
enabled=True with no config entry and no credentials. Add the
bool(required) guard to the enabled computation (per review suggestion)
and to the configured field, which came from the same all()-over-empty
expression and reported configured=True for the same shape — the
unscoped branch reports enabled=False / configured=False there, so the
scoped branch now agrees.
Adds tests/hermes_cli/test_web_server_scoped_enablement.py covering the
empty-required_env shapes, the explicit-enabled precedence, and the
credentials-present path.
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
The scoped branch of _platform_enablement consulted only config.yaml's
platforms: section, but the `hermes gateway setup` wizard writes .env
credentials and never a platforms: entry. The desktop always sends
?profile=default (normalizeProfileKey maps the primary profile to
`default`), so the Settings - Messaging page showed a working bot as
"Disabled" while /api/status reported it connected (#104614).
Mirror _enable_from_env (gateway/config_env.py): env credentials alone
enable a platform, an explicit enabled: false still wins. Only the
profile's own .env (env_on_disk) is consulted, so the root install's
os.environ credentials still never leak into a profile's state.
Fixes#104614