`hermes plugins doctor` defaults its target to `.`, and
resolve_plugin_path accepted any directory that existed. Doctor then
copytree'd the resolved path into a temporary HERMES_HOME *before* the
runtime got to reject it, so running the command from a directory that
is not a plugin copied that whole tree.
Run from $HOME on macOS this copies the home directory, and because
`Library/CloudStorage` is not excluded it also materializes every
cloud-only Google Drive/iCloud placeholder. Observed locally: 461 GB
written to /private/var/folders and still growing when the process was
killed, on a machine with 49 GB free.
Resolve now requires a manifest before returning a path, mirroring
PluginManager._scan_directory: `plugin.yaml`/`plugin.yml`/`plugin.json`
in the directory itself, or in one immediate subdirectory for the
category layout. Plugin-id candidates are only tried for values that can
be an id, since joining `.` onto a plugins root resolves to the root and
would hand Doctor every installed plugin at once.
Non-plugin targets now fail with a clear message and no copy.
/simplify-code efficiency reviewer (verified): cleanup_stale=False does
NOT deliver the exclusion the salvaged fix intended — get_running_pid
returns None whenever a record fails liveness VALIDATION (start-time
mismatch, argv drift, lock hiccup) regardless of the flag, which only
controls unlinking. In exactly the at-risk scenario the recorded PID
still never joined the exclusion set.
Exclusion evidence now comes from the RAW pidfile + lock records (no
validation, no unlink side effects); the validated non-destructive
probe is kept for the runtime-status fallback PID. For a KILL exclusion
list this is strictly safer: a stale recorded PID at worst spares one
process for one sweep, while a validation false-negative would
TerminateProcess a live gateway. Regression test reworked to drive the
real function semantics (raw record present, validation rejects);
mutation-checked: removing the raw-record read fails the test.
Follow-up on the salvaged #87158: the reaper-side fix (probe with
cleanup_stale=False so the sweep never deletes its own exclusion
evidence) fully protects a healthy standalone gateway, so the
web_server.py skip-sweep gate is dropped — a stale-but-present
registration must not veto the #77276 orphan reap that motivated the
sweep in the first place.
New regression test drives the exact failure: a registration that
fails liveness validation only surfaces its PID when probed
non-destructively; with the destructive default the standalone gateway
would be hard-killed (TerminateProcess, no drain). Mutation-checked:
reverting the probe to get_running_pid() fails the test.
The threading cluster (#99429) branched before the auth cluster (#99427)
added auth_tag to _exec_buzz's signature; merging both left two
fake_exec doubles in test_buzz_thread_topology.py without the kwarg,
red on main. Sibling-test blast radius, no behavior change.
Compose the two salvaged approaches (#78065 + #78511):
- Keep #78065's terminal-only scrub-path exemption (first-party prefix
predicate in _make_run_env / _sanitize_subprocess_env, plain env values
never scope-resolved, snapshot exclusion for cross-profile isolation,
every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
an import-time blocklist discard: the blocklist is shared by every
scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
process env (Buzz Desktop buzz-acp harness, #76243) OR the live
session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
session on a host that also runs a Buzz gateway does NOT get the
signing key in its terminal children (maintainer triage note on
#76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
strips) and positive test (buzz session platform enables carve-out);
docs updated accordingly.
Closes#78026, closes#76243.
The bundled buzz platform plugin registers BUZZ_PRIVATE_KEY / BUZZ_RELAY_URL /
BUZZ_AUTH_TAG as messaging credentials, so the terminal env blocklist strips
them unconditionally. But when Hermes runs as a Buzz-ACP managed agent, the
buzz-acp harness hands exactly those vars to the agent process ON PURPOSE:
the platform's reply-delivery contract is the model driving the buzz CLI
(block/buzz#2698 — final session text is not delivered to the channel).
Stripping them made every reply die with 'auth error: BUZZ_PRIVATE_KEY is
required' and the channel stayed empty while activity panels streamed the
answer. Mirror the CLAUDE_CODE_OAUTH_TOKEN exception, gated on
BUZZ_MANAGED_AGENT (set only by the Buzz Desktop harness), so gateway/CLI/
kanban contexts keep stripping the path-3 platform secret.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 065e41c0cceb3a373b5025449301c8e930d8a8f0)
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.
Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.
Follow-up hardening from review:
- First-party matches use the merged env value directly instead of
_resolve_passthrough_value: under multiplex with no profile secret scope
installed the resolver raised UnscopedSecretError (fail-closed) at call
sites like the webhook-filter script runner, a regression where the
script previously ran without the var. The vars are the process's own
env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
shared login-shell snapshot (_additional_profile_scoped_passthrough_names
override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
exclusion set (env_passthrough refuses blocklisted names), so without
this a multiplexed gateway would dump profile A's key into
hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
would source it — a cross-profile nsec leak. The names are now excluded
from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
like ddgs, computer-use driver, user-script runners) that also receive
first-party platform vars.
Fixes#78026
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Real adapter imports with synthetic NIP-10 relay events; covers the composed
cluster behavior end to end: in-thread triggers anchor replies to the thread
root (never nesting), top-level triggers still open exactly one thread,
reply_in_thread/reply_to_mode opt-outs post flat on send/send_image/
standalone cron send, progress thread resolution honors the opt-out, and
buzz no longer inherits verbose _GLOBAL_DEFAULTS (#95841). Sabotage-checked:
the display-tier tests fail against origin/main's display_config.py.
Compose/fix-up on top of the cherry-picked cluster commits:
- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
through _apply_yaml_config and honored by send(), send_image(), and
_standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
_progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
so the synthetic-thread fallback and the progress reply anchor are both
suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
_extract_thread_root (marked root > reply > legacy positional e-tag)
instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
edit_message now implemented, accumulate-style progress works, but without
the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
Buzz has no native thread_id; channel threading is entirely --reply-to on
the triggering event. Interim commentary and progress bubbles only passed
the anchor via metadata.reply_to_message_id (or not at all), so most Kathy
posts landed as new top-level messages and cluttered channels.
- Honor metadata.reply_to_message_id in BuzzAdapter.send
- Pass reply_to on stream commentary sends
- Treat buzz like slack/mattermost for progress thread resolution
- Set _progress_reply_to to the trigger event for buzz
- Add unit tests for adapter metadata and progress routing
Every Buzz reply opened a fresh thread, including when the user was
already replying inside one. A threaded client fills up with an endless
ladder of one-message threads and the conversation becomes unreadable.
The adapter itself never had threading logic; the behaviour comes from
the generic gateway default. `_reply_anchor_for_event()` in
gateway/platforms/base.py returns `event.message_id`, which is right for
reply-style platforms (Telegram/Discord "reply to this message") but
wrong for a thread-style one: anchoring to the message you are answering
nests a new sub-thread under every single turn.
Buzz threads are NIP-10, so the information needed is already on the
inbound event. The adapter now records each inbound message's thread
root from its `e` tags and resolves the outbound anchor to that root, so
a reply joins the thread the user is typing in. When the trigger was
itself top-level there is no root and the anchor passes through
unchanged, preserving the existing behaviour of opening exactly one
thread from a top-level message.
Fixed in the adapter rather than in `_reply_anchor_for_event()`: Buzz is
a plugin-supplied platform, and its NIP-10 tag semantics do not belong in
core. Root extraction prefers an explicit `root` marker, falls back to a
lone `reply` marker (a message bearing only `reply` started the thread,
so that parent is the root for everything after it), and treats a legacy
unmarked `e` tag as the parent. A mention-only `p` tag is not a reply.
The root cache is an OrderedDict bounded at 512 entries with FIFO
eviction so a long-lived gateway cannot leak, and the resolver is applied
to the image send path as well as `send()`. Both helpers tolerate a
missing `_thread_roots` attribute, since the standalone/cron send path
constructs an adapter without running `__init__`.
Tests cover root extraction (top-level, thread opener, nested, legacy
unmarked tag), the top-level passthrough that guards the existing
behaviour, unknown/None anchors, cache bounding and eviction, and an
end-to-end assertion through `send()` that `--reply-to` carries the root.
Verified against the real event shapes returned by a live hosted relay.
The Buzz adapter appended --reply-to unconditionally, so every agent reply
threaded onto its parent event id with no way to turn it off.
reply_to_mode is already a generic PlatformConfig field, parsed for any
platform from gateway.platforms.<name>.reply_to_mode, and the Discord and
Telegram adapters both honor it. Buzz never read it, so setting it was a
silent no-op.
Read it in __init__ (BUZZ_REPLY_TO_MODE overrides config.yaml, matching how
require_mention and transport already work in this adapter) and skip the
--reply-to append when it is "off", at all three send paths: send(),
send_image(), and the out-of-process _standalone_send() used for
deliver=buzz cron delivery.
Default is unchanged ("first"), so existing installs keep threading.
The gateway already streams by sending a first partial message and re-editing
it as tokens arrive, falling back to that path when an adapter does not support
native drafts. The Buzz adapter never implemented edit_message, so it inherited
the base stub that returns success=False and every reply was delivered in one
block when the turn finished, however long the turn took.
buzz-cli already exposes `messages edit` and `messages delete`, so no new
mechanism is needed.
One detail worth calling out for review: buzz-cli reports a NEW event id for
each edit, but the edit TARGET stays the original id, and the stream consumer
holds a single message_id for the whole stream. edit_message therefore returns
the id it was given rather than the one the CLI reports. Returning the CLI's id
would make every edit after the first address a message that was never sent.
delete_message is included because the consumer's fresh-final cleanup path
calls it when it replaces a preview rather than editing in place.
Tested: 10 new cases in tests/gateway/test_buzz_adapter.py covering the edit
target, stdin content, the returned id, echo suppression, finalize being inert,
both no-op guards, retryable vs non-retryable CLI failures, and delete. The
file goes from 33 passing to 6 failing if the adapter change is reverted while
the tests stay.
The 5s waitFor deadline trips on loaded CI runners: on 2026-08-31 the same
UI-shard flake hit a plugins-only main push and two unrelated PRs, always as
'expected vi.fn() to be called at least once' in gateway-settings /
messaging / session-unread-tile / toolset-config-panel. Success still
resolves the moment the assertion holds; 12s only absorbs starvation and
stays under the 15s testTimeout so real hangs keep their distinct failure
shape.
A named profile has no local gateway.pid, so cron warned that jobs
would not fire and recommended hermes gateway install — which the
start guard then refuses with exit 78. Share the multiplexer-serving
probe with the start guard and count it as liveness.
The #82871 symptom — gateway default-denies every Buzz user because no
central env allowlist exists and the adapter's config.extra.allowed_users
was never consulted — is fixed by the plugin-platform extra.allowed_users
fallback salvaged from PR #98748 (registry-gated, normalize_user_id-aware).
That PR's tests cover the multiplex profile path; these pin the plain
single-profile gateway path from the issue's repro: npub-only and hex-only
config lists authorize, unlisted senders and empty lists stay default-deny.
Sabotage-verified: all four fail on origin/main's authz_mixin.
The WS NIP-42 auth path now prefers the connect()-resolved _auth_tag
(credentials-file aware, #79514) and falls back to a lazy scope-aware
_resolve_auth_tag() so a bare adapter re-auth stays profile-correct
(#98738): scoped multiplex profiles fail closed instead of borrowing
the default profile's tag from os.environ. _exec_buzz fakes updated
for the auth_tag kwarg introduced by the #83155 salvage.
Regression for the cron/standalone path: credentials-file auth_tag must be
passed into _exec_buzz, and a bare BUZZ_PRIVATE_KEY must not invent a tag
from ambient credential files.
Reapplied from PR #83155 head (original commit carried a placeholder
local identity).
Co-authored-by: Alex P. Günsberg <alex@gunsberg.fi>
Remove the stray 'return val if val is not None else default' tail left
in _unscoped_profile_secrets() when the new return was added, and note
in the docstring that the process-global cache is startup-gate-only
(review feedback on #95224).
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.
Summary:
The gateway's central allowlist check compared the inbound Buzz sender's
64-char hex pubkey against the raw BUZZ_ALLOWED_USERS entries. An operator
who listed only their npub saw every message rejected with
"Unauthorized user: <hex pubkey>" (gateway drops the message). npub
entries are now decoded to hex before the comparison, so npub and hex
forms of the same identity are equivalent.
Root Cause:
The Buzz adapter's own intake check already normalizes npub→hex via
_normalize_user_ref when building _allowed_pubkeys, but the gateway
applies BUZZ_ALLOWED_USERS centrally as well (authz_mixin._is_user_authorized
via the platform registry's allowed_users_env). That central path did a
raw string comparison of the allowlist entries against the hex user_id,
so an npub-only entry never matched.
Change:
- gateway/authz_mixin.py: add a pure-stdlib bech32 npub→hex decoder
(mirroring plugins/platforms/buzz/adapter.py) and normalize the buzz
allowlist set in _is_user_authorized: each npub1… entry is decoded and
its hex form added; hex entries pass through unchanged, so existing
hex-only allowlists keep working. Comparison stays fail-closed for
unrelated senders.
- tests/gateway/test_buzz_authz.py: new tests covering npub-only,
hex-only, mixed, uppercase-npub, and denial of unrelated users, plus
unit tests for the decoder/helper.
Verification:
- 43 passed (test_buzz_authz.py + test_buzz_adapter.py +
test_pairing_allowlist_bypass.py)
- 9 passed (test_multiplex_profile_authz.py + test_buzz_websocket.py)
Closes#78428
One BUZZ_* read survived the #98738 sweep unscoped: the NIP-42 WebSocket
auth path read BUZZ_AUTH_TAG with a bare os.getenv. Under
gateway.multiplex_profiles the process env holds the default profile's
bridge/.env output, so a scoped secondary profile without its own tag
signed its relay auth event with the default profile's NIP-OA
owner-attestation tag. Reproduced on f3845a72af before the fix; the same
repro now attaches no tag (fail-closed).
The read goes through _get_scoped_secret: scoped multiplex profiles fail
closed to "", while single-profile and unscoped default-profile reads
keep the legacy env behavior. Adversarial coverage added for the leak
itself, the scoped positive control, unscoped precedence, partial-extra
adapter config, scoped validate_config, scoped standalone-send target
resolution, central-authz wildcard/blank-entry/normalization semantics,
and adapter-intake vs central-authz agreement on the same allowlist.
Fixes#98738
Signed-off-by: Kosta Gorod <35299380+KostaGorod@users.noreply.github.com>
The four new #98738 authorization tests failed on CI shards that had
never looked Buzz up: plugin platforms have no static Platform member —
Platform._missing_ creates one on demand and caches it in _member_map_,
so attribute access (Platform.BUZZ) only works after an earlier value
lookup in the same process. Local full-suite runs happened to register
it first, which is why this only surfaced on a fresh shard.
Resolve the member once via Platform("buzz") at module import and use
that constant in the runner/source helpers.
Under gateway.multiplex_profiles the default profile's YAML-to-env bridge
writes BUZZ_* values into os.environ, and every Buzz read gave that env
precedence over the secondary profile's PlatformConfig — so each secondary
adapter connected as the default identity, watched its channels, and
resolved its credentials file (#98738).
- Add _profile_scoped()/_scoped_platform_setting(): inside a secondary
profile scope extra is authoritative and env is not consulted (a missing
key fails closed to its default instead of borrowing the default
profile's value); single-profile and unscoped/default-profile reads keep
the legacy env-over-config precedence.
- Apply the scoped read to BuzzAdapter.__init__ (relay, CLI path, channels,
home channel, poll interval, require_mention, transport, allowed users),
_resolve_private_key (BUZZ_CREDENTIALS_FILE), validate_config,
_standalone_send, and check_requirements (which now consults the
profile's own config.yaml via the scoped home override).
- _env_enablement() returns None inside a profile scope and
_apply_yaml_config() skips the env bridge there, so the default profile's
env cannot fabricate Buzz for a profile that never configured it and a
secondary profile's YAML cannot be pinned into the process env
(first-writer-wins, #72348 Telegram/Discord mirror).
- Central authorization now consults a plugin platform's live-adapter
config.extra.allowed_users (gated on the registry entry declaring
allowed_users_env, with an optional normalize_user_id hook so Buzz npub
entries match hex-pubkey user ids) — under multiplex only the default
profile's list ever reached the env var, so listed secondary-profile
users were default-denied (#82871). Empty/absent lists change nothing;
default-deny is preserved.
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.
Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).
Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
The roster click's fronted-tab shortcut trusted the persisted session-tile
bucket (Local Storage 'hermes.desktop.sessionTiles.v2') unconditionally: a
persisted 'Bot Chat' tile naming a session the canonical registry no longer
resolves to — a superseded row from the retired ui_meta pointer design, a
re-minted canonical chat, a stale finished (often hidden) session — was
fronted on every click and remembered as the zone's active pane, so the row's
click target stuck to that stale session forever while its preview/age
described the live one, and clearing Local Storage only healed until the next
click re-persisted the same tile.
Per the Desktop guide the backend is authoritative for session state and the
renderer copy is a cache that must reconcile. focusWorkspaceOwnerSessionTile
now takes an optional staleness probe: tiles the probe rejects are discarded
(same no-undo rationale as discardSessionTile — resurrecting one would just
front the stale session again) and never fronted. The roster click supplies
the probe: a canonical-titled tile whose stored id matches neither the
server-resolved canonical_session registry row nor its compression-lineage
tip is stale, so the click falls through to the authoritative name-registry
open. Side-chat tabs carry no registry identity and are never judged; older
gateways without canonical_session (and shells without the probe) keep the
previous behavior unchanged.
Two growth leaks closed:
1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
trees, ~18GB on the reporting box): managed installs fetch with a
single-branch refspec, so pushed PR branches never get refs/remotes/*
entries and read as 'unpushed' forever. When a clean tree's branch head
EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
is redundant: reap the TREE, keep the BRANCH ref (shielded from the
orphaned-branch pass). Anything diverged/unverifiable stays preserved.
Applied to both the startup pruner and hermes worktree prune/list.
2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
gateway-driven boxes accumulated trees for days. The scheduler tick now
dispatches the same conservative pruner on a daemon thread, throttled
to once per 6h, against the install checkout + job-workdir repos that
have a .worktrees/ dir.
The Discord adapter resolves username allowlist entries to numeric IDs at
connect and mirrors them into os.environ — but the gateway's per-turn .env
hot-reload (load_hermes_dotenv(override=True)) restores the raw usernames
from the file. From the second agent turn onward, _is_user_authorized
compared numeric user_ids against username strings and dropped every
message from the operator as 'Unauthorized user' while the adapter layer
still admitted them (bot reacted, never replied).
Fix: gateway authz unions the adapter's resolved numeric IDs
(DiscordAdapter.resolved_allowlist_user_ids()) into the env-derived
allowlist. Union only fires when an env allowlist is configured (never a
widening; fail-closed branch unchanged), is duck-typed + isinstance-guarded
against mock adapters, and filters non-numeric entries so unresolved
usernames and '*' can't leak through adapter memory.
Live repro: symptom fired on origin/main (authorized=False after reload),
passes with fix; stranger + empty-allowlist + raising-resolver negatives
hold. Sabotage run: incident test fails on unfixed authz_mixin.
Since the #98790 heartbeat guard, a never-ticked gateway prints the YELLOW
first-heartbeat notice (which also contains 'NOT fire'). The lock-first
test's real contract (#87033) is that an active runtime lock suppresses the
RED 'Gateway is not running' false alarm — assert that directly.
The late-result close callback (#72782) retrieves the future's SessionDB
and closes it; passing an explicit db_path kwarg broke the hanging-init
test's mock shape. contextvars.copy_context().run(SessionDB) alone is
sufficient — SessionDB resolves its default path from get_hermes_home(),
which reads the profile ContextVar.
Repairs #98790 where
✓ Gateway is running — cron jobs will fire automatically
PID: 4165
Ticker heartbeat: 39s ago
4 active job(s)
Next run: 2026-08-30T22:50:18.762041+03:00 in profile B incorrectly reports
that jobs will fire based on profile A's gateway process.
Root causes:
1. executed ,
which enumerated the entire systemd fleet regardless of ,
violating the docstring "only PIDs belonging to the current profile".
2. checked when
(no heartbeat file) should trigger a warning — instead, it fell through
to the "✓ Gateway is running" green branch.
Changes:
- hermes_cli/gateway.py::_get_service_pids: pattern = get_service_name()
when all_profiles=False, filtering to the current profile's systemd unit.
- hermes_cli/cron.py::cron_status: guard hb_age is None first with an
explicit yellow warning: "ticker has not reported a heartbeat".
Regression test suite guards both systemd scoping (default + all_profiles)
and heartbeat branching (None vs fresh vs stale).
Under gateway.multiplex_profiles the primary gateway's in-process ticker
fires satellite-profile jobs and delivers through the primary's live
adapters (#69377) — the satellite home intentionally holds no platform
credentials (its own token would be a duplicate_credential fatal).
_preflight_check_delivery loads the gateway config of the job's OWN home,
so a profile_routes-routed platform reads as unconnected there and the
job is permanently blocked before any LLM call with a misleading
"not connected" error (#97476).
When the own-home config reports a platform unconnected, consult the
primary home's profile_routes: an enabled route matching the platform
that points at the profile currently being served means delivery is the
primary gateway's to make — pass the check. The primary config.yaml is
read directly (both top-level and nested gateway. forms) instead of via
load_gateway_config() so no primary platform config leaks into the
satellite process's environment. Lookup failures and missing configs
fail closed (the block stands).