Commit Graph

354 Commits

Author SHA1 Message Date
kshitijk4poor 16fe50b1a2 refactor(gateway): inline the user-bus adoption gate at the run_gateway call site
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
2026-09-08 02:28:21 +05:30
HexLab98 802f0f97ad fix(gateway): adopt the user D-Bus session when systemd starts the gateway
A system-level unit (/etc/systemd/system, User=<someone>) is exec'd with neither
XDG_RUNTIME_DIR nor DBUS_SESSION_BUS_ADDRESS, and a process environment is fixed
at exec time. 'systemd-run --user --scope' therefore fails for the whole lifetime
of that gateway even after the user manager is up and /run/user/<uid>/bus is
reachable. That is the seam every restart-safe worker crosses
(restart_safe_gateway_child_argv), and it fails closed by design — so on headless
systemd installs every agent-driven cron job and every Kanban dispatch died at
launch, ~26ms in, with nothing but 'error' on the job row.

_ensure_user_systemd_env() already derives both values from our own uid and adopts
them only when the runtime dir is really ours and the socket really exists; it was
just wired exclusively to the systemctl management paths, never to the gateway's
own boot. Call it from run_gateway() — the single in-process boot every entry point
goes through — so the adoption precedes every worker-environment snapshot (cron
builds its env after the scope check, Kanban before it, so fixing this at the
dispatch seam would only fix one of them).

The fail-closed posture is unchanged: with no user manager at all the probe still
reports unavailable and dispatch still refuses. That refusal now names the remedy
in the message the operator actually reads (it is stored as the cron execution's
error), instead of only the symptom.

Fixes #104893
2026-09-08 02:28:21 +05:30
Edizzier 2d52ddc07c fix(cli): planned systemd restarts no longer trigger failure alerts
Salvaged from #104272. Preserve restart and fatal-exit policy while classifying the planned restart code as success. Earlier analysis in #13604 by Justin Kausel.
2026-09-07 07:11:06 -07:00
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 1e6cfaa0d0 simplify(compat): setup/model_switch/nous_subscription — drop 34 re-exports, repoint 14 callers + 26 test files (~100 sites) 2026-09-03 13:33:10 -07:00
Teknium 2031c819fe simplify(compat): gateway — re-land 92d0bd0d73 (reverted by stale-index commit b818085298)
Re-applies the gateway compat removal byte-for-byte; see 92d0bd0d73 for the
full inventory (30 re-exports/aliases + 2 shim modules dropped, 3 shim-only
names re-removed, 24 callers + 34 test files repointed). No new changes.
2026-09-03 13:12:50 -07:00
Teknium b818085298 simplify(compat): doctor/status — drop 13 re-exports + the doctor_* globals() facade (97 names), repoint 6 callers / 13 tests 2026-09-03 13:10:52 -07:00
Teknium 92d0bd0d73 simplify(compat): gateway — drop 30 re-exports/aliases + 2 shim modules, re-remove 3 shim-only names, repoint 24 callers + 34 test files
Per COMPAT_REMOVAL.md (internal import paths are not a stable API):

Re-exports removed
- gateway/run.py: atomic_json_write, load_dotenv, resolve_delivery_transport,
  TurnRunner, merge_pending_message_event, _arm_loop_floor_timer,
  start_loop_liveness_watchdog, DEFAULT_GATEWAY_POST_INTERRUPT_GRACE_TIMEOUT,
  _UNSET (9) — run_* mixins and tests now import from the defining module
  (gateway.delivery / gateway.run_turn_runner / gateway.platforms.base /
  gateway.shutdown_watchdog / gateway.restart / utils).
- gateway/session.py: SessionResetPolicy, normalize_whatsapp_identifier,
  TranscriptReadError, auto_continue_freshness_window (+ "_now & co." noqa
  facade) — gateway/__init__ takes SessionResetPolicy from .config; callers
  take TranscriptReadError from gateway.session_transcript.
- gateway/kanban_watchers.py: _wake_scope_id + "tests import via origin" noqa
  facade; tests import from kanban_watchers_common / _notifier.
- gateway/slash_commands.py: _model_switch_skew_guard, HISTORY_UNREADABLE.
- gateway/stream_consumer.py: escape_code_fences_for_display.
- gateway/platforms/api_server.py: "re-exported" RunIdempotencyStore comment;
  tui_gateway + tests import gateway.platforms.api_server_run_idempotency.
- gateway/platforms/__init__.py: PEP 562 __getattr__/__dir__ lazy QQAdapter /
  YuanbaoAdapter facade (no in-tree importer).
- gateway/startup_watchdog.py: whole re-export shim module deleted; the three
  in-tree callers import hermes_startup_watchdog directly.

Aliases removed
- gateway/platforms/signal.py: SignalAdapter._markdown_to_signal.
- gateway/shutdown_forensics.py: _parse_systemd_duration_to_us.
- gateway/platforms/yuanbao.py: OutboundManager.start_slow_notifier /
  cancel_slow_notifier / get_chat_lock / _chat_locks / CHAT_DICT_MAX_SIZE
  delegates; module-level get_active_adapter / send_yuanbao_direct;
  MarkdownProcessor has_unclosed_fence / ends_with_table_row /
  split_at_paragraph_boundary static pass-throughs (chunk_markdown_text stays —
  it carries yuanbao's chunking policy). tools/send_message_senders +
  tools/yuanbao_tools call YuanbaoAdapter.get_active() / sender.send_direct().

Shim-only names re-removed (earlier review-fix round 96c104c903)
- CapabilityDescriptor.from_platform_entry (+ tests/gateway/relay/test_descriptor_from_entry.py)
- SessionTurnLeaseRegistry.__len__ (+ its test)
- is_relay_media_url KEPT: download() uses it, real internal helper.

Tests that pinned a shim (startup_watchdog re-export identity, api_server
RunIdempotencyStore identity, signal wrapper parity) are dropped; tests that
pinned live behavior are repointed at the implementation.

Note: the gateway/run.py hunk of this change was swept into 00a3cfe5c9 by a
concurrent commit on the shared worktree; this commit carries the rest.

Verified in an isolated worktree (HEAD + this change only): ruff clean,
import-smoke of all 33 touched modules under a fresh HERMES_HOME,
605 passed / 1 skipped across the 39 covering test files.
2026-09-03 13:10:49 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 5a0efda992 review-fix(comments): restore reviewer-named rationale blocks (#4926, #81335, #54220/#56747, #100436, #5057/#6252/#10370/#4665) 2026-09-03 09:32:25 -07:00
Teknium b72baecf69 refactor(hermes_cli): merge simp/r3-19-A1 (round 3) — systemd/launchd docstring compaction, _positive_pid/_launchd_ok helpers 2026-09-02 22:47:13 -07:00
Teknium 8f882b4ba3 refactor(hermes_cli): gateway.py compact docstrings in process-management and run-guard regions (WHYs kept) 2026-09-02 22:42:28 -07:00
Teknium dff07c09d8 refactor(hermes_cli): gateway.py setup wizard — drop redundant local import, join choice lists 2026-09-02 22:39:44 -07:00
Teknium 6824c0a685 refactor(hermes_cli): gateway.py _positive_pid + _launchd_ok helpers; join short wrapped statements 2026-09-02 22:35:25 -07:00
Teknium 8136c1a27f refactor(hermes_cli): gateway.py compact systemd/launchd docstrings + WHY comments; wrap long lines 2026-09-02 22:30:15 -07:00
Teknium 8fceffb842 refactor(hermes_cli): gateway.py compact respawn-watcher template 2026-09-02 22:22:07 -07:00
Teknium 4a1243c8ff refactor(hermes_cli): gateway.py drop blank lines after in-function imports 2026-09-02 22:19:03 -07:00
Teknium b6567d7096 refactor(hermes_cli): merge simp/r3-19-A2 — setup wizard + subcommand dispatch simplification (worker A2) 2026-09-02 22:10:14 -07:00
Teknium a8209adfa1 refactor(hermes_cli): merge simp/r3-19-A1 (round 2) — inline one-caller forwarders, s6 snapshot helper, launchd degrade helper 2026-09-02 22:03:38 -07:00
Teknium 8a2896c0a5 refactor(hermes_cli): gateway.py drop blank lines after local imports in service region 2026-09-02 21:57:13 -07:00
Teknium d95fac5f40 refactor(hermes_cli): merge simp/r3-19-A1 — systemd/launchd region simplification (worker A1) 2026-09-02 21:55:11 -07:00
Teknium 3b17cafe86 refactor(hermes_cli): gateway.py collapse restart/status ladders onto installed-kind helper, pack message tables, compact docstrings 2026-09-02 21:54:59 -07:00
Teknium 52b82a0b67 refactor(hermes_cli): gateway.py extract _s6_gateway_snapshot; table-driven install-scope prompt; compact reaper exclusions 2026-09-02 21:50:40 -07:00
Teknium 65fe6eb78a refactor(hermes_cli): gateway.py setup wizard + subcommands: service-backend dispatch, message tables, phase helpers 2026-09-02 21:48:42 -07:00
Teknium 65cacdd558 refactor(hermes_cli): gateway.py inline one-caller systemd/launchd forwarders; compact identity/linger helpers 2026-09-02 21:38:29 -07:00
Teknium 9f0766f5ec refactor(hermes_cli): gateway.py launchd helpers (_launchd_degrade_or_raise, _launchd_reload_budget, _service_venv_dir); compact reload script 2026-09-02 21:30:36 -07:00
Teknium 8537fc4e7c refactor(hermes_cli): gateway.py orphan reaper / kill helpers — fold survivor filter, compact exclusion comments 2026-09-02 21:16:27 -07:00
Teknium 6adf237876 refactor(hermes_cli): gateway.py systemd lifecycle preamble + graceful-restart phase helper; dedupe scope hint strings 2026-09-02 21:12:00 -07:00
Teknium d75c6f05e7 refactor(hermes_cli): gateway.py multiplexer probe + run guards — collapse branches, compact WHY comments 2026-09-02 21:09:12 -07:00
Teknium 5044c19a58 refactor(hermes_cli): gateway.py profile/linger/legacy-unit helpers (_profile_name_from_home, _user_runtime_dir, _loginctl_enable_linger) 2026-09-02 21:04:55 -07:00
Teknium 259f529f09 refactor(hermes_cli): gateway.py run_gateway region — shared tty/detached probes, collapsed exit-code ladder, storm-backoff env parsing 2026-09-02 21:02:19 -07:00
Teknium 2a16e105db refactor(hermes_cli): gateway.py systemd probe/restart-wait helpers (_parse_kv_pairs, _runtime_state_pid, _systemd_cli_bits) 2026-09-02 20:58:19 -07:00
Teknium acbc228684 refactor(hermes_cli): gateway.py process-management region — /proc scan helper, collapsed guards, watcher comment compaction 2026-09-02 20:56:40 -07:00
Teknium 3e0c47b24b refactor(hermes_cli): gateway.py _PLATFORMS table compact data layout (literal-equal) 2026-09-02 20:35:29 -07:00
Teknium fc87534461 refactor(hermes_cli): gateway.py join wrapped statements that fit in 110 cols 2026-09-02 20:27:34 -07:00
Teknium cc9ef82355 refactor(gateway-cli): dispatch-table subcommands and shared service helpers in hermes_cli/gateway.py
hermes_cli/gateway.py (9178 -> 6906):
- `hermes gateway` subcommand routing: if/elif chain -> _GATEWAY_SUBCOMMANDS dispatch table of
  _cmd_* handlers; _stop_installed_service and _refuse_from_inside_gateway unify the
  stop/restart/uninstall service-stop and self-target guards.
- One systemd unit template; _systemctl_show, systemctl reset+action, is-active probes,
  _installed_service_kind ladder (3 sites), launchctl-list PID probe, launchd bootstrap+kickstart
  helper, legacy-unit removal loop, loop-tick witness ping, ps/wmic line parsers, planned-stop
  marker helper, _CAPTURE_TEXT subprocess kwargs (21 sites), _gw_windows() accessor (15 lazy
  imports).
- run_gateway startup helpers, reaper exclusion set, Windows process listing extracted; wizard
  service actions and platform-setup prompts unified (_setup_service_action).
- Dead: _windows_scheduled_task_running (test-only), _container_systemd_operational (inlined),
  whatsapp/email/matrix built-in status branches (those platforms are registry plugins).
- try/except-pass around single statements -> contextlib.suppress; comments/docstrings compacted
  to the WHY (bounded_probe_run vs subprocess.run on Windows, pythonw launcher-stub filtering,
  raw-record vs validated-probe exclusion in the orphan reaper, Scheduled-Task Ready-vs-Running).
- `hermes gateway --help` byte-identical.
2026-09-02 13:31:44 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
OmniaZ1 d7bda2ad89 fix(gateway): arm the loop-scheduling witness on Windows via TCP loopback
`asyncio.start_unix_server` does not exist on Windows (no AF_UNIX event-loop
support in asyncio), so arming the loop-tick witness in
`loop_heartbeat_forever` raised AttributeError on every native-Windows
gateway start. The broad except swallowed it and recorded
`loop_tick_socket=False`, so every stale-heartbeat probe classified the
gateway as UNKNOWN — never WEDGED, never ALIVE-with-stalled-write. The
two-witness interlock from a1c83ef9 (issue #90502 follow-up) has been
effectively disabled on Windows since it landed: a wedged native-Windows
gateway could never be detected, and an alive one could never be
distinguished from a stalled heartbeat write.

On non-POSIX platforms the witness now arms over a TCP loopback server on
127.0.0.1 (OS-assigned dynamic port) instead:

- same protocol — connect, read one byte "1"
- same semantics — pure in-memory, zero disk I/O, answered only while the
  loop is dispatching, armed by the loop task itself (an awaited
  `asyncio.start_server` is structurally loop-owned exactly like the Unix
  variant, so a wedged loop cannot keep answering pings)
- the assigned port is published in the heartbeat payload as
  `loop_tick_tcp_port`, and `probe_gateway_loop_liveness` prefers the TCP
  witness when the producer published a port, falling back to the AF_UNIX
  socket for POSIX/legacy producers

POSIX behavior is unchanged: the AF_UNIX arm (including the stale-node
sweep) stays gated behind `os.name == "posix"` so the missing attribute can
never raise on Windows again. Legacy heartbeats without `loop_tick_tcp_port`
keep the existing socket-node contract untouched.

Tested end-to-end on native Windows: witness arms, port is published,
`_probe_loop_tick_tcp` answers from an external thread while the loop
dispatches, and the existing loop-liveness suite passes unchanged (the
AF_UNIX structural test still passes — the Unix arm text is preserved
inside the POSIX branch).

Adds two tests pinning the new behavior: an E2E test that arms the TCP
witness and probes it (skipped on POSIX, where the Unix arm is the real
witness), and a structural test that the TCP arm stays awaited on the loop
task and the AF_UNIX arm stays POSIX-gated.
2026-09-02 00:03:36 -07:00
Teknium ab9866bc64 fix(gateway): survive Windows Job-Object teardown across gateway restarts (#48820)
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.

Three surgical changes:

1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
   The inlined watcher now routes the respawned gateway's stray
   stdout/stderr to the same sidecar log gateway_windows._spawn_detached
   uses (DEVNULL only as fallback), so a gateway killed moments after
   respawn leaves a trace. Direct implementation of the 4th repro's
   hardening suggestion (1).

2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
   canonical _spawn_detached, so the respawned gateway's exit-diag /
   lifecycle records show whether it escaped the parent Job Object — a
   job-teardown kill is no longer indistinguishable from any other silent
   death.

3. Post-update resume verifies liveness before vouching
   (hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
   runs the same provisional-hit + 2s-confirmation liveness poll every
   other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
   with all_profiles= for the fleet) before printing ✓, writes the #91675
   start attestation for the verified PIDs, and fails the resume with a
   "restart could not be verified" warning + recovery hint when no stable
   gateway appears. Suggestion (2) of the 4th repro; closes the last
   silent-success hole in the family (#84185 fixed the cold-start leg,
   #91675 the direct-start leg; this is the relaunch leg).

Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.

Fixes the Bug-1 relaunch-trust leg of #48820.
2026-09-01 11:43:36 -07:00
Kshitij Kapoor 44eeef9a83 fix(gateway): apply startup-watchdog config to the already-armed handle
/simplify-code quality+efficiency reviewers (converged, verified): on
the standard 'hermes gateway run' path the argv fast-path arms BEFORE
run_gateway's config bridge executes, and arm_startup_watchdog() is
idempotent — so gateway.startup_watchdog: false and
startup_watchdog_timeout_seconds were dead knobs (env bridged, live
handle untouched). run_gateway now applies the config to the live
handle: disarm on disable; disarm+re-arm on a bridged config timeout so
the fresh handle covers the remaining pre-loop startup with the
configured deadline. config_defaults comment updated to match reality.

E2E (real module, fast-path armed first): disable path disarms the live
handle; timeout path re-arms a fresh handle at 123s.
2026-08-31 14:01:39 -07:00
Kshitij Kapoor d2c3c38e98 fix(gateway): config.yaml surface for the startup watchdog + precise argv arming
Review follow-ups on the salvaged #89750:

- gateway.startup_watchdog / gateway.startup_watchdog_timeout_seconds in
  config_defaults, bridged to the internal HERMES_STARTUP_WATCHDOG env
  vars in run_gateway() (the argv fast-path arms before config can load,
  so env remains the mechanism; config.yaml is the user-facing surface
  per policy — explicit env values still win as operator override).
- hermes_cli/main.py argv sniff now requires the ADJACENT token pair
  'gateway run' instead of independent membership, so unrelated commands
  mentioning both words can't arm a 300s hard-exit timer; profile-flagged
  invocations (-p work gateway run) still arm.
2026-08-31 14:01:39 -07:00
Shannon Sands f5bb1e144d fix(gateway): address startup-watchdog review findings (OOF-298, PR #89750)
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.

P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.

P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.

P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.

P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.

P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.

Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.

Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
2026-08-31 14:01:39 -07:00
Shannon Sands 8a3b6f374d fix(gateway): startup-liveness watchdog for pre-event-loop deadlocks (OOF-298)
A hosted gateway (hermes-doubleam-2568) deadlocked at startup with every
thread parked in futex_wait_queue before the asyncio loop came alive:
zero log lines, /health unreachable — but s6 saw a live PID so it never
respawned the process, and a stale gateway_state.json from the previous
life told every status surface "draining" for ~30 hours.

Every existing liveness backstop assumes startup succeeded: the
loop-liveness watchdog is armed inside the running loop's startup path,
the shutdown watchdog arms at stop(), and the heartbeat file is written
by an asyncio task. None can fire when the process wedges before the
loop exists.

New gateway/startup_watchdog.py: a plain daemon OS thread armed at
process entry (both gateway.run.main() and the `hermes gateway run` CLI
wrapper), disarmed the moment GatewayRunner confirms a live event loop —
the point where the existing loop-liveness watchdog takes over. If
startup neither reaches that milestone nor exits within the deadline
(default 300s; slowest legitimate pre-loop work is the 120s-bounded MCP
discovery wait), the watchdog:

* dumps all-thread stacks via faulthandler,
* appends a JSON record to logs/gateway-startup-watchdog.log,
* records the exit in the NS-608 lifecycle ledger
  (reason=startup_liveness_watchdog) so the next boot classifies it
  instead of reporting an unclean SIGKILL/OOM death,
* os._exit(75) so s6/systemd respawn the process.

Config is env-only (HERMES_STARTUP_WATCHDOG=0 to disable,
HERMES_STARTUP_WATCHDOG_TIMEOUT_S to tune, floor-clamped to 30s):
the watchdog must be armed before config.yaml is loaded — a wedge
during config parsing is exactly in scope — so it cannot depend on
config for its own enablement. Everything is best-effort; a watchdog
failure never affects the startup it observes.

Arm sites are placed after the PID-file/--replace conflict guards so a
--replace loser exiting early never arms a watchdog. Disarm happens even
when the loop guards are config-disabled (gateway.loop_watchdog: false)
— the startup watchdog only covers the pre-loop window, never adapter
connects or steady-state, so WhatsApp pairing / npm cold installs are
unaffected.

Tests: tests/gateway/test_startup_watchdog.py (29 tests — config
resolution, arm/disarm idempotency, fire path with captured exit,
lifecycle-ledger marking, dump record, disable knob).

Fixes OOF-298.
2026-08-31 14:01:39 -07:00
Teknium b0acc558b1 fix(cli): supervised gateway launches skip the sticky active_profile redirect
Generalize the HERMES_S6_SUPERVISED_CHILD supervisor-marker mechanism so
ANY supervised gateway launch (systemd, launchd, Windows Scheduled Task,
external supervisor) skips the active_profile redirect in
_apply_profile_override(). Previously only the s6 container marker was
honored, so a systemd-launched default-profile gateway with
HERMES_HOME=<root> followed the sticky active_profile file and silently
assumed another profile's identity — logging under that profile's tree
and connecting with its Telegram bot token (double-polling a token owned
by that profile's own live gateway).

- hermes_cli/main.py: honor HERMES_SUPERVISED_CHILD (new generalized
  marker), HERMES_S6_SUPERVISED_CHILD (back-compat), INVOCATION_ID
  (systemd; gateway commands only, since it leaks into every descendant
  of systemd-launched processes), and HERMES_GATEWAY_EXTERNAL_SUPERVISOR.
- hermes_cli/gateway.py: export HERMES_SUPERVISED_CHILD=1 in generated
  systemd units (user + system) and the launchd plist.
- hermes_cli/gateway_windows.py: export it from the Scheduled-Task cmd/vbs
  launchers and the windowless respawn env overlay.
- hermes_cli/service_manager.py: export it alongside the s6 sentinel.
- tests: regression coverage for all markers + non-gateway INVOCATION_ID
  neutrality + generated-unit marker presence.

Fixes #74872
2026-08-31 12:05:07 -07:00
webtecnica e0abbc4d18 fix(gateway): isolate PID check and credentials per profile (#74872)
Add _pid_record_belongs_to_current_profile() helper that verifies a
PID record's persisted hermes_home matches the current process. Use
it in get_running_pid() and get_runtime_status_running_pid() so the
default-profile gateway never mistakes another profile's gateway PID
as its own.

In _apply_profile_override(), clear HERMES_HOME instead of returning
early when it points to a profile directory but no --profile flag was
given, letting the sticky active_profile logic resolve the right one.

In _guard_existing_gateway_process_conflict(), detect stale PID files
from other profiles and emit a warning.
2026-08-31 12:05:07 -07:00
Teknium 90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
Nio Thomas ba7743b076 fix(state-db): report corruption instead of "session not found", detect it early
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.

1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
   errors via the existing is_malformed_db_error() and raises 503 at all five
   call sites. delete_session_endpoint was the worst — an unresolvable id
   counted as idempotent success, so DELETE reported it had removed a session
   that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
   quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
   verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
   for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
   before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
   not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
   mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
   _await_gateway_exit() that also re-checks after the final sleep (a PID
   exiting in the last interval must not be SIGKILLed — PID-reuse hazard).

NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.

Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
2026-08-31 09:56:54 -07:00
Kshitij Kapoor 9488950a9d fix(gateway): build reaper exclusion from raw registration records
/simplify-code efficiency reviewer (verified): cleanup_stale=False does
NOT deliver the exclusion the salvaged fix intended — get_running_pid
returns None whenever a record fails liveness VALIDATION (start-time
mismatch, argv drift, lock hiccup) regardless of the flag, which only
controls unlinking. In exactly the at-risk scenario the recorded PID
still never joined the exclusion set.

Exclusion evidence now comes from the RAW pidfile + lock records (no
validation, no unlink side effects); the validated non-destructive
probe is kept for the runtime-status fallback PID. For a KILL exclusion
list this is strictly safer: a stale recorded PID at worst spares one
process for one sweep, while a validation false-negative would
TerminateProcess a live gateway. Regression test reworked to drive the
real function semantics (raw record present, validation rejects);
mutation-checked: removing the raw-record read fails the test.
2026-08-31 21:09:43 +05:30