Commit Graph

1106 Commits

Author SHA1 Message Date
ethernet 283c4f058c fix(packaging): derive root py-modules from the tree in setup.py
The static py-modules list in pyproject.toml drifted from the source
tree each time the root layout changed. An installed wheel then
raised ModuleNotFoundError on import: hermes_state_common, and every
gateway or CLI start failed. The list missed hermes_state_holders,
mini_swe_runner, and the 15 new hermes_state_* modules from this
branch.

setup.py now derives py_modules from the source tree at build time.
setuptools package discovery sees only directories with an
__init__.py, so root single-file modules need py_modules in every
wheel build. setup() kwargs merge with pyproject.toml, and setup.py
is the only legitimate wheel or sdist builder, so the derived list is
the single source of truth.

The nix build is the only wheel consumer. Its source filter keeps
every root .py file, so the build sandbox derives the same set as the
checkout. Editable installs do not read py_modules: build_editable
never runs bdist_wheel.

Verified: wheel and sdist built with HERMES_NIX_BUILD=1 carry all 38
root modules. The guard still blocks builds without the nix env var.
The sealed uv2nix venv from nix build .#default imports
hermes_state_holders, hermes_state_sessions, mini_swe_runner, and
hermes_startup_watchdog.

(cherry picked from commit e23d467b523d50ad088b89b578017996709589cf)

Invariant test updated to pin the derived list (no static py-modules; every root module packaged code imports ships).
2026-09-03 12:22:50 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium c257c0ea2c refactor(cli-main): phase helpers for _has_any_provider_configured, cmd_update preflight, --in/@foreign resume; restore _warn_orphaned_update_autostashes lazy export (update_cmd reaches it via _m()) 2026-09-02 22:12:57 -07:00
Teknium 2d213ce777 refactor(cli-main): drop dead lazy export _warn_orphaned_update_autostashes 2026-09-02 21:37:13 -07:00
Teknium 2f86ce2eb0 refactor(cli-main): split cmd_chat (first-run guard, --query-file, passthrough table), tighten provider/container/cost-guard helpers, dispatch tables for completion/backup; compact comments and docstrings (every WHY kept); repoint xai test to the flow's real module 2026-09-02 21:27:17 -07:00
Teknium 1f2000d959 refactor(cli-main): _session_db ctx manager, _latest_session_id, receipt/light-parser/resume helpers; split _apply_profile_override, cmd_dashboard, select_provider_and_model into phase helpers 2026-09-02 21:06:26 -07:00
Teknium ac99d5fa69 refactor(cli-main): drop 60 dead re-exports, trivial _startup_fast wrappers, orphan comments; table-drive oneshot cleanup 2026-09-02 20:16:55 -07:00
Teknium d1df0f21ab refactor(hermes_cli/main): extract memory/acp/tools/insights/monitoring/skills handlers to main_agent_cmds.py 2026-09-02 17:00:17 -07:00
Teknium aa4ffc98e9 refactor(hermes_cli/main): extract WhatsApp/Slack/Skill-Sync setup wizards to main_platform_setup.py 2026-09-02 16:59:52 -07:00
Teknium 9994d2b113 refactor(hermes_cli/main): unify lazy re-export tables into _LAZY_ATTR_SOURCES; drop 26 unreferenced re-exports 2026-09-02 16:58:15 -07:00
Teknium 40230dcca7 refactor(hermes_cli/main): replace 23 thin cmd_* forwarders with a _forward_command factory table 2026-09-02 16:39:30 -07:00
Teknium 6eae28cb6d refactor(hermes_cli): split cmd_chat prologue into helpers; split build_top_level_parser into flags + chat builders 2026-09-02 16:33:36 -07:00
Teknium ffa08e6591 refactor(hermes_cli/main): extract dashboard/serve support cluster to main_dashboard.py 2026-09-02 16:24:55 -07:00
Teknium d482bd534d refactor(hermes_cli/main): split select_provider_and_model + cmd_dashboard into named helpers; drop unused imports 2026-09-02 16:19:17 -07:00
Teknium 0d32adbbb5 refactor(hermes_cli/main): extract provider-setup helpers to main_provider_setup.py 2026-09-02 16:11:38 -07:00
Teknium 07cd3ba7c9 refactor(hermes_cli/main): extract install/update recovery cluster to main_install_repair.py 2026-09-02 16:08:08 -07:00
Teknium b98a4246c3 refactor(hermes_cli/main): extract desktop (Electron) cluster to main_desktop.py 2026-09-02 15:56:56 -07:00
Teknium 01242369b1 refactor(hermes_cli/main): extract web-UI build cluster to main_web_build.py 2026-09-02 15:45:23 -07:00
Teknium b131e0d87f refactor(hermes_cli/main): extract TUI launcher cluster to main_tui_launch.py 2026-09-02 15:30:43 -07:00
Mike Smith 5edc0c492b fix(cli): skip wrapper-side MCP discovery when chat launches the TUI
Each TUI instance spawned three stdio MCP server copies: one in the
CLI wrapper, one in tui_gateway.entry, one in the slash worker. The
wrapper's copy is dead weight — _launch_tui blocks in subprocess.call
until the TUI exits, so its registered MCP tools are never invoked,
yet the server process (35-85 MB) lives for the whole session.

Root cause: _is_tui_chat_launch() only detected --tui / HERMES_TUI=1,
so bare `hermes` with display.interface: tui fell through to
background MCP discovery in the wrapper while the TUI gateway
(spawned moments later) ran a second discovery.

Fix: _is_tui_chat_launch() now consults _resolve_use_tui() — the exact
TUI-vs-classic decision cmd_chat makes — for chat commands only
(command in {None, "chat"}), leaving mcp serve / gateway / acp / cron
discovery behavior untouched.

Verified: unit tests (RED->GREEN); E2E with a canary stdio MCP server
in a scratch HERMES_HOME counted 2 spawned copies pre-fix vs 1
post-fix (gateway's only), and the wrapper's RSS dropped ~43 MB.

Related: #71928 (same per-process duplication class), #11115 (lazy
non-core discovery).
2026-09-03 03:48:39 +05:30
kshitijk4poor af019a3716 fix(cli): apply the -t/--toolsets MCP spawn filter on every discovery path
The cherry-picked commit added the allowed_mcp_names filter to
discover_mcp_tools(). Since then CLI startup grew a second discovery path —
start_background_mcp_discovery / the deferred desktop start in
hermes_cli/mcp_startup.py — so wiring the filter only into the inline call
would leave `hermes chat -t terminal` (the default backgrounded path) still
spawning every server.

Store the filter once in mcp_startup (set_mcp_server_filter, called from
_prepare_agent_startup from args.toolsets; `all`/`*`/empty clears it) and
have both the inline and the background discovery honor it. The unfiltered
call shape is unchanged so zero-arg test stubs keep working.

Dropped from the original PR: the atexit/SIGINT/SIGTERM oneshot MCP reap —
main already does this in _cleanup_oneshot_runtime() -> shutdown_mcp_servers().

E2E (3 configured stdio servers, real subprocesses, 5 runs median):
no filter 3 spawned / 2.0 s; `-t terminal` 0 spawned / 1 ms.
2026-09-03 03:16:04 +05:30
Teknium ab2fb71b23 refactor(cli): drop 18 update_cmd lazy re-exports with zero references outside update_cmd itself 2026-09-02 13:29:44 -07:00
Teknium 5c125d3bf6 refactor(cli): lift cmd_gui's install+build+stage/swap region into _build_desktop_app; drop 2 unused imports
Zero-back-ref region (inputs desktop_dir/source_mode/npm/env, single output
packaged_executable) becomes a helper returning the swapped-in executable.
Removes the unused functools/_add_accept_hooks_flag imports left by the
parser-builder migration.
2026-09-02 13:29:44 -07:00
Teknium 64f301dfc0 refactor(cli): dedupe the 3x --oneshot dispatch block and chat-arg defaults into helpers
_run_oneshot_from_args replaces three identical confirm+run_and_exit blocks
(main, fast chat, Termux fast cli); _default_to_chat reuses the existing
_set_chat_arg_defaults instead of a second attr table. Also restores the
'bare hermes profile' note as _profile_status's docstring.
2026-09-02 13:29:44 -07:00
Teknium ed8aa5e3e4 refactor(cli): move session browse/status/relative-time helpers from main.py into sessions_cmd.py
_relative_time, _session_status_tag, _annotate_session_statuses,
_session_browse_picker and _size_delta_label (410 LOC) replace the
call-time delegating wrappers in sessions_cmd; main.py re-exports them via
_LAZY_COMMAND_EXPORTS so hermes_cli.main.<name> imports/patches still work.
2026-09-02 13:29:44 -07:00
Teknium 277d51bbc9 refactor(cli): route select_provider_and_model flows through a _PROVIDER_MODEL_FLOWS table
The 20-way if/elif on selected_provider becomes a dict of uniform
flow(config, current_model, args) lambdas plus a _GENERIC_API_KEY_PROVIDERS
frozenset; custom-slug / remove-custom / api-key fallthroughs keep their
order. Lambdas resolve _model_flow_* by name at call time so existing
hermes_cli.main monkeypatches still intercept.
2026-09-02 13:29:44 -07:00
Teknium 853646f143 refactor(cli): move cmd_profile to hermes_cli/profile_cmd.py as a PROFILE_ACTIONS dispatch table
The 604-line 14-way if/elif on profile_action becomes one _profile_<action>
handler each, with per-handler imports derived from the branch's free names.
main.py re-exports cmd_profile and _render_distribution_plan so existing
imports/monkeypatches resolve unchanged.
2026-09-02 13:29:44 -07:00
Teknium 41b658d2b4 refactor(cli): unify duplicate reasoning-effort config helpers into hermes_cli.setup 2026-09-02 13:29:44 -07:00
Teknium 060039d169 refactor(cli): split main() into _build_cli_parser/_parse_cli_args/_default_to_chat helpers
main() is now a ~110-line orchestrator: startup prologue, parser build,
container routing, bpo-9338-safe parse, --version/--yolo/--oneshot, chat
default, dispatch. The two identical plugin add_parser blocks collapse into
_attach_plugin_cli_command; the two default-to-chat attr loops merge. Parser
tree, --help output and set_defaults are unchanged (399 parsers byte-diffed).
2026-09-02 13:29:44 -07:00
Teknium b81ab901f3 refactor(cli): extract the last 16 inline parser groups from main() into hermes_cli/subcommands/
moa, fallback, worktree, browser, secrets, egress, migrate, whatsapp-cloud,
checkpoints, bundles, curator, pets, journey, computer-use, sessions and
completion each become a build_<group>_parser() builder. Closure handlers
that only closed over their own parser moved verbatim; sessions/completion
take the handler by injection. --help/usage/defaults byte-identical for all
399 parsers in the tree (in-process dump before/after).
2026-09-02 13:29:44 -07:00
Teknium 0ccf6714be refactor(cli): drop zero-ref startup-fast wrappers and _build_provider_choices from main.py 2026-09-02 13:29:43 -07:00
Teknium 1398c0f5ca fix(update): stage-and-swap the Desktop rebuild so a failed pack never removes the working app
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).

Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.

- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
  `_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
  retry purge and the integrity self-heal only ever clear the staging tree,
  never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.

Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.

Closes #86443

Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
2026-09-02 06:53:20 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium 16370ae539 fix(tui_gateway): setup.status / setup.runtime_check answer for the requested profile
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.

`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.

The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.

Supersedes #94147 (python half)

Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
Gille 5180601a6a perf(cli): dispatch serve without the full parser tree 2026-09-02 06:06:47 -07:00
Jeffrey Quesnelle c56f8cdd48 Merge pull request #100667 from NousResearch/feat/local-models-squash
feat: local models — managed llama.cpp runtime with one-click desktop  setup
2026-09-01 17:53:28 -04:00
kshitijk4poor 6b46725a17 fix(update): surface leftover update autostashes older than 7 days (#63717)
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
2026-09-02 01:50:21 +05:30
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
JoaoMarcos44 86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
Teknium be597fc730 fix: extend corrupt-config fail-closed guard to gateway, serve, and cron surfaces
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
  construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
  on every surface
2026-09-01 07:00:22 -07:00
embwl0x 6f85df97fd fix(cli): keep quiet prompts interactive 2026-09-01 07:00:22 -07:00
embwl0x 55e7ecd260 test(cli): cover config guard edge cases 2026-09-01 07:00:22 -07:00
embwl0x c335dc734a fix(cli): reject corrupt config in noninteractive runs 2026-09-01 07:00:22 -07:00
lesseradmin 779482598f fix(desktop): pass --disable-setuid-sandbox on the userns launch path
When chrome-sandbox is present but not root-owned 4755, Chromium can still
abort via setuid_sandbox_host even though the namespace sandbox works.
After the userns probe skips sudo, append --disable-setuid-sandbox so
.desktop/no-TTY launches keep the namespace sandbox without a privilege
prompt. Does not add --no-sandbox.

Fixes #51327
2026-09-01 02:32:39 -07:00
4dlt 3a7f2234a6 fix(cli): use Chromium's namespace sandbox when userns is available on Linux
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).

On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.

Fixes #88032
Fixes #51327
Fixes #58593

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 02:32:39 -07:00
kshitijk4poor 26fb8f60e6 fix: anchor checkout detection to the export path, not cwd
Follow-up to the salvaged #92689:

- _profile_export_directory() now proves safety on the export dir's OWN
  ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
  heuristic missed the checkout whenever HERMES_HOME sat inside one but
  the process ran from elsewhere (cron, service manager) — the export
  landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
  violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
  and catch OSError too — a bad profile name or read-only home printed a
  raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
  pollution in the tests/hermes_cli sweep can't divorce monkeypatches
  from the code under test; add regression tests for the cwd-independent
  topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
2026-09-01 01:00:23 -07:00
joaomarcos c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
Kshitij Kapoor d2c3c38e98 fix(gateway): config.yaml surface for the startup watchdog + precise argv arming
Review follow-ups on the salvaged #89750:

- gateway.startup_watchdog / gateway.startup_watchdog_timeout_seconds in
  config_defaults, bridged to the internal HERMES_STARTUP_WATCHDOG env
  vars in run_gateway() (the argv fast-path arms before config can load,
  so env remains the mechanism; config.yaml is the user-facing surface
  per policy — explicit env values still win as operator override).
- hermes_cli/main.py argv sniff now requires the ADJACENT token pair
  'gateway run' instead of independent membership, so unrelated commands
  mentioning both words can't arm a 300s hard-exit timer; profile-flagged
  invocations (-p work gateway run) still arm.
2026-08-31 14:01:39 -07:00