Commit Graph

7648 Commits

Author SHA1 Message Date
teknium1 27df47af5e fix: count service-supervised gateways as running for per-profile cold-start
The per-profile cold-start probe took its running set from the socket-paused
ordinary gateways only. A profile whose gateway is alive under an SCM service
is skipped by the socket pause, so it looked "not running"; with an empty
current-PID list its live start attestation read as dead and the profile
landed in cold_start_profiles. On resume the service was restarted AND a
second, unsupervised gateway was spawned for the same profile.

Build the running set from every profile that had ANY live gateway at
discovery time: paused profiles, profile-mapped processes, and service
gateways' profiles.

Review finding: service-supervised running profile was cold-started beside
its restarted SCM service (double gateway).
2026-09-15 04:13:44 -07:00
Hermes Agent 8fbae2813b fix(update-windows): per-profile cold-start runs after the relaunch and fails loud
- Run `_cold_start_attested_profiles` AFTER the paused profiles are relaunched,
  so a sibling that fails to cold-start can never keep the profiles that were
  running from coming back.
- A profile that stays down raises like the active-profile cold-start does;
  the merged outcome then marks the update incomplete instead of printing
  success with a gateway still offline.
- Trim the salvaged tests to the two invariants (dead-attested default beside
  a live beta is cold-started under its own home; every dead-attested profile
  is cold-started when nothing runs, active first). The "no attestation →
  token unchanged" case is the pre-existing behaviour already covered by
  test_pause_skips_cold_start_plan_when_desktop_owns_lifecycle.
2026-09-15 04:13:44 -07:00
kshitijk4poor c728e6583e fix(update-windows): per-profile cold-start probes its own home and runs after the active spawn
Review follow-ups on the per-profile obligation:
- Order: the active profile's cold-start guard is fleet-wide (any live
  gateway ⇒ done), so a sibling spawned first would have left the active
  profile down. Per-profile spawns now run after it.
- A token is built even when the active plan owes nothing (clean exit,
  autostart not installed), so a dead-attested sibling still rides on it.
- Readiness for a per-profile spawn is probed in THAT home's identity files
  (`_live_gateway_pids(home=)` → `get_running_pid(home/"gateway.pid")`), not
  the fleet: a still-running sibling no longer vouches for a dead spawn. The
  wait/confirm helpers share one probe function.
- The new PID is attested in the profile home (`_write_start_attestation(...,
  home=)`) so a death after the CLI exits reaches the next update, and an
  already-live profile is not spawned twice.
2026-09-15 04:13:44 -07:00
kshitijk4poor c8686224d3 fix(update-windows): cold-start every dead-but-attested profile, not only when nothing runs
`_pause_windows_gateways_for_update` built the attested cold-start plan only when the
ALL-profile running PID list was empty, so a default gateway that died after a ✓ beside a
still-running `beta` never received a cold-start obligation and stayed down after the
update (#110959, fifth review thread on #110020).

- gateway_windows: thread `home` through `_start_attestation_path` → `_read_start_attestation`
  → `_attested_pid_exited_cleanly`/`_attested_dead` → `attested_death_generation(pids, home=)`
  and `_consume_start_attestation(gen, home=)`; `_spawn_detached(home=)` builds the argv/env
  for that profile home (`--profile` derived by `_launcher_settings`). Defaults unchanged.
- update_cmd_windows: `_record_attested_cold_start_profiles` evaluates every
  `profiles_to_serve(multiplex=True)` profile that is not running (Desktop-owned installs only,
  active profile left to the existing plan so nothing is spawned twice) and records
  `token["cold_start_profiles"] = {name: generation}` — a sibling key, so `profiles[name]`
  stays an int for relaunch/verify. `_cold_start_attested_profiles` runs on resume before the
  ordinary relaunch, spawns under the profile home, waits for readiness, consumes exactly that
  generation; one profile's failure never aborts the others.
- tests: default dead-attested + beta running → token carries the obligation, resume spawns
  under the default home and consumes its marker while beta is relaunched; no attested
  profile → token and resume unchanged.
2026-09-15 04:13:44 -07:00
teknium1 6dc6c92ea2 fix: release the profile MCP stderr handle before rename too
rename_profile moves the profile directory while this process may still
hold the cached per-profile mcp-stderr.log handle (left behind by a
completed probe or a running server). On Windows a directory containing
an open file cannot be renamed, the same WinError class delete_profile
now avoids. On other platforms the stale handle stayed cached under the
old home key, so a later probe on the renamed profile opened a second
handle and a new profile re-created under the old name wrote its MCP
stderr into the renamed profile's log. Release the scoped handle next to
the multiplexer unroute, mirroring delete_profile.

Review finding: rename_profile missed the sibling surface of the
delete_profile handle release.
2026-09-15 04:12:41 -07:00
tarkilhk 68dea35e0e fix(mcp): release profile stderr handles before deletion 2026-09-15 04:12:41 -07:00
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
Hermes Agent d3c5539056 refactor(linux-desktop-entry): collapse the fall-through into one guard
The salvaged fix replaced the early return with a `pass` + `elif`;
fold it into a single `if primary and rerouted is not None` so the
probe is reached by construction and the WHY lives in one comment.
2026-09-15 04:11:55 -07:00
Octopustank 93116dc820 linux-desktop-entry: note the None-path equivalence inline
resolve_hermes_bin's rerun hides argv[0], so a non-None reroute could
only come from PATH — which the first call would already have returned.
primary is None therefore implies rerouted is None, and the fall-through
is equivalent to return rerouted or primary; only the durable-wrapper
probe below can still find anything on that branch. (KeyArgo's review
asked why the branch was merged rather than dropped.)
2026-09-15 04:11:55 -07:00
Octopustank a2a7cf922a fix(linux): converge the desktop entry Exec on the durable wrapper
_resolve_hermes_bin_for_desktop_entry returned the resolver's None
outright when neither argv[0] nor PATH yielded a launcher (a cold
relaunch under `python -m` with a stripped PATH), skipping the
known-wrapper probe its own docstring describes. The persisted
`Exec=` then flipped to the bare module form, while a DE- or
terminal-launched context renders the wrapper form.

Every flip rewrites ~/.local/share/applications/hermes.desktop on the
next launch. gnome-shell 50.x removes a ShellApp from id_to_app on any
.desktop content change without checking its state; if the app was
still STARTING, the last strong reference drops and a later GC-triggered
dispose trips shell-app.c's `state == STOPPED` assertion — taking the
whole Wayland session down (observed crash, 50.4-1.fc44).

Probe the installer's known wrapper locations on the None path too, so
the entry converges on the durable wrapper wherever one exists and only
falls back to the module form when none does. A missing entry being
(re)created can still change the file regardless — this removes the
gratuitous rewrites, not the rescan trigger itself.
2026-09-15 04:11:55 -07:00
teknium1 37e4cd5faa fix(updater): a probe that cannot be spawned stays advisory instead of failing the update
bounded_probe_run collapses spawn failure and timeout into one None, so a venv
interpreter that exists but cannot be executed (PermissionError, ENOEXEC, fork
failure) was reported as "timed out before reporting import health" -- a bogus
warning on the git path and sys.exit(1) on the ZIP path -- and the
`except OSError: return {}` branch beneath it was dead. Let the Popen error
propagate (opt-in flag, default unchanged for the other probes) so "we could not
run our own probe" says nothing about the checkout and no longer blocks the
update, while a spawned child that hangs is still a verdict.

Restores the non-fatal test the earlier rewrite dropped, now driving the real
bounded_probe_run/Popen: red on the previous head, green here.

Review finding: spawn failure of the import probe reported as a fatal timeout.
2026-09-15 04:11:11 -07:00
teknium1 da1fb702c3 fix(updater): import probe children never resolve external secret sources
The critical-module import probe (`_critical_module_import_failures`) imports
`run_agent`, whose module-level `load_hermes_dotenv()` resolves every enabled
external secret source. The main-process skip only lived in `hermes_cli.main`'s
own dotenv call (`load_external_secrets=sys.argv[1:2] != ["update"]`), so the
probe child ran op/bws/command helpers with a 120s per-source budget inside the
120s probe and a healthy install reported "critical-module probe still fails to
import after updating: timed out before reporting import health" (#110823).

Move the argv check into `_early_recovery._should_skip_external_secret_sources`,
which every dotenv load already consults, and stamp `sys.argv = ['hermes',
'update']` into the probe so its imports inherit the updater contract.
Invariant test spawns the real probe against a home whose configured helper
touches a marker: red on main, green here.
2026-09-15 04:11:11 -07:00
KoNit-K e74c29de55 fix(updater): bound import probe teardown 2026-09-15 04:11:11 -07:00
teknium1 8e0b1a2e47 fix(update): keep the config-migration purge and reload inside the fail-open guard
_purge_stale_hermes_modules evicts hermes_cli.* but not root modules
(hermes_constants, utils, toolsets, ...), so the post-purge
`from hermes_cli.config import ...` re-executes the NEW config.py against
OLD cached root modules. When a pull adds a root-module symbol that
config.py imports at module level, that raises ImportError. The import
sat outside the step's try/except, so the error escaped
_run_post_update_maintenance and aborted the fleet restart -- where
origin/main printed the "Could not check config version / run hermes
config migrate" fallback and continued. Move the purge, reload and
import inside the existing try so the step stays fail-open.

Test: a cached hermes_constants lacking get_process_hermes_home (the
symbol hermes_cli/config.py imports at module level) must print the
fallback and return instead of raising. Red before, green after.

Review finding: fail-open -> fail-closed flip; post-purge hermes_cli.config import outside the try let ImportError abort post-update maintenance.
2026-09-15 04:10:27 -07:00
teknium1 24a623e557 fix(update): purge stale Hermes modules before post-pull config migrations
The pinned-list reload (tools_config) fixes the reported symbol; any
future symbol a pull adds to any module a migration imports at call time
(agent.skill_utils, hermes_cli.toolset_scope, ...) would fail the same
way. Evict every cached Hermes module at the migration entry point with
_purge_stale_hermes_modules — the class fix the fleet-restart phase
already relies on — so the migration graph is rebuilt from the new tree.

Tests: the reporter's exact shape (cached tools_config lacking
_configurable_keys, on-disk v44 config with an explicit platform
toolset list) must still migrate v44 -> v45 through the updater's own
entry point. Existing entry-point tests stub the purge like the rest of
the update suite does.
2026-09-15 04:10:27 -07:00
liuhao1024 91a3ab4c85 fix(update): reload the cached tools_config before post-pull migrations
The updater process is the pre-pull process: hermes_cli.tools_config cached in sys.modules lacks symbols the pull added, so a migration that imports them at call time fails with ImportError and the config silently stays at the old version (#111271).
2026-09-15 04:10:27 -07:00
teknium1 e43f2f6816 fix(cli): retire the 0.0 monotonic sentinel in the remaining input-mode throttles
Review follow-up on #91651: _recover_terminal_input_modes and the termios
drift check used the same `now - 0.0 < interval` idiom as the repaint
throttles. Their windows (0.5s / 1.0s) are unreachable in practice, but
converting them to the None sentinel retires the bug class instead of the
instance, so nobody "simplifies" a None back to 0.0 later. Also adds the
missing regression test for _invalidate, the throttle the PR title is about.
2026-09-15 04:08:53 -07:00
Teknium eab377a407 fix(cli): first repaint no longer swallowed when monotonic clock is small
time.monotonic() counts from an arbitrary epoch (boot on Linux). Two CLI
repaint throttles used 0.0 as the never-fired sentinel, so on a freshly
booted VM (CI runners, containers) now - 0.0 < min_interval suppressed
the FIRST repaint ever requested:

- _schedule_focus_regain_redraw: min_interval=60 suppressed the first
  focus-regain redraw whenever uptime < 60s — the exact failure in CI
  run 32494557030 (test_focus_regain_redraw_is_rate_limited, both
  attempts red on a fresh runner, green everywhere else).
- _invalidate: same 0.0 sentinel; a first spinner/stream repaint inside
  the first 250ms of uptime was droppable the same way.

Both now use None as the never-fired sentinel. Regression test pins
monotonic()=3.0 with min_interval=60 and asserts the first redraw fires.
2026-09-15 04:08:53 -07:00
teknium1 b2577df807 fix(dashboard): accept command-scoped NOPASSWD sudo for system gateway actions
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".

Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.

Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
2026-09-15 04:08:00 -07:00
teknium1 05fb879609 fix(dashboard): share the fleet's sudo posture; trim to two invariant tests
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
2026-09-15 04:08:00 -07:00
Baris Sencan eaf700c67e fix(dashboard): elevate system-scope gateway lifecycle actions (#110820)
The Restart Gateway button spawned `hermes gateway restart` as the dashboard's
own user. On a systemd *system* install the CLI refuses every lifecycle verb
below root (`_require_root_for_system_service`), so the button could never work:
the refusal landed in `~/.hermes/logs/gateway-restart.log` while the endpoint
reported a started action. `start`/`stop` shared the defect.

`_spawn_hermes_action` now prefixes `sudo -n` for gateway restart/start/stop when
the action resolves to the system unit. Scope comes from the CLI's own picker
(`_select_systemd_scope`) evaluated against the profile the action addresses, so
a host with a user unit installed keeps the unprivileged path and root never
shells out to sudo. With no passwordless path the request fails with an
actionable message instead of reporting a started action whose child refuses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 04:08:00 -07:00
teknium1 a3d3daf7e9 fix: hand back a custom-unit gateway even when identify answers systemd/launchd
gateway_declares_external_supervisor treated the control-socket identify answer
as final whenever `supervisor` was non-empty and only accepted "external".
_detect_supervisor() answers "systemd" as soon as INVOCATION_ID is set (and
"launchd" under the XPC name), so a gateway run by a custom, non-canonical
systemd unit / launchd agent with `ExecStart=... gateway run --external-supervisor`
— the documented contract — was classified non-external and `hermes gateway
restart` still took the stop + foreground run_gateway path that stamps the CLI's
PID and wedges every respawn (#110637). The update path's argv check on the same
gateway already said external-supervisor, so the two restart paths disagreed.

Any self-declared supervisor other than "manual" now means hand back (the
declaring supervisor owns the respawn); otherwise the argv/state-file marker
decides as before. The drain-failure message no longer claims a timeout when
SIGUSR1 could not be sent at all.

Review finding: socket-first detection classified a custom systemd unit's
--external-supervisor gateway as manual and took the wedge path.
2026-09-15 04:07:13 -07:00
teknium1 8731bb91f5 refactor(gateway): move the supervised-restart handback into gateway_supervised_restart.py
The handback logic was appended to the hermes_cli/gateway.py facade; it now lives in a
topical sibling. Supervisor detection also reads the gateway's own declaration (control
socket `identify` -> supervisor: "external", then the live argv marker, then the argv the
gateway stamped into gateway_state.json) so a gateway whose command line cannot be read via
psutil is still handed back rather than SIGTERMed and shadowed by a foreground run.

Tests trimmed to the two invariants (handback with fresh-PID success; either failure branch
never takes ownership) plus the plain-manual control. Docs: `hermes gateway restart` is now
part of the --external-supervisor contract.
2026-09-15 04:07:13 -07:00
liuhao1024 9ef2dd0be1 fix(cli): keep the supervisor the sole restart owner on both handback branches
Address review feedback on the SIGUSR1 handback:

- A graceful exit no longer reports success by itself: the pidfile is
  polled for the supervisor's replacement (bounded wait, identity via
  get_running_pid's lock+PID liveness check, freshness via != old PID)
  before the restart is called done. An unloaded/broken supervisor now
  surfaces as a failed restart instead of success printed over a dead
  gateway — the same contract launchd_restart enforces through
  _wait_for_launchd_service_pid.

- A drain timeout no longer falls back to SIGTERM + foreground run: that
  fallback recreated the competing-owner wedge (#110637) this handback
  exists to remove. Both failure branches keep the supervisor as the
  sole restart owner and fail loudly (exit 1) instead.

Credit: both lifecycle branches flagged by @ehz0ah in review.
2026-09-15 04:07:13 -07:00
liuhao1024 8c286af7ef fix(cli): hand externally-supervised gateways back to their supervisor on restart
A custom launchd agent (plist outside the canonical ai.hermes.gateway path)
is invisible to _installed_service_kind_for, so 'hermes gateway restart'
fell through to the manual stop + foreground run_gateway fallback. The
foreground run stamps the restart CLI's own PID into gateway.pid, and every
KeepAlive respawn of 'gateway run --external-supervisor' then refuses with
'Gateway already running (PID <restart>)' — the gateway stays down until
the restart process is killed (#110637).

Mirror the update path (_prepare_profile_gateway_update_restart): trust the
--external-supervisor argv marker on the live gateway and SIGUSR1 it
(_graceful_restart_via_sigusr1) so it drains, exits, and lets the supervisor
relaunch it. Drain failure falls back to the existing stop/restart path.

Fixes #110637
2026-09-15 04:07:13 -07:00
teknium1 127214a66b fix: keep fleet-restart marker while a receipt-owed gateway is down
_marker_only_restart_obsolete cleared the marker when every row from
collect_fleet_versions() was current. At CLI startup the probe runs without
pre_restart_pids, so a gateway the restart phase stopped and never brought
back produces no row at all (dead-pid records are skipped, no DOWN
classification possible). With alpha current and beta down the marker was
discharged and beta's catch-up restart never happened.

Before clearing, also require every gateway identity latest.json owes
(plan.runtimes / fleet entries) to be covered by a current row; the
rows-only rule applies only when the receipt names no gateways. The
identity extraction is shared with _live_fleet_covers_receipt.

Review finding: marker discharged with a DOWN sibling gateway absent from the startup probe.
2026-09-15 04:06:08 -07:00
teknium1 2fcb50c50e test(update): trim marker-discharge tests to two invariants; document the identity rule
The five #105417 tests collapse into one positive (fleet current at expected_sha
discharges the marker) and one parametrized negative (stale row / empty probe /
unknown identity / checkout moved past the marker keep the warning). Both are red
on origin/main. collect_fleet_versions' docstring now names why a live PID alone
is not a `current` row (#110420): write_runtime_status re-stamps pid/code_sha for
whatever process writes it, so the state-file fallback must pass
live_gateway_pid_for_home like the inventory already does (#109680).
2026-09-15 04:06:08 -07:00
linmukong 8130274b37 fix(update): discharge fleet_restart_pending when the fleet provably serves expected_sha
_pending_fleet_restart_needed() returns True on marker existence alone. A
supervisor-level gateway restart (`systemctl --user restart`, launchctl, an ops
script) never passes through _clear_fleet_restart_pending_marker(), so a marker
written by a pulled update survives a restart that DID bring every live gateway
to the new code — and every later CLI call prints the interrupted-update warning
forever, a permanent false positive that trains operators to ignore it.

The comment on that branch is right that "an older receipt cannot discharge that
unknown obligation" — but the obligation is not unknown. The marker records its
own expected_sha, so it can be verified against the live fleet directly, with no
receipt involved. _live_fleet_covers_receipt() cannot substitute: it is
receipt-anchored and returns False when `owed` is empty, which is exactly the
marker-only case.

Hold the marker to the same evidence bar _live_fleet_covers_receipt() applies to
a receipt: at least one row, and every row a `current` gateway under a known
profile whose code_sha equals the marker's expected_sha, with the checkout HEAD
not moved past the marker.

Still keeps the marker (warns) on:
  - stale / down rows — the restart genuinely is owed
  - an all-`unknown` fleet — pre-code-identity gateways cannot prove currency
    (same conservatism as #88848/#74973)
  - a marker with no expected_sha — nothing to verify against
  - a newer pull that moved the checkout — it owns a fresh obligation
  - a probe that raises or answers empty — no proof either way

Discharging deletes the marker only after every row passes; the historical
receipt is left intact, since a supervisor restart is still not a successful
update.

Tests: discharge on a provably-current fleet; keep on stale, empty probe,
unknown identity, and newer-pull-moved-checkout (which asserts the fleet is
never probed once HEAD has moved).
2026-09-15 04:06:08 -07:00
KoNit-K dab6fe2933 fix(update): drop untrusted fleet fallback version
When the state-file writer cannot be verified as the profile gateway, null code_version with the SHA so the row does not display a self-reported identity claim.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-15 04:06:08 -07:00
KoNit-K 0eaa51674f fix(update): verify fleet state fallback gateway identity 2026-09-15 04:06:08 -07:00
kshitijk4poor 288fdc1a4c fix(auth): accept a non-production Portal's own inference host when the operator selected it
A token minted by a non-production Portal is meant to be spent at that environment's own
inference gateway, and the Portal's refresh response names that host. The allowlist applied
to Portal-returned inference URLs was production-only, so the value was refused as "not in
allowlist" and healed to the production host — a token the production Portal never issued,
sent to the production gateway, which 401s it. Every hosted non-production instance hit
this on every gateway turn once #108319 made the deploy-wide NOUS_INFERENCE_BASE_URL
invisible inside a routed profile scope (by design, #65941).

The widening is keyed on the operator's trusted HERMES_PORTAL_BASE_URL override, never on
the stored portal_base_url: when that override names a Portal outside the production
allowlist, any https host under the Nous domain is accepted; otherwise the strict production
set stands. So a poisoned auth.json cannot widen the set, a production-Portal session that
finds a foreign inference URL in its state is still refused and healed, and the bearer can
only ever go to a Nous-owned host. No environment is named in code. Because the override is
read through the profile scope (previous commit), each multiplexed profile decides for
itself.

Validation: 4 invariant tests (accepted only under a non-production override; look-alike
domains, dotless suffix and http still refused; stored portal alone does not widen; the
decision follows the profile scope under multiplex) — the new-behaviour ones red on the
previous commit. Main's existing validation tests are unchanged and green. Live receipt for
the symptom and the fixed chain on a hosted instance: #111589.

Based on #102863 and its rebase onto the decomposed auth_nous.py in #111589, whose
portal-keyed pairing this replaces with the same behaviour and no environment literals.

Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-09-15 16:31:53 +05:30
kshitijk4poor 9366f59533 fix(auth): read the Portal env override through the profile scope
_nous_portal_env_override() is the Portal twin of _nous_inference_env_override(): it is
reached on every routed turn (_nous_effective_routing -> _NousRuntimeResolve._reload_routing)
and still read os.environ, so under gateway.multiplex_profiles a secondary profile POSTed
its refresh token to the DEFAULT profile's HERMES_PORTAL_BASE_URL while its inference
override (scoped in #108319) was correctly isolated. Same shape as that fix: get_secret with
the UnscopedSecretError -> environ fallback for the unscoped CLI / default-profile path.

The two inline os.getenv fallbacks in _nous_effective_routing were dead — the
env_portal_override branch two lines later always wins when the variable is set — so they
are removed rather than scoped twice. The CLI-only readers (_nous_device_code_login,
nous_billing, the dashboard OAuth router) run single-profile and stay on environ.

Deployments that set HERMES_PORTAL_BASE_URL only in the process env must carry it in each
served profile's .env, the rule #108319 established for NOUS_INFERENCE_BASE_URL.

Flagged by @webtecnica in the #108319 review thread. One invariant test: scoped value wins,
scoped miss stays a miss, unscoped keeps environ; red on main.
2026-09-15 16:31:53 +05:30
teknium1 bb745a0e9b fix(goals): failed quality gates re-run every boundary instead of replaying a status fingerprint
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
2026-09-15 03:59:50 -07:00
teknium1 b55767be82 fix: judge wait_on_pid race no longer raises out of evaluate_after_turn
`_apply_wait_directive` probed `_pid_alive` and then called `wait_on`,
which re-checks liveness and raises ValueError. A pid exiting between
the two checks surfaced as an exception from `evaluate_after_turn`; all
three callers swallow it, so the goal silently lost that boundary's
continuation. Do a single check by catching wait_on's own ValueError
and falling through to the continue decision.

Review finding: TOCTOU between _pid_alive pre-check and wait_on re-check drops the continuation.
2026-09-15 03:59:10 -07:00
teknium1 ee07fcd4d7 fix(goals): a judge wait_on_pid naming an unobservable pid continues instead of parking
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
2026-09-15 03:59:10 -07:00
KoNit-K c70db196eb fix(goals): reject dead wait-on PIDs 2026-09-15 03:59:10 -07:00
teknium1 9f30a0a269 fix(import-sync): never clobber a locally edited imported skill; keep digest on errors
The docs promised "skills you created or modified under the import category
yourself are never clobbered", but sync replaced any destination whose name was
in `imported_skills`, regardless of what was there now — an imported skill the
user had since edited was silently overwritten on the next source change.

The manifest now stores `imported_skills` as {name: digest-of-the-copy-we-wrote}
and `--sync` refreshes a destination only while it still matches that digest;
a locally modified copy records a `conflict` ("modified locally — not
refreshed") and is skipped. Pre-digest manifests (a plain list) keep the old
trusted behaviour for one more cycle and are upgraded on the next import.

`sync_imported_agents` also refreshed the source digest after a run with
errors, so the failed items were never retried; the previous digest is kept
whenever the report has errors.
2026-09-15 03:58:44 -07:00
Teknium 6fcd011c01 Inspired by ChatGPT Work: keep imported agent setups in sync (hermes import-agent --sync)
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:

- Every successful import registers its source + a content digest of
  everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
  files changed since the last run (digest compare; unchanged = no-op).
  Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
  sync; user-created skills under the import category keep conflict
  semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
  trigger (or leak into) a sync.

Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
2026-09-15 03:58:44 -07:00
teknium1 aa75d3724f fix(stream-json): verbatim text deltas, closed protocol on init failure, per-call tool keys
- on_text_delta dropped whitespace-only deltas, so concatenating the `text`
  events no longer reproduced the answer (a newline between paragraphs was
  lost). Only None/"" (the turn-end sentinel) is skipped now.
- The emitter was attached only after credentials + agent init succeeded, so a
  missing key or unknown provider exited 1 with an EMPTY stdout and the
  provider error rendered through ChatConsole (stdout). The emitter is now
  built before _ensure_runtime_credentials/_init_agent; that path closes the
  protocol with init + a failed `result` (exit_code 1, error) and the
  credential error goes to stderr whenever stdout is machine-readable
  (tool_progress_mode == "off", i.e. -Q and stream-json).
- _tool_started was keyed by tool name, so concurrent same-name calls
  clobbered each other's start time; key on tool_call_id when the caller
  passes one and surface it on tool_use/tool_result.

Live: `hermes chat -q … --format stream-json` with no provider and with a dead
custom base_url both yield pure JSONL (`system` + `result`, exit 1).
2026-09-15 03:53:13 -07:00
Alan 1657a1ce2d feat(cli): add --format stream-json for structured JSONL output
Adds a --format flag to hermes chat single-query mode. stream-json
emits newline-delimited JSON events (init, text, tool_use, tool_result,
result envelope with token stats + exit code) to stdout for CI
pipelines and external tooling. Session ID stays on stderr.

Salvaged from PR #12278 by @ProDrifterDK onto current main, including
the follow-up commit enforcing the single-query contract (implies
quiet, rejects --tui, emits a final result record with exit code 130
on interrupt).
2026-09-15 03:53:13 -07:00
teknium1 616a3ce036 fix(kanban): actor sentinel for synthesized runs; two invariant tests
A plain ``profile=None`` default could not tell "caller named the actor"
from "read the card" — an unassigned card is a legitimate None actor, and
the row re-read would silently kick back in for it. Use a module sentinel
so only callers that did not pass an actor fall back to the card row.

Trims the salvaged test file to two invariants in the existing review
lifecycle suite: the never-claimed handoff names the implementer (red on
main) and an unassigned card's synthesized run keeps ``profile=NULL``.
2026-09-15 03:51:28 -07:00
Kevin Rajan ff8f725c5a fix(kanban): attribute synthesized review-handoff run to the implementer
request_review captures the implementer before rewriting tasks.assignee to
the reviewer, but _synthesize_ended_run re-reads assignee off the mutated
row, so the zero-duration review_requested run names the reviewer instead
of the handoff's actor. Pass the captured implementer through a new
optional profile keyword on _end_or_synthesize_run/_synthesize_ended_run.

Fixes #111064
2026-09-15 03:51:28 -07:00
teknium1 7abe9502ee fix(kanban): worker liveness and kills require the spawn-time start fingerprint
tasks.worker_pid outlives a reboot; afterwards the number can belong to any
process. _pid_alive answered from bare existence, so reclaim_stale_claims kept
extending the claim of a "live" stranger (stuck running task), and
enforce_max_runtime / _terminate_reclaimed_worker SIGTERM'd then SIGKILL'd it.

_set_worker_pid now records gateway.status.get_process_start_time(pid) as
tasks.worker_started_at (additive column, NULL on legacy rows). _worker_alive
(pid, started_at) is the liveness check every reader uses (reclaim, defer,
reconcile, crash sweep, max-runtime, archive, reopen invalidation); a live pid
whose fingerprint disagrees is a recycled PID: treated as dead, never
signalled (termination reports pid_recycled). Legacy rows without a
fingerprint keep the existence answer until their next spawn.

Row hermes_cli/kanban_db_dispatch.py:458 (lane4_high) confirmed by tracing:
the SIGKILL at :461 was gated only on _pid_alive.
2026-09-15 03:47:15 -07:00
teknium1 5ca670b398 fix(dashboard): profile-routed routers run under the profile's secret scope; console send never writes the process env
_config_profile_scope bound only HERMES_HOME, so GET /api/config?profile=B
expanded B's `${VAR}` refs to the dashboard (DEFAULT) profile's plaintext
credentials, WS /api/console commands for B saw the default's keys wherever B
lacked one, and the audio speak-stream synthesis thread re-resolved the TTS key
unscoped. Console `send` for B went further: send_cmd._load_hermes_env copied
B's .env into the shared os.environ with override=True, so every later
default-profile read saw B's tokens.

- web_server_profiles._config_profile_scope binds home + hydrated secret scope
  for a named profile and flips the process to fail-closed multi-profile
  hosting (same activation as the tui_gateway); the dashboard's own profile then
  runs under its frozen-launch-env scope. A single-profile dashboard stays
  unscoped (systemd / op-run credential injection keeps working).
- GET /api/tools/toolsets computes the per-toolset "configured" flag inside
  the profile scope (it was read after the scope closed).
- web_routers/mcp._profile_secret_scope now just delegates (no second copy of
  the composition); audio `_produce` runs the whole synthesis body under the
  requesting profile's scope, not only the resolve step.
- send_cmd._load_hermes_env targets the installed secret scope when one is
  active (gateway.config._getenv reads the scope first), os.environ only for
  the standalone CLI.
2026-09-15 03:46:29 -07:00
teknium1 fdd5995ecb fix(tui_gateway): serve fails closed when hosting a second profile home; profile RPCs bind the full runtime scope
`hermes serve` / the Desktop backend hosted many profile homes (session
profile_home, the `profile` RPC param, hosted rooms) but never called
agent.secret_scope.set_multiplex_active, so every unscoped get_secret read for
a secondary silently returned the LAUNCH profile's os.environ value, and
@_profile_scoped bound only HERMES_HOME: `config.get full` for profile B
expanded B's `${VAR}` refs to the default profile's plaintext credentials,
model.options listed the default's env-keyed providers, llm.oneshot billed the
default's auxiliary key.

- tui_gateway/launch_profile_policy.py (was launch_terminal_policy.py): the
  first time _profile_home registers a non-launch home the process freezes
  the launch env and flips get_secret to fail closed
  (activate_multi_profile_hosting); launch_secret_scope composes the launch
  profile's .env + external sources over that frozen env so systemd / op-run
  injection survives the flip while a secondary never sees it.
- model_switch._profile_runtime_scope_tokens is the ONE composer for
  home + secret + terminal scope: a named profile binds its own files; the
  launch profile binds its frozen-env scope once multiplexing is active and
  stays unscoped in a single-profile process (legacy os.environ precedence).
  _profile_scoped, _profile_scoped_rpc, _session_profile_runtime_scope,
  _bind_build_profile_scopes and _prepare_turn_input all go through it.
- Hosted-room / Group Chat turns for a DEFAULT-profile member in a
  `multiplex_profiles: true` gateway no longer die at agent build with
  UnscopedSecretError: `profile_home is None` was treated as "no scope"
  in _start_agent_build._build and _prepare_turn_input.
- llm.oneshot runs under the session's (or params.profile's) scope;
  _lap_builtin_rows / _overlay_has_creds / _provider_has_credentials read
  provider keys through _scoped_key_env instead of raw os.environ;
  methods_groups._profile_execution_policy resolves the hosted-room policy
  (which reads provider credentials) under the profile's full scope.

Live repro (real `hermes serve`, two homes, config.get {key: full, profile: b}):
base  a_ref: <A_VALUE>  b_ref: ${B_ONLY_TOKEN}  env_ref: <ENV_INJECTED>
head  a_ref: ${A_ONLY_TOKEN}  b_ref: <B_VALUE>  env_ref: ${ENV_INJECTED_TOKEN}
Control (one home, --single): launch config still resolves env_ref from os.environ.
2026-09-15 03:46:29 -07:00
teknium1 8a8c3634e8 fix(kanban): scope the delegated-child write fence to the lineage's board root
HERMES_DELEGATED_CHILD_CONTEXT=1 is deliberately carried into every shell/
execute_code subprocess a delegate_task child spawns (the fence must survive
exec so a grandchild `hermes kanban complete` cannot promote itself). But the
readers treated the bare flag as "fence every Kanban DB": kanban_db_connect
opened ANY board ?mode=ro and write_txn refused ANY mutation. A subagent
running a Kanban reproduction against a scratch HERMES_HOME therefore got a
silently read-only board with a misleading "descendants require an
initialized board" error; only one lane in the retrospective ever discovered
why (deleg_15dac332), every earlier kanban repro ran degraded.

The marker's value is now the fenced board ROOT (kanban_home() at spawn) and
readers deny only paths under that root or the dispatcher-pinned
HERMES_KANBAN_DB (kanban_path_is_fenced). In-process children and a legacy
"1" marker still fence everything; an inherited path marker is never
re-derived, so a grandchild that moved HERMES_HOME cannot unfence the real
board. Owner-gate tests (test_kanban_descendant_scope, cron env isolation,
kanban CLI exit status) are unchanged and green.
2026-09-15 03:45:41 -07:00
teknium1 87ce653d1d feat(relay): migrate legacy HERMES_NEMO_RELAY_ATIF_*/ATOF_* vars into a validated relay-plugins.toml
3fad83df31 (Aug 11) moved Relay exporter config to a plugins.toml
selected by HERMES_NEMO_RELAY_PLUGINS_TOML. A .env still carrying the legacy
exporter vars and no TOML logs ONE warning and initialises no exporters, so
users who followed the earlier docs lost every trace silently (the
maintainer's stopped Aug 20, noticed Sep 14; five multiplexed profiles on
the same box carry the same eight vars today).

- `hermes_cli/relay_plugin_migrate.py`: build the document from the
  `nemo_relay.observability` dataclasses (`ComponentSpec(...).to_dict()`,
  so the `type = "file"` sink discriminator is emitted), validate it by
  activating it through `nemo_relay.plugin.initialize` + `clear_async`,
  write `<home>/relay-plugins.toml` (tomli_w when installed, minimal emitter
  otherwise), set HERMES_NEMO_RELAY_PLUGINS_TOML in that .env, and comment
  the legacy lines out (never delete). Defaults mirror the removed plugin so
  files land where they used to.
- `hermes update` runs it for the default home AND every live named profile
  (each writes its own TOML) as a best-effort post-update step, with a loud
  notice; `hermes migrate relay [--all-profiles] [--no-validate]` runs it on
  demand.
- The runtime WARNING and the `hermes doctor` finding now say "NO traces
  are being exported" and name the exact command and file path.
- Docs: environment-variables.md + built-in-plugins.md carry the migration
  note and a complete plugins.toml example including `type = "file"`.
2026-09-15 03:44:36 -07:00
teknium1 395e4248d0 fix(gateway): surface secondary WhatsApp/Relay skipped under multiplex instead of a silent continue
Under gateway.multiplex_profiles, `_start_one_profile_adapters` skipped
Platform.RELAY / Platform.WHATSAPP for secondaries with a bare `continue`,
and the startup "not being served" WARNING only covered platforms the
PRIMARY skipped. Four secondaries on one live box had WHATSAPP_ENABLED=true
and nothing in the log, status file, or `hermes gateway status` said the
channel was dead.

- `_note_unserved_secondary_platform`: one INFO per (profile, platform)
  naming the reason (shared process-level ingress owned by the default) and
  the remedy (enable it on the default profile, or disable it here), plus a
  `<profile>:<platform>` runtime-status stamp (state=disabled,
  error_code=multiplex_shared_ingress).
- `_start_secondary_profiles` folds those platforms into the loud WARNING
  when NO profile (default included) runs them.
- `hermes gateway status --profile X` prints
  `whatsapp: not served under multiplex (shared ingress owned by default)`
  from that stamp; /api/status excludes `disabled` entries from the
  platforms degraded verdict (informational, not a fault).
- Docs: multi-profile-gateways.md gets the shared-ingress rule.
2026-09-15 03:44:36 -07:00
teknium1 fd303c0137 fix(gateway): 'decline' survives the config.yaml load path; Telegram forwards it; wizard offers it
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.

Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.

`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
2026-09-15 03:44:13 -07:00
teknium1 8b9df066e2 test(gateway): trim block-loop wording coverage to two invariants; CLI keeps the needs_input distinction
The salvaged parametrized set collapsed to one neutral case (transient) plus
the needs_input positive control. hermes kanban block now mirrors the
notifier: 'needs a human decision' only when the block was typed
needs_input, 'orchestration attention needed' otherwise.
2026-09-15 03:42:00 -07:00