Commit Graph

16712 Commits

Author SHA1 Message Date
teknium1 0ea0c53e89 fix: record a known exit status when the reader's wait() raises
The early-EOF reaper made _finish_reader return without publishing when
wait() raises, so the session stays tracked for later reconciliation.
That is right for the pipe path (_reconcile_local_exit can still reap via
session.process), but PTY sessions have no session.process: when
ptyprocess.wait raises (waitpid ECHILD after isalive() already reaped the
child) the exitstatus is known, yet poll() reported "running" forever.

Only leave the session tracked when exit_code() is still None; otherwise
record the known status and finish as before.

Review finding: PTY session whose pty.wait raises stays in _running forever (fail-open regression vs main)
2026-09-15 04:23:32 -07:00
teknium1 dc4a2aeb10 test(process): trim early-EOF reaper test to its two invariants
Drop the wait-timeout call-shape assertion; the invariants are that the
session records the real exit code and that a failed reap does not publish a
false completion.
2026-09-15 04:23:32 -07:00
fangliquan 1bfbeff4e5 fix(process): reap children after early stdout EOF 2026-09-15 04:23:32 -07:00
teknium1 8bc5894a3a docs(skills): ip-as-logo follows the modern section order; add skill tests
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.

Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
2026-09-15 04:20:15 -07:00
teknium1 b531023622 fix(gateway): drop the redundant whole-block lease; one invariant test per atom; document the watchdog env vars
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.

Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.

Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
2026-09-15 04:19:19 -07:00
Ayush Nangia bcab2dd6eb fix(gateway): persist startup watchdog stacks before stderr
A detached or service-managed gateway can have a blocked stderr. The watchdog exit escort then hard-exits after ten seconds before the file-based traceback is reached, leaving only the metadata record. Write the durable dump first and pin the ordering with a regression test.
2026-09-15 04:19:19 -07:00
Kevin Rajan cef69276ae fix(gateway): hold startup-watchdog progress leases across state.db auto-maintenance
Construction-time maybe_auto_archive / maybe_auto_prune_and_vacuum ran
synchronously with no report_startup_progress lease. A multi-minute
VACUUM of a large state.db accrues near-zero CPU, so the startup
watchdog misread it as a parked deadlock and killed the attempt with
exit 75, live-locking gateway restarts. Renew the lease per long step
(prune, orphan sweep, VACUUM, archive) since leases clamp at 900 s,
plus one lease around the gateway maintenance block.

Fixes #111092
2026-09-15 04:19:19 -07:00
teknium1 5358780264 fix: keep the carried pre-admission input across the waited lease reload
carry_unadmitted_user_message appends the interrupted turn's user row to the
early result's in-memory history only; it is never persisted because that turn
never owned the lease. When the follow-up turn also has to wait for the lease
(the common case: the other process that made the first turn wait is usually
still busy), admit_durable_turn_lease reloads conversation_history from the DB
after admission and replaced the caller's list wholesale, so the tagged row was
dropped from both the model input and state.db. Re-append the tagged rows that
have no _row_id after the reload so the follow-up turn sees and flushes them.

Review finding: waited-reload branch of admit_durable_turn_lease discarded the
_persist_after_admission_interrupt row carried from the aborted turn.
2026-09-15 04:18:34 -07:00
teknium1 fae3030fa6 refactor: carry the unadmitted user message from the lease sibling, not the facade
Move the pre-admission carry-forward out of run_conversation (facade) into
agent/turn_facade_lease.py::carry_unadmitted_user_message next to the early
result it repairs, and drop the extra turn-start flush in build_turn_context:
the follow-up turn's normal turn-start persist already writes the marked row
because _db_flush_collect no longer stamps it as durable (live probe: exactly
one A row in state.db after two flushes). Trim to two invariant tests
(carry-forward with metadata; flushed exactly once); the hard-stop negative is
covered by the E2E probe in the PR body.
2026-09-15 04:18:34 -07:00
Ayush Nangia 24ae31f7e7 fix(gateway): retain pre-admission interrupted input 2026-09-15 04:18:34 -07:00
teknium1 9efa50071a fix: match find/read-tool dynamic words only as unquoted command arguments
The dynamic-shell-word rules fired on any `-del*`/`-exec*` substring after a
`find` anywhere in the segment, so quoted predicate arguments the shell never
expands were flagged: `find . -name 'log-del*'`, `find . -name 'pre-exec*.sh'`,
`find src -path '*-exec[0-9]*'`, and `echo find . -{delete,print}`.

Three changes close that:
- `find` must be the command word (_CMDPOS anchored, same as mkfs/rm/dd) and
  the dynamic word must start a whitespace-delimited token (`(?<!\S)`), for
  both the find rule and the rg/sort/ag/man program-option rule.
- Both rules now scan the quote-masked variant (_QUOTE_MASKED_DANGEROUS_
  DESCRIPTIONS, same _mask_quoted_prose used by the positionless hardline
  rules) so glob characters inside quotes are data, not expansion.
- _iter_shell_command_starts no longer treats the `{` inside a brace-expansion
  word (`-{delete,print}`) as a brace-group opener; it split the word across a
  marked start so `echo x; find . -{delete,print}` matched nothing. A brace
  group opener is `{` as its own word (after whitespace or a separator).

The quoted-name cases join the inert parametrized test; the separator case
joins the dangerous one. Still two parametrized functions.
2026-09-15 04:17:42 -07:00
Teknium 06a3a98751 fix: gate dynamic shell words in approval checks
Port from openai/codex#39159: require approval when shell expansion could synthesize destructive find flags or program-executing read-tool options.
2026-09-15 04:17:42 -07:00
teknium1 bc3df8a4d5 fix: NT-namespace guard fires before every sibling resolve (checkpoint, ACP bridge, @file:)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.

The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
2026-09-15 04:17:31 -07:00
teknium1 07e6461dfe fix: guard search_tool's root and trim the NT-namespace tests to two invariants
search_tool resolved its root via _resolve_path_for_task BEFORE the
NT-namespace check saw it (review finding): on Windows the resolve is the
SMB-auth trigger, on POSIX the task-base join hides the prefix from the
resolved-path denylist. The raw-string guard now runs first there too,
and the guard rides the existing top-level agent.file_safety import
instead of two function-local imports.

The tool-layer chokepoint test now covers all four entries and proves
none of them touched Path/_resolve_path_for_task/realpath before the
refusal; the 12 form-matrix tests collapse into one blocked/allowed
invariant over both classifiers.
2026-09-15 04:17:31 -07:00
Teknium faf71eb4c1 Inspired by Claude Code: file tools reject Windows NT-namespace paths (NTLM leak hardening)
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:

- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
  raw-string check (never resolves — resolving IS the leak trigger).
  Wired as the first check in get_read_block_error() and the write
  denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
  _check_sensitive_path (covers write_file_tool + patch_tool), before
  the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
  local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
  forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
2026-09-15 04:17:31 -07:00
teknium1 aaf38bb2fb test: service stubs carry the profile the pause path now reads
The fix-pass made _pause_windows_gateways_for_update read s.profile from
every discovered service gateway; two quarantine tests build services with
SimpleNamespace stubs that lacked the field the real WindowsGatewayService
always has.
2026-09-15 04:13:44 -07:00
teknium1 27df47af5e fix: count service-supervised gateways as running for per-profile cold-start
The per-profile cold-start probe took its running set from the socket-paused
ordinary gateways only. A profile whose gateway is alive under an SCM service
is skipped by the socket pause, so it looked "not running"; with an empty
current-PID list its live start attestation read as dead and the profile
landed in cold_start_profiles. On resume the service was restarted AND a
second, unsupervised gateway was spawned for the same profile.

Build the running set from every profile that had ANY live gateway at
discovery time: paused profiles, profile-mapped processes, and service
gateways' profiles.

Review finding: service-supervised running profile was cold-started beside
its restarted SCM service (double gateway).
2026-09-15 04:13:44 -07:00
teknium1 b19e5786e2 test(update-windows): spawn stubs accept the per-profile home argument
_spawn_detached now forwards home= to _build_gateway_argv so the post-update
cold-start can launch another profile's gateway; the two windows_only spawn
tests stubbed _build_gateway_argv with a zero-arg lambda and raised TypeError
on the Windows runner.
2026-09-15 04:13:44 -07:00
Hermes Agent 8fbae2813b fix(update-windows): per-profile cold-start runs after the relaunch and fails loud
- Run `_cold_start_attested_profiles` AFTER the paused profiles are relaunched,
  so a sibling that fails to cold-start can never keep the profiles that were
  running from coming back.
- A profile that stays down raises like the active-profile cold-start does;
  the merged outcome then marks the update incomplete instead of printing
  success with a gateway still offline.
- Trim the salvaged tests to the two invariants (dead-attested default beside
  a live beta is cold-started under its own home; every dead-attested profile
  is cold-started when nothing runs, active first). The "no attestation →
  token unchanged" case is the pre-existing behaviour already covered by
  test_pause_skips_cold_start_plan_when_desktop_owns_lifecycle.
2026-09-15 04:13:44 -07:00
kshitijk4poor c728e6583e fix(update-windows): per-profile cold-start probes its own home and runs after the active spawn
Review follow-ups on the per-profile obligation:
- Order: the active profile's cold-start guard is fleet-wide (any live
  gateway ⇒ done), so a sibling spawned first would have left the active
  profile down. Per-profile spawns now run after it.
- A token is built even when the active plan owes nothing (clean exit,
  autostart not installed), so a dead-attested sibling still rides on it.
- Readiness for a per-profile spawn is probed in THAT home's identity files
  (`_live_gateway_pids(home=)` → `get_running_pid(home/"gateway.pid")`), not
  the fleet: a still-running sibling no longer vouches for a dead spawn. The
  wait/confirm helpers share one probe function.
- The new PID is attested in the profile home (`_write_start_attestation(...,
  home=)`) so a death after the CLI exits reaches the next update, and an
  already-live profile is not spawned twice.
2026-09-15 04:13:44 -07:00
kshitijk4poor c8686224d3 fix(update-windows): cold-start every dead-but-attested profile, not only when nothing runs
`_pause_windows_gateways_for_update` built the attested cold-start plan only when the
ALL-profile running PID list was empty, so a default gateway that died after a ✓ beside a
still-running `beta` never received a cold-start obligation and stayed down after the
update (#110959, fifth review thread on #110020).

- gateway_windows: thread `home` through `_start_attestation_path` → `_read_start_attestation`
  → `_attested_pid_exited_cleanly`/`_attested_dead` → `attested_death_generation(pids, home=)`
  and `_consume_start_attestation(gen, home=)`; `_spawn_detached(home=)` builds the argv/env
  for that profile home (`--profile` derived by `_launcher_settings`). Defaults unchanged.
- update_cmd_windows: `_record_attested_cold_start_profiles` evaluates every
  `profiles_to_serve(multiplex=True)` profile that is not running (Desktop-owned installs only,
  active profile left to the existing plan so nothing is spawned twice) and records
  `token["cold_start_profiles"] = {name: generation}` — a sibling key, so `profiles[name]`
  stays an int for relaunch/verify. `_cold_start_attested_profiles` runs on resume before the
  ordinary relaunch, spawns under the profile home, waits for readiness, consumes exactly that
  generation; one profile's failure never aborts the others.
- tests: default dead-attested + beta running → token carries the obligation, resume spawns
  under the default home and consumes its marker while beta is relaunched; no attested
  profile → token and resume unchanged.
2026-09-15 04:13:44 -07:00
teknium1 ef44c1b73f fix: restore TestBridgeDispatch, trim #39797 tests, drop dead execution_guidance_text param
Why: the rebase conflict resolution in tests/tools/test_model_tools.py
deleted the unrelated TestBridgeDispatch class (3 tests from 73163e3);
it is restored verbatim from main with TestBrowserRetrievalHints after it.

The fix had five tests for one invariant: the OPENAI_MODEL_EXECUTION_GUIDANCE
check duplicated test_phantom_tool_references, and the two static-schema
"toolset-neutral" checks are now folded into test_silent_without_web_tools,
which runs _apply_dynamic_schemas over the real browser_navigate/browser_cdp
schemas so the rendered descriptions are what is asserted.

execution_guidance_text() no longer takes valid_tool_names: the guidance is
toolset-neutral, so the parameter was ignored; the single caller in
agent/system_prompt.py and its test are updated.
2026-09-15 04:13:13 -07:00
teknium1 ac63d0eea5 fix(agent): execution guidance and browser hints drop the web_search stripper; tests assert the invariant
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (3733e4aff5) matched
nothing and were dead; the function now returns the neutral text for
every toolset and its phantom-tool test asserts "no web tool named"
instead of the removed sentence. model_tools ports the PR's hint layer
into main's _DYNAMIC_SCHEMA_REWRITERS table (browser_navigate +
browser_cdp) rather than a second pass after it.

Tests: the two browser_cdp registry tests were re-added by the PR but
main pruned them in 39975613b13b4; replaced with one schema-neutrality
invariant. Exact-wording assertions ("lightweight retrieval tool",
"appropriate permitted retrieval/search tool") were change detectors and
are dropped. tools-reference.md row updated to the new schema text.
2026-09-15 04:13:13 -07:00
Kevin Yin ccf380f634 fix(agent): respect permitted web retrieval guidance 2026-09-15 04:13:13 -07:00
teknium1 6dc6c92ea2 fix: release the profile MCP stderr handle before rename too
rename_profile moves the profile directory while this process may still
hold the cached per-profile mcp-stderr.log handle (left behind by a
completed probe or a running server). On Windows a directory containing
an open file cannot be renamed, the same WinError class delete_profile
now avoids. On other platforms the stale handle stayed cached under the
old home key, so a later probe on the renamed profile opened a second
handle and a new profile re-created under the old name wrote its MCP
stderr into the renamed profile's log. Release the scoped handle next to
the multiplexer unroute, mirroring delete_profile.

Review finding: rename_profile missed the sibling surface of the
delete_profile handle release.
2026-09-15 04:12:41 -07:00
tarkilhk 68dea35e0e fix(mcp): release profile stderr handles before deletion 2026-09-15 04:12:41 -07:00
teknium1 2500f4ee56 test(cli): sessions open-failure test creates the store so the read-only empty path does not short-circuit 2026-09-15 04:12:13 -07:00
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
Octopustank a2a7cf922a fix(linux): converge the desktop entry Exec on the durable wrapper
_resolve_hermes_bin_for_desktop_entry returned the resolver's None
outright when neither argv[0] nor PATH yielded a launcher (a cold
relaunch under `python -m` with a stripped PATH), skipping the
known-wrapper probe its own docstring describes. The persisted
`Exec=` then flipped to the bare module form, while a DE- or
terminal-launched context renders the wrapper form.

Every flip rewrites ~/.local/share/applications/hermes.desktop on the
next launch. gnome-shell 50.x removes a ShellApp from id_to_app on any
.desktop content change without checking its state; if the app was
still STARTING, the last strong reference drops and a later GC-triggered
dispose trips shell-app.c's `state == STOPPED` assertion — taking the
whole Wayland session down (observed crash, 50.4-1.fc44).

Probe the installer's known wrapper locations on the None path too, so
the entry converges on the durable wrapper wherever one exists and only
falls back to the module form when none does. A missing entry being
(re)created can still change the file regardless — this removes the
gratuitous rewrites, not the rescan trigger itself.
2026-09-15 04:11:55 -07:00
teknium1 f876ba60ff fix(logging): unavailable-log notice is re-armed by a successful write, not by open()
The "named once" guarantee only held when open() itself failed. In the reported
case open() succeeds and write/seek/flush raise EIO, so every record went
handleError -> stream=None -> next emit reopens -> _open() reset
_unavailable_reported -> the path line was printed once per record (25 for 25).
Reset the flag in emit() only after a record actually reached the stream; that
is the moment the destination has recovered.

Adds the open-succeeds/I-O-fails case to the test file: red on the previous
head, green here.

Review finding: EIO on write after a successful reopen printed the path per record.
2026-09-15 04:11:11 -07:00
teknium1 37e4cd5faa fix(updater): a probe that cannot be spawned stays advisory instead of failing the update
bounded_probe_run collapses spawn failure and timeout into one None, so a venv
interpreter that exists but cannot be executed (PermissionError, ENOEXEC, fork
failure) was reported as "timed out before reporting import health" -- a bogus
warning on the git path and sys.exit(1) on the ZIP path -- and the
`except OSError: return {}` branch beneath it was dead. Let the Popen error
propagate (opt-in flag, default unchanged for the other probes) so "we could not
run our own probe" says nothing about the checkout and no longer blocks the
update, while a spawned child that hangs is still a verdict.

Restores the non-fatal test the earlier rewrite dropped, now driving the real
bounded_probe_run/Popen: red on the previous head, green here.

Review finding: spawn failure of the import probe reported as a fatal timeout.
2026-09-15 04:11:11 -07:00
teknium1 b847c1ea5e fix(logging): name an unavailable log file once instead of silently dropping records
Builds on the salvaged EIO suppression: the reporter asked that logging
"degrade gracefully and identify the affected path". Print one stderr line
naming the file and errno when the stream first fails, reset the flag when
`_open()` succeeds again so a later failure is reported anew. The contributor's
test is replaced by one invariant: five records through a stream raising EIO
produce zero tracebacks and exactly one path mention, and the next emit
recovers into the real file.
2026-09-15 04:11:11 -07:00
teknium1 da1fb702c3 fix(updater): import probe children never resolve external secret sources
The critical-module import probe (`_critical_module_import_failures`) imports
`run_agent`, whose module-level `load_hermes_dotenv()` resolves every enabled
external secret source. The main-process skip only lived in `hermes_cli.main`'s
own dotenv call (`load_external_secrets=sys.argv[1:2] != ["update"]`), so the
probe child ran op/bws/command helpers with a 120s per-source budget inside the
120s probe and a healthy install reported "critical-module probe still fails to
import after updating: timed out before reporting import health" (#110823).

Move the argv check into `_early_recovery._should_skip_external_secret_sources`,
which every dotenv load already consults, and stamp `sys.argv = ['hermes',
'update']` into the probe so its imports inherit the updater contract.
Invariant test spawns the real probe against a home whose configured helper
touches a marker: red on main, green here.
2026-09-15 04:11:11 -07:00
KoNit-K e74c29de55 fix(updater): bound import probe teardown 2026-09-15 04:11:11 -07:00
teknium1 8e0b1a2e47 fix(update): keep the config-migration purge and reload inside the fail-open guard
_purge_stale_hermes_modules evicts hermes_cli.* but not root modules
(hermes_constants, utils, toolsets, ...), so the post-purge
`from hermes_cli.config import ...` re-executes the NEW config.py against
OLD cached root modules. When a pull adds a root-module symbol that
config.py imports at module level, that raises ImportError. The import
sat outside the step's try/except, so the error escaped
_run_post_update_maintenance and aborted the fleet restart -- where
origin/main printed the "Could not check config version / run hermes
config migrate" fallback and continued. Move the purge, reload and
import inside the existing try so the step stays fail-open.

Test: a cached hermes_constants lacking get_process_hermes_home (the
symbol hermes_cli/config.py imports at module level) must print the
fallback and return instead of raising. Red before, green after.

Review finding: fail-open -> fail-closed flip; post-purge hermes_cli.config import outside the try let ImportError abort post-update maintenance.
2026-09-15 04:10:27 -07:00
teknium1 24a623e557 fix(update): purge stale Hermes modules before post-pull config migrations
The pinned-list reload (tools_config) fixes the reported symbol; any
future symbol a pull adds to any module a migration imports at call time
(agent.skill_utils, hermes_cli.toolset_scope, ...) would fail the same
way. Evict every cached Hermes module at the migration entry point with
_purge_stale_hermes_modules — the class fix the fleet-restart phase
already relies on — so the migration graph is rebuilt from the new tree.

Tests: the reporter's exact shape (cached tools_config lacking
_configurable_keys, on-disk v44 config with an explicit platform
toolset list) must still migrate v44 -> v45 through the updater's own
entry point. Existing entry-point tests stub the purge like the rest of
the update suite does.
2026-09-15 04:10:27 -07:00
liuhao1024 91a3ab4c85 fix(update): reload the cached tools_config before post-pull migrations
The updater process is the pre-pull process: hermes_cli.tools_config cached in sys.modules lacks symbols the pull added, so a migration that imports them at call time fails with ImportError and the config silently stays at the old version (#111271).
2026-09-15 04:10:27 -07:00
teknium1 e43f2f6816 fix(cli): retire the 0.0 monotonic sentinel in the remaining input-mode throttles
Review follow-up on #91651: _recover_terminal_input_modes and the termios
drift check used the same `now - 0.0 < interval` idiom as the repaint
throttles. Their windows (0.5s / 1.0s) are unreachable in practice, but
converting them to the None sentinel retires the bug class instead of the
instance, so nobody "simplifies" a None back to 0.0 later. Also adds the
missing regression test for _invalidate, the throttle the PR title is about.
2026-09-15 04:08:53 -07:00
Teknium eab377a407 fix(cli): first repaint no longer swallowed when monotonic clock is small
time.monotonic() counts from an arbitrary epoch (boot on Linux). Two CLI
repaint throttles used 0.0 as the never-fired sentinel, so on a freshly
booted VM (CI runners, containers) now - 0.0 < min_interval suppressed
the FIRST repaint ever requested:

- _schedule_focus_regain_redraw: min_interval=60 suppressed the first
  focus-regain redraw whenever uptime < 60s — the exact failure in CI
  run 32494557030 (test_focus_regain_redraw_is_rate_limited, both
  attempts red on a fresh runner, green everywhere else).
- _invalidate: same 0.0 sentinel; a first spinner/stream repaint inside
  the first 250ms of uptime was droppable the same way.

Both now use None as the never-fired sentinel. Regression test pins
monotonic()=3.0 with min_interval=60 and asserts the first redraw fires.
2026-09-15 04:08:53 -07:00
teknium1 6467172db9 fix(webhook): validate the coalesce block on hot-reloaded dynamic routes
_reload_dynamic_routes runs from _handle_webhook and admitted routes through
_dynamic_route_allowed without validate_coalesce_config, so a dynamic route
with a malformed coalesce block (no key, non-numeric window) raised inside
the request handler on its first event instead of being rejected up front
like a static route. The admission check now runs the same validator and
skips (warns on) the offending route. The parametrized rejection test covers
the dynamic path too; upstream-source reference dropped from the module
docstring.
2026-09-15 04:08:12 -07:00
Teknium 3aeb1736c5 Port from RooCodeInc/Roomote#1478: per-route webhook event coalescing
Rapid distinct events on the same logical entity (five pushes to one PR,
a burst of ticket edits, a flapping alert) each carry a fresh delivery ID,
so the idempotency cache cannot suppress them and every event wakes a
separate agent run. Roomote solved this for PR review tasks by keeping one
durable review task per PR and superseding stale heads; this ports the
same debounce-and-supersede pattern to the generic webhook adapter.

New opt-in per-route 'coalesce' block: events group by a payload-derived
key, each new event replaces the pending one and re-arms a quiet-window
timer (window_seconds, default 30), bounded by max_wait_seconds (default
300) past the group's first event so a steady stream cannot starve
dispatch. The settled group dispatches ONE agent run on the latest
event's payload/prompt/delivery templates, with a note when earlier
events were superseded. Pending groups flush on disconnect. Startup
validation rejects missing keys, non-positive windows, and the
deliver_only+coalesce combination.

Rebase onto the decomposed webhook adapter (salvage, #92066):
- Coalescing lives in a topical sibling, gateway/platforms/webhook_coalesce.py
  (WebhookCoalescer + validate_coalesce_config); webhook.py only wires it in
  (__init__, _validate_route, disconnect, _handle_webhook) and splits main's
  _dispatch_agent_run into the HTTP-response wrapper plus _spawn_agent_run,
  shared by the immediate and coalesced paths.
- Review finding (unresolved key fields collapsed unrelated entities into one
  group): an event whose rendered key still contains a {placeholder} is now
  dispatched immediately instead of coalesced; documented.
- Review finding (flush-on-disconnect vs process exit): disconnect() awaits
  the handoff of flushed runs; the docs claim is scoped to adapter disconnect
  and states that a hard kill loses the current window's buffer.
- cron_job + coalesce is rejected like deliver_only + coalesce (cron_job
  landed on main after the PR branched).
- Tests trimmed from 17 to 4 (validation parametrized; debounce/supersede/
  independent groups/duplicate-first in one behavioural test; max-wait +
  unresolved-key; flush-on-disconnect).
2026-09-15 04:08:12 -07:00
teknium1 b2577df807 fix(dashboard): accept command-scoped NOPASSWD sudo for system gateway actions
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".

Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.

Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
2026-09-15 04:08:00 -07:00
teknium1 05fb879609 fix(dashboard): share the fleet's sudo posture; trim to two invariant tests
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
2026-09-15 04:08:00 -07:00
Baris Sencan eaf700c67e fix(dashboard): elevate system-scope gateway lifecycle actions (#110820)
The Restart Gateway button spawned `hermes gateway restart` as the dashboard's
own user. On a systemd *system* install the CLI refuses every lifecycle verb
below root (`_require_root_for_system_service`), so the button could never work:
the refusal landed in `~/.hermes/logs/gateway-restart.log` while the endpoint
reported a started action. `start`/`stop` shared the defect.

`_spawn_hermes_action` now prefixes `sudo -n` for gateway restart/start/stop when
the action resolves to the system unit. Scope comes from the CLI's own picker
(`_select_systemd_scope`) evaluated against the profile the action addresses, so
a host with a user unit installed keeps the unprivileged path and root never
shells out to sudo. With no passwordless path the request fails with an
actionable message instead of reporting a started action whose child refuses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 04:08:00 -07:00
teknium1 a3d3daf7e9 fix: hand back a custom-unit gateway even when identify answers systemd/launchd
gateway_declares_external_supervisor treated the control-socket identify answer
as final whenever `supervisor` was non-empty and only accepted "external".
_detect_supervisor() answers "systemd" as soon as INVOCATION_ID is set (and
"launchd" under the XPC name), so a gateway run by a custom, non-canonical
systemd unit / launchd agent with `ExecStart=... gateway run --external-supervisor`
— the documented contract — was classified non-external and `hermes gateway
restart` still took the stop + foreground run_gateway path that stamps the CLI's
PID and wedges every respawn (#110637). The update path's argv check on the same
gateway already said external-supervisor, so the two restart paths disagreed.

Any self-declared supervisor other than "manual" now means hand back (the
declaring supervisor owns the respawn); otherwise the argv/state-file marker
decides as before. The drain-failure message no longer claims a timeout when
SIGUSR1 could not be sent at all.

Review finding: socket-first detection classified a custom systemd unit's
--external-supervisor gateway as manual and took the wedge path.
2026-09-15 04:07:13 -07:00
teknium1 8731bb91f5 refactor(gateway): move the supervised-restart handback into gateway_supervised_restart.py
The handback logic was appended to the hermes_cli/gateway.py facade; it now lives in a
topical sibling. Supervisor detection also reads the gateway's own declaration (control
socket `identify` -> supervisor: "external", then the live argv marker, then the argv the
gateway stamped into gateway_state.json) so a gateway whose command line cannot be read via
psutil is still handed back rather than SIGTERMed and shadowed by a foreground run.

Tests trimmed to the two invariants (handback with fresh-PID success; either failure branch
never takes ownership) plus the plain-manual control. Docs: `hermes gateway restart` is now
part of the --external-supervisor contract.
2026-09-15 04:07:13 -07:00
liuhao1024 9ef2dd0be1 fix(cli): keep the supervisor the sole restart owner on both handback branches
Address review feedback on the SIGUSR1 handback:

- A graceful exit no longer reports success by itself: the pidfile is
  polled for the supervisor's replacement (bounded wait, identity via
  get_running_pid's lock+PID liveness check, freshness via != old PID)
  before the restart is called done. An unloaded/broken supervisor now
  surfaces as a failed restart instead of success printed over a dead
  gateway — the same contract launchd_restart enforces through
  _wait_for_launchd_service_pid.

- A drain timeout no longer falls back to SIGTERM + foreground run: that
  fallback recreated the competing-owner wedge (#110637) this handback
  exists to remove. Both failure branches keep the supervisor as the
  sole restart owner and fail loudly (exit 1) instead.

Credit: both lifecycle branches flagged by @ehz0ah in review.
2026-09-15 04:07:13 -07:00
liuhao1024 8c286af7ef fix(cli): hand externally-supervised gateways back to their supervisor on restart
A custom launchd agent (plist outside the canonical ai.hermes.gateway path)
is invisible to _installed_service_kind_for, so 'hermes gateway restart'
fell through to the manual stop + foreground run_gateway fallback. The
foreground run stamps the restart CLI's own PID into gateway.pid, and every
KeepAlive respawn of 'gateway run --external-supervisor' then refuses with
'Gateway already running (PID <restart>)' — the gateway stays down until
the restart process is killed (#110637).

Mirror the update path (_prepare_profile_gateway_update_restart): trust the
--external-supervisor argv marker on the live gateway and SIGUSR1 it
(_graceful_restart_via_sigusr1) so it drains, exits, and lets the supervisor
relaunch it. Drain failure falls back to the existing stop/restart path.

Fixes #110637
2026-09-15 04:07:13 -07:00
teknium1 127214a66b fix: keep fleet-restart marker while a receipt-owed gateway is down
_marker_only_restart_obsolete cleared the marker when every row from
collect_fleet_versions() was current. At CLI startup the probe runs without
pre_restart_pids, so a gateway the restart phase stopped and never brought
back produces no row at all (dead-pid records are skipped, no DOWN
classification possible). With alpha current and beta down the marker was
discharged and beta's catch-up restart never happened.

Before clearing, also require every gateway identity latest.json owes
(plan.runtimes / fleet entries) to be covered by a current row; the
rows-only rule applies only when the receipt names no gateways. The
identity extraction is shared with _live_fleet_covers_receipt.

Review finding: marker discharged with a DOWN sibling gateway absent from the startup probe.
2026-09-15 04:06:08 -07:00
teknium1 f39b034db9 test(update): fleet-matrix live-pid tests verify pytest as the home's gateway
collect_fleet_versions now classifies a live pid from the state file only when the
home's identity resolver verifies it (#110420); these two tests write the record about
the pytest process itself, so they verify it explicitly. Dead-pid rows are unchanged.
2026-09-15 04:06:08 -07:00