6 Commits

Author SHA1 Message Date
Teknium 0e78694c72 simplify(compat): hermes_cli.main — drop 198 re-exports/aliases (155 eager + 45 lazy PEP 562 + _warn_stale_dashboard_processes alias + _time/_self/_LAZY_* machinery), repoint 17 source callers + 91 test files 2026-09-03 15:07:39 -07:00
Teknium 19631b5292 fix(update): tighten gateway-side unit gates to exact/hyphenated shape
Mirror the strict unit-name shape from the hermes-serve gate (review on
PR #83595) on the gateway side too: the discovery gate and the SIGUSR1
eligibility helper now accept only `hermes-gateway.service` or the
`hermes-gateway-<profile>` family, so a near-prefix unit like
`hermes-gatewayd.service` can neither enter the restart path nor be sent
a SIGUSR1 it does not handle.
2026-08-16 10:51:50 -07:00
chelsealong 0b42cae068 fix(update): tighten hermes-serve unit gate, dedupe fleet/cleanup restarts
Review on #83595 flagged two service-lifecycle gaps in the hermes-serve
restart support:

- The unit-name gate accepted anything starting with "hermes-serve",
  which also matched the unrelated hermes-server.service. Require the
  exact base unit or the hyphenated profile family instead.
- The fleet-restart loop and _finish_dashboard_update_cleanup() could
  both restart the same hermes-serve unit — the loop restarts it
  directly, then cleanup's PID scan finds the fresh process and
  restarts its owning unit again. Thread the fleet loop's restarted
  unit names through to _kill_stale_dashboard_processes() so it skips
  units already handled.
2026-08-16 10:51:50 -07:00
chelsealong f97cc250da fix(update): restart hermes-serve systemd units alongside gateways
hermes update discovered and restarted hermes-gateway* systemd units but
never looked for hermes-serve* — the Desktop app's backend — so it kept
running stale pre-update code until the user restarted it by hand (#83438).

Extend the systemd unit discovery/restart loop to also match hermes-serve*
units. They don't wire SIGUSR1 to a graceful drain (only gateway/run.py
does), so restart eligibility for the graceful path is now gated on unit
name via a small, directly-tested helper; hermes-serve units fall straight
to the existing blunt systemctl restart path, matching the workaround the
issue already documents.
2026-08-16 10:51:50 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
HexLab98 b7b0c37ef7 test(update): cover fleet restart timeout isolation (#68523) 2026-07-22 10:35:50 +05:30