`connect()` caches every path it has initialized in the process-local
`_INITIALIZED_PATHS` set and then skips all first-open work for it —
header validation, integrity probe, `SCHEMA_SQL`, additive migrations.
That cache is keyed on a path, but the schema it stands for lives in a
file, and the two can drift apart: delete or replace `kanban.db` under a
live gateway/dispatcher/dashboard process and the next `connect()` takes
the fast path, lets SQLite create a fresh empty database, and hands back
a connection with no tables in it.
Nothing notices. Every query then fails with `no such table: tasks`,
`plugin_api._conn()` logs its init warning and carries on, and the board
renders empty. Because the cache entry survives, the process re-creates
the same schema-less ~4 KB file on every restart of the desktop app in
front of it — only killing the backing process clears it.
Verify the sentinel table on the fast path and self-heal when it is gone:
drop the stale cache entry and fall through to the existing init path,
which re-runs the probes and the schema script under the cross-process
init lock. The check is one `sqlite_master` lookup on the already-resident
page 1, so the steady-state path stays lock-free (#36644) and does no
schema work.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The unconditional 2s join before InterruptedError delayed interrupt
detection when Relay managed execution was not active (CI:
tests/run_agent/test_interrupt_propagation.py — detection took 2.34s
against a <1.0s budget, because the mocked worker sleeps 5s and there
is no Relay scope to unwind).
Extract the join into _join_worker_for_relay_teardown(), which no-ops
unless a Relay runtime exists AND managed execution consumers are
registered — the only case where an orphaned physical scope can corrupt
the LIFO stack (#81521). Applied at all three interrupt sites
(streaming, non-streaming, Bedrock streaming). The regression test now
simulates a live runtime so the join path stays covered.
Follow-up to HexLab98's salvaged commits:
- Apply the same bounded worker join before raising InterruptedError at
the two sibling interrupt sites that share the raise-without-join
shape: the non-streaming API poll loop and the Bedrock streaming poll
loop. Both workers run Relay-managed physical attempts, so raising
immediately allowed turn teardown to race a still-open physical scope
exactly as in the streaming path.
- Address the #81601 review finding (egilewski): the pinned nemo-relay
binding's get_scope_stack() returns a native ScopeStack object which
scope.pop rejects with TypeError, so the orphan drain never drained
under the real binding. current_top() now prefers the version-correct
scope.get_handle() accessor and falls back to the old list-unwrap for
fake/legacy shapes. Handle comparisons go through same_handle(),
comparing by uuid, because native ScopeHandle instances do not
implement value equality.
- Add a real-binding regression test that reproduces the orphaned-scope
session close against the pinned native wheel (skips where the native
binding is unavailable), alongside the existing fake-based coverage.
_to_openai_base_url() matched ZAI (open.bigmodel.cn, api.z.ai, bare
"bigmodel") and Kimi (api.kimi.com) via `substring in url`, so any custom
gateway whose base_url happened to contain one of those strings as a path
segment (e.g. a reverse-proxy prefix like /proxy/bigmodel-fallback/) was
silently misrouted to the wrong OpenAI-wire endpoint shape.
This is the same false-positive class 6f33f510e8 just fixed for the
MiniMax branch in the same function by switching to base_url_host_matches()
(hostname-anchored). Apply the same fix to the ZAI and Kimi branches, which
that commit didn't touch. Drops the bare "bigmodel" substring check since
open.bigmodel.cn is the only canonical bigmodel-family host referenced
anywhere else in the codebase (agent/model_metadata.py, hermes_cli/auth.py).
Added regression tests mirroring the MiniMax marker-in-path tests added in
the same commit.
The gateway auto-restart phase used to swallow every exception at debug
level, so tests driving cmd_update end-to-end never noticed it touching
real gateway discovery. With #78574 surfacing an aborted restart as a
failed update, an unmocked find_gateway_pids on a box with a live
gateway hits the conftest live-system guard and turns into a spurious
sys.exit(1).
Add an autouse fixture in test_cmd_update.py (discovery returns nothing,
systemd unsupported) and the same seams in test_update_head_moved_gate's
helper so the phase is a clean no-op for tests that do not assert on
gateway restarts.
Reviewer egilewski found the original defer was circular (#83590 comment):
the self-lock preflight wrote .update-incomplete and exited, but the next
launch only ran the full recovery AFTER main.py's third-party imports —
so a healthy venv's probes made the early pass a no-op, main.py imported
cryptography eagerly, the .pyd got mapped again, and the deferred install
re-hit the exact self-lock it was meant to escape.
Close the loop by making the marker guarantee the install runs BEFORE any
native extension can be imported:
- hermes_cli/_install_repair.py (new, stdlib-only): single source of truth
for the core .[all] reinstall — ensurepip bootstrap, uv-pip/pip
resolution with VIRTUAL_ENV, Termux env stripping, Windows hermes*.exe
quarantine, per-extra fallback ladder, and fd1→fd2 routing for acp
safety. Deliberately free of managed_uv/hermes_constants imports so it
stays importable in the corrupted-venv state it exists to repair.
- hermes_cli/_early_recovery.py: recover_if_needed now completes a pending
.update-incomplete install BEFORE the import probes, on every launch
that sees the marker (unless argv is update). Success clears the
marker; failure bumps an attempts counter inside the marker body and
keeps it. A 3-attempt ceiling stops a persistently-failing install
from reinstall-hammering every launch (hermes acp included) — past the
ceiling the late post-import recovery takes over with its manual
recovery instructions. Single-flight lock shared with the late path.
- hermes_cli/main.py: _recover_core_update_marker_locked delegates the
install to the shared executor (no duplicated logic); ensure_uv stays
in the late path so a venv whose uv vanished mid-update still
bootstraps it.
- tests: 7 new regressions — the reviewer's exact case (marker + healthy
venv → install runs while sys.modules has no cryptography), failure
keeps marker + increments attempts, retry ceiling, lazy marker does
not trigger core install (#58004 invariant), argv-update skip, and
corrupt/missing marker bodies. The key test was sabotage-verified:
removing the pre-import branch makes it fail with zero install calls,
while a lone-lazy-marker test still passes; restoring the branch makes
it pass again.
Refs #83569
Two gaps left every Windows git-checkout install unable to recover from
the exact failure state #83569 reports:
1. Self-lock detection. _detect_venv_python_processes() always excludes
the calling process by design — a CLI hermes update IS the venv python.
An updater that had already imported a native venv extension (the
canonical one being cryptography.hazmat.bindings._rust, mapped while
hermes_cli.main resolved external secret sources) passed every
preflight and then died mid-sync with os error 5 when uv tried to
rewrite the mapped .pyd, stranding the venv half-updated. A new
preflight now refuses the sync before touching the checkout, writes
the update-incomplete marker so the next fresh launch completes the
install, and exits 2. Verified on a live Windows 11 host: after
importing hermes_cli.main, tasklist /m _rust.pyd shows the .pyd mapped
in the caller, and a peer process cannot open it read-write
(Permission denied) — while a rename succeeds, matching how uv/pip
actually fail (truncate+write, not rename).
2. Early-recovery install path. _early_recovery._run_repair_install used
sys.executable -m pip unconditionally. Windows git checkouts install
on a uv-managed base interpreter (python-build-standalone), whose
EXTERNALLY-MANAGED marker makes plain pip abort with
externally-managed-environment — the repair no-oped and the venv
stayed broken. The repair now detects the PEP 668 marker, prefers
uv pip install with VIRTUAL_ENV pointed at the project venv, and
falls back to pip --break-system-packages when no uv binary exists.
Both fixes ship with subprocess/unit regressions (sabotage-verified):
the new tests fail on pre-fix code and pass with it. Complements #77517,
which keeps the updater from importing cryptography in the first place;
this PR is the defence-in-depth when any future path loads it anyway.
Fixes#83569
_cold_start_windows_gateway_after_update() printed the success line off a
successful Popen return alone, which only proves CreateProcess succeeded,
not that the child survived. On Windows, a job object denying
CREATE_BREAKAWAY_FROM_JOB hard-kills the child during updater teardown
before it logs anything, yet the updater still printed "Starting Windows
gateway after update (PID ...)" — leaving Telegram/Discord/etc. offline
with no indication anything failed (#84185).
Route the success report through gateway_windows._report_gateway_start(),
the same post-spawn liveness poll every other _spawn_detached() caller
already uses, so a dead child is reported as a failure with a
manual-recovery hint instead of a false success.
Review follow-up (#78574): the aborted-restart handler only flagged the fleet
stale when the post-failure survivor probe was None or non-empty. A positive
empty probe was treated as proof-of-safety — but `[]` is only safe when
nothing was running before the phase. If a gateway was discovered, stopped
(SIGTERM/drain), and its replacement never came back, the probe is empty at
exactly that unsafe moment and the update reported success — the fail-open
contract this fix exists to close.
Snapshot the pre-restart gateway PIDs before any stop/drain and route the
handler decision through a pure _restart_phase_failure_is_incomplete() helper
that fails closed on an empty survivor set whenever a gateway existed
pre-restart (or the pre-state could not be read). Add decision-level regression
tests covering the stopped-without-replacement gap, unknown pre-state, and the
truly-no-gateway positive control.
The gateway auto-restart phase in `hermes update` was wrapped in a blanket
`except Exception` that only logged at debug level. When the phase raised
early — e.g. importing `hermes_cli.gateway` from the freshly pulled checkout
inside a process that already loaded pre-update modules — every drain and
restart line vanished from the update output, the update printed
"Update complete!" and exited 0, and the still-running gateway kept serving
pre-update modules against replaced source files. The next Telegram turn died
with `ImportError: cannot import name 'is_trivial_prompt'`.
The handler now probes for surviving gateway processes and, unless it can
positively prove none are running, prints the cause plus a manual recovery
command and marks the fleet restart incomplete — which exits nonzero and
writes the gateway-mode exit-code marker, matching the existing
failed-or-stale-unit path.
Fixes#78574
A detached/pinned checkout can report 'N new commit(s)' against origin,
run the ff-only merge successfully, and still sit on the old commit
afterward (the branch-switch step re-detaches to the raw SHA). Before
this guard 'hermes update' printed '✓ Code updated!' and reinstalled
deps + rebuilt the desktop app against the stale tree - no error, no
warning, 'hermes doctor' healthy.
Compare pre-pull and post-pull HEAD; if they match, fail loudly with a
reattach hint instead of claiming success.
The shared load-boundary flatten added for dict-valued model.default only
ran on the user/default merge; _load_config_impl then deep-merged the raw
managed overlay without normalizing, so a managed model.default:
{provider, model} still reached status/fallback/runtime readers as a dict.
Normalize the managed overlay (same _normalize_root_model_keys pass, plus
the bare model-string -> model.default promotion used by
managed_scope.apply_managed_overlay) before expanding and merging, so
every overlay is canonical before load_config returns.
Adds load_config() regressions for a nested managed default and a bare
managed model string.
Extends the fix to the config-load chokepoint so every reader sees plain
strings, not just the interactive CLI paths. _normalize_root_model_keys
now flattens a dict-valued model.default/model.model into a string default
plus the nested provider (promoted to model.provider when no explicit
outer provider or "auto" is set), covering the residual readers the
review flagged: doctor, status/dump, fallback picker, prompt-size, and
the context-switch guard — all of which called .strip()/flowed the raw
value and would crash or misroute on a nested dict.
Adds _normalize_root_model_keys regression coverage for the flatten
(precedence, auto-override, explicit-provider-win, alias shape, flat
strings untouched).
The prior dict coercion converted a dict-valued model.default to a plain
model string but dropped the nested provider. On the interactive CLI path
requested_provider then fell back to the outer merged model.provider
(typically "auto", authoritative at runtime resolution), so the model
could be routed through the wrong active provider.
Canonicalize both halves at the shared boundary: _split_model_config_default
flattens a dict-valued default into (model, provider) and HermesCLI.__init__
feeds the nested provider into the requested_provider chain (still below an
explicit --provider argument). new_session reuses the same helper.
Adds regression coverage asserting the nested provider stays paired with the
model, that flat string defaults keep the outer provider behavior, and that
an explicit provider argument still wins.
Review finding on PR #83194 (egilewski): Install-Venv committed the venv
transaction as soon as the replacement had a working interpreter, deleting
the parked previous venv. Install-Dependencies is a separate later stage
(a separate process under the stage-per-process bootstrap) and every
dependency tier or the baseline-import gate can still fail after that
point - a failed update could still leave Hermes and the blocker probe
unusable with no rollback source.
Now:
- Install-Venv records the parked backup in venv.pending-backup instead
of deleting it, and excludes it from the venv.stale.* sweep.
- Install-Dependencies wraps the dependency tiers + baseline-import gate
in the transaction: Restore-VenvBackup on failure (parks the failed
replacement as venv.failed.*, renames the previous venv back), and
Complete-VenvTransaction only after the imports prove the replacement
usable.
- Source-contract regression tests for the boundary
(tests/test_install_ps1_venv_transaction_boundary.py).
The initial (pre-running) connect awaited during gateway startup now uses
a capped 45s budget for Telegram instead of the full 180s (#67498) budget.
On timeout the platform is queued for the reconnect watcher, which retries
with the full budget and is_reconnect=True (preserving the offline update
queue, #46621). Combined with the parallel startup connects, an unreachable
Telegram no longer holds the whole gateway out of the running state.
The previous concurrency assertion (slow_start < fast_end) was true under
BOTH the serial and parallel implementations, so it proved nothing -- it
even passed against the old serial code on main. The only assertion that
distinguishes the two is that the fast platform finishes before the slow
one (fast_end before slow_end), which is only possible when the connects
overlap.
Switch the test to record connect start/end events in arrival order
(clock-resolution independent) and assert fast_end precedes slow_end. This
also fixes the Windows failure @zuowen7 reported: time.monotonic() has only
~15 ms resolution there, so two parallel connects could land on the same
tick and defeat any wall-clock comparison -- event ordering cannot.
Verified the new test fails against origin/main (serial) and passes against
this branch (parallel).
GatewayRunner.start() previously awaited each platform's connect() (with its
own timeout) in a serial for-loop. A single slow/failing platform (e.g.
Telegram behind a dead proxy) delayed every later platform's connect by a full
timeout window, cascading one platform's failure onto WeChat/QQ/etc.
Now the slow connect() calls run concurrently via asyncio.gather while the
serial pre-filter (checks, adapter creation, handler wiring) and the
single-threaded result aggregation (shared-state mutation, error handling)
are unchanged. A failing platform no longer blocks the others.
Adds regression tests proving connect() calls overlap and that one failing
platform leaves the others connected.
Layer 2 of the #81163 / #78050 fix: _get_platform_tools computed
plugin_ts_keys = _get_plugin_toolset_keys() but only used
CONFIGURABLE_TOOLSETS in the explicit-config filter, so a user-listed
plugin key like `a2a` in `platform_toolsets.cli: [hermes-cli, a2a]` was
silently dropped. The filter now unions configurable and plugin toolset
keys when evaluating has_explicit_config and when admitting per-key
entries.
Cherry-picked from PR #81190 (Layer 2 hunks only; Layer 1 is covered by
the provides_tools mechanism from PR #78842).
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.
Re-anchored accordingly:
- Discovery-time pre-registration, module reuse, and the `provides_tools`
opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
since those tools registered before `registration_start` and the slice
cannot see them.
- A failed materialization no longer carries attribution across. The
failure path now sweeps the whole ownership ledger for the plugin key,
not just the `registration_start:` slice, so the pre-registered tools
are disposed along with the adapter. Attribution and the registry now
agree at zero instead of reporting tools the process is not serving.
tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up on the salvaged #85764 commits, addressing review findings:
- _session_left_live_context now allowlists end_reason == 'compression'
or a fresh reset (_FRESH_RESET_END_REASONS) instead of accepting any
non-None end_reason. The wide predicate let 'branched' parents — whose
transcript /branch verbatim-copies into the child — surface as
same-lineage recall hits, returning content already in the caller's
live context (verified empirically vs main).
- _FRESH_RESET_END_REASONS is now derived from the canonical
hermes_state_common._RESET_END_REASONS (plus CLI 'new_session') instead
of a third hand-maintained copy, per that tuple's anti-drift comment.
Import verified cycle-free.
- Browse drops the Python re-check of parent_session_id rows:
list_sessions_rich (include_children=False) already applies the
canonical _LISTABLE_CHILD_SQL classifier, and the Python re-check
re-hid legacy pre-marker reset children the SQL same-key heuristic
deliberately admits. _has_reset_from_marker (now orphaned) removed.
- Tests: branched-parent exclusion regression guard (mutation-checked:
fails on the overbroad predicate) + legacy pre-marker reset child
browse guard. 48/48 pass.
The salvaged regression from #86582 predates the claim_job_for_fire
owner-fencing that landed with #70638; mock the claim and heartbeat so
the healthy job actually runs through the fenced flow.
Agent crons resolve OAuth credentials before the agent loop. A short
macOS/WARP DNS blip raised httpx.ConnectError ([Errno 8] nodename nor
servname provided) from xai-oauth token refresh, and the scheduler only
walked fallback_providers on AuthError — so Daily Focus Kickoff died
even when XAI_API_KEY / Anthropic were healthy.
Treat ConnectError/DNS OSError (and cause-chain equivalents) like
AuthError when selecting the fallback chain. Keep provider+model atomic.
Regression test covers the ConnectError path.
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:
- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
status-less paths now attach error_context {billing_unverified,
possible_content_filter}. Reason stays FailoverReason.billing (rotation +
fallback remain the right recovery either way); ClassifiedError grows a
billing_unverified property.
- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
unverified billing exhaustion gets the short transient cooldown instead of
the one-hour bench, regardless of pool size: a content-filter rejection
leaves the credential healthy and fails identically on every key, and the
hour-long sole-credential latch is what replayed the stored error and made
real fixes look ineffective. A true 402 keeps the full bench. The marker
persists with the entry so a restart cannot upgrade it back to a bench.
- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
threads billing_unverified and hands the pool 'billing_unverified' as the
persisted failure_reason.
- agent/conversation_loop.py: the fallback-switch status, max-retries status,
terminal label, and both structured terminal results hedge when the verdict
is unverified. New _billing_terminal_label + _billing_failure_result build
the returned terminal response in one place; the result dict now carries
billing_unverified and the billing_block gains 'unverified': true. The
confirmed-billing path (a real 402 or an API-key credit depletion) keeps
the original assertive wording, so the caveat no longer dilutes it.
Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.
Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.
Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).
Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:
- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
reporter verified returns 200. Meaning, the skill_manage reference, and the
## Skill Safety Rule block are all preserved. The reword is empirically
validated rather than understood, so a comment records the bisect and warns
that any rewrite must be re-verified against an OAuth token, not an API key.
- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
longer asserts exhaustion as fact. It hedges the opening line, names the
content-filter alternative, and gives the operator a way to tell the two apart
(if the usage page still shows quota, suspect a content rejection). It also
points at `hermes auth reset anthropic`, because the credential exhaustion
latch replays the stored error for ~60 min without issuing a request — which
makes a real fix look like it did not work.
- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
not an API key, despite auth_type="api_key". It stays in api_key_env_vars
because that tuple doubles as the credential-discovery list; removing it would
stop Hermes finding a `claude setup-token` credential at all.
Docs updated to match the reworded prompt.
Fixes#82154
Simplify-code pass: gateway/run.py called estimate_messages_tokens_rough
6x on the same data in the anti-growth guard (condition + warning f-string).
Bind to _hyg_in_toks/_hyg_out_toks locals like the conversation_compression.py
guard already does. Also trim the comment from 10 lines to 4 (keep the WHY,
drop the WHAT) and remove an extra blank line before TestCompactedTurnsStaySearchable.
The gateway rotation guard (#83339) only protects the rotate path, but
in-place compaction commits inside compress_context() via
archive_and_compact — before the gateway can inspect the result. Add the
anti-growth check at the commit site so both paths are covered: a
compression whose rough output exceeds its input is a strict no-op
(original transcript kept durable, session identity untouched).
Covers the observed failure where session hygiene persisted 426 -> 426
messages and ~379K -> ~688K tokens.
The relative exclude-newer = "14 days" cutoff bricks installs whenever the
resolver cannot see (or accept) a package's upload date:
- defusedxml / python-olm / unpaddedbase64 (#80387, #79434): ancient frozen
releases (2021-2023) whose upload dates are often absent from mirror
indexes and stale uv HTTP caches. uv then filters them entirely
("there are no versions of defusedxml"), breaking [youtube]/[wecom]/
[matrix] resolution and daily `uv sync --locked` runs.
- setuptools / pillow / mcp (#78227, #75992, #76020): exact-pinned deps.
When the pinned version's upload date is invisible, the resolver filters
the ONLY acceptable candidate — setuptools==83.0.0 in
[build-system].requires meant the project could not even be built from a
git checkout on released v0.20.0. Exempting an exact pin costs nothing:
the version cannot float without a reviewed pin bump.
Changes:
- pyproject.toml: add all six to the existing exclude-newer-package
whitelist, with rationale comments per class.
- uv.lock: regenerated; diff is the whitelist metadata only (verified
zero version drift, still 249 packages).
- tests/test_packaging_metadata.py: new standing guard
test_build_system_requires_exempt_from_exclude_newer — every
[build-system].requires package must be whitelisted while a relative
exclude-newer cutoff is configured. Verified both directions (fails
when setuptools is removed from the whitelist).
- scripts/install.sh: fix the stale tier-name comparison ("all (with
RL/matrix extras)" vs actual "all") that mislabeled every successful
Tier-1 install as a fallback-tier install (#79434 bonus finding).
Verification: uv lock --check green on uv 0.11.19 and 0.12.5;
uv sync --extra all --locked green; uv pip install -e '.[all]' resolves;
whitelist mechanism A/B-proven on a minimal project (unsatisfiable ->
resolves; build-requires variant: uv build fails -> succeeds).
Reported-by: MichaelClawHub (#80387), liujianqiu (#79434), maxonliu (#78227)
Pre-creates a 0o6755 target, imports a member over it, and asserts the
published file is 0o755 with both elevated bits gone — plus that the
staged temp file never carried them either, so there is no window where
archive content sits behind an elevated mode.
The existing coverage in this class cannot see the failure: every mode
assertion masks with ``& 0o777``, which discards exactly the bits at
issue, and the fixtures chmod their targets to ordinary modes that never
had them set. Without the mask on the preserved mode this test reports
the published file still holding S_ISUID.
Skipped where the platform or filesystem refuses setuid on a user-owned
file, so the assertion never depends on running as root.
Follow-up on the atomic-import restore, delegating both metadata concerns to
the shared helpers instead of half-handling them locally.
Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace`
publishes a temp file owned by the *writing* user, so `sudo hermes import`
re-owned every restored file to root — on the disaster-recovery path, and on
exactly the Docker/NAS volume installs `utils._restore_file_owner` was added
for. `_extract_member_atomically` now captures `_preserve_file_owner(target)`
before staging and calls `_restore_file_owner` after the replace, before the
mode restore (chown clears setuid/setgid, so the mode has to go back last).
Mode handling was also only half applied before the replace: the `os.fchmod`
branch applied it to the temp fd, but the platforms without `fchmod` fell
through to a best-effort post-replace chmod, leaving the published file at
mkstemp's 0600 until that chmod landed — permanently if the process died in
between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback
copy 0600 onto the target. The mode is now applied to the temp file on both
branches, with the post-replace `_restore_file_mode` kept as the belt-and-
braces path.
This is the same shape `atomic_write_text` and `atomic_yaml_write` already
carry after 3556728a5 and 43fc86562; capture and restore now reuse
`utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` /
`_restore_file_owner` rather than re-deriving them, which also drops the local
`import stat`.
Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites):
- test_restore_preserves_existing_file_owner — forces a uid/gid so it does not
need root; asserts chown fires once, with the captured owner, on the
pre-existing file only (a newly created member has no prior owner).
Mutation-checked: dropping only the `_restore_file_owner` call reds it.
- test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr`
on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600
without the fix, 0o644 with it. Mutation-checked the same way.
`hermes import` wrote every zip member with `open(target, "wb")` followed by
`dst.write(src.read())`, at both restore sites in `run_import`. Opening for
write truncates the user's existing file to zero *before* any replacement
bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves
`config.yaml`, `.env`, or an external provider config (e.g.
`~/.honcho/config.json`) empty with nothing behind it — during the
disaster-recovery path the user is running precisely because they already
lost something. The `_external/` branch writes outside HERMES_HOME, into
third-party configs under the user's home, so the blast radius is not
confined to Hermes state.
Both sites now stage the member into the target's own directory, fsync it,
and publish with `utils.atomic_replace`, so the target only ever moves from
its old contents to the complete new contents.
`atomic_replace` rather than a bare `os.replace`: it resolves a symlinked
target first, so deployments that link `config.yaml` into a dotfiles repo
keep the link instead of having it silently swapped for a regular file
(#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for
cross-device and bind-mount installs. Members stream through
`shutil.copyfileobj` instead of being read whole into memory. The temp file
is removed on any failure so a partial import leaves no residue, and
permission bits are carried across the replace so mkstemp's 0600 does not
silently tighten restored files.
This extends the module's own established idiom — `backup.py` already
publishes atomically via `os.replace` in `_atomic_output_path` and in the
snapshot writer — into the one path that still overwrote user files in place.
mark_running_jobs_interrupted skipped legacy fires without a registered
durable owner entirely — correct for the persisted last_status write
(no owner fence to protect a replacement run), but the gateway shutdown
path also uses the returned ID list to deliver interrupted-cron notices
while adapters are still connected (#82232). Keep the persistence skip,
but include the job in the returned list so the user is still told.
When the shutdown drain times out and kills an in-flight cron job, the
job's owner is never told. The cron worker does try: `_is_interrupted()`
forces the failure path with an honest "interrupted by gateway shutdown"
error, and failed jobs always deliver. But that worker is a thread, it
reaches `_deliver_result()` asynchronously, and by then
`_bounded_adapter_teardown()` has closed the transport. The reporter of
Worse, the loss is silent twice over: `_consume_interrupted_flag()`
returns True — the gateway already wrote `last_status` — so
`mark_job_run()` is skipped, and the `delivery_error` from the failed
send is discarded with it. The run's only trace is a generic line in
jobs.json.
The gateway already owns the right window. `_notify_active_sessions_of_
shutdown()` runs while adapters are up, precisely so shutdown messages
can be sent — but it iterates `_running_agents`, and cron work lives on
the scheduler's own thread pool. Same structural blindness already fixed
for counting (#60432) and draining (#63529), never fixed for notifying.
So notify from the post-interrupt phase, which is the last point where
the transport is still up: `_kill_tool_subprocesses()` now returns the
job IDs it marked, and `_notify_interrupted_cron_jobs()` sends each one's
owner a notice on the job's own resolved delivery targets. Adapter
teardown order is untouched — it is load-bearing for #53175 and #8202.
Jobs with `deliver: local`, and `deliver: origin` jobs with no resolvable
origin (#43014), resolve to zero targets and stay silent. Per-platform
`gateway_restart_notification: false` is honoured, matching the chat
path. Every failure is swallowed so a wedged adapter cannot extend
shutdown.
Second, when the interrupted flag short-circuits `mark_job_run()`, the
delivery failure is now persisted on its own via `update_job()`, so a
notice that still cannot be sent is at least recorded. `update_job()`
rather than a second `mark_job_run()`: the latter also advances
`next_run_at` and the repeat counter, and running that twice for one run
would skip a fire or auto-delete the job early.
Fixes#82232. Related: #82161, #82224.
CI slice 5/12 caught two ways the new cron budget broke `_stop_impl_body`
for callers that are not real GatewayRunner instances:
- `_FakeGateway` in test_shutdown_cache_cleanup.py borrows `_stop_impl`
without subclassing, so it never picked up the class-level
`_cron_drain_timeout` default and raised AttributeError. Read it through
the getattr-guard convention the same function already uses for its
liveness-guard machinery.
- The same double overrides `_drain_active_agents(self, timeout)`, so
passing the cron budget raised "takes 2 positional arguments but 3 were
given". The double now mirrors the real optional parameter. It is the
only override in the tree; test_startup_restart_race.py uses AsyncMock,
which accepts any signature.
Verified against a stashed clean tree: the 22 gateway test files that
still fail locally fail identically with and without this branch (80 = 80,
empty set difference both ways) — they are pre-existing Windows-only
failures (setsid, POSIX modes) unrelated to this change.