Commit Graph

22616 Commits

Author SHA1 Message Date
Teknium f8d75db026 test(gateway): accept keyword args in _connect_adapter_with_timeout mocks 2026-08-14 21:57:41 -07:00
Teknium c22815fcaa fix(gateway): cap Telegram cold-start connect so startup reaches running fast (#85993)
CI / Docs Site (push) Has been cancelled
CI / Detect affected areas (push) Has been cancelled
CI / Python tests (push) Has been cancelled
CI / OS-specific tests (push) Has been cancelled
CI / Python lints (push) Has been cancelled
CI / JS & TS checks (push) Has been cancelled
CI / Installer tests (push) Has been cancelled
CI / Desktop E2E (push) Has been cancelled
CI / Deny unrelated histories (push) Has been cancelled
CI / Check contributors (push) Has been cancelled
CI / Check uv.lock (push) Has been cancelled
CI / Check no committed infographics (push) Has been cancelled
CI / package-lock.json diff (push) Has been cancelled
CI / Lint Docker scripts (push) Has been cancelled
CI / Supply-chain scan (push) Has been cancelled
CI / Review label gate (push) Has been cancelled
CI / OSV scan (push) Has been cancelled
CI / All required checks pass (push) Has been cancelled
CI / CI timing report (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
Docker Build, Test, and Publish / Detect affected areas (push) Has been cancelled
Docker Build, Test, and Publish / build (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / build (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / publish (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / publish (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / merge (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
The initial (pre-running) connect awaited during gateway startup now uses
a capped 45s budget for Telegram instead of the full 180s (#67498) budget.
On timeout the platform is queued for the reconnect watcher, which retries
with the full budget and is_reconnect=True (preserving the offline update
queue, #46621). Combined with the parallel startup connects, an unreachable
Telegram no longer holds the whole gateway out of the running state.
2026-08-14 21:57:41 -07:00
EvanProgramming 42a4e86239 fix(test): genuinely verify parallel startup connects (#83791)
The previous concurrency assertion (slow_start < fast_end) was true under
BOTH the serial and parallel implementations, so it proved nothing -- it
even passed against the old serial code on main. The only assertion that
distinguishes the two is that the fast platform finishes before the slow
one (fast_end before slow_end), which is only possible when the connects
overlap.

Switch the test to record connect start/end events in arrival order
(clock-resolution independent) and assert fast_end precedes slow_end. This
also fixes the Windows failure @zuowen7 reported: time.monotonic() has only
~15 ms resolution there, so two parallel connects could land on the same
tick and defeat any wall-clock comparison -- event ordering cannot.

Verified the new test fails against origin/main (serial) and passes against
this branch (parallel).
2026-08-14 21:57:41 -07:00
EvanProgramming d86c67dc7e fix(gateway): connect messaging platforms in parallel at startup (#83791)
GatewayRunner.start() previously awaited each platform's connect() (with its
own timeout) in a serial for-loop. A single slow/failing platform (e.g.
Telegram behind a dead proxy) delayed every later platform's connect by a full
timeout window, cascading one platform's failure onto WeChat/QQ/etc.

Now the slow connect() calls run concurrently via asyncio.gather while the
serial pre-filter (checks, adapter creation, handler wiring) and the
single-threaded result aggregation (shared-state mutation, error handling)
are unchanged. A failing platform no longer blocks the others.

Adds regression tests proving connect() calls overlap and that one failing
platform leaves the others connected.
2026-08-14 21:57:41 -07:00
Teknium 366fd70f98 fix(gateway): registered_names() honors profile scope like is_registered()
The salvaged registered_names() from PR #71582 predates the scoped
platform registry: it read only the process-global _entries/_deferred
maps, but plugin platforms register their deferred loaders under a
profile scope. Result: `hermes tools enable a2a --platform a2a` still
rejected the platform. Union the current-scope maps with the global
ones, mirroring is_registered()'s semantics, under the registry lock.
2026-08-14 21:56:33 -07:00
Teknium 25857671f7 chore: map contributor email for attribution audit 2026-08-14 21:56:33 -07:00
tachyon-r 7a6b8917f7 fix(tools): recognize discovered plugin platforms 2026-08-14 21:56:33 -07:00
Chen Jin 7224301856 fix(toolsets): admit explicitly-configured plugin toolset keys in _get_platform_tools (#81163)
Layer 2 of the #81163 / #78050 fix: _get_platform_tools computed
plugin_ts_keys = _get_plugin_toolset_keys() but only used
CONFIGURABLE_TOOLSETS in the explicit-config filter, so a user-listed
plugin key like `a2a` in `platform_toolsets.cli: [hermes-cli, a2a]` was
silently dropped. The filter now unions configurable and plugin toolset
keys when evaluating has_explicit_config and when admitting per-key
entries.

Cherry-picked from PR #81190 (Layer 2 hunks only; Layer 1 is covered by
the provides_tools mechanism from PR #78842).
2026-08-14 21:56:33 -07:00
Eman e42db348c9 fix(plugins): register deferred platform client tools at discovery (#78050)
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.

Re-anchored accordingly:

- Discovery-time pre-registration, module reuse, and the `provides_tools`
  opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
  since those tools registered before `registration_start` and the slice
  cannot see them.
- A failed materialization no longer carries attribution across. The
  failure path now sweeps the whole ownership ledger for the plugin key,
  not just the `registration_start:` slice, so the pre-registered tools
  are disposed along with the adapter. Attribution and the registry now
  agree at zero instead of reporting tools the process is not serving.

tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 21:56:33 -07:00
kshitij 7a390f7c43 fix: __del__ delegates to close() for full cleanup
The original __del__ only closed _conn (the writer connection),
skipping the read-only connection pool, token writer thread, and
atexit unregister. Delegates to self.close() instead so all
cleanup paths run. Uses __dict__.get('_conn') guard to stay
safe on partially-constructed instances and during interpreter
shutdown.
2026-08-15 10:26:24 +05:30
RelaxJonh 1db4801063 fix: close leaked SessionDB connections on exception paths (#83226)
Two call sites create SessionDB instances without closing them on error:

1. gateway/slash_commands.py: /insights command — db.close() was on the
   success path but not in a finally block, so exceptions between
   SessionDB() and db.close() leak the connection.

2. hermes_cli/sessions_cmd.py: sessions repair — SessionDB() created
   inline with no .close() at all, leaking the FD on every call.

Additionally, add a __del__ safety net to SessionDB itself so that
instances orphaned by callers who forget .close() are cleaned up when
garbage collected, rather than pinning FDs alive until process exit
via the atexit hook.

Fixes #83226
2026-08-15 10:26:24 +05:30
kshitij ce658e82ff fix(session-search): narrow lineage escape to reset/compression; trust SQL child classifier
Follow-up on the salvaged #85764 commits, addressing review findings:

- _session_left_live_context now allowlists end_reason == 'compression'
  or a fresh reset (_FRESH_RESET_END_REASONS) instead of accepting any
  non-None end_reason. The wide predicate let 'branched' parents — whose
  transcript /branch verbatim-copies into the child — surface as
  same-lineage recall hits, returning content already in the caller's
  live context (verified empirically vs main).
- _FRESH_RESET_END_REASONS is now derived from the canonical
  hermes_state_common._RESET_END_REASONS (plus CLI 'new_session') instead
  of a third hand-maintained copy, per that tuple's anti-drift comment.
  Import verified cycle-free.
- Browse drops the Python re-check of parent_session_id rows:
  list_sessions_rich (include_children=False) already applies the
  canonical _LISTABLE_CHILD_SQL classifier, and the Python re-check
  re-hid legacy pre-marker reset children the SQL same-key heuristic
  deliberately admits. _has_reset_from_marker (now orphaned) removed.
- Tests: branched-parent exclusion regression guard (mutation-checked:
  fails on the overbroad predicate) + legacy pre-marker reset child
  browse guard. 48/48 pass.
2026-08-15 10:25:19 +05:30
Teknium 2b5a3fb8ac test(cron): expect tick to contain create_execution failure per-job 2026-08-14 21:55:14 -07:00
Teknium 0ea79484d2 chore: map contributor emails for cron salvage 2026-08-14 21:55:14 -07:00
webtecnica 2d81236f7f fix(cli): report background-dispatch cron runs without false 'failed' (#83340) 2026-08-14 21:55:14 -07:00
Teknium 5733ec097b test(cron): adapt guard-leak regression to owner-fenced dispatch on main
The salvaged regression from #86582 predates the claim_job_for_fire
owner-fencing that landed with #70638; mock the claim and heartbeat so
the healthy job actually runs through the fenced flow.
2026-08-14 21:55:14 -07:00
yuric 569b7a34b1 fix(cron): bound post-run cleanup 2026-08-14 21:55:14 -07:00
Teknium a5bb1bcde3 fix(cron): make transient-DNS classification platform-safe
The errno literal set {8, 7, 11, 51, 60, 61, 65} mixed macOS getaddrinfo
constants with errno values: on Linux socket.EAI_NONAME is -2 and
EAI_AGAIN is -3, so genuine DNS failures were missed while unrelated
OSErrors carrying errno 8/11 (ENOEXEC/EAGAIN semantics differ) could be
misclassified as transient. Classify socket.gaierror against the EAI_*
constants and plain OSError against named errno constants instead.

Addresses the platform-portability review on #83977.
2026-08-14 21:55:14 -07:00
Andrew Fiebert c2a1179a34 fix(cron): fall back on transient DNS during provider resolve
Agent crons resolve OAuth credentials before the agent loop. A short
macOS/WARP DNS blip raised httpx.ConnectError ([Errno 8] nodename nor
servname provided) from xai-oauth token refresh, and the scheduler only
walked fallback_providers on AuthError — so Daily Focus Kickoff died
even when XAI_API_KEY / Anthropic were healthy.

Treat ConnectError/DNS OSError (and cause-chain equivalents) like
AuthError when selecting the fallback chain. Keep provider+model atomic.
Regression test covers the ConnectError path.
2026-08-14 21:55:14 -07:00
fangliquanflq db696c798a test(tests): cover cron fallback delivery idempotency and interrupt skip 2026-08-14 21:55:14 -07:00
fangliquanflq 4668750fad fix(cron): deliver alerts for escaped run failures 2026-08-14 21:55:14 -07:00
zhao ec1aa8977c [verified] fix: clear running-job lock when execution creation fails 2026-08-14 21:55:14 -07:00
Jack Lau f57209bc9f fix(agent): carry the ambiguity of Anthropic's 'out of extra usage' 400 through classification, cooldown, and terminal surfaces
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:

- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
  status-less paths now attach error_context {billing_unverified,
  possible_content_filter}. Reason stays FailoverReason.billing (rotation +
  fallback remain the right recovery either way); ClassifiedError grows a
  billing_unverified property.

- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
  unverified billing exhaustion gets the short transient cooldown instead of
  the one-hour bench, regardless of pool size: a content-filter rejection
  leaves the credential healthy and fails identically on every key, and the
  hour-long sole-credential latch is what replayed the stored error and made
  real fixes look ineffective. A true 402 keeps the full bench. The marker
  persists with the entry so a restart cannot upgrade it back to a bench.

- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
  threads billing_unverified and hands the pool 'billing_unverified' as the
  persisted failure_reason.

- agent/conversation_loop.py: the fallback-switch status, max-retries status,
  terminal label, and both structured terminal results hedge when the verdict
  is unverified. New _billing_terminal_label + _billing_failure_result build
  the returned terminal response in one place; the result dict now carries
  billing_unverified and the billing_block gains 'unverified': true. The
  confirmed-billing path (a real 402 or an API-key credit depletion) keeps
  the original assertive wording, so the caveat no longer dilutes it.

Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.

Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
2026-08-14 21:54:56 -07:00
Jack Lau 6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00
kshitij 24573b396b refactor: bind gateway anti-growth guard to locals, trim overlong comment
Simplify-code pass: gateway/run.py called estimate_messages_tokens_rough
6x on the same data in the anti-growth guard (condition + warning f-string).
Bind to _hyg_in_toks/_hyg_out_toks locals like the conversation_compression.py
guard already does. Also trim the comment from 10 lines to 4 (keep the WHY,
drop the WHAT) and remove an extra blank line before TestCompactedTurnsStaySearchable.
2026-08-15 10:22:47 +05:30
Teknium 3bb83a9a51 chore: map contributor email for salvage 2026-08-15 10:22:47 +05:30
dhruv kejriwal ffaa63f887 fix(compressor): never commit a compression that grows the transcript (in-place path)
The gateway rotation guard (#83339) only protects the rotate path, but
in-place compaction commits inside compress_context() via
archive_and_compact — before the gateway can inspect the result. Add the
anti-growth check at the commit site so both paths are covered: a
compression whose rough output exceeds its input is a strict no-op
(original transcript kept durable, session identity untouched).

Covers the observed failure where session hygiene persisted 426 -> 426
messages and ~379K -> ~688K tokens.
2026-08-15 10:22:47 +05:30
dhruv kejriwal e7418f1621 fix(gateway): never persist a hygiene compression that grows the transcript
Session hygiene could persist a compressed transcript LARGER than the
original (observed: 427K -> 598K), when the generated summary was bigger
than the middle it replaced. Compare like-for-like (both rough estimates)
before persisting a rotated transcript; on growth, keep the original
unchanged so a failed compression is a strict no-op, never a net increase.
2026-08-15 10:22:47 +05:30
Teknium bc36d7d6c8 fix(deps): exempt no-upload-date and exact-pinned packages from exclude-newer bricking
The relative exclude-newer = "14 days" cutoff bricks installs whenever the
resolver cannot see (or accept) a package's upload date:

- defusedxml / python-olm / unpaddedbase64 (#80387, #79434): ancient frozen
  releases (2021-2023) whose upload dates are often absent from mirror
  indexes and stale uv HTTP caches. uv then filters them entirely
  ("there are no versions of defusedxml"), breaking [youtube]/[wecom]/
  [matrix] resolution and daily `uv sync --locked` runs.

- setuptools / pillow / mcp (#78227, #75992, #76020): exact-pinned deps.
  When the pinned version's upload date is invisible, the resolver filters
  the ONLY acceptable candidate — setuptools==83.0.0 in
  [build-system].requires meant the project could not even be built from a
  git checkout on released v0.20.0. Exempting an exact pin costs nothing:
  the version cannot float without a reviewed pin bump.

Changes:
- pyproject.toml: add all six to the existing exclude-newer-package
  whitelist, with rationale comments per class.
- uv.lock: regenerated; diff is the whitelist metadata only (verified
  zero version drift, still 249 packages).
- tests/test_packaging_metadata.py: new standing guard
  test_build_system_requires_exempt_from_exclude_newer — every
  [build-system].requires package must be whitelisted while a relative
  exclude-newer cutoff is configured. Verified both directions (fails
  when setuptools is removed from the whitelist).
- scripts/install.sh: fix the stale tier-name comparison ("all (with
  RL/matrix extras)" vs actual "all") that mislabeled every successful
  Tier-1 install as a fallback-tier install (#79434 bonus finding).

Verification: uv lock --check green on uv 0.11.19 and 0.12.5;
uv sync --extra all --locked green; uv pip install -e '.[all]' resolves;
whitelist mechanism A/B-proven on a minimal project (unsatisfiable ->
resolves; build-requires variant: uv build fails -> succeeds).

Reported-by: MichaelClawHub (#80387), liujianqiu (#79434), maxonliu (#78227)
2026-08-14 21:50:48 -07:00
Alvis cf8b505531 fix(desktop-update): make posix hand-off survive Electron quit teardown (macOS)
The Desktop-spawned hand-off consistently died during Electron's quit
teardown on macOS: the orchestrator process group was terminated right
after `running: hermes update ...`, so no exit code, result file, bundle
swap, or relaunch ever happened, and the loopback shim window surfaced
the death as ERR_CONNECTION_REFUSED or "Aw, Snap!" error code 15
(reproductions in #66753).

- Re-exec the orchestrator through a one-shot setsid child and let the
  direct Electron child exit immediately; the real orchestrator is owned
  by launchd (PPID 1), outside Electron's teardown, same marker/result
  protocol.
- Hold TERM ignored across the `hermes update` invocation and
  log-and-ignore the single teardown TERM that can still arrive after
  the desktop PID dies (durable SIGNAL breadcrumb for diagnosis).
- Delay start_ui until the desktop PID is gone plus 1s so the shim
  server/window are never born inside the teardown window.
- Run both UI processes in their own sessions; keep SIGTERM/SIGHUP
  ignored in the shim server and stop it with SIGKILL, so a stray TERM
  can no longer leave the progress window on a dead loopback URL while
  the update continues.

Verified on a production git install (macOS arm64, Darwin 27.0,
v0.20.1): six consecutive Desktop-triggered/production-shape updates
completed end-to-end including a full desktop rebuild + codesign; the
shim survived a deliberately injected TERM+HUP mid-update and a full
`hermes desktop --force-build --build-only` running alongside it.

Fixes the macOS reproductions in #66753.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 21:48:22 -07:00
Christopher 9b3823d08b fix(install): provision Node 26 so managed npm satisfies engines
Node 22 ships npm 11.16.0, which engines.npm rejects (11.10–11.16
ignore min-release-age-exclude). Fresh Hermes-managed installs then
fail npm ci with EBADENGINE. Node 26 ships 11.17.0.
2026-08-14 21:48:14 -07:00
briandevans 73b35b8b2b test(backup): assert an imported member cannot keep a setuid target's bits
Pre-creates a 0o6755 target, imports a member over it, and asserts the
published file is 0o755 with both elevated bits gone — plus that the
staged temp file never carried them either, so there is no window where
archive content sits behind an elevated mode.

The existing coverage in this class cannot see the failure: every mode
assertion masks with ``& 0o777``, which discards exactly the bits at
issue, and the fixtures chmod their targets to ordinary modes that never
had them set. Without the mask on the preserved mode this test reports
the published file still holding S_ISUID.

Skipped where the platform or filesystem refuses setuid on a user-owned
file, so the assertion never depends on running as root.
2026-08-14 21:47:23 -07:00
briandevans 66356fe24b fix(backup): drop setuid/setgid from the mode restored onto imported files
``_extract_member_atomically`` carries the replaced file's permissions
across the publish so that routing through mkstemp does not change what
the caller would otherwise have produced. But ``_preserve_file_mode``
returns ``stat.S_IMODE``, which is all twelve bits, and this restore is
deliberate on both sides of the replace: the mode is fchmod'd onto the
temp before ``atomic_replace`` and re-applied afterwards because chown
clears the elevated bits. So a target sitting at 0o4755 comes out of
``hermes import`` still at 0o4755 — with contents supplied by the zip.

That is a regression introduced by the atomic rewrite rather than a
pre-existing one. The overwrite it replaced was an in-place
``open(target, "wb")``, and an in-place write by a process without
CAP_FSETID has the elevated bits stripped by the kernel, so the old path
left 0o4755 as 0o755.

The blast radius is not limited to Hermes' own state: the ``_external/``
branch of ``run_import`` publishes members anywhere under ``$HOME``, and
this is the path that documents ``sudo`` use so ownership survives a
restore. An archive that happens to contain a member matching some
existing privileged file would take over the identity that file runs as.

Mask the two bits off the preserved mode. The masking happens once,
before the temp file is chmod'd, so there is no transient elevation
either. The sticky bit is kept — it is inert on a regular file. The
ordinary permission bits are unaffected, so the Docker/NAS installs the
preservation exists for still get their broader modes back.

This is the one write path in the repo where the bytes are untrusted;
the ``utils`` writers that preserve the full mode re-serialize content
the process itself produced, and are correct as they stand.
2026-08-14 21:47:23 -07:00
briandevans 60f86662e3 docs(backup): note that atomic_replace's cross-device fallback still truncates
The atomicity claim in _extract_member_atomically's docstring holds on the
os.replace path but not on atomic_replace's EXDEV/EBUSY fallback, which uses
shutil.copyfile and so opens the destination 'wb'. That is pre-existing
behaviour shared by every atomic writer in the repo, and it is reachable here
for a symlinked target whose real file lives on another filesystem. Scope the
docstring to what the helper actually guarantees instead of overstating it;
the fallback itself is a utils.atomic_replace change.
2026-08-14 21:47:23 -07:00
briandevans 1c3c1f4d71 fix(backup): preserve owner on atomic import writes and close the 0600 transit window
Follow-up on the atomic-import restore, delegating both metadata concerns to
the shared helpers instead of half-handling them locally.

Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace`
publishes a temp file owned by the *writing* user, so `sudo hermes import`
re-owned every restored file to root — on the disaster-recovery path, and on
exactly the Docker/NAS volume installs `utils._restore_file_owner` was added
for. `_extract_member_atomically` now captures `_preserve_file_owner(target)`
before staging and calls `_restore_file_owner` after the replace, before the
mode restore (chown clears setuid/setgid, so the mode has to go back last).

Mode handling was also only half applied before the replace: the `os.fchmod`
branch applied it to the temp fd, but the platforms without `fchmod` fell
through to a best-effort post-replace chmod, leaving the published file at
mkstemp's 0600 until that chmod landed — permanently if the process died in
between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback
copy 0600 onto the target. The mode is now applied to the temp file on both
branches, with the post-replace `_restore_file_mode` kept as the belt-and-
braces path.

This is the same shape `atomic_write_text` and `atomic_yaml_write` already
carry after 3556728a5 and 43fc86562; capture and restore now reuse
`utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` /
`_restore_file_owner` rather than re-deriving them, which also drops the local
`import stat`.

Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites):
- test_restore_preserves_existing_file_owner — forces a uid/gid so it does not
  need root; asserts chown fires once, with the captured owner, on the
  pre-existing file only (a newly created member has no prior owner).
  Mutation-checked: dropping only the `_restore_file_owner` call reds it.
- test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr`
  on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600
  without the fix, 0o644 with it. Mutation-checked the same way.
2026-08-14 21:47:23 -07:00
briandevans e88c9f0ef2 fix(backup): restore import members atomically so a failed import can't erase config
`hermes import` wrote every zip member with `open(target, "wb")` followed by
`dst.write(src.read())`, at both restore sites in `run_import`. Opening for
write truncates the user's existing file to zero *before* any replacement
bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves
`config.yaml`, `.env`, or an external provider config (e.g.
`~/.honcho/config.json`) empty with nothing behind it — during the
disaster-recovery path the user is running precisely because they already
lost something. The `_external/` branch writes outside HERMES_HOME, into
third-party configs under the user's home, so the blast radius is not
confined to Hermes state.

Both sites now stage the member into the target's own directory, fsync it,
and publish with `utils.atomic_replace`, so the target only ever moves from
its old contents to the complete new contents.

`atomic_replace` rather than a bare `os.replace`: it resolves a symlinked
target first, so deployments that link `config.yaml` into a dotfiles repo
keep the link instead of having it silently swapped for a regular file
(#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for
cross-device and bind-mount installs. Members stream through
`shutil.copyfileobj` instead of being read whole into memory. The temp file
is removed on any failure so a partial import leaves no residue, and
permission bits are carried across the replace so mkstemp's 0600 does not
silently tighten restored files.

This extends the module's own established idiom — `backup.py` already
publishes atomically via `os.replace` in `_atomic_output_path` and in the
snapshot writer — into the one path that still overwrote user files in place.
2026-08-14 21:47:23 -07:00
Teknium 5ecad87d1e fix(cron): report ownerless interrupted fires so their notices still send
mark_running_jobs_interrupted skipped legacy fires without a registered
durable owner entirely — correct for the persisted last_status write
(no owner fence to protect a replacement run), but the gateway shutdown
path also uses the returned ID list to deliver interrupted-cron notices
while adapters are still connected (#82232). Keep the persistence skip,
but include the job in the returned list so the user is still told.
2026-08-14 21:47:16 -07:00
joaomarcos baf7034850 fix(gateway): deliver interrupted-cron notices before adapters disconnect
When the shutdown drain times out and kills an in-flight cron job, the
job's owner is never told. The cron worker does try: `_is_interrupted()`
forces the failure path with an honest "interrupted by gateway shutdown"
error, and failed jobs always deliver. But that worker is a thread, it
reaches `_deliver_result()` asynchronously, and by then
`_bounded_adapter_teardown()` has closed the transport. The reporter of

Worse, the loss is silent twice over: `_consume_interrupted_flag()`
returns True — the gateway already wrote `last_status` — so
`mark_job_run()` is skipped, and the `delivery_error` from the failed
send is discarded with it. The run's only trace is a generic line in
jobs.json.

The gateway already owns the right window. `_notify_active_sessions_of_
shutdown()` runs while adapters are up, precisely so shutdown messages
can be sent — but it iterates `_running_agents`, and cron work lives on
the scheduler's own thread pool. Same structural blindness already fixed
for counting (#60432) and draining (#63529), never fixed for notifying.

So notify from the post-interrupt phase, which is the last point where
the transport is still up: `_kill_tool_subprocesses()` now returns the
job IDs it marked, and `_notify_interrupted_cron_jobs()` sends each one's
owner a notice on the job's own resolved delivery targets. Adapter
teardown order is untouched — it is load-bearing for #53175 and #8202.

Jobs with `deliver: local`, and `deliver: origin` jobs with no resolvable
origin (#43014), resolve to zero targets and stay silent. Per-platform
`gateway_restart_notification: false` is honoured, matching the chat
path. Every failure is swallowed so a wedged adapter cannot extend
shutdown.

Second, when the interrupted flag short-circuits `mark_job_run()`, the
delivery failure is now persisted on its own via `update_job()`, so a
notice that still cannot be sent is at least recorded. `update_job()`
rather than a second `mark_job_run()`: the latter also advances
`next_run_at` and the repeat counter, and running that twice for one run
would skip a fire or auto-delete the job early.

Fixes #82232. Related: #82161, #82224.
2026-08-14 21:47:16 -07:00
joaomarcos 4b06d9e9b1 fix(gateway): keep the cron drain floor compatible with shutdown test doubles
CI slice 5/12 caught two ways the new cron budget broke `_stop_impl_body`
for callers that are not real GatewayRunner instances:

- `_FakeGateway` in test_shutdown_cache_cleanup.py borrows `_stop_impl`
  without subclassing, so it never picked up the class-level
  `_cron_drain_timeout` default and raised AttributeError. Read it through
  the getattr-guard convention the same function already uses for its
  liveness-guard machinery.
- The same double overrides `_drain_active_agents(self, timeout)`, so
  passing the cron budget raised "takes 2 positional arguments but 3 were
  given". The double now mirrors the real optional parameter. It is the
  only override in the tree; test_startup_restart_race.py uses AsyncMock,
  which accepts any signature.

Verified against a stashed clean tree: the 22 gateway test files that
still fail locally fail identically with and without this branch (80 = 80,
empty set difference both ways) — they are pre-existing Windows-only
failures (setsid, POSIX modes) unrelated to this change.
2026-08-14 21:47:16 -07:00
joaomarcos 45bb486b26 fix(gateway): give in-flight cron work its own drain floor
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.

A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.

Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.

The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.

Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.

Relates to #82161 (complements #82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
2026-08-14 21:47:16 -07:00
Jeremy 1f6f86119f fix(cli): stop hermes update from respawning orphan serve --port 0 (#78821)
Filter manual dashboard/serve respawn candidates after update: skip
ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines,
and cap one restart per profile/HERMES_HOME so orphan counts no longer
grow across successive updates.
2026-08-14 21:46:32 -07:00
joaomarcos 69d1843512 fix(packaging): ship bundled plugin manifests 2026-08-14 21:45:08 -07:00
Teknium 169ff2db4f chore: add contributor email mapping for arccat-114 2026-08-14 21:44:28 -07:00
Teknium 0dba3316b2 fix(gateway): generalize supervised-gateway exemption in orphan reaper to all platforms
Compose the service-PID exclusion (#85743, RelaxJonh) and the recorded-PID +
parent-chain exemption (#86100, arccat-114) into one cross-platform rule:

- _get_service_pids() exclusion now runs unconditionally, not only under
  is_macos() — it is the authoritative "supervised" signal for launchd and
  any systemd unit visible on a host that got past the systemd gate.
- The recorded-healthy-gateway (get_running_pid()) + parent-chain exemption
  now runs on every platform, not only Windows. A recorded, liveness-verified
  gateway is by definition not an orphan "the pidfile/runtime record can't
  see", so the reaper must never target it — this covers Windows Scheduled
  Task / Startup VBS supervision, standalone launcher-started gateways
  (the case #85743 alone would miss), and macOS/WSL equivalents.

True orphans (no service registration, no valid runtime record) are still
found and reaped, preserving the #51325/#75936 duplicate-port protection.

Existing macOS regression tests updated to pin get_running_pid to None for
their scenario; Windows regression tests from #86100 carry over unchanged.

Bug class: #83683 (root), #86287, #86098, #85738, #85368, #85344, #85044,
#84855, #84824, #84200.
2026-08-14 21:44:28 -07:00
arccat-114 102369c5f6 fix(gateway): spare Scheduled-Task-supervised gateway from orphan reaper on Windows
The orphan reaper kills a healthy gateway (and its Scheduled-Task bootstrap
parent chain) every time the Desktop backend starts on Windows, because
_get_service_pids() only implements systemd/launchd and returns an empty
set on Windows — a supervised gateway is therefore indistinguishable from
an unsupervised orphan.

Exempt the recorded healthy gateway PID and its parent chain from the
orphan scan on Windows, mirroring the macOS launchd exemption (#85913).
The Scheduled-Task bootstrap's argv matches the gateway scan, so without
exempting the parent chain killing the bootstrap takes the detached
gateway down with it.

Fixes #86098
2026-08-14 21:44:28 -07:00
RelaxJonh ac9b058ef4 fix(gateway): exclude service-managed PIDs from orphan reaping
_reap_unsupervised_gateway_orphans() kills every gateway PID found by
find_gateway_pids() on hosts without systemd (macOS launchd, Windows
Scheduled Task). This includes service-managed gateways that are NOT
orphans — they are supervised by launchd/systemd and should never be
killed during a stale-process sweep.

Add own |= _get_service_pids() to the exclusion set before scanning,
so launchd/systemd-supervised gateways are preserved. True orphans
(reparented leftovers not present in launchctl/systemctl) are still
found and reaped, preserving the #77276 protection.

Fixes #85344 (macOS launchd gateway killed by desktop serve startup)
Fixes #85044 (Windows Scheduled Task gateway killed by desktop serve)
Fixes #84855 (Permission denied to kill orphaned gateway PID)
Fixes #85368 (gateway process repeatedly killed, messaging offline)
2026-08-14 21:44:28 -07:00
worlldz 547043a4d8 fix(telegram): honor fallback disable during connect 2026-08-14 21:43:06 -07:00
Teknium dac3c44afc test: fix salvage test imports; drop WAL worker-thread test superseded by read pool
- tests/cron/test_sessiondb_init_hang.py: add threading/time imports the
  salvaged late-close regression tests rely on.
- tests/test_hermes_state.py: drop
  test_close_closes_wal_read_connection_created_on_worker_thread — main
  replaced per-thread WAL reader ownership with the pooled read-connection
  design (permits + checkout/return), so cross-thread reader draining no
  longer exists in the form the test asserted.
2026-08-14 21:41:26 -07:00
Tranquil-Flow 38709ae6f2 fix(cron): close leaked SessionDB connection when init outlives the timeout-abandoned worker (#72782)
run_job() submits SessionDB() to a one-worker executor and abandons the
worker (shutdown(wait=False)) when init exceeds the cron timeout. If the
constructor later completes inside that abandoned worker, the Future's
result — an open SessionDB holding .db/WAL/SHM handles — was orphaned and
never closed, leaking descriptors until EMFILE. Attach a done-callback on
the timeout path that retrieves and closes any eventual late result.

Salvage note: the lazy-recall ownership half of #72822 (_owns_session_db
tracked on AIAgent, owned handle closed in close()) already landed on main;
this carries the remaining cron timeout-abandon half with its regression
test.
2026-08-14 21:41:26 -07:00
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00