Commit Graph

5936 Commits

Author SHA1 Message Date
HexLab98 18b442cdeb fix(install): abort Windows venv recreate when rename-aside fails
When Rename-Item on the live venv is denied, do not fall back to an
in-place Remove-Item that can gut site-packages and leave no rollback.
Also mark venv-blocker probe failures with probe_failed so they cannot
be read as a clear scan (#83149).
2026-08-14 21:58:09 -07:00
tachyon-r 7a6b8917f7 fix(tools): recognize discovered plugin platforms 2026-08-14 21:56:33 -07:00
Chen Jin 7224301856 fix(toolsets): admit explicitly-configured plugin toolset keys in _get_platform_tools (#81163)
Layer 2 of the #81163 / #78050 fix: _get_platform_tools computed
plugin_ts_keys = _get_plugin_toolset_keys() but only used
CONFIGURABLE_TOOLSETS in the explicit-config filter, so a user-listed
plugin key like `a2a` in `platform_toolsets.cli: [hermes-cli, a2a]` was
silently dropped. The filter now unions configurable and plugin toolset
keys when evaluating has_explicit_config and when admitting per-key
entries.

Cherry-picked from PR #81190 (Layer 2 hunks only; Layer 1 is covered by
the provides_tools mechanism from PR #78842).
2026-08-14 21:56:33 -07:00
Eman e42db348c9 fix(plugins): register deferred platform client tools at discovery (#78050)
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.

Re-anchored accordingly:

- Discovery-time pre-registration, module reuse, and the `provides_tools`
  opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
  since those tools registered before `registration_start` and the slice
  cannot see them.
- A failed materialization no longer carries attribution across. The
  failure path now sweeps the whole ownership ledger for the plugin key,
  not just the `registration_start:` slice, so the pre-registered tools
  are disposed along with the adapter. Attribution and the registry now
  agree at zero instead of reporting tools the process is not serving.

tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 21:56:33 -07:00
webtecnica 2d81236f7f fix(cli): report background-dispatch cron runs without false 'failed' (#83340) 2026-08-14 21:55:14 -07:00
Jack Lau 6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00
briandevans 66356fe24b fix(backup): drop setuid/setgid from the mode restored onto imported files
``_extract_member_atomically`` carries the replaced file's permissions
across the publish so that routing through mkstemp does not change what
the caller would otherwise have produced. But ``_preserve_file_mode``
returns ``stat.S_IMODE``, which is all twelve bits, and this restore is
deliberate on both sides of the replace: the mode is fchmod'd onto the
temp before ``atomic_replace`` and re-applied afterwards because chown
clears the elevated bits. So a target sitting at 0o4755 comes out of
``hermes import`` still at 0o4755 — with contents supplied by the zip.

That is a regression introduced by the atomic rewrite rather than a
pre-existing one. The overwrite it replaced was an in-place
``open(target, "wb")``, and an in-place write by a process without
CAP_FSETID has the elevated bits stripped by the kernel, so the old path
left 0o4755 as 0o755.

The blast radius is not limited to Hermes' own state: the ``_external/``
branch of ``run_import`` publishes members anywhere under ``$HOME``, and
this is the path that documents ``sudo`` use so ownership survives a
restore. An archive that happens to contain a member matching some
existing privileged file would take over the identity that file runs as.

Mask the two bits off the preserved mode. The masking happens once,
before the temp file is chmod'd, so there is no transient elevation
either. The sticky bit is kept — it is inert on a regular file. The
ordinary permission bits are unaffected, so the Docker/NAS installs the
preservation exists for still get their broader modes back.

This is the one write path in the repo where the bytes are untrusted;
the ``utils`` writers that preserve the full mode re-serialize content
the process itself produced, and are correct as they stand.
2026-08-14 21:47:23 -07:00
briandevans 60f86662e3 docs(backup): note that atomic_replace's cross-device fallback still truncates
The atomicity claim in _extract_member_atomically's docstring holds on the
os.replace path but not on atomic_replace's EXDEV/EBUSY fallback, which uses
shutil.copyfile and so opens the destination 'wb'. That is pre-existing
behaviour shared by every atomic writer in the repo, and it is reachable here
for a symlinked target whose real file lives on another filesystem. Scope the
docstring to what the helper actually guarantees instead of overstating it;
the fallback itself is a utils.atomic_replace change.
2026-08-14 21:47:23 -07:00
briandevans 1c3c1f4d71 fix(backup): preserve owner on atomic import writes and close the 0600 transit window
Follow-up on the atomic-import restore, delegating both metadata concerns to
the shared helpers instead of half-handling them locally.

Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace`
publishes a temp file owned by the *writing* user, so `sudo hermes import`
re-owned every restored file to root — on the disaster-recovery path, and on
exactly the Docker/NAS volume installs `utils._restore_file_owner` was added
for. `_extract_member_atomically` now captures `_preserve_file_owner(target)`
before staging and calls `_restore_file_owner` after the replace, before the
mode restore (chown clears setuid/setgid, so the mode has to go back last).

Mode handling was also only half applied before the replace: the `os.fchmod`
branch applied it to the temp fd, but the platforms without `fchmod` fell
through to a best-effort post-replace chmod, leaving the published file at
mkstemp's 0600 until that chmod landed — permanently if the process died in
between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback
copy 0600 onto the target. The mode is now applied to the temp file on both
branches, with the post-replace `_restore_file_mode` kept as the belt-and-
braces path.

This is the same shape `atomic_write_text` and `atomic_yaml_write` already
carry after 3556728a5 and 43fc86562; capture and restore now reuse
`utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` /
`_restore_file_owner` rather than re-deriving them, which also drops the local
`import stat`.

Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites):
- test_restore_preserves_existing_file_owner — forces a uid/gid so it does not
  need root; asserts chown fires once, with the captured owner, on the
  pre-existing file only (a newly created member has no prior owner).
  Mutation-checked: dropping only the `_restore_file_owner` call reds it.
- test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr`
  on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600
  without the fix, 0o644 with it. Mutation-checked the same way.
2026-08-14 21:47:23 -07:00
briandevans e88c9f0ef2 fix(backup): restore import members atomically so a failed import can't erase config
`hermes import` wrote every zip member with `open(target, "wb")` followed by
`dst.write(src.read())`, at both restore sites in `run_import`. Opening for
write truncates the user's existing file to zero *before* any replacement
bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves
`config.yaml`, `.env`, or an external provider config (e.g.
`~/.honcho/config.json`) empty with nothing behind it — during the
disaster-recovery path the user is running precisely because they already
lost something. The `_external/` branch writes outside HERMES_HOME, into
third-party configs under the user's home, so the blast radius is not
confined to Hermes state.

Both sites now stage the member into the target's own directory, fsync it,
and publish with `utils.atomic_replace`, so the target only ever moves from
its old contents to the complete new contents.

`atomic_replace` rather than a bare `os.replace`: it resolves a symlinked
target first, so deployments that link `config.yaml` into a dotfiles repo
keep the link instead of having it silently swapped for a regular file
(#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for
cross-device and bind-mount installs. Members stream through
`shutil.copyfileobj` instead of being read whole into memory. The temp file
is removed on any failure so a partial import leaves no residue, and
permission bits are carried across the replace so mkstemp's 0600 does not
silently tighten restored files.

This extends the module's own established idiom — `backup.py` already
publishes atomically via `os.replace` in `_atomic_output_path` and in the
snapshot writer — into the one path that still overwrote user files in place.
2026-08-14 21:47:23 -07:00
joaomarcos 45bb486b26 fix(gateway): give in-flight cron work its own drain floor
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.

A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.

Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.

The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.

Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.

Relates to #82161 (complements #82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
2026-08-14 21:47:16 -07:00
Jeremy 1f6f86119f fix(cli): stop hermes update from respawning orphan serve --port 0 (#78821)
Filter manual dashboard/serve respawn candidates after update: skip
ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines,
and cap one restart per profile/HERMES_HOME so orphan counts no longer
grow across successive updates.
2026-08-14 21:46:32 -07:00
Teknium 0dba3316b2 fix(gateway): generalize supervised-gateway exemption in orphan reaper to all platforms
Compose the service-PID exclusion (#85743, RelaxJonh) and the recorded-PID +
parent-chain exemption (#86100, arccat-114) into one cross-platform rule:

- _get_service_pids() exclusion now runs unconditionally, not only under
  is_macos() — it is the authoritative "supervised" signal for launchd and
  any systemd unit visible on a host that got past the systemd gate.
- The recorded-healthy-gateway (get_running_pid()) + parent-chain exemption
  now runs on every platform, not only Windows. A recorded, liveness-verified
  gateway is by definition not an orphan "the pidfile/runtime record can't
  see", so the reaper must never target it — this covers Windows Scheduled
  Task / Startup VBS supervision, standalone launcher-started gateways
  (the case #85743 alone would miss), and macOS/WSL equivalents.

True orphans (no service registration, no valid runtime record) are still
found and reaped, preserving the #51325/#75936 duplicate-port protection.

Existing macOS regression tests updated to pin get_running_pid to None for
their scenario; Windows regression tests from #86100 carry over unchanged.

Bug class: #83683 (root), #86287, #86098, #85738, #85368, #85344, #85044,
#84855, #84824, #84200.
2026-08-14 21:44:28 -07:00
arccat-114 102369c5f6 fix(gateway): spare Scheduled-Task-supervised gateway from orphan reaper on Windows
The orphan reaper kills a healthy gateway (and its Scheduled-Task bootstrap
parent chain) every time the Desktop backend starts on Windows, because
_get_service_pids() only implements systemd/launchd and returns an empty
set on Windows — a supervised gateway is therefore indistinguishable from
an unsupervised orphan.

Exempt the recorded healthy gateway PID and its parent chain from the
orphan scan on Windows, mirroring the macOS launchd exemption (#85913).
The Scheduled-Task bootstrap's argv matches the gateway scan, so without
exempting the parent chain killing the bootstrap takes the detached
gateway down with it.

Fixes #86098
2026-08-14 21:44:28 -07:00
RelaxJonh ac9b058ef4 fix(gateway): exclude service-managed PIDs from orphan reaping
_reap_unsupervised_gateway_orphans() kills every gateway PID found by
find_gateway_pids() on hosts without systemd (macOS launchd, Windows
Scheduled Task). This includes service-managed gateways that are NOT
orphans — they are supervised by launchd/systemd and should never be
killed during a stale-process sweep.

Add own |= _get_service_pids() to the exclusion set before scanning,
so launchd/systemd-supervised gateways are preserved. True orphans
(reparented leftovers not present in launchctl/systemctl) are still
found and reaped, preserving the #77276 protection.

Fixes #85344 (macOS launchd gateway killed by desktop serve startup)
Fixes #85044 (Windows Scheduled Task gateway killed by desktop serve)
Fixes #84855 (Permission denied to kill orphaned gateway PID)
Fixes #85368 (gateway process repeatedly killed, messaging offline)
2026-08-14 21:44:28 -07:00
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
RelaxJonh 0d91ab8889 fix: close leaked SessionDB connections on /insights and sessions-repair exception paths (#83226)
Two call sites create SessionDB instances without closing them on error:

1. gateway/slash_commands.py: /insights command - db.close() was on the
   success path but not in a finally block, so exceptions between
   SessionDB() and db.close() leak the connection.

2. hermes_cli/sessions_cmd.py: sessions repair - SessionDB() created
   inline with no .close() at all, leaking the FD on every call.

Salvage note: the original PR (#83237) also added a __del__ safety-net
finalizer to SessionDB; review showed the atexit hook registered by
queue_token_counts() strongly retains the instance, so the finalizer never
fires for the leak class it claimed to cover. Dropped here in favor of the
deterministic constructor-finally ownership repair salvaged from #83620.
2026-08-14 21:41:26 -07:00
Teknium f45813ea77 fix(sessions): run state.db schema migration eagerly at backend startup and stop swallowing locked ALTERs
After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).

Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):

1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
   writable open, typically the user's first NEW session. The dashboard
   backend now schedules one writable open of its own state.db from the
   lifespan (daemon thread, never blocks the ready-probe socket, never
   raises), so the store is brought current before the first session-
   list poll on every `hermes serve` / `hermes dashboard` / Desktop
   headless entrypoint.

2. _reconcile_columns caught sqlite3.OperationalError around every
   ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
   orphaned sibling backends made the ALTER fail silently — startup
   "succeeded" with a half-reconciled schema, and the open-time lock
   patience (#74478) never saw the error because it was swallowed
   inside first. Now: "duplicate column" races stay at DEBUG,
   locked/busy re-raises so _connect_and_init_with_lock_patience
   retries the whole idempotent init with jittered backoff, and any
   other failure (e.g. un-ADDable NOT NULL) logs at WARNING.

Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.

Fixes #79531
Fixes #80037

Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
2026-08-14 21:36:29 -07:00
David Metcalfe e4d0e4c3d8 fix(win): never rewrite the in-use managed Node tree (#80926)
The Hermes-managed Node tree at %HERMES_HOME%\node is destructively
rewritten while the desktop app's Node processes execute from it:
the Node-26 heal did shutil.rmtree + move, the EBADENGINE repair ran
npm install --global --prefix into the tree, and install.ps1's
Test-Node did Remove-Item + Move-Item. Windows rejects those writes
with PermissionError: [WinError 5] on npm.cmd.

- _heal_managed_node_windows: stage the fully-downloaded tree in a
  sibling node.new-* dir, then rename-swap (live tree -> node.old-*,
  staged -> node). The live tree is never deleted before its
  replacement is ready, so an interrupted heal cannot gut it; a
  refused rename is the OS-level in-use signal and defers (returns
  None) instead of forcing the write.
- heal_hermes_managed_node: an in-use deferral does not record the
  once-per-process attempt, so the heal retries once the tree is free.
- managed_node_tree_in_use: cheap psutil pre-check (Windows only) that
  avoids pointless 30-50MB re-downloads in long-lived processes.
- upgrade_managed_npm: defer the in-place npm self-upgrade while the
  tree is in use, with a notice.
- install.ps1: Test-ManagedNodeInUse guard around Update-ManagedNpm and
  the Test-Node install branch, which now rename-swaps instead of
  delete-then-move.

An in-use-but-outdated tree keeps serving the old runnable Node (old
Node beats no Node), and every npm resolution re-evaluates the heal, so
the upgrade applies automatically on the next update with the app
closed.
2026-08-14 21:35:30 -07:00
zuowen7 f71f91a39b fix(desktop): surface compaction-archived messages in transcript reads (#80680) 2026-08-14 21:08:14 -07:00
spfcraze 5f619cfa0e fix(gateway): don't restart supervised services on clean exit
The s6 finish script for profile-gateway services restarted on ANY exit
except EX_CONFIG (78) — including clean exit 0. Restart-on-normal-exit
turns an intentional stop into a reconnect loop: the ashriel-discord
storm in #76435 made 1,000+ connections and got the bot token reset by
Discord.

The finish script now exits 125 (permanent failure, no restart) for both
clean exit 0 and EX_CONFIG; only non-zero, non-78 exits (genuine crashes)
restart normally.

Scope note: #76435 bundles a second, separable symptom (Windows desktop
updater showing the literal 'managed outside dashboard' sentinel). Its
root cause is undiagnosed and #22733 covers the dialog-explanation path;
this PR is the gateway half only.

Tests: behavioral — the rendered finish script is executed via sh for
exit codes 0/78/1/137, asserting no-restart for clean stop and fatal
config, restart for crashes. Pre-existing EX_CONFIG test still passes.
2026-08-14 21:07:21 -07:00
Teknium bc5805c35f fix: compare base-URL hostnames, not substrings, in provider-identity checks
Port of the bug class from earendil-works/pi#7933 (DeepSeek base-URL
detection matched by raw substring, missing case variants and matching
lookalike URLs). Hermes had the same class at five sites:

- cli_agent_setup_mixin.py: keyless-custom-endpoint detection treated any
  URL containing the OpenRouter host substring (path segment, lookalike
  domain) as OpenRouter, and missed case variants of the real host.
- models.py validate_requested_model: same substring check for routing an
  openrouter provider with a custom base_url to the custom catalog.
- runtime_provider.py: local-endpoint autodetect matched the string
  localhost anywhere in the URL, including remote hostnames containing it.
- gateway/run.py: /status endpoint display, same local-host substring.
- agent_runtime_helpers.py: Nous Portal cache-layout detection matched
  the nousresearch substring anywhere in the URL.

All sites now use the existing base_url_host_matches / base_url_hostname
helpers (exact host or subdomain, case-insensitive). Regression tests
proven to fail against the old predicates.
2026-08-14 20:47:13 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00
Aleks Clark 62eefff697 perf(desktop): bound long-running app resource use
Persist backend ownership for reliable cleanup, park inactive panes, and
evict unreferenced transcripts so Desktop stays responsive over long sessions.

💘 Generated with Crush

Assisted-by: Crush:gpt-5.6
2026-08-14 20:23:07 -07:00
Teknium 50d98fc1f3 feat(delegation): raise subagent iteration cap default 50 -> 250 (+migration) (#86506)
delegation.max_iterations is the per-subagent tool-call budget. The old
default of 50 truncated substantial delegated work: leaf agents spend
~15-20 turns on reconnaissance before producing output, then ran out of
budget mid-task and returned 'completed but unfinished' summaries. 250
gives real delegated work room to finish.

Changes:
- config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36
- tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in
  sync with the shipped default to prevent drift)
- config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly
  the OLD default 50 -> 250 on update, so existing installs inherit the new
  headroom. Any other explicit value (deliberate override) is preserved;
  unset inherits 250 at read time.
- cli-config.yaml.example: doc the new default

The cap is per-child and children run concurrently (max_concurrent_children
default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds
(default 0 = off) remains available as a wall-clock guardrail, and users can
still pin a lower max_iterations explicitly.

Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset
untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
2026-08-14 16:50:30 -07:00
Justin Bennington 167dd48b40 fix(models): trust certifi for credentialed catalogs (E-1002) 2026-08-14 16:43:37 -07:00
SHL0MS b21e0bd8c9 fix(providers): honor per-provider TLS on custom /models and pricing probes
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:

- the requests-based metadata/pricing probe
  (agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
  (hermes_cli/models.py::probe_api_models)

A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.

This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:

- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
  ssl_ca_cert before falling back to the env vars. Callers with no base_url
  keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
  passes it through open_credentialed_url, which gains an ssl_context seam on
  the cloned secure opener. Unmatched or public endpoints pass None and keep
  urllib's default policy.

Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
2026-08-14 16:37:36 -07:00
SHL0MS cbfe186da8 fix(config): canonicalize legacy api_mode spellings instead of silently discarding them
Earlier releases accepted api_mode: openai on custom provider entries.
The canonical transport set is now {chat_completions, codex_responses,
anthropic_messages, bedrock_converse, codex_app_server}, and an
unrecognized value was silently ignored at both consumption sites
(_normalize_custom_provider_entry passes the raw string through and
agent_init's accepted-set check drops it; _parse_api_mode returns None),
falling through to hostname-based detection.

For hosts with a detection rule the provider silently switches
transports after an update. Observed live: a custom entry for
api.actual.inc with api_mode: openai (valid when written) flipped to
codex_responses via the hostname rule, and every reasoning-bearing
request to the relay's /v1/responses failed with a wrapped non-JSON
error while /v1/chat/completions worked throughout.

Fix: one shared alias map (_canonical_api_mode) consulted by both
sites. openai/openai_chat -> chat_completions, responses ->
codex_responses, anthropic/messages -> anthropic_messages, bedrock ->
bedrock_converse. Canonical names and unknown values pass through
unchanged, so invalid-config behavior is untouched.

Tests: alias map contract (every alias lands in _VALID_API_MODES),
normalizer canonicalization incl. the transport: key alias, and the
runtime gate accepting legacy spellings while still rejecting unknowns.
2026-08-14 16:36:39 -07:00
Buff Pesos 56f1afc834 feat(dashboard-auth): extend RFC 8252 native sign-in to password providers (system-browser autofill) (#75808)
* feat(dashboard-auth): extend RFC 8252 native sign-in to password providers

The desktop app runs password sign-in for gated gateways in an embedded
Electron BrowserWindow, where OS password managers (macOS Passwords /
iCloud Keychain autofill) cannot reach the form — Chromium-in-Electron
has no bridge to them, so users retype credentials by hand even though
the /login form already carries the right autocomplete attributes.

The existing RFC 8252 native flow (system browser + loopback + PKCE)
solves exactly this for OAuth providers, but was explicitly disabled for
password providers on the grounds that they have "no IDP round trip to
broker". The brokering is still worth having: it moves the credential
form into the system browser, where password-manager autofill just works.

Gateway-only change; the desktop needs no changes (runNativeLogin is
already page-agnostic), and older desktop builds pick the capability up
automatically once the gateway advertises it:

* /auth/native/authorize now accepts a supports_password provider:
  register the pending broker authorization as usual, then 302 the
  system browser to the interactive /login form with the opaque
  broker_state in the gateway's PKCE cookie (the same server-controlled
  channel the OAuth branch uses) instead of an IDP redirect.
* /auth/password-login: when the server-set PKCE cookie carries a
  broker handle, a successful credential check completes the pending
  authorization exactly like the /auth/callback native branch — mint
  the one-time loopback code, return the loopback redirect (validated
  loopback-only at authorize time) as `next`, clear the PKCE cookie,
  and set NO session cookies. A lapsed broker is a clean 400 telling
  the user to restart sign-in; a failed credential attempt leaves the
  pending entry intact so the user can retype.
* /api/status now advertises "native_pkce" whenever any interactive
  session provider is registered (previously only for non-password
  providers), so the desktop selects the system-browser strategy for
  password-only gateways.

Security posture is unchanged from the existing flow: loopback-literal
redirect_uri enforcement, PKCE S256 binding, single-use short-TTL codes,
constant-time comparison, and the same rate limiter on password attempts.

Tests: full authorize → /login → password-login → loopback → token →
bearer round trip, wrong-password keeps the pending entry, lapsed broker
→ 400, no-broker browser login keeps minting cookies, and the /api/status
advertisement for password-only gateways.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dashboard-auth): bind native password completion to the authorize-time provider

Review follow-ups for #75808:

* /auth/password-login now enforces that body.provider matches the
  provider recorded in the server-set PKCE cookie by
  /auth/native/authorize before completing a pending native
  authorization. /login renders a form for every session provider, so
  without this a native flow started for provider A could be completed
  with provider B's credentials, binding B's session into A's pending
  entry. The mismatch is rejected BEFORE credential verification (no
  session minted, no oracle) and preserves both the pending entry and
  the cookie, so the user can still submit the correct provider's form.
  Covered by a two-password-provider E2E regression test.

* Update the two docs spots that still said password-only providers do
  not advertise native_pkce (website desktop-native-signin guide and the
  auth_flows type comment in web/src/lib/api.ts).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: map contributor email for #75808 (buffpesos)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
2026-08-14 21:39:01 +00:00
algf 2912093e06 fix(dashboard): redraw TUI after PTY reattach 2026-08-14 13:51:26 -07:00
Teknium 20e5d51bea fix(browser): managed-first browser-use CLI resolution
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.

- _find_cli(): probe order flipped to managed bin -> PATH ->
  user-level tool dir (then uvx across the same order). A user's own
  uv tool install can no longer shadow the Hermes-managed copy with a
  drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
  install — only the managed copy does, so selecting any backend
  provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
  and always delegates to install_cli(), the single owner of the
  managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
  HERMES_HOME/bin/browser-use counts as installed.

Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.

Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.
2026-08-14 13:51:07 -07:00
andyst-dev 898d786ad7 fix(update): count real behind commits in SSH fast path
Fixes #84851

The SSH fast path in _check_via_local_git compared only the exact tip SHA
of local HEAD against upstream main. When a local carried commit makes the
SHAs differ, it returned 1 ('behind') without checking ancestry, so an
ahead-of-origin checkout was misreported as '1 commit behind' — nudging the
user to run 'hermes update' and wipe their carried work. Fall back to
git rev-list --count HEAD..origin/main (mirroring the full-clone path) when
the SHAs differ.
2026-08-14 13:46:33 -07:00
teknium1 f79440e0f4 feat: /loop — recurring in-session wakeups (Claude Code parity)
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).

Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.

Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.

Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
  process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
  post-turn tick completion, and a supervised loop_wakeup_watcher that
  injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
  notification-poller wakeup driver + post-turn completion in the turn
  dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
  ticks defer until it finishes, pauses, or parks; real user input always
  wins over both

Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
2026-08-14 13:40:19 -07:00
Teknium d8d7cc068d fix(update): stop reporting bogus 'Found 9980 new commit(s)' on shallow installs
The hermes update APPLY path still ran an unconditional
rev-list --count HEAD..origin/<branch> — on a depth-1 installer checkout
that walks the truncated graph and reports the entire remote ancestry
(#53479's 'Found 9980 new commit(s)' on Windows 11). The zero/nonzero gate
stays (a 0 count is trustworthy on any graph); when the count is positive
on a shallow repo, recover the real number via the GitHub compare API
(added in PR #86257) and print count-free wording when that fails.
ahead_by==0 (local-ahead) falls through to the up-to-date path.

Completes the class fix from PR #86257 on its last remaining site.
2026-08-14 13:36:42 -07:00
Teknium 9442a718da fix(update-check): recover the real behind-count via the GitHub compare API
The honesty half (no fabricated counts) leaves shallow installs permanently
count-less. The compare API knows the full graph regardless of local clone
depth: GET /repos/<o>/<r>/compare/<current>...<target> returns ahead_by —
exactly the behind count the shallow boundary lost.

- hermes_cli/banner.py: _github_compare_behind() (bounded, unauthenticated,
  best-effort); wired into _check_via_rev and the shallow branch of
  _check_via_local_git. ahead_by==0 with differing tips = local-ahead => 0.
- hermes_cli/update_cmd.py: hermes update --check shallow path prints the
  exact count when recoverable, presence-only wording otherwise.
- apps/desktop/electron/update-count.ts: compareApiUrl() +
  parseCompareBehindCount() pure helpers; main.ts fetches the count when
  resolveBehindCount() returns null, and the SSH-official passive path stops
  fabricating behind:1 (uses compare API + updateAvailable flag).
- apps/desktop/src/lib/version-status.ts: updateAvailable now applies to the
  client target too, so a shallow desktop install shows '(update)' instead of
  nothing (or the old frozen '(+1)').

Fixes #84591; CLI siblings of #78253 / #53479 behavior.

E2E: live compare API returned 61/62 for real 61/62-commit gaps and 0 for the
reversed (local-ahead) pair; real shallow-clone fixture (depth-1 clone +
depth-1 fetch, merge-base broken) recovers the exact count with the API and
falls back to the honest sentinel offline.
2026-08-14 13:09:44 -07:00
gkd2323c 294071bf08 fix(banner): stop fabricating '1 commit behind' on SSH-official remotes
The SSH-official-remote path in _check_via_local_git was hard-coded to
return 1 when _check_via_rev reported UPDATE_AVAILABLE_NO_COUNT, so
'hermes --version' and the CLI banner surfaced a stable but false
'Update available: 1 commit behind — run hermes update' message. The
count never grew: whether upstream was 1 commit or 100 commits ahead,
the banner always said '1 commit behind'.

Root cause: an ls-remote probe against the upstream URL can only tell
us tip SHAs, not a real commit count. Returning the sentinel
UPDATE_AVAILABLE_NO_COUNT (-1) already means 'update exists, count
unknown' — the exact right shape for this path.

The dashboard/desktop UI does not depend on the fabricated 1:

- The REST /api/hermes/update/check endpoint
  (hermes_cli/web_server.py::check_hermes_update) treats any nonzero
  behind as update_available=true, and its docstring explicitly
  documents -1 as a legitimate value.
- The desktop store (apps/desktop/src/store/updates.ts::mapBackendCheck)
  clamps behind<=0 to 0 and reads updateAvailable as a separate boolean
  field.

So restoring the sentinel is a strict improvement: CLI banner and
hermes --version now say 'Update available' honestly instead of
inventing a count, and every REST/desktop consumer keeps working.

Changes:
- hermes_cli/banner.py: drop the 'return 1' override in the SSH branch;
  propagate the sentinel unchanged.
- hermes_cli/main.py::_print_version_info: render the sentinel as
  'Update available — run <cmd>' (without a count).
- tests/hermes_cli/test_update_check.py: update the SSH-official test
  to assert on the sentinel; add 3 new tests covering the CLI
  renderer's -1 / >0 / 0 branches.
2026-08-14 13:09:44 -07:00
angeon 88a1a9fd95 fix(cli): let redraw recovery rebuild scrollback 2026-08-14 13:09:37 -07:00
Teknium 1169fb50a4 fix: install Browser Use CLI for every browser backend except Camofox
The Browser Use CLI 3.0 is the primary driver engine for all browser
backends except Camofox, but only the explicit 'Browser Use' picker row
ran the install hook. Local Browser, Browserbase, Firecrawl, and the
Nous-managed cloud rows left the CLI uninstalled, so those selections
depended on the uvx zero-install fallback (first-use PyPI download
inside the tool-call timeout) or silently downgraded to the built-in
browser tools where uvx was unavailable.

- Extract the install logic into _ensure_browser_use_cli() and run it
  from the agent_browser/browserbase post_setup branch too (Firecrawl
  and the Nous cloud row both declare post_setup: browserbase).
- Camofox is untouched: Firefox-based, no CDP surface, cannot be driven
  by the CDP-only browser-use harness.
- Failure stays non-fatal: uvx fallback, then built-in tools.

Tests pin the contract: every browser post_setup key except camofox
attempts the install; camofox never does; install failure never raises.
2026-08-14 13:05:14 -07:00
kshitij d6a5cb9725 Merge pull request #85147 from kshitijk4poor/feat/unified-deadline-layer
feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
2026-08-15 01:10:40 +05:30
Soheil Fakour 0c6761c511 fix(gateway): restart_after_turn_timeout default 6h -> 30min (#79133)
The 21600s (6h) default shipped in #77184 makes an interactive
'hermes gateway restart' block for up to six hours when a turn wedges
(hung tool call, wedged event loop, stuck provider stream) — the exact
scenario the cap exists for. The intent (don't force-kill an agent
mid-turn) is sound, but the default must be a safety valve for hung
agents, not a target latency.

Lower to 1800s (30 min): still protects the overwhelming majority of
long autonomous turns (tool calls have their own timeouts well below
that), keeps worst-case interactive restart latency human-tolerable, and
users running very long unattended turns can raise it in config.yaml.

RED: new contract test fails on old 21600 default. GREEN: 4/4.
2026-08-14 11:22:05 -07:00
rob-maron 0e4e8baf63 add deepseek v5 pro 0813 to model catalog (#86256) 2026-08-14 18:14:27 +00:00
Austin Pickett 29d0cc2602 fix(dashboard): treat Ctrl+C serve shutdown as a clean exit (supersedes #52970) (#85711)
* fix(dashboard): suppress Ctrl+C shutdown traceback

* fix(dashboard): extend clean Ctrl+C exit to the Windows serve branch

The Windows loop-factory branch (and its pre-0.36 asyncio.run fallback)
runs under the same uvicorn capture_signals() re-raise as the POSIX
path, so console Ctrl+C leaked the identical KeyboardInterrupt
traceback there. Guard both serve calls with the same clean-exit
contract, keeping the import-resolution try/except comment accurate
(genuine serve-time errors still propagate).

Also ports the reworded POSIX-test docstring (the serve path is no
longer 'byte-for-byte unchanged'), wraps the POSIX KI test in
pytest.fail so a regression reports red instead of aborting the pytest
session, and adds the windows_only sibling test.

Extends #52970 to the whole bug class.

* chore: map contributor email for @wangs1203

* test(dashboard): actually exercise the pre-0.36 Windows fallback KI contract

Copilot review caught that patching uvicorn._compat.asyncio_run with
raising=False makes the import succeed, so _runner is non-None and the
extra asyncio.run patch never covered the fallback. Split it out: a
dedicated windows_only test halts the _compat import (None in
sys.modules) so the fallback branch is genuinely selected, then asserts
the same clean-KI contract on bare asyncio.run.

---------

Co-authored-by: Emiya·Leon <wangs.coder@gmail.com>
2026-08-14 12:02:05 -04:00
Teknium 6f53373eb2 fix(cli): route startup guard through the selection-guard registry; cover the light oneshot fast-path
Follow-ups on top of the salvaged #70324:
- _confirm_startup_expensive_model_override evaluates the unified
  registry (combined_selection_warning) so id-keyed guards like the
  data-training-tier warning fire at startup too, not just the cost guard.
- The Termux-adjacent light oneshot fast-path (added after the PR
  branched) ran _run_and_exit_oneshot without the guard — same bug
  class, third sibling site now covered.
2026-08-14 01:31:31 -07:00
lkz-de 83d373aae6 fix(cli): guard expensive startup model overrides
Run the expensive-model warning for explicit startup `-m` / `--provider`
overrides before the chat loop starts, and fail closed for non-interactive
invocations that select an expensive or known-confusing model.

Also classify Nous paid-model 404s that say credits are required as billing
exhaustion so they fail fast with billing guidance.

Tested:
- scripts/run_tests.sh tests/hermes_cli/test_cli_startup_model_cost_guard.py tests/hermes_cli/test_model_cost_guard.py tests/agent/test_error_classifier.py -- --tb=short -q
2026-08-14 01:31:31 -07:00
Dustin Persek 54cc39aa15 fix(models): don't trust foreign catalog pricing for custom/unknown providers
Custom providers (custom:xxx) serve their own pricing; models.dev stores
OpenRouter prices for the same model ids. The cost guard fired on that
foreign pricing and blocked composer/CLI model switches on custom
providers with a wildly wrong warning (#54348).

expensive_model_warning now only trusts model_info/models.dev pricing
when the provider maps to a models.dev provider and the info's
provider_id matches, and only consults the pricing-entry lookup when
the billing route is known. Salvaged from #54422; the PR's desktop-hook
half predates the use-model-controls rewrite and is superseded by the
hook's existing rollback handling.
2026-08-14 01:31:26 -07:00
Teknium 9166530942 feat(models): unify selection-time guards into one registry across all surfaces
Adds hermes_cli/model_selection_guards.py: a single evaluation point that
runs every selection guard (cost + the new data-policy guard) and returns
the warnings that fired. All seven model-selection surfaces (CLI picker,
cli.py TUI modal, gateway typed /model, dashboard web_server, TUI gateway,
Telegram and Discord pickers) now call the registry instead of importing
model_cost_guard directly — so the data-training-tier warning from
PR #81416 fires everywhere at once, and future guards need zero surface
wiring.

Guard modules keep their public APIs; existing mock patch points
(hermes_cli.model_cost_guard.expensive_model_warning) remain valid.
2026-08-14 01:06:13 -07:00
Beto de Paola a06f1d7617 feat(models): warn on data-training tiers at model selection
muse-spark-1.2-contributor is heavily discounted BECAUSE Meta uses your
prompts and completions to train future models. Selecting it for the price
without realising the data trade-off is a footgun.

Add hermes_cli/model_data_policy_guard.py (mirrors model_cost_guard):
data_training_warning(model_id, provider, base_url) -> DataTrainingWarning|None,
driven by a vendor-agnostic rule table. The status is not machine-readable on
/v1/models or models.dev, so the v1 rule keys on the documented '-contributor'
model id (fires regardless of provider, so it also covers custom/gateway
routes). Message mirrors Meta's pricing-doc language and figures
(https://dev.meta.ai/docs/pricing-rate-limits/).

Wire it into the CLI model picker's confirm flow (auth.py) as a [y/N]
disclosure, chained after the expensive-model cost guard. Fires only on the
contributor tier; silent on muse-spark-1.1/1.2 and all other models.
2026-08-14 01:06:13 -07:00
Shedrack Eze ce20857f5a fix: exclude launchd-managed gateway from orphan reaper on macOS
`_reap_unsupervised_gateway_orphans()` short-circuits on Linux hosts
with systemd via `supports_systemd_services()`, but returns `False` on
macOS — there is no systemd.  This means the orphan reaper runs
unconditionally on macOS and treats the launchd-managed gateway as an
unsupervised orphan, SIGTERM-ing it.

When Hermes Desktop opens, `hermes serve` calls this function during
startup (web_server.py line ~251, gated on `HERMES_DESKTOP == "1"`).
The launchd gateway is killed, launchd restarts it via `KeepAlive: true`,
and the user sees a spurious gateway restart every time they reopen the
Desktop app.

Fix: exclude PIDs managed by launchd (`_get_service_pids()` already
returns launchd-managed PIDs on macOS) from the orphan scan, the same
way systemd PIDs are excluded on Linux.

Tested on macOS 26.5 (Tahoe) with Hermes Desktop 0.16+ and a
launchd-managed gateway (`ai.hermes.gateway` plist with `KeepAlive`).
Before the fix, quitting and reopening Hermes Desktop restarted the
gateway every time. After the fix, the gateway stays running across
Desktop quit/reopen cycles.
2026-08-14 13:15:56 +05:30
Teknium 1c971769ec feat(gateway): concise background process notifications by default
Background process completions on messaging platforms now default to a
one-line status message (✅/❌ + command + duration; failures append a
short output tail) instead of dumping the raw output buffer into the
chat. New display.background_process_notifications mode 'concise' is
the default; 'all' keeps the old raw-dump behavior for anyone who wants
it. Config migration v35 moves users still on the old implicit default
'all' to 'concise' on their next update; explicit result/error/off
choices are preserved.
2026-08-14 00:26:11 -07:00
brooklyn! 529ee80ac0 Recover from a half-replaced desktop bundle instead of requiring a reinstall (#85887)
* fix(desktop): load the intact renderer bundle when an update tears one copy

index.html and the hashed chunks it names are one generation. A packaged app
ships that bundle twice (inside app.asar and, via asarUnpack, beside it in
app.asar.unpacked), so an update that replaces the app while its files are
locked can leave the two copies from different generations. resolveRendererIndex
took the first index.html that existed, so it could pick the torn one and the
window died on its first lazy import with "Failed to fetch dynamically imported
module" -- with no way out, because every relaunch reloaded the same copy.

Check each candidate's declared modules and prefer a complete generation; when
both are torn, log which files are missing and how to repair instead of leaving
the crash unexplained.

* fix(cli): rebuild the desktop app when its renderer bundle is half-replaced

The content stamp hashes the SOURCE tree, which an interrupted update leaves
intact, so `hermes desktop` reported "up to date" and skipped the rebuild that
would repair a torn bundle -- the app relaunched into the same crash and
reinstalling looked like the only option.

Treat a bundle whose index.html names missing chunks as stale regardless of the
stamp, and say so on the way into the rebuild.
2026-08-14 02:21:20 -05:00