Fixes the two live E2E blockers @ctaylor86 found on PR #77189 (macOS 26.3.1,
OpenSSL 3.6.3):
- retry the PKCS#12 export with -legacy when security import rejects the
OpenSSL 3 default format ('MAC verification failed during PKCS12 import')
- trust the self-signed root for the codeSign policy (security
add-trusted-cert -r trustRoot -p codeSign) — an imported-but-untrusted cert
is invisible to find-identity -v and unusable by codesign
- gate success on find-identity -v -p codesigning (postcondition), and use
the same -v probe for idempotency so an untrusted leftover cert is repaired
instead of reported as done
Tests rewritten as stateful fakes (valid only after import+trust), plus new
coverage for the -legacy retry, trust failure, postcondition gate, and the
untrusted-cert repair path; sabotage-verified (reverting to the name-in-output
probe fails 4 tests). Docs: manual fallback now includes the Trust step.
Adds a one-shot `hermes desktop --setup-tcc-identity` command that creates a
self-signed code-signing certificate in the login keychain (openssl +
security import), grants codesign access to it, writes
desktop.macos_signing_identity to config.yaml, and re-signs the packaged app
with a certificate-anchored Designated Requirement.
macOS persists permission grants (Full Disk Access, Accessibility, Files and
Folders, microphone) against the app's code-signing identity, not its path.
The default identifier-pinned ad-hoc signature is stable across rebuilds, but
a certificate-anchored identity is the strongest guarantee — the same
mechanism yabai/skhd rely on. Previously users had to create the certificate
manually in Keychain Access; this command automates the whole flow and is
idempotent (re-run after updates).
Docs: desktop.md TCC section now leads with the command, keeps manual steps.
Tests: 4 new — fresh cert creation path, idempotent reuse, non-macOS no-op,
cmd_gui early-exit before build.
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):
1. The parallel batch planner now peels the tool_call bridge wrapper and
decides admission on the underlying tool — supports_parallel_tool_calls
works again when deferral is active. Unparseable wrappers stay
sequential barriers; bridged calls get exactly the admission the same
call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
service-name queries reach tools whose own name omits the service; the
dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).
Salvaged from #92693 by @alt-glitch with authorship preserved.
On Windows, the pre-update concurrent-instance gate aborted with exit 2
whenever ANY other process held the venv hermes.exe shim — including the
gateway itself, which _pause_windows_gateways_for_update() stops moments
later and the post-update restart phase brings back. Users with a running
gateway were forced into a manual taskkill dance before every update.
The gate now filters gateway runtimes out of the abort list and proceeds
when nothing else is concurrent. Classification delegates to
_is_pausable_gateway -> gateway.status.looks_like_gateway_command_line
(the canonical shlex-tokenized, profile-selector-aware matcher shared by
the Desktop preflight exemption and the venv-holder guard fallback), so
the gate's exemption and the pause machinery cannot drift apart. Anything
not positively identified as a gateway — REPLs, dashboard, Desktop
backend children, gateway MANAGEMENT commands like 'gateway status',
unreadable cmdlines — still aborts exactly as before, and the abort
message now lists only the PIDs that are actually the user's problem.
Surgical reapply of PR #37039 by @damadorPL onto current main (the gate
moved from hermes_cli/main.py to hermes_cli/update_cmd.py in the main.py
decomposition); his substring classifier was replaced with the canonical
matcher, which also fixes the 'hermes gateway status' misclassification
flagged in review.
Co-authored-by: Hermes <hermes@nousresearch.com>
Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.
- _restart_launchd_gateway_after_update() (his extraction, adapted):
plist-exists is the ONLY gate; launchd_restart() owns the
bootout/bootstrap/kickstart ladder for every plist-present state;
every failure path is loud and names the manual recovery command.
The gate-error 'except: pass' (the second silent variant) now counts
the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
supervised PID), composing his fix with the returned-is-not-supervised
guard.
- His regression suite adapted to the (restarted, failed) contract; the
old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
bug.
A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
Merge follow-up on top of the salvaged commits: fold the npm>=10 reduced
hidden-lockfile comparison (intersection of non-null fields) and the
annotation-field exclusions into a single entries_differ() helper, keeping
the workspace-closure scoping intact.
_tui_need_npm_install compared every field of the root package-lock.json
against node_modules/.package-lock.json. npm>=10/11 writes a reduced hidden
lockfile that omits declarative fields (version/dependencies/dev) and adds
extraneous, so nearly every package looked 'changed'; workspace link entries
("link": true, paths outside node_modules/) are never materialized by the
partial --workspace install. Both made the check return True forever, so
hermes --tui re-ran npm install (and dirtied package-lock.json) on every
launch (#84617).
Compare only the keys both sides record with non-null values (resolved,
integrity, ...), ignore workspace link entries and non-node_modules paths in
the missing-entry check, and treat extraneous as an npm runtime annotation.
Real skew (lockfile bumped while node_modules is behind) is still detected.
_tui_need_npm_install() compared ALL fields between
package-lock.json and .package-lock.json, treating npm's
intentionally-stripped metadata (version, license, engines,
dependencies, etc.) as real skew. This caused 'Installing
TUI dependencies…' + rebuild on every launch.
Two changes:
1. Compare only the *intersection* of fields between the
two lockfiles — a field present in the root lock but
absent in npm's hidden actualized lock is a normal npm
artefact, not a real dependency change.
2. Add , , ,
to _NPM_LOCK_RUNTIME_KEYS — these boolean annotations
are written non-deterministically between the two locks
and never indicate a real dependency skew.
Real version/dependency changes still change /
, which are present in both locks and caught by
the intersection comparison.
Fixes: 347 false-positive lockfile mismatches → 0.
On Termux the launch install also selects ui-tui's child packages/*
workspaces (include_child_workspaces=True), so npm installs each child's
devDependencies. The freshness closure only followed devDependencies for
the ui-tui workspace itself, so a devDependency unique to a selected child
was dropped from the closure and a genuine missing package slipped past
_tui_need_npm_install.
Derive the closure from every workspace the install path selects, following
devDependencies for each. _npm_lock_workspace_closure now accepts the set of
selected workspace keys (dev-included roots); _tui_selected_workspace_keys
mirrors _make_tui_argv (ui-tui, plus child packages/* on Termux). Adds a
child-workspace-devDependency regression test (installs on Termux, ignored
off Termux) plus a closure-level dev-scope test.
_tui_need_npm_install compared the full multi-workspace root
package-lock.json against the hidden .package-lock.json, but the launch
install is scoped with npm install --workspace ui-tui and only writes the
ui-tui dependency closure. Every dep belonging solely to another workspace
(apps/desktop, web, ...) was therefore reported as missing, so the check
returned True and printed "Installing TUI dependencies..." on every launch.
Restrict the comparison to the ui-tui workspace's dependency closure,
computed from the root lock's packages map (following npm's node-resolution
walk and workspace symlinks). Standalone / own-lockfile layouts and any
case where the workspace can't be located fall back to the full comparison,
so drift on a genuine ui-tui dependency is still detected.
Fixes#66978
Sites under active development but tested over the public internet
(Vercel previews, ngrok tunnels, staging domains) are public DNS, so
the local-dev never-cache rule can't catch them. web.cache_exempt_hosts
lists hosts whose pages are always fetched live: exact, "*.wildcard",
or domain-suffix matching (label-boundary aware — mysite.dev covers
preview.mysite.dev but never evilmysite.dev). Checked on both store
and lookup, so adding an exemption takes effect immediately even for
entries cached before the config change.
Repeat searches (same normalized query + provider) within a 20-minute
TTL are served from an in-process memo, and concurrent identical
queries are single-flighted so a parallel subagent fan-out pays for
one vendor request instead of N. Requested limits bucket up to
10/20/50/100 so near-identical requests share an entry; callers get
their requested count sliced from the bucket.
Repeat extracts of the same URL are served from the existing
cache/web full-text store (previously written for read_file paging
but never read back), via a small JSON sidecar index. Disk-backed, so
CLI, gateway, cron, and subagents share it. Cached extracts re-run
the normal truncate pipeline, so per-call char_limit still works.
Both caches sit after every safety gate (secret-URL, SSRF, policy,
provider resolution) and directly around the paid vendor call — hits
skip only the network request. Only successful responses cache;
rescue-served responses are never cached (one-shot rescue must stay
one-shot). Config: web.cache_enabled (default on),
web.cache_ttl_minutes (default 20, clamped 1-1440).
Idea credit: query coalescing + num-bucketing pattern observed in
Apodex FrontierAgent (Apache-2.0).
Closes the two config-UI halves of the exclude-mode review (GottZ findings
5-7 on #94513):
- hermes mcp configure: on an exclude-mode server, unchecking a tool now
APPENDS a literal exclude and re-checking drops it — glob patterns are
preserved so future vendor tools keep getting filtered. Previously one
uncheck converted the whole config to a frozen include list (globs
silently deleted, new vendor tools invisible). Re-checked tools still
shadowed by a kept glob get an explicit warning instead of a silent
no-op.
- hermes tools MCP checklist: same exclude-mode write-back, plus display
now matches excludes via matches_name_filter (fnmatch) — glob excludes
previously rendered as if nothing were excluded.
- klaviyo manifest: post_install no longer tells users to append a param
the URL already pins; now documents how to get the FULL surface.
Live-verified through the real cmd_mcp_configure path: exclude-mode server
with ['*_secret_*', 'docs'], uncheck beta + re-check docs ->
exclude becomes ['*_secret_*', 'beta'], no include written.
171 tests green across test_mcp_catalog/test_mcp_config/test_mcp_tool.
The probe-fail path ignored prior state entirely: with no manifest
default_enabled it wrote include=None, which pops the whole tools
block. For exclude-mode manifests default_enabled is necessarily unset,
and for the 30+ OAuth entries the entry rewrite precedes first auth —
so the common reinstall-while-unreachable case removed the curated
excludes and enabled every tool on next connect. The fallback order is
now: prior include > prior exclude > manifest default > no filter.
The "Probing '<name>' for available tools..." line printed before the
exclude-mode short-circuit, which deliberately never probes (the test
suite asserts _probe_tools must not run there). Move the announcement
next to the actual probe call so install output matches what happened.
Review blockers (independent reviewer on #94513):
1. Reinstalling an exclude-mode catalog entry wiped the user's edited
tools.exclude, replacing it with manifest defaults. install_entry now
reads the prior exclude (like it already did for include) and re-writes
it verbatim on reinstall. Regression test added + sabotage-verified
(fails on old behavior); include-priority test added too.
2. aws-knowledge: exclude aws___retrieve_skill — vendor SKILL.md loader is
a vendor skill layer (live tools/list confirmed the tool exists).
3. betterstack: exclude list rewritten to cover the snake_case wire names
(vendor's own header examples show remove_dashboard) via globs alongside
the doc display-labels; caveat documented in the manifest — server is
OAuth-gated so pre-auth enumeration is impossible.
4. railway: exclude railway-agent (opaque server-side agent delegation,
acts outside Hermes's per-tool approval loop).
5. twelve-data: exclude oauth plumbing pseudo-tools + quota probe.
6. betterstack post_install no longer claims a fully-checked checklist —
exclude-mode bypasses the checklist; text now describes the applied
exclude list.
Live E2E: fresh temp HERMES_HOME — install applies manifest excludes,
user edit survives reinstall. 33/33 catalog tests green.
The cloudflare entry's 3,320-endpoint surface is ~43% product families a
personal/dev account never touches (Zero Trust org-fleet suite, Magic
Transit/WAN, Cloudforce One, Radar analytics, API Shield, legacy
migration surfaces). Ship a 34-pattern curated exclude list in the
manifest: 3,320 -> 1,905 tools kept, and everything Cloudflare adds
later stays enabled by default.
Mechanism, two small extensions:
- tools/mcp_tool.py: tools.include/exclude entries containing glob
metacharacters now match via fnmatch (plain names stay exact-match),
so a product family is one pattern instead of hundreds of stale
literals.
- hermes_cli/mcp_catalog.py: manifests may declare
tools.default_excluded (mutually exclusive with default_enabled);
install writes it to tools.exclude and skips the probe/checklist —
a 3,320-row curses checklist is not a UX. Prior user include
selections still win on reinstall.
Verified by replaying the real filter functions over the live-probed
3,320-tool list: 1,415 excluded, zero overmatch against a per-product
target audit; DNS/Workers/R2/D1/tunnels/Access/AI kept.
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.
This adds an opt-in, provider-agnostic checkpoint contract:
- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
(default false, documented in cli-config.yaml.example). When enabled,
compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
uncompressed transcript is preserved) unless a checkpoint-capable provider
confirms the durable checkpoint. Providers receive normalized direct
user/assistant evidence: tool rows, system messages, tool-call wrappers,
and prior compaction summaries are filtered host-side into one stable
contract. codex_app_server compaction is rejected under the gate because
it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
migration via _reconcile_columns) so summary provenance survives process
restarts; only the resume model history carries the marker, keeping
get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
(skip_memory=False) so a required checkpoint also guards those rewrites.
The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.
Refs #93986
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.
- hermes-agent skill is now essential: cannot be disabled (config reads
strip it, hermes tools writes drop it), cannot be deleted by
skill_manage, is re-seeded past curator suppression, and is seeded
even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
file+terminal+vision+skills: read_file cannot read images and points
at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
tool isn't loaded.
All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.
- tools/browser_tool.py: remove _extract_relevant_content and
_get_extraction_model; oversized snapshots always truncate at line
boundaries, store the full tree to cache/web, and append a read_file
pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
pattern as session_search/PR #27590), cli.py defaults + env bridge,
gateway/run.py bridged keys, hermes config display, hermes model picker,
dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
LLM path is gone and stored files are secret-redacted
The docstring and all four caller comments still said the helper
refuses only / and top-level directories. Since #93757 it also refuses
the entire hermes-agent install tree. Bring the docstring and the
comments at the four credential-write call sites in line with the
actual behavior so future changes are not misled by a stale safety
description.
Follow-up to #93757.
OpenRouter's :nitro, :floor, :exacto, and :online suffixes are request-time
routing modifiers valid on any model id — /models lists only the base model.
validate_requested_model() compared the full suffixed id against the listing,
so a valid variant was either rejected outright or fuzzy-auto-corrected to
the base id, silently stripping the user's routing opt-in.
Now, for OpenRouter only, a recognized variant suffix validates the BASE id
against the live listing (and the curated-catalog soft-accept and static-
catalog fallback paths) while preserving the suffixed id for persistence and
API requests — checked BEFORE fuzzy correction. :free/:batch/:thinking
remain direct catalog SKUs and keep exact-match semantics; unknown suffixes
and unknown bases are still rejected.
Reported by JEB (Jakob's Hermes Agent) via Discord.
Per-site decisions:
1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
via a function-local import; keeps the swallow-everything fail-open
contract and the (proc) -> None signature (agent/shell_hooks.py imports
it by name; _kill_git_process_tree alias preserved). The old body is kept
verbatim as _legacy_kill_process_tree and used as fallback when the
delegation import/call fails. A final proc.kill() is retained on the
happy path so Popen bookkeeping sees the exit (matches old behavior).
2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
(delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
body sent SIGTERM then SIGKILL with zero grace between them; the shared
primitive sends SIGKILL only. With no grace period the observable effect
is identical, and the psutil descendant sweep now also reaches
agent-browser's setsid'd daemon grandchild, which killpg alone missed.
tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
the legacy fallback (its assertions describe the fallback's internals).
3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
kill when escalate=True) — expressed as two delegated calls:
kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
body terminated children before the parent; the shared primitive
signals the group atomically (child is a session leader via
start_new_session=True) plus an identity-aware descendant sweep —
strictly wider coverage, same signals.
4. gateway/status.py: KEPT BOTH SITES.
- terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
incompatible with the shared primitive — it must RAISE OSError with
taskkill's stderr on non-zero exit (callers branch on that), falls back
to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
fail-soft primitive would invert the error contract.
- reap_gateway_children (~l2029): NOT migrated. It operates on a
pre-snapshotted child list from a parent that is already dead
(psutil.Process(pid) on the parent would fail), and every signal is
wrapped in identity/ownership checks the primitive lacks: is_running()
identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
The coupling is the feature; migrating would delete the safety logic.
5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
Dev tooling that intentionally kills by CAPTURED pgid because the direct
child is usually already reaped (psutil/pid-based primitive cannot find
it), and it avoids the psutil import on the test-runner hot path. Its
docstring already documents why psutil is the wrong tool there.
New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
A held port made 'hermes serve' print only uvicorn's bare
'ERROR: [Errno 98/10048] error while attempting to bind on address'
and exit 1 — indistinguishable from a broken backend for the desktop
spawn and wrapping scripts.
- Preflight bind probe (matching uvicorn's SO_REUSEADDR bind flags)
before uvicorn.Server; on conflict print machine-readable
'BACKEND_PORT_IN_USE port=<port>' + a human hint naming likely
holders, exit 75 (EX_TEMPFAIL — existing repo convention, see
gateway/restart.py, kanban_db.py).
- Probe-to-bind race covered: SystemExit(1) from uvicorn's own bind
failure is re-checked and translated on both POSIX and Windows
runner paths.
- --port 0 (ephemeral) short-circuits the probe: unchanged behavior.
- HERMES_BACKEND_READY contract untouched.
- Tests: real held-socket repro (sentinel + exit 75, sabotage-proven
to fail as bare exit 1 without the fix), free-port boot regression,
ephemeral-port regression, probe/classification units.
- Docs: port-conflict paragraph under 'hermes serve' in
reference/cli-commands.md.
Simplify-pass follow-ups on the #87033 fix:
- _gateway_liveness_notice(plural=) authors both wording variants at one
site; removes the exact-substring .replace() that would silently no-op
if the create-path text is ever edited.
- Collapse the operator-precedence-trap conditional in list to a plain
'if jobs' — an empty list has nothing inert and now skips the probe.
- Fix docstring/code mismatch (builder returns gateway_running: True on
the happy path) and drop the dead try/except in
_warn_if_gateway_not_running (the helper never raises).
Follow-ups for the salvaged #93098:
- Move the tri-state liveness heuristic into hermes_cli.cron
(_builtin_gateway_liveness) so the CLI warning and the cronjob tool
share one implementation instead of two drifting copies.
- Surface gateway_running/warning on the list action too — an agent
inspecting jobs in a gateway-less environment has the same silent-
inert-job failure mode (#87033) as create. Empty lists stay quiet.
A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).
- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
that mismatch a declared prefix, logging a WARNING naming the env var and
expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
are never rejected. Valid env keys still win over the pool (precedence
unchanged).
Fixes#93593
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.
pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.
Fixes#93518.
The re-exec'd venv child spawned by
_reexec_dependency_sync_off_windows_shim completes every update step —
the receipt records success / "completed at command boundary" — but then
hangs in interpreter shutdown on a leftover non-daemon thread, freezing
the PowerShell window for minutes after "Update complete!". On the
hand-off path only (HERMES_UPDATE_REEXEC=1), after the receipt is
finalized, the update lock released, and stdio restored, flush and
os._exit(code) instead of unwinding — the same treatment #79040's cron
workaround applies. SystemExit codes (including early refusals)
propagate to the hard exit; real exceptions keep the normal raise path
so tracebacks still print. Non-hand-off invocations are untouched: the
marker env is set solely when the shim spawns the child.
Fixes#93581
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.
Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries
Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.
Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.
Builds on RelaxJonh's #93410. Fixes#93406
collect_fleet_versions() swallows every probe exception via
logger.debug() and returns whatever accumulated — which can be an
empty list. print_fleet_version_matrix([]) returns False (no rows
to report), so the update exits 0 with "success" even though no
gateway was actually verified.
After the restart phase touches live gateways (restarted_services or
killed_pids is truthy), an empty fleet snapshot means verification
failed, not that everything is healthy. Treat it as incomplete so
the receipt records "partial" and the exit code is 1.
Fixes#93406
Mirror the full update_cmd._holder_value_flags precedent: derive both
top-level value-flag sets from build_top_level_parser() with a cached
frozenset, and fall back to a handwritten snapshot if parser
introspection ever fails, so argv classification keeps working on a
broken tree. Parity test pins the derived sets against the live parser
so drift fails CI.
Builds on #93551 (fangliquanflq) and #93570 (aniruddhaadak80) for #93530.
Follow-up cleanup for salvaged PR #92838:
- Collapse redundant context_back and effective_allow_back into a single
allow_back variable (they were always identical, never reassigned)
- Update _run_curses_menu docstring to reflect draw_header's actual
signature including search and back_enabled kwargs
Decode Ghostty/Kitty enhanced selection and cancellation keys, make setup cancellation terminal, and add cross-terminal previous-step navigation to setup and model flows.
Refs #92833
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
Widen the exception guard from OSError to Exception (re-raising
KeyboardInterrupt/EOFError first) so any prompt_toolkit runtime
failure degrades to input() — matching the established pattern in
masked_secret_prompt. ValueError and RuntimeError can arise from
exotic stream wrappers or event-loop issues with the same root cause:
prompt_toolkit cannot attach stdin on the terminal.
Add test_line_input_falls_back_to_input_on_any_prompt_toolkit_failure
covering the ValueError case.
line_input() only guarded against a missing prompt_toolkit (ImportError),
not against prompt_toolkit failing at runtime. On some terminals isatty()
returns True but the asyncio event-loop selector rejects registering stdin
(macOS kqueue raises OSError EINVAL / 'Invalid argument' for fd 0), so
prompt_toolkit's Application.run() crashes while attaching its input.
This aborted 'hermes setup' at the first plain text prompt. Telegram hit it
first because its automatic/manual selection uses prompt() rather than the
curses-based prompt_choice() the other platforms use, but every text prompt
shared the same failure.
Catch OSError from the prompt_toolkit path and fall back to the built-in
input() reader, which needs no selector and works in cooked mode. The
prompt_toolkit raw-mode context manager restores terminal state on the way
out, so the fallback reads cleanly.