Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
'invalid response' errors (#7264) into the same empty-content abort
carve-out so those shapes also preserve the session (#94459's wider
classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
_COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.
- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py
Fixes#94448
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).
Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'
The 400 replays forever until the user manually starts a new session.
Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError). Applied
at both replay sites: the chat-message converter and the preflight
choke-point. Live tool-definition names are left untouched — they must
match the dispatch registry exactly. Pairing is by call_id, so
renaming a replayed function_call is safe.
call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.
Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).
Fixes#31666
Anonymous probes (_anonymous_json) are designed to probe server identity
before disclosing credentials. Sending the Hermes version on these probes
would fingerprint the exact version to an untrusted/MITM endpoint.
Keep User-Agent on authenticated requests (_headers) and multipart uploads
(_multipart_headers), which already send credentials.
Address maintainer review feedback (PR #66215, comment by @teknium1):
> `tests/test_desktop_mac_entitlements.py:47` reads `apps/desktop/package.json`
> from pytest. `AGENTS.md:1319-1329` requires assertions about `package.json`
> and JS-side artifacts to be in the JS/Vitest suite; otherwise CI
> classification can skip the regression test on a JS-only change.
The CI change classifier (`scripts/ci/classify_changes.py`) marks
`apps/desktop/package.json` as `_FRONTEND` (in `_PY_SKIP`), so a Python test
that reads it would be skipped on a JS-only PR — regression goes green on
the PR, red on main.
Move the regression to `tests-js/desktop-mac-usage-descriptions.test.ts`,
following the same convention as commit dbf86b923 ("test: port macOS
entitlements test from Python to vitest"), which ports an earlier Python
entitlements regression into `tests-js/desktop-mac-entitlements.test.ts`
for the identical reason. The new file is a sibling of that one — both
pin Desktop macOS manifest contracts, but they assert against different
files (`entitlements.mac.plist` vs `build.mac.extendInfo` in package.json).
The Vitest port mirrors the original assertions 1:1: every
`NS*UsageDescription` key pinned (parametrized over key + required
substring + reason), no leading/trailing whitespace or newlines in any
`extendInfo` string, and a drift-protection assertion that fails when a
new privacy key is added to the build config without a matching row.
A runtime type guard on `extendInfo` ensures a non-string plist scalar
raises a clean assertion error here ("`X` in build.mac.extendInfo must
be a string (got boolean)") rather than crashing the test runner with
`value.trim is not a function` deep in the whitespace test — caught by
Flash + GPT-OSS cross-vendor review.
Verified:
- `cd tests-js && npm run check` → typecheck clean, 14/14 tests pass
(4 files including the new one with 5 tests).
- Mutation: removing `NSAppleMusicUsageDescription` from
`apps/desktop/package.json` flips 1 test red with the exact symptom
("Info.plist privacy usage description \`NSAppleMusicUsageDescription\`
is missing"). Restore → 14/14 green.
- Mutation: adding an unpinned `NSSpeechRecognitionUsageDescription` with
whitespace flips 2 tests red (drift-protection + whitespace).
- Mutation: adding a non-string `CFBundleBooleanTest: true` flips the
whole file red with the clean "must be a string (got boolean)"
assertion (no downstream crash).
- `apps/desktop` Electron Vitest project still passes (42 files,
432 tests + 1 skipped).
Closes the maintainer comment thread on PR #66215.
Fixes#54551
The Hermes Desktop renderer initializes Chromium's audio stack on user
gesture (completion chimes via Web Audio API in completion-sound.ts,
voice TTS via voice-playback.ts, mic capture via use-mic-recorder.ts,
and an eager AudioContext prime in haptics-provider.tsx). On macOS 26+,
that initialization registers the helper with the MediaLibrary TCC
service (kTCCServiceMediaLibrary), which surfaces to the user as a
"Hermes wants to access Music" permission prompt even though Hermes
never reads or writes the Apple Music library.
The Info.plist (built from apps/desktop/package.json's build.mac.extendInfo)
already declares NSAudioCaptureUsageDescription and
NSMicrophoneUsageDescription, but NSAppleMusicUsageDescription was missing
from the desktop app entirely. macOS therefore shows a system-default or
generic prompt for the MediaLibrary bucket instead of an honest description
from the app.
Fix
---
Add NSAppleMusicUsageDescription to build.mac.extendInfo with copy that
disclaims Music library access while explaining the system audio stack
uses voice, TTS, and completion sounds.
Add tests/test_desktop_mac_entitlements.py to pin every NS*UsageDescription
key declared in the Desktop build config. The test:
- parametrized over a (key, required_substring, reason) table
- asserts no leading/trailing whitespace and no newline chars in any usage
string (electron-builder passes them through verbatim; control chars
render as broken prompt text)
- asserts drift-protection: a new NS*UsageDescription key added to the
build config without a matching test row causes a hard failure
Pattern reference: PR #59486 ("fix(desktop): add macOS contacts privacy
strings") is the open canonical for the same shape of fix for Contacts;
PR #64582 / PR #65220 extend it for Reminders. The closed duplicate PRs
Related, not in this PR
-----------------------
- PR #62601 (sounddevice on macOS) is the gateway/CLI side of the same
kTCCServiceMediaLibrary trigger.
- PR #45952 (macOS permission broker foundation) is architectural work
for centralized TCC handling; this fix does not depend on it.
- PR #52839 (browser automation Chrome launch) mutes Chromium audio in
a different surface; the same pattern is recorded there.
Fixes#54551
The legacy ad-hoc fallback signed and verified successfully but still
fell through to return False, contradicting the fixup's documented
contract. The success witness codified the contradiction. Return True
on the verified success path; the caller ignores the return value, so
no behavior change beyond the contract correction.
Addresses round-2 review feedback on #90961. The previous commits
scoped the keychain deletion to the legacy ad-hoc fallback, but the
reviewer correctly held the blocker: the fallback ran codesign with
check=False, ignored the result, and unconditionally deleted 'Hermes
Safe Storage' — permanently orphaning gateway and native OAuth
credentials even when signing failed or a configured identity had
failed and routed into the fallback.
This commit removes the deletion entirely:
- _desktop_macos_reset_keychain_safe_storage is gone; no code path
touches the keychain item anymore.
- The legacy fallback now checks the codesign result and runs
codesign --verify --deep --strict; any failure leaves the item
untouched and prints a warning.
- The keychain prompt after an ad-hoc re-sign is recoverable
(Always Allow updates the ACL partition list and preserves the
key); deletion is not. The durable proof-carrying migration
belongs in Electron (safeStorage can read the old key) and is
tracked as a follow-up.
Tests: 4 witnesses (stable path, default no-config success, fallback
failure, fallback success) all mutation-verified against both the
deletion regression and the ignored-codesign-result regression.
The fixup no-ops on non-macOS (sys.platform guard), so the new
regression tests must carry the same @pytest.mark.macos_only marker
as their siblings (test_relaunchable_fixup_falls_back_to_legacy_adhoc_on_failure).
Without it the legacy-adhoc test failed on the Linux CI runner where
the fixup returns True before reaching the reset path.
The previous commit deleted the 'Hermes Safe Storage' keychain item after
every successful re-sign, including the stable certificate-anchored
identity path. On that path the designated requirement is stable across
rebuilds, so after the first launch under the new identity the keychain
ACL already matches; deleting the item on every update permanently
orphaned gateway-token and native-OAuth credentials that were working
fine (both are safeStorage-backed: electron/main.ts connection config
and native-oauth-tokens.json).
Addresses review feedback on #90961:
- Rename _desktop_macos_update_keychain_acl -> _desktop_macos_reset_keychain_safe_storage (it deletes, it does not update an ACL).
- Only invoke it on the legacy ad-hoc fallback path, where every rebuild
produces a new cdhash so the ACL can never match and the alternative
is a recurring prompt. The trade-off (re-enter credentials once per
update) is documented; the durable fix is a stable signing identity.
- Add regression tests: stable path must NOT reset, ad-hoc fallback MUST.
A listed profile-less row in a multi-profile setup must not DELETE/archive against the primary backend. Unresolved ownership keeps the row, pins, and unread state and never calls the mutation.
Review follow-up for the salvaged #94848: the ProfileRouteRejected marker
looked write-only inside the primary handler. Document that
_handle_message's ingress gate reads the same marker to drop the message
fail-closed, and add a regression test showing the rejected route is
stamped once, dispatch falls back to the default home, and routing is not
re-run on redelivery.
cua-driver maps the agent-cursor overlay as a fullscreen, always-on-top,
all-workspaces X11 window (save-unders composited). When a computer-use
session ends uncleanly — an agent interrupted mid-capture, a stale target
window, a driver error — that window can be left stuck above every app on
every workspace, wedging desktop input until the app is restarted. This
is the same failure class as the HUD's transparent always-on-top window on
Mutter/X11 (#83473), and it bit a real user: an interrupted capture froze
the desktop, the app had to be force-restarted, and the overlay window was
still mapped fullscreen afterwards.
The overlay is cosmetic (a tinted cursor sprite); the driver, captures,
and synthetic input all work without it. `_cua_no_overlay()` already
defaulted it off on macOS (idle CPU redraw loop, #28152/#47032) and
headless/WSL2 Linux; this extends the same auto-detect to X11 desktop
sessions, where raw X11 stacking has no compositor-owned surface to tear
down with the driver's connection. Wayland keeps the overlay: the
compositor owns the layer-surface lifecycle, so a dead driver cannot leave
a stuck top window.
Behavior contract unchanged: an explicit `computer_use.no_overlay: false`
still restores the cursor on any platform, and `true` forces it off.
Tests: X11 (DISPLAY set, no Wayland env) and XDG_SESSION_TYPE=x11
auto-detect off; Wayland keeps the overlay; explicit false overrides
auto-detection on X11.
Review follow-up to the AI code-review pass on PR #92074: the bridge task's
handoff into _schedule_secondary_profile_reconnect was unguarded at both call
sites inside the parked coroutine. The scheduler touches live registries
(_profile_failed_platforms slot creation, background-task registration), so an
unexpected raise there would kill the parked task as an unretrieved-task
exception — logged only at GC time via "Task exception was never retrieved",
where no operator ever looks. A fix whose entire purpose is to stop a platform
dying silently should not contain its own silent-death path; both handoff sites
now wrap the scheduler call with logger.exception so the failure lands in
gateway.log with profile and platform context.
The early-exit branch (gateway already _running when the bridge starts) had the
identical exposure and is guarded the same way — same bug class, fixed together.
Regression test drives a handoff raise end-to-end through the real bridge task:
the await completes cleanly, the error is captured in gateway.run's logger, and
no adapter or failed-platform slot leaks behind the failed handoff.
When gateway.multiplex_profiles is active, a secondary profile whose platform
adapter fails its initial connect at startup was silently given up on: the
failure branches in _start_one_profile_adapters() logged and disconnected, but
never scheduled recovery. One unlucky connect window during a Telegram API
outage left the profile permanently silent until manual restart (~80 min in
the observed incident), while the mid-run fatal path already recovers via
_handle_profile_adapter_fatal_error() -> _schedule_secondary_profile_reconnect().
The same gap hit both failure shapes: a clean False return from
_connect_initial_adapter_with_timeout() and an exception escaping it.
Fix: call _schedule_secondary_profile_startup_reconnect() from both startup
failure branches after _safe_adapter_disconnect(). Because secondary adapters
are started mid-start(), before self._running flips True, the regular
scheduler's not-self._running guard would silently drop the request — so the
new bridge parks a background task until startup completes (or shutdown
begins) and then hands off to _schedule_secondary_profile_reconnect()
verbatim: backoff, fresh-adapter rebuild under the profile runtime scope,
slot dedupe, and shutdown cancellation all come from the existing path.
Non-retryable failures are dropped at scheduling time exactly as the regular
scheduler drops them, keeping duplicate-credential/auth-failed startups dead
instead of looping.
Fixes#92064
The #87196/#87720 conflict resolution kept the bounded-drain helper and
its constant; the windows_only kill-tree test still asserted the dropped
_CUA_INSTALLER_REAP_TIMEOUT name. Same 2-communicate contract, surviving
constant.
Re-enables the routine confirmed-upgrade path on Windows that #95008
deferred wholesale, now that every unattended-hostile surface is closed:
- stdin=DEVNULL (salvaged #79871): upstream's Read-Host consent prompt
can't block a hidden console.
- Bounded post-kill drain (salvaged #87720): a kill-surviving descendant
holding the stdout pipe can't strand the update past its ceiling.
- 120s background ceiling (salvaged #87196): safe now that a legitimate
600s lock wait can't occur on this path.
- NEW lock preflight: upstream's install lock held by a live process ->
skip in ~0s instead of eating its 600s stale-lock window (the actual
11-minute hang observed 2026-08-25; UAC was a red herring — base
install is no-admin by upstream design).
- NEW 5s network preflight: github.com unreachable -> skip immediately.
- Windows unattended runs pass -NoAutoStart, skipping the ONLY
install.ps1 branch that self-elevates (autostart task re-registration).
- Timeout diagnosability: partial installer output is logged on kill so
the next hang names its stage instead of dying silently.
Contract repairs and fresh installs stay interactive-only (SmartScreen /
first-time elevation legitimately need a human).
On Windows, `hermes update` can hang past its own 660s cua-driver timeout
until the user kills an orphaned PowerShell by hand. The timeout ceiling is
not the problem; the code that runs after it is.
`_run_cua_driver_installer` handles `TimeoutExpired` by killing the process
tree and then draining the pipes with a bare `proc.communicate()`. The kill
is best-effort by construction: every `psutil.Error` in `_kill_installer_tree`
is logged at debug level and stepped over, on the reasoning that a partly
killed tree beats none. That is the right call, but it means the drain has to
survive a partial kill, and an unbounded drain does not.
The concrete case is the one reported. `install.ps1` self-elevates through
`Start-Process -Verb RunAs`, so the descendant runs at High integrity and a
medium-integrity `child.kill()` raises `AccessDenied`. The per-child handler
logs it and continues. That survivor is still holding the `stdout=PIPE` write
handle it inherited, so the following `communicate()` waits for an EOF that
arrives only when somebody kills that process manually. A bounded 660s wait
becomes an unbounded one, after the warning has already printed.
Bound the drain instead. A kill that landed closes the pipe immediately, so
this costs nothing on the normal path; a kill that did not costs 15s rather
than forever. The original `TimeoutExpired` is re-raised either way, so the
existing manual re-run hint still prints and the update unwinds. Losing the
tail of a timed-out installer's log is the cheaper half of that trade, and it
is only lost in the case where the run already failed.
The drain deliberately does not close the pipe handles. `communicate()`'s
reader threads are still blocked on them and closing underneath them races;
they are daemon threads, so abandoning them does not hold the interpreter
open.
Both timeout handlers (streaming and captured) now go through one helper.
The streaming child inherits the console rather than a pipe, so it is much
harder to stall there, but the two branches should not drift on a rule this
small.
Tests: 5, in a new `TestInstallerTimeoutDrainIsBounded`. Two fail without the
fix, including the reported scenario end to end (a child kill refused with
`psutil.AccessDenied`, asserting the drain still carries a deadline). The
deadline is asserted as a kwarg rather than by timing, because a test that
proved the hang by hanging would be the same defect wearing a test's name.
Scope note: this does not touch the `stdin` inheritance that lets
`install.ps1`'s `Read-Host` block in the first place. That is #79684 and open
PR #79871 already carries the one-line `stdin=DEVNULL` fix; the two are
independent and neither subsumes the other, since `DEVNULL` cannot unblock a
UAC elevation dialog.
Fixes#87703
profiles.list opened every profile state.db as a writable SessionDB,
which waits out write-lock patience while that profile's backend is
mid-turn. The desktop RPC timed out and Bot Mode's infinite React
Query retry kept the sidebar on a spinner.
Inspect those DBs read-only and bound roster retries so names still
paint.
Two follow-ups on top of the #86391 salvage:
- check_macos_tcc_grants: a certificate-anchored DR (hermes desktop
--setup-tcc-identity, or a notarized release) now reports as stable in its
own class instead of falling into the identifier-pinned message; the
identifier-pinned message points at --setup-tcc-identity for the strongest
anchor.
- collect_relay_plugin_cutover_findings: only merge process-level env vars
when env_map is None (run_doctor's live path). An explicit env_map is a
complete environment description — merging os.environ on top made
report_deprecated_config_and_env non-hermetic on boxes exporting legacy
relay vars (10 findings vs the expected 2 in
test_report_does_not_count_as_blocking_issue).
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
so a hanging codesign degrades to the unreadable-DR warning, never crashing
the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
GPT-OSS review: an empty codesign output would fall through to the
'stable identity' branch and false-positive. Guard with and
cover the empty-string case. Flash review: the non-macOS silence test
mocked the bundle to None, so it never exercised the platform guard;
mock a real path instead.
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.
- hermes doctor: new check_macos_tcc_grants() reports the desktop
bundle's DR class (cdhash-pinned → grants reset on every update;
identifier-pinned → stable) and prints the exact stale-grant repair
(tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
documents the one-time re-grant for pre-fix grants.
Closes#86385
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.
- New config key delegation.surface_child_process_notifications
(default false = suppress). Flag true restores the previous behavior
exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
watch_disabled events whose task_id starts with 'sa-' when the flag
is false, logging at debug with session_id+task_id for diagnosis.
Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
crash the drain loop.
- Docs: delegation.md + configuration.md.
Follow-up to the salvaged #94296: the two guards covered the repair and
confirmed-update branches, but when cua-driver is enabled yet not
installed at all, control still reached _run_cua_driver_installer() and
an automatic 'hermes update' would launch the interactive install.ps1
anyway. Add the same defer before the installer run, keep POSIX
behavior unchanged, and give the confirmed-update message a natural
fallback when latest_version is unknown.
Two interaction seams between the #92693 salvage (merged as #95050) and
this branch: the source-label indexing test now compares in token space
(the stemmer shortens 'catalogsource' to 'catalogsourc'), and the
unregistered-core-name describe test forces the unregistered condition
via monkeypatch instead of depending on which sibling test file imported
model_tools first.
The parallel determinism test warms _stem's lru_cache after ~11 distinct
stems, so almost no iterations reach the underlying stemmer and a shared
(non-thread-local) instance survives it. New test bypasses the cache with
per-iteration unique tokens via _stem.__wrapped__, so thousands of stems
run concurrently: a shared stemmer's mutable parse state fails it within
2,000 calls (verified — the mutant dies 8/8 runs; healthy runs stay green).
tool_search now takes queries: string[] (searched independently against
the same catalog, limit applies per query, default 5 / max 25) and
returns the split shape: per-query groups carry tool names only, one
shared tools map holds each matched tool's source, description (400-char
cap) and required parameter names once. When some queries miss, a single
top-level available_sources + hint block replaces the old per-response
fallback.
tool_describe now takes names: string[] and returns a map keyed by name;
unknown names collect in not_found (with the refresh hint) and
non-deferrable names keep their per-name spelling-check error in errors,
so one bad name no longer fails the whole call. Duplicates dedupe
silently.
The shared tokenizer now applies Snowball stemming (english, exact-pinned
snowballstemmer) at both index and query time, closing the measured
plural/singular miss where 'issues' failed to return create_issue. The
inline BM25 is unchanged. Stemmer instances are thread-local (they carry
mutable parse state and bridge dispatch can run on parallel tool-call
threads).
New config knobs under tools.tool_search: max_queries / max_describe_names
(default 10 each, floor 1, no upper clamp) bound the per-call array
inputs; over-cap calls error so the model repairs in one round-trip.
No backward compatibility with the single query/name shapes, by decision.
scripts/analyze_livetest.py renders both shapes since transcripts on disk
may predate this change.
Fixes the two live E2E blockers @ctaylor86 found on PR #77189 (macOS 26.3.1,
OpenSSL 3.6.3):
- retry the PKCS#12 export with -legacy when security import rejects the
OpenSSL 3 default format ('MAC verification failed during PKCS12 import')
- trust the self-signed root for the codeSign policy (security
add-trusted-cert -r trustRoot -p codeSign) — an imported-but-untrusted cert
is invisible to find-identity -v and unusable by codesign
- gate success on find-identity -v -p codesigning (postcondition), and use
the same -v probe for idempotency so an untrusted leftover cert is repaired
instead of reported as done
Tests rewritten as stateful fakes (valid only after import+trust), plus new
coverage for the -legacy retry, trust failure, postcondition gate, and the
untrusted-cert repair path; sabotage-verified (reverting to the name-in-output
probe fails 4 tests). Docs: manual fallback now includes the Trust step.
Adds a one-shot `hermes desktop --setup-tcc-identity` command that creates a
self-signed code-signing certificate in the login keychain (openssl +
security import), grants codesign access to it, writes
desktop.macos_signing_identity to config.yaml, and re-signs the packaged app
with a certificate-anchored Designated Requirement.
macOS persists permission grants (Full Disk Access, Accessibility, Files and
Folders, microphone) against the app's code-signing identity, not its path.
The default identifier-pinned ad-hoc signature is stable across rebuilds, but
a certificate-anchored identity is the strongest guarantee — the same
mechanism yabai/skhd rely on. Previously users had to create the certificate
manually in Keychain Access; this command automates the whole flow and is
idempotent (re-run after updates).
Docs: desktop.md TCC section now leads with the command, keeps manual steps.
Tests: 4 new — fresh cert creation path, idempotent reuse, non-macOS no-op,
cmd_gui early-exit before build.
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):
1. The parallel batch planner now peels the tool_call bridge wrapper and
decides admission on the underlying tool — supports_parallel_tool_calls
works again when deferral is active. Unparseable wrappers stay
sequential barriers; bridged calls get exactly the admission the same
call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
service-name queries reach tools whose own name omits the service; the
dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).
Salvaged from #92693 by @alt-glitch with authorship preserved.
On Windows, the pre-update concurrent-instance gate aborted with exit 2
whenever ANY other process held the venv hermes.exe shim — including the
gateway itself, which _pause_windows_gateways_for_update() stops moments
later and the post-update restart phase brings back. Users with a running
gateway were forced into a manual taskkill dance before every update.
The gate now filters gateway runtimes out of the abort list and proceeds
when nothing else is concurrent. Classification delegates to
_is_pausable_gateway -> gateway.status.looks_like_gateway_command_line
(the canonical shlex-tokenized, profile-selector-aware matcher shared by
the Desktop preflight exemption and the venv-holder guard fallback), so
the gate's exemption and the pause machinery cannot drift apart. Anything
not positively identified as a gateway — REPLs, dashboard, Desktop
backend children, gateway MANAGEMENT commands like 'gateway status',
unreadable cmdlines — still aborts exactly as before, and the abort
message now lists only the PIDs that are actually the user's problem.
Surgical reapply of PR #37039 by @damadorPL onto current main (the gate
moved from hermes_cli/main.py to hermes_cli/update_cmd.py in the main.py
decomposition); his substring classifier was replaced with the canonical
matcher, which also fixes the 'hermes gateway status' misclassification
flagged in review.
Co-authored-by: Hermes <hermes@nousresearch.com>
Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.
- _restart_launchd_gateway_after_update() (his extraction, adapted):
plist-exists is the ONLY gate; launchd_restart() owns the
bootout/bootstrap/kickstart ladder for every plist-present state;
every failure path is loud and names the manual recovery command.
The gate-error 'except: pass' (the second silent variant) now counts
the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
supervised PID), composing his fix with the returned-is-not-supervised
guard.
- His regression suite adapted to the (restarted, failed) contract; the
old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
bug.
A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>