mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.
The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.
Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.
Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.
Two fixes:
1. _check_via_local_git now detects fetch failure (returncode != 0 or
exception) and returns None instead of falling through to stale
refs. The caller treats None as 'check could not run' rather than
'up to date'.
2. check_for_updates no longer caches None results. Previously, a
None from a failed check was cached for 6 hours, suppressing
retries until the cache expired. Now only conclusive results
(0 or >=1) are cached, so the next check attempt runs immediately
on the next call.
Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
Follow-up to the salvaged #95207 fix, completing the misreport bug class:
- restore() now also drops delete_targets whose unlink failed (OSError
swallowed) from restored_files — the sibling of the kept-oversize
misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
(slash_commands + gateway.rollback.kept_oversize locale key in all 17
catalogs) now tells the user which files were kept because the size
cap excluded them from every checkpoint; previously the file was
correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
Integration fixups so #82529's generation signing and #95131's interpreter
anchor cover each other's gaps (without these, each fix leaves the other's
rotation path broken):
- macos_tcc_anchor store detection now recognizes .hermes-runtime/python/
generation-* stores: repair_vulnerable_runtime() rebuilds the venv against
a generation interpreter, replacing the anchored bin/python with a fresh
symlink — previously the anchor then read 'not uv-managed' and NEVER
re-anchored, so every SQLite CVE repair silently orphaned terminal TCC
grants (the exact #82427 scenario, path-keyed).
- _install_anchor signs the anchor copy with the same identifier-pinned DR
(via managed_uv._macos_sign_managed_python) before it goes live: copy2
carries the source build's cdhash-based signature, so an unsigned refresh
would still change the stored csreq on every patch bump/repair despite
the stable path. Best-effort, never blocks the anchor.
- Tests: generation-store recognition + repair-generation anchoring +
sign-on-install call (sabotage-verified: dropping the generation root
marker fails both new tests).
Follow-up on the cherry-picked #90030: candidate.resolve() breaks the fix
on the standard Linux venv layout, where venv/bin/python is a symlink to
the base interpreter. Resolving makes the venv python compare equal to
the dependency-less base (so the swap never happens), and returning the
resolved target would spawn the bare base interpreter, bypassing
pyvenv.cfg — the fix would silently not fix#90026 on the exact platform
it was reported from. Compare and return normalized UNRESOLVED paths:
the venv path IS the interpreter's identity. Adds the symlink-layout
regression test; live-E2E'd with a real dependency-less base runtime.
Under an SSH remote backend the web server is launched by running the uv
BASE interpreter with the venv's site-packages injected into sys.path at
startup, so sys.executable is a dependency-less python. Detached
dashboard actions spawned from it (Update now, restart, anything routed
through _spawn_hermes_action) inherited neither the injected path nor a
PYTHONPATH and died on the first third-party import — 'Update now'
always failed instantly with ModuleNotFoundError: No module named 'yaml'
while 'hermes update' from the venv worked (#90026).
_dashboard_spawn_executable now prefers the install's own venv
interpreter (venv/bin/python, venv/Scripts/python.exe) when it differs
from sys.executable, resolving the same dependency set the venv launcher
provides. Same-interpreter launches return sys.executable unchanged,
preserving the Windows console-ownership behavior verbatim, and layouts
without an install venv keep the old fallback.
Hardening on top of the TCC daemon-identity salvage:
- _validate_cua_driver_app_signature: codesign -dv gate requiring EXACT
Identifier=com.trycua.driver and the official team (4YEC26S9KF) before
/usr/bin/open hands the bundle to LaunchServices — the identity fix must
not double as a launcher for arbitrary/impostor bundles (suffixed
identifiers and wrong teams rejected; unsigned dev builds only via
computer_use.allow_unsigned_driver: true in config.yaml).
- _resolve_cua_driver_app_path: derive the bundle ONLY from the resolved
driver binary — the /Applications fallback could launch a DIFFERENT
install than the manifest resolved.
- open -n -g: don't activate/steal focus when launching the daemon.
- 7 new tests incl. sabotage-verified exact-match assertions.
Grafted from #76433's review direction (@Chadmc9889's original fail-closed
validation requirement).
Co-authored-by: Chadmc9889 <Chadmc9889@users.noreply.github.com>
The legacy ad-hoc fallback signed and verified successfully but still
fell through to return False, contradicting the fixup's documented
contract. The success witness codified the contradiction. Return True
on the verified success path; the caller ignores the return value, so
no behavior change beyond the contract correction.
Addresses round-2 review feedback on #90961. The previous commits
scoped the keychain deletion to the legacy ad-hoc fallback, but the
reviewer correctly held the blocker: the fallback ran codesign with
check=False, ignored the result, and unconditionally deleted 'Hermes
Safe Storage' — permanently orphaning gateway and native OAuth
credentials even when signing failed or a configured identity had
failed and routed into the fallback.
This commit removes the deletion entirely:
- _desktop_macos_reset_keychain_safe_storage is gone; no code path
touches the keychain item anymore.
- The legacy fallback now checks the codesign result and runs
codesign --verify --deep --strict; any failure leaves the item
untouched and prints a warning.
- The keychain prompt after an ad-hoc re-sign is recoverable
(Always Allow updates the ACL partition list and preserves the
key); deletion is not. The durable proof-carrying migration
belongs in Electron (safeStorage can read the old key) and is
tracked as a follow-up.
Tests: 4 witnesses (stable path, default no-config success, fallback
failure, fallback success) all mutation-verified against both the
deletion regression and the ignored-codesign-result regression.
The previous commit deleted the 'Hermes Safe Storage' keychain item after
every successful re-sign, including the stable certificate-anchored
identity path. On that path the designated requirement is stable across
rebuilds, so after the first launch under the new identity the keychain
ACL already matches; deleting the item on every update permanently
orphaned gateway-token and native-OAuth credentials that were working
fine (both are safeStorage-backed: electron/main.ts connection config
and native-oauth-tokens.json).
Addresses review feedback on #90961:
- Rename _desktop_macos_update_keychain_acl -> _desktop_macos_reset_keychain_safe_storage (it deletes, it does not update an ACL).
- Only invoke it on the legacy ad-hoc fallback path, where every rebuild
produces a new cdhash so the ACL can never match and the alternative
is a recurring prompt. The trade-off (re-enter credentials once per
update) is documented; the durable fix is a stable signing identity.
- Add regression tests: stable path must NOT reset, ad-hoc fallback MUST.
The self-updater rebuilds the desktop app locally via electron-builder after
every update (4aa9f738ce). On macOS, the rebuilt app gets ad-hoc signed,
producing a different cdhash than the original CI-signed build. macOS ties
the 'Hermes Safe Storage' keychain item's ACL to the code signature, so
the new signature doesn't match → macOS re-prompts for keychain access on
every launch.
After re-signing, delete the existing keychain item so Electron recreates
it with the correct ACL for the newly-signed app on next launch. The
trade-off: previously encrypted tokens become unreadable (the user
re-enters the gateway token once), but the keychain prompt stops appearing
on every launch.
The delete-generic-password command doesn't require reading the secret
(no ACL check), so it runs without prompting.
Follow-ups on the #85416 salvage:
- _store_bin_names()/_sibling_names() derive the versioned interpreter name
from the RUNNING interpreter instead of a hardcoded 3.11-3.13 list, with a
sorted python3.* glob fallback for stores/aliases built by a different
Python minor (major bumps, fixtures) — no future version bump can silently
leave an alias resolving back into the versioned uv store.
- _repoint_alias_symlinks unions expected alias names with the versioned
symlinks actually on disk before repointing.
Re-enables the routine confirmed-upgrade path on Windows that #95008
deferred wholesale, now that every unattended-hostile surface is closed:
- stdin=DEVNULL (salvaged #79871): upstream's Read-Host consent prompt
can't block a hidden console.
- Bounded post-kill drain (salvaged #87720): a kill-surviving descendant
holding the stdout pipe can't strand the update past its ceiling.
- 120s background ceiling (salvaged #87196): safe now that a legitimate
600s lock wait can't occur on this path.
- NEW lock preflight: upstream's install lock held by a live process ->
skip in ~0s instead of eating its 600s stale-lock window (the actual
11-minute hang observed 2026-08-25; UAC was a red herring — base
install is no-admin by upstream design).
- NEW 5s network preflight: github.com unreachable -> skip immediately.
- Windows unattended runs pass -NoAutoStart, skipping the ONLY
install.ps1 branch that self-elevates (autostart task re-registration).
- Timeout diagnosability: partial installer output is logged on kill so
the next hang names its stage instead of dying silently.
Contract repairs and fresh installs stay interactive-only (SmartScreen /
first-time elevation legitimately need a human).
On Windows, `hermes update` can hang past its own 660s cua-driver timeout
until the user kills an orphaned PowerShell by hand. The timeout ceiling is
not the problem; the code that runs after it is.
`_run_cua_driver_installer` handles `TimeoutExpired` by killing the process
tree and then draining the pipes with a bare `proc.communicate()`. The kill
is best-effort by construction: every `psutil.Error` in `_kill_installer_tree`
is logged at debug level and stepped over, on the reasoning that a partly
killed tree beats none. That is the right call, but it means the drain has to
survive a partial kill, and an unbounded drain does not.
The concrete case is the one reported. `install.ps1` self-elevates through
`Start-Process -Verb RunAs`, so the descendant runs at High integrity and a
medium-integrity `child.kill()` raises `AccessDenied`. The per-child handler
logs it and continues. That survivor is still holding the `stdout=PIPE` write
handle it inherited, so the following `communicate()` waits for an EOF that
arrives only when somebody kills that process manually. A bounded 660s wait
becomes an unbounded one, after the warning has already printed.
Bound the drain instead. A kill that landed closes the pipe immediately, so
this costs nothing on the normal path; a kill that did not costs 15s rather
than forever. The original `TimeoutExpired` is re-raised either way, so the
existing manual re-run hint still prints and the update unwinds. Losing the
tail of a timed-out installer's log is the cheaper half of that trade, and it
is only lost in the case where the run already failed.
The drain deliberately does not close the pipe handles. `communicate()`'s
reader threads are still blocked on them and closing underneath them races;
they are daemon threads, so abandoning them does not hold the interpreter
open.
Both timeout handlers (streaming and captured) now go through one helper.
The streaming child inherits the console rather than a pipe, so it is much
harder to stall there, but the two branches should not drift on a rule this
small.
Tests: 5, in a new `TestInstallerTimeoutDrainIsBounded`. Two fail without the
fix, including the reported scenario end to end (a child kill refused with
`psutil.AccessDenied`, asserting the drain still carries a deadline). The
deadline is asserted as a kwarg rather than by timing, because a test that
proved the hang by hanging would be the same defect wearing a test's name.
Scope note: this does not touch the `stdin` inheritance that lets
`install.ps1`'s `Read-Host` block in the first place. That is #79684 and open
PR #79871 already carries the one-line `stdin=DEVNULL` fix; the two are
independent and neither subsumes the other, since `DEVNULL` cannot unblock a
UAC elevation dialog.
Fixes#87703
When `hermes update` runs the cua-driver installer non-interactively,
stdout is captured (PIPE) but stdin is inherited from the parent process.
The upstream install.ps1 prompts `[Y/n]` when it detects a stale
cua-driver daemon, but the prompt goes to captured stdout (invisible)
while stdin waits for input — causing an 11-minute hang until the
watchdog timeout.
Fix: redirect stdin to subprocess.DEVNULL on the non-verbose path so the
installer reads EOF immediately instead of blocking. The installer exits
quickly with a non-zero code, and the existing handler shows the manual
re-run hint with the installer's captured output.
Fixes#79684
Two follow-ups on top of the #86391 salvage:
- check_macos_tcc_grants: a certificate-anchored DR (hermes desktop
--setup-tcc-identity, or a notarized release) now reports as stable in its
own class instead of falling into the identifier-pinned message; the
identifier-pinned message points at --setup-tcc-identity for the strongest
anchor.
- collect_relay_plugin_cutover_findings: only merge process-level env vars
when env_map is None (run_doctor's live path). An explicit env_map is a
complete environment description — merging os.environ on top made
report_deprecated_config_and_env non-hermetic on boxes exporting legacy
relay vars (10 findings vs the expected 2 in
test_report_does_not_count_as_blocking_issue).
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
so a hanging codesign degrades to the unreadable-DR warning, never crashing
the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
GPT-OSS review: an empty codesign output would fall through to the
'stable identity' branch and false-positive. Guard with and
cover the empty-string case. Flash review: the non-macOS silence test
mocked the bundle to None, so it never exercised the platform guard;
mock a real path instead.
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.
- hermes doctor: new check_macos_tcc_grants() reports the desktop
bundle's DR class (cdhash-pinned → grants reset on every update;
identifier-pinned → stable) and prints the exact stale-grant repair
(tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
documents the one-time re-grant for pre-fix grants.
Closes#86385
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.
- New config key delegation.surface_child_process_notifications
(default false = suppress). Flag true restores the previous behavior
exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
watch_disabled events whose task_id starts with 'sa-' when the flag
is false, logging at debug with session_id+task_id for diagnosis.
Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
crash the drain loop.
- Docs: delegation.md + configuration.md.
Follow-up to the salvaged #94296: the two guards covered the repair and
confirmed-update branches, but when cua-driver is enabled yet not
installed at all, control still reached _run_cua_driver_installer() and
an automatic 'hermes update' would launch the interactive install.ps1
anyway. Add the same defer before the installer run, keep POSIX
behavior unchanged, and give the confirmed-update message a natural
fallback when latest_version is unknown.
tool_search now takes queries: string[] (searched independently against
the same catalog, limit applies per query, default 5 / max 25) and
returns the split shape: per-query groups carry tool names only, one
shared tools map holds each matched tool's source, description (400-char
cap) and required parameter names once. When some queries miss, a single
top-level available_sources + hint block replaces the old per-response
fallback.
tool_describe now takes names: string[] and returns a map keyed by name;
unknown names collect in not_found (with the refresh hint) and
non-deferrable names keep their per-name spelling-check error in errors,
so one bad name no longer fails the whole call. Duplicates dedupe
silently.
The shared tokenizer now applies Snowball stemming (english, exact-pinned
snowballstemmer) at both index and query time, closing the measured
plural/singular miss where 'issues' failed to return create_issue. The
inline BM25 is unchanged. Stemmer instances are thread-local (they carry
mutable parse state and bridge dispatch can run on parallel tool-call
threads).
New config knobs under tools.tool_search: max_queries / max_describe_names
(default 10 each, floor 1, no upper clamp) bound the per-call array
inputs; over-cap calls error so the model repairs in one round-trip.
No backward compatibility with the single query/name shapes, by decision.
scripts/analyze_livetest.py renders both shapes since transcripts on disk
may predate this change.
Fixes the two live E2E blockers @ctaylor86 found on PR #77189 (macOS 26.3.1,
OpenSSL 3.6.3):
- retry the PKCS#12 export with -legacy when security import rejects the
OpenSSL 3 default format ('MAC verification failed during PKCS12 import')
- trust the self-signed root for the codeSign policy (security
add-trusted-cert -r trustRoot -p codeSign) — an imported-but-untrusted cert
is invisible to find-identity -v and unusable by codesign
- gate success on find-identity -v -p codesigning (postcondition), and use
the same -v probe for idempotency so an untrusted leftover cert is repaired
instead of reported as done
Tests rewritten as stateful fakes (valid only after import+trust), plus new
coverage for the -legacy retry, trust failure, postcondition gate, and the
untrusted-cert repair path; sabotage-verified (reverting to the name-in-output
probe fails 4 tests). Docs: manual fallback now includes the Trust step.
Adds a one-shot `hermes desktop --setup-tcc-identity` command that creates a
self-signed code-signing certificate in the login keychain (openssl +
security import), grants codesign access to it, writes
desktop.macos_signing_identity to config.yaml, and re-signs the packaged app
with a certificate-anchored Designated Requirement.
macOS persists permission grants (Full Disk Access, Accessibility, Files and
Folders, microphone) against the app's code-signing identity, not its path.
The default identifier-pinned ad-hoc signature is stable across rebuilds, but
a certificate-anchored identity is the strongest guarantee — the same
mechanism yabai/skhd rely on. Previously users had to create the certificate
manually in Keychain Access; this command automates the whole flow and is
idempotent (re-run after updates).
Docs: desktop.md TCC section now leads with the command, keeps manual steps.
Tests: 4 new — fresh cert creation path, idempotent reuse, non-macOS no-op,
cmd_gui early-exit before build.
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):
1. The parallel batch planner now peels the tool_call bridge wrapper and
decides admission on the underlying tool — supports_parallel_tool_calls
works again when deferral is active. Unparseable wrappers stay
sequential barriers; bridged calls get exactly the admission the same
call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
service-name queries reach tools whose own name omits the service; the
dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).
Salvaged from #92693 by @alt-glitch with authorship preserved.
On Windows, the pre-update concurrent-instance gate aborted with exit 2
whenever ANY other process held the venv hermes.exe shim — including the
gateway itself, which _pause_windows_gateways_for_update() stops moments
later and the post-update restart phase brings back. Users with a running
gateway were forced into a manual taskkill dance before every update.
The gate now filters gateway runtimes out of the abort list and proceeds
when nothing else is concurrent. Classification delegates to
_is_pausable_gateway -> gateway.status.looks_like_gateway_command_line
(the canonical shlex-tokenized, profile-selector-aware matcher shared by
the Desktop preflight exemption and the venv-holder guard fallback), so
the gate's exemption and the pause machinery cannot drift apart. Anything
not positively identified as a gateway — REPLs, dashboard, Desktop
backend children, gateway MANAGEMENT commands like 'gateway status',
unreadable cmdlines — still aborts exactly as before, and the abort
message now lists only the PIDs that are actually the user's problem.
Surgical reapply of PR #37039 by @damadorPL onto current main (the gate
moved from hermes_cli/main.py to hermes_cli/update_cmd.py in the main.py
decomposition); his substring classifier was replaced with the canonical
matcher, which also fixes the 'hermes gateway status' misclassification
flagged in review.
Co-authored-by: Hermes <hermes@nousresearch.com>
Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.
- _restart_launchd_gateway_after_update() (his extraction, adapted):
plist-exists is the ONLY gate; launchd_restart() owns the
bootout/bootstrap/kickstart ladder for every plist-present state;
every failure path is loud and names the manual recovery command.
The gate-error 'except: pass' (the second silent variant) now counts
the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
supervised PID), composing his fix with the returned-is-not-supervised
guard.
- His regression suite adapted to the (restarted, failed) contract; the
old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
bug.
A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
Merge follow-up on top of the salvaged commits: fold the npm>=10 reduced
hidden-lockfile comparison (intersection of non-null fields) and the
annotation-field exclusions into a single entries_differ() helper, keeping
the workspace-closure scoping intact.
_tui_need_npm_install compared every field of the root package-lock.json
against node_modules/.package-lock.json. npm>=10/11 writes a reduced hidden
lockfile that omits declarative fields (version/dependencies/dev) and adds
extraneous, so nearly every package looked 'changed'; workspace link entries
("link": true, paths outside node_modules/) are never materialized by the
partial --workspace install. Both made the check return True forever, so
hermes --tui re-ran npm install (and dirtied package-lock.json) on every
launch (#84617).
Compare only the keys both sides record with non-null values (resolved,
integrity, ...), ignore workspace link entries and non-node_modules paths in
the missing-entry check, and treat extraneous as an npm runtime annotation.
Real skew (lockfile bumped while node_modules is behind) is still detected.
_tui_need_npm_install() compared ALL fields between
package-lock.json and .package-lock.json, treating npm's
intentionally-stripped metadata (version, license, engines,
dependencies, etc.) as real skew. This caused 'Installing
TUI dependencies…' + rebuild on every launch.
Two changes:
1. Compare only the *intersection* of fields between the
two lockfiles — a field present in the root lock but
absent in npm's hidden actualized lock is a normal npm
artefact, not a real dependency change.
2. Add , , ,
to _NPM_LOCK_RUNTIME_KEYS — these boolean annotations
are written non-deterministically between the two locks
and never indicate a real dependency skew.
Real version/dependency changes still change /
, which are present in both locks and caught by
the intersection comparison.
Fixes: 347 false-positive lockfile mismatches → 0.
On Termux the launch install also selects ui-tui's child packages/*
workspaces (include_child_workspaces=True), so npm installs each child's
devDependencies. The freshness closure only followed devDependencies for
the ui-tui workspace itself, so a devDependency unique to a selected child
was dropped from the closure and a genuine missing package slipped past
_tui_need_npm_install.
Derive the closure from every workspace the install path selects, following
devDependencies for each. _npm_lock_workspace_closure now accepts the set of
selected workspace keys (dev-included roots); _tui_selected_workspace_keys
mirrors _make_tui_argv (ui-tui, plus child packages/* on Termux). Adds a
child-workspace-devDependency regression test (installs on Termux, ignored
off Termux) plus a closure-level dev-scope test.
_tui_need_npm_install compared the full multi-workspace root
package-lock.json against the hidden .package-lock.json, but the launch
install is scoped with npm install --workspace ui-tui and only writes the
ui-tui dependency closure. Every dep belonging solely to another workspace
(apps/desktop, web, ...) was therefore reported as missing, so the check
returned True and printed "Installing TUI dependencies..." on every launch.
Restrict the comparison to the ui-tui workspace's dependency closure,
computed from the root lock's packages map (following npm's node-resolution
walk and workspace symlinks). Standalone / own-lockfile layouts and any
case where the workspace can't be located fall back to the full comparison,
so drift on a genuine ui-tui dependency is still detected.
Fixes#66978
Sites under active development but tested over the public internet
(Vercel previews, ngrok tunnels, staging domains) are public DNS, so
the local-dev never-cache rule can't catch them. web.cache_exempt_hosts
lists hosts whose pages are always fetched live: exact, "*.wildcard",
or domain-suffix matching (label-boundary aware — mysite.dev covers
preview.mysite.dev but never evilmysite.dev). Checked on both store
and lookup, so adding an exemption takes effect immediately even for
entries cached before the config change.
Repeat searches (same normalized query + provider) within a 20-minute
TTL are served from an in-process memo, and concurrent identical
queries are single-flighted so a parallel subagent fan-out pays for
one vendor request instead of N. Requested limits bucket up to
10/20/50/100 so near-identical requests share an entry; callers get
their requested count sliced from the bucket.
Repeat extracts of the same URL are served from the existing
cache/web full-text store (previously written for read_file paging
but never read back), via a small JSON sidecar index. Disk-backed, so
CLI, gateway, cron, and subagents share it. Cached extracts re-run
the normal truncate pipeline, so per-call char_limit still works.
Both caches sit after every safety gate (secret-URL, SSRF, policy,
provider resolution) and directly around the paid vendor call — hits
skip only the network request. Only successful responses cache;
rescue-served responses are never cached (one-shot rescue must stay
one-shot). Config: web.cache_enabled (default on),
web.cache_ttl_minutes (default 20, clamped 1-1440).
Idea credit: query coalescing + num-bucketing pattern observed in
Apodex FrontierAgent (Apache-2.0).
Closes the two config-UI halves of the exclude-mode review (GottZ findings
5-7 on #94513):
- hermes mcp configure: on an exclude-mode server, unchecking a tool now
APPENDS a literal exclude and re-checking drops it — glob patterns are
preserved so future vendor tools keep getting filtered. Previously one
uncheck converted the whole config to a frozen include list (globs
silently deleted, new vendor tools invisible). Re-checked tools still
shadowed by a kept glob get an explicit warning instead of a silent
no-op.
- hermes tools MCP checklist: same exclude-mode write-back, plus display
now matches excludes via matches_name_filter (fnmatch) — glob excludes
previously rendered as if nothing were excluded.
- klaviyo manifest: post_install no longer tells users to append a param
the URL already pins; now documents how to get the FULL surface.
Live-verified through the real cmd_mcp_configure path: exclude-mode server
with ['*_secret_*', 'docs'], uncheck beta + re-check docs ->
exclude becomes ['*_secret_*', 'beta'], no include written.
171 tests green across test_mcp_catalog/test_mcp_config/test_mcp_tool.
The probe-fail path ignored prior state entirely: with no manifest
default_enabled it wrote include=None, which pops the whole tools
block. For exclude-mode manifests default_enabled is necessarily unset,
and for the 30+ OAuth entries the entry rewrite precedes first auth —
so the common reinstall-while-unreachable case removed the curated
excludes and enabled every tool on next connect. The fallback order is
now: prior include > prior exclude > manifest default > no filter.
The "Probing '<name>' for available tools..." line printed before the
exclude-mode short-circuit, which deliberately never probes (the test
suite asserts _probe_tools must not run there). Move the announcement
next to the actual probe call so install output matches what happened.
Review blockers (independent reviewer on #94513):
1. Reinstalling an exclude-mode catalog entry wiped the user's edited
tools.exclude, replacing it with manifest defaults. install_entry now
reads the prior exclude (like it already did for include) and re-writes
it verbatim on reinstall. Regression test added + sabotage-verified
(fails on old behavior); include-priority test added too.
2. aws-knowledge: exclude aws___retrieve_skill — vendor SKILL.md loader is
a vendor skill layer (live tools/list confirmed the tool exists).
3. betterstack: exclude list rewritten to cover the snake_case wire names
(vendor's own header examples show remove_dashboard) via globs alongside
the doc display-labels; caveat documented in the manifest — server is
OAuth-gated so pre-auth enumeration is impossible.
4. railway: exclude railway-agent (opaque server-side agent delegation,
acts outside Hermes's per-tool approval loop).
5. twelve-data: exclude oauth plumbing pseudo-tools + quota probe.
6. betterstack post_install no longer claims a fully-checked checklist —
exclude-mode bypasses the checklist; text now describes the applied
exclude list.
Live E2E: fresh temp HERMES_HOME — install applies manifest excludes,
user edit survives reinstall. 33/33 catalog tests green.
The cloudflare entry's 3,320-endpoint surface is ~43% product families a
personal/dev account never touches (Zero Trust org-fleet suite, Magic
Transit/WAN, Cloudforce One, Radar analytics, API Shield, legacy
migration surfaces). Ship a 34-pattern curated exclude list in the
manifest: 3,320 -> 1,905 tools kept, and everything Cloudflare adds
later stays enabled by default.
Mechanism, two small extensions:
- tools/mcp_tool.py: tools.include/exclude entries containing glob
metacharacters now match via fnmatch (plain names stay exact-match),
so a product family is one pattern instead of hundreds of stale
literals.
- hermes_cli/mcp_catalog.py: manifests may declare
tools.default_excluded (mutually exclusive with default_enabled);
install writes it to tools.exclude and skips the probe/checklist —
a 3,320-row curses checklist is not a UX. Prior user include
selections still win on reinstall.
Verified by replaying the real filter functions over the live-probed
3,320-tool list: 1,415 excluded, zero overmatch against a per-product
target audit; DNS/Workers/R2/D1/tunnels/Access/AI kept.