A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
- The direct adapter.send() confirmation path in /approve and /deny is only
needed on native-streaming platforms (WeCom) where the reply stream is
already finalized; other platforms keep the return-text contract (fixes
4 approve/deny regression tests, guarded with 'is not True' against
MagicMock auto-attributes).
- Remove a stray [DEBUG] logger.info left in _deliver_media_from_response.
- website/docs wecom.md: replace the 'does not stream' notes with the native
msgtype:stream behavior and document the stream keepalive extra keys.
Per the 'when in doubt, optional' rule — site publishing is an
on-request capability, not a weekly daily-driver for most users.
Joins cloudflare-temporary-deploy/page-agent under
optional-skills/web-development (existing category, existing
DESCRIPTION.md kept; the new bundled category dir is dropped).
Install via: hermes skills install official/web-development/publish-site
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
When either unfurl key is set, media captions post as a separate
message before the file (the upload API cannot carry unfurl controls)
and native draft streaming falls back to edit-based delivery. Surface
both side effects in the config reference table.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.
- close_browser_holding_profile: graceful terminate → kill → poll until the
cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
unrelated same-name process on a different dir).
- Config key + docs admonition.
Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.
Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
Live windows-latest proof settled it: a running Chrome opens its cookie DB
deny-all (even CreateFile with FILE_SHARE_READ|WRITE|DELETE fails; sqlite
mode=ro/immutable/nolock all 'unable to open'), so copy-while-running is
impossible on Windows without VSS/admin — and the prior code HUNG ~24min on the
locked file.
Fix: a fast up-front lock probe (_profile_is_locked: one open() of the active
profile's cookie DB; PermissionError = locked) runs BEFORE any copy in
snapshot_real_profile. If locked, bail immediately with 'fully quit the browser
(incl. background/tray) and retry, or turn browser.use_real_profile off'. Never
hangs, never a silent signed-out copy. POSIX has no mandatory locking so the
probe never trips there — copy-while-running still works on macOS/Linux.
Docs: admonition stating Windows needs the browser fully closed (background
apps included); the live-drive-while-running path is #95669.
Tests: lock-probe unit coverage (readable/no-db/PermissionError), snapshot
fails-fast-no-copytree when locked. Windows live E2E asserts the fast-fail
contract (returns <30s with the quit message) + the read-strategy diagnostic.
Addresses the round-3 findings from @Adolanium + @kshitijk4poor on #95620:
1. Overlay-before-reuse race (blocker): _real_profile_cdp ran snapshot_real_profile
BEFORE the session-reuse check, so a cold resolve that ends in reuse rewrote
Cookies/Login Data under a live Chromium holding the user-data-dir open (torn
DBs, locked txns, phantom logouts). Now: resolve copy dir as a PATH, probe
reuse first, return early on a hit; snapshot/overlay only on the relaunch
path when no live browser owns the dir.
2. Torn first copy poisoned freshness forever: freshness keyed on isdir(Default),
so a half-written copy (disk full / Ctrl+C) was treated as populated and only
ever got auth overlays. Now gated on a .hermes-snapshot-complete marker
written only after a full copy succeeds; a torn copy is rebuilt from scratch.
3. Consent revocation left copied credentials on disk: turning use_real_profile
off now deletes ~/.hermes/browser-profile/ on next browser use
(cleanup_real_profile_snapshots), so cookies/logins don't outlive consent.
4. Stale non-active profile copies: only the ACTIVE profile (last_used) is copied
into the copy's Default now — other Chrome profiles are never snapshotted
(smaller copy, no stale credential dirs lingering).
5. Docs/config/desktop wording aligned to actual behavior (active-profile only,
refresh on fresh session, consent-off cleanup).
Tests: overlay-skipped-on-reuse + overlay-runs-on-relaunch, done-marker gating +
torn-copy rebuild, active-only copy, consent-off cleanup (removes store +
idempotent + triggered from _real_profile_cdp). 206 browser + 222
backup/file_safety pass. Live: reuse skips re-snapshot; direct launch on the
active-only copy loads the real signed-in Gmail inbox.
Copy the user's default-Chromium profile (auth state only) into a managed
snapshot, launch Hermes' packaged Chromium on it via agent-browser, and hand
the CDP endpoint to the Browser Use CLI (and built-in tools) to drive. The
snapshot is a non-default dir, so it sidesteps Chrome 136+'s default-profile
remote-debugging block and never contends with the user's running browser;
launched without mock-keychain switches so keyring-encrypted cookies decrypt.
- consent-gated browser_exec 'local' arg (schema only appears with consent)
- fail-closed on non-Chromium default / snapshot failure
- stale-session guard: reuse only when the live session is on our copy dir
- snapshot excludes extensions/service-workers (renderer wedge) + caches
Salvage adjustments to PR #94392 per review:
- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
recovery child now probes 'systemctl --user is-active' after each relaunch;
only an observed-active systemd unit is reported 'verified'. A relaunch that
merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
update_inventory serve collector) are no longer silently skipped: the
recovery pass records them (and manual gateways) as skipped-with-reason in
the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
(requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
fresh interpreter (sitecustomize shim intercepts the grandchild
'gateway restart' and systemctl probes).
`completed` already puts the worker's summary inside the synthetic wake
turn, so the woken creator sees what was done. `review_requested` did
not: the summary rode the passive ping only, and the wake turn said just
"handed off for review", forcing the woken reviewer to re-read the board
(and losing the PR link the worker had already written).
Reuse the same first-line handoff the `completed` branch builds, so the
existing `gateway.kanban.wake.handoff` string renders it — no new locale
keys, no change to the passive message.
`review_requested` and `block_loop_detected` are terminal event kinds that
hand a decision back to the origin subscriber, but neither was listed in the
gateway notifier's `_WAKE_KINDS`. A `notify+wake` subscription therefore got
the passive ping only and the origin agent never took a turn — so an agent
that delegated implementation work slept through the "ready for review"
handoff and through a task being routed to triage, while the equivalent
`blocked` event woke it.
Add both kinds to the wake set, add their status strings to the synthetic
wake message in every locale, and document which events wake.
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.
- Every registered gateway's profiles sit on the one strip, in registry order
(This device first, then by label), each group headed by that gateway's
kind glyph. The active gateway's squares are unchanged; the others are
"at rest" (dimmed) with tooltips/accessible names qualified by machine
(`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
the statusbar switcher, landing on that exact (gateway, profile):
`selectConnection(id, { profile })`. The spinner sits on the clicked
square; the previous source stays painted until the target answers.
Groups keep their slots whichever gateway is active, so a square never
moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
SOUL.md / Delete, executed on the owning gateway (renameProfile,
getProfileSoul and updateProfileSoul accept the same scope deleteProfile
already had); the delete confirmation names the machine. The legacy
per-profile "Connect to a remote host…" item is hidden on multi-gateway
setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
two registrations of one backend collapse to one group; past thirteen
squares across the fleet the strip condenses into a menu sectioned by
gateway. Roster is fetched on mount / focus / registry change only — no
periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.
Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.
Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.
Docs: multi-connection-desktop.md describes the fleet rail.
Refs #89304, #92384, #91047, #94724
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.
Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.
Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).
Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).
E2E counterfactual (real imports, 1M window @ 0.85):
main default: legacy, tail 170,000 (ceiling 255,000)
head default: lean, tail 25,000 (ceiling 37,500)
head legacy: 170,000 (opt-out intact)
update_model to 400K: 10,000 (lean preserved across switch)
The last piece of the macOS permissions campaign (#52010 follow-up): macOS
prompts per-folder (Desktop, then Downloads, then Documents, ...) as the
agent touches each one — a drip-feed of dialogs on first use. ONE Full Disk
Access grant covers all of them permanently, and with the stable signing
identities merged this week it survives every update. Nothing in Hermes
taught users that.
- hermes doctor: check_macos_full_disk_access() — prompt-free probe (the
FDA-gated TCC db dir returns EPERM without a dialog; TCC only prompts on
protected-CATEGORY paths), reports granted state or prints the one-switch
setup with the Privacy_AllFiles deep link.
- hermes setup: same probe at the end of onboarding — the moment users are
primed to do system setup — silent when already granted, indeterminate,
or non-macOS.
- docs: desktop.md TCC section now leads with the one-switch guidance.
- 7 tests (granted / denied / indeterminate / non-macOS, both surfaces).
The feature page still said 'only the origin chat is ever touched', which
this change makes stale. Documents the three attach-eligible shapes and
that the global flag never activates explicit targets (review nit).
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.
- hermes doctor: new check_macos_tcc_grants() reports the desktop
bundle's DR class (cdhash-pinned → grants reset on every update;
identifier-pinned → stable) and prints the exact stale-grant repair
(tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
documents the one-time re-grant for pre-fix grants.
Closes#86385
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.
- New config key delegation.surface_child_process_notifications
(default false = suppress). Flag true restores the previous behavior
exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
watch_disabled events whose task_id starts with 'sa-' when the flag
is false, logging at debug with session_id+task_id for diagnosis.
Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
crash the drain loop.
- Docs: delegation.md + configuration.md.
tool_search now takes queries: string[] (searched independently against
the same catalog, limit applies per query, default 5 / max 25) and
returns the split shape: per-query groups carry tool names only, one
shared tools map holds each matched tool's source, description (400-char
cap) and required parameter names once. When some queries miss, a single
top-level available_sources + hint block replaces the old per-response
fallback.
tool_describe now takes names: string[] and returns a map keyed by name;
unknown names collect in not_found (with the refresh hint) and
non-deferrable names keep their per-name spelling-check error in errors,
so one bad name no longer fails the whole call. Duplicates dedupe
silently.
The shared tokenizer now applies Snowball stemming (english, exact-pinned
snowballstemmer) at both index and query time, closing the measured
plural/singular miss where 'issues' failed to return create_issue. The
inline BM25 is unchanged. Stemmer instances are thread-local (they carry
mutable parse state and bridge dispatch can run on parallel tool-call
threads).
New config knobs under tools.tool_search: max_queries / max_describe_names
(default 10 each, floor 1, no upper clamp) bound the per-call array
inputs; over-cap calls error so the model repairs in one round-trip.
No backward compatibility with the single query/name shapes, by decision.
scripts/analyze_livetest.py renders both shapes since transcripts on disk
may predate this change.
Fixes the two live E2E blockers @ctaylor86 found on PR #77189 (macOS 26.3.1,
OpenSSL 3.6.3):
- retry the PKCS#12 export with -legacy when security import rejects the
OpenSSL 3 default format ('MAC verification failed during PKCS12 import')
- trust the self-signed root for the codeSign policy (security
add-trusted-cert -r trustRoot -p codeSign) — an imported-but-untrusted cert
is invisible to find-identity -v and unusable by codesign
- gate success on find-identity -v -p codesigning (postcondition), and use
the same -v probe for idempotency so an untrusted leftover cert is repaired
instead of reported as done
Tests rewritten as stateful fakes (valid only after import+trust), plus new
coverage for the -legacy retry, trust failure, postcondition gate, and the
untrusted-cert repair path; sabotage-verified (reverting to the name-in-output
probe fails 4 tests). Docs: manual fallback now includes the Trust step.
Adds a one-shot `hermes desktop --setup-tcc-identity` command that creates a
self-signed code-signing certificate in the login keychain (openssl +
security import), grants codesign access to it, writes
desktop.macos_signing_identity to config.yaml, and re-signs the packaged app
with a certificate-anchored Designated Requirement.
macOS persists permission grants (Full Disk Access, Accessibility, Files and
Folders, microphone) against the app's code-signing identity, not its path.
The default identifier-pinned ad-hoc signature is stable across rebuilds, but
a certificate-anchored identity is the strongest guarantee — the same
mechanism yabai/skhd rely on. Previously users had to create the certificate
manually in Keychain Access; this command automates the whole flow and is
idempotent (re-run after updates).
Docs: desktop.md TCC section now leads with the command, keeps manual steps.
Tests: 4 new — fresh cert creation path, idempotent reuse, non-macOS no-op,
cmd_gui early-exit before build.
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):
1. The parallel batch planner now peels the tool_call bridge wrapper and
decides admission on the underlying tool — supports_parallel_tool_calls
works again when deferral is active. Unparseable wrappers stay
sequential barriers; bridged calls get exactly the admission the same
call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
service-name queries reach tools whose own name omits the service; the
dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).
Salvaged from #92693 by @alt-glitch with authorship preserved.
Electron safeStorage parks a per-app key ('Hermes Key') in the macOS login
keychain; on machines with a locked/missing/corrupted default keychain that
turned every Hermes Desktop launch into a blocking 'Keychain Not Found' /
password dialog. Keychain-backed encryption is now an explicit opt-in:
- electron/secret-storage-policy.ts: standalone policy seam (default OFF,
strict === true coercion, one-shot migration flag) + unit tests
- default path never calls any safeStorage API (including
isEncryptionAvailable, which itself touches the keychain)
- one-shot legacy migration decrypts existing safeStorage blobs to plain
0600 files at first launch; undecryptable blobs are kept but read as
absent afterward (classify 'drop') so a dead keychain prompts at most once
- Settings -> Gateway toggle (all 5 locales) re-encodes every stored secret
store in place when flipped (v1 connection.json, v2 connections.json,
native-oauth-tokens.json)
- e2e: at-rest spec now covers both postures (opted-in unchanged contract,
default saves without secure storage, owner-only bits, restart round-trip)
- docs: multi-connection-desktop + desktop-native-signin updated
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
Sites under active development but tested over the public internet
(Vercel previews, ngrok tunnels, staging domains) are public DNS, so
the local-dev never-cache rule can't catch them. web.cache_exempt_hosts
lists hosts whose pages are always fetched live: exact, "*.wildcard",
or domain-suffix matching (label-boundary aware — mysite.dev covers
preview.mysite.dev but never evilmysite.dev). Checked on both store
and lookup, so adding an exemption takes effect immediately even for
entries cached before the config change.
Dev servers, hot-reload builds, and chat-GUI artifact previews live on
localhost/private addresses and change on every save — a 20-minute
cached copy would show a stale build exactly when freshness is the
point of fetching. The extract cache now declines loopback, private,
link-local, *.local, *.localhost, and single-label LAN hostnames on
both put and get. Hostname heuristics only (no DNS) — this is a
freshness carveout, not a security boundary; SSRF enforcement is
unchanged in tools/url_safety.py.
Public URLs keep the full TTL.
Review fixes for #94618 (all three blockers reproduced by the reviewer
through the real web_extract_tool):
1. Cache lookup moved AFTER provider resolution and strict-selection
validation, and gated per-URL on the website blocklist policy — a
blocklist-blocked or misconfigured-backend call now behaves exactly
as it would without a cache instead of serving cached content.
2. Rescue-served extract batches are never cached (mirrors the search
memo's exclusion), keeping one-shot rescue one-shot.
3. Cache entries now get dedicated per-(url, format, provider) files
instead of sharing the URL-keyed truncate-store file — html and
markdown (or two backends') copies of one URL no longer overwrite
each other, and switching extract backends within the TTL never
serves the old backend's rendering.
Also from review: per-process index tmp filename (cross-process writers
can no longer truncate each other mid-write) and held flight locks are
never evicted from the bounded lock table (eviction could have allowed
a duplicate paid request).
New regression tests for formats/provider keying; E2E harness extended
with policy-block, strict-selection, rescue-two-call, and dual-format
scenarios — 6/6 pass; original 13/13 still pass.
Repeat searches (same normalized query + provider) within a 20-minute
TTL are served from an in-process memo, and concurrent identical
queries are single-flighted so a parallel subagent fan-out pays for
one vendor request instead of N. Requested limits bucket up to
10/20/50/100 so near-identical requests share an entry; callers get
their requested count sliced from the bucket.
Repeat extracts of the same URL are served from the existing
cache/web full-text store (previously written for read_file paging
but never read back), via a small JSON sidecar index. Disk-backed, so
CLI, gateway, cron, and subagents share it. Cached extracts re-run
the normal truncate pipeline, so per-call char_limit still works.
Both caches sit after every safety gate (secret-URL, SSRF, policy,
provider resolution) and directly around the paid vendor call — hits
skip only the network request. Only successful responses cache;
rescue-served responses are never cached (one-shot rescue must stay
one-shot). Config: web.cache_enabled (default on),
web.cache_ttl_minutes (default 20, clamped 1-1440).
Idea credit: query coalescing + num-bucketing pattern observed in
Apodex FrontierAgent (Apache-2.0).
The cloudflare entry's 3,320-endpoint surface is ~43% product families a
personal/dev account never touches (Zero Trust org-fleet suite, Magic
Transit/WAN, Cloudforce One, Radar analytics, API Shield, legacy
migration surfaces). Ship a 34-pattern curated exclude list in the
manifest: 3,320 -> 1,905 tools kept, and everything Cloudflare adds
later stays enabled by default.
Mechanism, two small extensions:
- tools/mcp_tool.py: tools.include/exclude entries containing glob
metacharacters now match via fnmatch (plain names stay exact-match),
so a product family is one pattern instead of hundreds of stale
literals.
- hermes_cli/mcp_catalog.py: manifests may declare
tools.default_excluded (mutually exclusive with default_enabled);
install writes it to tools.exclude and skips the probe/checklist —
a 3,320-row curses checklist is not a UX. Prior user include
selections still win on reinstall.
Verified by replaying the real filter functions over the live-probed
3,320-tool list: 1,415 excluded, zero overmatch against a per-product
target audit; DNS/Workers/R2/D1/tunnels/Access/AI kept.
Adds optional-mcps/cloudflare — Cloudflare's managed remote MCP server
(mcp.cloudflare.com/mcp) fronting the entire Cloudflare API (2,500+
endpoints across DNS, Workers, R2, KV, D1, Zero Trust, WAF, Pages)
through two Code Mode tools, search() and execute(), at a fixed ~1k-token
schema footprint. HTTP transport + native MCP OAuth 2.1 with DCR — no
install block, nothing to pin. post_install documents the scoped OAuth
grant, the bearer-token path for headless/CI, and Cloudflare's
product-specific servers for narrower surfaces.
Docs: mention Cloudflare in the hosted-OAuth MCP examples.