The early-EOF reaper made _finish_reader return without publishing when
wait() raises, so the session stays tracked for later reconciliation.
That is right for the pipe path (_reconcile_local_exit can still reap via
session.process), but PTY sessions have no session.process: when
ptyprocess.wait raises (waitpid ECHILD after isalive() already reaped the
child) the exitstatus is known, yet poll() reported "running" forever.
Only leave the session tracked when exit_code() is still None; otherwise
record the known status and finish as before.
Review finding: PTY session whose pty.wait raises stays in _running forever (fail-open regression vs main)
The dynamic-shell-word rules fired on any `-del*`/`-exec*` substring after a
`find` anywhere in the segment, so quoted predicate arguments the shell never
expands were flagged: `find . -name 'log-del*'`, `find . -name 'pre-exec*.sh'`,
`find src -path '*-exec[0-9]*'`, and `echo find . -{delete,print}`.
Three changes close that:
- `find` must be the command word (_CMDPOS anchored, same as mkfs/rm/dd) and
the dynamic word must start a whitespace-delimited token (`(?<!\S)`), for
both the find rule and the rg/sort/ag/man program-option rule.
- Both rules now scan the quote-masked variant (_QUOTE_MASKED_DANGEROUS_
DESCRIPTIONS, same _mask_quoted_prose used by the positionless hardline
rules) so glob characters inside quotes are data, not expansion.
- _iter_shell_command_starts no longer treats the `{` inside a brace-expansion
word (`-{delete,print}`) as a brace-group opener; it split the word across a
marked start so `echo x; find . -{delete,print}` matched nothing. A brace
group opener is `{` as its own word (after whitespace or a separator).
The quoted-name cases join the inert parametrized test; the separator case
joins the dangerous one. Still two parametrized functions.
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:
- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
raw-string check (never resolves — resolving IS the leak trigger).
Wired as the first check in get_read_block_error() and the write
denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
_check_sensitive_path (covers write_file_tool + patch_tool), before
the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
INLINE_TOOL_EXECUTORS["session_search"] is the production dispatch on every surface
(tool_executor + agent_runtime_helpers) and maps schema args to kwargs explicitly, so the
schema advertised after/before/exclude_session_ids while the executor silently dropped them
and every time-bounded call ran unbounded. Add the three mappings.
Also drop the stranded `as_exclusive_end` parameter on _parse_iso_bound (accepted, never
read; exclusivity is the SQL `<` predicate) and the Python re-check of the time window over
FTS rows in _discover — _search_filter_clauses already bounds sessions.started_at on every
route, so the only place the window still needs a Python check is the title-match branch,
which bypasses that query.
Amp's thread feed supports relative time filters (`after:7d`,
`updated_before:7d`) alongside ISO dates. Extend the salvaged
after/before bounds (PR #86067 by @Moodtuner997) the same way:
- `_parse_iso_bound()` now accepts relative durations `Nh`/`Nd`/`Nw`
(case-insensitive) meaning "now minus N", alongside ISO
dates/datetimes. Clearer error message names both accepted forms.
- Forward after/before/exclude_session_ids through the public
`session_search()` wrapper (the PR predates the wrapper/impl split;
without this the SQL bounds were unreachable from the registry
handler — same class as the earlier `detail` forwarding fix).
Appended after `detail` to preserve positional compatibility.
- Tool schema descriptions teach both forms.
- Tests: relative after/before against the discovery shape, unit
checks for h/d/w math, case-insensitivity, and bad-unit rejection.
- Docs: tools-reference row mentions time bounds + exclude_session_ids.
Push the session-start bounds into search_messages so FTS LIMIT
cannot be filled by out-of-range hits. Covers FTS5, CJK, trigram,
LIKE fallback, and the unindexed-gap supplement.
Refs #86021.
Discovery-only filters for issue #86021. sort remains a ranking
bias. Date-only before is an exclusive midnight UTC bound.
exclude_session_ids drops the named session and its lineage (cap 20).
The stale-overwrite refusal made write_file permanently unusable for any
existing file it could not show in one read_file page: every >2000-line
(or >100K-char) page was recorded as partial, no full baseline ever
existed, and the refusal told the model to "re-read the whole file", which
the tool cannot do. Track the line ranges each task pages through per path
at one mtime; contiguous pages from line 1 to total_lines are a full read
(a new mtime between pages restarts the coverage). The same gap hit two
siblings: the extracted-document branch (.ipynb, text-authorable) returned
before any read bookkeeping, so an existing notebook could never be
overwritten; and reset_file_dedup dropped every baseline on compaction
while keeping read_timestamps, so every write after compaction was refused
even for files unchanged on disk. Baselines now survive compaction exactly
like the dedup mtime map does — only while the recorded mtime still matches.
Refusal texts no longer embed the pre-PR "Warning: … Consider re-reading"
copy inside "Refusing to overwrite", and every refusal names a recovery the
model can perform: read the remaining pages, or use patch.
Require an explicit full-file baseline before replacing existing host-visible files with write_file, and fail closed when that baseline is stale. This prevents stale conversation context from clobbering manual or external edits.\n\nRefs #65604
Widening the _presence() clearing from single-query to every unattended
context also cleared is_ask for platform=api_server. That surface answers
approvals through the /v1/runs bridge (approval.request ->
POST /v1/runs/{id}/approval), so a dangerous command that used to park in
waiting_for_approval became an instant BLOCK with no approval.request.
Restrict the clearing to single-query + cron, where nobody can answer.
Stripping HERMES_INTERACTIVE/HERMES_GATEWAY_SESSION/HERMES_EXEC_ASK from
the external worker env also made check_cronjob_requirements() False, so
the cronjob toolset vanished for every job on a managed-systemd gateway
even with cron.allow_agent_scheduling: true. Accept the existing
HERMES_CRON_SESSION marker (set by run_one_job's context) as well.
Review finding: _presence() over-widening broke the /v1/runs approval bridge; env strip hid the cronjob toolset in external workers.
Widen the cron-only clearing to `_unattended_contexts()`: a webhook /
api_server session running inside a gateway inherits HERMES_EXEC_ASK=1
exactly like an external cron worker does, and `_presence()` returning
is_ask=True sent it to the gateway-decision branch with no notifier — a
pending card nobody can answer — instead of `approvals.unattended_mode`.
Same class as #110932, one predicate.
Test trimmed to two invariants (cron / webhook leak → cleared; interactive
keeps presence); the launch-path comment in cron/scheduler.py names the
env-fallback consumers instead of an internal incident log.
_presence() cleared is_cli/is_gateway/is_ask for single-query sessions but
not for cron, so a cron worker that inherited HERMES_INTERACTIVE /
HERMES_EXEC_ASK from its launching gateway resolved as an interactive CLI
and blocked on an approval card nobody could answer (measured: 31
pending_approval hangs/hour, 6 stranded claims — #110932).
Mirror the single-query clearing for _is_cron_approval_context(), matching
the cron exclusion already inside _is_gateway_approval_context(). Layer 1
(#110942) strips the vars at the launch path; this makes the gate robust
to any other leak route.
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.
normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).
Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
Port from Kilo-Org/kilocode#13224: fixed waits belong in foreground terminal calls, while background mode is reserved for independently running processes.
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:
- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.
tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.
Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
`hermes serve` / the Desktop backend hosted many profile homes (session
profile_home, the `profile` RPC param, hosted rooms) but never called
agent.secret_scope.set_multiplex_active, so every unscoped get_secret read for
a secondary silently returned the LAUNCH profile's os.environ value, and
@_profile_scoped bound only HERMES_HOME: `config.get full` for profile B
expanded B's `${VAR}` refs to the default profile's plaintext credentials,
model.options listed the default's env-keyed providers, llm.oneshot billed the
default's auxiliary key.
- tui_gateway/launch_profile_policy.py (was launch_terminal_policy.py): the
first time _profile_home registers a non-launch home the process freezes
the launch env and flips get_secret to fail closed
(activate_multi_profile_hosting); launch_secret_scope composes the launch
profile's .env + external sources over that frozen env so systemd / op-run
injection survives the flip while a secondary never sees it.
- model_switch._profile_runtime_scope_tokens is the ONE composer for
home + secret + terminal scope: a named profile binds its own files; the
launch profile binds its frozen-env scope once multiplexing is active and
stays unscoped in a single-profile process (legacy os.environ precedence).
_profile_scoped, _profile_scoped_rpc, _session_profile_runtime_scope,
_bind_build_profile_scopes and _prepare_turn_input all go through it.
- Hosted-room / Group Chat turns for a DEFAULT-profile member in a
`multiplex_profiles: true` gateway no longer die at agent build with
UnscopedSecretError: `profile_home is None` was treated as "no scope"
in _start_agent_build._build and _prepare_turn_input.
- llm.oneshot runs under the session's (or params.profile's) scope;
_lap_builtin_rows / _overlay_has_creds / _provider_has_credentials read
provider keys through _scoped_key_env instead of raw os.environ;
methods_groups._profile_execution_policy resolves the hosted-room policy
(which reads provider credentials) under the profile's full scope.
Live repro (real `hermes serve`, two homes, config.get {key: full, profile: b}):
base a_ref: <A_VALUE> b_ref: ${B_ONLY_TOKEN} env_ref: <ENV_INJECTED>
head a_ref: ${A_ONLY_TOKEN} b_ref: <B_VALUE> env_ref: ${ENV_INJECTED_TOKEN}
Control (one home, --single): launch config still resolves env_ref from os.environ.
When a child's final answer still missed its output_schema after the one
bounded retry, the result entry flipped to status=failed with the error
"Final answer does not satisfy the declared output_schema" — the completion
line printed ✗ and orchestrators read a finished audit as a failure. Five
audits of 413-4103 s were lost this way in the Sep 10-14 retrospective and
the parent had to mine the live transcripts; in four of them the "violation"
was a ```json fence around a valid array, which the candidate extractor
sliced to its first..last object.
Now: status stays completed, `summary` is the child's raw final text,
`schema_valid: false` + `schema_errors` carry the verdict and a `schema_note`
says the text is unvalidated; the sync completion line shows ⚠ with the
reason. The extractor tries the earliest-opening bracket span and keeps the
first that parses (fenced arrays validate). The OUTPUT CONTRACT the child
sees now says "ONLY the JSON value — no prose, no code fence" and what a miss
costs. One bounded retry is unchanged.
A delegated child's execute_code kernel was keyed correctly
(<owner>::child::<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.
A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
_handle_session_expired_and_retry only reached the at-most-once guard when a
reconnectable server record existed; without one (server torn down, MCP loop
not running) a write-capable call fell through to the generic "MCP call
failed" error, which invites the model to replay a write that may already
have landed. The session-expired classification now runs first and a
write-capable call always gets the outcome_uncertain error; the reconnect is
attempted only when a server can be signalled.
_track_inflight_rpc's teardown RuntimeError said "retry the request on the
rebuilt session" for every op; for a write-capable tools/call it now says the
request may already have been dispatched and must be verified first, so the
wording matches the at-most-once contract the recoverer enforces.
Docs: the readOnlyHint row explains that the same hint gates auto-retry after
a mid-call session expiry, and that unannotated tools on an idle-TTL
Streamable-HTTP server return outcome_uncertain on the first call after idle
instead of being transparently replayed.
A 'session expired' / transport-closed failure can arrive AFTER the server
already accepted and executed the request (proxy-synthesized 404s, pod
rotation, ClosedResourceError firing mid-response). Auto-retrying a
write-capable tool in that window risks a duplicate side effect that MCP
offers no way to undo.
The session-expired recovery path now consults the discovery-time
readOnlyHint capture (same data the trust gate uses): only tools whose
annotation is exactly True keep the reconnect+retry-once behavior. Write-
capable calls still get the transport healed (reconnect, breaker reset on
success) but return a structured outcome_unknown error telling the model
to verify with a read before re-invoking.
The OAuth 401 path keeps its retry for all tools: a 401 means the server
demanded authorization before dispatch, so the call never executed —
matching the upstream classifier's McpAuthRequiredError => safe rule.
Fails safe: missing/unknown annotations classify as write-capable.
The PR added four test functions for one fix. The host-side short-write and
multibyte cases are the same invariant as the sandbox size probe (an archive
that is not byte-exact is never referenced to the model), so they become two
parametrized rows of test_size_probe_decides_lossless, which now also uses
multibyte content so every row pins the byte-vs-char comparison. Net new
tests for the fix: 2 functions.
Also names in _write_to_sandbox why the +1 tolerance is keyed on heredoc
mode only: the payload backend delivers stdin verbatim and is expected to be
byte-exact. The three bare write_text() calls the footgun scanner flags in
this file gain encoding= while it is being touched.
Five near-identical MagicMock scripts pinned the same contract (byte-exact
or discard, heredoc gets exactly one extra byte); one parametrized test
plus the unprobeable-backend case cover it. Drop the host-side happy-path
test that duplicated test_multibyte_content_verified_by_byte_count.
Docstring now names the real heredoc wrapper
(BaseEnvironment._embed_stdin_heredoc) and notes the extra exec RTT per
oversized result (review point).
Oversized tool results are archived to disk and replaced in-context with a
'Full output saved to: <path>' reference. Until now the write was trusted
blind: a partially-flushed host file (ENOSPC/quota races) or a lossy sandbox
write (API-body truncation on payload backends) still produced the archive
reference, so the model was told the full result was recoverable when bytes
had silently vanished.
Both persistence paths now round-trip-verify size before building the
reference and fail closed to the bounded inline truncation otherwise:
- _write_to_spillover: byte-count check via os.stat after write; mismatched
archives are deleted and the caller falls through to inline truncation.
- _write_to_sandbox: wc -c probe after the cat; heredoc-mode backends get a
+1 byte tolerance (wrap_modal_stdin_heredoc appends one newline by
construction), unprobeable backends stay best-effort success.
Regression tests fail without the fix (verified by stashing the source
change: 4 failed). E2E-verified against a temp HERMES_HOME with real file
I/O including multibyte content and a simulated short write.
The dashboard endpoint GET /api/skills/hub/search passes its user-supplied
`source` straight into parallel_search_sources and never applied the merged
provider cut, so ?source=nvidia returned a mixed set. That was the fourth
caller of the walker; the cut was copy-pasted at three of them and missing
at the fourth.
parallel_search_sources already computes the normalized provider filter, so
the cut now lives there — applied per source before results are counted and
merged. Every caller (CLI search via unified_search, CLI browse, TUI-gateway
browse, dashboard router) sees the same rule with no provider logic of its
own, source_counts stop reporting rows that are then dropped, and the three
duplicated call-site cuts are deleted. do_browse keeps its provider-specific
"No skills found for provider" message.
Also:
- HermesIndexSource.search now treats a whitespace-only provider_filter as
"no filter", matching GitHubSource.search (the two adapters previously
disagreed on the same keyword argument; unreachable through the walker,
which pre-normalizes).
- The regression-test fixture seeds tap caches by github_provider_for label
instead of case-sensitive repo literals, and serializes metas through
_skill_meta_to_dict, so a DEFAULT_TAPS casing change can no longer silently
unseed the fixture.
Validation: 121 targeted tests green; disabling the walker cut fails the
pre-existing test_unified_search_provider_filter_keeps_index_source with the
expected clawhub leak; 4/4 regression cases still red on unpatched main.
Follow-up polish on the provider-filter-before-limit fix:
- GitHubSource.search now skips taps whose repo maps to a different provider
instead of enumerating every tap and filtering afterwards. A tap's repo fixes
the provider of every result it yields (github_provider_for is the only source
of extra.provider in this adapter), so the skip is lossless and avoids up to 23
useless tap enumerations per provider-filtered search — real GitHub API calls
against the 60/hr unauthenticated budget and the 30s overall timeout when the
index is unavailable. The now-redundant post-loop filter is dropped.
- _provider_filter_of() is the single owner of "does --source name a provider";
it replaces the four inline copies of the strip/lower/membership idiom in
_select_active_sources, parallel_search_sources, unified_search and do_browse.
- _tap_cache_key() is shared by _list_skills_in_repo and the regression test so
the seeded tap cache can never drift from the production key format.
- _entry_provider() dedupes the raw-index provider extraction used by both the
pre-ranking filter and the scoring loop in HermesIndexSource.search.
- browse_skills (the TUI-gateway browse path) now applies the same merged
provider cut as do_browse; it accepted a provider value but returned
unfiltered results.
Validation: 121 targeted tests green; the regression tests go red on both
adapters when either the tap skip or the index pre-filter is neutralized, and
4/4 red on unpatched main; live CLI repro returns 0 results on main and 3/3
provider matches on this stack.
The gateway loopback handler re-inlined tools.mcp_oauth._parse_redirect_query;
that copy is exactly how the gateway relay lost `iss` while the CLI path
kept it. Use the helper so the four callback keys have one owner, and
point the docstring at it instead of repeating the RFC 9207 rationale.
mcp 2.x rejects an authorization response that omits the RFC 9207 `iss`
parameter when the authorization server advertised
`authorization_response_iss_parameter_supported`. Cloudflare advertises it
AND sends it; the CLI loopback handler has always forwarded it, but every
other callback producer parsed only code/state/error, so the SDK raised:
OAuthFlowError: Authorization response missing iss parameter
advertised by the authorization server
and the server parked. Same machine, same config, `hermes mcp login <name>`
from a terminal succeeded — the failure is specific to the non-CLI relays.
Forward `iss` on every producer, matching `_make_callback_handler()`:
- tools/mcp_dashboard_oauth.py: `deliver_callback()` accepts `iss`;
`wait_for_callback()` returns `(code, state, iss)`. The bridge in
tools/mcp_oauth.py already splats that tuple into
`_authorization_code_result(code, state, iss)`, so it needs no change.
- tui_gateway/mcp_oauth_sessions.py: the gateway-hosted loopback listener
parses `iss`, and `deliver_callback_flow()` forwards it.
- tui_gateway/methods_tools.py: the `oauth.callback` RPC passes `iss`.
- hermes_cli/web_routers/mcp.py: the dashboard callback route accepts it.
- apps/desktop/electron/mcp-oauth-callback-ipc.ts: the one-shot listener
reads `iss` off the redirect (the renderer already spreads the whole
callback object into the RPC, so it flows through unchanged).
Providers that omit `iss` round-trip as `None`/`null` rather than being
dropped, so servers that do not advertise RFC 9207 keep working.
Verified live on Windows against mcp.cloudflare.com, whose metadata sets
`authorization_response_iss_parameter_supported: true`: the server that
previously parked on the missing-iss error now reports
`Authenticated — 3452 tool(s) available` and `hermes mcp test cloudflare`
connects. State-mismatch and replay rejection are unchanged.
Tests (each fails on base, passes with the fix):
- test_dashboard_flow_preserves_rfc9207_iss
- test_deliver_callback_forwards_iss (client-redirect relay)
- test_loopback_listener_forwards_iss (real HTTP redirect)
- two vitest cases on the Electron listener, incl. the iss-absent case
Refs #92758, #99984. PR #92765 fixes the dashboard route and the loopback
listener but not the client-redirect relay
(`deliver_callback_flow` / `oauth.callback` / the Electron listener), which
is the path Desktop drives against a remote backend.
Widening secret_source_names() to include skipped_existing names silently changed
tools/mcp_tool_config.py::_build_safe_env, an untouched consumer that forwards every
returned name into MCP stdio child envs. That consumer wants only names a source
actually APPLIED (pre-stack semantics), so secret_source_names() goes back to
tuple(_SECRET_SOURCES).
The routed-child scrub in strip_launch_profile_env is the one site that must also see
names a source supplied but lost to a pre-existing process value, so it reads the new
source_supplied_names() accessor instead. tools/mcp_tool_config.py is byte-identical to
origin/main.
_run_job_script popped the launch profile's secret-source names in an
inline loop right above strip_launch_profile_env, so only the no_agent
child got that protection; the four other callers of the same helper
(the external cron worker, scheduler_delivery, kanban dispatch and the
byterover plugin) still handed a served profile the launch vault or
1Password names. Fold the names into the helper's residue set, which is
already gated on multiplex and on the target not being the launch
profile, and keeps administrator-managed keys.
Review findings on f5f88d5058. Three are defects the previous round introduced.
Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.
Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.
Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.
Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.
Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.
(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
Review findings on d8c467f223, each reproduced through its production path.
Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.
Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.
Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).
Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.
(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
The 400-recovery reload installs a disk pair and only then re-runs issuer
binding on it. When that pair was minted by a different issuer the enforcer
strips its refresh token, and the reload must report "no recovery" so the
session is cleared like any other dead grant. No test pinned that verdict:
a reload that ignored the install result would keep a stripped pair in the
context and return True. The new invariant drives a real 400 through
_handle_refresh_response against a foreign-bound disk pair and asserts the
result is False, the context is cleared, and the foreign refresh token does
not survive on disk.
_hermes_live_ttl read expires_in off current_tokens without checking for
None; getattr(None, ...) happened to yield the default and report "live".
Both current callers install a pair first, but the helper is now explicit
that an empty context is never live, so a future call site cannot adopt
nothing.
Two defects in the peer-adoption path of the refresh fence:
1. `_hermes_install_disk_pair` raised `_RefreshCompletedByPeer` when issuer
binding stripped the candidate's refresh token. That is the right outcome
for `_refresh_token` (restart the flow so the SDK lands in 401 -> full
auth), but `_hermes_reload_tokens_after_refresh_failure` shares the helper
and must instead treat the candidate as rejected: restore the previous
pair and return False so the caller clears state and prompts. The helper
now returns whether a refresh token survived binding and each caller
decides; the one-shot adopt wrapper is inlined into `_refresh_token`.
2. `_hermes_live_ttl` treated `expires_in is None` as expired. RFC 6749 makes
`expires_in` optional, the SDK's `is_token_valid()` is True with no expiry,
and `_rebase_expires_in` preserves None on read, so a peer's rotated pair
without an expiry was never adopted and we POSTed its refresh token
anyway, burning a generation on single-use providers. None now counts as
live; the try/except around a pydantic `int | None` field is dropped.
One new test drives the real auth flow against a peer pair with no
`expires_in` and asserts the pair is adopted with zero POSTs.
flock is bound to an inode, not a path. Unlinking `<srv>.json.refresh.lock`
from `remove()` while a peer still holds the fence lets the next acquirer
open and lock a brand-new inode, so two processes hold "the" fence at once
and the single-use refresh token can be consumed twice. On Windows the unlink
of a locked file raises PermissionError straight out of `remove()`/`restore()`.
A 0-byte 0600 sidecar in a 0700 directory is harmless, so leave it in place.
It stays out of `_state_paths()` so `restore(only_if_absent=True)` still keys
off real token state only.
The SDK drives a refresh as a generator (request yielded from
_refresh_token, response consumed in _handle_refresh_response), so the
fence was a hand-driven @asynccontextmanager: __aenter__ in one method,
__aexit__ in another, generator object stashed on the provider. A plain
`acquire_refresh_fence(path, timeout) -> fd` / `release_refresh_fence(fd)`
pair says what actually happens and leaves nothing half-entered to leak.
The descriptor is opened with os.open at 0600 and closed on every
acquisition failure.
The `_refresh_token` release-on-exception stays: `_refresh_token` is also
reachable outside `async_auth_flow` (tests call it directly), and the
wrapper's finally only covers the generator-driven path.
HermesTokenStorage.remove() now unlinks the `.refresh.lock` sibling so
logout leaves no stray file. It is deliberately NOT added to
_state_paths(): snapshot()/restore(only_if_absent=True) treat any
existing state path as "newer state exists", and a lingering lock file
would silently veto a rollback.
tests/tools/test_mcp_oauth.py: drop the unused `Path` import (`time` is
still used by the socket poll helper).
The poll loop swallowed every OSError from the lock syscall as contention,
so a filesystem that cannot take advisory locks at all (ENOLCK on some
network mounts, EMFILE, ...) stalled for the full 60 s deadline and then
blamed a peer. Only EWOULDBLOCK/EAGAIN/EACCES/EDEADLK mean "held by
someone else"; anything else now raises RefreshFenceTimeout immediately
with the real errno. Still fails closed -- the refresh is never POSTed
without ownership -- but the failure is diagnosable and instant.
The errno set mirrors cron.scheduler._is_lock_contention_errno; it is
duplicated rather than imported because importing the scheduler pulls in
the whole cron module graph for a four-value tuple.
Both fence paths that pull a peer's pair off disk now go through
_hermes_rotated_candidate (different, non-empty refresh token + non-empty
access token) and _hermes_install_disk_pair, which runs
enforce_refresh_token_issuer on the installed pair. Before, the adopt path
skipped the issuer check entirely, so a pair minted by a different issuer
could be POSTed straight to the new one.
The adopt path installs the candidate even when its access token has
already expired: the POST we are about to build needs the new refresh
token, and skipping the POST (_RefreshCompletedByPeer) is only correct
when the peer's access token is live with a positive TTL, mirroring the
reload path's clamp-to-zero guard. When the issuer enforcer strips the
refresh token there is nothing to refresh with, so the flow restarts into
401 -> full authorization instead of failing with OAuthTokenError.
The reload helper drops its outer except-Exception: get_tokens already
returns None for absent or corrupt files, so the blanket catch only hid
programming errors. Comment updated: this path exists for writers outside
the fence (interactive login, pre-fence Hermes), not for a fenced peer.
`_token_store_lock` serialized a single get_tokens()/set_tokens() call and
then released. Its two justifications no longer hold:
- torn reads: `_write_json` goes through `atomic_json_write` (write to a
sibling, rename), so a reader can never observe a half-written token
file, locked or not;
- the read-modify-write of a single-use refresh token: a lock released
between the read and the POST cannot close that race. `_refresh_fence`
now spans read -> POST -> persist, and every store access on the refresh
path (adopt-from-disk read, post-failure reload, `_store_tokens` write)
runs inside it.
The remaining unfenced accesses are the cold `_initialize` read and the
authorization-code exchange's full overwrite -- neither is a
read-modify-write, so neither needs mutual exclusion. Keeping a second,
fail-open lock layer only adds a 10 s stall on a stale lock file with no
correctness gain. Tests that exercised the removed lock go with it.
The fence is entered from the SDK's coroutine-driven auth flow, so its
acquire loop spun on time.sleep(0.05) and blocked the whole event loop
for up to 60 s while a peer finished its network round trip. The fence
is now an async context manager that polls a non-blocking flock with
asyncio.sleep, and it creates the token directory itself (the parent may
not exist yet on a first refresh; secure_parent_dir only chmods).
The in-process RLock layer is dropped: an advisory lock on a fresh
descriptor already excludes sibling tasks and threads of the same
process, and a thread RLock is reentrant across asyncio tasks on one
thread, so it excluded nothing there anyway.
After acquiring the fence the provider re-reads the store; when a peer
already rotated the pair and the adopted access token is valid, it
raises _RefreshCompletedByPeer instead of building the refresh request.
The auth-flow wrapper restarts the SDK flow so the original request goes
out with the winner's access token. Previously the loser adopted the
new pair but still presented its stale refresh token, burning a
generation on every single-use provider.
Design lifted from #71715.
Co-authored-by: Kevin Yin <182213728+yinkev@users.noreply.github.com>
The token-store lock is per-operation: get_tokens() and set_tokens() each
take it and release it. With a provider that issues single-use refresh
tokens, two processes can therefore both read R1, both POST it, and the
loser gets invalid_grant on a session that was healthy:
A: get_tokens() -> R1 (lock taken and RELEASED)
B: get_tokens() -> R1 (lock taken and RELEASED)
A: POST R1 -> 200, receives R2
B: POST R1 -> 400, credential already burned
Add _refresh_fence(), held across read -> POST -> persist so exactly one
process consumes a refresh generation. It fails CLOSED: unlike the
token-store lock it raises RefreshFenceTimeout instead of degrading to
unlocked, because proceeding without ownership is the race itself. It
locks a .refresh.lock sibling rather than the token file, since
flock/msvcrt locks are per-descriptor and nesting one path would
self-deadlock on Windows and silently no-op on POSIX.
The provider takes the fence before its final read, re-reads under it so
a peer rotation is adopted instead of overwritten, and releases in
_handle_refresh_response. async_auth_flow also releases on abandonment:
a cancelled generator never reaches the handler, which would strand the
fence and turn the race into a deadlock. That wrapper delegates send and
throw manually -- async generators have no yield from, and `async for`
would feed the SDK response to the inner generator as None.
tests/tools/test_mcp_oauth_refresh_fence.py is an acceptance test, not a
unit test: two real OS processes refresh against a single-use-token
authorization server that audits every redemption. It asserts R1 is
presented exactly once, neither process clears the session, and disk
converges on the newest token. Verified to FAIL without the fence
(audit=[rt-1, rt-1]); a threading-only lock cannot catch this.
(cherry picked from commit c1508a1d47e2cbd94e05fa507684f3716a1e988c)
The desktop app spawns 'serve' while the scheduled task runs 'gateway run',
so two backends routinely share one HERMES_HOME. With a provider that issues
single-use refresh tokens, both could POST the same token and the loser's
refresh was rejected, clearing an otherwise healthy session.
Guard the token file's read-modify-write with a bounded advisory file lock
(fcntl on Unix, msvcrt on Windows, in-process only where neither exists),
mirroring how cron/jobs.py guards jobs.json. Acquisition is non-blocking with
a 10s ceiling: a briefly-contended refresh beats a permanently stuck client.
(cherry picked from commit 16f0d811de446a66ed5fd061fd7594fca3d230da)
The timeout branch of _run_hook_callback_bounded unconditionally added
gate_key to _hook_abandoned. A worker that finishes between done.wait()
returning False and the caller taking the lock has already popped its
token via _release_token, so nothing would ever clear that entry: the
callback stayed blocked for every later call id until reload with no
thread behind it. Guard the insert on the worker still being registered.
The new test makes the race deterministic by swapping the module's
threading.Event for one whose wait() lets the worker finish and then
reports a timeout, and asserts a fresh call id still runs.
Also pass tool_call_id inline from terminal_tool_result instead of the
conditional dict plumbing: an empty id is already treated as "no
identity" by _hook_call_identity and unknown fields are withheld from
narrow-signature callbacks (same shape as _fire_approval_hook). Update
the stale "(hook_name, id(cb))" comment above _hook_running_callbacks.
Hook callbacks are gated per (hook, callback, call identity). Two hooks
on the tool-loop path fired without any identity, so concurrent terminal
calls in one turn, or overlapping turns, still collapsed onto a single
gate key and the second invocation was skipped as if a callback had hung.
transform_terminal_output now forwards the tool_call_id bound in the
approval context around dispatch (only when set); transform_llm_output
forwards the turn_id already in scope. Payloads are additive: the
dispatcher withholds unknown fields from narrow-signature callbacks.