The salvaged fix gave `_find_missing` and `_flatten_messages` each their own
visited-set loop, next to the one `_is_session_expired_error` already had —
three copies of the same idiom in one module. Collapse them into
`_iter_exception_nodes` (pre-order, left-to-right, each node once, bounded by
`_EXC_TRAVERSAL_MAX_NODES`) and read all three scans off that list. Acyclic
output is byte-identical: the missing-executable search keeps its depth-first
order and a message-less leaf still renders as its class name.
Tests move from the issue-numbered file into `tests/tools/test_mcp_tool_errors.py`
(mirror of the source module): a two-node cycle renders the real messages, and a
missing stdio binary wrapped deeper than the recursion limit with the chain
looping back to the top is still reported as the missing executable. Both are
red on origin/main (RecursionError).
Co-authored-by: Stephan Mongstad <stephan@users.noreply.github.com>
Review finding on #112198: _mask_prose_link_destinations matched
_FENCE_LINE against the raw line, so a fence behind a CommonMark
container prefix (`- ```sh`, `1. ```sh`, `> ```sh`, nested) was not
seen and its body was scored as prose with link destinations masked.
Strip the container prefix before fence matching (open and close).
Bundled-skill rescan vs origin/main: 208 skills, 1447 findings on
both, no new/gone findings, no verdict changes.
Follow-up to the two cherry-picked contributor commits.
The picked fence tracker never checked for a closing fence once a block was
open (the closer test sat inside the not-in-code branch), so every prose link
after any code block was scanned verbatim again and the #111254 documentation
link exemption was lost; a fence line carrying an info string was also accepted
as a closer, which handed the scanner back to prose mode mid-block. Rewrite the
loop around CommonMark fence semantics: a block opens on 3+ backticks/tildes
indented at most 3 spaces (backtick info strings may not contain a backtick)
and closes only on a fence with the same marker, at least as long, and nothing
after it; tab- or 4-space-indented lines are code; an unclosed fence stays
code to EOF. plugin_guard inherits the behaviour through scan_file.
The temp-root exemption in destructive_root_rm now also refuses a parent
segment reached through an empty path segment or followed by a shell
separator, which the first cut let through.
Tests trimmed to one invariant per fix: the fence test covers the six code
shapes plus the prose-link-after-fence control that the picked version broke;
the rm test gains the two residual shapes.
Part of #111334Fixes#112129Fixes#111335
replace the boolean fence toggle in _mask_prose_link_destinations with
proper (marker_char, opener_length) tracking so a mismatched-markdown-fence
body or an indented code block cannot re-enable prose-masking over live
command lines. closes an exploitable bypass in the community-source
install path; plugin_guard inherits the fix through scan_file.
also tighten is_indented_code to treat any tab indent (single or double)
as code, per CommonMark §4.4.
`hermes mcp login <server> --flow device` took `authorization_servers[0]`
from the protected-resource metadata and failed when that entry was a
browser-only or issuer-inconsistent server, even though a later entry was
the issuer-bound device_code server meant for headless clients (Higgsfield
advertises exactly this shape: a PKCE server first, the device server second).
Discovery now tries each advertised server in order and binds to the first
whose metadata issuer matches its advertised URL and that offers device
authorization. Issuer validation (RFC 8414 / SEP-2468) is unchanged per
server; a single-server resource raises exactly the error it raised before,
and a multi-server resource with no usable entry reports every attempt.
The browser path (`tools/mcp_oauth_manager.py` pre-flight) is deliberately
left on the SDK's own first-entry selection: the SDK's 401-branch discovery
re-selects `authorization_servers[0]` itself, so a divergent pre-flight pick
would only desynchronise the cached metadata from what the SDK authorizes against.
`terminal(background=true, notify_on_complete=true)` appended its watcher descriptor to
`process_registry.pending_watchers`, which only the post-turn hooks drain. A process that
finished while the turn that launched it was still running (an agent sleep-polling for
hours) had no watcher task at all: the completion_queue entry sat inert, nothing was
injected, and the chat stayed mute until that turn ended (#112033).
- `_register_completion_watcher` arms the watcher on the live gateway loop at registration
(`GatewayRunner.arm_process_watcher`, via the existing `_gateway_runner_ref` /
`_gateway_loop` seam that send_message and cron already use); `pending_watchers` stays
the fallback while the gateway is not serving (checkpoint recovery at startup, shutdown).
- The agent-notify branch of `_run_process_watcher` keeps its design (the agent's next turn
is the user-facing report) but, when the launching turn is still active at process exit,
the injection only queues a follow-up — so the concise receipt is sent to the chat right
away instead of never. The busy check is taken before injection because the injected turn
itself installs the adapter's session guard.
Live probe (real process, real GatewayRunner loop, fake telegram adapter, busy session):
before — pending_watchers=1 after exit, 0 watcher tasks, 0 injections, 0 receipts;
after — pending_watchers=0, watcher task armed at launch, 1 injection, 1 concise receipt.
Control (idle session): 1 injection, 0 receipts, unchanged.
Slimmer redo of #112038 by @KoNit-K: same two gaps closed, without a second scheduler
registry / loop attribute on ProcessRegistry and GatewayRunner.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
CI runs the suite as an unprivileged user with HOME=/root in one fixture;
Path.is_dir() raised PermissionError from _user_local_bin_entries and the
run-env builder crashed. An unreadable home has no usable ~/.local/bin, so
treat the OSError as absent.
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.
Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.
Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.
Fixes#111778
clear_session (/new, /reset, auto-reset boundary) stamped entry.result="deny"
before waking the wait, and an interrupted coalesced leader published the
same deny to its followers, so both still rendered outcome="denied" /
"denied by user". Carry the cause on the entry (entry.cancelled) and let
_cancel_cause map a result-less wake to a withdrawn prompt; the wait still
unwinds fail-closed and the leader's own decision is unchanged.
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).
The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:
- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
category — no string matching) for the interrupted state and marks a
notifier-unregister wake (event set, result None) as "the turn ended before
the prompt was answered". Both the direct and the coalesced-follower wait
return `cancelled=<cause>`; the post_approval_response hook fires
choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
"BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
with outcome="cancelled" and its own user_summary; an explicit /deny is
untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
tool_reason ("parent delegation ended"; the late-child mirror forwards the
parent's own category) so a child's pending approval can tell teardown from a
user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
instruction-file gate and MCP elicitation consume the same key instead of
reporting "denied by the user" / "decline".
Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
terminal_tool's approval gate answers `status: pending_approval` with an
EMPTY `error` (#28323) and no `session_id`, so _spawn_delivery's specific
branch (`if parsed.get("error")`) was skipped and every unanswered
approval fell through to "Delivery to X failed to start: no process id
returned" — blaming the spawn for an approval nobody in a non-interactive
turn (api_server, `hermes peer dm`, cron) could grant.
- _spawn_delivery: the pending shape gets its own message (the runner
command needs terminal approval nobody in this turn can grant); a
local/peer DM adds "nothing was sent — approve it or add it to
command_allowlist and send again". Ownership is never transferred, so
the existing finally still reclaims the plaintext DM file.
- _try_relay_delivery: the envelope is queued on disk BEFORE the reply
waiter spawns and the Desktop drains it independently, so ANY waiter
spawn failure is a lost wake-up, not a failed delivery; reporting it as
an error made the sender resend and deliver the message twice. The
relay path now returns the shape _start_delivery's live-owner branch
already uses (status queued + notification_error + "Do NOT resend")
instead of inventing a new status value nothing reads.
Slimmer redo of #92971 by @jonpol01 (same diagnosis, same relay/local
split on `dm_file is None`); the source-text contract test and the
`sent_no_reply_wake` status were dropped.
Fixes#111716
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Two defects in tools/async_delegation.py:
- _push_completion_event called _persist_completion unguarded before
publishing onto completion_queue. One sqlite3 error (locked/full
state.db) dropped the completion event, left the record parked on
"finalizing" (a permanently leaked max_concurrent_children slot) and
let recover_abandoned_delegations later rewrite a succeeded unit as
"unknown". The write is now try/except: the failure is logged and the
event is still delivered, so _finalize flips the status and frees the
slot. A lost durable row is acceptable degradation; a lost result and
a leaked slot are not.
- _prune_completed_locked treated anything != "running" as finished,
while the module's own _LIVE_STATES also names stalling/finalizing.
A stalling record has no completed_at, so it sorted oldest and was the
first eviction candidate once the retained cap overflowed; its late
runner return then hit the missing-record path and the real result was
dropped. The predicate is now `status not in _LIVE_STATES`.
Slim redo of #76606 (earliest fix) and #112031: the converge/shield/
delete-row machinery both PRs built around the write is dropped as
defense-in-depth; the two core hunks are ported as-is.
Fixes#76605Fixes#112030
Co-authored-by: luckystar2026 <1393268817@qq.com>
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.
Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
Reuse tests/tools/test_delegate_output_schema.py's _StubChild instead of a
new one-test file with its own double; the invariant (the retry turn sees
is_delegated_child_context() True and the flag is restored afterwards) is
unchanged. Trim the source comment to the WHY.
_validate_child_output_schema issues a second run_conversation on the child when
the first answer fails the declared output_schema. The main child turn is wrapped
in delegated_child_context; this one was not. It runs on the parent worker's
thread, where HERMES_KANBAN_TASK is set and nothing marks the execution as a
child, so every identity gate keyed on is_delegated_child_context() fails open.
The visible effect is the kanban stop guard: it nudges the child to call
kanban_complete or kanban_block. A child owns no board task and carries no kanban
toolset, so it cannot, and the nudge text ("do not narrate intent", "finish any
remaining deliverable") displaces the structured answer the retry exists to
produce. The retry then fails the same schema and delegate_task reports an error
for a child whose work was already complete.
Observed with four children, each nudged during its retry:
[subagent-0] Kanban worker tried to exit without kanban_complete/kanban_block
[subagent-2] Kanban worker tried to exit without kanban_complete/kanban_block
[subagent-3] Kanban worker tried to exit without kanban_complete/kanban_block
[subagent-1] Kanban worker tried to exit without kanban_complete/kanban_block
4/4 - Final answer does not satisfy the declared output_schema (after 1 retry)
Wrap the retry the same way the main turn is wrapped. The context is entered and
exited around the single call, so nothing outside the retry sees it.
Signed-off-by: moep90 <volleyballlive@googlemail.com>
_expand_parent_toolsets built the parent's tool surface from each
toolset's declared `tools` only, so a composite parent's `includes` were
invisible: a child of a `debugging` parent (terminal/process_manage +
includes web/file) asking for `file` or `web` was refused, and `safe` /
`hermes-gateway` parents could grant nothing but their own name. Same
root cause as the `_strip_blocked_tools` fix in the previous commit
(#111700, "Related" section).
Both sides of the subset check now use the resolved static surface
(`resolve_toolset(name, include_registry=False)`), so a child may request
any toolset whose real tools the parent genuinely holds, and still never
gains a tool the parent lacks. Candidates that resolve to nothing are not
expanded into (they cannot be a meaningful subset).
Co-authored-by: DresvyanskiyDenis <dresvyanskiydenis@gmail.com>
`_sanitize_node` deleted the `required` key whenever the pruned list came
out empty. Four built-in tools (skills_list, todo, delegate_task,
session_search) declare `required: []`, so they left the sanitizer with no
key at all. Strict OpenAI-compatible proxies read the missing key as
`null` and 400 the whole request ("null is not of type array"), which is
non-retryable and kills the session on its first call.
An empty array is valid for every backend; the pruning was added (34c3e67)
to drop names that are not in `properties`, not to delete the key. Keep the
key with the filtered list, even when that list is empty.
Fixes#111684Fixes#59386
Co-authored-by: Cr4ckMe <jiqing.liu@whu.edu.cn>
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).
Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.
Fixes#111764
faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2
dlopens the CUDA runtime on the FIRST encode, which happens while the segments
are iterated in `_join_confident_segments()` — outside the try/except that
implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an
NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux
`libcublas.so.12`) the model loads fine, the error escapes the guard, and every
voice note fails with "Local transcription failed: Library cublas64_12.dll is
not found or cannot be loaded" until the user pins `stt.local.device: cpu`.
Materialize the segments inside the guarded block (first attempt and CPU retry)
so the dlopen failure reaches the existing evict-and-retry-on-CPU path.
Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class).
Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`,
`cudnn64_`, `cudart64_`) — the Windows message already matches the existing
"cannot be loaded" marker, proven by the live probe with the reporter's exact
string; the 6-test file was reduced to 2 invariant tests in the existing suite.
Fixes#111929Fixes#105295
Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate)
Co-authored-by: atmaksri <sri.atmakur@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>
Checkpoints were enabled by default from 9e845a6e (2026-03-16) until #20709
(2026-05-06) flipped the default back to off; the migration of that window
wrote `checkpoints.enabled: true` into user configs, where a later default
flip cannot reach it. Users who never type /rollback have carried a GB-scale
`~/.hermes/checkpoints/store` since (one live install: 1.2 GB across 250
projects, mostly disposable worktrees and /tmp dirs), and the cap cannot
bring it down because every project keeps at least one snapshot.
Silently flipping the key back is indistinguishable from overriding a real
opt-in, so this surfaces it instead: `checkpoint_footprint_notice()` returns
one line when checkpoints are enabled AND the store is at or above
`max_total_size_mb`, naming the opt-out (`hermes config set
checkpoints.enabled false` + `hermes checkpoints clear`) and the
retention knob. `hermes update` prints it with the post-update notices;
`hermes doctor` reports it as a warning after the state.db check.
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.
Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).
Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.
The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.
- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
calls it from the housekeeping tick (last chore), the CLI from a daemon
thread. Nothing on either startup path waits for git.
Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
Figma's authorization-server metadata advertises
authorization_response_iss_parameter_supported and its redirect omits iss,
so the mcp SDK's RFC 9207 check discarded every valid code and login never
finished. For that one issuer the provider fills a missing iss with the
discovered issuer and warns; a mismatching iss still fails and every other
server keeps the strict rule.
Fixes#111135
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The live branch of _run_delivery returns from _wait_live_dm before the
try/finally that removes <dm>.txt, so a settled delivery dropped its
.live.json intent but left the sibling .txt — the same plaintext — on disk.
_wait_live_dm now takes the dm file and removes both on settled; pending
and failed outcomes still keep both for the retry.
Review finding: settled live path unlinked <dm>.live.json but the sibling <dm>.txt survived.
A live DM's intent file (<dm file>.live.json — owner, delivery id and the
message plaintext) has to outlive its runner so a retry replays the same
delivery id instead of minting a second one. Nothing ever removed it: the
runner unlinks only the dm file, and cleanup_bot_dm_cache sweeps only *.txt,
so every live DM left its plaintext in the DM cache directory indefinitely.
Now the sweep reaps intents past the stale cutoff, and a delivery the owner
settled drops its intent immediately — nothing retries a settled delivery.
Follow-up to the salvaged #110786 commit:
- `_is_bot_mode_session` mirrors the system-prompt gate (`agent._session_title_hint`
first, then the live DB title) instead of reading `pending_title`/`title` off the
session dict: `pending_title` is cleared after turn 1 and the record never carries
`title`, so the contributor's gate matched only the very first Bot Chat turn.
- `tools/bot_mode_dm.py::_run_local_turn` (the `hermes -p X chat -c "Bot Chat" -Q`
transport behind `message_agent` when no live owner holds the target) re-emits ""
for a successful bare marker — the third delivery path of the same class.
- Tests trimmed to one invariant per surface (live completion, relay RPC, one-shot
transport), each proven red on origin/main sources.
- Bot Mode docs gain a "Staying silent" line pointing at the shared token list.
scoped_spawn_lost_user_bus() re-derived the user bus from an empty base env,
so it only ever looked under /run/user/<uid>. The worker itself is launched
with systemd_user_bus_env(worker_env), which honours a configured
XDG_RUNTIME_DIR. On a host whose bus lives outside the default runtime dir,
any unrelated systemd-run exit was therefore misread as "bus gone": the job
error named a missing bus that was still there, the cached scope verdict
flipped to False, and the next 60s of cron fires dispatched without cgroup
isolation.
The check now takes the spawn env, drops the bus address the spawn already
carried, re-derives from that, and decides on the DBUS_SESSION_BUS_ADDRESS
key rather than on dict truthiness (a non-empty env with no bus was never
"bus present").
Review finding: scoped_spawn_lost_user_bus used systemd_user_bus_env({}) instead of the worker's spawn env, so a configured XDG_RUNTIME_DIR made unrelated wrapper exits flip the scope cache to unscoped dispatch.
One TTL for both probe verdicts (the success-TTL constant collapses into
`_SYSTEMD_SCOPE_PROBE_TTL_SECONDS`), and the ack-wait loop consults
`scoped_spawn_lost_user_bus()` when a scoped dispatch exits before the
worker acknowledged: with `/run/user/<uid>/bus` gone the job error names
the missing bus and the enable-linger remedy instead of the wrapper's bare
`exit 1`, and the cached True flips so the next fire degrades to a direct
external subprocess rather than consuming another occurrence on a dead
wrapper. Contributor test trimmed to the revalidation invariant.
read_file/search_files passed file_read=True, which folded into code_file=True and skipped
the ENV/JSON/YAML assignment passes, so an opaque prefix-less credential under a
credential-shaped key reached the model in cleartext from a secret-bearing file — the
file-read half of the #110228 gate (#110567).
Two defects on that path, both fixed here:
- The rendered line-number gutter ("5| ADS_API_TOKEN: ..." from read_file,
"6: ADS_API_TOKEN: ..." from grep -n / cat -n) defeated the line-anchored patterns,
so the real rendered read leaked exactly what the raw text masked. A gutter-free fixture
cannot see this, which is why the tool-level tests carry the real render shape.
- _is_secret_file_arg() could not see the RESOLVED Hermes home: the default home's basename
is an installation detail (".hermes" on POSIX, "hermes" under AppData/Local on Windows) and
a resolved path never spells $HERMES_HOME, so the managed Windows home's config.yaml was
classified as ordinary YAML on both the file-read and the terminal surface.
Changes:
- redact_sensitive_text(): secret_file= re-enables the assignment passes for content the
caller classified with _is_secret_file_arg, keeping code_file behaviour everywhere else.
It is authoritative over code_file, so a caller cannot be fail-open on the security flag
by setting both.
- _redact_assignments(): mask_nonreusable selects the non-reusable sentinel for file reads,
so the #35519 write-back hazard stays closed.
- _should_redact_assignment(): no longer re-masks an already-masked value, which was erasing
the vendor label the sentinel deliberately keeps.
- _is_secret_file_arg(): consult the resolved Hermes home for the config.yaml arm.
- _CFG_ANCHORED_RE / _YAML_ASSIGN_RE: tolerate a rendered line-number gutter.
- file_tools.py: classify the resolved path at all three file-read call sites.
Closes#110567
create_task(parents=[open parent]) — the reporter's actual incident path —
parked the card in todo with only a `created` event, and kanban_create's
payload carried no `gated`, so the board still showed an unexplained todo
while only the link surface was fixed. create_task now appends the same
dependency_wait {reason: parent_not_done, parent} event and kanban_create
returns gated/gated_by, mirroring kanban_link.
link_tasks gated on `status != 'done'`, but _parents_satisfied and
recompute_ready treat `archived` as terminal: linking a ready child under an
archived parent demoted it to todo with a false parent_not_done event and the
next recompute promoted it straight back. Gate on not in ('done','archived').
Review finding: create_task(parents=...) emitted no dependency_wait/gated; link under an archived parent flapped ready->todo->ready with a false reason.
A ready child linked under an unfinished parent drops to todo with no
event and no operator signal; the only trace used to be claim_rejected
after a forced promote. Record a dependency_wait event when the demotion
fires, return the gate from link_tasks, warn in the CLI link command,
report gated in the kanban_link tool, and document the gate.
The websockets dials got proxy=None but the three HTTP /json/version dials
(CDP override discovery, is_browser_debug_ready used by Lightpanda and the
real-profile readiness check, and surviving-Chrome detection) still resolved
via getproxies(), so under a system/env proxy discovery fell back to the raw
http:// URL, readiness never fired and /browser connect reported not ready.
loopback_request_kwargs() sits next to loopback_connect_kwargs() and is used
at all three sites (ProxyHandler({}) opener for the urllib one).
add_loopback_no_proxy turned an operator NO_PROXY=* into '*,127.0.0.1,...',
which urllib/requests no longer treat as the wildcard, flipping bypass-all
configs into proxy-all. A wildcard in either casing now leaves env untouched.
is_loopback_host also accepts any loopback IP literal (127.x, ::ffff:127.0.0.1).
Review finding: HTTP /json/version dials still proxied loopback; NO_PROXY=* wildcard broken by append.
Move the loopback NO_PROXY merge from browser_tool into agent/proxy_bypass.py (the
module that already owns NO_PROXY semantics) and reuse no_proxy_entries() so comma-
and whitespace-separated operator values are both preserved. Add
loopback_connect_kwargs() and pass proxy=None on the two in-process websockets
dials to loopback CDP endpoints (browser_cdp_tool._cdp_call, BrowserSupervisor._run):
those never see the child env, so the env merge alone left them routed through a
macOS system proxy. Remote CDP URLs keep the default proxy behaviour.
Tests trimmed to two invariants: the built child env appends loopback to an
operator NO_PROXY in both casings, and only loopback URLs get proxy=None.
Sibling helper in tools/browser_use_cli (#110570) is redundant once the shared
env carries the entries.
websockets>=14 defaults to proxy=True and resolves proxies via
urllib.request.getproxies(), which reads the macOS/Windows system
proxy config even with no *_proxy env vars set. Local CDP endpoints
(ws://127.0.0.1:<port>/devtools/...) were therefore dialed through
the system proxy and the handshake failed with "did not receive a
valid HTTP response" (#110565).
Append 127.0.0.1/localhost/::1 to NO_PROXY/no_proxy (both casings)
in _build_browser_env so every browser subprocess (Browser Use CLI,
agent-browser, Chromium, Lightpanda) bypasses proxies for loopback.
Operator-provided NO_PROXY entries are preserved.
Fixes#110565
record_created now stamps created_by="learn" on foreground creates
(e.g. /learn) instead of leaving it unset, and the learning-graph filter
honors "learn" alongside "agent"/used. "learn" is a learning-signal
marker only: curator management stays keyed strictly on "agent"
(_is_curator_managed_record), so user-taught skills appear in /journey
without becoming eligible for autonomous curation.
Fixes#111317.
---
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; the contributor reviewed the diff and ran the tests.
An interrupted first download leaves refs/main plus a snapshot folder
without model.bin. snapshot_download(local_files_only=True) returns that
folder rather than raising LocalEntryNotFoundError, so ctranslate2 fails
with RuntimeError "Unable to open file 'model.bin'" and the online path
never ran, leaving STT permanently broken with a misleading error. Fall
through to the download when the local attempt fails that way. The
loading tests are also gated on faster_whisper being installed, so the
module collects when the voice extra is absent.
Review finding: partial-cache RuntimeError bypassed the online fallback.