check_for_skill_updates() fetched every lock-file entry remotely, even
when the entry's install directory no longer existed, and each fetch had
no wall-clock bound — a few dead sources turned a routine
`hermes skills update` into a multi-minute stall (#104291).
- Entries whose recorded install_path resolves but does not exist are
reported as "orphaned" and skipped without a remote fetch;
unresolvable paths keep the previous fetch behavior.
- Each fetch now runs under a daemon helper thread with a hard timeout
(default 30 s) and degrades to "unavailable" when abandoned.
- `hermes skills check` prints a removal hint for orphaned entries.
Fixes#104291
mcp.servers.status now rides the shared _mcp_rpc decorator (profile scope, 4064,
5024 with the real message) instead of a hand-rolled try/finally with a blanket
except. Drop the _MCPConnectErrorText str subclass and reason taxonomy: the
existing status/error fields already carry the state, and a whitelist on the RPC
keeps error text out of the wire. The Desktop connections.health contribution
contract is held back until its consumer plugin is public. Tests trimmed to the
scope invariants (per-profile runtime visibility, scoped shutdown clears only its
own status, launch runtime never leaks into another profile).
Older desktop builds appended a frozen "## Messaging other agents" section (roster
included) to SOUL.md. Since the server started injecting the live section into Bot Chat
sessions, that copy did two wrong things: every CLI/TUI/messenger session paid ~600 tok
for a bot-only protocol, and in Bot Chat itself the probe went silent when SOUL carried
the heading, so bots saw the stale roster instead of the live one.
- load_soul_md strips the legacy section at read time (covers un-migrated profiles and
the ambient-home edge cases the same way the SOUL isolation fix does)
- bot_mode_probe drops the SOUL-carries-heading suppression; a SOUL-era stored Bot Chat
prompt now counts as legacy and is upgraded once (stamped, so it cannot loop)
- config migration v41 rewrites SOUL.md across the default + every profile once
fix(approval): a grep inside "$(...)" no longer trips the hardline malformed block, and a quoted substitution body keeps its command boundaries (546 false blocks; review-found bypass closed)
fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum
A bare os.killpg/signal.SIGKILL trips the Windows-footgun lane (the module is
imported on Windows even though the native lane never runs there). The local
environment already has the POSIX group killer with the TERM→KILL escalation
and setsid-escapee sweep; use it.
Three call sites each repeated `if native: _run_rg_native(...) else: _exec(... | head -n N)`.
The choice now lives in _run_rg_bounded; callers pass the words, the bound, and the
one thing the native lane cannot express (a cd prefix → native_ok=False). The grep/find
pipeline keeps its explicit shell form because of the column cap.
Test file: one module-scoped LocalEnvironment instead of thirteen (~0.8 s each).
_native_rg_enabled was a pass-through to _native_read_enabled (a wrapper with
no behaviour); call the gate directly and note in its docstring that it
covers search too. The rg --files invocation was assembled twice (argv list
for the native lane, string for the shell lane) with room to drift; build the
command string once and hand it to either transport.
The first version checked the deadline only after a line arrived, so an rg
that produced nothing for 60 s (huge tree, no hits yet) pinned the caller
past the timeout and ignored the interrupt flag that the shell path honours
via _wait_for_process. Drain on a daemon thread; the waiter owns deadline
(124) and interrupt (130) and kills the process group, so no rg or child
survives the return. Probe: silent 10 s process, timeout=2 → 2.0 s / 124;
interrupt at 0.5 s → 0.5 s / 130; zero stray processes afterwards.
read_file already bypasses the backend shell on a local POSIX environment
(_read_file_native); search_files still paid two bash spawns per call — the
`test -e` existence probe and `set -o pipefail; rg ... | head -n N` — plus one
per zero-match probe. Measured on macOS against the repo's tools/ tree:
content search 85 ms → 15 ms, no-match search (three probes) 179 ms → 44 ms,
file-name search 73 ms → 12 ms; raw `rg` argv is ~15 ms, so the remainder was
transport.
Same gate and kill switch as reads (_native_read_enabled: LocalEnvironment,
not win32, HERMES_NATIVE_FILE_READ=0 disables). The argv builders and
_parse_search_output are unchanged and shared: _run_rg_native shlex-splits
the already-quoted words, streams stdout and stops after fetch_limit lines
like `head` would, and reports exit 0/1/2 (124 with partial output on
timeout) so the parser sees the shell contract. grep/find fallbacks, remote
backends, Windows and the multi-root cd form keep the shell path.
Two shell-observer tests in test_search_zero_match_and_multipath pin the
shell lane explicitly; they assert on command text, not behaviour.
When results were truncated, search_files appended the pagination hint
as plain text after the serialized payload ("{...}\n\n[Hint: ...]"),
so the tool result was no longer parseable JSON — downstream tool-message
handling on providers strict about tool-content formatting could reject
or mishandle it, contributing to 400 upstream errors in sessions with
truncated search output (#90322).
Move the hint into the payload as a structured _hint field, matching the
existing _omitted/_warning side-channel convention in the same function.
The model-facing guidance (explicit next offset) is unchanged.
Fixes#90322
The `hermes-real-profile` agent-browser daemon (the attach lane for consented
real-profile browsing) ran with plain `_build_browser_env()`: no
`AGENT_BROWSER_SOCKET_DIR`, so it lived in agent-browser's default dir and no
reap path could see it. A wedged daemon + headless Chrome survived 47h across
two gateway restarts (#100855), and on macOS the genuine Chrome binary it held
made "Chrome won't open" for the user.
Give the attach lane the same contract every other lane already has:
`_prepare_session_socket_dir()` (per-session socket dir + `owner_pid` claim)
and `_agent_browser_command_env()`, and add the named dir to the orphan
reaper's scan. The existing `_reap_socket_dir` then applies its owner-liveness
and start-time-fingerprint rules unchanged; the daemon is listed as tracked so
the untracked-idle escape hatch never fires under a live user (per-task `rp_*`
sessions drive it over `--cdp`, so its own dir shows no activity). The daemon-side idle
timeout is NOT inherited: Chrome is launched by Hermes, not the daemon, so a
self-exiting daemon would leave Chrome holding the copy dir while the next
attach re-runs the snapshot overlay over it.
When a reaped daemon's Chrome (Hermes-launched, own process group) still
holds the copy dir, `_real_profile_cdp` re-attaches to it instead of running
the snapshot overlay over a live profile. DevToolsActivePort outlives a
crashed Chrome and its port can be recycled, so the file's browser id must
match `/json/version` before it is trusted; an attach failure on a live
Chrome fails closed rather than overlaying.
Tests: attach/get/close commands carry the reaper-visible socket dir and
owner_pid and no idle timeout; a dead-owner real-profile daemon is reaped by
`_reap_orphaned_browser_sessions`; a surviving Chrome is re-attached, never
overlaid (all red on main).
The salvaged fix ran the injector against a copy.copy(agent) whose
tools/valid_tool_names pointed at the staged pair, wrapped in a blanket
try/except. A shallow copy of a live AIAgent (locks, DB handle, in-flight
attribute writes from the late-binding thread) is a workaround for the gate
mutating in place, and the except turned any failure into a silently
published snapshot WITHOUT message_agent.
Extract the gate as tools.bot_mode_dm.message_agent_authorized(agent) (the
same predicate ensure_message_agent_tool already used), and have the
snapshot builder append message_agent_tool_schema() to the staged list
directly when it passes. No copy, no swallow; same tests, same live
behaviour (compaction / between-turns / resume keep the tool; ordinary
sessions scrubbed).
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
200K-400K caps within 5% of each other in cost once cache prefixes are intact,
because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
measured; at 500K a 1M child compacts roughly never.
What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
Since #104299 a background delegate_task call is split into completion units
(one per `group`, one per ungrouped task). A multi-child unit still joined on
all its children before anything was written durably, so an owner crash
between the first and last child lost the finished work and replayed the whole
unit as "outcome unknown" — the restart-granularity gap that #104233 (Xipong's
#76228/#76229 direction) solved with a second row per child.
Each finished child of a detached unit is now recorded on the unit's OWN row
(`record_unit_child` → result_json {results, partial}) as its future lands;
the real result overwrites it at finalize. `recover_abandoned_delegations`
replays recorded children with their real summaries and marks only the
unfinished ones unknown, naming the count. No new rows, no new consumer shape.
`task_indexes` is persisted so recovery knows a split unit's members.
A background delegate_task call used to be ONE async unit: the runner joined on
every child and a single consolidated message re-entered the conversation when
the SLOWEST finished. Fifteen independent PR reviews therefore waited on the
fifteenth before the parent could act on the first.
Each task now carries an optional `group`. `_units_of` partitions the call's
children into units — one per distinct group, one per ungrouped task — and each
unit is dispatched to the async registry on its own, so its results re-enter the
conversation as soon as THAT unit is done. Tasks that must be compared or merged
share a group and still return together.
Capacity is unchanged: every unit of one call joins the first unit's pool slot
(`slot_key` in `async_delegation._dispatch`), so splitting never consumes more of
`delegation.max_concurrent_children` than the call did. Unit ids suffix the
call's id (`deleg_xxxx-1`, `-2`, …) so live transcripts stay under one dir; the
completion block names the group and notes that sibling units report separately;
`active_task_count` counts a unit's own tasks.
#76996 raised DEFAULT_READ_LIMIT, the read_file_tool signature and the schema
default from 500 to 2000, but the registry dispatch handler `_handle_read_file`
kept its own literal `args.get("limit", 500)`. Since models omit `limit` on most
reads, every dispatched read still stopped at 500 lines with `truncated: true`,
so the default flip never reached production traffic.
All three sites now read the single DEFAULT_READ_LIMIT constant, so the schema,
the Python default and the dispatch fallback cannot drift again.
Live probe (registry.dispatch on a 1500-line file, no limit): before 501 lines /
truncated=true; after 1501 lines / truncated=false.
`format_batch_tag` handed out `set N` ordinals from one process-wide table,
so every conversation on a shared backend and every child's nested fan-out
advanced the same counter. A user's second wave of 15 lanes rendered as
`[set 20 · 13/15]`, which reads like 20 batches were spawned.
Scope the ordinal table by the parent conversation (`parent_agent.session_id`)
and thread the parent through the three render sites (batch header,
completion lines, child tree-line prefix via the shared session_ref). The
first fan-out in a conversation is `set 1`, the next `set 2`; sibling
conversations and nested child batches no longer inflate it.
Invariant test proven red on origin/main, green with the fix.
OOMPolicy is a service-unit property; systemd-run --user --scope
rejects it with 'Unknown assignment: OOMPolicy=kill'. The availability
probe therefore always failed in supervised Linux gateways, making
restart-safe cron worker dispatch (systemd scope) permanently
unavailable and falling back to in-cgroup workers.
MemoryMax + MemoryAccounting remain; OOMPolicy adds nothing for a
transient scope (no service manager to act on the OOM event).
Review follow-up on the salvaged #103421 guard: the USER.md/MEMORY.md ternary
was a third copy (also _path_for and memory_tool._memory_target_error); use the
path's name. "profile" → "store" in the error since MEMORY.md is not a profile.
apply_batch() committed an empty USER.md/MEMORY.md as a normal
successful write when a consolidation batch removed the last entry.
Refuse all-or-nothing with live entries so background consolidation
keeps at least one entry; a deliberate wipe stays a manual file edit.
A delegated child copies the parent's prompt_caching.cache_ttl. The 1h tier
is priced for a person who steps away between turns (2x write vs 1.25x for
5m); a subagent calls every few seconds for minutes and is gone, so it paid
2x on every tool result and never collected the retention. Live measurement
(Sep 5, 40 concurrent Fable 5.1 children via OpenRouter, 1h markers on the
wire): cache writes billed at $20/M against $12.51/M for the same run at
5m, ~60% of a write-dominated bill.
_apply_child_cache_ttl runs right after child construction: 1h -> 5m,
5m stays, disabled stays disabled, parent untouched. Tests: unit (markers
on the wire drop ttl; parent's 1h layout still differs) and the real spawn
path with cache_ttl: 1h in a temp HERMES_HOME.
- probe_only is consumed inside __init__ only; no instance attribute.
- _sync_manager is assigned only on the probe branch, after the connection is
up, so a normal SSHEnvironment whose constructor fails still behaves exactly
as before (its __del__ cleanup does not reach the shared socket teardown).
- _remote_home has no reader on the probe path; not assigned.
- prompt_builder passes probe_only=True unconditionally: which backends honor
it is the builder table's decision, not a caller-side env_type check.
- Tests trimmed to the probe_only contract: the cleanup-raises case was green
on main (pre-existing try/except), the socket-name length and double-cleanup
assertions covered pre-existing behaviour.
The prompt-time backend probe built a normal SSHEnvironment just to run a
one-line `uname`. That constructor detects the remote home, creates the
~/.hermes tree, force-uploads every sync file and snapshots a login session;
when the throwaway object was later garbage-collected, __del__ -> cleanup()
ran sync_back() and `ssh -O exit` against the ControlMaster socket the
agent's real environment shares (keyed by user@host:port).
Add an internal probe_only construction path for SSH: an isolated,
same-length ControlMaster socket (keyed by the instance's session id), no
remote dir setup, no FileSyncManager, no session snapshot. The probe now
tears its own connection down explicitly on success, non-zero exit and
exception, without replacing the probe result when cleanup fails. Normal
SSH callers and non-SSH backends are unchanged.
Salvaged from #77933 onto the facade/sibling layout (the probe body moved to
_run_backend_probe, _create_environment to tools/terminal_tool_backends.py).
A message typed while the agent runs (CLI busy_input_mode=interrupt, gateway
priority redirect, ACP redirect) goes through AIAgent.redirect(). During tool
execution redirect() degrades to steer(), whose delivery rides the tool result
— so a long foreground command (a `sleep 285` CI poller, a build) parked the
user's message until it exited. The UI printed "Redirected current turn" while
nothing happened for minutes.
redirect() now also asks the tool workers to YIELD (tools/interrupt.request_yield).
The local terminal backend's wait loop honours it: the drain thread is stopped, the
still-running Popen is adopted by the process registry as a notify_on_complete
background session (ProcessRegistry.adopt_local — output so far seeds the buffer,
the registry reader continues from the pipe), and the tool returns immediately with
status "yielded_to_background" + session_id. The command is never killed; the
completion notification arrives as usual and process(poll/wait/log/kill) work on it.
Non-local backends and internal env.execute() consumers pass no yield_handler and
are unaffected; a stale yield bit is cleared with the interrupt bit per worker tid.
DockerEnvironment.cleanup() runs docker stop + docker rm -f on a daemon thread and
keeps the handle on the env. The idle reaper pops the env out of _active_environments
BEFORE calling cleanup(), and the atexit drain only iterated that registry — so a
detached env's worker was unreachable and died with the interpreter, leaving a
stopped (or, if the exit came fast enough, still running) labeled container behind
while the log said the environment was cleaned.
Every teardown worker is now also recorded in a module-level set in docker.py;
_atexit_cleanup joins that set after the registry pass via
DockerEnvironment.wait_for_all_teardowns (re-snapshotting each pass, since the
reaper can start a worker while the drain runs). Finished workers are dropped from
the set on each drain so it cannot grow across a long gateway life.
Mechanism from #86344 by @PRATHAMESH75; this is the slim redo — one shared set
plus one static drain hooked into the existing terminal_tool atexit, instead of a
second atexit registration in docker.py.
On hosts where Docker ships as a snap (Ubuntu cloud images / Azure VMs), the
snap's AppArmor confinement turns two sandbox hardening flags into a dead
container at start: `--init` fails with "exec /sbin/docker-init: operation not
permitted" and `--security-opt no-new-privileges` then fails every exec the
same way ("exec /usr/bin/sleep: operation not permitted"). This is snapd
LP#1908448 — not probeable from the client, and docker_extra_args cannot remove
flags we add.
`terminal.docker_snap_compat: true` drops exactly those two flags; cap-drop ALL,
the tmpfs hardening, PID limits and the privdrop caps are unchanged, and a
warning is logged at container start. Bridged everywhere the other docker_*
keys are (CLI env map, gateway env map, `hermes config set` sync, terminal_tool
env read, the shared container_config shaper, DEFAULT_CONFIG).