DEFAULT_OUTPUT_DIR was resolved once at import time, so long-lived
multi-profile runtimes (dashboard console, TUI/Desktop backend, cron,
kanban workers) kept writing synthesized audio into the launch
profile's cache/audio instead of the requesting profile's (#98749).
Same bug class and fix as skills_tool (f8723c478) and skills_sync
(#65828): keep the legacy module attribute for tests and external
patchers, but re-resolve from the live profile-scoped HERMES_HOME on
every synthesis call.
With multiplexing on, check_fns evaluated before any profile secret scope
exists fail closed by design: get_secret raises UnscopedSecretError and the
tool re-probes on the first scoped turn. _run_check_fn_uncached logged that
expected signal like a crashed check_fn (WARNING + exc_info), so every
multiplexed gateway start printed three full tracebacks that drowned real
check_fn failures.
Split the handler: an unscoped read reported while the profile cache scope
was unresolved logs one debug line without a traceback; the same error with
the scope resolved is a genuinely lost scope and keeps the loud
warning + traceback.
Fixes#100697
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.
bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.
On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.
Closes#100523
Supersedes #100544, #100542
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
Extends the TTS lease from #100912 (4e3feb8bbb) beyond built-in local
engines: when the configured tts.provider is user-declared, acquiring
the first lease and releasing the last one now reach it, so a
self-hosted TTS server can preload its model when read-aloud / voice
conversation turns on and unload when it turns off (Discord request).
- agent/tts_provider.py: TTSProvider gains concrete no-op warm() /
release() (not abstract — existing plugins are unaffected).
- tools/tts_tool.py: _signal_user_tts_provider() forwards the lease
hook; plugin providers get warm()/release(), command providers run
optional `warm_command` / `release_command` (config.yaml, under
tts.providers.<name>) through the existing _run_command_tts helper
on a daemon thread — best-effort, output discarded, failures at
debug. warm_tts_provider() and release_tts_provider() call it.
- tests/tools/test_tts_lifecycle_leases.py: fake plugin provider and
fake command provider observe warm/release through acquire/release
lease (both fail on main with action == "noop").
- docs: features/tts.md — lease section, command-provider optional
keys table, plugin optional hooks.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.
- model_metadata: memoize successful remote /models probes on disk
(cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
in-memory cache, so authority semantics are unchanged (reconciliation
still lands within 5 minutes) but the answer is shared across
processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
250ms instead of every 2s — up to 2s of dead air on every relayed reply.
Nothing here changes turn ordering: DMs and group rounds stay serial.
Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.
- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
workers by exact delegation_id (heuristic shape/time grouping kept for
older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
returned by the dispatch and used for cache/delegation/live/<id>/.
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
A manual cronjob(action='run') derived success from last_status == 'ok'
and read the error from last_error — so a run that now records
delivery_failed came back as success=False with error=None, an unexplained
failure. Surface last_delivery_error as the error in that case (the
#84006 direction, re-applied on the delivery_failed status), and pin the
manual-run completion summary to say 'Result: FAILED' over an undelivered
run. Document the status in the cron user guide.
Co-authored-by: webtecnica <webtecnica@gmail.com>
Main grew claim_job_for_fire(job_id, return_job=True) — a claimed
snapshot dict instead of a bool — while this branch sat on an older
base. The merge-ref CI ran the hybrid: the wiring tests still mocked
return_value=True, which fails isinstance(claimed_job, dict) and fell
into the 'already being fired' branch, so every dispatch assert failed.
Mock the claim to return the job snapshot (the API's success shape),
read the summary's deliver from the claimed snapshot the run actually
executes, and keep the dispatch-result failure renderer. Rebased onto
current main; cron suite 710 passed.
Review follow-up on the #83993 fix: a stored falsy deliver ("", JSON
null) fell through the local check and produced 'output was delivered
there by the job itself' for a target that does not exist — the exact
false-delivery-claim class the PR removes. Fire time already normalizes
falsy deliver to local (no delivery, output persisted in last_output,
no delivery error), so the summary now canonicalizes with the
scheduler's own _normalize_deliver_value and reads saved-locally.
Whitespace-only deliver is deliberately not folded in: fire time
records 'no delivery target resolved' for it, and the error-driven
FAILED wording must stay visible.
The _execute_job_now completion notice unconditionally claimed
"(output was delivered there by the job itself)" for non-local
delivery targets, even when the job record's last_delivery_error
showed the delivery failed (#83993). Derive the note from the
refreshed job record so a failed delivery is reported honestly to
the calling agent.
Follow-up to the #100829 salvage. uv pip install writes no __pycache__ by
default (pip does), so --compile-bytecode covers the whole install including
transitive deps, which the per-spec warm never sees. Also skip *.dist-info /
*.egg-info roots in _installed_dist_roots — they own no importable code.
Live: fresh cpython-3.12.13 venv, real uv install of anthropic==0.87.0 via
_venv_pip_install: main -> 0 pyc, first import 0.468s; after -> 1212 pyc
(546 anthropic), first import 0.205s.
Refs #100461
A pip/uv install writes .py sources and no __pycache__ — and reinstalling
the same version still deletes the cache the previous copy had. Nothing in
Hermes compiles them, so the whole compile is paid by whoever imports the
package next. For a lazily installed backend that is the foreground of a
user request, with nothing printed while it runs.
Measured for anthropic==0.87.0 (541 modules) on cpython-3.12.13: the first
import after an install costs 2.2-2.7s against 0.7-1.0s warm, and 10.5s
under concurrent load. N per-profile daemons cold-starting together each
pay it in full, because none of them has written the cache yet.
Compile the freshly installed distributions in _venv_pip_install instead,
on the success path of both the uv and pip tiers. The caller is already
waiting on an installer there and can see why. Package directories are
resolved from each distribution's own file list, so specs whose import
name differs from their package name (python-telegram-bot -> telegram)
are covered. Best-effort: a compile failure never invalidates an install
that succeeded, and sys.dont_write_bytecode is honored.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MhAnkrFktFdmZwf64fUYLE
A stdio server that never answers the optional ping (no -32601, no
response at all) produced a bare TimeoutError that _keepalive_probe
classified as a dead transport, tearing down and respawning a healthy
subprocess on every keepalive tick. On a first ping timeout, confirm with
list_tools before declaring death; if it answers, latch _ping_unsupported
and use list_tools from then on. If both fail, propagate as before.
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.
Fixes#73175
Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.
The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.
Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.
The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.
Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
exhausted, parked, tools deregistered -- no respawn loop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/proc/<pid>/task/<tid>/children is per-thread. stdio_client() spawns the
MCP subprocess from the background loop thread, so the main-thread-only
read returned an empty set on Linux and _stdio_child_pids/_stdio_pids
never tracked the child: the #81995 dead-child fast-fail, the #96452
respawn signal and the killpg shutdown sweep were all no-ops. Union the
children of every task instead.
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').
- gateway/run.py: run both hygiene executor hops (detached-agent path and
codex app-server path) inside copy_context().run, keeping the default
executor so a fence-cancelled hung summary never occupies a gateway
agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
class failure — abort and preserve the session instead of dropping the
middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
run_in_executor); UnscopedSecretError classified as access failure.
Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
Some bundled CPython runtime builds strip stdlib ThreadPoolExecutor's
copy_context() propagation, so work submitted to the daemon pool runs in a
bare context. Under the multiplexed gateway this dropped the profile
secret scope in pool workers: the context-compression timeout fence
resolved auxiliary provider keys (SURPLUS_API_KEY) with
UnscopedSecretError, silently degrading LLM compression to lossy
deterministic summaries and driving re-read loops in affected sessions.
Restore stdlib semantics in submit() by snapshotting the caller's context
and running the callable inside it (a no-op re-application on runtimes
that already propagate). Mirrors the gateway's
_run_in_executor_with_context pattern.
Tests: daemon pool worker sees caller contextvars; scoped get_secret works
in a daemon-pool worker under multiplex while scoped misses still fail
closed (no env leak).
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
The fast-fail gate probed the stdio child watcher by CALLING it —
inspect.isawaitable(_watch_children()) — creating a fresh coroutine on
every stdio MCP tool call that was never awaited (RuntimeWarning spam +
gc churn). Inspect the function instead of invoking it.
Salvaged (unique hunk only) from PR #96044; the bundled
_stdio_children_dead polarity fix was already on main via #94339.
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).
Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
the shared client's pooled sockets (force_close_tcp_sockets — FD release
stays with the owning worker), abort+poison of the cached per-request
openai/anthropic wire clients, Codex app-server request_interrupt(), and
the inline _active_request_abort hook. Never client.close(), never
socket.close().
- delegate timeout path: after registering the deferred-close callback,
run one immediate drain plus one 5s re-sweep (covers a connection opened
between the interrupt and the first sweep). The settled read (EOF/EPIPE)
lets the worker unwind, which triggers the deferred close on the worker's
own thread — the only safe FD-release boundary. A worker that still never
settles retains its resources rather than risking a cross-thread close.
Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.
Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.
Closes#94248
#100540 added a REMOVED_BACKENDS startup warning keyed on tavily; with the
backend restored, that entry would warn on a working provider. The registry
stays (empty) for future removals; migration tests now pin the machinery via
a synthetic entry plus a guard asserting no live provider is ever listed as
removed.
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).
- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
removed_backend_note(); selection_error() swaps in the specific
removal explanation (removed in v0.21.0, keyless alternatives) while
keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
search_backend / extract_backend against the registry and emits a
startup warning (deduped per stale value), surfaced by the existing
print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
per-capability keys, dedupe, healthy-config negative, live-backend
failure text preserved.
_run_chrome_fallback_command creates agent-browser-<session> in the shared
tmpdir and then opens stdout/stderr files inside it, but never writes the
<session>.owner_pid marker. _reap_orphaned_browser_sessions rmtree's any
agent-browser-* dir that carries no live owner and is not tracked in the
calling process, so a second hermes process — or a parallel test worker —
deletes the directory between the makedirs and the first os.open, and the
command dies with FileNotFoundError on _stdout_open.
Write the owner marker immediately after creating the directory, which is
what every other socket-dir user already does.
Deterministic repro on main: run any test that exercises the fallback while
a second process calls _reap_orphaned_browser_sessions() in a loop —
5/5 fail before, 6/6 pass after.
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
Salvaged from PR #13996 (@giwaov, issue #13963), modernized onto current
main: config key renamed to skills.create_dir (per PR #81002's naming),
resolution centralized in agent/skill_utils.get_skill_create_dir() with
~/${VAR} expansion and HERMES_HOME-relative paths, and the directory is
folded into get_all_skills_dirs() so created skills are discovered,
trusted, findable, and patchable like local ones. Out-of-root creations
report their absolute path instead of crashing relative_to().
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):
- Gate every scope-path branch on a new _IS_LINUX constant instead of
'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
never touches systemd code — no probe subprocess, no scope argv,
byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
legacy argv byte-identical, no unit recorded) and probe-returns-False
off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
wired into the on-demand windows-venv-e2e lane): real spawn_local on
windows-latest asserting jobs run exactly as before — spawned, output
captured, exit code correct, systemd path never reached even under
faked gateway identity.
Refs #70716, #71378.
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.
- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
install-state spinners off $hubActions, and an OfficialSkillDetail pane
(hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
Off should mean the model is never told the tool exists. A switch that
only made the call fail leaves Hermes offering walkthroughs it cannot
give and promising to point at things it cannot point at, which reads as
a broken agent rather than a respected preference.
Both gate on the switch through a shared desktop_ui.user_enabled helper,
which is the reactions check_fn generalized — same config read, same
reason it has to be config rather than an env var: the switch belongs to
the session's client, and the client may be on another machine.
docker run/exec argv previously carried -e KEY=VALUE pairs for every
forwarded/passthrough variable. On Linux /proc/<pid>/cmdline is
world-readable regardless of process owner, so every allowlisted secret
was visible to all local users via plain ps for the duration of every
terminal call.
Emit name-only -e KEY flags and supply values via the docker client
subprocess env instead: the docker CLI resolves valueless --env KEY from
its own environment (documented docker/podman behavior), moving secrets
from /proc/*/cmdline (0444) to /proc/*/environ (0400). Covers the docker
run container-start path, the recreation/recovery path, the init-seeding
exec path, and the per-command runtime exec path.
Reported by @sashalab. Fixes#96268
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):
- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
false-positives), and an explicit start-time expectation is now honored
on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
ProcessRegistry._terminate_host_pid (previously unverified), and the
session-close path now runs the same daemon identity verification as
the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
probes (real spawned processes, real psutil ancestry) wired into the
on-demand windows-latest wine2e lane.
Fixes#98814Fixes#89614
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.
The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.
Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.