Review feedback on #88136 (monerostar): a profile-scoped `hermes update`
sets HERMES_HOME to <root>/profiles/<name>, but the Hermes-managed
PortableGit tree lives under the SHARED root (<root>/git/...). The locator
checked get_hermes_home() only, so a broken trampoline during a
profile-scoped update was not swapped and fell through to ZIP.
Extract _portable_git_candidates() (shared root first, profile home as
fallback) and add a regression test for the profile layout.
A Git-for-Windows trampoline launcher (bin\git.exe / cmd\git.exe shim,
~46KB) that fails to re-exec the real git-core binary refuses every git
call with a "BUG (fork bomb)" guard instead of running it (#87876).
Detect the trampoline up front via `git --version`, locate a real git
binary (Git for Windows or Hermes-managed PortableGit locations), and
rebuild the git command with it so fetch/pull/checkout keep working with
a real git instead of degrading to the ZIP fallback. When no real binary
can be found, leave the command untouched so the existing fetch-failure
handler still falls back to the ZIP path on Windows (#88046).
Builds on webtecnica's escape-aware _split_key_path (#84152, cherry-picked
with authorship preserved; earliest fix in the family was RelaxJonh's #80253
greedy-match approach — both behaviors now ship together):
- _greedy_literal_match: when navigating an EXISTING mapping, prefer an
existing literal key equal to the dot-join of the next N path segments
(longest match wins). Dotted model IDs are the norm, so the common
unescaped command (config set providers.p.models.grok-4.6.supports_vision
true) now hits the real key across set/get/unset instead of creating a
phantom sibling. Plain dotted paths with no dotted-key collision split
exactly as before.
- _phantom_sibling + ValueError in _set_nested: refuse to CREATE a new
intermediate mapping that would shadow an existing dotted literal sibling
(Soju06's fail-loudly suggestion on #84064); set_config_value surfaces it
as a clean CLI error with the escaped spelling to use.
- utils.py::atomic_roundtrip_yaml_update (the second split site, #91607 —
/model + TUI persistence) now uses the same escape-aware split + greedy
literal matching.
- CFG-04 empty-segment guard now splits escape-aware so escaped keys are
not misclassified.
- Tests for every repro shape in the family: #84064 provider model keys,
#80006 Matrix room IDs, #91095 dotted models under custom_providers list
index (incl. escaped creation-when-absent), #91607 model_overrides via
atomic_roundtrip_yaml_update, #99124 dotted leaf keys; plus
backward-compat coverage. Also fixed the carrier's one stale assertion
(structured-value coercion landed on main after #84152 branched) and
removed its dead _MCP_SECRETS_CONFIG fixture flagged in review.
- Docs: 'Dots inside key names' section in website/docs/reference/cli-commands.md.
Fixes#84064, fixes#80006, fixes#91095, fixes#91607, fixes#99124
Generalize the HERMES_S6_SUPERVISED_CHILD supervisor-marker mechanism so
ANY supervised gateway launch (systemd, launchd, Windows Scheduled Task,
external supervisor) skips the active_profile redirect in
_apply_profile_override(). Previously only the s6 container marker was
honored, so a systemd-launched default-profile gateway with
HERMES_HOME=<root> followed the sticky active_profile file and silently
assumed another profile's identity — logging under that profile's tree
and connecting with its Telegram bot token (double-polling a token owned
by that profile's own live gateway).
- hermes_cli/main.py: honor HERMES_SUPERVISED_CHILD (new generalized
marker), HERMES_S6_SUPERVISED_CHILD (back-compat), INVOCATION_ID
(systemd; gateway commands only, since it leaks into every descendant
of systemd-launched processes), and HERMES_GATEWAY_EXTERNAL_SUPERVISOR.
- hermes_cli/gateway.py: export HERMES_SUPERVISED_CHILD=1 in generated
systemd units (user + system) and the launchd plist.
- hermes_cli/gateway_windows.py: export it from the Scheduled-Task cmd/vbs
launchers and the windowless respawn env overlay.
- hermes_cli/service_manager.py: export it alongside the s6 sentinel.
- tests: regression coverage for all markers + non-gateway INVOCATION_ID
neutrality + generated-unit marker presence.
Fixes#74872
Add _pid_record_belongs_to_current_profile() helper that verifies a
PID record's persisted hermes_home matches the current process. Use
it in get_running_pid() and get_runtime_status_running_pid() so the
default-profile gateway never mistakes another profile's gateway PID
as its own.
In _apply_profile_override(), clear HERMES_HOME instead of returning
early when it points to a profile directory but no --profile flag was
given, letting the sticky active_profile logic resolve the right one.
In _guard_existing_gateway_process_conflict(), detect stale PID files
from other profiles and emit a warning.
HermesCLI construction imports helpers from hermes_cli.config before the agent-setup mixin can run, so a mixed-version tree crashed with a raw traceback. Catch that ImportError on the chat entry path and tell the user to run hermes update.
Follow-up to salvaged PR #89129: both auto-correct sites now call
_with_preset_suffix() so a future correction path can't forget to
re-attach the @preset/<slug> routing suffix.
tarfile.open(archive_path, "w:gz") truncates the destination the instant
it opens. If tf.add() fails partway (disk full, permission loss,
interruption), whatever was previously at that path is gone — including
an existing profile or board export the caller chose to overwrite. This
is the same failure shape a7e7de6407 just fixed for the desktop gateway
file-save path, one commit earlier in the same window, but it was never
propagated to this shared archive-writing primitive even though board
export gained a new caller into it in that same window.
make_targz now writes into a sibling temp file (mkstemp, same directory
as the destination so the final step is a same-volume rename) and only
replaces the destination via os.replace() after the archive is fully
written and closed, mirroring the mkstemp+os.replace pattern already
used throughout this codebase (agent/secret_sources/_cache.py,
cron/jobs.py, gateway/status.py, etc). The temp file is unlinked on any
failure.
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes#83612.
The orphan-backend classifier fingerprinted candidates via
gateway.status.get_process_start_time, which prefers /proc/<pid>/stat —
the HOST process table, in clock ticks. Under the fake-psutil test harness
(and any containerized run where the PID number happens to exist on the
host) that returns the WRONG process's fingerprint in the WRONG units,
while pid_is_hermes verifies via psutil centiseconds at kill time: the
guard would then refuse every legitimate reap. Read create_time() from the
same psutil handle used for classification, quantized exactly like
gateway.status does on Windows, so the fingerprint round-trips.
Also covers the Windows-lane sibling: test_uses_netstat_and_taskkill_on_windows
now pins the guarded call path, plus a new refusal test for a non-bridge
listener PID (#89614 class).
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):
- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
false-positives), and an explicit start-time expectation is now honored
on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
ProcessRegistry._terminate_host_pid (previously unverified), and the
session-close path now runs the same daemon identity verification as
the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
probes (real spawned processes, real psutil ancestry) wired into the
on-demand windows-latest wine2e lane.
Fixes#98814Fixes#89614
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).
Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden
Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.
Refs #90471, #89614
Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
hermes chat -q wired the interactive prompt_toolkit clarify callback
unconditionally, but a -q turn never builds the prompt_toolkit
application — the modal can never be painted or answered, so the turn
polls its response queue until agent.clarify_timeout expires (default
3600 s, 0 = unlimited). The gateway, cron jobs, the kanban dispatcher
and inter-agent wakeups all deliver work as -q turns. Route the
single-query case to a headless callback at the agent-construction site
that already knows _single_query_mode, mirroring _oneshot_clarify_callback
on the -z path (#94943; third member of the family after #86909 and
#88013).
validate_platform_toolsets() accumulated a single valid_count across every
platform, so the "zero valid toolsets" safety net was suppressed as soon as any
one platform carried a valid toolset. A platform wiped to [] — the active one,
typically cli — therefore produced no warning at all.
resolve_enabled_toolsets() honours that empty list verbatim ([] is a list, so
the platform-default fallback is skipped), leaving the agent with zero tool
schemas. The model then has nothing to call and emits the tool call as
assistant text with finish_reason=stop: no error, no warning, no log entry.
That is the silent-failure mode this module was written to prevent (#38798).
Note the asymmetry this leaves intact: a malformed *string* value is not a list,
so it falls back to the platform default and fails open (#78103); an empty list
fails closed. The fail-closed resolution is deliberate (the explicit_empty_
selection contract in tools_config.py, and #82010 wants it persistable), so this
only adds the missing warning and does not change resolution semantics.
Fixes#89050
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.
The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.
Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
Rewrites resolve_exec_command so the generated .desktop Exec no longer
depends on how the installer happened to be launched: fixes the bare
repo-script form whose shebang escapes the venv, and the symlinked-venv
form that .resolve() dereferenced into the base interpreter store.
Salvaged squashed from PR #94874 (24 commits) after the original branch
was found to carry stray __pycache__/.gitignore payload.
Co-authored-by: Gökhan <gkhn.yldrmlr@gmail.com>
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:
Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
rejected the session token.
The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.
Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.
Fixes#96490
_denormalize_config_from_web only wrote model_context_length into the
on-disk model dict inside the branch gated on `model` also being present
in the payload. That was harmless when the frontend always sent the full
config, but the prior commit switched Settings autosave to send only the
diff (diffConfig), so editing the Context Window control alone omits
`model` from the payload and the context-length edit is silently thrown
away. The mirror case regressed too: editing `model` alone now omits
model_context_length from the diff, and the old code treated that missing
key the same as an explicit 0, wiping an existing context_length override
that the user never touched.
Track whether model_context_length was actually present in the payload
and only mutate context_length when it was, independent of whether
`model` also changed.
A venv ever touched by sudo pip / sudo hermes contains root-owned files
(classically site-packages/*.dist-info/INSTALLER). A later normal-user
'hermes update' pulls code fine, then 'uv pip install -e .' dies with
'Permission denied (os error 13)' mid-mutation — venv/bin/hermes already
deleted, CLI bricked.
Add a bounded, pure-stat ownership preflight (_venv_foreign_owned_paths)
that runs after the code pull and immediately before the dependency
install. If foreign-owned paths are found it refuses up front, names the
offending paths + owner uid, prints the exact recovery command
(sudo chown -R $(id -un): <root>), and confirms the venv is untouched.
Windows (no os.geteuid) and root skip entirely. Never raises, capped at
~2000 stat calls, no subprocess use (update tests mock subprocess.run).
Same refuse-before-mutate philosophy as the contended-venv gate (#87331).
Fixes#83529
Diagnosis and documented recovery by @eabase.
Widens the salvaged cron doctor with the highest-value fleet check:
an active job whose next_run_at is parked >15min in the past is not
firing (dead ticker, downed gateway, wedged fire-claim). Also registers
doctor in the docs (cron guide + CLI reference) and resolves the salvage
onto current main alongside runs/incidents/notepad.
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.
tick() now checks, BEFORE acquiring the tick lock:
skew detected (boot fingerprint != disk revision)
AND this process does not own the gateway runtime lock
AND that lock is held (a fresh gateway is alive)
-> raise CronTickYielded, skipping the tick entirely
Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
proceed; yielding is a certainty claim, never a guess
The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.
Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.
gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
On a fresh boot with no resume_pending sessions, _finish_startup_restore
opened the inbound gate almost immediately while the agent-side turn
machinery (run_agent import graph, tool schemas + check_fn probes,
context-file tier) was still cold. A message arriving in that window was
served with a skeleton system prompt (~1.7K tokens vs ~14.6K healthy):
no AGENTS.md/context tier, no tool schemas, memory provider initializing
mid-turn.
Fix: start a background turn-machinery warm-up when the startup gate
closes (overlapping the network-bound platform connects) and have
_finish_startup_restore await it — BOUNDED by
agent.gateway_startup_warmup_timeout (default 20s, 0 disables) — before
draining the queue and opening the gate. On timeout the gate opens
anyway and the warm-up finishes in the background, so a wedged init can
never make the gateway permanently unavailable.
Reported by @yhfmstr in #99373.
Fixes#99373
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.
1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
errors via the existing is_malformed_db_error() and raises 503 at all five
call sites. delete_session_endpoint was the worst — an unresolvable id
counted as idempotent success, so DELETE reported it had removed a session
that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
_await_gateway_exit() that also re-checks after the final sleep (a PID
exiting in the last interval must not be SIGKILLed — PID-reuse hazard).
NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.
Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
Follow-ups on the salvage: regular-file guard before the zeroed byte-probe (a FIFO at the state.db path would block startup forever — #98017 review P2), plus an on-main-reproducing UnicodeDecodeError fixture for #98924 (raw bytes in sqlite_master, not messages.content, are what reach pysqlite error-message decode).
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:
- web_server._open_session_db_at_path: the one-writable-open heal only
caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
to decode SQLite's own error message over corrupt file bytes) bypassed
it, so the heal documented for malformed schema never fired (#98924
Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
vtables whose probe raised UnicodeDecodeError, the same too-narrow
catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
could not open, so prompt.submit streamed the turn while persisting
nothing (#98924 Failure 2). It now returns False and prompt.submit
fails the RPC with code 5072 so desktop maps it to a toast, mirroring
the disk-full/5070 convention. session.create stays silent per its
pinned degraded-mode contract.
- Guard against concurrent-opener race where newly created 0-byte state.db was falsely quarantined before first schema write
- Wrap startup in quarantine_cross_process_lock when database is uninitialized or zeroed
- Guard is_zeroed_sqlite_file and is_zeroed_state_db against active live connections in current process
- Add concurrent-opener and live-connection regression tests
- lightpanda_engine_status: check use_real_profile before the cloud
provider, matching browser_exec's actual resolution order (real-profile
resolution runs before backend resolution), so /browser status and
hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
(find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
_using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.
- browser_use_cli: when the engine is lightpanda and nothing with higher
precedence claimed the session, get a session from _get_session_info()
and export its endpoint as BU_CDP_URL; the browser is private to the
session key, so the own-tab preamble is skipped. The browser_exec
description gains a Lightpanda header (text-first, new_tab once then
goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
--host 127.0.0.1 --port <free>` per session key (new
tools/browser_lightpanda.py), reusing the session cache, inactivity
reaper and atexit cleanup; a dead process is respawned on the next call;
orphans from a crashed Hermes are reaped through per-process records in
$HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
(cloud_provider: local + engine: lightpanda; "Local Browser" resets the
engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
is shadowed, the reason.
Follow-up to the salvaged manifest guard (#90859): the doctor's temporary
HERMES_HOME now enters the ExitStack before the staging copytree, so
ENOSPC / KeyboardInterrupt / any exception during staging deterministically
removes the hermes-plugin-doctor-* directory instead of relying on the
TemporaryDirectory GC finalizer (which cannot run while the traceback pins
the frame). Regression test proven via sabotage run against the old code.
`hermes plugins doctor` defaults its target to `.`, and
resolve_plugin_path accepted any directory that existed. Doctor then
copytree'd the resolved path into a temporary HERMES_HOME *before* the
runtime got to reject it, so running the command from a directory that
is not a plugin copied that whole tree.
Run from $HOME on macOS this copies the home directory, and because
`Library/CloudStorage` is not excluded it also materializes every
cloud-only Google Drive/iCloud placeholder. Observed locally: 461 GB
written to /private/var/folders and still growing when the process was
killed, on a machine with 49 GB free.
Resolve now requires a manifest before returning a path, mirroring
PluginManager._scan_directory: `plugin.yaml`/`plugin.yml`/`plugin.json`
in the directory itself, or in one immediate subdirectory for the
category layout. Plugin-id candidates are only tried for values that can
be an id, since joining `.` onto a plugins root resolves to the root and
would hand Doctor every installed plugin at once.
Non-plugin targets now fail with a clear message and no copy.