Review feedback on #88136 (monerostar): a profile-scoped `hermes update`
sets HERMES_HOME to <root>/profiles/<name>, but the Hermes-managed
PortableGit tree lives under the SHARED root (<root>/git/...). The locator
checked get_hermes_home() only, so a broken trampoline during a
profile-scoped update was not swapped and fell through to ZIP.
Extract _portable_git_candidates() (shared root first, profile home as
fallback) and add a regression test for the profile layout.
A Git-for-Windows trampoline launcher (bin\git.exe / cmd\git.exe shim,
~46KB) that fails to re-exec the real git-core binary refuses every git
call with a "BUG (fork bomb)" guard instead of running it (#87876).
Detect the trampoline up front via `git --version`, locate a real git
binary (Git for Windows or Hermes-managed PortableGit locations), and
rebuild the git command with it so fetch/pull/checkout keep working with
a real git instead of degrading to the ZIP fallback. When no real binary
can be found, leave the command untouched so the existing fetch-failure
handler still falls back to the ZIP path on Windows (#88046).
Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.
Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.
Reported by @Artemonim in #57275 (residual claim 5).
The MINIMUM_CONTEXT_LENGTH floor in _compute_threshold_tokens only
degraded to the 85% trigger when it met or exceeded the effective
window exactly (#14690). Near-minimum windows slipped through: at
context_length=65536 the threshold passed through at 64,000 — 97.7%
of the window, ~1.5K tokens of output room — so pre-API compaction
effectively could not fire.
Providers that silently truncate over-window prompts instead of
rejecting them (e.g. ollama's OpenAI-compatible /v1 endpoint) never
deliver the reactive context-overflow backstop either. Observed live
on a 65,536-token local model: the session rode into the window
ceiling and each length-continuation retry re-sent a window-filling
prompt (65,120 -> 65,273 prompt tokens, 263 output tokens of room)
until the turn died with "Response remained truncated after 4
continuation attempts" — every retry paying a full multi-minute
prefill.
Cap the floored threshold at _MIN_CTX_TRIGGER_RATIO (85%) of the
effective input budget whenever the floor is the binding term. An
explicit threshold_percent above 85% is user intent and stays
uncapped; windows where the floor lands at/below the cap are
unchanged.
Builds on webtecnica's escape-aware _split_key_path (#84152, cherry-picked
with authorship preserved; earliest fix in the family was RelaxJonh's #80253
greedy-match approach — both behaviors now ship together):
- _greedy_literal_match: when navigating an EXISTING mapping, prefer an
existing literal key equal to the dot-join of the next N path segments
(longest match wins). Dotted model IDs are the norm, so the common
unescaped command (config set providers.p.models.grok-4.6.supports_vision
true) now hits the real key across set/get/unset instead of creating a
phantom sibling. Plain dotted paths with no dotted-key collision split
exactly as before.
- _phantom_sibling + ValueError in _set_nested: refuse to CREATE a new
intermediate mapping that would shadow an existing dotted literal sibling
(Soju06's fail-loudly suggestion on #84064); set_config_value surfaces it
as a clean CLI error with the escaped spelling to use.
- utils.py::atomic_roundtrip_yaml_update (the second split site, #91607 —
/model + TUI persistence) now uses the same escape-aware split + greedy
literal matching.
- CFG-04 empty-segment guard now splits escape-aware so escaped keys are
not misclassified.
- Tests for every repro shape in the family: #84064 provider model keys,
#80006 Matrix room IDs, #91095 dotted models under custom_providers list
index (incl. escaped creation-when-absent), #91607 model_overrides via
atomic_roundtrip_yaml_update, #99124 dotted leaf keys; plus
backward-compat coverage. Also fixed the carrier's one stale assertion
(structured-value coercion landed on main after #84152 branched) and
removed its dead _MCP_SECRETS_CONFIG fixture flagged in review.
- Docs: 'Dots inside key names' section in website/docs/reference/cli-commands.md.
Fixes#84064, fixes#80006, fixes#91095, fixes#91607, fixes#99124
Follow-up to the salvaged #98571: forward the interrupt to the compute
host whenever the parent 'running' mirror is stale, but only for
sessions that actually have hosted activity — HostSupervisor.interrupt()
calls start(), so an unconditional forward would spawn a compute-host
child just to deliver an interrupt for an idle lazy session.
Adds a regression test asserting the idle-lazy-session no-spawn path.
Refs #92916
reconcileRegistryDrift only healed remote/cloud v1 routes. A v1 global
mode:'ssh' route (host, no url) written by Settings after the one-shot
migration had no registry identity: resolvedConnectionId returned null,
primary stayed 'local', and every launch re-homed the window onto a
fresh local backend. Because the heal skipped SSH entirely, the two
config files re-drifted after every update relaunch instead of
converging once.
Normalize the v1 SSH descriptor into a v2 kind:'ssh' entry (via the
same validated normalizeConnectionInput path the editor uses) and align
primary/lastUsed, with the same narrow-heal rules as remote: already-
registered targets and deliberate primary picks are left alone, and
unusable hosts never touch the registry.
Diagnosis credit: mgallmur-glitch (root cause) and jakewvincent
(re-drift after update relaunch) on #93888.
A no-mux tunnel is a single persistent `ssh -N -L` child. On main, ANY
death of that child after readiness immediately set tunnel.alive=false,
which poisons SshConnection.isAlive() forever; upstream lifecycle probes
then treat the whole SSH connection as dead, tear down the scope, and
SIGTERM a perfectly healthy backend (~10s after HERMES_BACKEND_READY in
the #96266 logs: '[ssh] connection closed (no-mux tunnels killed)' ->
'Ignoring stale Hermes backend exit (SIGTERM)' -> 90s port-announcement
timeout, with retry/repair looping the same failure).
Now a post-readiness child death is a tunnel FLAP: the child is
restarted with a bounded budget (5 attempts, 1s delay by default,
injectable for tests) and only an exhausted budget marks the tunnel —
and thus the connection — unhealthy. Deliberate teardown (cancelForward
/ close) sets tunnel.stopping, cancels any pending restart timer, and
never restarts. Pre-readiness deaths keep failing fast with classified
stderr (auth/bind errors unchanged).
Fixes the kill chain of #96266.
The isinstance(primary, Path) branch in _candidate_sentinel_paths was dead
weight: the surrounding except Exception already covers non-Path test
doubles, and .resolve() failing on them falls through to the plain
inequality comparison. Verified the pre-existing fail-safe stat fixture
(test_is_engaged_fails_safe_on_stat_error) still passes without it.
Module docstring still claimed 'a single os.stat'; the fleet-root check
makes it one or two stats. Updated.
Profile processes launch with HERMES_HOME=~/.hermes/profiles/<name>, so
`hermes pause` at the fleet root did not bind fleet-analyst dispatch
(t_7b65ff88). Check/resume both the process home and the fleet root.
Generalize the HERMES_S6_SUPERVISED_CHILD supervisor-marker mechanism so
ANY supervised gateway launch (systemd, launchd, Windows Scheduled Task,
external supervisor) skips the active_profile redirect in
_apply_profile_override(). Previously only the s6 container marker was
honored, so a systemd-launched default-profile gateway with
HERMES_HOME=<root> followed the sticky active_profile file and silently
assumed another profile's identity — logging under that profile's tree
and connecting with its Telegram bot token (double-polling a token owned
by that profile's own live gateway).
- hermes_cli/main.py: honor HERMES_SUPERVISED_CHILD (new generalized
marker), HERMES_S6_SUPERVISED_CHILD (back-compat), INVOCATION_ID
(systemd; gateway commands only, since it leaks into every descendant
of systemd-launched processes), and HERMES_GATEWAY_EXTERNAL_SUPERVISOR.
- hermes_cli/gateway.py: export HERMES_SUPERVISED_CHILD=1 in generated
systemd units (user + system) and the launchd plist.
- hermes_cli/gateway_windows.py: export it from the Scheduled-Task cmd/vbs
launchers and the windowless respawn env overlay.
- hermes_cli/service_manager.py: export it alongside the s6 sentinel.
- tests: regression coverage for all markers + non-gateway INVOCATION_ID
neutrality + generated-unit marker presence.
Fixes#74872
Add _pid_record_belongs_to_current_profile() helper that verifies a
PID record's persisted hermes_home matches the current process. Use
it in get_running_pid() and get_runtime_status_running_pid() so the
default-profile gateway never mistakes another profile's gateway PID
as its own.
In _apply_profile_override(), clear HERMES_HOME instead of returning
early when it points to a profile directory but no --profile flag was
given, letting the sticky active_profile logic resolve the right one.
In _guard_existing_gateway_process_conflict(), detect stale PID files
from other profiles and emit a warning.
shutdown_mcp_servers() blocks on future.result(timeout=15) which, called
from the gateway event-loop thread during SIGTERM teardown, freezes the
loop for up to 15s when the MCP loop and its stdio children are torn down
concurrently. Supervisors with a shorter kill grace (s6-overlay: 3s)
SIGKILL the gateway before lifecycle_ledger.mark_exited() runs, producing
phantom 'exited UNCLEANLY' reports on every subsequent boot.
Run the sync shutdown on a daemon thread and poll via _await_thread_exit
with a 5s budget; proceed with teardown if it wedges. Fixes#82874;
completes the shutdown half of #64155.
Real byte-flip fixtures (no mocks) proving the #96038/#98090-class fix
end to end, closing the acceptance gate on issue #97940:
- test_canonical_btree_corruption_fails_closed: checkpoint the WAL,
clobber every messages-table B-tree leaf page header, then assert a
live append raises the genuine bare SQLITE_CORRUPT, the classifier
refuses the FTS route, no rebuild/detach/stale-marker side effects
occur, and the field incident's misdiagnosis log line ('canonical
message rows are preserved') never appears.
- test_fts_only_corruption_still_self_heals: contrast case — a real
messages_fts_data shadow-table stomp raises SQLITE_CORRUPT_VTAB (267),
is classified as FTS-scoped, and the write path still self-heals with
canonical rows intact.
Sabotage-verified: reverting the classifier fix (96739033c4) makes the
canonical-corruption test fail by entering the FTS self-heal route.
Credits @fangliquanflq (PR #98090) for the production timeline analysis
and @diatche (PR #96038) for the classifier fix these tests gate on.
Refs #97940, #98077.
Untagged session rows come from the ambient primary backend. Do not synthesize a local owner for them, and clear stale explicit hints before issuing an id-only resume.
HermesCLI construction imports helpers from hermes_cli.config before the agent-setup mixin can run, so a mixed-version tree crashed with a raw traceback. Catch that ImportError on the chat entry path and tell the user to run hermes update.
findElectron() probed exactly one path, and got three things wrong at
once for anyone not on a hoisted POSIX install:
* It looked only under the REPO ROOT. This is an npm workspaces repo and
npm hoists a dependency only when nothing conflicts, so `electron`
installing into apps/desktop/node_modules is an ordinary outcome, not a
broken tree.
* It joined a bare `electron`. On Windows the dist file is
`electron.exe`, so the probe could never match there.
* Its PATH fallback spawned `which`, which is not a command on Windows,
so the fallback failed for a reason unrelated to whether electron is on
PATH.
The three combine into a misleading error: the suite refuses to start
with 'Run "npm install" from the repo root' on a tree that has electron
installed. Reproduced on Windows 11 against this repo, where
apps/desktop/node_modules/electron/dist/electron.exe exists and the old
body throws that message; the reporter on #88036 hit the same thing on
Linux and had to hand-symlink the package before the suite would run.
Resolution now asks the installed `electron` package for its own path
first (its main export IS the absolute executable, resolved from
path.txt and honouring ELECTRON_OVERRIDE_DIST_PATH), then falls back to
explicit dist probes for each root, then to PATH with the platform's
lookup command. The error message lists what was searched.
The rules live in e2e/electron-binary.ts so they can be unit-tested
without importing the Playwright runner, with the platform passed in
rather than read from process.platform: reading it would leave every
Windows rule untested on the Linux CI runner.
Wiring: the vitest `electron` project picks up e2e/**/*.unit.test.ts and
Playwright ignores the same pattern, so helper unit tests run in exactly
one runner and the specs are untouched.
Verified: 5 unit tests pass; mutation-checked one rule at a time
(hardcoding the binary name fails 2, reversing the probe order fails 1,
hardcoding `which` fails 1). tsc -p . and tsc -p tsconfig.e2e.json
clean.
This is the environment blocker called out in #88036, not its rendering
bug, so it is deliberately a subset.
Refs #88036
The _CFG_SECRET_WORD_RE pre-gate only skips secret-FREE text. A compaction
payload containing one real secret assignment plus a long opaque dotted run
still reaches _CFG_DOTTED_RE's backtrackable '*' prefix, which re.sub retries
from every byte of the run — quadratic while holding the GIL (same class as
the _ENV_ASSIGN_LOWER_RE fix in this branch, #99255).
Anchor each attempt to the start of a key run with a negative lookbehind.
Match set is unchanged: any match starting mid-run implies a leftmost match
at the run start, verified 20/20 identical over a dotted-config corpus.
30k-char adversarial run: 102s -> 0.015s.
The Codex auxiliary Responses adapter enforced a single absolute
deadline (300s floor for compression). A dead stream held the entire
budget before fallback ran, and repeated compression attempts stacked
those waits into 20+ minute 'Summarizing thread' stalls (masoria debug
bundle, Aug 31 2026). Meanwhile a healthy-but-slow reasoning summary
was killed at the same absolute deadline even while producing tokens.
Replace the absolute kill with progress-aware deadlines:
- 60s no-progress window for the first substantive payload AND between
payloads; keepalive/lifecycle frames do not re-arm (mirrors the
commit-fence gating, #96707)
- a live stream re-arms per token and is bounded only by
_aux_stream_total_ceiling() (max(600s, 4x configured timeout)), the
same backstop the streamed chat.completions path already uses
- the compression critical-path retry gate now distinguishes failure
cost: a cheap first-token no-progress failure retries the same
provider once; mid-stream stalls and ceiling hits still skip straight
to provider fallback (#54465 semantics preserved)
Live A/B (real OpenAI SDK against a local SSE server, real adapter):
dead keepalive-only stream: main waits the full budget; fixed fails
over at the window. Slow-but-alive stream (tokens past the configured
timeout): main kills it mid-generation; fixed completes.
Widen #96038's fail-closed classifier to the gateway transcript retry
path: SessionStore._is_fts_corruption_error no longer treats a generic
'database disk image is malformed' as FTS-only damage. It now delegates
to SessionDB._is_fts_write_corruption_error (SQLITE_CORRUPT_VTAB result
code or explicit fts5 corrupt-structure text) and only keeps the
messages_fts-named cases. Structural corruption falls through to the
bounded retry/backoff path instead of rebuilding FTS and retrying writes
against a damaged database.
Sibling site spotted in PR #98090 by @fangliquanflq.
docker run/exec argv previously carried -e KEY=VALUE pairs for every
forwarded/passthrough variable. On Linux /proc/<pid>/cmdline is
world-readable regardless of process owner, so every allowlisted secret
was visible to all local users via plain ps for the duration of every
terminal call.
Emit name-only -e KEY flags and supply values via the docker client
subprocess env instead: the docker CLI resolves valueless --env KEY from
its own environment (documented docker/podman behavior), moving secrets
from /proc/*/cmdline (0444) to /proc/*/environ (0400). Covers the docker
run container-start path, the recreation/recovery path, the init-seeding
exec path, and the per-command runtime exec path.
Reported by @sashalab. Fixes#96268
teknium flagged it as scope creep on PR #98691 review: no callers in the diff.
_is_explicit_fork_child_row (the row-based helper actually used) is unchanged.
Follow-up to the per-file budget in the previous commit, which closed the
scaling axis it measured and left three others open.
* A per-file ceiling still lets the cost grow with the PROFILE count: a
multiplexed gateway serves N profiles from one process and each has its own
state.db, so `_READ_POOL_MAX` bounded each file while the process total went
unbounded — the per-instance bug one level out. `_READ_POOL_PROCESS_MAX`
(three files' worth) now bounds the process, and a miss reclaims an idle
connection from ANY path before degrading: a profile quiet for an hour must
not hold descriptors the profile being served right now needs.
* Hermes's SQLite descriptors are only ever a share of the fd table. In #98573
the ~20 state.db handles were not the whole 256 — they were the share that
pushed httpx sockets and terminal subprocess pipes over, and the EMFILE
surfaced in tools/terminal_tool.py rather than here. New read connections are
now refused when the process is within `_FD_HEADROOM_RESERVE` of its soft
RLIMIT_NOFILE, measured from /proc/self/fd or /dev/fd and cached briefly. The
guard fails OPEN where it cannot measure (Windows has neither the fd
directory nor RLIMIT_NOFILE, and a CRT limit in the thousands) and CLOSED on
evidence — including a probe that could not get a descriptor of its own.
`_read_open_denied_fd_headroom` makes it diagnosable from a running process.
* Writer connections cannot be rationed the way read connections can: a
SessionDB without one cannot write. Their only real bound is not opening
redundant handles, so a process that accumulates more than
`_HANDLES_PER_PATH_WARN` handles on one file now says so once, and the next
duplicate is visible before it is an incident instead of inferred from an
lsof after one.
`_READ_POOL_MAX` itself is deliberately unchanged at 8. Retuning that constant
is #98585's subject; with a process ceiling above it and the headroom guard
in front of it, the value is no longer the binding constraint.
Refs #98573
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Issue #98573 reports a long-lived gateway holding ~20 `state.db` descriptors
that never shrink, walking into the 256 soft RLIMIT_NOFILE a launchd/systemd
service manager hands the process. The cause named there — a per-thread
`threading.local()` read connection — is already gone (87aedbe7b6 pooled the
read connections, 0472c31aa1 added the peak permit). Measured on main: one
SessionDB with 40 concurrent reader threads peaks at 9 live connections, not 40.
The symptom survives one layer up. `_READ_POOL_MAX` was enforced by a
BoundedSemaphore owned by each SessionDB, which bounds the wrong noun: the
descriptors are spent on a FILE, so every additional handle on one state.db got
its own allowance and peak scaled as `instances x (1 + _READ_POOL_MAX)`.
Two changes, both needed:
* The permits move to a per-path `_PathReadBudget`, shared by every SessionDB
in the process that points at that file. A permit miss first reclaims an IDLE
pooled connection from a peer handle before degrading to the writer lock —
without that, whichever handle warmed up first would pin the whole budget and
permanently demote every later one (a cron job's transient handle, a second
profile's store) to the locked writer connection.
* `GatewayRunner` borrows `SessionStore`'s handle instead of opening its own.
Both caches resolve the same `_default_db_path()`, so the process was holding
two writer connections and two read pools against one file for no reason, and
doubling again per profile on a multiplexed gateway. The store owns the
connection and sweeps it at shutdown; the runner's cache now holds only the
async wrapper and its sweep skips borrowed handles.
Measured peak live connections against one file, 40 reader threads, by handle
count 1/2/4/8:
before: 9 / 18 / 36 / 51 (51 not 72 only because the sample window ended
before every pool filled)
after: 9 / 10 / 12 / 16 (read connections capped at 8 in total; the
remainder is one writer per handle, and the
gateway's per-profile pair is now one)
Fixes#98573
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to salvaged PR #89129: both auto-correct sites now call
_with_preset_suffix() so a future correction path can't forget to
re-attach the @preset/<slug> routing suffix.
tarfile.open(archive_path, "w:gz") truncates the destination the instant
it opens. If tf.add() fails partway (disk full, permission loss,
interruption), whatever was previously at that path is gone — including
an existing profile or board export the caller chose to overwrite. This
is the same failure shape a7e7de6407 just fixed for the desktop gateway
file-save path, one commit earlier in the same window, but it was never
propagated to this shared archive-writing primitive even though board
export gained a new caller into it in that same window.
make_targz now writes into a sibling temp file (mkstemp, same directory
as the destination so the final step is a same-volume rename) and only
replaces the destination via os.replace() after the archive is fully
written and closed, mirroring the mkstemp+os.replace pattern already
used throughout this codebase (agent/secret_sources/_cache.py,
cron/jobs.py, gateway/status.py, etc). The temp file is unlinked on any
failure.
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes#83612.
Adapt five scenarios from @liuliu0223's regression suite in #99021:
- paired-mode (empty allowlist) positive paths for approval and
update-prompt cards, the DM breakage this fix resolves
- fail-closed rejection of clicks with an empty operator identity
- chat-mismatch rejection when an approval card is forwarded
The synchronous card-action handlers and the update-prompt resolver
authorized clicks with _allow_group_message(), which answers "may this
sender chat in this group?" — with group_policy=open it returns True
for everyone. The approval resolver already used the correct operator
gate (_is_interactive_operator_authorized), so the three code paths
disagreed: with an open group policy an out-of-allowlist click on an
update-prompt card was fully executed, and approval clicks returned a
resolved-looking card before being rejected asynchronously.
Authorize all three paths with _is_interactive_operator_authorized(),
which checks membership of admins ∪ allowed_group_users (wildcard and
the empty pairing-mode allowlist keep their existing allow semantics,
matching _admit's DM pairing default). A missing operator identity now
fails closed on the update-prompt resolver instead of skipping the
check.
Fixes#96045
The pre-release line-walk (#96601 salvage) moved the download/probe body
into install_node_line() and rejects an unstartable binary BEFORE
adoption via node_satisfies_build on the extracted tree. The sandboxed
driver now inlines all three functions, and the broken-node test pins
the stronger pre-adoption rejection instead of post-adoption cleanup.
install_node() picks the newest tarball out of
nodejs.org/dist/latest-v${NODE_VERSION}.x/ and installs it without ever asking
whether the binary inside is usable. That index currently serves
node-v26.8.0-<os>-<arch>.tar.xz -- a final-looking filename -- whose binary
reports v26.8.0-alpha.0.0.0. Node publishes the headers tarball named by
process.release.headersUrl only for final releases, so node-gyp cannot compile
against that build and every native module fails to install.
Probe the extracted tree before it replaces anything on disk, and fall back to
an older release line when the probe rejects it, instead of leaving the install
with an unbuildable runtime. Mirror the guard in node-bootstrap.sh, and let
_managed_node_tree_outdated() treat a pre-release tree as outdated so an
already-broken install heals itself -- the existing heal only fires below the
target major, and a pre-release sits above it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GrEaXSjvFoBXKxAHTjnUbS