Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).
POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.
Fixes#12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.
Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.
Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.
Fixes#9748
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes#9636
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.
A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.
Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.
Fixes#9516
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.
Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.
Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
Ports #63762 forward onto current main per teknium1's review.
refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).
1. refresh_launchd_plist_if_needed(): return False after retry
exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
of the success message.
Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.
Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.
5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.
Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.
Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
Leaving the context-length prompt blank in the custom-endpoint wizard said
"will auto-detect" and then went silent, so users could not tell whether their
endpoint runs on a detected window or the runtime's default fallback (which
shapes compression and prompt-cache behaviour). After the save prompt, run the
same resolver the runtime uses (with the endpoint's URL and key) and print
either "auto-detected N tokens" or "not detected — using the default N tokens".
Feedback only: the probe result is not persisted, and a failing probe never
blocks the save.
Fixes#2513. Approach from PR #2522 (@ygd58) and PR #85499 (@Luna161), both
written against the pre-decomposition wizard module.
Co-authored-by: Luna161 <268031236+Luna161@users.noreply.github.com>
A throttled GitHub fetch also yields index-metadata-without-bundle, so the new
stale-entry verdict would tell users a skill "no longer exists upstream" when
it does. Check the adapters' rate-limit flag first and keep the existing
rate-limit hint for that case (the keep_open review concern on #3261).
The per-search "results may be stale" note is dropped: it fires on every
skills.sh search whether or not anything is stale, and the install-time error
now names the condition precisely where it happens.
_resolve_source_meta_and_bundle already distinguishes index-hit-without-
files from unknown identifiers, but do_install printed the same generic
'Could not fetch' for both, sending users off to re-check spellings for
what is actually a stale skills.sh entry. Split the message, and add a
staleness caveat to do_search results from skills.sh.
Fixes#3259. Supersedes #3261 (stale since July — re-applied onto the
current _print_fetch_failure helper).
BotFather rejects setMyCommands descriptions containing em/en dashes
(U+2012-U+2015, U+2212). Fold them to ASCII hyphen at the two Telegram
sinks (telegram_bot_commands, telegram_menu_commands) so core, plugin,
and skill entries are all covered.
Fixes#2925.
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:
- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
rev-list --all --reflog, so a boundary the repair just restored (one a
reflog-only commit still needs) is never dropped again; previously the
production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
message body is prose; _batch_missing_parents stops at the blank line
ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
not --batch-all-objects: unrelated object loss (a deleted parent of a
locally-created commit) is no longer re-labelled as shallow history;
fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
shallow.lock, so a depth-1 fetch between read and write fails fast
instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
(_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
test file footguns fixed (encoding=, as_uri()) and the missing
repair->prune end-to-end regression added, mutation-checked.
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.
Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.
This complements PR #108290, which owns the prevention half.
Refs #108286
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
- /p/<profile>/line/webhook is verified with the NAMED profile's channel secret under
its runtime scope; another profile's secret is 401 at that URL; the bare path is
untouched; unknown profile / profile without the adapter is 404.
- A shared-listener adapter binds no port and records its /p/<profile>/ ingress_url
in runtime status; LINE media URLs use the shared prefix.
- Runner: a secondary's port-binders are constructed in shared-listener mode instead
of refusing the whole profile; api_server/webhook are skipped as mirrors.
- Dashboard: only the mirrored pair is refused (409) on a secondary.
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.
Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.
Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
Trims the nine salvaged tests to four, one per contract: SUDO_USER's default
home keeps the bare name across the unit sync; a home the unit does not pin
keeps its suffix (the #105525 guard); a bare unit pinning a profiles/<name>
home beats the profile branch; the production sync itself preserves the name.
Review follow-ups to the unit-anchored service identity.
`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.
Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.
Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.
The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.
Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.
Refs #108674
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).
The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.
Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.
The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.
Refs #108674
The turn finalizer now releases only its own generation
(`_release_running_agent_state(key, run_generation=...)`); two test doubles
were zero-kwarg lambdas and raised TypeError inside the finally.
_load_global_auth_store is memoised on (path, st_mtime_ns). A write-through
to the root (borrowed Codex cooldown clear, xAI/Anthropic root rotation)
followed by a fallback read in the same mtime tick — coarse-mtime
filesystems (exFAT, some network/overlay mounts) — kept serving the
pre-write store, so the resolve path could log "quota restored" and then
raise quota_exhausted from the stale memo. _save_auth_store(target_path=...)
now drops the memo.
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.
Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
With reads now inheriting the global-root pool, a profile hitting a stale
root cooldown probes quota, sees it restored, and calls
clear_codex_pool_quota_cooldowns() — which only ever edited the (empty)
profile store, so the next resolve raised quota_exhausted again forever.
Pick the store the same way agent/credential_pool.py persists borrowed
rows (_profile_owns_pool_provider / _borrowed_single_use_pool_root) and
lock/save against that path. The fallback test now also binds
profile-wins precedence; one new test pins the root write.
- tests/gateway/test_multiplex_residue_parity.py: served profile reads its own
sessions.*; per-turn bridge skips secondary scope; resolve_proxy_url reads the
routed scope (and does not borrow the default's on a miss); a stale served
turn never recreates an archived profile; MCP discovery runs per profile home.
- test_config.py: v43 migration removes multiplex_profile_allowlist.
- Allowlist tests deleted (feature removed) or rewritten to "served set = all
live profiles"; the unserved cases now use a tombstone / missing dir.
- MCP discovery fixtures converted to the per-home set/dict slot.
_spawn_hermes_action copied the dashboard's os.environ verbatim into every
detached `hermes ...` action. The dashboard runs inside the gateway and has
loaded its own profile's .env into the process environment, so
`hermes -p <other> gateway restart` (and every other dashboard-driven
profile action) started with the DEFAULT profile's platform credentials
and ports already present. load_hermes_dotenv does not override keys that
are already set, so the named profile's own .env could not displace them:
an A2A-only profile ended up claiming the default Discord bot token and
binding the default API server / BlueBubbles ports.
For actions carrying a profile selector (`-p X`, `--profile X`,
`--profile=X`), build the child env from the standard scrubbed subprocess
environment, drop _PROFILE_MANAGED_ENV_KEYS plus every key defined by the
dashboard/default profile's .env and its hydrated secret sources, and pin
HERMES_HOME to the target profile so the child's normal startup loads that
profile's .env. Only the leading selector is inspected: argv after the
subcommand may legitimately contain -p for a nested process. Actions
without a selector keep the historical environment byte-for-byte.
Test: the new case asserts the leaked keys are gone, benign keys survive,
and (by running the real dotenv loader in a fresh interpreter with the
captured env) that the target profile's values load without reviving any
default-profile value.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A profile served by the default multiplexer owns no gateway.pid / gateway_state.json,
so every surface that reads per-profile identity files called it stopped while the CLI
status surfaces (hermes -p X status / gateway status / cron status) said "running via the
default-profile multiplexer":
- `/api/status?profile=X` and `/api/messaging/platforms?profile=X` reported
gateway_running=false / state=None / "gateway_stopped" in the same body that listed X
under gateways[].served_profiles. The shared ladder `resolve_gateway_liveness` gains a
fourth rung for a named profile_dir: the live default multiplexer that records X in
served_profiles IS X's gateway (pid = multiplexer pid, runtime = its record, X's
`<X>:<platform>` entries re-keyed to the standalone shape).
- `POST /api/gateway/stop?profile=X` spawned `hermes -p X gateway stop`, which printed
"No gateway running for this profile" (exit 0) into the action log while the UI flipped
to stopped and the multiplexer kept serving X; `/api/gateway/restart?profile=X` spawned
a `-p X gateway restart` that only exits 78. start/stop now answer 409 with the
multiplexer explanation (one helper shared with the existing start refusal) and restart
targets the multiplexer, the process that actually serves X. A `--force`-started
separate gateway for X (own pid file) keeps normal per-profile management.
- CLI `hermes -p X gateway stop` refuses with exit 78 like run/start/install/restart when
X has no gateway of its own, instead of a contradictory exit-0 "not running".
Docs: multi-profile-gateways.md §1 and §5 describe stop + the dashboard behaviour.
noninteractive_git_env() now spawns `git config --get-all safe.directory` before
building the env. Eight tests fake subprocess.run/Popen with a fixed sequence of
expected git calls (update check, plugin pull, MCP install, bounded probe) and the
extra spawn tripped them in CI. An autouse fixture stubs the read to "no entries";
the two carve-out invariant tests opt back in with @pytest.mark.real_safe_directory
(and were confirmed to still exercise the real read: the ordering test would fail
against the stub).
contributors/emails: pry@privacydied.net -> privacydied (check-attribution).
Tests: fold the ambient GIT_CONFIG_KEY_n negative and the isolation-still-in-force
assertions into the ordering test, and drop the two positive-only cases it subsumes.
What remains pins the injected sequence to git's own `config -z --get-all` output and
asserts the real trust decision (named repo usable, unrelated cross-owner repo refused).
Cost: noninteractive_git_env() runs on every internal git call, including the banner
startup probe, and the two `git config` children added ~10 ms per call against ~0.2 ms
before. Memoise per process on the inputs that select the config files plus the
system/global candidates' mtimes, so an edit to ~/.gitconfig is picked up without a
restart (verified live: 128 -> 0 within one process after appending safe.directory).
Review of #107748 found the carry path serialized the user's trust policy
incorrectly, and the mistake could WIDEN trust rather than merely reformat it.
safe.directory is an ordered multi-valued protected setting: an empty value
resets every entry seen so far. That is the documented mechanism for revoking a
system-wide `safe.directory=*` and then naming only the repositories you
actually trust. The previous helper broke all three properties that make it work:
* read `--global` before `--system` (git's precedence is system, then global)
* dropped empty entries via `if entry:`, deleting the reset marker
* de-duplicated values, though it is a sequence and not a set
Reproduced with real git (2.55.0 here, 2.47.3 by the reviewer). Given
system `safe.directory=*`, global `safe.directory=` then `/trusted/only`:
real git, user's own policy -> rc=128 (dubious ownership)
previous head's sequence `/trusted/only`, `*` -> rc=0 (trust widened)
this head's sequence `*`, ``, `/trusted/only` -> rc=128 (matches real git)
So a user who had deliberately revoked a machine-wide wildcard silently got it
back inside Hermes's internal git calls. That contradicted the "read-only and
non-widening" claim the original change rested on.
Fix: read scopes lowest-precedence first (system, then global) and replay every
value verbatim -- no de-duplication, no dropping of empty reset markers. Read
with `git config -z` so a value containing whitespace or a newline stays the one
entry git reads it as, instead of being split into several bogus trust entries
by splitlines()/strip(). The trailing field of a -z stream is always empty and
is discarded; interior empty fields are real resets and survive.
Two invariant tests, both proven red on the previous implementation:
* the injected sequence equals `git config -z --get-all safe.directory` for
the same two config files, asserting git's documented reset shape
* an E2E trust decision under GIT_TEST_ASSUME_DIFFERENT_OWNER=1: the named
repo stays usable and an unrelated cross-owner repo is still refused,
proving the reset still revokes the wildcard
13 passed in tests/hermes_cli/test_noninteractive_git.py, 69 in
tests/security/test_gitspawn_config_injection.py and
tests/tools/test_checkpoint_manager.py.
`noninteractive_git_env()` blanks GIT_CONFIG_GLOBAL/SYSTEM to /dev/null so a
user's config cannot hang Hermes's internal git plumbing with pagers, hooks or
credential prompts. Sound intent, but it also discards `safe.directory` — and
git honours that key ONLY from global/system config (it is rejected from
repo-level config by design, so a hostile repo cannot self-authorise).
Result: every internal git call fails on a repo whose st_uid != geteuid():
$ hermes -w
✗ Failed to create worktree: fatal: detected dubious ownership in
repository at '/mnt/nas/py/repo'
That hits NFS/CIFS mounts without idmapping, shared checkouts, and containers
with a remapped uid. The user's own `git config --global --add safe.directory`
is correctly set and their interactive git works — Hermes throws the setting
away before git reads it, so the error's own suggested remedy can never fix it.
There is no config or env escape hatch: the blanking is unconditional.
Reproducer (any repo where the checkout uid differs from the caller's):
git rev-parse HEAD # works
GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null \
GIT_CONFIG_NOSYSTEM=1 git rev-parse HEAD # dubious ownership
Fix: read the user's real `safe.directory` values before the isolation is
applied, then re-inject them over the GIT_CONFIG_KEY_n channel, which survives
GIT_CONFIG_GLOBAL=/dev/null. Isolation is unchanged — global/system config stay
pointed at /dev/null and every hardening override still applies, since a later
key of the same name wins in git's config order.
Read-only and non-widening: only values already present in the user's own
config are carried, so this grants no trust they had not granted. Ambient
GIT_CONFIG_KEY_n injection is still stripped first, so a caller cannot launder
an attacker-controlled path in this way — covered by a regression test.
Tests: three cases in TestNoninteractiveGitEnv — entries carried past
isolation, no entries injected when the user configured none, and ambient
injection not trusted. All pin GIT_CONFIG_SYSTEM at an empty file, since a real
/etc/gitconfig on the test host can otherwise leak entries and mask the
assertions.
Verified on Arch Linux, git 2.x, repo on an NFSv4 mount (uid 1024 vs caller
1000): `git worktree add` under the patched env returns rc=0 where it
previously failed. 80 passed in tests/hermes_cli/test_noninteractive_git.py,
tests/security/test_gitspawn_config_injection.py and
tests/tools/test_checkpoint_manager.py, including the pre-existing assertions
that the isolation stays in force.
turn_in_flight read only the dashboard session table; an in-process cron run
never registers there, so the watchdog reported "no running turn" and exited
mid-job (tool calls then failed with "cannot schedule new futures after
interpreter shutdown", the execution was marked unknown, the slot lost). The
probe now also consults cron.scheduler.get_running_job_ids — the ledger the
gateway shutdown drain already uses.
Addresses #107485