- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.
Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.
- `hermes profile migrate-identity <old> <new>`: retries the migration —
delegates to the gateway control verb while a multiplexer is live, performs the
durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
naming the offending database on a collision, a lock, or a partial failure. Only
the name format and the existence of the new profile are checked; the old profile
directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
`click`: a failed second database raised NameError instead of printing its warning.
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
backup.py imports hermes_cli.profiles only lazily and profiles.py never imports backup, so
there was no cycle to justify two literals. LOCAL_RUNTIME_ROOT_DIRS now feeds both
backup._EXCLUDED_ROOT_DIRS and the clone-all root gate; one invariant test pins the identity.
The named-profile control passed names that did not exist on disk, so
_non_exportable_entries flagged them as vanished-mid-walk and the ignore set
came back non-empty; the assertion also compared a set to a list. Seed the
trees the assertion names and assert the ignore set is empty.
Trim the three contributor tests from #101340 to two invariants: the
root-only ignore test now also asserts that a named profile used as
source keeps its own models/ (the exclusion is gated on the default
root), and the end-to-end create_profile(clone_all=True) test stays.
Same assertions, two tests — the salvage bar.
Docs: profile-commands.md and profiles.md list models/, runtimes/ and
node/ among the --clone-all exclusions so users know why a clone from
the default profile does not carry the local-model weights.
43e67d872 (local models) put three machine-scoped trees under the default
HERMES_HOME — models/ (GGUF weights, tens of GB), runtimes/ (managed
llama.cpp binaries) and node/ (managed Node) — and taught backup.py to
exclude them (_EXCLUDED_ROOT_DIRS). `hermes profile create X --clone-all`
from the default profile did not follow: it copied all three into the new
profile, which never reads them (models_dir() resolves from the default
root only) and cannot use them (the binaries are re-downloaded on demand),
turning a clone into a tens-of-GB copy.
Add the three names to _CLONE_ALL_DEFAULT_EXCLUDE_ROOT. The set is gated
on the source being the default profile, so a named profile that really
carries a models/ directory of its own keeps it; and the exclusion applies
at the source root only, so a skill's nested models/ directory is copied
as user data. Kept as a separate literal from backup.py's set on purpose
(backup also matches profiles/<name>/ and importing it here would be a
circular import) with a comment tying the two together.
get_profile_dir("default") re-resolves the platform-native root at call time; when pytest's
basetemp sits inside the operator's Hermes home the per-test sandbox counts as "under the
native home" and these two tests overwrote the live config.yaml / .env / MEMORY.md with their
stubs. Write through the profile_env fixture's home instead.
Taken from #111096 by @ruochu88s (the autouse native-home override and the session tripwire
from that PR are not included; the basetemp relocation in tests/conftest.py closes the class).
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.
`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
Move _non_exportable_entries next to its first caller, fold the .pyc/.pyo
suffixes into it (the clone-all closure kept its own copy), and route the
last un-ignored profile copytree (the skills/ copy in _bootstrap_profile_dir)
through it. Cut the three repeated "sockets abort copytree" comments down to
the helper docstring. Split the clone-all special-file case into its own
POSIX-marked test so the cron-jobs assertion keeps its Windows coverage, and
use monkeypatch.chdir in the socket-binding test helper.
Route the --clone-all copytree ignore through _non_exportable_entries so a
live source profile holding a gateway or agent-browser socket (or a FIFO)
no longer aborts the clone with [Errno 6] No such device or address.
.pyc/.pyo and the root exclude sets keep their existing handling.
hermes_cli/profile_distribution.py:_copy_dist_payload is left alone: it
copies from a freshly extracted distribution archive (a staged temp tree),
never from a live profile, so it cannot meet a socket.
Extends test_clone_all_does_not_copy_cron_jobs to cover the clone path.
Gate the tombstone+notify on "a live default gateway has recorded a served set"
(recorded_served_profiles() is not None) rather than on the per-profile
_served_by_running_multiplexer probe: a multiplexer serves every dir under
profiles/, the signal is cheap, and the narrower probe falls back to config
derivation the CLI process cannot see. Trim the salvaged tests to two
invariants — ordering (unroute while the old home still exists and a stale
mkdir_under_hermes_home of it is refused; hot-serve after the move; no
tombstone left) and no-signal-without-multiplexer. Rollback on a failed move is
kept and covered by the same code path.
Builds on #109269 (xielevi). Fixes#109267.
get_profile_dir now raises ValueError for a name that is not a valid profile id
(#108346). The tools/MCP scoped-RPC wrapper resolved the profile inside a broad
try that mapped resolve-time errors to the handler's own fail code (or re-raised
for mcp.servers.*), so a traversal-shaped name produced a different error than a
missing profile. Map it to the same 4064 the other profile-scoped surfaces use,
and drop the salvaged change-detector test that froze the profiles-root path.
Widens #108348 (Adolanium) to its remaining sibling surface.
A WS 'profile' param like '../../foo' normalized to a path component that
escaped the profiles root, letting a connected client bind an arbitrary
existing directory as a profile home (state.db opened there, and session
delete chains into per-id file cleanup under <dir>/sessions/).
get_profile_dir now validates the canonical name against the profile id
regex before joining it under profiles/. The regex only, not the reserved
list, so pre-reserved-list dirs like profiles/hermes keep resolving.
Callers that probe existence (profile_exists, _profile_home, the 4064
resolvers) treat ValueError as 'not found'.
If old_dir.rename(new_dir) raises (cross-device EXDEV, permissions, a
racing writer) after the pre-move tombstone + unroute, the profile was
left tombstoned-but-present — enumeration treats it as deleted, so the
profile silently vanishes (worse than the ghost this PR fixes). Undo the
unroute on failure: clear the tombstone and re-notify the multiplexer to
re-serve the old name before re-raising. Adds a regression test (proven
red on the base of this branch).
A multiplexed secondary profile has no gateway.pid of its own, so
rename_profile's _check_gateway_running(old_dir) reported it stopped and
skipped teardown. Unlike delete_profile, rename never tombstoned the old
name nor notified the multiplexer, so at the moment old_dir.rename(new_dir)
ran the default gateway still held the old profile's adapters, cron ticker,
logging and SQLite handles. Those live components immediately re-mkdir'd the
old home (no .deleted tombstone -> mkdir_under_hermes_home does not refuse
it) and the periodic reconcile re-adopted the resurrected dir as a ghost
served profile.
Give rename the same unroute-before-mutate discipline delete already has:
when the old name is served by a live multiplexer, tombstone + notify before
the move so its adapters stop and handles release into old_dir; clear the
stale tombstone after the move; then notify for the new name to hot-serve it
(mirrors create). Non-multiplexed renames are untouched.
Fixes#109267
- tests/gateway/test_multiplex_residue_parity.py: served profile reads its own
sessions.*; per-turn bridge skips secondary scope; resolve_proxy_url reads the
routed scope (and does not borrow the default's on a miss); a stale served
turn never recreates an archived profile; MCP discovery runs per profile home.
- test_config.py: v43 migration removes multiplex_profile_allowlist.
- Allowlist tests deleted (feature removed) or rewritten to "served set = all
live profiles"; the unserved cases now use a tombstone / missing dir.
- MCP discovery fixtures converted to the per-home set/dict slot.
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.
Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.
--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
Creating a bot from the desktop dialog builds the profile tree but no
config.yaml, so the profile resolves no provider and its first turn dies
with "No LLM provider configured" — created, but unable to run. Every bot
made that way was dead on arrival.
Seed the active profile's model block at creation. It is a copy, not a
link: profiles stay independent islands and editing either afterwards never
touches the other. "Fresh" means fresh skills and SOUL, not unreachable.
External review (Fable) caught a real false-positive widening in the
original commit: the new argv[1] script-name check reused the loose
`script_name == "hermes" or script_name.startswith("hermes")` pattern
(copy-pasted from the exe_name check above it), but argv[1] can be ANY
user-invoked python script path when argv[0] is a bare interpreter --
unlike a directly-resolved executable name, where a false match on the
substring is rare. A user's own script named e.g. "hermes-notes.py" or
"hermes-unrelated-tool" run via `python3 <script>` would be misidentified
as the console-script shim and become killable by profile delete.
Match against the actual known console-script entry points instead
(pyproject.toml [project.scripts]: hermes, hermes-agent, hermes-acp),
stripping the script's extension before comparing.
Added 2 regression tests: one confirms the false-positive case is now
rejected (fails against the pre-fix loose-match code, confirmed via a
scripted revert), the other confirms the other two real entry points
(hermes-agent, hermes-acp) still match via the shebang-exec path.
Tests: tests/hermes_cli/test_profiles.py -- 158 passed (156 previous + 2
new).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two independent bugs let a deleted profile reappear / leave orphaned
resources on next launch:
1. hermes_cli/profiles.py's backend-process scanner required argv[0] to
resolve to an executable literally named "hermes". Electron's
pool-backend spawn resolves the hermes console-script shim's path and
execs it via the interpreter directly (python3 /path/to/hermes ...), so
argv[0] reports as "python3" and the scanner never matched the running
backend -- delete removed the profile's files but left its live backend
process running (still bound to a port via uvicorn), which
accumulates across repeated delete/recreate cycles.
2. The desktop sidebar's ProfileRail only refreshed its cached profile
list once, on mount, so a delete/create/rename from another surface
(another window, or the CLI) left a stale ghost entry until something
unrelated triggered a refetch. Note: a delete via this window's own
Manage-Profiles view already refreshes the shared $profiles atom
ProfileRail subscribes to (confirmed by reading refreshProfiles() and
handleConfirmDelete()) -- this fix only covers the cross-window/cross-
process staleness gap, not a duplicate of the already-merged
#57329's Manage-Profiles rail-refresh work.
Fix 1: recognize a python-interpreter argv[0] exec'ing a hermes-named
console-script shim via argv[1]. Fix 2: refresh the profile list on window
focus/visibilitychange, matching the existing pattern used elsewhere in
the sidebar (sidebar/index.tsx, use-background-sync.ts, star-map.tsx,
use-gateway-boot.ts all use the same focus+visibilitychange pattern).
## Related work already on main
PR #57329 (merged) fixed the *headline* symptom from issue #52279
(deleted profile respawns) via a different, non-overlapping mechanism:
routing profile-delete through the primary backend instead of spawning a
fresh pool backend, plus a separate recreation guard in
ensure_hermes_home() (#49435, merged) that makes a backend spawned into a
deleted profile's directory raise FileNotFoundError instead of silently
recreating it.
This PR is NOT a duplicate of that fix. Verified: even with both of those
merged, a backend process that survives because of gap #1 above still
holds a bound port via uvicorn -- it just can no longer resurrect the
profile directory. That's real resource-hygiene, not a symptom already
covered. Gap #2 touches a different file/component (ProfileRail /
profile-switcher.tsx) than #57329's rail-refresh half (which touched the
Manage-Profiles view's own $profiles.ts / index.tsx) and covers a
distinct staleness path (cross-window/cross-process, not same-window
delete-then-refresh).
Tests: tests/hermes_cli/test_profiles.py -- 156 passed (existing +
regression coverage for the argv[0] python-interpreter detection case).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Same behavior, same coverage, less boilerplate (test file 1691 -> 1512
lines; PR diff unchanged in semantics).
Production (mechanical only):
- Extract _strip_hermes_owned_pythonpath_and_runtime_markers(): the three
builders (_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env)
ran the identical strip-then-pop-markers sequence in the same order
(ordering is load-bearing for VIRTUAL_ENV validation); the helper makes
that explicit once instead of three times.
Tests:
- Non-owned preservation: 11 single-shape tests -> one parametrized matrix
(user/Nix/other-version/python2.7/pythonX.Y-contained/raw spelling/empty
component/empty PYTHONPATH) + one runtime-shaped matrix (other-version SP,
venv-SP descendant, repo direct child, repo deep child).
- Owned stripping: venv SP, repo root (independent parents[2] computation),
duplicates, all-owned key removal, mixed ordering -> one matrix.
- Builder integration: _make_run_env/_sanitize_subprocess_env/
hermes_subprocess_env venv-SP stripping -> one parametrized test;
same for the four PYTHONHOME builders (incl. build_subprocess_env).
- Junction: same-named non-owned negative control now covers both the
configured-root location and an unrelated location; shared
_physical_repo_root helper; profile resolution matrix (root->named,
profile-shaped->named no nesting, profile-shaped->default, custom root).
- Every independent proof preserved: home-level junction, repo-level
junction, profile interaction, negative identity control, uv-base lexical
VIRTUAL_ENV, validated/unrelated VIRTUAL_ENV, no-scrub escape hatch,
#84500 same-env/external-env composition, PYTHONHOME removal, real
Windows-only semantics, POSIX fail-closed backslash paths.
The junction fix made resolve_profile_env preserve the configured
HERMES_HOME spelling as the launch root. Cover the four pre-existing
resolution invariants so the spelling-preservation never regresses them:
- root env + named profile -> <root>/profiles/<name>
- profile-shaped env + named profile -> <root>/profiles/<name> (no nesting)
- profile-shaped env + default -> <root>
- custom root env never falls back to the platform default
Plus existence/validation semantics (missing named profile still raises
FileNotFoundError) and the unset-env fallback contract.
many tests patched sys.platform or a module's _IS_WINDOWS flag, then
ran on linux ci. the patch selects the branch under test, but the host
does not have the behavior the branch exists for. the test proves the
patch, not the platform. some gated assertions never ran on any host.
this commit adds three markers: linux_only, macos_only, windows_only.
a conftest hook skips a marked test on the other hosts, with a clear
reason. no test fakes a host now. two documented fakes remain
(android/termux, freebsd) because no ci runner exists for them.
each fake site got one of four treatments:
- gate it: the real host supplies the platform; mocks cover real
dependencies only, never host identity
- patch the module's own probe when the subject is the probe's consumer
- assert against the real host when the fake stood in for any non-x host
- delete the patch when it set the value the host already has
bare skipif(sys.platform != ...) guards became markers too. the lane
model skips these on linux and never imports them on windows, so they
ran on no host. platform parametrize tables are now one marked test
per os.
running on real hosts found real errors: a chrome-sandbox failure in
test_gui_command that main hides, and two windows failures fixed here.
the agents.md testing section now documents the policy.
`write_profile_meta` and `hermes skin set` are both read-modify-write
helpers that rewrite a user-visible YAML file with a bare truncating
write, bypassing `utils.atomic_yaml_write` — the shared helper whose
docstring states that "every destructive file rewrite in the codebase
shares one implementation".
Both read halves swallow a parse error and fall back to `{}`, so a
truncated file is not transient corruption. The next call reads `{}` and
silently, permanently drops every field the caller did not explicitly
pass:
* `write_profile_meta` promises "unspecified fields preserve existing
values". After an interrupted write, a follow-up call that only sets
`description_auto` erases the profile's `description` — it vanishes
from `hermes profile list` and never comes back.
* `_skin_set` exists so that "changing one token never disturbs the rest
of the look". `path.write_text(...)` neither fsyncs nor swaps
atomically, so a crash or power loss can leave `<skin>.yaml`
zero-length; the next tweak then rewrites from empty and the whole
palette is gone. The gateway's skin watcher repaints live surfaces
from this file within ~1s, so a half-written file is observable.
Routing both through `atomic_yaml_write` gives temp file + fsync +
`atomic_replace`, which also preserves a symlinked target (GitHub
#16743) and restores owner/mode, and emits emoji descriptions as real
UTF-8 instead of `\UXXXXXXXX` escapes (GitHub #51356).
Supersedes #51808, which fixed the unicode-escaping symptom alone by
adding `allow_unicode=True` to the same `yaml.safe_dump` call.
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.
Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
(DEFAULT_CONFIG smart-approval leaked in) — pinned approval
mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
silently probed live provider catalogs (~2s/test) — stubbed
cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
MagicMock agents — interrupt.side_effect now clears _running_agents
(22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
(24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
with rationale comment — the watchdog-dump assertion needs the loop
blocked well past deadline+grace under parallel load (flaked once in
the 40-worker verification run at a 0.2s margin).
Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
`hermes profile export default` crashed with `shutil.Error` when
HERMES_HOME pointed outside ~/.hermes (common in Docker deployments)
and the workspace contained broken symlinks. Two root causes:
1. `copytree` defaults to `symlinks=False` and follows link targets;
broken ones crash. #58397 (liuhao1024) drafted a minimal
`symlinks=True` flag fix; this PR adopts that change.
2. `copytree` was invoked against the entire HERMES_HOME root (which
doubles as cwd in Docker layouts). The post-hoc blacklist at
`_DEFAULT_EXPORT_EXCLUDE_ROOT` is a fixed-length enumerate-and-pray
list that can't anticipate every unrelated sibling directory
(`x11-dev/`, etc.). Replaced with a positive allow-list at
`_DEFAULT_EXPORT_INCLUDE_ROOT` enumerating the known Hermes profile
artifacts (config, persona, skills, cron, scripts, sessions,
plugins, memories, knowledge, preferences). Sensitive runtime
surfaces (`state.db`, `logs/`, auth files, other profiles) are
intentionally not in the allow-list so the export stays a
portable, credential-free snapshot of the user-facing surface —
which means the existing `test_export_default_excludes_infrastructure`
regressions remain green.
Adds two regression tests:
* test_export_default_uses_allowlist_for_unrelated_dirs — >x11-dev<
sibling directories must not leak into the archive.
* test_export_default_handles_broken_symlinks — symlinks inside
allowed artifacts survive instead of crashing the export.
closing that PR as superseded once this lands.
Closes#58394
shutil.copytree() defaults to symlinks=False which follows symlinks and
crashes on broken ones. In Docker/custom HERMES_HOME deployments,
unrelated directories may contain stale symlinks that break export.
Add symlinks=True to both copytree() calls in export_profile() so
broken symlinks are preserved as symlink entries in the archive.
Fixes#58394
delete_profile stopped only the process named in gateway.pid, but a Desktop
app spawns a headless `serve`/`dashboard` backend per profile that holds the
profile's SQLite connection open and keeps writing sessions/WAL/sandbox files.
That backend is never in gateway.pid, so a CLI `hermes profile delete` run
while the Desktop app is up left it writing into the tree — rmtree's final
rmdir then failed with ENOTEMPTY (#47368 "Bug 2"), and pre-guard it also
resurrected the directory.
- _profile_bound_backend_pids(): find running Hermes backends bound to this
profile via a `--profile <name>` selector or a HERMES_HOME env resolving to
the profile dir. Tightly scoped — current-user only, backend subcommands
(serve/dashboard/gateway) only so an interactive chat is never killed, and
never this process or its ancestors.
- _stop_profile_backends(): terminate them (graceful, then force), best-effort
so it can never make delete worse.
- _rmtree_with_retry(): a few spaced retries absorb the ENOTEMPTY / Windows
file-lock race from a just-terminated writer's in-flight -wal/-shm/sandbox
writes instead of failing the whole delete on a race the next attempt wins.
Complements the recreation guard (deleted profiles no longer reappear) and the
Desktop teardown-before-delete flow; this is the CLI-side convergence fix for a
delete run while a Desktop-managed backend is live.
Part of #47368.
`hermes profile alias <profile> --name <custom>` accepted arbitrary
strings and used them verbatim as a filename under ~/.local/bin. Because
normalize_profile_name only lowercases/strips (no regex gate), a value
like `../../.bashrc` escaped the wrapper directory and clobbered
arbitrary user-writable files. remove_wrapper_script had the same sink.
Add validate_alias_name (reusing the profile-id regex, which forbids
`/`, `.`, and `..`) and wire it into check_alias_collision,
create_wrapper_script, remove_wrapper_script, and the CLI alias action so
the rejection surfaces a clear "Invalid alias name" error instead of
silently writing or unlinking outside the wrapper dir.
Co-authored-by: Gutslabs <gutslabsxyz@gmail.com>
Co-authored-by: Xowiek <xowiekk@gmail.com>
PR #52151 hardened the runtime-status liveness check to trust a readable
live process command line over stale gateway_state.json argv, so a recycled
PID now owned by an s6 supervisor no longer counts as a running gateway.
That fix is correct but incomplete for the reported symptom: the web
dashboard showed a named profile's gateway green while
`hermes -p <name> gateway status` showed it stopped. Two further issues:
1. Cross-profile PID reuse. In per-profile Docker supervision, one profile's
stale `gateway_state.json` can record a PID the OS later recycled onto a
DIFFERENT profile's live gateway. That PID's command line still
`looks_like_gateway`, so the dead profile was reported running. The
recorded argv has its `-p <name>` selector stripped in-process by
`_apply_profile_override`, so it cannot disambiguate; the live `/proc`
cmdline still carries it. `get_runtime_status_running_pid` now accepts an
`expected_home` and validates the live command line belongs to THAT
profile (mirroring `hermes_cli.gateway._matches_current_profile`, the
logic the CLI scan path already uses — which is why the CLI was correct).
`_check_gateway_running` passes the enumerated profile dir.
2. The existing regression test `test_gateway_running_check_falls_back_to_
runtime_state` used the live pytest PID with a gateway-shaped record; once
the live cmdline became authoritative it no longer looked like a gateway.
Updated to mock the live cmdline to the real separate-process scenario it
describes.
The active-profile path (`get_running_pid`) is intentionally left unscoped:
it is lock-verified and any live gateway cmdline is acceptable there. Multiplex
mode is unaffected — `running` state is only ever written to a gateway's own
home, never a secondary served profile's.
Adds coverage for: cross-profile PID reuse (named + default), matching
profile cmdline (`-p`, `--profile`, explicit HERMES_HOME=), the bare default
gateway, and the unreadable-cmdline cross-platform fallback. Each new
cross-profile assertion fails without the profile scope and passes with it.
Co-authored-by: helix4u <4317663+helix4u@users.noreply.github.com>
Selective --clone / --clone-from / --clone-config copied .env but not
auth.json, silently dropping the credential pool — including OAuth tokens
(Anthropic `claude /login`, Codex, xAI) that never land in .env. A profile
cloned from an OAuth-authenticated default therefore resolved a different
provider (or none) than the source under provider: auto. --clone-all already
carried auth.json via the full copytree; only the selective path missed it.
Add auth.json to _CLONE_CONFIG_FILES and tighten it to 0o600 after copy,
matching .env semantics.
The dashboard Profiles view showed "Gateway stopped" for a gateway that
is in fact running — while the sidebar status strip and `hermes gateway
status` (CLI) both correctly showed it running. Reported on v0.17.0
running the gateway + dashboard in one Docker container.
Root cause: three liveness surfaces with three detection strengths, all
reading the same `gateway.pid`:
- `hermes gateway status` -> find_gateway_pids() (process-table scan)
- sidebar /api/status -> get_running_pid() + gateway_state.json PID
fallback + health-URL probe
- Profiles view -> _check_gateway_running() = get_running_pid()
ONLY, no fallback
`get_running_pid()` short-circuits to None the moment the runtime lock
(`gateway.lock`) doesn't register as held by the *calling* process —
which is always true when the reader is a separate process from the
gateway (the dashboard is its own s6 service in the container), and also
for any launch-service-managed gateway that left a fresh
`gateway_state.json` but no live PID file. So the Profiles view alone
reported the live gateway as stopped.
Fix: give _check_gateway_running the same fallback the sidebar already
has — after the pid-file/lock check misses, validate the PID recorded in
that profile's gateway_state.json against the live process table via the
existing get_runtime_status_running_pid(). read_runtime_status() gains an
optional path arg so a profile's state file can be read without mutating
the process-global HERMES_HOME (preserving the contextvar-based profile
isolation the dashboard relies on). Backward compatible: every existing
caller passes no argument.
Tests: a regression test that fails pre-fix (live gateway, lock check
returns None -> must still report running) and a guard test that a
'stopped' state file is never reported running even with a live PID.
Foundations for serving multiple profiles from one gateway process, inert
when off:
- gateway.multiplex_profiles config flag (default false), round-trips through
GatewayConfig and load_gateway_config (top-level + nested gateway.* form).
- hermes_cli.profiles.profiles_to_serve(multiplex): the single chokepoint for
which (profile, HERMES_HOME) pairs the gateway serves. Lightweight dir scan;
active-profile-only when off, default + all named profiles when on.
- build_session_key gains a profile= namespace slot. Default/None reuse the
historical 'agent:main:...' literal BYTE-IDENTICALLY (no session migration,
positional parsers unaffected); a named profile becomes 'agent:<profile>:...'
so two profiles on the same platform/chat never collide.
- SessionStore._resolve_profile_for_key + _session_key_for_source fallback
resolve the namespace from the flag (legacy when off, active profile when on).
Tests: byte-identical-when-off (parametrized), namespace isolation, positional
layout preserved, config round-trip, profiles_to_serve enumeration.
Profiles created before #44792 have no .env. Now that the Channels/Keys
endpoints are profile-scoped (no os.environ fallback), those profiles
would show everything as unconfigured. hermes update now copies the
default install's .env into each named profile that lacks one (0600,
never overwrites, placeholder fallback when the root has no .env), so
existing users keep the credentials they were effectively running with.
--clone-all copied the source profile's state.db, sessions/, backups/,
state-snapshots/, and checkpoints/ into the new profile. These are
per-profile history: a 49GB copy in practice (15GB snapshots + 11GB
backup archives + 16GB state.db + 6.4GB sessions), and restoring a
copied backup inside the clone would resurrect the SOURCE profile's
state. A clone is a fresh workspace; history stays with the source.
New _CLONE_ALL_HISTORY_EXCLUDE_ROOT set, applied at root level for ANY
source profile (named profiles accumulate the same artifacts), unlike
the default-gated infrastructure excludes. Nested same-name dirs still
copy. Docs and the post-create CLI message updated to match; profile
export / hermes backup remain the full-history paths.
Two halves of the same community report (dashboard Profile Builder):
1. A fresh dashboard/CLI-created profile got no .env file unless cloned,
so it silently inherited API keys and messaging tokens from the shell
environment / root install. create_profile() now seeds a placeholder
.env (0600) for non-clone profiles, matching the SOUL.md seeding.
2. The Channels endpoints (/api/messaging/platforms GET/PUT/test) were
not profile-scoped: they read/wrote the dashboard process's own .env
via load_env()/save_env_value() regardless of the global profile
switcher. They now accept the standard optional profile param (body
beats query on the PUT, matching other scoped writes) and run inside
_profile_scope(). When scoped, the payload no longer falls back to
os.environ or load_gateway_config()'s env-override layer — both carry
the ROOT install's credentials and would misreport them as the
profile's. /api/messaging/platforms added to PROFILE_SCOPED_PREFIXES
so the sidebar switcher scopes the Channels page automatically.
profile list and profile show assumed the wrapper script is always named
after the profile (wrapper_dir / name). When a custom alias exists — e.g.
`hermes profile alias steve --name qiaobusi` creates ~/.local/bin/qiaobusi
pointing at `hermes -p steve` — the display silently showed the profile
name (or nothing) instead of the alias the user actually typed.
The custom-alias *creation* path (create_wrapper_script(name, target)) was
added later; the *display* path was never updated to match.
Add find_alias_for_profile() — a reverse lookup that scans the wrapper dir
for our own wrappers (alias-named file containing 'hermes -p <profile>'),
prefers a custom alias over the profile-named one, strips .bat on Windows,
and sorts for deterministic output. Populate ProfileInfo.alias_name and wire
it into the three display sites (profile describe, list, show).
Credit: salvages the intent of #11506 by wss434631143, reimplemented on
current main against the post-#11506 custom-alias (--name/target) mechanism.
Tests: 6 new (profile-named, custom-name, none, unrelated-file rejection,
windows .bat strip, list_profiles surfacing). All 123 in test_profiles pass.
E2E verified against the real CLI for both custom and profile-named aliases.
Self-hosted Honcho setup had four sharp edges:
- local/cloud URLs ending in /vN double-prefixed by the SDK (/v3/v3/... 404)
- authenticated local servers had no setup prompt for a JWT/bearer token
- profile-derived host keys could be dot-containing workspace IDs Honcho rejects
- memory-provider config files with API keys written world-readable per umask
This keeps existing behavior but makes those paths safer:
- strip a trailing /vN version segment from any configured baseUrl before SDK
init (the SDK's route builders always prepend their own version prefix);
auth-skipping stays loopback-only
- add an optional local JWT/bearer prompt in honcho setup, stored under
hosts.<host>.apiKey
- derive new profile host keys with underscores, still reading legacy
hermes.<profile> blocks
- write memory-provider config files atomically with 0600 via a shared
utils.atomic_json_write(mode=) arg (honcho/hindsight/mem0/supermemory)
- skip honcho.json parsing in gateway cache-busting unless Honcho is the active
memory provider; memoize by honcho.json mtime when active
- bust the gateway agent cache on memory.provider change
- add a hermes memory setup <provider> one-liner so fresh installs can configure
a named provider without the picker (the per-provider hermes <provider>
subcommand only registers once that provider is active)
Closes#20688, #29885, #26459, #30246, #33382, #32244.
Co-authored-by: BROCCOLO1D
The profile alias --name path in main.py rewrote the wrapper with a
hardcoded #!/bin/sh script right after create_wrapper_script(), clobbering
the .bat on Windows and reintroducing the exact bug for custom aliases.
create_wrapper_script() now takes an optional target so the alias file is
named after the alias while the -p content references the profile — one
platform-aware code path, no post-hoc rewrite.
On Windows, hermes profile create produced a #!/bin/sh script that the
shell cannot execute. Now creates a .bat file with @echo off + %* on
Windows, and keeps the POSIX shell script on macOS/Linux.
Also fixes check_alias_collision to use 'where' instead of 'which' on
Windows, and remove_wrapper_script to find .bat files.
Fixes#34708
Remove unused imports (F401) and duplicate/shadowed import
redefinitions (F811) across the codebase using ruff's safe
autofixes. No behavioral changes -- imports only.
- ~1400 safe autofixes applied across 644 files (net -1072 lines)
- __init__.py re-exports preserved (excluded from F401 removal so
public re-export surfaces stay intact)
- Re-exports that are imported or monkeypatched by tests but look
unused in their defining module are kept with explicit # noqa:
F401 (gateway/run.py load_dotenv; run_agent re-exports from
agent.message_sanitization, agent.context_compressor,
agent.retry_utils, agent.prompt_builder, agent.process_bootstrap,
agent.codex_responses_adapter)
- Unsafe F841 (unused-variable) fixes deliberately skipped -- those
can change behavior when the RHS has side effects
- ruff lints remain disabled in pyproject.toml (only PLW1514 is
selected); this is a one-time cleanup, not a config change
Verification:
- python -m compileall: clean
- pytest --collect-only: all 27161 tests collect (zero import errors)
- core entry points import clean (run_agent, model_tools, cli,
toolsets, hermes_state, batch_runner, gateway)
- static scan: every name any test imports directly from an edited
module still resolves