Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
backup.py imports hermes_cli.profiles only lazily and profiles.py never imports backup, so
there was no cycle to justify two literals. LOCAL_RUNTIME_ROOT_DIRS now feeds both
backup._EXCLUDED_ROOT_DIRS and the clone-all root gate; one invariant test pins the identity.
43e67d872 (local models) put three machine-scoped trees under the default
HERMES_HOME — models/ (GGUF weights, tens of GB), runtimes/ (managed
llama.cpp binaries) and node/ (managed Node) — and taught backup.py to
exclude them (_EXCLUDED_ROOT_DIRS). `hermes profile create X --clone-all`
from the default profile did not follow: it copied all three into the new
profile, which never reads them (models_dir() resolves from the default
root only) and cannot use them (the binaries are re-downloaded on demand),
turning a clone into a tens-of-GB copy.
Add the three names to _CLONE_ALL_DEFAULT_EXCLUDE_ROOT. The set is gated
on the source being the default profile, so a named profile that really
carries a models/ directory of its own keeps it; and the exclusion applies
at the source root only, so a skill's nested models/ directory is copied
as user data. Kept as a separate literal from backup.py's set on purpose
(backup also matches profiles/<name>/ and importing it here would be a
circular import) with a comment tying the two together.
`hermes_cli/profiles.py` imported SYNC_MANIFEST_NAME from agent_import_sync at
module top, which pulled yaml/utils into every startup that resolves a profile;
the import now happens only inside _bootstrap_profile_dir when --sync-imports
is used (verified: importing hermes_cli.profiles no longer loads
hermes_cli.agent_import_sync).
`profile create --clone-all --sync-imports` accepted the flag (a full copy
carries import-sync.json regardless) but printed no notice; the hint is now
printed for both --clone and --clone-all.
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.
`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
rename_profile moves the profile directory while this process may still
hold the cached per-profile mcp-stderr.log handle (left behind by a
completed probe or a running server). On Windows a directory containing
an open file cannot be renamed, the same WinError class delete_profile
now avoids. On other platforms the stale handle stayed cached under the
old home key, so a later probe on the renamed profile opened a second
handle and a new profile re-created under the old name wrote its MCP
stderr into the renamed profile's log. Release the scoped handle next to
the multiplexer unroute, mirroring delete_profile.
Review finding: rename_profile missed the sibling surface of the
delete_profile handle release.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
Move _non_exportable_entries next to its first caller, fold the .pyc/.pyo
suffixes into it (the clone-all closure kept its own copy), and route the
last un-ignored profile copytree (the skills/ copy in _bootstrap_profile_dir)
through it. Cut the three repeated "sockets abort copytree" comments down to
the helper docstring. Split the clone-all special-file case into its own
POSIX-marked test so the cron-jobs assertion keeps its Windows coverage, and
use monkeypatch.chdir in the socket-binding test helper.
Route the --clone-all copytree ignore through _non_exportable_entries so a
live source profile holding a gateway or agent-browser socket (or a FIFO)
no longer aborts the clone with [Errno 6] No such device or address.
.pyc/.pyo and the root exclude sets keep their existing handling.
hermes_cli/profile_distribution.py:_copy_dist_payload is left alone: it
copies from a freshly extracted distribution archive (a staged temp tree),
never from a live profile, so it cannot meet a socket.
Extends test_clone_all_does_not_copy_cron_jobs to cover the clone path.
Named-profile export staged its copy with an ignore callable that only
excluded credential files, so any Unix socket in the profile (e.g. a stale
agent-browser control socket under home/.agent-browser/) made
shutil.copytree collect "[Errno 6] No such device or address" and raise
shutil.Error, failing the entire export. The default-profile export
already excluded *.sock by suffix, but a socket without that suffix (or a
FIFO, or a device node) failed it the same way.
Extract the universal exclusions into _non_exportable_entries(), which
keeps the __pycache__/*.sock/*.tmp name rules and additionally drops any
entry that is not a regular file, directory, or symlink (os.lstat mode
check), and use it in both the default and named export branches.
Gate the tombstone+notify on "a live default gateway has recorded a served set"
(recorded_served_profiles() is not None) rather than on the per-profile
_served_by_running_multiplexer probe: a multiplexer serves every dir under
profiles/, the signal is cheap, and the narrower probe falls back to config
derivation the CLI process cannot see. Trim the salvaged tests to two
invariants — ordering (unroute while the old home still exists and a stale
mkdir_under_hermes_home of it is refused; hot-serve after the move; no
tombstone left) and no-signal-without-multiplexer. Rollback on a failed move is
kept and covered by the same code path.
Builds on #109269 (xielevi). Fixes#109267.
A WS 'profile' param like '../../foo' normalized to a path component that
escaped the profiles root, letting a connected client bind an arbitrary
existing directory as a profile home (state.db opened there, and session
delete chains into per-id file cleanup under <dir>/sessions/).
get_profile_dir now validates the canonical name against the profile id
regex before joining it under profiles/. The regex only, not the reserved
list, so pre-reserved-list dirs like profiles/hermes keep resolving.
Callers that probe existence (profile_exists, _profile_home, the 4064
resolvers) treat ValueError as 'not found'.
If old_dir.rename(new_dir) raises (cross-device EXDEV, permissions, a
racing writer) after the pre-move tombstone + unroute, the profile was
left tombstoned-but-present — enumeration treats it as deleted, so the
profile silently vanishes (worse than the ghost this PR fixes). Undo the
unroute on failure: clear the tombstone and re-notify the multiplexer to
re-serve the old name before re-raising. Adds a regression test (proven
red on the base of this branch).
A multiplexed secondary profile has no gateway.pid of its own, so
rename_profile's _check_gateway_running(old_dir) reported it stopped and
skipped teardown. Unlike delete_profile, rename never tombstoned the old
name nor notified the multiplexer, so at the moment old_dir.rename(new_dir)
ran the default gateway still held the old profile's adapters, cron ticker,
logging and SQLite handles. Those live components immediately re-mkdir'd the
old home (no .deleted tombstone -> mkdir_under_hermes_home does not refuse
it) and the periodic reconcile re-adopted the resurrected dir as a ghost
served profile.
Give rename the same unroute-before-mutate discipline delete already has:
when the old name is served by a live multiplexer, tombstone + notify before
the move so its adapters stop and handles release into old_dir; clear the
stale tombstone after the move; then notify for the new name to hot-serve it
(mirrors create). Non-multiplexed renames are untouched.
Fixes#109267
Post-merge review of #109502 (gaoanze888) on current main, findings 1-3 and 5-11
(finding 4, the migrate manifest ordering, was already fixed by f9e47aa6fe).
Why:
- `--clone-all` used copytree(symlinks=True); a symlinked source `.env` was then
edited THROUGH the link by the channel strip, deleting the SOURCE's bot token.
Root files the clone edits (.env, config.yaml, auth.json, SOUL.md) are now
materialized as private copies before any write.
- The final profiles/<name> existed during the copy; the multiplexer rescans
profiles/ on create (#109239) and could adopt the half-copied tree and start
adapters on credentials not yet stripped. Clones are built in
profiles/.<name>.staging-<pid> (a leading dot never matches _PROFILE_ID_RE, so
profiles_to_serve never lists it) and published with one os.rename after the
strip; a failed create removes the staging tree.
- The live-multiplexer refusal for --clone-channels lived only in the CLI; REST
(POST /api/profiles) and the TUI (profiles.create) bypassed it. It now lives in
create_profile as profile_channels.clone_channels_refusal, raising ValueError
which every surface already maps to a 400 / 4062. --clone-channels without a
clone flag is an error instead of a silent no-op.
- Channel inventory is ownership-based, evaluated in the SOURCE profile's plugin
scope: private platform plugins under <source>/plugins/ contribute their keys
(previously discovered under the ambient HERMES_HOME); GATEWAY_ALLOW_ALL_USERS /
GATEWAY_ALLOWED_USERS and GATEWAY_RELAY_ID/SECRET/DELIVERY_KEY are channel
settings; alias prefixes SUPPLEMENT the canonical <PLATFORM>_ prefix
(WECOM_DM_POLICY, SMS_WEBHOOK_PORT now stripped). Prefixes shared with tools
(HASS_, TWILIO_, EMAIL_) are stripped only when the source's gateway would run
that adapter (enabled in config, or complete credentials and not explicitly
disabled); their allowlist/port keys are always channel-only.
- --clone-all state removal handled files only; Google Chat's directory-shaped
google_chat_user_tokens/ survived. Directories are removed too, without
following a copied symlink.
Two kill/relaunch predicates decided identity by argv substring, the bug class root AGENTS.md
forbids: hermes_cli/dashboard_procs.py::_is_desktop_local_serve_cmdline (`"serve" not in cmd`,
on the orphan-reap KILL path) and hermes_cli/update_cmd_windows.py::_is_backend_argv
(`" serve" in argv_low`, in the very file that defines _hermes_holder_subcommand). Both now ask
the canonical token classifier; host/port are read as flag values, not substrings.
hermes_cli/profiles.py::_check_gateway_running open-coded rungs 1/3 of
gateway.status.resolve_gateway_liveness and skipped the multiplexer rung; it is now that ladder
scoped to the profile dir (pid probe keeps cleanup_stale=False so a probe for another profile
never unlinks its PID file). The gateway/status.py ladder itself is untouched.
Behavior change: `hermes kanban --preserve-cache --host 127.0.0.1 --port 0` and
`-m dashboard serve`-style argv are no longer classified as serve backends (never killed /
relaunched as one); a named profile served by the live default multiplexer now reads as
running from _check_gateway_running (previously only via the separate
_served_by_running_multiplexer OR at some call sites).
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).
Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.
`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.
The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
The multiplexing default gateway now serves default + every live named profile
under profiles/. profiles_to_serve(multiplex=True) is a pure directory read
(tombstoned profiles skipped, never mkdir); every reader — gateway served set,
/p/<profile>/ prefixes for api_server + webhook, the named-profile standalone
guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a
running gateway is unchanged) — drops the allowlist parameter.
Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG,
GatewayConfig and the top-level yaml bridge no longer carry it.
BREAKING: anyone who set an allowlist now has their excluded profiles served.
Archive or delete a profile you do not want served (Teknium approved).
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.
Co-authored-by: Cursor <cursoragent@cursor.com>
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.
Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.
--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.
Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.
Salvage of #69118 rebased onto the shared helper (which post-dates it).
Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
`hermes profile delete` read the target profile's gateway.pid raw and
SIGTERMed it. When that pid file was poisoned by a sibling profile's gateway
(the #89315 shape), deleting profile A killed profile B's running gateway.
- gateway/status.py: `_pid_record_belongs_to_profile()` helper — a pid
record whose recorded home differs from the expected profile home is not
ours; legacy records without a home prove nothing and are left alone.
- hermes_cli/profiles.py: `_stop_gateway_process` refuses (and says so)
when the record belongs to another profile; still stops its own gateway.
The stop/restart paths in hermes_cli/gateway.py did not need a guard:
`get_running_pid()` already filters cross-profile records and unlinks the
poisoned pid file before any kill can happen — verified live; the test for
that path now pins the real contract (returns False, other process alive,
poisoned pid file gone).
Live repro (unpatched main): `_stop_gateway_process(tim_home)` -> "Gateway
stopped (PID ...)" and the OTHER profile's process exits -15. After: "Refusing
to stop PID ..." and the process stays alive. 8 tests; sabotage (guard
removed) fails 1.
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:
1. `hermes profile create --clone-all` and the dashboard/TUI
`mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
`providers.<id>` device-code blocks, and the PKCE singleton file; API
keys are still copied. The clone reads the root grant through the
existing credential-pool root fallback.
2. A named profile with no local rows BORROWS the root grant via
`read_credential_pool()`'s fallback, but every persist
(`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
rows into the profile's own auth.json — materializing a fork on the first
rotation. `persist_pool_entries()` now routes borrowed single-use rows
back to the root store (update-only, under the root lock; never falls
back to a local copy). A borrowed `hermes_pkce` rotation commits its
singleton to the root `.anthropic_oauth.json`, the borrower never prunes
root-seeded rows it cannot see the backing file for, and
`hermes -p <profile> auth add` persists only the profile's own rows.
Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.
Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).
Closes#100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Final-diff review findings (/simplify-code on the full 3-commit stack):
- Temp-dir fallback uses a per-uid name (hermes-profile-exports-<uid>) and
get_profile_export_path refuses a pre-existing symlink or a directory
owned by another user — a fixed /tmp/hermes-profile-exports is a
predictable shared path a local attacker could pre-create to receive the
secret-bearing archive. Regression test mutation-checked.
- Fail-closed message reworded interface-neutrally (the web API surfaces it
as HTTP 400 detail where '-o' alone made no sense).
- Docs now cover the temp fallback and the fail-closed refusal.
- _profile_export_directory(): when the managed store, the home-sibling
store, AND the temp dir all resolve inside Git checkouts, raise a clear
ValueError instead of warning and proceeding — a stderr warning would not
stop a scripted export from staging a secret-bearing archive in a source
tree, which is the exact #92457 incident class. All three callers already
surface ValueError cleanly (CLI/TUI print Error: + exit, API returns 400).
- .dockerignore: drop the /default.tar.gz line made redundant by the global
*.tar.gz pattern this PR adds.
- hermes profile export -o help text: stop advertising the old
<name>.tar.gz cwd default.
- Tests: cwd-in-unrelated-checkout topology (the second production shape
from the blocking review) and the fail-closed path. Mutation-checked:
both fail on the pre-fix helper.
Follow-up to the salvaged #92689:
- _profile_export_directory() now proves safety on the export dir's OWN
ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
heuristic missed the checkout whenever HERMES_HOME sat inside one but
the process ran from elsewhere (cron, service manager) — the export
landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
and catch OSError too — a bad profile name or read-only home printed a
raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
pollution in the tests/hermes_cli sweep can't divorce monkeypatches
from the code under test; add regression tests for the cwd-independent
topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
Profile export/import owns the only hardened tar handling in the tree:
GNU-format writing (PAX fractional mtimes make macOS Archive Utility
throw "Error 94"), plus an extractor that rejects absolute paths, `..`
components, and non-regular members.
Kanban board transfer needs exactly that, and a second copy is how the
weaker of two extractors eventually ships. Move the four helpers to
hermes_cli/archive_safe and point profiles at them; no behavior change
beyond dropping a provably-unreachable fallback in the root-listing
helper, whose condition is a strict subset of the comprehension above it.
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
Creating a bot from the desktop dialog builds the profile tree but no
config.yaml, so the profile resolves no provider and its first turn dies
with "No LLM provider configured" — created, but unable to run. Every bot
made that way was dead on arrival.
Seed the active profile's model block at creation. It is a copy, not a
link: profiles stay independent islands and editing either afterwards never
touches the other. "Fresh" means fresh skills and SOUL, not unreachable.
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.
- hermes-agent skill is now essential: cannot be disabled (config reads
strip it, hermes tools writes drop it), cannot be deleted by
skill_manage, is re-seeded past curator suppression, and is seeded
even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
file+terminal+vision+skills: read_file cannot read images and points
at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
tool isn't loaded.
All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.
Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.
Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.
Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.
Fixes#91583 (defect 2). Repro and live validation by @kubaboski.
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
The desktop's main serve process opens memory_store.db for every known
profile and nothing closed those connections before delete_profile's
rmtree — on Windows the open SQLite handles make the removal fail with
WinError 32 for both the CLI and the DELETE /api/profiles/<name> route
(#88347). POSIX unlinking of open files hid the same leak.
MemoryStore.close() is refcount-driven, so a live holder keeps the
handle forever; add MemoryStore.release_all_under(directory) to
force-close every shared connection under a directory, and call it in
delete_profile after stopping the profile backends. Inside serve the
handles live in that very process and get released; from the CLI it is
a no-op.
Fixes#88347
External review (Fable) caught a real false-positive widening in the
original commit: the new argv[1] script-name check reused the loose
`script_name == "hermes" or script_name.startswith("hermes")` pattern
(copy-pasted from the exe_name check above it), but argv[1] can be ANY
user-invoked python script path when argv[0] is a bare interpreter --
unlike a directly-resolved executable name, where a false match on the
substring is rare. A user's own script named e.g. "hermes-notes.py" or
"hermes-unrelated-tool" run via `python3 <script>` would be misidentified
as the console-script shim and become killable by profile delete.
Match against the actual known console-script entry points instead
(pyproject.toml [project.scripts]: hermes, hermes-agent, hermes-acp),
stripping the script's extension before comparing.
Added 2 regression tests: one confirms the false-positive case is now
rejected (fails against the pre-fix loose-match code, confirmed via a
scripted revert), the other confirms the other two real entry points
(hermes-agent, hermes-acp) still match via the shebang-exec path.
Tests: tests/hermes_cli/test_profiles.py -- 158 passed (156 previous + 2
new).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>