Commit Graph

111 Commits

Author SHA1 Message Date
teknium1 c1e58f4cb2 fix(process-identity): desktop reap, Windows updater and profile liveness use the canonical matchers
Two kill/relaunch predicates decided identity by argv substring, the bug class root AGENTS.md
forbids: hermes_cli/dashboard_procs.py::_is_desktop_local_serve_cmdline (`"serve" not in cmd`,
on the orphan-reap KILL path) and hermes_cli/update_cmd_windows.py::_is_backend_argv
(`" serve" in argv_low`, in the very file that defines _hermes_holder_subcommand). Both now ask
the canonical token classifier; host/port are read as flag values, not substrings.

hermes_cli/profiles.py::_check_gateway_running open-coded rungs 1/3 of
gateway.status.resolve_gateway_liveness and skipped the multiplexer rung; it is now that ladder
scoped to the profile dir (pid probe keeps cleanup_stale=False so a probe for another profile
never unlinks its PID file). The gateway/status.py ladder itself is untouched.

Behavior change: `hermes kanban --preserve-cache --host 127.0.0.1 --port 0` and
`-m dashboard serve`-style argv are no longer classified as serve backends (never killed /
relaunched as one); a named profile served by the live default multiplexer now reads as
running from _check_gateway_running (previously only via the separate
_served_by_running_multiplexer OR at some call sites).
2026-09-13 05:21:02 -07:00
teknium1 3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
teknium1 acbecf588a fix(profiles): --clone leaves messaging channels behind; --clone-channels opts in
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).

Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.

`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.

The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
2026-09-12 18:35:21 -07:00
teknium1 d1dbb0ac9e feat(gateway): multiplexer hot-serves profiles created while it runs, unroutes deleted ones
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.

The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
  (new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
  (`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
  with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
  updated, MCP discovery + log routing run for it. Other profiles' adapters are never
  touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
  create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
  pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
  memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
  multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
  that did not pick the profile up (older build / signal failed).
2026-09-12 08:49:16 -07:00
Teknium 9848e22ed6 feat(multiplex)!: drop gateway.multiplex_profile_allowlist — serve every profile
The multiplexing default gateway now serves default + every live named profile
under profiles/. profiles_to_serve(multiplex=True) is a pure directory read
(tombstoned profiles skipped, never mkdir); every reader — gateway served set,
/p/<profile>/ prefixes for api_server + webhook, the named-profile standalone
guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a
running gateway is unchanged) — drops the allowlist parameter.

Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG,
GatewayConfig and the top-level yaml bridge no longer carry it.

BREAKING: anyone who set an allowlist now has their excluded profiles served.
Archive or delete a profile you do not want served (Teknium approved).
2026-09-11 19:50:46 -07:00
Sora-bluesky f1ee8e0f39 fix(cli): release SessionDB handles when deleting a profile
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 06:24:11 -07:00
Teknium 48465c3933 fix(profiles): --clone-all no longer copies cron jobs into the new profile
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.

Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.

--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
2026-09-09 04:27:24 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium ca2675d9d6 refactor(hermes_cli): profiles cluster — share _canon_valid/_existing_profile_dir with profile_distribution, fold single-item brackets 2026-09-02 21:06:03 -07:00
Teknium 62124aa693 refactor(hermes_cli): profiles cluster — contextlib.suppress for best-effort blocks, _existing_profile_dir helper, drop intra-function blanks 2026-09-02 21:00:35 -07:00
Teknium 5264289b4f refactor(hermes_cli): profiles cluster — AST-neutral argument/collection packing 2026-09-02 20:47:59 -07:00
Teknium 02226f9ef3 refactor(hermes_cli): profiles cluster — AST-neutral bracket/string folding 2026-09-02 20:41:10 -07:00
Teknium d533b037d0 refactor(hermes_cli): profiles.py — split create_profile phases, unify canon+validate, compact docs 2026-09-02 20:27:08 -07:00
Teknium cfa327e5dc refactor(hclib): remaining hermes_cli library modules — dead code, unified helpers, flattened branches 2026-09-02 15:03:42 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium d3e2ace1dd fix(profiles): profile delete refuses to kill another profile's gateway (#89315)
`hermes profile delete` read the target profile's gateway.pid raw and
SIGTERMed it. When that pid file was poisoned by a sibling profile's gateway
(the #89315 shape), deleting profile A killed profile B's running gateway.

- gateway/status.py: `_pid_record_belongs_to_profile()` helper — a pid
  record whose recorded home differs from the expected profile home is not
  ours; legacy records without a home prove nothing and are left alone.
- hermes_cli/profiles.py: `_stop_gateway_process` refuses (and says so)
  when the record belongs to another profile; still stops its own gateway.

The stop/restart paths in hermes_cli/gateway.py did not need a guard:
`get_running_pid()` already filters cross-profile records and unlinks the
poisoned pid file before any kill can happen — verified live; the test for
that path now pins the real contract (returns False, other process alive,
poisoned pid file gone).

Live repro (unpatched main): `_stop_gateway_process(tim_home)` -> "Gateway
stopped (PID ...)" and the OTHER profile's process exits -15. After: "Refusing
to stop PID ..." and the process stays alive. 8 tests; sabotage (guard
removed) fails 1.
2026-09-02 04:53:59 -07:00
Teknium 3038493ee6 fix(auth): never fork single-use OAuth grants across profiles (#100339)
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:

1. `hermes profile create --clone-all` and the dashboard/TUI
   `mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
   verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
   drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
   `providers.<id>` device-code blocks, and the PKCE singleton file; API
   keys are still copied. The clone reads the root grant through the
   existing credential-pool root fallback.

2. A named profile with no local rows BORROWS the root grant via
   `read_credential_pool()`'s fallback, but every persist
   (`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
   rows into the profile's own auth.json — materializing a fork on the first
   rotation. `persist_pool_entries()` now routes borrowed single-use rows
   back to the root store (update-only, under the root lock; never falls
   back to a local copy). A borrowed `hermes_pkce` rotation commits its
   singleton to the root `.anthropic_oauth.json`, the borrower never prunes
   root-seeded rows it cannot see the backing file for, and
   `hermes -p <profile> auth add` persists only the profile's own rows.

Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.

Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).

Closes #100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 00:58:29 -07:00
kshitijk4poor 801466765b fix: harden temp fallback against predictable-path attacks; neutral error text; docs
Final-diff review findings (/simplify-code on the full 3-commit stack):

- Temp-dir fallback uses a per-uid name (hermes-profile-exports-<uid>) and
  get_profile_export_path refuses a pre-existing symlink or a directory
  owned by another user — a fixed /tmp/hermes-profile-exports is a
  predictable shared path a local attacker could pre-create to receive the
  secret-bearing archive. Regression test mutation-checked.
- Fail-closed message reworded interface-neutrally (the web API surfaces it
  as HTTP 400 detail where '-o' alone made no sense).
- Docs now cover the temp fallback and the fail-closed refusal.
2026-09-01 01:00:23 -07:00
kshitijk4poor 525dd12da7 fix: fail closed when no safe export destination exists; polish salvage edges
- _profile_export_directory(): when the managed store, the home-sibling
  store, AND the temp dir all resolve inside Git checkouts, raise a clear
  ValueError instead of warning and proceeding — a stderr warning would not
  stop a scripted export from staging a secret-bearing archive in a source
  tree, which is the exact #92457 incident class. All three callers already
  surface ValueError cleanly (CLI/TUI print Error: + exit, API returns 400).
- .dockerignore: drop the /default.tar.gz line made redundant by the global
  *.tar.gz pattern this PR adds.
- hermes profile export -o help text: stop advertising the old
  <name>.tar.gz cwd default.
- Tests: cwd-in-unrelated-checkout topology (the second production shape
  from the blocking review) and the fail-closed path. Mutation-checked:
  both fail on the pre-fix helper.
2026-09-01 01:00:23 -07:00
kshitijk4poor 26fb8f60e6 fix: anchor checkout detection to the export path, not cwd
Follow-up to the salvaged #92689:

- _profile_export_directory() now proves safety on the export dir's OWN
  ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
  heuristic missed the checkout whenever HERMES_HOME sat inside one but
  the process ran from elsewhere (cron, service manager) — the export
  landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
  violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
  and catch OSError too — a bad profile name or read-only home printed a
  raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
  pollution in the tests/hermes_cli sweep can't divorce monkeypatches
  from the code under test; add regression tests for the cwd-independent
  topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
2026-09-01 01:00:23 -07:00
joaomarcos c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
Brooklyn Nicholson 110ecd238e refactor(profiles): extract the safe tar.gz primitives into archive_safe
Profile export/import owns the only hardened tar handling in the tree:
GNU-format writing (PAX fractional mtimes make macOS Archive Utility
throw "Error 94"), plus an extractor that rejects absolute paths, `..`
components, and non-regular members.

Kanban board transfer needs exactly that, and a second copy is how the
weaker of two extractors eventually ships. Move the four helpers to
hermes_cli/archive_safe and point profiles at them; no behavior change
beyond dropping a provably-unreachable fallback in the root-listing
helper, whose condition is a strict subset of the comprehension above it.
2026-08-28 22:38:01 -05:00
Brooklyn Nicholson 21cc469eb1 Merge remote-tracking branch 'origin/main' into bb/bot-mode-design-system
# Conflicts:
#	apps/desktop/src/plugins/hermes-bots/plugin.js
#	apps/desktop/src/plugins/hermes-bots/tests/group-chat.test.mjs
#	apps/desktop/src/sdk/profile-routing.test.ts
2026-08-28 12:05:11 -05:00
EndeavorYen af2dc685fc fix(profiles): honor tombstones in exists/backfill and tighten named-home detection
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
2026-08-28 05:27:15 -07:00
EndeavorYen 7e34522454 fix(profiles): tombstone deleted named profiles so logging cannot resurrect them
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
2026-08-28 05:27:15 -07:00
Brooklyn Nicholson 01a3e9a44c fix(cli): give a profile created without a clone a usable model block
Creating a bot from the desktop dialog builds the profile tree but no
config.yaml, so the profile resolves no provider and its first turn dies
with "No LLM provider configured" — created, but unable to run. Every bot
made that way was dead on arrival.

Seed the active profile's model block at creation. It is a copy, not a
link: profiles stay independent islands and editing either afterwards never
touches the other. "Fresh" means fresh skills and SOUL, not unreachable.
2026-08-27 19:08:28 -05:00
Teknium 3733e4aff5 fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
2026-08-24 20:25:10 -07:00
Teknium 6cb1085d3d fix(gateway): /p/<profile>/ on a non-multiplex gateway fails closed instead of serving the owner profile
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.

Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.

Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.

Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.

Fixes #91583 (defect 2). Repro and live validation by @kubaboski.
2026-08-23 03:57:47 -07:00
Teknium a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00
Teknium a1682376ca feat(profiles): rename any agent — the default profile gets a display name (#45624)
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.

Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").

Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
2026-08-18 02:27:18 -07:00
liuhao1024 4f354c27b7 fix(profiles): release memory-store handles before rmtree on profile delete
The desktop's main serve process opens memory_store.db for every known
profile and nothing closed those connections before delete_profile's
rmtree — on Windows the open SQLite handles make the removal fail with
WinError 32 for both the CLI and the DELETE /api/profiles/<name> route
(#88347). POSIX unlinking of open files hid the same leak.

MemoryStore.close() is refcount-driven, so a live holder keeps the
handle forever; add MemoryStore.release_all_under(directory) to
force-close every shared connection under a directory, and call it in
delete_profile after stopping the profile backends. Inside serve the
handles live in that very process and get released; from the CLI it is
a no-op.

Fixes #88347
2026-08-17 16:33:52 -07:00
Adam Durham 70598f52e1 fix(desktop): tighten shebang-exec script-name match to known console scripts
External review (Fable) caught a real false-positive widening in the
original commit: the new argv[1] script-name check reused the loose
`script_name == "hermes" or script_name.startswith("hermes")` pattern
(copy-pasted from the exe_name check above it), but argv[1] can be ANY
user-invoked python script path when argv[0] is a bare interpreter --
unlike a directly-resolved executable name, where a false match on the
substring is rare. A user's own script named e.g. "hermes-notes.py" or
"hermes-unrelated-tool" run via `python3 <script>` would be misidentified
as the console-script shim and become killable by profile delete.

Match against the actual known console-script entry points instead
(pyproject.toml [project.scripts]: hermes, hermes-agent, hermes-acp),
stripping the script's extension before comparing.

Added 2 regression tests: one confirms the false-positive case is now
rejected (fails against the pre-fix loose-match code, confirmed via a
scripted revert), the other confirms the other two real entry points
(hermes-agent, hermes-acp) still match via the shebang-exec path.

Tests: tests/hermes_cli/test_profiles.py -- 158 passed (156 previous + 2
new).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-08-17 02:11:27 -07:00
Adam Durham 19c68e8c96 fix(desktop): work profile deletion silently reverted after app restart
Two independent bugs let a deleted profile reappear / leave orphaned
resources on next launch:

1. hermes_cli/profiles.py's backend-process scanner required argv[0] to
   resolve to an executable literally named "hermes". Electron's
   pool-backend spawn resolves the hermes console-script shim's path and
   execs it via the interpreter directly (python3 /path/to/hermes ...), so
   argv[0] reports as "python3" and the scanner never matched the running
   backend -- delete removed the profile's files but left its live backend
   process running (still bound to a port via uvicorn), which
   accumulates across repeated delete/recreate cycles.
2. The desktop sidebar's ProfileRail only refreshed its cached profile
   list once, on mount, so a delete/create/rename from another surface
   (another window, or the CLI) left a stale ghost entry until something
   unrelated triggered a refetch. Note: a delete via this window's own
   Manage-Profiles view already refreshes the shared $profiles atom
   ProfileRail subscribes to (confirmed by reading refreshProfiles() and
   handleConfirmDelete()) -- this fix only covers the cross-window/cross-
   process staleness gap, not a duplicate of the already-merged
   #57329's Manage-Profiles rail-refresh work.

Fix 1: recognize a python-interpreter argv[0] exec'ing a hermes-named
console-script shim via argv[1]. Fix 2: refresh the profile list on window
focus/visibilitychange, matching the existing pattern used elsewhere in
the sidebar (sidebar/index.tsx, use-background-sync.ts, star-map.tsx,
use-gateway-boot.ts all use the same focus+visibilitychange pattern).

## Related work already on main

PR #57329 (merged) fixed the *headline* symptom from issue #52279
(deleted profile respawns) via a different, non-overlapping mechanism:
routing profile-delete through the primary backend instead of spawning a
fresh pool backend, plus a separate recreation guard in
ensure_hermes_home() (#49435, merged) that makes a backend spawned into a
deleted profile's directory raise FileNotFoundError instead of silently
recreating it.

This PR is NOT a duplicate of that fix. Verified: even with both of those
merged, a backend process that survives because of gap #1 above still
holds a bound port via uvicorn -- it just can no longer resurrect the
profile directory. That's real resource-hygiene, not a symptom already
covered. Gap #2 touches a different file/component (ProfileRail /
profile-switcher.tsx) than #57329's rail-refresh half (which touched the
Manage-Profiles view's own $profiles.ts / index.tsx) and covers a
distinct staleness path (cross-window/cross-process, not same-window
delete-then-refresh).

Tests: tests/hermes_cli/test_profiles.py -- 156 passed (existing +
regression coverage for the argv[0] python-interpreter detection case).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-08-17 02:11:27 -07:00
Xinyu Du 57d94dd8dc fix(tools): keep junction lexical root across profile re-home
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling.  tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).

Two narrow changes, no heuristics, no new env vars:

- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
  the configured spelling IS the launch root (junction-transparent,
  physically identical dirs); keep it instead of re-deriving the native
  default.  This is the same producer contract _preserve_hermes_home_path
  already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
  configured home is a profile home (<root>/profiles/<name>), also derive
  the root spelling lexically (parent of the profiles component, same
  rule get_default_hermes_root uses) and run the exact-ownership mapping
  against it, so the launcher's lexical root is recovered after re-home
  without ever matching arbitrary descendants of HERMES_HOME.

Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
2026-08-17 02:10:32 -07:00
Trevin Chow c8f235a106 feat(gateway): allow selective multiplex profile serving 2026-08-10 22:48:24 -07:00
Brooklyn Nicholson 6c5cb2db4a fix(profiles): scrub secret-shaped strings from export archives
Shareable profile tarballs already drop auth.json/.env, but keys pasted
into skills, SOUL.md, or memories still shipped in plaintext. Force-run
the same redact_sensitive_text pass sessions export --redact uses on the
staged copy so the live profile is never rewritten.
2026-08-10 16:21:17 -05:00
kshitij a6e1e270b1 refactor: trim verbose comments + drop redundant default_flow_style kwarg
simplify-code follow-up: collapse 9-12 line inline comments to 3 lines
(keeping the loss-chain WHY + issue refs), and remove redundant
default_flow_style=False (atomic_yaml_write already defaults to False).
2026-08-05 11:59:33 +05:30
briandevans 649ce1f811 fix(cli): make profile.yaml and skin writes atomic to stop silent field loss
`write_profile_meta` and `hermes skin set` are both read-modify-write
helpers that rewrite a user-visible YAML file with a bare truncating
write, bypassing `utils.atomic_yaml_write` — the shared helper whose
docstring states that "every destructive file rewrite in the codebase
shares one implementation".

Both read halves swallow a parse error and fall back to `{}`, so a
truncated file is not transient corruption. The next call reads `{}` and
silently, permanently drops every field the caller did not explicitly
pass:

* `write_profile_meta` promises "unspecified fields preserve existing
  values". After an interrupted write, a follow-up call that only sets
  `description_auto` erases the profile's `description` — it vanishes
  from `hermes profile list` and never comes back.
* `_skin_set` exists so that "changing one token never disturbs the rest
  of the look". `path.write_text(...)` neither fsyncs nor swaps
  atomically, so a crash or power loss can leave `<skin>.yaml`
  zero-length; the next tweak then rewrites from empty and the whole
  palette is gone. The gateway's skin watcher repaints live surfaces
  from this file within ~1s, so a half-written file is observable.

Routing both through `atomic_yaml_write` gives temp file + fsync +
`atomic_replace`, which also preserves a symlinked target (GitHub
#16743) and restores owner/mode, and emits emoji descriptions as real
UTF-8 instead of `\UXXXXXXXX` escapes (GitHub #51356).

Supersedes #51808, which fixed the unicode-escaping symptom alone by
adding `allow_unicode=True` to the same `yaml.safe_dump` call.
2026-08-05 11:59:33 +05:30
Brooklyn Nicholson 1d6606d2cf fix(profiles): exported archives open in Finder (GNU tar, not PAX)
shutil.make_archive writes PAX with fractional-mtime records, which macOS
Archive Utility rejects ("Error 94 - Bad message") on double-click. Write
the profile archive with tarfile in GNU format instead: integer mtimes,
longlink for deep paths, extracts under Finder, bsdtar, and gnutar alike.
Verified against /usr/bin/tar (bsdtar) with >100-char member paths.
2026-08-04 13:23:33 -06:00
Brooklyn Nicholson d1196750c0 feat(profiles): REST export/import + extra_files overlay hook
export_profile() accepts extra_files (root-relative filename -> text) so a
caller can stage companion files into the archive; the desktop uses it for
desktop.json, its appearance/interface overlay, now part of the default
profile's export allow-list.

New routes wrapping the existing hermes profile export/import machinery:
- POST /api/profiles/{name}/export  (extra_files + optional output path)
- POST /api/profiles/import         (returns the bundled desktop overlay)
- GET  /api/profiles/{name}/desktop-overlay

Paths cross the API, not bytes - the desktop's native dialogs and its
local/pooled backends share a filesystem.
2026-08-04 12:07:29 -06:00
teknium1 ed33ebca1d refactor: canonical config loaders for behavioral reads + guarded raw-read primitive (kills the managed-scope/env-expansion drift class)
The disease: ~15 scattered raw yaml.safe_load(config.yaml) reads that
silently miss managed-scope overlay, ${ENV_VAR} expansion, profile-aware
pathing, and root-model normalization. Every new config feature needed an
N-site sweep (incident chain 9cbcc0c9c8 → 732293cf87 → b0e47a98f9 →
1928aa0443). This commit assigns every raw read to an owner and adds a
lint-guard test so the class cannot regrow.

New primitive (additive-only change to hermes_cli/config.py):
  read_user_config_raw(path=None) — reads the user file EXACTLY as
  written; docstring states it is ONLY legal for write-back round-trips
  and raw-file diagnostics. Behavioral reads must use
  load_config()/load_config_readonly().

BEHAVIOR FIXES (class-a sites migrated to a canonical loader — these
previously read values that could DIFFER from the effective config):

  gateway/run.py _try_resolve_fallback_provider → _load_gateway_runtime_config
    keys: fallback_providers/fallback_model (provider, model, base_url,
    api_key). Drift fixed: a managed-pinned fallback chain was ignored;
    an api_key of "${OPENROUTER_API_KEY}" reached the resolver unexpanded.
  gateway/run.py GatewayRunner._load_provider_routing → same loader
    key: provider_routing. Drift fixed: managed-pinned routing prefs and
    ${VAR} templates were ignored.
  gateway/run.py GatewayRunner._load_fallback_model → same loader
    keys: fallback chain. Same drift as above.
  gateway/run.py GatewayRunner._refresh_fallback_model
    keeps the raw primitive (its last-known-good-on-parse-failure contract
    forbids the fail-open loader, which returns {} on a torn write) but now
    applies managed overlay + env expansion inline. Drift fixed: chain
    edits under managed scope / env templates were previously frozen out.
  tui_gateway/server.py _load_cfg (72 behavioral call sites)
    now = raw read + managed overlay (pre-existing) + NEW ${VAR} expansion,
    split from a new _load_cfg_raw() write-back primitive. Drift fixed:
    e.g. custom_prompt: "hello ${VAR}", agent.system_prompt, model,
    api_key/base_url templates reached sessions unexpanded. DEFAULT_CONFIG
    is deliberately NOT merged (callers treat missing keys as unset;
    `_load_cfg() == {}` sentinels and _save_cfg round-trips depend on it).
  tui_gateway/server.py _profile_configured_cwd
    keys: terminal.cwd of a NON-launch profile. Drift fixed: managed
    overlay + ${VAR} expansion now apply (load_config() would resolve the
    wrong profile's home, so the raw primitive + inline pipeline is used).
  plugins/platforms/telegram/adapter.py _reload_dm_topics_from_config
    → load_config_readonly(). keys: platforms.telegram.extra.dm_topics.
    Drift fixed: managed overlay + profile-aware pathing + expansion.
  plugins/memory/holographic _load_plugin_config → load_config_readonly().
    keys: plugins.hermes-memory-store.*. Same drift class.

WRITE-BACK ROUND-TRIPS (class-b: stay raw BY DESIGN via read_user_config_raw;
merging defaults/overlay would pollute the saved user file):
  gateway/slash_commands.py: model persist x2, _save_gateway_config_key,
    memory/skills write_approval toggles
  gateway/platforms/yuanbao.py auto-sethome
  tui_gateway/server.py _write_config_key + all cfg→_save_cfg blocks
    (reasoning show/hide/full/clamp, details_mode[.section], prompt)
    → new _load_cfg_raw()
  plugins/memory/holographic save_config

RAW-FILE DIAGNOSTICS + presence-sensitive bridges (class-c: stay raw,
now via the shared primitive with an explanatory comment):
  hermes_cli/doctor.py x5 (model validation, stale-root-keys, .env drift,
    deprecation sweep, memory-provider probe — the latter two keep their
    inline managed overlay where they had one)
  gateway/run.py _bridge_max_turns_from_config and the module-level
    TERMINAL_*/HERMES_* env bridge (bridging merged defaults would export
    all of DEFAULT_CONFIG into the environment; both keep their inline
    overlay + expansion)
  hermes_cli/send_cmd.py env bridge (same presence-sensitivity)
  hermes_cli/gateway.py multiplex-conflict probe (reads the DEFAULT root's
    config, not the active profile's — load_config is the wrong owner)
  hermes_cli/profiles.py / hermes_cli/web_server.py / tools/wake_word.py
    multi-profile reads (load_config targets only the ACTIVE profile home)
  cron/jobs.py _resolve_default_model_snapshot and cron/scheduler.py
    run_job config read keep their existing inline overlay+expansion but
    now share the primitive (their fail-open + last-value semantics and
    the deliberate no-defaults merge are preserved exactly).

Failure-semantics audit: every migrated site preserves its exact previous
behavior on missing file ({} / early return) and parse failure (raise into
the caller's existing except, warn, last-known-good, or fail-open) —
read_user_config_raw intentionally mirrors bare open()+safe_load semantics
(raises on parse errors, {} only on FileNotFoundError/non-dict root).

Guard: tests/hermes_cli/test_config_read_guard.py scans the tree for
yaml.safe_load within 6 lines of a 'config.yaml' reference outside an
explicit ALLOWLIST (hermes_cli/config.py, gateway/config.py, gateway/run.py
fallback path, hermes_cli/managed_scope.py which reads the MANAGED file,
gateway/readiness.py parse-health probe) and fails on new offenders.

E2E: tests/hermes_cli/test_config_loader_e2e.py runs a subprocess with a
temp HERMES_HOME (config.yaml containing ${E2E_PROMPT_SUFFIX}) plus a
HERMES_MANAGED_DIR overlay pinning agent.reasoning_effort, asserting
tui _load_cfg resolves "hello world"/"high" while _load_cfg_raw +
_save_cfg round-trip the template and user value verbatim with no
managed/default leakage.
2026-07-29 10:53:29 -07:00
AlexFucuson9 411a686de2 fix: add encoding="utf-8" to Path.write_text() calls (P1)
Path.write_text() without encoding defaults to system locale encoding.
On Windows (cp1252), this silently corrupts non-ASCII content written
to JSON files, config files, and cache files.

This is the write-side counterpart to the read_text() encoding fix
(PR #56115). PLW1514 only covers open() calls — Path methods are
unguarded by ruff.

39 instances across 16 files, all passing py_compile.

Files changed:
- agent/copilot_acp_client.py (1)
- tools/web_tools.py (1)
- tools/xai_http.py (1)
- tools/skills_hub.py (8)
- gateway/slash_commands.py (1)
- gateway/run.py (5)
- gateway/dead_targets.py (1)
- gateway/delivery.py (2)
- gateway/platforms/qqbot/adapter.py (1)
- hermes_cli/gateway.py (1)
- hermes_cli/banner.py (1)
- hermes_cli/service_manager.py (5)
- hermes_cli/container_boot.py (5)
- hermes_cli/uninstall.py (1)
- hermes_cli/main.py (2)
- hermes_cli/profiles.py (3)
2026-07-24 17:10:39 -07:00
AlexFucuson9 cf35fd6de5 fix(core,cli,gateway,plugins): add encoding='utf-8' to read_text() calls
Path.read_text() without an explicit encoding uses the platform's
default encoding. On Windows this is typically cp1252 or mbcs, which
causes UnicodeDecodeError or silent data corruption when reading
UTF-8 content (JSON files, user text, config with non-ASCII chars).

This is the read-side companion to the write_text() encoding fix.
Fixed the most critical locations that read JSON data, user content,
and config files across 14 files with 31 call sites.

Pattern: .read_text() → .read_text(encoding='utf-8')
         json.loads(path.read_text()) → json.loads(path.read_text(encoding='utf-8'))
2026-07-24 17:10:39 -07:00
Stoltemberg c89481db5e fix: add explicit UTF-8 encoding to all subprocess text=True calls (#53428)
On Windows with Chinese locale (GBK), subprocess.run(text=True) without
explicit encoding causes UnicodeDecodeError crashes. This fix adds
encoding='utf-8', errors='replace' to all subprocess.run() and
subprocess.Popen() calls that use text=True across 76 non-test Python files.

Fixes #53428 (master tracker for Windows GBK locale crash).

Note: credential_pool.py and electron changes excluded per reviewer request —
those will be submitted as separate focused PRs.
2026-07-24 11:45:57 -07:00
Teknium 55e3ee1ab8 fix: remove dead f-string prefixes via ruff F541 (216 sites) (#52336)
ruff check --fix --select F541 . on current main. Pure prefix removals;
adjacent-string concatenations keep the f only on interpolating fragments.
No string content or live placeholder altered.
2026-07-05 13:42:46 -07:00
Matt Van Horn b7192b1cb0 fix(profiles): preserve symlinks in clone-all and skills clone paths
Widens the symlinks=True fix to the create_profile clone sites so a
symlink pointing at a parent directory can't recurse infinitely during
'hermes profile create <name> --clone-all' (#11560). Export paths were
covered by the salvaged #58397/#58445 commits; this carries the clone
half of open PR #11573.

Fixes #11560
2026-07-05 00:48:50 -07:00
Ahmett101 8d9684c9da fix(profiles): allowlist default-export paths + preserve symlinks (#58394)
`hermes profile export default` crashed with `shutil.Error` when
HERMES_HOME pointed outside ~/.hermes (common in Docker deployments)
and the workspace contained broken symlinks. Two root causes:

1. `copytree` defaults to `symlinks=False` and follows link targets;
   broken ones crash. #58397 (liuhao1024) drafted a minimal
   `symlinks=True` flag fix; this PR adopts that change.
2. `copytree` was invoked against the entire HERMES_HOME root (which
   doubles as cwd in Docker layouts). The post-hoc blacklist at
   `_DEFAULT_EXPORT_EXCLUDE_ROOT` is a fixed-length enumerate-and-pray
   list that can't anticipate every unrelated sibling directory
   (`x11-dev/`, etc.). Replaced with a positive allow-list at
   `_DEFAULT_EXPORT_INCLUDE_ROOT` enumerating the known Hermes profile
   artifacts (config, persona, skills, cron, scripts, sessions,
   plugins, memories, knowledge, preferences). Sensitive runtime
   surfaces (`state.db`, `logs/`, auth files, other profiles) are
   intentionally not in the allow-list so the export stays a
   portable, credential-free snapshot of the user-facing surface —
   which means the existing `test_export_default_excludes_infrastructure`
   regressions remain green.

Adds two regression tests:
  * test_export_default_uses_allowlist_for_unrelated_dirs — >x11-dev<
    sibling directories must not leak into the archive.
  * test_export_default_handles_broken_symlinks — symlinks inside
    allowed artifacts survive instead of crashing the export.

closing that PR as superseded once this lands.

Closes #58394
2026-07-05 00:48:50 -07:00
liuhao1024 b6b9bcd2a1 fix(profiles): preserve symlinks during profile export
shutil.copytree() defaults to symlinks=False which follows symlinks and
crashes on broken ones.  In Docker/custom HERMES_HOME deployments,
unrelated directories may contain stale symlinks that break export.

Add symlinks=True to both copytree() calls in export_profile() so
broken symlinks are preserved as symlink entries in the archive.

Fixes #58394
2026-07-05 00:48:50 -07:00