Commit Graph

35799 Commits

Author SHA1 Message Date
teknium1 92d64465bf docs(multiplex): routed-child credential rule and the frozen launch env restart requirement
kvnloo (#111620 review) asked for the operator-visible statement that once a serve /
dashboard process hosts a second profile, the launch profile's env-only credentials are
frozen at activation and a rotation in the process env needs a restart. Also states the
routed-child rule with and without the multiplex flag (#111617). Adds encoding='utf-8'
to the probe child in the child-env authority test (windows-footguns lint).
2026-09-16 00:35:00 -07:00
teknium1 db54f5448d fix(kanban): worker fingerprint carries a boot witness; an uncaptured fingerprint never authorizes a signal
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):

- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
  is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
  a reboot, and that counter does not, so an unrelated process on a later boot with the
  same PID and the same tick value passed _start_times_agree(). The fingerprint is now
  "<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
  the witness the drain marker already uses); both halves must match. Integer values on
  rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
  row and falls back to bare PID existence - a new spawn silently recreated the #89614/
  #99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
  claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
  by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
  once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
  the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
  and must not outlive its pid.

Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.

Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
2026-09-16 00:35:00 -07:00
teknium1 8609743389 fix(serve): launch-profile scope decided at entry; send keeps scope authority; per-reset release
Three edges of the fail-closed multi-profile host (#111620 review, andrexibiza P1 + P2,
kvnloo finding 1):

- `send` under a routed profile's scope `update()`d the installed scope from raw `.env`,
  reversing build_profile_secret_scope's precedence (user .env, then external secret
  sources) for the rest of the request; a stale user value beat the secret-manager one.
  The installed scope is authoritative as-is; only the config.yaml setdefault bridge runs.
- The launch profile's body was scoped only when `is_multiplex_active()` was already true
  at entry, while get_secret consults that global on every read. A launch RPC / dashboard
  request entering single-profile and resuming after a concurrent first `?profile=B`
  activation raised UnscopedSecretError mid-request. The launch profile's secret scope
  (its .env + external sources over the launch env: live while single-profile, the frozen
  snapshot once multiplexing is active) is now bound for every launch-profile body, so the
  credential source is fixed at entry. The terminal policy overlay stays multiplex-only
  (standalone terminal execution keeps its os.environ bridge). _publish_env_value mirrors a
  same-request .env write into that scope AND os.environ for the launch profile, only into
  the scope for a routed one (serves_routed_profile).
- _release_profile_runtime_scope_tokens reset terminal → secret → home in sequence under
  one outer suppress; a failing terminal reset left the previous profile's secrets and
  HERMES_HOME installed for the next body in that context. Each reset is now independent;
  the first failure is re-raised after every scope is released.

tests/tui_gateway/test_multi_profile_hosting_transitions.py: manager-vs-dotenv precedence
through _load_hermes_env, TUI-RPC and dashboard barrier tests (launch enters single-profile,
B activates on another thread, launch resumes and still resolves its injected credential,
never B's), forced terminal-reset failure still releases secret + home. 4/4 red on base.
2026-09-16 00:35:00 -07:00
teknium1 3fe8e5e443 fix(multiplex): routed children never inherit launch-only credentials, with or without the multiplex flag
Two authority gaps in served_profile_child_env (#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):

- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
  provider credentials; strip_launch_profile_env only knows names with .env/source
  provenance, so a key systemd/Compose/the shell injected into the launch process
  survived into profile B's child whenever B did not define the same name. Now a ROUTED
  target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
  before B's own scope is overlaid (the child boundary gets get_secret's contract: a
  scoped miss is no credential, never ambient fallback). The launch profile's own child
  keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
  dashboard backends serve ?profile=B by installing the HERMES_HOME override without
  that flag, so B's slash worker / helper children kept A's .env and settings. The
  authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
  to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
  Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
  under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
  thread); it now raises UnscopedSecretError like get_secret.

tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
2026-09-16 00:35:00 -07:00
teknium1 7dde7a2424 fix(gateway): migrate --multiplex resumes from live state, compensates the whole destructive phase, and refuses an unknown default principal
Three P1 findings from the review of #111062 (fixes #110850 remainder):

- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
  installed-but-dead unit; `systemd_install` writes the unit before the start that
  can still be killed, so an apply interrupted at default/start (or mid-restart of
  a stopped default) was reported as "already multiplexed" and never recovered.
  `interrupted` is now derived from the manifest vs the LIVE default's recorded
  served_profiles: flag on + manifest + not every migrated profile served = resume.

- the compensation `try` began after `_remove_secondary_gateways` and the flag
  write, so a later secondary's stop, a unit's daemon-reload or the config write
  failing left the first secondary removed with no rollback. The flag write,
  removals and default bring-up now all sit inside one boundary that rolls back
  through the manifest; `_preflight_apply` refuses failures knowable from the plan
  (root system unit without a recorded User=, unresolvable recorded user, an
  unwritable config.yaml) before any working gateway is stopped.

- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
  `default_uid is None`; a default system unit whose User= this host cannot
  resolve is the same boundary from the other side and now blocks the update
  hook (same known uid still folds, different known uid still refuses).
2026-09-16 00:33:03 -07:00
Sora-bluesky 9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
teknium1 5d97d5ed6d refactor(profiles): move rename identity migration off the facades; trim tests to invariants
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
  migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
  now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
  profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
  gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
  hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
  tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.

Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
2026-09-16 00:32:15 -07:00
Konstantin Khlopkov 7f7d229f83 fix(utils): atomic writers refuse to resurrect a deleted named profile home
Background writers that still carry a tombstoned profile as their Hermes
home (reasoning-caps warm thread, models cache, models.dev ETag, gateway
lifecycle ledger, MCP OAuth tokens, memory store) re-created
profiles/<name>/ with a bare mkdir right before an atomic write. Route the
parent-dir creation through mkdir_under_hermes_home so a deleted named
profile raises FileNotFoundError and stays gone, matching the tombstone
contract already enforced for logging and state.
2026-09-16 00:32:15 -07:00
KoNit-K 91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
xielevi 81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00
xielevi 4ba717df12 fix(profiles): migrate session/routing identity on profile rename
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.

The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:

- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
  tables, matching the agent:<name>: namespace by exact prefix (substr, not
  LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
  the profile inside routing/origin JSON, and REFUSING on a target collision
  (routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
  (keys + origin.profile) then persist — the half a DB write cannot reach.
  Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
  params only to handlers that declare them, bare handlers unchanged) so a
  live gateway rekeys its in-memory copy AND both durable stores (routing home
  + the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
  does NOT fall back to a racing CLI-side write: it prints a warning telling
  the operator to restart the gateway and retry. With no live gateway it
  performs the durable rewrite itself (safe: nothing else holds the store
  open).

Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.

Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
2026-09-16 00:32:15 -07:00
brooklyn! 2e320e6d1a fix(desktop): keep inline edit placeholders off entered text 2026-09-16 02:04:58 -05:00
brooklyn! 64a9b43261 docs: run worktree desktops side by side without port eviction 2026-09-16 01:54:36 -05:00
brooklyn! 2b7292ff0f style(desktop): frame inline subagents with a subtle outline 2026-09-16 01:44:33 -05:00
brooklyn! 3bdd4cc5fd fix(desktop): keep background continuations in one visual response 2026-09-16 01:44:33 -05:00
brooklyn! 57d34b14c2 fix(desktop): keep completed subagents from reviving stale activity 2026-09-16 01:44:33 -05:00
kshitijk4poor f0b2ba4160 test(tui_gateway): teardown tests prove agent.close() reached the row finalizer
_teardown_session swallows agent.close() failures, so a "row stays open"
assertion alone is green even when close() aborted before
_finalize_owned_session_row. Assert the one side effect only that finalizer
produces (_owns_session_db cleared), give the stub the _memory_manager attr
close() reads directly, and assert the flag is untouched on the control
paths. Rename the sole-owner cross-process test to what it now asserts.
2026-09-16 12:13:05 +05:30
kshitijk4poor dfc40a50a9 fix(tui_gateway): a spared durable row survives agent.close() too
_finalize_session decides whether the TUI backend ends the state.db row
(gateway-owned source → no; another backend holds the lease → no; automatic
Desktop reclaim → no, #105588). _teardown_session then calls agent.close(),
whose _finalize_owned_session_row ends the same row as "agent_close" through
the agent's own handle whenever _end_session_on_close is still True — so every
spare decision was undone one call later. Verified with real SessionDB +
AIAgent.close(): with only the #105620 guard a Desktop ws_orphan_reap still
lands ended_at + end_reason="agent_close" (the #105588 follow-up report).

One seam after the lifecycle decision now clears agent._end_session_on_close
whenever the row was spared, replacing the gateway-owned-only assignment from
#111145 so the lease-held and Desktop-reclaim cases get the same treatment.
The five mocked _finalize_session tests from #105620 are replaced by two
real-DB _teardown_session tests beside the #111145 harness (red on both
origin/main and the bare cherry-picks, green on the stack).
2026-09-16 12:13:05 +05:30
Eva a8b8556108 fix(tui): a gateway-owned session survives agent.close() 2026-09-16 12:13:05 +05:30
ClintonEmok d3a753df30 fix(gateway): don't end durable session row on automatic Desktop cleanup
ws_orphan_reap, idle_timeout, lru_evict, ws_disconnect, and tui_shutdown
are runtime/connection GC — not user intent. Ending the durable session
row during these automatic cleanup reasons confuses GC with user action,
causing secondary bot chats to vanish from the sidebar even though the
transcript is intact in state.db.

The canonical Bot Chat already has resurrection paths for accidental
ws_orphan_reap ends, but non-canonical secondary chats do not, so they
are hit harder.

Skip db.end_session() when _desktop_automatic_cleanup is True (automatic
cleanup reason + Desktop source). Runtime is still reclaimed, the
session.reclaimed event still fires, and explicit user close/archive/reset
still ends the durable row normally.

Fixes #105588
2026-09-16 12:13:05 +05:30
kshitijk4poor 5d07f7fe66 test(agent): fold the aux create_client() seam tests to two contracts
Two parametrized invariant tests (native client sync/async; None or raising
profile falls back to the standard client) replace four near-duplicates, and
the hook's kwargs are pinned exactly to the mapping openai.OpenAI would have
received. The fixture isolates both provider registries through monkeypatch
instead of a hand-rolled snapshot/restore, drops the HERMES_HOME override that
tests/conftest.py already provides, and replaces the tuple-truthiness lambda
with a plain function. Call-site comment trimmed to the ordering WHY.
2026-09-16 11:57:56 +05:30
liuhao1024 4cd2eb013c fix(agent): honor ProviderProfile.create_client() for api_key aux routes
The auxiliary api_key branch built openai.OpenAI directly, so an
out-of-tree provider registered with auth_type="api_key" lost its
native transport for auxiliary tasks even though the main-agent path
(_provider_supplied_client) and the external_process branch honor the
same hook. Consult the profile's create_client() before the built-in
gemini/OpenAI ladder; None (the default) falls through untouched, and
a raising profile is logged and skipped (#112384).
2026-09-16 11:57:56 +05:30
teknium1 2cfb655d52 test(kanban): connect via kanban_db_connect, not the facade compat pointer 2026-09-15 22:47:03 -07:00
teknium1 6e6550e181 docs(kanban): document worker_output on crashed / protocol_violation events 2026-09-15 22:47:03 -07:00
chelsealong 3b9c1118cd fix(kanban): surface the worker's own last output on a dead-worker reap
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.

Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
2026-09-15 22:47:03 -07:00
teknium1 c2e5c94cd7 fix(tui): tolerate agents without session_cwd in _register_session_cwd; adapt stubs to the cwd kwarg
Workspace moves stamp agent.session_cwd so a lazily started Codex thread
starts in the moved-to directory. Agents that never had the attribute
(test doubles, slotted objects) must keep working, so stamp only when the
attribute exists. Test stubs of _set_session_context mirror the new cwd
kwarg, and the Codex gateway test asserts the contract (no pinned session
cwd) instead of the attribute's absence.
2026-09-15 22:30:11 -07:00
teknium1 3317b8c1e1 test(honcho): trim salvage coverage to the invariants; drop duplicate ACP session_cwd stamp
Salvage of #93452 (@outpoints). Keep one invariant per fix:
resolver-level "automatic title never remaps a strategy session",
integration "provider routes by logical workspace, not process cwd",
agent-level "title provenance + cwd reach the provider", deferred
Desktop/TUI build threads the session cwd, seeded branch titles are
derived, and the workspace-move E2E. Drop the plumbing/legacy-shape
tests that re-assert the same contract.

acp_adapter: AIAgent(cwd=...) now stamps session_cwd itself, so the
direct assignment after construction was a duplicate.
2026-09-15 22:30:11 -07:00
outpoints 94efd22177 fix(honcho): preserve title provenance through upstream peer routing 2026-09-15 22:30:11 -07:00
outpoints 1c8658bf0d test(honcho): [verified] preserve profile and title provenance coverage 2026-09-15 22:30:11 -07:00
outpoints c2f743f9a1 fix(honcho): [verified] preserve seeded branch title provenance 2026-09-15 22:30:11 -07:00
outpoints 86b3695248 fix(tui): synchronize runtime cwd after workspace moves 2026-09-15 22:30:11 -07:00
outpoints 9e6c79f5be docs(honcho): document workspace routing and provider context 2026-09-15 22:30:11 -07:00
outpoints 44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
outpoints 5237cab756 fix(honcho): thread logical cwd through agent construction
(cherry picked from commit b1d7207c45311be658592c6ad34ee84634fed0ee)
2026-09-15 22:30:11 -07:00
outpoints 231e1c1204 fix(honcho): resolve sessions against agent cwd, not process cwd
The Honcho provider resolved per-repo/per-directory session names from
os.getcwd(), which is the backend process launch directory on
Desktop/gateway hosts (typically $HOME), not the user's workspace. With
a manual sessions map entry for the home directory, every Desktop
conversation landed in that fallback bucket instead of the project's
per-repo session.

Use agent.runtime_cwd.resolve_agent_cwd() — the same single source of
truth already used for system-prompt and context-file discovery — so
Honcho session routing agrees with everything else about where the
agent logically lives. It honors the pinned session cwd, then
TERMINAL_CWD, then the launch directory; CLI sessions launched inside a
project resolve identically either way.

Adds a regression covering the Desktop-style case: backend launched in
$HOME, workspace elsewhere, home-directory manual map present.

Refs #24740

(cherry picked from commit cfb32757de2509dd404acf98311e357d73fa741e)
2026-09-15 22:30:11 -07:00
outpoints 3cbdc32565 fix(honcho): don't let auto-generated session titles override sessionStrategy
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.

Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.

Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.

Fixes #24740

(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
2026-09-15 22:30:11 -07:00
teknium1 3abeca16e6 fix(kanban): 5xx and timeouts requeue the worker instead of spending its retry budget
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
2026-09-15 22:05:38 -07:00
teknium1 bf1c28b480 ci: fail lint when a test fakes macOS without @pytest.mark.macos_only (#111866)
scripts/ci/check_os_marker_fakes.py flags test files that make the
interpreter believe it is on macOS (is_macos -> True, sys.platform ->
"darwin", platform.system -> "Darwin") while carrying no `macos_only`
marker: the macOS lane imports only marked files, so such a file is green
on Linux over a faked branch and never runs on the host it exists for.

A `# os-marker: ok — <why>` comment opts a host-independent line out.
_BASELINE holds the files that already faked macOS when the check landed;
a stale entry fails the check so the list can only burn down. Wired into
lint.yml next to the compat-pointer check; two invariant tests cover a
flagged fake vs marked/opted-out files and host-honest platform reads.
2026-09-15 21:49:33 -07:00
teknium1 f0fd0650b5 test(update): autostash suite keeps the launchd restart scope off the host (#111866)
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
2026-09-15 21:49:33 -07:00
teknium1 b7ccc82d82 docs(config): match the dropped unknown-name note 2026-09-15 21:47:48 -07:00
teknium1 08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1 fd74bf6f58 fix(config): drop the runtime read of the docs env registry
Production code must not depend on website/docs being present; the
"Hermes does not read this name" note now keys only on the code
registry (OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS / setup-hidden suffixes).
2026-09-15 21:47:48 -07:00
teknium1 f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1 db25a7852e fix(plugins): install refuses to ship an unreadable plugin tree (#111804)
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.

After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.

Fixes #111804 (its discovery half landed in #112293).
2026-09-15 21:47:18 -07:00
teknium1 23863ccbaf fix(cli): one-shot chat -q exits non-zero on failure; 75 covers upstream 429 and overload
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
2026-09-15 21:46:41 -07:00
teknium1 2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1 3c3ab69abb fix(tui_gateway): stop the bot mailbox poll from flooding the log on installs that never received a delivery
The per-session notification poller ran `_poll_bot_live_delivery_once` every
0.5 s. Once a "Bot Chat" session exists, each pass opens state.db and takes
the exclusive active-session registry lock; on Windows (`msvcrt` LK_LOCK
gives up after 10 s of contention) that raised
`RuntimeError: active session file lock unavailable` and the loop logged
`Bot live-owner delivery poll failed` on every attempt — 8,838 warnings in
three days, 91% of one install's WARNING output (#111719).

Two changes:
- `tools/bot_live_delivery.has_mailbox`: the mailbox directory is created
  only when a delivery is first admitted, so a profile without it has nothing
  to claim — the poll now returns before the state.db open / registry lock.
  This keeps cron→Bot Chat and Bot Mode DM delivery intact on installs with
  no messaging platform configured (both deliver through this mailbox), which
  is why the poll is gated on the mailbox rather than on connected platforms
  (PR #111733's guard would have broken those).
- `_poll_bot_live_delivery_guarded`: a failing poll backs off 5 s before the
  next attempt and is logged at WARNING once per 60 s window (with the count
  of suppressed repeats), DEBUG otherwise.

Live probe (real poller loop, temp HERMES_HOME with a Bot Chat row and the
registry lock made unavailable, 3 s):
before: owner_lookups=6 WARNING=6 (with or without a mailbox)
after:  no mailbox -> owner_lookups=0 WARNING=0; mailbox -> owner_lookups=1 WARNING=1

Fixes #111719
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:52:13 -07:00
teknium1 2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural 7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
teknium1 1b2a75fa27 chore: map dmelkk-secondbrain contributor email 2026-09-15 19:28:32 -07:00