4381 Commits

Author SHA1 Message Date
kshitijk4poor 9395f2f0d4 refactor(update): the purge scan has no fallback; trim to two invariants
The checkout root is where the update just pulled into, so "root unreadable" cannot
happen after a successful pull — drop the OSError fallback and the static five-name
tuple it fell back to (the tuple was the drift that caused the bug). Drop the phantom
`hermes_cli.hermes_logging` protection entry: no such module exists; the real
`hermes_logging` is root-level and now protected by name. Keep two tests: the stale
root `utils` scenario (red on base) and the hermes_logging protection.
2026-09-15 21:57:59 +05:30
Hubert Nimitanakit 29758a4eb2 fix(cli): purge every top-level checkout module, not just five packages
`_STALE_PURGE_PREFIXES` listed five package names, so every top-level module in
the checkout root survived the post-pull purge. `hermes update` then imported
new source against a cached pre-pull `utils`, `hermes_constants`, `plugins` or
`providers`.

Field failure (2026-09-12, macOS): updating 0.20.6 -> 0.21.2 crossed 3145986c20,
which added `base_url_origin` to `utils.py`. The restart phase's up-front
`from hermes_cli.gateway import ...` pulled `agent.auxiliary_client`, whose
`from utils import base_url_origin` hit the cached 0.20.6 `utils`:

    Update incomplete - gateway auto-restart failed: cannot import name
    'base_url_origin' from 'utils' (.../hermes-agent/utils.py)

The purge docstring already claims it evicts EVERY cached Hermes module; the
hardcoded tuple was the same "re-fixed per symptom" shape it replaced. Scan
`PROJECT_ROOT` instead: top-level `.py` files plus directories with an
`__init__.py`. 50 names here, and a newly added module can no longer drift out.

Two exclusions, both deliberate:

- `hermes_logging` joins `_STALE_PURGE_PROTECTED`. Its queue listener, handler
  list and `_logging_initialized` flag are module globals, so a fresh copy
  starts a second QueueListener over the same log files while the first runs.
- `tests` is never purged. pytest resolves fixtures through the identity of its
  already-imported test modules.

Falls back to the old tuple when the root is unreadable.
2026-09-15 21:57:59 +05:30
kshitijk4poor 78d338b9ee test(cli): drop the tautological mcp_servers assertion from the seed test
`_config_mcp_servers` is `self.config.get("mcp_servers") or {}`; with the
test's own `{"mcp_servers": {}}` it is `{}` whether or not the import fix
is present, so the assertion discriminated nothing. One invariant per fix.
2026-09-15 20:47:06 +05:30
mr-r0b0t 3d2842d84f fix(cli): import file_signature in TUI run-state init
#111408 widened the MCP config watcher seed from mtime to
utils.file_signature but omitted the import in cli_tui_mixin.
The NameError only fires when config.yaml exists, so isolated-home
tests short-circuited past it and every real CLI launch crashed.

Signed-off-by: mr-r0b0t <adam.manning@gmail.com>
2026-09-15 20:47:06 +05:30
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes d89cacc25f chore(free-tier): keep the rehearsal server out of the repo; the docs page explains the stand-in instead
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 24fd22b94d test(cli): trim #110737 re-queue coverage to two invariants on the mixin
Drive CLIChatTurnMixin._chat_render_turn through a minimal stub instead of a full
HermesCLI (prompt_toolkit stubs, cli reload): the invariants are that a (text, images)
payload reaches _pending_input intact and that several queued parts join their text
and merge their images. Both are red on origin/main.
2026-09-15 06:36:55 -07:00
Kevin Rajan 40b22806f8 fix(cli): preserve image attachments when re-queueing interrupt messages
_clui mixin bundles Enter submissions carrying images as (text, images)
tuples, and _tui_enter_while_busy puts those tuples on _interrupt_queue in
interrupt mode. _chat_render_turn then did "\n".join(all_parts) over the raw
payloads, raising TypeError on the tuple — swallowed by the outer chat()
handler, so the interrupt message was silently lost and no next turn started.

Unpack tuple payloads in the re-queue block: join the text parts and re-attach
all images as a single (text, images) tuple, matching the shape
_tui_process_one_input already accepts. Plain-string payloads are unchanged.

authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; all changes reviewed and approved by the contributor
2026-09-15 06:36:55 -07:00
teknium1 7a1d278528 fix(dashboard): skip the stalled-input close when attach detached itself
attach() now returns False both when the client dropped mid-replay (session
already detached, socket dead) and when the force-redraw write stalled
(socket still attached, worth closing with 1013). Only close the socket in
the second case; the registry detach stays as the guard for both.

Also trims the salvaged tests to two invariants, both red on origin/main:
a mid-replay drop leaves the session reap_idle()-reclaimable, and a drain
send failure detaches the current socket without touching a replacement
that attached during the send (#110849).
2026-09-15 06:36:19 -07:00
MKanso 941d8b59a7 fix(pty): handle replay send failures cleanly 2026-09-15 06:36:19 -07:00
MKanso 6d622d8eba fix(pty): recover sessions after socket send failures 2026-09-15 06:36:19 -07:00
teknium1 182ec5c28d fix: widen the remaining config.yaml stat caches to file_signature
Three caches that read config.yaml still keyed change detection on
st_mtime (or st_mtime_ns + st_size), so a same-size replacement that
keeps the old timestamp (cp -p, rsync -t, a timestamp-pinning writer)
was never noticed:

- model_tools._tool_defs_cache_key: get_tool_definitions kept serving
  stale dynamic tool schemas / mcp_servers for the process lifetime.
- CLI mcp_servers auto-reload watcher (cli_tui_mixin seed +
  cli_info_mixin._check_config_mcp_changes): the replaced mcp_servers
  section was never reloaded. The seed is now _config_sig.
- tui_gateway/server._load_cfg_raw / _save_cfg (_cfg_mtime -> _cfg_sig):
  the raw-config cache served the stale document and the next _save_cfg
  would write it back over the on-disk file.

All three now use utils.file_signature like the rest of the PR. Tests that
reset the renamed module/instance attributes follow the rename; one
pinned-mtime replacement test per cache, red on the previous head.
Also corrects two stale type/comment annotations in hermes_cli/config.py
(_env_cache key shape, _RAW_CONFIG_CACHE record shape).

Review finding: three sibling config.yaml caches (tool-defs memo, CLI mcp watcher, TUI-gateway raw cfg) still compared mtime/size only.
2026-09-15 06:29:50 -07:00
teknium1 a1e7f74e64 fix: stat signatures cover inode + ctime everywhere config identity is cached
The cherry-picked commit extends the config/profile/MCP/managed-scope/completer/
OAuth/skills-manifest signatures. This commit finishes the class and trims it:

- `file_signature()` lives in `utils.py` next to the other stat/metadata helpers
  instead of `hermes_cli.managed_scope` (gateway/ and agent/ callers no longer
  reach into the managed-scope module for a generic stat helper).
- `hermes_cli/config_effective.py` was left comparing 2-/4-wide prefixes against
  the widened `_RAW_CONFIG_CACHE` / `_load_config_cache_sig` records, so
  `load_user_config_effective()` re-parsed on every call (3 parses for 3 calls on
  an unchanged file, 1 before); index by the new widths.
- Sibling caches keyed on the same (mtime, size) shape and reading the SAME files
  now use the helper: `load_env()` memo, `agent/skill_utils` raw-config and
  external-dirs caches, `hermes_cli/model_switch` alias identity, `agent/moa_loop`
  preset stamp, `hermes_cli/auth` global auth-store memo.
- Tests trimmed to one invariant each (pinned-mtime replacement invalidates; an
  unchanged file still hits), both red on origin/main.

Left alone on purpose: `tools/registry.py`, `tools/skills_tool_dedup.py`,
`gateway/status.py`, `hermes_cli/banner.py`, `hermes_cli/main.py`,
`hermes_cli/session_recovery.py` — those fingerprint source files, PID/lock files
or write to persisted on-disk caches shared across processes, where an inode/ctime
key would churn on every checkout/copy rather than catch a replaced config.
2026-09-15 06:29:50 -07:00
Kevin Rajan b797e9d7b3 fix(config): detect file replacements in stat-based cache signatures
Config caches, the profile re-scan watcher, the MCP reconciler, managed-scope
reads, the completer memo, OAuth heal marks, and the skill-manifest snapshot
keyed change detection on (st_mtime_ns, st_size) only, so a replacement that
preserved both (cp -p, rsync -t, timestamp-pinning scripts, sync clients) was
treated as unchanged and stale values were served until restart.

Add st_ino (fresh inode on atomic replace) and st_ctime_ns (cannot be
backdated via os.utime) to every signature via a shared
hermes_cli.managed_scope.file_signature() helper.

Fixes #111105
2026-09-15 06:29:50 -07:00
teknium1 97d3d75fdc fix: ignore every cron SQLite store and per-launch marker on flat installs
The flat-install block only ignored the bare cron/executions.db. All three
cron stores (executions, deliveries, notepad) are opened through
sqlite_util.open_db in WAL mode, so executions.db-wal/-shm exist whenever
the scheduler is live, and deliveries.db / notepad.db were not ignored at
all. `git stash push --include-untracked` therefore still unlinked the
live WAL/SHM and the whole deliveries/notepad stores under the running
gateway - the same mechanism this PR closes for the root state.db.

Switch the cron entries to the by-class shape already used at the root
(/cron/*.db plus -wal/-shm/-journal/retired-wal sidecars) and also ignore
the cron lock/heartbeat/output files and the per-launch root markers
(.update_check, gateway-starts.log, .clean_shutdown, active_profile,
.hermes_history, slack_tokens.json) that otherwise force every update
into the stash step. `git ls-files -i -c --exclude-standard` is unchanged
(no tracked file newly hidden). The test's FLAT_INSTALL_RUNTIME_STATE
gains one representative per new class; red before, green after.

Review finding: cron/executions.db-wal, cron/deliveries.db and cron/notepad.db were unignored and swept by the flat-install autostash.
2026-09-15 06:29:09 -07:00
teknium1 a172d76f2d fix(update): also ignore backups/ and the vault on flat installs; document the flat-install rule
`backups/` is where the pre-update snapshot the updater restores a swept
state.db FROM lives — leaving it unignored means the recovery copy is
swept together with the live database. `vault.key`/`vault.json.enc` are
the local secret vault.

Docs: website/docs/getting-started/updating.md explains that on a flat
install (checkout root == $HERMES_HOME) runtime state is git-ignored and
never enters the autostash.
2026-09-15 06:29:09 -07:00
teknium1 6f24245532 fix(kanban): gate create-with-parents like link; archived parent is terminal
create_task(parents=[open parent]) — the reporter's actual incident path —
parked the card in todo with only a `created` event, and kanban_create's
payload carried no `gated`, so the board still showed an unexplained todo
while only the link surface was fixed. create_task now appends the same
dependency_wait {reason: parent_not_done, parent} event and kanban_create
returns gated/gated_by, mirroring kanban_link.

link_tasks gated on `status != 'done'`, but _parents_satisfied and
recompute_ready treat `archived` as terminal: linking a ready child under an
archived parent demoted it to todo with a false parent_not_done event and the
next recompute promoted it straight back. Gate on not in ('done','archived').

Review finding: create_task(parents=...) emitted no dependency_wait/gated; link under an archived parent flapped ready->todo->ready with a false reason.
2026-09-15 06:25:42 -07:00
teknium1 89ef145254 fix(kanban): dashboard link reports the gate; docs; trim to two invariant tests
The dashboard's POST /links is the fourth writer of link_tasks (CLI, tool,
dashboard, plus the graph builder); return the same ``gated`` flag so every
surface that can create the deadlock can see it. Document the
``dependency_wait`` payload the link path emits and the delegation rule the
reporter derived (never link a support card under the card it unblocks).

Drops the CLI output test (a change-detector on prose); the two DB-level
invariants (event emitted on demotion / none for a done parent) stay.
2026-09-15 06:25:42 -07:00
Konstantin Khlopkov 35b1609fc3 fix(kanban): surface the link-time demotion of a ready child to todo
A ready child linked under an unfinished parent drops to todo with no
event and no operator signal; the only trace used to be claim_rejected
after a forced promote. Record a dependency_wait event when the demotion
fires, return the gate from link_tasks, warn in the CLI link command,
report gated in the kanban_link tool, and document the gate.
2026-09-15 06:25:42 -07:00
teknium1 fa915784cb fix: merge shipped skill categories instead of replacing them on update
The per-root merge in _copy_dist_payload treated the first level under
skills/ as the replace unit, but skills live at skills/<category>/<skill>.
A distribution shipping one skill inside a category rmtree'd the whole
category directory, wiping every skill the user (or hermes update) had
placed there — the exact scenario of issue #25120 that this PR claims to
fix. A shipped directory with no files of its own is now a container:
its children are merged one level down, and the replace unit becomes the
nearest directory that holds a file (a skill dir always holds SKILL.md).

A symlinked category (skills/devops -> shared dir) is a container too,
so the pre-write symlink check now covers it instead of silently
unlinking the link and dropping the shipped copy in its place.

Review finding: categorised skill payload rmtree'd skills/<category>/,
destroying sibling user and bundled skills; symlinked category was
unlinked rather than refused.
2026-09-15 06:22:15 -07:00
teknium1 2bf8b143bf test(distribution): trim salvage tests to two invariants
Fold the skills/cron/force-install cases into one per-root merge test and keep the
symlink-refusal test; the three separate tests asserted the same mechanism.
2026-09-15 06:22:15 -07:00
kshitijk4poor d4fa373017 fix(distribution): refuse symlinked targets before the first write
The symlink refusal lived in `_real_dir`, which runs per entry inside the copy
loop. By the time it fired for `skills/`, `SOUL.md`/`mcp.json`/`cron` had already
been replaced and `write_manifest` had not run yet, leaving a half-updated
profile that fails identically on every retry. Add a pre-flight pass over
`_owned_entries` that walks each target path (whole chain for directories,
parent chain for files) and raises before any mutation; `_real_dir` keeps its
check as the backstop.

The `.env.template` branch used a bare `shutil.copy2`, which writes through an
existing `.env.EXAMPLE` symlink into whatever it points at. Route it through
`_replace_entry` so the link is unlinked first, like every other owned entry.

The error text now tells the user what they can actually do (remove the link or
replace it with a real directory) instead of pointing at `distribution_owned`,
which is the author's manifest, not the user's.

The existing symlink test now changes every other shipped entry upstream and
asserts the profile is byte-for-byte untouched after the refusal.
2026-09-15 06:22:15 -07:00
kshitijk4poor cfc11f217a fix(distribution): refuse to replace a symlinked owned container
A user who points `profiles/x/skills` at a shared directory did so on
purpose. The merge path replaced such a symlink with a real directory
silently, discarding that configuration, where the base rmtree at least
failed loudly. Raise DistributionError naming the link instead; the
file->directory replacement stays because a stray file there is never
deliberate configuration.

`_remove_existing` now removes anything that lexists but is not a real
directory via unlink, so fifos and sockets no longer slip through.

Reviewer P1 on the original PR (symlink container must survive intact).
2026-09-15 06:22:15 -07:00
kshitijk4poor b2ef4cb8b1 test(distribution): keep the three payload-merge invariants
Keep the tests that pin behaviour a regression would break: user-added and
stale skill roots survive an update while shipped roots are replaced, a
forced reinstall keeps a user skill, and a user-added cron file survives an
update. The allowlist and symlink-transition cases exercise the same code
path (`_real_dir` + `_replace_entry`) and were duplicates of those.

Refs #110933
2026-09-15 06:22:15 -07:00
kshitijk4poor cc3517fbbc refactor(distribution): merge every owned directory per root, not just skills
The skills-only special case (a `preserve_skills` kwarg plus nested closures
keyed on the literal "skills") left cron/ and any future owned directory on
the old rmtree-then-copytree path, which deletes user-added files on update
and on --force reinstall. The rule is the same for all of them: a top-level
owned directory is a container of roots; replace only the roots the payload
ships, never the whole destination.

Apply that rule to every owned directory, drop the kwarg and the string
check, and lift the helpers to module level. `_remove_existing` keeps its
symlink-no-follow semantics and `_real_dir` replaces any symlink or file on
the destination path so the payload is never written outside the profile.
The fresh-install path had nothing to preserve, so it needs no flag.

Restore "skills" in DEFAULT_DIST_OWNED; removing it was never needed for the
fix. Add one invariant test: a user-added cron/ file survives an update.

Refs #110933
2026-09-15 06:22:15 -07:00
MKanso a4f0a5e035 fix(distribution): preserve skill roots safely
(cherry picked from commit c60656df80956e2a09670d0f7ad63b96d5d82ade)
2026-09-15 06:22:15 -07:00
MKanso 5029a38234 fix(distribution): preserve profile skills on update
(cherry picked from commit 67b1d1542d4e5b0d772cea991886afad5bce418d)
2026-09-15 06:22:15 -07:00
teknium1 95987fb85a fix: bypass the proxy on loopback HTTP CDP discovery and keep NO_PROXY=* intact
The websockets dials got proxy=None but the three HTTP /json/version dials
(CDP override discovery, is_browser_debug_ready used by Lightpanda and the
real-profile readiness check, and surviving-Chrome detection) still resolved
via getproxies(), so under a system/env proxy discovery fell back to the raw
http:// URL, readiness never fired and /browser connect reported not ready.
loopback_request_kwargs() sits next to loopback_connect_kwargs() and is used
at all three sites (ProxyHandler({}) opener for the urllib one).

add_loopback_no_proxy turned an operator NO_PROXY=* into '*,127.0.0.1,...',
which urllib/requests no longer treat as the wildcard, flipping bypass-all
configs into proxy-all. A wildcard in either casing now leaves env untouched.
is_loopback_host also accepts any loopback IP literal (127.x, ::ffff:127.0.0.1).

Review finding: HTTP /json/version dials still proxied loopback; NO_PROXY=* wildcard broken by append.
2026-09-15 06:16:38 -07:00
teknium1 c12a3397b1 fix(kanban): read the worker's card title from the board instead of a new env var
Follow-up to the salvaged commit from #111169 (@KoNit-K):

- Drop HERMES_KANBAN_TASK_TITLE. The worker already has HERMES_KANBAN_TASK
  and HERMES_KANBAN_BOARD/HERMES_KANBAN_DB pinned in its env, so
  maybe_auto_title reads the card title from the board itself (no new
  HERMES_* env var for non-secret config; the dispatcher and the
  delegation scrub list stay untouched).
- Unreadable or missing card: the session is named `Kanban task <id>`
  with zero auxiliary calls (the fallback the issue asked for; the
  #109743 seed left such workers untitled).
- The card title persists at `llm` authority via set_auto_title, so a
  manual /title still wins and the upgrade thread never starts.
- Tests trimmed to two invariants against a real board + SessionDB
  (card title, unreadable-card fallback), both red on origin/main.
2026-09-15 06:08:20 -07:00
KoNit-K 56555f88df fix(kanban): name worker sessions from task title 2026-09-15 06:08:20 -07:00
teknium1 738d63a34d fix(kanban): automatic stale-claim reclaims count toward the failure breaker
A claim that expired without a worker ever spawning (worker_pid NULL) was
reclaimed and immediately re-claimed on every dispatcher tick, with
consecutive_failures stuck at 0 — nothing could trip the breaker. Route the
reclaim through _record_task_failure (own txn after the reclaim commit, same
shape as enforce_max_runtime) instead of the salvaged raw counter increment,
so per-task max_retries / kanban.failure_limit and the gave_up event apply
and last_failure_error carries the stale lock. reclaim_task (operator path)
still resets the counter; the live-worker extend path never reaches it.

Trims the salvaged tests to one invariant that walks the breaker to its trip.
2026-09-15 06:06:32 -07:00
Kevin Rajan 48bcdf94e9 fix(kanban): stale-claim reclaim advances consecutive_failures
release_stale_claims() reclaimed expired claims without counting the
failure, so a claim that was never spawned (worker_pid NULL) spun forever
with consecutive_failures stuck at 0 and the failure_threshold breaker
could never trip (#111306). The reclaim transaction now increments
consecutive_failures atomically. Extend-live-worker and manual operator
reclaim paths are untouched (extend is not a failure; reclaim_task
deliberately resets the counter).

Tests: 3 new regression tests in tests/hermes_cli/test_kanban_db.py —
reclaim without spawn increments (red on base), reclaim with dead worker
increments (red on base), live-worker extend does not. Related kanban
suites: 79 passed, 1 skipped.

authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
2026-09-15 06:06:32 -07:00
teknium1 c810bc771f test(moa): trim to two invariants; map teo-nex attribution
Replace the CLI-subprocess parametrization (6 hermes processes per run) with a
direct normalizer invariant: non-finite cadence -> default cadence, non-finite
temperatures -> None while finite values (0, 0.75) survive.
2026-09-15 05:34:29 -07:00
teo-nex 0c492361fa fix(moa): reject non-finite temperatures before provider requests 2026-09-15 05:34:29 -07:00
teo-nex 367bbc3793 fix(moa): tolerate overflowing fanout configuration 2026-09-15 05:34:29 -07:00
teknium1 c1b7f28693 test(local-runtime): trim idle-probe salvage to two invariants
Keep the two tests that were red on origin/main (probe failure keeps the clock and
unloads once telemetry recovers; confirmed busy still resets). Drop the is_idle() bool
contract test — it pins behaviour that did not change — and fold the probe-failure log
call onto two lines.
2026-09-15 05:32:58 -07:00
Konstantin Khlopkov b83ee209d9 fix(local_runtime): keep the idle clock across failed telemetry probes
is_idle() treated any /slots or /metrics probe error as "busy", so sweep_idle()
popped the idle clock on every failed probe — one transient failure per sweep
and a resident model never reached IDLE_UNLOAD_S (21 GB pinned for hours with
zero requests).

Split the probe into a tri-state: _probe_idle() returns True (confirmed idle),
False (confirmed busy), None (probe failed). The sweep now resets the clock
only on a confirmed busy sighting; a failed probe keeps the existing clock and
logs at INFO, and unload still requires confirmed idleness past the threshold.
The public is_idle() bool contract is unchanged (probe failure reads as not
idle), so no caller outside the sweep ever unloads on a telemetry hiccup.

Closes #111154
2026-09-15 05:32:58 -07:00
Hukla b039e679ea fix(local-runtime): spawn llama-server with --no-ui
llama.cpp b10964 (the pinned build) renamed --no-webui to --no-ui, so the
router refused to start with an unknown-argument error. The direct-I/O flag
is already selected per-build by _direct_io_args on main, so this reduces
to the one remaining rename and asserts it in the existing spawn test.

Refs #111323
2026-09-15 05:31:16 -07:00
teknium1 02b460a929 test(cli): neutralize the Windows supervisor scope in the no-gateways fleet-restart simulation too
_run_pending_fleet_restart has three supervisor branches; #110709 covers launchd, this covers
the Windows branch (gateway_windows.is_installed -> restart) so a Windows developer with an
installed gateway service does not restart it from the test suite.
2026-09-15 05:26:25 -07:00
funky-xamarin 86b809934a test(local-models): isolate quickstart success paths from host memory 2026-09-15 05:26:25 -07:00
funky-xamarin 8c523923d9 test(update): isolate cold-start liveness from Desktop ownership 2026-09-15 05:26:25 -07:00
liuhao1024 4f10a59f53 test(cli): cover the launchd scope in the no-gateways fleet-restart simulation
test_run_pending_restart_true_when_no_gateways patches the PID scan and
the systemd unit listings but leaves the macOS launchd restart live, so
on a machine with a real hermes launchd fleet the restart phase drains
those units, reports the fleet restart incomplete, and the assertion
fails. CI never sees this because it runs Linux.

Patch _restart_macos_launchd_gateways to a no-op (the same pattern the
file's _patch_update_deps helper already uses) so the no-gateways state
covers the launchd supervisor scope too.

Fixes #110701
2026-09-15 05:26:25 -07:00
teknium1 55266997e7 test(config_home): drop the unavailable-parent-link case already covered by the existing diagnostics test
test_unavailable_directory_links_are_diagnosed_without_creating_targets already proves a link
with a missing target is refused, not materialized; keep the two new invariants (aliased
parent -> 0o700; operator link at the home boundary still owns its mode).
2026-09-15 05:25:39 -07:00
KeyArgo 56d2438a45 fix(security): secure HERMES_HOME when only an ancestor path is a symlink
`_directory_links()` walks every parent up to `/`, and `initialize_home()` skipped
`_secure_dir()` whenever that list was non-empty, so any symlinked ancestor disabled
home hardening and left `~/.hermes` plus its `cron`/`sessions`/`logs`/`memories`
subdirectories at the mkdir default `0o755`.

macOS makes that the normal case: the default temp root is `/var/folders/...` and `/var`
is a symlink to `/private/var` (as are `/tmp` and `/etc`), so any home reached through
those prefixes was never secured. Reported on macOS arm64 as `493 != 448` (0o755 vs 0o700)
in tests/cron/test_file_permissions.py; the same numbers reproduce on Linux by aliasing a
parent directory.

Only links at or below the home boundary are operator-owned. The home itself, or a link
inside it, still transfers permission ownership to the operator — behaviour pinned by
test_initialization_preserves_external_directory_modes and unchanged here. Availability
diagnosis is untouched: a link above the home whose target is missing is still refused
instead of being materialized on the underlying filesystem.
2026-09-15 05:25:39 -07:00
ruochu88s aa0b228c2c test(profiles): default-export tests write through the fixture home, not get_profile_dir("default")
get_profile_dir("default") re-resolves the platform-native root at call time; when pytest's
basetemp sits inside the operator's Hermes home the per-test sandbox counts as "under the
native home" and these two tests overwrote the live config.yaml / .env / MEMORY.md with their
stubs. Write through the profile_env fixture's home instead.

Taken from #111096 by @ruochu88s (the autouse native-home override and the session tripwire
from that PR are not included; the basetemp relocation in tests/conftest.py closes the class).
2026-09-15 05:24:39 -07:00
teknium1 6f69d170e5 fix: classify .env-routed keys, hyphenated headers and ${VAR} placeholders in config get
`_is_secret_config_key` matched only an exact set plus four suffixes, so
credentials that `hermes config set` routes to .env but end in bare `_KEY`
(FAL_KEY, VOICE_TOOLS_OPENAI_KEY, API_SERVER_KEY) or AWS_SECRET_ACCESS_KEY
printed raw with redaction on. Every `_is_env_config_key`-routed key is now a
credential unless its suffix names a non-secret shape (_URL/_HOST/_USER/_ID/
_DOMAIN/_SCHEME); `_key`/`_access_key` join the leaf suffixes. Header names are
folded `-`→`_` before matching so `mcp_servers.<s>.headers.X-API-Key` masks on
get and on the set echo. Bare `auth` leaves the exact set: `mcp_servers.<s>.auth`
is the documented `oauth` mode enum and rendered as `***`. Unresolved `${VAR}`
placeholders are printed as-is so the operator can see which env var the config
references.

Review finding: `config get` printed FAL_KEY / AWS_SECRET_ACCESS_KEY / X-API-Key raw, masked `auth: oauth`, and hid `${VAR}` placeholders.
2026-09-15 05:08:55 -07:00
teknium1 a9a8a3fa2e fix(config): hermes config get masks credentials on every path; --raw opts out
`hermes config get providers`, `config get providers.<p>.api_key`, `config get
<PROVIDER>_API_KEY` (the .env-routed branch) and `config get mcp_servers.<s>.env.X_API_KEY`
all printed the full credential. The agent runs this command from sessions whose transcripts
persist and get forwarded (a Gemini key surfaced in a Discord DM log), so `print` output is a
leak path the logging redactor never sees.

`get_config_value` now applies the structural masker used by `config show` before printing,
honouring `security.redact_secrets` (default on), with a `--raw` flag for operators/scripts
that need the real value. `_is_secret_config_key` extends the exact-name set with the same
`*_API_KEY / *_TOKEN / *_SECRET / *_PASSWORD` suffixes `_is_env_config_key` already routes to
.env, so env-map leaves under `mcp_servers.*.env` mask too, and the `config set` echo uses the
same predicate.

Slim redo of #84153 by @webtecnica (same direction: mask in get_config_value; dropped the
redact_url_query_params re-export and the separate redaction-enabled reader in favour of
agent.redact._redact_enabled, which already resolves the profile-scoped policy).

Fixes #110758
Fixes #84106
2026-09-15 05:08:55 -07:00
teknium1 60d94fd8f4 fix(state): warn when an existing WAL state.db sits on a virtiofs/9p mount; doctor + docs
d8dcdfd620 (v2026.9.14) made apply_wal_with_fallback refuse to ENABLE WAL
on a fresh database whose directory is on a cross-VM bind mount (virtiofs/9p),
but a database that was already WAL on such a mount kept WAL — correctly, we
never live-downgrade under other openers — and emitted nothing. The operator
in #110848 ran exactly that shape (Podman applehv virtiofs bind mount) and got
"database disk image is malformed" within a minute with no prior signal.

- apply_wal_with_fallback: in the on-disk-WAL branch, log a once-per-process
  ERROR ('cross_vm_fs_existing_wal') when the DB file is on a cross-VM
  filesystem, naming the two remedies (offline PRAGMA journal_mode=DELETE
  after stopping every process + database.journal_mode: delete, or move the
  database to a native/named volume). The fresh-DB refusal is unchanged.
- hermes doctor: _report_database_journal_modes flags a WAL database on a
  cross-VM filesystem with check_warn and the same remedy (ranked above the
  WAL-reset exposure warning; the exposure bookkeeping is kept).
- docs: docker.md gains "Filesystem requirements for state.db in containers";
  configuration.md's database comment no longer implies operators must set
  delete by hand on virtiofs.

Detection stays /proc/self/mountinfo-based (runs inside the Linux container
on macOS/Windows hosts). locking_mode=EXCLUSIVE is deliberately not adopted:
gateway, cron and workers open state.db concurrently.

Fixes #110848
2026-09-15 05:02:30 -07:00