Commit Graph

2761 Commits

Author SHA1 Message Date
teknium1 43e7e830fd fix(tools): fold ~/.local/bin into the POSIX PATH completion siblings, tests + docs
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.

Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.

Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.

Fixes #111778
2026-09-15 18:49:29 -07:00
DavidMetcalfe 3e833fd56f docs(browser): cover logins across scheduled and unattended runs
Real-profile browsing and cron's per-job toolsets are both documented, but
nothing connected them: browser.md never mentioned scheduled runs and cron.md
never mentioned login state, so the constraints of driving a login-gated site
from a cron job were only discoverable in source.

Adds a Scheduled and unattended runs subsection to the real-profile section:
the browser.use_real_profile prerequisite (off by default), the credentials a
login form or a fresh 2FA challenge needs saved ahead of time, the auth re-sync
behaviour and its Windows consequence. Plus a pointer from cron.md's toolset
section.
2026-09-15 18:46:54 -07:00
teknium1 1cfa892db1 fix(desktop): make the zsh probe test legs visible and run them on CI
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.

- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
  passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
  probe's grandchildren) in the SSH connection guide.
2026-09-15 18:46:32 -07:00
teknium1 6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1 80f76edfaf fix(web): rescue eligibility asks the provider whether the ring was walked
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
2026-09-15 18:43:05 -07:00
teknium1 45a4db2225 fix(web): key extract cache on metadata.sourceURL too, pin redirect case
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.

Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
2026-09-15 18:42:38 -07:00
teknium1 339fa6d918 fix(gateway): bounded redacted result preview on tool.completed run events
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.

The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.

Salvages #111821 (@KoNit-K), part of #111815.

Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
2026-09-15 18:42:10 -07:00
teknium1 4c4d54554d docs(kanban): stop-nudge scope covers delegate children and in-process cron
State that the turn-end guard fires only for the dispatcher-owned worker, not
for delegate_task children or cron runs that inherit HERMES_KANBAN_TASK.
2026-09-15 18:41:19 -07:00
teknium1 996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
teknium1 0f38867a1b docs(dashboard): say lifetime-capping proxies still close the chat socket
The keepalive only defeats idle timeouts; proxies that cap total socket
lifetime (some tunnels) still close it and the chat reattaches on its own.
2026-09-15 18:38:16 -07:00
teknium1 a8a36c461b docs(dashboard): describe the chat PTY keepalive and hidden-tab reconnect pause 2026-09-15 18:38:16 -07:00
teknium1 bd63866253 fix(kanban): give a finished worker a grace window before the terminal reaper signals it
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.

Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
2026-09-15 18:35:32 -07:00
teknium1 aa5817d9be fix(kanban): reap workers that outlive their finished run
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).

Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.

Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.

Fixes #111791
2026-09-15 18:35:32 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1 c7f4bc5bd7 fix(kanban): text dispatch output and both "dispatcher stuck" warnings name the hold reason
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.

- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
  task plus rate_limited / skipped_locked / memory_pressure for one or more
  DispatchResults, so the CLI daemon and gateway warnings share one wording:
  `Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
  tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.

Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>

Part of #111910
2026-09-15 18:34:11 -07:00
teknium1 f7ea39481a fix(kanban): dashboard estimate calls declare a relay-affinity key too
The dashboard's estimate endpoints make the same headless auxiliary call
as specify/decompose but never bound an affinity scope, so they still sent
no x-opencode-session and the OpenCode Go relay answered 400
MissingSessionID (#112043). Declare kanban:<task_id> for an existing task
and a stable kanban:estimate key for the create dialog (no task yet),
unless a scope is already bound.

Test: _run_estimate captured header None before; now kanban:t_1 /
kanban:estimate and nothing leaks past the call.
2026-09-15 18:33:43 -07:00
teknium1 1a8d922003 docs(providers): note the per-task x-opencode-session key for headless Kanban aux calls 2026-09-15 18:33:43 -07:00
teknium1 9013fcdc87 fix(cron): an unreadable cron toolset restriction fails the run instead of granting every tool
_resolve_cron_enabled_toolsets returned None when _get_platform_tools
raised, and AIAgent reads None as "load every toolset": a malformed
platform_toolsets block (or a stale-module import error after an update)
turned the operator's cron restriction into the full default set, with
only a log warning. Unattended jobs process untrusted text, so that is a
privilege widening, not a safety net (#111380).

The resolver now raises a RuntimeError naming the cause; run_job's
existing failure path records it on the job (last_error, failure streak,
incident) and the agent is never constructed. Per-job enabled_toolsets
(unknown names included) and the MCP merge path never touch the
platform resolver and are unchanged; the disabled-toolset resolver has no
fail-open branch.

Live: platform_toolsets: oops -> before: run ok, enabled_toolsets=None,
98 tool names selected; after: run fails "Cron toolset resolution
failed, so this run was refused rather than given every tool", agent
never constructed. normal / unknown-per-job / mcp-merge shapes: identical
before and after.

Co-authored-by: Austin Bell <10687162+robertaustinbell@users.noreply.github.com>
2026-09-15 18:31:55 -07:00
teknium1 336227bf00 refactor(gateway): move the on-demand s6 slot helper into service_manager, trim tests
Reshape of the cherry-picked fix from #104194:

- The SOUL.md gate + `register_profile_gateway(start_now=False)` now live in
  `hermes_cli/service_manager.py::register_unregistered_profile_gateway`, next to the s6
  manager and `_profile_dir_for_gateway_service` it needs, instead of a private reach-in
  from the 6.5k-line `hermes_cli/gateway.py` facade. The facade only decides "start
  repairs, stop/restart re-raise" and keeps ONE error handler (S6Error is a RuntimeError;
  register's ValueError/RuntimeError/OSError surface as the same `✗` + exit 1).
- Tests trimmed from four to two invariants: start on a real profile registers `down`
  and then starts; stop on an unregistered profile / start on a directory without
  SOUL.md keep the original error and mint nothing (parametrized). Dropped: the
  registration-failure traceback test (covered by the single except clause) and the
  duplicate no-marker/stop split. The test now resolves the profile dir through the real
  HERMES_HOME mapping instead of monkeypatching `_profile_dir_for_gateway_service`.
- Docs: docker.md multi-profile section says `gateway start` inside the container
  registers a slot for a profile created from the host.
2026-09-15 18:30:05 -07:00
teknium1 49b9bbb6fc fix(desktop): an exit without a window reveal no longer writes hermes.desktop
gnome-shell moves a launched ShellApp from STARTING to STOPPED when the
startup-notification sequence completes or times out (mutter, ~15 s), not
when the process exits. finish() healing right after an exit-without-reveal
(boot crash, early quit) therefore wrote the entry during STARTING — the
exact #111906 arming condition. Only the reveal byte from Electron heals
now; the wake byte finish() writes just unblocks the reader. A skipped heal
is picked up by the next terminal/updater launch or revealed grid launch.
2026-09-15 18:29:37 -07:00
teknium1 5d2acf7066 test(desktop): pin deferred launcher-entry writes and the terminal-launch control
- cmd_gui with DESKTOP_STARTUP_ID: the entry is installed only after the fake Electron
  writes the reveal byte to HERMES_DESKTOP_READY_FD (spawn precedes install).
- cmd_gui without DESKTOP_STARTUP_ID (control): install still precedes the spawn.
- DeferredDesktopEntryInstall.finish(): an app that exits without revealing heals
  exactly once, after the exit.
- linux-launcher-ready: one byte, fd closed, variable consumed; garbage/absent = no-op.

Docs: user-guide/desktop.md explains when the entry write happens for grid launches.
2026-09-15 18:29:37 -07:00
teknium1 da9810387d docs(config): scope the .env routing claim to registered env settings
Only names in OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS and the platform *_HOME_CHANNEL /
*_ALLOWED_USERS suffix family route to .env; other documented ALL-CAPS names still land in
config.yaml as top-level scalars with a notice. Say so instead of 'every documented
environment variable' (#111848 stays open for the remaining names).
2026-09-15 18:28:49 -07:00
teknium1 0e63a1bc5c fix(config): refuse an unknown path under a known section before writing
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.

Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.

`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.

Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:28:49 -07:00
teknium1 70e4938c07 fix(config): route every registered env setting through .env from config set/get/unset
`hermes config set FEISHU_HOME_CHANNEL oc_x` wrote the top level of config.yaml
while the platform setup flows and /sethome write the same name to .env via
save_env_value, so two writers fed two readers: the gateway bridges the yaml copy
into the environment only when .env lacks the name, one-shot CLI readers never
bridge, and the two copies diverged silently (#111848). Only credential-shaped
names were routed to .env because `_is_env_config_key` is the provider-credential
predicate.

Follow-up to KoNit-K's cherry-picked fix (#111850), which routed the
`setup_hidden_env` suffix family: the predicate now lives in the topical sibling
`hermes_cli/config_env_routing.py` and covers every bare name Hermes itself
registers as an environment variable (OPTIONAL_ENV_VARS, _EXTRA_ENV_KEYS — "env
var names written to .env" — plus the setup-hidden suffixes for plugin adapters
nobody enumerated), so `*_ALLOWED_USERS`, `WHATSAPP_MODE`, `MATRIX_PASSWORD` and
the rest of the adapter-saved family take the same file. `set` and `unset` also
drop a stale same-named top-level config.yaml copy so the reporter's drift cannot
come back, and `get` resolves .env first then that copy — the gateway's own read
order. Provider credentials keep the credential_lifecycle rotation path.

Docs: environment-variables.md tip, hermes_cli/AGENTS.md config rule.
2026-09-15 18:28:49 -07:00
John Paul Soliva b4a6958129 fix(doctor): warn when database.journal_mode=delete never applied to a WAL database
`hermes doctor` printed the same informational "WAL journal mode" line for a database
that is still WAL although the operator configured `database.journal_mode: delete`. The
runtime never live-downgrades an existing WAL database (a downgrade under open
connections corrupts it) and says so only once per process in the gateway log, so the
one surface operators check told them they were protected when every process was still
writing WAL on the filesystem they configured `delete` for (#111729).

Doctor now compares the configured mode (`resolve_journal_mode`) with the on-disk header
and warns `<db> is in WAL mode despite database.journal_mode=delete`, naming the
never-live-downgraded rule and the one-time offline `PRAGMA journal_mode=DELETE`
remedy; a vulnerable SQLite still counts the database as WAL-reset exposed. This check
outranks the cross-VM hint (whose remedy, "set journal_mode: delete", is already
applied). A configured `wal` is unchanged.

Core hunk ported from #104714 (@jonpol01); its holder enumeration
(`foreign_state_db_holders(include_scan_gaps=True)` + per-PID report) is not included
here — that half is a 600-line change under separate review.
2026-09-15 18:28:26 -07:00
teknium1 e2fb6cc765 docs(update): receipts record the SQLite runtime repair step 2026-09-15 18:28:03 -07:00
teknium1 71aa0d635e fix: setup gateway skips the standalone service for a multiplex-served profile
`hermes -p <profile> setup gateway` (and `hermes setup` / `hermes import`) reach the
service step through `ensure_gateway_service`, which only knew "is THIS profile's
unit running" — a satellite served by the default multiplexer has no unit of its
own, so the step printed "Installing the gateway background service ..." and
registered a launchd plist / systemd unit that the #97120 start guard then refused,
leaving a stray dead service the user had to find and remove by hand (#111958).

Route both setup surfaces through one shared predicate: `_served_profile_needs_no_service`
wraps `named_profile_served_by_running_multiplexer` (the same probe `profile create`,
cron liveness and the run/start/install guards use), prints the "already served"
note and returns True so `ensure_gateway_service` and the `hermes gateway setup`
wizard (#111962's hunk) skip the install. Default profile and non-multiplex hosts
are unchanged.

Adds the invariant for the `ensure_gateway_service` path (served → no install,
unserved → still installs). Docs: multi-profile-gateways.md names the skipped step.

Co-authored-by: kvnloo <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:27:35 -07:00
teknium1 0a3e792942 fix(worktree): judge no-remote repos against the local trunk instead of reaping everything
`hermes worktree prune` (and the startup/cron pruner) classified every clean tree in a
repository with no remote as "clean and fully merged/pushed" and force-deleted its branch,
even when the branch carried commits that exist nowhere else. `audit_branches` returned []
in the same repos, so a unique local-only branch was invisible to the audit as well.

Root cause: `_worktree_has_unpushed_commits` answered False when `refs/remotes` was empty
("nothing to be unpushed against") and both consumers — `worktree_gc._classify_tree` and
`worktree_ops._classify_prune_candidates` — read False as "safe to reap".

The preceding commit (#111897) flips that branch to True, which is safe but also means a
no-remote repo can never reclaim anything (a tree sitting at trunk, or squash-merged into
it, stays "unpushed" forever because `_worktree_commits_all_merged_upstream` finds no
origin/* base). This commit replaces the unconditional True with a real baseline:

- `_worktree_local_trunk`: `main`/`master`, else the branch checked out in the main
  worktree; None when no trunk exists.
- `_worktree_merge_base_ref`: origin/HEAD|origin/main|origin/master, falling back to the
  local trunk ONLY when the repo has no remote-tracking refs at all. Single resolver used
  by `_worktree_commits_all_merged_upstream` and `worktree_gc.audit_branches`.
- `_worktree_has_unpushed_commits`: with no remote refs, `git log HEAD --not <trunk>`;
  no trunk -> True (preserve), matching the docstring's fail-safe promise.

Net effect in a no-remote repo: unique work is kept ("unpushed commits not found
upstream"), trees at/merged into the local trunk still reap, branch audit reports unique
branches as keep and merged ones as delete, and `git branch -D` can only run on a branch
whose every commit is reachable from or patch-equivalent to the trunk.

Tests: the salvaged no-remote keep test now uses a shared `local_repo` fixture; a control
test proves trunk-merged trees still reclaim and the branch audit reports in the same
repo; `test_merged_predicate_fails_safe_without_upstream` now pins both halves of the
contract (trunk resolves -> merged; no trunk at all -> False/preserve).
2026-09-15 18:27:06 -07:00
teknium1 0e0692240a test: fold the named-source control into one clone-all test; document the skipped trees
Trim the three contributor tests from #101340 to two invariants: the
root-only ignore test now also asserts that a named profile used as
source keeps its own models/ (the exclusion is gated on the default
root), and the end-to-end create_profile(clone_all=True) test stays.
Same assertions, two tests — the salvage bar.

Docs: profile-commands.md and profiles.md list models/, runtimes/ and
node/ among the --clone-all exclusions so users know why a clone from
the default profile does not carry the local-model weights.
2026-09-15 18:26:38 -07:00
teknium1 5f20bb8b63 test: trim timezone validation tests to two invariants; document the doctor check
Collapse the five contributor tests from #101345 into two: one for the
error class (unloadable IANA name, non-string value) and one for the
silent class (valid, blank, missing, None, and no tz database at all).
Same coverage, half the surface — the salvage bar is two invariant tests
per fix.

Docs: the `timezone` section now says `hermes doctor` reports a value the
runtime cannot load, so users know where the typo surfaces.
2026-09-15 18:26:15 -07:00
teknium1 536a07bec2 fix: sessions list prints a footer when --limit cuts the listing
`hermes sessions list` applied --limit inside the SQL query and rendered
whatever came back, so a user with 35 sessions saw 20 rows and a prompt
and had no way to tell the page was cut. The lister now asks for one row
past the cap, drops that probe row, and ends a truncated page with
"… more not shown (use --limit N to see more)". `--limit 0` is LIMIT 0
(no rows), so the copy never suggests it.

The footer goes through one shared helper, `cli_output.print_truncated`,
and the three "... N more" / "… N more" / "… and N more" variants already
in sessions_cmd.py (export dry-run preview, never-active cleanup, prune/
archive preview) now use it too, so every capped listing in the command
reads the same. Migrating checkpoints / curator / skills search / console
listers is follow-up work.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:25:46 -07:00
teknium1 d1bd778a5e fix(checkpoints): clear-legacy trims to two real-path tests, aligns failure line with the log
Replace the three monkeypatch-heavy tests from the salvaged commit (which
faked clear_legacy's return dict, so they could not catch the manager hunk
regressing) with two invariant tests that drive the real cmd_clear_legacy
against a temp checkpoint base: an undeletable legacy-* dir yields exit 2
plus the "Could not delete" line while the archive stays on disk; a clean
sweep keeps exit 0 and the unchanged success line. The green-path guard is
harvested from #111789.

Reword the CLI failure line to "Could not delete N archive(s) (see logs)."
so it matches the manager's WARNING wording and the text proposed in
#111776, and document the exit code in the CLI reference.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:24:22 -07:00
teknium1 cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
teknium1 b027a4658e fix: trim DeepInfra reasoning salvage to the invariant and document it
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.

- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
  config -> top-level field table, and the transport main-turn path with
  ``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
  two-directional reasoning control

Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.

Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
2026-09-15 18:22:40 -07:00
teknium1 ee49b7d25d fix(agent): file-mutation footer states failed edits, not "files were NOT modified"
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.

- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
  and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
  (mtime_ns, size) snapshot; a later success clears every entry with the same
  identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
  entries whose file changed since the failed call, so a receipt-less
  mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
  way the file tools resolved them.

Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.

Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:22:12 -07:00
teknium1 0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1 c1bbcf9712 fix(agent): hint-preview truncation log names the real remedy; invariant tests
Follow-up to the two salvaged commits (#111777, #111781 by @KoNit-K):

- agent/prompt_builder.py::_truncate_content — with queue_warning=False the
  logged line no longer tells the operator to "pin a larger
  context_file_max_chars, or use a larger-context model": the subdirectory
  hint cap is a constant neither knob raises. It now points at the read_file
  recovery the marker already discloses.
- tests/gateway/test_startup_environment_probe.py — replace the
  call-detection test with the behavioural invariant: an oversized SOUL.md in
  HERMES_HOME and a warm-up leave the truncation-warning queue empty for the
  next default-executor task (the api_server turn path runs on that executor
  without copy_context, which is how the boot warning reached a foreign
  session).
- tests/agent/test_subdirectory_hints.py — fold the new drain assertion into
  the existing oversized-hint test (same fixture) and pin that the log carries
  no context_file_max_chars advice.
- agent/AGENTS.md, website/docs/.../context-files.md — the hint cap is 32,000
  (docs said 8,000) and is fixed; document that it is logged, not surfaced as a
  chat warning.
2026-09-15 18:19:26 -07:00
teknium1 9b3042902e fix: trim hygiene bound salvage to two invariant tests and document the knob
Salvage follow-up to the cherry-picked #112055 hunk (@Finn763):

- tests: fold the landed-commit control into the turn-hold-miss test and drop the
  two extra end-to-end tests (sub-limit identity is already pinned by the pure
  contract test) so the fix carries exactly two invariant tests.
- gateway/run_turn.py::bound_model_input_without_hygiene: drop the isinstance /
  limit<=0 defences — the caller always passes the transcript list and a
  _knob()-validated positive int, so the checks could never fire.
- docs: configuration.md now says hygiene_hard_message_limit also bounds the
  model payload on any turn where hygiene did not land (payload-only, disk
  untouched, landed summaries adopted as-is).

Co-authored-by: Noor-Sol <187156657+Noor-Sol@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:18:58 -07:00
teknium1 c5c71ea1ad fix(docs): document the 10 s SSE keepalive comment for custom parsers 2026-09-15 18:18:30 -07:00
teknium1 eba9b5551c docs: note the automatic non-streaming fallback for contentless SSE frames
The streaming section documented only the manual model.streaming escape
hatch; users hitting a degraded gateway need to know the session flips to
non-streaming on its own and why the warning appears.
2026-09-15 18:18:08 -07:00
teknium1 eb562b10ad docs(plugin-catalog): allow maintainer-curated sweep entries alongside owner submissions
Teknium ruled that maintainers may add batches of community plugins from a
reviewed sweep instead of waiting for each owner to submit. Rule 5 and the
user-guide checklist now say so, and give authors the explicit right to adjust
or remove a swept-in entry via their own PR.
2026-09-15 12:55:47 -07:00
teknium1 651168d7b2 docs(credential-pools): env: rows may name any variable, not only the declared one 2026-09-15 11:50:48 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1 804707bea6 fix: checkpoint store gc never runs inside a tool call or gateway startup
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.

The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.

- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
  drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
  pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
  so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
  calls it from the housekeeping tick (last chore), the CLI from a daemon
  thread. Nothing on either startup path waits for git.

Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
2026-09-15 10:57:16 -07:00
teknium1 d84ece48b8 fix(mcp): Figma OAuth login completes despite the omitted iss parameter
Figma's authorization-server metadata advertises
authorization_response_iss_parameter_supported and its redirect omits iss,
so the mcp SDK's RFC 9207 check discarded every valid code and login never
finished. For that one issuer the provider fills a missing iss with the
discovered issuer and warns; a mismatching iss still fails and every other
server keeps the strict rule.

Fixes #111135
2026-09-15 09:29:00 -07:00
teknium1 043632e283 feat(desktop): hideable profile rail with a statusbar profile dropdown stand-in
For people who run profiles as bots, the colored profile strip at the sidebar
foot duplicates the sessions list (community request). Add a persisted
`hermes.desktop.profileRailVisible` preference (on by default) toggled from the
Sessions view menu ("Profile rail"), the shell right-click menu, ⌘K
("Toggle profile rail") and an unbound `view.toggleProfileRail` keybind.

While the rail is hidden the statusbar grows a `ProfileSwitcher` dropdown
beside the gateway switcher ("This device ⌄ · Profiles ⌄") offering the same
choices the rail does: this gateway's profiles, All profiles, every other
gateway's agents in fleet mode, New / Import / Manage. It also answers the
`profile.create` hotkey the rail used to own, so no door is lost.
2026-09-15 08:50:18 -07:00
teknium1 6bc3628ce8 fix(desktop): drag gateway/profile groups by their header, not only the hidden handle
The Gateway & profile sidebar sections carried dnd-kit listeners on the lead
glyph alone, and that glyph only reveals its grabber on hover — so a press on
the row's name or empty space did nothing and users fell back to the ⋯ menu's
Move up / Move down. Bind the sortable to the whole header (same shape as a
project row), with the handle and the ⋯/caret cluster keeping their own
gestures; a sub-threshold press on the label is still the fold click.

Live: Playwright pointer drag on the header label reorders sections (before:
unchanged, after: reordered); handle-drag still works. Test drives dnd-kit's
keyboard sensor on the label and is red on the old binding.
2026-09-15 08:50:18 -07:00
Robin Fernandes 1034215ae8 docs(free-tier): drop the rehearsal page and its references
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes d89cacc25f chore(free-tier): keep the rehearsal server out of the repo; the docs page explains the stand-in instead
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30