Commit Graph

5936 Commits

Author SHA1 Message Date
Solitud1nem 1e8f829f1a fix(model): show copilot-acp in pickers when its executable resolves
The overlay loop in list_authenticated_providers() checks every way a
provider might be authenticated — env keys, the auth store, the
credential pool, even Claude Code's external token files — but never
asks the one question that matters for an external_process provider:
does the executable resolve? copilot-acp has no key or token by design
(the spawned `copilot --acp --stdio` brings its own auth), so has_creds
stayed False and the filter dropped it from every picker. Funny enough,
five lines further down the same loop has a dedicated copilot-acp
branch for fetching its model ids — it just never got a chance to run.

Availability now comes from get_auth_status(), the same source
`hermes model` and the auth status endpoints already use, so the CLI
and GUI agree on what 'configured' means for external-process
providers.

Fixes #63662
2026-09-02 20:51:07 +05:30
kshitijk4poor 95f62ca3bf fix(cron): validate failure_deliver at preflight and dashboard update lanes
Follow-up to the failure_deliver salvage (#100375):

- _preflight_check_delivery also checks the failure lane, so a typo'd
  failure_deliver platform blocks at config-validation time instead of
  surfacing only when a failure occurs — exactly when the notice must
  not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
  deliver (text normalization; empty clears the optional override
  instead of coalescing), closing the one update path that could write
  an unnormalized value into jobs.json.

4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
2026-09-02 20:16:14 +05:30
Victor Kyriazakos c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium 73f68362b3 fix(sessions): auto-prune state.db by default (90d) and gate VACUUM on freelist ratio (#54189)
Flip the state.db retention defaults per Teknium's decision on #54189:

- sessions.auto_prune: false -> true. A stock install now prunes ENDED
  sessions inactive for retention_days at CLI/gateway/cron startup
  (at most once per min_interval_hours). Open, pinned and mid-turn
  sessions are never deleted; the only open rows touched are stale
  automation sessions (#100903 sweep), which are closed, not deleted,
  and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
  the file: PRAGMA freelist_count / page_count must exceed 25%
  (AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
  min_vacuum_interval_days throttle. Pruning a few small sessions on a
  dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
  Unknown ratio (pragma read failure) falls back to the time throttle.

Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.

Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
2026-09-02 07:26:52 -07:00
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
chelsealong 600b9c7e7e fix(gateway): scope force-reload hook re-registration to its own home
re_register_config_hooks() cleared the entire process-global idempotence
set on every force-reload, so a profile-local plugin force-reload dropped
another live profile's ledger key without touching its still-registered
callback — the next registration call for that profile then appended a
duplicate. Scope the clear to the reloading profile's own home, and give
outbound webhooks the same force-reload restoration shell hooks already
had, since unload() wipes both from the shared _hooks dict.
2026-09-02 07:00:13 -07:00
Teknium bb5a1b7b80 fix(plugins): platform_actions fails closed when the active profile cannot be resolved
Follow-up to #85245: on a profile-resolution exception the facade fell
through to `_authorization_adapter(platform, None)` — the DEFAULT bot —
which is exactly the fallthrough the profile-aware lookup exists to
prevent. Return `adapter_not_registered` instead. Tests now drive the
real `GatewayAuthorizationMixin` ladder (no hand-written stand-in) and
are trimmed to one positive + one parametrized fail-closed case.

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
pierrenode 956f493cf1 fix(plugins): route platform_actions through profile-aware adapter resolution
PlatformActions._resolve_adapter() unconditionally read
runner.adapters — the default profile's adapter registry — with no
awareness of multiplex/Team-Gateway secondary profiles. Every other
adapter-resolution path in this codebase (GatewayAuthorizationMixin.
_authorization_adapter, and the plugin message-injection path that
uses it) is careful about this split: a secondary profile's adapters
live in runner._profile_adapters[profile], and a stamped secondary
profile with no registry entry must fail closed rather than fall back
to the default profile's adapter of the same platform — falling back
would send replies out the wrong bot.

ctx.platform_actions had none of that: a plugin scoped to one profile
calling add_reaction/set_thread_title in a Team-Gateway deployment
would silently act through the DEFAULT profile's adapter/bot identity
instead of its own profile's, whenever the default profile also ran
that platform.

Route _resolve_adapter() through the same _authorization_adapter
lookup, resolving the calling profile via
hermes_cli.profiles.get_active_profile_name() (which reflects the
per-task HERMES_HOME override multiplex profiles already propagate).
Falls back to the old bare runner.adapters lookup only when the
runner predates _authorization_adapter (defensive, not expected in
practice).
2026-09-02 06:59:11 -07:00
Teknium 1398c0f5ca fix(update): stage-and-swap the Desktop rebuild so a failed pack never removes the working app
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).

Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.

- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
  `_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
  retry purge and the integrity self-heal only ever clear the staging tree,
  never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.

Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.

Closes #86443

Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
2026-09-02 06:53:20 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium 62a4599f89 fix(cron): keep profile scope on the standalone fallback pool; desktop ticker stands down for profiles with their own gateway (#100489)
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
2026-09-02 06:27:24 -07:00
berg 357062f39e fix(dashboard): multiplex cron fire URL uses the default profile's listener port
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
2026-09-02 06:27:24 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium 2475335443 fix(config): keep .env publishes inside the routed profile scope under multiplex (#88441)
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.

Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
2026-09-02 06:19:03 -07:00
Teknium b53c50cf8a fix(config): resolve ${VAR} config refs through the profile secret scope (#84079)
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).

Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
2026-09-02 06:19:03 -07:00
Patryk Kopycinski 1eb71756af toolchain-self-improve: route API_SERVER_KEY to profile env 2026-09-02 06:19:03 -07:00
Teknium 0a6aa7cce1 fix(env_loader): log routed-scope dotenv skip once per home; port single-profile control test
Follow-up to the #77592 salvage: emit a once-per-home debug line where the
multiplex guard skips the process-global dotenv load (requested on #77562),
and port the single-profile control test from #77970 so the guard is pinned
to the multiplex flag rather than the home override alone.

Co-authored-by: DonShelly <25538402+DonShelly@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Lester Liang 1aa62ceb45 fix(security): isolate multiplex dotenv reloads 2026-09-02 06:19:03 -07:00
Teknium 16370ae539 fix(tui_gateway): setup.status / setup.runtime_check answer for the requested profile
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.

`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.

The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.

Supersedes #94147 (python half)

Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
Teknium af86ad0479 fix(doctor): reuse resolved memory config at Memory Provider section
Follow-up to salvaged #100677: the file checks read the memory section via the
run's hermes_home while the Memory Provider section re-read config with no
argument (module-global HERMES_HOME). Resolve once and reuse so both sections
report against the same config.
2026-09-02 06:17:25 -07:00
fangliquanflq 21cebbfd68 fix(doctor): honor disabled built-in memory stores
Salvaged from #100677. Fixes #100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
2026-09-02 06:17:25 -07:00
Teknium 632078bca7 feat(cli): OSC 9 + Warp OSC 777 notifications ride on the bell flags
Extend _ring_bell() so display.bell_on_prompt / bell_on_complete also
emit terminal-native desktop notifications from the same six call sites
(clarify, clarify batch, approval incl. computer_use, sudo password,
secret capture, turn complete). No new config keys.

- OSC 9 (ESC ] 9 ; body BEL): Ghostty / iTerm2 / Kitty / WezTerm raise an
  OS notification; unknown terminals drop it. Body is "Hermes: <context>"
  with C0 controls and DEL stripped. Written to /dev/tty (prompt_toolkit's
  stdout wrapper can buffer/strip raw escapes) with a sys.stdout fallback.
- Warp OSC 777 warp://cli-agent (agent "hermes", event permission_request
  / stop, compact JSON mirroring build-payload.sh). Gated on
  TERM_PROGRAM=WarpTerminal + WARP_CLI_AGENT_PROTOCOL_VERSION + the
  should-use-structured.sh broken-build floor (stable/preview builds at or
  before v0.2026.03.25.08.24.*_05 rejected). Never raises.

Salvages #58957 and #100805.

Co-authored-by: glitchbunny0 <glitchbunny0@proton.me>
Co-authored-by: harsha-usethread <harsha@usethread.io>
2026-09-02 06:17:10 -07:00
Teknium c2954c8934 feat(model-catalog): picker catalogs refresh every 20 minutes, gateway keeps them warm
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.

- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
  an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
  disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
  it off-thread every TTL window, so every surface on the machine reads
  a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
2026-09-02 06:16:54 -07:00
Gille 5180601a6a perf(cli): dispatch serve without the full parser tree 2026-09-02 06:06:47 -07:00
jinglun010 7aff724e56 feat(desktop): add display.resume_last_session config toggle (#60812)
Config default (true) plus the Appearance-settings strings for a
"Reopen Last Chat on Launch" switch. Salvaged from PR #60816 onto
current main (defaults moved to config_defaults.py since the PR).
2026-09-02 05:56:54 -07:00
Teknium 552159d222 feat(cli,tui): collapse bell_on_clarify/approval into display.bell_on_prompt
One key covers every blocking prompt modal: clarify (single + batch),
dangerous-command approval (incl. computer_use), sudo password, and
secret capture. CLI gets a _ring_bell() helper shared with
bell_on_complete; TUI rings on clarify/approval/sudo/secret .request
events (isTTY-gated). 'hermes config' Bell summary shows both flags.
2026-09-02 05:34:35 -07:00
Turgut Kural 3082a34669 feat(cli,tui): add display.bell_on_approval + fix eslint error
- display.bell_on_approval (default false): same BEL mechanism as
  bell_on_complete, rings when a dangerous-command approval prompt
  opens (_approval_callback / approval.request event). Complements
  bell_on_clarify from the previous commit.
- fix(ui-tui): eslint curly error in useConfigSync.applyDisplay
  (if without braces) that failed the CI JS & TS checks job.
2026-09-02 05:34:35 -07:00
Turgut Kural ef6d3367a6 feat(cli,tui): add display.bell_on_clarify — terminal bell on clarify prompts
Same BEL mechanism as display.bell_on_complete (\a / \x07), gated by
display.bell_on_clarify (default false). CLI rings in _clarify_callback
and _clarify_callback_batch before _paint_now(); TUI rings on
clarify.request when bellOnClarify && stdout.isTTY. Docs in
cli-config.yaml.example and website/docs/user-guide/configuration.md.
2026-09-02 05:34:35 -07:00
1052326311 ea9924b09c fix(container): honor config.yaml multiplex_profiles at boot
hermes_cli/container_boot.py resolved multiplexing from the
GATEWAY_MULTIPLEX_PROFILES env var only, while the gateway runtime
resolves env -> config.yaml -> default. A deployment enabling
multiplex_profiles via config.yaml alone therefore auto-started every
named profile's gateway slot at boot, which then crash-looped in the
double-bind guard against the multiplexing default gateway.

Resolve through load_gateway_config().multiplex_profiles (the shared
resolver, so env override precedence is preserved) and fall back to the
env var only when config loading fails.

Salvage of #85437 (test module trimmed to config-only + env-override).

Fixes #85413
2026-09-02 05:34:28 -07:00
Teknium aaa34b0e08 fix(desktop): model picker no longer hardcodes --global; one persist policy for every surface (#90235)
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).

Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).

Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
  session-only rule: when neither model.default nor model.provider is
  configured yet, persist. This preserves #86414's first-pick motivation
  for CLI, gateway and Desktop alike. With a default configured, a plain
  pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
  Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
  lists --global.
- Docs: desktop.md picker note + slash-commands /model row.

Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
2026-09-02 05:33:33 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Aldo 6ff9426d22 fix(anthropic): track the current fast-mode model matrix (Opus 4.8 / Opus 5)
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):

- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
  only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
  requests silently run at standard speed and bill standard rates
  (usage.speed: 'standard'). Today's allowlist therefore shows 4.6
  users a fast toggle that does nothing, while denying it to the two
  models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
  select fast inference via the model field and are explicitly
  excluded from the param gate.

Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.

## How to test

scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q

113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
2026-09-02 05:33:13 -07:00
Teknium 44a57921c5 fix(providers): hide phantom -cn picker rows lit only by shared intl keys; give alibaba-token-plan-cn its own key var
- alibaba-coding-plan-cn / alibaba-token-plan-cn keep the shared intl key vars
  as ordered fallbacks after their dedicated *_CN_API_KEY, so users who set
  ALIBABA_CODING_PLAN_API_KEY / ALIBABA_TOKEN_PLAN_API_KEY for the CN endpoint
  keep working (the PR as filed dropped them).
- list_authenticated_providers hides a '-cn' row whose only lit key vars are
  ones it shares with its non-CN sibling, unless that CN provider is the
  configured model.provider. With only the shared key: one row, not two;
  DASHSCOPE_API_KEY alone: 3 alibaba rows, not 4.
- Docs: environment-variables.md, providers.md.
2026-09-02 05:32:20 -07:00
Teknium a81e503866 fix(kanban): reap notify subscriptions for stale blocked tasks too
purge_stale_done_notify_subs only matched status='done', so a task the
circuit breaker parked in 'blocked' kept its notify-sub rows forever on
boards that never archive. Widen the predicate to done OR blocked while
keeping the existing age clause; backlog/ready cards are idle, not
abandoned, and stay exempt (test_gc_spares_reopened_task_even_when_old).
Watcher comment/log and docs updated to say done/blocked.

Closes #100955

Co-authored-by: itsflownium <itsflownium@users.noreply.github.com>
2026-09-02 05:32:10 -07:00
Teknium 5c6cbbc1be fix(loops): pause /loop --until on a blocked verdict; trim redundant gate condition and duplicate test
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
2026-09-02 05:32:01 -07:00
itsflownium 1bd9fce6cb fix(kanban): judge unachievable goals as blocked, never done 2026-09-02 05:32:01 -07:00
Teknium d3e2ace1dd fix(profiles): profile delete refuses to kill another profile's gateway (#89315)
`hermes profile delete` read the target profile's gateway.pid raw and
SIGTERMed it. When that pid file was poisoned by a sibling profile's gateway
(the #89315 shape), deleting profile A killed profile B's running gateway.

- gateway/status.py: `_pid_record_belongs_to_profile()` helper — a pid
  record whose recorded home differs from the expected profile home is not
  ours; legacy records without a home prove nothing and are left alone.
- hermes_cli/profiles.py: `_stop_gateway_process` refuses (and says so)
  when the record belongs to another profile; still stops its own gateway.

The stop/restart paths in hermes_cli/gateway.py did not need a guard:
`get_running_pid()` already filters cross-profile records and unlinks the
poisoned pid file before any kill can happen — verified live; the test for
that path now pins the real contract (returns False, other process alive,
poisoned pid file gone).

Live repro (unpatched main): `_stop_gateway_process(tim_home)` -> "Gateway
stopped (PID ...)" and the OTHER profile's process exits -15. After: "Refusing
to stop PID ..." and the process stays alive. 8 tests; sabotage (guard
removed) fails 1.
2026-09-02 04:53:59 -07:00
Teknium 45b0d8cab5 feat(gateway): one gateway.trust_env key controls aiohttp proxy-env honoring at every adapter site (#48820 bug 3)
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.

- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
  resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
  auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
  teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
  bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.

Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
2026-09-02 04:13:02 -07:00
Teknium 2f44998353 fix(dashboard-auth): a non-JWT bearer is "not my token", not "provider unreachable" (#94558)
NousDashboardAuthProvider._verify_jwt (and the identical hunk in the
self-hosted OIDC provider) folded EVERY PyJWKClient failure into
ProviderError, which the gate translates to HTTP 503
{"detail":"Auth provider 'nous' unreachable"}. That branch fires for
jwt.DecodeError('Not enough segments') — i.e. the bearer is not a JWT at all
(an opaque peer key, a legacy token, garbage) — and for PyJWKSetError (JWKS
fetched fine, foreign kid). Neither involves reaching Portal, which is why
the hosted sjc agents in #94558 returned a fast, well-formed 503 that
survived token re-mint and instance restart while Portal was healthy.

Add one shared classifier, hermes_cli.dashboard_auth.classify_jwks_lookup_error:
only PyJWKClientConnectionError (transport) and an unexpected bare
PyJWKClientError stay ProviderError; DecodeError / PyJWKSetError /
InvalidTokenError become InvalidCodeError so verify_session() returns None
and the middleware proceeds to the next provider / refresh / 401 exactly as
the protocol documents. Both providers now use it.

Live repro (real NousDashboardAuthProvider against a local reachable JWKS
server; and the real gated web_server app): before — opaque bearer ->
ProviderError "JWKS lookup failed: DecodeError('Not enough segments')" ->
503 unreachable; after — verify_session() -> None, gated GET /api/auth/me
with the opaque bearer -> 401; a real JWT against an unreachable JWKS still
-> ProviderError (503).

This does not add /api/v1/message to the public-path allowlist (#94579):
that route has no verifier in this repo, so bypassing the gate would leave a
state-changing ingress fail-open. The correct fix is classification, which
also covers every other opaque-bearer surface.

Refs #94558
2026-09-02 01:15:58 -07:00
kshitijk4poor 2adb1a4ea6 refactor(agent): clone the review snapshot once at the spawn chokepoint
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.

Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
2026-09-02 13:28:41 +05:30
kshitijk4poor 26f0de23cf fix(agent): clone the /refine snapshot too, not just the automatic review
Widen #100802 to the two explicit review entry points. The CLI and gateway
/refine handlers built their own snapshot with a shallow list(), which
aliases the nested tool_calls/content containers of the live history. The
review fork sanitizes its transcript in place (sanitize_tool_call_arguments
rewrites function["arguments"]), so a /refine could rewrite the parent's
persisted transcript exactly like the automatic review could (#100795).

Both sites now use _clone_background_review_messages, the same structural
clone the automatic review uses. Regression tests drive the real handlers
and assert the snapshot shares no containers with the live transcript.
2026-09-02 13:28:41 +05:30
Teknium 37f3ba110a fix(auth): auto-heal single-use OAuth grants already forked across profiles (#100339)
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.

`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.

Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.

Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
2026-09-02 00:58:29 -07:00
Teknium 3038493ee6 fix(auth): never fork single-use OAuth grants across profiles (#100339)
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:

1. `hermes profile create --clone-all` and the dashboard/TUI
   `mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
   verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
   drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
   `providers.<id>` device-code blocks, and the PKCE singleton file; API
   keys are still copied. The clone reads the root grant through the
   existing credential-pool root fallback.

2. A named profile with no local rows BORROWS the root grant via
   `read_credential_pool()`'s fallback, but every persist
   (`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
   rows into the profile's own auth.json — materializing a fork on the first
   rotation. `persist_pool_entries()` now routes borrowed single-use rows
   back to the root store (update-only, under the root lock; never falls
   back to a local copy). A borrowed `hermes_pkce` rotation commits its
   singleton to the root `.anthropic_oauth.json`, the borrower never prunes
   root-seeded rows it cannot see the backing file for, and
   `hermes -p <profile> auth add` persists only the profile's own rows.

Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.

Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).

Closes #100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 00:58:29 -07:00
Teknium 00a7115a02 fix(cron): make cron push-notify configurable (cron.delivery.notify) and surface UNVERIFIED live deliveries in cron list/doctor
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.

An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.

Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
2026-09-02 00:56:52 -07:00
Teknium df4b3733ba fix(cron): every last_status consumer renders delivery_failed explicitly (dashboard badge, Desktop inspector, /cron list, docs)
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):

- web dashboard CronPage: last_status was never rendered at all — a
  delivery_failed job showed a green 'scheduled' badge and only a small red
  'delivery: ...' line. New pure cronLastResult() helper maps the closed
  literal set to tones (ok=success, delivery_failed/blocked_config=warning,
  error/unknown=destructive) and the card now shows an amber
  'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
  literal; routineLastResult() spells out each one ('Ran, but delivery
  failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
  appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
  on this branch; no consumer compared == 'ok' for success apart from the
  cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
  detail field carries the reason.

Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
2026-09-02 00:52:58 -07:00
Justin Wilson 8fd76fd1d6 fix(cron): surface delivery_failed instead of last_status ok
A successful agent run whose delivery failed used to persist
last_status=ok and bury the failure in last_delivery_error. CLI list
painted that as green and the run looked identical to a quiet success.

Record last_status=delivery_failed instead, keep last_delivery_error,
do not increment failure_streak, and teach cron list/doctor not to
treat it as ok.

Fixes #83993
2026-09-02 00:52:58 -07:00
Teknium f9bca5a0d0 fix(update): reconcile serve/dashboard runtimes in their own vocabulary and escalate survivors (#100479)
Widen the two salvaged fixes (#100490, #100493) to the whole class:

- match_runtime_outcomes: serve/dashboard rows never borrow gateway
  bookkeeping at ANY site — not just the bare hermes-gateway unit name
  (#100490) but also relaunched_profiles / externally_supervised_profiles
  and the profile-substring unit match (hermes-gateway-work credited the
  'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
  units (exact names, scope prefix tolerated) or, when the caller passes
  the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
  feed the Phase-2 reconciliation, so a surviving unmanaged serve is
  'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
  remedy instead of 'hermes gateway restart', which cannot reach it.

Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
2026-09-02 00:42:53 -07:00
TwotNguyenVN 5733d55f77 fix(update): warn surviving pre-update serve and dashboard runtimes on success (#100479) 2026-09-02 00:42:53 -07:00
chelsealong 806878612a fix(update): stop crediting unmanaged serve runtimes with a gateway's restart
match_runtime_outcomes() treats any default-profile runtime as covered
once the bare "hermes-gateway" unit restarts, regardless of the
runtime's own kind. An sshd-spawned `serve --isolated` backend (no
systemd unit, supervisor "manual-serve") shares the default profile
and gets silently marked "restarted" even though its own PID was never
touched — so the #91277 Phase 2 unaccounted-runtime tripwire never
fires for it and `hermes update` reports success while it keeps
running pre-update code (#100479).

Restrict the "hermes-gateway" special case to kind == "gateway" so a
serve/dashboard runtime under the same profile falls through to
"unaccounted" instead of borrowing the gateway's outcome.
2026-09-02 00:42:53 -07:00
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00