Commit Graph

5161 Commits

Author SHA1 Message Date
Teknium fdd8d75ba0 fix(gateway): make ws keepalive and orphan-reap grace config-driven (#79635)
- New dashboard.ws_ping_interval / dashboard.ws_ping_timeout defaults
  (20.0/20.0) in DEFAULT_CONFIG; hermes_cli/web_server.py reads them for
  non-loopback binds. Loopback keeps ws_ping=None (event-loop stalls must
  never kill a healthy local connection).
- New dashboard.ws_orphan_reap_grace_s (20.0): tui_gateway/server.py's
  _WS_ORPHAN_REAP_GRACE_S now resolves from config via
  _resolve_ws_orphan_reap_grace(); the HERMES_TUI_WS_ORPHAN_REAP_GRACE_S
  env var is kept as an internal override for backward compat and wins
  when set.
- tests/test_ws_keepalive_config.py: real load_config against a temp
  HERMES_HOME yaml — defaults, propagation, deep-merge, env override,
  invalid-value fallback.
2026-08-23 17:43:39 -07:00
Teknium 12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
fangliquan bd13d593a9 fix(dashboard-auth): correct authentication docs anchor 2026-08-23 16:09:14 -07:00
fangliquan 7cc92cb134 fix(dashboard-auth): replace stale insecure guidance 2026-08-23 16:09:14 -07:00
nbxuhk 0e038425db fix: gateway lifecycle guards gate on process ownership, not inherited env
The terminal tool lifecycle guard and the gateway stop/restart CLI
guards keyed on the raw _HERMES_GATEWAY=1 env marker, which every
gateway descendant inherits (and importing gateway.run sets it too).
CLI/TUI agent sessions were falsely blocked from documented gateway
management commands. Gate on _is_supervised_gateway_process() instead,
which requires owning the live gateway PID file.

Salvaged from PR #92196 (guard half) by @nbxuhk. Fixes #92560.
2026-08-23 15:49:52 -07:00
Teknium d861fbe550 fix(plugins): dispose persistent auth registrations on plugin disable and re-discovery drop
Follow-up to the #91701 salvage: persistent registrations survive a routine
unload-all, but must not outlive their plugin.

- Targeted unload (plugin disable/uninstall) now gathers persistent rows
  from the ownership ledger and disposes them.
- Unload-all parks live persistent handles in _persistent_carryover;
  discover_and_load(force=True) evicts the ones whose plugin did not
  re-register the same (kind, key) — superseded handles are dropped
  without disposal so a same-object re-registration stays live.
2026-08-23 15:49:44 -07:00
rainbowgits b2ade2388d fix(dashboard-auth): keep the provider registry alive across per-home plugin-manager unloads
The dashboard auth registry is process-global, but a bundled auth provider
was registered under the per-home plugin manager's scope and enrolled in that
manager's reverse-order teardown. A per-home manager is unloaded routinely
(profile-scoped dashboard activity, forced re-discovery), and that teardown
disposed the registration — emptying the auth registry for the whole process
and permanently disabling sign-in until restart.

Register dashboard-auth providers in the process-global slot as persistent
host-owned registrations kept out of per-home manager teardown, so a routine
unload can no longer disable authentication process-wide. Registration upserts,
so a forced re-discovery (e.g. a password change) still rotates the provider in
place. The test-only manager reset now clears the auth registry too, since
persistent registrations deliberately survive unload.

Fixes #91701
2026-08-23 15:49:44 -07:00
Gille b7cb321223 fix(dashboard): preserve placeholder cwd fallback 2026-08-23 15:21:48 -07:00
Gille 3a81721147 fix(dashboard): preserve exported terminal overrides 2026-08-23 15:21:48 -07:00
Gille 5a85c4a77b fix(dashboard): scope terminal config to selected profile 2026-08-23 15:21:48 -07:00
kshitijk4poor 58a8cc7dd1 chore: consolidate turn_wait_seconds into the merged bot_mode config section
Rebase onto main (post-#93102) left two bot_mode dicts in
config_defaults.py — later duplicate key silently wins in a Python
literal, so envelope_ttl_seconds would have shadowed turn_wait_seconds'
section. Single section now carries both keys.
2026-08-24 02:08:53 +05:30
kshitijk4poor a07ada235c fix(test): deterministic delivery-spawn sentinel + fold bot_mode into agent config tab
Two CI failures: (1) the target_busy test's global subprocess.run patch
recorded unrelated gateway-init git calls (rev-parse/ls-remote) as the
delivery spawn — now local_delivery_command is monkeypatched to a
sentinel argv so only the real delivery path counts; (2) the new
bot_mode config section is single-field, tripping the dashboard
no-single-field-categories rule — folded into the agent tab via
_CATEGORY_MERGE (same as #93102's fix).
2026-08-24 02:08:53 +05:30
kshitijk4poor ac3f9a2dc4 feat(bot-mode): per-profile turn lock — concurrent deliveries queue instead of racing (#93091) 2026-08-24 02:08:53 +05:30
kshitijk4poor 3ac6306655 fix(web): fold single-field bot_mode config section into agent tab
The new bot_mode.envelope_ttl_seconds default created a one-field
dashboard category, tripping test_no_single_field_categories. Merge it
into the agent tab via _CATEGORY_MERGE like code_execution et al.
2026-08-24 01:05:08 +05:30
kshitijk4poor b96369212c feat(bot-mode): envelope TTL + offline fast-fail for bot relay (#93091 item 2) 2026-08-24 01:05:07 +05:30
webtecnica f293e7206b fix(dashboard): detect stale code after hermes update and refuse model picker with clear 503 (#86207) 2026-08-23 06:36:57 -07:00
Teknium 706f33d424 feat(update): sibling profiles' configs migrate with the fleet — no more silent version drift (#20438/#54926/#79048)
The shared checkout serves every profile, but hermes update migrated
only the active profile's config.yaml. Siblings kept their old
_config_version until their (correctly restarted, post-#91378) gateway
hit a config shape the new code couldn't read — the last unabsorbed
substance from the Phase-2 restart-swarm audit (#20438 earliest, 2026
field repro on #79048: sibling at v33 vs v37).

_migrate_sibling_profile_configs(): per sibling home, scope config
reads/writes via the context-local HERMES_HOME override (ContextVar —
never os.environ), check version, run the NON-INTERACTIVE safe
migration; prompt-requiring settings stay for the profile's own next
interactive session (same contract as gateway-mode). Broken profiles
are skipped without blocking the sweep; override always reset.

Sabotage-verified; live E2E in a fresh process with real drifted
config files: v12→v38 and v25→v38 on disk, provider preserved, the
documented #81946 personality-reset migration correctly applied to
siblings too, never-configured profile untouched, active home
untouched, second run idempotent.
2026-08-23 05:12:59 -07:00
Teknium 18b7fc82b6 feat(update): the plan is now the restart worklist — every planned runtime must be accounted for (#91277 Phase 2)
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).

- update_inventory: restart_via becomes a machine-readable mechanism id
  (systemd|launchd|desktop|manual) — THE policy table as data; display
  derived via describe_restart_mechanism. match_runtime_outcomes()
  reconciles every planned runtime against the restart phase's
  bookkeeping (restarted/stopped/failed/unaccounted);
  report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
  land in the receipt (runtime_outcomes); any unaccounted runtime
  escalates exactly like a STALE/DOWN fleet row (exit 1).

Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
2026-08-23 04:45:29 -07:00
Jack Lau 1bf93660f1 fix(update): verify launchd is supervising the gateway after a restart
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).

_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.

That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.

Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.

A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.

Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.

tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.
2026-08-23 04:45:29 -07:00
Jack Lau dfcef70061 fix(update): stop a gateway we cannot relaunch instead of leaving it on stale code
Fixes #88654.

After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:

    restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
    if restart_mode is None:
        continue

A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:

    cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'

with no operator signal of any kind.

Two changes.

_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.

When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
2026-08-23 04:25:18 -07:00
Teknium 6cb1085d3d fix(gateway): /p/<profile>/ on a non-multiplex gateway fails closed instead of serving the owner profile
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.

Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.

Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.

Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.

Fixes #91583 (defect 2). Repro and live validation by @kubaboski.
2026-08-23 03:57:47 -07:00
Teknium 9e18197745 fix(cli): one-shot runs linger for notify_on_complete background processes so Bot Mode replies survive parent exit
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).

Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.

- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
  — bounded, interrupt-safe wait over pending notify_on_complete
  sessions; reconciles orphaned-pipe exits (#17327) each pass so a
  wedged reader cannot burn the full bound; KeyboardInterrupt aborts
  the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
  session flush / cleanup (covers -q and -Q, i.e. the DM recipient
  shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
  kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.

Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.

Fixes #90879
2026-08-23 03:56:37 -07:00
Teknium 231e613d3d fix(peer): resolve hidden canonical Bot Chats in hermes peer dm
Bot Mode always hides canonical 'Bot Chat' sessions, but _find_bot_chat's
GET /api/sessions listing used the default include_hidden=False path, so
the existing hidden row was invisible, _ensure_bot_chat tried to create a
duplicate, and the peer DB's UNIQUE(title) guard rejected it — DM failed.

- api_server: GET /api/sessions now accepts an exact-title lookup
  (?title=...) and honors include_hidden=1 ONLY alongside a title filter,
  so canonical hidden rows resolve without exposing a blanket hidden
  listing on the client surface. The title needle is pushed into SQL
  (search_query) so old hidden rows outside the recency window are found.
- peer dm client: _find_bot_chat sends title + include_hidden=1; older
  peers ignore the unknown params and degrade to today's behavior.
- Clear diagnosable error on the older-peer duplicate-create rejection,
  naming the hidden canonical chat and the PATCH hidden:false workaround.
- Unit tests (hidden resolution, no duplicate create, older-peer error,
  older-peer visible fallback) + real-gateway E2E over a real state.db.

Root-cause analysis and regression recipe by @kubaboski in #91583.

Fixes #91583
2026-08-23 03:56:28 -07:00
Teknium 0c14f060db fix: import managed_python_env at the git-path site; assert the managed-env contract in the repair test
The salvaged commit called managed_python_env() at the git-path sync
without an in-scope import (UnboundLocalError on every git update — CI
red). The repair test pinned the raw {**os.environ, VIRTUAL_ENV} dict, a
change-detector on exactly the construction #83914 replaces; it now
asserts the managed-env contract.
2026-08-23 03:55:14 -07:00
Teknium fbfdb9312b fix(update): widen UV-env isolation to the sibling dependency-sync sites
The salvaged fix covered the git-path sync; the same raw-os.environ
construction existed at the main update path and the interrupted-install
recovery path. All three now build their uv env via managed_python_env()
(#83914 class — same bug, all sites).

A/B-proven with real uv: poisoned UV_PYTHON/UV_SYSTEM_PYTHON steers the
merge-base construction into the hijacker's interpreter (VERDICT:
HIJACKED); the managed construction installs into the install's venv
(VERDICT: ISOLATED). Compose-checked with #92824's stale-VIRTUAL_ENV pin:
isolation + pin together install into the running interpreter on the
site-packages shape.
2026-08-23 03:55:14 -07:00
suntech-wang 6ce145f38f test(update): lock managed uv-env isolation regression
Address review feedback:
- Add two unit tests asserting the update's uv_env contract: third-party
  UV_PYTHON_INSTALL_DIR is dropped, managed pins (UV_MANAGED_PYTHON=1,
  UV_NO_CONFIG=1) are set, VIRTUAL_ENV points at this install's venv, and
  the managed store stays under .hermes-runtime.
- Drop the inline dated comment in favor of intent description.
2026-08-23 03:55:14 -07:00
suntech-wang 08f5a0a98b fix(update): isolate pip install from third-party UV env vars
uv respects UV_PYTHON_INSTALL_DIR from the process environment. When a
third-party app (e.g. WorkBuddy) sets a User-level UV_PYTHON_INSTALL_DIR,
the update's uv pip install can target the wrong interpreter and fail
installing extras, leaving the venv entry-point shims missing. Use the
official managed_python_env() isolation (drops VIRTUAL_ENV/PYTHONPATH/
UV_PYTHON, forces UV_PYTHON_INSTALL_DIR to .hermes-runtime/python,
UV_NO_CONFIG=1) and then point VIRTUAL_ENV at this install's venv.
2026-08-23 03:55:14 -07:00
Teknium 10f0d2278b feat(desktop): client-direct voice — use the active profile's STT/TTS keys from the desktop, no audio relay
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.

Backend:
- tools/voice_client_config.py: single resolver; per-provider client
  wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
  elevenlabs-tts). Server-host-only providers (local whisper, edge,
  command/plugin) and missing credentials resolve to {mode: relay}.
  xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
  same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).

Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
  profile) with 60s TTL, provider-direct STT + TTS calls, sentence
  cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
  first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
  startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
  below it. Barge-in via the same stopVoicePlayback sequence bump.

Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.

Docs: voice-mode.md client-direct section ships in this PR.
2026-08-23 03:54:52 -07:00
Teknium 5c1a304ce8 fix: derive the pinned interpreter's Scripts dir via venv_bin_dir (#76105 lint)
The salvaged _interpreter_scripts_dir hand-rolled the Scripts/bin layout,
which the AST lint-test in test_update_zip_two_phase forbids — route it
through the canonical hermes_constants.venv_bin_dir instead, with the
interpreter's own dir as fallback for non-venv layouts.
2026-08-23 02:16:33 -07:00
mrmixx-max 33e813da17 fix(update): address review on stale-VIRTUAL_ENV pin
- _is_uv_command: detect 'python -m uv'/'python -m uvx' and launcher
  wrappers, not just a uv basename (review: naive check missed module form)
- _insert_python_pin: never duplicate a caller-supplied --python (review:
  last-wins ambiguity)
- _interpreter_scripts_dir: when pinning to sys.executable on Windows with
  no project venv, quarantine the running interpreter's Scripts dir so the
  hermes.exe shims uv rewrites are actually unlocked (review: quarantine
  path diverged from pinned interpreter)
- tests: rewritten to repo English convention; added python -m uv,
  --python-guard and Windows quarantine-target cases (5 total)
2026-08-23 02:16:33 -07:00
mrmixx-max 13f9d18e33 fix(update): pin uv installs to the running interpreter when VIRTUAL_ENV is stale
When Hermes is installed via pip / site-packages (e.g. the Windows
installer), PROJECT_ROOT is the interpreter's site-packages directory and
PROJECT_ROOT/venv is never created. The update and interrupted-install
recovery paths still set VIRTUAL_ENV=PROJECT_ROOT/venv, so uv fails with
'Failed to inspect Python interpreter from active virtual environment'
before installing anything — leaving the install partially updated.

Detect the nonexistent VIRTUAL_ENV in the shared dependency-install helper
and pin uv to the running interpreter (uv pip install --python
sys.executable) instead, matching the fix already applied to lazy-deps
(#83335) and the ZIP update path (#71510).
2026-08-23 02:16:33 -07:00
Teknium 7b89e17774 fix(model): single exact eligibility predicate for -900k variants; reject ineligible aliases
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
  synthesis, context resolution, /model validation, and wire stripping.
  Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
  plus date-shaped 5.6 snapshots — family-prefix matching removed, so
  non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
  aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
  API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
  hidden-slug soft-accept, and accepts valid variants missing from a
  stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
  openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
  ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
  wire model.
2026-08-23 02:14:35 -07:00
Teknium 63a9c26fbe feat(model): Codex GPT slugs default back to 272K; explicit -900k picker variants opt into the verified large window
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).

- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
  advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
  gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
  the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
  gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
  the wire (main transport + auxiliary Responses adapter), and pricing
  aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
2026-08-23 02:14:35 -07:00
hsearcy 503d863fcd fix(install): never strand hermes.exe when a Windows update fails
On Windows the updater renames the live `hermes*.exe` shims aside
(`hermes.exe.old.<unix-ms>`) so uv can write replacements. When that quarantine
succeeds but the install then fails, the recovery path could leave the install
with no `hermes` on PATH at all — unrecoverable in place, because the command
that would repair it IS `hermes update` (#75584).

Restoring a quarantined shim happens at three sites: the updater, the
early-recovery installer, and the startup sweep's orphan rescue. Each was a
single un-retried rename whose OSError was swallowed in silence, while the
OUTBOUND quarantine rename already retried a lock. That is backwards — a failed
quarantine merely aborts an update, a failed restore removes `hermes` from
PATH — and the two sites that had messages had already drifted apart.

- `_early_recovery.restore_quarantined_shims()` is now the single
  implementation: retry ladder, one recovery message, returns the pairs it
  could not restore. It lives in the stdlib-only module that both `main` and
  `_install_repair` already import, so the layers cannot drift again. A pair is
  not a failure when the original reappeared or the quarantine file vanished —
  two processes sweeping the same orphan must not produce a spurious error.

- `_cleanup_quarantined_exes` unlinked every `*.exe.old.*` on each invocation.
  When the original shim was already missing, that .old file was the ONLY
  surviving copy — deleting it converted a one-rename recovery into a full
  reinstall. It now rescues the orphan through the shared helper instead, and
  leaves anything inside a 15-minute grace window alone so it cannot destroy a
  concurrent update's in-flight quarantine.

- Ordering is by the PARSED `.old.<unix-ms>` stamp, not the raw filename.
  Lexicographic ordering only tracks recency while every stamp shares a digit
  width; a stray `.old.999` sorts above a 13-digit epoch-ms stamp and would be
  the copy rescued onto the live shim name.

- Names whose suffix does not parse as int-ms are not ours: never rescued,
  never deleted. The sweep should not destroy files whose provenance it cannot
  establish, and they are not produced by the quarantiner.

The stamp is read from the filename rather than st_mtime because `rename`
preserves the original shim's mtime, which records when uv wrote the shim —
days earlier, in general — not when it was quarantined. A regression test pins
that distinction.

Messages go to stderr: the sweep runs on EVERY hermes invocation and
`hermes acp` speaks JSON-RPC on stdout.

Scope note: `_quarantine_running_hermes_exe` is deliberately byte-identical to
main here. Why the outbound rename fails in the first place (the launcher
holding its own image without FILE_SHARE_DELETE) is #88121's subject; this is
the net underneath, covering the case where quarantine SUCCEEDS and the install
dies afterwards. The two touch disjoint functions and can merge in either order.

Reproduced and verified on Windows 11 (26200), Python 3.11.15: stranded the
shims, confirmed a normal `hermes` invocation now rescues the orphan instead of
deleting it, and confirmed an exhausted rescue prints the recovery command.
16 new tests; 35 pass across the four quarantine suites.
2026-08-23 01:47:26 -07:00
Teknium 3c44cd0c67 docs: record why the reap grace window exists in the reaper docstring
Follow-up for salvaged PR #91994 -- without this note a future cleanup
pass could read the age gate as dead weight.
2026-08-23 00:52:50 -07:00
Ishman82 b44c2bdab2 fix(desktop): spare concurrently starting backends
Desktop writes backend.lock.json only after a remote profile backend announces readiness. Concurrent Bot Mode profile starts could therefore see their young siblings as unowned PPID-1 processes and mutually reap them, causing SSH reconnect storms and stale turn leases.\n\nProtect unregistered backends for a bounded startup grace period, fail closed when age cannot be read, and cover young, unknown-age, boundary, lock-owned, and old-orphan behavior.
2026-08-23 00:52:50 -07:00
Teknium 8804e78354 feat(dashboard): Desktop and dashboard read the update receipt instead of inferring success (#91277 Phase-1 bullet 3)
Builds on @mrsucesso's durable-marker recovery (previous commit):

- GET /api/hermes/update/receipt — the full durable receipt (steps,
  skips, gateway restart outcome, fleet matrix) + compact summary; the
  authoritative update-outcome record (written by every run since
  #91283, including refused/failed).
- /api/actions/hermes-update/status now attaches the receipt summary,
  and when BOTH the in-memory registries and the update.log marker are
  gone (dashboard restarted + log rotated — the #81193 state), a
  finished receipt reports the outcome: success→0, partial→1. A
  still-running receipt proves nothing (clients keep polling).
- Desktop (updates.ts): the apply poll reads the attached receipt — a
  finished receipt whose run started at/after this apply is
  authoritative, replacing timeout-based failure inference across the
  update's restart gap ('Backend update failed' on successful updates,
  #81193; 'boot failed' during update restarts, #87359).

Live-verified: real uvicorn server + real UpdateReceipt writer (the
exact code hermes update runs) over real HTTP — receipt endpoint 200
with summary; #81193 state (no registries, no marker) reports success
from the receipt alone; partial receipt with a DOWN fleet row maps to
exit 1 (no false success).
2026-08-23 00:50:41 -07:00
Mauricio Ruiz 091106092b fix(dashboard): recover update success after restart
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
2026-08-23 00:50:41 -07:00
Franci Penov e366df6889 fix(cli): treat a fork's upstream sync as an update
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.

Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.

Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.

steps still being skipped afterwards.

Refs #73108
2026-08-23 00:19:46 -07:00
Teknium 1684877868 fix(update): a gateway killed by the restart phase and never replaced now fails the fleet check (DOWN row)
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.

- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
  'down' row only when it was alive at update start AND its runtime
  status still claims a running state. Rollout-safe: no snapshot (old
  callers), clean stops, startup failures, and stale records from
  long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
  (exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.

Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
2026-08-22 23:46:06 -07:00
Josh Holt 5f0a8f8739 fix(picker): harden keyless provider gate logging and credentials path validation 2026-08-22 23:25:43 -07:00
Josh Holt b9f17ba3f1 fix(picker): scope Vertex explicit-config to Hermes signals, not ambient ADC
Address review feedback: the gate reused has_vertex_credentials(), which also
returns True for an ambient GOOGLE_APPLICATION_CREDENTIALS path. That var is
commonly set globally for unrelated GCP work, so a user who never configured
Hermes for Vertex would see it in the explicit-only picker and could spend
against those credentials — weakening the explicit-configuration guarantee the
gate documents (mirrors the existing _IMPLICIT_ENV_VARS carve-out).

Add has_explicit_vertex_config() in agent/vertex_adapter.py that checks only
Hermes-scoped signals — VERTEX_PROJECT_ID / vertex.project_id (project
override) or a resolvable VERTEX_CREDENTIALS_PATH — and NOT
GOOGLE_APPLICATION_CREDENTIALS. Route the auth gate through it.

Adds a regression test asserting an ambient GOOGLE_APPLICATION_CREDENTIALS
path alone does not mark Vertex explicit, and updates the existing test to
drive the real config signal instead of mocking has_vertex_credentials().
2026-08-22 23:25:43 -07:00
Ajad van Wyk 3503c06d80 fix: surface Bedrock in explicit-only model pickers when AWS env credentials are set
is_provider_explicitly_configured() only checked provider env vars for
auth_type="api_key" providers. Bedrock is registered with
auth_type="aws_sdk" and an empty api_key_env_vars tuple, so a user who
sets AWS_BEARER_TOKEN_BEDROCK (or an AWS_ACCESS_KEY_ID +
AWS_SECRET_ACCESS_KEY pair) in .env was never counted as having
explicitly configured the provider.

Symptom: the desktop model picker (and any consumer of
build_models_payload(explicit_only=True)) silently hid the Bedrock row
even though list_authenticated_providers had discovered credentials and
built a full model list for it. Reproduced on main:

    build_models_payload(ctx, explicit_only=False)
      -> ['moa', 'nous', 'bedrock', ...]        # row exists, 132 models
    build_models_payload(ctx, explicit_only=True)
      -> ['nous', ...]                          # bedrock filtered out

Fix: aws_sdk-type providers now count as explicitly configured when
Bedrock-relevant env credentials are present. Deliberately env-var-only:
ambient sources (AWS_PROFILE / SSO, EC2 IMDS, container credentials)
still do NOT auto-surface, consistent with the gate's purpose (#56974)
and with a lone AWS_ACCESS_KEY_ID (no secret) not counting.

Tests: six behavior-contract cases in test_auth_provider_gate.py
covering bearer token, key pair, lone key id, ambient AWS_PROFILE,
no-credential baseline, and non-leakage into other providers.
2026-08-22 23:25:43 -07:00
Kevin Yin 9ce46a0968 fix(models): bind Anthropic pool key to its endpoint 2026-08-22 23:25:29 -07:00
Kevin Yin 8ede2e1472 fix(models): discover Anthropic pool API keys 2026-08-22 23:25:29 -07:00
Teknium 4f0e466e5d fix: accept pool-only Anthropic OAuth entries in the desktop picker filter
The salvaged carve-out covered the PKCE file and Claude Code credentials
but missed the canonical wired-token location — auth.json
credential_pool.anthropic oauth entries. Discovery accepts those via
pool.has_credentials(), so the filter must too. Read-only dict access;
api_key pool entries intentionally stay excluded (E2E case 4).
2026-08-22 22:00:19 -07:00
Viktor 4ec57d56a9 fix(inventory): keep Anthropic OAuth logins visible in desktop pickers
The desktop explicit-only picker filter drops any provider row where
is_provider_explicitly_configured() is False. That gate only recognizes
active_provider, model.provider, and API-key env vars — so a user who
authenticated Anthropic via OAuth (Hermes device flow or Claude Code
~/.claude/.credentials.json) had the Anthropic row silently hidden from
the desktop model picker even though list_authenticated_providers()
accepted those same credentials when building it (model_switch.py has an
equivalent special-case in its discovery path).

Unlike ambient CLI tokens (gh -> copilot), an OAuth access token only
exists after an interactive login, so its presence is deliberate user
configuration. Add a narrow carve-out for the anthropic slug that keeps
the row when either OAuth reader finds an accessToken; the strict gate
still governs every other provider, and is_provider_explicitly_configured
itself stays untouched so PR #4210's consent guard for auxiliary tasks
keeps its current behavior.
2026-08-22 22:00:19 -07:00
Teknium 8b86097a62 fix(gateway): honor env-configured local backends and persist Tool Gateway declines
Follow-ups on top of the salvaged #92665 commits, closing the two gaps
called out in #92647:

- SEARXNG_URL / CAMOFOX_URL now count as direct web/browser configuration
  in _get_gateway_direct_credentials(), so env-only keyless local setups
  are offered unchecked instead of pre-checked.
- Submitting the checklist with unconfigured tools left unchecked records
  them in tool_gateway_declined_tools (new known root key) and stops
  pre-checking them on subsequent Nous model swaps; opting in later clears
  the decline. Cancel (Ctrl-C/ESC) records nothing.
- has_direct labels mention SearXNG/Camofox when that's what's detected.
- Docs: tool-gateway.md gains an enablement-checklist section.
2026-08-22 21:36:55 -07:00
chelsealong ce9ddd35e7 fix(gateway): fix 3-tuple/4-tuple arity crash in get_gateway_eligible_tools
The early "fail closed" returns (account fetch error, not entitled,
non-nous provider) still returned 3-element tuples after the function's
happy path and every caller moved to a 4-tuple, crashing
prompt_enable_tool_gateway with ValueError for any logged-in Nous
account that isn't paid or pool-entitled — the common case, hit
unconditionally from `hermes model`.
2026-08-22 21:36:55 -07:00
chelsealong 4edb24276f fix(tools): don't pre-check keyless local backends in Tool Gateway checklist
get_gateway_eligible_tools() classified a tool as "unconfigured" whenever
it found no direct API-key credential, ignoring an explicit non-nous
selection already stored in config (e.g. web.backend: searxng, a keyless
self-hosted backend). Every unconfigured tool is pre-checked in the
`hermes model` Tool Gateway checklist, so a single Enter during Nous model
setup silently rewrote web.backend (and similarly browser.cloud_provider,
tts/stt/image_gen.provider) to "nous" for users who had deliberately
configured a keyless local backend.

get_gateway_eligible_tools() now also resolves each tool's stored
selection via the existing _selected_provider() helper and routes an
explicit non-nous selection into a new explicit_configured bucket instead
of unconfigured. prompt_enable_tool_gateway() never offers those tools,
so they can no longer be pre-checked or accidentally overwritten.

Fixes #92647
2026-08-22 21:36:55 -07:00