Commit Graph

4753 Commits

Author SHA1 Message Date
Teknium fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00
Teknium ce996d4057 feat(delegation): raise max_concurrent_children default 3 -> 10 (+migration) (#86745)
delegation.max_concurrent_children caps how many delegated children run in
parallel per batch (and concurrent background delegation units). The old default
of 3 needlessly serialized independent fan-outs (e.g. reviewing/​investigating N
PRs or issues at once), so large batches ran in slow chunks of 3.

Raise the shipped default to 10, which sits at/below the existing high-cost
advisory threshold (>10), so the default never trips the warning. Each child
still consumes API tokens independently, so this is a throughput/latency win the
user pays for in parallel token spend — the floor stays 1 and there is no
ceiling, so anyone can tune it down or up.

- config_defaults.py: default 3 -> 10; _config_version 36 -> 37.
- delegate_tool.py: _DEFAULT_MAX_CONCURRENT_CHILDREN 3 -> 10 (+ docstring).
- config_migrations.py: _migrate_to_37 lifts configs pinned at exactly the old
  default 3 to 10 (deliberate non-3 overrides preserved; unset inherits 10).
- cli-config.yaml.example: documented default updated.

Verified: default/fallback read 10, version 37, and the migration lifts 3->10,
preserves an explicit 5, and leaves unset untouched.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:20:32 -07:00
EvanProgramming 30c469b153 fix(gateway): spare pidfile-less Scheduled-Task gateways from the orphan reaper on Windows (#83683)
On Windows _get_service_pids() is empty (no systemd/launchd query), so a
Scheduled-Task-supervised gateway whose gateway.pid record is missing or
stale is invisible to both the service-PID and recorded-PID exclusions the
reaper already applies (#86658) — and gets SIGTERM'd on every desktop open
(#86098 class, pidfile-less path).

Add a Windows-only backstop: any reaper candidate whose parent chain
reaches services.exe (the Task Scheduler launches tasks under the services
tree) is spared even with no pidfile.

The backstop is deliberately inert on POSIX: every process there has PID 1
(launchd/init/systemd) in its ancestry — and a genuine orphan is reparented
directly to PID 1 — so supervisor-name ancestry carries zero supervision
signal and would disable the reaper entirely on macOS/WSL (#51325, #75936).
POSIX supervised gateways are already covered pidfile-independently by the
_get_service_pids() exclusion.

Known limitation (fail-open, documented): if the Task-launched bootstrap
parent has already exited, Windows does not reparent the gateway, the chain
breaks before services.exe, and the gateway is treated as an orphan.

Salvaged from #86702 by @EvanProgramming (authorship preserved); reduced to
the genuinely-new Windows backstop — the PR's other two hunks were already
merged on main via #86658 (one in a strictly stronger full-parent-chain
form) and its POSIX ancestry checks were dropped as unsound (verified
empirically: a true double-fork orphan's psutil parent IS launchd).
2026-08-15 12:04:29 +05:30
Teknium 471c687c2b test(managed_uv): cover explicit-patch fallback on the next minor line; dedupe retried versions
Follow-up to the salvaged #76252 addressing both review gaps:

- New TestMinorLineFallForward class with a direct test of the
  explicit-patch fallback branch: bare '3.12' resolves to a VULNERABLE
  build while an explicit 3.12.x patch is fixed, so recovery must go
  through _list_available_patches on the next minor line. Asserts the
  exact `uv python install` request sequence.
- New all-minors-exhausted test: everything vulnerable on 3.11-3.13
  returns None with per-line attempts bounded by _MAX_PATCH_RETRIES and
  no requests beyond 3.13 (requires-python is <3.14).
- test_retry_is_bounded_by_max_retries_constant now actually uses its
  counting wrapper and asserts the collected install calls (the
  previous version collected them into a dead variable).

Also dedupes the fallback loop the same way the same-minor loop does:
_attempt_install_generation can now record the probed candidate version
into a caller-supplied tried_versions set, so the explicit-patch pass
skips the version the bare-minor request already resolved to and
rejected -- previously that wasted a full download+install+probe+delete
cycle per minor line re-trying a known-vulnerable build.
2026-08-14 22:37:46 -07:00
RelaxJonh 2bccd6ad08 fix(managed_uv): fall forward to next Python minor when current line has no fixed SQLite build
When every patch on the current minor line (e.g. 3.11) still links a
vulnerable SQLite (e.g. 3.50.4 on Windows), the provisioner now tries
the next supported minor line (3.12, then 3.13) before giving up.

Previously, _install_safe_python_generation only tried patches within
the same minor line. On Windows, where python-build-standalone may not
publish a fixed build for the installed patch, users were stuck with a
repeated warning on every `hermes update` with no path forward.

The requires-python constraint (>=3.11,<3.14) and the downstream
import smoke test already gate compatibility, so the minor-line
upgrade is safe.

Adds allow_minor_upgrade parameter to _attempt_install_generation to
relax the same-minor-line version guard when called from the fallback
path.

Fixes #76106
2026-08-14 22:37:46 -07:00
Tachi d1df111ccd fix(update): restore Hermes Tools dependencies 2026-08-14 22:33:44 -07:00
Tachi 979a20052f fix(update): preserve activated extras across runtime rebuilds 2026-08-14 22:33:44 -07:00
konsisumer 4aa9f738ce fix(update): rebuild Desktop after release artifact loss 2026-08-14 22:27:37 -07:00
joaomarcos fa72a1edf8 fix(kanban): re-create the schema when a cached DB path loses it (#83445)
`connect()` caches every path it has initialized in the process-local
`_INITIALIZED_PATHS` set and then skips all first-open work for it —
header validation, integrity probe, `SCHEMA_SQL`, additive migrations.
That cache is keyed on a path, but the schema it stands for lives in a
file, and the two can drift apart: delete or replace `kanban.db` under a
live gateway/dispatcher/dashboard process and the next `connect()` takes
the fast path, lets SQLite create a fresh empty database, and hands back
a connection with no tables in it.

Nothing notices. Every query then fails with `no such table: tasks`,
`plugin_api._conn()` logs its init warning and carries on, and the board
renders empty. Because the cache entry survives, the process re-creates
the same schema-less ~4 KB file on every restart of the desktop app in
front of it — only killing the backing process clears it.

Verify the sentinel table on the fast path and self-heal when it is gone:
drop the stale cache entry and fall through to the existing init path,
which re-runs the probes and the schema script under the cross-process
init lock. The check is one `sqlite_master` lookup on the already-resident
page 1, so the steady-state path stays lock-free (#36644) and does no
schema work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:25:23 -07:00
Teknium 4b7b2b0049 fix: widen base-URL hostname identity class to remaining substring sites
Follow-up to #85737, which migrated five provider-identity sites onto
utils.base_url_host_matches()/base_url_hostname(). This completes the class
sweep (never-patch-predicates: one owner, every site) and folds in the two
open contributor PRs attacking individual sites:

- agent/auxiliary_client.py ZAI/Kimi OpenAI-wire rewrite (PR #85715,
  pierrenode): 'bigmodel'/'api.z.ai'/'api.kimi.com' substring checks
  rewrote proxy paths containing those markers.
- hermes_cli/runtime_provider.py Azure endpoint detection (PR #74721,
  RelaxJonh, issue #74312): 'azure.com' substring picked the Azure key
  for non-Azure hosts whose path contained the text.
- run_agent.py: _is_azure_openai_url, _is_copilot_url, Anthropic
  credential-refresh azure guard, _anthropic_preserve_dots host
  allowlist, OpenRouter/mistral reasoning gates.
- agent/chat_completion_helpers.py: nousresearch / nvidia detection.
- agent/conversation_loop.py: GitHub Models 413 hint.
- agent/usage_pricing.py: localhost billing-route detection.
- hermes_cli/model_switch.py: api.openai.com catalog fallback and
  localhost custom-provider detection.
- cli.py: local-model autodetect and Ollama/LM Studio context-length
  hints (port-anchored instead of '11434' in URL).
- tools/mcp_oauth.py: Figma remote-MCP detection.
- tools/skills_hub.py: raw.githubusercontent.com source-URL check.

Regression tests extend tests/hermes_cli/test_base_url_host_identity.py
(azure/copilot/dotted-model/figma proxy-path + lookalike cases) and
tests/agent/test_minimax_auxiliary_url.py (ZAI/Kimi path false positives).

Closes #74312. Salvages #85715 and #74721 with authorship preserved.
2026-08-14 22:04:16 -07:00
RelaxJonh 198e2f2746 fix(routing): use hostname match for azure.com endpoint detection (#74312)
Replace raw substring checks ("azure.com" in full_url) with the existing
base_url_host_matches() helper at two sites in runtime_provider.py.

The substring approach misclassified URLs whose path (not hostname)
contained "azure.com" — e.g. https://example.invalid/proxy/azure.com/v1 —
causing the wrong credential (Azure key instead of explicit Anthropic token)
to be selected, and potentially leaking a more-privileged Azure key across
a trust boundary.

base_url_host_matches() parses the URL and validates only the hostname
against allowed Azure suffixes with proper boundary rules.

Fixes #74312
2026-08-14 22:04:16 -07:00
Teknium 42a1db4c64 fix(update): use canonical venv_bin_dir in _install_repair (no open-coded Scripts/bin) 2026-08-14 22:03:56 -07:00
Halldrix 19cff89300 fix(update): complete pending core install before any native import (self-lock loop fix)
Reviewer egilewski found the original defer was circular (#83590 comment):
the self-lock preflight wrote .update-incomplete and exited, but the next
launch only ran the full recovery AFTER main.py's third-party imports —
so a healthy venv's probes made the early pass a no-op, main.py imported
cryptography eagerly, the .pyd got mapped again, and the deferred install
re-hit the exact self-lock it was meant to escape.

Close the loop by making the marker guarantee the install runs BEFORE any
native extension can be imported:

- hermes_cli/_install_repair.py (new, stdlib-only): single source of truth
  for the core .[all] reinstall — ensurepip bootstrap, uv-pip/pip
  resolution with VIRTUAL_ENV, Termux env stripping, Windows hermes*.exe
  quarantine, per-extra fallback ladder, and fd1→fd2 routing for acp
  safety.  Deliberately free of managed_uv/hermes_constants imports so it
  stays importable in the corrupted-venv state it exists to repair.
- hermes_cli/_early_recovery.py: recover_if_needed now completes a pending
  .update-incomplete install BEFORE the import probes, on every launch
  that sees the marker (unless argv is update).  Success clears the
  marker; failure bumps an attempts counter inside the marker body and
  keeps it.  A 3-attempt ceiling stops a persistently-failing install
  from reinstall-hammering every launch (hermes acp included) — past the
  ceiling the late post-import recovery takes over with its manual
  recovery instructions.  Single-flight lock shared with the late path.
- hermes_cli/main.py: _recover_core_update_marker_locked delegates the
  install to the shared executor (no duplicated logic); ensure_uv stays
  in the late path so a venv whose uv vanished mid-update still
  bootstraps it.
- tests: 7 new regressions — the reviewer's exact case (marker + healthy
  venv → install runs while sys.modules has no cryptography), failure
  keeps marker + increments attempts, retry ceiling, lazy marker does
  not trigger core install (#58004 invariant), argv-update skip, and
  corrupt/missing marker bodies.  The key test was sabotage-verified:
  removing the pre-import branch makes it fail with zero install calls,
  while a lone-lazy-marker test still passes; restoring the branch makes
  it pass again.

Refs #83569
2026-08-14 22:03:56 -07:00
Halldrix c6a71294b6 fix(update): detect updater self-lock on Windows + repair venvs whose base interpreter is uv-managed
Two gaps left every Windows git-checkout install unable to recover from
the exact failure state #83569 reports:

1. Self-lock detection. _detect_venv_python_processes() always excludes
   the calling process by design — a CLI hermes update IS the venv python.
   An updater that had already imported a native venv extension (the
   canonical one being cryptography.hazmat.bindings._rust, mapped while
   hermes_cli.main resolved external secret sources) passed every
   preflight and then died mid-sync with os error 5 when uv tried to
   rewrite the mapped .pyd, stranding the venv half-updated. A new
   preflight now refuses the sync before touching the checkout, writes
   the update-incomplete marker so the next fresh launch completes the
   install, and exits 2. Verified on a live Windows 11 host: after
   importing hermes_cli.main, tasklist /m _rust.pyd shows the .pyd mapped
   in the caller, and a peer process cannot open it read-write
   (Permission denied) — while a rename succeeds, matching how uv/pip
   actually fail (truncate+write, not rename).

2. Early-recovery install path. _early_recovery._run_repair_install used
   sys.executable -m pip unconditionally. Windows git checkouts install
   on a uv-managed base interpreter (python-build-standalone), whose
   EXTERNALLY-MANAGED marker makes plain pip abort with
   externally-managed-environment — the repair no-oped and the venv
   stayed broken. The repair now detects the PEP 668 marker, prefers
   uv pip install with VIRTUAL_ENV pointed at the project venv, and
   falls back to pip --break-system-packages when no uv binary exists.

Both fixes ship with subprocess/unit regressions (sabotage-verified):
the new tests fail on pre-fix code and pass with it. Complements #77517,
which keeps the updater from importing cryptography in the first place;
this PR is the defence-in-depth when any future path loads it anyway.

Fixes #83569
2026-08-14 22:03:56 -07:00
chelsealong 49d72a02f6 fix(update): verify Windows gateway cold-start survives before reporting success
_cold_start_windows_gateway_after_update() printed the success line off a
successful Popen return alone, which only proves CreateProcess succeeded,
not that the child survived. On Windows, a job object denying
CREATE_BREAKAWAY_FROM_JOB hard-kills the child during updater teardown
before it logs anything, yet the updater still printed "Starting Windows
gateway after update (PID ...)" — leaving Telegram/Discord/etc. offline
with no indication anything failed (#84185).

Route the success report through gateway_windows._report_gateway_start(),
the same post-spawn liveness poll every other _spawn_detached() caller
already uses, so a dead child is reported as a failure with a
manual-recovery hint instead of a false success.
2026-08-14 22:03:56 -07:00
PRATHAMESH75 517151ee4a fix(install): fail closed when a stopped gateway leaves an empty survivor probe
Review follow-up (#78574): the aborted-restart handler only flagged the fleet
stale when the post-failure survivor probe was None or non-empty. A positive
empty probe was treated as proof-of-safety — but `[]` is only safe when
nothing was running before the phase. If a gateway was discovered, stopped
(SIGTERM/drain), and its replacement never came back, the probe is empty at
exactly that unsafe moment and the update reported success — the fail-open
contract this fix exists to close.

Snapshot the pre-restart gateway PIDs before any stop/drain and route the
handler decision through a pure _restart_phase_failure_is_incomplete() helper
that fails closed on an empty survivor set whenever a gateway existed
pre-restart (or the pre-state could not be read). Add decision-level regression
tests covering the stopped-without-replacement gap, unknown pre-state, and the
truly-no-gateway positive control.
2026-08-14 22:03:56 -07:00
PRATHAMESH75 95018b6bba fix(install): surface aborted gateway restart during hermes update
The gateway auto-restart phase in `hermes update` was wrapped in a blanket
`except Exception` that only logged at debug level. When the phase raised
early — e.g. importing `hermes_cli.gateway` from the freshly pulled checkout
inside a process that already loaded pre-update modules — every drain and
restart line vanished from the update output, the update printed
"Update complete!" and exited 0, and the still-running gateway kept serving
pre-update modules against replaced source files. The next Telegram turn died
with `ImportError: cannot import name 'is_trivial_prompt'`.

The handler now probes for surviving gateway processes and, unless it can
positively prove none are running, prints the cause plus a manual recovery
command and marks the fleet restart incomplete — which exits nonzero and
writes the gateway-mode exit-code marker, matching the existing
failed-or-stale-unit path.

Fixes #78574
2026-08-14 22:03:56 -07:00
Soheil Fakour bdfdd4392f fix(update): gate 'Code updated!' on HEAD actually moving (#79678)
A detached/pinned checkout can report 'N new commit(s)' against origin,
run the ff-only merge successfully, and still sit on the old commit
afterward (the branch-switch step re-detaches to the raw SHA). Before
this guard 'hermes update' printed '✓ Code updated!' and reinstalled
deps + rebuilt the desktop app against the stale tree - no error, no
warning, 'hermes doctor' healthy.

Compare pre-pull and post-pull HEAD; if they match, fail loudly with a
reattach hint instead of claiming success.
2026-08-14 22:03:56 -07:00
kshitij b58fa89cd7 refactor: centralize dict-valued model.default coercion via shared helper
Promote _split_model_config_default to hermes_cli/config.py as the single
shared helper for flattening dict-valued model.default/model.model config.
All 8 defense-in-depth sites now route through it instead of inlining
their own isinstance checks with inconsistent key orders.

Changes:
- Add split_model_config_default() to hermes_cli/config.py (public)
- cli.py: _split_model_config_default delegates to shared helper
- Fix key extraction order: agent_runtime_helpers.py was reversed
  (default->model); now consistent (model->default) across all sites
- Remove provider-as-model-name fallback from main.py, oneshot.py,
  model_tools.py, cli.py — provider is a routing key, not a model ID
- Add 'name' to _normalize_root_model_keys flattening loop and
  _has_nested_default detection to cover the deprecated model.name alias

Tests: 118 passed + 1 skipped (cli_init, managed_scope, config).
E2E: 31/31 passed (config chokepoint, managed scope, crash site,
negative cases, edge cases).
2026-08-15 10:29:05 +05:30
mariobgsp be708ff1b9 fix: normalize managed config overlay before merge in load_config
The shared load-boundary flatten added for dict-valued model.default only
ran on the user/default merge; _load_config_impl then deep-merged the raw
managed overlay without normalizing, so a managed model.default:
{provider, model} still reached status/fallback/runtime readers as a dict.

Normalize the managed overlay (same _normalize_root_model_keys pass, plus
the bare model-string -> model.default promotion used by
managed_scope.apply_managed_overlay) before expanding and merging, so
every overlay is canonical before load_config returns.

Adds load_config() regressions for a nested managed default and a bare
managed model string.
2026-08-15 10:29:05 +05:30
mariobgsp 998329a621 fix: flatten dict-valued model.default at config load boundary
Extends the fix to the config-load chokepoint so every reader sees plain
strings, not just the interactive CLI paths. _normalize_root_model_keys
now flattens a dict-valued model.default/model.model into a string default
plus the nested provider (promoted to model.provider when no explicit
outer provider or "auto" is set), covering the residual readers the
review flagged: doctor, status/dump, fallback picker, prompt-size, and
the context-switch guard — all of which called .strip()/flowed the raw
value and would crash or misroute on a nested dict.

Adds _normalize_root_model_keys regression coverage for the flatten
(precedence, auto-override, explicit-provider-win, alias shape, flat
strings untouched).
2026-08-15 10:29:05 +05:30
Ario Bagus Prakusa cb4daf23f3 fix: coerce dict-valued model/default config back to string across resolution paths
A dict-valued model.default (e.g. {provider:..., model:...}) in config.yaml
was leaking into agent.model and crashing the agent at init:

  AttributeError: 'dict' object has no attribute 'lower'
    agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy

This manifested on the Telegram gateway as an infinite reset loop: every
turn built an agent with model=dict, crashed during init, the gateway
treated the failed turn as a session needing reset, and /reset rebuilt the
agent and crashed again.

Coerce dict -> string at every model-resolution entry point so the value
is normalized once and never reaches a .lower() call as a dict:
- agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy (the crash site)
- agent/agent_init.py: configured default model resolution
- cli.py: CLI config model + _normalize_model_for_provider
- hermes_cli/main.py: _has_any_provider_configured
- hermes_cli/oneshot.py: _run_agent model resolution
- hermes_cli/runtime_provider.py: _get_model_config default handling
- model_tools.py: _resolve_active_context_length
2026-08-15 10:29:05 +05:30
HexLab98 18b442cdeb fix(install): abort Windows venv recreate when rename-aside fails
When Rename-Item on the live venv is denied, do not fall back to an
in-place Remove-Item that can gut site-packages and leave no rollback.
Also mark venv-blocker probe failures with probe_failed so they cannot
be read as a clear scan (#83149).
2026-08-14 21:58:09 -07:00
tachyon-r 7a6b8917f7 fix(tools): recognize discovered plugin platforms 2026-08-14 21:56:33 -07:00
Chen Jin 7224301856 fix(toolsets): admit explicitly-configured plugin toolset keys in _get_platform_tools (#81163)
Layer 2 of the #81163 / #78050 fix: _get_platform_tools computed
plugin_ts_keys = _get_plugin_toolset_keys() but only used
CONFIGURABLE_TOOLSETS in the explicit-config filter, so a user-listed
plugin key like `a2a` in `platform_toolsets.cli: [hermes-cli, a2a]` was
silently dropped. The filter now unions configurable and plugin toolset
keys when evaluating has_explicit_config and when admitting per-key
entries.

Cherry-picked from PR #81190 (Layer 2 hunks only; Layer 1 is covered by
the provides_tools mechanism from PR #78842).
2026-08-14 21:56:33 -07:00
Eman e42db348c9 fix(plugins): register deferred platform client tools at discovery (#78050)
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.

Re-anchored accordingly:

- Discovery-time pre-registration, module reuse, and the `provides_tools`
  opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
  since those tools registered before `registration_start` and the slice
  cannot see them.
- A failed materialization no longer carries attribution across. The
  failure path now sweeps the whole ownership ledger for the plugin key,
  not just the `registration_start:` slice, so the pre-registered tools
  are disposed along with the adapter. Attribution and the registry now
  agree at zero instead of reporting tools the process is not serving.

tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 21:56:33 -07:00
webtecnica 2d81236f7f fix(cli): report background-dispatch cron runs without false 'failed' (#83340) 2026-08-14 21:55:14 -07:00
Jack Lau 6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00
briandevans 66356fe24b fix(backup): drop setuid/setgid from the mode restored onto imported files
``_extract_member_atomically`` carries the replaced file's permissions
across the publish so that routing through mkstemp does not change what
the caller would otherwise have produced. But ``_preserve_file_mode``
returns ``stat.S_IMODE``, which is all twelve bits, and this restore is
deliberate on both sides of the replace: the mode is fchmod'd onto the
temp before ``atomic_replace`` and re-applied afterwards because chown
clears the elevated bits. So a target sitting at 0o4755 comes out of
``hermes import`` still at 0o4755 — with contents supplied by the zip.

That is a regression introduced by the atomic rewrite rather than a
pre-existing one. The overwrite it replaced was an in-place
``open(target, "wb")``, and an in-place write by a process without
CAP_FSETID has the elevated bits stripped by the kernel, so the old path
left 0o4755 as 0o755.

The blast radius is not limited to Hermes' own state: the ``_external/``
branch of ``run_import`` publishes members anywhere under ``$HOME``, and
this is the path that documents ``sudo`` use so ownership survives a
restore. An archive that happens to contain a member matching some
existing privileged file would take over the identity that file runs as.

Mask the two bits off the preserved mode. The masking happens once,
before the temp file is chmod'd, so there is no transient elevation
either. The sticky bit is kept — it is inert on a regular file. The
ordinary permission bits are unaffected, so the Docker/NAS installs the
preservation exists for still get their broader modes back.

This is the one write path in the repo where the bytes are untrusted;
the ``utils`` writers that preserve the full mode re-serialize content
the process itself produced, and are correct as they stand.
2026-08-14 21:47:23 -07:00
briandevans 60f86662e3 docs(backup): note that atomic_replace's cross-device fallback still truncates
The atomicity claim in _extract_member_atomically's docstring holds on the
os.replace path but not on atomic_replace's EXDEV/EBUSY fallback, which uses
shutil.copyfile and so opens the destination 'wb'. That is pre-existing
behaviour shared by every atomic writer in the repo, and it is reachable here
for a symlinked target whose real file lives on another filesystem. Scope the
docstring to what the helper actually guarantees instead of overstating it;
the fallback itself is a utils.atomic_replace change.
2026-08-14 21:47:23 -07:00
briandevans 1c3c1f4d71 fix(backup): preserve owner on atomic import writes and close the 0600 transit window
Follow-up on the atomic-import restore, delegating both metadata concerns to
the shared helpers instead of half-handling them locally.

Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace`
publishes a temp file owned by the *writing* user, so `sudo hermes import`
re-owned every restored file to root — on the disaster-recovery path, and on
exactly the Docker/NAS volume installs `utils._restore_file_owner` was added
for. `_extract_member_atomically` now captures `_preserve_file_owner(target)`
before staging and calls `_restore_file_owner` after the replace, before the
mode restore (chown clears setuid/setgid, so the mode has to go back last).

Mode handling was also only half applied before the replace: the `os.fchmod`
branch applied it to the temp fd, but the platforms without `fchmod` fell
through to a best-effort post-replace chmod, leaving the published file at
mkstemp's 0600 until that chmod landed — permanently if the process died in
between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback
copy 0600 onto the target. The mode is now applied to the temp file on both
branches, with the post-replace `_restore_file_mode` kept as the belt-and-
braces path.

This is the same shape `atomic_write_text` and `atomic_yaml_write` already
carry after 3556728a5 and 43fc86562; capture and restore now reuse
`utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` /
`_restore_file_owner` rather than re-deriving them, which also drops the local
`import stat`.

Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites):
- test_restore_preserves_existing_file_owner — forces a uid/gid so it does not
  need root; asserts chown fires once, with the captured owner, on the
  pre-existing file only (a newly created member has no prior owner).
  Mutation-checked: dropping only the `_restore_file_owner` call reds it.
- test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr`
  on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600
  without the fix, 0o644 with it. Mutation-checked the same way.
2026-08-14 21:47:23 -07:00
briandevans e88c9f0ef2 fix(backup): restore import members atomically so a failed import can't erase config
`hermes import` wrote every zip member with `open(target, "wb")` followed by
`dst.write(src.read())`, at both restore sites in `run_import`. Opening for
write truncates the user's existing file to zero *before* any replacement
bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves
`config.yaml`, `.env`, or an external provider config (e.g.
`~/.honcho/config.json`) empty with nothing behind it — during the
disaster-recovery path the user is running precisely because they already
lost something. The `_external/` branch writes outside HERMES_HOME, into
third-party configs under the user's home, so the blast radius is not
confined to Hermes state.

Both sites now stage the member into the target's own directory, fsync it,
and publish with `utils.atomic_replace`, so the target only ever moves from
its old contents to the complete new contents.

`atomic_replace` rather than a bare `os.replace`: it resolves a symlinked
target first, so deployments that link `config.yaml` into a dotfiles repo
keep the link instead of having it silently swapped for a regular file
(#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for
cross-device and bind-mount installs. Members stream through
`shutil.copyfileobj` instead of being read whole into memory. The temp file
is removed on any failure so a partial import leaves no residue, and
permission bits are carried across the replace so mkstemp's 0600 does not
silently tighten restored files.

This extends the module's own established idiom — `backup.py` already
publishes atomically via `os.replace` in `_atomic_output_path` and in the
snapshot writer — into the one path that still overwrote user files in place.
2026-08-14 21:47:23 -07:00
joaomarcos 45bb486b26 fix(gateway): give in-flight cron work its own drain floor
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.

A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.

Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.

The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.

Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.

Relates to #82161 (complements #82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
2026-08-14 21:47:16 -07:00
Jeremy 1f6f86119f fix(cli): stop hermes update from respawning orphan serve --port 0 (#78821)
Filter manual dashboard/serve respawn candidates after update: skip
ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines,
and cap one restart per profile/HERMES_HOME so orphan counts no longer
grow across successive updates.
2026-08-14 21:46:32 -07:00
Teknium 0dba3316b2 fix(gateway): generalize supervised-gateway exemption in orphan reaper to all platforms
Compose the service-PID exclusion (#85743, RelaxJonh) and the recorded-PID +
parent-chain exemption (#86100, arccat-114) into one cross-platform rule:

- _get_service_pids() exclusion now runs unconditionally, not only under
  is_macos() — it is the authoritative "supervised" signal for launchd and
  any systemd unit visible on a host that got past the systemd gate.
- The recorded-healthy-gateway (get_running_pid()) + parent-chain exemption
  now runs on every platform, not only Windows. A recorded, liveness-verified
  gateway is by definition not an orphan "the pidfile/runtime record can't
  see", so the reaper must never target it — this covers Windows Scheduled
  Task / Startup VBS supervision, standalone launcher-started gateways
  (the case #85743 alone would miss), and macOS/WSL equivalents.

True orphans (no service registration, no valid runtime record) are still
found and reaped, preserving the #51325/#75936 duplicate-port protection.

Existing macOS regression tests updated to pin get_running_pid to None for
their scenario; Windows regression tests from #86100 carry over unchanged.

Bug class: #83683 (root), #86287, #86098, #85738, #85368, #85344, #85044,
#84855, #84824, #84200.
2026-08-14 21:44:28 -07:00
arccat-114 102369c5f6 fix(gateway): spare Scheduled-Task-supervised gateway from orphan reaper on Windows
The orphan reaper kills a healthy gateway (and its Scheduled-Task bootstrap
parent chain) every time the Desktop backend starts on Windows, because
_get_service_pids() only implements systemd/launchd and returns an empty
set on Windows — a supervised gateway is therefore indistinguishable from
an unsupervised orphan.

Exempt the recorded healthy gateway PID and its parent chain from the
orphan scan on Windows, mirroring the macOS launchd exemption (#85913).
The Scheduled-Task bootstrap's argv matches the gateway scan, so without
exempting the parent chain killing the bootstrap takes the detached
gateway down with it.

Fixes #86098
2026-08-14 21:44:28 -07:00
RelaxJonh ac9b058ef4 fix(gateway): exclude service-managed PIDs from orphan reaping
_reap_unsupervised_gateway_orphans() kills every gateway PID found by
find_gateway_pids() on hosts without systemd (macOS launchd, Windows
Scheduled Task). This includes service-managed gateways that are NOT
orphans — they are supervised by launchd/systemd and should never be
killed during a stale-process sweep.

Add own |= _get_service_pids() to the exclusion set before scanning,
so launchd/systemd-supervised gateways are preserved. True orphans
(reparented leftovers not present in launchctl/systemctl) are still
found and reaped, preserving the #77276 protection.

Fixes #85344 (macOS launchd gateway killed by desktop serve startup)
Fixes #85044 (Windows Scheduled Task gateway killed by desktop serve)
Fixes #84855 (Permission denied to kill orphaned gateway PID)
Fixes #85368 (gateway process repeatedly killed, messaging offline)
2026-08-14 21:44:28 -07:00
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
RelaxJonh 0d91ab8889 fix: close leaked SessionDB connections on /insights and sessions-repair exception paths (#83226)
Two call sites create SessionDB instances without closing them on error:

1. gateway/slash_commands.py: /insights command - db.close() was on the
   success path but not in a finally block, so exceptions between
   SessionDB() and db.close() leak the connection.

2. hermes_cli/sessions_cmd.py: sessions repair - SessionDB() created
   inline with no .close() at all, leaking the FD on every call.

Salvage note: the original PR (#83237) also added a __del__ safety-net
finalizer to SessionDB; review showed the atexit hook registered by
queue_token_counts() strongly retains the instance, so the finalizer never
fires for the leak class it claimed to cover. Dropped here in favor of the
deterministic constructor-finally ownership repair salvaged from #83620.
2026-08-14 21:41:26 -07:00
Teknium f45813ea77 fix(sessions): run state.db schema migration eagerly at backend startup and stop swallowing locked ALTERs
After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).

Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):

1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
   writable open, typically the user's first NEW session. The dashboard
   backend now schedules one writable open of its own state.db from the
   lifespan (daemon thread, never blocks the ready-probe socket, never
   raises), so the store is brought current before the first session-
   list poll on every `hermes serve` / `hermes dashboard` / Desktop
   headless entrypoint.

2. _reconcile_columns caught sqlite3.OperationalError around every
   ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
   orphaned sibling backends made the ALTER fail silently — startup
   "succeeded" with a half-reconciled schema, and the open-time lock
   patience (#74478) never saw the error because it was swallowed
   inside first. Now: "duplicate column" races stay at DEBUG,
   locked/busy re-raises so _connect_and_init_with_lock_patience
   retries the whole idempotent init with jittered backoff, and any
   other failure (e.g. un-ADDable NOT NULL) logs at WARNING.

Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.

Fixes #79531
Fixes #80037

Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
2026-08-14 21:36:29 -07:00
David Metcalfe e4d0e4c3d8 fix(win): never rewrite the in-use managed Node tree (#80926)
The Hermes-managed Node tree at %HERMES_HOME%\node is destructively
rewritten while the desktop app's Node processes execute from it:
the Node-26 heal did shutil.rmtree + move, the EBADENGINE repair ran
npm install --global --prefix into the tree, and install.ps1's
Test-Node did Remove-Item + Move-Item. Windows rejects those writes
with PermissionError: [WinError 5] on npm.cmd.

- _heal_managed_node_windows: stage the fully-downloaded tree in a
  sibling node.new-* dir, then rename-swap (live tree -> node.old-*,
  staged -> node). The live tree is never deleted before its
  replacement is ready, so an interrupted heal cannot gut it; a
  refused rename is the OS-level in-use signal and defers (returns
  None) instead of forcing the write.
- heal_hermes_managed_node: an in-use deferral does not record the
  once-per-process attempt, so the heal retries once the tree is free.
- managed_node_tree_in_use: cheap psutil pre-check (Windows only) that
  avoids pointless 30-50MB re-downloads in long-lived processes.
- upgrade_managed_npm: defer the in-place npm self-upgrade while the
  tree is in use, with a notice.
- install.ps1: Test-ManagedNodeInUse guard around Update-ManagedNpm and
  the Test-Node install branch, which now rename-swaps instead of
  delete-then-move.

An in-use-but-outdated tree keeps serving the old runnable Node (old
Node beats no Node), and every npm resolution re-evaluates the heal, so
the upgrade applies automatically on the next update with the app
closed.
2026-08-14 21:35:30 -07:00
zuowen7 f71f91a39b fix(desktop): surface compaction-archived messages in transcript reads (#80680) 2026-08-14 21:08:14 -07:00
spfcraze 5f619cfa0e fix(gateway): don't restart supervised services on clean exit
The s6 finish script for profile-gateway services restarted on ANY exit
except EX_CONFIG (78) — including clean exit 0. Restart-on-normal-exit
turns an intentional stop into a reconnect loop: the ashriel-discord
storm in #76435 made 1,000+ connections and got the bot token reset by
Discord.

The finish script now exits 125 (permanent failure, no restart) for both
clean exit 0 and EX_CONFIG; only non-zero, non-78 exits (genuine crashes)
restart normally.

Scope note: #76435 bundles a second, separable symptom (Windows desktop
updater showing the literal 'managed outside dashboard' sentinel). Its
root cause is undiagnosed and #22733 covers the dialog-explanation path;
this PR is the gateway half only.

Tests: behavioral — the rendered finish script is executed via sh for
exit codes 0/78/1/137, asserting no-restart for clean stop and fatal
config, restart for crashes. Pre-existing EX_CONFIG test still passes.
2026-08-14 21:07:21 -07:00
Teknium bc5805c35f fix: compare base-URL hostnames, not substrings, in provider-identity checks
Port of the bug class from earendil-works/pi#7933 (DeepSeek base-URL
detection matched by raw substring, missing case variants and matching
lookalike URLs). Hermes had the same class at five sites:

- cli_agent_setup_mixin.py: keyless-custom-endpoint detection treated any
  URL containing the OpenRouter host substring (path segment, lookalike
  domain) as OpenRouter, and missed case variants of the real host.
- models.py validate_requested_model: same substring check for routing an
  openrouter provider with a custom base_url to the custom catalog.
- runtime_provider.py: local-endpoint autodetect matched the string
  localhost anywhere in the URL, including remote hostnames containing it.
- gateway/run.py: /status endpoint display, same local-host substring.
- agent_runtime_helpers.py: Nous Portal cache-layout detection matched
  the nousresearch substring anywhere in the URL.

All sites now use the existing base_url_host_matches / base_url_hostname
helpers (exact host or subdomain, case-insensitive). Regression tests
proven to fail against the old predicates.
2026-08-14 20:47:13 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00
Aleks Clark 62eefff697 perf(desktop): bound long-running app resource use
Persist backend ownership for reliable cleanup, park inactive panes, and
evict unreferenced transcripts so Desktop stays responsive over long sessions.

💘 Generated with Crush

Assisted-by: Crush:gpt-5.6
2026-08-14 20:23:07 -07:00
Teknium 50d98fc1f3 feat(delegation): raise subagent iteration cap default 50 -> 250 (+migration) (#86506)
delegation.max_iterations is the per-subagent tool-call budget. The old
default of 50 truncated substantial delegated work: leaf agents spend
~15-20 turns on reconnaissance before producing output, then ran out of
budget mid-task and returned 'completed but unfinished' summaries. 250
gives real delegated work room to finish.

Changes:
- config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36
- tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in
  sync with the shipped default to prevent drift)
- config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly
  the OLD default 50 -> 250 on update, so existing installs inherit the new
  headroom. Any other explicit value (deliberate override) is preserved;
  unset inherits 250 at read time.
- cli-config.yaml.example: doc the new default

The cap is per-child and children run concurrently (max_concurrent_children
default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds
(default 0 = off) remains available as a wall-clock guardrail, and users can
still pin a lower max_iterations explicitly.

Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset
untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
2026-08-14 16:50:30 -07:00
Justin Bennington 167dd48b40 fix(models): trust certifi for credentialed catalogs (E-1002) 2026-08-14 16:43:37 -07:00
SHL0MS b21e0bd8c9 fix(providers): honor per-provider TLS on custom /models and pricing probes
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:

- the requests-based metadata/pricing probe
  (agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
  (hermes_cli/models.py::probe_api_models)

A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.

This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:

- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
  ssl_ca_cert before falling back to the env vars. Callers with no base_url
  keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
  passes it through open_credentialed_url, which gains an ssl_context seam on
  the cloned secure opener. Unmatched or public endpoints pass None and keep
  urllib's default policy.

Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
2026-08-14 16:37:36 -07:00
SHL0MS cbfe186da8 fix(config): canonicalize legacy api_mode spellings instead of silently discarding them
Earlier releases accepted api_mode: openai on custom provider entries.
The canonical transport set is now {chat_completions, codex_responses,
anthropic_messages, bedrock_converse, codex_app_server}, and an
unrecognized value was silently ignored at both consumption sites
(_normalize_custom_provider_entry passes the raw string through and
agent_init's accepted-set check drops it; _parse_api_mode returns None),
falling through to hostname-based detection.

For hosts with a detection rule the provider silently switches
transports after an update. Observed live: a custom entry for
api.actual.inc with api_mode: openai (valid when written) flipped to
codex_responses via the hostname rule, and every reasoning-bearing
request to the relay's /v1/responses failed with a wrapped non-JSON
error while /v1/chat/completions worked throughout.

Fix: one shared alias map (_canonical_api_mode) consulted by both
sites. openai/openai_chat -> chat_completions, responses ->
codex_responses, anthropic/messages -> anthropic_messages, bedrock ->
bedrock_converse. Canonical names and unknown values pass through
unchanged, so invalid-config behavior is untouched.

Tests: alias map contract (every alias lands in _VALID_API_MODES),
normalizer canonicalization incl. the transport: key alias, and the
runtime gate accepting legacy spellings while still rejecting unknowns.
2026-08-14 16:36:39 -07:00