Conflict in hermes_cli/model_setup_flows.py: took simp/cli-models's
simplified version and re-applied simp/cli-main's 2-line change (import
_current_reasoning_effort/_set_reasoning_effort from hermes_cli.setup
instead of hermes_cli.main, since cli-main removed the main.py copies).
simp/cli-models had dropped those two helpers from hermes_cli.setup.py as
dead (their only base caller was main.py's own copy); restored them there
so the cli-main import path resolves.
hermes_cli/plugins.py (7193 -> 4961):
- _register_scoped_provider: one body for the 8 scope-keyed register_*_provider methods;
_track_callback/_track_mapping_entry unify manager-mapping lease tracking (4 sites);
_track_scoped_registration for the ownership ledger. Public method names, signatures,
return values and warning strings unchanged.
- Manifest v2 type checks table-driven (_manifest_field_of_type); _unload_scoped split into
_unload_target_keys + _reset_after_unload_all; _gate_manifest/_record_placeholder out of
_discover_and_load_inner; _track_tool_override_policy/_attribute_registrations out of
_load_plugin_scoped; bounded hook worker lifted into _run_hook_callback_bounded;
prompt-section rendering extracted; resolve_pre_tool_block delegates to
_dispatch_pre_tool_call_hooks; _evict_modules (3 sites); _remove_name_if_unowned and
_nowait_plugin_set unify unowned-name cleanup and *_nowait probes.
- Dead (zero references outside their own test): unload_plugins, has_portable_mcp_servers,
_classify_entrypoint_kind, _reset_event_bus, get_plugin_subscriptions,
get_telegram_handler_factories, pass-through _restore_* helpers; _env_enabled is now an alias
of utils.env_var_enabled (plugins/memory still imports it).
- Docstrings/comments compacted; contract sentences and rationale kept (multi-profile ledger
keying, persistent auth-provider registration, capability declaration is not a grant).
- Drop dead code: _telegram_effective_priority, _prioritize_telegram_menu_commands,
discord_skill_commands, _TG_NAME_LIMIT/_clamp_telegram_names compat aliases,
the empty _SLACK_PRIORITY_ALIASES pinning pass (and tests that only pinned them).
- Unify the Telegram and Discord skill collectors behind _iter_gateway_skills
(one eligibility/root-matching implementation) and _truncate_desc.
- Table-drive Telegram menu priority modes (_TELEGRAM_PRIORITY_TIERS) and the
completer's per-command dynamic completions (_DYNAMIC_COMPLETIONS).
- Unify path and @file:/@folder: directory-listing completions (_dir_completions),
command-completion construction (_short_desc), /tools candidate rows, the
Slack canonical/alias passes and the derived COMMANDS/COMMANDS_BY_CATEGORY loops.
- Compact docstrings/comments, keeping every rule, invariant and rationale.
Behavior-neutral: registry-derived outputs, menus, manifests and completions
byte-identical against origin/main on a fixture sweep.
moa, fallback, worktree, browser, secrets, egress, migrate, whatsapp-cloud,
checkpoints, bundles, curator, pets, journey, computer-use, sessions and
completion each become a build_<group>_parser() builder. Closure handlers
that only closed over their own parser moved verbatim; sessions/completion
take the handler by injection. --help/usage/defaults byte-identical for all
399 parsers in the tree (in-process dump before/after).
model_setup_flows.py (3313 -> 2848):
- _load_config_model_section, _begin/_commit_model_config, _ensure_flow_api_key,
_pick_model_or_prompt, _run_login, _models_dev_merged, _copilot_model_list,
_show_curated replace ~15 copies of config-save / api-key / picker boilerplate.
- _gemini_tier_ok and _api_key_provider_model_list lift the two inline blocks
out of _model_flow_api_key_provider; five-way provider branch -> early returns.
- Comments compacted, keeping every rationale (Bedrock geo routing, key_env
hygiene, discover_models semantics, Nous free/paid partition, etc.).
Also drops two tests that only asserted the existence of setup.py helpers
removed in the next commit.
Follow-up to the off-loop move: once the credential-pool handlers run on
worker threads, the dashboard's periodic /api/credentials/pool polls can
overlap, and during a DNS outage each poll would have started its own
exchange and abandoned its own hung resolver thread.
- Per-fingerprint threading.Lock around the exchange: concurrent callers
wait on the one in-flight attempt, then hit the positive or negative
cache (bounded worker count, no duplicate network calls).
- _urlopen_bounded: when the hard cap fires and the abandoned worker later
succeeds, close the HTTPResponse instead of leaking the socket.
- Tests (none shipped with the original PR): hard cap + late-close,
single-flight success and failure paths, and the pool endpoint running
off-loop / keeping the loop responsive under a 200 ms blocking read.
The module already binds run_in_threadpool (used by list_profiles_endpoint)
and every sibling router uses the same starlette helper; the nine new
loop.run_in_executor(None, _run) sites now go through that alias so the
file has one offload idiom. Behaviour-identical (both hand the callable to
a worker thread).
Also sweeps the one endpoint the PR left synchronous:
update_profile_model_endpoint's _write_profile_model reads and rewrites
the profile's config.yaml on the event loop.
Two assertions per offloaded site:
- a loop probe, where the stubbed callee records whether an event loop is
running in its own thread — the idiom already used by
tests/hermes_cli/test_cron_dashboard_off_loop.py; and
- a concurrency proof, where the stubbed callee blocks on a threading.Event
while an unrelated request is timed. On the unfixed handlers that request
waits out the whole block; served off the loop it returns in
milliseconds.
The concurrency proof needs a single event loop across requests, so the
client fixture enters the TestClient context manager: that pins one
blocking portal for the whole fixture, where a bare TestClient(app) would
spin up a fresh loop per request and pass even unfixed.
Also covers the status-code mapping through the executor hop (404 on a
missing profile, 400 on a rename collision, 404 from the resolve that stays
on the loop ahead of describe-auto) and the _MISSING sentinel cases: a
desktop.json holding `null` still reports exists=true, an absent one
reports exists=false, and an empty SOUL.md is still distinguishable from a
missing one.
The client fixtures read web_server._SESSION_TOKEN from the module rather
than pinning a literal. web_server resolves that token once at import, so
whichever test file imports it first fixes the value for the session and a
later monkeypatch.setenv is silently ignored — two files hardcoding
different tokens would 401 depending on collection order.
Extends the off-loop suite to the third and last blocking method on the
`UpstreamAdapter` contract, the 401/429 rotation.
As with the two existing pairs, the primary assertion is **thread identity**,
not latency: a latency assertion measured by an HTTP client on the blocked
loop is vacuous, because the client's own timer cannot advance until the
block ends and it therefore reports a fast response on provably frozen code.
* `test_get_retry_credential_runs_off_the_event_loop` records
`threading.get_ident()` inside the fake adapter and compares it to the
loop thread, and checks the rotation still works end to end (rejected
bearer forwarded first, rotated bearer second).
* `test_event_loop_keeps_running_while_the_retry_credential_resolves`
samples a loop-side heartbeat counter from inside the stalled adapter. On
the unfixed handler it records exactly 0 loop iterations across a 0.5s
rotation.
* `test_retry_credential_failure_still_returns_the_upstream_rejection`
guards the error contract the change must leave alone: a raising rotation
is still swallowed and the upstream's own 401 is streamed back, with no
second forward.
A new `_build_rejecting_upstream` harness drives the `status in {401, 429}`
branch by rejecting every bearer except the rotated one.
Extends `test_proxy_off_loop.py` with the `/health` half, using the same
two-assertion shape as the credential tests:
- `test_is_authenticated_runs_off_the_event_loop` compares the thread the
adapter's `is_authenticated` ran on against the loop thread. Before the
fix they are the same ident.
- `test_event_loop_keeps_running_while_health_resolves_auth_state` reads a
loop-side heartbeat counter sampled by the adapter across its own stall.
Before the fix exactly 0 iterations run across 0.5s.
Both also assert the response is unchanged (`200`, `authenticated: true`),
so the offload cannot quietly alter what `/health` reports.
Adds `tests/hermes_cli/test_proxy_off_loop.py`, mirroring the harness in
`test_proxy.py`: the proxy and a fake upstream run as real aiohttp
servers on ephemeral ports under a single `asyncio.run`, guarded by
`pytest.importorskip("aiohttp")` — no pytest-aiohttp dependency.
The primary assertion is thread identity, not latency. A latency
assertion measured with an HTTP client on the blocked loop is vacuous:
the client's own timer cannot advance until the block ends, so it reports
a fast response on code that was provably frozen.
- `test_get_credential_runs_off_the_event_loop` records
`threading.get_ident()` inside the adapter and compares it to the loop
thread. Before the fix both are the same ident.
- `test_event_loop_keeps_running_while_credentials_resolve` runs a
heartbeat task on the loop and has the adapter sample its counter on
entry and exit, so the reading is taken from the loop rather than
through a client that shares it. Before the fix exactly 0 iterations
run across a 0.5s stall; after it, ~50.
- `test_credential_failure_still_maps_to_401` pins the error contract
across the change of call form. It is deliberately not in the
red-before set — it guards behaviour the fix must leave alone.
GitHub answers anonymous fetches with HTTP 401 during outages (and for
renamed/private repos). git then prompts `Username for 'https://github.com':`
on the inherited terminal and `hermes update` sits there — users read it as
Hermes demanding a GitHub login.
Every network git call in the updater (fetch/pull/push, apply + --check +
fork sync) now runs with GIT_TERMINAL_PROMPT=0 / stdin=DEVNULL, so the 401
fails fast into the fetch-failure classifier, which now reports it as a
GitHub-side rejection (likely outage) rather than blaming the user's
credentials. Credential helpers/askpass are left configured so private-fork
origins still authenticate.
Live repro: PTY-attached update --check against a 401 origin hung 15s+ on
the prompt before; exits rc=1 in 0.2s with the diagnosis after.
Same class as #73751 (@Frowtek, pre-main.py decomposition); passive banner
half salvaged from #101421 (@RobbertC5).
Addresses review feedback on #84928.
The tick interval was exposed as HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS.
AGENTS.md reserves .env for credentials and puts behavioural thresholds in
config.yaml, so the knob moves to `nous.keepalive_interval_seconds`,
following the existing `vertex:` section's precedent for non-secret
provider settings. The env var is dropped rather than bridged: it was never
released, so nothing depends on it. Adding a key to a new section is handled
by the deep-merge, so no _config_version bump is required.
test_keepalive_interval_fits_inside_the_token_lifetime asserted
`900 < 899 * 4 - 120`, which is true for any realistic interval and could
never fail. It also tested the wrong value: the configured constant is only
a ceiling, while the schedule that ships is the derived tick. Replaced with
an assertion over the derived tick for each observed lifetime, which does
fail if the derivation constants regress -- verified against both
TICKS_PER_LIFETIME=1 and MIN_INTERVAL_SECONDS=5000.
Also adds coverage for an unreadable config.yaml, which must fall back to
the module default rather than take the keepalive thread down.
pytest tests/hermes_cli/test_nous_auth_keepalive.py -> 9 passed
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the interval fix: ticking faster narrowed the gap but did not
close it, and the hardcoded interval was wrong for short-lifetime accounts.
Two independent problems, both needed:
1. Lifetime is not a constant. Installs have been observed issuing ~3594s
and ~899s (see #35752, which reports expires_in: 899 while this account
reports 3594). Any hardcoded interval is right for one and wrong for the
other. The tick now derives from the lifetime the server actually issued,
capped by the configured interval and floored at 60s. The access token and
the invoke agent key carry separate lifetimes, so the shorter one governs.
2. Ticking faster alone never closes the gap. The refresh only fires once a
credential is within ACCESS_TOKEN_REFRESH_SKEW_SECONDS (120s) of expiry,
so a tick spaced wider than that window steps straight over it. With a 900s
tick against a 3594s lifetime the last tick before expiry still saw 894s
remaining, declined to refresh, and the credential died 6s before the next
one. The keepalive now asks "will this outlive my next tick?" (tick + skew)
rather than reusing the request path's bare skew.
Measured against both observed lifetimes:
lifetime 3594s: was never refreshed proactively; now refreshes 900s early
lifetime 899s: was never refreshed proactively; now refreshes 227s early
The lifetime is re-read every pass rather than cached, since it can change
when the account, plan, or server-side policy does -- exactly the case the
keepalive exists to cover.
min_access_ttl_seconds is threaded through as an optional parameter, so the
request path keeps its existing 120s behaviour and only the keepalive widens
the window.
Nous Portal access tokens carry a one-hour lifetime, and the keepalive only
refreshes once a token is within ACCESS_TOKEN_REFRESH_SKEW_SECONDS (120s) of
expiry. The tick interval was 6 hours, so it could only land inside that
2-minute window by coincidence. In practice every hour rolled over untouched
and the next inference call paid a 401 plus a re-auth round trip.
Observed on a production install: 71 "refreshed Nous runtime credentials
after 401, retrying" events across current logs.
Drop the default interval to 15 minutes, which gives four ticks per token
lifetime and leaves ample margin under the TTL-minus-skew ceiling of 3480s.
Add HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS so the interval is tunable without
a source edit, matching the existing HERMES_NOUS_TIMEOUT_SECONDS convention.
Zero still disables the thread.
Callers pass no interval, so the default is what actually shipped; the
signature now resolves at call time rather than binding the constant at
import.
Cuts the 41 contributor tests down to 8 pinning the before/after contracts
(out-of-tree provider resolves end to end, copilot-acp unchanged, broken
plugin falls through, flat-install discovery + non-provider kinds untouched).
Adds the create_client hook and process_* fields to the model-provider
plugin developer guide.
An external-process provider is an agent CLI Hermes drives over stdio rather
than an HTTP endpoint. Three things about it were spelled out for one vendor,
and each was a hard stop for any other:
* ``resolve_provider()`` gates on ``PROVIDER_REGISTRY``. Its auto-extend from
``providers/`` covered api-key providers only, so an external-process profile
never entered it and ``hermes -m <that provider>`` died with "Unknown
provider" before a client was ever built.
* ``resolve_runtime_provider()`` keyed the external-process branch on the
literal ``"copilot-acp"``, so anything else silently fell through to the
OpenRouter default instead of its own runtime.
* ``resolve_external_process_provider_credentials()`` hardcoded the binary
(``copilot``), the argv (``--acp --stdio``), the env var names and the
placeholder api_key — so a third-party provider would have been handed
another vendor's CLI.
Now the profile carries what only the provider knows — ``process_command``,
``process_args``, ``process_command_env_vars``, ``process_args_env_var`` — and
the three core paths key on ``auth_type == "external_process"`` instead of a
name. copilot-acp's values move into its profile verbatim, so
``HERMES_COPILOT_ACP_COMMAND`` / ``COPILOT_CLI_PATH`` /
``HERMES_COPILOT_ACP_ARGS`` and its ``copilot-acp`` api_key placeholder behave
exactly as before; the new tests assert that alongside the out-of-tree case at
every step.
The error for a missing binary now names the provider and its own env override
instead of telling every user to install GitHub Copilot CLI.
Co-Authored-By: Junie <junie@jetbrains.com>
Three gaps between the copilot-acp picker row and what the user's
subscription actually serves (reported: picker showed the stale curated
list while the Copilot CLI offered Sonnet 5 / Opus 5 / GPT-5.6):
1. _resolve_copilot_catalog_api_key() never looked at the Copilot CLI's
own token store (~/.copilot/config.json copilotTokens). A user whose
only credential is 'copilot login' got no catalog key, the live fetch
401'd, and copilot-acp silently fell back to the stale curated list.
Add it as resolution source 3, JSONC-tolerant, with each candidate
validated and exchanged like pool entries.
2. The existing credential-pool branch unpacked exchange_copilot_token()
into two names, but it returns (api_token, expires_at, base_url) —
the ValueError was swallowed by the enclosing except, disabling that
entire resolution path. Latent since the base_url return was added.
3. GitHub now returns model_picker_enabled: false for EVERY model on
some accounts/token types, so honoring the flag rejected the whole
live catalog. Treat the flag as a display hint: when it empties the
result, refilter without it (chat/endpoint checks still exclude
embeddings and non-chat rows).
Verified live: catalog resolves 44 models for a copilot-login-only
account, matching the CLI's own picker (claude-sonnet-5, claude-opus-5,
gpt-5.6-sol/terra, gemini, kimi).
Two follow-up gaps found by actually running 'copilot login' end-to-end:
1. The CLI (without an OS keychain) stores its token in
~/.copilot/config.json under copilotTokens — a JSONC file with
//-comment header lines. Add it as an auth-evidence source in
_external_process_auth_evidence(), parsed comment-tolerantly and
counting only a non-empty copilotTokens map (config.json exists after
first launch even when logged out).
2. The desktop chat picker requests explicit_only rows, and
_filter_explicit_provider_rows() dropped copilot-acp because a CLI
login leaves no trace in active_provider, model.provider, or env vars
— exactly the Anthropic-OAuth carve-out case. Keep external_process
rows when their CLI credentials are verified (auth_verified), while
still dropping ambient executable-on-PATH-only rows so the filter's
narrower contract holds.
Net effect: after 'copilot login', copilot-acp appears in the desktop
picker and the Accounts card reads signed in; a machine with only the
binary installed keeps today's hidden-until-configured behavior.
Pin the fix class from the previous commits: auth_type-based dispatch in
get_auth_status(), positive-only auth_verified semantics (supported env
token yes, classic ghp_* PAT no, populated hosts.json yes, empty store no),
and the Accounts-tab cli_command (valid 'copilot login' default, configured
executable substitution, non-external providers untouched).
Review follow-up: the sweeper is right that the fixture's premise leaked.
It clears the tokens and both command variables, but get_auth_status()
treats an `acp+tcp://` base URL as configured on its own — no executable
required — so on a host that sets COPILOT_ACP_BASE_URL the
missing-executable test was answering a question about the host instead
of about the code.
Verified by handing the test the hostile value it was vulnerable to:
with COPILOT_ACP_BASE_URL=acp+tcp://127.0.0.1:9999 in the environment,
test_copilot_acp_hidden_when_executable_missing fails before this commit
and all three tests pass after it.
The overlay loop in list_authenticated_providers() checks every way a
provider might be authenticated — env keys, the auth store, the
credential pool, even Claude Code's external token files — but never
asks the one question that matters for an external_process provider:
does the executable resolve? copilot-acp has no key or token by design
(the spawned `copilot --acp --stdio` brings its own auth), so has_creds
stayed False and the filter dropped it from every picker. Funny enough,
five lines further down the same loop has a dedicated copilot-acp
branch for fetching its model ids — it just never got a chance to run.
Availability now comes from get_auth_status(), the same source
`hermes model` and the auth status endpoints already use, so the CLI
and GUI agree on what 'configured' means for external-process
providers.
Fixes#63662
re_register_config_hooks() cleared the entire process-global idempotence
set on every force-reload, so a profile-local plugin force-reload dropped
another live profile's ledger key without touching its still-registered
callback — the next registration call for that profile then appended a
duplicate. Scope the clear to the reloading profile's own home, and give
outbound webhooks the same force-reload restoration shell hooks already
had, since unload() wipes both from the shared _hooks dict.
Follow-up to #85245: on a profile-resolution exception the facade fell
through to `_authorization_adapter(platform, None)` — the DEFAULT bot —
which is exactly the fallthrough the profile-aware lookup exists to
prevent. Return `adapter_not_registered` instead. Tests now drive the
real `GatewayAuthorizationMixin` ladder (no hand-written stand-in) and
are trimmed to one positive + one parametrized fail-closed case.
Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
PlatformActions._resolve_adapter() unconditionally read
runner.adapters — the default profile's adapter registry — with no
awareness of multiplex/Team-Gateway secondary profiles. Every other
adapter-resolution path in this codebase (GatewayAuthorizationMixin.
_authorization_adapter, and the plugin message-injection path that
uses it) is careful about this split: a secondary profile's adapters
live in runner._profile_adapters[profile], and a stamped secondary
profile with no registry entry must fail closed rather than fall back
to the default profile's adapter of the same platform — falling back
would send replies out the wrong bot.
ctx.platform_actions had none of that: a plugin scoped to one profile
calling add_reaction/set_thread_title in a Team-Gateway deployment
would silently act through the DEFAULT profile's adapter/bot identity
instead of its own profile's, whenever the default profile also ran
that platform.
Route _resolve_adapter() through the same _authorization_adapter
lookup, resolving the calling profile via
hermes_cli.profiles.get_active_profile_name() (which reflects the
per-task HERMES_HOME override multiplex profiles already propagate).
Falls back to the old bare runner.adapters lookup only when the
runner predates _authorization_adapter (defensive, not expected in
practice).
The Windows-only test still modelled the old in-place pack (corrupt exe already
at release/win-unpacked + a .bak to roll back to). Under stage-and-swap the pack
writes into -c.directories.output=<staging>; the integrity gate runs there and a
failure discards staging while the live tree is never touched. Verified on Linux
with sys.platform patched to win32 inside the test; sabotage (skip the discard)
fails it.
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).
Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.
- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
`_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
retry purge and the integrity self-heal only ever clear the staging tree,
never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.
Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.
Closes#86443
Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.
Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.
Salvage of #69118 rebased onto the shared helper (which post-dates it).
Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).
Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.
Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.
Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.
Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).
Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
Salvaged from #100677. Fixes#100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.
- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
it off-thread every TTL window, so every surface on the machine reads
a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.