Commit Graph

27345 Commits

Author SHA1 Message Date
Brooklyn Nicholson fdd399dfd3 test(sessions): pin display-projection parity across compaction
Assert the invariant — every display read of one session returns the same
transcript — plus the two things that must not grow with it: the model-fed
projection stays compressed, and Undo/Rewind rows stay hidden. Six of these
fail on the unfixed read paths.
2026-09-02 12:11:15 -05:00
Brooklyn Nicholson 86b204081d fix(gateway): read the compacted transcript for a warm session switch
_live_visible_history backs the payload a tab switch repaints from. Reading
the active-only projection made a compacted chat collapse to its summary on
switch while REST still served the whole thing.
2026-09-02 12:11:15 -05:00
Brooklyn Nicholson ab281990b8 fix(sessions): include compaction-archived rows in every display projection
The gateway's three display reads — session.resume, the ancestor lineage
prefix, and the warm-session payload — filtered active = 1, so a compacted
conversation rendered as its summary plus the carried-forward tail. The REST
transcript read has included the archived rows since #80680, so the same
session read two ways gave two different answers.

Extract the generation-dedupe policy the REST read carried inline into a
shared _dedupe_display_generations() and point every display projection at
it. The model-fed history stays active-only: compaction still compresses the
working context, it just no longer erases the user's transcript.
2026-09-02 12:11:15 -05:00
Teknium 0cbc6e37ac test/docs: trim seam tests to invariants, document create_client and external-process fields
Cuts the 41 contributor tests down to 8 pinning the before/after contracts
(out-of-tree provider resolves end to end, copilot-acp unchanged, broken
plugin falls through, flat-install discovery + non-provider kinds untouched).
Adds the create_client hook and process_* fields to the model-provider
plugin developer guide.
2026-09-02 09:57:39 -07:00
Alexander Prendota 2674d368e6 fix(providers): discover a provider plugin installed by hermes plugins install
`hermes plugins install owner/repo` — and the plugin index behind
`hermes plugins search` — clones into `$HERMES_HOME/plugins/<name>/`, flat, one
directory per plugin. Provider discovery only ever scanned
`$HERMES_HOME/plugins/model-providers/<name>/`.

Nothing joined the two. `PluginManager` does not close the gap either: it
classifies `kind: model-provider` and deliberately skips importing it, because
provider lifecycle is owned by `providers/__init__.py` — which was not looking
in the directory the installer writes to.

So the documented install path half-worked. The CLI reported success, wrote its
install metadata, and the provider silently did not exist: `hermes -m <it>` said
"Unknown provider" and `/model` never listed it. Verified before the fix — a
plugin at `~/.hermes/plugins/<name>/` was NOT FOUND while the identical plugin
at `~/.hermes/plugins/model-providers/<name>/` was discovered.

Discovery now also walks the flat directory, importing only entries whose
manifest declares `kind: model-provider`. Everything else there belongs to
`PluginManager`, which owns its lifecycle and consent flow — importing it here
would run third-party code behind its back, so the tests assert we don't (with
fixtures that register on import, since a fixture that merely raised would be
swallowed by `_import_plugin_dir` and prove nothing).

Manifests are parsed with PyYAML when present and a line scan otherwise, so
provider discovery gains no hard dependency; an unreadable manifest is skipped
rather than allowed to blank the registry.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Alexander Prendota 1bd0b8ed41 feat(providers): let an external-process provider ship out of tree
An external-process provider is an agent CLI Hermes drives over stdio rather
than an HTTP endpoint. Three things about it were spelled out for one vendor,
and each was a hard stop for any other:

* ``resolve_provider()`` gates on ``PROVIDER_REGISTRY``. Its auto-extend from
  ``providers/`` covered api-key providers only, so an external-process profile
  never entered it and ``hermes -m <that provider>`` died with "Unknown
  provider" before a client was ever built.
* ``resolve_runtime_provider()`` keyed the external-process branch on the
  literal ``"copilot-acp"``, so anything else silently fell through to the
  OpenRouter default instead of its own runtime.
* ``resolve_external_process_provider_credentials()`` hardcoded the binary
  (``copilot``), the argv (``--acp --stdio``), the env var names and the
  placeholder api_key — so a third-party provider would have been handed
  another vendor's CLI.

Now the profile carries what only the provider knows — ``process_command``,
``process_args``, ``process_command_env_vars``, ``process_args_env_var`` — and
the three core paths key on ``auth_type == "external_process"`` instead of a
name. copilot-acp's values move into its profile verbatim, so
``HERMES_COPILOT_ACP_COMMAND`` / ``COPILOT_CLI_PATH`` /
``HERMES_COPILOT_ACP_ARGS`` and its ``copilot-acp`` api_key placeholder behave
exactly as before; the new tests assert that alongside the out-of-tree case at
every step.

The error for a missing binary now names the provider and its own env override
instead of telling every user to install GitHub Copilot CLI.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Alexander Prendota 1131b22856 feat(providers): let a provider profile supply its own client
``create_openai_client`` was a hardcoded if-ladder: copilot-acp builds an ACP
stdio shim, gemini builds a native client, everything else gets an
``openai.OpenAI``. There was no extension point, so a provider whose wire
protocol is not OpenAI-over-HTTP could only be added by editing this function —
which is exactly why an ACP provider cannot ship outside this tree today, even
though ``providers/__init__.py`` has discovered out-of-tree profiles from
``~/.hermes/plugins/model-providers/`` and pip entry points for a while.

``ProviderProfile.create_client(**client_kwargs)`` closes that gap. It returns
``None`` by default, so every provider that wants the standard client is
unaffected and the existing ladder still runs as the fallback. copilot-acp is
migrated onto it — its hardcoded branch is gone and its profile supplies the
client in three lines, which is the same three lines an external package writes.

Resolution goes by provider name first, then by ``base_url`` prefix, so a
runtime configured only by URL still reaches its profile — matching what the
replaced ``startswith("acp://copilot")`` branch did. A profile that raises is
logged and skipped: a third-party plugin can fail to provide a client, but it
cannot take the turn down.

Also replaces the two ``isinstance`` checks in ``agent/auxiliary_client.py``
that mean "this client is complete, do not wrap it" with capability flags the
client class declares — ``HERMES_SKIP_TRANSPORT_WRAP`` and
``HERMES_SKIP_ASYNC_WRAP``, mirroring ``SUPPORTS_HERMES_TOOL_CALLS`` in
``background_review.py``. Two in-tree consumers (the ACP shim and the Gemini
native client), an out-of-tree client is covered by the same declaration, and
the hot path no longer imports those modules just to type-test.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Teknium 0ee98eda52 feat(models): add google/gemini-3.8-flash to nous + openrouter catalogs
Slots above gemini-3.7-flash (kept) in OPENROUTER_MODELS and
_PROVIDER_MODELS["nous"]; openrouter plugin fallback_models bumped
3.7 -> 3.8; model-catalog.json regenerated.

Verified live with test completions on both Nous Portal and OpenRouter
(model echo + billed). Same 1,048,576 window / 65,536 output / pricing
as 3.7-flash, so provider-agnostic metadata resolves via the existing
gemini entries and both routes bill live (official_models_api) — no
pricing snapshot needed.

Scoped to the two named providers: vertex/gemini/kilocode/gmi curated
lists, setup.py samples, and aux defaults untouched.
2026-09-02 09:46:15 -07:00
Teknium 3312947e14 fix(auth): concurrent Nous 401 recovery adopts a peer's refresh instead of re-rotating the shared grant
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.

- resolve_nous_runtime_credentials(stale_access_token=): under the store
  lock, skip the refresh POST when the on-disk token differs from the one
  that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
  pass the failed bearer through, and treat a lock TimeoutError as
  'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
  refresh/0 unrecovered.
2026-09-02 09:33:05 -07:00
Ayush Nangia 02103a2104 fix(gateway): drain detached hygiene workers (#98973) 2026-09-02 21:54:01 +05:30
Brooklyn Nicholson 195c3e5a3b fix(desktop): refuse every resume for a chat the user is deleting
The leftover-4001 fix stopped the dispatcher from CREATING a rebind
after a tombstone, but a request queued before it still fired, and the
push path (markRuntimeGone) never checked at all. Both funnel through
requestSessionResume, so guard there: a removal-pending id never gets
queued, and no consumer has to re-derive whether an id is doomed.

isSessionRemovalPending is now the single predicate. resumeSession keeps
its entry guard for requests queued before the tombstone; the session
tile gets the same guard via shouldResumeSessionTile, closing the tile
twin where a 4001 racing a delete unbound the runtime and re-armed the
resume effect against a dead id.

Drops the !freshDraftReady clause on explicitlyRequested: a gateway
switch also stages a fresh draft while leaving the URL on /:sid, where
an explicit request is the only remaining resume lever, so that clause
silently dropped plugin and SDK reselects after a connection apply.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-09-02 11:23:50 -05:00
Brooklyn Nicholson 86acfee949 refactor(desktop): give session tombstones their own store
The delete/archive tombstone atoms lived in store/projects.ts, which
imports store/session. Resume lives on the other side of that edge, so
consulting the tombstones from the resume path would have closed an
import cycle. Move the atoms and their mutators to store/session-removal
and repoint every consumer; no behavior change.
2026-09-02 11:23:50 -05:00
xxxigm 66109d3bbd test(desktop): leftover 4001 rebind must not revive a deleted session
Cover the delete-transition race and the tombstone early-return so the Resume failed toast cannot come back through those paths.
2026-09-02 11:23:50 -05:00
xxxigm 990879d688 fix(desktop): don't rebind a deleted chat from a leftover 4001 resume
A queued requestSessionResume still fired during the /:sid -> /new tick after delete, re-selected the doomed id, and toasted Resume failed / Session not found.
2026-09-02 11:23:50 -05:00
unsupportedpastels afc3d9d34c fix(copilot-acp): prefer stable session config for model selection
Use the ACP v1 session config contract advertised by session/new: locate the category=model option and apply the selected value through session/set_config_option. Retain session/set_model only as compatibility fallback for pre-configOptions agents. Reject unknown and policy-disabled values before prompting.

Verified against the installed Copilot ACP server: its model config option advertises the account-authorized choices, session/set_config_option returns the updated state, and live prompts route gpt-5.6-terra to Terra and claude-sonnet-5 to Sonnet 5.
2026-09-02 20:51:07 +05:30
unsupportedpastels a94b68ad40 fix(copilot-acp): stop substituted models impersonating the requested one
Follow-up to the session/set_model wiring, caught in live use: picking an
org-policy-disabled model (claude-fable-5) produced a response claiming to
BE that model while Copilot actually served its default (Claude Sonnet 5).
Two causes:

1. The prompt preamble injected 'Hermes requested model hint: <id>', so
   whatever model actually served the session parroted the requested name
   back as its identity. Remove the line entirely — the model is applied
   for real via session/set_model now, and identity must come from the
   backend, not prompt suggestion.

2. session/new advertises policy-disabled ids alongside enabled ones
   (_meta.copilotEnablement: 'disabled'); selecting one is accepted but
   silently serves the default. Exclude disabled ids from the offered set
   so the degrade-with-warning path handles them.

Verified live: requesting claude-fable-5 logs the does-not-offer warning
listing the 23 genuinely enabled models, serves the default, and the
response truthfully self-identifies as Claude Sonnet 5.
2026-09-02 20:51:07 +05:30
unsupportedpastels 426dab7de0 fix(copilot-acp): apply the picker-selected model via session/set_model
Selecting a model on the copilot-acp provider had no effect: the model id
never left Hermes. _create_chat_completion() dropped the model argument
before _run_prompt(), so the selection survived only as prompt text
('Hermes requested model hint: ...') and Copilot answered with its own
session default — a user picking gpt-5.6-terra visibly got Claude Sonnet 5.

Live-probing 'copilot --acp --stdio' shows the CLI validates but IGNORES
its --model spawn flag in ACP mode, while session/new advertises
models.availableModels and the ACP-native session/set_model call actually
switches the session. Wire that in: forward the model into _run_prompt,
and after session/new send session/set_model when the id is advertised
(or the server reports no list). Unknown ids degrade to the session
default with a warning instead of failing the turn; the provider-level
virtual slug 'copilot-acp' is never forwarded.

Verified live against the real CLI: requesting gpt-5.6-terra answers as
GPT-5.6 Terra and claude-sonnet-5 answers as Claude Sonnet 5.
2026-09-02 20:51:07 +05:30
unsupportedpastels 5487222658 fix(models): live Copilot catalog for CLI-login users; unbreak picker-flag-empty catalogs
Three gaps between the copilot-acp picker row and what the user's
subscription actually serves (reported: picker showed the stale curated
list while the Copilot CLI offered Sonnet 5 / Opus 5 / GPT-5.6):

1. _resolve_copilot_catalog_api_key() never looked at the Copilot CLI's
   own token store (~/.copilot/config.json copilotTokens). A user whose
   only credential is 'copilot login' got no catalog key, the live fetch
   401'd, and copilot-acp silently fell back to the stale curated list.
   Add it as resolution source 3, JSONC-tolerant, with each candidate
   validated and exchanged like pool entries.

2. The existing credential-pool branch unpacked exchange_copilot_token()
   into two names, but it returns (api_token, expires_at, base_url) —
   the ValueError was swallowed by the enclosing except, disabling that
   entire resolution path. Latent since the base_url return was added.

3. GitHub now returns model_picker_enabled: false for EVERY model on
   some accounts/token types, so honoring the flag rejected the whole
   live catalog. Treat the flag as a display hint: when it empties the
   result, refilter without it (chat/endpoint checks still exclude
   embeddings and non-chat rows).

Verified live: catalog resolves 44 models for a copilot-login-only
account, matching the CLI's own picker (claude-sonnet-5, claude-opus-5,
gpt-5.6-sol/terra, gemini, kimi).
2026-09-02 20:51:07 +05:30
unsupportedpastels 6b2d32d3f6 fix(picker): keep signed-in copilot-acp visible in explicit-only desktop pickers
Two follow-up gaps found by actually running 'copilot login' end-to-end:

1. The CLI (without an OS keychain) stores its token in
   ~/.copilot/config.json under copilotTokens — a JSONC file with
   //-comment header lines. Add it as an auth-evidence source in
   _external_process_auth_evidence(), parsed comment-tolerantly and
   counting only a non-empty copilotTokens map (config.json exists after
   first launch even when logged out).

2. The desktop chat picker requests explicit_only rows, and
   _filter_explicit_provider_rows() dropped copilot-acp because a CLI
   login leaves no trace in active_provider, model.provider, or env vars
   — exactly the Anthropic-OAuth carve-out case. Keep external_process
   rows when their CLI credentials are verified (auth_verified), while
   still dropping ambient executable-on-PATH-only rows so the filter's
   narrower contract holds.

Net effect: after 'copilot login', copilot-acp appears in the desktop
picker and the Accounts card reads signed in; a machine with only the
binary installed keeps today's hidden-until-configured behavior.
2026-09-02 20:51:07 +05:30
unsupportedpastels 15f003e0b9 test(auth): cover external-process dispatch, auth evidence, sign-in command
Pin the fix class from the previous commits: auth_type-based dispatch in
get_auth_status(), positive-only auth_verified semantics (supported env
token yes, classic ghp_* PAT no, populated hosts.json yes, empty store no),
and the Accounts-tab cli_command (valid 'copilot login' default, configured
executable substitution, non-external providers untouched).
2026-09-02 20:51:07 +05:30
unsupportedpastels 323168a289 fix(dashboard): correct copilot-acp sign-in command, honest status card
The Accounts-tab card told users to run 'copilot /login', which is not a
valid invocation — slash-commands only exist inside an interactive session.
Use 'copilot login', the CLI's device-code login subcommand.

The card's status_fn also hardcoded logged_in: False with a static label.
Wire it to get_external_process_provider_status(): claim logged_in only on
positive credential evidence (auth_verified), show which executable Hermes
resolved when merely configured, and say so when the CLI is missing from
PATH entirely.

The rendered cli_command now substitutes the executable the user actually
configured (HERMES_COPILOT_ACP_COMMAND / COPILOT_CLI_PATH) so a custom
binary path gets a copy-pasteable command that matches what Hermes spawns.
2026-09-02 20:51:07 +05:30
unsupportedpastels fd439ac1b8 fix(auth): dispatch external-process providers by auth_type, add positive auth evidence
get_auth_status() special-cased the literal slug 'copilot-acp'; any other
external_process provider (the pending kiro/devin/junie ACP backends) fell
through to {'logged_in': False}. Dispatch on
PROVIDER_REGISTRY[target].auth_type == 'external_process' instead so the
whole class gets a real status.

get_external_process_provider_status() equated 'logged_in' with 'the
executable resolves', which says nothing about whether the Copilot CLI is
actually signed in. Add auth_verified/auth_source: positive-only evidence
from supported env tokens (validated via copilot_auth, classic ghp_* PATs
excluded) or known on-disk GitHub Copilot credential stores. No evidence
means unknown — never presented as signed out, because the CLI may keep its
session in an OS keychain. Deliberately subprocess-free to avoid re-creating
the gh-auth-token cold-start stall (#60800).
2026-09-02 20:51:07 +05:30
Solitud1nem 6b00567718 test(model): clear COPILOT_ACP_BASE_URL in the copilot-acp fixture
Review follow-up: the sweeper is right that the fixture's premise leaked.
It clears the tokens and both command variables, but get_auth_status()
treats an `acp+tcp://` base URL as configured on its own — no executable
required — so on a host that sets COPILOT_ACP_BASE_URL the
missing-executable test was answering a question about the host instead
of about the code.

Verified by handing the test the hostile value it was vulnerable to:
with COPILOT_ACP_BASE_URL=acp+tcp://127.0.0.1:9999 in the environment,
test_copilot_acp_hidden_when_executable_missing fails before this commit
and all three tests pass after it.
2026-09-02 20:51:07 +05:30
Solitud1nem 1e8f829f1a fix(model): show copilot-acp in pickers when its executable resolves
The overlay loop in list_authenticated_providers() checks every way a
provider might be authenticated — env keys, the auth store, the
credential pool, even Claude Code's external token files — but never
asks the one question that matters for an external_process provider:
does the executable resolve? copilot-acp has no key or token by design
(the spawned `copilot --acp --stdio` brings its own auth), so has_creds
stayed False and the filter dropped it from every picker. Funny enough,
five lines further down the same loop has a dedicated copilot-acp
branch for fetching its model ids — it just never got a chance to run.

Availability now comes from get_auth_status(), the same source
`hermes model` and the auth status endpoints already use, so the CLI
and GUI agree on what 'configured' means for external-process
providers.

Fixes #63662
2026-09-02 20:51:07 +05:30
joaomarcos 11942f6de5 fix(gateway): fail closed when a shutdown worker survives the quiesce budget
andrexibiza's review on #101118 pointed out that the timeout branch
still ran the SessionDB close/checkpoint even when _shutdown_executor()
reported a live worker -- the exact sequence that produces the
wrong-page-number corruption in #101093. The close block now only runs
when _exec_live == 0; a surviving worker skips it entirely and leaves
the handle open for SQLite to recover from its own WAL on next open.

Adds test_stuck_worker_skips_the_session_db_close to prove the converse
of the existing ordering test: a worker that outlives the budget must
never be raced by close().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019JujDvCo2vfpEiizidAS2U
2026-09-02 20:43:06 +05:30
joaomarcos af52474aba fix(gateway): quiesce the thread pool before closing state.db at shutdown
`_shutdown_executor()` ran *after* the SessionDB close block in `_stop_impl`,
and it never waited. That left two ways for blocking DB work to outlive
`SessionDB.close()`:

  (a) `_executor_closing` was still False during the close, so a coroutine
      reaching `_run_in_executor_with_context` minted a brand-new pool and ran
      more blocking DB work against handles that had just been closed;
  (b) `cancel_futures` only drops work that has not started, and cancelling
      `self._background_tasks` does not stop the worker thread behind a
      `run_in_executor` future that is already running.

`SessionDB.close()` checkpoints the WAL and lets SQLite unlink the sidecar. A
write that lands after it silently reopens the handle (#94736) and mints a
fresh WAL generation behind that checkpoint, so teardown checkpoints the same
file a second time from a connection the shutdown log never accounts for --
the close-time page-write damage in #101093 and the split WAL generation in
#101064.

The quiesce now runs before the close and waits for the running workers. The
wait is bounded by `_EXECUTOR_QUIESCE_TIMEOUT` (2s) and clamped to what is left
of the shutdown watchdog leash minus a second for the close itself, so a stuck
worker can never cost the post-close cleanup window (#82161). Workers still
alive after the budget are logged as a warning instead of being waited on.

`_shutdown_executor()` keeps its no-argument fire-and-forget contract and now
returns the number of workers still running.

Refs #101093

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DGpGPnvz5FFFH999i59Xeb
2026-09-02 20:43:06 +05:30
kshitijk4poor 95f62ca3bf fix(cron): validate failure_deliver at preflight and dashboard update lanes
Follow-up to the failure_deliver salvage (#100375):

- _preflight_check_delivery also checks the failure lane, so a typo'd
  failure_deliver platform blocks at config-validation time instead of
  surfacing only when a failure occurs — exactly when the notice must
  not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
  deliver (text normalization; empty clears the optional override
  instead of coalescing), closing the one update path that could write
  an unnormalized value into jobs.json.

4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
2026-09-02 20:16:14 +05:30
Victor Kyriazakos fd35e1ec5a fix(cron): delivery bookkeeping reads the failure lane it actually routed through
Review findings (Salt, NS-788):

B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.

S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.

S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.

Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
2026-09-02 20:16:14 +05:30
Victor Kyriazakos c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium 73f68362b3 fix(sessions): auto-prune state.db by default (90d) and gate VACUUM on freelist ratio (#54189)
Flip the state.db retention defaults per Teknium's decision on #54189:

- sessions.auto_prune: false -> true. A stock install now prunes ENDED
  sessions inactive for retention_days at CLI/gateway/cron startup
  (at most once per min_interval_hours). Open, pinned and mid-turn
  sessions are never deleted; the only open rows touched are stale
  automation sessions (#100903 sweep), which are closed, not deleted,
  and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
  the file: PRAGMA freelist_count / page_count must exceed 25%
  (AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
  min_vacuum_interval_days throttle. Pruning a few small sessions on a
  dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
  Unknown ratio (pragma read failure) falls back to the time throttle.

Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.

Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
2026-09-02 07:26:52 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
joaomarcos 65b0f00002 fix(agent): stop the between-turns tool refresh from forking the cached prefix
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:

* a tool whose `check_fn` merely flapped (headless browser probe, expired
  credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.

Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.

`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.

Refs #100336
2026-09-02 07:22:59 -07:00
Teknium 2e25b47210 chore: map contributor email for #80825 salvage 2026-09-02 07:01:23 -07:00
Teknium febd2af391 fix(gateway): skip credential-less WhatsApp on secondary multiplex profiles
_start_one_profile_adapters skipped only Platform.RELAY as shared
process-level ingress. WhatsApp is the same shape: the bridge is one
authenticated session tied to a single phone number, so a secondary
profile has no credential of its own to bring; constructing an adapter
for it only produced a connect/retry loop that stalled startup for every
profile queued behind it. Treat WhatsApp like Relay -- the active profile
owns the connection and route-stamped source.profile fans inbound turns
out to secondary profiles.

Salvage of #69042 (narrowed by its author to this one behavioral line);
test re-expressed on the current secondary-startup fixtures.

Co-authored-by: sshawn <28279366+lsshawn@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Teknium 4fa1d5498e docs(multiplex): list teams among port-binding platforms 2026-09-02 07:01:23 -07:00
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
Teknium 001b8abbd4 fix(matrix): pin the E2EE crypto store per profile at connect(), not import
The multiplex gateway imports plugins/platforms/matrix/adapter.py once, so
the module-level _STORE_DIR/_CRYPTO_DB_PATH resolved against the root
HERMES_HOME for every profile: all bots' Olm identities landed in one
crypto.db and inbound E2EE failed with "no session found" (#89168).

connect() runs inside _profile_runtime_scope, so resolve the store dir
there via get_hermes_dir (honors the context-local HERMES_HOME) and cache
it on the instance -- diagnostics and error-log paths read outside the
scope then still report the store actually in use. Mirrors the
pairing-store fix (a6397c379).

Salvage of #89169 (per-call resolvers collapsed into one cached resolve;
dead `_CRYPTO_DB_PATH = None` alias dropped -- no external importers).
Also routes the last raw MATRIX_HOMESERVER read in check_matrix_requirements
through _startup_env_secret like its token/password neighbours (#69943).

Fixes #89168

Co-authored-by: Michael Short <18595461+mjshorty@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Teknium 9be1168cd3 fix(wecom): scope WECOM_WEBSOCKET_URL like its neighbours
Same class as #100627's WECOM_BOT_ID: the one remaining raw os.getenv in
WeComAdapter.__init__ let a secondary multiplex profile pick up the
default profile's bridged websocket URL. Route it through
_get_scoped_secret; folded into the existing scoped-miss test.
2026-09-02 07:01:23 -07:00
Teknium 2bcbdb61a7 fix(simplex): scope SIMPLEX_* reads to the active profile under multiplexing
SimplexAdapter.__init__ (auto_accept, group_allowed), the registry gates
check_requirements/validate_config/is_connected, _env_enablement and
_standalone_send all read SIMPLEX_* via raw os.getenv. Under
gateway.multiplex_profiles those paths run inside a secondary profile's
scope where os.environ holds the DEFAULT profile's YAML-to-env bridge
output -- so a secondary profile that never configured SimpleX was
auto-enabled on the default's daemon URL and inherited its group
allowlist / auto-accept setting.

Route every read through the module-local `_get_scoped_secret` wrapper
(get_secret; UnscopedSecretError -> os.getenv for the default profile,
which constructs unscoped) -- the same helper the IRC/ntfy/Photon/
Mattermost siblings use. Unlike the extra-only `_scoped_platform_setting`
shape proposed in #100241, this honors BOTH the secondary profile's own
.env (the scope) and its config.yaml extra, and needs no config.yaml
re-read in check_requirements.

Rewrite of #100241.

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
nftpoetrist 07ee457a21 fix(wecom): scope WECOM_BOT_ID reads to the active profile under multiplexing
WeComAdapter.__init__ read WECOM_BOT_ID via a raw os.getenv() call, while
the immediately adjacent line for WECOM_SECRET already used the module's
_get_scoped_secret() helper. Under gateway.multiplex_profiles, a secondary
profile's adapter is constructed inside a scoped context where os.environ
still holds the DEFAULT profile's env-bridge output -- so a secondary
profile's bot would silently connect using the default profile's bot_id
while (correctly) using its own secret, or vice versa on a scope miss.

Switch the bot_id read to _get_scoped_secret(), matching the sibling
_secret/_dm_policy/_group_policy/allow_from reads in the same __init__
that were already migrated in #76664/#93545. _standalone_send's
out-of-process fallback branch constructs a fresh WeComAdapter(pconfig)
and therefore inherits this fix automatically -- no separate change
needed there.

Adds two regression tests to the existing TestWeComAdapterAuthzScope
class (already covering dm_policy/allow_from scoping per #93522),
mirroring its established fixture/assertion style. Mutation-verified:
both fail against the pre-fix code (asserting the default profile's
bot_id leaks into a secondary profile's scope) and pass with the fix.
2026-09-02 07:01:23 -07:00
nftpoetrist 56d869d2d9 fix(mattermost): scope url/reply_mode/require_mention/free_response_channels/allowed_channels to the active profile under multiplexing
MattermostAdapter.__init__, validate_mattermost_config, _standalone_send,
and _handle_ws_event's mention-gating block all read MATTERMOST_URL/
MATTERMOST_REPLY_MODE/MATTERMOST_REQUIRE_MENTION/
MATTERMOST_FREE_RESPONSE_CHANNELS/MATTERMOST_ALLOWED_CHANNELS via raw
os.getenv -- only MATTERMOST_TOKEN was already scoped via
_get_scoped_secret. _apply_yaml_config additionally wrote
MATTERMOST_REQUIRE_MENTION/MATTERMOST_FREE_RESPONSE_CHANNELS/
MATTERMOST_ALLOWED_CHANNELS into the process-global os.environ
unconditionally (guarded only by `not os.getenv(...)`, first-writer-wins),
the same apply_yaml_config_fn bug class already fixed for the
Discord/Telegram/WhatsApp/DingTalk adapters in this series.

Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's
env-bridge output. A secondary profile with its own (or no) Mattermost
config could silently connect to the default profile's server, thread
its replies per the default profile's reply_mode, or -- since
_handle_ws_event's mention-gating block runs on every LIVE inbound
message, not just at construction -- have its require_mention/
free_response_channels/allowed_channels decisions driven by the default
profile's settings for the adapter's entire runtime lifetime.

Fix, mirroring the WhatsApp/DingTalk apply_yaml_config_fn pattern:
- Add _profile_scoped_config_load() (same helper as DingTalk).
- Rewrite _apply_yaml_config to skip the env-bridge write under a
  multiplexed secondary profile's scope, and instead return the YAML
  values as a dict merged into this profile's own PlatformConfig.extra.
- Make require_mention/free_response_channels read extra first (matching
  the existing allowed_channels precedent), falling back to
  _get_scoped_secret() instead of raw os.getenv when extra is absent --
  fixing a residual gap the DingTalk fix (#100615, this series' item 6)
  left in its own analogous extra-first-with-raw-fallback read sites
  (_dingtalk_require_mention et al. still fall back to bare os.getenv).
- Switch __init__'s url/reply_mode, validate_mattermost_config's url, and
  _standalone_send's url to _get_scoped_secret().
- Leave check_mattermost_requirements() (no longer reads any MATTERMOST_*
  var on current main -- just an aiohttp-importability probe) and
  _is_connected() (already scope-aware via hermes_cli.gateway.get_env_value,
  which itself routes through agent.secret_scope.get_secret) untouched.

Adds a new TestMultiplexProfileScope class to tests/gateway/test_mattermost.py
(7 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope, plus two
tests exercising _apply_yaml_config's new seeded-dict return directly.
Mutation-verified: stashed the production fix and confirmed 5 of 7 new
tests fail against pre-fix code (the other 2 are non-differentiating
regression guards -- extra-wins-over-env and unscoped-default-profile-
precedence -- which correctly pass either way). Restored the fix;
all 30 tests in the file, the plugin-setup test, and the full 75-test
tests/gateway/test_adapter_startup_secret_scope.py suite pass.
2026-09-02 07:01:23 -07:00
nftpoetrist c36def6aea fix(photon): scope project_id/node_bin/require_mention/reactions/sidecar config to the active profile under multiplexing
PhotonAdapter.__init__, check_requirements, validate_config,
_env_enablement, _markdown_enabled, _reactions_enabled, and
_standalone_send in adapter.py, plus load_project_credentials and
load_dashboard_project_id in auth.py, all read PHOTON_PROJECT_ID/
PHOTON_NODE_BIN/PHOTON_SIDECAR_PORT/PHOTON_SIDECAR_AUTOSTART/
PHOTON_PROBE_*/PHOTON_REQUIRE_MENTION/PHOTON_MENTION_PATTERNS/
PHOTON_REACTIONS/PHOTON_MARKDOWN/PHOTON_HOME_CHANNEL(_NAME)/
PHOTON_DASHBOARD_PROJECT_ID via raw os.getenv -- only
PHOTON_PROJECT_SECRET and PHOTON_SIDECAR_TOKEN were already scoped via
_get_scoped_secret.

Notably __init__'s project_id read was a stronger variant of the bug
(like the IRC fix in this series, item 11): the original
`os.getenv("PHOTON_PROJECT_ID") or extra.get("project_id") or stored_id`
ordering let a raw env read override even an explicitly configured
config.yaml extra -- a secondary profile that set its own project_id via
extra would still silently authenticate against the default profile's
Spectrum project, because the default profile's project id is always
bridged to os.environ under multiplex and env was checked first.

_reactions_enabled() and the require_mention/mention_patterns reads in
__init__ are exercised on every live inbound message / tapback, not just
at construction, so a secondary profile's reaction/mention-gating
behavior would be driven by the default profile's settings for the
adapter's entire runtime lifetime.

Switch every raw PHOTON_* read (except the two already scoped) to
_get_scoped_secret(), matching the module's existing helper (already
defined identically in both adapter.py and auth.py). Left
_dashboard_host()/_spectrum_host() and the interactive device-login flow
functions in auth.py untouched -- these are CLI-only management-plane
calls (`hermes photon login`/`setup`), not part of the gateway's
per-profile adapter construction/connection lifecycle, so they are not
reachable under a multiplexed secondary profile's scope; noted as a
"Scope note" in the PR body rather than silently expanding scope to
unreachable call sites.

Adds a new tests/plugins/platforms/photon/test_multiplex_profile_scope.py
(9 tests, two classes covering auth.py and adapter.py separately)
mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope, reusing
test_auth.py's tmp_hermes_home isolation pattern so tests don't depend on
the real ~/.hermes/auth.json fallback. Mutation-verified: stashed the
production fix and confirmed 7 of 9 new tests fail against pre-fix code
(the other 2 are non-differentiating regression guards -- unscoped-
default-profile-precedence, one per class -- which correctly pass either
way). Restored the fix; all 140 tests in tests/plugins/platforms/photon/,
the 10 photon-related parametrized tests in
test_adapter_startup_secret_scope.py, and the broader
test_multiplex_adapter_registry.py / test_adapter_connect_classification.py
suites (45 tests) pass.
2026-09-02 07:01:23 -07:00
nftpoetrist 327e9043e9 fix(ntfy): scope server/topic/publish_topic reads to the active profile under multiplexing
NtfyAdapter.__init__, _env_enablement, check_requirements, validate_config,
is_connected, and _standalone_send all read NTFY_SERVER_URL/NTFY_TOPIC/
NTFY_PUBLISH_TOPIC/NTFY_MARKDOWN/NTFY_HOME_CHANNEL(_NAME) via raw os.getenv
-- only NTFY_TOKEN already went through the module's _get_scoped_secret
helper. Under gateway.multiplex_profiles, env_enablement_fn/check_fn/
is_connected all run inside the registry-enablement loop in
load_gateway_config() (confirmed in gateway/config.py, lines ~2704-2820,
inside _profile_runtime_scope for secondary profiles), and adapter
construction likewise runs scoped -- so os.environ there still holds the
DEFAULT profile's env-bridge output. A secondary profile with its own (or
no) ntfy topic configured could silently:

- get auto-enabled via _env_enablement()/is_connected() using the default
  profile's topic, even though it never configured ntfy itself
- have its adapter subscribe to / publish on the default profile's topic
  and server instead of (or in addition to) its own
- deliver cron/send_message_tool messages via _standalone_send to the
  wrong topic

Switch every raw NTFY_* read (except the two secret-material fields
already scoped: NTFY_TOKEN) to _get_scoped_secret(), matching the
established helper already defined in this module and used for
NTFY_TOKEN, and the same pattern applied to the sibling LINE/DingTalk/
Teams/SMS/WeCom adapters in this series.

Adds a new TestMultiplexProfileScope class to tests/gateway/test_ntfy_plugin.py
(7 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope. Mutation-
verified: stashed the production fix and confirmed 5 of the 7 new tests
fail against pre-fix code (the other 2 are non-differentiating regression
guards -- extra-wins-over-env and unscoped-default-profile-precedence --
which correctly pass either way); restored the fix and confirmed all 37
tests in the file, plus the file's 5 parametrized _get_scoped_secret tests
in test_adapter_startup_secret_scope.py, pass.
2026-09-02 07:01:23 -07:00
nftpoetrist 05aa749977 fix(irc): scope server/port/nickname/channel/use_tls reads to the active profile under multiplexing
IRCAdapter.__init__, check_requirements, validate_config, is_connected,
_env_enablement, and _standalone_send all read IRC_SERVER/IRC_PORT/
IRC_NICKNAME/IRC_CHANNEL/IRC_USE_TLS via raw os.getenv -- only
IRC_SERVER_PASSWORD/IRC_NICKSERV_PASSWORD already went through the
module's _get_scoped_secret helper. Under gateway.multiplex_profiles,
env_enablement_fn/check_fn/is_connected all run inside the registry-
enablement loop in load_gateway_config() (gateway/config.py, ~lines
2704-2820), scoped for secondary profiles via _profile_runtime_scope,
and adapter construction runs scoped the same way -- so os.environ there
still holds the DEFAULT profile's env-bridge output.

Notably, __init__'s original `os.getenv("IRC_SERVER") or extra.get(...)`
ordering let a raw env read override even an explicitly configured
config.yaml extra -- a secondary profile that set its own server/channel
via config.yaml extra would still silently connect to the default
profile's IRC server/channel/nick if the default profile bridged its own
config to env (which it always does under multiplex). This is a stronger
variant of the same bug fixed for the sibling LINE/DingTalk/Teams/SMS/
WeCom/ntfy adapters in this series -- there, extra already won because
of the `extra.get(...) or os.getenv(...)` order.

Switch every raw IRC_* read (except IRC_SERVER_PASSWORD/
IRC_NICKSERV_PASSWORD, already scoped) to _get_scoped_secret(), matching
the module's existing helper. Also collapses a double os.getenv("IRC_USE_TLS")
read in __init__ into a single _get_scoped_secret() call (same behavior,
one scope lookup instead of two).

Adds a new TestMultiplexProfileScope class to tests/gateway/test_irc_adapter.py
(6 tests) mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope. Mutation-
verified: stashed the production fix and confirmed 5 of 6 new tests fail
against pre-fix code -- including the "extra wins" test, since IRC's
original env-first ordering meant even an explicit extra config was not
a safe differentiator boundary before the fix (only the DEFAULT-profile-
unscoped-precedence test is a non-differentiating regression guard that
correctly passes either way). Restored the fix; all 23 tests in the file
pass, plus the file's 5 parametrized _get_scoped_secret tests in
test_adapter_startup_secret_scope.py.
2026-09-02 07:01:23 -07:00
Teknium bd81bf0327 chore(contributors): map tky.juani@gmail.com -> JuaniLezcano 2026-09-02 07:00:13 -07:00
Teknium ee0e234a2c fix(gateway): discover and reload MCP servers per profile under multiplex
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.

- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
  served profile inside `_profile_runtime_scope`, carried into the
  executor via `copy_context()` (same shape as
  `_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
  caller (e.g. button-confirm callback) did not; shut down / rediscover /
  report only that profile's servers; refresh only that profile's cached
  agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
  `_server_scope_keys` ownership map; leaves the shared MCP loop running
  while other profiles' servers are live. Unscoped call keeps the full
  historical behavior.
- MCP tools register into the owning profile's registry overlay
  (`registry.register(scope=...)`), and `registry.deregister()` gains a
  matching `scope=` kwarg. Plugin callers still cannot name another
  profile's scope; the plugin-vs-global guard is unchanged for them.

Fixes #95518

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium dc7e1b7ab9 fix(webhook): load URL-resolved profile's skills under multiplex
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.

- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
  when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
  otherwise, same helper the runner uses) and wrap the script / render /
  skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
  `SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
  listed default's skills; the #88023 home-keyed cache alone could not fix
  that. Use the call-time `_skills_dir()` there and at the two other
  SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
  root, so a routed profile's absolute skill_dir was rejected by
  `skill_view` ("must be a relative path within the skills directory").
  Resolve against `_skills_dir()` — the root `skill_view` itself enforces.

Fixes #67277

Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium cd7811a7a7 fix(memory/hindsight): propagate profile scope into background threads under multiplex
Under multiplex_profiles the Hindsight provider's writer, daemon-start and
prefetch threads were spawned as bare threading.Thread, so they started with
an empty contextvars Context: no profile secret scope and no HERMES_HOME
override. get_secret() fails closed there, so the local_embedded daemon never
booted and every retain raised UnscopedSecretError, even though the spawning
thread (initialize()/sync_turn() inside the gateway's copy_context'd turn) had
the scope all along.

Spawn each thread with contextvars.copy_context().run so the child inherits
the spawner's scope + home override. No environ fallback, no re-parsed .env.
The shared hindsight-loop thread needs no wrap: coroutines submitted via
run_coroutine_threadsafe already run in the submitter's context per call.

Fixes #92608
Fixes #94933

Co-authored-by: KIAgent01 <297567825+KIAgent01@users.noreply.github.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium 4e7aa48716 fix(browser): reap idle multiplexed sessions under their owner profile scope
The inactivity janitor is one process-global thread started by whichever
profile first opens a browser, so under `gateway.multiplex_profiles` it runs
with no secret scope: `cleanup_browser` -> `is_camofox_mode` ->
`get_secret("CAMOFOX_URL")` raises UnscopedSecretError, the session entry is
never removed, and the same failure repeats every 30s while the Chromium
daemon leaks.

- `_update_session_activity` records the owning Hermes home per session;
  `_cleanup_inactive_browser_sessions` re-enters that owner's
  `set_hermes_home_override` + `build_profile_secret_scope` around each
  teardown (`_session_owner_scope`, mirroring `_profile_runtime_scope`).
  copy_context at thread spawn would pin the first profile's secrets onto
  every other profile's teardown; there is no os.environ fallthrough.
- 3 consecutive failures -> `_force_reap_browser_session`, which skips the
  failing `close` round-trips but still closes the cloud provider session
  and kills the local daemon via the shared `_release_session_resources`
  tail (extracted from `_cleanup_single_browser_session`, unchanged).
  An activity touch does not reset the failure budget.

Fixes #86402
Fixes #100738

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-02 07:00:13 -07:00
Teknium 98c27aa25c feat(webhooks): stamp the emitting profile on outbound webhook payloads
Under a multiplexed gateway every profile's outbound webhooks share one
delivery worker, so receivers could not tell which profile fired an
event. Add a top-level `profile` field to the payload, resolved at fire
time from the bound Hermes home via get_active_profile_name() ("default"
outside profiles). Documents the field in the wire-format section.

Reported by @vszgdcn8cj-ctrl.

Fixes #92674
2026-09-02 07:00:13 -07:00