488 Commits

Author SHA1 Message Date
m4 3ca91c55f3 feat(desktop): openExternalFileForIpc opens files via OS handler
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-09-17 22:28:07 +08:00
teknium1 92d64465bf docs(multiplex): routed-child credential rule and the frozen launch env restart requirement
kvnloo (#111620 review) asked for the operator-visible statement that once a serve /
dashboard process hosts a second profile, the launch profile's env-only credentials are
frozen at activation and a rotation in the process env needs a restart. Also states the
routed-child rule with and without the multiplex flag (#111617). Adds encoding='utf-8'
to the probe child in the child-env authority test (windows-footguns lint).
2026-09-16 00:35:00 -07:00
teknium1 8609743389 fix(serve): launch-profile scope decided at entry; send keeps scope authority; per-reset release
Three edges of the fail-closed multi-profile host (#111620 review, andrexibiza P1 + P2,
kvnloo finding 1):

- `send` under a routed profile's scope `update()`d the installed scope from raw `.env`,
  reversing build_profile_secret_scope's precedence (user .env, then external secret
  sources) for the rest of the request; a stale user value beat the secret-manager one.
  The installed scope is authoritative as-is; only the config.yaml setdefault bridge runs.
- The launch profile's body was scoped only when `is_multiplex_active()` was already true
  at entry, while get_secret consults that global on every read. A launch RPC / dashboard
  request entering single-profile and resuming after a concurrent first `?profile=B`
  activation raised UnscopedSecretError mid-request. The launch profile's secret scope
  (its .env + external sources over the launch env: live while single-profile, the frozen
  snapshot once multiplexing is active) is now bound for every launch-profile body, so the
  credential source is fixed at entry. The terminal policy overlay stays multiplex-only
  (standalone terminal execution keeps its os.environ bridge). _publish_env_value mirrors a
  same-request .env write into that scope AND os.environ for the launch profile, only into
  the scope for a routed one (serves_routed_profile).
- _release_profile_runtime_scope_tokens reset terminal → secret → home in sequence under
  one outer suppress; a failing terminal reset left the previous profile's secrets and
  HERMES_HOME installed for the next body in that context. Each reset is now independent;
  the first failure is re-raised after every scope is released.

tests/tui_gateway/test_multi_profile_hosting_transitions.py: manager-vs-dotenv precedence
through _load_hermes_env, TUI-RPC and dashboard barrier tests (launch enters single-profile,
B activates on another thread, launch resumes and still resolves its injected credential,
never B's), forced terminal-reset failure still releases secret + home. 4/4 red on base.
2026-09-16 00:35:00 -07:00
teknium1 3fe8e5e443 fix(multiplex): routed children never inherit launch-only credentials, with or without the multiplex flag
Two authority gaps in served_profile_child_env (#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):

- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
  provider credentials; strip_launch_profile_env only knows names with .env/source
  provenance, so a key systemd/Compose/the shell injected into the launch process
  survived into profile B's child whenever B did not define the same name. Now a ROUTED
  target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
  before B's own scope is overlaid (the child boundary gets get_secret's contract: a
  scoped miss is no credential, never ambient fallback). The launch profile's own child
  keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
  dashboard backends serve ?profile=B by installing the HERMES_HOME override without
  that flag, so B's slash worker / helper children kept A's .env and settings. The
  authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
  to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
  Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
  under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
  thread); it now raises UnscopedSecretError like get_secret.

tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
2026-09-16 00:35:00 -07:00
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
kshitijk4poor f0b2ba4160 test(tui_gateway): teardown tests prove agent.close() reached the row finalizer
_teardown_session swallows agent.close() failures, so a "row stays open"
assertion alone is green even when close() aborted before
_finalize_owned_session_row. Assert the one side effect only that finalizer
produces (_owns_session_db cleared), give the stub the _memory_manager attr
close() reads directly, and assert the flag is untouched on the control
paths. Rename the sole-owner cross-process test to what it now asserts.
2026-09-16 12:13:05 +05:30
kshitijk4poor dfc40a50a9 fix(tui_gateway): a spared durable row survives agent.close() too
_finalize_session decides whether the TUI backend ends the state.db row
(gateway-owned source → no; another backend holds the lease → no; automatic
Desktop reclaim → no, #105588). _teardown_session then calls agent.close(),
whose _finalize_owned_session_row ends the same row as "agent_close" through
the agent's own handle whenever _end_session_on_close is still True — so every
spare decision was undone one call later. Verified with real SessionDB +
AIAgent.close(): with only the #105620 guard a Desktop ws_orphan_reap still
lands ended_at + end_reason="agent_close" (the #105588 follow-up report).

One seam after the lifecycle decision now clears agent._end_session_on_close
whenever the row was spared, replacing the gateway-owned-only assignment from
#111145 so the lease-held and Desktop-reclaim cases get the same treatment.
The five mocked _finalize_session tests from #105620 are replaced by two
real-DB _teardown_session tests beside the #111145 harness (red on both
origin/main and the bare cherry-picks, green on the stack).
2026-09-16 12:13:05 +05:30
Eva a8b8556108 fix(tui): a gateway-owned session survives agent.close() 2026-09-16 12:13:05 +05:30
ClintonEmok d3a753df30 fix(gateway): don't end durable session row on automatic Desktop cleanup
ws_orphan_reap, idle_timeout, lru_evict, ws_disconnect, and tui_shutdown
are runtime/connection GC — not user intent. Ending the durable session
row during these automatic cleanup reasons confuses GC with user action,
causing secondary bot chats to vanish from the sidebar even though the
transcript is intact in state.db.

The canonical Bot Chat already has resurrection paths for accidental
ws_orphan_reap ends, but non-canonical secondary chats do not, so they
are hit harder.

Skip db.end_session() when _desktop_automatic_cleanup is True (automatic
cleanup reason + Desktop source). Runtime is still reclaimed, the
session.reclaimed event still fires, and explicit user close/archive/reset
still ends the durable row normally.

Fixes #105588
2026-09-16 12:13:05 +05:30
teknium1 c2e5c94cd7 fix(tui): tolerate agents without session_cwd in _register_session_cwd; adapt stubs to the cwd kwarg
Workspace moves stamp agent.session_cwd so a lazily started Codex thread
starts in the moved-to directory. Agents that never had the attribute
(test doubles, slotted objects) must keep working, so stamp only when the
attribute exists. Test stubs of _set_session_context mirror the new cwd
kwarg, and the Codex gateway test asserts the contract (no pinned session
cwd) instead of the attribute's absence.
2026-09-15 22:30:11 -07:00
teknium1 3317b8c1e1 test(honcho): trim salvage coverage to the invariants; drop duplicate ACP session_cwd stamp
Salvage of #93452 (@outpoints). Keep one invariant per fix:
resolver-level "automatic title never remaps a strategy session",
integration "provider routes by logical workspace, not process cwd",
agent-level "title provenance + cwd reach the provider", deferred
Desktop/TUI build threads the session cwd, seeded branch titles are
derived, and the workspace-move E2E. Drop the plumbing/legacy-shape
tests that re-assert the same contract.

acp_adapter: AIAgent(cwd=...) now stamps session_cwd itself, so the
direct assignment after construction was a duplicate.
2026-09-15 22:30:11 -07:00
outpoints 94efd22177 fix(honcho): preserve title provenance through upstream peer routing 2026-09-15 22:30:11 -07:00
outpoints 1c8658bf0d test(honcho): [verified] preserve profile and title provenance coverage 2026-09-15 22:30:11 -07:00
outpoints c2f743f9a1 fix(honcho): [verified] preserve seeded branch title provenance 2026-09-15 22:30:11 -07:00
outpoints 86b3695248 fix(tui): synchronize runtime cwd after workspace moves 2026-09-15 22:30:11 -07:00
outpoints 44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
outpoints 5237cab756 fix(honcho): thread logical cwd through agent construction
(cherry picked from commit b1d7207c45311be658592c6ad34ee84634fed0ee)
2026-09-15 22:30:11 -07:00
teknium1 3c3ab69abb fix(tui_gateway): stop the bot mailbox poll from flooding the log on installs that never received a delivery
The per-session notification poller ran `_poll_bot_live_delivery_once` every
0.5 s. Once a "Bot Chat" session exists, each pass opens state.db and takes
the exclusive active-session registry lock; on Windows (`msvcrt` LK_LOCK
gives up after 10 s of contention) that raised
`RuntimeError: active session file lock unavailable` and the loop logged
`Bot live-owner delivery poll failed` on every attempt — 8,838 warnings in
three days, 91% of one install's WARNING output (#111719).

Two changes:
- `tools/bot_live_delivery.has_mailbox`: the mailbox directory is created
  only when a delivery is first admitted, so a profile without it has nothing
  to claim — the poll now returns before the state.db open / registry lock.
  This keeps cron→Bot Chat and Bot Mode DM delivery intact on installs with
  no messaging platform configured (both deliver through this mailbox), which
  is why the poll is gated on the mailbox rather than on connected platforms
  (PR #111733's guard would have broken those).
- `_poll_bot_live_delivery_guarded`: a failing poll backs off 5 s before the
  next attempt and is logged at WARNING once per 60 s window (with the count
  of suppressed repeats), DEBUG otherwise.

Live probe (real poller loop, temp HERMES_HOME with a Bot Chat row and the
registry lock made unavailable, 3 s):
before: owner_lookups=6 WARNING=6 (with or without a mailbox)
after:  no mailbox -> owner_lookups=0 WARNING=0; mailbox -> owner_lookups=1 WARNING=1

Fixes #111719
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:52:13 -07:00
teknium1 490ee1607a fix(tui): heartbeat ownership follows the gateway's live routing index, not the row source
_notif_gateway_owns_heartbeat decided by the immutable sessions.source column,
so a heartbeat on an ARCHIVED gateway-sourced row (Telegram /reset, idle/daily
auto-reset, compression rotation) was skipped by the Desktop poller and never
registered by the gateway either — restore_heartbeat_watches only claims a key
whose current session_id is that row. The tick belonged to nobody and stayed
due forever, where origin/main's Desktop fired it.

Ownership now uses the same predicate the gateway does: a gateway_routing entry
whose current session_id is this session, with an origin and not suspended
(SessionDB.gateway_routing_entry_for_session; both the session's profile store
and the launch store are consulted so multiplexed and per-profile gateways are
covered). No entry is fail-open, as on main. The check runs after the cheap
is_active/is_due gate so idle sessions never touch the DB per poll.

Probe (SessionStore telegram -> force_new, heartbeat on the archived sid):
before: desktop fired False / gateway watches [] / still due True;
after: desktop fired True / still due False; the current gateway sid is still
left to the gateway (desktop fired False) — the hijack fix stays intact.
2026-09-15 18:57:09 -07:00
teknium1 8a86c56ddb fix(tui): session-owner poller leaves gateway-routed /loop ticks to the gateway
Sibling of the heartbeat fix: a /loop set from a messaging chat carries the gateway's
pinned ``route`` (platform + chat_id), and ``gateway/run_goals.py::_loop_wakeup_fire_one``
already defers route-less CLI/TUI loops to their own schedulers. The TUI/Desktop poller
never returned the favor, so a Desktop viewer of the same session could fire the wakeup
on its own surface and the reply never reached the chat. Mirror the rule: a routed loop
is skipped before the session is claimed, so the tick stays due for the gateway scanner.

Live probe (temp HERMES_HOME, qqbot-routed loop, Desktop viewer of the same session):
before -> loop.ticks=1 awaiting=True (consumed on Desktop); after -> ticks=0, still due.
Desktop-owned control loop still fires.
2026-09-15 18:57:09 -07:00
KoNit-K cf1e727067 fix(gateway): preserve routed heartbeat ownership 2026-09-15 18:57:09 -07:00
teknium1 6f0be923bc fix(tui_gateway): refuse a turn whose session has no agent instead of crashing the turn thread
`_admit_prompt_turn` is the gate every turn source crosses (prompt.submit,
auto-continue, heartbeat/loop ticks, bot mailbox, compute host). When the
record's agent is None (a deferred build that attached nothing, see previous
commit) it used to hand the None agent through; `_invoke_agent` raised
AttributeError, `_recover_turn_exception` emitted a frame, and then the
`finally` dereferenced `st.agent.interim_assistant_callback` again — the
turn thread died before `running = False`, so the prompt vanished and the
session stayed "busy" for every later prompt.

Now the gate releases `running`, emits the same retryable
`agent_init_failed` terminal frame `_run_after_agent_ready` uses (carrying
the recorded build reason when there is one), logs the refusal, and returns
None so `_run_prompt_submit` reports the turn as not started. The turn body
therefore never sees a None agent and needs no guard in its `finally`.

Live probe (real `_run_prompt_submit`, inline thread, agent None):
before: `CRASH out of _run_prompt_submit: AttributeError ... running: True`
after:  `returned: False ... error_surface {'layer': 'runtime', 'code':
'agent_init_failed', 'retryable': True} running: False`

Fixes #111531
Co-authored-by: zqy1-1 <185413950+zqy1-1@users.noreply.github.com>
2026-09-15 18:52:43 -07:00
teknium1 cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
finn763 a7dde8a57d fix(sessions): keep a stream-interrupt recovery inside the original session
A stream that dies mid-answer could leave an orphan session behind: source='unknown', its
first message an assistant message and no user prompt anywhere — invisible to the startup
orphan sweep, unrepairable by the session's own creator. Three links made it permanent:

* the token-accounting guard (hermes_state_usage.update_token_counts, the only writer that
  mints source='unknown') mints whenever the row is missing — which is exactly the state a
  recovery dispatch resumed from: _run_prompt_submit (the crash auto-continue, the
  queued-prompt drain) went straight into the turn without persisting the session's own row,
  unlike the prompt.submit handler, so the first durable writer for that session was the
  accounting side effect, and the turn's prompt could not be written at all (the messages FK
  needs the row);
* _insert_session_row's upsert deliberately keeps what the first writer set, so the real
  creator could never repair that placeholder;
* _ORPHAN_SWEEP_SOURCES skipped 'unknown', so such a row stayed ended_at IS NULL forever.

Every dispatch now binds its own row (original session_key, real source) before the turn
writes anything; the upsert repairs the placeholder source when the session's real creator
arrives; the sweep collects a phantom an older build already left on disk. Regression tests
(red before, green after) in tests/tui_gateway/test_stream_interrupt_recovery_orphan.py.

Refs #111999
2026-09-15 18:23:07 -07:00
KoNit-K 5910de20bc fix(gateway): setup.status / setup.runtime_check scope the launch profile under multiplex
`_readiness_check` bound a secret scope only for a named non-launch
profile and used nullcontext for the launch profile. Once the process
multiplexes (`set_multiplex_active(True)`) `get_secret` fails closed, so
the launch profile's `setup.runtime_check` died on the first profile-scoped
read inside `resolve_runtime_provider` (`HERMES_CODEX_BASE_URL` in
`_pool_entry_mode_and_url`, added by b62bb2a3d5) and the Desktop showed
onboarding while sessions — which resolve under
`_session_profile_runtime_scope` — worked fine.

Route the launch profile through the same helper: `profile_home=None`
binds the launch profile's frozen `.env` scope only when multiplexing is
active (`_profile_runtime_scope_tokens` returns None otherwise, keeping
the single-profile `os.environ` fallthrough for systemd / `op run`
injection). The unknown-profile `ok:False` answer is untouched — no
`@_profile_scoped`, whose `_profile_home` raise would turn it into an
error.

Slimmer shape than the PR's `scope_launch_profile` flag: the flag guarded
nothing the helper does not already decide, and setup.status reads the
same `.env`-derived state.

Fixes #112061
2026-09-15 14:24:49 -07:00
teknium1 57c9e92ae3 test(tui): setup.runtime_check agrees with the session fallback chain
Invariant for #111775 on the registered RPC: primary AuthError + one complete
fallback entry -> ok:True with the fallback provider and model (what
_make_agent builds); explicit `provider` still answers the primary's strict
failure. Red on origin/main (ok:False, "No Anthropic credentials found.").
2026-09-15 14:24:49 -07:00
KoNit-K 6f975b768e fix(tui): setup.runtime_check resolves like session creation and reports the model
Without an explicit `provider`, `setup.runtime_check` called the strict
`resolve_runtime_provider(requested=None)` while `_make_agent` goes through
`_resolve_agent_model_runtime` (startup model + provider pin, then the
configured fallback chain). With the primary blocked and a working fallback
entry the probe answered ok:False (the primary's auth error) and the Desktop
showed onboarding for a backend whose sessions built fine. The probe also
reported `model: null` because the resolver never populates that key.

The default case now runs the session builder's resolver; an explicit
`provider` stays a strict single-provider check (onboarding verifies the
provider just connected — another provider's fallback must not mask a failed
connection) and resolves against the startup model like the session would.
Both branches report the selected model.

Fixes #111775
2026-09-15 14:24:49 -07:00
brooklyn! 43bd5db429 fix(state): bound legacy transcript hydration memory
Co-authored-by: Benjamin Brumbaugh <benbrumbaugh@gmail.com>
2026-09-15 12:33:45 -05:00
kshitijk4poor 8f391f7d48 test(free-tier): reset the boot record and mint memo around each free_tier RPC test
`free_tier.provision` routes through the boot record when one exists, so a
has_identity record left by another file makes it skip the mint and the
lifecycle test fails whenever the two files share a process (CI's per-file
runner hid it). Reset both process memos before and after each test.
2026-09-15 20:44:42 +05:30
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 04e3e89d49 fix(tui): hold back Bot Chat stream deltas that could still be a silence marker
The completion-time emptying only takes effect when nothing streamed: the
Desktop keeps already-streamed text when message.complete carries an empty
text (#95514), so a bare NO_REPLY that arrived via message.delta still
rendered in the Bot Chat bubble. Mirror gateway.stream_consumer: for Bot
Chat sessions withhold deltas while the accumulated buffer satisfies
is_partial_silence_marker; flush the held text verbatim once the buffer
diverges from every marker, so a marker never streams and prose that merely
starts like one is delivered intact.

Review finding: streamed NO_REPLY still visible in Desktop Bot Chat because message.delta carried the marker verbatim.
2026-09-15 06:31:31 -07:00
teknium1 4eb69cbce5 fix(bot-mode): Bot Chat identity + one-shot DM transport apply the silence rule; trim tests
Follow-up to the salvaged #110786 commit:

- `_is_bot_mode_session` mirrors the system-prompt gate (`agent._session_title_hint`
  first, then the live DB title) instead of reading `pending_title`/`title` off the
  session dict: `pending_title` is cleared after turn 1 and the record never carries
  `title`, so the contributor's gate matched only the very first Bot Chat turn.
- `tools/bot_mode_dm.py::_run_local_turn` (the `hermes -p X chat -c "Bot Chat" -Q`
  transport behind `message_agent` when no live owner holds the target) re-emits ""
  for a successful bare marker — the third delivery path of the same class.
- Tests trimmed to one invariant per surface (live completion, relay RPC, one-shot
  transport), each proven red on origin/main sources.
- Bot Mode docs gain a "Staying silent" line pointing at the shared token list.
2026-09-15 06:31:31 -07:00
KoNit-K 50a834fa2a fix(tui): suppress Bot Mode silence markers 2026-09-15 06:31:31 -07:00
teknium1 182ec5c28d fix: widen the remaining config.yaml stat caches to file_signature
Three caches that read config.yaml still keyed change detection on
st_mtime (or st_mtime_ns + st_size), so a same-size replacement that
keeps the old timestamp (cp -p, rsync -t, a timestamp-pinning writer)
was never noticed:

- model_tools._tool_defs_cache_key: get_tool_definitions kept serving
  stale dynamic tool schemas / mcp_servers for the process lifetime.
- CLI mcp_servers auto-reload watcher (cli_tui_mixin seed +
  cli_info_mixin._check_config_mcp_changes): the replaced mcp_servers
  section was never reloaded. The seed is now _config_sig.
- tui_gateway/server._load_cfg_raw / _save_cfg (_cfg_mtime -> _cfg_sig):
  the raw-config cache served the stale document and the next _save_cfg
  would write it back over the on-disk file.

All three now use utils.file_signature like the rest of the PR. Tests that
reset the renamed module/instance attributes follow the rename; one
pinned-mtime replacement test per cache, red on the previous head.
Also corrects two stale type/comment annotations in hermes_cli/config.py
(_env_cache key shape, _RAW_CONFIG_CACHE record shape).

Review finding: three sibling config.yaml caches (tool-defs memo, CLI mcp watcher, TUI-gateway raw cfg) still compared mtime/size only.
2026-09-15 06:29:50 -07:00
teknium1 34184f314e test: browser.manage probe stubs patch the opener, not urlopen
is_browser_debug_ready now dials loopback /json/version through a
ProxyHandler({}) opener so a system proxy cannot capture the probe
(#110565); the TUI-gateway browser.manage tests stubbed urlopen, which
that path no longer calls, so every connect returned False.
2026-09-15 06:16:38 -07:00
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1 928e5993ee test(tui_gateway): drop removed GoalGate.last_failed_fingerprint from control-card fixture
The fix removed the fingerprint field from GoalGate (failed gates re-run every
boundary), so the structured-read fixture can no longer construct one with it.
No other reader/writer of last_failed_fingerprint/workspace_fingerprint remains
in tests/, tui_gateway/, apps/desktop or web/.
2026-09-15 03:59:50 -07:00
teknium1 8b181940b4 fix(multiplex): background threads and teardown paths carry the turn's profile scope
The profile scope (HERMES_HOME override, secret scope, terminal policy) is a
contextvar bundle bound per turn. A bare threading.Thread / Timer / gRPC
callback starts with an empty context and resolves the LAUNCH profile:

- agent/title_generator.py: the auto-title thread read
  auxiliary.title_generation (model, language, provider key) from the default
  profile's config and billed the default's key for a secondary's session.
  Spawn via agent.memory_provider.spawn_context_thread (copy_context).
- tui_gateway/session_lifecycle.py: every teardown caller is a bare Timer
  (ws-orphan reap), the idle-reaper thread, atexit _shutdown_sessions, the
  session.close pool RPC, superseded_by_resume or compute_host flush - none
  carries a scope, yet on_session_end / commit_memory_session / agent.close ->
  shutdown_memory_provider read the provider's config + credentials at call
  time. Under multiplex they failed closed (tail never committed, #110622
  class); on the Desktop backend a secondary's transcript went to the launch
  profile's memory tenant. _finalize_session and _teardown_session now bind
  _session_profile_runtime_scope(session) around those blocks, which covers
  every spawn site through the single chokepoint.
- plugins/platforms/google_chat/adapter.py: Pub/Sub callbacks run on the gRPC
  SubscriberClient's threads and run_coroutine_threadsafe copies THAT empty
  context onto the loop task, so _dispatch_message and everything under it
  (attachment cache, per-user OAuth token store via _acquire_user_chat_api ->
  _load_per_user_chat_api, TTS keys, delivery ledger, bot-id cache) resolved
  the launch profile. connect() captures its scope; _on_pubsub_message and
  _submit_on_loop run under a per-callback copy of it.

spawn_context_thread gains a kwargs passthrough for the title thread's
callbacks.
2026-09-15 03:47:15 -07:00
teknium1 d1794d5539 fix(multiplex): children spawned for a served profile start from that profile's env
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:

- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
  base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
  BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.

tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.

Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
2026-09-15 03:47:15 -07:00
teknium1 b66e7f2149 test(multiplex): invariant tests for fail-closed serve hosting and full-scope profile RPCs / routers
tests/tui_gateway/test_multi_profile_hosting_fail_closed.py — through the
registered RPC handlers: config.get for a secondary resolves only its own
secrets and flips the process fail-closed; llm.oneshot / model.options bodies
see the profile's home + secrets; a launch-profile agent build binds its own
scope once multiplexing is active; CONTROL: a single-profile serve keeps the
os.environ fall-through.

tests/hermes_cli/test_web_multi_profile_scope.py — through the real FastAPI
app: GET /api/config?profile=b expands only B's refs and never mutates
os.environ; console `send` for B lands B's .env in the request scope, not the
process env.

All red on origin/main by source swap (control green), green on head.
test_profile_terminal_scope_entrypoints follows the launch_profile_policy
rename and releases the launch turn's new secret scope.

tests/conftest.py resets the process-global hosting latch
(_MULTIPLEX_ACTIVE, the frozen launch env, _served_profile_homes) per test:
one request routed to a named profile otherwise left every later test in the
file fail-closed.
2026-09-15 03:46:29 -07:00
kshitijk4poor 8bdac1b17a fix(desktop): omit a null iss from the oauth.callback relay; hoist the loopback parse import
The Electron listener now always emits iss (null when the server sent none)
and McpOauthCallbackParams is extra="forbid", so a new Desktop against a
backend without this change would fail every remote MCP OAuth login with a
4000 - including providers that never send iss. Send the key only when set.
Also drop the deliver_callback_flow test the RPC test subsumes.
2026-09-15 13:00:12 +05:30
kshitijk4poor 1c243f86de fix(tui_gateway): relay RFC 9207 iss through the oauth.callback RPC
The oauth.callback handler parsed `iss` but never passed it to deliver_callback_flow, and McpOauthCallbackParams (extra="forbid") had no `iss` field, so the desktop renderer sending `iss: null` was rejected with 4000 "unknown key" — breaking every Desktop→remote-gateway MCP OAuth login. Add the field, forward it, and regenerate the OpenRPC/TS contract artifacts via scripts/gen_gateway_contracts.py.

Also update tests/hermes_cli/test_mcp_dashboard_oauth.py for the 3-tuple callback shape introduced by the cherry-picked commit (it was red on the stack).
2026-09-15 13:00:12 +05:30
OOOOOAO 1a6503a520 fix(mcp): thread RFC 9207 iss through every OAuth callback relay
mcp 2.x rejects an authorization response that omits the RFC 9207 `iss`
parameter when the authorization server advertised
`authorization_response_iss_parameter_supported`. Cloudflare advertises it
AND sends it; the CLI loopback handler has always forwarded it, but every
other callback producer parsed only code/state/error, so the SDK raised:

    OAuthFlowError: Authorization response missing iss parameter
    advertised by the authorization server

and the server parked. Same machine, same config, `hermes mcp login <name>`
from a terminal succeeded — the failure is specific to the non-CLI relays.

Forward `iss` on every producer, matching `_make_callback_handler()`:

- tools/mcp_dashboard_oauth.py: `deliver_callback()` accepts `iss`;
  `wait_for_callback()` returns `(code, state, iss)`. The bridge in
  tools/mcp_oauth.py already splats that tuple into
  `_authorization_code_result(code, state, iss)`, so it needs no change.
- tui_gateway/mcp_oauth_sessions.py: the gateway-hosted loopback listener
  parses `iss`, and `deliver_callback_flow()` forwards it.
- tui_gateway/methods_tools.py: the `oauth.callback` RPC passes `iss`.
- hermes_cli/web_routers/mcp.py: the dashboard callback route accepts it.
- apps/desktop/electron/mcp-oauth-callback-ipc.ts: the one-shot listener
  reads `iss` off the redirect (the renderer already spreads the whole
  callback object into the RPC, so it flows through unchanged).

Providers that omit `iss` round-trip as `None`/`null` rather than being
dropped, so servers that do not advertise RFC 9207 keep working.

Verified live on Windows against mcp.cloudflare.com, whose metadata sets
`authorization_response_iss_parameter_supported: true`: the server that
previously parked on the missing-iss error now reports
`Authenticated — 3452 tool(s) available` and `hermes mcp test cloudflare`
connects. State-mismatch and replay rejection are unchanged.

Tests (each fails on base, passes with the fix):
- test_dashboard_flow_preserves_rfc9207_iss
- test_deliver_callback_forwards_iss (client-redirect relay)
- test_loopback_listener_forwards_iss (real HTTP redirect)
- two vitest cases on the Electron listener, incl. the iss-absent case

Refs #92758, #99984. PR #92765 fixes the dashboard route and the loopback
listener but not the client-redirect relay
(`deliver_callback_flow` / `oauth.callback` / the Electron listener), which
is the path Desktop drives against a remote backend.
2026-09-15 13:00:12 +05:30
brooklyn! b79107c565 fix(gateway): include command context in sudo password requests 2026-09-15 02:22:22 -05:00
teknium1 23ef64081c fix(gateway): bind the session profile once around command.dispatch and bundle routing (#110695)
Follow-up on the salvaged #110698 (which scoped `_dispatch_skill` alone):

- One `_session_home_scope(session)` binding around the whole stage loop in
  `command.dispatch`, so quick commands (`_load_cfg` → `_active_config_path`
  honours the override), bundles (`skill-bundles/` is home-relative) and skills
  all resolve against the SAME profile the routing guard used.
- `slash.exec` resolves `_bundle_key_for(base)` under the same scope, so a
  bundle that exists only in the secondary profile is routed to dispatch at all.
- `_dispatch_skill` uses the home-keyed `get_skill_commands()` (the guard's
  reader) instead of an unconditional `scan_skill_commands()`.
- Test (red on origin/main): a bundle only under profile B's `skill-bundles/`
  is dispatched for a session bound to B.
2026-09-14 16:14:33 -07:00
KoNit-K 4c306e9e8c fix(gateway): scope skill dispatch to session profile 2026-09-14 16:14:33 -07:00
Siddharth Balyan ee2f5629b8 Desktop connect runs on the connection operation: one card, no link to the model, no renderer polling (NS-868) (#110574)
* refactor(connectors): cut comments that restate the code

Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.

* feat(connectors): managed connect runs on the connection operation

Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).

Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.

What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
  transition table. `operation.transition()` enforces it; a card cannot claim a managed
  target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
  reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
  settle -> result). The managed `observe` hook polls the gateway list once per tick for the
  whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
  shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
  are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
  the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
  docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
  added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
  is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
  `connect_card_available` instead of the link when the session platform is `desktop`.

Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.

* feat(desktop): connector card subscribes to the connection operation

The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.

- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
  `applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
  entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
  is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
  snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
  `Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
  failed / expired reissues through `connectors.connect` on the open operation, Not now is a
  per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
  with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
  `HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
  backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
  and compiles unchanged; PR3 moves it onto the operation.

anti-slop: no net-new findings (17 touched files vs 11d1a12472).

* fix(connectors): the card never parks the tool thread; every update carries the snapshot

Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).

- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
  the tool thread on a private request-id Event until a `_respond` that no longer exists for
  this event. `connection.respond` settled the operation but the tool waited its full deadline
  before the watcher loop even started. The callback now only emits the card; the operation's
  own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
  through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
  (state, link, detail). The initial mint happened before the card existed, so the renderer
  never saw the links and Connect stayed disabled; a Continue settlement stamped
  `not_connected` on the backend while the card still showed `initiated`. The store now
  overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
  once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
  `_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
  and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
  `interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
  (the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.

tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.

* docs(connectors): prompts and docs describe the operation, not the deleted wait verb

The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.

* fix(connectors): the panel re-mints only a dead link

Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.

* test(connectors): the local-batch test answers the operation the way the card does

The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.

* ci: retrigger

* fix(connectors): the desktop card appears outside guided onboarding

Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.

The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").

The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.

ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.

message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.

* style(connectors): shorter comments, no module mock in the card router test

The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.

anti-slop: no net-new findings (25 touched files)

* fix(connectors): Connect on a waiting row opens the stored link

ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.

The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.

* fix(connectors): a settled card stays dead; the card binds to its tool call only

A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.

The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.

`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.

`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.

Sid's rule of record: a resolved card is fully dead; no path brings it back.

* fix(connectors): the watch loop settles once, on time, and never raises into the result

Three findings from the live review, one loop.

Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.

Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.

Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.

Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.

* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking

run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.

The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.

* fix(connectors): a failed Try again shows the failure, not the old dead link

The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.

One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.

* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected

`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.

A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.

* fix(connectors): the operation registers under the gateway session key

The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.

The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.

* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed

Three follow-ups from the verification of the fix pass.

The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.

Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.

`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.

`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
2026-09-15 00:41:14 +05:30
Siddharth Balyan d105376b21 Connector code lives in one package, tools/connectors/ (move only; NS-868 prep) (#110368)
* refactor(tools): discovery also scans tools/<pkg>/tool.py

A tool family that is a whole package had no way to register: discovery
globbed tools/*.py only and derived the module name from the filename.
Now the candidate list is tools/*.py plus tools/*/tool.py, merged and
sorted once so import order does not depend on depth (register() lets a
same-name duplicate overwrite silently), and the module name comes from
the path relative to tools/. Only tool.py is scanned inside a package,
so its siblings are libraries by construction. A package without an
__init__.py is skipped with a warning rather than registering from a
checkout and vanishing from the installed wheel.

The AST prefilter and the (mtime, size) disk cache are per absolute path
and work unchanged. The two hand-rolled tools/*.py enumerators in tests
now use the same candidate helper.

* refactor(connectors): one package for the connector domain, tools/connectors/

The connector code was spread across six flat files and a root-level
module that was a sibling of model_tools.py only by address:

  tools/connections_tool.py            -> tools/connectors/tool.py     (schema, register, dispatcher)
                                          tools/connectors/managed.py  (the managed leg, split out)
  tools/connections_tool_mcp.py        -> tools/connectors/mcp.py      (validation split out ->)
                                          tools/connectors/targets.py  (normalize_targets, validate_action)
  tools/connections_tool_operation.py  -> tools/connectors/operation.py
  tools/connector_search.py            -> tools/connectors/search.py
  model_tools_connectors.py            -> tools/connectors/dispatch.py
  tools/tool_gateway/                  -> tools/connectors/gateway/

Move only; every function body is unchanged. tools/connectors/__init__.py
is the door: nine names, the whole cross-package surface. model_tools and
tool_search deep-import a few helpers past it on purpose and the docstring
says so. The two split files make the import graph one-directional
(tool -> mcp -> targets, tool -> managed) where the old layout had
connections_tool importing validation out of the MCP file.

One behaviour-neutral seam change: the _connectors_available try/except
wrapper is gone. connectors_available() already fails closed, and both
the registry handler and the inline executor now read it as a module
attribute (gateway.config.connectors_available), so tests patch it in
one place instead of two. _default_client lives in managed.py, the only
module that calls it.

tools/managed_tool_gateway.py and tools/managed_gateway_auth.py stay:
they are gateway identity shared by tts, transcription, image and modal.

Test files follow their modules. No docs referenced the old paths; no
compat pointer is added (in-tree moves get none).

* ci: retrigger (zero-job dispatch on 142466de6b)
2026-09-15 00:41:13 +05:30
Siddharth Balyan e0ef0eb9c3 manage_connections covers local MCP servers; setup_mcp leaves the schema (NS-867, PR1) (#109517)
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema

One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.

MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.

Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.

`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.

Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.

`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.

The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.

* wip(desktop): connection.request store, resume restore, card routing for MCP targets

Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.

* fix(config): hermes update turns on the connections toolset for saved toolset lists

`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.

Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.

`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.

* refactor: anti-slop pass on the desktop slice; shorten added comments

Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.

* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed

The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.

session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.

lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.

vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.

* chore: drop __pycache__ files swept in by an over-broad git add

* fix(desktop): correlate the connection.request row with the model's tool call by reason

The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.

* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds

* fix(connections): settle reason derives from target state, never from the renderer

A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.

* fix(desktop): a pending connection card re-arms on resume and activate

The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.

* style: literal wording in added comments, docstrings and docs

* fix: shared gateway-event contract and config-schema category for the connection events

connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.

* style: import order (perfectionist) in the desktop and shared files this PR touches

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-15 00:41:13 +05:30