Commit Graph

4844 Commits

Author SHA1 Message Date
teknium1 5358780264 fix: keep the carried pre-admission input across the waited lease reload
carry_unadmitted_user_message appends the interrupted turn's user row to the
early result's in-memory history only; it is never persisted because that turn
never owned the lease. When the follow-up turn also has to wait for the lease
(the common case: the other process that made the first turn wait is usually
still busy), admit_durable_turn_lease reloads conversation_history from the DB
after admission and replaced the caller's list wholesale, so the tagged row was
dropped from both the model input and state.db. Re-append the tagged rows that
have no _row_id after the reload so the follow-up turn sees and flushes them.

Review finding: waited-reload branch of admit_durable_turn_lease discarded the
_persist_after_admission_interrupt row carried from the aborted turn.
2026-09-15 04:18:34 -07:00
teknium1 fae3030fa6 refactor: carry the unadmitted user message from the lease sibling, not the facade
Move the pre-admission carry-forward out of run_conversation (facade) into
agent/turn_facade_lease.py::carry_unadmitted_user_message next to the early
result it repairs, and drop the extra turn-start flush in build_turn_context:
the follow-up turn's normal turn-start persist already writes the marked row
because _db_flush_collect no longer stamps it as durable (live probe: exactly
one A row in state.db after two flushes). Trim to two invariant tests
(carry-forward with metadata; flushed exactly once); the hard-stop negative is
covered by the E2E probe in the PR body.
2026-09-15 04:18:34 -07:00
Ayush Nangia 24ae31f7e7 fix(gateway): retain pre-admission interrupted input 2026-09-15 04:18:34 -07:00
teknium1 bc3df8a4d5 fix: NT-namespace guard fires before every sibling resolve (checkpoint, ACP bridge, @file:)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.

The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
2026-09-15 04:17:31 -07:00
Teknium faf71eb4c1 Inspired by Claude Code: file tools reject Windows NT-namespace paths (NTLM leak hardening)
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:

- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
  raw-string check (never resolves — resolving IS the leak trigger).
  Wired as the first check in get_read_block_error() and the write
  denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
  _check_sensitive_path (covers write_file_tool + patch_tool), before
  the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
  local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
  forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
2026-09-15 04:17:31 -07:00
teknium1 ef44c1b73f fix: restore TestBridgeDispatch, trim #39797 tests, drop dead execution_guidance_text param
Why: the rebase conflict resolution in tests/tools/test_model_tools.py
deleted the unrelated TestBridgeDispatch class (3 tests from 73163e3);
it is restored verbatim from main with TestBrowserRetrievalHints after it.

The fix had five tests for one invariant: the OPENAI_MODEL_EXECUTION_GUIDANCE
check duplicated test_phantom_tool_references, and the two static-schema
"toolset-neutral" checks are now folded into test_silent_without_web_tools,
which runs _apply_dynamic_schemas over the real browser_navigate/browser_cdp
schemas so the rendered descriptions are what is asserted.

execution_guidance_text() no longer takes valid_tool_names: the guidance is
toolset-neutral, so the parameter was ignored; the single caller in
agent/system_prompt.py and its test are updated.
2026-09-15 04:13:13 -07:00
teknium1 ac63d0eea5 fix(agent): execution guidance and browser hints drop the web_search stripper; tests assert the invariant
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (3733e4aff5) matched
nothing and were dead; the function now returns the neutral text for
every toolset and its phantom-tool test asserts "no web tool named"
instead of the removed sentence. model_tools ports the PR's hint layer
into main's _DYNAMIC_SCHEMA_REWRITERS table (browser_navigate +
browser_cdp) rather than a second pass after it.

Tests: the two browser_cdp registry tests were re-added by the PR but
main pruned them in 39975613b13b4; replaced with one schema-neutrality
invariant. Exact-wording assertions ("lightweight retrieval tool",
"appropriate permitted retrieval/search tool") were change detectors and
are dropped. tools-reference.md row updated to the new schema text.
2026-09-15 04:13:13 -07:00
Kevin Yin ccf380f634 fix(agent): respect permitted web retrieval guidance 2026-09-15 04:13:13 -07:00
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1 18da11720d fix(session_search): forward after/before/exclude_session_ids through the inline executor
INLINE_TOOL_EXECUTORS["session_search"] is the production dispatch on every surface
(tool_executor + agent_runtime_helpers) and maps schema args to kwargs explicitly, so the
schema advertised after/before/exclude_session_ids while the executor silently dropped them
and every time-bounded call ran unbounded. Add the three mappings.

Also drop the stranded `as_exclusive_end` parameter on _parse_iso_bound (accepted, never
read; exclusivity is the SQL `<` predicate) and the Python re-check of the time window over
FTS rows in _discover — _search_filter_clauses already bounds sessions.started_at on every
route, so the only place the window still needs a Python check is the title-match branch,
which bypasses that query.
2026-09-15 03:58:05 -07:00
Teknium 225d53f953 Port from cline/cline#12876: classify Anthropic output-cap errors 2026-09-15 03:54:58 -07:00
teknium1 123db98635 docs(gemini): scope the base-URL normalization claim to the Google host and TTS
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
2026-09-15 03:54:01 -07:00
Teknium f030c03970 Port from cline/cline#13329: normalize host-root Gemini base URLs to /v1beta
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.

normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).

Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
2026-09-15 03:54:01 -07:00
teknium1 cb84e7d94e fix(codex): import kill_process_tree so the SIGTERM-timeout path actually kills
close() escalated to kill_process_tree() but never imported it; the NameError
was swallowed by contextlib.suppress, so on the TimeoutExpired path neither the
kill nor the post-kill wait ran and a root codex ignoring SIGTERM leaked (a
regression vs the previous self._proc.kill()). Import it from agent.deadline
and add a test forcing the timeout path that asserts the tree kill and the
follow-up wait both run. Also drop the upstream product reference from the
test docstring (credit stays in the PR body) and pass encoding= to the PID
file reads flagged by the Windows footgun scanner.
2026-09-15 03:51:44 -07:00
Teknium a180274a55 fix(codex): reap app-server descendant processes
Port from openclaw/openclaw#126285: snapshot Codex app-server descendants before root retirement and sweep the proven process identities after close so independently grouped stdio MCP children cannot survive client shutdown.
2026-09-15 03:51:44 -07:00
teknium1 8b181940b4 fix(multiplex): background threads and teardown paths carry the turn's profile scope
The profile scope (HERMES_HOME override, secret scope, terminal policy) is a
contextvar bundle bound per turn. A bare threading.Thread / Timer / gRPC
callback starts with an empty context and resolves the LAUNCH profile:

- agent/title_generator.py: the auto-title thread read
  auxiliary.title_generation (model, language, provider key) from the default
  profile's config and billed the default's key for a secondary's session.
  Spawn via agent.memory_provider.spawn_context_thread (copy_context).
- tui_gateway/session_lifecycle.py: every teardown caller is a bare Timer
  (ws-orphan reap), the idle-reaper thread, atexit _shutdown_sessions, the
  session.close pool RPC, superseded_by_resume or compute_host flush - none
  carries a scope, yet on_session_end / commit_memory_session / agent.close ->
  shutdown_memory_provider read the provider's config + credentials at call
  time. Under multiplex they failed closed (tail never committed, #110622
  class); on the Desktop backend a secondary's transcript went to the launch
  profile's memory tenant. _finalize_session and _teardown_session now bind
  _session_profile_runtime_scope(session) around those blocks, which covers
  every spawn site through the single chokepoint.
- plugins/platforms/google_chat/adapter.py: Pub/Sub callbacks run on the gRPC
  SubscriberClient's threads and run_coroutine_threadsafe copies THAT empty
  context onto the loop task, so _dispatch_message and everything under it
  (attachment cache, per-user OAuth token store via _acquire_user_chat_api ->
  _load_per_user_chat_api, TTS keys, delivery ledger, bot-id cache) resolved
  the launch profile. connect() captures its scope; _on_pubsub_message and
  _submit_on_loop run under a per-callback copy of it.

spawn_context_thread gains a kwargs passthrough for the title thread's
callbacks.
2026-09-15 03:47:15 -07:00
teknium1 d1794d5539 fix(multiplex): children spawned for a served profile start from that profile's env
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:

- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
  base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
  BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.

tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.

Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
2026-09-15 03:47:15 -07:00
teknium1 8a8c3634e8 fix(kanban): scope the delegated-child write fence to the lineage's board root
HERMES_DELEGATED_CHILD_CONTEXT=1 is deliberately carried into every shell/
execute_code subprocess a delegate_task child spawns (the fence must survive
exec so a grandchild `hermes kanban complete` cannot promote itself). But the
readers treated the bare flag as "fence every Kanban DB": kanban_db_connect
opened ANY board ?mode=ro and write_txn refused ANY mutation. A subagent
running a Kanban reproduction against a scratch HERMES_HOME therefore got a
silently read-only board with a misleading "descendants require an
initialized board" error; only one lane in the retrospective ever discovered
why (deleg_15dac332), every earlier kanban repro ran degraded.

The marker's value is now the fenced board ROOT (kanban_home() at spawn) and
readers deny only paths under that root or the dispatcher-pinned
HERMES_KANBAN_DB (kanban_path_is_fenced). In-process children and a legacy
"1" marker still fence everything; an inherited path marker is never
re-derived, so a grandchild that moved HERMES_HOME cannot unfence the real
board. Owner-gate tests (test_kanban_descendant_scope, cron env isolation,
kanban CLI exit status) are unchanged and green.
2026-09-15 03:45:41 -07:00
teknium1 87ce653d1d feat(relay): migrate legacy HERMES_NEMO_RELAY_ATIF_*/ATOF_* vars into a validated relay-plugins.toml
3fad83df31 (Aug 11) moved Relay exporter config to a plugins.toml
selected by HERMES_NEMO_RELAY_PLUGINS_TOML. A .env still carrying the legacy
exporter vars and no TOML logs ONE warning and initialises no exporters, so
users who followed the earlier docs lost every trace silently (the
maintainer's stopped Aug 20, noticed Sep 14; five multiplexed profiles on
the same box carry the same eight vars today).

- `hermes_cli/relay_plugin_migrate.py`: build the document from the
  `nemo_relay.observability` dataclasses (`ComponentSpec(...).to_dict()`,
  so the `type = "file"` sink discriminator is emitted), validate it by
  activating it through `nemo_relay.plugin.initialize` + `clear_async`,
  write `<home>/relay-plugins.toml` (tomli_w when installed, minimal emitter
  otherwise), set HERMES_NEMO_RELAY_PLUGINS_TOML in that .env, and comment
  the legacy lines out (never delete). Defaults mirror the removed plugin so
  files land where they used to.
- `hermes update` runs it for the default home AND every live named profile
  (each writes its own TOML) as a best-effort post-update step, with a loud
  notice; `hermes migrate relay [--all-profiles] [--no-validate]` runs it on
  demand.
- The runtime WARNING and the `hermes doctor` finding now say "NO traces
  are being exported" and name the exact command and file path.
- Docs: environment-variables.md + built-in-plugins.md carry the migration
  note and a complete plugins.toml example including `type = "file"`.
2026-09-15 03:44:36 -07:00
teknium1 88f2844d46 fix(agent): cover the remaining refusal-only surfaces and fold the tests
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
  stream watchdog sees progress on a refusal-only stream instead of timing it
  out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
  parts; without it an aux refusal-only turn parsed to content=None and hit the
  empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
  alongside-content) so there is one test per surface; drop upstream product
  references from docstrings (credit stays in the PR body); pass encoding= to
  the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
  content_filter result, not an empty response to retry.
2026-09-15 03:39:07 -07:00
Hermes Agent e27f16365d fix(agent): preserve streamed refusals as text (port of anomalyco/opencode#43343)
A model that declines mid-stream delivers the explanation on the
structured refusal channel (chat_completions delta.refusal; Responses
response.refusal.delta / refusal content parts) and leaves content
empty. The streaming accumulators dropped that channel entirely, so a
streamed refusal assembled into an empty message and fell into the
empty/invalid-response retry loops - burning paid retries reproducing a
deterministic refusal - while the non-streaming path had already fixed
this class in #46013.

- chat_completions streaming: accumulate delta.refusal (incl.
  model_extra), expose message.refusal on the assembled mock response so
  ChatCompletionsTransport.normalize_response applies the existing
  sole-payload -> content_filter promotion; count refusal deltas in the
  zero-chunk guard; carry refusal in the Relay final-response dict.
- Codex Responses stream consumer: collect response.refusal.delta as
  answer text so a refusal-only stream no longer raises 'did not emit a
  terminal response' with zero usable content.
- Responses normalizer: read type=refusal content parts in
  _extract_responses_message_text (attr and dict shapes).

Sabotage-verified: each new test fails with its wiring line disabled.
E2E: refusal-only stream -> terminal content_filter with explanation;
refusal-alongside-content stays a normal usable turn; plain-text
streams unchanged.
2026-09-15 03:39:07 -07:00
teknium1 e860b8e4e4 fix(context): compute-host /context and session.context_breakdown carry the per-file manifest; report blocked files
Why: the tui_gateway live formatter (`_format_live_context_output`, used when
the session runs on a compute host) renders its own summary and never got the
"Context files" block, and `session.context_breakdown` had no structured rows,
so Desktop's popover could not show them. The formatter now appends
render_context_file_lines() with the session cwd bound (the RPC thread has no
session context, so the discovery walk would key on the backend's cwd), and
the RPC payload gains a `context_files` list (contract + generated TS/OpenRPC
+ Desktop type). The docs sentence is scoped to the surfaces that render it.

A file whose content _scan_context_content replaces with a BLOCKED marker was
reported "loaded"; the manifest now runs the same scan and reports `blocked`.
The module docstring names the frontmatter-strip / chain-cap approximations
and drops the product-name attribution (credit stays in the PR body).
2026-09-15 03:37:49 -07:00
teknium1 f271ba09b0 fix(context): derive the /context file listing from the builder's own discovery walk
Review follow-up (Enough1122) on the salvaged #91272: the original
list_context_file_sources() hand-mirrored the priority ladder inside
build_context_files_prompt, so the two would drift the moment the builder
gained a context type or changed precedence — misreporting what the prompt
holds is worse than not showing it.

Now prompt_builder exposes one candidate finder per context type
(_CONTEXT_FILE_CANDIDATES → discover_context_files) and BOTH the loaders and
the manifest walk it. The manifest lives in the new sibling
agent/context_file_sources.py (not appended to the facade) and:
- reports empty / unreadable files truthfully instead of "✓ 0 tokens",
- mirrors the install-tree guard ("suppressed") so a Desktop session that
  fell back into the Hermes tree sees why nothing loaded,
- lists every .cursor/rules/*.mdc as loaded, matching the builder which
  concatenates all of them,
- measures truncation on the rendered "## label" section like the builder.

The block now renders on every surface that shows the /context category
table: CLI/TUI (hermes_cli/cli_info_mixin.py) and the messaging gateway
(gateway/slash_commands_status.py). The Desktop popover consumes the raw
session.context_breakdown payload (no text table) and is left as-is.

Tests trimmed to the two invariants: manifest/prompt parity across every
context type at once, and truncated/suppressed follow the builder.
2026-09-15 03:37:49 -07:00
teknium1 286e723db8 fix(skills): auto-load resolves under the agent's own home and skips internal forks
Why: build_auto_load_prompt read config via ambient load_config_readonly()
and looked skills up under the ambient SKILLS_DIR. Gateway bot threads lose
the HERMES_HOME ContextVar, so a bot profile's pinned skills came from the
launch profile — the docs promise profile scoping. _auto_load_parts now
passes home_override=_agent_home(agent) and build_auto_load_prompt binds it
for config, disabled-list and <home>/skills lookup, the same seam
_skills_prompt uses via skills_dir_override.

_auto_load_parts was unconditional and injected pinned SKILL.md bytes into
delegate children, curator/background_review forks and gateway hygiene
agents; it now mirrors _skills_prompt's gate (nothing without the skills
toolset) and returns [] when skip_context_files is set.

cli.py's HERMES_IGNORE_RULES check used == "1" while system_prompt used
is_truthy_value; both use is_truthy_value now.

Tests stay at 4: the build test asserts the home-scoped resolution, the
ignore-rules test also covers the subagent / no-skills-toolset gates.
2026-09-15 03:34:20 -07:00
Carl Taylor 075a256597 fix(skills): auto-load resolves once per agent and dedupes against -s
The system prompt must stay byte-stable for the life of a conversation:
`_auto_load_skills_result` is seeded in `_SESSION_STATE` and filled on
the FIRST prompt build only (HERMES_IGNORE_RULES captured then too), so
model switches, compression and static-prefix restoration reuse the
exact rendered bytes rather than re-reading config or skill files.

CLI: auto_load renders in the existing background `--skills` preload
thread (real session id for ${HERMES_SESSION_ID}), `-s` names dedupe
against the auto-loaded canonical names via
`build_preloaded_skills_prompt(excluded_loaded_names=)`, the activated
skills line shows auto_load first, and the lazily built agent is seeded
with the pre-resolved bytes. `--ignore-rules` skips auto-load with the
rest of the auto-injected context.

Re-implementation of #74060 by @ctaylor86 against current main.
2026-09-15 03:34:20 -07:00
ArcherQAQ 1976869c01 feat(skills): skills.auto_load pins skills into every new session's prompt
`skills.auto_load: [name, ...]` in config.yaml renders the listed skills
as fully loaded skill blocks in the system prompt of every new agent —
CLI, TUI, gateway, cron and API — the persistent counterpart of `-s`.
Missing or operator-disabled names are warned about and skipped; a
config typo never blocks session start.

Re-implementation of #26840 by @ArcherQAQ (via #74060) against current
main — the original patch targeted the pre-decomposition cli.py /
system_prompt.py god files; design and diagnosis preserved. The loader
reuses `_load_skill_blocks` (same disabled gate and Curator usage bump
as `-s`) instead of a parallel loop.
2026-09-15 03:34:20 -07:00
kshitijk4poor a55c972e09 refactor(gemini): prefix rule only for no-id slot matching; share the provider-id guard
Slot arguments are always complete json.dumps output (Gemini re-sends full args), so the mid-stream JSON check could never fire; the id key needs no tool name; _new_call_id and the slot lookup now share _provider_call_id.
2026-09-15 13:27:24 +05:30
jmiguellucas 91665b5e79 fix(gemini): give each native tool call its own streaming slot
Two different calls to the same tool arriving in separate stream events
collided in one accumulator slot: Gemini 2.5 sends no call id and
part_index restarts at 0 per event, so the second call's arguments were
emitted as a delta on the first call's index and concatenated downstream
into unparseable JSON, dropping a call. Gemini 3 ids are now the slot
identity (part_index and thought signature drift across events of one
call); without an id, a call whose arguments are not a continuation or
resend of the slot's accumulated JSON opens its own slot, kept reachable
as key#N so a later resend lands on it.

Re-applied by hand onto the collapsed translate_stream_event on main from
#75528 (9371874010 + f4c8863cdc). #24676 by cdbartholomew (May 13) was
the first fix for this collision (value-based slot matching without the
id key) and is credited as co-author.

Co-authored-by: Chris Bartholomew <chris.bartholomew@vectorize.io>
2026-09-15 13:27:24 +05:30
kshitijk4poor b82c79ac50 refactor(agent): share the shutdown-only socket primitives with the pool sweep
The stale-attempt socket shutdown re-implemented two blocks that already
live in agent_runtime_helpers: the settimeout(0)+shutdown(SHUT_RDWR)
body of force_close_tcp_sockets (now _shutdown_socket) and the
candidate->socket lookup of _iter_pool_sockets (now _socket_from_candidate).
The hand-unrolled _httpcore_stream unwrapping is dead since
_connection_candidates walks _stream/_httpcore_stream itself, so the helper
starts from the network_stream extension and the response stream only.

Also add ReadError to _TRANSIENT_TRANSPORT_ERRORS (the third classifier of
the same abort-induced read; the other two were already updated), drop the
incidental gettimeout() assertion from the shape test, and cut the E2E
from ~3 s to ~1.5 s (stale budget 1 s, serve_forever poll 50 ms).
2026-09-15 12:46:27 +05:30
kshitijk4poor e625602a67 test(agent): trim stale-kill unwedge tests to the two invariants
Keep the real httpx 0.28 wrapper-shape test (proves the shutdown reaches
the socket through BoundSyncStream/ResponseStream/PoolByteStream) and the
loopback E2E (a parked reader unwinds within its stale budget and the
retry lands). The other four were narrower restatements of the same paths.

Also treat httpx.ReadError as a transport error in codex_runtime: it is the
same abort-induced-read class the streaming retry loop now recovers from.
2026-09-15 12:46:27 +05:30
kshitijk4poor 13f9206e52 fix(agent): drop the stale-aligned read-timeout cap from _stream_timeouts
Capping ``read`` to ``stale`` silently overrode an explicit
HERMES_STREAM_READ_TIMEOUT and the local-endpoint ``read = base`` branch,
and let the stale-kill E2E pass via ReadTimeout alone rather than through
the socket-shutdown unwedge this change is about. The shutdown path is
sufficient on its own: the E2E still passes with the cap removed.
2026-09-15 12:46:27 +05:30
finn763 2355d593f3 fix(agent): stale-killed stream unwedges its reader and reconnects (#110769)
The stale-stream monitor aborted a wedged provider stream only via
force_close_tcp_sockets() -> shutdown(SHUT_RDWR). That is best-effort: a
parked body read is not unblocked on every platform (Windows keeps the
pending recv parked) and the sweep can miss the socket. The worker then
stayed blocked in the provider read, so the retry loop never retried; the
monitor re-killed every stale interval and the call only ended at the
byte-read timeout, far past the stale budget - the reported
"No response from provider for 180-240s ... Reconnecting" loop ending in
"The model server is not responding".

- _kill_stale_stream now also closes the killed attempt's own provider
  response (identity-guarded self._attempt_stream_response), which is what
  actually unblocks a parked reader; a racing retry's fresh response is
  never touched.
- an abort-induced httpx.ReadError counts as a transient connection error,
  so the aborted attempt reconnects instead of ending the turn.

Reproduced with a local SSE server: before, the worker stayed parked and no
second request was issued (recovery only at the byte-read timeout); after,
the kill unblocks the reader at the stale budget and the retry lands.
2026-09-15 12:46:27 +05:30
kshitijk4poor cedf4a3d78 fix(secret-scope): compose the managed .env into every profile secret scope
Answers the P1 review on #111187: build_profile_secret_scope() held only
<profile>/.env plus that profile's external-source snapshot, never the
administrator-managed .env. The launch process applies that file LAST with
override (_apply_managed_env), so a managed key beats the user's own value in
os.environ. Under multiplex semantics get_secret() stops falling back to
os.environ on a scope miss, so inside a routed cron fire (and equally inside a
real multiplex gateway turn, which builds its scope through the same function
via gateway/run.py::_load_profile_secret_scope) a managed-only credential
resolved as absent and a managed-vs-user collision resolved to the USER value:
reversed precedence.

Fix at the source: build_profile_secret_scope() overlays load_managed_env()
last, after the profile .env and external sources, skipping process-global
names exactly as it does for the other two layers. Every multiplex-authoritative
scope (gateway turn, routed desktop fire, external worker env build) is built
here, so managed authority is composed once instead of restored per consumer.
No generic ambient-env fallback is reintroduced: only the managed file's own
keys enter the scope, and only with the managed file's values.

Regression (parametrized, two invariants): inside a routed fire a managed-only
key resolves through get_secret(); a managed-vs-user collision yields the
managed value.
2026-09-15 11:03:39 +05:30
kshitijk4poor 0916191ca1 refactor(env-loader): one helper records source-supplied names; scope refresh never empties
_hydrate_profile_secret_sources and _apply_external_secret_sources each
rebuilt the same "applied + skipped_existing" set and pushed it into
_SOURCE_SUPPLIED_NAMES; the routed-child scrub depends on both sites
agreeing, so give them one helper.

refresh_installed_secret_scope cleared the installed scope before
refilling it, which left a window where a concurrent reader of the same
fire saw no credentials at all. Update first, then pop the names the
rebuild no longer supplies; stale values still disappear.
2026-09-15 11:03:39 +05:30
John Paul Soliva 0943e77136 fix(cron): strip launch external-source names too, and make the scope refresh replace
Two credential-isolation gaps found in review of the previous head.

1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.

2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.

Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.

(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
2026-09-15 11:03:39 +05:30
John Paul Soliva dbede34f6e fix(cron): a routed profile's cron fire in the desktop backend runs under multiplex semantics
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).

Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).

Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
  the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
  passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
  and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
  the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
  `_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.

Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.

(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
2026-09-15 11:03:39 +05:30
kshitijk4poor 14a93d67eb refactor(codex): take the aux adapter's xAI flag from the shared route classifier
`_CodexCompletionsAdapter._build_responses_kwargs` already asks
`classify_responses_route` for the GitHub flag but hand-rolled its own
x.ai host match for `is_xai`. Two classifiers for one route drift; use
`route.is_xai_responses` so the aux path agrees with the main transport.
`is_copilot` (used for `is_github_responses` replay stripping) is left
as-is.
2026-09-15 10:49:19 +05:30
kshitijk4poor 00f0d92be0 fix(codex): canonicalise persisted issuer stamps via the route-identity owner
`_classify_responses_issuer` reimplemented endpoint canonicalisation with
its own urlsplit/urlunsplit pass. The repo already owns that logic in
`hermes_cli/route_identity.py::normalize_route_base_url` (stdlib-only,
used by agent/backend_identity.py), so delegate to it.

Reasoning items persisted before canonicalisation were stamped with the
raw `other:<agent.base_url>` (trailing slash, host case). Comparing them
verbatim against the now-canonical `current_issuer_kind` marked them
foreign and dropped them on the very endpoint that minted them. Run the
persisted stamp through the same canonicaliser (`_canonical_issuer_kind`,
non-`other:` kinds untouched) before comparing.

Also fixes the `_chat_messages_to_responses_input` docstring, which still
stated the pre-stack rule that legacy endpoint-stamped items drop when the
current model is known; the stack replays them on a matching issuer.
2026-09-15 10:49:19 +05:30
kshitijk4poor 16986c4bff fix(codex): canonicalise the custom-endpoint issuer kind
The openai SDK appends a trailing slash to `client.base_url`, so the aux
adapter stamped `other:https://h/v1/` while the main transport stamped
`other:https://h/v1`. On custom Responses endpoints every aux call
(compression, flush_memories) therefore dropped all main-minted reasoning
items as "foreign".

`_classify_responses_issuer` now strips whitespace and trailing slashes and
lowercases scheme+netloc before stamping. The aux adapter also derives its
route flags from `classify_responses_route` — the single owner of the
codex/xai/github predicates — instead of an inline chatgpt.com host check,
and reuses the same flags for the effort clamp.
2026-09-15 10:49:19 +05:30
kshitijk4poor cb49660bb4 fix(codex): replay legacy endpoint-stamped reasoning without a model stamp
Native compaction checkpoints and reasoning items persisted before model
stamping existed carry only `_issuer_kind`. Treating a missing `_issuer_model`
as foreign dropped every such item once the current model was known, which
wiped existing sessions' native-compaction context on upgrade (four consumer
tests in test_native_compaction / test_native_preflight_estimate /
test_413_compression went red on the stack).

Trust the endpoint stamp when no model stamp is present, as main does today.
Items minted after this change carry the model stamp and still drop on a
same-endpoint model switch; a wrong guess on a legacy item is caught by the
invalid_encrypted_content 400 classifier and the replay kill switch.
2026-09-15 10:49:19 +05:30
Fangliquan 51ebdff570 fix(codex): scope encrypted-reasoning replay to the issuing model
Encrypted reasoning blobs are sealed to the model that minted them, not
only to the endpoint. Switching models on the same custom Responses
endpoint therefore replayed blobs the new model cannot decrypt and the
turn failed with HTTP 400.

Stamp captured reasoning items with `_issuer_model` (the canonical wire
model) alongside `_issuer_kind`, and replay an item only when both the
issuer kind and the model match the current request. Endpoint-stamped
legacy items without model provenance are dropped once the current
model is known (fail closed); ordinary assistant text stays replayable.
The transport threads the effective wire model (request_overrides win)
into conversion and normalization; the auxiliary Codex adapter stamps
and filters against its own model rather than the main agent's. The
400 classifier also recognises the custom-endpoint wording
"encrypted content could not be decrypted or parsed" so recovery strips
the replay state instead of aborting.

Hand-grafted from #95849 (final head d9cf6bcc08) onto current main; the
middleware-model-rewrite half is intentionally left out.

Closes #95834
2026-09-15 10:49:19 +05:30
fangliquanflq f7b6a2b59f fix(codex): drop foreign replay message ids
(cherry picked from commit b58b94e16ccf5c8e01813f25f6a1a94dfb17f79b)
2026-09-15 10:49:19 +05:30
kshitijk4poor 4bd38ec9fd fix(plugins): give output-transform hooks a call identity for the callback gate
Hook callbacks are gated per (hook, callback, call identity). Two hooks
on the tool-loop path fired without any identity, so concurrent terminal
calls in one turn, or overlapping turns, still collapsed onto a single
gate key and the second invocation was skipped as if a callback had hung.

transform_terminal_output now forwards the tool_call_id bound in the
approval context around dispatch (only when set); transform_llm_output
forwards the turn_id already in scope. Payloads are additive: the
dispatcher withholds unknown fields from narrow-signature callbacks.
2026-09-15 10:48:49 +05:30
Siddharth Balyan ee2f5629b8 Desktop connect runs on the connection operation: one card, no link to the model, no renderer polling (NS-868) (#110574)
* refactor(connectors): cut comments that restate the code

Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.

* feat(connectors): managed connect runs on the connection operation

Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).

Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.

What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
  transition table. `operation.transition()` enforces it; a card cannot claim a managed
  target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
  reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
  settle -> result). The managed `observe` hook polls the gateway list once per tick for the
  whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
  shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
  are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
  the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
  docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
  added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
  is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
  `connect_card_available` instead of the link when the session platform is `desktop`.

Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.

* feat(desktop): connector card subscribes to the connection operation

The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.

- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
  `applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
  entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
  is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
  snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
  `Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
  failed / expired reissues through `connectors.connect` on the open operation, Not now is a
  per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
  with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
  `HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
  backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
  and compiles unchanged; PR3 moves it onto the operation.

anti-slop: no net-new findings (17 touched files vs 11d1a12472).

* fix(connectors): the card never parks the tool thread; every update carries the snapshot

Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).

- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
  the tool thread on a private request-id Event until a `_respond` that no longer exists for
  this event. `connection.respond` settled the operation but the tool waited its full deadline
  before the watcher loop even started. The callback now only emits the card; the operation's
  own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
  through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
  (state, link, detail). The initial mint happened before the card existed, so the renderer
  never saw the links and Connect stayed disabled; a Continue settlement stamped
  `not_connected` on the backend while the card still showed `initiated`. The store now
  overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
  once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
  `_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
  and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
  `interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
  (the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.

tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.

* docs(connectors): prompts and docs describe the operation, not the deleted wait verb

The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.

* fix(connectors): the panel re-mints only a dead link

Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.

* test(connectors): the local-batch test answers the operation the way the card does

The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.

* ci: retrigger

* fix(connectors): the desktop card appears outside guided onboarding

Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.

The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").

The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.

ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.

message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.

* style(connectors): shorter comments, no module mock in the card router test

The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.

anti-slop: no net-new findings (25 touched files)

* fix(connectors): Connect on a waiting row opens the stored link

ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.

The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.

* fix(connectors): a settled card stays dead; the card binds to its tool call only

A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.

The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.

`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.

`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.

Sid's rule of record: a resolved card is fully dead; no path brings it back.

* fix(connectors): the watch loop settles once, on time, and never raises into the result

Three findings from the live review, one loop.

Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.

Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.

Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.

Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.

* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking

run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.

The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.

* fix(connectors): a failed Try again shows the failure, not the old dead link

The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.

One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.

* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected

`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.

A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.

* fix(connectors): the operation registers under the gateway session key

The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.

The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.

* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed

Three follow-ups from the verification of the fix pass.

The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.

Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.

`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.

`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
2026-09-15 00:41:14 +05:30
Siddharth Balyan d105376b21 Connector code lives in one package, tools/connectors/ (move only; NS-868 prep) (#110368)
* refactor(tools): discovery also scans tools/<pkg>/tool.py

A tool family that is a whole package had no way to register: discovery
globbed tools/*.py only and derived the module name from the filename.
Now the candidate list is tools/*.py plus tools/*/tool.py, merged and
sorted once so import order does not depend on depth (register() lets a
same-name duplicate overwrite silently), and the module name comes from
the path relative to tools/. Only tool.py is scanned inside a package,
so its siblings are libraries by construction. A package without an
__init__.py is skipped with a warning rather than registering from a
checkout and vanishing from the installed wheel.

The AST prefilter and the (mtime, size) disk cache are per absolute path
and work unchanged. The two hand-rolled tools/*.py enumerators in tests
now use the same candidate helper.

* refactor(connectors): one package for the connector domain, tools/connectors/

The connector code was spread across six flat files and a root-level
module that was a sibling of model_tools.py only by address:

  tools/connections_tool.py            -> tools/connectors/tool.py     (schema, register, dispatcher)
                                          tools/connectors/managed.py  (the managed leg, split out)
  tools/connections_tool_mcp.py        -> tools/connectors/mcp.py      (validation split out ->)
                                          tools/connectors/targets.py  (normalize_targets, validate_action)
  tools/connections_tool_operation.py  -> tools/connectors/operation.py
  tools/connector_search.py            -> tools/connectors/search.py
  model_tools_connectors.py            -> tools/connectors/dispatch.py
  tools/tool_gateway/                  -> tools/connectors/gateway/

Move only; every function body is unchanged. tools/connectors/__init__.py
is the door: nine names, the whole cross-package surface. model_tools and
tool_search deep-import a few helpers past it on purpose and the docstring
says so. The two split files make the import graph one-directional
(tool -> mcp -> targets, tool -> managed) where the old layout had
connections_tool importing validation out of the MCP file.

One behaviour-neutral seam change: the _connectors_available try/except
wrapper is gone. connectors_available() already fails closed, and both
the registry handler and the inline executor now read it as a module
attribute (gateway.config.connectors_available), so tests patch it in
one place instead of two. _default_client lives in managed.py, the only
module that calls it.

tools/managed_tool_gateway.py and tools/managed_gateway_auth.py stay:
they are gateway identity shared by tts, transcription, image and modal.

Test files follow their modules. No docs referenced the old paths; no
compat pointer is added (in-tree moves get none).

* ci: retrigger (zero-job dispatch on 142466de6b)
2026-09-15 00:41:13 +05:30
Siddharth Balyan e0ef0eb9c3 manage_connections covers local MCP servers; setup_mcp leaves the schema (NS-867, PR1) (#109517)
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema

One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.

MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.

Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.

`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.

Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.

`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.

The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.

* wip(desktop): connection.request store, resume restore, card routing for MCP targets

Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.

* fix(config): hermes update turns on the connections toolset for saved toolset lists

`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.

Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.

`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.

* refactor: anti-slop pass on the desktop slice; shorten added comments

Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.

* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed

The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.

session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.

lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.

vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.

* chore: drop __pycache__ files swept in by an over-broad git add

* fix(desktop): correlate the connection.request row with the model's tool call by reason

The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.

* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds

* fix(connections): settle reason derives from target state, never from the renderer

A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.

* fix(desktop): a pending connection card re-arms on resume and activate

The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.

* style: literal wording in added comments, docstrings and docs

* fix: shared gateway-event contract and config-schema category for the connection events

connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.

* style: import order (perfectionist) in the desktop and shared files this PR touches

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-15 00:41:13 +05:30
kshitijk4poor afeaabf629 refactor(agent_init): reuse anthropic_adapter._TOOL_STREAMING_BETA
_FINE_GRAINED_BETA duplicated the identical beta string; import the
adapter constant lazily (same direction agent_init already imports
build_anthropic_client). Header output byte-identical.
2026-09-14 21:13:54 +05:30
kshitijk4poor a19160bbf0 refactor(agent): one outbound-kwargs sanitizer seam for the main loop and the summary
The summary path had grown a verbatim copy of turn_api_request's 3-line
surrogate/ASCII chokepoint — the same drift class this PR removes for the
hand-rolled kwargs builder. Move the two lines and the #50959 rationale into
`message_sanitization.sanitize_outbound_kwargs` and call it from both sites,
so the next sanitizer step added to the main loop cannot miss the summary.
Tighten two comments: "same kwargs builder" (cache_control redecoration is
not re-applied here) and a `_summary_text` note that is true for all three
summary branches, not just the chat one.
2026-09-14 20:35:28 +05:30
kshitijk4poor 8c6cb6f995 refactor(agent): drop the Optional import orphaned by the LM Studio helper removal 2026-09-14 20:35:28 +05:30
kshitijk4poor f4442e44b9 refactor(agent): keep _summary_text; the tool-call log is the only new behaviour
`_summary_text_with_scrub` was the old helper plus a warning — `.content`
never carried tool calls, so nothing was being discarded that was not
already ignored. Restore the original name, keep the warning with the WHY
(the request now carries tools, this path never executes a call), and make
the tool-only test assert the log line so it actually binds the new code
path instead of passing on the pre-fix retry behaviour. Trim the builder
comment to the WHY.
2026-09-14 20:35:28 +05:30