Files
hermes-agent/tui_gateway/agent_callbacks.py
T
Siddharth Balyan ee2f5629b8 Desktop connect runs on the connection operation: one card, no link to the model, no renderer polling (NS-868) (#110574)
* refactor(connectors): cut comments that restate the code

Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.

* feat(connectors): managed connect runs on the connection operation

Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).

Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.

What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
  transition table. `operation.transition()` enforces it; a card cannot claim a managed
  target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
  reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
  settle -> result). The managed `observe` hook polls the gateway list once per tick for the
  whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
  shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
  are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
  the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
  docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
  added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
  is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
  `connect_card_available` instead of the link when the session platform is `desktop`.

Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.

* feat(desktop): connector card subscribes to the connection operation

The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.

- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
  `applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
  entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
  is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
  snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
  `Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
  failed / expired reissues through `connectors.connect` on the open operation, Not now is a
  per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
  with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
  `HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
  backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
  and compiles unchanged; PR3 moves it onto the operation.

anti-slop: no net-new findings (17 touched files vs 11d1a12472).

* fix(connectors): the card never parks the tool thread; every update carries the snapshot

Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).

- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
  the tool thread on a private request-id Event until a `_respond` that no longer exists for
  this event. `connection.respond` settled the operation but the tool waited its full deadline
  before the watcher loop even started. The callback now only emits the card; the operation's
  own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
  through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
  (state, link, detail). The initial mint happened before the card existed, so the renderer
  never saw the links and Connect stayed disabled; a Continue settlement stamped
  `not_connected` on the backend while the card still showed `initiated`. The store now
  overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
  once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
  `_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
  and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
  `interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
  (the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.

tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.

* docs(connectors): prompts and docs describe the operation, not the deleted wait verb

The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.

* fix(connectors): the panel re-mints only a dead link

Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.

* test(connectors): the local-batch test answers the operation the way the card does

The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.

* ci: retrigger

* fix(connectors): the desktop card appears outside guided onboarding

Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.

The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").

The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.

ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.

message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.

* style(connectors): shorter comments, no module mock in the card router test

The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.

anti-slop: no net-new findings (25 touched files)

* fix(connectors): Connect on a waiting row opens the stored link

ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.

The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.

* fix(connectors): a settled card stays dead; the card binds to its tool call only

A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.

The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.

`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.

`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.

Sid's rule of record: a resolved card is fully dead; no path brings it back.

* fix(connectors): the watch loop settles once, on time, and never raises into the result

Three findings from the live review, one loop.

Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.

Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.

Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.

Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.

* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking

run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.

The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.

* fix(connectors): a failed Try again shows the failure, not the old dead link

The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.

One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.

* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected

`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.

A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.

* fix(connectors): the operation registers under the gateway session key

The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.

The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.

* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed

Three follow-ups from the verification of the fix pass.

The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.

Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.

`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.

`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
2026-09-15 00:41:14 +05:30

476 lines
24 KiB
Python

"""Agent callback wiring: child-session live mirror, per-session agent callbacks, personality
overlay, background/preview agent kwargs, agent reset. Bodies are rebound onto server.py's
globals at install time (method_ctx.bind_module), so they reference server.py globals bare."""
from __future__ import annotations
import json
import contextlib
import threading
from .method_ctx import bind_module
# Child-session live mirror: a delegated child's activity reaches the gateway only as
# relayed ``subagent.*`` events on the PARENT sid; translate them into native stream
# events on the CHILD sid (write_json routes by sid) so its own window is not silent.
_child_mirrors: dict[str, dict] = {}
_child_mirrors_lock = threading.Lock()
# Child sids with a run in flight (refreshed per relayed event, popped on complete) so a
# lazy watch resume reports running=true during a silent long tool.
_active_child_runs: dict[str, float] = {}
# Anything quiet this long lost its completion event — don't pin "running".
_CHILD_RUN_STALE_S = 3600.0
_CHILD_DELTA_EVENTS = {"subagent.thinking": "reasoning.delta", "subagent.text": "message.delta",
"subagent.start": "message.delta"}
def _child_run_active(child_key: str) -> bool:
ts = _active_child_runs.get(child_key)
return ts is not None and (time.time() - ts) < _CHILD_RUN_STALE_S
def _mirror_subagent_to_child(event_type: str, payload: dict) -> None:
child_key = str(payload.get("child_session_id") or "")
if not child_key:
return
# Liveness registry first: accurate with no window open (one opened mid-run knows busy).
if event_type == "subagent.complete":
_active_child_runs.pop(child_key, None)
else:
_active_child_runs[child_key] = time.time()
# Mirror only into a live watch session NOT upgraded to a full agent (an upgraded one owns
# a real native stream). Either way drop state so a reopened window starts fresh.
live = _find_live_session_by_key(child_key)
if live is None or live[1].get("agent") is not None:
with _child_mirrors_lock:
_child_mirrors.pop(child_key, None)
return
csid = live[0]
text = str(payload.get("text") or "")
with _child_mirrors_lock:
st = _child_mirrors.setdefault(child_key, {"seq": 0, "open_tool": None, "started": False})
if not st["started"]:
st["started"] = True
_emit("message.start", csid)
# thinking/text/start (the child's goal, as a one-time header) are plain deltas.
if event_type in _CHILD_DELTA_EVENTS:
if text:
_emit(_CHILD_DELTA_EVENTS[event_type], csid,
{"text": f"{text}\n" if event_type == "subagent.start" else text})
return
if event_type not in ("subagent.tool", "subagent.complete"):
return
if st["open_tool"]:
_emit("tool.complete", csid, st["open_tool"])
if event_type == "subagent.tool":
st["seq"] += 1
tool = {"name": str(payload.get("tool_name") or "tool"),
"tool_id": f"submirror:{child_key}:{st['seq']}", "args": {}}
if preview := str(payload.get("tool_preview") or payload.get("text") or ""):
tool["preview"] = preview
st["open_tool"] = tool
_emit("tool.start", csid, tool)
else:
summary = str(payload.get("summary") or payload.get("text") or "")
_emit("message.complete", csid, {"text": summary})
_child_mirrors.pop(child_key, None)
def _agent_cbs(sid: str) -> dict:
def _read_block(method: str, timeout: int):
# read_terminal / read_preview (desktop GUI): server request like clarify; the preview
# read gets longer since a URL tab extracts text from a live page.
return lambda start=None, count=None: _ask(
method, sid, {k: v for k, v in (("start", start), ("count", count)) if v is not None},
timeout=timeout)
callbacks = {
"tool_start_callback": lambda tc_id, name, args: _on_tool_start(sid, tc_id, name, args),
"tool_complete_callback": lambda tc_id, name, args, result: _on_tool_complete(sid, tc_id, name, args, result),
"tool_progress_callback": lambda event_type, name=None, preview=None, args=None, **kwargs: _on_tool_progress(
sid, event_type, name, preview, args, **kwargs),
"tool_gen_callback": lambda name: _tool_progress_enabled(sid) and _emit("tool.generating", sid, {"name": name}),
"thinking_callback": lambda text: _emit("thinking.delta", sid, {"text": text}),
# Affection reaction (ily / <3 / good bot) → hearts; core-detected so TUI/desktop share it.
"reaction_callback": lambda kind: _emit("reaction", sid, {"kind": kind}),
"reasoning_callback": lambda text: _emit(
"reasoning.delta", sid, {"text": text, **({"verbose": True} if _session_verbose(sid) else {})}),
"status_callback": lambda kind, text=None: _status_update(sid, str(kind), None if text is None else str(text)),
# Credits/notice spine: AgentNotice → notification.show; recovery → notification.clear.
"notice_callback": lambda n: _emit(
"notification.show", sid,
{"text": n.text, "level": n.level, "kind": n.kind, "ttl_ms": n.ttl_ms, "key": n.key, "id": n.id}),
"notice_clear_callback": lambda key: _emit("notification.clear", sid, {"key": key}),
"clarify_callback": lambda q, c, multi_select=False, questions=None: (
_clarify_block(sid, q, c, multi_select=multi_select, questions=questions)),
"read_terminal_callback": _read_block("terminal.read", 30),
"read_preview_callback": _read_block("preview.read", 45),
# drive_preview / annotate_preview (desktop GUI): same budget as the preview read it ends with.
"drive_preview_callback": lambda payload: _ask("preview.act", sid, dict(payload), timeout=45),
# read_window_below (desktop GUI): main process enumerates native windows.
"read_window_below_callback": lambda: _ask("window.read", sid, {}, timeout=30),
# manage_connections card. Fire-and-forget: the tool thread waits on its own operation
# (tools/connectors/run.py), and the card drives it through connection.respond by op_id.
"connection_callback": lambda payload: _emit("connection.request", sid, dict(payload)) and None,
# tour (desktop GUI): renderer drives driver.js and answers the ``tour`` request.
"tour_callback": lambda payload: _tour_request(sid, payload)}
# Interim assistant commentary (text alongside tool calls), gated on display.interim_assistant_
# messages; _run_prompt_submit overwrites it per turn and clears it so a stale closure can't fire.
if _load_interim_assistant_messages():
callbacks["interim_assistant_callback"] = lambda text, *, already_streamed=False: _emit(
"message.interim", sid, {"text": str(text), "already_streamed": bool(already_streamed)})
return callbacks
def _apply_project_workspace(task_id: str, path: str, _name: str = "") -> None:
"""Intentional workspace move from the project_* tools: re-anchor the live session's cwd
and push session.info. The ONLY auto-cwd path — an explicit tool call, never a `cd`."""
if not path:
return
# task_id is the durable session_key; _sessions (and desktop event routing) key by sid.
key = str(task_id or "")
with _sessions_lock:
sid, session = (key, _sessions[key]) if key in _sessions else next(
((s, c) for s, c in _sessions.items()
if c.get("session_key") == key or getattr(c.get("agent"), "session_id", None) == key),
("", None))
resolved = os.path.abspath(os.path.expanduser(str(path)))
if session is None or not os.path.isdir(resolved):
return
# explicit switch supersedes a settle-adopted cwd
session.update(cwd=resolved, explicit_cwd=True, cwd_from_settle=False)
_register_session_cwd(session)
_persist_session_cwd_and_schedule_git_meta(session, resolved)
try:
agent = session.get("agent")
info = _session_info(agent, session) if agent is not None else {
"cwd": resolved, "branch": git_probe.branch(resolved),
"project": _project_info_for_cwd(resolved), "lazy": True}
_emit("session.info", sid, info)
except Exception:
logger.debug("failed to emit session.info after project workspace move", exc_info=True)
def _wire_callbacks(sid: str):
from tools.terminal_tool import set_sudo_password_callback
from tools.skills_tool import set_secret_capture_callback
from tools.project_tools import set_project_workspace_callback
def secret_cb(env_var, prompt, metadata=None):
pl = {"prompt": prompt, "env_var": env_var, **({"metadata": metadata} if metadata else {})}
val = _ask("secret", sid, pl)
if not val:
return {"success": True, "stored_as": env_var, "validated": False, "skipped": True, "message": "skipped"}
from hermes_cli.config import save_env_value_secure
return {**save_env_value_secure(env_var, val), "skipped": False, "message": "ok"}
set_sudo_password_callback(lambda: _ask("sudo", sid, {}, timeout=120))
set_project_workspace_callback(_apply_project_workspace)
set_secret_capture_callback(secret_cb)
# External password-manager unlock: the renderer shows a masked master-password card; the
# answer is consumed by the manager CLI on stdin and only a session token stays in memory.
from agent.vault_backends.unlock import (set_code_prompt_callback, set_current_session_id,
set_save_login_prompt_callback, set_unlock_prompt_callback)
set_current_session_id(sid) # an unlock made on this turn belongs to this session (released with it)
set_unlock_prompt_callback(lambda backend, display_name: _ask(
"vault.unlock_prompt", sid, {"backend": backend, "display_name": display_name}, timeout=120))
def save_login_cb(origin, site):
# The renderer shows identifier + masked password; the JSON answer goes straight to the vault store.
raw = _ask("vault.save_login", sid, {"origin": origin, "site": site}, timeout=180)
try:
data = json.loads(raw) if raw else None
except ValueError:
return None
return data if isinstance(data, dict) and data.get("password") else None
set_save_login_prompt_callback(save_login_cb)
set_code_prompt_callback(lambda site, hint: _ask(
"vault.code", sid, {"site": site, "hint": hint}, timeout=180))
def _available_personalities(cfg: dict | None = None) -> dict:
"""Built-ins + user overrides, via hermes_cli.personality (single owner)."""
from hermes_cli.personality import available_personalities
return available_personalities(_load_cfg() if cfg is None else cfg)
def _validate_personality(value: str, cfg: dict | None = None) -> tuple[str, str]:
"""(name, prompt) for a requested personality or ValueError; like resolve_personality but
via the module-level _available_personalities so tests keep a single patch point."""
from hermes_cli.personality import normalize_personality_name, render_personality_prompt
if not (name := normalize_personality_name(value)):
return "", ""
personalities = _available_personalities(cfg)
if name not in personalities:
names = ", ".join(f"`{n}`" for n in sorted(personalities))
raise ValueError(f"Unknown personality: `{str(value).strip()}`.\n\nAvailable: `none`, {names}")
return name, render_personality_prompt(personalities[name])
def _prompt_text(value) -> str:
"""Normalize config prompt values from YAML for AIAgent (hermes_cli.personality owns this)."""
from hermes_cli.personality import prompt_text
return prompt_text(value)
def _apply_personality_to_session(
sid: str, session: dict, new_prompt: str, personality: str = "") -> tuple[bool, dict | None]:
"""Apply a personality change without resetting history: the ephemeral system prompt is
updated in place (appended at API-call time, so prompt-cache hits survive) plus a pivot
marker so the model stops pattern-matching its earlier tone. Returns (False, info)."""
if not session:
return False, None
session["personality"] = personality
if not (agent := session.get("agent")):
return False, None
agent.ephemeral_system_prompt = new_prompt or None
marker = (
"[System: The user has changed the assistant's personality. "
"From this point forward, adopt the following persona and respond "
f"accordingly: {new_prompt}]"
if new_prompt else
"[System: The user has cleared the personality overlay. "
"From this point forward, respond in your normal default style.]")
# Like the model-switch marker: role=user so strict providers accept it mid-conversation,
# but `display_kind` keeps it out of the `truncate_before_user_ordinal` addressing space
# (untagged, every rewind would land one turn early and hard-delete the difference).
# Untagged, it counts as a real user turn on the gateway side while no client counts it, so every later
# rewind resolves one turn too early and `replace_messages` hard-deletes the difference (#82756).
with session["history_lock"]:
session["history"].append({"role": "user", "content": marker, "display_kind": "personality_switch"})
session["history_version"] = int(session.get("history_version", 0)) + 1
info = _session_info(agent)
_emit("session.info", sid, info)
return False, info
def _cfg_max_turns(cfg: dict, default: int) -> int:
from hermes_cli.config import resolve_turn_limit as _resolve_turn_limit
# Env override wins; resolve_turn_limit makes "none"/"unlimited"/0 first-class spellings.
if env_val := os.environ.get("HERMES_TUI_MAX_TURNS"):
return _resolve_turn_limit(env_val, default=default)
raw = (cfg.get("agent") or {}).get("max_turns")
if raw is None:
raw = cfg.get("max_turns")
return default if raw is None else _resolve_turn_limit(raw, default=default)
def _parse_tui_skills_env() -> list[str]:
raw = os.environ.get("HERMES_TUI_SKILLS", "")
return list(dict.fromkeys(p.strip() for p in raw.replace("\n", ",").split(",") if p.strip()))
def _load_fallback_model():
"""Configured fallback chain via the shared ``get_fallback_chain`` (parity with
HermesCLI/gateway: ``fallback_providers`` first, legacy ``fallback_model`` merged after)."""
from hermes_cli.fallback_config import get_fallback_chain
return get_fallback_chain(_load_cfg())
def _background_agent_kwargs(agent, task_id: str) -> dict:
cfg = _load_cfg()
def g(name, default=None):
return getattr(agent, name, default)
# Don't rehydrate a deliberately empty fallback chain.
if hasattr(agent, "_fallback_chain"):
fallback = agent._fallback_chain or []
else:
fallback = (agent._fallback_model if hasattr(agent, "_fallback_model")
else _load_fallback_model())
# Detached tasks declare platform="tui" (no UI sid for renderer-routed events), so resolve
# toolsets against it — never GUI schema they can't use.
return {
**{k: g(k) or None for k in ("base_url", "api_key", "provider", "api_mode", "acp_command",
"acp_args", "ephemeral_system_prompt")},
**{k: g(k) for k in ("providers_allowed", "providers_ignored", "providers_order", "provider_sort",
"provider_data_collection", "openrouter_min_coding_score")},
"model": g("model") or _resolve_model(), "max_iterations": _cfg_max_turns(cfg, 25),
"enabled_toolsets": g("enabled_toolsets") or _load_enabled_toolsets("tui"),
"quiet_mode": True, "verbose_logging": False,
"provider_require_parameters": g("provider_require_parameters", False), "session_id": task_id,
"reasoning_config": g("reasoning_config") or _load_reasoning_config(str(g("model", "") or "")),
"service_tier": g("service_tier") or _load_service_tier(),
"request_overrides": dict(g("request_overrides", {}) or {}),
# The side agent persists into the PARENT's store: a named-profile chat's ``bg_*`` rows
# belong to that profile's state.db, not the launch handle.
"platform": "tui", "session_db": getattr(agent, "_session_db", None) or _get_db(), "fallback_model": fallback}
def _ephemeral_preview_agent_kwargs(agent, task_id: str) -> dict:
return {**_background_agent_kwargs(agent, task_id),
"enabled_toolsets": ["terminal", "file"], "session_db": None, "skip_memory": True}
@contextlib.contextmanager
def _side_agent_session_db(parent_db):
"""A side agent's OWN registry reference on the parent's store for the duration of its turn.
Handing the parent's object across is not enough: the parent releases its reference from
``AIAgent.close()`` / a session reset, and when it was the last holder the registry tears the
connection down under the still-running background turn (the delegated-child path acquires
the same way, ``tools/delegate_tool._open_child_session_db``). Released on exit."""
path = getattr(parent_db, "db_path", None)
if parent_db is None or path is None:
yield parent_db
return
from hermes_state_registry import acquire, release_or_close
db = acquire(path)
try:
yield db
finally:
release_or_close(db)
def _preview_restart_history(session: dict, max_messages: int = 24, max_tool_chars: int = 1200) -> list[dict]:
"""Distill recent parent history for the ephemeral preview-restart agent (else it guesses
app/cwd/port from the bare URL): last ``max_messages`` back to the last user turn, tool
results truncated to ``max_tool_chars``."""
try:
with session["history_lock"]:
history = list(session.get("history") or [])
except Exception:
history = list(session.get("history") or [])
if not history:
return []
last_user = next((i for i in range(len(history) - 1, -1, -1) if history[i].get("role") == "user"), None)
start = max(0, len(history) - max_messages)
if last_user is not None:
start = min(start, last_user)
trimmed: list[dict] = []
for msg in history[start:]:
if not isinstance(msg, dict) or msg.get("role") not in ("user", "assistant", "tool", "system"):
continue
copy = {k: v for k, v in msg.items() if k != "reasoning"}
content = copy.get("content")
if msg.get("role") == "tool" and isinstance(content, str) and len(content) > max_tool_chars:
copy["content"] = content[:max_tool_chars] + f"\n... (truncated, original {len(content)} chars)"
trimmed.append(copy)
return trimmed
def _preview_tool_result_preview(name: str, result: str) -> str:
try:
data = json.loads(result)
except Exception:
data = None
if not isinstance(data, dict):
return ""
if name == "terminal":
if output := str(data.get("output") or "").strip():
return output[-1200:]
if data.get("session_id"):
return f"Background process started: {data.get('session_id')}"
if data.get("exit_code") is not None:
return f"terminal exited with code {data.get('exit_code')}"
return str(data.get("error") or "").strip()[:1200]
def _preview_restart_callbacks(parent: str, task_id: str) -> dict:
started_at: dict[str, float] = {}
def progress(message: str, level: str = "info") -> None:
if text := str(message or "").strip():
_emit("preview.restart.progress", parent, {"task_id": task_id, "level": level, "text": text})
def tool_start(tool_call_id: str, name: str, args: dict) -> None:
started_at[tool_call_id] = time.time()
ctx = _tool_ctx(name, args)
progress(f"Running {name}{f': {ctx}' if ctx else ''}")
def tool_complete(tool_call_id: str, name: str, _args: dict, result: str) -> None:
duration_s = time.time() - started_at.get(tool_call_id, time.time())
summary = _tool_summary(name, result, duration_s) or f"Finished {name}{f' in {_fmt_tool_duration(duration_s)}' if duration_s else ''}"
output = _preview_tool_result_preview(name, result)
progress(summary + (f"\n{output}" if output else ""))
def tool_progress(event_type: str, name: str | None = None, preview: str | None = None, **_kwargs) -> None:
if preview or name:
progress(str(preview) if preview else f"{event_type.replace('.', ' ')}: {name}")
return {
"tool_start_callback": tool_start, "tool_complete_callback": tool_complete,
"tool_progress_callback": tool_progress,
"tool_gen_callback": lambda name: progress(f"Preparing {name}"),
"status_callback": lambda kind, text=None: progress(text if text is not None else kind)}
def _rebuild_session_agent(sid: str, session: dict, **kwargs):
"""Prepare and install a replacement on the session's profile, then transfer DB ownership.
An unscoped _make_agent defaults to the launch store: named-profile Bot Chat turns then disappear
from the profile's replay even though they were successfully written to another database (#104079).
"""
old_agent = session.get("agent")
profile_home = session.get("profile_home")
session_db = getattr(old_agent, "_session_db", None)
# No live agent to inherit from (rebuild before the deferred build ran): open the profile's store the
# same FAIL-CLOSED way _start_agent_build does rather than letting _make_agent reach for the launch db.
opened = session_db is None and bool(profile_home)
scopes = _bind_build_profile_scopes(profile_home) if profile_home else None
try:
# Resolve fallible config before allocating a replacement or moving its handle.
config_model_seen = _config_model_target()
if opened:
session_db = _open_profile_session_db(profile_home)
agent = _make_agent(sid, session["session_key"], session_db=session_db, **kwargs)
except BaseException:
if opened and session_db is not None:
with contextlib.suppress(Exception):
session_db.close()
raise
finally:
if scopes is not None:
_release_build_profile_scopes(scopes)
# Only a DEDICATED handle carries ownership; the shared launch handle outlives every agent and
# _transfer_db_to_agent refuses it.
with _sessions_lock:
session.update(agent=agent, config_model_seen=config_model_seen)
owned = opened or bool(getattr(old_agent, "_owns_session_db", False))
if owned and _transfer_db_to_agent(agent, session_db):
if old_agent is not None:
old_agent._owns_session_db = False
elif opened:
with contextlib.suppress(Exception):
session_db.close()
return agent
def _reset_session_agent(sid: str, session: dict) -> dict:
updates = dict(
attached_images=[], queued_prompt=None,
_queued_prompt_generation=int(session.get("_queued_prompt_generation", 0)) + 1,
edit_snapshots={}, image_counter=0, running=False, show_reasoning=_load_show_reasoning(),
tool_progress_mode=_load_tool_progress_mode(), tool_started_at={})
tokens = _set_session_context(session["session_key"])
try:
# /new is a full conversation boundary: session-scoped runtime overrides (/model,
# /reasoning, /fast) do NOT carry forward and the pins are cleared so a rebuild can't
# resurrect them. Global process state is never touched (see _apply_model_switch).
for k in ("model_override", "create_reasoning_override", "create_service_tier_override", "one_turn_model_restore"):
session.pop(k, None)
new_agent = _rebuild_session_agent(
sid, session, session_id=session["session_key"],
platform_override=_session_source(session),
context_cwd_is_launch_artifact=_context_cwd_is_launch_artifact(session))
finally:
_clear_session_context(tokens)
session.update(updates)
session.pop("queued_prompts", None)
with session["history_lock"]:
session["history"] = []
session["history_version"] = int(session.get("history_version", 0)) + 1
info = _session_info(new_agent, session)
_emit("session.info", sid, info)
_restart_slash_worker(sid, session)
return info
def register(server) -> None:
"""Publish this module's helpers + handlers onto ``server``, rebound to its globals."""
bind_module(globals(), server, skip=("_",))