The rename sweep in the base commit missed the sibling-test blast radius
(18 red files on CI). Three classes, all fixed:
1. Stale old names in tests (todo/cronjob/process/tour/tip) — updated to
todo_list/cronjob_manage/process_manage/gui_tour/show_tip at every
registry.get_entry/dispatch/coerce/preview/allowlist call site, plus
the coding-brief sentence in agent/coding_context.py now names
todo_list (and its gating test).
2. Missed rename in production: AGENT_RUNTIME_POST_HOOK_TOOL_NAMES still
held 'tour' — post-hook ownership would have double-emitted for
gui_tour via the bridge path.
3. Tests pinning pre-deferral assembly (blank-slate surface, modal
sandbox resolution, desktop diet, HUD note) now pin their ACTUAL
contract under the legacy defer:[] override, or assert on granted
tool names instead of visible schemas.
Also fixes a pre-existing ordering flake surfaced by the sweep:
test_holds_exactly_the_gui_affordances depended on whether an earlier
test had imported apply_layout_tool (registry-registered, not in the
static desktop_ui list) — now forces discovery and pins the full set.
649 tests green locally across all touched files, both orderings.
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
open_preview and read_preview could drive the page, but nothing could close
the pane. Same desktop_ui / session-source gate as the rest of the GUI tools.
Ten behavior tests for target discovery, stable-selector ordering, step
paging, recovery hints, and the self-containment contract the preview
injection depends on. Docs cover data-tour markup and the curated-tour
entry point alongside the tool itself.
The toolset inventory and the post-hook ownership contract both
enumerate the GUI tools; the new tool joins both lists (and the
emit-once parametrization actually exercises its executor path).
The desktop_ui and post-hook ownership contract tests enumerate their tool
sets exactly — add read_window_below to both (plus the executor-path
parametrize case). Lint: sorted type import, explicit GetWindowsModule type
instead of an import() annotation, curly + blank-line style.
The pane, in-app browser, and reaction tools were gated on HERMES_DESKTOP=1 —
an env var set only on backends Electron spawns itself (local and SSH). A
desktop client connected to a plain URL gateway or Hermes Cloud lost all six:
they were stripped from the schema before the model saw them, on the same
backend whose platform hint was telling it "you are chatting inside the Hermes
desktop app". open_preview, read_preview, read_terminal, close_terminal,
focus_pane, and react_to_message were all silently absent.
The client is not the host. Capability now resolves from the session's own
source, which session.create already carries:
- The six tools move into a `desktop_ui` toolset, off _HERMES_CORE_TOOLS so no
other platform pays their schema.
- _gui_surface_toolsets(platform) folds `desktop_ui` (and the existing
`project` tools) into the GUI gateway's resolution when the session's
platform is the desktop app — the same answer on every topology.
- check_fn drops the env probe. It kept the one thing that is genuinely a
per-process/user fact: react_to_message's display.message_reactions opt-in,
which the desktop mirrors onto whichever gateway it is connected to.
react_to_message was doubly broken: it read that toggle behind the env gate, so
even a local-backend user's Settings toggle could not reach a remote session.
The embedded terminal pane keeps working correctly the other way round: it runs
`hermes --tui` against a desktop-spawned backend, and a tui-sourced session
gets no GUI tools even though HERMES_DESKTOP=1 is set on that process.