The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.
`tui_gateway/server_requests.py` owns the one mechanism:
send() block the agent thread until the response frame
(`srq-<n>` ids; ints belong to the client)
send_async() fire-and-callback variant (bot relay)
cancel*() withdraw with ONE `request.cancel {id, method, reason}`
event (timeout / interrupt / process exit /
answered elsewhere) instead of per-kind *.expire
open_requests() the still-open frames, replayed by session.resume,
session.activate and session.events.since so a
reconnecting client re-renders every kind, not two
clarify.lock stays a real client→server RPC (locks one batch
answer early); locked answers merge into the final
set even when the closing response carries only the
tail the user answered last
A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.
Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.
Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.
- gateway/platforms/webhook.py: cron_job route mode — after the same
HMAC auth / rate limit / filters / script / idempotency as agent
routes, the rendered prompt becomes transient per-run context and the
job fires through execute_job_for_event on a worker thread (202
Accepted immediately). Startup validation rejects cron_job +
deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
the shared claimed-run body (_execute_job_now), so event fires share
at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
subscriptions; job ref validated (and canonicalized to the job ID) at
create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
CLI tests in test_webhook_cli.py.
The shell `find ~/.hermes/skills ~/.hermes/hermes-agent/optional-skills ...`
snippet hardcoded a dev-clone path that does not exist on user installs; the
loader already substitutes ${HERMES_SKILL_DIR} (skills.template_vars), which
is what every other bundled skill uses. Drop the upstream-agent path mention.
Regenerated the docs page so it matches SKILL.md (also fixes the comfyui
related-skill link, which now lives under optional).
Optional skill: scroll-as-timeline landing pages on a deterministic
CSS/JS engine, with interview → page grammar → signature move workflow
and screenshot-based scroll verification. Engine and scripts vendored
verbatim; asset generation re-anchored on image_generate with the
upstream kie.ai flow kept as an optional path.
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.
Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.
Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.
Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).
A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.
Docs: relevance-floor bullet added to tool-search.md implementation
details.
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro
(fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key,
aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and
negative_prompt real, no seed/resolution keys per the published llms.txt
schemas. Payload shapes pinned in tests; docs mention updated.
omo's omo-agent-toolkit worktree-sweep (their PR #7151) added three
capabilities our hermes worktree command lacked:
- --json on list and prune: machine-readable audit/result payloads so
scripts and agents can consume verdicts without scraping table output.
- --older-than DAYS: an age floor that only ever RESTRICTS reaping
(young-but-reapable trees are kept); it never widens eligibility, so
the existing safety invariants are untouched.
- External-tree visibility: linked worktrees registered outside
.worktrees/ are now reported read-only in the audit (branch, locked,
missing) instead of being invisible, and registrations whose
directory has vanished are dropped via git worktree prune (metadata
only, no files touched) during prune.
Not ported: omo's ancestor-of-default-branch merge test (our git cherry
patch-equivalence is strictly stronger under rebase/squash merges), and
their hardcoded external-root exclusion list (we exclude by location:
everything outside .worktrees/ is hands-off).
Tests: 9 new contracts in tests/hermes_cli/test_worktree_gc.py (age gate
restrict-only, external trees never reaped, stale-registration prune
dry-run/real, JSON shapes, negative --older-than rejected). Live E2E on
a scratch repo verified all three flags end to end.
Tests collapse 13 change-detectors into 7 invariants (silent below threshold /
without context, fires above, same-model re-select silent, config override and
0-disables, registry threading, agent context derivation). The legacy 5-arg
guard test goes with the TypeError fallback it covered: that fallback was
defence for a case nobody has (every in-tree guard and test double is
*args-tolerant) and would re-run a guard whose real TypeError it masked, so the
rebased port passes the context positionally like every other argument.
Docs no longer claim the confirm fires on the Telegram/Discord pickers or the
dashboard: those surfaces call combined_selection_warning() without a live
agent, so the context-cache guard is (correctly) silent there.
Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).
- hermes_cli/model_selection_guards.py: new context_cache guard +
SelectionContext carrier + selection_context_for_agent() helper;
registry threads live-session facts to guards (6-arg signature with a
TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
Port from openai/codex#39615: the authorization server discovered for an
MCP server can change (protected-resource metadata edit, server
migration, DNS takeover). Without binding, Hermes would send the stored
refresh token to whatever issuer the server now advertises — handing a
long-lived credential to a different authorization server.
- HermesTokenStorage records hermes_issuer alongside cached tokens
(stripped before OAuthToken.model_validate; never sent on the wire).
- Both provider classes (tools/mcp_oauth.py legacy path and
tools/mcp_oauth_manager.py managed path) stamp the discovered issuer
on every token save and enforce the binding on _initialize.
- On mismatch: refresh token is stripped from memory and disk; the
unexpired access token keeps working; full re-auth happens at expiry.
- Legacy token files without an issuer adopt the current one once
(no forced re-login for existing installs — deliberate divergence
from Codex, which requires reauth).
Validated: 12 new tests + 163 existing MCP OAuth tests green; sabotage
run confirms the new tests fail without the enforcement; E2E against
the real manager provider class with a temp HERMES_HOME confirms
mismatch strips and match preserves.
Droid v0.203 (Aug 25 2026) added 'edit queued messages' — a queued steering
message can be pulled back and changed before it is sent. Hermes /queue could
only append blindly: no way to see, fix, drop, or reorder queued prompts.
/queue now supports management subcommands in the CLI:
- /queue — list pending prompts (bare prompt still enqueues)
- /queue list — same
- /queue edit N <p> — replace item N (keeps voice sentinel, #65827)
- /queue rm N — remove item N
- /queue move A B — reorder
- /queue clear — drop everything
- /queue add <p> — force-enqueue prompts starting with a management word
Queue mutations hold queue.Queue's mutex and rebuild unfinished_tasks so
join()/task_done bookkeeping stays consistent. Paste references expand on
enqueue and edit, matching the old inline path.
Reimplementation of PR #18833 by @abhinav11082001-stack (commit was authored
under a fabricated 'Hermes Agent' noreply identity that cannot be carried
into history; engineering credit is theirs), hardened for current main:
voice-sentinel-aware previews/edit, paste-reference expansion, queue
bookkeeping asserts, out-of-range no-op tests, and docs.
Two-direction contract test (tests/website/test_slash_commands_doc_parity.py):
every CommandDef must be documented under its name or an alias, and every
doc table row must resolve to a registered command. Ported from IronClaw's
doc-fact contract tests (nearai/ironclaw#7378), adapted from their clap
--help parser to our COMMAND_REGISTRY single source of truth.
Real drift it caught, fixed here: /loop (alias /proactive) shipped with a
full feature page (user-guide/features/loops.md) and CLI+gateway handlers
but never got a row in the slash-commands reference. Added to both the CLI
Session table and the messaging table, plus the both-surfaces note.
Backend half of the per-task effort control, on today's layout: POST /api/model/set
distinguishes omitted (leave the task's override alone) from explicit null (clear →
inherit) via model_fields_set, canonicalises a level through parse_reasoning_effort
(400 on an unknown one), and "Reset all to main" also drops every override. GET
/api/model/auxiliary returns reasoning_effort per task and the row summary shows it.
The inherit row reads "inherit · main model effort" (own i18n key in all six locales)
rather than reusing the provider's "auto · use main model" copy — the two mean
different things and the reused string read as "use the main model" for the effort.
Runtime already consumes auxiliary.<task>.reasoning_effort (agent/auxiliary_client.py)
and hermes model writes the same key (#110346), so Desktop and CLI now edit one value.
Closes#89259. Salvages #90649 by @higgs1729.
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.
One request now carries a model pick AND its effort on every surface:
- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
<level>` (validated against `parse_reasoning_effort`; unknown level ->
`MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
`ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
high` applies the effort AFTER the agent swap (`switch_model` re-resolves
`reasoning_config` from config.yaml, so an earlier write is clobbered) with the
pick's scope (session; config on `--global`; `--once` snapshots and restores it).
The `/model` picker gains a third stage, "Reasoning effort for <model>", built
from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
`config.set model "X --reasoning high"` applies after the swap; session pin
(`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
one-turn restore carries `reasoning_config`; re-emits `session_info` so the
status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
`<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
`_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
the Copilot-only inline prompt; Copilot keeps its per-model level set via
`github_model_reasoning_efforts`, other routes get the ladder, catalog
`supports_reasoning=False` skips it) plus a "Reasoning effort for the current
model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
with the same step (+ "Provider default"), stored as
`auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
the task list ("openrouter · model · high"), cleared by "Reset all to auto";
tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
it.
Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
"Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
errors "Model names cannot contain spaces"; after switches and `config.get
reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
"reasoning: high", status bar "fable 5.1 high".
The salvaged flag was spelled --profile, which collides with the global
-p/--profile that hermes_cli.main scans BEFORE argparse: `hermes webhook
subscribe x --profile compta` would switch this CLI process to compta's
HERMES_HOME and write the subscription into compta's webhook_subscriptions.json
— a file the default gateway's webhook adapter never reads — while the route
still lacked the profile key. #109020 special-cased the scanner for the webhook
subcommand; naming the flag --route-profile removes the ambiguity without
touching _scan_profile_flag: -p picks the gateway whose subscriptions file is
written, --route-profile picks which /p/<profile>/ prefix may hit the route.
Docs: cli-commands reference row, multi-profile-gateways webhook section, the
route `profile` field. Builds on #109020 (fangliquanflq). Fixes#109016.
Extends the multi-profile MCP paragraph with the identity rules landed in
this branch: OAuth tokens live per profile so each profile's calls run as
its own account; client_cert/client_key are part of the connection
identity; trust and supports_parallel_tool_calls are the consuming
profile's policy even when it shares another profile's connection.
Wording of the per-profile OAuth account guarantee follows the docs draft
in PR #109574 (its token-file fingerprint code path was not taken).
Co-authored-by: ly6751 <99090550+ly6751@users.noreply.github.com>
`main` is a valid profile name (only hermes/default/test/tmp/root/sudo are
reserved), but _session_key_namespace mapped it to `agent:main` — the default
profile's namespace. Both profiles then built byte-identical keys: one routing
entry, one cached agent, and, since 75ae2859b9 pinned default-namespace
keys to the launch store, profiles/main's scoped sessions were written into
the ROOT state.db instead of profiles/main/state.db.
Key the `main` profile as `agent:main~` (`~` is outside the profile-id
alphabet, so the marked form cannot be any other profile's id) and give the
namespace slot one inverse, profile_from_session_key_namespace, used by the
store's key parser, _parse_session_key, the update-marker profile reader and
the profile-delete eviction prefix. Default keys stay byte-identical.
The raw-output watcher modes (all/result/error) and the interim running
update sent the bracketed debug wrapper with the internal process id
(`[Background process proc_… finished with exit code N~ Here's the final
output: …]`) to Telegram/Discord/Slack chats. Reuse the concise one-line
status header for every mode and append the bounded, ANSI-stripped output
tail in a code block; the running update gets the same shape.
Salvaged from #54266 (rebased onto the post-#102117 run_notifications
sibling; the concise mode had landed in between, so the header is shared
rather than reimplemented). Also covers #13122 (ANSI stripping).
Reshape the two salvaged commits onto current main (#109954):
- Move the boundary guard out of the gateway_migrate facade into a new sibling
hermes_cli/gateway_migrate_guards.py as a table of guard functions
(_AUTO_MIGRATION_GUARDS: service domain, UNIX user, HERMES_HOME tree) plus the
identity resolver. The facade grows by ~20 lines only (uid/runtime_home on
ProfileGateway, one seam, the hook wiring).
- Compare uids, not strings: live pid owner via /proc (ps fallback only on
macOS, where /proc does not exist), else the system unit's User= via
_read_systemd_user_from_unit (root when absent), else the home directory's
owner. None means unknown and never blocks.
- The home-tree guard reads the HERMES_HOME the installed unit pins, not the
directory the plan enumerated: that is where the gateway really runs and is
exactly the "stale copies under profiles/" shape from the report.
- When the default is detached, a service-managed secondary is a different
domain for the AUTO path (it must not elect the secondary's manager); the
explicit command keeps electing it as before.
- The explicit command surfaces the same findings as notices (dry run shows
them) and is never blocked by them; only the update hook refuses.
- Rename the opt-out key to gateway.auto_multiplex_migration (nested only, no
top-level alias) and read it before a plan is built, so false prints nothing
and touches nothing. The explicit command ignores it.
- Tests trimmed to the invariants: one parametrized boundary test that exercises
the real hook end to end (refuses, touches nothing, dry run shows the notice),
one "same user / same scope still migrates" control, one opt-out test.
- Docs: boundary table + renamed opt-out section in multi-profile-gateways.md;
one line in hermes_cli/AGENTS.md.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Athena <athena@olympus.local>
`hermes update` folds an eligible multi-profile install onto one multiplexed
gateway on its own, and there is currently no way to say no. The only lever,
`gateway.multiplex_profiles: false`, is also the default: `_read_multiplex_flag`
returns `False` for "absent" and for an explicit `false` alike, so an operator
who has already decided to stay on per-profile gateways has no way to record
that decision. The migration runs again on the next update.
Add `gateway.auto_migrate` (bool, default `true`). Read from the default
profile's config, it gates the automatic path only:
- absent or `true` -> today's behaviour exactly, no change
- `false` -> `maybe_auto_migrate_after_update()` returns before
building a plan; no output, no changes
`hermes gateway migrate --multiplex` is an explicit request and still migrates
regardless of the flag, so it stays the supported way to opt back in.
One early return, one schema entry with the reasoning inline, one invariant
test (opt-out blocks the hook, absent/true do not, explicit command still
applies), one section in the multi-profile gateways guide.
`hermes sessions export --session-id X <dir>/` crashed with IsADirectoryError
because jsonl/html/trace opened the positional as a file while --help called it
an "output path" and md/qmd really do take a directory. An existing directory
(or one spelled with a trailing separator) now receives a default-named file
(`hermes_session_<id>.<fmt>`), and the help text spells out per-format what
OUTPUT means.
The troubleshooting entry told users to raise A2A_REPLY_TIMEOUT for long tasks,
which did nothing against the hardcoded 300s orphan sweep (#106972). Now that the
sweep derives its grace from the reply window and skips tasks with a live waiter,
state that contract next to the variable.
Skipping `tests/`, `spec/`, ... in EXCLUDED_DIRS made those trees
invisible to the guard, but `plugins_loader._load_directory_module`
sets `submodule_search_locations=[plugin_dir]`, so a plugin
`__init__.py` doing `from .tests import evil` imports and runs whatever
lives there: a `tests/evil.py` with a destructive root remove scanned
`dangerous` on main and `safe` on this branch. `_walk` also matched the
names at any depth, so `src/spec/handler.py` — plain runtime code — went
unscanned.
Keep scanning everything; instead cap a critical finding located under a
ROOT-level test dir at `high`, so the verdict is `caution` (confirmation
required, `--force` overridable) rather than the un-overridable
`dangerous`. Fixture strings still cannot brick an install, which was
the reported problem, while a critical in any runtime file (`setup.sh`,
`src/spec/...`) still yields `dangerous`. Trade-off stated in the PR
body: hostile code deliberately placed under `tests/` is now
force-installable rather than blocked outright.
Docs no longer claim test code never runs.
User-visible scanner behaviour changed in this PR (test trees skipped, the block
reason names the critical rule ids), so the plugin docs say so in the same PR.
The served-set paragraph documented create and delete as live operations; rename now
follows the same unroute-before-mutate protocol, so say so where operators look for it.
The multiplex guide now says client_cert/client_key are part of the
"same credentials" test and states the OAuth rule as its own sentence;
the MCP config reference names the per-profile token directory and the
never-shared-across-profiles rule next to the OAuth behaviour list.
Co-authored-by: ly6751 <99090550+ly6751@users.noreply.github.com>
`kanban_request_review(artifacts=[...])` now preserves scratch deliverables
the same way `kanban_complete` does; say so where the scratch-workspace
lifecycle is documented.
Post-merge review of #109502 (gaoanze888) on current main, findings 1-3 and 5-11
(finding 4, the migrate manifest ordering, was already fixed by f9e47aa6fe).
Why:
- `--clone-all` used copytree(symlinks=True); a symlinked source `.env` was then
edited THROUGH the link by the channel strip, deleting the SOURCE's bot token.
Root files the clone edits (.env, config.yaml, auth.json, SOUL.md) are now
materialized as private copies before any write.
- The final profiles/<name> existed during the copy; the multiplexer rescans
profiles/ on create (#109239) and could adopt the half-copied tree and start
adapters on credentials not yet stripped. Clones are built in
profiles/.<name>.staging-<pid> (a leading dot never matches _PROFILE_ID_RE, so
profiles_to_serve never lists it) and published with one os.rename after the
strip; a failed create removes the staging tree.
- The live-multiplexer refusal for --clone-channels lived only in the CLI; REST
(POST /api/profiles) and the TUI (profiles.create) bypassed it. It now lives in
create_profile as profile_channels.clone_channels_refusal, raising ValueError
which every surface already maps to a 400 / 4062. --clone-channels without a
clone flag is an error instead of a silent no-op.
- Channel inventory is ownership-based, evaluated in the SOURCE profile's plugin
scope: private platform plugins under <source>/plugins/ contribute their keys
(previously discovered under the ambient HERMES_HOME); GATEWAY_ALLOW_ALL_USERS /
GATEWAY_ALLOWED_USERS and GATEWAY_RELAY_ID/SECRET/DELIVERY_KEY are channel
settings; alias prefixes SUPPLEMENT the canonical <PLATFORM>_ prefix
(WECOM_DM_POLICY, SMS_WEBHOOK_PORT now stripped). Prefixes shared with tools
(HASS_, TWILIO_, EMAIL_) are stripped only when the source's gateway would run
that adapter (enabled in config, or complete credentials and not explicitly
disabled); their allowlist/port keys are always channel-only.
- --clone-all state removal handled files only; Google Chat's directory-shaped
google_chat_user_tokens/ survived. Directories are removed too, without
following a copied symlink.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
The cherry-picked commit added an `event_silence` probe dimension stamped from
`on_socket_raw_receive`. Two verified problems make it a regression rather than a fix:
- discord.py 2.7.1 dispatches `socket_raw_receive` only when the client is built with
`enable_debug_events=True` (client.py:330, gateway.py:410-412; the default
`log_receive` is a no-op). The adapter never sets it, so the stamp only ever moves at
`on_ready` and every healthy connection reads `event_silence` 300s later — a forced
reconnect every ~5 min. Live-verified against a real `commands.Bot` +
`DiscordWebSocket.received_message`: 6 frames delivered, stamp unchanged, probe unhealthy.
- discord.py already keeps a per-frame clock (`KeepAliveHandler._last_recv`) and closes the
socket itself after `heartbeat_timeout` without frames; and because ACKs are frames,
`ack_stale` (60s) always trips before `event_silence` (300s). A raw-frame stamp cannot
detect the "ESTAB + ACKing + zero events" incident by construction.
Kept and tightened the warning half: bool values (`float(True) == 1.0` silently enabled a
knob at 1s), negative ints, and unparsable strings now warn; an explicit `0` is the documented
opt-out and stays silent. Tests trimmed to the two invariant contracts (warn / don't warn),
proven red on origin/main. Docs updated to match.
The Gateway WS health probe sampled only transport state — ready, open,
heartbeat-ACK age, latency. A socket that stays ESTAB and keeps ACKing while
zero gateway frames arrive (the #109521 "connected-but-deaf" incident) read
healthy indefinitely, and the adapter went silent for hours with no log line
and no watchdog firing.
Two defects fixed:
1. Dispatch-side dimension. `on_socket_raw_receive` now stamps
`_last_gateway_frame_at` for every inbound raw gateway frame — heartbeats
and ACKs included, so a legitimately quiet server is not flagged. The
health check gains an `event_silence` reason with its own bound,
`websocket_event_max_silence_seconds` (default 300s, 0 disables the
dimension alone). The stamp resets on `on_ready` so a reconnect never
inherits pre-restart silence. Trip path is unchanged: consecutive
failures -> retryable `discord_websocket_health_stale` -> the existing
reconnect watcher builds a fresh adapter.
2. Silent probe disable. `_finite_positive_config_float` / `_config_int`
mapped anything `float()` rejects ("15s", "nan", "true") to 0.0 with no
log line, permanently disabling the watchdog invisibly. Unparsable and
non-positive values now log one WARNING naming the knob and raw value.
The new key rides the existing `_YAML_WEBSOCKET_LIVENESS_KEYS` seeding and is
documented in the Discord guide's liveness section.
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which
changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb →
#3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR
body said no preset VALUE changed. The ladder is now the desktop's original
algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to
1.0001, re-mixed from the source colour — with `step` as a parameter. The
only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so
the terminal palette is byte-identical too.
Test: apps/desktop context.test.tsx iterates every builtin preset × mode,
paints it through ThemeProvider and asserts --dt-primary-solid equals the
value a reference copy of the old desktop algorithm computes. Sabotage
(default step 0.05): 11/30 rows fail. Docs: the SDK table now lists
contrastRatio as `number | null` under sRGB measures, not OKLCH.
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.
The composer showed "<model> · Med" in one truncating pill and the only way
to change the effort was to open the model menu, find the active model's
row, and hover it for the per-row options submenu. Users read the pill as
"this model is medium only" and never found the submenu.
- New `ReasoningPill` next to the model pill: shows the active model's live
effort (session value, else the profile default) and opens the same
Thinking / Fast / Effort rows the catalog submenu offers, for the active
model only. Hidden when the catalog reports `reasoning: false`; stays
while capabilities are unknown so it never flickers during the fetch.
Folds away with the model pill in the compact composer stages.
- `useModelMenuController` (shell sibling) now owns the session write /
preset / optimistic-store / rollback logic that lived inside
`ModelMenuPanel`; the model menu and the new `ReasoningMenuPanel` share it
so an edit from either surface is one code path. Tiles get their own
pill bound to their SessionView, primary or tile — never the globals.
- `ModelOptionsContent` (the submenu body) is exported container-free so
the pill's top-level menu renders it without a Radix Sub wrapper.
- The model pill drops the effort suffix (`formatModelPillLabel`: name +
Fast); `formatModelStatusLabel` had no other caller and is removed.
- `currentModelCapabilities()` in lib/model-options resolves the active
pick's caps through `catalogProviderMatches` (aliases, custom slugs).
Live (headless Electron + worktree `hermes serve`, CDP): before — one
pill "Deepseek V4 Flash · Low", no effort control; after — "Deepseek V4
Flash" + "Low" pill; pick High → `config.get reasoning` on the live
session returns high; a `reasoning:false` cap unmounts the pill; the
catalog row submenu still writes through and the pill mirrors it.
Credit: the dedicated-pill direction was proposed independently in
composer selector on current main with the shared-controller shape.
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.
Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).
Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.
- mcp-research-data/: 224K of July tool-search bench result rows. The
harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
gitignored dir; the rows were committed by hand once and the headline
numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
WebResearchEnv that no longer exists; the yaml paths point at a
configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
for a resolved bug, two RFCs whose work shipped, an unimplemented
profile-builder proposal, the kanban dialog mock HTML and the kanban v1
spec PDF. profile-routing.md duplicated the profile_routes section of
website/docs/user-guide/multi-profile-gateways.md.
Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
Three TypeScript clients each declared their own copy of the tui_gateway wire
types and had drifted apart: apps/shared had a partial GatewayEventName union
with a `(string & {})` escape hatch, ui-tui/gatewayTypes.ts a 150-line
discriminated union, and apps/desktop an `RpcEvent<T>` that was field-for-field
the shared GatewayEvent with `type: string`. None matched the emitter:
message.complete lacked warning/status/error/recoverable/error_surface,
tool.start/tool.complete lacked args/result, SessionResumeResponse lacked
session_key/messages_omitted/hydrating/auto_continue/todo_state, three
different ModelOptionProvider shapes disagreed on fields, and all three unions
handled a `tool.progress` event that no Python emitter has ever produced.
Now:
* `apps/shared/src/gateway-events.ts` is the single home: payload interfaces
typed from the Python emitters (file::symbol cited per interface),
`BackendGatewayEventMap` (89 backend names) + `ClientLocalGatewayEventMap`
(5 TUI-synthetic transport events, clearly marked, excluded from the
contract) merged into `GatewayEventMap`; `GatewayEvent<K>` is discriminated
on `type` with `seq` typed. RPC shapes shared by 2+ surfaces live beside it
(ModelOptionProvider = union of every field hermes_cli/inventory.py sets,
incl. pricing_pending/free_tier_pending; SessionResumeResponse<Info>;
SessionListItem with resolved_id; Usage).
* `JsonRpcGatewayClient.on<K>` is keyed by event name; the gateway.ready
heartbeat/replay_epoch and per-frame `seq` reads are typed instead of cast.
* ui-tui and apps/desktop import the shared names; their local duplicates are
deleted (no re-export shims — importers are repointed; the desktop plugin
SDK barrel keeps its public `RpcEvent` name as an alias of GatewayEvent).
web/src repoints ModelOptionProvider/ModelOptionsResponse.
* `tool.progress` handling is removed from the TUI handler/turnController,
desktop event sets/tools handler, shared union, tests, and two docs
(`grep '"tool.progress"' tui_gateway/` = 0 hits; the `display.tool_progress`
config mode is unrelated and untouched).
* `message.complete.warning` (history-commit note from
prompt_turn.py::_complete_turn_payload) is typed and surfaced on both
surfaces through their existing notice paths (TUI pushActivity 'warn',
desktop notify kind 'warning').
Contract: `apps/shared/src/gateway-events.json` is the sorted list of
backend-emitted names. `tests/tui_gateway/test_gateway_event_contract.py`
collects names from the Python emitter side (emit-helper literals, the
`.request → .expire` table, change-watcher table, child delta mirror,
subagent relay, desktop_ui tool emitters, gateway.ready/setup.ready/
browser-controller frames) and asserts emitted == JSON in both directions.
`apps/shared/src/gateway-events.test.ts` asserts BACKEND_EVENT_NAMES (which
the map type is `satisfies`-checked against) == JSON. Sabotage-verified: a
fake JSON name fails both tests; a fake TS name fails tsc + vitest; a fake
Python `_emit("...")` fails pytest.
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.
_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.
Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
headless -q, plain library use): destructive actions are now REFUSED
with a BLOCKED error and never reach the backend. Previously they
silently ran. cron honors approvals.cron_mode, unattended platforms
approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
command_allowlist entry (`cua:click:background`) and is scoped to that
action+mode — the old blanket "always_approve unlocks everything for
the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
"BLOCKED: Action timed out ..."); the error JSON keeps `action`.
Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
The developer guide showed the exact bare threading.Thread pattern the previous
commits removed from six in-tree providers; an external plugin copying it loses
the profile ContextVar under multiplex. Point it at
agent.memory_provider.spawn_context_thread and the canonical JSON/config writers.