In the Bots pane the Cronjobs rows were inert. The only interactive
controls were the enable switch and the hover-only delete button, so
clicking a cronjob to see what it runs, when it runs next, or why it
stopped did nothing at all — while the same job on the main Cron page
opens a full detail panel.
The gateway already ships every one of those facts with
`cron.manage list` (schedule, repeat, next/last run, last status,
delivery target, model, workdir, prompt preview, and the
fire/delivery/pause failures). None of it had a surface in Bot Mode: a
job failing every run reads exactly like a healthy paused one.
The row title becomes a real button that opens a read-only inspector
rendered from the record the pane is already holding — no extra RPC,
and no second mutation path beside the row's own switch and delete. The
switch and delete button stay siblings of the opener, so a toggle can
never be swallowed by the open. The inspector tracks the job by id
rather than by object, so the 20s poll keeps an open panel live instead
of freezing the snapshot it opened with.
- Clear the accidental end stamp on resurrection (at the lineage tip):
a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
archive auto-resurrect on the next lookup — the user could never retire
the canonical chat. Test pins the resurrect -> deliberate-archive ->
stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
compressed lineage carries end_reason='compression', so tip-stamped
accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
dm resolution) filtered archived rows out via list_sessions_rich and
still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
Final-diff pass: trimGroupChatLog drops entries from the FRONT once a room
crosses the history cap, so slicing the post-turn log at the pre-turn
LENGTH could overshoot after a mid-turn trim, read an empty tail, and
silently commit a stale turn — re-opening #93127's double delivery in
long-history rooms exactly. Anchor on the last pre-turn entry's id; if the
anchor itself was trimmed, every surviving entry is newer, so scanning the
whole log stays exact.
- Cross-thread supersession no longer discards finished work: an epoch bump
from a send in ANOTHER thread doesn't re-drive this thread's members (delta
filters are thread-scoped), so dropping the finished reply lost completed
work until someone revisited the old thread. shouldCommitMemberTurn now
drops only when a newer USER entry landed in the same thread; the caller
computes that from the log tail past the pre-turn length.
- '@all stop' now holds every member — it parsed to everyone:true with no
mentions and silently held nobody, the asymmetric twin of the tested
'@all resume'. classifyGroupHoldDirective gains holdAll; the send path
passes the room's member keys for expansion.
- Tests pin both: cross-thread commit preserved, @all-stop holds all
(mutation-checked: reverting either guard fails its test).
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
wrapping it in --run-delivery would make the child contend with its
parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
loaded CI runners; the wait-duration message assert matches ~Ns generally.
A user 'stop @member' was just log text: the next room delta (receipt
round completing, any later turn) re-dispatched the member and it
re-claimed the very task it was told to stop. Holds are now durable
room state: set by an explicit user stop mention, checked by the round
loop before dispatch (skip consumes the delta exactly once — no spin),
released only by an explicit resume, @all resume, or a direct non-stop
mention of the held member. Holds persist and rehydrate with the same
durability as room watermarks, and the activity feed shows WHY a held
bot is silent (⏸ held glyph + hint) the first time it is skipped.
Conservative parse documented in-code: any standalone stop/halt/pause
next to a mention holds — a wrongly-held bot is one mention away from
release; a wrongly-running one keeps doing forbidden work.
Rebase onto main (post-#93102) left two bot_mode dicts in
config_defaults.py — later duplicate key silently wins in a Python
literal, so envelope_ttl_seconds would have shadowed turn_wait_seconds'
section. Single section now carries both keys.
Two CI failures: (1) the target_busy test's global subprocess.run patch
recorded unrelated gateway-init git calls (rev-parse/ls-remote) as the
delivery spawn — now local_delivery_command is monkeypatched to a
sentinel argv so only the real delivery path counts; (2) the new
bot_mode config section is single-field, tripping the dashboard
no-single-field-categories rule — folded into the agent tab via
_CATEGORY_MERGE (same as #93102's fix).
Final-diff pass: pin both envelope mtimes via os.utime relative to the
watermark (write_text alone is wall-clock/FS dependent), and reset
relayDrainRerun in stopBotRelay so a rerun remembered mid-drain can't
leak one stale drain into the next start/stop cycle.
Review follow-ups:
- A push signal landing while drainRelayOutboxes is mid-flight hit the
relayDrainBusy early-return and was gone forever — the gateway signature
is monotone (one event per new envelope, never re-broadcast), so the
envelope waited out the full 4s poll, exactly the latency the push path
removes. relayDrainRerun remembers the race and schedules one debounced
follow-up pass after the drain finishes.
- test_new_envelope_after_drain_fires_pending_again pins the untested half
of the monotone contract: the watermark must not eat genuinely NEW
envelopes (write -> drain -> write-newer fires twice). Mutation-checked:
a stale-signature regression fails it while the other three still pass.
Cross-connection DMs were pure polling: the Desktop drains every gateway's
bot_relay outbox on a 4s interval, so each hop eats up to 4s outbound plus
4s for the reply leg (#92760 'bots reply slowly').
Emission point: the gateway's existing change watcher (_CHANGE_WATCHES in
tui_gateway/server.py). Envelopes are written by the AGENT process
(message_agent -> tools.bot_relay.enqueue_envelope), not the gateway, so no
gateway RPC is on the enqueue path and an in-process emit is impossible.
That is exactly the situation the change watcher already solves for the
pairing store (pairing.changed: 'written by a different process; the files
are the only shared signal') - so a new bot_relay.outbox.pending entry in
the existing watch table is the smallest correct diff: one cheap 1s-interval
stat probe folded into the existing 0.5s watcher tick, no new thread, no new
RPC, and _broadcast_global_event fans it to every connected WS client for
free. The signature is monotone (newest envelope mtime ever seen) so a
drain emptying outbox/ never re-fires the event.
Desktop (hermes-bots plugin): subscribe via the existing host.onEvent tap
(feature-detected - older shells lack it) and run drainRelayOutboxes through
a 250ms trailing debounce so a burst of signals collapses to one drain.
The 4s interval poll is intentionally UNCHANGED as the backstop: the event
tap only hears the active gateway socket, so per-connection push detection
would be complex and wrong to trade the poll against - push simply makes
the common case near-instant while older backends keep working exactly as
before.
Tests: 3 new watcher contracts (fires on enqueue, monotone across drain,
silent with no outbox) and a new relay-push-drain.test.mjs (debounce burst
-> one drain, re-arm after window, disposed no-op, poll backstop intact).
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
The new bot_mode.envelope_ttl_seconds default created a one-field
dashboard category, tripping test_no_single_field_categories. Merge it
into the agent tab via _CATEGORY_MERGE like code_execution et al.
Review follow-up: bare \b401\b / \b402\b / \b429\b / \b5xx\b matched any
3-digit token in error text ('line 502', 'took 429 ms'), and server_error
misfires feed AUTO_RETRYABLE — a supervisor could auto-retry a permanent
local failure. Numeric rules now require an 'error code:'/'status:'/'http'
prefix; phrase alternatives (rate limit, server error, overloaded, out of
funds) unchanged. Adds parametrize rows for the false-positive guards and
the previously untested branches (bare 'status: 401', 'upstream server
error', 'model_not_found').
Review follow-up: the relay drain records attention under
'<connectionId>::<profile>', but local/unannotated roster rows carry no
bot.connectionId — botRosterKey gives 'legacy::name' and botSelectionKey
bare 'name', so a failed relay DM to a bot on the ACTIVE connection never
rendered its badge. BotRow now also checks
'<bot.connectionId || activeConnectionId>::<name>', covering exactly the
rows the user is most likely looking at. Test pins all three lookup shapes.
Post-merge review follow-up for #93080: the isDisabled guard test
documented 'flipping back also notifies' but never asserted it — a
regression making the true->false transition silent (e.g. gating
disabledChanged on truthiness) would still pass. Pin the flip-back
notify with a beforeFlipBack capture (mutation-checked: gating the
seed on newDisabled truthiness now fails this test).
Review follow-up: __internal_setAdapter assigned this.isDisabled before
the fast path but never fed it into the new 'changed' flag, so an
isDisabled-only flip on an otherwise-identical adapter swap would have
been silently swallowed. Seed 'changed' with the isDisabled comparison
and add a guard test (mutation-checked: reverting the seed fails it).
Two independent desktop-boot/runtime bugs found driving the app over CDP
against current main, each pinned by a regression test:
1) Adapter no-op notify loop: IncrementalExternalStoreThreadRuntimeCore.
__internal_setAdapter's fast path (same isRunning + same messageRepository)
called _notifySubscribers() unconditionally. ChatRuntimeBoundary passes a
fresh adapter literal every render, so any subscriber whose notification
re-renders the boundary loops render->setAdapter->notify->render until
React kills the tile with 'Maximum update depth exceeded' (reproduced live
on every bot-profile switch; session tile dies behind its error boundary).
Now the fast path notifies only when extras/suggestions/capabilities
actually changed.
2) Dual-venv interpreter mismatch: findPythonForRoot() prefers .venv over
venv, but createPythonBackend() hardcoded venvRoot=root/venv for
PYTHONPATH. A checkout with BOTH venvs (dev .venv 3.12 + install venv
3.11) got a 3.12 interpreter with 3.11-compiled native wheels on
PYTHONPATH and died on the first import (pydantic_core) before the
gateway bound - the renderer then showed 'Gateway offline' on every
profile. venvRootForPython() now maps the selected interpreter back to
ITS venv; root/venv remains the fallback for system pythons only.
The real fix for Bot Mode 'session not found' / endless hang: dispatch
session-scoped RPCs on the OWNING profile's local gateway, using the route
the chat tile already carries — the same multi-connection machinery Sessions
mode uses, which has never had this problem.
Root cause chain:
- A bot chat is a persisted tile that records its exact owner (connectionId +
profile) in tile.ownerRoute; requestForSessionProfile already dispatches on
any (connectionId, profile) via the per-profile local gateway pool.
- But wiring's requestGateway resolved the owner via rememberedSessionProfile,
a $sessions row lookup. Canonical Bot Chats are born hidden (never listed),
so the lookup missed and fell back to the ACTIVE profile -> prompt.submit hit
the launch backend that never owned the session -> 4001, and the resume
ladder re-resolved through the same blind spot, so it hung.
- It also keyed off $selectedStoredSessionId, but a bot chat renders in a TILE
whose id is $focusedStoredSessionId (selected stays the primary pane), so
even the row path was reading the wrong session.
Fix:
- wiring requestGateway: resolve owner from the FOCUSED stored id, preferring
the tile's persisted ownerRoute; fall back to the list-derived profile only
when no tile route exists. One resolver, every session RPC (submit, resume,
attach, interrupt, compress) inherits it.
- sdk openSession: synthesize a local ownerRoute from for bot opens
that carry no explicit cross-connection route, so LOCAL bot tiles carry their
owner too (previously only remote routes did). Strictly routing metadata:
the dial path, all-profiles view, and the route-registry retry check all
still key off the EXPLICIT route, so a plain local open behaves exactly as
before (no registry-secondary dial, no forced all-profiles view).
Fixes already-open chats (tile route is persisted, needs no fresh open) and
survives relaunch. 3 tests for sessionTileOwnerRoute. tsc 0 errors.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
When the gateway reaps the runtime behind the open bot chat (idle TTL,
LRU cap, or the WS-orphan mass reap that killed every background bot's
handle at once in the Aug 23 incident), the plugin now hears
session.reclaimed and re-resumes the canonical chat immediately, instead
of leaving the dead handle for the user's next send to trip over.
Matched on the stored id against both claim identities; guarded by the
open generation so a user action mid-re-resume wins; a failed re-resume
is swallowed — the next-send recovery ladder (#92928) stays the
backstop. Feature-detected on host.onEvent; disposed with the other
listeners.
groupSpeakerLabel resolved friendly identity for exactly one case: the
literal profile name 'default' → 'Hermes'. A renamed default (core
display_name via 'hermes profile rename', e.g. Lucy) or a Bot Mode title
never reached the room's working line, activity feed, or transcript
speaker prefix — the community report was Lucy's group turns still
reading 'hermes thinking'.
The label now walks the same rungs as displayName(): Bot Mode title
first, then the ACTIVE gateway roster row's display_name (remote/thin
rows are skipped so another connection's default can't lend its name),
then the existing default→Hermes fallback.
Validation: group-chat.test.mjs 87/87, full hermes-bots suite 474/474.
* fix(desktop): route hidden Bot Chat RPCs to the owning profile backend
Bot Mode chats failed with 4001 'session not found' (then hung on retry)
for every bot except the launch profile. Root cause is an identity gap in
the session-RPC router for HIDDEN sessions:
- wiring's requestGateway resolves the owning profile via
rememberedSessionProfile($sessions, selectedStoredSessionId, active).
- Canonical Bot Chats are born hidden (hermes-agent#86797), so the sidebar
aggregator NEVER lists them; whenever the in-memory row is absent the
lookup misses and the resolver silently falls back to the ACTIVE profile.
- prompt.submit then lands on the launch backend, which never owned the
session -> 4001. withSessionNotFoundResume's session.resume ALSO resolves
the profile through the same blind spot, so the recovery re-registers the
wrong backend too and the ladder dies without ever reaching the bot's own
(healthy) gateway. Log fingerprint: the bot backend shows ws accepts with
messages=0 and no 'tui prompt accepted' after the reap, while the launch
backend answers 4001.
Fix at the resolver, so every session-scoped RPC (submit, resume, attach,
interrupt, compress) gets the same answer:
- rememberedSessionProfile: when no session row matches, consult the
session-owner hint (targetProfile over profile) before falling back to
the active profile.
- sdk openSession: record an owner hint for LOCAL plugin opens that carry
an explicit profile (remote routes already did via ownerRoute). Hidden
sessions have no sidebar row, so this hint is the only durable owner
record the router can consult.
3 focused tests: hint fallback for a hidden session, targetProfile
preference, and row-over-hint precedence. tsc 0 errors (baseline-equal).
* test: add activeGatewayConnectionId to the full-replacement gateway mock
sdk/index.ts now imports it for the local-open owner hint; the full
vi.mock('@/store/gateway') in profile-routing.test.ts must export every
symbol the module under test imports or all 18 suite tests fail at load.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
The shared checkout serves every profile, but hermes update migrated
only the active profile's config.yaml. Siblings kept their old
_config_version until their (correctly restarted, post-#91378) gateway
hit a config shape the new code couldn't read — the last unabsorbed
substance from the Phase-2 restart-swarm audit (#20438 earliest, 2026
field repro on #79048: sibling at v33 vs v37).
_migrate_sibling_profile_configs(): per sibling home, scope config
reads/writes via the context-local HERMES_HOME override (ContextVar —
never os.environ), check version, run the NON-INTERACTIVE safe
migration; prompt-requiring settings stay for the profile's own next
interactive session (same contract as gateway-mode). Broken profiles
are skipped without blocking the sweep; override always reset.
Sabotage-verified; live E2E in a fresh process with real drifted
config files: v12→v38 and v25→v38 on disk, provider preserved, the
documented #81946 personality-reset migration correctly applied to
siblings too, never-configured profile untouched, active home
untouched, second run idempotent.
The dialog body's grid used the implicit column, which sizes its track to
unbreakable content: the nowrap view-link <code> forced the track wider than
the dialog, so the description and URL were clipped and the Copy button sat
past the right edge behind a horizontal scrollbar. Both DialogContent body
boxes now pin the column to minmax(0,1fr) so children truncate instead of
widening the track — this hardens every dialog against long unbreakable
content, not just this one.
The view link is also now a real anchor (system browser on click, link
context menu on right-click) inside a data-selectable-text row so the URL
can be highlighted and copied by hand; truncation clips the paint only,
selection still carries the full URL.
session.create intentionally persists no state.db row until the first
prompt, but session.resume only looked in the database — so resuming a
live lazy session by its stored key or pending title hard-404'd. Bot
Mode hits this on every fresh non-default bot: the canonical Bot Chat
is created lazily on the profile, the open/send resumes it, and the
user gets 'session not found' on their first message to that bot.
session.resume now falls back to the in-memory session registry,
matching by stored key or pending title scoped to the SAME profile
home. Cross-profile lookups still fail closed; unknown ids still 404.
The Bots home landing appeared on EVERY first click of a bot whose
canonical Bot Chat had been compressed; only a second click got through.
openRosterBot claimed the center with the durable registry id, but the
session-focus edge fired by the open itself reports the compression-
lineage TIP. releaseStaleOpenBotChat compared tip !== registry id,
declared the claim stale, released it, and the home reasserted over the
freshly opened chat. The second click worked only because the tip was
already focused — no new focus edge fired to sabotage it.
openBotCanonicalChat now returns both identities (registryId + openedId);
the claim carries both; a focus edge matching EITHER keeps it. Foreign
sessions still release, and the legacy no-id draft claim is unchanged.
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).
- update_inventory: restart_via becomes a machine-readable mechanism id
(systemd|launchd|desktop|manual) — THE policy table as data; display
derived via describe_restart_mechanism. match_runtime_outcomes()
reconciles every planned runtime against the restart phase's
bookkeeping (restarted/stopped/failed/unaccounted);
report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
land in the receipt (runtime_outcomes); any unaccounted runtime
escalates exactly like a STALE/DOWN fleet row (exit 1).
Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).
_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.
That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.
Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.
A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.
Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.
tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.