Follow-up to the renderer-dial retry:
- Drop the `!connectionDescriptorResolved` gate so `bootFailureIsRetryable()`
still decides every failure main can see (#82679 contract preserved for
post-descriptor mint failures); the renderer-owned dial is OR'd in as the
one failure main cannot classify. The budget check now short-circuits
before the IPC, as on main.
- Four booleans → one `stage` variable; the `isGatewayReauthRequired` term
is dropped (reauth is raised before the dial stage, so it was unreachable).
- `isGatewayWebSocketUrl` moves to `apps/shared/src/json-rpc-gateway.ts`
and `JsonRpcGatewayClient.connect()` uses it, so the hook's "valid dial"
predicate cannot drift from what `connect()` accepts.
- Tests: keep the three that bind behaviour (first dial retries under a
stale ready snapshot; invalid URL terminal; post-connect failure
terminal), drop the four that pinned the removed gate or duplicated the
existing #82679 bound test. 46/46 green; reverting the hook to main makes
the retry test fail.
The old tests encoded the skip (agent never constructed, [drift_skip] text,
alert-once bit). New invariants, proven red on origin/main: an unpinned job
with model_snapshot=A / provider_snapshot=P runs with AIAgent(model=A) and
resolve_runtime_provider(requested=P) after the global default moved to B/Q;
an explicit job pin still beats the snapshot; cron.model / cron.model_provider
still beat the snapshot; a legacy job without a snapshot follows the global
default. test_cron_drift_alert_once.py is deleted with the bit it tested; the
config-notice and impact-summary tests drop guard_enabled / model_drift_guard;
the desktop toast test asserts the informational copy.
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).
The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.
Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".
Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).
The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
With the startup config read now retried (previous commits, #76158), a retry
can resolve after the user has already chosen a language in Settings and
would repaint the stale on-disk value over it. Mark the explicit pick in a
ref and let every retry outcome (success or exhausted fallback) defer to it.
Idea and guard from #105470 (webtecnica), folded onto the #76158 base.
Co-authored-by: webtecnica <webtecnica@gmail.com>
Address review feedback on the startup-race fix:
- Restore the permanent-failure contract: a rejected config load settles
on English (DEFAULT_LOCALE) so the UI stays usable, as enforced by
context.test.tsx.
- Bound the retry loop to MAX_LOCALE_RETRIES (10) at 3s intervals instead
of retrying forever — permanent auth/backend failures reach a terminal
English state instead of scheduling requests indefinitely.
Tests: add coverage for transient-failure recovery (first attempt rejects,
retry applies persisted display.language) and for bounded-budget exhaustion
(no retry fires after the budget is spent).
The renderer mounts before the desktop's own backend is ready, so the
first GET /api/config times out after 60s. The catch handler fell back
to DEFAULT_LOCALE (en) permanently, leaving the UI stuck in English
until a manual language switch, even when config.yaml persisted
display.language: zh.
Retry with a 3s delay on failure instead of locking the default locale.
The backend comes up, the next attempt succeeds, and the persisted
language takes effect. Cleanup clears the pending retry timer on unmount.
Both steer sites now build the row through one helper, prompt_builder.steer_user_row:
a role:user row with display_kind="steer" and no leading blank lines. The alternation
repair (_merge_consecutive_users) skips a steer-typed prev row, so a run that ended
right after a steered batch (Ctrl-C, interrupt) does not get the next real prompt
merged INTO the already-persisted steer row — which would have rewritten it in place
and re-broken live≠replay parity, the exact class this PR fixes.
TUI/desktop history projects the steer row as the user's own words instead of the
model-facing marker wrapper; 'steer' joins the display_kind union. The compression
anchor scan keeps its tool-row branch for transcripts persisted before this change and
its docstring says so.
Port Adolanium's focused-turn pose from Hermes-Bot-Mode#101 and
hermes-agent#88134 to the current typed Bot Mode implementation.
Match the busy signal's connection-qualified focused owner rather than
the gateway socket, retain worker activity, and ease transitions in
elapsed time on the existing shared face clock.
Includes owner-isolation and animated-pose invariants, both proven red
on origin/main, and native Electron before/after verification against
a real temporary Hermes backend with held loopback inference.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Add Move up/down controls for actual rooms without changing bot or folder
ordering. Preserve default pin/activity ordering until an explicit move,
retain hidden room slots, and persist Desktop-local order through room
updates, mirror merges, and hydration. No membership or routing writes.
Adapted narrowly from the group ordering idea in archived
NousResearch/Hermes-Bot-Mode#105 by @onuraycicek; rename already exists.
Co-authored-by: Onur Aycicek <onur.m.aycicek@gmail.com>
The importer stays in the command palette. The labeled nav row was
clutter next to New session / Capabilities / Messaging.
Co-authored-by: Cursor <cursoragent@cursor.com>
hideOnly chrome pinned the Sessions/Bots strip on at any tab count, so
never was a silent no-op. The panes stay; ⌘⌥T brings the strip back.
Co-authored-by: Cursor <cursoragent@cursor.com>
Complete the PR #101452 salvage, preserving sprmn24 authorship. Credit @kokhlo PR #101591 for independently diagnosing the in-flight rename race; use scoped authoritative room bindings rather than persistent aliases. Keep mention attention independent and retire late work after disband.
$groupNeedsYou was written by two independent sources: an @user mention
(appendGroupChatEntry) and, until now, syncGroupClarify whenever a member
blocked on a clarify or approval. Nothing kept the two in sync, so every
path that consumed a clarify -- answering it, the server resolving it,
disbanding or renaming the room -- left a stale badge lit with nothing
behind it, because none of them cleared the copy syncGroupClarify had
written. A naive fix (writing false back into $groupNeedsYou on every
clarify cleanup path) would have created the opposite bug: clearing an
unrelated, still-unresolved @user mention.
$groupClarify is already the source of truth for pending clarify/
approval attention. syncGroupClarify no longer writes $groupNeedsYou; a
new groupHasPendingClarify(clarifies, group) derives the same signal
from a $groupClarify snapshot. It's intentionally pure -- the caller
(roster-pane) subscribes to $groupClarify itself via useValue and passes
the live snapshot in, so the subscription actually drives the
recalculation instead of existing only to force a re-render. A prompt
resolving, being answered, or its room being disbanded all self-correct
through the existing $groupClarify cleanup paths, with nothing left to
keep in sync.
Rename needed its own fix: the old clearGroupClarify(oldName) call on
rename dropped a clarify-only room's attention entirely, since clarify no
longer lives in $groupNeedsYou to be carried over by the existing
old->new key swap there. New renameGroupClarify(oldName, newName) re-keys
matching mirrors onto the new name in two passes -- unrelated entries
first, migrated entries last -- so a stale mirror already stranded at the
destination key (left behind by a known, separate in-flight-poll race)
can never clobber the just-migrated current prompt.
The roster now reads groupNeedsYou[group] || groupHasPendingClarify(...)
and subscribes to both stores so either one repaints the row.
Tests: multi-clarify sequencing, mention+clarify independence (resolving
one must not clear the other), server-side resolution cleanup, rename
migration, rename destination-collision (red-green verified against the
prior single-pass implementation), a GroupRow badge render test, and a
subscription harness using the real stores/useValue/helper end-to-end.
(cherry picked from commit e5038f12ccc1a4d1fb56ecc5f0e1a884f1e8b795)
Separate saved Cloud instances from the live window source, use the existing
registry activation path instead of repeating sign-in/apply, and persist the
chosen instance name while retaining registry identity and custom labels.
Naming metadata adapted from IAvecilla's contribution in PR #103224.
Co-authored-by: IAvecilla <ignacio.avecilla@lambdaclass.com>