A route without its own base_url (persisted route / SessionDB row / plain config) no longer
replaces the _resolve_runtime_agent_kwargs read in _resolve_gateway_model_context, so the
custom endpoint is still probed and the matching model.context_length pin survives; only a
/model switch carrying its own endpoint bypasses the default runtime read.
Two diagnostics diverged after a session-only `/model` switch (#111436).
/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).
Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).
Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Slim the salvaged #112111 mechanism (issue #112109) while keeping its
behaviour: the boot pass and the reconnect hook both run
_replay_pending_planned_restart_notification, which sends to every home
channel still owed an online notice, records delivered targets in
.restart_pending.json and unlinks the marker only when every owed target
(configured home with gateway_restart_notification=true) has been reached.
Dropped from the contributor diff:
- the per-target on_delivered checkpoint callback and pending_targets
field: delivered targets are written once after the pass. Residual is a
benign duplicate notice only if the process dies mid send-loop.
- getattr-based lazy lock -> class attribute default, same idiom as
run_profile_reconcile._reconcile_lock.
- _clear_planned_restart_notification in gateway/run.py: no production
caller remained; the roundtrip test unlinks the path directly.
- tests trimmed to two invariants: offline-at-boot is replayed once on
reconnect (with live-at-boot control), and partial delivery is persisted
so a fresh process does not re-notify and an opted-out home never keeps
the marker alive.
Live probe (temp HERMES_HOME, Discord home, adapter absent at boot then
reconnected): base consumed the marker with 0 sends; fixed head retains it
and sends the online notice exactly once on reconnect, then clears it.
_notif_gateway_owns_heartbeat decided by the immutable sessions.source column,
so a heartbeat on an ARCHIVED gateway-sourced row (Telegram /reset, idle/daily
auto-reset, compression rotation) was skipped by the Desktop poller and never
registered by the gateway either — restore_heartbeat_watches only claims a key
whose current session_id is that row. The tick belonged to nobody and stayed
due forever, where origin/main's Desktop fired it.
Ownership now uses the same predicate the gateway does: a gateway_routing entry
whose current session_id is this session, with an origin and not suspended
(SessionDB.gateway_routing_entry_for_session; both the session's profile store
and the launch store are consulted so multiplexed and per-profile gateways are
covered). No entry is fail-open, as on main. The check runs after the cheap
is_active/is_due gate so idle sessions never touch the DB per poll.
Probe (SessionStore telegram -> force_new, heartbeat on the archived sid):
before: desktop fired False / gateway watches [] / still due True;
after: desktop fired True / still due False; the current gateway sid is still
left to the gateway (desktop fired False) — the hijack fix stays intact.
Users pointed at the Desktop app as a workaround-breaker ("don't keep the shared session open"); state
the ownership rule so the behaviour is discoverable: a chat-registered heartbeat or /loop is fired by the
gateway and replies into the chat even while a TUI / Desktop viewer has the same session open.
Sibling of the heartbeat fix: a /loop set from a messaging chat carries the gateway's
pinned ``route`` (platform + chat_id), and ``gateway/run_goals.py::_loop_wakeup_fire_one``
already defers route-less CLI/TUI loops to their own schedulers. The TUI/Desktop poller
never returned the favor, so a Desktop viewer of the same session could fire the wakeup
on its own surface and the reply never reached the chat. Mirror the rule: a routed loop
is skipped before the session is claimed, so the tick stays due for the gateway scanner.
Live probe (temp HERMES_HOME, qqbot-routed loop, Desktop viewer of the same session):
before -> loop.ticks=1 awaiting=True (consumed on Desktop); after -> ticks=0, still due.
Desktop-owned control loop still fires.
A /handoff into a Telegram forum supergroup created a topic and bound the CLI session under
telegram🧵<chat>:<topic>, while the Telegram adapter keys every topic reply
telegram:group:<chat>:<topic> — the same restart-orphaning shape as the Slack case. Non-private
Telegram homes now use chat_type group; private-chat DM topics are unchanged.
scope_id_for_chat only consulted the channel→team map, which is empty right after boot (and
after a reconnect) until an inbound event from that channel arrives. A /handoff into a Slack home
without a stored scope_id (SLACK_HOME_CHANNEL env homes, or config homes never re-set via
/sethome) therefore built a key without the team while every thread reply carries it — the
handed-off thread was still orphaned across a restart (#111896).
When the map has no entry and the channel is not known to be shared across workspaces, fall back
to the single authenticated workspace (filled by auth.test at connect); multi-workspace installs
keep returning None.
`/handoff slack` bound the CLI session under `slack🧵<channel>:<ts>` while every inbound
reply in that thread resolves (via the Slack adapter's source shape) to
`slack:dm|group:<team>:<channel>:<ts>`. After a gateway restart the reply key found no
binding and the gateway opened a fresh empty session, orphaning the handed-off one (#111896).
The handoff destination now mirrors the adapter: chat_type `dm` for a D… home channel, else
`group`, plus the workspace scope_id (home channel provenance, falling back to the adapter's
channel→team map). Channel handoffs were equally affected (`thread` vs `group`), so the fix
covers both, not only DMs. Discord and Telegram destinations are unchanged.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Follow-up to the salvaged #111938 commit: `_slack_response_payload` already normalizes a
SlackResponse/dict body, so the new `_slack_api_error_code` helper and the two-branch
logger.error were redundant. One log line now always carries `api_error=<code|none>` so an
HTTP 200 + ok=false failure (e.g. message_not_found) is readable without exc_info.
Tests trimmed to one invariant per fix (session key on the response-ready line; API error
code on the edit failure); the `session=unknown` fallback test was a change-detector.
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.
Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.
Fixes#111695
Fold the three re-arm helpers (_typing_retrigger_state, _clear_typing_retrigger,
_send_typing_quietly) into _retrigger_typing itself; the semantic change from #111886
is unchanged: schedule sendChatAction as a tracked background task instead of awaiting
it on the send path, one in-flight re-arm per chat, at most one per
typing_retrigger_min_interval_seconds (2s default, matching _keep_typing), and honour
typing_indicator: false, which previously only gated the refresh loop.
Tests: keep the two invariants that are red on origin/main — an intermediate send()
returns while sendChatAction is stalled, and a 20-chunk stream costs one sendChatAction
per chat — and drop the eight change-detector variants.
Co-authored-by: aurel282 <aurelien.gek@gmail.com>
`_retrigger_typing` awaited `sendChatAction` inline on the send path, and
streaming re-arms after *every* intermediate send. `sendChatAction` is a
fire-and-forget UI hint whose result nobody reads, but awaiting it ran its
TLS round-trip on the same event loop as the `getUpdates` long-polls.
With several agents streaming concurrently the loop stayed pinned, the
long-polls were never serviced, and they decayed into CLOSE-WAIT while the
adapter still reported `connected` — a gateway that is deaf but healthy, which
`Restart=always` cannot recover because the process never exits.
py-spy put 20 of 20 MainThread samples in `send_typing` -> `send_chat_action`
-> `start_tls`. The handshakes are what cost: with `max_keepalive_connections=4`,
a re-arm per chunk churns the pool so most calls pay a fresh TLS handshake on
the loop thread.
Three changes, all in the re-arm path:
- Schedule the re-arm as a tracked task rather than awaiting it, so a
round-trip never delays a send or a poll. It joins `_background_tasks`, so
shutdown cancels it and it cannot outlive the adapter.
- One in-flight re-arm per chat, and at most one per
`typing_retrigger_min_interval_seconds` (default 2s, `extra` knob; 0 restores
a call per send). Telegram's bubble lasts ~5s and `_keep_typing` already
refreshes every 2s, so the re-arm only has to cover the gap left by a landed
message.
- Honour `typing_indicator: false`. Only `_keep_typing` consulted it, so the
documented workaround still paid for a `sendChatAction` on every
intermediate send.
Simulating 200 streamed chunks with a 10ms loop-blocking handshake:
200 `sendChatAction` calls and 2037ms of send-path time before, 1 call and
13ms after.
Fixes#111727
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Keep: overlapping cache misses share one probe (never two `gh` at once);
a refresh issued mid-probe after a login flip gets the fresh answer and
still never overlaps. Dropped: the bounded-helper call-shape detector (the
descendant cleanup is proven by the live wrapper probe, not by asserting
the helper's name), and the two tests of pre-existing behaviour (fresh
cache short-circuit, missing `gh`).
With the shared single-flight probe (previous commit) a `refresh=true`
request that landed while a probe was already running simply joined it. That
probe may have started before `gh auth login` completed, so the refresh
returned "not authenticated" and cached it for the full 5-minute TTL — the
composer pill kept offering /github-auth right after a successful login.
A refresh now accepts only a probe that started at or after the refresh was
requested: it awaits the in-flight one, then starts (or joins) the next.
Still only one `gh` runs at a time. A finished task whose done-callback has
not run yet is treated as absent so the loop cannot spin on it.
Live probe (real route, fake `gh` reading login state at start, 1 s answer):
PR head: refresh=true right after login -> authenticated False, cached False
fixed: refresh=true right after login -> authenticated True, cached True;
5 concurrent requests -> peak concurrent probes 1
Co-authored-by: aron-intframe <aron-intframe@users.noreply.github.com>
_start_desktop_cron_ticker now hands the built-in scheduler a callable that
re-enumerates profiles every tick (so a deleted profile stops being ticked
without an app restart). The sibling scope test in tests/cron still compared
the kwarg to a list snapshot and went red in CI; it now asserts the callable
resolves to the same homes, which is the contract the scheduler consumes.
Drop the extra assertLocalProfileCanStart call added ahead of the slot
queue: the deleted-profile fence already runs at the spawn boundary below
and is not the mechanism behind the #111338 retry storm.
Keep the two discriminating cases (a sessions.changed tick reloads the archived
set while the view is open; a failed refresh retains the last good rows) and
drop the closed-view control, which passes on the unfixed base as well.
Seven per-case tests collapsed to two: one per behaviour class (our own
close reasons → copy, server reasons verbatim; mic DOMException → recorder
copy, non-mic failures untouched). Source is byte-identical to the
contributor commit. Dropped: the `closed`/blank-reason case, the
`close_requested` filter (pre-existing behaviour), and the per-name
DOMException matrix — the mapping table lives in `micError`, asserting one
name proves the live path routes through it.
The session-end toast printed the raw wire reason (`connection_lost (127s)`,
`closed`) and a denied `getUserMedia` reached the toast as bare `DOMException`
text, while the recorder path already had friendly copy for those names.
- `liveEndedMessage` (new, exported for the tests) maps our own close reasons —
`connection_lost`, `closed`, plus a blank reason — to i18n copy; server-sent
reasons stay verbatim (unbounded, no redaction claim).
- `micError` is exported and reused on the live start path, so a mic
DOMException gets the recorder's copy; an unmapped DOMException name now falls
back to `microphoneStartFailed` instead of its raw text.
- The start catch maps DOMExceptions only: non-mic failures ('GPT-Live session
already started', 'Missing local SDP offer', API errors) keep their message.
New i18n keys: `notifications.voice.liveEndedConnectionLost` / `liveEndedClosed`
(en base + zh translation; the other locales fall back to en).
Tests: `use-voice-live-conversation.test.ts` drives the real toast store through
a fake transport — reason copy, blank/unknown reasons, `close_requested` filter,
mic DOMException names, and untouched non-mic failures.
Agent-less producers (session.activate of a lazily-resumed session via
_fallback_session_info, session.info events from agent-less cwd switches)
replace state.info with payloads that omit stored_session_id, so the exit
handler's recovery target went stale/null after such a switch. Keep the
durable id in a dedicated ui.storedSid field written at every sid transition
(create/resume/activate), read it from there in the exit handler, and when a
session.info payload omits it, carry the tracked id forward onto info.
Invariant for #111594: after drain() on mount, every later transport
generation still emits gateway.ready live, and the reconnect delay grows
across consecutive failures instead of restarting at the base delay.
An attached (dashboard-embedded) Ink TUI whose WebSocket dropped never
recovered even though the backend stayed alive, and a spawned gateway that
crashed resumed the wrong session.
- GatewayClient no longer resets `subscribed` on each transport generation:
the renderer drain()s once on mount, so every post-reconnect event
(gateway.ready included) stayed buffered forever.
- clearReconnect() keeps the attempt counter; it is reset on gateway.ready
(and kill()), so backoff actually grows across failed reconnects.
- useMainApp's exit handler no longer calls start() for an attached socket
closure — GatewayClient owns that reconnect; it only respawns a still-owned
child, and plans the resume with the durable stored_session_id (what
session.resume takes) instead of the process-local runtime sid.
- session.create's stored_session_id is carried into ui state / the active
session file so recovery and the exit epilogue target the durable id.
- The recovery target is cleared only after resumeById resolves into a live
sid, so a second disconnect during setup/history loading keeps it.
- Stale-socket identity guards on 'open'/'message'.
Salvaged from #111599 (@gustavosmendes) with trims: kept the client-side
backoff reconnect for spawned children too (the "keeps trying to reconnect
in the background" copy depends on it), kept the RPC-triggered reconnect
during backoff, kept the spawn-mode "reply in progress was lost" wording
(true for a dead child) and added attached-mode copy in userMessages.ts,
dropped the config_warning contract regen and the SessionCreateResponse
re-export (local type gains stored_session_id instead).
Fixes#111594
`_admit_prompt_turn` is the gate every turn source crosses (prompt.submit,
auto-continue, heartbeat/loop ticks, bot mailbox, compute host). When the
record's agent is None (a deferred build that attached nothing, see previous
commit) it used to hand the None agent through; `_invoke_agent` raised
AttributeError, `_recover_turn_exception` emitted a frame, and then the
`finally` dereferenced `st.agent.interim_assistant_callback` again — the
turn thread died before `running = False`, so the prompt vanished and the
session stayed "busy" for every later prompt.
Now the gate releases `running`, emits the same retryable
`agent_init_failed` terminal frame `_run_after_agent_ready` uses (carrying
the recorded build reason when there is one), logs the refusal, and returns
None so `_run_prompt_submit` reports the turn as not started. The turn body
therefore never sees a None agent and needs no guard in its `finally`.
Live probe (real `_run_prompt_submit`, inline thread, agent None):
before: `CRASH out of _run_prompt_submit: AttributeError ... running: True`
after: `returned: False ... error_surface {'layer': 'runtime', 'code':
'agent_init_failed', 'retryable': True} running: False`
Fixes#111531
Co-authored-by: zqy1-1 <185413950+zqy1-1@users.noreply.github.com>
A deferred build whose session record was closed or replaced while it ran
leaves through `_build` early, yet the `finally` still sets `agent_ready`
with `agent` None and `agent_error` None. Record the cause on the record so
the turn that was waiting on it can refuse with a real reason instead of
running against a missing agent (#111531).
Trimmed from PR #111534: the `_wait_agent_for_prompt` readiness hunk and the
turn-body guard are superseded by a single refusal at `_admit_prompt_turn`
(next commit), which every turn source crosses.
The repo's eslint config forbids `typeof import(...)` type annotations
(@typescript-eslint/consistent-type-imports); derive the module type from the
dynamic-import thunk instead so `npm run check:lint` stays at 0 errors.
The previous commit claims every outbox before delivering and runs
deliveries per target profile, but `drainBusy` still spanned the delivery
phase: a `bot_relay.outbox.pending` push that arrived while one lane ran a
long turn (up to RELAY_DELIVER_TIMEOUT_MS) only set `drainRerun`, and the new
envelope was claimed after that turn — the reporter's step 3 (bot C mails D
while A→B runs) still ended in `queued_expired`, because the gateway checks
the TTL at the claim.
Scope `drainBusy` to the claim phase and make the delivery lanes module
state: `relayLanes` maps `target_connection::target_profile` to the tail of
that target's in-flight deliveries, so a later drain appends to the running
lane (same target stays ordered, one turn at a time) or starts a new one
(other targets run now). Lane entries drop once idle; stopBotRelay clears
them so a restart begins fresh.
Tests: keep the contributor's red-on-base test (claims every outbox first,
delivers to different targets concurrently) and replace the ordering-only
test — green on base — with one that pins the mid-delivery claim plus the
same-target ordering (red on base AND on the previous commit alone).
Docs: bot-mode.md states the delivery concurrency contract.
Part of #111587 (with the previous commit: Fixes#111587)
drainRelayOutboxes drained one gateway's outbox and delivered its envelopes
before draining the next gateway, and delivered every envelope one after
another. One long turn — bot_relay.deliver may take up to
RELAY_DELIVER_TIMEOUT_MS, 25 minutes — therefore held every other bot's mail:
a sibling gateway's envelope sat unclaimed in its outbox, and since the
gateway checks the envelope's age against bot_mode.envelope_ttl_seconds
(15 minutes) at the claim, the sender's waiter received queued_expired for a
message nothing was wrong with; envelopes that did get claimed still waited
their turn behind unrelated deliveries, against a finite waiter.
Claim every gateway's outbox first, then deliver in lanes keyed by target
connection and profile: a lane runs its envelopes in order, one turn at a
time (the target gateway serialises that profile's turns behind its turn
lock anyway), and lanes run concurrently.
The composer hands key bursts to the parent on a 16ms timer. When
history navigation, a slash completion or a submit clears/replaces the
draft while such a flush is still armed, the [value] effect reset the
local buffer correctly but left the timer running, so 16ms later the
stale burst was handed to the parent and overwrote the external value
(second hole in #111934).
Disarm the pending flush in the external branch of the [value] effect:
once the parent has replaced the draft, a burst typed against the old
draft can never be the newer value. Second invariant test covers it
alongside the stale own-echo case.
A deferred key-burst flush can still be in flight when the parent's
re-render lands: the echoed value is the one we emitted, older than
vRef because the user typed past it. The [value] effect treated any
non-equal incoming value as an external assignment and rewound local
state — the cursor jumped backward and freshly typed letters were
overwritten (#111934).
Track the last value handed to onChange; an echo matching it stays on
the own-change path, and the pending flush for the newer local value
converges the parent on its next timer.
createMainWindow walked the entire renderer generation twice before it
could call loadWindowUrl:
const rendererIndex = DEV_SERVER ? null : resolveRendererIndex()
const tornAssets = rendererIndex ? missingRendererAssets(rendererIndex) : []
resolveRendererIndex already computes exactly that list while choosing the
copy — it needs it to decide whether a copy is torn — and then throws it
away. missingRendererAssets is a BFS that readFileSync's every present
chunk whole and regex-scans it for the inline __vite__mapDeps table, so on
a release tree it is not a stat walk: measured against the real
apps/desktop/dist (252 chunks, 28.7 MiB of JS), one walk is 162
readFileSync calls reading 28.23 MiB, 576 existsSync calls, and 56.6 ms
median (min 55.8, 9 reps, warm page cache, darwin-arm64). Both walks run
synchronously on the main thread before the window gets its URL.
Return the list alongside the index. resolveRendererIndexWithMissing()
carries the existing body and hands back { index, missing }; the
path-only resolveRendererIndex() stays as a one-line wrapper so the nine
other call sites are untouched. The primary-window path takes one
resolution.
Semantics are unchanged in every branch: the same candidate is chosen, the
same log lines are emitted, and the missing list always describes the copy
actually returned. The all-copies-torn branch now reuses the first
candidate's list, captured on the first loop iteration, rather than
recomputing it for present[0] — recomputing there would have reintroduced
the second walk in exactly the case that matters most, and using the
loop's last value would have described a bundle we do not load.
Net effect on every primary-window boot: one fewer full walk, so 162 fewer
readFileSync calls, 28.23 MiB less synchronous reading, 288 fewer
existsSync calls, and ~57 ms of main-thread blocking removed before
loadURL. The win lands on packaged and --prod launches; DEV_SERVER skips
the walk entirely, so `vite dev` is unaffected.
`.dt-portal-scrollbar` is the same themed bar as `.scrollbar-dt`, applied
to overlays that portal under document.body (dropdown/context menus, the
command palette, the session and connection switchers). Widening only the
#root theme (#111634) would have left those lists on the 4px bar that was
too thin to grab; keep the two variants on one width (0.5rem = 8px).
The app-wide .scrollbar-dt theme (on #root) rendered 4px scrollbars
everywhere, including the conversation window, making them nearly
impossible to see or grab (#111634). Bump the webkit track size to 8px
so the thumb is actually hittable while staying a slim themed bar.
Fixes#111634