POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.
Refs #94724
The smart-approval guardian (`_smart_approve`) gates every flagged
terminal command with a synchronous auxiliary LLM call, but it never
passes `timeout=` and logs nothing on the normal path. In production a
stalled provider response silently froze the agent turn for 62 minutes
with zero log output; the gateway kill-switch eventually fired, and only
an unrelated error surfaced afterwards (#82846; watchdog-style fix in
#72500). The call was invisible by design — nothing logs at the hang
point.
Changes in tools/approval.py:
- Resolve the same configured timeout the client would use internally
(`auxiliary.approval.timeout` via `_get_task_timeout("approval")`) and
pass it explicitly to `call_llm`, so the deadline cannot be lost if the
internal default resolution changes or is misconfigured.
- Log the assessment call and its duration (DEBUG), and promote the
failure branch from DEBUG to WARNING with elapsed time + exception
class, so a wedged guardian call is visible in the logs instead of
silent.
- Failure still returns "escalate" (fail open to the human/pattern
gate) — behavior unchanged, observability only.
Complements #72500 (watchdog hard ceiling) rather than duplicating it:
explicit timeout is the root-cause hardening, logging closes the
silence gap; the watchdog remains the safety net if the SDK-level
timeout itself is defeated.
Tests: explicit timeout forwarded to call_llm (revert-fails), failure
logs WARNING + escalates. 49 approval-adjacent tests pass; one unrelated
test_approval.py failure is pre-existing (fails on clean main too).
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.
The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.
Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
lost" path orphaned setsid grandchildren the same way (whole-bug-class
rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
that exits right at the deadline doesn't log a spurious "no signal"
warning (mirrors _terminate_cron_script_process); pinned by
test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
CI load could eat the whole 1s window before the spawner wrote its
pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
old path gave a 1s SIGTERM grace window; intended for a deadline-
expiry hard stop (both docstrings say "hard stop")
Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.
Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
Apply-ready delta distilled by @andrexibiza: deterministic watcher-consumer
tests (watcher times out while a child is alive, resolves when all are dead),
psutil-unavailable fail-open pin, and probe-failure fail-open handling in
_stdio_children_dead (unknown is never proof that every child exited).
Local: 8 passed on tests/tools/test_mcp_stdio_children_dead.py
_stdio_children_dead returned True ('all children dead') on the first LIVE
pid — the intended False was dead code right below it. Every spawn path
that captures child PIDs (observed in hermes -z oneshots) then failed the
#81995 pre-call fast-fail with 'TimeoutError: MCP stdio subprocess ... has
exited' on every tools/call while the subprocess was demonstrably alive.
Long-lived gateway/dashboard sessions were unaffected only when
_stdio_child_pids was empty (the not-pids short-circuit).
Return False on the first live pid and drop the unreachable line.
TestGetHermesHome.test_default_path asserted ~/.hermes unconditionally,
but the native Windows default is %LOCALAPPDATA%\hermes (see
hermes_constants._get_platform_default_hermes_home). Branch the
assertion by platform so the test passes everywhere.
Salvaged from PR #96003 by @Aoshi-Dev (the parse-guard half of that PR
was superseded by #96169); authorship preserved.
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):
- config_command: catch RuntimeError for set/unset and print a clean
one-line error + exit(1) instead of a raw traceback on the primary
'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
ConsoleCommandError so 'hermes console' and the dashboard console
report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
branch — the old fallthrough claimed 'falling back to default config'
even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.
Fail closed when config.yaml is unparseable or non-mapping before set/unset writes, reuse the readable-config guard to return the parsed mapping, and cover refuse/empty-mapping paths with regression tests.
- Reuse the existing _commit_status variable for the terminal-edge gate
instead of the parallel _compaction_succeeded boolean (derived state).
- Give the commit_fence_cancelled abort the same force_terminal=True
terminal edge as the lock-contended abort, and reword the closure
comment that overstated the lock contender as 'the one exception'.
- Inline the codex app-server path's lifecycle closure: after gating on
success it reduced to a single success-site emit, so the scaffolding
(done-flag + closure + two no-op failure-path calls) was dead.
The send_media lane (08b95c3) had no committed regression test: cover
explicit-bool stamping on media frames, the descriptor-platform
fallback when _platform_by_chat is empty (post-restart proactive
sends), and the omitted-key absence case.
When either unfurl key is set, media captions post as a separate
message before the file (the upload API cannot carry unfurl controls)
and native draft streaming falls back to edit-based delivery. Surface
both side effects in the config reference table.
The send and send_media lanes resolved the platform only from
_platform_by_chat, which is empty until an inbound frame arrives (e.g.
after a gateway restart). A proactive send to a Slack chat then missed
the unfurl stamp. Mirror the streaming gate and delivery resolver:
fall back to the negotiated descriptor's platform.
Relay-plane parity: hermes config set / Railway persist YAML booleans
as strings, and _slack_unfurl_kwargs silently dropped them — so
'unfurl_links: "false"' was a no-op on native while working on relay.
Coerce recognized string booleans exactly as _slack_unfurl_hints does;
unrecognized values still drop so junk config keeps Slack's default
instead of accidentally suppressing previews.
Replaces test_send_ignores_non_boolean_unfurl_options (which froze the
dropped-string behavior) with coercion + junk-drop tests.
Live staging (Coatue Slack):
- hermes config set / Railway knobs persist "true" as a string; bots that
omit unfurl_links do NOT inherit the human default, so dropping the
string looked like suppression.
- chat.startStream cannot carry unfurl_*. Native SlackAdapter already
falls back to chat.postMessage; the relay now matches.
Relay-fronted Slack reads platforms.relay.extra.slack.unfurl_links/unfurl_media
and stamps explicit booleans onto the frame metadata; the connector forwards
them to chat.postMessage with no config of its own (mirrors reply_in_thread).
Covers send, send_for_platform (cron/scheduled), and send_media lanes.
* fix(relay): map wire media[] → event.media_types; accept message_type voice
A relayed voice note arrived as MessageType.AUDIO with media_types=[] —
the STT gate (_event_media_is_stt_input) excludes AUDIO unconditionally
and its per-attachment MIME rescue was unreachable, so STT never fired
and the agent fell back to the "user sent an audio file attachment"
context note (live-verified on staging 2026-08-26, Discord + Telegram).
Two wire-boundary fixes, both additive within contract_version 1:
- "voice" parses to MessageType.VOICE: the enum already had it — pinned
by test so a future refactor can't collapse the two.
- media[] is now mapped into event.media_types (positional alignment
with media_urls; mime-less entries keep their slot as ""). This is
what run.py's per-attachment classifiers key off, so EVERY relayed
attachment — image vs document, audio vs voice — now routes like its
native-adapter equivalent, not just voice notes.
Behaviour pinned: new-connector voice → STT-eligible; legacy
audio-typed events unchanged (no STT); music uploads never STT-eligible
(direct _event_media_is_stt_input assertions on real wire-parsed
events, not mocks).
Pairs with the gateway-gateway PR that puts "voice" on the wire.
* review: pin the STT gate by test; fail safe on media/media_urls mismatch
Addresses independent review of #95274.
1. The PR's acceptance criterion is STT ROUTING, but no committed test
called _event_media_is_stt_input — it was only asserted ad-hoc. Adds
TestSttGate: voice→eligible, voice-without-media_types→eligible
(the new-connector/old-gateway shape), legacy audio-typed voice
note→not eligible, music→not eligible. Mutation-verified: removing
the VOICE branch from the gate turns these RED.
2. media_urls and media[] are INDEPENDENT wire fields that consumers
index by the same i. Mapping MIMEs positionally without checking
agreement means a disagreeing producer misassociates a MIME with the
wrong URL and mis-routes that attachment — strictly worse than no
MIME, which degrades safely to message-level classification.
_media_types_from_wire() now maps only when the lengths agree, warns
and returns [] otherwise.
Note for the record: MessageType.VOICE predates this PR and the gate's
VOICE branch ignores media_types, so a NEW connector against an OLD
gateway ALREADY fires STT. That is desirable, but it is not "unchanged"
— the PR body's rollout matrix said otherwise and is corrected.
* fix(relay): send a User-Agent on relay media requests (Discord CDN 403)
Discord's CDN rejects urllib's default "Python-urllib/x.y" User-Agent
with HTTP 403, and RelayMediaClient never set one. Every Discord CDN
pass-through download therefore failed; _localize_inbound_media then
kept the raw URL (its "a public URL still has value" branch), and the
consumer tried to open a URL as a FILE PATH:
WARNING gateway.relay.media: relay media download failed for
https://cdn.discordapp.com/...voice-message.ogg: HTTP Error 403
INFO gateway.run: Voice transcription failed for https://cdn.discord...
: Audio file not found: https://cdn.discordapp.com/...
This killed ALL Discord relay media inbound — voice notes, images and
documents alike — not just the voice lane. Telegram/WhatsApp were
unaffected because their media is connector-re-hosted (/relay/media/{id},
fetched from our own host) and localizes to real /tmp paths.
Reproduced from a clean shell against a live CDN URL:
curl (own UA) -> 200
urllib, no UA -> 403 Forbidden
urllib + descriptive UA -> 200, 14583 bytes, OggS magic
Fix: a module-level _MEDIA_USER_AGENT sent on both download() and
upload(). upload() only ever targets our own connector so it was not
broken, but a single client should identify itself consistently.
Validated on staging: hot-patched hermes-agent-stg-test-6698, restarted
the gateway service, and Ben's Discord voice note transcribed
successfully — zero new 403s and zero new transcription failures after
the patch (last 403 predates it).
Test is mutation-verified: removing the UA from download() turns it RED
while the other five media tests stay green.
* fix(relay): keep url↔mime pairing through media localization
Addresses a blocking review finding on my own change: mapping media[]
into media_types created a POSITIONAL contract that the rest of the
inbound path then broke.
1. _localize_inbound_media (adapter.py) filtered media_urls without
filtering media_types. Dropping a dead connector re-host is a NORMAL
best-effort path, so every surviving attachment inherited its
neighbour's mime. Reproduced through the real functions:
before urls [.../relay/media/dead, .../kept.png]
types [application/pdf, image/png]
after urls [.../kept.png]
types [application/pdf, image/png] <-- PNG reads as PDF
_event_media_is_image(ev, 0) -> False
The loop now carries (url, mime) as PAIRS, so a dropped URL drops its
mime with it.
2. _media_types_from_wire compared LENGTHS only, which is not alignment:
equal-length-but-reordered wire fields were accepted and paired
wrongly, and an absent media_urls skipped the check entirely while
still emitting types. Resolution is now BY URL (url -> mime lookup
over media_urls); an unmatched URL degrades to "" and falls back to
message-level classification.
Tests: 4 new cases driving the real chain (wire parse -> localization ->
run.py classifier), incl. the dropped-first-attachment case the existing
localization test could not catch (it builds events without
media_types). The obsolete length-mismatch test now asserts the stronger
by-url guarantee. Both fixes mutation-verified: reinstating the URL-only
filter fails 1 test, reverting to positional resolution fails 3.
Relay suite 258 passed; media/voice/stt selection 685 passed; ruff clean;
cross-repo integration payload re-verified.
* fix(relay): media_types is always one slot per media_url
Self-review after two review rounds flagged this bug class in adjacent
seams: I checked the function I edited, not every consumer of the
parallel arrays I created. Grepping ALL writers found a third instance.
merge_pending_message_event (gateway/platforms/base.py:2725-2735)
EXTENDS media_urls and media_types together when a second media message
merges into a pending one. My mapping could emit a POPULATED media_urls
with an EMPTY media_types (an older connector sends media_urls but no
media[]), so extend() concatenated lists of different lengths:
A urls [old1.png, old2.png] types []
B urls [new.pdf] types [application/pdf]
merged urls [old1.png, old2.png, new.pdf]
types [application/pdf]
-> old1.png reads as application/pdf; the real PDF gets ''
Fix: media_types is now ALWAYS len(media_urls), padded with '' — the
url-keyed lookup runs even when media[] is absent, and the localizer
rewrites the list unconditionally (no short-circuit that
could leave a stale/short list behind).
Tests: 4 new cases — padding with no media[], the merge shift above
driven through the real merge_pending_message_event, localization
preserving the invariant while dropping an entry, and normalization of
a short/empty media_types arriving from a non-wire source. All
mutation-verified: removing the padding fails 4; restoring the
guard fails 1.
Relay 262 passed; media/voice/stt selection 689 passed; ruff clean;
cross-repo integration payload re-verified.
The reconcile-to-guarded-model interaction test's requestGateway mock
must only answer config.set with the confirm handshake — the panel's
model.options read rides the same dispatcher and was eating the
first mocked response.
The Bots editor's model write (profiles.configure) was the one switch
surface that bypassed the data-policy / expensive-model selection guard:
a guarded pick (e.g. muse-spark contributor tier) was applied silently,
with no confirm flow anywhere — the #95293 remainder after the core
picker's confirm handshake landed in use-model-controls.
Gateway: profiles.configure now answers confirm_required +
confirm_message for a guarded model (same handshake as config.set
model) and writes NOTHING until the client resends with
confirm_expensive_model: true. Other sections still apply; the pending
model section is not reported as failed.
Desktop: the confirm flow is extracted out of use-model-controls into
one shared applier (lib/guarded-model-switch.ts, exported through the
plugin SDK) — warning toast, staleness-guarded Confirm, single
confirmed resend, never a retry loop. The core picker and the Bots
editor now consume the SAME handler; the Bots editor's Confirm resends
only the model section with confirm_expensive_model: true.
Fixes#95293 (Bots surface remainder).
Refresh Models only updated the catalog cache, so the composer kept showing a model that was no longer in the new group list.
Co-authored-by: Cursor <cursoragent@cursor.com>
The Bots model picker's catalog read rode the bot's own socket with no
deadline: a wedged dial left the query pending forever, so the picker
spun indefinitely. On top of that every fetch forced refresh:true,
bypassing the staleTime cache, so each Bots view remount (tab re-front,
dialog reopen, pane visibility flip) knocked the picker back into its
loading state and wiped the staged provider/model pick mid-edit.
Bound every attempt to 20s (rejection falls through to the picker's
existing free-text fallback), drop the forced refresh so the read
participates in the cache like every other surface's catalog, and pin
the contract with a red->green regression suite.
Slim renderer UI for the managed SSH remote update engine (#95942),
adapted from #93042's renderer unit with the deferred canary/rollout
scope stripped. Adds a per-connection store (idle/updating/terminal
states, receipt, managed-update-in-progress busy envelope) and a
'Managed updates' section on the Gateways settings page with an Update
button, progress line, and correlated receipt per registered
Desktop-managed SSH connection. Fails closed when the Electron main
lacks connections.updateManaged.
The renderer pings each pool backend every 60s (`hermes:backend:touch` →
`touchPoolBackend` → updates `lastActiveAt`). The LRU eviction cap used a
keepalive-fresh window of 90s — only 1.5× the ping cadence — to decide
whether a backend was "plausibly still alive". One missed or delayed ping
pushed a live backend past the threshold and the cap-driven eviction killed
the active profile's backend mid-session, restarting the gateway and
re-minting runtime ids. On WSL2, where the renderer→Electron IPC roundtrips
through 9p, brief 9p hiccups commonly stretch a ping to seconds of observed
silence, producing the ~80–90s exit / ~2 min cycle reported in #95189
(122 gateway starts on 2026-08-26 alone, driving renderer OOM via reconnect
churn at ~5GB/day).
Widen POOL_KEEPALIVE_FRESH_MS to 4 minutes (3× ping cadence + IPC stall
headroom, still bounded well below POOL_IDLE_MS=10min). Backends with one or
even two missed pings are now spared; truly idle backends (multiple lapses,
minutes idle) remain eligible for eviction by the cap and the idle reaper.
The constant is also overridable via HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS
to make this tunable without a rebuild.
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:
- _is_actionable_user_turn (tail anchor) only checked role/content, so a
notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
operational notices leaked into the compact focus hint.
Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.
Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.
Fixes#92703
Group rooms persist source-qualified members. After Desktop switches to
the built-in This-device source, a dead loopback row still listed next
to the live profile and looked like a second agent. Collapse only the
sidebar tiles; $lastRoster, group seats, and mentions keep every
(connectionId, profile) identity.