Salvage adjustments to PR #94392 per review:
- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
recovery child now probes 'systemctl --user is-active' after each relaunch;
only an observed-active systemd unit is reported 'verified'. A relaunch that
merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
update_inventory serve collector) are no longer silently skipped: the
recovery pass records them (and manual gateways) as skipped-with-reason in
the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
(requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
fresh interpreter (sitecustomize shim intercepts the grandchild
'gateway restart' and systemctl probes).
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.
clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).
Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
* refactor(clarify): halve the schema (880->436 tok/call) — same rules, half the words
* refactor(clarify): one question field — questions[] is the only advertised shape (single = one-entry array; legacy shape stays handler-accepted)
* refactor(clarify): unadvertise per-question id — response rows already carry question text + order
The repo's own override pinned nanoid@3.3.17 — the exact version
GHSA-2v37-7h3g-55p8 / CVE-2026-67213 flags (custom generators loop
indefinitely on size 0) — so every fresh install and every npm audit
shipped/reported the vulnerable pin no matter what transitives wanted.
3.3.18 is the patched release on the same major. Lockfile re-resolved;
npm audit now reports zero nanoid findings, and 3.3.18's zero-size
generator returns instead of hanging (verified live).
Addresses the upstream-pin quarter of #91931 (mechanism 1 of the
reporter's four); the updater-side skip/verify mechanisms and the
uv.lock staleness half (#91424) remain tracked there.
Addresses the two follow-up notes from review: document that
pre_restart_pids is a bare PID set (not (pid, start_time) pairs), so a
recycled PID from one gateway landing in another's stale record could
still mislabel it as down; and add a companion test asserting a
matching start_time still yields the live/current row.
collect_fleet_versions()'s gateway_state.json fallback path only checked
_pid_exists(pid) to decide whether a recorded gateway was still running.
On Windows, a paused gateway's PID can be recycled by an unrelated
process spawned during the update's own churn (npm/git/python
subprocesses) before the record is refreshed, so the dead gateway's
stale code_sha still gets compared against HEAD and reported STALE for a
PID that no longer belongs to it (#93258).
Switch to runtime_status_pid_is_live(), the existing (pid, start_time)
PID-reuse guard already used elsewhere in gateway/status.py, so a
recycled PID is treated the same as a dead one (DOWN row, or no row, per
the existing rollout-safety rules) instead of a false STALE.
#91277 Phase 2's plan-vs-execution reconciliation (match_runtime_outcomes)
cross-checks every runtime collect_runtime_inventory() saw against
restarted_services / relaunched_profiles / externally_supervised_profiles /
killed_pids — the systemd/launchd restart phase's bookkeeping. That
inventory is cross-platform (control-socket / PID-file based), so it
includes Windows gateways too, but Windows's own pause/resume mechanism
(_pause_windows_gateways_for_update / _resume_windows_gateways_after_update)
never wrote into any of that bookkeeping.
Result: a Windows gateway that was correctly stopped and relaunched by
_resume_windows_gateways_after_update was still classified "unaccounted" by
the reconciliation (the plan saw it and no bookkeeping mentions it) —
report_unaccounted_runtimes() escalates that into sys.exit(1), and in
gateway_mode also writes ".update_exit_code"="1". Every successful
`hermes update` on Windows with a running gateway reported itself as
failed, unconditionally (the sys.exit(1) is not gated to gateway_mode).
_resume_windows_gateways_after_update now records the profiles it
successfully relaunched onto the resume token; _cmd_update_impl merges
that into the shared relaunched_profiles list right before reconciliation
runs. A profile whose relaunch genuinely fails is deliberately left off
the list, so it still surfaces as unaccounted — Windows has no watcher to
recover a failed relaunch, so that escalation is the correct signal.
Regression tests exercise _resume_windows_gateways_after_update directly
(records successes, omits failures) and reproduce the reconciliation-level
bug end to end: the same plan row resolves "unaccounted" without the merge
and "restarted" with it. Mutation-verified: with the fix reverted, three of
the four new tests fail (KeyError on the token / wrong outcome).
The wrapper now getattr-defaults _processed_message_ts (object.__new__
adapters in sibling suites lack it), and the reaction-guard source pin
reads _handle_slack_message_impl where the production expression lives.
Follow-up for the #95417 salvage, addressing the review finding: the entry
claim closes the unfurl race but a handler that raises mid-enrichment would
hold the claim forever — neither a Slack retry nor a user edit could ever
re-drive the message. _handle_slack_message is now a thin guard around the
impl that releases only claims taken by the failed invocation itself, with a
warning log so swallowed turns are traceable. Pre-existing claims from a
successful turn are never released. Two failure-path tests pin both sides.
Slack emits `message_changed` for a link unfurl carrying a DIFFERENT event ts
than the original message. That ts legitimately misses the `_dedup` check, so
`_processed_message_ts` is the only guard against it becoming a second user
turn -- but it was only populated at the END of `_handle_slack_message`, after
thread context, permalink resolution and file downloads had all awaited.
An unfurl landing inside that window found the guard empty and was promoted to
a duplicate turn: a spurious "Interrupting current task" banner plus the same
answer posted twice.
Production capture (adminbot, 2026-08-22 02:23:30-31Z, channel C0BF1EYUA9H):
02:23:30.718 message ts=1787365409.908499 dedup_hit=False
02:23:30.737 app_mention ts=1787365409.908499 dedup_hit=True
02:23:31.675 message ts=1787365411.012100 dedup_hit=False <- leaked
subtype=message_changed
The original copy was still resolving two Slack permalinks when the unfurl
arrived 957ms later.
Claim the message ts once every filter has passed and the event is certain to
be delivered, before the slow enrichment awaits. Claiming any earlier (right
after the dedup check) also claims messages the handler then discards, which
breaks summoning the bot by editing "@bot" into a previously ignored message
(tests/gateway/test_slack.py::TestMessageRouting::
test_message_edit_with_new_mention_processed).
Eviction logic is extracted to `_remember_processed_message_ts` so both call
sites share one bounded implementation.
Two profiles can hold sessions with the SAME stored id (restored
backups, copied state.dbs, cross-profile imports). mergeSessionPage
keyed rows by bare id, so the twins collapsed into one sidebar row
whose title/activity carry stitched one profile's content onto the
other's route — clicking a row previewing profile A resumed profile B
and wedged on an eternal 'Waking up…' when the mismatched resume never
completed (live-reproduced on main).
- mergeSessionPage: identity + lineage + dedupe keys are now
(profile, id); a kept twin in another profile survives the incoming
page dedupe. Local profile normalize (importing @/store/profile would
be circular).
- Sidebar clicks carry the ROW as the identity: onResumeSession passes
the clicked SessionInfo, and wiring pins the row's own
(connection, profile) as the resume owner via requestSessionResume
before navigating. Untagged rows keep the id-only path.
A backend.lock.json that exists but doesn't match what this build writes
(unknown/future schemaVersion, truncated JSON, missing or foreign
ownershipId, malformed shape) was previously indistinguishable from 'no
lockfile': connect() would spawn a fresh backend on top of it and
overwrite the record, and cleanup paths could drop foreign state —
disarming the #78872 ownership guard exactly when another (e.g. forked)
desktop build shares the remote. readLockfile now returns a skew
sentinel for existing-but-foreign lockfiles; connect() refuses with a
'remote-lockfile-skew' error and a skew warning instead of
reaping/overwriting, disconnect() and cleanupStale() skip entirely.
Refs #95532
The Electron shell boots by fetching / and extracting
window.__HERMES_SESSION_TOKEN__ to authenticate /api/ws
(dashboard-token.ts adoptServedDashboardToken). Headless serve 404'd
every path, so when the renderer's spawn token drifted from the
backend's live token — e.g. hermes update replaced the backend and the
env pin no longer matched — the renderer had no way to adopt the served
token, the WebSocket handshake failed, and the primary window
white-screened (#95575).
Serve a minimal token-only HTML page at the exact root path in
mount_spa()'s headless branch, matching the renderer's extraction regex.
Gate it on app.state.auth_required read at request time: a gated
(non-loopback / remote public_url) serve keeps returning the 404 JSON so
the session token never leaks past the loopback boundary. Every other
path stays 404 JSON — the SPA remains unserved.
Regression tests: TestHeadlessServeTokenPage (3 cases) — verified to
fail against the pre-fix headless branch.
JSON.stringify does not escape '<' — a reloadUrl containing
'</script><script>…' would terminate the inline <script> element of the
data: error page and let an attacker-controlled URL inject markup/script.
Escape <, >, & (and U+2028/U+2029) as \uXXXX sequences after stringify;
add regression test.
A torn renderer bundle (update replaced the app while its files were
locked, e.g. antivirus or a still-running instance) loads index.html
fine and then dies on the first lazy import — a white screen with only
a desktop.log line. A main-frame load failure (missing index.html,
blocked file) was likewise log-only.
- resolveRendererIndex() already detects torn bundles; the primary
window now refuses to load one and shows a visible repair page
(error code, missing assets, 'hermes desktop --force-build', Reload)
instead of a blank window.
- did-fail-load on the main frame now gets bounded auto-reload through
the shared rolling reload budget (transient failures self-heal) and,
once the budget is exhausted, surfaces the visible error page.
ERR_ABORTED and sub-frame failures stay log-only, and helper windows
(OAuth/portal) keep their log-only policy (opt-in via
reloadOnFailedLoad).
Regression tests cover the policy decisions (reload / abort /
budget-exhausted surface), budget sharing with render-process-gone,
and the error page content + data: URL loading.
fal's post-trained H3 variant — #1-ranked quality/prompt adherence/
aesthetics, 5s 768p video in under 3 seconds, $0.04/s launch pricing.
- New minimax-h3-max family: minimax/h3-max/{text,image}-to-video
- Inherits base-H3 wire quirks (integer duration, i2v drops
aspect_ratio) but caps at 768P (480P/768P enums, no 2K/4K) and
declares seed on both endpoints
- New generic static_payload family flag: constant keys the endpoint
requires on every request (H3 Max lists prompt_expansion_mode in its
required array; sent as 'balanced')
Payload asserted against the endpoint OpenAPI schema; 73/73 targeted
tests green (surface matrix auto-covers the new family).
hermes_cli/memory_setup.py::_write_env_vars() wrote provider-controlled
.env entries with a direct Path.write_text() + post-hoc chmod, bypassing
the denylist/regex/CRLF-stripping/atomic-replace validation that
hermes_cli/config.py::save_env_value() already provides for every other
.env writer in the codebase. A malicious or buggy memory-provider plugin
declaring a crafted env-var name/value in its setup schema could inject
arbitrary lines into .env.
Routes memory-provider env writes through save_env_value(), and fixes a
regression this surfaced in plugins/memory/supermemory/__init__.py::
post_setup(), which called the old two-parameter _write_env_vars(env_path,
values) signature — restores the caller via context-local
hermes_constants.set_hermes_home_override()/reset_hermes_home_override()
instead of a removed env_path parameter, so explicit HERMES_HOME overrides
during setup still resolve correctly.
Adds test_env_file_created_with_secure_permissions, guarded on Windows
(POSIX mode bits aren't enforced there, mirroring the existing skip in
test_openviking_provider.py / test_supermemory_provider.py) since
save_env_value's atomic-replace path creates the temp file at 0o600 before
writing content, closing the TOCTOU window the old direct-write + chmod
implementation had.
Deliver changes_requested review outcomes through kanban subscriptions and
wake the origin for notify+wake / wake modes. Review-specific only: no task
mutation. Reasons are redacted, path-scrubbed and truncated before delivery.
Salvaged from #88694; conflicts with the #87733 wake-kinds expansion resolved
keep-both.
* refactor(skills): diet skill_manage schema (924->567 tok/call) — drop prose duplicated in system prompt + validation errors
* refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)
* Revert "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"
This reverts commit 9988b33f177fa82a145d5243d556f113a7e2fd2b.
* Reapply "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"
This reverts commit 49a0a8f18b478824398767d7ae681913cee00640.
* refactor(skills): unadvertise absorbed_into — curator-only vocabulary leaves the shared schema
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".
Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
Tailwind v4 wraps hover:/group-hover: in @media (hover: hover). Windows
hosts with a digitizer often answer false even with a mouse, so those
controls stay opacity-0 — clickable, invisible. Trust :hover itself.
Co-authored-by: xxxigm <54813621+xxxigm@users.noreply.github.com>
The live thinking body pins to the newest tokens and keeps max-h-40
after the turn settles so the transcript doesn't jump. overflow-hidden
let that pin work in JS but clipped the rest of the thought. overflow-auto
makes the cap a real scroller; overscroll-contain keeps the wheel from
chaining into the transcript.
Supersedes #73757.
Co-authored-by: Dan Latimer <latdani@gmail.com>
The OAuth complete-with-model path is headless: a confirm toast would
hang with nobody to click it. Skip the prompt and surface the backend
message instead.
Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
Settings → Model → Apply treated confirm_required as a red error, so
contributor-tier models like muse-spark-1.2-contributor could never be
saved. Prompt and retry with confirm_expensive_model, matching the
in-session picker handshake.
Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
`main` is red. 68518c1f9b added a dedicated `renderTypingSync` harness for the
typing-aware deferral tests, but its params object omits `updateSessionState`,
which `BackgroundSyncParams` requires:
src/app/contrib/hooks/use-background-sync.test.ts(660,25): error TS2345:
Property 'updateSessionState' is missing in type '{ ... }' but required
in type 'BackgroundSyncParams'.
`npm run typecheck` exits 2, so `apps/desktop :: check:lint` fails and the
JS & TS checks job goes red on every open PR regardless of its contents.
Vitest did not catch this because it transpiles without typechecking, so the
new suite passes while `tsc -p .` fails.
Supply the missing prop. The updater runs against a throwaway state — this
harness never exercises the transcript path — and it is added to the existing
`stable` object rather than inline, because the harness's own comment requires
every param to keep a stable identity across the tick-driven re-renders; a
fresh `vi.fn()` per render would re-run the connect-reseed effect and
re-subscribe the throttle, polluting the very counts the tests observe.
Verified: all three tsc projects clean (`tsc -p .`, `tsconfig.electron.json`,
`tsconfig.e2e.json`), use-background-sync suites 26/26, eslint clean.
The starvation cap in the original patch re-ran the heavy pass at the
same ~10s mark the freeze is measured at. Hold until the keyboard is
quiet, then land one coalesced pass.
new BrowserWindow({ icon }) and app.dock.setIcon() decode the icon file
synchronously on the main process and throw on undecodable bytes. The
icon ladder was resolved with statSync().isFile(), which only proves a
file exists — a truncated or zero-byte PNG inside a packaged app.asar
(interrupted electron-builder run, partial copy) killed the main process
inside createWindow(): the window never appeared, running turns lost
their renderer, and the desktop log showed 'Uncaught exception: Error:
Failed to load image from path .../app.asar/public/apple-touch-icon.png
at createWindow'.
Resolution now runs through a decoding probe (nativeImage.createFromPath
must yield a non-empty image); a candidate that exists but does not
decode is skipped like a missing one, so the app falls through to the
next rung or starts with the platform default icon instead of dying.
The ladder and probe live in a pure module (electron/app-icon.ts) so
precedence is unit-testable without a running Electron app; window
factories re-resolve per call exactly as before.
Regression tests cover: skip-first-undecodable, all-fail -> undefined,
first-pass wins, missing/empty/directory rejection, and the unchanged
mac/Windows precedence ladder.
Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
The apt/docker CLI tests pinned exit 1; refusals are now exit 2
(refused-by-contract, distinct from errors). The web_server guards
patched the module-local detect_install_method alias, which the shared
admission gate no longer consults — patch hermes_cli.config directly.
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.
A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.
Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
Cherry-picked core of #92545: the image build writes a versioned,
non-secret marker (/etc/hermes/image-provenance.json) outside both the
bind-mountable checkout and the HERMES_HOME volume, and
hermes_cli/image_provenance.py reads it fail-closed — absence means
'not image-managed', any present-but-malformed marker still means
image-managed (an integrity defect is never permission to mutate the
image in place).
(#91277 Phase 3; salvaged from #92545 by @andrexibiza — marker bake +
reader only, the scoped carve-out.)
* fix(tui_gateway): ask before queueing a guarded model picked mid-turn
config.set model on a running session cannot swap the agent in place, so it
stashes the pick in session["pending_model_switch"] and applies it at the
next turn start. That branch answered confirm_required=False without ever
running the selection guards.
A client that implements the confirm round-trip was therefore told no
consent was needed and never prompted. One turn later
_apply_pending_model_switch ran the guards with the stashed (unconfirmed)
flag, saw the warning, and dropped the switch by design. The model reverted
with no confirm ever offered, because the only moment a round-trip was
possible had already passed.
Evaluate the guards before stashing, where the client still has a live
response to turn into a prompt. Nothing is queued for an unconfirmed
guarded pick, so the session is left exactly as it was and the re-send
carrying confirm_expensive_model queues it for real. The apply-time check
stays as the backstop for guards that can only decide after resolution.
The data-policy guard keys on the model id alone, which is all this branch
can see before resolution. The cost guard returns None when pricing is
unknown and its models.dev lookup is allow_network=False, so calling it
early can only under-fire and never blocks the RPC thread.
* test(tui_gateway): pin provider forwarding, name the canonical confirm field
Two review follow-ups, no behavior change.
_pending_switch_selection_warning forwards `provider=provider or None`, but
nothing asserted it: a guarded model id fires the data-policy guard on the
model alone, so the existing tests passed with `provider` dropped entirely.
Record the kwargs instead. Dropping the argument fails the first test;
removing the `or None` normalization fails the second.
The confirm responses carry `warning` and `confirm_message` with identical
text, which reads like an accident. Name which one clients should read
(`confirm_message`; `warning` is the pre-confirm-era alias that
_apply_pending_model_switch already treats as a fallback) so the two do not
drift apart later.
Both raised by @Enough1122 in review.
* fix(desktop): keep attachment close and code copy icons visible
Hover-only opacity-0 hid the composer remove control and code-block copy button, so they stayed clickable but invisible on Windows and other no-hover surfaces.
* test(desktop): pin attachment close and code copy visibility at rest
The remove chip and code-block copy control must stay in the tree without a hover class, so Windows and no-hover surfaces cannot hide them again.
slack_bolt builds a fresh AsyncWebClient for every inbound request and
copies proxy=app.client.proxy into its constructor, where slack_sdk reads
a None/blank proxy *argument* as "unspecified" and reloads HTTP(S)_PROXY
from the environment. aiohttp then treats that env value as an explicit
proxy and skips its own NO_PROXY check, so the adapter's resolved decision
to go direct - a NO_PROXY bypass, or a proxy scheme aiohttp cannot use -
holds on every client except the one authorization spends on auth.test.
The failure looks like a healthy bot: Socket Mode connects, outbound sends
keep working, and every inbound event is rejected with "Failed to authorize
with the given token" - forever, since a failed auth_test_result is not
cached and never retried differently.
Re-apply the resolved proxy through AsyncApp(before_authorize=...), which
bolt inserts before the authorization middleware: the request-scoped client
already exists there and has not been used yet. Assigning the attribute
post-construction is the only way to express "no proxy" to slack_sdk.
Co-authored-by: Junie <junie@jetbrains.com>
`completed` already puts the worker's summary inside the synthetic wake
turn, so the woken creator sees what was done. `review_requested` did
not: the summary rode the passive ping only, and the wake turn said just
"handed off for review", forcing the woken reviewer to re-read the board
(and losing the PR link the worker had already written).
Reuse the same first-line handoff the `completed` branch builds, so the
existing `gateway.kanban.wake.handoff` string renders it — no new locale
keys, no change to the passive message.