Commit Graph

25618 Commits

Author SHA1 Message Date
Teknium b2c011364e fix(update): conservative outcomes + serve-ledger coverage for fresh restart recovery
Salvage adjustments to PR #94392 per review:

- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
  recovery child now probes 'systemctl --user is-active' after each relaunch;
  only an observed-active systemd unit is reported 'verified'. A relaunch that
  merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
  coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
  update_inventory serve collector) are no longer silently skipped: the
  recovery pass records them (and manual gateways) as skipped-with-reason in
  the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
  (requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
  fresh interpreter (sitecustomize shim intercepts the grandchild
  'gateway restart' and systemctl probes).
2026-08-26 16:45:26 -07:00
joaomarcos f9135c189c fix(update): persist fresh recovery outcome 2026-08-26 16:45:26 -07:00
joaomarcos ccdd7f41ed fix(update): persist per-profile recovery outcomes 2026-08-26 16:45:26 -07:00
joaomarcos f0045c5383 fix(update): verify fresh restart recovery results 2026-08-26 16:45:26 -07:00
joaomarcos 5609ccbece fix(update): recover aborted gateway restart in a fresh process 2026-08-26 16:45:26 -07:00
Teknium df3d41ee67 fix(update): sweep aborted-fetch tmp_pack debris before it corrupts the pack directory (#93732)
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.

clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).

Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
2026-08-26 16:45:22 -07:00
Teknium 7d6c6ae4ae refactor(clarify): schema diet + single questions[] interface (880 → 335 tok/call, −62%) (#95907)
* refactor(clarify): halve the schema (880->436 tok/call) — same rules, half the words

* refactor(clarify): one question field — questions[] is the only advertised shape (single = one-entry array; legacy shape stays handler-accepted)

* refactor(clarify): unadvertise per-question id — response rows already carry question text + order
2026-08-26 16:42:44 -07:00
hermes-seaeye[bot] 77001a6be7 fmt(js): npm run fix on merge (#95924)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 23:21:02 +00:00
Teknium 8fdda828a8 fix(deps): bump the nanoid@^3 override past GHSA-2v37-7h3g-55p8 (#91931)
The repo's own override pinned nanoid@3.3.17 — the exact version
GHSA-2v37-7h3g-55p8 / CVE-2026-67213 flags (custom generators loop
indefinitely on size 0) — so every fresh install and every npm audit
shipped/reported the vulnerable pin no matter what transitives wanted.
3.3.18 is the patched release on the same major. Lockfile re-resolved;
npm audit now reports zero nanoid findings, and 3.3.18's zero-size
generator returns instead of hanging (verified live).

Addresses the upstream-pin quarter of #91931 (mechanism 1 of the
reporter's four); the updater-side skip/verify mechanisms and the
uv.lock staleness half (#91424) remain tracked there.
2026-08-26 16:14:51 -07:00
chelsealong 2812d6121b fix(cli): note pre_restart_pids' per-PID data model gap, pin the matching-start_time path
Addresses the two follow-up notes from review: document that
pre_restart_pids is a bare PID set (not (pid, start_time) pairs), so a
recycled PID from one gateway landing in another's stale record could
still mislabel it as down; and add a companion test asserting a
matching start_time still yields the live/current row.
2026-08-26 16:14:32 -07:00
chelsealong 5d7ed70eef fix(cli): guard the post-update fleet check against PID reuse
collect_fleet_versions()'s gateway_state.json fallback path only checked
_pid_exists(pid) to decide whether a recorded gateway was still running.
On Windows, a paused gateway's PID can be recycled by an unrelated
process spawned during the update's own churn (npm/git/python
subprocesses) before the record is refreshed, so the dead gateway's
stale code_sha still gets compared against HEAD and reported STALE for a
PID that no longer belongs to it (#93258).

Switch to runtime_status_pid_is_live(), the existing (pid, start_time)
PID-reuse guard already used elsewhere in gateway/status.py, so a
recycled PID is treated the same as a dead one (DOWN row, or no row, per
the existing rollout-safety rules) instead of a false STALE.
2026-08-26 16:14:32 -07:00
pierrenode de2a9de788 fix(update): feed Windows gateway relaunch outcome into fleet reconciliation
#91277 Phase 2's plan-vs-execution reconciliation (match_runtime_outcomes)
cross-checks every runtime collect_runtime_inventory() saw against
restarted_services / relaunched_profiles / externally_supervised_profiles /
killed_pids — the systemd/launchd restart phase's bookkeeping. That
inventory is cross-platform (control-socket / PID-file based), so it
includes Windows gateways too, but Windows's own pause/resume mechanism
(_pause_windows_gateways_for_update / _resume_windows_gateways_after_update)
never wrote into any of that bookkeeping.

Result: a Windows gateway that was correctly stopped and relaunched by
_resume_windows_gateways_after_update was still classified "unaccounted" by
the reconciliation (the plan saw it and no bookkeeping mentions it) —
report_unaccounted_runtimes() escalates that into sys.exit(1), and in
gateway_mode also writes ".update_exit_code"="1". Every successful
`hermes update` on Windows with a running gateway reported itself as
failed, unconditionally (the sys.exit(1) is not gated to gateway_mode).

_resume_windows_gateways_after_update now records the profiles it
successfully relaunched onto the resume token; _cmd_update_impl merges
that into the shared relaunched_profiles list right before reconciliation
runs. A profile whose relaunch genuinely fails is deliberately left off
the list, so it still surfaces as unaccounted — Windows has no watcher to
recover a failed relaunch, so that escalation is the correct signal.

Regression tests exercise _resume_windows_gateways_after_update directly
(records successes, omits failures) and reproduce the reconciliation-level
bug end to end: the same plan row resolves "unaccounted" without the merge
and "restarted" with it. Mutation-verified: with the fix reverted, three of
the four new tests fail (KeyError on the token / wrong outcome).
2026-08-26 16:14:27 -07:00
Teknium 7a7a371c59 fix: harden claim-release guard for bare test doubles; repoint source-pinning test at the impl
The wrapper now getattr-defaults _processed_message_ts (object.__new__
adapters in sibling suites lack it), and the reaction-guard source pin
reads _handle_slack_message_impl where the production expression lives.
2026-08-26 15:54:53 -07:00
Teknium 39a5838f07 fix(slack): release a failed handler's fresh ts claim so the turn isn't swallowed
Follow-up for the #95417 salvage, addressing the review finding: the entry
claim closes the unfurl race but a handler that raises mid-enrichment would
hold the claim forever — neither a Slack retry nor a user edit could ever
re-drive the message. _handle_slack_message is now a thin guard around the
impl that releases only claims taken by the failed invocation itself, with a
warning log so swallowed turns are traceable. Pre-existing claims from a
successful turn are never released. Two failure-path tests pin both sides.
2026-08-26 15:54:53 -07:00
Richard Hojun Jang 708f84c477 fix(slack): claim message ts before enrichment so link unfurls can't duplicate a turn
Slack emits `message_changed` for a link unfurl carrying a DIFFERENT event ts
than the original message. That ts legitimately misses the `_dedup` check, so
`_processed_message_ts` is the only guard against it becoming a second user
turn -- but it was only populated at the END of `_handle_slack_message`, after
thread context, permalink resolution and file downloads had all awaited.

An unfurl landing inside that window found the guard empty and was promoted to
a duplicate turn: a spurious "Interrupting current task" banner plus the same
answer posted twice.

Production capture (adminbot, 2026-08-22 02:23:30-31Z, channel C0BF1EYUA9H):

  02:23:30.718  message      ts=1787365409.908499  dedup_hit=False
  02:23:30.737  app_mention  ts=1787365409.908499  dedup_hit=True
  02:23:31.675  message      ts=1787365411.012100  dedup_hit=False   <- leaked
                subtype=message_changed

The original copy was still resolving two Slack permalinks when the unfurl
arrived 957ms later.

Claim the message ts once every filter has passed and the event is certain to
be delivered, before the slow enrichment awaits. Claiming any earlier (right
after the dedup check) also claims messages the handler then discards, which
breaks summoning the bot by editing "@bot" into a previously ignored message
(tests/gateway/test_slack.py::TestMessageRouting::
test_message_edit_with_new_mention_processed).

Eviction logic is extracted to `_remember_processed_message_ts` so both call
sites share one bounded implementation.
2026-08-26 15:54:53 -07:00
Teknium db127f7502 fix(desktop): session rows are identified by (profile, id) — twins in two profiles stop collapsing and mis-routing (#92454)
Two profiles can hold sessions with the SAME stored id (restored
backups, copied state.dbs, cross-profile imports). mergeSessionPage
keyed rows by bare id, so the twins collapsed into one sidebar row
whose title/activity carry stitched one profile's content onto the
other's route — clicking a row previewing profile A resumed profile B
and wedged on an eternal 'Waking up…' when the mismatched resume never
completed (live-reproduced on main).

- mergeSessionPage: identity + lineage + dedupe keys are now
  (profile, id); a kept twin in another profile survives the incoming
  page dedupe. Local profile normalize (importing @/store/profile would
  be circular).
- Sidebar clicks carry the ROW as the identity: onResumeSession passes
  the clicked SessionInfo, and wiring pins the row's own
  (connection, profile) as the resume owner via requestSessionResume
  before navigating. Untagged rows keep the id-only path.
2026-08-26 15:53:41 -07:00
Hermes Agent 3ee0c62224 fix(desktop): SSH orphan reaper fails CLOSED on remote lockfile schema/ownership skew
A backend.lock.json that exists but doesn't match what this build writes
(unknown/future schemaVersion, truncated JSON, missing or foreign
ownershipId, malformed shape) was previously indistinguishable from 'no
lockfile': connect() would spawn a fresh backend on top of it and
overwrite the record, and cleanup paths could drop foreign state —
disarming the #78872 ownership guard exactly when another (e.g. forked)
desktop build shares the remote. readLockfile now returns a skew
sentinel for existing-but-foreign lockfiles; connect() refuses with a
'remote-lockfile-skew' error and a skew warning instead of
reaping/overwriting, disconnect() and cleanupStale() skip entirely.

Refs #95532
2026-08-26 15:51:54 -07:00
Teknium 39a5aa91ed fix(serve): serve Desktop token page at / in headless mode (#94227)
The Electron shell boots by fetching / and extracting
window.__HERMES_SESSION_TOKEN__ to authenticate /api/ws
(dashboard-token.ts adoptServedDashboardToken). Headless serve 404'd
every path, so when the renderer's spawn token drifted from the
backend's live token — e.g. hermes update replaced the backend and the
env pin no longer matched — the renderer had no way to adopt the served
token, the WebSocket handshake failed, and the primary window
white-screened (#95575).

Serve a minimal token-only HTML page at the exact root path in
mount_spa()'s headless branch, matching the renderer's extraction regex.
Gate it on app.state.auth_required read at request time: a gated
(non-loopback / remote public_url) serve keeps returning the 404 JSON so
the session token never leaks past the loopback boundary. Every other
path stays 404 JSON — the SPA remains unserved.

Regression tests: TestHeadlessServeTokenPage (3 cases) — verified to
fail against the pre-fix headless branch.
2026-08-26 15:51:22 -07:00
Finn763 de34c746cc chore: map contributor emails 2026-08-26 15:51:22 -07:00
Finn763 c9d7b22e05 fix(desktop): escape reloadUrl in error page inline script (script-tag breakout)
JSON.stringify does not escape '<' — a reloadUrl containing
'</script><script>…' would terminate the inline <script> element of the
data: error page and let an attacker-controlled URL inject markup/script.
Escape <, >, & (and U+2028/U+2029) as \uXXXX sequences after stringify;
add regression test.
2026-08-26 15:51:22 -07:00
Finn763 4bc84d0499 fix(desktop): replace white screen after update with visible error + auto-reload (#95575)
A torn renderer bundle (update replaced the app while its files were
locked, e.g. antivirus or a still-running instance) loads index.html
fine and then dies on the first lazy import — a white screen with only
a desktop.log line. A main-frame load failure (missing index.html,
blocked file) was likewise log-only.

- resolveRendererIndex() already detects torn bundles; the primary
  window now refuses to load one and shows a visible repair page
  (error code, missing assets, 'hermes desktop --force-build', Reload)
  instead of a blank window.
- did-fail-load on the main frame now gets bounded auto-reload through
  the shared rolling reload budget (transient failures self-heal) and,
  once the budget is exhausted, surfaces the visible error page.
  ERR_ABORTED and sub-frame failures stay log-only, and helper windows
  (OAuth/portal) keep their log-only policy (opt-in via
  reloadOnFailedLoad).

Regression tests cover the policy decisions (reload / abort /
budget-exhausted surface), budget sharing with render-process-gone,
and the error page content + data: URL loading.
2026-08-26 15:51:22 -07:00
Teknium b0563d5ad3 feat: MiniMax H3 Max joins the FAL video picker (t2v + i2v)
fal's post-trained H3 variant — #1-ranked quality/prompt adherence/
aesthetics, 5s 768p video in under 3 seconds, $0.04/s launch pricing.

- New minimax-h3-max family: minimax/h3-max/{text,image}-to-video
- Inherits base-H3 wire quirks (integer duration, i2v drops
  aspect_ratio) but caps at 768P (480P/768P enums, no 2K/4K) and
  declares seed on both endpoints
- New generic static_payload family flag: constant keys the endpoint
  requires on every request (H3 Max lists prompt_expansion_mode in its
  required array; sent as 'balanced')

Payload asserted against the endpoint OpenAPI schema; 73/73 targeted
tests green (surface matrix auto-covers the new family).
2026-08-26 15:48:43 -07:00
pierrenode 6766732620 fix(memory-setup): route .env writer through save_env_value's validation gate
hermes_cli/memory_setup.py::_write_env_vars() wrote provider-controlled
.env entries with a direct Path.write_text() + post-hoc chmod, bypassing
the denylist/regex/CRLF-stripping/atomic-replace validation that
hermes_cli/config.py::save_env_value() already provides for every other
.env writer in the codebase. A malicious or buggy memory-provider plugin
declaring a crafted env-var name/value in its setup schema could inject
arbitrary lines into .env.

Routes memory-provider env writes through save_env_value(), and fixes a
regression this surfaced in plugins/memory/supermemory/__init__.py::
post_setup(), which called the old two-parameter _write_env_vars(env_path,
values) signature — restores the caller via context-local
hermes_constants.set_hermes_home_override()/reset_hermes_home_override()
instead of a removed env_path parameter, so explicit HERMES_HOME overrides
during setup still resolve correctly.

Adds test_env_file_created_with_secure_permissions, guarded on Windows
(POSIX mode bits aren't enforced there, mirroring the existing skip in
test_openviking_provider.py / test_supermemory_provider.py) since
save_env_value's atomic-replace path creates the temp file at 0o600 before
writing content, closing the TOCTOU window the old direct-write + chmod
implementation had.
2026-08-26 15:48:39 -07:00
Teknium 2f8d6b5552 fix(i18n): changes_requested + review_detail wake keys in all 17 locales
Follow-up for the #88694 salvage: the PR added the two new wake keys to
en.yaml only; main carries 17 locale files (parity rule).
2026-08-26 15:31:36 -07:00
DmytroVolodymyrson a699234f81 fix(kanban): wake controllers on review changes
Deliver changes_requested review outcomes through kanban subscriptions and
wake the origin for notify+wake / wake modes. Review-specific only: no task
mutation. Reasons are redacted, path-scrubbed and truncated before delivery.

Salvaged from #88694; conflicts with the #87733 wake-kinds expansion resolved
keep-both.
2026-08-26 15:31:36 -07:00
Teknium 3fd70ad1c9 refactor(skills): skill_manage 924 → 518 tok/call — dedup diet, retire 'edit', unadvertise curator-only absorbed_into (#95697)
* refactor(skills): diet skill_manage schema (924->567 tok/call) — drop prose duplicated in system prompt + validation errors

* refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)

* Revert "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"

This reverts commit 9988b33f177fa82a145d5243d556f113a7e2fd2b.

* Reapply "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"

This reverts commit 49a0a8f18b478824398767d7ae681913cee00640.

* refactor(skills): unadvertise absorbed_into — curator-only vocabulary leaves the shared schema
2026-08-26 15:13:30 -07:00
AlexGabbia b3e477f304 fix(update): wait for resumed Windows gateway before failing fleet check
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".

Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
2026-08-26 15:05:32 -07:00
hermes-seaeye[bot] d9e4b234a3 fmt(js): npm run fix on merge (#95861)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 21:50:14 +00:00
hermes-seaeye[bot] 74cb4cb80c fmt(js): npm run fix on merge (#95858)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 21:44:19 +00:00
ethernet a6d6060d61 desktop: remove js vestiges 2026-08-26 14:38:57 -07:00
Brooklyn Nicholson a2c3b54939 fix(desktop): put attachment close and code copy back on hover-reveal
#95611 painted those two always-on. With hover un-gated, restore the
corner reveal so they match the rest of the chrome.
2026-08-26 16:38:46 -05:00
Brooklyn Nicholson 2f87fa66ad fix(desktop): don't gate hover-reveal on the hover media query
Tailwind v4 wraps hover:/group-hover: in @media (hover: hover). Windows
hosts with a digitizer often answer false even with a mouse, so those
controls stay opacity-0 — clickable, invisible. Trust :hover itself.

Co-authored-by: xxxigm <54813621+xxxigm@users.noreply.github.com>
2026-08-26 16:38:46 -05:00
Brooklyn Nicholson be7eefec8d fix(desktop): restore scroll on the capped thinking preview
The live thinking body pins to the newest tokens and keeps max-h-40
after the turn settles so the transcript doesn't jump. overflow-hidden
let that pin work in JS but clipped the rest of the thought. overflow-auto
makes the cap a real scroller; overscroll-contain keeps the wheel from
chaining into the transcript.

Supersedes #73757.

Co-authored-by: Dan Latimer <latdani@gmail.com>
2026-08-26 16:38:40 -05:00
Brooklyn Nicholson 19f9d1badb fix(desktop): fail closed when onboarding hits a model guard
The OAuth complete-with-model path is headless: a confirm toast would
hang with nobody to click it. Skip the prompt and surface the backend
message instead.

Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
2026-08-26 16:38:22 -05:00
Brooklyn Nicholson 8fe4816edd fix(desktop): confirm guarded Settings model applies
Settings → Model → Apply treated confirm_required as a red error, so
contributor-tier models like muse-spark-1.2-contributor could never be
saved. Prompt and retry with confirm_expensive_model, matching the
in-session picker handshake.

Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
2026-08-26 16:38:22 -05:00
Trevor Nash-Keller 15b673d178 fix(desktop): restore the typecheck by completing the typing-sync harness
`main` is red. 68518c1f9b added a dedicated `renderTypingSync` harness for the
typing-aware deferral tests, but its params object omits `updateSessionState`,
which `BackgroundSyncParams` requires:

    src/app/contrib/hooks/use-background-sync.test.ts(660,25): error TS2345:
    Property 'updateSessionState' is missing in type '{ ... }' but required
    in type 'BackgroundSyncParams'.

`npm run typecheck` exits 2, so `apps/desktop :: check:lint` fails and the
JS & TS checks job goes red on every open PR regardless of its contents.

Vitest did not catch this because it transpiles without typechecking, so the
new suite passes while `tsc -p .` fails.

Supply the missing prop. The updater runs against a throwaway state — this
harness never exercises the transcript path — and it is added to the existing
`stable` object rather than inline, because the harness's own comment requires
every param to keep a stable identity across the tick-driven re-renders; a
fresh `vi.fn()` per render would re-run the connect-reseed effect and
re-subscribe the throttle, polluting the very counts the tests observe.

Verified: all three tsc projects clean (`tsc -p .`, `tsconfig.electron.json`,
`tsconfig.e2e.json`), use-background-sync suites 26/26, eslint clean.
2026-08-26 16:34:23 -05:00
Brooklyn Nicholson 68518c1f9b fix(desktop): hold the sessions list refresh for the whole typing burst
The starvation cap in the original patch re-ran the heavy pass at the
same ~10s mark the freeze is measured at. Hold until the keyboard is
quiet, then land one coalesced pass.
2026-08-26 14:40:43 -05:00
Bruno Bza 01e9b9abb1 fix(desktop): defer heavy sessions.changed list refresh while typing 2026-08-26 14:40:43 -05:00
Leon Phull f0187332d1 fix(desktop): decode-probe app icon candidates instead of existence-only (#94806)
new BrowserWindow({ icon }) and app.dock.setIcon() decode the icon file
synchronously on the main process and throw on undecodable bytes. The
icon ladder was resolved with statSync().isFile(), which only proves a
file exists — a truncated or zero-byte PNG inside a packaged app.asar
(interrupted electron-builder run, partial copy) killed the main process
inside createWindow(): the window never appeared, running turns lost
their renderer, and the desktop log showed 'Uncaught exception: Error:
Failed to load image from path .../app.asar/public/apple-touch-icon.png
at createWindow'.

Resolution now runs through a decoding probe (nativeImage.createFromPath
must yield a non-empty image); a candidate that exists but does not
decode is skipped like a missing one, so the app falls through to the
next rung or starts with the platform default icon instead of dying.
The ladder and probe live in a pure module (electron/app-icon.ts) so
precedence is unit-testable without a running Electron app; window
factories re-resolve per call exactly as before.

Regression tests cover: skip-first-undecodable, all-fail -> undefined,
first-pass wins, missing/empty/directory rejection, and the unchanged
mac/Windows precedence ladder.

Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
2026-08-26 18:44:46 +00:00
Teknium 847af7301a test: re-pin refusal exit codes and gate patch points to the shared contract
The apt/docker CLI tests pinned exit 1; refusals are now exit 2
(refused-by-contract, distinct from errors). The web_server guards
patched the module-local detect_install_method alias, which the shared
admission gate no longer consults — patch hermes_cli.config directly.
2026-08-26 11:41:04 -07:00
Teknium 8bf4e5a7ab docs: image provenance marker and the shared refusal gate (#91277 Phase 3) 2026-08-26 11:41:04 -07:00
Teknium 4860978115 feat(update): image/package-managed installs refuse in-place updates through one shared gate (#91277 Phase 3)
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.

A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.

Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
2026-08-26 11:41:04 -07:00
Andrex Ibiza, MBA 82a702bf67 feat(update): bake authoritative image provenance into the Docker image
Cherry-picked core of #92545: the image build writes a versioned,
non-secret marker (/etc/hermes/image-provenance.json) outside both the
bind-mountable checkout and the HERMES_HOME volume, and
hermes_cli/image_provenance.py reads it fail-closed — absence means
'not image-managed', any present-but-malformed marker still means
image-managed (an integrity defect is never permission to mutate the
image in place).

(#91277 Phase 3; salvaged from #92545 by @andrexibiza — marker bake +
reader only, the scoped carve-out.)
2026-08-26 11:41:04 -07:00
Jack Lau 2552579912 fix(tui_gateway): ask before queueing a guarded model picked mid-turn (#91043)
* fix(tui_gateway): ask before queueing a guarded model picked mid-turn

config.set model on a running session cannot swap the agent in place, so it
stashes the pick in session["pending_model_switch"] and applies it at the
next turn start. That branch answered confirm_required=False without ever
running the selection guards.

A client that implements the confirm round-trip was therefore told no
consent was needed and never prompted. One turn later
_apply_pending_model_switch ran the guards with the stashed (unconfirmed)
flag, saw the warning, and dropped the switch by design. The model reverted
with no confirm ever offered, because the only moment a round-trip was
possible had already passed.

Evaluate the guards before stashing, where the client still has a live
response to turn into a prompt. Nothing is queued for an unconfirmed
guarded pick, so the session is left exactly as it was and the re-send
carrying confirm_expensive_model queues it for real. The apply-time check
stays as the backstop for guards that can only decide after resolution.

The data-policy guard keys on the model id alone, which is all this branch
can see before resolution. The cost guard returns None when pricing is
unknown and its models.dev lookup is allow_network=False, so calling it
early can only under-fire and never blocks the RPC thread.

* test(tui_gateway): pin provider forwarding, name the canonical confirm field

Two review follow-ups, no behavior change.

_pending_switch_selection_warning forwards `provider=provider or None`, but
nothing asserted it: a guarded model id fires the data-policy guard on the
model alone, so the existing tests passed with `provider` dropped entirely.
Record the kwargs instead. Dropping the argument fails the first test;
removing the `or None` normalization fails the second.

The confirm responses carry `warning` and `confirm_message` with identical
text, which reads like an accident. Name which one clients should read
(`confirm_message`; `warning` is the pre-confirm-era alias that
_apply_pending_model_switch already treats as a fallback) so the two do not
drift apart later.

Both raised by @Enough1122 in review.
2026-08-26 13:38:53 -05:00
xxxigm b519ce29ad fix(desktop): keep attachment close and code copy icons visible (#95611)
* fix(desktop): keep attachment close and code copy icons visible

Hover-only opacity-0 hid the composer remove control and code-block copy button, so they stayed clickable but invisible on Windows and other no-hover surfaces.

* test(desktop): pin attachment close and code copy visibility at rest

The remove chip and code-block copy control must stay in the tree without a hover class, so Windows and no-hover surfaces cannot hide them again.
2026-08-26 13:38:19 -05:00
Gille d0fdbfd655 fix(desktop): show repo-root-only sessions in project drill-in (#94552) 2026-08-26 13:37:43 -05:00
Pedro Fontana b0dbf72f76 Merge pull request #91716 from NousResearch/fix/scale-to-zero-no-pointless-quiesce
fix(gateway): don't quiesce for a suspend this platform can't schedule
2026-08-26 14:58:38 -03:00
Nikita Barkov 2e80d7fa05 fix(slack): keep the resolved proxy on bolt's per-request client
slack_bolt builds a fresh AsyncWebClient for every inbound request and
copies proxy=app.client.proxy into its constructor, where slack_sdk reads
a None/blank proxy *argument* as "unspecified" and reloads HTTP(S)_PROXY
from the environment. aiohttp then treats that env value as an explicit
proxy and skips its own NO_PROXY check, so the adapter's resolved decision
to go direct - a NO_PROXY bypass, or a proxy scheme aiohttp cannot use -
holds on every client except the one authorization spends on auth.test.

The failure looks like a healthy bot: Socket Mode connects, outbound sends
keep working, and every inbound event is rejected with "Failed to authorize
with the given token" - forever, since a failed auth_test_result is not
cached and never retried differently.

Re-apply the resolved proxy through AsyncApp(before_authorize=...), which
bolt inserts before the authorization middleware: the request-scoped client
already exists there and has not been used yet. Assigning the attribute
post-construction is the only way to express "no proxy" to slack_sdk.

Co-authored-by: Junie <junie@jetbrains.com>
2026-08-26 10:35:05 -07:00
Nikita Barkov bac960e23d fix(slack): stop injecting thread roots as reply context 2026-08-26 10:26:00 -07:00
Nikita Barkov 1f92c5d4ce feat(kanban): carry the review handoff summary into the wake turn
`completed` already puts the worker's summary inside the synthetic wake
turn, so the woken creator sees what was done. `review_requested` did
not: the summary rode the passive ping only, and the wake turn said just
"handed off for review", forcing the woken reviewer to re-read the board
(and losing the PR link the worker had already written).

Reuse the same first-line handoff the `completed` branch builds, so the
existing `gateway.kanban.wake.handoff` string renders it — no new locale
keys, no change to the passive message.
2026-08-26 10:25:33 -07:00