Commit Graph

12957 Commits

Author SHA1 Message Date
Teknium b2c011364e fix(update): conservative outcomes + serve-ledger coverage for fresh restart recovery
Salvage adjustments to PR #94392 per review:

- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
  recovery child now probes 'systemctl --user is-active' after each relaunch;
  only an observed-active systemd unit is reported 'verified'. A relaunch that
  merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
  coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
  update_inventory serve collector) are no longer silently skipped: the
  recovery pass records them (and manual gateways) as skipped-with-reason in
  the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
  (requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
  fresh interpreter (sitecustomize shim intercepts the grandchild
  'gateway restart' and systemctl probes).
2026-08-26 16:45:26 -07:00
joaomarcos f9135c189c fix(update): persist fresh recovery outcome 2026-08-26 16:45:26 -07:00
joaomarcos ccdd7f41ed fix(update): persist per-profile recovery outcomes 2026-08-26 16:45:26 -07:00
joaomarcos f0045c5383 fix(update): verify fresh restart recovery results 2026-08-26 16:45:26 -07:00
joaomarcos 5609ccbece fix(update): recover aborted gateway restart in a fresh process 2026-08-26 16:45:26 -07:00
Teknium df3d41ee67 fix(update): sweep aborted-fetch tmp_pack debris before it corrupts the pack directory (#93732)
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.

clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).

Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
2026-08-26 16:45:22 -07:00
Teknium 7d6c6ae4ae refactor(clarify): schema diet + single questions[] interface (880 → 335 tok/call, −62%) (#95907)
* refactor(clarify): halve the schema (880->436 tok/call) — same rules, half the words

* refactor(clarify): one question field — questions[] is the only advertised shape (single = one-entry array; legacy shape stays handler-accepted)

* refactor(clarify): unadvertise per-question id — response rows already carry question text + order
2026-08-26 16:42:44 -07:00
chelsealong 2812d6121b fix(cli): note pre_restart_pids' per-PID data model gap, pin the matching-start_time path
Addresses the two follow-up notes from review: document that
pre_restart_pids is a bare PID set (not (pid, start_time) pairs), so a
recycled PID from one gateway landing in another's stale record could
still mislabel it as down; and add a companion test asserting a
matching start_time still yields the live/current row.
2026-08-26 16:14:32 -07:00
chelsealong 5d7ed70eef fix(cli): guard the post-update fleet check against PID reuse
collect_fleet_versions()'s gateway_state.json fallback path only checked
_pid_exists(pid) to decide whether a recorded gateway was still running.
On Windows, a paused gateway's PID can be recycled by an unrelated
process spawned during the update's own churn (npm/git/python
subprocesses) before the record is refreshed, so the dead gateway's
stale code_sha still gets compared against HEAD and reported STALE for a
PID that no longer belongs to it (#93258).

Switch to runtime_status_pid_is_live(), the existing (pid, start_time)
PID-reuse guard already used elsewhere in gateway/status.py, so a
recycled PID is treated the same as a dead one (DOWN row, or no row, per
the existing rollout-safety rules) instead of a false STALE.
2026-08-26 16:14:32 -07:00
pierrenode de2a9de788 fix(update): feed Windows gateway relaunch outcome into fleet reconciliation
#91277 Phase 2's plan-vs-execution reconciliation (match_runtime_outcomes)
cross-checks every runtime collect_runtime_inventory() saw against
restarted_services / relaunched_profiles / externally_supervised_profiles /
killed_pids — the systemd/launchd restart phase's bookkeeping. That
inventory is cross-platform (control-socket / PID-file based), so it
includes Windows gateways too, but Windows's own pause/resume mechanism
(_pause_windows_gateways_for_update / _resume_windows_gateways_after_update)
never wrote into any of that bookkeeping.

Result: a Windows gateway that was correctly stopped and relaunched by
_resume_windows_gateways_after_update was still classified "unaccounted" by
the reconciliation (the plan saw it and no bookkeeping mentions it) —
report_unaccounted_runtimes() escalates that into sys.exit(1), and in
gateway_mode also writes ".update_exit_code"="1". Every successful
`hermes update` on Windows with a running gateway reported itself as
failed, unconditionally (the sys.exit(1) is not gated to gateway_mode).

_resume_windows_gateways_after_update now records the profiles it
successfully relaunched onto the resume token; _cmd_update_impl merges
that into the shared relaunched_profiles list right before reconciliation
runs. A profile whose relaunch genuinely fails is deliberately left off
the list, so it still surfaces as unaccounted — Windows has no watcher to
recover a failed relaunch, so that escalation is the correct signal.

Regression tests exercise _resume_windows_gateways_after_update directly
(records successes, omits failures) and reproduce the reconciliation-level
bug end to end: the same plan row resolves "unaccounted" without the merge
and "restarted" with it. Mutation-verified: with the fix reverted, three of
the four new tests fail (KeyError on the token / wrong outcome).
2026-08-26 16:14:27 -07:00
Teknium 7a7a371c59 fix: harden claim-release guard for bare test doubles; repoint source-pinning test at the impl
The wrapper now getattr-defaults _processed_message_ts (object.__new__
adapters in sibling suites lack it), and the reaction-guard source pin
reads _handle_slack_message_impl where the production expression lives.
2026-08-26 15:54:53 -07:00
Teknium 39a5838f07 fix(slack): release a failed handler's fresh ts claim so the turn isn't swallowed
Follow-up for the #95417 salvage, addressing the review finding: the entry
claim closes the unfurl race but a handler that raises mid-enrichment would
hold the claim forever — neither a Slack retry nor a user edit could ever
re-drive the message. _handle_slack_message is now a thin guard around the
impl that releases only claims taken by the failed invocation itself, with a
warning log so swallowed turns are traceable. Pre-existing claims from a
successful turn are never released. Two failure-path tests pin both sides.
2026-08-26 15:54:53 -07:00
Richard Hojun Jang 708f84c477 fix(slack): claim message ts before enrichment so link unfurls can't duplicate a turn
Slack emits `message_changed` for a link unfurl carrying a DIFFERENT event ts
than the original message. That ts legitimately misses the `_dedup` check, so
`_processed_message_ts` is the only guard against it becoming a second user
turn -- but it was only populated at the END of `_handle_slack_message`, after
thread context, permalink resolution and file downloads had all awaited.

An unfurl landing inside that window found the guard empty and was promoted to
a duplicate turn: a spurious "Interrupting current task" banner plus the same
answer posted twice.

Production capture (adminbot, 2026-08-22 02:23:30-31Z, channel C0BF1EYUA9H):

  02:23:30.718  message      ts=1787365409.908499  dedup_hit=False
  02:23:30.737  app_mention  ts=1787365409.908499  dedup_hit=True
  02:23:31.675  message      ts=1787365411.012100  dedup_hit=False   <- leaked
                subtype=message_changed

The original copy was still resolving two Slack permalinks when the unfurl
arrived 957ms later.

Claim the message ts once every filter has passed and the event is certain to
be delivered, before the slow enrichment awaits. Claiming any earlier (right
after the dedup check) also claims messages the handler then discards, which
breaks summoning the bot by editing "@bot" into a previously ignored message
(tests/gateway/test_slack.py::TestMessageRouting::
test_message_edit_with_new_mention_processed).

Eviction logic is extracted to `_remember_processed_message_ts` so both call
sites share one bounded implementation.
2026-08-26 15:54:53 -07:00
Teknium 39a5aa91ed fix(serve): serve Desktop token page at / in headless mode (#94227)
The Electron shell boots by fetching / and extracting
window.__HERMES_SESSION_TOKEN__ to authenticate /api/ws
(dashboard-token.ts adoptServedDashboardToken). Headless serve 404'd
every path, so when the renderer's spawn token drifted from the
backend's live token — e.g. hermes update replaced the backend and the
env pin no longer matched — the renderer had no way to adopt the served
token, the WebSocket handshake failed, and the primary window
white-screened (#95575).

Serve a minimal token-only HTML page at the exact root path in
mount_spa()'s headless branch, matching the renderer's extraction regex.
Gate it on app.state.auth_required read at request time: a gated
(non-loopback / remote public_url) serve keeps returning the 404 JSON so
the session token never leaks past the loopback boundary. Every other
path stays 404 JSON — the SPA remains unserved.

Regression tests: TestHeadlessServeTokenPage (3 cases) — verified to
fail against the pre-fix headless branch.
2026-08-26 15:51:22 -07:00
pierrenode 6766732620 fix(memory-setup): route .env writer through save_env_value's validation gate
hermes_cli/memory_setup.py::_write_env_vars() wrote provider-controlled
.env entries with a direct Path.write_text() + post-hoc chmod, bypassing
the denylist/regex/CRLF-stripping/atomic-replace validation that
hermes_cli/config.py::save_env_value() already provides for every other
.env writer in the codebase. A malicious or buggy memory-provider plugin
declaring a crafted env-var name/value in its setup schema could inject
arbitrary lines into .env.

Routes memory-provider env writes through save_env_value(), and fixes a
regression this surfaced in plugins/memory/supermemory/__init__.py::
post_setup(), which called the old two-parameter _write_env_vars(env_path,
values) signature — restores the caller via context-local
hermes_constants.set_hermes_home_override()/reset_hermes_home_override()
instead of a removed env_path parameter, so explicit HERMES_HOME overrides
during setup still resolve correctly.

Adds test_env_file_created_with_secure_permissions, guarded on Windows
(POSIX mode bits aren't enforced there, mirroring the existing skip in
test_openviking_provider.py / test_supermemory_provider.py) since
save_env_value's atomic-replace path creates the temp file at 0o600 before
writing content, closing the TOCTOU window the old direct-write + chmod
implementation had.
2026-08-26 15:48:39 -07:00
DmytroVolodymyrson a699234f81 fix(kanban): wake controllers on review changes
Deliver changes_requested review outcomes through kanban subscriptions and
wake the origin for notify+wake / wake modes. Review-specific only: no task
mutation. Reasons are redacted, path-scrubbed and truncated before delivery.

Salvaged from #88694; conflicts with the #87733 wake-kinds expansion resolved
keep-both.
2026-08-26 15:31:36 -07:00
Teknium 847af7301a test: re-pin refusal exit codes and gate patch points to the shared contract
The apt/docker CLI tests pinned exit 1; refusals are now exit 2
(refused-by-contract, distinct from errors). The web_server guards
patched the module-local detect_install_method alias, which the shared
admission gate no longer consults — patch hermes_cli.config directly.
2026-08-26 11:41:04 -07:00
Teknium 4860978115 feat(update): image/package-managed installs refuse in-place updates through one shared gate (#91277 Phase 3)
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.

A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.

Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
2026-08-26 11:41:04 -07:00
Jack Lau 2552579912 fix(tui_gateway): ask before queueing a guarded model picked mid-turn (#91043)
* fix(tui_gateway): ask before queueing a guarded model picked mid-turn

config.set model on a running session cannot swap the agent in place, so it
stashes the pick in session["pending_model_switch"] and applies it at the
next turn start. That branch answered confirm_required=False without ever
running the selection guards.

A client that implements the confirm round-trip was therefore told no
consent was needed and never prompted. One turn later
_apply_pending_model_switch ran the guards with the stashed (unconfirmed)
flag, saw the warning, and dropped the switch by design. The model reverted
with no confirm ever offered, because the only moment a round-trip was
possible had already passed.

Evaluate the guards before stashing, where the client still has a live
response to turn into a prompt. Nothing is queued for an unconfirmed
guarded pick, so the session is left exactly as it was and the re-send
carrying confirm_expensive_model queues it for real. The apply-time check
stays as the backstop for guards that can only decide after resolution.

The data-policy guard keys on the model id alone, which is all this branch
can see before resolution. The cost guard returns None when pricing is
unknown and its models.dev lookup is allow_network=False, so calling it
early can only under-fire and never blocks the RPC thread.

* test(tui_gateway): pin provider forwarding, name the canonical confirm field

Two review follow-ups, no behavior change.

_pending_switch_selection_warning forwards `provider=provider or None`, but
nothing asserted it: a guarded model id fires the data-policy guard on the
model alone, so the existing tests passed with `provider` dropped entirely.
Record the kwargs instead. Dropping the argument fails the first test;
removing the `or None` normalization fails the second.

The confirm responses carry `warning` and `confirm_message` with identical
text, which reads like an accident. Name which one clients should read
(`confirm_message`; `warning` is the pre-confirm-era alias that
_apply_pending_model_switch already treats as a fallback) so the two do not
drift apart later.

Both raised by @Enough1122 in review.
2026-08-26 13:38:53 -05:00
Pedro Fontana b0dbf72f76 Merge pull request #91716 from NousResearch/fix/scale-to-zero-no-pointless-quiesce
fix(gateway): don't quiesce for a suspend this platform can't schedule
2026-08-26 14:58:38 -03:00
Nikita Barkov 2e80d7fa05 fix(slack): keep the resolved proxy on bolt's per-request client
slack_bolt builds a fresh AsyncWebClient for every inbound request and
copies proxy=app.client.proxy into its constructor, where slack_sdk reads
a None/blank proxy *argument* as "unspecified" and reloads HTTP(S)_PROXY
from the environment. aiohttp then treats that env value as an explicit
proxy and skips its own NO_PROXY check, so the adapter's resolved decision
to go direct - a NO_PROXY bypass, or a proxy scheme aiohttp cannot use -
holds on every client except the one authorization spends on auth.test.

The failure looks like a healthy bot: Socket Mode connects, outbound sends
keep working, and every inbound event is rejected with "Failed to authorize
with the given token" - forever, since a failed auth_test_result is not
cached and never retried differently.

Re-apply the resolved proxy through AsyncApp(before_authorize=...), which
bolt inserts before the authorization middleware: the request-scoped client
already exists there and has not been used yet. Assigning the attribute
post-construction is the only way to express "no proxy" to slack_sdk.

Co-authored-by: Junie <junie@jetbrains.com>
2026-08-26 10:35:05 -07:00
Nikita Barkov bac960e23d fix(slack): stop injecting thread roots as reply context 2026-08-26 10:26:00 -07:00
Nikita Barkov 1f92c5d4ce feat(kanban): carry the review handoff summary into the wake turn
`completed` already puts the worker's summary inside the synthetic wake
turn, so the woken creator sees what was done. `review_requested` did
not: the summary rode the passive ping only, and the wake turn said just
"handed off for review", forcing the woken reviewer to re-read the board
(and losing the PR link the worker had already written).

Reuse the same first-line handoff the `completed` branch builds, so the
existing `gateway.kanban.wake.handoff` string renders it — no new locale
keys, no change to the passive message.
2026-08-26 10:25:33 -07:00
Nikita Barkov 7700d3a011 fix(kanban): wake the origin on review handoffs and triage escalations
`review_requested` and `block_loop_detected` are terminal event kinds that
hand a decision back to the origin subscriber, but neither was listed in the
gateway notifier's `_WAKE_KINDS`. A `notify+wake` subscription therefore got
the passive ping only and the origin agent never took a turn — so an agent
that delegated implementation work slept through the "ready for review"
handoff and through a task being routed to triage, while the equivalent
`blocked` event woke it.

Add both kinds to the wake set, add their status strings to the synthetic
wake message in every locale, and document which events wake.
2026-08-26 10:25:33 -07:00
Nikita Barkov 5538bd1f93 fix(slack): prevent duplicate rich-text message content
Slack sends an authored message twice: flat in `event.text` and structurally
in `event.blocks`. The blocks are rendered so quoted and forwarded content is
not lost, and whatever the render carries beyond the flat text is appended to
the message. That comparison had several ways to fail on the *same* sentence,
each of which showed the author their own words a second time:

1. HTML entities — the flat copy escapes `&`/`<`/`>` while `blocks[].link.url`
   stays raw, so any link with query parameters (every "Copy link" on a
   thread) mismatched.
2. Permalink unfurls — the live inbound path skips `is_msg_unfurl`
   attachments, thread/parent hydration did not, so the linked message's body
   was appended again.
3. The Block Kit dump — it serialized the authored `rich_text` alongside the
   UI blocks it exists for, and its allowlist drops `url`, so the sentence
   reappeared with every link removed.
4. Unknown inline elements — the renderer knew eight types and silently
   dropped the rest. A pasted message permalink arrives as `message_mention`,
   so the link vanished from the render and the sides stopped comparing equal.
5. `message_mention` without a url — `url` is optional on that element while
   `channel_id` and `message_ts` are not, so the element rendered as nothing
   and the sentence came back with a blank in the link's place.
6. `date` elements — `fallback` and `url` are both optional, and the flat
   `<!date^…>` form was never read down to what the rich text renders.
7. Labelled mentions — Slack may attach a label (`<@U…|name>`,
   `<#C…|general>`, `<!subteam^S…|@marketing>`, `<!here|@here>`) in the flat
   text while the blocks carry the bare id. The bot's own mention is one of
   these, and stripping only its bare form left it in the flat copy.
8. Autolink schemes — only `https` and `mailto` were matched, so a `tel:` link
   kept its angle brackets and mismatched too.

Unknown inline types are now read by their `url`/`text`/`fallback` so a type
Slack adds later still renders, and `team`, `color` and a fallback-less `date`
render into the flat form Slack sends. Every field is read as a string or
not at all: Block Kit carries text as an object in many places, and a
non-string one reaches the renderer's `str.join` and raises there, which
costs the whole message. `channel_id` and `message_ts` are the
permalink's own components, so a url-less `message_mention` renders the
permalink's tail; the workspace host and the thread query cannot be rebuilt
from the element, so a permalink on either side is reduced to that same tail.
Canonicalization is used for matching only -- the authored text still reaches
the agent verbatim, so a mistake here can cost an unrendered element, never an
altered or missing message.

An element carrying neither a url nor a label still renders as nothing, and a
message containing one is still appended twice. Suppressing such a render was
tried and is worse: an app message whose body lives only in the blocks
disappears, and a forwarded quote is dropped. Genuinely additional content --
quotes, lists, code blocks, attachments, interactive bot blocks -- is
unaffected throughout.

Tests cover both merge sites (live inbound and thread hydration) and the
negative cases.
2026-08-26 10:10:51 -07:00
Alexander Prendota 613164dadc fix(acp): key the ACP runtime exclusions on the scheme, not on one vendor
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.

Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
2026-08-26 10:10:11 -07:00
Alexander Prendota 37fd61d13b fix(background-review): skip the fork when the provider can't emit tool calls
The review fork's entire job is to emit `memory` / `skill_manage` tool calls,
and by default it inherits the parent's live runtime. A provider that IS an
autonomous agent reaches Hermes through a client shim; if that shim cannot
carry Hermes tool calls back, the fork is a guaranteed no-op that still pays
for a full agent spawn — a whole CLI process, sometimes a JVM — on every
review cadence.

A client declares `SUPPORTS_HERMES_TOOL_CALLS = False` and the fork is
skipped with a warning naming the `auxiliary.background_review.{provider,model}`
override that routes the review to a normal model instead. Anything that says
nothing is assumed capable, so ordinary providers are untouched.

The check runs before the thread-scoped silence so the warning is not
swallowed, and only resolves the review runtime once the cheap capability
test has already failed, so the normal path does not resolve it twice.
2026-08-26 10:10:11 -07:00
Alexander Prendota 07200e9cd6 feat(agent): fold an agent-as-provider's own tool work back into the turn
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.

Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:

- the self-improvement loop, which distils memories and skills by replaying
  `messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
  Hermes tool iterations, of which there are none.

So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.

The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
2026-08-26 10:10:11 -07:00
Alexander Prendota 083c5920e5 refactor(acp): share one OpenAI bridge between the ACP clients
ACP has no OpenAI `tools`/`tool_calls` channel: a prompt is text and a
response is text plus the agent's own tool notifications. Hermes' agentic
surface — memory, todo, skill_manage — is dispatched from OpenAI-shaped
tool_calls, so on an ACP provider it only works if the schemas travel into
the prompt as text and the calls are parsed back out of the response text.

copilot-acp already carried that bridge as private module-level helpers.
Lift it verbatim into `agent/acp_openai_bridge.py` so every ACP client
shares one implementation of the wire contract instead of re-deriving it —
`agent/claude_code_acp_client.py` (#81375) is currently a third copy of the
same four functions, and each copy is a place the `<tool_call>` contract can
drift.

copilot-acp is migrated onto it as the in-tree consumer and loses 176 lines
of duplication; its prompt shape is unchanged, which the new tests pin.

Two things the shared version adds over the copy:

- `render_tool_bridge_sections(..., allowlist=)`. A CLI with no tools of its
  own forwards Hermes' whole toolset (copilot, unchanged: no allowlist). A
  CLI that *is* an autonomous agent must forward only Hermes' agent-level
  tools — re-offering the overlapping read/edit/execute ones makes Hermes
  re-run work the agent already finished.
- `StreamChunks`, a list subclass that keeps response-level attributes.
  Hermes reads provider extras off the object returned by
  `chat.completions.create`; the old plain-list return silently dropped them
  whenever a caller asked for `stream=True`.
2026-08-26 10:10:11 -07:00
Teknium 03537d69dc feat(gateway): updaters pause gateways over the control socket instead of tree-killing them (#92091 step 2)
Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.

- gateway/run.py: pause-for-update verb handler registered on the
  existing control server; marshals onto the loop thread, ACKs with
  {pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
  no-answer (older gateway / no socket), so every caller keeps the
  legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
  per mapped profile gateway before the drain wait; positive ACKs extend
  the wait to the gateway's own declared drain budget (+ teardown grace)
  so a mid-turn gateway isn't force-killed at the end of a too-short
  local default. Marker write + force-kill ladder retained verbatim as
  the fallback.

Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).
2026-08-26 09:59:17 -07:00
unsupportedpastels c4f376c19a fix(config): block generic Copilot ACP controls 2026-08-26 09:54:31 -07:00
unsupportedpastels 5425ba14f2 fix(config): harden MCP env policy on Windows 2026-08-26 09:54:31 -07:00
unsupportedpastels 08cf4fea5d fix(mcp): restrict catalog environment writes 2026-08-26 09:54:31 -07:00
pefontana b2c493c173 Assert the skip branch ran in the no-quiesce watcher test
The three existing assertions are absence checks, so the test also
passed when the loop got no iteration inside the sleep window (with
interval=5.0 it passes without the gate ever executing). Checking
_scale_to_zero_no_suspend_logged proves the branch was taken.
2026-08-26 13:02:04 -03:00
pefontana 85816595d2 Merge remote-tracking branch 'origin/main' into fix/scale-to-zero-no-pointless-quiesce 2026-08-26 13:02:03 -03:00
Teknium 30749ed9dc test(tools): guard _ever_connected set in reconnect regression mock so it bites on pre-fix code (#94671 hardening) 2026-08-26 08:40:07 -07:00
chelsealong 8f517f5ca6 chore: address AI-review nits on _ever_connected fix
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
2026-08-26 08:40:07 -07:00
chelsealong e8dc0af5b1 fix(tests): set _ever_connected in reconnect-scenario test mocks
These pre-existing tests fake a successful first connect by calling
only _ready.set(), which is what the real code did before this PR.
Now that run() gates the initial-vs-reconnect branch on the new sticky
_ever_connected flag instead, their later simulated reconnect failures
were misclassified as never-connected and hit the 3-attempt ladder,
failing test_reconnect_counter_resets_after_successful_session,
test_parked_server_self_probes_and_revives, and
test_retry_attempts_log_debug_transitions_warn in CI. Set the flag
alongside _ready.set() to mirror the real success sites, same as the
new test added in tools/mcp_tool.py's own PR.
2026-08-26 08:40:07 -07:00
chelsealong c7673f322b fix(tools): stop treating a post-registration reconnect drop as an initial-connect failure
MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.

Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.

Fixes #94654
2026-08-26 08:40:07 -07:00
liuhao1024 57309c0cbb fix(update): never respawn backends from a foreign HERMES_HOME (#94030)
The stale-dashboard sweep at the end of hermes update snapshots each killed
backend's HERMES_HOME (_hermes_home_for_pid) but only used it as the per-profile
dedupe key. _respawn_dashboard_processes replays the argv with no env=, so a
backend belonging to a second install (e.g. a launchd KeepAlive sidecar) came
back running on the updating install's default home and stole the sidecar's
fixed port: the supervisor crash-looped on EADDRINUSE and clients on that port
silently talked to the wrong backend.

Drop such candidates in _filter_dashboard_respawn_candidates: a backend whose
captured HERMES_HOME differs from the updater's own get_hermes_home() is not
replayed at all — its own supervisor/user owns its lifecycle. Homes are
normalized the same way _profile_key_for_respawn normalizes home: keys, so
symlinked roots compare equal. An unreadable home (None) stays eligible,
keeping the pre-fix fail-open behaviour.
2026-08-26 08:39:04 -07:00
fangliquan 858916acc4 fix(update): preserve SSH ownership only during updates 2026-08-26 08:39:04 -07:00
fangliquan 1676c614b3 fix(update): preserve SSH-owned backends during cleanup 2026-08-26 08:39:04 -07:00
Teknium 306a096b4a test: re-pin --status output to the serve-inclusive contract (#81564)
The old assertions pinned the phrasing that HID serve backends — the
exact asymmetry #81564 reports. Re-pinned to the new message and
strengthened: a serve-mode row must now appear, tagged [serve].
2026-08-26 07:57:04 -07:00
Teknium 27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
beplee 19d8b87234 test(desktop): harden rebind assertions per #94417 review
Enough1122 review points on #94417:
1. Precedence hazard fixed: the busy-guard assertion now locates the
   rebind helper body precisely and asserts the guard INSIDE it, instead
   of a 2000-char window with an (m and X) or Y precedence trap.
2. stored_session_id guarantee: documented + pinned — the gateway always
   stamps it ('stored_session_id': session_key or "" in server.py), and
   the rebind's typeof check refuses non-string/empty values, so an
   unnamed rebuilt runtime is never adopted as lineage proof.
3. New third assertion pins that refusal contract.

Structural smoke tests remain structural by design; the behavior
contract for the rebind is exercised end-to-end by the model-switch
manual repro path — a vitest harness driving handleSessionInfoEvent is
the follow-up candidate noted in the reply.
2026-08-26 07:28:10 -07:00
beplee ec8ca8f2cb fix(desktop): re-bind open pane to rebuilt runtime after model switch
A mid-conversation model/provider switch rebuilds the agent runtime. The
rebuilt runtime emits session.info (and all later events) under a NEW
explicit session_id while the pane still holds the dead one as its
active id — isActiveEvent is false for the same conversation from that
moment on, so view-scoped updates stop and the chat freezes until a
full resume (#93942 scenario B; backend even logs 'client should resume
the stored session', but the client never does).

Fix: when a session.info event lineage-matches the selected conversation
(sessionMatchesStoredId over stored_session_id) but carries a different
runtime id, adopt the new runtime id as the active session id — keeping
the durable selection untouched — so every subsequent isActiveEvent gate
keeps matching without a resume. Guarded: the old runtime must show no
live turn (not busy/awaiting/streaming) or the adoption is refused, so
an overlapping manual switch can never split one conversation across
two panes.

The existing compression-rotation path does not cover this case: it
fires when the SAME runtime's stored id rotates, while a rebuild
produces a NEW runtime with a NEW stored id.

Together with #94255 (tile reconcile on sessions.changed), closes
#93942.

Regression tests verified failing pre-fix on 41447a6d70.
2026-08-26 07:28:10 -07:00
beplee db8ff4eb75 fix(desktop): reconcile workspace-tile transcripts on sessions.changed
Bot canonical chats open as workspace tiles (workspaceMode: 'bots') and
are deliberately hidden from $sessions/$messagingSessions, so the
sessions.changed transcript refresh skipped them twice over: it covers
only the main pane's selection, and its resolveSession() bails on hidden
sessions. A background delivery (bot-to-bot DM via bot_relay.deliver, a
cron run's output, another machine) therefore never reached an open bot
chat — the roster updated but the pane stayed stale until remount
(#93942 scenario A).

Fix: the sessions.changed tick now also reconciles every visible
workspace tile through a dedicated signature-gated path. Each tile
carries its own stored↔runtime id pair so no resolution step is needed;
per-tile signatures make no-change ticks free; busy tiles are skipped
(their own stream owns the view); closed/superseded tiles discard their
in-flight read.

Slice 1 of 2 for #93942 (scenario A only). Scenario B (stream re-key
after mid-conversation model switch) follows separately.

Fixes part of #93942
2026-08-26 07:28:10 -07:00
Teknium f4df86fe1a test: adapt summary-continuity + rotation-flush fixtures to the lean default
Continuity tests pin tail_mode=legacy (they assert the raw LLM text
terminates the stored summary; lean's verbatim-user appendix follows it by
design and the contract under test is mode-independent). The #57491
rotation fixture grows 200→2000 chars/message: at ~2.5K total tokens the
old fixture fit entirely inside lean's 10K tail floor, so the no-growth
guard correctly refused the rotation the test exercises.
2026-08-26 07:16:04 -07:00
Teknium 6e5413844e feat(compression): lean tail retention is the default — compaction keeps 10-25K verbatim, not 100-240K
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.

Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.

Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).

Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).

E2E counterfactual (real imports, 1M window @ 0.85):
  main default:  legacy, tail 170,000 (ceiling 255,000)
  head default:  lean,   tail  25,000 (ceiling  37,500)
  head legacy:   170,000 (opt-out intact)
  update_model to 400K: 10,000 (lean preserved across switch)
2026-08-26 07:16:04 -07:00
Teknium 84b91a1dc5 test(ssh-ownership): remove process-global patches that crashed sibling threads
De-flakes tests/hermes_cli/test_ssh_ownership_endpoint.py, which failed CI
twice on PR #95563 with teardown-time daemon-thread excepthook crashes — a
different test in the file each attempt, always green in isolation. Root
cause: three PROCESS-GLOBAL monkeypatches leaked into every other thread
sharing the per-file worker:
- monkeypatch.setattr(web_server.os, 'stat', ...) — web_server.os IS the os
  module; any daemon thread from an earlier test that stat()ed during the
  patch window got the fake 2-field stat and died in its excepthook, which
  fired at interpreter teardown.
- monkeypatch.setattr('builtins.open', ...) — same class, worse blast radius.
- monkeypatch.setattr(web_server.sysconfig, 'get_paths', ...) — sysconfig is
  process-global too.

Fixes, none of which weaken coverage:
- replaced-runtime test: a REAL tmp_path purelib whose recorded inode
  deliberately mismatches (st_ino + 1) — real os.stat, same code path.
- readonly-purelib test: chmod 0o555 on the real directory instead of an
  open() interceptor — exercises the genuine OSError branch (root-skipped,
  where mode bits aren't enforced).
- sysconfig patches swapped for a SimpleNamespace on the web_server module
  attribute — module-scoped, invisible to other threads.

Verified: 14 consecutive full-file runs green; sabotaging
_ssh_runtime_intact to always-True still fails 2 tests (coverage intact).
2026-08-26 07:03:04 -07:00