Commit Graph

15410 Commits

Author SHA1 Message Date
chelsealong 21d5df4100 fix(session-state): keep the canonical Bot Chat hidden when pinned
set_session_pinned's unconditional hidden-clear (0be7f931) also
unhides the canonical Bot Chat, which the desktop contract requires
to stay hidden and reachable only through the bot row. Exposing it
breaks the sidebar and disables the rename guard that protects its
identity title.

Skip the unhide when the pinned row is hidden and carries the exact
canonical title; ordinary hidden sessions are unaffected.
2026-09-09 10:52:21 -07:00
chelsealong c59aa2d041 fix(session-state): clear hidden flag when a session is pinned
A bot-mode session is created with hidden=true. Pinning it only wrote
pinned=1 and left hidden=1, so the row was filtered out of both the
default listing (`s.hidden = 0`) and the pinned back-fill (which reuses
the same WHERE), making a pinned session vanish from the sidebar
entirely. set_session_pinned now clears hidden across the session's
compression lineage whenever pinned is set to true.

Fixes #106171
2026-09-09 10:52:21 -07:00
KeyArgo 070128f056 test(agent): provider-neutral name for 429 server-overload regression test 2026-09-09 10:47:49 -07:00
KeyArgo 43ecad8fe6 fix(agent): classify Novita 'server overload' 429 as overloaded
Novita returns HTTP 429 with message 'server overload, please try again
later' and error type 'server_overload' when its server is genuinely busy.
Neither phrase was in _OVERLOADED_PATTERNS, so the 429 fell through to
_V_RATE_LIMIT and set should_fallback=True + should_rotate_credential=True —
rotating the credential / falling back early instead of retrying the same
key. Add 'server overload' and 'server_overload' to the overload tuple so
this reaches the existing _V_OVERLOADED verdict (retryable, no rotation).
Closes #106205
2026-09-09 10:47:49 -07:00
Teknium e22e04692a test(matrix): trim LaTeX salvage to two invariant tests, tidy _markdown_to_html
Keep the two tests that fail without the fix: inline + display math reach
formatted_body as data-mx-maths markup with the TeX HTML-escaped while the
plain body keeps the raw TeX; unpaired dollars and text colliding with the
sentinel format pass through unchanged (no IndexError). The 17 unit tests
from #106452 are dropped per the salvage bar (≤ 2 invariant tests).

Also drop the dead `_tex_store = []` pre-assignment and the `_fb`/`html`
temporaries in `_markdown_to_html` — pure tidy, no behaviour change.
2026-09-09 10:43:58 -07:00
romanovzky 199f4d5b96 feat(matrix): render LaTeX math via Element data-mx-maths markup
Element (feature_latex_maths) typesets <div|span data-mx-maths="TEX">
elements at display time, but the outbound HTML sanitizer allowlists tags
and attributes, so data-mx-maths markup sent by the gateway never reaches
Element intact - messages containing $...$ render as raw dollars.

Convert $...$ (inline) and $$...$$ (display) to opaque sentinel tokens
before Markdown conversion and expand them to data-mx-maths markup after
sanitization. Tokens are printable text with no HTML/Markdown meaning, so
neither the converter nor the sanitizer touches the TeX. Unpaired dollars
(prices, literals) are untouched, and adversarial text colliding with the
sentinel format passes through verbatim (index-checked expansion).

(cherry picked from commit eb73aafe0e08502c54e63dbd03e99796c3b7fee3)
2026-09-09 10:43:58 -07:00
Teknium 1a500a43a7 fix(gateway): keep launchd_stop's bootout quiet too; prove the fd-2 invariant with a fake launchctl
Widens #106272 to the one remaining sibling: `launchd_stop()` boots out with check=True and
already handles exit 3/113/125 (job unloaded) and 5/125 (domain unmanageable) by falling through
to the PID kill, yet inherited stderr — so `hermes gateway stop` against an unloaded job printed
"Boot-out failed: 3: No such process" next to "✓ Service stopped". Same `_CAPTURE_TEXT` kwargs as
the sibling calls; an unexpected exit still raises with `e.stderr` populated.

Replaces the contributor's two kwarg-assertion tests (`capture_output is True` on a mocked
`subprocess.run`) with two invariant tests that run a real fake `launchctl` on PATH and read fd 2
through `capfd`: restart-on-unloaded prints only its ↻/✓ lines and drives
kickstart→bootout→bootstrap→kickstart; stop-on-unloaded is silent, while a real bootout failure
(exit 1) still raises with the captured stderr. Both red on origin/main, green here.
2026-09-09 10:42:32 -07:00
teknium1 79d73e8f63 test(telegram): trim bots_require_mention coverage to the two invariants
Keep the loop-breaker test (bot quote-reply processed on main, dropped with the
flag, explicit @mention still passes) and the human quote-reply test (unaffected
with the flag on). Drop the defaults-off snapshot, the /cmd@botname variant, the
plain-chatter duplicate and the own-echo helper test: they re-cover paths
(_message_mentions_bot, _is_own_message) already pinned elsewhere in this file.
2026-09-09 10:37:37 -07:00
Konstantin Khlopkov 1a981d8e83 fix(telegram): bots_require_mention gates bot quote-replies behind an explicit @mention (closes #106430)
(cherry picked from commit a67fa8c4b21e31a44019d850aedc5be9015a546d)
2026-09-09 10:37:37 -07:00
teknium1 8d93081971 fix(desktop): stop flagging local/LAN auxiliary pins as stale
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.

- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
  the verdict of the one canonical classifier
  (`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
  private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
  `staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
  appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
  IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
  a global-scope address (`2607:f8b0::1`) is not local while `::1`,
  ULA and link-local still are via the `ipaddress` scope checks.

Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.

Refs #106228

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 10:33:00 -07:00
teknium1 113304199e refactor(desktop): gate the WSLg D3D12 selection inside the helper and test the real launch env
Move the WSL / /dev/dxg / d3d12_dri.so probes into _prefer_wsl_d3d12 with
the probed paths as module constants, so the launcher call site is a single
line and a test can lay out a fake WSLg host without touching real
/dev or /usr/lib. The two tests now run the real _desktop_launch_env end to
end (selected under WSL+dxg+driver; untouched with an explicit Mesa
override, off WSL, without /dev/dxg, or without the driver file) instead of
unit-testing the helper with a precomputed boolean.

Docs: one paragraph in the Desktop guide on the automatic selection and
the env vars that keep an explicit choice authoritative.

Follow-up to Xipong's fix for #106117 (salvaged from #106118).
2026-09-09 10:28:52 -07:00
Xipong 51fa04d49d fix(desktop): select installed WSL D3D12 driver before Electron exec 2026-09-09 10:28:52 -07:00
teknium1 a785db3672 fix(cli): inline the $HERMES_HOME/npmrc lookup, trim tests to two, document it
Fold the helper into _npm_lifecycle_env itself: the whole fix is one
is_file() check plus a setdefault, so a separate function, the
try/except around get_hermes_home() (it never raises) and the
dict-returning indirection were shape-gate violations. Build the path
with os.fspath so it is correct on Windows too.

Tests: keep the two invariants (file present -> NPM_CONFIG_USERCONFIG
points at it; explicit process/caller value wins and a missing file
sets nothing), fold the other two into the negative test.

Docs: one paragraph in the desktop troubleshooting page next to
ELECTRON_MIRROR describing $HERMES_HOME/npmrc.

Refs #106373
2026-09-09 10:26:06 -07:00
Konstantin Khlopkov ce4a33a9f7 fix(cli): keep $HERMES_HOME/npmrc npm config across updates (#106373)
(cherry picked from commit 339368f2215ffcfd04acf295c81806bf1e850770)
2026-09-09 10:26:06 -07:00
teknium1 d3dcc064df fix(cli): detect ssl.SSLError by type in the Codex login hint; trim tests; add the openssl.cnf snippet to docs
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
  __cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
  SSLEOFError whose text httpx did not repeat still gets the hint. The
  hint now names the TLS 1.2 diagnostic and links the providers docs
  note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
  the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
  classic-groups snippet (EN + existing zh-Hans copy).

Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
2026-09-09 10:14:58 -07:00
liuhao1024 d8eb177c93 fix(cli): keep SSL detail and add middlebox hint on Codex device-login transport errors
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.

- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
  (device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
  workaround hint; KeyboardInterrupt handling is unchanged

(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
2026-09-09 10:14:58 -07:00
Teknium 42d28d64f0 test(hermes_state): trim compacted-paging tests to the invariants that fail on base
Six new tests collapsed into three that each go red without the indexed
projection: bounded VM steps for a page read and for an append (the
symptom), composite user-handoff identity + first-row position + internal
columns never leaking + signed-zero timestamp identity (the parity
contract), and a legacy store read-only then lazily migrated (the
persistence contract). The parametrized symmetry/visibility-toggle test
exercised the same key through raw SQL edits no production path performs
and passed on base, so it was a change-detector for the trigger shape.
2026-09-09 10:05:59 -07:00
Xipong 1c6683e8e0 fix: index compacted display identity writes 2026-09-09 10:05:59 -07:00
Xipong 49e6d661a0 fix: bound compacted display history paging 2026-09-09 10:05:59 -07:00
Eva 2536772301 fix: preserve summary boundaries when restoring model replay 2026-09-09 09:59:13 -07:00
teknium1 5de30f36f1 refactor(gateway): reuse recorded_gateway_home_conflicts for the scoped PID home check
The salvaged fix added `_pid_record_matches_home`, a near-copy of
`recorded_gateway_home_conflicts(record, expected_home=...)` (already the
scoped-home predicate used by the #89315 stop guard). Reuse it instead of
carrying a second helper with the same semantics (legacy records without
`hermes_home` are accepted by both).

Also trims the salvaged tests to the two invariants that prove the symptom
(scoped probe reports a live foreign profile's PID and leaves its
identity files intact; a dead-PID scoped record is still cleaned). The
third-home-claim variant exercised the same `saw_live_pid` branch.
2026-09-09 09:55:52 -07:00
liuhao1024 a23319d0f7 fix(gateway): scoped PID queries validate against the probed home (#106406)
get_running_pid(pid_path) validated records against the serve process's
HERMES_HOME, so a dashboard scoped status poll (?profile=) rejected a live
foreign profile's record, force-unlinked its gateway.pid/gateway.lock, and
reported running=false. Validate scoped queries against the probed home
(pid_path.parent, like the expected_home sibling) and never cleanup-unlink
a live record's identity files; dead-PID records still clean up.

(cherry picked from commit 472ca1b8d87df5246d83864e9b1492d06e99e8e5)
2026-09-09 09:55:52 -07:00
Teknium b24a781b6f test(agent): trim the restart-bound tests to two invariants
Collapse the four class-based tests into two parametrized invariants over both
refunding restart flags and move them to tests/agent/ (the phase modules live in
agent/): a single restart still refunds-and-continues; a re-armed restart breaks
after max_retries refunds. The stub grows the redirect seam the follow-up commit
uses so the queued-correction contract is covered by the same test.
2026-09-09 09:51:31 -07:00
yoyodine-industries e9312da68b fix(agent): bound redirect/rebuilt restart refunds so a runaway turn can't hold the session lease
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).

Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
2026-09-09 09:51:31 -07:00
Halldrix 78de23053f fix(gateway): report media-only turns as SUCCESS when attachments deliver
Thread record_delivery through _deliver_attachments, _deliver_media_attachments
and _send_image_batch so attachment sends feed the turn outcome tracker.
send_multiple_images (base default and Signal override) now returns SendResult
(success when at least one image/batch was accepted); _send_attachment_batch
returns bool. Legacy native-batch overrides that still return None record
nothing, keeping their previous behavior until migrated.

Fixes #106153
2026-09-09 09:45:54 -07:00
Teknium 00bcef9b5f test(kanban): install the reviewer profile the review-surface fixture hands off to
kanban_request_review now rejects reviewers that are not installed profiles
(#106163); the cross-surface lifecycle test used a bare "reviewer" name with no
profile behind it, which is exactly the phantom the guard exists to catch.
2026-09-09 09:45:13 -07:00
teknium1 e2763baf1c refactor(kanban): route the reviewer guard through _check and tighten its tests
Use the module's `_check`/`_Reject` idiom instead of an inline
`return tool_error(...)` so every kanban_request_review validation
failure renders through the same path, and drop the unreachable
`or "none"` (list_profile_names() always contains "default").

Tests: compare the task's (status, assignee, run) tuple and the event
log before/after instead of the unordered 6-assert block, use the
context-managed kanban_db_connect.connect (the kb.connect alias is a
plugin-compat pointer — scripts/check_compat_pointers.py flagged it),
and reference #106163 in the invariant's docstring.

Salvage note vs #106214 (@gaoanze888): that PR guards the same condition
inside hermes_cli/kanban_db.py::request_review, but the DB primitive is
also the chokepoint for `hermes kanban request-review` and the
dashboard's drag-to-review, both operator surfaces where a non-profile
assignee (external/human review lane) is a documented board shape
(website/docs/user-guide/features/kanban-worker-lanes.md) — and it forced
five unrelated test fixtures to monkeypatch profile_exists to True. The
model-facing tool wrapper is the layer where a typo'd string is a bug,
so the guard lives there.
2026-09-09 09:45:13 -07:00
auroracapital 1d89286b36 fix(kanban): reject phantom worker reviewers
kanban_request_review(reviewer=<name>) reassigned the task to whatever
string the model supplied. A non-profile value (e.g. the literal
"reviewer") parked the card in `review` on an assignee the dispatcher
can never spawn, with no error to the worker — the chain stalled
silently (#106163). Validate the explicit reviewer against installed
profiles before touching the board and return a tool error listing the
installed profiles so the model can self-correct.

Salvage of #97429: the kanban_diagnostics `review_reopened` hunk and its
tests were dropped (main already replaced that loop with
`_latest_event_ts`; the tests targeted the PR's pre-refactor base).
Re-authored from the placeholder identity `regen <regen@local>` to the
PR author's GitHub noreply address (misconfigured local git, not malice).
2026-09-09 09:45:13 -07:00
Teknium e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
Teknium 3026f4a993 test(compression): trim batch clarify coverage to two invariants
Drop the sentinel-only batch test: a batch sentinel is already rejected by the
shared _is_clarify_non_response_sentinel list check that the existing sentinel
tests pin, so the case adds no new contract. Also add the contributor email
mapping for the cherry-picked commit so release CI can attribute it.
2026-09-09 09:21:55 -07:00
gaoanze888 8af248042c fix(compression): preserve batch clarify answers in summarize pass
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.

Closes #106077.
2026-09-09 09:21:55 -07:00
Teknium bca7cd0eb0 fix(state): projected compression tip inherits the root's title when the tip is untitled
`hermes peer dm` resolves the target's canonical Bot Chat with
GET /api/sessions?title=Bot%20Chat&include_hidden=1. list_sessions_rich
admits the hidden root via the chain search, then _project_compression_tips
overwrites every surfaced field — title included — with the live tip's. The
title is carried root->tip by the agent AFTER publish_compression_child's
transaction; a rotation cut off in between (crash, closed app, the tip's
title write failing) leaves "Bot Chat" on the ended root and NULL on the tip,
so the projected row carries title=None, the handler's exact-title filter
drops it, the peer POSTs a duplicate and the UNIQUE(title) guard answers
400 "Title already in use" (#106165).

Fix at the projection: fall back to the root's title only when the tip has
none (a titled tip keeps winning). Same COALESCE in the bounded recent-
sessions lister, the other place that projects a lineage onto its tip. This
replaces the handler-level fallback in PR #106365 (a second lookup path
bolted onto _handle_list_sessions with try/except: pass) with a 5-line fix
at the one place the title is lost, so every list consumer sees the name.

Salvage of #106365 by @finn763.
2026-09-09 09:21:46 -07:00
finn763 0c604d7e80 test(gateway): peer dm e2e against a hidden Bot Chat that rotated through compression (#106165)
Regression test from PR #106365 (one of its two tests kept: the full peer-dm
e2e; the listing-shape test asserted the same row and was dropped). The
handler-level fallback that shipped with the test is replaced by a SessionDB
fix in the follow-up commit, so this commit carries only the test.
2026-09-09 09:21:46 -07:00
1052326311 17fe7bc75f test(compression): cover cooldown rollback on a concurrently deleted session (#106271)
Regression tests carried from #106277 (its fix hunk is redundant with #106276's, picked before this).

Salvaged from #106277.
2026-09-09 09:21:39 -07:00
teknium1 02005cfe20 fix(kanban): promote refuses undone parents instead of a false --force success
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.

The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.

- drop `--force` from `promote` (parser, CLI handler, `promote_task`
  kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
  the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
  fake `ready`; the flag no longer parses

Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
2026-09-09 09:21:29 -07:00
Teknium ae65399d8f test(cron): trim kind-flip repeat tests to the two invariants
Keep the two tests that fail without the fix (one-shot -> recurring drops the
implicit times=1, recurring -> one-shot gains it) and fold the explicit-repeat
case into the first; drop the same-kind and explicit-finite cases, which only
pin behaviour the fix never touches.
2026-09-09 09:20:20 -07:00
Konstantin Khlopkov 111aea1a01 fix(cron): re-derive repeat defaults when a schedule update flips the kind 2026-09-09 09:20:20 -07:00
Totoro-qaq 5d6d5fb223 fix(cli): refresh TERMINAL_CWD when --in re-homes the session
`--in DIR` only chdir'd. Every cwd consumer (resolve_agent_cwd -> Codex
app-server thread cwd, the terminal tool, context-file discovery) prefers
TERMINAL_CWD over the process cwd, so a value inherited from a parent
Hermes surface, the shell or .env survived the chdir and the session kept
running in the old directory. The local backend was rescued by cli.py's
force-export at import time; docker/ssh backends and the TUI launch path,
which never imports cli.py, were not.

Refresh TERMINAL_CWD to the --in target when it is already set. An unset
variable stays unset so the backends keep deriving from the new process
cwd and no host path is pre-seeded into ssh/container backends.

Fixes #106220
2026-09-09 09:20:12 -07:00
teknium1 060cd7f9bb fix(state): trim the futile-holder FTS diagnostic to shape and fix the remedy text
Slim redo of the mechanism from #106410 on top of its pick (no wrappers, no
persisted "kind" enum, no process-local flag that dies with the process):

- Futility = the SAME holder PID set has blocked >= _FTS_HOLDER_FUTILE_ATTEMPTS
  (10) deferrals over >= _FTS_HOLDER_FUTILE_SECONDS (30 min); tracked as
  holders_since/holders_attempts in the persisted fts_rebuild_deferral record
  and reset whenever the holder set changes. The 3-deferral/60 s escalate
  window is the orphan-reap gate and stays as is.
- ONE escalated ERROR line names each holder pid + cmdline and the remedy that
  can actually be followed from inside a gateway session: stop ONLY the other
  holder; this process's own retry admits the rebuild within 60 s. The old
  "with the gateway stopped" advice was unrunnable from a gateway-hosted
  session (the gateway is the session) and is gone from both log and doctor.
- hermes doctor renders the futile record distinctly.
- retry_deferred_fts_recovery: a capped backoff earned by holder set X no
  longer applies once the live holder set differs from X, so stopping the
  other service is followed by a retry on the next tick, not up to an hour
  later (the issue's 16-min wait).
- Tests trimmed from 5 to 2 invariants (futile line + doctor entry after N
  same-holder deferrals; backoff reset when the holder set changes); the
  contributor's control tests for changing PIDs / orphan reap are covered by
  the existing test_repeated_deferrals_reap_inactive_orphan_then_rebuild.

The "canonical writes and LIKE search remain available" WARNING is kept
because it is true on origin/main: a stale open drops every FTS trigger, so
the messages INSERT succeeds (probed live with a real state.db + a second
process holding it). Writes fail only when a peer re-publishes triggers over
the corrupt index — a separate class, not this diagnostic.

Refs #106393
2026-09-09 09:19:57 -07:00
KoNit-K e7ca9b47ac fix(state): diagnose futile FTS deferral from a permanent holder
A supervised peer never satisfies the orphan reap, so stale-FTS repair retried forever with a misleading "canonical writes remain available" warning.

(cherry picked from commit e57f3a975d311aa44da1e92c5e727eba7c8cff70)
2026-09-09 09:19:57 -07:00
Teknium b8b2278440 test: import pytest in test_stderr_timestamp (marker needs it)
The spawns_gateway_lookalike marker was added to a module that never imported
pytest; collection failed with NameError in CI.
2026-09-09 09:19:54 -07:00
Teknium ca16cafee4 test(guard): block spawning a real gateway runtime from tests
Tests that exercise the dashboard's gateway-restart path can end up
spawning a REAL `python -m hermes_cli.main gateway restart` child when
the spawn seam is not intercepted. `_spawn_hermes_action` launches it
with start_new_session=True, so it outlives the pytest worker; the
child inherits the pytest-tmp HERMES_HOME, which is not a profile and
hashes to no service suffix, so `get_service_name()` resolves the
DEVELOPER's `hermes-gateway` unit, `systemd_restart` restarts the live
gateway, and without systemd the fallback runs `run_gateway()`
in-process forever and squats the webhook port.

Live repro on this machine (origin/main): an unintercepted spawn of
["gateway", "restart"] from a test restarted the production gateway
(MainPID 136820 -> 1689090, NRestarts=1). On 2026-09-03 a sibling
refactor moved `_spawn_hermes_action` from the `hermes_cli.web_server`
facade to `hermes_cli.web_server_gateway` ~10 minutes before the tests
were repointed; runs in that window patched a name production never
read and left 39 orphans alive for six days.

The live-system guard now rejects any subprocess primitive whose
command line the canonical matcher (`gateway.status.
_gateway_command_subcommand`) classifies as `gateway run|start|restart`.
Argv substrings are never consulted, so `gateway status`, `gateway
--help`, `hermes_cli.main serve`, etc. pass through. Three files that
deliberately spawn and reap a stub child with a gateway-shaped argv
(flock holders, sleep sleepers with an argv tail) opt out with the new
`spawns_gateway_lookalike` marker, which lifts only this check and keeps
os.kill guarded. Two canary tests pin the block and the pass-through.
2026-09-09 09:19:54 -07:00
Teknium 19f2f19987 fix(gateway): refuse to uninstall a systemd unit pinned to another HERMES_HOME
Defence at the exact boundary the incident crossed: systemd_uninstall() and
uninstall._remove_systemd_gateway() unlinked whatever get_systemd_unit_path()
returned. Before stop/disable/unlink, read the unit's own
Environment="HERMES_HOME=..." line (the parser status/refresh already use)
and, when it names a different home than this process, warn with both paths
and leave the unit alone. A unit without the line (hand-written) is still
removed as before.
2026-09-09 09:19:36 -07:00
Teknium 4746e34448 fix(gateway): foreign HERMES_HOME no longer resolves to the default hermes-gateway unit
_profile_suffix() compared HERMES_HOME against get_default_hermes_root(),
which treats ANY home outside ~/.hermes (Docker /opt/data, a mktemp dir) as
"the root itself". Every such home therefore collapsed to the bare
`hermes-gateway` service name and the default profile's unit path
(~/.config/systemd/user/hermes-gateway.service); the documented
"else a short hash of the path" branch was unreachable.

A parity harness run with HERMES_HOME=$(mktemp -d) called
uninstall_gateway_service(), resolved to the production unit, ran
`systemctl --user stop/disable`, unlinked it and daemon-reloaded. With the
unit gone Restart= could not revive it: all cron jobs and every messaging
platform were down for 6.5 days.

Compare against the platform-native default home (~/.hermes) for the bare
name; keep the profile name for <root>/profiles/<name>; everything else
(temp dirs, Docker /opt/data) gets its sha256[:8] suffix as the docstring
always promised. The Docker image supervises with s6 (`gateway-<profile>`
slots), not systemd/launchd, so the bare host-service name was never load-
bearing there.
2026-09-09 09:19:36 -07:00
kshitijk4poor f3ae63162e test(auth): keep the persisted clean-mark tests to the two invariants
A fresh process with an unchanged store takes zero auth-store locks, and a
root store that gains a forked grant invalidates the persisted mark so the
heal re-runs. The corrupt-mark fallback, same-mtime size change and
no-secrets checks were pinning implementation details of the same cache.
2026-09-09 21:17:35 +05:30
John Paul Soliva 57e03e2be6 perf(auth): persist the forked-OAuth clean mark so a fresh process skips the heal's locks
`_heal_forked_single_use_oauth_grants()` runs on every `load_pool()`, and its
clean mark lived only in memory. Every fresh `hermes` invocation and every new
worker therefore re-took the heal's two nested EXCLUSIVE auth-store locks just
to rediscover a store it had already cleared — 4 acquisitions per process on a
two-provider profile, every one of them finding nothing to consolidate. Behind
a sibling process holding those locks that costs a full
`AUTH_LOCK_TIMEOUT_SECONDS` per provider before the process can do anything at
all: measured 30.1s for two providers.

Persist the mark next to the store it describes (`<profile>/cache/
oauth_heal_clean.json`, 0600, paths and stat data only — no credential
material) and consult it BEFORE taking any lock.

Outliving the process means the mark needs a stronger key than the in-memory
one did:

- The ROOT store joins the fingerprint. This heal consolidates root → profile,
  so root acquiring a counterpart turns a row the heal deliberately KEPT into a
  fork it must strip. The in-memory mark could ignore root because it died with
  the process; a persisted mark would keep skipping a heal that has become
  necessary.
- File sizes join it too, so a metadata-preserving rewrite (`rsync -t`,
  `tar -p`, a restore) cannot leave a stale mark looking current indefinitely
  rather than for one process.

Measured on an isolated HERMES_HOME with two OAuth providers:

    lock acquisitions per fresh process    4     -> 0
    contended load_pool() x2               30.1s -> 0.00s
    mark-file reads per 100 load_pool()    -     -> 2 (one per provider)

The mark stays a cache: absent, unreadable, corrupt or wrong-shaped content all
mean "unknown" and fall through to the locked heal, and a failed write only
means the next process re-runs it — the behaviour before this cache existed.
2026-09-09 21:17:35 +05:30
kshitijk4poor e2bd400233 fix(models): only cache unreachability, not HTTP errors; key the entry via base_url_origin
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.

_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
2026-09-09 21:16:48 +05:30
kshitijk4poor 20762c67eb test(models): trim probe negative-cache tests to two invariants 2026-09-09 21:16:48 +05:30
finn763 48b8528e7c fix(desktop): stop UI freeze on unreachable provider Closes #81123 2026-09-09 21:16:48 +05:30
kshitijk4poor a4113eb994 test(dashboard): keep the read-coalescing tests to the two admission invariants
The contributor's 15-test module pinned encoding details (frozen-arg shapes,
postponed-annotation resolution, monkeypatch seams). The two behaviour
contracts that matter survive: a 12-request /api/profiles burst admits ONE
worker and leaves /api/status responsive on a 2-token pool (this file), and
the kanban board read shares one worker per key (test_kanban_read_admission).
Both go red when the coalescing wrapper is removed from the route.
2026-09-09 21:16:40 +05:30