Commit Graph

3340 Commits

Author SHA1 Message Date
Mariano Nicolini 3c4e84c166 fix(models): peek past expired and superseded pricing entries 2026-09-01 16:35:47 -03:00
Mariano Nicolini d7520b2822 fix(aux): seed the shared Nous catalog entry with the pickers' arguments 2026-09-01 16:34:48 -03:00
Mariano Nicolini 79972c6781 fix(models): expire the Nous catalog so policy changes land
A cached catalog was held for the life of the process, so a long-lived
gateway or desktop kept offering models the org had since blocked until
restart. Opt-in TTL — other providers keep no-expiry caching.
2026-08-31 16:18:07 -03:00
Mariano Nicolini e681decfae fix(aux): policy-check the whole auxiliary model ladder
Only the catalog step was filtered. With no fast-family match in the allowed
catalog it returned empty and the ladder fell through to a public
recommendation, which could hand titling a model the org blocks.
2026-08-31 15:58:27 -03:00
Mariano Nicolini 6e20ec4101 fix(nous): apply the org policy before the free/paid tier split
Rescuing an empty list after partitioning put paid models back into a
free-tier user's selectable list, and the dashboard could pick one as the
silent default. Narrowing first also drops the separate unavailable-list
filter.
2026-08-31 15:40:49 -03:00
Mariano Nicolini e89f0087b4 fix(models): key the pricing cache per credential, not per auth state 2026-08-31 15:29:53 -03:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
Mariano Nicolini da3c2435e2 fix(nous): only rescue an empty list where emptiness means "filtered out"
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
2026-08-28 13:34:27 -03:00
Mariano Nicolini 04647f15c8 fix(nous): only fall back to the reachable set when the overlap is empty
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.

Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:39:29 -03:00
Mariano Nicolini 117e7fef88 fix(nous): surface allowed models the curated list does not carry
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.

When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.

Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:24:55 -03:00
Mariano Nicolini c38d62aefe feat(nous): tell a governed org its model choice is restricted
The gateway omits a policy-blocked model from `/v1/models` rather than
marking it, so after the preceding commits a restricted model is simply
absent from the pickers. That reads as "Hermes does not support this"
instead of "your organization disallows it".

Show one line when the org is governed, in the two flows where a user
picks a model. It enumerates nothing: model policy is an allowlist, so an
org admitting a handful of models blocks the whole rest of the catalog,
and graying hundreds of rows would be a worse UI than omitting them.

Driven by the `policy_present` claim, which is tri-state — the line shows
only when it is explicitly true, because an absent claim means an older
mint rather than an unrestricted org. The claim is stamped at mint time,
so the line can lag a policy change by up to the access token's lifetime.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:17:56 -03:00
Mariano Nicolini 9fc43919cc perf(nous): stop prefetching a catalog nothing reads
The `/model` picker warms `provider_models_cache.json` in parallel before
its serial build loop, and Nous was collected into that prefetch because
the credential scan treats any auth.json providers entry as credentials
regardless of auth type.

Nothing reads the result. The picker's nous branch builds from the curated
list rather than `cached_provider_model_ids`, and Nous cannot reach the
api_key-only unified pathway that would call it. Because the prefetch
forces a refresh it also skips the cache read, so the entry is written and
never read — a live authenticated /v1/models round trip per picker open
for nothing.

Exclude it. Also add the plan this and the preceding commits implement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:14:38 -03:00
Mariano Nicolini 35e0d15861 fix(aux): read the Nous fast-model catalog with credentials, and filter it
`_fast_model_from_catalog` treats the catalog's keys as a source of ids,
scanning them for a cheap model to use for side tasks like titling. Two
problems for Nous.

The credential lookup goes through `resolve_api_key_provider_credentials`,
which raises for Nous because it is OAuth. The read then went out
anonymous and came back with the full catalog rather than the one the org
may reach, so a policy-hidden model could be selected and then refused at
request time with `model_blocked_by_org_policy`.

Fall back to the Nous credential resolver when the api-key path raises,
and narrow the resulting ids by the org policy the same way the pickers'
lists are narrowed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:14:22 -03:00
Mariano Nicolini b1ea9196f7 fix(nous): narrow every model list to the org's policy
Four surfaces list Nous models, and none of them was filtered. All four
seed from the docs-hosted curated manifest and union the Portal's
`recommended-models` endpoint; neither source is authenticated, so org
policy had no effect on the model a user picks — which is the model they
then use. The Portal endpoint compounds it, serving one globally
CDN-cached payload for the whole platform, invalidated only by admin
pricing edits and never by a policy change, so it can put a hidden model
straight back into a list.

Narrow all four against the authenticated catalog:

  - `_login_nous`, which chooses the model the session starts on
  - `_model_flow_nous`, the `hermes model` picker
  - `list_authenticated_providers`, the `/model` picker
  - `/api/model/recommended-default`, dashboard onboarding

The list stays curated and curated-ordered — the policy set only ever
subtracts. Replacing a list with the catalog's keys would swap a curated
agentic list for a large alphabetical dump of vendor-prefixed models,
which is the regression the picker's nous branch already exists to avoid.

The `/model` picker's filter sits outside the try that wraps the Portal
union, so a Portal outage still yields a policy-filtered curated list.
`_login_nous` and `_model_flow_nous` also narrow their unavailable lists,
so a policy-hidden model is not offered as a free-tier upsell either.

For an org with no policy — the common case — the filter is a no-op and
every list is what it was.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:13:53 -03:00
Mariano Nicolini c248d5356c feat(nous): read the org model policy and expose it as a list filter
A Nous team admin can restrict which models and which serving providers
their org may use. The inference gateway applies that policy to
`GET /v1/models`, omitting blocked rows with no marker field, so the keys
of an authenticated catalog read are the reachable set.

Add the two pieces the pickers need:

`nous_policy_present()` reads the `policy_present` claim off the OAuth
access token, which costs no request. `/api/oauth/account` does not carry
the claim, so this reads the token rather than going through
`get_nous_portal_account_info`. The claim is tri-state — absent means an
older mint, which is not the same as "no policy" and must not be reported
as one.

`nous_policy_allowed_ids()` turns the authenticated pricing response into
that set, reusing the cache entry a caller asking for pricing already
populates rather than issuing a second round trip. It returns None —
"leave the list alone" — for an org with no policy, for an anonymous read
whose catalog is unfiltered, and for an empty read, each of which would
otherwise narrow a list on evidence that cannot support it.

`restrict_to_nous_policy()` applies the set while preserving the caller's
order, and keeps a `:free` sibling whose base model is reachable. The
gateway admits a row when any of its requestable ids passes and treats
anything unknown as a keep, on the grounds that over-listing costs a 403
from the authoritative gate while hiding a row the gate would serve is
unrecoverable from the client. This mirrors that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:13:13 -03:00
Mariano Nicolini 4caeb02735 fix(models): key the pricing cache on auth state, not just the base URL
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.

That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.

Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.

`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:12:58 -03:00
Gille 0dfba37b11 fix(dashboard): trust configured reverse proxies (#94126)
* fix(dashboard): trust configured reverse proxies

* fix(dashboard): trust IPv6 loopback proxies
2026-08-27 10:35:22 -07:00
Brooklyn Nicholson 51e67babca fix(cli): keep Desktop liveness leases when the session cap is off
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
kshitijk4poor e941be7a81 test(gateway): adapt witness-composition harness to the SIGUSR1 in-place drain path
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
2026-08-27 22:06:17 +05:30
kshitijk4poor 5abe2e1880 docs+test(gateway): pin Windows witness-absent behavior and the probe-budget math
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
  there, so the witness is permanently absent, the payload records
  loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
  WEDGED — deliberate fail-safe (graceful drain remains the backstop).
  WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
  stays inside the documented probe budget so retuning can't silently
  blow past the 10s subprocess query tier.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor b48540701c refactor(gateway): simplify witness-probe ambiguity arms; sweep stale tick-socket nodes; clean test tempdirs
/simplify-code follow-ups on the 90502 salvage:

- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
  byte-identical — saw_node was effectively write-only. Collapsed to one
  arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
  from dead PIDs at arm time (POSIX-only liveness probe; Windows never
  creates AF_UNIX nodes) so state/ does not accumulate nodes across
  os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
  was DISPROVED for this call site — asyncio's create_unix_server
  os.remove()s an existing node before binding — but the contract is now
  pinned by test_producer_rebinds_over_stale_socket_node (a live
  producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
  longer leaks a directory per test run.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor db154edbaf test: short tmp_path for loop-tick witness sockets (macOS AF_UNIX limit)
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
2026-08-27 22:06:17 +05:30
rodrigo ca4a9ec686 fix(gateway): a single tick-socket miss must not authorize the wedge kill
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.

WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.

New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
  required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
  the loop is frozen for longer than one tick timeout but shorter than
  the wedge window -> probe is ALIVE and launchd_restart drains, never
  escalates; loop frozen for longer than the window -> WEDGED.

The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
2026-08-27 22:06:17 +05:30
rodrigo a1c83ef901 fix(gateway): interlock the stale-heartbeat wedge verdict with a loop-scheduling witness
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.

The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).

The classifier is now two-witness:
- socket answers            -> ALIVE (file age irrelevant: a stalled write
  or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
  longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
  agree the loop stopped scheduling)
- legacy payload (no flag)  -> unchanged single-witness contract: the
  legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity    -> UNKNOWN, never escalate

Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
2026-08-27 22:06:17 +05:30
Brooklyn Nicholson 2119ed7b4a test(model): cover numeric YAML provider keys in picker and CRUD
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
kshitijk4poor f3cbb262c1 fix(update): valid --ignored=matching mode; rename-only path split; shared preserve constant
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):

- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
  ignored mode'); with it, every ZIP update was refused as 'could not
  check the working tree'. The mocked tests could not see this — a new
  real-git test creates an actual repo + .gitignore and asserts the guard
  runs clean, blocks on an ignored user file, and exempts ignored
  preserved entries. --ignored=matching also reports an ignored dir as
  one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
  rename/copy status codes. Porcelain v1 does not quote plain filenames
  with spaces, so an ignored file literally named 'venv -> node_modules'
  parsed as two preserved tops and slipped past the guard into the
  destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
  instead of a comment-synced duplicate set (change-detector test added).
2026-08-27 21:07:34 +05:30
joaomarcos e64db76982 fix(update): gitignored user files also block the ZIP overlay
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.

Credit: @JoaoMarcos44, whose #87392 included this hardening.
2026-08-27 21:07:34 +05:30
liuhao1024 b0a8d16c60 test(gateway): end-to-end rendered-unit coverage for the cron drain floor
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
2026-08-27 20:39:11 +05:30
HexLab98 1282362803 test(gateway): cover TimeoutStopSec including the cron drain floor 2026-08-27 20:39:11 +05:30
kshitijk4poor b26a359de8 polish: reuse launchd domain probe, fast-observe in wedged tests, blank-line cleanup
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
  up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
  observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
2026-08-27 20:06:04 +05:30
kshitijk4poor 8872cd137c fix: verify launchd replacement PID before trusting KeepAlive (review follow-up)
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
2026-08-27 20:06:04 +05:30
Joby Ellington 7a76046a86 fix(gateway): use SIGUSR1 graceful restart on launchd, not bare SIGTERM
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.

`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:

1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
   on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
   was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
   gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
   `_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
   used by `systemd_restart()` and the updater — had no launchd call site.

2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
   0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
   systemd branch uses `_get_restart_exit_wait_budget()`
   (drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
   documents that callers falling back to a hard kill must cover both phases
   or they reintroduce #77184.

The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.

Observed on macOS 27.0 / Hermes 0.20.4:

    → Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
    ⚠ Gateway PID 49787 still running after 0.0s — restart may fail
    ⚠ Gateway drain timed out after 0s — forcing launchd restart

Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.

The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.

Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 20:06:04 +05:30
x7peeps 1348e65e26 fix(gateway): rebase launchd --replace removal onto main
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.

Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.

Refs: #79096
2026-08-27 20:06:04 +05:30
Heath Harris a71be9852e fix(kanban): close half-open tracked connection when busy_timeout PRAGMA fails
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.

Salvaged from PR #96290 (kanban slice) with regression test.
2026-08-27 19:07:15 +05:30
kshitijk4poor 8d95ab1b37 fix(serve): review follow-ups — never-raise sentinel fallback, DEVNULL stderr in split-stream test
- _write_machine_sentinel_line: wrap the print() fallback so a closed
  redirected stream (ValueError, not OSError) can't propagate out of the
  ready path and kill a healthy serve; document that pythonw port
  discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
  redirect active all server logging lands on stderr, and an unread
  stderr pipe can fill and block the child before the sentinel, flaking
  the test at the 120s timeout
2026-08-27 17:11:56 +05:30
Kitson Kelly f2dd32d3e5 fix(serve): announce READY sentinel on fd 1, not the redirected sys.stdout
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).

Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.

Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
2026-08-27 17:11:56 +05:30
Adolanium a9611f3c6f feat(models): add GLM-5.3-Flash to z.ai and OpenCode Go pickers
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
2026-08-27 04:14:31 -07:00
fangliquanflq 091cc0e8be fix(hermes_cli): scope hook timeouts and fail closed on pre_tool_call
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
2026-08-27 16:13:45 +05:30
Teknium 01f7ce5b76 feat(sessions): one-shot single-match owner backfill for legacy NULL-profile rows (#94724)
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.

Refs #94724
2026-08-27 02:17:56 -07:00
Aoshi-Dev 5ab04c764c test(hermes_cli): make test_default_path pass on native Windows
TestGetHermesHome.test_default_path asserted ~/.hermes unconditionally,
but the native Windows default is %LOCALAPPDATA%\hermes (see
hermes_constants._get_platform_default_hermes_home). Branch the
assertion by platform so the test passes everywhere.

Salvaged from PR #96003 by @Aoshi-Dev (the parse-guard half of that PR
was superseded by #96169); authorship preserved.
2026-08-27 14:02:06 +05:30
kshitijk4poor 93a29d110d fix(hermes_cli): surface fail-closed config write refusals cleanly
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):

- config_command: catch RuntimeError for set/unset and print a clean
  one-line error + exit(1) instead of a raw traceback on the primary
  'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
  ConsoleCommandError so 'hermes console' and the dashboard console
  report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
  branch — the old fallthrough claimed 'falling back to default config'
  even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.
2026-08-27 13:27:24 +05:30
fangliquanflq df8b841c0c test(hermes_cli): align malformed YAML set/unset expectations with RuntimeError 2026-08-27 13:27:24 +05:30
fangliquanflq 77d4d23fbc fix(hermes_cli): stop config set/unset from wiping user overrides on invalid YAML
Fail closed when config.yaml is unparseable or non-mapping before set/unset writes, reuse the readable-config guard to return the parsed mapping, and cover refuse/empty-mapping paths with regression tests.
2026-08-27 13:27:24 +05:30
Gille 6defe7eb6c fix(config): preserve lossy decimal values as strings 2026-08-27 11:49:58 +05:30
Cursor Agent 8246c4f92a fix(cli): repair interrupted update fleet restart
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.

Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
2026-08-26 21:38:20 -07:00
fangliquanflq cb54576b1a fix(gateway): isolate control routes from default executor 2026-08-26 21:38:20 -07:00
Teknium 824f7e081a chore: remove Windows real-profile PROOF workflow + live test
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
2026-08-26 19:25:33 -07:00
Teknium e4451ec6e5 feat(browser): close-with-approval flow for Windows real-profile (toggle arms, agent asks, blocked if still locked) [proof do-not-merge]
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
   browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
   always returns the [profile-locked] signal and the copy is refused. A later
   attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
    (new CLI subcommand) runs
   close_browser_holding_profile only when the agent has the user's OK. The
   locked error tells the agent to ask first, then run it, then retry.

- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
  message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py:  subcommand (identity+binding-verified
  tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.

Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
2026-08-26 19:25:33 -07:00
Teknium 00d5632249 chore: remove Windows real-profile PROOF workflow + live test
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 75f402c302 test(ci): assert cookie DB at either location + dump Default contents on miss [do-not-merge]
Auto-close live test asserted the legacy Default/Cookies path, but modern Chrome
writes Default/Network/Cookies. Accept either; on miss, print the copy's Default
listing so a real copy gap (vs a path-assertion bug) is visible.
2026-08-26 19:25:33 -07:00