A cached catalog was held for the life of the process, so a long-lived
gateway or desktop kept offering models the org had since blocked until
restart. Opt-in TTL — other providers keep no-expiry caching.
Only the catalog step was filtered. With no fast-family match in the allowed
catalog it returned empty and the ladder fell through to a public
recommendation, which could hand titling a model the org blocks.
Rescuing an empty list after partitioning put paid models back into a
free-tier user's selectable list, and the dashboard could pick one as the
silent default. Narrowing first also drops the separate unavailable-list
filter.
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.
Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.
When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.
Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gateway omits a policy-blocked model from `/v1/models` rather than
marking it, so after the preceding commits a restricted model is simply
absent from the pickers. That reads as "Hermes does not support this"
instead of "your organization disallows it".
Show one line when the org is governed, in the two flows where a user
picks a model. It enumerates nothing: model policy is an allowlist, so an
org admitting a handful of models blocks the whole rest of the catalog,
and graying hundreds of rows would be a worse UI than omitting them.
Driven by the `policy_present` claim, which is tri-state — the line shows
only when it is explicitly true, because an absent claim means an older
mint rather than an unrestricted org. The claim is stamped at mint time,
so the line can lag a policy change by up to the access token's lifetime.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The `/model` picker warms `provider_models_cache.json` in parallel before
its serial build loop, and Nous was collected into that prefetch because
the credential scan treats any auth.json providers entry as credentials
regardless of auth type.
Nothing reads the result. The picker's nous branch builds from the curated
list rather than `cached_provider_model_ids`, and Nous cannot reach the
api_key-only unified pathway that would call it. Because the prefetch
forces a refresh it also skips the cache read, so the entry is written and
never read — a live authenticated /v1/models round trip per picker open
for nothing.
Exclude it. Also add the plan this and the preceding commits implement.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_fast_model_from_catalog` treats the catalog's keys as a source of ids,
scanning them for a cheap model to use for side tasks like titling. Two
problems for Nous.
The credential lookup goes through `resolve_api_key_provider_credentials`,
which raises for Nous because it is OAuth. The read then went out
anonymous and came back with the full catalog rather than the one the org
may reach, so a policy-hidden model could be selected and then refused at
request time with `model_blocked_by_org_policy`.
Fall back to the Nous credential resolver when the api-key path raises,
and narrow the resulting ids by the org policy the same way the pickers'
lists are narrowed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four surfaces list Nous models, and none of them was filtered. All four
seed from the docs-hosted curated manifest and union the Portal's
`recommended-models` endpoint; neither source is authenticated, so org
policy had no effect on the model a user picks — which is the model they
then use. The Portal endpoint compounds it, serving one globally
CDN-cached payload for the whole platform, invalidated only by admin
pricing edits and never by a policy change, so it can put a hidden model
straight back into a list.
Narrow all four against the authenticated catalog:
- `_login_nous`, which chooses the model the session starts on
- `_model_flow_nous`, the `hermes model` picker
- `list_authenticated_providers`, the `/model` picker
- `/api/model/recommended-default`, dashboard onboarding
The list stays curated and curated-ordered — the policy set only ever
subtracts. Replacing a list with the catalog's keys would swap a curated
agentic list for a large alphabetical dump of vendor-prefixed models,
which is the regression the picker's nous branch already exists to avoid.
The `/model` picker's filter sits outside the try that wraps the Portal
union, so a Portal outage still yields a policy-filtered curated list.
`_login_nous` and `_model_flow_nous` also narrow their unavailable lists,
so a policy-hidden model is not offered as a free-tier upsell either.
For an org with no policy — the common case — the filter is a no-op and
every list is what it was.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A Nous team admin can restrict which models and which serving providers
their org may use. The inference gateway applies that policy to
`GET /v1/models`, omitting blocked rows with no marker field, so the keys
of an authenticated catalog read are the reachable set.
Add the two pieces the pickers need:
`nous_policy_present()` reads the `policy_present` claim off the OAuth
access token, which costs no request. `/api/oauth/account` does not carry
the claim, so this reads the token rather than going through
`get_nous_portal_account_info`. The claim is tri-state — absent means an
older mint, which is not the same as "no policy" and must not be reported
as one.
`nous_policy_allowed_ids()` turns the authenticated pricing response into
that set, reusing the cache entry a caller asking for pricing already
populates rather than issuing a second round trip. It returns None —
"leave the list alone" — for an org with no policy, for an anonymous read
whose catalog is unfiltered, and for an empty read, each of which would
otherwise narrow a list on evidence that cannot support it.
`restrict_to_nous_policy()` applies the set while preserving the caller's
order, and keeps a `:free` sibling whose base model is reachable. The
gateway admits a row when any of its requestable ids passes and treats
anything unknown as a keep, on the grounds that over-listing costs a 403
from the authoritative gate while hiding a row the gate would serve is
unrecoverable from the client. This mirrors that.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.
That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.
Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.
`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
apply_durability_barriers() is the guest-connection entry point for
secondary state.db users (async delegation ledger) that must not run
journal-mode setup. database.synchronous (#90892) rides on the
journal-mode/pragma paths guests now skip, so apply it here directly —
otherwise guest connections silently run at the compile-time default.
Follow-up to the #93012 salvage.
Review feedback on #89681: the restore issued the switch pragma
directly via _set_journal_mode_no_wait, which bypassed the
vulnerable-SQLite WAL-reset gate — on the reporter's own runtime
(SQLite 3.50.4) it re-enabled WAL on the one file that had just
demonstrated it can corrupt, exactly what the gate exists to prevent
(a rebuilt file is a new database). It also skipped the WAL companions
(size limit, checkpoint barrier, synchronous=FULL) and used the
leave-WAL helper for an enter-WAL switch.
The restore now calls apply_wal_with_fallback, inheriting the gate,
the macOS-NFS silent-refusal handling and the companions from the
front door, and the pre-surgery mode is probed (best-effort; a
malformed file may refuse) and recorded in the report so the
before/after WARNING only fires when the comparison is honest.
Corruption can drop the WAL bit from the database header, and every
repair strategy rebuilds or rewrites the file in place — so a repaired
store comes back in the default journal mode (delete). The WAL-reset
gate at open time never sees the flip because it happens inside the
repair path, not at open (the open-time flip #89393 warns about is a
different door), leaving the operator no signal that a WAL store moved
to DELETE (#89674).
After a successful repair, re-apply the canonical database.journal_mode
through _set_journal_mode_no_wait (concurrent openers abort the flip
instead of sneaking it between transactions) and log a WARNING naming
the old and restored modes. Best-effort: a refused restore is logged,
never raised — the repair itself already succeeded. Probing the damaged
file for a pre-repair mode is deliberately not trusted: the corruption
itself is what drops the WAL bit, so the configured setting is the only
reliable target.
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:
* state=completed: a one-shot that genuinely has no more occurrences,
ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
fails to produce a next occurrence (e.g. the croniter package is
missing at runtime). _mark_job_run_locked's own comment is explicit:
"Recurring jobs must NEVER be silently disabled" (issue #16265) — the
job is left enabled=True specifically so it keeps being a live,
recoverable job once the underlying issue resolves.
Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:
* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
its own is_terminal_job() check) never runs, because the check itself
skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
and claim steps the scheduler's own dispatch loop calls immediately
after get_due_jobs() for anything that DOES make it into the due list —
both also refuse the job, so even a manually-recovered next_run_at
would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
broken recurring job through the normal path.
The only way out was deleting the job and recreating it.
Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.
Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.
New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.
Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Automatic Desktop ends (ws_orphan_reap, disconnect, idle, LRU, shutdown)
now drop the local runtime but keep the state.db row and durable-key
delegations when another live lease still owns the session.
Co-authored-by: metamindedu <metamind@kakao.com>
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.
Co-authored-by: metamindedu <metamind@kakao.com>
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
there, so the witness is permanently absent, the payload records
loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
WEDGED — deliberate fail-safe (graceful drain remains the backstop).
WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
stays inside the documented probe budget so retuning can't silently
blow past the 10s subprocess query tier.
/simplify-code follow-ups on the 90502 salvage:
- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
byte-identical — saw_node was effectively write-only. Collapsed to one
arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
from dead PIDs at arm time (POSIX-only liveness probe; Windows never
creates AF_UNIX nodes) so state/ does not accumulate nodes across
os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
was DISPROVED for this call site — asyncio's create_unix_server
os.remove()s an existing node before binding — but the contract is now
pinned by test_producer_rebinds_over_stale_socket_node (a live
producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
longer leaks a directory per test run.
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.
WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.
New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
the loop is frozen for longer than one tick timeout but shorter than
the wedge window -> probe is ALIVE and launchd_restart drains, never
escalates; loop frozen for longer than the window -> WEDGED.
The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.
The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).
The classifier is now two-witness:
- socket answers -> ALIVE (file age irrelevant: a stalled write
or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
agree the loop stopped scheduling)
- legacy payload (no flag) -> unchanged single-witness contract: the
legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity -> UNKNOWN, never escalate
Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
`loop_heartbeat_forever` wrote the heartbeat inline on the gateway loop. That
write ends in `atomic_json_write` -> `os.fsync`, and on a stalling filesystem the
fsync blocks whichever thread runs it — which was the loop the liveness watchdog
exists to monitor. So the watchdog timed out its probe
(DEFAULT_LOOP_WATCHDOG_TIMEOUT_S = 10s, MAX_STRIKES = 3, a ~90-120s budget) and
took the hard exit, for a loop that was unresponsive because it was blocked
inside the watchdog's own liveness write.
Measured on the install that prompted this: a WSL2 VHDX under io pressure
(/proc/pressure/io full avg300 = 7.60) stalled a trivial stat-and-fsync probe at
p99 31s and max 112s — longer than the entire watchdog budget. Two of that day's
three gateway restarts carry byte-identical stack dumps parked in the heartbeat
writer.
The write now goes to a thread. Awaited, not fire-and-forget, and that distinction
is the whole design: the docstring promises that a frozen loop lets the file age,
because that staleness is how an external supervisor notices. The loop still
initiates and awaits the write, so a wedged loop still stops refreshing the file —
while a blocked fsync no longer stops the loop from answering the probe. One write
in flight at a time, so a 112s stall cannot queue a thread per interval behind it.
Scope: the timeout and strike defaults are untouched. Raising them would only
delay the same kill, and picking a budget above this box's p99 is a deployment
decision, not a fix.
2 tests. The behavioural one patches a slow write and asserts the loop still
completes ~15 ticks while it blocks; it is bounded by a fixed sleep rather than an
Event handshake so a regression fails on the tick count instead of hanging. The
second pins that the write stays awaited, since fire-and-forget would pass the
first test while destroying the staleness signal.
Verified by mutation: reverting only gateway/shutdown_watchdog.py fails both new
tests and leaves the 5 existing ones green. 15 passed across
test_loop_liveness_watchdog / test_shutdown_watchdog /
test_systemd_watchdog_lifecycle / test_watchdog_review_76354.
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.
Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.
Credit: @Finn763
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.
Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.
Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):
- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
ignored mode'); with it, every ZIP update was refused as 'could not
check the working tree'. The mocked tests could not see this — a new
real-git test creates an actual repo + .gitignore and asserts the guard
runs clean, blocks on an ignored user file, and exempts ignored
preserved entries. --ignored=matching also reports an ignored dir as
one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
rename/copy status codes. Porcelain v1 does not quote plain filenames
with spaces, so an ignored file literally named 'venv -> node_modules'
parsed as two preserved tops and slipped past the guard into the
destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
instead of a comment-synced duplicate set (change-detector test added).
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.
Credit: @JoaoMarcos44, whose #87392 included this hardening.
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.
Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.
Original PR #32352 by @hbentel, salvaged onto current main.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).
Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.
Fixes#93820
Carried from #94034 (closed as duplicate of #94033): explicit regression
that a jobs.json-edited stale next_run_at with NO manual_run_at marker
still re-anchors without firing — the direct #93049 protection case.
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
cli-config.yaml.example said macOS is "always held at FULL regardless",
which reads as "your setting is ignored on this platform" and would talk
an operator out of choosing EXTRA. _apply_synchronous_pragma only refuses
values BELOW FULL on Darwin; EXTRA is applied normally.
The existing doc guard only asserted the key's presence, so it could not
have caught this. Pin the distinction instead.
Reported by @Enough1122 in review.