Commit Graph

25886 Commits

Author SHA1 Message Date
Mariano Nicolini 3c4e84c166 fix(models): peek past expired and superseded pricing entries 2026-09-01 16:35:47 -03:00
Mariano Nicolini d7520b2822 fix(aux): seed the shared Nous catalog entry with the pickers' arguments 2026-09-01 16:34:48 -03:00
pefontana 9d5bb7a807 map mariano.nicolini@lambdaclass.com to entropidelic
The contributor check failed because the PR author's commit email had no
mapping under contributors/emails/.
2026-09-01 12:41:28 -03:00
Mariano Nicolini 79972c6781 fix(models): expire the Nous catalog so policy changes land
A cached catalog was held for the life of the process, so a long-lived
gateway or desktop kept offering models the org had since blocked until
restart. Opt-in TTL — other providers keep no-expiry caching.
2026-08-31 16:18:07 -03:00
Mariano Nicolini e681decfae fix(aux): policy-check the whole auxiliary model ladder
Only the catalog step was filtered. With no fast-family match in the allowed
catalog it returned empty and the ladder fell through to a public
recommendation, which could hand titling a model the org blocks.
2026-08-31 15:58:27 -03:00
Mariano Nicolini 6e20ec4101 fix(nous): apply the org policy before the free/paid tier split
Rescuing an empty list after partitioning put paid models back into a
free-tier user's selectable list, and the dashboard could pick one as the
silent default. Narrowing first also drops the separate unavailable-list
filter.
2026-08-31 15:40:49 -03:00
Mariano Nicolini e89f0087b4 fix(models): key the pricing cache per credential, not per auth state 2026-08-31 15:29:53 -03:00
Mariano Nicolini 705a10850d refactor(nous): trim comments and drop unused code 2026-08-28 17:00:26 -03:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
Mariano Nicolini a51df3864e refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:01 -03:00
Mariano Nicolini da3c2435e2 fix(nous): only rescue an empty list where emptiness means "filtered out"
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
2026-08-28 13:34:27 -03:00
Mariano Nicolini 04647f15c8 fix(nous): only fall back to the reachable set when the overlap is empty
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.

Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:39:29 -03:00
Mariano Nicolini bafaac5e61 docs(nous): correct the subtract-only claim in the policy plan
The plan stated the policy set should only ever subtract from a list. That is
wrong when an allowlist names a model the curated manifest lacks, which empties
the picker instead of narrowing it — the behaviour fixed in 117e7fef88.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:26:30 -03:00
Mariano Nicolini 117e7fef88 fix(nous): surface allowed models the curated list does not carry
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.

When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.

Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:24:55 -03:00
Mariano Nicolini c38d62aefe feat(nous): tell a governed org its model choice is restricted
The gateway omits a policy-blocked model from `/v1/models` rather than
marking it, so after the preceding commits a restricted model is simply
absent from the pickers. That reads as "Hermes does not support this"
instead of "your organization disallows it".

Show one line when the org is governed, in the two flows where a user
picks a model. It enumerates nothing: model policy is an allowlist, so an
org admitting a handful of models blocks the whole rest of the catalog,
and graying hundreds of rows would be a worse UI than omitting them.

Driven by the `policy_present` claim, which is tri-state — the line shows
only when it is explicitly true, because an absent claim means an older
mint rather than an unrestricted org. The claim is stamped at mint time,
so the line can lag a policy change by up to the access token's lifetime.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:17:56 -03:00
Mariano Nicolini 9fc43919cc perf(nous): stop prefetching a catalog nothing reads
The `/model` picker warms `provider_models_cache.json` in parallel before
its serial build loop, and Nous was collected into that prefetch because
the credential scan treats any auth.json providers entry as credentials
regardless of auth type.

Nothing reads the result. The picker's nous branch builds from the curated
list rather than `cached_provider_model_ids`, and Nous cannot reach the
api_key-only unified pathway that would call it. Because the prefetch
forces a refresh it also skips the cache read, so the entry is written and
never read — a live authenticated /v1/models round trip per picker open
for nothing.

Exclude it. Also add the plan this and the preceding commits implement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:14:38 -03:00
Mariano Nicolini 35e0d15861 fix(aux): read the Nous fast-model catalog with credentials, and filter it
`_fast_model_from_catalog` treats the catalog's keys as a source of ids,
scanning them for a cheap model to use for side tasks like titling. Two
problems for Nous.

The credential lookup goes through `resolve_api_key_provider_credentials`,
which raises for Nous because it is OAuth. The read then went out
anonymous and came back with the full catalog rather than the one the org
may reach, so a policy-hidden model could be selected and then refused at
request time with `model_blocked_by_org_policy`.

Fall back to the Nous credential resolver when the api-key path raises,
and narrow the resulting ids by the org policy the same way the pickers'
lists are narrowed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:14:22 -03:00
Mariano Nicolini b1ea9196f7 fix(nous): narrow every model list to the org's policy
Four surfaces list Nous models, and none of them was filtered. All four
seed from the docs-hosted curated manifest and union the Portal's
`recommended-models` endpoint; neither source is authenticated, so org
policy had no effect on the model a user picks — which is the model they
then use. The Portal endpoint compounds it, serving one globally
CDN-cached payload for the whole platform, invalidated only by admin
pricing edits and never by a policy change, so it can put a hidden model
straight back into a list.

Narrow all four against the authenticated catalog:

  - `_login_nous`, which chooses the model the session starts on
  - `_model_flow_nous`, the `hermes model` picker
  - `list_authenticated_providers`, the `/model` picker
  - `/api/model/recommended-default`, dashboard onboarding

The list stays curated and curated-ordered — the policy set only ever
subtracts. Replacing a list with the catalog's keys would swap a curated
agentic list for a large alphabetical dump of vendor-prefixed models,
which is the regression the picker's nous branch already exists to avoid.

The `/model` picker's filter sits outside the try that wraps the Portal
union, so a Portal outage still yields a policy-filtered curated list.
`_login_nous` and `_model_flow_nous` also narrow their unavailable lists,
so a policy-hidden model is not offered as a free-tier upsell either.

For an org with no policy — the common case — the filter is a no-op and
every list is what it was.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:13:53 -03:00
Mariano Nicolini c248d5356c feat(nous): read the org model policy and expose it as a list filter
A Nous team admin can restrict which models and which serving providers
their org may use. The inference gateway applies that policy to
`GET /v1/models`, omitting blocked rows with no marker field, so the keys
of an authenticated catalog read are the reachable set.

Add the two pieces the pickers need:

`nous_policy_present()` reads the `policy_present` claim off the OAuth
access token, which costs no request. `/api/oauth/account` does not carry
the claim, so this reads the token rather than going through
`get_nous_portal_account_info`. The claim is tri-state — absent means an
older mint, which is not the same as "no policy" and must not be reported
as one.

`nous_policy_allowed_ids()` turns the authenticated pricing response into
that set, reusing the cache entry a caller asking for pricing already
populates rather than issuing a second round trip. It returns None —
"leave the list alone" — for an org with no policy, for an anonymous read
whose catalog is unfiltered, and for an empty read, each of which would
otherwise narrow a list on evidence that cannot support it.

`restrict_to_nous_policy()` applies the set while preserving the caller's
order, and keeps a `:free` sibling whose base model is reachable. The
gateway admits a row when any of its requestable ids passes and treats
anything unknown as a keep, on the grounds that over-listing costs a 403
from the authoritative gate while hiding a row the gate would serve is
unrecoverable from the client. This mirrors that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:13:13 -03:00
Mariano Nicolini 4caeb02735 fix(models): key the pricing cache on auth state, not just the base URL
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.

That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.

Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.

`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:12:58 -03:00
kshitijk4poor 28ee6ac043 chore: map contributor emails for attribution (#95433, #94996) 2026-08-28 02:29:27 +05:30
kshitijk4poor ba7df6a4bf refactor: extract shared _coerce_positive_timeout helper (#95433)
The timeout validation (isinstance(raw, (int, float)) and not
isinstance(raw, bool) and raw > 0 → float(raw)) was duplicated
between auxiliary_client.py:_fallback_entry_timeout and
conversation_compression.py:resolve_compression_fallback_route.
Extracted into _coerce_positive_timeout to prevent drift if the
config schema evolves.
2026-08-28 02:29:27 +05:30
Shaun Eccles 6151e59d65 fix(compression): surface an unpublished stall-fallback fence at WARNING
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.

Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
2026-08-28 02:29:27 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
kshitij f7f45e5cfe Merge pull request #96623 from kshitijk4poor/chore/author-map-compression-contributors
chore: AUTHOR_MAP — shauneccles (#95433) + fedosis (#94996)
2026-08-28 02:29:19 +05:30
fangliquanflq dd401e0f15 fix(bot-mode): keep delivery runner on host backend 2026-08-27 13:49:04 -07:00
Finn763 253b9d78c1 fix(desktop): keep bot chat focused when clicking the Bots pane (#96062)
Clicking a bot row moved the layout interaction tracker to the sidebar
group, so $focusedStoredSessionId fell back to the primary selection —
which is null in Bot Mode, because bot chats open as tiles and never set
$selectedStoredSessionId. The Bots plugin reads that null 'focused'
edge as 'the chat lost the center', releases its open claim, and the
Bots home re-asserts over the still-visible chat: the UI jumps to the
list instead of staying in the chat.

$focusedStoredSessionId now answers from the main zone's active tile in
Bot Mode before falling back to the selection, so a chrome-sidebar
click no longer fabricates a null edge; a genuinely closed chat (no
tile in main) still surfaces null and the home returns as before.

Regression tests cover the sidebar-click case (red before the fix),
plus guards for the closed-chat and sessions-mode derivations.
2026-08-27 13:48:57 -07:00
Teknium db63afb9d7 fix(state): journal-mode probe/restore go through _connect_repair_durable
The write-durability contract test (test_repair_path_has_no_bare_connects)
enforces that every repair-path connection carries the macOS write
barriers; the restore rewrites the file header, so it belongs under the
same rule.
2026-08-27 13:48:51 -07:00
Teknium 4882184e95 fix(state): guest durability barriers also apply configured database.synchronous
apply_durability_barriers() is the guest-connection entry point for
secondary state.db users (async delegation ledger) that must not run
journal-mode setup. database.synchronous (#90892) rides on the
journal-mode/pragma paths guests now skip, so apply it here directly —
otherwise guest connections silently run at the compile-time default.

Follow-up to the #93012 salvage.
2026-08-27 13:48:51 -07:00
fangliquanflq 7e6eda7bbb fix(delegation): expose state durability barriers 2026-08-27 13:48:51 -07:00
fangliquanflq 6548177eb1 fix(delegation): restore ledger durability barriers 2026-08-27 13:48:51 -07:00
fangliquanflq 4a8b4d43a3 fix(delegation): preserve state database journal mode 2026-08-27 13:48:51 -07:00
liuhao1024 e40f1be759 fix(state): route the post-repair journal-mode restore through the canonical path
Review feedback on #89681: the restore issued the switch pragma
directly via _set_journal_mode_no_wait, which bypassed the
vulnerable-SQLite WAL-reset gate — on the reporter's own runtime
(SQLite 3.50.4) it re-enabled WAL on the one file that had just
demonstrated it can corrupt, exactly what the gate exists to prevent
(a rebuilt file is a new database). It also skipped the WAL companions
(size limit, checkpoint barrier, synchronous=FULL) and used the
leave-WAL helper for an enter-WAL switch.

The restore now calls apply_wal_with_fallback, inheriting the gate,
the macOS-NFS silent-refusal handling and the companions from the
front door, and the pre-surgery mode is probed (best-effort; a
malformed file may refuse) and recorded in the report so the
before/after WARNING only fires when the comparison is honest.
2026-08-27 13:48:51 -07:00
liuhao1024 786e65bf4f fix(state): re-apply the configured journal mode after corruption repair
Corruption can drop the WAL bit from the database header, and every
repair strategy rebuilds or rewrites the file in place — so a repaired
store comes back in the default journal mode (delete). The WAL-reset
gate at open time never sees the flip because it happens inside the
repair path, not at open (the open-time flip #89393 warns about is a
different door), leaving the operator no signal that a WAL store moved
to DELETE (#89674).

After a successful repair, re-apply the canonical database.journal_mode
through _set_journal_mode_no_wait (concurrent openers abort the flip
instead of sneaking it between transactions) and log a WARNING naming
the old and restored modes. Best-effort: a refused restore is logged,
never raised — the repair itself already succeeded. Probing the damaged
file for a pre-repair mode is deliberately not trusted: the corruption
itself is what drops the WAL bit, so the configured setting is the only
reliable target.
2026-08-27 13:48:51 -07:00
Teknium 9a9e9074cb style: sort MINIMIZED_TRACK import (perfectionist lint) for salvaged #95956 2026-08-27 13:48:45 -07:00
Teknium 9522c4e8e2 chore: map contributor email for salvage of #95956 2026-08-27 13:48:45 -07:00
Thomas Bekkers dbca7a4f02 fix(hermes-bots): keep the Cronjobs tile registered while it holds focus in Bot Mode
Clicking the Cronjobs tile shifts focus onto the tile itself, momentarily
dropping bot-chat workspace ownership — syncRoutinesPane then unregistered
the pane out from under the user's own click, with no way back. Keep the
tile while Bot Mode is on screen and the tile is the focused surface;
leaving Bot Mode still unregisters as designed. Live-verified.
2026-08-27 13:48:45 -07:00
Thomas Bekkers 584f3a748b fix(desktop): keep a restore tab when a pane or strip collapses (#91223)
Hiding the Sessions/Bots strip, or tapping the header of a lone docked
tile (Cronjobs and any plugin pane beside the workspace), left no mouse
path back: the restore menu lived on chrome the gesture just unmounted,
and a row-collapsed rail could size to 0px.

Treat hide-only chrome as stranded so `never` cannot hide those chips.
Stop collapsing on header tap (chevron only). Size a minimized zone to
MINIMIZED_TRACK and keep the horizontal strip when two or more tabs
remain.
2026-08-27 13:48:45 -07:00
kshitijk4poor 939dec1348 fix: harden _is_recoverable_error_job against schedule=None
job.get("schedule", {}).get("kind") crashes with AttributeError when
schedule is present but explicitly None (disk corruption edge case).
Use (job.get("schedule") or {}).get("kind") instead, which safely
returns False for None. This pattern is already used at other sites
in the file (e.g. cron/scheduler.py line 170).
2026-08-28 02:18:25 +05:30
pierrenode ba4c2d5253 fix(cron): make a recurring job stuck in state=error recoverable again
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:

* state=completed: a one-shot that genuinely has no more occurrences,
  ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
  fails to produce a next occurrence (e.g. the croniter package is
  missing at runtime). _mark_job_run_locked's own comment is explicit:
  "Recurring jobs must NEVER be silently disabled" (issue #16265) — the
  job is left enabled=True specifically so it keeps being a live,
  recoverable job once the underlying issue resolves.

Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:

* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
  its own is_terminal_job() check) never runs, because the check itself
  skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
  ... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
  itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
  and claim steps the scheduler's own dispatch loop calls immediately
  after get_due_jobs() for anything that DOES make it into the due list —
  both also refuse the job, so even a manually-recovered next_run_at
  would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
  broken recurring job through the normal path.

The only way out was deleting the job and recreating it.

Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.

Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.

New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.

Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
2026-08-28 02:18:25 +05:30
kshitijk4poor 29d1c02699 chore: map contributor emails — shauneccles (#95433) + fedosis (#94996)
Adds contributor email→username mappings via the new
contributors/emails/ system (one file per email) for two
compression contributors whose salvage PRs need attribution CI.
2026-08-28 02:10:33 +05:30
Brooklyn Nicholson a24c12d14f fix(desktop): gate transcript budget cap so Show earlier works
The render-phase cap snapped a visible pane's Show-earlier growth back
on the next render, so the button did nothing. Clamp only hot-hidden
panes, and grow the DOM budget when expanding the store window too.

Supersedes #87686.

Co-authored-by: Kirk <317508070+chukirk-svg@users.noreply.github.com>
Co-authored-by: Per0 <175494353+Per0-1@users.noreply.github.com>
2026-08-27 14:51:13 -05:00
Teknium d91e4376f1 Merge pull request #65108 from NousResearch/hermes/hermes-793f4fd9
feat(skills): rewrite AgentMail optional skill CLI-first (salvages #60811)
2026-08-27 12:45:01 -07:00
hope b39d76d902 feat(tools): session-persistent kernels for execute_code (kernel_mode: session) (#94647)
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)

execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.

Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.

Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.

Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.

Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).

* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority

Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.

1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.

2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.

Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.

Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).

* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)

---------

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-08-27 12:26:31 -07:00
Gille 0dfba37b11 fix(dashboard): trust configured reverse proxies (#94126)
* fix(dashboard): trust configured reverse proxies

* fix(dashboard): trust IPv6 loopback proxies
2026-08-27 10:35:22 -07:00
Brooklyn Nicholson 39f1e1881a fix(tui-gateway): spare durable rows while a sibling backend holds them
Automatic Desktop ends (ws_orphan_reap, disconnect, idle, LRU, shutdown)
now drop the local runtime but keep the state.db row and durable-key
delegations when another live lease still owns the session.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
Brooklyn Nicholson 51e67babca fix(cli): keep Desktop liveness leases when the session cap is off
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
kshitijk4poor f6f707b783 chore: suppress posix-gated os.kill probe in footgun scan; map rodrigogs in AUTHOR_MAP
- The stale tick-socket sweep's os.kill(pid, 0) liveness probe sits inside
  an explicit os.name == 'posix' gate (AF_UNIX nodes never exist on
  Windows) — suppress with the standard inline marker.
- contributors/emails/: rodrigo.smscom@gmail.com -> rodrigogs (author of
  the salvaged #92315 commits), unblocking check-attribution.
2026-08-27 22:06:17 +05:30
kshitijk4poor e941be7a81 test(gateway): adapt witness-composition harness to the SIGUSR1 in-place drain path
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
2026-08-27 22:06:17 +05:30
kshitijk4poor 5abe2e1880 docs+test(gateway): pin Windows witness-absent behavior and the probe-budget math
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
  there, so the witness is permanently absent, the payload records
  loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
  WEDGED — deliberate fail-safe (graceful drain remains the backstop).
  WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
  stays inside the documented probe budget so retuning can't silently
  blow past the 10s subprocess query tier.
2026-08-27 22:06:17 +05:30