_set_worker_pid on a fake PID (54321) now persists the 'unverified' marker, under which
the archive path correctly refuses to signal; the test pins the verified-spawn behaviour,
so it stubs the fingerprint capture.
kvnloo (#111620 review) asked for the operator-visible statement that once a serve /
dashboard process hosts a second profile, the launch profile's env-only credentials are
frozen at activation and a rotation in the process env needs a restart. Also states the
routed-child rule with and without the multiplex flag (#111617). Adds encoding='utf-8'
to the probe child in the child-env authority test (windows-footguns lint).
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):
- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
a reboot, and that counter does not, so an unrelated process on a later boot with the
same PID and the same tick value passed _start_times_agree(). The fingerprint is now
"<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
the witness the drain marker already uses); both halves must match. Integer values on
rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
row and falls back to bare PID existence - a new spawn silently recreated the #89614/
#99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
and must not outlive its pid.
Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.
Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
Three edges of the fail-closed multi-profile host (#111620 review, andrexibiza P1 + P2,
kvnloo finding 1):
- `send` under a routed profile's scope `update()`d the installed scope from raw `.env`,
reversing build_profile_secret_scope's precedence (user .env, then external secret
sources) for the rest of the request; a stale user value beat the secret-manager one.
The installed scope is authoritative as-is; only the config.yaml setdefault bridge runs.
- The launch profile's body was scoped only when `is_multiplex_active()` was already true
at entry, while get_secret consults that global on every read. A launch RPC / dashboard
request entering single-profile and resuming after a concurrent first `?profile=B`
activation raised UnscopedSecretError mid-request. The launch profile's secret scope
(its .env + external sources over the launch env: live while single-profile, the frozen
snapshot once multiplexing is active) is now bound for every launch-profile body, so the
credential source is fixed at entry. The terminal policy overlay stays multiplex-only
(standalone terminal execution keeps its os.environ bridge). _publish_env_value mirrors a
same-request .env write into that scope AND os.environ for the launch profile, only into
the scope for a routed one (serves_routed_profile).
- _release_profile_runtime_scope_tokens reset terminal → secret → home in sequence under
one outer suppress; a failing terminal reset left the previous profile's secrets and
HERMES_HOME installed for the next body in that context. Each reset is now independent;
the first failure is re-raised after every scope is released.
tests/tui_gateway/test_multi_profile_hosting_transitions.py: manager-vs-dotenv precedence
through _load_hermes_env, TUI-RPC and dashboard barrier tests (launch enters single-profile,
B activates on another thread, launch resumes and still resolves its injected credential,
never B's), forced terminal-reset failure still releases secret + home. 4/4 red on base.
Two authority gaps in served_profile_child_env (#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):
- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
provider credentials; strip_launch_profile_env only knows names with .env/source
provenance, so a key systemd/Compose/the shell injected into the launch process
survived into profile B's child whenever B did not define the same name. Now a ROUTED
target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
before B's own scope is overlaid (the child boundary gets get_secret's contract: a
scoped miss is no credential, never ambient fallback). The launch profile's own child
keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
dashboard backends serve ?profile=B by installing the HERMES_HOME override without
that flag, so B's slash worker / helper children kept A's .env and settings. The
authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
thread); it now raises UnscopedSecretError like get_secret.
tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
Three P1 findings from the review of #111062 (fixes#110850 remainder):
- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
installed-but-dead unit; `systemd_install` writes the unit before the start that
can still be killed, so an apply interrupted at default/start (or mid-restart of
a stopped default) was reported as "already multiplexed" and never recovered.
`interrupted` is now derived from the manifest vs the LIVE default's recorded
served_profiles: flag on + manifest + not every migrated profile served = resume.
- the compensation `try` began after `_remove_secondary_gateways` and the flag
write, so a later secondary's stop, a unit's daemon-reload or the config write
failing left the first secondary removed with no rollback. The flag write,
removals and default bring-up now all sit inside one boundary that rolls back
through the manifest; `_preflight_apply` refuses failures knowable from the plan
(root system unit without a recorded User=, unresolvable recorded user, an
unwritable config.yaml) before any working gateway is stopped.
- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
`default_uid is None`; a default system unit whose User= this host cannot
resolve is the same boundary from the other side and now blocks the update
hook (same known uid still folds, different known uid still refuses).
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.
Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
Background writers that still carry a tombstoned profile as their Hermes
home (reasoning-caps warm thread, models cache, models.dev ETag, gateway
lifecycle ledger, MCP OAuth tokens, memory store) re-created
profiles/<name>/ with a bare mkdir right before an atomic write. Route the
parent-dir creation through mkdir_under_hermes_home so a deleted named
profile raises FileNotFoundError and stays gone, matching the tombstone
contract already enforced for logging and state.
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.
- `hermes profile migrate-identity <old> <new>`: retries the migration —
delegates to the gateway control verb while a multiplexer is live, performs the
durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
naming the offending database on a collision, a lock, or a partial failure. Only
the name format and the existence of the new profile are checked; the old profile
directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
`click`: a failed second database raised NameError instead of printing its warning.
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
_teardown_session swallows agent.close() failures, so a "row stays open"
assertion alone is green even when close() aborted before
_finalize_owned_session_row. Assert the one side effect only that finalizer
produces (_owns_session_db cleared), give the stub the _memory_manager attr
close() reads directly, and assert the flag is untouched on the control
paths. Rename the sole-owner cross-process test to what it now asserts.
_finalize_session decides whether the TUI backend ends the state.db row
(gateway-owned source → no; another backend holds the lease → no; automatic
Desktop reclaim → no, #105588). _teardown_session then calls agent.close(),
whose _finalize_owned_session_row ends the same row as "agent_close" through
the agent's own handle whenever _end_session_on_close is still True — so every
spare decision was undone one call later. Verified with real SessionDB +
AIAgent.close(): with only the #105620 guard a Desktop ws_orphan_reap still
lands ended_at + end_reason="agent_close" (the #105588 follow-up report).
One seam after the lifecycle decision now clears agent._end_session_on_close
whenever the row was spared, replacing the gateway-owned-only assignment from
#111145 so the lease-held and Desktop-reclaim cases get the same treatment.
The five mocked _finalize_session tests from #105620 are replaced by two
real-DB _teardown_session tests beside the #111145 harness (red on both
origin/main and the bare cherry-picks, green on the stack).
ws_orphan_reap, idle_timeout, lru_evict, ws_disconnect, and tui_shutdown
are runtime/connection GC — not user intent. Ending the durable session
row during these automatic cleanup reasons confuses GC with user action,
causing secondary bot chats to vanish from the sidebar even though the
transcript is intact in state.db.
The canonical Bot Chat already has resurrection paths for accidental
ws_orphan_reap ends, but non-canonical secondary chats do not, so they
are hit harder.
Skip db.end_session() when _desktop_automatic_cleanup is True (automatic
cleanup reason + Desktop source). Runtime is still reclaimed, the
session.reclaimed event still fires, and explicit user close/archive/reset
still ends the durable row normally.
Fixes#105588
Two parametrized invariant tests (native client sync/async; None or raising
profile falls back to the standard client) replace four near-duplicates, and
the hook's kwargs are pinned exactly to the mapping openai.OpenAI would have
received. The fixture isolates both provider registries through monkeypatch
instead of a hand-rolled snapshot/restore, drops the HERMES_HOME override that
tests/conftest.py already provides, and replaces the tuple-truthiness lambda
with a plain function. Call-site comment trimmed to the ordering WHY.
The auxiliary api_key branch built openai.OpenAI directly, so an
out-of-tree provider registered with auth_type="api_key" lost its
native transport for auxiliary tasks even though the main-agent path
(_provider_supplied_client) and the external_process branch honor the
same hook. Consult the profile's create_client() before the built-in
gemini/OpenAI ladder; None (the default) falls through untouched, and
a raising profile is logged and skipped (#112384).
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.
Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
Workspace moves stamp agent.session_cwd so a lazily started Codex thread
starts in the moved-to directory. Agents that never had the attribute
(test doubles, slotted objects) must keep working, so stamp only when the
attribute exists. Test stubs of _set_session_context mirror the new cwd
kwarg, and the Codex gateway test asserts the contract (no pinned session
cwd) instead of the attribute's absence.
Salvage of #93452 (@outpoints). Keep one invariant per fix:
resolver-level "automatic title never remaps a strategy session",
integration "provider routes by logical workspace, not process cwd",
agent-level "title provenance + cwd reach the provider", deferred
Desktop/TUI build threads the session cwd, seeded branch titles are
derived, and the workspace-move E2E. Drop the plumbing/legacy-shape
tests that re-assert the same contract.
acp_adapter: AIAgent(cwd=...) now stamps session_cwd itself, so the
direct assignment after construction was a duplicate.
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.
(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
The Honcho provider resolved per-repo/per-directory session names from
os.getcwd(), which is the backend process launch directory on
Desktop/gateway hosts (typically $HOME), not the user's workspace. With
a manual sessions map entry for the home directory, every Desktop
conversation landed in that fallback bucket instead of the project's
per-repo session.
Use agent.runtime_cwd.resolve_agent_cwd() — the same single source of
truth already used for system-prompt and context-file discovery — so
Honcho session routing agrees with everything else about where the
agent logically lives. It honors the pinned session cwd, then
TERMINAL_CWD, then the launch directory; CLI sessions launched inside a
project resolve identically either way.
Adds a regression covering the Desktop-style case: backend launched in
$HOME, workspace elsewhere, home-directory manual map present.
Refs #24740
(cherry picked from commit cfb32757de2509dd404acf98311e357d73fa741e)
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.
Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.
Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.
Fixes#24740
(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.
- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
into config.yaml (--force included); the env writer's denylist
(HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
bridging the value into os.environ; a name neither registered nor in the
environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
and lowercase bare keys are unchanged.
Fixes#111848 (first half landed in #112250).
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.
After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.
Fixes#111804 (its discovery half landed in #112293).
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
The per-session notification poller ran `_poll_bot_live_delivery_once` every
0.5 s. Once a "Bot Chat" session exists, each pass opens state.db and takes
the exclusive active-session registry lock; on Windows (`msvcrt` LK_LOCK
gives up after 10 s of contention) that raised
`RuntimeError: active session file lock unavailable` and the loop logged
`Bot live-owner delivery poll failed` on every attempt — 8,838 warnings in
three days, 91% of one install's WARNING output (#111719).
Two changes:
- `tools/bot_live_delivery.has_mailbox`: the mailbox directory is created
only when a delivery is first admitted, so a profile without it has nothing
to claim — the poll now returns before the state.db open / registry lock.
This keeps cron→Bot Chat and Bot Mode DM delivery intact on installs with
no messaging platform configured (both deliver through this mailbox), which
is why the poll is gated on the mailbox rather than on connected platforms
(PR #111733's guard would have broken those).
- `_poll_bot_live_delivery_guarded`: a failing poll backs off 5 s before the
next attempt and is logged at WARNING once per 60 s window (with the count
of suppressed repeats), DEBUG otherwise.
Live probe (real poller loop, temp HERMES_HOME with a Bot Chat row and the
registry lock made unavailable, 3 s):
before: owner_lookups=6 WARNING=6 (with or without a mailbox)
after: no mailbox -> owner_lookups=0 WARNING=0; mailbox -> owner_lookups=1 WARNING=1
Fixes#111719
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.
`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.
`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.
This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.
Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.
Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
The salvaged commit shipped two tests: a disconnect-mid-stream test that reads
the reply back through GET /v1/runs/{run_id}, and a structural guard that
opened api_server.py and searched its text for `output=`. Tests that read
source are change-detectors, not invariants; the behavioural test already
fails the moment `output` disappears from the terminal write, so the guard is
gone. The inline comment shrinks to the WHY (parity with /v1/runs, the
recovery path) and drops the narrative.
`POST /api/sessions/{id}/chat/stream` mints a run_id and writes run status,
but its terminal write omitted `output=` — the reply text went only onto the
SSE queue. `POST /v1/runs` has always recorded it.
The asymmetry costs a caller its answer. A client whose socket dies mid-turn —
a sleeping laptop, a dropped WiFi link, a peer DM over a flaky LAN — sees the
run reach "completed" through GET /v1/runs/{run_id} and has no way to learn
what the agent said. Worse, the disconnect path interrupts the agent, which
unwinds and *returns* a partial result, so that partial answer is recorded as
a clean "completed" and then discarded: indistinguishable from a run that
produced nothing, and equally unrecoverable.
Record `output` on this route too. Terminal statuses are already retained for
_RUN_STATUS_TTL (3600s), so the existing GET /v1/runs/{run_id} becomes a
recovery path for any client that loses its stream, without a new endpoint,
without touching the run registries, and with no change to the streaming
contract.
Deliberately nothing else: the detached turn stays out of _active_run_tasks
(it is already counted via _inflight_agent_runs, and a task entry would
double-count it in the shutdown drain), and the run stays out of
_run_streams_created, so the orphan sweeper's 300s reap still cannot see it.
Tests: a disconnect mid-stream now leaves the reply readable both in the run
record and through _handle_get_run; a structural guard asserts BOTH routes
still pass output=, anchored on the status write rather than the SSE payload's
"completed": True key, so the asymmetry cannot quietly return.
The MoA aggregator and test stand-ins never merge extra_body; the bypass
must hand them the kwargs untouched or the conversation would be sent
empty. Fold that control into the existing rail test (still two tests).
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
The cherry-picked helper deleted 'tools' from the typed kwargs, so the SDK's
post-transform extra_body merge appended it after the caller's extra_body keys —
equal dict, different bytes (byte-keyed prompt caches would miss). keep_slots=True
leaves [] placeholders that the merge overwrites in place. Drop the invented
HERMES_CHAT_SDK_TRANSFORM env var; the pre-existing HERMES_CODEX_SDK_TRANSFORM
hatch from #93650 now disables both API families. Tests trimmed to the two
invariants (byte-identity incl. caller extra_body precedence; escape hatch).
`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. #93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.
Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:
101 msgs / 76 KB 13.6 ms -> 1.2 ms
401 msgs / 190 KB 48.6 ms -> 2.0 ms
1601 msgs / 650 KB 188.7 ms -> 5.8 ms
and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.
The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.
Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.
Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.
The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.