Commit Graph

3 Commits

Author SHA1 Message Date
Teknium 09adb5e2bd refactor(gateway/slash): unify cached-agent lookup, session-db reply, approval delivery, checkpoint mgr, approval setters; table-drive /busy /voice /rollback /fast 2026-09-02 15:32:49 -07:00
Teknium c5c9aa8d44 fix(gateway): hygiene compaction keeps the profile secret scope under multiplexing
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').

- gateway/run.py: run both hygiene executor hops (detached-agent path and
  codex app-server path) inside copy_context().run, keeping the default
  executor so a fence-cancelled hung summary never occupies a gateway
  agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
  class failure — abort and preserve the session instead of dropping the
  middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
  ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
  stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
  run_in_executor); UnscopedSecretError classified as access failure.

Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
2026-09-01 22:28:52 -07:00
Teknium ff3835a630 fix(gateway): compact the live codex thread instead of no-op mirror rewrites (#73503)
On the codex_app_server runtime the model's real working context is the
app-server's server-side thread: CodexAppServerSession is constructed with
no history and each turn submits only the new user message
(agent/codex_runtime.py), so Hermes' transcript is a mirror that is never
replayed into a thread. Every out-of-turn compression call site (gateway
session hygiene, gateway /compress) built a DETACHED agent whose
_codex_session was None, so the codex route bailed at its "no active codex
thread" guard and returned the transcript unchanged ("compressed 150 ->
150 msgs") — and hygiene's finally-clause then evicted the cached live
agent, destroying the only real context: the next turn spawned an empty
thread while Hermes still mirrored a full history.

Fix, per the documented compression.codex_app_server_auto contract:

* Session hygiene now routes codex_app_server sessions to
  run_codex_hygiene_compaction(): in 'hermes' mode it compacts the LIVE
  cached agent's thread via thread/compact/start (through the existing
  codex route in _compress_context) and KEEPS that agent cached; 'native'
  and 'off' skip cleanly with no eviction and no local fallback. A wedged
  compaction records the persistent failure cooldown; success resets the
  hygiene failure streak.
* Gateway /compress detects the codex_app_server runtime before building
  a temporary compression agent and compacts the live thread with
  force=True instead (a manual compress is an explicit user decision in
  every mode). No live thread -> honest "nothing to compact" reply
  instead of a mirror rewrite plus eviction.
* No mode ever runs the local transcript compressor on this runtime:
  rewriting the mirror cannot shrink the thread, so the #73715-style
  local fallback (including its force=True leak into native/off) is
  deliberately not adopted.

Diagnosis of the mode-gate/no-thread deadlock builds on PR #73715.

Closes #73503

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-08-30 19:52:43 -07:00