The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
The salvaged _close_scope_handle replaced direct scope.pop calls that
were bounded by _SCOPE_OP_TIMEOUT with an unbounded run_in_session
callback, regressing the bounded-finalization contract (CI:
tests/agent/test_relay_runtime_bounded_scope_ops.py — end_turn /
close_session / finish_logical_calls hung when the native pop wedged).
Pass timeout=_SCOPE_OP_TIMEOUT so the whole drain+close costs at most
one span and never blocks turn or session completion.
Follow-up to HexLab98's salvaged commits:
- Apply the same bounded worker join before raising InterruptedError at
the two sibling interrupt sites that share the raise-without-join
shape: the non-streaming API poll loop and the Bedrock streaming poll
loop. Both workers run Relay-managed physical attempts, so raising
immediately allowed turn teardown to race a still-open physical scope
exactly as in the streaming path.
- Address the #81601 review finding (egilewski): the pinned nemo-relay
binding's get_scope_stack() returns a native ScopeStack object which
scope.pop rejects with TypeError, so the orphan drain never drained
under the real binding. current_top() now prefers the version-correct
scope.get_handle() accessor and falls back to the old list-unwrap for
fake/legacy shapes. Handle comparisons go through same_handle(),
comparing by uuid, because native ScopeHandle instances do not
implement value equality.
- Add a real-binding regression test that reproduces the orphaned-scope
session close against the pinned native wheel (skips where the native
binding is unavailable), alongside the existing fake-based coverage.
Empty-stream stalls that trip interrupt were raising InterruptedError
before the stream worker closed its physical LLM scope, corrupting the
Relay LIFO stack and cascading into a CLI EIO redraw storm (#81521).
notify_session_compacted closed the old session scope immediately on a
legacy rotating compaction. A compaction can complete while a turn is
still live on the old session; closing then pops the session scope under
the live turn scope, violating the stack's LIFO order — the exact
invariant the rest of the segmentation feature protects.
Now: when the old session has an active turn, set close_pending instead;
that turn's end_turn consumes the flag after its own turn scope pops and
it unregisters from the active-turn table. Sabotage-verified: the new
test fails without the fix.
Continuous gateway sessions keep the Relay session scope open for days;
close-driven export means the session root span and out-of-turn marks
never export until /new or idle-end, and a crash loses the open segment
entirely.
Opt-in segmentation (both defaults OFF => scope lifecycle byte-identical
to today):
gateway.telemetry.session_segments.on_compaction: false
gateway.telemetry.session_segments.max_turns: 0
Rotation closes the current session scope and pushes the next segment
(same session_id attribute, plus hermes.session.segment=N and
segment_reason=compaction|max_turns) ONLY at a turn boundary in
begin_turn — never mid-turn (scope stack is LIFO). Compaction completion
just flags rotate_pending (observer semantics, nothing on the compaction
critical path); legacy rotating compaction closes the orphaned old
session scope so its segment exports. Both native calls ride the
existing bounded scope-op executor: a wedged rotation costs one segment
span, never the agent. Segment bookkeeping advances even on native
failure so a degraded rotation cannot retry every turn.
CI caught the file hanging AFTER '6 passed in 4.32s' until the runner's
300s SIGKILL. Two defects, same class the PR fixes:
1. The executor-refused (interpreter shutdown) fallback ran the native
call UNBOUNDED on the calling thread — a wedged pipeline would block
process exit forever. Now runs on a bounded daemon exit-thread with
the same timeout/abandon semantics as the executor lane.
2. The wedge tests left daemon workers parked on Event.wait() and live
sessions registered on the atexit shutdown hook; exit re-ran the
wedged pops (bounded, 10s each) and the per-file runner timed out.
Autouse teardown now releases every wedge and drains each runtime.
Canonical runner: 4.4s (was 300s file-timeout kill). Bare pytest was a
false green for this class — it exits before atexit replay cost shows.
The NeMo Relay native binding's scope.pop/push are synchronous and
unbounded ('returns after the scope is closed successfully'). When the
native pipeline cannot make progress, the session coordinator's turn and
session finalization block forever inside run_conversation: delegated
children finish their turns but never return, and delegation batches die
on the stall watchdog. Proven live 2026-08-10 on the staging fleet — a
falsification probe (plugin disabled, identical config) completed the
same delegation batch that wedged with the plugin active.
Bound every scope lifecycle operation that gates turn/session completion
(session push, turn push, turn pop, logical-LLM pops, session pop,
subscriber flush) by running the native call on a shared
DaemonThreadPoolExecutor and honoring a 10s result timeout. On breach a
TimeoutError propagates into each call site's existing exception
handling — warn, retain the unclosed-prefix diagnostics, continue — so
the worst case is one lost span, never a blocked agent. timeout=None
preserves byte-identical synchronous behavior for all other callers, and
interpreter-shutdown paths fall back to the synchronous call so the
atexit flush still exports.
Observability must never block the product.