AST-driven, body-identical move of 359 GatewayRunner methods into cohesive
mixin modules (gateway/run_{voice,adapters,topics,turn,shutdown,busy,
config_loaders,startup,watchers,notifications,inbound,goals,agent_cache}.py)
plus TurnRunner -> gateway/run_turn_runner.py. run.py-internal symbols are
imported lazily inside method bodies so patch('gateway.run.X') keeps
intercepting; neutral deps are top-level; logger name stays 'gateway.run'.
_UNSET moved to leaf gateway/run_common.py (def-time default-arg sentinel).
Whole-module inspect.getsource(gateway_run) AST-walker tests repointed to
the module that now holds the walked code.
andrexibiza's review on #101118 pointed out that the timeout branch
still ran the SessionDB close/checkpoint even when _shutdown_executor()
reported a live worker -- the exact sequence that produces the
wrong-page-number corruption in #101093. The close block now only runs
when _exec_live == 0; a surviving worker skips it entirely and leaves
the handle open for SQLite to recover from its own WAL on next open.
Adds test_stuck_worker_skips_the_session_db_close to prove the converse
of the existing ordering test: a worker that outlives the budget must
never be raced by close().
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019JujDvCo2vfpEiizidAS2U
`_shutdown_executor()` ran *after* the SessionDB close block in `_stop_impl`,
and it never waited. That left two ways for blocking DB work to outlive
`SessionDB.close()`:
(a) `_executor_closing` was still False during the close, so a coroutine
reaching `_run_in_executor_with_context` minted a brand-new pool and ran
more blocking DB work against handles that had just been closed;
(b) `cancel_futures` only drops work that has not started, and cancelling
`self._background_tasks` does not stop the worker thread behind a
`run_in_executor` future that is already running.
`SessionDB.close()` checkpoints the WAL and lets SQLite unlink the sidecar. A
write that lands after it silently reopens the handle (#94736) and mints a
fresh WAL generation behind that checkpoint, so teardown checkpoints the same
file a second time from a connection the shutdown log never accounts for --
the close-time page-write damage in #101093 and the split WAL generation in
#101064.
The quiesce now runs before the close and waits for the running workers. The
wait is bounded by `_EXECUTOR_QUIESCE_TIMEOUT` (2s) and clamped to what is left
of the shutdown watchdog leash minus a second for the close itself, so a stuck
worker can never cost the post-close cleanup window (#82161). Workers still
alive after the budget are logged as a warning instead of being waited on.
`_shutdown_executor()` keeps its no-argument fire-and-forget contract and now
returns the number of workers still running.
Refs #101093
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DGpGPnvz5FFFH999i59Xeb
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:
* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
agent.tools from live probes with no predecessor to preserve. Persist
the session's resolved tool-name order in a new `sessions.tool_names`
JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
alongside the system prompt and re-pinned on every published refresh
(so /reload-mcp and compaction naturally reset it; /new mints a new
row). On restore-for-existing-session the fresh definitions are folded
onto the saved order via the SAME `_merge_preserving_prefix` helper —
a probe-flipped tool is carried forward from the registry schema, a
deregistered one dropped, new tools appended at the tail.
* /reload-mcp (CLI, gateway, TUI RPC) now also calls
`reprobe_tool_availability()` — drops the check_fn verdict cache and the
get_tool_definitions memo — so a user can consciously pick up a
credential/daemon that appeared mid-session. Docs updated.
_start_one_profile_adapters skipped only Platform.RELAY as shared
process-level ingress. WhatsApp is the same shape: the bridge is one
authenticated session tied to a single phone number, so a secondary
profile has no credential of its own to bring; constructing an adapter
for it only produced a connect/retry loop that stalled startup for every
profile queued behind it. Treat WhatsApp like Relay -- the active profile
owns the connection and route-stamped source.profile fans inbound turns
out to secondary profiles.
Salvage of #69042 (narrowed by its author to this one behavioral line);
test re-expressed on the current secondary-startup fixtures.
Co-authored-by: sshawn <28279366+lsshawn@users.noreply.github.com>
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.
- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
served profile inside `_profile_runtime_scope`, carried into the
executor via `copy_context()` (same shape as
`_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
caller (e.g. button-confirm callback) did not; shut down / rediscover /
report only that profile's servers; refresh only that profile's cached
agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
`_server_scope_keys` ownership map; leaves the shared MCP loop running
while other profiles' servers are live. Unscoped call keeps the full
historical behavior.
- MCP tools register into the owning profile's registry overlay
(`registry.register(scope=...)`), and `registry.deregister()` gains a
matching `scope=` kwarg. Plugin callers still cannot name another
profile's scope; the plugin-vs-global guard is unchanged for them.
Fixes#95518
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
Multiplex gateway startup only ever calls agent.shell_hooks/
outbound_webhooks register_from_config() once, against the
root/default profile's config, before any profile scope exists.
_start_one_profile_adapters() discovers Python plugins per profile
but never registered that profile's own declarative `hooks:` block,
so a secondary profile's shell hooks (e.g. a deny-writes gate) and
outbound webhooks silently never fire.
Load and register each profile's own config inside its
_profile_runtime_scope, and key the module-level idempotence sets in
shell_hooks.py/outbound_webhooks.py by resolved Hermes home so two
profiles configuring an identical hook/webhook both register on
their own plugin manager instead of the second being dropped as a
duplicate of the first.
Fixes#92672
Voice state was keyed `<platform>:<chat_id>` with no profile namespace,
so two bots in one Discord channel shared one /voice mode; every
`_voice_input_callback` was the bare `_handle_voice_channel_input`, which
(like `_handle_voice_timeout_cleanup` and the /voice slash handler) always
picked `self.adapters[DISCORD]` — a secondary profile's voice transcripts
were dispatched through the default profile's bot.
- `_voice_key(platform, chat_id, profile=None)`: named profiles get a
`<profile>:` prefix; default keeps the legacy shape (persisted state valid).
- `_voice_key_for_source` keys by the transport-OWNING profile
(`_adapter_profile_for_source`), matching what `_sync_voice_mode_state_to_adapter`
now restores per adapter via `_owner_profile`.
- `_bind_voice_input_callback` binds the capturing adapter into the
transcript handler (functools.partial); used at primary connect,
primary reconnect, /voice channel join, and `_configure_profile_adapter`.
- `_handle_voice_timeout_cleanup` takes the adapter it was bound to.
- /voice, join, leave and `_should_send_voice_reply` resolve the adapter via
`_adapter_for_source` (fail-closed) instead of `self.adapters[platform]`.
- #84872: `_start_one_profile_adapters` now calls
`_sync_voice_mode_state_to_adapter` on secondary INITIAL connect, as the
primary path and both reconnect paths already did.
Co-authored-by: davidxyuan <124700534+davidxyuan@users.noreply.github.com>
_run_secondary_profile_reconnect now pre-hydrates external secret sources
in a worker thread (PR #99519); _start_one_profile_adapters entered
_profile_runtime_scope on the event loop three times per profile, each
running the same synchronous network-bound hydration under
_SECRET_SOURCE_CACHE_LOCK. Hydrate once via asyncio.to_thread and enter
every scope in that method with hydrate_secrets=False.
The reconnect test is parametrized over both entry points; the startup
reconnect handoff waits are deadline-based since the runner now hops to a
worker thread before publishing the replacement adapter.
Co-authored-by: GoBeromsu <37897508+GoBeromsu@users.noreply.github.com>
setup_logging(mode="gateway") binds agent.log / errors.log / gateway.log
to the launch home, so under multiplex_profiles every secondary profile's
records (emitted inside _profile_runtime_scope) landed in the DEFAULT
profile's files. Main already has the routing primitives from #99440
(record.hermes_home factory + _ProfileRoutingFileHandler +
enable_profile_log_routing) but only the Desktop cron ticker used them.
Enable them at gateway startup for the served profile set; single-profile
gateways are untouched (routing is a no-op below two homes).
Supersedes #84954, which introduced a parallel routing mechanism; the
cross-profile isolation test is translated from its suite.
Co-authored-by: Michał Dziwisz <michal@dziwisz.net>
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.
Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.
Salvaged from #70831 (Hudson). First reported in #68332.
Same class as the Feishu _app_id gap (#76793): Teams and WeCom authenticate
with an id/secret pair and store no token attribute, so
_adapter_credential_fingerprint returned None and cloned profiles started
competing adapters against one app. Add both ids to the attr tuple.
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.
Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.
Fixes#89161
Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.
Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.
Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.
- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
it off-thread every TTL window, so every surface on the machine reads
a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
`_authorization_adapter` compared a stamped profile against
`_active_profile_name()`, which reads the per-turn HERMES_HOME override.
Inside a secondary profile's `_profile_runtime_scope` (cron, restored or
hand-built sources without transport provenance) that reported the
secondary itself, so it was handed the DEFAULT bot for egress instead of
the fail-closed None. Capture the launch identity once in `__init__`
(`_primary_profile_name`) and compare against that; the
`_active_profile_name()` fallback remains for partial fixtures.
`gateway/channel_directory.py` resolved `DIRECTORY_PATH` /
`CHANNEL_ALIASES_PATH` at import time, pinning every multiplexed profile's
directory to whichever home imported the module first. Resolve lazily
from the current home; the module attributes stay as explicit overrides
(tests patch them) and default to None.
Extracted from #87240 (topic-table half handled separately via #76487).
Co-authored-by: cherryb16 <166878179+cherryb16@users.noreply.github.com>
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.
- `_make_adapter_auth_check`: for the shared primary adapter under
multiplex, mirror the inbound message path exactly — stamp the
`profile_routes` match so the routed profile's pairing store is
consulted, and authorize under the TRANSPORT home via
`_is_user_authorized_for_source` (same split as
`_make_default_profile_message_handler`, 2afed50863). A rejected route
fails closed like the ingress gate. Retain the receiving adapter as
`_transport_adapter_ref` so config.yaml policy reads stay on it.
Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
`thread_id` as keywords only when set, so legacy 3-positional callbacks
keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
introspection class — fall back to the injected `gateway_runner` and
the adapter's owner profile.
Fixes#86296Fixes#92840
Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
Same class as the handoff watcher (#100014): the secondary/primary message
handlers and the scoped inbound-preprocess hop entered
_profile_runtime_scope synchronously on the loop thread, so a slow
profile .env read or external secret-source hydration stalled every other
adapter's traffic. Route those three async sites through
_async_profile_runtime_scope (asyncio.to_thread load, then the existing
sync scope with prepared_secret_scope=). _format_session_info_scoped is
already called via asyncio.to_thread and stays sync.
The handoff watcher entered _profile_runtime_scope synchronously on the
event loop each tick; hydrate_profile_secret_sources + build_profile_secret_scope
do blocking file/secret-source IO, so a slow profile secret read stalled every
adapter (#100014). Load the secret scope via asyncio.to_thread, then enter
the existing sync scope with prepared_secret_scope=.
Fixes#100014
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
_process_handoff caught a config load failure for a secondary profile,
logged a warning, and kept going with self.config, which is the primary
profile's config. The handoff then went out through the right bot to the
primary's home channel and the row was reported completed. That is the
exact wrong delivery the multi-profile handoff work exists to prevent,
and the same fail closed posture the no-live-adapters branch already
takes.
A load failure now logs an error and raises, which marks the row failed
so the CLI can report and retry it. The default profile path is
untouched, it never reloaded config.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:
* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
environment failures that polling cannot fix — `_acquire_db_flock` and
both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
now defer immediately with the real errno instead of burning the full
120s / holder timeout and then logging a fake "held by another process".
* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
lock) stayed `_fts_stale` — LIKE-only search — until the process
reopened state.db. Short-lived CLIs reopen every run; the gateway opens
once and stays up for days, so the deferral was effectively permanent
(#100108). The retry runs from the EXISTING gateway housekeeping tick
(`_start_gateway_housekeeping`, 60s) against the shared SessionDB
instances via `hermes_state_registry.live_shared_session_dbs()`:
non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
bounded backoff 60s -> 1h, no new thread, still fails closed on live
holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.
* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
line can be matched to the interpreter that actually linked it
(#100108 point 3).
Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Sibling site of the same loss class. The #72680 shutdown flush only
serialised the adapter slot (_pending_messages); the FIFO tail parked in
SessionState.conversation.queued_events was discarded with the process,
so every follow-up queued behind the head at restart time vanished the
same way the idle-orphan did. flush_overflow_to_file writes one payload
per overflow event in the slot-flush shape (plus seq for arrival order),
so the existing recover_pending_to_db startup replay inserts them with no
new reader. Wired into _stop_impl beside the slot flush.
Follow-up to the salvaged #99912 rescue. The original helper left the
rescued orphan IN the adapter slot while the caller also swapped it in as
the current turn, so the post-turn _dequeue_pending_event ran the same
follow-up a second time (live repro: TURNS=['Sent','C','C','D']). The
helper now pops the oldest orphan and returns it to run as this turn,
stages the NEXT orphan in the slot so the drain continues the chain in
arrival order, and the call site parks the incoming message behind the
chain via _enqueue_fifo (slot when free, overflow otherwise) instead of
always appending to overflow. The rescued event's own source drives the
turn so reply anchors point at the message actually being answered.
Tests: contract updated for the new return type; added the 2-orphan chain
case and the single-orphan-then-new-message slot case (both fail against
the original helper shape).
Review note on #99912: rescued = 1 followed by if rescued: is a constant
conditional — the log block runs unconditionally now that staging is
single-orphan by design.
When a follow-up is demoted to /queue during compression-in-flight,
it lands in SessionState.conversation.queued_events (overflow) with
the slot event in adapter._pending_messages. After the slot's turn
completes, _promote_queued_event should move the overflow head into
the slot for the recursive drain. When that drain never runs — the
#99882 shape: busy window ended through an exit that skipped the
promotion site — the overflow is silently orphaned: never dispatched,
never persisted, never logged. A 170-char Telegram follow-up vanished
without a trace; its re-send also vanished for the same reason.
Fix: _rescue_orphaned_overflow stages one orphan into the empty slot
on the next idle arrival, and the new message is enqueued behind it
so FIFO order (#28503) holds — oldest orphan runs as this turn, the
rest drain in order, the new message last. The helper is best-effort
(slot occupied or no overflow → no-op) and logs at WARNING when it
fires so a future drain regression is visible.
Tests (tests/gateway/test_fifo_overflow_rescue.py, 4 cases on the real
GatewayRunner FIFO):
- moves overflow head to empty slot
- no-op when slot occupied
- no-op when no overflow
- FIFO preserved: orphan-1, orphan-2, new-msg in exact arrival order
Existing queue suites pass unchanged (test_queue_consumption — 5 passed).
Fixes#99882