Files
hermes-agent/gateway/run_turn_runner.py
T
Ben Barclay 866332bfb5 fix(relay): authorize send_message targets and surface egress declines (P5) (#99220)
* fix(relay): authorize send_message targets and surface egress declines

P5 of the relay egress-authorization workstream. The relay path
authenticated the SENDER but never authorized the DESTINATION, and the
gateway compounded it from both ends.

(a) send_message could silently name an arbitrary relay target. Its
`target` parameter is free-form ('platform:chat_id'), so a model could
name ANY chat id and the gateway would emit an outbound frame for it.
gateway/relay/egress.py adds an attestation floor: a relay-routed
destination must have a provenance this gateway can show -- the
operator's home channel, the channel directory, or its own gateway
session origins. Anything else is refused HERE, with a visible tool
error naming the target, before a frame is written. Non-relay platforms
and platforms served by a live native adapter in this process are
untouched (same precedence resolve_delivery_transport applies).

(b) Connector declines were swallowed into apparent successes. The
connector's egress floor answers an unauthorized destination with a
DEFINITE failure whose text is deliberately uniform (F-005). Several
relay lanes degrade a *transport drop* by design and were degrading an
*authorization refusal* the same way:

  - _send_media returned None, sending the caller into
    BasePlatformAdapter's text fallback -- a DIFFERENT op re-addressed at
    the very chat the connector had just refused.
  - _send_prompt returned None, so exec-approval / slash-confirm /
    clarify reported "relay prompt op unavailable" (a wrong reason) and
    ran their numbered-text fallbacks into the refused chat.
  - task_card_stop discarded the error entirely.
  - typing / delete / react / thread ops degraded silently at debug.

is_egress_decline() classifies THAT a decline happened (never why --
the uniform text is not parsed for reasons) and requires a definite,
non-ambiguous failure, so a lost-ack retry is still a transport
outcome. Lanes with an error-carrying contract now report the decline
verbatim; cosmetic bool/None lanes still degrade but log it at WARNING.

Advisory progress drops that legitimately degrade are unchanged: the
task_card send lane, the draft ambiguous/except branches, and every
transport-exception path keep their existing fail-open behaviour.

Tests: 21 mutations of the production source, all KILLED.

* fix(relay): authorize the RESOLVED target; declines must not fall back

Review round 1 (independently confirmed by a second reviewer) found three
blockers. Two are fixed here; the third (B-2, Telegram @username) is a policy
decision left open deliberately.

B-1 — THE FIX CAUSED THE OUTAGE IT PREVENTED (tools/send_message_tool.py)

The P5(a) guard ran ABOVE Slack user->DM resolution, so it authorized the
internal pseudo-id `_parse_target_ref` emits (`user_name:ben`, `user:U...`).
Provenances only ever hold RESOLVED conversation ids, so a fully attested DM
was compared as a handle against a set of `D...` ids and refused:

  base  slack:@ben  SENT        head(before)  slack:@ben  REFUSED

Every Slack DM by handle was broken. Moved the guard below resolution; it now
authorizes the destination that is actually sent to, and the refusal names the
resolved id. Position is load-bearing, so it is commented as such and pinned:
reverting the move turns exactly the four new cases red.

B-3 — A DECLINE IS NOT A LANE FAILURE (gateway/run.py)

`_approval_send_outcome` had only sent/failed/ambiguous, so a connector
decline collapsed into `failed` — which is the cue to run the plain-text
fallback into the chat the connector had just refused. The adapter fix in the
previous commit improved the error STRING while user-visible behaviour stayed
identical to base; the commit message overstated it. Fixed properly:

  - new `declined` verdict, recognised via the shared `is_egress_decline`
    contract (not string sniffing at the call site)
  - exec-approval returns without the text fallback
  - slash-confirm suppresses the text reply AND clears the registration, so a
    card that never rendered cannot capture the user's next message

`send_clarify` was already correct (returns early inside the adapter).

MUTATIONS (production source; both directions)
  classifier never returns 'declined'        -> KILLED (4 cases)
  ALL failures classified as 'declined'      -> KILLED (2 cases)
  guard moved back above Slack resolution    -> KILLED (4 cases)
  decline CODE changed (review M05)          -> KILLED
  marker match made case-sensitive (M10)     -> KILLED

M05 was a tautology: the test asserted the imported constant against itself,
so changing the constant could not fail it. The wire contract is now pinned as
a literal, because the connector stamps that exact string and a one-sided
change is a silent cross-repo break.

REGRESSION CHECK: the 12 failures + 1 collection error in this test selection
are PRE-EXISTING cross-test contamination — the identical set fails at
7cf86188ac. Verified by diffing the failing sets: no new failures, 363 -> 374
passed.

NOT FIXED (deliberate): B-2, Telegram `@username`. The Bot API resolves handles
at send time, so there is no id to compare and no canonicalization exists yet.
That is a policy decision, not a code move.

* fix(relay): fail CLOSED on guard faults; classify the structured decline

Third independent review. Two more blockers, both reproduced before fixing.

1. THE GUARD ITSELF FAILED OPEN (tools/send_message_tool.py:158)

`_authorize_relay_target` wrapped BOTH the import and the call in one
`except Exception: return None` — and None means AUTHORIZED at every call site.
So any runtime bug inside the guard silently switched the entire P5(a) boundary
off. Reproduced: with the guard raising, an unattested target sent.

The docstring already stated the correct intent ("must not fail closed on its
own IMPORT error") and the code did something broader. The two failures are not
the same: a missing gateway package means there is no relay egress to
authorize; a fault inside the guard means authorization did not happen. The
import is tolerated, the call is not — a guard that cannot answer refuses.

2. THE STRUCTURED DECLINE WAS THROWN AWAY (gateway/run.py)

The adapter preserves the connector's dict in `SendResult.raw_response`. My
previous commit rebuilt a dict from the error STRING, which loses two
contracts:

  * a decline carrying `code: egress_declined` and NO text renders as
    "relay egress declined" — no marker colon — so it classified as `failed`,
    which is exactly the cue to run the fallback into the refused chat;
  * `ambiguous: True` (lost ack) was flattened into a DEFINITE failure,
    re-sending a card that may already be on the user's screen. That is the
    duplicate-card bug the ambiguous verdict exists to prevent, reintroduced
    by the fix meant to harden the same path.

Both call sites now classify `raw_response` when present, ambiguity first, and
fall back to the wire sentence only for connectors that send no structured
response.

I had fixed the text-marker path and tested only the text-marker path. Worth
naming: the review's probe was a shape my tests never produced.

MUTATIONS (production source)
  guard fault returns None (fail open again)   -> KILLED
  classifier ignores raw_response              -> KILLED (3 cases)
  ambiguous treated as a definite failure      -> KILLED (2 cases)

40 focused tests pass. Regression check vs be321faf27: identical 13-item
failing set (pre-existing cross-test contamination), no new failures.

STILL OPEN: B-2 / finding 3, Telegram `@username`. The reviewer is right that
this is a REGRESSION of an existing contract (#53573 added Bot API username
support), not merely an unspecified input, since relay provenance stores the
numeric chat id. Fixing it means resolving the handle before authorization, or
explicitly revoking the contract. That is a policy decision, not a code move,
and it is Ben's call.

* test(relay): pin M21 and M25, the survivors whose comments called them load-bearing

Round-2 review reported six unpinned survivors from round 1. Two guard real
behaviour and are now covered; the other four are cosmetic-lane warnings and
fail-open branches I am leaving documented rather than pretending to close.

M25 — thread-qualified session ids. `_session_ids` adds BOTH "chat:thread" and
the bare chat, because the connector authorizes the CHAT. Without the split a
gateway whose session origin is `-100999:77` cannot send to `-100999`, the chat
it is demonstrably already talking in. KILLED.

M21 — the generic `relay` plane must union every fronted platform, since a
relay session is filed under its LOGICAL platform. KILLED.

MY FIRST M21 TEST WAS THE DEFECT IT WAS TESTING FOR. I patched `_relay_fronted`
— the very function the mutation empties — so emptying it changed nothing the
test could see, and the mutation SURVIVED against a green test. Rewritten to
drive the real `relay_fronted_platforms()` through its env source
(`GATEWAY_RELAY_PLATFORMS`), which is how production learns it.

That is the same "the test verifies my stand-in" failure I have spent this
workstream removing from the connector harnesses, reproduced here in three
lines of Python. The tell was identical: a mutation that survives a test
written specifically to kill it.

334 tests pass.

NOT PINNED, deliberately: M03 (success-guard on a malformed dict), M24
(empty-target allowance — the one fail-open branch, reachable only when the
bare-platform path already resolved a home channel), M35/M36 (decline WARNINGs
on cosmetic lanes). All four are observability or defence-in-depth rather than
authorization, and the review agrees they are non-blocking.

* fix(relay): defer Telegram @username authorization to the connector (B-2)

Closes the last blocker. Two reviewers independently called this a REGRESSION
of the public-channel username support added in #53573, not an unspecified
input, and they were right: provenance stores RESOLVED numeric chat ids, so
comparing `@channel` against them could only ever refuse.

WHY THE GATEWAY CANNOT ANSWER IT. The guard fires only when there is no live
native adapter — i.e. relay-fronted deployments — and on exactly those the
CONNECTOR holds the bot token, not this process. There is no local way to turn
a handle into the numeric id. Refusing here is not "fail closed", it is "fail
always".

WHY DEFERRING IS SAFE. The destination is still authorized one layer out: the
connector's Telegram egress floor (gg#238, merged 743a7c2) classifies and
refuses unauthorized destinations after ITS resolution — the layer that closed
the reported vulnerability in the first place. Handles go from two guards to
one, the authoritative one, not to zero.

The carve-out is deliberately narrow and its EDGES are pinned, because the
failure mode of an exemption is silent widening:

  telegram `@handle`        -> deferred            (the regression case)
  telegram numeric id       -> still guarded
  matrix `@user:server`     -> still guarded       (telegram-only)
  bare name, no `@`         -> still guarded
  attested handle           -> normal path, attestation still consulted

MUTATIONS
  carve-out widened to all platforms   -> KILLED
  carve-out widened to every target    -> KILLED
  carve-out removed (regression back)  -> KILLED
  carve-out checked BEFORE attestation -> KILLED

THE ORDERING MUTANT SURVIVED MY FIRST TEST. Both orderings return None, so
asserting the verdict could not tell them apart — the test asserted the claim
instead of the mechanism. Rewritten to observe that attestation is actually
consulted. Same defect class as the M21 test earlier in this branch: a
mutation surviving a test written specifically to kill it means the test is
measuring the wrong thing.

341 tests pass.

FOLLOW-UP (option 2, Ben's call, deliberately NOT done here): resolve the
handle before authorizing so BOTH layers apply. That needs a resolution
round-trip through the connector — new wire surface — so it belongs in its own
phase rather than bolted onto this one. Recorded in the code comment at the
carve-out, not just here.

* fix(relay): close two fail-open boundaries; test the code-only decline for real

Both blockers from review, each REPRODUCED before fixing.

1. STRUCTURED DECLINE HAD NO GUARD. Deleting `raw_response=result` from both
   `_send_prompt` return branches left all 34 tests green — a surviving,
   non-equivalent security mutant. The `code` field is the documented
   PREFERRED signal precisely because a connector may send no prose, and a
   caller rebuilding `{"success": False, "error": ...}` cannot see it.

   Cause: every existing case declines with marker TEXT. The evidence for the
   code-only path was a hand-built SimpleNamespace in a different file — a
   stand-in for the adapter, so it verified my fixture instead of production.

   Fixed with a CodeOnlyDecliningConnector driving the real
   `send_exec_approval` -> `_send_prompt`, feeding the REAL SendResult to the
   REAL `_approval_send_outcome`, plus the same shape on the media lane.
       drop raw_response  SURVIVED (34 passed) -> KILLED

2. TWO FAIL-OPEN BOUNDARIES, both "absence" and "fault" sharing a return.

   `_relay_fronted` swallowed EVERY exception and returned an empty set, which
   `relay_routed_platform` reads as "not relay-routed" — skipping the guard.
   Probe, with a positive control in the same run:
       positive_control_denied   = True
       discovery_fault_denied    = False   <- unattested target AUTHORIZED

   `_authorize_relay_target` caught every exception during IMPORT as "no
   gateway package". A module that exists and fails to initialize is a fault,
   not an absence, and returning None there means authorized.

   Now: ImportError alone is absence; anything else raises RelayRouteUnknown
   and `authorize_relay_target` converts it to a REFUSAL STRING (not a raised
   exception — every caller treats the return value as the verdict, so raising
   would trade a fail-open for a crash).

   Kept the converse under test so "fail closed" does not silently become
   "refuse everything in CLI/cron", which is the outage the broad except
   existed to prevent.
       discovery fault -> empty set        KILLED
       RelayRouteUnknown -> authorized     KILLED
       import fault -> authorized          KILLED

397 passed (was 392, +5 new cases), zero failures.

* fix(relay): close all seven review-round-3 blockers

Every finding reproduced before fixing; every fix mutation-checked after.

CONTENT LEAKS (the decline was laundered into a different op, same chat)

#1 A declined DRAFT SEAL replayed as a plain send. On stream-is-the-message
   platforms the turn-final becomes draft(final=True); `_seal_open_draft`
   dropped the structured body, so `_absorb_into_open_draft` read a REFUSAL as
   a lane failure and fell through. Probe, Slack descriptor:
       before: draft(partial) -> draft(final,SECRET) -> send(SECRET)
       after:  draft(partial) -> draft(final,SECRET)
   My first probe of this used a discord descriptor and showed no seal at all —
   the leak is real, my probe was wrong (streams only arm for Slack).

#6 Task-card PROGRESS had the same defect one lane over: a bare failed
   SendResult reads as "card lane unavailable", and TurnRunner then sends the
   task text to the same chat. Both card methods now carry raw_response and
   the caller suppresses the fallback on a decline.

AUTHORIZATION BYPASSES

#2 `except ImportError` was NOT the fix I claimed last round. ImportError also
   covers a broken dependency inside an INSTALLED gateway; review probed
   `ImportError.name = "gateway.relay.dependency"` and got an authorized
   verdict. Now only a name identifying the gateway relay module itself is
   absence. An ImportError with NO name stays absence — refusing on a fault we
   cannot attribute would trade an unidentifiable bug for a real CLI/cron
   outage, and an existing test caught exactly that when I first got it wrong.

#3 `relay_routed_platform` lowercases the requested platform; `_relay_fronted`
   returned configured names verbatim. A platform configured as "Discord"
   missed the membership test, looked native, and skipped the guard:
       'discord' => refused    'Discord' => ALLOWED    'DISCORD' => ALLOWED
   An attestation bypass on a string comparison.

UNDELIVERABLE PROMPTS THAT HUNG

#4 `_clarify_send_disposition` handled `failed` and `ambiguous` but not
   `declined`, so a REFUSED clarify card fell through to wait_for_response and
   blocked until clarify_timeout — indefinitely when configured non-positive.
   A decline is more definitive than a failure, not less.

#5 The exec-approval decline branch returned quietly, which suppressed the text
   fallback (right) but left the CENTRAL approval entry pending (wrong) — the
   dangerous command stayed blocked until the approval timeout. My comment
   claimed the registration was torn down; only RelayAdapter's private map was.
   It now raises `_ExecApprovalDeclined`, which propagates to
   `_await_gateway_decision`'s existing notify-failure path (drops the entry,
   unblocks the tool). A dedicated type, re-raised past the local
   `except Exception` that would otherwise have restored the leak.

#7 THE GAP THAT LET ALL OF THIS SHIP. Both caller-level suppressions were
   unfalsifiable: deleting either branch left 36/38 tests green. The suites
   drove `_approval_send_outcome` and `RelayAdapter` but never the real
   TurnRunner / busy-session callers, so nothing observed whether a text send
   FOLLOWED a decline — which is the whole property.
   tests/gateway/test_decline_fallback_suppression.py drives both real callers
   and records every send. Each decline case is paired with an ordinary-FAILURE
   control, because without one a caller that never falls back would also pass.

MUTATIONS (all on production source, anchors count-checked, restored after)

  #1  seal decline -> plain send                KILLED
  #1b seal drops raw_response                   KILLED
  #2  nested ImportError -> authorized          KILLED
  #3  fronted set not normalized                KILLED
  #4  clarify declined branch removed           KILLED
  #5  approval decline returns not raises       KILLED
  #6  task_card drops raw_response              KILLED
  #7  slash-confirm suppression removed         KILLED

#7's two were the reviewer's SURVIVORS (36/38 passing); both now die.

425 passed, zero failures.

* fix(relay): close the three round-4 blockers

Round 4 confirmed six of seven round-3 fixes and found three more. Each
reproduced before fixing, each mutation-checked after.

1. A NAMELESS ImportError still authorized. Last round I admitted it as
   "absence" to protect the CLI/cron path. That reasoning was WRONG and the
   interpreter says so:

       import gateway.relay.nope  -> ModuleNotFoundError, name="gateway.relay.nope"
       import totally_absent_pkg  -> ModuleNotFoundError, name="totally_absent_pkg"

   Genuine absence is ALWAYS ModuleNotFoundError with `.name` set, so the
   CLI/cron path never produces a bare ImportError and nothing legitimate was
   being protected. A plain or nameless ImportError comes from an import hook
   or a module that failed while initializing — an unattributable FAULT.
   Now: absence is ModuleNotFoundError naming gateway / gateway.relay /
   gateway.relay.egress; everything else refuses. Two existing tests raised a
   bare ImportError to simulate absence and were corrected to the real shape.

2. SESSION ATTESTATION INVENTED IDS. `_session_ids` split every id on the first
   colon to recover "chat" from "chat:thread". Matrix ids contain a colon
   natively, so `!room:server.org` attested a bare `!room` — the guard
   vouching for a destination on its own fabrication. The split now applies
   only to platforms whose ids genuinely carry a `:thread` suffix (allow-list;
   unknown platforms are treated as un-splittable, which can only refuse more).
   Kept a Slack control: dropping the split entirely would refuse legitimate
   thread replies, which is the outage the split exists to prevent.

3. THE TASK-CARD FIX WAS UNFALSIFIABLE — my own round-3 mistake, and the same
   one round 3 caught me making. I added the production branch AND a test, but
   the test stopped at RelayAdapter: it proved `raw_response` is carried and
   never called `TurnRunner._task_card_publish`, which owns the property.
   Deleting the real branch left 30 tests green. Now driven through the real
   caller, with an ordinary-failure control.

   The lesson generalises: proving the DATA reaches the boundary is not proving
   the CALLER acts on it. Every one of these decline fixes has two halves and
   the second half is where the security lives.

Also closed the round-4 non-blocking finding: `gateway/relay/egress.py` has its
OWN import boundary, and the existing test intercepted the earlier import in
tools/send_message_tool.py, so it was never exercised. Mutating that classifier
to treat every ImportError as absence now dies.

MUTATIONS (production source, anchors count-checked, restored after)

  R4-1 nameless ImportError -> authorized        KILLED
  R4-2 session split unconditional               KILLED
  R4-3 task-card caller branch removed           KILLED  (was SURVIVED)
  egress classifier: any ImportError = absence   KILLED

Also probed and found NOT a leak: a refused OPENING draft frame disarms the
stream and the turn-final goes out via `send`. That send is itself guarded and
the connector refuses it too, so no content is delivered — unlike the seal case
(round 3, #1) where the seal was the only check on that path.

452 passed, zero failures.

* fix(relay): recover the thread parent from thread_id, not a colon split

Round 4 blocker 2 was closed with an allow-list of platforms whose ids have no
native colon. Reviewing my own fix while round 5 ran, the allow-list is the
wrong mechanism: it NARROWS a guess instead of removing it, and it still gets
Matrix wrong the moment a Matrix session is thread-qualified
(`!room:server.org:$thr` -> split yields `!room`).

The structured field was there all along. `_session_entry_id` composes the id
as f"{chat_id}:{thread_id}" and the entry still carries `thread_id`
separately, so the parent is knowable EXACTLY: strip the known suffix, or add
nothing. No platform list, no guessing, correct for ids that contain colons.

Mutations:
  back to splitting on the first colon        KILLED
  thread parent never recovered (over-refuse) KILLED

Both directions matter: the first invents attestations, the second refuses
legitimate thread replies.

One existing test (M25) asserted the right PROPERTY with a fixture that omitted
`thread_id` — a shape real entries never have. Fixture corrected, assertions
untouched.

453 passed.

* fix(relay): close the four round-5 blockers

Each reproduced before fixing, each mutation-checked after.

R5-1 A DISABLED NATIVE ADAPTER BYPASSED AUTHORIZATION. `_has_live_native_adapter`
     treated any entry in the adapter map as native; `resolve_delivery_transport`
     ignores a native adapter whose config is disabled and routes over Relay.
     Two independent routing classifiers, disagreeing:
         guard says native: True   delivery routes relay: True
     So the guard skipped authorization for a send that went over the relay.
     The guard now applies the router's enabled-state rule; probed both
     configurations and they agree.

R5-2 THREAD IDS WERE NEVER AUTHORIZED. The parser splits chat_id and thread_id;
     only chat_id reached the guard. On Discord the thread IS the destination —
     `POST /channels/{thread_id}/messages` — so an attested parent channel
     authorized an arbitrary caller-supplied thread. `authorize_relay_target`
     now takes thread_id and requires its own attestation (bare id or the
     `chat:thread` form a session origin produces); both call sites forward it.

R5-3 A DECLINED **INITIAL** DRAFT WAS RETRIED AS A PLAIN SEND. Round 3 fixed the
     declined SEAL; the declined OPEN was a different path. `send_draft`
     returned a bare failure, so the stream consumer read "draft transport
     unusable", disabled drafts and fell through to `_first_send`. Measured
     through the real adapter and real StreamTransportMixin:
         before: ops ['draft', 'send']      after: ops ['draft']
     send_draft now carries raw_response; a decline is terminal for the run and
     the guard sits in `_first_send`, where every fallback path converges.

R5-4 MY ROUND-4 TASK-CARD FIX SUPPRESSED EXACTLY ONE UPDATE. It set
     `native_failed`, which the entry gate already uses for an ordinary broken
     lane, so the next progress event skipped the decline branch and went
     straight to the text fallback:
         after first publish: []      after second: ['send']
     Terminal declines are now a separate `egress_declined` state checked at the
     entry gate. A refusal does not expire after one tick.

MUTATIONS

  R5-1  disabled native counts as native        KILLED
  R5-2  thread_id not authorized                KILLED
  R5-2b tool does not forward thread_id         KILLED (was SURVIVED)
  R5-3  initial-draft decline not terminal      KILLED
  R5-3b _first_send guard removed               KILLED
  R5-4  declined state not persistent           KILLED

R5-2b is the same gap that produced findings 3 and 4 of the last two rounds, a
third time: every test called `authorize_relay_target` directly, so dropping the
argument from the TOOL WRAPPER changed nothing. Testing the callee never proves
the caller uses it — now pinned explicitly.

Each fix ships with an ordinary-failure control, because every one of these
makes the guard refuse MORE, and over-refusal is now the larger risk.

474 passed, zero failures.

* refactor(relay): declare the terminal-decline state where it lives

Both terminal-decline flags were set dynamically. They worked (neither class is
frozen or slotted) but an undeclared attribute hides the state from anyone
reading the class, and this one is security-relevant.

  _TaskCardState.egress_declined  — declared dataclass field
  StreamConsumer._egress_declined — initialised in __init__

Lifetime verified while checking whether a refusal can leak ACROSS turns and
mute a healthy destination: it cannot. _TaskCardState is constructed per
progress-drain (run_turn_runner.py:420) and the consumer's flags per run
(stream_consumer.py:163), so both are fresh each turn.

Also verified the guard's blast radius after adding thread authorization: the
ONLY callers of authorize_relay_target are the two model-facing send_message
call sites. Gateway-internal sends — notably the handoff path, which creates a
thread and immediately posts to it with no session provenance yet — go through
transport.adapter directly and are unaffected. That was the most plausible
over-refusal, and it does not reach this guard.

461 passed.

* fix(relay): close the four round-6 blockers — the edit lane

R6-1 MY OWN R5-1 FIX REINTRODUCED THE BYPASS IT CLOSED. I wrote
     `except Exception: return True` around the config lookup, so a config read
     fault declared the platform native while the ROUTER, reading the real
     config, sends over the relay:
         guard_has_live_native True   guard_verdict None   router relay
     Routing we cannot determine is UNKNOWN. It now raises RelayRouteUnknown,
     which the outer handler must re-raise rather than flatten to False, and
     `authorize_relay_target` turns into a refusal. This is the second time a
     convenience `except` in this function created a bypass; there is now no
     permissive return left in it.

R6-2/3/4 THE NINTH LANE: `edit`. ONE dropped field, THREE leaks.
     `RelayAdapter.edit_message` discarded the connector response, and three
     independent callers read a bare edit failure as "editing is unavailable"
     and re-send the content as a NEW message to the same chat:

       stream edit fallback   ['edit', 'edit', 'send']  the unseen tail
       queued reconciliation  ['edit', 'send']          the WHOLE response
       task-card fallback     ['edit', 'send']          the task text again

     Fixed at the source (edit_message carries raw_response) plus each caller:
     `_on_edit_failure` — the single funnel for stream edit failures — makes a
     decline terminal for the run, `_send_fallback_final` refuses to deliver a
     continuation after one, the queued reconciler returns instead of sending,
     and the task-card fallback sets the same terminal state R5-4 introduced.

     R5-4 fixed the native task-card op and I did not check its sibling
     fallback path. The pattern across rounds 3-6 is consistent: the fix goes
     where the decline is OBSERVED, and the leak lives wherever someone else
     later decides to retry.

MUTATIONS

  R6-1  config fault -> assume native            KILLED
  R6-1b RelayRouteUnknown swallowed as False     KILLED
  R6-2  edit drops raw_response                  KILLED
  R6-2b edit-failure decline not terminal        KILLED
  R6-3  queued reconcile falls back on decline   KILLED
  R6-4  task-card fallback edit decline          KILLED

Each with an ordinary-failure control: a genuinely un-editable message must
still be delivered, and a broken card lane must still reach the user.

481 passed, zero failures.

* fix(relay): add a terminal-decline latch at the adapter choke point

THE STRUCTURAL FIX, not a twelfth local check.

Rounds 3-6 of review found ONE defect in eleven lanes: the connector refuses an
op, and some caller downstream reads that as 'this lane is unavailable' and
retries the same content through a DIFFERENT op against the SAME chat. Media,
prompt, draft-open, draft-seal, native task card, task-card fallback edit,
slash-confirm, exec-approval, clarify, stream edit, queued reconciliation.

Each was closed by adding a check at one more call site. That approach cannot
converge: gateway/ has ~60 outbound call sites, every one of them a place a
future change can reintroduce this, and four consecutive review rounds each
found another. The reviewer's own count of lanes is the argument against the
per-site design.

Every relay frame from every one of those callers passes through
_transport.send_outbound. One latch there covers them all: once the connector
refuses a chat, this adapter stops emitting CONTENT frames for that chat.

Proven to subsume the local checks: with the stream-edit per-site check
DISABLED, the leak probe still reports blocked=true — the frame never reaches
the wire. The local checks stay as defence in depth and for their better error
messages, but they are no longer the only thing standing between a decline and
a re-addressed send.

Scope is deliberately narrow, and each limit is mutation-pinned:
  per CHAT       - a refusal must not mute other conversations
  CONTENT ops    - typing/delete carry nothing; latching them would leave a
                   stuck typing indicator for no security gain
  self-healing   - cleared when the connector accepts that chat again, so a
                   transient policy change does not need a restart

Mutations:
  latch never set                KILLED
  latch never consulted          KILLED
  latch is global, not per-chat  KILLED
  latch never clears             KILLED

485 passed.

* fix(relay): one route source; the latch already covered round 7's lanes

Round 7 reviewed 573e41e294 — one commit BEFORE the terminal-decline latch —
and independently reached the same conclusion I had: 'The per-call-site
approach is structurally wrong. Use one turn-scoped choke point.' That is the
latch in 6dbc004594.

Its four 'still broken' lanes (tool-progress edit, progress-overflow edit,
long-running heartbeat edit, stale streamed-final reconciliation) all share the
shape edit_message->declined->adapter.send(same chat, same content), and NONE
has a local check. Probed all four against the latch:

  tool_progress      ops ['edit']  blocked
  progress_overflow  ops ['edit']  blocked
  heartbeat          ops ['edit']  blocked
  stale_final        ops ['edit']  blocked

That is the argument for the choke point, measured: lanes nobody patched are
safe anyway. Pinned by a parametrized test named for those four lanes.

R7-1 IS A REAL BYPASS THE LATCH DOES NOT COVER, and it is fixed here. The guard
rebuilt routing from GATEWAY_RELAY_PLATFORMS while resolve_delivery_transport
asks the CONNECTED adapter (fronts_platform, from the handshake identity set).
Different snapshots: with env discovery stale or momentarily empty, the guard
said 'native' and the router sent over the relay, skipping authorization.

  before: guard_relay_routed False / delivery relay
  after:  guard_relay_routed True  / delivery relay / unattested target refused

The guard now asks the live adapter first and falls back to config only when
there is no runner (CLI/cron) — pinned in both directions.

R7-5 (non-blocking, and a fair hit): my stream-fallback test asserted
_egress_declined and never drove _send_fallback_final, so removing that early
return SURVIVED. The test now calls the real fallback and asserts the wire is
untouched; the mutation dies.

Mutations:
  R7-1 guard ignores the live adapter    KILLED (was SURVIVED)
  R7-5 fallback early return removed     KILLED (was SURVIVED)
  latch not consulted                    KILLED

491 passed.

* fix(relay): close three holes found by attacking my own latch

Round 8's brief told the reviewer to attack the latch. I did the same in
parallel and found three real holes in it before the review returned.

1. send_for_platform BYPASSED THE LATCH ENTIRELY. It builds and posts its frame
   directly rather than through _outbound — and it is the delivery resolver's
   OWN entry point, so it is the single most important caller.
       before: ops ['edit', 'send']   after: ops ['edit']
   gateway/AGENTS.md states the rule I had just broken: 'Seal-interception
   exists at BOTH egress doors (send() and send_for_platform()); a new egress
   door needs the same two checks.' The latch is a third such check and I had
   wired it to one door.

2. A COSMETIC SUCCESS CLEARED THE LATCH. Clearing on ANY success meant a
   typing indicator — routinely allowed for a chat whose content is refused —
   re-opened the door for the very next send:
       ops ['edit', 'typing', 'send']
   Only a CONTENT op the connector accepted may clear it now.

3. A THREAD INSIDE A REFUSED CHAT WAS NOT COVERED. A thread lives inside its
   parent, so the same content reached the same conversation one level down:
       ops ['edit', 'send']
   The latch key now strips the thread suffix.

Also normalised int/str chat ids (callers pass both; a type mismatch would
silently unlatch).

MUTATIONS
  send_for_platform not latched          KILLED
  cosmetic success clears the latch      KILLED
  thread suffix not stripped             KILLED
  draft-seal retry not latched           SURVIVED — EQUIVALENT, proven:
       is unreachable while latched (a declined edit before the seal
      produces ZERO seal frames, measured). Kept as defence in depth because it
      posts directly, and documented at the site rather than covered by a
      test that could not fail.

One self-inflicted bug on the way: a blanket replace put 1Password CLI brings 1Password to your terminal.

Turn on the 1Password app integration and sign in to get started. Run
'op signin --help' to learn more.

For more help, read our documentation:
https://www.1password.dev/cli

1Password CLI is built using open-source software. View our credits and
licenses:
https://downloads.1password.com/op/credits/stable/credits.html

Usage:  op [command] [flags]

Management Commands:
  account         Manage your locally configured 1Password accounts
  connect         Manage Connect server instances and tokens in your 1Password account
  document        Perform CRUD operations on Document items in your vaults
  events-api      Manage Events API integrations in your 1Password account
  group           Manage the groups in your 1Password account
  item            Perform CRUD operations on the 1Password items in your vaults
  plugin          Manage the shell plugins you use to authenticate third-party CLIs
  service-account Manage service accounts
  user            Manage users within this 1Password account
  vault           Manage permissions and perform CRUD operations on your 1Password vaults

Commands:
  completion      Generate shell completion information
  inject          Inject secrets into a config file
  read            Read a secret reference
  run             Pass secrets as environment variables to a process
  signin          Sign in to a 1Password account
  signout         Sign out of a 1Password account
  update          Check for and download updates.
  whoami          Get information about a signed-in account

Global Flags:
      --account account    Select the account to execute the command by account shorthand, sign-in address, account ID, or user ID. For a list
                           of available accounts, run 'op account list'. Can be set as the OP_ACCOUNT environment variable.
      --cache              Store and use cached information. Caching is enabled by default on UNIX-like systems. Caching is not available on
                           Windows. Options: true, false. Can also be set with the OP_CACHE environment variable. (default true)
      --config directory   Use this configuration directory.
      --debug              Enable debug mode. Can also be enabled by setting the OP_DEBUG environment variable to true.
      --encoding type      Use this character encoding type. Default: UTF-8. Supported: SHIFT_JIS, gbk.
      --format string      Use this output format. Can be 'human-readable' or 'json'. Can be set as the OP_FORMAT environment variable.
                           (default "human-readable")
  -h, --help               Get help for op.
      --iso-timestamps     Format timestamps according to ISO 8601 / RFC 3339. Can be set as the OP_ISO_TIMESTAMPS environment variable.
      --no-color           Print output without color.
      --session token      Authenticate with this session token. 1Password CLI outputs session tokens for successful 'op signin' commands when
                           1Password app integration is not enabled.
  -v, --version            version for op

Run 'op [command] --help' for more information on the command. into
send_for_platform, which has no such variable. Two existing unfurl tests caught
it — NameError at adapter.py:1407.

504 passed.

* fix(relay): Telegram handle exemption + a turn boundary for the latch

Round 8 blockers. Two of its four were already closed by 93750e351a (it
reviewed the commit before it); these two are real and both are mine.

B1 — THE TELEGRAM @HANDLE EXEMPTION COVERED A NATIVE SEND.

_is_unresolved_handle exempts telegram @handles from attestation because
"the connector resolves and authorizes it". That justification is FALSE
whenever the gateway holds its own token: _send_to_platform calls
_send_telegram(pconfig.token, ...) directly and no connector is involved.
So an unattested @handle went out under the gateway's own credential
while the numeric control was correctly refused.

The exemption now requires that no native credential exists. A probe
fault WITHDRAWS the exemption (falls back to the ordinary attestation
check) rather than granting it.

Shipped with the converse control: relay-only config still exempts
@handles, and numeric targets stay guarded in both modes.

B4 — THE LATCH HAD NO BOUNDARY, SO IT WAS AN OUTAGE MECHANISM.

My own regression, and worse than reported. Removing "clear on cosmetic
success" (correctly) removed the ONLY way the latch could ever clear: a
content op can never reach the connector to succeed, because the latch
blocks it locally first. A refusal at 09:00 muted that chat forever.

A new inbound message for a chat is the generation marker — the natural
teardown point. Suppression still holds for the whole turn.

    same_turn_blocked: true     next_turn_delivered: true

MUTATIONS (all killed)
  handle exemption ignores native credential
  native-credential fault GRANTS the exemption
  no turn boundary (latch never clears)
  teardown clears ALL chats not just this one
  teardown ignores the chat

The last two SURVIVED first: I tested _clear_declined_for_turn directly
and never proved _on_inbound calls it — the caller-level gap that has now
produced four blockers on this branch. Added a test driving the real
inbound entry point.

One self-inflicted bug, caught by my own fault test: the probe imported
load_config, which does not exist (it is load_gateway_config), so it
always threw and returned the fault default. The test that pinned fault
behaviour is what exposed it.

510 passed.

* fix(relay): correct latch identity and boundary; one config snapshot

Round 9, four blockers, all reproduced.

B1+B4 — THE TEARDOWN WAS AT THE WRONG PLACE, twice over.

It sat on the adapter's raw _on_inbound, which runs BEFORE profile
routing, the ignored-channel guard, plugin hooks and user authorization.
An unauthorized or dropped event could therefore clear a refusal
belonging to an active turn, and stale content then went out as a
different op. The same placement missed Discord interaction passthrough,
which builds its own MessageEvent and calls handle_message directly, so
slash commands and modal submits stayed muted after an earlier decline.

Both are one mistake: I picked a lane instead of a boundary. Teardown now
runs immediately after _hm_admit_event, the single admission gate every
entry path shares.

  dropped event  -> latch survives, stale send blocked
  admitted event -> latch clears

B2 — THE LATCH KEY SPLIT ON ':', WHICH IS A MISTAKE I ALREADY FIXED ONCE.

_latch_key did str(chat_id).split(":", 1)[0], so !room:tenant-a and
!room:tenant-b both keyed !room: a decline in one Matrix room muted
another, and inbound from one cleared the other's refusal. egress.py
::_session_ids stopped doing exactly this in round 4 and I reintroduced
it three rounds later.

Parent identity is never recoverable from identifier TEXT. Thread
coverage is now structural: _thread_parent looks the relationship up in
the recorded auto-thread map.

B3 — AUTHORIZATION AND DISPATCH USED DIFFERENT CONFIG SNAPSHOTS.

_handle_send retains one pconfig; the guard independently reloaded
config. Across a transition the authorization snapshot could see a
connector-only setup (exemption granted) while dispatch still held the
native token and sent the unattested @handle itself. The guard now takes
native_token from the SAME snapshot dispatch will use. A caller that
omits it does not silently look like "no token".

NB-1/2/3 also closed: real-object snapshot tests, an exception shield
that faces a real exception, and send_follow_up no longer discards the
connector's verdict (that discard is exactly how the edit lane laundered
declines).

MUTATIONS (all killed)
  latch key splits on colon again
  thread parent lookup disabled
  dispatch token ignored by guard
  tool drops the snapshot token
  admission teardown removed
  teardown moved BEFORE admission
  exception shield removed
  follow_up drops raw_response

"admission teardown removed" SURVIVED first: I had tested the helper, not
_handle_message. Added a test driving production _handle_message with
admission stubbed both ways. Fifth caller-level gap on this branch.

One self-inflicted bug caught before commit: I passed pconfig.token in
_handle_react, which has no pconfig — a NameError on every reaction.

516 passed.

* docs(relay): pin the latch's thread coverage limit as a deliberate trade

_thread_parent only sees connector auto-threads, and that map is capped at
256 entries, so a user-created or evicted thread does not inherit its
parent's latch. Documented at the site and asserted by a test, because the
alternative - deriving parents from identifier text - is exactly what muted
unrelated Matrix rooms in round 9.

The primary control is unaffected: authorize_relay_target takes thread_id as
part of the destination and attests it on every send (6 thread tests).

* refactor(relay): one SendResult decline classifier for all 8 gateway lanes

The extraction found a DEFECT, not just repetition.

Eight gateway lanes each hand-rolled the unwrapping of a decline from a
SendResult, and they did not agree. Six checked only raw_response. Two
also checked the error text. A connector that answers with the uniform
decline SENTENCE and no structured code - the documented contract for
older connectors, per _approval_send_outcome - was therefore classified
as an ordinary failure by those six lanes, so each treated a refusal as
"editing unavailable" and retried through another op.

Measured:

    text-only decline    six-site check False    two-site check True
    structured decline   six-site check True     two-site check True

No content leaked, because the adapter latch classifies the transport
dict directly and catches both shapes (verified: text-only decline still
latches C1 and keeps SECRET off the wire). The cost was wrong verdicts
and futile retries, not disclosure.

declined_send(result) in gateway/relay/egress.py now owns this. It checks
raw_response when structured, else the error text, and preserves the
ambiguous exclusion - an ambiguous result is a transport outcome, so it
must never read as a refusal.

run.py keeps its own shape deliberately: that lane has three verdicts
(ambiguous / declined / failed), so it checks ambiguous first and then
delegates the boolean.

MUTATIONS (all killed)
  helper drops the text-only branch
  helper drops the structured branch
  ambiguous no longer excluded
  draft lane decline check removed
  edit-failure lane decline check removed
  prompt verdict lane check removed
  slash-confirm lane check removed
  draft lane goes terminal on ANY failure   (over-refusal direction)

"draft lane decline check removed" SURVIVED first: _send_draft_frame had
no test driving an unsuccessful send_draft at all. Added one, with an
ordinary-failure control so the fix cannot silently become "one flaky
frame mutes the chat". A non-unique anchor also masked the edit-failure
lane on the first pass - the trap my own skill warns about.

This closes the duplication that caused four of nine rounds of blockers:
a new lane now calls one classifier instead of copying three lines.

519 passed.

* fix(relay): latch identity, new-turn boundary, seal arming, ambiguity

Round 10, four blockers, each reproduced before fixing. Two are my own
regressions from the previous two rounds.

B1 - ADMISSION IS NOT A NEW-TURN BOUNDARY.

Round 9 moved teardown to just after _hm_admit_event. That is only an
ADMISSION gate: an authorized message can be steered into a running
session, answer a pending prompt, run a busy slash command, or be refused
by the pause/drain gates - all without starting a turn. Each of those
cleared the ACTIVE turn's refusal, and a later fallback from that turn
reached the wire (probe: latch emptied, wire ops ['edit', 'send']).

Teardown now runs after _claim_active_session_slot, the first point the
runner OWNS a new turn. The new test drives production _handle_message
through all four non-turn lanes plus the real new-turn path.

B2 - LATCH IDENTITY OMITTED THE LOGICAL PLATFORM.

One relay adapter fronts several platforms, so native ids collide. A
Discord refusal for chat 42 was cleared by clear_egress_latch("telegram",
"42") - the method took a platform and ignored it - and the Discord
fallback then reached the connector. Keyed by normalized platform plus
exact chat id; thread-parent expansion keeps the platform component.

B3 - THE DIRECT DRAFT-SEAL PATH DID NOT ARM THE LATCH.

_seal_open_draft posts through _attempt directly rather than _outbound,
so a definite decline logged and returned but never latched. The
immediate plain-send fallback was suppressed by the caller's own check;
later same-turn sends were not (wire ['draft', 'draft', 'send'], the
third frame carrying refused content).

B4 - MY OWN REFACTOR MADE AMBIGUOUS RESULTS TERMINAL.

send_draft's ambiguous projection discarded raw_response, so
declined_send fell through to the error-text branch - and an ambiguous
result whose text carries the decline marker ("... egress declined: ack
lost") read as a DEFINITE refusal and terminated the run. Ambiguous means
the frame may well have been delivered: a transport outcome, never an
authorization one.

Fixed on both layers: the projection carries the body (and the seal's
ambiguous return is now explicit too), and declined_send's text-only
branch - which cannot see the ambiguous flag - treats ack-lost text as
transport ambiguity. Audited every SendResult projection in adapter.py
for the same shape.

MUTATIONS (all killed)
  latch key drops the platform
  clear_egress_latch ignores platform
  draft seal does not arm the latch
  ambiguous projection drops raw body
  declined_send infers decline from ack-lost text
  teardown back at admission

523 passed.

* refactor(relay): split the terminal-decline latch out of the guard PR

The latch moves to feat/p5-egress-decline-latch (pushed at 3cf45736d7,
which retains the full history) for redesign. This PR keeps the
authorization guard and the per-site decline checks.

WHY. Across eleven review rounds the two halves behaved very differently.
The guard is a PURE FUNCTION of the destination - its blockers were all
"you asked the wrong question" (case sensitivity, nested ImportError,
missing thread_id, config snapshot skew), each a one-line correction that
then stayed fixed. Rounds 7-10 found nothing new in it.

The latch is MUTABLE STATE WITH A LIFETIME living on RelayAdapter - an
object registered once per process that holds the WebSocket and has no
concept of a turn. Nine of its blockers reduce to three questions the
adapter cannot answer: when does it end, who arms it, what is it keyed
on. Every answer so far has been a proxy (a successful op, an inbound
message, an admitted event, a claimed session slot) and every proxy was
wrong in a lane found later.

The per-site checks hold identical information on `st` - a PER-TURN
object - and have produced zero blockers, because the state dies with the
turn and nobody has to decide when it ends.

The no-relaunder property does NOT depend on the latch. Measured on the
real consumer path with the latch absent: a declined draft frame sets
_egress_declined and puts nothing on the wire.

Removal verified structurally rather than by eye: an AST diff of every
symbol between HEAD and this tree reports only latch symbols gone,
nothing added. That check caught two over-deletions my strip made -
_on_inbound (consumed by a "next def" boundary) and _SEEN_INBOUND_MAX
(a class constant inside the removed span). Both restored; 19 failures
went to 0.

ALSO: RESTORED A TEST I WRONGLY REPORTED AS PASSING.

test_tool_guard_forwards_thread_id never made it into the repo - `git log
-S` finds it in no commit - though round 5 recorded its mutant as killed.
Dropping thread_id from the guard call therefore survived the entire
tests/tools suite (146 passed). Written properly this time, driving the
real _handle_send far enough to reach the guard. It now KILLS that
mutant.

MUTATIONS on this tree
  guard fault authorizes instead of refusing      KILLED
  thread_id dropped from the guard call           KILLED  (was SURVIVED)
  handle exemption ignores native credential      KILLED
  draft lane decline check removed                KILLED
  prompt verdict lane check removed               KILLED
  slash-confirm lane check removed                KILLED

503 passed.

* test(relay): close the phantom-coverage gaps the guard audit found

The thread_id test that was reported as killing a round-5 mutant turned
out never to have been committed. That is a reason to distrust the other
claimed kills, so I re-ran every guard mutation against the COMMITTED
tree instead of trusting the earlier reports.

Result: 9 of 11 killed, and the two "SKIPPED" ones had non-unique
anchors hiding SIX separate sites. Mutating those individually found
three real survivors.

CASE NORMALISATION (round 3, finding 3) WAS HALF-COVERED.

test_relay_fronted_matching_is_case_insensitive varies the CONFIGURED
name but always requests lowercase "discord", so it pins _relay_fronted's
normalisation and nothing else. The REQUESTED name's `.lower()` was
covered by nothing at all. Probe with it removed:

    relay_routed("Discord") -> False
    authorize("Discord", unattested) -> AUTHORIZED

which is exactly the bypass round 3 reported, alive again and untested.

Two further sites were untested in the OVER-REFUSAL direction: the
attested store is keyed lowercase, so a mixed-case request missed its own
attested set and refused legitimate traffic. attested_relay_targets' own
normalisation was invisible to every existing test because they all
monkeypatch that function away; it is now asserted against the real
function with only its leaf sources stubbed.

Three tests added. All six case sites now die when mutated.

I also re-did the three fail-closed RelayRouteUnknown mutations properly.
The first pass swapped whole lines and produced IndentationErrors, so
"KILLED" there proved nothing but a syntax error. Neutralising each raise
at correct indentation: all three genuinely KILLED.

FINAL AUDIT ON THIS TREE — 17 mutations, zero survivors
  guard: thread_id dropped at the call site
  guard: react path unguarded
  guard: handle exemption ignores native credential
  guard: 3x fail-closed raise neutralised
  guard: 6x case-normalisation site
  classifier: ambiguous treated as a decline
  classifier: text-only decline branch removed
  lane: draft / stream-edit / prompt / slash-confirm checks removed

511 passed.

* test(relay): make the stream-edit test fail for the right reason

Review of 45835a282d raised one blocking issue and three non-blocking
ones. All four are addressed; none was a production defect.

BLOCKING — the stream-edit test failed on the double, not on a leak.

test_declined_stream_edit_does_not_send_the_unseen_tail implemented only
the GUARDED path in its consumer double. Removing either guard therefore
raised AttributeError inside the fake before any send could be observed:

  guard 1 removed -> AttributeError: no attribute '_is_flood_error'
  guard 2 removed -> AttributeError: no attribute '_clean_for_display'

Red, but for the wrong reason — the test could not have caught the leak
it is named for. My own docstring claimed it drove the fallback and
checked the wire; it did neither.

The double now implements everything the UNGUARDED path reaches
(_is_flood_error, _flood_strikes, _current_edit_interval, _last_edit_time,
_notify_new_message, _try_strip_cursor, _clean_for_display,
_fallback_prefix, _metadata_for_send). Both mutations now fail on real
assertions:

  guard 1 removed -> assert consumer._egress_declined is True
  guard 2 removed -> AssertionError: the unseen tail reached the wire:
                     ['send']

NON-BLOCKING 1 — a docstring claimed more than the test exercises.

test_requested_platform_name_is_also_normalised described a mixed-case
send_message(target="Discord:999") bypass. That entry point cannot reach
it: _resolve_tool_target lowercases the platform at
tools/send_message_tool.py:47 before the guard runs. The test still pins
a real contract — the helpers must not assume a lowercased argument, for
the gateway lanes and any future non-normalising caller — so the claim is
narrowed to that rather than the test removed.

NON-BLOCKING 2 — the module docstring said "every lane drives the REAL
RelayAdapter". The stream tests drive mixin doubles by design, because
the behaviour under test belongs to the adapter's CALLER. Docstring now
distinguishes the two kinds.

NON-BLOCKING 3 — latch-deletion residue in gateway/relay/adapter.py:418:

      return None
      return latched if surface_declines else None

The second line was unreachable and referenced a name deleted with the
latch. Removed, along with the 20-line comment block describing the latch
as "the structural fix" — that mechanism now lives on
feat/p5-egress-decline-latch, not here.

The reviewer independently confirmed the large deletion: an AST census
between 3cf45736d7 and f57a2298fa reports only latch symbols removed and
nothing added.

511 passed.

* docs(relay): correct three claims that outran the code

Review of 41ce3cc765 found no new production defect but three overstated
claims, one of them in my own commit message.

1. THE LATCH COMMENTARY WAS STILL THERE. My previous commit message said
   it removed "the 20-line comment block describing the latch as the
   structural fix". It removed only the unreachable statement. Twenty
   lines at adapter.py:361-380 still described a per-chat latch, a choke
   point and its scope rules - none of which exist on this branch. In a
   refusal-sensitive module that reads as coverage this branch does not
   have. Now removed for real.

   This is the same defect class as the tests: a claim that outran what
   the code does. I made it while fixing that class.

2. THE STREAM-TEST DOCSTRING OVERSTATED BOTH MUTANTS. It said the
   mutation "now fails on the assertion that a send reached the wire" -
   true of one guard, not both. Verified separately:

     remove the _on_edit_failure check  -> dies on _egress_declined,
                                           never reaches the fallback
     remove the fallback early return   -> dies on the wire: ['send']

   Both are valid behavioural failures, which is what the blocker asked
   for; they are different observables and the docstring now says so.

3. Duplicate `from types import SimpleNamespace` from an earlier scripted
   insert; imports reordered.

112 tests pass in the four focused files.

* fix(relay): close two authorization defects found in review

Both were reproduced before fixing and both mutants are pinned.

1. A LIVE relay adapter whose fronts_platform() raised degraded into the
   config fallback. `_live_relay_fronted` returned None for every failure,
   and None means "no live adapter, use the config snapshot" — so a faulting
   adapter plus an empty/stale snapshot made the guard conclude "not
   relay-routed" and authorize an unattested destination, while
   resolve_delivery_transport asks that same adapter and still routes over
   the relay. Measured: relay_routed=False, verdict None for chat 999.

   Absence and fault now have separate return values: None only when there
   is no runner or no relay adapter; a live adapter that cannot answer
   raises RelayRouteUnknown. This is the third instance of this bug class in
   this file, and the first two were also mine.

2. An attested chat whose id equalled the requested THREAD id vouched for
   that thread. The `thread in attested` arm proved nothing about parentage.
   Measured: attested {"-100A", "7"} authorized (-100A, thread 7).

   Only the bound `parent:thread` form is accepted now. Nothing legitimate
   needed the bare arm — _session_entry_id records a threaded origin as
   f"{chat_id}:{thread_id}", and a thread addressed as its own channel
   arrives as chat_id and passes the parent check.

The existing test blessed the bare form via parametrize, so it PINNED the
defect. Corrected, plus negative controls for the sibling-chat and
other-parent cases and a positive control proving genuine absence still
takes the config path (otherwise fix 1 would break native-only deploys).

Merged origin/main (was 22 behind). 428 passed via scripts/run_tests.sh;
full 10-row mutation ledger re-killed on the merged tree, none dying on an
exception rather than an assertion.

* fix(relay): only a missing adapter is absence; everything else is a fault

Reviewer BLOCKER, reproduced before fixing. Two more paths where a PRESENT
relay adapter still degraded into the config snapshot:

1. `fronts_platform` may be a property or descriptor, so the ATTRIBUTE
   LOOKUP can raise — and the lookup sat inside the absence handler. Probed
   with a raising property plus an empty snapshot: live=None, routed=False,
   verdict=None, i.e. an unattested target authorized. The previous test made
   an already-retrieved METHOD raise, so it could not reach this.

2. A present adapter with no usable `fronts_platform` returned None for the
   same reason. An adapter that cannot say what it fronts is broken, not
   absent, so it now raises too.

Also found by my own spot-check while the review ran: the nested imports of
`gateway.config` / `gateway.run` inside the live probe shared the broad
handler, so a broken installation degraded to the snapshot as well. Probed
with a healthy-adapter positive control in the same run — healthy refused
the unattested target, faulted authorized it. `_relay_fronted` one function
below already drew this exact distinction for its own import.

The boundary is now: `relay is None` is the ONLY absence. Everything about a
present adapter — attribute access, callability, the call itself, and the
imports needed to reach it — is a fault and raises RelayRouteUnknown.

This is the fourth variant of absence-vs-fault in this file and all four
were mine. The lesson is in the code as a comment rather than in a commit
message nobody re-reads.

Four controls keep genuine absence benign: no runner, no relay adapter in
the runner, a real ModuleNotFoundError naming the gateway package, and the
configured-attested-target-still-sends case.

434 passed via scripts/run_tests.sh; 9-row mutation ledger re-killed
including both new guards, none dying on an exception.

* fix(relay): invert the live probe to fail closed by default

Reviewer BLOCKER round 2, reproduced: reading the adapter registry can also
raise. A runner whose `adapters.get()` raised gave relay_present=True,
live=None, routed=False, verdict=None — unattested discord:999 authorized.

That was the FIFTH boundary in one function with the same defect: the call,
the attribute lookup, a non-callable attribute, the nested imports, and now
the registry lookup. Each round I patched the reported boundary and the
defect moved one statement up. The cause was the shape, not the statements:
the function asked "did something go wrong?" and answered None, and None
MEANS "no live adapter, use the config snapshot" — so every statement was a
new chance to fail open, and every new statement would have been too.

Inverted rather than patched a sixth time. Each `return None` now sits
behind an explicit narrow check that cannot itself be the fault (no runner,
no adapters, no relay key, gateway package genuinely absent), and one outer
handler turns anything else into RelayRouteUnknown. A statement added inside
this function is now fail-CLOSED by default.

Verified all six fault shapes raise (call, attribute, missing method,
registry .get, .adapters property, runner ref) and all five absence shapes
stay benign, plus a liveness control where the config snapshot disagrees
with a healthy adapter and the adapter still wins.

Four new tests, including the two absence controls that keep native-only and
CLI deployments working. 438 passed via scripts/run_tests.sh. Mutation
ledger: 8 killed. One survivor recorded as a proven equivalent mutant —
widening `if not registry` to `or {}` is behaviourally identical because
`{}.get()` returns None, i.e. the same absence; it is a readability guard.
2026-09-08 17:51:35 +10:00

1787 lines
101 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Per-turn callback runner (progress/status/voice/run_sync) for the gateway agent turn.
``TurnRunner`` owns the per-turn callbacks ``GatewayRunner._run_agent_inner`` binds. ``gateway.run``
internals are imported lazily inside method bodies (import cycle), so ``patch("gateway.run.X")``
keeps intercepting them at call time.
"""
from __future__ import annotations
import asyncio
import dataclasses
import json
import logging
import queue
import re
import threading
import time
from contextlib import suppress
from datetime import datetime
from typing import TYPE_CHECKING, Any, Dict, List, Optional
from agent.interrupt_compat import _accepts_keyword
from agent.replay_cleanup import strip_stale_dangerous_confirmations
from gateway.config import Platform
from gateway.media_repair import repair_explicit_computer_use_media_paths
from gateway.platforms.base import BasePlatformAdapter
from gateway.turn_context import TurnContext
from hermes_cli.config import cfg_get
from utils import is_truthy_value
if TYPE_CHECKING: # string annotations only; never imported at runtime (cycle)
from gateway.run import GatewayRunner # noqa: F401
# Log-record parity with the origin module.
logger = logging.getLogger("gateway.run")
class _ExecApprovalDeclined(RuntimeError):
"""The connector refused the approval card's destination.
Raised (not returned) so it propagates out of `_approval_notify_sync` to
`_await_gateway_decision`, whose notify-failure path drops the central
approval queue entry and unblocks the waiting tool. A plain return
suppressed the text fallback but left that entry pending.
"""
class TurnRunner:
"""Per-turn collaborator carrying ``GatewayRunner._run_agent_inner``'s tool-progress callbacks."""
def __init__(self, runner: "GatewayRunner", ctx: TurnContext) -> None:
self._runner = runner
self._ctx = ctx
# ── shared thread→loop plumbing ─────────────────────────────────────────────────────────
def _schedule(self, coro, log_message: str, loop=None):
"""Hop a coroutine from the agent's sync worker thread onto the gateway loop."""
from gateway.run import safe_schedule_threadsafe
return safe_schedule_threadsafe(
coro, self._ctx._loop_for_step if loop is None else loop, logger=logger, log_message=log_message,
)
def _agent_interrupted(self) -> bool:
"""True once the user sent `stop` (agent_holder[0] is the shared agent handle)."""
try:
agent = self._ctx.agent_holder[0] if self._ctx.agent_holder else None
return bool(agent is not None and getattr(agent, "is_interrupted", False))
except Exception:
return False
def _stream_consumer(self):
holder = self._ctx.stream_consumer_holder
return holder[0] if holder else None
def _drain_progress_queue(self) -> None:
q = self._ctx.progress_queue
with suppress(Exception):
while not q.empty():
q.get_nowait()
def _track_progress_result(self, result) -> None:
"""Remember a delivered progress/status message id for end-of-turn cleanup."""
ctx = self._ctx
if ctx._cleanup_progress and getattr(result, "success", False) and getattr(result, "message_id", None):
ctx._cleanup_msg_ids.append(str(result.message_id))
def _track_future_cleanup_id(self, fut) -> None:
try:
res = fut.result()
except Exception:
return
self._track_progress_result(res)
# ── progress_callback (agent thread → progress queue) ───────────────────────────────────
def progress_callback(self, event_type: str, tool_name: str = None, preview: str = None, args: dict = None, **kwargs):
"""Callback invoked by agent on tool lifecycle events."""
ctx = self._ctx
# Failed subagent → one clean user-facing notice, handled FIRST, before every progress-queue
# gate: platforms with tool_progress off must still hear about a dead delegation.
if event_type == "subagent.complete":
self._progress_subagent_notice(preview, kwargs)
return
self._progress_live_status(event_type, tool_name, args)
# "log" mode: append tool.started lines to the log queue, silent in chat. Handled before
# the progress_queue guard because log mode runs without a chat progress queue.
if ctx.log_queue is not None and event_type == "tool.started" and tool_name and tool_name != "_thinking":
ts = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
preview_str = f' "{preview}"' if preview else ""
ctx.log_queue.put(f"{ts} {tool_name}:{preview_str}".rstrip())
if not ctx.progress_queue or not ctx._run_still_current():
return
if event_type == "tool.completed" and not ctx.long_tool_hint_fired[0]:
self._progress_onboarding_hint(kwargs)
return
# "_thinking" is assistant scratch text between tool calls, never ordinary tool progress:
# only relayed when the platform explicitly opted into thinking_progress.
if event_type == "_thinking" or tool_name == "_thinking":
thinking_text = (preview if tool_name == "_thinking" else tool_name) if ctx._thinking_enabled else None
if thinking_text:
ctx.progress_queue.put(f"💬 {thinking_text}")
return
# Native task cards consume the ID-bearing tool_start/tool_complete callbacks instead;
# name-correlated text events would duplicate cards and mispair concurrent same-tool calls.
if ctx._native_slack_task_cards and event_type in {"tool.started", "tool.completed"}:
return
# tool_progress off → only _thinking passes (above). Only tool.started renders. clarify:
# send_clarify IS the user-facing rendering (a bubble would duplicate it, and verbose mode
# would dump the raw args JSON right under the prompt). Post-`stop`: N parallel tool calls
# fire N tool.started events before the interrupt check, so a late stop must not render them.
if (
not ctx.tool_progress_enabled
or event_type != "tool.started"
# The adapter's send_clarify IS the user-facing rendering (interactive buttons or the
# numbered-text fallback), so a progress bubble is pure duplication — and in verbose mode it
# dumps the raw tool-call args JSON ({"question": ..., "choices": [...]}) into the chat. Because
# the progress queue drains on a background task, that raw JSON typically lands right underneath
# the rendered prompt (#52374).
or tool_name == "clarify"
or self._agent_interrupted()
):
return
# "new" mode: only report when tool changes
if ctx.progress_mode == "new" and tool_name == ctx.last_tool[0]:
return
ctx.last_tool[0] = tool_name
msg = self._progress_build_message(tool_name, preview, args)
if msg is not None:
self._progress_emit(msg)
def _progress_subagent_notice(self, preview, kwargs: dict) -> None:
"""Only terminal failure statuses render (same notice rail as credit warnings)."""
ctx = self._ctx
status = kwargs.get("status")
try:
from tools.delegate_tool import SUBAGENT_FAILURE_STATUSES, format_subagent_failure_line
if status in SUBAGENT_FAILURE_STATUSES and ctx._run_still_current():
line = format_subagent_failure_line(
kwargs.get("goal"), status, error=kwargs.get("summary") or preview,
duration_seconds=kwargs.get("duration_seconds"),
)
self._schedule(self._runner._deliver_platform_notice(ctx.source, line), "subagent failure notice scheduling error")
except Exception:
logger.debug("subagent failure notice failed", exc_info=True)
def _progress_live_status(self, event_type: str, tool_name, args) -> None:
"""Live status line (Slack assistant status): stash the tool phrase on the adapter; the
_keep_typing refresh renders it. Plain dict write, safe from the sync worker thread."""
ctx = self._ctx
adapter = ctx._live_status_adapter
if adapter is None or ctx._live_status_mode == "off" or tool_name == "_thinking":
return
try:
if event_type == "tool.started" and tool_name and ctx._run_still_current():
from agent.display import build_status_phrase
adapter.set_status_text(ctx.source.chat_id, build_status_phrase(tool_name, args if ctx._live_status_mode == "full" else None))
elif event_type == "tool.completed":
# Between tools the model is genuinely "thinking" again — revert to the static default.
adapter.set_status_text(ctx.source.chat_id, None)
except Exception as err:
logger.debug("live status update failed: %s", err)
def _progress_onboarding_hint(self, kwargs: dict) -> None:
"""First-touch onboarding: the first time a tool exceeds _LONG_TOOL_THRESHOLD_S while
streaming every tool (progress_mode == "all"), append a one-time /verbose hint."""
from gateway.run import _hermes_home, _load_gateway_config
ctx = self._ctx
try:
if (kwargs.get("duration") or 0) >= ctx._LONG_TOOL_THRESHOLD_S and ctx.progress_mode == "all":
from agent.onboarding import TOOL_PROGRESS_FLAG, is_seen, mark_seen, tool_progress_hint_gateway
cfg = _load_gateway_config()
gate_on = is_truthy_value(cfg_get(cfg, "display", "tool_progress_command"), default=False)
if gate_on and not is_seen(cfg, TOOL_PROGRESS_FLAG):
ctx.long_tool_hint_fired[0] = True
ctx.progress_queue.put(tool_progress_hint_gateway())
mark_seen(_hermes_home / "config.yaml", TOOL_PROGRESS_FLAG)
except Exception as err:
logger.debug("tool-progress onboarding hint failed: %s", err)
@staticmethod
def _preview_cap() -> int:
"""tool_preview_length (default 40): the one-line preview budget for "all"/"new" modes."""
from agent.display import get_tool_preview_max_len
pl = get_tool_preview_max_len()
return pl if pl > 0 else 40
def _progress_terminal_blocks(self, adapter, tool_name, args, emoji):
"""(full, short) fenced blocks for a terminal command on markdown platforms, else (None, None).
No language tag: Slack mrkdwn renders it as a literal first code line. Verbose shows the FULL
command; "all"/"new" truncate to one line capped at ``tool_preview_length``. Consecutive
terminal calls drop the repeated header so back-to-back commands render as adjacent blocks.
"""
if not (
getattr(adapter, "supports_code_blocks", False) and tool_name == "terminal" and isinstance(args, dict)
and isinstance(args.get("command"), str) and args["command"].strip()
):
return None, None
cmd_full = args["command"].rstrip()
header = "" if self._ctx.last_was_terminal_block[0] else f"{emoji} {tool_name}\n"
cap = self._preview_cap()
lines = cmd_full.splitlines()
cmd_short = lines[0] if lines else cmd_full
if len(cmd_short) > cap:
cmd_short = cmd_short[:cap - 3] + "..."
elif len(lines) > 1:
cmd_short += " ..."
return f"{header}```\n{cmd_full}\n```", f"{header}```\n{cmd_short}\n```"
def _progress_build_message(self, tool_name, preview, args) -> Optional[str]:
"""Render the progress line. Verbose mode queues directly (no dedup) and returns None."""
ctx = self._ctx
from agent.display import get_tool_emoji
emoji = get_tool_emoji(tool_name, default="⚙️")
try:
adapter = self._runner._adapter_for_source(ctx.source)
except Exception:
adapter = None
code_full, code_short = self._progress_terminal_blocks(adapter, tool_name, args, emoji)
verbose = ctx.progress_mode == "verbose"
code = code_full if verbose else code_short
ctx.last_was_terminal_block[0] = code is not None
if verbose:
if code is None and args:
from agent.display import get_tool_preview_max_len
pl = get_tool_preview_max_len()
args_str = json.dumps(args, ensure_ascii=False, default=str)
# tool_preview_length 0 (default) = no truncation in verbose mode; the user asked
# for full detail and platform message-length limits handle the rest.
if pl > 0 and len(args_str) > pl:
args_str = args_str[:pl - 3] + "..."
code = f"{emoji} {tool_name}({list(args.keys())})\n{args_str}"
elif code is None:
code = f"{emoji} {tool_name}: \"{preview}\"" if preview else f"{emoji} {tool_name}..."
ctx.progress_queue.put(code)
return None
if code is not None:
return code
if not preview:
return f"{emoji} {tool_name}..."
from agent.display import get_tool_verb, prepare_tool_preview, tool_verb_connector, verb_drops_preview
prepared = prepare_tool_preview(tool_name, args, fallback=preview, max_len=self._preview_cap())
preview = adapter.format_tool_preview(prepared) if adapter is not None else prepared.text
# Friendly labels: human-phrased line for built-in tools ("🔍 Searching the web for ...")
# by prefixing the verb onto the computed preview, so the command/url/query is kept.
verb = get_tool_verb(tool_name)
if not verb:
return f"{emoji} {tool_name}: \"{preview}\""
return f"{emoji} {verb}" if verb_drops_preview(tool_name) else f"{emoji} {verb}{tool_verb_connector(tool_name)}{preview}"
def _progress_emit(self, msg: str) -> None:
"""Dedup consecutive identical lines (execute_code boilerplate), then route to the native
stream bubble when the consumer accepts tool progress, else the progress queue."""
ctx = self._ctx
sc = self._stream_consumer()
native = sc is not None and getattr(sc, "accepts_tool_progress", False)
if msg == ctx.last_progress_msg[0]:
ctx.repeat_count[0] += 1
if native:
sc.on_tool_progress(f"{msg} (×{ctx.repeat_count[0] + 1})")
else:
ctx.progress_queue.put(("__dedup__", msg, ctx.repeat_count[0]))
return
ctx.last_progress_msg[0], ctx.repeat_count[0] = msg, 0
if native:
sc.on_tool_progress(msg)
else:
ctx.progress_queue.put(msg)
# ── Slack-native task cards (progress-queue drain) ──────────────────────────────────────
@dataclasses.dataclass
class _TaskCardState:
"""Task-card rail state for ``_send_native_task_card_progress``."""
adapter: Any
tasks: Dict[str, Dict[str, str]] = dataclasses.field(default_factory=dict)
task_order: List[str] = dataclasses.field(default_factory=list)
fallback_msg_id: Optional[str] = None
native_failed: bool = False
# TERMINAL authorization refusal, distinct from native_failed: the
# connector refused this destination, so no later publication in this
# turn may re-deliver the task text through the text fallback. Declared
# rather than set dynamically so the state is visible where it lives.
egress_declined: bool = False
anonymous_seq: int = 0
@staticmethod
def _compact(value: Any, limit: int = 120) -> str:
text = re.sub(r"\s+", " ", str(value or "")).strip()
return text if len(text) <= limit else text[: limit - 3].rstrip() + "..."
def visible_tasks(self) -> List[Dict[str, str]]:
return [self.tasks[task_id] for task_id in self.task_order[-8:]]
def fallback_text(self) -> str:
labels = {"in_progress": "running", "complete": "complete", "error": "error"}
lines = [f"- {t['title']} - {labels.get(t['status'], t['status'])}" for t in self.visible_tasks()]
return "Hermes is working\n" + "\n".join(lines)
def _upsert(self, call_id: str, title: str) -> Dict[str, str]:
if call_id not in self.tasks:
self.task_order.append(call_id)
self.tasks[call_id] = {"id": call_id, "title": self._compact(title), "status": "in_progress"}
return self.tasks[call_id]
def apply_event(self, raw: Any) -> bool:
event_type = raw.get("type") if isinstance(raw, dict) else None
if event_type not in {"tool.started", "tool.completed"}:
return False
call_id = str(raw.get("tool_call_id") or "")
if not call_id:
self.anonymous_seq += 1
call_id = f"anonymous_{self.anonymous_seq}"
tool_name = str(raw.get("tool_name") or "tool")
if event_type == "tool.started":
preview = self._compact(raw.get("preview"), 64)
self._upsert(call_id, f"{tool_name} - {preview}" if preview else tool_name)
return True
# Completion-only events are rare but valid on some runtimes; keep their real ID instead
# of guessing a same-name pending call.
task = self.tasks.get(call_id) or self._upsert(call_id, tool_name)
task["status"] = "error" if raw.get("is_error") else "complete"
return True
async def _task_card_send_or_edit_fallback(self, st) -> None:
ctx = self._ctx
text = st.fallback_text()
from gateway.relay.egress import declined_send
if getattr(st, "egress_declined", False):
return
if st.fallback_msg_id:
result = await st.adapter.edit_message(
chat_id=ctx.source.chat_id, message_id=st.fallback_msg_id, content=text, metadata=ctx._progress_metadata,
)
if getattr(result, "success", False):
return
# P5(b): R5-4 made a declined native CARD terminal but left this
# editable-text fallback: a declined edit fell through to
# _send_progress_text and re-sent the same task text to the refused
# chat. The decline must set the terminal state here too.
if declined_send(result):
logger.warning(
"Task-card fallback edit DECLINED by the connector's egress "
"guard; suppressing progress delivery for the rest of this "
"turn (the destination is not approved)"
)
st.egress_declined = True
return
result = await self._send_progress_text(st, text)
if getattr(result, "success", False) and getattr(result, "message_id", None):
st.fallback_msg_id = str(result.message_id)
async def _task_card_publish(self, st) -> None:
ctx = self._ctx
if not st.tasks:
return
if getattr(st, "egress_declined", False):
# The connector refused this destination earlier in the turn; every
# later publication would re-deliver the same task text there.
return
if not st.native_failed:
result = await st.adapter.send_native_task_card_progress(
chat_id=ctx.source.chat_id, tasks=st.visible_tasks(), title="Hermes is working",
reply_to=ctx._progress_reply_to, metadata=ctx._progress_metadata, fallback_text=st.fallback_text(),
)
if getattr(result, "success", False):
return
# P5(b): an AUTHORIZATION decline is not a broken card lane. The
# fallback below sends the same task text to the same chat, which
# turns a refused card into delivered plain text. Stop the lane
# without re-delivering; the refusal is already logged.
from gateway.relay.egress import declined_send
if declined_send(result):
# TERMINAL, and stored SEPARATELY from native_failed. Reusing
# native_failed suppressed exactly ONE update: the next progress
# event skipped this branch (the lane is already "failed") and
# went straight to the text fallback. A refusal does not expire
# after one tick.
st.egress_declined = True
st.native_failed = True
logger.warning(
"Slack native task-card progress DECLINED by the connector's "
"egress guard — suppressing the text fallback for the rest "
"of this turn (the destination is not approved)"
)
return
st.native_failed = True
logger.warning(
"Slack native task-card progress failed; falling back "
"to an editable text update: %s", getattr(result, "error", "unknown error"),
)
# Once the native rail fails, every later lifecycle event edits the same fallback message.
await self._task_card_send_or_edit_fallback(st)
def _task_card_drain(self, st) -> bool:
changed = False
try:
while True:
changed = st.apply_event(self._ctx.progress_queue.get_nowait()) or changed
except queue.Empty:
pass
except Exception:
logger.debug("Slack native progress queue drain failed", exc_info=True)
return changed
async def _send_native_task_card_progress(self, adapter) -> None:
"""Drain the progress queue into Slack-native plan/task cards; on any native failure, fall
back to an editable in-thread message so progress stays live.
See #29483.
"""
ctx = self._ctx
st = self._TaskCardState(adapter)
try:
while ctx._run_still_current():
try:
raw = ctx.progress_queue.get_nowait()
except queue.Empty:
await asyncio.sleep(0.1)
continue
if not self._agent_interrupted() and st.apply_event(raw):
await self._task_card_publish(st)
except asyncio.CancelledError:
if self._task_card_drain(st) and ctx._run_still_current() and not self._agent_interrupted():
await self._task_card_publish(st)
finally:
if hasattr(adapter, "stop_native_task_card_progress"):
# Best-effort on the turn-cleanup path: an escaping transport exception would skip
# final-delivery logic (cleanup awaits catch only CancelledError).
try:
await adapter.stop_native_task_card_progress(
ctx.source.chat_id, reply_to=ctx._progress_reply_to, metadata=ctx._progress_metadata,
)
except asyncio.CancelledError:
raise
except Exception:
logger.debug("task-card stop failed during turn cleanup", exc_info=True)
# ── editable progress bubbles (progress-queue drain) ────────────────────────────────────
@dataclasses.dataclass
class _ProgressEditState:
"""Mutable editable-bubble state shared by ``send_progress_messages`` and its helpers."""
adapter: Any
progress_lines: list
progress_msg_id: Any
can_edit: bool
_progress_len_fn: Any
_PROGRESS_TEXT_LIMIT: int
_edit_accepts_metadata: bool
def _progress_edit_state(self, adapter) -> "TurnRunner._ProgressEditState":
ctx = self._ctx
len_fn = adapter.message_len_fn if isinstance(adapter, BasePlatformAdapter) else len
try:
raw_limit = int(getattr(adapter, "MAX_MESSAGE_LENGTH", 4000) or 4000)
except Exception:
raw_limit = 4000
# Per-chat resolution (relay adapter fronting N platforms): cap and length unit follow the
# chat's underlying platform; native adapters return their scalar/property unchanged.
if isinstance(adapter, BasePlatformAdapter):
with suppress(Exception):
raw_limit = int(adapter.max_message_length_for_chat(ctx.source.chat_id) or 4000)
len_fn = adapter.message_len_fn_for_chat(ctx.source.chat_id)
return self._ProgressEditState(
adapter=adapter, progress_lines=[], progress_msg_id=None,
# "separate" = one message per tool (pre-v0.9 behavior)
can_edit=ctx.progress_grouping != "separate",
_progress_len_fn=len_fn,
# Leave room for platform quirks / formatting; tiny test adapters keep a usable limit.
_PROGRESS_TEXT_LIMIT=max(1, raw_limit - (64 if raw_limit > 128 else 0)),
# Overflow edits pass metadata (Telegram topic/thread routing) only when edit_message takes it.
_edit_accepts_metadata=bool(ctx._progress_metadata) and _accepts_keyword(adapter.edit_message, "metadata"),
)
async def _edit_progress_message(self, st, message_id: str, content: str):
ctx = self._ctx
kwargs = {"chat_id": ctx.source.chat_id, "message_id": message_id, "content": content}
if getattr(st.adapter, "REQUIRES_EDIT_FINALIZE", False):
kwargs["finalize"] = True
if st._edit_accepts_metadata:
kwargs["metadata"] = ctx._progress_metadata
return await st.adapter.edit_message(**kwargs)
@staticmethod
def _progress_text(lines: list) -> str:
return "\n".join(str(line) for line in lines)
def _split_progress_groups(self, st, lines: list) -> list[list]:
"""Partition progress lines into platform-sized editable bubbles."""
groups: list[list] = []
current: list = []
for line in lines:
candidate = current + [line]
if current and st._progress_len_fn(self._progress_text(candidate)) > st._PROGRESS_TEXT_LIMIT:
groups.append(current)
candidate = [line]
current = candidate
return groups + ([current] if current else [])
async def _send_progress_text(self, st, text: str):
ctx = self._ctx
result = await st.adapter.send(
chat_id=ctx.source.chat_id, content=text, reply_to=ctx._progress_reply_to, metadata=ctx._progress_metadata,
)
self._track_progress_result(result)
return result
async def _roll_progress_overflow_if_needed(self, st) -> bool:
"""Start fresh editable progress bubbles before a bubble exceeds limit.
Returns True when it delivered/split the buffer or a transient edit failure left it
intact for retry — either way the caller skips the normal send/edit path this tick.
"""
if not st.progress_lines or not st.can_edit:
return False
groups = self._split_progress_groups(st, st.progress_lines)
if len(groups) <= 1:
return False
if st.progress_msg_id is not None:
result = await self._edit_progress_message(st, st.progress_msg_id, self._progress_text(groups[0]))
if not result.success:
if getattr(result, "retryable", False):
logger.debug("[%s] Transient overflow edit failure — keeping can_edit=True", st.adapter.name)
return True
st.can_edit = False
# Fall back to the existing non-edit behavior.
return False
groups = groups[1:]
for group in groups:
result = await self._send_progress_text(st, self._progress_text(group))
if result.success and result.message_id:
st.progress_msg_id = result.message_id
# The newest continuation is the only mutable bubble: keep just its lines so later
# edits update it instead of replaying the full transcript into new messages.
st.progress_lines = groups[-1]
return True
@staticmethod
def _is_reset_marker(raw) -> bool:
return isinstance(raw, tuple) and len(raw) >= 1 and raw[0] == "__reset__"
def _reset_progress_bubble(self, st) -> None:
"""Content bubble landed — close the tool-progress bubble so the next tool starts fresh
below it; else tool edits hit the ORIGINAL message above (out of order)."""
st.progress_msg_id, st.progress_lines = None, []
self._ctx.last_progress_msg[0], self._ctx.repeat_count[0] = None, 0
def _progress_absorb(self, st, raw) -> Any:
"""Fold a queue item into the bubble buffer; returns the line to render this tick."""
if isinstance(raw, tuple) and len(raw) == 3 and raw[0] == "__dedup__":
_, base_msg, count = raw
if not st.progress_lines:
return base_msg
st.progress_lines[-1] = f"{base_msg} (×{count + 1})"
return st.progress_lines[-1]
st.progress_lines.append(raw)
return raw
async def _flush_progress_edit(self, st) -> None:
if st.can_edit and st.progress_lines and st.progress_msg_id:
with suppress(Exception):
await self._edit_progress_message(st, st.progress_msg_id, self._progress_text(st.progress_lines))
async def _drain_progress_on_cancel(self, st) -> None:
ctx = self._ctx
with suppress(Exception):
while not ctx.progress_queue.empty():
raw = ctx.progress_queue.get_nowait()
if self._is_reset_marker(raw):
# Content-bubble marker during drain: close the current progress bubble
# and start a fresh one for tool lines that arrived after.
await self._roll_progress_overflow_if_needed(st)
await self._flush_progress_edit(st)
self._reset_progress_bubble(st)
else:
self._progress_absorb(st, raw)
await self._roll_progress_overflow_if_needed(st)
# Final edit with all remaining tools (only if editing works)
if st.can_edit and st.progress_lines and st.progress_msg_id:
await self._roll_progress_overflow_if_needed(st)
await self._flush_progress_edit(st)
async def _progress_restore_typing(self, st) -> None:
ctx = self._ctx
await asyncio.sleep(0.3)
if ctx._run_still_current():
await st.adapter.send_typing(ctx.source.chat_id, metadata=ctx._progress_metadata)
async def _progress_send_or_edit(self, st, msg) -> bool:
"""Deliver this tick's bubble. Returns False on a transient edit failure (retry next tick).
Transient network errors (ConnectError, timeouts) must not disable editing; only permanent
failures (not found, permissions) set can_edit=False. Flood control backs off but keeps editing.
"""
if st.can_edit and st.progress_msg_id is not None:
result = await self._edit_progress_message(st, st.progress_msg_id, "\n".join(st.progress_lines))
if result.success:
return True
if getattr(result, "retryable", False):
logger.debug("[%s] Transient edit failure — keeping can_edit=True", st.adapter.name)
return False
if any(w in (getattr(result, "error", "") or "").lower() for w in ("flood", "retry after")):
logger.info("[%s] Progress edit flood control, backing off", st.adapter.name)
else:
st.can_edit = False
await self._send_progress_text(st, msg)
return True
# First tool: send all accumulated text as a new message; editing unsupported: just this line.
result = await self._send_progress_text(st, "\n".join(st.progress_lines) if st.can_edit else msg)
if result.success and result.message_id:
st.progress_msg_id = result.message_id
return True
async def send_progress_messages(self):
ctx = self._ctx
adapter = self._runner._adapter_for_source(ctx.source) if ctx.progress_queue else None
if not adapter:
return
if ctx._native_slack_task_cards and hasattr(adapter, "send_native_task_card_progress"):
await self._send_native_task_card_progress(adapter)
return
# Skip tool progress for platforms that can't edit messages (e.g. iMessage/BlueBubbles):
# each update would be a separate bubble. getattr, not attribute access: duck-typed
# adapters (test fakes, minimal plugins) may lack edit_message — treated as "can't edit".
adapter_edit = getattr(type(adapter), "edit_message", None)
if adapter_edit is None or adapter_edit is BasePlatformAdapter.edit_message:
self._drain_progress_queue()
return
st = self._progress_edit_state(adapter)
last_edit_ts = 0.0
EDIT_INTERVAL = 1.5 # Minimum seconds between edits (Telegram flood control)
while True:
try:
if not ctx._run_still_current():
self._drain_progress_queue()
return
raw = ctx.progress_queue.get_nowait()
# Drain silently when interrupted: events queued in the window between tool parse
# and interrupt processing should not render as bubbles.
if self._agent_interrupted():
await asyncio.sleep(0)
continue
if self._is_reset_marker(raw):
self._reset_progress_bubble(st)
continue
msg = self._progress_absorb(st, raw)
if not await self._roll_progress_overflow_if_needed(st):
# Throttle edits: batch rapid tool updates into fewer API calls (grammY pattern:
# proactively rate-limit rather than react to 429s). Loop back to drain further
# queued messages before sending a single batched edit.
remaining = EDIT_INTERVAL - (time.monotonic() - last_edit_ts)
if remaining > 0:
await asyncio.sleep(remaining)
continue
if not ctx._run_still_current():
return
if not await self._progress_send_or_edit(st, msg):
continue
last_edit_ts = time.monotonic()
await self._progress_restore_typing(st)
except queue.Empty:
await asyncio.sleep(0.3)
except asyncio.CancelledError:
await self._drain_progress_on_cancel(st)
return
except Exception as e:
logger.error("Progress message error: %s", e)
await asyncio.sleep(1)
# ── ID-bearing lifecycle callbacks (agent thread) ───────────────────────────────────────
def voice_ack_callback(self, call_id, tool_name, args):
"""tool_start_callback: speak a one-time ack in the voice channel."""
ctx = self._ctx
if ctx._voice_ack_fired[0] or ctx._voice_ack_guild[0] is None or not ctx._run_still_current():
return
ctx._voice_ack_fired[0] = True
adapter = self._runner.adapters.get(Platform.DISCORD)
if adapter is None or not hasattr(adapter, "play_ack_in_voice"):
return
try:
self._schedule(
adapter.play_ack_in_voice(ctx._voice_ack_guild[0]), "voice ack scheduling error", loop=ctx._voice_ack_loop,
)
except Exception as err:
logger.debug("voice ack schedule failed: %s", err)
# Slack-native task cards ride agent.tool_start_callback / tool_complete_callback so start and
# completion correlate by the REAL tool-call id; name-correlated progress_callback text events
# would duplicate cards and mispair concurrent calls.
def _native_card_gate(self) -> bool:
ctx = self._ctx
return bool(ctx.progress_queue) and ctx._run_still_current() and not self._agent_interrupted()
# ── Slack-native task cards: ID-bearing lifecycle callbacks (#29483) ── These ride
# agent.tool_start_callback / agent.tool_complete_callback so start/completion events correlate by the
# REAL tool-call id — the name-correlated text events in progress_callback would duplicate cards and
# mispair concurrent calls to the same tool.
def native_tool_start_callback(self, call_id, tool_name, args):
"""Queue an ID-correlated native progress start from the agent thread."""
if not self._native_card_gate():
return
from agent.display import build_tool_preview
name = str(tool_name or "tool")
self._ctx.progress_queue.put({
"type": "tool.started", "tool_call_id": str(call_id or ""), "tool_name": name,
"preview": build_tool_preview(name, args or {}, max_len=64) or "",
})
def native_tool_complete_callback(self, call_id, tool_name, args, result):
"""Queue the matching native completion using the real tool-call ID."""
if not self._native_card_gate():
return
from agent.display import _detect_tool_failure
name = str(tool_name or "tool")
is_error, _ = _detect_tool_failure(name, result)
self._ctx.progress_queue.put({
"type": "tool.completed", "tool_call_id": str(call_id or ""), "tool_name": name, "is_error": bool(is_error),
})
def combined_tool_start_callback(self, call_id, tool_name, args):
"""Compose the voice ack + native task-card start consumers."""
if self._ctx._voice_ack_guild[0] is not None:
self.voice_ack_callback(call_id, tool_name, args)
if self._ctx._native_slack_task_cards:
self.native_tool_start_callback(call_id, tool_name, args)
# ── hook / status bridges (agent thread → gateway loop) ────────────────────────────────
def _step_callback_sync(self, iteration: int, prev_tools: list) -> None:
ctx = self._ctx
if not ctx._run_still_current():
return
# prev_tools may be list[str] or list[dict] with "name"/"result" keys. Normalise so
# "tool_names" stays backward-compatible for user hooks that do ', '.join(tool_names).
names = [(t.get("name") or "") if isinstance(t, dict) else str(t) for t in (prev_tools or [])]
self._schedule(
ctx._hooks_ref.emit("agent:step", {
"platform": ctx.source.platform.value if ctx.source.platform else "",
"user_id": ctx.source.user_id, "session_id": ctx.session_id,
"iteration": iteration, "tool_names": names, "tools": prev_tools,
}),
"agent:step hook scheduling error",
)
def _event_callback_sync(self, event_type: str, context: dict) -> None:
ctx = self._ctx
try:
asyncio.run_coroutine_threadsafe(ctx._hooks_ref.emit(event_type, context), ctx._loop_for_step)
except Exception as e:
logger.debug("event_callback hook error: %s", e)
def _status_live(self) -> bool:
"""Status adapter present and this run is still the current generation."""
return bool(self._ctx._status_adapter) and self._ctx._run_still_current()
def _send_status_text(self, text: str, metadata, log_message: str) -> None:
ctx = self._ctx
self._schedule(ctx._status_adapter.send(ctx._status_chat_id, text, metadata=metadata), log_message)
def _attach_session_title_callback(self, agent, ctx) -> None:
"""Wire the platform thread-rename lane onto the agent as `_on_session_title`.
The titler runs in the turn prologue, so attach before the run, not after it.
"""
try:
# Gateway auto-title failures are not user-actionable, so never surface them as messages;
# overriding the failure sink keeps CLI on _emit_auxiliary_failure while gateway logs debug.
agent._title_failure_callback = lambda task, exc: logger.debug(
"Gateway auto-title failure suppressed (not user-visible): %s: %s", task, exc,
)
session_id = getattr(agent, "session_id", None)
source = ctx.source
runner = self._runner
# Both lanes spend a rate-limited platform call per title, so they use the model's title
# only (TitleCallback); renaming twice burns Discord's 2-per-10-min budget on a throwaway.
# Relay Discord predicate is shape-only: whether the connector auto-threaded our reply is
# only knowable AFTER delivery, so register eagerly and let the rename lane look up the
# cache at fire time — gating registration on the cache read meant it never registered.
if runner._is_telegram_topic_lane(source):
lane = "_schedule_telegram_topic_title_rename"
elif runner._is_discord_auto_thread_lane(source) or runner._is_relay_discord_channel_lane(source):
lane = "_schedule_discord_semantic_thread_rename"
else:
return
agent._on_session_title = lambda title, title_source: (
title_source == "llm" and getattr(runner, lane)(source, session_id, title)
)
except Exception:
logger.debug("Failed to attach session title callback", exc_info=True)
def _status_callback_sync(self, event_type: str, message: str) -> None:
from gateway.run import _prepare_gateway_status_message, _redact_gateway_user_facing_secrets, _send_or_update_status_coro
ctx = self._ctx
if not self._status_live():
return
prepared = _prepare_gateway_status_message(ctx.source.platform, event_type, message)
if prepared is None:
logger.debug(
"status_callback suppressed for %s/%s: %s",
ctx.source.platform.value if ctx.source.platform else "unknown", event_type,
_redact_gateway_user_facing_secrets(str(message or ""))[:160],
)
return
fut = self._schedule(
_send_or_update_status_coro(ctx._status_adapter, ctx._status_chat_id, event_type, prepared, ctx._status_thread_metadata),
f"status_callback ({event_type}) scheduling error",
)
if fut is not None and ctx._cleanup_progress:
fut.add_done_callback(self._track_future_cleanup_id)
# ── stream consumer / interim commentary wiring ─────────────────────────────────────────
def _setup_stream_consumer(self, platform_key):
ctx = self._ctx
stream_consumer = None
# The streaming-TTS consumer is created on the outer loop thread before run_sync launches;
# run_sync only reads it via the holder for delta-callback wiring.
stts = ctx.streaming_tts_consumer_holder[0]
scfg = getattr(getattr(self._runner, 'config', None), 'streaming', None)
if scfg is None:
from gateway.config import StreamingConfig
scfg = StreamingConfig()
# display.platforms.<plat>.streaming may disable streaming per platform; None = follow global.
plat_streaming = ctx.resolve_display_setting(ctx.user_config, platform_key, "streaming")
want_stream_deltas = (
scfg.enabled and scfg.transport != "off" if plat_streaming is None else bool(plat_streaming)
)
want_interim_messages = ctx.interim_assistant_messages_enabled
if want_stream_deltas or want_interim_messages:
try:
from gateway.stream_consumer import GatewayStreamConsumer
adapter = self._runner._adapter_for_source(ctx.source)
if adapter:
consumer_cfg, pause_typing_before_finalize = self._runner._build_stream_consumer_config(
ctx.source, scfg, adapter, on_missing_cursor="raise",
)
stream_consumer = GatewayStreamConsumer(
adapter=adapter, chat_id=ctx.source.chat_id, config=consumer_cfg,
metadata=ctx._status_thread_metadata,
on_new_message=(
(lambda: ctx.progress_queue.put(("__reset__",))) if ctx.progress_queue is not None else None
),
on_before_finalize=pause_typing_before_finalize,
initial_reply_to_id=ctx.event_message_id, run_still_current=ctx._run_still_current,
)
ctx.stream_consumer_holder[0] = stream_consumer
except Exception as err:
logger.debug("Could not set up stream consumer: %s", err)
# Deltas tee to the stream consumer (when text streaming is on) and to streaming TTS.
delta_sinks = [sc for sc in ((stream_consumer if want_stream_deltas else None), stts) if sc is not None]
stream_delta_cb = None
if delta_sinks:
def stream_delta_cb(text: str) -> None:
if ctx._run_still_current():
for sink in delta_sinks:
sink.on_delta(text)
def interim_assistant_cb(text: str, *, already_streamed: bool = False) -> None:
if not ctx._run_still_current():
return
if stream_consumer is not None:
stream_consumer.on_segment_break() if already_streamed else stream_consumer.on_commentary(text)
elif not already_streamed and ctx._status_adapter and str(text or "").strip():
self._send_status_text(text, ctx._status_thread_metadata, "interim_assistant_callback scheduling error")
return stream_consumer, stream_delta_cb, interim_assistant_cb, want_interim_messages
# ── agent resolution (cache reuse vs fresh build) ───────────────────────────────────────
@dataclasses.dataclass
class _CachedAgentLookup:
agent: Any = None
reused: bool = False
evicted: Any = None # agent evicted under the lock; released off-lock on a daemon thread
def _skip_context_files(self, platform_key) -> bool:
"""gateway.platforms.<plat>.skip_context_files: messaging platforms may opt out of
filesystem-heavy context-file discovery (SOUL.md, AGENTS.md, .cursorrules)."""
platforms_cfg = (self._ctx.user_config.get("gateway") or {}).get("platforms") or {}
# ``hermes gateway setup`` writes ``gateway.platforms`` as a LIST of enabled platform names,
# not a dict; treat any non-dict shape as "no per-platform overrides" rather than crashing.
if not isinstance(platforms_cfg, dict):
return False
return bool((platforms_cfg.get(platform_key) or {}).get("skip_context_files"))
def _cached_sid_is_dead(self, cache_lock, cache) -> tuple:
"""(peeked cached session_id, is_dead) — checked OUTSIDE the cache lock. "cached sid != current
sid" normally means an intentional switch (reuse), but the routing-key self-heal yields the same
shape with an agent bound to a DEAD session; reusing it re-binds the dead sid and loops."""
ctx = self._ctx
peek_sid = None
if cache_lock and cache is not None:
with cache_lock:
entry = cache.get(ctx.session_key)
if entry and len(entry) > 3:
peek_sid = entry[3]
dead = False
if peek_sid is not None and ctx.session_id is not None and peek_sid != ctx.session_id:
with suppress(Exception):
dead = self._runner.session_store._is_session_ended_in_db(peek_sid)
return peek_sid, dead
def _current_message_count(self):
"""Cross-process write guard input: the session's current DB message_count (or None)."""
ctx = self._ctx
if self._runner._session_db is None or not ctx.session_id:
return None
count = None
with suppress(Exception):
# run_sync is off-loop (executor); sync DB is fine.
row = self._runner._session_db._db.get_session(ctx.session_id)
if row:
count = row.get("message_count", 0)
return count
def _pop_cached_agent_for_eviction(self):
"""Evict under the lock but DEFER release (release_clients can block on memory-provider /
socket teardown while the idle sweeper waits on this lock). The turn rebuilds a fresh agent, so
the caller does a SOFT release that keeps sandbox / browser / bg processes."""
from gateway.run import _AGENT_PENDING_SENTINEL
evicted = self._runner._agent_cache.pop(self._ctx.session_key, None)
agent = evicted[0] if isinstance(evicted, tuple) and evicted else None
return agent if agent and agent is not _AGENT_PENDING_SENTINEL else None
def _lookup_cached_agent(self, sig, cache_lock, cache, max_iterations, peek_sid, dead, msg_count):
ctx = self._ctx
out = self._CachedAgentLookup()
if not (cache_lock and cache is not None):
return out
with cache_lock:
cached = cache.get(ctx.session_key)
if not (cached and cached[1] == sig):
return out
# cached[2] = message_count at cache time (stale when a second process appended rows);
# cached[3] = the session_id the snapshot was taken for.
cached_mc = cached[2] if len(cached) > 2 else None
cached_sid = cached[3] if len(cached) > 3 else None
# Same session_key, other conversation: the counts track DIFFERENT DB rows, so the
# comparison is meaningless — REUSE rather than bust the prompt cache on every switch.
sid_mismatch = cached_sid is not None and ctx.session_id is not None and cached_sid != ctx.session_id
# Re-validate the outside-lock dead-session peek against the tuple read under THIS lock:
# a stale "dead" verdict must never be applied to a different (possibly live) agent.
if sid_mismatch and dead and cached_sid == peek_sid:
logger.info(
"Agent cache invalidated for session %s: "
"cached agent's session_id %s is ended in "
"state.db (stale self-heal artifact, "
"#54878 x #54947) — discarding instead of "
"reusing across the routing recovery", ctx.session_key, cached_sid,
)
elif not sid_mismatch and cached_mc is not None and msg_count is not None and msg_count != cached_mc:
logger.info(
"Agent cache invalidated for session %s: "
"message_count changed (%s -> %s), "
"possible cross-process write", ctx.session_key, cached_mc, msg_count,
)
else:
out.agent = cached[0]
# Refresh LRU order so cap enforcement evicts truly-oldest entries.
if hasattr(cache, "move_to_end"):
with suppress(KeyError):
cache.move_to_end(ctx.session_key)
self._runner._init_cached_agent_for_turn(out.agent, ctx._interrupt_depth)
# Cached agent may have been created with old config.
out.agent.max_iterations = max_iterations
logger.debug("Reusing cached agent for session %s", ctx.session_key)
out.reused = True
return out
out.evicted = self._pop_cached_agent_for_eviction()
return out
def _release_evicted_agent(self, agent) -> None:
"""Off-lock soft release on a daemon thread so teardown never blocks the gateway loop."""
self._runner._spawn_release_thread(
self._runner._release_evicted_agent_soft, (agent,), f"agent-xproc-evict-{str(self._ctx.session_key)[:24]}",
inline_fallback=True,
)
def _build_fresh_agent(self, turn_route, platform_key, combined_ephemeral, max_iterations,
reasoning_config, pr, skip_context_files):
from gateway.run import _checkpoint_agent_kwargs
ctx = self._ctx
runner = self._runner
src = ctx.source
return ctx.AIAgent(
model=turn_route["model"], **turn_route["runtime"], **_checkpoint_agent_kwargs(ctx.user_config),
max_iterations=max_iterations, quiet_mode=True, verbose_logging=False,
enabled_toolsets=ctx.enabled_toolsets, disabled_toolsets=ctx.disabled_toolsets,
ephemeral_system_prompt=combined_ephemeral or None,
prefill_messages=runner._prefill_messages or None,
reasoning_config=reasoning_config, service_tier=runner._service_tier,
request_overrides=turn_route.get("request_overrides"),
providers_allowed=pr.get("only"), providers_ignored=pr.get("ignore"), providers_order=pr.get("order"),
provider_sort=pr.get("sort"), provider_require_parameters=pr.get("require_parameters", False),
provider_data_collection=pr.get("data_collection"),
session_id=ctx.session_id, platform=platform_key,
user_id=src.user_id, user_id_alt=src.user_id_alt, user_name=src.user_name,
chat_id=src.chat_id, chat_name=src.chat_name, chat_type=src.chat_type, thread_id=src.thread_id,
gateway_session_key=ctx.session_key,
session_db=getattr(runner._session_db, "_db", runner._session_db),
# Reload from disk — do not reuse the startup snapshot.
# See #60955.
fallback_model=self._runner._refresh_fallback_model(),
skip_context_files=skip_context_files,
# Keep the persona even with minimal context: soul identity is one small file.
load_soul_identity=True,
)
def _resolve_turn_agent(self, turn_route, platform_key, combined_ephemeral, max_iterations, reasoning_config, pr):
"""Reuse this session's cached AIAgent (frozen system prompt + tool schemas → prompt cache
hits) or build a fresh one. Returns (agent, reused_cached_agent)."""
ctx = self._ctx
runner = self._runner
skip_context_files = self._skip_context_files(platform_key)
sig = runner._agent_config_signature(
turn_route["model"], turn_route["runtime"], ctx.enabled_toolsets, combined_ephemeral,
cache_keys=runner._extract_cache_busting_config(ctx.user_config),
user_id=getattr(ctx.source, "user_id", None),
user_id_alt=getattr(ctx.source, "user_id_alt", None),
skip_context_files=skip_context_files,
)
cache_lock = getattr(runner, "_agent_cache_lock", None)
cache = getattr(runner, "_agent_cache", None)
peek_sid, dead = self._cached_sid_is_dead(cache_lock, cache)
msg_count = self._current_message_count()
found = self._lookup_cached_agent(sig, cache_lock, cache, max_iterations, peek_sid, dead, msg_count)
agent = found.agent
# Lock released — refresh the reused agent's fallback chain from disk OUTSIDE the cache lock
# (disk I/O under the lock stalls the idle-sweep watcher and Discord heartbeats). A chain
# configured after caching must reach the next turn; per-session serialization keeps it safe.
if found.reused and agent is not None:
self._runner._apply_fallback_chain_to_agent(agent, runner._refresh_fallback_model())
if found.evicted is not None:
self._release_evicted_agent(found.evicted)
if agent is None:
agent = self._build_fresh_agent(
turn_route, platform_key, combined_ephemeral, max_iterations, reasoning_config, pr, skip_context_files,
)
if cache_lock and cache is not None:
with cache_lock:
# Record the snapshot's session_id with message_count so the cross-process guard
# can skip the meaningless count comparison if the active session_id switches.
cache[ctx.session_key] = (agent, sig, msg_count, ctx.session_id)
runner._enforce_agent_cache_cap()
logger.debug("Created new agent for session %s (sig=%s)", ctx.session_key, sig)
return agent, found.reused
# ── per-turn agent wiring ───────────────────────────────────────────────────────────────
def _notice_callback_sync(self, notice) -> None:
"""Credits / out-of-band notices (usage bands, depletion, restored) fire from the agent's
sync worker thread; hop onto the gateway loop. Fired-once latch lives on the cached agent."""
from gateway.run import render_notice_line
if not self._status_live():
return
try:
line = render_notice_line(notice)
except Exception:
logger.debug("render_notice_line failed", exc_info=True)
return
if line:
self._schedule(self._runner._deliver_platform_notice(self._ctx.source, line), "notice_callback delivery scheduling error")
def _make_bg_review_callbacks(self):
"""(send, release): background-review messages ("💾 Memory updated") are held until the
adapter's post-delivery hook releases them after the main response lands."""
from gateway.run import _interim_metadata, _non_conversational_metadata
ctx = self._ctx
release_evt = threading.Event()
pending: list[str] = []
pending_lock = threading.Lock()
def deliver(message: str) -> None:
if self._status_live():
self._send_status_text(
message,
_interim_metadata(_non_conversational_metadata(ctx._status_thread_metadata, platform=ctx.source.platform)),
"background_review_callback scheduling error",
)
def release() -> None:
release_evt.set()
with pending_lock:
queued = list(pending)
pending.clear()
for message in queued:
deliver(message)
def send(message: str) -> None:
if not self._status_live():
return
if not release_evt.is_set():
with pending_lock:
if not release_evt.is_set():
pending.append(message)
return
deliver(message)
return send, release
@staticmethod
def _merge_turn_request_overrides(agent, turn_route) -> None:
"""Merge, never overwrite: init-time request overrides (e.g. a custom provider's extra_body)
must survive every reused-agent turn. Drop only the PREVIOUS turn's routing overrides before
layering this turn's, so stale per-turn values never linger."""
overrides = dict(getattr(agent, "request_overrides", {}) or {})
for key, value in (getattr(agent, "_gateway_turn_request_overrides", {}) or {}).items():
if overrides.get(key) == value:
overrides.pop(key, None)
turn_overrides = dict(turn_route.get("request_overrides") or {})
overrides.update(turn_overrides)
agent.request_overrides = overrides
agent._gateway_turn_request_overrides = turn_overrides
def _wire_turn_agent_callbacks(self, agent, turn_route, reasoning_config,
stream_delta_cb, interim_assistant_cb, want_interim_messages):
"""Per-message state — callbacks and reasoning config change every turn, so they aren't
baked into the cached agent."""
ctx = self._ctx
runner = self._runner
# ALWAYS attached (never gated to None): its body gates each event class, and subagent-
# failure notices must fire even with tool_progress/thinking off.
agent.tool_progress_callback = ctx.progress_callback
# Discord's one-time voice ack and Slack's task cards both ride the authoritative start
# callback, so neither infers identity from tool names.
agent.tool_start_callback = (
(ctx.native_tool_start_callback or ctx.voice_ack_callback)
if (ctx._voice_ack_guild[0] is not None or ctx._native_slack_task_cards) else None
)
agent.tool_complete_callback = ctx.native_tool_complete_callback if ctx._native_slack_task_cards else None
agent.step_callback = ctx._step_callback_sync if ctx._hooks_ref.loaded_hooks else None
agent.stream_delta_callback = stream_delta_cb
agent.interim_assistant_callback = interim_assistant_cb if want_interim_messages else None
agent.status_callback, agent.notice_callback = ctx._status_callback_sync, self._notice_callback_sync
agent.notice_clear_callback = None # sends can't be retracted
agent.event_callback = ctx._event_callback_sync
agent.reasoning_config, agent.service_tier = reasoning_config, runner._service_tier
self._merge_turn_request_overrides(agent, turn_route)
# Must-deliver notes for THIS turn ride the current user message (api_content sidecar), never
# the system prompt. Assigned unconditionally so a reused agent never replays a stale note.
agent._gateway_turn_context_notes = "\n\n".join(runner._consume_pending_turn_sidecar_notes(ctx.session_key))
agent.background_review_callback, bg_release = self._make_bg_review_callbacks()
# Register the release hook on the adapter so base.py's finally block fires it after the
# main response is delivered.
if ctx._status_adapter and ctx.session_key:
if getattr(type(ctx._status_adapter), "register_post_delivery_callback", None) is not None:
ctx._status_adapter.register_post_delivery_callback(ctx.session_key, bg_release, generation=ctx.run_generation)
else:
pdc = getattr(ctx._status_adapter, "_post_delivery_callbacks", None)
if pdc is not None:
pdc[ctx.session_key] = bg_release
# display.memory_notifications: off | on (generic "💾 Memory updated", default) | verbose.
mem_notif = ctx.user_config.get("display", {}).get("memory_notifications")
if isinstance(mem_notif, bool):
mem_notif = "on" if mem_notif else "off"
agent.memory_notifications = str(mem_notif).lower() if mem_notif else "on"
agent.clarify_callback = self._clarify_callback_sync
# Thinking between tool calls is independent of tool_progress mode (Mattermost opts in
# per platform so global scratch-text doesn't leak into threads).
agent.thinking_progress = ctx._thinking_enabled
ctx.agent_holder[0] = agent # interrupt support
# The titler fires from the turn prologue, so attach the rename lane before the run.
self._attach_session_title_callback(agent, ctx)
# Publish turn ownership for /stop, /new, disconnect and shutdown interrupts; older session
# processes are outside this baseline and remain alive.
agent._gateway_turn_process_task_id, agent._gateway_turn_process_baseline = ctx.process_task_id, ctx.process_baseline
ctx.tools_holder[0] = getattr(agent, "tools", None) # transcript logging
# ── blocking prompts from the agent thread (approval / clarify) ─────────────────────────
def _close_native_stream_boundary(self, reason: str, placeholder: str | None = None, reopen: bool = False) -> bool:
"""Native-streaming platforms (e.g. WeCom): an interrupting interaction (approval or clarify
prompt) must finalize the current stream first, or post-interaction output keeps updating the
OLD bubble above the prompt. Runs on the agent thread; the consumer serializes via its queue."""
sc = self._stream_consumer()
if not (sc and getattr(sc, "_use_native_streaming", False)):
return True
cancelled_flag = None
try:
boundary = sc.close_for_approval_prompt(placeholder, reason=reason, reopen=reopen)
# Returns (future, cancelled_flag) or just a future.
if isinstance(boundary, tuple):
boundary, cancelled_flag = boundary
if not hasattr(boundary, "result"):
return True
ok = boundary.result(timeout=10)
if not ok:
logger.warning(
"%s boundary failed to close stream properly — "
"prompt may still appear in typing bubble", reason,
)
return bool(ok)
except (TimeoutError, Exception) as err:
if cancelled_flag is not None:
cancelled_flag["cancelled"] = True
logger.warning("%s boundary timed out or failed: %s", reason, err)
return False
def _clarify_callback_sync(self, question: str, choices, multi_select: bool = False) -> str:
"""Present a clarify prompt and block on a response (clarify_tool's synchronous contract):
schedule send_clarify on the gateway loop, block on the primitive's threading.Event with a
timeout. Returns the response string, or a sentinel when none arrived."""
from gateway.run import _clarify_send_then_wait
from tools import clarify_gateway as clarify_mod
import uuid
ctx = self._ctx
if not ctx._status_adapter:
return ""
session_key = ctx.session_key or ""
clarify_id = uuid.uuid4().hex[:10]
choices = list(choices) if choices else None
clarify_mod.register(
clarify_id=clarify_id, session_key=session_key, question=question, choices=choices,
multi_select=bool(multi_select),
)
# Unlike approval, clarify passes reopen=True so the continuation re-opens a native stream
# below the question; if the re-seed fails the consumer degrades to send() automatically.
self._close_native_stream_boundary("Clarify", "💬 等待你的选择...", reopen=True)
# Pause typing: a "thinking..." status must not obscure the prompt or block an "Other" reply
# on platforms that disable input while typing (Slack Assistant).
with suppress(Exception):
ctx._status_adapter.pause_typing_for_chat(ctx._status_chat_id)
# Ordering barrier: flush buffered assistant prose BEFORE the poll, which goes out on a
# separate agent-thread-blocking path and would otherwise render ABOVE its own explanation.
# Best-effort + short timeout so the agent thread never hangs if the consumer isn't running.
flush = getattr(self._stream_consumer(), "flush_pending_sync", None)
try:
if callable(flush):
flush(timeout=3.0)
except Exception:
logger.debug("Stream-consumer flush before clarify prompt failed", exc_info=True)
fut = self._schedule(
ctx._status_adapter.send_clarify(
chat_id=ctx._status_chat_id, question=question, choices=choices, clarify_id=clarify_id,
session_key=session_key, metadata=ctx._status_thread_metadata,
),
"Clarify send failed to schedule",
)
# Boundary rule (see _approval_send_outcome): a send timeout is AMBIGUOUS — the card may
# have posted with a late ack. Only a definitive failure tears down the registration;
# ambiguous falls through to the bounded wait so a late reply resolves.
response = _clarify_send_then_wait(fut, clarify_id=clarify_id, session_key=session_key, clarify_mod=clarify_mod)
# Only re-arm typing when the user actually answered — the undeliverable sentinel and the
# timeout/cancellation strings start with '[' and must pass through untouched.
if not (isinstance(response, str) and response.startswith("[")):
# Reopen typing IMMEDIATELY, not on the LLM's first post-answer token (native streaming
# otherwise re-seeds lazily on the first delta: ~48s of dead air). request_reopen_seed is
# a no-op outside the reopen-pending native state.
sc = self._stream_consumer()
if sc is not None:
try:
sc.request_reopen_seed()
except Exception:
logger.debug("request_reopen_seed after clarify answer failed", exc_info=True)
try:
ctx._status_adapter.resume_typing_for_chat(ctx._status_chat_id)
except Exception:
logger.debug("resume_typing_for_chat after clarify answer failed", exc_info=True)
return response
def _approval_notify_sync(self, approval_data: dict) -> None:
"""Send the approval request from the agent thread: the adapter's interactive button
approvals (``send_exec_approval``) when available, else plain text with ``/approve`` steps."""
from gateway.run import _approval_send_outcome, _format_exec_approval_fallback, _interim_metadata, _redact_approval_command
ctx = self._ctx
adapter = ctx._status_adapter
# Slack's assistant_threads_setStatus disables the compose box, so the user can't type
# /approve while "is thinking..." shows. Pausing stops _keep_typing re-setting it; resumed
# in approve/deny.
adapter.pause_typing_for_chat(ctx._status_chat_id)
self._close_native_stream_boundary("Approval")
# Redact credentials before display: Tirith's findings are already redacted, but the raw
# command string still leaks secrets. Both the button and plain-text paths use this value.
cmd = _redact_approval_command(approval_data.get("command", ""))
desc = approval_data.get("description", "dangerous command")
flags = {k: approval_data.get(k, d) for k, d in (("allow_permanent", True), ("allow_session", True), ("smart_denied", False))}
# Check the *class*, not the instance — MagicMock auto-creates attributes in tests.
if getattr(type(adapter), "send_exec_approval", None) is not None:
try:
fut = self._schedule(
adapter.send_exec_approval(
chat_id=ctx._status_chat_id, command=cmd, session_key=ctx.session_key or "",
description=desc, metadata=ctx._status_thread_metadata, **flags,
),
"send_exec_approval scheduling error",
)
if fut is None:
raise RuntimeError("send_exec_approval: loop unavailable")
outcome = _approval_send_outcome(fut, timeout=15)
if outcome == "sent":
return
if outcome == "ambiguous":
# Timeout ≠ failure: the card may have posted with a late ack. The prompt
# registration stays alive so a tap still resolves; re-sending made duplicate
# cards + orphaned "/approve: nothing pending".
logger.warning(
"Button-based approval send timed out — treating "
"as possibly-delivered (no re-send; the prompt "
"stays armed for a late tap)"
)
return
if outcome == "declined":
# P5(b): the connector AUTHORIZED this destination and
# refused it. The text fallback below re-sends the same
# content to the same chat, which would turn a refused
# button card into a delivered plain-text one — the exact
# leak the egress guard exists to stop. A decline is
# definitive, so unlike `ambiguous` the registration is
# torn down; unlike `failed`, nothing is re-sent.
logger.warning(
"Button-based approval DECLINED by the connector's "
"egress guard — not falling back to text (the "
"destination is not approved for this connection)"
)
# RAISE, do not return. This function is the notify_cb for
# `_await_gateway_decision`, which already has a correct
# undeliverable path: a raising notify drops the queue entry
# and returns `notify_failed`, unblocking the tool. Returning
# quietly suppressed the text fallback (right) but left the
# CENTRAL approval entry pending (wrong) — the dangerous
# command then blocked until the approval timeout. My earlier
# comment claimed the registration was torn down; only the
# adapter's private prompt map was.
raise _ExecApprovalDeclined(
"exec approval undeliverable: connector egress declined "
"this destination"
)
logger.warning("Button-based approval failed (send returned error), falling back to text")
except _ExecApprovalDeclined:
# Must escape this handler: the fallback below is a text send to
# the destination the connector just refused.
raise
except Exception as e:
logger.warning("Button-based approval failed, falling back to text: %s", e)
# Plain-text prompt with the adapter's typed prefix (e.g. `!approve`): typed "/" is blocked
# in Slack threads and reserved by Matrix clients.
msg = _format_exec_approval_fallback(cmd, desc, getattr(adapter, "typed_command_prefix", "/"), **flags)
try:
# Mark as approval prompt so WeCom routes through the control lane.
metadata = {**(ctx._status_thread_metadata or {}), "is_approval_prompt": True}
fut = self._schedule(
adapter.send(ctx._status_chat_id, msg, metadata=_interim_metadata(metadata)), "Approval text-send scheduling error",
)
if fut is not None:
fut.result(timeout=15)
except Exception as e:
logger.error("Failed to send approval request: %s", e)
# ── run_sync phases ─────────────────────────────────────────────────────────────────────
def _load_turn_history(self, agent, reused_cached_agent):
from gateway.run import (
_build_gateway_agent_history, _collect_history_media_paths, _message_timestamps_enabled,
_select_cached_agent_history,
)
ctx = self._ctx
# Transcript rows ({role, content, timestamp}) lose timestamps; interrupt-path agent messages
# (tool_calls/tool_call_id/reasoning) pass through intact so the API sees valid assistant→tool
# sequences. Telegram observed=True rows are withheld from replayable history and attached to
# the current addressed message as API-only context.
agent_history, observed_group_context = _build_gateway_agent_history(
ctx.history, channel_prompt=ctx.channel_prompt, inject_timestamps=_message_timestamps_enabled(ctx.user_config),
)
# FTS write-corruption guard: if persistence failed silently the reloaded transcript is stale
# while the SAME cached agent still holds the live conversation (same-session amnesia). Only
# for a reused agent bound to this exact session_id.
# Replacing the live transcript with that shorter copy causes immediate same-session amnesia. See
# #50502.
if reused_cached_agent and getattr(agent, "session_id", None) == ctx.session_id:
selected = _select_cached_agent_history(agent_history, getattr(agent, "_session_messages", None))
if selected is not agent_history:
logger.warning(
"Persisted transcript lagged live cached history for "
"session %s (disk=%d, memory=%d); preserving live "
"conversation context (possible FTS write corruption)",
ctx.session_key, len(agent_history), len(selected),
)
# The live history bypassed _build_gateway_agent_history's cleanup — re-apply the
# stale-confirmation expiry so a dangerous confirmation can't slip through.
agent_history = strip_stale_dangerous_confirmations(selected, now=time.time())
# MEDIA paths already in history are excluded from this turn's extraction (compression-safe).
return agent_history, observed_group_context, _collect_history_media_paths(agent_history)
def _prepend_pending_note(self, attr: str) -> None:
"""Consume a one-shot per-session note (model switch, /reload-skills) into the NEXT user
message. Nothing hits the transcript out-of-band, so alternation stays intact."""
ctx = self._ctx
notes = getattr(self._runner, attr, None)
note = notes.pop(ctx.session_key, None) if notes and ctx.session_key and ctx.session_key in notes else None
if note:
ctx.message = note + "\n\n" + ctx.message
def _resume_note_interactive(self) -> bool:
"""Interactive platforms report the restore and ask what next; event platforms (webhook,
API server) continue the work — nobody is present to answer."""
return bool(getattr(self._runner._adapter_for_source(self._ctx.source), "interactive_resume", True))
def _prepare_turn_message(self, agent_history):
"""Prepend recovery/notice guidance to ``ctx.message``.
Returns (persist_user_message_override, persist_user_timestamp_override): real user text is
kept separate from API-only recovery guidance so stale guidance never replays as user text.
"""
from gateway.run import (
_auto_continue_freshness_window, _is_fresh_gateway_interruption,
_last_transcript_timestamp, _prepare_resume_pending_message, build_resume_recovery_note,
)
ctx = self._ctx
persist_override: Optional[Any] = ctx.persist_user_message
self._prepend_pending_note("_pending_model_notes")
# Auto-continue: history ending with a tool result means the previous turn was cut off
# (restart, crash, SIGTERM). Session-level resume_pending (drain-timeout shutdown) uses
# stronger reason-aware wording that subsumes this case. Both gate on the age of
# ``history[-1]`` (not agent_history, which stripped tool-row timestamps); no stamp = fresh.
window = _auto_continue_freshness_window()
interruption_is_fresh = _is_fresh_gateway_interruption(_last_transcript_timestamp(ctx.history), window_secs=window)
entry = None
if ctx.session_key:
with suppress(Exception):
entry = self._runner.session_store._entries.get(ctx.session_key)
resume_pending = entry is not None and getattr(entry, "resume_pending", False)
resume_reason = (getattr(entry, "resume_reason", None) or "restart_timeout") if resume_pending else None
# resume_pending freshness ALSO uses the restart watchdog's ``last_resume_marked_at`` (the true
# interruption stamp): the transcript clock can be hours older for an active thread, and the
# startup auto-resume turn has empty text, so gating on it alone yields a blank user message.
mark_is_fresh = resume_pending and _is_fresh_gateway_interruption(
getattr(entry, "last_resume_marked_at", None), window_secs=window,
)
if resume_pending and (interruption_is_fresh or mark_is_fresh):
# Empty message = the startup auto-resume turn; there is no NEW user message.
ctx.message, persist_override = _prepare_resume_pending_message(
resume_reason, ctx.message, interactive=self._resume_note_interactive(),
)
elif agent_history and agent_history[-1].get("role") == "tool" and interruption_is_fresh:
persist_override = ctx.message
ctx.message = (
"[System note: A new message has arrived. The conversation "
"history contains pending tool outputs from an interrupted turn. "
"IGNORE those pending results. Address the user's NEW message "
"below FIRST. Do NOT re-execute old tool calls from the history.]\n\n"
+ ctx.message
)
self._prepend_pending_note("_pending_skills_reload_notes")
# Safety net: a startup auto-resume event carries empty text; if the resume_pending branch
# did not fire (freshness signals disagreed, marker cleared) we must NOT hand the model a blank
# user turn. Restricted to resume_pending sessions so caption-less image turns are untouched.
if isinstance(ctx.message, str) and not ctx.message.strip() and resume_pending:
ctx.message = build_resume_recovery_note(resume_reason, "", interactive=self._resume_note_interactive())
return persist_override, ctx.persist_user_timestamp
def _native_image_run_message(self):
"""Wrap the user turn as an OpenAI-style multimodal content list when
_prepare_inbound_message_text buffered image paths; consume-and-clear so later turns on the
same runner never re-attach stale images. Falls back to plain text when nothing is readable."""
ctx = self._ctx
native_imgs = self._runner._consume_pending_native_image_paths(ctx.session_key)
if not native_imgs:
return ctx.message
try:
from agent.image_routing import build_native_content_parts
parts, skipped = build_native_content_parts(ctx.message, native_imgs)
if skipped:
logger.warning("Native image attachment: skipped %d unreadable path(s): %s", len(skipped), skipped)
if any(p.get("type") == "image_url" for p in parts):
return parts
except Exception as exc:
logger.warning("Native image attachment failed, falling back to text: %s", exc)
return ctx.message
def _run_conversation_with_approval(self, agent, agent_history, observed_group_context,
persist_user_message_override, persist_user_timestamp_override):
"""Run the turn with the per-session gateway approval callback registered: dangerous-command
approval blocks the agent thread (mirrors CLI input()); the callback bridges sync→async."""
from gateway.run import _wrap_current_message_with_observed_context
from tools.approval import register_gateway_notify, unregister_gateway_notify
from tools.approval_context import reset_current_session_key, set_current_session_key
ctx = self._ctx
session_key = ctx.session_key or ""
token = set_current_session_key(session_key)
register_gateway_notify(session_key, self._approval_notify_sync)
try:
api_message = _wrap_current_message_with_observed_context(self._native_image_run_message(), observed_group_context)
kwargs = {"conversation_history": agent_history, "task_id": ctx.session_id}
if persist_user_message_override is not None:
kwargs["persist_user_message"] = persist_user_message_override
elif observed_group_context:
kwargs["persist_user_message"] = ctx.message
if ctx.persist_user_display_kind:
# Internal self-injected turn: type the persisted user row so UIs render it as a
# timeline notice, not a user bubble (stripped from provider payloads downstream).
kwargs["persist_user_display_kind"] = ctx.persist_user_display_kind
if ctx.persist_user_display_metadata:
kwargs["persist_user_display_metadata"] = ctx.persist_user_display_metadata
if ctx.moa_config is not None:
kwargs["moa_config"] = ctx.moa_config
if persist_user_timestamp_override is not None:
kwargs["persist_user_timestamp"] = persist_user_timestamp_override
# The RAW inbound id (not event_message_id, the reply anchor) rides the persisted user
# turn so a restart-interrupted turn is recorded WITH its id for drain-window dedup.
if ctx.inbound_message_id is not None:
kwargs["persist_user_platform_id"] = str(ctx.inbound_message_id)
return agent.run_conversation(api_message, **kwargs)
finally:
unregister_gateway_notify(session_key)
# Cancel pending clarify entries so blocked agent threads don't hang past the end of the
# run (interrupt, completion, gateway shutdown). Idempotent.
with suppress(Exception):
from tools.clarify_gateway import clear_session
clear_session(session_key)
reset_current_session_key(token)
def _finish_stream_consumer(self, result, agent_history, stream_consumer):
ctx = self._ctx
# Canonicalize a model-emitted computer-use screenshot path at the common result boundary so
# the streaming finalizer and the non-streaming delivery path see the same response.
if isinstance(result, dict) and isinstance(result.get("final_response"), str):
result["final_response"] = repair_explicit_computer_use_media_paths(
result["final_response"], result.get("messages", []), history_offset=len(agent_history),
)
ctx.result_holder[0] = result
if stream_consumer is None:
return
# Pass final_response as the authoritative finalize payload: it includes post-stream
# augmentation (verifier footer, explainer) the accumulator never saw. Adopt ONLY a genuinely
# completed final: interrupt paths return {interrupted: True, completed: False} with a
# DIAGNOSTIC final_response — adopting it would seal the partial answer over with the
# diagnostic AND suppress the gateway's own error delivery.
_final_for_stream = None
if (
isinstance(result, dict) and not result.get("failed") and not result.get("interrupted")
and result.get("completed") is not False
):
fr = result.get("final_response")
if isinstance(fr, str) and fr.strip() and fr != "(empty)":
_final_for_stream = fr
if _final_for_stream is None:
stream_consumer.finish()
return
# Duck-type safe: test doubles / older consumers may expose a zero-arg finish().
try:
stream_consumer.finish(_final_for_stream)
except TypeError:
stream_consumer.finish()
def _restore_telegram_thread_id_after_split(self, agent_session_id) -> None:
"""Telegram DM whose source.thread_id was lost in the session split (synthetic/recovered
event): restore it from the binding so _thread_metadata_for_source yields the right
message_thread_id instead of the General thread (non-fatal)."""
ctx = self._ctx
try:
# run_sync is off-loop (executor); sync DB is fine.
binding = self._runner._session_db._db.get_telegram_topic_binding_by_session(session_id=agent_session_id)
if binding and binding.get("thread_id"):
ctx.source.thread_id = str(binding["thread_id"])
logger.debug(
"Restored source.thread_id=%s from binding after session split %s → %s",
ctx.source.thread_id, ctx.session_id, agent_session_id,
)
except Exception:
logger.debug("Failed to restore thread_id from binding after session split", exc_info=True)
def _sync_session_after_run(self, agent_history):
"""Sync session_id right after run_conversation(): compression can rotate before a follow-up
model call fails, and the failure return must still point at the compressed child.
Returns (compacted_in_place, effective_session_id, effective_history_offset)."""
ctx = self._ctx
runner = self._runner
agent = ctx.agent_holder[0]
# In-place compaction compacts the transcript WITHOUT rotating the id, so the id-change diff
# can't see it; compress_context() sets this flag and the gateway re-baselines as for a split.
compacted_in_place = bool(getattr(agent, "_last_compaction_in_place", False)) if agent else False
agent_session_id = getattr(agent, 'session_id', ctx.session_id) if agent else ctx.session_id
session_was_split = bool(agent and ctx.session_key and agent_session_id != ctx.session_id)
if session_was_split:
logger.info("Session split detected: %s → %s (compression)", ctx.session_id, agent_session_id)
entry = runner.session_store._entries.get(ctx.session_key)
persisted = False
if entry:
entry_session_id = getattr(entry, "session_id", None)
if not ctx._run_still_current():
logger.info(
"Skipping session split sync for stale run %s — "
"generation %s is no longer current",
ctx.session_key or "?", ctx.run_generation,
)
elif entry_session_id == agent_session_id:
persisted = True
elif entry_session_id != ctx.session_id:
logger.info(
"Skipping session split sync for %s because the "
"session binding moved from %s to %s before "
"compression finished",
ctx.session_key or "?", ctx.session_id, entry_session_id,
)
else:
entry.session_id = agent_session_id
runner.session_store._save()
runner.session_store._record_gateway_session_peer(agent_session_id, ctx.session_key, ctx.source)
persisted = True
# Only after this run published its split — a stale /stop→/new predecessor must not
# mutate routing state.
if persisted:
src = ctx.source
if (
getattr(src, "platform", None) == Platform.TELEGRAM and getattr(src, "chat_type", None) == "dm"
and getattr(src, "thread_id", None) is None and runner._session_db is not None
):
self._restore_telegram_thread_id_after_split(agent_session_id)
runner._sync_telegram_topic_binding(src, entry, reason="agent-run-compression")
runner._sync_session_model_from_agent(agent_session_id, agent)
# history_offset=0 whenever the agent's message list lost the original history prefix
# (split OR in-place compaction): the returned `messages` is the compacted set, persist all
# of it; slicing past the pre-compaction length would drop everything.
offset = 0 if (session_was_split or compacted_in_place) else len(agent_history)
return compacted_in_place, agent_session_id, offset
def _combined_ephemeral_prompt(self) -> str:
"""Platform context + YAML channel_prompts hint + channel_overrides system_prompt (or global
ephemeral) + the gateway ephemeral prompt."""
ctx = self._ctx
combined = ctx.context_prompt or ""
for extra in (
(ctx.channel_prompt or "").strip(),
self._runner._get_system_prompt_for_channel(
ctx.source.platform, ctx.source.chat_id or "", thread_id=getattr(ctx.source, "thread_id", None),
parent_id=getattr(ctx.source, "parent_chat_id", None),
),
):
if extra:
combined = (combined + "\n\n" + extra).strip()
return combined
def _append_auto_media_tags(self, final_response: str, result, agent_history, history_media_paths) -> str:
"""Append MEDIA:<path> tags from tool results (e.g. TTS) that the model's final text omits, so
extract_media() delivers each file once. Scoped to THIS turn (slice at len(agent_history)) so
a stale MEDIA: path from an earlier turn never rides a later reply; the history-path dedup is
the secondary guard — and the sole one when mid-run compression shrank the list."""
from gateway.run import _collect_auto_append_media_tags
if "MEDIA:" in final_response:
return final_response
# Scan tool results for MEDIA:<path> tags that need to be delivered as native audio/file
# attachments. The TTS tool embeds MEDIA: tags in its JSON response, but the model's final text
# reply usually doesn't include them. We collect unique tags from tool results and append any that
# aren't already present in the final response, so the adapter's extract_media() can find and
# deliver the files exactly once. Scope the scan to THIS turn's tool results only. ``agent_history``
# was passed into run_conversation as ``conversation_history``, so the agent's returned ``messages``
# list is ``agent_history`` followed by the messages produced this turn. Slicing at
# ``len(agent_history)`` isolates the current turn precisely, so a stale MEDIA: path emitted by a
# tool several turns earlier (still present in the full message list) can never leak onto a later
# text-only reply. (Fixes #34608) Path-based deduplication against _history_media_paths (collected
# before run_conversation) is retained as a secondary guard. It is also the sole guard on the
# fallback branch taken when mid-run context compression shrinks the message list below the original
# history length, preserving the compression-safe behaviour of #160.
media_tags, has_voice_directive = _collect_auto_append_media_tags(
result.get("messages", []), history_offset=len(agent_history), history_media_paths=history_media_paths,
)
if not media_tags:
return final_response
unique_tags = (["[[audio_as_voice]]"] if has_voice_directive else []) + list(dict.fromkeys(media_tags))
return final_response + "\n" + "\n".join(unique_tags)
def run_sync(self):
"""Executor-thread body of the turn; returns the gateway result dict.
The turn message lives on the shared TurnContext (``ctx.message``) so ``_run_agent_inner`` sees
every rebind. session_key propagates via contextvars (_set_session_env / set_current_session_key)
— never os.environ["HERMES_SESSION_KEY"], which would misroute approvals across sessions.
"""
from gateway.run import _current_max_iterations, _normalize_empty_agent_response, _sanitize_gateway_final_response
ctx = self._ctx
runner = self._runner
# Platform.LOCAL ("local") maps to the "cli" hint key the agent understands.
# session_key is propagated via contextvars in _set_session_env() (_SESSION_KEY) and via
# set_current_session_key() (_approval_session_key) below — both concurrency-safe and inherited by
# tool worker threads. We deliberately do NOT write os.environ["HERMES_SESSION_KEY"] here:
# os.environ is process-global, so concurrent gateway sessions (e.g. two Discord threads) would
# clobber each other's value, and a tool thread whose contextvar is unset would fall back to
# os.environ and read the wrong session key — misrouting command-approval prompts to the wrong
# thread (#24100). The non-gateway surfaces don't depend on this write: CLI and cron bind the
# session via contextvars (set_current_session_key / session context), and only the TUI slash-worker
# *subprocess* exports HERMES_SESSION_KEY (from its own --session-key argv, a separate process) — so
# removing this in-process gateway write does not affect any of them.
platform_key = "cli" if ctx.source.platform == Platform.LOCAL else ctx.source.platform.value
combined_ephemeral = self._combined_ephemeral_prompt()
max_iterations = _current_max_iterations()
try:
model, runtime_kwargs = runner._resolve_session_agent_runtime(
source=ctx.source, session_key=ctx.session_key, user_config=ctx.user_config,
)
logger.debug(
"run_agent resolved: model=%s provider=%s session=%s",
model, runtime_kwargs.get("provider"), ctx.session_key or "",
)
except Exception as exc:
return {"final_response": f"⚠️ Provider authentication failed: {exc}", "messages": [], "api_calls": 0, "tools": []}
pr = runner._provider_routing
reasoning_config = runner._resolve_session_reasoning_config(source=ctx.source, session_key=ctx.session_key, model=model)
runner._reasoning_config = reasoning_config
runner._service_tier = runner._resolve_session_service_tier(source=ctx.source, session_key=ctx.session_key)
stream_consumer, stream_delta_cb, interim_cb, want_interim = self._setup_stream_consumer(platform_key)
turn_route = runner._resolve_turn_agent_config(ctx.message, model, runtime_kwargs)
agent, reused_cached_agent = self._resolve_turn_agent(
turn_route, platform_key, combined_ephemeral, max_iterations, reasoning_config, pr,
)
self._wire_turn_agent_callbacks(agent, turn_route, reasoning_config, stream_delta_cb, interim_cb, want_interim)
agent_history, observed_group_context, history_media_paths = self._load_turn_history(agent, reused_cached_agent)
persist_msg, persist_ts = self._prepare_turn_message(agent_history)
result = self._run_conversation_with_approval(agent, agent_history, observed_group_context, persist_msg, persist_ts)
self._finish_stream_consumer(result, agent_history, stream_consumer)
# The streaming-TTS consumer's finish() runs on the outer loop thread after the executor
# returns, so early run_sync returns are also finalised.
# See the outer finally/completion section below. See #60671.
final_response = result.get("final_response")
# Actual token counts from the agent instance used for this run.
agent = ctx.agent_holder[0]
has_comp = bool(agent) and hasattr(agent, "context_compressor")
comp = agent.context_compressor if has_comp else None
usage = {
"last_prompt_tokens": getattr(comp, "last_prompt_tokens", 0) if has_comp else 0,
"input_tokens": getattr(agent, "session_prompt_tokens", 0) if has_comp else 0,
"output_tokens": getattr(agent, "session_completion_tokens", 0) if has_comp else 0,
"model": getattr(agent, "model", None) if agent else None,
"context_length": (getattr(comp, "context_length", 0) or 0) if has_comp else 0,
}
compacted_in_place, effective_session_id, history_offset = self._sync_session_after_run(agent_history)
# failure_reason must survive the empty-response path too (TUI billing, transient-failure
# persistence). compression_deferred (soft lock-contention defer) is distinct from
# compression_exhausted so the gateway never auto-resets a session a concurrent compressor is
# about to shrink.
common = {
"messages": result.get("messages", []), "api_calls": result.get("api_calls", 0),
"failed": result.get("failed", False), "failure_reason": result.get("failure_reason"),
"partial": result.get("partial", False), "completed": result.get("completed"),
"interrupted": result.get("interrupted", False), "interrupt_message": result.get("interrupt_message"),
"error": result.get("error"),
"compression_exhausted": result.get("compression_exhausted", False),
"compression_deferred": result.get("compression_deferred", False),
"tools": ctx.tools_holder[0] or [],
"history_offset": history_offset, "compacted_in_place": compacted_in_place, "session_id": effective_session_id,
**usage,
}
if not final_response:
final_response = _normalize_empty_agent_response(result, final_response or "", history_len=len(agent_history))
final_response = _sanitize_gateway_final_response(ctx.source.platform, final_response)
if not final_response:
final_response = f"⚠️ {result['error']}" if result.get("error") else ""
# NOTE: deliberately omits agent_persisted/last_reasoning/response_* — the caller
# defaults agent_persisted differently when the key is absent.
return {"final_response": final_response, **common}
final_response = self._append_auto_media_tags(final_response, result, agent_history, history_media_paths)
# Auto-titling runs at TURN START (agent/turn_context.py) from the user's message alone, so a
# failed/interrupted turn is still titled.
return {
"final_response": final_response, "last_reasoning": result.get("last_reasoning"), **common,
"response_previewed": result.get("response_previewed", False),
"response_transformed": result.get("response_transformed", False),
# Lets the persistence block tell whether the codex app-server path self-persisted (it
# didn't — see codex_runtime.py); default True keeps skip-db for the standard runtime.
"agent_persisted": result.get("agent_persisted", True),
}