Fold the transient/permanent sync-loop tests into one parametrized table
and add the two classifier branches that had no loop-level coverage: a
401 whose body was rewritten to HTML by a reverse proxy (errcode dropped,
so only http_status can stop the loop) and a 429 M_LIMIT_EXCEEDED that
must be retried because neither errcode nor status is an auth signal.
Deleting the http_status fallback in _is_permanent_matrix_auth_error now
fails the 401-html case.
Hoist _sync_error to module level so parametrize can call it directly
instead of the staticmethod.__func__ workaround. Trim the
test_ws_auth_retry docstring, which still described a Matrix test class
that moved to test_matrix.py.
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.
With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).
The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.
A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.
Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.
Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.
(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.
I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.
I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.
While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.
Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.
On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.
(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.
I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.
Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.
(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)
The zero-knob test only asserted that ack_stale still trips with the knob
at 0. Because _read_websocket_health evaluates ack age before the
event-silence dimension, the test stayed green even with the
`_event_max_silence_seconds > 0` guard deleted: it never observed the
stale stamp being ignored. Assert (True, "healthy") with knob=0, a stale
stamp and a green transport first, then make the ACK stale and keep the
ack_stale assertion. Dropping the guard now fails this test.
Also correct the on_socket_event_type comment: discord.py dispatches
socket_event_type before the op-code switch but only for a non-null `t`;
heartbeat ACK frames carry `t: null`, they do not "return before" it.
The event-silence tests duplicated `_make_adapter`/`_connect` from
test_discord_liveness.py and carried a bespoke handler-wait loop. Reuse
the sibling helpers instead: `_make_adapter` grows an optional
`max_event_silence` that is only written into `extra` when given, so the
sibling's own tests keep the adapter default and stay unchanged.
Two invariants now bind the fix:
- deaf socket: a `None` stamp (no DISPATCH yet) reads healthy across
several probe intervals, and once armed, transport-green + event
silence trips `event_silence` through `_liveness_loop`.
- `websocket_event_max_silence_seconds: 0` disables only that
dimension: with a stale stamp and a stale heartbeat ACK the probe must
still run and trip on `ack_stale`. This goes red if the knob is moved
into `_start_liveness_probe`'s all-or-nothing guard (#109782).
The YAML->extra passthrough test also asserts the new key, so dropping
its `_YAML_WEBSOCKET_LIVENESS_KEYS` entry fails.
Trim the new event-silence file from 8 tests to the 2 that bind the fix:
the deaf-socket e2e through `_liveness_loop` (transport green, no
DISPATCH → probe trips with the `event_silence` reason) and the
None-stamp window reading healthy (no false trip on quiet reconnects).
The removed cases re-checked knob parsing, defaults and the YAML seed
loop already covered by the sibling liveness knobs' tests, or restated
the two kept invariants from a different angle.
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.
This adds the dispatch-side signal that was requested instead:
- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
every parsed DISPATCH frame with no debug gate (verified live against
the real received_message path: 4/4 frames fired with
enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
incident report's field-proven operator bound); 0 opts out of this
dimension ONLY — the #109782 review failure put the knob in
_start_liveness_probe's all-or-nothing guard, killing the whole
watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out
Fixes#109521
(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
Follow-up on the salvaged #110928 (CLI `--parent-chat-id` / `--guild-id`):
- `_claim_for_sub` skipped a thread-shaped row that matched no `profile_routes`
entry at DEBUG on every tick. Legacy rows written before the flags existed can
never match a channel-level route (no `parent_chat_id`), so the notifier now
logs ONE WARNING per row naming the task, the thread and the re-subscribe
command. Still fail-closed: the events stay unclaimed.
- Docs: the kanban user guide explains the anchors and shows the Discord-thread
subscribe command under `profile_routes`.
- Test (red on origin/main): two collects → exactly one WARNING, events unseen.
Under `gateway.multiplex_profiles: true` every gateway-hosted session that ended
inside the gateway process logged `Memory provider 'openviking' on_session_end
failed: get_secret('OPENVIKING_API_KEY') called with no profile secret scope
active` and skipped its end-of-session commit. The provider lifecycle hooks
(`flush_pending` -> `on_session_end` -> provider teardown -> `close`) read
credentials and home at call time; the stop-time finalize pass and the
idle-cache sweep run on the main loop outside any adapter handler, so
`_run_in_executor_with_context` copied an EMPTY scope into the worker.
Route `_cleanup_agent_resources_off_loop` and `_finalize_session_off_loop`
through `_run_release_in_profile_scope` (the seam cache eviction already uses):
a scoped caller keeps its scope; an unscoped caller passes the session key and
the owner's profile scope is entered from it. The stop path passes the key for
both active and idle-cached agents. Cron's post-run cleanup thread is the
sibling seam and is fixed by the cherry-picked #110634.
Live repro (fake OpenViking recording the X-API-Key header, alpha profile,
multiplex on): base - WARNING + traceback, no commit; head - no warning,
`POST /api/v1/sessions/<sid>/commit` arrives with alpha's key.
Fixes#110622.
Review follow-up: a sentinel written before the create_time stamp has only
the claim time, but that is epoch seconds too, and the owner was born before
it claimed while a reuser was born after the owner died. One-sided compare
instead of trusting any live PID. Dead monkeypatch line removed from the test.
`detect_unclean_exit` decided "live owner mid-handover" by comparing the
sentinel's `start_time` (the ledger claim, `time.time()` seconds) with
`gateway.status.get_process_start_time(pid)` (proc clock ticks on Linux,
centiseconds elsewhere — its own docstring says it is only comparable with
itself). The two never matched, so every `--replace` takeover whose old owner
was still tearing down read as a crash and was logged/persisted as one.
The sentinel now carries the psutil `create_time` (stamped at claim since
8c3a35b69d), so the guard compares that with the live PID's create time —
same producer, same unit. A pre-stamp sentinel cannot disambiguate PID
reuse; a live PID is taken as the owner, as the psutil-silent case already
was. Test drives all three cases; it fails on the previous ledger.
Found while reviewing the start-attestation follow-ups (#110958).
Review follow-ups on the identity binding:
- The sentinel's `start_time` is `time.time()` at `record_startup`, seconds
after the process was born once imports finish, so comparing it with
psutil's create_time within 2 s would have read every real gateway as
undecidable and silently stopped the #109538 cold-start. `record_startup`
now stamps `create_time` (psutil birth via the existing
`process_identity._process_create_time`), `mark_exited` carries it, and
the attestation compares birth to birth. A sentinel from a gateway older
than the stamp falls back to the PID-only rule.
- A resume token written by pre-generation code and resumed by this code
probes the marker again instead of skipping the spawn.
- Horizon allows a 60 s backwards clock step; the unused `now` parameter is
gone; the create-time tolerance is a named constant; the read-then-unlink
in `_consume_start_attestation` is documented as best-effort.
`_attested_pid_exited_cleanly` matched the lifecycle sentinel by numeric PID
only, so a stale marker for PID 111 flipped from "clean exit" to "crash" once
an unrelated PID 222 lifecycle overwrote the sentinel, and a reused PID's clean
exit could vouch for a different life (#110020 review, gateway_windows.py:937).
`_write_start_attestation` now records `create_times: {pid: create_time}` via
the existing `process_identity._process_create_time`; `mark_exited` carries the
running sentinel's `start_time` onto the exited sentinel; the attested probes
fail closed for a bound PID whenever the sentinel cannot be shown to describe
that incarnation (other PID, start time off by > 2s, or no start time) —
"unknown" never reads as "dead". A missing sentinel still reads as dead, and
markers without `create_times` keep the PID-only rule.
Tests: two attestation tests (stale marker vs. moved-on sentinel → no
authority; own incarnation keeps authority / clean exit / legacy marker) and a
ledger test for the carried `start_time`. Mutation: with HEAD's prod files the
no-authority test and the ledger test fail.
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.
Test proven red on the salvaged head (adapter.sent == [] with `all`).
Review findings (Salt, adversarial pass on the two preceding commits):
- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
counted as an explicit mode because the gate tested key presence, while
the display resolver skips None and inherits. Null resolved to Slack's
tier default `off` and disabled cards, which is the default-off trap the
change exists to avoid. Explicit intent is now a non-None value (or the
env bridge). Tests cover null at each level plus null-over-global-all;
mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
destinations, so the name no longer described the field. Renamed to
`publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
described the opt-in as independent of tool_progress. Rewritten: cards
follow an operator-written off (including /verbose), null inherits, an
un-threaded chat with the card lane active shows no tool progress, other
native failures keep the editable fallback.
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.
Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.
Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.
Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.
`tui_gateway/server_requests.py` owns the one mechanism:
send() block the agent thread until the response frame
(`srq-<n>` ids; ints belong to the client)
send_async() fire-and-callback variant (bot relay)
cancel*() withdraw with ONE `request.cancel {id, method, reason}`
event (timeout / interrupt / process exit /
answered elsewhere) instead of per-kind *.expire
open_requests() the still-open frames, replayed by session.resume,
session.activate and session.events.since so a
reconnecting client re-renders every kind, not two
clarify.lock stays a real client→server RPC (locks one batch
answer early); locked answers merge into the final
set even when the closing response carries only the
tail the user answered last
A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.
Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.
Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
Fold the whitespace-only case into the empty-file miss test and drop
test_overwrite_is_all_or_nothing: it wrote sequentially, so it exercised
nothing about os.replace atomicity that the publish test does not.
CI runs the suite as ~96 concurrent per-file pytest subprocesses and does
not install filelock, so the adapter-guard cache in
tests/gateway/conftest.py runs lock-free there. The verdict was published
with Path.write_text (open-truncate-write): a sibling reading between the
truncate and the write saw an empty file and raised pytest.UsageError(""),
failing a healthy tests/gateway/relay/* file with a bare 'ERROR:' line
(flaky-retry frames in CI runs 32930243538, 32932008269, 32933117564).
Fix the class, not the window:
- _write_guard_cache_atomic: publish via staging file + os.replace so
readers see the old state or the complete verdict, never a torn write.
Staging name (.tmp-*) is invisible to the stale-fingerprint eviction
globs so a sibling's eviction can't unlink it mid-publish.
- _read_guard_cache: empty/vanished cache is a miss (rescan), never a
violation verdict.
Repro: 24 concurrent lock-free subprocesses over 3 relay files fired the
empty-ERROR failure 8x in 8 rounds on origin/main; 0x in 10 rounds after
the fix. Regression tests sabotage-verified.
A /p/<profile>/ cron_job route resolved and fired the job from the gateway's
default home: `_fire_cron_job` ran `execute_job_for_event` outside
`_profile_scope`, so `cron/jobs.json` lookups (`resolve_job_ref`,
`claim_job_for_fire`) and the run's config/secrets came from the wrong
profile — the agent-mode path already scopes its run, this path did not.
`asyncio.to_thread` copies contextvars, so wrapping the call is enough.
Route-level `skills` were also being injected into the rendered prompt on
cron_job routes even though the docs say they are ignored (the job's own
skills apply); skip `_apply_skills` for cron_job routes so the per-run
context stays plain event text.
Test proven red on the pre-fix tree (home resolved to the default profile),
green after.
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.
- gateway/platforms/webhook.py: cron_job route mode — after the same
HMAC auth / rate limit / filters / script / idempotency as agent
routes, the rendered prompt becomes transient per-run context and the
job fires through execute_job_for_event on a worker thread (202
Accepted immediately). Startup validation rejects cron_job +
deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
the shared claimed-run body (_execute_job_now), so event fires share
at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
subscriptions; job ref validated (and canonicalized to the job ID) at
create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
CLI tests in test_webhook_cli.py.
The rebase moved the inbound attachment loop into SlackAdapter._append_link_unfurls,
so the nested-table hunk now lives there and is asserted directly. Drop the source-
provenance references and duplicate cell-level cases; one ragged/malformed-row test
covers raw_text, rich_text, None and unknown cell types.
Port from qwibitai/nanoclaw#3666: Slack represents a pasted table as
'table' blocks — usually nested in attachments[].blocks[], sometimes
top-level. They appear in neither the message text nor the file list,
so the agent received the sentence before the table and nothing else.
- _render_slack_table_block(): projects rows as 'cell | cell' lines,
collecting text leaves from raw_text/rich_text cell subtrees; capped
at 20k chars with a visible '[table truncated]' marker.
- Wired into all three ingestion paths: _extract_text_from_slack_blocks
(thread history + attachment-nested blocks), the live inbound
attachment loop, and _extract_additional_text_from_slack_blocks
(top-level blocks on live messages).
- _serialize_slack_blocks_for_agent skips 'table' blocks — the
allowlist drops 'rows', so it only emitted an empty husk.
Parametrize the reject case over send_video/send_document/send_voice
(each asserts no file upload, base fallback never runs, error names the
file, size and limit), fold the all-oversized batch case into the batch
test, and drop the duplicated fixture setup.
Discord raised the default file upload limit from 10 MiB to 20 MiB for
users, bots, webhooks and interaction responses (developer changelog,
Sep 3 2026). The 25 MiB constant here predates the preflight salvage and
never matched the platform; more importantly discord.py 2.7.1 still
reports 10 MiB via guild.filesize_limit for unboosted guilds, so the
guild-aware path under-reported the cap and rejected 10-20 MiB files
Discord now accepts. Floor the guild value at the platform default so a
stale library constant can only widen, never shrink, the preflight.
The salvaged preflight (#67040) covers _send_file_attachment, but
send_multiple_images opened local files straight into discord.File with
no size check — one oversized image 413'd the whole chunk and dumped its
siblings into the per-image fallback. Preflight each local file against
the same boost-aware limit, skip oversized ones with a user-visible
notice, and still deliver the rest of the chunk.
Sibling-site widening for lobehub-scout salvage of PR #67040 (#50846).
The adapter facade is ~6.6k lines; new behaviour belongs in a topical sibling per the
facade+siblings layout. expand_link_entities() now lives in telegram_entities.py and
reuses the encode/decode UTF-16 slicing the adapter already uses for entity spans.
Also: skip inlining when the anchor text already is the URL (no 'url (url)' duplication),
trim the test file to the invariants and point it at the sibling.
- Keep main's media_write_timeout=60s (PR's HERMES_* env var dropped per
.env-is-secrets-only policy; main already fixed the timeout half).
- Replace stdlib imghdr (removed in Python 3.13) with a magic-byte sniff.
- Exclude GIFs: JPEG conversion flattens animations to one frame.
- Fix transparent-PNG handling: RGBA hit the len(getbands())==4 branch
before the white-background composite, rendering transparency black.
- Clean up temp JPEGs after send (both single and media-group paths);
the docstring promised caller cleanup that neither call site did.
- Write temp files via tempfile default dir instead of an undefined
DEFAULT_OUTPUT_DIR (NameError at runtime in the original PR).
- Add real-Pillow regression tests incl. a sabotage-verified
white-background test.
Twelve credential-driven env branches were routed through _enable_from_env on
main (867e4158f0, #48820), which honors the loader's `_enabled_explicit`
marker. The flag-driven WhatsApp step was the one survivor: `_whatsapp` still
set `wa_cfg.enabled = True` on WHATSAPP_ENABLED=true regardless of an explicit
YAML disable. The dashboard's disable action writes only
`platforms.whatsapp.enabled: false` and leaves the env flag on disk, so the
Baileys bridge reconnected to real contacts on the next full restart
(reported live on #73289).
Route the truthy branch through _enable_from_env like every other platform;
WHATSAPP_ENABLED=false still forces a disable. Register the flag in
_ENV_ENABLE_CREDENTIALS so the one-time explicit-disable WARNING can name it.
Remaining scope of #96557 (the other ~20 sites) landed on main in 867e4158f0
and the config_env.py extraction; this is the delta. Fix direction from
@CryptoDombili in #73303.
Co-authored-by: Professor Dombili <Cryptodombili@gmail.com>
main now passes persist_user_platform_id (and friends) into run_conversation; the fakes
only need the message, so swallow the rest with **_kwargs instead of pinning today's signature.
A message that arrives mid-turn is parked in the adapter's pending slot and drained in-band
by the runner's recursive _run_agent, a path that never touches on_processing_start /
on_processing_complete — so queued, interrupting and steer-demoted messages never got the
read-receipt reaction idle-session messages get, on every adapter implementing the hooks.
Bracket the drain in TurnRunner._run_agent_queued_followup with the hooks, resolving the
adapter from the follow-up's own source (multiplex-safe). Fires only for real inbound
platform events (message_id or raw_message present) and only for adapters that override
on_processing_start, so complete-only adapters (Google Chat, webhook) are not handed an
early completion. Cancels are classified like _process_message_background does.
Re-ported onto current main: the drain moved from gateway/run.py to
gateway/run_turn.py::_run_agent_queued_followup (#102117 decomposition) and the helpers
now live in the topical sibling gateway/run_turn_followup_ack.py.
Original commits (8828a3db37c2, f3c4f7a68c8d) by Mira Solari <268252643+mira-solari@users.noreply.github.com>:
--- 8828a3db37c2 fix(gateway): fire processing hooks for runner-drained queued follow-ups
A message that arrives while a turn is already running never gets the
processing-start acknowledgement — the 👀 read receipt on Slack, and the
equivalent in-progress reaction on Discord, Telegram, Feishu, Matrix,
Signal and Photon. It is not added-then-removed; the hook is never called.
`_run_processing_hook("on_processing_start", …)` has exactly one call site,
inside `BasePlatformAdapter._process_message_background` (base.py:5403).
`handle_message` takes the busy branch at base.py:5165, parks the event and
returns at base.py:5316 — above `_start_session_processing`, which is the
only thing that spawns `_process_message_background`. The parked event is
then drained in-band by the runner (`_dequeue_pending_event`, run.py:23277)
and replayed through a recursive `_run_agent` that touches no adapter hooks.
Because `get_pending_message` pops, the adapter's own drain (base.py:5823)
— the one path that would fire the hook — finds an empty slot. The gap is
structural, not a race, and it is shared by every mode (`queue`,
`interrupt`, steer-demoted-to-queue), by `/queue`, by photo-burst and
text-debounce flushes, and by voice drains.
Fire the existing hook pair around the recursive call. The follow-up now
gets the same lifecycle an idle-session message already gets, on every
platform, through one shared site.
Details that shaped the placement:
- Fired after the depth-cap requeue (run.py:23372) and after every discard
and early return in the block, so no path can strand an in-progress
marker: from that point on, control either reaches the recursion or
raises, and both close the hook.
- Fired before `_refresh_agent_cache_message_count` so that re-baseline
stays adjacent to the recursive call it exists to protect — inserting a
reaction round-trip between them would widen the window in which the
cross-process coherence guard (#45966) can trip on our own writes and
rebuild the agent, destroying the prompt-cache prefix #46237 preserves.
- The hook adapter is resolved from the follow-up's own source, not the
completing turn's: a multiplexed gateway can route it to a different
profile's adapter, and only that instance holds the per-message reaction
state.
- Gated on a truthy `message_id`, which is what every adapter's own hook
already checks. Synthetic drains (`/goal` continuations, wake-ups, CLI
hand-offs) carry no id and stay silent; `interrupt_message` and leftover
`/steer` carry no event at all.
- Cancellation maps to CANCELLED rather than FAILURE, matching
_process_message_background — Telegram clears the marker on CANCELLED and
Signal deliberately leaves it, so the distinction is load-bearing.
Outcome is SUCCESS unless the recursion raises. That mirrors the existing
non-queued contract, where a run returning `failed: True` still delivers a
diagnostic message and reports SUCCESS; making the outcome track agent
failure is a separate change that should apply to both paths at once.
No new config, no new env var, no new hook, no change to message
construction or role alternation.
Fixes#72502
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--- f3c4f7a68c8d fix(gateway): don't hand a completion to complete-only adapters; cover raw-envelope events
Self-review of the previous commit found three defects in it. All three are
about the *guard*, not the mechanism.
1. Complete-only adapters were handed an unpaired completion. Google Chat
(google_chat/adapter.py:2766) and webhook (webhook.py:901) implement
on_processing_complete WITHOUT on_processing_start, and theirs is
end-of-cycle teardown, not a reaction: Google Chat reaps the typing card,
patching it to "(no reply)"/"(interrupted)", and webhook ends the
per-delivery session. The drain fires before the follow-up's reply is
delivered — delivery happens after the whole chain unwinds back into
_process_message_background — so a Google Chat space would get a permanent
"(no reply)" tombstone on every queued follow-up, and the real answer would
then land as a separate message. Now gated on the adapter actually
overriding on_processing_start: we bracket, so both halves must be ours.
2. The message_id gate made the fix a no-op on Signal, which the previous
commit message claimed to fix. SignalAdapter never sets message_id
(signal.py:749-766) — its hook keys off raw_message["sender"] and
["timestamp_ms"] via _extract_reaction_target — and Discord's start hook
reads raw_message too. Gate is now message_id OR raw_message; synthetic
drains still carry neither, so /goal continuations stay silent.
3. Cancellation was classified unconditionally as CANCELLED. base.py:5862-5868
only reports CANCELLED for a task in _expected_cancelled_tasks (/stop, /new,
/reset, adapter cleanup) and downgrades anything else to FAILURE. Signal and
Matrix deliberately LEAVE the marker in place on CANCELLED, so the previous
version would strand exactly the marker this PR exists to clear. Now mirrors
base.py via _followup_cancel_outcome().
Also moves _refresh_agent_cache_message_count inside the try. It awaits DB I/O
and guards it with `except Exception`, which does not catch cancellation, so a
/stop landing there escaped both handlers and stranded the marker. Ordering is
unchanged, so the prompt-cache adjacency the previous commit describes still
holds.
Two new tests, both failing against the previous commit:
test_complete_only_adapter_is_left_alone and
test_raw_envelope_only_followup_is_acknowledged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swallowing FeatureUnavailable returned a bare False, so a hosted operator
saw the same generic "requirements not met / run hermes setup" line the
original report started with. The registry's _probe already logs a raised
exception with its message, so propagating the error is what puts
"target not writable" / "quarantine 404" in gateway.log.
Hosted/Docker images lock /opt/hermes/.venv, so --install-deps writing
site-packages fails with Permission denied and the adapter never starts.
Route Google Chat through lazy_deps (HERMES_LAZY_INSTALL_TARGET) and bake
the extra into the published image so a configured gateway can connect.
Under gateway.multiplex_profiles the routed handler runs inside
_profile_runtime_scope, but the adapter's delivery side
(_process_message_background -> _extract_response_content, weixin's own
send()) extracts and validates the reply's MEDIA: / bare-path attachments
after that scope was reset. Docker translation in platforms/base.py
(_docker_sandbox_dir_candidates via get_active_profile_name,
_parse_docker_volume_mounts via the scope-aware TERMINAL_DOCKER_VOLUMES)
therefore resolved a secondary's /output or /root path against the DEFAULT
profile's sandbox and mounts: dropped as "not found on this host", or a
same-named file from the default's mount delivered instead (#109024).
Add GatewayRunner._media_delivery_scope_for_source (home + terminal policy,
no secret hydration: path validation reads no credentials and runs on the
loop) and enter it from BasePlatformAdapter._media_delivery_scope around
the two extraction sites that run outside the turn scope. The streamed path
(_deliver_media_from_response), the background task and cron delivery
already run inside their profile scope.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.
One request now carries a model pick AND its effort on every surface:
- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
<level>` (validated against `parse_reasoning_effort`; unknown level ->
`MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
`ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
high` applies the effort AFTER the agent swap (`switch_model` re-resolves
`reasoning_config` from config.yaml, so an earlier write is clobbered) with the
pick's scope (session; config on `--global`; `--once` snapshots and restores it).
The `/model` picker gains a third stage, "Reasoning effort for <model>", built
from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
`config.set model "X --reasoning high"` applies after the swap; session pin
(`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
one-turn restore carries `reasoning_config`; re-emits `session_info` so the
status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
`<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
`_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
the Copilot-only inline prompt; Copilot keeps its per-model level set via
`github_model_reasoning_efforts`, other routes get the ladder, catalog
`supports_reasoning=False` skips it) plus a "Reasoning effort for the current
model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
with the same step (+ "Provider default"), stored as
`auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
the task list ("openrouter · model · high"), cleared by "Reset all to auto";
tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
it.
Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
"Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
errors "Model names cannot contain spaces"; after switches and `config.get
reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
"reasoning: high", status bar "fable 5.1 high".
`write_runtime_status` re-stamps the previous writer's `gateway_state.json` in place and
only `_record_served_profiles` (multiplex on) ever wrote `served_profiles`, so a
multiplexer's list survived into a later non-multiplex run of the same home. Every
`hermes -p X` surface then kept treating X as served by that live default gateway: exit
78 on start/install, "running via the default-profile multiplexer" on status (review of
#108352, finding D, second half). The secondary-profile phase now writes an empty list
when multiplexing is off; an empty list is the authoritative "serves nobody else" the
readers already honour.
Both default-profile process matchers (`gateway.status._command_line_belongs_to_profile`
and `hermes_cli.gateway._scan_gateway_pids._matches_current_profile`) rejected a named
gateway with a substring test for `--profile ` / ` -p `, which the equals spelling the
CLI pre-parser accepts (`--profile=ops`) slipped past. The default home's identity check
then adopted that gateway's PID, and a default-profile `gateway stop` with no pid file
scanned the process table and could SIGTERM the named gateway (review of #108352,
finding E). Both sites now ask `profile_flag_value()`, the same tokenizer the named
branch already uses.
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.
One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.