Commit Graph

3468 Commits

Author SHA1 Message Date
kshitijk4poor 935cc237db test(matrix): cover status-only 401 and rate-limit paths in the sync-loop test
Fold the transient/permanent sync-loop tests into one parametrized table
and add the two classifier branches that had no loop-level coverage: a
401 whose body was rewritten to HTML by a reverse proxy (errcode dropped,
so only http_status can stop the loop) and a 429 M_LIMIT_EXCEEDED that
must be retried because neither errcode nor status is an auth signal.
Deleting the http_status fallback in _is_permanent_matrix_auth_error now
fails the 401-html case.

Hoist _sync_error to module level so parametrize can call it directly
instead of the staticmethod.__func__ workaround. Trim the
test_ws_auth_retry docstring, which still described a Matrix test class
that moved to test_matrix.py.
2026-09-15 10:50:45 +05:30
kshitijk4poor 257288ede1 test(matrix): make the 502 SVG fixture actually embed "403"
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
2026-09-15 10:50:45 +05:30
kshitijk4poor d5430fb0c1 refactor(matrix): classify sync auth errors on errcode/http_status only
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.

With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
2026-09-15 10:50:45 +05:30
Stephen Chin 9969a995f6 fix(matrix): correct sync result-object comment and route it through the classifier
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).

The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.

A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.

Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.

Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.

(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
2026-09-15 10:50:45 +05:30
Stephen Chin 6bd9cf8018 fix(matrix): use a genuinely discriminating fixture for the sync-loop test
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.

I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.

I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.

While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.

Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
2026-09-15 10:50:45 +05:30
Stephen Chin f8085f2e03 test(matrix): add loop-level and attribute-narrowing coverage
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.

On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.

(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
2026-09-15 10:50:45 +05:30
Stephen Chin 104889ace6 fix(matrix): tighten sync error classifier
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.

I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.

Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.

(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)
2026-09-15 10:50:45 +05:30
kshitijk4poor 7959e3b0ff test(discord): bind the "0 disables event-silence" half of the knob test
The zero-knob test only asserted that ack_stale still trips with the knob
at 0. Because _read_websocket_health evaluates ack age before the
event-silence dimension, the test stayed green even with the
`_event_max_silence_seconds > 0` guard deleted: it never observed the
stale stamp being ignored. Assert (True, "healthy") with knob=0, a stale
stamp and a green transport first, then make the ACK stale and keep the
ack_stale assertion. Dropping the guard now fails this test.

Also correct the on_socket_event_type comment: discord.py dispatches
socket_event_type before the op-code switch but only for a non-null `t`;
heartbeat ACK frames carry `t: null`, they do not "return before" it.
2026-09-15 10:48:58 +05:30
kshitijk4poor 54041e98cf test(discord): bind event-silence liveness to its config surface and probe gate
The event-silence tests duplicated `_make_adapter`/`_connect` from
test_discord_liveness.py and carried a bespoke handler-wait loop. Reuse
the sibling helpers instead: `_make_adapter` grows an optional
`max_event_silence` that is only written into `extra` when given, so the
sibling's own tests keep the adapter default and stay unchanged.

Two invariants now bind the fix:

- deaf socket: a `None` stamp (no DISPATCH yet) reads healthy across
  several probe intervals, and once armed, transport-green + event
  silence trips `event_silence` through `_liveness_loop`.
- `websocket_event_max_silence_seconds: 0` disables only that
  dimension: with a stale stamp and a stale heartbeat ACK the probe must
  still run and trip on `ack_stale`. This goes red if the knob is moved
  into `_start_liveness_probe`'s all-or-nothing guard (#109782).

The YAML->extra passthrough test also asserts the new key, so dropping
its `_YAML_WEBSOCKET_LIVENESS_KEYS` entry fails.
2026-09-15 10:48:58 +05:30
kshitijk4poor fff69b7dfc test(discord): keep the two binding event-silence invariants
Trim the new event-silence file from 8 tests to the 2 that bind the fix:
the deaf-socket e2e through `_liveness_loop` (transport green, no
DISPATCH → probe trips with the `event_silence` reason) and the
None-stamp window reading healthy (no false trip on quiet reconnects).
The removed cases re-checked knob parsing, defaults and the YAML seed
loop already covered by the sibling liveness knobs' tests, or restated
the two kept invariants from a different angle.
2026-09-15 10:48:58 +05:30
salch-cred 101861c7f7 fix(discord): dispatch-side liveness dimension detects an ACKing-but-deaf gateway socket (#109521)
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.

This adds the dispatch-side signal that was requested instead:

- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
  every parsed DISPATCH frame with no debug gate (verified live against
  the real received_message path: 4/4 frames fired with
  enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
  incident report's field-proven operator bound); 0 opts out of this
  dimension ONLY — the #109782 review failure put the knob in
  _start_liveness_probe's all-or-nothing guard, killing the whole
  watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
  stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out

Fixes #109521

(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
2026-09-15 10:48:58 +05:30
teknium1 a21747fe4e fix(kanban): an anchorless thread subscription warns once instead of vanishing (#110919)
Follow-up on the salvaged #110928 (CLI `--parent-chat-id` / `--guild-id`):

- `_claim_for_sub` skipped a thread-shaped row that matched no `profile_routes`
  entry at DEBUG on every tick. Legacy rows written before the flags existed can
  never match a channel-level route (no `parent_chat_id`), so the notifier now
  logs ONE WARNING per row naming the task, the thread and the re-subscribe
  command. Still fail-closed: the events stay unclaimed.
- Docs: the kanban user guide explains the anchors and shows the Discord-thread
  subscribe command under `profile_routes`.
- Test (red on origin/main): two collects → exactly one WARNING, events unseen.
2026-09-14 16:14:33 -07:00
teknium1 90a65d6e38 fix(gateway): gateway-stop teardown fires memory-provider lifecycle hooks under the owning profile
Under `gateway.multiplex_profiles: true` every gateway-hosted session that ended
inside the gateway process logged `Memory provider 'openviking' on_session_end
failed: get_secret('OPENVIKING_API_KEY') called with no profile secret scope
active` and skipped its end-of-session commit. The provider lifecycle hooks
(`flush_pending` -> `on_session_end` -> provider teardown -> `close`) read
credentials and home at call time; the stop-time finalize pass and the
idle-cache sweep run on the main loop outside any adapter handler, so
`_run_in_executor_with_context` copied an EMPTY scope into the worker.

Route `_cleanup_agent_resources_off_loop` and `_finalize_session_off_loop`
through `_run_release_in_profile_scope` (the seam cache eviction already uses):
a scoped caller keeps its scope; an unscoped caller passes the session key and
the owner's profile scope is entered from it. The stop path passes the key for
both active and idle-cached agents. Cron's post-run cleanup thread is the
sibling seam and is fixed by the cherry-picked #110634.

Live repro (fake OpenViking recording the X-API-Key header, alpha profile,
multiplex on): base - WARNING + traceback, no commit; head - no warning,
`POST /api/v1/sessions/<sid>/commit` arrives with alpha's key.

Fixes #110622.
2026-09-14 16:13:51 -07:00
kshitijk4poor e6e4213dbf fix(lifecycle-ledger): pre-stamp sentinels still tell an owner from a PID reuser
Review follow-up: a sentinel written before the create_time stamp has only
the claim time, but that is epoch seconds too, and the owner was born before
it claimed while a reuser was born after the owner died. One-sided compare
instead of trusting any live PID. Dead monkeypatch line removed from the test.
2026-09-14 21:12:32 +05:30
kshitijk4poor 40087fba7d fix(lifecycle-ledger): a --replace handover is no longer reported as an unclean death
`detect_unclean_exit` decided "live owner mid-handover" by comparing the
sentinel's `start_time` (the ledger claim, `time.time()` seconds) with
`gateway.status.get_process_start_time(pid)` (proc clock ticks on Linux,
centiseconds elsewhere — its own docstring says it is only comparable with
itself). The two never matched, so every `--replace` takeover whose old owner
was still tearing down read as a crash and was logged/persisted as one.

The sentinel now carries the psutil `create_time` (stamped at claim since
8c3a35b69d), so the guard compares that with the live PID's create time —
same producer, same unit. A pre-stamp sentinel cannot disambiguate PID
reuse; a live PID is taken as the owner, as the psutil-silent case already
was. Test drives all three cases; it fails on the previous ledger.

Found while reviewing the start-attestation follow-ups (#110958).
2026-09-14 21:12:32 +05:30
kshitijk4poor 8c3a35b69d fix(gateway-windows): bind the attestation to process birth, not the ledger claim
Review follow-ups on the identity binding:
- The sentinel's `start_time` is `time.time()` at `record_startup`, seconds
  after the process was born once imports finish, so comparing it with
  psutil's create_time within 2 s would have read every real gateway as
  undecidable and silently stopped the #109538 cold-start. `record_startup`
  now stamps `create_time` (psutil birth via the existing
  `process_identity._process_create_time`), `mark_exited` carries it, and
  the attestation compares birth to birth. A sentinel from a gateway older
  than the stamp falls back to the PID-only rule.
- A resume token written by pre-generation code and resumed by this code
  probes the marker again instead of skipping the spawn.
- Horizon allows a 60 s backwards clock step; the unused `now` parameter is
  gone; the create-time tolerance is a named constant; the read-then-unlink
  in `_consume_start_attestation` is documented as best-effort.
2026-09-14 20:51:37 +05:30
kshitijk4poor faf6e6889b fix(gateway-windows): bind the start attestation to each PID's process create time
`_attested_pid_exited_cleanly` matched the lifecycle sentinel by numeric PID
only, so a stale marker for PID 111 flipped from "clean exit" to "crash" once
an unrelated PID 222 lifecycle overwrote the sentinel, and a reused PID's clean
exit could vouch for a different life (#110020 review, gateway_windows.py:937).

`_write_start_attestation` now records `create_times: {pid: create_time}` via
the existing `process_identity._process_create_time`; `mark_exited` carries the
running sentinel's `start_time` onto the exited sentinel; the attested probes
fail closed for a bound PID whenever the sentinel cannot be shown to describe
that incarnation (other PID, start time off by > 2s, or no start time) —
"unknown" never reads as "dead". A missing sentinel still reads as dead, and
markers without `create_times` keep the PID-only rule.

Tests: two attestation tests (stale marker vs. moved-on sentinel → no
authority; own incarnation keeps authority / clean exit / legacy marker) and a
ledger test for the carried `start_time`. Mutation: with HEAD's prod files the
no-authority test and the ledger test fail.
2026-09-14 20:51:37 +05:30
teknium1 4748caff76 fix(gateway): explicit tool_progress new/all keeps text progress in un-cardable Slack chats
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.

Test proven red on the salvaged head (adapter.sent == [] with `all`).
2026-09-14 07:46:51 -07:00
Victor Kyriazakos a4e2a82a6d fix(gateway): preflight task-card destination before transport fallback 2026-09-14 07:46:51 -07:00
Victor Kyriazakos dab7eebf47 fix(gateway): resolve progress mode and intent from the same source 2026-09-14 07:46:51 -07:00
Victor Kyriazakos bc125d59d5 fix(gateway): null tool_progress inherits; name the task-card suppression latch
Review findings (Salt, adversarial pass on the two preceding commits):

- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
  counted as an explicit mode because the gate tested key presence, while
  the display resolver skips None and inherits. Null resolved to Slack's
  tier default `off` and disabled cards, which is the default-off trap the
  change exists to avoid. Explicit intent is now a non-None value (or the
  env bridge). Tests cover null at each level plus null-over-global-all;
  mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
  destinations, so the name no longer described the field. Renamed to
  `publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
  described the opt-in as independent of tool_progress. Rewritten: cards
  follow an operator-written off (including /verbose), null inherits, an
  un-threaded chat with the card lane active shows no tool progress, other
  native failures keep the editable fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos 3412490ad1 fix(gateway): no text tool progress when a Slack chat cannot host a task card
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.

Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos ed25a40917 fix(gateway): explicit tool_progress off disables Slack task cards
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.

Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.

Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
2026-09-14 07:46:51 -07:00
teknium1 9f7f2f28c0 feat(gateway): server→client JSON-RPC requests replace the *.request/*.respond event pairs (#110521)
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.

`tui_gateway/server_requests.py` owns the one mechanism:

  send()          block the agent thread until the response frame
                  (`srq-<n>` ids; ints belong to the client)
  send_async()    fire-and-callback variant (bot relay)
  cancel*()       withdraw with ONE `request.cancel {id, method, reason}`
                  event (timeout / interrupt / process exit /
                  answered elsewhere) instead of per-kind *.expire
  open_requests() the still-open frames, replayed by session.resume,
                  session.activate and session.events.since so a
                  reconnecting client re-renders every kind, not two
  clarify.lock    stays a real client→server RPC (locks one batch
                  answer early); locked answers merge into the final
                  set even when the closing response carries only the
                  tail the user answered last

A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.

Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.

Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
2026-09-14 06:02:05 -07:00
teknium1 2178b3ebe8 test: trim guard-cache race tests to the invariants
Fold the whitespace-only case into the empty-file miss test and drop
test_overwrite_is_all_or_nothing: it wrote sequentially, so it exercised
nothing about os.replace atomicity that the publish test does not.
2026-09-13 21:31:18 -07:00
Teknium 8c71fd78d5 fix(tests): gateway adapter-guard cache no longer fails healthy files with a blank ERROR under the parallel runner
CI runs the suite as ~96 concurrent per-file pytest subprocesses and does
not install filelock, so the adapter-guard cache in
tests/gateway/conftest.py runs lock-free there. The verdict was published
with Path.write_text (open-truncate-write): a sibling reading between the
truncate and the write saw an empty file and raised pytest.UsageError(""),
failing a healthy tests/gateway/relay/* file with a bare 'ERROR:' line
(flaky-retry frames in CI runs 32930243538, 32932008269, 32933117564).

Fix the class, not the window:
- _write_guard_cache_atomic: publish via staging file + os.replace so
  readers see the old state or the complete verdict, never a torn write.
  Staging name (.tmp-*) is invisible to the stale-fingerprint eviction
  globs so a sibling's eviction can't unlink it mid-publish.
- _read_guard_cache: empty/vanished cache is a miss (rescan), never a
  violation verdict.

Repro: 24 concurrent lock-free subprocesses over 3 relay files fired the
empty-ERROR failure 8x in 8 rounds on origin/main; 0x in 10 rounds after
the fix. Regression tests sabotage-verified.
2026-09-13 21:31:18 -07:00
teknium1 ee51e8bf8b chore(webhook): drop external-product attribution from code, tests and docs 2026-09-13 21:30:02 -07:00
teknium1 90c3131246 fix(webhook): run cron_job triggers under the routed profile's scope; skip route skills
A /p/<profile>/ cron_job route resolved and fired the job from the gateway's
default home: `_fire_cron_job` ran `execute_job_for_event` outside
`_profile_scope`, so `cron/jobs.json` lookups (`resolve_job_ref`,
`claim_job_for_fire`) and the run's config/secrets came from the wrong
profile — the agent-mode path already scopes its run, this path did not.
`asyncio.to_thread` copies contextvars, so wrapping the call is enough.

Route-level `skills` were also being injected into the rendered prompt on
cron_job routes even though the docs say they are ignored (the job's own
skills apply); skip `_apply_skills` for cron_job routes so the per-run
context stays plain event text.

Test proven red on the pre-fix tree (home resolved to the default profile),
green after.
2026-09-13 21:30:02 -07:00
Teknium 45ab5e3b5d Inspired by ChatGPT Work: event-triggered cron jobs via webhook routes
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.

- gateway/platforms/webhook.py: cron_job route mode — after the same
  HMAC auth / rate limit / filters / script / idempotency as agent
  routes, the rendered prompt becomes transient per-run context and the
  job fires through execute_job_for_event on a worker thread (202
  Accepted immediately). Startup validation rejects cron_job +
  deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
  the shared claimed-run body (_execute_job_now), so event fires share
  at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
  handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
  subscriptions; job ref validated (and canonicalized to the job ID) at
  create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
  cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
  CLI tests in test_webhook_cli.py.
2026-09-13 21:30:02 -07:00
teknium1 aea84cbb1a test(slack): collapse pasted-table tests to five invariants, cover live unfurl path
The rebase moved the inbound attachment loop into SlackAdapter._append_link_unfurls,
so the nested-table hunk now lives there and is asserted directly. Drop the source-
provenance references and duplicate cell-level cases; one ragged/malformed-row test
covers raw_text, rich_text, None and unknown cell types.
2026-09-13 20:59:55 -07:00
Teknium a74e0155b6 feat(slack): pasted tables now reach the agent instead of silently vanishing
Port from qwibitai/nanoclaw#3666: Slack represents a pasted table as
'table' blocks — usually nested in attachments[].blocks[], sometimes
top-level. They appear in neither the message text nor the file list,
so the agent received the sentence before the table and nothing else.

- _render_slack_table_block(): projects rows as 'cell | cell' lines,
  collecting text leaves from raw_text/rich_text cell subtrees; capped
  at 20k chars with a visible '[table truncated]' marker.
- Wired into all three ingestion paths: _extract_text_from_slack_blocks
  (thread history + attachment-nested blocks), the live inbound
  attachment loop, and _extract_additional_text_from_slack_blocks
  (top-level blocks on live messages).
- _serialize_slack_blocks_for_agent skips 'table' blocks — the
  allowlist drops 'rows', so it only emitted an empty husk.
2026-09-13 20:59:55 -07:00
teknium1 a2b1c4cf44 test(discord): collapse the preflight tests to four invariants
Parametrize the reject case over send_video/send_document/send_voice
(each asserts no file upload, base fallback never runs, error names the
file, size and limit), fold the all-oversized batch case into the batch
test, and drop the duplicated fixture setup.
2026-09-13 20:59:17 -07:00
Teknium 6019a39efa fix(discord): raise default upload preflight to 20 MiB (Sep 3 2026 API change)
Discord raised the default file upload limit from 10 MiB to 20 MiB for
users, bots, webhooks and interaction responses (developer changelog,
Sep 3 2026). The 25 MiB constant here predates the preflight salvage and
never matched the platform; more importantly discord.py 2.7.1 still
reports 10 MiB via guild.filesize_limit for unboosted guilds, so the
guild-aware path under-reported the cap and rejected 10-20 MiB files
Discord now accepts. Floor the guild value at the platform default so a
stale library constant can only widen, never shrink, the preflight.
2026-09-13 20:59:17 -07:00
Teknium 20a7a274d1 fix(discord): widen upload-size preflight to batch image sends
The salvaged preflight (#67040) covers _send_file_attachment, but
send_multiple_images opened local files straight into discord.File with
no size check — one oversized image 413'd the whole chunk and dumped its
siblings into the per-image fallback. Preflight each local file against
the same boost-aware limit, skip oversized ones with a user-visible
notice, and still deliver the rest of the chunk.

Sibling-site widening for lobehub-scout salvage of PR #67040 (#50846).
2026-09-13 20:59:17 -07:00
Slobaka e95f4fcf2f fix(discord): preflight attachment size before upload
Reject oversized local attachments before channel.send(file=...) so
users get an explicit notice instead of a doomed 413 round-trip.

Fixes #50846
2026-09-13 20:59:17 -07:00
teknium1 d7ceee19a8 refactor(telegram): move text_link expansion to telegram_entities sibling
The adapter facade is ~6.6k lines; new behaviour belongs in a topical sibling per the
facade+siblings layout. expand_link_entities() now lives in telegram_entities.py and
reuses the encode/decode UTF-16 slicing the adapter already uses for entity spans.

Also: skip inlining when the anchor text already is the URL (no 'url (url)' duplication),
trim the test file to the invariants and point it at the sibling.
2026-09-13 20:58:39 -07:00
WS f6cfbd2b16 fix(telegram): expose hidden text-link URLs 2026-09-13 20:58:39 -07:00
Teknium 333bf120a5 fix(telegram): harden image pre-compression for salvage of #74893
- Keep main's media_write_timeout=60s (PR's HERMES_* env var dropped per
  .env-is-secrets-only policy; main already fixed the timeout half).
- Replace stdlib imghdr (removed in Python 3.13) with a magic-byte sniff.
- Exclude GIFs: JPEG conversion flattens animations to one frame.
- Fix transparent-PNG handling: RGBA hit the len(getbands())==4 branch
  before the white-background composite, rendering transparency black.
- Clean up temp JPEGs after send (both single and media-group paths);
  the docstring promised caller cleanup that neither call site did.
- Write temp files via tempfile default dir instead of an undefined
  DEFAULT_OUTPUT_DIR (NameError at runtime in the original PR).
- Add real-Pillow regression tests incl. a sabotage-verified
  white-background test.
2026-09-13 20:57:40 -07:00
Professor Dombili 73a9c34529 fix(gateway): WHATSAPP_ENABLED=true no longer overrides an explicit platforms.whatsapp.enabled: false (#73289, salvage #73303)
Twelve credential-driven env branches were routed through _enable_from_env on
main (867e4158f0, #48820), which honors the loader's `_enabled_explicit`
marker. The flag-driven WhatsApp step was the one survivor: `_whatsapp` still
set `wa_cfg.enabled = True` on WHATSAPP_ENABLED=true regardless of an explicit
YAML disable. The dashboard's disable action writes only
`platforms.whatsapp.enabled: false` and leaves the env flag on disk, so the
Baileys bridge reconnected to real contacts on the next full restart
(reported live on #73289).

Route the truthy branch through _enable_from_env like every other platform;
WHATSAPP_ENABLED=false still forces a disable. Register the flag in
_ENV_ENABLE_CREDENTIALS so the one-time explicit-disable WARNING can name it.

Remaining scope of #96557 (the other ~20 sites) landed on main in 867e4158f0
and the config_env.py extraction; this is the delta. Fix direction from
@CryptoDombili in #73303.

Co-authored-by: Professor Dombili <Cryptodombili@gmail.com>
2026-09-13 20:53:54 -07:00
teknium1 876bdc857d test: accept the widened run_conversation kwargs in the queued-followup agent doubles
main now passes persist_user_platform_id (and friends) into run_conversation; the fakes
only need the message, so swallow the rest with **_kwargs instead of pinning today's signature.
2026-09-13 20:52:02 -07:00
Mira Solari 606ea7f4ff fix(gateway): fire processing hooks for runner-drained queued follow-ups (#72502, salvage #72503)
A message that arrives mid-turn is parked in the adapter's pending slot and drained in-band
by the runner's recursive _run_agent, a path that never touches on_processing_start /
on_processing_complete — so queued, interrupting and steer-demoted messages never got the
read-receipt reaction idle-session messages get, on every adapter implementing the hooks.

Bracket the drain in TurnRunner._run_agent_queued_followup with the hooks, resolving the
adapter from the follow-up's own source (multiplex-safe). Fires only for real inbound
platform events (message_id or raw_message present) and only for adapters that override
on_processing_start, so complete-only adapters (Google Chat, webhook) are not handed an
early completion. Cancels are classified like _process_message_background does.

Re-ported onto current main: the drain moved from gateway/run.py to
gateway/run_turn.py::_run_agent_queued_followup (#102117 decomposition) and the helpers
now live in the topical sibling gateway/run_turn_followup_ack.py.

Original commits (8828a3db37c2, f3c4f7a68c8d) by Mira Solari <268252643+mira-solari@users.noreply.github.com>:
--- 8828a3db37c2 fix(gateway): fire processing hooks for runner-drained queued follow-ups
A message that arrives while a turn is already running never gets the
processing-start acknowledgement — the 👀 read receipt on Slack, and the
equivalent in-progress reaction on Discord, Telegram, Feishu, Matrix,
Signal and Photon. It is not added-then-removed; the hook is never called.

`_run_processing_hook("on_processing_start", …)` has exactly one call site,
inside `BasePlatformAdapter._process_message_background` (base.py:5403).
`handle_message` takes the busy branch at base.py:5165, parks the event and
returns at base.py:5316 — above `_start_session_processing`, which is the
only thing that spawns `_process_message_background`. The parked event is
then drained in-band by the runner (`_dequeue_pending_event`, run.py:23277)
and replayed through a recursive `_run_agent` that touches no adapter hooks.
Because `get_pending_message` pops, the adapter's own drain (base.py:5823)
— the one path that would fire the hook — finds an empty slot. The gap is
structural, not a race, and it is shared by every mode (`queue`,
`interrupt`, steer-demoted-to-queue), by `/queue`, by photo-burst and
text-debounce flushes, and by voice drains.

Fire the existing hook pair around the recursive call. The follow-up now
gets the same lifecycle an idle-session message already gets, on every
platform, through one shared site.

Details that shaped the placement:

- Fired after the depth-cap requeue (run.py:23372) and after every discard
  and early return in the block, so no path can strand an in-progress
  marker: from that point on, control either reaches the recursion or
  raises, and both close the hook.
- Fired before `_refresh_agent_cache_message_count` so that re-baseline
  stays adjacent to the recursive call it exists to protect — inserting a
  reaction round-trip between them would widen the window in which the
  cross-process coherence guard (#45966) can trip on our own writes and
  rebuild the agent, destroying the prompt-cache prefix #46237 preserves.
- The hook adapter is resolved from the follow-up's own source, not the
  completing turn's: a multiplexed gateway can route it to a different
  profile's adapter, and only that instance holds the per-message reaction
  state.
- Gated on a truthy `message_id`, which is what every adapter's own hook
  already checks. Synthetic drains (`/goal` continuations, wake-ups, CLI
  hand-offs) carry no id and stay silent; `interrupt_message` and leftover
  `/steer` carry no event at all.
- Cancellation maps to CANCELLED rather than FAILURE, matching
  _process_message_background — Telegram clears the marker on CANCELLED and
  Signal deliberately leaves it, so the distinction is load-bearing.

Outcome is SUCCESS unless the recursion raises. That mirrors the existing
non-queued contract, where a run returning `failed: True` still delivers a
diagnostic message and reports SUCCESS; making the outcome track agent
failure is a separate change that should apply to both paths at once.

No new config, no new env var, no new hook, no change to message
construction or role alternation.

Fixes #72502

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

--- f3c4f7a68c8d fix(gateway): don't hand a completion to complete-only adapters; cover raw-envelope events
Self-review of the previous commit found three defects in it. All three are
about the *guard*, not the mechanism.

1. Complete-only adapters were handed an unpaired completion. Google Chat
   (google_chat/adapter.py:2766) and webhook (webhook.py:901) implement
   on_processing_complete WITHOUT on_processing_start, and theirs is
   end-of-cycle teardown, not a reaction: Google Chat reaps the typing card,
   patching it to "(no reply)"/"(interrupted)", and webhook ends the
   per-delivery session. The drain fires before the follow-up's reply is
   delivered — delivery happens after the whole chain unwinds back into
   _process_message_background — so a Google Chat space would get a permanent
   "(no reply)" tombstone on every queued follow-up, and the real answer would
   then land as a separate message. Now gated on the adapter actually
   overriding on_processing_start: we bracket, so both halves must be ours.

2. The message_id gate made the fix a no-op on Signal, which the previous
   commit message claimed to fix. SignalAdapter never sets message_id
   (signal.py:749-766) — its hook keys off raw_message["sender"] and
   ["timestamp_ms"] via _extract_reaction_target — and Discord's start hook
   reads raw_message too. Gate is now message_id OR raw_message; synthetic
   drains still carry neither, so /goal continuations stay silent.

3. Cancellation was classified unconditionally as CANCELLED. base.py:5862-5868
   only reports CANCELLED for a task in _expected_cancelled_tasks (/stop, /new,
   /reset, adapter cleanup) and downgrades anything else to FAILURE. Signal and
   Matrix deliberately LEAVE the marker in place on CANCELLED, so the previous
   version would strand exactly the marker this PR exists to clear. Now mirrors
   base.py via _followup_cancel_outcome().

Also moves _refresh_agent_cache_message_count inside the try. It awaits DB I/O
and guards it with `except Exception`, which does not catch cancellation, so a
/stop landing there escaped both handlers and stranded the marker. Ordering is
unchanged, so the prompt-cache adjacency the previous commit describes still
holds.

Two new tests, both failing against the previous commit:
test_complete_only_adapter_is_left_alone and
test_raw_envelope_only_followup_is_acknowledged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 20:52:02 -07:00
teknium1 afe06f21f4 fix(google-chat): let ensure_deps_fn surface the lazy-install reason
Swallowing FeatureUnavailable returned a bare False, so a hosted operator
saw the same generic "requirements not met / run hermes setup" line the
original report started with. The registry's _probe already logs a raised
exception with its message, so propagating the error is what puts
"target not writable" / "quarantine 404" in gateway.log.
2026-09-13 19:19:36 -07:00
xxxigm 1468e98e48 test(google-chat): pin hosted install wiring and document the lazy target 2026-09-13 19:19:36 -07:00
xxxigm f3dbb19707 fix(google-chat): install optional deps into the sealed-image lazy target
Hosted/Docker images lock /opt/hermes/.venv, so --install-deps writing
site-packages fails with Permission denied and the adapter never starts.
Route Google Chat through lazy_deps (HERMES_LAZY_INSTALL_TARGET) and bake
the extra into the published image so a configured gateway can connect.
2026-09-13 19:19:36 -07:00
teknium1 28138f2524 fix(gateway): validate a secondary's Docker MEDIA paths under its own profile scope
Under gateway.multiplex_profiles the routed handler runs inside
_profile_runtime_scope, but the adapter's delivery side
(_process_message_background -> _extract_response_content, weixin's own
send()) extracts and validates the reply's MEDIA: / bare-path attachments
after that scope was reset. Docker translation in platforms/base.py
(_docker_sandbox_dir_candidates via get_active_profile_name,
_parse_docker_volume_mounts via the scope-aware TERMINAL_DOCKER_VOLUMES)
therefore resolved a secondary's /output or /root path against the DEFAULT
profile's sandbox and mounts: dropped as "not found on this host", or a
same-named file from the default's mount delivered instead (#109024).

Add GatewayRunner._media_delivery_scope_for_source (home + terminal policy,
no secret hydration: path validation reads no credentials and runs on the
loop) and enter it from BasePlatformAdapter._media_delivery_scope around
the two extraction sites that run outside the turn scope. The streamed path
(_deliver_media_from_response), the background task and cron delivery
already run inside their profile scope.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-13 19:18:26 -07:00
teknium1 2c0bec33f9 feat(model-pickers): reasoning effort selection on every model picker
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.

One request now carries a model pick AND its effort on every surface:

- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
  <level>` (validated against `parse_reasoning_effort`; unknown level ->
  `MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
  `ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
  high` applies the effort AFTER the agent swap (`switch_model` re-resolves
  `reasoning_config` from config.yaml, so an earlier write is clobbered) with the
  pick's scope (session; config on `--global`; `--once` snapshots and restores it).
  The `/model` picker gains a third stage, "Reasoning effort for <model>", built
  from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
  inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
  `config.set model "X --reasoning high"` applies after the swap; session pin
  (`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
  one-turn restore carries `reasoning_config`; re-emits `session_info` so the
  status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
  `<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
  label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
  `_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
  the Copilot-only inline prompt; Copilot keeps its per-model level set via
  `github_model_reasoning_efforts`, other routes get the ladder, catalog
  `supports_reasoning=False` skips it) plus a "Reasoning effort for the current
  model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
  with the same step (+ "Provider default"), stored as
  `auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
  the task list ("openrouter · model · high"), cleared by "Reset all to auto";
  tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
  it.

Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
  "Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
  and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
  writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
  errors "Model names cannot contain spaces"; after switches and `config.get
  reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
  effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
  "reasoning: high", status bar "fable 5.1 high".
2026-09-13 16:43:50 -07:00
luinbytes ed6f19f14f fix(gateway): format scoped MCP server names during reload 2026-09-13 15:41:01 -07:00
teknium1 aed011d88a fix(gateway): a single-profile gateway start clears an inherited served_profiles list
`write_runtime_status` re-stamps the previous writer's `gateway_state.json` in place and
only `_record_served_profiles` (multiplex on) ever wrote `served_profiles`, so a
multiplexer's list survived into a later non-multiplex run of the same home. Every
`hermes -p X` surface then kept treating X as served by that live default gateway: exit
78 on start/install, "running via the default-profile multiplexer" on status (review of
#108352, finding D, second half). The secondary-profile phase now writes an empty list
when multiplexing is off; an empty list is the authoritative "serves nobody else" the
readers already honour.
2026-09-13 15:41:01 -07:00
teknium1 e03d3d00bf fix(gateway): --profile=ops gateway is never matched as the default profile's
Both default-profile process matchers (`gateway.status._command_line_belongs_to_profile`
and `hermes_cli.gateway._scan_gateway_pids._matches_current_profile`) rejected a named
gateway with a substring test for `--profile ` / ` -p `, which the equals spelling the
CLI pre-parser accepts (`--profile=ops`) slipped past. The default home's identity check
then adopted that gateway's PID, and a default-profile `gateway stop` with no pid file
scanned the process table and could SIGTERM the named gateway (review of #108352,
finding E). Both sites now ask `profile_flag_value()`, the same tokenizer the named
branch already uses.
2026-09-13 15:41:01 -07:00
teknium1 9188e708b3 fix(gateway): served_profiles bind to a verified gateway identity, not bare PID existence
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.

One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
2026-09-13 15:41:01 -07:00