Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
Review found _create_entry_from_recovered_row builds a minimal entry:
replacing the live object would silently drop model_override, token/cost
counters, resume_pending/queued-work markers, and metadata. Keep the
original entry (routing is unchanged, so no sessions.json rewrite either)
and log at INFO — this is a success path, not a corrective action.
Regression test now asserts state preservation, reopen_session call, and
no save.
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.
`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:
1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
`_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
used by `systemd_restart()` and the updater — had no launchd call site.
2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
systemd branch uses `_get_restart_exit_wait_budget()`
(drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
documents that callers falling back to a hard kill must cover both phases
or they reintroduce #77184.
The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.
Observed on macOS 27.0 / Hermes 0.20.4:
→ Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
⚠ Gateway PID 49787 still running after 0.0s — restart may fail
⚠ Gateway drain timed out after 0s — forcing launchd restart
Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.
The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.
Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.
Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.
Refs: #79096
- The direct adapter.send() confirmation path in /approve and /deny is only
needed on native-streaming platforms (WeCom) where the reply stream is
already finalized; other platforms keep the return-text contract (fixes
4 approve/deny regression tests, guarded with 'is not True' against
MagicMock auto-attributes).
- Remove a stray [DEBUG] logger.info left in _deliver_media_from_response.
- website/docs wecom.md: replace the 'does not stream' notes with the native
msgtype:stream behavior and document the stream keepalive extra keys.
Fixes two related native-streaming bubble defects surfaced in production:
1. Duplicate bubble on long turns — when keep-alive already refreshed the
6-min reply window, the Layer-2 clock fallback still declined the finalize
frame and forced a proactive send(), duplicating the message. Skip the
clock fallback while keep-alive is active; intermediate-frame failures are
now fully fire-and-forget (only a failed FINAL frame falls back to send()).
2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
notice) called _reset_segment_state(), clearing the cumulative
_accumulated so the next delta + finalize frame carried only the few
chars accumulated after the reset. Native streaming now skips that
reset (commentary still posts as its own message via send()).
b) The adapter-side _BlockChunker.update() 'only grow' guard silently
dropped any cumulative snapshot shorter than its high-water mark, so
after a baseline reset the leading characters were stranded before
_emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
layer entirely; intermediate frames are pure identity-dedup, matching
the fire-and-forget model.
Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.
Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.
Includes the machinery intrinsic to native streaming on WeCom:
- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
frames update it, a finalize frame closes it; native-streaming adapters are
let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
long-connection mode has no documented edit-rate limit); an adapter-level
frame cap is retained. (Early builds gated frames behind a char throttle;
removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
(control vs normal) plus a per-chat token bucket to stay under WeCom's rate
limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
delivered the moment it is emitted; failures logged, not re-sent; delivery
marked once per turn), per-req_id reply queue with ack tracking, and the
timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
the prompt is the last thing on screen and never traps a lingering bubble;
eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
messages; image+text double-callback merged into one turn.
Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.
When send_message is invoked from the agent's worker thread (a different
event loop than the gateway's), awaiting the WeCom adapter directly can hang
because the adapter enqueues onto the gateway loop. Dispatch via
run_coroutine_threadsafe onto the gateway loop when the caller loop differs,
with caller-cancellation shielded so an already-enqueued send is not cancelled
mid-flight (which would otherwise cause a false-failure retry -> duplicate).
Recognizes WeCom native chat IDs as explicit send targets and whitelists WeCom
for media delivery. Part of the async queue design this branch introduces.
Add WeCom to the per-platform streaming defaults (DEFAULT_CONFIG display
plumbing) so native streaming is enabled by default for the WeCom adapter,
alongside the existing per-platform flags. Non-secret config lives in
config.yaml (no HERMES_* env vars).
Per the 'when in doubt, optional' rule — site publishing is an
on-request capability, not a weekly daily-driver for most users.
Joins cloudflare-temporary-deploy/page-agent under
optional-skills/web-development (existing category, existing
DESCRIPTION.md kept; the new bundled category dir is dropped).
Install via: hermes skills install official/web-development/publish-site
Deduplicate the explanation across the constant comment, count_empty_sessions
docstring, and delete_empty_sessions docstring. Keep the incident refs
(#70516/#80763/#82756/#95868) and the core WHY on the constant; cross-reference
from the methods.
Precision pass on the comments added by the previous commit: prompt.submit
reaches replace_messages(archive_dropped=True) with an empty prefix on a
confirmed ordinal-0 rewind, which is the production shape that lands a
populated session on message_count = 0. archive_and_compact normally
publishes at least a summary row, so it is pinned as defense in depth
rather than claimed as an equally reachable trigger.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:
* `replace_messages(..., archive_dropped=True)` — the rewind / edit /
regenerate mode added in #82756 so a taken-back turn stays recoverable.
* `archive_and_compact` — in-place compaction, which archives the
pre-compaction transcript under the same session id (#38763).
A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.
Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.
Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.
Fixes#95868
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
Replaces the stale MCP-first AgentMail skill with CLI-first guidance:
self-signup + OTP verification, inbox/message/thread/label/attachment
flows, webhook and WebSocket delivery references, and MCP as an
alternative path. Declares AGENTMAIL_API_KEY (optional) so a stored key
reaches the sandboxed terminal while self-signup stays viable without
one.
Salvaged from PR #60811 — kept in optional-skills/ per the March 2026
decision that third-party-API-key skills are not bundled.
Review follow-ups on the #96290 salvage:
- The inline env->config->default ladder was the third copy of the pattern;
extract it next to _get_script_timeout/_get_media_send_timeout. Using
load_config() (deep-merge) also removes the default-drift hazard flagged
in review: cron.session_db_timeout_seconds now resolves from
DEFAULT_CONFIG (config_defaults.py) instead of relying on the hardcoded
10.0 staying in sync with it, and drops the distant-state coupling to
run_job's raw _cfg local.
- Trim the relocated comment's stale claim about _submit_with_guard (at the
new position the store init happens inside the guarded worker, not before
it).
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.
Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.
Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.
Salvaged from PR #96290 (kanban slice) with regression test.
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.
Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
httpx timeout exceptions (ReadTimeout, WriteTimeout) stringify to "",
which defeats _is_timeout_error's first-line guard (if not error: return
False). The base-layer plain-text fallback then re-sends an already-
delivered message — the user receives it twice.
Replace error=str(exc) with error=str(exc) or type(exc).__name__ at every
httpx-based adapter boundary so the existing matcher ("readtimeout",
"writetimeout") still fires. ConnectTimeout intentionally stays
unmatched: if the connection never opened the message was not delivered,
so retry/fallback remains correct.
Applies to BlueBubbles (send + _create_chat_for_handle), WhatsApp Cloud
(text + interactive + media), QQ Bot (send chunk + keyboard + media), and
Yuanbao media handler — the same latent bug exists in every adapter that
stores error=str(exc) from an httpx call.
- _write_machine_sentinel_line: wrap the print() fallback so a closed
redirected stream (ValueError, not OSError) can't propagate out of the
ready path and kill a healthy serve; document that pythonw port
discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
redirect active all server logging lands on stderr, and an unread
stderr pipe can fill and block the child before the sentinel, flaking
the test at the 120s timeout
The same stdout redirect that rerouted the READY sentinel (#96282) also
reroutes the machine-parsed BACKEND_PORT_IN_USE sentinel printed by
_report_port_in_use() — both preflight and probe-to-bind-race callers run
after tui_gateway.server's sys.stdout=sys.stderr swap. Extract the fd-1
write into _write_machine_sentinel_line() and use it at both sentinel
sites; human-facing hint lines stay on print().
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).
Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.
Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
Live /api/v1/models probe (2026-08-27) confirms the id is gone from the
catalog, so the curated picker entry was a dead pick. Manifest
regenerated. No provider-agnostic metadata existed for the slug.
Delist credit: @orouge97 flagged this in PR #80036.
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
The agent loop writes an internal "(empty)" sentinel when the
nudge/prefill/empties/fallback ladder all fail. The gateway converts it
into a user-friendly notice at delivery, but the desktop group-chat
bridge appended the raw sentinel into the room log (seen posting
"(empty)" in a Bot Mode group room), and it synced to the shared
ui_meta for mobile.
Normalize at the single choke point, appendGroupChatEntry, mirroring
gateway/run.py substitution so group chat and gateway surfaces show the
same text. (pass)/empty silence semantics unchanged. Includes a
regression test proven to fail on the pre-fix code.
Fixes#94308
Address review feedback on #94386: the new tests only exercised
runGroupChatMemberTurn's use of pickGroupTurnReply. Add the analogous
case for harvestStrandedGroupReply (substantive answer -> synthetic
continuation nudge -> (pass) tail) and document the pass-only tie-break
(newest wins) in pickGroupTurnReply's docstring.
runGroupChatMemberTurn (and harvestStrandedGroupReply) selected only the
last assistant message in a finished turn. A Codex intent-ack continuation
nudge can land a complete, substantive room answer and then get a
synthetic "(pass)" reply to the nudge itself — the terminal message picked
by the old scan, which silently discarded the real answer (#94376).
Both call sites now scan the messages appended this turn for the last
substantive (non-pass) assistant reply, falling back to a pass only when
no substantive answer exists in that window.
Source-contract tests (the group-room-ux pattern): the workspace renders
the Stop button only while room.running, wires it to stopGroupThread
(not the #94570 per-member interrupt spray), and carries no hardcoded
CJK label.
Salvaged from #94570 (@ShonnQ): the Activity bar gains a Stop button
while a round is running (room.running), plus an inline Stop on the
expanded 'working' activity row. Rewired from the original per-member
session.interrupt spray onto the stopGroupThread primitive so the round
loop actually stops (epoch bump + holds + on-turn interrupt) instead of
marching to the next member; labels are plain English like the rest of
the plugin's UI strings.
Co-authored-by: Hermes Agent <agent@nousresearch.com>
stopGroupThread(group, thread, members?) is the room's first true
cancellation primitive: it bumps the room epoch (the driving loop bails
at its next member boundary), sets #93129 holds for every member (no
future turns until an explicit release), records a 'stopped' activity
event on the new epoch, and sends session.interrupt to the member
currently on turn via its own route — previously the plugin issued zero
interrupt RPCs, so 'stop' meant waiting out the in-flight model call.
The runGroupChatMemberTurnLeased poll loop now abandons a turn whose
dispatch epoch went stale WHILE its member is held — the stop signature.
An ordinary newer-send epoch bump without a hold still polls to
completion so late work keeps landing (#93127 commit check unchanged).