Commit Graph

14240 Commits

Author SHA1 Message Date
liuhao1024 23f597a8f5 fix(cron): verify a persisted final assistant message before booking complete
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).

Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.

Fixes #93820
2026-08-27 20:39:30 +05:30
liuhao1024 ad4e61f4c9 test(cron): stale hand-edit without manual marker still re-anchors
Carried from #94034 (closed as duplicate of #94033): explicit regression
that a jobs.json-edited stale next_run_at with NO manual_run_at marker
still re-anchors without firing — the direct #93049 protection case.
2026-08-27 20:39:20 +05:30
fangliquanflq 7e64a48303 fix(cron): preserve recurring manual run intent 2026-08-27 20:39:20 +05:30
liuhao1024 b0a8d16c60 test(gateway): end-to-end rendered-unit coverage for the cron drain floor
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
2026-08-27 20:39:11 +05:30
HexLab98 1282362803 test(gateway): cover TimeoutStopSec including the cron drain floor 2026-08-27 20:39:11 +05:30
Jack Lau 1944570708 docs(state): call the macOS synchronous rule a floor, not a pin
cli-config.yaml.example said macOS is "always held at FULL regardless",
which reads as "your setting is ignored on this platform" and would talk
an operator out of choosing EXTRA. _apply_synchronous_pragma only refuses
values BELOW FULL on Darwin; EXTRA is applied normally.

The existing doc guard only asserted the key's presence, so it could not
have caught this. Pin the distinction instead.

Reported by @Enough1122 in review.
2026-08-27 07:52:26 -07:00
Jack Lau ef29fc63d7 fix(state): make state.db synchronous configurable on every platform
`apply_database_pragmas()` reads five sizing pragmas from `database:` and no
durability one, and `_enforce_macos_synchronous_full()` returns early when
`sys.platform != "darwin"`. Between them, nothing in the process ever executes
`PRAGMA synchronous` against state.db on Linux or Windows.

The effective level there is therefore `SQLITE_DEFAULT_WAL_SYNCHRONOUS`, a
compile-time constant of whichever SQLite the interpreter links. Debian and
Ubuntu builds commonly ship it as NORMAL; the bundled build and a plain
source build use FULL. So the durability of state.db is decided by which
python3 the installer found, is invisible from config, and cannot be pinned.

#90837 is three weeks of corruption forensics conducted on Ubuntu under the
stated premise `synchronous=FULL`, with every other cause eliminated live.
That premise is not something the reporter could have verified from config,
because there was no config key to set and no log line to read back.

- `resolve_synchronous_level()` maps the spellings operators actually write
  (OFF/NORMAL/FULL/EXTRA, any case, or 0-3) to the PRAGMA integer, and
  returns None for anything else. Kept out of the sizing loop on purpose: an
  unrecognised `cache_size` harmlessly falls back to a default, an
  unrecognised durability level must not.
- `_apply_synchronous_pragma()` applies it, and on Darwin refuses to lower
  below FULL. `_enforce_macos_synchronous_full()` runs during
  `apply_wal_with_fallback()`, which is earlier than `apply_database_pragmas()`,
  so without an explicit floor a configured NORMAL would silently undo #64355
  by the accident of running last. Raising to EXTRA on macOS is allowed.
- Unset changes nothing, so no existing install moves.

Tests: 34, covering the parser, application, the unset path, the typo path,
a guardrail that #77630's five keys still apply, and the Darwin floor in both
directions. Removing the wiring fails 6; removing the floor alone fails 1.

Related to #90837
2026-08-27 07:52:26 -07:00
Jack Lau d5d42b96ef fix(state): warn when an existing database's journal_mode is flipped to WAL
apply_wal_with_fallback treats an on-disk WAL database as authoritative and
says so twice: it never live-downgrades one. The mirror case had no
protection at all. When the on-disk mode is DELETE and the configured mode
is wal, the function flips the database and logs nothing.

journal_mode is a property of the FILE, so that rewrites the header and
persists after the process exits. Setting the mode directly on the file is
something operators do; it was the documented mitigation for the SQLite
3.50.4 WAL-reset bug. A config key that makes the choice durable already
exists (database.journal_mode, #68545), but nothing named it at the moment
the PRAGMA was being undone.

#89293 reports the cost: after upgrading past the vulnerable SQLite,
is_sqlite_wal_reset_vulnerable() stopped short-circuiting into
_apply_delete_for_wal_reset_bug, the flip path went live, and 4 of 5
databases silently returned to WAL with no log line anywhere.

Add a deduped WARNING at both points where the switch succeeds, decided
before the pragma runs since both inputs are only readable while the file is
still in its original state. Log-only: the flip still happens, the return
value is unchanged, and the never-live-downgrade rule is untouched.

WARNING rather than ERROR is deliberate. The reverse direction is ERROR
because dropping to DELETE costs concurrency; this direction is normally the
desirable one (managed_uv treats a database stuck on DELETE as a bug worth
repairing on update). The problem was never the change, it was that the
change was invisible.

The page_count guard is the load-bearing half. A brand-new database also
reports journal_mode=delete and is also about to be switched to WAL, and
every opener applies WAL before creating schema, so without it the warning
would fire on the first run of every install.

Refs #89293
2026-08-27 07:52:26 -07:00
Paolo Antinori 9aeb582e38 fix(state): warn when configured journal_mode=delete is overridden by on-disk WAL
When database.journal_mode=delete is configured but the on-disk DB is already
WAL, apply_wal_with_fallback honors the never-live-downgrade rule and keeps WAL.
That is correct (a live downgrade under open connections causes mixed-mode
corruption), but the operator's configured mode silently has no effect, and on
a WAL-incompatible filesystem (virtiofs/NFS/SMB) the DB then corrupts on the
next crash/sleep exactly what they configured to prevent.

Two code paths return WAL in this situation; both now emit a once-per-process-
per-db_label ERROR telling the operator the config did not apply and they must
convert the DB header offline (stop connections, PRAGMA journal_mode=DELETE):

1. The WAL-reset-vulnerable path (_apply_delete_for_wal_reset_bug): previously
   warned only about the vulnerability with an "upgrade SQLite" remedy, which
   does not help when the real cause is the filesystem. Emitted after that
   warning so the actionable message is last.
2. The read-only probe path (non-vulnerable runtime): previously returned WAL
   with no signal at all.

The never-live-downgrade behavior is unchanged (existing test now also asserts
the warning). New tests cover both paths, the per-db_label dedup, and the
require_wal=True edge case.

Real-world impact: a Hermes deployment with state.db on a Podman virtiofs
bind-mount (or any NFS/SMB home) that upgrades across a version where WAL was
the default, then sets journal_mode=delete, sees no corruption protection
until the DB header is converted. This makes the gap visible. See #68545.
2026-08-27 07:52:26 -07:00
Teknium 34393c32aa feat(plugins): wire plugin platform handlers into a2a, buzz, and qqbot adapters
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
2026-08-27 07:51:37 -07:00
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
teknium1 c96f830252 feat(plugins): let plugins register Telegram PTB handlers via ctx.register_telegram_handler
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.

Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
2026-08-27 07:51:37 -07:00
konsisumer c760143935 fix(tui-gateway): claim disconnect sessions before teardown 2026-08-27 07:51:22 -07:00
Teknium 9dfbde19db refactor(delegate_task): tasks-only interface + depth-derived delegation (1,201 → 773 tok/call, −36%) (#96424)
* refactor(delegate_task): depth-derived delegation (role param retired), session-filtered restrictions, background unadvertised — 1,201->819 tok/call

* refactor(delegate_task): tasks[] is the only advertised shape — single task = one-entry array (legacy goal/context/output_schema stay handler-accepted)
2026-08-27 07:38:53 -07:00
kshitijk4poor 726f0ce1b5 fix: keep original entry object on same-id recovery — preserve live state (review follow-up)
Review found _create_entry_from_recovered_row builds a minimal entry:
replacing the live object would silently drop model_override, token/cost
counters, resume_pending/queued-work markers, and metadata. Keep the
original entry (routing is unchanged, so no sessions.json rewrite either)
and log at INFO — this is a success path, not a corrective action.
Regression test now asserts state preservation, reopen_session call, and
no save.
2026-08-27 20:06:16 +05:30
Jackal991 ff3f25e041 fix(gateway): keep sessions.json entry when startup recovery succeeds with same session id
Closes #95957
2026-08-27 20:06:16 +05:30
kshitijk4poor b26a359de8 polish: reuse launchd domain probe, fast-observe in wedged tests, blank-line cleanup
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
  up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
  observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
2026-08-27 20:06:04 +05:30
kshitijk4poor 8872cd137c fix: verify launchd replacement PID before trusting KeepAlive (review follow-up)
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
2026-08-27 20:06:04 +05:30
Joby Ellington 7a76046a86 fix(gateway): use SIGUSR1 graceful restart on launchd, not bare SIGTERM
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.

`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:

1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
   on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
   was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
   gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
   `_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
   used by `systemd_restart()` and the updater — had no launchd call site.

2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
   0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
   systemd branch uses `_get_restart_exit_wait_budget()`
   (drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
   documents that callers falling back to a hard kill must cover both phases
   or they reintroduce #77184.

The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.

Observed on macOS 27.0 / Hermes 0.20.4:

    → Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
    ⚠ Gateway PID 49787 still running after 0.0s — restart may fail
    ⚠ Gateway drain timed out after 0s — forcing launchd restart

Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.

The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.

Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 20:06:04 +05:30
x7peeps 1348e65e26 fix(gateway): rebase launchd --replace removal onto main
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.

Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.

Refs: #79096
2026-08-27 20:06:04 +05:30
wansui 9faa953cd1 fix(wecom): eliminate duplicate + split bubbles in native streaming
Fixes two related native-streaming bubble defects surfaced in production:

1. Duplicate bubble on long turns — when keep-alive already refreshed the
   6-min reply window, the Layer-2 clock fallback still declined the finalize
   frame and forced a proactive send(), duplicating the message. Skip the
   clock fallback while keep-alive is active; intermediate-frame failures are
   now fully fire-and-forget (only a failed FINAL frame falls back to send()).

2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
   a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
      notice) called _reset_segment_state(), clearing the cumulative
      _accumulated so the next delta + finalize frame carried only the few
      chars accumulated after the reset. Native streaming now skips that
      reset (commentary still posts as its own message via send()).
   b) The adapter-side _BlockChunker.update() 'only grow' guard silently
      dropped any cumulative snapshot shorter than its high-water mark, so
      after a baseline reset the leading characters were stranded before
      _emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
      layer entirely; intermediate frames are pure identity-dedup, matching
      the fire-and-forget model.

Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.

Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).
2026-08-27 07:33:36 -07:00
wansui 2ecb544551 feat(wecom): native reply streaming (per-turn isolation, dedup-safe delivery, interaction boundaries)
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.

Includes the machinery intrinsic to native streaming on WeCom:

- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
  frames update it, a finalize frame closes it; native-streaming adapters are
  let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
  long-connection mode has no documented edit-rate limit); an adapter-level
  frame cap is retained. (Early builds gated frames behind a char throttle;
  removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
  isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
  (control vs normal) plus a per-chat token bucket to stay under WeCom's rate
  limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
  delivered the moment it is emitted; failures logged, not re-sent; delivery
  marked once per turn), per-req_id reply queue with ack tracking, and the
  timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
  846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
  the prompt is the last thing on screen and never traps a lingering bubble;
  eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
  stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
  turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
  messages; image+text double-callback merged into one turn.

Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.
2026-08-27 07:33:36 -07:00
wansui 42dc0dea70 fix(send_message): cross-loop dispatch to live WeCom adapter
When send_message is invoked from the agent's worker thread (a different
event loop than the gateway's), awaiting the WeCom adapter directly can hang
because the adapter enqueues onto the gateway loop. Dispatch via
run_coroutine_threadsafe onto the gateway loop when the caller loop differs,
with caller-cancellation shielded so an already-enqueued send is not cancelled
mid-flight (which would otherwise cause a false-failure retry -> duplicate).
Recognizes WeCom native chat IDs as explicit send targets and whitelists WeCom
for media delivery. Part of the async queue design this branch introduces.
2026-08-27 07:33:36 -07:00
wansui 81aa4f18a8 feat(gateway): per-platform streaming config default for WeCom
Add WeCom to the per-platform streaming defaults (DEFAULT_CONFIG display
plumbing) so native streaming is enabled by default for the WeCom adapter,
alongside the existing per-platform flags. Non-secret config lives in
config.yaml (no HERMES_* env vars).
2026-08-27 07:33:36 -07:00
teknium1 034c6ac77b refactor(skills): move publish-site to optional-skills
Per the 'when in doubt, optional' rule — site publishing is an
on-request capability, not a weekly daily-driver for most users.
Joins cloudflare-temporary-deploy/page-agent under
optional-skills/web-development (existing category, existing
DESCRIPTION.md kept; the new bundled category dir is dropped).

Install via: hermes skills install official/web-development/publish-site
2026-08-27 07:30:15 -07:00
teknium1 5c57e77541 feat(skills): publish-site — versioned website publishing to GitHub/Cloudflare/Netlify Pages
Zero-core-footprint equivalent of ChatGPT Work Sites: preview, version-before-deploy, provider ladder, rollback, verification.
2026-08-27 07:30:15 -07:00
joaomarcos 59993be6e9 docs(sessions): name the reachable ordinal-0 rewind path in the #95868 guards
Precision pass on the comments added by the previous commit: prompt.submit
reaches replace_messages(archive_dropped=True) with an empty prefix on a
confirmed ordinal-0 rewind, which is the production shape that lands a
populated session on message_count = 0. archive_and_compact normally
publishes at least a summary row, so it is pinned as defense in depth
rather than claimed as an equally reachable trigger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
2026-08-27 19:54:32 +05:30
joaomarcos 0bfe715e4d fix(sessions): don't let the empty-session sweep delete an archived transcript
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:

  * `replace_messages(..., archive_dropped=True)` — the rewind / edit /
    regenerate mode added in #82756 so a taken-back turn stays recoverable.
  * `archive_and_compact` — in-place compaction, which archives the
    pre-compaction transcript under the same session id (#38763).

A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.

Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.

Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.

Fixes #95868

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
2026-08-27 19:54:32 +05:30
Haakam Aujla 14320d18df feat(skills): rewrite AgentMail optional skill CLI-first
Replaces the stale MCP-first AgentMail skill with CLI-first guidance:
self-signup + OTP verification, inbox/message/thread/label/attachment
flows, webhook and WebSocket delivery references, and MCP as an
alternative path. Declares AGENTMAIL_API_KEY (optional) so a stored key
reaches the sandboxed terminal while self-signup stays viable without
one.

Salvaged from PR #60811 — kept in optional-skills/ per the March 2026
decision that third-party-API-key skills are not bundled.
2026-08-27 07:09:26 -07:00
Heath Harris 3113f6056b fix(cron): open the session store only after wake-gate and validation early-returns
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.

Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.

Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
2026-08-27 19:07:15 +05:30
Heath Harris a71be9852e fix(kanban): close half-open tracked connection when busy_timeout PRAGMA fails
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.

Salvaged from PR #96290 (kanban slice) with regression test.
2026-08-27 19:07:15 +05:30
fangliquanflq 6662b3618d fix(computer-use): accept current notarised CUA Driver 2026-08-27 20:21:34 +08:00
nftpoetrist 08b4875f4a fix(deadline): remove the dead second SuspectableBackend class shadowing the Phase 3a Protocol
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.

Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
2026-08-27 17:22:02 +05:30
Tranquil-Flow 6cbb7b6115 fix(gateway): preserve exception type when error string is empty (#78183)
httpx timeout exceptions (ReadTimeout, WriteTimeout) stringify to "",
which defeats _is_timeout_error's first-line guard (if not error: return
False).  The base-layer plain-text fallback then re-sends an already-
delivered message — the user receives it twice.

Replace error=str(exc) with error=str(exc) or type(exc).__name__ at every
httpx-based adapter boundary so the existing matcher ("readtimeout",
"writetimeout") still fires.  ConnectTimeout intentionally stays
unmatched: if the connection never opened the message was not delivered,
so retry/fallback remains correct.

Applies to BlueBubbles (send + _create_chat_for_handle), WhatsApp Cloud
(text + interactive + media), QQ Bot (send chunk + keyboard + media), and
Yuanbao media handler — the same latent bug exists in every adapter that
stores error=str(exc) from an httpx call.
2026-08-27 17:21:26 +05:30
kshitijk4poor 8d95ab1b37 fix(serve): review follow-ups — never-raise sentinel fallback, DEVNULL stderr in split-stream test
- _write_machine_sentinel_line: wrap the print() fallback so a closed
  redirected stream (ValueError, not OSError) can't propagate out of the
  ready path and kill a healthy serve; document that pythonw port
  discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
  redirect active all server logging lands on stderr, and an unread
  stderr pipe can fill and block the child before the sentinel, flaking
  the test at the 120s timeout
2026-08-27 17:11:56 +05:30
Kitson Kelly f2dd32d3e5 fix(serve): announce READY sentinel on fd 1, not the redirected sys.stdout
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).

Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.

Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
2026-08-27 17:11:56 +05:30
fangliquanflq a65ad15636 fix(agent): honor explicit free OpenRouter models 2026-08-27 04:37:36 -07:00
Adolanium a9611f3c6f feat(models): add GLM-5.3-Flash to z.ai and OpenCode Go pickers
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
2026-08-27 04:14:31 -07:00
fangliquanflq 091cc0e8be fix(hermes_cli): scope hook timeouts and fail closed on pre_tool_call
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
2026-08-27 16:13:45 +05:30
kshitijk4poor 3a94101524 fix(browser): suspect-session recycle + wedged-daemon tree-kill on timeout (#85125 3b-browser/4c)
Builds on e11187208f (salvage of #72206 by @luyifan, authorship
preserved) which added the post-timeout session reset. This commit
completes the Phase 3b contract:

- suspect flag: a command timeout marks the session key suspect;
  the NEXT _get_session_info for that key health-checks and recycles
  the session instead of handing back the poisoned handle (#72205)
- flag cleared on fresh-session store so a healthy new session is
  never spuriously recycled; cross-key leakage fixed
- wedged-vs-alive rule (#68139): after a timeout, if the daemon's
  socket still answers, recycle the session only; if it is
  unresponsive (or no socket exists to probe — conservative), tree-
  kill via agent.deadline.kill_process_tree and evict
- negative probe: successful commands never mark or recycle

Tests: tests/tools/test_browser_suspect_recycle.py (20 tests:
mark-once, recycle-then-succeed, success-never-recycles, tree-kill
invoked on the wedged path with pid assertion, flag lifecycle).

Co-authored-by: luyifan <al3060388206@gmail.com>
2026-08-27 16:11:16 +05:30
luyifan 0be9b7cd57 fix(browser): reset sessions after command timeout 2026-08-27 16:11:16 +05:30
kshitijk4poor 1e5fb70fb5 fix(cron): widen the lock-first liveness check to 'hermes cron status'
Sibling site of the salvaged #95947 fix (same file): cron_status
declared 'Gateway is not running — cron jobs will NOT fire' from a bare
find_gateway_pids() miss even while the runtime lock proved the gateway
alive. Now the not-running verdict requires both the scan AND the lock
to read dead; when only the lock answers, the pid line falls back to
the recorded gateway pid (or is omitted).

Two regression tests pin the false-alarm suppression and the genuine
not-running warning.
2026-08-27 16:10:57 +05:30
kshitijk4poor f2d043eb08 test(cron): pin lock-first liveness + harden lock-probe failure
Follow-ups to the salvaged #95947 cron commit:

- Wrap the lock probe in its own try/except: a crashing probe is
  'unknown', not 'dead' — the pid scan still decides instead of the
  whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
  held + empty pid scan -> alive (the reported false alarm); lock
  inactive -> pid-scan fallback both ways; crashing lock probe still
  falls back.
- patch_liveness now pins the lock probe inactive by default so the
  pre-existing pid-scan tests stay deterministic on machines where a
  real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-27 16:10:57 +05:30
Teknium 9b44273c05 fix: follow-up for salvaged PRs #93250 + #96234
- move minimax/minimax-m3:free into the Free tier section (house
  convention: :free SKUs group together, matching glm-5.2:free and the
  nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
  metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
  otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
  same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
  reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
  ':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
  family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
2026-08-27 03:24:20 -07:00
Teknium ca753b96cb fix(tui-gateway): unset semantics for every live-adopted compression/model key
Independent review finding on the merged #95980:
_apply_live_compression_config only acted on PRESENT keys, so removing
tail_mode / model.context_length / target_ratio / model_thresholds /
proactive_prune_* / protect_last_n / min_tail_user_messages / threshold /
idle_compact_after_seconds from config.yaml left stale values active in
live sessions forever (probe-verified: all six stale after applying
empty mappings).

Absence now restores the normalized default — or the model-derived
value — through the SAME derivation the construction path uses:

- ContextCompressor ctor defaults read off its real __init__ signature
  (no hardcoded copies to drift)
- compression.threshold removal re-derives via agent_init's
  _resolve_compression_threshold (Codex gpt-5.4/5.5 + spark autoraise
  included)
- model.context_length removal drops the config override and forces
  re-inference through the deferred get_model_context_length resolution,
  which also re-applies the small-context threshold floor
- model_thresholds removal clears stale per-model overrides from the
  live threshold; tail_mode falls back to the ctor's 'lean' (the old
  present-key path normalized invalid values to 'legacy', diverging
  from the compressor's own fallback)

Also fixes proactive_prune_min_reclaim_tokens's present-but-null default
(was 0; the real default is 4096).

Refs #94724
2026-08-27 02:18:16 -07:00
Teknium 6d4e851d80 fix(serve): bounded flush-on-SIGTERM + periodic incremental session flush
A hermes serve killed mid-update lost every un-flushed in-memory session
(#94724 item 2, reported by @ruangraung): the next RPC failed with
'session-scoped RPC rejected: not in memory (detached/reaped runtime)'
and no store held the transcript. #95576 made serves survive future
updates; this closes the kill path itself:

- install chaining SIGTERM/SIGINT handlers (hermes serve / dashboard
  startup, before uvicorn's capture_signals) that first persist
  in-memory session transcripts to state.db — bounded by
  HERMES_TUI_EXIT_FLUSH_BUDGET_S (default 5s, daemon worker + join) so
  a hung SQLite write can never block exit
- _shutdown_sessions (atexit) runs the same bounded flush FIRST, before
  the slow per-session teardown a supervisor may SIGKILL mid-way
- the idle-reaper scan piggybacks a periodic incremental flush
  (marker-deduped agent._persist_session, running sessions skipped) so
  even a SIGKILL loses at most one flush interval — no new timer
  subsystem

Refs #94724
2026-08-27 02:18:16 -07:00
Teknium 01f7ce5b76 feat(sessions): one-shot single-match owner backfill for legacy NULL-profile rows (#94724)
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.

Refs #94724
2026-08-27 02:17:56 -07:00
yflmq001 349e6611a1 fix(approval): enforce explicit timeout on smart-approval guardian call and log its outcome
The smart-approval guardian (`_smart_approve`) gates every flagged
terminal command with a synchronous auxiliary LLM call, but it never
passes `timeout=` and logs nothing on the normal path. In production a
stalled provider response silently froze the agent turn for 62 minutes
with zero log output; the gateway kill-switch eventually fired, and only
an unrelated error surfaced afterwards (#82846; watchdog-style fix in
#72500). The call was invisible by design — nothing logs at the hang
point.

Changes in tools/approval.py:
- Resolve the same configured timeout the client would use internally
  (`auxiliary.approval.timeout` via `_get_task_timeout("approval")`) and
  pass it explicitly to `call_llm`, so the deadline cannot be lost if the
  internal default resolution changes or is misconfigured.
- Log the assessment call and its duration (DEBUG), and promote the
  failure branch from DEBUG to WARNING with elapsed time + exception
  class, so a wedged guardian call is visible in the logs instead of
  silent.
- Failure still returns "escalate" (fail open to the human/pattern
  gate) — behavior unchanged, observability only.

Complements #72500 (watchdog hard ceiling) rather than duplicating it:
explicit timeout is the root-cause hardening, logging closes the
silence gap; the watchdog remains the safety net if the SDK-level
timeout itself is defeated.

Tests: explicit timeout forwarded to call_llm (revert-fails), failure
logs WARNING + escalates. 49 approval-adjacent tests pass; one unrelated
test_approval.py failure is pre-existing (fails on clean main too).
2026-08-27 14:28:22 +05:30
Ayush Nangia 2f0f01192d fix(cron): tree-kill script timeout descendants via agent.deadline.kill_process_tree
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.

The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.

Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
  lost" path orphaned setsid grandchildren the same way (whole-bug-class
  rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
  that exits right at the deadline doesn't log a spurious "no signal"
  warning (mirrors _terminate_cron_script_process); pinned by
  test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
  CI load could eat the whole 1s window before the spawner wrote its
  pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
  old path gave a 1s SIGTERM grace window; intended for a deadline-
  expiry hard stop (both docstrings say "hard stop")

Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.

Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
2026-08-27 14:28:06 +05:30
kshitij f45477b4a8 Merge pull request #86412 from kshitijk4poor/salvage/83225-overflow-clamp
fix(approval): oversized approvals.timeout crashes parallel tool batches — clamp at config read (salvage #83225/#83298, #85125 2b)
2026-08-27 14:26:38 +05:30