The initial (pre-running) connect awaited during gateway startup now uses
a capped 45s budget for Telegram instead of the full 180s (#67498) budget.
On timeout the platform is queued for the reconnect watcher, which retries
with the full budget and is_reconnect=True (preserving the offline update
queue, #46621). Combined with the parallel startup connects, an unreachable
Telegram no longer holds the whole gateway out of the running state.
The previous concurrency assertion (slow_start < fast_end) was true under
BOTH the serial and parallel implementations, so it proved nothing -- it
even passed against the old serial code on main. The only assertion that
distinguishes the two is that the fast platform finishes before the slow
one (fast_end before slow_end), which is only possible when the connects
overlap.
Switch the test to record connect start/end events in arrival order
(clock-resolution independent) and assert fast_end precedes slow_end. This
also fixes the Windows failure @zuowen7 reported: time.monotonic() has only
~15 ms resolution there, so two parallel connects could land on the same
tick and defeat any wall-clock comparison -- event ordering cannot.
Verified the new test fails against origin/main (serial) and passes against
this branch (parallel).
GatewayRunner.start() previously awaited each platform's connect() (with its
own timeout) in a serial for-loop. A single slow/failing platform (e.g.
Telegram behind a dead proxy) delayed every later platform's connect by a full
timeout window, cascading one platform's failure onto WeChat/QQ/etc.
Now the slow connect() calls run concurrently via asyncio.gather while the
serial pre-filter (checks, adapter creation, handler wiring) and the
single-threaded result aggregation (shared-state mutation, error handling)
are unchanged. A failing platform no longer blocks the others.
Adds regression tests proving connect() calls overlap and that one failing
platform leaves the others connected.
mark_running_jobs_interrupted skipped legacy fires without a registered
durable owner entirely — correct for the persisted last_status write
(no owner fence to protect a replacement run), but the gateway shutdown
path also uses the returned ID list to deliver interrupted-cron notices
while adapters are still connected (#82232). Keep the persistence skip,
but include the job in the returned list so the user is still told.
When the shutdown drain times out and kills an in-flight cron job, the
job's owner is never told. The cron worker does try: `_is_interrupted()`
forces the failure path with an honest "interrupted by gateway shutdown"
error, and failed jobs always deliver. But that worker is a thread, it
reaches `_deliver_result()` asynchronously, and by then
`_bounded_adapter_teardown()` has closed the transport. The reporter of
Worse, the loss is silent twice over: `_consume_interrupted_flag()`
returns True — the gateway already wrote `last_status` — so
`mark_job_run()` is skipped, and the `delivery_error` from the failed
send is discarded with it. The run's only trace is a generic line in
jobs.json.
The gateway already owns the right window. `_notify_active_sessions_of_
shutdown()` runs while adapters are up, precisely so shutdown messages
can be sent — but it iterates `_running_agents`, and cron work lives on
the scheduler's own thread pool. Same structural blindness already fixed
for counting (#60432) and draining (#63529), never fixed for notifying.
So notify from the post-interrupt phase, which is the last point where
the transport is still up: `_kill_tool_subprocesses()` now returns the
job IDs it marked, and `_notify_interrupted_cron_jobs()` sends each one's
owner a notice on the job's own resolved delivery targets. Adapter
teardown order is untouched — it is load-bearing for #53175 and #8202.
Jobs with `deliver: local`, and `deliver: origin` jobs with no resolvable
origin (#43014), resolve to zero targets and stay silent. Per-platform
`gateway_restart_notification: false` is honoured, matching the chat
path. Every failure is swallowed so a wedged adapter cannot extend
shutdown.
Second, when the interrupted flag short-circuits `mark_job_run()`, the
delivery failure is now persisted on its own via `update_job()`, so a
notice that still cannot be sent is at least recorded. `update_job()`
rather than a second `mark_job_run()`: the latter also advances
`next_run_at` and the repeat counter, and running that twice for one run
would skip a fire or auto-delete the job early.
Fixes#82232. Related: #82161, #82224.
CI slice 5/12 caught two ways the new cron budget broke `_stop_impl_body`
for callers that are not real GatewayRunner instances:
- `_FakeGateway` in test_shutdown_cache_cleanup.py borrows `_stop_impl`
without subclassing, so it never picked up the class-level
`_cron_drain_timeout` default and raised AttributeError. Read it through
the getattr-guard convention the same function already uses for its
liveness-guard machinery.
- The same double overrides `_drain_active_agents(self, timeout)`, so
passing the cron budget raised "takes 2 positional arguments but 3 were
given". The double now mirrors the real optional parameter. It is the
only override in the tree; test_startup_restart_race.py uses AsyncMock,
which accepts any signature.
Verified against a stashed clean tree: the 22 gateway test files that
still fail locally fail identically with and without this branch (80 = 80,
empty set difference both ways) — they are pre-existing Windows-only
failures (setsid, POSIX modes) unrelated to this change.
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.
A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.
Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.
The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.
Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.
Relates to #82161 (complements #82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.
- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
adopt only when tip != old_id AND the tip row is live, retry the flush
exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
lookup with the same tip + liveness contract, so gateway transcript
reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
recovery preflight now resolves via get_compression_tip with the same
liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
explanation names compression rotation and tells the client to refresh the
session id instead of blaming a full disk.
Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).
Closes#82001
Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
On macOS (256 soft fd limit), routing the weixin/email pollers through a
local HTTP proxy leaked one TCP socket per failed poll/connect cycle
until the gateway hit `[Errno 24] Too many open files` and crashed
(launchd respawn loop). Live capture showed 216 of 256 fds pinned on
connections to the proxy, ~214 of them abandoned.
Code-side gaps fixed:
- email adapter, `connect()`: no try/finally around the IMAP test
connection — a failure in login/ID/select/search abandoned the
connected socket with no owner. Every reconnect-watcher retry builds
a fresh adapter, so each retry against an unreachable/proxied host
leaked another fd. Teardown now runs in `finally`.
- email adapter, IMAP teardown: `imaplib.IMAP4.logout()` only swallows
`OSError` internally; on a broken connection `LOGOUT` raises
`IMAP4.abort` before the internal `shutdown()`, leaving the socket
open. New `_close_imap()` helper chases a failed `logout()` with an
unconditional `shutdown()`; used in `connect()` and
`_fetch_new_messages()`.
- weixin adapter: repeated poll failures through a proxy strand
sockets in the aiohttp connector where the tight keepalive reaper
never sees them. The poll loop now recycles its ClientSession
(swap-then-close, safe for concurrent `_process_message` tasks)
after each MAX_CONSECUTIVE_FAILURES streak, tearing down the
connector and every socket it holds.
Targeted tests: tests/gateway/test_poller_fd_lifecycle.py (9 tests).
Reported by @EthanHunter1229 with measured fd captures.
The `_owner_alive` fallback in gateway/delivery_ledger.py (taken whenever a
process start time is unreadable) probed liveness with a raw
`os.kill(pid, 0)`. On Windows that is NOT a no-op: CPython maps sig=0 to
`GenerateConsoleCtrlEvent(0, pid)` (bpo-14484), so probing a LIVE pid whose
start time psutil could not read would Ctrl+C the target's entire console
group. The prior `# windows-footgun: ok` annotation only justified the
EPERM-means-alive exception semantics, not the Ctrl+C side effect.
Route the probe through `gateway.status._pid_exists` (psutil-first,
ctypes OpenProcess fallback on Windows), preserving EPERM-means-alive.
A POSIX-only raw-probe fallback remains for the unreachable case where
gateway.status cannot be imported; on Windows that path reports dead
rather than firing a sig-0 probe.
Tests patch `gateway.status._pid_exists` per the windows-native-support
pattern, including a regression guard asserting os.kill is never used
for the probe. Sabotage-verified (revert → red, restore → green).
Part of #41662 (the os.kill half; watchdog half tracked separately).
Teknium review on #75732: releasing the pending clarify whenever
resolve_text_response_for_session returned False also cancelled
retryable multi-select invalid selections (out-of-range numbers,
unrecognised comma-lists).
Classify rejected typed replies in clarify_gateway:
- rejected_prose → cancel clarify, fall through busy routing (deadlock break)
- rejected_selection → keep clarify armed so the user can retry
Add native multi-select gateway regressions for both paths.
Fixes#82975.
The adapter-level clarify reply bypass in gateway/platforms/base.py's
handle_message() built its session_key via build_session_key(...)
without a profile= argument, defaulting to the legacy agent:main
namespace. The runner registers pending clarifies under
SessionStore._generate_session_key()'s key, which DOES include
profile=self._resolve_profile_for_key(source). Under a named-profile
multiplex these diverge, so the bypass lookup at
clarify_gateway.get_pending_for_session(session_key, ...) misses --
the user's answer to a pending clarify() gets routed to the adapter's
busy-session queue instead of resolving it. The turn then hangs until
the clarify's 3600s timeout, with no inbound message: log line and no
"Gateway intercepted clarify text response" log line, matching the
reported Telegram symptom exactly.
Verified the divergence directly: _resolve_profile_for_key() returns
None when multiplex_profiles is off (default) -- byte-identical to
the prior implicit profile=None, so this only changes behavior for
multiplexed deployments, matching the issue's exact reported scope.
Fixed by using the same self._session_store._resolve_profile_for_key()
the runner's key generator calls, guarded with getattr() + a None
fallback since _session_store is set via a setter and can be unset
for adapters that never call set_session_store() -- preserving prior
behavior for any such adapter rather than introducing a new crash.
Added a regression test alongside the existing bypass coverage: with a
mocked session_store configured for profile multiplexing, a clarify
registered under the profile-namespaced key must still be found and
resolved (not routed to the busy queue). Verified as a genuine
regression by reverting the fix and confirming the new test fails
with the exact reported symptom (the message handler never gets
awaited -- the clarify lookup misses).
21/21 pass across the five directly related clarify test files;
16/16 across the broader multiplex/clarify-progress test files (no
regression).
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
Address review on #83878:
- Permanent fatal fences all hold producers and discards pending maps on
teardown instead of re-populating a queue that can never drain.
- Any hold created while connected schedules a tracked redispatch (cancel-
after-pop no longer orphans until a future reconnect).
- Redispatch failures re-hold current + remainder without tight-looping.
Regression coverage for the three residual paths, plus the interaction
with OOF-156's connect-failure classification: the retryable network
path (telegram_connect_error) must NOT clear the hold queue — reconnect
is precisely what drains it; only non-retryable fatals discard.
The disconnect drop-guard (#55971) correctly prevents dispatch into a
torn-down session. Destroying the event was wrong: by enqueue/flush time
python-telegram-bot has already acked the update and advanced the polling
offset, so Telegram never redelivers. Result: silent permanent loss, no
log, no error.
Hold inbound events (text/photo/media-group) when the drop-guard fires,
salvage pending batch maps on teardown, cancel+await the redispatch task
in the delivery cancel map (lifecycle-tracked), and redispatch from
_mark_connected after reconnect. Cap the hold queue (default 64), dedupe
by object identity, discard on non-retryable fatal. Cancel-after-pop in
flush paths also holds.
Distinct from #72037 (cancel-after-pop during follow-up supersession) and
#81528 (boundary discard). Tests use delay=0 and entered/release Events —
no wall-clock races; includes production terminal-step coverage.
Follow-up on the #83854 salvage: prepend $HERMES_HOME/bin ahead of the
venv and user-local bin dirs, matching the managed-first Browser Use
CLI resolution policy — the worker resolves the same canonical binary
the agent process does.
`hermes gateway setup` writes `gateway.platforms` as a LIST of
enabled platform names (e.g. `- telegram`), not a dict. Treat any
non-dict shape as "no per-platform overrides" instead of crashing
on `.get()` for every incoming turn (#83185).
Co-authored-by: SeashoreShi <seashore.shi@gmail.com>
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).
Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.
Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.
Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
post-turn tick completion, and a supervised loop_wakeup_watcher that
injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
notification-poller wakeup driver + post-turn completion in the turn
dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
ticks defer until it finishes, pauses, or parks; real user input always
wins over both
Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
set_session_metadata() and advance_compression_session()'s repoint both
stamped entry.updated_at = now. updated_at is the user-activity clock that
drives idle/daily reset policy and the restart-resume freshness gate
(suspend_recently_active, #85709), so a background metadata write (e.g.
Slack thread watermark) or a background compression repoint on a long-idle
session could make it look freshly active and get it falsely resume_pending
after a gateway restart.
These are the last internal stamp sites after 784f733cf (recover) and
5462f689b (touch_activity gating): drop the stamps, keep the durable save.
Follow-up to #85895 (closed) — credit @GodsBoy for the report-side push and
@chelsealong for the analysis on #85709.
When the inactivity reaper interrupts a timed-out turn, the interrupt
frees the blocked frame — destroying the only evidence of where the
turn was wedged. The Aug 2026 zombie-turn incident (WhatsApp session,
Relay-corrupted scope stack) wedged every turn for exactly the 1800s
timeout somewhere between 'Turn ended' and run_sync returning, and the
wedge point was unprovable post-mortem.
The reaper now logs the stack of every thread with turn-machinery
frames BEFORE interrupting, so the next occurrence names the exact
blocked line. Best-effort, bounded (8 threads, 25 frames), pure
in-process, never raises into the reaper.
A turn idle past agent.gateway_timeout (the same threshold the turn
reaper uses) no longer defers an in-band restart. Restart is usually
the remedy for a wedged turn; waiting restart_after_turn_timeout on
one inverts the graceful path's purpose — a wedged WhatsApp turn
pinned 'hermes update' in draining until SIGTERM was sent manually
(Aug 2026). stop()'s bounded drain interrupts wedged turns instead.
gateway_timeout=0 (unbounded turns) disables wedge detection; cron
and API-server work has no per-turn activity clock and is never
counted as wedged; unreadable activity summaries fail open.
The 21600s (6h) default shipped in #77184 makes an interactive
'hermes gateway restart' block for up to six hours when a turn wedges
(hung tool call, wedged event loop, stuck provider stream) — the exact
scenario the cap exists for. The intent (don't force-kill an agent
mid-turn) is sound, but the default must be a safety valve for hung
agents, not a target latency.
Lower to 1800s (30 min): still protects the overwhelming majority of
long autonomous turns (tool calls have their own timeouts well below
that), keeps worst-case interactive restart latency human-tolerable, and
users running very long unattended turns can raise it in config.yaml.
RED: new contract test fails on old 21600 default. GREEN: 4/4.
Plain type=completion events built in _run_process_watcher carried only
session_key (chat/thread routing) with no spawning-session stamp, so after
/new (or a session switch) a completion notification from the OLD session
was injected into the chat's NEW session. Main already solved this exact
class for async delegations via the _classify_completion_target pre-flight
(_USER_BOUNDARY_END_REASONS drop on user-closed sessions, deliver on
idle-ends, follow the compression-tip chain), but the gate only ran for
type=async_delegation events.
Kernel salvage of #16455:
- Stamp the spawning conversation's session-db id (HERMES_SESSION_ID via
session-scoped env) on the ProcessSession and the pending_watchers entry
at spawn time in tools/terminal_tool.py; persist it through the process
registry checkpoint/restore so recovered watchers keep the stamp.
- Thread the stamp into the completion_evt built by _run_process_watcher
(watcher entry first, ProcessSession fallback for recovered watchers).
- In _deliver_completion_notification, run the SAME pre-flight classifier
for stamped type=completion events: terminal -> drop with a log (output
stays available via process(action='log')), retry -> False so the
watcher re-polls, deliver -> proceed. The policy has exactly one owner
(_classify_completion_target); nothing is forked. Unstamped legacy
events keep today's deliver-always behavior, and the async-delegation
path is untouched.
Based on the session-boundary approach from #16455 by @Tosko4 (original PR
was over-scoped across adapters/slash-commands/cron; this lands the kernel
only).
Tests: completion from a /new-closed session is dropped; completion after
an idle-end still delivers; unstamped legacy event delivers; retry verdict
returns retryable False without adapter injection; async_delegation gate
unchanged; stamp survives checkpoint recovery.
The async-delegation watcher drained the completion queue as a batch but
then delivered each event as its own synthetic turn, flooding the session
when a fan-out of background subagents finished together. Builds on the
per-process completion batching salvaged from PR #71898 (thanks
@yuzilongleif-collab) which coalesces concurrent _run_process_watcher
completions behind a short per-route fan-in window.
This commit adds the async-delegation half: group the drained batch by
full routing key (session_key + parent_session_id + platform/chat/thread/
user) and inject ONE consolidated turn per group. Durable-ack handling
stays honest: sibling rows are claimed up front via claim_event_delivery;
rows another consumer owns are excluded from the consolidated text (no
double-delivery); sibling claims are acknowledged only after adapter
acceptance and released (still pending) on failure. Events for different
sessions never coalesce, and a single-event group rides the existing
per-event path unchanged (latency and text identical).
Tests: 3 same-tick events -> exactly one adapter.handle_message carrying
all 3 results with all 3 durable rows delivered; 2 sessions -> 2 turns;
single-event path unchanged; failed batch releases claims and retries;
foreign-claimed sibling excluded and left pending.
Async-delegation batch completions and background watch notifications
re-enter the gateway as synthetic MessageEvent(internal=True) turns via
_inject_watch_notification, but were persisted as bare role='user' rows —
indistinguishable from real user input in transcripts and the desktop UI.
Thread the event's internal flag through to persistence: when
event.internal is set, the turn's persisted user row is stamped
display_kind='internal_notification' (the existing DB-only presentation
sidecar used by auto_continue / model_switch rows). Wired through
_run_agent → _run_agent_inner → TurnContext → run_conversation's
persist_user_display_kind, and onto the three gateway-side fallback user
rows (transient failure, no-new-messages, pre-run crash), whose
append_to_transcript writer now forwards display_kind/display_metadata
to SessionDB.append_message.
Invariants preserved: role stays 'user' (alternation untouched), no new
injections, no past-context mutation, and display_kind is already popped
from every provider-bound copy in conversation_loop, so replayed sessions
never leak the marker to the API.
Regression tests: internal turn marked, real user turn unmarked, fallback
rows marked/unmarked per event, and a DB round-trip proving replay keeps
role/content intact while the provider copy drops the marker.
format_process_notification had no case for watch_overflow_tripped /
watch_overflow_released, so a watch-pattern notification flood surfaced
as '[IMPORTANT: Background process exited (exit code ?)]' — a phantom
exit notification for a process that never existed — while the actual
'watch flood, N notifications suppressed' summary in the event's
message field was silently dropped. The gateway delivery path was
worse: _drain_gateway_watch_events retained only watch_match and
watch_disabled, discarding overflow events entirely before formatting.
Route both event types through the message field in the shared
formatter and the gateway formatter, and retain them in the gateway
drain.
Background process completions on messaging platforms now default to a
one-line status message (✅/❌ + command + duration; failures append a
short output tail) instead of dumping the raw output buffer into the
chat. New display.background_process_notifications mode 'concise' is
the default; 'all' keeps the old raw-dump behavior for anyone who wants
it. Config migration v35 moves users still on the old implicit default
'all' to 'concise' on their next update; explicit result/error/off
choices are preserved.
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).
- gateway/disk_status.py (new): collect_disk_status() samples
shutil.disk_usage(HERMES_HOME) and classifies pressure
(critical: <256 MB free or >=95% used; elevated: <512 MB free).
Never raises — degrades to pressure="unknown" with null telemetry,
same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
resource banner with worst-first triggers (disk critical > memory
critical > OOM restart > disk elevated > memory elevated) and
cascading dismissals — hiding the top trigger surfaces the next one
instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
with English fallback per existing pattern).
Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
Addresses the human review findings on the memory-pressure feature:
* [P2] Dismissal hid later incidents of the same kind. The gateway now
publishes `boot_id` (the lifecycle sentinel's started_at — changes on
every gateway life) in the /api/status memory block, and the dashboard
keys OOM-restart dismissal on it: acknowledging one restart no longer
mutes the NEXT one (the OOM-loop case this banner exists for). Live
pressure dismissals now also reset once pressure is demonstrably back
to "ok" — "unknown" (stale heartbeat) is absence of evidence and
clears nothing. Dismissal storage moved to a JSON list; old bare-string
entries fail JSON.parse and degrade to a clean reset.
* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
heartbeat), not proof the OOM killer acted — banner copy now says
"restarted unexpectedly, most likely because it ran out of memory"
instead of stating OOM as fact.
* [P3] Mobile header clearance was applied per-banner (mt-14 on both
MemoryPressureBanner and ProfileScopeBanner) AND on the content
(pt-14), double/triple-stacking 56px gaps when banners were visible.
Replaced with a single h-14 spacer above the banner stack.
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.
Carry the worker's completion handoff into the synthetic creator wake
turn and label it as an automatic notification with inspect-the-board /
don't-recreate guidance, so a woken orchestrator doesn't re-decompose
work that already exists (#70752).
Salvaged from PR #71100 by @yinkev; ported onto the restructured wake
region (delivery_mode gating, scope_id, sub chat_id destinations). The
auto_subscribe_on_create config-default half of the original PR was
dropped as already superseded on main.
Widen #62804's class fix: _get_goal_manager_for_event and
_get_heartbeat_manager_for_event also call get_or_create_session on
behalf of the triggering event; when that event is internal the lookup
must not advance the user-activity clock either.
For a push-adapter subscription with delivery_mode='wake' the visible text
ping is intentionally skipped (the send_passive gate), so the wake injection
IS the sole delivery — yet the event cursor advanced BEFORE the wake, which
then ran best-effort with its failure swallowed. A single failed wake
permanently lost the event.
Apply the same ordering the non-push (api_server) self-post branch already
uses: attempt the wake BEFORE advancing the cursor; on failure rewind the
claim (_kanban_rewind) and bump the per-sub failure counter so the next tick
retries; on success reset the counter; drop the subscription after
MAX_SEND_FAILURES consecutive failures like text sends do. notify+wake mode
is unchanged: the text ping is the delivery and the wake stays best-effort
after the cursor advance.
Extracts the residual delivery-ordering insight from closed PR #84191.
Co-authored-by: MaximCrabbe <crabbemaxim@gmail.com>