Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
Carve the history-only implementation and regression from #104754
(b4bfa76facb43111a863b1256c89875e3f01fde0) by PLASMA-FR; omit
the unrelated locale-picker changes. Complement merged #104523.
Real serve + WebSocket reconnect: SQLite retains two rows; before,
resume/activate/history return only the user; after, both survive.
Plain-content control retains both rows in both arms. Native macOS
sleep and the full desktop symptom are not established by this probe.
Refs #68321
Wire the existing consent-aware, profile-keyed hook registrar at agent
construction, where the correct session home is already bound. This covers
serve and TUI agent construction without a startup-only registration or a
new helper that swallows registration failures.
Live isolated serve/WebSocket probes reproduce the missing registration on
base for both write_file and terminal, then verify each consented profile
blocks its configured tool while the unapproved profile still runs normally.
Repeated alpha construction does not duplicate hook callbacks.
Slim implementation of the agent-build placement proposed in #57020;
thanks also to the profile-scoped analysis in #102691.
Co-authored-by: grimmjoww578 <willies578@gmail.com>
The real closed-WebSocket resident variant still wedges after the missing-timer repair. Re-enter existing transport cleanup before rearming orphan timers, preserving viewer transfer and the reconnect/delegation fence rather than deleting registry rows from a stale snapshot. Extend the same invariant to both dead-transport shapes. Live class investigation informed by #104710; no direct-vouch reclaim machinery imported.
Salvage #104704 (e9423d2d0bbe3e795c5eaccb86a913f1d95444ba, 5fa2b98fc833fc4e1ee7f1aaa7eb45cdd3cd7cde). Reuse guarded orphan teardown instead of deleting ownership fences. Real two-backend WebSocket probe reproduces the missing-timer wedge on base and proves reconnect/delegation protection and recovery. Add reusable probe and user documentation.
Salvage #101453 (03a3f466d38134ba416764185884b3d655197a1d). Preserve its opt-out and first-run behavior; replace predicate-mocked tests with one native config/filesystem invariant and clarify XDG docs. Real venv/XDG probe: base clobbers custom entry, fix preserves it; targeted suite 94 passed.
_sync_bot_capabilities (Bot Chat capability refresh, run at every turn start) and
_reset_session_agent (/new, tools.set) swap a fresh AIAgent into a LIVE session but
called _make_agent with neither the session's state.db handle nor its HERMES_HOME.
_make_agent resolves prompt/skills/toolsets through get_hermes_home() and defaults
session_db to the process-wide LAUNCH _get_db() handle, so a named-profile session
silently migrated onto the launch profile: every later turn appended to
~/.hermes/state.db under the same session id while the desktop replayed
profiles/<name>/state.db and showed a stale transcript.
Route both rebuilds through _rebuild_session_agent, which binds the session's
profile scopes and inherits the outgoing agent's handle (same session, same file)
so no second refcount is taken and teardown still releases it exactly once.
Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.
Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.
Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.
Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
The redelivery hook keyed on the canonical flood_control:<seconds> result, so
two real refusals slipped past it and armed no timer, leaving the reply for the
next restart. A short wait that outlived the send retries raised instead of
failing closed, and an edit refused again after its inline wait returned the
platform's raw text. Both now fail closed canonically, the second carrying the
new delay rather than the first refusal's.
The ledger also accepts a row still carrying the platform's own wording, so a
row persisted by an unnormalized path is dated from the delay it states instead
of the generic default. Without that a boot sweep claims it at once and spends
its one attempt inside the penalty. Matching requires the flood wording as well
as a delay, so an unrelated retry suggestion is never read as a flood.
Six new assertions fail without this change. 770 passed across the ledger,
Telegram, send-retry and queued suites.
Round-2 review. Making _arm_flood_timers_for_waiting_rows read the ledger in a
worker thread introduced a race: a flood refusal recorded during that thread's
yield had its own _schedule_flood_redelivery request declined (the timer still
held the slot), and the timer then cleared its slot on a snapshot taken before
the new row committed, leaving that reply with no timer until the next reconnect
or restart. The read is a single indexed SELECT; doing it synchronously keeps the
arm decision atomic with respect to concurrent schedule requests. A test drives a
refusal during the timer's own redelivery send and asserts the same timer arms it.
Second review pass on the flood-retry change.
- sweep_recoverable returns a dead owner's not-yet-due flood row flagged
`adopted` (with its `not_before`) instead of dropping it from the result,
so _claim_pending_obligations clears its session's resume_pending flag like
every other claimed row; the answer is in the ledger and the turn must not
be re-run at boot. _redeliver_claimed_obligations skips adopted rows and
still arms the timer for them.
- A legacy row without adapter_profile is normalised to 'default' when the
boot sweep claims or adopts it: the runtime sweep matches profiles exactly
and the timer asks for 'default', so an adopted NULL row could only wake the
timer without ever being sent.
- A sleeping timer is replaced by a shorter refusal's timer; the shorter one
re-arms for the longer row after its sweep. A timer already sweeping is
never cancelled from outside.
- pending_flood_retries is read in a worker thread like the other ledger calls.
- Adoption tests seed a distinct dead-owner stamp and assert ownership moves.
- Scratch tags and long comment lines removed.
Review on the first cut reproduced two gaps.
The raw UTF-16 length of the stored reply said nothing about how many
requests the adapter made: MarkdownV2 escaping turns 3000 dots into 6000
units, two Telegram messages, and the send result does not say which chunk
the platform refused. Drop the length-based certainty (and its helper and
constant) and mark every flood redelivery with the rate-limit marker, at
runtime and at boot. The cost is a marker on a reply nothing of which was
delivered; the alternative was a silent duplicate.
Releasing an unsent runtime claim always wrote send_path_degraded, so a flood
row whose resume flag could not be cleared (or whose adapter vanished before
dispatch) left the flood timer's list and stayed stranded until a reconnect or
restart. The claimed row now carries its pre-claim error and the release
writes it back, so the row waits the platform's figure once more and is sent
on the next timer.
The Telegram adapter fails a flood-controlled final send closed as
flood_control:<seconds> so the send coroutine never sleeps a long penalty
(#91969), on the understanding that the delivery ledger owns the wait. The
ledger did not: sweep_failed_for_runtime only replayed send_path_degraded
rows, so a flood-refused row sat in 'failed' until the next restart, whose
sweep_recoverable then redelivered it hours late prefixed with "Recovered
reply, the gateway restarted during delivery, so this may be a duplicate".
Observed: a reply refused at 08:58 UTC (both the MarkdownV2 send and the
plain fallback got flood_control:185) arrived at 12:15 UTC after a restart,
labelled as a possible duplicate although the platform had never accepted it.
Ledger (gateway/delivery_ledger.py):
- flood_control rows are runtime-retryable, but only once their own deadline
has passed: the refusal's updated_at plus the platform's wait
(flood_not_before). Neither an early timer nor a reconnect sweep spends a
redelivery attempt inside the penalty window.
- A flood refusal of a reply that fits in one Telegram message (4096 UTF-16
units) proves non-delivery, so that redelivery carries no duplicate marker.
A chunked reply may have had its first chunk accepted before the refusal
and keeps the marker.
- Claiming a flood row clears the stale refusal (last_error NULL, state
'attempting'), so a resend interrupted before mark_delivered is seen as
uncertain by the next boot and gets the marker.
- At boot, a dead owner's flood row that is not yet due is adopted (owner
re-stamped, no attempt spent) instead of being resent early.
- pending_flood_retries() lists this process's waiting flood rows per adapter
identity with the earliest deadline.
Runner (gateway/run_startup.py):
- _schedule_flood_redelivery arms one timer per adapter identity that runs
the existing runtime sweep after the wait (capped at 15 minutes per sleep;
the row's deadline, not the timer, decides eligibility, so a capped timer
wakes early, sends nothing, and re-arms for the remainder). The slot stays
occupied until the timer ends and only the running timer may arm its
successor into it, so a refusal during the sweep can never strand the row.
- _arm_flood_timers_for_waiting_rows runs after every redelivery pass (boot
and runtime), covering adopted rows, rows skipped as not yet due, and rows
refused again.
Adapter (gateway/platforms/base.py): _finalize_delivery_obligation arms the
timer on a flood_control failure, best-effort inside the existing try.
Tests: 35 in tests/gateway/test_delivery_ledger_flood_retry.py, with a
controllable clock shared by ledger and runner. Each guard was mutation
tested. Existing ledger, reconnect and redelivery suites (131 tests) pass.
Keep two invariants covering empty create/cwd/save/fork, genuine content and existing-row metadata. Preserve original authorship and avoid source-only legacy pruning: an empty ACP row does not prove its owner is dead. Native ACP wire plus a local streaming model fixture verifies the first turn and nonempty fork remain durable.
Fixes#104724
The ACP wire contract has session/new as a separate round trip from
session/prompt precisely so a client can open a session before it knows a
prompt is coming, and at least one shipping client opens sessions it will
never prompt: bb discovers a model catalog by spawning a throwaway
`hermes acp`, sending initialize + session/new, reading
NewSessionResponse.models, and killing the process. Its discovery cache
TTL is 60s, so an open editor re-probes continuously.
_persist() created the state.db row the moment a session was created, so
every such probe left a permanent message_count=0 shell. Measured on one
workstation: 116 accumulated, appearing on a time cadence rather than a
per-conversation one (26 real ACP user turns vs 13 shells in a day; exactly
120.0-minute spacing overnight with zero user activity).
The shells are indistinguishable from real chats in the session list and
`hermes sessions prune` cannot remove them: the rows are never ended, so
ended_at stays NULL and prune skips them by design, leaving
`hermes sessions delete <id>` one id at a time as the only remedy.
Defer row creation until the session has history. Nothing is lost for a
genuine conversation: AIAgent._ensure_db_session() creates the row on the
first turn and the post-prompt save_session() lands the ACP metadata on
top, which makes the create-time write redundant for every session except
the one case that should not be recorded at all.
Gate on state.history rather than message_count so fork_session, which
deep-copies a non-empty history into a fresh id, still persists at once.
Verified by differential execution of an identical probe against both
trees: on 693641aa8b an unprompted session/new adds a row, with this
change it adds none, and a session that does receive a prompt persists
in both. tests/acp: 138 passed. Reverting only the source change leaves
the two new assertions failing, confirming they exercise the fix.
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.
Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.
Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.
Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
Slim rework of despotak's modified-keypad fix in #97290. Mirror existing
non-keypad mappings for modified keypad keys, including lock-state variants,
so Alt+keypad Enter reaches the existing newline handler rather than leaking
[57414;3u into the draft. Preserve installed twin mappings before consulting
pending aliases, matching first-writer-wins registration.
Replace the source PR's keyed branch ladder with a format table and verify
parser parity plus real buffer insertion with two invariant tests. Document
keypad multiline support in English and Chinese.
Live PTY: the exact doubled leak after a real collapsed paste reproduces on
main; all 21 editor cases pass with the fix, including ordinary Enter and
legacy Alt+Enter controls. Whitespace also reproduces on main: adjacent
characters are not the root cause.
Co-authored-by: Christos Despotakis <christos@despotak.is>