Commit Graph

28271 Commits

Author SHA1 Message Date
kshitijk4poor c394b005fc fix: publish watchdog settlement only after the abort commits
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.

- Split the surface: `_surface_stall` is now observational only ("no
  progress for Ns; attempting recovery"), and the definitive
  aborted/lease-stopped settlement moves to a new
  `_surface_committed_abort` that runs only after `_commit_abort`
  succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
  turn whose aborts keep declining no longer re-logs an ERROR and
  re-warns the user every poll interval.
- Add the committed-path regression test
  (`test_watchdog_publishes_definitive_settlement_only_after_commit`)
  and extend the declined-path witness
  (`...resumes_during_warning`) to assert no committed-abort or
  definitive pre-commit claim appears when the abort is vetoed. Both
  fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
  unconditionally, no generation claim — losing the lease means the
  process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
  WHY, drop the round numbering), and drop the dead
  `cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.

On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
2026-09-01 03:19:59 +05:30
Finn763 0fe7abe37a fix(agent): surface silent turn stalls with a bounded turn-liveness watchdog (#95548, #95663)
Add a turn-liveness watchdog keyed to the agent activity clock: a turn
that stalls mid-flight while the durable lease keeps renewing is logged
loudly, surfaced to the UI, force-interrupted, and — when the hard
interrupt cannot unwind the wedge — lease renewal is stopped so
stale-turn cleanup can reclaim the session.

Race safety (rounds 3/4/6 of the #95663 review, all folded into this
squashed commit):
- AIAgent.interrupt(require_generation=G) re-validates the generation
  claim at the last instant before the hammer; a stale claim abandons
  the abort and the turn continues.
- The claim is reserved under the activity lock, invalidated by any real
  progress in _touch_activity(), and consumed immediately before the
  first observable interrupt publication; exceptional paths fail closed.
- Claim consumption and the first interrupt publication are atomic
  inside one _liveness_activity_lock() critical section; unclaimed
  interrupts publish lock-free so AIAgent stand-ins without the liveness
  seam keep working (CI 33096454629 regression, fixed here).

Deterministic race regressions (written red-first) in
tests/run_agent/test_turn_liveness_watchdog.py cover the
post-revalidation window, the consume-to-publication window, the
exceptional path, and atomic claim consumption.

Round-7 rebuild: single squashed commit on current origin/main; the
former four-commit lineage (241f8e484..299122558 on merge base
6defe7eb6c) no longer exists, so no surviving commit carries a red
exact-object CI record, and no empty CI-trigger commit was added.
2026-09-01 03:19:59 +05:30
Teknium d95682cd04 fix(gateway): announce declined 'all' scope in /sessions and /resume listings
A non-admin '/sessions all' or '/resume --all' silently downgraded to
chat-scoped listing with zero feedback, which reads as 'my session
vanished' (community Telegram report). Both surfaces now append a notice
that cross-chat listing requires a configured admin. Follows up the
salvaged current-session '(current)' marker (PR #68556, fixes #68547):
sibling tests updated to pin the new contract, new i18n key
gateway.resume.all_requires_admin added to all 17 locales, docs updated.
2026-08-31 14:40:09 -07:00
Xuezhao Lan ca177221b1 chore: map contributor email for PR 68556 2026-08-31 14:40:09 -07:00
Xuezhao Lan dc65b75d73 fix(gateway): mark current session in listings 2026-08-31 14:40:09 -07:00
Teknium 8dbf07e950 fix: satisfy the no-locked-pure-readers gate and close the DB before rmtree
_record_db_file_identity's PRAGMA fallback is a pure read — route it
through _read_ctx() instead of the writer lock (Pattern C gate).
test_codex_turn_persists_each_message_exactly_once leaked a live
SessionDB into shutil.rmtree, racing the WAL sidecars ('Directory not
empty' on CI); close the handle first and rmtree with ignore_errors.
2026-08-31 14:02:55 -07:00
Teknium a9b6b979e9 fix: take the writer lock inside the file-identity stamp helpers
The repo-wide conn-lock audit (tests/test_hermes_state_conn_lock_audit.py)
requires every self._conn statement outside construction to hold
self._lock — an unlocked statement races SessionDB.close() inside
pysqlite's statement cache (#99349). Wrap _ensure_db_file_generation's
statements and _record_db_file_identity's PRAGMA fallback in the lock;
both call sites run outside the lock, so no re-entry.
2026-08-31 14:02:55 -07:00
Teknium 0685bf1992 test: align file-identity guard tests with post-18ac3c4 FTS recovery contract
Main removed the live-path runtime FTS rebuild (18ac3c4fb6) and #99652
narrowed the corruption classifier, after the salvaged PR #89364 branched.
Update its tests to assert the current contract: no fail-open detach on a
replaced file (fts_enabled stays True, no stale marker), fail-open (not
rebuild) for genuine same-file FTS corruption, and FTS-provenance errors
for the guard probe since generic malformed no longer reaches fail-open.
2026-08-31 14:02:55 -07:00
rainbowgits 71256dfd01 fix(state): fail loudly when state.db is replaced under a live process
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:02:55 -07:00
Shannon Sands 5a264f9a58 docs+test: pin lease worst-case trade-off and cover all arm sites (OOF-298 review follow-ups)
Addresses Enough1122's two non-blocking review notes on PR #92316:

1. Lease vs deadline arithmetic: the state.db init/migration/repair leases
   (600-900s) are authoritative against the 300s default deadline by design.
   Add explicit comments at all three lease sites pinning the trade-off:
   single lease is deliberate (clamped to _MAX_LEASE_S=900), honest worst
   case is up to the lease duration of zombie time on a wedged DB phase,
   accepted over per-chunk renewal complexity in the migration loops.

2. Arm-site coverage: add a structural contract test asserting every
   documented entry point (hermes_cli/main.py argv fast-path,
   hermes_cli/gateway.py config-bridge re-arm, gateway/run.py backstop,
   cli.py legacy --gateway) actually calls arm_startup_watchdog (or its
   aliased import), so a future entry point can't silently ship unwatched.
2026-08-31 14:01:39 -07:00
Kshitij Kapoor 44eeef9a83 fix(gateway): apply startup-watchdog config to the already-armed handle
/simplify-code quality+efficiency reviewers (converged, verified): on
the standard 'hermes gateway run' path the argv fast-path arms BEFORE
run_gateway's config bridge executes, and arm_startup_watchdog() is
idempotent — so gateway.startup_watchdog: false and
startup_watchdog_timeout_seconds were dead knobs (env bridged, live
handle untouched). run_gateway now applies the config to the live
handle: disarm on disable; disarm+re-arm on a bridged config timeout so
the fresh handle covers the remaining pre-loop startup with the
configured deadline. config_defaults comment updated to match reality.

E2E (real module, fast-path armed first): disable path disarms the live
handle; timeout path re-arms a fresh handle at 123s.
2026-08-31 14:01:39 -07:00
Kshitij Kapoor d2c3c38e98 fix(gateway): config.yaml surface for the startup watchdog + precise argv arming
Review follow-ups on the salvaged #89750:

- gateway.startup_watchdog / gateway.startup_watchdog_timeout_seconds in
  config_defaults, bridged to the internal HERMES_STARTUP_WATCHDOG env
  vars in run_gateway() (the argv fast-path arms before config can load,
  so env remains the mechanism; config.yaml is the user-facing surface
  per policy — explicit env values still win as operator override).
- hermes_cli/main.py argv sniff now requires the ADJACENT token pair
  'gateway run' instead of independent membership, so unrelated commands
  mentioning both words can't arm a 300s hard-exit timer; profile-flagged
  invocations (-p work gateway run) still arm.
2026-08-31 14:01:39 -07:00
Shannon Sands 852db61abe fix(startup-watchdog): bounded hard-exit escort + phase-owned progress leases
Addresses the two class-level review blockers on PR #89750:

1. Bounded hard-exit seam (escort thread). The forensic fire path
   (logger.critical, dump record, faulthandler, lifecycle ledger) can
   itself wedge — the parked main thread may hold the logging handler
   lock, or the disk may be full/hung. _fire() now starts an exit-escort
   daemon thread BEFORE any forensics; it is free of log handlers,
   filesystem access, module loads and application locks, and hard-exits
   with the restart code after _FIRE_EXIT_BOUND_S unless the normal fire
   path signals completion. Adversarial tests hold the logging handler
   lock / hang the dump write at fire time and assert the exit seam is
   still reached.

2. Phase-owned progress leases (report_startup_progress). Process CPU
   time proves process activity, not startup progress: an unrelated busy
   thread could extend forever while startup sits parked (false
   negative), and I/O-bound repair/backup accrues ~zero CPU and would be
   killed (false positive). Long synchronous startup phases now declare
   authoritative, clamped (_MAX_LEASE_S), renewable progress leases:
   state.db _init_schema + the version-gated data-migration chain
   (hermes_state_schema) and repair_state_db_schema (hermes_state) are
   wired. CPU progress remains only as a bounded fallback, capped at
   _MAX_CPU_EXTENSIONS, with leases outranking the cap. Adversarial
   tests cover both directions (lease saves zero-CPU legitimate work;
   capped CPU noise no longer hides a parked deadlock).

Fire-path dump record now includes lease_count/last_lease_phase for
forensics. gateway/startup_watchdog.py shim re-exports
report_startup_progress.

OOF-298
2026-08-31 14:01:39 -07:00
Shannon Sands f5bb1e144d fix(gateway): address startup-watchdog review findings (OOF-298, PR #89750)
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.

P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.

P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.

P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.

P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.

P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.

Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.

Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
2026-08-31 14:01:39 -07:00
Shannon Sands 8a3b6f374d fix(gateway): startup-liveness watchdog for pre-event-loop deadlocks (OOF-298)
A hosted gateway (hermes-doubleam-2568) deadlocked at startup with every
thread parked in futex_wait_queue before the asyncio loop came alive:
zero log lines, /health unreachable — but s6 saw a live PID so it never
respawned the process, and a stale gateway_state.json from the previous
life told every status surface "draining" for ~30 hours.

Every existing liveness backstop assumes startup succeeded: the
loop-liveness watchdog is armed inside the running loop's startup path,
the shutdown watchdog arms at stop(), and the heartbeat file is written
by an asyncio task. None can fire when the process wedges before the
loop exists.

New gateway/startup_watchdog.py: a plain daemon OS thread armed at
process entry (both gateway.run.main() and the `hermes gateway run` CLI
wrapper), disarmed the moment GatewayRunner confirms a live event loop —
the point where the existing loop-liveness watchdog takes over. If
startup neither reaches that milestone nor exits within the deadline
(default 300s; slowest legitimate pre-loop work is the 120s-bounded MCP
discovery wait), the watchdog:

* dumps all-thread stacks via faulthandler,
* appends a JSON record to logs/gateway-startup-watchdog.log,
* records the exit in the NS-608 lifecycle ledger
  (reason=startup_liveness_watchdog) so the next boot classifies it
  instead of reporting an unclean SIGKILL/OOM death,
* os._exit(75) so s6/systemd respawn the process.

Config is env-only (HERMES_STARTUP_WATCHDOG=0 to disable,
HERMES_STARTUP_WATCHDOG_TIMEOUT_S to tune, floor-clamped to 30s):
the watchdog must be armed before config.yaml is loaded — a wedge
during config parsing is exactly in scope — so it cannot depend on
config for its own enablement. Everything is best-effort; a watchdog
failure never affects the startup it observes.

Arm sites are placed after the PID-file/--replace conflict guards so a
--replace loser exiting early never arms a watchdog. Disarm happens even
when the loop guards are config-disabled (gateway.loop_watchdog: false)
— the startup watchdog only covers the pre-loop window, never adapter
connects or steady-state, so WhatsApp pairing / npm cold installs are
unaffected.

Tests: tests/gateway/test_startup_watchdog.py (29 tests — config
resolution, arm/disarm idempotency, fire path with captured exit,
lifecycle-ledger marking, dump record, disable knob).

Fixes OOF-298.
2026-08-31 14:01:39 -07:00
686f6c61 37fc92a3d6 fix(desktop): keep SSH serve teardown across re-entrant quit
Window X calls app.quit(); backendShutdown.finally() calls it again.
teardownSshConnection deletes the map entry before SSH exec kill, so
the second before-quit saw an empty map and exited while disconnect
was still running. Latch the in-flight teardown so the re-entrant quit
still waits.
2026-08-31 13:59:51 -07:00
Teknium 3783fd9ffe fix(update): defer ledger-verified serve/dashboard holders to the updater instead of dead-ending the Desktop hand-off
On Windows, the Desktop update preflight scan (hermes_cli._scan_venv_blockers)
classified only python -m http.server as a stoppable blocker. A hermes serve
or dashboard backend that survived the Desktop teardown (or was launched
manually) either dead-ended the hand-off with venv-blocked, or slipped past
it and kept venv\Scripts\hermes.exe mapped, so the updater quarantine
failed with os error 32 (#98336).

The CLI updater downstream already owns exactly this holder class with
positive-identity rungs: _ledger_reapable_backend_pids reaps dead-spawner
orphans and _ledger_manual_serve_holders stops manual serves and relaunches
them on their recorded host/port. Mirror the existing pausable-gateway
exemption: defer ledger-verified serve/dashboard holders to those rungs
instead of reporting them as blockers.

Identity is positive-only, per the #99558 guard contract: token-parsed
subcommand (never substring, #90778) + live-verified (pid, create_time)
ledger entry with matching purpose + provable ownership (spawner dead,
unrecorded, or the hand-off Desktop itself — an ancestor of the scan,
verified by pid+create_time). A backend supervised by any other live
process keeps blocking; ledger unreadable fails closed.

Fixes #98336
2026-08-31 13:11:54 -07:00
Teknium fb9b2c893f feat(agent): escalate repeated transcript-sanitiser heals with a one-time user notice (#96870)
Builds the escalation layer on top of HexLab98's heal-log windowing
(salvaged from PR #96916):

- Per-session heal counters (heal events + messages healed) tracked by the
  repair path in agent_runtime_helpers.py, session totals preserved across
  10-minute log windows.
- Threshold escalation: after N heals in a session window (default 3,
  configurable via agent.sanitizer_heal_escalation_threshold in
  config.yaml, 0 = off) log ONE ERROR carrying session id + heal pattern
  (events/messages/window/threshold), then stay quiet.
- ONE-TIME out-of-band user notice queued at the threshold and delivered by
  the conversation loop through _emit_warning (status callback -> gateway
  status message / CLI print). Never injected into conversation context or
  the wire copy: prompt caching, role alternation, and durable history are
  untouched. Never re-arms on a new window; scoped per session.
- Counters visible in diagnostics: get_sanitizer_heal_stats() rendered in
  the /debug share // hermes debug report, and the config key surfaced in
  hermes dump overrides. errors.log carries the ERROR line for `hermes logs
  errors`.
2026-08-31 13:11:41 -07:00
HexLab98 a779f527fa test(agent): cover empty-transcript projection fill and heal-log escalation (#96870) 2026-08-31 13:11:41 -07:00
HexLab98 20fd5d0b25 fix(agent): stop empty-transcript sanitizer from warning on every send (#96870)
Fill empty non-final user/assistant turns on the wire copy during send-time projection so the sanitizer does not re-heal the same poisoned row every call. When a caller still hits the owner, log WARNING then one ERROR per session window instead of flooding errors.log.
2026-08-31 13:11:41 -07:00
theo 8207862212 fix(compression): stop timeout paths from blocking retries 2026-08-31 13:00:33 -07:00
theo 0a8b25e0ab fix(auxiliary): forward service tier on Codex Responses 2026-08-31 13:00:33 -07:00
Teknium 13c3958df5 fix(auxiliary): credit-limited 402s clamp to the affordable budget instead of failing
Second leg of the masoria 20-minute 'Summarizing' stall (Aug 31 2026
bundle): after the Codex timeout, compression fell back to OpenRouter,
which defaulted the omitted output cap to the model's full 65,536-token
window and rejected with '402 ... can only afford 7117' — on an account
whose balance easily covered a summary. Three fallbacks, three 402s,
zero summaries.

- _create_with_progress: when a 402 names an affordable budget, retry
  ONCE with that cap (minus 64-token margin, 512-token floor); plain
  exhaustion 402s and within-budget 402s re-raise unchanged. Single
  funnel covers primary, fallback, and retry call sites.
- extends the #41055 OpenRouter max_tokens preservation onto the current
  _build_call_kwargs gate (explicit caps survive; None still omitted).

Live A/B against a mock OpenRouter enforcing a 7117-token budget with
the real SDK + real adapter: main fails with the exact bundle 402;
fixed branch retries clamped and returns the summary.
2026-08-31 13:00:24 -07:00
liuhao1024 b26a1eae8f fix(auxiliary): preserve max_tokens for OpenRouter to prevent free-tier 402
_build_call_kwargs strips max_tokens for non-Anthropic providers to
avoid wire-format issues (Copilot, ZAI, GPT-5). However, OpenRouter
free/limited-credit tiers need max_tokens because the model's full
output window exceeds the credit budget, causing HTTP 402.

Without max_tokens, the 402 triggers fallback to a text-only model
which then fails with 'unknown variant image_url, expected text'.

Include max_tokens when provider is 'openrouter' or base_url contains
'openrouter.ai'.

Fixes #41035
2026-08-31 13:00:24 -07:00
Teknium 8a766c3f9f fix(auxiliary): stranger-thread timeout also wakes the attempt stream
Socket shutdown() releases readers blocked on a real transport, but the
owner can be blocked inside the SDK event stream (or a transportless
double). Close the attempt-owned stream from the Timer — the same
attempt-scoped wake the hard-cancel branch uses — so the owner unwinds
and performs the real FD release in its finally. Fixes the
test_codex_timeout_and_explicit_cancel_have_one_linearized_outcome red.
2026-08-31 13:00:18 -07:00
Teknium 093d1ebd0b test(auxiliary): watchdog timeout asserts owner-thread FD release
Adapt the #99660 blocked-before-first-event watchdog test to the FD-ownership
contract: the stranger-thread Timer marks the timeout but never close()s;
the owning thread surfaces the TimeoutError and releases the FDs on unwind.
2026-08-31 13:00:18 -07:00
dsad 8a5b49d86a fix(auxiliary): never release Codex client FDs from the timeout Timer
_CodexCompletionsAdapter.create arms a daemon threading.Timer that calls
client.close() when the aux Responses stream exceeds its timeout. On a
stalled stream -- the failure the timeout exists for -- the Timer is the
only thing that fires, so the close runs on a thread that does not own
the in-flight httpx connection.

That is the FD-ownership violation the repo already fixed twice on the
main transport (#29507, #67142, #70773): close() releases the raw TLS fd
while the owner's OpenSSL BIO still caches that integer, the kernel
recycles it into the next open() in the process -- a SessionDB or
kanban.db handle -- and the owner's unwinding TLS flush writes an
application-data record into that database file.

agent/auxiliary_client.py had no thread-ownership machinery at all: the
guarded twins (_retire_shared_openai_client, _abort_request_openai_client)
live in run_agent.py and are unreachable from this adapter, which holds no
AIAgent reference.

Dispatch on ownership the way chat_completion_helpers already does: from a
stranger thread only force_close_tcp_sockets() (shutdown(SHUT_RDWR), which
is FD-safe from any thread), and let the owning thread release the FDs when
it unwinds. The owner-thread caller (_check_cancelled) keeps closing
directly. Cache eviction (#23432) is unchanged.
2026-08-31 13:00:18 -07:00
Teknium 0f16e413e6 fix(update): run pending fleet-restart catchup before the runtime-verification exit gate
The #95294/#91277 fleet contract requires the pending-restart check to
execute on every already-up-to-date pass; the verification exit(1) now
fires only after the catchup runs, so a vulnerable runtime demotes the
outcome to partial without stranding the fleet on stale code.
2026-08-31 12:58:14 -07:00
Teknium d6bf89de40 fix(update): treat an unprobeable post-update interpreter as non-blocking
Verification-features-must-not-false-positive-on-rollout rule: only a
POSITIVE vulnerable SQLite probe demotes success to partial. A dev
checkout without a venv (or a failed probe subprocess) keeps the
success banner and the fleet-restart catchup path — the CI-red
test_update_fleet_restart_pending already-up-to-date trio pinned this.
2026-08-31 12:58:14 -07:00
fangliquanflq 9caf31394f fix(update): verify repaired current-checkout runtime 2026-08-31 12:58:14 -07:00
fangliquanflq 4f19808915 fix(update): fail current-checkout runtime verification 2026-08-31 12:58:14 -07:00
fangliquanflq 7e22725bb4 fix(update): verify SQLite runtime remediation 2026-08-31 12:58:14 -07:00
Teknium 18c9048070 ci(wine2e): include TestWindowsRuntimeSelfLock in the on-demand windows-latest lane 2026-08-31 12:50:29 -07:00
Finn763 57f653db90 test(managed_uv): make self-lock regression assert carry its message
The previous form was `mock_install.assert_not_called(), ("msg")` — a bare
tuple expression whose parenthetical never surfaces as a failure message.
Switch to `assert mock_install.call_count == 0, "msg"` so the diagnostic
actually appears when the guard regresses (review feedback on #93163).
2026-08-31 12:50:29 -07:00
Finn763 d0503b500f fix(update): defer Windows runtime repair when the updater holds the venv
On Windows, hermes update launched from the install's own venv can never
complete the managed-runtime repair: Windows keeps the image of the
updater's venv\Scripts\python.exe (and the waiting hermes.exe launcher
ancestor) mapped until exit, so the park rename in _cut_over_candidate
always fails with ERROR_ACCESS_DENIED. The pre-flight holder scan
deliberately excludes the calling process and its ancestors, so the guard
passes and the repair burns its retries on a structurally unwinnable
rename - forever, since the failure is non-fatal and the 'next update
will retry' message is misleading for this case.

Detect the self-lock (sys.executable or a launcher ancestor inside the
live venv) before provisioning and defer with actionable guidance instead
of walking into the doomed cutover. The deferral happens pre-provisioning,
so the incomplete generation-* leftovers the reporter observed are no
longer produced. No-op off Windows: POSIX renames work while the updater
maps the tree. Mirrors the existing _defer_update_for_self_lock pattern.

Regression tests prove the fix bites: neutralized guard -> repair proceeds
to provisioning (red); restored guard -> deferred before provisioning
(green). Verified on Windows 11 against a real venv.

Closes #93032
2026-08-31 12:50:29 -07:00
Teknium 7beb0de676 fix: follow-up for salvaged PR #71904 — source raw inbound id, strip persistence key from the wire
- Thread event.message_id (raw inbound id) as TurnContext.inbound_message_id
  instead of reusing event_message_id, which is the reply/thread anchor and
  can be the replied-to message on Slack/Mattermost/Buzz or None for
  Telegram topics.
- Add platform_message_id to the schema-foreign strip sets in
  ChatCompletionsTransport.convert_messages and the summary path so strict
  providers never see the persistence-only key.
2026-08-31 12:46:32 -07:00
Kyzcreig c1bd0511cf fix(gateway): persist the platform message id on every user turn 2026-08-31 12:46:32 -07:00
Teknium 5505042f40 fix(sessions): fail closed on ownership uncertainty and fence every turn source
Follow-ups on top of the #94595 cherry-pick, implementing the maintainer
review's two blockers:

Blocker 1 (turn-admission chokepoint): _run_prompt_submit itself now runs
the ownership admission, so synthesized turns that never pass through the
prompt.submit RPC handler (crash auto-continue from cold session.resume,
wake-ups) are fenced too. Auto-continue additionally checks ownership
BEFORE emitting message.start and leaves the marker in place, closing the
#94778 shape where backend B resumed a session backend A was actively
running and auto-continued A's fresh interrupted-turn marker into a
duplicate concurrent turn.

Blocker 2 (fail-closed registry semantics): try_acquire_active_session no
longer converts an unreadable/corrupt registry into an untracked go-ahead.
Ownership uncertainty is a distinct typed refusal —
SESSION_COORDINATION_UNAVAILABLE — because "could not prove ownership" must
never be collapsed into "no owner exists". The TUI gateway claim helper
fails closed on claim exceptions for every surface, not just desktop.

Also: empty session ids short-circuit to a no-op lease (nothing to fence,
and the strict registry schema rejects empty ids), and the existing
fail-open tests were updated to assert the new fail-closed contract.
2026-08-31 12:36:33 -07:00
Futahua 58c1876d37 test(sessions): falsify the per-session fence across real processes
The unit tests share one interpreter, so they cannot exercise the failure the
fence exists to prevent: two SEPARATE gateway processes, each holding its own
snapshot of a conversation, both writing to it. That is how the defect was found
and it is the only way to show it is closed.

This drives two real `python -m tui_gateway.entry` processes over stdio and
checks the whole sequence, including the parts that are easy to get wrong:

  session.create claims nothing        an idle composer must not hold a session
  the lease keys on the STORED id      a lease keyed on the runtime handle would
                                       fence nothing, since two processes
                                       resuming one conversation have different
                                       runtime ids by construction
  B may still RESUME                   reading is never fenced; only writing is
  B's submit -> SESSION_NOT_OWNED      typed, and the registry is unchanged
  A killed, B retries -> accepted      a dead owner is pruned, not permanent

No provider is needed. The fence is checked before the agent is built, so a
submit that later fails for want of a model still proves who owns the session --
which keeps the probe free of credentials and of inference cost.

Against the parent commit it stops at the second check with an empty registry,
which is the defect stated exactly: with no cap configured, nothing was recorded
and therefore nothing could be refused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 12:36:33 -07:00
Futahua a5f0fbb262 fix(sessions): per-session exclusivity is correctness, not a capacity policy
Cherry-picked from PR #94595 (author: Futahua) onto current main, with the
maintainer-review revision points folded in during the rebase:

- the lease engages UNCONDITIONALLY: try_acquire_active_session no longer
  returns a disabled no-op lease when max_concurrent_sessions is unset;
  the concurrency cap stays an orthogonal, optional policy checked second
- ownership uncertainty fails CLOSED (SESSION_COORDINATION_UNAVAILABLE)
  instead of degrading to an untracked go-ahead: a corrupt/unreadable
  registry must not be collapsed into 'no owner exists' (review blocker 2)
- the ownership admission sits at the _run_prompt_submit chokepoint that
  EVERY fresh turn source crosses, and crash auto-continue acquires (or
  bails) BEFORE emitting message.start — closing the #94778 bypass where
  backend B's auto-continue ran a duplicate turn while backend A was live
  (review blocker 1)
- the TUI gateway claim helper fails closed on claim exceptions for every
  surface, not just desktop
- CLI and messaging-gateway call sites pass live_session_id metadata so
  the (pid, live id) re-entrancy identity protects them from self-fencing
  on a leaked lease

Co-authored-by: teknium1 <teknium1@users.noreply.github.com>
2026-08-31 12:36:33 -07:00
fangliquanflq 53c0df6de9 fix(agent): stop compression retries after host timeout (#98722)
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:

- conversation_compression.py: threading.local + per-agent lock timeout
  state (mark/reset/read helpers) upgrading #98424's simple attribute
  where overlapping automatic/manual compression entrypoints matter;
  the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
  overflow (413/400 context_length_exceeded) recovery path end the turn
  with the typed compression_exhausted recovery contract instead of
  re-sending the unchanged oversized request and re-entering compression
  in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
  attempt; the #98424 turn-start check reads it through the typed helper.

Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
2026-08-31 12:36:02 -07:00
Dev 532ee1a199 fix(gateway): persist gateway routing identity on lazily-created session rows
When the default/global state.db is corrupt at gateway startup,
SessionStore degrades (_db=None) and record_gateway_session_peer never
self-heals. Under multiplexed profile routes the AIAgent's lazy
_ensure_db_session was then the ONLY durable write for the session, and
it created the row identity-less (session_key/chat_id/chat_type/
thread_id/user_id/origin_json all NULL) — unrecoverable by
find_latest_gateway_session_for_peer, so Telegram chats forgot prior
turns.

Persist the routing identity the agent already carries into
create_session; plain CLI sessions keep the old identity-less shape.

Salvaged leg 1 of #88804; the transcript-recovery leg is covered by the
scope-aware session DB resolution already on main.
2026-08-31 12:32:33 -07:00
Teknium 29112bef09 chore: release v0.21.0 (2026.8.31) 2026-08-31 12:29:27 -07:00
Teknium 7790c8f4c4 fix(telegram): bound stale-client cleanup and add Windows CLOSE-WAIT live probes (#87057)
Follow-ups on top of the salvaged commits from PR #87111 (@HexLab98) and
PR #87265 (@JoaoMarcos44):

- keep main's #92991 stall watchdog (150s progress-based) as the single
  steady-state liveness probe instead of adding a second overlapping one
- orphaned-client aclose() cleanup uses the wall-clock thread deadline and
  is tracked in _background_tasks so a wedged close can neither hang nor
  leak one task per reconnect attempt (from #87265's review findings)
- merge #87265's no-keepalive getUpdates pool (max_keepalive_connections=0)
  with #87111's TCP-keepalive socket options on all transports
- add tests/gateway/test_telegram_closewait_windows_live.py: live probes
  against a real half-closing HTTP server, skipif non-win32, wired into
  the on-demand windows-venv-e2e lane (wine2e/**)
2026-08-31 12:28:49 -07:00
joaomarcos b02b7122f8 fix(telegram): prevent Windows long-poll socket reuse deadlock
Prevent the dedicated getUpdates pool from reusing server-closed connections and replace a polling HTTP client left open after a timed-out CLOSE-WAIT drain. Keep the general Bot API pool reusable so concurrent sends and edits are unaffected. Add regression coverage for both transport limits and stale-client replacement. Fixes #87057
2026-08-31 12:28:49 -07:00
HexLab98 505f64b798 test(telegram): cover CLOSE-WAIT drain rebuild and getUpdates liveness 2026-08-31 12:28:49 -07:00
HexLab98 a06c0d0a2f fix(telegram): recover Windows CLOSE-WAIT getUpdates deadlock
After updater.stop() times out, HTTPXRequest.initialize() is a no-op unless
the client is already closed, so start_polling reused the wedged socket and
the gateway stayed alive but deaf. Rebuild the polling client after a hung
drain, watch getUpdates I/O independently of get_me(), and enable TCP
keepalive on the fallback transport.
2026-08-31 12:28:49 -07:00
Teknium 7cefa87ea7 fix(agent_init): reserve Gemini's default maxOutputTokens in the compressor when max_tokens is unset
The native generateContent adapter never runs uncapped: when
model.max_tokens is unset it sends maxOutputTokens=65,535
(GEMINI_DEFAULT_MAX_OUTPUT_TOKENS) because Gemini treats an omitted cap
as a low internal default. The context compressor's trigger is
pct×(window − max_tokens), and constructing it with max_tokens=None
reserved 0 — so on a 128K Gemma window the trigger landed at 98,304
while the real safe input budget was 65,537, and the provider 400'd
before compaction fired.

Live repro (real imports, temp HERMES_HOME, native Gemini base_url,
window=131072, max_tokens unset):
  before: compressor.max_tokens=None, threshold_tokens=98304,
          wire maxOutputTokens=65535 → trigger ABOVE the safe budget
  after:  compressor.max_tokens=65535, threshold_tokens=64000 → below it

Scoped to the native Gemini wiring (provider names + native base_url via
is_native_gemini_base_url; the /openai compat endpoint is excluded). The
generic provider-default reservation gap remains tracked in #63839.

Reported by @Artemonim in #57275 (residual claim 4).
2026-08-31 12:22:55 -07:00
Teknium cb71d5f1b1 fix(agent_init): clamp compressor window to Ollama num_ctx resolved after construction
model.ollama_num_ctx is resolved AFTER the context compressor is
constructed, so a config that sets only ollama_num_ctx (without
model.context_length) ran every request at the smaller served num_ctx
while the compressor still targeted the probed GGUF window (e.g. 256K
Gemma metadata). The compaction trigger then sat several times above the
window the server actually serves and never fired — reproducing the
original #57275 'blows past the limit' symptom on current main.

Live repro (real imports, temp HERMES_HOME, config = {model:
{ollama_num_ctx: 65536}}, probed window 262144):
  before: _ollama_num_ctx=65536, compressor.context_length=262144,
          threshold_tokens=196608 (300% of the served window)
  after:  compressor.context_length=65536, threshold below the window

The clamp is one-directional (a num_ctx larger than the resolved window
never inflates the compressor) and reuses update_model() so every
threshold-derived budget recalibrates. Overlaps #60103 (silent-clamp
dead zone) — this is the init-order half.

Reported by @Artemonim in #57275 (residual claim 3).
2026-08-31 12:22:17 -07:00
Teknium 7acc65a399 test(windows): live E2E for the git trampoline self-heal on the wine2e lane
Real windows-latest coverage for the #88136 salvage: probes drive the
actual _git_is_trampoline/_locate_real_git/_ensure_non_trampoline_git
helpers against the runner's genuine Git-for-Windows install plus a real
fork-bomb-guard trampoline stand-in. Wired into the on-demand
windows-venv-e2e lane (wine2e/** pushes only).
2026-08-31 12:21:46 -07:00