Commit Graph

14240 Commits

Author SHA1 Message Date
joaomarcos 45bb486b26 fix(gateway): give in-flight cron work its own drain floor
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.

A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.

Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.

The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.

Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.

Relates to #82161 (complements #82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
2026-08-14 21:47:16 -07:00
Jeremy 1f6f86119f fix(cli): stop hermes update from respawning orphan serve --port 0 (#78821)
Filter manual dashboard/serve respawn candidates after update: skip
ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines,
and cap one restart per profile/HERMES_HOME so orphan counts no longer
grow across successive updates.
2026-08-14 21:46:32 -07:00
joaomarcos 69d1843512 fix(packaging): ship bundled plugin manifests 2026-08-14 21:45:08 -07:00
Teknium 0dba3316b2 fix(gateway): generalize supervised-gateway exemption in orphan reaper to all platforms
Compose the service-PID exclusion (#85743, RelaxJonh) and the recorded-PID +
parent-chain exemption (#86100, arccat-114) into one cross-platform rule:

- _get_service_pids() exclusion now runs unconditionally, not only under
  is_macos() — it is the authoritative "supervised" signal for launchd and
  any systemd unit visible on a host that got past the systemd gate.
- The recorded-healthy-gateway (get_running_pid()) + parent-chain exemption
  now runs on every platform, not only Windows. A recorded, liveness-verified
  gateway is by definition not an orphan "the pidfile/runtime record can't
  see", so the reaper must never target it — this covers Windows Scheduled
  Task / Startup VBS supervision, standalone launcher-started gateways
  (the case #85743 alone would miss), and macOS/WSL equivalents.

True orphans (no service registration, no valid runtime record) are still
found and reaped, preserving the #51325/#75936 duplicate-port protection.

Existing macOS regression tests updated to pin get_running_pid to None for
their scenario; Windows regression tests from #86100 carry over unchanged.

Bug class: #83683 (root), #86287, #86098, #85738, #85368, #85344, #85044,
#84855, #84824, #84200.
2026-08-14 21:44:28 -07:00
arccat-114 102369c5f6 fix(gateway): spare Scheduled-Task-supervised gateway from orphan reaper on Windows
The orphan reaper kills a healthy gateway (and its Scheduled-Task bootstrap
parent chain) every time the Desktop backend starts on Windows, because
_get_service_pids() only implements systemd/launchd and returns an empty
set on Windows — a supervised gateway is therefore indistinguishable from
an unsupervised orphan.

Exempt the recorded healthy gateway PID and its parent chain from the
orphan scan on Windows, mirroring the macOS launchd exemption (#85913).
The Scheduled-Task bootstrap's argv matches the gateway scan, so without
exempting the parent chain killing the bootstrap takes the detached
gateway down with it.

Fixes #86098
2026-08-14 21:44:28 -07:00
worlldz 547043a4d8 fix(telegram): honor fallback disable during connect 2026-08-14 21:43:06 -07:00
Teknium dac3c44afc test: fix salvage test imports; drop WAL worker-thread test superseded by read pool
- tests/cron/test_sessiondb_init_hang.py: add threading/time imports the
  salvaged late-close regression tests rely on.
- tests/test_hermes_state.py: drop
  test_close_closes_wal_read_connection_created_on_worker_thread — main
  replaced per-thread WAL reader ownership with the pooled read-connection
  design (permits + checkout/return), so cross-thread reader draining no
  longer exists in the form the test asserted.
2026-08-14 21:41:26 -07:00
Tranquil-Flow 38709ae6f2 fix(cron): close leaked SessionDB connection when init outlives the timeout-abandoned worker (#72782)
run_job() submits SessionDB() to a one-worker executor and abandons the
worker (shutdown(wait=False)) when init exceeds the cron timeout. If the
constructor later completes inside that abandoned worker, the Future's
result — an open SessionDB holding .db/WAL/SHM handles — was orphaned and
never closed, leaking descriptors until EMFILE. Attach a done-callback on
the timeout path that retrieves and closes any eventual late result.

Salvage note: the lazy-recall ownership half of #72822 (_owns_session_db
tracked on AIAgent, owned handle closed in close()) already landed on main;
this carries the remaining cron timeout-abandon half with its regression
test.
2026-08-14 21:41:26 -07:00
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
Teknium 8dc5608a78 fix(compression): adopt live continuation tip at flush across multi-hop chains
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.

- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
  tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
  adopt only when tip != old_id AND the tip row is live, retry the flush
  exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
  lookup with the same tip + liveness contract, so gateway transcript
  reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
  recovery preflight now resolves via get_compression_tip with the same
  liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
  explanation names compression rotation and tells the client to refresh the
  session id instead of blaming a full disk.

Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).

Closes #82001

Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-14 21:39:44 -07:00
kshitij 70704b962f fix(delegate): child's dedicated SessionDB must follow the parent's db_path
A bare SessionDB() resolves the launch profile's default state.db, but
parents can hold non-default per-profile handles (tui_gateway opens
SessionDB(db_path=<profile_home>/state.db) for non-launch profiles and
hands them to agents via _transfer_db_to_agent). A child of such a
parent would write its transcript into the WRONG database — cross-
profile leakage that breaks parent_session_id lineage and
session_search. Open the dedicated handle at the parent handle's
db_path instead (AsyncSessionDB forwards .db_path via __getattr__, so
the gateway wrapper path works too). Regression test verified RED on
the pre-fix code.
2026-08-14 21:39:31 -07:00
thatssoheil eef43e3160 fix(delegate): close the dedicated SessionDB if child construction fails; test degradation
Review follow-up (cc3f18197): if AIAgent() raises inside _build_child_agent
the freshly-opened dedicated handle has no owner and no child close() will
ever run — release it on the exception path so the sqlite fds don't
outlive the failed spawn. Also pin the degradation contract with a test:
a parent without a SessionDB still yields session_db=None children.
2026-08-14 21:39:31 -07:00
thatssoheil 65e005d0e7 fix(delegate): subagents get a dedicated SessionDB, not the parent's (#81267)
Cron run_job closes its per-job SessionDB in its finally block while a
fire-and-forget background delegation subagent is still flushing on a
daemon thread. The child shared the parent's SessionDB object, so every
subsequent flush hit the closed handle ('NoneType' object has no
attribute 'execute') and the child's whole transcript was silently
dropped. The same teardown-while-child-alive shape exists on gateway
session end and /new mid-delegation.

Each child now opens its own SessionDB connection (owned flag set at
construction so child.close() releases it), so no parent teardown can
close the child's handle out from under it.

Regression test proves the child gets a distinct live handle that
survives the parent's close().
2026-08-14 21:39:31 -07:00
Teknium 480342232a fix(gateway): close leaked poller sockets in weixin/email adapters (#79889)
On macOS (256 soft fd limit), routing the weixin/email pollers through a
local HTTP proxy leaked one TCP socket per failed poll/connect cycle
until the gateway hit `[Errno 24] Too many open files` and crashed
(launchd respawn loop). Live capture showed 216 of 256 fds pinned on
connections to the proxy, ~214 of them abandoned.

Code-side gaps fixed:

- email adapter, `connect()`: no try/finally around the IMAP test
  connection — a failure in login/ID/select/search abandoned the
  connected socket with no owner. Every reconnect-watcher retry builds
  a fresh adapter, so each retry against an unreachable/proxied host
  leaked another fd. Teardown now runs in `finally`.
- email adapter, IMAP teardown: `imaplib.IMAP4.logout()` only swallows
  `OSError` internally; on a broken connection `LOGOUT` raises
  `IMAP4.abort` before the internal `shutdown()`, leaving the socket
  open. New `_close_imap()` helper chases a failed `logout()` with an
  unconditional `shutdown()`; used in `connect()` and
  `_fetch_new_messages()`.
- weixin adapter: repeated poll failures through a proxy strand
  sockets in the aiohttp connector where the tight keepalive reaper
  never sees them. The poll loop now recycles its ClientSession
  (swap-then-close, safe for concurrent `_process_message` tasks)
  after each MAX_CONSECUTIVE_FAILURES streak, tearing down the
  connector and every socket it holds.

Targeted tests: tests/gateway/test_poller_fd_lifecycle.py (9 tests).

Reported by @EthanHunter1229 with measured fd captures.
2026-08-14 21:38:54 -07:00
burak33bb 453e6d8b95 fix(moa): preserve facade across client rebuilds 2026-08-14 21:36:41 -07:00
RelaxJonh 6bf3d39499 fix(agent): strip _moa_prepared_request before dispatching to native client
After a client replacement (credential rotation, dead-connection cleanup,
or fallback+restore), agent.client may become a native OpenAI client
while agent.provider stays "moa".  The _moa_prepared_request key was
passed through to the native SDK, causing TypeError on every turn.

Pop the key at the dispatch point (chat_completion_helpers.py:509).
The MoAClient facade already handles a missing key by falling through
to its normal resolution path.

Closes #78382
2026-08-14 21:36:41 -07:00
Drexuxux ab879a1e22 fix(agent): stop the MoA prepared request reaching a swapped-in native client
`_moa_prepared_request` is a private handshake between the conversation
loop and MoAChatCompletions.create. It is attached whenever
agent.provider == "moa", on the assumption that agent.client is still the
in-process MoA facade.

Credential rotation, provider fallback and dead-connection cleanup all
rebuild agent.client from _client_kwargs between attempts, and
pending_moa_prepared_request deliberately carries a prepared request
across exactly that boundary. The rebuilt client is a native OpenAI
client while provider stays "moa", so the key reaches an SDK that has
never heard of it:

    TypeError: Completions.create() got an unexpected keyword argument
    '_moa_prepared_request'

That error is non-retryable, so every remaining turn on the session
fails. Both dispatch paths are affected: the non-streaming one calls
agent.client directly, and _create_request_openai_client returns
agent.client unchanged for provider "moa".

Re-check the live client at the point the key is attached, which covers
both paths at once. When the facade is gone, send the prepared prompt
without the handshake and log the downgrade.
2026-08-14 21:36:41 -07:00
Teknium f45813ea77 fix(sessions): run state.db schema migration eagerly at backend startup and stop swallowing locked ALTERs
After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).

Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):

1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
   writable open, typically the user's first NEW session. The dashboard
   backend now schedules one writable open of its own state.db from the
   lifespan (daemon thread, never blocks the ready-probe socket, never
   raises), so the store is brought current before the first session-
   list poll on every `hermes serve` / `hermes dashboard` / Desktop
   headless entrypoint.

2. _reconcile_columns caught sqlite3.OperationalError around every
   ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
   orphaned sibling backends made the ALTER fail silently — startup
   "succeeded" with a half-reconciled schema, and the open-time lock
   patience (#74478) never saw the error because it was swallowed
   inside first. Now: "duplicate column" races stay at DEBUG,
   locked/busy re-raises so _connect_and_init_with_lock_patience
   retries the whole idempotent init with jittered backoff, and any
   other failure (e.g. un-ADDable NOT NULL) logs at WARNING.

Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.

Fixes #79531
Fixes #80037

Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
2026-08-14 21:36:29 -07:00
David Metcalfe 542f180e7a test: pin Test-Node managed-Node swap to same-directory renames
The Test-Node stage-and-swap relies on Rename-Item's same-directory
carve-out: -NewName accepts a path only when it shares the directory of
-Path (FileSystemProvider strips the directory and keeps the leaf), and
all four swap calls rename between $HermesHome\node and sibling
node.new-* / node.old-* paths. Pin that invariant so a future refactor
cannot introduce a cross-directory rename, which would throw on every
Windows install and read as a false "in use" deferral.

Source-level probe, matching the other tests/test_install_ps1_*.py
regressions (Linux CI cannot execute the Windows installer).
2026-08-14 21:35:30 -07:00
David Metcalfe e4d0e4c3d8 fix(win): never rewrite the in-use managed Node tree (#80926)
The Hermes-managed Node tree at %HERMES_HOME%\node is destructively
rewritten while the desktop app's Node processes execute from it:
the Node-26 heal did shutil.rmtree + move, the EBADENGINE repair ran
npm install --global --prefix into the tree, and install.ps1's
Test-Node did Remove-Item + Move-Item. Windows rejects those writes
with PermissionError: [WinError 5] on npm.cmd.

- _heal_managed_node_windows: stage the fully-downloaded tree in a
  sibling node.new-* dir, then rename-swap (live tree -> node.old-*,
  staged -> node). The live tree is never deleted before its
  replacement is ready, so an interrupted heal cannot gut it; a
  refused rename is the OS-level in-use signal and defers (returns
  None) instead of forcing the write.
- heal_hermes_managed_node: an in-use deferral does not record the
  once-per-process attempt, so the heal retries once the tree is free.
- managed_node_tree_in_use: cheap psutil pre-check (Windows only) that
  avoids pointless 30-50MB re-downloads in long-lived processes.
- upgrade_managed_npm: defer the in-place npm self-upgrade while the
  tree is in use, with a notice.
- install.ps1: Test-ManagedNodeInUse guard around Update-ManagedNpm and
  the Test-Node install branch, which now rename-swaps instead of
  delete-then-move.

An in-use-but-outdated tree keeps serving the old runnable Node (old
Node beats no Node), and every npm resolution re-evaluates the heal, so
the upgrade applies automatically on the next update with the app
closed.
2026-08-14 21:35:30 -07:00
fangliquanflq 4642f9630d fix(gateway): reject unsafe ordinal-only truncation 2026-08-14 21:33:31 -07:00
SHL0MS bec7df1bba fix(anthropic): coerce blank system text blocks at extraction (#70909)
Residual from PR #70910 after #77509 landed the message-list scrub: a
whitespace-only system content block carrying a cache_control marker
still reached the wire and 400'd the whole request ("text content
blocks must contain non-whitespace text"), wedging the session on every
retry. The block cannot be dropped (it carries the cache breakpoint),
so coerce its text to the shared non-whitespace placeholder when
extracting the system param, copying the block so caller message dicts
are never mutated.

Adds SHL0MS's request-level regression suite from #70910; four of its
five cases already pass on main via #77509 — the system-block case
fails without this fix.
2026-08-14 21:27:05 -07:00
Turgut Kural 071d295a6e fix(test): rename misleading test — non-empty array codex field stripping, not empty array
Reviewer noted the test name suggested an empty-array case but the fixture
has one tool call; renamed to match actual behavior.
2026-08-14 21:25:48 -07:00
Turgut Kural c464001da8 fix(transport): strip empty/null tool_calls on assistant messages
Strict OpenAI-compatible providers (onerouter / Qwen, DeepSeek v4) reject
an assistant message carrying tool_calls: [] (or null) with HTTP 400
'Empty tool_calls is not supported in message.'

The pre-API sanitizer in agent_runtime_helpers.sanitize_api_messages already
drops these on the conversation_loop path, but auxiliary / custom-provider
routes that bypass that sanitizer can still reach the wire with an invalid
empty array and abort the whole session (non-retryable 400).

Normalize at the transport layer too: detect an empty-list / null
tool_calls on assistant messages, strip the key on the per-call copy (never
mutate the stored history), and keep real tool_calls untouched. Includes
unit tests covering empty-list, null, real-call preservation, mixed batches,
user-role non-mutation, copy-on-write, and cross-provider parity.

Follow-up to #58755.
2026-08-14 21:25:48 -07:00
webtecnica f316f7d086 fix(session): drop empty tool_calls in repair_message_sequence (#77921) 2026-08-14 21:25:48 -07:00
liuhao1024 b2453b5894 fix(sanitize): drop tool_calls key when dedup removes all calls
The dedup pass in sanitize_api_messages (introduced by #58327) can
produce an empty tool_calls array when all tool_call_ids in a message
are duplicates of earlier messages in a long conversation history.

DeepSeek v4 and newer OpenAI reject empty tool_calls with HTTP 400:
'Invalid messages[N].tool_calls: empty array'.

When kept_tcs is empty after dedup, drop the tool_calls key entirely
instead of writing tool_calls: [].

Fixes #64335
2026-08-14 21:25:48 -07:00
Teknium 11ccbb4b6f fix(gateway): route delivery-ledger owner-liveness probe through _pid_exists (#41662)
The `_owner_alive` fallback in gateway/delivery_ledger.py (taken whenever a
process start time is unreadable) probed liveness with a raw
`os.kill(pid, 0)`. On Windows that is NOT a no-op: CPython maps sig=0 to
`GenerateConsoleCtrlEvent(0, pid)` (bpo-14484), so probing a LIVE pid whose
start time psutil could not read would Ctrl+C the target's entire console
group. The prior `# windows-footgun: ok` annotation only justified the
EPERM-means-alive exception semantics, not the Ctrl+C side effect.

Route the probe through `gateway.status._pid_exists` (psutil-first,
ctypes OpenProcess fallback on Windows), preserving EPERM-means-alive.
A POSIX-only raw-probe fallback remains for the unreachable case where
gateway.status cannot be imported; on Windows that path reports dead
rather than firing a sig-0 probe.

Tests patch `gateway.status._pid_exists` per the windows-native-support
pattern, including a regression guard asserting os.kill is never used
for the probe. Sabotage-verified (revert → red, restore → green).

Part of #41662 (the os.kill half; watchdog half tracked separately).
2026-08-14 21:25:30 -07:00
Fangliquan 86a8928711 fix(tui_gateway): read profile DB for live session display payload
Warm/live reuse was hard-coding the launch SessionDB, so app-global remote
profile sessions fell back to collapsed in-memory history and dropped
verification candidates that eager profile resume still showed.
2026-08-14 21:25:03 -07:00
deacon-botdoctor 9bff109783 fix(gateway): cancel native clarify only on free prose
Teknium review on #75732: releasing the pending clarify whenever
resolve_text_response_for_session returned False also cancelled
retryable multi-select invalid selections (out-of-range numbers,
unrecognised comma-lists).

Classify rejected typed replies in clarify_gateway:
- rejected_prose → cancel clarify, fall through busy routing (deadlock break)
- rejected_selection → keep clarify armed so the user can retry

Add native multi-select gateway regressions for both paths.
2026-08-14 21:24:36 -07:00
deacon-botdoctor 28b62b069a fix(gateway): keep first clarify resolution 2026-08-14 21:24:36 -07:00
loulanyue c1c3557723 fix(gateway): release rejected native clarifies before steering 2026-08-14 21:24:36 -07:00
HexLab98 1a06e70e1e test(session-search): cover /new-reset discovery, scroll, and browse
Regress the #85756 chain: a session_reset parent must surface from the
empty child, title/scroll must follow, browse must list that parent, and
live delegation children must stay hidden.
2026-08-14 21:22:56 -07:00
Ufonik 54aad0d670 fix(tools): refresh activity heartbeat while a tool call is in flight (#84491)
The gateway turn-inactivity watchdog (gateway/run.py::_watch_gateway_turn_inactivity)
abandons a turn once seconds_since_activity exceeds the inactivity timeout
(default 30 min). Activity was only stamped when a tool started and when it
completed, so a tool call that runs silently for 30+ minutes (quiet builds,
long pytest suites, large downloads, network waits with no output) froze the
clock and the watchdog hard-abandoned a turn that was still making progress,
reaping the tool's processes mid-execution (issue #84491).

Add a daemon-thread heartbeat inside _run_agent_tool_execution_middleware that
touches agent._touch_activity every 30s while the tool is in flight, until the
call returns. Both the sequential and concurrent execution paths funnel through
this single middleware, so one heartbeat covers every tool. The thread is
stopped in a try/finally so it always tears down even if execute() raises, and
is never started when a guardrail/authorization block short-circuits before
dispatch. A genuinely hung tool remains bounded by the tool layer's own
timeouts (terminal default 180s, concurrent batch deadline ~420s), so the
heartbeat only extends the turn's life while the call is legitimately running.

Verified by an independent reviewer (no security/logic defects); 5 unit/integration
tests pass on Python 3.12 (upstream CI). The 30-min gateway backstop remains
for turns whose agent loop itself stalls.
2026-08-14 21:20:03 -07:00
HexLab98 c5788b5ea3 test(agent): cover inline hard timeout and in-flight pool-request abort
Pin that cron/subagent non-streaming calls receive a read timeout matching the stale budget, that an explicit timeout is left alone, and that force_close_tcp_sockets finds sockets on httpcore PoolRequest.connection and clears the socket timeout before shutdown without close().
2026-08-14 21:19:55 -07:00
Teknium dc2fe99ecf feat(delegation): mark max_iterations-truncated subagent results for the parent (#86641)
A delegated subagent that exhausts its per-child iteration budget
(delegation.max_iterations) still returns a summary, so the result carries
status='completed' even though the child's exit_reason is 'max_iterations' and
its work was cut off mid-task. The parent then reads 'completed', trusts the
partial summary, and only discovers the truncation by parsing the prose (where
the child happens to mention 'hit the iteration limit'). That wastes parent
turns and risks acting on incomplete work.

exit_reason is already computed authoritatively and threaded to every
parent-visible surface; it just wasn't reflected anywhere the parent reads at a
glance. This surfaces it:

- delegate_tool.py: add a parent-visible boolean 'truncated' (= exit_reason ==
  'max_iterations') to each task entry, alongside the existing exit_reason.
- process_registry._format_async_delegation: for both the batch and single-task
  paths, when truncated -> use a warning icon, append
  'TRUNCATED: hit max_iterations — work may be incomplete' to the header/Status
  line, and prefix the summary with an unmissable truncation notice. status
  semantics are left unchanged (stays 'completed') so existing icon/summary
  branch logic and ~10 tests asserting status=='completed' stay valid.

Tests: single-task truncated -> banner; single-task clean -> no banner; batch
marks only the truncated task, not its clean sibling. 23/23 in the async-
delegation suite.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-14 21:19:15 -07:00
Teknium 0643105611 fix(file-ops): skip per-file tsc when an ancestor tsconfig.json exists (#86640)
The post-write lint runs `npx tsc --noEmit <file>` on a single .ts file with
no `-p tsconfig`. tsc ignores tsconfig.json for explicit file args, so for any
file in a real TS project it floods phantom diagnostics — unresolved path
aliases (@/... -> TS2307) and ambient globals (Window.hermesDesktop -> TS2339)
that the project config defines. The delta filter sees the same phantom errors
pre- and post-edit, finds no NEW ones, and returns the misleading
'pre-existing lint errors ... the file is still broken' on a perfectly correct
one-line edit — wasting the caller's turns chasing nonexistent breakage.

Existing code already skips shell tsc when an LSP server claims the file, but
that only fires with LSP configured+enabled (not the default), leaving the
common LSP-disabled case fully exposed.

Fix: when an ancestor tsconfig.json exists (local host only), skip the per-file
shell tsc for .ts — its verdict carries zero signal for project files. Real
diagnostics still come from the LSP tier or an explicit `tsc -p tsconfig.json`.
Best-effort ancestor walk; remote/sandbox backends fall back to running the
linter as before. (.tsx already returns skipped via the ext-not-in-LINTERS
branch.)

Tests: ancestor-tsconfig .ts -> skipped even with LSP off; standalone .ts with
no ancestor tsconfig -> shell tsc still runs. 7/7 in the LSP-skip suite.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-14 21:17:00 -07:00
Teknium eeb2d23b86 fix(memory): accept new_text as an alias for content (#86642)
memory replace/remove target an entry by old_text but supply the replacement
via content — an asymmetric pairing. Callers naturally reach for new_text to
mirror old_text (it's exactly the patch tool's old_string/new_string shape),
which left content empty and errored 'content is required'. The failure also
rendered tersely, making the cause easy to miss and costing a retry.

Accept new_text as an alias for content on both shapes:
- single-op: coalesce content = content or new_text in memory_tool() + a
  new_text param wired through the registry handler.
- batch ops: content = op.content or op.new_text (and in the approval-gate
  preview builder).
- schema: document the alias on content and add a new_text property to the
  single-op params and batch item props so strict validators accept it.
content wins if both are set.

Tests: new_text alias on single add/replace, batch add/replace, and
content-wins-over-new_text. 43/43 across memory tool + schema suites.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-14 21:15:08 -07:00
zuowen7 bbb6cc7e99 test: cover tool-message dedupe and latest paging for include_compacted (#80680)
Dedupe key now includes tool_call_id/tool_calls/tool_name: compaction
copies carry those fields verbatim, so identical tool messages across
generations still collapse, while distinct tool calls sharing
role/content/timestamp are never merged. Add endpoint-level coverage
for the desktop's real read path (limit + order=latest +
include_compacted=true).
2026-08-14 21:08:14 -07:00
zuowen7 f71f91a39b fix(desktop): surface compaction-archived messages in transcript reads (#80680) 2026-08-14 21:08:14 -07:00
spfcraze 5f619cfa0e fix(gateway): don't restart supervised services on clean exit
The s6 finish script for profile-gateway services restarted on ANY exit
except EX_CONFIG (78) — including clean exit 0. Restart-on-normal-exit
turns an intentional stop into a reconnect loop: the ashriel-discord
storm in #76435 made 1,000+ connections and got the bot token reset by
Discord.

The finish script now exits 125 (permanent failure, no restart) for both
clean exit 0 and EX_CONFIG; only non-zero, non-78 exits (genuine crashes)
restart normally.

Scope note: #76435 bundles a second, separable symptom (Windows desktop
updater showing the literal 'managed outside dashboard' sentinel). Its
root cause is undiagnosed and #22733 covers the dialog-explanation path;
this PR is the gateway half only.

Tests: behavioral — the rendered finish script is executed via sh for
exit codes 0/78/1/137, asserting no-restart for clean stop and fatal
config, restart for crashes. Pre-existing EX_CONFIG test still passes.
2026-08-14 21:07:21 -07:00
David Metcalfe ebb2859132 fix(clarify): preserve resolved answers when clear_session races a button response
Session-boundary cleanup (gateway/run.py: run-finalization and prompt
delivery-failure paths) calls clear_session to cancel pending clarifies.
It unconditionally overwrote every entry's response with the empty
cancellation sentinel, even entries already resolved by a button callback
or text intercept. A waiter that had already observed the resolved event
would then return the empty sentinel instead of the real answer, silently
discarding the user's response on /new, gateway shutdown, or cached-agent
eviction.

First-writer-wins: clear_session now cancels only entries whose event is
not yet set; already-resolved entries keep their response. Mirrors the
guard resolve_gateway_clarify gained for the same contract. Reported by
doryani-ai on PR #75732; regression test covers the button-then-cleanup
interleaving.
2026-08-14 20:58:42 -07:00
ygd58 525ca9ca1c fix(gateway): match the profile-namespaced session key in the clarify bypass lookup
Fixes #82975.

The adapter-level clarify reply bypass in gateway/platforms/base.py's
handle_message() built its session_key via build_session_key(...)
without a profile= argument, defaulting to the legacy agent:main
namespace. The runner registers pending clarifies under
SessionStore._generate_session_key()'s key, which DOES include
profile=self._resolve_profile_for_key(source). Under a named-profile
multiplex these diverge, so the bypass lookup at
clarify_gateway.get_pending_for_session(session_key, ...) misses --
the user's answer to a pending clarify() gets routed to the adapter's
busy-session queue instead of resolving it. The turn then hangs until
the clarify's 3600s timeout, with no inbound message: log line and no
"Gateway intercepted clarify text response" log line, matching the
reported Telegram symptom exactly.

Verified the divergence directly: _resolve_profile_for_key() returns
None when multiplex_profiles is off (default) -- byte-identical to
the prior implicit profile=None, so this only changes behavior for
multiplexed deployments, matching the issue's exact reported scope.

Fixed by using the same self._session_store._resolve_profile_for_key()
the runner's key generator calls, guarded with getattr() + a None
fallback since _session_store is set via a setter and can be unset
for adapters that never call set_session_store() -- preserving prior
behavior for any such adapter rather than introducing a new crash.

Added a regression test alongside the existing bypass coverage: with a
mocked session_store configured for profile multiplexing, a clarify
registered under the profile-namespaced key must still be found and
resolved (not routed to the busy queue). Verified as a genuine
regression by reverting the fix and confirming the new test fails
with the exact reported symptom (the message handler never gets
awaited -- the clarify lookup misses).

21/21 pass across the five directly related clarify test files;
16/16 across the broader multiplex/clarify-progress test files (no
regression).
2026-08-14 20:57:37 -07:00
HexLab98 39f3bb9312 test(compression): cover degenerate compress_end same-session handoff keep
Regression for #83248: handoff beyond compress_end must not clear a valid
same-session _previous_summary via the cross-session discard path.
2026-08-14 20:52:16 -07:00
Teknium bc5805c35f fix: compare base-URL hostnames, not substrings, in provider-identity checks
Port of the bug class from earendil-works/pi#7933 (DeepSeek base-URL
detection matched by raw substring, missing case variants and matching
lookalike URLs). Hermes had the same class at five sites:

- cli_agent_setup_mixin.py: keyless-custom-endpoint detection treated any
  URL containing the OpenRouter host substring (path segment, lookalike
  domain) as OpenRouter, and missed case variants of the real host.
- models.py validate_requested_model: same substring check for routing an
  openrouter provider with a custom base_url to the custom catalog.
- runtime_provider.py: local-endpoint autodetect matched the string
  localhost anywhere in the URL, including remote hostnames containing it.
- gateway/run.py: /status endpoint display, same local-host substring.
- agent_runtime_helpers.py: Nous Portal cache-layout detection matched
  the nousresearch substring anywhere in the URL.

All sites now use the existing base_url_host_matches / base_url_hostname
helpers (exact host or subdomain, case-insensitive). Regression tests
proven to fail against the old predicates.
2026-08-14 20:47:13 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00
Evgenii acaafcc6bb fix(cron): make immediate execution race-safe
- claim_job_for_fire returns the atomically claimed snapshot with a unique
  fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
  terminal writes by expected_fire_owner so a stale worker cannot record
  over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
  event (ownership loss OR external cancel) into run_job; the agent path
  is interrupted cooperatively and script-based jobs (no_agent + pre-run
  scripts) are hard-stopped with a process-tree kill (POSIX killpg
  SIGTERM then SIGKILL for surviving group members; Windows
  taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
  descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
  the bare job ID, so a replacement run of the same job never consumes a
  stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
  with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
  completed one-shot retention (#80624), blocked_config preflight
  (T1-26), and the advance_next_runs batch on top of current main.
2026-08-14 20:46:50 -07:00
Teknium 5d9e4aaaf2 fix: raise ProviderStreamError for choiceless error chunks + regression tests 2026-08-14 20:46:42 -07:00
Teknium 5a3b593230 fix(desktop): order mid-turn user messages after the assistant output that predates them
A message typed while a turn streamed rendered ABOVE assistant output the
user had already watched arrive (#73793), and the retired
insert-before-the-active-reply fallback could splice the bubble mid-thread
— halfway up the chat — when the stream id was missing or stale (#83151).

Fix the class at every path that assigns a transcript position to a
mid-turn user message:

- New shared appendMidTurnUserMessage (rewind.ts): seal the live stream
  bubble in place (interim), append the correction at the live tail, and
  clear streamId so post-redirect deltas seed a fresh bubble BELOW the
  correction. Used by both the primary composer redirect path
  (use-prompt-actions) and the session-tile steer path
  (session-tile-actions), replacing the insert-before splice and its
  last-assistant mid-thread fallback.
- appendLiveSessionProjection now projects the resume/reload turn in
  arrival order (prompt → streamed output → correction → post-redirect
  output) instead of prompt → corrections → reply, so the projection
  agrees with the live transcript and messages no longer jump upward on
  reconnect. With the gateway's new correction_offsets the flat dump is
  split at each accepted-correction boundary; without offsets the
  corrections follow the projected reply.
- tui_gateway/server.py records correction_offsets (assistant text length
  at each accepted correction) on the inflight turn and carries them in
  _inflight_snapshot, only when complete, so resume can rebuild true
  arrival order. Older gateways/clients degrade cleanly.
- preserveLocalPendingTurnMessages and the projection's latest-user-run
  matcher now treat a live-tail assistant row between the prompt and its
  correction as part of the same turn's run, so arrival-ordered runs
  survive refreshes without dropping the prompt.

Fixes #73793. Fixes #83151.
2026-08-14 20:25:07 -07:00
Teknium f8cc6d082e test(todo): make JSON-string coercion test order-agnostic
The type-coercion test pinned index order of todos, which #42649's
_normalize_order intentionally changes (in_progress lifts ahead of
earlier pending rows). Assert coercion by id instead of position.
2026-08-14 20:24:42 -07:00
Teknium 8bd83a9f7c test(gateway): pin get_update_result in metadata-mirror snapshot test
The test compares two _session_info snapshots taken at different times; the background update-check thread can complete between them and flip update_behind (None -> -1), making the equality assertion flaky once the suite runs long enough. Pin the value via monkeypatch.
2026-08-14 20:24:42 -07:00