Commit Graph

11645 Commits

Author SHA1 Message Date
Teknium fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00
Teknium d2672a349b feat(gateway): optional profile param on cron.manage RPC (#86796)
cron.manage resolved its jobs store from the process HERMES_HOME, so a profile
whose cron lives in ~/.hermes/profiles/<name>/cron/ was invisible to the default
gateway (and any bot/plugin querying per-profile routines saw 'no cron jobs').

Add an optional 'profile' param that scopes the whole action via
set_hermes_home_override, exactly mirroring the adjacent skills.manage handler:
resolve get_profile_dir(profile), 404 (err 4064) if missing, override in a
try/finally that always reset_hermes_home_override. Omitted/None keeps the
launch-profile behavior, so existing callers are unaffected. cronjob() itself is
unchanged (it already keys off HERMES_HOME).

Enables the Hermes-Bot-Mode plugin to show a bot's real routines
(NousResearch/Hermes-Bot-Mode#37). Needs a SERVE-backend gateway restart to take
effect live. 2/2 in the new focused test.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:25:00 -07:00
adikpb 688abc585f test(vision): assert no max_tokens cap in browser and video aux kwargs
Sweeper follow-up: the browser-screenshot and video kwargs captures now
also assert max_tokens is absent, protecting the central auxiliary
no-cap policy against refactors that would restore the hardcoded caps.
2026-08-15 12:53:37 +05:30
adikpb ec470d9db2 test(vision): assert vision aux calls carry no max_tokens cap
Covers the max-tokens-knob contract: vision call_kwargs omit max_tokens
entirely (configured values, defaults, and even an explicit
auxiliary.vision.max_tokens config entry must never be forwarded), so
providers use their full output budget.
2026-08-15 12:53:37 +05:30
kshitij 8b58f9f68f test(bedrock): pin stream-path cap omission; document truthiness edge
Self-review follow-up: cover call_converse_stream's max_tokens=None path
(same builder, previously unpinned) and document why the shim reads the
caller cap with truthiness rather than 'is None' (parity with the
Anthropic shim's reading).
2026-08-15 12:47:27 +05:30
kshitij 5ef52273cd fix(bedrock): let aux calls omit the Converse maxTokens cap
The Bedrock Converse shim hardcoded 'else 4096' when the caller passed no
max_tokens, so auxiliary vision descriptions stayed capped at 4096 tokens
on the Bedrock wire even after #75253 removed the vision call sites' own
caps (#10809 was only partially fixed there).

Converse's inferenceConfig.maxTokens is optional; when omitted, Bedrock
defaults to the model's maximum allowed output. Thread an explicit
max_tokens=None through build_converse_kwargs/call_converse to omit the
field, and drop an all-empty inferenceConfig from the wire request
entirely. The 4096 default is unchanged for every existing caller (main
transport passes params.get('max_tokens', 4096) explicitly), so only
no-cap aux calls opt in.

Surfaced during review of #75253.
2026-08-15 12:47:27 +05:30
EvanProgramming 30c469b153 fix(gateway): spare pidfile-less Scheduled-Task gateways from the orphan reaper on Windows (#83683)
On Windows _get_service_pids() is empty (no systemd/launchd query), so a
Scheduled-Task-supervised gateway whose gateway.pid record is missing or
stale is invisible to both the service-PID and recorded-PID exclusions the
reaper already applies (#86658) — and gets SIGTERM'd on every desktop open
(#86098 class, pidfile-less path).

Add a Windows-only backstop: any reaper candidate whose parent chain
reaches services.exe (the Task Scheduler launches tasks under the services
tree) is spared even with no pidfile.

The backstop is deliberately inert on POSIX: every process there has PID 1
(launchd/init/systemd) in its ancestry — and a genuine orphan is reparented
directly to PID 1 — so supervisor-name ancestry carries zero supervision
signal and would disable the reaper entirely on macOS/WSL (#51325, #75936).
POSIX supervised gateways are already covered pidfile-independently by the
_get_service_pids() exclusion.

Known limitation (fail-open, documented): if the Task-launched bootstrap
parent has already exited, Windows does not reparent the gateway, the chain
breaks before services.exe, and the gateway is treated as an orphan.

Salvaged from #86702 by @EvanProgramming (authorship preserved); reduced to
the genuinely-new Windows backstop — the PR's other two hunks were already
merged on main via #86658 (one in a strictly stronger full-parent-chain
form) and its POSIX ancestry checks were dropped as unsound (verified
empirically: a true double-fork orphan's psutil parent IS launchd).
2026-08-15 12:04:29 +05:30
kshitij e3fab0437e refactor(cache): never-raising scope resolver shared by both call sites
/simplify-code finding: turn_context evaluated resolve_prompt_cache_scope()
inside set_runtime_main's argument list under the umbrella try/except — a
resolution failure would silently skip the ENTIRE runtime binding
(provider/model/base_url/api_key/session_id for all aux calls that turn),
not just the cache scope.

- prompt_cache_scope: add resolve_prompt_cache_scope_safe() (never raises,
  returns None on failure/empty).
- turn_context: resolve the scope into a local via the safe variant BEFORE
  the set_runtime_main call, so a failure can only lose the scope.
- chat_completion_helpers: _prompt_cache_scope_for_agent delegates to the
  shared safe variant (guarded import retained).
- tests: +1 (hostile-property agent -> None; normal/empty passthrough).
2026-08-15 11:09:56 +05:30
kshitij 96cdf19a0b refactor(cache): fold self-review findings on the rotation-scope fix
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
  _session_db re-resolves instead of staying pinned to the physical id);
  _persist_disabled agents (background-review forks that never get a DB row)
  memoize the fallback instead of re-querying the lineage per API call;
  module docstring cross-references get_conversation_root and why the two
  lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
  _prompt_cache_scope_for_agent(agent) call to a single local above the
  OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
  don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
  key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
2026-08-15 11:09:56 +05:30
kshitij cee2446222 fix(cache): keep prompt_cache_key warm across compression session rotation
Legacy compaction mode (compression.in_place: false) rotates the physical
session_id mid-conversation. The prompt-cache scope introduced in #79161 was
derived from that physical id, so every rotation moved the same conversation
into a fresh cache bucket - the prompt cache went cold at every rotation
boundary (#79017).

Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT
of the current session (SessionDB.get_compression_lineage, fork-aware
post-#79193) - once per turn, memoized per transcript segment, and prefer it
over the physical session_id at every prompt_cache_key derivation site:

- agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) -
  lineage-root walk with per-segment memo; falls back to the physical id
  when no DB is attached or the walk fails, degrading to pre-fix behavior.
- transports/codex.py: build_kwargs accepts cache_scope_id and prefers it
  for the body prompt_cache_key, the xAI x-grok-conv-id header, and the
  Codex x-client-request-id routing header. The Codex session_id header
  keeps the raw physical id (transcript identity, #57012 contract).
- transports/chat_completions.py: _add_prompt_cache_key accepts
  cache_scope_id with the same precedence.
- chat_completion_helpers.py: build_api_kwargs threads the resolved scope
  into all three build_kwargs call sites (codex, profile, legacy).
- auxiliary_client.py: set_runtime_main carries cache_scope; the aux
  Responses cache-key site prefers it over the physical session_id.
- turn_context.py: resolves the scope once per turn and threads it through
  set_runtime_main (no DB walk on the per-API-call hot path).

Scope semantics preserved from #79161: /new starts a fresh scope (new
lineage), /branch children, delegate subagents, and tool children stay
isolated (explicit-fork exclusion in get_compression_lineage), unrelated
sessions keep distinct buckets, and cron per-fire timestamps still
normalize via _cache_scope_from_session_id.

Default installs compact in place (session_id never rotates), so they hit
the memo and produce byte-identical keys to before.

Fixes #79017
2026-08-15 11:09:56 +05:30
Teknium 471c687c2b test(managed_uv): cover explicit-patch fallback on the next minor line; dedupe retried versions
Follow-up to the salvaged #76252 addressing both review gaps:

- New TestMinorLineFallForward class with a direct test of the
  explicit-patch fallback branch: bare '3.12' resolves to a VULNERABLE
  build while an explicit 3.12.x patch is fixed, so recovery must go
  through _list_available_patches on the next minor line. Asserts the
  exact `uv python install` request sequence.
- New all-minors-exhausted test: everything vulnerable on 3.11-3.13
  returns None with per-line attempts bounded by _MAX_PATCH_RETRIES and
  no requests beyond 3.13 (requires-python is <3.14).
- test_retry_is_bounded_by_max_retries_constant now actually uses its
  counting wrapper and asserts the collected install calls (the
  previous version collected them into a dead variable).

Also dedupes the fallback loop the same way the same-minor loop does:
_attempt_install_generation can now record the probed candidate version
into a caller-supplied tried_versions set, so the explicit-patch pass
skips the version the bare-minor request already resolved to and
rejected -- previously that wasted a full download+install+probe+delete
cycle per minor line re-trying a known-vulnerable build.
2026-08-14 22:37:46 -07:00
RelaxJonh 2bccd6ad08 fix(managed_uv): fall forward to next Python minor when current line has no fixed SQLite build
When every patch on the current minor line (e.g. 3.11) still links a
vulnerable SQLite (e.g. 3.50.4 on Windows), the provisioner now tries
the next supported minor line (3.12, then 3.13) before giving up.

Previously, _install_safe_python_generation only tried patches within
the same minor line. On Windows, where python-build-standalone may not
publish a fixed build for the installed patch, users were stuck with a
repeated warning on every `hermes update` with no path forward.

The requires-python constraint (>=3.11,<3.14) and the downstream
import smoke test already gate compatibility, so the minor-line
upgrade is safe.

Adds allow_minor_upgrade parameter to _attempt_install_generation to
relax the same-minor-line version guard when called from the fallback
path.

Fixes #76106
2026-08-14 22:37:46 -07:00
Teknium 62014d8dd3 fix(guard): steer live-checkout block message to disk-backed scratch clones
The guard's "use a separate worktree or temporary clone" advice sent
agents to /tmp by default. /tmp is RAM-backed tmpfs on most distros, and
parallel salvage clones each running npm ci (~1.6GB per clone) filled a
32GB tmpfs to 97% during a 15-subagent campaign, ENOSPC-ing sibling test
runs. The message now recommends `git clone --shared <root> ~/.hermes/scratch/<task>`
(honoring HERMES_HOME), warns that dependency installs belong on real
disk, and tells the agent to delete the clone once the branch is pushed.
2026-08-14 22:34:51 -07:00
Teknium 3af56c2203 fix(install.ps1): surface uv installer errors and add GitHub + existing-uv fallbacks (#69216)
Install-Uv piped the astral installer's entire output to Out-Null, so any
real failure (proxy block, AV quarantine, permissions) surfaced only as the
generic "uv installed but not found" message, and astral.sh was the sole
install source even though corporate proxies commonly block it while the
byte-identical GitHub releases installer downloads fine.

Three-rung ladder, all inside Install-Uv:
1. astral.sh installer with output captured via Tee-Object.
2. GitHub releases installer mirror (same UV_INSTALL_DIR).
3. Salvage an existing uv.exe (Get-Command uv, or the astral default
   %USERPROFILE%\.local\bin\uv.exe) by copying it into $HermesHome\bin so
   the managed-first invariant holds.

On total failure, print the last 15 lines of captured installer output plus
the existing manual-install pointer.

Reported by @BitBernd; proxy diagnosis by @gakugaku; Out-Null suppression
first identified by @webtecnica in #69366.

Closes #69216
2026-08-14 22:33:58 -07:00
Tachi d1df111ccd fix(update): restore Hermes Tools dependencies 2026-08-14 22:33:44 -07:00
Tachi b67f202184 fix(update): honor lazy install opt-out during restore 2026-08-14 22:33:44 -07:00
Tachi 979a20052f fix(update): preserve activated extras across runtime rebuilds 2026-08-14 22:33:44 -07:00
konsisumer 4aa9f738ce fix(update): rebuild Desktop after release artifact loss 2026-08-14 22:27:37 -07:00
joaomarcos fa72a1edf8 fix(kanban): re-create the schema when a cached DB path loses it (#83445)
`connect()` caches every path it has initialized in the process-local
`_INITIALIZED_PATHS` set and then skips all first-open work for it —
header validation, integrity probe, `SCHEMA_SQL`, additive migrations.
That cache is keyed on a path, but the schema it stands for lives in a
file, and the two can drift apart: delete or replace `kanban.db` under a
live gateway/dispatcher/dashboard process and the next `connect()` takes
the fast path, lets SQLite create a fresh empty database, and hands back
a connection with no tables in it.

Nothing notices. Every query then fails with `no such table: tasks`,
`plugin_api._conn()` logs its init warning and carries on, and the board
renders empty. Because the cache entry survives, the process re-creates
the same schema-less ~4 KB file on every restart of the desktop app in
front of it — only killing the backing process clears it.

Verify the sentinel table on the fast path and self-heal when it is gone:
drop the stale cache entry and fall through to the existing init path,
which re-runs the probes and the schema script under the cross-process
init lock. The check is one `sqlite_master` lookup on the already-resident
page 1, so the steady-state path stays lock-free (#36644) and does no
schema work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:25:23 -07:00
konsisumer 4b0c1031db fix(desktop-update): wait for rebuilt executable before relaunch 2026-08-14 22:23:00 -07:00
Teknium 0cbc4ce83b fix(streaming): gate the interrupt worker join on live Relay managed execution
The unconditional 2s join before InterruptedError delayed interrupt
detection when Relay managed execution was not active (CI:
tests/run_agent/test_interrupt_propagation.py — detection took 2.34s
against a <1.0s budget, because the mocked worker sleeps 5s and there
is no Relay scope to unwind).

Extract the join into _join_worker_for_relay_teardown(), which no-ops
unless a Relay runtime exists AND managed execution consumers are
registered — the only case where an orphaned physical scope can corrupt
the LIFO stack (#81521). Applied at all three interrupt sites
(streaming, non-streaming, Bedrock streaming). The regression test now
simulates a live runtime so the join path stays covered.
2026-08-14 22:11:30 -07:00
Teknium 3537ef9d01 fix(streaming): widen #81521 interrupt join to sibling paths and use version-correct Relay top accessor
Follow-up to HexLab98's salvaged commits:

- Apply the same bounded worker join before raising InterruptedError at
  the two sibling interrupt sites that share the raise-without-join
  shape: the non-streaming API poll loop and the Bedrock streaming poll
  loop. Both workers run Relay-managed physical attempts, so raising
  immediately allowed turn teardown to race a still-open physical scope
  exactly as in the streaming path.

- Address the #81601 review finding (egilewski): the pinned nemo-relay
  binding's get_scope_stack() returns a native ScopeStack object which
  scope.pop rejects with TypeError, so the orphan drain never drained
  under the real binding. current_top() now prefers the version-correct
  scope.get_handle() accessor and falls back to the old list-unwrap for
  fake/legacy shapes. Handle comparisons go through same_handle(),
  comparing by uuid, because native ScopeHandle instances do not
  implement value equality.

- Add a real-binding regression test that reproduces the orphaned-scope
  session close against the pinned native wheel (skips where the native
  binding is unavailable), alongside the existing fake-based coverage.
2026-08-14 22:11:30 -07:00
HexLab98 81c4f3a143 test(streaming): cover interrupt join, orphan drain, and EIO paint freeze 2026-08-14 22:11:30 -07:00
Teknium 4b7b2b0049 fix: widen base-URL hostname identity class to remaining substring sites
Follow-up to #85737, which migrated five provider-identity sites onto
utils.base_url_host_matches()/base_url_hostname(). This completes the class
sweep (never-patch-predicates: one owner, every site) and folds in the two
open contributor PRs attacking individual sites:

- agent/auxiliary_client.py ZAI/Kimi OpenAI-wire rewrite (PR #85715,
  pierrenode): 'bigmodel'/'api.z.ai'/'api.kimi.com' substring checks
  rewrote proxy paths containing those markers.
- hermes_cli/runtime_provider.py Azure endpoint detection (PR #74721,
  RelaxJonh, issue #74312): 'azure.com' substring picked the Azure key
  for non-Azure hosts whose path contained the text.
- run_agent.py: _is_azure_openai_url, _is_copilot_url, Anthropic
  credential-refresh azure guard, _anthropic_preserve_dots host
  allowlist, OpenRouter/mistral reasoning gates.
- agent/chat_completion_helpers.py: nousresearch / nvidia detection.
- agent/conversation_loop.py: GitHub Models 413 hint.
- agent/usage_pricing.py: localhost billing-route detection.
- hermes_cli/model_switch.py: api.openai.com catalog fallback and
  localhost custom-provider detection.
- cli.py: local-model autodetect and Ollama/LM Studio context-length
  hints (port-anchored instead of '11434' in URL).
- tools/mcp_oauth.py: Figma remote-MCP detection.
- tools/skills_hub.py: raw.githubusercontent.com source-URL check.

Regression tests extend tests/hermes_cli/test_base_url_host_identity.py
(azure/copilot/dotted-model/figma proxy-path + lookalike cases) and
tests/agent/test_minimax_auxiliary_url.py (ZAI/Kimi path false positives).

Closes #74312. Salvages #85715 and #74721 with authorship preserved.
2026-08-14 22:04:16 -07:00
pierrenode 2d9f116351 fix(agent): anchor ZAI/Kimi base_url host matching to avoid substring false positives
_to_openai_base_url() matched ZAI (open.bigmodel.cn, api.z.ai, bare
"bigmodel") and Kimi (api.kimi.com) via `substring in url`, so any custom
gateway whose base_url happened to contain one of those strings as a path
segment (e.g. a reverse-proxy prefix like /proxy/bigmodel-fallback/) was
silently misrouted to the wrong OpenAI-wire endpoint shape.

This is the same false-positive class 6f33f510e8 just fixed for the
MiniMax branch in the same function by switching to base_url_host_matches()
(hostname-anchored). Apply the same fix to the ZAI and Kimi branches, which
that commit didn't touch. Drops the bare "bigmodel" substring check since
open.bigmodel.cn is the only canonical bigmodel-family host referenced
anywhere else in the codebase (agent/model_metadata.py, hermes_cli/auth.py).

Added regression tests mirroring the MiniMax marker-in-path tests added in
the same commit.
2026-08-14 22:04:16 -07:00
Teknium cd34661e9c test(update): mock gateway discovery now that the restart phase is surfaced
The gateway auto-restart phase used to swallow every exception at debug
level, so tests driving cmd_update end-to-end never noticed it touching
real gateway discovery. With #78574 surfacing an aborted restart as a
failed update, an unmocked find_gateway_pids on a box with a live
gateway hits the conftest live-system guard and turns into a spurious
sys.exit(1).

Add an autouse fixture in test_cmd_update.py (discovery returns nothing,
systemd unsupported) and the same seams in test_update_head_moved_gate's
helper so the phase is a clean no-op for tests that do not assert on
gateway restarts.
2026-08-14 22:03:56 -07:00
Halldrix 19cff89300 fix(update): complete pending core install before any native import (self-lock loop fix)
Reviewer egilewski found the original defer was circular (#83590 comment):
the self-lock preflight wrote .update-incomplete and exited, but the next
launch only ran the full recovery AFTER main.py's third-party imports —
so a healthy venv's probes made the early pass a no-op, main.py imported
cryptography eagerly, the .pyd got mapped again, and the deferred install
re-hit the exact self-lock it was meant to escape.

Close the loop by making the marker guarantee the install runs BEFORE any
native extension can be imported:

- hermes_cli/_install_repair.py (new, stdlib-only): single source of truth
  for the core .[all] reinstall — ensurepip bootstrap, uv-pip/pip
  resolution with VIRTUAL_ENV, Termux env stripping, Windows hermes*.exe
  quarantine, per-extra fallback ladder, and fd1→fd2 routing for acp
  safety.  Deliberately free of managed_uv/hermes_constants imports so it
  stays importable in the corrupted-venv state it exists to repair.
- hermes_cli/_early_recovery.py: recover_if_needed now completes a pending
  .update-incomplete install BEFORE the import probes, on every launch
  that sees the marker (unless argv is update).  Success clears the
  marker; failure bumps an attempts counter inside the marker body and
  keeps it.  A 3-attempt ceiling stops a persistently-failing install
  from reinstall-hammering every launch (hermes acp included) — past the
  ceiling the late post-import recovery takes over with its manual
  recovery instructions.  Single-flight lock shared with the late path.
- hermes_cli/main.py: _recover_core_update_marker_locked delegates the
  install to the shared executor (no duplicated logic); ensure_uv stays
  in the late path so a venv whose uv vanished mid-update still
  bootstraps it.
- tests: 7 new regressions — the reviewer's exact case (marker + healthy
  venv → install runs while sys.modules has no cryptography), failure
  keeps marker + increments attempts, retry ceiling, lazy marker does
  not trigger core install (#58004 invariant), argv-update skip, and
  corrupt/missing marker bodies.  The key test was sabotage-verified:
  removing the pre-import branch makes it fail with zero install calls,
  while a lone-lazy-marker test still passes; restoring the branch makes
  it pass again.

Refs #83569
2026-08-14 22:03:56 -07:00
Halldrix c6a71294b6 fix(update): detect updater self-lock on Windows + repair venvs whose base interpreter is uv-managed
Two gaps left every Windows git-checkout install unable to recover from
the exact failure state #83569 reports:

1. Self-lock detection. _detect_venv_python_processes() always excludes
   the calling process by design — a CLI hermes update IS the venv python.
   An updater that had already imported a native venv extension (the
   canonical one being cryptography.hazmat.bindings._rust, mapped while
   hermes_cli.main resolved external secret sources) passed every
   preflight and then died mid-sync with os error 5 when uv tried to
   rewrite the mapped .pyd, stranding the venv half-updated. A new
   preflight now refuses the sync before touching the checkout, writes
   the update-incomplete marker so the next fresh launch completes the
   install, and exits 2. Verified on a live Windows 11 host: after
   importing hermes_cli.main, tasklist /m _rust.pyd shows the .pyd mapped
   in the caller, and a peer process cannot open it read-write
   (Permission denied) — while a rename succeeds, matching how uv/pip
   actually fail (truncate+write, not rename).

2. Early-recovery install path. _early_recovery._run_repair_install used
   sys.executable -m pip unconditionally. Windows git checkouts install
   on a uv-managed base interpreter (python-build-standalone), whose
   EXTERNALLY-MANAGED marker makes plain pip abort with
   externally-managed-environment — the repair no-oped and the venv
   stayed broken. The repair now detects the PEP 668 marker, prefers
   uv pip install with VIRTUAL_ENV pointed at the project venv, and
   falls back to pip --break-system-packages when no uv binary exists.

Both fixes ship with subprocess/unit regressions (sabotage-verified):
the new tests fail on pre-fix code and pass with it. Complements #77517,
which keeps the updater from importing cryptography in the first place;
this PR is the defence-in-depth when any future path loads it anyway.

Fixes #83569
2026-08-14 22:03:56 -07:00
chelsealong 49d72a02f6 fix(update): verify Windows gateway cold-start survives before reporting success
_cold_start_windows_gateway_after_update() printed the success line off a
successful Popen return alone, which only proves CreateProcess succeeded,
not that the child survived. On Windows, a job object denying
CREATE_BREAKAWAY_FROM_JOB hard-kills the child during updater teardown
before it logs anything, yet the updater still printed "Starting Windows
gateway after update (PID ...)" — leaving Telegram/Discord/etc. offline
with no indication anything failed (#84185).

Route the success report through gateway_windows._report_gateway_start(),
the same post-spawn liveness poll every other _spawn_detached() caller
already uses, so a dead child is reported as a failure with a
manual-recovery hint instead of a false success.
2026-08-14 22:03:56 -07:00
PRATHAMESH75 517151ee4a fix(install): fail closed when a stopped gateway leaves an empty survivor probe
Review follow-up (#78574): the aborted-restart handler only flagged the fleet
stale when the post-failure survivor probe was None or non-empty. A positive
empty probe was treated as proof-of-safety — but `[]` is only safe when
nothing was running before the phase. If a gateway was discovered, stopped
(SIGTERM/drain), and its replacement never came back, the probe is empty at
exactly that unsafe moment and the update reported success — the fail-open
contract this fix exists to close.

Snapshot the pre-restart gateway PIDs before any stop/drain and route the
handler decision through a pure _restart_phase_failure_is_incomplete() helper
that fails closed on an empty survivor set whenever a gateway existed
pre-restart (or the pre-state could not be read). Add decision-level regression
tests covering the stopped-without-replacement gap, unknown pre-state, and the
truly-no-gateway positive control.
2026-08-14 22:03:56 -07:00
PRATHAMESH75 95018b6bba fix(install): surface aborted gateway restart during hermes update
The gateway auto-restart phase in `hermes update` was wrapped in a blanket
`except Exception` that only logged at debug level. When the phase raised
early — e.g. importing `hermes_cli.gateway` from the freshly pulled checkout
inside a process that already loaded pre-update modules — every drain and
restart line vanished from the update output, the update printed
"Update complete!" and exited 0, and the still-running gateway kept serving
pre-update modules against replaced source files. The next Telegram turn died
with `ImportError: cannot import name 'is_trivial_prompt'`.

The handler now probes for surviving gateway processes and, unless it can
positively prove none are running, prints the cause plus a manual recovery
command and marks the fleet restart incomplete — which exits nonzero and
writes the gateway-mode exit-code marker, matching the existing
failed-or-stale-unit path.

Fixes #78574
2026-08-14 22:03:56 -07:00
Soheil Fakour bdfdd4392f fix(update): gate 'Code updated!' on HEAD actually moving (#79678)
A detached/pinned checkout can report 'N new commit(s)' against origin,
run the ff-only merge successfully, and still sit on the old commit
afterward (the branch-switch step re-detaches to the raw SHA). Before
this guard 'hermes update' printed '✓ Code updated!' and reinstalled
deps + rebuilt the desktop app against the stale tree - no error, no
warning, 'hermes doctor' healthy.

Compare pre-pull and post-pull HEAD; if they match, fail loudly with a
reattach hint instead of claiming success.
2026-08-14 22:03:56 -07:00
mariobgsp be708ff1b9 fix: normalize managed config overlay before merge in load_config
The shared load-boundary flatten added for dict-valued model.default only
ran on the user/default merge; _load_config_impl then deep-merged the raw
managed overlay without normalizing, so a managed model.default:
{provider, model} still reached status/fallback/runtime readers as a dict.

Normalize the managed overlay (same _normalize_root_model_keys pass, plus
the bare model-string -> model.default promotion used by
managed_scope.apply_managed_overlay) before expanding and merging, so
every overlay is canonical before load_config returns.

Adds load_config() regressions for a nested managed default and a bare
managed model string.
2026-08-15 10:29:05 +05:30
mariobgsp 998329a621 fix: flatten dict-valued model.default at config load boundary
Extends the fix to the config-load chokepoint so every reader sees plain
strings, not just the interactive CLI paths. _normalize_root_model_keys
now flattens a dict-valued model.default/model.model into a string default
plus the nested provider (promoted to model.provider when no explicit
outer provider or "auto" is set), covering the residual readers the
review flagged: doctor, status/dump, fallback picker, prompt-size, and
the context-switch guard — all of which called .strip()/flowed the raw
value and would crash or misroute on a nested dict.

Adds _normalize_root_model_keys regression coverage for the flatten
(precedence, auto-override, explicit-provider-win, alias shape, flat
strings untouched).
2026-08-15 10:29:05 +05:30
mariobgsp 86054eff62 fix: keep nested model.default provider paired through HermesCLI
The prior dict coercion converted a dict-valued model.default to a plain
model string but dropped the nested provider. On the interactive CLI path
requested_provider then fell back to the outer merged model.provider
(typically "auto", authoritative at runtime resolution), so the model
could be routed through the wrong active provider.

Canonicalize both halves at the shared boundary: _split_model_config_default
flattens a dict-valued default into (model, provider) and HermesCLI.__init__
feeds the nested provider into the requested_provider chain (still below an
explicit --provider argument). new_session reuses the same helper.

Adds regression coverage asserting the nested provider stays paired with the
model, that flat string defaults keep the outer provider behavior, and that
an explicit provider argument still wins.
2026-08-15 10:29:05 +05:30
Teknium aa5a960675 fix(installer): hold the venv rollback source through dependency validation (#83149)
Review finding on PR #83194 (egilewski): Install-Venv committed the venv
transaction as soon as the replacement had a working interpreter, deleting
the parked previous venv. Install-Dependencies is a separate later stage
(a separate process under the stage-per-process bootstrap) and every
dependency tier or the baseline-import gate can still fail after that
point - a failed update could still leave Hermes and the blocker probe
unusable with no rollback source.

Now:
- Install-Venv records the parked backup in venv.pending-backup instead
  of deleting it, and excludes it from the venv.stale.* sweep.
- Install-Dependencies wraps the dependency tiers + baseline-import gate
  in the transaction: Restore-VenvBackup on failure (parks the failed
  replacement as venv.failed.*, renames the previous venv back), and
  Complete-VenvTransaction only after the imports prove the replacement
  usable.
- Source-contract regression tests for the boundary
  (tests/test_install_ps1_venv_transaction_boundary.py).
2026-08-14 21:58:09 -07:00
HexLab98 737d6339b7 test(install): lock no in-place venv gut on rename failure
Cover Install-Venv abort-on-rename-fail, probe_failed JSON when psutil
is missing, and update the process-sweep contract for rename-aside.
2026-08-14 21:58:09 -07:00
JoaoMarcos44 be1e0c89c9 test(installer): align Windows venv process guard with transactional parking (#83149) 2026-08-14 21:58:09 -07:00
JoaoMarcos44 6dab0318bf test(installer): reproduce destructive venv fallback (#83149) 2026-08-14 21:58:09 -07:00
Teknium f8d75db026 test(gateway): accept keyword args in _connect_adapter_with_timeout mocks 2026-08-14 21:57:41 -07:00
Teknium c22815fcaa fix(gateway): cap Telegram cold-start connect so startup reaches running fast (#85993)
CI / Docs Site (push) Has been cancelled
CI / Detect affected areas (push) Has been cancelled
CI / Python tests (push) Has been cancelled
CI / OS-specific tests (push) Has been cancelled
CI / Python lints (push) Has been cancelled
CI / JS & TS checks (push) Has been cancelled
CI / Installer tests (push) Has been cancelled
CI / Desktop E2E (push) Has been cancelled
CI / Deny unrelated histories (push) Has been cancelled
CI / Check contributors (push) Has been cancelled
CI / Check uv.lock (push) Has been cancelled
CI / Check no committed infographics (push) Has been cancelled
CI / package-lock.json diff (push) Has been cancelled
CI / Lint Docker scripts (push) Has been cancelled
CI / Supply-chain scan (push) Has been cancelled
CI / Review label gate (push) Has been cancelled
CI / OSV scan (push) Has been cancelled
CI / All required checks pass (push) Has been cancelled
CI / CI timing report (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
Docker Build, Test, and Publish / Detect affected areas (push) Has been cancelled
Docker Build, Test, and Publish / build (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / build (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / publish (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / publish (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / merge (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
The initial (pre-running) connect awaited during gateway startup now uses
a capped 45s budget for Telegram instead of the full 180s (#67498) budget.
On timeout the platform is queued for the reconnect watcher, which retries
with the full budget and is_reconnect=True (preserving the offline update
queue, #46621). Combined with the parallel startup connects, an unreachable
Telegram no longer holds the whole gateway out of the running state.
2026-08-14 21:57:41 -07:00
EvanProgramming 42a4e86239 fix(test): genuinely verify parallel startup connects (#83791)
The previous concurrency assertion (slow_start < fast_end) was true under
BOTH the serial and parallel implementations, so it proved nothing -- it
even passed against the old serial code on main. The only assertion that
distinguishes the two is that the fast platform finishes before the slow
one (fast_end before slow_end), which is only possible when the connects
overlap.

Switch the test to record connect start/end events in arrival order
(clock-resolution independent) and assert fast_end precedes slow_end. This
also fixes the Windows failure @zuowen7 reported: time.monotonic() has only
~15 ms resolution there, so two parallel connects could land on the same
tick and defeat any wall-clock comparison -- event ordering cannot.

Verified the new test fails against origin/main (serial) and passes against
this branch (parallel).
2026-08-14 21:57:41 -07:00
EvanProgramming d86c67dc7e fix(gateway): connect messaging platforms in parallel at startup (#83791)
GatewayRunner.start() previously awaited each platform's connect() (with its
own timeout) in a serial for-loop. A single slow/failing platform (e.g.
Telegram behind a dead proxy) delayed every later platform's connect by a full
timeout window, cascading one platform's failure onto WeChat/QQ/etc.

Now the slow connect() calls run concurrently via asyncio.gather while the
serial pre-filter (checks, adapter creation, handler wiring) and the
single-threaded result aggregation (shared-state mutation, error handling)
are unchanged. A failing platform no longer blocks the others.

Adds regression tests proving connect() calls overlap and that one failing
platform leaves the others connected.
2026-08-14 21:57:41 -07:00
tachyon-r 7a6b8917f7 fix(tools): recognize discovered plugin platforms 2026-08-14 21:56:33 -07:00
Chen Jin 7224301856 fix(toolsets): admit explicitly-configured plugin toolset keys in _get_platform_tools (#81163)
Layer 2 of the #81163 / #78050 fix: _get_platform_tools computed
plugin_ts_keys = _get_plugin_toolset_keys() but only used
CONFIGURABLE_TOOLSETS in the explicit-config filter, so a user-listed
plugin key like `a2a` in `platform_toolsets.cli: [hermes-cli, a2a]` was
silently dropped. The filter now unions configurable and plugin toolset
keys when evaluating has_explicit_config and when admitting per-key
entries.

Cherry-picked from PR #81190 (Layer 2 hunks only; Layer 1 is covered by
the provides_tools mechanism from PR #78842).
2026-08-14 21:56:33 -07:00
Eman e42db348c9 fix(plugins): register deferred platform client tools at discovery (#78050)
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.

Re-anchored accordingly:

- Discovery-time pre-registration, module reuse, and the `provides_tools`
  opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
  since those tools registered before `registration_start` and the slice
  cannot see them.
- A failed materialization no longer carries attribution across. The
  failure path now sweeps the whole ownership ledger for the plugin key,
  not just the `registration_start:` slice, so the pre-registered tools
  are disposed along with the adapter. Attribution and the registry now
  agree at zero instead of reporting tools the process is not serving.

tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 21:56:33 -07:00
kshitij ce658e82ff fix(session-search): narrow lineage escape to reset/compression; trust SQL child classifier
Follow-up on the salvaged #85764 commits, addressing review findings:

- _session_left_live_context now allowlists end_reason == 'compression'
  or a fresh reset (_FRESH_RESET_END_REASONS) instead of accepting any
  non-None end_reason. The wide predicate let 'branched' parents — whose
  transcript /branch verbatim-copies into the child — surface as
  same-lineage recall hits, returning content already in the caller's
  live context (verified empirically vs main).
- _FRESH_RESET_END_REASONS is now derived from the canonical
  hermes_state_common._RESET_END_REASONS (plus CLI 'new_session') instead
  of a third hand-maintained copy, per that tuple's anti-drift comment.
  Import verified cycle-free.
- Browse drops the Python re-check of parent_session_id rows:
  list_sessions_rich (include_children=False) already applies the
  canonical _LISTABLE_CHILD_SQL classifier, and the Python re-check
  re-hid legacy pre-marker reset children the SQL same-key heuristic
  deliberately admits. _has_reset_from_marker (now orphaned) removed.
- Tests: branched-parent exclusion regression guard (mutation-checked:
  fails on the overbroad predicate) + legacy pre-marker reset child
  browse guard. 48/48 pass.
2026-08-15 10:25:19 +05:30
Teknium 2b5a3fb8ac test(cron): expect tick to contain create_execution failure per-job 2026-08-14 21:55:14 -07:00
webtecnica 2d81236f7f fix(cli): report background-dispatch cron runs without false 'failed' (#83340) 2026-08-14 21:55:14 -07:00
Teknium 5733ec097b test(cron): adapt guard-leak regression to owner-fenced dispatch on main
The salvaged regression from #86582 predates the claim_job_for_fire
owner-fencing that landed with #70638; mock the claim and heartbeat so
the healthy job actually runs through the fenced flow.
2026-08-14 21:55:14 -07:00