Commit Graph

22992 Commits

Author SHA1 Message Date
Teknium 31ca1200ef feat(compression): field-proven summarizer prompt upgrades
- anti-injection rule in preamble (gemini-cli state_snapshot pattern)
- verbatim security-constraint preservation in Constraints & Preferences
  (claude-code rule)
- Errors & Fixes section with user-correction quoting (claude-code sections
  4/  + CompInt user-feedback emphasis)
2026-08-15 17:01:22 -07:00
Teknium c4bbb14e52 feat(compression): mechanical anchor index + region-scoping tripwire
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
  file paths, error strings, handles, URLs from the compacted region into a
  bounded indexed summary section. LLM-free, so needle identifiers cannot be
  paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
  golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
  summarizer input carries ONLY the compacted region (head/tail sentinels
  never reach the serialized turns body) in both legacy and lean modes.
2026-08-15 17:01:22 -07:00
Teknium 7a82457ede feat(compression): digest noise filter + FTS5 recovery sim + digest-aware query hints
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
  digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
  session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
  mine anchor identifiers
2026-08-15 17:01:22 -07:00
Teknium 8fe9025abd feat(compression): lean tail mode + recovery-aware eval arm
Lean mode (tail_mode='lean', default stays 'legacy'):
- tail budget = clamp(2.5% of window, 10K, 25K) instead of 0.20*window
- stale tail tool results demoted to session_search recovery stubs
- chunked identifier-preserving digests of the compacted region (map-reduce,
  pristine pre-prune tool contents)
- verbatim user messages embedded in summary (codex retention-by-role rule)
- deterministic session_search recovery footer

Eval: policies matrix gains lean + a '+recovery' arm giving the answerer one
simulated session_search round-trip against the archived region.
2026-08-15 17:01:22 -07:00
Teknium 33242d5ee0 feat(evals): compaction recall eval harness
Measures recall accuracy vs tokens retained across compaction policies.
Real transcripts in, LLM-generated recall exam from the summarized region,
per-policy answer+judge passes, scorecard out.
2026-08-15 17:01:21 -07:00
Teknium 9c58a78a7d feat(desktop): Capabilities-wide profile scoping + one-click hub installs on the Skills tab
Extends the Capabilities "Configuring:" profile selector (#86548) from
Tools/MCP to the WHOLE view — Skills, Tools, MCP, and Browse Hub now all
read and write the same selected profile — and brings Bot Mode's
one-click Skills Hub picker into the main Capabilities -> Skills tab.

Scope widening:
- skills/index.tsx: the selector renders once above whichever tab is
  active. Skills list, toggles, bulk ops, editor, and archive are scoped
  via the trailing-profile pattern; the skills RQ key gains the scope key.
  Toolsets analytics (usage badges) load per scope. SkillsHub and the
  hub picker are keyed/remounted per scope. Scope changes drop the open
  editor/archive dialog and in-flight analytics (same hazards as an
  app-wide profile switch); an app profile switch clears the override.
- hermes.ts: getSkills, setSkillEnabled, get/edit/deleteLearningNode,
  getUsageAnalytics, and all seven skills-hub fetchers take the optional
  trailing profile? (omitting preserves exact app-wide behavior).
- store/hub-actions.ts: runHubAction threads profile through spawn and
  getActionStatus polling so install/uninstall/update and their logs run
  against the scoped backend.
- hub.tsx: sources/search/preview queries keyed+scoped per profile;
  install/uninstall/update/scan route to the scoped profile.
- archive-skill-confirm-dialog.tsx: optional profile prop.

One-click hub installs (from Hermes-Bot-Mode):
- skills/embedded-hub-picker.tsx: collapsible, resizable iframe of the
  live Skills Hub (hermes-agent.nousresearch.com/docs/skills?embed=picker)
  on the Skills tab. Origin-checked hermes-skill-pick postMessages route
  through the standard hub action pipeline (background action, tailed
  log, optimistic flip, Skills list + slash-completion invalidation),
  scoped to the selected profile.
- i18n: skills.hub.picker* keys (en + zh; others fall back).

Tests: index.test.tsx — new case asserts picking a profile on the Skills
tab refetches skills scoped to it and routes toggles there (6/6);
toolset-config-panel 28/28. Full typecheck (3 tsconfigs) + eslint clean.
2026-08-15 16:21:56 -07:00
Teknium 951ae62ffc test(computer-use): pin 0.17+ split refs/content_refs merge behavior
Regression test for the _ref_map merge (salvaged from #79515): the live
0.19.3 driver splits action refs into refs[] while content_refs re-lists
every node with empty actions; the empty entries must not clobber the
action-bearing ones. Caught live: every typed click refused with
browser_ref_stale until the merge fix.
2026-08-15 15:36:19 -07:00
weisiwu bbd3462e25 fix(computer-use): merge refs+content_refs in _ref_map for cua-driver 0.17
cua-driver >= 0.17 splits the semantic_v2 snapshot payload: action-bearing
refs live in the `refs` array while `content_refs` carries every node
with EMPTY action lists. _ref_map only absorbed content_refs, so every
click/pointer/type ref was registered with no declared actions and all
typed-browser mutations failed with browser_ref_stale.

Merge refs + content_refs + snapshot.refs with set union so action info
is never dropped by an empty content entry. Verified against the 0.17
split format, the legacy refs-only format, the transitional dict format,
and the snapshot.refs fallback.
2026-08-15 15:36:19 -07:00
hermes-seaeye[bot] 17c7b0bebd fmt(js): npm run fix on merge (#87296)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-15 22:15:27 +00:00
Teknium 2e9dcb7c55 fix(gateway): session.history ships durable row_id stamps
The Desktop's content-based truncation-target resolution (and reactions)
address persisted turns by row_id, but session.history loaded the
transcript without include_row_ids=True, so _history_to_messages had no
stamp to forward and the projection silently stripped the one durable
address clients can use. Discovered live-testing the #87294 client flow:
resolveDurableRowId saw 0 stamped rows and degraded every edit to a
plain resubmit.
2026-08-15 15:09:11 -07:00
Teknium 3e8ab06107 fix(desktop): never send ordinal-only truncation — resolve durable row ids by content
Client half of #87059. The gateway now fails ordinal-only truncation
closed for durable sessions (#87150), which turned the mis-aimed cut into
a visible edit-resend error for any bubble without a bound rowId (edit
after an interrupted turn, unstamped resume). Make the Desktop always
produce a durable address or degrade safely:

- runRewindSubmit: when a truncation request lacks a durable address,
  resolve the target's row id by exact content against session.history
  (which ships row_id per persisted row). Resolution is
  exact-or-nothing: a unique text match wins; ambiguity is accepted only
  when the target is provably the newest persisted turn (the
  edit-after-interrupt shape). Anything else degrades to a PLAIN
  resubmit — never a guessed cut. The client ordinal is dropped either
  way (its space can diverge from the gateway's — the #87059 root).
- planReload/planRestore: degrade failed turns to a plain resubmit
  (extends the #86623 pattern to regenerate/restore) and carry the
  turn's persisted sourceText as the content key.
- rebindSurvivorRowIds: iterate the same failed-turn-aware ordinal
  space as the truncate math.
- session-tile-actions: reload goes through the shared runRewindSubmit
  primitive instead of a raw prompt.submit, so the tile surface gets the
  same discipline.
2026-08-15 15:09:11 -07:00
vondelomlo c2a50a8662 fix(desktop): skip failed turns in the backend-facing user ordinal space
A user turn whose submit failed keeps its optimistic bubble but never
reached the gateway, so counting it makes every later
truncate_before_user_ordinal overshoot the backend index (refused 4018,
regenerate dead for the rest of the session). Skip failed turns in the
one shared visible-user ordinal space (visibleUserMessageIndices) used by
truncate ordinals, ordinal->index resolution, and survivor-rowId
rebinding.

Based on #41275 by @vondelomlo, relocated onto the split
use-prompt-actions/ modules and widened from visibleUserOrdinal to the
shared index helper.
2026-08-15 15:09:11 -07:00
Teknium 20cf326bd1 fix(computer-use): align browser authorization with live-verified cua-driver 0.19.3 contract
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):

- bounded serve flags corrected: the daemon accepts
  --session-policy/--approve-session-policy, not the docs'
  --capability-manifest names (which it rejects). Verified end-to-end:
  a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
  TTY) and its token is a legacy compatibility path disabled by default
  on current drivers (per the live browser_prepare schema). Kept as a
  passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
  with cua-driver's trusted-launcher grant. config opt-in
  computer_use.grant_existing_profile: true appends
  --grant existing-profile to the standard-mode MCP spawn (MCP
  initialize verified accepting the flag). Default false = attachment
  keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
  ladder: config grant > bounded manifest > YOLO; token = legacy.
2026-08-15 15:04:32 -07:00
Teknium 48dd9c87cf feat(computer-use): user-facing authorization for cua-driver browser attachment
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:

- hermes computer-use browser-approve: CLI passthrough that mints
  cua-driver's five-minute single-use attachment token for one exact
  (pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
  browser_route), forwarded only for existing_profile and only as a
  non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
  private per-session embedded daemon launched with
  --capability-manifest/--approve-capability-manifest; missing manifest
  fails loudly. 'unrestricted' is deliberately NOT a config value —
  it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
  rungs and the isolated-profile-first default.

E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
2026-08-15 15:04:32 -07:00
Teknium fe0a56ed16 fix(nemo_relay): bound plugin Relay marks so a wedged native pipeline cannot stall the agent
The plugin's _Runtime.run_in_session wrapper serves every mark/event it
emits (turn start/end, approvals, subagent marks) and runs synchronously
on the agent's conversation thread. It passed no timeout, so the host's
run_in_session default (timeout=None) made each mark an UNBOUNDED native
call. With a wedged native Relay pipeline the agent blocked between API
calls with zero activity ticks — observed live 2026-08-15: two cron jobs
died at the 600s inactivity kill and a gateway chat session at 1800s,
all with last_activity="API call #N completed".

The core's scope push/pop/flush/close sites were bounded with
_SCOPE_OP_TIMEOUT after the 2026-08-10 delegation stall; the plugin's
event marks were the missed sibling class.

Changes:
- plugins/observability/nemo_relay: the wrapper always passes
  timeout=relay_runtime._SCOPE_OP_TIMEOUT (10s) to the host. A breach
  costs one telemetry span, never the agent; it also sets scope_errored
  (so close_session skips the ATIF export for the wedged session) and
  warns once so the sick pipeline is visible.
- tests/plugins/test_nemo_relay_bounded_marks.py: proves the budget
  reaches the host (fails on the pre-fix code — sabotage-verified),
  a TimeoutError flags the session and disables its export, and the
  generic error path keeps its scope_errored contract.
2026-08-15 14:31:32 -07:00
Teknium 92c998c86c fix(update): restore quarantined hermes.exe shims after no-op installs on Windows
Both quarantine wrappers (_run_quarantined_install in main.py and
_run_install_cmd in _install_repair.py) renamed live hermes*.exe shims
aside before invoking the installer, but only renamed them back on
FAILURE. A SUCCESSFUL install that never rewrites entry points — uv
audits an already-satisfied editable install as a no-op — left the
shims quarantined as hermes.exe.old.<ms> and `hermes` disappeared from
PATH after a green install (#75584; reproduced live on a Windows
install recovering from the #86735 self-lock deferral).

Switch both sites from except/re-raise to try/finally so restore runs
on every path. _restore_quarantined_exes already skips shims the
installer actually replaced, so fresh output is never clobbered and
failure behavior is unchanged.

Regression tests cover both wrappers x {no-op success, rewriting
success, failure}; the no-op cases fail on the previous code.
2026-08-15 14:24:35 -07:00
Teknium 763b10c320 fix(gateway): run session-finalize plugin hooks off-loop and bounded
Session-finalize hooks ran synchronously on the gateway event loop from
three call sites (shutdown drain, session-expiry watcher, /new reset).
A plugin hook doing heavy blocking work froze the whole loop: adapter
heartbeats stopped, the drain machinery could not run, and systemd
eventually SIGKILLed the process mid-export. Observed live on a
multi-day 4.7G session where the nemo_relay observability plugin
serialized a full-session ATIF trace inside on_session_finalize.

Changes:
- gateway/run.py: new GatewayRunner._finalize_session_off_loop()
  dispatches hermes_cli.lifecycle.finalize_session via the gateway
  executor under asyncio.wait_for (10s budget), mirroring
  _cleanup_agent_resources_off_loop (#53175). Shutdown finalize and
  the session-expiry watcher now use it.
- gateway/slash_commands.py: /new reset path uses the same helper.
- plugins/observability/nemo_relay: ATIF export is now bounded
  (HERMES_NEMO_RELAY_ATIF_EXPORT_TIMEOUT_S, default 30s) and skipped
  entirely for sessions whose Relay scope operations already errored
  (their exporter state is unreliable and the export can be
  pathologically slow).
- tests/gateway/test_finalize_session_off_loop.py: regression tests
  proving the loop stays live under a wedged hook and the budget is
  enforced.
2026-08-15 14:20:27 -07:00
hermes-seaeye[bot] 0c50bdbdea fmt(js): npm run fix on merge (#87257)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-15 20:44:31 +00:00
hermes-seaeye[bot] 9c7f92bf93 fmt(js): npm run fix on merge (#87251)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-15 20:38:02 +00:00
fangliquanflq 3863de3155 revert(gateway): keep profile truncation routing out of scope 2026-08-15 13:36:25 -07:00
fangliquanflq 0640fe7119 fix(gateway): route truncation writes to profile database 2026-08-15 13:36:25 -07:00
fangliquanflq 79b7d969d3 fix(gateway): reject unstamped durable ordinal rewinds 2026-08-15 13:36:25 -07:00
Yingliang Zhang eec4d5ec19 style: order deep-parent fixture import before siblings (perfectionist/sort-imports) 2026-08-15 13:35:32 -07:00
Yingliang Zhang 9f78a0d37f fix(desktop): preserve turn-elapsed timer across session switches
Rebased onto latest origin/main. Resolved conflicts in:
- use-session-actions.test.tsx: kept both HEAD's image-attachment test
  and PR's turn-clock restoration test (orthogonal features)
- use-session-actions/index.ts, gateway-event.ts, server.py,
  test_tui_gateway_server.py, test_protocol.py: adapted to HEAD's
  refactored structure while preserving PR's turn-origin tracking
2026-08-15 13:35:32 -07:00
Teknium 9859e8852f chore: map 807847218@qq.com -> Tommy00748 for attribution audit 2026-08-15 13:31:35 -07:00
Tommy00748 93a9b2318f feat(desktop): show per-turn wall-clock duration in the transcript
Each assistant reply now carries a small time badge below the message text
showing how long its turn took (message.start -> message.complete), so
users can gauge task latency at a glance without hovering.

The duration is computed renderer-side from the per-session turnStartedAt
timestamp the app already tracks and stamped onto the ChatMessage at
completion (successful and failed turns alike). It is not persisted
backend-side, so messages hydrated from history have no badge — matching
how reasoning-block durations already behave.

Also adds the assistant.thread.turnDuration i18n key across all five
locale files.
2026-08-15 13:31:35 -07:00
ducky_56789 a525bbed0e fix(update): avoid cryptography self-lock on Windows 2026-08-15 13:28:14 -07:00
Teknium 4b583e4476 fix(desktop): discriminate backend-confirmed turns with turnLive so the settle gate survives submit-time clock seeding
Follow-up hardening for the #74163 salvage: the no-payload settle gate used
turnStartedAt as "backend reported the turn live", but since the turn clock is
now optimistically seeded at submit (#86923), that signal is ambiguous.
Introduce ClientSessionState.turnLive, set on message.start, the running=true
session.info edge, and resume-onto-running paths; cleared by every settle.
The pre-start bail now gates on turnLive so a running=false heartbeat in the
submit gap still keeps the spinner up, while a genuinely started turn that
dies without a payload settles and unbricks the session.
2026-08-15 13:27:08 -07:00
briandevans 3e46389e40 fix(desktop): settle a turn that ends with no assistant payload
A turn that finishes without ever producing an assistant payload never
reaches message.complete, so session.info with running=false is the only
event that can release it. The busy=false branch bailed out of the state
update whenever awaitingResponse was still set and no payload had been
seen, so awaitingResponse and busy stayed latched until the app was
restarted.

That is not a cosmetic indicator. The per-session busy flag is
authoritative for isTargetSessionBusy, so submitPrompt and the slash
dispatcher silently returned false: the user typed, pressed Enter, and
nothing happened, with no error. Per-session state does not self-heal on
a session switch, so the session was effectively bricked. It reproduces
on a gateway crash mid-stream, a provider error before the first delta,
and an agent-build failure.

The bail still has a real job: submit arms busy/awaitingResponse
optimistically, so a running=false heartbeat landing in the gap before
the turn spins up is a pre-start report, not a finished turn, and
settling on it would drop the spinner and re-open the send guard
mid-flight. Gate the bail on turnStartedAt, which is stamped only once
the backend reports the turn live and cleared by every settle: null means
no turn was ever reported running, so keep waiting; non-null means the
turn started and is now reported finished, so settle.

On recovery, catch up the surfaces the missing message.complete would
have refreshed. The sidebar refresh stays unscoped so a background
session's working dot clears without the user opening it, and it fires on
the recovery edge only because the unchanged-state guard short-circuits
every later heartbeat. The transcript hydrate is scoped to the active
session so an idle background session does not cost a REST call.
2026-08-15 13:27:08 -07:00
kshitij 165c889e5b fix(cli): stop pushing Kitty keyboard protocol that breaks Ctrl+C
Commit 2ae7884ffa added _EXTENDED_ENTER_KEYS_SEQ which pushes both the
Kitty keyboard protocol (CSI >1u) and xterm modifyOtherKeys level 2
(CSI >4;2m) on supported terminals (Ghostty, iTerm2, WezTerm, kitty).

Under the Kitty keyboard protocol, Ctrl+C is encoded as \x1b[99;5u
(codepoint 99='c', modifier 5=Ctrl) instead of \x03 (ETX). prompt_toolkit
3.x has no mapping for \x1b[99;5u, so the sequence leaks as literal
text '[99;5u' on screen. Worse, the kernel's INTR mechanism looks for
the raw \x03 character, so SIGINT never fires either — Ctrl+C is
completely dead.

Fix: drop the CSI >1u push from _EXTENDED_ENTER_KEYS_SEQ, keeping only
modifyOtherKeys (CSI >4;2m). Shift+Enter still works via the
\x1b[27;2;13~ sequence that modifyOtherKeys produces and prompt_toolkit
already maps (Keys.ControlM). The exit reset sequence still pops both
modes for safety.

Refs #56684.
2026-08-15 20:57:18 +05:30
hermes-seaeye[bot] 45af7a71fc fmt(js): npm run fix on merge (#86935)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-15 11:39:10 +00:00
Teknium bda83a4737 style: satisfy desktop eslint (curly braces, import order) 2026-08-15 04:33:47 -07:00
Teknium 07161e1da4 chore: map contributor email for @RGerrish 2026-08-15 04:33:47 -07:00
Teknium ed4f91c4eb fix(update): widen stale-lock self-heal to banner and desktop; align with compare-API status
Follow-up on the cherry-picked gitlock work (#80501 by @RGerrish, covering
the #75133 / #75168 wedge first reported and fixed by @RelaxJonh):

- Drop the PR's ancestor-check halves in banner.py, update-count.ts and
  main.ts: superseded by the compare-API status recovery that landed in
  #86257/#86331 (ahead_by == 0 already reports local-ahead as up to date).
  The salvaged update_cmd.py check path keeps main's compare-API structure
  instead of the PR's tip-SHA-plus-ancestry print.
- Keep and wire clear_stale_git_locks() at the remaining wedge sites the
  original PR targeted: hermes update apply, hermes update --check, and the
  passive banner check.
- Add the desktop counterpart (electron/gitlock.ts) so checkUpdates() heals
  the same wedge instead of reporting fetch-failed forever; mirrored
  age + git-process guards; vitest coverage.

E2E verified: real --depth 1 clone with an aged .git/shallow.lock reproduces
"Unable to create '.git/shallow.lock': File exists"; clear_stale_git_locks
removes it and the fetch succeeds; a fresh lock (in-flight fetch) is
preserved.
2026-08-15 04:33:47 -07:00
RGerrish 7fe3bf042b fix(update): self-heal stale git locks and stop false 'update available' on shallow clones
Two related failure modes after a crashed/interrupted fetch on a shallow
clone (git clone --depth 1 installs):

1. STALE LOCK WEDGES EVERY FETCH. A killed fetch can leave .git/shallow.lock
   behind; every later 'git fetch' then fails with 'Unable to create
   .../shallow.lock: File exists'. 'hermes update --check' reported a hard
   fetch failure, and the passive banner check swallowed the exception and
   compared stale refs. Add hermes_cli.gitlock.clear_stale_git_locks(), a
   guarded sweep (age + git-process check so a live fetch is never yanked)
   wired into the check path, the apply path, and the banner's passive check.

2. SHALLOW TIP-SHA COMPARE FALSE-POSITIVES. On a shallow clone the check
   cannot count commits, so it compares tip SHAs. Local cherry-picks on top
   of the remote tip (e.g. re-applied local patches) make HEAD differ from
   origin/main even though HEAD already contains it — a false 'update
   available' banner. Add hermes_cli.gitlock.is_ancestor_of_head() and use
   'git merge-base --is-ancestor' in the CLI check and banner paths before
   reporting an update. Mirror in the desktop (update-count.ts gains an
   isAncestor input; main.ts probes merge-base --is-ancestor).

Tests: tests/test_gitlock.py (9) covering stale/young/no-lock/no-repo sweeps
and ancestry true/false; update-count.test.ts +3 for the isAncestor path.
2026-08-15 04:33:47 -07:00
Teknium c9a806e9d2 fix(browser): pin named sessions to their own tab on shared browsers
Follow-up to #86916. That fix gave named sessions their own daemon
(socket/log/pid) and their own provider browser — but on a SHARED local
Chrome / CDP browser, a fresh named daemon still attaches to the first
existing page, the same page a sibling daemon may hold. A named session
that never calls new_tab() could still stomp another's tab.

browser_exec now prepends a small preamble to the model's code for named
sessions on shared browsers: once per daemon process (marker keyed by
uid + BU_NAME + daemon pid), it creates a fresh tab via
Target.createTarget and switch_tab()s onto it before any model code
runs. Private per-name browsers (provider-keyed bu-named-<name>, or
direct-API Browser Use cloud) skip the preamble via an internal env
sentinel popped before launch — there's nobody to collide with, and the
extra tab would leak.

Best-effort by design: if the preamble's CDP calls fail, behavior
degrades to pre-fix, never blocks the exec.

E2E against a shared headless Chrome with the STOCK harness: two named
sessions issuing bare js() writes (no new_tab) kept distinct state
(EDGE-A/EDGE-B read back intact); the sabotage run without the preamble
reproduced the clobber (both read EDGE-B). Removes the dependency on the
upstream browser-harness tab-isolation PR for correctness.
2026-08-15 04:23:26 -07:00
Teknium f70277bc70 fix(desktop): arm the turn progress timer at submit instead of waiting for message.start
The progress box's timer (turnStartedAt) was only seeded by the backend's
message.start event, so the submit RPC -> gateway accept -> WS round trip
(seconds under load) showed no timer at all. Seed the per-session clock in
seedOptimistic at Enter-time; message.start now keeps an existing seed
(?? Date.now()) so backend-originated turns still arm there, the active-
session mirror reuses the seeded value instead of snapping to accept-time,
and the abort/failure paths retire the seed with the turn. Adds a
console.debug submit->accept latency probe at message.start.
2026-08-15 04:22:54 -07:00
Teknium 56e5385e96 fix(desktop): stop blocking the image submit path on vision pre-analysis
Desktop's send path pre-analyzed every attached image with the auxiliary
vision model serially, BEFORE dispatching the turn (_enrich_with_attached_
images). Users saw the progress box sit idle 25s-4min for messages that
take ~4s in the CLI; failures were silently swallowed, and touching
another session during the window killed the turn with zero API calls
(#83291). The prepended description also poisoned session auto-titles
(#82339).

Replace pre-analysis with _build_image_ref_message: reference the image
paths in the message and let the agent analyze them in-loop with
vision_analyze — its own retries, visible tool progress, and the turn
starts immediately. This is exactly how the @folder: reference path
already behaves, which responds in seconds for the same images.

Native-vision routing is unchanged; only the "text" mode (non-vision
main model / codex_app_server) loses the blocking submit-path calls.

Tests: tests/tui_gateway/test_image_ref_message.py (6 cases) including
a guard asserting the submit path never invokes the vision tool;
sabotage-verified (restoring the old blocking body fails 5/6).
2026-08-15 04:19:59 -07:00
Teknium 12859e9eb5 docs: add Connecting Desktop to Many Hermes Instances guide
New user-guide page for the multi-connection registry (Settings →
Connections): connection kinds + auth table, unique device names, v1
migration, union agent roster with @name-device handles, lazy sockets /
ssh connect-on-demand, fleet-wide updates (cloud excluded), plugin SDK
surface (host.connections/agents/ensureAgent/warmAgent, Bot Mode as
reference consumer), troubleshooting. Registered in sidebars.ts;
cross-linked from desktop.md and multi-profile-gateways.md.
Docusaurus build validated.
2026-08-15 04:16:31 -07:00
Teknium fe63353cbb fix(tools-config): stop clobbering image_gen.use_gateway on Nous-managed FAL picks
_select_plugin_image_gen_provider hardcoded image_gen.use_gateway = False.
The managed (Nous-subscription) flow writes use_gateway = True via
_write_provider_config, then this selector runs AFTER it — so picking FAL
through Nous Portal silently persisted provider: fal, use_gateway: false
and every generation billed the user's personal FAL_KEY instead of the
subscription (real incident: key drained to zero-balance lock while the
managed route sat unused).

Fix the class, not the site:
- _select_plugin_image_gen_provider gains the same use_gateway kwarg its
  video twin (_select_plugin_video_gen_provider) already had; all four
  call sites pass use_gateway=bool(managed_feature), matching the video
  call sites, TTS, STT, browser, and web.
- Active-provider detection (the checkmark in `hermes tools`): the
  image_gen_plugin_name branch now defers managed entries to the
  managed_feature branch and requires use_gateway OFF for direct-key
  entries — mirroring the video branch's existing guard, so a managed
  FAL pick and a direct-key FAL pick no longer both report active.

Runtime side (prefers_gateway("image_gen")) was already correct; the bug
was purely the setup-time writer.

Tests: new tests/hermes_cli/test_imagegen_managed_gateway.py (3 cases:
managed flag survives, direct pick still clears, image/video selector
contract parity). Sabotage-verified: restoring the hardcoded False fails
2/3. Neighboring hermes_cli provider/managed suites: 180 passed.
2026-08-15 04:16:31 -07:00
Teknium bb4f680f22 fix(browser): named browser_exec sessions compose with every backend
session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.

Now a named session composes with whatever browser source is configured:

- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
  pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
  too; previously a named daemon ignored it and fell back to scanning
  local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
  keys its own provider browser via the shared _get_session_info cache
  (bu-named-<name>), so each name gets its own cloud browser, the same
  name reuses one across calls and tasks, and unnamed calls keep the
  per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
  (provider resolution would double-session and double-bill).

Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.

E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
2026-08-15 04:05:57 -07:00
hermes-seaeye[bot] 3d5f450781 fmt(js): npm run fix on merge (#86915)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-15 11:02:12 +00:00
Teknium bfd9cef389 fix(worktree): deepen shallow clones so worktree cleanup can verify push state
The installer clones with --depth 1, so every default install is shallow.
In a shallow repo, an older worktree HEAD (a past snapshot of main) is
disconnected from current origin/main by the shallow boundary, so
'git log HEAD --not --remotes' misreports thousands of already-public
commits as unpushed. The fail-safe unpushed guard then preserves every
aged 'hermes -w' worktree forever, and the git-cherry squash-merge
escape hatch never rescues them (22k 'ahead' >> max_ahead=20).
Real incident: 21 of 25 hermes-* worktrees stuck on one install.

Fix at the root, one owner:
- _deepen_shallow_repo(): one-time blobless unshallow
  (fetch --unshallow --filter=blob:none; plain --unshallow fallback)
  run from the background startup pruner thread before classification,
  so history verdicts become correct and the backlog self-clears on the
  next 'hermes -w' startup. Fail-soft offline: keep preserving.
- _cleanup_worktree(): when the unpushed verdict comes from a shallow
  clone, say 'Shallow clone — cannot verify push state' instead of the
  misleading 'has unpushed commits' message.
- Document the shallow caveat on _worktree_has_unpushed_commits (the
  primitive stays conservative on purpose).

Tests: real shallow clone over file:// reproducing the disconnect shape,
covering detection, deepen+verdict flip, pruner E2E reap, offline
fail-soft preserve, full-clone noop, and genuine-unpushed-work survival.
Sabotage-verified: the E2E test fails with the deepen call disabled.
2026-08-15 04:02:07 -07:00
Teknium 3671529e9d feat(sdk): export route-decoupled McpTab + ToolsetConfigPanel for plugins (#86896)
* feat(sdk): export route-decoupled McpTab + ToolsetConfigPanel for plugins

Runtime plugins can only import from @hermes/plugin-sdk, but the real
Capabilities components (the full per-toolset config panel and the full MCP
tab with OAuth/API-key setup) were never exported there — so a plugin could
only reimplement bare checkbox lists. Export both, route-decoupled so they
render safely outside the Settings react-router context:

- toolset-config-panel.tsx: useOptionalNavigate() wraps useNavigate in try/catch
  (returns null with no router); the 'manage keys' deep link becomes a no-op
  when embedded outside Settings. In-Settings behavior unchanged.
- use-deep-link-highlight.ts: useOptionalSearchParams() degrades to inert params
  with no router (shared by 4 in-Settings callers incl. McpTab; identical there).
- sdk/index.ts: export { McpTab }, export { ToolsetConfigPanel }, export type
  HermesGateway, and host.getGateway() returning the live $gateway instance
  (McpTab takes a HermesGateway prop; plugins had no way to get the instance).

Both components are already profile-aware (profile?: null|string, #86548), so a
plugin can scope them to a specific bot profile. tsc: 0 errors (unchanged from
baseline). Enables Hermes-Bot-Mode to show the real Tools+MCP config in the bot
editor instead of checkbox stand-ins.

* fix(lint): sort the new SDK exports into perfectionist/sort-exports order

CI check:lint failed — the capabilities exports were grouped by comment instead
of interleaved into the file's path-sorted export list. Place them at their
natural-ascending positions: ToolsetConfigPanel (@/app/settings) after @/app/routes,
McpTab (@/app/skills) after @/app/shell/*, HermesGateway (@/hermes) after
@/contrib/types. Verified 0 adjacent-unsorted export pairs.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 03:55:51 -07:00
fangliquanflq b70bd03b3c fix(updater): mark PID probe POSIX-only 2026-08-15 03:54:07 -07:00
fangliquanflq d528f4da00 fix(updater): defer native parser imports after recovery 2026-08-15 03:54:07 -07:00
fangliquanflq 97051703ae fix(updater): keep recovered retries native-safe 2026-08-15 03:54:07 -07:00
fangliquanflq 8f5e5e49a2 fix(updater): recover deferred installs on update retry 2026-08-15 03:54:07 -07:00
Teknium efdd715f8a fix(gateway): resolve the profile-aware scheduled-task name in the reaper guard
Follow-up to the salvaged #86823: the guard queried a hardcoded
"HermesGateway" task, but `hermes gateway install` registers
Hermes_Gateway (Hermes_Gateway_<profile> for named profiles) via
gateway_windows.get_task_name(). Query that name so the supervisor
guard is active on standard installs; fall back to the default literal
if the module import fails. Test now asserts the profile-aware name is
what reaches the task-state query.

Also corrects the cherry-picked commit's placeholder author email to
the contributor's GitHub noreply address.
2026-08-15 03:53:51 -07:00
hutao562 2795b2ab9f fix(gateway): treat a Running HermesGateway scheduled task as a supervisor
The orphan-reap sweep (_reap_unsupervised_gateway_orphans) must not kill a
gateway that Windows Task Scheduler is actively managing. The existing
services.exe parent-chain backstop fails open: when the Task-launched conhost
bootstrap has already exited, Windows does not reparent the gateway, the
chain breaks, and the supervised gateway is treated as an orphan. The reaper
then writes the planned-stop marker, the gateway exits cleanly with code 0,
and the scheduler never restarts it (RestartCount only fires on non-zero
exit) — silently killing A2A/messaging on every desktop-app launch.

Querying the task's own state is the authoritative signal and closes the gap
without depending on process ancestry: if HermesGateway is Running, skip the
reap entirely. Uses PowerShell Get-ScheduledTask (English State enum,
locale-stable) rather than schtasks (localized output + codepage mangling).
2026-08-15 03:53:51 -07:00