Commit Graph

25029 Commits

Author SHA1 Message Date
chelsealong ec013b76db fix(desktop): split strict connection routing from passive roster lookup
botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.

Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.
2026-08-24 03:22:28 -07:00
chelsealong 09529afdd2 fix(desktop): group chats no longer crash when a member's connection is deleted
botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).

botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.

Fixes #93492
2026-08-24 03:22:28 -07:00
Teknium e8d5660bae fix(desktop): bound the boot and gateway-switch resolveGatewayWsUrl awaits too (#93454)
Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await
exists on the soft gateway-switch path and the initial boot() path. Bound
both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged
IPC round-trip fails into the existing retry paths instead of hanging the
switch or the 'Starting Hermes…' screen forever.
2026-08-24 03:22:03 -07:00
chelsealong 5ef205d855 fix(desktop): bound the revalidateConnection() await too (#93454)
attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger #93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.

Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.
2026-08-24 03:22:03 -07:00
chelsealong 17f8e24c81 fix(desktop): bound reconnect awaits so a stuck IPC round-trip can't latch the UI frozen
After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.

Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.

Fixes #93454
2026-08-24 03:22:03 -07:00
chelsealong 9c013eaaf8 fix(dashboard): follow scroll on implicit active-session resume (#93518)
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes #93518.
2026-08-24 03:21:49 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 85cd576b06 test(bot-relay): match delivery CLI by basename in argv filters
CI runners have a real hermes sibling next to the venv python, so
local_delivery_command now resolves an absolute path there — the exact
argv filters in the retry-policy fakes and the relay-methods pins must
match by basename instead of the literal "hermes", mirroring the
_delivery_lock matcher.
2026-08-24 03:21:37 -07:00
liuhao1024 c099ef05de fix(bot-relay): Windows path SyntaxError in waiter + PATH-less delivery ENOENT
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):

1. waiter_command embeds the reply path in generated python -c source
   with !r. repr escapes each backslash, but the Windows execution layer
   folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
   and SyntaxErrors the whole waiter script. Raw-string literals keep
   the folded single backslash a literal; POSIX paths have no
   backslashes so the prefix is a no-op there, and \' inside a raw
   literal still cannot terminate the string, keeping the #93091
   injection defense intact.

2. local_delivery_command hardcoded "hermes", relying on PATH — absent
   in service contexts (systemd units, desktop launchers, non-login SSH
   shells), so delivery died with ENOENT. It now resolves the CLI next
   to this gateway's own interpreter (venv bin/Scripts sibling,
   hermes.exe on Windows) with a bare-name fallback. The #93091
   per-profile turn-lock recognition in bot_mode_dm now matches the CLI
   element by basename (split on both separators) so resolved absolute
   paths still take the lock instead of silently bypassing it.

Fixes #93590
2026-08-24 03:21:37 -07:00
liuhao1024 8d17060249 fix(cli): hard-exit the Windows update hand-off child once work is durable
The re-exec'd venv child spawned by
_reexec_dependency_sync_off_windows_shim completes every update step —
the receipt records success / "completed at command boundary" — but then
hangs in interpreter shutdown on a leftover non-daemon thread, freezing
the PowerShell window for minutes after "Update complete!". On the
hand-off path only (HERMES_UPDATE_REEXEC=1), after the receipt is
finalized, the update lock released, and stdio restored, flush and
os._exit(code) instead of unwinding — the same treatment #79040's cron
workaround applies. SystemExit codes (including early refusals)
propagate to the hard exit; real exceptions keep the normal raise path
so tracebacks still print. Non-hand-off invocations are untouched: the
marker env is set solely when the shim spawns the child.

Fixes #93581
2026-08-24 03:21:28 -07:00
Teknium 93bf6f7225 fix(cli): fail closed on empty fleet probe across all pre-update liveness signals (#93406)
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.

Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
  same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries

Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.

Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.

Builds on RelaxJonh's #93410. Fixes #93406
2026-08-24 03:21:18 -07:00
RelaxJonh d74bbb9bd9 fix(cli): treat empty fleet probe as incomplete when gateways were restarted (#93406)
collect_fleet_versions() swallows every probe exception via
logger.debug() and returns whatever accumulated — which can be an
empty list.  print_fleet_version_matrix([]) returns False (no rows
to report), so the update exits 0 with "success" even though no
gateway was actually verified.

After the restart phase touches live gateways (restarted_services or
killed_pids is truthy), an empty fleet snapshot means verification
failed, not that everything is healthy.  Treat it as incomplete so
the receipt records "partial" and the exit code is 1.

Fixes #93406
2026-08-24 03:21:18 -07:00
aniruddhaadak80 a26154aceb fix(batch_runner): teach the resume content scan to honor discard tombstones
Complete the #93527 fix: the tombstone now carries the human prompt
text via _entry_prompt_text (handling flat prompt, ShareGPT, and
chat-style shapes), _scan_completed_prompts_by_content counts
discarded rows as completed instead of only reading ShareGPT
conversations, and the merge step reports excluded tombstones in the
combined-count summary. Adds a dedicated regression suite covering the
tombstone round-trip, the all-discarded-batch resume path, and merge
exclusion.

Salvaged from #93579 (issue reporter's PR), building on #93542.
Fixes #93527
2026-08-24 03:21:07 -07:00
chelsealong 316d52faf2 fix(batch_runner): write a discard tombstone so resume skips no-reasoning prompts
The no-reasoning discard branch in _process_batch_worker continued
before writing any JSONL row, so run(resume=True) — which filters
solely via _scan_completed_prompts_by_content over batch_*.jsonl —
never saw discarded prompts and re-ran them at full cost on every
resume. Write a tombstone row on discard, exclude tombstones from the
trajectories.jsonl merge, and report discarded_no_reasoning in
final statistics.

Salvaged from #93542.
Fixes #93527
2026-08-24 03:21:07 -07:00
Teknium 4b622bbc4b test(model_metadata): lock in max_tokens last-resort fallback + cache self-heal
Adjust the #93423 max_tokens-only regression test to the merged policy:
max_tokens stays as an explicit LAST-RESORT fallback (some local servers
report nothing else) instead of being dropped entirely, and add coverage
that _reconcile_local_cached_context_length rewrites a cache entry
poisoned by the old probe (393216) upward to the real window (1048576)
once the probe is fixed.

Co-authored-by: pju-hoge <grkt@ppmz.com>
Co-authored-by: re-ITRT <1940428933@qq.com>
2026-08-24 03:20:57 -07:00
re-ITRT a0c802c02c fix(model_metadata): stop misreading max_tokens as context length in local probe
The local-endpoint context probe (_query_local_context_length_uncached)
treated max_tokens — an output-completion cap — as a candidate for the
model's context window. For OpenAI-compatible gateways that advertise a
1M context via context_size / max_input_tokens alongside a smaller
max_tokens output cap (e.g. TokenHub serving deepseek-v4-flash:
context_size=1048576, max_input_tokens=1048576, max_tokens=393216),
Hermes mis-detected the window as 393,216 and — because loopback
endpoints are reconciled against a live probe — actively overwrote a
previously-correct 1M cache entry.

- Add context_size and max_input_tokens to both /v1/models probe
  candidate lists (single-model detail and list branches).
- Remove max_tokens from the context-length candidates; it remains
  handled separately as an output cap (_MAX_COMPLETION_KEYS).

Adds regression tests covering context_size/max_input_tokens priority
over max_tokens and the max_tokens-only (no real context key) case.
2026-08-24 03:20:57 -07:00
Kyzcreig 4d729e4b31 fix(model-metadata): local ctx probe must not read max_tokens as the context window
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.

Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
2026-08-24 03:20:57 -07:00
carlotestor 394f0f0902 fix(context): prefer max_input_tokens over max_tokens for Anthropic proxies
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:

  max_input_tokens = context window (e.g. 1M for claude-fable-5)
  max_tokens       = max output     (e.g. 128k)

That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).

Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
2026-08-24 03:20:57 -07:00
Teknium f28718c331 chore: map contributor email for shanthans-es 2026-08-24 03:20:46 -07:00
Shanthan Subramaniam e210fd8c1f fix(gateway): resolve PairingStore's default pairing dir lazily, not at import time
PairingStore(profile=None) resolved its storage directory from the
module-level PAIRING_DIR constant, which was computed exactly once, at
module import time. A long-lived process (the gateway, started once at
container/process boot) can import this module before HERMES_HOME or a
profile's context is fully established, freezing PAIRING_DIR to a wrong
value for the rest of that process's lifetime -- even though a freshly
started, short-lived process (e.g. the `hermes pairing` CLI) re-imports
the module later with the environment already correct.

That asymmetry is exactly what made pending pairing codes issued by the
gateway process unrecoverable (the pending-code write landed under the
stale, wrong directory) while CLI-invoked writes to the same nominal
directory kept working -- see #93449 for the full writeup and a live
reproduction. tests/hermes_cli/test_dashboard_admin_endpoints.py already
carried a comment acknowledging this exact staleness in passing ("the
module-level PAIRING_DIR is bound at import"), and
TestProfileScopedStorage::test_default_store_uses_global_dir's own
comment describes working around it rather than it being intentional
behavior -- this fixes the underlying cause both were compensating for.

The profile-scoped branch already resolved its directory lazily inside
__init__ (matching this docstring's claim that resolution is lazy); this
brings the non-profile branch in line with it.

Fix keeps PAIRING_DIR as the same test seam already used throughout the
test suite (`patch("gateway.pairing.PAIRING_DIR", tmp_path)`, ~30 call
sites) unchanged: it's now a None sentinel instead of an eagerly computed
path, and a new _default_pairing_dir() helper resolves it fresh on every
call, honoring a patched (non-None) value when one is set. No existing
test needed to change.

Added a regression test that does not patch PAIRING_DIR directly and
instead exercises the real lazy-resolution path across two different
HERMES_HOME values in the same process -- confirmed it fails on the
pre-fix code (gets stuck with whatever the first PairingStore() call in
the test session happened to see) and passes with the fix.

Verified: tests/gateway/test_pairing.py (39, incl. the new one),
tests/hermes_cli/test_pairing.py, and tests/tools/test_pr_6656_regressions.py
all pass unmodified.
2026-08-24 03:20:46 -07:00
Teknium 905edf37a4 refactor(cli): add static fallback to parser-derived value-flag helper
Mirror the full update_cmd._holder_value_flags precedent: derive both
top-level value-flag sets from build_top_level_parser() with a cached
frozenset, and fall back to a handwritten snapshot if parser
introspection ever fails, so argv classification keeps working on a
broken tree. Parity test pins the derived sets against the live parser
so drift fails CI.

Builds on #93551 (fangliquanflq) and #93570 (aniruddhaadak80) for #93530.
2026-08-24 03:20:37 -07:00
aniruddhaadak80 1007296ca0 fix(cli): add --reasoning to both top-level value-flag sets
--reasoning takes a value (metavar=LEVEL in _parser.py) but was absent
from _TOP_LEVEL_VALUE_FLAGS (used by _first_positional_argv) and from
_apply_profile_override's value_flags set. Every invocation like
"hermes --reasoning high chat ..." therefore misclassified "high" as the
first positional, and _plugin_cli_discovery_needed() forced full eager
plugin CLI discovery at argparse-setup time - the documented startup
cost paid on every use of the reasoning override.

Add --reasoning to both sets, and add a parser-derived parity regression
test so future drift between the hand-maintained sets and
build_top_level_parser() fails CI instead of silently degrading startup
(the exact drift class AGENTS.md bans).

Fixes #93530

(cherry picked from commit 9280617ab8b308f9a8cf947f1617276a6bc8eb4f)
2026-08-24 03:20:37 -07:00
fangliquanflq 694550e486 fix(cli): derive top-level value flags from parser
(cherry picked from commit f2e5a1388115615fafde49a5f2144e5855a81d89)
2026-08-24 03:20:37 -07:00
fangliquanflq 3963fc6f21 fix(config): stop reporting stripped v15 defaults
(cherry picked from commit 4c6b67ec371b16c15e9ffbb91bfb47a504e913fe)
2026-08-24 03:20:28 -07:00
Teknium a7aa814c42 fix(tools): widen the command-position anchor to the whole hardline class
#93392 was not just one pattern: every hardline rule with a bare \b anchor
fired on its token anywhere in the command line, including inside quoted
prose handed to echo, git commit -m, or gh --body. Anchor the
command-name-token rules and quote-mask the positionless ones:

- dd-to-block-device and kill -1 get the same _CMDPOS anchor as the
  format/rm/shutdown families, keeping their argument tails.
- redirect-to-block-device and the fork bomb have no command-name token to
  anchor (`>` appears mid-command; the bomb is a function definition), so
  they now match a quote-masked variant (_mask_quoted_prose) where quoted
  string content is blanked. $() and backtick spans inside double quotes
  stay raw (the shell executes them), and any command whose command-position
  words include a shell carrier (sh/bash/zsh/ksh/dash -c, eval, source, .)
  is scanned unmasked -- quoting is not a bypass. bash/sh -c payloads also
  still surface as raw detection variants via _execution_flag_findings.

Regression tests cover both directions for every touched pattern: quoted
prose passes, and every true-positive shape (bare, ; && | separators,
sudo/env prefix, $(), backticks, sh -c/bash -c/eval payloads) stays on the
unconditional floor.
2026-08-24 03:20:14 -07:00
liuhao1024 8163c8731b fix(tools): anchor the mkfs hardline pattern to command position
mkfs was the only HARDLINE_PATTERNS entry without a _CMDPOS anchor, so
the unconditional floor blocked any command that merely mentioned the
token inside quoted prose — `echo "does this workflow use mkfs
anywhere?"` was refused outright (#93392) instead of running the echo.

Anchor mkfs to command position like every sibling entry (rm root-
delete, shutdown family, dd): it matches at the start of a command,
after separators, or behind sudo/env/exec/nohup/setsid wrappers, and
no longer fires on argument-position mentions. The quote-aware
_mark_command_starts pass already keeps separators inside quoted
strings from looking like command starts, and \b still protects
mkfs_helper-style names.
2026-08-24 03:20:14 -07:00
aniruddhaadak80 7befc1d2dd fix(gateway): route platform authorization reads through the profile secret scope
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):

- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
  WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
  GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
  GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
  hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
  _sig_secret helper; empty scoped group list previously meant "drop all
  groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
  WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
  while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
  startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
  though its sibling dm/group reads already used the scoped _getenv

All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.

Fixes #93522
2026-08-24 03:20:06 -07:00
chelsealong d7e4204e77 fix(gateway): scope multiplex-profile authorization reads (weixin/yuanbao/wecom)
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.

Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.

Fixes #93522.
2026-08-24 03:20:06 -07:00
Teknium 8eba0d2fd6 chore: map dougatbuck contributor email 2026-08-24 03:19:47 -07:00
dougatbuck 855e191d89 test(desktop): document connection-scoped filesystem keys 2026-08-24 03:19:47 -07:00
dougatbuck 9abeb89a49 fix(desktop): guard stale files refreshes 2026-08-24 03:19:47 -07:00
dougatbuck 30150fd082 fix(desktop): isolate files across connections 2026-08-24 03:19:47 -07:00
kshitijk4poor 0eda2ba0c8 fix: remove dead code, deduplicate error constants, fix skill key check
Follow-up to PR #92189 salvage:
- Remove unused job_no_agent_without_script() function (dead code)
- Replace inline NO_AGENT_WITHOUT_SCRIPT_ERROR string in _validate_job_mode_invariants with the constant
- Replace scheduler inline reason string with EMPTY_PAYLOAD_ERROR constant
- Add 'skill' (singular) to job_payload_is_empty 'in job' presence check
2026-08-24 15:47:42 +05:30
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
kshitijk4poor b57530afee refactor: collapse context_back/effective_allow_back, fix draw_header docstring
Follow-up cleanup for salvaged PR #92838:
- Collapse redundant context_back and effective_allow_back into a single
  allow_back variable (they were always identical, never reassigned)
- Update _run_curses_menu docstring to reflect draw_header's actual
  signature including search and back_enabled kwargs
2026-08-24 15:45:52 +05:30
fyzanshaik 8345effc4e fix(cli): support reliable setup menu navigation
Decode Ghostty/Kitty enhanced selection and cancellation keys, make setup cancellation terminal, and add cross-terminal previous-step navigation to setup and model flows.

Refs #92833
2026-08-24 15:45:52 +05:30
Teknium d9a48f656a fix(desktop): scheduled jobs on sleeping profiles keep firing
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
2026-08-24 03:14:30 -07:00
Teknium 94af11a920 chore: map contributor email for jeremyrandria-debug 2026-08-24 03:14:24 -07:00
Teknium da57f49237 fix(desktop-update): shim window only opens in the user's own browser family
A Safari/Firefox/Helium user who merely had Chrome installed watched Chrome open on every desktop update (community report). start_ui now checks the system default browser (LaunchServices https handler on macOS, xdg-settings on Linux) and skips the shim window unless the default is Chromium-family; notify_fallback and the durable result file still carry the outcome. Detection is best-effort: any failure keeps the old behavior.
2026-08-24 03:14:24 -07:00
jeremyrandria-debug 60bb2bb719 fix(update): auto-close desktop-update shim window after error/manual outcomes
On error/manual outcomes stop_ui('leave-window') kept the browser shim
window open indefinitely, so an aborted update left a Chrome window on
screen until the user closed it by hand; repeated update attempts piled
up more windows.

stop_ui now always closes the shim. leave-window paths keep it up for a
short grace period (HERMES_UPDATE_SHIM_GRACE_SECONDS, default 15) so a
watching user can read the message, then close it. The success path is
unchanged. The error/manual outcome is durably written to
.hermes-update-result.json and surfaced in a dialog on the next Desktop
boot, so closing the shim loses no information.
2026-08-24 03:14:24 -07:00
liuhao1024 eb21740b06 Also skip Brave: its P3A bar paints over the update shim window
Brave renders its own P3A privacy-notice bar ("Got it" / "Disable" /
"Learn more") at the top of the throwaway-profile window the posix shim
opens, cramped to unreadability at the shim's small size - the same
window-pollution class as Edge's MSA sync notice, and equally immune to
the throwaway --user-data-dir (#88682). Drop Brave from the candidates
on both platforms; Chrome and Chromium stay.

Covers #88682 on top of #88410
2026-08-24 03:14:24 -07:00
liuhao1024 f329f9e40d fix(desktop-update): never render the posix update shim in Edge
Edge's OS-level Microsoft-account integration signs even a fresh
throwaway profile into the user's MSA and renders its own "syncing
your browsing data" notification — the user's MSA email included —
inside the update window, which is titled "Hermes" (#88410). The
throwaway --user-data-dir start_ui already passes cannot block that
OS-account path, so the only reliable protection is to not pick Edge
at all: drop it from the browser candidates on both macOS and Linux.
The update UI is a best-effort layer — with no other Chromium-family
browser installed, start_ui falls back to its existing
"no renderer; skipping UI" path and the update itself is unaffected.

Fixes #88410
2026-08-24 03:14:24 -07:00
Teknium b90289b046 fix(desktop): update flow nudges a gateway reconnect; wake path probes instead of blind-closing
After a remote backend update restarts the gateway, the window WebSocket often dies without a close event (SSH/tailscale tunnels) and users force-quit to recover. finishBackendApply now nudges the registered reconnect handler, which rides the new ping liveness probe: healthy sockets are left untouched, dead ones are force-closed and re-dialed. The old blind gateway.close() on every wake signal is removed in favor of the probe.
2026-08-24 03:14:19 -07:00
Owenz-creator cdd37035d9 fix(desktop): probe half-open gateway socket on wake and reconnect
macOS sleep/wake (or a silent network drop) can leave the renderer's
WebSocket half-open: no close event fires, so connectionState stays
'open' while every RPC hangs until its per-call timeout. prompt.submit's
timeout is 30 minutes, so the user's next message reads as "enter does
nothing until I restart the app".

- Add a minimal ping RPC (tui_gateway/server.py) answered synchronously
  on the WS reader thread.
- On wake signals, reconnectNow now probes the open-looking socket with a
  5s-bounded ping and force-closes it on failure, letting the existing
  reconnect machinery (backoff, tile rebinding, session refresh) take
  over. A pre-ping backend answering -32601 is treated as healthy.
- Tests: half-open socket force-reconnects; healthy socket untouched;
  method-not-found backend untouched; backend ping envelope contract.
2026-08-24 03:14:19 -07:00
Teknium 29c5a12e04 fix(cron): warn loudly when the due-scan removes a consumed one-shot that already ran (#93524)
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
2026-08-24 03:13:30 -07:00
Teknium a1c5e515b7 fix(desktop): SkillsView tests no longer cascade-fail on slow CI runners
The whole test file legitimately runs ~14s on CI (heavy dynamic import paid
by the first test), brushing the global 15s per-test budget. Slow runners
tip the first test over and cascade-fail all 11 — hit twice in a row on PR
#93612 and on a main run in the same hour. Raise the file's describe-level
timeout to 60s; individual tests still run in milliseconds locally.

Sabotage-verified: timeout:1 fails all 11, 60s passes all 11.
2026-08-24 03:13:24 -07:00
kshitijk4poor 45aa0dc33e fix(cli): widen prompt_toolkit fallback to catch any runtime failure
Widen the exception guard from OSError to Exception (re-raising
KeyboardInterrupt/EOFError first) so any prompt_toolkit runtime
failure degrades to input() — matching the established pattern in
masked_secret_prompt.  ValueError and RuntimeError can arise from
exotic stream wrappers or event-loop issues with the same root cause:
prompt_toolkit cannot attach stdin on the terminal.

Add test_line_input_falls_back_to_input_on_any_prompt_toolkit_failure
covering the ValueError case.
2026-08-24 15:42:58 +05:30
wo-o a88cbe10fb test(cli): add regression test for line_input OSError fallback
Cover the prompt_toolkit runtime-failure path added in the fix commit: a
tty-reporting stdin where prompt_toolkit raises OSError(22) (macOS kqueue
EINVAL on fd 0 under curl|bash installs) must degrade to input() instead
of aborting the setup wizard.
2026-08-24 15:42:58 +05:30
wo-o e8ab3b075d fix(cli): fall back to input() when prompt_toolkit can't attach stdin
line_input() only guarded against a missing prompt_toolkit (ImportError),
not against prompt_toolkit failing at runtime. On some terminals isatty()
returns True but the asyncio event-loop selector rejects registering stdin
(macOS kqueue raises OSError EINVAL / 'Invalid argument' for fd 0), so
prompt_toolkit's Application.run() crashes while attaching its input.

This aborted 'hermes setup' at the first plain text prompt. Telegram hit it
first because its automatic/manual selection uses prompt() rather than the
curses-based prompt_choice() the other platforms use, but every text prompt
shared the same failure.

Catch OSError from the prompt_toolkit path and fall back to the built-in
input() reader, which needs no selector and works in cooked mode. The
prompt_toolkit raw-mode context manager restores terminal state on the way
out, so the fallback reads cleanly.
2026-08-24 15:42:58 +05:30
kshitijk4poor 9857bcba5c fix: widen secure_parent_dir to skip entire install tree
Replace hardcoded /opt/hermes check with dynamic install-tree detection
using Path(__file__).resolve().parent. This catches ALL install paths
(Docker /opt/hermes, apt /usr/local/lib/hermes-agent, git clone, custom)
instead of just the Docker image path. Also covers subdirectories of the
install tree, not just the top-level dir.

Add regression test test_install_tree_skipped to verify both the install
root and subdirectories are excluded from chmod.

Add contributor email mapping for bradmarshall987.

Follow-up to PR #93050 by @bradmarshall987.
2026-08-24 15:42:09 +05:30