Commit Graph

14240 Commits

Author SHA1 Message Date
Teknium 32fb12a235 fix(hindsight): let hindsight_retain convey event time via occurred_at
Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the
hindsight_retain tool schema, threaded into the retain item's timestamp
field. When absent, the item timestamp defaults to the configured event
clock (base from PR #82928 by @ragingbulld, authorship preserved) so the
Hindsight server can resolve relative time phrases; previously no item
timestamp was ever sent and temporal memories landed with null
occurred_start/occurred_end.

Fixes #93568. Salvages #82928.
2026-08-24 03:23:52 -07:00
Fabiao De Gongniu 497d6d5a66 fix(hindsight): harden event timestamps 2026-08-24 03:23:52 -07:00
Fabiao De Gongniu 97850afa33 fix(hindsight): send configured event timestamps
Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.
2026-08-24 03:23:52 -07:00
Teknium c9e2a46df6 fix(pricing): support Gemini context-tiered rates in pricing snapshot (#93469)
The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).

- Add optional tier fields to PricingEntry: tier_threshold_tokens,
  input/output/cache_read_cost_per_million_above (None = flat, falls
  back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
  request once usage.prompt_tokens (input + cache read + cache write)
  exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
  gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
  above 200k).
- Flat entries are untouched: no threshold means no behavior change.

Reported and tier-field shape designed by @tornike14 (#93469).

Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
2026-08-24 03:23:07 -07:00
Teknium a87d314e44 fix(auth): malformed OpenRouter env key no longer shadows valid credential-pool key
A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).

- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
  that mismatch a declared prefix, logging a WARNING naming the env var and
  expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
  only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
  are never rejected. Valid env keys still win over the pool (precedence
  unchanged).

Fixes #93593
2026-08-24 03:22:57 -07:00
Teknium d8d1e18ab9 test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
2026-08-24 03:22:48 -07:00
chelsealong b3730153c3 fix(terminal): cap watch_patterns notifications over a process's lifetime
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
2026-08-24 03:22:48 -07:00
chelsealong 9c013eaaf8 fix(dashboard): follow scroll on implicit active-session resume (#93518)
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes #93518.
2026-08-24 03:21:49 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 85cd576b06 test(bot-relay): match delivery CLI by basename in argv filters
CI runners have a real hermes sibling next to the venv python, so
local_delivery_command now resolves an absolute path there — the exact
argv filters in the retry-policy fakes and the relay-methods pins must
match by basename instead of the literal "hermes", mirroring the
_delivery_lock matcher.
2026-08-24 03:21:37 -07:00
liuhao1024 c099ef05de fix(bot-relay): Windows path SyntaxError in waiter + PATH-less delivery ENOENT
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):

1. waiter_command embeds the reply path in generated python -c source
   with !r. repr escapes each backslash, but the Windows execution layer
   folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
   and SyntaxErrors the whole waiter script. Raw-string literals keep
   the folded single backslash a literal; POSIX paths have no
   backslashes so the prefix is a no-op there, and \' inside a raw
   literal still cannot terminate the string, keeping the #93091
   injection defense intact.

2. local_delivery_command hardcoded "hermes", relying on PATH — absent
   in service contexts (systemd units, desktop launchers, non-login SSH
   shells), so delivery died with ENOENT. It now resolves the CLI next
   to this gateway's own interpreter (venv bin/Scripts sibling,
   hermes.exe on Windows) with a bare-name fallback. The #93091
   per-profile turn-lock recognition in bot_mode_dm now matches the CLI
   element by basename (split on both separators) so resolved absolute
   paths still take the lock instead of silently bypassing it.

Fixes #93590
2026-08-24 03:21:37 -07:00
liuhao1024 8d17060249 fix(cli): hard-exit the Windows update hand-off child once work is durable
The re-exec'd venv child spawned by
_reexec_dependency_sync_off_windows_shim completes every update step —
the receipt records success / "completed at command boundary" — but then
hangs in interpreter shutdown on a leftover non-daemon thread, freezing
the PowerShell window for minutes after "Update complete!". On the
hand-off path only (HERMES_UPDATE_REEXEC=1), after the receipt is
finalized, the update lock released, and stdio restored, flush and
os._exit(code) instead of unwinding — the same treatment #79040's cron
workaround applies. SystemExit codes (including early refusals)
propagate to the hard exit; real exceptions keep the normal raise path
so tracebacks still print. Non-hand-off invocations are untouched: the
marker env is set solely when the shim spawns the child.

Fixes #93581
2026-08-24 03:21:28 -07:00
Teknium 93bf6f7225 fix(cli): fail closed on empty fleet probe across all pre-update liveness signals (#93406)
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.

Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
  same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries

Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.

Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.

Builds on RelaxJonh's #93410. Fixes #93406
2026-08-24 03:21:18 -07:00
aniruddhaadak80 a26154aceb fix(batch_runner): teach the resume content scan to honor discard tombstones
Complete the #93527 fix: the tombstone now carries the human prompt
text via _entry_prompt_text (handling flat prompt, ShareGPT, and
chat-style shapes), _scan_completed_prompts_by_content counts
discarded rows as completed instead of only reading ShareGPT
conversations, and the merge step reports excluded tombstones in the
combined-count summary. Adds a dedicated regression suite covering the
tombstone round-trip, the all-discarded-batch resume path, and merge
exclusion.

Salvaged from #93579 (issue reporter's PR), building on #93542.
Fixes #93527
2026-08-24 03:21:07 -07:00
chelsealong 316d52faf2 fix(batch_runner): write a discard tombstone so resume skips no-reasoning prompts
The no-reasoning discard branch in _process_batch_worker continued
before writing any JSONL row, so run(resume=True) — which filters
solely via _scan_completed_prompts_by_content over batch_*.jsonl —
never saw discarded prompts and re-ran them at full cost on every
resume. Write a tombstone row on discard, exclude tombstones from the
trajectories.jsonl merge, and report discarded_no_reasoning in
final statistics.

Salvaged from #93542.
Fixes #93527
2026-08-24 03:21:07 -07:00
Teknium 4b622bbc4b test(model_metadata): lock in max_tokens last-resort fallback + cache self-heal
Adjust the #93423 max_tokens-only regression test to the merged policy:
max_tokens stays as an explicit LAST-RESORT fallback (some local servers
report nothing else) instead of being dropped entirely, and add coverage
that _reconcile_local_cached_context_length rewrites a cache entry
poisoned by the old probe (393216) upward to the real window (1048576)
once the probe is fixed.

Co-authored-by: pju-hoge <grkt@ppmz.com>
Co-authored-by: re-ITRT <1940428933@qq.com>
2026-08-24 03:20:57 -07:00
re-ITRT a0c802c02c fix(model_metadata): stop misreading max_tokens as context length in local probe
The local-endpoint context probe (_query_local_context_length_uncached)
treated max_tokens — an output-completion cap — as a candidate for the
model's context window. For OpenAI-compatible gateways that advertise a
1M context via context_size / max_input_tokens alongside a smaller
max_tokens output cap (e.g. TokenHub serving deepseek-v4-flash:
context_size=1048576, max_input_tokens=1048576, max_tokens=393216),
Hermes mis-detected the window as 393,216 and — because loopback
endpoints are reconciled against a live probe — actively overwrote a
previously-correct 1M cache entry.

- Add context_size and max_input_tokens to both /v1/models probe
  candidate lists (single-model detail and list branches).
- Remove max_tokens from the context-length candidates; it remains
  handled separately as an output cap (_MAX_COMPLETION_KEYS).

Adds regression tests covering context_size/max_input_tokens priority
over max_tokens and the max_tokens-only (no real context key) case.
2026-08-24 03:20:57 -07:00
Kyzcreig 4d729e4b31 fix(model-metadata): local ctx probe must not read max_tokens as the context window
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.

Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
2026-08-24 03:20:57 -07:00
carlotestor 394f0f0902 fix(context): prefer max_input_tokens over max_tokens for Anthropic proxies
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:

  max_input_tokens = context window (e.g. 1M for claude-fable-5)
  max_tokens       = max output     (e.g. 128k)

That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).

Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
2026-08-24 03:20:57 -07:00
Shanthan Subramaniam e210fd8c1f fix(gateway): resolve PairingStore's default pairing dir lazily, not at import time
PairingStore(profile=None) resolved its storage directory from the
module-level PAIRING_DIR constant, which was computed exactly once, at
module import time. A long-lived process (the gateway, started once at
container/process boot) can import this module before HERMES_HOME or a
profile's context is fully established, freezing PAIRING_DIR to a wrong
value for the rest of that process's lifetime -- even though a freshly
started, short-lived process (e.g. the `hermes pairing` CLI) re-imports
the module later with the environment already correct.

That asymmetry is exactly what made pending pairing codes issued by the
gateway process unrecoverable (the pending-code write landed under the
stale, wrong directory) while CLI-invoked writes to the same nominal
directory kept working -- see #93449 for the full writeup and a live
reproduction. tests/hermes_cli/test_dashboard_admin_endpoints.py already
carried a comment acknowledging this exact staleness in passing ("the
module-level PAIRING_DIR is bound at import"), and
TestProfileScopedStorage::test_default_store_uses_global_dir's own
comment describes working around it rather than it being intentional
behavior -- this fixes the underlying cause both were compensating for.

The profile-scoped branch already resolved its directory lazily inside
__init__ (matching this docstring's claim that resolution is lazy); this
brings the non-profile branch in line with it.

Fix keeps PAIRING_DIR as the same test seam already used throughout the
test suite (`patch("gateway.pairing.PAIRING_DIR", tmp_path)`, ~30 call
sites) unchanged: it's now a None sentinel instead of an eagerly computed
path, and a new _default_pairing_dir() helper resolves it fresh on every
call, honoring a patched (non-None) value when one is set. No existing
test needed to change.

Added a regression test that does not patch PAIRING_DIR directly and
instead exercises the real lazy-resolution path across two different
HERMES_HOME values in the same process -- confirmed it fails on the
pre-fix code (gets stuck with whatever the first PairingStore() call in
the test session happened to see) and passes with the fix.

Verified: tests/gateway/test_pairing.py (39, incl. the new one),
tests/hermes_cli/test_pairing.py, and tests/tools/test_pr_6656_regressions.py
all pass unmodified.
2026-08-24 03:20:46 -07:00
aniruddhaadak80 1007296ca0 fix(cli): add --reasoning to both top-level value-flag sets
--reasoning takes a value (metavar=LEVEL in _parser.py) but was absent
from _TOP_LEVEL_VALUE_FLAGS (used by _first_positional_argv) and from
_apply_profile_override's value_flags set. Every invocation like
"hermes --reasoning high chat ..." therefore misclassified "high" as the
first positional, and _plugin_cli_discovery_needed() forced full eager
plugin CLI discovery at argparse-setup time - the documented startup
cost paid on every use of the reasoning override.

Add --reasoning to both sets, and add a parser-derived parity regression
test so future drift between the hand-maintained sets and
build_top_level_parser() fails CI instead of silently degrading startup
(the exact drift class AGENTS.md bans).

Fixes #93530

(cherry picked from commit 9280617ab8b308f9a8cf947f1617276a6bc8eb4f)
2026-08-24 03:20:37 -07:00
fangliquanflq 694550e486 fix(cli): derive top-level value flags from parser
(cherry picked from commit f2e5a1388115615fafde49a5f2144e5855a81d89)
2026-08-24 03:20:37 -07:00
fangliquanflq 3963fc6f21 fix(config): stop reporting stripped v15 defaults
(cherry picked from commit 4c6b67ec371b16c15e9ffbb91bfb47a504e913fe)
2026-08-24 03:20:28 -07:00
Teknium a7aa814c42 fix(tools): widen the command-position anchor to the whole hardline class
#93392 was not just one pattern: every hardline rule with a bare \b anchor
fired on its token anywhere in the command line, including inside quoted
prose handed to echo, git commit -m, or gh --body. Anchor the
command-name-token rules and quote-mask the positionless ones:

- dd-to-block-device and kill -1 get the same _CMDPOS anchor as the
  format/rm/shutdown families, keeping their argument tails.
- redirect-to-block-device and the fork bomb have no command-name token to
  anchor (`>` appears mid-command; the bomb is a function definition), so
  they now match a quote-masked variant (_mask_quoted_prose) where quoted
  string content is blanked. $() and backtick spans inside double quotes
  stay raw (the shell executes them), and any command whose command-position
  words include a shell carrier (sh/bash/zsh/ksh/dash -c, eval, source, .)
  is scanned unmasked -- quoting is not a bypass. bash/sh -c payloads also
  still surface as raw detection variants via _execution_flag_findings.

Regression tests cover both directions for every touched pattern: quoted
prose passes, and every true-positive shape (bare, ; && | separators,
sudo/env prefix, $(), backticks, sh -c/bash -c/eval payloads) stays on the
unconditional floor.
2026-08-24 03:20:14 -07:00
liuhao1024 8163c8731b fix(tools): anchor the mkfs hardline pattern to command position
mkfs was the only HARDLINE_PATTERNS entry without a _CMDPOS anchor, so
the unconditional floor blocked any command that merely mentioned the
token inside quoted prose — `echo "does this workflow use mkfs
anywhere?"` was refused outright (#93392) instead of running the echo.

Anchor mkfs to command position like every sibling entry (rm root-
delete, shutdown family, dd): it matches at the start of a command,
after separators, or behind sudo/env/exec/nohup/setsid wrappers, and
no longer fires on argument-position mentions. The quote-aware
_mark_command_starts pass already keeps separators inside quoted
strings from looking like command starts, and \b still protects
mkfs_helper-style names.
2026-08-24 03:20:14 -07:00
aniruddhaadak80 7befc1d2dd fix(gateway): route platform authorization reads through the profile secret scope
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):

- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
  WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
  GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
  GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
  hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
  _sig_secret helper; empty scoped group list previously meant "drop all
  groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
  WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
  while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
  startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
  though its sibling dm/group reads already used the scoped _getenv

All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.

Fixes #93522
2026-08-24 03:20:06 -07:00
chelsealong d7e4204e77 fix(gateway): scope multiplex-profile authorization reads (weixin/yuanbao/wecom)
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.

Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.

Fixes #93522.
2026-08-24 03:20:06 -07:00
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
fyzanshaik 8345effc4e fix(cli): support reliable setup menu navigation
Decode Ghostty/Kitty enhanced selection and cancellation keys, make setup cancellation terminal, and add cross-terminal previous-step navigation to setup and model flows.

Refs #92833
2026-08-24 15:45:52 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Teknium d9a48f656a fix(desktop): scheduled jobs on sleeping profiles keep firing
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
2026-08-24 03:14:30 -07:00
Owenz-creator cdd37035d9 fix(desktop): probe half-open gateway socket on wake and reconnect
macOS sleep/wake (or a silent network drop) can leave the renderer's
WebSocket half-open: no close event fires, so connectionState stays
'open' while every RPC hangs until its per-call timeout. prompt.submit's
timeout is 30 minutes, so the user's next message reads as "enter does
nothing until I restart the app".

- Add a minimal ping RPC (tui_gateway/server.py) answered synchronously
  on the WS reader thread.
- On wake signals, reconnectNow now probes the open-looking socket with a
  5s-bounded ping and force-closes it on failure, letting the existing
  reconnect machinery (backoff, tile rebinding, session refresh) take
  over. A pre-ping backend answering -32601 is treated as healthy.
- Tests: half-open socket force-reconnects; healthy socket untouched;
  method-not-found backend untouched; backend ping envelope contract.
2026-08-24 03:14:19 -07:00
Teknium 29c5a12e04 fix(cron): warn loudly when the due-scan removes a consumed one-shot that already ran (#93524)
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
2026-08-24 03:13:30 -07:00
kshitijk4poor 45aa0dc33e fix(cli): widen prompt_toolkit fallback to catch any runtime failure
Widen the exception guard from OSError to Exception (re-raising
KeyboardInterrupt/EOFError first) so any prompt_toolkit runtime
failure degrades to input() — matching the established pattern in
masked_secret_prompt.  ValueError and RuntimeError can arise from
exotic stream wrappers or event-loop issues with the same root cause:
prompt_toolkit cannot attach stdin on the terminal.

Add test_line_input_falls_back_to_input_on_any_prompt_toolkit_failure
covering the ValueError case.
2026-08-24 15:42:58 +05:30
wo-o a88cbe10fb test(cli): add regression test for line_input OSError fallback
Cover the prompt_toolkit runtime-failure path added in the fix commit: a
tty-reporting stdin where prompt_toolkit raises OSError(22) (macOS kqueue
EINVAL on fd 0 under curl|bash installs) must degrade to input() instead
of aborting the setup wizard.
2026-08-24 15:42:58 +05:30
kshitijk4poor 9857bcba5c fix: widen secure_parent_dir to skip entire install tree
Replace hardcoded /opt/hermes check with dynamic install-tree detection
using Path(__file__).resolve().parent. This catches ALL install paths
(Docker /opt/hermes, apt /usr/local/lib/hermes-agent, git clone, custom)
instead of just the Docker image path. Also covers subdirectories of the
install tree, not just the top-level dir.

Add regression test test_install_tree_skipped to verify both the install
root and subdirectories are excluded from chmod.

Add contributor email mapping for bradmarshall987.

Follow-up to PR #93050 by @bradmarshall987.
2026-08-24 15:42:09 +05:30
Ayush Nangia b392591d7a test(conformance): crash/resume persistence cells — Phase 1 skeleton (#80921)
Deterministic, LLM-free conformance cells against the real SessionDB with
real SIGKILL mid-write, per the tracking issue's spot-probe method:

- cell 1: acknowledged-append durability + recovery determinism (adapted
  from the issue's 29.5K probe, scaled kill window, identical assertions)
- cell 2: consume-once under 8-process concurrent claim_handoff
- cell 3 (new): compression-rotation atomicity — never a compression-ended
  parent without a continuation (#80337 contract; #80487 recovery context)
- cells 4-5: documented stubs interlocked with #82956-#82959 and
  #83197/#83557

Journal-mode matrix (resolver default / DELETE / WAL-with-skip-gate) per
cell; every wait deadline-bounded; writers asserted alive at kill time.
2026-08-24 15:37:26 +05:30
BrunoBza 2a27e1ffb3 fix(gateway): prove --replace ownership from the bound pid record
Review v2 of #93084: the readable-cmdline path used substring matching
as destructive authority. That fails open on prefix collisions
(--profile timothy vs our profile tim) and same-name profiles under
different roots, so a poisoned record could still reach SIGTERM.

Ownership is now decided by the persisted identity record ALONE — exact
_same_hermes_home equality, bound to the live target by exact pid +
start-time. Missing/legacy/unbound/foreign records all refuse. A
readable argv feeds only a token-exact consistency check
(_looks_like_profile_conflict_from_cmdline via shlex tokens) that
refuses explicit contradictions like --profile timothy under tim; bare
or matching argv adds nothing.

Also fixes the review's source-of-record concern: the guard validates
the record that authorizes get_running_pid()'s answer rather than
assuming {HERMES_HOME}/gateway.pid is always the source.

Adds the requested signal-boundary regression: start_gateway(replace=
True) with unprovable ownership returns False without calling
terminate_pid or writing a takeover marker; the bound same-home
counterpart still reaches the replace flow. Legacy replace-flow tests
updated to stage a valid bound record for their legitimate-replace
fixtures.
2026-08-24 15:37:22 +05:30
kshitijk4poor 9975544101 style: ruff format test_codex_sdk_transform_bypass.py 2026-08-24 15:30:30 +05:30
kchernev 10a070bd49 fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.

Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 15:29:45 +05:30
kshitijk4poor dc50f02090 fix(adoption): re-check donor growth at retire time, not just at export
Closes the TOCTOU window flagged in review on #93369 (merged via
#93430): the divergence guard compared EXPORT-TIME message counts, but
another backend can append donor messages between the export snapshot
and the retire loop — that growth would be stamped behind the
non-recoverable adopted_by_profile archive, the exact H2 class the
guard exists to prevent, just via a narrower race.

The retire loop now re-reads live donor vs local counts immediately
before end_session and leaves the donor unretired (donor_retired=False,
warn-logged) on any donor-ahead signal; the next resume's export-time
guard then handles the divergence normally. Equal-count CONTENT
divergence (donor rewind+rewrite) remains invisible to count comparison
— documented as accepted: bytes stay in the donor store either way.

New red-first-verified regression simulates the exact race by appending
to the donor from inside an export_session_lineage wrapper.
adoption+ownership suites: 25 passed; ruff clean.
2026-08-24 15:07:18 +05:30
kshitijk4poor f93b350711 fix: align cache-policy pre-gate identity with the capability matcher
Follow-ups on top of the salvaged #92785 commit:

- Pre-gate now matches base URLs via normalize_route_base_url and
  provider ids via custom_provider_aliases, mirroring the semantics of
  get_custom_provider_model_capability. The raw string comparison
  silently dropped declarations whose config spelling differed only by
  host case or trailing slash (proven empirically: …/v1/ vs …/v1 with a
  non-matching provider name returned (False, False) despite an explicit
  prompt_caching: true).
- get_provider(..., allow_network=False) in the early-init/stub branch:
  the policy runs per request destination (MoA aggregator, auxiliary
  replans via blank_cache_policy_stub, early agent init) and a cold
  models.dev cache triggered a measured ~450 ms foreground registry
  fetch from the send path. A catalog miss degrades to the conservative
  side.
- Debug-log the previously silent provider-lookup exception fallback.
- Tests: _make_agent defaults _custom_providers=[] (post-init reality;
  keeps built-in-route tests off the catalog/config fallback), the two
  early-init tests delete the attr explicitly, and three regression
  tests pin the URL-drift, spaced-legacy-name, and no-network contracts
  (all three fail on the unfixed commit).
2026-08-24 14:58:42 +05:30
Blood Shot 0204e4898e fix(agent): normalize custom provider route identity 2026-08-24 14:58:42 +05:30
Blood Shot 0a3b7efec5 fix(agent): honor prompt_caching for custom providers
Apply explicit per-model prompt_caching capabilities to custom
chat-completions routes, rather than limiting them to recognized providers,
hosts, or model families.

Keep undeclared routes conservative, derive the marker layout from the wire
transport, and leave Responses and Bedrock caching paths unchanged.
2026-08-24 14:58:42 +05:30
leosiedler a0ca7c1920 feat(cron): add explicit one-shot re-arm 2026-08-24 00:34:26 -07:00
Ben Barclay 9a91f7058c Merge remote-tracking branch 'origin/main' into fix/pkce-samesite-none-salvage 2026-08-24 16:48:06 +10:00
Teknium ed8ee9a871 fix(cron): misfire backstop honors the one-shot grace window (#93526)
The hosted-provider misfire catch-up (fire_overdue_jobs) fired any runnable
overdue job with no one-shot grace check, so a stored past-due one-shot
bypassed ONESHOT_GRACE_SECONDS and executed arbitrarily late after downtime.
Sibling site of the due-scan gate from #89571; pins both directions with
tests.
2026-08-23 23:33:29 -07:00
spfcraze b37a5bc0df fix(cron): due-scan must not dispatch a one-shot past its grace window
create_job / update_job / resume_job all reject a one-shot whose run time is
more than ONESHOT_GRACE_SECONDS in the past ("will never fire"), and
_recoverable_oneshot_run_at never recovers such a schedule — but
_get_due_jobs_locked dispatched ANY one-shot whose *persisted* next_run_at was
in the past, even hours later (gateway down past the window, host asleep,
hand-edited jobs.json). A wall-clock one-shot then ran hours late, violating
the "will never fire" contract enforced everywhere else.

- Grace gate: a once-kind job whose next_run_dt is more than
  ONESHOT_GRACE_SECONDS in the past is never appended to the due list.
- If no run_claim/fire_claim exists (nothing was ever dispatched), retire the
  record with a diagnostic file so it stops being scanned and the miss is
  operator-visible.
- If a (possibly stale) claim exists, a run may still be in flight in another
  process: skip this scan but KEEP the record so its mark_job_run can land
  (avoids re-introducing mid-flight record deletion).
- Manual re-trigger still works: trigger_job sets next_run_at=now (inside
  grace) so an explicitly re-run stale one-shot fires.

Tests (tests/cron/test_oneshot_grace_due_scan.py): stale-not-due+retired,
within-grace-still-due, stale+claim-skipped-but-kept, retriggered-is-due, and
recurring-jobs-unaffected.
2026-08-23 23:33:29 -07:00
Gille 21b92d2687 fix(agent): bypass response cache for empty retries 2026-08-23 23:31:54 -07:00
Teknium ec44116d59 test(tools): shared sanitizer contract + singularity overlay coverage
Behavior-contract tests for sanitize_task_id_for_path (colon/separator
removal, verbatim pass-through for existing safe ids, determinism,
collision-freedom incl. the a:b vs a_b digest case, traversal and
oversized-id bounds) and for the singularity persistent overlay path
(sanitized, verbatim for safe ids, distinct dirs for colon-vs-underscore
ids).

Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
2026-08-23 21:12:32 -07:00