Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium 4d030a37c7 fix(desktop): stack subsequent preview tiles as tabs instead of new right splits
Every opened file registered its preview pane with dock dir 'right', so
each open split a new zone off the right edge — three file opens made
three ever-narrower columns (#93610). The first preview still opens its
own zone docked beside main; every subsequent preview now anchors to an
existing preview-tile pane with dir 'center', so it stacks as a tab in
the same preview zone. Covers files, artifacts, and the Browser tab
alike (all flow through openPreview/$previewTabs); session tiles are
untouched.

Fixes #93610
2026-08-24 03:23:42 -07:00
Teknium 57ece81101 fix(desktop): exclude cron sessions from the titlebar unread badge
Cron runs finish unwatched by design, so counting them in
$unreadSessionCount turned the titlebar badge into a permanently-lit
cron run counter (#93552). The badge now counts regular + messaging
sessions only; cron unread state stays visible on the sidebar cron
section rows, and 'Mark all as read' (markAllSessionsRead +
ackAllSessionsRead, which iterates cron rows) still clears them.

Fixes #93552
2026-08-24 03:23:33 -07:00
Teknium 06fc941d23 fix(desktop): stop bot-relay drain loop from redialing a WebSocket per connection per tick
The bot relay's drain loop RPCs every registered connection through
requestGatewayForAgent's per-request lease. With no other consumer
holding the route, the refcount hit 0 after every tick and the pooled
secondary was disposed — a fresh WebSocket dial + teardown per
connection every 4s, flooding the gateway logs with connect/disconnect
pairs (#93594).

Two changes, both directions from the issue:

- Retained relay-route secondaries: retainGatewayForRelay pins a
  route's pooled socket with a counted retention (never clobbering the
  foreground 'retained' flag) for the relay's active lifetime, reusing
  the existing scheduleReconnect/full-jitter machinery on drops. The
  plugin pins each registered connection once via the new feature-
  detected host.retainProfileSocket door, reconciles pins with the
  current connection set on every drain, and releases everything in
  stopBotRelay/dispose. Local routes (null/'local') are exempt so the
  idle reaper can still reclaim spawned local backends. The live-work
  pruner also respects the pin.

- RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (#93091,
  bot_relay.outbox.pending) carries envelope latency, so the poll is
  purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS.

Tests: relay-push-drain updated to the new backstop semantics; new
gateway-relay-retention.test.ts proves one socket construction across
5 drain ticks (vs 3 constructions for 3 unretained ticks) and that
release/prune/local-exemption behave; new relay-socket-retention
plugin test pins the pin-once / release-on-departure / stop-releases
contracts.
2026-08-24 03:23:15 -07:00
Teknium c9e2a46df6 fix(pricing): support Gemini context-tiered rates in pricing snapshot (#93469)
The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).

- Add optional tier fields to PricingEntry: tier_threshold_tokens,
  input/output/cache_read_cost_per_million_above (None = flat, falls
  back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
  request once usage.prompt_tokens (input + cache read + cache write)
  exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
  gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
  above 200k).
- Flat entries are untouched: no threshold means no behavior change.

Reported and tier-field shape designed by @tornike14 (#93469).

Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
2026-08-24 03:23:07 -07:00
Teknium a87d314e44 fix(auth): malformed OpenRouter env key no longer shadows valid credential-pool key
A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).

- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
  that mismatch a declared prefix, logging a WARNING naming the env var and
  expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
  only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
  are never rejected. Valid env keys still win over the pool (precedence
  unchanged).

Fixes #93593
2026-08-24 03:22:57 -07:00
Teknium d8d1e18ab9 test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
2026-08-24 03:22:48 -07:00
chelsealong b3730153c3 fix(terminal): cap watch_patterns notifications over a process's lifetime
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
2026-08-24 03:22:48 -07:00
Teknium c7f1f4d6e9 fix(desktop): teach the torn-bundle guard to see missing lazy chunks
missingRendererAssets only checked the module refs index.html itself names
(<script type=module> + modulepreload), so a torn install whose boot-critical
files were intact but whose lazy chunks were gone passed the generation check
and died minutes later on the first React.lazy() route with 'Failed to fetch
dynamically imported module' (#93479: syntax-diff-*, shiki-*, mermaid-embed-*).

Walk the generation's module graph: for every present JS chunk, parse its
inline __vite__mapDeps filename table (the lazy-import manifest Vite bakes
into each chunk) and check those files too, transitively and cycle-safe.
resolveRendererIndex now skips a lazy-chunk-torn candidate in favor of the
intact copy instead of shipping a delayed crash.

Tests cover the mapDeps parser (definition table vs index-only call sites,
CDN refs), the exact #93479 tear shape, transitive/cyclic walks, and the
torn-vs-intact preference end to end.
2026-08-24 03:22:38 -07:00
Teknium 73af57a2ae fix(desktop): prefer the unpacked web dist over the asar-internal renderer index when packaged
The renderer index resolver tried APP_ROOT/dist/index.html — inside app.asar
when packaged — before the app.asar.unpacked copy that asarUnpack (dist/**)
ships and that resolveWebDist() already prefers for the embedded dashboard.
Loading the asar-internal index is how lazily imported chunks (syntax-diff-*,
shiki-*, mermaid-embed-*) end up fetched from a path that cannot serve them,
killing the workspace pane (#93479).

Reorder the candidate ladder to prefer the unpacked web dist when packaged,
following the unpackedPathFor/resolveWebDist precedent. All window loaders
(main, overlay, quick) share resolveRendererIndex, so one reorder covers
every surface. Dev behavior is unchanged: outside an asar both candidates
collapse to APP_ROOT/dist and keep the original order.
2026-08-24 03:22:38 -07:00
chelsealong b614e2656b fix(desktop): isolate a failed lazy syntax-diff import from the workspace pane
React.lazy(() => import('./syntax-diff')) only has its pending state
covered by Suspense. When the dynamic import rejects (e.g. a packaged
app whose renderer window resolves to the app.asar copy of dist/ while
the chunk exists only in app.asar.unpacked, #93479), the rejection
throws past Suspense to the nearest error boundary, which is the whole
workspace ContribBoundary. One missing highlighter chunk then blanks
the entire chat transcript instead of just the diff falling back to
the plain colored DiffBody, the way markdown-text.tsx already isolates
this failure class for markdown.

Wraps LazySyntaxDiff in a local ErrorBoundary that renders DiffBody on
catch, so a failed highlight chunk degrades in place.
2026-08-24 03:22:38 -07:00
Teknium 40ab950ae3 fix(desktop): guard render-reachable route lookups against orphaned rows
Audit of the remaining unguarded botConnectionRoute() callers a pane
render can reach (#93492 follow-up to the botRosterMeta split). Each now
uses the non-throwing resolveBotConnectionRoute() and degrades on an
owner_removed row instead of throwing into the pane's error boundary:

- botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility
  listener, Bots home open, and roster context menus recompute these on
  passive UI edges; an orphaned selection now yields the name-keyed owner
  and the blocked workspace target.
- durableGroupChatMembers: rebuilt on every group send over the whole
  seated roster; one orphaned member no longer aborts the room update,
  and a swept member's degraded mark now survives the rebuild.
- useModelOptions: hook body runs during render; the query is disabled
  for an orphaned row and the picker paints its error/disabled state.
- AdvancedProfileConfig: dialog falls back to the bot's own name scope.

Strict dispatch callers (requestForBot, session creation, deleteBot,
duplicateBot, ensureBotMetadata, routines) intentionally keep the
fail-closed throw — remote-routing-races.test.mjs still asserts it.

Adds orphaned-connection-members.test.mjs covering the removed-connection
sweep, the hydrate annotate (with/without a readable registry), the
degraded 'Gateway removed' rendering of swept rows, and every guarded
caller.
2026-08-24 03:22:28 -07:00
Teknium 725cfe2909 fix(desktop): annotate group-chat members already orphaned before hydrate
Rows poisoned before the removed-connection sweep existed (their
connection was deleted while an older Desktop ran, so no lifecycle push
ever swept them) are what made #93492 survive app restarts. After the
persisted 'group-chats' hydrate, run a pure annotate pass over the rooms:

- a descriptor that lost its connectionId (route unresolvable — the exact
  shape that threw on render) is always marked;
- a descriptor whose connectionId is absent from the live connection
  registry is marked only when the registry could actually be read —
  an unavailable registry must not read as 'everything is orphaned'.

Marked rows keep their identity and degrade to the existing 'Gateway
removed' state; nothing is deleted.
2026-08-24 03:22:28 -07:00
Teknium c509af689f fix(desktop): sweep group-chat rosters when a connection is removed
Root cause of #93492: deleting a cloud/remote connection disposed its
gateways (store/gateway.ts) but never touched the persisted 'group-chats'
storage, so every member descriptor referencing the deleted connection
stayed behind as a poisoned row (remoteSource: true, connection gone) that
render-path route lookups tripped over forever.

Subscribe to the connection registry's 'removed' lifecycle push
(window.hermesDesktop.connections.onChanged, feature-detected — older
Electron mains don't emit it) and annotate every persisted group-chat
member owned by the deleted connection. Rows are marked
(sourceMissing/sourceReachable), never silently deleted: the member keeps
its identity and panes render the existing degraded 'Gateway removed'
botSourceStatus state. Writes ride updateGroupChat so the durable record
keeps its full shape, and the listener unbinds on plugin dispose.
2026-08-24 03:22:28 -07:00
chelsealong ec013b76db fix(desktop): split strict connection routing from passive roster lookup
botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.

Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.
2026-08-24 03:22:28 -07:00
chelsealong 09529afdd2 fix(desktop): group chats no longer crash when a member's connection is deleted
botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).

botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.

Fixes #93492
2026-08-24 03:22:28 -07:00
Teknium e8d5660bae fix(desktop): bound the boot and gateway-switch resolveGatewayWsUrl awaits too (#93454)
Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await
exists on the soft gateway-switch path and the initial boot() path. Bound
both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged
IPC round-trip fails into the existing retry paths instead of hanging the
switch or the 'Starting Hermes…' screen forever.
2026-08-24 03:22:03 -07:00
chelsealong 5ef205d855 fix(desktop): bound the revalidateConnection() await too (#93454)
attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger #93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.

Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.
2026-08-24 03:22:03 -07:00
chelsealong 17f8e24c81 fix(desktop): bound reconnect awaits so a stuck IPC round-trip can't latch the UI frozen
After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.

Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.

Fixes #93454
2026-08-24 03:22:03 -07:00
chelsealong 9c013eaaf8 fix(dashboard): follow scroll on implicit active-session resume (#93518)
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes #93518.
2026-08-24 03:21:49 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 85cd576b06 test(bot-relay): match delivery CLI by basename in argv filters
CI runners have a real hermes sibling next to the venv python, so
local_delivery_command now resolves an absolute path there — the exact
argv filters in the retry-policy fakes and the relay-methods pins must
match by basename instead of the literal "hermes", mirroring the
_delivery_lock matcher.
2026-08-24 03:21:37 -07:00
liuhao1024 c099ef05de fix(bot-relay): Windows path SyntaxError in waiter + PATH-less delivery ENOENT
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):

1. waiter_command embeds the reply path in generated python -c source
   with !r. repr escapes each backslash, but the Windows execution layer
   folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
   and SyntaxErrors the whole waiter script. Raw-string literals keep
   the folded single backslash a literal; POSIX paths have no
   backslashes so the prefix is a no-op there, and \' inside a raw
   literal still cannot terminate the string, keeping the #93091
   injection defense intact.

2. local_delivery_command hardcoded "hermes", relying on PATH — absent
   in service contexts (systemd units, desktop launchers, non-login SSH
   shells), so delivery died with ENOENT. It now resolves the CLI next
   to this gateway's own interpreter (venv bin/Scripts sibling,
   hermes.exe on Windows) with a bare-name fallback. The #93091
   per-profile turn-lock recognition in bot_mode_dm now matches the CLI
   element by basename (split on both separators) so resolved absolute
   paths still take the lock instead of silently bypassing it.

Fixes #93590
2026-08-24 03:21:37 -07:00
liuhao1024 8d17060249 fix(cli): hard-exit the Windows update hand-off child once work is durable
The re-exec'd venv child spawned by
_reexec_dependency_sync_off_windows_shim completes every update step —
the receipt records success / "completed at command boundary" — but then
hangs in interpreter shutdown on a leftover non-daemon thread, freezing
the PowerShell window for minutes after "Update complete!". On the
hand-off path only (HERMES_UPDATE_REEXEC=1), after the receipt is
finalized, the update lock released, and stdio restored, flush and
os._exit(code) instead of unwinding — the same treatment #79040's cron
workaround applies. SystemExit codes (including early refusals)
propagate to the hard exit; real exceptions keep the normal raise path
so tracebacks still print. Non-hand-off invocations are untouched: the
marker env is set solely when the shim spawns the child.

Fixes #93581
2026-08-24 03:21:28 -07:00
Teknium 93bf6f7225 fix(cli): fail closed on empty fleet probe across all pre-update liveness signals (#93406)
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.

Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
  same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries

Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.

Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.

Builds on RelaxJonh's #93410. Fixes #93406
2026-08-24 03:21:18 -07:00
RelaxJonh d74bbb9bd9 fix(cli): treat empty fleet probe as incomplete when gateways were restarted (#93406)
collect_fleet_versions() swallows every probe exception via
logger.debug() and returns whatever accumulated — which can be an
empty list.  print_fleet_version_matrix([]) returns False (no rows
to report), so the update exits 0 with "success" even though no
gateway was actually verified.

After the restart phase touches live gateways (restarted_services or
killed_pids is truthy), an empty fleet snapshot means verification
failed, not that everything is healthy.  Treat it as incomplete so
the receipt records "partial" and the exit code is 1.

Fixes #93406
2026-08-24 03:21:18 -07:00
aniruddhaadak80 a26154aceb fix(batch_runner): teach the resume content scan to honor discard tombstones
Complete the #93527 fix: the tombstone now carries the human prompt
text via _entry_prompt_text (handling flat prompt, ShareGPT, and
chat-style shapes), _scan_completed_prompts_by_content counts
discarded rows as completed instead of only reading ShareGPT
conversations, and the merge step reports excluded tombstones in the
combined-count summary. Adds a dedicated regression suite covering the
tombstone round-trip, the all-discarded-batch resume path, and merge
exclusion.

Salvaged from #93579 (issue reporter's PR), building on #93542.
Fixes #93527
2026-08-24 03:21:07 -07:00
chelsealong 316d52faf2 fix(batch_runner): write a discard tombstone so resume skips no-reasoning prompts
The no-reasoning discard branch in _process_batch_worker continued
before writing any JSONL row, so run(resume=True) — which filters
solely via _scan_completed_prompts_by_content over batch_*.jsonl —
never saw discarded prompts and re-ran them at full cost on every
resume. Write a tombstone row on discard, exclude tombstones from the
trajectories.jsonl merge, and report discarded_no_reasoning in
final statistics.

Salvaged from #93542.
Fixes #93527
2026-08-24 03:21:07 -07:00
Teknium 4b622bbc4b test(model_metadata): lock in max_tokens last-resort fallback + cache self-heal
Adjust the #93423 max_tokens-only regression test to the merged policy:
max_tokens stays as an explicit LAST-RESORT fallback (some local servers
report nothing else) instead of being dropped entirely, and add coverage
that _reconcile_local_cached_context_length rewrites a cache entry
poisoned by the old probe (393216) upward to the real window (1048576)
once the probe is fixed.

Co-authored-by: pju-hoge <grkt@ppmz.com>
Co-authored-by: re-ITRT <1940428933@qq.com>
2026-08-24 03:20:57 -07:00
re-ITRT a0c802c02c fix(model_metadata): stop misreading max_tokens as context length in local probe
The local-endpoint context probe (_query_local_context_length_uncached)
treated max_tokens — an output-completion cap — as a candidate for the
model's context window. For OpenAI-compatible gateways that advertise a
1M context via context_size / max_input_tokens alongside a smaller
max_tokens output cap (e.g. TokenHub serving deepseek-v4-flash:
context_size=1048576, max_input_tokens=1048576, max_tokens=393216),
Hermes mis-detected the window as 393,216 and — because loopback
endpoints are reconciled against a live probe — actively overwrote a
previously-correct 1M cache entry.

- Add context_size and max_input_tokens to both /v1/models probe
  candidate lists (single-model detail and list branches).
- Remove max_tokens from the context-length candidates; it remains
  handled separately as an output cap (_MAX_COMPLETION_KEYS).

Adds regression tests covering context_size/max_input_tokens priority
over max_tokens and the max_tokens-only (no real context key) case.
2026-08-24 03:20:57 -07:00
Kyzcreig 4d729e4b31 fix(model-metadata): local ctx probe must not read max_tokens as the context window
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.

Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
2026-08-24 03:20:57 -07:00
carlotestor 394f0f0902 fix(context): prefer max_input_tokens over max_tokens for Anthropic proxies
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:

  max_input_tokens = context window (e.g. 1M for claude-fable-5)
  max_tokens       = max output     (e.g. 128k)

That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).

Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
2026-08-24 03:20:57 -07:00
Teknium f28718c331 chore: map contributor email for shanthans-es 2026-08-24 03:20:46 -07:00
Shanthan Subramaniam e210fd8c1f fix(gateway): resolve PairingStore's default pairing dir lazily, not at import time
PairingStore(profile=None) resolved its storage directory from the
module-level PAIRING_DIR constant, which was computed exactly once, at
module import time. A long-lived process (the gateway, started once at
container/process boot) can import this module before HERMES_HOME or a
profile's context is fully established, freezing PAIRING_DIR to a wrong
value for the rest of that process's lifetime -- even though a freshly
started, short-lived process (e.g. the `hermes pairing` CLI) re-imports
the module later with the environment already correct.

That asymmetry is exactly what made pending pairing codes issued by the
gateway process unrecoverable (the pending-code write landed under the
stale, wrong directory) while CLI-invoked writes to the same nominal
directory kept working -- see #93449 for the full writeup and a live
reproduction. tests/hermes_cli/test_dashboard_admin_endpoints.py already
carried a comment acknowledging this exact staleness in passing ("the
module-level PAIRING_DIR is bound at import"), and
TestProfileScopedStorage::test_default_store_uses_global_dir's own
comment describes working around it rather than it being intentional
behavior -- this fixes the underlying cause both were compensating for.

The profile-scoped branch already resolved its directory lazily inside
__init__ (matching this docstring's claim that resolution is lazy); this
brings the non-profile branch in line with it.

Fix keeps PAIRING_DIR as the same test seam already used throughout the
test suite (`patch("gateway.pairing.PAIRING_DIR", tmp_path)`, ~30 call
sites) unchanged: it's now a None sentinel instead of an eagerly computed
path, and a new _default_pairing_dir() helper resolves it fresh on every
call, honoring a patched (non-None) value when one is set. No existing
test needed to change.

Added a regression test that does not patch PAIRING_DIR directly and
instead exercises the real lazy-resolution path across two different
HERMES_HOME values in the same process -- confirmed it fails on the
pre-fix code (gets stuck with whatever the first PairingStore() call in
the test session happened to see) and passes with the fix.

Verified: tests/gateway/test_pairing.py (39, incl. the new one),
tests/hermes_cli/test_pairing.py, and tests/tools/test_pr_6656_regressions.py
all pass unmodified.
2026-08-24 03:20:46 -07:00
Teknium 905edf37a4 refactor(cli): add static fallback to parser-derived value-flag helper
Mirror the full update_cmd._holder_value_flags precedent: derive both
top-level value-flag sets from build_top_level_parser() with a cached
frozenset, and fall back to a handwritten snapshot if parser
introspection ever fails, so argv classification keeps working on a
broken tree. Parity test pins the derived sets against the live parser
so drift fails CI.

Builds on #93551 (fangliquanflq) and #93570 (aniruddhaadak80) for #93530.
2026-08-24 03:20:37 -07:00
aniruddhaadak80 1007296ca0 fix(cli): add --reasoning to both top-level value-flag sets
--reasoning takes a value (metavar=LEVEL in _parser.py) but was absent
from _TOP_LEVEL_VALUE_FLAGS (used by _first_positional_argv) and from
_apply_profile_override's value_flags set. Every invocation like
"hermes --reasoning high chat ..." therefore misclassified "high" as the
first positional, and _plugin_cli_discovery_needed() forced full eager
plugin CLI discovery at argparse-setup time - the documented startup
cost paid on every use of the reasoning override.

Add --reasoning to both sets, and add a parser-derived parity regression
test so future drift between the hand-maintained sets and
build_top_level_parser() fails CI instead of silently degrading startup
(the exact drift class AGENTS.md bans).

Fixes #93530

(cherry picked from commit 9280617ab8b308f9a8cf947f1617276a6bc8eb4f)
2026-08-24 03:20:37 -07:00
fangliquanflq 694550e486 fix(cli): derive top-level value flags from parser
(cherry picked from commit f2e5a1388115615fafde49a5f2144e5855a81d89)
2026-08-24 03:20:37 -07:00
fangliquanflq 3963fc6f21 fix(config): stop reporting stripped v15 defaults
(cherry picked from commit 4c6b67ec371b16c15e9ffbb91bfb47a504e913fe)
2026-08-24 03:20:28 -07:00
Teknium a7aa814c42 fix(tools): widen the command-position anchor to the whole hardline class
#93392 was not just one pattern: every hardline rule with a bare \b anchor
fired on its token anywhere in the command line, including inside quoted
prose handed to echo, git commit -m, or gh --body. Anchor the
command-name-token rules and quote-mask the positionless ones:

- dd-to-block-device and kill -1 get the same _CMDPOS anchor as the
  format/rm/shutdown families, keeping their argument tails.
- redirect-to-block-device and the fork bomb have no command-name token to
  anchor (`>` appears mid-command; the bomb is a function definition), so
  they now match a quote-masked variant (_mask_quoted_prose) where quoted
  string content is blanked. $() and backtick spans inside double quotes
  stay raw (the shell executes them), and any command whose command-position
  words include a shell carrier (sh/bash/zsh/ksh/dash -c, eval, source, .)
  is scanned unmasked -- quoting is not a bypass. bash/sh -c payloads also
  still surface as raw detection variants via _execution_flag_findings.

Regression tests cover both directions for every touched pattern: quoted
prose passes, and every true-positive shape (bare, ; && | separators,
sudo/env prefix, $(), backticks, sh -c/bash -c/eval payloads) stays on the
unconditional floor.
2026-08-24 03:20:14 -07:00
liuhao1024 8163c8731b fix(tools): anchor the mkfs hardline pattern to command position
mkfs was the only HARDLINE_PATTERNS entry without a _CMDPOS anchor, so
the unconditional floor blocked any command that merely mentioned the
token inside quoted prose — `echo "does this workflow use mkfs
anywhere?"` was refused outright (#93392) instead of running the echo.

Anchor mkfs to command position like every sibling entry (rm root-
delete, shutdown family, dd): it matches at the start of a command,
after separators, or behind sudo/env/exec/nohup/setsid wrappers, and
no longer fires on argument-position mentions. The quote-aware
_mark_command_starts pass already keeps separators inside quoted
strings from looking like command starts, and \b still protects
mkfs_helper-style names.
2026-08-24 03:20:14 -07:00
aniruddhaadak80 7befc1d2dd fix(gateway): route platform authorization reads through the profile secret scope
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):

- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
  WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
  GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
  GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
  hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
  _sig_secret helper; empty scoped group list previously meant "drop all
  groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
  WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
  while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
  startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
  though its sibling dm/group reads already used the scoped _getenv

All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.

Fixes #93522
2026-08-24 03:20:06 -07:00
chelsealong d7e4204e77 fix(gateway): scope multiplex-profile authorization reads (weixin/yuanbao/wecom)
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.

Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.

Fixes #93522.
2026-08-24 03:20:06 -07:00
Teknium 8eba0d2fd6 chore: map dougatbuck contributor email 2026-08-24 03:19:47 -07:00
dougatbuck 855e191d89 test(desktop): document connection-scoped filesystem keys 2026-08-24 03:19:47 -07:00
dougatbuck 9abeb89a49 fix(desktop): guard stale files refreshes 2026-08-24 03:19:47 -07:00
dougatbuck 30150fd082 fix(desktop): isolate files across connections 2026-08-24 03:19:47 -07:00
kshitijk4poor 0eda2ba0c8 fix: remove dead code, deduplicate error constants, fix skill key check
Follow-up to PR #92189 salvage:
- Remove unused job_no_agent_without_script() function (dead code)
- Replace inline NO_AGENT_WITHOUT_SCRIPT_ERROR string in _validate_job_mode_invariants with the constant
- Replace scheduler inline reason string with EMPTY_PAYLOAD_ERROR constant
- Add 'skill' (singular) to job_payload_is_empty 'in job' presence check
2026-08-24 15:47:42 +05:30
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
kshitijk4poor b57530afee refactor: collapse context_back/effective_allow_back, fix draw_header docstring
Follow-up cleanup for salvaged PR #92838:
- Collapse redundant context_back and effective_allow_back into a single
  allow_back variable (they were always identical, never reassigned)
- Update _run_curses_menu docstring to reflect draw_header's actual
  signature including search and back_enabled kwargs
2026-08-24 15:45:52 +05:30
fyzanshaik 8345effc4e fix(cli): support reliable setup menu navigation
Decode Ghostty/Kitty enhanced selection and cancellation keys, make setup cancellation terminal, and add cross-terminal previous-step navigation to setup and model flows.

Refs #92833
2026-08-24 15:45:52 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00