Commit Graph

14240 Commits

Author SHA1 Message Date
HexLab98 eef7a10756 test(docker): cover session-key sandbox paths and their collision boundary
Drives the real DockerEnvironment constructor with a Telegram DM session key
and asserts every persistent -v spec is a two-field bind whose source holds no
colon — the assertion that reproduces exit 125 on the unfixed path.

The derivation's own contract is covered separately: ids that already work stay
verbatim (no sandbox migration), docker's separator and the path separators
never survive, ids differing only in rewritten characters keep distinct
directories, the mapping is stable across calls so cross-process container
reuse still resolves, pathological keys stay inside the per-component length
limit, and "."/".."/empty cannot resolve to the docker sandbox root.
2026-08-23 21:12:32 -07:00
Teknium 03c3554fc2 fixup(curator): align #93002 test stubs with #93149 set_pinned bool contract
Combining both PRs for issue #92993: #93149 makes set_pinned() return a
bool and _cmd_pin/_cmd_unpin exit 1 on a no-op write; #93002's tests
stubbed set_pinned with a None-returning lambda, which the combined
_cmd_pin now reads as failure. The stub reports True (write landed) so
#93002's messaging assertions exercise the intended success path.
2026-08-23 21:12:20 -07:00
liuhao1024 ef882a5595 fix(curator): say what pin actually does on an unmanaged skill
`hermes curator pin` guarded on is_agent_created (a filesystem-shape
check), but the flag only matters when the skill carries the
curator-management marker: curated_report() walks marker-carrying skills
only, so auto-transitions never consider an unmanaged (pre-marker)
skill at all. Pinning one recorded the flag and then printed
"will bypass auto-transitions" — an effect that does not exist.

Keep the write (the flag becomes meaningful after `hermes curator
adopt`) and branch the message on is_curator_managed: unmanaged pins
now say the skill is unmanaged and point at adopt. Unpin gets the
symmetric wording.
2026-08-23 21:12:20 -07:00
beplee dd20c30dec fix(curator): check unpin result, guard status ghost rows, tighten test
Review feedback on #93149:
- _cmd_unpin now checks set_pinned's return (same false-success defect
  existed symmetrically on the unpin path)
- curated_report() pinned-visibility branch requires a local skill dir,
  so stale records for deleted dirs don't render as ghost rows
- test 2 asserts rc==0 unconditionally instead of vacuous-passing
- error message points to list-unmanaged (status doesn't render reasons)
2026-08-23 21:12:20 -07:00
beplee 7caa731e80 fix(curator): report pin failures instead of false success and surface pinned unmanaged skills
`hermes curator pin <skill>` printed success even when the underlying
write never landed. set_pinned() routes through _mutate() with
require_curation_eligible=True, which silently returns None for skills
that pass is_agent_created() but fail is_curation_eligible() — e.g. a
user-created skill named "plan", which PROTECTED_BUILTIN_SKILLS blocks
by name. The CLI then announced a pin that does not exist (#92993).

Also, a pin that DID land on an eligible-but-unmanaged skill (no
created_by marker) was invisible: curated_report() only iterated
list_agent_created_skill_names(), which requires the management marker,
so the skill showed up under 'unmanaged' with no trace of its pin.

- set_pinned() now returns bool write success; _cmd_pin() checks it,
  exits nonzero and explains the refusal when the write did not land
- curated_report() additionally includes curation-eligible skills whose
  usage record carries pinned=true, so their pins are visible in status

Fixes #92993
2026-08-23 21:12:20 -07:00
Teknium c584d15cdc feat(bots): typed failure reasons reach the sending agent on A2A calls (#93091)
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:

- Desktop relay drain forwards bot_relay.deliver's error.data.reason
  into bot_relay.reply (and prefers it for the attention badge over
  free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
  text, so the completion notification the sending agent receives is
  machine-branchable.

Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
2026-08-23 20:07:21 -07:00
Teknium 0c3a507535 fix(classifier): 429 quota walls route to billing across providers; reset signals stay rate-limited
Consolidates the 429-quota-classifier cluster on top of the merged #93419
Anthropic core. Three independent contributor findings salvaged into one
coherent change to the single 429 branch:

- Broaden the 429 usage-limit check from the narrow 'usage limit' string to
  the full _USAGE_LIMIT_PATTERNS ('quota', 'limit exceeded', 'key limit
  exceeded') and add _BILLING_PATTERNS detection on 429 ('insufficient
  credits' wrapped in a 429 instead of 402), guarded by a _RATE_LIMIT_PATTERNS
  exclusion so an explicit 'Rate limit exceeded' never promotes to
  non-retryable billing. (credit @Pluviobyte, #39441 — earliest submitter)
- Add 'resets in' to the transient signals: Codex's 'Weekly usage limit
  reached. Resets in 6hr 29min.' wrongly read as terminal billing because
  main only had 'reset in' (no substring match). (credit @LeonSGP43, #63021)
- Add 'reset after' / 'available in' / 'per minute' / 'per second' transient
  signals. (credit @jtstothard, #74785)

Supersedes #65633 (defective branch placement, no tests). The aux-client
path already covers these shapes (_is_payment_error catches weekly/quota
walls; _is_rate_limit_error treats 'resets in' as transient), so no change
there.

Tests: 6 new cases (generic quota wall, insufficient-credits 429, rate-limit
guard, Codex resets-in, extra transient phrases). Guard sabotage-verified.

Co-authored-by: Pluviobyte <Pluviobyte@users.noreply.github.com>
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
Co-authored-by: jtstothard <jtstothard@users.noreply.github.com>
2026-08-23 20:02:07 -07:00
fangliquanflq fe24525605 fix(state): reap only proven database holders 2026-08-23 20:01:41 -07:00
fangliquanflq 8f3a82f96a fix(state): recover FTS after orphan holder deferrals 2026-08-23 20:01:41 -07:00
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
fangliquanflq 030edf9774 fix(auth): canonicalize configured provider display names 2026-08-23 20:00:53 -07:00
fangliquanflq 3a7c094582 fix(auth): preserve configured provider compatibility 2026-08-23 20:00:53 -07:00
fangliquanflq c527b2c0a4 fix(auth): normalize configured provider pool keys 2026-08-23 20:00:53 -07:00
Adolanium 2912c36aa4 fix(gateway): stop multiplex allowlist leak and bot-relay python -c injection
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.

bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
2026-08-23 20:00:30 -07:00
Teknium 57649294be test(bots): turn-lock fake Proc gains stdout/stderr attrs
_run_delivery now captures output to drive the retry policy; the
turn-lock test's minimal _P fake predates that contract. Sibling-test
blast radius fix, no behavior change.
2026-08-23 20:00:18 -07:00
Teknium b274b346d8 feat(bots): retry session policy — resume transient turns, compress-and-resume on context overflow (#93091 item 5)
Maintainer ruling (2026-08-23): a retried bot turn never mints a fresh
session. retry_action() maps the #93091 item-1 reason enum to one of
resume / compress_then_resume / none:

- transient classes (runtime_offline, delivery_timeout, rate limit,
  server error) re-run the same Bot Chat session once;
- context_overflow also re-runs the same session — the retried turn
goes through the pre-API compaction pass in conversation_loop.py,
  which compacts the over-threshold transcript first (the one
  sanctioned context mutation); no fresh-session escape hatch exists;
- auth/quota/config/model classes never auto-retry.

Wired at both delivery surfaces (fix the class, not one site):
bot_relay.deliver (relay handler) and _run_delivery (local
message_agent runner). Failed deliveries now carry the classified
reason in the structured error payload (error.data.reason).

Sabotage-verified: with the retry blocks removed, 3 consumer tests
fail; with them present, 22/22 pass.
2026-08-23 20:00:18 -07:00
joaomarcos f5a9ba9ee6 perf(bluebubbles): move attachment reads off the event loop 2026-08-23 20:00:07 -07:00
Teknium b03b8ac51d fix(dashboard): name the exact gate trigger in fail-closed refusals
When the bind is loopback and the only gate trigger is
dashboard.public_url, the startup refusal now says so explicitly and
gives both exits (configure a dashboard auth provider, or remove
dashboard.public_url if the proxy no longer exists). Prevents the
stale-public_url mystery-locked-dashboard upgrade trap.

Adds a truth-table regression suite for should_require_auth and the
fail-closed message shape.
2026-08-23 19:55:17 -07:00
e-macgregor d3df14a7e3 fix(dashboard): secure loopback public URL proxy mode 2026-08-23 19:55:17 -07:00
briandevans 608a56ed7f fix(state): stop rebuilding the whole FTS index on every open when the trigram tokenizer is missing
`_init_schema` decided whether the FTS triggers needed repair by comparing
the live trigger count against `len(_FTS_TRIGGERS)`, the full six-name set.
Three of those six are the `messages_fts_trigram_*` triggers, and they are
declared only inside `FTS_TRIGRAM_SQL` / `LEGACY_FTS_TRIGRAM_SQL`, whose
`CREATE VIRTUAL TABLE ... tokenize='trigram'` needs a tokenizer SQLite only
gained in 3.34.

On an older build `_ensure_fts_schema` soft-fails that DDL by design (via
`_is_trigram_unavailable_error`) and returns False, so those three triggers
can never be created. The count is therefore pinned at 3, `3 < 6` is
permanently true, and the repair path ran on every single `SessionDB` open,
forever, while holding the SQLite write lock. It never converged: every
`hermes` command, gateway start, dashboard request and cron tick paid a full
re-index of the message corpus. That is ordinary LTS territory — Ubuntu
20.04 ships 3.31, RHEL/CentOS 8 and Alibaba Cloud Linux ship 3.26, and
Hermes has no minimum-SQLite gate precisely because it is supposed to
degrade gracefully here.

The v23 repair also ends by clearing `fts_rebuild_high_water` and
`fts_rebuild_progress`, which is correct after a genuine full rebuild but
means an interrupted `hermes sessions optimize-storage` silently lost its
resume point on the next open, restarting the chunked backfill from zero
every time.

Fix: keep `_FTS_TRIGGERS` as the single source of truth and derive two
subsets from it, then measure each half against the DDL that can actually
create it. `_fts_trigger_count` takes an optional `names` sequence
(defaulting to the full set, so no caller changes), and both branches gate
on `base_triggers_missing or (trigram_enabled and trigram_triggers_missing)`.
The counts are still taken before the DDL runs so they describe the
pre-repair state, while `trigram_enabled` is only known afterwards — hence
the combination at the `if` rather than at the assignment.

Behaviour is unchanged wherever the tokenizer exists: a genuinely missing
trigram trigger on a capable host still triggers the rebuild. Only the
permanently unsatisfiable comparison changes.
2026-08-23 19:31:35 -07:00
Finn763 bf15b050b1 fix(telegram): watchdog silent long-poll death via last getUpdates progress (#92991) 2026-08-23 19:26:41 -07:00
Teknium 3f5d37568e fix: managed-runtime guard no longer trips on sdist/build copies in the workspace
The bare-which() scanner rglobs the repo root; a CI job that builds the
wheel leaves an sdist extraction (hermes_agent-<version>/) in the
workspace, and the scanner re-found every already-exempted call site
under that versioned prefix — which can never match an _ALLOWED key —
failing the guard on untouched code (flaked PR #93420's Python-tests
job). _source_files now skips build/, dist/, *.egg-info, and any
top-level dir carrying PKG-INFO.

A/B: planted a fake hermes_agent-9.9.9/ sdist with a which('node')
site — old scanner 1 failed, fixed scanner 7 passed, clean tree
unchanged.
2026-08-23 19:15:28 -07:00
Teknium 7526bd39a8 feat: every subagent's prompt embeds the workspace's project context files
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.

The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.

Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 23fb949f2c feat: /review briefing embeds the workspace's project context files
load_workspace_context() resolves the parent's workspace via the same
_resolve_workspace_hint used for child prompts (explicit sources only —
TERMINAL_CWD / agent cwd hints, never a bare getcwd fallback, so the
#64590 install-tree-leak guard concern doesn't apply) and runs it
through agent.prompt_builder.build_context_files_prompt — the exact
discovery/priority/cap logic the main system prompt uses (.hermes.md >
AGENTS.md chain > CLAUDE.md > .cursorrules; SOUL.md skipped). The
result is embedded in the reviewer briefing as binding review
standards. Subagents are built with skip_context_files=True, so without
this the reviewer judged repo work without the repo's own conventions.

5 new tests incl. real-filesystem AGENTS.md discovery through the real
loader. Docs updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 22381edc11 feat: /review briefing carries the parent's loaded skills
The reviewer subagent now inherits the primary agent's working skill
context: collect_parent_loaded_skills() gathers launch-preloaded skills
(from the activation notes in ephemeral_system_prompt) and mid-session
skill_view loads (from assistant tool_calls in history), deduped and
capped at 8, and the briefing instructs the reviewer to skill_view each
and treat their conventions as binding for the assessment.

Reference-file reads (file_path=...) don't count as loads; full-skill
injection was rejected as too costly (a single dev skill can be 40KB+).

Docs: delegation.md /review flow updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 580060ffd8 fix: reuse first-observed sequence when announced items land via output_item.done
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.

Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
2026-08-23 19:02:30 -07:00
cxxCoolStar 4f3ae189a3 fix(agent): harden pending Responses tool call settlement 2026-08-23 19:02:30 -07:00
cxxCoolStar 720344cfba fix(codex): settle pending Responses tool calls when output_item.done is omitted
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
2026-08-23 19:02:30 -07:00
Teknium 081cdd9911 fix(terminal): subagents no longer hijack the tty with an interactive sudo prompt
delegate_task children run on worker threads of the parent process and
inherit the process-wide HERMES_INTERACTIVE=1 the CLI sets at startup.
_transform_sudo_command's interactive gate therefore fired inside
children with no sudo callback registered, falling through to the raw
/dev/tty password prompt: a password box printed mid-TUI from a
background thread, parallel children racing for the tty, and each child
blocked for the full 45s timeout.

Gate the prompt (and the sibling 'you will be prompted again' message
after an auth failure) on agent.delegation_context.is_delegated_child_context(),
the ContextVar set around every child run and propagated through
contextvars.copy_context onto the executor thread. Children now behave
as headless for sudo: configured SUDO_PASSWORD, the session cache, and
the NOPASSWD probe still work; otherwise the command fails gracefully
with a subagent-specific tip.

A/B verified: 3 regression tests fail on merge-base, 7/7 pass at head.
2026-08-23 19:00:47 -07:00
Teknium 9d0727d49b fix(state): single fail-closed cross-process authority for all full FTS rebuilds
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:

- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
  30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
  _rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()

Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.

Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
2026-08-23 19:00:36 -07:00
Christian Pompa d6bc3f2bca fix(tui_gateway): enable TCP keepalive on websocket sockets (dead-peer detection)
Without SO_KEEPALIVE a silently-dropped client (SSH tunnel reset, laptop
sleep, NAT timeout) leaves the TCP leg half-open forever: receive_text()
blocks indefinitely and the disconnect teardown (detach, orphan reap,
resume replay) never runs. The server then leaks the session and never
reclaims its orphans.

_disable_nagle already reaches the raw socket, so enable keepalive there:
SO_KEEPALIVE on, plus TCP_KEEPIDLE=30s / TCP_KEEPINTVL=10s /
TCP_KEEPCNT=3 on Linux and TCP_KEEPALIVE=30s on macOS. A dead peer is now
detected in ~60s instead of never. Best-effort like the Nagle tuning —
any failure to reach the socket is logged at debug and skipped.

Tests: new tests/tui_gateway/test_ws_keepalive.py fakes the socket and
pins SO_KEEPALIVE + the platform-specific idle tuning, plus the
no-transport no-raise path. tests/tui_gateway: 336 passed.
2026-08-23 18:58:40 -07:00
Teknium a7977771a6 fix(tui-gateway): revalidate transport ownership before sentinel-parking on WS disconnect
Reimplements the concept from #77129 on the current structure (viewer
rebinding from #83716 and _client_gone_interrupt_requested clearing are
preserved).

_close_sessions_for_transport snapshots owned sessions under
_sessions_lock, then wrote session['transport'] = _detached_ws_transport
WITHOUT re-checking that the session still pointed at the disconnecting
transport. A session.resume that rebinds the session to a new live
transport between the snapshot and the stomp got knocked back onto the
drop sentinel with an orphan-reap Timer armed against a client that is
attached right now.

The park now happens under _sessions_lock and first revalidates
ownership: if the session already moved to a different live transport,
the disconnect has nothing to tear down — skip the sentinel park AND the
reap scheduling. Regression test simulates the rebind landing between
snapshot and stomp.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 18:58:40 -07:00
Teknium 47d6ce78a2 fix(tui): make startup_orphan_reap recoverable and move its config onto dashboard.*
Follow-up to the #65422 salvage:

- startup_orphan_reap joins _RECOVERABLE_END_REASONS (kept distinct from
  ws_orphan_reap for forensics): every recovery fence
  (find_latest_gateway_session_for_peer, unarchive_recoverable_session,
  promote_to_session_reset) now treats a startup-swept row as an
  accidental end, so a sweep never makes a session unresumable.
- Config key moves from sessions.orphan_reaper to
  dashboard.startup_orphan_sweep in DEFAULT_CONFIG, next to its siblings
  ws_ping_interval / ws_ping_timeout / ws_orphan_reap_grace_s; the raw
  loader in tui_gateway.server reads the new key (fail-open on missing).
  cli-config.yaml.example and website/docs/user-guide/configuration.md
  follow the dashboard.* documentation pattern.
- New regression test: a stranded 'active' row (ended_at NULL, no live
  runtime) is swept AND still recoverable via peer-keyed lookup and fully
  revivable via reopen_session afterward.
2026-08-23 18:58:40 -07:00
halaprix d3e4b50e68 fix(tui): sweep orphaned tui/desktop/subagent session rows at gateway startup
Close session rows left ended_at IS NULL when the in-process websocket
orphan timer dies with the process (#65194). Dual-clock staleness
(started_at AND newest message), desktop included, live in-memory
sessions excluded, scheduled once from both entry.main and the WS
sidecar so desktop/dashboard boots also run the sweep.
2026-08-23 18:58:40 -07:00
A2chitect c305839442 fix(tui): log 4001 session-not-found rejections for diagnosability
Messages sent into a session whose in-memory runtime was detached on WS
disconnect and orphan-reaped vanished silently: _sess_nowait returned
4001 with no log line, so 'request arrived and was rejected' was
indistinguishable from 'request never arrived' in a 'message vanished'
report. Log a WARNING with the session id and request id on every
session-scoped RPC rejected against an unknown runtime id.

Adds a regression test asserting the 4001 response and the warning.

Closes #90428
2026-08-23 18:58:40 -07:00
kshitijk4poor 80b202f53a harden(adoption): review findings — exact-id donors only, divergence guard, honest donor_retired
Review batch (3 reviewers) on the final diff surfaced:
- H1: title-based donor matching could adopt AND non-recoverably retire
  an UNRELATED default-store conversation (bot titles collide by design;
  get_session_by_title has no archived filter/ordering). Donor probe is
  now exact-id only — the stranded repro always has the id.
- H2: re-adoption after a partial run could retire a donor that had
  accumulated NEWER messages than the profile copy (skip-based
  idempotency never merges). New divergence guard compares message
  counts and refuses retirement when the donor is ahead (still adopts).
- M1: donor_retired reported True even when every retirement step
  failed under suppress. Now per-segment tracked + warn-logged;
  True only when all applied.
- M3: adopted=False (e.g. import validation limits) was silent — now
  warn-logged with import errors.
- M4: archived donors are never re-adopted (no cross-profile cloning).
- Dead 'from pathlib import Path' dropped; contextlib no longer needed.

5 new red-first-verified regressions (title-collision immunity,
archived-donor immunity, non-vacuous owns_db gating with a real donor
seeded, divergent-donor retirement refusal, donor_retired truthfulness).
tests/tui_gateway: 578 passed. ruff clean.
2026-08-23 18:58:40 -07:00
kshitijk4poor 26a4f89ada fix(gateway): adopt stranded bot sessions from the default store on profile resume
Pre-#93296, the desktop routed session RPCs by the focused tile, so a
profile bot's turns executed on the default backend and its canonical
session accumulated in the DEFAULT profile's state.db. Post-fix, the
profile backend correctly receives the resume — but its store has never
seen the session, so the same chat 4001s forever (unreachable instead
of misrouted). Live repro: Teknium's Developer bot, session c93770.

- hermes_state_portability: SessionDB.adopt_session_lineage_from() —
  composes the existing export_session_lineage()/import_sessions()
  primitives; donor rows are archived (never deleted) with
  end_reason=adopted_by_profile, which is deliberately NOT in
  RECOVERABLE_END_REASONS so canonical-lookup resurrection cannot undo
  an adoption. Idempotent (already-present ids skip).
- tui_gateway/methods_session: profile-scoped session.resume falls back
  to adoption from the default store right before the 4007; ids unknown
  to BOTH stores still 4007 exactly as before, and launch-profile
  resumes never consult the fallback.
- tests: 10 new (7 unit on the primitive incl. compression-lineage
  unit adoption + non-resurrectable archive; 3 handler-level through
  server.handle_request incl. the live repro shape); db-ownership
  leak test taught that the shared launch handle probe is by design.

Follow-up to #93296/#93311; part of #93091.
2026-08-23 18:58:40 -07:00
fangliquanflq 654d537088 fix(agent): honor structured quota reset signals 2026-08-23 18:43:12 -07:00
fangliquanflq c2090ba6b4 fix(desktop): distinguish provider quota exhaustion 2026-08-23 18:43:12 -07:00
RickyYii 6b3a7af73d fix(security): cover privilege wrappers and command-string options
Review follow-up on #84203. Both points reproduce; neither was a regression
from the first pass, but both are live bypasses of the same guard.

**Privilege and namespace wrappers were missing.** The allowlist covered the
coreutils-shaped wrappers but not the privilege ones, so each of these ran a
lifecycle script straight past the walk:

    pkexec bash ~/restart.sh
    runuser -u root -- bash ~/restart.sh
    setpriv --reuid=0 -- bash ~/restart.sh
    systemd-run --scope bash ~/restart.sh
    nsenter --target 1 --mount bash ~/restart.sh
    unshare -r bash ~/restart.sh

Added `pkexec`, `su`, `runuser`, `setpriv`, `systemd-run`, `nsenter` and
`unshare`, each with the value-taking options that would otherwise be
mistaken for the command (`nsenter -t 1`, `systemd-run -p X=1`,
`runuser -u root`, …).

**An option can carry a command STRING, not an argv tail.** `env -S` and
`su`/`runuser` `-c` take shell source. The peel treated the operand as an
opaque value and skipped it, so `env -S 'bash ~/restart.sh'` was never
scanned — the string went unread rather than being recursed into.

`_STRING_COMMAND_OPTIONS` now names those options and their values are
re-scanned as shell source, the same treatment `sh -c` payloads already get.
They are read at the ORIGINAL command token, before the transparent-prefix
peel, because peeling past `su`/`env` would discard the very option carrying
the command. `--opt value` and `--opt=value` are both handled.

Scope, stated plainly: this is an enumerated allowlist, not a general
solution to "wrapper that execs its tail". A wrapper outside the set, or a
value-taking option outside these tables, still resolves to no reference —
that fails open, exactly as it did before this PR, and it is a miss rather
than a false block. The reviewer offered "extend the set with tests, or
document that the list is heuristic"; this does the first and states the
second.

Tests: 23 new cases (220 in the file) — every added wrapper against a script
reference including the value-operand option forms, both command-string
option spellings for env/su/runuser, and the same wrappers around ordinary
work (`pkexec systemctl status nginx`, `su -c 'ls -la'`, `env -S 'echo hi'`,
`nsenter -t 1 -m ps aux`) which must stay allowed. 15 fail on the tree
before this commit.

False positives re-checked at scale: the 9,258 command lines from this
repo's own scripts and docs give an identical verdict set before and after —
0 new false positives, 0 lost detections, 0 exceptions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
RickyYii a19e1bae10 fix(cron): stop a relative path from disabling the data-sink exemption
`_mask_data_sink_arguments` exempts lifecycle text living in the arguments of
executables that cannot run them (`grep`, `rg`, `journalctl`, `sqlite3`, …),
so hunting for a restart string in logs is diagnostics rather than a command.
The exemption is dropped when an argument looks like an escape back into
execution — including anything starting with a dot, because sqlite3 spells
its escapes as dot-commands (`.shell`, `.system`).

But `.`, `./x` and `../x` are ordinary path operands, and

    grep -r 'systemctl restart hermes-gateway' .

is the most ordinary recursive search there is. The leading-dot test treated
its `.` operand as a sqlite3 escape, disabled masking for the whole segment,
and blocked the command outright — the exact false-positive class the
exemption exists to prevent, on the shape most likely to hit it. Searching a
relative subdirectory (`./logs`, `../archive`) fails the same way, as does a
relative sqlite3 database path (`sqlite3 ./stats.db "SELECT ..."`).

Require a dot followed by a NAME character (`^\.[A-Za-z]`) so a dot-command
still defeats the exemption while a relative path stays a path. A dotfile
operand (`.env`) still reads as a dot-command — conservative, and unchanged
from today's behavior.

This narrows a security guard in the permissive direction, so the escape
hatches are pinned explicitly: with a relative-path operand present,
`.shell`/`.system`, psql's `\!`, a pipe into `sh`/`bash`/`sudo sh`/`xargs`,
command substitution, and a `;`/`&&` continuation all still block. Only the
segment's own data arguments are masked, and only when nothing in it can
reach execution.

Tests: 18 new cases in tests/hermes_cli/test_gateway_restart_loop.py — the
relative-path shapes that must now be allowed, plus the ten escape-hatch
shapes that must still block. The allow cases fail on the unfixed tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
RickyYii 5921ba8c06 fix(security): see through wrapper prefixes in the gateway lifecycle guards
`sudo`, `env`, `nohup`, `timeout` and friends exec their argument tail, so
the command that actually runs sits further right. Three guards read only the
first token of a segment, saw the wrapper, and never inspected what it runs:

  bash ~/restart.sh                      → blocked
  sudo bash ~/restart.sh                 → allowed
  launchctl submit -l com.x -- helper    → blocked
  sudo launchctl submit -l com.x -- helper → allowed

Same foot-gun, one word of prefix. That reaches both enforcement points —
`cron.jobs.create_job` and `tools/terminal_tool.py` under `_HERMES_GATEWAY=1`
— and defeats the label-independent submit block that #62891 added precisely
because a persistent helper is the indirect route to a restart loop.

`_peel_transparent_prefixes()` walks past a bounded chain of these wrappers,
skipping their own options, their value-taking options (`sudo -u deploy`,
`stdbuf -o0`), `VAR=value` assignments, a `--` end-of-options separator, and
`timeout`'s duration operand, then returns the index of the real command. It
is applied to the referenced-script walk, the `sh -c` payload walk, and the
`launchctl submit`/`bootstrap` block.

In the referenced-script walk the peel is ADDITIVE — the segment is read at
the original token and again at the peeled one — because peeling must never
remove a reference the un-peeled read would have found. A local script named
`./timeout` is a script, not the coreutils wrapper, and consuming it as a
prefix would have silently stopped scanning it. (The other two call sites
need no such care: no wrapper name is also a shell name or `launchctl`, so
peeling there can only add.) That split is why the per-index logic now lives
in `_references_at()`.

This is not a new reading of shell syntax for this module — `_PIPE_TO_INTERPRETER`
already treats `sudo ` as transparent for the pipe case (`... | sudo sh`).
This generalises the same reading to the command position.

Deliberately NOT applied to the data-sink masking in
`_mask_data_sink_arguments`: peeling there would widen an exemption, and the
conservative reading is the safe one.

No false positives: peeling only changes which token is treated as the
command, so a wrapper around ordinary work resolves to a non-shell executable
and yields nothing, exactly as before (`sudo apt-get update`,
`timeout 60 curl ...`, `nice -n 10 make -j4`, a bare `env`).

Tests: 40 new cases in tests/hermes_cli/test_gateway_restart_loop.py — every
wrapper form against a script reference, a dot-source, a nested `sh -c`
payload and `launchctl submit`, plus the benign wrapped commands, a wrapped
clean script, and the `./timeout`-style lookalike names that pin the additive
reading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
Teknium f2639f8872 fix(cron): preserve map keys as ids and skip junk values when flattening id-keyed jobs.json
Harden the id-keyed-map flatten with an id-preserving merge:
{**value, "id": value.get("id") or key} — an inline "id" wins,
otherwise the map key is adopted (external tools often key by id and
omit the inline copy; plain list(values) would emit id-less records
that collide or get dropped downstream). Non-dict junk values are
skipped with a warning instead of crashing the load. The self-heal
rewrite persists the id-merged, junk-free records.

Tests: key adopted when no inline id (and inline id wins over a
differing key), non-dict junk skipped with warning + list_jobs
survives + self-heal persists only valid records, all-junk map
flattens to [].
2026-08-23 18:27:32 -07:00
Teknium ec4b3bc06a fix(cron): self-heal id-keyed jobs.json to canonical list form on load
Layer on the load-boundary flatten: when load_jobs() encounters an
ID-keyed jobs map ({"jobs": {"<job_id>": {...}, ...}} — written by
external tools or hand edits, never by save_jobs()), it now not only
flattens to the list contract but persists the canonical
{"jobs": [...]} form back to disk via the existing auto-repair path
(save_jobs), so the store self-heals and subsequent reads are
idempotent.

Note: _peek_jobs_unlocked() intentionally does NOT tolerate the dict
shape — it returns None so the save path never shrink-merges against
an unrepaired baseline. The flatten + repair live only at the
load_jobs() boundary.

Regression tests cover the flatten, the reported list_jobs() traceback
path, idempotent on-disk repair, and the empty-map edge case.

Salvaged from PR #92994.

Co-authored-by: a-yeyang <88581400+a-yeyang@users.noreply.github.com>
2026-08-23 18:27:32 -07:00
WK Wong 5a24dcf4f2 fix(cron): normalize id-keyed jobs stores on load 2026-08-23 18:27:32 -07:00
liuhao1024 203f111c98 fix(install.ps1): initialize LastResolver before the resolved-path report
ConvertTo-LongPath short-circuits for ordinary long paths (no ~\d alias),
so $script:LastResolver is only assigned when a short path actually needs
expansion. The ResolvedPathReport block read it unconditionally, which is
fatal under Set-StrictMode before any install stage runs (#93017: fresh
installs died at line 367 through three different invocation styles).

Initialize it to 'none' — the resolver's own value for "nothing ran" — at
script scope before Set-LongProfileEnvVars can invoke a resolver.
2026-08-23 18:27:21 -07:00
Teknium 3e3e6f94a0 test(cua): pin PATH-preservation contract, not byte equality
The CUA spawn-env tests froze PATH == '/usr/bin:/bin' verbatim.
_sanitize_subprocess_env now (intentionally) prepends the hermes
console-script dir for all sanitized children (#92998), so these
assertions flip to the contract: original entries preserved as
suffix, hermes bin dir first when prepended.
2026-08-23 18:27:15 -07:00
Teknium ec06e706f1 test(cron): e2e regression — scrubbed child env resolves bare hermes under minimal parent PATH
Exercises the real build_subprocess_env()/_resolve_hermes_bin_dir chain (no
helper mocks) under a simulated systemd/cron minimal PATH, the exact call
path cron/scheduler._run_job_script uses. Companion to #93082.
2026-08-23 18:27:15 -07:00
UniversePeak b0001f45a2 fix(cron): keep hermes console script on child PATH 2026-08-23 18:27:15 -07:00
Teknium 04dd2bb233 test(agent): drain truncation warnings before and after each prompt-builder test
Follow-up to the ContextVar-leak fix: the autouse fixture now drains on
both sides (drain(); yield; drain()) so earlier files can't pollute this
file's assertions either.
2026-08-23 18:27:12 -07:00