Commit Graph

14240 Commits

Author SHA1 Message Date
the3asic 18ac3c4fb6 fix(state): defer corrupt FTS rebuilds past live operations 2026-08-31 12:08:30 -07:00
devops ad08a5813c fix(estop): honor canonical ~/.hermes/ESTOP from profile gateways
Profile processes launch with HERMES_HOME=~/.hermes/profiles/<name>, so
`hermes pause` at the fleet root did not bind fleet-analyst dispatch
(t_7b65ff88). Check/resume both the process home and the fleet root.
2026-09-01 00:35:29 +05:30
Teknium b0acc558b1 fix(cli): supervised gateway launches skip the sticky active_profile redirect
Generalize the HERMES_S6_SUPERVISED_CHILD supervisor-marker mechanism so
ANY supervised gateway launch (systemd, launchd, Windows Scheduled Task,
external supervisor) skips the active_profile redirect in
_apply_profile_override(). Previously only the s6 container marker was
honored, so a systemd-launched default-profile gateway with
HERMES_HOME=<root> followed the sticky active_profile file and silently
assumed another profile's identity — logging under that profile's tree
and connecting with its Telegram bot token (double-polling a token owned
by that profile's own live gateway).

- hermes_cli/main.py: honor HERMES_SUPERVISED_CHILD (new generalized
  marker), HERMES_S6_SUPERVISED_CHILD (back-compat), INVOCATION_ID
  (systemd; gateway commands only, since it leaks into every descendant
  of systemd-launched processes), and HERMES_GATEWAY_EXTERNAL_SUPERVISOR.
- hermes_cli/gateway.py: export HERMES_SUPERVISED_CHILD=1 in generated
  systemd units (user + system) and the launchd plist.
- hermes_cli/gateway_windows.py: export it from the Scheduled-Task cmd/vbs
  launchers and the windowless respawn env overlay.
- hermes_cli/service_manager.py: export it alongside the s6 sentinel.
- tests: regression coverage for all markers + non-gateway INVOCATION_ID
  neutrality + generated-unit marker presence.

Fixes #74872
2026-08-31 12:05:07 -07:00
Teknium 11ba76c0fb fix(gateway): run MCP shutdown off-loop with a bounded wait on the shutdown path
shutdown_mcp_servers() blocks on future.result(timeout=15) which, called
from the gateway event-loop thread during SIGTERM teardown, freezes the
loop for up to 15s when the MCP loop and its stdio children are torn down
concurrently. Supervisors with a shorter kill grace (s6-overlay: 3s)
SIGKILL the gateway before lifecycle_ledger.mark_exited() runs, producing
phantom 'exited UNCLEANLY' reports on every subsequent boot.

Run the sync shutdown on a daemon thread and poll via _await_thread_exit
with a 5s budget; proceed with teardown if it wedges. Fixes #82874;
completes the shutdown half of #64155.
2026-08-31 12:04:10 -07:00
Teknium 51773a7733 test(state): physical-corruption acceptance tests for the fail-closed classifier
Real byte-flip fixtures (no mocks) proving the #96038/#98090-class fix
end to end, closing the acceptance gate on issue #97940:

- test_canonical_btree_corruption_fails_closed: checkpoint the WAL,
  clobber every messages-table B-tree leaf page header, then assert a
  live append raises the genuine bare SQLITE_CORRUPT, the classifier
  refuses the FTS route, no rebuild/detach/stale-marker side effects
  occur, and the field incident's misdiagnosis log line ('canonical
  message rows are preserved') never appears.
- test_fts_only_corruption_still_self_heals: contrast case — a real
  messages_fts_data shadow-table stomp raises SQLITE_CORRUPT_VTAB (267),
  is classified as FTS-scoped, and the write path still self-heals with
  canonical rows intact.

Sabotage-verified: reverting the classifier fix (96739033c4) makes the
canonical-corruption test fail by entering the FTS self-heal route.

Credits @fangliquanflq (PR #98090) for the production timeline analysis
and @diatche (PR #96038) for the classifier fix these tests gate on.
Refs #97940, #98077.
2026-08-31 12:02:58 -07:00
andyst-dev 4d3e1e4d13 fix(desktop): prevent venv scan timeout on busy Windows hosts 2026-08-31 12:00:33 -07:00
Mariano Nicolini e681decfae fix(aux): policy-check the whole auxiliary model ladder
Only the catalog step was filtered. With no fast-family match in the allowed
catalog it returned empty and the ladder fell through to a public
recommendation, which could hand titling a model the org blocks.
2026-08-31 15:58:27 -03:00
HexLab98 85118cc9d4 test(cli): cover mixed-config ImportError recovery hint on chat startup (#96900) 2026-09-01 00:26:55 +05:30
kshitijk4poor fe0cfdf99c fix(redact): keep dotted config-key scans linear past the keyword pre-gate
The _CFG_SECRET_WORD_RE pre-gate only skips secret-FREE text. A compaction
payload containing one real secret assignment plus a long opaque dotted run
still reaches _CFG_DOTTED_RE's backtrackable '*' prefix, which re.sub retries
from every byte of the run — quadratic while holding the GIL (same class as
the _ENV_ASSIGN_LOWER_RE fix in this branch, #99255).

Anchor each attempt to the start of a key run with a negative lookbehind.
Match set is unchanged: any match starting mid-run implies a leftmost match
at the run start, verified 20/20 identical over a dotted-config corpus.
30k-char adversarial run: 102s -> 0.015s.
2026-09-01 00:25:45 +05:30
Teknium f50b5bb0fa fix(compression): dead Codex summary streams fail over in 60s instead of stacking 5-minute waits
The Codex auxiliary Responses adapter enforced a single absolute
deadline (300s floor for compression). A dead stream held the entire
budget before fallback ran, and repeated compression attempts stacked
those waits into 20+ minute 'Summarizing thread' stalls (masoria debug
bundle, Aug 31 2026). Meanwhile a healthy-but-slow reasoning summary
was killed at the same absolute deadline even while producing tokens.

Replace the absolute kill with progress-aware deadlines:
- 60s no-progress window for the first substantive payload AND between
  payloads; keepalive/lifecycle frames do not re-arm (mirrors the
  commit-fence gating, #96707)
- a live stream re-arms per token and is bounded only by
  _aux_stream_total_ceiling() (max(600s, 4x configured timeout)), the
  same backstop the streamed chat.completions path already uses
- the compression critical-path retry gate now distinguishes failure
  cost: a cheap first-token no-progress failure retries the same
  provider once; mid-stream stalls and ceiling hits still skip straight
  to provider fallback (#54465 semantics preserved)

Live A/B (real OpenAI SDK against a local SSE server, real adapter):
dead keepalive-only stream: main waits the full budget; fixed fails
over at the window. Slow-but-alive stream (tokens past the configured
timeout): main kills it mid-generation; fixed completes.
2026-08-31 11:52:51 -07:00
embwl0x 221312834e fix(gateway): bound signal interrupt grace 2026-08-31 11:47:16 -07:00
Teknium f680a4dc7d fix(gateway): require FTS provenance before transcript rebuild-and-retry
Widen #96038's fail-closed classifier to the gateway transcript retry
path: SessionStore._is_fts_corruption_error no longer treats a generic
'database disk image is malformed' as FTS-only damage. It now delegates
to SessionDB._is_fts_write_corruption_error (SQLITE_CORRUPT_VTAB result
code or explicit fts5 corrupt-structure text) and only keeps the
messages_fts-named cases. Structural corruption falls through to the
bounded retry/backoff path instead of rebuilding FTS and retrying writes
against a damaged database.

Sibling site spotted in PR #98090 by @fangliquanflq.
2026-08-31 11:42:23 -07:00
Pavel Diatchenko 96739033c4 fix(state): fail closed on unscoped corruption 2026-08-31 11:42:23 -07:00
Mariano Nicolini 6e20ec4101 fix(nous): apply the org policy before the free/paid tier split
Rescuing an empty list after partitioning put paid models back into a
free-tier user's selectable list, and the dashboard could pick one as the
silent default. Narrowing first also drops the separate unavailable-list
filter.
2026-08-31 15:40:49 -03:00
Mariano Nicolini e89f0087b4 fix(models): key the pricing cache per credential, not per auth state 2026-08-31 15:29:53 -03:00
teknium1 d10ef89ee5 fix(docker): keep forwarded secret values out of world-readable argv
docker run/exec argv previously carried -e KEY=VALUE pairs for every
forwarded/passthrough variable. On Linux /proc/<pid>/cmdline is
world-readable regardless of process owner, so every allowlisted secret
was visible to all local users via plain ps for the duration of every
terminal call.

Emit name-only -e KEY flags and supply values via the docker client
subprocess env instead: the docker CLI resolves valueless --env KEY from
its own environment (documented docker/podman behavior), moving secrets
from /proc/*/cmdline (0444) to /proc/*/environ (0400). Covers the docker
run container-start path, the recreation/recovery path, the init-seeding
exec path, and the per-command runtime exec path.

Reported by @sashalab. Fixes #96268
2026-08-31 11:29:17 -07:00
joaomarcos 0e9b57d310 fix(state): cap read connections per PROCESS and yield when the fd table is tight
Follow-up to the per-file budget in the previous commit, which closed the
scaling axis it measured and left three others open.

* A per-file ceiling still lets the cost grow with the PROFILE count: a
  multiplexed gateway serves N profiles from one process and each has its own
  state.db, so `_READ_POOL_MAX` bounded each file while the process total went
  unbounded — the per-instance bug one level out. `_READ_POOL_PROCESS_MAX`
  (three files' worth) now bounds the process, and a miss reclaims an idle
  connection from ANY path before degrading: a profile quiet for an hour must
  not hold descriptors the profile being served right now needs.

* Hermes's SQLite descriptors are only ever a share of the fd table. In #98573
  the ~20 state.db handles were not the whole 256 — they were the share that
  pushed httpx sockets and terminal subprocess pipes over, and the EMFILE
  surfaced in tools/terminal_tool.py rather than here. New read connections are
  now refused when the process is within `_FD_HEADROOM_RESERVE` of its soft
  RLIMIT_NOFILE, measured from /proc/self/fd or /dev/fd and cached briefly. The
  guard fails OPEN where it cannot measure (Windows has neither the fd
  directory nor RLIMIT_NOFILE, and a CRT limit in the thousands) and CLOSED on
  evidence — including a probe that could not get a descriptor of its own.
  `_read_open_denied_fd_headroom` makes it diagnosable from a running process.

* Writer connections cannot be rationed the way read connections can: a
  SessionDB without one cannot write. Their only real bound is not opening
  redundant handles, so a process that accumulates more than
  `_HANDLES_PER_PATH_WARN` handles on one file now says so once, and the next
  duplicate is visible before it is an incident instead of inferred from an
  lsof after one.

`_READ_POOL_MAX` itself is deliberately unchanged at 8. Retuning that constant
is #98585's subject; with a process ceiling above it and the headroom guard
in front of it, the value is no longer the binding constraint.

Refs #98573

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 11:23:48 -07:00
joaomarcos e8c41568ca fix(state): bound state.db read connections per FILE, and stop opening two gateway handles
Issue #98573 reports a long-lived gateway holding ~20 `state.db` descriptors
that never shrink, walking into the 256 soft RLIMIT_NOFILE a launchd/systemd
service manager hands the process. The cause named there — a per-thread
`threading.local()` read connection — is already gone (87aedbe7b6 pooled the
read connections, 0472c31aa1 added the peak permit). Measured on main: one
SessionDB with 40 concurrent reader threads peaks at 9 live connections, not 40.

The symptom survives one layer up. `_READ_POOL_MAX` was enforced by a
BoundedSemaphore owned by each SessionDB, which bounds the wrong noun: the
descriptors are spent on a FILE, so every additional handle on one state.db got
its own allowance and peak scaled as `instances x (1 + _READ_POOL_MAX)`.

Two changes, both needed:

* The permits move to a per-path `_PathReadBudget`, shared by every SessionDB
  in the process that points at that file. A permit miss first reclaims an IDLE
  pooled connection from a peer handle before degrading to the writer lock —
  without that, whichever handle warmed up first would pin the whole budget and
  permanently demote every later one (a cron job's transient handle, a second
  profile's store) to the locked writer connection.

* `GatewayRunner` borrows `SessionStore`'s handle instead of opening its own.
  Both caches resolve the same `_default_db_path()`, so the process was holding
  two writer connections and two read pools against one file for no reason, and
  doubling again per profile on a multiplexed gateway. The store owns the
  connection and sweeps it at shutdown; the runner's cache now holds only the
  async wrapper and its sweep skips borrowed handles.

Measured peak live connections against one file, 40 reader threads, by handle
count 1/2/4/8:

  before: 9 / 18 / 36 / 51   (51 not 72 only because the sample window ended
                              before every pool filled)
  after:  9 / 10 / 12 / 16   (read connections capped at 8 in total; the
                              remainder is one writer per handle, and the
                              gateway's per-profile pair is now one)

Fixes #98573

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 11:23:48 -07:00
mbac 3e912874bf fix(models): support OpenRouter preset references 2026-08-31 11:19:26 -07:00
nftpoetrist 4d4cffd118 fix(profiles): make_targz writes to a temp file and renames, not the destination directly
tarfile.open(archive_path, "w:gz") truncates the destination the instant
it opens. If tf.add() fails partway (disk full, permission loss,
interruption), whatever was previously at that path is gone — including
an existing profile or board export the caller chose to overwrite. This
is the same failure shape a7e7de6407 just fixed for the desktop gateway
file-save path, one commit earlier in the same window, but it was never
propagated to this shared archive-writing primitive even though board
export gained a new caller into it in that same window.

make_targz now writes into a sibling temp file (mkstemp, same directory
as the destination so the final step is a same-volume rename) and only
replaces the destination via os.replace() after the archive is fully
written and closed, mirroring the mkstemp+os.replace pattern already
used throughout this codebase (agent/secret_sources/_cache.py,
cron/jobs.py, gateway/status.py, etc). The temp file is unlinked on any
failure.
2026-08-31 11:19:05 -07:00
anhtahaylove 99d037eeb9 fix(compression): count streamed reasoning details as progress 2026-08-31 11:18:53 -07:00
anhtahaylove 4d95d8eca8 test(context): cover preflight timeout provider boundary 2026-08-31 11:18:53 -07:00
anhtahaylove de49e1ef05 fix(context): fail closed when preflight compression stalls 2026-08-31 11:18:53 -07:00
fangliquanflq ba0f5839d4 fix(redact): keep lowercase assignment scans linear 2026-08-31 11:18:41 -07:00
RickyYii 3145986c20 fix(cli): honour model_aliases api_key, stop cross-provider key leak (#83612)
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes #83612.
2026-08-31 10:59:45 -07:00
liuhao1024 b8d3c8aa07 test(feishu): cover DM paired-mode card clicks and fail-closed identity checks
Adapt five scenarios from @liuliu0223's regression suite in #99021:
- paired-mode (empty allowlist) positive paths for approval and
  update-prompt cards, the DM breakage this fix resolves
- fail-closed rejection of clicks with an empty operator identity
- chat-mismatch rejection when an approval card is forwarded
2026-08-31 10:54:58 -07:00
liuhao1024 a89706e4cb fix(feishu): gate approval/update-prompt card clicks on operator allowlist, not group policy
The synchronous card-action handlers and the update-prompt resolver
authorized clicks with _allow_group_message(), which answers "may this
sender chat in this group?" — with group_policy=open it returns True
for everyone. The approval resolver already used the correct operator
gate (_is_interactive_operator_authorized), so the three code paths
disagreed: with an open group policy an out-of-allowlist click on an
update-prompt card was fully executed, and approval clicks returned a
resolved-looking card before being rejected asynchronously.

Authorize all three paths with _is_interactive_operator_authorized(),
which checks membership of admins ∪ allowed_group_users (wildcard and
the empty pairing-mode allowlist keep their existing allow semantics,
matching _admit's DM pairing default). A missing operator identity now
fails closed on the update-prompt resolver instead of skipping the
check.

Fixes #96045
2026-08-31 10:54:58 -07:00
Teknium c29485fa64 test(install): repin #87460 probe tests on the line-walk contract
The pre-release line-walk (#96601 salvage) moved the download/probe body
into install_node_line() and rejects an unstartable binary BEFORE
adoption via node_satisfies_build on the extracted tree. The sandboxed
driver now inlines all three functions, and the broken-node test pins
the stronger pre-adoption rejection instead of post-adoption cleanup.
2026-08-31 10:53:18 -07:00
semao0 39540a03fc fix(install): never adopt a pre-release Node.js build
install_node() picks the newest tarball out of
nodejs.org/dist/latest-v${NODE_VERSION}.x/ and installs it without ever asking
whether the binary inside is usable. That index currently serves
node-v26.8.0-<os>-<arch>.tar.xz -- a final-looking filename -- whose binary
reports v26.8.0-alpha.0.0.0. Node publishes the headers tarball named by
process.release.headersUrl only for final releases, so node-gyp cannot compile
against that build and every native module fails to install.

Probe the extracted tree before it replaces anything on disk, and fall back to
an older release line when the probe rejects it, instead of leaving the install
with an unbuildable runtime. Mirror the guard in node-bootstrap.sh, and let
_managed_node_tree_outdated() treat a pre-release tree as outdated so an
already-broken install heals itself -- the existing heal only fires below the
target major, and a pre-release sits above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GrEaXSjvFoBXKxAHTjnUbS
2026-08-31 10:53:18 -07:00
teknium1 d3a45a9ce4 test(agent): update enqueue-after-close contract to the #94736 self-heal
The old contract (write after close() raises AttributeError and drops
the token delta) is superseded: the persistence boundary now reopens
the connection, so the delta lands. Assert the new, stronger contract.
2026-08-31 10:51:19 -07:00
teknium1 9db48053e8 fix(state): self-heal SessionDB writes after close() races an in-flight worker
Subagent/cron sessions died mid-run with "Session DB append_message
failed: 'NoneType' object has no attribute 'execute'": a teardown owner
(cron run_job finally, delegate timeout owner, agent close()) called
SessionDB.close() — nulling _conn — while a still-unwinding worker had
one more transcript flush to land. The flush then hit None.execute, the
turn force-ended as session_persistence_failed, and the session tail was
silently dropped while cron delivery reported last_status: ok.

Fix at the shared persistence boundary: _execute_write and the _read_ctx
writer-lock fallback detect the closed handle under self._lock and
reopen a connection to the same database file with a loud WARNING naming
the race. Read-only handles never reopen — they raise an explicit
'was closed' error. A failed reopen raises an OperationalError naming
the teardown race so classify_persistence_error gets a real cause.

Closes #94736
2026-08-31 10:51:19 -07:00
HexLab98 97f2a571b5 test(terminal): cover hung-wait bound, parent-tid interrupt, and cron inactivity watchdog
Pin that execute() returns at the wall-clock deadline when the inner wait
never returns, that /stop on the tool-worker tid still kills the subprocess,
that the cron inactivity helper fires while the caller thread is blocked,
and that ContextVars plus the activity callback reach the deadline worker.
2026-08-31 10:42:39 -07:00
Teknium ce942ab125 fix(update): fingerprint orphan backends from the classification psutil handle
The orphan-backend classifier fingerprinted candidates via
gateway.status.get_process_start_time, which prefers /proc/<pid>/stat —
the HOST process table, in clock ticks. Under the fake-psutil test harness
(and any containerized run where the PID number happens to exist on the
host) that returns the WRONG process's fingerprint in the WRONG units,
while pid_is_hermes verifies via psutil centiseconds at kill time: the
guard would then refuse every legitimate reap. Read create_time() from the
same psutil handle used for classification, quantized exactly like
gateway.status does on Windows, so the fingerprint round-trips.

Also covers the Windows-lane sibling: test_uses_netstat_and_taskkill_on_windows
now pins the guarded call path, plus a new refusal test for a non-bridge
listener PID (#89614 class).
2026-08-31 10:41:54 -07:00
Teknium 90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
Ayush Nangia c923b53913 fix(update): refuse gateway ancestor tree-kill on Windows 2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
gebilaowang404 cdd063528f fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).

Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden

Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.

Refs #90471, #89614

Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
2026-08-31 10:41:54 -07:00
Teknium d8f8a07ee3 fix(compression): truncated summaries no longer become compaction checkpoints (port of earendil-works/pi#7048)
A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.

Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
  main-model fallback (a larger output budget may finish the summary), and
  on terminal failure ABORTS compression preserving the session unchanged
  (new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
  exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
  recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
  existing retry/backoff loop.

_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.

Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.

Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).
2026-08-31 10:39:55 -07:00
fangliquanflq da090aa4ba fix(prompt): preserve resumed workspace provenance 2026-08-31 10:10:25 -07:00
fangliquanflq c6ee4e0809 fix(prompt): skip bundled AGENTS.md for desktop launch cwd 2026-08-31 10:10:25 -07:00
Teknium 793690a0a0 test(cron): pin _REDACT_ENABLED in incident redaction test
test_redaction_applied_to_incident_error asserted real redaction while
relying on the ambient HERMES_REDACT_SECRETS default. agent.redact
snapshots _REDACT_ENABLED at import time; when a co-collected module
(tests/cron/test_codex_execution_paths.py) imports the gateway chain at
COLLECTION time under a shell exporting HERMES_REDACT_SECRETS=false, the
snapshot freezes False before the conftest env scrub runs, and the test
fails only in full-directory runs. Pin the flag via monkeypatch like the
~30 other redaction tests do.

Bisect evidence: pytest tests/cron/test_codex_execution_paths.py
tests/cron/test_cron_incidents.py -k redaction_applied -> 1 failed on
main under HERMES_REDACT_SECRETS=false; passes with the pin.
2026-08-31 10:09:59 -07:00
liuhao1024 db2fd5f59a fix(cli): answer clarify headless in single-query turns
hermes chat -q wired the interactive prompt_toolkit clarify callback
unconditionally, but a -q turn never builds the prompt_toolkit
application — the modal can never be painted or answered, so the turn
polls its response queue until agent.clarify_timeout expires (default
3600 s, 0 = unlimited). The gateway, cron jobs, the kanban dispatcher
and inter-agent wakeups all deliver work as -q turns. Route the
single-query case to a headless callback at the agent-construction site
that already knows _single_query_mode, mirroring _oneshot_clarify_callback
on the -z path (#94943; third member of the family after #86909 and
#88013).
2026-08-31 10:09:42 -07:00
Michael Nguyen 808a22ea00 fix(discord): gate relay-only thread rename kwargs 2026-08-31 10:09:23 -07:00
Teknium b7ebe6456f fix(xai): request-local alias provenance + collision-safe wire aliasing
Hardens the two #95003 alias carriers per review feedback on #95019/#95011:

- _alias_reserved_tools / _rename_tool_search_bridge_for_xai now return the
  alias map THIS request emitted; the transport stashes it
  (_last_wire_aliases) and normalize_response reverses ONLY those aliases.
  A real user/plugin/MCP tool named hermes_tool_search is never silently
  dispatched as tool_search when no alias was sent.
- Collision safety: if a real tool already occupies the alias name, the
  bridge takes hermes_tool_search_2/_3 — no duplicate wire declarations.
- Legacy static reverse map retained only for normalize-only call sites
  that never built a request on the transport instance.
- chat_completion_helpers resets provenance per request so stale maps from
  a prior request can't leak into the next response's dispatch.

Refs #95003
2026-08-31 10:09:04 -07:00
liuhao1024 5e2f8b9865 fix(xai): alias the reserved tool_search bridge name on chat completions
xAI's chat-completions API reserves the function name tool_search for
its native server-side tool and rejects the whole request when the
client Tool Search bridge declares it (HTTP 400 'The function name
tool_search is reserved for the tool_search tool', #95003) — Grok
providers were unusable whenever the bridge assembled into the payload
(default tools.tool_search: auto). Mirror the web_search treatment in
transports/codex.py: rename the bridge's wire declaration to
hermes_tool_search for xAI targets (deep-copied first, #27907 lesson)
and map the alias back to tool_search in normalize_response so dispatch
is unchanged. Alias matches the Codex-side fix for the same class
(#83122).
2026-08-31 10:09:04 -07:00
David Zhang de123be524 fix(xai): alias the reserved tool_search bridge on the wire (#95003)
xAI reserves the function name `tool_search` for Grok's native
server-side Tool Search and rejects the client declaration outright:

    HTTP 400 {"code":"invalid-argument","error":"The function name
    tool_search is reserved for the tool_search tool"}

Hermes's progressive-disclosure bridge registers exactly that literal
(`TOOL_SEARCH_NAME` in tools/tool_search.py) and assembly is not
provider gated, so with the default `tools.tool_search.enabled: auto`
every grok turn fails the moment the catalog crosses the threshold —
mid-session, which reads to the user as a session reset.

Same treatment as the two collisions already handled on this
transport (xAI `web_search` #48108, OpenCode reserved names #85589):
alias to `hermes_tool_search` on the wire in build_kwargs, map back in
normalize_response so Hermes dispatch and the bridge contract are
untouched. `tool_describe` / `tool_call` are not reserved by xAI and
are left alone.

Folds the per-provider rename helpers into one `_alias_reserved_tools`
owner parameterized by the reserved-name tuple, and extends the
existing `_RESERVED_ALIAS_TO_NAME` reverse map so the dispatch-side
un-aliasing needs no new branch.

Scope note: this covers the Responses transport, which is where every
api.x.ai route lands by default (`_fallback_api_mode` maps api.x.ai →
codex_responses, and the xai provider profile declares it). An xAI
model forced onto `api_mode: chat_completions` would still hit the
400; that path has no provider-specific tool rewriting today and would
need the symmetric hook in agent/transports/chat_completions.py. Happy
to add it here if you'd rather have both in one change.

Tests: new TestXaiReservedToolSearchAlias covering the wire alias,
non-xAI backends keeping the canonical name, composition with the
native web_search swap, and the normalize_response round trip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012vLaAmnsdii3Gm9jMDs5gw
2026-08-31 10:09:04 -07:00
Teknium 92574ac508 test(cli): lock buffer-level Shift+letter coverage onto the KeyPress.data fix (#92343)
Follow-up to the salvaged #88097: the same normalization covers the
Shift+letter class reported in #92343 (xterm modifyOtherKeys and both
kitty CSI-u codepoint forms), plus a guard that plain ASCII typing
never triggers the ESC-prefix predicate.
2026-08-31 10:08:48 -07:00
webtecnica 839de43d52 fix(cli): stop raw CSI bytes from Shift+Space leaking into buffer (#88071) 2026-08-31 10:08:48 -07:00
phi hu 081030a7dd fix(config): warn for empty platform toolsets 2026-08-31 10:08:31 -07:00
phihu ccd32a9f0b fix(config): warn when a platform_toolsets entry is an empty list
validate_platform_toolsets() accumulated a single valid_count across every
platform, so the "zero valid toolsets" safety net was suppressed as soon as any
one platform carried a valid toolset. A platform wiped to [] — the active one,
typically cli — therefore produced no warning at all.

resolve_enabled_toolsets() honours that empty list verbatim ([] is a list, so
the platform-default fallback is skipped), leaving the agent with zero tool
schemas. The model then has nothing to call and emits the tool call as
assistant text with finish_reason=stop: no error, no warning, no log entry.
That is the silent-failure mode this module was written to prevent (#38798).

Note the asymmetry this leaves intact: a malformed *string* value is not a list,
so it falls back to the platform default and fails open (#78103); an empty list
fails closed. The fail-closed resolution is deliberate (the explicit_empty_
selection contract in tools_config.py, and #82010 wants it persistable), so this
only adds the missing warning and does not change resolution semantics.

Fixes #89050

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:08:31 -07:00