Commit Graph

35183 Commits

Author SHA1 Message Date
teknium1 0ea0c53e89 fix: record a known exit status when the reader's wait() raises
The early-EOF reaper made _finish_reader return without publishing when
wait() raises, so the session stays tracked for later reconciliation.
That is right for the pipe path (_reconcile_local_exit can still reap via
session.process), but PTY sessions have no session.process: when
ptyprocess.wait raises (waitpid ECHILD after isalive() already reaped the
child) the exitstatus is known, yet poll() reported "running" forever.

Only leave the session tracked when exit_code() is still None; otherwise
record the known status and finish as before.

Review finding: PTY session whose pty.wait raises stays in _running forever (fail-open regression vs main)
2026-09-15 04:23:32 -07:00
teknium1 dc4a2aeb10 test(process): trim early-EOF reaper test to its two invariants
Drop the wait-timeout call-shape assertion; the invariants are that the
session records the real exit code and that a failed reap does not publish a
false completion.
2026-09-15 04:23:32 -07:00
fangliquan 1bfbeff4e5 fix(process): reap children after early stdout EOF 2026-09-15 04:23:32 -07:00
teknium1 8bc5894a3a docs(skills): ip-as-logo follows the modern section order; add skill tests
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.

Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
2026-09-15 04:20:15 -07:00
teknium1 febea3e28d docs(skills): pin ip-as-logo upstream attribution to author, repo and commit
The ported skill carried the upstream MIT text but the LICENSE file did not
say where it came from, and the frontmatter only pointed at the repo via a
non-standard `homepage:` key. Reviewers asked for proper attribution.

- LICENSE: header naming the upstream repo, the pinned upstream commit
  (b1bf517c54a4…) and the copyright holder (s1dashu) above the verbatim MIT text.
- SKILL.md: `metadata.hermes.upstream: <repo> (pinned b1bf517c)` — the same
  shape mono-color and pr-lens use — and the adaptation-notes blockquote now
  links the upstream repo and commit so the generated docs page links the source.
- Regenerated website/docs/user-guide/skills/optional/creative/creative-ip-as-logo.md
  with website/scripts/generate-skill-docs.py (scoped to this skill).
2026-09-15 04:20:15 -07:00
Teknium 398279fd6a feat(skills): add ip-as-logo optional skill (minimal cute IP mascot marks)
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.

Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
  (square aspect, main-prompt constraints mode — no negative_prompt
  parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
  own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
  proposal-round skip for pre-authorized batches, dimensions-reporting
  rule when the backend returns only a URL, limbless-subject note

Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.

Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).

Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
2026-09-15 04:20:15 -07:00
teknium1 1432a3b47a catalog: add hermes-memory-wiki (Memory Wiki dashboard plugin)
The Memory Wiki from #89940 (salvage of #31244 by @Araja119) ships as a
standalone plugin in NousResearch/hermes-memory-wiki instead of landing in
core; this entry lists it in the catalog at the repo's reviewed commit.

The pin is the repo's first commit (younger than the two-week maturity
window) and needs the repo owner's explicit waiver to merge.
2026-09-15 04:19:27 -07:00
teknium1 b531023622 fix(gateway): drop the redundant whole-block lease; one invariant test per atom; document the watchdog env vars
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.

Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.

Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
2026-09-15 04:19:19 -07:00
Ayush Nangia bcab2dd6eb fix(gateway): persist startup watchdog stacks before stderr
A detached or service-managed gateway can have a blocked stderr. The watchdog exit escort then hard-exits after ten seconds before the file-based traceback is reached, leaving only the metadata record. Write the durable dump first and pin the ordering with a regression test.
2026-09-15 04:19:19 -07:00
Kevin Rajan cef69276ae fix(gateway): hold startup-watchdog progress leases across state.db auto-maintenance
Construction-time maybe_auto_archive / maybe_auto_prune_and_vacuum ran
synchronously with no report_startup_progress lease. A multi-minute
VACUUM of a large state.db accrues near-zero CPU, so the startup
watchdog misread it as a parked deadlock and killed the attempt with
exit 75, live-locking gateway restarts. Renew the lease per long step
(prune, orphan sweep, VACUUM, archive) since leases clamp at 900 s,
plus one lease around the gateway maintenance block.

Fixes #111092
2026-09-15 04:19:19 -07:00
teknium1 5358780264 fix: keep the carried pre-admission input across the waited lease reload
carry_unadmitted_user_message appends the interrupted turn's user row to the
early result's in-memory history only; it is never persisted because that turn
never owned the lease. When the follow-up turn also has to wait for the lease
(the common case: the other process that made the first turn wait is usually
still busy), admit_durable_turn_lease reloads conversation_history from the DB
after admission and replaced the caller's list wholesale, so the tagged row was
dropped from both the model input and state.db. Re-append the tagged rows that
have no _row_id after the reload so the follow-up turn sees and flushes them.

Review finding: waited-reload branch of admit_durable_turn_lease discarded the
_persist_after_admission_interrupt row carried from the aborted turn.
2026-09-15 04:18:34 -07:00
teknium1 fae3030fa6 refactor: carry the unadmitted user message from the lease sibling, not the facade
Move the pre-admission carry-forward out of run_conversation (facade) into
agent/turn_facade_lease.py::carry_unadmitted_user_message next to the early
result it repairs, and drop the extra turn-start flush in build_turn_context:
the follow-up turn's normal turn-start persist already writes the marked row
because _db_flush_collect no longer stamps it as durable (live probe: exactly
one A row in state.db after two flushes). Trim to two invariant tests
(carry-forward with metadata; flushed exactly once); the hard-stop negative is
covered by the E2E probe in the PR body.
2026-09-15 04:18:34 -07:00
Ayush Nangia 24ae31f7e7 fix(gateway): retain pre-admission interrupted input 2026-09-15 04:18:34 -07:00
teknium1 9efa50071a fix: match find/read-tool dynamic words only as unquoted command arguments
The dynamic-shell-word rules fired on any `-del*`/`-exec*` substring after a
`find` anywhere in the segment, so quoted predicate arguments the shell never
expands were flagged: `find . -name 'log-del*'`, `find . -name 'pre-exec*.sh'`,
`find src -path '*-exec[0-9]*'`, and `echo find . -{delete,print}`.

Three changes close that:
- `find` must be the command word (_CMDPOS anchored, same as mkfs/rm/dd) and
  the dynamic word must start a whitespace-delimited token (`(?<!\S)`), for
  both the find rule and the rg/sort/ag/man program-option rule.
- Both rules now scan the quote-masked variant (_QUOTE_MASKED_DANGEROUS_
  DESCRIPTIONS, same _mask_quoted_prose used by the positionless hardline
  rules) so glob characters inside quotes are data, not expansion.
- _iter_shell_command_starts no longer treats the `{` inside a brace-expansion
  word (`-{delete,print}`) as a brace-group opener; it split the word across a
  marked start so `echo x; find . -{delete,print}` matched nothing. A brace
  group opener is `{` as its own word (after whitespace or a separator).

The quoted-name cases join the inert parametrized test; the separator case
joins the dangerous one. Still two parametrized functions.
2026-09-15 04:17:42 -07:00
Teknium 06a3a98751 fix: gate dynamic shell words in approval checks
Port from openai/codex#39159: require approval when shell expansion could synthesize destructive find flags or program-executing read-tool options.
2026-09-15 04:17:42 -07:00
teknium1 bc3df8a4d5 fix: NT-namespace guard fires before every sibling resolve (checkpoint, ACP bridge, @file:)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.

The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
2026-09-15 04:17:31 -07:00
teknium1 07e6461dfe fix: guard search_tool's root and trim the NT-namespace tests to two invariants
search_tool resolved its root via _resolve_path_for_task BEFORE the
NT-namespace check saw it (review finding): on Windows the resolve is the
SMB-auth trigger, on POSIX the task-base join hides the prefix from the
resolved-path denylist. The raw-string guard now runs first there too,
and the guard rides the existing top-level agent.file_safety import
instead of two function-local imports.

The tool-layer chokepoint test now covers all four entries and proves
none of them touched Path/_resolve_path_for_task/realpath before the
refusal; the 12 form-matrix tests collapse into one blocked/allowed
invariant over both classifiers.
2026-09-15 04:17:31 -07:00
Teknium faf71eb4c1 Inspired by Claude Code: file tools reject Windows NT-namespace paths (NTLM leak hardening)
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:

- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
  raw-string check (never resolves — resolving IS the leak trigger).
  Wired as the first check in get_read_block_error() and the write
  denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
  _check_sensitive_path (covers write_file_tool + patch_tool), before
  the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
  local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
  forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
2026-09-15 04:17:31 -07:00
teknium1 ce04a6f189 docs: match room picture and member-session wording to the shipped UI
The roster row no longer fans member faces (GroupRow renders the room
image or a single group glyph since 5afa487e9), so describe the picture
as replacing the default glyph. Member sessions are titled by roomId,
not by display name (group-turns.ts), so drop the `Group: <name>`
literal and say "room session" everywhere the doc mentioned it.
2026-09-15 04:15:28 -07:00
Teknium 4f6e8c7345 docs: Bot Mode group chats — editable name and room picture
Documents PR #89371: room picture at creation (upload/generate),
Group settings dialog (rename + picture) after creation, rename
keeps history/sessions and rejects collisions.
2026-09-15 04:15:28 -07:00
teknium1 9c689bee1a fix(bot-mode): an empty member seat surfaces an error instead of swallowing the group send
Ported from the pre-TSX PR onto the current module layout. main's
group-chat-view.tsx already restores the composer draft when
sendToGroupChat returns null (feat(desktop): retain Bot group drafts by
room, b42d8279ed), so the draft-loss half of the original fix is
FIXED_ON_MAIN and not re-applied here. What remained: sendToGroupChat in
group-rounds.ts still folded "no text" and "no members" into one silent
`return null`, so a fully typed message into a room whose roster had not
hydrated (or a legacy room record without member descriptors) was
rejected with no thread, no log entry and no error.

The guards are split: empty content stays silent, an empty member seat
raises host.notify with a Bot Mode i18n string (en/ja/zh/zh-hant). The
wording no longer promises a retry will help (review: a legacy room with
no member descriptors never recovers by retrying) and points at the two
real remedies.

The original source-regex .mjs tests are dropped per review; the
contract is pinned by a behavioural vitest in group-rounds.test.ts that
drives sendToGroupChat through the scripted room harness: empty members
→ null + one error toast + no log entry; blank text → null, no toast.
2026-09-15 04:14:09 -07:00
teknium1 aaf38bb2fb test: service stubs carry the profile the pause path now reads
The fix-pass made _pause_windows_gateways_for_update read s.profile from
every discovered service gateway; two quarantine tests build services with
SimpleNamespace stubs that lacked the field the real WindowsGatewayService
always has.
2026-09-15 04:13:44 -07:00
teknium1 27df47af5e fix: count service-supervised gateways as running for per-profile cold-start
The per-profile cold-start probe took its running set from the socket-paused
ordinary gateways only. A profile whose gateway is alive under an SCM service
is skipped by the socket pause, so it looked "not running"; with an empty
current-PID list its live start attestation read as dead and the profile
landed in cold_start_profiles. On resume the service was restarted AND a
second, unsupervised gateway was spawned for the same profile.

Build the running set from every profile that had ANY live gateway at
discovery time: paused profiles, profile-mapped processes, and service
gateways' profiles.

Review finding: service-supervised running profile was cold-started beside
its restarted SCM service (double gateway).
2026-09-15 04:13:44 -07:00
teknium1 b19e5786e2 test(update-windows): spawn stubs accept the per-profile home argument
_spawn_detached now forwards home= to _build_gateway_argv so the post-update
cold-start can launch another profile's gateway; the two windows_only spawn
tests stubbed _build_gateway_argv with a zero-arg lambda and raised TypeError
on the Windows runner.
2026-09-15 04:13:44 -07:00
Hermes Agent 8fbae2813b fix(update-windows): per-profile cold-start runs after the relaunch and fails loud
- Run `_cold_start_attested_profiles` AFTER the paused profiles are relaunched,
  so a sibling that fails to cold-start can never keep the profiles that were
  running from coming back.
- A profile that stays down raises like the active-profile cold-start does;
  the merged outcome then marks the update incomplete instead of printing
  success with a gateway still offline.
- Trim the salvaged tests to the two invariants (dead-attested default beside
  a live beta is cold-started under its own home; every dead-attested profile
  is cold-started when nothing runs, active first). The "no attestation →
  token unchanged" case is the pre-existing behaviour already covered by
  test_pause_skips_cold_start_plan_when_desktop_owns_lifecycle.
2026-09-15 04:13:44 -07:00
kshitijk4poor c728e6583e fix(update-windows): per-profile cold-start probes its own home and runs after the active spawn
Review follow-ups on the per-profile obligation:
- Order: the active profile's cold-start guard is fleet-wide (any live
  gateway ⇒ done), so a sibling spawned first would have left the active
  profile down. Per-profile spawns now run after it.
- A token is built even when the active plan owes nothing (clean exit,
  autostart not installed), so a dead-attested sibling still rides on it.
- Readiness for a per-profile spawn is probed in THAT home's identity files
  (`_live_gateway_pids(home=)` → `get_running_pid(home/"gateway.pid")`), not
  the fleet: a still-running sibling no longer vouches for a dead spawn. The
  wait/confirm helpers share one probe function.
- The new PID is attested in the profile home (`_write_start_attestation(...,
  home=)`) so a death after the CLI exits reaches the next update, and an
  already-live profile is not spawned twice.
2026-09-15 04:13:44 -07:00
kshitijk4poor c8686224d3 fix(update-windows): cold-start every dead-but-attested profile, not only when nothing runs
`_pause_windows_gateways_for_update` built the attested cold-start plan only when the
ALL-profile running PID list was empty, so a default gateway that died after a ✓ beside a
still-running `beta` never received a cold-start obligation and stayed down after the
update (#110959, fifth review thread on #110020).

- gateway_windows: thread `home` through `_start_attestation_path` → `_read_start_attestation`
  → `_attested_pid_exited_cleanly`/`_attested_dead` → `attested_death_generation(pids, home=)`
  and `_consume_start_attestation(gen, home=)`; `_spawn_detached(home=)` builds the argv/env
  for that profile home (`--profile` derived by `_launcher_settings`). Defaults unchanged.
- update_cmd_windows: `_record_attested_cold_start_profiles` evaluates every
  `profiles_to_serve(multiplex=True)` profile that is not running (Desktop-owned installs only,
  active profile left to the existing plan so nothing is spawned twice) and records
  `token["cold_start_profiles"] = {name: generation}` — a sibling key, so `profiles[name]`
  stays an int for relaunch/verify. `_cold_start_attested_profiles` runs on resume before the
  ordinary relaunch, spawns under the profile home, waits for readiness, consumes exactly that
  generation; one profile's failure never aborts the others.
- tests: default dead-attested + beta running → token carries the obligation, resume spawns
  under the default home and consumes its marker while beta is relaunched; no attested
  profile → token and resume unchanged.
2026-09-15 04:13:44 -07:00
teknium1 ef44c1b73f fix: restore TestBridgeDispatch, trim #39797 tests, drop dead execution_guidance_text param
Why: the rebase conflict resolution in tests/tools/test_model_tools.py
deleted the unrelated TestBridgeDispatch class (3 tests from 73163e3);
it is restored verbatim from main with TestBrowserRetrievalHints after it.

The fix had five tests for one invariant: the OPENAI_MODEL_EXECUTION_GUIDANCE
check duplicated test_phantom_tool_references, and the two static-schema
"toolset-neutral" checks are now folded into test_silent_without_web_tools,
which runs _apply_dynamic_schemas over the real browser_navigate/browser_cdp
schemas so the rendered descriptions are what is asserted.

execution_guidance_text() no longer takes valid_tool_names: the guidance is
toolset-neutral, so the parameter was ignored; the single caller in
agent/system_prompt.py and its test are updated.
2026-09-15 04:13:13 -07:00
teknium1 ac63d0eea5 fix(agent): execution guidance and browser hints drop the web_search stripper; tests assert the invariant
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (3733e4aff5) matched
nothing and were dead; the function now returns the neutral text for
every toolset and its phantom-tool test asserts "no web tool named"
instead of the removed sentence. model_tools ports the PR's hint layer
into main's _DYNAMIC_SCHEMA_REWRITERS table (browser_navigate +
browser_cdp) rather than a second pass after it.

Tests: the two browser_cdp registry tests were re-added by the PR but
main pruned them in 39975613b13b4; replaced with one schema-neutrality
invariant. Exact-wording assertions ("lightweight retrieval tool",
"appropriate permitted retrieval/search tool") were change detectors and
are dropped. tools-reference.md row updated to the new schema text.
2026-09-15 04:13:13 -07:00
Kevin Yin ccf380f634 fix(agent): respect permitted web retrieval guidance 2026-09-15 04:13:13 -07:00
teknium1 6dc6c92ea2 fix: release the profile MCP stderr handle before rename too
rename_profile moves the profile directory while this process may still
hold the cached per-profile mcp-stderr.log handle (left behind by a
completed probe or a running server). On Windows a directory containing
an open file cannot be renamed, the same WinError class delete_profile
now avoids. On other platforms the stale handle stayed cached under the
old home key, so a later probe on the renamed profile opened a second
handle and a new profile re-created under the old name wrote its MCP
stderr into the renamed profile's log. Release the scoped handle next to
the multiplexer unroute, mirroring delete_profile.

Review finding: rename_profile missed the sibling surface of the
delete_profile handle release.
2026-09-15 04:12:41 -07:00
teknium1 325236d801 chore: map tarkil@gmail.com to @tarkilhk for contributor attribution 2026-09-15 04:12:41 -07:00
tarkilhk 68dea35e0e fix(mcp): release profile stderr handles before deletion 2026-09-15 04:12:41 -07:00
teknium1 2500f4ee56 test(cli): sessions open-failure test creates the store so the read-only empty path does not short-circuit 2026-09-15 04:12:13 -07:00
teknium1 7c6b29e1e4 fix(gateway): 401 hint cites hermes auth add, not the removed hermes login 2026-09-15 04:12:13 -07:00
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1 7fcade0100 chore: map octopustank@foxmail.com to @Octopustank for contributor attribution 2026-09-15 04:11:55 -07:00
Hermes Agent d3c5539056 refactor(linux-desktop-entry): collapse the fall-through into one guard
The salvaged fix replaced the early return with a `pass` + `elif`;
fold it into a single `if primary and rerouted is not None` so the
probe is reached by construction and the WHY lives in one comment.
2026-09-15 04:11:55 -07:00
Octopustank 93116dc820 linux-desktop-entry: note the None-path equivalence inline
resolve_hermes_bin's rerun hides argv[0], so a non-None reroute could
only come from PATH — which the first call would already have returned.
primary is None therefore implies rerouted is None, and the fall-through
is equivalent to return rerouted or primary; only the durable-wrapper
probe below can still find anything on that branch. (KeyArgo's review
asked why the branch was merged rather than dropped.)
2026-09-15 04:11:55 -07:00
Octopustank a2a7cf922a fix(linux): converge the desktop entry Exec on the durable wrapper
_resolve_hermes_bin_for_desktop_entry returned the resolver's None
outright when neither argv[0] nor PATH yielded a launcher (a cold
relaunch under `python -m` with a stripped PATH), skipping the
known-wrapper probe its own docstring describes. The persisted
`Exec=` then flipped to the bare module form, while a DE- or
terminal-launched context renders the wrapper form.

Every flip rewrites ~/.local/share/applications/hermes.desktop on the
next launch. gnome-shell 50.x removes a ShellApp from id_to_app on any
.desktop content change without checking its state; if the app was
still STARTING, the last strong reference drops and a later GC-triggered
dispose trips shell-app.c's `state == STOPPED` assertion — taking the
whole Wayland session down (observed crash, 50.4-1.fc44).

Probe the installer's known wrapper locations on the None path too, so
the entry converges on the durable wrapper wherever one exists and only
falls back to the module form when none does. A missing entry being
(re)created can still change the file regardless — this removes the
gratuitous rewrites, not the rescan trigger itself.
2026-09-15 04:11:55 -07:00
teknium1 f876ba60ff fix(logging): unavailable-log notice is re-armed by a successful write, not by open()
The "named once" guarantee only held when open() itself failed. In the reported
case open() succeeds and write/seek/flush raise EIO, so every record went
handleError -> stream=None -> next emit reopens -> _open() reset
_unavailable_reported -> the path line was printed once per record (25 for 25).
Reset the flag in emit() only after a record actually reached the stream; that
is the moment the destination has recovered.

Adds the open-succeeds/I-O-fails case to the test file: red on the previous
head, green here.

Review finding: EIO on write after a successful reopen printed the path per record.
2026-09-15 04:11:11 -07:00
teknium1 37e4cd5faa fix(updater): a probe that cannot be spawned stays advisory instead of failing the update
bounded_probe_run collapses spawn failure and timeout into one None, so a venv
interpreter that exists but cannot be executed (PermissionError, ENOEXEC, fork
failure) was reported as "timed out before reporting import health" -- a bogus
warning on the git path and sys.exit(1) on the ZIP path -- and the
`except OSError: return {}` branch beneath it was dead. Let the Popen error
propagate (opt-in flag, default unchanged for the other probes) so "we could not
run our own probe" says nothing about the checkout and no longer blocks the
update, while a spawned child that hangs is still a verdict.

Restores the non-fatal test the earlier rewrite dropped, now driving the real
bounded_probe_run/Popen: red on the previous head, green here.

Review finding: spawn failure of the import probe reported as a fatal timeout.
2026-09-15 04:11:11 -07:00
teknium1 7020a0081a docs(secrets): hermes update and its probes never resolve external sources 2026-09-15 04:11:11 -07:00
teknium1 b847c1ea5e fix(logging): name an unavailable log file once instead of silently dropping records
Builds on the salvaged EIO suppression: the reporter asked that logging
"degrade gracefully and identify the affected path". Print one stderr line
naming the file and errno when the stream first fails, reset the flag when
`_open()` succeeds again so a later failure is reported anew. The contributor's
test is replaced by one invariant: five records through a stream raising EIO
produce zero tracebacks and exactly one path mention, and the next emit
recovers into the real file.
2026-09-15 04:11:11 -07:00
teknium1 da1fb702c3 fix(updater): import probe children never resolve external secret sources
The critical-module import probe (`_critical_module_import_failures`) imports
`run_agent`, whose module-level `load_hermes_dotenv()` resolves every enabled
external secret source. The main-process skip only lived in `hermes_cli.main`'s
own dotenv call (`load_external_secrets=sys.argv[1:2] != ["update"]`), so the
probe child ran op/bws/command helpers with a 120s per-source budget inside the
120s probe and a healthy install reported "critical-module probe still fails to
import after updating: timed out before reporting import health" (#110823).

Move the argv check into `_early_recovery._should_skip_external_secret_sources`,
which every dotenv load already consults, and stamp `sys.argv = ['hermes',
'update']` into the probe so its imports inherit the updater contract.
Invariant test spawns the real probe against a home whose configured helper
touches a marker: red on main, green here.
2026-09-15 04:11:11 -07:00
KoNit-K e74c29de55 fix(updater): bound import probe teardown 2026-09-15 04:11:11 -07:00
teknium1 8e0b1a2e47 fix(update): keep the config-migration purge and reload inside the fail-open guard
_purge_stale_hermes_modules evicts hermes_cli.* but not root modules
(hermes_constants, utils, toolsets, ...), so the post-purge
`from hermes_cli.config import ...` re-executes the NEW config.py against
OLD cached root modules. When a pull adds a root-module symbol that
config.py imports at module level, that raises ImportError. The import
sat outside the step's try/except, so the error escaped
_run_post_update_maintenance and aborted the fleet restart -- where
origin/main printed the "Could not check config version / run hermes
config migrate" fallback and continued. Move the purge, reload and
import inside the existing try so the step stays fail-open.

Test: a cached hermes_constants lacking get_process_hermes_home (the
symbol hermes_cli/config.py imports at module level) must print the
fallback and return instead of raising. Red before, green after.

Review finding: fail-open -> fail-closed flip; post-purge hermes_cli.config import outside the try let ImportError abort post-update maintenance.
2026-09-15 04:10:27 -07:00
teknium1 24a623e557 fix(update): purge stale Hermes modules before post-pull config migrations
The pinned-list reload (tools_config) fixes the reported symbol; any
future symbol a pull adds to any module a migration imports at call time
(agent.skill_utils, hermes_cli.toolset_scope, ...) would fail the same
way. Evict every cached Hermes module at the migration entry point with
_purge_stale_hermes_modules — the class fix the fleet-restart phase
already relies on — so the migration graph is rebuilt from the new tree.

Tests: the reporter's exact shape (cached tools_config lacking
_configurable_keys, on-disk v44 config with an explicit platform
toolset list) must still migrate v44 -> v45 through the updater's own
entry point. Existing entry-point tests stub the purge like the rest of
the update suite does.
2026-09-15 04:10:27 -07:00
liuhao1024 91a3ab4c85 fix(update): reload the cached tools_config before post-pull migrations
The updater process is the pre-pull process: hermes_cli.tools_config cached in sys.modules lacks symbols the pull added, so a migration that imports them at call time fails with ImportError and the config silently stays at the old version (#111271).
2026-09-15 04:10:27 -07:00
teknium1 e43f2f6816 fix(cli): retire the 0.0 monotonic sentinel in the remaining input-mode throttles
Review follow-up on #91651: _recover_terminal_input_modes and the termios
drift check used the same `now - 0.0 < interval` idiom as the repaint
throttles. Their windows (0.5s / 1.0s) are unreachable in practice, but
converting them to the None sentinel retires the bug class instead of the
instance, so nobody "simplifies" a None back to 0.0 later. Also adds the
missing regression test for _invalidate, the throttle the PR title is about.
2026-09-15 04:08:53 -07:00