Today's spill/cache-writer hardening (tools/spill_safety.py,
write_text_exclusive/ensure_spill_dir with O_CREAT|O_EXCL|O_NOFOLLOW)
migrated tools/web_tools.py::_store_full_text() — which writes to the same
cache/web directory with the same content-hash filename scheme — but left
its near-identical sibling, tools/browser_tool.py::_store_full_snapshot(),
on the pre-fix plain open()/write_text() pattern. A pre-planted symlink at
the content-hash path redirected the write onto an arbitrary user-owned
file, same as the sites that commit fixed.
Reproduced live: with a symlink planted at the exact
browser-snapshot-<digest>.txt path (predictable from the snapshot content
hash), the pre-fix write followed the link and overwrote the link's
target with the (secret-redacted but otherwise user/page-controlled)
snapshot content.
Fix mirrors _store_full_text's exact usage: ensure_spill_dir(private=False)
+ write_text_exclusive(private=False, overwrite=True) — not private since
cache/web is bind-mounted into remote backends whose container UID must
read it; overwrite=True because re-snapshotting the same page state
legitimately reuses the same content-hash name (the overwrite path
lstat-unlinks the link itself, never following it to write through).
Added a regression test planting a symlink at the exact digest path and
asserting the link's target is untouched (only the link itself gets
safely replaced by a real file). Mutation-verified: with the fix stashed,
the pre-fix code wrote the snapshot content into the symlink's target
file, reproducing the vulnerability exactly.
The diff-lines error-boundary test mocked './syntax-diff' with a factory
that THREW. vitest hoists the factory and registers its module promise in
the mocker registry; a throwing factory leaves rejected promises there,
and under CI load one escapes as "Vitest caught 1 unhandled error during
the test run" attributed to whichever sibling file the worker is running
(user-message-edit.test.tsx in run 32803716726) — an intermittent js-tests
red on green code.
Rework: the factory now resolves to a component that throws the fetch
error during render — the same way React surfaces a rejected lazy payload
— so no rejected promise ever sits in the registry.
Guard proof (sabotage A/B): with the local syntax-diff ErrorBoundary
removed from diff-lines.tsx, the reworked test still fails (workspace
fallback renders), so the #93479 regression pin is intact. Full ui
project: 596 files / 5729 tests green, no unhandled errors in 3 full-run
repetitions.
Follow-up to the cherry-picked gate-parity fix: the silent 'return 0'
in inject_memory_provider_tools made #81014 undiagnosable — a
configured provider looked half-on with no hint which config key
suppressed its tools. Now an INFO line names the withheld providers
and the gating keys.
The external memory provider's `system_prompt_block()` was injected
unconditionally into the system prompt, while the provider's tools were
gated by `memory_provider_tools_enabled()` via platform_toolsets or
disabled_toolsets. Result: the agent received instructions to call
`mnemosyne_remember`, `mnemosyne_recall`, etc., that did not exist in
its tool surface.
Centralize the gating into `memory_provider_tools_exposed(agent)`, use
it from both `inject_memory_provider_tools` and the system prompt
assembly path, and add regression tests covering:
* memory toolset enabled -> both tools and prompt block exposed,
* memory in disabled_toolsets -> neither exposed,
* memory not in enabled_toolsets and not built-in -> neither exposed,
* the built-in "memory" tool present as an opt-in -> both exposed,
* parity between `inject_memory_provider_tools` and
`memory_provider_tools_exposed`.
The /goal post-turn hook constructs a real SessionDB on an executor
thread at the turn boundary. On a cold or loaded CI runner that
state.db init can exceed send_and_capture's 2s poll window, so the
send lands after the assertion and the test reports the bare
'Expected mock to have been called once. Called 0 times.' (#92130).
Mock _run_post_turn_hooks in the e2e runner — these tests exercise
gateway command dispatch, not goal hooks.
Also scrub TELEGRAM_GROUP_ALLOWED_CHATS / *_GROUP_ALLOWED_USERS / QQ
allowlist env vars in the hermetic conftest: a developer shell with
those set flips _get_unauthorized_dm_behavior to 'ignore' and fails
the pairing e2e test locally.
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.
- hermes-agent skill is now essential: cannot be disabled (config reads
strip it, hermes tools writes drop it), cannot be deleted by
skill_manage, is re-seeded past curator suppression, and is seeded
even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
file+terminal+vision+skills: read_file cannot read images and points
at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
tool isn't loaded.
All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
The #94219 replay was a production no-op: the server returned full
JSON-RPC envelopes from session.events.since while the client's replay
loop dispatches only elements with a top-level 'type' — every replayed
event was silently skipped. Each side's tests validated its own
assumption, so both suites stayed green.
- server: events_since() now returns bare event objects (the frame's
params), the exact shape the live dispatch path consumes; ring stores
params directly; cross-language contract test added on both sides.
- client: live frames racing an in-flight replay are parked and flushed
seq-gated afterward — no double dispatch of deltas, no gap-skip from
a watermark advanced past the replay window.
- restart poisoning: seq counters are in-process, so a backend restart
reset them while clients kept high watermarks (replay forever empty,
truncated=false). New replay_epoch advertised in gateway.ready and
echoed by session.events.since; the client clears watermarks on epoch
change.
- methods_session no longer reaches into event_replay privates
(is_truncated() accessor).
Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client
gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff
clean.
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.
- tools/browser_tool.py: remove _extract_relevant_content and
_get_extraction_model; oversized snapshots always truncate at line
boundaries, store the full tree to cache/web, and append a read_file
pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
pattern as session_search/PR #27590), cli.py defaults + env bridge,
gateway/run.py bridged keys, hermes config display, hermes model picker,
dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
LLM path is gone and stored files are secret-redacted
Follow-up to lkz-de's adapter chunking commit: long Signal messages no
longer truncate on ANY delivery path.
- tools/send_message_tool.py: register Signal's 8000-char limit in
_MAX_LENGTHS (imported from the adapter module so the two paths can't
drift) so hermes send / cron standalone / MCP sends split via the
shared truncate_message() pass instead of signal-cli rejecting them.
Standalone-path idea credited to @5L-hermes01 (#67279).
- tests: regression test proving standalone Signal sends chunk at the
adapter limit with no truncation footer (fails on pre-fix main).
- docs: Long Messages section on the Signal page (en + zh-Hans).
Both fixes verified by sabotage A/B (tests fail with the respective
half reverted to origin/main) and a real-import E2E: 27k-char message
with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles
in-range UTF-16, lossless reassembly.
Drives the reported path rather than the helper: set a non-default
scale, then navigate to routes Chromium holds no zoom record for, which
is what opening a new session looks like to the per-URL store. Keeps the
Cmd/Ctrl+N case alongside it.
Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com>
Desktop is a HashRouter over one file:// document, so every route is a
distinct URL to Chromium's per-URL zoom store. A route the user never
zoomed on has no record at all and resolves to the host default (100%) —
that is every fresh session and every never-visited settings tab.
In-page navigation fires neither did-finish-load nor any window event,
so nothing re-asserted the persisted level. The window dropped to 100%
while the Appearance control kept reading the chosen scale, because the
renderer only learns of zoom changes through 'hermes:zoom:changed',
which never fired. Touching the setting sent a fresh apply, which is
why it appeared to fix itself.
Re-assert the persisted level on main-frame did-navigate-in-page.
Verified on real Electron 40.10.2 / Chromium 144 (win32): a recordless
hash route reports 100% at the event, so the existing drift-guard sees
the drop and re-applies, and still no-ops when the route's record
already matches.
Fixes#48658Fixes#38854Fixes#79863
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
The dom context menu disabled Paste unless a renderer-side
readClipboard() probe reported text when the menu opened. The items
action never consumes that probe: editableCommand("paste") dispatches
webContents.paste() in main - the same Chromium path Ctrl+V takes, which
resolves the system clipboard itself. On Windows the Win32
clipboard.readText() bridge can return empty while that path succeeds,
so Paste stayed grayed out even though pasting would have worked; probe
errors were swallowed the same way (.catch(() => undefined)).
Fail open instead: drop the gate and the now-unused clipboardHasText
fact from the dom menu shape, so opening an editable menu no longer
makes the IPC round-trip at all. Pasting with an empty clipboard is a
harmless no-op, matching Chromiums own menu, which keeps Paste enabled
for editables. The terminal paste item keeps its gate - its action
inserts the readClipboard() text into the PTY directly, so there the
probe and the action share one mechanism and the gate stays honest.
Fixes#91553
The overlay window's playVibeHearts() only fires on a reaction forwarded by
burstVibeHearts, so the single gate covers it — these tests pin that so a
future direct caller shows up as a red test.
Floating affection hearts were always on with no off switch. Message
Reactions in Appearance looks related but only gates message-row
tapbacks. Add a separate Vibe Hearts preference (default on) next to it.
`read_window_below` answers "could not enumerate windows on this system" on
macOS and Windows whatever went wrong, and the three failure paths behind it
discarded their errors — so a report where the HUD could see nothing had no
way to distinguish the module failing to load, the helper failing to spawn,
and the OS answering with nothing. Three different fixes, one sentence.
Enumeration now returns the reason, the tool's error carries it, and the HUD's
game-overlay watch logs it once before it gives up (it retries twice and then
goes quiet forever, which is the other half of why the log said nothing).
Linux keeps its environment-derived advice, which is more actionable than the
raw exception.
The frost is the whole window rectangle and the `[data-hud-glass]` scrim is
what makes it readable, but the two ran on different gates: the scrim is
focus-only, while the caller widened the frost to "recent or held" — i.e. for
the whole of a turn. Thinking with the composer unfocused therefore raised a
bare native material with no scrim over it, which on a light theme is a white
slab under the band's unconditionally white ink.
Put the frost back on the scrim's gate, and re-run it on the window's own
focus changes: clicking away to another app fires no focusout, so the scrim
would go while the frost stayed behind.
The Dockerfile fix in #93757 only helps newly built images, and an
image upgrade (container recreate) resets the permission because
/opt/hermes lives in the image layer. The one stranded case is an old
image whose container was stopped and restarted after the lockout: it
keeps the 0700 install dir and runs code without the guard.
Document the one-line in-place recovery (chmod 0755 /opt/hermes as
root) in the Docker troubleshooting section.
Follow-up to #93757.
The docstring and all four caller comments still said the helper
refuses only / and top-level directories. Since #93757 it also refuses
the entire hermes-agent install tree. Bring the docstring and the
comments at the four credential-write call sites in line with the
actual behavior so future changes are not misled by a stale safety
description.
Follow-up to #93757.
The install-tree exclusion in secure_parent_dir() (#93757) returned
silently. A credential file being written inside the install tree is
exactly the misconfiguration signal that produced the production
lockouts the exclusion guards against, and it also means a previously
hardened path (e.g. a hermes home nested inside a git clone) silently
loses its 0700 parent tightening.
Emit a single warning naming the skipped directory and the install
root so the condition is diagnosable from logs.
Follow-up to #93757.
The install-tree exclusion added in #93757 has a positive test (paths
inside the tree are skipped) but no negative boundary test. The guard
compares path components, so a prefix-named sibling like
/opt/hermes-data must still be chmod'd 0700 — but a rewrite to a
string-prefix match would silently drop that hardening with the suite
staying green.
Add test_install_tree_siblings_still_hardened covering a prefix-named
sibling and an ordinary sibling of the install root. Verified by
mutation: replacing the guard with str(parent).startswith(...) turns
the new test red.
Follow-up to #93757.
GSC page-level data shows configuration (171K impr, pos 5.3), quickstart
(135K, 5.3), providers (123K, 5.3), web-dashboard (68K, 6.5), docker
(37K, 6.3) and the desktop app page (98K, 3.9) all losing ranking
headroom because their title/H1 are bare nouns instead of the terms
people search.
- Frontmatter title + H1 now carry the query terms on all six pages
- Desktop docs page links back to the new marketing /desktop product
page, joining the two official properties Google sees for the query
Done by Hermes Agent (deepseek-v4-pro via nous), Nous Research.
Server: per-session monotonic seq on every routed event frame, bounded
512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs —
session.events.since (replay newer-than-watermark, reports latest_seq +
truncated so clients detect gaps) and session.events.stats (telemetry).
Client: per-session seq watermarks recorded from live frames; after any
successful reconnect a fire-and-forget fetchReplay() drains missed events
through the normal dispatch path (recordSeq ignores non-increasing seqs,
so stale replay can never regress a watermark); focus-triggered reconnect
nudge in use-gateway-boot for the Electron unfocused case where macOS wake
skips visibilitychange.
Replay failures are swallowed by design: lossless resume is an upgrade
over the previous lossy reconnect, never a new failure mode.
OpenRouter's :nitro, :floor, :exacto, and :online suffixes are request-time
routing modifiers valid on any model id — /models lists only the base model.
validate_requested_model() compared the full suffixed id against the listing,
so a valid variant was either rejected outright or fuzzy-auto-corrected to
the base id, silently stripping the user's routing opt-in.
Now, for OpenRouter only, a recognized variant suffix validates the BASE id
against the live listing (and the curated-catalog soft-accept and static-
catalog fallback paths) while preserving the suffixed id for persistence and
API requests — checked BEFORE fuzzy correction. :free/:batch/:thinking
remain direct catalog SKUs and keep exact-match semantics; unknown suffixes
and unknown bases are still rejected.
Reported by JEB (Jakob's Hermes Agent) via Discord.
An MCP call_tool deadline hit left the cua-driver session wedged for
all later computer-use calls. Mark the session suspect on a
concurrent.futures.TimeoutError and tear down + recreate it before the
next non-lifecycle call; healthy sessions are never restarted.
Fail-closed: the timed-out action may still have taken effect on the
remote screen, so it is never silently replayed — the error result
carries structuredContent.code=timeout_outcome_unknown with
next_step=fresh_state.
Informed by #74877 by BlackishGreen33.
Co-authored-by: BlackishGreen33 <s5460703@gmail.com>
Replay single-question and batch snapshots from session.activate/resume,
including locked answers, and extract the helper so the session-actions
god-file is not the only owner of that wire shape.
Co-authored-by: ClintonEmok <54935030+ClintonEmok@users.noreply.github.com>
Co-authored-by: frendo <frendo.wu@gmail.com>
A hydrated Ask/clarify row stays complete after session or bot switch, so
the live card never mounts. Re-arm the existing transcript row and keep
the provider tool id instead of appending a duplicate at the tail.
Co-authored-by: frendo <frendo.wu@gmail.com>
Clarify remounted as a tool row once session.info flipped running=false, thinking previews collapsed their body, and the duration line grew the footer.
LocalEnvironment._kill_process kills the process GROUP (SIGTERM ->
1s wait -> SIGKILL -> 2s wait), but a descendant that called setsid
escapes the group and survives — the #71148 orphan class, terminal
flavor (issue #84967's local sibling).
Fix: snapshot the descendant set via psutil BEFORE the first SIGTERM
(children reparent to init once the wrapper dies, so a later parent
walk finds nothing — same snapshot-before-signal design as
agent/deadline.py kill_process_tree), then after the existing group
escalation completes, SIGKILL any snapshotted survivor whose pgid is
no longer the (now-dead) group. The TERM->KILL grace window for
in-group members is preserved (interrupts use this path too), the
Windows branch is untouched, and the snapshot is fully guarded — a
broken psutil never breaks the kill path (unit-tested).
Tests: live_system_guard_bypass acceptance test spawning a setsid
grandchild and forcing the timeout path (RED on unmodified file,
GREEN after), plus a psutil-failure unit test.
Docker design note (#84967 open question 1, condensed; full note at
/tmp/4b-docker-design-note.md): the docker backend inherits
base.py:1378 _kill_process, which only proc.kill()s the HOST-side
`docker exec` client — the in-container tree (child of containerd-
shim, not the client) survives every timeout entirely. Option A,
`docker exec <cid> kill -- -<pgid>` with TERM->KILL escalation using
a PGID captured at command start, is surgical and preserves container
state but needs a live container + shell and still misses in-container
setsid escapees. Option B, container restart, is absolute (PID-
namespace teardown kills everything) but destroys all in-container
state mid-session and punishes every other consumer of the shared
persistent container. Recommendation: Option A as a best-effort
_kill_process override (degrade to today's behavior on failure);
reserve restart for the existing container-gone recovery path.
Per-site decisions:
1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
via a function-local import; keeps the swallow-everything fail-open
contract and the (proc) -> None signature (agent/shell_hooks.py imports
it by name; _kill_git_process_tree alias preserved). The old body is kept
verbatim as _legacy_kill_process_tree and used as fallback when the
delegation import/call fails. A final proc.kill() is retained on the
happy path so Popen bookkeeping sees the exit (matches old behavior).
2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
(delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
body sent SIGTERM then SIGKILL with zero grace between them; the shared
primitive sends SIGKILL only. With no grace period the observable effect
is identical, and the psutil descendant sweep now also reaches
agent-browser's setsid'd daemon grandchild, which killpg alone missed.
tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
the legacy fallback (its assertions describe the fallback's internals).
3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
kill when escalate=True) — expressed as two delegated calls:
kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
body terminated children before the parent; the shared primitive
signals the group atomically (child is a session leader via
start_new_session=True) plus an identity-aware descendant sweep —
strictly wider coverage, same signals.
4. gateway/status.py: KEPT BOTH SITES.
- terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
incompatible with the shared primitive — it must RAISE OSError with
taskkill's stderr on non-zero exit (callers branch on that), falls back
to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
fail-soft primitive would invert the error contract.
- reap_gateway_children (~l2029): NOT migrated. It operates on a
pre-snapshotted child list from a parent that is already dead
(psutil.Process(pid) on the parent would fail), and every signal is
wrapped in identity/ownership checks the primitive lacks: is_running()
identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
The coupling is the feature; migrating would delete the safety logic.
5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
Dev tooling that intentionally kills by CAPTURED pgid because the direct
child is usually already reaped (psutil/pid-based primitive cannot find
it), and it avoids the psutil import on the test-runner hot path. Its
docstring already documents why psutil is the wrong tool there.
New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).