The two halves of tips were behind one switch, which meant the app
volunteering commentary at idle shipped on by default. Split them along
the line that matters: the rotation talks unprompted, so it now waits to
be asked for, while an agent tip stays ungated like the tour it mirrors
— Hermes raises one mid-conversation, in answer to something the user
said.
Drops the tool's config gate along with the config key it read. The
renderer mirrored that key with config.set, which has no branch for it
and answered "unknown config key" into a swallowed catch, so the opt-out
never reached the backend in the first place.
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.
Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.
Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
_stdio_children_dead() returned True ('all children dead') when a child
process was ALIVE — the liveness predicate was inverted:
for pid in pids:
if not psutil.pid_exists(pid):
continue # dead — skip
return True # BUG: an ALIVE child reported as 'all dead'
Consequence: every tools/call on a stdio MCP server with a healthy
subprocess failed instantly (~0.01-0.3s) with
'MCP stdio subprocess has exited; failing the call fast' (#81995
fast-fail path), while hermes mcp test kept passing (it never reaches
tools/call). Servers appeared dead regardless of restarts.
Fix: return False as soon as one tracked child is alive; True only when
every child has exited:
for pid in pids:
if psutil.pid_exists(pid):
return False # at least one child alive
return True
Also: on a genuinely-dead stdio (session object still present), signal
reconnect instead of a bare fast-fail TimeoutError, so the manager
respawns the subprocess instead of stranding the call slot.
Verified: alive child -> False, all exited -> True, no tracked pids ->
False (unknown, don't fail fast).
When send_message is invoked from the agent's worker thread (a different
event loop than the gateway's), awaiting the WeCom adapter directly can hang
because the adapter enqueues onto the gateway loop. Dispatch via
run_coroutine_threadsafe onto the gateway loop when the caller loop differs,
with caller-cancellation shielded so an already-enqueued send is not cancelled
mid-flight (which would otherwise cause a false-failure retry -> duplicate).
Recognizes WeCom native chat IDs as explicit send targets and whitelists WeCom
for media delivery. Part of the async queue design this branch introduces.
Builds on e11187208f (salvage of #72206 by @luyifan, authorship
preserved) which added the post-timeout session reset. This commit
completes the Phase 3b contract:
- suspect flag: a command timeout marks the session key suspect;
the NEXT _get_session_info for that key health-checks and recycles
the session instead of handing back the poisoned handle (#72205)
- flag cleared on fresh-session store so a healthy new session is
never spuriously recycled; cross-key leakage fixed
- wedged-vs-alive rule (#68139): after a timeout, if the daemon's
socket still answers, recycle the session only; if it is
unresponsive (or no socket exists to probe — conservative), tree-
kill via agent.deadline.kill_process_tree and evict
- negative probe: successful commands never mark or recycle
Tests: tests/tools/test_browser_suspect_recycle.py (20 tests:
mark-once, recycle-then-succeed, success-never-recycles, tree-kill
invoked on the wedged path with pid assertion, flag lifecycle).
Co-authored-by: luyifan <al3060388206@gmail.com>
The smart-approval guardian (`_smart_approve`) gates every flagged
terminal command with a synchronous auxiliary LLM call, but it never
passes `timeout=` and logs nothing on the normal path. In production a
stalled provider response silently froze the agent turn for 62 minutes
with zero log output; the gateway kill-switch eventually fired, and only
an unrelated error surfaced afterwards (#82846; watchdog-style fix in
#72500). The call was invisible by design — nothing logs at the hang
point.
Changes in tools/approval.py:
- Resolve the same configured timeout the client would use internally
(`auxiliary.approval.timeout` via `_get_task_timeout("approval")`) and
pass it explicitly to `call_llm`, so the deadline cannot be lost if the
internal default resolution changes or is misconfigured.
- Log the assessment call and its duration (DEBUG), and promote the
failure branch from DEBUG to WARNING with elapsed time + exception
class, so a wedged guardian call is visible in the logs instead of
silent.
- Failure still returns "escalate" (fail open to the human/pattern
gate) — behavior unchanged, observability only.
Complements #72500 (watchdog hard ceiling) rather than duplicating it:
explicit timeout is the root-cause hardening, logging closes the
silence gap; the watchdog remains the safety net if the SDK-level
timeout itself is defeated.
Tests: explicit timeout forwarded to call_llm (revert-fails), failure
logs WARNING + escalates. 49 approval-adjacent tests pass; one unrelated
test_approval.py failure is pre-existing (fails on clean main too).
Apply-ready delta distilled by @andrexibiza: deterministic watcher-consumer
tests (watcher times out while a child is alive, resolves when all are dead),
psutil-unavailable fail-open pin, and probe-failure fail-open handling in
_stdio_children_dead (unknown is never proof that every child exited).
Local: 8 passed on tests/tools/test_mcp_stdio_children_dead.py
_stdio_children_dead returned True ('all children dead') on the first LIVE
pid — the intended False was dead code right below it. Every spawn path
that captures child PIDs (observed in hermes -z oneshots) then failed the
#81995 pre-call fast-fail with 'TimeoutError: MCP stdio subprocess ... has
exited' on every tools/call while the subprocess was demonstrably alive.
Long-lived gateway/dashboard sessions were unaffected only when
_stdio_child_pids was empty (the not-pids short-circuit).
Return False on the first live pid and drop the unreachable line.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
Addresses the round-3 findings from @Adolanium + @kshitijk4poor on #95620:
1. Overlay-before-reuse race (blocker): _real_profile_cdp ran snapshot_real_profile
BEFORE the session-reuse check, so a cold resolve that ends in reuse rewrote
Cookies/Login Data under a live Chromium holding the user-data-dir open (torn
DBs, locked txns, phantom logouts). Now: resolve copy dir as a PATH, probe
reuse first, return early on a hit; snapshot/overlay only on the relaunch
path when no live browser owns the dir.
2. Torn first copy poisoned freshness forever: freshness keyed on isdir(Default),
so a half-written copy (disk full / Ctrl+C) was treated as populated and only
ever got auth overlays. Now gated on a .hermes-snapshot-complete marker
written only after a full copy succeeds; a torn copy is rebuilt from scratch.
3. Consent revocation left copied credentials on disk: turning use_real_profile
off now deletes ~/.hermes/browser-profile/ on next browser use
(cleanup_real_profile_snapshots), so cookies/logins don't outlive consent.
4. Stale non-active profile copies: only the ACTIVE profile (last_used) is copied
into the copy's Default now — other Chrome profiles are never snapshotted
(smaller copy, no stale credential dirs lingering).
5. Docs/config/desktop wording aligned to actual behavior (active-profile only,
refresh on fresh session, consent-off cleanup).
Tests: overlay-skipped-on-reuse + overlay-runs-on-relaunch, done-marker gating +
torn-copy rebuild, active-only copy, consent-off cleanup (removes store +
idempotent + triggered from _real_profile_cdp). 206 browser + 222
backup/file_safety pass. Live: reuse skips re-snapshot; direct launch on the
active-only copy loads the real signed-in Gmail inbox.
Real-profile browsing routes all in-page browser work through the Browser Use
CLI (browser_exec), which obsoletes the cua-driver typed-browser surface baked
into computer_use. Remove it so computer_use is a pure DESKTOP-control tool
(screenshots / mouse / keyboard / window management) and every call's schema
drops ~24 browser-only params + 9 actions.
- schema.py: 9 cua_browser_* actions and the typed-browser param block removed;
14 desktop actions + shared params kept; description drops the browser rung.
- tool.py: cua_browser entries out of _SAFE/_DESTRUCTIVE_ACTIONS; the whole
cua_browser dispatch block deleted; {"type","cua_browser_type"} → "type"
(desktop typing untouched); _config_preauthorized (browser-prepare-only, a
no-op for every desktop action) and the browser-page escalation hint removed.
- browser_route.py deleted (no importers outside the package); cua_backend.py
drops the import + typed_browser_* methods; backend.py drops the non-abstract
defaults.
- tests: browser-route/contract suites removed; browser assertions trimmed.
Desktop control unchanged. 233 computer_use tests pass; the 1 remaining failure
(test_gateway_session_key_yolo_maps_to_unrestricted_mode) is a pre-existing
cross-test state leak — fails identically on origin/main, passes in isolation.
Addresses the five findings from @kshitijk4poor + @GottZ on #95620:
1. macOS 26 LSHandlers parser returned a version number ('7559.97') from the
nested LSHandlerPreferredVersions block instead of the bundle id — detection
returned None on a machine whose default IS Chrome. Strip the nested block
before the role regex.
2. Wrong profile launched (the LinkedIn/Gmail 'logged out' bug): Chrome opens
Default, but the session lives in Local State profile.last_used (e.g.
'Profile 6'). Resolve last_used and mirror its auth files into the copy's
Default on both fresh and refresh paths, so the launched browser is signed in.
3. Private-URL sidecar carried the real cookie jar to arbitrary LAN hosts:
_create_local_session gains allow_real_profile (default True); the
force_local sidecar passes False → always a throwaway profile, and a
real-profile resolve failure no longer breaks private-URL routing.
4. Snapshot permissions were set once (fresh only): now secure the snapshot dir
AND its browser-profile parent on every consented launch.
5. browser.engine=lightpanda + consent gave an unactionable error: guard with
_using_lightpanda_engine() before detection, naming the setting and the fix.
Tests: last_used mirroring (fresh+refresh+fallback), sidecar throwaway + error
isolation, macOS26 parser + detect, perms-on-refresh, lightpanda guard. 198
browser tests pass. Live: Profile-6 cookie DB lands in copy Default (file-level);
real Gmail (Default profile) still signed in.
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:
Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
(_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
(honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod
Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots
Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.
Copy the user's default-Chromium profile (auth state only) into a managed
snapshot, launch Hermes' packaged Chromium on it via agent-browser, and hand
the CDP endpoint to the Browser Use CLI (and built-in tools) to drive. The
snapshot is a non-default dir, so it sidesteps Chrome 136+'s default-profile
remote-debugging block and never contends with the user's running browser;
launched without mock-keychain switches so keyring-encrypted cookies decrypt.
- consent-gated browser_exec 'local' arg (schema only appears with consent)
- fail-closed on non-Chromium default / snapshot failure
- stale-session guard: reuse only when the live session is on our copy dir
- snapshot excludes extensions/service-workers (renderer wedge) + caches
_navigation_session_key returned the ::local sidecar key for local_browser
before the cloud-provider and auto_local_for_private_urls checks. Two
consequences:
- Every private-URL gate in browser_navigate (credential-bearing query,
_is_safe_url pre-nav, post-redirect) is keyed off the sidecar key, so a
model-supplied browser_navigate(url, local_browser=True) opened LAN
addresses in a host-side Chromium — with the real profile's cookies — even
when the user had set browser.auto_local_for_private_urls: false. Consent to
the profile is not consent to override the LAN routing opt-out; a private
URL now follows auto_local_for_private_urls exactly as without the flag.
- Without a cloud provider the bare session already is the local Chromium
(and already carries the real profile when consented), so the flag created a
second session for the same task on the same user-data-dir, which Chromium's
process singleton refuses. The flag is now a no-op there.
The existing consent/CDP-precedence tests keep their assertions; the sidecar
test now states the cloud-provider precondition it silently relied on.
* fix(terminal): stop claiming a Linux environment — point at the env section; near-neutral tokens
* feat(terminal): default GIT_PAGER/PAGER=cat in session env; drop schema lines the runtime already enforces; fix pty backend claim
* refactor(terminal): unify notify_on_complete+watch_patterns into notify (bool|list); trim pipe-masking prose (runtime hint owns it)
* fix(terminal): background param referenced the unadvertised legacy arg name
* refactor(terminal): background-only modifiers (pty, notify) fail loud on foreground calls with corrected shape
* fix(execute_code): block the new notify arg in the sandbox terminal stub (foreground-only)
Stdio MCP helper subprocesses (npx/binary servers) never import Hermes
code, so they could not self-register in the machine spawn ledger and an
unclean parent exit left them running invisibly forever.
- process_identity.register_child(pid, purpose): ledger mirror of
register_self for spawned children — records the CHILD (pid,
create_time) with this process as spawner. Refuses pid-only entries a
PID reuse could forge. Writes go through the single _append_entry
path under _LEDGER_LOCK (prune + atomic tmp/replace unchanged).
- 'mcp-helper' added to REAPABLE_PURPOSES so the updater's
_ledger_reapable_backend_pids rung flows helpers through its existing
spawner_is_dead gate (live spawner => never reaped).
- tools/mcp_tool.py: best-effort register_child(pid, 'mcp-helper') at
the post-spawn PID capture; never breaks MCP startup.
- reap_orphaned_mcp_helpers(): startup sweep mirroring
_reap_orphaned_desktop_local_serves but ledger-driven — kills only
helpers whose recorded spawner is PROVABLY dead, with a create_time
re-check at kill time. Wired next to the desktop serve reap in
web_server.py.
* refactor(clarify): halve the schema (880->436 tok/call) — same rules, half the words
* refactor(clarify): one question field — questions[] is the only advertised shape (single = one-entry array; legacy shape stays handler-accepted)
* refactor(clarify): unadvertise per-question id — response rows already carry question text + order
* refactor(skills): diet skill_manage schema (924->567 tok/call) — drop prose duplicated in system prompt + validation errors
* refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)
* Revert "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"
This reverts commit 9988b33f177fa82a145d5243d556f113a7e2fd2b.
* Reapply "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"
This reverts commit 49a0a8f18b478824398767d7ae681913cee00640.
* refactor(skills): unadvertise absorbed_into — curator-only vocabulary leaves the shared schema
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.
Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.
Fixes#94654
Two fixups the #75785 review required before landing:
- grep fallback no longer uses --exclude-dir for protected dirs: grep
matches exclude-dir globs against BASENAMES anywhere in the tree, so
--exclude-dir=Downloads silently skipped every nested directory named
Downloads (a repo's own Downloads/ included). Protected-dir searches now
route through find's path-scoped -prune (same traversal-prevention the
find backend uses) feeding grep via -exec. Regression test proves a
nested work/repo/Downloads/notes.txt is still found while ~/Downloads is
not (live filesystem, real find+grep).
- exclusions gated on env.is_local (new BaseEnvironment flag, True on
LocalEnvironment): sys.platform/Path.home() describe the controller, not
the execution host — a macOS controller driving a Linux SSH/container
backend must not prune the remote's unprotected Downloads. Environments
without the flag default to local semantics (warning-carrying skip,
never data loss).
Both sabotage-verified: restoring basename --exclude-dir fails 2 tests.
Folds review findings: surface failed_deletes in CLI and gateway
/rollback output (new gateway.rollback.failed_deletes locale key, 17
locales), emit skipped_oversize on the nothing-to-restore early return
too, document all three report keys in the restore() docstring, and pin
the failed_deletes contract from both sides in tests.
Follow-up to #95491. The restore result dict had two inconsistent
reporting surfaces: skipped_oversize was only present when non-empty
(unlike skipped_user_edits), and failed_deletes was filtered from
restored_files but never surfaced to the user at all (debug-level log
only). Both are the same silent-omission class #95491 fixed for
oversize files; this completes the cleanup.
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.
_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.
Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.
Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
Follow-up to the salvaged #95207 fix, completing the misreport bug class:
- restore() now also drops delete_targets whose unlink failed (OSError
swallowed) from restored_files — the sibling of the kept-oversize
misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
(slash_commands + gateway.rollback.kept_oversize locale key in all 17
catalogs) now tells the user which files were kept because the size
cap excluded them from every checkpoint; previously the file was
correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.
The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.
Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.
No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.
Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.
Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.
Reproduced on main with shipped defaults:
corpus.jsonl (2 MB), agent appends to it, then /rollback 1
safe_restore_plan restore=['corpus.jsonl', 'notes.py']
restore ok=True restored_files=['corpus.jsonl', 'notes.py']
notes.py exists=True content="v1 = 'original source'"
corpus.jsonl exists=False <- deleted, was in no checkpoint
Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.
The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.
The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.
The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.
Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.
Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."
Fixes#84802
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).
The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.
Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
gate: origin unchanged; origin_fallback eligible under the same flags
as origin (per-job attach_to_session wins, else global
cron.mirror_delivery); explicit platform:chat targets eligible ONLY
under per-job attach_to_session — the global flag never activates
them, so it cannot start writing transcript entries into arbitrary
explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
session keys are user-isolated, so a seed without a user_id (origin-
less job into a shared channel) would create an orphan session no
reply resolves to — those targets fall back to the plain mirror. DM
targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.
Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.
15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.