The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
A Safari/Firefox/Helium user who merely had Chrome installed watched Chrome open on every desktop update (community report). start_ui now checks the system default browser (LaunchServices https handler on macOS, xdg-settings on Linux) and skips the shim window unless the default is Chromium-family; notify_fallback and the durable result file still carry the outcome. Detection is best-effort: any failure keeps the old behavior.
On error/manual outcomes stop_ui('leave-window') kept the browser shim
window open indefinitely, so an aborted update left a Chrome window on
screen until the user closed it by hand; repeated update attempts piled
up more windows.
stop_ui now always closes the shim. leave-window paths keep it up for a
short grace period (HERMES_UPDATE_SHIM_GRACE_SECONDS, default 15) so a
watching user can read the message, then close it. The success path is
unchanged. The error/manual outcome is durably written to
.hermes-update-result.json and surfaced in a dialog on the next Desktop
boot, so closing the shim loses no information.
Brave renders its own P3A privacy-notice bar ("Got it" / "Disable" /
"Learn more") at the top of the throwaway-profile window the posix shim
opens, cramped to unreadability at the shim's small size - the same
window-pollution class as Edge's MSA sync notice, and equally immune to
the throwaway --user-data-dir (#88682). Drop Brave from the candidates
on both platforms; Chrome and Chromium stay.
Covers #88682 on top of #88410
Edge's OS-level Microsoft-account integration signs even a fresh
throwaway profile into the user's MSA and renders its own "syncing
your browsing data" notification — the user's MSA email included —
inside the update window, which is titled "Hermes" (#88410). The
throwaway --user-data-dir start_ui already passes cannot block that
OS-account path, so the only reliable protection is to not pick Edge
at all: drop it from the browser candidates on both macOS and Linux.
The update UI is a best-effort layer — with no other Chromium-family
browser installed, start_ui falls back to its existing
"no renderer; skipping UI" path and the update itself is unaffected.
Fixes#88410
After a remote backend update restarts the gateway, the window WebSocket often dies without a close event (SSH/tailscale tunnels) and users force-quit to recover. finishBackendApply now nudges the registered reconnect handler, which rides the new ping liveness probe: healthy sockets are left untouched, dead ones are force-closed and re-dialed. The old blind gateway.close() on every wake signal is removed in favor of the probe.
macOS sleep/wake (or a silent network drop) can leave the renderer's
WebSocket half-open: no close event fires, so connectionState stays
'open' while every RPC hangs until its per-call timeout. prompt.submit's
timeout is 30 minutes, so the user's next message reads as "enter does
nothing until I restart the app".
- Add a minimal ping RPC (tui_gateway/server.py) answered synchronously
on the WS reader thread.
- On wake signals, reconnectNow now probes the open-looking socket with a
5s-bounded ping and force-closes it on failure, letting the existing
reconnect machinery (backoff, tile rebinding, session refresh) take
over. A pre-ping backend answering -32601 is treated as healthy.
- Tests: half-open socket force-reconnects; healthy socket untouched;
method-not-found backend untouched; backend ping envelope contract.
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
The whole test file legitimately runs ~14s on CI (heavy dynamic import paid
by the first test), brushing the global 15s per-test budget. Slow runners
tip the first test over and cascade-fail all 11 — hit twice in a row on PR
#93612 and on a main run in the same hour. Raise the file's describe-level
timeout to 60s; individual tests still run in milliseconds locally.
Sabotage-verified: timeout:1 fails all 11, 60s passes all 11.
Widen the exception guard from OSError to Exception (re-raising
KeyboardInterrupt/EOFError first) so any prompt_toolkit runtime
failure degrades to input() — matching the established pattern in
masked_secret_prompt. ValueError and RuntimeError can arise from
exotic stream wrappers or event-loop issues with the same root cause:
prompt_toolkit cannot attach stdin on the terminal.
Add test_line_input_falls_back_to_input_on_any_prompt_toolkit_failure
covering the ValueError case.
Cover the prompt_toolkit runtime-failure path added in the fix commit: a
tty-reporting stdin where prompt_toolkit raises OSError(22) (macOS kqueue
EINVAL on fd 0 under curl|bash installs) must degrade to input() instead
of aborting the setup wizard.
line_input() only guarded against a missing prompt_toolkit (ImportError),
not against prompt_toolkit failing at runtime. On some terminals isatty()
returns True but the asyncio event-loop selector rejects registering stdin
(macOS kqueue raises OSError EINVAL / 'Invalid argument' for fd 0), so
prompt_toolkit's Application.run() crashes while attaching its input.
This aborted 'hermes setup' at the first plain text prompt. Telegram hit it
first because its automatic/manual selection uses prompt() rather than the
curses-based prompt_choice() the other platforms use, but every text prompt
shared the same failure.
Catch OSError from the prompt_toolkit path and fall back to the built-in
input() reader, which needs no selector and works in cooked mode. The
prompt_toolkit raw-mode context manager restores terminal state on the way
out, so the fallback reads cleanly.
Replace hardcoded /opt/hermes check with dynamic install-tree detection
using Path(__file__).resolve().parent. This catches ALL install paths
(Docker /opt/hermes, apt /usr/local/lib/hermes-agent, git clone, custom)
instead of just the Docker image path. Also covers subdirectories of the
install tree, not just the top-level dir.
Add regression test test_install_tree_skipped to verify both the install
root and subdirectories are excluded from chmod.
Add contributor email mapping for bradmarshall987.
Follow-up to PR #93050 by @bradmarshall987.
secure_parent_dir() is called before credential file writes to harden
the parent dir. Its existing safety check refuses only paths with fewer
than 3 path parts, but /opt/hermes is exactly 3 parts, so it passes and
gets chmod'd to 0700. UID 10000 (hermes) cannot then traverse the
install dir and every new exec fails with 'Permission denied' until
manual chmod 0755 /opt/hermes.
This change:
- Adds an explicit refusal for /opt/hermes parents in secure_parent_dir()
- Adds chmod 0755 /opt/hermes to the Dockerfile install step next to
the existing bin chmod, so the dir starts traversable
Reproducer: any auth write to a file directly under /opt/hermes
(e.g. auth.json when HERMES_HOME resolves there). Observed in
production 2026-07-06 and 2026-08-22. See #25821 for context.
Fixes#25821 follow-up.
Deterministic, LLM-free conformance cells against the real SessionDB with
real SIGKILL mid-write, per the tracking issue's spot-probe method:
- cell 1: acknowledged-append durability + recovery determinism (adapted
from the issue's 29.5K probe, scaled kill window, identical assertions)
- cell 2: consume-once under 8-process concurrent claim_handoff
- cell 3 (new): compression-rotation atomicity — never a compression-ended
parent without a continuation (#80337 contract; #80487 recovery context)
- cells 4-5: documented stubs interlocked with #82956-#82959 and
#83197/#83557
Journal-mode matrix (resolver default / DELETE / WAL-with-skip-gate) per
cell; every wait deadline-bounded; writers asserted alive at kill time.
Review v2 of #93084: the readable-cmdline path used substring matching
as destructive authority. That fails open on prefix collisions
(--profile timothy vs our profile tim) and same-name profiles under
different roots, so a poisoned record could still reach SIGTERM.
Ownership is now decided by the persisted identity record ALONE — exact
_same_hermes_home equality, bound to the live target by exact pid +
start-time. Missing/legacy/unbound/foreign records all refuse. A
readable argv feeds only a token-exact consistency check
(_looks_like_profile_conflict_from_cmdline via shlex tokens) that
refuses explicit contradictions like --profile timothy under tim; bare
or matching argv adds nothing.
Also fixes the review's source-of-record concern: the guard validates
the record that authorizes get_running_pid()'s answer rather than
assuming {HERMES_HOME}/gateway.pid is always the source.
Adds the requested signal-boundary regression: start_gateway(replace=
True) with unprovable ownership returns False without calling
terminate_pid or writing a takeover marker; the bound same-home
counterpart still reaches the replace flow. Legacy replace-flow tests
updated to stage a valid bound record for their legitimate-replace
fixtures.
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.
Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes the TOCTOU window flagged in review on #93369 (merged via
#93430): the divergence guard compared EXPORT-TIME message counts, but
another backend can append donor messages between the export snapshot
and the retire loop — that growth would be stamped behind the
non-recoverable adopted_by_profile archive, the exact H2 class the
guard exists to prevent, just via a narrower race.
The retire loop now re-reads live donor vs local counts immediately
before end_session and leaves the donor unretired (donor_retired=False,
warn-logged) on any donor-ahead signal; the next resume's export-time
guard then handles the divergence normally. Equal-count CONTENT
divergence (donor rewind+rewrite) remains invisible to count comparison
— documented as accepted: bytes stay in the donor store either way.
New red-first-verified regression simulates the exact race by appending
to the donor from inside an export_session_lineage wrapper.
adoption+ownership suites: 25 passed; ruff clean.
Follow-ups on top of the salvaged #92785 commit:
- Pre-gate now matches base URLs via normalize_route_base_url and
provider ids via custom_provider_aliases, mirroring the semantics of
get_custom_provider_model_capability. The raw string comparison
silently dropped declarations whose config spelling differed only by
host case or trailing slash (proven empirically: …/v1/ vs …/v1 with a
non-matching provider name returned (False, False) despite an explicit
prompt_caching: true).
- get_provider(..., allow_network=False) in the early-init/stub branch:
the policy runs per request destination (MoA aggregator, auxiliary
replans via blank_cache_policy_stub, early agent init) and a cold
models.dev cache triggered a measured ~450 ms foreground registry
fetch from the send path. A catalog miss degrades to the conservative
side.
- Debug-log the previously silent provider-lookup exception fallback.
- Tests: _make_agent defaults _custom_providers=[] (post-init reality;
keeps built-in-route tests off the catalog/config fallback), the two
early-init tests delete the attr explicitly, and three regression
tests pin the URL-drift, spaced-legacy-name, and no-network contracts
(all three fail on the unfixed commit).
Apply explicit per-model prompt_caching capabilities to custom
chat-completions routes, rather than limiting them to recognized providers,
hosts, or model families.
Keep undeclared routes conservative, derive the marker layout from the wire
transport, and leave Responses and Bedrock caching paths unchanged.
Merging main brought in the RFC 8252 native sign-in path for password
providers (#75808), added while this PR was open. Its loopback-code
branch calls clear_pkce_cookie() without use_https, which is now a
required keyword-only argument — so /auth/native/password-login raised
TypeError on the success path.
This is the same call-site class the PR already fixed at the other three
sites: the deletion must mirror the shape the setter emitted for the
active origin, or the browser keeps the stale PKCE cookie.
Caught by CI running the merge commit against main's newer
test_dashboard_auth_native_flow.py suite, which does not exist on the
branch. Three tests failed there and pass with this change.
The SameSite=None change updated the website docs but left two
source-level contracts asserting the opposite:
- base.py: LoginStart.cookie_payload said cookies set there "MUST"
be SameSite=Lax.
- cookies.py: the module docstring said all three cookies are
SameSite=Lax.
Both now describe the actual behaviour: session cookies stay Lax, the
short-lived PKCE cookie is SameSite=None; Secure over HTTPS and Lax
over plain HTTP. A provider author following the old base.py contract
would have had a documented reason to undo the fix.
Also records the forwarded_allow_ips caveat in cookies.py: uvicorn only
honours X-Forwarded-Proto from a peer inside forwarded_allow_ips
(default 127.0.0.1), so a TLS terminator reaching the dashboard from a
non-loopback address (a reverse proxy in its own container) leaves the
request looking like HTTP and the cookies written in their HTTP shape.
Docstrings only; no behaviour change.
macOS Chinese pinyin IME: pressing Enter to confirm a candidate word in
the group-chat composer submitted the draft as a message mid-composition.
The GroupMentionInput onKeyDown checked only `event.key === 'Enter' &&
!event.shiftKey` with no IME guard, unlike the core composer which guards
isComposing + keyCode 229 (#44135).
Add the same guard to the three Enter handlers in the bots plugin:
- GroupMentionInput (group composer + reply box) — the reported bug
- GroupClarifyCard free-text answer input — same premature-submit
- skill-hub search input — same premature-trigger
Closes#93528
The hosted-provider misfire catch-up (fire_overdue_jobs) fired any runnable
overdue job with no one-shot grace check, so a stored past-due one-shot
bypassed ONESHOT_GRACE_SECONDS and executed arbitrarily late after downtime.
Sibling site of the due-scan gate from #89571; pins both directions with
tests.
create_job / update_job / resume_job all reject a one-shot whose run time is
more than ONESHOT_GRACE_SECONDS in the past ("will never fire"), and
_recoverable_oneshot_run_at never recovers such a schedule — but
_get_due_jobs_locked dispatched ANY one-shot whose *persisted* next_run_at was
in the past, even hours later (gateway down past the window, host asleep,
hand-edited jobs.json). A wall-clock one-shot then ran hours late, violating
the "will never fire" contract enforced everywhere else.
- Grace gate: a once-kind job whose next_run_dt is more than
ONESHOT_GRACE_SECONDS in the past is never appended to the due list.
- If no run_claim/fire_claim exists (nothing was ever dispatched), retire the
record with a diagnostic file so it stops being scanned and the miss is
operator-visible.
- If a (possibly stale) claim exists, a run may still be in flight in another
process: skip this scan but KEEP the record so its mark_job_run can land
(avoids re-introducing mid-flight record deletion).
- Manual re-trigger still works: trigger_job sets next_run_at=now (inside
grace) so an explicitly re-run stale one-shot fires.
Tests (tests/cron/test_oneshot_grace_due_scan.py): stale-not-due+retired,
within-grace-still-due, stale+claim-skipped-but-kept, retriggered-is-due, and
recurring-jobs-unaffected.
Behavior-contract tests for sanitize_task_id_for_path (colon/separator
removal, verbatim pass-through for existing safe ids, determinism,
collision-freedom incl. the a:b vs a_b digest case, traversal and
oversized-id bounds) and for the singularity persistent overlay path
(sanitized, verbatim for safe ids, distinct dirs for colon-vs-underscore
ids).
Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
Hoist the sandbox-directory sanitizer into tools/environments/base.py as
sanitize_task_id_for_path() and route BOTH host-path consumers through it:
the docker persistent sandbox (get_sandbox_dir()/docker/<id>) and the
singularity persistent overlay (hermes-overlays/overlay-<id>). One helper,
one mapping, whole bug class fixed in one place instead of per-backend
copies (#92414, #92640, #93044).
docker.py keeps _sandbox_dir_name as an alias of the shared helper so the
sanitized mapping (safe ids verbatim, digest suffix on rewrite for
collision safety) is unchanged for existing sandboxes.
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
Drives the real DockerEnvironment constructor with a Telegram DM session key
and asserts every persistent -v spec is a two-field bind whose source holds no
colon — the assertion that reproduces exit 125 on the unfixed path.
The derivation's own contract is covered separately: ids that already work stay
verbatim (no sandbox migration), docker's separator and the path separators
never survive, ids differing only in rewritten characters keep distinct
directories, the mapping is stable across calls so cross-process container
reuse still resolves, pathological keys stay inside the per-component length
limit, and "."/".."/empty cannot resolve to the docker sandbox root.
With terminal.backend: docker and container_persistent: true, every gateway
session failed on its first tool call: docker run exited 125 with
"invalid spec ... too many colons" and no command could execute.
_resolve_container_task_id() returns "session:<key>" whenever a session key
is present, and gateway session keys are colon-delimited
(session:agent:main:telegram:dm:<chat_id>). DockerEnvironment joined that id
into the persistent sandbox path verbatim, so the -v spec became
".../docker/session:agent:main:telegram:dm:<id>/home:/root" — docker splits a
spec on ':', read the extra fields as extra mount options, and refused the
run. The container label a few lines below already guards this exact value
class via _sanitize_label_value(); the bind-mount source did not.
Derive the directory name through _sandbox_dir_name() instead. Ids that are
already bind-mountable are returned verbatim, so the shared "default" sandbox
and RL/benchmark rollouts keep their existing directory and no installed
package or /root state moves; only ids that could never have produced a
working mount are rewritten. A rewrite carries a digest of the original id,
because ':' -> '_' alone is not injective and would otherwise collapse two
chats onto one persistent /root.