The app-wide .scrollbar-dt theme (on #root) rendered 4px scrollbars
everywhere, including the conversation window, making them nearly
impossible to see or grab (#111634). Bump the webkit track size to 8px
so the thumb is actually hittable while staying a slim themed bar.
Fixes#111634
CI runs the suite as an unprivileged user with HOME=/root in one fixture;
Path.is_dir() raised PermissionError from _user_local_bin_entries and the
run-env builder crashed. An unreadable home has no usable ~/.local/bin, so
treat the OSError as absent.
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.
Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.
Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.
Fixes#111778
The ghost-prune loop in reconcileUnifiedDesktopHalves used existsSync on
the marker's source, which answers false for EACCES/EPERM as well as
ENOENT, so a mode-000 / ACL-denied package folder was treated as an
uninstall and its materialized half rm -rf'd with no warning.
Distinguish a genuinely missing source (ENOENT/ENOTDIR) from one the app
is not allowed to stat: keep the half and warn. materializeDesktopHalf
now also warns on a non-ENOENT stat failure instead of swallowing it.
Part of #111804
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.
Part of #111804
`reconcileUnifiedDesktopHalves` let a stat/copy failure on one package
(`EPERM: lstat` on Windows in #111804; EACCES on a mode-000 file here)
reject the whole pass. The `hermes:fs:desktopPluginsRoot` IPC runs that
reconcile before returning the root, so the renderer never got a root and
every disk desktop plugin silently stopped loading. Warn about the one
package and keep materializing the siblings.
Part of #111804
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.
Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.
Part of #111804
Once the launchd file carried the macos_only marker, the first real macOS
lane run showed TestIncompleteWarningMentionsLaunchctl asserting the
systemd-host branch (kickstart hint, systemctl line) that never runs on
macOS. The two Linux-arm tests move to their own linux_only file; the
launchd file keeps the macOS arm (bootstrap/list hint, no systemctl).
With the macos_only marker the win32 skipif is redundant (the marker already
skips every non-darwin host) and the is_macos() fake in _wire contradicts the
repo rule that host-specific behaviour is tested on that host, not by making
the interpreter believe it is elsewhere. Drop both.
_get_service_pids(all_profiles=True) also runs a bare `launchctl list` prefix
scan, so on a macOS box with a live ai.hermes.gateway* fleet the scoping
assertions picked up real PIDs (the one failure the reporter could clear by
booting the gateway out). Stub subprocess.run in _wire so the tests assert on
routing, not on the developer machine.
The cherry-picked tests/ci change-detector (asserting one specific file is in
the macos_only list) is dropped: the selector test suite already covers the
mechanism, and the marker is now the file's declaration.
The example comment said the guard refuses writes to instruction files;
the code (tools/file_tools_write_guards.py) always prompts a human, even
under --yolo, and refuses only when nobody can answer.
The corrected invariant still read as if "inactive loops" were the thing
being bounded. Say what the watchdog measures (idle time, not wall-clock),
what that means for operators sizing work (an active job is never cut off),
and that attached scripts have their own bound (_DEFAULT_SCRIPT_TIMEOUT),
which is the second misreading #111633 calls out.
agent/agent_init.py::_configure_ollama_num_ctx caps only the auto-detected
value (the cap is skipped when an explicit override is set). The salvaged
example wording said context_length caps the explicit value too, which
would send readers chasing a cap that does not apply.
Real-profile browsing and cron's per-job toolsets are both documented, but
nothing connected them: browser.md never mentioned scheduled runs and cron.md
never mentioned login state, so the constraints of driving a login-gated site
from a cron job were only discoverable in source.
Adds a Scheduled and unattended runs subsection to the real-profile section:
the browser.use_real_profile prerequisite (off by default), the credentials a
login form or a fresh 2FA challenge needs saved ahead of time, the auth re-sync
behaviour and its Windows consequence. Plus a pointer from cron.md's toolset
section.
Two default claims in the example contradict the code:
- agent.verify_on_stop: the example says the default is "auto"; the default
in config_defaults.py is False, agent/verification_stop.py treats "auto"
as the explicit opt-in for the surface-aware mode, and the website
already says off. The comment and the sample value now say so.
- display.show_reasoning: the example marks false as the default and ships
show_reasoning: false; DEFAULT_CONFIG and cli.py's own defaults are True
("Default ON ... with this off the user stares at a spinner"). Anyone who
copied the example silently turned reasoning display off. The marker and
the sample value now match the code.
Six keys the runtime reads but the example never mentioned:
- security.protected_instruction_files / protected_instruction_extra_patterns
(tools/file_tools.py) — a write guard for AGENTS.md/CLAUDE.md/SOUL.md and
friends with no documentation anywhere.
- browser.restrict_evaluate / browser.allow_unsafe_evaluate — both in
DEFAULT_CONFIG, only the former mentioned in the website browser page.
- tts.delivery_profiles.<platform>.{max_file_bytes,safety_ratio}
(tools/tts_tool.py) — per-platform audio upload limits over the built-in
discord/telegram/default table.
- gateway restart_after_turn_timeout (config_defaults.py, 1800) — the
in-band restart knob its own sibling comment tells readers to prefer.
- model.ollama_num_ctx (agent/agent_init.py) — the documented-in-code VRAM
cap for Ollama's num_ctx.
display.streaming was left alone on purpose: the CLI builds its config from
its own defaults (streaming: True) rather than DEFAULT_CONFIG (False), so
the example's "(default)" marker is correct for the surface it documents.
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.
- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
probe's grandchildren) in the SSH connection guide.
The user-visible failure was the ownership capability probe reporting a
current remote as unsupported when sshd ran it under a non-interactive zsh.
Pin the probe itself, not only the wrapper: run remoteSupportsSshOwnership
through `zsh -c` against a fixture CLI that advertises both flags; skipped
where no zsh is installed. Fails on the pre-fix wrapper (empty capture).
Co-authored-by: the-repeter <50600051+the-repeter@users.noreply.github.com>
requestGatewayForProfile always dialed with the 'background' spawn priority,
so the Vault tab under the Settings 'Applies to' selector kept the #111651
infinite-spinner path on a cold profile even after the REST side moved to
scopedDialPriority. Let requestGatewayForProfile take a spawnPriority option
and pass 'foreground' from the Vault panel; ambient callers are unchanged.
Part of #111651
Fold the salvaged per-call ternaries in api/config.ts into a single
scopedDialPriority(scope) helper on api/client.ts and apply it to the
Capabilities scope selector's cold-start reads (getSkills / getToolsets /
getMcpCatalog), which hit the same background-capped pool queue when the
selector targets a stopped profile.
Drop the salvaged main.ts source-text change-detector test; keep only the
string update the existing #90812 wiring test needs. The renderer seam test
(hermes-capability-scope) is the invariant: an explicit scope carries
priority 'foreground', the ambient path stays untagged.
Part of #111651 (salvage #111672)
In connect()'s post-spawn catch, a cleanupStale rejection (indeterminate
ownership probe) replaced the boot failure's message and kind. Catch it,
attach it as error.cleanupCause, and rethrow the original error.
The Desktop SSH bootstrap proves the remote `hermes serve --isolated` alive
(`kill -0 … && echo ALIVE || echo DEAD`) and owned (argv probe printing
OWNED/FOREIGN) over the same SSH channel that is often mid-teardown right
after the served token was resolved. Both probes treated ANY answer that was
not the positive sentinel as the negative verdict, so an exec that resolved
with empty output was read as death: the boot failed with "remote dashboard
exited while its served token was being resolved", the post-spawn cleanup ran
the ownership probe on the same channel, read the lost answer as FOREIGN,
skipped the kill and removed the lockfile — one orphaned ~144 MB backend per
failed attempt, ten in one session on a 1 GB host (#111810).
`execProbeVerdict` now runs both probes: an answer that is neither sentinel is
indeterminate and is retried over a short bounded window (3 attempts, 500 ms);
with no definite answer it fails closed with a `transient-transport-error`, so
`cleanupStale` keeps the ownership record for the next connect to reap instead
of leaving a lockless orphan. The post-spawn catch probes liveness first so a
child that genuinely died at startup skips the ownership proof and the
original error ("exited before announcing") stays visible.
Slimmer redo of #111828 by @kokhlo: same mechanism, without the parallel
`verifyRemotePidAlive`/`verifyPidOwnership`/`readOwnershipVerdict` layer that
left `remotePidAlive` and the original probe as dead duplicates.
Fixes#111810
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
The salvaged comment promised the path component stays within
[A-Za-z0-9._-], but sanitizePartitionComponent() also emits the other
characters encodeURIComponent leaves alone (!~*'()). State the real
invariant — nothing Electron percent-escapes, pinned by the test as
"no ':' and no '%'" — and record the migration decision: one partition
name on every platform, so a non-primary cookie-auth remote signed in
on macOS/Linux under the old `:conn:` name is re-prompted once, and the
old `%3Aconn%3A` folder stays on disk, inert. apps/desktop/AGENTS.md gets
the rule so the next partition does not repeat #92183's folder name.
Electron escapes ':' in a session partition name as '%3A' for the on-disk
profile folder. On Windows, a profile folder whose name contains '%3A' gets a
cookie store the network stack can neither read nor write: a jar placed there
reads back zero cookies, `cookies.set()` never reaches disk, and every request
goes out with no cookies at all — the gateway answers 401 `no_cookie` and the
desktop concludes the user is signed out.
Per-connection partitions (#92183) are the only ones carrying a colon beyond
the 'persist:' prefix, so every NON-primary cookie-auth remote hits this: the
dial mints no ticket, the reauth copy fires ("Remote Hermes gateway uses
OAuth, but you are not signed in…"), and the login window — riding the same
partition — writes into the same invisible jar. Nothing ever persists, so the
prompt returns on every connect. The registry primary and the v1 remote keep
the legacy `persist:hermes-remote-oauth` jar, whose folder has no '%3A', which
is why a single-gateway setup never shows the symptom.
Verified against the app's own Electron runtime (40.10.2, Windows): identical
jar bytes seeded into two partition folders read 3 cookies and mint a
ws-ticket (200) from `…-conn-…`, and 0 cookies / 401 from `…%3Aconn%3A…`.
Keep the partition path component inside [A-Za-z0-9._-]; a regression test
pins that it never contains ':' or '%'.
clear_session (/new, /reset, auto-reset boundary) stamped entry.result="deny"
before waking the wait, and an interrupted coalesced leader published the
same deny to its followers, so both still rendered outcome="denied" /
"denied by user". Carry the cause on the entry (entry.cancelled) and let
_cancel_cause map a result-less wake to a withdrawn prompt; the wait still
unwinds fail-closed and the leader's own decision is unchanged.
_YOLO_MODE_FROZEN is read from HERMES_YOLO_MODE at import; a host shell running yolo
auto-approved the gate under test so the prompt was never enqueued.
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).
The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:
- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
category — no string matching) for the interrupted state and marks a
notifier-unregister wake (event set, result None) as "the turn ended before
the prompt was answered". Both the direct and the coalesced-follower wait
return `cancelled=<cause>`; the post_approval_response hook fires
choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
"BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
with outcome="cancelled" and its own user_summary; an explicit /deny is
untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
tool_reason ("parent delegation ended"; the late-child mirror forwards the
parent's own category) so a child's pending approval can tell teardown from a
user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
instruction-file gate and MCP elicitation consume the same key instead of
reporting "denied by the user" / "decline".
Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
terminal_tool's approval gate answers `status: pending_approval` with an
EMPTY `error` (#28323) and no `session_id`, so _spawn_delivery's specific
branch (`if parsed.get("error")`) was skipped and every unanswered
approval fell through to "Delivery to X failed to start: no process id
returned" — blaming the spawn for an approval nobody in a non-interactive
turn (api_server, `hermes peer dm`, cron) could grant.
- _spawn_delivery: the pending shape gets its own message (the runner
command needs terminal approval nobody in this turn can grant); a
local/peer DM adds "nothing was sent — approve it or add it to
command_allowlist and send again". Ownership is never transferred, so
the existing finally still reclaims the plaintext DM file.
- _try_relay_delivery: the envelope is queued on disk BEFORE the reply
waiter spawns and the Desktop drains it independently, so ANY waiter
spawn failure is a lost wake-up, not a failed delivery; reporting it as
an error made the sender resend and deliver the message twice. The
relay path now returns the shape _start_delivery's live-owner branch
already uses (status queued + notification_error + "Do NOT resend")
instead of inventing a new status value nothing reads.
Slimmer redo of #92971 by @jonpol01 (same diagnosis, same relay/local
split on `dm_file is None`); the source-text contract test and the
`sent_no_reply_wake` status were dropped.
Fixes#111716
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Two defects in tools/async_delegation.py:
- _push_completion_event called _persist_completion unguarded before
publishing onto completion_queue. One sqlite3 error (locked/full
state.db) dropped the completion event, left the record parked on
"finalizing" (a permanently leaked max_concurrent_children slot) and
let recover_abandoned_delegations later rewrite a succeeded unit as
"unknown". The write is now try/except: the failure is logged and the
event is still delivered, so _finalize flips the status and frees the
slot. A lost durable row is acceptable degradation; a lost result and
a leaked slot are not.
- _prune_completed_locked treated anything != "running" as finished,
while the module's own _LIVE_STATES also names stalling/finalizing.
A stalling record has no completed_at, so it sorted oldest and was the
first eviction candidate once the retained cap overflowed; its late
runner return then hit the missing-record path and the real result was
dropped. The predicate is now `status not in _LIVE_STATES`.
Slim redo of #76606 (earliest fix) and #112031: the converge/shield/
delete-row machinery both PRs built around the write is dropped as
defense-in-depth; the two core hunks are ported as-is.
Fixes#76605Fixes#112030
Co-authored-by: luckystar2026 <1393268817@qq.com>
Review follow-up on the keyless Exa UTF-8 fix:
- `_parse_mcp_body` split SSE frames on `\n` only, so a body using bare CR
line terminators (permitted by the SSE spec, and parsed fine before via
`splitlines()`) failed with "Unrecognized MCP response shape". Split on
the SSE terminators CRLF / CR / LF instead.
- `mcp_call` decoded the body as UTF-8 unconditionally, ignoring a charset
the server did declare. Decode with the declared charset when the
Content-Type carries one and fall back to UTF-8 only when it is absent.
- The HTTP >= 400 branch and `_keenable_request` still surfaced
`response.text`, so a non-ASCII error body from a charset-less text/*
response reached the user as ISO-8859-1 mojibake. Both now go through the
same decode helper as the success path.
Exa's MCP endpoint answers `text/event-stream` without a charset, so
`requests` decoded `.text` as ISO-8859-1: every non-ASCII character came back
as mojibake, and for CJK results the UTF-8 continuation byte 0x85 became
U+0085, which `str.splitlines()` treats as a line break — the `data:` JSON
line was cut in two and the perfectly valid result surfaced as "Unrecognized
MCP response shape", after which the ring fell through to keyless Firecrawl
(403). Decode the bytes as UTF-8 (JSON-RPC and SSE are UTF-8 by spec) and
split SSE frames on newlines only. A parsed envelope that really carries no
text is now reported as "no text content" instead of an unrecognized shape.
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
The redirect test used a single requested URL, so main's positional
zip(fetch_urls, results) happened to key it correctly and the test only
pinned. Omit the first requested URL and redirect the second: positional
pairing now files the page under the wrong requested URL, while the
metadata.sourceURL key still lands on the right one.
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.
Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
_detect_tool_failure now classifies dict results as failures, so the
concurrent worker's failure log line sliced result[:200] on a dict and
raised TypeError; the worker died and the model saw "thread did not
return a result" instead of the tool's own error payload. Stringify the
preview like the sequential path does.
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.
The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.
Salvages #111821 (@KoNit-K), part of #111815.
Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
Replace the cherry-picked 3.14-shape test with two interpreter-agnostic
invariants. The original test did `del pool._initializer` (raises on 3.14,
where the attribute never exists) and drove `_WorkItem.run()` with no
`ctx` (TypeError on 3.14), so it could only ever pass on 3.11-3.13 — the
exact interpreters where the bug does not occur.
The fake worker now records the args tuple and resolves the work item's
future directly, so the tests assert only on the shape the executor picks:
`(ref, ctx, queue)` when `_create_worker_context` exists (3.14+),
`(ref, queue, initializer, initargs)` when the legacy fields do (3.11-3.13).
Both run green on 3.11 and 3.14; the 3.14-shape test is red on main under
either interpreter (`AttributeError: ... no attribute '_initializer'`).
Refs #58596, #111813.