websockets>=14 defaults to proxy=True and resolves proxies via
urllib.request.getproxies(), which reads the macOS/Windows system
proxy config even with no *_proxy env vars set. Local CDP endpoints
(ws://127.0.0.1:<port>/devtools/...) were therefore dialed through
the system proxy and the handshake failed with "did not receive a
valid HTTP response" (#110565).
Append 127.0.0.1/localhost/::1 to NO_PROXY/no_proxy (both casings)
in _build_browser_env so every browser subprocess (Browser Use CLI,
agent-browser, Chromium, Lightpanda) bypasses proxies for loopback.
Operator-provided NO_PROXY entries are preserved.
Fixes#110565
Kanban cards have no length limit, but the session title store rejects
titles past SessionDB.MAX_TITLE_LENGTH with ValueError. _persist_session_title
reads that as a unique-title collision, retries with a "#N" suffix (longer
still), and the caller suppresses the second failure - so a worker spawned on
a >100-char card ended up with no title at all, where main at least gave it a
derived one. Trim the card title (with room for the "#N" retry suffix) before
persisting; a retried card now gets "<trimmed> #2" within the cap.
Review finding: >100-char card title left the kanban worker session untitled.
Follow-up to the salvaged commit from #111169 (@KoNit-K):
- Drop HERMES_KANBAN_TASK_TITLE. The worker already has HERMES_KANBAN_TASK
and HERMES_KANBAN_BOARD/HERMES_KANBAN_DB pinned in its env, so
maybe_auto_title reads the card title from the board itself (no new
HERMES_* env var for non-secret config; the dispatcher and the
delegation scrub list stay untouched).
- Unreadable or missing card: the session is named `Kanban task <id>`
with zero auxiliary calls (the fallback the issue asked for; the
#109743 seed left such workers untitled).
- The card title persists at `llm` authority via set_auto_title, so a
manual /title still wins and the upgrade thread never starts.
- Tests trimmed to two invariants against a real board + SessionDB
(card title, unreadable-card fallback), both red on origin/main.
The future-instant guard in claim_job_for_fire dropped the occurrence
identity for ANY claim ahead of the stored next_run_at. A hosted/webhook
fire for the armed slot that arrives a few seconds early (the fire
scheduler's clock runs ahead of ours) was therefore treated as an
off-tick run: it ran occurrence-free, mark_job_run recomputed the same
cron slot from a now still before it, and the tick/misfire backstop then
ran the slot a second time.
Only claims at least FIRE_CLAIM_SKEW_SECONDS (60 s) ahead of the slot are
now classified off-tick, so dashboard/manual far-future fires stay
occurrence-free while a skewed early fire keeps the slot identity.
completed_occurrence honours the same window so the early run's
completion row (finished just before the slot) still proves the slot
done instead of being discarded as poison.
Review finding: claim_job_for_fire future-instant guard had no skew tolerance; an early hosted fire for the armed slot ran twice.
tests/cron/test_manual_fire_occurrence.py (from #106970) overlapped the two smaller
invariants salvaged from #110419: an unclassified off-tick claim stays occurrence-free
(test_claim_job_for_fire.py) and a completion recorded before its occurrence is not
proof (test_scheduled_occurrence.py). On-time / late binding is already pinned by
test_scheduled_occurrence.py::test_ledger_migration_and_completion_identity.
Direction (ii) of the t_a662ca7a fix menu, ported from the live install
(ffffab1bd6 on the deployed checkout) and layered on top of the already-
landed direction (i) (ac10770894, manual= flag): claim_job_for_fire() now
also declines to bind an occurrence whose instant is still in the future.
A scheduled tick only fires when now >= next_run_at (_evaluate_due_job
returns False while the stored occurrence is in the future), so a claim
arriving BEFORE the stored next occurrence cannot be the tick that owns
it: it is a manual / dashboard / webhook fire (including the
claim_ttl_seconds reclaim after an expired lease, which arrives with no
manual flag) and stays occurrence-free - ledger row scheduled_instant=NULL,
next_run_at untouched. On-time and late (catch-up) ticks keep the exact
at-most-once occurrence identity.
Live case 2026-09-08: job d88fde170fb4 [bot:carl] Daily Plan Evening
(0 19 * * *), manual run at 19:53 stamped scheduled_instant=
2026-09-10T02:00:00+00:00 (= next day 19:00 PDT) status='completed';
completed_occurrence() then refused every later manual fire ("Job is
already being fired by the scheduler; not run again.") and the next
scheduled tick dedupe-skipped its real delivery. Pre-fix poison rows in
existing ledgers are not repaired by this commit; repair remains
audit-preserving UPDATE-only via the retired reference script.
Regression test (tests/cron/test_manual_fire_occurrence.py): two
invariant tests proven red on base b88e6776f4 (mid-cycle manual fire and
webhook claim_fire bind the future instant) plus three green-on-base
guard nets (on-time tick binding, late catch-up binding, forced-fire
occurrence-free). Red on base: 2 failed / 3 passed; green post-fix:
5 passed. tests/cron suite: 1232 passed, 1 skipped, 1 pre-existing
environmental failure in
test_run_job_cron_execute_code_deny_does_not_pollute_later_gateway_execute_code
that fails identically on base.
Kanban: t_d36f3b54 / t_2a050d21 (defect t_a662ca7a)
A claim that expired without a worker ever spawning (worker_pid NULL) was
reclaimed and immediately re-claimed on every dispatcher tick, with
consecutive_failures stuck at 0 — nothing could trip the breaker. Route the
reclaim through _record_task_failure (own txn after the reclaim commit, same
shape as enforce_max_runtime) instead of the salvaged raw counter increment,
so per-task max_retries / kanban.failure_limit and the gave_up event apply
and last_failure_error carries the stale lock. reclaim_task (operator path)
still resets the counter; the live-worker extend path never reaches it.
Trims the salvaged tests to one invariant that walks the breaker to its trip.
release_stale_claims() reclaimed expired claims without counting the
failure, so a claim that was never spawned (worker_pid NULL) spun forever
with consecutive_failures stuck at 0 and the failure_threshold breaker
could never trip (#111306). The reclaim transaction now increments
consecutive_failures atomically. Extend-live-worker and manual operator
reclaim paths are untouched (extend is not a failure; reclaim_task
deliberately resets the counter).
Tests: 3 new regression tests in tests/hermes_cli/test_kanban_db.py —
reclaim without spawn increments (red on base), reclaim with dead worker
increments (red on base), live-worker extend does not. Related kanban
suites: 79 passed, 1 skipped.
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
The budget in _sweep_agent_cache_under_pressure comes from the gateway's own
cgroup memory.high/memory.max, but read_anon_rss_mb() read /proc/self/status
RssAnon: only the main process. Every child in the unit (execute_code kernels,
terminal commands) is charged against the same limit, so a kernel at 4.7 GiB
pushed the unit to MemoryHigh while the valve saw <1.6 GiB and never fired;
systemd's stop then SIGKILLed the gateway mid-flush (the #80764 signature).
Under a capped cgroup v2, read own memory.stat `anon` (the same scope as the
budget); uncapped, or when the file is unreadable, keep the self reading.
_finite_limit is factored out of _cgroup_limit_bytes so the cap check and
the budget agree on what "unlimited" means.
Fixes#110549
The mid-turn /steer marker is delivered as a standalone role:"user"
message right after the newest tool result (steer_user_row /
apply_pending_steer_to_tool_results). But STEER_CHANNEL_NOTE (the
model-facing briefing) and the module comment still said Hermes
'appends their message to the end of a tool result'.
That stale claim briefs the model to expect the marker INSIDE tool
output, so a real standalone user row carrying the marker can read as
off-channel — the under-trust half of #110979. Correct the briefing and
comment to describe actual delivery; add a contract test tying the note
wording to the steer_user_row delivery mechanism (proven red on the old
wording).
#108319 moved hooks.outbound[].secret_env from os.environ to get_secret, which fails closed
outside a profile scope while multiplexing is on. The launch profile's hooks: block is
registered from start(), before any turn scope exists, so the first secret_env target raised
UnscopedSecretError, register_from_config propagated it, and _register_config_hooks swallowed
it at DEBUG — every launch-profile outbound webhook (signed or not) was silently dropped on a
multiplexed gateway. Secondaries were unaffected: they register inside their own scope.
The launch profile now gets the same treatment: _register_launch_profile_config_hooks enters
_profile_runtime_scope(get_process_hermes_home()) when multiplexing is active, then registers.
Not the environ fallback — a scopeless multi-profile read has no authority over the launch
env; the spawn site is where the scope belongs. The swallow logs at WARNING: a dropped hook
block is not debug noise.
Two invariant tests: multiplex on, no scope, secret in the launch .env -> both targets
register, the signed one with the launch secret (0 register on main); a registration failure
surfaces at WARNING.
#108319 made the release thread carry the caller's context and enter the owner's profile
scope only when no scope was present. But the LRU-cap eviction runs inside the REQUESTING
turn (run_turn_runner: _enforce_agent_cache_cap / _release_evicted_agent), and the LRU
agent it evicts may belong to another profile — so with profile A's scope active, profile
B's end-of-session memory commit ran under A's scope and A's home: B's transcript extracted
into A's provider namespace with A's credentials. Before #108319 the same commit ran under
the launch profile; the fix moved the leak, it did not close it.
_run_release_in_profile_scope now resolves the owner from the session key first — a named
profile's home, else the DEFAULT profile at the root Hermes dir (agent:main: keys), which
is not the launch profile when the gateway runs under `hermes -p x` — and enters that
profile's scope whenever the current one is not already the owner's (compared via
hermes_home_key). Same-profile in-turn evictions and the unscoped housekeeping sweep behave
as before. A failed owner lookup is logged at WARNING instead of silently swallowed, since
it reassigns the commit to the default profile.
Three invariant tests: A's turn evicting B's agent commits under B; a secondary's turn
evicting a default-profile agent commits under the default home; the same with the gateway
launched under a named profile still commits under the root. All red on main.
The share-code codec encodes createdBy through a fixed table that only
knew none/agent/user, so idxOf returned 0 for 'learn' and a learned
skill round-tripped as createdBy: null. The 2-bit field has a free slot,
so 'learn' now occupies it, and metaBadges labels agent- and
learn-created skills alike as 'learned' — matching the Python side,
where both values count as learning signals.
Review finding: learn nodes lost createdBy in share-code round-trip.
Replace the predicate unit test with an invariant on the real builder: a skill recorded
by a foreground create is in build_learning_graph() with zero uses, and an unmarked
never-used local skill is not. `hermes curator list-unmanaged` prints the actual
marker (created_by:learn) instead of hard-coding created_by:null. Docs: curator.md and
memory.md describe the learn marker and what the journey shows.
record_created now stamps created_by="learn" on foreground creates
(e.g. /learn) instead of leaving it unset, and the learning-graph filter
honors "learn" alongside "agent"/used. "learn" is a learning-signal
marker only: curator management stays keyed strictly on "agent"
(_is_curator_managed_record), so user-taught skills appear in /journey
without becoming eligible for autonomous curation.
Fixes#111317.
---
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; the contributor reviewed the diff and ran the tests.
An interrupted first download leaves refs/main plus a snapshot folder
without model.bin. snapshot_download(local_files_only=True) returns that
folder rather than raising LocalEntryNotFoundError, so ctranslate2 fails
with RuntimeError "Unable to open file 'model.bin'" and the online path
never ran, leaving STT permanently broken with a misleading error. Fall
through to the download when the local attempt fails that way. The
loading tests are also gated on faster_whisper being installed, so the
module collects when the voice extra is absent.
Review finding: partial-cache RuntimeError bypassed the online fallback.
The cache-first loader imported LocalEntryNotFoundError unconditionally; in
environments without the optional huggingface_hub dependency every local
transcription raised ModuleNotFoundError. Resolve the cache-miss exception
lazily and fall back to its OSError base class when the package is absent.
The salvaged fallback caught bare Exception around the online load, which
would relabel a CUDA runtime error or an invalid model size as a download
problem and defeat the CUDA → CPU fallback above it. huggingface_hub raises
every network/Hub failure as an OSError subclass (LocalEntryNotFoundError,
HfHubHTTPError), so catch that class only. The negative test now feeds the
real LocalEntryNotFoundError('Got: ConnectTimeout ...') shape the reporter
saw instead of a synthetic RuntimeError.
The salvaged commit only swept `hermes-sync-back-*.tar`. The extraction
staging dir (`tempfile.TemporaryDirectory(prefix="hermes-sync-back-")`) is
leaked by the same hard kill, so the sweep now reclaims both shapes and
returns the count. The window drops from 24 h to 6 h (the reporter's
value): archives appear every few minutes on a busy gateway and a live
transfer is never hours old. Tests trimmed to two invariants: the sweep
touches only stale prefixed entries, and a real sync_back reclaims a
leaked archive while writing its own tar under the identifiable prefix.
Replace the CLI-subprocess parametrization (6 hermes processes per run) with a
direct normalizer invariant: non-finite cadence -> default cadence, non-finite
temperatures -> None while finite values (0, 0.75) survive.
Drop the session-only negative (session grants were never consulted on the unattended
path, so it pins pre-existing behaviour rather than the fix). Keep the permanent-key
positive and the Tirith-not-bypassed negative. Docs: command_allowlist rule keys are
honored in cron/-q/unattended sessions.
Keep the two tests that were red on origin/main (probe failure keeps the clock and
unloads once telemetry recovers; confirmed busy still resets). Drop the is_idle() bool
contract test — it pins behaviour that did not change — and fold the probe-failure log
call onto two lines.
is_idle() treated any /slots or /metrics probe error as "busy", so sweep_idle()
popped the idle clock on every failed probe — one transient failure per sweep
and a resident model never reached IDLE_UNLOAD_S (21 GB pinned for hours with
zero requests).
Split the probe into a tri-state: _probe_idle() returns True (confirmed idle),
False (confirmed busy), None (probe failed). The sweep now resets the clock
only on a confirmed busy sighting; a failed probe keeps the existing clock and
logs at INFO, and unload still requires confirmed idleness past the threshold.
The public is_idle() bool contract is unchanged (probe failure reads as not
idle), so no caller outside the sweep ever unloads on a telemetry hiccup.
Closes#111154
llama.cpp b10964 (the pinned build) renamed --no-webui to --no-ui, so the
router refused to start with an unknown-argument error. The direct-I/O flag
is already selected per-build by _direct_io_args on main, so this reduces
to the one remaining rename and asserts it in the existing spawn test.
Refs #111323
Corrections on top of #110617: hermes_state._ensure_test_isolation raises RuntimeError (the
draft said 'warning'); the older Database Location section still told readers the default
path is ~/.hermes/state.db, which the new section says never to hard-code. Mirror the new
section into the zh-Hans page.
The TemporaryDirectory was closed right after adapter construction, but
_persist_seen_message_ids mkdirs the store's parent back on the first
write, so every run left an empty scratch dir behind in the system temp
dir. Holding the context around the whole test body lets it clean up.
Review finding: test_concurrent_dedup_persists_land_in_order leaks an empty tmp dir per run.
_run_pending_fleet_restart has three supervisor branches; #110709 covers launchd, this covers
the Windows branch (gateway_windows.is_installed -> restart) so a Windows developer with an
installed gateway service does not restart it from the test suite.
test_blocks_private_dns_answer_at_connect_time clears the six proxy env
vars, but httpx resolves environment proxies through
urllib.request.getproxies(), which on macOS falls back to the System
Configuration (scutil --proxy) when no proxy env var is set. On a runner
with a system-wide proxy the request dials the proxy rather than the
rebinding host, and the test fails with:
AssertionError: TCP connect attempted for 127.0.0.1:7890
httpx binds that name at import time (httpx/_utils.py does
``from urllib.request import getproxies``), so patching
urllib.request.getproxies has no effect; patch the httpx-side binding.
Test-only change: production proxy semantics are untouched, and the
mutation check still holds (removing the connect-time guard in
tools/url_safety.py turns this test red again).
One of the test-isolation gaps reported in #110728 (the system-proxy one).
test_run_pending_restart_true_when_no_gateways patches the PID scan and
the systemd unit listings but leaves the macOS launchd restart live, so
on a machine with a real hermes launchd fleet the restart phase drains
those units, reports the fleet restart incomplete, and the assertion
fails. CI never sees this because it runs Linux.
Patch _restart_macos_launchd_gateways to a no-op (the same pattern the
file's _patch_update_deps helper already uses) so the no-gateways state
covers the launchd supervisor scope too.
Fixes#110701
test_concurrent_dedup_persists_land_in_order clears os.environ, which drops the per-test
HERMES_HOME; FeishuAdapter then resolves feishu_seen_message_ids.json under the operator's
real ~/.hermes and every id a running gateway has persisted lands in writes[-1] (265 extra
elements in the report). Keep HERMES_HOME through the wipe so the store stays in the sandbox.
Slim redo of #110708 (same fix, without re-indenting the test body).
test_unavailable_directory_links_are_diagnosed_without_creating_targets already proves a link
with a missing target is refused, not materialized; keep the two new invariants (aliased
parent -> 0o700; operator link at the home boundary still owns its mode).
`_directory_links()` walks every parent up to `/`, and `initialize_home()` skipped
`_secure_dir()` whenever that list was non-empty, so any symlinked ancestor disabled
home hardening and left `~/.hermes` plus its `cron`/`sessions`/`logs`/`memories`
subdirectories at the mkdir default `0o755`.
macOS makes that the normal case: the default temp root is `/var/folders/...` and `/var`
is a symlink to `/private/var` (as are `/tmp` and `/etc`), so any home reached through
those prefixes was never secured. Reported on macOS arm64 as `493 != 448` (0o755 vs 0o700)
in tests/cron/test_file_permissions.py; the same numbers reproduce on Linux by aliasing a
parent directory.
Only links at or below the home boundary are operator-owned. The home itself, or a link
inside it, still transfers permission ownership to the operator — behaviour pinned by
test_initialization_preserves_external_directory_modes and unchanged here. Availability
diagnosis is untouched: a link above the home whose target is missing is still refused
instead of being materialized on the underlying filesystem.
The fallback root when the system temp dir sits inside the operator's
Hermes home was PROJECT_ROOT/.pytest_cache, but the default install checks
the repo out inside that very home (~/.hermes/hermes-agent,
%LOCALAPPDATA%\hermes\hermes-agent), so the relocated basetemp still
resolved under the native home and get_default_hermes_root() pointed the
sandbox back at the live install. mkdtemp under native.parent is outside
the home by construction, and a loud assertion now fails collection if the
chosen basetemp ever resolves inside it.
Review finding: fallback basetemp under PROJECT_ROOT/.pytest_cache stays inside the native home when the repo lives in ~/.hermes.