- logger.warning when the 6h ceiling stops the heartbeat (matches the
delegate_task stale-stop precedent) so the eventual watchdog reap is
explainable from logs instead of silent
- new mutation-checked test: past the ceiling the heartbeat stops while
the job still completes
- clearer assertion messages (surface res on failure)
- heartbeat loop continues past a raising activity callback instead of
silently stopping (matches delegate_task / touch_activity_if_due
swallow-and-continue semantics) — one transient error must not drop
watchdog protection for the rest of a long job
- hard 6h elapsed ceiling so a wedged job under HERMES_CRON_TIMEOUT=0
(unlimited child watchdog) cannot mask the gateway watchdog forever
- public get_activity_callback() accessor in tools/environments/base.py
instead of importing the private _get_activity_callback cross-module
- tests: deterministic heartbeat test (event-gated, no timing sleep),
no-callback test now asserts the thread is truly never created, new
exception-survival guard; dead started event removed
- fix comment: delegate_task heartbeat cadence is 30s, not 10s
/simplify-code review found _poll_for_token has a second caller:
web_server._nous_poller (dashboard/desktop device login), which surfaces
str(e) as the UI error_message — so wrapping only in
_nous_device_code_login left the dashboard showing the bare timeout.
Move the enrichment into _poll_for_token's deadline raise so every
caller inherits the guidance, and drop the now-redundant try/except
wrap in the CLI login. Add a source-level regression test driving the
real poll loop (authorization_pending stub client) to the deadline.
A bare 'Timed out waiting for device authorization' gives the user
nothing to act on. The most common cause is Portal sign-in failing in
the opened browser tab (including the server-side CAPTCHA loop from
issue #20605), so point at the Portal login page and the hermes portal
retry command.
Salvaged from PR #75290 by @HexLab98 (timeout-guidance kernel only).
The URL-rewrite portion of that PR was dropped: the live Portal has no
/device route (verified 404 with a real user_code), so rewriting the
manage-subscription verification URL would break login entirely.
Guidance text reworded to reference only real URLs.
/simplify-code findings on the salvage stack:
- extract _category_skill_dirs() as the single category detector; the
install guard and hermes_cli._existing_categories() now share it
(third copy of the heuristic eliminated)
- filter rglob hits through is_excluded_skill_path so vendored /
support-dir SKILL.md files (node_modules, references/pkg) no longer
misclassify a plain directory as a category and block install
- fix inaccurate WHY comment (lock-file check only guards hub-installed
skills, not hand-authored dirs), drop underscore prefixes on locals,
fold the file-collision guard under the single exists() check
Follow-up to the salvaged #76000 guard:
- refuse installing a skill INTO an existing skill directory (hybrid
skill-plus-category dirs whose later update/uninstall rmtree would
destroy the nested skill — sibling case of #75983)
- refuse a stray regular file at the install path with the caller's
ValueError contract instead of an uncaught NotADirectoryError
- regression tests: nested-only category (skills at depth >= 2),
category-inside-skill, file collision
The salvaged hydrate_profile_secret_sources (#74549) seeded its
profile-local env from <home>/.env only, but the documented 1Password
bootstrap flow puts OP_SERVICE_ACCOUNT_TOKEN in the gitignored
<home>/.op.env (mirrored from load_hermes_dotenv). A cold profile using
that flow still failed 1Password hydration — the one unaddressed item
from the sweeper review on #74549. Seed .op.env via setdefault so .env
values win; never touches os.environ. Two regression tests.
Follow-ups on the #75382 salvage (review findings):
- _wenv/_get_wsecret now catch UnscopedSecretError and fall back to
os.getenv for the DEFAULT profile's adapter, which constructs and sends
outside any _profile_runtime_scope under multiplexing — a bare
get_secret would crash its WhatsApp path (fixing one profile by
breaking another). Same pattern as Slack SLACK_APP_TOKEN (#59739) and
the Matrix recovery key. Scoped misses still return the default — no
cross-profile borrow.
- bridge_env overlay extended to the full WHATSAPP_* set bridge.js
consumes (DEBUG, FORWARD_OWNER_MESSAGES, REPLY_PREFIX,
MAX_MESSAGE_LENGTH, CHUNK_DELAY_MS, SEND_TIMEOUT_MS).
- Removed the always-true conditional on WHATSAPP_MODE injection.
Fix#75349
Root cause:
Under multiplex_profiles, secondary profiles run inside
_profile_runtime_scope which installs a per-profile secret scope via
set_secret_scope. The WhatsApp adapter (and the shared
WhatsAppBehaviorMixin + Cloud API adapter) read WHATSAPP_MODE,
WHATSAPP_DM_POLICY, etc. via raw os.getenv(), bypassing the secret
scope. Since os.environ doesn't contain secondary profile .env values,
the bridge silently falls back to 'self-chat' and rejects all inbound
messages with self_chat_mode_rejects_non_self.
Fix:
- Add _wenv() helper in adapter.py that reads WHATSAPP_* vars through
get_secret() (agent.secret_scope), which honors the active scope.
- Replace all os.getenv('WHATSAPP_*') calls in adapter.py,
whatsapp_common.py, and whatsapp_cloud.py with get_secret()-based
equivalents.
- Inject resolved WHATSAPP_* values into the bridge subprocess
environment so the Node.js bridge (which reads process.env) sees the
profile's own configuration.
Changes:
- plugins/platforms/whatsapp/adapter.py: 37 lines (+ helper, bridge_env
injection, 2 os.getenv→_wenv)
- gateway/platforms/whatsapp_common.py: 13 lines (6 os.getenv→_get_wsecret)
- gateway/platforms/whatsapp_cloud.py: 21 lines (9 os.getenv→_get_wsecret)
- New regression test: 6 test cases covering scope isolation, fallback,
and cross-profile non-leakage.
The Matrix adapter read MATRIX_RECOVERY_KEY via os.getenv, so under
gateway.multiplex_profiles every profile resolved the default profile's
key. That produced "recovery key verification failed: Key MAC does not
match" and broke E2EE for secondary profiles (#69090).
Route the read through agent.secret_scope.get_secret, which honors the
active profile's scope, with an os.getenv fallback for an unscoped read
under multiplex (default-profile startup loop) — mirroring the Slack
app-token pattern (#59739). Applied to both the startup verification
site and the status diagnostic.
Fixes#69090
Multiplexed gateways resolve credentials through the fail-closed
per-profile secret scope (agent/secret_scope.py, Workstream A): any
get_secret() read outside a set_secret_scope(...) block raises
UnscopedSecretError. The agent turn installs the scope via _run_agent's
profile-scoping wrapper, but slash-command dispatch does not — so manual
/compress reached provider resolution unscoped and every invocation on a
gateway.multiplex_profiles: true deployment failed with:
Manual compress failed: get_secret('OPENROUTER_BASE_URL') called with
no profile secret scope active while multiplexing is on.
Same bug class as the cron scheduler (#57692) and the /v1/runs agent
path — an un-migrated call site the fail-closed design is meant to catch.
Two changes, both required:
- _handle_compress_command becomes a profile-scoping wrapper around the
existing handler (renamed _handle_compress_command_inner), mirroring
_run_agent: gated on multiplex_profiles, resolves the source profile's
home and runs the whole handler inside _profile_runtime_scope. Covers
the coroutine-side read (_resolve_session_agent_runtime).
- The compressor call switches from a bare loop.run_in_executor(None, …)
to the existing _run_in_executor_with_context helper, so the scope
contextvar survives the thread hop into _compress_context, where the
aux-client provider resolution reads credentials.
Single-profile gateways take the pass-through branch — zero behavior
change (pinned by test).
Tests: 2 added (scoped read inside the executor under fail-closed
multiplexing reproduces the field failure pre-fix; single-profile
pass-through). 162 gateway compress/multiplex-scope tests green.
The Playwright suite fails identically on every PR regardless of diff
(verified on a Python-only PR and a docs-only PR): the mock-backend
Electron window never gets a title, so boot/chat/setup/interim specs all
fail; only the dead-backend boot-failure path still passes. Breakage
window matches the Aug 1 night engines/npm churn (#76499/#76562/#76575).
Gated with 'false &&' in the job condition — delete that to re-enable.
Root-fix + re-enable tracked in #76627 (Ari).
- Merge _aux_free_only() + _aux_openrouter_model() into single
_aux_openrouter_settings() that reads config once via
load_config_readonly (avoids double deepcopy).
- Remove 15-line block comment and 5-line inline comment that
restated what the code already says.
- Trim module docstring from 10 lines to 3.
- Update test patches to target load_config_readonly.
Desktop edit / regenerate / restore-checkpoint send truncate_before_user_ordinal
on prompt.submit. The handler rewrote session['history'] first, then called
replace_messages, and on failure only printed to stderr and still started the
turn. When the durable write fails, in-memory history is already truncated
and history_version is bumped while state.db still holds the pre-edit tail.
The agent flush is append-only for history-dict identities, so the new
exchange is appended on top of the 'undone' turns — durable zombie history.
Fix: persist first; only then mutate memory. On failure return 5008 and
leave memory/DB unchanged.
Based on #72876 by @necoweb3. Adapted: the handler moved from
tui_gateway/server.py to tui_gateway/methods_prompt.py since the PR's base.
Start from main's 13 tests (renamed test_openai_available_reflects_key
to test_openai_available_reflects_audio_key_resolution, added 4 new
tests for xai oauth, elevenlabs secret resolver, openai configured
key, stream cap). Append 12 new regression tests from PR #71084 for
the prefetch pipeline, PCM misalignment, and PortAudio resilience.
Patch platform.system in stream-path tests for main's macOS guard.
Replaces the synchronous per-sentence streamer.stream() loop in
stream_tts_to_speaker with a per-sentence prefetch pipeline. Each
sentence gets its own background thread that fires the HTTP request
immediately, buffering PCM chunks into a per-segment queue (capped at
3 concurrent via semaphore). A single playback worker drains segments
in FIFO order with PortAudio error recovery (reinit up to 3 attempts,
then temp-file fallback). Also fixes PCM chunk alignment for odd-byte
HTTP chunks and increases the worker join timeout from 30s to 300s.
Salvage of PR #71084 onto current main.
Closes#71084
Follow-ups from review of #76113:
- Extract cache_ttl_means_disabled() as the single disable-synonym
predicate; agent_init and prompt_caching_disabled_from_config both use
it so the two detection sites can no longer drift (drift would recreate
the #76085 bug class).
- Mirror _run_reference's not-None injection guard in
aggregate_moa_context (stamping None was a harmless no-op copy).
- Replace a vacuous trailing test assertion with the intended
input-non-mutation check; drop a stray blank line.
- Add a predicate-parity regression test (unknown TTL values keep
caching enabled, matching historical agent_init semantics).
Absorb the useful deltas from the parallel #76121 approach: a single
blank_cache_policy_stub factory so _cache_disabled cannot be left off
hand-rolled SimpleNamespaces, and pin the live agent disable onto MoA
advisor fan-out and one-shot aggregate_moa_context decoration so those
paths track conversation state rather than a fresh config re-read.
Keeps the earlier tri-state prepared-aggregator no-agent fix. Adds
factory and synthesis/advisor regressions.
Coordinates with #76121 / #76085.
Co-authored-by: JoaoMarcos44 <87440198+JoaoMarcos44@users.noreply.github.com>
Prepared-aggregator facades built via __new__ lack _agent. Accessing
self._agent raised inside the planner try and bool-coercion of a missing
snapshot forced False, suppressing config fallback for cache_ttl=off.
Pass a tri-state value and add a no-agent/config-off regression.
Blank SimpleNamespace stubs used by MoA decoration and
plan_cache_sections_for_destination never set _cache_disabled, so
anthropic_prompt_cache_policy re-injected cache_control markers after
operators turned caching off. Stamp the disable onto those stubs from
an explicit flag or the live config, and pass the agent flag from the
MoA aggregator path.
Fixes#76085
Extend phone/JID alias matching to whatsapp_cloud and treat a removed
allowlist env key as empty so sole-entry revoke cannot revive a stale
adapter snapshot.
Track which source seeded the DM allowlist so live intake does not let a
stale env carrier override explicit config, while env-seeded adapters still
reread pairing mutations.
After cleaning up an expired cloud browser session, re-check under
_cleanup_lock whether another thread has already created a replacement.
If so, return the live replacement instead of falling through to create
yet another session (which would be orphaned, or worse, a second
concurrent cleanup could destroy the replacement).
Follow-up to helix4u's expired-session renewal fix in PR #76356.
`hermes desktop` still failed with EBADENGINE demanding Node >=26 after
#76562, on a machine whose `apps/desktop/package.json` already said
`^20.19.0 || >=22.12.0`. #76562 fixed the manifest but not its mirror in
`package-lock.json`, and `npm ci` reads engines from the lockfile:
package.json apps/desktop -> {'node': '^20.19.0 || >=22.12.0'}
package-lock apps/desktop -> {'node': '>=26.0.0'} <- what gated
Chasing that exposed a second, pre-existing problem: the floor #76562
declared was too generous. Running the real `npm ci` against the whole
workspace on Node 22.21.1 fails on a transitive dependency —
npm error notsup Not compatible with your version of node/npm:
react-router@8.3.0
npm error notsup Required: {"node":">=22.22.0"}
react-router 8.3.0 (a direct dependency of both `apps/desktop` and `web`)
declares `>=22.22.0`, which is tighter than Vite's `^20.19 || >=22.12` and
excludes all of Node 20. So `>=20.0.0` promised support the tree cannot
deliver: an install on Node 20 or early 22 passed the installer's gate and
then died inside `npm ci` on someone else's package.
All four engine declarations now state the floor the dependency tree
actually has, `>=22.22.0`: root `package.json`, `apps/desktop/package.json`,
and both of their `package-lock.json` mirrors. The installer gates move with
them (`node_satisfies_build` in install.sh, `Test-NodeVersionOk` in
install.ps1) so a too-old system Node is replaced with the managed one
*before* npm runs, and the failure a user does see names hermes-agent rather
than a transitive package. NODE_VERSION stays 22 — latest-v22.x is 22.23.2,
comfortably above the floor.
The invariant test gains the case that would have caught the mirror drift on
its own: the desktop assertion now pins the tightest floor a dependency
actually declares, and the managed-runtime check compares majors, since
install.sh fetches latest-v{major}.x rather than {major}.0.0.
Verified with real `npm ci --dry-run` over the full workspace:
- node 22.23.2 (what install.sh provisions) -> 1258 packages
- node 26.5.1 -> 1189 packages
- node 22.21.1 (below the floor) -> EBADENGINE naming
hermes-agent, i.e. our own manifest, not react-router
`hermes update` aborted its managed-Python runtime repair with an error
that reads like a contradiction:
⚠ Managed Python runtime repair skipped: cannot import name
'venv_python_path' from 'hermes_constants'
(/home/teknium/.hermes/hermes-agent/hermes_constants.py)
The named file does contain the symbol. The module in memory does not.
main.py imports hermes_constants from the OLD checkout, `git pull` then
replaces that file on disk, and the freshly-pulled managed_uv runs its
lazy `from hermes_constants import venv_python_path` against the module
object Python already cached in sys.modules — the pre-upgrade one. The
ImportError reports the path of the new file, so it looks like the symbol
is missing from a file that plainly has it.
Same update-boundary class already documented on `_UvResult` for the
ensure_uv() arity skew. It fires on the first update from any release
older than 83314ca38, which introduced the symbol.
Both lazy call sites now reload the module from disk on ImportError:
- hermes_cli/managed_uv.py::_venv_python — the reported path
- hermes_cli/gateway.py::get_python_path — same flaw, same fix; a gateway
restarted mid-update hits it identically
Reload rather than a local fallback on purpose. Recomputing the layout
inline would hand-roll `Scripts`/`bin` a second time — exactly what #76105
deduped into venv_bin_dir()/venv_python_path(), and what
test_no_open_coded_venv_layout_remains_in_hermes_cli bans. Reloading fixes
the actual problem (a stale module) and keeps one owner for the layout.
hermes_cli/update_cmd.py imports the symbol at module scope, which is a
different failure mode (the module fails to import at all) and is already
covered by the installer's retry-once for the update-boundary crash.
Tests reproduce the stale-module state by deleting the attribute from the
imported hermes_constants: recovery resolves through the reloaded shared
helper (asserted via a sentinel, so an open-coded copy cannot pass), and
the normal path never reloads. Against origin/main's managed_uv they fail
with the exact reported ImportError.
Fresh installs and `hermes update` both fail at the first `npm ci`:
npm error code EBADENGINE
npm error notsup Required: {"node":">=26.0.0","npm":">=12.0.0"}
npm error notsup Actual: {"node":"v24.15.0","npm":"11.12.1"}
`.npmrc` sets engine-strict=true, so `engines` is a hard gate on every
install. The floor was raised to npm >=12 — but no Node release bundles
npm 12: Node 26 ships 11.17.0, 24 ships 11.16.0, 22 ships 10.9.8. The
requirement is unsatisfiable by any stock toolchain, so the installer
provisions a Node and is immediately unable to install with it.
engines.npm becomes `<11.10.0 || >=11.17.0`. That still excludes the band
the strictness was actually for: npm 11.10-11.16 honor `min-release-age`
but ignore `min-release-age-exclude`, both set in .npmrc, so they apply the
14-day gate to packages we exempted. Verified rather than assumed — npm
11.12.1 fails `ETARGET ... vite@8.2.0 with a date before 7/18/2026` while
11.17.0 installs it.
engines.node returns to >=20.0.0 and the toolchain floor to Node 22.
Nothing in the tree needs 26: Vite 8.2.0 declares `^20.19.0 || >=22.12.0`
and Electron 40 declares `>=12.20.55`. Requiring 26 force-migrated every
working install for no dependency reason. apps/desktop drops to Vite's own
floor for the same reason; the desktop bundle builds clean on Node 22.
install.sh gained a second gate: a system Node was accepted on version
alone, so a machine with Node 24 + its bundled npm 11.16.0 (the bad band)
passed the check and then failed `npm ci`. npm_supports_npmrc() now rejects
that band and installs the managed Node instead.
CI, Docker and nix are hermetic and keep pinning Node 26 / npm 12 — they
provision their own toolchain, and both satisfy the relaxed range.
tests/test_engines_satisfiable.py encodes the invariants that would have
caught this: the npm floor must be met by an npm some shipping Node
bundles, the node floor by the runtime install.sh provisions, the desktop
floor by its own build toolchain, and the lockfile mirror must match.
Restoring the broken values fails 5 of them with the reason stated.
Verified end-to-end (real downloads, temp HERMES_HOME):
- fresh install: managed node v22.23.2 / npm 10.9.8 -> npm ci, 209 packages
- existing managed tree (v22.22.3 / npm 10.9.8) -> npm ci, 209 packages
- node 26.5.1 + bundled npm 11.17.0 -> npm ci, 208 packages
- system npm 11.12.1 (the reported case) -> EBADENGINE, recovery provisions
a managed tree and retries green
- apps/desktop `npm run build` on Node 22 -> dist built, assert passes