Commit Graph

4721 Commits

Author SHA1 Message Date
teknium1 be67e1d31e fix(tools): tolerate an untraversable HOME when probing ~/.local/bin
CI runs the suite as an unprivileged user with HOME=/root in one fixture;
Path.is_dir() raised PermissionError from _user_local_bin_entries and the
run-env builder crashed. An unreadable home has no usable ~/.local/bin, so
treat the OSError as absent.
2026-09-15 18:49:29 -07:00
teknium1 43e7e830fd fix(tools): fold ~/.local/bin into the POSIX PATH completion siblings, tests + docs
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.

Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.

Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.

Fixes #111778
2026-09-15 18:49:29 -07:00
KoNit-K c68e306ea4 fix(tools): include user local bin in POSIX PATH 2026-09-15 18:49:29 -07:00
teknium1 1e2cb57973 fix(approval): session teardown and interrupted leaders withdraw the prompt instead of denying it
clear_session (/new, /reset, auto-reset boundary) stamped entry.result="deny"
before waking the wait, and an interrupted coalesced leader published the
same deny to its followers, so both still rendered outcome="denied" /
"denied by user". Carry the cause on the entry (entry.cancelled) and let
_cancel_cause map a result-less wake to a withdrawn prompt; the wait still
unwinds fail-closed and the leader's own decision is unchanged.
2026-09-15 18:44:46 -07:00
teknium1 6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1 54d7f75590 fix(bot-mode): a pending command approval is not a failed DM delivery
terminal_tool's approval gate answers `status: pending_approval` with an
EMPTY `error` (#28323) and no `session_id`, so _spawn_delivery's specific
branch (`if parsed.get("error")`) was skipped and every unanswered
approval fell through to "Delivery to X failed to start: no process id
returned" — blaming the spawn for an approval nobody in a non-interactive
turn (api_server, `hermes peer dm`, cron) could grant.

- _spawn_delivery: the pending shape gets its own message (the runner
  command needs terminal approval nobody in this turn can grant); a
  local/peer DM adds "nothing was sent — approve it or add it to
  command_allowlist and send again". Ownership is never transferred, so
  the existing finally still reclaims the plaintext DM file.
- _try_relay_delivery: the envelope is queued on disk BEFORE the reply
  waiter spawns and the Desktop drains it independently, so ANY waiter
  spawn failure is a lost wake-up, not a failed delivery; reporting it as
  an error made the sender resend and deliver the message twice. The
  relay path now returns the shape _start_delivery's live-owner branch
  already uses (status queued + notification_error + "Do NOT resend")
  instead of inventing a new status value nothing reads.

Slimmer redo of #92971 by @jonpol01 (same diagnosis, same relay/local
split on `dm_file is None`); the source-text contract test and the
`sent_no_reply_wake` status were dropped.

Fixes #111716

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-15 18:44:23 -07:00
fangliquanflq 27e3fc51ff fix(tools): keep async delegation results past a failed durable write and never prune live records
Two defects in tools/async_delegation.py:

- _push_completion_event called _persist_completion unguarded before
  publishing onto completion_queue. One sqlite3 error (locked/full
  state.db) dropped the completion event, left the record parked on
  "finalizing" (a permanently leaked max_concurrent_children slot) and
  let recover_abandoned_delegations later rewrite a succeeded unit as
  "unknown". The write is now try/except: the failure is logged and the
  event is still delivered, so _finalize flips the status and frees the
  slot. A lost durable row is acceptable degradation; a lost result and
  a leaked slot are not.

- _prune_completed_locked treated anything != "running" as finished,
  while the module's own _LIVE_STATES also names stalling/finalizing.
  A stalling record has no completed_at, so it sorted oldest and was the
  first eviction candidate once the retained cap overflowed; its late
  runner return then hit the missing-record path and the real result was
  dropped. The predicate is now `status not in _LIVE_STATES`.

Slim redo of #76606 (earliest fix) and #112031: the converge/shield/
delete-row machinery both PRs built around the write is dropped as
defense-in-depth; the two core hunks are ported as-is.

Fixes #76605
Fixes #112030
Co-authored-by: luckystar2026 <1393268817@qq.com>
2026-09-15 18:43:56 -07:00
teknium1 80f76edfaf fix(web): rescue eligibility asks the provider whether the ring was walked
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
2026-09-15 18:43:05 -07:00
KoNit-K 911eea567f fix(web): rescue failed Nous gateway searches 2026-09-15 18:43:05 -07:00
teknium1 45a4db2225 fix(web): key extract cache on metadata.sourceURL too, pin redirect case
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.

Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
2026-09-15 18:42:38 -07:00
KoNit-K ffd02f37e8 fix(web): cache extracts by returned URL 2026-09-15 18:42:38 -07:00
KoNit-K c9fa191334 fix(tools): support daemon pool workers on Python 3.14
Cherry-picked from #111814. The same feature-detected fix was proposed earlier in
#58699, #65182 and #57459 (final form).

Co-authored-by: nankingjing <76432572+nankingjing@users.noreply.github.com>
Co-authored-by: TheNeuralVault <jdkabattles@gmail.com>
Co-authored-by: gongyi <yigongsyl@gmail.com>
2026-09-15 18:41:46 -07:00
teknium1 8e16bde491 test: fold the schema-retry context test into the output-schema module
Reuse tests/tools/test_delegate_output_schema.py's _StubChild instead of a
new one-test file with its own double; the invariant (the retry turn sees
is_delegated_child_context() True and the flag is restored afterwards) is
unchanged. Trim the source comment to the WHY.
2026-09-15 18:41:19 -07:00
moep90 e002cdb92d fix(delegation): run the schema-retry turn in the delegated-child context
_validate_child_output_schema issues a second run_conversation on the child when
the first answer fails the declared output_schema. The main child turn is wrapped
in delegated_child_context; this one was not. It runs on the parent worker's
thread, where HERMES_KANBAN_TASK is set and nothing marks the execution as a
child, so every identity gate keyed on is_delegated_child_context() fails open.

The visible effect is the kanban stop guard: it nudges the child to call
kanban_complete or kanban_block. A child owns no board task and carries no kanban
toolset, so it cannot, and the nudge text ("do not narrate intent", "finish any
remaining deliverable") displaces the structured answer the retry exists to
produce. The retry then fails the same schema and delegate_task reports an error
for a child whose work was already complete.

Observed with four children, each nudged during its retry:

  [subagent-0] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-2] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-3] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-1] Kanban worker tried to exit without kanban_complete/kanban_block
  4/4 - Final answer does not satisfy the declared output_schema (after 1 retry)

Wrap the retry the same way the main turn is wrapped. The context is entered and
exited around the single call, so nothing outside the retry sees it.

Signed-off-by: moep90 <volleyballlive@googlemail.com>
2026-09-15 18:41:19 -07:00
teknium1 b121a02416 fix(delegate): composite parents can grant their included toolsets to children
_expand_parent_toolsets built the parent's tool surface from each
toolset's declared `tools` only, so a composite parent's `includes` were
invisible: a child of a `debugging` parent (terminal/process_manage +
includes web/file) asking for `file` or `web` was refused, and `safe` /
`hermes-gateway` parents could grant nothing but their own name. Same
root cause as the `_strip_blocked_tools` fix in the previous commit
(#111700, "Related" section).

Both sides of the subset check now use the resolved static surface
(`resolve_toolset(name, include_registry=False)`), so a child may request
any toolset whose real tools the parent genuinely holds, and still never
gains a tool the parent lacks. Candidates that resolve to nothing are not
expanded into (they cannot be a meaningful subset).

Co-authored-by: DresvyanskiyDenis <dresvyanskiydenis@gmail.com>
2026-09-15 18:40:52 -07:00
KoNit-K 3d4c4cc24d fix(delegate): retain composite child toolsets 2026-09-15 18:40:52 -07:00
KoNit-K ae5666f7fc fix(tools): quiet expected unavailable toolsets 2026-09-15 18:40:29 -07:00
teknium1 decf8e3f26 fix(tools): keep an empty required: [] in sanitized tool schemas
`_sanitize_node` deleted the `required` key whenever the pruned list came
out empty. Four built-in tools (skills_list, todo, delegate_task,
session_search) declare `required: []`, so they left the sanitizer with no
key at all. Strict OpenAI-compatible proxies read the missing key as
`null` and 400 the whole request ("null is not of type array"), which is
non-retryable and kills the session on its first call.

An empty array is valid for every backend; the pruning was added (34c3e67)
to drop names that are not in `properties`, not to delete the key. Keep the
key with the filtered list, even when that list is empty.

Fixes #111684
Fixes #59386
Co-authored-by: Cr4ckMe <jiqing.liu@whu.edu.cn>
2026-09-15 18:39:09 -07:00
Konstantin Khlopkov 0724a6a0fd fix(tools): persist the kill outcome when the reader thread finalises first 2026-09-15 18:38:44 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
KoNit-K 0aec64c874 fix(skills): honor profile-scoped readiness secrets 2026-09-15 18:33:15 -07:00
Sahil Vishnalya afe9e25c57 fix(stt): consume lazy whisper segments inside the CUDA→CPU retry guard
faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2
dlopens the CUDA runtime on the FIRST encode, which happens while the segments
are iterated in `_join_confident_segments()` — outside the try/except that
implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an
NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux
`libcublas.so.12`) the model loads fine, the error escapes the guard, and every
voice note fails with "Local transcription failed: Library cublas64_12.dll is
not found or cannot be loaded" until the user pins `stt.local.device: cpu`.

Materialize the segments inside the guarded block (first attempt and CPU retry)
so the dlopen failure reaches the existing evict-and-retry-on-CPU path.

Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class).
Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`,
`cudnn64_`, `cudart64_`) — the Windows message already matches the existing
"cannot be loaded" marker, proven by the live probe with the reporter's exact
string; the 6-test file was reduced to 2 invariant tests in the existing suite.

Fixes #111929
Fixes #105295
Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate)

Co-authored-by: atmaksri <sri.atmakur@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>
2026-09-15 18:25:17 -07:00
KoNit-K 065bc91846 fix(checkpoints): report legacy archive deletion failures 2026-09-15 18:24:22 -07:00
teknium1 f9c3a8a186 feat: hermes update and doctor tell you when /rollback checkpoints are on and large
Checkpoints were enabled by default from 9e845a6e (2026-03-16) until #20709
(2026-05-06) flipped the default back to off; the migration of that window
wrote `checkpoints.enabled: true` into user configs, where a later default
flip cannot reach it. Users who never type /rollback have carried a GB-scale
`~/.hermes/checkpoints/store` since (one live install: 1.2 GB across 250
projects, mostly disposable worktrees and /tmp dirs), and the cap cannot
bring it down because every project keeps at least one snapshot.

Silently flipping the key back is indistinguishable from overriding a real
opt-in, so this surfaces it instead: `checkpoint_footprint_notice()` returns
one line when checkpoints are enabled AND the store is at or above
`max_total_size_mb`, naming the opt-out (`hermes config set
checkpoints.enabled false` + `hermes checkpoints clear`) and the
retention knob. `hermes update` prints it with the post-update notices;
`hermes doctor` reports it as a warning after the state.db check.
2026-09-15 12:06:08 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1 804707bea6 fix: checkpoint store gc never runs inside a tool call or gateway startup
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.

The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.

- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
  drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
  pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
  so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
  calls it from the housekeeping tick (last chore), the CLI from a daemon
  thread. Nothing on either startup path waits for git.

Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
2026-09-15 10:57:16 -07:00
teknium1 d84ece48b8 fix(mcp): Figma OAuth login completes despite the omitted iss parameter
Figma's authorization-server metadata advertises
authorization_response_iss_parameter_supported and its redirect omits iss,
so the mcp SDK's RFC 9207 check discarded every valid code and login never
finished. For that one issuer the provider fills a missing iss with the
discovered issuer and warns; a mismatching iss still fails and every other
server keeps the strict rule.

Fixes #111135
2026-09-15 09:29:00 -07:00
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 14bd358f5e fix: unlink the dm plaintext file too once a live delivery settled
The live branch of _run_delivery returns from _wait_live_dm before the
try/finally that removes <dm>.txt, so a settled delivery dropped its
.live.json intent but left the sibling .txt — the same plaintext — on disk.
_wait_live_dm now takes the dm file and removes both on settled; pending
and failed outcomes still keep both for the retry.

Review finding: settled live path unlinked <dm>.live.json but the sibling <dm>.txt survived.
2026-09-15 06:34:25 -07:00
John Paul Soliva 1238dbc09d fix(bot-mode): live-delivery intents are swept, and dropped once the owner settled
A live DM's intent file (<dm file>.live.json — owner, delivery id and the
message plaintext) has to outlive its runner so a retry replays the same
delivery id instead of minting a second one. Nothing ever removed it: the
runner unlinks only the dm file, and cleanup_bot_dm_cache sweeps only *.txt,
so every live DM left its plaintext in the DM cache directory indefinitely.
Now the sweep reaps intents past the stale cutoff, and a delivery the owner
settled drops its intent immediately — nothing retries a settled delivery.
2026-09-15 06:34:25 -07:00
teknium1 4eb69cbce5 fix(bot-mode): Bot Chat identity + one-shot DM transport apply the silence rule; trim tests
Follow-up to the salvaged #110786 commit:

- `_is_bot_mode_session` mirrors the system-prompt gate (`agent._session_title_hint`
  first, then the live DB title) instead of reading `pending_title`/`title` off the
  session dict: `pending_title` is cleared after turn 1 and the record never carries
  `title`, so the contributor's gate matched only the very first Bot Chat turn.
- `tools/bot_mode_dm.py::_run_local_turn` (the `hermes -p X chat -c "Bot Chat" -Q`
  transport behind `message_agent` when no live owner holds the target) re-emits ""
  for a successful bare marker — the third delivery path of the same class.
- Tests trimmed to one invariant per surface (live completion, relay RPC, one-shot
  transport), each proven red on origin/main sources.
- Bot Mode docs gain a "Staying silent" line pointing at the shared token list.
2026-09-15 06:31:31 -07:00
teknium1 14ebd48c64 fix: derive the lost-bus check from the env the scoped worker was spawned with
scoped_spawn_lost_user_bus() re-derived the user bus from an empty base env,
so it only ever looked under /run/user/<uid>. The worker itself is launched
with systemd_user_bus_env(worker_env), which honours a configured
XDG_RUNTIME_DIR. On a host whose bus lives outside the default runtime dir,
any unrelated systemd-run exit was therefore misread as "bus gone": the job
error named a missing bus that was still there, the cached scope verdict
flipped to False, and the next 60s of cron fires dispatched without cgroup
isolation.

The check now takes the spawn env, drops the bus address the spawn already
carried, re-derives from that, and decides on the DBUS_SESSION_BUS_ADDRESS
key rather than on dict truthiness (a non-empty env with no bus was never
"bus present").

Review finding: scoped_spawn_lost_user_bus used systemd_user_bus_env({}) instead of the worker's spawn env, so a configured XDG_RUNTIME_DIR made unrelated wrapper exits flip the scope cache to unscoped dispatch.
2026-09-15 06:28:07 -07:00
teknium1 a313a211d7 fix(cron): a scoped worker whose user bus vanished names the cause and re-probes
One TTL for both probe verdicts (the success-TTL constant collapses into
`_SYSTEMD_SCOPE_PROBE_TTL_SECONDS`), and the ack-wait loop consults
`scoped_spawn_lost_user_bus()` when a scoped dispatch exits before the
worker acknowledged: with `/run/user/<uid>/bus` gone the job error names
the missing bus and the enable-linger remedy instead of the wrapper's bare
`exit 1`, and the cached True flips so the next fire degrades to a direct
external subprocess rather than consuming another occurrence on a dead
wrapper. Contributor test trimmed to the revalidation invariant.
2026-09-15 06:28:07 -07:00
KoNit-K c143ec4d88 fix(tools): revalidate systemd scope availability 2026-09-15 06:28:07 -07:00
shehjaddev d205cef418 fix(redact): mask assignments in secret-bearing file reads
read_file/search_files passed file_read=True, which folded into code_file=True and skipped
the ENV/JSON/YAML assignment passes, so an opaque prefix-less credential under a
credential-shaped key reached the model in cleartext from a secret-bearing file — the
file-read half of the #110228 gate (#110567).

Two defects on that path, both fixed here:

- The rendered line-number gutter ("5|      ADS_API_TOKEN: ..." from read_file,
  "6:      ADS_API_TOKEN: ..." from grep -n / cat -n) defeated the line-anchored patterns,
  so the real rendered read leaked exactly what the raw text masked. A gutter-free fixture
  cannot see this, which is why the tool-level tests carry the real render shape.
- _is_secret_file_arg() could not see the RESOLVED Hermes home: the default home's basename
  is an installation detail (".hermes" on POSIX, "hermes" under AppData/Local on Windows) and
  a resolved path never spells $HERMES_HOME, so the managed Windows home's config.yaml was
  classified as ordinary YAML on both the file-read and the terminal surface.

Changes:

- redact_sensitive_text(): secret_file= re-enables the assignment passes for content the
  caller classified with _is_secret_file_arg, keeping code_file behaviour everywhere else.
  It is authoritative over code_file, so a caller cannot be fail-open on the security flag
  by setting both.
- _redact_assignments(): mask_nonreusable selects the non-reusable sentinel for file reads,
  so the #35519 write-back hazard stays closed.
- _should_redact_assignment(): no longer re-masks an already-masked value, which was erasing
  the vendor label the sentinel deliberately keeps.
- _is_secret_file_arg(): consult the resolved Hermes home for the config.yaml arm.
- _CFG_ANCHORED_RE / _YAML_ASSIGN_RE: tolerate a rendered line-number gutter.
- file_tools.py: classify the resolved path at all three file-read call sites.

Closes #110567
2026-09-15 06:26:29 -07:00
teknium1 6f24245532 fix(kanban): gate create-with-parents like link; archived parent is terminal
create_task(parents=[open parent]) — the reporter's actual incident path —
parked the card in todo with only a `created` event, and kanban_create's
payload carried no `gated`, so the board still showed an unexplained todo
while only the link surface was fixed. create_task now appends the same
dependency_wait {reason: parent_not_done, parent} event and kanban_create
returns gated/gated_by, mirroring kanban_link.

link_tasks gated on `status != 'done'`, but _parents_satisfied and
recompute_ready treat `archived` as terminal: linking a ready child under an
archived parent demoted it to todo with a false parent_not_done event and the
next recompute promoted it straight back. Gate on not in ('done','archived').

Review finding: create_task(parents=...) emitted no dependency_wait/gated; link under an archived parent flapped ready->todo->ready with a false reason.
2026-09-15 06:25:42 -07:00
Konstantin Khlopkov 35b1609fc3 fix(kanban): surface the link-time demotion of a ready child to todo
A ready child linked under an unfinished parent drops to todo with no
event and no operator signal; the only trace used to be claim_rejected
after a forced promote. Record a dependency_wait event when the demotion
fires, return the gate from link_tasks, warn in the CLI link command,
report gated in the kanban_link tool, and document the gate.
2026-09-15 06:25:42 -07:00
teknium1 95987fb85a fix: bypass the proxy on loopback HTTP CDP discovery and keep NO_PROXY=* intact
The websockets dials got proxy=None but the three HTTP /json/version dials
(CDP override discovery, is_browser_debug_ready used by Lightpanda and the
real-profile readiness check, and surviving-Chrome detection) still resolved
via getproxies(), so under a system/env proxy discovery fell back to the raw
http:// URL, readiness never fired and /browser connect reported not ready.
loopback_request_kwargs() sits next to loopback_connect_kwargs() and is used
at all three sites (ProxyHandler({}) opener for the urllib one).

add_loopback_no_proxy turned an operator NO_PROXY=* into '*,127.0.0.1,...',
which urllib/requests no longer treat as the wildcard, flipping bypass-all
configs into proxy-all. A wildcard in either casing now leaves env untouched.
is_loopback_host also accepts any loopback IP literal (127.x, ::ffff:127.0.0.1).

Review finding: HTTP /json/version dials still proxied loopback; NO_PROXY=* wildcard broken by append.
2026-09-15 06:16:38 -07:00
teknium1 8f6f92d901 fix(browser): one loopback proxy-bypass helper covers child envs and in-process CDP dials
Move the loopback NO_PROXY merge from browser_tool into agent/proxy_bypass.py (the
module that already owns NO_PROXY semantics) and reuse no_proxy_entries() so comma-
and whitespace-separated operator values are both preserved. Add
loopback_connect_kwargs() and pass proxy=None on the two in-process websockets
dials to loopback CDP endpoints (browser_cdp_tool._cdp_call, BrowserSupervisor._run):
those never see the child env, so the env merge alone left them routed through a
macOS system proxy. Remote CDP URLs keep the default proxy behaviour.

Tests trimmed to two invariants: the built child env appends loopback to an
operator NO_PROXY in both casings, and only loopback URLs get proxy=None.
Sibling helper in tools/browser_use_cli (#110570) is redundant once the shared
env carries the entries.
2026-09-15 06:16:38 -07:00
liuhao1024 00a68a8768 fix(browser): bypass proxies for loopback hosts in browser child envs
websockets>=14 defaults to proxy=True and resolves proxies via
urllib.request.getproxies(), which reads the macOS/Windows system
proxy config even with no *_proxy env vars set. Local CDP endpoints
(ws://127.0.0.1:<port>/devtools/...) were therefore dialed through
the system proxy and the handshake failed with "did not receive a
valid HTTP response" (#110565).

Append 127.0.0.1/localhost/::1 to NO_PROXY/no_proxy (both casings)
in _build_browser_env so every browser subprocess (Browser Use CLI,
agent-browser, Chromium, Lightpanda) bypasses proxies for loopback.
Operator-provided NO_PROXY entries are preserved.

Fixes #110565
2026-09-15 06:16:38 -07:00
Kevin Rajan 7d292af875 fix(journey): show foreground-created skills in /journey via learn provenance marker
record_created now stamps created_by="learn" on foreground creates
(e.g. /learn) instead of leaving it unset, and the learning-graph filter
honors "learn" alongside "agent"/used. "learn" is a learning-signal
marker only: curator management stays keyed strictly on "agent"
(_is_curator_managed_record), so user-taught skills appear in /journey
without becoming eligible for autonomous curation.

Fixes #111317.

---
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; the contributor reviewed the diff and ran the tests.
2026-09-15 05:38:31 -07:00
teknium1 8ecc3dd71e fix(stt): treat a partial whisper cache as a cache miss
An interrupted first download leaves refs/main plus a snapshot folder
without model.bin. snapshot_download(local_files_only=True) returns that
folder rather than raising LocalEntryNotFoundError, so ctranslate2 fails
with RuntimeError "Unable to open file 'model.bin'" and the online path
never ran, leaving STT permanently broken with a misleading error. Fall
through to the download when the local attempt fails that way. The
loading tests are also gated on faster_whisper being installed, so the
module collects when the voice extra is absent.

Review finding: partial-cache RuntimeError bypassed the online fallback.
2026-09-15 05:36:37 -07:00
teknium1 0bb460b488 fix(stt): tolerate a missing huggingface_hub when checking the whisper cache
The cache-first loader imported LocalEntryNotFoundError unconditionally; in
environments without the optional huggingface_hub dependency every local
transcription raised ModuleNotFoundError. Resolve the cache-miss exception
lazily and fall back to its OSError base class when the package is absent.
2026-09-15 05:36:37 -07:00
teknium1 489ad83907 fix(stt): only rewrap Hub/network failures on the whisper download path
The salvaged fallback caught bare Exception around the online load, which
would relabel a CUDA runtime error or an invalid model size as a download
problem and defeat the CUDA → CPU fallback above it. huggingface_hub raises
every network/Hub failure as an OSError subclass (LocalEntryNotFoundError,
HfHubHTTPError), so catch that class only. The negative test now feeds the
real LocalEntryNotFoundError('Got: ConnectTimeout ...') shape the reporter
saw instead of a synthetic RuntimeError.
2026-09-15 05:36:37 -07:00
fangliquan ebb6dc6e70 fix(stt): prefer cached local whisper models 2026-09-15 05:36:37 -07:00
teknium1 bae9f8ab85 fix(file-sync): sweep stale sync-back dirs too and tighten the stale window to 6 h
The salvaged commit only swept `hermes-sync-back-*.tar`. The extraction
staging dir (`tempfile.TemporaryDirectory(prefix="hermes-sync-back-")`) is
leaked by the same hard kill, so the sweep now reclaims both shapes and
returns the count. The window drops from 24 h to 6 h (the reporter's
value): archives appear every few minutes on a busy gateway and a live
transfer is never hours old. Tests trimmed to two invariants: the sweep
touches only stale prefixed entries, and a real sync_back reclaims a
leaked archive while writing its own tar under the identifiable prefix.
2026-09-15 05:35:34 -07:00
KoNit-K 6a03d5a94c fix(file-sync): clean stale sync-back archives 2026-09-15 05:35:34 -07:00
Hukla 05e7e89175 fix(approval): honor pattern-key allowlists when unattended 2026-09-15 05:33:41 -07:00
fangliquan 7851d1d3ea fix(skills): ignore package-owned legacy markdown 2026-09-15 05:32:14 -07:00
teknium1 5117e3b3a0 fix: key the check_fn cache by the same served-profile predicate as the MCP registry scope
_mcp_registry_scope() became profile-keyed for served profiles with the
multiplex flag off, but check_fn_cache_scope() still returned None in that
mode, so the process-wide availability cache stayed keyed (fn, None) across
profiles. A served profile whose mcp__x__* check_fn now correctly resolves to
its own (absent) connection cached False for the TTL window and the launch
profile that owns the live connection lost its tools for that window.

Both sites now call one helper, agent.secret_scope.serves_routed_profile()
(multiplex on, or a HERMES_HOME override naming a home other than the
process home), so the registry scope and the cache key can no longer drift.

Review finding: served profile B's check_fn verdict shadowed the launch profile's live mcp tools via the unscoped check_fn cache.
2026-09-15 04:56:40 -07:00