_detect_tool_failure now classifies dict results as failures, so the
concurrent worker's failure log line sliced result[:200] on a dict and
raised TypeError; the worker died and the model saw "thread did not
return a result" instead of the tool's own error payload. Stringify the
preview like the sequential path does.
agent/kanban_stop.py::kanban_stop_nudge_enabled tested only HERMES_KANBAN_TASK,
which in-process delegate_task children (and cron runs fired inside a worker)
inherit from the worker's process environment. Those executions own no board
task and have the kanban toolset withheld, so the turn-end nudge ordered them to
call a tool they cannot reach — burning attempts, and in production driving
children to complete the parent's card through the CLI.
Gate on agent/delegation_context.py::is_dispatcher_owned_worker_context, the
predicate every other HERMES_KANBAN_* identity gate already uses. The real
worker and the HERMES_KANBAN_STOP_NUDGE opt-out are unchanged.
Salvaged from PR #84656 by @jerryhjones (re-applied onto the current facade
shape); the same gate was first proposed in PR #80023 by @webdevfrancisco
using the narrower delegated-child predicate.
Co-authored-by: webdevfrancisco <franciscombautista2015@gmail.com>
The concurrent completion line logged len(result) directly, so a native-path
vision_analyze envelope dict reported "4 chars" — its key count — while the
sequential path already logs the serialized length. Mirror the sequential
measurement so parallel multimodal calls stop looking truncated in logs.
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves#76593; supersedes the config-key
approach of #87835.
Pinned vision ids rot silently (glm-5v-turbo was retired from the Coding Plan endpoints
while still valid on pay-as-you-go), so name the data behind the new pin — the only
image-capable GLM id on every Z.AI surface — and why the pin cannot simply be dropped in
favour of ProviderProfile.default_vision_model() (ZaiProfile returns None, which would route
vision to a text-only chat model).
Co-authored-by: POWERFULMOVES <142271328+POWERFULMOVES@users.noreply.github.com>
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.
Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.
Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.
Refs #111999
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.
- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
(mtime_ns, size) snapshot; a later success clears every entry with the same
identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
entries whose file changed since the failed call, so a receipt-less
mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
way the file tools resolved them.
Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.
Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.
The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.
`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.
CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.
Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
After a mid-turn /model switch while a stream was stalled, the streaming
retry loop re-sent the request it had captured at construction time. That
payload still named the OLD model, but every stream (re)open builds its
request client from the LIVE agent, so the new provider's base_url received
a foreign model slug: 404 "Not found the model ...", then the turn sat in
the provider's rate-limit hold (#112121).
_StreamingCall now records the route (model, provider, base_url, api_mode)
its api_kwargs were built for. When a retry is about to be issued and the
live route differs, the streamer stops and hands the transient error back
to the turn loop instead. The turn loop already rebuilds the request per
attempt for the CURRENT route (turn_api_request.build_api_request: model,
wire shape, prompt-cache decoration, provider request overrides), so a
re-keyed model alone would still have shipped a payload shaped for the old
provider. Non-streaming requests have no in-process retry, and the
fallback / restore-primary paths go through the same turn-loop rebuild, so
this is the only site that replayed a captured route.
Fixes#112121
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
`_extract_pricing`'s generic path copied catalog values verbatim, while
usage_pricing unconditionally applies OpenRouter's per-token convention and
multiplies by 1e6. A provider quoting USD per 1M tokens (Neosantara 0.6/M,
Crof cost.input 0.04/M) or declaring `unit: per_1m_tokens` therefore priced
at $600,000/M and corrupted estimated_cost_usd in state.db and every cost
report summing across providers.
Normalize at the producer, where Novita/DeepInfra unit handling already
lives: an explicit `unit` beside the rates wins (per_token / per_1k_tokens /
per_1m_tokens); without one, a token rate at or above $0.001 per token
($1,000/MTok - no real model) can only be a per-million quote. Output keeps
the per-token-string contract, so the consumer is untouched; `request` fees
and per-token catalogs pass through unchanged.
The $0.001/token magnitude threshold is the one proposed in #34263 by
@Bartok9 (earliest fix); #112036 by @kvnloo proposed the same heuristic at
the consumer.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.
Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.
Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
Follow-up to the two salvaged commits (#111777, #111781 by @KoNit-K):
- agent/prompt_builder.py::_truncate_content — with queue_warning=False the
logged line no longer tells the operator to "pin a larger
context_file_max_chars, or use a larger-context model": the subdirectory
hint cap is a constant neither knob raises. It now points at the read_file
recovery the marker already discloses.
- tests/gateway/test_startup_environment_probe.py — replace the
call-detection test with the behavioural invariant: an oversized SOUL.md in
HERMES_HOME and a warm-up leave the truncation-warning queue empty for the
next default-executor task (the api_server turn path runs on that executor
without copy_context, which is how the boot warning reached a foreign
session).
- tests/agent/test_subdirectory_hints.py — fold the new drain assertion into
the existing oversized-hint test (same fixture) and pin that the log carries
no context_file_max_chars advice.
- agent/AGENTS.md, website/docs/.../context-files.md — the hint cap is 32,000
(docs said 8,000) and is fixed; document that it is logged, not surfaced as a
chat warning.
Drop the raw JSONDecodeError belt from _is_provider_stream_empty_frame_error:
every chat stream iteration goes through _iter_provider_stream_chunks, which
already translates the decode failure, so the belt guarded a path that does
not exist. Drop the explicit classifier entry for the new code: unknown codes
already resolve to the same retryable unknown verdict (verified live: identical
ClassifiedError with and without the entry).
An empty `data:` frame (or a frame carrying only `event:` / `id:`) is a legal SSE
no-op, but the OpenAI SDK still hands it to `json.loads`, which raises
`JSONDecodeError` with an empty document. The streaming helper translated that
into `ProviderStreamError(provider_stream_non_json_data)`, nothing recognised it
as recoverable, and the main loop retried streaming -- identically -- until the
retry budget ran out: "API call failed after 3 retries: Provider stream returned
non-JSON SSE data". A degraded gateway answers EVERY streaming request that way,
so the retries only repeated the failure while the same request sent
non-streaming succeeded seconds later.
An empty document means no payload, which is a different fact from a malformed
payload: give it its own code, and when it appears before any delta switch the
session to non-streaming (the existing `_disable_streaming` mechanism, already
used for "stream not supported", Bedrock IAM denials and adapter-returned final
responses) so the retry goes out on a channel the degraded gateway can answer.
Non-empty payloads keep today's fatal semantics; failures after deltas keep the
existing stream-drop handling.
Tests: the turn-level recovery test drives the real SDK decoder over a real
httpx response and is red on base (three identical streaming attempts, turn
lost); the boundary test pins the malformed-payload path so a future "ignore bad
frames" change cannot swallow genuine provider errors.
The openrouter branch of _seed_from_env returns before the generic loop,
so a persisted `env:OPENROUTER_API_KEY_2` row stayed empty exactly like
the registry-provider case #103067 fixed. Fold the "declared vars + on-disk
env rows" union into one helper both branches use.
Restores the env-source seeding loop in _seed_from_env that was dropped
by upstream refactors. Without this, any env-source entry whose env var
name isn't in the provider registry's hardcoded api_key_env_vars tuple
stays empty forever, gets filtered out of rotation by _available_entries,
and round-robin silently degrades to single-key behavior.
Forward-port of ba71e00db07c6263ee8d44b27dfbce2a92e6b39c onto current main
f1ccf436a2. Narrow insertion after final
env_vars list in _seed_from_env: scan only existing entries whose source
begins with env:, extract/dedupe the named variable, and let the existing
seeding path hydrate it via get_env_prefer_dotenv (preserving _Seeder,
secret-scope/profile isolation, borrowed-secret sanitization, and
suppression behavior).
Tests: 10 new tests in test_credential_pool_seed_existing_env_sources.py
covering three-key hydration, round_robin cycling, dotenv/secret-scope
resolution without os.environ bypass, duplicate dedupe, manual rows
untouched, unset remains unavailable, and sanitized persistence.
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.
Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).
Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
The "says something else" guard reuses the billing table, which carries the
Nous gateway's own free-tier phrases ("not available on the free tier",
"model_not_supported_on_free_tier"). On the welcome route those words mean
exactly "the tier refused"; routing them to billing prints a credits check to
an anonymous session that has none. Leave those two phrases on the
tier_disabled path.
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
matches neither the content-policy nor the billing patterns, so a safety
refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
overwrites a record that has one (the loop racing the user's click).
Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
_welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
of freezing whole sentences.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A welcome-tier 403 classifies as auth_permanent, so the desktop's error
surface mapped it to "Your Nous Portal sign-in expired" with a Nous Portal
re-login button — the chat sentence never reached the user. Terminal results
on the free route now carry a structured free_tier block (kind + the chat
sentence); agent/error_surface.py turns it into a free_tier_<kind> code on
the provider layer with the sentence as `message`. The desktop gives those
codes their own titles, shows the backend sentence as the body, and offers
"Sign in with a Nous account" (the free-tier dialog) instead of the OAuth
re-login, with Retry only where a later send can succeed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The cherry-picked commit extends the config/profile/MCP/managed-scope/completer/
OAuth/skills-manifest signatures. This commit finishes the class and trims it:
- `file_signature()` lives in `utils.py` next to the other stat/metadata helpers
instead of `hermes_cli.managed_scope` (gateway/ and agent/ callers no longer
reach into the managed-scope module for a generic stat helper).
- `hermes_cli/config_effective.py` was left comparing 2-/4-wide prefixes against
the widened `_RAW_CONFIG_CACHE` / `_load_config_cache_sig` records, so
`load_user_config_effective()` re-parsed on every call (3 parses for 3 calls on
an unchanged file, 1 before); index by the new widths.
- Sibling caches keyed on the same (mtime, size) shape and reading the SAME files
now use the helper: `load_env()` memo, `agent/skill_utils` raw-config and
external-dirs caches, `hermes_cli/model_switch` alias identity, `agent/moa_loop`
preset stamp, `hermes_cli/auth` global auth-store memo.
- Tests trimmed to one invariant each (pinned-mtime replacement invalidates; an
unchanged file still hits), both red on origin/main.
Left alone on purpose: `tools/registry.py`, `tools/skills_tool_dedup.py`,
`gateway/status.py`, `hermes_cli/banner.py`, `hermes_cli/main.py`,
`hermes_cli/session_recovery.py` — those fingerprint source files, PID/lock files
or write to persisted on-disk caches shared across processes, where an inode/ctime
key would churn on every checkout/copy rather than catch a replaced config.
Config caches, the profile re-scan watcher, the MCP reconciler, managed-scope
reads, the completer memo, OAuth heal marks, and the skill-manifest snapshot
keyed change detection on (st_mtime_ns, st_size) only, so a replacement that
preserved both (cp -p, rsync -t, timestamp-pinning scripts, sync clients) was
treated as unchanged and stale values were served until restart.
Add st_ino (fresh inode on atomic replace) and st_ctime_ns (cannot be
backdated via os.utime) to every signature via a shared
hermes_cli.managed_scope.file_signature() helper.
Fixes#111105
The optional line-number gutter was written as `^(?:[ \t]*GUTTER)?[ \t]*`,
stacking two adjacent whitespace runs so _CFG_ANCHORED_RE / _YAML_ASSIGN_RE
went quadratic on any indented line with a secret keyword and no `=` (2s per
5k spaces, 50s at 20k) — and these run on every terminal output and file read.
The gutter now carries the single optional group and callers own the one
leading whitespace run. It also accepts `cat -n` / `nl` (number + TAB) and
`grep -A/-B/-C` (`5-`) gutters, which the comment claimed but leaked.
The strong-key branch masked SSH_AUTH_SOCK=$HOME/..., DOCKER_AUTH_CONFIG=/...
on shell rc reads, leaving the agent unable to edit them; a value starting with
$, / or ~ is a variable or path reference and is now kept unless the key is
password-class. _is_secret_file_arg only extends config.yaml to the
backups/config .good./.corrupt. copies, not config.yaml.pdf/.bak.
Review finding: quadratic gutter regex; cat -n/nl TAB and grep context gutters leaked; over-redaction of rc path references.
search_files over $HERMES_HOME returned the token from backups/config/config.yaml.good.<stamp>
in cleartext while config.yaml itself was masked; the predicate only matched the exact basename.
Same predicate serves the terminal side, so `cat backups/config/config.yaml.good.*` is covered too.
read_file/search_files passed file_read=True, which folded into code_file=True and skipped
the ENV/JSON/YAML assignment passes, so an opaque prefix-less credential under a
credential-shaped key reached the model in cleartext from a secret-bearing file — the
file-read half of the #110228 gate (#110567).
Two defects on that path, both fixed here:
- The rendered line-number gutter ("5| ADS_API_TOKEN: ..." from read_file,
"6: ADS_API_TOKEN: ..." from grep -n / cat -n) defeated the line-anchored patterns,
so the real rendered read leaked exactly what the raw text masked. A gutter-free fixture
cannot see this, which is why the tool-level tests carry the real render shape.
- _is_secret_file_arg() could not see the RESOLVED Hermes home: the default home's basename
is an installation detail (".hermes" on POSIX, "hermes" under AppData/Local on Windows) and
a resolved path never spells $HERMES_HOME, so the managed Windows home's config.yaml was
classified as ordinary YAML on both the file-read and the terminal surface.
Changes:
- redact_sensitive_text(): secret_file= re-enables the assignment passes for content the
caller classified with _is_secret_file_arg, keeping code_file behaviour everywhere else.
It is authoritative over code_file, so a caller cannot be fail-open on the security flag
by setting both.
- _redact_assignments(): mask_nonreusable selects the non-reusable sentinel for file reads,
so the #35519 write-back hazard stays closed.
- _should_redact_assignment(): no longer re-masks an already-masked value, which was erasing
the vendor label the sentinel deliberately keeps.
- _is_secret_file_arg(): consult the resolved Hermes home for the config.yaml arm.
- _CFG_ANCHORED_RE / _YAML_ASSIGN_RE: tolerate a rendered line-number gutter.
- file_tools.py: classify the resolved path at all three file-read call sites.
Closes#110567
The salvaged fix added a second directory tuple and a second loop for the same
predicate. _HERMES_PROTECTED_SUBPATHS already expresses "HERMES_HOME subpaths the
file tools must not rewrite", so the two secret directories join it and the
duplicate loop goes. Tests trimmed to two invariants: every read-denied secret
store is also write-denied on both the profile and the global root, and the
#45947 control files plus same-named paths outside HERMES_HOME stay writable.
Narrow the fix to secret material and stop reverting #45947.
The earlier revision added every read-denied name to the write denylist,
which re-blocked auth.json and webhook_subscriptions.json and rewrote the
test guarding them. #45947 freed those control files deliberately:
containment belongs in Docker/remote backends and OS permissions, not an
expanding hardcoded denylist.
What #45947 kept blocked is secret material, and that list had drifted:
auth/google_oauth.json (OAuth token store), the plaintext Bitwarden cache,
vault/ (key + ciphertext side by side) and browser-profile/ (copied
cookies / Login Data) were writable via write_file / patch.
- Write-deny those four; control files stay writable and read-denied.
- _WRITE_DENIED_SECRET_DIRS is its own tuple rather than _READ_DENIED_DIRS,
so a future read-only convenience deny cannot silently become a write deny.
- tests/tools/test_write_deny.py is restored unchanged from main.
- New tests pin both halves: secrets denied, control files writable, and
write denies stay a subset of read denies.
Fixes#110464
Read guards already refuse auth.json, webhook HMAC secrets,
google_oauth.json, the Bitwarden plaintext cache, vault/, and
browser-profile/. The write denylist only covered .env, the Anthropic
PKCE store, and bws_cache.enc.json, so write_file/patch could replace
the rest.
Drive the write denylist from the same credential names and
read-denied directories. Keep the extra write-only entries
(bws_cache.enc.json, state.db, sessions, pairing).
Fixes#110464
Related: #108716
NVIDIA NIM's Rust gateway 400s on list-type tool message content with a serde
error that names the enum rather than the field: "data did not match any
variant of untagged enum ChatCompletionRequestToolMessageContent". None of the
existing _MULTIMODAL_TOOL_CONTENT_PATTERNS match that wording, so the error
classifies as format_error (non-retryable) and the existing recovery —
downgrade the image-bearing tool message to text, remember (provider, model),
retry once — never fires: every retry re-sends the same shape and fails
identically in a session bound to that model.
Add the serialized enum name to the pattern tuple so the NVIDIA wording routes
to multimodal_tool_content_unsupported like the MiMo/Alibaba/Console Go
wordings already do (#111231).
The websockets dials got proxy=None but the three HTTP /json/version dials
(CDP override discovery, is_browser_debug_ready used by Lightpanda and the
real-profile readiness check, and surviving-Chrome detection) still resolved
via getproxies(), so under a system/env proxy discovery fell back to the raw
http:// URL, readiness never fired and /browser connect reported not ready.
loopback_request_kwargs() sits next to loopback_connect_kwargs() and is used
at all three sites (ProxyHandler({}) opener for the urllib one).
add_loopback_no_proxy turned an operator NO_PROXY=* into '*,127.0.0.1,...',
which urllib/requests no longer treat as the wildcard, flipping bypass-all
configs into proxy-all. A wildcard in either casing now leaves env untouched.
is_loopback_host also accepts any loopback IP literal (127.x, ::ffff:127.0.0.1).
Review finding: HTTP /json/version dials still proxied loopback; NO_PROXY=* wildcard broken by append.
Move the loopback NO_PROXY merge from browser_tool into agent/proxy_bypass.py (the
module that already owns NO_PROXY semantics) and reuse no_proxy_entries() so comma-
and whitespace-separated operator values are both preserved. Add
loopback_connect_kwargs() and pass proxy=None on the two in-process websockets
dials to loopback CDP endpoints (browser_cdp_tool._cdp_call, BrowserSupervisor._run):
those never see the child env, so the env merge alone left them routed through a
macOS system proxy. Remote CDP URLs keep the default proxy behaviour.
Tests trimmed to two invariants: the built child env appends loopback to an
operator NO_PROXY in both casings, and only loopback URLs get proxy=None.
Sibling helper in tools/browser_use_cli (#110570) is redundant once the shared
env carries the entries.
Kanban cards have no length limit, but the session title store rejects
titles past SessionDB.MAX_TITLE_LENGTH with ValueError. _persist_session_title
reads that as a unique-title collision, retries with a "#N" suffix (longer
still), and the caller suppresses the second failure - so a worker spawned on
a >100-char card ended up with no title at all, where main at least gave it a
derived one. Trim the card title (with room for the "#N" retry suffix) before
persisting; a retried card now gets "<trimmed> #2" within the cap.
Review finding: >100-char card title left the kanban worker session untitled.
Follow-up to the salvaged commit from #111169 (@KoNit-K):
- Drop HERMES_KANBAN_TASK_TITLE. The worker already has HERMES_KANBAN_TASK
and HERMES_KANBAN_BOARD/HERMES_KANBAN_DB pinned in its env, so
maybe_auto_title reads the card title from the board itself (no new
HERMES_* env var for non-secret config; the dispatcher and the
delegation scrub list stay untouched).
- Unreadable or missing card: the session is named `Kanban task <id>`
with zero auxiliary calls (the fallback the issue asked for; the
#109743 seed left such workers untitled).
- The card title persists at `llm` authority via set_auto_title, so a
manual /title still wins and the upgrade thread never starts.
- Tests trimmed to two invariants against a real board + SessionDB
(card title, unreadable-card fallback), both red on origin/main.
The mid-turn /steer marker is delivered as a standalone role:"user"
message right after the newest tool result (steer_user_row /
apply_pending_steer_to_tool_results). But STEER_CHANNEL_NOTE (the
model-facing briefing) and the module comment still said Hermes
'appends their message to the end of a tool result'.
That stale claim briefs the model to expect the marker INSIDE tool
output, so a real standalone user row carrying the marker can read as
off-channel — the under-trust half of #110979. Correct the briefing and
comment to describe actual delivery; add a contract test tying the note
wording to the steer_user_row delivery mechanism (proven red on the old
wording).
record_created now stamps created_by="learn" on foreground creates
(e.g. /learn) instead of leaving it unset, and the learning-graph filter
honors "learn" alongside "agent"/used. "learn" is a learning-signal
marker only: curator management stays keyed strictly on "agent"
(_is_curator_managed_record), so user-taught skills appear in /journey
without becoming eligible for autonomous curation.
Fixes#111317.
---
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; the contributor reviewed the diff and ran the tests.
The first compaction of a session runs check_compression_model_feasibility()
inside compress_context() BEFORE _announce_compression_start(). That probe is
network-bound (live model catalog / provider lookups; connect timeouts stack up
through proxies and slow remote gateways), so every automatic entrypoint that
reaches compress_context without its own pre-emit — the post-tool gate in
agent/turn_preflight.py::compress_after_tool_results, overflow recovery,
manual /compress — left the client with no `kind="compacting"` status for the
whole probe. On Desktop that is a bare working-row spinner with no
"Summarizing thread" label (#111294).
Move the announcement ahead of the probe in the one choke point so every
caller is covered, and retire the announced phase (force_terminal) when the
probe's hard rejection propagates so the client never stays "compacting" for
an attempt that never started.
Live repro: /tmp probe driving compress_after_tool_results with a 2 s
feasibility stand-in — before: first compacting status at t=2.29 s (after the
block); after: t=0.13 s.
_mcp_registry_scope() became profile-keyed for served profiles with the
multiplex flag off, but check_fn_cache_scope() still returned None in that
mode, so the process-wide availability cache stayed keyed (fn, None) across
profiles. A served profile whose mcp__x__* check_fn now correctly resolves to
its own (absent) connection cached False for the TTL window and the launch
profile that owns the live connection lost its tools for that window.
Both sites now call one helper, agent.secret_scope.serves_routed_profile()
(multiplex on, or a HERMES_HOME override naming a home other than the
process home), so the registry scope and the cache key can no longer drift.
Review finding: served profile B's check_fn verdict shadowed the launch profile's live mcp tools via the unscoped check_fn cache.
Binding every saved URI made a URI the user explicitly set to "Never"
match (match=5) a valid fill target — wider than the manager's own policy.
Filter those out; all other URIs keep exact-origin semantics.
Review finding: match=5 (Never) URIs became fill origins.