Commit Graph

4911 Commits

Author SHA1 Message Date
teknium1 1d14418ea2 fix(agent): concurrent worker survives a dict error result
_detect_tool_failure now classifies dict results as failures, so the
concurrent worker's failure log line sliced result[:200] on a dict and
raised TypeError; the worker died and the model saw "thread did not
return a result" instead of the tool's own error payload. Stringify the
preview like the sequential path does.
2026-09-15 18:42:10 -07:00
KoNit-K c89209d564 fix(gateway): report structured tool failures in run events 2026-09-15 18:42:10 -07:00
jerryhjones 2bd0f1c59e fix(kanban): scope the stop nudge to the dispatcher-owned worker
agent/kanban_stop.py::kanban_stop_nudge_enabled tested only HERMES_KANBAN_TASK,
which in-process delegate_task children (and cron runs fired inside a worker)
inherit from the worker's process environment. Those executions own no board
task and have the kanban toolset withheld, so the turn-end nudge ordered them to
call a tool they cannot reach — burning attempts, and in production driving
children to complete the parent's card through the CLI.

Gate on agent/delegation_context.py::is_dispatcher_owned_worker_context, the
predicate every other HERMES_KANBAN_* identity gate already uses. The real
worker and the HERMES_KANBAN_STOP_NUDGE opt-out are unchanged.

Salvaged from PR #84656 by @jerryhjones (re-applied onto the current facade
shape); the same gate was first proposed in PR #80023 by @webdevfrancisco
using the narrower delegated-child predicate.

Co-authored-by: webdevfrancisco <franciscombautista2015@gmail.com>
2026-09-15 18:41:19 -07:00
Konstantin Khlopkov 699037176c fix(agent): log the serialized size of multimodal results in the concurrent executor
The concurrent completion line logged len(result) directly, so a native-path
vision_analyze envelope dict reported "4 chars" — its key count — while the
sequential path already logs the serialized length. Mirror the sequential
measurement so parallel multimodal calls stop looking truncated in logs.
2026-09-15 18:40:03 -07:00
teknium1 996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
teknium1 dc7e52874e docs(agent): record why the Z.AI vision default is glm-5.3-flash
Pinned vision ids rot silently (glm-5v-turbo was retired from the Coding Plan endpoints
while still valid on pay-as-you-go), so name the data behind the new pin — the only
image-capable GLM id on every Z.AI surface — and why the pin cannot simply be dropped in
favour of ProviderProfile.default_vision_model() (ZaiProfile returns None, which would route
vision to a text-only chat model).

Co-authored-by: POWERFULMOVES <142271328+POWERFULMOVES@users.noreply.github.com>
2026-09-15 18:23:55 -07:00
KoNit-K c93f2e1d59 fix(agent): update zai vision fallback 2026-09-15 18:23:55 -07:00
teknium1 cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
teknium1 ee49b7d25d fix(agent): file-mutation footer states failed edits, not "files were NOT modified"
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.

- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
  and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
  (mtime_ns, size) snapshot; a later success clears every entry with the same
  identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
  entries whose file changed since the failed call, so a receipt-less
  mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
  way the file tools resolved them.

Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.

Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:22:12 -07:00
teknium1 0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1 66c9440826 fix: stream retry no longer replays the old model on the switched-to provider
After a mid-turn /model switch while a stream was stalled, the streaming
retry loop re-sent the request it had captured at construction time. That
payload still named the OLD model, but every stream (re)open builds its
request client from the LIVE agent, so the new provider's base_url received
a foreign model slug: 404 "Not found the model ...", then the turn sat in
the provider's rate-limit hold (#112121).

_StreamingCall now records the route (model, provider, base_url, api_mode)
its api_kwargs were built for. When a retry is about to be issued and the
live route differs, the streamer stops and hands the transient error back
to the turn loop instead. The turn loop already rebuilds the request per
attempt for the CURRENT route (turn_api_request.build_api_request: model,
wire shape, prompt-cache decoration, provider request overrides), so a
re-keyed model alone would still have shipped a payload shaped for the old
provider. Non-streaming requests have no in-process retry, and the
fallback / restore-primary paths go through the same turn-loop rebuild, so
this is the only site that replayed a captured route.

Fixes #112121

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:21:17 -07:00
teknium1 d54ae02606 fix(pricing): custom-provider /models prices already per-million are no longer inflated 1e6x
`_extract_pricing`'s generic path copied catalog values verbatim, while
usage_pricing unconditionally applies OpenRouter's per-token convention and
multiplies by 1e6. A provider quoting USD per 1M tokens (Neosantara 0.6/M,
Crof cost.input 0.04/M) or declaring `unit: per_1m_tokens` therefore priced
at $600,000/M and corrupted estimated_cost_usd in state.db and every cost
report summing across providers.

Normalize at the producer, where Novita/DeepInfra unit handling already
lives: an explicit `unit` beside the rates wins (per_token / per_1k_tokens /
per_1m_tokens); without one, a token rate at or above $0.001 per token
($1,000/MTok - no real model) can only be a per-million quote. Output keeps
the per-token-string contract, so the consumer is untouched; `request` fees
and per-token catalogs pass through unchanged.

The $0.001/token magnitude threshold is the one proposed in #34263 by
@Bartok9 (earliest fix); #112036 by @kvnloo proposed the same heuristic at
the consumer.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:20:51 -07:00
teknium1 5435ac8cc4 fix(aux): clamp extra_body.reasoning effort too so auxiliary.<task>.reasoning_effort ultra never reaches the wire 2026-09-15 18:20:13 -07:00
teknium1 9e45a90488 fix: clamp aux reasoning effort once before profile projection
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.

Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.

Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
2026-09-15 18:20:13 -07:00
KoNit-K c02db64077 fix(agent): clamp auxiliary ultra reasoning effort 2026-09-15 18:20:13 -07:00
teknium1 f55d4f6747 fix(aux): bare-custom AuthError yields no custom endpoint instead of a stale env OPENAI_BASE_URL 2026-09-15 18:19:50 -07:00
teknium1 c1bbcf9712 fix(agent): hint-preview truncation log names the real remedy; invariant tests
Follow-up to the two salvaged commits (#111777, #111781 by @KoNit-K):

- agent/prompt_builder.py::_truncate_content — with queue_warning=False the
  logged line no longer tells the operator to "pin a larger
  context_file_max_chars, or use a larger-context model": the subdirectory
  hint cap is a constant neither knob raises. It now points at the read_file
  recovery the marker already discloses.
- tests/gateway/test_startup_environment_probe.py — replace the
  call-detection test with the behavioural invariant: an oversized SOUL.md in
  HERMES_HOME and a warm-up leave the truncation-warning queue empty for the
  next default-executor task (the api_server turn path runs on that executor
  without copy_context, which is how the boot warning reached a foreign
  session).
- tests/agent/test_subdirectory_hints.py — fold the new drain assertion into
  the existing oversized-hint test (same fixture) and pin that the log carries
  no context_file_max_chars advice.
- agent/AGENTS.md, website/docs/.../context-files.md — the hint cap is 32,000
  (docs said 8,000) and is fixed; document that it is logged, not surfaced as a
  chat warning.
2026-09-15 18:19:26 -07:00
KoNit-K fc7cbc7d8e fix(agent): suppress subdirectory hint truncation warning 2026-09-15 18:19:26 -07:00
teknium1 857133babd refactor: trim empty-frame salvage to the translated error only
Drop the raw JSONDecodeError belt from _is_provider_stream_empty_frame_error:
every chat stream iteration goes through _iter_provider_stream_chunks, which
already translates the decode failure, so the belt guarded a path that does
not exist. Drop the explicit classifier entry for the new code: unknown codes
already resolve to the same retryable unknown verdict (verified live: identical
ClassifiedError with and without the entry).
2026-09-15 18:18:08 -07:00
moxian 71281fb266 fix(agent): contentless SSE keepalive frames retry without streaming
An empty `data:` frame (or a frame carrying only `event:` / `id:`) is a legal SSE
no-op, but the OpenAI SDK still hands it to `json.loads`, which raises
`JSONDecodeError` with an empty document. The streaming helper translated that
into `ProviderStreamError(provider_stream_non_json_data)`, nothing recognised it
as recoverable, and the main loop retried streaming -- identically -- until the
retry budget ran out: "API call failed after 3 retries: Provider stream returned
non-JSON SSE data". A degraded gateway answers EVERY streaming request that way,
so the retries only repeated the failure while the same request sent
non-streaming succeeded seconds later.

An empty document means no payload, which is a different fact from a malformed
payload: give it its own code, and when it appears before any delta switch the
session to non-streaming (the existing `_disable_streaming` mechanism, already
used for "stream not supported", Bedrock IAM denials and adapter-returned final
responses) so the retry goes out on a channel the degraded gateway can answer.
Non-empty payloads keep today's fatal semantics; failures after deltas keep the
existing stream-drop handling.

Tests: the turn-level recovery test drives the real SDK decoder over a real
httpx response and is red on base (three identical streaming attempts, turn
lost); the boundary test pins the malformed-payload path so a future "ignore bad
frames" change cannot swallow genuine provider errors.
2026-09-15 18:18:08 -07:00
teknium1 423bc7e4e4 fix(credential-pool): hydrate on-disk env rows on the openrouter branch too
The openrouter branch of _seed_from_env returns before the generic loop,
so a persisted `env:OPENROUTER_API_KEY_2` row stayed empty exactly like
the registry-provider case #103067 fixed. Fold the "declared vars + on-disk
env rows" union into one helper both branches use.
2026-09-15 11:50:48 -07:00
McClean-Newton 23afade67b fix(credential-pool): seed env-source entries not in registry tuple
Restores the env-source seeding loop in _seed_from_env that was dropped
by upstream refactors. Without this, any env-source entry whose env var
name isn't in the provider registry's hardcoded api_key_env_vars tuple
stays empty forever, gets filtered out of rotation by _available_entries,
and round-robin silently degrades to single-key behavior.

Forward-port of ba71e00db07c6263ee8d44b27dfbce2a92e6b39c onto current main
f1ccf436a2. Narrow insertion after final
env_vars list in _seed_from_env: scan only existing entries whose source
begins with env:, extract/dedupe the named variable, and let the existing
seeding path hydrate it via get_env_prefer_dotenv (preserving _Seeder,
secret-scope/profile isolation, borrowed-secret sanitization, and
suppression behavior).

Tests: 10 new tests in test_credential_pool_seed_existing_env_sources.py
covering three-key hydration, round_robin cycling, dotenv/secret-scope
resolution without os.environ bypass, duplicate dedupe, manual rows
untouched, unset remains unavailable, and sanitized persistence.
2026-09-15 11:50:48 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
kshitijk4poor 683ca44046 fix(classifier): a welcome-host 403 naming the free tier itself is the tier refusing, not a billing wall
The "says something else" guard reuses the billing table, which carries the
Nous gateway's own free-tier phrases ("not available on the free tier",
"model_not_supported_on_free_tier"). On the welcome route those words mean
exactly "the tier refused"; routing them to billing prints a credits check to
an anonymous session that has none. Leave those two phrases on the
tier_disabled path.
2026-09-15 20:44:42 +05:30
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 2e3b8ecc82 fix(free-tier): the desktop card body leaves the "To sign in" tail off; the button is the door
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 16def7b8cc fix(free-tier): the desktop renders a free-tier refusal as its own card, not an OAuth re-login
A welcome-tier 403 classifies as auth_permanent, so the desktop's error
surface mapped it to "Your Nous Portal sign-in expired" with a Nous Portal
re-login button — the chat sentence never reached the user. Terminal results
on the free route now carry a structured free_tier block (kind + the chat
sentence); agent/error_surface.py turns it into a free_tier_<kind> code on
the provider layer with the sentence as `message`. The desktop gives those
codes their own titles, shows the backend sentence as the body, and offers
"Sign in with a Nous account" (the free-tier dialog) instead of the OAuth
re-login, with Retry only where a later send can succeed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 a1e7f74e64 fix: stat signatures cover inode + ctime everywhere config identity is cached
The cherry-picked commit extends the config/profile/MCP/managed-scope/completer/
OAuth/skills-manifest signatures. This commit finishes the class and trims it:

- `file_signature()` lives in `utils.py` next to the other stat/metadata helpers
  instead of `hermes_cli.managed_scope` (gateway/ and agent/ callers no longer
  reach into the managed-scope module for a generic stat helper).
- `hermes_cli/config_effective.py` was left comparing 2-/4-wide prefixes against
  the widened `_RAW_CONFIG_CACHE` / `_load_config_cache_sig` records, so
  `load_user_config_effective()` re-parsed on every call (3 parses for 3 calls on
  an unchanged file, 1 before); index by the new widths.
- Sibling caches keyed on the same (mtime, size) shape and reading the SAME files
  now use the helper: `load_env()` memo, `agent/skill_utils` raw-config and
  external-dirs caches, `hermes_cli/model_switch` alias identity, `agent/moa_loop`
  preset stamp, `hermes_cli/auth` global auth-store memo.
- Tests trimmed to one invariant each (pinned-mtime replacement invalidates; an
  unchanged file still hits), both red on origin/main.

Left alone on purpose: `tools/registry.py`, `tools/skills_tool_dedup.py`,
`gateway/status.py`, `hermes_cli/banner.py`, `hermes_cli/main.py`,
`hermes_cli/session_recovery.py` — those fingerprint source files, PID/lock files
or write to persisted on-disk caches shared across processes, where an inode/ctime
key would churn on every checkout/copy rather than catch a replaced config.
2026-09-15 06:29:50 -07:00
Kevin Rajan b797e9d7b3 fix(config): detect file replacements in stat-based cache signatures
Config caches, the profile re-scan watcher, the MCP reconciler, managed-scope
reads, the completer memo, OAuth heal marks, and the skill-manifest snapshot
keyed change detection on (st_mtime_ns, st_size) only, so a replacement that
preserved both (cp -p, rsync -t, timestamp-pinning scripts, sync clients) was
treated as unchanged and stale values were served until restart.

Add st_ino (fresh inode on atomic replace) and st_ctime_ns (cannot be
backdated via os.utime) to every signature via a shared
hermes_cli.managed_scope.file_signature() helper.

Fixes #111105
2026-09-15 06:29:50 -07:00
teknium1 979576d938 fix: linear gutter anchors, cat -n/grep-context gutters, keep rc path references readable
The optional line-number gutter was written as `^(?:[ \t]*GUTTER)?[ \t]*`,
stacking two adjacent whitespace runs so _CFG_ANCHORED_RE / _YAML_ASSIGN_RE
went quadratic on any indented line with a secret keyword and no `=` (2s per
5k spaces, 50s at 20k) — and these run on every terminal output and file read.
The gutter now carries the single optional group and callers own the one
leading whitespace run. It also accepts `cat -n` / `nl` (number + TAB) and
`grep -A/-B/-C` (`5-`) gutters, which the comment claimed but leaked.

The strong-key branch masked SSH_AUTH_SOCK=$HOME/..., DOCKER_AUTH_CONFIG=/...
on shell rc reads, leaving the agent unable to edit them; a value starting with
$, / or ~ is a variable or path reference and is now kept unless the key is
password-class. _is_secret_file_arg only extends config.yaml to the
backups/config .good./.corrupt. copies, not config.yaml.pdf/.bak.

Review finding: quadratic gutter regex; cat -n/nl TAB and grep context gutters leaked; over-redaction of rc path references.
2026-09-15 06:26:29 -07:00
teknium1 f5907fd052 fix(redact): config.yaml backup copies under HERMES_HOME are secret-bearing too
search_files over $HERMES_HOME returned the token from backups/config/config.yaml.good.<stamp>
in cleartext while config.yaml itself was masked; the predicate only matched the exact basename.
Same predicate serves the terminal side, so `cat backups/config/config.yaml.good.*` is covered too.
2026-09-15 06:26:29 -07:00
shehjaddev d205cef418 fix(redact): mask assignments in secret-bearing file reads
read_file/search_files passed file_read=True, which folded into code_file=True and skipped
the ENV/JSON/YAML assignment passes, so an opaque prefix-less credential under a
credential-shaped key reached the model in cleartext from a secret-bearing file — the
file-read half of the #110228 gate (#110567).

Two defects on that path, both fixed here:

- The rendered line-number gutter ("5|      ADS_API_TOKEN: ..." from read_file,
  "6:      ADS_API_TOKEN: ..." from grep -n / cat -n) defeated the line-anchored patterns,
  so the real rendered read leaked exactly what the raw text masked. A gutter-free fixture
  cannot see this, which is why the tool-level tests carry the real render shape.
- _is_secret_file_arg() could not see the RESOLVED Hermes home: the default home's basename
  is an installation detail (".hermes" on POSIX, "hermes" under AppData/Local on Windows) and
  a resolved path never spells $HERMES_HOME, so the managed Windows home's config.yaml was
  classified as ordinary YAML on both the file-read and the terminal surface.

Changes:

- redact_sensitive_text(): secret_file= re-enables the assignment passes for content the
  caller classified with _is_secret_file_arg, keeping code_file behaviour everywhere else.
  It is authoritative over code_file, so a caller cannot be fail-open on the security flag
  by setting both.
- _redact_assignments(): mask_nonreusable selects the non-reusable sentinel for file reads,
  so the #35519 write-back hazard stays closed.
- _should_redact_assignment(): no longer re-masks an already-masked value, which was erasing
  the vendor label the sentinel deliberately keeps.
- _is_secret_file_arg(): consult the resolved Hermes home for the config.yaml arm.
- _CFG_ANCHORED_RE / _YAML_ASSIGN_RE: tolerate a rendered line-number gutter.
- file_tools.py: classify the resolved path at all three file-read call sites.

Closes #110567
2026-09-15 06:26:29 -07:00
teknium1 b98ff81978 refactor(file-safety): fold vault/ and browser-profile/ into the existing protected-subpath table
The salvaged fix added a second directory tuple and a second loop for the same
predicate. _HERMES_PROTECTED_SUBPATHS already expresses "HERMES_HOME subpaths the
file tools must not rewrite", so the two secret directories join it and the
duplicate loop goes. Tests trimmed to two invariants: every read-denied secret
store is also write-denied on both the profile and the global root, and the
#45947 control files plus same-named paths outside HERMES_HOME stay writable.
2026-09-15 06:20:15 -07:00
NUXER 1c0d95badb fix(security): write-deny HERMES_HOME secret stores, keep control files writable
Narrow the fix to secret material and stop reverting #45947.

The earlier revision added every read-denied name to the write denylist,
which re-blocked auth.json and webhook_subscriptions.json and rewrote the
test guarding them. #45947 freed those control files deliberately:
containment belongs in Docker/remote backends and OS permissions, not an
expanding hardcoded denylist.

What #45947 kept blocked is secret material, and that list had drifted:
auth/google_oauth.json (OAuth token store), the plaintext Bitwarden cache,
vault/ (key + ciphertext side by side) and browser-profile/ (copied
cookies / Login Data) were writable via write_file / patch.

- Write-deny those four; control files stay writable and read-denied.
- _WRITE_DENIED_SECRET_DIRS is its own tuple rather than _READ_DENIED_DIRS,
  so a future read-only convenience deny cannot silently become a write deny.
- tests/tools/test_write_deny.py is restored unchanged from main.
- New tests pin both halves: secrets denied, control files writable, and
  write denies stay a subset of read denies.

Fixes #110464
2026-09-15 06:20:15 -07:00
NUXER e7cd1848c9 fix(security): deny writes to read-blocked Hermes credential stores
Read guards already refuse auth.json, webhook HMAC secrets,
google_oauth.json, the Bitwarden plaintext cache, vault/, and
browser-profile/. The write denylist only covered .env, the Anthropic
PKCE store, and bws_cache.enc.json, so write_file/patch could replace
the rest.

Drive the write denylist from the same credential names and
read-denied directories. Keep the extra write-only entries
(bws_cache.enc.json, state.db, sessions, pairing).

Fixes #110464
Related: #108716
2026-09-15 06:20:15 -07:00
KeyArgo 0968fa2631 fix(agent): classify NVIDIA NIM serde rejection of list-type tool content
NVIDIA NIM's Rust gateway 400s on list-type tool message content with a serde
error that names the enum rather than the field: "data did not match any
variant of untagged enum ChatCompletionRequestToolMessageContent". None of the
existing _MULTIMODAL_TOOL_CONTENT_PATTERNS match that wording, so the error
classifies as format_error (non-retryable) and the existing recovery —
downgrade the image-bearing tool message to text, remember (provider, model),
retry once — never fires: every retry re-sends the same shape and fails
identically in a session bound to that model.

Add the serialized enum name to the pattern tuple so the NVIDIA wording routes
to multimodal_tool_content_unsupported like the MiMo/Alibaba/Console Go
wordings already do (#111231).
2026-09-15 06:18:31 -07:00
KoNit-K 7d03ea3adf fix(agent): recover opencode zen encrypted replay 2026-09-15 06:18:31 -07:00
teknium1 95987fb85a fix: bypass the proxy on loopback HTTP CDP discovery and keep NO_PROXY=* intact
The websockets dials got proxy=None but the three HTTP /json/version dials
(CDP override discovery, is_browser_debug_ready used by Lightpanda and the
real-profile readiness check, and surviving-Chrome detection) still resolved
via getproxies(), so under a system/env proxy discovery fell back to the raw
http:// URL, readiness never fired and /browser connect reported not ready.
loopback_request_kwargs() sits next to loopback_connect_kwargs() and is used
at all three sites (ProxyHandler({}) opener for the urllib one).

add_loopback_no_proxy turned an operator NO_PROXY=* into '*,127.0.0.1,...',
which urllib/requests no longer treat as the wildcard, flipping bypass-all
configs into proxy-all. A wildcard in either casing now leaves env untouched.
is_loopback_host also accepts any loopback IP literal (127.x, ::ffff:127.0.0.1).

Review finding: HTTP /json/version dials still proxied loopback; NO_PROXY=* wildcard broken by append.
2026-09-15 06:16:38 -07:00
teknium1 8f6f92d901 fix(browser): one loopback proxy-bypass helper covers child envs and in-process CDP dials
Move the loopback NO_PROXY merge from browser_tool into agent/proxy_bypass.py (the
module that already owns NO_PROXY semantics) and reuse no_proxy_entries() so comma-
and whitespace-separated operator values are both preserved. Add
loopback_connect_kwargs() and pass proxy=None on the two in-process websockets
dials to loopback CDP endpoints (browser_cdp_tool._cdp_call, BrowserSupervisor._run):
those never see the child env, so the env merge alone left them routed through a
macOS system proxy. Remote CDP URLs keep the default proxy behaviour.

Tests trimmed to two invariants: the built child env appends loopback to an
operator NO_PROXY in both casings, and only loopback URLs get proxy=None.
Sibling helper in tools/browser_use_cli (#110570) is redundant once the shared
env carries the entries.
2026-09-15 06:16:38 -07:00
teknium1 1b7355d7fa fix: cap kanban card titles so an overlong card still names its worker
Kanban cards have no length limit, but the session title store rejects
titles past SessionDB.MAX_TITLE_LENGTH with ValueError. _persist_session_title
reads that as a unique-title collision, retries with a "#N" suffix (longer
still), and the caller suppresses the second failure - so a worker spawned on
a >100-char card ended up with no title at all, where main at least gave it a
derived one. Trim the card title (with room for the "#N" retry suffix) before
persisting; a retried card now gets "<trimmed> #2" within the cap.

Review finding: >100-char card title left the kanban worker session untitled.
2026-09-15 06:08:20 -07:00
teknium1 c12a3397b1 fix(kanban): read the worker's card title from the board instead of a new env var
Follow-up to the salvaged commit from #111169 (@KoNit-K):

- Drop HERMES_KANBAN_TASK_TITLE. The worker already has HERMES_KANBAN_TASK
  and HERMES_KANBAN_BOARD/HERMES_KANBAN_DB pinned in its env, so
  maybe_auto_title reads the card title from the board itself (no new
  HERMES_* env var for non-secret config; the dispatcher and the
  delegation scrub list stay untouched).
- Unreadable or missing card: the session is named `Kanban task <id>`
  with zero auxiliary calls (the fallback the issue asked for; the
  #109743 seed left such workers untitled).
- The card title persists at `llm` authority via set_auto_title, so a
  manual /title still wins and the upgrade thread never starts.
- Tests trimmed to two invariants against a real board + SessionDB
  (card title, unreadable-card fallback), both red on origin/main.
2026-09-15 06:08:20 -07:00
KoNit-K 56555f88df fix(kanban): name worker sessions from task title 2026-09-15 06:08:20 -07:00
KoNit-K af4a3eba0a fix(agent): skip corrupted Gemini thought signatures 2026-09-15 05:42:35 -07:00
AltenLi 645b9526a2 fix(agent): steer briefing must match standalone-user-row delivery (#110979)
The mid-turn /steer marker is delivered as a standalone role:"user"
message right after the newest tool result (steer_user_row /
apply_pending_steer_to_tool_results). But STEER_CHANNEL_NOTE (the
model-facing briefing) and the module comment still said Hermes
'appends their message to the end of a tool result'.

That stale claim briefs the model to expect the marker INSIDE tool
output, so a real standalone user row carrying the marker can read as
off-channel — the under-trust half of #110979. Correct the briefing and
comment to describe actual delivery; add a contract test tying the note
wording to the steer_user_row delivery mechanism (proven red on the old
wording).
2026-09-15 05:39:31 -07:00
Kevin Rajan 7d292af875 fix(journey): show foreground-created skills in /journey via learn provenance marker
record_created now stamps created_by="learn" on foreground creates
(e.g. /learn) instead of leaving it unset, and the learning-graph filter
honors "learn" alongside "agent"/used. "learn" is a learning-signal
marker only: curator management stays keyed strictly on "agent"
(_is_curator_managed_record), so user-taught skills appear in /journey
without becoming eligible for autonomous curation.

Fixes #111317.

---
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; the contributor reviewed the diff and ran the tests.
2026-09-15 05:38:31 -07:00
teknium1 d6d9e67f54 fix(compression): announce the compacting status before the lazy feasibility probe
The first compaction of a session runs check_compression_model_feasibility()
inside compress_context() BEFORE _announce_compression_start(). That probe is
network-bound (live model catalog / provider lookups; connect timeouts stack up
through proxies and slow remote gateways), so every automatic entrypoint that
reaches compress_context without its own pre-emit — the post-tool gate in
agent/turn_preflight.py::compress_after_tool_results, overflow recovery,
manual /compress — left the client with no `kind="compacting"` status for the
whole probe. On Desktop that is a bare working-row spinner with no
"Summarizing thread" label (#111294).

Move the announcement ahead of the probe in the one choke point so every
caller is covered, and retire the announced phase (force_terminal) when the
probe's hard rejection propagates so the client never stays "compacting" for
an attempt that never started.

Live repro: /tmp probe driving compress_after_tool_results with a 2 s
feasibility stand-in — before: first compacting status at t=2.29 s (after the
block); after: t=0.13 s.
2026-09-15 05:19:57 -07:00
teknium1 5117e3b3a0 fix: key the check_fn cache by the same served-profile predicate as the MCP registry scope
_mcp_registry_scope() became profile-keyed for served profiles with the
multiplex flag off, but check_fn_cache_scope() still returned None in that
mode, so the process-wide availability cache stayed keyed (fn, None) across
profiles. A served profile whose mcp__x__* check_fn now correctly resolves to
its own (absent) connection cached False for the TTL window and the launch
profile that owns the live connection lost its tools for that window.

Both sites now call one helper, agent.secret_scope.serves_routed_profile()
(multiplex on, or a HERMES_HOME override naming a home other than the
process home), so the registry scope and the cache key can no longer drift.

Review finding: served profile B's check_fn verdict shadowed the launch profile's live mcp tools via the unscoped check_fn cache.
2026-09-15 04:56:40 -07:00
teknium1 8a2996503a fix: skip Bitwarden URIs marked match=Never when collecting fill origins
Binding every saved URI made a URI the user explicitly set to "Never"
match (match=5) a valid fill target — wider than the manager's own policy.
Filter those out; all other URIs keep exact-origin semantics.

Review finding: match=5 (Never) URIs became fill origins.
2026-09-15 04:56:01 -07:00