Commit Graph

4933 Commits

Author SHA1 Message Date
m4 3ca91c55f3 feat(desktop): openExternalFileForIpc opens files via OS handler
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-09-17 22:28:07 +08:00
Konstantin Khlopkov 7f7d229f83 fix(utils): atomic writers refuse to resurrect a deleted named profile home
Background writers that still carry a tombstoned profile as their Hermes
home (reasoning-caps warm thread, models cache, models.dev ETag, gateway
lifecycle ledger, MCP OAuth tokens, memory store) re-created
profiles/<name>/ with a bare mkdir right before an atomic write. Route the
parent-dir creation through mkdir_under_hermes_home so a deleted named
profile raises FileNotFoundError and stays gone, matching the tombstone
contract already enforced for logging and state.
2026-09-16 00:32:15 -07:00
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
kshitijk4poor 5d07f7fe66 test(agent): fold the aux create_client() seam tests to two contracts
Two parametrized invariant tests (native client sync/async; None or raising
profile falls back to the standard client) replace four near-duplicates, and
the hook's kwargs are pinned exactly to the mapping openai.OpenAI would have
received. The fixture isolates both provider registries through monkeypatch
instead of a hand-rolled snapshot/restore, drops the HERMES_HOME override that
tests/conftest.py already provides, and replaces the tuple-truthiness lambda
with a plain function. Call-site comment trimmed to the ordering WHY.
2026-09-16 11:57:56 +05:30
liuhao1024 4cd2eb013c fix(agent): honor ProviderProfile.create_client() for api_key aux routes
The auxiliary api_key branch built openai.OpenAI directly, so an
out-of-tree provider registered with auth_type="api_key" lost its
native transport for auxiliary tasks even though the main-agent path
(_provider_supplied_client) and the external_process branch honor the
same hook. Consult the profile's create_client() before the built-in
gemini/OpenAI ladder; None (the default) falls through untouched, and
a raising profile is logged and skipped (#112384).
2026-09-16 11:57:56 +05:30
outpoints 44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
outpoints 5237cab756 fix(honcho): thread logical cwd through agent construction
(cherry picked from commit b1d7207c45311be658592c6ad34ee84634fed0ee)
2026-09-15 22:30:11 -07:00
outpoints 3cbdc32565 fix(honcho): don't let auto-generated session titles override sessionStrategy
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.

Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.

Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.

Fixes #24740

(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
2026-09-15 22:30:11 -07:00
kshitijk4poor a12b3c7aa3 refactor(agent): import the transform bypass from its defining module, no re-export shim
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
2026-09-15 19:23:53 -07:00
kshitijk4poor af7b60e8ec fix(agent): keep moved chat fields as slot placeholders so the wire body is byte-identical; one shared escape hatch
The cherry-picked helper deleted 'tools' from the typed kwargs, so the SDK's
post-transform extra_body merge appended it after the caller's extra_body keys —
equal dict, different bytes (byte-keyed prompt caches would miss). keep_slots=True
leaves [] placeholders that the merge overwrites in place. Drop the invented
HERMES_CHAT_SDK_TRANSFORM env var; the pre-existing HERMES_CODEX_SDK_TRANSFORM
hatch from #93650 now disables both API families. Tests trimmed to the two
invariants (byte-identity incl. caller extra_body precedence; escape hatch).
2026-09-15 19:23:53 -07:00
John Paul Soliva 1e39c93710 perf(agent): keep bulk chat-completions payloads out of the SDK request transform
`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. #93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.

Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:

    101 msgs /  76 KB   13.6 ms -> 1.2 ms
    401 msgs / 190 KB   48.6 ms -> 2.0 ms
   1601 msgs / 650 KB  188.7 ms -> 5.8 ms

and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.

The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.

Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.

Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.

The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.
2026-09-15 19:23:53 -07:00
KoNit-K a457e91a50 fix(secrets): preserve OP_CONFIG_DIR for 1Password 2026-09-15 19:10:00 -07:00
teknium1 204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
Shenrui Ma 6973f2622d fix(codex): target hermes-tools in Kanban worker overrides
Dispatcher-owned Kanban workers on the codex app-server runtime injected their
HERMES_KANBAN_* scope into `mcp_servers.hermes-mcp.env.*`, but the runtime
migration registers Hermes' MCP callback as `[mcp_servers.hermes-tools]`. The
override therefore materialised a second, env-only server entry that codex
rejects at bootstrap ("invalid transport in `mcp_servers.hermes-mcp`"), so no
worker could initialize. Point the overrides at the entry that actually exists.

Salvaged from #107337 (its 147-line parametrized test file is replaced by two
invariant tests in a follow-up commit). #111711 proposed the identical two lines.

Fixes #111707

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:05:29 -07:00
teknium1 837acdc91d fix: share the control-frame opener list and cover [System:/[IMPORTANT:/[PRIOR CONTEXT/[CONTEXT SUMMARY]
Review finding on #112260: the hosted-room member relabel regex was a third
hand-copied opener list that missed the frames context_compressor and
title_generator already treat as harness input, so a member reply starting
with "[System: ..." or "[IMPORTANT: 2 background processes ..." reached peers
in its exact trusted shape. agent.prompt_builder.CONTROL_FRAME_OPENERS is now
the single source; the gateway regex is built from it and the desktop TS
literal mirrors it byte-for-byte.
2026-09-15 19:04:59 -07:00
teknium1 54ed7cbb7b fix: validate the skill name before opening its lock; key lock files on a digest
Review finding on #112218 (major): `_skill_lock_path` opened `<skills>/.locks/<name>.lock`
before the name was validated, so `skill_manage(action='create', name='a'*300)` raised
OSError (File name too long) and a NUL name raised ValueError instead of the handler's
JSON error, and every rejected name ('../../etc', '') left a residue lock file.

- tools/skill_manager_tool.py: lock filename is sha256(basename).lock (fixed width, no
  filesystem limit reachable; `foo` and `category/foo` still share one lock), the redundant
  `_find_skill` rglob is gone, and `skill_manage` runs `_validate_name` on the name
  (create) / basename (other actions) before the lock is opened.
- '.locks' joins the skills-dir exclusion sets (EXCLUDED_SKILL_DIRS, ledger
  _NON_PACKAGE_TOPS, learning-graph/skill-commands skip parts, curator backup excludes).
- tests: 2 invariants in TestSkillMutationLock (rejected names -> JSON + no .locks residue;
  digest-keyed lock shared across name forms), red on the old head.
2026-09-15 19:03:33 -07:00
Zheqing Zeng 8b58562614 docs(curator): align the consolidation prompt with the delete guard
_curator_consolidation_delete_guard refuses background deletes whose
absorbed_into is missing OR empty (fail-closed, #29912): pruning with no
forwarding target belongs to the deterministic staleness pass. The prompt
still instructed the model to pass absorbed_into="" for exactly that case —
a dead-end instruction that burned tool iterations on guaranteed refusals.
Tell the model the rule the guard actually enforces.
2026-09-15 19:01:11 -07:00
Zheqing Zeng 9b003f201f fix(curator): seed shared read-marks store in the LLM consolidation fork
The read-before-write guard requires a skill_view mark from the SAME review
run before any skill_manage write. mark_background_review_skill_read
auto-creates a store when the ContextVar is unset, but tool workers run on
copied contexts, so marks recorded in one worker stayed invisible to the
others: every patch was refused with "current SKILL.md content has not been
loaded in this review turn" even after fresh full reads, and the
consolidation pass burned its iterations retrying a dead-end write.

The background-review fork already seeds a shared store before
run_conversation (agent/background_review.py); do the same in the curator
fork so every copied worker context shares one store.
2026-09-15 19:01:11 -07:00
teknium1 60f436b5f6 fix: redact '/'- and '~'-led secrets whose segments cannot be a path
Review finding (agent/redact.py::_should_redact_assignment): a '/'-prefixed
secret with a second '/' (`AWS_SECRET_ACCESS_KEY=/wJalrXUtnFEMIK7MDENG/bPx…`)
still parsed as a multi-segment path and leaked under a strong key; the same
held for a '~'-led value. Apply the opaque bar per segment (16+ chars, no
'.', mixed case and digits) to every '/' or '~' value instead of only to
single-segment ones, so `/home/u/.docker`, `~/.ssh/id_rsa` and
`S.gpg-agent.ssh`-style paths stay readable while base64 secrets mask.
2026-09-15 18:59:55 -07:00
teknium1 34067a7b6d fix: keep $(cmd) substitutions readable under strong-key assignments
Review finding (agent/redact.py::_PATH_OR_VAR_VALUE_RE): anchoring the
path/var exemption regressed `export SSH_AUTH_SOCK=$(gpgconf --list-dirs
agent-ssh-socket)` vs main — the `$(gpgconf` token no longer parsed as a
reference and was masked. Accept a leading `$(` as a reference atom in the
grammar and pin the gpg-agent line in
test_real_path_and_var_references_stay_readable.
2026-09-15 18:59:55 -07:00
teknium1 93a269c026 fix(redact): keep $VAR interpolations inside rc path values readable
The anchored path/variable grammar from the salvaged fix only allowed one
leading $VAR; a strong-key rc line such as SSH_AUTH_SOCK=/run/user/$UID/ssh or
SSH_AUTH_SOCK=$XDG_RUNTIME_DIR/agent.$USER.sock no longer parsed as a reference
and was masked, undoing the readability contract from 979576d938 for exactly
the lines it was written for.

Allow $VAR / ${VAR...} anywhere in the value (and ':' list separators). Crypt
digests still fall through to the credential checks: their '$' fields start
with a digit or carry '=' / ',', which the grammar rejects.
2026-09-15 18:59:55 -07:00
beardthelion 93a0ec705a fix(redact): anchor the path/var exemption so leading-/ and $ secrets still mask
_PATH_OR_VAR_VALUE_RE was an unanchored character class, so re.match made it a
first-character test: any assignment value beginning with '$', '/', or '~'
returned early from _should_redact_assignment, ahead of the strong-key and
opaque-credential checks. AWS secret access keys (~1 in 64 begin with '/') and
argon2/bcrypt digests (always '$'-prefixed) leaked verbatim through
redact_sensitive_text.

Anchor the pattern on both ends so the exemption only fires on a complete
$VAR/${VAR}/~/path//abs/path reference, and require a single-segment absolute
path — indistinguishable by shape from a high-entropy secret — to clear the
opaque-credential bar first. $VAR and ~/ references stay exempt
unconditionally, preserving the rc-readability contract that motivated the
exemption (SSH_AUTH_SOCK=$HOME/.ssh/agent.sock,
DOCKER_AUTH_CONFIG=/home/u/.docker).
2026-09-15 18:59:55 -07:00
teknium1 1d14418ea2 fix(agent): concurrent worker survives a dict error result
_detect_tool_failure now classifies dict results as failures, so the
concurrent worker's failure log line sliced result[:200] on a dict and
raised TypeError; the worker died and the model saw "thread did not
return a result" instead of the tool's own error payload. Stringify the
preview like the sequential path does.
2026-09-15 18:42:10 -07:00
KoNit-K c89209d564 fix(gateway): report structured tool failures in run events 2026-09-15 18:42:10 -07:00
jerryhjones 2bd0f1c59e fix(kanban): scope the stop nudge to the dispatcher-owned worker
agent/kanban_stop.py::kanban_stop_nudge_enabled tested only HERMES_KANBAN_TASK,
which in-process delegate_task children (and cron runs fired inside a worker)
inherit from the worker's process environment. Those executions own no board
task and have the kanban toolset withheld, so the turn-end nudge ordered them to
call a tool they cannot reach — burning attempts, and in production driving
children to complete the parent's card through the CLI.

Gate on agent/delegation_context.py::is_dispatcher_owned_worker_context, the
predicate every other HERMES_KANBAN_* identity gate already uses. The real
worker and the HERMES_KANBAN_STOP_NUDGE opt-out are unchanged.

Salvaged from PR #84656 by @jerryhjones (re-applied onto the current facade
shape); the same gate was first proposed in PR #80023 by @webdevfrancisco
using the narrower delegated-child predicate.

Co-authored-by: webdevfrancisco <franciscombautista2015@gmail.com>
2026-09-15 18:41:19 -07:00
Konstantin Khlopkov 699037176c fix(agent): log the serialized size of multimodal results in the concurrent executor
The concurrent completion line logged len(result) directly, so a native-path
vision_analyze envelope dict reported "4 chars" — its key count — while the
sequential path already logs the serialized length. Mirror the sequential
measurement so parallel multimodal calls stop looking truncated in logs.
2026-09-15 18:40:03 -07:00
teknium1 996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
teknium1 dc7e52874e docs(agent): record why the Z.AI vision default is glm-5.3-flash
Pinned vision ids rot silently (glm-5v-turbo was retired from the Coding Plan endpoints
while still valid on pay-as-you-go), so name the data behind the new pin — the only
image-capable GLM id on every Z.AI surface — and why the pin cannot simply be dropped in
favour of ProviderProfile.default_vision_model() (ZaiProfile returns None, which would route
vision to a text-only chat model).

Co-authored-by: POWERFULMOVES <142271328+POWERFULMOVES@users.noreply.github.com>
2026-09-15 18:23:55 -07:00
KoNit-K c93f2e1d59 fix(agent): update zai vision fallback 2026-09-15 18:23:55 -07:00
teknium1 cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
teknium1 ee49b7d25d fix(agent): file-mutation footer states failed edits, not "files were NOT modified"
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.

- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
  and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
  (mtime_ns, size) snapshot; a later success clears every entry with the same
  identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
  entries whose file changed since the failed call, so a receipt-less
  mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
  way the file tools resolved them.

Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.

Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:22:12 -07:00
teknium1 0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1 66c9440826 fix: stream retry no longer replays the old model on the switched-to provider
After a mid-turn /model switch while a stream was stalled, the streaming
retry loop re-sent the request it had captured at construction time. That
payload still named the OLD model, but every stream (re)open builds its
request client from the LIVE agent, so the new provider's base_url received
a foreign model slug: 404 "Not found the model ...", then the turn sat in
the provider's rate-limit hold (#112121).

_StreamingCall now records the route (model, provider, base_url, api_mode)
its api_kwargs were built for. When a retry is about to be issued and the
live route differs, the streamer stops and hands the transient error back
to the turn loop instead. The turn loop already rebuilds the request per
attempt for the CURRENT route (turn_api_request.build_api_request: model,
wire shape, prompt-cache decoration, provider request overrides), so a
re-keyed model alone would still have shipped a payload shaped for the old
provider. Non-streaming requests have no in-process retry, and the
fallback / restore-primary paths go through the same turn-loop rebuild, so
this is the only site that replayed a captured route.

Fixes #112121

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:21:17 -07:00
teknium1 d54ae02606 fix(pricing): custom-provider /models prices already per-million are no longer inflated 1e6x
`_extract_pricing`'s generic path copied catalog values verbatim, while
usage_pricing unconditionally applies OpenRouter's per-token convention and
multiplies by 1e6. A provider quoting USD per 1M tokens (Neosantara 0.6/M,
Crof cost.input 0.04/M) or declaring `unit: per_1m_tokens` therefore priced
at $600,000/M and corrupted estimated_cost_usd in state.db and every cost
report summing across providers.

Normalize at the producer, where Novita/DeepInfra unit handling already
lives: an explicit `unit` beside the rates wins (per_token / per_1k_tokens /
per_1m_tokens); without one, a token rate at or above $0.001 per token
($1,000/MTok - no real model) can only be a per-million quote. Output keeps
the per-token-string contract, so the consumer is untouched; `request` fees
and per-token catalogs pass through unchanged.

The $0.001/token magnitude threshold is the one proposed in #34263 by
@Bartok9 (earliest fix); #112036 by @kvnloo proposed the same heuristic at
the consumer.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:20:51 -07:00
teknium1 5435ac8cc4 fix(aux): clamp extra_body.reasoning effort too so auxiliary.<task>.reasoning_effort ultra never reaches the wire 2026-09-15 18:20:13 -07:00
teknium1 9e45a90488 fix: clamp aux reasoning effort once before profile projection
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.

Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.

Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
2026-09-15 18:20:13 -07:00
KoNit-K c02db64077 fix(agent): clamp auxiliary ultra reasoning effort 2026-09-15 18:20:13 -07:00
teknium1 f55d4f6747 fix(aux): bare-custom AuthError yields no custom endpoint instead of a stale env OPENAI_BASE_URL 2026-09-15 18:19:50 -07:00
teknium1 c1bbcf9712 fix(agent): hint-preview truncation log names the real remedy; invariant tests
Follow-up to the two salvaged commits (#111777, #111781 by @KoNit-K):

- agent/prompt_builder.py::_truncate_content — with queue_warning=False the
  logged line no longer tells the operator to "pin a larger
  context_file_max_chars, or use a larger-context model": the subdirectory
  hint cap is a constant neither knob raises. It now points at the read_file
  recovery the marker already discloses.
- tests/gateway/test_startup_environment_probe.py — replace the
  call-detection test with the behavioural invariant: an oversized SOUL.md in
  HERMES_HOME and a warm-up leave the truncation-warning queue empty for the
  next default-executor task (the api_server turn path runs on that executor
  without copy_context, which is how the boot warning reached a foreign
  session).
- tests/agent/test_subdirectory_hints.py — fold the new drain assertion into
  the existing oversized-hint test (same fixture) and pin that the log carries
  no context_file_max_chars advice.
- agent/AGENTS.md, website/docs/.../context-files.md — the hint cap is 32,000
  (docs said 8,000) and is fixed; document that it is logged, not surfaced as a
  chat warning.
2026-09-15 18:19:26 -07:00
KoNit-K fc7cbc7d8e fix(agent): suppress subdirectory hint truncation warning 2026-09-15 18:19:26 -07:00
teknium1 857133babd refactor: trim empty-frame salvage to the translated error only
Drop the raw JSONDecodeError belt from _is_provider_stream_empty_frame_error:
every chat stream iteration goes through _iter_provider_stream_chunks, which
already translates the decode failure, so the belt guarded a path that does
not exist. Drop the explicit classifier entry for the new code: unknown codes
already resolve to the same retryable unknown verdict (verified live: identical
ClassifiedError with and without the entry).
2026-09-15 18:18:08 -07:00
moxian 71281fb266 fix(agent): contentless SSE keepalive frames retry without streaming
An empty `data:` frame (or a frame carrying only `event:` / `id:`) is a legal SSE
no-op, but the OpenAI SDK still hands it to `json.loads`, which raises
`JSONDecodeError` with an empty document. The streaming helper translated that
into `ProviderStreamError(provider_stream_non_json_data)`, nothing recognised it
as recoverable, and the main loop retried streaming -- identically -- until the
retry budget ran out: "API call failed after 3 retries: Provider stream returned
non-JSON SSE data". A degraded gateway answers EVERY streaming request that way,
so the retries only repeated the failure while the same request sent
non-streaming succeeded seconds later.

An empty document means no payload, which is a different fact from a malformed
payload: give it its own code, and when it appears before any delta switch the
session to non-streaming (the existing `_disable_streaming` mechanism, already
used for "stream not supported", Bedrock IAM denials and adapter-returned final
responses) so the retry goes out on a channel the degraded gateway can answer.
Non-empty payloads keep today's fatal semantics; failures after deltas keep the
existing stream-drop handling.

Tests: the turn-level recovery test drives the real SDK decoder over a real
httpx response and is red on base (three identical streaming attempts, turn
lost); the boundary test pins the malformed-payload path so a future "ignore bad
frames" change cannot swallow genuine provider errors.
2026-09-15 18:18:08 -07:00
teknium1 423bc7e4e4 fix(credential-pool): hydrate on-disk env rows on the openrouter branch too
The openrouter branch of _seed_from_env returns before the generic loop,
so a persisted `env:OPENROUTER_API_KEY_2` row stayed empty exactly like
the registry-provider case #103067 fixed. Fold the "declared vars + on-disk
env rows" union into one helper both branches use.
2026-09-15 11:50:48 -07:00
McClean-Newton 23afade67b fix(credential-pool): seed env-source entries not in registry tuple
Restores the env-source seeding loop in _seed_from_env that was dropped
by upstream refactors. Without this, any env-source entry whose env var
name isn't in the provider registry's hardcoded api_key_env_vars tuple
stays empty forever, gets filtered out of rotation by _available_entries,
and round-robin silently degrades to single-key behavior.

Forward-port of ba71e00db07c6263ee8d44b27dfbce2a92e6b39c onto current main
f1ccf436a2. Narrow insertion after final
env_vars list in _seed_from_env: scan only existing entries whose source
begins with env:, extract/dedupe the named variable, and let the existing
seeding path hydrate it via get_env_prefer_dotenv (preserving _Seeder,
secret-scope/profile isolation, borrowed-secret sanitization, and
suppression behavior).

Tests: 10 new tests in test_credential_pool_seed_existing_env_sources.py
covering three-key hydration, round_robin cycling, dotenv/secret-scope
resolution without os.environ bypass, duplicate dedupe, manual rows
untouched, unset remains unavailable, and sanitized persistence.
2026-09-15 11:50:48 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
kshitijk4poor 683ca44046 fix(classifier): a welcome-host 403 naming the free tier itself is the tier refusing, not a billing wall
The "says something else" guard reuses the billing table, which carries the
Nous gateway's own free-tier phrases ("not available on the free tier",
"model_not_supported_on_free_tier"). On the welcome route those words mean
exactly "the tier refused"; routing them to billing prints a credits check to
an anonymous session that has none. Leave those two phrases on the
tier_disabled path.
2026-09-15 20:44:42 +05:30
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 2e3b8ecc82 fix(free-tier): the desktop card body leaves the "To sign in" tail off; the button is the door
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 16def7b8cc fix(free-tier): the desktop renders a free-tier refusal as its own card, not an OAuth re-login
A welcome-tier 403 classifies as auth_permanent, so the desktop's error
surface mapped it to "Your Nous Portal sign-in expired" with a Nous Portal
re-login button — the chat sentence never reached the user. Terminal results
on the free route now carry a structured free_tier block (kind + the chat
sentence); agent/error_surface.py turns it into a free_tier_<kind> code on
the provider layer with the sentence as `message`. The desktop gives those
codes their own titles, shows the backend sentence as the body, and offers
"Sign in with a Nous account" (the free-tier dialog) instead of the OAuth
re-login, with Retry only where a later send can succeed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30