Commit Graph

35297 Commits

Author SHA1 Message Date
teknium1 f8eb2912ce fix: route component log files added after profile routing is already on
Adoption of a second home only wrapped the file handlers that existed at
adoption time. A later setup_logging(hermes_home=<known>, mode="gateway")
skipped _adopt_secondary_home (the home is already served) and appended a
bare _ManagedRotatingFileHandler for gateway.log, which carries no home
filter and so took every profile's gateway records. When a router is
already queued, the new handler is now wrapped in a _ProfileRoutingFileHandler
over the union of the live routers' homes.

Review finding: setup_logging(A); setup_logging(B); setup_logging(A, mode="gateway") wrote gw-b into A's gateway.log.
2026-09-15 05:23:12 -07:00
John Paul Soliva c7a942e833 fix(logging): a second Hermes home in one process gets a routed log, not an unfiltered handler
setup_logging(hermes_home=X) for a home other than the one this process already
logs for added another queued file handler with no home filter. In a dashboard
or serve backend that builds agents for several profiles, every profile's
records were therefore written into every served profile's agent.log and
errors.log — the same lines, byte for byte — and with profile routing already
on (multiplexed gateway, Desktop cron ticker) a duplicate writer for the
secondary home sat on top of the router. Adopt the new home into profile
routing instead: enable it for the union of homes when no router exists yet,
widen the live routers otherwise, and let _add_rotating_handler recognise a
router that already covers a home.
2026-09-15 05:23:12 -07:00
teknium1 d6d9e67f54 fix(compression): announce the compacting status before the lazy feasibility probe
The first compaction of a session runs check_compression_model_feasibility()
inside compress_context() BEFORE _announce_compression_start(). That probe is
network-bound (live model catalog / provider lookups; connect timeouts stack up
through proxies and slow remote gateways), so every automatic entrypoint that
reaches compress_context without its own pre-emit — the post-tool gate in
agent/turn_preflight.py::compress_after_tool_results, overflow recovery,
manual /compress — left the client with no `kind="compacting"` status for the
whole probe. On Desktop that is a bare working-row spinner with no
"Summarizing thread" label (#111294).

Move the announcement ahead of the probe in the one choke point so every
caller is covered, and retire the announced phase (force_terminal) when the
probe's hard rejection propagates so the client never stays "compacting" for
an attempt that never started.

Live repro: /tmp probe driving compress_after_tool_results with a 2 s
feasibility stand-in — before: first compacting status at t=2.29 s (after the
block); after: t=0.13 s.
2026-09-15 05:19:57 -07:00
teknium1 ee6fb0ed22 fix: clear the relay roster per sole connection, not once per below-two regime
rosterCleared was a single boolean latched the first time the peer set fell
below two connections. When the sole connection a was replaced by c between
ticks (still length 1) the flag stayed set, so c's gateway never received
the empty-roster push and kept the stale roster. Track the id of the
connection that got the clear instead and push whenever the sole
connection's id differs; reset once two or more connections relay again.

Review finding: routes [a] → tick → routes [c] → tick pushed no roster clear to c.
2026-09-15 05:19:12 -07:00
teknium1 c02a269351 fix(desktop): an empty relay route list does not spend the one roster clear
relayConnections() returns [] before the registry loads (and whenever the
host bridge is missing). The below-two clear must not latch on that empty
list, or the single connection that arrives on the next tick never gets its
roster cleared. Gate the clear on exactly one connection; formats the
salvaged test file.
2026-09-15 05:19:12 -07:00
John Paul Soliva 8b091e539c fix(desktop): the relay clears the remaining gateways' remote roster when the peer set drops below two
syncRelayRosters returned early with fewer than two connections, so a machine
removed from the registry stayed in every remaining gateway's
bot_relay/roster.json — still listed in each bot's system-prompt roster and
still a message_agent target — until a second connection appeared again.
Push the now-empty roster once when the set shrinks below two (and once at
start with a single connection, for a roster left behind by an earlier peer);
union pushes resume when a peer returns.
2026-09-15 05:19:12 -07:00
teknium1 33cd423d56 fix: remote probe watchdog kills the probe's process group, not just its child
The watchdog SIGKILLed only the direct child of the wrapper. A remote
`hermes` launcher that runs the CLI without exec leaves the hung grandchild
alive after its parent dies — exactly the broken-launcher class from
#110478 — so the probe still orphaned a process on the remote. Enabling
job control (`set -m`) around the probe start puts it in its own process
group; the watchdog now kills `-$__htp` first and the direct pid as the
fallback for shells that cannot turn job control on without a tty (dash),
where behaviour is unchanged.

Review finding: watchdog SIGKILLs only its direct child; a non-exec launcher leaves the hung grandchild orphaned.
2026-09-15 05:18:28 -07:00
teknium1 b06071c0c9 test(desktop): make the remote-watchdog probe test host-safe
The live shell leg skipped nothing on Windows (no POSIX sh) and its orphan
sweep matched ANY `sleep 30` on the host, so a busy dev box or a parallel
test failed it. Skip on win32 like the other shell-shape tests here, and
sleep a per-run unique duration so the sweep can only see our own child.
The lockfile-skew guard now rejects a kill of any literal pid instead of
only `kill -9`, so the probe watchdog's own `kill -9 $__htp` no longer
weakens it.
2026-09-15 05:18:28 -07:00
Kevin Rajan 5724cb3002 fix(desktop): kill hung remote SSH probes via a POSIX remote watchdog
runSsh SIGKILLs the LOCAL ssh child on timeout, but the remote command keeps
running as an orphan (ppid=1). A hung remote CLI (e.g. a wedged
`hermes --version`) therefore accumulates orphans on every timed-out probe.

Wrap the two remote Hermes CLI probes (--version and the serve --help
ownership probe) in withRemoteTimeout(), a pure-POSIX remote watchdog
(macOS remotes have no GNU `timeout`): the probe runs as the watchdog's
direct child and is kill -9'd remotely after 15s, before the local
20s exec timeout fires. The sleeper's stdio is detached so its orphaned
sleep cannot hold the ssh channel open on the healthy path.

Fixes #110478
2026-09-15 05:18:28 -07:00
teknium1 bc59c39cfd docs(desktop): point the copy-control inset comment at the real scroller
The salvaged comment cited tool/fallback.tsx, which does not exist; name ExpandableBlock/CodeCardBody (the .scrollbar-overlay scroller) and log-tail.tsx (the 12px sibling) instead.
2026-09-15 05:15:42 -07:00
David Metcalfe 6e04bf8e59 fix(desktop): keep the code-block copy control clear of the scrollbar
The fenced code card's scroller (CodeCardBody / ExpandableBlock) carries
`.scrollbar-overlay`, which hands the card's right edge back to the
platform's ~15-17px scrollbar lane. The hover copy control sat at a 6px
inset with a 10px icon, inside that lane and on top of the bar itself.
Inset it 16px and use the 12px icon the other corner copy controls use
(log-tail.tsx).

Salvaged from #110548 without its className change-detector test; the fix
is pure CSS (Tailwind inset/icon size) and is verified by a static trace
against the .scrollbar-overlay scroller.
2026-09-15 05:15:42 -07:00
KoNit-K fccdcdaea8 fix(desktop): expose Codex compression auto-raise 2026-09-15 05:14:59 -07:00
KoNit-K ec5c975902 fix(desktop): refetch vault sources after closed-to-open remount
Per-query staleTime: 0 so a fresh installed:false cache still refetches when the Passwords page remounts disabled and the gateway opens later.
2026-09-15 05:14:15 -07:00
KoNit-K c89a4404fd fix(desktop): refresh vault source detection on mount 2026-09-15 05:14:15 -07:00
DavidMetcalfe 6201a8236f fix(desktop): drop unsaved credential edits when the settings target changes
The shared Settings "Applies to" target re-fetches `vars`, but the in-flight
edit and revealed maps are keyed by var name alone and were not reset with it.
A value typed while targeting one profile therefore survived a switch to
another, where the still-live Save button wrote it into the profile then being
targeted — the credential landed in the wrong profile.
2026-09-15 05:13:32 -07:00
teknium1 7c47e5517b test(desktop): trim the live-draft helper coverage to two invariants
Keep the case the fix exists for (editor holds text the mirror has not
seen yet) and the pre-mount fallback; the whitespace, stale-non-empty and
cleared-editor cases all reduce to "the helper returns composerPlainText
of the editor", which the first test already pins.
2026-09-15 05:12:33 -07:00
DavidMetcalfe 7679556d90 test(desktop): tighten the live-draft race coverage and dedupe the sibling reads
Tighten the empty-editor case against a stale non-empty mirror, add JSDOM body cleanup for editorWith, and route the bare-Enter and Cmd/Ctrl+Enter live reads through liveComposerDraft.
2026-09-15 05:12:33 -07:00
DavidMetcalfe bc62b270ed fix(desktop): read the live composer draft for the recall guard
draftRef is a once-per-frame mirror, so an ArrowUp in the same frame as a keystroke or paste saw the pre-keystroke text and let the sent-message recall overwrite what the user just wrote.
2026-09-15 05:12:33 -07:00
teknium1 e5d57c7fc5 refactor(desktop): workspace lanes read the show-all preference at the leaf
The salvaged fix threaded `showAllSessions` through three components
(sessions-section → entered-content → RepoFlatSection → workspace-group) as
a prop. The repo's TypeScript rule is that a leaf subscribes to the shared
atom instead of state being threaded through intermediaries, and
overview-row.tsx already reads `$sidebarShowAllSessions` that way. Drop
the prop plumbing and let SidebarWorkspaceGroup `useStore` the atom, which
also covers the bare `groups.map(SidebarWorkspaceGroup)` branch in
sessions-section.tsx that the prop version left paged.

The regression test flips the atom instead of re-rendering with a prop.
Still red against origin/main source, green here.
2026-09-15 05:11:46 -07:00
KoNit-K 38a4e909c6 fix(desktop): honor show all sessions in projects 2026-09-15 05:11:46 -07:00
teknium1 6f69d170e5 fix: classify .env-routed keys, hyphenated headers and ${VAR} placeholders in config get
`_is_secret_config_key` matched only an exact set plus four suffixes, so
credentials that `hermes config set` routes to .env but end in bare `_KEY`
(FAL_KEY, VOICE_TOOLS_OPENAI_KEY, API_SERVER_KEY) or AWS_SECRET_ACCESS_KEY
printed raw with redaction on. Every `_is_env_config_key`-routed key is now a
credential unless its suffix names a non-secret shape (_URL/_HOST/_USER/_ID/
_DOMAIN/_SCHEME); `_key`/`_access_key` join the leaf suffixes. Header names are
folded `-`→`_` before matching so `mcp_servers.<s>.headers.X-API-Key` masks on
get and on the set echo. Bare `auth` leaves the exact set: `mcp_servers.<s>.auth`
is the documented `oauth` mode enum and rendered as `***`. Unresolved `${VAR}`
placeholders are printed as-is so the operator can see which env var the config
references.

Review finding: `config get` printed FAL_KEY / AWS_SECRET_ACCESS_KEY / X-API-Key raw, masked `auth: oauth`, and hid `${VAR}` placeholders.
2026-09-15 05:08:55 -07:00
teknium1 a9a8a3fa2e fix(config): hermes config get masks credentials on every path; --raw opts out
`hermes config get providers`, `config get providers.<p>.api_key`, `config get
<PROVIDER>_API_KEY` (the .env-routed branch) and `config get mcp_servers.<s>.env.X_API_KEY`
all printed the full credential. The agent runs this command from sessions whose transcripts
persist and get forwarded (a Gemini key surfaced in a Discord DM log), so `print` output is a
leak path the logging redactor never sees.

`get_config_value` now applies the structural masker used by `config show` before printing,
honouring `security.redact_secrets` (default on), with a `--raw` flag for operators/scripts
that need the real value. `_is_secret_config_key` extends the exact-name set with the same
`*_API_KEY / *_TOKEN / *_SECRET / *_PASSWORD` suffixes `_is_env_config_key` already routes to
.env, so env-map leaves under `mcp_servers.*.env` mask too, and the `config set` echo uses the
same predicate.

Slim redo of #84153 by @webtecnica (same direction: mask in get_config_value; dropped the
redact_url_query_params re-export and the separate redaction-enabled reader in favour of
agent.redact._redact_enabled, which already resolves the profile-scoped policy).

Fixes #110758
Fixes #84106
2026-09-15 05:08:55 -07:00
teknium1 2034126e0d fix: warn about existing WAL on cross-VM fs on the vulnerable-SQLite path too
apply_wal_with_fallback returns via _apply_delete_for_wal_reset_bug on
WAL-reset-vulnerable SQLite builds (Debian 12 / Ubuntu 22.04 system Pythons,
the reporter's pinned image) before the #110848 existing-WAL cross-VM check,
so exactly the deployment class the fix targets still got zero startup signal
while doctor flagged it. Share one helper between both early-return paths and
key the once-per-process dedupe on the DB path instead of the label so a
gateway serving several profiles hears about each database. The docker doc
remedy now uses the image's python3 (it ships libsqlite3 but no sqlite3 shell).

Review finding: vulnerable-SQLite early return skipped the cross-VM ERROR; dedupe per label; sqlite3 shell absent from image.
2026-09-15 05:02:30 -07:00
teknium1 60d94fd8f4 fix(state): warn when an existing WAL state.db sits on a virtiofs/9p mount; doctor + docs
d8dcdfd620 (v2026.9.14) made apply_wal_with_fallback refuse to ENABLE WAL
on a fresh database whose directory is on a cross-VM bind mount (virtiofs/9p),
but a database that was already WAL on such a mount kept WAL — correctly, we
never live-downgrade under other openers — and emitted nothing. The operator
in #110848 ran exactly that shape (Podman applehv virtiofs bind mount) and got
"database disk image is malformed" within a minute with no prior signal.

- apply_wal_with_fallback: in the on-disk-WAL branch, log a once-per-process
  ERROR ('cross_vm_fs_existing_wal') when the DB file is on a cross-VM
  filesystem, naming the two remedies (offline PRAGMA journal_mode=DELETE
  after stopping every process + database.journal_mode: delete, or move the
  database to a native/named volume). The fresh-DB refusal is unchanged.
- hermes doctor: _report_database_journal_modes flags a WAL database on a
  cross-VM filesystem with check_warn and the same remedy (ranked above the
  WAL-reset exposure warning; the exposure bookkeeping is kept).
- docs: docker.md gains "Filesystem requirements for state.db in containers";
  configuration.md's database comment no longer implies operators must set
  delete by hand on virtiofs.

Detection stays /proc/self/mountinfo-based (runs inside the Linux container
on macOS/Windows hosts). locking_mode=EXCLUSIVE is deliberately not adopted:
gateway, cron and workers open state.db concurrently.

Fixes #110848
2026-09-15 05:02:30 -07:00
teknium1 8fecfeea05 test(state): reaction test reopens a store carrying the broad display-identity trigger
The narrowed CREATE TRIGGER IF NOT EXISTS DDL only reaches fresh stores by
itself; _execute_ddl_skipping_settled_triggers drops and recreates any trigger
whose stored body differs, which is what upgrades existing stores. Pin that
path: install the pre-narrowing trigger text, reopen, assert the narrowed body
before measuring the reaction write.
2026-09-15 05:01:48 -07:00
fangliquan 7e4180d1e8 fix(state): avoid display metadata backfill rewrites 2026-09-15 05:01:48 -07:00
teknium1 5ffdf823ae fix: compare legacy reset children against the parent's started_at, not ended_at
The reopen backfill guard `child.started_at >= parent.ended_at` compared against
the parent's CURRENT end boundary. A genuine markerless legacy reset child whose
parent was later reopened and re-ended has started_at earlier than that second
boundary, so it was no longer frozen with `_reset_from`; once end_reason cleared
it dropped out of /sessions as ephemeral — the multi-cycle gateway-peer shape the
issue describes. Comparing against the parent's started_at still rejects
children that predate the parent while keeping every earlier-boundary reset
child; the marker and source exclusions are unchanged.

Review finding: cycled parent's earlier reset child lost its `_reset_from` stamp and vanished from the session list on reopen.
2026-09-15 05:00:59 -07:00
teknium1 61e8f1cd82 test(sessions): trim the reopen backfill regression to the two red-on-base invariants
The positive case (a markerless legacy reset child is still frozen with _reset_from on
reopen) is already covered by tests/gateway/test_resume_command.py::
test_sessions_full_keeps_legacy_reset_child_after_parent_resume, so the third test was a
change-detector duplicate.
2026-09-15 05:00:59 -07:00
KoNit-K a5aa2379f0 fix(sessions): preserve branch lineage on reopen 2026-09-15 05:00:59 -07:00
teknium1 1f08821bd0 fix: derive suggest globs from the raw command, class-key when redaction hits them
Redaction ran before glob derivation, so a mined `GITHUB_TOKEN=ghp_… git push`
became the pattern `GITHUB_TOKEN=*** git *` and `--apply` persisted it to
config.yaml. In the permanent-allowlist matcher `***` is three fnmatch
wildcards, so that entry pre-approved any `GITHUB_TOKEN=… git …` command
(`sudo git push --force`, `chmod -R 777 /etc git x`) ahead of the dangerous
command detector. Globs now come from the raw normalized command; when the
redacted form would yield a different glob the command is proposed under its
dangerous-class key instead. Redaction stays for example rendering only.

Review finding: masked `***` inside a persisted glob widened the allowlist to arbitrary `KEY=… git …` commands.
2026-09-15 04:57:29 -07:00
Teknium c60fab351f fix(cli): hoist redactor import, trim suggest redaction tests, document masking
- Move the `agent.redact` import out of the per-record loop in build_proposals.
- Keep the two display-boundary invariant tests (rendered `e.g.` line, --json
  examples), drop the unit-level duplicate that asserted the same masking.
- website/docs: state that mined examples are masked at display time only;
  state.db itself is unchanged.
- contributors/emails: map kokhlo's commit email.
2026-09-15 04:57:29 -07:00
kokhlo a9db7ed5a8 fix(cli): mask credentials in approvals suggest proposals
Mined terminal commands can carry credentials (URL userinfo, env
assignments, bearer tokens). Patterns and examples are printed to the
operator and can be persisted into config.yaml, so pass them through
redact_sensitive_text like every other display boundary. Danger
classification still sees the raw command.
2026-09-15 04:57:29 -07:00
teknium1 5117e3b3a0 fix: key the check_fn cache by the same served-profile predicate as the MCP registry scope
_mcp_registry_scope() became profile-keyed for served profiles with the
multiplex flag off, but check_fn_cache_scope() still returned None in that
mode, so the process-wide availability cache stayed keyed (fn, None) across
profiles. A served profile whose mcp__x__* check_fn now correctly resolves to
its own (absent) connection cached False for the TTL window and the launch
profile that owns the live connection lost its tools for that window.

Both sites now call one helper, agent.secret_scope.serves_routed_profile()
(multiplex on, or a HERMES_HOME override naming a home other than the
process home), so the registry scope and the cache key can no longer drift.

Review finding: served profile B's check_fn verdict shadowed the launch profile's live mcp tools via the unscoped check_fn cache.
2026-09-15 04:56:40 -07:00
teknium1 57c4e1d963 fix(mcp): served profiles keep their own MCP connections without the multiplex flag
A dashboard/desktop backend (and the per-profile cron ticker) serves sessions of
several profiles through the HERMES_HOME contextvar override while
gateway.multiplex_profiles stays off. _mcp_registry_scope() keyed every MCP
connection by the bare server name in that mode, so the first profile to
discover `zernio` owned the only connection and every later served profile —
including one whose config carries a different Authorization header — called
the server through it and got the other account's data back (#111151).

The registry scope now follows the served home: a routed profile (an override
naming a home other than the process home) gets the same per-profile overlay
the multiplexer uses, so a same-named server with other credentials is a
separate connection, discovery for profile B is a connect candidate instead of
"already connected", and status/tool views stay per profile. Single-profile
processes (no override) keep bare keys, byte-identical to before.

Fixes #111151

Credit: #111158 by @KoNit-K located the inert flag on hermes_cli surfaces; its
fix (activating fail-closed multiplex secret scoping from config.yaml on the
dashboard) is not taken — the connection-key seam, not the secret-scope mode,
is what leaks the connection, and flipping the process-wide mode from the
dashboard would change credential resolution for every code path in it.
2026-09-15 04:56:40 -07:00
teknium1 8a2996503a fix: skip Bitwarden URIs marked match=Never when collecting fill origins
Binding every saved URI made a URI the user explicitly set to "Never"
match (match=5) a valid fill target — wider than the manager's own policy.
Filter those out; all other URIs keep exact-origin semantics.

Review finding: match=5 (Never) URIs became fill origins.
2026-09-15 04:56:01 -07:00
teknium1 0d1a3e704a docs(vault): manager items fill on every saved website origin 2026-09-15 04:56:01 -07:00
liuhao1024 5fff41e52e fix(vault): bind manager logins to every saved web origin
1Password/Bitwarden items can carry several websites, but both backends
collapsed the item's urls[]/uris[] to the first origin that normalizes,
so browser_vault_fill refused every other explicitly saved origin with
origin_mismatch. Reordering the URLs in the manager just moved which
single origin worked.

VaultItemMeta now carries allowed_origins (every normalized, deduped
web origin; origin stays the first/primary one). Fill matching stays
exact-origin against that list — no wildcard, parent-domain or subdomain
inference — and the in-page synchronous check pins the origin actually
matched via build_fill_js(expected_origin=page_origin). App URIs such as
androidapp:// never widen the fill set.
2026-09-15 04:56:01 -07:00
liuhao1024 ddd084a24c test(mcp): cover a properties-map entry literally named required (#110530) 2026-09-15 04:55:14 -07:00
liuhao1024 798cc60f4c fix(mcp): repair schema-map keywords per-entry in _repair_object_shape
_repair_object_shape recursed over every dict value as a schema node,
including the properties map itself. When one of the map's keys was
literally named properties/required, the missing-type heuristic fired on
the map and injected a bogus "type": "object" string as a parameter,
400ing the whole tool array on strict providers.

Recurse into map values only for the mapping-valued keywords
(properties, patternProperties, $defs, definitions, dependentSchemas),
matching the in-tree precedent in _rewrite_local_refs.

Fixes #110530
2026-09-15 04:55:14 -07:00
KoNit-K 20d80bb36f fix(mcp): refresh expired cold-loaded OAuth tokens 2026-09-15 04:54:30 -07:00
teknium1 417b707f59 fix: refresh the credential that 401'd on the Codex /usage retry
Two holes in the forced retry. With an explicit live-agent api_key the
retry dropped it and re-resolved from the singleton/pool, so a 401 on pool
entry B rendered account A's usage — the cross-account leak the surrounding
comments say the tiering prevents. The retry now force-refreshes the pool
entry that issued the key (or the singleton when it IS the singleton token)
and fails open when neither matches.

In a pool-only setup (empty singleton) resolve_codex_runtime_credentials
returned the pool token before consulting force_refresh, so the retry resent
the identical revoked bearer. The pool-only branch now rotates the pool
entry through try_refresh_matching when force_refresh is set.

Also: when the forced refresh itself raises inside redeem_codex_reset_credit,
surface the original 401 (re-login hint) instead of the refresh error.

Review finding: explicit-key force refresh re-resolved another account; pool-only tier ignored force_refresh.
2026-09-15 04:53:38 -07:00
fangliquan 2a6a9b0427 fix(auth): retry Codex usage after 401 2026-09-15 04:53:38 -07:00
teknium1 9336fb11cd fix: honour HERMES_CODEX_BASE_URL on the raw Codex client too
The raw_codex branch of _resolve_openai_codex_branch (main agent built without
explicit creds via agent_init._routed_client_kwargs, and the mid-turn fallback
chain) still hardcoded the official Codex endpoint while the pooled/aux/singleton
paths honoured the override, so a proxy user's main agent silently bypassed it.
Hoist the profile-scoped read into _codex_base_url_override and use it in both
builders; the Cloudflare identity headers follow the resolved base_url.

Review finding: raw_codex client ignored HERMES_CODEX_BASE_URL while pooled/aux/singleton honoured it.
2026-09-15 04:52:47 -07:00
teknium1 70ff4863d7 fix(aux): read HERMES_CODEX_BASE_URL through the profile secret scope
The aux Codex client already resolves its API-key env vars via _scoped_key_env
so a multiplexed profile never borrows a sibling's value; the endpoint override
follows the same rule instead of a bare os.getenv.
2026-09-15 04:52:47 -07:00
joaomarcos b62bb2a3d5 fix(codex): honor HERMES_CODEX_BASE_URL on pooled and aux codex client resolution 2026-09-15 04:52:47 -07:00
KoNit-K 7e06625687 fix(auth): avoid auto provider resolution recursion 2026-09-15 04:52:00 -07:00
teknium1 d4e0df1e8d fix: keep the OpenCode keyless Authorization blank on the async aux client
_to_async_client rebuilds default_headers from scratch and dropped the blank
Authorization set by _create_openai_client, so every async aux call (the path
aux tasks actually take) for a free-tier OpenCode model still shipped the
keyless placeholder as a bearer and 401'd. Re-apply the same keyless policy on
the async twin.

Review finding: fourth client builder (_to_async_client) missed the keyless header policy.
2026-09-15 04:51:15 -07:00
teknium1 e77060e081 fix(agent): keyless OpenCode placeholder blanks Authorization under the paid profile too
A free slug (deepseek-v4-flash-free) re-resolved under the paid `opencode` profile — via
`/model --provider opencode`, a fallback switch, or a config default — resolves to the
keyless placeholder, but the primary client builder blanked the Authorization header only
when `agent.provider == "opencode-free"`. The placeholder then went on the wire as a
bearer, the Zen relay 401'd every request, and the credential pool (size 0) had nothing to
rotate: the loop reported in #110831.

Key the header policy on the placeholder as well as the provider, matching agent_init and
auxiliary_client which already do so.
2026-09-15 04:51:15 -07:00
teknium1 21d28dfe00 fix: key the ACP named-provider swap on is_user_defined, not slug-name collision
build_model_state dropped every inventory row whose slug name matched a
named `providers:` entry and relabelled the current provider `custom:<key>`
unconditionally. When a `providers:` key shadows a canonical provider name
(providers.openrouter: -> proxy with a models: list), a session genuinely
running on openrouter.ai lost all canonical rows and its current id became
custom:openrouter:..., which resolves to the proxy base_url — picking the
"current" row silently re-routed the session to a different endpoint.

Only rows flagged is_user_defined are replaced by the named catalogs now,
and the current provider is promoted to custom:<key> only when the session
base_url matches that entry's api_url or no canonical row for the raw key
exists (the reported relay/MixedCase class keeps its custom:<key> id).

Review finding: named entry shadowing a canonical provider dropped canonical rows and re-routed the current session to the proxy.
2026-09-15 04:50:31 -07:00
KoNit-K dcaf346856 fix(acp): deduplicate configured provider model ids 2026-09-15 04:50:31 -07:00