Commit Graph

28271 Commits

Author SHA1 Message Date
Carry00 edac49e473 fix(skills): stop syncing bookkeeping dirs to sandboxes
iter_skills_files() walked the skills tree with a bare rglob("*"), so the
.hub download cache, .archive, curator backups, and any node_modules/.git
under a skill package were uploaded to the sandbox on every sync. The
sandbox never reads them: skill content is resolved host-side.

EXCLUDED_SKILL_DIRS is already the canonical exclusion set, honoured by
discovery and backup. Apply it to the sync path too, across all three
roots iter_skills_files() walks (local, external, project-local), and add
.curator_backups to the set.

Measured on a local install: 900 files / 67.3 MB -> 771 files / 8.4 MB.

This is not just wasted bandwidth on the SSH backend, where the oversized
payload can exceed the 120s _ssh_bulk_upload deadline and surface as the
agent hanging on every tool call.

The filter intentionally does not reuse is_excluded_skill_path(), which
also prunes references/, templates/, assets/ and scripts/ -- those hold
support files and bundled scripts the sandbox does read and execute.
2026-09-03 01:33:18 +05:30
kshitijk4poor c4e394cdf8 refactor(auxiliary): one predicate for the critical-path retry skip
The sync and async retry sites each re-derived the same three-clause
decision (critical task + full-budget timeout + not a no-progress fail) with
their own copy of the rationale — which is exactly how the async site drifted
in the first place. _should_skip_same_provider_retry() now owns the rule and
its carve-out next to _TIMEOUT_NO_RETRY_TASKS; both sites call it.

Behavior-preserving: same clauses, same exception object, same outer guard.
2026-09-03 01:33:00 +05:30
kshitijk4poor 8115ff897a fix(auxiliary): mirror the no-progress carve-out on the async timeout skip
f50b5bb0fa taught the sync retry site to keep the cheap same-provider retry
when a Codex stream dies inside the 60s no-progress window (zero output),
skipping straight to fallback only on a stall or hard-ceiling timeout. The
async site never got that carve-out, so after widening the skip to vision
(#97572) an async vision call on a stillborn stream would have jumped to
fallback where the sync path retries. Both sites now apply the same rule.

Adds the async twin of the vision-skip test and a no-progress-still-retries
guard for the async site.
2026-09-03 01:33:00 +05:30
xmhua 827cf6fa01 fix(auxiliary): skip same-provider retry on a vision full-budget timeout
Issue #54465 established that a same-provider retry after a full-budget
timeout costs a second whole `timeout` window before the fallback chain is
reached, doubling the user-visible stall, and that compression must not pay
it because it sits on a critical path. The guard added for that is spelled
`task == "compression"`, so vision — which sits on the interactive path —
still retries.

The cost is the same and the stall is more visible: the turn holding the
image cannot answer, and because turns are serialised the following user
messages queue behind it. Two sequential full-budget timeouts on an
unhealthy vision provider is a long stall for something the fallback chain
could have served immediately.

Replaces the string comparison at both retry sites (sync `call_llm` and
`async_call_llm`) with `_TIMEOUT_NO_RETRY_TASKS = {"compression", "vision"}`,
so the two paths cannot drift again. Behaviour is unchanged for every other
task: fast blips (a streaming-close or a 5xx) still retry, and only
full-budget timeouts on those two tasks skip straight to fallback.

Tests: vision now falls straight through to fallback with the primary tried
exactly once, and a non-critical task still gets its one same-provider
retry, so the change stays scoped. Reverting the source change fails the
vision test and leaves the scoping test green.

Not the same as #51513, which fixes five separate defects in the vision
fallback chain (capability detection, sync/async client misuse, geo-block
and RemoteProtocolError classification, and chain iteration). This is about
what happens before that chain is reached.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 01:33:00 +05:30
Teknium 30b83ab7b1 chore: map contributor email for RobbertC5 2026-09-02 12:34:26 -07:00
Teknium 10088c569e fix(update): hermes update no longer hangs on a GitHub username prompt
GitHub answers anonymous fetches with HTTP 401 during outages (and for
renamed/private repos). git then prompts `Username for 'https://github.com':`
on the inherited terminal and `hermes update` sits there — users read it as
Hermes demanding a GitHub login.

Every network git call in the updater (fetch/pull/push, apply + --check +
fork sync) now runs with GIT_TERMINAL_PROMPT=0 / stdin=DEVNULL, so the 401
fails fast into the fetch-failure classifier, which now reports it as a
GitHub-side rejection (likely outage) rather than blaming the user's
credentials. Credential helpers/askpass are left configured so private-fork
origins still authenticate.

Live repro: PTY-attached update --check against a 401 origin hung 15s+ on
the prompt before; exits rc=1 in 0.2s with the diagnosis after.

Same class as #73751 (@Frowtek, pre-main.py decomposition); passive banner
half salvaged from #101421 (@RobbertC5).
2026-09-02 12:34:26 -07:00
RobbertC5 6e775907d7 fix(cli): keep passive update checks noninteractive 2026-09-02 12:34:26 -07:00
Teknium 1cb3ab6173 fix(gateway): /goal gate add now requires an explicitly configured admin (salvage #91677)
A goal quality gate is a shell command persisted by `/goal gate add` and
later executed with `subprocess.run(shell=True)` at every goal turn
boundary (run_gate), with no approval prompt. On the messaging gateway,
slash access is backward-compatible: with no `allow_admin_from` list
configured (the default), every *allowed* chat user is treated as
unrestricted. So an allowed but non-admin remote sender could add an
arbitrary shell command and get authenticated RCE as the Hermes process
account.

Gate ONLY the shell-creating operation (`gate add`) behind the existing
fail-closed explicit-admin check (`_resume_caller_is_admin`, the same one
that guards cross-origin `/resume`). `gate list` / `remove` / `clear`
stay open so a non-admin can still inspect and recover. CLI/TUI/Desktop
`gate add` is local (the user already has host access) and is unchanged.

Salvaged from #91677 by @unsupportedpastels — narrowed to the single
shell-creating op and reusing the existing admin helper rather than
renaming it. Contributor's regression tests preserved.

Co-authored-by: unsupportedpastels <theoldwizard123@pm.me>
2026-09-02 10:55:57 -07:00
Teknium e1d773a6a6 chore: map contributor email for @globalvet2025 2026-09-02 10:55:46 -07:00
globalvet2025 b36489be37 fix: gate Nous auth-refresh message behind verbose mode
The 'Nous agent key refreshed after 401' message used a bare print(),
making it always visible. The equivalent xAI/Codex, Copilot, and Anthropic
auth-refresh messages all use _buffer_vprint() (verbose-gated). This brings
the Nous path in line with the others so the routine ~15min OAuth key
refresh no longer prints noise on every retry.
2026-09-02 10:55:46 -07:00
Teknium 98c45f8c74 test(aux-client): keep only the two end-to-end 401 eviction invariants 2026-09-02 10:55:46 -07:00
cryptoyasenka 3265801900 fix(aux-client): evict expired auxiliary client on 401 by refreshing under the lookup cache key
On the default Nous config, call_llm acquires the auxiliary client via
_get_cached_client(resolved_model=None), so the cache key's model
element is "". On a 401, _refresh_nous_auxiliary_client rebuilt the
client but keyed the new entry on the resolved wire model (final_model,
e.g. "Hermes-4-405B"). The fresh client therefore landed under a
different key than the lookup, and the stale expired-credential client
under "" was never overwritten: every auxiliary call kept hitting the
dead client, 401ing and forcing a credential portal round-trip on each
request instead of self-healing after the first refresh.

The auto-provider dimensions had the same divergence: call_llm and
async_call_llm dropped task at both acquisition sites, and the async
path additionally dropped main_runtime at acquisition and at both of
its refresh sites, so the refreshed client shadowed the stale one under
a divergent (provider, task, model, runtime) key.

Pass the original lookup model (which may be None) into the refresh as
a separate lookup_model argument used only to build the cache key,
while the resolved model is still stored as the entry's usable model
and returned to the caller. Thread task into both acquisition sites and
main_runtime into the async acquisition and both async refresh sites,
so sync and async compute the same cache key on acquire and on refresh.
The stale client is now overwritten in place instead of lingering under
an orphaned key, preserving the per-model cache keying introduced in
on the default config.

Add end-to-end regression tests that drive the real call_llm and
async_call_llm through the real client cache; the existing 401 tests
patch _get_cached_client wholesale and so cannot observe
acquire/refresh key divergence.
2026-09-02 10:55:46 -07:00
hermes-seaeye[bot] 4e66da6eeb fmt(js): npm run fix on merge (#101497)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 17:45:59 +00:00
Brooklyn Nicholson 6ef6691960 fix(desktop): a saved comment crop shows its own marker, and only its own
Two faults in the crop that "Add N comments" attaches, both visible in a
two-comment batch on one page.

The marker was missing. showDraft sets the marker's style and resolves, but
resolving only means the property is set — the compositor has not drawn it.
capturePage then photographed the frame before the marker existed, so crops
arrived outlined in blue with no number, while the prompt line said "Image N
marks the target in blue". beginCapture now waits two animation frames, which
puts the shot after the paint.

The wrong marker could appear. Saved pins stay drawn on the page, so any pin
within the crop padding of the new element landed inside the shot: a comment
on a heading came back carrying the marker belonging to the comment on the
paragraph below it, pointing the agent at the wrong element. Saved pins and
hover chrome are hidden for the duration of the shot and restored after.

Restoring runs in a finally, so a capture that throws cannot leave every saved
pin invisible on the page, and the guest calls are best-effort — a torn-down
overlay degrades to the old unbracketed shot rather than failing the capture.
2026-09-02 12:40:00 -05:00
Teknium f6234d00c5 fix(security): close GitSpawn RCE class — malicious repo .git/config no longer executes on context gathering (GHSA-7x36-8jrh-v4pw)
Hermes gathers workspace context by running git against the session
directory automatically — the coding-workspace snapshot, gateway
project-tree build, /diff, @diff|@staged context refs, goal-gate
fingerprint, and -w startup worktree add — before any prompt, tool call,
approval, or trust gate. Those probes ran the system git without
stripping the repository's own config, so a repo delivered as files with
its .git directory intact (a shared zip, sync folder, or USB stick;
git clone never transfers .git/config) could set an execution-sink git
setting and get arbitrary host code execution as the user with nothing
on screen.

- core.fsmonitor / core.hooksPath / pager / editor / credential helper:
  neutralized by routing every automatic probe through
  noninteractive_git_env(), which pins those keys to inert values via
  GIT_CONFIG_* and ignores global/system config. bounded_git_probe (the
  reported sink, coding_context._git + tui_gateway.git_probe) now defaults
  to that env; worktree-add, working_diff, web_git, context_references,
  goals, and subagent_worktree route through it too.
- Attribute-scoped [diff "x"] command=/textconv= drivers: the attacker
  names the driver in .gitattributes, so GIT_CONFIG_KEY overrides can't
  enumerate them. Added harden_git_argv(), which inserts
  --no-ext-diff --no-textconv on diff-rendering subcommands (diff/show/
  log/blame) only — status et al reject the flags. Both flags required
  (verified empirically; each alone leaves the other live).

Builds on the noninteractive_git_env config-scrubbing from the
gemini-cli #28792 port. Real-git E2E regression suite arms a malicious
repo and asserts every automatic path neutralizes fsmonitor, hooks,
external-diff, and textconv; a baseline test proves the repo is armed.
2026-09-02 10:33:43 -07:00
Teknium 413b6ba3dd Port from google-gemini/gemini-cli#28792: harden internal git env
(cherry picked from commit 09bb9c3b2f20f22bdb1688b6efcb78b24f1f19d6)
2026-09-02 10:33:43 -07:00
Brooklyn Nicholson 06a4f4ab31 feat(desktop): group a comment batch by page region so it lands as a few tasks
Twenty-three comments arrived as twenty-three flat blocks, so the agent made
twenty-three todos and ground through them one at a time. They now arrive
grouped by where they sit in the page, with a line telling the agent to work
the groups rather than the comments.

The renderer groups on structure, not meaning. Whether a comment is a UI nit
or a functional bug is a judgment only the model can make, and prose-matching
it here would be wrong constantly; which pins share a DOM subtree is something
the selector already answers. That split is also the one that makes parallel
work safe — grouping by theme instead ("all the spacing ones") cuts across the
same components and puts several workers in the same files, so the guidance
says to hand out whole groups and never to regroup by theme.

Grouping compares ancestor paths, so a heading and a paragraph in one card
stay together instead of becoming two singletons. Depth is derived rather than
tuned: descend the shared prefix until it stops being shared, then sub-split
any group still holding more than a third of the batch — without that pass a
normal page buries every section under `main`. Batches under four comments,
and batches that all land in one region, stay flat.

Grouping is advice in the prompt, never an action: the renderer does not spawn
or delegate anything. That stays the agent's call.
2026-09-02 12:29:25 -05:00
Brooklyn Nicholson e4bda3ff77 feat(desktop): browser comments carry the element's selector, markup, and styles
Comment mode shipped the crop and the note, so an agent got a picture of the
problem and had to grep for the element it showed. Each element comment now
also names its CSS selector, its markup, and the computed styles that decide
layout, which is what the agent needs to land in the right file.

The target line stays prose — it is what the user pointed at — and the DOM
detail rides labelled lines beneath it. Area pins have no element, so they
still get only the crop and the note.

Markup is redacted in the guest before it crosses to the host: password and
hidden input values, and any attribute reading as a key/token/secret, are
replaced with [redacted] on a clone, so a page's secrets never reach the
composer or the model. It is clipped to a 600-char budget so one comment
cannot paste a whole section.

AnnotateIdentity was a hand-copy of CompactIdentity that had already drifted;
it is now an alias, so the guest, the pin, and the packer cannot disagree
about the shape again.
2026-09-02 12:29:25 -05:00
Teknium 1b49ad9be9 fix(dashboard): fold the one-field nous config section into the agent settings tab 2026-09-02 10:22:30 -07:00
Teknium 5c85f9de2d chore: map contributor email for @olopez25 2026-09-02 10:22:30 -07:00
Teknium 457d9b73f6 test/docs(nous): trim keepalive tests to the two invariants; document nous.keepalive_interval_seconds 2026-09-02 10:22:30 -07:00
olopez25 c6c8c74c70 Move the keepalive interval to config.yaml and tighten the schedule guard
Addresses review feedback on #84928.

The tick interval was exposed as HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS.
AGENTS.md reserves .env for credentials and puts behavioural thresholds in
config.yaml, so the knob moves to `nous.keepalive_interval_seconds`,
following the existing `vertex:` section's precedent for non-secret
provider settings. The env var is dropped rather than bridged: it was never
released, so nothing depends on it. Adding a key to a new section is handled
by the deep-merge, so no _config_version bump is required.

test_keepalive_interval_fits_inside_the_token_lifetime asserted
`900 < 899 * 4 - 120`, which is true for any realistic interval and could
never fail. It also tested the wrong value: the configured constant is only
a ceiling, while the schedule that ships is the derived tick. Replaced with
an assertion over the derived tick for each observed lifetime, which does
fail if the derivation constants regress -- verified against both
TICKS_PER_LIFETIME=1 and MIN_INTERVAL_SECONDS=5000.

Also adds coverage for an unreadable config.yaml, which must fall back to
the module default rather than take the keepalive thread down.

pytest tests/hermes_cli/test_nous_auth_keepalive.py -> 9 passed

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 10:22:30 -07:00
olopez25 fd05b97913 Derive Nous keepalive tick from the issued credential lifetime
Follow-up to the interval fix: ticking faster narrowed the gap but did not
close it, and the hardcoded interval was wrong for short-lifetime accounts.

Two independent problems, both needed:

1. Lifetime is not a constant. Installs have been observed issuing ~3594s
   and ~899s (see #35752, which reports expires_in: 899 while this account
   reports 3594). Any hardcoded interval is right for one and wrong for the
   other. The tick now derives from the lifetime the server actually issued,
   capped by the configured interval and floored at 60s. The access token and
   the invoke agent key carry separate lifetimes, so the shorter one governs.

2. Ticking faster alone never closes the gap. The refresh only fires once a
   credential is within ACCESS_TOKEN_REFRESH_SKEW_SECONDS (120s) of expiry,
   so a tick spaced wider than that window steps straight over it. With a 900s
   tick against a 3594s lifetime the last tick before expiry still saw 894s
   remaining, declined to refresh, and the credential died 6s before the next
   one. The keepalive now asks "will this outlive my next tick?" (tick + skew)
   rather than reusing the request path's bare skew.

Measured against both observed lifetimes:

  lifetime 3594s: was never refreshed proactively; now refreshes 900s early
  lifetime  899s: was never refreshed proactively; now refreshes 227s early

The lifetime is re-read every pass rather than cached, since it can change
when the account, plan, or server-side policy does -- exactly the case the
keepalive exists to cover.

min_access_ttl_seconds is threaded through as an optional parameter, so the
request path keeps its existing 120s behaviour and only the keepalive widens
the window.
2026-09-02 10:22:30 -07:00
olopez25 a301d19ba4 Fix reactive Nous 401s by ticking keepalive inside the token lifetime
Nous Portal access tokens carry a one-hour lifetime, and the keepalive only
refreshes once a token is within ACCESS_TOKEN_REFRESH_SKEW_SECONDS (120s) of
expiry. The tick interval was 6 hours, so it could only land inside that
2-minute window by coincidence. In practice every hour rolled over untouched
and the next inference call paid a 401 plus a re-auth round trip.

Observed on a production install: 71 "refreshed Nous runtime credentials
after 401, retrying" events across current logs.

Drop the default interval to 15 minutes, which gives four ticks per token
lifetime and leaves ample margin under the TTL-minus-skew ceiling of 3480s.
Add HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS so the interval is tunable without
a source edit, matching the existing HERMES_NOUS_TIMEOUT_SECONDS convention.
Zero still disables the thread.

Callers pass no interval, so the default is what actually shipped; the
signature now resolves at call time rather than binding the constant at
import.
2026-09-02 10:22:30 -07:00
hermes-seaeye[bot] bcef556b00 fmt(js): npm run fix on merge (#101481)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 17:18:47 +00:00
Teknium 7840a0e2d9 feat: delegation batch tags read "set N" instead of a hex id slice
Interleaved subagent fan-outs were tagged with the first 4 hex chars of the
delegation id ([b2ac 3/9]), which is attributable but unreadable. Batches are
now numbered in order of appearance per process: [set 1 · 3/9], [set 2 · 1/7].
Desktop /agents already labels groups "Delegation N", so its duplicate hex
badge is dropped.
2026-09-02 10:12:54 -07:00
Brooklyn Nicholson fdd399dfd3 test(sessions): pin display-projection parity across compaction
Assert the invariant — every display read of one session returns the same
transcript — plus the two things that must not grow with it: the model-fed
projection stays compressed, and Undo/Rewind rows stay hidden. Six of these
fail on the unfixed read paths.
2026-09-02 12:11:15 -05:00
Brooklyn Nicholson 86b204081d fix(gateway): read the compacted transcript for a warm session switch
_live_visible_history backs the payload a tab switch repaints from. Reading
the active-only projection made a compacted chat collapse to its summary on
switch while REST still served the whole thing.
2026-09-02 12:11:15 -05:00
Brooklyn Nicholson ab281990b8 fix(sessions): include compaction-archived rows in every display projection
The gateway's three display reads — session.resume, the ancestor lineage
prefix, and the warm-session payload — filtered active = 1, so a compacted
conversation rendered as its summary plus the carried-forward tail. The REST
transcript read has included the archived rows since #80680, so the same
session read two ways gave two different answers.

Extract the generation-dedupe policy the REST read carried inline into a
shared _dedupe_display_generations() and point every display projection at
it. The model-fed history stays active-only: compaction still compresses the
working context, it just no longer erases the user's transcript.
2026-09-02 12:11:15 -05:00
Teknium 0cbc6e37ac test/docs: trim seam tests to invariants, document create_client and external-process fields
Cuts the 41 contributor tests down to 8 pinning the before/after contracts
(out-of-tree provider resolves end to end, copilot-acp unchanged, broken
plugin falls through, flat-install discovery + non-provider kinds untouched).
Adds the create_client hook and process_* fields to the model-provider
plugin developer guide.
2026-09-02 09:57:39 -07:00
Alexander Prendota 2674d368e6 fix(providers): discover a provider plugin installed by hermes plugins install
`hermes plugins install owner/repo` — and the plugin index behind
`hermes plugins search` — clones into `$HERMES_HOME/plugins/<name>/`, flat, one
directory per plugin. Provider discovery only ever scanned
`$HERMES_HOME/plugins/model-providers/<name>/`.

Nothing joined the two. `PluginManager` does not close the gap either: it
classifies `kind: model-provider` and deliberately skips importing it, because
provider lifecycle is owned by `providers/__init__.py` — which was not looking
in the directory the installer writes to.

So the documented install path half-worked. The CLI reported success, wrote its
install metadata, and the provider silently did not exist: `hermes -m <it>` said
"Unknown provider" and `/model` never listed it. Verified before the fix — a
plugin at `~/.hermes/plugins/<name>/` was NOT FOUND while the identical plugin
at `~/.hermes/plugins/model-providers/<name>/` was discovered.

Discovery now also walks the flat directory, importing only entries whose
manifest declares `kind: model-provider`. Everything else there belongs to
`PluginManager`, which owns its lifecycle and consent flow — importing it here
would run third-party code behind its back, so the tests assert we don't (with
fixtures that register on import, since a fixture that merely raised would be
swallowed by `_import_plugin_dir` and prove nothing).

Manifests are parsed with PyYAML when present and a line scan otherwise, so
provider discovery gains no hard dependency; an unreadable manifest is skipped
rather than allowed to blank the registry.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Alexander Prendota 1bd0b8ed41 feat(providers): let an external-process provider ship out of tree
An external-process provider is an agent CLI Hermes drives over stdio rather
than an HTTP endpoint. Three things about it were spelled out for one vendor,
and each was a hard stop for any other:

* ``resolve_provider()`` gates on ``PROVIDER_REGISTRY``. Its auto-extend from
  ``providers/`` covered api-key providers only, so an external-process profile
  never entered it and ``hermes -m <that provider>`` died with "Unknown
  provider" before a client was ever built.
* ``resolve_runtime_provider()`` keyed the external-process branch on the
  literal ``"copilot-acp"``, so anything else silently fell through to the
  OpenRouter default instead of its own runtime.
* ``resolve_external_process_provider_credentials()`` hardcoded the binary
  (``copilot``), the argv (``--acp --stdio``), the env var names and the
  placeholder api_key — so a third-party provider would have been handed
  another vendor's CLI.

Now the profile carries what only the provider knows — ``process_command``,
``process_args``, ``process_command_env_vars``, ``process_args_env_var`` — and
the three core paths key on ``auth_type == "external_process"`` instead of a
name. copilot-acp's values move into its profile verbatim, so
``HERMES_COPILOT_ACP_COMMAND`` / ``COPILOT_CLI_PATH`` /
``HERMES_COPILOT_ACP_ARGS`` and its ``copilot-acp`` api_key placeholder behave
exactly as before; the new tests assert that alongside the out-of-tree case at
every step.

The error for a missing binary now names the provider and its own env override
instead of telling every user to install GitHub Copilot CLI.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Alexander Prendota 1131b22856 feat(providers): let a provider profile supply its own client
``create_openai_client`` was a hardcoded if-ladder: copilot-acp builds an ACP
stdio shim, gemini builds a native client, everything else gets an
``openai.OpenAI``. There was no extension point, so a provider whose wire
protocol is not OpenAI-over-HTTP could only be added by editing this function —
which is exactly why an ACP provider cannot ship outside this tree today, even
though ``providers/__init__.py`` has discovered out-of-tree profiles from
``~/.hermes/plugins/model-providers/`` and pip entry points for a while.

``ProviderProfile.create_client(**client_kwargs)`` closes that gap. It returns
``None`` by default, so every provider that wants the standard client is
unaffected and the existing ladder still runs as the fallback. copilot-acp is
migrated onto it — its hardcoded branch is gone and its profile supplies the
client in three lines, which is the same three lines an external package writes.

Resolution goes by provider name first, then by ``base_url`` prefix, so a
runtime configured only by URL still reaches its profile — matching what the
replaced ``startswith("acp://copilot")`` branch did. A profile that raises is
logged and skipped: a third-party plugin can fail to provide a client, but it
cannot take the turn down.

Also replaces the two ``isinstance`` checks in ``agent/auxiliary_client.py``
that mean "this client is complete, do not wrap it" with capability flags the
client class declares — ``HERMES_SKIP_TRANSPORT_WRAP`` and
``HERMES_SKIP_ASYNC_WRAP``, mirroring ``SUPPORTS_HERMES_TOOL_CALLS`` in
``background_review.py``. Two in-tree consumers (the ACP shim and the Gemini
native client), an out-of-tree client is covered by the same declaration, and
the hot path no longer imports those modules just to type-test.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Teknium 0ee98eda52 feat(models): add google/gemini-3.8-flash to nous + openrouter catalogs
Slots above gemini-3.7-flash (kept) in OPENROUTER_MODELS and
_PROVIDER_MODELS["nous"]; openrouter plugin fallback_models bumped
3.7 -> 3.8; model-catalog.json regenerated.

Verified live with test completions on both Nous Portal and OpenRouter
(model echo + billed). Same 1,048,576 window / 65,536 output / pricing
as 3.7-flash, so provider-agnostic metadata resolves via the existing
gemini entries and both routes bill live (official_models_api) — no
pricing snapshot needed.

Scoped to the two named providers: vertex/gemini/kilocode/gmi curated
lists, setup.py samples, and aux defaults untouched.
2026-09-02 09:46:15 -07:00
Teknium 3312947e14 fix(auth): concurrent Nous 401 recovery adopts a peer's refresh instead of re-rotating the shared grant
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.

- resolve_nous_runtime_credentials(stale_access_token=): under the store
  lock, skip the refresh POST when the on-disk token differs from the one
  that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
  pass the failed bearer through, and treat a lock TimeoutError as
  'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
  refresh/0 unrecovered.
2026-09-02 09:33:05 -07:00
Ayush Nangia 02103a2104 fix(gateway): drain detached hygiene workers (#98973) 2026-09-02 21:54:01 +05:30
Brooklyn Nicholson 195c3e5a3b fix(desktop): refuse every resume for a chat the user is deleting
The leftover-4001 fix stopped the dispatcher from CREATING a rebind
after a tombstone, but a request queued before it still fired, and the
push path (markRuntimeGone) never checked at all. Both funnel through
requestSessionResume, so guard there: a removal-pending id never gets
queued, and no consumer has to re-derive whether an id is doomed.

isSessionRemovalPending is now the single predicate. resumeSession keeps
its entry guard for requests queued before the tombstone; the session
tile gets the same guard via shouldResumeSessionTile, closing the tile
twin where a 4001 racing a delete unbound the runtime and re-armed the
resume effect against a dead id.

Drops the !freshDraftReady clause on explicitlyRequested: a gateway
switch also stages a fresh draft while leaving the URL on /:sid, where
an explicit request is the only remaining resume lever, so that clause
silently dropped plugin and SDK reselects after a connection apply.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-09-02 11:23:50 -05:00
Brooklyn Nicholson 86acfee949 refactor(desktop): give session tombstones their own store
The delete/archive tombstone atoms lived in store/projects.ts, which
imports store/session. Resume lives on the other side of that edge, so
consulting the tombstones from the resume path would have closed an
import cycle. Move the atoms and their mutators to store/session-removal
and repoint every consumer; no behavior change.
2026-09-02 11:23:50 -05:00
xxxigm 66109d3bbd test(desktop): leftover 4001 rebind must not revive a deleted session
Cover the delete-transition race and the tombstone early-return so the Resume failed toast cannot come back through those paths.
2026-09-02 11:23:50 -05:00
xxxigm 990879d688 fix(desktop): don't rebind a deleted chat from a leftover 4001 resume
A queued requestSessionResume still fired during the /:sid -> /new tick after delete, re-selected the doomed id, and toasted Resume failed / Session not found.
2026-09-02 11:23:50 -05:00
unsupportedpastels afc3d9d34c fix(copilot-acp): prefer stable session config for model selection
Use the ACP v1 session config contract advertised by session/new: locate the category=model option and apply the selected value through session/set_config_option. Retain session/set_model only as compatibility fallback for pre-configOptions agents. Reject unknown and policy-disabled values before prompting.

Verified against the installed Copilot ACP server: its model config option advertises the account-authorized choices, session/set_config_option returns the updated state, and live prompts route gpt-5.6-terra to Terra and claude-sonnet-5 to Sonnet 5.
2026-09-02 20:51:07 +05:30
unsupportedpastels a94b68ad40 fix(copilot-acp): stop substituted models impersonating the requested one
Follow-up to the session/set_model wiring, caught in live use: picking an
org-policy-disabled model (claude-fable-5) produced a response claiming to
BE that model while Copilot actually served its default (Claude Sonnet 5).
Two causes:

1. The prompt preamble injected 'Hermes requested model hint: <id>', so
   whatever model actually served the session parroted the requested name
   back as its identity. Remove the line entirely — the model is applied
   for real via session/set_model now, and identity must come from the
   backend, not prompt suggestion.

2. session/new advertises policy-disabled ids alongside enabled ones
   (_meta.copilotEnablement: 'disabled'); selecting one is accepted but
   silently serves the default. Exclude disabled ids from the offered set
   so the degrade-with-warning path handles them.

Verified live: requesting claude-fable-5 logs the does-not-offer warning
listing the 23 genuinely enabled models, serves the default, and the
response truthfully self-identifies as Claude Sonnet 5.
2026-09-02 20:51:07 +05:30
unsupportedpastels 426dab7de0 fix(copilot-acp): apply the picker-selected model via session/set_model
Selecting a model on the copilot-acp provider had no effect: the model id
never left Hermes. _create_chat_completion() dropped the model argument
before _run_prompt(), so the selection survived only as prompt text
('Hermes requested model hint: ...') and Copilot answered with its own
session default — a user picking gpt-5.6-terra visibly got Claude Sonnet 5.

Live-probing 'copilot --acp --stdio' shows the CLI validates but IGNORES
its --model spawn flag in ACP mode, while session/new advertises
models.availableModels and the ACP-native session/set_model call actually
switches the session. Wire that in: forward the model into _run_prompt,
and after session/new send session/set_model when the id is advertised
(or the server reports no list). Unknown ids degrade to the session
default with a warning instead of failing the turn; the provider-level
virtual slug 'copilot-acp' is never forwarded.

Verified live against the real CLI: requesting gpt-5.6-terra answers as
GPT-5.6 Terra and claude-sonnet-5 answers as Claude Sonnet 5.
2026-09-02 20:51:07 +05:30
unsupportedpastels 5487222658 fix(models): live Copilot catalog for CLI-login users; unbreak picker-flag-empty catalogs
Three gaps between the copilot-acp picker row and what the user's
subscription actually serves (reported: picker showed the stale curated
list while the Copilot CLI offered Sonnet 5 / Opus 5 / GPT-5.6):

1. _resolve_copilot_catalog_api_key() never looked at the Copilot CLI's
   own token store (~/.copilot/config.json copilotTokens). A user whose
   only credential is 'copilot login' got no catalog key, the live fetch
   401'd, and copilot-acp silently fell back to the stale curated list.
   Add it as resolution source 3, JSONC-tolerant, with each candidate
   validated and exchanged like pool entries.

2. The existing credential-pool branch unpacked exchange_copilot_token()
   into two names, but it returns (api_token, expires_at, base_url) —
   the ValueError was swallowed by the enclosing except, disabling that
   entire resolution path. Latent since the base_url return was added.

3. GitHub now returns model_picker_enabled: false for EVERY model on
   some accounts/token types, so honoring the flag rejected the whole
   live catalog. Treat the flag as a display hint: when it empties the
   result, refilter without it (chat/endpoint checks still exclude
   embeddings and non-chat rows).

Verified live: catalog resolves 44 models for a copilot-login-only
account, matching the CLI's own picker (claude-sonnet-5, claude-opus-5,
gpt-5.6-sol/terra, gemini, kimi).
2026-09-02 20:51:07 +05:30
unsupportedpastels 6b2d32d3f6 fix(picker): keep signed-in copilot-acp visible in explicit-only desktop pickers
Two follow-up gaps found by actually running 'copilot login' end-to-end:

1. The CLI (without an OS keychain) stores its token in
   ~/.copilot/config.json under copilotTokens — a JSONC file with
   //-comment header lines. Add it as an auth-evidence source in
   _external_process_auth_evidence(), parsed comment-tolerantly and
   counting only a non-empty copilotTokens map (config.json exists after
   first launch even when logged out).

2. The desktop chat picker requests explicit_only rows, and
   _filter_explicit_provider_rows() dropped copilot-acp because a CLI
   login leaves no trace in active_provider, model.provider, or env vars
   — exactly the Anthropic-OAuth carve-out case. Keep external_process
   rows when their CLI credentials are verified (auth_verified), while
   still dropping ambient executable-on-PATH-only rows so the filter's
   narrower contract holds.

Net effect: after 'copilot login', copilot-acp appears in the desktop
picker and the Accounts card reads signed in; a machine with only the
binary installed keeps today's hidden-until-configured behavior.
2026-09-02 20:51:07 +05:30
unsupportedpastels 15f003e0b9 test(auth): cover external-process dispatch, auth evidence, sign-in command
Pin the fix class from the previous commits: auth_type-based dispatch in
get_auth_status(), positive-only auth_verified semantics (supported env
token yes, classic ghp_* PAT no, populated hosts.json yes, empty store no),
and the Accounts-tab cli_command (valid 'copilot login' default, configured
executable substitution, non-external providers untouched).
2026-09-02 20:51:07 +05:30
unsupportedpastels 323168a289 fix(dashboard): correct copilot-acp sign-in command, honest status card
The Accounts-tab card told users to run 'copilot /login', which is not a
valid invocation — slash-commands only exist inside an interactive session.
Use 'copilot login', the CLI's device-code login subcommand.

The card's status_fn also hardcoded logged_in: False with a static label.
Wire it to get_external_process_provider_status(): claim logged_in only on
positive credential evidence (auth_verified), show which executable Hermes
resolved when merely configured, and say so when the CLI is missing from
PATH entirely.

The rendered cli_command now substitutes the executable the user actually
configured (HERMES_COPILOT_ACP_COMMAND / COPILOT_CLI_PATH) so a custom
binary path gets a copy-pasteable command that matches what Hermes spawns.
2026-09-02 20:51:07 +05:30
unsupportedpastels fd439ac1b8 fix(auth): dispatch external-process providers by auth_type, add positive auth evidence
get_auth_status() special-cased the literal slug 'copilot-acp'; any other
external_process provider (the pending kiro/devin/junie ACP backends) fell
through to {'logged_in': False}. Dispatch on
PROVIDER_REGISTRY[target].auth_type == 'external_process' instead so the
whole class gets a real status.

get_external_process_provider_status() equated 'logged_in' with 'the
executable resolves', which says nothing about whether the Copilot CLI is
actually signed in. Add auth_verified/auth_source: positive-only evidence
from supported env tokens (validated via copilot_auth, classic ghp_* PATs
excluded) or known on-disk GitHub Copilot credential stores. No evidence
means unknown — never presented as signed out, because the CLI may keep its
session in an OS keychain. Deliberately subprocess-free to avoid re-creating
the gh-auth-token cold-start stall (#60800).
2026-09-02 20:51:07 +05:30
Solitud1nem 6b00567718 test(model): clear COPILOT_ACP_BASE_URL in the copilot-acp fixture
Review follow-up: the sweeper is right that the fixture's premise leaked.
It clears the tokens and both command variables, but get_auth_status()
treats an `acp+tcp://` base URL as configured on its own — no executable
required — so on a host that sets COPILOT_ACP_BASE_URL the
missing-executable test was answering a question about the host instead
of about the code.

Verified by handing the test the hostile value it was vulnerable to:
with COPILOT_ACP_BASE_URL=acp+tcp://127.0.0.1:9999 in the environment,
test_copilot_acp_hidden_when_executable_missing fails before this commit
and all three tests pass after it.
2026-09-02 20:51:07 +05:30
Solitud1nem 1e8f829f1a fix(model): show copilot-acp in pickers when its executable resolves
The overlay loop in list_authenticated_providers() checks every way a
provider might be authenticated — env keys, the auth store, the
credential pool, even Claude Code's external token files — but never
asks the one question that matters for an external_process provider:
does the executable resolve? copilot-acp has no key or token by design
(the spawned `copilot --acp --stdio` brings its own auth), so has_creds
stayed False and the filter dropped it from every picker. Funny enough,
five lines further down the same loop has a dedicated copilot-acp
branch for fetching its model ids — it just never got a chance to run.

Availability now comes from get_auth_status(), the same source
`hermes model` and the auth status endpoints already use, so the CLI
and GUI agree on what 'configured' means for external-process
providers.

Fixes #63662
2026-09-02 20:51:07 +05:30