Commit Graph

28271 Commits

Author SHA1 Message Date
Mariano Nicolini d7520b2822 fix(aux): seed the shared Nous catalog entry with the pickers' arguments 2026-09-01 16:34:48 -03:00
Sergey Tiraspolsky f6c9cb7b90 fix(desktop): re-score heuristic active-profile.json pins
_migrated:true files were skipped on later boots, so #100576 installs
stayed stuck on the named profile. Re-evaluate those files only; leave
user-selected pins (no _migrated) alone. If default now wins, write
{profile:null} instead of pinning default.
2026-09-01 15:24:33 -04:00
Sergey Tiraspolsky c90f8e8cf6 fix(desktop): score default ~/.hermes/state.db in first-boot profile migration
First-boot migrateActiveProfileIfMissing only listed ~/.hermes/profiles/*
and scored profiles/<name>/state.db. Default's real DB is ~/.hermes/state.db,
so a tiny named profile could be pinned after an update.

Always candidate default, score/pid-check it at HERMES_HOME, and do not
write active-profile.json when the winner is default.

Fixes #100576
2026-09-01 15:24:33 -04:00
Teknium fab36436b2 fix(gateway): keep one Telegram bubble when the finalize edit fails under reply_to_mode=first (#71047)
Problem B of #71047: with streaming + reply_to_mode='first', the streamed
preview is a reply-quote of the user's message. When the turn-final edit
hits flood control, the empty-tail fresh-commit resend either (a) also got
flood-capped -> the consumer reported 'failed', the gateway's normal final
send fired, and the never-deleted preview + the fresh final left TWO
visible bubbles, or (b) succeeded but as a plain non-reply message that
didn't match the preview's anchor.

- preserve the turn's reply anchor (initial_reply_to_id) on the
  empty-fallback fresh-commit resend so the replacement message quotes the
  user's message exactly like the preview and the non-streaming path
- retry a flood-rejected preview deleteMessage once (delete_message
  returns False rather than raising) so the stale preview doesn't linger
  next to the fresh final; still best-effort, and the preview is only ever
  deleted AFTER the replacement send succeeded
- regression tests for the anchor, the delete retry, and the
  flood-capped-resend single-bubble suppression decision

Builds on @fangliquanflq's PR #96097 ('preview' verdict for flood-rejected
fresh commits), cherry-picked as the previous commit with the conflict
against fd998120c1 resolved (record the payload AND keep
_delivery_ambiguous only for real timeouts).
2026-09-01 12:08:34 -07:00
fangliquanflq d7e6461b5f fix(gateway): preserve streamed final after flood rejection 2026-09-01 12:08:34 -07:00
Teknium 419232d49b fix(codex): extend Happy-Eyeballs racing to Codex OAuth/auth clients; pin async native racing
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:

- hermes_cli/auth.py Codex OAuth clients (token refresh at
  auth.openai.com/oauth/token, device-code login, token exchange, usage
  probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
  every connect eats the full timeout per AAAA before IPv4 is tried, so
  auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
  had no explicit racing wired.

Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
  installs the existing _HappyEyeballsSyncBackend on a ready-built sync
  httpx.Client's direct transports (default transport + mounts), skipping
  proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
  host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
  five Codex OAuth/probe endpoints. Best-effort: falls back to default
  serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
  RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
  no custom backend needed. Documented in build_keepalive_http_client and
  pinned by tests (contract test on the anyio signature + a live
  regression test where a blackholed 100::1 IPv6 addr hangs and local
  IPv4 wins in ~250ms instead of the serial connect timeout).

network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.

Refs #13834; follows #94388 (9cce8725).
2026-09-01 12:08:11 -07:00
Teknium e793653503 chore(contributors): map ileocorp@gmail.com -> ComicBit (PR #90889 cherry-pick) 2026-09-01 12:07:52 -07:00
Teknium 9387bf929c fix(delegate): drain abandoned-worker transports FD-safely on child timeout
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).

Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
  the shared client's pooled sockets (force_close_tcp_sockets — FD release
  stays with the owning worker), abort+poison of the cached per-request
  openai/anthropic wire clients, Codex app-server request_interrupt(), and
  the inline _active_request_abort hook. Never client.close(), never
  socket.close().
- delegate timeout path: after registering the deferred-close callback,
  run one immediate drain plus one 5s re-sweep (covers a connection opened
  between the interrupt and the first sweep). The settled read (EOF/EPIPE)
  lets the worker unwind, which triggers the deferred close on the worker's
  own thread — the only safe FD-release boundary. A worker that still never
  settles retains its resources rather than risking a cross-thread close.

Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.

Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.

Closes #94248
2026-09-01 12:07:52 -07:00
Leandro Piccione aa1d22670e fix(delegate): defer timed-out child teardown 2026-09-01 12:07:52 -07:00
Teknium 33797073bb fix: harden startup route salvage — aggregator-slug guard, alias credential ownership, oneshot dedup
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
  routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
  URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
  provider label never carries the vendor token to the alias host (#28660);
  route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
  already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
2026-09-01 12:07:35 -07:00
joaomarcos 4c870951e2 fix(cli): resolve startup model routes before provider defaults
Resolve configured aliases and provider/model inputs before HermesCLI attaches the configured default provider. Keep aggregator namespaces intact, cover oneshot startup, and document the supported CLI forms.
2026-09-01 12:07:35 -07:00
liuhao1024 f82d2f1301 fix(models): scope prefix routing to user-configured providers only 2026-09-01 12:07:35 -07:00
liuhao1024 4033f3fc5f fix(models): honor vendor/model prefix and dict model.aliases in provider detection (#87189) 2026-09-01 12:07:35 -07:00
teknium1 37f5f1ff98 test(redact): corpus-level before/after coverage for value-aware gating (#96607) 2026-09-01 12:07:16 -07:00
fangliquanflq a8ddb231aa fix(redaction): gate ambiguous assignment values 2026-09-01 12:07:16 -07:00
David Metcalfe e4c35e397c test: pin /api/model/options to _config_profile_scope for selected profiles
Regression for #58576: _profile_scope holds _SKILLS_PROFILE_LOCK across
the payload build, which can block up to 15s on a models.dev cache miss
and starve concurrent /api/config on the same lock. The test records
which scope the handler enters for a selected profile and asserts only
the config-only (contextvar) scope is used.
2026-09-01 12:07:00 -07:00
David Metcalfe 6545812c86 fix(dashboard): use config-only scope for /api/model/options to prevent lock-contention freeze
_profile_scope holds _SKILLS_PROFILE_LOCK (threading.RLock) across the
entire context-manager yield.  When get_model_options' worker thread
blocks on fetch_models_dev → requests.get() (up to 15s on a models.dev
cache miss), the lock stays held for the full duration.  Concurrent
requests to /api/config (get_config also enters _profile_scope) then
block the main event-loop thread on the RLock, freezing the server.

Switch to _config_profile_scope which uses only the contextvar-based
HERMES_HOME override (thread-safe, no lock) — sufficient for the config
reads + credential checks that build_model_options_payload needs, and
already used by other await-safe endpoints.

Refs #58576
2026-09-01 12:07:00 -07:00
Teknium c64054a26b test(desktop-e2e): run the mock-provider suite gate-free (approvals off)
The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.

Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.

Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
2026-09-01 12:04:29 -07:00
Teknium eb1b14b952 test(desktop-e2e): fix spec drift accrued while the lane was disabled
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:

- fixtures.ts: title generation now rides the main model (#83636), firing a
  background completion at the mock after every turn — it contains the whole
  conversation (trigger keywords included), advancing scripted-turn indices
  and tripping hold-for-prompt matchers. Disabled by default in the mock
  provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
  hidden-history-messages: busy-state and transcript assertions updated to
  the current composer aria-labels, interim-message semantics, and
  verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
  selectors updated for the Bot Mode rework (tabs keyed by
  connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
  (steps() keyframes + layer promotion probed via the animation registry
  instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
  release handles replace wall-clock polls that lost races on loaded
  runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
  waits for the thread viewport before evaluating; failure path now dumps
  per-surface pane state.

Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
2026-09-01 12:04:29 -07:00
teknium1 9ce95929a2 ci: re-enable the Desktop E2E lane — harness root-fixed by #99671
The lane was disabled Aug 2 2026 (#76627) because the mock-backend
Electron window never got a title after the Aug 1 engines/npm churn
(#76499/#76562/#76575), failing every PR identically. #99671 fixed the
root cause: per-platform/layout Electron binary resolution in the e2e
harness (apps/desktop/e2e/electron-binary.ts). The suite is green again
on Node 26 + npm 12 — delete the temporary `false &&` guard and update
the stale comment block.

Fixes #76627
2026-09-01 12:04:29 -07:00
Teknium ab9866bc64 fix(gateway): survive Windows Job-Object teardown across gateway restarts (#48820)
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.

Three surgical changes:

1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
   The inlined watcher now routes the respawned gateway's stray
   stdout/stderr to the same sidecar log gateway_windows._spawn_detached
   uses (DEVNULL only as fallback), so a gateway killed moments after
   respawn leaves a trace. Direct implementation of the 4th repro's
   hardening suggestion (1).

2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
   canonical _spawn_detached, so the respawned gateway's exit-diag /
   lifecycle records show whether it escaped the parent Job Object — a
   job-teardown kill is no longer indistinguishable from any other silent
   death.

3. Post-update resume verifies liveness before vouching
   (hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
   runs the same provisional-hit + 2s-confirmation liveness poll every
   other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
   with all_profiles= for the fleet) before printing ✓, writes the #91675
   start attestation for the verified PIDs, and fails the resume with a
   "restart could not be verified" warning + recovery hint when no stable
   gateway appears. Suggestion (2) of the 4th repro; closes the last
   silent-success hole in the family (#84185 fixed the cold-start leg,
   #91675 the direct-start leg; this is the relaunch leg).

Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.

Fixes the Bug-1 relaunch-trust leg of #48820.
2026-09-01 11:43:36 -07:00
chelsealong 82e6c46b94 fix(desktop): stop the HUD transcript-band probe once the viewport mounts
The 500ms setInterval in HudShell's band-measurement effect polled
forever, contradicting its own comment ("poll briefly until it exists,
then let the ResizeObserver own it") — the viewport was never actually
checked, so the timer never cleared. It kept re-running measure()
(DOM queries + getBoundingClientRect + a style write) every 500ms for
the life of the HUD window, one of several sustained per-window timers
reported in #98394 as sustained idle renderer CPU / repeated re-renders.

Extracted the effect into useHudTranscriptBand() (matching the
existing per-concern hook split in this file: useHudGlass,
useHudClickThrough, useHudThreadFocus) and made the interval check for
the viewport before re-measuring, clearing itself once found so the
ResizeObserver takes over as the comment always said it would.
2026-09-01 11:33:34 -07:00
Teknium 9f069a1175 feat(models): add anthropic/claude-fable-5.1 to OpenRouter and Nous catalogs
Curated picker lists (OPENROUTER_MODELS + _PROVIDER_MODELS['nous']) gain
claude-fable-5.1 above claude-fable-5 per newest-first ordering; manifest
regenerated via scripts/build_model_catalog.py.

Provider-agnostic metadata verified as already resolving for the 5.1 slug
(no new entries needed): DEFAULT_CONTEXT_LENGTHS fuzzy-matches the
claude-fable-5 prefix (1,000,000), reasoning stale-timeout floor fires
(600s), and both routes bill via official_models_api (live pricing, no
snapshot entry required).
2026-09-01 11:31:58 -07:00
Teknium 95d4265602 fix(install): provision supported Python from TUR on Termux and update docs 2026-09-01 11:25:10 -07:00
gustbr 8581120011 fix(install): enforce Termux Python upper bound 2026-09-01 11:25:10 -07:00
Teknium 7cd91114b4 fix(web): drop tavily from the removed-backend registry after the restore
#100540 added a REMOVED_BACKENDS startup warning keyed on tavily; with the
backend restored, that entry would warn on a working provider. The registry
stays (empty) for future removals; migration tests now pin the machinery via
a synthetic entry plus a guard asserting no live provider is ever listed as
removed.
2026-09-01 10:56:49 -07:00
Lakshya Agarwal 89ca5e614b fix(tavily): update Tavily provider documentation 2026-09-01 10:56:49 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium f8f4d056f5 test(gateway): pre-warm the goals SessionDB cache in async goal tests
GoalManager.set() on an event-loop thread only waits the bounded
_DB_BOOTSTRAP_INIT_WAIT_S window for the background SessionDB bootstrap
(deliberate: an unbounded init starved the gateway loop watchdog). On a
loaded CI runner the cold init overruns that window, the goal write is
silently dropped by design, and the test flakes downstream: /loop showed
no active-goal note and the goal continuation was never enqueued (both
FLAKY on main run 33455779041).

Fix the class: every async goal test fixture that clears goals._DB_CACHE
now pre-warms it via _get_session_db() from sync context (unbounded init
path), so the bounded-window degradation can never fire mid-test. Applied
to all four gateway goal/loop test files; sync-only goal tests are
unaffected by construction.

Live repro: slowing SessionDB.__init__ past the window reproduces the
dropped write deterministically without the pre-warm and never with it.
2026-09-01 10:52:58 -07:00
Teknium 75bf672a78 test: managed-runtime source scan survives a vanishing sdist dir (TOCTOU)
Path.rglob raises FileNotFoundError when a directory disappears between
listing and scandir — a sibling CI job creating/removing its sdist
extraction (hermes_agent-<ver>/) killed test_allowlist_has_no_stale_entries
on run 33531869442. Switch to os.walk (tolerates vanishing dirs) with
top-level pruning of exempt and packaging dirs; file set is byte-identical
(886 files verified old==new) and a 30-scan churn harness that reliably
exercised the window shows zero errors.
2026-09-01 10:52:42 -07:00
Teknium 83cde7f31d test(gateway): widen pending-drain chain wait budget (4s -> 20s)
Same loaded-runner class: the 12-turn drain chain completed only 11
turns inside the 400x0.01s poll budget on main run 33455779041. The
loop still exits early on success, so the wider budget costs nothing
on healthy runs.
2026-09-01 10:52:42 -07:00
Teknium 28834a2098 test: raise tight wall-clock bounds that flaked on loaded CI runners
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).

Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
2026-09-01 10:52:42 -07:00
Teknium 67de93862c fix(tui-gateway): gate the ws-orphan interrupt of running turns on activity staleness
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).

The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).

Fixes #98028
Fixes #100325
2026-09-01 10:52:23 -07:00
Teknium 894fc35337 fix(state): break provably-orphaned repair/FTS-rebuild locks left by dead holders (#100108) 2026-09-01 10:52:06 -07:00
Teknium 09b88bab88 fix(state): stop the on-write identity probe cancelling our own POSIX locks (#100368) 2026-09-01 10:51:52 -07:00
teknium1 46e7ad8e12 fix(gateway): gate record-less visible-text match on _already_sent
Draft frames set _last_sent_text for dedupe without setting
_already_sent (they are ephemeral); an ungated has_delivered_text match
let a draft-only preview count as durable delivery and regressed
test_relay_seal_failure's dead-transport guarantee on CI.
2026-09-01 10:51:37 -07:00
teknium1 fd998120c1 fix(gateway): judge delivery success against final content, not flag trust (#95382, #98552)
A record-less delivery flag (final_response_sent /
final_content_delivered set with no recorded turn-final payload) was
trusted blindly by delivered_final_matches (None -> legacy trust), so a
first-edit prefix or a truncated finalize suppressed the gateway's
corrective send — silent partial delivery.

- delivered_final_matches: record-less flags are now reconciled against
  the FINAL content via has_delivered_text; only the explicitly-marked
  ambiguous-timeout path (_delivery_ambiguous) keeps legacy trust.
- _try_fresh_final and the native-streaming optimistic finalize now
  record their delivered payload (the last record-less flag setters);
  the optimistic record rolls back on definitive dispatch failure.
- Discord adapter: dead-transport send failures (client gone, WS
  closed/reset) are classified as send_path_degraded (retryable) so the
  delivery-obligation ledger's reconnect sweep replays the stranded
  final response instead of losing it until a process restart.

Fixes #95382; closes the #98552 false-positive class.
2026-09-01 10:51:37 -07:00
Teknium 043c258ac2 fix(web): stale removed-backend config warns at startup and errors by name
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).

- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
  removed_backend_note(); selection_error() swaps in the specific
  removal explanation (removed in v0.21.0, keyless alternatives) while
  keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
  search_backend / extract_backend against the registry and emits a
  startup warning (deduped per stale value), surfaced by the existing
  print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
  per-capability keys, dedupe, healthy-config negative, live-backend
  failure text preserved.
2026-09-01 10:43:20 -07:00
rainbowgits 622883bad7 fix(agent): accept marker-only finish_reason after stream supersession
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:27:06 -07:00
kshitijk4poor b81383ec21 fix(auth): compare pool-identity callers against all candidate keys
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:

- _prune_replaced_custom_model_config_credentials skipped only the
  preferred key, so a keyed provider's own legacy-named pool
  (custom:b.ai) was false-pruned of its current model_config credential
  when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
  key, so a legacy-named pool stopped being seeded from model.api_key.

Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.

Follow-up to #100413.
2026-09-01 22:42:27 +05:30
xxxigm 0bee5ff408 fix(auth): look up keyed custom providers by durable pool slug
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
2026-09-01 22:42:27 +05:30
xxxigm 43470980bf test(auth): cover keyed providers.<key> credential-pool lookup
New-style providers store keys under the durable config slug, but
runtime still looks up custom:<display-name> and sends a placeholder.
2026-09-01 22:42:27 +05:30
Teknium 6ddafd34f0 test(proxy): cover multi-line SSE data joins, truthy lastOne, EOF-without-blank-line dispatch 2026-09-01 10:12:21 -07:00
Teknium 93591eccb5 fix(proxy): harden SSE DONE tracker — spec multi-line data joins, truthy lastOne, EOF-write guard
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
  consecutive data: lines per the SSE spec (a split JSON event no longer
  reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
  proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
  widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
2026-09-01 10:12:21 -07:00
Jon Nielsen d304422b3d fix(streaming): extract finish_reason/usage before content-shape continues
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.

Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.

Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
2026-09-01 10:12:21 -07:00
loulanyue 66d42e0dba fix(stream): do not misclassify stream with final usage chunk as mid-stream drop (#91373)
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.

If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.

Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
2026-09-01 10:12:21 -07:00
rainbowgits ce7f805869 fix(proxy): append SSE [DONE] when Nous streams omit the sentinel
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:12:21 -07:00
Hermes Agent 219cb714a1 fix(gateway): persist api_server async-delegation completions as deliveries, not user-turn wakes
On the stateless api_server platform, a background delegate_task
completion was delivered by self-POSTing /v1/chat/completions with
role=user after the parent turn had already completed (event.complete,
finish_reason=stop). That starts an unauthorized new agent turn the
client never sent, persists the completion as an ordinary role=user row
(display_kind NULL), and can blow through a pending human-confirmation
gate — the exact skip-ahead reported in #85957.

Fix: async_delegation completions targeting a non-push api_server
session are now written into the session transcript as a durable
DELIVERY row (role=user + display_kind=async_delegation_complete +
display metadata — the same bookkeeping shape the TUI/desktop delivery
path persists). No agent turn runs; clients polling
GET /api/sessions/{id}/messages see the result immediately, and the
next real client turn carries it as context. Persist failures return
False so the durable claim is released and the completion retried.

Watch-pattern notifications keep the existing self-post wake behavior;
push-capable adapters are untouched.

Fixes #85957
2026-09-01 10:11:07 -07:00
Teknium e9541d213f fix(gateway): never print ✓ for a Windows gateway that dies after the liveness poll
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.

Two layers:

1. _wait_for_gateway_ready now treats the first hit as provisional: the
   gateway must stay visible through a 2s confirmation window
   (_confirm_gateway_stable) before it is reported ready; a death during
   confirmation resumes polling until the deadline. Failure output is an
   honest ✗ with the Job Object explanation and the schtasks /Run
   recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
   state/gateway.start-attestation.json with the vouched-for PIDs. The
   next `gateway start`/`gateway status` invocation checks it — if the
   attested PIDs are gone with no clean-exit record in the lifecycle
   ledger, the CLI reports (once) that the previous ✓ was false and
   prints the schtasks recovery hint. `gateway stop` and a clean
   lifecycle-ledger exit clear the marker silently.

Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.

Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.

Fixes #91675
2026-09-01 10:06:06 -07:00
hermes-seaeye[bot] e5e71d8c46 fmt(js): npm run fix on merge (#100517)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 16:59:53 +00:00