_migrated:true files were skipped on later boots, so #100576 installs
stayed stuck on the named profile. Re-evaluate those files only; leave
user-selected pins (no _migrated) alone. If default now wins, write
{profile:null} instead of pinning default.
First-boot migrateActiveProfileIfMissing only listed ~/.hermes/profiles/*
and scored profiles/<name>/state.db. Default's real DB is ~/.hermes/state.db,
so a tiny named profile could be pinned after an update.
Always candidate default, score/pid-check it at HERMES_HOME, and do not
write active-profile.json when the winner is default.
Fixes#100576
Problem B of #71047: with streaming + reply_to_mode='first', the streamed
preview is a reply-quote of the user's message. When the turn-final edit
hits flood control, the empty-tail fresh-commit resend either (a) also got
flood-capped -> the consumer reported 'failed', the gateway's normal final
send fired, and the never-deleted preview + the fresh final left TWO
visible bubbles, or (b) succeeded but as a plain non-reply message that
didn't match the preview's anchor.
- preserve the turn's reply anchor (initial_reply_to_id) on the
empty-fallback fresh-commit resend so the replacement message quotes the
user's message exactly like the preview and the non-streaming path
- retry a flood-rejected preview deleteMessage once (delete_message
returns False rather than raising) so the stale preview doesn't linger
next to the fresh final; still best-effort, and the preview is only ever
deleted AFTER the replacement send succeeded
- regression tests for the anchor, the delete retry, and the
flood-capped-resend single-bubble suppression decision
Builds on @fangliquanflq's PR #96097 ('preview' verdict for flood-rejected
fresh commits), cherry-picked as the previous commit with the conflict
against fd998120c1 resolved (record the payload AND keep
_delivery_ambiguous only for real timeouts).
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:
- hermes_cli/auth.py Codex OAuth clients (token refresh at
auth.openai.com/oauth/token, device-code login, token exchange, usage
probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
every connect eats the full timeout per AAAA before IPv4 is tried, so
auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
had no explicit racing wired.
Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
installs the existing _HappyEyeballsSyncBackend on a ready-built sync
httpx.Client's direct transports (default transport + mounts), skipping
proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
five Codex OAuth/probe endpoints. Best-effort: falls back to default
serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
no custom backend needed. Documented in build_keepalive_http_client and
pinned by tests (contract test on the anyio signature + a live
regression test where a blackholed 100::1 IPv6 addr hangs and local
IPv4 wins in ~250ms instead of the serial connect timeout).
network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.
Refs #13834; follows #94388 (9cce8725).
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).
Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
the shared client's pooled sockets (force_close_tcp_sockets — FD release
stays with the owning worker), abort+poison of the cached per-request
openai/anthropic wire clients, Codex app-server request_interrupt(), and
the inline _active_request_abort hook. Never client.close(), never
socket.close().
- delegate timeout path: after registering the deferred-close callback,
run one immediate drain plus one 5s re-sweep (covers a connection opened
between the interrupt and the first sweep). The settled read (EOF/EPIPE)
lets the worker unwind, which triggers the deferred close on the worker's
own thread — the only safe FD-release boundary. A worker that still never
settles retains its resources rather than risking a cross-thread close.
Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.
Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.
Closes#94248
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
provider label never carries the vendor token to the alias host (#28660);
route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
Regression for #58576: _profile_scope holds _SKILLS_PROFILE_LOCK across
the payload build, which can block up to 15s on a models.dev cache miss
and starve concurrent /api/config on the same lock. The test records
which scope the handler enters for a selected profile and asserts only
the config-only (contextvar) scope is used.
_profile_scope holds _SKILLS_PROFILE_LOCK (threading.RLock) across the
entire context-manager yield. When get_model_options' worker thread
blocks on fetch_models_dev → requests.get() (up to 15s on a models.dev
cache miss), the lock stays held for the full duration. Concurrent
requests to /api/config (get_config also enters _profile_scope) then
block the main event-loop thread on the RLock, freezing the server.
Switch to _config_profile_scope which uses only the contextvar-based
HERMES_HOME override (thread-safe, no lock) — sufficient for the config
reads + credential checks that build_model_options_payload needs, and
already used by other await-safe endpoints.
Refs #58576
The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.
Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.
Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:
- fixtures.ts: title generation now rides the main model (#83636), firing a
background completion at the mock after every turn — it contains the whole
conversation (trigger keywords included), advancing scripted-turn indices
and tripping hold-for-prompt matchers. Disabled by default in the mock
provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
hidden-history-messages: busy-state and transcript assertions updated to
the current composer aria-labels, interim-message semantics, and
verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
selectors updated for the Bot Mode rework (tabs keyed by
connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
(steps() keyframes + layer promotion probed via the animation registry
instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
release handles replace wall-clock polls that lost races on loaded
runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
waits for the thread viewport before evaluating; failure path now dumps
per-surface pane state.
Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
The lane was disabled Aug 2 2026 (#76627) because the mock-backend
Electron window never got a title after the Aug 1 engines/npm churn
(#76499/#76562/#76575), failing every PR identically. #99671 fixed the
root cause: per-platform/layout Electron binary resolution in the e2e
harness (apps/desktop/e2e/electron-binary.ts). The suite is green again
on Node 26 + npm 12 — delete the temporary `false &&` guard and update
the stale comment block.
Fixes#76627
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.
Three surgical changes:
1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
The inlined watcher now routes the respawned gateway's stray
stdout/stderr to the same sidecar log gateway_windows._spawn_detached
uses (DEVNULL only as fallback), so a gateway killed moments after
respawn leaves a trace. Direct implementation of the 4th repro's
hardening suggestion (1).
2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
canonical _spawn_detached, so the respawned gateway's exit-diag /
lifecycle records show whether it escaped the parent Job Object — a
job-teardown kill is no longer indistinguishable from any other silent
death.
3. Post-update resume verifies liveness before vouching
(hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
runs the same provisional-hit + 2s-confirmation liveness poll every
other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
with all_profiles= for the fleet) before printing ✓, writes the #91675
start attestation for the verified PIDs, and fails the resume with a
"restart could not be verified" warning + recovery hint when no stable
gateway appears. Suggestion (2) of the 4th repro; closes the last
silent-success hole in the family (#84185 fixed the cold-start leg,
#91675 the direct-start leg; this is the relaunch leg).
Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.
Fixes the Bug-1 relaunch-trust leg of #48820.
The 500ms setInterval in HudShell's band-measurement effect polled
forever, contradicting its own comment ("poll briefly until it exists,
then let the ResizeObserver own it") — the viewport was never actually
checked, so the timer never cleared. It kept re-running measure()
(DOM queries + getBoundingClientRect + a style write) every 500ms for
the life of the HUD window, one of several sustained per-window timers
reported in #98394 as sustained idle renderer CPU / repeated re-renders.
Extracted the effect into useHudTranscriptBand() (matching the
existing per-concern hook split in this file: useHudGlass,
useHudClickThrough, useHudThreadFocus) and made the interval check for
the viewport before re-measuring, clearing itself once found so the
ResizeObserver takes over as the comment always said it would.
Curated picker lists (OPENROUTER_MODELS + _PROVIDER_MODELS['nous']) gain
claude-fable-5.1 above claude-fable-5 per newest-first ordering; manifest
regenerated via scripts/build_model_catalog.py.
Provider-agnostic metadata verified as already resolving for the 5.1 slug
(no new entries needed): DEFAULT_CONTEXT_LENGTHS fuzzy-matches the
claude-fable-5 prefix (1,000,000), reasoning stale-timeout floor fires
(600s), and both routes bill via official_models_api (live pricing, no
snapshot entry required).
#100540 added a REMOVED_BACKENDS startup warning keyed on tavily; with the
backend restored, that entry would warn on a working provider. The registry
stays (empty) for future removals; migration tests now pin the machinery via
a synthetic entry plus a guard asserting no live provider is ever listed as
removed.
GoalManager.set() on an event-loop thread only waits the bounded
_DB_BOOTSTRAP_INIT_WAIT_S window for the background SessionDB bootstrap
(deliberate: an unbounded init starved the gateway loop watchdog). On a
loaded CI runner the cold init overruns that window, the goal write is
silently dropped by design, and the test flakes downstream: /loop showed
no active-goal note and the goal continuation was never enqueued (both
FLAKY on main run 33455779041).
Fix the class: every async goal test fixture that clears goals._DB_CACHE
now pre-warms it via _get_session_db() from sync context (unbounded init
path), so the bounded-window degradation can never fire mid-test. Applied
to all four gateway goal/loop test files; sync-only goal tests are
unaffected by construction.
Live repro: slowing SessionDB.__init__ past the window reproduces the
dropped write deterministically without the pre-warm and never with it.
Path.rglob raises FileNotFoundError when a directory disappears between
listing and scandir — a sibling CI job creating/removing its sdist
extraction (hermes_agent-<ver>/) killed test_allowlist_has_no_stale_entries
on run 33531869442. Switch to os.walk (tolerates vanishing dirs) with
top-level pruning of exempt and packaging dirs; file set is byte-identical
(886 files verified old==new) and a 30-scan churn harness that reliably
exercised the window shows zero errors.
Same loaded-runner class: the 12-turn drain chain completed only 11
turns inside the 400x0.01s poll budget on main run 33455779041. The
loop still exits early on success, so the wider budget costs nothing
on healthy runs.
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).
Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).
The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).
Fixes#98028Fixes#100325
Draft frames set _last_sent_text for dedupe without setting
_already_sent (they are ephemeral); an ungated has_delivered_text match
let a draft-only preview count as durable delivery and regressed
test_relay_seal_failure's dead-transport guarantee on CI.
A record-less delivery flag (final_response_sent /
final_content_delivered set with no recorded turn-final payload) was
trusted blindly by delivered_final_matches (None -> legacy trust), so a
first-edit prefix or a truncated finalize suppressed the gateway's
corrective send — silent partial delivery.
- delivered_final_matches: record-less flags are now reconciled against
the FINAL content via has_delivered_text; only the explicitly-marked
ambiguous-timeout path (_delivery_ambiguous) keeps legacy trust.
- _try_fresh_final and the native-streaming optimistic finalize now
record their delivered payload (the last record-less flag setters);
the optimistic record rolls back on definitive dispatch failure.
- Discord adapter: dead-transport send failures (client gone, WS
closed/reset) are classified as send_path_degraded (retryable) so the
delivery-obligation ledger's reconnect sweep replays the stranded
final response instead of losing it until a process restart.
Fixes#95382; closes the #98552 false-positive class.
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).
- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
removed_backend_note(); selection_error() swaps in the specific
removal explanation (removed in v0.21.0, keyless alternatives) while
keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
search_backend / extract_backend against the registry and emits a
startup warning (deduped per stale value), surfaced by the existing
print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
per-capability keys, dedupe, healthy-config negative, live-backend
failure text preserved.
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.
Co-authored-by: Cursor <cursoragent@cursor.com>
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:
- _prune_replaced_custom_model_config_credentials skipped only the
preferred key, so a keyed provider's own legacy-named pool
(custom:b.ai) was false-pruned of its current model_config credential
when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
key, so a legacy-named pool stopped being seeded from model.api_key.
Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.
Follow-up to #100413.
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
consecutive data: lines per the SSE spec (a split JSON event no longer
reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.
Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.
Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.
If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.
Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.
Co-authored-by: Cursor <cursoragent@cursor.com>
On the stateless api_server platform, a background delegate_task
completion was delivered by self-POSTing /v1/chat/completions with
role=user after the parent turn had already completed (event.complete,
finish_reason=stop). That starts an unauthorized new agent turn the
client never sent, persists the completion as an ordinary role=user row
(display_kind NULL), and can blow through a pending human-confirmation
gate — the exact skip-ahead reported in #85957.
Fix: async_delegation completions targeting a non-push api_server
session are now written into the session transcript as a durable
DELIVERY row (role=user + display_kind=async_delegation_complete +
display metadata — the same bookkeeping shape the TUI/desktop delivery
path persists). No agent turn runs; clients polling
GET /api/sessions/{id}/messages see the result immediately, and the
next real client turn carries it as context. Persist failures return
False so the durable claim is released and the completion retried.
Watch-pattern notifications keep the existing self-post wake behavior;
push-capable adapters are untouched.
Fixes#85957
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.
Two layers:
1. _wait_for_gateway_ready now treats the first hit as provisional: the
gateway must stay visible through a 2s confirmation window
(_confirm_gateway_stable) before it is reported ready; a death during
confirmation resumes polling until the deadline. Failure output is an
honest ✗ with the Job Object explanation and the schtasks /Run
recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
state/gateway.start-attestation.json with the vouched-for PIDs. The
next `gateway start`/`gateway status` invocation checks it — if the
attested PIDs are gone with no clean-exit record in the lifecycle
ledger, the CLI reports (once) that the previous ✓ was false and
prints the schtasks recovery hint. `gateway stop` and a clean
lifecycle-ledger exit clear the marker silently.
Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.
Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.
Fixes#91675