Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.
Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.
Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
The Codex OAuth backend (chatgpt.com/backend-api/codex) intermittently
injects prompt_cache_retention into its own upstream call and then rejects
it, returning HTTP 400 invalid_parameter. Hermes never sends that field on
this route (see agent/transports/codex.py::_default_prompt_cache_retention_
for_request, which only sets it for api.meta.ai and bedrock-mantle hosts).
Reproduced live: a minimal 1-message request carrying no cache parameters
at all failed 4/20 (20%) with this error, so the rejection is not
deterministic and retrying the identical request is the correct recovery.
Previously the catch-all in _classify_400 returned format_error/
retryable=False, which tripped the is_client_error abort gate in
conversation_loop and killed the turn on the first attempt - burning an
entire large-context request (~550k tokens) per failure.
Classify these as retryable server_error (should_compress=False - the
request shape was never the problem). The same guard is applied to the
sibling 5xx request-validation branch, where a fronting proxy can surface
the identical rejection.
Deliberately narrow: keyed on parameters we only send on specific routes,
and skipped when the current provider is one that legitimately sends them,
so a genuine client-side bad parameter (max_tokens on GPT-5) still fails
fast as a format_error.
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.
The fix:
- _is_summary_item delegates to the canonical
agent.context_compressor.is_compaction_summary_message provenance check
(not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
through to the pruner, so it can read summary content directly from the
source when the Responses conversion shape is lossy (tool-result carrier
becomes function_call_output, or stale codex_message_items replay shadows
merged content)
Fixes#90975.
Salvage of #90976 by @JoaoMarcos44.
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:
Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
(absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free
Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
+ laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
keyless Ox Alpha; keyed — Go relay 401s anonymous requests)
Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
glm-5.2:free 256K (the free variant is capped below the 1M paid entry)
Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
MEMBERSHIP in the verified opencode-free catalog, not the -free
suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
suffix-based healing would have routed it to a Zen relay that
doesn't serve it (verified: zen 401s 'not supported', go 401s
'Missing API key'). New regression test pins this.
Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
A keyless provider has no credential to lack, but every auth-gated
surface treated 'no key' as 'not authenticated', so opencode-free was
invisible in /model, provider:model listing, and the desktop model
pickers unless a user had unrelated OpenCode env vars set.
One policy, three gates, all derived from the HermesOverlay keyless
flag (#91358):
- auth.py get_api_key_provider_status: keyless providers report
configured/logged_in=True with key_source 'keyless' — flows through
get_auth_status to every status consumer (hermes status, dashboards,
list_available_providers).
- model_switch.py list_authenticated_providers: keyless overlay rows
get has_creds=True before any env/pool/auth-store checks — this is
the source for /model, the TUI picker, and the desktop
/api/model/options payload.
- inventory.py explicit-only filter (desktop chat pickers): keyless
providers are kept — there is nothing to 'explicitly configure', and
hiding a zero-setup provider defeats its purpose.
E2E (temp HERMES_HOME, all keys stripped): get_auth_status logged_in,
list_available_providers authenticated, picker row with 6 models,
desktop payload default AND explicit_only both include the provider,
and the full switch pipeline (parse free:x-preview-f-free →
switch_model) resolves to the keyless runtime. 4 new tests.
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.
- hermes_cli/update_inventory.py (new): side-effect-free runtime
inventory — install kind via detect_install_method (git / docker / nix
/ apt, updatable-in-place or not, with the correct external update
command for image/package-managed installs), all profiles, every live
gateway with its supervisor (systemd / launchd / manual via the
fleet-wide _get_service_pids), running code_sha/code_version from the
#91283 gateway_state.json stamps, and the restart mechanism each
runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
docker/nix refusal gates so image-managed installs get a useful
'not updatable in place + right command' report instead of a bare
refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
('plan' key) and prints a one-line fleet summary, so post-mortems can
compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
never-raises, JSON round-trip for the receipt, print output shapes,
receipt integration.
tests/gateway/conftest.py already installs a comprehensive telegram mock
at collection time (line 330), before any test module's imports run.
The per-file copies were fully redundant — each was a simpler subset
(plain strings, setdefault, fewer error classes) of the conftest version
(which uses _fake_str_enum for PTB-faithful StrEnum semantics, sys.modules
overwrite to win over partial/broken imports, and a full error hierarchy
including BadRequest, Forbidden, RetryAfter, Conflict, InvalidToken).
Removed: function def + module-level call + now-unused imports (sys,
MagicMock where no longer referenced) + dangling comment blocks that
referenced the deleted mock, in 27 test files.
Left untouched: tests/gateway/conftest.py (canonical source) and
tests/e2e/conftest.py (separate conftest tree that may run in isolation).
Found by /simplify-code review of PR #90560.
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:
- _restart_macos_launchd_gateways(): the invoking profile keeps the
existing launchd_restart() path; every other gateway of this install
is drained via SIGUSR1 (same as systemd siblings), then hard-
kickstarted unless KeepAlive already respawned it, then verified on a
fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
counts toward failed_or_stale_units — including timeouts during
liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
derives labels from THIS install's profiles (get_default_hermes_root),
not by globbing the shared per-user ~/Library/LaunchAgents — a
sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
installs) must never enumerate, let alone restart, another install's
fleet. This also keeps the hermetic test suite blind to a dev
machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
liveness, kickstart, and fresh-PID verification all use the domain the
service was actually located in (gui/<uid> vs user/<uid> probed per
label via `launchctl print`). This addresses the #41403 review defect:
the process-wide _launchd_domain() cache resolves the current profile's
domain and must never be reused for a sibling. _launchd_domain() itself
becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
sweep excludes every gateway service PID (mirror of the systemd
hermes-gateway* pattern) so it cannot mistake a freshly respawned
sibling service for a stale manual gateway. Default-scope callers
(gateway status, cron checks, stop_profile_gateway's orphan reaper —
which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
hints for launchd labels alongside the systemctl ones.
Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).
Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.
- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
persistent memory
Review on #91283: begin_update_receipt() fires early in _cmd_update_impl,
but finalization only existed on the success/ZIP/CalledProcessError
paths. Early sys.exit paths (Windows concurrent-instance preflight,
venv-holder refusal, head-pinned no-op, fetch failure) terminated with
the receipt started but never written — losing exactly the refused/
failed runs the receipt matters most for.
- update_receipt.py: finalize_pending_update_receipt(exit_code,
stop_reason) — boundary safety net; maps exit 2 → 'refused', other
non-zero → 'failed'; records exit_code + stop_reason. Exactly-once by
construction (singleton popped in finalize_update_receipt), so runs
the inner paths already finalized are untouched.
- main.py cmd_update: SystemExit/BaseException/else arms around
_cmd_update_impl persist any still-open receipt with the real exit
code, then re-raise unchanged. Future early exits are covered without
per-site finalize patches.
- 5 regression tests incl. end-to-end through the real cmd_update
wrapper (exit 2 preserved, outcome 'refused', stop reason recorded,
singleton cleared, exactly one receipt file).
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
Widens the salvaged #91323 fix (@vinsew): the effort vocabulary moves to
agent.reasoning_effort (OX_ALPHA_EFFORTS/OVERRIDES, the declared-policy
home every other model vocabulary lives in), and the translation is
shared between the opencode-zen profile and the keyless opencode-free
profile — Ox Alpha is reachable through both, and the free profile
previously dropped effort entirely.
Live-verified: medium clamps to low (raw medium 400s: 'This model always
engages in thinking... use low, high, or max'), xhigh rounds to max, and
full agent turns with effort=medium complete on BOTH providers.
OpenCode documents x-preview-f-free as accepting low, high, and max reasoning effort on its Zen Chat Completions endpoint. Hermes previously resolved the user's per-model override to max but the plain Zen provider profile discarded it, so successful calls silently ran at the server default.
Introduce an OpenCodeZenProfile scoped only to x-preview-f-free. It forwards the normalized top-level reasoning_effort, maps xhigh to max, preserves server defaults when unset or disabled, and leaves every other Zen model untouched.
Add profile and full transport tests that prove max reaches the outgoing request and that non-target models are unaffected. Also correct the nanoid security-pin comment to match the already-locked 3.3.18 release.
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.
Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
opencode-free broke two provider-surface contract tests: 'api_key
providers must expose a credential env var' and 'GUI ⊇ hermes model
universe'. Both premises assume a credential exists. Add a keyless flag
to HermesOverlay + ProviderDescriptor (same derived-exemption pattern as
virtual providers) so any future anonymous provider is covered without
hardcoded slugs. Nothing to configure = no Providers-tab card, by design;
the model picker remains the selection surface.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
It exercises build_prompt_cache_plan's direct_native_tool_cache fallback,
not repeated apply on pre-decorated input, so it belongs with the other
plan-layout tests rather than in TestApplyIdempotency.
Follow-up to the #90972 salvage:
- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
is copy-on-write on content parts by contract (pops the top-level key,
rebuilds content lists/part dicts fresh), so a shallow top-level copy
preserves the caller-non-mutation guarantee — verified for all four
marker shapes — and removes the redundant second deepcopy the re-mark
path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
tests/agent/test_prompt_caching.py (where this module's tests live) as
TestApplyIdempotency; dropped the three tests that duplicated existing
coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
test_copies_sections_and_keeps_canonical_tools_plain which already
asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
no-tools fallback pins == 3 (marker loss can no longer masquerade as
safety); added the one new _can_carry_marker assertion (native=True
empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
where part-dict aliasing is the risk); fails on pre-fix base with
marker accumulation (9 > 4), passes with the fix.
apply_anthropic_cache_control never stripped pre-existing cache_control
markers before placing new ones, so calling it twice (or handing it
messages a prior call already marked) accumulated markers past
Anthropic's 4-breakpoint limit and produced HTTP 400
'cache_control can only be specified up to 4 times'.
Strip any pre-existing markers from per-message copies before marking,
mirroring the strip-then-mark pattern build_prompt_cache_plan already
uses. Only messages that already carry a marker pay the copy cost; the
copy-on-write contract (caller-owned messages are never mutated) is
preserved. Repeated calls now converge to byte-identical output.
Salvaged from #90972 by @JoaoMarcos44 (net diff of the PR's commit
stack, intermediate reverts collapsed).
Related: #90971
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
on macOS/BSD ps, making the fallback silently return [] on every macOS
machine. The matcher only needs argv (not env vars), so e is
unnecessary. -ww keeps unlimited-width output on both BSD and
procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
on macOS, enumerate every ai.hermes.gateway* launchd agent across
profiles via bare launchctl list instead of only the current
profile's label. This prevents the update sweep from misclassifying
sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
_get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
pass all_profiles=True so the update fleet sweep excludes every
service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
launchctl list <label>, all_profiles uses bare launchctl list
with prefix filtering, handles empty/broken output gracefully, and
preserves systemd behavior.
Tranquil-Flow
The Zen relay serves *-free models (x-preview-f-free / Ox Alpha) ONLY
anonymously: any Authorization bearer it doesn't recognize is a 401
'Invalid API key' — including our no-key-required placeholder and valid
OpenCode GO subscription keys. The Go relay doesn't serve the free tier
at all ('Model x is not supported'). So the free model failed for every
Hermes user: keyless setups got the placeholder bearer, and OpenCode
subscribers sent a Go key to a relay that rejects it.
Fix (class-wide for all 8 current *-free Zen slugs, not just Ox Alpha):
- hermes_cli/models.py: is_opencode_zen_free_model / opencode_zen_free_runtime
/ opencode_zen_free_headers — one shared policy: free slugs pin to the
Zen relay with a keyless placeholder and an empty Authorization header
that overrides the OpenAI SDK's 'Bearer <key>'.
- runtime_provider.py: free slugs route through the keyless runtime before
the credential-pool/explicit/api_key paths (no key required; Go
selections heal to Zen). Paid models still fail closed without a key.
- agent_init.py + auxiliary_client.py: the placeholder key swaps in the
empty-Authorization headers at both client-build chokepoints.
Verified live (2026-08-21): anonymous chat/completions 200 incl. tools,
streaming, parallel; bad bearer 401; full E2E AIAgent turn with a real
terminal tool round-trip completes keyless under both opencode-zen and
opencode-go providers. Sabotage run: routing tests fail without the fix.
* fix(gateway): persist prompt.submit truncation to the session's own profile DB
`_get_db()` returns the LAUNCH profile's SessionDB handle. App-global
remote mode gives a session its own profile (`session["profile_home"]`)
whose transcript lives in that profile's `state.db`, so a write keyed on
`session_key` that goes through `_get_db()` addresses the wrong database.
In the `prompt.submit` truncate branch that has two consequences. The
edit/resend never sticks — `session.resume` reopens the profile db and
resurrects the undone turns — and when the launch profile happens to hold
a row under the same session id, the truncated transcript is inserted
into a profile the session does not belong to.
It also silently voids the branch's own fail-closed contract. The handler
persists before it rewrites `session["history"]` precisely so that a
failed write refuses the turn and leaves memory and DB aligned; that only
holds if the handle it checks is the one that owns the row.
`_session_db(session)` is the profile-aware resolver that already exists
for this: the profile's `state.db` when `profile_home` is set, otherwise
the shared launch handle. Non-profile sessions are unaffected —
`_session_db` borrows the same shared handle and leaves it open.
`active_only=True` and `archive_dropped=True` are carried through
unchanged; only the handle the call is made against changes.
* fix(gateway): resolve the /undo command against the session's own profile DB
`command.dispatch`'s `/undo` branch opened the launch profile's handle via
`_get_db()`, but every read and write under it is scoped by session id:
`list_recent_user_messages`, `rewind_to_message` and the
`get_messages_as_conversation` reload all key on `session_key`.
For a session with its own profile (`session["profile_home"]`) the rows
live in that profile's `state.db`, so against the launch handle
`list_recent_user_messages` returns nothing and the command fails closed
with `4018 "no user messages to undo"` — for the entire session, on every
invocation, even though the transcript is right there in the profile db.
Route the whole branch through `_session_db(session)`, which yields the db
that owns the session's row and closes a profile handle on exit. Sessions
without a profile keep borrowing the shared launch handle exactly as
before, so this is behaviourally identical for them.
* fix(gateway): read /history and /context from the session's own profile DB
`_format_live_history_output` and `_format_live_context_output` rebuild the
transcript from the database rather than from `session["history"]`, because
the in-memory list is empty for a session this process did not run itself.
Both reads are scoped by session id but were issued against `_get_db()`,
the launch profile's handle.
A session with its own profile (`session["profile_home"]`) keeps its rows
in that profile's `state.db`, so both reads come back empty and the
commands under-report: `/history` renders "No conversation history yet."
and `/context` falls back to the empty in-memory list and reports a
conversation of zero messages. Both swallow their exceptions, so there is
no error either — just a wrong answer about the user's own transcript.
Resolve both through `_session_db(session)`, the profile-aware resolver
used by the rest of the session-scoped paths.
* test(gateway): cover session-scoped transcript ops against a profile DB
Regression coverage for the three session-scoped sites that resolved
against the launch profile's handle instead of the db owning the
session's row. Each test drives the real JSON-RPC entry point with a
session carrying `profile_home`, seeds the transcript into the profile's
own `state.db`, and asserts against both databases.
Per site, with the production change reverted to its pre-fix form:
- `prompt.submit` truncation — `test_truncation_persists_to_the_profile_db`
and `test_truncation_does_not_copy_rows_into_the_launch_profile` fail.
The second seeds a row under the same session id in the launch db so the
foreign write succeeds instead of failing a key check, which is the case
that copies a transcript into a profile it does not belong to.
- `/undo` — `test_undo_rewinds_the_profile_transcript` fails with
`4018 "no user messages to undo"`.
- `/history` and `/context` — `test_history_reads_the_profile_transcript`
and `test_context_reads_the_profile_transcript` fail, reporting an empty
conversation.
`test_undo_still_uses_the_shared_handle_without_a_profile` and
`test_truncation_without_a_profile_uses_the_shared_handle` pin the
unchanged path: with no `profile_home` the resolver must borrow the shared
launch handle and leave it open. Both stay green in every direction, so a
future change cannot satisfy the profile cases by abandoning the shared
one.
* fix(desktop): aim truncations by durable id alone on tail-only transcripts
The cold-open transcript is a newest-first prefetch page
(LATEST_SESSION_MESSAGES_LIMIT = 120) with the resume RPC sent
omit_messages — older rows only arrive via "Show earlier" backfill.
planEdit/planReload/planRestore still counted truncate ordinals over
that windowed list, so every edit/reload/restore in a session longer
than the prefetch page sent a window-relative ordinal alongside the
durable row/message id. The gateway's #82959 cross-check resolved the
durable id to its full-history ordinal, read the offset as drift, and
refused with 4030 — making the Edit affordance permanently dead in
long sessions.
When the transcript may be tail-only (the transcript-tail
bookkeeping's possiblyTruncated), drop the client ordinal and address
the truncation by durable id alone — the same rule runRewindSubmit
already applies to content-resolved row ids (#87059). The ordinal
tripwire stays on whenever the transcript is complete.
Closes#88082
* fix(desktop): drop client rewind ordinal whenever a durable id is present
#88092 gated the drop on tail-only prefetch. After in-place compact the
live scrollback is treated as complete, so Restore still sent a
display-lineage ordinal next to a resolved row id and the gateway
refused with 4030 (#89244). prefix_user_count is structurally 0 on
in-place because get_ancestor_display_prefix is cross-session.
Same choke point: if a durable truncate_before_row_id or a real
truncate_before_message_id is present, omit the client ordinal.
confirm_empty_truncate is still carried from a caller ordinal of 0.
Unknown ids still fail closed at 4018.
Closes#89244
---------
Co-authored-by: briandevans <252620095+briandevans@users.noreply.github.com>
Co-authored-by: zengzheqing <yuntianqing@yahoo.com>
* fix(docker): stage2 API_SERVER_KEY bootstrap no longer depends on .env existing (OOF-285)
Fleet sweep found 144/351 started hosted instances (41%) on v2026.8.13+
with no API_SERVER_KEY: the loopback gateway api_server (which serves
/api/cron/fire on :8642) never started, so every scheduled cron fire was
silently lost until the NAS retry budget exhausted.
Root cause chain:
- .dockerignore excludes .env.example (image-size optimization), so
/opt/hermes/.env.example does not exist in shipped images
- stage2's first-boot seed `seed_one ".env" ".env.example"` is a silent
no-op when the source is missing -> fresh volumes never get a .env
- the API_SERVER_KEY generation added in #84339 was gated on
`[ -f "$HERMES_HOME/.env" ]` -> never ran on those instances
Fixes:
- stage2-hook.sh: keygen now creates an owner-only .env when missing
instead of requiring it to exist; still append-only w.r.t. operator
keys, still refuses symlinked paths
- .dockerignore: re-include .env.example (negation after the .env.*
exclusion) so the first-boot template seed works again
- tests: new tests/tools/test_stage2_hook_api_server_keygen.py covers
create-when-missing, append-without-clobber, operator-key preservation,
symlink refusal, and a .dockerignore contract test for .env.example
* fix(docker): container-provided API_SERVER_KEY wins over stage2 keygen (review)
The bootstrap generated a key whenever .env lacked one, without checking
the inherited container environment. That broke the documented
`docker run -e API_SERVER_KEY=...` flow: Hermes loads $HERMES_HOME/.env
with override=True (hermes_cli/env_loader.py), so the generated key
silently shadowed the operator's env key and 401'd existing clients.
- stage2-hook.sh: skip generation when API_SERVER_KEY is present in the
container environment; if BOTH the env and .env carry keys, warn that
the .env value wins at runtime and touch nothing
- tests: regression tests for the env-provided path (skip + no .env
write; env+file conflict warns without clobbering); sandbox runner now
pins/unsets API_SERVER_KEY explicitly so results don't depend on the
host environment
* fix(docker): drop stale empty API_SERVER_KEY= line when container env provides the key
A leftover empty 'API_SERVER_KEY=' assignment in .env clobbers a
container-provided key at runtime (.env loads with override=True and
python-dotenv sets the empty string), so the api_server startup guard
fails and every scheduled cron fire is silently lost — the exact
symptom class this PR fixes, reintroduced in the env-key branch.
Remove the stale empty line (behind the existing symlink guard) before
skipping generation, so the operator's env key actually wins. Addresses
the IMPORTANT finding both reviewers converged on.
Test: env-key + stale-empty-line combination now covered; strict
removal assertion gated on GNU sed (BSD sed on macOS dev hosts skips
the -i invocation, same caveat as the append test).
* fix(docker): warn at boot when a container-provided API_SERVER_KEY is too weak to start the api_server
The startup guard refuses keys under 16 chars. Now that a
container-provided key suppresses stage2 generation, a weak
`docker run -e API_SERVER_KEY=...` value means the api_server stays
down (cron fires unavailable) instead of clients getting 401s against
a generated key. Say so in the boot log, where the operator will look.
* fix(docker): create .env under umask 077 instead of touch+chmod
touch created the file with the inherited umask (typically 0644), then
a silenced chmod tightened it to 0600 — a brief group/world-readable
window, and no warning if the chmod failed. Creating under umask 077
makes the file owner-only from the first instant with no dependence on
a second command succeeding. Covered by the existing 0600 mode
assertion in test_keygen_creates_env_when_missing.
* fix(docker): guard the API_SERVER_KEY append so a read-only .env degrades to a warning, not a failed boot
stage2 runs under set -eu; the unguarded printf append meant a keyless
.env on a read-only volume (or full disk) aborted the whole cont-init
phase and the container boot. Guard it and emit the same loud warning
the create-failure path uses.
Test harness now runs the extracted block under set -eu to match
production (it ran set -u only, so it could not see this defect class);
new read-only regression test verified RED against the unguarded
append via mutation.
* fix(docker): only warn about a weak container API_SERVER_KEY when it is actually the effective key
The <16-chars warning fired before the .env inspection, so a weak
container key alongside a strong .env key produced a false boot-log
claim that the api_server 'will refuse to start' — immediately followed
by the both-keys warning saying the .env value wins, and the server in
fact starts. Move the check into the branch where the env key really is
the effective key on this boot (round-2 review finding, verified by
execution against python-dotenv last-wins semantics).
---------
Co-authored-by: Ben Barclay <ben@nousresearch.com>
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
sol-reviewer round-4 IMPORTANT (reproduced by execution): the enroll
warning read multiplex_profiles via load_gateway_config() under the
SECONDARY profile's HERMES_HOME, but the flag normally lives in the
DEFAULT root's config.yaml — so the warning never fired in the real
topology, preserving the round-3 defect it claimed to fix.
The topology decision now mirrors the multiplexer-conflict guard in
hermes_cli/gateway.py: secondary detection is the resolved-path
relationship to <default_root>/profiles/ (not a directory-name
heuristic — also fixes the round-4 MINOR false positive on unrelated
dirs named 'profiles'), and the multiplex flag comes from the
GATEWAY_MULTIPLEX_PROFILES env override or a raw read of the default
root's config.yaml. The raw read also avoids running the full
enablement pass (round-4 MINOR: load_gateway_config() emitted the
relay-exclusive sweep's own warnings into enroll output).
The warning now replaces the generic 'restart to pick up the new env'
line instead of following it (round-4 NIT: the two messages were
contradictory), and the helper returns whether it fired.
New test file pins all six topology cases, including the exact
false-negative reproduction (flag in default root only) and env
override in both directions.
sol-reviewer round-3 findings: the regression tests stopped at the
config boundary, so config/registration agreement was only manually
verified. New TestConfigRegistrationAgreementUnderMultiplexScope
exercises both sides under an active profile scope: a process-env
stamp yields Platform.RELAY enabled AND relay_url()/
register_relay_adapter() agreement AND a constructed RelayAdapter with
a live transport; a profile-only .env stamp is inert on BOTH sides (no
half-enabled state). Also narrows the test_config docstring that
overclaimed profile .env stamps are unsupported — the launch profile's
.env still activates relay via load_hermes_dotenv's os.environ export;
only an isolated multiplex scope is never consulted.
sol-reviewer round-2 IMPORTANT: relay env vars had no scope
classification, so the two readers disagreed under a multiplexed
profile scope — gateway/config.py (scope-aware getenv) dropped a
process-env GATEWAY_RELAY_URL during the scoped runner reload while
gateway/relay's relay_url()/register_relay_adapter()/self-provision
(direct os.environ reads) still saw it. Result: adapter registered but
Platform.RELAY absent from config, so the connect loop never dialed
and direct adapters stayed up. The inverse split (profile-only stamp:
config enables RELAY, registration finds no URL) was equally dead.
GATEWAY_RELAY_* ROUTING stamps (URL, ENDPOINT, ALLOW_DIRECT_PLATFORMS,
PLATFORMS, BOT_IDS, ROUTE_KEYS, INSTANCE_ID, WAKE_URL, DISPLAY_NAME)
are now in _GLOBAL_ENV_EXACT: deployment config read from os.environ
under any scope, exactly like the API_SERVER listener settings
(#69379), so every reader resolves the same value. Relay AUTH material
(SECRET, ID, DELIVERY_KEY, IDP_*) is deliberately NOT global — it
stays profile-scoped with the fail-closed multiplex guard, mirroring
the non-secret/secret line the terminal env blocklist already draws
(tools/environments/local.py).
The round-1 multiplex regression test asserted the now-rejected
semantic (profile-scoped stamps win); it is inverted to pin the
global-stamp contract: a process-env stamp survives the profile scope
(sweep runs, matching registration), and a profile-only .env stamp
does NOT activate relay.
The named-custom-provider runtime path returned a static api_mode, so a
providers: entry like opencode-go-bridge -> https://opencode.ai/zen/go/v1
sent responses-only models (grok-4.5, gpt-5.6-luna) to /chat/completions
and got HTTP 503 (#85589 repro). Now: when the provider name is in the
OpenCode family or the base_url is hosted on opencode.ai, derive api_mode
from the effective model and run the symmetric /v1 normalization — unless
the user declared an explicit transport, which stays authoritative.
5 new regression tests against a real temp HERMES_HOME config.
Builds on @Lesnak1's #85619 (issue #85589):
- New opencode_provider_family() single-owner predicate in
hermes_cli/models.py — resolves built-in AND custom family providers
(opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
x4 from the salvaged commits) plus 4 sibling sites the PR missed:
cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
model_normalize.py flat-namespace strip, model_switch.py base_url
normalization.
- Responses transport: alias OpenCode-reserved function names
(web_search, search_files -> hermes_*) on the wire and map them back on
dispatch — same pattern as the xAI web_search collision fix. Matches
family providers and any base_url on opencode.ai. Fixes the HTTP 400
'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
Route the model-supplied target through _bound_error_text so a huge
bogus target can't bloat context, and restore the "Use 'memory' or
'user'" hint. Follow-up to HexLab98's review note on the salvage.
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
A single urlopen(timeout=5) TimeoutError from the PS runspace listener
failed the test on a loaded runner (run 32440286339) even though the
listener recovered moments later — a transient stall is not the hang
this test guards. /progress sampling now retries until a deadline
(only a persistently unresponsive listener fails), the self-test hold
grows 10s -> 30s so retry time cannot push sampling past the held
stage, and the exit wait gets matching headroom.
The CLI (hermes cron create/edit) routes through cronjob(); removing the
parameter outright broke that lane (CI slices 6/9). The parameter is back
on the function, but CRONJOB_SCHEMA and the registry handler still omit
it — same pattern as the intentional model/provider/base_url omission.
New test proves a hallucinated reasoning_effort arg through the model
dispatch is dropped.
Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.