Commit Graph

14240 Commits

Author SHA1 Message Date
abundantbeing fdf0111410 fix(agent): guard merged assistant compaction handoffs
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.

Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.

Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
2026-08-21 21:52:55 +05:30
Troy Rowe 6e53628338 fix(classifier): retry provider-injected parameter 400s instead of aborting
The Codex OAuth backend (chatgpt.com/backend-api/codex) intermittently
injects prompt_cache_retention into its own upstream call and then rejects
it, returning HTTP 400 invalid_parameter. Hermes never sends that field on
this route (see agent/transports/codex.py::_default_prompt_cache_retention_
for_request, which only sets it for api.meta.ai and bedrock-mantle hosts).

Reproduced live: a minimal 1-message request carrying no cache parameters
at all failed 4/20 (20%) with this error, so the rejection is not
deterministic and retrying the identical request is the correct recovery.

Previously the catch-all in _classify_400 returned format_error/
retryable=False, which tripped the is_client_error abort gate in
conversation_loop and killed the turn on the first attempt - burning an
entire large-context request (~550k tokens) per failure.

Classify these as retryable server_error (should_compress=False - the
request shape was never the problem). The same guard is applied to the
sibling 5xx request-validation branch, where a fronting proxy can surface
the identical rejection.

Deliberately narrow: keyed on parameters we only send on specific routes,
and skipped when the current provider is one that legitimately sends them,
so a genuine client-side bad parameter (max_tokens on GPT-5) still fails
fast as a format_error.
2026-08-21 21:43:39 +05:30
poisdahl 13fcf2fe38 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-21 16:02:59 +02:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
JoaoMarcos44 fb27614add fix(native_compaction): preserve compression summary messages during pre-checkpoint pruning
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.

The fix:
- _is_summary_item delegates to the canonical
  agent.context_compressor.is_compaction_summary_message provenance check
  (not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
  through to the pruner, so it can read summary content directly from the
  source when the Responses conversion shape is lossy (tool-result carrier
  becomes function_call_output, or stale codex_message_items replay shadows
  merged content)

Fixes #90975.

Salvage of #90976 by @JoaoMarcos44.
2026-08-21 17:24:22 +05:30
Teknium 624723130b fix: catalog drift sync — dead free slugs out, live OpenRouter free models in, ox-alpha-free on Go
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:

Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
  (absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
  s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free

Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
  + laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
  keyless Ox Alpha; keyed — Go relay 401s anonymous requests)

Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
  nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
  glm-5.2:free 256K (the free variant is capped below the 1M paid entry)

Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
  MEMBERSHIP in the verified opencode-free catalog, not the -free
  suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
  suffix-based healing would have routed it to a Zen relay that
  doesn't serve it (verified: zen 401s 'not supported', go 401s
  'Missing API key'). New regression test pins this.

Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
2026-08-21 04:44:15 -07:00
Teknium 2a2307e68f feat: keyless providers count as authenticated everywhere — opencode-free appears in /model and desktop pickers with zero setup
A keyless provider has no credential to lack, but every auth-gated
surface treated 'no key' as 'not authenticated', so opencode-free was
invisible in /model, provider:model listing, and the desktop model
pickers unless a user had unrelated OpenCode env vars set.

One policy, three gates, all derived from the HermesOverlay keyless
flag (#91358):
- auth.py get_api_key_provider_status: keyless providers report
  configured/logged_in=True with key_source 'keyless' — flows through
  get_auth_status to every status consumer (hermes status, dashboards,
  list_available_providers).
- model_switch.py list_authenticated_providers: keyless overlay rows
  get has_creds=True before any env/pool/auth-store checks — this is
  the source for /model, the TUI picker, and the desktop
  /api/model/options payload.
- inventory.py explicit-only filter (desktop chat pickers): keyless
  providers are kept — there is nothing to 'explicitly configure', and
  hiding a zero-setup provider defeats its purpose.

E2E (temp HERMES_HOME, all keys stripped): get_auth_status logged_in,
list_available_providers authenticated, picker row with 6 models,
desktop payload default AND explicit_only both include the provider,
and the full switch pipeline (parse free:x-preview-f-free →
switch_model) resolves to the keyless runtime. 4 new tests.
2026-08-21 04:39:31 -07:00
Teknium 0aecadc17c feat(update): hermes update --plan — read-only fleet inventory + plan phase in every update
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.

- hermes_cli/update_inventory.py (new): side-effect-free runtime
  inventory — install kind via detect_install_method (git / docker / nix
  / apt, updatable-in-place or not, with the correct external update
  command for image/package-managed installs), all profiles, every live
  gateway with its supervisor (systemd / launchd / manual via the
  fleet-wide _get_service_pids), running code_sha/code_version from the
  #91283 gateway_state.json stamps, and the restart mechanism each
  runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
  docker/nix refusal gates so image-managed installs get a useful
  'not updatable in place + right command' report instead of a bare
  refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
  ('plan' key) and prints a one-line fleet summary, so post-mortems can
  compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
  cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
  dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
  never-raises, JSON round-trip for the receipt, print output shapes,
  receipt integration.
2026-08-21 04:23:13 -07:00
kshitij 17d1095443 Merge pull request #91401 from kshitijk4poor/refactor/dedupe-telegram-mock
refactor: remove 19 duplicate _ensure_telegram_mock() copies from gateway tests
2026-08-21 16:34:54 +05:30
kshitijk4poor c1693d7dcc refactor: remove 27 duplicate _ensure_telegram_mock() copies from gateway tests
tests/gateway/conftest.py already installs a comprehensive telegram mock
at collection time (line 330), before any test module's imports run.
The per-file copies were fully redundant — each was a simpler subset
(plain strings, setdefault, fewer error classes) of the conftest version
(which uses _fake_str_enum for PTB-faithful StrEnum semantics, sys.modules
overwrite to win over partial/broken imports, and a full error hierarchy
including BadRequest, Forbidden, RetryAfter, Conflict, InvalidToken).

Removed: function def + module-level call + now-unused imports (sys,
MagicMock where no longer referenced) + dangling comment blocks that
referenced the deleted mock, in 27 test files.
Left untouched: tests/gateway/conftest.py (canonical source) and
tests/e2e/conftest.py (separate conftest tree that may run in isolation).

Found by /simplify-code review of PR #90560.
2026-08-21 16:25:24 +05:30
Teknium ff6186dc60 test(gateway): re-pin _get_service_pids tests to the label-derived locate + prefix-scan union 2026-08-21 03:51:03 -07:00
PT f29ee96dd3 fix(update): restart all macOS launchd gateways on hermes update
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:

- _restart_macos_launchd_gateways(): the invoking profile keeps the
  existing launchd_restart() path; every other gateway of this install
  is drained via SIGUSR1 (same as systemd siblings), then hard-
  kickstarted unless KeepAlive already respawned it, then verified on a
  fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
  counts toward failed_or_stale_units — including timeouts during
  liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
  derives labels from THIS install's profiles (get_default_hermes_root),
  not by globbing the shared per-user ~/Library/LaunchAgents — a
  sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
  installs) must never enumerate, let alone restart, another install's
  fleet. This also keeps the hermetic test suite blind to a dev
  machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
  liveness, kickstart, and fresh-PID verification all use the domain the
  service was actually located in (gui/<uid> vs user/<uid> probed per
  label via `launchctl print`). This addresses the #41403 review defect:
  the process-wide _launchd_domain() cache resolves the current profile's
  domain and must never be reused for a sibling. _launchd_domain() itself
  becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
  sweep excludes every gateway service PID (mirror of the systemd
  hermes-gateway* pattern) so it cannot mistake a freshly respawned
  sibling service for a stale manual gateway. Default-scope callers
  (gateway status, cron checks, stop_profile_gateway's orphan reaper —
  which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
  hints for launchd labels alongside the systemctl ones.

Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).

Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 03:51:03 -07:00
Teknium ef04d846e9 feat(cron): cron agents now run with memory enabled like every other agent
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.

- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
  from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
  and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
  memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
  persistent memory
2026-08-21 03:46:37 -07:00
Teknium 40b4a3bfe1 fix(update): every begun update receipt is now persisted — command-boundary finalization
Review on #91283: begin_update_receipt() fires early in _cmd_update_impl,
but finalization only existed on the success/ZIP/CalledProcessError
paths. Early sys.exit paths (Windows concurrent-instance preflight,
venv-holder refusal, head-pinned no-op, fetch failure) terminated with
the receipt started but never written — losing exactly the refused/
failed runs the receipt matters most for.

- update_receipt.py: finalize_pending_update_receipt(exit_code,
  stop_reason) — boundary safety net; maps exit 2 → 'refused', other
  non-zero → 'failed'; records exit_code + stop_reason. Exactly-once by
  construction (singleton popped in finalize_update_receipt), so runs
  the inner paths already finalized are untouched.
- main.py cmd_update: SystemExit/BaseException/else arms around
  _cmd_update_impl persist any still-open receipt with the real exit
  code, then re-raise unchanged. Future early exits are covered without
  per-site finalize patches.
- 5 regression tests incl. end-to-end through the real cmd_update
  wrapper (exit 2 preserved, outcome 'refused', stop reason recorded,
  singleton cleared, exactly one receipt file).
2026-08-21 03:44:55 -07:00
kshitijk4poor b883756b79 fix: foreground priority for background review cancel timeout
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
2026-08-21 16:12:57 +05:30
qixuancao 1b92a94962 refactor(agent): simplify background review run state 2026-08-21 16:12:57 +05:30
qixuancao 37da0d4d50 fix(agent): synchronize background review cancellation 2026-08-21 16:12:57 +05:30
Teknium d4d04098a5 fix: Ox Alpha reasoning effort reaches the wire clamped — shared across zen and free providers
Widens the salvaged #91323 fix (@vinsew): the effort vocabulary moves to
agent.reasoning_effort (OX_ALPHA_EFFORTS/OVERRIDES, the declared-policy
home every other model vocabulary lives in), and the translation is
shared between the opencode-zen profile and the keyless opencode-free
profile — Ox Alpha is reachable through both, and the free profile
previously dropped effort entirely.

Live-verified: medium clamps to low (raw medium 400s: 'This model always
engages in thinking... use low, high, or max'), xhigh rounds to max, and
full agent turns with effort=medium complete on BOTH providers.
2026-08-21 03:04:06 -07:00
vinsew 54227416ce fix(opencode): send Ox Alpha reasoning effort through Zen
OpenCode documents x-preview-f-free as accepting low, high, and max reasoning effort on its Zen Chat Completions endpoint. Hermes previously resolved the user's per-model override to max but the plain Zen provider profile discarded it, so successful calls silently ran at the server default.

Introduce an OpenCodeZenProfile scoped only to x-preview-f-free. It forwards the normalized top-level reasoning_effort, maps xhigh to max, preserves server defaults when unset or disabled, and leaves every other Zen model untouched.

Add profile and full transport tests that prove max reaches the outgoing request and that non-target models are unaffected. Also correct the nanoid security-pin comment to match the already-locked 3.3.18 release.
2026-08-21 03:04:06 -07:00
Adolanium fc9cbc872d fix(cron): do not load MEMORY.md into scheduled jobs
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.

Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
2026-08-21 13:24:43 +05:30
Teknium 930be347f1 feat: keyless flag on the provider catalog — GUI contract tests exempt anonymous providers
opencode-free broke two provider-surface contract tests: 'api_key
providers must expose a credential env var' and 'GUI ⊇ hermes model
universe'. Both premises assume a credential exists. Add a keyless flag
to HermesOverlay + ProviderDescriptor (same derived-exemption pattern as
virtual providers) so any future anonymous provider is covered without
hardcoded slugs. Nothing to configure = no Providers-tab card, by design;
the model picker remains the selection surface.
2026-08-21 00:24:32 -07:00
Teknium ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal 28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
kshitijk4poor 8e77d03188 test(prompt_caching): home empty-tools fallback test in TestPromptCachePlan
It exercises build_prompt_cache_plan's direct_native_tool_cache fallback,
not repeated apply on pre-decorated input, so it belongs with the other
plan-layout tests rather than in TestApplyIdempotency.
2026-08-21 12:23:47 +05:30
kshitijk4poor c26357ad6a refactor(prompt_caching): shallow strip copy, exact-count guards, dedupe idempotency tests
Follow-up to the #90972 salvage:

- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
  is copy-on-write on content parts by contract (pops the top-level key,
  rebuilds content lists/part dicts fresh), so a shallow top-level copy
  preserves the caller-non-mutation guarantee — verified for all four
  marker shapes — and removes the redundant second deepcopy the re-mark
  path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
  tests/agent/test_prompt_caching.py (where this module's tests live) as
  TestApplyIdempotency; dropped the three tests that duplicated existing
  coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
  test_copies_sections_and_keeps_canonical_tools_plain which already
  asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
  never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
  no-tools fallback pins == 3 (marker loss can no longer masquerade as
  safety); added the one new _can_carry_marker assertion (native=True
  empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
  where part-dict aliasing is the risk); fails on pre-fix base with
  marker accumulation (9 > 4), passes with the fix.
2026-08-21 12:23:47 +05:30
joaomarcos 0fc52b055f fix(prompt_caching): make apply_anthropic_cache_control idempotent on pre-decorated input
apply_anthropic_cache_control never stripped pre-existing cache_control
markers before placing new ones, so calling it twice (or handing it
messages a prior call already marked) accumulated markers past
Anthropic's 4-breakpoint limit and produced HTTP 400
'cache_control can only be specified up to 4 times'.

Strip any pre-existing markers from per-message copies before marking,
mirroring the strip-then-mark pattern build_prompt_cache_plan already
uses. Only messages that already carry a marker pay the copy cost; the
copy-on-write contract (caller-owned messages are never mutated) is
preserved. Repeated calls now converge to byte-identical output.

Salvaged from #90972 by @JoaoMarcos44 (net diff of the PR's commit
stack, intermediate reverts collapsed).

Related: #90971
2026-08-21 12:23:47 +05:30
xxxigm a86569bd11 test(teams-pipeline): cover quoted-user Graph @odata.id paths
Lock in users('{id}')/onlineMeetings('{id}') parsing, job creation from that
notification shape, and replay that re-reads a stored transcript id.
2026-08-21 12:15:36 +05:30
Teknium 15328b4db7 test(gateway): accept all_profiles kwarg in macOS reaper _get_service_pids mocks 2026-08-20 23:41:00 -07:00
Tranquil-Flow d8047c303b fix(gateway): BSD-compatible ps flags and all-profile launchd pid discovery (#74075)
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
  on macOS/BSD ps, making the fallback silently return [] on every macOS
  machine. The matcher only needs argv (not env vars), so e is
  unnecessary. -ww keeps unlimited-width output on both BSD and
  procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
  on macOS, enumerate every ai.hermes.gateway* launchd agent across
  profiles via bare launchctl list instead of only the current
  profile's label. This prevents the update sweep from misclassifying
  sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
  _get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
  pass all_profiles=True so the update fleet sweep excludes every
  service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
  not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
  launchctl list <label>, all_profiles uses bare launchctl list
  with prefix filtering, handles empty/broken output gracefully, and
  preserves systemd behavior.

Tranquil-Flow
2026-08-20 23:41:00 -07:00
Teknium 1017a56274 fix: OpenCode Zen free-tier models (Ox Alpha) work keyless — any bearer 401s them
The Zen relay serves *-free models (x-preview-f-free / Ox Alpha) ONLY
anonymously: any Authorization bearer it doesn't recognize is a 401
'Invalid API key' — including our no-key-required placeholder and valid
OpenCode GO subscription keys. The Go relay doesn't serve the free tier
at all ('Model x is not supported'). So the free model failed for every
Hermes user: keyless setups got the placeholder bearer, and OpenCode
subscribers sent a Go key to a relay that rejects it.

Fix (class-wide for all 8 current *-free Zen slugs, not just Ox Alpha):
- hermes_cli/models.py: is_opencode_zen_free_model / opencode_zen_free_runtime
  / opencode_zen_free_headers — one shared policy: free slugs pin to the
  Zen relay with a keyless placeholder and an empty Authorization header
  that overrides the OpenAI SDK's 'Bearer <key>'.
- runtime_provider.py: free slugs route through the keyless runtime before
  the credential-pool/explicit/api_key paths (no key required; Go
  selections heal to Zen). Paid models still fail closed without a key.
- agent_init.py + auxiliary_client.py: the placeholder key swaps in the
  empty-Authorization headers at both client-build chokepoints.

Verified live (2026-08-21): anonymous chat/completions 200 incl. tools,
streaming, parallel; bad bearer 401; full E2E AIAgent turn with a real
terminal tool round-trip completes keyless under both opencode-zen and
opencode-go providers. Sabotage run: routing tests fail without the fix.
2026-08-20 23:05:57 -07:00
brooklyn! 02e270a47e fix: editing a message in an old session fails (profile DB + window-relative ordinal) (#91302)
* fix(gateway): persist prompt.submit truncation to the session's own profile DB

`_get_db()` returns the LAUNCH profile's SessionDB handle. App-global
remote mode gives a session its own profile (`session["profile_home"]`)
whose transcript lives in that profile's `state.db`, so a write keyed on
`session_key` that goes through `_get_db()` addresses the wrong database.

In the `prompt.submit` truncate branch that has two consequences. The
edit/resend never sticks — `session.resume` reopens the profile db and
resurrects the undone turns — and when the launch profile happens to hold
a row under the same session id, the truncated transcript is inserted
into a profile the session does not belong to.

It also silently voids the branch's own fail-closed contract. The handler
persists before it rewrites `session["history"]` precisely so that a
failed write refuses the turn and leaves memory and DB aligned; that only
holds if the handle it checks is the one that owns the row.

`_session_db(session)` is the profile-aware resolver that already exists
for this: the profile's `state.db` when `profile_home` is set, otherwise
the shared launch handle. Non-profile sessions are unaffected —
`_session_db` borrows the same shared handle and leaves it open.

`active_only=True` and `archive_dropped=True` are carried through
unchanged; only the handle the call is made against changes.

* fix(gateway): resolve the /undo command against the session's own profile DB

`command.dispatch`'s `/undo` branch opened the launch profile's handle via
`_get_db()`, but every read and write under it is scoped by session id:
`list_recent_user_messages`, `rewind_to_message` and the
`get_messages_as_conversation` reload all key on `session_key`.

For a session with its own profile (`session["profile_home"]`) the rows
live in that profile's `state.db`, so against the launch handle
`list_recent_user_messages` returns nothing and the command fails closed
with `4018 "no user messages to undo"` — for the entire session, on every
invocation, even though the transcript is right there in the profile db.

Route the whole branch through `_session_db(session)`, which yields the db
that owns the session's row and closes a profile handle on exit. Sessions
without a profile keep borrowing the shared launch handle exactly as
before, so this is behaviourally identical for them.

* fix(gateway): read /history and /context from the session's own profile DB

`_format_live_history_output` and `_format_live_context_output` rebuild the
transcript from the database rather than from `session["history"]`, because
the in-memory list is empty for a session this process did not run itself.
Both reads are scoped by session id but were issued against `_get_db()`,
the launch profile's handle.

A session with its own profile (`session["profile_home"]`) keeps its rows
in that profile's `state.db`, so both reads come back empty and the
commands under-report: `/history` renders "No conversation history yet."
and `/context` falls back to the empty in-memory list and reports a
conversation of zero messages. Both swallow their exceptions, so there is
no error either — just a wrong answer about the user's own transcript.

Resolve both through `_session_db(session)`, the profile-aware resolver
used by the rest of the session-scoped paths.

* test(gateway): cover session-scoped transcript ops against a profile DB

Regression coverage for the three session-scoped sites that resolved
against the launch profile's handle instead of the db owning the
session's row. Each test drives the real JSON-RPC entry point with a
session carrying `profile_home`, seeds the transcript into the profile's
own `state.db`, and asserts against both databases.

Per site, with the production change reverted to its pre-fix form:

- `prompt.submit` truncation — `test_truncation_persists_to_the_profile_db`
  and `test_truncation_does_not_copy_rows_into_the_launch_profile` fail.
  The second seeds a row under the same session id in the launch db so the
  foreign write succeeds instead of failing a key check, which is the case
  that copies a transcript into a profile it does not belong to.
- `/undo` — `test_undo_rewinds_the_profile_transcript` fails with
  `4018 "no user messages to undo"`.
- `/history` and `/context` — `test_history_reads_the_profile_transcript`
  and `test_context_reads_the_profile_transcript` fail, reporting an empty
  conversation.

`test_undo_still_uses_the_shared_handle_without_a_profile` and
`test_truncation_without_a_profile_uses_the_shared_handle` pin the
unchanged path: with no `profile_home` the resolver must borrow the shared
launch handle and leave it open. Both stay green in every direction, so a
future change cannot satisfy the profile cases by abandoning the shared
one.

* fix(desktop): aim truncations by durable id alone on tail-only transcripts

The cold-open transcript is a newest-first prefetch page
(LATEST_SESSION_MESSAGES_LIMIT = 120) with the resume RPC sent
omit_messages — older rows only arrive via "Show earlier" backfill.
planEdit/planReload/planRestore still counted truncate ordinals over
that windowed list, so every edit/reload/restore in a session longer
than the prefetch page sent a window-relative ordinal alongside the
durable row/message id. The gateway's #82959 cross-check resolved the
durable id to its full-history ordinal, read the offset as drift, and
refused with 4030 — making the Edit affordance permanently dead in
long sessions.

When the transcript may be tail-only (the transcript-tail
bookkeeping's possiblyTruncated), drop the client ordinal and address
the truncation by durable id alone — the same rule runRewindSubmit
already applies to content-resolved row ids (#87059). The ordinal
tripwire stays on whenever the transcript is complete.

Closes #88082

* fix(desktop): drop client rewind ordinal whenever a durable id is present

#88092 gated the drop on tail-only prefetch. After in-place compact the
live scrollback is treated as complete, so Restore still sent a
display-lineage ordinal next to a resolved row id and the gateway
refused with 4030 (#89244). prefix_user_count is structurally 0 on
in-place because get_ancestor_display_prefix is cross-session.

Same choke point: if a durable truncate_before_row_id or a real
truncate_before_message_id is present, omit the client ordinal.
confirm_empty_truncate is still carried from a caller ordinal of 0.
Unknown ids still fail closed at 4018.

Closes #89244

---------

Co-authored-by: briandevans <252620095+briandevans@users.noreply.github.com>
Co-authored-by: zengzheqing <yuntianqing@yahoo.com>
2026-08-21 06:02:19 +00:00
Gille 790c850144 fix(telegram): preserve rich finals after DM drafts 2026-08-20 21:58:18 -07:00
shannonsands 7a17a1b8a6 fix(docker): stage2 API_SERVER_KEY bootstrap no longer depends on .env existing (OOF-285) (#88926)
* fix(docker): stage2 API_SERVER_KEY bootstrap no longer depends on .env existing (OOF-285)

Fleet sweep found 144/351 started hosted instances (41%) on v2026.8.13+
with no API_SERVER_KEY: the loopback gateway api_server (which serves
/api/cron/fire on :8642) never started, so every scheduled cron fire was
silently lost until the NAS retry budget exhausted.

Root cause chain:
- .dockerignore excludes .env.example (image-size optimization), so
  /opt/hermes/.env.example does not exist in shipped images
- stage2's first-boot seed `seed_one ".env" ".env.example"` is a silent
  no-op when the source is missing -> fresh volumes never get a .env
- the API_SERVER_KEY generation added in #84339 was gated on
  `[ -f "$HERMES_HOME/.env" ]` -> never ran on those instances

Fixes:
- stage2-hook.sh: keygen now creates an owner-only .env when missing
  instead of requiring it to exist; still append-only w.r.t. operator
  keys, still refuses symlinked paths
- .dockerignore: re-include .env.example (negation after the .env.*
  exclusion) so the first-boot template seed works again
- tests: new tests/tools/test_stage2_hook_api_server_keygen.py covers
  create-when-missing, append-without-clobber, operator-key preservation,
  symlink refusal, and a .dockerignore contract test for .env.example

* fix(docker): container-provided API_SERVER_KEY wins over stage2 keygen (review)

The bootstrap generated a key whenever .env lacked one, without checking
the inherited container environment. That broke the documented
`docker run -e API_SERVER_KEY=...` flow: Hermes loads $HERMES_HOME/.env
with override=True (hermes_cli/env_loader.py), so the generated key
silently shadowed the operator's env key and 401'd existing clients.

- stage2-hook.sh: skip generation when API_SERVER_KEY is present in the
  container environment; if BOTH the env and .env carry keys, warn that
  the .env value wins at runtime and touch nothing
- tests: regression tests for the env-provided path (skip + no .env
  write; env+file conflict warns without clobbering); sandbox runner now
  pins/unsets API_SERVER_KEY explicitly so results don't depend on the
  host environment

* fix(docker): drop stale empty API_SERVER_KEY= line when container env provides the key

A leftover empty 'API_SERVER_KEY=' assignment in .env clobbers a
container-provided key at runtime (.env loads with override=True and
python-dotenv sets the empty string), so the api_server startup guard
fails and every scheduled cron fire is silently lost — the exact
symptom class this PR fixes, reintroduced in the env-key branch.

Remove the stale empty line (behind the existing symlink guard) before
skipping generation, so the operator's env key actually wins. Addresses
the IMPORTANT finding both reviewers converged on.

Test: env-key + stale-empty-line combination now covered; strict
removal assertion gated on GNU sed (BSD sed on macOS dev hosts skips
the -i invocation, same caveat as the append test).

* fix(docker): warn at boot when a container-provided API_SERVER_KEY is too weak to start the api_server

The startup guard refuses keys under 16 chars. Now that a
container-provided key suppresses stage2 generation, a weak
`docker run -e API_SERVER_KEY=...` value means the api_server stays
down (cron fires unavailable) instead of clients getting 401s against
a generated key. Say so in the boot log, where the operator will look.

* fix(docker): create .env under umask 077 instead of touch+chmod

touch created the file with the inherited umask (typically 0644), then
a silenced chmod tightened it to 0600 — a brief group/world-readable
window, and no warning if the chmod failed. Creating under umask 077
makes the file owner-only from the first instant with no dependence on
a second command succeeding. Covered by the existing 0600 mode
assertion in test_keygen_creates_env_when_missing.

* fix(docker): guard the API_SERVER_KEY append so a read-only .env degrades to a warning, not a failed boot

stage2 runs under set -eu; the unguarded printf append meant a keyless
.env on a read-only volume (or full disk) aborted the whole cont-init
phase and the container boot. Guard it and emit the same loud warning
the create-failure path uses.

Test harness now runs the extracted block under set -eu to match
production (it ran set -u only, so it could not see this defect class);
new read-only regression test verified RED against the unguarded
append via mutation.

* fix(docker): only warn about a weak container API_SERVER_KEY when it is actually the effective key

The <16-chars warning fired before the .env inspection, so a weak
container key alongside a strong .env key produced a false boot-log
claim that the api_server 'will refuse to start' — immediately followed
by the both-keys warning saying the .env value wins, and the server in
fact starts. Move the check into the branch where the env key really is
the effective key on this boot (round-2 review finding, verified by
execution against python-dotenv last-wins semantics).

---------

Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-08-21 14:56:28 +10:00
Teknium 1d74833d8d feat(update): structured update receipts + post-update fleet version verification
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.

- hermes_cli/build_info.py: get_code_identity() — process-cached code
  identity (git sha for source installs, baked .hermes_build_sha for
  Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
  code_sha/code_version into gateway_state.json, so a running gateway's
  actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
  update run (steps, skips with reasons, gateway restart outcome, fleet
  snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
  pointer for the dashboard/desktop; plus collect_fleet_versions() /
  print_fleet_version_matrix() comparing every live profile gateway
  against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
  git, ZIP, and hard-failure paths; after the restart phase, prints the
  fleet version matrix and escalates provably-stale gateways into the
  existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
  gateways report 'unknown' and never fail the update (no false
  positives during rollout).

Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
2026-08-20 21:54:25 -07:00
TheTom 3b83b5e1fa fix(gateway): serialize shared bot room updates 2026-08-20 21:54:18 -07:00
Ben Barclay efb6b40f94 Merge pull request #91237 from NousResearch/fix/relay-env-exclusive-messaging
fix(gateway): GATEWAY_RELAY_URL env stamp disables direct messaging platforms
2026-08-21 14:33:29 +10:00
Ben Barclay 6360113a8b fix(cli): read multiplex topology from the default root in enroll warning
sol-reviewer round-4 IMPORTANT (reproduced by execution): the enroll
warning read multiplex_profiles via load_gateway_config() under the
SECONDARY profile's HERMES_HOME, but the flag normally lives in the
DEFAULT root's config.yaml — so the warning never fired in the real
topology, preserving the round-3 defect it claimed to fix.

The topology decision now mirrors the multiplexer-conflict guard in
hermes_cli/gateway.py: secondary detection is the resolved-path
relationship to <default_root>/profiles/ (not a directory-name
heuristic — also fixes the round-4 MINOR false positive on unrelated
dirs named 'profiles'), and the multiplex flag comes from the
GATEWAY_MULTIPLEX_PROFILES env override or a raw read of the default
root's config.yaml. The raw read also avoids running the full
enablement pass (round-4 MINOR: load_gateway_config() emitted the
relay-exclusive sweep's own warnings into enroll output).

The warning now replaces the generic 'restart to pick up the new env'
line instead of following it (round-4 NIT: the two messages were
contradictory), and the helper returns whether it fired.

New test file pins all six topology cases, including the exact
false-negative reproduction (flag in default root only) and env
override in both directions.
2026-08-21 14:13:28 +10:00
Ben Barclay d97cb8f253 test(gateway): pin config/registration relay-URL agreement under multiplex scope
sol-reviewer round-3 findings: the regression tests stopped at the
config boundary, so config/registration agreement was only manually
verified. New TestConfigRegistrationAgreementUnderMultiplexScope
exercises both sides under an active profile scope: a process-env
stamp yields Platform.RELAY enabled AND relay_url()/
register_relay_adapter() agreement AND a constructed RelayAdapter with
a live transport; a profile-only .env stamp is inert on BOTH sides (no
half-enabled state). Also narrows the test_config docstring that
overclaimed profile .env stamps are unsupported — the launch profile's
.env still activates relay via load_hermes_dotenv's os.environ export;
only an isolated multiplex scope is never consulted.
2026-08-21 13:35:20 +10:00
Ben Barclay f8c3635d40 fix(gateway): classify relay routing stamps as process-global deployment config
sol-reviewer round-2 IMPORTANT: relay env vars had no scope
classification, so the two readers disagreed under a multiplexed
profile scope — gateway/config.py (scope-aware getenv) dropped a
process-env GATEWAY_RELAY_URL during the scoped runner reload while
gateway/relay's relay_url()/register_relay_adapter()/self-provision
(direct os.environ reads) still saw it. Result: adapter registered but
Platform.RELAY absent from config, so the connect loop never dialed
and direct adapters stayed up. The inverse split (profile-only stamp:
config enables RELAY, registration finds no URL) was equally dead.

GATEWAY_RELAY_* ROUTING stamps (URL, ENDPOINT, ALLOW_DIRECT_PLATFORMS,
PLATFORMS, BOT_IDS, ROUTE_KEYS, INSTANCE_ID, WAKE_URL, DISPLAY_NAME)
are now in _GLOBAL_ENV_EXACT: deployment config read from os.environ
under any scope, exactly like the API_SERVER listener settings
(#69379), so every reader resolves the same value. Relay AUTH material
(SECRET, ID, DELIVERY_KEY, IDP_*) is deliberately NOT global — it
stays profile-scoped with the fail-closed multiplex guard, mirroring
the non-secret/secret line the terminal env blocklist already draws
(tools/environments/local.py).

The round-1 multiplex regression test asserted the now-rejected
semantic (profile-scoped stamps win); it is inverted to pin the
global-stamp contract: a process-env stamp survives the profile scope
(sweep runs, matching registration), and a profile-only .env stamp
does NOT activate relay.
2026-08-21 13:21:46 +10:00
Teknium 49780391a2 fix(runtime): per-model api_mode + /v1 healing for custom OpenCode-family providers
The named-custom-provider runtime path returned a static api_mode, so a
providers: entry like opencode-go-bridge -> https://opencode.ai/zen/go/v1
sent responses-only models (grok-4.5, gpt-5.6-luna) to /chat/completions
and got HTTP 503 (#85589 repro). Now: when the provider name is in the
OpenCode family or the base_url is hosted on opencode.ai, derive api_mode
from the effective model and run the symmetric /v1 normalization — unless
the user declared an explicit transport, which stays authoritative.

5 new regression tests against a real temp HERMES_HOME config.
2026-08-20 20:21:12 -07:00
Teknium a83c3915a3 fix(opencode): family-wide provider predicate + reserved tool-name aliases for custom opencode-* providers
Builds on @Lesnak1's #85619 (issue #85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
2026-08-20 20:21:12 -07:00
Lesnak1 f9cd51eb66 fix(providers): support custom opencode-go-* provider routing and grok models 2026-08-20 20:21:12 -07:00
Teknium 7c9285aa14 fix(memory): bound invalid-target error and restore recovery hint
Route the model-supplied target through _bound_error_text so a huge
bogus target can't bloat context, and restore the "Use 'memory' or
'user'" hint. Follow-up to HexLab98's review note on the salvage.
2026-08-20 20:20:23 -07:00
kshitijk4poor 2cf7b36e11 fix(memory): enforce independent built-in store permissions
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
2026-08-20 20:20:23 -07:00
kshitijk4poor c809d964d4 fix(memory): parse boolean config values consistently
Use Hermes's shared truthy-value parser so quoted false memory flags disable both built-in stores as expected.
2026-08-20 20:20:23 -07:00
kshitijk4poor 5a5d6b966d fix(memory): unify store flags and bypass stale availability cache
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
2026-08-20 20:20:23 -07:00
Teknium 18a15a46d8 test(update): Windows progress self-test survives transient /progress socket stalls
A single urlopen(timeout=5) TimeoutError from the PS runspace listener
failed the test on a loaded runner (run 32440286339) even though the
listener recovered moments later — a transient stall is not the hang
this test guards. /progress sampling now retries until a deadline
(only a persistently unresponsive listener fails), the self-test hold
grows 10s -> 30s so retry time cannot push sampling past the held
stage, and the exit wait gets matching headroom.
2026-08-20 19:56:28 -07:00
Teknium 43c6dace56 fix(cron): restore reasoning_effort on cronjob() for the CLI lane — model dispatch still drops it
The CLI (hermes cron create/edit) routes through cronjob(); removing the
parameter outright broke that lane (CI slices 6/9). The parameter is back
on the function, but CRONJOB_SCHEMA and the registry handler still omit
it — same pattern as the intentional model/provider/base_url omission.
New test proves a hallucinated reasoning_effort arg through the model
dispatch is dropped.
2026-08-20 19:56:14 -07:00
Teknium 991af03f4c refactor(cron): keep reasoning_effort off the model-facing cronjob tool schema
Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.
2026-08-20 19:56:14 -07:00
Victor Kyriazakos 58258af0cf test(cron): scrub tracker references from reasoning-effort test docstring 2026-08-20 19:56:14 -07:00