The composer middleware is now identification-only: it resolves the
user's @tags against the live roster and annotates the draft with who
they refer to (profile, friendly title, device for cross-connection
rows). The agent decides whether to contact them and does it through
its message_agent tool — one send path, composed messages only.
Deleted the renderer's entire parallel delivery transport:
deliverRemoteRosterMentions / pollRemoteDmReply /
ensureRemoteCanonicalChat and the injected shellout instructions
('[@mention handoff — run hermes -p …]' and 'Desktop is delivering …
over Connections'). This retires the whole invocation bug class at the
source instead of sanitizing it: no verbatim user text is ever
forwarded by the renderer (#91397), and no shell command is ever
composed from prompt text (#91304, #91339 shape).
Tests: mention-identification.test.mjs replaces the two delivery-era
files — identification note shape, no-shellout/no-delivery containment
(sabotage-verified: re-adding a renderer delivery call fails 2 tests),
poisoned-title inertness, pass-through for unknown @s, and a source
contract pinning the deleted machinery. hide-bots + roster-cache-key
harnesses re-pinned to the new contract. 390/390 green.
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.
Sessions running on provider 'nous' get a 'Nous support' action on the
failed-turn card, opening the portal help hub
(https://portal.nousresearch.com/help — docs, Discord, GitHub) in the
external browser. All five locales + docs updated.
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.
Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
Bot Mode agents now DM teammates through a real tool instead of
hand-assembled shell commands. message_agent(target, message) validates
the target against the live roster, applies the sender's attribution
prefix server-side, and delivers over the existing proven transports
(hermes -p ... --query-file for local teammates, hermes peer dm for
peer gateways) as a tracked background process with notify-on-complete
— fire-and-forget, the reply wakes the sender on a later turn.
Containment: the schema is injected per-turn ONLY into a bot's
canonical 'Bot Chat' session on Bot-Mode-managed installs (same gate as
the protocol section); it is never registered in the tool registry or
any toolset, and dispatch re-gates on the session title so a forged
call from any other session refuses. The gate is session-stable, so the
tool list stays byte-identical across turns (prompt-cache safe).
The protocol section is rewritten to teach the tool and now carries the
teammate roster WITH ROLES (Bot Mode title + profile description), so
bots know who does what before picking a recipient. Roles and a
protocol version salt join the capability fingerprint: existing eternal
Bot Chats adopt the v2 protocol + tool with one epoch refresh, and a
rename/description edit refreshes the roster on the next message.
Adapts the in-place branch update from PR #89507 (@willfrombr) onto the
switch-by-default behavior: the deterministic switch path remains the
default so non-interactive updates (desktop, gateway, cron) never dead-end
on a merge conflict, and deliberate custom-branch users opt in with
updates.parked_branch_strategy: update_in_place. --switch-branch overrides
the in-place strategy for one run (deep feature branches that must not
accumulate update merge commits). Docs + config comments + tests cover
all three routes.
Co-authored-by: Willian Santos <285090322+willfrombr@users.noreply.github.com>
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.
Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
Address review feedback on #65076:
- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
config-first region resolution (bedrock.region in config.yaml, then
AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
branch) to the new helper. Previously it derived its region with bare
resolve_bedrock_region() (env-first), so when config.yaml pinned
bedrock.region to a different region than the ambient AWS env, auxiliary
calls (compression, memory, vision) left the primary runtime's region.
Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
uses the OpenAI-compatible endpoint, which the Mantle route made stale.
Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
Converse), the Mantle auth model (bearer token or SigV4), and add the
GPT-5.5/5.6 model IDs to the models table.
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.
The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.
Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.
Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.
Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.
Fixes#88655
`hermes backup` already skips `backups/` so a full zip never re-ships
earlier pre-update zips. `state-snapshots/` (written by `hermes backup
--quick`, `/snapshot create`, and the pre-update safety net) has the same
shape — every retained snapshot holds its own copy of state.db — but was
not in `_EXCLUDED_DIRS`, so a full backup shipped the DB once per
retained snapshot on top of the live one.
Two places hit this in practice:
- `hermes update` in `full` mode takes the quick snapshot *before* the
full zip, so the pre-update zip always nests the snapshot it just made
(state.db twice in every pre-update-*.zip).
- Any recurring `hermes backup --quick` (default keep=20) makes a daily
`hermes backup` grow by roughly one compressed state.db per retained
snapshot; a 750 MB state.db with two snapshots on disk pushed a daily
zip from 1.8 GB to 2.3 GB.
Add `_QUICK_SNAPSHOTS_DIR` to `_EXCLUDED_DIRS` (moving the constant up
next to the exclusion rules so there is one source of truth). Both walk
sites and `_should_exclude` share the set, so `hermes backup`, the
pre-update zip and the auto-backup path all pick it up. Restoring
snapshots after a machine move was never the point of the full backup —
`profiles.py` already excludes `state-snapshots/` from `--clone-all` for
the same reason.
Tests: unit case next to the `backups/` one, plus two end-to-end cases
that use the real `create_quick_snapshot` producer and assert the zip
carries exactly one state.db (full backup and pre-update-order).
Follow-up to the salvaged GLM-5.3 support commit: drop z-ai/glm-5.1 from
both curated lists per Teknium's direction (glm-5.2 keeps the 'default'
tag), and regenerate the docs manifest. glm-5.1 remains available via
live discovery and on out-of-scope surfaces (zai plugin, setup defaults,
opencode-go) — named leftovers, not silently swept.
Home Manager separates an installation from a daemon. This module put
both under `services.hermes-agent`, and `installPackage` added a program
to the PATH from a service module.
`programs.hermes-agent` now installs the command line application and
the desktop application. `services.hermes-agent` keeps the state, the
configuration and the daemons, and stays the authority: the new module
reads `hermesHome` and the backend address from it. A person can enable
one without the other, which is a machine with an application and no
gateway, or a headless gateway with no display.
The desktop application needs this split to work correctly. A launcher
that starts from the desktop menu reads no shell profile, thus the
HERMES_HOME that `home.sessionVariables` exports reaches an interactive
shell only. Home Manager writes `systemd.user.sessionVariables` to
environment.d, and this module puts no HERMES_HOME there, because that
file applies to each user unit. The application then opens ~/.hermes
while the services use `hermesHome`, and the person sees no sessions and
no keys. Thus the launcher carries the value itself, through a new
`extraEnv` argument on the desktop package.
The application also gets the Nix agent package, with
HERMES_DESKTOP_HERMES. The usual distribution of the Electron
application carries its own Hermes runtime and downloads more at the
first start. `hermesDesktop` is a passthru of the agent and pins
`finalAttrs.finalPackage`, so an override of `extraPythonPackages` or
`extraDependencyGroups` reaches both. One machine thus has one runtime.
`backend.sessionTokenFile` connects the application to the backend of
the service. Without it the module runs `hermes serve` and the
application starts a backend of its own, which gives two backends on one
HERMES_HOME. The backend reads the file into
HERMES_DASHBOARD_SESSION_TOKEN. The launcher reads the same file into
HERMES_DESKTOP_REMOTE_TOKEN, beside a HERMES_DESKTOP_REMOTE_URL that
names the address of the service.
Measurements against a live `hermes serve` on loopback show why that
shape is the correct one:
- `_resolve_session_token()` reads HERMES_DASHBOARD_SESSION_TOKEN, and
`_has_valid_session_token` accepts that value as a Bearer credential.
A request without it gets 401, and a request with the wrong value
gets 401.
- The /api/ws socket accepts a query parameter only. A header gets 403,
and `?token=` connects. Hermes Desktop builds exactly that URL, in
`apps/desktop/electron/connection-config.ts`. Thus a test of the HTTP
leg alone is a false positive.
- `resolveDesktopRemoteRoute` throws when the URL is set and the token
is not. Thus the two variables travel together or not at all.
The token enters no Nix store path. `makeWrapper --set` and a systemd
`Environment=` value both write a literal into the store, which all
users can read. Thus each side reads the file at start time. The
launcher does it through a new `extraRun` argument on the desktop
package, and the backend through the launcher script that
`backend.waitFor` already uses. launchd has no EnvironmentFile, so a
script is the one shape that works on Linux and on Darwin.
`backendArgv` gives the plain argv only when nothing must run before
the backend.
`services.hermes-agent.installPackage` is removed. It defaulted to true,
so a person who never named it still got the command line. A silent
removal thus gives them a machine with no `hermes` and no message. The
module refuses a configuration that sets it, and the text names the
exact replacement for the value they gave.
Checks:
- the launcher carries HERMES_HOME
- the launcher reports HERMES_MANAGED only when the services own the
configuration, because no activation writes a marker without them
- the launcher pins the agent package that `programs.enable` installs
- the launcher names the backend of the service, and gives a token
beside the URL
- the backend reads the session token
- each side reads the file at start time, and the token is no `--set`
value
- `programs.enable` alone starts no service
- `installPackage` is refused, with a message that names the
replacement, and its absence evaluates
Each check reads the wrapper of the real package, and not an option
value. Each one was tested with a mutation that breaks the behavior it
asserts.
Third surface for the Ox Alpha stealth reasoning model (after the
OpenCode Zen rollout in #91250 and the OpenRouter listing in #91284).
Adds stealth/ox-alpha to the curated Nous list and regenerates the docs
manifest. Free on the portal ($0/$0), 1M context, 131K max output —
verified against the live inference-api.nousresearch.com/v1/models.
Provider-agnostic metadata already resolves via the bare ox-alpha slug
(DEFAULT_CONTEXT_LENGTHS 1,048,576; reasoning_timeouts 300s floor), and
the nous route bills via official_models_api, so no pricing snapshot is
needed.
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.
- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
snapshot set, per-file 1GiB cap, and keep policy as the invoking
profile (no partial tier, no new restore-coherence class), each into
the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
runs the #34600 cron-loss safety net per profile against its OWN
snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
profile's (best-effort, receipt-recorded); post-update cron restore
extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
Generic Hermes callers still use the existing browser backend when extension control is disabled or no server-bound controller identity exists.
Once the gateway binds a controller identity, missing scope, disconnect, or capability loss now fail closed instead of silently switching a control-this-tab request to another local or cloud browser. Covers the schema-build to dispatch disconnect race.
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.
Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.
Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:
Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
(absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free
Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
+ laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
keyless Ox Alpha; keyed — Go relay 401s anonymous requests)
Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
glm-5.2:free 256K (the free variant is capped below the 1M paid entry)
Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
MEMBERSHIP in the verified opencode-free catalog, not the -free
suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
suffix-based healing would have routed it to a Zen relay that
doesn't serve it (verified: zen 401s 'not supported', go 401s
'Missing API key'). New regression test pins this.
Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.
- hermes_cli/update_inventory.py (new): side-effect-free runtime
inventory — install kind via detect_install_method (git / docker / nix
/ apt, updatable-in-place or not, with the correct external update
command for image/package-managed installs), all profiles, every live
gateway with its supervisor (systemd / launchd / manual via the
fleet-wide _get_service_pids), running code_sha/code_version from the
#91283 gateway_state.json stamps, and the restart mechanism each
runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
docker/nix refusal gates so image-managed installs get a useful
'not updatable in place + right command' report instead of a bare
refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
('plan' key) and prints a one-line fleet summary, so post-mortems can
compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
never-raises, JSON round-trip for the receipt, print output shapes,
receipt integration.
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.
- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
persistent memory
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
Adds OpenRouter's free "Ox Alpha" stealth reasoning model
(stealth/ox-alpha) to the OpenRouter fallback snapshot, plus the
provider-agnostic metadata it needs:
- OPENROUTER_MODELS: free-tier entry (1M ctx)
- DEFAULT_CONTEXT_LENGTHS: ox-alpha -> 1,048,576 (verified against
OpenRouter live /api/v1/models; without this the slug fell through
to no match)
- reasoning_timeouts.py: 300s stale floor for ox-alpha and the
OpenCode Zen twin slug x-preview-f-free (reasoning model,
long-horizon agentic work per its model card)
- model-catalog.json regenerated
Pricing snapshot skipped: openrouter bills via official_models_api
(live pricing; model is free anyway).
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
`useTheme`, the accent override, the retint helper and the OKLCH math reached
the SDK without reaching this page, which still documented `THEMES_AREA` alone.
Registering a theme only lists it in the picker, so the natural reading was that
plugins cannot switch themes at all — and the one person who tried concluded
exactly that and patched the app instead.
Documents the selection half: the hook for components, `requestTheme` for
callbacks with no component around them, and a Theming row in the export table.
The agent-facing reference gets the same note, since it never covered themes.
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').
New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.
Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
Addresses teknium1 sweeper review on PR #67696:
1. Gateway bridge: Skip str(None) bridging when YAML value is Python None
(from or bare ). Previously str(None) → None → unlimited
instead of default 90. Now clears stale env var so resolver applies default.
2. TUI: Route _cfg_max_turns through resolve_turn_limit instead of bare
int(). Old code crashed on none/unlimited and swallowed 0 via
. HERMES_TUI_MAX_TURNS env var also routed through
resolver.
3. Docs: Document unlimited spellings (none/unlimited/infinite/0/-1) in
configuration.md.
4. Tests: Add TestGatewayBridgeNullHandling (4 tests) and TestTUIResolver
(8 tests) covering null handling, string spellings, env var override,
and legacy root-level config.
All 50 tests pass.
#90732 removed right-click → Sessions (one forever-chat per bot is the
product contract); #90756 cleaned the last in-app copy. This removes the
remaining docs bullet describing the dead affordance.
Every drive_preview action answered with the entire inventory — around 120
elements of ref, role, label, and an up-to-eight-rung `:nth-child` selector
chain. On a real app shell that was ~24.5k characters, re-sent after every
click, so a ten-step task paid for ten copies of a page that had barely moved.
Handles are now durable and legible. An element is named after what it is and
what it says — `btn-sign-in`, `inp-email`, `srch-search-projects` — minted once
per page and never reused, with duplicates disambiguated as `btn-edit`,
`btn-edit-1`. Each one remembers a stable attribute, its role, its accessible
name, and the nearest landmark it sits in, so when a framework destroys the
node and builds a new one the handle moves across and the agent is told
`rebound` rather than being handed a removal it has to react to and an addition
it has to re-read. The re-bind ladder is anchortree's (Apache-2.0), minus its
geometry rung, which can never clear the threshold on its own.
Because the handles hold, the first look at a page returns the inventory and
every look after it returns only what moved. `changed` carries the ref and
whichever of label/value/disabled actually shifted — role and selector are
absent by construction, since a change in either would mean the re-bind ladder
was looking at a different element. A delta gives way to a full re-read when
half the page is new, where there is nothing left to reuse.
The selector column is gone with it. It was 74% of the inventory on an
85-element page, nothing downstream ever read it, and a positional chain is
wrong the moment a sibling appears. An `#id` or `[data-testid]` survives when
the page offers one; everything else is addressed by handle.
Legibility is what makes the delta work rather than a nicety. `+ btn-sign-in`
on turn nine reads on its own, where `+ @e42` sends the model back to an
inventory twenty thousand tokens ago.
Measured on an 85-element app shell: 18,693 -> 4,930 characters for a baseline,
and a steady turn that moved two things costs ~200.
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
'hermes --version' (and -V) now prints the full version report — banner
version line with upstream SHA, install directory, authoritative install
method, Python and OpenAI SDK versions, and update status — making the
separate 'hermes version' subcommand redundant. The subcommand is removed.
- _startup_fast.print_fast_version_info() is now THE canonical version
printer: static lines print instantly from stdlib probes, then the
banner label, install-method resolver, and update check lazy-import
after the first line is on screen (each degrades gracefully).
- main.py _print_version_info() delegates to it (used by /version in the
CLI chat surface and the --version flag path); the old duplicate
implementation is deleted.
- hermes_cli/subcommands/version.py removed; parser wiring, subcommand
sets, console-engine extraction entry, and tests updated. Hermes
Console keeps a 'version' command wired to the shared printer.
- Termux fast paths now include update status too (previously
check_updates=False).
- Docs/i18n, CONTRIBUTING, SECURITY, and nix checks updated to
'hermes --version'.
When the chosen/keyed backend fails a web_search or web_extract call
(bad key, upstream outage, 5xx, raised exception), that single call
retries on the keyless free-tier ring instead of erroring. The next
call attempts the chosen backend again — no sticky failover, no state.
Resolves the keyed half of #78984/#32159 (keyless half landed in the
ring PR).
- tools/web_tools.py: _rescue_eligible (keyed ring vendors + non-ring
backends eligible; keyless-mode calls excluded — they already walked
the ring), _rescue_search/_rescue_extract (search annotates
rescued_from + backend_error naming the original failure and the
retry-next-call semantics; extract rescues only whole-batch failures,
partial failures pass through untouched; rescue failure preserves the
ORIGINAL backend error with the rescue note appended)
- both dispatchers wrap the provider call: failure-results AND raised
exceptions rescue; ineligible paths re-raise unchanged
- web.keyless_rescue config key (default true; implicitly off when
keyless_fallback is off); docs updated
Live E2E: keyed Tavily with an invalid key 401'd and the call was
served by the real ring with the rescue annotation; a second call
re-attempted Tavily first (statelessness proven); whole-batch extract
rescue returned real page content. 13 new tests; 67 green across the
keyless suites.
Remote-mode installs had every update affordance (About panel Update now,
⌘K Update Hermes, the update-ready toast) pointed at the BACKEND only, so
users updated their VPS forever while the desktop app itself sat weeks
stale — with no signal it was behind (the skew warning only fired the
other way). Reported by Santiago Sarceda: mac app on v0.20.0 kept
repro'ing UI bugs fixed on main because 'update' never touched the app.
- store/updates.ts: applyEverythingUpdate() orchestrates all targets —
active backend first (detailed progress), every other eligible
registered gateway via the existing Electron fan-out (cloud rows skip),
the client LAST (its apply relaunches the app). startActiveUpdate/
requestActiveUpdate route through it whenever more than one update
target exists; single-machine installs keep the one-button flow.
- After ANY successful backend update, the client version is re-checked
and a one-click 'Update desktop app' warning fires if the GUI is still
behind — the reverse-skew signal that didn't exist.
- electron: hermes:connections:update-all accepts optional excludeIds so
the flow doesn't double-dispatch the active backend / local runtime.
- i18n: 7 new updates.* keys across en/zh/zh-hant/ja/ar.
- docs: desktop.md Updating section + multi-connection guide.
- tests: 10 new cases (gating, ordering, exclusions, failure isolation,
memoization, nudge on/off).