Upgrades yesterday's #99310 skip-guard to full pointer-carry from
PR #90484: model assignment and custom-endpoint activation now write
key_env or the raw ${VAR} template into model config instead of
dropping the credential reference entirely, so the model entry keeps
resolving at runtime with zero plaintext in config.yaml. Applied
surgically onto current main (the PR branch predates newer
web_server.py changes); key_env carry made independent of the
expanded api_key guard, tests updated to pin pointer-carry.
POST /api/model/set copied the load_config()-resolved plaintext of a
${VAR}/key_env provider entry into model.api_key, writing the secret
into config.yaml and recreating it on every re-apply. The mirror now
checks the RAW on-disk entry and skips env-referencing entries;
literal keys keep the existing behavior.
_env_line_defines_key() decides which .env lines the writers may rewrite or
drop. It matched on the `KEY=` prefix, but load_env() splits on the first
`=` and strips the name:
key, _, value = line.partition('=')
env_vars[key.strip()] = _parse_env_value(value)
so `OPENAI_API_KEY = sk-...` is a live assignment — the key resolves, the
provider works, and every UI shows it as set. The writers did not see it.
This is the same resurrection hole #40041 fixed for `export KEY=`, still
open for the whitespace form:
- DELETE /api/env 404s ("not found in .env") while the credential stays
active — a key the user revoked through the UI is never actually revoked
- PUT /api/env appends a SECOND line instead of replacing; a later delete
removes the appended line and the original value silently comes back
Rotate-then-delete on a spaced line therefore restores exactly the key the
user rotated away from.
Match load_env()'s parse instead of prefix-matching, so the writers accept
precisely what the reader accepts: skip blank/comment/no-'=' lines, strip an
`export ` prefix, then compare the stripped name. Commented-out lines stay
untouched and `KEY_EXTRA=`/`MY_KEY=` still do not match `KEY`.
Verified against the real dashboard endpoints on a temp HERMES_HOME: the
spaced line is now removed, rotation replaces it in place with no duplicate,
and a parity check asserts the writer matches a line iff load_env() does.
`_scrub_config_yaml_mirrors` reconciles the config.yaml copies of a credential
when it is rotated or removed through the dashboard. It walks `model`,
`auxiliary.<task>`, and `custom_providers.<name>` — but not the keyed
`providers` schema.
`providers` is not a niche section: `get_compatible_custom_providers` documents
it as "the newer keyed schema" (v12+), and it is exactly where the dashboard /
desktop write a custom endpoint's inline key —
`_write_custom_endpoint` sets `providers.<id>.api_key`. That value is a real
credential: the runtime resolver reads it (`runtime_provider` /
`hermes_cli.main` / `model_switch` all read `entry.get("api_key")` off a
`providers` entry), and an inline key outranks the env var.
So the section the scrub skips is the one the dashboard writes to, and both
callers break on it:
- save_provider_env_credential (rotation, #62269): a stale
`providers.<id>.api_key` is left at the OLD value and, being
higher-precedence than the freshly-rotated env var, shadows the rotation —
the "persistent 401 with a key the UI no longer shows" that #62269 fixed,
reintroduced for the newer schema.
- remove_provider_env_credential: its contract is to "remove a credential from
EVERY store it lives in", yet the `providers` copy survives, leaving the
secret in config.yaml after the user asked to delete it.
Walk `providers.<id>` too. The scrub stays value-matched, so an unrelated
endpoint's key is untouched. Only `api_key` is scrubbed here: in the keyed
`providers` schema `api` is the base_url alias, not a credential (unlike
model/auxiliary/custom_providers), so `_fix` takes an explicit field list and
this section passes `("api_key",)` — a provider's endpoint URL is never
rewritten even if it happened to equal the credential string.
tests/hermes_cli/test_credential_lifecycle.py: drive the real PUT/DELETE
/api/env endpoints against a `providers.<id>.api_key` mirror — rotation moves
it to the new key, delete clears it, and a `providers.<id>.api` base_url alias
is preserved. The two scrub tests fail on main (stale key survives); the
base_url guard passes on main as a control. 15 pass here; 345 pass across the
credential-lifecycle + web-server suites (the one failing honcho-merge test
fails identically on clean main).
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
The PKCE payload is a flat 'provider=...;state=...;verifier=...;next=...'
string. A raw ';' is a cookie-attribute terminator, so Python's
http.cookies emits the value in RFC 6265 quoted form with each ';'
escaped as the backslash-octal '\073'. Mainstream browsers echo that
form back verbatim and Python parsers decode it — the browser round
trip is fine. But '"' and '\' are outside the plain cookie-octet set,
and non-Python hops that re-serialize the Cookie header reject the
value and drop the cookie entirely: Go's net/http (Traefik middleware,
Authentik outposts, other gateways) refuses any cookie value
containing a backslash. The OIDC callback then 400s with "Missing
PKCE state cookie" even though the browser sent the cookie.
Field reproduction: support thread "Still unable to use Authentik for
signin with traefik" — devtools showed the browser sending the intact
quoted \073 cookie on /auth/callback while Hermes logged
missing_pkce_cookie behind a Traefik+Authentik chain.
Fix: URL-encode the whole payload in set_pkce_cookie (quote(payload,
safe='') — ';' becomes '%3B') so the wire value contains only
cookie-octets and no parser in the chain has anything to reject, and
decode through a single shared inverse, cookies.parse_pkce_payload(),
in BOTH readers: the OAuth /auth/callback and the native
password-login path (routes.login_submit), whose broker/provider
binding check would otherwise parse zero segments from the
newly-encoded value and silently disable itself.
Regression coverage: the wire-shape test pins the full cookie-octet
set (the '"'/'\' assertions are the ones a Go-parser hop fails
pre-fix), the round-trip tests drive the real /auth/login →
/auth/callback path, and the next= test pins the exact post-login
redirect byte shape. Native-flow broker assertions updated to decode
through parse_pkce_payload instead of substring-matching the raw wire
value.
Salvaged from #84065 (rebased onto current main, which gained the
SameSite=None PKCE attrs and the RFC 8252 native password flow since
the PR branched): kept main's _pkce_attrs cookie shape, extended the
fix to the login_submit reader the original PR predated, and reframed
the rationale — browsers do NOT truncate at the first ';' (there is
no literal ';' on the wire in the quoted form); the failing hop is a
strict middlebox cookie parser.
Closes#83832
Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
The snapshot dirs were 0700 but every file inside landed umask-wide:
shutil.copy2 preserves Chrome's own 0644 profile-file modes and
sqlite3.connect creates the online-backup destinations as plain umask
files — so the copied Cookies / Login Data / Web Data (the user's live
session credentials) sat 0644. The 0700 parents contain it by default,
but the documented HERMES_HOME_MODE traversal hatch makes group/world-
readable children a real exposure.
snapshot_real_profile now reconciles every file (0600) and nested dir
(0700) inside the snapshot through the house helpers (_secure_file /
_secure_dir — managed-mode and container carve-outs included) at the
end of every pass, so snapshots written by older builds heal on their
next launch. Best-effort, never blocks a launch.
Tests: owner-only walk under umask 022 (fails on the pre-fix code —
sabotage-verified) + heal-on-refresh for a pre-existing 0644 Cookies.
The issue's other two findings are already fixed on main: mock-keychain
flags eliminated by the direct native-binary launch (#98249, salvage of
#96763); the 'Device not configured' TTY failure is superseded by the
same launch-path rework.
hermes skills install impeccable (and the docs-page install button) now
installs the impeccable frontend-design skill as an official optional-skills
entry. The local optional-skills/creative/impeccable/ dir is a catalog STUB:
its frontmatter declares metadata.hermes.upstream (repo + path), and
OptionalSkillSource.fetch() pulls the real 163-file bundle live from
pbakaus/impeccable:.hermes/skills/impeccable — the Hermes-native bundle
upstream maintains and verifies. Nothing vendored, never stale.
New mechanism (generic, not impeccable-specific):
- OptionalSkillSource._upstream_pointer(): parses/validates the upstream
pointer (owner/name repo, clean relative path, traversal rejected).
- _fetch_from_upstream(): delegates to GitHubSource.fetch(), relabels the
bundle official/<rel> at trust 'trusted' (curated endorsement, but
third-party content — dangerous scan verdicts still block).
- The live-repo fallback path redirects stubs the same way, so stale local
checkouts behave identically.
Three real gaps this surfaced, all fixed:
- GitHubSource.fetch() only downloaded SKILL.md plus paths linked from a
canonical support dir (references/, scripts/, ...). Impeccable keeps its
playbooks under reference/ (singular) and links scripts only from
reference files, so fetch shipped 1 of 163 files. fetch() now downloads
the full skill directory via the git tree (same approach as the
optional-skills live fetch), still rejecting symlinks/hidden/unsafe paths
and still failing on a missing SKILL.md-linked references/ path.
- The five env_exfil_* scanner patterns flagged loopback requests as
critical exfiltration: impeccable's live mode polls
http://localhost:PORT/status?token=TOKEN and scored two CRITICALs.
Scheme-anchored loopback exemption added; evil.com/?u=localhost decoys
still fire (10-case regex matrix in tests).
- unified_search() truncated to limit before ranking, so official catalog
entries got crowded out by skills.sh mirrors and bare-name installs
stalled on an ambiguity table. Results now stable-sort by trust rank
before the cut, and _resolve_short_name prefers a sole official exact
match over community mirrors.
Also fixes pre-existing test pollution: TestInstallPathSafety's fixture
monkeypatched the PEP 562 dynamic SKILLS_DIR, permanently shadowing dynamic
resolution and breaking the served_repo E2E tests in any combined run
(reproducible on main).
Validation: live E2E do_install("impeccable") against real GitHub —
resolves to official/creative/impeccable, verdict SAFE, 163 files on disk,
skill loads, /impeccable slash command registers. 128/128 targeted tests;
full-dir fetch test sabotage-verified. Docs: optional-skills catalog row,
generated skill page, sidebar.
The bundled plan skill's auto-generated slash command fell off the capped
Telegram/Discord command menus for most installs (skills are the only tier
trimmed at the platform caps, alphabetically — 'plan' sat past the cutoff at
index 57 of 82 bundled skills). Converting it to a first-class CommandDef
gives it a guaranteed core-tier menu slot on every platform.
- agent/plan_prompt.py: build_plan_prompt() — plan-mode rules + authoring
craft distilled from the retired skill; prompt-injection pattern like
/learn and /init (no engine, no model-tool footprint, cache-safe).
- CLI: _handle_plan_command mixin handler (pending-input injection).
- Gateway: /plan branch rewrites event.text and falls through (role
alternation preserved).
- TUI: command.dispatch branch ('plan' was already in
_PENDING_INPUT_COMMANDS).
- Removed skills/software-development/plan/ + docs pages (EN + zh-Hans),
catalog rows, sidebar entry, related_skills references.
- PROTECTED_BUILTIN_SKILLS is now empty (mechanism kept); dependent
curator/usage tests moved to monkeypatched sentinels.
Salvages #67292 by @webtecnica (credit: first /plan command submission,
issue #67264); reworked from inline planning prompt to the prompt-injection
pattern with workspace-saved plans. Closes#67264, closes#36821 (empty
/plan infers task from conversation context).
Completes the #90953 salvage on post-#98237 main:
- New _merge_request_overrides helper defines the precedence contract:
explicit delegation.request_overrides merges OVER runtime/parent-derived
overrides — explicit top-level keys win; extra_body is deep-merged one
level so runtime extra_body keys survive unless redefined. Inputs are
copy.deepcopy'd so transport-side mutation can't leak into config or the
provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
provider-alongside-base_url runtime overrides instead of being a separate
return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
(previously only when override_provider was set), enabling the inherit
branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
proofs, explicit-over-runtime precedence on the provider-alongside-base_url
path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
Follow-up to the previous commit (which fixed the default/fallback
provider path). A mid-session `/model` switch stores a per-session
override bundle in `_session_model_overrides` that omitted
`request_overrides`, and the two consumers
(`_resolve_session_agent_runtime` fast path and
`_apply_session_model_override`) only copied
provider/api_key/base_url/api_mode. So switching *to* a custom provider
via `/model` did not apply its `extra_body`.
- `ModelSwitchResult` gains a `request_overrides` field, derived for the
switched provider via `_get_named_custom_provider` /
`_custom_provider_request_overrides` (the same overrides
`resolve_runtime_provider` surfaces for the default path).
- Both `/model` override-storage sites in slash_commands.py persist it.
- Both consumers apply it; `_apply_session_model_override` also clears a
stale value when switching to a provider that has none.
Extends tests/gateway/test_turn_request_overrides.py (3 new cases).
Carry provider-derived request_overrides through runtime resolution,
fallback projection, session /model state, restart rehydration, and
turn-route merge so named custom providers keep extra_body and related
overrides.
- _terminate_real_profile_chrome(): directly-launched real browsers are ours
to reap (agent-browser only attaches); wired into the atexit emergency
cleanup and both launch-failure paths so orphaned Chrome processes can't
accumulate.
- Display-less Linux gate: append --headless=new (shares the profile's normal
cookie store, unlike legacy headless) so the direct-launch path doesn't
regress servers without DISPLAY/WAYLAND_DISPLAY.
- Register browser.real_profile_pin in config_defaults.py and document the
new launch model + pin in website/docs/user-guide/features/browser.md.
- Drop unused tempfile import from the cherry-picked commit.
The Default dir in the snapshot holds the SOURCE profile cookies, but
info_cache['Default'] kept the source user-data-dir own Default entry
(a different person). Chrome saw cookies that belong to profile B while
its profile metadata said profile A, demanded a 'Continue as <name>'
profile-sign-in reconciliation on every launch, and treated the profile
as mid-sign-in. Use the source profile info_cache entry (name + Google
account) for the copy Default.
Four fixes for real-profile browsing (browser.use_real_profile), found and
verified end-to-end on macOS with a live Chrome:
1. _copy_auth_file: sqlite3.connect('file:...?mode=ro') on a live Chrome
auth DB can block indefinitely inside lock negotiation - the busy
timeout never fires, so the 'fail fast' path hangs the launch forever.
Try immutable=1 first (reads instantly, correct for a committed
snapshot of a file another process owns); mode=ro stays as fallback.
2. Launch shape: agent-browser's own launch injects --use-mock-keychain /
--password-store=basic / --headless=new. On macOS the mock keychain
makes Chrome treat every keychain-encrypted cookie as undecryptable
and drop it - the copied profile launches signed out (~3 anonymous
cookies instead of the full jar). Launch the user's real browser
binary directly on the copy (no mock-keychain switches), wait for
DevToolsActivePort, then attach agent-browser via --cdp.
3. Snapshot copy: Local State was copied verbatim, still naming the
SOURCE profile (last_used='Profile 2', info_cache listing several)
while the copy only contains Default. Chrome opens the missing profile
dir and starts signed out. Normalize the copy's Local State to
Default-only.
4. CDP resolution: the agent-browser daemon may report the endpoint of a
browser IT spawned (throwaway temp profile) instead of the real
browser we launched on the copy. Trust the port our browser wrote to
DevToolsActivePort.
Also adds browser.real_profile_pin (optional): pin which source Chromium
profile dir is snapshotted instead of following profile.last_used - on a
machine with a work profile and a personal one, last-used roulette can
silently give the agent the wrong identity. A pin naming a missing dir
fails closed (signed out) rather than falling back to last_used.
Tests: 4 new pin tests + 3 launch tests reshaped to the direct-launch
contract (Popen the real binary, agent-browser attaches). 77 passing.
- Add rolling status bar metrics:
- cache hit ratio (◈) delta since model/compression reset
(hit = cache_read / prompt, verified against live logs)
- avg latency (◷) and throughput (↑ t/s) over last 10 API calls
(deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
idle, focus, yolo, stash, battery, title
Missing/null -> all enabled (backward compat). Unknown keys ignored.
Title gated via right-align; stash/battery also gated.
- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
overflow trim preserved. Battery also respects display.battery.
No private data; mock data in tests.
Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.
Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.
When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.
total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.
Closes#41909
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).
Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.
A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.
Fixes#87887, #87888, #87889.
Follow-up to the #77848 salvage: mirror the curated model lists onto
the alibaba-cn / alibaba-coding-plan-cn / alibaba-token-plan-cn
profiles registered in #73345, and add all Alibaba variants to the
qwen provider group so they appear in the drill-down picker.
Token Plan (Personal Edition) model catalog for hermes model /
provider pickers, verified against a live Token Plan subscription
(2026-08-03). Provider profiles landed separately in #73345; this
carries the picker-list half of #77848.
Co-salvaged-from: PR #77848
Extends the real-profile machinery (PR #95620) to Brave Origin — Brave's
standalone paid build with a fully separate install identity:
- new canonical key 'brave-origin' in _CHROMIUM_BROWSERS
- Windows: BraveOHTML ProgId -> brave-origin; channel ProgIds BraveOBHTML/
BraveODHTML/BraveOSHTM fail closed (identifiers from brave-core
install_static)
- macOS: com.brave.Browser.origin bundle id (exact match); .beta/.dev/
.nightly channel bundles fail closed; /Applications/Brave Origin.app
- Linux: brave-origin.desktop matched BEFORE the bare 'brave' fragment
(substring scan would otherwise resolve an Origin default to stable
Brave and drive the wrong profile — #95549 wrong-principal invariant);
brave-origin-{beta,nightly,dev} fail closed
- profile dirs: BraveSoftware/Brave-Origin on all three OSes (per
brave-core kProductPathName + Homebrew cask zap paths)
- /browser connect launch tables: Brave Origin split into its OWN group
so a 'brave' executable lookup can never resolve to the Origin binary
- user-facing strings/docs/desktop tooltip updated
Tests: progid/bundle/desktop map params + data-dir resolution for all
three OSes; 125 passed in the three browser test files.
The synced identity text carries em-dashes, but install.ps1 must stay
pure ASCII (Windows PowerShell 5.1 reads BOM-less .ps1 in the ANSI code
page; a non-ASCII byte in a string literal desyncs the parser — see
tests/test_install_ps1_ascii_only.py, issues #66994/#67000). Seed the
ASCII-dashed variant there instead, and register that variant in
_LEGACY_TEMPLATE_SOULS so Windows installs converge onto the canonical
em-dash text on first run.
DEFAULT_AGENT_IDENTITY was rewritten in agent/prompt_builder.py (behavior
spec, exploration-thrift line deliberately removed) but the actual seed
written to disk on first run, hermes_cli/default_soul.py's
DEFAULT_SOUL_MD, was never updated. ensure_hermes_home() writes
DEFAULT_SOUL_MD into SOUL.md on every fresh install before the agent's
first turn, so virtually all real users end up as "SOUL.md users" seeded
with the pre-rewrite text -- including the exact "targeted and efficient
exploration" line the rewrite explicitly banned -- while the new
DEFAULT_AGENT_IDENTITY fallback essentially never serves the "fresh
install" audience its own PR body named as the target.
- DEFAULT_SOUL_MD now matches DEFAULT_AGENT_IDENTITY exactly.
- The pre-rewrite text is added to _LEGACY_TEMPLATE_SOULS so installs
already seeded with it self-heal via the existing upgrade-in-place
mechanism (same guarantee as the comment-only scaffold entries: the
string carries zero user intent, so it's safe to replace).
- Synced the other places install.sh's own comment says "MUST match
DEFAULT_SOUL_MD": scripts/install.sh, scripts/install.ps1,
docker/SOUL.md, and the docs/i18n pages that quote the fallback text
verbatim.
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
Flip the salvaged --start-now behavior (PR #97958) into the unconditional
default: /loop's first iteration is due the moment the loop is set, then
recurs on the normal cadence. The flag is dropped — it was never released,
so there is nothing to deprecate.
- LoopManager.set(): next_due_at = now for both cadence modes
- drop --start-now parsing, the persisted LoopState.start_now field, and
the flag from help text; confirmation now always says the first wakeup
fires now
- tests updated to pin the new default (incl. the TUI not-due test, which
now has to push next_due_at out explicitly)
- docs: quick-start and command table describe the immediate first run
/loop [interval] <prompt> currently schedules the first wakeup one full
interval after the command runs (next_due_at = now + interval). When the
user just told Hermes what to check, waiting the whole interval before
any output feels like the command was ignored.
Add an opt-in --start-now flag that keeps Claude Code parity as the
default but lets the user run the first iteration immediately, then
continue on the cadence:
/loop 1h check the deploy status # first run in 1h (unchanged)
/loop 1h --start-now check the deploy # first run now, then hourly
- parse_loop_args(): parse and strip --start-now (leading or trailing)
- LoopState: new persisted start_now field (default False, survives
serialization round-trip and old rows missing the field)
- LoopManager.set(): next_due_at = now when start_now, for both fixed
interval and self-paced modes
- dispatch_loop_command(): wire start_now through, update help text, and
report "First wakeup fires now" in the confirmation
- website/docs: document the flag in the /loop guide
- tests: parse (trailing/leading/absent/self-paced/combo/prompt-word),
tick lifecycle (due immediately vs after interval), serde round-trip,
and dispatch-level confirmation
The inline set-comprehension re-implemented modelName extraction that
_extract_model_name() already provides (and that both
union_with_portal_free/paid_recommendations already use). Beyond the
duplication, the inline str(entry.get("modelName", "")) stringified
non-string values — a malformed Portal entry with modelName 5 would have
produced a garbage "5" match that discard("") does not filter. The helper
isinstance-checks and returns None for those, so routing through it makes
the validation tier semantically identical to the union helpers.
Adds test_non_string_model_name_entries_ignored locking the behavior
(mutation-checked: fails on the raw-stringify form, passes on the helper).
Fixes#71312 (duplicate #71313).
When selecting a model via the Telegram /model picker (or any other
messaging-platform slash command, since they all share
validate_requested_model() through gateway/slash_commands.py ->
model_switch.switch_model()), a model available via Nous Portal's live
/api/nous/recommended-models endpoint but not yet in the hardcoded
curated catalog (_PROVIDER_MODELS["nous"]) was rejected with "was not
found in this provider's model listing" -- even though the exact same
model works fine via `hermes chat -m <model> --provider nous`.
Root cause: `hermes chat` merges Portal recommendations into its model
list via union_with_portal_free_recommendations() /
union_with_portal_paid_recommendations() at model-list build time
(hermes_cli/auth.py, web_server.py, model_setup_flows.py,
model_switch.py), so the model already appears "known" by the time
validation runs for that path. validate_requested_model() itself,
which every per-message /model command goes through, only checked the
live /v1/models listing and the curated catalog (_model_in_provider_catalog) --
never the Portal recommendations feed -- so a model that exists only
in Portal Recommendations was rejected on that path specifically.
Fix: add a Nous-specific fallback tier in validate_requested_model(),
checked after the curated-catalog fallback and before the final
rejection, reading the same fetch_nous_recommended_models() feed
(free + paid tiers) the CLI union helpers already use. Scoped to
provider == "nous" only; short-circuits before the network call when
an earlier tier already accepted the model; fails closed (rejects,
doesn't crash) if the Portal feed is unreachable.
Reported two issues filed 3 minutes apart with identical content by
the same author (#71312, #71313) -- commented on #71313 marking it a
duplicate of #71312 and pointing to this fix (could not close it
directly, no admin rights on the repo from this token).
6/6 new tests pass in TestValidateRequestedModelNousPortalRecommendations;
95/95 in the full tests/hermes_cli/test_model_validation.py file;
87/87 in tests/hermes_cli/test_models.py (unaffected, confirmed).